Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,270 characters · 15 sections · 47 citation commands
Extracting Mechanisms from Heterogeneous Effects: An Identification Strategy for Mediation Analysis
\singlespacing
\thispagestyle{empty} \doublespacing
\doparttoc \faketableofcontents
\setcounter{page}{1}
“How and why does the treatment affect the outcome?” This causal mechanism question lies at the core of social science research, guiding both theoretical development and the generalization of empirical findings. Causal mediation analysis answers this question by quantifying the portion of a treatment’s effect that operates through intermediate variables between the treatment and the outcome vanderweele2015explanation. The dominant approach in social science relies on the assumption of sequential ignorability imai2011unpacking. Although numerous methods have been developed for causal inference more broadly, the methodological toolkit for causal mediation analysis remains limited.\footnote{Only around $10\%$ of published papers have the opportunity to employ formal causal mediation analysis, according to blackwell2024assumption.} This study proposes an alternative identification strategy that explicitly quantifies causal mediation effects by leveraging heterogeneous treatment effects (HTEs).
Mediation analysis has been approached through two main methods. In explicit mediation analysis, researchers aim to identify the exact causal mediation effect—that is, the causal effects mediated by a specific proposed mediator. Using either the counterfactual or structural framework, as illustrated in Figure (ref), the total treatment effect is decomposed into direct and indirect effects. Because of the existence of a mediator in the indirect effect, mediation analysis requires stronger assumptions than what is required for the identification of total treatment. The identification of treatment effects requires addressing confounding between the treatment and the outcome ($U_1$ and $U_3$ in the right panel), while mediation analysis additionally requires the ignorability of the mediator ($U_2$, $U_4$, and $U_1$) imai2010identification. In some empirical studies, satisfying and justifying the second ignorability assumptions is challenging, even in the traditional randomized experiment.
In implicit mediation analysis, researchers do not seek to identify the exact causal mediation effect but rather aim to derive qualitative insights into the underlying mechanism bullock2021failings. For example, they may estimate the correlation between the treatment and mediator or between the mediator and the outcome blackwell2024assumption. The most popular approach utilizes HTEs. The underlying intuition is that if a specific mechanism is active for certain units, it will generate HTEs with respect to a particular covariate. If the mechanism is inert, it will fail to produce HTEs fu2023heterogeneous. This method is widely used because it does not require measuring the mediator or relying on strong identification assumptions.
Our identification strategy incorporates the strength of two approaches and bypass some ignobility assumptions. We first decompose the total treatment effect by highlighting the treatment effect on the mediator. As we show in section (ref), the new decomposition has the form of simple linear regression, where the `dependent variable' is the average treatment effect on the outcome and the `independent variable' is the average total treatment effect on the mediator. To non-parametrically identify the causal mediation effects (i.e. the slope), we assume that the average treatment effect on the mediator is not correlated with the error term. Under this assumption, even if confounders exist between the mediator and the outcome (so that sequential ignorability fails), we can still identify the causal mediation effect. In the instrumental variables literature, a related identification assumption appears in kolesar2015identification and in the Mendelian randomization framework bowden2015mendelian. Actually, fu2023heterogeneous shows that this kind of isolated mechanism assumption is also generically necessary for HTEs to reflect qualitative information about mechanism activation, even if not in a quantitative sense. In this respect, our approach does not impose additional requirements; if this assumption fails to hold for all covariates, HTEs cannot be used to study mechanism activation in the usual way.
To implement the strategy, we rely on HTEs on the mediator as our primary data sources. If the same research is conducted several times in different locations, we naturally obtain different treatment effects from those studies. In section (ref), we propose two more research designs to exploit those HTEs. In the first design, we suggest researchers use pre-treatment covariates to identify those subgroups, probably with data-driven methods like causal tree or forest wager2018estimation. Next, for each subgroup, we obtain the required data (average treatment effects on the outcome and on the mediator), which we refer to as Heterogeneous Subgroup Design. It is worth emphasizing that our goal is not to identify the true underlying subgroups, but rather to obtain sufficient variation in the data. The second design explores variation in treatment intensity. For example, in a Get-Out-The-Vote (GOTV) experiment, researchers might randomly assign different numbers of canvassing mails. By manipulating the magnitude or intensity of the treatment across different arms, we can induce varying treatment effects on the mediator, a design we refer to as the Multiple Treatment Meta Design.
The estimation and inference are discussed in section (ref). We then use Monte Carlo simulation to examine the performance of those estimators in section (ref). Our method demonstrates greater robustness and efficiency. We apply our methodology to estimate causal mediation effects in two empirical studies using actual data. The first study is about `Governance on Resources'. We demonstrate the estimation of mediation effects using just six group-level estimates derived from six different sites. The second examines the impact of information on voting behavior, where we calculate the mediation effect using individual-level data. We quantify the causal mediation effect of “bad/good information” about legislators on voting choice through voters’ perceptions of politicians’ effort. In section (ref), we offer a range of valuable extensions designed to assist applied researchers in meeting our core identification assumptions through both the design stage and the data analysis stage.
In general, we introduce an alternative identification strategy for causal mediation analysis. Causal mediation analysis is challenging, and each method relies on its own identification assumptions. We hope our method provides an additional useful tool for applied researchers who are eager to learn about causal mechanisms. Instead of relying on unconfoundedness assumptions, our method leverages HTEs and the theoretical structure of mechanisms. No single method is suitable for all study types: identification based on "exogeneity" or "ignorability" should be applied when researchers have sufficient covariates and understanding of potential confounders. Conversely, when confounders are unclear or unmeasured, our approach—relying on HTEs and knowledge of mechanisms—may be an effective alternative strategy.
In this section, we review basic definitions and notations for mediation analysis, if necessary, with the help of directed acyclic graphs (DAG).\footnote{Although there exist some debates about the potential outcome approach and graphic approach (see imbens2020potential, and Pearl's reply: http://causality.cs.ucla.edu/blog/index.php/2020/01/29/on-imbens-comparison-of-two-approaches-to-empirical-economics/), we still find both approaches have their own particular merits.}
Throughout this paper, we will use the Get-Out-The-Vote experiments as running examples to illustrate various concepts. Consider the scenario where researchers have observed that treatments such as door-to-door canvassing, phone calls, or mailings can significantly increase voter turnout green2019get. A critical question arises: through what mechanisms, do these treatments influence turnout? Mediation analysis helps us quantify the average causal effects along these pathways.
To better understand the mechanisms involved in this example, we draw on political economy theory. According to riker1968theory, there are four elements that influence a voter’s decision to abstain or vote, as shown in Figure (ref). (1) Cost of Voting ($C$): This includes costs related to information gathering, registration, and transportation, among others. Generally, higher costs discourage turnout. (2) Civic Duty ($D$): In democracies, voting is often portrayed as a civic responsibility. $D$ captures the utility derived from fulfilling this civic duty; fail to fulfill the duty may lead to social pressure as dis-utility. (3) Utility Difference ($B$): Voters are more likely to turn out if the utility gained from their preferred candidate winning significantly outweighs the utility loss if they lose. (4) Pivotality ($p_i$): This refers to the probability that a voter’s vote could decisively impact the election outcome, with $p_i \in [0,1]$ representing the probability. A lower likelihood of affecting the result may decrease the motivation to vote.
Consequently, a voter prefers to vote over abstaining if and only if $p_i B + D - C \ge 0$; that is, the utility from voting outweighs the costs. These four elements—cost, civic duty, utility difference, and pivotality—represent the mechanisms that can influence turnout. As the formula clearly demonstrates, Mechanisms 3 and 4 are correlated. In our causal mediation question, we specifically focus on the civic duty mechanism, measured by social pressure, to quantify its causal effect on voter turnout.
In the counterfactual approach of causal inference, causation and mediation are interpreted and decomposed with potential outcomes holland1986statistics. Let $T$ be the binary treatment and $Y$ be the outcome variable. The overall treatment effect of $T$ on $Y$ for individual $i$, denoted by $\tau^i$, is represented by $$ \tau^i = Y^i(1,M^i_1(1))- Y^i(0,M^i_1(0)). $$ For the total effects, mediators $M$ should consider potential outcomes under the treatment status. For instance, $M_1^i(1)$ represents the potential social pressure when the treatment $T=1$.
Total treatment effects can be decomposed into natural direct and indirect effects robins1992identifiability. We call $\delta^i(t)= Y^i(1,M^i(t))-Y^i(0,M^i(t))$ the natural direct effect for $t=0, 1$, where the mediator is set to the value it would have been under treatment $t$. It captures the effects that are not transmitted by the mediator of interest. \footnote{The terminology “natural” is in contrast to “controlled.” For controlled direct effect, we fix the mediator at a certain value $m$, rather than their potential outcomes under a given treatment assignment. Therefore, the controlled direct effect can be defined as $Y^i(1, m)- Y^i(0, m)$ acharya2016explaining.} Similarly, $\eta^i(t) = Y^i(t,M^i(1))-Y^i(t,M^i(0))$ is the natural indirect effect for $t=0, 1$. It denotes the treatment effect through the mediating variable. In our example, this refers to how the decision to turn out changes in response to the shift in social pressure from what it would be under the control condition ($M^i(0)$) to the treatment condition ($M^i(1)$), while holding the treatment status constant at $t$. \footnote{See more discussions in supplementary materials (SI) (ref).}
Due to the fundamental problem of causal inference, we are confined to the estimation of aggregate effects, typically averages. We therefore define $\tau$ as the average overall treatment effect: $\tau := \mathbb{E}[\tau^i]$, $\delta:=\mathbb{E}[\delta^i]$ as the average direct effect, and $\eta:= \mathbb{E} [\eta^i]$ as the average indirect effect for mechanism represented by the mediator $M$. As is standard, illustrated in the left DAG in Figure (ref), we can then decompose the total causal effect into the sum of natural direct and indirect effects: $\tau = \delta(t)+\eta(1-t)$.
Historically, path analysis and structural equation modeling (SEM) is the predominantly used framework for conducting mediation analysis mackinnon2012introduction,hong2015causality. As a special case, baron1986moderator develop the linear additive model using a single treatment, mediator, and outcome variable. It comprises two main equations:
Replacing $M$ (equation (ref)) in (ref), we obtain the reduced form as follows:
Traditionally, parameter before the treatment $T$, i.e., $\tau=\delta+\tilde{\beta} \gamma$ in (ref) is interpreted as the total treatment effect; $\delta$ is interpreted as the direct effect; and $\tilde{\beta} \gamma$ is interpreted as the indirect effect. The right part of Figure (ref) illustrates this DGP. Unlike the counterfactual approach, structural models implicitly incorporate parametric assumptions, such as linearity and constant effects.
Explicit mediation analysis aims to quantify the exact mediation effects. This essentially requires the sequential ignorability to non-parametrically identify the causal mediation effect. Multiple versions of sequential ignorability assumption exist. One of the most concise versions is given by imai2010identification. Formally, it has two important parts. The first part is similar to the unconfoundedness assumption in causal inference. Essentially, it requires treatment assignment to be ignorable given the observed pretreatment confounders $X$:
Notably, the treatment value is different for outcome $Y$ and mediator $M$. Hence, it specifies the full joint distribution of all the potential outcome and mediator variables ten2012review. The second part entails the mediator is ignorable given the observed treatment and pre-treatment confounders:
In the assumption (ref), the mediator takes the value at the “current” treatment assignment $t$, but the potential outcome is under treatment assignment $t'$. For example, it requires that the potential social pressure under treatment condition is independent of the potential turnout under control condition, given the individual is under the treatment status and has pre-treatment variables $X^i=x$. This cross-world argument makes it challenging to be satisfied in some studies.
One important limitation of both assumptions is that all covariates $X$ must be pre-treatment; generally, the natural indirect effect is not identified even if we have data on the post-treatment confounders avin2005identifiability. There have been several attempts to relax these constraints by introducing additional assumptions. robins2003semantics proposes the finest fully randomized causally interpreted structured tree graph (FRCISTG) model. Under this semantic framework for causal DAGs, the “cross-world/indices” property can be relaxed, allowing for post-treatment confounders. However, an additional no-interaction assumption is required to nonparametrically identify the causal mediation effect. Similar assumption is also discussed in glynn2012product. rudolph2023efficient and tchetgen2014identification explore how a monotonicity assumption can help identify causal mediation effects when a confounder is affected by treatment. Recently, a new strand of literature has proposed another estimand, the interventional indirect effect, which circumvents these strong assumptions diaz2021nonparametric,miles2023causal.
Typically, to non-parametrically identify natural indirect effects, sequential ignorability requires us to account for all pre-treatment confounders affecting treatment, mediator, and outcome. Practically, researchers hardly observe and measure all confounders, and it is challenging to ensure all confounders are under control. Because of the challenge, in practice, researchers often rely on modeling assumptions to estimate the causal mediation effect. The traditional choice is the linear regression model (equation (ref)-(ref)). Several parametric assumptions are required to identify parameters and thus the indirect effect mackinnon2012introduction. In particular, we need two assumptions:
Those function-form assumptions can also be interpreted by counterfactual languages jo2008causal,sobel2008identification. Generally, explicit mediation analysis under the linear regression model still requires several “exogeneity” assumptions to identify parameters in regression equations ($\tilde{\beta}$ and $\gamma$ or $\tau$ and $\delta$). As shown in Figure (ref), both non-parametric and model-based assumptions require controlling $U_1$ and $U_2$. However, in practice, it is difficult to measure and control all such confounders.
Given these challenges, researchers frequently rely on implicit mediation analysis bullock2021failings. Although it does not necessarily reveal the causal mediation effect, it can provide useful qualitative insights into the mechanism. However, many practices incorporate unstated assumptions.
The dominant approach to implicit mediation analysis is to examine HTEs to assess whether a proposed mechanism is active. Although mediation is conceptually distinct from moderation, moderation analyses are often used to provide evidence about underlying causal mechanisms. The intuition is straightforward. For example, if a mechanism explains the causal relationship, we would expect individuals with different values of a given covariate to exhibit different treatment effects, as the covariate is assumed to moderate the effect. A major advantage of this approach is that it does not require direct measurement of the mediator. However, fu2023heterogeneous find that valid inference of a mechanism from HTEs requires an exclusion assumption, ruling out the possibility of the covariate moderating through alternative pathways. This is intuitive because, otherwise, the observed HTEs cannot be attributed to the underlying mechanism.
Moreover, in some cases, we cannot directly observe the outcome variable that is affected by the mechanism implied by the theory; instead, we can only measure a related outcome. For example, information about corruption involving the incumbent may influence voters’ relative evaluations of the incumbent and other candidates. While we cannot directly observe changes in utility, we can observe their behavioral manifestation in voting choices. In such cases, even when the exclusion assumption holds, the presence of HTEs does not necessarily indicate activation of the underlying mechanism.
When the mediator is measured, another approach is to test whether the treatment affects the proposed mediator; blackwell2024assumption clarify that a monotonicity assumption is needed to draw sharper conclusions.
In this section, we introduce the new identification strategy, which starts with synthetic causal decomposition under the counterfactual approach but emphasizes the mechanical process as the structural approach. Under the new decomposition, we then convert the difficult mediation problem into a simple linear regression problem. It turns out to be quite general and simple to identify the causal mediation under this new structure. We then compare our identification strategy with existing methods, emphasizing that it provides a useful alternative in settings where the sequential ignorability assumption is violated.
Recall that total causal effect can be decomposed into direct and indirect effects. As is standard, we begin with the “no interaction effect” situation in the main text ($\delta=\delta(1)=\delta(0)$ and $\eta=\eta(1)=\eta(0)$).\footnote{See SI (ref) for the discussion on the interaction effect.} The identification strategy starts with a straightforward transformation of this decomposition.
The first two lines are two decompositions. In the line (ref),we multiply and divide the average indirect effect $\eta=\mathbb{E} [Y^i(0,M^i(1)-Y^i(0,M^i(0))]$ by the same term $\gamma=\mathbb{E}[M^i(1)-M^i(0)]$. It is the average effect of treatment on the mediator of interests. We define $\frac{\eta}{\gamma}=\frac{\mathbb{E} [Y^i(0,M^i(1)-Y^i(0,M^i(0))]}{\mathbb{E}[M^i(1)-M^i(0)]}$ by $\beta$, which denotes the ratio of how pure indirect effect changes according to one unit change of $\gamma$. If researchers plan to use SEM to conduct mediation analysis, under the linear model (ref) and (ref), we can easily see that $\beta=\tilde{\beta}$. Therefore, $\beta$ can be interpreted as the effect of the mediator $M$ on the outcome $Y$ \footnote{ Under linear SEM, we implicitly assume $\varepsilon_1$ and $\varepsilon_2$ are independent. From equation $\eqref{equ:sem1}$ and $\eqref{equ:sem2}$, we observe
Therefore, $\eta=\mathbb{E}[Y^i(0,M^i(1))]-\mathbb{E}[Y^i(0,M^i(0))]=\tilde{\beta} \gamma$, where $\mathbb{E}[M^i(1)-M^i(0)]=(\alpha_2+\gamma)-\alpha_2=\gamma.$}. Finally, we use simple notation to represent the final decomposition $\tau=\delta+\beta \gamma$.
If we can identify $\beta$ and $\gamma$, equivalently we can identify $\eta=\beta\gamma$. In most empirical studies, $\gamma$ and $\tau$ are easy to identify if treatment is as if random through careful research designs. The remaining part is to identify the parameter $\beta$.
A critical insight in causal inference is the recognition that causal effects vary across populations and even among individuals. This implies that the $\gamma$ is a random variable (We will discuss how to get this sample in the section (ref)). Therefore, equation (ref) can be written as $\tau_k=\delta_k+\beta\gamma_k$. We use subscript $k$ to emphasize that they are random rather than fixed values. Next, we complete the transformation by adding and subtracting the expectation of $\delta_k$.
In the Line (ref), we add and subtract the expectation of $\delta_k$; in line (ref), we define $\varepsilon_k=(\delta_k-\mathbb{E}[\delta_k])$. In the structural equation, $\beta$, the ratio of the indirect effect to the treatment effect on the mediator is assumed to be constant. In a more general case, $\beta$ can also be random, and is denoted by $\beta_k$. Here, we focus directly on the effects. For strict frequentists, one can view this framework as analogous to a random-effects model in meta-analysis, which reconciles the “random” nature of $\gamma_k$ and $\tau_k$. A direct analogue is provided in Section (ref).
Equation (ref) should be familiar to readers: it is a simple linear regression model. The key assumption for identifying $\beta$ is that the direct effect $\delta$ is uncorrelated with the effect of treatment on mediator $\gamma$.
The assumption requires that, $\gamma$, the average treatment effect on the mediator of interest, is not correlated with the average direct effect. If there exist multiple mechanisms, generally, we should interpret $\delta_k$ as effects from all other possible mechanisms, see SI (ref) and (ref). Therefore, this assumption implies that the mechanism of interest is somewhat isolated from other mechanisms. A similar assumption has been proposed in the multiple instrumental variables literature kolesa2013estimation, bowden2015mendelian. A related condition is also required for implicit mediation analysis using heterogeneous treatment effects fu2023heterogeneous. Intuitively, if the covariate moderates multiple mechanisms, the observed HTEs may not provide information specifically about the mechanism of interest. Similarly, we require that the HTE in $\gamma$ reveals information solely about $\gamma$, rather than about other mechanisms. Technically, this assumption is equal to that of the traditional simple linear regression assumption $\mathbb{E}[\gamma_k\epsilon_k]=0$ and implies $Cov(\gamma_k,\varepsilon_k)=0$.\footnote{To see this, $\mathbb{E}[\gamma_k\epsilon_k]=\mathbb{E}[\gamma_k(\delta_k-\mathbb{E}\delta_k)]=Cov(\gamma_k,\delta_k)=0$ and $Cov(\gamma_k,\varepsilon_k)=Cov(\gamma_k,\delta_k-\mathbb{E}\delta_k)=Cov(\gamma_k,\delta_k)=0$. }
The validity of this assumption depends on the theoretical framework in question. To clarify, let us revisit our ongoing example. As we introduced earlier, there are four mechanisms—cost, civic duty, utility difference, and pivotality—that influence the decision to vote. If a researcher aims to examine the impact of mailings that encourage voters with the message "DO YOUR CIVIC DUTY—VOTE!" it is plausible to assume that such an intervention mainly affects the civic duty utility. Moreover, according to the theory, since $p_i b+D - C > 0$, other mediators are distinct and capture different mechanisms. Therefore, it is reasonable to assume that the effect of the canvassing treatment on civic duty is uncorrelated with its effects on other mechanisms. However, in general, assumption (ref) requires the belief that no other variable concurrently moderates both the mechanism of interest and other mechanisms. We can, however, allow for correlations among other mechanisms, such as the correlation between utility difference and pivotality, as illustrated in (ref). In the section (ref), we will provide more discussions and techniques on meeting the identification assumption in application.
In Proposition (ref), we propose a simple estimator $\hat{\beta}$ to estimate the unknown $\beta$.
The estimator $\hat{\beta}$ is exactly the simple OLS estimator for the slope. The assumption $Var(\gamma_k)>0$ is technical; it guarantees that we have “random” observations of $\gamma$. The proposition indicates that if we assume the treatment effect on the mediator of interest is not correlated with the effects of other mechanisms, then $\beta$ can be consistently estimated. Combined with the information of $\gamma$, the overall average mediation effect $\mathbb{E}[\eta_k]$ can be estimated by $\hat{\beta} \sum_{k=1}^K \gamma_k \mathbb{P}(\gamma_k)$, where $\mathbb{P}(\gamma_k)$ can be consistently estimated by the proportion of sample size $k$ relative to the total sample size. The \href{https://github.com/Jiawei-Fu/mechte}{R package} will return all essential statistics. It is important to note that the estimates produced by other mediation analysis techniques, which do not account for heterogeneity, can be interpreted as average effects over implicit heterogeneous effects.
Readers may be concerned that we are assuming a linear relationship between $\tau_k$ and $\gamma_k$. Actually, we do not assume such linearity; $\tau_k=\mathbb{E}[\delta_k] + \beta\gamma_k + \varepsilon_k$ is the structural model derived from causal decomposition. There is an important distinction between this structural model and statistical linear regression models. First, in most cases, people assume the linear statistical relationship between data, that is, $\tau_k$ and $\gamma_k$ here. Nevertheless, our model $\tau_k=\mathbb{E}[\delta_k] + \beta\gamma_k + \varepsilon_k$ is naturally guaranteed by the nature of the causal effect. In the counterfactual framework, we can always additively decompose the total causal effect into two pieces. Second, in statistical applications, people assume the expectation of the error term in their population model is zero: $\mathbb{E}[\varepsilon_k]=0$. However, here, this property is guaranteed by construction, not by assumption: $\mathbb{E}[\varepsilon_k]=\mathbb{E}[\delta_k]-\mathbb{E}[\delta_k]=0$. Because of this property, in contrast to OLS, the unbiasedness of our estimator requires a slightly weaker mean independence assumption ($\mathbb{E}[\delta_k|\gamma_k]=\mathbb{E}[\delta_k]$), rather than the zero conditional mean assumption ($\mathbb{E}[\delta_k|\gamma_k]=0$). We summarize this result in SI (ref).
What are the main advantages of our identification strategy? First, it does not require that mediator is ignorable. In other words, we allow unobserved confounders that simultaneously affect the mediator and the outcome variable (i.e., $U_2$ in the Figure (ref)). As mentioned in the section (ref), current methods cannot efficiently address this unconfoundedness problem without further assumptions.
Second, we allow researchers to simultaneously estimate both treatment and mediation effects. The causal mediation, is simply a byproduct after identifying the treatment effects ($\tau$ and $\gamma$). We do not need other advanced techniques to identify the indirect effect except simple OLS. We will introduce exact estimation methods and research designs in the next section. We can use both aggregate-level data and individual-level data to get the causal mediation effect. Therefore, we believe our methods can be applied in a variety of empirical studies when sequential ignobility does not hold.
To implement the strategy, we need a random sample of treatment effects, $(\gamma_K,\tau_K)$. It is common for the same treatment to generate varying (average) treatment effects across different populations, such as those in different countries or areas. This is similar to getting multiple effects when conducting meta-analysis. Rigorously, we assume there are $k$ independent trials that generate $k$ study-level effects $(\gamma_k,\tau_k)$, where $\gamma_k$ is assumed to come from a super-population (normal) distribution. In the random-effects model, $\gamma_k= \mu + \Delta_k +e_k$, where $\mu$ is the overall effect, $\Delta_k$ is the deviation of $k's$-effect from the overall effect, and $e_k$ is the sampling error. As is standard in meta-analysis, we require that $\Delta_k$ be independent of study $k$, or exchangeable in the language of Bayesian Statistics higgins2009re. This assumption is critical for statistical inference; therefore, we will proceed under the premise that it holds in the subsequent analysis.
Under the above model, practically, how can we find study-level effects $(\gamma_k,\tau_k)$? If the same research is conducted several times in different locations, we naturally obtain data from those studies. If not, it is still possible to `create' multiple studies within a single study. In this section, we introduce two general research designs.
Multiple Treatment Meta Design. The key idea is that we can adjust the treatment to induce heterogeneous effects. For example, in the Get-Out-The-Vote (GOTV) experiments, researchers could modify the treatment by having some groups receive one mail, while other groups receive two or three mails. See Figure (ref). Given the same treatment, varying the intensity or magnitude of the treatment can likely induce different effects. Formally, we still use $G_k \in \{T_1,T_2,...,T_l\}$ to denote different sub-types of the treatment. If individual $i$ belongs $G_k=T_j$, it means the individual receives treatment intensity $T_j$.
Heterogeneous Subgroup Design. This design exploits the heterogeneous effects from different population. Naturally, the same treatment may generate different (average) treatment effects for different population, for example, population in different country or area. To be concrete, for example, suppose researchers desire to understand how the mailing affects turnout through social pressure in the GOTV experiment gerber2008social, as shown in Figure (ref). Formally, each individual is characterized by a vector of pre-treatment covariates $X=(X_1,X_2,...,X_l)$ that can moderate treatment effects on the outcome and the mediator. We can subsequently define several subgroups $G_k$, where $k \in\{1,2,...,K\}$ according to $X$. Suppose $X_1$ is gender, and $X_2$ is age. We can define group $G_1=\{X_1=Male, X_2>30\}$, comprising individuals who are male and older than 30. Each individual $i$ should belong to only one group. An assumption is that for each group, treatment generates different and independent average treatment effects $\tau$ and $\gamma$. How should we identify these groups? If several similar studies are conducted in geographically different areas, similar to meta-analysis, then those studies automatically generate a sample of effects. Alternatively, a data-driven method such as causal tree or forest can be used wager2018estimation \footnote{Technically, the data obtained from such a design represent conditional average effects. The analysis remains valid under the assumption on $\beta$, as we take the average across subgroups.}.
Synthetic Methods. People may incorporate these two designs and define finer subgroups. The incorporated subgroup $G_k = \{T\in \textbf{T},X_1 \in \textbf{X}_1,X_2 \in \textbf{X}_2,...,X_l\in \textbf{X}_l\}$ is defined by treatment types and covariates, where $\textbf{X}_l$ denotes a set possible values of $X_l$. For example, in the GOTV design, for each intensity of treatment, we can find subgroups defined by covariates. A possible subgroup could be $G_k=\{Single Mail, X_1=Male, X_2>30\}$. If individual $i$ is in this group, it implies that individual $i$ is male, older than 30, and receive treatment phone call.
Regardless of the method used, the objective is the same: to obtain data suitable for regression analysis. Our goal is not to identify the true, unobserved subgroups. Rather, we seek to generate variation in the estimated effects $\hat{\gamma}_k$, which—by the causal decomposition—correspond to associated estimates $\hat{\tau}_k$. If observations are partitioned into subgroups at random, the resulting $\hat{\gamma}_k$ are likely to be similar across $k$. Although such estimates can still be used, the limited variation in $\hat{\gamma}_k$ leads to large standard errors. We therefore prefer subgroup constructions that induce substantial heterogeneity across $k$. Importantly, these subgroups need not correspond to the true underlying subpopulations. Identifying the correct latent subgroups is an interesting and important research problem in its own right, but it is not required for our approach.
Once we obtain the heterogeneous effects $(\hat{\gamma}_k,\hat{\tau}_k)$, we are ready to estimate $\beta$ and, consequently, the causal mediation effects. Because the second-stage regression uses estimated group-level effects, \((\widehat\tau_k,\widehat\gamma_k)\), rather than the infeasible quantities \((\tau_k,\gamma_k)\), the estimator is a generated-regressor estimator. To ensure that the first stage has no first-order effect on the second stage, we require the first-stage heterogeneous effects $(\hat{\gamma}_k,\hat{\tau}_k)$ to be regular \(n_k^{1/2}\)-consistent treatment-effect estimators. Formally, for each group $k$, let $\widehat\gamma_k = \gamma_k + u_{k}$, and $\widehat\tau_k = \tau_k + v_{k}$, where $u_k$ and $v_k$ denote the estimation errors.
If the groups are identified ex ante, then this rate condition is generally satisfied under standard regularity conditions. When subgroups are selected using a data-driven method, honesty or sample splitting provides a convenient sufficient condition. Specifically, we partition the data into two disjoint subsamples, $I_1 \cup I_2$. The first subsample $I_1$ is used to identify the $K$ subgroups, while the second subsample $I_2$ is used to estimate $(\hat{\gamma}_k,\hat{\tau}_k,\hat{\sigma}^2_{uk})$, where $\hat{\sigma}^2_{uk}$ denotes the estimated asymptotic variance of $\hat{\gamma}_k$. In practice, $k$-fold cross-fitting is recommended, as is standard in the current DML literature. This separation ensures that, conditional on the selected partition, the group-level estimators behave like standard treatment-effect estimators.
The rate condition $\sum_{k=1}^K \frac{1}{n_k} \to 0$ ensures that this total first-stage noise vanishes asymptotically. In the balanced case, the condition $ \sum_{k=1}^K \frac{1}{n_k} \to 0$ reduces to \(K/n_k\to 0\). In general, this condition requires that the sample size within each group grow sufficiently fast so that the accumulated first-stage estimation error disappears asymptotically. A consistent heteroskedasticity-robust estimator of \(V_\beta\) is
where $ \widehat\varepsilon_k = \widehat\tau_k-\widehat\alpha_K-\widehat\beta_K\widehat\gamma_k$, and $ \widehat\alpha_K = \bar{\widehat\tau}-\widehat\beta_K\bar{\widehat\gamma}$.
Proposition (ref) is stated under this asymptotic framework, in which subgroup sizes grow such that the first-stage estimation error in \(\widehat\gamma_k\) and \(\widehat\tau_k\) vanishes. In finite samples, however, when group sizes are small, one may still be concerned about measurement-error attenuation arising from first-stage estimation noise. In such settings, we suggest using the well-known simulation--extrapolation (SIMEX) method cook1994simulation. In SI (ref), we generalize and adapt standard asymptotic normality results to our setting, which involves two-step estimation with non–i.i.d.\ data.
Using either approach, we obtain a well-behaved estimator for $\beta$. Our primary estimand, however, is the overall average indirect effect, defined as the weighted average of $\beta_k \gamma_k$ across the $k$ subgroups. The variance of the estimator can be derived through Delta method. As shown in SI (ref), under further sample splitting, the asymptotic variance equals $(\frac{1}{K}\sum_{k=1}^K \hat{\gamma}_k)^2V_{\beta}$ under a further sample splitting. To achieve better finite-sample performance, we instead employ an intersection–union test. For simplicity, let $\lambda_k$ denote the population proportion associated with $\gamma_k$, and let $\hat{\lambda}_k$ represent its consistent estimator. The following proposition summarizes how to test the null effect of the overall average indirect effect based on $p$-values. Let $\hat{\sigma}_{\beta}$ be the standard error. $\Phi_z$ is the cumulative distribution function of the standard normal distribution.
In fact, $p_{\beta}$ and $p_{\gamma}$ are $p$-values for two sub-parts in the intersection-union test. From the above proposition (ref), it follows that the overall $p$-value associated with the null effect of the overall indirect effect is defined as $\max [p_{\beta}, p_{\gamma}]$. The next proposition provides an conservative confidence interval based on the intersection-union test. We use $\hat{\gamma}_0$ to denote $\sum_{k=1}^K\hat{\gamma}_k \hat{\lambda}_k$.
Based on both propositions, conducting sub-group inference is also straightforward. We encapsulate estimation and inference in the \href{https://github.com/Jiawei-Fu/mechte}{R package}. In the package, we also include Cochran's Q and Higgins & Thompson’s $I^2$ from meta-analysis to assess whether $\gamma_k$ values exhibit true heterogeneity rather than random error.
In this section, we employ Monte Carlo simulations to evaluate the effectiveness of our methods by comparing them with the current methods under sequential ignorability. Furthermore, we apply our methodology to real data from two distinct experiments—one using aggregate data and the other using individual-level data—to illustrate its application in real studies.
We generate heterogeneous treatment effects for each individual $i$ by assuming $10$ subgroups using the following simple model:
$$
$$ Treatment $T_i$ is randomly drawn from a standard normal distribution, along with two error terms $\epsilon$. The average treatment effect on the mediator $M_i$ varies across groups, with $\gamma_k \in \{1,2,3,4,5,6,7,8,9,10\}$. The effect of the mediator on the outcome is $\beta$. We also consider an `unobserved' confounder $u_i$ that simultaneously affects the mediator and outcome, with magnitude $\omega \in \{0,0.5,1,1.5,2\}$.
We first set $\beta=0$ so that the AMCE is zero. In Table (ref), we report the empirical distribution of p-values and the $95\%$ confidence intervals (averaged over simulations) for both methods across different values of $\omega$. Confidence intervals are constructed using the method introduced in Section (ref). It is evident that when $\omega=0$, the sequential ignorability assumption holds. In the first row, therefore, both methods yield estimates that are very close to the theoretical AMCE of 0. As $\omega$ increases, the sequential ignorability assumption is violated. Our method remains robust to this violation, in the sense that the confidence interval continues to cover 0 and the Type I error rate in the first column remains close to the nominal level. However, for methods relying on sequential ignorability, the rejection rate increases and the confidence interval no longer covers the true value of zero.
Next, we study the power of our method. To do so, we increase $\beta$ so that the AMCE changes from 10% to 30% of the ATE. We do not further increase this percentage because the power already reaches 1. As shown in Figure (ref), when the AMCE exceeds 20% of the ATE, the power of our method reaches approximately 80%.This suggests that our method has strong power to detect relatively small AMCEs in practice.
We also explore how the number of groups and the group size influence statistical power using a similar DGP as before. The results are shown in Figure (ref). In general, both increasing the number of groups and increasing the group size enhance statistical power. However, in this simulation, increasing the number of groups has a more pronounced effect than increasing the group size. This suggests that heterogeneity provides more information than merely increasing the precision of the estimates. In practice, researchers are advised to conduct similar power analyses to optimize their research designs.
Evidence in Governance and Politics (EGAP) \footnote{https://egap.org/} funds and coordinates multiple field experiments on different topics across countries. This collaborative research model is called “Metaketa Initiative.” In Metaketa III, they examine the effect of community monitoring on common pool resources (CPR) governance. To causally answer this question, slough2021adoption conducted six harmonized experiments with the same `meta' treatment (community monitoring) but heterogeneous CPRs and treatment sub-types, as shown in the table (ref).
In their study, the authors report effects on multiple outcome variables, which include resource use, user satisfaction, user knowledge about community’s CPRs, and resource stewardship. They also investigate the underlying mechanism: how monitoring affects those outcomes through different channels. However, their analysis is limited to examining the treatment effects on mediators, which does not necessarily delineate the precise causal mediation effects. We intend to use their data to illustrate how to apply our causal mediation analysis with aggregate-level data; to be specific, we ask “How does the treatment (i.e. monitoring) affect user knowledge about community CPRs through altered perceptions of sanction likelihood for CPR misuse?"
The six experiments naturally provide us with six subgroups. While it is feasible for researchers to further segment subgroups within each experimental site, our focus will be on these six primary subgroups. To estimate ACME, we require specific data: (1) the average treatment effects on both the outcome and the mediator, and (2) the standard errors associated with these effects. These data points are represented in the two-dimensional Figure (ref), where each dot represents an estimate and each line indicates the $95\%$ confidence interval (CI). Generally, most confidence intervals cross zero, particularly for the treatment effects on the mediator (sanction), as illustrated in Figure (ref). Here, no single estimate significantly deviates from zero at the 0.05 level.
However, the data reveal a clear pattern: an increased treatment effect on the perception of sanctions correlates with an increased effect on knowledge about the resource. To quantify the mediation effect, we estimate $\beta$, with the estimated slope being $\beta = 0.68$ and the $p$-value =0.049. This estimate is also depicted by the green line in Figure (ref). The figure is also useful for determining whether $\beta$ is constant. If it is indeed constant, we should expect the data points to display a relatively linear relationship. As previously mentioned, in the structural approach to mediation analysis, we can generally interpret $\beta$ as the effect of mediator (sanction) on the outcome (knowledge). This confirms our expectation that an increased perception of sanctions enhances the incentive to gain more knowledge about common-pool resources.
To obtain the average causal mediation effect, we need to multiply $\beta$ by $\gamma$ (the ATE on the mediator) and possibly weight this by the sample size in each experiment. We find that the estimated ACME is $0.01$. Since all $\gamma$ in six experiments are not statistically significant at $0.05$ level, it is also challenging to achieve a significant ACME. However, the positive ACME does suggest that the monitoring enhances knowledge about CPRs can be explained by a mediation effect through the perception for CPR misuse.
Accountability is a cornerstone of democracies and is fundamental to good governance. However, in reality, voters often lack sufficient information about politicians' performance. Many organization and civil society groups have dedicated efforts to disseminate such information to the electorate. A pivotal question arises: “Do informational interventions influence voters' behaviors and thereby promote accountability? If yes, what is the key mechanism?" Numerous field and survey experiments have sought to quantify this treatment effect; yet the findings are inconclusive incerti2020corruption,dunning2019voter. Furthermore, our understanding of how information influences voting behavior is limited. In a few experiments, researchers have measured intermediate outcomes and explore potential mechanisms. Nevertheless, these intermediate outcomes are not ignorable given the treatment status, making the estimation of mediation effects challenging. In this section, we will illustrate the use of our method in one field experiment from Benin, demonstrating how to identify and estimate the causal mediation effect in an information experiment using individual-level data.
Around 2015 National Assembly elections in Benin, as a part of Metaketa I, adida2019under randomly disseminated information about the performance of incumbent legislators to voters through videos. These videos provided official data on four key performance dimensions: (1) attendance rate at legislative sessions, (2) frequency of posing questions during these sessions, (3) committee attendance rate, and (4) productivity of committee work. One of the primary outcome variables was individual voting choice, which was captured via baseline and endline surveys. The surveys also gathered intermediate variables, such as voters' perceptions of the incumbents' effort/hardworking. Overall, the intervention did not significantly affect the incumbents' vote shares, aligning with the results from most other field experiments dunning2019voter. However, a subsequent meta-analysis highlighted a notable correlation between voters' perceptions of effort and support for incumbents dunning2019information. As emphasized by authors, this correlation does not illuminate any causal relationship due to the design. Nevertheless, it indicates a potential indirect effect of information on voting behavior mediated by perceptions of hard work. Thus, we intend to apply our method to their individual-level data to directly estimate this causal mediation effect.
To apply our method to individual-level data, the initial step involves identifying potential heterogeneous subgroups. This identification can be achieved through data-driven methods. We employ the widely-used causal tree approach to detect these subgroups, estimating the heterogeneous treatment effect on the mediator (effort) using individual pre-treatment covariates, such as age, gender, wealth, and political attitudes.\footnote{For details, please refer to the replication materials.} As depicted in Figure (ref), the informational effect on the perception of effort varies according to factors like coethnic, education, age, wealth status etc. With different algorithms, the identified subgroups may vary. For example, we can obtain alternative subgroup classifications, as shown in the SI (ref). However, if the identification assumptions hold, $\gamma_k$ and $\tau_k$ should capture the same underlying information, and thus the estimate of $\beta$ remains nearly identical. This finding is also empirically confirmed.
Next, we estimate the average treatment effect on vote choice across various subgroups, focusing on the “bad information” arm, where the information reveals poor performance by the incumbent. These estimates are then used to calculate the indirect effect using SIMEX, treating them as aggregate-level data. Before that, it is helpful to plot the data and check for a clear linear relationship. If the data show linearity, it provides evidence that $\beta$ is constant, and we do not need to account for non-linearity. The final results are illustrated in Figure (ref). We found a clear linear trend and that the estimated $\beta$ is 0.3, significant at the 0.1 level according to the SIMEX analysis. This suggests that the perception of high effort by the incumbent increases potential votes. The average causal mediation effect is $-0.08$, also significant at the 0.1 level. Consequently, we deduce that while the overall effect of information may not be substantial enough to detect in the field experiment, there is a significant indirect effect of bad information through voters' perceptions of politicians' effort. Specifically, bad information leads to fewer votes for the incumbent due to perceived lower effort. However, no significant results were detected for the “good information” arm, although the average mediation effect is positively aligned with our expectations.
Our identification strategy hinges on a pivotal assumption: $Cov(\gamma_k,\delta_k)=0$, implying that the treatment effect on the mediator is not correlated with other mechanisms. How can we assess this assumption in practice? The new identification strategy provides two opportunities to address the identification problem—during the design stage and the data analysis stage.
First, researchers can assess this assumption during the design stage, before any experiments or data collection. (1) This can be achieved through a theoretical examination. Understanding the underlying mechanisms necessitates a theoretical foundation. For instance, considering the voter turnout example in section (ref), theory identifies four major mechanisms that may influence a voter's decision to abstain or vote: relative payoff from the favored candidate, civic duty, the probability of being pivotal, and voting cost. If theory suggests these mechanisms are uncorrelated, then our assumption is valid. In the turnout example, most theories propose that these mechanisms represent distinct aspects of voting calculus and lack evident correlation. For example, the probability of being pivotal is determined by the size of the population, whereas civic duty encompasses all moral considerations, which are unlikely to influence preferences for the candidate. These considerations are reflected in the formula $p_i b + D - C > 0$, where additivity implies a form of independence among the mediators. Given a clear treatment, if the mediators capture distinct mechanisms, the identification assumptions are more likely to hold. (2) If conducting an experiment, researchers can control the treatment to minimize potential correlations between $\gamma$ and $\delta$. For instance, when varying treatment intensity, it is advisable to limit the treatment elements to prevent correlations with other mechanisms.
Second, in the data analysis phase, we can leverage our novel decomposition, which takes the form $\tau_k=\mathbb{E} [\delta_k] + \beta \gamma_k +\epsilon_k$. This formulation highlights a close analogy to the traditional linear regression model. Should the assumption fail, variables denoted by $X$, moderating both the treatment effects on the mediator and other mechanisms, may exist. Thus, akin to the approach in traditional linear regression, we control for $X$ to `purify' the error term. As our analysis is based on aggregate-level data (i.e., expected treatment effects), we need to account for $\mathbb{E}[X_k]$, the average of $X_k$ in the respective groups, in the linear regression model. More advanced and robust techniques, such as double machine learning, can also be employed. Cross-fitting is additionally recommended to reduce overfitting.
Owing to the challenge of obtaining a large number of observations of average treatment effects $(\gamma_k,\delta_k)$ in practice, indiscriminately adding all covariates in the linear regression model without careful consideration is not advised. This underscores the crucial role of theory in mediation analysis. The decision to include a covariate hinges on theoretical considerations regarding its potential to modify other mechanistic pathways. With theory, researchers can readily identify the most `important' omitted moderators, as measured by $R^2$, and incorporate them into the regression to enhance the robustness of their estimates oster2019unobservable. With the extension to linear regression, researchers may also apply traditional sensitivity analysis to evaluate the robustness of their estimates against unobserved confounders cinelli2020making.
Causal mediation analysis is inherently challenging. We hope our method provides an additional useful tool for applied researchers who seek to better understand causal mechanisms. In this study, we propose an alternative identification assumption and strategy that can enable researchers easily estimate causal mediation effects. The method converts the intricate mediation problem into a simple linear regression problem. Based on the isolated mechanism assumption, once researchers identify the treatment effects on the mediator and the outcome, our approach can consistently estimate the indirect effect. The proposed strategy enables researchers to address the identification problem in both the design and analysis stages.
While our method reduces the reliance on unconfoundedness assumptions, it may require greater data collection efforts to obtain precise estimates. Thankfully, in the era of big data, this should not be a major problem nowadays. Moreover, it is important to recognize that researchers should select appropriate mediation analysis techniques based on the specific design of their studies. All methods come with their own set of assumptions. There also remain several avenues for future research. Questions like post-treatment confounders, extending the method to non-binary treatments, identifying necessary assumptions for correlated mechanisms or integrating other causal identification strategies beyond instrumental variables for individual-level estimators remain open. Lastly, our method suggests a promising avenue to bridge causal mediation with causal moderation, indicating the potential for discovering other effective methodological combinations.
\renewbibmacro{in:} \printbibliography[]
\setcounter{page}{1}