Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,945 characters · 22 sections · 90 citation commands
$\ $
\fontsize{12}{14pt plus.8pt minus .6pt}\selectfont \centerline{\bf TOTAL-EFFECT TEST MAY ERRONEOUSLY REJECT} \centerline{\bf SO-CALLED “FULL" OR “COMPLETE" MEDIATION}
\centerline{Tingxuan Han$^1$, Luxi Zhang$^2$, Xinshu Zhao$^2$ and Ke Deng$^1$} \centerline{\it $^1$Tsinghua University, $^2$University of Macau} \fontsize{9}{11.5pt plus.8pt minus.6pt} \selectfont
\let\nofiles\relax
\fontsize{12}{14pt plus.8pt minus .6pt}\selectfont
The procedure to establish mediation, i.e., how an independent variable $X$ affects a dependent variable $Y$ through some mediator $M$, has been under debate. The classic causal-steps procedure requires that the total effect ($c$), i.e., the effect of $X$ on $Y$ without controlling $M$, be significant, now known as statistically acknowledged baron1986. It has been shown that the total-effect test can erroneously reject competitive mediation and is superfluous for establishing complementary mediation jiang2021total,hayes2009beyond, zhao2010reconsidering. Little is known about the third and the last type, the indirect-only mediation, in which the indirect ($ab$) path passes the statistical test while the direct-and-remainder ($d$) path fails.
Roughly equivalent to “full mediation” aka “complete mediation" in the classic quasi-typology of full, partial, and no mediation, indirect-only mediation is believed to be the strongest form of mediation baron1986,Hayes2022IntroMediation. While a revised procedure of causal steps kenny1998data, kenny2008reflections, kenny2021mediation allows researchers to “suspend” or “relax” the total-effect test when suppression, aka inconsistent or competitive mediation, is suspected, full (indirect-only) mediation and partial (complementary) mediation do not qualify for the relief. As of today, the total-effect test and, most importantly, the underlying conception of “mediation" and “effect" remain at the core of the criteria for establishing mediation across disciplines and languages (e.g. kenny2021mediation; mathieu2006clarifying; rose2004mediator; wen2004testing; wen2005compare; wen2014analyses; wen2022review). Section 2 below provides a brief review of the debate over the total-effect test.
This study is assigned several tasks. 1a) Provide a mathematical proof that the total-effect test can erroneously reject indirect-only mediation, including both sub-types, assuming least square estimation (LSE) and $F$-test. 1b) Provide derivation to show that the same results can be obtained assuming Sobel test. 2) Provide a simulation to duplicate the mathematical proofs and extend the conclusion to LAD-$Z$ test. 3) Provide two real-data examples, one for each sub-type of the indirect-only mediation, to illustrate the mathematical proof and the simulation outcomes. 4a) In light of the mathematical findings, propose revisions to the concepts, theories, and techniques of mediation analysis and other causal dissection analyses. 4b) Introduce the principles of a more comprehensive, i.e., more encompassing and more informative, alternative, process-and-product analysis (PAPA).
Mediation model suggests a causal chain where an independent variable $X$ affects a dependent variable $Y$ through a third variable $M$, known as a mediator. In a classic work that influenced generations of researchers, baron1986 defines mediation as a linear regression model:
where the errors are assumed to follow independent normal distributions: \[
\]
As shown in Figure (ref), the model involves two paths: 1) the indirect path $``X \rightarrow M \rightarrow Y"$ indicates the mediated effect of $X$ on $Y$ via mediator $M$, which equals to $a \times b$, and 2) the so-called direct path $``X \rightarrow Y \mid M"$ indicates the direct-and-remainder effect of $X$ on $Y$ while $M$ is controlled, represented by $d$. Reorganizing models (ref) and (ref), we have:
where $i_Y^* = i_Y + bi_M$, $\varepsilon_Y^* = \varepsilon_Y + b\varepsilon_M$, and $c=a \times b + d$ represents the total effect of $X$ on $Y$.
A formal typology has been established that features three types of mediation zhao2010reconsidering,zhao2011does: (1) Complementary mediation, where mediated effect and direct effect both pass the statistical partition tests and bear the same sign, i.e., $a\times b\times d>0$; (2) Competitive mediation, where mediated effect and direct effect both pass the tests and bear opposite signs, i.e., $a\times b\times d<0$; and (3) Indirect-only mediation, where mediated effect passes the test while direct effect fails, i.e., $a\times b\neq 0$ but $d=0$.
Almost all experts accept and adopt the above definition of mediation, namely a statistically acknowledged $a \times b$. The “causal-steps” procedure, however, adds another test, a statistically significant total effect, $c = a \times b + d$. This “total-effect test" is necessary because, according to this dominant doctrine, c represents the effect to be mediated; a statistically non-significant c indicates there is nothing to mediate hence no mediation is possible. Therefore, in the causal-steps procedure, if c fails to pass the statistical test, the mediation hypothesis is declared a failure and further analysis is stopped.
Although the causal-steps doctrine dominated mediation analysis across disciplines, whether and under what conditions the total-effect test should be required became a subject of discussion and debate. The opinions may be organized into three groups.
1) Complete acceptance: The seminal baron1986 requires the total-effect test as the first bar to pass and allows no exception for establishing mediation of any type. The procedure and the total-effect test as the first criterion have been recommended time and again by mediation experts across disciplines (e.g., judd1981,mathieu2006clarifying,rose2004mediator). As we write, the total-effect test, and more importantly the underlying conception of “mediation" and “effect," remain part of the standard guidelines for establishing mediation across disciplines and languages (e.g. kenny2021mediation; mathieu2006clarifying,rose2004mediatorwen2004testing,wen2005compare,wen2014analyses,wen2022review).
2) Conditional Suspension. Even before baron1986, statisticians had recognized “suppression”, aka “inconsistent models”, “confounding”, and, more recently, ‘‘competition", where the mediated path $a\times b$ and the direct-and-remainder path $d$ have the opposite signs breslow1980statistical,davis1985logic,judd1981estimating,lord1968statistical,mcfatter1979use,velicer1978suppressor. The subject came up more often after baron1986, as shown in numerous studies discussing it cliff1994all,cohen2013applied,conger1974revised, collins1998alternative,hamilton1987sometimes,hayes2009beyond,horst1941role,kenny1998data,kenny2021mediation,mackinnon2000equivalence,rucker2011mediation,sharpe1997relationship,shrout2002mediation,tzelgov1991suppression,zhao2010reconsidering.
It was collins1998alternative who proposed suspending the total-effect test for inconsistent mediation. Only a special type of the “inconsistent" models -- those with dichotomous variables -- would qualify for the relief. Other authors, at about the same time or shortly after, suggested suspending or dropping the total-effect test when there is a priori belief of suppression kenny1998data,kenny2021mediation,mackinnon2000contrasts,mackinnon2000equivalence,mackinnon2002comparison,shrout2002mediation. shrout2002mediation also added a type, by proposing to relax the total-effect test for distal processes and expectedly small effect sizes.
These authors often stressed the importance of retaining the total-effect test for all other types of mediation. shrout2002mediation, for example, argued that the total-effect test has conceptual usefulness, because “clearly, (researchers) need to first establish that the effect exists”.
3) Complete repeal: hayes2009beyond pointed out that suppression is a regular occurrence and recommended “researchers not require a significant total effect”. zhao2010reconsidering re-conceptualized “suppression” as “competitive mediation”, and presented a real-data example in which the competition caused a non-significant c even though the mediated path was strong. They hence concluded “There need not be a significant zero-order effect of $X$ on $Y,\ldots,$ to establish mediation”. rucker2011mediation agreed, concluding that “focusing on the significance of the $X\rightarrow Y$ relationship before or after examining a mediator might be unnecessarily restrictive”.
The discussions and debates, however, were conducted mostly on conceptual and philosophical levels without mathematically rigorous evidence. In fact, the subjects under discussion were not formulated as mathematical problems until recently.
zhao2010reconsidering and zhao2011does replaced the traditional one-dimension conception of mediation with the two-dimension framework, which allowed the authors to replace the dominant full-partial-no quasi-typology with the five-type typology. Through the prism of the new typology, and armed with the new vernaculars, the authors observed that the total-effect test may 1) be superfluous for establishing complementary mediation 2) erroneously reject competitive mediation, and 3) erroneously reject indirect-only mediation.
Although not backed by rigorous proof or systematic evidence other than the one real-data example zhao2010reconsidering, the three observations fostered three questions: Does the total-effect test help, harm, or neither for establishing each of the three types, i.e., complementary, competitive, or indirect-only mediation?
jiang2021total turned one observation, that the total-effect test is superfluous, into a mathematical conjecture, and produced a proof for the conjecture. They did so after transforming the task into a series of mathematical problems, which is to verify the geometry of rejection regions of hypothesis tests for $a$, $b$, $d$ and $c$. Employing theoretical analyses, mathematical proofs, Monte Carlo simulation, and real-data examples, jiang2021total demonstrated that the total-effect test is indeed superfluous for establishing complementary mediation when the paths are estimated by the least square estimators and tested by $F$- or Sobel test.
We are to extend the advancement to the other two types. The geometric analysis developed by jiang2021total potentially can be applied to build a mathematical framework for analyzing mediation of all types. The potential, however, has yet to be realized. This study is to fill the gap. Section 3 reviews the key elements of the geometric analysis, and extends the analysis to establishing complementary mediation under the LSE-Sobel frameworks. Sections 4 and 5 utilize the geometric approach to analyze indirect-only and competitive mediation under LSE-$F$ and LSE-Sobel frameworks. The results are validated numerically in Section 6 via simulation. Section 7 applies the main results to two real-data examples. Section 8 summarizes and discusses the main findings.
In data analysis, researchers use hypothesis tests to determine whether the direct, indirect, and total effects each pass the partition threshold kenny1998data, zhao2010reconsidering. Denoting the rejection regions of appropriate hypothesis tests with significance level $\alpha$ as $\mathcal{R}_{a\times b}(\alpha)$, $\mathcal{R}_d(\alpha)$, and $\mathcal{R}_c(\alpha)$, and using estimators $(\hat a, \hat b, \hat d, \hat c)$, we can claim a statistically significant causal effect based on the sign of its estimator. For example, a positive direct causal effect is claimed when $\hat d > 0$ and the observed data fall within $\mathcal{R}_d(\alpha)$. Similar definitions apply to the total and indirect effects.
The process-and-product approach (PAPA) defines three types of mediation, namely complementary ($\bm C_+$), indirect-only ($\bm C_0$), and competitive ($\bm C_-$), at a given significance level $\alpha$ jiang2021total, Liu2023Electronic, zhao2010reconsidering,zhao2011does. The causal steps approach, however, requires also a statistically significant total effect (c) to claim any type of mediation baron1986. The two sets of requirements together would imply the following:
where $\mathcal{D}$ represents the observed data, $\bar\mathcal{R}$ stands for the complementary set of a rejection region $\mathcal{R}$. If the total-effect test fails, the mediation fails to validate. However, the PAPA analysts propose alternative criteria without considering the total effect $c$:
Therefore, the need for the total-effect test in establishing mediation can be rationalized by examining the geometric relationships of the corresponding rejection regions. Specifically, if we demonstrate that
we establish $\bm C_+\Leftrightarrow\bm C_+^*$, indicating that the total-effect test is unnecessary for complementary mediation. Similarly, if we show that
we have $\bm C_0\nLeftrightarrow\bm C_0^*$, suggesting that the total effect test may lead to misleading results and erroneously reject indirect-only mediation when we consider criterion $\bm C^*_0$ to be more appropriate. Likewise, if we demonstrate that
we have $\bm C_-\nLeftrightarrow\bm C_-^*$, implying that the total effect test may incorrectly reject competitive mediation. In the following sections, we will provide a detailed implementation of the above geometric analysis. {
The LSE-$F$ framework proposed by judd1981 estimates the parameters $(a, b, d, c)$ using least squares estimators (LSEs) and tests their significance using $F$-tests. The following equations define the LSEs:
where $\mathcal{D}_{1,X} = (\bm 1, \bm X)$ and $\mathcal{D}_{1,M,X} = (\bm 1, \bm M, \bm X)$ are data matrices with or without the mediator, respectively. The rejection regions for the $F$-tests on $(a, b, d, c)$ with significance level $\alpha$ are defined as follows:
where $\bm Y=(Y_1,\ldots,Y_n)^T$, $\bm Y_{\bm X}$ represents the projection of $\bm Y$ onto the linear space spanned by $\bm X$, and $\lambda_{t,s}(\alpha)$ is the $\alpha$-quantile of $F$-distribution with degrees of freedom $(t,s)$. Additionally, to claim the significance of the indirect effect $a\times b$, we reject the hypothesis test when both $a$ and $b$ are rejected, i.e., $\mathcal{R}_{a\times b}(\alpha)$ is replaced with $\mathcal{R}_a(\alpha)\cap\mathcal{R}_b(\alpha)$.
Alternatively, baron1986 suggested the LSE-Sobel framework, which is similar to the LSE-$F$ framework but uses the Sobel test to examine the indirect effect $a \times b$. The test statistic $S$ is defined as:
Under the null hypothesis of $a\times b=0$, $S$ asymptotically follows a standard normal distribution. The rejection region for the Sobel test is given by:
where $z_\alpha$ represents the $\alpha$-quantile of standard normal distribution. The LSE-Sobel framework provides a direct inference of the indirect effect $a \times b$ using a single test, but is limited by not being an exact test as the exact distribution of the test statistic $S$ depends on the values of $a$ and $b$.
The LSE-based frameworks assume normal distribution for the noise terms $\varepsilon_M$ and $\varepsilon_Y$. In case of non-normality, an alternative is the LAD-$Z$ framework pollard1991. It uses the test statistic $z = |\check{\beta}|/sd(\check{\beta})$ compared to the standard normal distribution. Here, $\check{\beta}$ is the least absolute deviance (LAD) estimator of $\beta$ for $\beta \in \{a,b,d,c\}$, and $sd(\check{\beta})$ is the estimated standard deviation of $\check{\beta}$.
The LSE estimators $(\hat{a},\hat{b},\hat{d},\hat{c})$ and the corresponding rejection regions in the LSE-$F$ and LSE-Sobel frameworks involve complex components. However, we have discovered a simpler mathematical formulation by properly transforming the original data matrix, inspired by jiang2021total.
Lemma 1 in jiang2021total demonstrated that the LSE estimators $(\hat{a},\hat{b},\hat{d},\hat{c})$ and the rejection regions for $F$-tests are invariant to scale and orthogonal transformations of the observed data. The following lemma extends the original lemma in jiang2021total by including the invariance of the Sobel test for $a\times b$, and its proof can be found in Section 1 (S1) of the Supplementary Material.
The above lemma suggests that we can transform the original data matrix to obtain simpler rejection regions. As highlighted by jiang2021total, for a classic mediation model with the data matrix $\mathcal{D}$ with rank$(\mathcal{D}) = 4$, there always exists an $n \times n$ real orthogonal matrix $\Gamma$ and a global scale parameter $\gamma > 0$ s.t. the transformed data matrix $\Tilde{\mathcal{D}}= (\Tilde{\bm 1}, \Tilde{\bm M},\Tilde{\bm X},\Tilde{\bm Y}) $ satisfies $\Tilde{\bm 1} = (1,0,\ldots,0)^T$, $\Tilde{\bm X} = (x_1,x_2,0,\ldots,0)^T$, $\Tilde{\bm M} = (m_1,m_2,m_3,0,\ldots,0)^T$, $\Tilde{\bm Y} = (y_1,y_2,y_3,y_4,0,\ldots,0)^T$ with $x_2 > 0$, $m_3 > 0$ and $y_4>0$.
Apparently, the transformed data $\Tilde{\mathcal{D}}$ simplifies the LSE estimators and rejection regions. The following lemma summarizes the results, including the explicit form of $\mathcal{R}_{a\times b}$, which was not provided in jiang2021total.
Using the simplified formulas in Lemma (ref), jiang2021total showed that $\mathcal{R}_{a}(\alpha)\cap\mathcal{R}_{b}(\alpha)\cap\mathcal{R}_d(\alpha)\subseteq \mathcal{R}_c(\alpha)\ \mbox{whenever}\ \hat a\times\hat b\times\hat d>0$ under mild conditions. This implies that total-effect test is superfluous for establishing complementary mediation under the LSE-$F$ framework. Additionally, they showed that the total-effect test is also unnecessary asymptotically under the LSE-Sobel framework. However, their analysis within the LSE-Sobel framework lacks a geometric perspective, and can be enhanced by the following theorem, with the proof detailed in Section 1 (S1) of the Supplementary Material.
Theorem (ref) implies that as the sample size $n \rightarrow \infty$, $ \mathcal{R}_{a\times b}(\alpha) \cap \mathcal{R}_{d}(\alpha)\subseteq {\mathcal{R}}_{c}(\alpha)$ holds asymptotically. This provides an alternative perspective supporting the argument that the total-effect test is superfluous for testing complementary mediation under the LSE-Sobel framework.
jiang2021total mentioned that the total-effect test can be misleading when testing indirect-only mediation under the LSE-$F$ framework. However, no technical proof was provided to support the observation, and the observation does not cover the LSE-Sobel framework. In this section, we will fill these gaps through explicit theoretical analysis that shows the potentially misleading nature of the total-effect test for establishing indirect-only mediation.
To show that the total-effect test may erroneously reject indirect-only mediation under the LSE-$F$ framework, we need to verify condition (ref): $$\mathcal{R}_a(\alpha) \cap \mathcal{R}_b(\alpha) \cap \Bar{\mathcal{R}}_d(\alpha) \cap \Bar{\mathcal{R}}_c(\alpha) \not= \emptyset,\ \text{for all } \alpha \in (0,1).$$ This condition can be equivalently expressed as:
where $\mathcal{R}_\beta(\alpha|r)$ represents the intersection of $\mathcal{R}_{\beta}(\alpha)$ and the $p$-$q$ plane $\mathcal{P}_r$ for $\beta \in \{a,b,d,c\}$. Since $\mathcal{R}_a(\alpha\mid r) = \mathcal{P}_r \cap \{r>r_{n,\alpha}\} = \emptyset$ for $0 < r \leq r_{n,\alpha}$, we will focus on $r > r_{n,\alpha}$. We verify the argument by considering two sub-types of indirect-only mediation separately: the case where $\hat{a}\hat{b}\hat{d}> 0$, representing directionally complementary mediation, and the case where $\hat{a}\hat{b}\hat{d} < 0$, representing directionally competitive mediation.
In the case where $\hat{a}\hat{b}\hat{d} > 0$, we can observe the following relationships. Corollary 1 in jiang2021total implies that $\hat{a}\hat{b}\hat{c} > 0$ and $q > rp$. Thus, we have $\Bar{\mathcal{R}}_d(\alpha\mid r)=\left\{rp < q \leq rp+ p_{n,\alpha}(r^2+1)^{1/2}\right\}$. Additionally, $\mathcal{R}_b(\alpha\mid r) = \{p>p_{n,\alpha}\}$, and $\Bar{\mathcal{R}}_c(\alpha\mid r)=\left\{0 \leq q \leq r_{n,\alpha}(p^2+1)^{1/2},p \geq 0\right\}$. The geometry of these regions is depicted in Figure (ref). According to Theorem 1 in jiang2021total, we have $r_{n,\alpha}(p^2+1)^{1/2} < rp + p_{n,\alpha}(r^2+1)^{1/2}$ for $r > r_{n,\alpha}$. Therefore, the intersection $\mathcal{R}_a(\alpha\mid r) \cap \mathcal{R}_b(\alpha\mid r) \cap \Bar{\mathcal{R}}_d(\alpha\mid r) \cap \Bar{\mathcal{R}}_c(\alpha\mid r)$ is $$\left\{rp < q \leq r_{n,\alpha}(p^2+1)^{1/2},p_{n,\alpha} < p < r_{n,\alpha}/(r^2-r_{n,\alpha}^2)^{1/2}\right\},$$ which can be verified to be not empty for $r_{n,\alpha} < r < r_{n,\alpha}(1+1/p_{n,\alpha}^2)^{1/2}$. Figure (ref) (D) provides a graphical demonstration of this intersection.
In the case where $\hat{a}\hat{b}\hat{d} < 0$, the sign of $\hat{a}\hat{b}\hat{c}$ is indeterminate. The regions $\mathcal{R}_a(\alpha\mid r)$, $\mathcal{R}_b(\alpha\mid r)$, and $\Bar{\mathcal{R}}_c(\alpha\mid r)$ remain the same as in the case where $\hat{a}\hat{b}\hat{d} > 0$. The only difference lies in the region $\Bar{\mathcal{R}}_d(\alpha\mid r)$. The following lemma helps define ${\mathcal{R}}_d(\alpha\mid r)$ when $\hat{a}\hat{b}\hat{c} \geq 0$.
Using Lemma (ref), It can be verified that \[
\] and the intersection of $\mathcal{R}_b(\alpha\mid r)$, $\Bar{\mathcal{R}}_c(\alpha\mid r)$ and $\Bar{\mathcal{R}}_d(\alpha\mid r)$ is not empty for any $r>r_{n,\alpha}$. The geometry of $\Bar{\mathcal{R}}_d(\alpha\mid r)$ and the intersection of interest under the directionally competitive mediation are shown in Figure (ref).
Above all, we validate the argument (ref). Figures (ref) and (ref) imply that $\mathcal{R}_a(\alpha\mid r) \cap \mathcal{R}_b(\alpha\mid r) \cap \Bar{\mathcal{R}}_d(\alpha\mid r) \cap \Bar{\mathcal{R}}_c(\alpha\mid r)$ has larger support under directionally competitive mediation, indicating a higher probability of observing an insignificant total effect when $\hat{a}\hat{b}$ and $\hat{d}$ have opposite signs and $\hat{d}$ is not statistically significant.
{ The following theorem summarizes the above analysis and depicts the rejection regions based on a nice geometry under a mild condition.
}
Similarly, the following Theorem shows that the total-effect test can be erroneous for establishing indirect-only mediation under the LSE-Sobel framework { with large sample size}. The details of proof can be found in {Section S1 of the Supplementary Material}.
Similar analyses can be applied to competitive mediation. The following theorems demonstrate that the total-effect test can lead to erroneously rejection of competitive mediation under LSE-$F$ and LSE-Sobel frameworks, respectively, for statistical partitioning. While previous studies have shown the possibility of erroneous rejection using bootstrap tests mackinnon2000equivalence, mackinnon2007mediation, hayes2009beyond, zhao2010reconsidering through derivations and examples, the theorems below provide rigorous proofs assuming LSE-$F$ and LSE-Sobel tests.
The procedure for proving the indirect-only mediation was applied to proving the two theorems above, of which the details are documented in Section 1 (S1) of the Supplementary Material.
jiang2021total conducted simulations to show that the total-effect test is unnecessary for establishing complementary mediation under LSE-$F$ and LSE-Sobel frameworks and can be misleading with the LAD-$Z$ test. The study focused on indirect-only and competitive mediation, presenting the results for indirect-only mediation in the main text and the results for competitive mediation in Section 2 (S2) of the Supplementary Material.
To validate Theorem (ref), we generate the simulated data from model (ref) and (ref) as follows: \[
\] A total of $10,000$ independent datasets of different sample sizes were simulated. For each dataset, we calculated the LSEs $(\hat{a},\hat{b},\hat{c},\hat{d})$ and $p$-values $(p_a,p_b,p_c,p_d)$ under the LSE-$F$ framework. We checked if, for any fixed $\alpha \in (0,1)$, $\max(p_a,p_b) < \alpha$ and $p_d \geq \alpha$ imply $\{p_c \geq \alpha\} \not= \emptyset$.
Figure (ref) (A) checks the $p$-value condition when $\alpha=0.1$. Each simulated dataset is represented by a point in a 3-dimensional space with $\max(p_a,p_b)$, $p_d$, and $p_c$ as the $X$, $Y$, and $Z$ axes, respectively. The solid circles represents datasets satisfying $\max(p_a,p_b) < \alpha$ and $p_d \geq \alpha$, gray crossings represent data sets with $\max(p_a,p_b) \geq \alpha$ or $p_d < \alpha$, and the dark gray dashed plane represents $p_c = \alpha$. The solid circles above the plane $p_c = \alpha$ indicate the empirical set $\{p_c \geq \alpha\}$. We observe that when $\max(p_a,p_b) < \alpha = 0.1$ and $p_d \geq \alpha$, the set $\{p_c \geq \alpha\}$ is not empty.
Figure (ref) (B) presents the proportion of datasets satisfying $p_c \geq \alpha$ for 1000 evenly spaced values of $\alpha$ in the range of $(0.01,0.99)$ when $\max(p_a,p_b) < \alpha$ and $p_d \geq \alpha$ under the LSE-$F$ framework. It shows that for significance levels $\alpha$ smaller than 0.1, which is commonly used in practice, the proportion of cases where $p_c > \alpha$ is greater than 40%. This indicates a high probability of erroneous rejection of indirect-only mediation by the total-effect test.
The proportion of erroneous total-effect test results for both directionally complementary and directionally competitive indirect-only mediation cases is depicted in Figure (ref). The plot demonstrates that the total-effect test can lead to incorrect conclusions regarding the presence of indirect-only mediation in both cases. Interestingly, the erroneous judgments are more frequent when the signs of $\hat{a}\hat{b}$ and $\hat{d}$ are opposite, which is in line with expectations established in the analysis of Theorem (ref).
To investigate whether a similar result holds for other frameworks in establishing indirect-only mediation, we conducted a similar analysis using the LSE-Sobel framework and LAD-$Z$ framework with the same set of simulated datasets. Under the LSE-Sobel framework, we additionally calculated the the p-value $p_{ab}$ for the Sobel test of $a\times b$. If a similar result holds, we could expect to see $\{p_c \geq \alpha\} \not= \emptyset$ for any fixed $\alpha \in (0,1)$ when $p_{ab} < \alpha$ and $p_d \geq \alpha$, which is supported by Figures (ref) and (ref). Moreover, graphical verification of results under the LAD-$Z$ framework are shown in Figures (ref) and (ref), supporting similar conclusions.
We illustrate the conclusions of the mathematical derivation and simulation presented above with two real-data examples. The example data came from the Health Information National Trends Survey (HINTS, \url{http://hints.cancer.gov/}), which is conducted regularly by the National Cancer Institute on representative samples of United States adults to track changes in health behavior and communication jiang2020digital,liu2023communication,finney2020data. This study used the 2020 version of the postal-mail survey (HINTS 5 Cycle 4) with 3,865 participants.
Two models are presented below. Model 1 is a directionally competitive indirect-only (d-petitive IO) mediation depicting an effect of caregiving (CG) on smoking (SM) through psychological distress (PD) (Figure (ref)). Model 2 is a directionally complementary indirect-only (d-plementary IO) mediation describing the effect of employment (EM) on physical activity (PA) through psychological distress (PD) (Figure (ref)). In each model, the mediated effect passed the statistical threshold ($p<.05$) while the total effect failed ($p\geq .05$).
In both examples, the total-effect ($c$) test would have concluded there was no “effect to be mediated", which would be equal to “no mediated effect" in the causal-steps doctrine, thereby requiring a full-stop of all further analysis, when the mediated effect ($ab$) would have passed the statistical test. Thus, each model is a real-data example of the total-effect test erroneously rejecting indirect-only mediation.
The two example models fit the definition of “full mediation" aka “complete mediation" under the quasi typology of {full, partial, and no mediation} baron1986. The terms were meant to connote the strongest form of mediation Hayes2022IntroMediation. Hopefully, the total-effect test erroneously rejecting the perceived strongest mediation demonstrates the pitfalls of the total-effect test and the pitfalls of the causal-steps approach.
When presenting the examples, we employ process-and-product analysis (PAPA) emerging in several disciplines jiang2021total,liu2023effect, Peng20203rdperson, Liu2023Electronic,zhao2014emerging,zhao1994media,zhao2010reconsidering. PAPA approach sees mediation as a process and total effect (c) as the product of the process. While the causal-steps approach is focused on one task, which is to “establish mediation", PAPA is given three tasks, 1) testing effect hypotheses, 2) classifying effect types, and 3) analyzing effect sizes, all for the ultimate mission of better understanding the relationship between parts, process, and product.
To estimate effect sizes, PAPA employs the percentage coefficient ($b_p$), the regression coefficient with the dependent variable (DV) and independent variable (IV) both on a $0\sim 1$ percentage scale ($p_s$) jiang2021total,liu2022effects,zhao2016enough, Liu2023Electronic. Scaled such, $b_p$ indicates the percentage change in DV associated with a 100$\%$ whole-scale increase in IV or, in other words, the change in DV measured by the percentage of a point associated with an increase in IV by one percentage point. Thus, $b_p$ is interpretable and comparable assuming scale equivalence zhao2014emerging,jiang2021total. The two features make $b_p$ a generic indicator of effect sizes that is easy to interpret and efficient to compare. Table (ref) provides scale details and univariate descriptions of variables, and Eq. 1 of Table (ref) is the formula for percentizing the scales.
To help dissect the product, discern the parts and divine the process, PAPA calculates percent contribution ($c_p$), the contribution of each elemental part, i.e., the $a$, $b$, $ab$, $d$, or $c$ path, to the $IV \rightarrow DV$ total effect, $c$, as detailed in Table (ref). To reduce overuse and misuse of $p$ values, we strive to practice what we consider the best practice, 1) refraining from the term “statistical significance” and “statistical non-significance”; 2) referring to $p<.05$ as “statistical acknowledgment”, benjamini2021asa,nature2019s,siegfried2015p,wilkinson1999statistical and 3) referring to $p\geq .05$ as “statistical inconclusiveness”. Such practices indicate that $p<.05$ is merely a pretest yardstick or partition threshold under the principles of functionalism, passing which would allow for classifying effect types and analyzing effect sizes liu2023covid,zhao2022interrater,liu2023effect. It's hoped that such practices, including the application of effect-size indicators such as $b_p$ and $c_p$, benefit from and contribute to the “effect size movement” Kelley2012effectsize,preacher2011effect, wilkinson1999statistical, robinson2002effect,jiang2021total,Schmidt1996significance.
Competitive mediation, aka suppression or inconsistent mediation mackinnon2000equivalence,mackinnon2007mediation, defined as a model with statistically acknowledged $ab$ and $d$ paths at the opposite directions, is widely known as the type of mediation that can be erroneously rejected by the total-effect test xiao2018social,busse2016abc,gopalakrishnan2019client. The mathematical derivation above has provided proof that the total-effect test can also erroneously reject indirect-only (IO) mediation, which includes two subtypes, directionally competitive (d-petitive) and directionally complementary (d-plementary). The following is a real-data example for the first subtype, D-petitive IO mediation, defined as a model with a statistically acknowledged $ab$ path and a statistically inconclusive $d$ path in the opposite direction. The example shows that this model of mediation is erroneously rejected by the total-effect test.
\noindentKey Variables.\\ \noindentDependent variable: Smoking frequency (SF) was measured by four items that asked responders how often they smoke cigarettes and e-cigarettes. The composite variable ranged from 0 to 1 where 0 represents not smoking and 1 represents smoking every day.\\ \noindentMediating variable: Psychological distress (PD) was the sum of four items (Cronbach’s $\alpha= .871$) measuring the frequency by which the respondents experienced four symptoms of psychological distress in the past two weeks, feeling little interest in doing things, being emotionally down, hopeless, and anxious. Again, it was transferred into a $0 \sim 1$ where 0 means never feeling any symptoms and 1 means feeling four symptoms every day. \\ \noindentIndependent variable: \textit{Caregiving (CG)} was the sum of five items, measuring whether the respondents took five types of caregiving currently. The caregiving types include caring for children, partners, parents, relatives, and friends. The compound variable CG ranges $0 \sim 1$ where 0 indicates no caregiving responsibilities and 1 indicates the respondent needs to take all types of caregiving.
\noindentControl variables: Age, income, and education were reported in Table (ref) and included as control variables for the mediation analysis. To simplify the presentation, the control variables were omitted from Table (ref) and Figures (ref) $\&$ (ref) that report the outcomes.
\noindentMediation analysis.\\ Table (ref) and Figure (ref) summarize Model 1 findings. They show a positive and statistically acknowledged indirect effect, and a negative but statistically inconclusive direct-and-remainder (di-remainder) effect, producing a directionally competitive indirect-only (d-petitive IO) mediation. Figure (ref) reports the contribution of each path. The $ab$ path ($b_p=.0165\%$, $c_p\approx 117,857\%$) and $d$ path ($b_p=-.0167$, $c_p\approx -119,286\%$) contributed about equal percentages in opposite directions, leading to the nearly-zero total effect ($c$ path, $b_p=.000014$). Because the competition between $ab$ and $d$ paths was about equal (118K$\%$ v.s. -119K$\%$), they offset each other to produce a near-zero total effect ($b_p=.000014$). Consequently, the contribution of each part appeared huge percentage-wise. In such cases of small total effect due to even competition, the comparative sizes may be more important than the sizes themselves.
The indirect path ($ab$) passed the statistical threshold ($p<.001$) while the direct-and-remainder ($d$) path failed ($p=.6411$), making it a directionally competitive indirect-only (d-petitive IO) mediation by our standards zhao2011does,zhao2010reconsidering,hayes2009beyond,rucker2011mediation. Nevertheless, the total effect ($c$) failed ($p=.9997$). If passing the total-effect test remains a necessary condition for establishing mediation as some experts continued to prescribe, this model would have been disqualified as mediation wen2004testing,wen2014analyses,rose2004mediator.
In the above example, the competition between the indirect ($ab$) and the direct-and-remainder ($d$) paths was clearly a main contributor to the near-zero total effect ($c$) and the statistical inconclusiveness. The following is a real-data example that a second subtype, a directionally complementary indirect-only (d-plementary IO) mediation, is erroneously rejected by the total-effect test, thereby showing that the total-effect test can erroneously reject mediation without competition.
\noindentKey Variables.\\ \noindentDependent variable: Physical activity (PA) was measured by two items that asked responders how many minutes per day and how many days per week they usually did physical activity xie2020electronic,kontos2014predictors. The two items were multiplied to compute the weekly physical activity the respondents conducted. The composite variable was then linearly transformed to a 0-1 percentage scale where 1 represents the highest weekly physical activity and 0 represents not conducting weekly physical activity at all.
\noindentMediating variable: Psychological distress (PD) in Model 2 was the same mediating variable in Model 1.
\noindentIndependent variable: Employment (EM) measured whether the respondents were employed or not, e.g., unemployed, retired, or being students, recoded 1 for employed and 0 for not employed.
\noindentControl variable: Age, gender, and education were controlled in mediation analysis but omitted from Table (ref) and Figures (ref) $\&$ (ref) that report the outcomes.
\noindentMediation analysis.\\ Model 2 of Table (ref) show that the indirect ($ab$) path was positive and statistically acknowledged, while the dire-mainder ($d$) path was also positive but failed to pass the statistical threshold ($p=.2704$), making it a directionally complementary indirect-only (d-plementary IO) mediation. The total effect ($c$) also failed the statistical test ($p=.1169$), demonstrating that total-effect test can also erroneously reject this subtype of mediation. This could be among the first documented real-data examples that the total-effect test erroneously rejects mediation without competition, aka suppression.
A question arises: Given that the indirect path was statistically acknowledged, the direct-and-remainder path complemented, and the total effect was the sum of the two paths, what caused the statistical inconclusiveness of the total effect?
The mathematical derivations provided above would point fingers at the estimated variance (SE) of the $d$ path. That is, for this subtype, the large variance of the $d$ path relative to the other parameters should be considered the largest factor contributing to the larger-than-threshold $p$-value of the total effect ($c$). The process-and-product analysis (PAPA) of the real-data example provides a non-mathematical illustration of the point.
As shown in Figure (ref), even though the $d$ path was statistically inconclusive ($p=.2704$), the estimated effect size of the $d$ path ($b_p=.0243$) more than doubled that of the statistically acknowledged $ab$ path ($b_p=.0102$, $p<.001$). The contribution of the d path accounted for 71$\%$ of the total effect ($c_p=71\%$). The effect size of the $d$ path is not small relative to the other main parameters. Rather, it was the relatively large variance of the $d$ path (SE = .022) that was a main contributor to the variance of the $c$ path (SE = .0218), which was, in turn, a main factor contributing to the statistical inconclusiveness of the total effect ($p=.1169$).
The previous sections proved a mathematical theorem and provided Monte Carlo simulations showing that the total-effect test can erroneously reject indirect-only (IO) mediation. This section added two real-data examples documenting that the test did erroneously reject this type of mediation. Two models were shown, recording erroneous rejection for each of the two major subtypes of IO mediation. Model 1 is a case where the total-effect test erroneously rejected directionally competitive indirect-only (d-petitive IO) mediation. Model 2 is a case where the total-effect test erroneously rejected directionally complementary (d-plementary IO) mediation.
As the d-plementary IO model involves no competition, Model 2 shows that the erroneous rejection can occur without competition, statistical or directional. The finding has implications for data analysts who struggle to interpret d-plementary IO models with statistically inconclusive total effects. It also illustrates an implication of Theorem (ref): The estimation variance, i.e., standard error, of direct-and-remainder ($d$) path tends to be large relative to other parameters of this subtype; the estimation variance of the direct-and-remainder path may be the more significant factor than effect sizes and other parameters contributing to the statistical inconclusiveness of the total effect ($c$).
An implication is that practicing researchers may need to focus more on effect sizes than on p-values or confidence intervals. For example, in a d-plementary IO model, the statistically inconclusive $p$-value for the $c$ path is attributed more to the large variance in the $d$ path. While the $d$ portion of the $c$ path can not be acknowledged, the direction of the $c$ path still can, due mainly to the direction and strength of the $ab$ path. The direction and strength of the effect are often theoretically and practically as important as, if not more important than, the variance of the effect.
This study provided a mathematical theorem, a {Monte Carlo} simulation, and two real-data examples to demonstrate that the total-effect test can and did erroneously reject indirect-only mediation. There are three and only three types of mediation, competitive, complementary, and indirect-only busse2016abc,jiang2021total,zhao2011does,zhao2010reconsidering. While prior studies have shown that the total-effect test is superfluous for establishing complementary mediation, this study shows that the test can erroneously reject indirect-only and competitive mediation. Thus, this study completes the argument that the total-effect test should be recanted, but not just relaxed or suspended, for establishing any type of mediation, to the extent that the traditional ordinary least square (LSE-$F$ or LSE-Sobel) procedures were applied to calculate $p$-values or confidence intervals for statistical tests. A similar conclusion can be reached under the LAD-$Z$ framework, which is verified with simulation studies.
Table (ref) displays the repercussions of imposing the total-effect test for establishing mediation. Altogether three types of mediation and the three types of tests make up the nine cells. The table shows two possible types of outcomes, i.e., repercussions, to be erroneous or superfluous. The total-effect test is erroneous in seven of the nine situations and is superfluous in the other two situations.
The table also lists the studies that have contributed to the knowledge and influenced this study the most directly. Of all the nine cells, “this study" is the only entry in three (C2, B3, C3), indicating it is the only contributor. “This study" also appears in two other cells (B1 and C1). If we count the nine cells as nine pieces of knowledge, this study emerges as the sole or main contributor to five of the nine pieces regarding the repercussions of the test.
Now, this study has provided proof that the total-effect test can produce erroneous outcomes, not only for competitive mediation but also for indirect-only mediation, both directionally competitive and complementary. Every cell of Table (ref) has been filled. The total-effect test harms or fails to help, whether LSE-$F$, LSE-Sobel, or LAD-$Z$ test is used. The burden of proof is now on the causal-step procedure to show that the total-effect test is harmless and helpful for establishing any type of mediation.
At the meantime, it might be appropriate to consider recanting the total-effect test for establishing mediation of all types, rather than imposing the test and then suspending or relaxing it for one special type. This means completely removing the test from the regular procedure instead of allowing exceptions in the cases of anticipated suppression or inconsistency, aka competitive mediation baron1986,kenny1998data,kenny2008reflections,kenny2021mediation,rose2004mediator,wen2004testing,wen2014analyses.
Admittedly, proofs assuming bootstrap tests are not yet available. But the available evidence, especially if adding this study, is overwhelmingly against the test. Given the striking imbalance between pro and con, it's time to set aside the total-effect test for establishing mediation unless and until new evidence emerges to show a benefit.
Why and how did the total-effect test dominate so many disciplines for so long baron1986? Over-dichotomization of the effect concept and oversimplification of the mediation concept may be among the root causes. In light of the revelation, it may be time to revisit and possibly revamp the objectives, concepts, theories, and techniques of mediation analysis. It may be time also to consider more broadly causal dissection models, which include moderation and curvilinearity in addition to mediation zhao2017MainEffect. Accordingly, Section 6 above are tasked to showcase two applications of process-and-product analysis (PAPA), whose missions are 1) testing the presence of mediation, 2) identifying types of mediation, and 3) analyzing sizes of the mediation and non-mediation effects jiang2021total,liu2022effects,Liu2023Electronic. Establishing mediation is a part, and only a part, of the first mission of PAPA.
While this study completes the argument about recanting total-effect test for establishing mediation using OLS tests, i.e., LSE-$F$, LSE-Sobel, and LAD-$Z$, future research may falsify or qualify the argument by testing the main theses assuming non-OLS tests such as bootstraps zhao2011does. The more challenging task is to overcome the spiral of inertia, resist over-dichotomization and oversimplification of fundamental concepts, and foster an understanding of mediation that is more analytical and comprehensive. It will take time and luck, but above all persistence. (c.f. zhao2022interrater; zhao2018WeAgreed)
Supplementary Material is attached at the end of this document to provide the detailed proof for Lemma (ref), (ref), (ref), Theorem (ref), (ref), (ref) and (ref) and results of simulation for competitive mediation under LSE-$F$, LSE-Sobel, and LAD-$Z$ frameworks.
This research was supported by the Beijing Natural Science Foundation [Z190021]; National Natural Science Foundation of China [Grant 11931001]; grants of University of Macau, including CRG2021-00002-ICI, ICI-RTO-0010-2021, CPG2021-00028-FSS and SRG2018-00143-FSS, ZXS PI; and Macau Higher Education Fund, HSS-UMAC-2020-02, ZXS PI. Xinshu Zhao and Ke Deng are co-corresponding authors.
\bibhang=1.7pc \bibsep=2pt \fontsize{9}{14pt plus.8pt minus .6pt}\selectfont \expandafter\ifx\csname natexlab\endcsname\relax\def\natexlab#1{#1}\fi \expandafter\ifx\csname url\endcsname\relax \def\url#1{#1}\fi \expandafter\ifx\csname urlprefix\endcsname\relax\fi
\vskip .65cm Tingxuan Han\\ Center for Statistical Science & Department of Industry Engineering, Tsinghua University, Haidian, Beijing, China. \vskip 2pt E-mail: [email removed] \vskip 2pt
Luxi Zhang \vskip 2pt Department of Communication, University of Macau, Macau, China. \vskip 2pt E-mail: [email removed]
Xinshu Zhao \vskip 2pt Department of Communication, University of Macau, Macau, China. \vskip 2pt E-mail: [email removed]
Ke Deng\\ Center for Statistical Science & Department of Industry Engineering, Tsinghua University, Haidian, Beijing, China. \vskip 2pt E-mail: [email removed]