Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
128,909 characters · 10 sections · 90 citation commands
Testing Mechanisms
Social scientists are often able to identify the causal effect of a treatment $D$ on some outcome of interest $Y$, either by explicitly randomizing $D$ or using some “quasi-experimental” variation in $D$. Once the causal effect of $D$ on $Y$ is established, a natural question is why does it work, i.e. what are the mechanisms by which $D$ affects $Y$?
To fix ideas, consider the setting of bursztyn_misperceived_2020, which will be one of our empirical applications below. The authors show that the vast majority of men in Saudi Arabia underestimate how open other men are to women working outside of the home. They then run an experiment in which some men are randomized to receive information about other men's beliefs. At the end of the experiment, all of the men are given the choice between signing their wives up for a job-search service or taking a gift card. The authors observe that the information treatment increases the probability that men sign their wives up for the job-search service, and also increases the probability that their wives apply for and interview for jobs over the subsequent five months. A natural question in interpreting these results is then whether the increase in longer-run outcomes (e.g. job applications) is explained by the short-run sign-up for the job-search service, or whether the treatment also affects labor market outcomes through other longer-run changes in behaviors.
The literature on mediation analysis (see huber_review_2019 for a review) provides formal methodology for disentangling how much of the average effect of a treatment $D$ (e.g. information about others' beliefs) on an outcome $Y$ (e.g. job applications) is explained by the indirect effect through some potential mediator $M$ (e.g. job-search service sign-up). A challenge, however, is that even if the treatment $D$ is randomly assigned, it will often be the case that the mediator of interest $M$ is not randomly assigned.\footnote{One exception is “mechanism experiments” ludwig_mechanism_2011, where the researcher explicitly randomizes an $M$ of interest. Our focus is on the common setting where $M$ was not randomized (e.g. due to lack of foresight, budget, or feasibility of randomization) and thus potentially endogenous.} Existing approaches typically make strong assumptions that allow for the identification of the causal effect of $M$ on $Y$ (see Related Literature below). A common assumption in the biostatistics literature, for example, is that $M$ is as good as randomly assigned given $D$ and some observable characteristics. This assumption will often be restrictive in applications---for example, we may worry that sign-up for the job-search service is correlated with unobservables related to women's labor supply.
In this paper, we develop methodology that sheds light on mechanisms without having to impose strong assumptions to identify the effect of $M$ on $Y$. We make progress by considering an easier question than what is typically studied in the literature on mediation analysis, but one that we think will still be informative in many applications. Rather than trying to identify how much of the average effect is explained by the indirect effect through $M$, we start by testing what we refer to as the sharp null of full mediation: is the effect of $D$ on $Y$ fully explained through its effect on $M$? In our motivating application, the sharp null asks whether the effect of treatment on job applications is fully explained by the short-run take-up of the job-search service. More precisely, letting $Y(d,m)$ be the potential outcome as a function of treatment $d$ and mediator $m$, the sharp null posits that $Y(d,m)$ depends only on $m$. If we can reject this null in our motivating example, then we can conclude that the treatment affects long-run outcomes through some change in behavior other than job-search service sign-up. In addition to testing the sharp null, our approach also provides useful information about the extent to which the null is violated. In particular, we develop lower bounds on the fraction of individuals whose outcome is affected by the treatment despite having the same value of $M$ under both treatments. In our motivating example, this means we can lower bound the fraction of women whose labor market outcomes are affected by the information treatment despite the treatment having no effect on whether they sign up for the job service.
Our main theoretical results impose two assumptions. First, we suppose that the treatment $D$ is as good as randomly assigned, i.e. $D$ is independent of the potential outcomes $Y(\cdot,\cdot)$ and potential mediators $M(\cdot)$. In our motivating example, this is guaranteed by design since $D$ is randomly assigned. (We consider extensions to “quasi-experimental” settings in (ref).) Second, we allow the researcher to impose restrictions on how the mediator $M$ responds to treatment. A leading example is the monotonicity assumption that the potential mediator $M(d)$ is increasing in $d$. In our motivating application, this imposes that providing men with information that other men are more open to women working outside of the home can only increase whether they sign up for the job-search service (in our main analysis, we restrict attention to the majority of men who initially underestimate others' openness, so the information plausibly updates beliefs in a common direction). We first consider the setting where monotonicity holds, and then introduce a more general framework that allows the researcher to impose arbitrary restrictions on the distribution of $(M(0),M(1))$, which nests monotonicity and relaxations thereof as special cases.
A key observation is that under the sharp null of full mediation and the independence and monotonicity assumptions just described, the treatment $D$ is a valid instrumental variable for the local average treatment effect (LATE) of $M$ on $Y$. In the case of binary $D$ and binary $M$, the LATE assumptions are known to have testable implications balke_bounds_1997,kitagawa_test_2015, huber_testing_2015,mourifie_testing_2017. Existing tools for testing the LATE assumptions can thus be used “off-the-shelf” for testing the sharp null of full mediation when $D$ and $M$ are binary, as we describe in more detail in (ref). In our motivating example, the testable implications of the sharp null appear to be violated (significant at the 5% level), and thus we can conclude that the effect of the information treatment does not operate entirely through job-search service sign-up.
While existing tools can be used to test the sharp null in the case of a binary mediator $M$ and a monotonicity assumption, several questions remain. First, we may be interested in testing that the treatment effect is explained by a non-binary $M$, or by a set of mechanisms---can the approach above be applied when $M$ is non-binary and potentially multi-dimensional? Second, in some applications we may be concerned about violations of the monotonicity assumption---can one test the sharp null of full mediation under relaxations of this assumption? Third, if we reject the sharp null then we know that mechanisms other than $M$ must matter, but how large is the contribution of the alternative mechanisms?
In (ref), we develop a general framework that enables us to tackle all of these questions. We allow the mediator $M$ to take on multiple values and to have multiple dimensions, so long as it has finite support $\{m_{0}, \dots, m_{K-1}\}$. We also allow the researcher to place arbitrary restrictions on $\theta_{lk} = P(M(0) =m_{l}, M(1)=m_{k})$, the fraction of individuals with $M(0) =m_{l}$ and $M(1) = m_{k}$. The monotonicity assumption in the case with scalar $M$ then corresponds to the special case where one imposes that $\theta_{lk} =0$ if $m_l > m_k$. Our framework allows the researcher to impose weaker versions of this requirement---e.g. by allowing for up to $\bar{d}$ share of the population to be defiers---or to completely eliminate the monotonicity requirement altogether. Our framework also allows for various extensions of monotonicity to the setting with multi-dimensional $M$---e.g. a partial monotonicity assumption that imposes that each dimension of $M$ is increasing in $d$.
We derive testable implications of the sharp null of full mediation in this general setting. These testable implications imply that for any set $A$ and any value of the mediator $m_k$, the treatment effect on the compound outcome $\tilde{Y} = 1\{Y \in A, M=m_k \}$ should be no larger than the number of “compliers” with $M(0) = m_{l}$ and $M(1)=m_{k}$ for some $l \neq k $. The intuition for this is that under the sharp null, there should be no effect of the treatment on the outcome for “always-takers” with $M(1) = M(0)$. It follows that the treatment effect on $\tilde{Y}$ can only be driven by compliers, and thus the treatment effect on $\tilde{Y}$ must be weakly smaller than the number of compliers. When $M$ is non-binary, a complication arises because the vector of shares of always-takers and compliers, denoted by $\theta$, is only partially-identified. The testable implication is therefore that there exists some shares $\tilde\theta$ consistent with the observable data such that the inequalities described above are satisfied. Since the identified set for $\theta$ is characterized by linear inequalities, it is simple to verify whether such a $\tilde\theta$ exists by solving a linear program; we also show that the solution to the linear program has a closed-form solution under monotonicity. We further show that these testable implications are sharp in the sense that they exhaust all of the testable information in the data: if they are satisfied, there exists a distribution of potential outcomes (and potential mediators) consistent with the observable data such that the sharp null holds.
We also provide lower bounds on the extent to which the sharp null is violated. In particular, our results imply lower bounds on the fraction of the individuals who have $M=m_{k}$ under both treatments (the “$k$-always-takers”) who are nevertheless affected by the treatment, $\nu_{k} = P(Y(1,m_{k}) \neq Y(0,m_{k}) \mid M(1) = M(0) = m_{k})$. The lower bounds on the $\nu_k$ are informative about the prevalence of alternative mechanisms: if the lower bound on $\nu_{k}$ is large, then alternative mechanisms matter for a high fraction of $k$-always-takers. In (ref), we also derive bounds on the average direct effect for $k$-always-takers, $ADE_{k} = E[Y(1,m_{k}) - Y(0,m_{k}) \mid M(1) = M(0) = m_{k}]$.
In (ref), we show how one can conduct inference on the sharp null of full mediation, exploiting results from the literature on moment inequalities andrews_inference_2023, cox_simple_2022, fang_inference_2023. In Monte Carlo simulations calibrated to our applications, we find good performance for the approach of cox_simple_2022, and thus recommend it in practice. Although for simplicity our main theoretical results focus on the case where $D$ is randomized, in (ref) we show that our results extend to other non-experimental settings, including settings with instrumental variables, conditional unconfoundedness, and distributional difference-in-differences.
\Copy{pgraph:use-cases}{We anticipate that our results will have several potential use-cases in applications, as highlighted by our empirical examples in (ref). First, in many settings, there is an obvious mechanism by which the treatment would be expected to affect the outcome---often referred to as a “mechanical effect”---and it is of economic interest to know whether there are other mechanisms at play. Our motivating example of bursztyn_misperceived_2020 is one such case, where there is the mechanical effect of the information on job applications through the job-search service, and we are interested in whether the information treatment also has an effect on other behavior outside of the lab. Our tests of the sharp null directly address whether the effect of the treatment is driven entirely by the mechanical effect: in bursztyn_misperceived_2020, we reject that the impact of the information treatment on job applications is driven entirely by the mechanical effect of job-search service sign-up. This conclusion is of economic interest, since it suggests that an information treatment that was not tied to a job-search service would also have some effect on labor market outcomes. Our results also help us to quantify the magnitude of the alternative mechanisms: our lower bounds suggest that at least \unskip percent of “never-takers” who would not enroll in the job-search service regardless of treatment status would nevertheless be induced to apply for jobs by receiving the information treatment (compared with an overall ATE of \unskip).
In other settings, there may not be a focal “mechanical effect”, but the researcher may observe that the treatment affects a particular $M$ (or set of $M$'s), and conjecture that this mediator explains the treatment effect. Our tests of the sharp null, along with accompanying lower bounds on $\nu_k$ and $ADE_k$, help to quantify the completeness of such conjectures. This is illustrated in our second application to baranov_maternal_2020, who find that cognitive behavioral therapy for new mothers has an impact on women's economic outcomes. They conjecture that this effect may operate through increased presence of a grandmother in the home and improved relationship quality with the husband. Our results help us to quantify the completeness of these conjectures. Our tests reject the null hypothesis that either of these mechanisms on its own fully explains the treatment, with our lower bounds suggesting that at least 10% of always-takers are affected by the treatment for each mediator. On the other hand, we cannot statistically reject that the two mechanisms together explain the treatment effect. This, of course, does not imply that these are the only two mechanisms, but rather that the data is statistically consistent with the hypothesis that the combination of these mechanisms explains the effect.\footnote{The literature on using short-run surrogates for long-run outcomes often justifies the statistical surrogacy assumption (in part) by arguing that the treatment affects the long-run outcome only through the short-run outcome athey_estimating_2024. In settings where both short- and long-run outcomes are available, our tests of the sharp null may also be useful for assessing the plausibility of these arguments.}}
We have developed the \href{https://github.com/jonathandroth/TestMechs}{TestMechs} R package to facilitate implementation of the methods in this paper.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Related literature.} Our work relates to a large literature on mediation analysis. We briefly overview a few relevant strands of the literature, with a non-exhaustive list of citations, and refer the reader to vanderweele_mediation_2016 and huber_review_2019 for more comprehensive reviews. Much of the mediation analysis literature focuses on identification of average direct effects and indirect effects robins_identifiability_1992,pearl_direct_2001. A key challenge is that even if the treatment $D$ is randomized, it is typically the case that the mediator $M$ is not, and thus it is difficult to identify the effect of $M$ on $Y$ (conditional on $D$). Various strands of the literature have identified the effect of $M$ on $Y$ by assuming conditional unconfoundedness for $M$ imai_identification_2010, using an instrument for $M$ frolich_direct_2017, or adopting difference-in-differences strategies deuchert_direct_2019, schenk_mediation_2023. In contrast, we focus on learning about mechanisms without imposing assumptions that identify the effect of $M$ on $Y$. The question we try to answer is different from most of the existing literature, however: rather than focus on average direct and indirect effects, we start by testing the sharp null that the effect of $D$ on $Y$ is fully explained by a particular mechanism (or set of mechanisms) $M$.\footnote{miles_causal_2023 also considers a sharp null. However, his sharp null is that either $Y(d,m)$ depends only on $d$ or $M(d)$ does not depend on $d$, whereas we consider the sharp null that $Y(d,m)$ depends only on $m$. His focus is also different: rather than testing this sharp null, he considers which measures of the indirect effect are zero when his sharp null is satisfied.} We further provide lower-bounds on the extent to which $M$ does not fully explain the effect of $D$ on $Y$ by lower-bounding the treatment effects for always-takers who have the same value of $M$ regardless of treatment status. We view our work as complementary to much of the literature on mediation analysis, as we impose different assumptions but also address different questions.
A key observation in our paper is that under the sharp null of full mediation, $D$ is an instrument for the effect of $M$ on $Y$. Thus, in the setting where $M$ is binary, existing tools for testing instrument validity with binary endogenous treatment can be used “off-the-shelf” to test the sharp null, both with monotonicity kitagawa_test_2015, huber_testing_2015,mourifie_testing_2017 and without monotonicity balke_bounds_1997,wang_falsification_2017, kedagni_generalized_2020.\footnote{wang_falsification_2017 consider tests of instrument validity when instrument $Z$, treatment $D$, and outcome $Y$ are all binary, and one does not impose monotonicity. They observe that the testable implications imply lower bounds on the average controlled direct effect (ACDE) of $Z$ on $Y$. Although their focus is testing instrument validity, they note in the conclusion that such lower bounds might also be used for “explaining causal mechanisms” in experiments. This observation is thus a precursor to the connections between tests for instrument validity and testing mechanisms derived in the more general setting in our paper.} One of the key technical contributions of our paper is to derive sharp testable implications of the sharp null in the setting where $M$ is potentially multi-valued or multi-dimensional, and where one places arbitrary restrictions on the type shares (e.g. monotonicity or relaxations thereof). Based on the equivalence between testing the sharp null and testing instrument validity described above, our results immediately imply sharp testable implications for settings with a binary instrument and multi-valued treatment, which may be of independent interest. Our testable implications build on the work of sun_instrument_2023, who derived non-sharp testable implications of instrument validity with multi-valued treatments under monotonicity.\footnote{Another related paper is kedagni_generalized_2020, who derive testable implications of instrument validity with potentially multi-valued treatments, without monotonicity. Their testable implications assume a weaker notion of independence, however, which when mapped to our context would imply that $D$ is independent of $Y(\cdot,\cdot)$ but not $M(\cdot)$. Under this weaker notion of independence, their testable implications are sharp in the special case of binary treatment and outcome, but may not be sharp otherwise.}
Our paper also relates to the literature on principal stratification frangakis_principal_2002,zhang_estimation_2003,lee_training_2009,flores_nonparametric_2010. In particular, note that the sub-population of $k$-always-takers corresponds to the so-called principal stratum with $M(1)=M(0)=m_{k}$. Our bounds on $ADE_k$, the average direct effect for $k$-always-takers, match those in the aforementioned papers in the special case where $M$ is binary and one imposes monotonicity. Our bounds on $ADE_k$ extend the existing results to cover settings with non-binary $M$ and/or relaxations of monotonicity. Our primary focus, however, is on testing the sharp null of full mediation, which implies that the fraction of always-takers affected should be zero (a Fisherian sharp null), which is stronger than the weak null of a zero average effect studied in the literature on principal stratification.
Finally, we note that in empirical economics, mechanisms are often studied more informally, rather than using the formal tools for mediation discussed above. One common approach is to show the effects of $D$ on a variety of intermediate outcomes, and to conjecture that a particular intermediate outcome $M$ may be an important mechanism if $D$ has an effect on $M$ (see our application to baranov_maternal_2020 below for an example). The tools developed in this paper give formal methodology for testing the completeness of these conjectures: is the data consistent with the hypothesized $M$ fully explaining the treatment effect, and if not, how important are alternative mechanisms? A second common approach for evaluating mechanisms is heterogeneity analysis: is the treatment effect on $Y$ larger in observable subgroups of the population for which the effect of $D$ on $M$ is larger? Although heterogeneity is often analyzed informally, this approach is sometimes formalized with an over-identification test that evaluates the null that, across subgroups defined by covariate cells, the conditional average treatment effect of $D$ on $Y$ is linear in the conditional average treatment effect of $D$ on $M$ angrist_choice_2023, angrist_instrumental_2023. This approach provides a valid test of our sharp null under the additional assumption that the effect of $M$ on $Y$ for compliers is constant across sub-groups. By contrast, we derive testable implications of the sharp null that do not assume constant effects and do not require the presence of covariates.\footnote{Moreover, our results indicate that the typical over-identification test does not exploit all the information in the data even under the assumption of constant effects: not only can one test the relationship of the average effects across covariate cells, but under the sharp null the restrictions that we derive should also hold within covariate cells.}
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Set-up and notation.} Let $Y$ denote a scalar outcome, $D$ a binary treatment, and $M \in \mathbb{R}^p$ a $p$-dimensional vector of mediators with $K$ support points, $m_0,...,m_{K-1}$. We denote by $Y(d,m)$ the potential outcome under treatment $d$ and mediator $m$. Likewise, $M(d)$ denotes the potential mediator under treatment $d$. The researcher observes $(Y,M,D) = (Y(D,M(D)), M(D), D) \sim P_{obs}$. We use $P$ to denote the joint probability measure over observed variables and potential outcomes and mediators.
We first consider the special case with a binary mediator $M$, which helps us to develop intuition and illustrate connections to the existing literature on testing instrument validity. In the notation just introduced, this corresponds to $K=2$, with $m_0 =0$ and $m_1=1$, so that $M \in \{0,1\}$.
To fix ideas, consider the setting of bursztyn_misperceived_2020. The authors conduct a randomized controlled trial (RCT) in Saudi Arabia focused on women's economic outcomes. Their analysis is motivated by the descriptive fact that at baseline in their experiment, the vast majority of men in Saudi Arabia under-estimate how open other men are to allowing women to work outside the home. After eliciting beliefs, they randomly assign a treated group of men to receive information about the other men's opinions. At the end of the experiment, both treated and untreated men choose between signing their wives up for a job-search service or taking a gift card. bursztyn_misperceived_2020 find that the treatment has a positive effect on enrollment in the job-search service and on longer-run economic outcomes for women, such as applying and interviewing for jobs.
An important question in interpreting these results is whether the treatment increased long-run labor market outcomes solely by increasing take-up of the job-search service, or whether the information led men to change behavior in other ways. This question is important for understanding what might happen if one were to provide men with information about others' beliefs without offering the opportunity to sign up for the job-search service. bursztyn_misperceived_2020 write (p. 3017):
The authors provide some indirect evidence that the effects may not operate entirely through the job-search service---for example, there are effects on men's opinions in a follow-up survey---but they cannot directly link these long-run changes in opinions to economic outcomes. In what follows, we will show that in fact there is information in the data that is directly informative about the question of whether the effects on long-run labor market outcomes are driven solely by the job-search service.
For notation, let $D$ be a binary indicator for receiving the information treatment, $M$ a binary variable indicating job-search service sign-up, and $Y$ a binary variable indicating applying for jobs three to five months after the experiment (i.e., a longer-term labor supply outcome). We let $Y(d,m)$ denote whether a woman would apply for jobs as a function of treatment status $d$ and job-search service sign-up $m$, and let $M(d)$ denote job-search service sign-up as a function of treatment status. Since treatment is randomly assigned, it is reasonable to assume that it is independent of the potential outcomes and mediators, i.e. $D \mathrel{\perp\mspace{-10mu}\perp} (Y(\cdot,\cdot),M(\cdot))$. For our analysis in this section, we will also impose the monotonicity assumption that receiving the information treatment weakly increases job-search service sign-up, so that $M(1) \geq M(0)$ (almost surely). To make this assumption reasonable, we restrict our analysis to the majority of men who prior to the experiment under-estimate other men's openness, so that all men are provided with information that other men are more open than they initially expected, which we expect will increase job-search service sign-up. In the subsequent sections, we will show how this monotonicity assumption can be relaxed, but imposing it will make it easier to highlight the connections to instrumental variables.
We now formalize the null hypothesis that the information treatment only affects long-run outcomes through its effect on job-search service sign-up. In particular, we say that the sharp null of full mediation is satisfied if
i.e. the treatment impacts the outcome only through its impact on $M$. If the sharp null holds, signing up for the job-search service is the only mechanism that matters for long-run job applications. On the other hand, if we reject the sharp null, there is evidence that other mechanisms play a role for at least some people---i.e., there is some impact of changes in beliefs on long-run outcomes that does not operate purely through sign-up for the job-search service at the end of the experiment.
Our first main observation is that if the sharp null holds (together with our assumptions of independence and monotonicity), then $D$ is a valid instrument for the LATE of $M$ on $Y$. This implies that testing the sharp null in this setting is equivalent to testing the validity of the LATE assumptions when both the treatment and instrument are binary. However, prior work has shown that in settings with a binary instrument and treatment, the LATE assumptions have testable implications kitagawa_test_2015, huber_testing_2015, mourifie_testing_2017, and thus such tools can be used to test the sharp null.\footnote{More precisely, these tests are joint tests of the sharp null along with the independence and monotonicity assumptions. However, if we maintain that the latter two hold, then any violations must be due to violations of the sharp null. We explore relaxations of the monotonicity assumption in subsequent sections.} Applying the results in kitagawa_test_2015, with $M$ playing the role of treatment and $D$ the role of instrument, we obtain the following sharp testable implications:
for all Borel sets $A$.
To gain intuition for these testable restrictions, observe that under our monotonicity assumption we can divide the population into “always-takers” who enroll in the job-search service regardless of treatment ($M(0)=M(1)=1$), “never-takers” who do not enroll regardless of treatment ($M(0)=M(1)=0$), and “compliers” who enroll only if treated ($M(0)=0,M(1)=1$). Now, consider the compound outcome $\tilde{Y} = 1\{Y \in A, M=0\}$. For example, if $A= \{1\}$, then in our running example $\tilde{Y}$ is an indicator for the joint event of applying for a job and not signing up for the job-search service. Note that always-takers have $M=1$ under both treatments, and thus always have $\tilde{Y}=0$ regardless of treatment status. Next, observe that if the sharp null holds, then never-takers must also have the same value of $\tilde{Y}$ under both treatments: by definition, they have $M=0$ under both treatments, and hence under the sharp null that $D$ affects $Y$ only through $M$, they also have the same value of $Y$ under both treatments. Since $\tilde{Y}$ is just a compound outcome involving $Y$ and $M$, it follows that they have the same value of $\tilde{Y}$ under both treatments. Observe, further, that the treatment effect of $D$ on $\tilde{Y}$ for compliers must be weakly negative, since compliers have $M=1$ when they are treated, and thus when compliers are treated they have $\tilde{Y} = 0$. It follows that under the sharp null, there must be a weakly negative treatment effect of $D$ on $\tilde{Y}$, and hence $$P(Y\in A, M=0 \mid D=1) - P(Y \in A, M=0 \mid D=0) \leq 0 ,$$ which gives the first testable implication in (ref). If this implication is violated in the data, then we can conclude that---in violation of the sharp null---there must be an effect of the treatment for some never-takers. The second testable implication in (ref) can be derived analogously using the compound outcome of the form $1\{ Y \in A, M=1 \}$.
The argument above implies that if the sharp null holds in bursztyn_misperceived_2020, there should be a negative treatment effect on the compound outcome $1\{Y=1,M=0\}$, i.e. there should be fewer women in the treated group who both apply for jobs and don't use the job service. However, as shown in (ref), the empirical distribution shows that the opposite is true: there are more women who apply for jobs and do not sign up for the job-search service in the treated group ($\hat{P}(Y = 1, M=0 \mid D=1) > \hat{P}(Y = 1, M=0 \mid D=0)$), indicating a violation of the sharp null. This difference is statistically significant at the 5% level, as we will describe in more detail in (ref) after we describe methods for conducting inference.
The data thus reject the sharp null hypothesis that the impact of the information treatment on job applications operates purely through job-search service sign-up. In particular, the data suggest that some never-takers must have their outcome affected by the treatment. We can thus conclude that there is some impact of changes in beliefs on job applications that does not operate mechanically through signing up for the job-search service.
The analysis so far shows that tools originally developed for testing the LATE assumptions can be useful for testing hypotheses about mechanisms. However, several questions remain. First, our rejection of the null implies that the treatment affects the outcome through mechanisms other than job-search service sign-up, but how big are these alternative mechanisms? Second, our analysis relied on the monotonicity assumption that treatment increases job-search service sign-up, but what if we would like to relax this assumption? Third, while our motivating example had a binary $M$, in many cases we may be interested in testing that the treatment is explained by a non-binary mechanism, or by the combination of multiple mechanisms. Can the approach be extended to such cases?
In the subsequent section, we develop a general theoretical framework that allows us to address all of these questions. Our framework accommodates mechanisms $M$ that are potentially multi-valued or multi-dimensional, and allows for relaxations of the monotonicity assumption. Further, in addition to deriving testable implications of the sharp null, we also derive lower bounds on the extent to which the alternative mechanisms matter---in particular, we derive bounds on the fraction of always-takers (or never-takers) that are affected by the treatment, as well as the average effect of the treatment for these always-takers.
We now consider the general case where $M$ is a $p$-dimensional vector with finite support $\{m_0,...,m_{K-1}\}$. We denote by $G=lk$ the event that $M(0) = m_l$ and $M(1)=m_k$. We refer to individuals with $G=kk$ as the $k$-always-takers, and individuals with $G = lk$ for $l \neq k$ as the $lk$-compliers. (Note that the terms “always-taker” and “complier” are used somewhat broadly here. For example, a “never-taker” in the case where $M$ is binary would be referred to as $0$-always taker, and likewise a defier would be a $10$-complier.) We denote by $\theta_{lk} := P(M(0)=m_l,M(1)=m_k)$ the fraction of the population of type $G=lk$, and let $\theta \in \mathbb{R}^{K^2}$ be the vector in the $(K^2-1)$-dimensional simplex that collects the $\theta_{lk}$.
Extending the definition from the previous section, we say that the sharp null of full mediation holds if $$Y(0,m) = Y(1,m) \equiv Y(m) \text{ almost surely, for all } m \in \{m_0,...,m_{K-1}\}.$$ We note that if $M$ is multi-dimensional with, say, the first dimension corresponding to mechanism $A$ and the second corresponding to mechanism $B$, then the sharp null imposes that the treatment operates on $Y$ only through its joint effect on mechanisms $A$ and $B$.
For simplicity, we assume in this section that treatment assignment is independent of the potential outcomes and mediators. In (ref), we show how the results extend to other settings, such as instrumental variables, conditional unconfoundedness, and distributional difference-in-differences.
For our identification results, we allow for the researcher to place arbitrary restrictions on the shares of each compliance type.
We briefly review a few examples of restrictions on $\theta$ (i.e. choices of $R$) that may be natural in some applications.
In contrast to the special case in (ref), where the shares of each type were point-identified, in our general framework the vector of type shares $\theta$ may only be partially-identified. For example, if one relaxes the monotonicity imposed in (ref), then analogous to the setting of instrumental variables without monotonicity huber_sharp_2017, the share of defiers $\theta_{10}$ will generically be partially identified.\footnote{As a concrete example, suppose that $P(M=1 \mid D=1) = 0.5$ and $P(M=1 \mid D=0) = 0.3$. Then the data is consistent with there being no defiers (by setting $\theta_{11} = 0.3$, $\theta_{01} = 0.2$, $\theta_{00} = 0.5$, and $\theta_{10} =0$) but it is also consistent with up to 0.3 fraction of the population being defiers (by setting $\theta_{11} = 0$, $\theta_{01} = 0.5$, $\theta_{00} = 0.2$, $\theta_{10} = 0.3$).} Partial identification of $\theta$ can also arise when $M$ is multi-valued even if one imposes monotonicity. To see this, suppose that $M \in \{0,1,2\}$ and the marginal distributions of $M \mid D$ are as given in (ref), panel (a). As can be seen in the figure, the treated group has a 0.2 higher probability that $M=2$ and a 0.2 lower probability that $M=0$ relative to the control group. This is consistent with 20% of the population being $02$-compliers and there being no other complier types (i.e. $\theta_{02} = 0.2$, $\theta_{01}=\theta_{12}=0$), as shown in (ref), panel (b). However, it is also consistent with a “cascade” in which 20% of the population is $01$-compliers, and another 20% of the population is $12$-compliers (i.e. $\theta_{01}=\theta_{12}=0.2$, $\theta_{02}=0$), as shown in (ref), panel (c).
We will denote by $\Theta_I$ the set of possible values for $\theta$ (i.e. joint distributions on $(M(0),M(1))$) that are consistent with the observed distributions of $M \mid D$. Formally, we define the identified set $\Theta_I$ to be the set of values of $\tilde\theta$ such that
For clarity of notation, we will use $\theta$ for the “true” shares and use $\tilde\theta$ to denote a generic element of the identified set $\Theta_I$. It is worth noting that the first two restrictions above are linear in $\tilde\theta$. Thus, if $R$ is characterized by linear restrictions (as is the case in Examples (ref)-(ref) above), then $\Theta_I$ is characterized by linear constraints, and thus quantities such as $\max_{\tilde\theta \in \Theta_I} \tilde\theta_{kk}$ can be calculated by linear programming. This observation will be useful for practical implementation of the testable implications below, which involve optimizations over $\Theta_I$.
We now derive lower-bounds on the fraction of always-takers whose outcome is affected by the treatment despite having the same value of $M$ under both treatments.\footnote{In (ref), we derive bounds on the average effect of the treatment for the $k$-always-takers.} These lower bounds lead naturally to tests of the sharp null of full mediation, under which the fraction of always-takers affected should be zero. To be more precise, we define $$\nu_k := P(Y(1,m_k) \neq Y(0,m_k) \mid G =kk)$$ to be the fraction of $k$-always-takers whose outcome is affected by the treatment despite always having $M=m_k$ under both treatments. The $\nu_k$ are a measure of the strength of mechanisms other than $M$: they tell us what fraction of the $k$-always-takers has a direct effect of the treatment. Under the sharp null of full mediation, $Y(1,m_k) = Y(0,m_k)$ with probability 1, and thus $\nu_k =0$ for all $k$. By contrast, if $\nu_k$ is close to 1 for a particular $k$, then alternative mechanisms other than $M$ matter for nearly all $k$-always-takers.
Our first main result provides a lower bound on the $\nu_k$ as a function of the observable data and the type shares $\theta$. To simplify notation, let $$\Delta_k(A) := P(Y \in A, M=m_k \mid D=1) - P(Y \in A, M=m_k \mid D=0)$$ be the difference in the probability that $Y \in A$ and $M=m_k$ between the treated and control groups. Let $(x)_+ := \max\{x, 0\}$. We then have the following sharp lower bound on the fraction of $k$-always-takers affected by the treatment.
Equation (ref) gives a lower-bound on the fraction of $k$-always-takers whose outcome is affected by the treatment, $\nu_k$, involving the true type shares $\theta$ and functions of the observable data (the $\Delta_k$). The true $\theta$ will generically not be point-identified, and so (ref) cannot be used directly to give a feasible lower bound on $\nu_k$. However, we know that the true $\theta$ must lie in the identified set $\Theta_I$. This leads to the feasible implication given in (ref) which replaces $\theta$ in (ref) with some $\tilde\theta \in \Theta_I$. The second part of (ref) shows that the bound given in (ref) is sharp in the sense that there exists a distribution of primitives consistent with the observable data such that the lower bound holds with equality. It further shows that under this distribution of primitives, there is no direct effect of treatment for complier types; hence, we can only obtain non-trivial lower bounds on direct effects for the always-takers.
Recall that under the sharp null of full mediation, the fraction of always-takers whose outcome is affected by the treatment should be zero. We thus immediately obtain the following testable implications of the sharp null by setting $\nu_k = 0$ in (ref).
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Intuition.} Observe that since the treatment is randomly assigned, $\Delta_k(A)$ is simply the average effect of the treatment $D$ on the compound outcome $\tilde{Y} = 1\{Y \in A, M=m_{k}\}$. Note that under the sharp null of full mediation, the treatment effect of $D$ on $\tilde{Y}$ should be zero for always-takers, since they have the same value of $M$ and $Y$ under both treatments. This implies that under the sharp null the effect of $D$ on $\tilde{Y}$ is driven only by compliers and thus cannot be “too large”. This is precisely what is captured in (ref), which shows that under the sharp null, $\Delta_k(A)$ should be bounded above by the total mass of $lk$-compliers, $\sum_{l: l \neq k} \theta_{lk}$, regardless of the choice of $A$. If, in fact, the treatment effect on $\tilde{Y}$ is larger than the number of $lk$-compliers, then it must be that some $k$-always-takers had their outcome affected by the treatment, in violation of the sharp null. Indeed, the lower bound on the fraction of $k$-always-takers whose outcome is affected by the treatment ($\nu_k$) given in (ref) is proportional to the positive part of the difference between $\sup_A \Delta_k(A)$ and the number of $lk$-compliers.
A slightly more formal sketch of the argument is as follows. We can write $\Delta_k(A) = E[\tilde{\tau}]$, where $\tilde{\tau} $ is the individual-level treatment effect of $D$ on $\tilde{Y}$.\footnote{Formally, $\tilde\tau = \tilde{Y}(1) - \tilde{Y}(0)$ for $\tilde{Y}(d) = 1\left\{Y(d,M(d)) \in A, M(d) =m_{k} \right\}$.} Since $\tilde{Y}$ is a binary outcome, the treatment effect $\tilde{\tau}$ must be in $\{-1,0,1\}$. We now argue that $\tilde\tau$ can equal 1 only if an individual is either an $lk$-complier, or a $k$-always taker with $Y(1,m_{k}) \neq Y(0,m_{k})$. To see why this is a case, note that for $\tilde{\tau}$ to be 1, an individual must have $M(1) =m_{k}$, and thus must be either an $lk$-complier or a $k$-always taker. However, a $k$-always taker can have a treatment effect on $\tilde{Y}$ of 1 only if $Y(1,m_{k}) \in A$ and $Y(0,m_{k}) \not\in A$, which implies that $Y(1,m_{k}) \neq Y(0,m_{k})$. It follows that
Using the fact that $\theta_{lk} = P(G=lk)$ by definition, we can rewrite the inequality as
Rearranging terms, we obtain that $$\theta_{kk} P( Y(1,m_{k}) \neq Y(0,m_{k}) \mid G = kk) \geq \Delta_k(A) - \sum_{l: l \neq k} \theta_{lk} ,$$ which together with the fact that probabilities are non-negative yields (ref). $\blacktriangle$
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Computation of bounds.} Suppose we are interested in computing the lower-bound on $\nu_k$ for a particular $k$. Recall that for any $\tilde\theta \in \Theta_I$, we have $P(M=m_{k} \mid D=1) = \tilde\theta_{kk} + \sum_{l: l \neq k} \tilde\theta_{lk}$. It follows that we can re-write (ref) as $$\tilde{\theta}_{kk} \nu_k \geq \left(\sup_A \Delta_k(A) - (P(M=m_{k} \mid D=1) - \tilde\theta_{kk}) \right)_+ ,$$ where now the lower-bound depends on $\tilde\theta$ only through $\tilde\theta_{kk}$. To compute a lower-bound on $\nu_k$, we must minimize the lower-bound for $\nu_k$ given in the previous display over $\tilde\theta \in \Theta_I$. It can be shown (see (ref)) that the minimum is actually obtained at the minimum possible value of $\tilde\theta_{kk}$, i.e. by plugging in $\tilde{\theta}_{kk}^{min} := \inf_{\tilde\theta \in \Theta_{I}} \tilde\theta_{kk}$ into the expression in the previous display. When $R$ is a polyhedron (as in our examples above), the identified set $\Theta_I$ is characterized by linear inequalities, and thus $\tilde{\theta}_{kk}^{min}$ can be easily computed by solving a linear program. Assuming $\tilde\theta_{kk}^{min} > 0$, we then obtain the bound $$\nu_k \geq \frac{1}{\tilde\theta_{kk}^{min}} \left(\sup_A \Delta_k(A) - (P(M=m_{k} \mid D=1) - \tilde\theta_{kk}^{min}) \right)_+ .$$ Similarly, to test whether the observable data is compatible with the sharp null we must verify whether there is any $\tilde\theta \in \Theta_I$ such that $\sup_A \Delta_k(A) \leq \sum_{l: l \neq k} \tilde\theta_{lk}$ for all $k$. By the same argument as in the previous paragraph, this is equivalent to testing whether there is any $\tilde\theta \in \Theta_I$ such that $\sup_A \Delta_k(A) \leq P(M=m_{k} \mid D=1) - \tilde\theta_{kk}$ for all $k$. Such a $\tilde\theta \in \Theta_I$ exists if and only if the solution to the linear program
is weakly negative, and so given knowledge of the distribution of the observable data, testing the implications of the sharp null is equivalent to solving a linear program.\footnote{We note that linear programming has been used for tractability in a variety of related but distinct partial identification settings; see, e.g. mogstad_policy_2024, ji_model-agnostic_2024, yap_sensitivity_2025 for some recent contributions.}
The previous section derived testable implications of the sharp null of full mediation, as well as measures of the extent to which it is violated, which involved the distribution of the observable data $(Y,M,D) \sim P_{obs}$. We now derive methods for inference on the sharp null given a sample of $N$ $iid$ observations (or clusters) drawn from $P_{obs}$, $(Y_i,M_i,D_i)_{i=1}^{N}$. For simplicity of notation, we focus on testing the sharp null, although a simple adaptation of the described approach can be used to test null hypotheses of the form $H_0: \nu_k \leq \nu^{ub}_k \hspace{.1cm} \forall k$ for any $\nu^{ub}_k$ (with the sharp null the special case with $\nu^{ub}_k = 0$ for all $k$.)
We first comment on the non-standard nature of the inference problem. Recall that the testable implications of the sharp null are equivalent to whether the linear program (ref) has a weakly negative solution. However, functions of the observable data enter the constraints of the linear program, and it is well-known that the solution to a linear program can be non-differentiable in the constraints (shapiro_stochprog_1991). Second, the function of the observable data in the constraints, $\sup_A \Delta_k(A)$, is itself potentially non-differentiable in the underlying data-generating process. If the outcome $Y$ is discrete, for example, then $\sup_A \Delta_k(A) = \sum_{y} (f_{Y,M=m_k \mid D=1}(y) - f_{Y,M=m_k \mid D=0}(y))_+ $, where $(x)_+ = \max\{x,0\}$, which is clearly non-differentiable in the partial probability mass functions $f_{Y,M=m_k \mid D=d}(y) := P(Y=y, M=m_k \mid D=d)$ if $f_{Y,M=m_k \mid D=1}(y) = f_{Y,M=m_k \mid D=0}(y)$ for any $y$. Since bootstrap methods are generally invalid when the target parameter is non-differentiable in the underlying data-generating process fang_inference_2019, we cannot simply bootstrap the solution to (ref).
We now show that methods from the moment inequality literature can be used to circumvent these issues. We focus on the case where the distribution of $Y$ is discrete, with support points $y_1,...,y_Q$. As we discuss in (ref) below, if $Y$ is continuous, then the tests we derive remain valid if one uses a discretization of $Y$, although at the potential loss of sharpness. We also focus on the case where $R$ takes the polyhedral form $R = \{ \theta \in \Delta : B \theta \leq c\}$. To see the connection with moment inequalities, observe that with discrete $Y$, we have that $$\sup_A \Delta_k(A) = \sum_{q=1}^Q \left(P(Y=y_q, M=m_k \mid D=1) - P(Y=y_q, M=m_k \mid D=0) \right)_+$$
where again $(x)_+ = \max \left\{x,0 \right\}$. It follows that the inequality $\sup_A \Delta_k(A) \leq P(M=m_k \mid D=1) - \tilde\theta_{kk} $ holds if and only if there exist $\delta_{k1},...,\delta_{kQ}$ such that
Hence, the testable implications of the sharp null derived in (ref) are equivalent to the statement that there exists some $\tilde\theta \in \Theta_I$ and $\delta$ such that (ref)-(ref) hold for all $k=0,...,K-1$.
Observe, further, that $\delta,\tilde\theta$ and the observable probabilities enter the inequalities (ref)-(ref) linearly, and the same is true for the constraints that determine $\Theta_I$. Letting $\omega = (\tilde\theta',\delta')'$, it follows that we can write the testable implications of the sharp null as
where $C_1,C_2$ are known matrices (not depending on the data) and $p$ is a vector that collects probabilities of the forms $P(Y=y_q, M=m_k \mid D=d)$ and $P(M=m_k \mid D=d)$. A recent literature on moment inequalities has considered testing hypotheses of the above form---in which the nuisance parameter $\omega$ enters linearly and with known coefficients $C_1$---given estimates $\hat{p}$ such that $\sqrt{N}(\hat{p}-p) \to N(0,\Sigma)$ andrews_inference_2023, cox_simple_2022, fang_inference_2023, cho_simple_2024. Under mild conditions, the central limit theorem implies that the vector of conditional sample means $\hat{p}$ is asymptotically normal, and thus existing methods from the aforementioned papers can thus be used directly to test the sharp null of full mediation.
To evaluate the methods for inference described above, we conduct Monte Carlo simulations calibrated to our applications to bursztyn_misperceived_2020 and baranov_maternal_2020 in (ref) below. For simplicity, we focus on testing the sharp null under a monotonicity assumption.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Treatment, outcome, and mediator.} The treatment, outcome, and mediator in our simulations match those in our empirical applications. For bursztyn_misperceived_2020, the treatment is receiving information about other men's beliefs, the outcome is a binary indicator for applying for jobs outside of the home, and the mediator is a binary indicator for job-search service sign-up. For baranov_maternal_2020, the treatment is cognitive behavioral therapy and the outcome is an index of financial empowerment. We consider two mediators, a binary indicator for the presence of a grandmother in the household, and a relationship-quality score, which is a score on a 1-5 scale.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Sample sizes.} The sample used for our main analysis of bursztyn_misperceived_2020 contains \unskip people, with treatment assignment randomized at the individual level (approximately half (\unskip) were treated). For the simulations calibrated to bursztyn_misperceived_2020, we draw \unskip $iid$ observations to match the original sample size. In baranov_maternal_2020, treatment was assigned at the level of a cluster (i.e. at the Union Council level), with a total of 40 clusters (20 treated, 20 control), and a total sample size of approximately 600 individuals (\unskip or \unskip depending on the choice of $M$). For simulations calibrated to baranov_maternal_2020, we therefore draw 20 independent clusters from each treatment group. Given the small number of clusters, we expect this to be a relatively challenging setting for inference. To evaluate the impact of having a small number of clusters, we also consider alternative simulation designs where we sample 40 or 100 clusters of each treatment type, with a total of 80 and 200 clusters for each design.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Description of DGP.} In all of our simulations, we sample the distribution of $(Y,M)$ for control units (or clusters) from the empirical distribution of control units (or clusters) in our applications (i.e. from $(Y,M) \mid D=0$). For treated units in our simulations, we draw with probability $t$ from the empirical distribution of $(Y,M)$ for treated units, and with probability $1-t$ from the empirical distribution for control units, where $t \in \{0,0.5,1\}$ is a simulation parameter. Thus, when $t=1$, we are sampling both treated and control units in the simulation from the empirical distribution in the data, under which the sharp null is violated. This allows us to assess the power of the various tests. When $t=0$, on the other hand, the distribution of $(Y,M)$ for both treated and control units in the simulation is drawn from the empirical distribution for control units in the original data. This ensures that the testable implications of the sharp null and monotonicity are satisfied, which allows us to evaluate size control. (In fact, the design ensures that all of the implied moment inequalities hold with equality, which is generally a challenging setting for size control for moment inequality methods.) When $t=0.5$, the distribution of $(Y,M)$ for treated units is a mixture of the empirical distribution for treated and control units in the original data. Thus, the sharp null is violated, but the violation is smaller than under the case when $t=1$. Comparing across the cases $t=0.5$ and $t=1$ thereby allows us to evaluate how power changes with the size of the violation of the null.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Methods used.} To implement tests based on moment inequalities as described above, we consider the hybrid test proposed by andrews_inference_2023, the conditional conditional chi-squared test proposed by cox_simple_2022,\footnote{More precisely, CS propose a conditional chi-squared test and a “refined” version of this test. The refinement addresses the fact that the baseline test can be conservative. Since the refinement is computationally costly with many moments, and only matters when one moment is binding, we only implement the refinement in DGPs with a binary outcome, for which there are fewer moments.} and the test proposed by fang_inference_2023.\footnote{When $M$ is binary, we implement the formulation of the moment inequalities derived in (ref) without nuisance parameters. For non-binary $M$, we use the formulation in (ref).} For comparison to existing methods in the case where $M$ is binary, we consider the test for instrument validity proposed by kitagawa_test_2015.\footnote{For the DGPs based on baranov_maternal_2020, we use a modified version of kitagawa_test_2015 to account for clustering.} In the simulations calibrated to bursztyn_misperceived_2020, the outcome is binary, and thus no discretization of the outcome is needed. For the simulations calibrated to baranov_maternal_2020, where the outcome takes many values, for the moment inequality methods we consider a discretization of the outcome based on 5 bins in our main specification (see (ref)). We also consider alternative specifications using 2 and 10 bins. Since the K test does not require a discrete outcome, we use the original continuous outcome when implementing the K test. Implementation of the FSST test requires specifying the moment-selection tuning parameter $\lambda$. We consider the two choices recommended by FSST in their Remark 4.2, one of which is data-driven and the other is not. We refer to the resulting tests as FSSTdd and FSSTndd (where `dd' denotes data-driven). For CS and ARP, we use analytic estimates of the variance of the moments, assuming the data are drawn $iid$ in the simulations calibrated to bursztyn_misperceived_2020, or that clusters are drawn $iid$ in the simulations calibrated to baranov_maternal_2020. Since the K and FSST tests require bootstrap replicates, we use a non-parametric bootstrap at either the individual or cluster level, as appropriate.\footnote{We have verified that ARP and CS return similar results if we use an analogous bootstrap estimate of the variance rather than the analytic estimates.} All tests impose monotonicity as defined in (ref).\footnote{\Copy{mon-violated}{As described in the empirical section below, for the multi-valued $M$ in baranov_maternal_2020, the empirical distribution for $M \mid D$ is inconsistent with monotonicity (although the violation is not statistically significant). Our simulation design ensures that the data are consistent with monotonicity under the null DGP ($t=0$). However, the alternative DGPs ($t \in \{0.5,1\}$) are based on the empirical distribution and are therefore inconsistent with monotonicity. Hence, the reported power of tests imposing monotonicity under these alternatives corresponds to their power to jointly detect a violation of the sharp null and a relatively small violation of the monotonicity assumption. We found the ranking of power across methods was identical when we modified the tests to allow for the minimal relaxation of monotonicity consistent with the data, and therefore present the results imposing monotonicity for simplicity and consistency with the other specifications.}} All tests are implemented with nominal size of 5%.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Simulation results.} (ref) reports the results for simulations designs where we have a binary mediator. This includes the DGP based on bursztyn_misperceived_2020 (Panel A), and the DGPs that are based on baranov_maternal_2020 where the considered mediator is the binary indicator for the presence of a grandmother (Panels B-D). (ref) shows results calibrated to baranov_maternal_2020 using the non-binary relationship quality variable as the mediator. Both tables show the rejection probabilities for each of the methods described above under different simulation designs. To quantify the magnitude of the violations of the sharp null, the table also reports the lower-bound on the fraction of always-takers affected ($\bar{\nu}$).\footnote{For the simulations calibrated to baranov_maternal_2020 with multi-valued $M$, we compute the lower bound on $\bar{\nu}$ in the same way as described in (ref) in the application section below, which deals with the fact that the empirical distribution shows a small (but statistically insignificant) violation of monotonicity.}
We first evaluate size control. Recall that DGPs with $t=0$ impose the sharp null of full mediation. Across nearly all simulation designs, we find that the ARP, CS, and K tests have close to nominal size, with rejection probabilities no larger than 9% for a 5% test. The one notable exception is the simulations in Panel B of (ref), where there are only 40 independent clusters, in which case CS is somewhat over-sized, with a null rejection probability of 0.15. Doubling the number of clusters to 80 (Panel C) restores approximate size control, however. We find that the FSST tests often have reasonable size control for settings with a large number of independent observations or clusters, but can be substantially over-sized in settings with a small or moderate number of clusters using the two default choices of tuning parameters, particularly with multi-valued $M$ (e.g. rejection probabilities of 0.274 and 0.178 in (ref), Panel A).
We next evaluate power, focusing on the simulations with $t=0.5$ and $t=1$ under which the null is violated. Across all of the simulation designs, the CS test has power similar to or greater than that of ARP. The differences can be substantial in some cases, particularly with multi-valued $M$ (e.g. power of $0.96$ vs $0.16$ in Panel B of (ref)). Likewise, the power of the FSST tests is similar to or exceeds that of the CS test across nearly all simulation designs, although this comparison must be taken with some caution in cases where the FSST test appears to be over-sized. Finally, we note that in all of the simulations with binary $M$ ((ref)), the power of the three moment inequality tests (ARP, CS, FSST) is either similar to or exceeds that of the K test. This is the case both when the outcome is binary (Panel A), and when the outcome is continuous (Panels B-D). Recall that when the outcome is continuous, the moment inequality tests use a discretization of the outcome to 5 bins, whereas the K test does not use a discretization. The favorable power comparisons in Panels B-D thus suggest that discretization does not come at a large loss of power in this simulation design, although of course this conclusion may be specific to the particular DGP studied here.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Choice of bins.} In (ref), we present results for simulations calibrated to baranov_maternal_2020 using a discretization with 2 or 10 bins, rather than the 5 considered above. The comparisons of size control and power across the methods are similar to the results reported above. However, increasing the number of bins from 5 to 10 exacerbates the size control issues seen with CS in the simulations based on baranov_maternal_2020 with only 40 clusters and binary $M$, while decreasing the number of bins to 2 improves size control. This is intuitive, since the number of moments used increases with the bin size, and thus we expect the quality of the central limit theorem approximation to be worse with more bins. In (ref), we report the median number of independent observations (unique clusters) per $(Y^{disc},M,D)$ cell in each of the simulation designs. We find that CS exhibits close to nominal coverage in all specifications with 15 or more independent observations per cell, but exhibits moderate size distortions in several (but not all) of the specifications with fewer than 15 observations per cell. In terms of power, we do not find an obvious pattern across bin sizes, with power increasing in the number of bins for some tests/DGPs and decreasing for others. This reflects the fact that although the testable implications become sharper the more bins are used (see (ref)), the finite-sample power of moment inequality methods often decreases when increasing the number of moments. Based on our simulations, we heuristically recommend that researchers should try to have at least 15 independent observations per cell, acknowledging that there is a potential tradeoff between size control and power. This heuristic roughly aligns with that in andrews_cmi_2013, who recommend having 10-20 observations per cell in settings with conditional moment inequalities.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Choice of test.} Based on our simulations, CS strikes us a reasonable default choice for most empirical settings, given that it has approximate size control across most of our simulation designs and favorable power relative to ARP. However, ARP performs somewhat better in terms of size control in settings with a small number of clusters, and thus may be an attractive alternative for researchers concerned about size control in such settings, albeit at the loss of some power (particularly with multi-valued $M$). Likewise, FSST may offer power improvements relative to CS in settings with a large number of independent observations, so that size control is less of a concern. In our applications below, we report results for CS in the main text; analogous results for ARP and FSST are given in (ref).
Our results so far have relied on the assumption that the treatment is as good as randomly assigned ((ref)). The role of randomization was simply to identify the distribution of potential outcomes and potential mediators under each treatment: randomization ensures that the observable distribution $(Y,M) \mid D=d$ corresponds to the distribution of $(Y^{\text{tot}}(d), M(d))$, where $Y^{\text{tot}}(\cdot) := Y(\cdot, M(\cdot))$. In this section, we show that analogous results go through if the distributions of $(Y^{\text{tot}}(d),M(d))$ are identified through some other strategy. One simply substitutes the expressions involving $(Y,M) \mid D=d$ in our earlier results with the formulas for $(Y^{\text{tot}}(d), M(d))$ under the alternative identifying assumptions. We first provide a general result extending our results to the case where $(Y^{\text{tot}}(d), M(d))$ is identified, then discuss how this applies to the common settings of instrumental variables, conditional unconfoundedness, and difference-in-differences.
To be more precise, define $$\Delta_k^*(A) := P(Y^{\text{tot}}(1) \in A, M(1) =m_{k}) - P(Y^{\text{tot}}(0) \in A, M(0) =m_{k})$$ to be the treatment effect on the compound outcome $1\{Y \in A, M=m_{k}\}$. Note that under randomization, $\Delta_k^*(A) = \Delta_k(A)$. Likewise, define
and observe that $\Theta_I^* = \Theta_I$ under randomization. We then have the following result, analogous to (ref), which gives lower bounds on the $\nu_k$ in terms of probabilities involving $(Y^{\text{tot}}(d),M(d))$.
We likewise obtain the following corollary regarding the sharp null of full mediation, analogous to (ref).
(ref) and (ref) show that we can bound the fraction of always-takers affected by treatment and test the sharp null so long as the marginal distributions of $(Y^{\text{tot}}(d),M(d))$ are identified. We now outline several settings where these distributions are identified, possibly for some sub-population of interest (e.g. for compliers with respect to an instrument).
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Instrumental variables.} Suppose rather than $D$ being randomly assigned, we have a valid binary instrument $Z \in \{0,1\}$ for $D$. For example, in an experiment with imperfect compliance, $Z$ could be the randomized treatment assignment, and $D$ the realized treatment take-up. Suppose that $Z$ satisfies the standard instrument-monotonicity, relevance, exclusion, and independence assumptions imbens_identification_1994. To be precise, we assume that $D=D(Z)$, where $D(1) \geq D(0)$ (a.s.) and $P(D(1) > D(0)) > 0$, and that $(Y,M) = (Y(D(Z),M(D(Z))), M(D(Z)))$ for $Z \mathrel{\perp\mspace{-10mu}\perp} (Y(\cdot,\cdot),M(\cdot),D(\cdot))$. These assumptions allow us to identify the LATE of $D$ on $Y$, i.e. the treatment effect of $D$ on $Y$ for instrument-compliers.\footnote{We use the phrase “instrument-compliers” to refer to compliers with respect to the instrument, i.e. with $D(1) > D(0)$, in contrast to the use of “compliers” elsewhere in the paper, which refers to individuals with $M(1) \neq M(0)$.} They likewise allow us to identify the LATE of $D$ on $M$. We might then be interested in the extent to which the effect of $D$ on $Y$ for instrument-compliers operates through $M$. To apply the results in (ref) and (ref), we need to identify the distributions of $(Y^{\text{tot}}(d),M(d))$ for instrument-compliers. It is well-known, however, that the distributions of potential outcomes for instrument-compliers are non-parametrically identified (see, e.g., abadie_semiparametric_2003). In particular, if we define $C^z = 1\{ D(1) > D(0)\}$ to be an indicator for being an instrument-complier, then $P( Y^{\text{tot}}(1) \in A, M(1) =m_{k} \mid C^z =1 )$ is identified as
In words, the probability that $Y^{\text{tot}}(1) \in A$ and $M(1) =m_{k}$ for instrument-compliers corresponds to the population IV estimand using the compound outcome $D \cdot 1\{Y \in A, M=m_{k} \}$. We can likewise identify $P(Y^{\text{tot}}(0) \in A, M(0) =m_{k})$ using the population IV estimand with the compound outcome $-(1-D) \cdot 1\{Y \in A, M=m_{k} \}$.\footnote{Note that the instrument exclusion restriction implies that $Z$ affects $Y$ only through $D$, and the sharp null implies that $D$ affects $Y$ only through $M$. Hence, under the sharp null, $Z$ affects $Y$ only through $M$. A simple approach to testing the sharp null in IV settings is thus to apply the results in (ref) for experiments, relabeling the randomized treatment as $Z$ and ignoring the endogenous take-up $D$. In Section (ref) we show that this is equivalent to applying (ref) when one imposes monotonicity of $M(d)$ in $d$, but does not exhaust all of the information in the data without this monotonicity condition.}
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Conditional unconfoundedness.} Suppose that $D$ is as good as randomly assigned conditional on observable characteristics, $D \mathrel{\perp\mspace{-10mu}\perp} (Y(\cdot,\cdot), M(\cdot)) \mid X$. Under the overlap condition that $\eta < E[D \mid X] < 1-\eta$ for some $\eta>0$, the distributions of $(Y^{\text{tot}}(d),M(d))$ are non-parametrically identified by re-weighting the observed outcomes by the propensity score $p(X) := E[D \mid X]$,
Hence, one can apply the results in (ref) and (ref) to obtain lower-bounds on the fraction of always-takers affected by the treatment, and to test the sharp null.$\blacktriangle$
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Difference-in-differences.} Consider a two-period setting where no units are treated in the first period and units with $D=1$ are treated in the second period. Under a parallel trends assumption for $Y^{\text{tot}}(0)$, we can identify $E[Y_2^{\text{tot}}(1) - Y_2^{\text{tot}}(0) \mid D=1]$, the average treatment effect on the treated (ATT) in period 2 (where the 2 subscript denotes the second time period). Likewise, under a parallel trends assumption for $M(0)$, we can identify $E[M_2(1) - M_2(0) \mid D=1]$, the ATT on $M_2$. We may then be interested in the extent to which the ATT for $Y_2$ is driven by the effect of the treatment on $M_2$. However, the standard parallel trends assumptions for $Y^{\text{tot}}(0)$ and $M(0)$ identify only the counterfactual means for the treated group and not the counterfactual distributions, as would be required to apply the results in (ref) and (ref). However, several papers have developed extensions of the standard difference-in-differences approach that allow one to infer the full counterfactual distribution for the treated group athey_identification_2006, callaway_quantile_2019, roth_when_2023. These approaches could be applied to identify the distributions of $(Y^{\text{tot}}_{2}(d), M_2(d)) \mid D=1$, which could then be used in conjunction with (ref) and (ref) to examine the extent to which the ATT on $Y_2$ is driven by the effect on $M_2$. $\blacktriangle$ \\
The approach to inference described in (ref) naturally extends to these settings as well. In (ref), we considered inference based on a vector of estimates $\hat{p}$, where each element of $\hat{p}$ corresponded to an estimate of a probability of the form $P(Y^{\text{tot}}(d) \in A, M=m_{k})$ under the assumption of randomly assigned treatment. To test the sharp null under the identifying strategies described above, one simply replaces the $\hat{p}$ in (ref) with analogous estimates of $P(Y^{\text{tot}}(d) \in A, M=m_{k})$ derived under the alternative identifying assumptions. For example, in IV settings we could use two-stage least squares estimates based on the sample analog to the identification result in (ref); under conditional unconfoundedness, we could use inverse-probability weighting estimates based on sample analogs to equation (ref) (one could likewise use outcome-modeling or doubly-robust methods); and in difference-in-differences settings we could use estimates obtained from any of the distributional difference-in-differences approaches mentioned above.
We note that the approach introduced in this section uses only the information about the implied marginal distributions of $(Y^{\text{tot}}(d),M(d))$. Under conditional unconfoundedness, one could potentially use information on the conditional distributions $(Y^{\text{tot}}(d),M(d)) \mid X$.\footnote{The same applies to conditional exogeneity or conditional parallel trends in the IV and difference-in-differences settings, respectively.} Indeed, if treatment is as-good-as-randomly assigned conditional on $X$, then the implications of the sharp null in (ref) should hold conditional on $X$ (almost surely). If $X$ is discrete and takes on a small number of values (e.g. an indicator for gender), then it is straightforward to apply the testable content separately for each possible value of $X$, and inference can be conducted by simply stacking the moment inequalities for each value of $X$. When $X$ is continuously distributed, however, fully exploiting the covariates becomes complicated since it requires estimating conditional distributions and involves the infinite-dimensional conditional type shares, $\theta(X)$. The approach we described here is simpler, since it does not involve estimating full conditional distributions and has a finite-dimensional parameter $\theta$. However, it may be conservative since it does not exploit all the information available. carr_testing_2023 and farbmacher_instrument_2022 propose methods for testing IV validity conditional on covariates in the setting with binary endogenous treatments, which can be applied off-the-shelf in the setting in (ref) with binary $M$ and monotonicity. Whether these approaches can be extended to the setting with multi-valued $M$ and relaxations of monotonicity is an interesting question for future research.
We now revisit our application to bursztyn_misperceived_2020 from (ref). Recall that our treatment $D$ is random assignment to an information treatment about other men's beliefs about women working outside the home, $M$ is sign-up for the job-search service, and $Y$ is an indicator for whether the wife applies for jobs outside of the home. For our main specification, we restrict attention to the majority of men who at baseline under-estimate other men's beliefs, so that the monotonicity assumption that treatment weakly increases job-search service is plausible. (We find similar results when including all men; see (ref).)
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Statistical significance.} Recall from (ref) that the testable implications of the sharp null were rejected based on the empirical distribution. Using the approach to inference described above, we find these violations are in fact statistically significant, with a $p$-value of \unskip using the CS test.\footnote{Since the outcome is binary, no discretization is needed for this application. The $p$-value reported here is the smallest value of $\alpha$ for which the test rejects.} (We obtain similar results using the other tests; see (ref).) The data thus provides strong evidence that the impact of the information treatment on long-run labor market outcomes does not operate solely through the sign-up for the job-search service. In particular, there are some never-takers who would not sign up for the service under either treatment who are nevertheless induced to apply for jobs by the treatment. We thus see that, for at least some people, the information treatment has meaningful impact outside of the lab, beyond its impact on job-search service sign-up.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Magnitudes of alternative mechanisms.} How large are the effects of the information treatment for those who are not induced to sign-up for the job-search service? (ref) gives us a lower bound on the fraction of the always-takers/never-takers who are affected by the treatment despite having no effect on job-search service signup. Our estimates of the lower bounds suggest that at least \unskip percent of “never-takers” who would not be signed up for the job-search service under either treatment are nevertheless affected by the treatment. (We obtain a trivial lower-bound of 0 for the “always-takers”.) Applying the results in (ref), we also estimate lower and upper bounds on the average effect for these never-takers of \unskip to \unskip.\footnote{In this simple setting with a binary outcome, the lower bound for the average effect corresponds exactly to our lower bound on the fraction of always-takers affected.} For comparison, our estimate of the overall average treatment effect is \unskip. The effect for never-takers is thus of a fairly similar magnitude to that of the total population, despite the fact that they have no change in job-search service signup. If we were willing to assume that the direct effects (i.e. effects not through the job-search service) were similar between always-takers, never-takers, and compliers (granted, a strong assumption), this would imply that the majority of the total effect operates through the information treatment.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Robustness to monotonicity violations.} Our baseline results impose the monotonicity assumption that receiving the information that other men are more open to women working than one initially thought only increases job-search service sign-up. This could be violated if, for example, there is measurement error in the initial elicitation of beliefs, so that some men included in our sample actually initially over-estimated other men's beliefs. To explore robustness to violations of the monotonicity assumption, we re-compute our bounds on the fraction of never-takers affected allowing for up to $\bar{d}$ fraction of the population to be defiers. We find that the estimated lower-bound is positive for $\bar{d}$ up to \unskip, which corresponds to \unskip% of the population being defiers, or put otherwise, \unskip defiers for every complier.
We next examine the setting of baranov_maternal_2020. They present long-run results on an RCT that randomized access to a cognitive behavioral therapy (CBT) program intended to reduce depression for pregnant women and recent mothers. In a seven-year followup, they find that the program substantially reduced depression and increased measures of women's financial empowerment, such as having control over finances and working outside of the home. They are then interested in the mechanisms by which treating depression increases financial empowerment. They therefore examine a variety of intermediate outcomes. Two of the outcomes for which they find positive effects of the treatment are the presence of a grandmother in the household (a proxy for family support) and the women's self-reported relationship quality with the husband (on a 1-5 scale). They write (p. 849):
The tools developed above allow us to test the completeness of these conjectured mechanisms. Can the presence of a grandmother or improved relationship quality, either individually or together, explain the impact on financial empowerment, or must there be other mechanisms at play as well? We begin by analyzing each of these mechanisms separately, and then turn to studying the combination of the two.
Since the outcome in baranov_maternal_2020 is a continuous index, we rely on a discretization using 5 bins, which we found in our Monte Carlo simulations calibrated to this application led to a reasonable tradeoff between size control and power (although with moderate size distortions for the setting with binary $M$). This also roughly aligns with our heuristic for choosing the bin size in the setting with binary $M$, as it yields a median cell count of 14. In the setting with non-binary $M$, this leads to a median cell count of 8, which is somewhat below our heuristic threshold of 15; using 2 bins would deliver a count of 15. Nevertheless, in the main text we present results using 5 bins to maintain consistency in the outcome variable across different specifications, and because the Monte Carlo results suggest that this choice performs well in this application with multi-valued $M$ despite the small cell size. This also gives the $\nu_k$ naturally-interpretable units as the fraction of always-takers whose outcome changes quintile when treated. In (ref), we find qualitatively similar results using 2 or 10 bins.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Grandmother mechanism.} We first examine whether the effects of the intervention can be explained through the binary mechanism of whether a grandmother is present in the household (measured at the 7-year follow-up). (ref) shows estimates of $P(Y=y,M=0 \mid D=d)$ for both $d=1$ and $d=0$, similar to (ref) for our previous application. If one imposes monotonicity, then as derived in (ref) we should have that $P(Y=y, M=0 \mid D=1) \leq P(Y=y, M=0 \mid D=0)$ for all values of $y$. As shown in the figure, however, this inequality appears to be violated at large values of $y$, suggesting that the outcome for some never-takers improved when receiving the treatment. These violations of the sharp null are statistically significant (CS $p=$ \unskip). Our estimates of the lower bound derived in (ref) imply that at least \unskip percent of never-takers are affected by the treatment. Thus, we can reject that the entirety of the treatment effect operates through increased grandmother presence in the home.\footnote{Specifically, we can reject that the effect operates through increasing long-run grandmother presence, as measured at the 7-year follow-up. The results are less conclusive using the presence of a grandmother at the 1-year follow-up: we obtain $p=$\unskip, although the point estimates suggest that \unskip percent of never-takers are affected by treatment.} These conclusions rely on the monotonicity assumption that receiving CBT weakly increases the presence of the grandmother; this could be violated if, for example, some grandmothers were present when the mother was struggling but decided they were no longer needed as the mother improved. As before, we can explore robustness to allowing for defiers: our estimated lower bounds on the fraction of never-takers affected remain positive unless we allow for at least \unskip percent of the population to be defiers, or equivalently, \unskip defiers per complier.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Relationship quality mechanism.} We next examine relationship quality (as of the 7-year follow-up) as the mechanism, which is measured on a 1-5 scale. We can thus apply the methods for multi-valued $M$ developed in (ref). Under the monotonicity assumption that CBT improves the relationship with the husband, we reject the sharp null using CS ($p=$ \unskip); we obtain a point estimate of the lower bound on the fraction of always-takers affected (pooling across different values of $M$) of \unskip%.\footnote{The monotonicity assumption requires that the population CDF of $M \mid D=1$ is everywhere smaller than the population CDF of $M \mid D=0$. This is satisfied at three of the four support points of the empirical distribution. However, the empirical CDF in the treated group is \unskip larger at $M=4$, although this difference is not statistically significant from zero ($p$=\unskip). Thus, the empirical distribution violates monotonicity, although we cannot reject that monotonicity holds in the population. To compute our estimate of the lower bound on the fraction of always-takers affected using the empirical distribution, we therefore allow for the minimum number of defiers compatible with the empirical distribution of $M \mid D$ (\unskip). We apply an analogous approach when considering the grandmother and relationship-quality mechanisms jointly (using the minimal relaxation of the elementwise version of monotonicity).} There is thus some evidence that the effect of CBT on financial empowerment does not operate entirely through improvements in relationship quality. The lower bound on the fraction of always-takers affected remains positive allowing for up to \unskip% of the population to be defiers.
\@startsection {paragraph}{4}{\z@} {1.75ex \@plus .2ex \@minus .2ex} {-1em} {\normalfont}{Combinations of mechanisms.} Can the combination of the grandmother and relationship-quality mechanisms explain the improvement in financial empowerment? To evaluate this, we consider the case where $M$ is a vector containing both candidate mechanisms. Our estimated lower bound on the fraction of always-takers affected is \unskip%. However, our test of the sharp null, imposing the monotonicity assumption that treatment increases each of the elements of $M$, is not statistically significant at conventional levels (CS $p = $ \unskip). Thus, while the point estimate suggests a moderate violation of the sharp null, we do not significantly reject the null hypothesis that the combination of these two mechanisms, which the authors interpret broadly as proxies for “social support within the household”, can explain the effect of CBT on financial empowerment. This of course does not establish that no other mechanisms are at play, but rather that the data are statistically consistent with this null hypothesis at conventional levels.
This paper develops tests for the “sharp null of full mediation” that the effect of a treatment $D$ on an outcome $Y$ operates only through a conjectured set of mediators $M$. A key observation is that when $M$ is binary, existing tools for testing the validity of the LATE assumptions can be used for testing the null. We develop sharp testable implications in a more general setting that allows for multi-valued and multi-dimensional $M$, and allows for relaxations of the monotonicity assumption. Our results also provide lower bounds on the size of the alternative mechanisms for always-takers. We illustrate the usefulness of these tests in two empirical applications.
Future work might extend the analysis in this paper in several directions. First, our analysis focuses on the case where $M$ is discrete. Although one can discretize $M$ under the assumptions described in (ref), an interesting question for future work is whether one can impose alternative assumptions that allow for testing the sharp null directly when $M$ is continuous. One potentially fruitful direction is to explore whether methods for testing instrument validity with a continuous treatment dhaultfoeuille_testing_2024 can be adapted to this setting. Second, our current analysis allows the potential outcomes to depend arbitrarily on $M$, and does not impose any assumptions on how $M$ is assigned. In some settings, however, it may be reasonable to restrict the magnitude of the effect of $M$ on $Y$, or to restrict the degree of endogeneity of $M$. Incorporating such restrictions may lead to sharper testable implications. Finally, it may be interesting to extend our results to settings with non-binary treatments.
\FloatBarrier