Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,762 characters · 15 sections · 55 citation commands
Causal Mediation in Natural Experiments
\thispagestyle{empty}
\setcounter{page}{1} \onehalfspacing Economists use natural experiments to credibly answer social questions, when an experiment was infeasible. For example, does winning access to health insurance causally improve health and well-being finkelstein2008oregon? Natural experiments are settings which answer these questions, but give little indication of how these effects came about. Causal Mediation (CM) aims to estimate the mechanisms behind causal effects, by estimating how much of the treatment effect operates through a proposed mediator mechanism. For example, do causal gains from winning access to health insurance come mostly from physical use of healthcare, or plausible psychological gains from no longer having to worry about being uninsured? This study of mechanisms behind causal effects broadens the economic understanding of social settings studied with natural experiments. This paper shows that the conventional approach to estimating CM effects is inappropriate in a natural experiment setting, provides a theoretical framework for how bias operates, and develops an approach to correctly estimate CM effects under alternative assumptions. These methods contrast the current practice in applied economics of providing suggestive evidence of mechanisms, which gives necessary but not sufficient conditions for the mechanisms behind a treatment effect.
This paper starts by considering conventional CM methods in a natural experiment setting. Conventional CM methods rely on assuming the initial treatment, and the subsequent mediator mechanism, are both quasi-randomly assigned imai2010identification. Assuming the mediator is as-good-as-randomly assigned requires either (1) selection is fully captured by observabed control variables, or (2) that decisions are effectively random. The assumption can be plausible when a researcher has access to a very rich set of control variables, or mediator is directly manipulated. In economic applications, however, unconstrained take-up reflects choices based on costs and benefits, which complicates the argument for mediator quasi-random assignment. For example, winners of the Oregon Health Insurance Experiment wait-list were randomly chosen by a lottery, but made the choice to visit healthcare possibly factoring the effects of undiagnosed health conditions (unobserved by the researcher). In practice, the main observational setting where mediator quasi-random assignment becomes credible is when another natural experiment affected the mediator --- a rare occurrence given how difficult it is to find one source of random variation for a treatment, let alone another independent source for a mediator, at the same time.
Applied economics research often complements reduced-form causal estimates with suggestive evidence for mechanisms behind the treatment effect. These descriptive analyses are informative, but do not generally identify a mediating pathway without additional assumptions --- see also blackwell2024assumption,green2010enough. A new strand of the econometric literature has arisen in implicit acknowledgement that suggestive evidence of mechanisms, or a conventional approach to CM, can lead to biased inference and needs alternative methods for credible inference. These include identifying CM effects with overlapping quasi-experimental research designs deuchert2019direct,frolich2017direct, functional form restrictions heckman2015econometric,heckman2013understanding, partial identification flores2009identification, or a hypothesis test of full mediation through observed channels kwon2024testing --- see huber2019review for an overview.\footnote{ An alternative method to estimate CM effects is ensuring treatment and mediator quasi-random assignment holds by a running two randomised controlled trials for both treatment and mediator, at the same time. This set-up has been considered in the literature previously, in theory imai2013experimental and in practice ludwig2011mechanism. }
I develop a framework showing exactly how selection bias contaminates conventional CM estimates when mediator choices are driven by unobserved gains --- settings where none of the natural experiment research designs in the previously cited papers apply (i.e., the mediator is not ignorable). This provides a rigorous warning to applied economists against uncritically applying conventional CM methods to investigate mechanisms --- as is common in the fields of epidemiology, psychology, and medicine.
Selection based on costs and benefits is at odds with assuming a mediator is as-good-as randomly assigned in an observational setting, so I import methods grounded in labour economic theory to solve the identification problem. This approach identifies CM effects via the marginal effect of the mediator, and requires three main assumptions. (1) Mediator take-up must respond only positively to the initial treatment (monotonicity), which implies mediator selection follows a selection model. (2) Mediator take-up is motivated by mediator benefits. (3) A valid instrument for mediator take-up must exist, to avoid relying on parametric assumptions on unobserved selection. While these assumptions are strong, they are plausible in many applied settings. Mediator monotonicity aligns with conventional theories for selection-into-treatment, and is accepted widely in many applications using an instrumental variables research design. Selection based on costs and benefits is central to economic theory, and is the dominant concern for judging observational designs that identify causal effects. Access to valid instrumental variation is a strong condition, though is important to avoid further modelling assumptions; the most compelling example is using variation in mediator take-up costs as an instrument.
Applying the new methods to the Oregon Health Insurance Experiment shows that unobserved selection matters in an analysis of a real-world natural experiment. A substantial portion of the wait-list lottery's impact on subjective health and well-being is mediated indirectly through extra healthcare usage, after instrumenting for healthcare usage with respondents' usual provider. A conventional CM analysis would put this indirect mediated share at practically zero, so that my methods expose that negative selection into healthcare usage would be hiding evidence for this mechanism. These estimates replace claims on mechanisms with credible causal evidence that extra healthcare use mediates a sizeable share of the Medicaid-lottery's benefits, avoiding claims based only on suggestive conjecture.
The methods I propose for CM analyses are not perfect for every setting: the structural assumptions are strong, and are tailored to selection-into-mediator based on the economic principle of selection based on costs and benefits. Indeed, this approach provides no safe harbour for estimating CM effects if these structural assumptions do not hold true. This approach imports insights from the instrumental variables literature, connecting the influential imai2010identification approach to CM with the economics literature on selection-into-treatment and marginal treatment effects vytlacil2002independence,heckman2004using,heckman2005structural,florens2008identification,kline2019heckits. frolich2017direct have previously explored identification of CM effects with a control function in the context of two instruments (one each for treatment and mediator) and a continuous mediator. This paper considers CM effects via the marginal effect of a binary mediator, with a different identification analysis and estimation strategies.
This paper proceeds as follows. (ref) describes the dominant approach in economics for studying mechanisms behind treatment effects, illustrating with data from the Oregon Health Insurance Experiment. (ref) introduces the formal framework for CM, and develops expressions for bias in CM estimates in natural experiments. (ref) describes this bias in applied settings with (1) a regression framework, (2) a setting with selection based on costs and benefits. (ref) purges bias from CM estimates by identifying CM effects with a control function adjustment, via the marginal effect of the mediator. (ref) demonstrates how to estimate CM effects with this approach, with either parametric or semi-parametric methods, and gives simulation evidence. (ref) returns to the Oregon Health Insurance Experiment, providing credible estimates of effects on subjective health and well-being mediated through healthcare usage. (ref) concludes.
In the United States, healthcare is generally not provided directly by the government. Instead, consumers purchase health insurance to fund healthcare expenses, with the government providing insurance only for elderly individuals (Medicare) and for those with low-incomes (Medicaid). In 2004, the state of Oregon ceased accepting new applications for Medicaid due to budgetary constraints, and did not reopen applications until 2008. When the state resumed enrolment, 90,000 individuals applied, vastly exceeding the programme's capacity. Oregon therefore allocated the opportunity to apply for Medicaid via a lottery system among those on the wait-list. Winning this wait-list lottery significantly increased healthcare usage, plus self-reported health and well-being.
Winning the wait-list lottery increased the average health insurance coverage rate by 22 percentage points (pp), and self-reported visitation of any healthcare provider in the following 12 months by 5 pp. In addition, wait-list lottery winners agreed 6 pp more with the question “In general, would you say your health is excellent, very good, or good” (hereafter, self-reported health), and 5 pp for “How would you say things are these days-would you say that you are very or pretty happy” (hereafter, self-reported happiness or well-being). These numbers are calculated among the 11,126 people from the Oregon wait-list lottery who responded to a survey sent by finkelstein2008oregon one year later,\footnote{ This number restricts to those who gave non-missing answers to all relevant questions. } using anonymised data from the Oregon Health Insurance Experiment replication package icspr2014oregon. (ref) summarises these results.
These results show that winning the wait-list lottery led to large gains in subjective health and well-being. The economics, medicine, and health policy literatures have primarily focused on the health benefits --- often interpreted as healthcare benefits from new access to government provided health insurance.\footnote{ finkelstein2008oregon use the wait-list lottery as an IV because health insurance is not randomly assigned; this paper focuses on the average effects of winning the wait-list lottery (which is randomly assigned). } However, the original authors also noted other benefits, including complete elimination of catastrophic out-of-pocket medical debt among those with new access to Medicaid. These are plausibly income effects that benefit recipients directly, not only through increased use of healthcare, but also by reducing stress and improving financial security. These plausible direct effects have not been explored in the applied literature.
Accepted practice in applied economics is to investigate mechanisms behind causal effects with suggestive evidence. This involves estimating the average causal effect of the wait-lottery on a proposed mediator (healthcare usage) and separately estimating its effect on the final outcomes (self-reported health and well-being). When both estimates are positive, and the mediator precedes the outcome, it is taken as de facto evidence that the mediator transmits the treatment effect. In the case of the Oregon Health Insurance Experiment, this amounts to concluding that increased healthcare usage mediates the positive effects of winning the lottery on health and well-being. (ref) illustrates this approach, which is also prevalent in other social science fields --- see blackwell2024assumption,green2010enough.
This approach gives necessary, but not sufficient, identification of healthcare as a mediating mechanism. It is not sufficient because it provides no evidence for the effect of healthcare on health and well-being, so does not identify the causal mechanism. Studying this mechanism with suggestive evidence require an additional, hidden assumption that healthcare positively affects health outcomes. While this assumption is not unreasonable in general, it may not apply within the data under study; subjective well-being is measured only 12 months after the Medicaid lottery, which may not be enough time for health gains to accrue. Second, this approach does not quantify the mechanism effects. Healthcare could only have a very large effect on subjective health and well-being, or possibly a very large small effect --- it is a priori unclear. In addition, the mediatior mechanism effect refers not to the average effect of healthcare usage, but to the effect for Oregon residents who were induced to use more healthcare after winning the wait-list lottery (mediator compliers). This local effect could differ substantially from a population average, and potentially mislead conclusions about the magnitude or generality of the mechanism. Together, these concerns mean that a suggestive analysis of healthcare as a mediating mechanism are not dispositive; without additional assumptions, they neither identify nor quantify the mediator mechanism channel.
CM offers a compelling alternative framework, explicitly defining the average direct and indirect effects and clear assumptions under which they are identified. Moreover, it delivers quantitative answers to the key question: how much of a treatment effect operates through a specific mediator mechanism? CM is widely used in fields such as epidemiology, psychology, and medicine where researchers regularly decompose treatment effects into component pathways. However, CM methods have not yet been examined from an economic perspective to assess their applicability in observational causal research, such as natural experiments.
CM decomposes causal effects into two channels, through a mediating mechanism (indirect effect) and through all other paths (direct effect). To develop notation, write $Z_i = 0, 1$ for a binary treatment, $D_i = 0, 1$ a binary mediator mechanism, and $Y_i$ a continuous outcome for individuals $i = 1, \hdots, n$.\footnote{ This paper exclusively focuses on the binary case. See huber2020direct or frolich2017direct for a discussion of CM with continuous treatment and/or mediator, and the assumptions required. } $D_i$ and $Y_i$ are a sum of their potential outcomes,
Assume treatment $Z_i$ is quasi-randomly assigned,\footnote{ This assumption can hold conditional on a covariate vector, $\vec X_i$. To simplify notation in this section, leave the conditional part unsaid, as it changes no part of the identification framework. } \[ Z_i \, \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \, D_i(z'), Y_i(z, d'), \text{ for } z', z, d' = 0, 1. \]
There are only two average effects which are identified without additional assumptions.
$Z_i$ affects outcome $Y_i$ directly, and indirectly via the $D_i(Z_i)$ channel, with no reverse causality. (ref) visualises the design, where the direction arrows denote the causal direction. CM aims to decompose the ATE of $Z_i$ on $Y_i$ into these two separate pathways:
Estimating the AIE answers the following question: how much of the causal effect $Z_i$ on $Y_i$ goes through the $D_i$ channel? When studying the health gains of winning the Medicaid wait-list lottery finkelstein2008oregon, the AIE represents how much of the effect comes from using the hospital more often. Estimating the ADE answers the following equation: how much is left over after accounting for the $D_i$ channel?\footnote{ In a non-parametric setting it is not necessary that ADE $+$ AIE $=$ ATE. See imai2010identification for this point in full. } For the example, how much of the wait-list lottery effect is a direct effect, other than increased healthcare usage --- e.g., income effects of lower medical debt, or less worry over health shocks thanks to government support. An Instrumental Variables (IV) approach assumes this direct effect is zero for everyone (the exclusion restriction). CM is a similar, yet distinct, framework attempting to explicitly model the direct effect, and not assuming it is zero.
The ADE and AIE are not separately identified without further assumptions.
The conventional approach to estimating direct and indirect effects assumes both $Z_i$ and $D_i$ are quasi-randomly assigned, conditional on a vector of control variables $\vec X_i$.
Sequential quasi-random assignment assumes that the initial treatment $Z_i$ is quasi-randomly assigned conditional on $\vec X_i$ (as has already been assumed above). It then also assumes that, after $Z_i$ is assigned, that $D_i$ is quasi-randomly assigned conditional on $\vec X, Z_i$ (hereafter, mediator quasi-random assignment). If (ref)(ref) and (ref)(ref) hold, then the ADE and AIE are identified by two-stage mean differences conditioning on $\vec X_i$.\footnote{ In addition, a common support condition for both $Z_i, D_i$ (across $\vec X_i$) is necessary. imai2010identification show a general identification statement; I show identification in terms of two-stage regression. See \hyperref[appendix:identification]{Appendix \ref*{appendix:identification}}. }
\makebox[\textwidth]{\parbox{1.25\textwidth}{ \[ \mathbb{E}_{D_i, \vec X_i} \left[ \underbrace{\mathbb{E}_{} \left[ Y_i \, \middle\vert \, Z_i = 1, D_i, \vec X_i \right] - \mathbb{E}_{} \left[ Y_i \, \middle\vert \, Z_i = 0, D_i, \vec X_i \right]}_{\text{Second-stage regression, $Y_i$ on $Z_i$ holding $D_i, \vec X_i$ constant}} \right] = \underbrace{\mathbb{E}_{} \left[ Y_i(1, D_i(Z_i)) - Y_i(0, D_i(Z_i)) \right]}_{\text{Average Direct Effect (ADE)}} \] \[ \mathbb{E}_{Z_i, \vec X_i} \left[ \underbrace{\Big( \mathbb{E}_{} \left[ D_i \, \middle\vert \, Z_i = 1, \vec X_i \right] - \mathbb{E}_{} \left[ D_i \, \middle\vert \, Z_i = 0, \vec X_i \right] \Big)}_{\text{First-stage regression, $D_i$ on $Z_i$}} \times \underbrace{\Big( \mathbb{E}_{} \left[ Y_i \, \middle\vert \, Z_i, D_i = 1, \vec X_i \right] - \mathbb{E}_{} \left[ Y_i \, \middle\vert \, Z_i, D_i = 0, \vec X_i \right] \Big)}_{\text{Second-stage regression, $Y_i$ on $D_i$ holding $Z_i, \vec X_i$ constant}} \right] \] \[ = \underbrace{\mathbb{E}_{} \left[ Y_i(Z_i, D_i(1)) - Y_i(Z_i, D_i(0)) \right]}_{\text{Average Indirect Effect (AIE)}} \] }}
I refer to the estimands on the left-hand side as CM estimands, which are typically estimated by a composition of two-stage Ordinary Least Squares (OLS) estimates imai2010identification. While this is the most common approach in the applied literature, I do not assume the linear model for my identification analysis. Linearity assumptions are not necessary for identification, and it suffices to note that heterogeneous treatment effects and non-linear confounding can bias OLS estimates of CM estimands in the same manner that is well documented elsewhere (see e.g., angrist1998estimating,sloczynski2022interpreting). This section focuses on problems that plague conventional CM, regardless of estimation method.
Applied research often uses a natural experiment to study settings where treatment $Z_i$ is quasi-randomly assigned, justifying assumption (ref)(ref). Rarely do they also have access to an additional, overlapping natural experiment to isolate random variation in $D_i$ --- to justify mediator quasi-random assignment (ref)(ref). One might consider conventional CM methods in such a setting to learn about the mechanisms behind the causal effect $Z_i$ on $Y_i$. This approach leads to estimates at risk of bias, contaminating inference on direct and indirect effects.
Below I present the relevant selection bias and group difference terms, omitting the conditional on $\vec X_i$ notation for brevity.
For the direct effect: CM estimand $=$ ADE $+$ selection bias $+$ group differences.\footnote{ The bias terms here mirror those in heckman1998characterizing,angrist2009mostly for a single $D_i$ on $Y_i$ treatment effect, when $D_i$ is not quasi-randomly assigned: \[ \mathbb{E}_{} \left[ Y_i \, \middle\vert \, D_i =1 \right] - \mathbb{E}_{} \left[ Y_i \, \middle\vert \, D_i =0 \right] = \text{ATE} + \underbrace{\Big( \mathbb{E}_{} \left[ Y_i(.,0) \, \middle\vert \, D_i =1 \right] - \mathbb{E}_{} \left[ Y_i(.,0) \, \middle\vert \, D_i =0 \right] \Big)}_{ \text{Selection Bias}} + \underbrace{ \Pr\left( D_i=0 \right) (\text{ATT} - \text{ATU}) }_{ \text{Group-differences Bias}}. \] }
For the indirect effect: CM estimand $=$ AIE $+$ selection bias $+$ group differences.
The selection bias terms come from systematic differences between the groups taking or refusing the mediator ($D_i = 1$ versus $D_i = 0$), differences not fully unexplained by $\vec X_i$. These selection bias terms would equal zero if the mediator had been quasi-randomly assigned (ref)(ref), but do not necessarily average to zero if not. In the Oregon Health Insurance Experiment, the wait-list gave random variation in the treatment (the Medicaid wait-list lottery) but there was not a similar natural experiment for healthcare usage; correspondingly, the selection-on-observables approach to CM has selection bias.
The group differences represent the fact that a matching approach gives an average effect on the treated group, which is systematically different from the average effect if selection-on-observables does not hold. These terms are a non-parametric framing of the bias from controlling for intermediate outcomes, previously studied only in a linear setting (i.e., bad controls in cinelli2024crash, or M-bias in ding2015adjust).
The AIE group differences term is longer, because the indirect effect is comprised of the effect of $D_i$ local to $D_i(Z_i)$ compliers.
It is important to acknowledge the mediator compliers here, because the AIE is the treatment effect going through the $D_i(Z_i)$ channel, thus only refers to individuals pushed into mediator $D_i$ by initial treatment $Z_i$. If we had been using a population average effect for $D_i$ on $Y_i$, then this is losing focus on the definition of the AIE; it is not about the causal effect $D_i$ on $Y_i$, it is about the causal effect $D_i(Z_i)$ on $Y_i$.
The group difference bias term arises because the selection-on-observables approach assumes that this complier average effect is equal to the population average effect, which does not hold true if the mediator is not quasi-randomly assigned.
Unobserved confounding is particularly problematic when studying the mechanisms behind treatment effects. For example, in studying health gains from the Oregon wait-list lottery, we might expect that health gains came about because those who won access to Medicaid started visiting their healthcare provider more often, when in past they avoided it over financial concerns. Applying conventional CM methods to investigate this expectation would be dismissing unobserved confounders for how often individuals visit healthcare providers, leading to biased results.
The wider population does not have one uniform bill of health; many people are born predisposed to ailments, due to genetic variation or other unrelated factors. These conditions can exist for years before being diagnosed. People with severe underlying conditions may visit healthcare providers more often than the rest of the population, to investigate or begin treating the ill-effects. It stands to reason that people with more serve underlying conditions may gain more from more often attending healthcare providers once given health insurance. These underlying causes cannot be controlled for by researchers, as we cannot hope to observe and control for health conditions that are yet to even be diagnosed. This means underlying health conditions are an unobserved confounder, and will bias estimates of the ADE and AIE in this setting.
In this section, I further develop the issue of selection on unobserved factors in a general CM setting. First, I show the non-parametric bias terms from (ref) can be written as omitted variables bias in a random coefficients regression framework. Second, I show how selection bias operates in a basic model for selection-into-mediator based on costs and benefits.
Inference for CM effects can be written in a regression framework with random coefficients, showing how correlation between unobserved error terms and the mediator disrupts identification.
Start by writing potential outcomes $Y_i(., .)$ as a sum of observed and unobserved factors, following the notation of heckman2005structural. For each $z',d' = 0,1$, put $\mu_{d'}(z'; \vec X_i) = \mathbb{E}_{} \left[ Y_i(z', d') \, \middle\vert \, \vec X_i \right]$ and the corresponding error terms, $U_{d', i} = Y_i(z', d') - \mu_{d'}(z'; \vec X_i)$, so we have the following expressions: \[ Y_i(Z_i, 0) = \mu_{0}(Z_i; \vec X_i) + U_{0,i}, \;\; Y_i(Z_i, 1) = \mu_{1}(Z_i; \vec X_i) + U_{1,i}. \]
With this notation, observed data $Z_i, D_i, Y_i, \vec X_i$ have the following random coefficient outcome formulae --- which characterise direct effects, indirect effects, and selection bias.
This is not consequence of linearity assumptions; the outcome formulae allow for unconstrained heterogeneous treatment effects, because the coefficients are random. If either $Z_i, D_i$ were continuously distributed, then this representative would not necessarily hold true. First-stage (ref) is identified, with $\theta + \zeta(\vec X_i)$ the intercept, and $\bar \pi$ the first-stage average compliance rate (conditional on $\vec X_i$). Second-stage (ref) has the following definitions, and is not identified thanks to omitted variables bias. See \hyperref[appendix:regression-model]{Appendix \ref*{appendix:regression-model}} for the derivation.
The ADE and AIE are averages of the random coefficients:
The ADE is a simple sum of the coefficients, while the AIE includes a group differences term because it only refers to $D_i(Z_i)$ compliers.
By construction, $\vec U_i \coloneqq \left(U_{0, i}, U_{1, i} \right)$ is an unobserved confounder. The regression estimates of $\beta, \gamma, \delta$ in second-stage (ref) give unbiased estimates only if $D_i$ is also conditionally quasi-randomly assigned: $D_i \, \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \, \vec U_{i} $. If not, then estimates of CM effects suffer from omitted variables bias from failing to adjust for the unobserved confounder, $\vec U_i$.
CM is at risk of bias because $D_i \, \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \, \vec U_i$ is unlikely to hold in applied settings. A separate identification strategy could disrupt the selection-into-$D_i$ based on unobserved factors, and lend credibility to the mediator quasi-random assignment assumption. Without it, bias will persist, given how we conventionally think of selection-into-treatment.
Consider a model where individual $i$ selects into a mediator based on costs and benefits (in terms of outcome $Y_i$), after $Z_i, \vec X_i$ have been assigned. In a natural experiment setting, an external factor has disrupted individuals selecting $Z_i$ by choice (thus $Z_i$ is quasi-randomly assigned), but it has not disrupted the choice to take mediator (thus $D_i$ is not quasi-randomly assigned). In the Oregon Health Insurance Experiment, the treatment variation comes from the wait-list lottery, while healthcare usage was not subject to a similar lottery. Write $C_i$ for individual $i$'s costs of taking mediator $D_i$, and $\mathds{1}\left\{ . \right\}$ for the indicator function. The Roy model has $i$ taking the mediator if the benefits exceed the costs,
The Roy model provides an intuitive framework for analysing selection mechanisms because it captures the fundamental economic principle of decision-making based on costs and benefits in terms of the outcome under study roy1951some,heckman1990empirical. In the Oregon Health Insurance Experiment, this models choice to visit the doctor in terms of health and well-being benefits relative to costs.\footnote{ If the choice is considers over a sum of outcomes, then a simple extension to a utility maximisation model maintains this same framework with expected costs and benefits. See heckman1990empirical,eisenhauer2015generalized. } This makes the Roy model useful as a base case for CM, where selection-into-mediator may be driven by private information (unobserved by the researcher).
By using the Roy model as a benchmark, I explore the practical limits of the mediator quasi-random assignment assumption. If selection follows a Roy model, and the mediator is quasi-randomly assigned, then unobserved benefits can play no part in selection. The only driver of selection are individuals' differences in costs (and not benefits). If there are any selection-into-$D_i$ benefits unobserved to the researcher, then mediator quasi-random assignment cannot hold. \newtheorem{proposition}{Proposition}
This is an equivalence statement: selection based on costs and benefits is only consistent with mediator quasi-random assignment if the researcher observed every single source of mediator benefits. See \hyperref[appendix:roy-seq-ig]{Appendix \ref*{appendix:roy-seq-ig}} for the proof. This means than the vector of control variables $\vec X_i$ must be incredibly rich. Together, $\vec X_i$ and unobserved cost differences $U_{C,i}$ must explain selection-into-$D_i$ one hundred percent. In the Roy model framework, however, individuals make decisions about mediator take-up based on gains --- whether the researcher observes them or not. The unobserved gains are unlikely to be fully captured by an observed control set $\vec X_i$, except in very special cases.
In practice, the best setting to believe in the mediator quasi-random assignment assumption is to study a setting where the researcher has two causal research designs, one for treatment $Z_i$ and another for mediator $D_i$, at the same time. Absent this, mediator quasi-random assignment become hard to believe, and the corresponding conventional CM estimates are at risk of selection bias.
If your goal is to estimate CM effects, and you could control for unobserved selection terms $U_{0,i}, U_{1,i}$, then you would. This ideal (but infeasible) scenario would yield unbiased estimates for the ADE and AIE. A Control Function (CF) approach takes this insight seriously, providing conditions to model the implied confounding by $U_{0,i}, U_{1,i}$, and then controlling for it.
The main problem is that second-stage regression equation (ref) is not identified, because $U_{0,i},U_{1,i}$ are unobserved, and lead to omitted variables bias.
The CF approach models the contaminating terms in (ref), avoiding the bias from omitting them in regression estimates. CF methods were first devised to correct for sample selection problems heckman1974shadow, and were extended to a general selection problem of the same form as (ref) heckman1979sample. The approach works in the following manner: (1) assume that the variable of interest follows a selection model, where unexplained first-stage selection informs unobserved second-stage confounding; (2) extract information about unobserved confounding from the first-stage; and (3) incorporate this information as control terms in the second-stage equation to adjust for selection-into-mediator. Identification in CF methods typically relies an external instrument or distributional assumptions; the identification strategy here focuses exclusively on the case that an instrument is available. By explicitly accounting for the information contained in the first-stage selection model, CF methods enable consistent estimation of causal effects in the second-stage even when selection is driven by unobserved factors florens2008identification.
In the example of analysing health gains from the Oregon Health Insurance Experiment, a CF approach addresses the unobserved confounding by modelling unobserved effects of underlying health conditions. It does so by assuming that unobserved selection-into-healthcare use is informative for underlying health conditions, assuming people with more severe underlying conditions visit the doctor more often than those without. Then it uses this information in the second-stage estimation of how much the effect goes through increased healthcare usage, estimating the ADE and AIE after controlling for this confounding.
The following assumptions are sufficient to model the correlated error terms, identifying $\beta, \gamma, \delta$ in the second-stage regression (ref), and thus both the ADE and AIE.
\theoremstyle{definition} \newtheorem{assumptionCF}{Assumption}
Assumption (ref) is the monotonicity condition first used in an IV context imbens1994identification. Here, it is assuming that people respond to treatment, $Z_i$, by consistently taking or refusing the mediator $D_i$ (always or never-mediators), or taking the mediator $D_i$ if and only if assigned to the treatment $Z_i=1$ (mediator compliers). There are no mediator defiers.
The main implication of Assumption (ref) is that selection-into-mediator can be written as a selection model with ordered threshold crossing values that describe selection-into-$D_i$ vytlacil2002independence. \[ D_i(z') = \mathds{1}\left\{ V_i \leq \psi \big( z'; \vec X_i \big) \right\}, \;\text{ for } z'=0,1 \] where $V_i$ is a latent variable with continuous distribution and conditional cumulative density function $F_V(. \,|\vec X_i)$, and $\psi(. \,;\vec X_i)$ collects observed sources of mediator selection. $V_i$ could be assumed to follow a known distribution; the canonical Heckman selection model assumes $V_i$ is normally distributed (a “Heckit” model). The identification strategy here applies to the general case that the distribution of $V_i$ is unknown, without parametric restrictions.
I focus on the equivalent transformed model of heckman2005structural, \[ D_i(z') = \mathds{1}\left\{ U_i \leq \pi(z'; \vec X_i) \right\}, \;\;\; \text{for } z'=0,1 \] where $U_i \coloneqq F_V\left( V_i \mid \vec X_i \right)$ follows a uniform distribution, and $\pi(z'; \vec X_i) = F_V\big(\psi(z'; \vec X_i)\big) = \Pr\left( D_i = 1 \, \middle\vert \, Z_i = z', \vec X_i \right)$ is the mediator propensity score. $U_i$ are the unobserved mediator take-up costs. Note the maintained assumption that treatment $Z_i$ is quasi-randomly assigned conditional on $\vec X_i$ implies $Z_i \, \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \, U_i$ conditional on $\vec X_i$.
This selection model setup is equivalent to the monotonicity condition, and is importing a well-known equivalence result from the IV literature to the CM setting. The main conceptual difference is not assuming $Z_i$ is a valid instrument for identifying the $D_i$ on $Y_i$ effect among compliers; it is using the selection model representation to correct for selection bias. See \hyperref[appendix:cf-monotonicity]{Appendix \ref*{appendix:cf-monotonicity}} for a validation of the general vytlacil2002independence equivalence result in a CM setting, with conditioning covariates $\vec X_i$.
Assumption (ref) is stating that unobserved selection in mediator take-up ($U_i$) informs second-stage confounding, when refusing or taking the mediator ($U_{0,i}$ and $U_{1,i}$). If there is unobserved confounding in $Y_i$, then it can be measured in $D_i$.
This is a strong assumption, and will not hold in all examples. If people had been deciding to take $D_i$ by a Roy model, then this assumption holds because $V_i = U_{C,i} - \big( U_{1,i} - U_{0,i} \big)$. Individuals could be making decisions based on other outcomes, but as long as mediator benefits guide at least part of this decision (i.e., bounded away from zero), then this assumption will hold.
For notation purposes, suppose the vector of control variables $\vec X_i$ has at least two entries; denote $\vec X_i^{\text{IV}}$ as one entry in the vector, and $\vec X_i^-$ as the remaining.
Assumption (ref) is requiring at least one control variable guides selection-into-$D_i$ --- an IV. It assumes an instrument exists, which satisfies an exclusion restriction (i.e., not impacting mediator gains $\mu_1-\mu_0$), and has a non-zero influence on the mediator (i.e., strong IV first-stage). The exclusion restriction is untestable, and must be guided by domain-specific knowledge; IV first-stage strength is testable, and must be justified with data by methods common in the IV literature.
This assumption identifies the mediator propensity score separately from the direct and indirect effects, avoiding indeterminacy in the second-stage outcome equation. While not technically required for identification, it avoids relying entirely on an assumed distribution for unobserved error terms (and bias from inevitably breaking this assumption). The most compelling example of a mediator IV is using data on the cost of mediator take-up as a first-stage IV, if it varies between individuals for unrelated reasons and is strong in explaining mediator take-up.
Again, this set-up required no linearity assumptions, and treatment effects vary, because $Z_i, D_i$ are categorical and $\beta, \gamma, \delta , \varphi(\vec X_i)$ vary with $\vec X_i$. The CFs are functions which measure unobserved mediator gains, for those with unobserved mediator costs above or below a propensity score value. Following the IV notation of kline2019heckits, put $\mu_V = \mathbb{E}_{} \left[ F_V^{-1}\left( U_i \mid \boldsymbol{\mathit{X}}_i \right) \right]$, to give the following representation for the CFs:
All relevant parameters --- $\alpha, \beta, \gamma, \delta, \varphi(.)$ --- are identified once we control for selection bias through the CFs $\lambda_0, \lambda_1$, with $\pi(z';\vec X_i)$ identified separately in the first-stage thanks to the instrument(s) $\vec X_i^{\text{IV}}$. In the case that the CFs have an assumed functional form, then identification is complete. For example, in the canonical Heckman selection model, the error terms follow a normal distribution, so that $\lambda_0, \lambda_1$ are the inverse Mills ratio. If we do not know want to assume a specific distribution, then $\lambda_0, \lambda_1$ can be estimated separately with semi-parametric methods to avoid relying on parametric assumptions.\footnote{ This comes at the cost of $\alpha$ and $\rho_0, \rho_1$ no longer being separately identified from $\lambda_0, \lambda_1$ --- ultimately, however, this does not jeopardise identification and estimation of the ADE and AIE (see (ref)). }
This identification strategy is a Marginal Treatment Effect approach (MTE, bjorklund1987estimation,heckman2005structural) applied to a CM setting. One can see this by noting the connection to the marginal effect of the mediator,
The marginal effect of the mediator is identified under the (ref) assumptions, thanks to instrumental variation in $\vec X_i^{\text{IV}}$. The final step uses the corresponding CFs to extrapolate from $\vec X_i^{\text{IV}}$ compliers to mediator compliers, and thus identify the ADE and AIE.
\newtheorem{theoremCF}{Theorem}
This theorem provides a solution to the identification problem for CM effects when facing selection; rather than assuming away selection problems, it explicitly models them. The ADE is straightforward to calculate as an average of the direct effect parameters, while the AIE also includes an adjustment for unobserved complier gains to the mediator. Again, this is because the AIE only refers to individuals who were induced by treatment $Z_i$ into taking mediator $D_i$ (mediator compliers). The CFs allow us to measure both selection bias and complier differences, and thus purge persistent bias in identifying CM effects.
The ideal instrument $\vec X_i^{\text{IV}}$ for identification is continuous, and varies $\pi(z'; \vec X_i)$ between 0 and 1 for every possible value of $z', \vec X_i^-$ (identification at infinity). In practice, it is unlikely to find IV(s) that satisfy this condition. In this case, the brinch2017beyond restricted approach can be used --- even with a categorical instrument and no control variables. This amounts to assuming a limited specification for the respective CFs, limiting the number of parameters used to approximate $\lambda_0, \lambda_1$ to the number of discrete values that $\pi(z'; \vec X_i)$ takes minus one. E.g., if there are no control variables and $\vec X_i^{\text{IV}}$ is binary, then $\lambda_0, \lambda_1$ can only be identified up to 3 parameters each.\footnote{ The value of 3 comes from the cases that $Z_i, \vec X_i^{\text{IV}}$ each could take 2 values, so $\pi(z'; \vec X_i)$ has 4 possible values, giving the semi-parametric identification (and estimation) of each CF only 3 degrees of freedom to work with. If the CFs are instead assumed to have a known distribution (i.e., parametric CF), then those concerns do not matter. } Ultimately, this changes little to the identification strategy, and little to the estimation.
In a simulation with Roy selection-into-mediator based on unobserved error terms, the CF adjustment pushes conventional CM estimates back to the true value. (ref) shows how a CF adjustment corrects unadjusted CM effect estimates.
A conventional approach to estimating CM effects involves a two-stage approach to estimating the ADE and the AIE: the first-stage ($Z_i$ on $D_i$), and the second-stage ($Z_i, D_i$ on $Y_i$). A CF approach is a simple and intuitive addition to this approach: including the CF terms $\lambda_0, \lambda_1$ in the second-stage regression to address selection-into-mediator.
This section presents two practical estimation strategies. First, I demonstrate how to estimate CM effects with an assumed distribution of error terms, focusing on the Heckman selection model as the leading case. Second, I consider a more flexible semi-parametric approach that avoids distributional assumptions --- at the cost of semi-parametrically estimating the corresponding CFs. While both methods effectively address the selection bias issues detailed in previous sections, they differ in their implementation complexity, efficiency, and underlying assumptions.
A parametric CF solves the identification problem by assuming a distribution for the unobserved error terms in the first-stage selection model, and modelling selection based on this distribution. The Heckman selection model is the most pertinent example, assuming the normal distribution for unobserved errors heckman1979sample. A parametric CF using other distributions works in exactly the same manner, replacing the relevant density functions for an alternative distribution as needed. As such, this section focuses exclusively on the Heckman selection model. This estimation approach is the same as in the parametric selection model approach to MTEs, in bjorklund1987estimation.
The Heckman selection model assumes unobserved errors $V_i$ follow a normal distribution, so estimates the first-stage using a probit model. \[ \Pr\left( D_i = 1 \, \middle\vert \, Z_i, \vec X_i \right) = \Phi \big( \theta + \bar\pi Z_i + \vec\zeta' \vec X_i \big), \] where $\Phi(.)$ is the cumulative density function for the standard normal distribution, and $\theta, \bar\pi, \vec\zeta$ are parameters estimated with maximum likelihood. In the parametric case, an excluded instrument ($\vec X_i^{\text{IV}}$) is not technically necessary in the first-stage equation --- though not including one exposes the method to indeterminacy if the errors are not normally distributed. Thus, it is best practice to use this method with access to an instrument.
From this probit first-stage, construct the inverse Mills ratio terms to serve as the CFs. These terms capture the correlation between unobserved factors influencing both mediator selection and outcomes, when the errors are normally distributed. \[ \lambda_0(p') = \frac{\phi( - \Phi^{-1}(p') )}{\Phi( -\Phi^{-1}(p') )}, \;\;\;\; \lambda_1(p') = \frac{\phi( \Phi^{-1}(p') )}{\Phi( \Phi^{-1}(p') )}, \;\;\;\; \text{ for } p' \in (0,1) \] where $\phi(.)$ is the probability density function for the standard normal distribution.
Lastly, the second-stage is estimated with OLS, including the CFs with plug in estimates of the mediator propensity score, and $\vec \varphi'$ a linear approximation of nuisance function $\varphi(.)$.
where $\hat\pi \big(z';\vec X_i \big)$ are the predictions from the probit first-stage.
The resulting ADE and AIE estimates are composed from sample estimates of the terms in Theorem (ref), \[ \widehat{\text{ADE}} = \widehat{\gamma} + \widehat{\delta}\,\bar D, \;\;\;\; \widehat{\text{AIE}} = \widehat{\bar\pi}\; \Big( \widehat{\beta} + \widehat{\delta}\,\bar Z + \big(\hat \rho_1 - \hat \rho_0 \big) \frac1n \sum_{i=1}^n \Gamma \big( \hat\pi(0;\vec X_i), \hat\pi(1;\vec X_i)\big) \Big) \] where $\bar D = \frac1n \sum_{i=1}^n D_i$, $\bar Z = \frac1n \sum_{i=1}^n Z_i$, $\widehat{\bar\pi}$ is the estimate of the mean compliance rate, and $\frac1n \sum_{i=1}^n \Gamma(.,.)$ is the average of the complier adjustment term as a function of $\lambda_1$ with $\hat\pi \big(0; \vec X_i \big), \hat\pi \big(1; \vec X_i \big)$ values plugged in.
The standard errors for estimates can be computed using the delta method. Specifically, accounting for sampling variability in both first-stage mediator propensity score estimation and second-stage causal effects estimation. This approach yields $\sqrt{n}$-consistent estimates when the underlying error terms follow a bivariate normal distribution --- i.e., when $\pi(Z_i; \vec X_i)$ is correctly modelled by the probit first-stage. Errors can also be estimated by the bootstrap, by including estimation of both the first and second-stage within each bootstrap iteration.
In practice, a parametric CF approach is simple to implement using standard statistical packages. The key advantage is computational simplicity and efficiency, particularly in moderate-sized samples. However, this comes at the cost of strong distributional assumptions. For example, if the error terms deviate substantially from joint normality, the estimates may be biased.\footnote{ While this concern is immaterial in an IV setting estimating the LATE kline2019heckits, it is pertinent in this setting as the CF extrapolates from IV compliers to mediator compliers. }
For settings where researchers are not comfortable specifying a specific distribution for the error terms, a semi-parametric CF will nonetheless consistently estimate CM effects. This method maintains the same identification strategy but avoids assuming a specific error distribution. This estimation approach is used in the modern semi-parametric approach to estimating MTEs, for example in brinch2017beyond,heckman2007econometric.
The semi-parametric approach begins with flexible estimation of the first-stage, estimating the mediator propensity score, \[ \Pr\left( D_i = 1 \, \middle\vert \, Z_i, \vec X_i \right) = \pi\left(Z_i; \vec X_i \right), \] where $\vec X_i$ must include the instrument(s) $\vec X_i^{\text{IV}}$. This can be estimated using flexible methods, as long as the first-stage is estimated $\sqrt{n}$-consistently.\footnote{ If an estimate of the first-stage that is not $\sqrt{n}$-consistent is used (e.g., a modern machine learning estimator), then the resulting second-stage estimate will not be $\sqrt{n}$-consistent. } An attractive option is the klein1993efficient semi-parametric binary response model, which avoids relying on an assumed distribution of first-stage errors though requires a linear specification. If it is important to avoid a linear specification, then a probability forest avoids linearity assumptions athey2019generalized --- though is best used for cases with many columns in the $\vec X_i$ variables.
The second-stage is estimated with semi-parametric methods. Consider the subsamples of mediator refusers and takers separately,
The separated subsamples can be estimated, each individually, with semi-parametric methods. The linear parameters (including a linear approximation $\vec \varphi'$ of nuisance function $\varphi(.)$)\footnote{ Appropriate interactions between $Z_i, D_i$ and $\vec X_i$ can also flexibly control for $\vec X_i$, again avoiding linearity assumptions. } can be estimated with OLS, while $\rho_0 \lambda_0$ and $\rho_1 \lambda_1$ take a flexible semi-parametric specification with first-stage estimates $\hat \pi(Z_i; \vec X_i)$ plugged in. An attractive option is a series estimator, such as a spline specification, as this estimates the function without assuming a functional form but maintains $\sqrt n$-consistency.
The ADE is estimated by this approach as follows. Take $\hat \gamma$, the $D_i = 0$ subsample estimate of $\E \gamma$, and $(\widehat{\gamma + \delta})$, the $D_i = 1$ subsample estimate of $\mathbb{E}_{} \left[ \gamma + \delta \right]$, to give \[ \widehat{\text{ADE}}^{\text{CF}} = (1 - \bar D) \, \hat \gamma + \bar D \, (\widehat{\gamma + \delta}). \]
The AIE is less simple, for two reasons that differ from the parametric CF setting. First, the intercepts for each subsample, $\alpha$ and $(\alpha + \beta)$, are not separately identified from the CFs if the $\lambda_0, \lambda_1$ functions are flexibly estimated. Second, a semi-parametric specification for the CFs mean $\rho_0$ and $\lambda_0$ are no longer separately identified from each other (and same for $\rho_1,\lambda_1$). As such, it is not possible to directly use $\hat \lambda_0, \hat \lambda_1$ in estimating the complier adjustment term (as is done in the parametric case).
These problems can be avoided by estimating the AIE using its relation to the ATE. Write $\widehat{\text{ATE}}$ for the point-estimate of the ATE, and $\hat\delta = (\widehat{\gamma + \delta}) - \hat\gamma$ for the point estimate of $\E \delta$, to give the following estimator, \[ \widehat{\text{AIE}}^{\text{CF}} = \widehat{\text{ATE}} - (1 - \bar Z) \, \left( \frac 1n \sum_{i = 1}^n \hat\gamma + \hat \delta \, \hat\pi(1; \vec X_i) \right) - \bar Z \, \left( \frac 1n \sum_{i = 1}^n \hat\gamma + \hat \delta \, \hat\pi(0; \vec X_i) \right), \] where $\frac 1n \sum_{i = 1}^n \hat\gamma + \hat \delta \, \hat\pi(0; \vec X_i)$ estimates the ADE conditional on $Z_i = 0$, $\mathbb{E}_{} \left[ \gamma + \delta D_i(0) \right]$, and $\frac 1n \sum_{i = 1}^n \hat\gamma + \hat \delta \, \hat\pi(1; \vec X_i)$ estimates the ADE conditional on $Z_i = 1$, $\mathbb{E}_{} \left[ \gamma + \delta D_i(1) \right]$. \hyperref[appendix:semi-parametric]{Appendix \ref*{appendix:semi-parametric}} describes the reasoning for this estimator of the AIE, relative to estimates of the ATE and ADE, in further detail.
This semi-parametric approach achieves valid estimation of the CM effects, without specifying the distribution behind unobserved error terms, and achieves desirable properties as long as the first-stage correctly estimates the mediator propensity score, and the structural assumptions hold true. The standard errors for estimates can again be computed using the delta method, or estimated by the bootstrap --- again, across both first and second-stages within each bootstrap iteration. Note that relying on propensity score estimation requires assumptions that can be found wanting in real-world settings; a common support condition for the mediator is required, and a semi-/non-parametric first-stage may become cumbersome if there are many control variables or many rows of data.
The following simulation gives an example to show how these methods work in practice. Suppose data observed to the researcher $Z_i, D_i, Y_i, \vec X_i$ are drawn from the following data generating processes, for $i = 1, \hdots, N$, with $n = 5,000$ for this simulation. \[ Z_i \sim \text{Binom}\left(0.5 \right), \;\; \vec X_i^- \sim N(4, 1), \;\; \vec X_i^{\text{IV}} \sim \text{Uniform}\left( -1, 1 \right), \;\; \left( U_{0,i}, U_{1,i}, U_{C,i} \right) \sim N\left( \vec 0, \mat \Sigma \right) \] $\mat \Sigma$ is the matrix of parameters which controls the level of confounding from unobserved costs and benefits.\footnote{ The correlation and relative standard deviations for $U_{0,i}, U_{1,i}$ affect how large selection bias in conventional CM estimates; correlation for these with unobserved costs $U_{C,i}$ does not particularly matter, though increased variance in unobserved costs makes estimates less precise for both OLS and CF methods. }
Each $i$ chooses to take mediator $D_i$ by a Roy model, with following mean definitions for each $z', d' = 0, 1$
Following (ref), these data have the following first and second-stage equations:
Treatment $Z_i$ has a causal effect on outcome $Y_i$, and it operates partially through mediator $D_i$. Outcome mean $\mu_{D_i}\left( Z_i; \vec X_i \right)$ contains an interaction term, $Z_i D_i$, so while $Z_i,D_i$ have constant partial effects, the ATE depends on how many $i$ choose to take the mediator and there is treatment effect heterogeneity.
After $Z_i$ is assigned, $i$ chooses to take mediator $D_i$ by considering the costs and benefits --- which vary based on $Z_i$, demographic controls $\vec X_i$, and the (non-degenerate) unobserved error terms $U_{i,0}, U_{1,i}$. As a result, sequential quasi-random assignment does not hold; the mediator is not conditionally ignorable. Thus, a conventional approach to CM does not give an estimate for how much of the ATE goes through mediator $D$, but is contaminated by selection bias thanks to the unobserved error terms.
I simulate this data generating process 10,000 times, using $\mat\Sigma = \left(
\right)$,\footnote{ This choice of parameters has $Var_ \left( U_{0,i} \right) = 1, Var_ \left( U_{1,i} \right) = 2.25, Corr\big(U_{0,i}, U_{1,i}\big) = 0.5$ so that unobserved errors meaningfully confound conventional CM methods, with notable heteroscedasticity. Unobserved costs are uncorrelated with $U_{0,i}, U_{1,i}$ (although non-zero correlation would not meaningfully change the results), and $Var_ \left( U_{C,i} \right) = 0.25$ maintains uncertainty in unobserved costs. } and estimate CM effects with conventional CM methods (two-stage OLS) and the introduced CF methods. In this simulation $\Pr\left( D_i = 1 \right) = 0.379$, and $65.77%$ of the sample are mediator compliers (for whom $D_i(0)=0$ and $D_i(1)=1$). This gives an ATE value of 2.60, ADE 1.38, and AIE 1.22, respectively.\footnote{ Note that ATE $=$ ADE $+$ AIE in this setting. $\Pr\left( Z_i=1 \right) = 0.5$ ensures this equality, but it is not guaranteed in general. See \hyperref[appendix:semi-parametric]{Appendix \ref*{appendix:semi-parametric}}. }
(ref) shows how these estimates perform, with a parametric CF approach, relative to the true value. The OLS estimates' distribution do not overlap the true values for any standard level of significance; the distance between the OLS estimates and the true values are the underlying bias terms derived in (ref). The parametric CF approach perfectly reproduces the true values, as the probit first-stage correctly models the normally distributed error terms. The semi-parametric approach (not shown in (ref)) performs similarly, with a wider distribution; this is to be expected comparing a correctly specified parametric model with a semi-parametric one.
The parametric CF may not be appropriate in setting with non-normal error terms. I simulated the same data again, but transform $U_{0,i}, U_{1,i}$ to be correlated uniform errors (with the same standard deviations as previously). (ref) shows the resulting distribution of point-estimates, relative to the truth, for the parametric and semi-parametric approaches. The parametric CF is slightly off target, showing persistent bias from incorrectly specifying the error term distribution. The semi-parametric approach is centred exactly around the truth, with a slightly higher variance (as is expected).
The error terms determine the bias in OLS estimates of the ADE and AIE, so the bias varies for different values of the error-term parameters $\text{Corr}\big(U_{0,i}, U_{1,i}\big) \in [-1, 1]$ and $\text{Var}_{} \left( U_{0,i} \right), \text{Var}_{} \left( U_{1,i} \right) \geq 0$. The true AIE values vary, because $D_i(Z_i)$ compliers have higher average values of $U_{1,i} - U_{0,i}$ as $\text{Corr}\big(U_{0,i}, U_{1,i}\big)$ increases. (ref) shows CF estimates against estimates calculated by standard OLS, showing 95% confidence intervals calculated from 1,000 bootstraps. The point estimates of the CF do not exactly equal the true values, as they are estimates from one simulation (not averages across many generated datasets, as in (ref)). The CF approach improves on OLS estimates by correcting for bias, with confidence regions overlapping the true values.\footnote{ In the appendix, (ref) shows the same simulation while varying $\text{Var}_{} \left( U_{1,i} \right)$, with fixed $\text{Var}_{} \left( U_{0,i} \right) = 1, \text{Corr}\big(U_{0,i}, U_{1,i}\big) = 0.5$. The conclusion is the same as for varying the correlation coefficient, $\rho$, in (ref). } This correction did not come for free: the standard errors are significantly greater in a CF approach than OLS. In this manner, this simulation shows the pros and cons of using the CF approach to estimating CM effects in practice.
In the Oregon Health Insurance Experiment, winning the wait-list lottery significantly improved subjective health and happiness among participants. This study investigates the mechanisms behind these benefits, quantifying the extent to which improvements are mediated through increased healthcare usage.
To credibly address mediation concerns, I use the respondents' regular healthcare provider type as an IV for healthcare usage. Approximately 75.4% reported visiting a healthcare provider within the past year, but rates vary notably depending on their usual provider type: those attending hospital emergency rooms (A&E) and urgent care clinics reported significantly lower visitation rates (12.6 and 20 percentage points lower, respectively) than the 40% attending private clinics.\footnote{ The combined $F$ statistic for the categorical variable for healthcare usual location on healthcare usage is 124. } The IV validity arises from differential costs faced by individuals based on their usual care provider. Private clinics generally charge through health insurance and are more expensive without coverage, while A&E and urgent care often provide costly services but rarely follow up on unpaid bills, effectively creating variation in healthcare attendance costs. Additionally, individuals' choice of provider likely depend on neighbourhood-based access.
Initial results with unadjusted CM estimates suggest almost no mediating role for healthcare usage; the unadjusted estimates of the AIE are close to zero for both outcomes, contradicting intuitive suggestive evidence. These estimates remained robust when controlling for serious health conditions, such as kidney disease or diabetes, thus reinforcing the initially surprising conclusion. However, applying my CF methods reveals a much larger, positive AIE, restoring the mediating role of healthcare usage in line with suggestive intuition. This is because a correlational estimate of health and well-being gains to healthcare visits are practically zero, while the IV estimates restore positive gains and the CF methods pick this up with a larger AIE estimate. These numbers are reported in (ref), where panel A shows the CM effects with the binary outcome of subjective good overall health, and panel B the binary outcome of subjective overall well-being.
This reversal in conclusions highlights the importance of correcting for negative selection into healthcare usage. A conventional approach to CM fails to account for the fact that individuals with poorer underlying health tend to visit healthcare providers more frequently, generating negative selection bias that obscures the true positive AIE. By explicitly adjusting for this bias using the MTE approach, I isolate a credible positive indirect effect of healthcare usage on subjective health and well-being.
These findings offer credible evidence that improved healthcare access does yield meaningful subjective health and well-being benefits, despite previous research emphasising negligible effects on objective health measures such as blood pressure baicker2013oregon. Subjective measures likely reflect broader psychological and financial relief associated with reduced healthcare-related anxiety and diminished risk of catastrophic medical debt, thus producing more noticeable short-term subjective improvements.
Nevertheless, this analysis is subject to notable limitations. The IV is not ideal, and potentially more important mediators (such as explicit health insurance status) would require additional IVs beyond the wait-list lottery itself, presenting a challenging identification issue. Furthermore, the 95% confidence intervals for both ADE and AIE estimates (based on boostrapped SEs) remain large, though statistically significant and excluding zero. This uncertainty underscores common challenges in applied CM analyses, where statistical precision can be limited by data constraints.
This paper has studied a selection-on-observables approach to CM in a natural experiment setting. I have shown the pitfalls of using the most popular methods for estimating direct and indirect effects without a clear case for the mediator being quasi-randomly assigned. Using the Roy model as a benchmark, a mediator is unlikely to be quasi-randomly assigned in natural experiment settings, and the bias terms likely crowd out inference regarding CM effects.
This paper has also contributed to the growing CM literature in economics, connecting to MTE methods and developing a compelling way of estimating direct and indirect effects in a natural experiment setting. It has also recognised limitations in the common practice of suggestive evidence for mechanisms, and given credible CM estimates in a famous natural experiment setting well-known in the economics field. Further research could build on the approaches presented by suggesting efficiency improvements, adjustments for common statistical irregularities (say, cluster dependence), or integrating the MTE approach to the growing double robustness literature on CM farbmacher2022causal,bia2024double.
These findings do not provide a blanket endorsement for applied researchers to use CM methods. There are strong structural assumptions for adjusting identifying CM effects despite unobserved selection-into-mediator, and inference requires an IV for mediator take-up. If these assumptions do not hold true, then selection-adjusted estimates of CM effects will also be biased, and will not improve on an unadjusted conventional approach.
Yet, there are likely settings in which the structural assumptions are credible. Mediator monotonicity aligns well with economic theory in many cases, and it is plausible for researchers to study big data settings with external variation in mediator take-up costs. In these cases, this paper opens the door to identifying mechanisms behind treatment effects in natural experiment settings.
\singlespacing \onehalfspacing