EconBase
← Back to paper

A Double Machine Learning Approach to Combining Experimental and Observational Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,652 characters · 22 sections · 60 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Double Machine Learning Approach for Combining Experimental and Observational Studies

abstractExperimental and observational studies often lack validity due to untestable assumptions. We propose a double machine learning approach to combine experimental and observational studies, allowing practitioners to test for assumption violations and estimate treatment effects consistently. Our framework proposes a falsification test for external validity and ignorability under milder assumptions. We provide consistent treatment effect estimators even when one of the assumptions is violated. However, our no-free-lunch theorem highlights the necessity of accurately identifying the violated assumption for consistent treatment effect estimation. Through comparative analyses, we show our framework's superiority over existing data fusion methods. The practical utility of our approach is further exemplified by three real-world case studies, underscoring its potential for widespread application in empirical research.
keywordsData Fusion, Generalizability, External Validity, Observational Study, Machine Learning

Introduction

Experiments and observational studies are both indispensable for estimating treatment effects and investigating causal questions of interest. Experiments, or randomized control trials (RCTs), guarantee unbiased estimation of sample average treatment effects via the randomization of treatment across the experimental population. However, they are typically small, expensive, and slow to design and implement. Another concern involves the experimental units themselves, which may be recruited via a non-randomized process and therefore may not be representative of the superpopulation of interest; in such cases, we would say that the experiment lacks external validity. If this is the case, the conclusions of the experiment may not be generalizable to the superpopulation, which is oftentimes of primary interest. On the other hand, observational studies, in which units self-select into treatment, tend to be larger and include units that are representative of the superpopulation. However, analysis of observational studies is plagued by observed and unobserved confounders -- covariates that affect both units' propensity to self-select into treatment and their response to that treatment. While a variety of methods -- matching and weighting being two large classes among them -- can be used to adjust for confounders, they typically rely on the crucial assumptions that they can accurately model the treatment-covariate or response-covariate relationships and that the available covariates are sufficient for doing so. This latter assumption, which is often referred to as strong or conditional ignorability, is untestable using observational data alone rosenbaum1983assessing.

In this paper, we propose methods for jointly using experimental and observational data to (i) test for violations of at least one of validity and ignorability and (ii) estimate population average treatment effects even under such violations. To the best of our knowledge, no formal testing framework has been proposed for these violations, and as we discuss below, existing estimation approaches target similar estimands by assuming external validity and conditional ignorability. Crucially, our estimators are valid under violations of certain sets of assumptions that are often required in the literature, making our methods applicable in settings where other methods may be inappropriate. We provide a detailed discussion of the literature in Section (ref). We derive the influence functions for our scenario of interest. Further, we rigorously show that our estimators possess rate-double-robustness properties, ensuring their root-n consistency, asymptotic normality, and unbiasedness in Sections (ref) and (ref). The properties of these estimators allow us to address testing as well as estimation tasks. We emphasize that we accomplish these tasks by making use of both observational and experimental data, despite the fact that one of external validity or conditional ignorability may be violated.

After establishing the theoretical properties of our tests and estimators, we apply them to three real-world studies -- one of them is in Section (ref) while the other two are relegated to supplementary material. In doing so, we showcase the utility of our method for practitioners using observational and experimental data to answer their causal questions of interest. The analyses reach different conclusions about violations of ignorability and validity, illustrating the wide range of data that can be appropriately analyzed with our methods. In the first, we reanalyze data from Project STAR project_star_data,mosteller2014tennessee, which consists of experimental and observational data gathered to study the impact of small class size in early schooling on future test scores. We estimate a small, positive treatment effect (about 6 points on a test that ranges between 400-800 total points) attributable to a small class size on a standardized test, consistent with prior results on this dataset. Leveraging the theory developed in our paper, we can point-identify the amount of selection bias in the observational data using breakdown-frontiers analysis masten2020inference. Our estimates indicate that an unobserved confounder may have led to a selection bias of 16 points in the observational study. Second, we analyze data gathered by the Coronary Artery Surgery Study (CASS) CASS, sponsored by the National Heart, Lung, and Blood Institute, to compare the effects of coronary artery bypass surgery to non-surgical medical treatment. Here, our population ATE estimate is consistent with those in past literature, as is the fact that we find no evidence of unobserved confounding or violations of external validity. Lastly, we explore the application of our approach to the experimental data from the National Supported Work Demonstration (NSW) and observational controls from the Population Survey of Income Dynamics (PSID) lalonde1986,dehejia2002. Our analysis indicates that external validity is violated in the NSW sample. Interestingly, we find that an existing approach to treatment effect estimation proposed by malts can prune heavily confounded units and recover the experimental ATE using observational data.

The rest of the paper proceeds as follows. In Section (ref), we discuss related work. In Section (ref), we introduce our setting, notation, and various quantities of interest. Sections (ref) and (ref) outline our procedures for testing and estimation. Sections (ref) & (ref) apply our methods to the Synthetic and STAR datasets. Further, Appendices (ref) & (ref) discusses the application of our approach to CASS and Lalonde datasets. Section (ref) summarizes our contributions and discusses future extensions.

Literature Review

The advent of huge observational datasets has led to the rapid growth of several (mostly disjoint) subfields. In this section we aim to bridge those together and place our proposed methodology within this broader context. We concentrate on the following fields: (i) generalizability, (ii) combining observational and experimental datasets, (iii) violations of ignorability, (iv) negative controls, and (v) double robustness.

Generalizability. Though mild conditions ensure that treatment effects estimated from randomized control trials are unbiased, this unbiasedness is with respect to the distributions within the experiment; that is, the unbiasedness only necessarily holds for the average treatment effect within the experimental sample. One branch of work on generalizability defines the conditions under which effects can be appropriately generalized to larger populations. For example, egami_hartman_ex_validity categorize the types of validity that are necessary for an experiment's results to be appropriately extended to the population. Furthermore, they specify assumptions leading to these forms of validity and show identification of population treatment effects given they hold. Other work on generalizability deals with adjusting treatment effects derived from an experimental sample to extend them to the population. sens_w_bias_funs, for example, posit bias functions describing how lack of exchangeability between the units in experiment and population affects estimates, and they use these bias functions to perform sensitivity analysis. If the bias function is correctly specified, standard weighting, outcome-based, and doubly robust estimators can be modified to correctly debias estimates. pop_outcomes_adjust_exp use population data to learn outcome models that are used to estimate and residualize experimental outcomes; the residuals are used in conjunction with sampling weights to estimate the population average treatment effect (PATE).

Combining Experimental and Observational Studies. Observational and experimental data have recently been used together to improve the efficiency of treatment effect estimation procedures. In this setting, the idea is to take advantage of both the internal validity of the experimental sample and the external validity (and typically larger sample size) of the observational data. This field has grown rapidly in the past several years and here we briefly mention some key work. {For a broad review of this literature, see brantner2023methods and colnet2020causal.}

design_exp_from_obs use observational data to estimate confidence sets for the variance of within-strata potential outcomes and incorporate the sets into a regret minimization problem used to design an experiment. Conversely, emulate_rct seek to duplicate randomized control trials, isolating trial-mimicking populations and then conducting propensity score matching therein. shrinkage_combine construct a James-Stein type shrinkage estimator that, under mild conditions, is guaranteed to outperform estimates constructed using only one of the two data sources. Similarly, pscore_combine use a propensity score model trained on the observational data to stratify experimental data based off what their propensity scores would have been had they been in the other sample. The resulting stratified design is then used to estimate treatment effects using both samples. hartman_sekhon discuss the assumptions required for identification of the population average treatment effect on the treated (ATT) from experimental data and suggest a method for using baseline characteristics and outcomes available from non-randomized data to adjust experimental estimates appropriately. long_term_combine focus on the related problem of estimating treatment effects on a long-term outcome that is only observed in the experimental data. ts_combine use time-series data to learn a linear outcome model for each individual observational unit; the estimated parameters can then be incorporated into a model for the experimental data that may increase power of the ensuing analysis. mutual_debias propose a multi-study instrumental variable approach. They show how the approach allows for nonparametrically debiasing both experimental (biased due to lack of external validity) and observational estimates (biased due to unobserved variables affecting selection into treatment); in practice, an instrumented difference-in-differences is used to point-identify local average treatment effects. Triantafillou2021 use experimental data to learn feature sets that should be adjusted for when estimating treatment effects on observational data. Ghassami2022 focus on short- and long-term outcomes and proposes various assumption sets under which causal effects on the latter can be identified. Dang2022 estimate treatment effects using negative control outcomes -- outcomes that are unaffected by treatment, but affected by some unobserved confounder induced bias. power_likelihood augment experimental data with observational data, where the likelihood of the observational data is raised to some power $\eta < 1$ and $\eta$ is chosen to maximize the expected log pointwise predictive density on the experimental data. Likelihood-based inference is then performed to estimate treatment effects.

Our work broadens many of the above works in a fundamental way: while we also take advantage of having both observational and experimental data to better estimate population average treatment effects, we can do so in the presence of unobserved confounding, and even formally test for its impact on the treatment selection mechanism in the observational sample. Furthermore, our estimators are computationally efficient and do not depend on the presence of specialized structure in the data (e.g., existence of instrumental variables or proxy outcomes).

Violation of Ignorability. Our work is also related to the study of the impact of unobserved confounders on a treatment effect estimation procedure. In settings in which only observational data are available, most works (explicitly or implicitly) make an assumption equivalent to strong ignorability to identify causal quantities of interest. This is because, in many of these scenarios, it is impossible to identify or estimate the degree of unobserved confounding using observational data alone. One approach to such settings relies on sensitivity analyses to potential unobserved confounding rosenbaum_book, intro_to_sa. Such analyses typically posit parametric models (but see, e.g., sa_wo_assumptions for a nonparametric approach) for how an unobserved confounder affects treatment and outcome and then assess how large the effect of the unobserved confounder needs to be to change the treatment effect estimate by some specified amount. Crucially, however, such methods cannot suggest anything about the true degree of unobserved confounding present. In general, even if a sensitivity model is correctly specified, extreme sensitivity to the specified amount of unobserved confounding does not imply a high degree of unobserved confounding; neither does a lack of sensitivity imply a low degree.

In experimental settings, the randomization of treatment protects against any unaccounted selection-bias. There is currently a dearth of work concerning the study of unobserved confounders using both experimental and observational data. kallus_unobserved consider the problem of CATE estimation in this setting. Theoretical results rely upon assumptions of external validity and of a parametric structure for the impact of the unobserved confounder. Furthermore, the validity of the method depends on an identification assumption whose form depends on the posited parametric form. Lastly, it is not clear how this approach generalizes to estimation of average treatment effects. shrinkage_combine suggest a sensitivity analysis that yields an `implied' degree of unobserved confounding, based off an augmented inverse propensity weighted estimator. The method is restricted to this estimator, however, and, as with other sensitivity analyses, cannot truly identify the presence of an unobserved confounder. iv_test_unobserved describes a set of sufficient conditions on an instrument that allows for testing of unobserved confounding, though their method does not allow for identification of average causal effects under the presence of confounders.

We further draw connections to the literature on negative control and double robustness.

Negative Controls. Negative controls are variables that do not causally affect the outcome of interest and thus have a null effect by design lipsitch2010negative,tchetgen2020introduction, and have recently been used in complex causal inference settings. For example, unobserved confounders are integral to the notion of negative control outcomes in the work of Dang2022. However, the required dependencies between the unobserved confounder and the treatment, outcome, and negative control outcome must be assumed. Our work can also be interpreted in this context: Here, the variable informing about a unit's inclusion in an experimental sample can be considered as a negative control under the assumption that selection into an experimental sample has no direct causal path to the outcome. The key difference between negative control literature and our work here is that for units with $S=1$, the treatment assignment is independent of any unobserved confounders. The knowledge of the treatment assignment mechanism in the experiment not only allows testing for unobserved confounding in the corresponding observational study and external validity of the experimental sample but also estimation of population ATE even if one of these assumptions is violated.

Double Robustness. Our work also bears a resemblance to the literature on the doubly-robust estimation of causal effects robins_original_dr, summary_dr, where outcome and treatment-selection estimators are combined to create an estimator consistent under misspecification of either propensity or prognostic score models (but not both simultaneously). Our work distinguishes itself from the above in we consider violations of assumptions, not inappropriate model specifications. We will later show that our estimator is also doubly robust in the traditional sense; however, this is not the primary focus of this paper.

Preliminaries

figure[figure omitted — 682 chars of source]

Causal Inference: Notation

We consider settings with $n$ observed units, indexed by $i$, where we are interested in estimating the causal effect of a binary treatment $T$ on outcome $Y$. Each unit is associated with two real-valued potential outcomes, $Y_i(0)$ and $Y_i(1)$ corresponding to two treatment choices. While we assume binary treatment for simplicity, our results immediately extend to the cases of $M > 2$ potential outcomes. We take $Y_i(t) \in \mathcal{Y}$ for $t \in \{0, 1\}$, where $\mathcal{Y} \subseteq \mathbb{R}$ that we will further characterize shortly. Each unit is also assigned a treatment whose observed value is denoted by $T_i \in \{0, 1\}$. The variable $Y_i$ (note the absence of parentheses) denotes the observed outcome of unit $i$, which we assume to be the only potential outcome corresponding to their assigned treatment, i.e., $Y_i = Y_i(T_i)$. In addition, we also observe, for each unit, a vector of $d$ pre-treatment covariates $\mathbf{X}_i \in \mathcal{X}$, such that $\mathcal{X}$ is a compact subset of $\mathbb{R}^d$.

In our setting, the $n$ units are divided into two subsets: a set of observational units and a set of experimental units. We will use the variable $S_i = 1$ to denote unit $i$ is in the experimental sample and $S_i = 0$ to denote unit $i$ is in the observational sample. As we will shortly see, the statistical behavior of the variables associated with a unit in the observational set will be different than that of a unit in the experimental set. Lastly, we consider an unobserved confounder $U \in \mathcal{U} \subseteq \mathbb{R}$ that may causally affect $T$, $S$, and $Y$.

Throughout the rest of this paper, we will use capitalized letters to represent random variables, lowercase letters for values within the domain of the distribution of these random variables, and $f_A$ and $F_A$ to indicate the probability density function (pdf) and cumulative distribution function (CDF) of the random variable $A$, respectively. We will be studying several functionals of the joint distribution of $Y$, $X$, $T$, and $S$. We introduce some shorthand notation for these quantities of interest here:

align[align omitted — 467 chars of source]

with corresponding estimates denoted by hats: $\hat{\mu}(\mathbf{x}, t, 0), \hat{e}(\mathbf{x}, t, 1), \hat{p}(\mathbf{x})$, etc.

The first two quantities are the conditional expectations of the outcome given covariates $\mathbf{X}$ and treatment $t$ in the observational (Equation (ref)) and experimental samples (Equation (ref)), respectively. The following two quantities are the propensity scores for treatment level $t$ also in the observational (Equation (ref)) and experimental samples (Equation (ref)), respectively. Finally, $p(\mathbf{x}_i)$ denotes the probability that unit $i$ is assigned to the experimental sample, conditionally on its observed covariates (Equation (ref)). All these quantities are unknown, but estimable with the observed data. In an ideal experimental setting, both ${e(\mathbf{x}_i, t, 1)}$ and $p(\mathbf{x}_i)$ would be known by the analyst, but this need not be the case in our framework. We now discuss assumptions key to our framework and that will help us estimate the above quantities.

Causal Inference: Assumptions

Here, we state several assumptions that are common in the causal inference literature and which are of key importance in our work. In our setting, several of these become testable and certain violations do not hinder consistent estimation of causal effects.

\noindentA1 (Data Distribution) For notation's sake, define the random vector $\Omega := (Y, T, S, \mathbf{X})$. Let $f_{\Omega}$ be a probability distribution over the joint domain $\mathcal{S} = \mathcal{Y} \times \{0, 1\} \times \{0, 1\} \times \mathcal{X}$. We assume that the full data are iid random samples from this distribution, i.e.: $\{\Omega_i\}_{i=1}^n \overset{iid}{\sim} f_{\Omega}$, and that $\Omega$ (note the absence of indices) denotes an arbitrary draw from $f_{\Omega}$. We also maintain the following standard overlap assumption: for all $\omega \in \mathcal{S}$, we assume $0 < f_{\Omega}(\omega) < 1$. \\ A2 (Internal Validity of the Experiment) We assume that $(Y(0), Y(1)) \mathrel{\text{\scalebox{1.07}{$\perp\mkern-10mu\perp$}}} T \mid \mathbf{X}=\mathbf{x}, S = 1$, i.e., if $i$ is an experimental unit, then its treatment assignment is independent of its potential outcomes conditional on its covariates. \\A3 (External Validity of the Experiment) Adjusted for pre-sampling covariates $\mathbf{X}$, the conditional means of potential outcomes are exchangeable across study populations: $\mathbb{E}[Y(t) \mid \mathbf{X}=\mathbf{x}, S=1] = \mathbb{E}[Y(t) \mid \mathbf{X}=\mathbf{x}, S=0]$ for all $t$ and $\mathbf{x}$.\\ A4 (Conditional Ignorability) Adjusted covariates $\mathbf{X}$, in the observational study, the conditional means of potential outcomes are exchangeable across treatment arms: $\mathbb{E}[Y(t) \mid \mathbf{X}=\mathbf{x}, T=1, S=0] = \mathbb{E}[Y(t) \mid \mathbf{X}=\mathbf{x}, T=0, S=0]$ for all $t$ and $\mathbf{x}$. \\A5 (Sampling Ignorability) For all $t$ and almost surely over $f_{\Omega}$ we have: $Y \mathrel{\text{\scalebox{1.07}{$\perp\mkern-10mu\perp$}}} S \mid \mathbf{X} = \mathbf{x}, T=t, U=u$, i.e., being sampled into the experimental set is independent of the outcome conditionally on the covariates, unobserved confounder, and the treatment level. \\A6 (Nuisance Parameter Estimation) For all $t \in \{0, 1\}$ and $s \in \{0, 1\}$:

gather[gather omitted — 408 chars of source]

where $\left\Vertf\right\Vert_{P, q} = \left\Vertf(\Omega)\right\Vert_{P, q} = (\int |f(\omega)|^qd {P}(\omega))^{1/q}$ denotes the $L^q(P)$ norm with respect to the distribution of the data.

\paragraph{Discussion of Assumptions} Assumption A1 is the standard overlap assumption and a regularity condition required to operationalize our double machine learning procedure. Assumption A2, which states that experimental units are assigned treatment effectively at random conditional on the observed covariates, should hold in most cases where an experiment is properly conducted and treatment truly is randomized. As a consequence of A2, we have $P(Y(t) | \mathbf{X} =\mathbf{x}, T=t, S=1) = P(Y(t) | \mathbf{X} =\mathbf{x}, S=1)$. It may still be violated, however, in certain cases, such as if treated units with higher potential outcomes tend to drop out of the experiment and are excluded from further analysis. The knowledge of a unit's treatment status would give information about their potential outcomes and vice versa, violating A2.

Assumption A3 states that units are selected into the experimental set effectively at random conditional on the observed covariates. Importantly, if there exists an unobserved confounder that is not marginally independent of the potential outcomes, it must be marginally independent of $S$. A practical example of a setting where A2 and A3 both hold (barring violations of A2 akin to the above) is one in which a researcher first samples $n$ individuals from a superpopulation of interest, measures baseline covariates, $\mathbf{X}$, and chooses a subset to invite into a lab experiment based on $\mathbf{X}$, while the rest of the individuals are allowed to self-select into treatment. This is approximately the setup of the CASS study we will analyze later (Section (ref)). As a consequence of A3, we have $P(Y(t) | \mathbf{X} =\mathbf{x}, S=s) = P(Y(t) | \mathbf{X} =\mathbf{x})$.

Assumption A4 is the standard conditional ignorability assumption and asserts that there are no unobserved confounders affecting treatment and outcome. This assertion is untestable from observational data alone. However, A4 is typically assumed by default in observational analyses (though sensitivity analyses may be performed to assess the sensitivity of causal estimates to varying strengths of a posited unobserved confounder). As a consequence of A4, we have $P(Y(t) | \mathbf{X} =\mathbf{x}, T=t, S=0) = P(Y(t) | \mathbf{X} =\mathbf{x}, S=0)$. Later, we will show that having experimental data makes this assumption testable.

Assumption A5 assumes that sampling into the experiment does not have a direct causal path affecting the outcome i.e., in Figure (ref)(a) there is no arrow from S to Y. Such an assumption would be unreasonable if participation in the experiment in and of itself changed a unit's outcome. For example, consider an experiment in which individuals are paired and their conversations monitored for indications of politeness. If knowledge of being monitored leads one to speak more politely than they would otherwise, A5 would be violated. This assumption can also be seen as a form of the consistency or SUTVA assumption extended to our two population settings.

Assumptions A3 to A5 (or a version of them) are common across approaches that combine experimental and observational studies to estimate treatment effects rudolph2023improving, lu2019causal, wu2022integrative. Further, wu2022integrative considers estimating effects in a scenario where A4 is violated.

Lastly, Assumption A6 specifies the rates at which we must be able to estimate nuisance parameters. Importantly, achieving these rates does not require correct specification of a parametric model and can be accomplished with flexible nonparametric or machine learning methods. These rate conditions arise naturally from efficient influence function based estimators in semiparametric statistics, which underpins the construction of doubly robust and orthogonal estimators bickel1993efficient,robins1994estimation,rotnitzky2012improved,farrell2015robust,chernozhukov2018double,kennedy2024semiparametric.

Importantly, the above assumptions imply that some of our quantities of interest have the following key properties:

align[align omitted — 331 chars of source]

Equation (ref) is implied by A2, and is key in any causal inference framework as it permits identification of the conditional mean: potential outcomes can be consistently estimated for every experimental unit. Equation (ref) is also directly implied by A2 and is a consequence of the fact that assignment to treatment in the experiment does not depend on a unit's potential outcomes. Finally, by A3, Equation (ref) states that assignment to the experimental sample is independent of the potential outcomes, conditionally on covariates. Note that all these quantities are guaranteed to be non-degenerate by the distributional requirements in A1.

Impossibility of Double Resilience

We begin our exposition with a result that motivates the estimators we propose in the rest of our paper: while it is possible to leverage experimental and observational data to overcome violations of A3 or A4, combining experimental and observation data without the knowledge of violations of A3 and/or A4 may not yield the expected benefits and can result in a biased estimate.

This is important as it clarifies the requirements of estimation with experimental and observational data in a very general sense: users of these methodologies must possess knowledge of whether the experimental data generalize, or the observational set satisfies unconfoundedness, and knowledge of the fact that either of these assumptions may hold is not sufficient.

In order to prove the fact that knowledge of whether A3 or A4 is violated is necessary for estimation, we investigate whether an estimator for treatment effects might exist that is efficient under violations of either A3 or A4. We would call such an estimator doubly resilient as it would be resilient to violation of one of two assumptions. The terminology `double resiliency' contrasts with the term `double robustness'; the latter refers to the misspecification of models, not the violation of assumptions. We formally define this property below:

definition[Double Resilience] Given A1, A2, and A5, a doubly resilient estimator $g_{DR}(1, \mathbf{X})-g_{DR}(0, \mathbf{X})$ is an unbiased and consistent estimator of conditional average treatment effects if at least one of A3 or A4 is not violated. More precisely, we say that $g_{DR}(t, \mathbf{X})$ is doubly resilient if under assumptions (A1, A2, A3, A5) or under assumptions (A1, A2, A4, A5), then $\forall t,\mathbf{X},s$: $\mathbb{E}[g_{DR}(t, \mathbf{X})] = \mathbb{E}\left[ Y(t) \mid \mathbf{X} \right]$ and $\mathbb{E}[g_{DR}(t, \mathbf{X})\mid S=s] = \mathbb{E}\left[ Y(t) \mid \mathbf{X}, S=s \right]$.

Importantly, we have the following result:

theorem[Doubly Resilient Estimators Do Not Exist] There does not exist any doubly resilient estimator $g_{DR}(t, \mathbf{X})$.

The proof can be found in the supplement. The intuition behind the proof is that existence of such an estimator would allow one to distinguish between violations of A3 and A4 from the data alone, which in turn would imply the ability to learn about the unobserved confounder-induced selection bias, which we cannot characterize with the observed data by definition. Given that a doubly resilient estimator does not exist, we now continue with the exposition of our proposed estimator.

Falsification Test for External Validity and Conditional Ignorability

As stated initially, we would like to leverage the experimental sample in order to learn something about the observational sample, and vice versa. Here, we show that it is possible to test whether at least one of the assumptions A3 and/or A4 are violated, without necessarily knowing which assumption(s) are violated. This is done by demonstrating that the same test statistic provides us evidence against the null of A3 and A4 holding towards the following two alternatives:

enumerate• Given that internal validity (A2) and external validity (A3) of the experiment hold, we can test for violations of ignorability in the observational sample (A4), • Alternatively, given that internal validity of the experiment (A2), and ignorability of treatment (A4), we can test for violations of external validity in the experimental sample (A3).

To formalize this claim, we introduce the quantity $\theta(t) := \mathbb{E}_{\mathbf{X}}\left[ \mu(\mathbf{X},t,0) - \mu(\mathbf{X},t,1) \right]$ that can provide insight into the existence of an unobserved confounder. Our key claims are as follows:

theorem[Identification] \begin{enumerate} • Given the set of assumptions (A1, A2, A3), if there exists a $t$ such that $\theta(t) \neq 0$, then A4 (conditional ignorability in the observational sample) is violated. • Given the set of assumptions (A1, A2, A4, A5), if there exists a $t$ such that $\theta(t) \neq 0$, then A3 (external validity in the experimental sample) is violated. • $\mathbb{E}[Y(t)|\mathbf{X}=\mathbf{x}]$ is identified under either (A1, A2, A3) or (A1, A2, A4, A5). \end{enumerate}

The proof of this theorem is in Appendix (ref).

The above result implies that $\theta(t)$ is informative about the level of “confoundedness” of the data with respect to potential outcome $Y(t)$. To see this, we can expand the definition of $\theta(t)$ and use Equation (ref):

equation*[equation* omitted — 196 chars of source]

Therefore, $\theta(t)$ will be informative as to the expected difference between the distributions of the potential outcomes across the two samples. When this difference is large in expectation, then $\theta(t)$ will be large (in magnitude), and we can conclude that either sample may not be very informative as to the true value of the mean potential outcome of interest. This may be the case if the selection of units in the experimental sample or the treatment selection in the observational sample are affected by unobserved confounders. Thus, if an analyst believes external validity holds for their data, a large $\theta(t)$ would then suggest violation of conditional ignorability, and vice versa.

An Estimator for Confounding

lemmaGiven assumption A1, A2 and A5, for any $t \in \{0,1\}$, the efficient influence function for $\theta(t)$ under null is given as \begin{eqnarray} \phi(t) &:=& {\mu(\mathbf{x}, t, 0)} - {\mu(\mathbf{x}, t, 1)} \nonumber + \frac{T(t) (1-S)}{{e(\mathbf{x}, t, 0)}(1-p(\mathbf{x}))}(Y - {\mu(\mathbf{x}, t, 0)}) \&& \quad\quad\quad\quad\quad\quad\quad\quad\quad - \frac{T(t)(S)}{{e(\mathbf{x}, t, 1)}p(\mathbf{x})}(Y - {\mu(\mathbf{x}, t, 1)}). \end{eqnarray}

The efficient influence in Lemma (ref) can be derived via procedure described in hines2022demystifying and kennedy2024semiparametric. We now introduce a flexible two-stage estimator for $\theta(t)$ based on the efficient influence function in equation (ref). Our estimator leverages the following equality: $\mathbb{E}[\psi(t)] = \theta(t)$ for doubly robust and efficient estimation.

The asymptotic properties of our estimator can be derived specifically for the estimation problem as a two-stage procedure leveraging results on double ML in chernozhukov2018double. The first-stage parameters that our estimator of $\theta(t)$ depends on are estimates of all the quantities defined in Equations (ref) -- (ref). Our framework estimates these quantities using flexible machine-learning models, leading to the theoretical guarantees for our second-stage estimator.

Our proposed procedure is comprised of two stages:\\ Stage 1: Randomly partition the full sample into \(K\) approximately equal-sized folds, indexed by \(k = 1, \dots, K\). For each fold \(k\), proceed as follows:

enumerate• Fit a flexible ML model for predicting \(\mu(\mathbf{x}_i,t,0)\) using the observational units not in fold \(k\). Use the fitted model to obtain predictions \(\hat{\mu}(\mathbf{x}_i,t,0)\) for units in fold \(k\). • Fit a flexible ML model for predicting \(\mu(\mathbf{x}_i,t,1)\) using the experimental units not in fold \(k\). Use the fitted model to obtain predictions \(\hat{\mu}(\mathbf{x}_i,t,1)\) for units in fold \(k\). • Estimate the sampling propensity \(p(\mathbf{x}_i)\) either using true known propensities or by fitting a flexible ML model to sample indicators \(S_i\) using units not in fold \(k\), and obtain predictions \(\hat{p}(\mathbf{x}_i)\) for units in fold \(k\). • Fit a flexible ML model for predicting the treatment propensity \(e(\mathbf{x}_i,t,0)\) using observational units not in fold \(k\), and obtain predictions \(\hat{e}(\mathbf{x}_i,t,0)\) for units in fold \(k\). • Fit a flexible ML model for predicting the treatment propensity \(e(\mathbf{x}_i,t,1)\) using experimental units not in fold \(k\) (or use true experimental propensities if known), and obtain predictions \(\hat{e}(\mathbf{x}_i,t,1)\) for units in fold \(k\).

This crossfitting structure ensures that the model used to predict each quantity for a unit is trained on data excluding that unit, thus mitigating overfitting bias.

\noindentStage 2: After obtaining the first-stage estimates for all units via cross fitting, compute the following estimator:

equation[equation omitted — 101 chars of source]

where $\widehat{\phi}_i(t) =$

equation[equation omitted — 332 chars of source]

In this second stage, we aggregate the per-unit corrected estimates \(\widehat{\phi}_i(t)\) across all units to produce a final estimate of \(\theta(t)\). Each unit’s contribution \(\widehat{\phi}_i(t)\) is a bias-corrected difference between the observational and experimental response function estimates. If a unit belongs to the experimental sample, its observed outcome is used to correct the experimental model estimate; if it belongs to the observational sample, its observed outcome corrects the observational model estimate. These corrections are critical for achieving asymptotic efficiency.

Owing to the specific form of the estimator, and to the cross-fitting procedure at Stage 1 of our algorithm, the asymptotic properties of $\hat{\theta}(t)$ are given in the following theorem:

theoremUnder null such that Assumptions (A1-A6) are satisfied, then: \begin{align*} \sqrt{n}\hat{\theta}(t) \overset{d}{\rightarrow} \mathcal{N}(0, \sigma^2(t)), \end{align*} where: $\sigma^2(t) = \mathbb{E}\left[ \phi^2(t) \right] = \mathbb{E}\left[(\phi(t) - \theta(t))^2\right]$ and the expectation is taken over $\Omega$. \\In addition, the na\"ive estimator for the variance is consistent with the first-stage estimates used in place of the true nuisance parameters: $$\hat{\sigma}^2(t) = \frac{1}{n}\sum_{i=1}^N(\widehat{\phi}_i(t) - \hat{\theta}(t))^2 \overset{p}{\rightarrow} \sigma^2(t),$$ Finally, it follows that an uniformly valid asymptotic $1-\alpha$ confidence interval is: $$\left[\hat{\theta}(t) \pm \Phi^{-1}(1-\alpha/2)\sqrt{\hat{\sigma}^2(t)/n}\right]$$ where $\Phi$ is CDF of univariate standard normal distribution.

This theorem relies on the adaptation of Theorems 3.1 and 3.2, and Corollary 3.1 in chernozhukov2018double to our setting. Its proof, which can be found in the appendix, revolves around showing that the score yielding the estimator $\hat{\Psi}$ enjoys properties like unbiasedness, Neyman-orthogonality, and local insensitivity to the values of the nuisance parameters $\mu(\mathbf{x}, t, s)$, $e(\mathbf{x}, t, s)$, and $p(\mathbf{x})$. The theorem is useful as it permits us to estimate uncertainty around our estimate of $\theta(t)$, and, therefore, the level of confounding present in the observational dataset. It also allows for the construction of falsification tests regarding the presence of unobserved confounding, as we show next.

Notably, the above algorithm depends on a leave-one-out style procedure at Stage 1. This procedure can be slow for large datasets and complicated ML models for first stage parameters that take a long time to fit to data. Because of this, a sample-splitting procedure akin to K-fold cross validation is a computationally efficient alternative. Specifically, it is possible to split the data into $K$ non-overlapping folds at random, and to then run both Stage 1 and Stage 2 by fitting models on all data but the units in the $k^{th}$ fold, and then predicting on those units. The final output of this procedure will be $K$ estimators of the form in (ref), which can be averaged together to obtain a final estimate of $\theta(t)$. This procedure of sample-splitting for efficient estimation of semiparametric models, also known as repeated cross-fitting, has been well studied in the literature schick1986asymptotically, bickel1993efficient, zheng2010asymptotic, ayyagari2010applications, robins2013new,chernozhukov2018double. bickel1993efficient examined estimation of nuisance functions using a vanishingly small portion of the sample; schick1986asymptotically extended these ideas by introducing equal two-way sample splitting and discretization of the parameter space; and van der Vaart’s treatment in bickel1993efficient also employed such sample splitting and discretization to provide weak conditions for $k$-step estimators based on efficient scores, with the updates estimated on the split samples. Along similar lines, robins2013new proposed sample splitting to construct higher-order influence function adjustments in semiparametric estimation; and in the targeted maximum likelihood literature, zheng2010asymptotic operationalized sample splitting in the context of $k$-step updating.

Using $\hat{\theta}(t)$ as a test statistic

Our goal is to test if either A3 or A4 is violated. To do this, we note that Theorem (ref) implies that if Assumptions A3 and A4 are satisfied (along with A1, A2, A5, A6) then $Y$ and $S$ are mean independent conditional on $X,T$ i.e., $\mathbb{E}[Y|\mathbf{X}=\mathbf{x}_i, T=t, S=1] = \mathbb{E}[Y|\mathbf{X}=\mathbf{x}_i, T=t, S=0]$. Thus, violation of mean independence implies violation of either A3 or A4. This suggests the following specification of the null hypothesis of interest:

align[align omitted — 240 chars of source]

By definition, under $\mathbb{H}^0(t)$:

align*[align* omitted — 262 chars of source]

This allows us to use the results in Theorem (ref) to construct a test of this null. The second statement is a direct consequence of expanding $\sigma^2(t) = \mathbb{E}[(\phi_i(t) - \theta(t))^2|\mathbb{H}^0(t)]$ and ${\mu(\mathbf{x}_i, t, 0)} = {\mu(\mathbf{x}_i, t, 1)}$ under $\mathbb{H}^0(t)$. First we note that under the assumptions of Theorem (ref), a consistent estimator for the variance above is given by:

equation[equation omitted — 355 chars of source]

Since, under $\mathbb{H}^0(t)$, the required assumptions for the theorem are satisfied, this allows us to construct a Wald-type test statistic $\hat{\theta}(t)$: $Z_n(t) := \sqrt{n}\hat{\theta}(t)/\hat{\sigma}^2_r(t)$. By Theorem (ref), we know that $Z_n(t)$ is asymptotically normal under $\mathbb{H}^0(t)$. By construction, it is enough to reject the null for a single $t$ in order to have evidence against both A3 and A4 holding. In some settings only a single test need to be performed (Section (ref)) but when multiple tests are performed, a correction for multiple falsification testing, such as Bonferroni's correction, may become needed.

Treatment Effect Estimation

In the presence of an unobserved confounder, it is not possible to consistently estimate the ATE using just observational data without any additional untestable assumptions. This section presents a regular and asymptotically normal estimator that combines experimental and observational data to consistently estimate the ATE when conditional ignorability (A4) is violated but external validity (A3) holds.

Efficiency Bound & Asymptotic Variance

We start by first deriving an efficient influence function (EIF) in a scenario where (i) both A3 and A4 hold and (ii) A3 holds but A4 is violated (the scenario mentioned above).

The EIF is the canonical gradient of the pathwise derivative of the statistical estimand, capturing the most informative direction for local perturbations of the data distribution within the nonparametric model rudolph2023improving. As such, is used as a foundation for deriving the efficient estimator with the lowest possible asymptotic variance, as defined by the efficiency bound. This bound represents the smallest achievable variance for an unbiased regular and asymptotically linear (RAL) estimator under the given model constraints. We use this EIF to guide the construction of a robust double machine learning estimator.

The estimand of interest is the average treatment effect (ATE) $\tau:= \mathbb{E}_{\Omega}[Y(1) - Y(0)]$. For simplicity of notation, let $\nu(t) := \mathbb{E}[Y(t)]$ be the expected potential outcome for treatment $t$ and $\nu(\mathbf{x},t) := \mathbb{E}[Y(t) \mid \mathbf{X}=\mathbf{x}]$ be the conditional equivalent of the same. Then, the ATE can be rewritten $\tau = \nu(1) - \nu(0)$. Further, let $\forall \mathbf{x},t,s, Var(Y \mid \mathbf{X}=\mathbf{x}, T=t, S=s) = Var(Y \mid \mathbf{X}=\mathbf{x}, T=t)$ (cross-sample homoskedasticity) and $\sigma^2_Y(\mathbf{X},t) = Var(Y \mid \mathbf{X},T=t)$. We use the strategies delineated in hines2022demystifying and kennedy2024semiparametric to derive the efficient influence function for $\nu(1)$ and $\nu(0)$ and the corresponding efficiency bound.

theorem[EIF under A3] Given assumption A1-A3 and cross-sample homoskedasticity for any $\mathbf{x}, t$ \begin{enumerate} • If A4 holds then the efficient influence function for estimating $\nu(t)$ is given as \begin{eqnarray*} EIF(t) &=& {\mu(\mathbf{x}, t, 1)} p(\mathbf{x}) + {\mu(\mathbf{x}, t, 0)} (1-p(\mathbf{x})) - \nu(t) \&& + \left( \frac{S \mathbf{1}[T=t]}{{e(\mathbf{x}, t, 1)}}(Y - {\mu(\mathbf{x}, t, 1)}) + \frac{(1-S) \mathbf{1}[T=t]}{{e(\mathbf{x}, t, 0)}}(Y - {\mu(\mathbf{x}, t, 0)})\right) \end{eqnarray*} and the efficiency bound is given as follows \begin{eqnarray*} \mathbb{E}\left[ \frac{\sigma^2_Y(\mathbf{X},t) p(\mathbf{X}) }{e(\mathbf{X},t,1)} + \frac{\sigma^2_Y(\mathbf{X},t) (1-p(\mathbf{X}))}{e(\mathbf{X},t,0)} \right]+ \mathbb{E}\left[ (\nu(\mathbf{X},t) - \nu(t))^2\right] \end{eqnarray*} • If A4 is violated then the efficient influence function for estimating $\nu(t)$ is given as \begin{equation*} EIF(t) = {\mu(\mathbf{x}, t, 1)} + \frac{S}{p(\mathbf{x})}\left( \frac{\mathbf{1}[T=t]}{{e(\mathbf{x}, t, 1)}}(Y - {\mu(\mathbf{x}, t, 1)})\right) - \nu(t) \end{equation*} and the efficiency bound is given as follows \begin{equation*} \mathbb{E}\left[\frac{\sigma^2_Y(\mathbf{X},t)}{p(\mathbf{X})e(\mathbf{X},t,1)} \right] + \mathbb{E}\left[ (\nu(\mathbf{X},t) - \nu(t))^2\right] \end{equation*} \end{enumerate}

Theorem (ref) characterizes the efficiency bounds, i.e., the smallest achievable variance for unbiased, RAL estimators, under two distinct sets of assumptions. Comparing the efficiency bounds in Theorems (ref).1 and (ref).2, we find that the efficiency bound under assumption A3 alone (i.e., violation of A4) is larger than under both assumptions A3 and A4 -- this is as expected. Furthermore, under assumption A3 without A4, the EIF utilizes treatment-effect information exclusively from the experimental data and does not leverage outcome or treatment data from the observational study. This highlights that when conditional ignorability (A4) is violated, no asymptotic efficiency gain is achievable by incorporating observational data. There are no RAL estimators of any marginal treatment effect that achieve a lower efficiency bound than the AIPW estimator using data only from the RCT under assumption A3 (without assumption A4). Nevertheless, even in the absence of asymptotic benefits, observational study data may still offer finite-sample variance improvements (albeit at the cost of finite-sample bias), particularly if the observational dataset is considerably larger than the experimental one. In the next section, we introduce an estimator that effectively incorporates observational study data (under the violation of A4), which may result in finite-sample benefits.

Estimation under External Validity (A3)

We build on the insights from Theorem (ref) in order to propose an estimator for the ATE that incorporates both observational and experimental data. Consider ${\Lambda_i(t)} = {\mu(\mathbf{x}_i, t, 0)} + \frac{S_i T_i(t)}{p(\mathbf{x}_i) {e(\mathbf{x}_i, t, 1)}}(Y_i - {\mu(\mathbf{x}_i, t, 0)} )$, then under assumptions A1-A3 $\nu(t) = \mathbb{E}[{\Lambda_i(t)}]$. We first estimate $\nu(t)$ by estimating ${\Lambda_i(t)}$ for each $i$ and $t$, and then estimate the ATE as the difference of $\nu(1)$ and $\nu(0)$. We construct our cross-fitted estimator for $\nu(t)$ as:

equation[equation omitted — 104 chars of source]

where:

equation[equation omitted — 212 chars of source]

This estimator is structured around the doubly-robust estimator studied in, e.g., glynn2010introduction, farrell2015robust. First, conditional means and propensities are estimated for all units (both experimental and observational) by fitting flexible non-parametric ML models. Second, these estimates are combined for each unit (${\widehat{\Lambda}_i(t)}$), and the scores are then averaged to obtain an estimate $\hat{\nu}(t)$. An estimate of $\tau$ can be obtained simply by taking the difference of $\hat{\nu}(1)$ and $\hat{\nu}(0)$.

In our case, the estimator we propose is consistent whenever external validity (A3) is satisfied, even if conditional ignorability (A4) is violated. We provide a formal derivation in the supplement.

On top of this guarantee of consistency under violation of A4, the results for two-step estimators of chernozhukov2018double apply to our estimator for $\nu(t)$, and permit us to state the following properties of our method:

theoremGiven assumptions A1 -- A3 are satisfied, if $\hat{p}(\mathbf{X}) \rightarrow p(\mathbf{X}) $ and $\hat{e}(\mathbf{X},t,1) \rightarrow e(\mathbf{X},t,1) $ then $\hat{\nu}(t)$ is a consistent estimator of $\nu(t)$ i.e. $\hat{\nu}(t) \overset{p}\rightarrow \nu(t).$ Further, if Assumptions A4 and A6 are also satisfied, then $\sqrt{n}(\hat{\nu}(t) - \nu(t)) \overset{d}{\rightarrow} \mathcal{N}(0, \gamma^2(t)),$ where $\gamma^2(t) = \sigma^2_Y(t) \mathbb{E}\left[\frac{1}{p(\mathbf{X})e(\mathbf{X},t,1)} \right] + \mathbb{E}\left[ (\nu(\mathbf{X},t) - \nu(t))^2\right]$. In addition, the vanilla estimator for the variance is consistent with the first-stage estimates used in place of the true nuisance parameters, $\hat{\gamma}^2(t) = \frac{1}{n}\sum_{i=1}^N({\widehat{\Lambda}_i(t)} - \hat{\nu}(t))^2 \overset{p}{\rightarrow} \gamma^2(t),$ and it follows that an uniformly valid asymptotic $1-\alpha$ confidence interval is $\left[\hat{\nu}(t) \pm \Phi^{-1}(1-\alpha/2)\sqrt{\hat{\gamma}^2(t)/n}\right].$

The data-splitting procedure at step 1 of our algorithm permits our second-stage estimator to have an asymptotically normal distribution, with known variance. The theorem directly implies the following corollary for the observational ATE:

corollaryLet all the conditions in Theorem (ref) hold, with $\boldsymbol{\eta}$ and $\hat{\boldsymbol{\eta}}$ denoting estimands and respective estimators for two treatment levels, $t$ and $t'$. Then the estimator: $\hat{\tau} = \frac{1}{n}\sum_{i=1}^n\widehat{\Lambda}_i(1) - \widehat{\Lambda}_i(0)$ satisfies: $\sqrt{n}(\hat{\tau} - \tau) \overset{d}{\rightarrow}\mathcal{N}(0, \Gamma^2)$, with $\Gamma^2 = \mathbb{E}[(\Lambda_i(1) - \Lambda_i(0) - \tau)^2]$. Additionally, the estimator $\widehat{\Gamma}^2 = \frac{1}{n}\sum_{i=1}^n(\widehat{\Lambda}_i(1) - \widehat{\Lambda}_i(0) - \hat{\tau})^2$ is consistent for $\Gamma^2$.

The usefulness of this corollary is in that it permits us to construct approximate hypothesis tests and confidence intervals for $\tau$, with guaranteed uniform asymptotic coverage.

We provide the proof of Theorem (ref) in Appendix (ref). As for Theorem (ref), the proofs rely on showing that the score that has $\hat{\Lambda}$ as an estimator satisfies several key properties, including unbiasedness, Neyman-orthogonality, and local insensitivity to the values of the nuisance parameters.

Student Teacher Achievement Ratio (STAR) Project

Data Description

Project STAR (Student-Teacher Achievement Ratio) was a three-phase experiment designed to study the effect of class-size on short and long-term student performance project_star_data,mosteller2014tennessee. The project was run across 79 schools in the state of Tennessee across inner-city (17), urban (8), sub-urban (16), and rural (38) areas. A single cohort of students was studied within each school for four years. Within each STAR school, students in the study cohort were randomly assigned to one of three treatment arms: a small class (13 to 17 students), a regular class (22 to 25 students), or a regular class with a full-time teacher aide. Teachers were randomized across the treatment arms. The treatment arm was fixed for students who moved between STAR schools during the duration of the study. However, some students participated in the study for fewer than four years if they moved to a non-STAR study school. The student's gender, race, birth year, birth month, and free-lunch status were collected to look for any systematic difference between treatment arms. The primary study did not find any evidence of systematic differences across treatment arms on observable characteristics --- providing evidence against systematic failures in the randomization process, though balance on unobservables cannot be directly verified. Student performance was measured once a year, from 1986 to 1989, using standardized achievement tests, which included norm-references and criterion-referenced tests. Norm-referenced tests included Stanford achievement tests (SAT) which were developed by the Psychological Corporation. These tests are composed of separate reading, mathematics, and listening portions for grades K through 3. The observational comparison group for the study included 1780 students across grades 1 to 3 from 21 schools which were located within the same 13 districts as STAR schools. This was done to ensure a level of similarity between the comparison schools and the STAR schools in their respective districts. The same achievement tests were administered in 1987, 1988, and 1989.

Analysis and Result

Here, we study whether the Project STAR observational data has an unobserved confounder that affects students' selection into classrooms of different sizes. We also assess the degree of the associated selection bias and estimate the average treatment effect of class size on short-term standard test outcomes. Given that comparison, schools were chosen from the same 13 districts in Tennessee as STAR schools and that their similarity to STAR schools was built into the study design, it is reasonable to assume that the external validity of the experiment holds. Our framework thus allows us to test for violation of A4 and estimate the ATE even in the presence of unobserved confounding.

Point estimates of the test statistics for $\theta(1)$ and $\theta(0)$ are $22.96$ and $1.42$, and their distribution is plotted in Figure (ref). P-values for the two statistics are $0.0006$ and $0.2663$, respectively, meaning that the test rejects that $\theta(1) = 0$ at the (Bonferroni corrected) $0.05/2$ level. We thus reject the null and, given the plausibility of external validity discussed above, conclude there is strong evidence for an unobserved confounder affecting selection. To analyze the strength of selection bias we employ the idea of breakdown-frontiers: we specify a form for the selection bias, adjust our observed data according to that form, and perform our proposed test using the debiased outcomes masten2020inference. When the test does not reject, we have a plausible magnitude for the selection bias blackwell2014selection. The parametric form we choose is $q(X,T;\alpha) = \alpha(2T-1)$ --- that is, the treated potential outcome for treated units is $\alpha$ larger while the control potential outcome for control units is $\alpha$ smaller. We concentrate on $\theta(1)$ and find that the test fails to reject the null hypothesis for $\alpha \in [3, 29]$ (see Figure (ref)). Specifically, we observe that for $\alpha=16$, the p-value $p(1)$ peaks and it decreases as we further increase $\alpha$ (see Figure (ref)). To contextualize these values of $\alpha$, we note that the outcome is bounded between 400 and 800 and so the relative range of selection bias compared to the scale of $Y$ ranges from $0.75\%$ to $7.25\%$.

Lastly, we estimate the average treatment effect of small class size on grade 3 standardized test scores using the difference of means estimator for experimental data and compare it with our estimator discussed in Section (ref). We estimate the ATE to be $5.75$ , in agreement with the difference of means estimate from experimental data of $7.24 \pm 2.72$.

figure[figure omitted — 330 chars of source]
figure[figure omitted — 489 chars of source]

Data Description

The coronary artery surgery study (CASS) was initiated by National Heart, Lung and Blood institute (NHLBI) to study the effect of coronary bypass surgery in comparison to conventional medical therapy. The data were collected from 15 medical centers across the US and Canada from 1974 to 1979, yielding 24,989 patients. Of these, 2,099 eligible patients were selected for the randomized control trial (RCT) and were part of a comprehensive follow-up study. 780 of the 2,099 patients accepted the randomized treatment -- we refer to them as the experimental arm in our analysis. 1,319 patients refused the randomization and self-selected into treatment groups -- this is our observational arm. For our analysis, in the experimental arm, the treatment group is defined based on intent-to-treat by the original randomized assignment. Note that $23.5\%$ of patients assigned to medical therapy had bypass surgery within 5 years since their angina worsened. In the observational arm, a similar treatment group was identified as any patient who was selected for surgery within 90 days of enrollment or if their surgery was in the first year period (when 95% of CASS experimental arm surgeries were done). Any observational study patient who did not have early elective surgery were treated as the control group (or medical therapy arm). Here, we use all-cause mortality during the course of the study as the outcome of interest.

Analysis and Result

In this paper, we are interested in studying if the conditional ignorability of the observational sample as well as the external validity of the experimental sample holds. Furthermore, we are also interested in estimating the population average treatment effect of surgical intervention as compared to medical therapy.

Given that the same set of practitioners administers the treatments for both the observational as well as experimental samples, it is reasonable to assume that selection in the experimental (or observational) arm does not have a direct causal link to the outcome. That implies that Assumption A5 is likely to hold.

Our test allows us to study if either A3 or/and A4 are violated. If we find significant evidence for $\theta(t)\neq0$ for any $t\in\{0,1\}$, then either of these assumptions can be violated. The result is that our test fails to reject the null hypothesis and finds no evidence for the violation of these assumptions. The p-values corresponding to the tests are 0.290 and 0.915, respectively for $\theta(0)$ and $\theta(1)$. We also show the distributions of $\theta(0)$ and $\theta(1)$ in Figure (ref).

figure[figure omitted — 213 chars of source]

Note that failure to find evidence does not necessarily imply that there is no unobserved confounding. However, these results are in congruence with previous works dahabreh2020extending,olschewski1992analysis, suggesting that the experimental sample had external validity and the observational sample has conditional ignorability.

Next, we estimate the average treatment effect using our estimator and other estimators. Table (ref) presents the estimated average treatment effects using our estimator proposed in Equation (ref), the difference in means estimator for the experimental sample and an augmented inverse propensity score weighted (AIPW) estimator in the observational sample. Since our tests failed to reject, it is not surprising that the three estimates are effectively the same and consistent with the literature: the treatment effect is not significantly different than 0, i.e., the coronary bypass surgery neither helps nor hurts patients' chances of survival.

table[table omitted — 436 chars of source]

Data Description

The Lalonde data consist of a randomized control trial, the National Support Work Demonstration (NSW), that studied the effect of training programs on participant income levels lalonde1986. The NSW is often augmented with the Population Survey of Income Dynamics (PSID-2) dataset of observational control units dehejia1999; together, the two can serve as a benchmark for observational causal inference methods. While most estimation methods struggle with recovering the experimental point estimate using the joint dataset, some methods are able to recover the experimental ATE after pre-processing malts. We reanalyze the NSW and PSID-2 datasets to study if the experimental controls are comparable to observational controls; that is, if the experiment satisfies external validity (Assumption A3). We limit our analysis to the subsample of male household heads under the age of 55 and who are not retired by 1975. The outcome of interest is the income of participants in 1978. Furthermore, for both datasets, we use pre-treatment information about units' age, race, marital status, education, and income in 1975.

Analysis and Results

Unlike our analyses of the STAR and CASS datasets, the observational data in this analysis consists only of control units. Our analysis is thus limited to identifying a possible violation of experimental external validity (A3) by testing if $\theta(0)=0$. Our test rejects the null hypothesis (p-value = $6.2 \times 10^{-6}$), indicating that the NSW experiment lacks external validity (with respect to PSID-2). As such, it is not at all surprising that many observational causal inference approaches are unable to recover experimental ATE when using the observational control group. One notable exception to this behavior is using the matching algorithm MALTS malts.

We reanalyze the joint NSW and PSID-2 data using MALTS, a matching algorithm that learns a distance metric to guarantee tighter matches on more important covariates. In this approach, we estimate a matched group of size 10 for each unit in the sample and calculate a diameter of the matched group with respect to the learned metric. We plot these diameters in Figure (ref) and note the large gap between diameters of size 80 and 100. Using this heuristic we prune (eliminate) units from the analysis that have a diameter that's larger than 80; we emphasize that we do not use outcome information in order to prune. It is important to note that the pruned set has a significant number of control units from the observational arm along with units from the experimental arm as shown in Figure (ref) by the color of the marker.

After pruning, the test for violation of A3 then fails to reject the null (p-value = $0.79$ and Figure (ref) shows the point estimate of $\theta_0$ with respect to the reference null distribution before and after pruning). While the matched groups were pruned only based on the tightness of the match, the corresponding change in the test statistic suggests that it is possible that an unobserved confounder causing selection in the experiment is actually correlated with observed pre-treatment covariates. This provides a compelling explanation for why most observational approaches fail to recover the experimental ATE using an observational arm, while a flexible matching framework is able to do so.

figure[figure omitted — 441 chars of source]
figure[figure omitted — 278 chars of source]

Synthetic Data Experiment

In this section, we contrast our Double Machine Learning (DML) approach with the state-of-the-art data fusion methods that similarly aim to blend experimental and observational data for a more precise estimation of treatment effects, employing a straightforward yet illustrative synthetic data experiment. We begin by detailing the process of generating our synthetic data and describing the alternative baseline approaches. This is followed by an analytical comparison of these methodologies, highlighting how our DML strategy enhances the integration of insights from both experimental and observational studies.

\paragraph{Data Generative Procedure.} We generate the data that satisfies assumptions A1-A3 and A4 is violated. We consider the dataset with 3 pre-treatment covariates $X_0 \dots X_3$ and an unobserved covariate $U$ sampled identically and independently from the normal distribution $\mathcal{N}(1/2, 25)$. The study participation, corresponding treatment, and outcomes are determined as follows:

eqnarray*[eqnarray* omitted — 326 chars of source]

Note that, here, $X_1$ and $U$ are effect modifiers, and $X_1$ and $X_2$ is a confounder affecting the study participation $S$, treatment choice $T$, and outcome $Y$.

\paragraph{Baseline Approaches.} Here, we compared the performance of our double machine learning approach in estimating ATE with the 3 other baseline approaches: (i) augmented inverse probability weighted estimator just using the experimental data chernozhukov2018double, (ii) comprehensive cohort studies (CCS) described in lu2019causal, and (iii) integrative R-learner of heterogeneous treatment effects combining experimental and observational studies (integrative HTE) described in wu2022integrative.

\paragraph{Analysis.} We generate datasets of varying sample sizes from $n=250$ to $n=2000$. For the datasets with $n=2000$, our result in Figure (ref)(a) shows that, on average, our DML approach has the smallest mean squared error of all the other approaches. Further, while comparing the empirical bias and the corresponding variance, Figure (ref)(b) shows that our DML approach has the smallest spread of empirical bias centered around zero, corresponding to the tight estimates with the smallest standard errors.

figure[figure omitted — 493 chars of source]

Conclusion

With the expanding role of causal inference in all avenues of high-stakes decision-making, it is of fundamental importance that studies making causal claims are as robust to violations of critical assumptions as possible. Experimental studies, such as randomized trials, do not suffer from violations of ignorability assumptions but suffer from small samples and a lack of generalizability to larger populations of interest. On the other hand, while observational studies often involve larger and more representative samples, they are prone to violations of ignorability.

In this paper, we introduce methods that enable analysts to leverage both experimental and observational data simultaneously, allowing them to detect and correct violations of these assumptions. To detect the violations of ignorability in observational data or external validity in experimental data, we have proposed a statistical quantity that summarizes the extent of violations of ignorability, together with a root-n consistent estimator and falsification test for this quantity. To remedy the lack of conditional ignorability in observational data, we have proposed an estimator of treatment effects that simultaneously makes use of both the observational as well as experimental data. While our work proposes both a test and an estimator, they serve independent purposes, and we do not recommend users to use the same data for testing and estimation, as they can result in post-selection bias. Users may choose to split the dataset and use one half for testing and the other for estimation to ensure valid inference.

Under externally valid experimental data, our estimator is unbiased and consistent. Our methods take advantage of modern machine learning tools for outcome estimation and therefore can achieve an excellent degree of accuracy as shown by the synthetic data experiments. We have shown how our proposed tools can be applied to real-world data by re-analyzing the STAR project, and showing that observational samples are likely biased. We also demonstrate the performance of our approach by reanalyzing the CASS study, showing that the observational sample in this study likely satisfies conditional ignorability (see Appendix (ref)), and the Lalonde study data where the units in the experimental and observational studies are not exchangeable, in general (see Appendix (ref)).

Natural extensions of our methods include the formulation of alternative test statistics and quantities of interest that describe violations of ignorability, as well as considering methods to assess potential violations of other causal assumptions, such as SUTVA. Ultimately, the methods we have proposed are aimed at strengthening the trustworthiness and reliability of causal inference, as well as enabling analysts to take advantage of all the possible data sources they have access to.

Data Availability

The data used to demonstrate the functionality of our methodology is publicly available online. Readers can download (1) Project STAR data from \href{https://dataverse.harvard.edu/dataset.xhtml?persistentId=hdl