EconBase
← Back to paper

Combining Experimental and Observational Data for Identification and Estimation of Long-Term Causal Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

70,040 characters · 17 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Combining Experimental and Observational Data for Identification and Estimation of Long-Term Causal Effects

abstractWe study identifying and estimating the causal effect of a treatment variable on a long-term outcome using data from an observational and an experimental domain. The observational data are subject to unobserved confounding. Furthermore, subjects in the experiment are only followed for a short period; thus, long-term effects are unobserved, though short-term effects are available. Consequently, neither data source alone suffices for causal inference on the long-term outcome, necessitating a principled fusion of the two. We propose three approaches for data fusion for the purpose of identifying and estimating the causal effect. The first assumes equal confounding bias for short-term and long-term outcomes. The second weakens this assumption by leveraging an observed confounder for which the short-term and long-term potential outcomes share the same partial additive association with this confounder. The third approach employs proxy variables of the latent confounder of the treatment-outcome relationship, extending the proximal causal inference framework to the data fusion setting. For each approach, we develop influence function-based estimators and analyze their robustness properties. We illustrate our methods by estimating the effect of class size on 8th-grade SAT scores using data from the Project STAR experiment combined with observational data from the Early Childhood Longitudinal Study.\\ Keywords: Causal Inference; Data Fusion; Equi-Confounding Bias; Proximal Causal Inference; Influence Functions\\

Introduction

The gold standard for estimating the causal effect of a treatment variable on an outcome variable of interest is to conduct a randomized experiment. This is due to the fact that randomization ensures (conditional) exchangeability (also known as ignorability), which implies that the treatment and control groups are comparable. However, conducting such experiments can be costly and time-consuming, often resulting in limited or incomplete experimental data. In contrast, large-scale observational data is frequently available, yet there may be unobserved confounders of the treatment-outcome relationship in the setting, rendering causal inference impossible without positing extra assumptions. Given the limitations of both data sources, a natural question arises: Can we improve our causal inference by combining experimental and observational data?

In some applications, experiments can be run on the target variable of interest and data fusion is merely for the sake of improving estimation efficiency kallus2020role. However, in many settings, the experiment is run with another target variable different from the primary target. Especially, the practicalities of conducting an experiment with human subjects dictates that observations on subjects (and their compliance with the study protocol) only extends over a relatively short period of time from enrollment in the experiment. Hence, information is often missing on long-term outcomes (primary target) in randomized experiments. In this case, the experimental data alone cannot be used for estimating the causal effect of the treatment on the long-term outcome variable as it only includes the short-term outcome variable, and if the observational domain is confounded, the observational data alone cannot be used for causal inference either. This is the setup that we focus on in this work. An example of such a setup, discussed by athey2020combining, is the estimation of the effect of class size on eighth-grade test scores in New York schools, where we have access to observational data from New York schools, and Project STAR DVN/SIWH9F_2008 serves as the experimental dataset, including test scores only through third grade.

In their work, athey2020combining proposed a method for combining data from experimental and observational domains to enable causal inference. They showed that, under exchangeability-type assumptions for ensuring internal and external validity of the experimental data, along with an extra novel assumption termed latent unconfoundedness, the average treatment effect (ATE) on the long-term outcome in the observational data can be identified. In this paper, we first review their proposed approach and discuss the latent unconfoundedness assumption. We then propose three alternative approaches for data fusion to estimate ATE as well as the effect of treatment on the treated (ETT).

Our first proposed data fusion approach is based on assuming equal confounding bias for the short-term and long-term outcomes, which we refer to as the equi-confounding assumption. We consider both additive and quantile-quantile equi-confounding. Roughly speaking, the equi-confounding assumption posits that the magnitude of the confounding bias for the short-term and the long-term outcome variables are the same. This approach draws inspiration from the literature of difference-in-differences (DiD) framework card1990impact, angrist2008mostly, change-in-changes framework athey2006identification, and negative control-based causal inference lipsitch2010negative,tchetgen2014control,sofer2016negative. Our second proposed data fusion approach is based on assuming the existence of an observed confounder in the system called the bespoke instrumental variable (BSIV), such that the short-term and long-term potential outcome variables have the same partial additive association with that confounder. The existence of such a variable allows us to relax the equi-confounding assumption in the first approach by removing the need for assuming restrictions on the selection bias for the short-term and the long-term outcomes. Hence, in the presence of a BSIV in the setting, the researcher should prefer the use of this method over the equi-confounding method. This method builds on the BSIV causal inference framework recently introduced by richardson2022bespoke. Finally, our third proposed data fusion approach relies on the presence of a proxy variable of the latent confounder, which is independent of the outcome variables conditional on the treatment and all observed and unobserved confounder variables. This approach extends the proximal causal inference framework miao2018identifying,tchetgen2020introduction,cui2020semiparametric to the data fusion setting, but requires only a single proxy variable--unlike the standard proximal framework, which relies on two.

We formally establish that the unbiased information encoded in the experimental data enables the relaxation of key identification assumptions in standard DiD methods, the original BSIV approach, and proximal causal inference methods: Standard DiD methods applied to our setting would require that the treatment cannot causally impact the short-term outcome. However, as the short-term outcome is a post-treatment variable, this assumption may not hold. We demonstrate that by leveraging experimental data, we can relax this requirement by anchoring the short-term causal effect at that observed in the experimental sample, thus enabling identification of the treatment’s impact on the long-term outcome. Similarly, in the original BSIV causal inference framework, it is assumed that there exists a reference domain in which the treatment is not applied and hence, the outcome in that domain in fact represents the potential outcome under no treatments. In our setup, we relax this assumption by utilizing the internal validity of the experimental domain. Likewise, in the proximal approach, the standard framework would require exclusion restriction of no treatment effect on the short-term outcome. We demonstrate that by leveraging experimental data, such requirement can also be relaxed.

To the best of our knowledge, we are the first to propose nonparametric identification methods for long-term causal effects that allow for latent confounders influencing the treatment as well as both short- and long-term outcomes. After the release of the first draft of our work, imbens2022long also proposed an approach for identifying long-term causal effects based on the proximal causal inference framework. We discuss that work in the Supplementary Materials. Its main identification result relies on stronger assumptions; however, the authors also present an extension to their setup that resembles our proximal data fusion approach. Moreover, their method requires access to three short-term outcome variables, each influenced only by its immediate predecessor--an assumption that may not be feasible in many real-world applications.

The rest of the paper is organized as follows. We describe the model and parameters of interest in Section (ref). In Section (ref), we review the approach of athey2020combining and discuss their latent unconfoundedness assumption. Our proposed alternative methods, the equi-confounding, BSIV, and proximal causal data fusion approaches, are presented in Sections (ref), (ref), and (ref), respectively. In Section (ref), we focus on the estimation aspect of the parameters of interest, and propose influence function-based estimation strategies for each of our data fusion approaches and study the robustness properties of the proposed estimators. We evaluated our proposed methods on synthetic data in Section (ref). In Section (ref), we apply our proposed methods to estimate the effect of class size on long-term educational outcomes, measured by 8th-grade SAT scores, by combining data from the Project STAR experiment with observational data from the Early Childhood Longitudinal Study tourangeau2009eclsk. Our concluding remarks are provided in Section (ref). All the proofs are provided in the Supplementary Materials.

Problem Description

Let \( A \) denote a binary treatment variable, \( X \) the vector of pre-treatment covariates, \( M \) the short-term outcome, \( Y \) the long-term outcome, and \( U \) the set of unobserved (latent) confounders of the treatment--outcome relationship. Let \( G \in \{O, E\} \) indicate the data domain, with \( G = O \) corresponding to the observational domain and \( G = E \) to the experimental domain. We denote the set of all observed variables by \( V \). The potential short-term and long-term outcomes under treatment level \( a \in \{0,1\} \) are denoted by \( M^{(a)} \) and \( Y^{(a)} \), respectively. We observe independent and identically distributed (i.i.d.) data from \( \{A, X, M\} \) in the experimental domain, where treatment is (conditionally) randomized. In contrast, i.i.d. data from \( \{A, X, M, Y\} \) are available in the observational domain, where both short- and long-term outcomes are observed, but the treatment--outcome relationship is confounded by the unobserved variables \( U \). Thus, while the experimental data provide unconfounded information on the short-term outcome, only the observational data contain information about the long-term outcome, albeit subject to confounding.

Our objective is to identify two causal parameters of interest in the observational domain: the average treatment effect (ATE) in the observational population, \[ \theta_{\text{ATE}} = \mathbb{E}\left[Y^{(1)} - Y^{(0)} \mid G = O\right], \] and the effect of treatment on the treated (ETT), \[ \theta_{\text{ETT}} = \mathbb{E}\left[Y^{(1)} - Y^{(0)} \mid A = 1,\, G = O\right]. \] Clearly, these parameters are not identifiable from the experimental data alone, as the long-term outcome \( Y \) is unobserved in that domain, and also the information regarding the treatment effect in the experimental domain may be not relevant to the observational domain. Moreover, due to potential unobserved confounding, the parameters are also not identifiable from the observational data alone without additional assumptions. To proceed, we assume that the data-generating process satisfies the following conditions.

figure[figure omitted — 359 chars of source]
assumption[$\{X,U,G\}$--conditional exchangeability w.r.t. $A$] \[ A\perp\mkern-9.5mu\perp\{Y^{(a)},M^{(a)}\}\mid \{X,U,G\}\qquad \forall a\in\{0,1\}. \]
assumption[$X$--conditional exchangeability w.r.t. $A$ in the experimental domain] \[ A\perp\mkern-9.5mu\perp\{Y^{(a)},M^{(a)}\}\mid \{X,G=E\}\qquad \forall a\in\{0,1\}. \]
assumption[$X$--conditional exchangeability w.r.t. $G$] \[ G\perp\mkern-9.5mu\perp\{Y^{(a)},M^{(a)}\}\mid X\qquad \forall a\in\{0,1\}. \]

Assumption (ref) is a much milder version of the standard conditional exchangeability assumption hernan2020causal as it is stated conditional on any unobserved confounder $U$. Assumption (ref) is a context-specific conditional independence assumption. That is, a conditional independence which is only realized conditional on a certain event. We also posit the standard consistency and positivity assumptions. Figure (ref) represents a graphical model that satisfies Assumptions (ref)-(ref). Variables in gray circle are unobserved. Figures (ref) (a) and (ref) (b) represent the pooled dataset, Figures (ref) (c) and (ref) (d) represent the experimental dataset, which is the pooled data set conditioned on $G=E$, and Figures (ref) (e) and (ref) (f) represent the observational dataset, which is the pooled data set conditioned on $G=O$. Figures (ref) (b), (ref) (d), and (ref) (f) represent single world intervention graphs (SWIGs), which are graphical models obtained from the original graphs, which include potential outcome variables, and hence, can be used for representing conditional independences involving potential outcome variables. See richardson2013single for the definition and details.

Assumptions (ref) and (ref) are not sufficient for identification of the causal parameters of interest. In the following, we first review an approach proposed by athey2020combining for identification based on an extra assumption called latent unconfoundedness, and then propose our alternative approaches.

athey2020combining Approach

athey2020combining introduced an approach for combining experimental and observational data to identify the average treatment effect parameter, $\theta_{\text{ATE}}$. Their approach relies on Assumptions (ref) and (ref), along with an additional identifying condition stated below. As we show in this section, these assumptions also suffice to identify the effect of treatment on the treated, $\theta_{\text{ETT}}$.

assumption[Latent Unconfoundedness] \[ A \perp\mkern-9.5mu\perp Y^{(a)} \mid \{X, M^{(a)}, G = O\} \qquad \text{for all } a \in \{0,1\}. \]
theoremUnder Assumptions (ref)--(ref), the parameters $\theta_{\text{ATE}}$ and $\theta_{\text{ETT}}$ are identified. The identification formulae are presented in Supplementary Material (ref).
remarkOne can show that under Assumptions (ref)--(ref), the full distribution \( p(Y^{(a)} \mid G = O) \) is identified. Hence, the latent unconfoundedness assumption in athey2020combining is sufficiently strong to identify a broad class of causal estimands beyond mean contrasts, including functionals of the entire counterfactual distribution. The same holds for our proposed method in Section (ref) and the proximal data fusion approach in Section (ref).

{\bf Discussion of Assumption (ref).} Assume the data-generating process is governed by a nonparametric structural equation model. We consider two cases:

enumerate• The distribution is faithful spirtes2000causation to the causal model. That is, all conditional independence relations are reflected in causal relations among the variables and no conditional independence arises from specific cancellations or alignment of the modules in the causal model. Under faithfulness, a graphical model is a reliable diagnostic tool for the existence of a natural sequential data generating processes, in which edges represent direct causal relations. In this setting, Assumption (ref) implies the absence of a directed edge from the latent variable \( U \) to the outcome \( Y \). In other words, there must be no unobserved confounder that causally influences both \( A \) and \( Y \). Consequently, Assumption (ref) may be overly strong for many real-world settings, where violations of faithfulness are unlikely. • The distribution violates faithfulness. In this case, it is possible for \( A \) and \( Y \) to share a latent confounder, yet still satisfy Assumption (ref) due to structural cancellation. However, such scenarios are typically regarded as pathological or non-generic. An example of this case is discussed below.
exampleSuppose that in the observational domain, variables $M$ and $Y$ are generated via the following linear structural equation model: \[ M = \theta_M A+\gamma_M X+U,\qquad\qquad Y = \theta_Y A+\zeta_Y M+\gamma_Y X +\delta_Y U +\epsilon, \] such that $A\perp\mkern-9.5mu\perp\epsilon\mid \{X,U,G=O\}$. Importantly, the non-generic assumption here is excluding an independent exogenous noise from the structural equation corresponding to the variable $M$. Note that $Y^{(a)}=\theta_Y a+\zeta_Y M^{(a)}+\gamma_Y X+\delta_Y U+\epsilon = (\theta_Y+\zeta_Y\theta_M) a+(\gamma_Y+\zeta_Y\gamma_M) X+(\delta_Y+\zeta_Y) U+\epsilon$. Therefore, the assumption $A\perp\mkern-9.5mu\perp\epsilon\mid \{X,U,G=O\}$ implies that $A\perp\mkern-9.5mu\perp Y^{(a)}\mid \{X,U,G=O\}$. Also, $M^{(a)}=\theta_M a+\gamma_M X+U$ and hence $U=\theta_M a+\gamma_M X-M^{(a)}$. Therefore, $A\perp\mkern-9.5mu\perp Y^{(a)}\mid \{X,U,G=O\}$ implies that $A\perp\mkern-9.5mu\perp Y^{(a)}\mid \{X,M^{(a)},G=O\}$, which is Assumption (ref). Hence, the assumption basically implies that adjusting for $\{M^{(a)},X\}$ is equivalent to adjusting for $\{U,X\}$, i.e., roughly speaking, observing $\{M^{(a)},X\}$ is equivalent to observing $\{U,X\}$. Note that excluding an independent exogenous noise from the structural equation corresponding to the variable $M$ is essential for this equivalence to hold. Here, the deterministic dependence of $M$ on $X$, $A$ and $U$ violates faithfulness.

In the following three sections, we present our alternative identification approaches. In the main text, we focus on results for the parameter $\theta_{\text{ETT}}$, while the corresponding results for $\theta_{\text{ATE}}$ are provided in the Supplementary Materials.

Approach 1: Equi-Confounding Data Fusion

Our first proposed approach is based on the assumption that the conditional confounding bias is equal for the short-term and long-term outcomes. This restriction is inspired by assumptions used in identification strategies involving negative outcome controls lipsitch2010negative,tchetgen2014control,sofer2016negative, and is conceptually similar to the parallel trends assumption in the difference-in-differences (DiD) framework card1990impact,angrist2008mostly.

assumption[Conditional Additive Equi-Confounding Bias] With probability one, \begin{align*} \mathbb{E}[M^{(0)}\mid X,A=0,&G=O]-\mathbb{E}[M^{(0)}\mid X,A=1,G=O]\\ &=\mathbb{E}[Y^{(0)}\mid X,A=0,G=O]-\mathbb{E}[Y^{(0)}\mid X,A=1,G=O]. \end{align*}

Unlike the standard conditional exchangeability assumption, Assumption (ref) allows for the presence of latent confounders. It does not require that the treated and control groups have identical conditional potential outcome distributions. Rather, it asserts that the difference of the expected value of the short-term potential outcome across these two groups (that is the bias due to confounding on an additive scale) is the same as that of the long-term potential outcome variable. Equivalently, it implies that the expected change from the short-term to long-term potential outcome is the same in each stratum of $X$ across the treated and control groups.

exampleAssumption (ref) is satisfied if the data are generated from the following model: \[ M = g_M(A, X) + f_M(X, U) + \epsilon_M, \qquad \qquad Y = g_Y(A, X) + f_Y(X, M, U) + \epsilon_Y, \] where $\epsilon_M$ and $\epsilon_Y$ are independent noise terms and we have $\mathbb{E}[f(X,M,U)\mid X,A=1]=\mathbb{E}[f(X,M,U)\mid X,A=0]$, where $f(X,M,U)=f_Y(X,M,Y)-f_M(X,U)$, and $g_M$, $f_M$, $g_Y$, and $f_Y$ can be stochastic functions.
remarkAs noted above, Assumption (ref) is similar in flavor to the parallel trends assumption in the DiD framework. However, in that setting, the counterpart of $M$ is assumed to be a pre-treatment variable, and thus cannot be causally affected by the treatment. In contrast, in our setting, $M$ is post-treatment and may be influenced by $A$.

We now present our identification result for $\theta_{\text{ETT}}$ under conditional equi-confounding assumption.

theoremUnder Assumptions (ref), (ref), and (ref), the parameter $\theta_{\text{ETT}}$ is identified as: { \begin{equation} \begin{aligned} &\theta_{ETT} = \mathbb{E}[Y \mid A=1, G=O] + \mathbb{E}\left[ \frac{1}{p(A=1 \mid X, G=O)} \mathbb{E}[M \mid X, A=0, G=O] \mid A=1, G=O \right] \\ &\quad - \mathbb{E}\left[ \frac{1}{p(A=1 \mid X, G=O)} \mathbb{E}[M \mid X, A=0, G=E] + \mathbb{E}[Y \mid X, A=0, G=O] \mid A=1, G=O \right]. \end{aligned} \end{equation} }

The corresponding identification result for the parameter $\theta_{\text{ATE}}$ is provided in Supplementary Material (ref).

remarkWe have provided the unconditional counterpart of the conditional equi-confounding bias assumption in Assumption (ref) and the corresponding identification result in Supplementary Material (ref). Importantly, Assumptions (ref) and (ref) are not nested and do not imply one another. Each allows for different forms of heterogeneity and interactions among variables. Assumption (ref) posits that the confounding bias for the short-term and the long-term outcome variables are the same. However, it might be the case that the researcher does not believe that the bias for the two outcomes are equal marginally, but this equality holds in each stratum of $X$. In this case, Assumption (ref) is more appropriate. In this sense, Assumption (ref) may be viewed as a generally weaker assumption.

Quantile-Quantile Equi-Confounding Data Fusion

We note that the (conditional) additive equi-confounding bias assumption may be restrictive, as it requires the short-term and long-term outcomes to be measured on the same scale. While this is not a limitation in our specific application---where \( M \) and \( Y \) represent short- and long-term versions of the same outcome---it may be problematic in other settings. To address this, we propose a generalization of the additive equi-confounding framework, inspired by the changes-in-changes approach in panel data analysis athey2006identification and its analogue in the negative control inference literature sofer2016negative. This generalization also enables identification of causal parameters beyond mean-based quantities such as ATE and ETT. We present this extension in Supplementary Material (ref).

Approach 2: Bespoke IV Data Fusion

In this section, we show that the (conditional) equi-confounding bias assumption can be relaxed in the presence of a variable \( B \) among the observed pre-treatment covariates, provided it satisfies a specific condition regarding its association with the potential outcomes. The key idea is that, under such conditions, \( B \) can play a role analogous to that of an instrumental variable. This approach is inspired by the recently proposed Bespoke Instrumental Variable (BSIV) framework of richardson2022bespoke. Following that work, we refer to the variable \( B \) as a BSIV and refer to the resulting method as the BSIV data fusion approach. We present the framework for a binary BSIV, though as discussed in Remark (ref), the approach naturally extends to non-binary variables. To maintain consistency with previous notation, we denote the remaining observed covariates by \( X \). The requirements on the BSIV variable \( B \) are as follows.

assumption[BSIV Relevance] \[ \mathbb{E}[A \mid B=0, X, G=O] \neq \mathbb{E}[A \mid B=1, X, G=O]. \]
assumption[BSIV Partial Additive Equi-Association] With probability one, \begin{align*} \mathbb{E}[M^{(0)}\mid X,B=1,&G=O]-\mathbb{E}[M^{(0)}\mid X,B=0,G=O]\\ &=\mathbb{E}[Y^{(0)}\mid X,B=1,G=O]-\mathbb{E}[Y^{(0)}\mid X,B=0,G=O]. \end{align*}

Assumption (ref) is the standard, testable instrumental variable (IV) relevance condition, common to all IV frameworks. Assumption (ref) posits that the short-term and long-term potential outcomes share the same partial additive association with the covariate \( B \). This condition is a weaker version of Assumption (ref). While Assumption (ref) requires the selection bias to be equal for both \( M \) and \( Y \), Assumption (ref) does not constrain the selection bias directly. Instead, it focuses solely on the additive association of the potential outcomes with the observed variable \( B \), thereby giving the researcher flexibility to select a variable they believe is most likely to satisfy this condition. A similar assumption appears in the framework of richardson2022bespoke. However, their setting assumes that, in the reference population, treatment assignment is withheld through external intervention, so the counterfactual outcome under \( A = 0 \) coincides with the observed outcome.

Note that Assumption (ref) can be equivalently expressed as: \[ \mathbb{E}[M^{(0)} - Y^{(0)} \mid X, B=1, G=O] = \mathbb{E}[M^{(0)} - Y^{(0)} \mid X, B=0, G=O]. \] This resembles the restriction imposed on an instrumental variable in a standard IV framework when the outcome is defined as the difference \( Y - M \). However, unlike a classical IV, the covariate \( B \) in our setting does not satisfy the two key assumptions of standard IV analysis---namely, unconfoundedness and the exclusion restriction. This motivates the term bespoke instrumental variable for \( B \).

To investigate identification, we consider the following nonparametric reparametrization of the outcome regression function, which is a variation of the approach proposed in robins1994correcting, tchetgen2013alternative:

align*[align* omitted — 192 chars of source]

where

align*[align* omitted — 271 chars of source]

We use this reparametrization to identify the causal effect contrast function \( \beta_1 \), and in turn the parameter \( \theta_{\text{ETT}} \). An additional reparametrization is used to identify \( \theta_{\text{ATE}} \), as discussed in Supplementary Material (ref). By Assumption (ref), the term \( \mathbb{E}[Y^{(0)} - M^{(0)} \mid B = b, X, G = O] \) does not depend on \( b \). Therefore, for each value of \( X \), the reparametrization yields four equations and five unknowns: two for \( \beta_1(b, X) \), two for \( \gamma_0(b, X) \), and one for \( \mathbb{E}[Y^{(0)} - M^{(0)} \mid B = b, X, G = O] \). To achieve point identification, we impose one of two additional restrictions: either \( \beta_1 \) does not depend on \( b \), or \( \gamma_0 \) does not depend on \( b \). These are formalized below.

assumption[Partial Homogeneity of Causal Effect Contrast] With probability one, \begin{align*} &\mathbb{E}[Y^{(1)} - Y^{(0)} \mid A = 1, B = 1, X, G = O] - \mathbb{E}[M^{(1)} - M^{(0)} \mid A = 1, B = 1, X, G = O] \\ &= \mathbb{E}[Y^{(1)} - Y^{(0)} \mid A = 1, B = 0, X, G = O] - \mathbb{E}[M^{(1)} - M^{(0)} \mid A = 1, B = 0, X, G = O]. \end{align*}
assumption[Partial Homogeneity of Bias Contrast] With probability one, \begin{equation*} \begin{aligned} &\Big\{ \mathbb{E}[Y^{(0)} \mid A = 1, B = 1, X, G = O] - \mathbb{E}[Y^{(0)} \mid A = 0, B = 1, X, G = O] \Big\} \\ &\quad - \Big\{ \mathbb{E}[M^{(0)} \mid A = 1, B = 1, X, G = O] - \mathbb{E}[M^{(0)} \mid A = 0, B = 1, X, G = O] \Big\} \\ &= \Big\{ \mathbb{E}[Y^{(0)} \mid A = 1, B = 0, X, G = O] - \mathbb{E}[Y^{(0)} \mid A = 0, B = 0, X, G = O] \Big\} \\ &\quad - \Big\{ \mathbb{E}[M^{(0)} \mid A = 1, B = 0, X, G = O] - \mathbb{E}[M^{(0)} \mid A = 0, B = 0, X, G = O] \Big\}. \end{aligned} \end{equation*}

Note that neither Assumption (ref) nor Assumption (ref) restricts the set of causes or effects for the involved variables, and the variable \( B \) may act as a confounder of \( A \), \( M \), and \( Y \). Assumption (ref) posits that the difference between the conditional in-group causal effects for outcomes \( Y \) and \( M \) is invariant to \( B \). For instance, the ETT for the outcome \( Y - M \) does not vary with \( B \). A special case where this holds is when the conditional in-group causal effect for each of \( Y \) and \( M \) is individually independent of \( B \). Assumption (ref) requires that the selection bias for the contrast \( Y - M \) is constant across levels of \( B \). One special case where this assumption holds is when the selection bias for both \( Y \) and \( M \) is individually invariant to \( B \). Another special case is when the curly brackets on the right hand side are equal, and also the curly brackets on the left hand side are equal---i.e., both sides of the equation are zero. This situation corresponds precisely to the conditional equi-confounding bias condition stated in Assumption (ref). Therefore, Assumption (ref) is strictly weaker than Assumption (ref). Hence, in the presence of a BSIV the researcher should prefer using the BSIV data fusion approach over the equi-confounding approach.

theoremDefine $\pi(B,X) \coloneqq p(A=1 \mid B,X,G=O)$, and for $a,b \in \{0,1\}$ define \[ E_{ab}^O(X) \coloneqq \mathbb{E}[Y - M \mid A=a, B=b, X, G=O], \quad P_{ab}^O(X) \coloneqq p(A=a \mid B=b, X, G=O). \] \begin{itemize} • Under Assumptions (ref), (ref), (ref), (ref), and (ref), the parameter $\theta_{\text{ETT}}$ is identified by: { \begin{equation} \begin{aligned} \theta_{ETT} &= \mathbb{E}\Bigg[ \frac{\mathbb{E}[Y - M \mid B=1, X, G=O] - \mathbb{E}[Y - M \mid B=0, X, G=O]}{P_{11}^O(X) - P_{10}^O(X)} \\ &\quad - \frac{\mathbb{E}[M \mid A=0, B, X, G=E]}{\pi(B, X)} \ \Big| \ A=1, G=O \Bigg] + \frac{\mathbb{E}[M \mid G=O]}{p(A=1 \mid G=O)}. \end{aligned} \end{equation} } • Under Assumptions (ref), (ref), (ref), (ref), and (ref), the parameter $\theta_{\text{ETT}}$ is identified by: { \begin{equation} \begin{aligned} \theta_{ETT} &= \mathbb{E}\Bigg[ \big\{ E_{11}^O(X) - E_{01}^O(X) - E_{10}^O(X) + E_{00}^O(X) \big\} B + E_{10}^O(X) - E_{00}^O(X) \\ &\quad - \frac{E_{01}^O(X) \!-\! E_{00}^O(X)}{P_{01}^O(X) \!-\! P_{00}^O(X)} \!+\! \mathbb{E}[M \!\mid \!A=1, B, X, G=O] \!-\! \frac{\mathbb{E}[M \!\mid \!A=0, B, X, G=E]}{\pi(B, X)} \\ &\quad + \frac{\mathbb{E}[M \mid A=0, B, X, G=O](1 - \pi(B, X))}{\pi(B, X)} \ \Big| \ A=1, G=O \Bigg]. \end{aligned} \end{equation} } \end{itemize}

See Supplementary Material (ref) for the corresponding identification results for the ATE.

remarkAssumption (ref) is formulated as $A \perp\mkern-9.5mu\perp \{Y^{(a)}, M^{(a)}\} \mid B, X, G=E$, and Assumption (ref) is expressed as $G \perp\mkern-9.5mu\perp M^{(a)} \mid B, X$. However, it is worth noting that the identification results in Theorem (ref) only require the weaker conditions, e.g., $\mathbb{E}[M^{(a)} \mid X, B, G=O] = \mathbb{E}[M^{(a)} \mid X, B, G=E]$, which is an equality of conditional expectations rather than a full conditional independence assumption.
remarkAs noted earlier, the bespoke instrumental variable \( B \) need not be binary. In Supplementary Material (ref), we provide the counterparts of the assumptions in our BSIV data fusion framework for non-binary \( B \). The identification arguments extend straightforwardly to this setting with no substantive changes to the proofs.

Approach 3: Proximal Data Fusion

In this section, we present our third identification strategy, which leverages the presence of a variable that serves as a proxy for the latent confounder. This approach is inspired by the proximal causal inference framework miao2018identifying, tchetgen2020introduction, cui2020semiparametric. However, unlike that framework---which requires two distinct proxies---we only require a single proxy variable.\footnote{A comparison to other proximal setups is presented in Supplementary Material (ref).} Specifically, we assume the existence of a variable \( Z \) in the observational domain that satisfies the following condition.

assumption[Existence of Proxy Variable] There exists a proxy variable \( Z \) for the latent confounder in the observational domain such that \[ Z \perp\mkern-9.5mu\perp \{M, Y\} \mid \{U, X, A, G=O\}. \]

Figure (ref) provides an example of a causal graph that satisfies this proxy variable condition. Importantly, no assumptions are required about \( Z \) in the experimental domain---indeed, \( Z \) need not even exist in that domain.

remarkFigure (ref) illustrates just one of many possible graphical structures that can satisfy Assumption (ref). For example, the assumption still holds if the edge from \( Z \) to \( A \) is reversed, making \( Z \) a post-treatment variable. Alternatively, the edge between \( Z \) and \( A \) can be removed entirely, or \( Z \) and \( A \) can share a latent confounder. These alternatives illustrate the flexibility available to the researcher in selecting a variable to serve as the proxy \( Z \).

In order to achieve identifiability, we impose the following additional assumption.

assumption\begin{itemize} • For any square-integrable function \( g \), and for all \( a \) and \( x \), if \( \mathbb{E}[g(U) \mid Z, a, x, G=O] = 0 \) almost surely, then \( g(U) = 0 \) almost surely. • There exists a function \( h(m,a,x) \), called the outcome bridge function, that solves the following integral equation: \begin{equation} \mathbb{E}[Y \mid Z, A, X, G=O] = \mathbb{E}[h(M, A, X) \mid Z, A, X, G=O]. \end{equation} \end{itemize}

We now state the identification result.

theoremUnder Assumptions (ref)--(ref), (ref), and (ref), the parameter \( \theta_{\text{ETT}} \) is identified as: \begin{align} \theta_{ETT} &= \frac{\mathbb{E}[Y \mid G=O]}{p(A=1 \mid G=O)} - \frac{\mathbb{E}\big[\mathbb{E}[h(M,0,X) \mid A=0, X, G=E] \mid G=O\big]}{p(A=1 \mid G=O)}. \end{align}

See Supplementary Material (ref) for the corresponding identification results for the ATE.

figure[figure omitted — 1,227 chars of source]

An Alternative Proximal Identification Method

We now present an alternative proximal identification strategy based on a different set of completeness and bridge function assumptions.

assumption\begin{itemize} • For any square-integrable function \( g \), and for all \( a \) and \( x \), if \( \mathbb{E}[g(U) \mid M, a, x, G=O] = 0 \) almost surely, then \( g(U) = 0 \) almost surely. • There exists a function \( q(z, a, x) \), called the treatment bridge function, that solves the following integral equation: \begin{equation*} \mathbb{E}[q(Z, A, X) \mid M, A, X, G=O] = \frac{p(M \mid A, X, G=E)}{p(M \mid A, X, G=O) \cdot p(A \mid X, G=O)}. \end{equation*} \end{itemize}

The corresponding identification result is as follows.

theoremUnder Assumptions (ref)--(ref), (ref), and (ref), the parameter \( \theta_{\text{ETT}} \) is identified as: \begin{align} \theta_{ETT} &= \frac{\mathbb{E}[Y \mid G=O]}{p(A=1 \mid G=O)} - \frac{\mathbb{E}[I(A=0) Y q(Z, A, X) \mid G=O]}{p(A=1 \mid G=O)}. \end{align}

See Supplementary Material (ref) for the corresponding identification results for the ATE.

Estimation Strategies

In this section, we turn to the problem of estimation and propose influence function (IF)-based estimators for the parameters of interest under each of our data fusion frameworks. IF-based estimators are advantageous in that their bias is of second order with respect to the error in estimating the nuisance functions—an important property in the presence of complex data-generating processes. Furthermore, under certain convergence rate conditions on the nuisance function estimators—which are weaker than those required for non-IF-based estimators—the resulting IF-based estimator is consistent and asymptotically normal. This enables valid inference without requiring resampling methods such as the bootstrap.

An additional benefit is that our IF-based estimators exhibit multiple robustness: the estimator remains unbiased even when some of the nuisance functions are misspecified. While the asymptotic normality of IF-based estimators is now standard in the semiparametric literature (see, e.g., bickel1993efficient,newey1990semiparametric,tsiatis2007semiparametric,robins2017minimax,kennedy2024semiparametric), the analysis of multiple robustness is more context-specific and depends on the parameter being estimated. Therefore, we will focus on establishing multiple robustness results specific to each framework. As in the previous sections, we only provide the results regarding the ETT in the main text; the corresponding results for the ATE are provided in Supplementary Material (ref).

In the subsections that follow, we derive the influence functions and analyze the robustness properties of the proposed estimators under each data fusion framework. Given an IF-based moment function, we adopt the cross-fitting procedure of chernozhukov2018double to decouple the estimation of nuisance functions from that of the target parameter. This procedure allows for weaker regularity conditions on the nuisance estimators while still ensuring asymptotic normality. The estimation procedure proceeds as follows:

itemize• {\bf Step 1.} Partition the sample into \( L \) equally sized folds: \( \{I_1, \ldots, I_L\} \). • {\bf Step 2.} For each \( \ell \in \{1, \ldots, L\} \), estimate the set of nuisance functions \( \zeta^{\text{approach}} \) using the data from all folds except \( I_\ell \), yielding \( \hat\zeta_\ell^{\text{approach}} \). • {\bf Step 3.} For each \( \ell \in \{1, \ldots, L\} \), estimate the parameter of interest \( \hat\psi_\ell \) by solving: \[ \frac{1}{|I_\ell|}\sum_{i \in I_\ell} \Phi^{\text{approach}}(V_i; \hat\zeta_\ell^{\text{approach}}, \hat\psi_\ell) = 0, \] where \( \Phi^{\text{approach}} \) is the IF-based moment function, which differs across our approaches. • {\bf Step 4.} Define the final estimator of the parameter as the average across folds: \begin{center} $\hat\psi^{\text{approach}} = \frac{1}{L} \sum_{\ell=1}^L \hat\psi_\ell$. \end{center}

Estimation under Equi-Confounding Data Fusion

In this subsection, we consider estimation under the equi-confounding data fusion framework, focusing on the case of conditional equi-confounding. Recall from Theorem (ref) that, under the assumptions of this framework, for the parameter $\theta_{\text{ETT}}$, we only need to focus on the right hand side of Equation (ref), which is a functional of the observed data distribution. We denote this functional by $\psi^{\text{equi}}_{\text{ETT}}$. One can use this functional to construct a simple plug-in estimator. However, such an estimator generally suffers from first-order bias with respect to the error in the estimation of the nuisance functions. Hence, as mentioned earlier, we use an IF-based estimator instead.

We begin by deriving the influence function for the parameter $\psi^{\text{equi}}_{\text{ETT}}$. Let the collection of nuisance functions be defined as: $\zeta^{\text{equi}}:= \Big\{ \mu^{G=E}_M(A=\cdot,X=\cdot)\coloneqq \mathbb{E}[M\mid X=\cdot,A=\cdot,G=E], \mu^{G=O}_M(A=\cdot,X=\cdot)\coloneqq \mathbb{E}[M\mid X=\cdot,A=\cdot,G=O], \mu^{G=O}_Y(A=\cdot,X=\cdot)\coloneqq \mathbb{E}[Y\mid X=\cdot,A=\cdot,G=O], \pi^{G=E}(\cdot)\coloneqq p(A=1\mid X=\cdot,G=E), \pi^{G=O}(\cdot)\coloneqq p(A=1\mid X=\cdot,G=O), p(G=E\mid X=\cdot) \Big\}$.

theoremUnder a nonparametric model, the efficient influence function for the parameter $\psi^{\text{equi}}_{\text{ETT}}$ is given by: { \begin{align*} &IF_{\psi^{equi}_{ETT}}(V)= \frac{1}{p(A=1,G=O)}\Big\{ \frac{I(G=O)I(A=0)}{1-\pi^{G=O}(X)}\{M-\mu^{G=O}_M(A=0,X)\}\\ &\qquad-\frac{I(G=E)I(A=0)}{1-\pi^{G=E}(X)}\left\{\frac{1}{p(G=E\mid X)}-1\right\}\{M-\mu^{G=E}_M(A=0,X)\}\\ &\qquad+\frac{I(G=O)I(A=0)\pi^{G=O}(X)}{1-\pi^{G=O}(X)}\{Y-\mu^{G=O}_Y(A=0,X)\}\\ &\qquad+I(G=O)\Big[ \mu^{G=O}_M(A=0,X) \!-\!\mu^{G=E}_M(A=0,X)\!+\!I(A=1)\{Y\!-\!\mu^{G=O}_Y(A=0,X)\!-\!\psi^{equi}_{ETT}\} \Big]\Big\}. \end{align*} }

Let $\hat\zeta^{\text{equi}}$ be an estimator for $\zeta^{\text{equi}}$, and define the moment function $\Phi^{\text{equi}}(V; \hat\zeta^{\text{equi}}, \hat\psi)$ to be the expression for $IF_{\psi^{\text{equi}}_{\text{ETT}}}(V)$ with nuisance components replaced by their corresponding estimates from $\hat\zeta^{\text{equi}}$, and $\psi^{\text{equi}}_{\text{ETT}}$ replaced by $\hat\psi$. We use this moment function in the cross-fitting procedure to obtain the estimator $\hat\psi^{\text{equi}}_{\text{ETT}}$. We have the following multiple robustness result for $\hat\psi^{\text{equi}}_{\text{ETT}}$.

propositionThe IF-based estimator $\hat\psi^{\text{equi}}_{\text{ETT}}$ is multiply robust in the sense that it is unbiased if at least one of the following subsets of nuisance functions is correctly specified: (i) $\{\mu^{G=E}_M,\mu^{G=O}_M,\mu^{G=O}_Y\}$; (ii) $\{\pi^{G=E},\pi^{G=O},p(G=E\mid X=\cdot)\}$; (iii) $\{\mu^{G=E}_M,\pi^{G=O}\}$; (iv) $\{\pi^{G=E},\mu^{G=O}_M,\mu^{G=O}_Y,p(G=E\mid X=\cdot)\}$.

Estimation under Bespoke IV Data Fusion

In this subsection, we consider estimation under the BSIV data fusion framework. Recall from Theorem (ref) that, under the assumptions of this framework, for the parameter $\theta_{\text{ETT}}$, we only need to focus on the right hand sides of Equations (ref) and (ref), which are functionals of the observed data distribution. We denote these functionals by $\psi^{\text{bsiv1}}_{\text{ETT}}$ and $\psi^{\text{bsiv2}}_{\text{ETT}}$, respectively.

We begin by deriving the influence functions for the parameters. Let the collection of nuisance functions be defined as: $\zeta^{bsiv}:= \Big\{ E_{ab}^O(X):=\mathbb{E}[Y-M\mid A=a, B=b,X,G=O], e_{b}^O(X):=\mathbb{E}[Y-M\mid B=b,X,G=O], M_{a}^{E}(B,X):=\mathbb{E}[M\mid A=a,B,X,G=E], \mu_{b}^O(X):=\mathbb{E}[M\mid B=b,X,G=O], \pi^O(X):=p(A=1\mid X,G=O), P_{ab}^O(X):=p(A=a\mid B=b,X,G=O), P_{ab}^E(X):=p(A=a\mid B=b,X,G=E), \rho_{b}^O(X):=p(B=b\mid X,G=O), \rho_{b}^E(X):=p(B=b\mid X,G=E), \tau(B,X):=p(G=E\mid B,X) \Big\}$.

theoremUnder a nonparametric model, the efficient influence functions for the parameters $\psi^{\text{bsiv1}}_{\text{ETT}}$ and $\psi^{\text{bsiv2}}_{\text{ETT}}$ are given by $IF_{\psi^{\text{bsiv1}}_{\text{ETT}}}(V)$ and $IF_{\psi^{\text{bsiv2}}_{\text{ETT}}}(V)$, respectively, where the closed-form expressions for the influence functions are deferred to Supplementary Material (ref).

Let $\hat\zeta^{\text{bsiv}}$ be an estimator for $\zeta^{\text{bsiv}}$, and define the moment functions $\Phi^{\text{bsiv1}}(V; \hat\zeta^{\text{bsiv}}, \hat\psi)$ and $\Phi^{\text{bsiv2}}(V; \hat\zeta^{\text{bsiv}}, \hat\psi)$ to be the expressions for $IF_{\psi^{\text{bsiv1}}_{\text{ETT}}}(V)$ and $IF_{\psi^{\text{bsiv2}}_{\text{ETT}}}(V)$ with nuisance components replaced by their corresponding estimates from $\hat\zeta^{\text{bsiv}}$, and $\psi^{\text{bsiv1}}_{\text{ETT}}$ and $\psi^{\text{bsiv2}}_{\text{ETT}}$ replaced by $\hat\psi$. We use these moment functions in the cross-fitting procedure to obtain the estimators $\hat\psi^{\text{bsiv1}}_{\text{ETT}}$ and $\hat\psi^{\text{bsiv2}}_{\text{ETT}}$. We have the following multiple robustness result for these estimators.

propositionThe IF-based estimator $\hat\psi^{\text{bsiv1}}_{\text{ETT}}$ is multiply robust in the sense that it is unbiased if at least one of the following subsets of nuisance functions is correctly specified:\\ (i) $\{P_{AB}^O(X),\mathbb{E}[M\mid B,X,G=O],\mathbb{E}[Y\mid B,X,G=O],\mathbb{E}[M\mid A,B,X,G=E]\}$;\\ (ii) $\{\tau(B,X),\rho_{B}^O(X),\rho_{B}^E(X),P_{AB}^O(X),P_{AB}^E(X)\}$;\\ (iii) $\{\tau(B,X),P_{AB}^O(X),P_{AB}^E(X),\mathbb{E}[M\mid B,X,G=O],\mathbb{E}[Y\mid B,X,G=O]\}$;\\ (iv) $\{\rho_{B}^O(X),\rho_{B}^E(X),P_{AB}^O(X),\mathbb{E}[M\mid A,B,X,G=E]\}$.\\ Same result holds for the estimator $\hat\psi^{\text{bsiv2}}_{\text{ETT}}$ but with replacing outcome regression functions with $\mathbb{E}[M\mid A,B,X,G=E]$, $\mathbb{E}[M\mid A,B,X,G=O]$, and $\mathbb{E}[Y\mid A,B,X,G=O]\}$.

Estimation under Proximal Data Fusion

In this subsection, we consider estimation under the proximal data fusion framework. As in our other two frameworks, we propose IF-based estimator. This estimator requires the estimation of both bridge functions $h$ and $q$. As seen earlier, these nuisance functions are solutions to conditional moment equations, and hence they cannot be estimated by a simple standard regression. Therefore, in Subsection (ref), we also present simpler (yet not robust) estimation strategies which require the estimation of only one of the bridge functions. We will discuss the estimation of bridge functions in Subsection (ref).

Recall from Theorem (ref) that, under the assumptions of this framework, for the parameter $\theta_{\text{ETT}}$, we only need to focus on the right hand side of Equation (ref), which is a functional of the observed data distribution. We denote this functional by $\psi_{\text{ETT}}^{\text{proxy}}$. We begin by deriving the influence function for this parameter. We assume that the integral equations in Assumptions (ref) $(ii)$ and (ref) $(ii)$ have unique solutions $h(\cdot)$ and $q(\cdot)$, and let the collection of nuisance functions be defined as $\zeta^{\text{proxy}}:=\Big\{ h,q, p(M=\cdot\mid A=\cdot,X=\cdot,G=E), p(A=1\mid X=\cdot,G=E), p(G=E\mid X=\cdot)\Big\}$.

theoremUnder a semiparametric model that Assumptions (ref) $(ii)$ and (ref) $(ii)$ have unique solutions, an influence function of the parameter $\psi_{\text{ETT}}^{\text{proxy}}$ is given by { \begin{align*} &IF_{\psi_{ETT}^{proxy}}(V) =\frac{1}{p(G=O)p(A=1\mid G=O)}\Big\{I(G=O)\big\{Y-I(A=0)q(Z,0,X)\{Y-h(M,0,X)\}\\ &-\eta(0,X)-I(A=1)\psi_{ETT}^{proxy}\big\}\\ &\quad-\frac{I(G=E)I(A=0)}{1-p(A=1\mid X,G=E)} \{h(M,0,X)-\eta(0,X)\}\{\frac{1}{p(G=E\mid X)}-1\}\Big\}, \end{align*} } where $\eta(a,x)\coloneqq\mathbb{E}[h(M,A,X)\mid A=a,X=x,G=E]$.

Estimation Strategies for ETT

Let $\hat\zeta^{\text{proxy}}$ be an estimator for $\zeta^{\text{proxy}}$. Based on identification formulae (ref) and (ref), along with Theorem (ref), we propose the following estimation strategies for the parameter $\theta_{\text{ETT}}$. Strategies 1–3 require estimation of only one of the bridge functions, $h$ or $q$, whereas Strategy 4, the IF-based estimator, requires estimation of both $h$ and $q$.

itemize• {\bf Estimation Strategy 1.} From (ref), define the moment function as, { \begin{align*} \Phi^{proxy1}(V;\hat{\zeta}^{proxy},\hat\psi) \!:=\! \frac{I(G=O)}{p(G=O)p(A=1\!\mid \!G=O)}\{Y \!-\!\sum_m\hat{h}(m,0,X)\hat{p}(m\!\mid \!0,X,G=E)\}\!-\!\hat\psi. \end{align*} } • {\bf Estimation Strategy 2.} Also from (ref), define the moment function as, { \begin{align*} \Phi^{proxy2}(V;\hat{\zeta}^{proxy},\hat\psi) :=& \frac{1}{p(G=O)p(A=1\mid G=O)}\Big\{I(G=O)Y\\ &-\frac{I(G=E)I(A=0)}{1-\hat{p}(A=1\mid X,G=E)} \hat{h}(M,0,X)\{\frac{1}{\hat{p}(G=E\mid X)}-1\}\Big\}-\hat\psi. \end{align*} } • {\bf Estimation Strategy 3.} From (ref), define the moment function as, { \begin{align*} \Phi^{proxy3}(V;\hat{\zeta}^{proxy},\hat\psi) :=& \frac{I(G=O)}{p(G=O)p(A=1\mid G=O)}\{Y-I(A=0)Y\hat{q}(Z,0,X)\}-\hat\psi. \end{align*} } \begin{sloppypar} • {\bf Estimation Strategy 4 (IF-based Strategy).} Define the moment function $\Phi^{\text{proxy}}(V; \hat\zeta^{\text{proxy}}, \hat\psi)$ to be the expression for $IF_{\psi^{\text{proxy}}_{\text{ETT}}}(V)$ with nuisance components replaced by their corresponding estimates from $\hat\zeta^{\text{proxy}}$, and $\psi^{\text{proxy}}_{\text{ETT}}$ replaced by $\hat\psi$. \end{sloppypar}

These moment functions are implemented in the cross-fitting procedure to obtain $\hat\psi^{\text{proxy1}}_{\text{ETT}}$, $\hat\psi^{\text{proxy2}}_{\text{ETT}}$, $\hat\psi^{\text{proxy3}}_{\text{ETT}}$, and the IF-based estimator $\hat\psi^{\text{proxy}}_{\text{ETT}}$.

propositionThe IF-based estimator $\hat\psi^{\text{proxy}}_{\text{ETT}}$ is multiply robust in the sense that it is unbiased if at least one of the following subsets of nuisance functions is correctly specified: (i) $\{h,p(M=\cdot\mid A=0,X=\cdot,G=E))\}$; (ii) $\{h,p(A=1\mid X=\cdot,G=E),p(G=E\mid X=\cdot)\}$; (iii) $\{q,p(A=1\mid X=\cdot,G=E),p(G=E\mid X=\cdot)\}$.

Estimating the Bridge Functions

In all proposed estimation strategies, estimation of at least one of the nuisance functions \(h\) or \(q\) is required to estimate the target parameter. However, these nuisance functions are defined as solutions to conditional moment (integral) equations and therefore cannot be obtained via simple regression. A recent line of work proposes nonparametric adversarial (minimax) estimators for such equations dikkala2020minimax; these ideas have been adapted to the original semiparametric proximal causal inference framework in ghassami2022minimax,kallus2021causal. We employ the same technique to estimate the bridge functions \(h\) and \(q\) in our proximal data fusion setup.

To proceed, note that the bridge function \(q\) satisfies the following conditional moment equation, which we use to design an estimator for \(q\).

propositionThe bridge function \(q\) satisfies { \begin{equation} \mathbb{E}\!\left[ \frac{I(G=O)}{p(G=O\mid X)}\,q(Z,A,X) - \frac{I(G=E)}{p(A,G=E\mid X)} \,\Bigm|\, M,A,X \right] = 0. \end{equation} }

Let \(\mathcal{H}\), \(\mathcal{Q}\), and \(\mathcal{F}\) be normed function spaces. Based on the conditional moment equations (ref) and (ref), we propose the following regularized minimax estimators for the bridge functions \(h\) and \(q\): {

align*[align* omitted — 618 chars of source]

} We refer the reader to dikkala2020minimax,ghassami2022minimax for convergence analyses of these minimax estimators.

Simulation Studies

table[table omitted — 644 chars of source]
table[table omitted — 939 chars of source]
table[table omitted — 1,416 chars of source]

We evaluated our BSIV and proximal data fusion frameworks on synthetic data for estimating the ETT. We omit the equi-confounding framework because, in the presence of a BSIV, it is nested as a special case of the BSIV framework. Binary variables were generated from logistic (expit) models, and continuous variables from normal models. To enable a comparison between the proximal and BSIV frameworks, data-generating parameters were chosen so that the assumptions of both frameworks hold; see Supplementary Material (ref) for full details.

We considered sample sizes \(n \in \{1000, 2000, 4000\}\) and conducted 300 Monte Carlo replications for each \(n\). For cross-fitting, we used four folds. Hyper-parameters for nuisance estimators were tuned via cross-validation with a 75/25 train–validation split. For the BSIV approach, we report both plug-in and IF-based estimators; for the proximal approach, we consider the four estimation strategies described in Section (ref).

Table (ref) compares the IF-based estimators for the BSIV and proximal approaches. The BSIV estimator exhibits slightly lower bias but notably higher variance. Coverage for both methods approaches the nominal \(95\%\) as \(n\) increases.

We also conducted a robustness study; results for the BSIV and proximal approaches appear in Tables (ref) and (ref), respectively. Each scenario corresponds to a setting in which specific subsets of nuisance functions (matching the cases in Propositions (ref) and (ref)) are correctly specified; see Supplementary Material (ref) for details. As expected, the two IF-based estimators display multiple robustness: whenever the corresponding subset of nuisance functions is correctly specified, the estimators have small bias that diminishes with increasing \(n\).

Application: Effect of Class Size on SAT Scores

table[table omitted — 479 chars of source]

We applied our proposed methods to evaluate the effect of class size on long-term educational outcomes. The observational (target) data come from the Early Childhood Longitudinal Study, Kindergarten Class of 1998–99 (ECLS-K) tourangeau2009eclsk, which tracks a nationally representative cohort from kindergarten through eighth grade. As the experimental domain, we use the Project STAR experiment DVN/SIWH9F_2008—a large-scale randomized study conducted in Tennessee. Both datasets report class size; following the Project STAR design, we define small classes (treatment) as those with at most 19 students and regular classes (control) as those with more than 19 students. We use third-grade Mathematics SAT scores as short-term outcomes and eighth-grade Mathematics SAT scores as long-term outcomes. We used observed covariates gender, ethnicity, access to free lunch, and school location (rural, suburban, urban). Family socioeconomic status (SES) is a plausible unobserved confounder; school location is considered a potential proxy for SES, as higher-SES families are more likely to live in suburban or urban areas. To ensure support overlap across domains, we restrict the short-term score to the 570–630 range and discretize it into four categories. Restricting to complete cases yields 2,172 observations from ECLS-K and 2,584 from Project STAR.

sloppyparWe apply only proximal estimation strategies 1 and 2. We do not employ our BSIV approach because we lack a covariate that plausibly satisfies the BSIV conditions; for example, free-lunch status and school location are correlated with factors that drive changes in achievement over time, making equal partial associations with short- and long-term outcomes unlikely. We also do not use proximal methods relying on the bridge function \(q(\cdot)\) because Assumption (ref)(i) requires the third-grade SAT score to be sufficiently informative about SES, which may not hold. Since all covariates and the proxy variable are discrete, the outcome bridge \(h\) admits a closed-form solution miao2018identifying: $h(m,a,x)\;=\;\mathbb{E}\!\left[Y \,\middle|\, B, A=a, X=x, G=O\right]\; P\!\left(M \,\middle|\, B, A=a, X=x, G=O\right)^{\dagger}$, where \({\dagger}\) denotes the matrix pseudoinverse; here \(h(m,a,x)\) and \(\mathbb{E}[Y \mid B, A=a, X=x, G=O]\) are row vectors of sizes \(|\mathcal{M}|\) and \(|\mathcal{B}|\), respectively, and \(P(M \mid B, A=a, X=x, G=O)\) is a \(|\mathcal{M}|\times|\mathcal{B}|\) probability matrix.

We compare our estimators with those of athey2020combining, the equi-confounding estimator, and a naive IF-based estimator that assumes no unobserved confounding, and is based on the influence function of the parameter, $\psi_{\text{ETT, naive}}=\mathbb{E}[Y\mid A=1,G=O]-\mathbb{E}[\mathbb{E}[Y\mid X,A=0,G=O]\mid A=1,G=O]$ (which is equal to the true ETT if there were no latent confounders present in the system). Results are summarized in Table (ref). Our estimators indicate a positive effect of smaller class sizes on SAT scores, consistent with krueger1999experimental, which found that smaller classes improve student test performance relative to regular-sized classes. By contrast, the estimator of athey2020combining and the naive IF-based estimator yield a negative effect or an effect close to zero. Yet these are unlikely to be plausible due to the presence of unobserved confounder in the setting which is likely to be directly affecting the long-term outcome, and hence violating the assumptions of these two frameworks. Moreover, the equi-confounding IF-based estimator also yields a negative effect which is not plausible as the association with the potential outcome is likely to change over time.

Conclusion and Discussion

In many real-world settings, available observational data are confounded by latent variables and therefore cannot be used to identify the causal effect of a treatment on an outcome of interest. At the same time, experimental data may be available, but due to practical constraints, the observed outcome may only reflect a short-term version of the long-term outcome of interest. Individually, neither the observational nor the experimental dataset suffices to identify the causal parameter. This raises the central question: can we combine information from the two sources to achieve identification? In this work, we proposed three data fusion frameworks under which the long-term causal effect can be identified: (1) Equi-confounding method, which assumes equal confounding bias for the short-term and long-term outcomes; (2) BSIV method, which leverages an observed confounder for which the short-term and long-term potential outcomes share the same partial additive association; (3) Proximal method, which relies on a proxy variable of the latent confounder in the treatment-outcome relationship, extending the proximal causal inference framework to the data fusion setting. For each approach, we developed influence function-based estimation strategies and analyzed the robustness properties of the resulting estimators.

A natural question is how a practitioner should choose among our three proposed data-fusion strategies. The main requirements of both equi-confounding and BSIV approaches for connecting \(M\) and \(Y\) are of the equal additive-association nature. Equi-confounding approach requires this in terms of requiring equal association of the treatment variables with the short-term and long-term potential outcomes. On the other hand, the BSIV approach allows the treatment variable to be substituted with some confounder $B$ that the practitioner believes is more likely to satisfy the equal additive-association restriction. Moreover, as mentioned earlier, the partial homogeneity assumption of the BSIV approach is weaker than the assumption of the equi-confounding approach. Therefore, when a credible BSIV exists, practitioners should prefer BSIV over equi-confounding. The proximal data fusion approach takes on a more generalized method for connecting \(M\) and \(Y\) by modeling their potentially complex and non-linear relation beyond additive association via bridge functions that solve conditional moment (integral) equations. This yields substantial model flexibility and generality compared to the BSIV approach. The trade-offs are practical: identifying a valid proxy \(Z\) may be harder than finding a credible BSIV, and because bridge functions are not structural outcome models, parametric forms are difficult to justify—hence they are best estimated nonparametrically (as in Section (ref)). In practice, the main challenge is constructing and reliably estimating the bridge function. A table summarizing the comparison of the three proposed methods is provided in Supplementary Material (ref).

center[center omitted — 167 chars of source]

\allowdisplaybreaks