EconBase
← Back to paper

The Proximal Surrogate Index: Long-Term Treatment Effects under Unobserved Confounding

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,805 characters · 17 sections · 53 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Proximal Surrogate Index: Long-Term Treatment Effects under Unobserved Confounding

abstractWe study the identification and estimation of long-term treatment effects under unobserved confounding by combining an experimental sample, where the long-term outcome is missing, with an observational sample, where the treatment assignment is unobserved. While standard surrogate index methods fail when unobserved confounders exist, we establish novel identification results by leveraging proxy variables for the unobserved confounders. We further develop multiply robust estimation and inference procedures based on these results. Applying our method to the Job Corps program, we demonstrate its ability to recover experimental benchmarks even when unobserved confounders bias standard surrogate index estimates.

Introduction

\paragraph{Motivation}

Empirical researchers are often interested in evaluating the long-term effects of interventions or treatments. For example, labor economists are interested in evaluating job training programs on participants' long-term earnings and employment outcomes, public economists are interested in assessing in-kind transfer programs on recipients' long-term well-being, and health economists study medical treatments on patients' survival and quality of life.

\paragraph{Challenges}

Although randomized controlled trials (RCTs) are the gold standard for causal inference, they often face practical challenges in measuring long-term outcomes. These challenges include limited funding, high attrition rates, and lengthy follow-up periods. As a result, researchers may only have access to short-term outcomes from experimental studies, while long-term outcomes of interest are only available in observational datasets, such as administrative records or surveys, where the treatment assignment is unobserved.

\paragraph{Existing Approaches}

The challenge of identifying long-term effects from combined experimental and observational data has been addressed in prior literature. The most related work is athey2025surrogate, who propose a surrogate index approach. Their method leverages short-term surrogate outcomes to predict long-term outcomes across datasets, recovering the long-term treatment effect under certain identification assumptions. However, these assumptions require the absence of unobserved confounders that affect both the treatment assignment and the long-term outcome, or the surrogate outcomes and the long-term outcome. These assumptions may be violated in many practical settings. For example, in job training program evaluations, unobserved ability may affect both participation and long-term earnings; in healthcare studies, unobserved health consciousness may confound the relationship between short-term behaviors and long-term health. These concerns motivate our work.

\paragraph{Research Question}

In this paper, we study the nonparametric identification and estimation of long-term treatment effects under unobserved confounding, by combining experimental and observational data, allowing for unobserved confounding through the use of proxy variables.

\paragraph{Our Contributions}

This paper contributes to the literature in three ways. First, we study long-term treatment effects in a two-sample surrogate setting where the experimental sample lacks the long-term outcome and the observational sample lacks treatment status, allowing for unobserved confounding. We establish nonparametric identification by leveraging proxy variables and deriving two complementary bridge-based strategies: an outcome bridge that imputes the missing long-term outcome in the experimental sample, and a surrogate bridge that reweights the observational sample to the experimental target. Second, we combine these strategies to obtain a multiply robust identification formula that remains valid under four alternative sets of correctly specified nuisance components, and we characterize the corresponding efficient influence function (EIF) and semiparametric local efficiency bound. Third, we develop cross-fitted estimators based on double/debiased machine learning (DML) that accommodate flexible nuisance estimation, and we provide consistency, asymptotic normality, and efficiency results. In an application to Job Corps (JC), our approach closely matches experimental benchmarks in designs where standard surrogate index methods are biased.

\paragraph{Related Literature}

Our paper is situated at the intersection of several strands of literature. First, it is related to the literature on causal inference with surrogate outcomes weir2006statistical, vanderweele2013surrogate, joffe2009related. These works explore the use of intermediate outcomes as surrogates for long-term outcomes of interest. Many criteria have been proposed to ensure the validity of surrogates prentice1989surrogate,frangakis2002principal,lauritzen2004discussion. However, these criteria may produce “surrogate paradoxes” in the presence of unobserved confounding chen2007criteria. Our work addresses this limitation, offering new strategies to use surrogates.

Second, this paper adds to the growing body of work on combining different data sources for causal inference colnet2024causal, especially those that leverage short-term outcomes to estimate long-term treatment effects. Besides athey2025surrogate, athey2025experimental, chenSemiparametricEstimationLongterm2023, ghassamiCombiningExperimentalObservational2022b, and imbensLongtermCausalInference2025a, propose various methods to combine experimental and observational samples to estimate long-term effects, assuming both samples contain information on the treatment assignment, which differs from our setting where the treatment is unobserved in the observational sample.

Finally, our work contributes to the recent literature on proximal causal inference miaoIdentifyingCausalEffects2018, tchetgentchetgenIntroductionProximalCausal2024, cuiSemiparametricProximalCausal2024, which has been applied to various settings, including longitudinal studies imbens2021controlling,deaner2023proxycontrolspaneldata,liu2024proximal,ying2023proximal,tchetgentchetgenIntroductionProximalCausal2024, mediation analysis dukes2023proximal,ghassami2025causal, and optimal treatment regimes bai2025proximal,gao2025multiple,shen2023optimal,qi2024proximal. But most existing works focus on single-sample settings, except for ghassamiCombiningExperimentalObservational2022b,imbensLongtermCausalInference2025a, which also consider data combination settings but under different data structures and assumptions. Our paper provides an extension by adapting the proximal learning framework to a more challenging data combination structure, where key variables are missing across samples.

\paragraph{Organization of the Paper}

The rest of this paper is organized as follows. In (ref), we introduce the setup and notation. (ref) reviews the surrogate index approach proposed by athey2025surrogate. In (ref), three identification results are presented. The estimation and inference procedures are discussed in (ref). (ref) presents the empirical application. Finally, (ref) concludes.

Setup and Notation

We consider a setting where we have access to an experimental dataset ($\text{E}$) and an observational dataset ($\text{O}$), with sample sizes $N_{\text{E}}$ and $N_{\text{O}}$ respectively. It is convenient to view the pooled dataset as consisting of $N = N_{\text{E}} + N_{\text{O}}$ units, drawn from a superpopulation that mixes both experimental and observational distributions, where each unit belongs to either the experimental sample ($G_i = \text{E}$) or the observational sample ($G_i = \text{O}$). We denote the marginal probability of being in the experimental sample as $\pi_0 \equiv P(G_i = \text{E})$, with $0 < \pi_0 < 1$.

For each unit $i$, there is a binary treatment of interest, $A_{i}\in\{0,1\}$, a scalar primary outcome, $Y_{i}$, a set of intermediate outcomes, $S_{i}$ (which we refer to as surrogates), and a set of pre-treatment covariates, $X_{i}$. The data structure is characterized by differential availability across samples. In the experimental sample ($\text{E}$), the treatment $A_{i}$ is randomly assigned, and we observe $(A_i, S_i, X_i)$. Our goal is to recover the causal effect of $A_{i}$ on $Y_{i}$. However, the primary long-term outcome $Y_{i}$ is not recorded in this sample. To provide information on the long-term outcome, the observational dataset ($\text{O}$) is collected, which contains measurements of $(Y_i, S_i, X_i)$. Nevertheless, the treatment status $A_{i}$ is unobserved in the observational sample.

A key challenge in this surrogate setting, which we will address in (ref), arises when the identification assumptions required for the existing surrogate index approach are violated due to unobserved confounders $U_i$ athey2025surrogate. To overcome this, we draw insights from recent proximal inference literature miaoIdentifyingCausalEffects2018, tchetgentchetgenIntroductionProximalCausal2024, cuiSemiparametricProximalCausal2024. We therefore assume that we have access to two types of proxy variables: a set of surrogate-aligned proxies $Z_{i}$ and a set of outcome-aligned proxies $W_{i}$. While these proxies are not utilized in the standard surrogate index approach reviewed in (ref), they are essential to our proposed identification strategy in (ref). The list of variables and their availability across samples is summarized in (ref).

table[table omitted — 893 chars of source]

To formally define our causal estimand and identification assumptions, we adopt the potential outcomes framework rubin1974estimating, holland1986statistics, imbens2015causal. For any variable $V_{i}$ that can be affected by the treatment $A_{i}$, we denote $V_{i}(a)$ as its potential value if $A_i$ were set to $a\in\{0,1\}$. Similarly, we denote $V_{i}(s)$ as the potential value if the surrogate were set to $s\in\mathcal{S}$, and $V_{i}(a,s)$ if both were set. We also use the nested potential outcomes notation $V_{i}(a,S_{i}(a'))$ introduced by pearl2001direct and robins1992identifiability to denote the potential value of $V_i$ when the treatment is set to $a$ and the surrogate is set to its potential value $S_i(a')$.

Throughout the paper, we assume that all random variables are independent and identically distributed (IID) samples from a superpopulation, although some variables are only observed in one of the two samples. We use $P(\cdot)$ and $\E[\cdot]$ to denote the probability and expectation with respect to this superpopulation, and use $f(\cdot)$ to denote the corresponding probability density/mass function as appropriate. We drop the subscript $i$ when there is no ambiguity.

assumption[Random Sampling] We observe an IID sample \begin{align*} \mathcal{D} = \bc{D_i}_{i = 1}^N = \bc{\mathbf{1}_{\{G_i = O\}} Y_i, W_i, \mathbf{1}_{\{G_i = O\}} Z_i, S_i, \mathbf{1}_{\{G_i = E\}} A_i, X_i, G_i}_{i = 1}^N \end{align*} from the population described in (ref), where $N = N_{\text{E}} + N_{\text{O}}$ is the total sample size, and $\mathbf{1}_{\{G_i = g\}}$ indicates whether unit $i$ belongs to sample $g \in \{\text{E}, \text{O}\}$. We denote $\mathcal{D}_{\text{E}}$ and $\mathcal{D}_{\text{O}}$ as the subsets of $\mathcal{D}$ corresponding to the experimental and observational samples, respectively.

We further maintain the stable unit treatment value assumption (SUTVA) rubin1980discussion: for any variable $V$ affected by $A$ and/or $S$,

align*[align* omitted — 55 chars of source]

which links the observed variables to the potential outcomes.

Our primary estimand of interest is the average treatment effect (ATE) on the primary outcome $Y$ in the experimental sample:

align[align omitted — 79 chars of source]

where $a = 1$ and $a' = 0$.

The Surrogate Index and Its Limitations

In this section, we briefly review the surrogate index approach proposed by athey2025surrogate and analyze its limitations in the presence of unobserved confounding. This discussion motivates our proximal approach.

Identification via the Surrogate Index

athey2025surrogate rely on three key assumptions to identify the long-term ATE. First, the standard unconfoundedness assumption in the experimental sample ($A \ind (Y(a), S(a)) \mid X, G = \text{E}$) is assumed to hold. Second, the surrogacy assumption ($A \ind Y \mid S, X, G = \text{E}$), originally proposed by prentice1989surrogate, posits that the treatment $A$ provides no additional information about the primary outcome $Y$ beyond what is already contained in the surrogates $S$ and pre-treatment variables $X$ in the experimental sample. Third, the comparability assumption requires that the conditional outcome distribution is stable across samples, i.e., $G \perp Y \mid S, X$.

Under these assumptions, the mean potential outcome is identified by

align[align omitted — 179 chars of source]

where $\mu(S, X, g) \equiv \E[Y \mid S, X, G = g]$ is called the surrogate index.

Limitations and Motivation

While elegant, the surrogate index approach relies heavily on the validity of the surrogacy and comparability assumptions. (ref) illustrates three common scenarios where these assumptions can be violated.

The first scenario ((ref)) involves an unobserved confounder $U$ that affects both the treatment $A$ and the primary outcome $Y$. Structurally, this resembles the front-door model pearl1995causal. However, in our context, conditioning on $S$ opens the collider path $G \to A \leftarrow U \to Y$, violating the comparability assumption.

The second scenario ((ref)) involves an unobserved confounder $U$ that affects both the surrogates $S$ and the primary outcome $Y$, resembling an instrumental variable (IV) model wrightTariffAnimalVegetable1928,imbensIdentificationEstimationLocal1994,imbensInstrumentalVariablesEconometricians2014. Here, since $S$ is a collider, conditioning on $S$ opens the collider path $A \to S \leftarrow U \to Y$, violating the surrogacy and potentially leading to surrogate paradoxes chen2007criteria,frangakis2002principal. In addition, the sample indicator $G$ becomes associated with $Y$ after conditioning on $S$, violating the comparability assumption.

The third scenario ((ref)) involves a direct effect of $A$ on $Y$ that bypasses $S$, as in mediation analysis pearl2001direct, robins1992identifiability, vanderweele2015explanation. Since $A$ and $Y$ are never observed together, this effect is fundamentally unidentified. Like the standard surrogate index, our proximal approach must assume the absence of direct effects ((ref)).

We summarize the key features of these three scenarios in comparison to the surrogate index setting in (ref).

figure[figure omitted — 2,741 chars of source]

These limitations point to a clear strategy inspired by the front-door criterion. Recall that the front-door formula identifies the causal effect $A \to Y$ by composing the effect of $A$ the mediator $S$, and the effect of $S$ on $Y$. In our setting, the first component ($A \to S$) is readily identified from the experimental sample due to randomization. The challenge lies entirely in identifying the second component ($S \to Y$) from the observational sample, where the relationship is confounded by $U$. While the standard front-door approach assumes no unobserved $S$-$Y$ confounding, we relax this by leveraging proxy variable ($W, Z$) to adjust for $U$. Thus, our approach conceptually mirrors the front-door strategy but adapts it to the data combination setting: we combine the experimentally identified effect of $A$ on $S$ with the proximally identified effect of $S$ on $Y$ to recover the long-term ATE.

table[table omitted — 1,132 chars of source]

Restoring Identification via Proxies

As discussed in (ref), identifying the long-term ATE requires addressing the potential violations of surrogacy and comparability assumptions. In this section, we present a proximal identification strategy that allows for unobserved confounders $U$. To provide some intuition, we adopt the single-world intervention graph (SWIG) to visualize the causal structure under a hypothetical intervention richardson2013single, heckman2015causal.

We first maintain the following no direct effect assumption. Similar assumptions have been widely used in the IV literature, where it is often referred to as the exclusion restriction imbensIdentificationEstimationLocal1994,imbensInstrumentalVariablesEconometricians2014. This assumption requires that $A$ has no direct causal link to $Y$ other than through $S$, corresponding to the absence of an edge $a \to Y(a,s)$ in (ref).

assumption[No Direct Effect] \begin{align} Y(a, s) = Y(a', s) \quad for all a, a' \in \{0, 1\}, s \in \mathcal{S}. \end{align}

Next, we characterize the confounding structure. While unobserved confounders $U$ may bias the relationships in the observational sample, we assume that $U$, together with the observed covariates $X$, are the only sources of confounding. In terms of (ref), this means that there are no other unmeasured common causes between $A$ and $Y(a, s)$, and between $S$ and $Y(s)$, besides $U$.

assumption[Unconfoundedness in the Observational Sample] For all $a \in \{0, 1\}$ and $s \in \mathcal{S}$, \begin{align} A \ind \bp{\bp{Y(a)}_{a \in \{0, 1\}}, \bp{S(a)}_{a \in \{0, 1\}}} &\mid (X, U, G = O), \\ S \ind \bp{Y(s)}_{s \in \mathcal{S}} &\mid (A, X, U, G = O), \end{align} where $0 < P(A = 1 \mid X, U, G = \text{O}) < 1$ almost surely.

In addition, we assume that treatment assignment is randomized in the experimental sample, as is standard in RCTs. However, the surrogate $S$ is an intermediate outcome, not a randomized intervention. Even within the experiment, individuals with the same treatment status may have different surrogate values due to unobserved factors $U$ (e.g., motivation or ability), which also affect the long-term outcome $Y$. Similarly to (ref), we assume that $U$ is the only source of unobserved confounding in the experimental sample as well. The following assumption formalizes these conditions.

assumption[Unconfoundedness in the Experimental Sample] For all $a \in \{0, 1\}$ and $s \in \mathcal{S}$, \begin{align} A \ind \bp{\bp{Y(a)}_{a \in \{0, 1\}}, \bp{S(a)}_{a \in \{0, 1\}}, U, W} &\mid (X, G = E), \\ S \ind \bp{Y(s)}_{s \in \mathcal{S}} &\mid (A, X, U, G = E), \end{align} where $0 < e_0(X) < 1$ almost surely, and $e_0(X) \equiv P(A = 1 \mid X, G = \text{E})$.
remarkThe SWIG in (ref), is a pooled representation. The randomization of $A$ in the experimental sample is captured by an directed edge from $G$ to $A$, indicating that the treatment assignment mechanism depends on the sample indicator $G$. Although the pooled graph depicts a generic dependence $U \to A$, the experimental design explicitly breaks this dependence within the $G=\text{E}$ stratum.

To address the unobserved confounding, we need to leverage proxy variables that provide indirect information about the unobserved confounders $U$. We make the following assumption regarding the availability of such proxies.

assumption[Proxies Availability] We observe $W$ in both samples and $Z$ only in the observational sample. Moreover, \begin{align} Z \ind Y &\mid (S, X, U, G), \\ W \ind (A, S, Z) &\mid (X, U, G). \end{align}

(ref) illustrates the roles of these proxies. In this representation, $Z$ acts as a surrogate-aligned proxy. While the figure depicts $Z$ as a post-treatment variable (e.g., a post-program survey), the formal requirement is broader: $Z$ must be conditionally independent of $Y$ given $S$, $X$, and $U$, regardless of its temporal or causal positioning. Conversely, $W$ acts as an outcome-aligned proxy. Graphically, it behaves like a pre-treatment covariate that is a noisy measure of $U$. The key requirement is that $W$ is sufficiently informative about $U$ but is structurally independent of the $(A, S, Z)$ given covariates.

The proxies satisfying (ref) are ubiquitous in practical applications. For example, in job training programs, test scores collected before the training can serve as outcome-aligned proxies $W$, because they reflect the participants' underlying abilities $U$, but do not affect or are affected by the training $A$ or the short-term outcomes $S$ directly. Similarly, post-training surveys about job search self-efficacy or soft skills can serve as surrogate-aligned proxies $Z$. These measures capture the participants' latent motivation $U$ and the immediate psychological impact of the training $A$, but arguably do not directly determine long-term earnings $Y$ once the actual intermediate employment history $S$ is accounted for.

To generalize the relationships learned in the observational sample to the experimental sample, we require the mechanism underlying the data generation to be comparable across groups. In the proximal learning framework, this transportability must hold for both the primary outcome and the outcome-aligned proxy.

assumption[Transportability] \begin{align} Y &\ind G \mid (S, X, U), \\ W &\ind G \mid (X, U), \end{align} and \begin{align} \frac{P(G = E \mid U, X)}{P(G = O \mid U, X)} < \infty \quad almost surely. \end{align}

(ref) is analogous to the comparability assumption in athey2025surrogate, but allows for unobserved confounders $U$. It posits that given the surrogate and the complete set of confounders (both observed and unobserved), the distribution of the primary outcome $Y$ is stable across samples. (ref) is a new requirement specific to the proximal approach. It ensures that the outcome-aligned proxy $W$ relates to the latent confounder $U$ in an invariant manner across the experimental and observational samples. In our SWIG representation ((ref)), these assumptions correspond to the requirement that the group indicator $G$ has no directed edges on $Y$ or $W$. These transportability assumptions are closely related to the concept of S-admissibility in the data fusion literature bareinboim2016causal, pearl2014external.

Notably, our identification results do not require transportability for the surrogate-aligned proxy $Z$. While $Z$ is crucial for identifying the bridge functions within the observational sample, its distributional relationship with $U$ may vary across samples. Thus, we allow for arbitrary dependence between $G$ and $Z$, accommodating scenarios where the proxy measurement mechanism for $Z$ differs across samples. In the SWIG ((ref)), this is reflected by the presence of a directed edge from $G$ to $Z$.

figure[figure omitted — 1,624 chars of source]

Identification via Outcome Bridge Function

We first introduce the outcome bridge function, which helps to impute the missing long-term outcome $Y$ in the experimental sample.

assumption[Outcome Bridge Function] There exists a square-integrable function $h_0: \mathcal{W} \times \mathcal{S} \times \mathcal{X} \to \mathbb{R}$ for the observational sample $(G = \text{O})$ such that \begin{align} \E\bs{Y \mid Z, S, X, G = O} = \E\bs{h_0 (W, S, X) \mid Z, S, X, G = O}. \end{align}

We further assume that the surrogate-aligned proxy $Z$ is sufficiently informative about the unobserved confounders $U$ to ensure that we can learn the unobserved confounding relationship using $Z$.

assumption[Completeness for Outcome Bridge Function] For any $g \in L_2$, if $\E\bs{g(U) \mid Z = z, S = s, X = x} = 0$ for all $z, s, x$, then $g(U) = 0$ almost surely.

This completeness condition has been widely used in the nonparametric IV models neweyInstrumentalVariableEstimation2003,ai2003efficient,chernozhukov2005iv, measurement error models hu2008instrumental,an2012well, and the recent proximal inference literature miaoIdentifyingCausalEffects2018,tchetgentchetgenIntroductionProximalCausal2024,cuiSemiparametricProximalCausal2024,miaoConfoundingBridgeApproach2024,ghassamiCombiningExperimentalObservational2022b,imbensLongtermCausalInference2025a.

We establish the first identification result using the outcome bridge function.

theorem[Identification via Outcome Bridge Function] \begin{enumerate} • • Under (ref), the function $h_0$ in (ref) satisfies \begin{align*} \E\bs{Y \mid U, S, X, G = O} = \E\bs{h_0(W, S, X) \mid U, X, G = O}. \end{align*} • Suppose that the function $h_0$ satisfies (ref). Under (ref), the function $h_0$ also satisfies \begin{align*} \E\bs{Y \mid U, S, X, G = E} &= \E\bs{h_0(W, S, X) \mid U, X, G = E}. \end{align*} That is, the same function $h_0$ transports from the observational sample to the experimental sample. • Under (ref), the mean potential outcome is identified by \begin{align} \E\bs{Y(a) \mid X, G = E} = \E\bs{h_0(W, S, X) \mid A = a, X, G = E}, \end{align} where $h_0$ satisfies (ref), and we call $h_0$ the \emph{proximal surrogate index}. Thus, the ATE in the experimental sample is identified by \begin{align*} \tau_0 &= \E\bs{\bar{h}_0(1, X) - \bar{h}_0(0, X) \mid G = \text{E}} \\ &= \E\bs{\frac{A h_0(W, S, X)}{e_0(X)} - \frac{(1 - A) h_0(W, S, X)}{1 - e_0(X)} \;\middle|\; G = \text{E}}, \end{align*} where for any $h: \mathcal{W} \times \mathcal{S} \times \mathcal{X} \to \mathbb{R}$, we define \begin{align*} \bar{h}^{(h)}(A, X) \equiv \E\bs{h(W, S, X) \mid A, X, G = \text{E}}. \end{align*} In particular, $\bar{h}_0 \equiv \bar{h}^{(h_0)}$. \end{enumerate}
proofSee (ref).

(ref) shows that the missing long-term outcome $Y$ can be imputed in the experimental sample using the outcome bridge function $h$ learned from the observational sample. The long-term ATE in the experimental sample can then be identified using standard methods, such as outcome regression (OR) or inverse probability weighting (IPW), based on the imputed outcomes.

remark[Connection to athey2025surrogate] (ref) generalizes athey2025surrogate's surrogate index identification by allowing for unobserved confounding between treatment and outcome in the observational sample. See (ref) for more details.

Identification via Surrogate Bridge Function

An alternative identification strategy leverages the surrogate bridge function, as introduced below.

assumption[Surrogate Bridge Function] For each treatment level $a \in \{0, 1\}$, there exists a square-integrable function $q_{a, 0}: \mathcal{Z} \times \mathcal{S} \times \mathcal{X} \to \mathbb{R}$ such that \begin{align} \E\bs{q_{a, 0}(Z, S, X) \mid W, S, X, G = O} = \frac{f(W, S \mid A = a, X, G = E) f(X \mid G = E)}{f(W, S, X \mid G = O)}. \end{align}

To ensure that the surrogate bridge function ((ref)) can be used for reweighting, we impose the following completeness condition.

assumption[Completeness for Surrogate Bridge Function] For any $g \in L_2$, if $\E\bs{g(U) \mid W = w, S = s, X = x} = 0$ for all $w, s, x$, then $g(U) = 0$ almost surely.

This completeness condition mirrors that in (ref), guaranteeing the outcome-aligned proxy $W$ is sufficiently informative about the unobserved confounders $U$.

We now present an alternative identification result based on the surrogate bridge function.

theorem[Identification via Surrogate Bridge Function] \begin{enumerate} • • Under (ref), the function $q_a$ in (ref) satisfies \begin{align*} \E\bs{q_{a, 0}(Z, S, X) \mid U, S, X, G = O} = \frac{f(U, S \mid A = a, X, G = E) f(X \mid G = E)}{f(U, S, X \mid G = O)}. \end{align*} • Under (ref), the mean potential outcome is identified by \begin{align} \E\bs{Y(a) \mid G = E} = \E\bs{q_{a, 0}(Z, S, X) Y \mid G = O}, \end{align} where $q_{a, 0}$ satisfies (ref). Thus, the ATE in the experimental sample is identified by \begin{align*} \tau_0 = \E\bs{q_{1, 0}(Z, S, X) Y \mid G = \text{O}} - \E\bs{q_{0, 0}(Z, S, X) Y \mid G = \text{O}}. \end{align*} \end{enumerate}
proofSee (ref).

(ref) shows that the long-term ATE in the experimental sample can be identified by reweighting the observed outcomes $Y$ in the observational sample using the surrogate bridge functions $q_{a, 0}$'s.

remark(ref) is new to the literature, providing a novel form of bridge function ((ref)) for identifying the ATE in the data combination setting.
remark[Connection to athey2025surrogate] (ref) generalizes athey2025surrogate's surrogate score identification by allowing for unobserved confounding. See (ref) for more details.

Multiply Robust Identification

In (ref), we present two distinct identification strategies based on the outcome bridge functions and the surrogate bridge functions, respectively. Combining these two strategies, we obtain a multiply robust identification result as follows.

theorem[Multiply Robust Identification] Let $\tau_0$ be the true ATE defined in (ref). Let $\eta = (e, h, \bar{h}, q_0, q_1)$ be a set of possibly misspecified nuisance functions. Let $\eta_0 = (e_0, h_0, \bar{h}_0, q_{0,0}, q_{1,0})$ denote the true nuisance functions, where $e_0$ is the true propensity score, $h_0$ is the function satisfying (ref), $\bar{h}_0(A, X)$ is defined as $\E\bs{h_0(W, S, X) \mid A, X, G = \text{E}}$, and $q_{a, 0}$'s are the functions satisfying (ref). If any of the following conditions on nuisance functions holds: \begin{enumerate} • $h = h_0$, $\bar{h} = \bar{h}_0$; • $h = h_0$, $e = e_0$; • $q_a = q_{a, 0}$ for each $a \in \{0, 1\}$, $e = e_0$; • $q_a = q_{a, 0}$ for each $a \in \{0, 1\}$, $\bar{h} = \bar{h}^{(h)}$; \end{enumerate} then \begin{align} \tau_0 = \E\bs{\varphi(D; \eta)}, \end{align} where $\varphi(D; \eta) \equiv \varphi_{\text{E}}(D; \eta) + \varphi_{\text{O}}(D; \eta)$ is defined as \begin{align*} \varphi_{E}(D; \eta) &\equiv \frac{\mathbf{1}_{\bc{G = E}}}{\pi_0} \bp{\frac{(A - e(X))\bp{h(W, S, X) - \bar{h}(A, X)}}{e(X)(1 - e(X))} + \bar{h}(1, X) - \bar{h}(0, X)} \\ \varphi_{O}(D; \eta) &\equiv \frac{\mathbf{1}_{\bc{G = O}}}{1 - \pi_0} \bp{\bp{q_1(Z, S, X) - q_0(Z, S, X)} (Y - h(W, S, X))}. \end{align*}
proofSee (ref).
remark(ref) is multiply robust in the sense that it identifies the ATE $\tau_0$ if any one of the four sets of nuisance functions is correctly specified. This identification result is motivated by the EIF of $\tau_0$ under a certain nonparametric model (see (ref)), extending the influence function-based identification in athey2025surrogate to accommodate unobserved confounding. See (ref) for more details.

Estimation and Inference

Cross-Fitted Estimators

In this section, we propose a generic estimation procedure for the ATE $\tau_0$. This approach builds upon the identification results in (ref) and employs the cross-fitting technique from chernozhukov2018double to reduce overfitting bias when estimating nuisance parameters.

definition[Cross-Fitted Estimators] Fix an integer $K \geq 2$. Let $\mathcal{I} \equiv \{1, \dots, N\}$ be the index set, where $N$ is the total sample size. Define $\mathcal{I}_{\text{E}} \equiv \{i \in \mathcal{I}: G_i = \text{E}\}$ and $\mathcal{I}_{\text{O}} \equiv \{i \in \mathcal{I}: G_i = \text{O}\}$ as the index sets for the experimental and observational samples, respectively. Randomly partition $\mathcal{I}_{\text{E}}$ into $K$ disjoint subsets $\{\mathcal{I}_{\text{E}, k}\}_{k = 1}^K$. Similarly, randomly partition $\mathcal{I}_{\text{O}}$ into $K$ disjoint subsets $\{\mathcal{I}_{\text{O}, k}\}_{k = 1}^K$. Define $\mathcal{D}_{k} \equiv \{D_i: i \in \mathcal{I}_{\text{E}, k} \cup \mathcal{I}_{\text{O}, k}\}$ as the evaluation set for fold $k$, and $\mathcal{D}_{-k} \equiv \mathcal{D} \setminus \mathcal{D}_{k}$ as the training set for fold $k$. For each fold $k$, estimate the nuisance functions $\hat{\eta}_{k} \equiv (\hat{e}_{k}, \hat{h}_{k}, \hat{\bar{h}}_{k}, \hat{q}_{0, k}, \hat{q}_{1, k})$ using the data excluding fold $k$, i.e., $\mathcal{D}_{-k}$, where $\hat{\bar h}_k$ is fitted by regressing $\hat h_k(W,S,X)$ on $(A,X)$ using $\mathcal D_{\mathrm{E},-k}$. The cross-fitted estimators are defined as \begin{align*} \hat{\tau}_{OB-OR} &\equiv \frac{1}{N_{E}} \sum_{k = 1}^K \sum_{i \in \mathcal{I}_{E, k}} \bp{\hat{\bar{h}}_{k}(1, X_i) - \hat{\bar{h}}_{k}(0, X_i)}, \\ \hat{\tau}_{OB-IPW} &\equiv \frac{1}{N_{E}} \sum_{k = 1}^K \sum_{i \in \mathcal{I}_{E, k}} \bp{\frac{A_i \hat{h}_{k}(W_i, S_i, X_i)}{\hat{e}_{k}(X_i)} - \frac{(1 - A_i) \hat{h}_{k}(W_i, S_i, X_i)}{1 - \hat{e}_{k}(X_i)}}, \\ \hat{\tau}_{\text{SB}} &\equiv \frac{1}{N_{\text{O}}} \sum_{k = 1}^K \sum_{i \in \mathcal{I}_{\text{O}, k}} \bp{\hat{q}_{1, k}(Z_i, S_i, X_i) - \hat{q}_{0, k}(Z_i, S_i, X_i)} Y_i, \\ \hat{\tau}_{\text{MR}} &\equiv \frac{1}{N_{\text{E}}} \sum_{k = 1}^K \sum_{i \in \mathcal{I}_{\text{E}, k}} \bp{\frac{(A_i - \hat{e}_{k}(X_i))(\hat{h}_{k}(W_i, S_i, X_i) - \hat{\bar{h}}_{k}(A_i, X_i))}{\hat{e}_{k}(X_i)(1 - \hat{e}_{k}(X_i))}} \\ &\quad + \frac{1}{N_{\text{E}}} \sum_{k = 1}^K \sum_{i \in \mathcal{I}_{\text{E}, k}} \bp{\hat{\bar{h}}_{k}(1, X_i) - \hat{\bar{h}}_{k}(0, X_i)} \\ &\quad + \frac{1}{N_{\text{O}}} \sum_{k = 1}^K \sum_{i \in \mathcal{I}_{\text{O}, k}} \bp{(\hat{q}_{1, k}(Z_i, S_i, X_i) - \hat{q}_{0, k}(Z_i, S_i, X_i)) (Y_i - \hat{h}_{k}(W_i, S_i, X_i))}. \end{align*}

Asymptotic Properties

Before stating the main theorems, we introduce some additional definitions and assumptions. First, we define the linear operator $T: L_2(W, S, X) \to L_2(Z, S, X)$ as \[ (T h)(Z, S, X) = \E\bs{h(W, S, X) \mid Z, S, X, G = \text{O}}, \] and its adjoint operator $T^*: L_2(Z, S, X) \to L_2(W, S, X)$ as \[ (T^* q_a)(W, S, X) = \E\bs{q_a(Z, S, X) \mid W, S, X, G = \text{O}}. \]

Consistency

Following the standard approach in imbensLongtermCausalInference2025a, we impose the following convergence rate conditions on the nuisance estimators to establish consistency.

assumption[Convergence Rates] Let $\| \cdot \|_{2}$ denote the $L_2$ norm. For any square-integrable function $g(W, S, X)$, let $\tilde{\bar{h}}^{(g)}(A, X)$ denote the population limit of the estimator $\hat{\bar{h}}$ trained with pseudo-outcome $g(W, S, X)$ and predictors $(A, X)$ in the experimental sample. For each fold $k$, the nuisance estimators satisfy the following convergence rates: \begin{align*} \|\hat{h}_{k} - \tilde{h}\|_{2} &= O_P(\delta_{h, N}), \quad \|T(\hat{h}_{k} - \tilde{h})\|_{2} = O_P(\rho_{h, N}), \quad \|\hat{\bar h}_{k} - \tilde{\bar{h}}^{(\hat{h}_k)}\|_2 = O_P(\delta_{\bar h,N}), \\ \|\hat{q}_{a, k} - \tilde{q}_{a}\|_{2} &= O_P(\delta_{q, N}), \quad \|T^*(\hat{q}_{a, k} - \tilde{q}_{a})\|_{2} = O_P(\rho_{q, N}), \quad \|\hat e_{k} - \tilde{e}\|_{2} = O_P(\delta_{e, N}), \end{align*} where $\delta_{h, N}$, $\rho_{h, N}$, $\delta_{\bar{h}, N}$, $\delta_{q, N}$, $\rho_{q, N}$, $\delta_{e, N}$ are sequences converging to zero as $N \to \infty$, and the mapping $g \mapsto \tilde{\bar{h}}^{(g)}$ is Lipschitz continuous with respect to the $L_2$ norm.
theorem[Consistency] \begin{enumerate} • • If the conditions in (ref), (ref), $\tilde{h} = h_0$ and $\tilde{\bar{h}}^{(h_0)} = \bar{h}_0$, then $\hat{\tau}_{\text{OB-OR}}$ is consistent for $\tau_0$. • If the conditions in (ref), (ref), $\tilde{h} = h_0$ and $\tilde{e} = e_0$ hold, then $\hat{\tau}_{\text{OB-IPW}}$ is consistent for $\tau_0$. • If the conditions in (ref), (ref), $\tilde{q}_{a} = q_{a, 0}$ for each $a \in \{0, 1\}$ and $\tilde{e} = e_0$ hold, then $\hat{\tau}_{\text{SB}}$ is consistent for $\tau_0$. • If (ref) holds, and any of the following sets of conditions is satisfied: \begin{enumerate} • $\tilde{h} = h_0$ and $\tilde{\bar{h}}^{(h_0)} = \bar{h}_0$; • $\tilde{h} = h_0$ and $\tilde{e} = e_0$; • $\tilde{q}_{a} = q_{a, 0}$ for each $a \in \{0, 1\}$ and $\tilde{e} = e_0$; • $\tilde{q}_{a} = q_{a, 0}$ for each $a \in \{0, 1\}$ and $\tilde{\bar{h}}^{(\tilde{h})}(A, X) = \E[\tilde{h} \mid A, X, G=\text{E}]$, \end{enumerate} then $\hat{\tau}_{\text{MR}}$ is consistent for $\tau_0$. \end{enumerate}
proofSee (ref).

Asymptotic Normality

While consistency requires only that the nuisance estimators converge to their true values, establishing asymptotic normality requires stricter conditions on their convergence rates. Specifically, we need to ensure that the second-order bias terms vanish faster than $N^{-1/2}$, ensuring that the asymptotic distribution of $\hat{\tau}_{\text{MR}}$ is dominated by the first-order term. We impose the following assumption on the product rates of the nuisance estimators.

assumption[Product Rates of Convergence] Suppose that (ref) holds. The convergence rates satisfy the following conditions: \begin{align} \delta_{e, N} \delta_{\bar{h}, N} = o_P(N^{-1/2}), \\ \min \left\{ \delta_{q, N} \rho_{h, N}, \; \rho_{q, N} \delta_{h, N} \right\} = o_P(N^{-1/2}). \end{align}
remark(ref) is standard in the DML literature chernozhukov2018double, requiring the product of the convergence rates of the propensity score and the pseudo-outcome regression that uses the estimated outcome bridge function as the outcome to be $o_P(N^{-1/2})$. (ref) is specific to the proximal inference setting. Following imbensLongtermCausalInference2025a, we distinguish between two types of convergence rates: strong-metric errors ($\delta_{h,N}$ and $\delta_{q,N}$) measure $L_2$ deviation from the true bridge functions, and weak-metric errors ($\rho_{h,N}$ and $\rho_{q,N}$) measure violations of the bridge equations via operators $T$ and $T^*$. Because bridge-function estimation amounts to solving potentially ill-posed conditional moment equations, an estimator can exhibit a slow strong-metric rate while still achieving a fast weak-metric rate (e.g., via Tikhonov regularization or sieve methods). (ref) exploits this asymmetry, requiring only that the product of one strong-metric error and one weak-metric error be $o_P(N^{-1/2})$, a strictly weaker condition than requiring both strong rates to satisfy $\delta_{h,N} \delta_{q,N} = o_P(N^{-1/2})$.

The following theorem establishes the asymptotic normality of the cross-fitted estimator $\hat{\tau}_{\text{MR}}$ when the bridge functions and propensity score converge to their true values and the pseudo-outcome regression for $\hat{\bar h}_k$ is correctly specified.

theorem[Asymptotic Normality] Suppose (ref) holds with the correct limits \[ \tilde{h} = h_0, \quad \tilde{e} = e_0, \quad \tilde{q}_a = q_{a,0} \text{ for } a \in \{0, 1\}, \quad \text{and} \quad \tilde{\bar{h}}^{(h_0)} = \bar{h}_0. \] If (ref) also holds, then the cross-fitted estimator $\hat{\tau}_{\text{MR}}$ is asymptotically normal: \begin{align*} \sqrt{N}(\hat{\tau}_{MR} - \tau_0) \xrightarrow{d} \mathcal{N}(0, V), \end{align*}
proofSee (ref).

An immediate consequence of (ref) is that we can construct a consistent estimator for the asymptotic variance $V$ using the plug-in approach:

align*[align* omitted — 529 chars of source]

The following theorem establishes the consistency of this variance estimator and the validity of the corresponding confidence interval (CI).

theorem[Consistency of Variance Estimator] Under the conditions of (ref), the variance estimator $\hat{V}$ is consistent for $V$: \begin{align*} \hat{V} \xrightarrow{p} V. \end{align*} Consequently, the asymptotic $(1-\alpha)$ CI for $\tau_0$ is given by \begin{align*} CI_{1 - \alpha} = \left( \hat{\tau}_{MR} - \Phi^{-1}(1 - \alpha/2) \sqrt{\hat{V}/N}, \quad \hat{\tau}_{MR} + \Phi^{-1}(1 - \alpha/2) \sqrt{\hat{V}/N} \right), \end{align*} satisfying $\lim_{N \to \infty} P(\tau_0 \in \text{CI}_{1 - \alpha}) = 1 - \alpha$, where $\Phi^{-1}(\cdot)$ is the quantile function of the standard normal distribution.
proofSee (ref).

Semiparametric Efficiency

We now establish the local semiparametric efficiency of our cross-fitted estimator $\hat{\tau}_{\text{MR}}$ under a statistical model $\mathcal{P}$ that imposes no restrictions on the observed data law other than the existence of an outcome bridge function $h_0$ ((ref))

Following cuiSemiparametricProximalCausal2024 and imbensLongtermCausalInference2025a, to ensure the EIF is well-defined, we impose the following surjectivity condition, which guarantees the uniqueness of the bridge functions.

assumption[Surjectivity] The linear operators $T$ and $T^*$ are surjective.
proposition[Uniqueness of Bridge Functions] Under (ref), the outcome bridge function $h_0$ in (ref) and the surrogate bridge functions $q_{a, 0}$'s in (ref) are unique almost surely.
proofSee (ref).

The following theorem characterizes the EIF for the ATE $\tau_0$ and the corresponding semiparametric efficiency bound.

theorem[EIF] Consider the semiparametric model $\mathcal{P}$ evaluated at the submodel where the (ref) hold. The EIF for the ATE $\tau_0$ is given by \begin{align*} &\mathrel{\phantom{=}} \IF(D) \\ &\equiv \frac{\mathbf{1}_{\bc{G = E}}}{\pi_0} \bp{\frac{(A - e_0(X))\bp{h_0(W, S, X) - \bar{h}_0(A, X)}}{e_0(X)(1 - e_0(X))} + \bar{h}_0(1, X) - \bar{h}_0(0, X) - \tau_0} \\ &\quad + \frac{\mathbf{1}_{\bc{G = O}}}{1 - \pi_0} \bp{\bp{q_{1, 0}(Z, S, X) - q_{0, 0}(Z, S, X)} (Y - h_0(W, S, X))}. \end{align*} Therefore, the corresponding semiparametric local efficiency bound for estimating $\tau_0$ is $V_\text{eff} = \E\bs{\IF(D)^2}$.
proofSee (ref).

A direct consequence of (ref) and (ref) is that our cross-fitted estimator $\hat{\tau}_{\text{MR}}$ achieves the semiparametric local efficiency bound as stated in the following theorem.

theorem[Semiparametric Efficiency] Suppose the conditions of (ref) and (ref) hold. Then, \begin{align*} \sqrt{N} (\hat{\tau}_{MR} - \tau_0) \xrightarrow{d} \mathcal{N}(0, V_eff), \end{align*} where $V_\text{eff}$ is the semiparametric efficiency bound derived in (ref).
proofSee (ref).
comment\section{Numerical Studies} \subsection{A Simulation Study}

Real Data Application

Following the seminal approach of lalonde1986evaluating, we exploit the experimental data to construct a “ground truth” benchmark against which we evaluate the performance of different methods. We analyze data from the JC program, a large-scale randomized evaluation of job training programs for disadvantaged youth in the United States schochet2001national,schochet2008does.

\paragraph{Design}

Our evaluation design proceeds in two steps. First, we randomly split the experimental dataset into two disjoint subsets: an “Experimental Sample” ($\text{E}$) where the long-term outcome ($Y$) is masked, and an “Observational Sample” ($\text{O}$) where the treatment assignment ($A$) is masked.\footnote{ The dataset for this application is publicly available at the causalweight R package causalweight. } This setup mimics the data combination problem while ensuring, by construction, that there is no selection bias between the two samples (i.e., comparability or transportability assumptions hold). This allows us to isolate the performance of the estimators in handling the failure of the surrogacy assumption or the presence of unobserved confounding.

\paragraph{Variables}

We define the treatment ($A$) as random assignment and the long-term outcomes ($Y$) as weekly earnings and proportion of weeks employed in the fourth year. The surrogate vector ($S$) includes earnings and proportion of weeks employed in years 2 and 3. To address potential confounding, we select possession of a general education development (GED) degree ($W$) and English mother tongue ($Z$) as proxies. These variables are determined pre-treatment and likely provide information about unobserved confounders, such as cognitive skills or long-term career orientation. We also include baseline covariates ($X$), including age, gender, race, education level, and pre-program earnings.

\paragraph{Methods Compared}

We compare four methods for estimating the ATE. The first method is the RCT benchmark, which estimates the ATE directly from the experimental sample using a linear regression of $Y$ on $(1, A, X)$. The second method is the standard surrogate index estimator athey2025surrogate, which first fits a regression of $Y$ on $(1, S, X)$ in the observational sample, then predicts $Y$ in the experimental sample, and finally estimates the ATE using a linear regression of the predicted $Y$ on $(1, A, X)$. The third method is similar to the second method but includes the proxy variables ($W$ and $Z$) as additional covariates in both regressions. The fourth method is our proposed method, which first estimates an linear outcome bridge function $h(W, S, X; \beta)$ in the observational sample, then predicts $Y$ in the experimental sample using the fitted $h$, and finally estimates the ATE using a linear regression of the predicted $Y$ on $(1, A, X)$.

\paragraph{Results}

table[table omitted — 1,042 chars of source]

(ref) presents the estimates of the treatment effects in the experimental sample. While the true benchmark shows a significant positive effect, the standard surrogate index method substantially underestimates the effect for both outcomes, and including the proxies $W$ and $Z$ as additional covariates does not improve the estimates. In contrast, our proposed method that accounts for unobserved confounding via the proxy variables yields estimates much closer to the benchmark results, albeit with larger standard errors.

\paragraph{Discussion}

As found in schochet2008does, program participation temporarily suppresses short-term earnings and employment because participants spend time in training instead of working. Theerefore, the substantial underestimation by the standard surrogate index likely arises from unobserved confounding, where latent ability $U$ positively affects both short-term ($S$) and long-term ($Y$) outcomes. This confounding creates a positive bias that attenuates the true negative structural relationship between $S$ and $Y$, where sacrificing short-term earnings represents an investment in future growth. Consequently, the standard method underestimates the true long-term effect of the program. By leveraging proxies to adjust for $U$, our method distinguishes between low earnings due to investment versus low ability, thereby recovering the experimental benchmark.

\paragraph{Diagnostic Analysis}

We perform a diagnostic analysis on the experimental sample to investigate the validity of the standard surrogacy assumption and the role of unobserved confounding. The results are summarized in (ref).

First, since we have access to the long-term outcome $Y$ in the experimental sample, we can directly test the surrogacy assumption by regressing $Y$ on $(1, A, S, X)$ and examining the coefficient of $A$. The first rows from Panel A and Panel B of (ref) show that the coefficient of $A$ is statistically significant for both outcomes, indicating a potential violation of the surrogacy assumption. Second, to assess whether the proxies $W$ and $Z$ help mitigate unobserved confounding, we further fit an IV regression, where the first stage regresses $W$ on $(1, Z, A, S, X)$ and the second stage regresses $Y$ on $(1, A, S, X, \hat W)$, with $Z$ acting as the instrument for $W$. In this setup, $\hat W$ captures the variation coming from the unobserved confounders $U$, so including it in the regression helps adjust for unobserved confounding. The second rows from Panel A and Panel B of (ref) show that the coefficient of $A$ becomes statistically insignificant for both outcomes after adjusting for $\hat W$, suggesting that the proxies $W$ and $Z$ effectively mitigate unobserved confounding, which results in the improved performance in (ref).

table[table omitted — 833 chars of source]

Conclusion

We study the challenging problem of identifying and estimating long-term treatment effects by combining experimental and observational data in the presence of unobserved confounding. The standard surrogate index methods rely on strong assumptions that rule out such confounding, leading to potentially biased estimates in many empirical settings.

We develop new methods for relaxing these stringent assumptions. Our methods rely on proxy variables that are informative about the unobserved confounders and satisfy certain conditional independence properties. We establish nonparametric identification results based on bridge functions, which generalize existing identification strategies that assume no unobserved confounding. We further propose estimation and inference procedures and establish their asymptotic properties, including consistency, asymptotic normality, and semiparametric efficiency.

Applying our methods to a real-world job training program evaluation, we find that the standard surrogate index approach substantially underestimates the true long-term effects, while our proposed methods that account for unobserved confounding via proxy variables yield estimates much closer to the benchmark results from the experimental data.

\phantomsection \addcontentsline{toc}{section}{References}