EconBase
← Back to paper

A Sensitivity Analysis of the Surrogate Index Approach for Estimating Long-Term Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

73,619 characters · 16 sections · 91 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Sensitivity Analysis of the Surrogate Index Approach for Estimating Long-Term Treatment Effects

abstractThis paper develops a sensitivity analysis of the surrogacy assumption for the surrogate index approach in athey2025surrogate. We introduce “Weighted Surrogate Indices (WSIs)," the analog of the surrogate index under the surrogacy assumption. We show that under comparability, the ATE on WSI identifies the ATE on the long-term outcome when a copula of the treatment and the long-term outcome conditional on baseline covariates and surrogates is known. When the copula is unknown, we establish the identified set of the ATE on the long-term outcome. Furthermore, we construct debiased estimators of the ATE for any given copula and develop asymptotically valid inference in both point-identified and partially identified cases. Using data from a poverty alleviation program in Pakistan, we demonstrate the importance of sensitivity checks as well as the usefulness of our approach. Keywords: Copula; Debiased Estimation; Partial Identification; Weighted Surrogate Index.

Introduction

In a variety of applications, researchers are interested in the effect of a treatment on outcomes, where the outcomes take a substantial amount of time to mature. If so, it may be useful to measure the effect of treatment on a set of intermediate or “surrogate" outcomes, which are predictive of long-term outcomes. For example, when measuring the effect of a job training program on long-term employment and earnings outcomes, short-term employment and earnings may serve as useful surrogates.

A literature has characterized the assumptions required to identify average treatment effects of interventions through surrogates when researchers have access to two data samples, an experimental and an observational data sample (athey2025surrogate). In this setting, the experimental data sample contains observations labeled with treatment assignments and surrogates. The observational data sample contains observations labeled with surrogates and long-term outcomes of interest. Both samples contain information on a set of pre-treatment characteristics. The goal is to identify the average treatment effect (ATE) on long-term outcomes of interest, even though long-term outcomes and treatments are observed in different data samples.

athey2025surrogate introduce the Surrogate Index (SI) and show that under four assumptions, the ATE of the long-term outcome is identified as the ATE of the SI. The SI is the conditional expectation of the long-term outcome given the surrogates and pre-treatment characteristics in the observational data sample. A critical assumption in athey2025surrogate is the Prentice Criterion (prentice1989SurrogateEndpoints) or the surrogacy assumption. The surrogacy assumption requires that, conditional on surrogates and pre-treatment characteristics, treatment assignment is independent of long-term outcomes. If this assumption holds, the effect of the treatment is fully mediated by the surrogates. However, the surrogacy assumption is often not satisfied (freedman1992). In many cases, the ATE is only partially mediated by the surrogates. For example, in medical contexts, low-density lipoprotein cholesterol (LDL-c) is often used as a surrogate for long-term cardiovascular disease (CVD), given that low levels of LDL-c are correlated with better CVD outcomes. However, some hormone replacement therapies have been found to lower LDL-c but also increase CVD through other causal pathways (yetley2017SurrogateDiseaseMarkers). In education, smaller class size may lead to changes in non-cognitive traits not fully captured by standardized exams ({heckman2006EffectsCognitiveNoncognitive}). In online advertising, increasing customer ad-load may generate short-term ad revenue but also frustrate customers, leading to long-term purchase drops not well-proxied by short-term ad-revenue (hohnhold2015FocusingLongterm).

When the surrogacy assumption fails, researchers have two available strategies.\footnote{When treatment is observed in the observational data set, the surrogacy assumption is not needed, see athey2025ExperimentalSelectionCorrectionEstimator, chen_ritzwoller_2023, obradovic2024TemporalLinks, and park2024DataCombinationUnderDynamicSelection.} First, they can find different surrogates or add more of them. Better surrogates or adding proxies of confounding surrogates may meet the surrogacy assumption, see cai2024SurrogatesRepresentationLearning. However, there are many applications where this strategy fails (bernard2023EstimatingLongtermTreatment). Second, they can evaluate worst-case bounds on the ATE leading to the worst-case identified set, see athey2025surrogate for binary outcomes. The worst-case identified set shows how much ATEs can vary when the surrogate assumption fails.

This paper contributes a third strategy. This strategy extends the scope of applications of the SI approach in athey2025surrogate from full mediation to partial mediation. This is accomplished by modeling the joint distribution of the long-term outcome and the treatment via a copula, conditional on surrogates and pre-treatment covariates. When the copula is the independence copula, the surrogacy assumption holds. Conversely, a general non-independence copula implies failure of the surrogacy assumption. It characterizes partial mediation of the surrogate variables on the effect of treatment on the long-term outcome.

We extend the identification result in athey2025surrogate from full mediation (the independence copula) to partial mediation (any non-independence copula). Specifically, we introduce “weighted surrogate indices (WSIs)” and show that the ATE on the long-term outcome is identified as the “ATE on WSIs”, see (ref). The WSIs reduce to the SI in athey2025surrogate when the copula is the independence copula. Numerically, we provide evidence of the sensitivity of the ATE to the strength of the global dependence between treatment and long-term outcome (which parameterizes the copula). We also show the robustness of the ATE to the shape of the copula. Using the identification result, we establish the sign of the surrogacy bias for stochastically monotone copulas.

Our approach subsumes point identification of the ATE under the surrogacy assumption in athey2025surrogate and the worst case bounds as special cases. For example, at one extreme, when the long-term outcome and treatment are independent (conditional on surrogates and pre-treatment variables), the surrogacy assumption holds and our result reduces to the point identification of the ATE in athey2025surrogate. At the other extreme, when the copula varies between the lower and upper bound copulas, our approach leads to the worst-case bounds on the ATE for any outcome type including binary outcomes as studied in athey2025surrogate. Researchers can evaluate the risk of the surrogacy assumption failing on ATEs for their application by evaluating the worst-case bounds.

More importantly, our identification result allows researchers to explicitly model scenarios in-between these two extremes. This is especially helpful when an application allows researchers to rule out certain relationships a priori. For example, in an educational setting, it may be reasonable to assume smaller class sizes are non-negatively related to future earnings after conditioning on short-term test score surrogates. These surrogates do not fully control for improvements in non-cognitive measures, which also improve future earnings. This can be achieved by varying the copula from an independence copula to the upper bound copula. Restricting the joint distribution of class size and future earnings in this way (conditional on short-term test scores) reduces the range of relevant ATE bounds.

We develop a complete set of estimation and inference methods for both the point identified case, when the copula is known, as well as the partially identified case, when the copula is unknown. Specifically, we construct doubly-robust estimators of the ATE for any given copula including the worst-case bounds. We establish the asymptotic normality of our doubly-robust estimator for any given copula. We also establish the joint normality of the doubly-robust estimators of the worst-case bounds. The orthogonal moment functions for non-independence copulas including the lower and upper Fr\'echet copulas are more complicated than those in chen_ritzwoller_2023 and athey2025surrogate. To verify the general conditions in chernozhukov2018doubleML, we adapt the technical proofs in chen_ritzwoller_2023 , dorn2024dvds, and semenova2025generalizedleebounds to our setting.

Empirically, using data from a poverty alleviation program in Pakistan (obtained from banerjee2015MultifacetedProgramCauses), we demonstrate that relaxing the surrogacy assumption can substantially alter conclusions. In particular, this often reverses the sign of the estimated treatment effect assuming surrogacy, which highlights the importance of these robustness checks.

The rest of the paper is organized as follows. Section (ref) reviews the setup and identification of the ATE using the surrogacy assumption. Section (ref) introduces weighted surrogate indices, showing the conditions under which the ATE is point identified. It also presents a numerical illustration of how the ATE changes as a function of Kendall's tau, which describes the concordance relationship between the treatment and long-term outcome (conditional on surrogates). Section (ref) focuses on the worst-case bounds, their debiased estimation, and inference results for the ATE when the surrogacy assumption fails. Section (ref) generalizes results in Section (ref) to a general copula and partial identification. Section (ref) applies the proposed framework to a household poverty alleviation dataset banerjee2015MultifacetedProgramCauses, conducting sensitivity analysis and deriving partial identification bounds. Section (ref) concludes. The technical details for the main result and the proof of the results in Section (ref) are presented in a supplementary appendix.

Setup and Identification Under Surrogacy Assumption

To be self-contained, this section formally reviews the setup and three critical assumptions used in athey2025surrogate to identify the ATE on the primary (long-term) outcome. This identification is achieved using two different datasets 1) an experimental data sample, which contains information on treatment assignment and baseline characteristics (but not the primary outcome), and 2) an observational data sample, which contains information on baseline characteristics and the primary outcome (but not treatment assignment). We also restate the surrogate index representation of the ATE in athey2025surrogate.

Let $W_i$ denote the treatment status (binary) of unit $i$, $Y_i(0),Y_i(1)$ denote the potential primary (long-term) outcomes, $S_i(0),S_i(1)$ denote the potential short-term outcomes (surrogacy variables), and $X_i$ denote a vector of baseline covariates that are not affected by treatment. Further, let $Y_i:=W_iY_i(1)+(1-W_i)Y_i(0)$ and $S_i:=W_iS_i(1)+(1-W_i)S_i(0)$.

assumptionWe have a single random sample of size $N$ drawn from the joint distribution of $(P_i, X_i, S_i,W_i, Y_i)$, where we observe for each unit in the sample $Z_i$, where $Z_i:=(P_i, X_i, S_i,\mathds{1}_{P_i=E}W_i, \mathds{1}_{P_i=O}Y_i)$.

(ref) states that sample information comes from two data sets: one labeled observational data that contains observations on $(P_i=O, X_i, S_i, Y_i)$ and the other labeled experimental data that contains observations on $(P_i=E, X_i, S_i,W_i)$.

remarkWe point out that (ref) can be replaced with the following assumption: $\{X_i, S_i, Y_i\}_{P_i =O}$ and $\{X_i, W_i, S_i\}_{P_i =E}$ are two independent samples and each is a random sample. This makes it clear that treatment status is not found in the observational data sample.

The experimental data satisfy the following assumption.

assumption[Unconfounded Treatment Assignment/Strong Ignorability] (i) \[W_i \ \rotatebox[origin=c]{90}{$\models$} \ (Y_i(0),Y_i(1),S_i(0),S_i(1))|X_i,P_i=E; \] (ii) $0<\rho(x)<1$ for all $x\in \mathcal{X}$, where $\rho(x):=\mathbb{P}(W_i=1|X_i=x,P_i=E)$ is the propensity score.

Let $\tau$ denote the ATE on the primary outcome in the population from which the experimental sample is drawn:

equation[equation omitted — 73 chars of source]

We are interested in identifying $\tau$ using the sample information satisfying (ref).

athey2025surrogate adopts two additional assumptions.

assumption[Surrogacy] (i) $W_i \ \rotatebox[origin=c]{90}{$\models$} \ Y_i|S_i,X_i, P_i=E;$ (ii) $ 0<\rho(s,x)<1$ for all $s\in\mathcal{S}, x\in \mathcal{X}$ and $0<\mathbb{P}(P_i=E)<1$, where $\rho(s,x):=\mathbb{P}(W_i=1|S_i=s,X_i=x,P_i=E)$ is the surrogacy score.

(ref) (surrogacy) implies that the conditional distribution of $Y_i$ given $(W_i, S_i,X_i, P_i=E)$ is the same as that of $Y_i$ given $ (S_i,X_i, P_i=E)$. Thus, given $X_i$, the surrogate variable $S_i$ completely mediates the effect of $W_i$ on the primary outcome $Y_i$ for the experimental sample. For expositional simplicity, we refer to this scenario as full mediation.

assumption[Comparability of Samples] (i) $P_i \ \rotatebox[origin=c]{90}{$\models$} \ Y_i|S_i,X_i;$ (ii) $\varphi(s,x)<1$ for all $s\in\mathcal{S}, x\in \mathcal{X}$, where $\varphi(s,x)=\mathbb{P}(P_i=E|S_i=s,X_i=x)$.
definition[Surrogate Index, athey2025surrogate] The surrogate index (SI) is the conditional expectation of the primary outcome given the surrogate outcomes and the pre-treatment variables, conditional on the sample: \[\mu(s,x,p)=\mathbb{E}\left[Y_i|S_i=s,X_i=x,P_i=p\right], \mbox{ } p=O,E.\]

The SI $\mu(S_i,X_i,O)$ is identified from the observational data. Theorem 1 in athey2025surrogate shows that (ref)-(ref) imply that $\tau$ is identified as the ATE on the SI, $\mu(S_i,X_i,O)$. We restate this result in the lemma below. The proof relies on the following expressions:

align[align omitted — 290 chars of source]
lemma[athey2025surrogate] Under (ref)-(ref), $\tau$ is identified as \begin{align*} \tau& = \mathbb{E}\left[ \mu(S_i,X_i,O)\frac{W_i }{\rho(X_i)} \mid P_i = E \right] - \mathbb{E}\left[ \mu(S_i,X_i,O)\left(\frac{1-W_i }{1-\rho(X_i)}\right)\mid P_i = E \right]\\ &= \mathbb{E}\left[\frac{\rho(S_i,X_i) }{\rho(X_i)(1-\rho(X_i))} \mu(S_i,X_i,O) \mid P_i = E \right] - \mathbb{E}\left[ \frac{ 1 }{1-\rho(X_i)}\mu(S_i,X_i,O)\mid P_i = E \right]. \end{align*}
remarkSubsequent work such as yang2024TargetingLongTermOutcomes extends the SI approach in athey2025surrogate to CATE estimation and policy learning using SI $\mu(S_i,X_i,O)$. Specifically, the CATE denoted as $\tau(x)$ is identified as \begin{align*} \tau(x)=\mathbb{E}\left[\mu(S_i,x, O)\mid W_i=1,X_i=x, P_i=E \right]-\mathbb{E}\left[\mu(S_i,x, O)\mid W_i=0,X_i=x, P_i=E \right] \end{align*} and the first-best policy is $\mathds{1}(\tau(x)>0)$.

Point Identification Without Surrogacy Assumption

The surrogate index $\mu(S_i,X_i,O)$ identified from the observational data plays a critical role in athey2025surrogate: under (ref)-(ref), $\tau$ is identified as the ATE on the SI, $\mu(S_i,X_i,O)$. athey2025surrogate discuss the plausibility and implications of violation of either (ref) or (ref). Under (ref), athey2025surrogate state that without (ref), the ATE on the SI may not identify the ATE on the primary outcome. Motivated by this and the abundant empirical evidence on the possible failure of (ref), we propose a new framework to study identification of the ATE on the primary outcome relaxing the surrogacy assumption.

By the conditional version of Sklar's Theorem, there exists a copula $C_o(u,v|S_i,X_i, P_i=E)$, $(u,v)\in [0,1]^2$, such that

align[align omitted — 138 chars of source]
assumptionSuppose that $C_o$ in (ref) is known.

Since $W$ is discrete, $C_o$ is not unique. However, as we show later, for continous outcomes, the identification result depends on the unique sub-copula only. In the special case that the sub-copula is the independence sub-copula, (ref) is equivalent to (ref) and the surrogate $S_i$ fully mediates the effect of $W_i$ on $Y_i$ given $X_i$.

example[Families of Copulas] (i) A Gaussian copula with constant correlation $\vartheta$ takes the following form: \[C_o(u,v|S_i,X_i, P_i=E) = \Phi_\vartheta(\Phi^{-1}(u), \Phi^{-1}(v)) = C_o(u,v|P_i=E),\] where $\Phi(\cdot)$ is the cumulative distribution function of the standard normal distribution and $\Phi_\vartheta(\cdot, \cdot)$ is the cumulative distribution function of the standard bivariate normal distribution with correlation $\vartheta\in [0,1]$. (ii) Archimedean copulas are a widely used class of copulas defined by a generator function $\phi$: \[ C(u_1, \dots, u_n) = \phi\left(\phi^{-1}(u_1) + \cdots + \phi^{-1}(u_n)\right), \] where $\phi: [0, \infty) \to [0, 1]$ is a strictly decreasing, convex function satisfying $\phi(0) = 1$ and $\phi(\infty) = 0$. Notable examples of Archimedean copulas are Clayton with generator function $\phi(t) = (1 + t)^{-1/\vartheta}$ for $\vartheta\in[-1,\infty)\setminus\{0\}$; Gumbel copula with $\phi(t) = \exp(-t^{1/\vartheta})$ for $\vartheta\in[1,\infty)$; and Frank copula with $\phi(t) = -\frac{1}{\vartheta} \log\left(1 - (1 - e^{-\vartheta}) e^{-t}\right)$ for $\vartheta\in\mathbb{R}\setminus\{0\}$. (iii) Let $C_-(u,v):=\max(u+v-1,0)$ and $C_+(u,v):=\min(u,v)$ denote the Fr\'echet-Hoeffding lower and upper bound copulas. When $C_o=C_-$ or $C_o=C_+$, (ref) implies that $Y_i$ and $W_i$ are comonotonically dependent on each other given $S_i,X_i$ so that $S_i$ does not mediate any effect of $W_i$ on $Y_i$ given $X_i$.

In the rest of this section, we establish point identification of the ATE on the primary outcome under (ref), extending the identification under full mediation or (ref) in athey2025surrogate to known partial mediation or (ref).

For notational convenience, we sometimes ignore the conditional arguments in the copula and write $C_o(u,v)$ instead of $C_o(u,v|S_i,X_i, P_i=E)$ and denote the ATE as $\tau_{C_o}$ to indicate its dependence on the copula $C_o$.

Weighted Surrogate Indices

Our identification strategy relies on the Weighted Surrogate Index (WSI) introduced in the following.

definition[Weighted Surrogate Indices] Let $(U,V)\sim C_o(u,v)$, where $C_o$ is defined in (ref). For $\alpha\in (0,1)$, let $C_o(\alpha|u) := \Pr(V \le \alpha |U = u)$. We define two Weighted Surrogate Indices associated with $w=0,1$ for each $p=E,O$ as \begin{align*} \mu_{C_o, w}(S_i,X_i,p) = \int_{0}^{1} F_Y^{-1}(u|S_i,X_i, P_i=p)\sigma_{C_o, w}(u; 1- \rho(X_i,S_i)) d u, \end{align*} where $ \sigma_{C_o, w}(u; \alpha) =\left(w - C_o(\alpha|u)\right)/(w-\alpha).$

The WSIs in (ref) extend the SI in athey2025surrogate. We show in (ref) that the ATE of the primary outcome is identified as ATE of the WSIs defined in (ref). When the outcomes are continuous and $C_o$ is smooth in $u$, the WSIs depend only on the unique sub-copula. To see this, we note that $\mu_{C_o, 1}$ ($\mu_{C_o, 0}$) depends on $C_o(1-\rho(S_i, X_i)|u) = \partial_u C_o(u, 1-\rho(S_i, X_i)$ (or $C_o(u, 1-\rho(S_i, X_i)$) only: $C_o(u, 1-\rho(S_i, X_i)$ is unique because the sup-copula of $Y$ and $W$ given $S_i, X_i, P_i = E$ is unique. When the sub-copula is the independence sub-copula, $\sigma_{C_o, w}(u;\alpha) = 1$ and \[\mu_{C_o, w}(S_i,X_i,O) =\int_{0}^{1} F_Y^{-1}(u|S_i,X_i, P_i=O) d u=\mu(S_i,X_i,O), \mbox{ for } w=0,1\] where $\mu(S_i,X_i,O)$ is the SI in athey2025surrogate. For a non-independence sub-copula, the weights $\sigma_{C_o, 1}(u;\alpha)$ and $\sigma_{C_o, 0}(u;\alpha)$ are generally non-constant functions of $u\in[0,1]$ and are not equal leading to two different surrogate indices $\mu_{C_o, w}(S_i,X_i,O)$ for $w=0,1$.

remarkIt is interesting to observe that WSIs are closely related to two distinct classes of functions, one in finance and risk management and the other in social choice and welfare. Specifically, let \begin{align*} r := \int_0^1 F_Y^{-1}(u) \sigma(u) du. \end{align*} When $\sigma(\cdot)$ is a non-decreasing function, $r$ is a Distortion Risk Measure, where higher ranked $Y$'s are given higher weight, see pichler2015 and pflug2006SubdifferentialRepresentationsRisk. When $\sigma(\cdot)$ is a non-increasing function, $r$ is known as Rank-dependent Social Welfare Function, where higher ranked $Y$'s are given lower weight, see yaari1987DualTheoryChoice. For a general copula $C_o$, the weight $\sigma_{C_o, w}(u;\alpha)$ in WSIs may not be monotone.
example[AVaR and WSIs for Upper and Lower Bound Copulas] When $\sigma(u) = \mathds{1}(u \in (\alpha, 1])/(1-\alpha)$ for some $\alpha\in [0,1)$, $r$ is the well-known AVaR of $Y$ at level $\alpha$: \[AVaR_{\alpha}(Y)=\frac{1}{1-\alpha}\int_\alpha^1 F_Y^{-1}(u)du.\] It is insightful to examine the WSIs when the outcome and treatment are perfectly dependent on each other conditional on the surrogates and pre-treatment covariates. When $C_o=C_+$, for any $\alpha\in (0,1)$, it holds that \begin{align*} \sigma_{C_{+}, 1}(u;\alpha) = \frac{\mathds{1}(u \in (\alpha, 1])}{1-\alpha} and \sigma_{C_{+}, 0}(u;\alpha) = \frac{\mathds{1}(u \in [0, \alpha])}{\alpha}. \end{align*} Thus for $w=1$, the top $\rho(X_i,S_i)$ percentile of the conditional distribution of $Y_i$ is given equal positive weight and the bottom $1-\rho(X_i,S_i)$ percentile is given zero weight; for $w=0$, the bottom $1-\rho(X_i,S_i)$ percentile is given equal positive weight and the top $\rho(X_i,S_i)$ is given zero weight. Consequently, \begin{align*} \mu_{C_+, 1}(S_i,X_i,O) &=\int_{0}^{1} F_Y^{-1}(u|S_i,X_i, P_i=O)\frac{\mathds{1}(u \in (1- \rho(X_i,S_i), 1])}{\rho(X_i,S_i)} d u\\ &=AVaR_{1- \rho(X_i,S_i)}(Y_i\mid S_i,X_i,P_i=O) and\\ \mu_{C_+, 0}(S_i,X_i,O)&=-AVaR_{\rho(X_i,S_i)}(-Y_i\mid S_i,X_i,P_i=O). \end{align*} Similarly, for $C_o=C_-$, \begin{align*} \sigma_{C_-, 1}(u;\alpha) = \frac{\mathds{1}(u\in[0,1-\alpha])}{1-\alpha} and \sigma_{C_-, 0}(u;\alpha) = \frac{\mathds{1}(u\in(1-\alpha,1])}{\alpha}. \end{align*} As a result, $\mu_{C_-, 1}(S_i,X_i,O)$ is the conditional mean of $Y_i$ for the bottom $\rho(X_i,S_i)$ percentile of the conditional distribution of $Y_i$ and $\mu_{C_-, 0}(S_i,X_i,O)$ is the conditional mean of $Y_i$ for the top $1-\rho(X_i,S_i)$ percentile.

Identification of ATE

Theorem (ref) below shows that “the ATE on WSIs” defined on the right hand side of (ref) identifies the ATE of the primary outcome thus extending (ref) or Theorem 1 in athey2025surrogate under the surrogacy assumption to partial mediation.

Consider $\mathbb{E}[Y_i(1)|P_i=E]$. Under (ref) (unconfoundedness), (ref) implies that

equation*[equation* omitted — 154 chars of source]

Similarly, \[\mathbb{E}[Y_i(0)|P_i=E] =\mathbb{E}\left[\frac{1-\rho(X_i,S_i)}{1-\rho(X_i)}\mathbb{E}\left[Y_i|W_i=0,S_i,X_i,P_i=E \right]|P_i=E\right].\] Below we extend Proposition 2 (iii) in athey2025surrogate under the surrogacy assumption to any copula $C_o$.

propositionUnder (ref), it holds that $$ \mu_{C_o, w}(S_i,X_i,O)=\mu_{C_o, w}(S_i,X_i,E)=\mathbb{E}[Y_i | W_i=w, X_i, S_i, P_i=E] \mbox{ for } w = 0, 1.$$

(ref) implies that $\mathbb{E}\left[Y_i|W_i=w,S_i,X_i,P_i=E\right]=\mu_{C_o, w}(S_i,X_i,E)$ and under comparability, $\mu_{C_o, w}(S_i,X_i,E)$ is identified as $\mu_{C_o, w}(S_i,X_i,O)$ for any copula $C_o$ and hence (ref) holds.

theorem[Known Sub-copula] Suppose (ref), (ref), (ref), and (ref) hold. Then $\tau_{C_o}$ is identified as \begin{align} \tau_{C_o} & =\mathbb{E}\left[ \mu_{C_o, 1}(S_i,X_i,O)\frac{W_i }{\rho(X_i)} \mid P_i = E \right] - \mathbb{E}\left[ \mu_{C_o, 0}(S_i,X_i,O)\left(\frac{1-W_i }{1-\rho(X_i)}\right)\mid P_i = E \right]\\ &= \mathbb{E}\left[ \frac{\rho(X_i,S_i)}{\rho(X_i)[1-\rho(X_i)]} \mu_{C_o, 1}(S_i,X_i,O) \mid P_i = E \right] - \mathbb{E}\left[ \frac{1}{1-\rho(X_i)} \mu(S_i,X_i,O) \mid P_i = E \right].\nonumber \end{align} When the conditional distribution of $Y_i$ given $S_i,X_i,P_i=O$ is degenerate, $ \tau_{C_o}= \tau_{\Pi} $ and is thus identified even if $C_o$ is unknown.

When the outcomes are continuous and $C_o$ is smooth in $u$, the WSIs depend on the unique sub-copula only and $\tau_{C_o}$ is identified from the sub-copula. Equipped with the WSIs $\left(\mu_{C_o, 1}(S_i,X_i,O), \mu_{C_o, 0}(S_i,X_i,O)\right)$, (ref) allows to identify ATE of the primary outcome regardless of full or partial mediation of $S_i$ on the effect of $W_i$ on $Y_i$ as long as $C_o$ is known which includes the lower and upper bound copulas.

remarkAnalogously to yang2024TargetingLongTermOutcomes, under (ref), (ref), (ref), and (ref), the CATE denoted as $\tau_{C_o}(x)$ is identified as \begin{align*} \tau_{C_o}(x)=\mathbb{E}\left[\mu_{C_o,1}(S_i,x, O)\mid W_i=1,X_i=x, P_i=E \right]-\mathbb{E}\left[\mu_{C_o, 0}(S_i,x, O)\mid W_i=0,X_i=x, P_i=E \right] \end{align*} and the first-best policy is $\mathds{1}(\tau_{C_o}(x)>0)$.

Surrogacy Bias---Stochastically Monotone Copulas

Theorem 4 (ii) in athey2025surrogate provides an expression for the surrogacy bias. For a copula $C_o$, it follows directly from (ref) that

equation[equation omitted — 195 chars of source]

For the class of stochastically monotone copulas, we will establish the sign of the surrogacy bias.

definitionLet $(U,V)\sim C_o$. Then $V$ is stochastically increasing (decreasing) in $U$, if $C_o(\alpha|u) = \Pr(V \le \alpha |U = u)$ decreases (increases) in $u$ for every $\alpha$.

Commonly used copulas are monotone copulas. For example, the Gaussian and Frank copulas are monotonically increasing when $\vartheta > 0$ and monotonically decreasing when $\vartheta < 0$. The Clayton copula is monotonically increasing for $\vartheta > 0$ and monotonically decreasing for $\vartheta \in (-1,0)$. The Gumbel copula is monotonically increasing. (ref) implies that when $V$ is stochastically increasing (decreasing) in $U$, $\sigma_{C_o, 1}(u; \alpha)$ is an increasing (decreasing) function of $u\in [0,1]$ for every $\alpha$ and $\sigma_{C_o, 0}(u; \alpha)$ is a decreasing (increasing) function of $u\in [0,1]$ for every $\alpha$. As a result, we expect $\tau_{C_o}$ to be larger (smaller) than $\tau_{\Pi}$ when $V$ is stochastically increasing (decreasing) in $U$. We show in the rest of this section that this is indeed the case.

definition[Concordance Order] The copula $C_1$ is smaller than the copula $C_2$ in concordance order denoted as $C_1\prec_c C_2$ iff $C_1(u,v)\leq C_2(u,v)$ for all $(u,v)\in [0,1]^2$.

It follows from the proof of (ref) that for any copula $C$,

align*[align* omitted — 103 chars of source]

This and Theorem 1 of cambanis_1976_inequalities imply the statement for $\mu_{C, 1}$ in (ref) below. The statement for $\mu_{C, 0}$ follows from that of $\mu_{C, 1}$ and the following relation: $$(1-\rho(s,x))\mu_{C, 0}(s,x,O)= \mathbb{E}[Y] - \rho(s,x)\mu_{C, 1}(s,x,O).$$

propositionFor any $(s,x)\in \mathcal{S}\otimes\mathcal{X}$, if $C_1(.,.|s,x, E)\prec_c C_2(.,.|s, x, E)$, then $ \mu_{C_1, 1}(s,x,O) \leq \mu_{C_2, 1}(s,x,O) $ and $ \mu_{C_1, 0}(s,x,O) \geq \mu_{C_2, 0}(s,x,O) $.

If $\{C_o(\cdot|u)\}$ is stochastically increasing in $u$, then $\Pi\prec_c C_o$, see Sections 2.8 and 8.3 of joe2014DependenceModelingCopulas. The following corollary follows from (ref) and (ref).

corollary[The Sign of the Surrogacy Bias] Suppose (ref), (ref), and (ref) hold. If $\{C_o(\cdot|u)\}$ is stochastically increasing (decreasing) in $u$, then $\tau_{C_o}-\tau_{\Pi}> 0 \mbox{ }(<0)$.

Consequently, if $\{C_o(\cdot|u)\}$ is stochastically increasing (decreasing) in $u$, then the ATE of the SI under-estimates (over-estimates) $\tau_{C_o}$, the ATE of the long term outcome.

A Numerical Illustration

In this section, we use several parametric families of copulas to gauge the sensitivity of $\tau_{C_o}$ on the shape of $C_o$ and the strength of global dependence by varying their parameters.

We consider the Gaussian copula and several copulas from the Archimedean family. For simplicity, we assume that the conditional copula is the same as the unconditional copula. For common (bivariate) copulas such as the Gaussian, Clayton, Gumbel, and Frank copulas, the dependency parameter can be expressed in terms of Kendall's tau $\varrho_K$, a rank-based measure of monotonic dependence. This parameterization makes dependency strength interpretable and comparable across different copula families. Appendix (ref) in the supplementary appendix collects the mapping between $\varrho_K$ and $\vartheta$ for the aforementioned copulas.

For illustration, consider the following data generating process for both the experimental and observational data samples: \[X_i\sim U[0,1], \quad W_i\sim Bernoulli(\rho), \quad S_i = X_i + W_i + \eta_i^S, \text{ where } \eta_i^S\sim\mathcal{N}(0,1),\] \[Y_i = S_i + 0.5X_i + \eta_i^Y, \text{ where } \eta_i^Y\sim\mathcal{N}(0,1).\] We investigate $\rho\in\{0.1, 0.5, 0.9\}$. Following Assumption (ref), $W_i|S_i,X_i =I\{\epsilon_i > 1 - \rho(S_i, X_i)\}$, where $\epsilon_i | S_i, X_i, P_i=E \sim U[0, 1]$, and the copula structure is for $\epsilon_i$ and $\eta_i^Y$, independent of $S_i, X_i$. More specifically, we set $C_o(u,v;\vartheta|S_i, X_i, P_i=E) = C_o(u,v;\vartheta|P_i=E)$, for $C_o(u,v;\vartheta|P_i=E)$ being the Gaussian, Clayton, Gumbel, and Frank copulas. For each copula family, we solve for $\vartheta$ for each $\varrho_K$ in the corresponding grid of appropriate $\varrho_K$ values. Note that the Clayton copula is more commonly used to model positive dependence, while the Gumbel copula exclusively models positive dependence. Therefore, the range of $\varrho_K$ is adjusted to ensure the corresponding $\vartheta$ falls within its valid range.

In Figure (ref), we plot how $\tau_{C_o}$ in Theorem (ref) changes with $\varrho_K$, with computational details in Appendix (ref). Consistent with the DGP, Figure (ref) shows that for any $\rho\in(0,1)$, if $C_o(u, v; \vartheta \mid P_i = E)$ is the independence copula ($\varrho_K=0$), the true long-term ATE is $\tau_{\Pi}=1$. The surrogacy bias $\tau_{C_o}-\tau_{\Pi}$ increases in magnitude as $|\varrho_K|$ increases. When $\rho=0.5$, the threshold for $\tau$ to change sign is $\varrho_K\approx-0.55$, which is consistent across the copulas considered here that allow for negative dependence (Gaussian and Frank). In other words, if the negative dependence between $W$ and $Y$ is strong enough with Kendall's tau smaller than $-0.55$, assuming an independence copula will produce the wrong sign for the long-term ATE. For $\rho=0.1$ or $0.9$, the same threshold for the Gaussian copula is around $-0.41$. Note that the form of the copula is not crucial for $\rho = 0.5$. This is intuitive because values of $\rho(s, x)$ tend to be around $0.5$ as well, making the conditional dependence $C_o(1-\rho(s,x)|u)$ in $\sigma_{C_o,1}(\cdot;\cdot)$ relatively insensitive to the copula family. In contrast, when $\rho$ takes more extreme values like $0.1$ or $0.9$, differences in tail dependence across copula families lead to more pronounced variation in conditional behavior, thereby affecting $\tau_{C_o}$. In general, in a randomized controlled trial with relatively balanced treatment and control groups, if $S$ depends only weakly on $W$ (i.e., $S$ carries limited information about $W$), the choice of copula serves mainly as a functional tool for achieving the desired dependence level, while $\varrho_K$ captures the essential dependence information relevant for $\tau_{C_o}$.

figure[figure omitted — 515 chars of source]

Worst-Case Bounds, Debiased Estimation, and Inference

In practice, the copula $C_o$ is rarely known. (ref) provides a numerical illustration of the application of (ref) to check sensitivity of $\tau_{\Pi}$ to the violation of the surrogacy assumption by letting the copula $C_o$ deviate from the independence copula. (ref) can also be used to establish sharp bounds on $\tau_{C_o}$ when $C_o$ is unknown but lies between two known copulas.

The following corollary follows immediately from (ref) and (ref).

corollary[Identified Set] Suppose (ref), (ref), and (ref) hold. Furthermore, suppose the copula $C_o$ satisfies $C_L(\cdot, \cdot|s,x,E)\prec_c C_o(\cdot, \cdot|s,x,E)\prec_c C_U(\cdot, \cdot|s,x,E)$ for almost all $(s,x)\in\mathcal{S}\otimes\mathcal{X}$, where $C_L(\cdot, \cdot |s,x,E)$ and $C_U(\cdot, \cdot | s,x,E)$ are two known copula functions. Then $\tau_{C_o}$ is partially identified with the identified set $ [\tau_{C_L}, \tau_{C_U}]$. When the conditional distribution of $Y_i$ given $S_i,X_i,P_i=O$ is degenerate, the identified set is singleton: $\tau_{C_L}=\tau_{C_U}=\tau_{\Pi}$.

(ref) implies that the worst case bounds on $\tau_{C_o}$ are $\tau_{C_-}$ and $ \tau_{C_+}$. As a result, $\tau_{C_-}$ can be interpreted as the smallest ATE when the surrogacy assumption fails. Conversely, $\tau_{C_+}$ is the largest ATE when the surrogacy assumption fails. Another implication is that the ATE under the surrogacy assumption in athey2025surrogate is the least ATE among copulas dominating the independence copula in concordance order. It is also the greatest ATE among copulas dominated by the independence copula in concordance order.

Comparison with Lemma 1 in athey2025surrogate

We restate the expressions for $\tau_{C_-}, \tau_{C_+}$ in the following proposition which also shows that they are the same as Lemma 1 (ii) in athey2025surrogate for binary outcomes.

propositionSuppose (ref), (ref), and (ref) hold. Then the identified set of $\tau_{C_o}$ is $[\tau_{C_-}, \tau_{C_+}]$, where $\tau_{C_-},\tau_{C_+}\in\mathbb{R}$ are given by \begin{align*} \tau_{C_-} & =\mathbb{E}\left[ \mu_{C_-, 1}(S_i,X_i,O)\frac{W_i }{\rho(X_i)} \mid P_i = E \right] - \mathbb{E}\left[ \mu_{C_-, 0}(S_i,X_i,O)\left(\frac{1-W_i }{1-\rho(X_i)}\right)\mid P_i = E \right],\\ \tau_{C_+} & =\mathbb{E}\left[ \mu_{C_+, 1}(S_i,X_i,O)\frac{W_i }{\rho(X_i)} \mid P_i = E \right] - \mathbb{E}\left[ \mu_{C_+, 0}(S_i,X_i,O)\left(\frac{1-W_i }{1-\rho(X_i)}\right)\mid P_i = E \right], \end{align*} in which \begin{align*} \mu_{C_-,0}(S_i,X_i,O)&=AVaR_{\rho(X_i,S_i)}(Y_i\mid S_i,X_i,P_i=O),\\ \mu_{C_-,1}(S_i,X_i,O)&=-AVaR_{1-\rho(X_i,S_i)}(-Y_i\mid S_i,X_i,P_i=O),\\ \mu_{C_+,1}(S_i,X_i,O)&=AVaR_{1-\rho(X_i,S_i)}(Y_i\mid S_i,X_i,P_i=O), and \\ \mu_{C_+,0}(S_i,X_i,O)&=-AVaR_{\rho(X_i,S_i)}(-Y_i\mid S_i,X_i,P_i=O). \end{align*} When the outcome is binary, the identified set $[\tau_{C_-}, \tau_{C_+}]$ is the same as Lemma 1 (ii) in athey2025surrogate.

(ref) implies that the ATE is partially identified regardless of the outcome type/range and for binary outcomes. The identified interval in (ref) is the same as that in Lemma 1 (ii) in Section 5.2 of athey2025surrogate. However, it differs from Lemma 1 (i) in Section 5.2 of athey2025surrogate which states that “If the outcome can take on values on the whole real line, then there is no value for the average treatment effect $\tau$ that can be ruled out.” Their proof seems to have ignored the fact that $F_Y(\cdot\mid s,x,E)$ is identified from $F_Y(\cdot\mid s,x,O)$ under (ref). In contrast, our proof makes use of the identified $F_Y(\cdot\mid s,x,E)$ and the following expression for $\mu_{C_o,1}(S_i,X_i,O)$ in (ref):

align*[align* omitted — 107 chars of source]

Since both $F_Y(y|s,x, E)$ and $F_W(w|s,x, E)$ are point identified, Theorem 1 of cambanis_1976_inequalities implies that $\mu_{C_o, 1}(s,x,O)$ is partially identified.

Debiased Estimation

athey2025surrogate and chen_ritzwoller_2023 develop debiased estimation of $\tau_{\Pi}$ for which WSIs reduce to the SI $\mu(S_i,X_i,O)=\mathbb{E}\left(Y_i\mid S_i,X_i,P_i=O\right)$. For the worst-case bounds, the WSIs are related to conditional AVaR of $Y_i$ instead of the conditional mean of $Y_i$ and also depend on the surrogacy score $\rho(S_i,X_i)$. As a result, it is more challenging to derive the orthogonal moment functions for the worst-case bounds $\tau_{C_+}$ and $\tau_{C_-}$.

To proceed, we make use of the dual representations of the worst-case WSIs in terms of conditional means of the following newly defined functions:

align*[align* omitted — 172 chars of source]

Specifically, from the dual form of AVaR (c.f., rockafellar2002ConditionalValueatriskGeneral and acerbi2002es), it follows that

align*[align* omitted — 428 chars of source]

where

align*[align* omitted — 161 chars of source]

Let $\varphi(x) := \mathbb{P}(P_i=E | X_i = x)$, $ \varphi := \mathbb{P}(P_i = E)$, and for $w = 0, 1$,

align*[align* omitted — 211 chars of source]

Finally, let $\eta$ denote the collection of all the nuisance functions, i.e.,

align*[align* omitted — 259 chars of source]

Then $\tau_{C_+}$ satisfies the moment condition: $ \mathbb{E}[m_{C_+}(Z_i, \tau_{C_+}, \eta)] = 0,$ where

align*[align* omitted — 1,131 chars of source]

The first three terms in the orthogonal moment function $ m_{C_+}$ are analogous to those in Equation (4.4) of athey2025surrogate and Theorem 3.1 of chen_ritzwoller_2023, and the last two terms are new and correct for the effect of estimating $\rho(S_i, X_i)$ in the dual representations of the WSIs.

Similarly, $\tau_{C_-}$ satisfies: $ \mathbb{E}[m_{C_-}(Z_i, \tau_{C_-}, \eta)] = 0,$ where

align*[align* omitted — 1,135 chars of source]

Our debiased estimators $\widehat{\tau}_{C_-}$ and $\widehat{\tau}_{C_+}$ are defined as the solutions to $\frac{1}{n}\sum_{i=1}^n m_{C_-}(Z_i, \tau, \widehat{\eta})=0$ and $\frac{1}{n}\sum_{i=1}^n m_{C_+}(Z_i, \tau, \widehat{\eta})=0$, respectively, where $\widehat{\eta}$ is estimated by cross-fitting over $K$ even folds, which ensures that the nuisance estimators remain independent of the samples to which they are applied chernozhukov2018doubleML.

Our orthogonal moment conditions in debiased estimation allow for greater flexibility in estimating nuisance parameters, enabling the use of both parametric regressions and machine learning methods. For example, the estimation of $\rho(x)$, $\rho(s, x)$, $\varphi(x)$, and $\varphi(s, x)$ readily accommodates methods such as Lasso, random forests, gradient boosting, and neural networks. In contrast, estimating the worst-case WSIs ($\mu_{C_{-}, 0}$, $\mu_{C_{-}, 1}$, $\mu_{C_{+}, 0}$, and $\mu_{C_{+}, 1}$) relies on previously obtained cross-fitted estimates of $\rho(X_i, S_i)$ and is more involved due to the computation of conditional AVaRs. We follow the two-stage, locally robust approach of olma2021 with our chosen model specifications: the first stage requires estimating a conditional quantile, which we do nonparametrically using quantile forests; the second stage involves fitting a linear sieve model to a generated outcome variable whose derivative with respect to the conditional quantile, evaluated at the truth, is zero. With estimates of the worst-case WSIs as pseudo-outcomes, $\bar{\mu}_{C-, 0}$, $\bar{\mu}_{C-, 1}$, $\bar{\mu}_{C+, 0}$, and $\bar{\mu}_{C+, 1}$ can then be estimated either parametrically or nonparametrically. Finally, $q_{C_+}$ and $q_{C_-}$ are again estimated using quantile forests with cross-fitted $\widehat{\rho}(X_i,S_i)$. More details can be found in Algorithm (ref) of Appendix (ref) in the supplementary materials.

Asymptotic Theory

We adopt conditions similar to those of chen_ritzwoller_2023, dorn2024dvds, and semenova2025generalizedleebounds to establish the asymptotic joint normality of $\widehat{\tau}_{C_-}$ and $\widehat{\tau}_{C_+}$.

assumption[Regularity Conditions] (i) For a continuous outcome, we assume that it's distribution is absolutely continuous with bounded support and conditional density function $f(y|s, x, O)$ that is continuous with respect to y for each $s$ and $x$, and is uniformly bounded above and below by positive absolute constants; For a binary outcome, we assume that the random variable $\left[\mu(S_i, X_i, O) - \rho(S_i, X_i)\right]$ has bounded density. (ii) There exists some absolute constant $t$ such that for each $w \in \{0, 1\}$ and $b \in \{C_-, C_+\}$, either (1) or (2-1, 2-2) holds with probability one: \begin{gather*} (1) \mathbb{E}[(\bar{\mu}_{b, 1}(1, X_i) - \bar{\mu}_{b, 0}(0, X_i) - \tau_{b})^2|X_i, P_i = E] > t or \\ (2-1) \mathrm{Var}\left(\frac{1}{\rho(X_i)}(Y_i - q_{C_+}(S_i, X_i, O))_{+} + \frac{1}{1-\rho(X_i)}(Y_i - q_{C_+}(S_i, X_i, O))_{-} | S_i, X_i, O)\right) > t and \\ (2-2) \mathrm{Var}\left(\frac{1}{1-\rho(X_i)}(Y_i - q_{C_-}(S_i, X_i, O))_{+} + \frac{1}{\rho(X_i)}(Y_i - q_{C_-}(S_i, X_i, O)_{-} | S_i, X_i, O)\right) > t. \end{gather*}

The assumption of bounded support of $Y$ and boundedness of the conditional density function in (ref) (i) are similar to semenova2025generalizedleebounds. Assumption (1) or (2-1, 2-2) in (ref) (ii) ensures that the asymptotic variances of $\tau_{C_+}$ and $\tau_{C_-}$ are positive.

assumption[Realization Set] Let $q>2$ be a positive constant and $\epsilon$ be a constant such that $0 < \epsilon < 1/2$, $\omega = (\omega_{\rho}, \omega_{\varphi})$ with $\omega_{\rho} =(\rho(S_i, X_i), \rho(X_i)), \mbox{ } \omega_{\varphi}=(\varphi(S_i, X_i), \varphi )$, and \begin{align*} \kappa & = (\mu_{C_+, 1}(S_i, X_i, O), \mu_{C_+, 0}(S_i, X_i, O), \mu_{C_-, 1}(S_i, X_i, O), \mu_{C_-, 0}(S_i, X_i, O), \\ & \quad \quad \bar{\mu}_{C_+, 1}(1, X_i), \bar{\mu}_{C_+, 0}(0, X_i), \bar{\mu}_{C_-, 1}(1, X_i), \bar{\mu}_{C_-, 0}(0, X_i), q_{C_+}(S_i, X_i, O), q_{C_-}(S_i, X_i, O) ). \end{align*} For all probability measures $P$ satisfying Assumptions (ref), (ref), and (ref), the following condition holds; for some sequences $\Delta_n \to 0$ and $\delta_n \to 0$ with $\delta_n \ge n^{-1/2}$ with probability $1 - \Delta_n$, the estimator of nuisance parameter belongs to the realization set $R_n$ which contains $\widetilde{\eta}$ such that \begin{gather} \| \widetilde{\eta} - \eta \|_{P, q} \le C, \mathbb{P}(\epsilon \le \widetilde{\rho}(S_i, X_i) \le 1-\epsilon) = 1, \notag \\ \mathbb{P}(\epsilon \le \widetilde{\varphi}(S_i, X_i) \le 1-\epsilon) = 1, \| \widetilde{\eta} - \eta \|_{P, 2} \le \delta_n, \notag \\ \| \widetilde{\omega} - \omega \|_{P, 2} \times \| \widetilde{\kappa} - \kappa \|_{P, 2} \le \delta_n n^{-1/2}, \notag \\ \| \widetilde{\omega}_{\rho} - \omega_{\rho} \|_{P, 2} \times \| \widetilde{\omega}_{\varphi} - \omega_{\varphi} \|_{P, 2} \le \delta_n n^{-1/2}, \end{gather} and for continuous outcomes, \begin{align} \| \widetilde{q}_{C_+} - q_{C_+} \|_{P, 2}^2 \le \delta_n n^{-1/2} ; \end{align} for binary outcomes, \begin{align} \| \widetilde{\mu}(S_i, X_i, O) - \mu(S_i, X_i, O) \|_{\infty} \le \delta_n, \| \widetilde{\mu}(S_i, X_i, O) - \mu(S_i, X_i, 0) \|_{\infty}^2 \le \delta_n n^{-1/2}. \notag \end{align}

Most of the conditions in (ref) are similar to those in chen_ritzwoller_2023. Since our moment function has additional correction terms for $\rho$, we impose additional conditions on the rate of the cross-product in (ref). Condition (ref) and conditions for the binary case are similar to dorn2024dvds.

theoremSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Then \begin{align*} \sqrt{n} \begin{pmatrix} \widehat{\tau}_{C_+} - \tau_{C_+} \\ \widehat{\tau}_{C_-} - \tau_{C_-} \end{pmatrix} \to N \left(\begin{pmatrix} 0 \\ 0 \end{pmatrix}, \begin{pmatrix} \sigma_{C_+}^2 & \rho \sigma_{C_+} \sigma_{C_-} \\ \rho \sigma_{C_+} \sigma_{C_-} & \sigma_{C_-}^2 \end{pmatrix} \right), \end{align*} where $\sigma_{C_+}^2 = \mathbb{E}[m(Z_i, \tau_{C_+}, \eta_{0})^2] > 0$, $\sigma_{C_-}^2 = \mathbb{E}[m(Z_i, \tau_{C_-}, \eta_{0})^2] > 0$, and $-1 \le \rho \le 1$.

Our proof strategy builds on chen_ritzwoller_2023 and dorn2024dvds. Departing from chen_ritzwoller_2023 which checks the orthogonality condition in Assumption 3.1 and the statistical rate for second-order derivative of moment condition $r_N'$ in Assumption 3.2 (c) in chernozhukov2018doubleML, similar to dorn2024dvds, we directly verified the following high-level condition in the proof of Theorem 3.1 of chernozhukov2018doubleML:

align*[align* omitted — 149 chars of source]

Note that this high-level condition is satisfied when both the orthogonality condition in Assumption 3.1 in chernozhukov2018doubleML and the statistical rate for second-order derivative of moment condition $r_N'$ in Assumption 3.2 (c) in chernozhukov2018doubleML hold.

Let $V$ denote the asymptotic variance-covariance matrix in Theorem (ref). That is,

align*[align* omitted — 181 chars of source]

We provide a consistent estimator of $V$ by following Theorem 3.2 in chernozhukov2018doubleML.

theoremSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. In addition, we assume that $\delta_n \ge n^{-[((1-2/q) \wedge 1/2]}$ for all $n \ge 1$. Then, $V$ can be consistently estimated by \begin{align*} \frac{1}{K} \sum_{k=1}^K \mathbb{E}_{r, k} [(m_{C_+}(Z_i, \hat{\tau}_{C_+}, \hat{\eta}_{k}), m_{C_-}(Z_i, \hat{\tau}_{C_-}, \hat{\eta}_{k})) (m_{C_+}(Z_i, \hat{\tau}_{C_+}, \hat{\eta}_{k}), m_{C_-}(Z_i, \hat{\tau}_{C_-}, \hat{\eta}_{k}))^{\top}], \end{align*} where $K$ is the number of folds in K-fold cross-fitting, and $r := [n/K]$ is the number of observations in each fold, and $E_{r, k}$ is operator for sample expectation from empirical data in $k$-the fold, i.e., $E_{r, k}[g(X_i)] = r^{-1} \sum_{i \in \mathcal{F}_k} g(X_i)$ where $\mathcal{F}_k$ is the $k$-th fold of indices $\{1, \dotsc, n\}$.

Wald inference on each bound is straightfoward and inference on the true ATE can be carried out by applying the misspecification-adaptive confidence interval in stoye2020SimpleShortNeverEmpty which allows for the covariance matrix to be degenerate.

Debiased Estimation and Inference---General Copula

The debiased estimators of $\tau_{C_+}$ and $\tau_{C_-}$ constructed in (ref) rely critically on the dual representations of the WSIs (associated with $C_+$ and $C_-$) obtained from the existing dual representation of AVaR.

For a general copula $C_o$, we establish dual representation for $r$ introduced in (ref) when $\sigma$ is of bounded variation in (ref) below. Our proof builds on the proof of pichler2015 and Section 2.4.2. in pflug2007ModelingMeasuringManagingRIsk for non-decreasing function $\sigma(\cdot)$ and the dual representation for AVaR.

lemmaAssume that $Y$ is bounded and $\sigma(u)$ is of bounded variation on [0, 1]. Then, \begin{align} r & = \sigma(0) \mathbb{E}[Y] + \int_0^1 \left((1-u) F_Y^{-1}(u) + \mathbb{E}[[Y-F_Y^{-1}(u)]_+]\right) d \sigma(u) \\ & = \sigma(1) \mathbb{E}[Y] - \int_0^1 \left(u F_Y^{-1}(u) - \mathbb{E}[[Y-F_Y^{-1}(u)]_-]\right) d \sigma(u) . \end{align}

(ref) extends the dual representation for AVaR to $r$ with a general weight function $\sigma$. However, in contrast to AVaR for which $F_Y^{-1}(u) \in \operatorname{\mathop{\mathrm{argmin}}}_G \{(1-u) G(u) + \mathbb{E}[[Y-G(u)]_+\}$, $F_Y^{-1}$ may not be an argmin of the following minimization problem:

align[align omitted — 130 chars of source]

This is because the sign of the second-order derivative may not be positive in the minimization problem (ref).

Orthogonal Moment Function and Debiased Estimator

We construct a debiased estimator of $\tau_{C_o}$ from the dual representation in (ref) under the following assumption.

assumptionThe outcome variable is continuous variable, and it has a bounded support. In addition, $C_o(\alpha|\cdot)$ is of bounded variation.

(ref) implies that under (ref),

align[align omitted — 314 chars of source]

where

align*[align* omitted — 583 chars of source]

Similar to orthogonal moment functions in the worst case bounds in (ref), we construct the following orthogonal moment function for a general $C_o$ using the dual representation of the WSIs in (ref):

align[align omitted — 1,218 chars of source]

where

align*[align* omitted — 242 chars of source]

in which

align*[align* omitted — 205 chars of source]

Asymptotic Theory

assumption[Regularity Conditions] (i) There exists some absolute constant $t$ such that with probability 1, either (1) or (2) holds: { \begin{gather*} (1) \mathbb{E}[(\bar{\mu}_{C_o, 1}(1, X_i) - \bar{\mu}_{C_o, 0}(0, X_i) - \tau_{C_o})^2|X_i, P_i = E] > t; \\ (2) \mathrm{Var}\left(\frac {\rho(S_i, X_i)}{\rho(X_i)} h_{C_o, 1}(Y_i, F_Y^{-1}, 1-\rho(S_i, X_i)) - \frac {1-\rho(S_i, X_i)}{1-\rho(X_i)} h_{C_o, 0}(Y_i, F_Y^{-1}, 1-\rho(S_i, X_i)\Big|S_i, X_i, O) \right) > t. \end{gather*} } (ii) $C_o(\alpha|u)$ is continuously twice differentiable with respect to $u, \alpha$, and \begin{gather*} \sup_{\alpha \in [\epsilon, 1-\epsilon], u \in [0, 1]}\left|\frac{\partial C_o(\alpha|u)}{\partial u}\right|, \sup_{\alpha \in [\epsilon, 1-\epsilon], u \in [0, 1]}\left|\frac{\partial c_o(\alpha|u)}{\partial u}\right|, \sup_{\alpha \in [\epsilon, 1-\epsilon], u \in [0, 1]}\left|\frac{\partial^2 c_o(\alpha|u)}{\partial \alpha \partial u}\right|, \end{gather*} $\sup_{\alpha \in [\epsilon, 1-\epsilon], u \in [0, 1]} |c_o(\alpha|u)|$, and $\sup_{\alpha \in [\epsilon, 1-\epsilon], u \in [0, 1]}\left|\frac{\partial c_o(\alpha|u)}{\partial \alpha}\right|$ are all bounded from above by absolute positive constant for some $\epsilon > 0$.

The conditions in (ref) (i) are identical to those in (ref) (ii): Condition (2) in (ref) (ii) is reduced to Condition (2) in (ref) (ii) when $C_o$ is $C_+$ or $C_-$. The conditions in (ref) (iii) imply that $C_o(\alpha|u)$ and $c_o(\alpha|u)$ are absolute continuous with respect to $u$ so (ref) is applicable.

assumption[Realization Set] Let $q>2$ be a constant and $\epsilon$ be a constant such that $0 < \epsilon < 1/2$, $\omega = (\omega_{\rho}, \omega_{\varphi})$ in which $\omega_{\rho} =(\rho(S_i, X_i), \rho(X_i)), \omega_{\varphi}=(\varphi(S_i, X_i), \varphi )$, and \begin{align*} \kappa = (\mu_{C_o, 1}, \bar{\mu}_{C_o, 1}, \mu_{C_o, 0}, \bar{\mu}_{C_o, 0}, d_{C_o}(S_i, X_i, O), F_Y^{-1}(\cdot|S_i, X_i, O)). \end{align*} and $\eta = (\omega, \kappa)$. For all $P$ satisfying Assumptions (ref), (ref), and (ref), the following condition holds: for some sequences $\Delta_n \to 0$ and $\delta_n \to 0$ with $\delta_n \ge n^{-1/2}$ with probability $1 - \Delta_n$, the estimator of nuisance parameter belongs to the realization set $R_n$ which contains $\widetilde{\eta}$ such that \begin{gather*} \| \widetilde{\eta} - \eta \|_{P, q} \le C, \mathbb{P}(\epsilon \le \widetilde{\rho}(S_i, X_i) \le 1-\epsilon) = 1, \mathbb{P}(\epsilon \le \widetilde{\varphi}(S_i, X_i) \le 1-\epsilon) = 1, \\ \| \widetilde{\omega}_{\rho} - \omega_{\rho} \|_{P, 2} \times \| \widetilde{\omega}_{\varphi} - \omega_{\varphi} \|_{P, 2} \le \delta_n n^{-1/2}, \| \hat{\omega} - \omega \|_{P, 2} \times \| \hat{\kappa} - \kappa \|_{P, 2} \le \delta_n n^{-1/2}, \\ \|\widetilde{\rho}(S_i, X_i) - \rho(S_i, X_i)\|_{P, 2}^2 \le \delta_n n^{-1/2}, \\ \|\widetilde{F}_Y^{-1}(\cdot|S_i, X_i, O) - F_Y^{-1}(\cdot|S_i, X_i, O) \|_{P, 2}^2 \le \delta_n n^{-1/2}. \end{gather*} Here, with abuse of notation, we denote the $L_q$ norm of a function of the form of $G(u|S_i, X_i)$ where $u \in [0, 1]$ by $\| G(\cdot|S_i, X_i) \|_{P, q} = \left(\mathbb{E}[\int_0^1 G^q(u|S_i, X_i) du]\right)^{1/q}$.

We verify conditions related to Theorem 3.1 in chernozhukov2018doubleML. In particular, we verify (ref) in the appendix.

theoremUnder Assumptions (ref), (ref), (ref), (ref), (ref), (ref) (i), (ref), and (ref), $\sqrt{n} (\widehat{\tau}_{C_o} - \tau_{C_o}) \to N(0, \sigma_{C_o}^2)$, where $\sigma_{C_o}^2 = \mathbb{E}[m(Z_i, \tau_{C_0}, \eta)^2]$.
remark(i) When $C_o$ is known, Wald inference can be constructed from (ref). (ii) Suppose $C_o$ is unknown and satisfies $C_L(\cdot, \cdot|s,x,E)\prec_c C_o(\cdot, \cdot|s,x,E)\prec_c C_U(\cdot, \cdot|s,x,E)$ for almost all $(s,x)\in\mathcal{S}\otimes\mathcal{X}$, where $C_L(\cdot, \cdot |s,x,E)$ and $C_U(\cdot, \cdot | s,x,E)$ are two known copula functions. Then (ref) implies that $\tau_{C_o}$ is partially identified with the identified set $ [\tau_{C_L}, \tau_{C_U}]$. (ref) can be extended to the joint asymptotic normality of $(\widehat{\tau}_{C_L},\widehat{\tau}_{C_U})$ and inference for $\tau_{C_o}$ can be done in the same way as in Section (ref).

Empirical Application

In the same spirit as (ref), this section presents results from a sensitivity analysis using the household poverty alleviation dataset used in banerjee2015MultifacetedProgramCauses. In this data set, the treatment program allocated productive assets to randomly selected households. The baseline variables are welfare-related measurements taken before treatment. The short-term outcomes consist of the same set of measurements taken two years after treatment, while the long-term outcome is one of these measurements recorded three years after treatment. We take the data from Pakistan, which has 446 treated units and 408 control units. For illustrative purposes, consider $X$ and $S$ to include the following five welfare indicators: per capita consumption, the food security index, the household asset index, the total amount borrowed, and agricultural income. The long-term outcome, $Y$, represents the household asset value in dollars three years post-treatment.

The poverty alleviation data is both experimental and observational in that $W_i$ and $Y_i$ are observed for all households in the sample. To illustrate our method, we randomly and evenly split the Pakistan data into two parts, removing $Y_i$ from the experimental sample and $W_i$ from the observational sample. $\tau_{C_{\vartheta}}$ is estimated given $C_{\vartheta}$, where $C_{\vartheta}$ is a Frank copula whose dependence parameter $\vartheta$ is calibrated to match values of Kendall’s tau in the set $\{-0.9, -0.75, -0.5, -0.25, -0.1, 0, 0.1, 0.25, 0.5, 0.75, 0.9\}$. To estimate the nuisance components in the orthogonal moment function ((ref)), we employ quantile forests (quantile_forest from the grf package) to estimate the conditional quantile function $F_{Y}^{-1}(\cdot \mid S_i, X_i, O)$; use Lasso regressions (cv.glmnet) to estimate the conditional expectation of the WSI $\bar{\mu}_{C_o, w}(w, X_i)$; and apply logistic Lasso regressions to estimate the propensity and surrogacy scores $\rho(X_i)$, $\rho(S_i, X_i)$, as well as the selection probability $\varphi(S_i, X_i)$. The nuisance functions are cross-fitted with three data folds. Figure (ref) shows how $\hat{\tau}_{C_\vartheta}$ varies with Kendall's tau. The estimated worst-case bounds are [$-\$131.8$, $\$148.8$], with 95% confidence intervals of [$-\$158.9$, $-\$104.8$] for the lower bound and [$\$124.3$, $\$173.4$] for the upper bound.

figure[figure omitted — 237 chars of source]

We then repeat the analysis using the Plackett copula. The results, reported in Figure (ref), indicate that the estimated treatment effects are largely insensitive to the choice of copula family: for a given Kendall's tau, the point estimates under the Plackett copula closely align with those obtained using the Frank copula. This robustness is consistent with the structure of the empirical application. In our data, both the estimated propensity score $\widehat\rho(X_i)$ and the estimated surrogacy score $\widehat\rho(S_i,X_i)$ are tightly centered around $0.5$ (with means $\approx0.54$), implying that most observations enter the conditional copula $C_o(1-\widehat\rho(s,x)\mid u)$ in regions far from the tails where copula families differ most sharply in their dependence behavior. Hence, the copula family itself has little influence on $\widehat{\tau}_{C_\vartheta}$. What matters most for the sensitivity analysis is the overall dependence level captured by Kendall's tau. The wide identified interval underscores the need for caution when interpreting the results under the surrogacy assumption. Even moderate departures from the assumed dependence structure can lead to substantial variation in the long-term treatment effect.

figure[figure omitted — 246 chars of source]

To further examine the behavior of the welfare estimate near the surrogacy benchmark, Figure (ref) takes a closer look at small positive values of Kendall's tau near zero. Specifically, we conduct a local sensitivity analysis over a fine grid of small positive dependence levels, re-estimating the long-term treatment effect at each value. This zoomed-in analysis allows us to pinpoint the minimum degree of dependence at which the welfare estimate becomes statistically distinguishable from zero. For Pakistan, once Kendall's tau exceeds approximately 0.032, the confidence intervals no longer include zero, and the estimated effect remains significant thereafter. We interpret this Kendall's tau as a practical breakpoint, capturing the minimum strength of conditional dependence between treatment and outcome required for the welfare effect to turn significant.

figure[figure omitted — 211 chars of source]

Concluding Remarks

In this paper, we have extended the SI approach for identifying and estimating long-term treatment effects in athey2025surrogate from full mediation to partial mediation, substantially broadening the scope of application of the SI approach. Specifically, we develop two methodologies based on our identification result for a known copula: sensitivity analysis and partial identification analysis. The usefulness of both is illustrated via synthetic and real data. Complementing athey2025surrogate, we determine the sign of the surrogacy bias for stochastically monotone copulas and establish the worst-case bounds on the true ATE regardless of the type of the primary outcome. Our partial identification result applies to any copula bounds. When applied to copulas that dominate the independence copula in concordance order, the lower bound is the ATE under the surrogacy assumption in athey2025surrogate. This gives an alternative interpretation of ATE under the surrogacy assumption as the minimum ATE among all copulas that dominate the independence copula and thus are robust to such deviations from the surrogacy assumption.

Several extensions are worthwhile and are currently under investigation. First, in a companion article, the authors develop a sensitivity analysis to the comparability assumption. Second, the worst-case bounds are often wide, suggesting caution in making the surrogacy assumption. In addition to exploiting prior knowledge on the range of copulas to shrink the identified set, in specific applications, side information such as exclusion restrictions, monotone IV, monotone treatment response may be available. It is worthwhile exploring the possibility of tightening the worst-case bounds by exploring such information. Third, extensions to multi-valued treatments and continuous treatments would broaden the applicability of the SI approach further. Finally, on the technical side, it would be worthwhile exploring the possibility of extending the inference results to sub-copulas.