Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
220,544 characters · 37 sections · 68 citation commands
The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely
\thispagestyle{empty}
{Keywords:\ Potential Outcomes, Causality, Surrogate Outcomes, Surrogate Scpore, Surrogate Index, Mediators, Propensity Score, Principal Stratification, Job Training}
\baselineskip=22.5pt\setcounter{page}{1} \global\long\global\long
A fundamental challenge for evaluating interventions is that the primary outcomes of interest are often hard to measure. For example, researchers are often interested in the effect of the policy on some long-term outcome but do not observe that in their study. Instead, they observe a number of short-term outcomes that are all related to this primary outcome of interest.
One setting where this type of problem arises involves an educational policy maker evaluating a policy that would change class size. The ultimate goal may be to improve long-term labor market outcomes for the students. However, the decision regarding the class size policy needs to be made at a time when only short-term outcomes such as test scores or other educational achievement measures are available. Another setting involves policy makers considering labor market interventions such as job search assistance or human capital acquisition programs, where they may be primarily interested in the long-term labor market attachment of the participants, but in the short run they may only have access to outcomes such as employment records or earnings over a short period of time. In randomized experiments for medical interventions, the ultimate outcome of interest is often survival or quality-adjusted years of life. Survival rates may be high in the short run, and so typically such trials are evaluated in terms of surrogate measures, such as including tumor size or other measures of the progression of the disease, which can be measured earlier. In all of these types of setting, to make a timely decision, the policy maker needs to assess the programs based on short-term outcomes. These challenges also arise in business settings. In the context of experimentation in digital technology companies, a discussion of the most important challenges ranks as the top concern that “While most experiments in the industry run for 2 weeks or less, we are really interested in detecting the long-term effect of a change. How do long-term effects differ from short-term outcomes? How can we accurately measure those long- term factors without having to wait a long time in every case?” (gupta2019top, p. 21).
In these and many other examples, the researcher is faced with making recommendations regarding the future implementation of the intervention on the basis of measurements of its effect on a variety of sometimes disparate and possibly conflicting outcome measures. A key question is how to balance these different outcomes when making an overall assessment. In practice, researchers often deemphasize short-term outcomes for which they do not find statistically significant effects, instead making perhaps somewhat {\it ad hoc} qualitative assessments regarding the relative importance of the remaining short-term outcomes.
In this paper we lay out a framework for analyzing these issues. We consider the scenario in which researchers do not measure the primary outcome in the context of data containing information on the intervention. Instead, we assume that the researcher has a second, observational, dataset where the researcher observes the surrogates and the primary outcome but does not observe the treatment. In both samples the researcher may also observe variables not affected by the treatment, such as pre-treatment characteristics of the participants.
We make four main contributions. First, we articulate three key assumptions under which the average effect of the treatment on the primary outcome is identified from the combination of the experimental and observational samples: $(i)$ a standard assumption that the assignment in the experimental sample is Unconfounded; $(ii)$ a Surrogacy assumption which requires that the causal path from the treatment to the primary outcome goes through the surrogates prentice1989surrogate, day1996trial, begg2000, frangakis2002principal; and $(iii)$ a Comparability or external validity assumption, which requires that the observational and experimental samples are comparable in the sense that the outcome distributions conditional on surrogates and pre-treatment variables are identical.
Under these three assumptions, the average effect of the treatment on the primary outcome can be estimated as the average effect of the treatment on an aggregate of the surrogates, which we label the surrogate index. This index combines the individual surrogates through their predicted value of the primary outcome. For example, when studying the impact of class size, the primary outcome might be high school graduation, while the two surrogates might be mathematics and reading scores. For the special case of linear models, the proposal boils down to multiplying the causal effects of the intervention on the two scores (which can be estimated in the experimental sample) by the coefficients from a linear regression of the primary outcome on the two scores in the observational sample. The approach replaces a subjective assessment of the relative importance of the two short-term measures by an objective data-driven criterion, namely the predictive power of the scores for the outcome of interest.
In our second contribution, we derive the efficiency bound and propose various efficient estimators under various scenarios, including scenarios with a single sample or two samples, as well as with and without Surrogacy. This allows us to quantify the information content of the Surrogacy assumption.\footnote{We are grateful to Kevin Chen and David Ritzwoller for pointing out an error in one of our earlier efficiency bound calculations., see chen2023semiparametric for more details.}
In our third contribution, we provide bounds on the biases that arise in scenarios where either or both of Surrogacy or Comparability are violated. We show that even if these assumptions fail to hold (but unconfoundedness does hold), the proposed estimators still estimate a well-defined causal effect, by providing a principled way of combining short-term outcomes in a single measure through their predicted effect on the long-term outcome.
In our fourth contribution, we evaluate these methods in the context of a labor market program where we observe long-term (thirty-six quarters) outcomes in four locations. Following an approach popularized by lalonde, we put aside part of the data and investigate whether we could have estimated the long-term effects without having long-term experimental data. Specifically, we take one of the locations, Riverside, and put aside the long-term outcome for individuals from that location. Then we take the other three locations, Alameda, Los Angeles, and San Diego, and put aside the treatment assignment for that sample. We investigate whether these two samples allow us to recover the experimental long-term effects in Riverside using surrogates corresponding to the first $T$ quarters of outcomes (employment, earnings, and aid indicators). We find that combining six quarters of outcome data into a surrogate index suffices to obtain estimates close to the long run effects. Using the additional data that were put aside for the main analysis, we also directly test whether the critical assumptions, Surrogacy and Comparability, hold given various alternative sets of surrogates.
We recognize that the credibility of the Surrogacy assumption may be questioned in any given application, especially when viewed in isolation. Therefore, we view the best path forward as building a “library” of surrogate indices in which researchers systematically catalog across several studies the smallest set of surrogates that successfully match long-term outcomes of interest ({\it e.g.}, earnings, mortality, educational attainment). If one establishes, for instance, that six quarters of employment and earnings data are sufficient to predict the impacts of many different job training programs -- as our cross-site comparisons of the GAIN program suggest -- then the long-term impacts of future job training programs could be credibly estimated using the established six-quarter surrogate index. We view the empirical application in this paper as providing one element of such a library and hope future work will expand upon it by identifying surrogate indices that match estimated long-term impacts in other applications.
This study is related to three main bodies of literature, surrogacy, mediation, and missing data. We extend the literature on surrogacy (prentice1989surrogate, day1996trial, fleming1996surrogate,begg2000,xu2001evaluation,lauritzen2004discussion,d2006surrogate,qu2006quantifying,alonso2006unifying,gilbert2008evaluating, weir2006statistical) by formally including the presence of a second observational sample that is used to estimate the relationship between surrogates and the primary outcome and articulating the assumptions that justify doing so. In doing so we allow for uncertainty in the estimation of this surrogates/outcome relationship, whereas the previous literature took this relation as known. We also consider biases arising from violations of Surrogacy and Comparability.
In addition, this study builds on the literature on mediation (baron1986moderator,van2004estimation,imai2010general,zheng2012targeted,tchetgen2011semiparametric,vanderweele2015explanation), which considers the decomposition of an average treatment effect into the direct effect of a treatment on an outcome and indirect effects that flow through a mediator. In the mediation setup, all three key variables -- the outcome, the treatment, and the mediator -- are observed for the same units. The goal in the mediation literature is to determine the relative magnitudes of the direct and indirect effects. In our surrogacy analysis we focus on the case in which the direct effect is absent by assumption.
This paper is also related to the classical missing data literature in statistics (rubin1976inference, rubin2004multiple, little2014statistical). Our key assumptions are closely related to the Missing At Random (MAR) assumption. Our approach can be viewed as a special case of approaches that combine data sets, {\it e.g.}, ridder2007econometrics,chen2008semiparametric. In particular rassler2004data,rassler2012statistical refers to our setting, where one variable is missing in one part of a sample and a second variable missing in the remainder of the sample, as a “data fusion” setting. graham2016efficient discuss efficient estimation for a particular set of models defined by moment conditions in such a data fusion setting, where they allow the treatment to be a general random variable, rather than a binary indicator as in our setup.
The paper is organized as follows. Section (ref) sets up the problem and introduces the notation. Section (ref) discusses the critical assumptions and links the setup to the mediation and missing data literature. Section (ref) discusses identification and the efficiency bounds. Section (ref) presents formulas for bias when the surrogacy assumption fails and derives bounds on the degree of bias. Section (ref) discusses estimation. Section (ref) presents the empirical application. Section 8 concludes.
We define two samples, an Experimental (${\rm E}$) sample and an Observational (${\rm O}$) sample, with $N_{\rm E}$ and $N_{\rm O}$ units or individuals, respectively. It is convenient to view the data as consisting of a single sample of size $N=N_{\rm E}+N_{\rm O}$, with $P_{i}\in\{{\rm O},{\rm E}\}$ a binary indicator denoting the sample to which unit $i$ belongs.
For each unit, there is a binary treatment of interest, $W_{i}\in\{0,1\},$ and a scalar primary outcome, denoted by $Y_{i}$. This outcome is not observed for individuals in the experimental sample. In addition, there are intermediate or secondary outcomes, which we refer to as surrogates (to be defined precisely in Section (ref)), denoted by $S_{i}$ for each unit. Typically, the surrogate outcomes are vector-valued in order to make the properties we define plausible. Finally, we measure pre-treatment covariates $X_{i}$ for each unit, known not to be affected by the treatment.
Following the potential outcomes framework or Rubin Causal Model (rubin1974estimating,holland1986statistics,imbens2015causal), individuals in this group have two pairs of potential outcomes: $(Y_{i}(0),Y_{i}(1))$ and $(S_{i}(0),S_{i}(1))$. The realized outcomes are related to their respective potential outcomes as follows. \[ Y_{i}\equiv Y_{i}(W_{i})=\left\{
\right.\hskip1cm{\rm and}\ \ S_{i}\equiv S_{i}(W_{i})=\left\{
\right. \] Overall, the units are characterized by the values of the septuple $(Y_{i}(0),Y_{i}(1),S_{i}(0),S_{i}(1),X_{i},W_{i},P_{i})$. We do not observe the full septuple for any units. Rather, for units in the experimental sample we observe the triple $(X_{i},W_{i},S_{i})$ with support $(\mathbb{X}, \mathbb{W}, \mathbb{S})$ where $\mathbb{W}=\{0,1\}$. In the observational sample, we do not observe to which treatment each of the $N_{\rm O}$ individuals were assigned. We observe the triple $(X_{i},S_{i},Y_{i})$, with support $\mathbb{X}$, $\mathbb{S}$, and $\mathbb{Y}$ respectively. To simplify the exposition, we analyze the data as if we have a random sample from a population of units for which we observe the quintuple $(P_{i},X_{i},S_{i},\boldsymbol{1}_{P_{i}={\rm E}} W_{i},\boldsymbol{1}_{P_{i}={\rm O}} Y_{i})$, where we treat $P_{i}$ as a random variable taking on the values $\{{\rm O},{\rm E}\}$.
We summarize this data setup in Table (ref). The setup differs from those in athey2020combining and kallus2020role, where we would also observe the treatment in the observational sample, but in the experimental sample we would still not observe the primary outcome.
We are interested in the Average Treatment Effect (ATE) on the primary outcome in the population from which the experimental sample is drawn:
The same issues we study in the current paper apply to other estimands, such as the average treatment effect for the treated units, or the average for the observational sample.
An implicit assumption in our setup is that the two variables that are common to both samples, $S_{i}$ and $X_{i}$, measure the same underlying variables in both samples. In some cases it is possible that in one of the two samples, a coarser version is measured, for example age or education may be measured in multi-year categories rather than in years. In that case, a simple solution is to proceed by using the coarser version of the variables as corresponding to the surrogate of pre-treatment variable, thereby relying on strong assumptions. Another complication arises if the unit of observation differs in two samples, say individuals versus zipcodes. Again additional assumptions are required to link the variables between samples.
Table (ref) summarizes key definitions and notation.
In this section, we discuss the three key assumptions that together allow us to combine the observational and experimental samples and estimate the causal effect of the treatment on the primary outcome, exploiting the presence of the surrogates. The first assumption is {\it Unconfoundedness} or {\it Ignorability}, common in the program evaluation literature (rosenbaum1983central,imbens2015causal), which ensures that adjusting for pre-treatment variables leads to valid causal effects in the experimental sample. The second assumption is the {\it Surrogacy condition} due to prentice1989surrogate, that allows us to use the surrogate variables to proxy for the primary outcome. The third assumption is {\it Comparability}, which formalizes the connection between the two samples. This assumption is rarely stated formally, but plays an important role in our analysis.
For the individuals in the experimental group, the propensity score is the conditional probability of receiving the treatment: $\rho(x)\equiv {\rm pr}(W_{i}=1|X_{i}=x,P_{i}={\rm E}).$ We assume that for individuals in the experimental group, treatment assignment is unconfounded, and we have overlap in the distribution of pre-treatment variables between the treatment and control groups (rosenbaum1983central, imbens2015causal):
This assumption, widely used in the causal inference literature, implies that in the experimental sample, we can estimate the average causal effect of the treatment on the surrogates by adjusting for pre-treatment variables. We would also have been able to estimate the causal effect on the primary outcome had the primary outcome been measured in the experimental sample. In many applications of surrogacy approaches, the treatment in the experimental sample is assigned completely randomly. In that case this assumption is satisfied by design. However, unconfoundedness is all that is required.
Next we discuss the second critical assumption, surrogacy. We also introduce two concepts, the {\it surrogacy score}, similar to the propensity score, and the {\it surrogacy index}, to combine multiple surrogates.
prentice1989surrogate defines a surrogate as a post-treatment variable where conditioning on it makes the outcome and the treatment independent:
Surrogacy is often debated in empirical applications. freedman1992statistical argue that the surrogate may not mediate the full effect of the treatment in many settings. For example, reductions in class size may affect earnings through changes in non-cognitive skills that are not fully captured by standardized test scores (heckman2006effects,chetty2011does).
There are two scalar functions of the surrogates that play an important role in the analyses: the surrogate index and surrogate score.
The surrogacy score plays is similar to the role the propensity score plays in analyses under unconfoundedness rosenbaum1983central. Here if the surrogacy condition holds conditional on $(S_i,X_i)$, it also holds conditional on the surrogacy score.
All proofs are given in the Appendix.
One theme of this paper is that having multiple short-term variables can make a surrogacy approach more plausible, the same way multiple pre-treatment variables can make the unconfoundedness assumption more plausible. Here we discuss some illustrative examples.
The first example is illustrated in Figure 1.A. Suppose the treatment is an educational intervention. This treatment affects the outcome of interest, some labor market outcome, e.g., earnings, through a number of different channels corresponding to different skill sets. These channels may include mathematics skills, language skills, and social skills. Using only one of these variables as a surrogate would lead to biased estimates because they would ignore the other causal paths. In this case the set of three short-term variables collectively satisfy Unconfoundedness and Surrogacy.
The second case is illustrated in Figure 1.B. In this setup there is a variable, labeled “skills', that satisfies the critical assumptions for surrogacy. However, skills is not observed by the researcher. Instead we have two noisy measures of this surrogate, say both a written and an oral exam. Collectively these two variables may still not satisfy Surrogacy, since there may be impacts of skills on earnings not captured by the exams, but the bias from using both would be less than the bias from using only one candidate surrogate.
The third case is illustrated in Figure 1.C. Here there is a pathway from the treatment, an informative advertisement about an item, to the outcome, an indicator for the individual purchasing the advertised item, going through two variables that on their own could each serve as surrogates. These two variables are whether someone has interested in the item, and whether the individual engaged with the website where the item was sold. However, we only measure noisy versions of these surrogates. For the first surrogate we observe whether an individual clicked on the advertisement for the item, and for the second surrogate we observe the time spent on the website. Neither of these two observed variables is a valid surrogate, but the combination of the two generally removes more of the bias than a single one.
Figure 1.D illustrates further the concerns with only using a single surrogate, the possibility of focusing on treatments that improve the surrogate variable but not the primary outcome. Suppose that a researcher uses the single variable “click on ad” as a surrogate for the effect of the ad on purchases. If the observational sample was based on informative advertisements, there is likely a positive correlation between clicking on the advertisement and purchases. However, if the new treatment is uninformative, {\it e.g.,} clickbait advertisement, with no effect on the actual interest in the item, the surrogacy analysis using click behavior as the surrogate will be ineffective. Using both click behavior and time spent on the website as surrogates will likely reduce the bias. The same argument implies that using multiple tests as surrogates can reduce problems with “teaching to the test,” where the long-run impact of an intervention is not well captured by scores on a test.
Surrogacy and Unconfoundedness by themselves are not sufficient for consistent estimation of $\tau$ because they do not place restrictions on how the relationship between $Y_{i}$ and $S_{i}$ in the observational sample compares to that in the experimental sample. As far as we know, such restrictions were not previously articulated in the surrogacy literature because the setup is typically one with just the separate experimental sample. However, a comparability assumption is implicit in the way the postulated relationship between the surrogate and the primary outcome is used in that literature. Related assumptions about the possibility of using causal estimates in one location to predict causal effects in a second location on the basis of distributions of pre-treatment variables are discussed in hotz2005predicting and the literature on transportability, pearl2014external.
Let $\varphi\equiv {\rm pr}(P_{i}={\rm E})$ be the probability of a unit being part of the experimental sample. We introduce the Sampling Score, the propensity to be in the experimental sample:
The third key assumption we make is that the conditional distribution of $Y_{i}$ given $(S_{i},X_{i})$ in the observational sample is the same as the conditional distribution of $Y_{i}$ given $(S_{i},X_{i})$ in the experimental sample, and that the support of $(S_{i},X_{i})$ in the experimental sample is a subset of that in the observational sample. Formally,
Similar to Unconfoundedness and Surrogacy this is a strong assumption, but unlike those assumptions it is rarely discussed explicitly. As we show in Section (ref), by making it explicit we can discuss the biases arising from violations and improve the intuition when this assumption may be of concern. If the observational and experimental samples are substantially different in terms of the distribution of pre-treatment variables and surrogates, it would likely be more controversial to assume that conditional on those variables the outcome distributions are identical.
We let $\mu(s,w,x,p)$ denote the conditional expectation of the primary outcome given pre-treatment variables, surrogates, treatment, and sample:
Comparability and Surrogacy together allow us to impute the missing primary outcomes in the experimental sample, as shown by the following proposition.
Because we can estimate $\mu(s,x,{\rm O})=\mathbb{E}[Y_{i}|S_{i}=s,X_{i}=x,P_{i}={\rm O}]$, we can impute the missing $Y_{i}$ in the experimental sample as $\mu(S_{i},X_{i},{\rm O})$.
To provide context for the setup here and the key assumptions, it is useful to make a link to three related literatures, on mediation, instrumental variables, and missing data respectively. We describe the causal structures for surrogacy, mediation, and instrumental variables using a directed acyclical graph (DAG) pearl. The interpretations provided in this subsection are note essential to the main results in the next section.
The surrogacy, mediation, and instrumental variables literatures all study causal structures involving a causally linked sequence of three (sets) of variables. They differ in three key aspects: $(i)$ the assumptions they make on the causal structure, $(ii)$ the estimands that are the primary focus of the analysis, and $(iii)$ the data available for the analyses. The literatures also differ in the labels typically used for the three variables. In Table (ref) we list the labels, estimands, and some of the assumptions.
In Figures 2.A-2.C we show the differences in structures in DAG form in a single sample setting (so that we need not be concerned with the comparability assumption). Figure 2.A illustrates the surrogacy setup, with a causal link from the treatment to the surrogate and from the surrogate to the outcome. There is no unobserved confounder for the causal relation between treatment and surrogate, which would violate Assumption (ref) (Unconfounded Treatment Assignment / Strong Ignorability). There is no direct causal link from the treatment to the outcome. There are also no unobserved confounders for the causal relation between surrogate and the outcome. These two features of the DAG (no direct link between treatment and outcome and no unobserved confounder for the relation between surrogate and outcome imply Assumption ((ref)) (Surrogacy).
Figure 2.B shows a mediation example where Assumption (ref) is violated because there is a direct effect of the treatment on the outcome that does not pass through the surrogate. In this case $S_{i}$ is a typically labelled a mediator, rather than a surrogate. In the mediation case the direct effect of the treatment on the outcome is estimable because all three variables, treatment, mediator and outcome are observed in the same sample.
Figure 2.C shows a DAG representation of the standard instrumental variables (IV) model familiar to economists. The first difference from the surrogacy setup in Figure 2.A is that in the instrumental variables setting the interest is in the causal effect of the variable in the middle of the three variable chain (the surrogate $S$ in the surrogacy setting, and the treatment $W$ in the instrumental variables setting), on the outcome, whereas in the surrogacy setting the primary interest is in the effect of the first variable in the chain (the treatment $W$ in the surrogacy setting and the instrument $Z$ in the instrumental variables setting) on the outcome. In the instrumental variables case the surrogacy estimand is immediately identified as the intention-to-treat effect of the instrument, since the instrument and the surrogate are observed in the same sample. Under the assumptions of the surrogacy setup, the target for an instrumental variables analysis, the effect of the surrogate on the primary outcome, is immediately identified. The instrumental variables settings is characterized by the presence of an unobserved confounder that affects both the treatment of interest and the outcome. The presence of that unobserved confounder violates Surrogacy, even if the treatment has no direct effect on the long-term outcome (frangakis2002principal,rosenbaum1984consequences,joffe2009related,vanderweele2015explanation).
The presence of this unobserved confounder also violates the comparability assumption if the marginal distribution of the treatment $W$ differs between the observational and experimental samples, as will typically be the case. In both the surrogacy and the instrumental variables cases, we assume the absence of a direct effect of the first variable in the causal chain (the treatment $W$ in the surrogacy case and the instrument in the instrumental variables case) on the primary outcome. In the surrogacy setting, this assumption is part of the Surrogacy assumption, while in the instrumental variables setting this is typically referred to as the exclusion restriction angrist1996identification.
In the Online Appendix we also discuss a missing data interpretation of the surrogacy approach. Eessentially we show that the following joint conditional independence assumption,
implies both surrogacy and comparability.
This missing data characterization is useful because it allows one to use insights from the missing data literature, both for the current problems and for generalizations. Given ((ref)) we can use the conditional distribution of $Y_i$ given $(S_{i},X_{i})$ in the observational sample with $P_i={\rm O}$ to impute the missing outcomes in the experimental sample with $P_i={\rm E}$, and we can use the conditional distribution of $W_{i}$ given $(S_{i},X_{i})$ in the experimental sample wth $P_i={\rm E}$ to impute the missing treatments in the observational sample with $P_i={\rm O}$.
This observation directly extends to more general imputation problems. Suppose we have two samples where in one sample, indicated by $P_{i}={\rm E}$ we observe one set of variables, $(Z_{i1},Z_{i2})$ and in the second sample, indicated by $P_{i}={\rm O}$ we observe a partially overlapping set of variables, $(Z_{i2},Z_{i3})$. Then the analogous assumption that allows the imputation of all missing variables is $P_{i}\perp\!\!\!\perp Z_{i1}\perp\!\!\!\perp Z_{i3}|Z_{i2}.$
We now present our central identification result. We analyze three different representations of the average treatment effect that lead to three estimation strategies, somewhat similar to inverse propensity score weighting, regression, and influence function estimators for average treatment effects under unconfoundedness imbens2004. The motivation for developing the different representations is that estimators corresponding to those different representations can have different properties in finite samples, just like they do in the unconfoundedness setting. Estimators based on the first representation require estimation of the surrogate index, but not the surrogate score. Estimators based on the second representation instead require estimation of the surrogate score, but not the surrogate index. Estimators based on the third representation require estimation of both, but have attractive double robustness properties.
We define the following four objects, all functionals of distributions that are directly estimable from the data. First define the statistical estimand, the average difference in the surrogate index between treated and control, adjusted for pretreatment variablles, in the experimental sample:
\[ \left.\left. - \mathbb{E}\Bigl[ \mathbb{E}\left[\left. Y_{i} \right|S_{i},X_{i},P_{i}={\rm O}\right] \Bigr| W_{i}=0,X_{i},P_{i}={\rm E} \Bigr] \Bigr\}\right| P_{i}={\rm E}\right]. \] Next, with a surrogate index representation:
then a surrogate score representation,
\[ \hskip2cm\left.\left.-Y_{i}\cdot\frac{(1-\rho(S_{i},X_{i}))\cdot \varphi(S_{i},X_{i})\cdot(1-\varphi)}{(1-\rho(X_{i}))\cdot(1-\varphi(S_{i},X_{i}))\cdot \varphi}\right|P_{i}={\rm O}\right]. \] The third representation is based on the influence function. We first define \[ \mu(w,x)\equiv \mathbb{E}[\mu(S_{i},X_{i},{\rm O})|W_{i}=w,X_{i}=x,P_{i}={\rm E}].\] Then the influence function is
\[\hskip2cm+ \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,x) - \mu(0,x)-\tau \Bigr) \] \[\hskip2cm+\frac{\boldsymbol{1}_{p={\rm O}}}{\varphi} \frac{\varphi(s,x)}{1-\varphi(s,x)}\frac{(y-{\mu}(s,x,{\rm O}))\left({\rho}(s,x)-{\rho}(x)\right)}{{\rho}(x)(1-{\rho}(x))} \] with the estimand
In this subsection we present two pairs of semiparametric efficiency bound results bickel1993efficient,newey1990semiparametric for two different data configurations. The first directly refers to the main setup in this paper with the experimental and observational sample. This result is essentially shown in chen2023semiparametric which corrects a mistake in an earlier version of the current paper.
Next we consider the case where in a single sample we observe the treatment, primary outcome, surrogates and pre-treatment variables. In this single sample case we do not need the fifth variable, $P_i\in\{{\rm E},{\rm O}\}$. To maintain consistency with the other parts of the discussion and to avoid ambiguity, we keep the notation as before. In this case we can think of $P_{i}$ always taking the value $P_{i}={\rm E}$. We calculate the efficiency bound both without the assumption that surrogacy holds and with the assumption that surrogacy holds. We do so for a data generating process where surrogacy does hold, to see the information gain from that assumption.
The three critical assumptions, Unconfoundedness, Surrogacy, and Comparability, are strong. There is a large literature studying the sensitivity to unconfoundedness conditions rosenbaumrubin_sensitivity, imbens2003, cinelli2020making or bounds manski_bounds. Multiple studies have also raised concerns that in practice Surrogacy may not be satisfied begg2000, freedman1992statistical, frangakis2002principal,rosenbaum1984consequences,joffe2009related,vanderweele2015explanation, although we are not aware of formal sensitivity or bounds analyses. Violations of Comparability have not been explored because this assumption has not been previously formalized. In this section we examine the biases that arise from violations of Surrogacy and Comparability. We first characterize these biases and then derive estimable bounds on the magnitude of the biases that can arise from such violations.
We begin by characterizing the probability limit of estimators based on the representations of the estimand, $\tau^{{\rm E}}$, $\tau^{{\rm O}}$, and $\tau^{{\rm O},{\rm E}}$, in Theorem (ref) when the Surrogacy and Comparability assumptions are violated, as well as in cases where the surrogate index is misspecified. Throughout the section, we maintain Unconfoundedness in the experimental sample (Assumption (ref)), and the random sampling assumption (Assumption (ref)). We denote the probability limit of the estimators by $\underline{\tau}$ to differentiate it from the average treatment effect $\tau=\mathbb{E}[Y_i(1)-Y_i(0)|P_{i}={\rm E}].$
In this subsection we explore bounds on the parameter of interest. We show that in general these bounds are uninformative. However, if outcomes themselves are bounded, for example, if the outcomes are binary, informative bounds can be derived. Moreover, we present bounds given assumptions on the range of violations of the Surrogacy and Comparability assumptions.
In this section, we first present four estimators for the average treatment effect. The first, the surrogate index estimator, is related to previously proposed estimators with the difference that in the earlier literature the surrogate index was implicitly assumed to be known. We then discuss three new alternative estimators. The last of these new estimators is a matching estimator. Although matching estimators are generally not efficient in settings with unconfoundedness (rubin2006matched,abadie2006,abadie2016matching), they are widely applied, and it is instructive to see how a matching strategy can be used here.
Suppose we estimate the surrogate index as $\hat{\mu}(s,x,{\rm O})$ and the propensity score as $\hat{\rho}_{\rm E}(x)$. We take an average of the surrogate index in the experimental sample for the treatment and control groups, after adjusting for the propensity score. A natural estimator, corresponding to ((ref)), is the following difference of the two averages over the experimental sample:
\[ \hskip3cm-\frac{1}{\sum_{i=1}^{N_{{\rm E}}}(1-W_{i})/(1-\hat{\rho}(X_{i}))}\sum_{i=1}^{N_{{\rm E}}}\hat{\mu}(S_{i},X_{i},{\rm O})\cdot\frac{1-W_{i}}{1-\hat{\rho}(X_{i})}. \] We refer to this as the surrogate index estimator. Note that compared to the representation in Theorem (ref), we normalize the weights so that the weights sum up to one. This tends to improve the finite sample properties of related estimators in other settings substantially (hirano2003efficient,busso2014new).
In the case where the estimator for the surrogate index ${\mu}(s,x,{\rm O})$ was based on a linear specification for the regression of the primary outcome on the intermediate outcome, $\mu(s,x,{\rm O})=\gamma_{0}+\gamma_{S}'s+\gamma_{X}'x$, this leads to \[ \hat{\tau}^{{\rm E}}=\hat{\gamma}_{S}'\hat{\tau}_{S}, \] where $\hat{\tau}_{S}$ is an estimator for the average effect of the treatment on the surrogates, $\mathbb{E}\left[S_{i}(1)-S_{i}(0)\right].$ In the simplest case without pre-treatment variables and where the experimental sample is randomized, $\hat{\tau}_{S}=\overline{S}_{1}-\overline{S}_{0}$, where $\overline{S}_{1}$ and $\overline{S}_{0}$ are the average values of the surrogate outcomes. Here, the estimator simplifies to the difference in the estimated surrogate index in the treatment group and the control group: $ \hat{\tau}^{{\rm E}}=\hat{\gamma}_{S}'(\overline{S}_{1}-\overline{S}_{0})$. This expression is also familiar from the mediation literature (e.g., baron1986moderator) and the surrogacy literature day1996trial. However, we emphasize that in general, there may be interactions between the surrogates and pre-treatment variables, and in that case the linear specification need not be not adequate.
We now use the second representation for $\tau$ in the main theorem to derive an alternative estimator. Let $\hat{\rho}(x)$, $\hat{\rho}(s,x),\hat{\varphi}(s,x),\hat{\varphi}(x)$, and $\hat{\varphi}$, be estimators for $\rho(x)$, $\rho(s,x),\varphi(s,x),\varphi(x)$, and $\varphi$ respectively.
The surrogate score estimator is based on averaging the following expression over the observational sample:
where for $w=0,1$ the weights are
We can also base estimation on the efficient score given in ((ref)). Given estimators for the propensity score, the surrogate score, and the sampling score, we can estimate the average treatment effect as
\[+ \frac{\boldsymbol{1}_{P_{i}={\rm E}}}{\hat{\varphi}} \Biggl(\hat{\mu}(1,X_{i}) \left( 1 - \frac{W_{i}}{\hat{\rho}(X_{i})} \right) - \hat{\mu}(0,X_{i}) \left( 1 - \frac{1 - W_{i}}{1 - \hat{\rho}(X_{i})} \right) \Biggr) \] \[ \hskip2cm+\frac{\boldsymbol{1}_{P_{i}={\rm O}}}{1-\hat{\varphi}}\left(\frac{\hat{\varphi}(S_{i},X_{i})}{1-\hat{\varphi}(S_{i},X_{i})} \frac{1-\hat{\varphi}}{\hat{\varphi}}\right) \frac{(Y_{i}-\hat{\mu}(S_{i},X_{i},{\rm O}))\left(\hat{\rho}(S_{i},X_{i})-\hat{\rho}(X_{i})\right)}{\hat{\rho}(X_{i})(1-\hat{\rho}(X_{i}))} \Biggr\}. \] Based on the results in newey1994asymptotic, it follows that under standard conditions the two estimators above and the surrogate index estimator all reach the semi-parametric efficiency bound, and are first-order equivalent.
The recent literature on double robust estimation of average treatment effects under unconfoundedness chernozhukov2016double suggests that this estimator may have superior properties in small samples.
Consider unit $i$ in the experimental sample with $X_{i}=x$ and $S_{i}=s$, and suppose this is a treated unit with $W_{i}=1$. We need to find three matches for this unit. First, we need to find a unit with the opposite treatment in the same (experimental) sample. Specifically, we need to find the closest unit in the experimental sample, in terms of pre-treatment variables, among the units with $W_{i}=0$. Suppose this unit is unit $j$, with $W_{j}=0$, and the value of the pre-treatment variables for this unit are $X_{j}=x'$, and the surrogate outcomes are $S_{j}=s'$. As a result of the matching we should have $x\approx x'$, but potentially $s$ could be quite different from $s'$. Next, we need to find for each of the two units $i$ and $j$ a match in the observational sample. Find the unit in the observational sample closest to unit $i$, in terms of both pre-treatment variables and surrogates. Let $i'$ be the index for this unit, and let the value of the outcome for this unit be $Y_{i'}$, and the values of the pre-treatment variables and surrogates $X_{i'}$ and $S_{i'}$. Now as a result of the matching $X_{i}\approxX_{i'}$ and $S_{i}\approxS_{i'}$. Finally, find the unit in the observational sample closest to unit $j$, in terms of both pre-treatment variables and surrogates. Let the value of the outcome for this unit be $Y_{j'}$, and the values of the pre-treatment variables and surrogates $X_{j'}$ and $S_{j'}$, with $X_{j}\approxX_{j'}$ and $S_{j}\approxS_{j'}$.
Then we combine these matches to estimate the causal effect for unit $i$, $Y_{i}(1)-Y_{i}(0)$, as the difference in average outcomes for the two matches from the observational sample:
The matching estimator for $\tau$ would then be the average value of ((ref)) over the experimental sample. The double matching estimator is then \[ \hat\tau^{\rm match}=\frac{1}{N^{\rm E}}\sum_{i:P_{i}{\rm E}}\left\{ W_i \left( Y_{i'}-Y_{j'}\right) +(1-W_i) \left( Y_{j'}-Y_{i'}\right)\right\}. \]
In this section, we apply our method to estimate the causal effect of the Greater Avenues to Independence (GAIN) job training program on long-term labor market outcomes. GAIN was a job assistance program implemented in California in the 1980s to help welfare recipients find work (riccio1989gain,friedlander1995evaluating, hotz2006evaluating). MDRC conducted a randomized trial to evaluate the GAIN program's employment impacts in six counties in California in the late 1980s. We focus primarily on the GAIN trial in Riverside, which was widely heralded as the program that had the largest treatment effects on earnings. The Riverside program emphasized a “jobs first” approach to re-entry into the labor force, encouraging unemployed workers to take any job they find; in contrast, other sites focused more heavily on developing human capital through training programs (hotz2006evaluating).
We have available long-term outcomes for the four GAIN sites, including employment, earnings, and receipt of aid over the first thirty-six quarters after random assignment. We take the average of the thirty-six employment indicators and earnings in Riverside as our primary outcomes. We then investigate whether we could have predicted the long-term impact on these outcomes using only the first $T$ quarters of all outcomes (including employment, earnings, and aid) as surrogates, as well as using pre-treatment variables (characteristics of the individuals as well as lagged employment, earnings and aid outcomes). The Riverside data on the treatment, surrogates and pre-treatment variables play the role of of our experimental sample. We use the data from the combination of the other three locations (Alameda, Los Angeles, and San Diego) as our observational sample. For the observational sample we only use the information on the surrogates, pre-treatment variables, and outcome, but not the treatment assignment, nor the indicator for the location.
We begin by presenting a brief summary of the samples. We then describe how we construct our surrogate index. Next we illustrate our theoretical results by evaluating the magnitude of the gains from using surrogate indices in terms of time and precision relative to existing experimental estimates of the program's long-term impacts in Riverside. We also show how one can validate the surrogacy assumption using intermediate outcomes and bound the degree of bias arising from potential violations of surrogacy.
The GAIN treatment was randomly assigned to welfare (Aid for Families with Dependent Children) recipients, a very low-income population. The treatment group consisted of $N_{{\rm E},T}=4405$ participants, which the control group consisted of $N_{{\rm E},C}=1040$ participants who were not eligible for the additional services in the GAIN program. The data we use come from the hotz2006evaluating which followed study participants for nine years after assignment of the treatment, measuring quarterly employment rates and earnings\footnote{ All income variables were converted to 1999 dollars using cost-of-living deflators; see footnote 21 of hotz2000long for more information.} from the Unemployment Insurance database. They found that the treatment effects of the Riverside GAIN program on employment rates and earnings were initially large, but declined over time, as shown in Figure 3A, which plots employment rates by quarter for individuals in the experimental (Riverside) treatment and control groups, and in Figure 3B, which shows the correspond results for quarterly earnings.
In Riverside, the estimated causal effects on the primary outcomes were a 6.4 (s.e. = 1.2) percentage point (pp) increase in average quarterly employment rates, and an \$249 (s.e. \$84) increase in average quarterly earnings, in both cases averaged over the 36 quarter post-treatment. Our question is whether these impacts could have been estimated more quickly by using short-term employment, earnings and aid receipt as surrogates.
The observational sample includes the other three locations, Alameda, Los Angeles and San Diego, for a total of $N_{\rm O}= 13,725$ individuals.
In the online appendix Table (ref) presents information on the pre-treatment variables. Clearly the two samples, Riverside and the combination of the other three locations, are substantially different prior to the intervention in terms of permanent characteristics such as ethnicity, as well as in pre-treatment outcomes.
We discuss here the estimators for the average effect of the program We wish to consider different set of surrogates, indexed by the number of periods $t$ we want to use as surrogates. To capture this we index the surrogate for individual $i$, $S_i^t$, by the superscript $t$. $S_i^t$ contains the employment indicators, earnings outcomes and aid receipt indicators for the $t$ quarters after the intervention.
To construct the surrogacy index we estimate a linear regression model using least squares, for the individuals in the observational sample
The predicted value from this regression, which we denote by $\hat{Y}_{i}$, is our surrogate index for mean employment based on surrogates up to quarter $t$. We then compute this surrogate index for each of the individuals in the experimental sample and estimate the treatment effect based on the surrogate index as
If we use the all 36 quarters of employment indicators are surrogates, then the regression of $Y_i$ on the set of surrogates will fit perfectly, $\hat{Y}_i$ will be equal to $Y_i$, and the estimated effect will be identical to the original experimental estimate. The question is whether using a much more limited set of surrogates will get us close to the experimental benchmark.
For the surrogate score estimator we first estimate a logistic regression of the treatment indicator on the pretreatment variables and the surrogates. We specify \[ \ln\left( \frac{\rho(S_i^t,X_i)}{1-\rho(S_i^t,X_i)} \right)\equiv \ln\left( \frac{{\rm pr}(W_i=1|S_i^t,X_i,P_{i}={\rm E})}{1-{\rm pr}(W_i=1|S_i^t,X_i,P_{i}={\rm E})} \right)= \alpha_0+\alpha_{S}^\top S_{i}^t+\alpha_X^\top X_i,\] and estimate this on the experimental (Riverside) sample.
Next we estimate the propensity score, also as a logistic regression, \[ \ln\left( \frac{\rho(X_i)}{1-\rho(X_i)} \right)\equiv \ln\left( \frac{{\rm pr}(W_i=1|X_i,P_{i}={\rm E})}{1-{\rm pr}(W_i=1|X_i,P_{i}={\rm E})} \right)= \delta_0+\delta_X^\top X_i,\] and estimate this again on the experimental (Riverside) sample. In principle the random assignment implies that the $\delta_X$ should be close to zero in this case.
Finally we estimate the comparability score \[ \ln\left( \frac{\varphi(S_i^t,X_i)}{1-\varphi(S_i^t,X_i)} \right)\equiv \ln\left( \frac{{\rm pr}(P_{i}={\rm E}|X_i,S_i^t)}{1-{\rm pr}(P_{i}={\rm E}|X_i,S_i^t)} \right)= \gamma_0+\gamma_{S}^\top S_{i}^t+\gamma_X^\top X_i,\] and estimate this on the combined observational and experimental samples.
The surrogate score estimator is based on averaging the following expression over the observational sample:
where the weights are as before in Equation ((ref)).
For the influence function estimator we first estimate the surrogacy index, the surrogacy score, the propensity score, and the comparability score as before. We then plug those into the estimator in Equation ((ref)).
Here we discuss two sets of results. First the estimates for the average effect of the intervention on the two primary outcomes under various assumptions about the surrogates. Second, we test the Surrogacy and Comparability assumptions directly.
As the discussion after the surrogate index estimator shows, we recover the experimental estimates if we use all 36 quarters of employment indicators as surrogates. The question is whether we can do approximately as well with fewer than 36 quarters of surrogates. In Figures 4A and 4B we compare the experimental estimates of the effect on the primary outcomes (0.064 for the employment outcome, and \$249 for the earnings outcome) to the three sets of surrogate estimates, as a function of how many periods of surrogates we use, ranging from 1 quarter to 36 quarters. To put this in perspective we also include in these two figures what we label the “naive” estimator where we estimate the effect on the long-term outcome as the effect on the first $t$ quarters of the outcome. In Tables (ref) and (ref) we report a subset of the numbers underlying these estimates with the corresponding standard errors.
We see that the naive estimator does very poorly. It takes more than 25 quarters before the naive estimator is within two standard errors of the experimental estimate. In contrast all three surrogate-based estimators are all within two standard errors when the surrogates include 5 quarters of outcomes, for both outcomes.
Given the data available we can also test whether using $t$ quarters of surrogates is sufficient to satisfy Surrogacy and Comparability. To test Surrogacy we regress the primary outcome on the pre-treatment variables, the surrogates up to quarter $t$, and the indicator for the treatment; a finding that the treatment has an impact indicates a violation of Surrogacy. We estimate this regression using a logistic regression model, using only the data from the experimental (Riverside) sample. We report in Table (ref) and (ref) the results from these regressions for a number of different values for $t$, for the employment outcome and the earnings outcome. We report the point estimate, standard error and t-statistic. We see that point estimates for $t\leq 3$ are large and highly statistically significant. After that most of the t-statistics are less than 2, although there are some where the t-statistics are a little above 2, but the coefficient estimates are small.
We do a similar exercise for Comparability. We combine the experimental and observational samples and regress the final outcome on the surrogates, the pretreatment variables, and an indicator for the experimental sample, again using surrogates up to period $t$. We report the estimates on the indicator for the experimental sample, and the corresponding standard error. Here the the point estimates become smaller after $t=12$, but the t-statistics remain large even with a substantial number of surrogate periods, indicating a violation of Surrogacy.
If we are unwilling to make the surrogacy assumption we can still calculate bounds for the effect on employment, using the fact that this outcome is binary. For the case where the first six quarters of post-treatment data are used as surrogates, the lower and upper bound are estimated as -0.186 and 0.124. These are not very informative, because the data now do not allow us to estimate the indirect effect of the treatment on the outcome.
Similarly, we can calculate bounds for the average effect without assuming Comparability. With six quarters of surrogates the bounds are again wide at -0.076 and 0.194 respectively. Here the fact that the treatment effect on the surrogates is strong leads to substantial sensitivity to the comparability assumption as formalized in Lemma (ref).
Using the data for Riverside we can also assess the value of the Surrogacy assumption. Using the six quarters of data as surrogates, we find that the gain from knowledge of Surrogacy (the $\Delta$ in Theorem (ref)) is quite large. The standard error given Surrogacy, $\sqrt{\mathbb{V}}_{\rm s}$, is 0.33 times the standard error without knowledge that Surrogacy holds, $\sqrt{\mathbb{V}}_{\rm ns}$.
We develop new methods for combining intermediate outcomes to estimate the long-term impacts of treatments more rapidly and precisely. Our method requires estimating a “surrogate index” -- the conditional expectation of the long-term outcome given intermediate outcomes -- and then estimating the treatment effect on the surrogate index. The surrogate index can be estimated using parametric or nonparametric regression methods. We formalize conditions under which this method yields unbiased estimates, derive bounds for the degree of bias when those assumptions fail, and propose a simple out-of-sample validation approach using “hold out” intermediate outcomes. We show that surrogates can also greatly improve the precision of estimates even in settings where the treatment effect on the long-term outcome can be estimated directly, particularly when that outcome is rare or noisy.
Applying the method to analyze the impacts of the GAIN job training program in California, we find that using short-term earnings and employment rates to construct surrogate indices expedite the detection of long-term treatment effects on employment and earnings by several years and also substantially increases precision. Furthermore, a single surrogate index accurately predicts heterogeneity in the long-term treatment effects of different types of job training programs across sites, showing that surrogate indices estimated in a given setting may be generalizable to other settings. The success of the surrogate index in this application validates the use of short-term employment outcomes as surrogates for detecting longer-term impacts of job training programs, an empirical result that can be applied when analyzing ongoing programs.
Building on this application, it would be useful to systematically establish surrogate indices that match the long-term treatment effects estimated in other experiments and quasi-experiments. Over time, this would allow researchers to collectively build a public library of surrogate indices for long-term outcomes that could be used to expedite the analysis of future interventions.
\vskip0.5cm
\vskip0.5cm
{}
{ }
{A}.\ Additional Table
{B}.\ Related Literature
In the mediation literature (e.g., baron1986moderator,vanderweele2015explanation), the intermediate outcome that we refer to here as the surrogate $S_{i}$ is called a mediator. To emphasize its role as a causal variable in the mediation literature, we expand the notation and consider potential outcomes $Y_{i}(w,s)$ that are indexed by the treatment and the surrogate. (In terms of these potential outcomes the original potential outcomes defined in the previous section, $Y_{i}(w)$, indexed only by the treatment $W_{i}$, equals $Y_{i}(w)=Y_{i}(w,S_{i}(w))$, for $w\in\mathbb{W}$.) In the setting considered in the mediation literature, we observe the quadruple $(Y_{i},S_{i},W_{i},X_{i},P_{i})$ for all units in the sample and so there is not necessarily a distinction between the experimental sample and the observational sample. To capture that we focus in this section on the case where we only have the experimental sample, $P_{i}={\rm E}$, and where we observe the primary outcome $Y_{i}$ for this sample.
The focus of the mediation literature is on decomposing the causal effect of the treatment on the outcome into a direct effect that involves comparing potential outcomes where the surrogate remains fixed, and an indirect effect that passes through the mediator/surrogate. Three key estimands are the average total effect, \[ \tau^{{\rm total}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(1))-Y_{i}(0,S_{i}(0))\right], \] the average natural indirect effect, where we fix the treatment at $w=1$, but change the surrogate from $S_{i}(0)$ to $S_{i}(1)$, \[ \tau^{{\rm nie}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(1))-Y_{i}(1,S_{i}(0))\right], \] and the average natural direct effect, where we fix the surrogate at $S_{i}(0)$ and change the treatment from $W_{i}=0$ to $W_{i}=1$: \[ \tau^{{\rm nde}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(0))-Y_{i}(0,S_{i}(0))\right], \] with the latter two adding up to the first: $\tau^{{\rm total}}=\tau^{{\rm nie}}+\tau^{{\rm nde}}$.
These effects are identified in the mediation literature using assumptions similar to Assumptions (ref) and (ref). The first assumption in the mediation framework is a reformulation of the unconfoundedness assumption, Assumption (ref). It rules out the presence of unmeasured confounders between the treatment and the surrogate, and between the treatment and the outcome.
The second assumption typically made in the mediation literature is another unconfoundedness assumption that rules out the presence of unobserved confounders between the surrogate and the outcome, conditional on the treatment.
This assumption implies that comparisons of primary outcomes for units with different values for the surrogates but identical values for the treatment and pre-treatment variables can be given a causal interpretation.
To make the link to the surrogacy literature we need to add one key assumption that is not commonly made in the mediation literature. This assumption rules out any direct effect of the treatment on the outcome, allowing only for an indirect effect through the surrogate.
This assumption is similar to the exclusion restriction in instrumental variables settings, e.g., imbens1994,angrist1996identification. In combination with the previous assumption this implies that we can give comparisons in the primary outcome between units with different values for the surrogates but the same values for pre-treatment variables a causal interpretation, without knowing the treatment status.
The following proposition links the surrogacy and mediation assumptions.
This connection highlights that at the heart of the surrogacy assumption is a causal relation between the surrogate and the primary outcome that mediates the causal effect of the treatment on the outcome.
From a missing data perspective, Surrogacy and Comparability have parallels to the missingness at random (MAR) assumption common in the missing data literature (rubin1976inference,little2019statistical), and specifically the literature on combining samples with different sets of variables, (ridder2007econometrics,gelman1998not,rassler2004data, graham2016efficient). In particular rassler2012statistical focuses on a missing data structure closely related to ours.
In our two sample setting, we can think of the complete data as the quintuple $(Y_{i},S_{i},W_{i},X_{i},P_{i})$. Here, we view the sample as randomly drawn from a large population, so that we view $P_{i}$ as a stochastic missing data indicator. For the units in the sample we observe the incomplete data $(\boldsymbol{1}_{P_{i}={\rm O}}Y_{i},S_{i},X_{i},{\boldsymbol{1}}_{P_{i}={\rm E}}W_{i},P_{i})$, where for units with $P_{i}={\rm O}$ the treatment indicator $W_{i}$ is missing, and for units with $P_{i}={\rm E}$ the outcome $Y_{i}$ is missing. Now consider the following assumption.
This is slightly different from a standard MAR assumption in rubin1976inference where one would assume $P_{i}\perp\!\!\!\perp Y_{i}|S_{i},X_{i}$ and/or $P_{i}\perp\!\!\!\perp W_{i}|S_{i},X_{i}$. We need the stronger assumption to incorporate surrogacy, as the following proposition shows.
Note that even after we have dealt with the missing $Y_{i}$ and missing $W_{i}$ problems, we still have the missing potential outcomes, which is why we also need the unconfoundedness assumption.
C. Proofs
Proof of Proposition (ref): \[ {\rm pr}\left(W_{i}=1|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right)=\mathbb{E}\left[\left.W_{i}\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right] \] \[ \hskip1cm=\mathbb{E}\left[\left.\mathbb{E}\left[\left.W_{i}\right|Y_{i}=y,S_{i},X_{i},\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right]\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right] \] \[ \hskip1cm=\mathbb{E}\left[\left.\mathbb{E}\left[\left.W_{i}\right|Y_{i}=y,S_{i},X_{i},P_{i}={\rm E}\right]\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right] \] \[ \hskip1cm=\mathbb{E}\left[\left.\mathbb{E}\left[\left.W_{i}\right|S_{i},X_{i},P_{i}={\rm E}\right]\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right] \] \[ \hskip1cm=\mathbb{E}\left[\left.\rho(S_{i},X_{i})\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right]=\rho(S_{i},X_{i}), \] which proves the result. $\square$
\vskip0.5cm
Proof of Proposition (ref): Part $(i)$ follows directly from the definitions of $\mu(\cdot,{\rm E})$ and Assumption (ref). Part $(ii)$ follows directly from the definitions of $\mu(\cdot,{\rm E})$ and $\mu(\cdot,{\rm O})$ and Assumption (ref). Part $(iii)$ follows from parts $(i)$ and $(ii)$. $\square$
\vskip0.5cm
Proof of Proposition (ref): We wish to show that the three conditions
and
imply
Note that we leave out the conditioning in $P_{i}={\rm E}$ in the last two conditions because we are focused here on the one-sample case. Condition ((ref)) follows directly from ((ref)) because $Y_{i}(w)=Y_{i}(w,S_{i}(w))$.
Condition ((ref)) implies that we can write $Y_{i}(s)$ without ambiguity, and by ((ref)), we have $ W_{i}\ \perp\!\!\!\perp\ Y_{i}(s)\ \Bigr|\ X_{i}. $ By ((ref)) we have $S_{i}\ \perp\!\!\!\perp\ Y_{i}(s)\ \Bigr|\ X_{i},W_{i}. $ Combining these implies $ \Bigl(S_{i},W_{i}\Bigr)\ \perp\!\!\!\perp\ Y_{i}(s)\ \Bigr|\ X_{i}. $ This in turn implies $ W_{i}\ \perp\!\!\!\perp\ Y_{i}(s)\ \Bigr|\ S_{i},X_{i}, $ which in turn implies $ W_{i}\ \perp\!\!\!\perp\ Y_{i}(S_{i})\ \Bigr|\ S_{i},X_{i}. $ This is equivalent to the condition we set out to prove, $ W_{i}\ \perp\!\!\!\perp\ Y_{i}\ \Bigr|\ S_{i},X_{i}. $ $\square$
\vskip0.5cm
Proof of Proposition (ref): The first part of the Proposition is immediate. For the second part, note that we can identify from the data the distributions \[ f_{Y_i|S_i,X_i,P_i}(y|s,x,{\rm O}),\hskip1cmf_{W_i|S_i,X_i,P_i}(w|s,x,{\rm E}),\hskip1cm{\rm and}\ \ f_{P_i,S_i,X_i}(p,s,x), \] but no other distributions. That implies that the joint distribution of $(Y_i,S_i,W_i,X_i,P_i)$ implied by $ f_{Y_i|S_i,W_i,X_i,P_i}(y|s,w,x,p)=f_{Y_i|S_i,X_i,P_i}(y|s,x,{\rm O}), $ and $ f_{W_i|S_i,X_i,P_i}(w|s,x,{\rm O})=f_{W_i|S_i,X_i,P_i}(w|s,x,{\rm E}), $ for all $(y,s,s,w,x,p)$ is consistent with the data, and it also satisfies Assumption (ref). $\square$
\vskip0.5cm
Proof of Theorem (ref): We prove the case for $\mathbb{E}[Y_{i}(1)|P_{i}={\rm E}]$, specifically
The proof of $\mathbb{E}[Y_{i}(0)|P_{i}={\rm E}]$ is similar. The score function representation is immediate from these equalities. We note that equality (ref) uses Assumptions (ref)--(ref) and equalities (ref) and (ref) only use the overlap condition, Assumption (ref)$(ii)$.
Consider ((ref)). By Assumption (ref) (unconfoundedness), it follows that \[ \mathbb{E}[Y_{i}(1)|P_{i}={\rm E}]=\mathbb{E}\left[\left.Y_{i}\cdot\frac{W_{i}}{\rho(X_{i})}\right|P_{i}={\rm E}\right]. \] Using the law of iterated expectations, we can first condition on $S_{i}$ and $X_{i}$ to get \[ \mathbb{E}\left[\left.Y_{i}\cdot\frac{W_{i}}{\rho(X_{i})}\right|P_{i}={\rm E}\right]=\mathbb{E}\left[\left.\mathbb{E}\left[\left.Y_{i}\cdot\frac{W_{i}}{\rho(X_{i})}\right|S_{i},X_{i},P_{i}={\rm E}\right]\right|P_{i}={\rm E}\right]. \] By Assumption (ref) (surrogacy), we have \[ \mathbb{E}\left[\left.\mathbb{E}\left[\left.Y_{i}\cdot\frac{W_{i}}{\rho(X_{i})}\right|S_{i},X_{i},P_{i}={\rm E}\right]\right|P_{i}={\rm E}\right]=\mathbb{E}\left[\left.\mathbb{E}\left[Y_{i}|S_{i},X_{i},P_{i}={\rm E}\right]\cdot\frac{\mathbb{E}\left[W_{i}|S_{i},X_{i},P_{i}={\rm E}\right]}{\rho(X_{i})}\right|P_{i}={\rm E}\right] \] By Assumption (ref) (Comparability), $\mu(s,x,{\rm E})=\mu(s,x,{\rm O})$ so that this is equal to \[ \mathbb{E}\left[\left.\mu(S_{i},X_{i},{\rm O})\cdot\frac{\mathbb{E}\left[W_{i}|S_{i},X_{i},P_{i}={\rm E}\right]}{\rho(X_{i})}\right|P_{i}={\rm E}\right]=\mathbb{E}\left[\left.\mu(S_{i},X_{i},{\rm O})\cdot\frac{\rho(S_{i},(X_{i})}{\rhoX_{i})}\right|P_{i}={\rm E}\right] \] Undoing the law of iterated expectations gives us the desired equality.
Consider ((ref)). By the definition of $\varphi(s,x)$, we have \[ \frac{\varphi(s,x)}{(1-\varphi(s,x))}\cdot\frac{1-\varphi}{\varphi}=\frac{{\rm pr}\left(\left.S_{i}=s,X_{i}=x\right|P_{i}={\rm E}\right)}{{\rm pr}\left(\left.S_{i}=s,X_{i}=x\right|P_{i}={\rm O}\right)} \] where the common support condition assures $1-\varphi(s,x)$ is not zero. This leads to \[ \mathbb{E}\left[\left.Y_{i}\cdot\frac{\rho(S_{i},X_{i})\cdot t(S_{i},X_{i})\cdot(1-\varphi)}{\rho (X_{i})\cdot(1-t(S_{i},X_{i}))\cdot \varphi}\right|P_{i}={\rm O}\right]=\mathbb{E}\left[\left.Y_{i}\cdot\frac{\rho(S_{i},X_{i})}{\rho(X_{i})}\cdot\frac{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm E}\right)}{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm O}\right)}\right|P_{i}={\rm O}\right] \] Again, by the law of iterated expectations, conditioning on $S_{i}$ and $X_{i}$ leads to \[ \mathbb{E}\left[\left.Y_{i}\cdot\frac{\rho(S_{i},X_{i})}{\rho(X_{i})}\cdot\frac{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm E}\right)}{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm O}\right)}\right|P_{i}={\rm O}\right]=\mathbb{E}\left[\left.\mu(S_{i},X_{i},{\rm O})\frac{\rho(S_{i},X_{i})}{\rho(X_{i})}\cdot\frac{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm E}\right)}{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm O}\right)}\right|P_{i}={\rm O}\right] \] Using the definition of conditional expectations, we obtain
Consider ((ref)). By the law of iterated expectations conditional on $S_{i}$ and $X_{i}$, we obtain \[ \mathbb{E}\left[\left.Y_{i}\cdot\frac{\rho(S_{i},X_{i})\cdot \varphi(S_{i},X_{i})\cdot(1-\varphi)}{\rho(X_{i})\cdot(1-\varphi(S_{i},X_{i}))\cdot \varphi}\right|P_{i}={\rm O}\right]=\mathbb{E}\left[\left.\mu(S_{i},X_{i},{\rm O})\cdot\frac{\rho(S_{i},X_{i})\cdot \varphi(S_{i},X_{i})\cdot(1-\varphi)}{\rho(X_{i})\cdot(1-\varphi(S_{i},X_{i}))\cdot \varphi}\right|P_{i}={\rm O}\right] \] where the common support condition assures $1-\varphi(s,x)$ is not zero. Part $(iii)$ follows from Proposition (ref), which shows that Surrogacy and Comparability have no testable implications. Standard arguments then imply that unconfoundednes does not generate any testable implications. $\square$
\vskip0.5cm
Proof of Theorem (ref): For Part (i), we need to calculate the variance of the Efficient Influence Function (EIF) to obtain the efficiency bound. We provide the detailed calculation for completeness.\footnote{ While our influence function representation coincides with chen2023semiparametric, the variance calculation resulted in a slightly different expression.}
Given the EIF:
\[+\frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,x) - \mu(0,x) -\tau \Bigr) \] \[\hskip2cm+\frac{\boldsymbol{1}_{p={\rm O}}}{\varphi} \left(\frac{\varphi(s,x)}{1-\varphi(s,x)}\frac{(y-{\mu}(s,x,{\rm O}))\left({\rho}(s,x)-{\rho}(x)\right)}{{\rho}(x)(1-{\rho}(x))} \right)\]
\[ \mathbb{V} = \bigg[\psi(Y_i,S_i,W_i,X_i)^2 \bigg] \]
\[ =\mathbb{E} \bigg[ \left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi}\left(\frac{W_{i} \cdot (\mu(S_{i},X_{i},{\rm O})-\mu(1,x))}{{\rho}(X_{i})}-\frac{(1-W_{i})\cdot ({\mu}(S_{i},X_{i},{\rm O})- \mu(0,x) )}{1-{\rho}(X_{i})}\right) \right)^2 \] \[ + \left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,x) - \mu(0,x) -\tau \Bigr) \right)^2 \] \[ + \left( \frac{\boldsymbol{1}_{p={\rm O}}}{\varphi} \left(\frac{\varphi(S_{i},X_{i})}{1-\varphi(S_{i},X_{i})}\frac{(Y_{i}-{\mu}(S_{i},X_{i},{\rm O}))\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)}{{\rho}(X_{i})(1-{\rho}(X_{i}))} \right)\right)^2 \bigg] \]
Focusing on the first block \[\left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi}\left(\frac{w \cdot (\mu(s,x,{\rm O})-\mu(1,x))}{{\rho}(x)}-\frac{(1-w)\cdot ({\mu}(s,x,{\rm O})- \mu(0,x) )}{1-{\rho}(x)}\right) \right)^2, \] noting that $w(1-w)=0$ and hence the cross-term disappearing, we only have to take the expectation of \[ \left(\frac{\boldsymbol{1}_{p={\rm E}}}{\varphi}\frac{w \cdot (\mu(s,x,{\rm O})-\mu(1,x))}{{\rho}(x)}\right)^2 \qquad {\rm and }\quad \left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \frac{ (1-w)\cdot ({\mu}(s,x,{\rm O})- \mu(0,x) )}{1-{\rho}(x)}\right)^2 \] Note that \[ \mathbb{E} \bigg[ \left(\frac{\boldsymbol{1}_{p={\rm E}}}{\varphi}\frac{W_{i} \cdot (\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))}{{\rho}(X_{i})}\right)^2 \bigg] \] \[ = \mathbb{E} \bigg[ \frac{(\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2}{ {\rho}(X_{i})^2\varphi^2} \mathbb{E} \big[ \boldsymbol{1}_{p={\rm E}} W_{i} | S_{i},X_{i} \big] \bigg]\quad (\because \text{Tower Property} ) \] \[ = \mathbb{E} \bigg[ \frac{(\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2}{ {\rho}(X_{i})^2\varphi^2} \varphi(S_{i},X_{i}) \mathbb{E} \big[ W_{i} | S_{i},X_{i} ,P_{i} = {\rm E} \big] \bigg] \] \[ = \mathbb{E} \bigg[ \frac{(\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2}{ {\rho}(X_{i})^2\varphi^2} \varphi(S_{i},X_{i}) \rho(S_{i},X_{i}) \bigg] \] \[ = \mathbb{E} \bigg[ \frac{\varphi(S_{i},X_{i}) \rho(S_{i},X_{i}) }{ \varphi^2{\rho}(X_{i})^2} (\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2\bigg] \]
Likewise, we can derive \[ \mathbb{E} \bigg[\left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \frac{ (1-W_{i})\cdot ({\mu}(S_{i},X_{i},{\rm O})- \mu(0,X_{i}) )}{1-{\rho}(X_{i})}\right)^2 \bigg] \] \[ = \mathbb{E} \bigg[ \frac{\varphi(S_{i},X_{i}) (1-\rho(S_{i},X_{i})) }{ \varphi^2(1-\rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm O})-\mu(0,X_{i}))^2\bigg] \] Collectivizing the two term yields the first block: \[ \mathbb{E} \bigg[ \frac{\varphi(S_{i},X_{i}) }{ \varphi^2} \bigg( \frac{1-\rho(S_{i},X_{i}) }{(1-\rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm O})-\mu(0,X_{i}))^2 + \frac{\rho(S_{i},X_{i}) }{{\rho}(X_{i})^2} (\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2 \bigg) \bigg] \]
Next, for the second block \[\frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,x) - \mu(0,x) -\tau \Bigr) \] we can likewise derive by using the Tower Property with respect to $X_{i}$ that \[ \mathbb{E} \bigg[ \left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr) \right)^2 \bigg] = \mathbb{E} \bigg[ \frac{\varphi(X_{i})}{\varphi^2} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \bigg] \] Finally, for the third block \[ \left( \frac{\boldsymbol{1}_{p={\rm O}}}{\varphi} \left(\frac{\varphi(s,x)}{1-\varphi(s,x)}\frac{(y-{\mu}(s,x,{\rm O}))\left({\rho}(s,x)-{\rho}(x)\right)}{{\rho}(x)(1-{\rho}(x))} \right)\right)^2, \] note that \[ \mathbb{E} \bigg[ \left( \frac{\boldsymbol{1}_{P_{i}={\rm O}}}{\varphi} \left(\frac{\varphi(S_{i},X_{i})}{1-\varphi(S_{i},X_{i})}\frac{(Y_{i}-{\mu}(S_{i},X_{i},{\rm O}))\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)}{{\rho}(X_{i})(1-{\rho}(X_{i}))} \right)\right)^2 \bigg] \] \[ =\mathbb{E} \bigg[ \frac{(\varphi(S_{i},X_{i}))^2}{\varphi^2(1-\varphi(S_{i},X_{i}))^2} \frac{\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)^2 }{{((\rho}(X_{i})(1-{\rho}(X_{i})))^2)^2} \mathbb{E} \bigg[ \boldsymbol{1}_{P_{i}={\rm O}} (Y_{i}-{\mu}(S_{i},X_{i},{\rm O}))^2 | S_{i},X_{i} \bigg] \bigg] \] \[ =\mathbb{E} \bigg[ \frac{(\varphi(S_{i},X_{i}))^2}{\varphi^2(1-\varphi(S_{i},X_{i}))^2} \frac{\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)^2 }{{(\rho}(X_{i})(1-{\rho}(X_{i})))^2} \mathbb{E} \bigg[ (1-\varphi(S_{i},X_{i})) (Y_{i}-{\mu}(S_{i},X_{i},{\rm O}))^2 | S_{i},X_{i} , P_{i} = {\rm O} \bigg] \bigg] \] \[ =\mathbb{E} \bigg[ \frac{(\varphi(S_{i},X_{i}))^2}{\varphi^2(1-\varphi(S_{i},X_{i}))^2} \frac{\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)^2 }{{(\rho}(X_{i})(1-{\rho}(X_{i})))^2} (1-\varphi(S_{i},X_{i})) \sigma^2(S_{i},X_{i},{\rm O}) \bigg] \] \[ =\mathbb{E} \bigg[ \frac{(\varphi(S_{i},X_{i}))^2}{\varphi^2(1-\varphi(S_{i},X_{i}))} \frac{\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)^2 }{{(\rho}(X_{i})(1-{\rho}(X_{i})))^2} \sigma^2(S_{i},X_{i},{\rm O}) \bigg] \] \[ =\mathbb{E} \bigg[\frac{1-\varphi(S_{i},X_{i})}{\varphi^2} \left( \left( \frac{\varphi(S_{i},X_{i})}{1 -\varphi(S_{i},X_{i})} \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \right)\bigg] \]
Hence, adding up the three blocks (in the order from the third to the first block) yield the desired efficiency bound: \[ \mathbb{V}=\mathbb{E}[\psi(Y_{i},S_{i},W_{i},X_{i},P_{i})^2]\] \[\qquad=\mathbb{E}\bigg[ \frac{1-\varphi(S_{i},X_{i})}{\varphi^2} \left( \left( \frac{\varphi(S_{i},X_{i})}{1 -\varphi(S_{i},X_{i})} \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \right) \] \[+\frac{\varphi(X_{i})}{\varphi^2} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left. + \frac{\varphi(S_{i},X_{i})}{\varphi^2} \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right].\] \[\qquad=\mathbb{E}\bigg[ \frac{1}{\varphi^2} \frac{\varphi(S_{i},X_{i})^2}{1 -\varphi(S_{i},X_{i})}\left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \] \[+\frac{\varphi(X_{i})}{\varphi^2} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left. + \frac{\varphi(S_{i},X_{i})}{\varphi^2} \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right].\]
For part $(ii)$, first rewrite the variance bound, normalized by the square root of the expected size of the experimental sample, $\varphi N$, instead of normalized by the total sample size $N$, as \[\tilde \mathbb{V}=\mathbb{E}\bigg[ \frac{1}{\varphi} \frac{\varphi(S_{i},X_{i})^2}{1 -\varphi(S_{i},X_{i})}\left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \] \[+\frac{\varphi(X_{i})}{\varphi} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left. + \frac{\varphi(S_{i},X_{i})}{\varphi} \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right].\]
Next, we re-write the bound in terms of a conditional expectation in the experimental sample, rather than as the unconditional expectation, (this implies multiplying by $\varphi/\varphi(S_i,X_i)$ or $\varphi/\varphi(X_i)$ appropriately) as
\[\tilde \mathbb{V}=\mathbb{E}\bigg[ \frac{\varphi(S_{i},X_{i})}{1 -\varphi(S_{i},X_{i})}\left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \] \[+ \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left.\left. + \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right|P_{i}={\rm E}\right].\]
Now we consider a sequence of data generating processes, where the outcome distribution in the observational sample remains fixed, and the propensity and surrogate scores remain fixed, and only the functions $\varphi(s,x)$, $\varphi(x)$ and the scalar $\varphi$ change, in such a way that $\sup_{s,x}\varphi(s,x)\rightarrow 0$. The the first term converges to zero, leaving us with \[\bar\mathbb{V}=\mathbb{E}\bigg[\Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left.\left. + \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right|P_{i}={\rm E}\right].\] The final step is to note that $\rho(S_{i},X_{i})=\mathbb{E}[W_{i}|S_{i},X_{i},P_{i}={\rm E}]$ so we can write $\bar\mathbb{V}$ as \[\bar\mathbb{V}=\mathbb{E}\bigg[\Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left.\left. + \left( \frac{ (1-\mathbb{E}[W_{i}|S_{i},X_{i},P_{i}={\rm E}])(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\mathbb{E}[W_{i}|S_{i},X_{i},P_{i}={\rm E}](\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right|P_{i}={\rm E}\right]\] \[=\mathbb{E}\bigg[\Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left.\left. + \left( \frac{ (1-W_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{W_{i}(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right|P_{i}={\rm E}\right].\]
$\square$
\vskip0.5cm
Proof of Theorem (ref):
The first representation of the efficiency bound without surrogacy in part $(i)$ of the Theorem is essentially rewriting the efficiency bound in hahn1998role, and related results in robins1995semiparametric,robins1995analysis. The standard version of the efficiency bound is \[ \mathbb{V}=\mathbb{E} \left[ \frac{\sigma^2(1,X_i)}{\rho(X_i)}+\frac{\sigma^2(0,X_i)}{1-\rho(X_i)}+\left(\mu(1,X_i)-\mu(0,X_i)-\tau\right)^2 \right].\] The proof consists of showing that this is equal to the expression for $\mathbb{V}_{{\rm ns}}$ in Theorem (ref): \[ \mathbb{V}_{{\rm ns}}=\mathbb{E}\biggl[\sigma^{2}(S_{i},X_{i},{\rm E})\cdot\left(\frac{\rho(S_{i},X_{i})}{\rho(X_{i})^{2}}+\frac{1-\rho(S_{i},X_{i})}{(1-\rho(X_{i}))^{2}}\right) \] \[ \hskip2cm+\frac{\rho(S_{i},X_{i})}{\rho(X_{i})^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(1,X_{i})\right)^{2}+\frac{1-\rho(S_{i},X_{i})}{(1-\rho(X_{i}))^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(0,X_{i})\right)^{2} \] \[\hskip2cm +\left(\mu(1,X_i)-\mu(0,X_i)-\tau\right)^2 \biggr].\] which amounts to showing the equality of
and
\[ \hskip2cm+\frac{\rho(S_{i},X_{i})}{\rho(X_{i})^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(1,X_{i})\right)^{2}+\frac{1-\rho(S_{i},X_{i})}{(1-\rho(X_{i}))^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(0,X_{i})\right)^{2}\biggr]. \]
By unconfoundedness \[ \sigma^2(1,x)\equiv \mathbb{V}(Y_i(1)|X_i=x)=\mathbb{V}(Y_i|W_i=1,X_i=x),\] where as mentioned in the main text, we implicitly condition on the sampling indicator and abstract it from the notation when it does not lead to confusion.
By iterated expectations this is equal to \[ \mathbb{E}\left[\left.\mathbb{V}\left(Y_i|W_i=1,S_i,X_i=x\right)\right|W_i=1,X_i=x\right]+\mathbb{V}\left(\left.\mathbb{E}[Y_i|W_i=1,S_i,X_i]\right|W_i=1,X_i\right).\] By surrogacy the conditional distribution of $Y_i$ given $W_i$, $S_i$ and $X_i$ does not vary by $W_i$, so this is equal to \[ \mathbb{E}\left[\left.\mathbb{V}\left(Y_i|S_i,X_i=x\right)\right|W_i=1,X_i=x\right]+\mathbb{V}\left(\left.\mathbb{E}[Y_i|S_i,X_i]\right|W_i=1,X_i\right)\] \[ \hskip1cm =\mathbb{E}\left[\left.\sigma^2(S_i,X_i)\right|W_i=1,X_i=x\right]+\mathbb{V}\left(\left.\mu(S_i,X_i)\right|W_i=1,X_i\right).\]
For the first term, \[ \mathbb{E}\left[\left.\sigma^2(S_i,X_i)\right|W_i=1,X_i=x\right]=\mathbb{E}\left[\left.\frac{\sigma^2(S_i,X_i)\rho(S_i,X_i)}{\rho(X_i)}\right|X_i=x\right].\] For the second term, note that \[ \mathbb{E}\left[\left.\mu(S_i,X_i)\right|W_i=1,X_i\right] = \mathbb{E}\left[\left.\mathbb{E}[Y_i|S_i,X_i]\right|W_i=1,X_i\right] \] is by surrogacy equal to $ \mathbb{E}\left[\left.\mathbb{E}[Y_i|W_i=1,S_i,X_i]\right|W_i=1,X_i\right] ,$ which in turn by iterated expectations is equal to $ \mathbb{E}\left[\left.Y_i\right|W_i=1,X_i\right] =\mu(1,X_i).$ Hence the second term is \[ \mathbb{V}\left(\left.\mu(S_i,X_i)\right|W_i=1,X_i\right) = \mathbb{E}\left[\left.\left(\mu(S_i,X_i)-\mu(1,X_i)\right)^2\right| W_i=1,X_i\right] \] \[\hskip1cm = \mathbb{E}\left[ \left(\mu(S_i,X_i)-\mu(1,X_i)\right)^2 \frac{\rho(S_i,X_i)}{\rho(X_i)} \right].\] Combining the two terms and including the denominator $\rho(X_i)$, we have \[\mathbb{E} \left[ \frac{\sigma^2(1,X_i)}{\rho(X_i)} \right]=\mathbb{E}\left[\frac{\sigma^2(S_i,X_i)\rho(S_i,X_i)}{\rho(X_i)^2}\right]+ \mathbb{E}\left[ \left(\mu(S_i,X_i)-\mu(1,X_i)\right)^2 \frac{\rho(S_i,X_i)}{\rho(X_i)^2} \right]. \] By the same argument \[\mathbb{E} \left[ \frac{\sigma^2(0,X_i)}{1-\rho(X_i)} \right]=\mathbb{E}\left[\frac{\sigma^2(S_i,X_i)(1-\rho(S_i,X_i))}{(1-\rho(X_i))^2}\right]+ \mathbb{E}\left[ \left(\mu(S_i,X_i)-\mu(0,X_i)\right)^2 \frac{1-\rho(S_i,X_i)}{(1-\rho(X_i))^2} \right] \] Hence, adding up the two equalities above shows the desired equivalence of ((ref)) and ((ref)). This finishes the proof of part $(i)$ of the theorem.
Next, for part $(ii)$ of the theorem, we derive the efficiency bound for the case with surrogacy by first deriving the efficient influence function and then deriving its variance. To derive the efficient influence function, we follow the proof in chen2023semiparametric and newey1990semiparametric, specifically the following four steps: (1) constructing the tangent space, (2) deriving the pathwise derivative of the target estimand (i.e. the ATE under surrogacy), (3) showing that the conjectured efficient influence function (EIF) lies in the tangent space, and (4) showing that the pathwise derivative of the target estimand and the conjectured EIF satisfies a key condition in newey1990semiparametric.
First, to characterize the tangent space, considering the data density where the functions $f$ denote the density of random variables. \[ f_{Y_i,S_i,W_i,X_i}(y,s,w,x) = f_{Y_i \mid S_i,X_i} (y\mid s,x) f_{S_i \mid W_i,X_i}(s \mid w,x) f_{W_i \mid X_i}(w \mid x) f_{X_i}(x). \] We assume the data density satisfies the regularity and smoothness conditions in Definition (A.1) of Newey (1990).
Let $G^\epsilon$ be a parametric submodel parameterized by $\epsilon \in [0,1]$ where $G^{\epsilon =0} = G$ and $G$ is the true data generating model. Let $f^{\epsilon}$ be the corresponding density function for the parametric submodel. Then, the score of $f_{\epsilon}$ is
{
} We use $Q(\cdot)$'s to denote the score function, i.e. $Q(\cdot) = \frac{\delta}{\delta \epsilon} \log(f^\epsilon(\cdot))$. Evaluating the derivative at $\epsilon=0$ leads us to the score of the true model, i.e., \[ Q_{Y_i,S_i,W_i,X_i}(y,s,w,x) = Q_{Y_i\mid S_i,X_i}(y \mid s,x) + Q_{S_i \mid W_i,X_i}(s \mid w,x) + Q_{W_i\mid X_i}(w \mid x) + Q_{X_i}(x). \] The tangent space $\mathcal{T}$ is the mean closure of a linear combination of mean-zero, square-integrable functions $\overline{Q}_1,...,\overline{Q}_4$ that satisfy the following conditions:
Second, we derive the pathwise derivative of our estimand. With some abuse of the integral notation, our estimand can be written as follows:
The pathwise derivative of the estimand $\tau$ is
{
} The derivatives above use the chain rule from calculus and the fact that \[ \frac{\delta}{\delta \epsilon} f^{\epsilon} = \frac{\delta}{\delta \epsilon} \log(f^{\epsilon}) f^{\epsilon} = Q^{\epsilon} f^{\epsilon} \]
Let $\tau'$ denote evaluating the above derivative at $\epsilon = 0$, i.e. {
} Third, consider the conjectured efficient influence function (EIF). { \[ \psi(Y_i,S_i,W_i,X_i) = \frac{(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))} + \frac{W_i(\mu(S_i,X_i) - \mu(1,X_i))}{\rho(X_i)} - \frac{(1-W_i)(\mu(S_i,X_i) - \mu(0,X_i))}{1-\rho(X_i)} + \mu(1,X_i) - \mu(0,X_i) - \tau \]} We show that $\psi(Y_i,S_i,W_i,X_i)$ is an element of the tangent space $\mathcal{T}$ by showing that different parts of $\psi(Y_i,S_i,W_i,X_i)$ satisfies conditions for $\overline{Q}_1, \overline{Q}_2$, and $\overline{Q}_4$.
By setting $\overline{Q}_3 = 0$, we arrive at $\psi(Y_i,S_i,W_i,X_i) \in \mathcal{T}$.
Fourth, we show that $\tau'$ and $\psi(Y_i,S_i,W_i,X_i)$ satisfy the following relationship that all efficient influence functions must satisfy from Theorem 2.2 in newey1990semiparametric:
We break the proof of this equality into several steps.
Combining the four steps (a)-(d) arrives at the desired equality between $\tau'$ and $\psi(Y_i,S_i,W_i,X_i)$.
Finally, note that the $\mathbb{V}_s$ is obtained by calculating the variance of the EIF (already written in Theorem 3):\footnote{We henceforth explicitly show the conditioning $P_{i}={\rm E}$ to be consistent with the notation in our Theorem statement.}
{ \[ \psi(Y_i,S_i,W_i,X_i,P_{i}) = \frac{(Y_i - \mu(S_i,X_i,{\rm E})) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))}+ \frac{W_i(\mu(S_i,X_i,{\rm E}) - \mu(1,X_i))}{\rho(X_i)} \]} { \[ - \frac{(1-W_i)(\mu(S_i,X_i,{\rm E}) - \mu(0,X_i))}{1-\rho(X_i)} + \mu(1,X_i) - \mu(0,X_i) - \tau, \] } i.e., \[ \mathbb{V}_{{\rm s}} = \bigg[\psi(Y_i,S_i,W_i,X_i,P_{i})^2 \bigg] =\mathbb{E} \bigg[ \left( \frac{(Y_i - \mu(S_i,X_i,{\rm E})) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))} \right)^2 + \left( \frac{W_i(\mu(S_i,X_i,{\rm E}) - \mu(1,X_i))}{\rho(X_i)} \right)^2 \] \[ + \left( \frac{(1-W_i)(\mu(S_i,X_i,{\rm E}) - \mu(0,X_i))}{1-\rho(X_i)} \right)^2 + \left( \mu(1,X_i) - \mu(0,X_i) - \tau \right)^2 \bigg] \]
\[ =\mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2+ \left( \mu(1,X_{i}) - \mu(0,X_{i}) - \tau \right)^2 \right.\] \[\left. + \frac{W_{i}}{\rho(X_{i})^2} (\mu(S_{i},X_{i},{\rm E}) -\mu(1,X_{i}))^2 + \frac{1-W_{i}}{ (1 - \rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm E}) - \mu(0,X_{i}))^2 \right] \] \[ =\mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2+ \left( \mu(1,X_{i}) - \mu(0,X_{i}) - \tau \right)^2 \right.\] \[\left. + \frac{\rho(S_{i},X_{i})}{\rho(X_{i})^2} (\mu(S_{i},X_{i},{\rm E}) -\mu(1,X_{i}))^2 + \frac{1-\rho(S_{i},X_{i})}{ (1 - \rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm E}) - \mu(0,X_{i}))^2 \right] \] by the law of iterated expectations, and hence we have that
\[ \Delta=\mathbb{V}_{{\rm ns}} - \mathbb{V}_{{\rm s}} = \mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \left( \frac{\rho\left(S_i, X_i\right)}{\rho\left(X_i\right)^2}+\frac{1-\rho\left(S_i, X_i\right)}{\left(1-\rho\left(X_i\right)\right)^2} - \left(\frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))}\right)^2 \right) \right] \] \[ = \mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \frac{\rho\left(S_i, X_i\right)\left(1-\rho\left(S_i, X_i\right)\right)}{\rho\left(X_i\right)^2\left(1-\rho\left(X_i\right)\right)^2} \right] \]
$\square$
Proof of Theorem (ref): Consider part (i). By the law of iterated expectations conditional on $S_{i}$ and $X_{i}$, we have
By the proof of (ref) in Theorem (ref) where we don't use Surrogacy or Comparability, we get
The second equality in $\tau^{{\rm E}}=\tau^{{\rm O}}=\tau^{{\rm E},{\rm O}}$ is immediate based on only the law of iterated expectations. Finally, by the law of iterated expectations conditional on $X_{i}$, we have
By Assumption (ref) (unconfoundedness), we have
Undoing the law of iterated expectations give the desired result.
For parts (ii)-(iv), we prove (iv) first. By Assumption (ref) (unconfoundedness), we have \[ \tau=\mathbb{E}\left[\mathbb{E}\left[Y_{i}|W_{i}=1,X_{i},P_{i}={\rm E}\right]\mid P_{i}={\rm E}\right]-\mathbb{E}\left[\mathbb{E}\left[Y_{i}|W_{i}=0,X_{i},P_{i}={\rm E}\right]\mid P_{i}={\rm E}\right]. \] By iterated expectations, this is equal to
Thus, we have
We add and subtract \[ \mathbb{E}\left[\mathbb{E}\left[\mu(S_{i},X_{i},{\rm E})|W_{i}=1,X_{i},P_{i}={\rm E}\right]\mid P_{i}={\rm E}\right]-\mathbb{E}\left[\mathbb{E}\left[\mu(S_{i},X_{i},{\rm E})|W_{i}=0,X_{i},P_{i}={\rm E}\right]\mid P_{i}={\rm E}\right] \] to get
Rearranging the terms, we have
Next, by the definition of expectations,
Use this to write ((ref)) as {
Using the same argument we can write ((ref)) as
Combining the results for ((ref)) and ((ref)) leads to \[ \mathbb{E}\left[\left(\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\right)\cdot\frac{(1-\rho(S_{i},X_{i}))\cdot \rho(S_{i},X_{i})}{(1-\rho(X_{i}))\cdot \rho(X_{i})}\mid P_{i}={\rm E}\right] \] Collecting the last two terms, ((ref)) and ((ref)), we have }{
}{ Combining the terms together, we obtain the expression in (iv)
Finally for part (ii), under Assumption (ref) (Comparability), but not Assumption (ref) (Surrogacy), $\mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})=0$ and the result is immediate from (iv). For part (iii), under Assumption (ref) (Surrogacy), but not Assumption (ref) (Comparability), $\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})=0$ and the result is immediate from (iv). $\square$}
{\bf Proof of Lemma (ref)} We can identify, given overlap, the surrogate score $\rho(s,x)$, the propensity score $\rho(X)$, the surrogate index $\mu(s,x,{\rm O})$, and the joint distribution of $(S_{i},X_{i},P_{i})$. This implies that to derive upper and lower bounds we just need to derive upper and lower bounds for the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ for each value of $(s,x)$ and then integrate these bounds. We will demonstrate the sharpness of these bounds by showing that there exist data distributions consistent with all assumptions such that these bounds are achieved.
Part $(i)$: By Theorem (ref) the surrogacy bias can be characterized as \[\textrm{\rm surrogacy-bias}= \mathbb{E}\left[\left.\Bigl\{\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]. \] The data are not directly informative about the two conditional expectation $\mu(s,w,x,{\rm E})$ (because we do not observe the outcome in the experimental sample) beyond their relation to the surrogacy index: \[ \mu(s,x,{\rm O})=\rho(s,x) \mu(s,1,x,{\rm E})+(1-\rho(s,x)) \mu(s,0,x,{\rm E}),\qquad \forall s,x.\] This implies the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ can be written as \[\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})=\frac{\mu(s,x,{\rm O})}{\rho(s,x)}-\frac{\mu(s,0,x,{\rm E})}{\rho(s,x)}.\] Fixing $\mu(s,x,{\rm O})$, $\rho(s,x)$, and $\mu(s,0,x,{\rm E})$ this places no restrictions on the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ and thus no restrictions on the bias, and therefore any value for the treatment effect on the whole real line is consistent with the data in the absence of surrogacy.
Part $(ii)$: If the outcome is binary, then some values can be ruled out. Because $\mu(s,w,x,{\rm E})$ is the conditional expectation of the outcome given some conditioning variables, it obviously must be inside the interval $[0,1]$, and both $\mu(s,1,x,{\rm E})$ and $\mu(s,0,x,{\rm E})$ must lie inside the interval $[0,1]$. This directly implies that $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})\in[-1,1]$. However, we can sharpen these bounds exploiting the fact that $\mu(s,x,{\rm O})=\rho(s,x)\mu(s,1,x,{\rm E})+(1-\rho(s,x))\mu(s,0,x,{\rm E})$. This implies that
First consider the upper bound on $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$. The question is what the pairs of values $(\mu(s,1,x,{\rm E}),\mu(s,0,x,{\rm E}))$ are that both lie inside $[0,1]$, such that $\mu(s,x,{\rm O})=\rho(s,x)\mu(s,1,x,{\rm E})+(1-\rho(s,x))\mu(s,0,x,{\rm E})$ for given $\mu(s,x,{\rm O})$ and $\rho(s,x)$, and that maximize the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$. There are two possibilities. Either $\mu(s,x,{\rm O})\geq \rho(s,x)$ or $\mu(s,x,{\rm O})< \rho(s,x)$.
If $\mu(s,x,{\rm O})\geq \rho(s,x)$, then the smallest value for $\mu(s,0,x,{\rm E})$ such that the value for $\mu(s,x,{\rm E})$ implied by ((ref)) is less than or equal to one is $\mu(s,0,x,{\rm E})=(\mu(s,x,{\rm O})-\rho(s,x))/(1-\rho(s,x))$. This value has to be less than one by the assumption that there is a pair of values $(\mu(s,0,x,{\rm E}),\mu(s,1,x,{\rm E}))$ that satisfies ((ref)). In this case upper bound for the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ is equal to $(1-\mu(s,x,{\rm O}))/(1-\rho(s,x))$. If $\mu(s,x,{\rm O})\leq \rho(s,x)$, then the largest value for $\mu(s,1,x,{\rm E})$ such that $\mu(s,0,x,{\rm E})$ is nonnegative is $\mu(s,x,{\rm O})/\rho(s,x)$. In that case the upper bound for the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ is equal to $\mu(s,x,{\rm O})/\rho(s,x)$.
In summary, to demonstrate sharpness, consider the following data distributions:
If \(\mu(s, x, {\rm O}) \geq \rho(s, x)\), set \(\mu(s, 0, x, {\rm E}) = \frac{\mu(s, x, {\rm O}) - \rho(s, x)}{1 - \rho(s, x)}\) and \(\mu(s, 1, x, {\rm E}) = 1\).
If \(\mu(s, x, {\rm O}) < \rho(s, x)\), set \(\mu(s, 0, x, {\rm E}) = 0\) and \(\mu(s, 1, x, {\rm E}) = \frac{\mu(s, x, {\rm O})}{\rho(s, x)}\)
In both cases, these distributions are admissible under our assumptions, and also achieve the bounds, demonstrating that the bounds are sharp.
Therefore, the sharp upper bound is \[ \Delta^U_S(s,x)= \left\{
\right.\]\[ =\min\left(\frac{\mu(s,x,{\rm O})}{\rho(s,x)},\frac{1-\mu(s,x,{\rm O})}{1-\rho(s,x)}\right). \] The proof for the lower bound follows the same argument.
Part $(iii)$: \[\textrm{\rm surrogacy-bias}= \mathbb{E}\left[\left.\Bigl\{\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] \] \[\leq \mathbb{E}\left[\left.\left|\Bigl\{\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\Bigr\} \right|\cdot\left|\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|\right|P_{i}={\rm E}\right] \] \[\leq c \cdot \mathbb{E}\left[\left.\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] \] The upper bound can be achieved by setting $\mu(s,0,x,{\rm E})=\mu(s,x,{\rm O})-c\cdot \rho(s,x)$ and $\mu(s,1,x,{\rm E})=\mu(s,0,x,{\rm E})+c,$ These distributions are admissible under our assumptions, and hence sharpness is obtained. We can likewise obtain the lower bound.
$\Box$
{\bf Proof of Lemma (ref)} We show that the derived bounds are sharp by demonstrating that there exist data distributions consistent without assumptions that achieve these bounds. $(i)$ In the absence of Comparability the data imply no restrictions on the values for $\mu(s,x,{\rm E})$, and so as long as there is some difference between $\rho(s,x)$ and $\rho(x)$ there is no bound on the bias. \\ $(ii)$ If the outcomes are binary, the only restrictions implied on $\mu(s,x,{\rm E})$ are that all values lie inside $[0,1]$. The upper bound comes from imputing 1 for $\mu(s,x,{\rm E})$ if $\rho(s,x)>\rho(x)$ and $0$ if $\rho(s,x)<\rho(x)$, a choice of distribution that is admissible. This directly implies the bounds on the bias. \\ $(iii)$ \[\textrm{\rm comparability-bias} =\mathbb{E}\left[\left.\Bigl\{ \mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]. \] Then \[\left|\mathbb{E}\left[\left.\Bigl\{ \mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]\right| \] \[\leq\mathbb{E}\left[\left.\left|\Bigl\{ \mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})\Bigr\}\right|\cdot\frac{\left|\rho(S_{i},X_{i})-\rho(X_{i})\right|}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] \] \[\leq c\cdot \mathbb{E}\left[\left.\frac{\left|\rho(S_{i},X_{i})-\rho(X_{i})\right|}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] .\] The upper bound can be attained by setting \[ \mu(s,x,{\rm E})= \left\{
\right. \] and similarly for the lower bound. $\Box$
C. Illustration of Bias Bounds Calculation
We will provide a simple illustration of how the theoretical bias bounds we calculated in Section (ref) look like in practice. We focus on the employment outcome to illustrate the surrogacy bias and comparability bias bounds in the binary case (Case (ii)).
Table (ref) and (ref) show the bounds on the treatment effects using the Influence Function Estimator under potential violations of Surrogacy and Comparability, respectively.\footnote{If we are interested in conducting inference on the partial identification bounds, we can take the approach illustrated in, e.g., ImbensManski2004,molinari2020microeconometrics.} This demonstrates that in the binary outcome of employment, the sign can still be credibly inferred under the latter half even under surrogacy violation. The comparability bias seems to be non-negligible, part of our design of choosing Riverside (experimental data) due to its unique "jobs first" approach, in contrast to the "human capital" approach used in LA, San Diego, and Alameda counties (observational data). Further work must be done to ensure cases where comparability bias is minimal. We can similarly compute non-binary outcomes like Earnings, with some plausible range of user-specified parameter $c$ (Case (iii) in Section (ref)).