EconBase
← Back to paper

Estimating Treatment Effects using Multiple Surrogates: The Role of the Surrogate Score and the Surrogate Index

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

220,544 characters · 37 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely

\thispagestyle{empty}

abstractA common challenge in estimating the impact of interventions ({\it e.g.}, job training programs, educational programs) is that many outcomes of interest ({\it e.g.}, lifetime earnings or other labor market outcomes) are observed with a long delay. In biomedical settings this is often addressed by using short-term outcomes as so-called “surrogates” for the outcome of interest, {\it e.g.}, tumor size as a surrogate for mortality in cancer studies. We build on this literature by combining multiple, possibly qualitatively distinct, short-term outcomes ({\it e.g.}, short-run earnings and employment indicators) systematically into a “surrogate index.” Under the Prentice surrogacy assumption, which requires that the primary outcome is independent of the treatment conditional on the surrogates, we show that the average treatment effect on the surrogate index equals the treatment effect on the long-term outcome. We also relate the surrogacy assumption to a set of structural, causal assumptions. We then characterize the bias that arises from violations of each of the key assumptions, and we provide simple methods to validate these assumptions using additional observed outcomes. We apply our method to analyze the long-term impacts of a multi-site job training experiment in California. Rather than waiting a full nine years to directly observe the long-term impact, we show that it is possible to use short-term (the first six quarters) outcomes as surrogates. One could have estimated the program's long-term impacts on mean employment rates using the employment rates observed in the first six quarters, with a 35% reduction in standard errors.

{Keywords:\ Potential Outcomes, Causality, Surrogate Outcomes, Surrogate Scpore, Surrogate Index, Mediators, Propensity Score, Principal Stratification, Job Training}

\baselineskip=22.5pt\setcounter{page}{1} \global\long\global\long

Introduction

A fundamental challenge for evaluating interventions is that the primary outcomes of interest are often hard to measure. For example, researchers are often interested in the effect of the policy on some long-term outcome but do not observe that in their study. Instead, they observe a number of short-term outcomes that are all related to this primary outcome of interest.

One setting where this type of problem arises involves an educational policy maker evaluating a policy that would change class size. The ultimate goal may be to improve long-term labor market outcomes for the students. However, the decision regarding the class size policy needs to be made at a time when only short-term outcomes such as test scores or other educational achievement measures are available. Another setting involves policy makers considering labor market interventions such as job search assistance or human capital acquisition programs, where they may be primarily interested in the long-term labor market attachment of the participants, but in the short run they may only have access to outcomes such as employment records or earnings over a short period of time. In randomized experiments for medical interventions, the ultimate outcome of interest is often survival or quality-adjusted years of life. Survival rates may be high in the short run, and so typically such trials are evaluated in terms of surrogate measures, such as including tumor size or other measures of the progression of the disease, which can be measured earlier. In all of these types of setting, to make a timely decision, the policy maker needs to assess the programs based on short-term outcomes. These challenges also arise in business settings. In the context of experimentation in digital technology companies, a discussion of the most important challenges ranks as the top concern that “While most experiments in the industry run for 2 weeks or less, we are really interested in detecting the long-term effect of a change. How do long-term effects differ from short-term outcomes? How can we accurately measure those long- term factors without having to wait a long time in every case?” (gupta2019top, p. 21).

In these and many other examples, the researcher is faced with making recommendations regarding the future implementation of the intervention on the basis of measurements of its effect on a variety of sometimes disparate and possibly conflicting outcome measures. A key question is how to balance these different outcomes when making an overall assessment. In practice, researchers often deemphasize short-term outcomes for which they do not find statistically significant effects, instead making perhaps somewhat {\it ad hoc} qualitative assessments regarding the relative importance of the remaining short-term outcomes.

In this paper we lay out a framework for analyzing these issues. We consider the scenario in which researchers do not measure the primary outcome in the context of data containing information on the intervention. Instead, we assume that the researcher has a second, observational, dataset where the researcher observes the surrogates and the primary outcome but does not observe the treatment. In both samples the researcher may also observe variables not affected by the treatment, such as pre-treatment characteristics of the participants.

We make four main contributions. First, we articulate three key assumptions under which the average effect of the treatment on the primary outcome is identified from the combination of the experimental and observational samples: $(i)$ a standard assumption that the assignment in the experimental sample is Unconfounded; $(ii)$ a Surrogacy assumption which requires that the causal path from the treatment to the primary outcome goes through the surrogates prentice1989surrogate, day1996trial, begg2000, frangakis2002principal; and $(iii)$ a Comparability or external validity assumption, which requires that the observational and experimental samples are comparable in the sense that the outcome distributions conditional on surrogates and pre-treatment variables are identical.

Under these three assumptions, the average effect of the treatment on the primary outcome can be estimated as the average effect of the treatment on an aggregate of the surrogates, which we label the surrogate index. This index combines the individual surrogates through their predicted value of the primary outcome. For example, when studying the impact of class size, the primary outcome might be high school graduation, while the two surrogates might be mathematics and reading scores. For the special case of linear models, the proposal boils down to multiplying the causal effects of the intervention on the two scores (which can be estimated in the experimental sample) by the coefficients from a linear regression of the primary outcome on the two scores in the observational sample. The approach replaces a subjective assessment of the relative importance of the two short-term measures by an objective data-driven criterion, namely the predictive power of the scores for the outcome of interest.

In our second contribution, we derive the efficiency bound and propose various efficient estimators under various scenarios, including scenarios with a single sample or two samples, as well as with and without Surrogacy. This allows us to quantify the information content of the Surrogacy assumption.\footnote{We are grateful to Kevin Chen and David Ritzwoller for pointing out an error in one of our earlier efficiency bound calculations., see chen2023semiparametric for more details.}

In our third contribution, we provide bounds on the biases that arise in scenarios where either or both of Surrogacy or Comparability are violated. We show that even if these assumptions fail to hold (but unconfoundedness does hold), the proposed estimators still estimate a well-defined causal effect, by providing a principled way of combining short-term outcomes in a single measure through their predicted effect on the long-term outcome.

In our fourth contribution, we evaluate these methods in the context of a labor market program where we observe long-term (thirty-six quarters) outcomes in four locations. Following an approach popularized by lalonde, we put aside part of the data and investigate whether we could have estimated the long-term effects without having long-term experimental data. Specifically, we take one of the locations, Riverside, and put aside the long-term outcome for individuals from that location. Then we take the other three locations, Alameda, Los Angeles, and San Diego, and put aside the treatment assignment for that sample. We investigate whether these two samples allow us to recover the experimental long-term effects in Riverside using surrogates corresponding to the first $T$ quarters of outcomes (employment, earnings, and aid indicators). We find that combining six quarters of outcome data into a surrogate index suffices to obtain estimates close to the long run effects. Using the additional data that were put aside for the main analysis, we also directly test whether the critical assumptions, Surrogacy and Comparability, hold given various alternative sets of surrogates.

We recognize that the credibility of the Surrogacy assumption may be questioned in any given application, especially when viewed in isolation. Therefore, we view the best path forward as building a “library” of surrogate indices in which researchers systematically catalog across several studies the smallest set of surrogates that successfully match long-term outcomes of interest ({\it e.g.}, earnings, mortality, educational attainment). If one establishes, for instance, that six quarters of employment and earnings data are sufficient to predict the impacts of many different job training programs -- as our cross-site comparisons of the GAIN program suggest -- then the long-term impacts of future job training programs could be credibly estimated using the established six-quarter surrogate index. We view the empirical application in this paper as providing one element of such a library and hope future work will expand upon it by identifying surrogate indices that match estimated long-term impacts in other applications.

This study is related to three main bodies of literature, surrogacy, mediation, and missing data. We extend the literature on surrogacy (prentice1989surrogate, day1996trial, fleming1996surrogate,begg2000,xu2001evaluation,lauritzen2004discussion,d2006surrogate,qu2006quantifying,alonso2006unifying,gilbert2008evaluating, weir2006statistical) by formally including the presence of a second observational sample that is used to estimate the relationship between surrogates and the primary outcome and articulating the assumptions that justify doing so. In doing so we allow for uncertainty in the estimation of this surrogates/outcome relationship, whereas the previous literature took this relation as known. We also consider biases arising from violations of Surrogacy and Comparability.

In addition, this study builds on the literature on mediation (baron1986moderator,van2004estimation,imai2010general,zheng2012targeted,tchetgen2011semiparametric,vanderweele2015explanation), which considers the decomposition of an average treatment effect into the direct effect of a treatment on an outcome and indirect effects that flow through a mediator. In the mediation setup, all three key variables -- the outcome, the treatment, and the mediator -- are observed for the same units. The goal in the mediation literature is to determine the relative magnitudes of the direct and indirect effects. In our surrogacy analysis we focus on the case in which the direct effect is absent by assumption.

This paper is also related to the classical missing data literature in statistics (rubin1976inference, rubin2004multiple, little2014statistical). Our key assumptions are closely related to the Missing At Random (MAR) assumption. Our approach can be viewed as a special case of approaches that combine data sets, {\it e.g.}, ridder2007econometrics,chen2008semiparametric. In particular rassler2004data,rassler2012statistical refers to our setting, where one variable is missing in one part of a sample and a second variable missing in the remainder of the sample, as a “data fusion” setting. graham2016efficient discuss efficient estimation for a particular set of models defined by moment conditions in such a data fusion setting, where they allow the treatment to be a general random variable, rather than a binary indicator as in our setup.

The paper is organized as follows. Section (ref) sets up the problem and introduces the notation. Section (ref) discusses the critical assumptions and links the setup to the mediation and missing data literature. Section (ref) discusses identification and the efficiency bounds. Section (ref) presents formulas for bias when the surrogacy assumption fails and derives bounds on the degree of bias. Section (ref) discusses estimation. Section (ref) presents the empirical application. Section 8 concludes.

Setup and Notation

We define two samples, an Experimental (${\rm E}$) sample and an Observational (${\rm O}$) sample, with $N_{\rm E}$ and $N_{\rm O}$ units or individuals, respectively. It is convenient to view the data as consisting of a single sample of size $N=N_{\rm E}+N_{\rm O}$, with $P_{i}\in\{{\rm O},{\rm E}\}$ a binary indicator denoting the sample to which unit $i$ belongs.

For each unit, there is a binary treatment of interest, $W_{i}\in\{0,1\},$ and a scalar primary outcome, denoted by $Y_{i}$. This outcome is not observed for individuals in the experimental sample. In addition, there are intermediate or secondary outcomes, which we refer to as surrogates (to be defined precisely in Section (ref)), denoted by $S_{i}$ for each unit. Typically, the surrogate outcomes are vector-valued in order to make the properties we define plausible. Finally, we measure pre-treatment covariates $X_{i}$ for each unit, known not to be affected by the treatment.

Following the potential outcomes framework or Rubin Causal Model (rubin1974estimating,holland1986statistics,imbens2015causal), individuals in this group have two pairs of potential outcomes: $(Y_{i}(0),Y_{i}(1))$ and $(S_{i}(0),S_{i}(1))$. The realized outcomes are related to their respective potential outcomes as follows. \[ Y_{i}\equiv Y_{i}(W_{i})=\left\{

array[array omitted — 96 chars of source]

\right.\hskip1cm{\rm and}\ \ S_{i}\equiv S_{i}(W_{i})=\left\{

array[array omitted — 96 chars of source]

\right. \] Overall, the units are characterized by the values of the septuple $(Y_{i}(0),Y_{i}(1),S_{i}(0),S_{i}(1),X_{i},W_{i},P_{i})$. We do not observe the full septuple for any units. Rather, for units in the experimental sample we observe the triple $(X_{i},W_{i},S_{i})$ with support $(\mathbb{X}, \mathbb{W}, \mathbb{S})$ where $\mathbb{W}=\{0,1\}$. In the observational sample, we do not observe to which treatment each of the $N_{\rm O}$ individuals were assigned. We observe the triple $(X_{i},S_{i},Y_{i})$, with support $\mathbb{X}$, $\mathbb{S}$, and $\mathbb{Y}$ respectively. To simplify the exposition, we analyze the data as if we have a random sample from a population of units for which we observe the quintuple $(P_{i},X_{i},S_{i},\boldsymbol{1}_{P_{i}={\rm E}} W_{i},\boldsymbol{1}_{P_{i}={\rm O}} Y_{i})$, where we treat $P_{i}$ as a random variable taking on the values $\{{\rm O},{\rm E}\}$.

assumptionWe have a single random sample of size $N$ drawn from the joint distribution of $(P_{i},X_{i},S_{i}, W_{i}, Y_{i})$, where we observe for each unit in the sample $(P_{i},X_{i},S_{i},\boldsymbol{1}_{P_{i}={\rm E}} W_{i},\boldsymbol{1}_{P_{i}={\rm O}} Y_{i})$.

We summarize this data setup in Table (ref). The setup differs from those in athey2020combining and kallus2020role, where we would also observe the treatment in the observational sample, but in the experimental sample we would still not observe the primary outcome.

table[table omitted — 737 chars of source]

We are interested in the Average Treatment Effect (ATE) on the primary outcome in the population from which the experimental sample is drawn:

equation[equation omitted — 81 chars of source]

The same issues we study in the current paper apply to other estimands, such as the average treatment effect for the treated units, or the average for the observational sample.

An implicit assumption in our setup is that the two variables that are common to both samples, $S_{i}$ and $X_{i}$, measure the same underlying variables in both samples. In some cases it is possible that in one of the two samples, a coarser version is measured, for example age or education may be measured in multi-year categories rather than in years. In that case, a simple solution is to proceed by using the coarser version of the variables as corresponding to the surrogate of pre-treatment variable, thereby relying on strong assumptions. Another complication arises if the unit of observation differs in two samples, say individuals versus zipcodes. Again additional assumptions are required to link the variables between samples.

Table (ref) summarizes key definitions and notation.

table[table omitted — 2,234 chars of source]

The Critical Assumptions: Unconfoundedness, Surrogacy, and Comparability

In this section, we discuss the three key assumptions that together allow us to combine the observational and experimental samples and estimate the causal effect of the treatment on the primary outcome, exploiting the presence of the surrogates. The first assumption is {\it Unconfoundedness} or {\it Ignorability}, common in the program evaluation literature (rosenbaum1983central,imbens2015causal), which ensures that adjusting for pre-treatment variables leads to valid causal effects in the experimental sample. The second assumption is the {\it Surrogacy condition} due to prentice1989surrogate, that allows us to use the surrogate variables to proxy for the primary outcome. The third assumption is {\it Comparability}, which formalizes the connection between the two samples. This assumption is rarely stated formally, but plays an important role in our analysis.

Unconfoundedness

For the individuals in the experimental group, the propensity score is the conditional probability of receiving the treatment: $\rho(x)\equiv {\rm pr}(W_{i}=1|X_{i}=x,P_{i}={\rm E}).$ We assume that for individuals in the experimental group, treatment assignment is unconfounded, and we have overlap in the distribution of pre-treatment variables between the treatment and control groups (rosenbaum1983central, imbens2015causal):

assumption(Unconfounded Treatment Assignment / Strong Ignorability) \\ $(i)$ \[W_{i}\ \perp\!\!\!\perp\ \Bigl(Y_{i}(0),Y_{i}(1),S_{i}(0),S_{i}(1)\Bigr)\ \Bigr|\ X_{i},P_{i}={\rm E},\] $(ii)$ $0<\rho(x)<1\ {\rm for\ all}\ x\in\mathbb{X}.$

This assumption, widely used in the causal inference literature, implies that in the experimental sample, we can estimate the average causal effect of the treatment on the surrogates by adjusting for pre-treatment variables. We would also have been able to estimate the causal effect on the primary outcome had the primary outcome been measured in the experimental sample. In many applications of surrogacy approaches, the treatment in the experimental sample is assigned completely randomly. In that case this assumption is satisfied by design. However, unconfoundedness is all that is required.

Surrogacy

Next we discuss the second critical assumption, surrogacy. We also introduce two concepts, the {\it surrogacy score}, similar to the propensity score, and the {\it surrogacy index}, to combine multiple surrogates.

The Prentice Criterion

prentice1989surrogate defines a surrogate as a post-treatment variable where conditioning on it makes the outcome and the treatment independent:

assumption(Surrogacy, Prentice Criterion)\\ $(i)$ \[ W_{i}\ \perp\!\!\!\perp\ Y_{i}\ \Bigr|\ S_{i},X_{i},P_{i}={\rm E}. \] and $(ii)$ $0<\rho(s,x)<1, {\rm for\ all}\ s\in\mathbb{S},x\in\mathbb{X},\ \textrm{and}\ 0<{\rm pr}(P_{i}={\rm E})<1.$
remarkIf the quadruple $(Y_{i},S_{i},W_{i},X_{i})$ were observed for all units, surrogacy would be a testable condition. With $(S_{i},W_{i},X_{i})$ observed for units in the experimental sample, and $(Y_{i},S_{i},X_{i})$ observed for units in the observational sample, this assumption has no testable implications.
remarkNote that Surrogacy is formulated in terms of the realized outcome and surrogate values. In contrast we formulated the ignorability condition (Assumption (ref)) in terms of the potential outcomes. This is partly to connect our discussion to the surrogacy literature prentice1989surrogate, day1996trial.

Surrogacy is often debated in empirical applications. freedman1992statistical argue that the surrogate may not mediate the full effect of the treatment in many settings. For example, reductions in class size may affect earnings through changes in non-cognitive skills that are not fully captured by standardized test scores (heckman2006effects,chetty2011does).

The Surrogacy Index and the Surrogacy Score

There are two scalar functions of the surrogates that play an important role in the analyses: the surrogate index and surrogate score.

definition(The Surrogate Index) The surrogate index is the conditional expectation of the primary outcome given the surrogate outcomes and the pre-treatment variables, conditional on the sample: \[ \mu(s,x,p)\equiv \mathbb{E}\left[\left.Y_{i}\right|S_{i}=s,X_{i}=x,P_{i}=p\right]. \]
remarkThe surrogate index in the observational sample, $\mu(s,x,{\rm O})$, is identified because we observe the triple $(Y_{i},S_{i},X_{i})$ in the observational sample.
definition(The Surrogate Score) The surrogate score is the conditional probability of having received the treatment given the value for the surrogate outcomes and the covariates in the experimental sample: \[ \rho(s,x)\equiv{\rm pr}(W_{i}=1|S_{i}=s,X_{i}=x,P_{i}=E). \]

The surrogacy score plays is similar to the role the propensity score plays in analyses under unconfoundedness rosenbaum1983central. Here if the surrogacy condition holds conditional on $(S_i,X_i)$, it also holds conditional on the surrogacy score.

prop(Surrogate Score) Suppose Surrogacy (Assumption (ref)) holds. Then: \[ W_{i}\ \perp\!\!\!\perp\ Y_{i}\ \Bigr|\ \rho(S_{i},X_{i}),P_{i}={\rm E}. \]

All proofs are given in the Appendix.

The Benefits of Multiple Surrogates

One theme of this paper is that having multiple short-term variables can make a surrogacy approach more plausible, the same way multiple pre-treatment variables can make the unconfoundedness assumption more plausible. Here we discuss some illustrative examples.

The first example is illustrated in Figure 1.A. Suppose the treatment is an educational intervention. This treatment affects the outcome of interest, some labor market outcome, e.g., earnings, through a number of different channels corresponding to different skill sets. These channels may include mathematics skills, language skills, and social skills. Using only one of these variables as a surrogate would lead to biased estimates because they would ignore the other causal paths. In this case the set of three short-term variables collectively satisfy Unconfoundedness and Surrogacy.

The second case is illustrated in Figure 1.B. In this setup there is a variable, labeled “skills', that satisfies the critical assumptions for surrogacy. However, skills is not observed by the researcher. Instead we have two noisy measures of this surrogate, say both a written and an oral exam. Collectively these two variables may still not satisfy Surrogacy, since there may be impacts of skills on earnings not captured by the exams, but the bias from using both would be less than the bias from using only one candidate surrogate.

figure[figure omitted — 4,034 chars of source]
comment\begin{figure}[H] \begin{subfigure}[b]{0.45\textwidth} \scriptsizeFigure 1.a Surrogacy Assumption Satisfied{\scriptsize} \begin{tikzpicture}[>=stealth, node distance=0.8cm] \node[observed, label=above left:{\scriptsize\(\rm Education\)}] (1) ; \node[observed, right=1.5cm of 1, label=above:{\scriptsize\(\rm Language\ Skills\)}] (2) ; \node[observed, above=of 2, label={ {\scriptsize\(\rm Math\ Skills\)}} ] (3) ; \node[observed, below=of 2, label=below: { {\scriptsize\(\rm Social\ Skills\)}} ] (4) ; \node[observed, right=1.5cm of 2, label=above right:{\scriptsize\(\rm Wage\)}] (5) ; \draw [->, notouch] (1.east) -- (2.west); \draw [->, notouch] (2.east) -- (5.west); \draw [->, notouch] (1.south east) -- (3.north west); \draw [->, notouch] (3.north east) -- (5.south west); \draw [->, notouch] (1.south) -- (4.north west); \draw [->, notouch] (4.north east) -- (5.south); \end{tikzpicture} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \scriptsizeFigure 1.b Multiple Surrogates{\scriptsize} \begin{tikzpicture}[>=stealth, node distance=0.8cm] \node[observed, label=above:{\scriptsize\(\rm Education\)}] (1) ; \node[unobserved, right=1.3cm of 1, label=above right:{\scriptsize\(\rm Skills\)}] (2) ; \node[observed, above=of 2, label=above: {\scriptsize {\(\rm Written\ Exam\)}} ] (3) ; \node[observed, below=of 2, label=below:{\scriptsize {\(\rm Oral\ Exam\)}} ] (4) ; \node[observed, right=1.3cm of 2, label=above right:{\scriptsize\(\rm Wage\)}] (5) ; \draw [->, notouch] (1.east) -- (2.west); \draw [->, notouch] (2.east) -- (5.west); \draw [->, notouch] (2.north) -- (3.south); \draw [->, notouch] (2.south) -- (4.north); \end{tikzpicture} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \scriptsizeFigure 1.c: Multiple Surrogates, Scenario 1{\scriptsize} \begin{tikzpicture}[>=stealth, node distance=0.8cm] \node[observed, label=above left:{\scriptsize\(\rm Informative\ Ad\)}] (1) ; \node[unobserved, right=2.3cm of 1, label=below:{\scriptsize\(\rm Interested\ in\ item\)}] (2) ; \node[unobserved, right=2.3cm of 2, label=above:{\scriptsize\(\rm Engaged\ with\ Website\)}] (3) ; \node[observed, above=of 2, label=above:{\scriptsize {\(\rm Click\ on\ Ad\)}} ] (4) ; \node[observed, below=of 3, label=below:{\scriptsize {\(\rm Spend\ Time\ on\ Website\)}} ] (5) ; \node[observed, right=2.3cm of 3, label=above right:{\scriptsize\(\rm Purchase\ Item\)}] (6) ; \draw [->, notouch] (1.east) -- (2.west); \draw [->, notouch] (2.east) -- (3.west); \draw [->, notouch] (3.east) -- (6.west); \draw [->, notouch] (2.south) -- (4.north); \draw [->, notouch] (3.north) -- (5.south); \end{tikzpicture} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \textsf{\textbf{\scriptsizeFigure 1.d: Multiple Surrogates, Scenario 2}}{\scriptsize} \begin{tikzpicture}[>=stealth, node distance=0.8cm] \node[observed, label=above left:{\scriptsize\(\rm Click\ Bait\ Ad\)}] (1) ; \node[unobserved, right=2.3cm of 1, label=below:{\scriptsize\(\rm Interested\ in\ item\)}] (2) ; \node[unobserved, right=2.3cm of 2, label=above:{\scriptsize\(\rm Engaged\ with\ Website\)}] (3) ; \node[observed, above=of 2, label=above:{ \scriptsize{\(\rm Click\ on\ Ad\)}} ] (4) ; \node[observed, below=of 3, label=below:{\scriptsize {\(\rm Spend\ Time\ on\ Website\)}} ] (5) ; \node[observed, right=2.3cm of 3, label=above right:{\scriptsize\(\rm Purchase\ Item\)}] (6) ; \draw [->, notouch] (1.south east) -- (4.north west); \draw [->, notouch] (2.east) -- (3.west); \draw [->, notouch] (3.east) -- (6.west); \draw [->, notouch] (2.south) -- (4.north); \draw [->, notouch] (3.north) -- (5.south); \end{tikzpicture} \end{subfigure} \caption{Compilation of Surrogacy Assumptions and Scenarios} \end{figure}

The third case is illustrated in Figure 1.C. Here there is a pathway from the treatment, an informative advertisement about an item, to the outcome, an indicator for the individual purchasing the advertised item, going through two variables that on their own could each serve as surrogates. These two variables are whether someone has interested in the item, and whether the individual engaged with the website where the item was sold. However, we only measure noisy versions of these surrogates. For the first surrogate we observe whether an individual clicked on the advertisement for the item, and for the second surrogate we observe the time spent on the website. Neither of these two observed variables is a valid surrogate, but the combination of the two generally removes more of the bias than a single one.

Figure 1.D illustrates further the concerns with only using a single surrogate, the possibility of focusing on treatments that improve the surrogate variable but not the primary outcome. Suppose that a researcher uses the single variable “click on ad” as a surrogate for the effect of the ad on purchases. If the observational sample was based on informative advertisements, there is likely a positive correlation between clicking on the advertisement and purchases. However, if the new treatment is uninformative, {\it e.g.,} clickbait advertisement, with no effect on the actual interest in the item, the surrogacy analysis using click behavior as the surrogate will be ineffective. Using both click behavior and time spent on the website as surrogates will likely reduce the bias. The same argument implies that using multiple tests as surrogates can reduce problems with “teaching to the test,” where the long-run impact of an intervention is not well captured by scores on a test.

Comparability

Surrogacy and Unconfoundedness by themselves are not sufficient for consistent estimation of $\tau$ because they do not place restrictions on how the relationship between $Y_{i}$ and $S_{i}$ in the observational sample compares to that in the experimental sample. As far as we know, such restrictions were not previously articulated in the surrogacy literature because the setup is typically one with just the separate experimental sample. However, a comparability assumption is implicit in the way the postulated relationship between the surrogate and the primary outcome is used in that literature. Related assumptions about the possibility of using causal estimates in one location to predict causal effects in a second location on the basis of distributions of pre-treatment variables are discussed in hotz2005predicting and the literature on transportability, pearl2014external.

The Comparability Assumption

Let $\varphi\equiv {\rm pr}(P_{i}={\rm E})$ be the probability of a unit being part of the experimental sample. We introduce the Sampling Score, the propensity to be in the experimental sample:

definition(Sampling Score) \\ The sampling score is $\varphi(s,x)\equiv {\rm pr}(P_{i}={\rm E}|S_{i}=s,X_{i}=x).$

The third key assumption we make is that the conditional distribution of $Y_{i}$ given $(S_{i},X_{i})$ in the observational sample is the same as the conditional distribution of $Y_{i}$ given $(S_{i},X_{i})$ in the experimental sample, and that the support of $(S_{i},X_{i})$ in the experimental sample is a subset of that in the observational sample. Formally,

assumption(Comparability of Samples) \\ $(i) \hspace{.25cm}P_{i}\ \perp\!\!\!\perp\ Y_{i}\ \Bigr|\ S_{i},X_{i},$ \\$(ii) \varphi(s,x)<1\ \ {\rm for\ all}\ s\in\mathbb{S}\ \ {\rm and\ }x\in\mathbb{X}.$

Similar to Unconfoundedness and Surrogacy this is a strong assumption, but unlike those assumptions it is rarely discussed explicitly. As we show in Section (ref), by making it explicit we can discuss the biases arising from violations and improve the intuition when this assumption may be of concern. If the observational and experimental samples are substantially different in terms of the distribution of pre-treatment variables and surrogates, it would likely be more controversial to assume that conditional on those variables the outcome distributions are identical.

The Surrogate Index and the Sampling Score

We let $\mu(s,w,x,p)$ denote the conditional expectation of the primary outcome given pre-treatment variables, surrogates, treatment, and sample:

equation[equation omitted — 119 chars of source]

Comparability and Surrogacy together allow us to impute the missing primary outcomes in the experimental sample, as shown by the following proposition.

prop(Surrogate Index) $(i)$ Suppose Assumption (ref) (Surrogacy) holds. Then: \[ \mu(s,w,x,{\rm E})=\mu(s,x,{\rm E}),\hskip1cm{\rm for\ all}\ s\in\mathbb{S},\ x\in\mathbb{X},\ \ {\rm and}\ w\in\mathbb{W}. \] $(ii)$ Suppose Assumption (ref) (Comparability) holds. Then: \[ \mu(s,x,{\rm E})=\mu(s,x,{\rm O})\ \ \ {\rm for\ all}\ s\in\mathbb{S}, \ {\rm and}\ x\in\mathbb{X}. \] $(iii)$ Suppose Assumptions (ref) (Surrogacy) and (ref) (Comparability) hold. Then: \[ \mu(s,w,x,{\rm E})=\mu(s,x,{\rm O})\ \ \ {\rm for\ all}\ s\in\mathbb{S},x\in\mathbb{X},\ {\rm and}\ w\in\mathbb{W}. \]

Because we can estimate $\mu(s,x,{\rm O})=\mathbb{E}[Y_{i}|S_{i}=s,X_{i}=x,P_{i}={\rm O}]$, we can impute the missing $Y_{i}$ in the experimental sample as $\mu(S_{i},X_{i},{\rm O})$.

Surrogacy, Mediation, Instrumental Variables, Directed Acyclical Graphs, and Missing Data

To provide context for the setup here and the key assumptions, it is useful to make a link to three related literatures, on mediation, instrumental variables, and missing data respectively. We describe the causal structures for surrogacy, mediation, and instrumental variables using a directed acyclical graph (DAG) pearl. The interpretations provided in this subsection are note essential to the main results in the next section.

Directed Acyclical Graph Representations

The surrogacy, mediation, and instrumental variables literatures all study causal structures involving a causally linked sequence of three (sets) of variables. They differ in three key aspects: $(i)$ the assumptions they make on the causal structure, $(ii)$ the estimands that are the primary focus of the analysis, and $(iii)$ the data available for the analyses. The literatures also differ in the labels typically used for the three variables. In Table (ref) we list the labels, estimands, and some of the assumptions.

table[table omitted — 815 chars of source]

In Figures 2.A-2.C we show the differences in structures in DAG form in a single sample setting (so that we need not be concerned with the comparability assumption). Figure 2.A illustrates the surrogacy setup, with a causal link from the treatment to the surrogate and from the surrogate to the outcome. There is no unobserved confounder for the causal relation between treatment and surrogate, which would violate Assumption (ref) (Unconfounded Treatment Assignment / Strong Ignorability). There is no direct causal link from the treatment to the outcome. There are also no unobserved confounders for the causal relation between surrogate and the outcome. These two features of the DAG (no direct link between treatment and outcome and no unobserved confounder for the relation between surrogate and outcome imply Assumption ((ref)) (Surrogacy).

Figure 2.B shows a mediation example where Assumption (ref) is violated because there is a direct effect of the treatment on the outcome that does not pass through the surrogate. In this case $S_{i}$ is a typically labelled a mediator, rather than a surrogate. In the mediation case the direct effect of the treatment on the outcome is estimable because all three variables, treatment, mediator and outcome are observed in the same sample.

figure[figure omitted — 2,816 chars of source]
comment\begin{figure}[h] \begin{subfigure}[b]{0.45\textwidth} \begin{centering} \scriptsizeFigure 2.A. Surrogacy Assumption Satisfied{\scriptsize} \end{centering} \begin{tikzpicture}[ >=stealth, node distance=1.5cm ] \node[observed, label=above:{\scriptsize\({\rm Treatment}\)}] (1) ; \node[observed, right=of 1, label=above:{\scriptsize\({\rm Surrogate}\)}] (2) ; \node[observed, right=of 2, label=above:{\scriptsize\({\rm Outcome}\)}] (4) ; \draw [->, notouch] (1.east) -- (2.west); \draw [->, notouch] (2.east) -- (4.west); \draw [->, notouch] (1.east) -- (2.west); \end{tikzpicture} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \begin{centering} \scriptsizeFigure 2.B. Violation of Surrogacy due to Direct Effect (Mediation Setup){\scriptsize} \end{centering} \begin{tikzpicture}[ >=stealth, node distance=1.5cm ] \node[observed, label=below:{\scriptsize\({\rm Treatment}\)}] (1) ; \node[observed, right=of 1, label=above:{\scriptsize\({\rm Surrogate}\)}] (2) ; \node[observed, right=of 2, label=above right:{\scriptsize\({\rm Outcome}\)}] (4) ; \draw [->, notouch] (1.east) -- (2.west); \draw [->, notouch] (2.east) -- (4.west); \draw [->, notouch] (1.east) -- (2.west); \draw [->, notouch] (1.north east) to [out=55, in=125] (4.north west); \end{tikzpicture} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \begin{centering} \end{centering} \begin{centering} \scriptsizeFigure 2.C. Violation of Surrogacy Assumption due to Unobserved Confounder (Instrumental Variable Setup) {\scriptsize} \end{centering} \begin{tikzpicture}[ >=stealth, node distance=1.5cm ] \node[observed, label=above:{\scriptsize\({\rm Instrument}\)}] (1) ; \node[observed, right=of 1, label=below:{\scriptsize\({\rm Treatment}\)}] (2) ; \node[observed, right=of 2, label=below right:{\scriptsize\({\rm Outcome}\)}] (4) ; \node[unobserved, above right=of 2, label=above:{\scriptsize\({\rm Unobserved \ Confounder}\)}] (5) ; \draw [->, notouch] (1.east) -- (2.west); \draw [->, notouch] (2.east) -- (4.west); \draw [->, notouch] (1.east) -- (2.west); \draw [dashed, ->, notouch] (5.south west) -- (2.north east); \draw [dashed, ->, notouch] (5.south east) -- (4.north west); \end{tikzpicture} \end{subfigure} \vskip1cm \end{figure}

Figure 2.C shows a DAG representation of the standard instrumental variables (IV) model familiar to economists. The first difference from the surrogacy setup in Figure 2.A is that in the instrumental variables setting the interest is in the causal effect of the variable in the middle of the three variable chain (the surrogate $S$ in the surrogacy setting, and the treatment $W$ in the instrumental variables setting), on the outcome, whereas in the surrogacy setting the primary interest is in the effect of the first variable in the chain (the treatment $W$ in the surrogacy setting and the instrument $Z$ in the instrumental variables setting) on the outcome. In the instrumental variables case the surrogacy estimand is immediately identified as the intention-to-treat effect of the instrument, since the instrument and the surrogate are observed in the same sample. Under the assumptions of the surrogacy setup, the target for an instrumental variables analysis, the effect of the surrogate on the primary outcome, is immediately identified. The instrumental variables settings is characterized by the presence of an unobserved confounder that affects both the treatment of interest and the outcome. The presence of that unobserved confounder violates Surrogacy, even if the treatment has no direct effect on the long-term outcome (frangakis2002principal,rosenbaum1984consequences,joffe2009related,vanderweele2015explanation).

The presence of this unobserved confounder also violates the comparability assumption if the marginal distribution of the treatment $W$ differs between the observational and experimental samples, as will typically be the case. In both the surrogacy and the instrumental variables cases, we assume the absence of a direct effect of the first variable in the causal chain (the treatment $W$ in the surrogacy case and the instrument in the instrumental variables case) on the primary outcome. In the surrogacy setting, this assumption is part of the Surrogacy assumption, while in the instrumental variables setting this is typically referred to as the exclusion restriction angrist1996identification.

A Missing Data Representation

In the Online Appendix we also discuss a missing data interpretation of the surrogacy approach. Eessentially we show that the following joint conditional independence assumption,

equation[equation omitted — 108 chars of source]

implies both surrogacy and comparability.

This missing data characterization is useful because it allows one to use insights from the missing data literature, both for the current problems and for generalizations. Given ((ref)) we can use the conditional distribution of $Y_i$ given $(S_{i},X_{i})$ in the observational sample with $P_i={\rm O}$ to impute the missing outcomes in the experimental sample with $P_i={\rm E}$, and we can use the conditional distribution of $W_{i}$ given $(S_{i},X_{i})$ in the experimental sample wth $P_i={\rm E}$ to impute the missing treatments in the observational sample with $P_i={\rm O}$.

This observation directly extends to more general imputation problems. Suppose we have two samples where in one sample, indicated by $P_{i}={\rm E}$ we observe one set of variables, $(Z_{i1},Z_{i2})$ and in the second sample, indicated by $P_{i}={\rm O}$ we observe a partially overlapping set of variables, $(Z_{i2},Z_{i3})$. Then the analogous assumption that allows the imputation of all missing variables is $P_{i}\perp\!\!\!\perp Z_{i1}\perp\!\!\!\perp Z_{i3}|Z_{i2}.$

Identification and Semiparametric Efficiency Bounds

Three Identification Results

We now present our central identification result. We analyze three different representations of the average treatment effect that lead to three estimation strategies, somewhat similar to inverse propensity score weighting, regression, and influence function estimators for average treatment effects under unconfoundedness imbens2004. The motivation for developing the different representations is that estimators corresponding to those different representations can have different properties in finite samples, just like they do in the unconfoundedness setting. Estimators based on the first representation require estimation of the surrogate index, but not the surrogate score. Estimators based on the second representation instead require estimation of the surrogate score, but not the surrogate index. Estimators based on the third representation require estimation of both, but have attractive double robustness properties.

We define the following four objects, all functionals of distributions that are directly estimable from the data. First define the statistical estimand, the average difference in the surrogate index between treated and control, adjusted for pretreatment variablles, in the experimental sample:

equation[equation omitted — 205 chars of source]

\[ \left.\left. - \mathbb{E}\Bigl[ \mathbb{E}\left[\left. Y_{i} \right|S_{i},X_{i},P_{i}={\rm O}\right] \Bigr| W_{i}=0,X_{i},P_{i}={\rm E} \Bigr] \Bigr\}\right| P_{i}={\rm E}\right]. \] Next, with a surrogate index representation:

equation[equation omitted — 217 chars of source]

then a surrogate score representation,

equation[equation omitted — 208 chars of source]

\[ \hskip2cm\left.\left.-Y_{i}\cdot\frac{(1-\rho(S_{i},X_{i}))\cdot \varphi(S_{i},X_{i})\cdot(1-\varphi)}{(1-\rho(X_{i}))\cdot(1-\varphi(S_{i},X_{i}))\cdot \varphi}\right|P_{i}={\rm O}\right]. \] The third representation is based on the influence function. We first define \[ \mu(w,x)\equiv \mathbb{E}[\mu(S_{i},X_{i},{\rm O})|W_{i}=w,X_{i}=x,P_{i}={\rm E}].\] Then the influence function is

equation[equation omitted — 219 chars of source]

\[\hskip2cm+ \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,x) - \mu(0,x)-\tau \Bigr) \] \[\hskip2cm+\frac{\boldsymbol{1}_{p={\rm O}}}{\varphi} \frac{\varphi(s,x)}{1-\varphi(s,x)}\frac{(y-{\mu}(s,x,{\rm O}))\left({\rho}(s,x)-{\rho}(x)\right)}{{\rho}(x)(1-{\rho}(x))} \] with the estimand

align[align omitted — 103 chars of source]
remarkAn earlier version of the paper had a mistake in the representation of the influence function. We are grateful to Kevin Chen and David Ritzwoller for pointing this out. See chen2023semiparametric for details.
theorem(Identification) $(i)$ Suppose that Assumption (ref) holds. Then, assuming all expectations are finite, \[\tau^*= \tau^{{\rm E}}=\tau^{{\rm O}}=\tau^{{\rm O},{\rm E}}.\] $(ii)$ Suppose that Assumptions (ref)--(ref) hold. Then the average treatment effect is equal to the following three estimable functions of the data: \[ \tau\equiv\mathbb{E}[Y_{i}(1)-Y_{i}(0)|P_{i}={\rm E}]= \tau^*=\tau^{{\rm E}}=\tau^{{\rm O}}=\tau^{{\rm O},{\rm E}}, \] $(iii)$ Jointly Assumptions (ref), (ref)$(i)$, (ref)$(i)$ and (ref)$(i)$ have no testable implications.
remarkThe first part of the theorem implies that the four functionals of the joint distribution of $(\boldsymbol{1}_{P_{i}={\rm O}}Y_{i},S_{i},\boldsymbol{1}_{P_{i}={\rm E}} W_{i},X_{i}, P_{i} )$ are identical, irrespective of the Unconfoundedness, Surrogacy, and Comparability assumptions.
remarkJust like in the unconfoundedness case newey1994asymptotic,chernozhukov2016double, the influence function representation is doubly robust. chen2023semiparametric show that if the functions in the influence function that represent conditional expectations of the outcome, $\mu(s,x,{\rm O})$ and $\mu(w,x)$, are correctly specified, then the influence function has expectation zero irrespective of the functions used for the various propensity score, $\rho(s,x)$, $\rho(x),$ $\rho$, and the sampling score $\varphi(s,x)$. Similarly, if the various propensity score, $\rho(s,x)$, $\rho(x),$ $\rho$, and the sampling score $\varphi(s,x)$ are correct, the influence function has expectation zero, irrespective of the functions used for the conditional outcome expectations $\mu(s,x,{\rm O})$ and $\mu(w,x)$.

Semiparametric Efficiency Bounds

In this subsection we present two pairs of semiparametric efficiency bound results bickel1993efficient,newey1990semiparametric for two different data configurations. The first directly refers to the main setup in this paper with the experimental and observational sample. This result is essentially shown in chen2023semiparametric which corrects a mistake in an earlier version of the current paper.

theoremSuppose Assumptions (ref)--(ref) hold. Then\\ $(i)$ the semiparametric efficiency bound, normalized by the square root of the sample size $N$, is \[ \mathbb{V}=\mathbb{E}[\psi(Y_{i},S_{i},W_{i},X_{i},P_{i})^2]\] \[\qquad=\mathbb{E}\bigg[ \frac{1-\varphi(S_{i},X_{i})}{\varphi^2} \left( \left( \frac{\varphi(S_{i},X_{i})}{1 -\varphi(S_{i},X_{i})} \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \right) \] \[+\frac{\varphi(X_{i})}{\varphi^2} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left. + \frac{\varphi(S_{i},X_{i})}{\varphi^2} \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right].\] $(ii)$ If in addition the observational sample is large relative to the experimental sample, and $\sup_{s,x}\varphi(s,x)\rightarrow 0$, then the efficiency bound, now normalized by the expected sample size of the experimental sample, $\mathbb{E}[N_{\rm E}]=\varphi N$ simplifies to \[\mathbb{E}\left[ \left. \left( \mu(1,X_{i}) - \mu(0,X_{i}) - \tau \right)^2 + \frac{ (1-W_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{W_{i}(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right| P_{i}={\rm E} \right].\]
remarkThis variance in part $(ii)$ of Theorem (ref) is smaller than the effiency bound we would obtain in a randomized experiment where we do observe the primary outcome and did not observe the surrogate. The bound in that case is well known since hahn1998role, \[\mathbb{E}\left[ \left. \left( \mu(1,X_{i}) - \mu(0,X_{i}) - \tau \right)^2 + \frac{(1-W_{i}) (Y_{i} - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{W_{i}(Y_{i} - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right| P_{i}={\rm E} \right].\] This advantage in terms of asymptotic precision of using the (true) predicted outcome $\mu(S_i,X_i,{\rm O})$ rather than the actual outcome $Y_i$ has been noted previously in day1996trial in a setting with binary outcomes. In the general case this gain is equal to \[\mathbb{E}\left[ \left. \frac{ (1-W_{i})(Y_{i}-\mu(S_{i},X_{i},{\rm O}))^2}{(1 - \rho(X_{i}))^2} + \frac{W_{i}(Y_{i}-\mu(S_{i},X_{i},{\rm O}))^2}{\rho(X_{i})^2} \right| P_{i}={\rm E} \right].\]

Next we consider the case where in a single sample we observe the treatment, primary outcome, surrogates and pre-treatment variables. In this single sample case we do not need the fifth variable, $P_i\in\{{\rm E},{\rm O}\}$. To maintain consistency with the other parts of the discussion and to avoid ambiguity, we keep the notation as before. In this case we can think of $P_{i}$ always taking the value $P_{i}={\rm E}$. We calculate the efficiency bound both without the assumption that surrogacy holds and with the assumption that surrogacy holds. We do so for a data generating process where surrogacy does hold, to see the information gain from that assumption.

theoremSuppose Assumptions (ref) and (ref) hold. $(i)$ The variance bound without assuming surrogacy is \[ \mathbb{V}_{{\rm ns}}=\mathbb{E}\biggl[\sigma^{2}(S_{i},X_{i},{\rm E})\cdot\left(\frac{\rho(S_{i},X_{i})}{\rho(X_{i})^{2}}+\frac{1-\rho(S_{i},X_{i})}{(1-\rho(X_{i}))^{2}}\right) +\left(\mu(1,X_i)-\mu(0,X_i)-\tau\right)^2 \] \[ \hskip2cm+\frac{\rho(S_{i},X_{i})}{\rho(X_{i})^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(1,X_{i})\right)^{2}+\frac{1-\rho(S_{i},X_{i})}{(1-\rho(X_{i}))^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(0,X_{i})\right)^{2} \biggr].\] $(ii)$ The efficiency gain from assuming surrogacy is \[ \Delta=\mathbb{V}_{{\rm ns}} - \mathbb{V}_{{\rm s}} = \mathbb{E}\left[\sigma^2\left(S_i, X_i, E\right) \frac{\rho\left(S_i, X_i\right)\left(1-\rho\left(S_i, X_i\right)\right)}{\rho\left(X_i\right)^2\left(1-\rho\left(X_i\right)\right)^2}\right] \geq 0 , \] where $\mathbb{V}_{\rm s}$ is the variance bound for the case with surrogacy, \[ \mathbb{V}_{{\rm s}}=\mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2+ \left( \mu(1,X_{i}) - \mu(0,X_{i}) - \tau \right)^2 \right.\] \[\left. + \frac{\rho(S_{i},X_{i})}{\rho(X_{i})^2} (\mu(S_{i},X_{i},{\rm E}) -\mu(1,X_{i}))^2 + \frac{1-\rho(S_{i},X_{i})}{ (1 - \rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm E}) - \mu(0,X_{i}))^2 \right] \]
remarkThe expression for $\mathbb{V}_{\rm ns}$ is equivalent to the efficiency bound in hahn1998role. It is written here in terms of the surrogates to facilitate the comparison to the efficiency bound exploiting surrogacy.
comment\begin{remark} The difference between the two bounds, $\mathbb{V}_{{\rm ns}}-\mathbb{V}_{{\rm s}}$, is the efficiency gain from the Surrogacy assumption. \end{remark}
remarkNote that the variance bound in part $(ii)$ of Theorem (ref) differs from that in Theorem (ref)$(ii)$ which was derived under the same surrogacy assumption, but assuming that the observational sample was infinitely large, so the relation between the surrogates and the primary outcome was known without error. The result in $(ii)$ captures just the value of the surrogacy assumption.

Violations of the Surrogacy and Comparability Assumptions: Biases and Bounds

The three critical assumptions, Unconfoundedness, Surrogacy, and Comparability, are strong. There is a large literature studying the sensitivity to unconfoundedness conditions rosenbaumrubin_sensitivity, imbens2003, cinelli2020making or bounds manski_bounds. Multiple studies have also raised concerns that in practice Surrogacy may not be satisfied begg2000, freedman1992statistical, frangakis2002principal,rosenbaum1984consequences,joffe2009related,vanderweele2015explanation, although we are not aware of formal sensitivity or bounds analyses. Violations of Comparability have not been explored because this assumption has not been previously formalized. In this section we examine the biases that arise from violations of Surrogacy and Comparability. We first characterize these biases and then derive estimable bounds on the magnitude of the biases that can arise from such violations.

Biases

We begin by characterizing the probability limit of estimators based on the representations of the estimand, $\tau^{{\rm E}}$, $\tau^{{\rm O}}$, and $\tau^{{\rm O},{\rm E}}$, in Theorem (ref) when the Surrogacy and Comparability assumptions are violated, as well as in cases where the surrogate index is misspecified. Throughout the section, we maintain Unconfoundedness in the experimental sample (Assumption (ref)), and the random sampling assumption (Assumption (ref)). We denote the probability limit of the estimators by $\underline{\tau}$ to differentiate it from the average treatment effect $\tau=\mathbb{E}[Y_i(1)-Y_i(0)|P_{i}={\rm E}].$

theorem$(i)$ Suppose Assumption (ref) (Unconfoundedness) holds, but Assumptions (ref) (Surrogacy) and (ref) (Comparability) do not necessarily hold. Then \[ \underline{\tau}\equiv \tau^{{\rm O}}=\tau^{{\rm E}}=\tau^{{\rm E},{\rm O}}=\mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]. \] $(ii)$ Suppose Assumptions (ref) (Unconfoundedness) and (ref) (Comparability) hold, but Assumption (ref) (Surrogacy) does not necessarily hold. Then the difference between the average causal effect and the estimand is \[\textrm{\rm (surrogacy-bias)}\quad \tau-\underline{\tau}= \mathbb{E}\left[\left.\Bigl\{\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]. \] $(iii)$. Suppose Assumptions (ref) (Unconfoundedness) and (ref) (Surrogacy) hold, but Assumption (ref) (Comparability) does not necessarily hold. Then the difference between the average causal effect and the estimand is \[\textrm{\rm (comparability-bias)}\quad \tau-\underline{\tau}=\mathbb{E}\left[\left.\Bigl\{ \mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]. \] $(iv)$. Suppose Assumption (ref) (Unconfoundedness) holds, but Assumptions (ref) (Surrogacy) and (ref) (Comparability) do not necessarily hold. Then the difference between the average causal effect and the estimand is \begin{align*} \rm (total\ bias)\quad & \tau-\tau=\mathbb{E}\left[\left(\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\right)\cdot\frac{(1-\rho(S_{i},X_{i}))\cdot \rho(S_{i},X_{i})}{(1-\rho(X_{i}))\cdot \rho(X_{i})}\mid P_{i}={\rm E}\right]\\ & \quad+\mathbb{E}\left[\left(\mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})\right)\cdot\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{(1-\rho(X_{i}))\cdot \rho(X_{i})}\mid P_{i}={\rm E}\right]. \end{align*}
remarkTheorem (ref)$(i)$ shows that even without Surrogacy and Comparability, we estimate a valid average causal effect as long as unconfoundedness holds. The treatment effect we estimate is the average effect of the treatment on the surrogate index -- a principled aggregate of intermediate outcomes -- rather than the average effect on the primary outcome. This result also shows that the interpretation does not change with the choice of estimator (using the surrogate score approach, the surrogate index approach, or the influence function). Theorem (ref)$(ii-iv)$ show how violations of Comparability or Surrogacy affect the difference between what is being estimated and the average treatment effect on the primary outcome.
remarkThe bias from violations of Surrogacy (Theorem (ref)$(ii)$) consists of two factors. The first factor is small if the treatment does not explain much of the variation in $Y_{i}$ and therefore $\mu(s,1,x,{\rm E})$ and $\mu(s,0,x,{\rm E})$ are close. The second factor is small if the surrogate explains a large share of the variation in $W_{i}$, so that the surrogate score is close to zero or one and therefore $\mathbb{E}[\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))]$ is close to zero.
remarkThe bias from violations of Comparability (Theorem (ref)$(iii)$) also consists of two factors. The first is the difference between the surrogacy index $\mu(s,x,{\rm O})$ and its counterpart in the experimental sample, $\mu(s,x,{\rm E})$. The second factor depends on the deviation between the surrogacy score and the propensity score, $\rho(S_{i},X_{i})-\rho(X_{i})$. If the treatment does not have much effect on the surrogates, violations of Comparability do not generate much bias, because the bias that comes from a combination of the effect of the treatment on the surrogates and the effect of the surrogates on the outcome, will be small in that case.

Bounds on the Bias

In this subsection we explore bounds on the parameter of interest. We show that in general these bounds are uninformative. However, if outcomes themselves are bounded, for example, if the outcomes are binary, informative bounds can be derived. Moreover, we present bounds given assumptions on the range of violations of the Surrogacy and Comparability assumptions.

lemmaSuppose Assumptions (ref) (Unconfoundedness) and (ref) (Comparability) hold, but Assumption (ref) (Surrogacy) does not necessarily hold. Then:\\ $(i)$ If the outcome can take on values on the whole real line, then there is no value for the average treatment effect $\tau$ that can be ruled out.\\ $(ii)$ if the outcome is binary, then the average treatment effect $\tau$ is inside the interval \[ \biggl\{ \mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]+ \mathbb{E}\left[\left. \Delta^L_S(S_{i},X_{i}) \frac{\rho(S_{i},X_{i})(1-\rho(S_{i},X_{i}))}{\rho(X_{i})(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right], \] \[ \mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]+\mathbb{E}\left[\left. \Delta_S^U(S_{i},P_{i}) \frac{\rho(S_{i},X_{i})(1-\rho(S_{i},X_{i}))}{\rho(X_{i})(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] \biggr\}, \] where \[ \Delta^L_S(s,x)=- \min\left(\frac{1-\mu(s,x,{\rm O})}{\rho(s,x)},\frac{\mu(s,x,{\rm O})}{1-\rho(s,x)}\right) \qquad \Delta_S^U(s,x)= \min\left(\frac{\mu(s,x,{\rm O})}{\rho(s,x)},\frac{1-\mu(s,x,{\rm O})}{1-\rho(s,x)}\right), \] and these bounds on the bias are sharp.\\ $(iii)$ if the direct effect of the treatment on the outcome $\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})$ is bounded in absolute value by $c$, then the average treatment effect $\tau$ is inside the interval \[ \biggl\{ \mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]-c\cdot \mathbb{E}\left[\left. \frac{\rho(S_{i},X_{i})(1-\rho(S_{i},X_{i}))}{\rho(X_{i})(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right], \] \[ \mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]+c\cdot \mathbb{E}\left[\left. \frac{\rho(S_{i},X_{i})(1-\rho(S_{i},X_{i}))}{\rho(X_{i})(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] \biggr\}, \] and this bound is sharp.
remarkTo provide some intuition for the sharpness of the bounds, consider the surrogacy bias in Theorem (ref). The bias has two factors, with the second estimable from the data. The first factor is the difference $\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})$. The data are not directly informative about this difference beyond the fact that the weighted average $\rho(S_{i},X_{i})\mu(S_{i},1,X_{i},{\rm E})+(1-\rho(S_{i},X_{i}))\mu(S_{i},0,X_{i},{\rm E})$ is equal to the estimable quantity $\mu(S_{i},X_{i},{\rm O})$. In the absence of any restrictions on the outcome this implies there are no restrictions on $\mu(S_{i},w,X_{i},{\rm E})$ or on the difference $\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})$, and thus not on the bias or the average treatment effect. Given restrictions on the range of the outcome this representation directly leads to upper and lower bounds on the bias and the average treatment effect.
lemmaSuppose Assumptions (ref) (Unconfoundedness) and (ref) (surrogacy) hold, but Assumption (ref) (Comparability) does not necessarily hold. Then:\\ $(i)$ If the outcome can take on value on the whole real line, then there is no value for the average treatment effect $\tau$ that can be ruled out.\\ $(ii)$ if the outcome is binary, then the average treatment effect $\tau$ is inside the interval \[\left\{ \mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]+ \mathbb{E}\left[\left.\Bigl\{ \boldsymbol{1}_{\rho(S_{i},X_{i})<\rho(X_{i})}-\mu(S_{i},X_{i},{\rm O})\Bigr\}\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right],\right. \] \[\left. \mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]+ \mathbb{E}\left[\left.\Bigl\{ \boldsymbol{1}_{\rho(S_{i},X_{i})>\rho(X_{i})}-\mu(S_{i},X_{i},{\rm O})\Bigr\}\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]\right\}, \] with width \[2 \mathbb{E}\left[\left. \boldsymbol{1}_{\rho(S_{i},X_{i})>\rho(X_{i})}\cdot\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] , \] and these bounds on the bias are sharp.\\ $(iii)$ if $\mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})$ is bounded in absolute value by $c$, then the average treatment effect $\tau$ is inside the interval \[\left\{ \mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]-c\cdot \mathbb{E}\left[\left.\frac{|\rho(S_{i},X_{i})-\rho(X_{i})|}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right],\right. \] \[\left. \mathbb{E}\left[\left.\mu(S_{i}(1),X_{i},{\rm O})-\mu(S_{i}(0),X_{i},{\rm O})\right|P_{i}={\rm E}\right]+c\cdot \mathbb{E}\left[\left.\frac{|\rho(S_{i},X_{i})-\rho(X_{i})|}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]\right\}, \] and these bound are sharp.

Estimation

In this section, we first present four estimators for the average treatment effect. The first, the surrogate index estimator, is related to previously proposed estimators with the difference that in the earlier literature the surrogate index was implicitly assumed to be known. We then discuss three new alternative estimators. The last of these new estimators is a matching estimator. Although matching estimators are generally not efficient in settings with unconfoundedness (rubin2006matched,abadie2006,abadie2016matching), they are widely applied, and it is instructive to see how a matching strategy can be used here.

Surrogate Index

Suppose we estimate the surrogate index as $\hat{\mu}(s,x,{\rm O})$ and the propensity score as $\hat{\rho}_{\rm E}(x)$. We take an average of the surrogate index in the experimental sample for the treatment and control groups, after adjusting for the propensity score. A natural estimator, corresponding to ((ref)), is the following difference of the two averages over the experimental sample:

equation[equation omitted — 200 chars of source]

\[ \hskip3cm-\frac{1}{\sum_{i=1}^{N_{{\rm E}}}(1-W_{i})/(1-\hat{\rho}(X_{i}))}\sum_{i=1}^{N_{{\rm E}}}\hat{\mu}(S_{i},X_{i},{\rm O})\cdot\frac{1-W_{i}}{1-\hat{\rho}(X_{i})}. \] We refer to this as the surrogate index estimator. Note that compared to the representation in Theorem (ref), we normalize the weights so that the weights sum up to one. This tends to improve the finite sample properties of related estimators in other settings substantially (hirano2003efficient,busso2014new).

In the case where the estimator for the surrogate index ${\mu}(s,x,{\rm O})$ was based on a linear specification for the regression of the primary outcome on the intermediate outcome, $\mu(s,x,{\rm O})=\gamma_{0}+\gamma_{S}'s+\gamma_{X}'x$, this leads to \[ \hat{\tau}^{{\rm E}}=\hat{\gamma}_{S}'\hat{\tau}_{S}, \] where $\hat{\tau}_{S}$ is an estimator for the average effect of the treatment on the surrogates, $\mathbb{E}\left[S_{i}(1)-S_{i}(0)\right].$ In the simplest case without pre-treatment variables and where the experimental sample is randomized, $\hat{\tau}_{S}=\overline{S}_{1}-\overline{S}_{0}$, where $\overline{S}_{1}$ and $\overline{S}_{0}$ are the average values of the surrogate outcomes. Here, the estimator simplifies to the difference in the estimated surrogate index in the treatment group and the control group: $ \hat{\tau}^{{\rm E}}=\hat{\gamma}_{S}'(\overline{S}_{1}-\overline{S}_{0})$. This expression is also familiar from the mediation literature (e.g., baron1986moderator) and the surrogacy literature day1996trial. However, we emphasize that in general, there may be interactions between the surrogates and pre-treatment variables, and in that case the linear specification need not be not adequate.

Surrogate Score Estimator

We now use the second representation for $\tau$ in the main theorem to derive an alternative estimator. Let $\hat{\rho}(x)$, $\hat{\rho}(s,x),\hat{\varphi}(s,x),\hat{\varphi}(x)$, and $\hat{\varphi}$, be estimators for $\rho(x)$, $\rho(s,x),\varphi(s,x),\varphi(x)$, and $\varphi$ respectively.

The surrogate score estimator is based on averaging the following expression over the observational sample:

equation[equation omitted — 232 chars of source]

where for $w=0,1$ the weights are

equation[equation omitted — 278 chars of source]

Influence Function Estimator

We can also base estimation on the efficient score given in ((ref)). Given estimators for the propensity score, the surrogate score, and the sampling score, we can estimate the average treatment effect as

equation[equation omitted — 282 chars of source]

\[+ \frac{\boldsymbol{1}_{P_{i}={\rm E}}}{\hat{\varphi}} \Biggl(\hat{\mu}(1,X_{i}) \left( 1 - \frac{W_{i}}{\hat{\rho}(X_{i})} \right) - \hat{\mu}(0,X_{i}) \left( 1 - \frac{1 - W_{i}}{1 - \hat{\rho}(X_{i})} \right) \Biggr) \] \[ \hskip2cm+\frac{\boldsymbol{1}_{P_{i}={\rm O}}}{1-\hat{\varphi}}\left(\frac{\hat{\varphi}(S_{i},X_{i})}{1-\hat{\varphi}(S_{i},X_{i})} \frac{1-\hat{\varphi}}{\hat{\varphi}}\right) \frac{(Y_{i}-\hat{\mu}(S_{i},X_{i},{\rm O}))\left(\hat{\rho}(S_{i},X_{i})-\hat{\rho}(X_{i})\right)}{\hat{\rho}(X_{i})(1-\hat{\rho}(X_{i}))} \Biggr\}. \] Based on the results in newey1994asymptotic, it follows that under standard conditions the two estimators above and the surrogate index estimator all reach the semi-parametric efficiency bound, and are first-order equivalent.

The recent literature on double robust estimation of average treatment effects under unconfoundedness chernozhukov2016double suggests that this estimator may have superior properties in small samples.

Double Matching Estimator

Consider unit $i$ in the experimental sample with $X_{i}=x$ and $S_{i}=s$, and suppose this is a treated unit with $W_{i}=1$. We need to find three matches for this unit. First, we need to find a unit with the opposite treatment in the same (experimental) sample. Specifically, we need to find the closest unit in the experimental sample, in terms of pre-treatment variables, among the units with $W_{i}=0$. Suppose this unit is unit $j$, with $W_{j}=0$, and the value of the pre-treatment variables for this unit are $X_{j}=x'$, and the surrogate outcomes are $S_{j}=s'$. As a result of the matching we should have $x\approx x'$, but potentially $s$ could be quite different from $s'$. Next, we need to find for each of the two units $i$ and $j$ a match in the observational sample. Find the unit in the observational sample closest to unit $i$, in terms of both pre-treatment variables and surrogates. Let $i'$ be the index for this unit, and let the value of the outcome for this unit be $Y_{i'}$, and the values of the pre-treatment variables and surrogates $X_{i'}$ and $S_{i'}$. Now as a result of the matching $X_{i}\approxX_{i'}$ and $S_{i}\approxS_{i'}$. Finally, find the unit in the observational sample closest to unit $j$, in terms of both pre-treatment variables and surrogates. Let the value of the outcome for this unit be $Y_{j'}$, and the values of the pre-treatment variables and surrogates $X_{j'}$ and $S_{j'}$, with $X_{j}\approxX_{j'}$ and $S_{j}\approxS_{j'}$.

Then we combine these matches to estimate the causal effect for unit $i$, $Y_{i}(1)-Y_{i}(0)$, as the difference in average outcomes for the two matches from the observational sample:

equation[equation omitted — 89 chars of source]

The matching estimator for $\tau$ would then be the average value of ((ref)) over the experimental sample. The double matching estimator is then \[ \hat\tau^{\rm match}=\frac{1}{N^{\rm E}}\sum_{i:P_{i}{\rm E}}\left\{ W_i \left( Y_{i'}-Y_{j'}\right) +(1-W_i) \left( Y_{j'}-Y_{i'}\right)\right\}. \]

Application: Impacts of Job Training on Employment

In this section, we apply our method to estimate the causal effect of the Greater Avenues to Independence (GAIN) job training program on long-term labor market outcomes. GAIN was a job assistance program implemented in California in the 1980s to help welfare recipients find work (riccio1989gain,friedlander1995evaluating, hotz2006evaluating). MDRC conducted a randomized trial to evaluate the GAIN program's employment impacts in six counties in California in the late 1980s. We focus primarily on the GAIN trial in Riverside, which was widely heralded as the program that had the largest treatment effects on earnings. The Riverside program emphasized a “jobs first” approach to re-entry into the labor force, encouraging unemployed workers to take any job they find; in contrast, other sites focused more heavily on developing human capital through training programs (hotz2006evaluating).

We have available long-term outcomes for the four GAIN sites, including employment, earnings, and receipt of aid over the first thirty-six quarters after random assignment. We take the average of the thirty-six employment indicators and earnings in Riverside as our primary outcomes. We then investigate whether we could have predicted the long-term impact on these outcomes using only the first $T$ quarters of all outcomes (including employment, earnings, and aid) as surrogates, as well as using pre-treatment variables (characteristics of the individuals as well as lagged employment, earnings and aid outcomes). The Riverside data on the treatment, surrogates and pre-treatment variables play the role of of our experimental sample. We use the data from the combination of the other three locations (Alameda, Los Angeles, and San Diego) as our observational sample. For the observational sample we only use the information on the surrogates, pre-treatment variables, and outcome, but not the treatment assignment, nor the indicator for the location.

We begin by presenting a brief summary of the samples. We then describe how we construct our surrogate index. Next we illustrate our theoretical results by evaluating the magnitude of the gains from using surrogate indices in terms of time and precision relative to existing experimental estimates of the program's long-term impacts in Riverside. We also show how one can validate the surrogacy assumption using intermediate outcomes and bound the degree of bias arising from potential violations of surrogacy.

The GAIN Program

The GAIN treatment was randomly assigned to welfare (Aid for Families with Dependent Children) recipients, a very low-income population. The treatment group consisted of $N_{{\rm E},T}=4405$ participants, which the control group consisted of $N_{{\rm E},C}=1040$ participants who were not eligible for the additional services in the GAIN program. The data we use come from the hotz2006evaluating which followed study participants for nine years after assignment of the treatment, measuring quarterly employment rates and earnings\footnote{ All income variables were converted to 1999 dollars using cost-of-living deflators; see footnote 21 of hotz2000long for more information.} from the Unemployment Insurance database. They found that the treatment effects of the Riverside GAIN program on employment rates and earnings were initially large, but declined over time, as shown in Figure 3A, which plots employment rates by quarter for individuals in the experimental (Riverside) treatment and control groups, and in Figure 3B, which shows the correspond results for quarterly earnings.

comment\begin{figure}[htbp] \end{figure} \begin{figure}[htbp] \end{figure}
figure[figure omitted — 398 chars of source]

In Riverside, the estimated causal effects on the primary outcomes were a 6.4 (s.e. = 1.2) percentage point (pp) increase in average quarterly employment rates, and an \$249 (s.e. \$84) increase in average quarterly earnings, in both cases averaged over the 36 quarter post-treatment. Our question is whether these impacts could have been estimated more quickly by using short-term employment, earnings and aid receipt as surrogates.

The observational sample includes the other three locations, Alameda, Los Angeles and San Diego, for a total of $N_{\rm O}= 13,725$ individuals.

In the online appendix Table (ref) presents information on the pre-treatment variables. Clearly the two samples, Riverside and the combination of the other three locations, are substantially different prior to the intervention in terms of permanent characteristics such as ethnicity, as well as in pre-treatment outcomes.

Three Estimators

We discuss here the estimators for the average effect of the program We wish to consider different set of surrogates, indexed by the number of periods $t$ we want to use as surrogates. To capture this we index the surrogate for individual $i$, $S_i^t$, by the superscript $t$. $S_i^t$ contains the employment indicators, earnings outcomes and aid receipt indicators for the $t$ quarters after the intervention.

Surrogate Index Estimator

To construct the surrogacy index we estimate a linear regression model using least squares, for the individuals in the observational sample

equation[equation omitted — 120 chars of source]

The predicted value from this regression, which we denote by $\hat{Y}_{i}$, is our surrogate index for mean employment based on surrogates up to quarter $t$. We then compute this surrogate index for each of the individuals in the experimental sample and estimate the treatment effect based on the surrogate index as

equation[equation omitted — 197 chars of source]

If we use the all 36 quarters of employment indicators are surrogates, then the regression of $Y_i$ on the set of surrogates will fit perfectly, $\hat{Y}_i$ will be equal to $Y_i$, and the estimated effect will be identical to the original experimental estimate. The question is whether using a much more limited set of surrogates will get us close to the experimental benchmark.

Surrogate Score Estimator

For the surrogate score estimator we first estimate a logistic regression of the treatment indicator on the pretreatment variables and the surrogates. We specify \[ \ln\left( \frac{\rho(S_i^t,X_i)}{1-\rho(S_i^t,X_i)} \right)\equiv \ln\left( \frac{{\rm pr}(W_i=1|S_i^t,X_i,P_{i}={\rm E})}{1-{\rm pr}(W_i=1|S_i^t,X_i,P_{i}={\rm E})} \right)= \alpha_0+\alpha_{S}^\top S_{i}^t+\alpha_X^\top X_i,\] and estimate this on the experimental (Riverside) sample.

Next we estimate the propensity score, also as a logistic regression, \[ \ln\left( \frac{\rho(X_i)}{1-\rho(X_i)} \right)\equiv \ln\left( \frac{{\rm pr}(W_i=1|X_i,P_{i}={\rm E})}{1-{\rm pr}(W_i=1|X_i,P_{i}={\rm E})} \right)= \delta_0+\delta_X^\top X_i,\] and estimate this again on the experimental (Riverside) sample. In principle the random assignment implies that the $\delta_X$ should be close to zero in this case.

Finally we estimate the comparability score \[ \ln\left( \frac{\varphi(S_i^t,X_i)}{1-\varphi(S_i^t,X_i)} \right)\equiv \ln\left( \frac{{\rm pr}(P_{i}={\rm E}|X_i,S_i^t)}{1-{\rm pr}(P_{i}={\rm E}|X_i,S_i^t)} \right)= \gamma_0+\gamma_{S}^\top S_{i}^t+\gamma_X^\top X_i,\] and estimate this on the combined observational and experimental samples.

The surrogate score estimator is based on averaging the following expression over the observational sample:

equation[equation omitted — 216 chars of source]

where the weights are as before in Equation ((ref)).

Influence Function Estimator

For the influence function estimator we first estimate the surrogacy index, the surrogacy score, the propensity score, and the comparability score as before. We then plug those into the estimator in Equation ((ref)).

Results

Here we discuss two sets of results. First the estimates for the average effect of the intervention on the two primary outcomes under various assumptions about the surrogates. Second, we test the Surrogacy and Comparability assumptions directly.

Estimation Results

figure[figure omitted — 140 chars of source]
figure[figure omitted — 138 chars of source]

As the discussion after the surrogate index estimator shows, we recover the experimental estimates if we use all 36 quarters of employment indicators as surrogates. The question is whether we can do approximately as well with fewer than 36 quarters of surrogates. In Figures 4A and 4B we compare the experimental estimates of the effect on the primary outcomes (0.064 for the employment outcome, and \$249 for the earnings outcome) to the three sets of surrogate estimates, as a function of how many periods of surrogates we use, ranging from 1 quarter to 36 quarters. To put this in perspective we also include in these two figures what we label the “naive” estimator where we estimate the effect on the long-term outcome as the effect on the first $t$ quarters of the outcome. In Tables (ref) and (ref) we report a subset of the numbers underlying these estimates with the corresponding standard errors.

comment\begin{figure}[htbp] \begin{subfigure}[b]{0.6\textwidth} \caption{Figure 4A: Employment} \end{subfigure} \begin{subfigure}[b]{0.6\textwidth} \caption{Figure 4B: Earnings} \end{subfigure} \caption{Figure 4: Employment and Earnings} \end{figure}
table[table omitted — 1,253 chars of source]
comment========================================================================== 1 & 0.049 & 0.007 & 0.011 & 0.002 & 0.010 & 0.001 & 0.009 & 0.001\\ 2 & 0.087 & 0.008 & 0.033 & 0.002 & 0.032 & 0.002 & 0.032 & 0.002\\ 3 & 0.104 & 0.010 & 0.042 & 0.004 & 0.043 & 0.004 & 0.043 & 0.004\\ 4 & 0.110 & 0.009 & 0.047 & 0.005 & 0.050 & 0.006 & 0.050 & 0.006\\ 5 & 0.115 & 0.010 & 0.055 & 0.005 & 0.058 & 0.006 & 0.058 & 0.006\\ \addlinespace 6 & 0.117 & 0.010 & 0.061 & 0.005 & 0.063 & 0.006 & 0.064 & 0.007\\ 7 & 0.115 & 0.010 & 0.059 & 0.007 & 0.063 & 0.007 & 0.064 & 0.008\\ 8 & 0.114 & 0.010 & 0.064 & 0.007 & 0.068 & 0.008 & 0.068 & 0.008\\ 9 & 0.115 & 0.011 & 0.068 & 0.007 & 0.073 & 0.008 & 0.073 & 0.008\\ 10 & 0.113 & 0.011 & 0.067 & 0.008 & 0.072 & 0.009 & 0.072 & 0.009\\ \addlinespace 11 & 0.111 & 0.011 & 0.066 & 0.008 & 0.071 & 0.009 & 0.072 & 0.009\\ 12 & 0.108 & 0.011 & 0.065 & 0.008 & 0.071 & 0.010 & 0.072 & 0.010\\ 13 & 0.105 & 0.011 & 0.063 & 0.008 & 0.069 & 0.009 & 0.070 & 0.010\\ 14 & 0.103 & 0.011 & 0.066 & 0.008 & 0.072 & 0.010 & 0.073 & 0.010\\ 15 & 0.101 & 0.011 & 0.066 & 0.008 & 0.073 & 0.009 & 0.073 & 0.009\\ \addlinespace 16 & 0.098 & 0.011 & 0.064 & 0.008 & 0.071 & 0.009 & 0.071 & 0.009\\ 17 & 0.097 & 0.011 & 0.065 & 0.008 & 0.074 & 0.010 & 0.074 & 0.010\\ 18 & 0.095 & 0.010 & 0.065 & 0.009 & 0.073 & 0.010 & 0.073 & 0.010\\ 19 & 0.093 & 0.011 & 0.065 & 0.010 & 0.073 & 0.010 & 0.073 & 0.010\\ 20 & 0.091 & 0.011 & 0.066 & 0.010 & 0.073 & 0.010 & 0.073 & 0.010\\ \addlinespace 21 & 0.090 & 0.011 & 0.065 & 0.010 & 0.072 & 0.010 & 0.072 & 0.010\\ 22 & 0.088 & 0.011 & 0.065 & 0.009 & 0.071 & 0.010 & 0.072 & 0.010\\ 23 & 0.087 & 0.011 & 0.065 & 0.010 & 0.071 & 0.010 & 0.071 & 0.010\\ 24 & 0.085 & 0.011 & 0.064 & 0.010 & 0.070 & 0.011 & 0.071 & 0.011\\ 25 & 0.084 & 0.011 & 0.065 & 0.010 & 0.071 & 0.011 & 0.072 & 0.011\\ \addlinespace 26 & 0.081 & 0.011 & 0.062 & 0.010 & 0.068 & 0.010 & 0.069 & 0.011\\ 27 & 0.079 & 0.011 & 0.060 & 0.010 & 0.067 & 0.010 & 0.068 & 0.011\\ 28 & 0.077 & 0.011 & 0.059 & 0.010 & 0.066 & 0.010 & 0.067 & 0.011\\ 29 & 0.075 & 0.011 & 0.060 & 0.010 & 0.067 & 0.010 & 0.068 & 0.011\\ 30 & 0.073 & 0.011 & 0.059 & 0.010 & 0.067 & 0.010 & 0.068 & 0.011\\ \addlinespace 31 & 0.072 & 0.011 & 0.060 & 0.010 & 0.067 & 0.010 & 0.068 & 0.011\\ 32 & 0.071 & 0.011 & 0.060 & 0.010 & 0.067 & 0.010 & 0.068 & 0.011\\ 33 & 0.069 & 0.011 & 0.059 & 0.010 & 0.066 & 0.010 & 0.067 & 0.011\\ 34 & 0.067 & 0.012 & 0.058 & 0.010 & 0.065 & 0.010 & 0.066 & 0.011\\ 35 & 0.066 & 0.012 & 0.058 & 0.010 & 0.065 & 0.010 & 0.067 & 0.011\\ \addlinespace 36 & 0.064 & 0.012 & 0.058 & 0.010 & 0.065 & 0.010 & 0.066 & 0.011\\ \begin{table} \caption{Estimates for Figure 4A Table for Mean Employment Rate} \begin{tabular}[t]{rrrrrrrrr} \toprule \multicolumn{1}{c} & \multicolumn{2}{c}{Naive} & \multicolumn{2}{c}{Regression} & \multicolumn{2}{c}{Weighted} & \multicolumn{2}{c}{Double Robust} \\ \cmidrule(l{3pt}r{3pt}){2-3} \cmidrule(l{3pt}r{3pt}){4-5} \cmidrule(l{3pt}r{3pt}){6-7} \cmidrule(l{3pt}r{3pt}){8-9} t & Coefficient & Std. Dev. & Coefficient & Std. Dev. & Coefficient & Std. Dev. & Coefficient & Std. Dev.\\ \midrule 1 & 0.049 & 0.013 & 0.011 & 0.003 & 0.010 & 0.002 & 0.009 & 0.003\\ 2 & 0.087 & 0.012 & 0.033 & 0.003 & 0.032 & 0.003 & 0.032 & 0.004\\ 3 & 0.104 & 0.011 & 0.042 & 0.004 & 0.043 & 0.004 & 0.043 & 0.004\\ 4 & 0.110 & 0.011 & 0.047 & 0.005 & 0.050 & 0.005 & 0.050 & 0.005\\ 5 & 0.115 & 0.011 & 0.055 & 0.005 & 0.058 & 0.005 & 0.058 & 0.005\\ \addlinespace 6 & 0.117 & 0.010 & 0.061 & 0.006 & 0.063 & 0.006 & 0.064 & 0.006\\ 7 & 0.115 & 0.010 & 0.059 & 0.006 & 0.063 & 0.006 & 0.064 & 0.006\\ 8 & 0.114 & 0.010 & 0.064 & 0.006 & 0.068 & 0.006 & 0.068 & 0.007\\ 9 & 0.115 & 0.010 & 0.068 & 0.006 & 0.073 & 0.007 & 0.073 & 0.007\\ 10 & 0.113 & 0.010 & 0.067 & 0.007 & 0.072 & 0.007 & 0.072 & 0.007\\ \addlinespace 11 & 0.111 & 0.010 & 0.066 & 0.007 & 0.071 & 0.007 & 0.072 & 0.007\\ 12 & 0.108 & 0.010 & 0.065 & 0.007 & 0.071 & 0.008 & 0.072 & 0.008\\ 13 & 0.105 & 0.010 & 0.063 & 0.007 & 0.069 & 0.008 & 0.070 & 0.008\\ 14 & 0.103 & 0.010 & 0.066 & 0.008 & 0.072 & 0.008 & 0.073 & 0.008\\ 15 & 0.101 & 0.010 & 0.066 & 0.008 & 0.073 & 0.008 & 0.073 & 0.008\\ \addlinespace 16 & 0.098 & 0.010 & 0.064 & 0.008 & 0.071 & 0.009 & 0.071 & 0.009\\ 17 & 0.097 & 0.010 & 0.065 & 0.008 & 0.074 & 0.009 & 0.074 & 0.009\\ 18 & 0.095 & 0.010 & 0.065 & 0.008 & 0.073 & 0.009 & 0.073 & 0.009\\ 19 & 0.093 & 0.010 & 0.065 & 0.008 & 0.073 & 0.009 & 0.073 & 0.009\\ 20 & 0.091 & 0.010 & 0.066 & 0.009 & 0.073 & 0.009 & 0.073 & 0.009\\ \addlinespace 21 & 0.090 & 0.010 & 0.065 & 0.009 & 0.072 & 0.009 & 0.072 & 0.009\\ 22 & 0.088 & 0.010 & 0.065 & 0.009 & 0.071 & 0.009 & 0.072 & 0.009\\ 23 & 0.087 & 0.010 & 0.065 & 0.009 & 0.071 & 0.009 & 0.071 & 0.010\\ 24 & 0.085 & 0.010 & 0.064 & 0.009 & 0.070 & 0.010 & 0.071 & 0.010\\ 25 & 0.084 & 0.010 & 0.065 & 0.009 & 0.071 & 0.010 & 0.072 & 0.010\\ \addlinespace 26 & 0.081 & 0.010 & 0.062 & 0.009 & 0.068 & 0.010 & 0.069 & 0.010\\ 27 & 0.079 & 0.010 & 0.060 & 0.009 & 0.067 & 0.010 & 0.068 & 0.010\\ 28 & 0.077 & 0.010 & 0.059 & 0.009 & 0.066 & 0.010 & 0.067 & 0.010\\ 29 & 0.075 & 0.010 & 0.060 & 0.009 & 0.067 & 0.010 & 0.068 & 0.010\\ 30 & 0.073 & 0.010 & 0.059 & 0.009 & 0.067 & 0.010 & 0.068 & 0.010\\ \addlinespace 31 & 0.072 & 0.010 & 0.060 & 0.009 & 0.067 & 0.010 & 0.068 & 0.010\\ 32 & 0.071 & 0.010 & 0.060 & 0.009 & 0.067 & 0.010 & 0.068 & 0.010\\ 33 & 0.069 & 0.010 & 0.059 & 0.009 & 0.066 & 0.010 & 0.067 & 0.010\\ 34 & 0.067 & 0.010 & 0.058 & 0.009 & 0.065 & 0.010 & 0.066 & 0.010\\ 35 & 0.066 & 0.010 & 0.058 & 0.009 & 0.065 & 0.010 & 0.067 & 0.010\\ \addlinespace 36 & 0.064 & 0.010 & 0.058 & 0.009 & 0.065 & 0.010 & 0.066 & 0.010\\ \bottomrule \end{tabular} \end{table}
table[table omitted — 1,470 chars of source]
comment========================================================================== 1 & 122.426 & 28.233 & 41.765 & 12.470 & 29.639 & 8.547 & 20.329 & 7.990\\ 2 & 217.535 & 30.493 & 131.052 & 19.911 & 150.678 & 42.489 & 139.714 & 45.641\\ 3 & 260.571 & 41.781 & 154.463 & 31.112 & 186.779 & 39.321 & 177.021 & 40.050\\ 4 & 284.399 & 46.174 & 172.690 & 29.878 & 225.073 & 37.615 & 215.102 & 39.518\\ 5 & 306.479 & 47.445 & 209.587 & 26.957 & 253.358 & 33.919 & 244.004 & 36.119\\ \addlinespace 6 & 327.145 & 47.384 & 238.845 & 24.723 & 279.654 & 38.107 & 270.099 & 40.572\\ 7 & 333.936 & 47.523 & 231.950 & 32.088 & 277.504 & 39.270 & 267.307 & 41.653\\ 8 & 341.688 & 47.473 & 244.669 & 31.529 & 298.461 & 44.925 & 287.519 & 47.222\\ 9 & 352.244 & 45.729 & 265.792 & 37.528 & 321.788 & 42.905 & 311.348 & 44.589\\ 10 & 355.966 & 43.668 & 259.683 & 36.556 & 313.781 & 38.459 & 303.257 & 40.656\\ \addlinespace 11 & 356.451 & 43.739 & 259.301 & 39.177 & 313.031 & 42.415 & 302.224 & 44.043\\ 12 & 353.884 & 45.035 & 249.100 & 45.288 & 306.777 & 49.169 & 296.049 & 50.186\\ 13 & 348.460 & 46.275 & 238.394 & 54.170 & 300.341 & 57.710 & 289.735 & 58.682\\ 14 & 348.260 & 49.152 & 254.535 & 57.780 & 317.187 & 60.656 & 306.680 & 61.197\\ 15 & 345.738 & 52.694 & 250.325 & 65.528 & 314.697 & 66.418 & 304.178 & 67.196\\ \addlinespace 16 & 342.980 & 54.416 & 247.981 & 65.591 & 315.251 & 66.624 & 304.500 & 67.867\\ 17 & 342.156 & 56.451 & 255.201 & 70.224 & 325.395 & 69.357 & 314.787 & 70.774\\ 18 & 340.180 & 58.660 & 252.269 & 71.778 & 320.129 & 71.674 & 309.546 & 73.372\\ 19 & 338.607 & 60.475 & 258.711 & 73.660 & 321.593 & 71.103 & 311.364 & 72.887\\ 20 & 335.793 & 62.184 & 253.502 & 76.427 & 314.402 & 73.620 & 304.195 & 75.174\\ \addlinespace 21 & 332.970 & 63.645 & 253.144 & 77.736 & 308.410 & 74.700 & 297.784 & 75.997\\ 22 & 330.800 & 65.336 & 257.465 & 80.224 & 311.429 & 78.767 & 301.049 & 80.315\\ 23 & 329.014 & 67.350 & 260.156 & 81.915 & 312.367 & 80.926 & 302.176 & 82.411\\ 24 & 322.179 & 69.431 & 241.037 & 85.618 & 298.344 & 85.438 & 288.837 & 88.097\\ 25 & 317.707 & 70.935 & 243.568 & 84.771 & 300.464 & 83.548 & 291.561 & 86.114\\ \addlinespace 26 & 311.248 & 72.452 & 232.526 & 85.976 & 292.966 & 83.881 & 284.772 & 87.216\\ 27 & 303.903 & 74.111 & 220.795 & 86.276 & 285.994 & 83.400 & 278.207 & 86.612\\ 28 & 296.689 & 76.137 & 221.057 & 87.198 & 284.432 & 87.044 & 276.125 & 91.097\\ 29 & 291.707 & 77.108 & 224.803 & 84.670 & 289.344 & 88.893 & 281.051 & 94.430\\ 30 & 286.512 & 78.763 & 224.286 & 87.708 & 289.866 & 88.773 & 281.567 & 94.511\\ \addlinespace 31 & 280.321 & 80.278 & 220.149 & 88.532 & 282.984 & 102.169 & 274.693 & 112.954\\ 32 & 274.475 & 81.569 & 219.312 & 88.392 & 283.489 & 104.453 & 275.361 & 118.062\\ 33 & 268.127 & 81.826 & 215.150 & 86.063 & 277.929 & 109.303 & 270.007 & 125.876\\ 34 & 261.724 & 82.416 & 212.403 & 86.393 & 278.083 & 114.313 & 270.006 & 133.578\\ 35 & 255.689 & 82.910 & 211.632 & 86.137 & 276.789 & 115.863 & 268.435 & 135.770\\ \addlinespace 36 & 249.054 & 83.211 & 210.872 & 85.521 & 276.622 & 112.920 & 268.352 & 131.987\\ \begin{table} \caption{Estimates for Figure 4B Table for Mean Earnings} \begin{tabular}[t]{rrrrrrrrr} \toprule \multicolumn{1}{c} & \multicolumn{2}{c}{Naive} & \multicolumn{2}{c}{Regression} & \multicolumn{2}{c}{Weighted} & \multicolumn{2}{c}{Double Robust} \\ \cmidrule(l{3pt}r{3pt}){2-3} \cmidrule(l{3pt}r{3pt}){4-5} \cmidrule(l{3pt}r{3pt}){6-7} \cmidrule(l{3pt}r{3pt}){8-9} t & Coefficient & Std. Dev. & Coefficient & Std. Dev. & Coefficient & Std. Dev. & Coefficient & Std. Dev.\\ \midrule 1 & 122.426 & 28.936 & 41.765 & 13.448 & 29.639 & 9.450 & 20.329 & 12.572\\ 2 & 217.535 & 30.086 & 131.052 & 18.305 & 150.678 & 18.934 & 139.714 & 20.316\\ 3 & 260.571 & 31.786 & 154.463 & 23.414 & 186.779 & 24.128 & 177.021 & 24.925\\ 4 & 284.399 & 33.461 & 172.690 & 26.975 & 225.073 & 27.993 & 215.102 & 28.273\\ 5 & 306.479 & 35.162 & 209.587 & 29.501 & 253.358 & 30.269 & 244.004 & 30.706\\ \addlinespace 6 & 327.145 & 36.572 & 238.845 & 31.520 & 279.654 & 32.487 & 270.099 & 32.875\\ 7 & 333.936 & 38.016 & 231.950 & 33.837 & 277.504 & 36.277 & 267.307 & 36.993\\ 8 & 341.688 & 38.843 & 244.669 & 34.613 & 298.461 & 36.608 & 287.519 & 37.418\\ 9 & 352.244 & 39.520 & 265.792 & 35.509 & 321.788 & 37.605 & 311.348 & 38.250\\ 10 & 355.966 & 40.238 & 259.683 & 36.782 & 313.781 & 39.434 & 303.257 & 40.125\\ \addlinespace 11 & 356.451 & 40.599 & 259.301 & 37.554 & 313.031 & 40.706 & 302.224 & 41.457\\ 12 & 353.884 & 41.254 & 249.100 & 39.417 & 306.777 & 42.873 & 296.049 & 43.640\\ 13 & 348.460 & 41.843 & 238.394 & 41.150 & 300.341 & 43.871 & 289.735 & 44.570\\ 14 & 348.260 & 42.234 & 254.535 & 40.972 & 317.187 & 43.262 & 306.680 & 43.972\\ 15 & 345.738 & 42.720 & 250.325 & 42.655 & 314.697 & 44.894 & 304.178 & 45.645\\ \addlinespace 16 & 342.980 & 43.228 & 247.981 & 43.300 & 315.251 & 45.311 & 304.500 & 46.119\\ 17 & 342.156 & 43.524 & 255.201 & 43.617 & 325.395 & 45.350 & 314.787 & 46.243\\ 18 & 340.180 & 43.882 & 252.269 & 44.282 & 320.129 & 45.656 & 309.546 & 46.495\\ 19 & 338.607 & 44.347 & 258.711 & 46.187 & 321.593 & 46.998 & 311.364 & 48.053\\ 20 & 335.793 & 44.796 & 253.502 & 46.940 & 314.402 & 47.603 & 304.195 & 48.574\\ \addlinespace 21 & 332.970 & 45.173 & 253.144 & 47.465 & 308.410 & 47.704 & 297.784 & 48.708\\ 22 & 330.800 & 45.577 & 257.465 & 47.894 & 311.429 & 48.206 & 301.049 & 49.289\\ 23 & 329.014 & 45.987 & 260.156 & 48.867 & 312.367 & 48.918 & 302.176 & 50.032\\ 24 & 322.179 & 46.487 & 241.037 & 49.798 & 298.344 & 49.705 & 288.837 & 50.831\\ 25 & 317.707 & 46.825 & 243.568 & 49.633 & 300.464 & 49.864 & 291.561 & 50.914\\ \addlinespace 26 & 311.248 & 47.327 & 232.526 & 50.343 & 292.966 & 50.139 & 284.772 & 51.126\\ 27 & 303.903 & 47.718 & 220.795 & 50.689 & 285.994 & 50.768 & 278.207 & 51.695\\ 28 & 296.689 & 48.032 & 221.057 & 50.458 & 284.432 & 50.408 & 276.125 & 51.385\\ 29 & 291.707 & 48.220 & 224.803 & 50.439 & 289.344 & 50.812 & 281.051 & 51.861\\ 30 & 286.512 & 48.461 & 224.286 & 50.288 & 289.866 & 50.538 & 281.567 & 51.553\\ \addlinespace 31 & 280.321 & 48.688 & 220.149 & 50.302 & 282.984 & 50.870 & 274.693 & 51.943\\ 32 & 274.475 & 48.892 & 219.312 & 50.083 & 283.489 & 50.832 & 275.361 & 51.899\\ 33 & 268.127 & 49.042 & 215.150 & 50.070 & 277.929 & 51.178 & 270.007 & 52.191\\ 34 & 261.724 & 49.308 & 212.403 & 50.119 & 278.083 & 51.068 & 270.006 & 52.159\\ 35 & 255.689 & 49.612 & 211.632 & 50.200 & 276.789 & 51.512 & 268.435 & 52.683\\ \addlinespace 36 & 249.054 & 49.960 & 210.872 & 50.231 & 276.622 & 51.378 & 268.352 & 52.523\\ \bottomrule \end{tabular} \end{table} ==========================================================================

We see that the naive estimator does very poorly. It takes more than 25 quarters before the naive estimator is within two standard errors of the experimental estimate. In contrast all three surrogate-based estimators are all within two standard errors when the surrogates include 5 quarters of outcomes, for both outcomes.

Validation Results and Other Supplementary Analyses

Given the data available we can also test whether using $t$ quarters of surrogates is sufficient to satisfy Surrogacy and Comparability. To test Surrogacy we regress the primary outcome on the pre-treatment variables, the surrogates up to quarter $t$, and the indicator for the treatment; a finding that the treatment has an impact indicates a violation of Surrogacy. We estimate this regression using a logistic regression model, using only the data from the experimental (Riverside) sample. We report in Table (ref) and (ref) the results from these regressions for a number of different values for $t$, for the employment outcome and the earnings outcome. We report the point estimate, standard error and t-statistic. We see that point estimates for $t\leq 3$ are large and highly statistically significant. After that most of the t-statistics are less than 2, although there are some where the t-statistics are a little above 2, but the coefficient estimates are small.

We do a similar exercise for Comparability. We combine the experimental and observational samples and regress the final outcome on the surrogates, the pretreatment variables, and an indicator for the experimental sample, again using surrogates up to period $t$. We report the estimates on the indicator for the experimental sample, and the corresponding standard error. Here the the point estimates become smaller after $t=12$, but the t-statistics remain large even with a substantial number of surrogate periods, indicating a violation of Surrogacy.

table[table omitted — 1,119 chars of source]
comment====================================================================== 1 & 0.050 & 0.010 & 5.164 & 0.009 & 0.005 & 2.009\\ 2 & 0.032 & 0.009 & 3.429 & -0.004 & 0.004 & -0.828\\ 3 & 0.021 & 0.009 & 2.385 & -0.006 & 0.004 & -1.423\\ 4 & 0.016 & 0.009 & 1.789 & -0.007 & 0.004 & -1.738\\ 5 & 0.007 & 0.008 & 0.905 & -0.010 & 0.004 & -2.501\\ \addlinespace 6 & 0.002 & 0.008 & 0.289 & -0.011 & 0.004 & -2.994\\ 7 & 0.002 & 0.008 & 0.291 & -0.014 & 0.004 & -3.926\\ 8 & -0.002 & 0.007 & -0.213 & -0.016 & 0.004 & -4.561\\ 9 & -0.006 & 0.007 & -0.927 & -0.016 & 0.003 & -4.873\\ 10 & -0.006 & 0.007 & -0.864 & -0.015 & 0.003 & -4.802\\ \addlinespace 11 & -0.006 & 0.006 & -0.867 & -0.015 & 0.003 & -5.048\\ 12 & -0.006 & 0.006 & -0.928 & -0.015 & 0.003 & -4.955\\ 13 & -0.004 & 0.006 & -0.693 & -0.013 & 0.003 & -4.570\\ 14 & -0.007 & 0.005 & -1.272 & -0.012 & 0.003 & -4.674\\ 15 & -0.007 & 0.005 & -1.387 & -0.013 & 0.002 & -4.994\\ \addlinespace 16 & -0.005 & 0.005 & -1.100 & -0.010 & 0.002 & -4.281\\ 17 & -0.008 & 0.005 & -1.712 & -0.011 & 0.002 & -4.631\\ 18 & -0.007 & 0.004 & -1.714 & -0.009 & 0.002 & -4.060\\ 19 & -0.007 & 0.004 & -1.725 & -0.009 & 0.002 & -4.523\\ 20 & -0.008 & 0.004 & -2.114 & -0.006 & 0.002 & -3.485\\ \addlinespace 21 & -0.007 & 0.004 & -2.045 & -0.005 & 0.002 & -3.192\\ 22 & -0.007 & 0.003 & -2.114 & -0.005 & 0.002 & -3.428\\ 23 & -0.007 & 0.003 & -2.184 & -0.006 & 0.002 & -3.782\\ 24 & -0.005 & 0.003 & -1.958 & -0.005 & 0.001 & -3.596\\ 25 & -0.006 & 0.003 & -2.401 & -0.004 & 0.001 & -3.194\\ \addlinespace 26 & -0.003 & 0.002 & -1.415 & -0.003 & 0.001 & -2.795\\ 27 & -0.002 & 0.002 & -0.757 & -0.003 & 0.001 & -3.412\\ 28 & -0.001 & 0.002 & -0.334 & -0.002 & 0.001 & -2.747\\ 29 & -0.002 & 0.002 & -1.154 & -0.002 & 0.001 & -3.108\\ 30 & -0.002 & 0.001 & -1.264 & -0.001 & 0.001 & -1.900\\ \addlinespace 31 & -0.002 & 0.001 & -1.944 & 0.000 & 0.000 & -0.827\\ 32 & -0.002 & 0.001 & -2.111 & 0.000 & 0.000 & -0.824\\ 33 & -0.001 & 0.001 & -1.891 & 0.000 & 0.000 & -0.895\\ 34 & 0.000 & 0.000 & -0.752 & 0.000 & 0.000 & -1.040\\ 35 & 0.000 & 0.000 & -2.114 & 0.000 & 0.000 & -0.887\\ \addlinespace \begin{table} \caption{Assumptions Tests for T=35 Surrogates Using Mean Employment Rate} \begin{tabular}[t]{rrrrrrr} \toprule \multicolumn{1}{c} & \multicolumn{3}{c}{Surrogate Test} & \multicolumn{3}{c}{Comparability Test} \\ \cmidrule(l{3pt}r{3pt}){2-4} \cmidrule(l{3pt}r{3pt}){5-7} t & Coefficient & Std. Error & T-Statistic & Coefficient & Std. Error & T-Statistic\\ \midrule 1 & 0.052 & 0.010 & 5.392 & 0.008 & 0.005 & 1.664\\ 2 & 0.034 & 0.009 & 3.692 & -0.004 & 0.004 & -0.937\\ 3 & 0.024 & 0.009 & 2.631 & -0.006 & 0.004 & -1.481\\ 4 & 0.018 & 0.009 & 2.049 & -0.007 & 0.004 & -1.764\\ 5 & 0.010 & 0.008 & 1.156 & -0.010 & 0.004 & -2.517\\ \addlinespace 6 & 0.004 & 0.008 & 0.513 & -0.011 & 0.004 & -3.002\\ 7 & 0.004 & 0.008 & 0.501 & -0.014 & 0.004 & -3.907\\ 8 & 0.000 & 0.007 & -0.031 & -0.016 & 0.004 & -4.559\\ 9 & -0.005 & 0.007 & -0.744 & -0.016 & 0.003 & -4.867\\ 10 & -0.005 & 0.007 & -0.681 & -0.015 & 0.003 & -4.804\\ \addlinespace 11 & -0.004 & 0.006 & -0.661 & -0.015 & 0.003 & -5.043\\ 12 & -0.004 & 0.006 & -0.729 & -0.015 & 0.003 & -4.962\\ 13 & -0.003 & 0.006 & -0.504 & -0.013 & 0.003 & -4.567\\ 14 & -0.006 & 0.005 & -1.099 & -0.012 & 0.003 & -4.666\\ 15 & -0.006 & 0.005 & -1.216 & -0.013 & 0.002 & -4.990\\ \addlinespace 16 & -0.005 & 0.005 & -0.936 & -0.010 & 0.002 & -4.314\\ 17 & -0.007 & 0.005 & -1.559 & -0.011 & 0.002 & -4.671\\ 18 & -0.007 & 0.004 & -1.618 & -0.009 & 0.002 & -4.120\\ 19 & -0.007 & 0.004 & -1.618 & -0.009 & 0.002 & -4.548\\ 20 & -0.008 & 0.004 & -1.998 & -0.007 & 0.002 & -3.522\\ \addlinespace 21 & -0.007 & 0.004 & -1.915 & -0.005 & 0.002 & -3.192\\ 22 & -0.007 & 0.003 & -1.998 & -0.005 & 0.002 & -3.434\\ 23 & -0.006 & 0.003 & -2.047 & -0.006 & 0.002 & -3.765\\ 24 & -0.005 & 0.003 & -1.823 & -0.005 & 0.001 & -3.578\\ 25 & -0.006 & 0.003 & -2.281 & -0.004 & 0.001 & -3.163\\ \addlinespace 26 & -0.003 & 0.002 & -1.283 & -0.003 & 0.001 & -2.765\\ 27 & -0.001 & 0.002 & -0.646 & -0.003 & 0.001 & -3.384\\ 28 & 0.000 & 0.002 & -0.260 & -0.002 & 0.001 & -2.714\\ 29 & -0.002 & 0.002 & -1.100 & -0.002 & 0.001 & -3.076\\ 30 & -0.002 & 0.001 & -1.277 & -0.001 & 0.001 & -1.899\\ \addlinespace 31 & -0.002 & 0.001 & -1.941 & 0.000 & 0.000 & -0.815\\ 32 & -0.002 & 0.001 & -2.052 & 0.000 & 0.000 & -0.826\\ 33 & -0.001 & 0.001 & -1.805 & 0.000 & 0.000 & -0.873\\ 34 & 0.000 & 0.000 & -0.621 & 0.000 & 0.000 & -1.052\\ 35 & 0.000 & 0.000 & -1.980 & 0.000 & 0.000 & -0.831\\ \addlinespace 36 & 0.000 & 0.000 & 0.292 & 0.000 & 0.000 & 1.047\\ \bottomrule \end{tabular} \end{table} ======================================================================
table[table omitted — 1,024 chars of source]
comment====================================================================== 1 & 186.410 & 50.854 & 3.666 & -34.395 & 25.329 & -1.358\\ 2 & 130.408 & 49.914 & 2.613 & -64.214 & 24.493 & -2.622\\ 3 & 96.164 & 48.083 & 2.000 & -72.426 & 23.455 & -3.088\\ 4 & 67.097 & 46.293 & 1.449 & -66.778 & 22.438 & -2.976\\ 5 & 42.708 & 44.466 & 0.961 & -70.601 & 21.510 & -3.282\\ \addlinespace 6 & 13.339 & 42.027 & 0.317 & -73.464 & 20.486 & -3.586\\ 7 & 15.311 & 40.185 & 0.381 & -86.537 & 19.484 & -4.441\\ 8 & -0.744 & 38.736 & -0.019 & -88.033 & 18.740 & -4.697\\ 9 & -25.816 & 36.763 & -0.702 & -78.670 & 17.863 & -4.404\\ 10 & -22.147 & 35.291 & -0.628 & -72.442 & 17.036 & -4.252\\ \addlinespace 11 & -23.614 & 33.015 & -0.715 & -68.797 & 16.041 & -4.289\\ 12 & -19.398 & 31.279 & -0.620 & -63.828 & 15.293 & -4.174\\ 13 & -16.080 & 29.862 & -0.538 & -55.964 & 14.482 & -3.864\\ 14 & -34.704 & 27.897 & -1.244 & -47.942 & 13.673 & -3.506\\ 15 & -32.738 & 26.527 & -1.234 & -49.873 & 13.074 & -3.815\\ \addlinespace 16 & -34.959 & 25.248 & -1.385 & -40.474 & 12.293 & -3.292\\ 17 & -46.320 & 23.514 & -1.970 & -38.620 & 11.575 & -3.336\\ 18 & -42.893 & 22.038 & -1.946 & -29.700 & 10.833 & -2.742\\ 19 & -44.803 & 20.429 & -2.193 & -33.687 & 10.083 & -3.341\\ 20 & -41.055 & 18.978 & -2.163 & -31.907 & 9.270 & -3.442\\ \addlinespace 21 & -35.547 & 17.734 & -2.004 & -29.526 & 8.585 & -3.439\\ 22 & -38.443 & 16.189 & -2.375 & -30.449 & 7.911 & -3.849\\ 23 & -38.034 & 14.872 & -2.557 & -28.860 & 7.210 & -4.003\\ 24 & -21.380 & 13.618 & -1.570 & -26.629 & 6.669 & -3.993\\ 25 & -25.815 & 12.270 & -2.104 & -20.063 & 6.016 & -3.335\\ \addlinespace 26 & -15.305 & 10.780 & -1.420 & -17.903 & 5.337 & -3.354\\ 27 & -7.148 & 9.354 & -0.764 & -17.773 & 4.618 & -3.849\\ 28 & -5.179 & 8.304 & -0.624 & -12.169 & 3.987 & -3.052\\ 29 & -11.421 & 6.956 & -1.642 & -9.068 & 3.353 & -2.704\\ 30 & -10.905 & 5.876 & -1.856 & -4.565 & 2.831 & -1.613\\ \addlinespace 31 & -6.467 & 4.989 & -1.296 & -1.912 & 2.394 & -0.799\\ 32 & -6.417 & 3.931 & -1.633 & 0.250 & 1.908 & 0.131\\ 33 & -3.043 & 2.952 & -1.031 & 0.896 & 1.406 & 0.638\\ 34 & -1.492 & 1.907 & -0.783 & -0.431 & 0.908 & -0.474\\ 35 & -0.571 & 0.986 & -0.579 & -0.368 & 0.471 & -0.781\\ \addlinespace 36 & 0.000 & 0.000 & -1.375 & 0.000 & 0.000 & -2.582\\ \begin{table} \caption{Assumptions Tests for T=35 Surrogates Using Mean Earnings} \begin{tabular}[t]{rrrrrrr} \toprule \multicolumn{1}{c} & \multicolumn{3}{c}{Surrogate Test} & \multicolumn{3}{c}{Comparability Test} \\ \cmidrule(l{3pt}r{3pt}){2-4} \cmidrule(l{3pt}r{3pt}){5-7} t & Coefficient & Std. Error & T-Statistic & Coefficient & Std. Error & T-Statistic\\ \midrule 1 & 185.591 & 50.846 & 3.650 & -35.598 & 25.328 & -1.406\\ 2 & 129.687 & 49.909 & 2.598 & -65.041 & 24.495 & -2.655\\ 3 & 94.315 & 48.076 & 1.962 & -72.760 & 23.462 & -3.101\\ 4 & 66.255 & 46.287 & 1.431 & -67.481 & 22.443 & -3.007\\ 5 & 42.236 & 44.471 & 0.950 & -71.039 & 21.516 & -3.302\\ \addlinespace 6 & 12.644 & 42.037 & 0.301 & -73.901 & 20.496 & -3.606\\ 7 & 14.813 & 40.207 & 0.368 & -86.414 & 19.493 & -4.433\\ 8 & -1.069 & 38.758 & -0.028 & -88.522 & 18.748 & -4.722\\ 9 & -25.914 & 36.782 & -0.705 & -79.310 & 17.870 & -4.438\\ 10 & -22.179 & 35.309 & -0.628 & -73.109 & 17.043 & -4.290\\ \addlinespace 11 & -22.990 & 33.031 & -0.696 & -69.671 & 16.047 & -4.342\\ 12 & -19.070 & 31.285 & -0.610 & -65.189 & 15.295 & -4.262\\ 13 & -15.928 & 29.866 & -0.533 & -57.274 & 14.483 & -3.954\\ 14 & -34.218 & 27.903 & -1.226 & -49.196 & 13.674 & -3.598\\ 15 & -32.214 & 26.529 & -1.214 & -51.217 & 13.073 & -3.918\\ \addlinespace 16 & -34.161 & 25.250 & -1.353 & -41.955 & 12.292 & -3.413\\ 17 & -44.861 & 23.499 & -1.909 & -40.170 & 11.569 & -3.472\\ 18 & -41.837 & 22.032 & -1.899 & -31.244 & 10.827 & -2.886\\ 19 & -43.781 & 20.416 & -2.144 & -34.739 & 10.078 & -3.447\\ 20 & -39.759 & 18.957 & -2.097 & -33.118 & 9.263 & -3.575\\ \addlinespace 21 & -33.746 & 17.703 & -1.906 & -30.465 & 8.575 & -3.553\\ 22 & -37.096 & 16.163 & -2.295 & -31.396 & 7.900 & -3.974\\ 23 & -36.703 & 14.834 & -2.474 & -29.610 & 7.200 & -4.113\\ 24 & -20.185 & 13.586 & -1.486 & -27.410 & 6.660 & -4.116\\ 25 & -24.737 & 12.254 & -2.019 & -20.877 & 6.009 & -3.474\\ \addlinespace 26 & -14.549 & 10.759 & -1.352 & -18.560 & 5.332 & -3.481\\ 27 & -6.565 & 9.345 & -0.703 & -18.120 & 4.617 & -3.925\\ 28 & -4.800 & 8.300 & -0.578 & -12.444 & 3.987 & -3.121\\ 29 & -11.060 & 6.956 & -1.590 & -9.289 & 3.353 & -2.770\\ 30 & -10.793 & 5.877 & -1.837 & -4.655 & 2.832 & -1.644\\ \addlinespace 31 & -6.311 & 4.989 & -1.265 & -1.907 & 2.395 & -0.796\\ 32 & -6.234 & 3.928 & -1.587 & 0.266 & 1.909 & 0.140\\ 33 & -2.878 & 2.952 & -0.975 & 0.906 & 1.407 & 0.644\\ 34 & -1.369 & 1.906 & -0.719 & -0.451 & 0.909 & -0.496\\ 35 & -0.535 & 0.986 & -0.543 & -0.362 & 0.471 & -0.768\\ \addlinespace 36 & 0.000 & 0.000 & -1.340 & 0.000 & 0.000 & -2.543\\ \bottomrule \end{tabular} \end{table} $===================================$

If we are unwilling to make the surrogacy assumption we can still calculate bounds for the effect on employment, using the fact that this outcome is binary. For the case where the first six quarters of post-treatment data are used as surrogates, the lower and upper bound are estimated as -0.186 and 0.124. These are not very informative, because the data now do not allow us to estimate the indirect effect of the treatment on the outcome.

Similarly, we can calculate bounds for the average effect without assuming Comparability. With six quarters of surrogates the bounds are again wide at -0.076 and 0.194 respectively. Here the fact that the treatment effect on the surrogates is strong leads to substantial sensitivity to the comparability assumption as formalized in Lemma (ref).

Using the data for Riverside we can also assess the value of the Surrogacy assumption. Using the six quarters of data as surrogates, we find that the gain from knowledge of Surrogacy (the $\Delta$ in Theorem (ref)) is quite large. The standard error given Surrogacy, $\sqrt{\mathbb{V}}_{\rm s}$, is 0.33 times the standard error without knowledge that Surrogacy holds, $\sqrt{\mathbb{V}}_{\rm ns}$.

Conclusion

We develop new methods for combining intermediate outcomes to estimate the long-term impacts of treatments more rapidly and precisely. Our method requires estimating a “surrogate index” -- the conditional expectation of the long-term outcome given intermediate outcomes -- and then estimating the treatment effect on the surrogate index. The surrogate index can be estimated using parametric or nonparametric regression methods. We formalize conditions under which this method yields unbiased estimates, derive bounds for the degree of bias when those assumptions fail, and propose a simple out-of-sample validation approach using “hold out” intermediate outcomes. We show that surrogates can also greatly improve the precision of estimates even in settings where the treatment effect on the long-term outcome can be estimated directly, particularly when that outcome is rare or noisy.

Applying the method to analyze the impacts of the GAIN job training program in California, we find that using short-term earnings and employment rates to construct surrogate indices expedite the detection of long-term treatment effects on employment and earnings by several years and also substantially increases precision. Furthermore, a single surrogate index accurately predicts heterogeneity in the long-term treatment effects of different types of job training programs across sites, showing that surrogate indices estimated in a given setting may be generalizable to other settings. The success of the surrogate index in this application validates the use of short-term employment outcomes as surrogates for detecting longer-term impacts of job training programs, an empirical result that can be applied when analyzing ongoing programs.

Building on this application, it would be useful to systematically establish surrogate indices that match the long-term treatment effects estimated in other experiments and quasi-experiments. Over time, this would allow researchers to collectively build a public library of surrogate indices for long-term outcomes that could be used to expedite the analysis of future interventions.

\vskip0.5cm

comment\subsection{Additional Table} \begin{table}[!htbp] \caption{Summary Statistics of Covariates by Location} \begin{tabular}{@{\extracolsep{3pt}} cccccc} \\[-1.8ex]\hline \hline \\[-1.8ex] & \multicolumn{2}{c}{Riverside $(N_{\rm E}=5,445)$} & \multicolumn{2}{c}{Other Locations ($N_{\rm O}= 13,725$)}\\ & Mean & (Std. Dev.) & Mean & (Std. Dev.) & t-statistic\\ \hline \\[-1.8ex] \hline \\[-1.8ex] \end{tabular} \end{table} \subsubsection{The Critical Assumptions in the Mediation Literature and their Relation to Surrogacy} In the mediation literature (\textit{e.g.,} \citealt{baron1986moderator,vanderweele2015explanation}), the intermediate outcome that we refer to here as the surrogate $S_{i}$ is called a mediator. To emphasize its role as a causal variable in the mediation literature, we expand the notation and consider potential outcomes $Y_{i}(w,s)$ that are indexed by the treatment and the surrogate. (In terms of these potential outcomes the original potential outcomes defined in the previous section, $Y_{i}(w)$, indexed only by the treatment $W_{i}$, equals $Y_{i}(w)=Y_{i}(w,S_{i}(w))$, for $w\in\mathbb{W}$.) In the setting considered in the mediation literature, we observe the quadruple $(Y_{i},S_{i},W_{i},X_{i},P_{i})$ for all units in the sample and so there is not necessarily a distinction between the experimental sample and the observational sample. To capture that we focus in this section on the case where we only have the experimental sample, $P_{i}={\rm E}$, and where we observe the primary outcome $Y_{i}$ for this sample. The focus in the mediation literature is on decomposing the causal effect of the treatment on the outcome into a direct effect that involves comparing potential outcomes where the surrogate remains fixed, and an indirect effect that passes through the mediator/surrogate. Three key estimands are the average \textit{total effect,} \[ \tau^{{\rm total}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(1))-Y_{i}(0,S_{i}(0))\right], \] the average natural indirect effect, where we fix the treatment at $w=1$, but change the surrogate from $S_{i}(0)$ to $S_{i}(1)$, \[ \tau^{{\rm nie}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(1))-Y_{i}(1,S_{i}(0))\right], \] and the average natural direct effect, where we fix the surrogate at $S_{i}(0)$ and change the treatment from $W_{i}=0$ to $W_{i}=1$: \[ \tau^{{\rm nde}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(0))-Y_{i}(0,S_{i}(0))\right], \] with the the latter two adding up to the first: $\tau^{{\rm total}}=\tau^{{\rm nie}}+\tau^{{\rm nde}}$. These effects are identified in the mediation literature using assumptions similar to Assumptions (ref) and (ref). The first assumption in the mediation framework is a reformulation of the unconfoundedness assumption, Assumption (ref). It rules out the presence of unmeasured confounders between the treatment and the surrogate, and between the treatment and the outcome.\begin{assumption} (Unconfounded Treatment Assignment / Strong Ignorability) \\ $(i)$ $W_{i}\ \perp\!\!\!\perp\ \Bigl(S_{i}(0),S_{i}(1),Y_{i}(0,S_{i}(0)),Y_{i}(1,S_{i}(1))\Bigr)\ \Bigr|\ X_{i},P_{i}={\rm E}, $\\ $(ii)$ $0<\rho(x)<1\ {\rm for\ all}\ x\in\mathbb{X}.$ \end{assumption} The second assumption typically made in the mediation literature is another unconfoundedness assumption that rules out the presence of unobserved confounders between the surrogate and the outcome, conditional on the treatment. \begin{assumption} \[ S_{i}\ \perp\!\!\!\perp\ \Bigl(Y_{i}(W_{i},s)_{s\in\mathbb{S}}\Bigr)\ \Bigr|\ W_{i},X_{i},P_{i}. \] \end{assumption} This assumption implies that comparisons of primary outcomes for units with different values for the surrogates but identical values for the treatment and pre-treatment variables can be given a causal interpretation. To make the link to the surrogacy literature we need to add one key assumption that is not commonly made in the mediation literature. This assumption rules out any direct effect of the treatment on the outcome, allowing only for an indirect effect through the surrogate.\begin{assumption} For all $i$, $w,w'\in\mathbb{W},s\in\mathbb{S}$, \[ Y_{i}(w,s)=Y_{i}(w',s). \] \end{assumption} This assumption is similar to the exclusion restriction in instrumental variables settings, e.g., imbens1994,angrist1996identification. In combination with the previous assumption this implies that we can give comparisons in the primary outcome between units with different values for the surrogates but the same values for pre-treatment variables a causal interpretation, without knowing the treatment status. The following proposition links the surrogacy and mediation assumptions. \begin{prop} Suppose Assumptions (ref)-(ref) hold. Then Assumptions (ref) and (ref) hold. \end{prop} This connection highlights that at the heart of the surrogacy assumption is a causal relation between the surrogate and the primary outcome that mediates the causal effect of the treatment on the outcome. \subsubsection{Surrogacy and Comparability from a Missing Data Perspective} From a missing data perspective, the surrogacy and comparability assumptions we make have parallels to the missingness at random (MAR) assumption common in the missing data literature (rubin1976inference,little2019statistical), and specifically the literature on combining samples with different sets of variables, (ridder2007econometrics,gelman1998not,rassler2004data, graham2016efficient). In particular rassler2012statistical focuses on a missing data structure closely related to ours. In our two sample setting, we can think of the complete data as the quintuple $(Y_{i},S_{i},W_{i},X_{i},P_{i})$. Here, we view the sample as randomly drawn from a large population, so that we view $P_{i}$ as a stochastic missing data indicator. For the units in the sample we observe the incomplete data $(\boldsymbol{1}_{P_{i}={\rm O}}Y_{i},S_{i},X_{i},{\boldsymbol{1}}_{P_{i}={\rm E}}W_{i},P_{i})$, where for units with $P_{i}={\rm O}$ the treatment indicator $W_{i}$ is missing, and for units with $P_{i}={\rm E}$ the outcome $Y_{i}$ is missing. Now consider the following assumption. \begin{assumption}(Augmented Missing At Random Assumption)\\ Conditional on $(S_{i},X_{i})$, the three variables $P_{i}$, $Y_{i}$ and $W_{i}$ are jointly independent: \[ P_{i}\ \perp\!\!\!\perp\ Y_{i}\ \perp\!\!\!\perp\ W_{i}\ \Bigr|\ S_{i},X_{i}. \] \end{assumption} This is slightly different from a standard MAR assumption in rubin1976inference where one would assume $P_{i}\perp\!\!\!\perp Y_{i}|S_{i},X_{i}$ and/or $P_{i}\perp\!\!\!\perp W_{i}|S_{i},X_{i}$. We need the stronger assumption to incorporate surrogacy, as the following proposition shows. \begin{prop}\textsc{(Missing Data Model)}\\ $(i)$ Assumption (ref) implies Assumption (ref) (Surrogacy) \[ Y_{i}\ \perp\!\!\!\perp\ W_{i}\ \Bigr|\ S_{i},X_{i}, \] and Assumption (ref) (Comparability) \[ P_{i}\ \perp\!\!\!\perp\ Y_{i}\ \Bigr|\ S_{i},X_{i}. \] $(ii)$ Assumption (ref) has no testable implications. \end{prop} Note that even after we have dealt with the missing $Y_{i}$ and missing $W_{i}$ problems, we still have the missing potential outcomes, which is why we also need the unconfoundedness assumption.

\vskip0.5cm

{}

comment{ }

{ }

center[center omitted — 66 chars of source]

{A}.\ Additional Table

table[table omitted — 2,158 chars of source]

{B}.\ Related Literature

Critical Assumptions in the Mediation Literature and their Relation to Surrogacy

In the mediation literature (e.g., baron1986moderator,vanderweele2015explanation), the intermediate outcome that we refer to here as the surrogate $S_{i}$ is called a mediator. To emphasize its role as a causal variable in the mediation literature, we expand the notation and consider potential outcomes $Y_{i}(w,s)$ that are indexed by the treatment and the surrogate. (In terms of these potential outcomes the original potential outcomes defined in the previous section, $Y_{i}(w)$, indexed only by the treatment $W_{i}$, equals $Y_{i}(w)=Y_{i}(w,S_{i}(w))$, for $w\in\mathbb{W}$.) In the setting considered in the mediation literature, we observe the quadruple $(Y_{i},S_{i},W_{i},X_{i},P_{i})$ for all units in the sample and so there is not necessarily a distinction between the experimental sample and the observational sample. To capture that we focus in this section on the case where we only have the experimental sample, $P_{i}={\rm E}$, and where we observe the primary outcome $Y_{i}$ for this sample.

The focus of the mediation literature is on decomposing the causal effect of the treatment on the outcome into a direct effect that involves comparing potential outcomes where the surrogate remains fixed, and an indirect effect that passes through the mediator/surrogate. Three key estimands are the average total effect, \[ \tau^{{\rm total}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(1))-Y_{i}(0,S_{i}(0))\right], \] the average natural indirect effect, where we fix the treatment at $w=1$, but change the surrogate from $S_{i}(0)$ to $S_{i}(1)$, \[ \tau^{{\rm nie}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(1))-Y_{i}(1,S_{i}(0))\right], \] and the average natural direct effect, where we fix the surrogate at $S_{i}(0)$ and change the treatment from $W_{i}=0$ to $W_{i}=1$: \[ \tau^{{\rm nde}}\equiv\mathbb{E}\left[Y_{i}(1,S_{i}(0))-Y_{i}(0,S_{i}(0))\right], \] with the latter two adding up to the first: $\tau^{{\rm total}}=\tau^{{\rm nie}}+\tau^{{\rm nde}}$.

These effects are identified in the mediation literature using assumptions similar to Assumptions (ref) and (ref). The first assumption in the mediation framework is a reformulation of the unconfoundedness assumption, Assumption (ref). It rules out the presence of unmeasured confounders between the treatment and the surrogate, and between the treatment and the outcome.

assumption(Unconfounded Treatment Assignment / Strong Ignorability) \\ $(i)$ $W_{i}\ \perp\!\!\!\perp\ \Bigl(S_{i}(0),S_{i}(1),Y_{i}(0,S_{i}(0)),Y_{i}(1,S_{i}(1))\Bigr)\ \Bigr|\ X_{i},P_{i}={\rm E}, $\\ $(ii)$ $0<\rho(x)<1\ {\rm for\ all}\ x\in\mathbb{X}.$

The second assumption typically made in the mediation literature is another unconfoundedness assumption that rules out the presence of unobserved confounders between the surrogate and the outcome, conditional on the treatment.

assumption\[ S_{i}\ \perp\!\!\!\perp\ \Bigl(Y_{i}(W_{i},s)_{s\in\mathbb{S}}\Bigr)\ \Bigr|\ W_{i},X_{i},P_{i}. \]

This assumption implies that comparisons of primary outcomes for units with different values for the surrogates but identical values for the treatment and pre-treatment variables can be given a causal interpretation.

To make the link to the surrogacy literature we need to add one key assumption that is not commonly made in the mediation literature. This assumption rules out any direct effect of the treatment on the outcome, allowing only for an indirect effect through the surrogate.

assumptionFor all $i$, $w,w'\in\mathbb{W},s\in\mathbb{S}$, \[ Y_{i}(w,s)=Y_{i}(w',s). \]

This assumption is similar to the exclusion restriction in instrumental variables settings, e.g., imbens1994,angrist1996identification. In combination with the previous assumption this implies that we can give comparisons in the primary outcome between units with different values for the surrogates but the same values for pre-treatment variables a causal interpretation, without knowing the treatment status.

The following proposition links the surrogacy and mediation assumptions.

propSuppose Assumptions (ref)-(ref) hold. Then Assumptions (ref) and (ref) hold.

This connection highlights that at the heart of the surrogacy assumption is a causal relation between the surrogate and the primary outcome that mediates the causal effect of the treatment on the outcome.

Surrogacy and Comparability from a Missing Data Perspective

From a missing data perspective, Surrogacy and Comparability have parallels to the missingness at random (MAR) assumption common in the missing data literature (rubin1976inference,little2019statistical), and specifically the literature on combining samples with different sets of variables, (ridder2007econometrics,gelman1998not,rassler2004data, graham2016efficient). In particular rassler2012statistical focuses on a missing data structure closely related to ours.

In our two sample setting, we can think of the complete data as the quintuple $(Y_{i},S_{i},W_{i},X_{i},P_{i})$. Here, we view the sample as randomly drawn from a large population, so that we view $P_{i}$ as a stochastic missing data indicator. For the units in the sample we observe the incomplete data $(\boldsymbol{1}_{P_{i}={\rm O}}Y_{i},S_{i},X_{i},{\boldsymbol{1}}_{P_{i}={\rm E}}W_{i},P_{i})$, where for units with $P_{i}={\rm O}$ the treatment indicator $W_{i}$ is missing, and for units with $P_{i}={\rm E}$ the outcome $Y_{i}$ is missing. Now consider the following assumption.

assumption(Augmented Missing At Random Assumption)\\ Conditional on $(S_{i},X_{i})$, the three variables $P_{i}$, $Y_{i}$ and $W_{i}$ are jointly independent: \[ P_{i}\ \perp\!\!\!\perp\ Y_{i}\ \perp\!\!\!\perp\ W_{i}\ \Bigr|\ S_{i},X_{i}. \]

This is slightly different from a standard MAR assumption in rubin1976inference where one would assume $P_{i}\perp\!\!\!\perp Y_{i}|S_{i},X_{i}$ and/or $P_{i}\perp\!\!\!\perp W_{i}|S_{i},X_{i}$. We need the stronger assumption to incorporate surrogacy, as the following proposition shows.

prop(Missing Data Model)\\ $(i)$ Assumption (ref) implies Assumption (ref) (Surrogacy) \[ Y_{i}\ \perp\!\!\!\perp\ W_{i}\ \Bigr|\ S_{i},X_{i}, \] and Assumption (ref) (Comparability) \[ P_{i}\ \perp\!\!\!\perp\ Y_{i}\ \Bigr|\ S_{i},X_{i}. \] $(ii)$ Assumption (ref) has no testable implications.

Note that even after we have dealt with the missing $Y_{i}$ and missing $W_{i}$ problems, we still have the missing potential outcomes, which is why we also need the unconfoundedness assumption.

C. Proofs

Proof of Proposition (ref): \[ {\rm pr}\left(W_{i}=1|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right)=\mathbb{E}\left[\left.W_{i}\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right] \] \[ \hskip1cm=\mathbb{E}\left[\left.\mathbb{E}\left[\left.W_{i}\right|Y_{i}=y,S_{i},X_{i},\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right]\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right] \] \[ \hskip1cm=\mathbb{E}\left[\left.\mathbb{E}\left[\left.W_{i}\right|Y_{i}=y,S_{i},X_{i},P_{i}={\rm E}\right]\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right] \] \[ \hskip1cm=\mathbb{E}\left[\left.\mathbb{E}\left[\left.W_{i}\right|S_{i},X_{i},P_{i}={\rm E}\right]\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right] \] \[ \hskip1cm=\mathbb{E}\left[\left.\rho(S_{i},X_{i})\right|Y_{i}=y,\rho(S_{i},X_{i})=r,P_{i}={\rm E}\right]=\rho(S_{i},X_{i}), \] which proves the result. $\square$

\vskip0.5cm

Proof of Proposition (ref): Part $(i)$ follows directly from the definitions of $\mu(\cdot,{\rm E})$ and Assumption (ref). Part $(ii)$ follows directly from the definitions of $\mu(\cdot,{\rm E})$ and $\mu(\cdot,{\rm O})$ and Assumption (ref). Part $(iii)$ follows from parts $(i)$ and $(ii)$. $\square$

\vskip0.5cm

Proof of Proposition (ref): We wish to show that the three conditions

equation[equation omitted — 130 chars of source]
equation[equation omitted — 114 chars of source]

and

equation[equation omitted — 100 chars of source]

imply

equation[equation omitted — 117 chars of source]
equation[equation omitted — 80 chars of source]

Note that we leave out the conditioning in $P_{i}={\rm E}$ in the last two conditions because we are focused here on the one-sample case. Condition ((ref)) follows directly from ((ref)) because $Y_{i}(w)=Y_{i}(w,S_{i}(w))$.

Condition ((ref)) implies that we can write $Y_{i}(s)$ without ambiguity, and by ((ref)), we have $ W_{i}\ \perp\!\!\!\perp\ Y_{i}(s)\ \Bigr|\ X_{i}. $ By ((ref)) we have $S_{i}\ \perp\!\!\!\perp\ Y_{i}(s)\ \Bigr|\ X_{i},W_{i}. $ Combining these implies $ \Bigl(S_{i},W_{i}\Bigr)\ \perp\!\!\!\perp\ Y_{i}(s)\ \Bigr|\ X_{i}. $ This in turn implies $ W_{i}\ \perp\!\!\!\perp\ Y_{i}(s)\ \Bigr|\ S_{i},X_{i}, $ which in turn implies $ W_{i}\ \perp\!\!\!\perp\ Y_{i}(S_{i})\ \Bigr|\ S_{i},X_{i}. $ This is equivalent to the condition we set out to prove, $ W_{i}\ \perp\!\!\!\perp\ Y_{i}\ \Bigr|\ S_{i},X_{i}. $ $\square$

\vskip0.5cm

Proof of Proposition (ref): The first part of the Proposition is immediate. For the second part, note that we can identify from the data the distributions \[ f_{Y_i|S_i,X_i,P_i}(y|s,x,{\rm O}),\hskip1cmf_{W_i|S_i,X_i,P_i}(w|s,x,{\rm E}),\hskip1cm{\rm and}\ \ f_{P_i,S_i,X_i}(p,s,x), \] but no other distributions. That implies that the joint distribution of $(Y_i,S_i,W_i,X_i,P_i)$ implied by $ f_{Y_i|S_i,W_i,X_i,P_i}(y|s,w,x,p)=f_{Y_i|S_i,X_i,P_i}(y|s,x,{\rm O}), $ and $ f_{W_i|S_i,X_i,P_i}(w|s,x,{\rm O})=f_{W_i|S_i,X_i,P_i}(w|s,x,{\rm E}), $ for all $(y,s,s,w,x,p)$ is consistent with the data, and it also satisfies Assumption (ref). $\square$

\vskip0.5cm

Proof of Theorem (ref): We prove the case for $\mathbb{E}[Y_{i}(1)|P_{i}={\rm E}]$, specifically

align[align omitted — 583 chars of source]

The proof of $\mathbb{E}[Y_{i}(0)|P_{i}={\rm E}]$ is similar. The score function representation is immediate from these equalities. We note that equality (ref) uses Assumptions (ref)--(ref) and equalities (ref) and (ref) only use the overlap condition, Assumption (ref)$(ii)$.

Consider ((ref)). By Assumption (ref) (unconfoundedness), it follows that \[ \mathbb{E}[Y_{i}(1)|P_{i}={\rm E}]=\mathbb{E}\left[\left.Y_{i}\cdot\frac{W_{i}}{\rho(X_{i})}\right|P_{i}={\rm E}\right]. \] Using the law of iterated expectations, we can first condition on $S_{i}$ and $X_{i}$ to get \[ \mathbb{E}\left[\left.Y_{i}\cdot\frac{W_{i}}{\rho(X_{i})}\right|P_{i}={\rm E}\right]=\mathbb{E}\left[\left.\mathbb{E}\left[\left.Y_{i}\cdot\frac{W_{i}}{\rho(X_{i})}\right|S_{i},X_{i},P_{i}={\rm E}\right]\right|P_{i}={\rm E}\right]. \] By Assumption (ref) (surrogacy), we have \[ \mathbb{E}\left[\left.\mathbb{E}\left[\left.Y_{i}\cdot\frac{W_{i}}{\rho(X_{i})}\right|S_{i},X_{i},P_{i}={\rm E}\right]\right|P_{i}={\rm E}\right]=\mathbb{E}\left[\left.\mathbb{E}\left[Y_{i}|S_{i},X_{i},P_{i}={\rm E}\right]\cdot\frac{\mathbb{E}\left[W_{i}|S_{i},X_{i},P_{i}={\rm E}\right]}{\rho(X_{i})}\right|P_{i}={\rm E}\right] \] By Assumption (ref) (Comparability), $\mu(s,x,{\rm E})=\mu(s,x,{\rm O})$ so that this is equal to \[ \mathbb{E}\left[\left.\mu(S_{i},X_{i},{\rm O})\cdot\frac{\mathbb{E}\left[W_{i}|S_{i},X_{i},P_{i}={\rm E}\right]}{\rho(X_{i})}\right|P_{i}={\rm E}\right]=\mathbb{E}\left[\left.\mu(S_{i},X_{i},{\rm O})\cdot\frac{\rho(S_{i},(X_{i})}{\rhoX_{i})}\right|P_{i}={\rm E}\right] \] Undoing the law of iterated expectations gives us the desired equality.

Consider ((ref)). By the definition of $\varphi(s,x)$, we have \[ \frac{\varphi(s,x)}{(1-\varphi(s,x))}\cdot\frac{1-\varphi}{\varphi}=\frac{{\rm pr}\left(\left.S_{i}=s,X_{i}=x\right|P_{i}={\rm E}\right)}{{\rm pr}\left(\left.S_{i}=s,X_{i}=x\right|P_{i}={\rm O}\right)} \] where the common support condition assures $1-\varphi(s,x)$ is not zero. This leads to \[ \mathbb{E}\left[\left.Y_{i}\cdot\frac{\rho(S_{i},X_{i})\cdot t(S_{i},X_{i})\cdot(1-\varphi)}{\rho (X_{i})\cdot(1-t(S_{i},X_{i}))\cdot \varphi}\right|P_{i}={\rm O}\right]=\mathbb{E}\left[\left.Y_{i}\cdot\frac{\rho(S_{i},X_{i})}{\rho(X_{i})}\cdot\frac{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm E}\right)}{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm O}\right)}\right|P_{i}={\rm O}\right] \] Again, by the law of iterated expectations, conditioning on $S_{i}$ and $X_{i}$ leads to \[ \mathbb{E}\left[\left.Y_{i}\cdot\frac{\rho(S_{i},X_{i})}{\rho(X_{i})}\cdot\frac{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm E}\right)}{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm O}\right)}\right|P_{i}={\rm O}\right]=\mathbb{E}\left[\left.\mu(S_{i},X_{i},{\rm O})\frac{\rho(S_{i},X_{i})}{\rho(X_{i})}\cdot\frac{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm E}\right)}{{\rm pr}\left(\left.S_{i},X_{i}\right|P_{i}={\rm O}\right)}\right|P_{i}={\rm O}\right] \] Using the definition of conditional expectations, we obtain

align*[align* omitted — 755 chars of source]

Consider ((ref)). By the law of iterated expectations conditional on $S_{i}$ and $X_{i}$, we obtain \[ \mathbb{E}\left[\left.Y_{i}\cdot\frac{\rho(S_{i},X_{i})\cdot \varphi(S_{i},X_{i})\cdot(1-\varphi)}{\rho(X_{i})\cdot(1-\varphi(S_{i},X_{i}))\cdot \varphi}\right|P_{i}={\rm O}\right]=\mathbb{E}\left[\left.\mu(S_{i},X_{i},{\rm O})\cdot\frac{\rho(S_{i},X_{i})\cdot \varphi(S_{i},X_{i})\cdot(1-\varphi)}{\rho(X_{i})\cdot(1-\varphi(S_{i},X_{i}))\cdot \varphi}\right|P_{i}={\rm O}\right] \] where the common support condition assures $1-\varphi(s,x)$ is not zero. Part $(iii)$ follows from Proposition (ref), which shows that Surrogacy and Comparability have no testable implications. Standard arguments then imply that unconfoundednes does not generate any testable implications. $\square$

\vskip0.5cm

Proof of Theorem (ref): For Part (i), we need to calculate the variance of the Efficient Influence Function (EIF) to obtain the efficiency bound. We provide the detailed calculation for completeness.\footnote{ While our influence function representation coincides with chen2023semiparametric, the variance calculation resulted in a slightly different expression.}

Given the EIF:

equation*[equation* omitted — 204 chars of source]

\[+\frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,x) - \mu(0,x) -\tau \Bigr) \] \[\hskip2cm+\frac{\boldsymbol{1}_{p={\rm O}}}{\varphi} \left(\frac{\varphi(s,x)}{1-\varphi(s,x)}\frac{(y-{\mu}(s,x,{\rm O}))\left({\rho}(s,x)-{\rho}(x)\right)}{{\rho}(x)(1-{\rho}(x))} \right)\]

\[ \mathbb{V} = \bigg[\psi(Y_i,S_i,W_i,X_i)^2 \bigg] \]

\[ =\mathbb{E} \bigg[ \left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi}\left(\frac{W_{i} \cdot (\mu(S_{i},X_{i},{\rm O})-\mu(1,x))}{{\rho}(X_{i})}-\frac{(1-W_{i})\cdot ({\mu}(S_{i},X_{i},{\rm O})- \mu(0,x) )}{1-{\rho}(X_{i})}\right) \right)^2 \] \[ + \left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,x) - \mu(0,x) -\tau \Bigr) \right)^2 \] \[ + \left( \frac{\boldsymbol{1}_{p={\rm O}}}{\varphi} \left(\frac{\varphi(S_{i},X_{i})}{1-\varphi(S_{i},X_{i})}\frac{(Y_{i}-{\mu}(S_{i},X_{i},{\rm O}))\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)}{{\rho}(X_{i})(1-{\rho}(X_{i}))} \right)\right)^2 \bigg] \]

Focusing on the first block \[\left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi}\left(\frac{w \cdot (\mu(s,x,{\rm O})-\mu(1,x))}{{\rho}(x)}-\frac{(1-w)\cdot ({\mu}(s,x,{\rm O})- \mu(0,x) )}{1-{\rho}(x)}\right) \right)^2, \] noting that $w(1-w)=0$ and hence the cross-term disappearing, we only have to take the expectation of \[ \left(\frac{\boldsymbol{1}_{p={\rm E}}}{\varphi}\frac{w \cdot (\mu(s,x,{\rm O})-\mu(1,x))}{{\rho}(x)}\right)^2 \qquad {\rm and }\quad \left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \frac{ (1-w)\cdot ({\mu}(s,x,{\rm O})- \mu(0,x) )}{1-{\rho}(x)}\right)^2 \] Note that \[ \mathbb{E} \bigg[ \left(\frac{\boldsymbol{1}_{p={\rm E}}}{\varphi}\frac{W_{i} \cdot (\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))}{{\rho}(X_{i})}\right)^2 \bigg] \] \[ = \mathbb{E} \bigg[ \frac{(\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2}{ {\rho}(X_{i})^2\varphi^2} \mathbb{E} \big[ \boldsymbol{1}_{p={\rm E}} W_{i} | S_{i},X_{i} \big] \bigg]\quad (\because \text{Tower Property} ) \] \[ = \mathbb{E} \bigg[ \frac{(\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2}{ {\rho}(X_{i})^2\varphi^2} \varphi(S_{i},X_{i}) \mathbb{E} \big[ W_{i} | S_{i},X_{i} ,P_{i} = {\rm E} \big] \bigg] \] \[ = \mathbb{E} \bigg[ \frac{(\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2}{ {\rho}(X_{i})^2\varphi^2} \varphi(S_{i},X_{i}) \rho(S_{i},X_{i}) \bigg] \] \[ = \mathbb{E} \bigg[ \frac{\varphi(S_{i},X_{i}) \rho(S_{i},X_{i}) }{ \varphi^2{\rho}(X_{i})^2} (\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2\bigg] \]

Likewise, we can derive \[ \mathbb{E} \bigg[\left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \frac{ (1-W_{i})\cdot ({\mu}(S_{i},X_{i},{\rm O})- \mu(0,X_{i}) )}{1-{\rho}(X_{i})}\right)^2 \bigg] \] \[ = \mathbb{E} \bigg[ \frac{\varphi(S_{i},X_{i}) (1-\rho(S_{i},X_{i})) }{ \varphi^2(1-\rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm O})-\mu(0,X_{i}))^2\bigg] \] Collectivizing the two term yields the first block: \[ \mathbb{E} \bigg[ \frac{\varphi(S_{i},X_{i}) }{ \varphi^2} \bigg( \frac{1-\rho(S_{i},X_{i}) }{(1-\rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm O})-\mu(0,X_{i}))^2 + \frac{\rho(S_{i},X_{i}) }{{\rho}(X_{i})^2} (\mu(S_{i},X_{i},{\rm O})-\mu(1,X_{i}))^2 \bigg) \bigg] \]

Next, for the second block \[\frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,x) - \mu(0,x) -\tau \Bigr) \] we can likewise derive by using the Tower Property with respect to $X_{i}$ that \[ \mathbb{E} \bigg[ \left( \frac{\boldsymbol{1}_{p={\rm E}}}{\varphi} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr) \right)^2 \bigg] = \mathbb{E} \bigg[ \frac{\varphi(X_{i})}{\varphi^2} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \bigg] \] Finally, for the third block \[ \left( \frac{\boldsymbol{1}_{p={\rm O}}}{\varphi} \left(\frac{\varphi(s,x)}{1-\varphi(s,x)}\frac{(y-{\mu}(s,x,{\rm O}))\left({\rho}(s,x)-{\rho}(x)\right)}{{\rho}(x)(1-{\rho}(x))} \right)\right)^2, \] note that \[ \mathbb{E} \bigg[ \left( \frac{\boldsymbol{1}_{P_{i}={\rm O}}}{\varphi} \left(\frac{\varphi(S_{i},X_{i})}{1-\varphi(S_{i},X_{i})}\frac{(Y_{i}-{\mu}(S_{i},X_{i},{\rm O}))\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)}{{\rho}(X_{i})(1-{\rho}(X_{i}))} \right)\right)^2 \bigg] \] \[ =\mathbb{E} \bigg[ \frac{(\varphi(S_{i},X_{i}))^2}{\varphi^2(1-\varphi(S_{i},X_{i}))^2} \frac{\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)^2 }{{((\rho}(X_{i})(1-{\rho}(X_{i})))^2)^2} \mathbb{E} \bigg[ \boldsymbol{1}_{P_{i}={\rm O}} (Y_{i}-{\mu}(S_{i},X_{i},{\rm O}))^2 | S_{i},X_{i} \bigg] \bigg] \] \[ =\mathbb{E} \bigg[ \frac{(\varphi(S_{i},X_{i}))^2}{\varphi^2(1-\varphi(S_{i},X_{i}))^2} \frac{\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)^2 }{{(\rho}(X_{i})(1-{\rho}(X_{i})))^2} \mathbb{E} \bigg[ (1-\varphi(S_{i},X_{i})) (Y_{i}-{\mu}(S_{i},X_{i},{\rm O}))^2 | S_{i},X_{i} , P_{i} = {\rm O} \bigg] \bigg] \] \[ =\mathbb{E} \bigg[ \frac{(\varphi(S_{i},X_{i}))^2}{\varphi^2(1-\varphi(S_{i},X_{i}))^2} \frac{\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)^2 }{{(\rho}(X_{i})(1-{\rho}(X_{i})))^2} (1-\varphi(S_{i},X_{i})) \sigma^2(S_{i},X_{i},{\rm O}) \bigg] \] \[ =\mathbb{E} \bigg[ \frac{(\varphi(S_{i},X_{i}))^2}{\varphi^2(1-\varphi(S_{i},X_{i}))} \frac{\left({\rho}(S_{i},X_{i})-{\rho}(X_{i})\right)^2 }{{(\rho}(X_{i})(1-{\rho}(X_{i})))^2} \sigma^2(S_{i},X_{i},{\rm O}) \bigg] \] \[ =\mathbb{E} \bigg[\frac{1-\varphi(S_{i},X_{i})}{\varphi^2} \left( \left( \frac{\varphi(S_{i},X_{i})}{1 -\varphi(S_{i},X_{i})} \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \right)\bigg] \]

Hence, adding up the three blocks (in the order from the third to the first block) yield the desired efficiency bound: \[ \mathbb{V}=\mathbb{E}[\psi(Y_{i},S_{i},W_{i},X_{i},P_{i})^2]\] \[\qquad=\mathbb{E}\bigg[ \frac{1-\varphi(S_{i},X_{i})}{\varphi^2} \left( \left( \frac{\varphi(S_{i},X_{i})}{1 -\varphi(S_{i},X_{i})} \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \right) \] \[+\frac{\varphi(X_{i})}{\varphi^2} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left. + \frac{\varphi(S_{i},X_{i})}{\varphi^2} \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right].\] \[\qquad=\mathbb{E}\bigg[ \frac{1}{\varphi^2} \frac{\varphi(S_{i},X_{i})^2}{1 -\varphi(S_{i},X_{i})}\left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \] \[+\frac{\varphi(X_{i})}{\varphi^2} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left. + \frac{\varphi(S_{i},X_{i})}{\varphi^2} \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right].\]

For part $(ii)$, first rewrite the variance bound, normalized by the square root of the expected size of the experimental sample, $\varphi N$, instead of normalized by the total sample size $N$, as \[\tilde \mathbb{V}=\mathbb{E}\bigg[ \frac{1}{\varphi} \frac{\varphi(S_{i},X_{i})^2}{1 -\varphi(S_{i},X_{i})}\left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \] \[+\frac{\varphi(X_{i})}{\varphi} \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left. + \frac{\varphi(S_{i},X_{i})}{\varphi} \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right].\]

Next, we re-write the bound in terms of a conditional expectation in the experimental sample, rather than as the unconditional expectation, (this implies multiplying by $\varphi/\varphi(S_i,X_i)$ or $\varphi/\varphi(X_i)$ appropriately) as

\[\tilde \mathbb{V}=\mathbb{E}\bigg[ \frac{\varphi(S_{i},X_{i})}{1 -\varphi(S_{i},X_{i})}\left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2 \sigma^2(S_{i},X_{i},{\rm O}) \] \[+ \Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left.\left. + \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right|P_{i}={\rm E}\right].\]

Now we consider a sequence of data generating processes, where the outcome distribution in the observational sample remains fixed, and the propensity and surrogate scores remain fixed, and only the functions $\varphi(s,x)$, $\varphi(x)$ and the scalar $\varphi$ change, in such a way that $\sup_{s,x}\varphi(s,x)\rightarrow 0$. The the first term converges to zero, leaving us with \[\bar\mathbb{V}=\mathbb{E}\bigg[\Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left.\left. + \left( \frac{ (1-\rho(S_{i},X_{i}))(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\rho(S_{i},X_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right|P_{i}={\rm E}\right].\] The final step is to note that $\rho(S_{i},X_{i})=\mathbb{E}[W_{i}|S_{i},X_{i},P_{i}={\rm E}]$ so we can write $\bar\mathbb{V}$ as \[\bar\mathbb{V}=\mathbb{E}\bigg[\Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left.\left. + \left( \frac{ (1-\mathbb{E}[W_{i}|S_{i},X_{i},P_{i}={\rm E}])(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{\mathbb{E}[W_{i}|S_{i},X_{i},P_{i}={\rm E}](\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right|P_{i}={\rm E}\right]\] \[=\mathbb{E}\bigg[\Bigl(\mu(1,X_{i}) - \mu(0,X_{i}) -\tau \Bigr)^2 \] \[\left.\left. + \left( \frac{ (1-W_{i})(\mu(S_{i},X_{i},{\rm O}) - \mu(0,X_{i}))^2}{(1 - \rho(X_{i}))^2} + \frac{W_{i}(\mu(S_{i},X_{i},{\rm O}) - \mu(1,X_{i}))^2}{\rho(X_{i})^2} \right) \right|P_{i}={\rm E}\right].\]

$\square$

\vskip0.5cm

Proof of Theorem (ref):

The first representation of the efficiency bound without surrogacy in part $(i)$ of the Theorem is essentially rewriting the efficiency bound in hahn1998role, and related results in robins1995semiparametric,robins1995analysis. The standard version of the efficiency bound is \[ \mathbb{V}=\mathbb{E} \left[ \frac{\sigma^2(1,X_i)}{\rho(X_i)}+\frac{\sigma^2(0,X_i)}{1-\rho(X_i)}+\left(\mu(1,X_i)-\mu(0,X_i)-\tau\right)^2 \right].\] The proof consists of showing that this is equal to the expression for $\mathbb{V}_{{\rm ns}}$ in Theorem (ref): \[ \mathbb{V}_{{\rm ns}}=\mathbb{E}\biggl[\sigma^{2}(S_{i},X_{i},{\rm E})\cdot\left(\frac{\rho(S_{i},X_{i})}{\rho(X_{i})^{2}}+\frac{1-\rho(S_{i},X_{i})}{(1-\rho(X_{i}))^{2}}\right) \] \[ \hskip2cm+\frac{\rho(S_{i},X_{i})}{\rho(X_{i})^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(1,X_{i})\right)^{2}+\frac{1-\rho(S_{i},X_{i})}{(1-\rho(X_{i}))^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(0,X_{i})\right)^{2} \] \[\hskip2cm +\left(\mu(1,X_i)-\mu(0,X_i)-\tau\right)^2 \biggr].\] which amounts to showing the equality of

equation[equation omitted — 123 chars of source]

and

equation[equation omitted — 185 chars of source]

\[ \hskip2cm+\frac{\rho(S_{i},X_{i})}{\rho(X_{i})^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(1,X_{i})\right)^{2}+\frac{1-\rho(S_{i},X_{i})}{(1-\rho(X_{i}))^{2}}\cdot\left(\mu(S_{i},X_{i},{\rm E})-\mu(0,X_{i})\right)^{2}\biggr]. \]

By unconfoundedness \[ \sigma^2(1,x)\equiv \mathbb{V}(Y_i(1)|X_i=x)=\mathbb{V}(Y_i|W_i=1,X_i=x),\] where as mentioned in the main text, we implicitly condition on the sampling indicator and abstract it from the notation when it does not lead to confusion.

By iterated expectations this is equal to \[ \mathbb{E}\left[\left.\mathbb{V}\left(Y_i|W_i=1,S_i,X_i=x\right)\right|W_i=1,X_i=x\right]+\mathbb{V}\left(\left.\mathbb{E}[Y_i|W_i=1,S_i,X_i]\right|W_i=1,X_i\right).\] By surrogacy the conditional distribution of $Y_i$ given $W_i$, $S_i$ and $X_i$ does not vary by $W_i$, so this is equal to \[ \mathbb{E}\left[\left.\mathbb{V}\left(Y_i|S_i,X_i=x\right)\right|W_i=1,X_i=x\right]+\mathbb{V}\left(\left.\mathbb{E}[Y_i|S_i,X_i]\right|W_i=1,X_i\right)\] \[ \hskip1cm =\mathbb{E}\left[\left.\sigma^2(S_i,X_i)\right|W_i=1,X_i=x\right]+\mathbb{V}\left(\left.\mu(S_i,X_i)\right|W_i=1,X_i\right).\]

For the first term, \[ \mathbb{E}\left[\left.\sigma^2(S_i,X_i)\right|W_i=1,X_i=x\right]=\mathbb{E}\left[\left.\frac{\sigma^2(S_i,X_i)\rho(S_i,X_i)}{\rho(X_i)}\right|X_i=x\right].\] For the second term, note that \[ \mathbb{E}\left[\left.\mu(S_i,X_i)\right|W_i=1,X_i\right] = \mathbb{E}\left[\left.\mathbb{E}[Y_i|S_i,X_i]\right|W_i=1,X_i\right] \] is by surrogacy equal to $ \mathbb{E}\left[\left.\mathbb{E}[Y_i|W_i=1,S_i,X_i]\right|W_i=1,X_i\right] ,$ which in turn by iterated expectations is equal to $ \mathbb{E}\left[\left.Y_i\right|W_i=1,X_i\right] =\mu(1,X_i).$ Hence the second term is \[ \mathbb{V}\left(\left.\mu(S_i,X_i)\right|W_i=1,X_i\right) = \mathbb{E}\left[\left.\left(\mu(S_i,X_i)-\mu(1,X_i)\right)^2\right| W_i=1,X_i\right] \] \[\hskip1cm = \mathbb{E}\left[ \left(\mu(S_i,X_i)-\mu(1,X_i)\right)^2 \frac{\rho(S_i,X_i)}{\rho(X_i)} \right].\] Combining the two terms and including the denominator $\rho(X_i)$, we have \[\mathbb{E} \left[ \frac{\sigma^2(1,X_i)}{\rho(X_i)} \right]=\mathbb{E}\left[\frac{\sigma^2(S_i,X_i)\rho(S_i,X_i)}{\rho(X_i)^2}\right]+ \mathbb{E}\left[ \left(\mu(S_i,X_i)-\mu(1,X_i)\right)^2 \frac{\rho(S_i,X_i)}{\rho(X_i)^2} \right]. \] By the same argument \[\mathbb{E} \left[ \frac{\sigma^2(0,X_i)}{1-\rho(X_i)} \right]=\mathbb{E}\left[\frac{\sigma^2(S_i,X_i)(1-\rho(S_i,X_i))}{(1-\rho(X_i))^2}\right]+ \mathbb{E}\left[ \left(\mu(S_i,X_i)-\mu(0,X_i)\right)^2 \frac{1-\rho(S_i,X_i)}{(1-\rho(X_i))^2} \right] \] Hence, adding up the two equalities above shows the desired equivalence of ((ref)) and ((ref)). This finishes the proof of part $(i)$ of the theorem.

Next, for part $(ii)$ of the theorem, we derive the efficiency bound for the case with surrogacy by first deriving the efficient influence function and then deriving its variance. To derive the efficient influence function, we follow the proof in chen2023semiparametric and newey1990semiparametric, specifically the following four steps: (1) constructing the tangent space, (2) deriving the pathwise derivative of the target estimand (i.e. the ATE under surrogacy), (3) showing that the conjectured efficient influence function (EIF) lies in the tangent space, and (4) showing that the pathwise derivative of the target estimand and the conjectured EIF satisfies a key condition in newey1990semiparametric.

First, to characterize the tangent space, considering the data density where the functions $f$ denote the density of random variables. \[ f_{Y_i,S_i,W_i,X_i}(y,s,w,x) = f_{Y_i \mid S_i,X_i} (y\mid s,x) f_{S_i \mid W_i,X_i}(s \mid w,x) f_{W_i \mid X_i}(w \mid x) f_{X_i}(x). \] We assume the data density satisfies the regularity and smoothness conditions in Definition (A.1) of Newey (1990).

Let $G^\epsilon$ be a parametric submodel parameterized by $\epsilon \in [0,1]$ where $G^{\epsilon =0} = G$ and $G$ is the true data generating model. Let $f^{\epsilon}$ be the corresponding density function for the parametric submodel. Then, the score of $f_{\epsilon}$ is

{

align*[align* omitted — 552 chars of source]

} We use $Q(\cdot)$'s to denote the score function, i.e. $Q(\cdot) = \frac{\delta}{\delta \epsilon} \log(f^\epsilon(\cdot))$. Evaluating the derivative at $\epsilon=0$ leads us to the score of the true model, i.e., \[ Q_{Y_i,S_i,W_i,X_i}(y,s,w,x) = Q_{Y_i\mid S_i,X_i}(y \mid s,x) + Q_{S_i \mid W_i,X_i}(s \mid w,x) + Q_{W_i\mid X_i}(w \mid x) + Q_{X_i}(x). \] The tangent space $\mathcal{T}$ is the mean closure of a linear combination of mean-zero, square-integrable functions $\overline{Q}_1,...,\overline{Q}_4$ that satisfy the following conditions:

align*[align* omitted — 516 chars of source]

Second, we derive the pathwise derivative of our estimand. With some abuse of the integral notation, our estimand can be written as follows:

align*[align* omitted — 443 chars of source]

The pathwise derivative of the estimand $\tau$ is

{

align*[align* omitted — 2,285 chars of source]

} The derivatives above use the chain rule from calculus and the fact that \[ \frac{\delta}{\delta \epsilon} f^{\epsilon} = \frac{\delta}{\delta \epsilon} \log(f^{\epsilon}) f^{\epsilon} = Q^{\epsilon} f^{\epsilon} \]

Let $\tau'$ denote evaluating the above derivative at $\epsilon = 0$, i.e. {

align*[align* omitted — 1,223 chars of source]

} Third, consider the conjectured efficient influence function (EIF). { \[ \psi(Y_i,S_i,W_i,X_i) = \frac{(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))} + \frac{W_i(\mu(S_i,X_i) - \mu(1,X_i))}{\rho(X_i)} - \frac{(1-W_i)(\mu(S_i,X_i) - \mu(0,X_i))}{1-\rho(X_i)} + \mu(1,X_i) - \mu(0,X_i) - \tau \]} We show that $\psi(Y_i,S_i,W_i,X_i)$ is an element of the tangent space $\mathcal{T}$ by showing that different parts of $\psi(Y_i,S_i,W_i,X_i)$ satisfies conditions for $\overline{Q}_1, \overline{Q}_2$, and $\overline{Q}_4$.

enumerate• For $\overline{Q}_1$, we have $\mathbb{E}\biggl[\frac{(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))} \biggl| S_i=s,X_i=x\biggr] = 0$ by definition of $\mu(S_i,X_i)$ and $\mathbb{E} \biggl[\frac{(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))} \biggl| S_i=s,W_i=w,X_i=x\biggr] = 0$ by using statistical surrogacy. • For $\overline{Q}_2$, we have $\mathbb{E} \biggl[\frac{W_i(\mu(S_i,X_i) - \mu(1,X_i))}{\rho(X_i)} \biggl| W_i=w,X_i=x\biggr] = \frac{w}{\rho(x)} (\mathbb{E} [\mu(S_i,X_i) \mid W_i=w,X_i=x] -\mu(1,x)) = 0$ for any value of $w$. Similarly, $\mathbb{E} \biggl[\frac{(1-W_i)(\mu(S_i,X_i) - \mu(0,X_i))}{1-\rho(X_i)} \biggl| W_i=w,X_i=x\biggr] = \frac{1-w}{1-\rho(x)} \biggl(\mathbb{E} \biggl[h(S_i,X_i) \biggl| W_i=w,X_i=x\biggr] -\mu(0,x)\biggr) = 0$ for any value of $w$. • For $\overline{Q}_4$, we have $\mathbb{E}[ \mu(1,X_i) - \mu(0,X_i) - \tau] = 0$.

By setting $\overline{Q}_3 = 0$, we arrive at $\psi(Y_i,S_i,W_i,X_i) \in \mathcal{T}$.

Fourth, we show that $\tau'$ and $\psi(Y_i,S_i,W_i,X_i)$ satisfy the following relationship that all efficient influence functions must satisfy from Theorem 2.2 in newey1990semiparametric:

equation[equation omitted — 84 chars of source]

We break the proof of this equality into several steps.

enumerate[label=(\alph*)] • Let us consider the part of the $\psi(Y_i,S_i,W_i,X_i)$ concerning $\frac{(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))}$. We have { \begin{align*} &\mathbb{E}\biggl[\frac{(Y_i - h(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))} Q_{Y_i,S_i,W_i,X_i}(Y_i,S_i,W_i,X_i) \biggr] \\ =&\mathbb{E}\Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ (Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) Q_{Y_i,S_i,W_i,X_i}(Y_i,S_i,W_i,X_i) \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E}\Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} E \biggl[(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) Q_{Y_i\mid,S_i,X_i}(Y_i \mid S_i,X_i) \biggl| X_i \biggr] \Biggr] \\ &\quad + \mathbb{E}\Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i) \biggl| X_i \biggr] \Biggr] \\ &\quad + \mathbb{E}\Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) Q_{W_i \mid X_i}(W_i\mid X_i) \biggl| X_i \biggr] \Biggr] \\ &\quad + \mathbb{E} \Biggl[ \frac{Q_{X_i}(X_i)}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) \biggl| X_i \biggr] \Biggr] \end{align*} } The first equality uses the law of total expectation. The second equality uses the definition of $Q_{Y_i,S_i,W_i,X_i}$. We consider each term separately, starting from the bottom. For the $Q_{X_i}$ term, we have \begin{align*} &\mathbb{E}\Biggl[ \frac{Q_{X_i}(X_i)}{\rho(X_i)(1-\rho(X_i))} E \biggl[(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E}\Biggl[ \frac{Q_{X_i}(X_i)}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ (\rho(S_i,X_i) - \rho(X_i)) \mathbb{E} \bigl[(Y_i - \mu(S_i,X_i)) \bigl| S_i, X_i \biggr] \biggl| X_i \biggr] \Biggr] \\ =&\mathbb{E} \Biggl[ \frac{Q_4(X_i)}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ (\rho(S_i,X_i) - \rho(X_i)) \cdot 0 \biggl| X_i \biggr] \Biggr] \\ =& 0. \end{align*} The first equality uses the law of total expectation. The second equality uses the definition of $\mu(S_i,X_i)$. For the $Q_{W_i\mid X_i}$ term, we have { \begin{align*} &\mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) Q_{W_i\mid X_i}(W_i\mid X_i) \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ Q_{W_i\mid X_i}(W_i \mid X_i) \mathbb{E} \bigl[ (Y_i - h(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) \bigl| W_i,X_i \bigr] \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ Q_{W_i\mid X_i}(W_i \mid X_i) \mathbb{E} \biggl[ (\rho(S_i,X_i) - \rho(X_i)) \mathbb{E} \bigl[ (Y_i - \mu(S_i,X_i)) \bigl| S_i, W_i,X_i \bigr] \biggl| W_i,X_i \biggr] \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ Q_{W_i\mid X_i}(W_i \mid X_i) \mathbb{E} \biggl[ (\rho(S_i,X_i) - \rho(X_i)) (\mathbb{E} \bigl[Y_i \mid S_i,W_i,X_i\bigr] - \mu(S_i,X_i)) \biggl| W_i,X_i \biggr] \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ Q_{W_i\mid X_i}(W_i \mid X_i) \mathbb{E} \biggl[ (\rho(S_i,X_i) - \rho(X_i)) \cdot 0 \biggl| W_i,X_i \biggr] \biggl| X_i \biggr] \Biggr] \\ =& 0 \end{align*}} The first and second equalities use the law of total expectation. The third equality is algebra. The fourth equality uses statistical surrogacy. For the $Q_{S_i \mid W_i,X_i}$ term, we have \begin{align*} &\mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i) \biggl| X_i \biggr] \Biggr] \\ =&\mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ \mathbb{E} \biggl[ (Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i) \biggl| X_i,W_i \biggr] \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ \mathbb{E} \biggl[ (\rho(S_i,X_i) - \rho(X_i)) Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i) \mathbb{E} \biggl[ Y_i - \mu(S_i,X_i) \biggl| S_i,X_i,W_i \biggr] \biggl| X_i,W_i \biggr] \biggl| X_i \biggr] \Biggr] \\ =& 0 \end{align*} The first two equalities use the law of total expectation. The third equality uses statistical surrogacy. For the $Q_{Y_i\mid S_i,X_i}$ term, we have { \begin{align*} &\mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[(Y_i - \mu(S_i,X_i)) (\rho(S_i,X_i) - \rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| X_i \biggr] \Biggr] \\ =&\mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ Y_i (\rho(S_i,X_i) - \rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| X_i \biggr] \Biggr] \\ &\quad - \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ \mu(S_i,X_i) (\rho(S_i,X_i) - \rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| X_i \biggr] \Biggr] \end{align*}} The first term above is equal to $\mathbb{E} \biggl[\frac{Y_i(W_i -\rho(X_i))Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) }{\rho(X_i) (1-\rho(X_i))} \biggr]$ because \begin{align*} &\mathbb{E} \Biggl[ \frac{Y_i(W_i -\rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i)}{\rho(X_i) (1-\rho(X_i))} \biggr] \\ =& \mathbb{E} \biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[Y_i(W_i -\rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) \mid S_i,X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[Y_i Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| S_i,X_i \biggr] \mathbb{E} \biggl[ W_i - \rho(X_i) \biggl| S_i,X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[Y_i Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| S_i,X_i \biggr] (\rho(S_i,X_i) - \rho(X_i)) \biggr] \\ =& \mathbb{E} \Biggl[ \frac{Y_i (\rho(S_i,X_i) - \rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i)}{\rho(X_i)(1-\rho(X_i))} \Biggr] \end{align*} The first equality uses the law of total expectation. The second equality uses statistical surrogacy where $Y_i \perp W_i | S_i,X_i$ implies $Y_i,S_i,X_i \perp W_i,X_i | S_i,X_i$. The third equality is the definition of the surrogate score. The fourth equality uses the law of total expectation. The second term above simplifies to zero because { \begin{align*} &\mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ \mu(S_i,X_i) (\rho(S_i,X_i) - \rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ \mu(S_i,X_i) (\rho(S_i,X_i) - \rho(X_i)) \mathbb{E} \biggl[ Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| S_i,X_i \biggr] \biggl| X_i \biggr] \Biggr] \\ =& \mathbb{E} \Biggl[ \frac{1}{\rho(X_i)(1-\rho(X_i))} \mathbb{E} \biggl[ \mu(S_i,X_i) (\rho(S_i,X_i) - \rho(X_i)) \cdot 0 \biggl| X_i \biggr] \Biggr] \\ =& 0 \end{align*}} The first equality uses the law of total expectation. The second equality uses the the mean-zero property of the score function $Q_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i)$. Finally, we can rewrite $\frac{Y_i(W_i -\rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i)}{\rho(X_i) (1-\rho(X_i))}$ as { \[ \frac{Y_i(W_i -\rho(X_i)) Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i)}{\rho(X_i) (1-\rho(X_i))} = \frac{Y_iW_iQ_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i)}{\rho(X_i)} - \frac{Y_i(1-W_i)Q_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i)}{1-\rho(X_i)} \]} Also, in expectation, each term above equals to {\scriptsize \begin{align*} \mathbb{E} \biggl[ \frac{W_iY_iQ_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i)}{\rho(X_i)} \biggr] & = \mathbb{E} \biggl[ \frac{1}{\rho(X_i)} \mathbb{E} \biggl[ Y_iQ_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i) \mid W_i = 1, X_i \biggr] \rho(X_i) \biggr] = \mathbb{E} \biggl[ \mathbb{E} \biggl[ Y_iQ_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i) \mid W_i = 1, X_i \biggr] \biggr], \\ \mathbb{E} \biggl[ \frac{(1-W_i)Y_iQ_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i)}{1-\rho(X_i)} \biggr] &= \mathbb{E} \biggl[ \frac{1}{1-\rho(X_i)} \mathbb{E} \biggl[ Y_iQ_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i) \mid W_i = 0, X_i \biggr] (1-\rho(X_i)) \biggr] \\ &= \mathbb{E} \biggl[ \mathbb{E} \biggl[ Y_iQ_{Y_i\mid S_i,X_i} (Y_i\mid S_i,X_i) \mid W_i = 0, X_i \biggr] \biggr]. \end{align*} } The first equality uses the law of total expectation and the definition of the propensity score. The second equality is algebra. Overall, we have \begin{align*} &\mathbb{E} \biggl[\frac{(Y_i-\mu(S_i,X_i))(\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))} Q_{Y_i,S_i,W_i,X_i}(Y_i,S_i,W_i,X_i) \biggl] \\ =& \mathbb{E} \biggl[ \mathbb{E} \biggl[ Y_iQ_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i) \mid W_i = 1, X_i \biggr] \biggr] - \mathbb{E} \biggl[ \mathbb{E} \biggl[ Y_iQ_{Y_i\mid S_i,X_i}(Y_i\mid S_i,X_i) \mid W_i = 0, X_i \biggr] \biggr]. \end{align*} • Let's consider the part of the $\psi(Y_i,S_i,W_i,X_i)$ concerning $\frac{W_i(\mu(S_i,X_i) - \mu(1,X_i))}{\rho(X_i)}$. We have { \begin{align*} &\mathbb{E}\biggl[ \frac{W_i(\mu(S_i,X_i) - \mu(1,X_i))}{\rho(X_i)} Q_{Y_i,S_i,W_i,X_i}(Y_i,S_i,W_i,X_i)\biggr] \\ =& \mathbb{E} \biggl[ \frac{W_i}{\rho(X_i)} \mathbb{E} \biggl[ \biggl\{ \mu(S_i,X_i) - \mu(1,X_i) \biggr\} \biggl\{ Q_{Y_i \mid S_i,X_i}(Y_i \mid S_i,X_i) + Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i) + Q_{W_i \mid X_i}(W_i \mid X_i) + Q_{X_i}(X_i)\biggr\} \biggl| W_i,X_i\biggr] \biggr] \\ =& \mathbb{E} \biggl[ \frac{W_i}{\rho(X_i)} \mathbb{E} \biggl[ (\mu(S_i,X_i) - \mu(1,X_i)) Q_{Y_i \mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| W_i,X_i \biggr] \biggr] \\ &\quad + \mathbb{E} \biggl[ \frac{W_i}{\rho(X_i)} \mathbb{E} \biggl[ (\mu(S_i,X_i) - \mu(1,X_i)) Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i) \biggl| W_i,X_i \biggr] \biggr] \\ &\quad + \mathbb{E} \biggl[ \frac{W_i}{\rho(X_i)} Q_{W_i \mid X_i}(W_i \mid X_i) \mathbb{E} \biggl[ \mu(S_i,X_i) - \mu(1,X_i) \biggl| W_i,X_i \biggr] \biggr] \\ &\quad+\mathbb{E} \biggl[ \frac{W_i}{\rho(X_i)} Q_{X_i}(X_i) \mathbb{E} \biggl[ \mu(S_i,X_i) - \mu(1,X_i) \biggl| W_i,X_i \biggr] \biggr] \\ =& \mathbb{E} \biggl[ \frac{W_i}{\rho(X_i)} \mathbb{E} \biggl[ (\mu(S_i,X_i) - \mu(1,X_i)) \mathbb{E} \biggl[ Q_{Y_i \mid S_i,X_i}(Y_i \mid S_i,X_i) \biggl| S_i,W_i,X_i \biggr] \bigg| W_i,X_i \biggr] \biggr] \\ &\quad + \mathbb{E} \biggl[ \frac{W_i}{\rho(X_i)} \mathbb{E} \biggl[ (\mu(S_i,X_i) - \mu(1,X_i)) Q_{S_i\mid W_i,X_i}(S_i \mid W_i,X_i) \biggl| W_i,X_i \biggr] \biggr] \\ =& \mathbb{E} \biggl[ \frac{W_i(\mu(S_i,X_i) - \mu(1,X_i))Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i)}{\rho(X_i)} \biggr] \end{align*}} The first equality uses the law of total expectation. The third equality uses the relationship $\mathbb{E} [\mu(S_i,X_i)| W_i=w,X_i] = \mu(w,X_i)$. The fourth equality uses the mean-zero property of the score and the law of total expectation. We can further simplify the above expression by noticing that \begin{align*} \mathbb{E} \biggl[ \frac{W_i \mu(1,X_i)Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i)}{\rho(X_i)} \biggr] &= \mathbb{E} \biggl[ \frac{W_i \mu(1,X_i)}{\rho(X_i)} \mathbb{E} \biggl[ Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i)\biggl| W_i, X_i \biggr] \biggr] =0 \end{align*} The first equality uses the law of total expectation. The second equality uses the mean-zero property of the score function. Also, \begin{align*} \mathbb{E} \biggl[ \frac{W_i \mu(S_i,X_i) Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i)}{\rho(X_i)} \biggr] &= \mathbb{E} \biggl[ \mathbb{E} \biggl[\mu(S_i,X_i) Q_{S_i \mid W_i,X_i}(S_i \mid W_i=1,X_i)\biggl| W_i=1,X_i \biggr] \biggl| X_i \biggr] \end{align*} The first equality uses the law of total expectation and the definition of conditional expectation with the definition $\mathbb{E} [W_i \mid X_i] = \rho(X_i)$. Overall, we end up with the following expression \[ \mathbb{E} \biggl[ \frac{W_i(\mu(S_i,X_i) - \mu(1,X_i))}{\rho(X_i)} Q_{Y_i,S_i,W_i,X_i}(Y_i,S_i,W_i,X_i)\biggr] = \mathbb{E} \biggl[ \mathbb{E} \biggl[\mu(S_i,X_i) Q_{S_i \mid W_i,X_i}(S_i \mid W_i=1,X_i)\biggl| X_i,W_i=1 \biggr] \biggl| X_i \biggr] \] • Let's consider the part of the $\psi(Y_i,S_i,W_i,X_i)$ concerning $\frac{(1-W_i)(\mu(S_i,X_i) - \mu(0,X_i))}{1-\rho(X_i)}$. From the above exercise, we end up with { \[ \mathbb{E} \biggl[ \frac{(1-W_i)(\mu(S_i,X_i) - \mu(0,X_i))}{1-\rho(X_i)} Q_{Y_i,S_i,W_i,X_i}(Y_i,S_i,W_i,X_i)\biggr] = \mathbb{E} \biggl[ \mathbb{E} \biggl[ \mu(S_i,X_i)Q_{S_i \mid W_i,X_i}(S_i \mid W_i=0,X_i) \biggl| W_i=0,X_i] \biggl| X_i\biggr] \]} • Let's consider the part of the $\psi(Y_i,S_i,W_i,X_i)$ concerning $\mu(1,X_i) - \mu(0,X_i) - \tau$. We have \begin{align*} &\mathbb{E} [ (\mu(1,X_i) - \mu(0,X_i) - \tau)Q_{Y_i,S_i,W_i,X_i}(Y_i,S_i,W_i,X_i)] \\ =& \mathbb{E} [(\mu(1,X_i) - \mu(0,X_i) - \tau) \mathbb{E}[Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) + Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i) + Q_{W_i \mid X_i}(W_i \mid X_i) + Q_{X_i}(X_i) \mid X_i]] \\ =& \mathbb{E} [(\mu(1,X_i) - \mu(0,X_i) - \tau) \mathbb{E} [Q_{Y_i\mid S_i,X_i}(Y_i \mid S_i,X_i) + Q_{S_i \mid W_i,X_i}(S_i \mid W_i,X_i) + Q_{X_i}(X_i) \mid X_i]] \\ =& E[(\mu(1,X_i) - \mu(0,X_i) - \tau) Q_{X_i}(X_i)] \\ =& E[(\mu(1,X_i) - \mu(0,X_i)) Q_{X_i}(X_i)] \end{align*} The first equality uses the law of total expectation. The second equality uses the property of the score where $E[Q_{W_i \mid X_i}(W_i \mid X_i) \mid X_i] = 0$. The third equality uses both the law of total expectation and the property of the score where \[ \mathbb{E}[Q_{Y_i \mid S_i,X_i}(Y_i \mid S_i,X_i) \mid X_i] = \mathbb{E} [ \mathbb{E} [Q_{Y_i \mid S_i,X_i}(Y_i \mid S_i,X_i) \mid S_i,X_i] \mid X_i] = \mathbb{E} [0 \mid X_i] = 0 \]

Combining the four steps (a)-(d) arrives at the desired equality between $\tau'$ and $\psi(Y_i,S_i,W_i,X_i)$.

Finally, note that the $\mathbb{V}_s$ is obtained by calculating the variance of the EIF (already written in Theorem 3):\footnote{We henceforth explicitly show the conditioning $P_{i}={\rm E}$ to be consistent with the notation in our Theorem statement.}

{ \[ \psi(Y_i,S_i,W_i,X_i,P_{i}) = \frac{(Y_i - \mu(S_i,X_i,{\rm E})) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))}+ \frac{W_i(\mu(S_i,X_i,{\rm E}) - \mu(1,X_i))}{\rho(X_i)} \]} { \[ - \frac{(1-W_i)(\mu(S_i,X_i,{\rm E}) - \mu(0,X_i))}{1-\rho(X_i)} + \mu(1,X_i) - \mu(0,X_i) - \tau, \] } i.e., \[ \mathbb{V}_{{\rm s}} = \bigg[\psi(Y_i,S_i,W_i,X_i,P_{i})^2 \bigg] =\mathbb{E} \bigg[ \left( \frac{(Y_i - \mu(S_i,X_i,{\rm E})) (\rho(S_i,X_i) - \rho(X_i))}{\rho(X_i)(1-\rho(X_i))} \right)^2 + \left( \frac{W_i(\mu(S_i,X_i,{\rm E}) - \mu(1,X_i))}{\rho(X_i)} \right)^2 \] \[ + \left( \frac{(1-W_i)(\mu(S_i,X_i,{\rm E}) - \mu(0,X_i))}{1-\rho(X_i)} \right)^2 + \left( \mu(1,X_i) - \mu(0,X_i) - \tau \right)^2 \bigg] \]

\[ =\mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2+ \left( \mu(1,X_{i}) - \mu(0,X_{i}) - \tau \right)^2 \right.\] \[\left. + \frac{W_{i}}{\rho(X_{i})^2} (\mu(S_{i},X_{i},{\rm E}) -\mu(1,X_{i}))^2 + \frac{1-W_{i}}{ (1 - \rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm E}) - \mu(0,X_{i}))^2 \right] \] \[ =\mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \left( \frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))} \right)^2+ \left( \mu(1,X_{i}) - \mu(0,X_{i}) - \tau \right)^2 \right.\] \[\left. + \frac{\rho(S_{i},X_{i})}{\rho(X_{i})^2} (\mu(S_{i},X_{i},{\rm E}) -\mu(1,X_{i}))^2 + \frac{1-\rho(S_{i},X_{i})}{ (1 - \rho(X_{i}))^2} (\mu(S_{i},X_{i},{\rm E}) - \mu(0,X_{i}))^2 \right] \] by the law of iterated expectations, and hence we have that

\[ \Delta=\mathbb{V}_{{\rm ns}} - \mathbb{V}_{{\rm s}} = \mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \left( \frac{\rho\left(S_i, X_i\right)}{\rho\left(X_i\right)^2}+\frac{1-\rho\left(S_i, X_i\right)}{\left(1-\rho\left(X_i\right)\right)^2} - \left(\frac{\rho(S_{i},X_{i})- \rho(X_{i})}{\rho(X_{i})(1-\rho(X_{i}))}\right)^2 \right) \right] \] \[ = \mathbb{E}\left[\sigma^2(S_{i},X_{i},{\rm E}) \frac{\rho\left(S_i, X_i\right)\left(1-\rho\left(S_i, X_i\right)\right)}{\rho\left(X_i\right)^2\left(1-\rho\left(X_i\right)\right)^2} \right] \]

$\square$

Proof of Theorem (ref): Consider part (i). By the law of iterated expectations conditional on $S_{i}$ and $X_{i}$, we have

align*[align* omitted — 392 chars of source]

By the proof of (ref) in Theorem (ref) where we don't use Surrogacy or Comparability, we get

align*[align* omitted — 550 chars of source]

The second equality in $\tau^{{\rm E}}=\tau^{{\rm O}}=\tau^{{\rm E},{\rm O}}$ is immediate based on only the law of iterated expectations. Finally, by the law of iterated expectations conditional on $X_{i}$, we have

align*[align* omitted — 403 chars of source]

By Assumption (ref) (unconfoundedness), we have

align*[align* omitted — 667 chars of source]

Undoing the law of iterated expectations give the desired result.

For parts (ii)-(iv), we prove (iv) first. By Assumption (ref) (unconfoundedness), we have \[ \tau=\mathbb{E}\left[\mathbb{E}\left[Y_{i}|W_{i}=1,X_{i},P_{i}={\rm E}\right]\mid P_{i}={\rm E}\right]-\mathbb{E}\left[\mathbb{E}\left[Y_{i}|W_{i}=0,X_{i},P_{i}={\rm E}\right]\mid P_{i}={\rm E}\right]. \] By iterated expectations, this is equal to

align*[align* omitted — 588 chars of source]

Thus, we have

align*[align* omitted — 633 chars of source]

We add and subtract \[ \mathbb{E}\left[\mathbb{E}\left[\mu(S_{i},X_{i},{\rm E})|W_{i}=1,X_{i},P_{i}={\rm E}\right]\mid P_{i}={\rm E}\right]-\mathbb{E}\left[\mathbb{E}\left[\mu(S_{i},X_{i},{\rm E})|W_{i}=0,X_{i},P_{i}={\rm E}\right]\mid P_{i}={\rm E}\right] \] to get

align*[align* omitted — 1,126 chars of source]

Rearranging the terms, we have

align[align omitted — 1,132 chars of source]

Next, by the definition of expectations,

align*[align* omitted — 393 chars of source]

Use this to write ((ref)) as {

align*[align* omitted — 932 chars of source]

Using the same argument we can write ((ref)) as

align*[align* omitted — 1,188 chars of source]

Combining the results for ((ref)) and ((ref)) leads to \[ \mathbb{E}\left[\left(\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\right)\cdot\frac{(1-\rho(S_{i},X_{i}))\cdot \rho(S_{i},X_{i})}{(1-\rho(X_{i}))\cdot \rho(X_{i})}\mid P_{i}={\rm E}\right] \] Collecting the last two terms, ((ref)) and ((ref)), we have }{

align*[align* omitted — 1,928 chars of source]

}{ Combining the terms together, we obtain the expression in (iv)

align*[align* omitted — 514 chars of source]

Finally for part (ii), under Assumption (ref) (Comparability), but not Assumption (ref) (Surrogacy), $\mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})=0$ and the result is immediate from (iv). For part (iii), under Assumption (ref) (Surrogacy), but not Assumption (ref) (Comparability), $\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})=0$ and the result is immediate from (iv). $\square$}

{\bf Proof of Lemma (ref)} We can identify, given overlap, the surrogate score $\rho(s,x)$, the propensity score $\rho(X)$, the surrogate index $\mu(s,x,{\rm O})$, and the joint distribution of $(S_{i},X_{i},P_{i})$. This implies that to derive upper and lower bounds we just need to derive upper and lower bounds for the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ for each value of $(s,x)$ and then integrate these bounds. We will demonstrate the sharpness of these bounds by showing that there exist data distributions consistent with all assumptions such that these bounds are achieved.

Part $(i)$: By Theorem (ref) the surrogacy bias can be characterized as \[\textrm{\rm surrogacy-bias}= \mathbb{E}\left[\left.\Bigl\{\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]. \] The data are not directly informative about the two conditional expectation $\mu(s,w,x,{\rm E})$ (because we do not observe the outcome in the experimental sample) beyond their relation to the surrogacy index: \[ \mu(s,x,{\rm O})=\rho(s,x) \mu(s,1,x,{\rm E})+(1-\rho(s,x)) \mu(s,0,x,{\rm E}),\qquad \forall s,x.\] This implies the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ can be written as \[\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})=\frac{\mu(s,x,{\rm O})}{\rho(s,x)}-\frac{\mu(s,0,x,{\rm E})}{\rho(s,x)}.\] Fixing $\mu(s,x,{\rm O})$, $\rho(s,x)$, and $\mu(s,0,x,{\rm E})$ this places no restrictions on the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ and thus no restrictions on the bias, and therefore any value for the treatment effect on the whole real line is consistent with the data in the absence of surrogacy.

Part $(ii)$: If the outcome is binary, then some values can be ruled out. Because $\mu(s,w,x,{\rm E})$ is the conditional expectation of the outcome given some conditioning variables, it obviously must be inside the interval $[0,1]$, and both $\mu(s,1,x,{\rm E})$ and $\mu(s,0,x,{\rm E})$ must lie inside the interval $[0,1]$. This directly implies that $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})\in[-1,1]$. However, we can sharpen these bounds exploiting the fact that $\mu(s,x,{\rm O})=\rho(s,x)\mu(s,1,x,{\rm E})+(1-\rho(s,x))\mu(s,0,x,{\rm E})$. This implies that

equation[equation omitted — 112 chars of source]

First consider the upper bound on $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$. The question is what the pairs of values $(\mu(s,1,x,{\rm E}),\mu(s,0,x,{\rm E}))$ are that both lie inside $[0,1]$, such that $\mu(s,x,{\rm O})=\rho(s,x)\mu(s,1,x,{\rm E})+(1-\rho(s,x))\mu(s,0,x,{\rm E})$ for given $\mu(s,x,{\rm O})$ and $\rho(s,x)$, and that maximize the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$. There are two possibilities. Either $\mu(s,x,{\rm O})\geq \rho(s,x)$ or $\mu(s,x,{\rm O})< \rho(s,x)$.

If $\mu(s,x,{\rm O})\geq \rho(s,x)$, then the smallest value for $\mu(s,0,x,{\rm E})$ such that the value for $\mu(s,x,{\rm E})$ implied by ((ref)) is less than or equal to one is $\mu(s,0,x,{\rm E})=(\mu(s,x,{\rm O})-\rho(s,x))/(1-\rho(s,x))$. This value has to be less than one by the assumption that there is a pair of values $(\mu(s,0,x,{\rm E}),\mu(s,1,x,{\rm E}))$ that satisfies ((ref)). In this case upper bound for the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ is equal to $(1-\mu(s,x,{\rm O}))/(1-\rho(s,x))$. If $\mu(s,x,{\rm O})\leq \rho(s,x)$, then the largest value for $\mu(s,1,x,{\rm E})$ such that $\mu(s,0,x,{\rm E})$ is nonnegative is $\mu(s,x,{\rm O})/\rho(s,x)$. In that case the upper bound for the difference $\mu(s,1,x,{\rm E})-\mu(s,0,x,{\rm E})$ is equal to $\mu(s,x,{\rm O})/\rho(s,x)$.

In summary, to demonstrate sharpness, consider the following data distributions:

If \(\mu(s, x, {\rm O}) \geq \rho(s, x)\), set \(\mu(s, 0, x, {\rm E}) = \frac{\mu(s, x, {\rm O}) - \rho(s, x)}{1 - \rho(s, x)}\) and \(\mu(s, 1, x, {\rm E}) = 1\).

If \(\mu(s, x, {\rm O}) < \rho(s, x)\), set \(\mu(s, 0, x, {\rm E}) = 0\) and \(\mu(s, 1, x, {\rm E}) = \frac{\mu(s, x, {\rm O})}{\rho(s, x)}\)

In both cases, these distributions are admissible under our assumptions, and also achieve the bounds, demonstrating that the bounds are sharp.

Therefore, the sharp upper bound is \[ \Delta^U_S(s,x)= \left\{

array[array omitted — 195 chars of source]

\right.\]\[ =\min\left(\frac{\mu(s,x,{\rm O})}{\rho(s,x)},\frac{1-\mu(s,x,{\rm O})}{1-\rho(s,x)}\right). \] The proof for the lower bound follows the same argument.

Part $(iii)$: \[\textrm{\rm surrogacy-bias}= \mathbb{E}\left[\left.\Bigl\{\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] \] \[\leq \mathbb{E}\left[\left.\left|\Bigl\{\mu(S_{i},1,X_{i},{\rm E})-\mu(S_{i},0,X_{i},{\rm E})\Bigr\} \right|\cdot\left|\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|\right|P_{i}={\rm E}\right] \] \[\leq c \cdot \mathbb{E}\left[\left.\frac{\rho(S_{i},X_{i})\cdot(1-\rho(S_{i},X_{i}))}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] \] The upper bound can be achieved by setting $\mu(s,0,x,{\rm E})=\mu(s,x,{\rm O})-c\cdot \rho(s,x)$ and $\mu(s,1,x,{\rm E})=\mu(s,0,x,{\rm E})+c,$ These distributions are admissible under our assumptions, and hence sharpness is obtained. We can likewise obtain the lower bound.

$\Box$

{\bf Proof of Lemma (ref)} We show that the derived bounds are sharp by demonstrating that there exist data distributions consistent without assumptions that achieve these bounds. $(i)$ In the absence of Comparability the data imply no restrictions on the values for $\mu(s,x,{\rm E})$, and so as long as there is some difference between $\rho(s,x)$ and $\rho(x)$ there is no bound on the bias. \\ $(ii)$ If the outcomes are binary, the only restrictions implied on $\mu(s,x,{\rm E})$ are that all values lie inside $[0,1]$. The upper bound comes from imputing 1 for $\mu(s,x,{\rm E})$ if $\rho(s,x)>\rho(x)$ and $0$ if $\rho(s,x)<\rho(x)$, a choice of distribution that is admissible. This directly implies the bounds on the bias. \\ $(iii)$ \[\textrm{\rm comparability-bias} =\mathbb{E}\left[\left.\Bigl\{ \mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]. \] Then \[\left|\mathbb{E}\left[\left.\Bigl\{ \mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})\Bigr\}\cdot\frac{\rho(S_{i},X_{i})-\rho(X_{i})}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right]\right| \] \[\leq\mathbb{E}\left[\left.\left|\Bigl\{ \mu(S_{i},X_{i},{\rm E})-\mu(S_{i},X_{i},{\rm O})\Bigr\}\right|\cdot\frac{\left|\rho(S_{i},X_{i})-\rho(X_{i})\right|}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] \] \[\leq c\cdot \mathbb{E}\left[\left.\frac{\left|\rho(S_{i},X_{i})-\rho(X_{i})\right|}{\rho(X_{i})\cdot(1-\rho(X_{i}))}\right|P_{i}={\rm E}\right] .\] The upper bound can be attained by setting \[ \mu(s,x,{\rm E})= \left\{

array[array omitted — 126 chars of source]

\right. \] and similarly for the lower bound. $\Box$

C. Illustration of Bias Bounds Calculation

We will provide a simple illustration of how the theoretical bias bounds we calculated in Section (ref) look like in practice. We focus on the employment outcome to illustrate the surrogacy bias and comparability bias bounds in the binary case (Case (ii)).

Table (ref) and (ref) show the bounds on the treatment effects using the Influence Function Estimator under potential violations of Surrogacy and Comparability, respectively.\footnote{If we are interested in conducting inference on the partial identification bounds, we can take the approach illustrated in, e.g., ImbensManski2004,molinari2020microeconometrics.} This demonstrates that in the binary outcome of employment, the sign can still be credibly inferred under the latter half even under surrogacy violation. The comparability bias seems to be non-negligible, part of our design of choosing Riverside (experimental data) due to its unique "jobs first" approach, in contrast to the "human capital" approach used in LA, San Diego, and Alameda counties (observational data). Further work must be done to ensure cases where comparability bias is minimal. We can similarly compute non-binary outcomes like Earnings, with some plausible range of user-specified parameter $c$ (Case (iii) in Section (ref)).

table[table omitted — 477 chars of source]
table[table omitted — 486 chars of source]