Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,304 characters · 13 sections · 47 citation commands
The Experimental Selection Correction Estimator: Using Experiments to Remove Biases in Observational Estimates
\thispagestyle{empty}
\vskip0.5cm Keywords:\ causality, surrogates, observational studies, long-term outcomes, control functions
\baselineskip=20pt \setcounter{page}{1}
As observational data become more widely available, researchers seeking to estimate treatment effects increasingly have access to two types of data: (i) large observational datasets where treatments and a broad range of outcomes are observed, but treatment is not randomized and (ii) smaller experimental datasets where treatment is randomly assigned, but only a subset of outcomes are observed. For example, in the context of education, many analysts have been interested in identifying the causal effects of classroom sizes in elementary school on high school graduation rates. Observational data with information on class sizes and graduation rates are now widely available from school districts' administrative records. But causal inference using these data is challenging because of selection biases arising from non-random assignment to classrooms. Causal inference is more straightforward in experimental data – such as the widely studied Project STAR class size experiment (e.g., krueger1999experimental) – but experimental datasets often do not contain information on outcomes such as graduation rates because they are observed with long lags.
The most common method of identifying the causal effect of a treatment ({\it e.g.}, class size reduction) on the primary outcome of interest ({\it e.g.}, graduation rates) in such settings is to use secondary intermediate outcomes that are observed in the experimental data ({\it e.g.}, test scores) as statistical surrogates prentice1989surrogate, athey2019surrogate. The surrogate approach, illustrated in Figures 1a-b below, uses the observational dataset to estimate the relationship between the primary outcome ($Y^{\rm P}_i$) and secondary outcome ($Y^{\rm S}_i$), and then estimates the impact of the treatment of interest ($W_i$) on $Y^{\rm P}_i$ based on that relationship. Under the surrogacy assumptions that (i) $W_i$ only affects $Y^{\rm P}_i$ through its impact on $Y^{\rm S}_i$ and (ii) there are no unobserved confounders that affect the relationship between $Y^{\rm S}_i$ and $Y^{\rm P}_i$ in the observational sample, this approach provides an unbiased estimate of the effect of $W_i$ on $Y^{\rm P}_i$.
The surrogate approach has been applied in many fields, from economics to product testing to public health alonso2006unifying, adams2006overweight, d2006surrogate, gupta2019top. Yet there remains concern that the surrogacy assumptions may be violated in these settings. For example, test scores are a widely used surrogate in labor economics, but researchers have identified other pathways through which childhood interventions affect long-term outcomes outside test scores, such as non-cognitive skills heckman2006effects, chetty2011does.
In this paper, we develop an “Experimental Selection Correction” (ESC) estimator that identifies the effect of $W_i$ on $Y^{\rm P}_i$ even when the surrogacy assumptions are violated, as illustrated in Panels C and D of Figure 1. Our estimator relies on more information than the surrogacy approach: it requires that the observational dataset contains information not just on the primary and secondary outcomes, $Y^{\rm S}_i$ and $Y^{\rm P}_i$, but also on treatment $W_i$ (with variation in treatment across observations). With this additional information, we show how one can identify the effect of $W_i$ on $Y^{\rm P}_i$ under strictly weaker assumptions than those required for the surrogacy approach.
To illustrate the general information scheme we analyze, consider a setting with two datasets: (i) the Project STAR experimental data, where class size is randomized and we observe test scores ($Y^{\rm S}_i$) but not high school graduation rates ($Y^{\rm P}_i$), and (ii) observational data from the New York City school district, in which class size is observed but not randomized (and hence likely to be correlated with both observed and unobserved characteristics) and we observe both test scores and graduation rates.
In the STAR experimental data, we can estimate the treatment effect of small class size ($W_i$) on 3rd grade test scores by regressing test scores on an indicator for being assigned to a small class (with 7 fewer students on average) in 3rd grade. Column 1 of Table 1 shows that being assigned to a small class in 3rd grade increases students’ end-of-3rd-grade test scores by 0.19 standard deviations (SD). We cannot, however, estimate the effect of class size on high school graduation in the STAR data, because we do not observe graduation in the STAR sample.\footnote{Researchers attempted to follow the STAR students longitudinally, but were only able to collect information on high school graduation for 43% of students, whose characteristics are not representative of the experimental sample as a whole. This underscores the challenges of tracking primary outcomes in experiments and motivates the approach we take here of combining observational administrative records and experimental data.}
In the observational NYC sample, estimating an analogous OLS regression of test scores on an indicator for being assigned to a small class yields an estimate of $-0.12$ SD (s.e. 0.01, Column 2 of Table 1). Children in smaller classes are also 1.76 percentage points (s.e. 0.29) less likely to graduate from high school. These negative estimates of the causal effect of class size reductions are implausible both in the light of the positive experimental Project STAR estimates and based on {\it a priori} beliefs. Of course, the OLS estimates may be confounded because class size is not randomly assigned in NYC. For example, students with needs for additional educational support may be assigned to smaller classes. Our goal is to obtain an unconfounded estimate of $W_i$ on $Y^{\rm P}_i$, {\it i.e.}, to fill in the lower left box in Table 1.
We obtain an unconfounded estimate of the effect of $W_i$ on $Y^{\rm P}_i$ in the observational data by using the difference in the distribution of test scores ($Y^{\rm S}_i$) conditional on treatment in the observational and experimental samples to adjust for selection. In linear models, the ESC estimator can be implemented in three straightforward steps (see Appendix for code). First, we estimate the effect of class size on test scores ($\tau^{\rm S}$) in the experimental sample using a linear regression, as in Column 1 of Table 1.\footnote{We focus on the case where treatment is randomly assigned in an experiment, but any quasi-experimental research design that yields an unbiased estimate of the treatment effect on the secondary outcome $\tau^{\rm S}$ can be used to implement the ESC estimator.} Second, for all students in the observational sample, we calculate the difference between the secondary outcome (test score) and the predicted test score based on the student’s class size (the residual $\alpha^{\rm S}_i= Y^{\rm S}_i - \tau^{\rm S} W_i$), with the parameter $\tau^{\rm S}$ in the prediction model coming from the experimental sample. Finally, we regress the primary outcome (graduation rates) on treatment (class size) in the observational data, controlling for the residual $\alpha^{\rm S}_i$.
We show that this control function approach identifies the causal effect of $W_i$ on $Y^{\rm P}_i$ under three assumptions: (i) random assignment (or unconfoundedness) in the experimental sample; (ii) a standard external validity assumption; and (iii) a new assumption that we term latent unconfoundedness. External validity requires that (conditional on pretreatment observables), the treatment effect in the experimental sample is the same as the treatment effect in the population represented by the observational sample shadishcookcampbell, hotz2005predicting. Latent unconfoundedness requires that the unobserved confounders that affect the primary outcome (graduation rates) are the same as those that affect the secondary outcome (test scores). Under this assumption, the difference between the actual secondary outcome in the observational data and the predicted secondary outcome based on the experimental estimate ($\alpha^{\rm S}_i$) fully captures any selection bias that affects the primary outcome. Thus, controlling for $\alpha^{\rm S}_i$ is sufficient to identify the causal effect of $W_i$ on $Y^{\rm P}_i$. Intuitively, $\alpha^{\rm S}_i$ functions as a selection correction, similar to parametric selection correction approaches dating to heckman1979sample and control function methods (heckman1985alternative, imbens2009identification, wooldridge2015control).
The main theoretical result of this paper is that the treatment effect of $W_i$ on $Y^{\rm P}_i$ is point-identified under latent unconfoundedness, external validity, and random assignment in the experimental sample (without any functional form or distributional assumptions). We also present a control function approach to estimation for the general, nonlinear case. A corollary of our main result is that if an observational estimate of the treatment effect on the secondary outcome ($Y^{\rm S}_i$) is the same as the experimental estimate, then under linearity, latent unconfoundedness and external validity together imply that the observational estimator is unconfounded for the primary outcome. Many empirical studies show that an observational and experimental estimator yield similar estimates for secondary outcomes and then use this as a heuristic justification for estimating impacts on a broader range of primary outcomes using the observational estimator ({\it e.g.}, chetty2014onemeasuring, bleemer2022affirmative). Our analysis makes precise the conditions—most importantly, latent unconfoundedness across outcomes—under which this heuristic is justified.
We also propose falsification tests for the underlying identifying assumptions that make use of additional “holdout” post-treatment outcomes ($Y^H_i$) observed in both datasets. One could use these measures as additional secondary outcomes. An alternative is to use them to validate the estimator instead of using them to implement the selection correction itself. Tests of whether our ESC estimator that uses 3rd grade scores for selection correction matches experimental estimates for 4th and 5th grade test scores (holdout outcomes) serve as tests for whether the underlying latent unconfoundedness and external validity assumptions jointly hold.
We apply the experimental selection correction estimator to estimate the causal effects of 3rd grade class size reduction on high school graduation rates in the New York City data, using end-of-3rd-grade test scores as the secondary outcome for the selection correction. We use the Tennessee STAR sample as the experimental sample in which we estimate the treatment effect of class size reduction on 3rd grade test scores.\footnote{Naturally, one may have concerns about the external validity of the Tennessee sample for the New York City data. Both samples reflect a relatively low-income population with fairly similar demographic characteristics. Furthermore, we show that adjusting for remaining differences in observable demographics does not affect our conclusions meaningfully.}
As discussed above, standard OLS regression estimates of 3rd grade test scores on an indicator for small class size in the NYC data yield a treatment effect estimate of $-0.12$ SD. The ESC estimator (shown in Column 3 of Table 1) yields an estimate of $0.19$ SD, coinciding with the STAR experimental estimate by construction. The ESC treatment effect estimates for test scores in grades 4-8 (holdout outcomes) are nearly identical to the STAR experimental estimates (Figure (ref)). Notably, they capture the well-known “fadeout” pattern on test score impacts documented in prior work (deming2009early,chetty2011does,cascio2012knowledge). These results support the identification assumptions underlying our method and more broadly serve to validate the ESC approach.
Finally, the ESC estimator implies that assignment to a small class in 3rd grade (which has 25% fewer students on average) increases the probability of graduating from a New York City public high school by $0.69$ percentage points (pp), relative to a sample mean of $51.4\%$. This estimate is one of the first estimates of the causal effect of class size reduction on high-school graduation rates in the U.S.
In contrast, the standard OLS estimator yields significant negative estimates on test scores in later grades and on high school graduation rates. When we control for observable characteristics, the OLS estimates remain negative, while the ESC estimates remain similar to the experimental estimates (Figure (ref)). These findings demonstrate how our proposed experimental selection correction can detect and adjust for selection biases that are difficult to address with conventional methods in observational data without relying on strong surrogacy assumptions.
In addition to the literature on statistical surrogates, our analysis relates to other studies that have examined similar observation schemes, including rosenman2018propensity, rosenman2020combining, kallus2020role, mealli2013using and imbens2025long. rosenman2018propensity focuses on the problem where assignment is unconfounded in both samples and combining the samples increases precision. rosenman2020combining allow for unobserved confounders in the observational sample and consider shrinkage estimators to decrease bias. kallus2020role analyze the case where assignment in the combined experimental and observational sample is unconfounded, but not in each sample separately. kallus2018removing focus on a case where the same variables (including the primary outcome) are observed in the two samples, but where unconfoundedness is violated in the observational sample and the experimental sample is used to estimate bias. mealli2013using focuses on an instrumental variables setting where the presence of multiple outcomes improves estimates. bhattacharya2013evaluating proposes combining experimental and observational data for estimating policy rules in a different setting. Our approach is conceptually closely related to the Changes-in-Changes (CIC) estimator in athey2006identification. For a unit in the observational sample, the control function is essentially the rank of the secondary outcome in the distribution of secondary outcomes in the experimental sample with the same treatment. Under our maintained assumptions here, differences in the estimated effect of the treatment between the experimental and observational sample are attributed to violations of unconfoundedness in the observational data.
The paper is organized as follows. Section 2 analyzes a linear model that captures the intuition underlying our approach. Section 3 presents the identification result for the general case. Section 4 discusses estimation. Section 5 presents the application. Section 6 concludes.
In this section, we introduce our key identifying assumption and a control function estimator in the context of linear models. The linear case simplifies exposition and captures the key ideas that apply in more general models.
Using the potential outcome set up for observational studies introduced by rubin1974estimating (see imbens2015causal for a textbook discussion), let the pair of potential outcomes for the primary outcome for unit $i$ be denoted by $Y_{i}^{\rm P}(0)$ and $Y_{i}^{\rm P}(1)$, where the superscript “${\rm P}$” stands for “Primary”. In many applications, $Y_{i}^{\rm P}$ is a long-term outcome; in our application, it is a binary indicator for high school graduation. The treatment received by unit $i$ is $W_{i}\in\{0,1\}$. In our application, $W_i$ is an indicator for small class size in third grade, with $W_i=1$ indicating a small class size and $W_i=0$ indicating a regular class size. There is also a secondary outcome, with the pair of potential outcomes for unit $i$ denoted by $Y_{i}^{\rm S}(0)$ and $Y_{i}^{\rm S}(1)$, where the superscript “${\rm S}$” stands for “Secondary”. In our application, $Y_{i}^{\rm S}$ is a student's end-of-third-grade test score.
We focus in this section on the case where both the primary and secondary outcomes are scalars, but both may be vector-valued ({\it e.g.,} test scores in multiple grades could be used as secondary outcomes).
The realized values for the primary and secondary outcomes are $Y^{\rm P}_i\equiv Y_i^{\rm P}(W_i)$ and $Y^{\rm S}_i\equiv Y_i^{\rm S}(W_i)$. We may also observe pretreatment variables, denoted by $X_i$, that are known not to be affected by the treatment.
We focus on identifying the average treatment effect on the primary outcome,
although other estimands such as the average effect on the treated can be accommodated in our set up as well. The average treatment effect on the secondary outcome, $ \tau^{\rm S}\equiv \mathbb{E}\left[ Y^{\rm S}_i(1)-Y^{\rm S}_i(0)\right],$ is, for the purpose of the current study, not of intrinsic interest.
We have two samples to draw on to estimate $\tau^{\rm P}$, as in the literature on combining datasets, {\it e.g.,} hotz2005predicting, ridder2007econometrics, pearl2014external. The first is an observational study that is a random sample from the population of interest. For all units in this observational sample, we observe the quadruple $(W_i,Y^{\rm S}_i,Y^{\rm P}_i,X_i)$.
The second sample is a possibly selective sample from the same population, with random assignment of treatment $W_i$. For all units in this experimental sample, we observe the triple $(W_i,Y^{\rm S}_i,X_i)$, but not the primary outcome $Y^{\rm P}_i$.
Let $G_i\in\{{\rm E},{\rm O}\}$, be an indicator for the sample a unit is drawn from. Then we can conceptualize the combined sample as a random sample of size $N$ from an artificial super-population for which we observe the quintuple $(W_i,G_i,Y^{\rm S}_i,Y^{\rm P}_i\boldsymbol{1}_{G_i={\rm O}},X_i)$, where $\boldsymbol{1}_{G_i={\rm O}}$ is a binary indicator, equal to 1 if $G_i={\rm O}$ and equal to 0 if $G_i={\rm E}$.
Suppose we have a linear model for the secondary potential outcomes in combination with a constant treatment effect $\tau^{\rm S}$: \[ Y^{{\rm S}}_i(0)=X_i^\top\gamma^{\rm S}+\alpha_i^{\rm S},\hskip1cm Y^{\rm S}_i(1)=Y^{{\rm S}}_i(0)+\tau^{\rm S} .\] This model holds in both the experimental and observational samples. However, the properties of the unobserved component $\alpha^{\rm S}_i$ differ between the two samples. In the experimental sample, randomization guarantees the following conditional independence condition:\footnote{In fact the randomization implies an even stronger condition, $W_i \perp\!\!\!\perp\alpha_i^{\rm S},X_i|G_i={\rm E}$, but we do not need that condition here.} \[ W_i\ \perp\!\!\!\perp\ \alpha_i^{\rm S}\ \Bigl|\ X_i,G_i={\rm E}.\] In the observational study, the same conditional independence does not generally hold: \[ W_i\ \not\!\perp\!\!\!\perp\ \alpha_i^{\rm S}\ \Bigl|\ X_i,G_i={\rm O}.\] We specify a similar linear model for the primary outcome, but allow the coefficients to be different from those of the model for the secondary outcome: \[ Y^{{\rm P}}_i(0)=X_i^\top\gamma^{\rm P}+\alpha_i^{\rm P},\hskip1cm Y^{{\rm P}}_i(1)= Y^{{\rm P}}_i(0)+\tau^{\rm P}.\] Again, the unobserved component might be correlated with the treatment in the observational sample: \[W_i\ \not\!\perp\!\!\!\perp\ \alpha_i^{\rm P}\ \Bigl|\ X_i,G_i={\rm O},\] that is, $W_i$ is again endogenous.
To identify the treatment effect on the primary outcome $\tau^{\rm P}$ in the observational sample, we make the following assumption that links the endogeneity problems for the primary and secondary outcomes:
This assumption requires that the component of the residual in the primary outcome that is not explained by the residual in the secondary outcome, $\varepsilon_i^{\rm P}\equiv \alpha_i^{\rm P}-\mathbb{E}[\alpha^{\rm P}_i|\alpha^{\rm S}_i]$, is orthogonal to treatment. The key substantive restriction captured by this condition is that the unobserved confounders that affect the secondary outcome are the same as those that affect the primary outcome. This assumption, which we term latent unconfoundedness, is the key to identifying $\tau^{\rm P}$ in the general case below as well.
We now show how this latent unconfoundedness assumption allows us to identify $\tau^{\rm P}$ using a simple control function approach in the linear case. First, we exploit randomization in the experimental sample to estimate $\tau^{\rm S}$ and $\gamma^{\rm S}$ using ordinary least squares regression. Denote these least squares estimates by $\hat\tau^{{\rm S}}$ and $\hat\gamma^{{\rm S}}$.
Next, we estimate the residual $\alpha_i^{\rm S}$ for the units in the observational sample as
If the assignment to treatment in the observational sample were random (and assuming the linear model is correct), the population value of these residuals $\alpha^{\rm S}_i$ would be uncorrelated with the treatment indicator in the observational sample.
When treatment assignment is non-random, we can use the association between the secondary outcome residuals $\alpha_i^{\rm S}$ and the treatment to correct for selection bias in the estimating equation for the primary outcome. We do so by including $\alpha_i^{\rm S}$ as a control variable in an ordinary least squares regression of the primary outcome on treatment. To see why this yields a consistent estimate of $\tau^{\rm P}$, observe that we can use the linear representation in ((ref)) to write the primary outcome as:
Because the error term $\varepsilon_i^{\rm P}$ is orthogonal to treatment in this specification, estimating this equation using OLS yields a consistent estimator for $\tau^{\rm P}$ under our assumptions.
In this section, we generalize the linear example above to accommodate (i) non-linear models and (ii) multiple secondary outcomes.
We are interested in causal estimands defined for the population of interest. Such estimands include simple average treatment effects, but more generally also the average effect of a policy that assigns the treatment to individuals in this population on the basis of covariates ({\it e.g.}, manski2004statistical, dehejia2005program, hirano2009asymptotics, athey2017efficient, zhou2018offline). For expositional simplicity, we focus here on average treatment effects. Define
to be the average effect of the treatment on outcome $t\in\{{\rm S},{\rm P}\}$ for group $g\in\{{\rm O},{\rm E}\}$. The superscripts on the estimands denote the outcome, and subscripts denote the population. The primary estimand we focus on in this paper is the average effect of the treatment on the primary outcome in the observational study population:
where we drop the subscript and superscript to simplify the notation.
There are three key features of our set up. First, we are interested in the population that the units in the observational study were drawn from. That is, the observational study has external validity.
This can be thought of as simply defining the estimand in terms of the population distribution underlying the observational sample.
Second, we maintain throughout the paper the assumption that the treatment in the experimental sample is unconfounded.
Although internal validity of the experimental sample is guaranteed by design, external validity of the experimental study does not follow. We assume that conditional on the pretreatment variables we have external validity (hotz2005predicting):
This assumption implies that if we find systematic differences between in differences in average outcomes by treatment status conditional on covariates between the experimental and observational sample, these differences must arise from violations of unconfoundedness for the observational sample.
The first result is that these three maintained assumptions are in general not sufficient for point-identification of the average effect of interest. Of course this does not mean that these assumptions do not have any identifying power. They do in fact imply non-trivial identified sets in the spirit of the work by (manski_bounds).
The proof for this result is given in the appendix.
Next, let us briefly mention a common assumption that we do {\it not} wish to make in this context. Specifically, we consider the assumption that assignment in the observational study is unconfounded. For $w=0,1$,
This assumption is made, for example, in rosenman2018propensity. This assumption is sufficient for identification of $\tau$, but it is stronger than necessary. Intuitively it implies that we do not need the experimental sample for identification because under unconfoundedness the observational sample is sufficient for identification of the average treatment effect. However, the experimental sample may still be useful for precision.
Suppose that we reject the combination of Assumptions (ref)-(ref) and unconfoundedness ((ref)). If we maintain unconfoundedness in the experimental sample (Assumption (ref)), it must be that either conditional external validity in the experimental study (Assumption (ref)), or unconfoundedness in the observational study ((ref)) must be violated. In many cases we may wish to maintain conditional external validity and interpret the finding that the combination does not hold as evidence that unconfoundedness in ((ref){) does not hold for the observational study.
The fundamental idea behind our approach, although not the implementation, can be seen as related to that in a Difference-In-Differences (cardmariel, cardkrueger1, angrist2008mostly) set up where the initial (pre-treatment) differences between a treatment and control group are used to adjust post-treatment differences between the treatment and control group. More specifically, it relates to the Changes-In-Changes approach in athey2006identification where functional form assumptions are avoided. Here initial differences in treatment effects between an experimental and observational study are used to adjust subsequent treatment effects for the observational study.
They key additional assumption that links the biases between adjusted comparisons for the primary and secondary outcomes, is the following.
This assumption is both novel as well as critical in the current discussion, so we offer some remarks.
To highlight the link to the control function literature (heckman1979sample, heckman1985alternative, imbens2009identification, wooldridge2010econometric, athey2006identification, kline2019heckits, mogstad2018using, mogstad2018identification, wooldridge2015control), let us model the primary and secondary potential outcomes as \[Y^{\rm P}_i(w)=h^{\rm P}(w,\nu_i,X_i),\hskip1cm {\rm and}\ \ Y^{\rm S}_i(w)=h^{\rm S}(w,\eta_i,X_i) ,\] with the function $h^{\rm S}(w,\eta,x)$ strictly monotone in $\eta$. In the context of this model we can write the latent unconfoundedness assumption as \[ W_i\ \perp\!\!\!\perp\ \nu_i\ \Bigl|\ X_{i},\eta_i,G_i={\rm O}.\] Although it is not generally true that $W_i \perp\!\!\!\perp \nu_i| X_{i},G_i={\rm O}$ (without conditioning on $\eta_i$), adding $\eta_i$ to the conditioning set restores the exogeneity of $W_i$ in the observational sample.
It is useful to contrast this with a control function in a nonparametric instrumental variables setting ({\it e.g.,} imbens2009identification), where the two models are \[Y^{\rm P}_i(w)=h^{\rm P}(w,\nu_i,X_i),\hskip1cm {\rm and}\ \ W_i(z)=r(z,\eta_i,X_i) ,\] with $r(z,\eta,x)$ strictly monotone in $\eta$. The key assumption here is that \[ W_i\ \perp\!\!\!\perp\ \nu_i\ \Bigl|\ X_{i},\eta_i.\] The model relating the outcome of interest and the endogenous regressor is essentially the same in the two settings, $Y^{\rm P}_i(w)=h^{\rm P}(w,\nu_i,X_i)$. In both cases we address the endogeneity by conditioning on an additional variable, the control variable $\eta_i$. This control variable is estimated using an auxiliary model. This auxilliary model differs between the set up in the current paper and the instrumental variables setting imbens2009identification. In the Imbens-Newey nonparametric instrumental variables setting we model the relation between the endogenous regressor and an additional variable, the instrument, and deriving the control variable from that relation. In the current setting we model the relation between the secondary outcome and the endogenous regressor and deriving the control variable from that relation. In both cases the auxiliary model has a strict monotonicity assumption. This comparison shows one of the limitations of the approach: the unobserved confounder $\eta_i$ cannot have a dimension higher than that of the secondary outcome.
Formally, adding Assumption (ref) (latent unconfoundedness) to Assumptions (ref)-(ref) allows us to point-identify the average effect of interest. The following theorem states our main identification result.
There is an interesting connection between Assumptions (ref)-(ref) and the Missing At Random (MAR) assumption in the missing data literature (rubin1976inference, little2019statistical, rubin2004multiple).
Because $G_i={\rm E}$ is equivalent to an indicator that $Y^{\rm P}_i$ missing, and because $W_i$, $X_i$, and $Y^{\rm S}_i$ are observed for all individuals in the sample, the conditional independence in ((ref)) is equivalent to a MAR assumption. The result does not go the other way around. The MAR assumption by itself has no testable implications, but the combination of Assumptions (ref)-(ref) and (ref) does imply some inequality restrictions on the joint distribution of the observed variables. kallus2020role starts with a MAR assumption, and uses that in combination with an unconfoundedness assumption on the full sample to identify the average effect of the treatment for the full sample.
In this section, we discuss estimation and inference. There are multiple approaches here, some of which we discussed in the examples in Section (ref). These strategies include imputation, weighting, control function methods, and influence-function based methods. Because the model is just-identified, all four of these methods are first-order equivalent, although they will have different finite sample properties, see newey1994asymptotic, chen2018overidentification for a general discussion.
We focus here on the control function approach that is special to this setting with observational and experimental data. In the appendix, we discuss the imputation, weighting, and influence function approaches, which closely resemble their equivalents in standard unconfoundedness settings.
In the control function approach, we directly estimate the unobserved confounder. We then estimate the average treatment effect in the observational sample adjusting for both the observed covariates and the estimated confounder. The procedure consists of three steps.
In the first step, we estimate the conditional cumulative distribution function of the secondary outcome, conditional on the treatment and pre-treatment variables, in the experimental sample: \[ F_{Y^{\rm S}|W,X}(y^{\rm S}|w,x)\equiv {\rm pr}(Y^{\rm S}_i\leq y^{\rm S}|W_i=w,X_i=x,G_i={\rm E}).\] Note that if the secondary outcome is a vector, this is a vector of conditional cumulative distribution functions, demonstrating that multiple secondary outcomes can weaken the identifying assumptions.
In the second step, we calculate for all units in the observational sample the control variable as \[ \eta_i= F_{Y^{\rm S}|W,X}(Y_i^{\rm S}|W_i,X_i).\]
In the third step, we estimate the adjusted difference \[ \mathbb{E}\left[\left. \mathbb{E}\left[\left.Y^{\rm P}_i\right|W_i=1,X_i,G_i={\rm O}\right] - \mathbb{E}\left[\left.Y^{\rm P}_i\right|W_i=0,X_i,G_i={\rm O}\right] \right| G_i={\rm O}\right],\] which by the assumptions in Theorem (ref) is equal to the average causal effect $\tau$. Here we can use any of the conventional methods for estimating average treatment effects under unconfoundedness for the observational data (including matching, regression, inverse propensity score weighting, augmented inverse propensity score weighting, or doubly robust methods), where we use the combination of the pretreatment variables $X_i$ and the estimated control function $\eta_i$ as the variables to be adjusted for. For example, using a imputation/regression approach, one would estimate the conditional mean of the primary outcome in the observational sample given treatment status, control variable, and pre-treatment variables: \[\gamma(w,h,x)\equiv \mathbb{E}\left[\left. Y^{\rm P}_i\right|W_i=w, \eta_i=h,X_i=x,G_i={\rm O}\right].\] These estimated conditional means would then be use to estimate the average treatment effect $\tau$ as \[ \hat\tau^{\rm cf}=\frac{1}{N^{\rm O}_1}\sum_{i:G_i={\rm O}} W_i \hat\gamma(1,\hat\eta_i,X_i)- \frac{1}{N^{\rm O}_0}\sum_{i:G_i={\rm O}} (1-W_i) \hat\gamma(1,\hat\eta_i,X_i),\] where $N^g_w=\sum_{i=1}^N {\bf 1}_{W_i=w,G_i=g}.$ Under standard conditions ({\it e.g., } newey1994asymptotic, chen2018overidentification) this estimator will be semiparametrically efficient and asymptotically linear and normally distributed. From newey1994asymptotic it follows that the control function estimator is first order equivalent to the imputation, weighting, and influence function estimators, and so one can use the semiparametric efficiency bound for variance estimation. Alternatively, one can use the regular bootstrap.
In this section, we evaluate the performance of our approach by estimating the long-term impacts of reducing class sizes in elementary school. Many experimental and quasi-experimental studies have analyzed the impacts of educational inputs – such as class size, teacher quality, and resources – on test scores krueger1999experimental,kane2008estimating,biasi2025works. Meanwhile, observational data with information on educational inputs as well as long-term outcomes such as high school graduation rates from school districts’ administrative records have become widely available. We combine these two sets of information to estimate the effects of class size on high school graduation rates.\footnote{A few studies have linked experimental data to administrative data from tax records and other sources to measure impacts on outcomes such as college attendance rates and earnings (see e.g., Chetty et al. 2011 STAR, Dynarski et al STAR, Fredrikkson et al. QJE on class size effects). However, such linkages remain challenging and relatively rare. Our objective here is to show how one can make progress in identifying treatment effects of interest even when direct measurement of primary outcomes in the experimental sample is infeasible.} We first describe the data we use and then present results.
We combine information from two datasets: experimental data from Tennessee STAR and observational data from the New York City public school district.
Experimental Sample: Tennessee STAR. The STAR experiment was conducted at 79 low-income public schools in Tennessee between 1985-89. In the 1985-86 school year, 6,323 kindergarten students in participating schools were randomly assigned to a small (target size 13-17 students) or regular-sized (20-25 students) class within their schools. An additional 5,248 children joined the 1985-86 entry cohort at the participating schools after kindergarten in grades 1-3. These new entrants were also randomly assigned to small vs. large classrooms within school upon entry. Students were intended to remain in the same class type (small vs. large) through 3rd grade, at which point all students returned to regular class sizes.
In each year from grades 3-8, STAR students were administered standardized tests that measure performance in math and reading. We standardize the average of math and reading scores to have mean 0 and standard deviation 1 within each grade among students in the STAR sample. We also observe information on students’ race and ethnicity, sex, and eligibility for free or reduced-price lunch (an indicator for having low-income parents). For further information on the STAR experiment, see word1990state , krueger1999experimental, and chetty2011does.
Observational Sample: New York City. We obtain observational information from the administrative records of the New York City public school district for 1.76 million children in grades 3-8 between the 1991-2009 school years. Starting from the raw data, we impose the same sample restrictions as in chetty2014onemeasuring – such as excluding special education classrooms and classrooms with less than 10 or more than 50 students – and additionally limit the sample to students for whom we observe test scores throughout grades 3-8. For comparability to the STAR treatment of dichotomous assignment to small vs. large classes, we define a “small” class in NYC as one with 26 (the sample median) or fewer students in third grade.
We observe math and reading test scores at the end of grades 3-8, which we standardize to have mean 0 and standard deviation within each grade in the NYC sample.\footnote{Chetty et al. (2014) show that the within-grade variation in achievement in the NYC school district is comparable to the within-grade variation in other urban school districts nationwide, and hence is likely comparable to that in the STAR population, which exhibits broadly similar socioeconomic characteristics.} Critically, unlike in the STAR sample, we also observe an indicator for graduating from a NYC public high school by 2016, which we view as the primary outcome of interest.\footnote{We can only observe whether students graduated from a high school in the New York City public school district. 27% of students in our sample leave the NYC school district before the end of high school; we include these students in our analysis and code them as not graduating from an NYC high school. The estimated effects of class size reduction on graduation rates are larger in schools where fewer students leave the district, suggesting that this missing data issue leads us to understate the overall impact of class size reduction on graduating from any high school.} We also observe information on students' race and ethnicity, sex, and (after the 1999 school year) eligibility for free or reduced-price lunch. For further information on the New York City data, see chetty2014onemeasuring and mariano2024effects.
Summary Statistics. Appendix Table (ref) presents summary statistics for the two samples. Although they are from different time periods and geographic settings, the two samples overlap on key student characteristics. Both districts serve primarily low-income students, with 61% of students in the STAR sample and 81% of the students in the NYC sample eligible for free or reduced-price lunches. Approximately one-third of the students are Black in both datasets, while the New York City sample has a significantly larger share of Hispanic students than the STAR sample. On average, there are 7.0 fewer students in third grade classrooms defined as “small” in the New York City data and 6.7 fewer students in the classrooms of students assigned to small classes in the STAR sample. 51% of students in New York City public schools graduated from high school, consistent with official statistics nysed_downloads.
We begin with OLS regressions of 3rd grade test scores on an indicator for being assigned to a small class in the STAR and NYC datasets.\footnote{Throughout our analysis of the STAR data, we use initial assignment to small class (rather than actual realized class size) as the independent variable, thereby reporting intent-to-treat estimates. Compliance with treatment assignment was imperfect because principals had to re-balance classes on other dimensions, such as gender composition. The ITT estimates provide the appropriate scaling for comparison to the observational NYC sample because the difference in average class size between those initially assigned to small vs. large classes in STAR of 6.7 students is comparable to that in the New York City data.} We include school fixed effects in all regressions run in the STAR sample because randomization was conducted within schools among children who entered in a given birth cohort. We analogously include school and birth cohort fixed effects in the NYC sample to isolate within-school and cohort variation in class size.
Table 1 (in the introduction) reports estimates from these regressions. In the STAR sample (Column 1), small class assignment in third grade increases end-of-third-grade test scores by 0.19 SD (se= 0.04).\footnote{Students who entered STAR schools before 3rd grade and were assigned to small classes in 3rd grade were assigned to small classes in earlier grades as well. We find that assignment to a small class has similar effects of end-of-3rd-grade test scores for those who entered STAR schools in 3rd grade (and thus were treated for only one year) as for those who entered in earlier grades. This is a consequence of the rapid fade-out of treatment effects on subsequent test scores documented in Figure 2. We therefore interpret the treatment effect on test scores as the causal effect of being assigned to a small class in third grade when we construct an analogous observational estimate in the NYC sample.} In the NYC sample, the corresponding OLS estimate is -0.12 SD (se = 0.01). The difference between these estimates implies that the observational estimates are confounded under our maintained external validity assumption (Assumption (ref)).
Next, we implement the ESC estimator by estimating equation (ref). We first calculate the difference between actual 3rd grade test scores and predicted test scores based on the student’s class size ($\alpha^{\rm S}_i= Y^{\rm S}_i - \tau^{\rm S} W_i$) in the NYC sample. We then replicate the OLS specification in Column 2, additionally controlling for the residuals $\alpha^{\rm S}_i$. The ESC correction yields an estimated treatment effect of class size on third grade test scores in the NYC sample of 0.19 SD (se = 0.04), which coincides with the experimental estimate in the STAR sample by construction (Column 3 of Table 1).\footnote{We estimate standard errors for this and all other ESC estimates reported below using a bootstrap procedure, where we resample the student‐level observations with replacement $B=1{,}000$ times, re‐estimate both the first‐stage class‐size effect $\hat\tau^{\rm S}$ (to form residuals $\alpha_i^{\rm S}$) and the second‐stage ESC OLS controlling for $\alpha_i^{\rm S}$ in each replication, and compute standard errors as the empirical standard deviation of the resulting bootstrap estimates.}
We next use the ESC estimator to estimate treatment effects on subsequent outcomes. Figure (ref) plots treatment effects on test scores in grades 3-8. The STAR experimental estimates (shown in black squares) are positive in all grades, while the OLS estimates in the NYC data (orange triangles) are all negative. The ESC estimates in grades 4-8 are all positive and very similar in magnitude and temporal pattern to the STAR estimates. Notably, the ESC estimates capture the well-known “fadeout” pattern in the STAR estimates – where the effects of interventions in early grades on test scores diminish in later grades. The close correspondence between the ESC estimates and the experimental estimates for the holdout outcomes ($Y^H_i$) of test scores in grades 4-8 supports the latent unconfoundedness assumption and demonstrates the ability of our approach to adjust for selection.
Finally, in the right panel of Figure (ref), we turn to the primary outcome of interest – high school graduation – which we observe in the observational but not experimental sample. The ESC estimator implies that assignment to a small third grade class increases the probability of graduating from a NYC high school by 0.69 percentage points (Column 3 of Table 1). Small classes have 7 fewer students on average relative to a sample mean of 28 students; hence, a 25% reduction in class size in third grade increases high school graduation rates by 0.69 pp.
In contrast, the OLS estimator yields a negative association between small class assignment and high school graduation rates in the observational sample. Importantly, adjusting for selection on observables by controlling flexibly for key demographic covariates in the observational sample – the interaction of indicators for gender, race and ethnicity, eligibility for free or reduced-price lunch, and cohort – has little impact on the OLS estimates on test scores and graduation rates (Figure (ref), Appendix Table (ref)). This result demonstrates that the experimental selection correction can adjust for selection on dimensions that are typically unobserved in standard administrative datasets.
Robustness. We find very similar estimated impacts on high school graduation rates when focusing on specific demographic groups ({\it e.g.}, by race or sex), as shown in Figure (ref). These findings allay the concern that differences in the demographic distribution between the STAR and NYC samples may lead to violations of the external validity assumption. We also find that the estimated impacts on high school graduation remain similar when we correct for selection using all test scores from grades 3-8 instead of just 3rd grade scores (Figure (ref)). These findings are consistent with the finding that treatment effects on test scores in grades 4-8 are very similar in the NYC and STAR data once we adjust for selection using 3rd grade test scores (Figure (ref)).
Comparison to Surrogate Estimates. In Figure (ref), we compare the ESC estimates to estimates from a surrogate index approach athey2019surrogate. To construct the surrogate-based estimates, we multiply treatment effects on third grade test scores in the STAR sample by coefficients from OLS regressions of the outcome of interest on third grade scores in the NYC sample (including school and cohort fixed effects). The surrogate estimates on test scores in grades 4-8 are all higher than the experimental estimates (though not significantly so), while the ESC estimates match the experimental estimates more closely. This pattern is consistent with the education literature on fadeout, which finds that treatment effects of early interventions on test scores persist less than one would expect based on the serial correlation of test scores across grades, violating the assumption that test scores in earlier grades provide surrogates for later outcomes. Accordingly, the ESC estimate yields an estimated treatment effect on high school graduation that is about half as large as the surrogate-based estimate.
This paper has proposed a new method of combining experimental and observational data to improve causal inference about a primary outcome of interest. We leverage the internal validity of the experimental data to implement a selection correction based on secondary outcomes when estimating treatment effects on the primary outcome in the observational dataset. Our estimator relies on a key new assumption that we term latent unconfoundedness, which requires that the unobserved confounders that affect the primary and secondary outcomes are the same. Our approach strictly weakens the assumptions underlying the popular approach of using intermediate outcomes as surrogates, yielding more credible estimates of causal effects when both the treatment and primary outcome are observed in the observational dataset.
We apply our Experimental Selection Correction estimator to estimate the effects of 3\textsuperscript{rd} grade class size on test scores in later grades and high school graduation rates. In observational data, we find wrong-signed estimates that are likely biased by unobserved selection. The ESC estimator yields estimates that match holdout experimental estimates for test scores in later grades and provides one of the first estimates of the causal effects of class sizes on high school graduation rates in the U.S. -- showing that a 25% reduction in class size in third grade increases high school graduation rates by 0.7 pp.
As observational data become more widely available, it would be valuable to build on the ideas proposed here by developing approaches to using experiments to correct for selection in observational data. Recent econometric advances have extended the framework we propose here in both identification and estimation—e.g., meza2024,park2024bracketing, park2024informativeness,imbens2025long—but there remains substantial scope for further development. In particular, future work could seek to characterize and weaken the latent unconfoundedness condition, potentially by leveraging multiple intermediate variables or partial identification strategies.\footnote{For example, the latent unconfoundedness assumption we use requires that all variation in students’ test scores arises from unobservables that affect graduation rates as well, ruling out shocks that may affect test performance but not later outcomes such as illness or noise on the day of the test. In practice, test-retest reliability tends to be very high (exceeding 0.8), so such noise is likely minimal in our application. But in other settings, noise in the intermediate outcome may be more substantial; in such cases, it may be possible to use instrumental variables approaches to adjust for such noise.} Empirically, it would be useful to characterize settings where latent unconfoundedness is a good approximation using validation studies in order to guide future applications.
\parindent 0cm\setcounter{equation}{0} {A.\arabic{equation}} \setcounter{lemma}{0}{A.\arabic{lemma}} \setcounter{theorem}{0}{A.\arabic{theorem}} \setcounter{table}{0} \setcounter{figure}{0}