Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,274 characters · 23 sections · 123 citation commands
A Bracketing Relationship for Long-Term Policy Evaluation with Combined Experimental and Observational Data
\allowdisplaybreaks
In devising, implementing, and evaluating a policy, it is essential to precisely estimate its long-term impacts. Experimental data free of confounding have been successful to this end card1992does,hotz2005predicting. However, collecting long-term experimental data is costly and time-consuming.
There has been work to overcome this challenge by combining short-term experimental data and long-term observational data. This literature has heavily relied on the so-called surrogacy condition prentice1989surrogate, requiring the long-term outcome to be conditionally independent of the treatment given a surrogate. However, subsequent research has questioned the validity of this assumption, leading to discussions referred to as the surrogate paradox -- see vanderweele2013surrogate for a comprehensive exposition. While several approaches have been proposed to mitigate the surrogacy problem athey2019surrogate,ma2021individual, a particularly promising breakthrough was made by athey2020combining with their novel latent unconfoundedness (LU) condition, replacing the conventional surrogacy condition, to correct for endogenous treatment selection in observational data.
Together with the standard internal and external validity assumptions for experiments, the LU condition proposed by athey2020combining allows one to identify long-term treatment effects with combined experimental and observational data. More recently, ghassami2022combining propose an alternative assumption to correct for selection, called the equi-confounding bias (ECB) condition.\footnote{ghassami2022combining propose several alternative approaches including the ECB. The others (e.g., Bespoke instrumental variables and proximal data fusion) require additional information besides the baseline setup, and hence we will not compare them in this paper. See also imbens2022long who use additional data (multiple short-term outcomes with a sequential structure) to invoke new assumptions that tackle persistent confounding.} These two alternative assumptions are non-nested and lead to distinct identifying formulae in general. As such, assuming the ECB condition leads to biased estimates when the LU condition is true, and vice versa.
In this light, we investigate a bracketing relationship angrist2009mostly between the two distinct identifying formulae implied by the LU and ECB conditions. Namely, we show that the LU-based estimand is less than or equal to the ECB-based estimand. Thus, if either the LU condition or the ECB condition is true, then we enjoy doubly-robust bounds for the true long-term treatment effect, bounded below by the LU-based estimand and bounded above by the ECB-based estimand. This result implies that both the LU and ECB approaches are useful, and should not be treated as mutually exclusive alternatives in practice. If a researcher seeks a point estimand and puts a higher priority on Type I Error than Type II Error, then the LU-based estimand is conservatively more robust in the sense that it underestimates the true long-term treatment effect.
Through an empirical analysis using Project Star and extensive administrative data, we demonstrate that our derived bracketing relation indeed holds. Furthermore, an analysis a la lalonde1986evaluating suggests that the LU-based estimates accurately reflect experimental estimates in educational policy evaluation. On the other hand, the ECB-based estimates are fairly distant from the experimental estimates. This suggests the potential success of the LU-based and the potential failure of the ECB-based estimates in our data setting. However, we take a step further to provide qualitative reasoning for these empirical findings, further harnessing the information value of the Lalonde-style exercise. We provide economically interpretable necessary and sufficient conditions for the two assumptions through the lens of a nonparametric family of selection mechanisms. These conditions are expressed as general restrictions, rather than being tied to any particular parametric model of selection. We find that the LU condition and the ECB condition are essentially equivalent to an invertibility condition and a martingale condition, respectively. We thus connect the success of LU condition to the literature on childhood educational interventions chetty2014measuring1,chetty2011does documenting the student's past test score as a sufficient statistic chetty2009sufficient for latent unobservables. Crucially, in our new setting under data combination, it is the latent potential test score (as in the LU condition), rather than the observed test score (as in the conventional surrogacy condition), that plays the role analogous to a sufficient statistic. Likewise, we connect the failure of the ECB condition to the dubious assumption of student test scores evolving as a martingale process. Finally, using sensitivity analysis techniques, we embed these economic assumptions into our empirical application. Anchoring the hold-out experimental estimate as the baseline truth, we assess how much the qualitative assumptions need to be violated to recover the experimental estimate in our empirical application. We find that the data-generating process is more consistent with a sub-martingale process as opposed to a martingale process, explaining the failure of the ECB-based estimates to replicate the experiment in our context.
We highlight the significance of establishing the bracketing relation between the LU- and ECB-based estimands. First, to our knowledge, the LU and ECB conditions are the sole assumptions to point-identify long-term average treatment effects\footnote{There is another condition, called the quantile-quantile ECB condition ghassami2022combining It is useful for identifying distributional and quantile treatment effects, but it will not recover the average effects unless additional assumptions, such as the rank invariance, are imposed.} under our setting without additional data (e.g., sequential outcomes, instrumental variables, or proxy variables). Second, there are practical implications, since we anticipate the LU condition, as the pioneering assumption for this data setting, and the ECB condition, due to its reliance on the parallel trend assumptions, to be widely adopted in causal analysis in a combined data setting in the future. Third, despite the growing theoretical interest in combining experimental and observational data, empirical applications in economics remain limited. Demonstrating how informative the bracketing is in our application, while remaining agnostic on the choice of the estimand, may be a promising step to bridge this gap.
In the broader literature, the LU condition encompasses the lagged dependent variable (LDV) model and the ECB condition is analogous to the parallel trend (PT) condition in the context of the conventional policy evaluation without data combination -- see angrist2009mostly. In that classical context without data combination, there is also a well-known bracketing relation that the LDV regression weakly underestimates the true non-negative treatment effect under the PT assumption while the fixed-effect (or DID) regression weakly overestimates it under the LDV assumption angrist2009mostly -- also see ding2019bracketing for nonparametric bracketing results. Our contribution to this literature is to provide an analogous bracketing relation for the novel setting of combined experimental and observation data for long-term policy evaluation involving more complicated formulae.
This section introduces the setup following the literature athey2020combining,ghassami2022combining on long-term policy evaluation with combined experimental and observational data.
A researcher conducts a randomized experiment aimed at assessing the effects of a policy. For each individual indexed by $i$ in the experiment (indicated by $G_i=E$), the researcher observes pre-treatment covariates $X_i$, a binary indicator $W_i$ of treatment assignment, and a short-term post-treatment outcome $Y_{1i}$. The researcher is interested in the effect of the treatment $W_i$ on a long-term post-treatment outcome $Y_{2i}$ that is not measured in the experimental data (i.e., $G_i=E$). To complement this deficiency, the researcher obtains an auxiliary observational data set containing measurements for a separate population of individuals $i$ (indicated by $G_i=O$), which consists of the identical list of covariates $X_i$, treatment $W_i$, and short-term outcome $Y_{1i}$, as well as the long-term outcome $Y_{2i}$ that was missing in the experimental data. The following table summarizes the observability of each of the random variables/vectors in this setup.
The symbol `$\bigcirc$' indicates that the variable is observed by the researcher. The symbol `$\times$' indicates that the variable is not observed by the researcher.
Suppose that the latent random vector $(Y_{1i}(0),Y_{1i}(1),Y_{2i}(0),Y_{2i}(1),X_i',W_i,G_i)'$ is randomly drawn from a population distribution, where $Y_{1i}(w)$ (respectively, $Y_{2i}(w)$) denotes the short-term (respectively, long-term) potential outcome under treatment $w \in \{0,1\}$. The observed outcomes are constructed according to $Y_{1i} = (1-W_i)Y_{1i}(0) + W_iY_{1i}(1)$ for all $i$ and $Y_{2i} = (1-W_i)Y_{2i}(0) + W_iY_{2i}(1)$ for all $i$ such that $G_i=O$. With these notations, a researcher is interested in identifying some conditional averages ${\text{E}}[Y_{2i}(1)-Y_{2i}(0)|\mathcal{I}]$ of the long-term treatment effect $Y_{2i}(1)-Y_{2i}(0)$, given some information set $\mathcal{I}$ describing a subpopulation of interest. Throughout, we maintain the usual potential outcome assumptions such as SUTVA imbens2015causal.
We now introduce the assumptions commonly invoked in the literature of long-term policy evaluation based on data combination. For simplicity of notations, the subscript $i$ will be omitted hereafter except when it becomes necessary.
This assumption is plausibly satisfied by construction in cases where the experimental data are collected from a randomized control trial (RCT). The next assumption, on the other hand, is less commonly imposed and might be questionable in certain applications.
This external validity assumption requires that the distributions of the potential outcomes be identical between the experimental data $(G=E)$ and the observational data $(G=O)$, given pre-treatment covariates $X$.
Throughout the paper, we assume the overlap condition that ${\text{P}}(W=1|X,G=O) \in (0,1)$ holds almost surely with respect to the law of $X$ given $G=O$.
Section (ref) introduced the basic assumptions concerning the experimental design that are under the control of the researcher. This section, on the other hand, introduces the key assumptions concerning observational data that are generally not under the control of the researcher.
There are two alternative assumptions for point identification of the long-term average treatment effects in our data setting without additional data requirements -- see Footnotes (ref) and (ref). They are the latent unconfoundedness (LU) condition athey2020combining and the equi-confounding bias (ECB) condition ghassami2022combining, as formally stated below.
Observe that Assumption (ref) encompasses the lagged dependent variable (LDV) model as a special case, while Assumption (ref) is essentially the same as the parallel trend condition -- see angrist2009mostly for discussions of these two alternative frameworks in the context of program evaluation without data combination. To fix ideas, one can consider, for instance, a simple LDV model:
Assumption (ref) holds for $w=0$ if $e \perp \!\!\! \perp W |Y_1(0), X, G=O$. On the other hand, Assumption (ref) holds if $\rho=1$ and $e \perp \!\!\! \perp W|X,G=O$. The former is stronger in one direction, while the latter is stronger in another. Namely, the former requires conditioning on the LDV for the unconfoundedness, whereas the latter requires $\rho=1$. Through this illustration, we see that Assumptions (ref) and (ref) are not nested by each other.
Under Assumption (ref), athey2020combining nonparametrically identify the long-term average treatment effect ${\text{E}}[Y_2(1)-Y_2(0)|G=O]$. More recently, ghassami2022combining investigate various long-term treatment effect estimands under each of Assumptions (ref) and (ref). Since our objective in this paper is a deeper understanding of the relationships between the alternative assumptions and identifying formulae, instead of exploring a list of various estimands, we focus on a simple estimand, namely the long-term average treatment effect on the treated, defined by $ {\theta_{\text{ATT}}} = {\text{E}}[Y_2(1)-Y_2(0)|W=1,G=O]. $
Under Assumptions (ref), (ref), and (ref), ${\theta_{\text{ATT}}}$ is identified by
athey2020combining focus on the average treatment effect (ATE), but their proof strategy directly applies to the identification of ${\theta_{\text{ATT}}}$ by (ref) as well. On the other hand, Under Assumptions (ref), (ref), and (ref), ${\theta_{\text{ATT}}}$ is identified by
See Theorem 5 of ghassami2022combining.
An important point to note here is that the estimands, (ref) and (ref), are different from each other. If Assumption (ref) holds but Assumption (ref) does not, then ${\theta_{\text{ATT}}^{\text{LU}}}$ correctly identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{ECB}}}$ fails to identify ${\theta_{\text{ATT}}}$ in general. In contrast, if Assumption (ref) holds but Assumption (ref) does not, then ${\theta_{\text{ATT}}^{\text{ECB}}}$ correctly identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{LU}}}$ fails to identify ${\theta_{\text{ATT}}}$ in general. Since Assumptions (ref) and (ref) do not nest each other as argued above, there does not seem to exist a dominant strategy for a researcher as to which assumption and identifying formula are more robust to employ in practice.
This observation raises a couple of questions. First, given that a researcher may not know which of the alternative assumptions is more plausible for a given application, is there a way to bound the true ${\theta_{\text{ATT}}}$ by utilizing (ref) or (ref)? We are going to address this question theoretically in Section (ref) and empirically in Section (ref). Second, when a point estimate is required, how can we systematically accumulate credible and generalizable insights, guided by the exogenous variation of experiments with past empirical economic evidence, to make the right choice between (ref) and (ref) in a given application? This question will be addressed via lalonde1986evaluating-style exercises in Section (ref), and through the lens of a nonparametric selection model and sensitivity analysis in Sections (ref) and (ref).
As emphasized in Section (ref), neither of Assumptions (ref) and (ref) nests the other. These assumptions are not empirically testable. They are not under the researcher's control either. The researcher needs to make a decision where there is no dominantly more robust choice.
If the latent unconfoundedness (LU) condition (Assumption (ref)) does not hold, then the estimated ${\theta_{\text{ATT}}^{\text{LU}}}$ in (ref) fails to identify the true ${\theta_{\text{ATT}}}$. If the equi-confounding bias (ECB) condition (Assumption (ref)) does not hold, then the estimated ${\theta_{\text{ATT}}^{\text{ECB}}}$ in (ref) fails to identify the true ${\theta_{\text{ATT}}}$. Hence, in the absence of knowledge about the true data-generating process, there is some risk of bias associated with each of the two estimands. Therefore, we establish a bracketing relationship between the two estimands, ${\theta_{\text{ATT}}^{\text{LU}}}$ and ${\theta_{\text{ATT}}^{\text{ECB}}}$, analogous to angrist2009mostly and ding2019bracketing. While angrist2009mostly or ding2019bracketing do not consider data combination, our result is tailored to the new data combination setting introduced by athey2020combining.
We suppress pre-treatment covariates $X$ for clarity of exposition, as they do not play a role in the identification. Let us introduce the short-hand auxiliary notation $$ \Psi(y) = {\text{E}}[Y_2(0) - Y_1(0) | Y_1(0)=y,W=0,G=O]. $$ With this notation, we state two assumptions.
This assumption requires that the series $\{Y_t(0)\}_t$ of the potential outcomes be non-explosive. As one particular example, we can interpret Assumption (ref) in the context of the aforementioned lagged dependent variable (LDV) model angrist2009mostly. Specifically, consider the LDV model: $$ Y_2(0) = \alpha + \rho Y_1(0) + e, \quad {\text{E}}[e|Y_1(0),W=0,G=O]=0. $$ Since $\Psi(y) = \alpha - (1-\rho) y$ in this example, Assumption (ref) holds under the non-explosive condition $\rho \leq 1$, i.e., the weak sub-unit root condition. angrist2009mostly also make this assumption in deriving their bracketing result.
Assumption (ref) is empirically testable, as $F_{Y_1(0)|W=0,G=O}=F_{Y_1|W=0,G=O}$ is true and $F_{Y_1(0)|G=E}=F_{Y_1(0)|W=0,G=E}=F_{Y_1|W=0,G=E}$ is also true by the internal validity condition (Assumption (ref)). While it is testable from empirical data, we remark that condition (i) is more plausible in general. Condition (i) requires that the distribution of the short-run potential outcome $Y_1(0)$ with no treatment among those who voluntarily chose not to be treated in the observational data (i.e., $W=0$ and $G=O$) first-order stochastically dominates that among the average individual in the experimental group ($G=E$). This is a reasonable assumption given that rational agents who opt out from treatment may well tend to have higher potential outcomes $Y_1(0)$ with no treatment. angrist2009mostly also make a similar assumption in deriving their bracketing result.\footnote{In our notations, angrist2009mostly assume a non-positive covariance between $W_i$ and $Y_1(0)$. In other words, untreated individuals tend to have higher $Y_1(0)$ than treated individuals. Note that $Y_1=Y_1(0)$ is true for all individuals in their data setting, which is distinct from ours.} With this said, we once again emphasize the empirical testability of Assumption (ref) (i) and (ii). Indeed, we check that Assumption (ref) (i) is satisfied in our real data analyses to be presented in Section (ref).
Under these two restrictions, we obtain the following bracketing relationship.
A proof is relegated to Appendix (ref).
Note that we invoke the experimental internal validity condition (Assumption (ref)) in this theorem, in addition to Assumptions (ref) and (ref). As remarked in Section (ref), Assumption (ref) is plausibly satisfied by construction in cases where the experimental data are collected from a randomized control trial (RCT).
{\bf Implications of Theorem (ref):} As remarked below the statement of Assumption (ref), the conditions (i) and (ii) are empirically testable. That is, the researcher knows the direction of the inequality in the bracketing relationship stated in Theorem (ref). Suppose that condition (i) is true,\footnote{If condition (ii) is true, the subsequent discussions apply with the directions of the inequalities reversed.} as it is more plausible in general and is also the case with our real-data analysis to be presented in Section (ref). In this case, we have ${\theta_{\text{ATT}}^{\text{LU}}} \leq {\theta_{\text{ATT}}^{\text{ECB}}}$. If the LU condition (Assumption (ref)) is true, then ${\theta_{\text{ATT}}^{\text{LU}}} = {\theta_{\text{ATT}}} \leq {\theta_{\text{ATT}}^{\text{ECB}}}$ follows. If the ECB condition (Assumption (ref)) is true, on the other hand, then ${\theta_{\text{ATT}}^{\text{LU}}} \leq {\theta_{\text{ATT}}} = {\theta_{\text{ATT}}^{\text{ECB}}}$ follows. Hence, if either one of the alternative conditions is true, then we enjoy the doubly-robust bound
In other words, ${\theta_{\text{ATT}}^{\text{LU}}}$ weakly underestimates the true ${\theta_{\text{ATT}}}$, while ${\theta_{\text{ATT}}^{\text{ECB}}}$ weakly overestimates the true ${\theta_{\text{ATT}}}$ regardless of which of the two alternative conditions holds. This doubly-robust bracketing relation implies that both the LU-based and ECB-based approaches are useful, and should not be considered as mutually exclusive alternatives in practice. Nonetheless, researchers often want a point estimand rather than bounds. Since underestimation is a more conservative direction of bias than overestimation, in the sense that scientists normally put a higher priority on Type I Error than Type II Error, it is more robust to assume the LU condition (Assumption (ref)) and use ${\theta_{\text{ATT}}^{\text{LU}}}$ when non-negative treatment effects are expected.\footnote{Given the conventional bracketing relation angrist2009mostly, empirical researchers often take a stance that the LDV is conservatively more robust than the DID. See crozet2017should, glynn2017front, and marsh2022trauma for instance. Since the LDV-based (respectively, DID-based) estimand serves as the lower (respectively, upper) bound in their context, our statement made in the main text reflects the viewpoints of these empirical researchers.}
{\bf Contributions of Theorem (ref) to the Literature:} This bracketing result is analogous to that for the policy evaluation without data combination extensively discussed by angrist2009mostly. Namely, they show that the fixed-effect (or DID) estimator weakly overestimates the true treatment effect when the LDV assumption is true; and the LDV estimator weakly underestimates the true treatment effect when the parallel trend assumption is true. While angrist2009mostly present this relationship for linear models, a more recent paper by ding2019bracketing shows this relationship in the context of nonparametric models. Our Theorem (ref), therefore, can be considered as a counterpart of ding2019bracketing, where we focus on the novel framework of policy evaluations with combined experimental and observational data. Since this novel setting yields more complicated identifying formulae athey2020combining,ghassami2022combining than the conventional setting without data combination, the technical value as well as the substantive value added by our new bracketing relation is non-trivial relative to the conventional bracketing relationship.
This section demonstrates the bracketing relation (Theorem (ref)) using real data consisting of combined experimental and observational data. We use the same data sets as athey2020combining to this end but provide a brief description of them for completeness in Section (ref) below.
Project STAR (Student/Teacher Achievement Ratio) was a landmark educational experiment conducted in Tennessee from 1985 to 1989. The primary objective of the experiment was to understand the impact of class size on student achievement, particularly focusing on lower-income schools.
Project STAR was motivated by the need to explore whether smaller class sizes could improve educational outcomes in lower socioeconomic settings. In the initial year (1985--86), 6,323 kindergarten students across 79 schools were randomly assigned to small classes (13-17 students) or regular-sized classes (20-25 students). This random assignment was intended to persist through the third grade. The study faced attrition as students moved or were held back in grades. Moreover, students joining in grades 1--3 were also randomly assigned to classes, making the school-by-entry-grade the primary randomization pool. Both students and teachers were randomly assigned to classes. The study administered the Stanford Achievement Test annually to assess math and reading performance, as the state tests did not extend to early grades.
Our observational data are derived from the administrative records of a large urban district. This data set includes information on approximately two million children in grades 3--8, encompassing those born between 1966 and 2001.
The data set comprises around 15 million test scores in English language, arts, and math. Due to changes in the testing regime over the 20 years, including a shift from district-specific to statewide tests and variations in test timing, we have normalized test scores to have a mean of zero and a standard deviation of one by year and grade in line with prior research practices staiger2010searching. This normalization allows for comparability with other samples nationwide.
In our analysis, we focus on the following variables common between the experimental and observational data. First, the long-term outcome $Y_2$ represents students' test scores in the long term, specifically standardized test scores (averaging between mathematics and English) at the eighth grade. Second, the short-term outcome $Y_1$ represents students' test scores in the short term, specifically, the third-grade analog of $Y_2$. Third, $W$ is a binary indicator of treatment in the form of an assignment to a small class size. Besides, we use gender, race, and eligibility for free lunch to define subpopulations, where the internal and external validities are assumed within each subpopulation.
We are now going to examine our proposed bracketing relationship, as stated in Theorem (ref). While Assumption (ref) (i) is plausible as argued in Section (ref), we can empirically check if the stochastic dominance condition is actually satisfied for an application of interest. Figure (ref) illustrates the graphs of the empirical cumulative distribution functions for $F_{Y_1(0)|G=E}=F_{Y_1(0)|W=0,G=E}=F_{Y_1|W=0,G=E}$ (solid) and $F_{Y_1(0)|W=0,G=O}=F_{Y_1|W=0,G=O}$ (dashed). Within the target subpopulation of students with lower socio-economic status proxied by eligibility for free lunch, each of the four panels (1)--(4) focuses on a subpopulation characterized by gender and race.\footnote{We show these multiple results, instead of a single aggregated result, to showcase multiple numbers of empirical evidence as opposed to just one.}
Observe that $F_{Y_1(0)|W=0,G=O}(y) \leq F_{Y_1(0)|G=E}(y)$ is likely true for each of the four subpopulations, and Assumption (ref) (i) is plausibly satisfied. Thus, our main result, Theorem (ref), implies that the bracketing relation ${\theta_{\text{ATT}}^{\text{LU}}} \leq {\theta_{\text{ATT}}^{\text{ECB}}}$ should hold. To see if this theoretical prediction is true, we now compute estimates of ${\theta_{\text{ATT}}^{\text{LU}}}$ and ${\theta_{\text{ATT}}^{\text{ECB}}}$.
As we compute ${\theta_{\text{ATT}}}$ conditional on each of the subpopulations (1)--(4) characterized by gender and race, we continue to suppress $X$ in our notations. Thus, (ref) and (ref) can be estimated by
and
respectively. We can straightforwardly compute $\widehat{\text{E}}[Y_t|W=w,G=g]$ by the sample conditional mean of $Y_t$ given $W=w$ and $G=g$ for each $t \in \{1,2\}$, $w \in \{0,1\}$, and $g \in \{E,O\}$. Similarly, we can also straightforwardly compute $\widehat{\text{P}}(W=w|G=g)$ by the sample conditional probability of $W=w$ given $G=g$ for each $w \in \{0,1\}$ and $g \in \{E,O\}$. Finally, we compute $\widehat{\text{E}}[Y_2|Y_1,W=0,G=O]$ by the linear regression of $Y_2$ on $Y_1$ using the subsample of $G=O$ with $W=0$. Plugging in yields the estimates $\widehat{\theta_{\text{ATT}}^{\text{LU}}}$ and $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}$ reported in the middle two columns of Table (ref).
Observe that both of the estimates $\widehat{\theta_{\text{ATT}}^{\text{LU}}}$ and $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}$ are positive for each of the subpopulations (1)--(4). Furthermore, we have $\widehat{\theta_{\text{ATT}}^{\text{LU}}} \leq \widehat{\theta_{\text{ATT}}^{\text{ECB}}}$ for each (1)--(4). We thus obtained an empirical confirmation of Theorem (ref) with the direction of inequality consistent with the direction of the stochastic dominance (i.e., Assumption (ref) (i)) illustrated in Figure (ref).
Moreover, the bracketing relation is informative about the sign and magnitude of causal effects. To facilitate this discussion, consider the na\"ive observational estimates $$ \widehat{\theta^{\text{na\"ive}}} = \widehat{\text{E}}[Y_2|W=1,G=O] - \widehat{\text{E}}[Y_2|W=0,G=O], $$ where $\widehat{\text{E}}[Y_2|W=w,G=O]$ is the conditional sample mean of $Y_2$ given $W=w$ and $G=O$ for each $w \in \{0,1\}$, displayed in the first column of Table (ref). They are significantly negative for all of the four subpopulations. These results are to be contrasted with the significantly positive estimates by both $\widehat{\theta_{\text{ATT}}^{\text{LU}}}$ and $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}$. Since the true ${\theta_{\text{ATT}}}$ falls between ${\theta_{\text{ATT}}^{\text{LU}}}$ and ${\theta_{\text{ATT}}^{\text{ECB}}}$ as far as either one of the LU and ECB conditions is true, one key takeaway from the bracketing relation provides a credible implication on the sign of the causal effect. While the lower bound $\widehat{\theta_{\text{ATT}}^{\text{LU}}}$ ensures the sign as such, the upper bound $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}$ additionally informs that the magnitude of the treatment effect is at most a half standard deviation for row (1).
While Section (ref) successfully evidences the bracketing relation and its informativeness about the true causal effects, the large discrepancy between $\widehat{\theta_{\text{ATT}}^{\text{LU}}}$ and $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}$ could be concerning. Those results motivate us to investigate which of the two alternative assumptions, the LU condition and the ECB condition, is more plausible for this particular application.
To answer this question, we implement analyses a la lalonde1986evaluating. Since our experimental data cover the long-term outcome $Y_2$ as well as the short-term outcome $Y_1$, we can in fact compute estimates of the experimental causal effect
where the second equality follows from the experimental internal validity (Assumption (ref)). The last column of Table (ref) lists its estimate
for each of the subpopulations (1)--(4), where $\widehat{\text{E}}[Y_2|W=w,G=E]$ is the conditional sample mean of $Y_2$ given $W=w$ and $G=E$ for each $w \in \{0,1\}$.
Observe that $\widehat{\theta_{\text{ATT}}^{\text{E}}}$ is much closer to $\widehat{\theta_{\text{ATT}}^{\text{LU}}}$ than to $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}$. To make this observation formal, we conduct two-sided tests of the hypotheses $H_0: {\theta_{\text{ATT}}^{\text{LU}}}={\theta_{\text{ATT}}^{\text{E}}}$ and $H_0: {\theta_{\text{ATT}}^{\text{ECB}}}={\theta_{\text{ATT}}^{\text{E}}}$. Table (ref) summarizes the test results. For each of the four subpopulations, we fail to reject the hypothesis $H_0: {\theta_{\text{ATT}}^{\text{LU}}}={\theta_{\text{ATT}}^{\text{E}}}$ but we do reject the hypothesis $H_0: {\theta_{\text{ATT}}^{\text{ECB}}}={\theta_{\text{ATT}}^{\text{E}}}$. Hence, we draw the robust conclusion that ${\theta_{\text{ATT}}^{\text{LU}}} = {\theta_{\text{ATT}}^{\text{E}}} \neq {\theta_{\text{ATT}}^{\text{ECB}}}$ holds. Recall that our bracketing theory predicts ${\theta_{\text{ATT}}^{\text{LU}}} = {\theta_{\text{ATT}}} \leq {\theta_{\text{ATT}}^{\text{ECB}}}$ under the LU condition (Assumption (ref)) but ${\theta_{\text{ATT}}^{\text{LU}}} \leq {\theta_{\text{ATT}}} = {\theta_{\text{ATT}}^{\text{ECB}}}$ under the ECB condition (Assumption (ref)).
For this particular application, therefore, the LU condition may be more plausible than the ECB condition. In addition to our bracketing relationship implying the conservatively greater robustness of ${\theta_{\text{ATT}}^{\text{LU}}}$ over ${\theta_{\text{ATT}}^{\text{ECB}}}$, this evidence encourages the use of the LU condition when a researcher has to choose one between the two non-nested alternative conditions for evaluating the effects of policies on educational intervention. We leave explorations of this lalonde1986evaluating-style exercise in other applied contexts (e.g., job training programs) as important future work.
The lalonde1986evaluating-style exercise performed in Seciton (ref) evidences ${\theta_{\text{ATT}}^{\text{LU}}}={\theta_{\text{ATT}}^{\text{E}}} \neq {\theta_{\text{ATT}}^{\text{ECB}}}$. Since we make additional assumptions (Assumptions (ref) and (ref)), this observation neither formally validates the LU condition (Assumption (ref)) nor formally refutes the ECB condition (Assumption (ref)). With this said, it suggests that the LU condition (Assumption (ref)) is perhaps more plausible than the ECB condition (Assumption (ref)) in this particular application. However, we take a step further to provide qualitative reasoning for these empirical findings by asking what economic substantives justify the statistical assumption of the LU and ECB conditions. Our main result is to provide economically interpretable necessary and sufficient conditions for the two assumptions through the lens of a nonparametric family of selection mechanisms and are not tied to any particular parametric model of selection. For concreteness, however, we begin with two particular seminal selection mechanisms, namely ashenfelter1985susing's (ashenfelter1985susing) Model in Section (ref) and roy1951some's (roy1951some) Model in Section (ref).
Mathematically, we are going to examine
which corresponds to the latent unconfoundedness (LU) condition (Assumption (ref)) for $w=0$,\footnote{We focus on $w=0$ for ease of writing, but a similar argument follows for $w=1$ as well.} and
which corresponds to the equi-confounding bias (ECB) condition (Assumption (ref)).
ashenfelter1985susing propose
as one possible form of selection mechanism, where $\beta \in [0,1]$ is a discount factor.
Under the selection model (ref), the LU condition (ref) is written as
We can see that this condition generally fails, with $Y_2(0)$ being the factor of dependence. In the special case where $\beta=0$, however, it reduces to
and it is always satisfied. Hence, the LU condition (ref) is plausible under ashenfelter1985susing's (ashenfelter1985susing) selection model when individuals are perfectly myopic.
Under the selection model (ref), the ECB condition (ref) is written as
We can see that this condition generally fails, with both $Y_1(0)$ and $Y_2(0)$ being the factors of dependence. In the special case where $\beta=0$, it reduces to
and it is satisfied if $Y_2(0)-Y_1(0) \perp \!\!\! \perp Y_1(0) | G=O$. Hence, the ECB condition (ref) is plausible under ashenfelter1985susing's (ashenfelter1985susing) selection model when $\{Y_t(0)\}_t$ follows a martingale-type condition and individuals are perfectly myopic.
roy1951some proposes the selection mechanism of the form
Under the selection model (ref), the LU condition (ref) is written as
We can see that this condition generally fails with $Y_2(0)$ being the factor of dependence. Even if we impose the myopia assumption that $f$ is constant in the second argument, we still cannot rule out the dependence between $Y_1(1)$ in the first argument of $f$ and $Y_2(0)$.
Under the selection model (ref), the ECB condition (ref) is written as
We can see that this condition generally fails. In the special case where the potential outcomes take the two-way fixed-effect model of the form $$ Y_{t}(w) = \alpha_{0} + \lambda_{0t} + \alpha_{1} \lambda_{1t} + \delta_{t}w + \epsilon_{t}, \qquad {\text{E}}[\epsilon_{t} | \delta_{1},\delta_{2},G=O]=0 \text{ for all } t \in \{1,2\} $$ with random $(\alpha_0,\alpha_1,(\delta_t)_{t=1}^2,\varepsilon_t)$ and constant time effects $(\lambda_{0t},\lambda_{1t})_{t=1}^2$, it reduces to
This reduced restriction is satisfied if $\lambda_{11}=\lambda_{12}$, that is, the interactive time effects are invariant over time.
While the previous two subsections shed some light on the economic contents of the two identifying assumptions, some obvious concerns are misspecification of the selection models (also raised by lalonde1986evaluating) and logical laxness by being only sufficient but not necessary. In the current subsection, we will explore necessary and sufficient conditions under a nonparametric family of selection mechanisms, boosting credibility and sharpness of economic characterization. Specifically, we consider a class of selection models of imperfect foresight, motivated by our finding in Section (ref) that both of the alternative identifying assumptions relate to myopia.
Reversing the order, we first discuss the ECB condition (ref) in Section (ref), as we can take advantage of some results from the existing literature. This will be followed by our discussion of the LU condition (ref) in Section (ref).
As mentioned in Section (ref), the ECB condition (Assumption (ref)) is essentially the same as the parallel trend assumption in the context of conventional policy evaluation without data combination. Hence, the work by ghanem2022selection, which studies the parallel trend condition from the perspective of an economic selection model, is useful to characterize the ECB condition (ref).
ghanem2022selection consider the structure
and define $ \mathcal{G}_{\text{IF}} = \left\{ g : g(\alpha,\varepsilon_{1},\varepsilon_{2},\nu,\eta_{1},\eta_{2}) \text{ is constant in } (\varepsilon_{2},\eta_{2}) \right\} $. This class $\mathcal{G}_{\text{IF}}$ entails imperfect foresight in that the selection indicator $W$ does not account for long-term idiosyncratic components $(\varepsilon_{2},\eta_{2})$.\footnote{ghanem2022selection also consider other variants of selection mechanisms. Essentially, they obtain an impossibility result under a general class of selection functions in that the parallel trend condition implies time-invariant relative levels of the long- and short-term potential outcomes. The class $\mathcal{G}_{\text{IF}}$ we focus on is the most general class they consider among those that produce non-trivial characterizations.} We focus on this class of selection models with imperfect foresight in light of our discovery from Section (ref) that both of the identifying assumptions, namely the LU condition (ref) and the ECB condition (ref), entail myopia.
ghanem2022selection show necessary conditions for the parallel trend condition to hold for all $g \in \mathcal{G}_{\text{IF}}$. Let $\dot Y_{t}(0) = Y_{t}(0) - {\text{E}}[Y_{t}(0)|G=O]$. In our framework with combined experimental and observational data, their result directly translates into the following proposition.
This necessary condition also becomes sufficient under an additional condition ghanem2022selection. As discussed by ghanem2022selection, it can be interpreted as the martingale condition for $\{\dot Y_{t}(0)\}_t$ given $\alpha$ (and also given $G=O$ in our framework with data combination).
Unlike the ECB condition (ref) discussed in Section (ref), the LU condition (ref) has not been studied in the context of selection models like (ref)--(ref) to the best of our knowledge. Thus, we present a characterization of the LU condition (ref) under (ref)--(ref) in the current subsection.
To use $Y_{1}(0)$ as a control variable in the current subsection, we assume the functional form $f_1(\alpha,\varepsilon_{1})=\tilde f_1(\alpha)$ as in athey2020combining after suppressing the pre-treatment covariates. For technical purposes, we further assume that $\tilde f_1$ is measurable, $\mathcal{F}_2$ consists of all measurable functions $f_2$, and $\mathcal{G}_{\text{IF}}$ consists of all measurable functions $g$ with the same restriction from Section (ref). To characterize the LU condition (ref) under (ref)--(ref), we make the following generalization of the function invertibility.
In short, this generalized concept of invertibility requires $\tilde f_1$ to be invertible except on the domain of zero probability. This invertibility in the structure (ref) can be interpreted as follows. The production factor $\alpha$ represents the student's current ability in $t=1$, which monotonically determines the potential outcome $Y_1(0)$ in the short run. In the long run, on the other hand, this short-run ability $\alpha$ together with additional future factors $\varepsilon_2$, not necessarily a scalar, determines the potential outcome $Y_2(0)$ in a way that is not necessarily monotone. This sense of invertibility characterizes the latent unconfoundedness condition (ref), as formally stated in the following proposition.
A proof is relegated to Appendix (ref).
This proposition implies that the invertibility of $\tilde f_1$ in the sense of Definition (ref) is necessary and sufficient for the LU condition (ref) to hold. In particular, it does not require restrictions on the relative levels of the long- and short-term potential outcomes unlike the characterization (Proposition (ref)) presented in the previous subsection for the equi-confounding bias condition (ref). This feature of the latent unconfoundedness condition is analogous to the fact that the change-in-changes method athey2006identification does not require a restriction on the relative levels of the potential outcomes,\footnote{Indeed, athey2020combining describe the latent unconfoundedness as a condition that relates to the change-in-changes method, as is the case with general control function approaches. We move one step further by showing that a slightly modified version is indeed necessary and sufficient for the LU condition under general conditions.} unlike the difference-in-differences methods, in the conventional policy evaluations without data combination. In the novel context of long-term policy evaluation with data combination, Proposition (ref) provides a similar but different characterization of the LU condition (ref).
Recall that the LU condition uses the short-term potential outcome $Y_1(0)$, as opposed to the observed short-term outcome $Y_1$ as in the conventional surrogacy condition, to control the confounding factor $\alpha$. As characterized by Proposition (ref), the LU condition recovers $\alpha = \tilde f_1^{-1}(Y_1(0))$ from the potential outcome $Y_1(0)$ via the invertibility of $\tilde f_1$. On the other hand, $\alpha$ is hard to recover from the observed outcome $Y_1 = \tilde f_1(\alpha)+g(\alpha,\varepsilon_1,\varepsilon_2,\nu,\eta_1,\eta_2)(Y_1(1)-Y_1(0))$ due to the presence of potentially confounding factors other than $\alpha$, even if $\tilde f_1$ were invertible. This observation through the lens of the economic selection model explains why the LU condition may well be reasonable while the conventional surrogacy condition is not.
In closing this section, we emphasize that the invertibility condition (Definition (ref)) does not restrict the unobserved factor $\alpha$ to be a scalar. If a researcher obtains multi-dimensional $Y_1(0)$ then the invertibility condition accordingly admits multidimensional $\alpha$ to characterize the LU condition (ref). Examples include parental background and children's cognitive and non-cognitive skills. This idea is analogous to the idea of multi-dimensional surrogates proposed by athey2019surrogate.
We studied the alternative identifying assumptions through the lenses of two parametric classes and one nonparametric class of selection models. The following table summarizes the observations that we have made in Sections (ref), (ref) (ref), and (ref).
The LU condition (Assumption (ref) or (ref)) is weaker than the equi-confounding bias condition (Assumption (ref) or (ref)) under ashenfelter1985susing's model of selection, while the ECB condition (Assumption (ref) or (ref)) is weaker than the latent unconfoundedness (Assumption (ref) or (ref)) under roy1951some's model of selection. Under the nonparametric class of selection models with imperfect foresight, the two alternative conditions are characterized by a distinct set of necessary (and sufficient) conditions.
Recall that the lalonde1986evaluating-style exercise in Seciton (ref) suggests that the LU condition (Assumption (ref)) is consistent with the data. In other words, the equivalent invertibility condition (Definition (ref)) may be plausible. At first glance, this condition appears strong since it implies that the one-dimensional past test score alone captures all the unobserved factors relevant to selection. However, this observation is in line with empirical findings about the informativeness of student test scores reported in the existing literature. For instance, in the context of teacher value-added analysis, chetty2014measuring1 find that accounting for students' prior test scores provides unbiased forecasts of teachers’ impacts on student achievement. Moreover, using the same Project STAR data set as in our analysis, chetty2011does show that past test scores serve as excellent controls to predict children's earnings in adulthood. In this sense, our findings reconfirm those in the empirical economics literature. It is worth noting that the prior test score $Y_1$ per se is not sufficient as a surrogate, as noted by athey2020combining, but its latent value in the form of the potential outcome $Y_1(0)$ as in the LU condition may well be a sufficient control.
On the other hand, the lalonde1986evaluating-style exercise in Seciton (ref) suggests that the ECB condition (Assumption (ref)) is perhaps implausible. Thus, under the class of ashenfelter1985susing's model of selection and, more generally, in the nonparametric class of selection with imperfect foresight, the martingale condition is unlikely to hold. This conclusion is reasonable for the test score as it is supposed to be sub-martingale.\footnote{Note that a conditionally non-degenerate martingale process implies that the variance monotonically increases with $t$. But this contradicts the restriction that the test scores are bounded.} To embed this into our real-data application, we finally conduct a selection-based sensitivity analysis in the following section.
As discussed at the end of the previous section, the lalonde1986evaluating-style exercise in Seciton (ref) implies that the martingale condition, as the necessary (and sufficient) condition for the ECB condition (Assumption (ref)), is likely to fail. This motivates the following question. How much deviation from the martingale condition would allow the ECB-type approach to rationalize the experimental estimate $\widehat{\theta_{\text{ATT}}^{\text{E}}}$ in the lalonde1986evaluating-style exercise in Seciton (ref)? The current section investigates this question by adapting the idea of the selection-based sensitivity analysis, proposed by ghanem2022selection in the context of the conventional policy evaluation without data combination, to our novel context of the long-term policy evaluation with data combination.
Consider the following deviation of $\{\dot Y_t(0)\}_t$ from the martingale process, governed by a super-parameter $\overline\rho$.
For instance, we can think of the process specified by
In this concrete specification, $\overline\rho=1$ entails the martingale process as a special case, whereas $\overline\rho<1$ implies sub-martingale processes.
Let $\dot Y_1 = Y_1 - {\text{E}}[Y_1|W=0,G=E]$ for a short hand. The following proposition shows how the super-parameter $\overline\rho$ translates into the discrepancy between the ECB-based estimand ${\theta_{\text{ATT}}^{\text{ECB}}}$ and the true value of ${\theta_{\text{ATT}}}$.
A proof is relegated to Appendix (ref). In the concrete specification (ref) of $\phi$, the bias $\Delta(\overline\rho)$ takes the simple form
{\bf Contributions of Proposition (ref) to the Literature: } This result makes non-trivial contributions to the literature on selection-based sensitivity analysis. ghanem2022selection study the selection-based sensitivity in the context of the conventional policy evaluation without data combination. Following their framework directly, we would obtain
In their framework, $Y_1(0)=Y_1$ is observable for all units as the pre-treatment outcome under no anticipation, and hence observational data alone allow for their sensitivity analysis. On the other hand, in the novel context of long-term policy evaluation with data combination as in athey2020combining, even the short-run potential outcome $Y_1(0)$ is observed only for the untreated observations, and hence we cannot directly take advantage of the results from the existing literature. In view of our proof of Proposition (ref), the reader can see that the internal and external validity conditions (Assumptions (ref) and (ref)) together elegantly solve this unobservability problem.
Proposition (ref) motivates the $\overline\rho$-adjusted ECB-based estimate $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}(\overline\rho) = \widehat{\theta_{\text{ATT}}^{\text{ECB}}} - \widehat\Delta(\overline\rho)$, where $\widehat\Delta(\overline\rho)$ is the sample-counterpart estimator of $\Delta(\overline\rho)$ by replacing the conditional mean (respectively, probability) with the sample conditional mean (respectively, probability). With this variant $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}(\overline\rho)$ of the ECB-based estimate, we will now revisit the empirical application presented in Section (ref).
The left panel of Figure (ref) illustrates Table (ref) for sample (1) in terms of box plots. Indicated around the estimates in the middle are the interquartile ranges and 95% confidence intervals based on the limit normal approximation. As we observed in Sections (ref)--(ref), the LU-based estimate $\widehat{\theta_{\text{ATT}}^{\text{LU}}}$ is close to the experimental estimate $\widehat{\theta_{\text{ATT}}^{\text{E}}}$, but the ECB-based estimate $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}$ is far above $\widehat{\theta_{\text{ATT}}^{\text{E}}}$.
Using the specification (ref) for deviations from the martingale condition, the right panel of Figure (ref) illustrates $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}(\overline\rho)$ as a function of $\overline\rho \in [0.5,1.0]$. Recall that $\overline\rho=1$ implies the martingale process, which is consistent with the ECB condition (Assumption (ref)). Thus, $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}(1)$ exactly coincides with the baseline ECB-based estimate $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}$. As $\overline\rho$ becomes smaller, $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}(\overline\rho)$ accordingly gets smaller. It is when $\overline\rho=0.58$ (indicated by the dashed vertical line in the right panel of the figure) that $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}(\overline\rho)$ coincides with the experimental estimate $\widehat{\theta_{\text{ATT}}^{\text{E}}}$. In other words, for the ECB-type estimate $\widehat{\theta_{\text{ATT}}^{\text{ECB}}}(\overline\rho)$ to coincide with the experimental estimate, the martingale condition needs to be violated, and instead be replaced by a sub-martingale condition with the specific value of $\overline\rho=0.58$ in terms of point estimates. This number is approximately between the white noise process $(\overline\rho-0)$ and the Margingale process $(\overline\rho=1)$. Thus, the ECB condition, being equivalent to the martingale condition, was perhaps incompatible with this application.
The literature on long-term policy evaluation with data combination proposes alternative identifying assumptions for nonparametric identification of long-term treatment effects. First, athey2020combining propose the latent unconfoundedness (LU) condition. Second, ghassami2022combining propose the equi-confounding bias (ECB) condition, which is analogous to the parallel trend assumption. The LU and ECB conditions are mutually non-nested and lead to distinct identifying formulae, ${\theta_{\text{ATT}}^{\text{LU}}}$ and ${\theta_{\text{ATT}}^{\text{ECB}}}$, for the true long-term treatment effect ${\theta_{\text{ATT}}}$.
In this light, we develop a bracketing relation between the LU-based estimand and the ECB-based estimand. Specifically, we show that the LU-based estimand ${\theta_{\text{ATT}}^{\text{LU}}}$ is less than or equal to the ECB-based estimand ${\theta_{\text{ATT}}^{\text{ECB}}}$. Thus, we have ${\theta_{\text{ATT}}^{\text{LU}}} = {\theta_{\text{ATT}}} \leq {\theta_{\text{ATT}}^{\text{ECB}}}$ when the LU condition (Assumption (ref)) holds true, while we have ${\theta_{\text{ATT}}^{\text{LU}}} \leq {\theta_{\text{ATT}}} = {\theta_{\text{ATT}}^{\text{ECB}}}$ when the ECB condition (Assumption (ref)) holds true. More generally, if either the LU condition (Assumption (ref)) or the ECB condition (Assumption (ref)) is true, then we enjoy the doubly-robust bound ${\theta_{\text{ATT}}^{\text{LU}}} \leq {\theta_{\text{ATT}}} \leq {\theta_{\text{ATT}}^{\text{ECB}}}$. This result implies that both the LU and ECB approaches are useful, and should not be considered as mutually exclusive alternatives in practice. If a researcher seeks a point estimand and puts a higher priority on the Type I Error than Type II Error, then the LU-based estimand ${\theta_{\text{ATT}}^{\text{LU}}}$ is conservatively more robust than the ECB-based estimand ${\theta_{\text{ATT}}^{\text{LU}}}$ in the sense that ${\theta_{\text{ATT}}^{\text{LU}}}$ underestimates ${\theta_{\text{ATT}}}$.
The existing literature angrist2009mostly,ding2019bracketing provides bracketing relations in the context of policy evaluations without data combination. Our result contributes to this literature by providing a counterpart for long-term policy evaluation with combined experimental and observational data involving more complicated identifying formulae.
Combining data from Project STAR, a seminal experiment, and extensive administrative observational data, we demonstrate that our proposed bracketing relationship is indeed satisfied. The implied bounds are informative about the sign and magnitude of the causal effects. A lalonde1986evaluating-style exercise shows that the LU condition may be more plausible than the ECB condition for this particular application of evaluating policies on educational interventions.
We take additional steps to further harness the value of the exogenous variation from the experiment. To understand the economic substantives behind these empirical results, we characterize the LU and ECB conditions from the perspectives of the economic selection model while aiming to mitigate the concern of model misspecification. Under a general nonparametric class of selection models with imperfect foresight, the LU condition is equivalent to the invertibility of the short-term outcome production function, whereas the ECB condition is equivalent to the martingale process of potential outcomes.
We thus connect the success of LU condition to the literature on childhood educational interventions chetty2014measuring1,chetty2011does documenting the student's past test score as a sufficient statistic chetty2009sufficient for latent unobservables. Crucially, in our new setting under data combination, it is the latent potential test score (as in the LU condition), rather than the observed test score (as in the conventional surrogacy condition), that plays the role analogous to a sufficient statistic. Likewise, we connect the failure of the ECB condition to the dubious assumption of student test scores evolving as a martingale process.
Finally, anchoring on the hold-out experimental estimate as the baseline truth, we conducted a selection-based sensitivity analysis to find that the martingale condition would need to be violated substantially for the ECB-based estimates to match the true causal effect.
Table (ref) summarizes the LU and ECB conditions in terms of the bracketing result, its implied robustness, scale sensitivity, the results of our lalonde1986evaluating-style exercise, equivalent characterizations, dimension requirements, and recommendations of when to use.