EconBase
← Back to paper

The Informativeness of Combined Experimental and Observational Data under Dynamic Selection

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

122,350 characters · 19 sections · 165 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dynamic confounding and long-term treatment effect estimation by data combination: point and partial identification

abstract

Introduction

In devising, implementing, and evaluating a policy, estimating the long-term impacts fast and precisely is essential. Experimental data free of confounding have been successful in answering these causal questions; in assessing the long-term effect of classroom size on children's academic performance card1992does, Project Star played a key role; for debates on the long-term impact of job training on unemployment rates, California Gain program has been widely cited hotz2005predicting. However, at the same time, collecting experimental data is both costly and time-consuming. The situation exacerbates when long-term data is necessary and sometimes even infeasible for follow-up. Can we get a reliable long-term estimate circumventing the previous challenges?

This question has received considerable interest in the past, especially in the biostatistics literature weir2006statistical,vanderweele2013surrogate. There, the literature has mainly relied on the surrogacy condition. First proposed by prentice1989surrogate, the statistical surrogate criterion requires the primary outcome to be conditionally independent of the treatment given the surrogate. However,

subsequent literature has questioned the validity of the assumption, sometimes refered to as the surrogate paradox (frangakis2002principal,freedman1992statistical,buyse1998criteria,buyse2000validation,chen2007criteria;See vanderweele2013surrogate for a comprehensive exposition). While various approaches have been taken to mitigate the problem(e.g., proposing alternative surrogate criterionsma2021individual, introducing multi-dimensional surrogatesathey2019surrogate ), a particularly promising approahc was proposed by athey2020combining, which proved that assuming that the treatment indicator is observed in the short-term experimental data obviates the need for the surrogacy assumption, replaced by a novel latent unconfounding assumption. The latent unconfounding assumption says that conditional on a potential, rather than the observed short-term outcome in the observational data, unconfounding in the observational data holds. Following athey2020combining, ghassami2022combining proposed a distinct, non-nested equi-confounding bias assumption, which is a parallel-trends type assumption that can also identify the long-term ATE in the same setup.

Residual literature

Combining datasets Augmenting inference using auxiliary information itself has a relatively long history; imbenslancaster provide additional moment restrictions from auxiliary data; imbenshellerstein combine macro data with micro data, degroot1971matchmaking used multiple datasets on the same variables for dealing with measurement error; glass1976primary takes a meta-analysis approach to combine multiple data for checking the robustness of a particular study;

Recent literature that particularly focus on combining experimental and observational data to improve causal inference, other than identification(which is mostly the surrogacy litearutre) is for efficiency gains, (kallus2020role,zhang2022high,cheng2021robust ), partial identification( ridder2007econometrics,fan2014identifying)

Our paper is closely related to contemporary work like ghassami2022combining,imbens2022long. ghassami2022combining provide alternative assumptions like equi-confounding bias assumption or proxy methods assuming the availability of additional variables that satisfy certain independence assumptions. Similarly, imbens2022long mainly focus on persistent confounding: confounding induced by variables that simultaneously affect the treatment, short and long-term outcome, like intelligence, which can lead to the violation of the latent confounding assumption. They assume the existence of additional multiple short-term outcomes satisfying particular sequential independence conditions to overcome this challenge. However, none of these papers have been able to extend this approach of athey2020combining without the existence of additional data.

others( economic content of parallel trend, bracketing etc.

frangakis2002principal pointed out that this criterion and related methods (see, e.g., freedman1992statistical,buyse1998criteria,buyse2000validation, ignore the fact that surrogates can be affected by the treatment, so their parameters of interest may not have a valid causal interpretation. This can have problematic consequences, termed the surrogate paradox(chen2007criteria); even if the surrogate positively affects the primary outcome, violation of the surrogacy condition by the surrogacy-outcome confounding may lead to adverse treatment effects on the primary outcome. While various approaches have been taken to mitigate the problem(e.g., proposing alternative surrogate criterionsma2021individual, introducing multi-dimensional surrogatesathey2019surrogate ), the literature relying on surrogate conditions has struggled to provide a definite solution.

A recent breakthrough was proposed by athey2020combining, which proved that assuming that the treatment indicator is observed in the short-term experimental data obviates the need for the surrogacy assumption, replaced by a novel latent unconfounding assumption. The latent unconfounding assumption says that conditioning on a potential, rather than the observed short-term outcome in the observational data, unconfounding in the observational data holds. Following athey2020combining, ghassami2022combining proposed a distinct, non-nested equi-confounding bias assumption, which is a parallel-trends type assumption that can also identify the long-term ATE in the same setup. These recent discoveries raise an additional question of which assumption policy-makers should choose and what type of treatment effects, beyond the ATE, are identifiable under it. In addition, some strands of literaturedeaton2018understanding,manski2013public suggest that the relatively neglected external validity, opposed to internal validity assumptions is also essential for credible causal inference.

In this paper, we take a closer look into these questions on the assumptions motivated by a desire to estimate the long-term effect of job training programs on labor market outcomes without long-term experimental data. We take a semi-structural approach to this problem armed with two tools under-utilized in the data-combination literature; economic theory and partial identification. We develop ways to interpret, relax, and test these internal and external validity assumptions. We deploy seminal selection mechanisms like ashenfelter1985susing,roy1951some that have been successful in empirical economics to interpret and choose between these assumptions. When point identification is infeasible, we turn to partial identification, providing sharp identification bounds under restrictions motivated by economic theory in increasing strength. We extend the current research on long-term causal inference with data combination from average treatment effect (ATE) to variants of those. For each identification result, we accompany them with estimation methods using refined statistical tools like high-dimensional Gaussian approximation and semiparametric efficiency theory. In many instances, we refine state-of-the-art convergence rates of those methods, increasing their theoretical validity.

To carry out the agenda in the previous paragraph, in the main body, we divide our analysis into two main parts based on the treatment effect of interest. Part I is mainly interested in the canonical long-term ATE. Within this part, we first approach the external validity assumption and see how we can relax and test the external validity assumptions in athey2020combining. We next approach the observational-internal validity assumption and mainly focus on the latent-unconfounding assumption proposed by athey2020combining and the equi-confounding assumption proposed by ghassami2022combining. We show that justifying either assumption can be translated into a more concrete problem of justifying either the Ashenfelter-and-Card(AC) selection mechanism ashenfelter1985susing or the Roy selection mechanism, roy1951some based on treatment effects. This approach translates a problem with little prior knowledge into an empirically grounded economic problem, facilitating informed decision-making. Importantly, in this process, we discover how the time-varying confounding invalidates both these assumptions, highlighting the limitations of current methods. Finally, we complete our analysis with rigorous estimation and inference. In estimating the partially identified bounds, we utilize uniform confidence bands chernozhukov2014anti and techniques accustomed for possibly many moment inequalities chernozhukov2019inference. Along the way, We provide an improved Gaussian approximation result to the suprema of empirical processes of chernozhukov2013gaussian, the theoretical basis for the previous methods.

To highlight the challenge of dynamic selection, Part II focuses on a variant of the long-term ATE, namely the Average treatment effect on the treated survivors (ATETS). We can interpret ATETS as the long-term treatment effect on those who failed to transition to employment in the short term, hence has an interpretation of a hazard ratio. We show how the latent unconfounding assumption fails to identify the ATETS, again because of dynamic confounding. Therefore, we turn to alternative identification strategies, starting from point identification, and ending in partial identification. For point identification, under the potential outcome framework, identification requires stringent assumptions such as no state dependenceheckman1981heterogeneity. Therefore, we turn to structural approaches that allow more nuanced ways of embedding the economics of forward-looking optimizing agents. The structural approach allows for state dependence in contrast to the experimental approach, providing a reasonable estimate of how time-varying heterogeneity enters into the identification problem. The structural approach is complemented by experimental data that serves as a tool for validating the structural models. Finally, we look into what can be learned about the ATETS under minimal assumptions, building a dynamic potential outcome framework implied by a semi-structural model eschewing explicit parametric assumptions. Here, we impose constraints and information in increasing strength, starting from only the observational data, then adding the short-term experimental data, and finally, imposing nonparametric shape restrictions justified by the semi-structural model. We again complete our series of identification strategies with formal inference. We derive the efficient influence function for nonparametrically identified ATETS leading to Neyman-orthogonal robust moment functions. We also provide novel nonasymptotic Gaussian inference for structural models. Importantly, this approach to structural estimation remains valid even for high-dimensional state variables or non-regular semiparametric cases manski1975maximum exhibiting cube-root asymptotics kim1990cube.

In the Appendix, we shortly discuss equilibrium effect considerations in our context. Unlike the short-term outcome(e.g., test scores in 3rd grade), the ultimate interest of policy makers are often market outcomes( e.g., wages), which biases estimates that ignore interactions. Unlike the full-blown structural model approach of classical literature in economics(e.g., heckman1998general), we are interested in extending the nonparametric potential outcome framework by relaxing the SUTVA assumptionimbens2015causal. In our motivating example, assuming that price mediates the interaction among the market participants, which increase to infinity leads to identification of the equilibrium-augmented treatment effect in the mean-field regime, extending munro2021treatment,wager2021experimenting to long-term causal inference combining data.

Related Literature

description• \\ Combining data sets itself has a relatively long history; degroot1971matchmaking used multiple datasets on the same variables for dealing with measurement error; glass1976primary takes a meta-analysis approach to combine multiple data for checking the robustness of a particular study. On the other hand, when the goal is long-term causal inference, people have commonly deployed surrogate methods. A large body of biostatistics literature has discussed this approach of long-term estimation on surrogate outcomes; see reviews in weir2006statistical,vanderweele2013surrogate,joffe2009related However, these criteria can quickly run into logic paradoxes chen2007criteria or rely on unidentifiable quantities, showing the challenge of causal inference when the primary outcome is entirely missing. In recent years, athey2019surrogate proposed using multiple surrogates to make the surrogate condition hold. Along this line, some contemporary literature also combines experimental and observational data and rely on the statistical surrogate criterion, either to estimate cumulative treatment effects in dynamic settings battocchi2021estimating or to learn long-term optimal treatment policies cai2021coda. Still, many have questioned the validity of the surrogacy conditions. In this light, the novelty of athey2020combining and its latent unconfounding assumption can be seen as a way to relax the strong surrogacy assumption. We build on this line of research. Our paper is closely related to contemporary work like ghassami2022combining,imbens2022long. ghassami2022combining provide alternative assumptions like equi-confounding bias assumption or proxy methods assuming the availability of additional variables that satisfy certain independence assumptions. Similarly, imbens2022long mainly focus on persistent confounding: confounding induced by variables that simultaneously affect the treatment, short and long-term outcome, like intelligence, which can lead to the violation of the latent confounding assumption. They assume the existence of additional multiple short-term outcomes satisfying particular sequential independence conditions to overcome this challenge. However, none of these papers have been able to extend this approach of athey2020combining without the existence of additional data. Our paper is the first to use the economics of the selection-into-treatment and the inter-temporal optimizing agents to understand the assumptions and to deal with dynamic confounding. We are also the first, to our knowledge, to extend the point-identification framework of athey2020combining to partial identification. • \\ In program evaluation in economics, there has been a historical distinction between the 'experimental approach'imbens2010better,angrist2009mostly and the 'structural approach' deaton2010instruments,deaton2018understanding,heckman2010comparing, featuring different identification and estimation strategies. The 'experimental' approach or the 'program evaluation' approach focuses on 'effects' defined by experiments or surrogates for experiments as the objects of interest and not the parameters of explicit economic models heckman2010building. On the other hand, the 'structural' approach generally consists of a fully specified behavioral model, usually, not necessarily, parametric. The structural approach is often used to evaluate existing policies and perform counterfactual program/policy experiments, such as evaluating new hypothetical policies. Many debates have revolved around whether the structural or experimental approach is preferable for policy evaluation, as illustrated in papers like lalonde1986evaluating,dehejia1999causal,dehejia2002propensity,smith2005does,deaton2010instruments,deaton2018understanding,imbens2010better. Recently, however, several strands of important research have emerged, questioning the necessity of taking a simplistic binary approach for causal inference. Instead, they propose combining these two approaches to exploit the advantages and ameliorate the disadvantages of each. We resonate with and build on this approach. We take several steps forward in this direction, both from a point and partial identification perspective. We give a brief overview of the several strands and how our work fits within each strand. Firstly, the semi-structural approach that focuses on making minimal, and theory-motivated assumptions to answer the policy question has received increasing interest in the literature. From the seminal work of marschak1974economic and recently heckman2010building, they state that usually, only a combination of economic parameters is necessary for policy making. Notable works along this line include ichimura2000direct,ichimura2002semiparametric for estimating the effect of tuition subsidies using a semiparametric reduce form approach. Relatedly, in the public economics literature, chetty2009sufficient proposes a sufficient statistic approach that enables counterfactual welfare analysis with key statistical quantities, building on harberger1964measurement. Our novel contribution to this strand first appears in Part I. There, we face the challenge of choosing between non-nested observational internal validity assumptions with little empirical implementation and comparison. To make progress, instead of building a full-blown structural model, we use classical selection mechanisms that model the decision process of entering into the program as a intermediate decision step, that provides a concrete, economically interpretable grounding for decision making. For example, we can show that the moment we assume the AC mechanism, the equi-confounding assumption is invalidated, and the question becomes whether we can justify the latent unconfounding assumption by showing, for example, \textit{myopic} decision-making of the agents. Furthermore, in the Appendix, we consider general equilibrium effect considerations and take a mean-field approximation approach that essentially requires only a market clearing condition, contrary to the full-blown structural model of, e.g., heckman1998general in estimating effects of college-tuition subsidies. Secondly, research has used unconfounded experimental data to validate parametric structural models. For example, hausman2007social building a housing demand model, todd2006assessing a female labor supply model, both to evaluate program impacts, increasing the credibility model and the counterfactual analysis by using an RCT holdout sample. We employ this in Part II, Section (ref) to evaluate a job-search model using short-term data. Finally, and closest to our current theme, the approach of athey2020combining,athey2019surrogate,kallus2020role,imbens2022long synthesizes the complementary characteristics of experimental and observational data, leading to feasible and reliable inference for long-term treatment effects. Our contribution is to show how this approach is useful even for \textit{paial identification} of long-term estimands, \textit{beyond the canonical ATE}. In providing bounds on ATETS in Part II, we show how the short-term experimental data can significantly reduce the uncertainty of worst-case bounds attained from only observational data. We go further to refine those bounds by employing nonparametric assumptions. Importantly, we only employ assumptions that are \textit{justified by the underlying semi-structural problem}, a \textit{partial identification extension} of ichimura2000direct,ichimura2002semiparametric. • \mbox\\ Checking the robustness of the conclusions to imposed assumptions is crucial, which has motivated exciting lines of research in sensitivity analysis imbens2003sensitivity,cinelli2020making,imai2010identification,imai2013identification and partial identification. We will take the latter approach in this paper. Our method echoes the Law of Decreasing Credibility initially proposed by manski2003partial that the credibility of the conclusion decreases as the strength of the assumption increases. We employ this in Part I to create a uniformly valid test based on confidence bands. A similar approach has been deployed, for example, in blundell2007changes to test the gender differential wage gap in the UK or to overcome measurement problems; see manski2002inferenceon interval-valued outcome variables and ho2014hospital on errors in regressors in discrete choice models. For the use of inequalities to overcome selection problems, see blundell2007changes on changes in inequality and kreider2012identifying on take-up of SNAP (food stamps) andrews2019inference ;for problems in non-compliance in RCTs, see lee2009training. To our knowledge, few papers use the partial identification approach for assessing external validity. While manski2013public provides a framework of 'external assessment' using bounds, unlike our work, it takes a social decision theory approach. We will see what reasonable assumptions may be placed upon the study's locational design, introduce new partial identifying assumptions, and see how they translate to sharp identification bounds. To calculate those sharp identification bounds, methods in random set theory, introduced by beresteanu2008asymptotic,molchanov2018random to the partial identification literature, will be instrumental. In particular, it helps get \textit{uniformly valid} bounds for the distribution function that may be difficult to characterize otherwise. • \mbox\\ For high-dimensional Gaussian approximation, we heavily rely on the seminal works of chernozhukov2013gaussian,chernozhukov2017central, and its extensions to suprema of empirical processes chernozhukov2014gaussian. These papers rely on Slepian-Stein methods for Normal approximation that have suboptimal rates in low-dimensional regimes, but are well-suited for the high-dimensional case with only a log order of dependence on dimension. This method will be the fundamental theory underlying various methods that we use; the uniform confidence bands for estimating the bounds on the distributional functions in Part I; the inference of Neyman-orthogonal moment functions for nonparametrically point-identified ATETS in Part2; the high-dimensional M-estimation framework for high-dimensional, semiparametric structural models in Part2. We refine the approximation by refining the coupling inequality, for example, in chernozhukov2014gaussian based on the Stein-exchangeable pair approach. There have been recent works chernozhukov2019improved,chernozhukov2020nearly in refining the Gaussian approximation methods of maxima of high-dimensional vectors, for example, using the iterative randomized Lindeberg method in chernozhukov2019improved. Our refinement can be seen as an empirical process extension of those results. For semiparametric efficiency theory, this mainly relates to our Neyman-orthogonal moment function construction in Section(ref). Adding the influence functions of the estimators has been shown to lead to second-order bias properties, as in, for example, chernozhukov2022locally. We depart from the setting in chernozhukov2022locally by considering non-linear functionals of the nuisance parameters, naturally arising in nested expectations for long-term causal estimands by data combination. Therefore, we take the classical, but mechanical path-wise-derivative calculation approach, as in bickel1993efficient to calculate the efficient influence function for ATETS.

The rest of the paper is organized as follows. In Part I, we provide the formal setup and the identification results of athey2020combining which is the starting point of this paper. Then we touch external validity, and observational internal validity issues in the subsequent sections, concluding with estimation considerations. In Part II, we introduce and motivate the ATETS, and turn to nonparametic, then structural approach to point identification, and then to the bound analysis. We finally turn to estimation. Most of the proof is relegated to the Appendix.

\part{External and internal validity assumptions for estimating long-term ATE combining data}

Notation and Setup

The notation we deploy here is as follows. Let $\lambda$ be a $\sigma$-finite probability measure on the measurable space $(\Omega, \mathcal{F})$, and let $\mathcal{M}_\lambda$ be the set of all probability measures on $(\Omega, \mathcal{F})$ that are absolutely continuous with respect to $\lambda .$ For an arbitrary random variable $D$ defined on $(\Omega, \mathcal{F}, Q)$ with $Q \in \mathcal{M}_\lambda$, we let $E_Q[D]$ denote its expected value, where we leave the measure implicit when it is clear. We let $\|\cdot\|_{P, q}$ denote the $L^q(P)$ norm and $[n]$ denote the set $\{1, \ldots, n\}$. The quantity $p(d \mid E)$ will denote the density of the random variable $D$ at $d$ under $Q$ with respect to $\lambda$ conditional on the event.

Next, consider the follwing setting. A researcher conducts a randomized experiment aimed at assessing the effects of a job training program. For each individual in the experiment(with index $G=0$), they measure a q-vector $X_i$ of pre-treatment covariants, a binary variable $W_i$ denoting assignment to treatment, and a d-vector $Y_{1i}$ \footnote{to ease exposition in the analysis of partial identification bound, we assume without loss of generality that $Y_{1i}$ lie in the interval $[0,1]$}of short-term post-treatment outcomes, which will pertain to wage or unemployment rates in our context. The researcher is interested in the effect of the treatment on a scalar long-term post-treatment outcome $Y_{2i} $ that is not measured in the experimental data. They can obtain an auxiliary, observational data set containing measurements for a separate population of individuals(with index $G=1$), which consist of the same covariates $X_i $, treatment, $W_i$, short-term outcome $Y_{i1} $ and also the long-term outcome $Y_{i2}$. We will omit the individual subscript $i$ when it does not lead to confusion. To embed this into a potential outcome framework(see imbens2015causal) , we posit that under the true data generating process, we obtain $(Y_{2i}(1),Y_{2i}(0),Y_{1i}(1),Y_{1i}(0), W_i,X_i, G_i) $ drawn from a distribution $P_{\star} \in \mathcal{M}_\lambda$, where the potential outcomes are w.r.t. $W_i$, the treatment. On the other hand, for the observable data, we obtain for each tuple $(Y_{i1}, W_i, X_i,G_i=0)$ in the experimental data set while $ (Y_{i2}, Y_{i1}, W_i, X_i, G_i = 1)$ for the observational data. We maintain the usual potential outcome assumptions such as SUTVA in imbens2015causal in the main body. We consider relaxing this assumption under general equilibrium effect consideration in the Appendix.

center[center omitted — 389 chars of source]

The target estimand for Part I is $\tau = E[ Y_2(1) | G = 1 ] -E[ Y_2(0) | G = 1 ]$, which is the long-term difference in the treated and untreated potential outcome in the observational data, which is by definition the population of interest having external validity. We are interested in combining the above two datasets to establish point identification.

We introduce the four assumptions laid out in athey2020combining to establish identification of the long-term ATE. The first assumption, Experimental internal validity, will be satisfied by construction in cases where data is collected from a randomized control trial.

assumption[Experimental internal validity] $W \amalg Y_2(1), Y_2(0), Y_1(1), Y_1(0) | X, G=0$.

The next assumption is less commonly imposed and might be questionable, analyzed in depth in Section (ref).

assumption[External validity of experiment] $G \amalg Y_2(1), Y_2(0), Y_1(0), Y_1(0) | X$.

The strict overlap condition below should always be coupled with the external validity assumption that asserts conditional independence with the location and potential outcomes.

assumption[Strict overlap] The probability of being assigned to treatment or of being measured in the observational data set is strictly bounded away from zero and one, i.e., for each $w$ and $g$ in $\{0,1\}$, the conditional probabilities \begin{equation} P(W = w | Y_1(1), X, G = g ) \quadand\quad P(G=g | Y_1(1), X, W = w) \end{equation} are bounded between $\epsilon $ and $ 1 - \epsilon $, $ \lambda$-almost surely, for some fixed constant $ 0 < \epsilon < 1/2 $.

Finally, the following latent unconfounding assumption is the novel assumption proposed by athey2020combining in order to establish identification of the long-term treatment outcome. It states that treatment is independent of the long-term outcome in the observational data, conditioned on the potential short-term outcome along with baseline covariates. Its novelty is the potential outcome, which turns out to have a non-trivial difference from the usual observed treatment outcome conditioning. See the discussion in athey2020combining for details. The latent nature and post-treatment outcome conditioning of this assumption make its intuitive understanding slightly challenging, discussed in Section (ref)

assumption[Latent unconfounding] k $W \amalg Y_2(w) \mid Y_1(w), X, G=1\quad(w =0,1)$.

Theorem 1 of athey2020combining proves that the long-term outcome is nonparametrically identified with the above four assumptions.

proposition[athey2020combining Theorem 1] Assume the above four assumptions. Then $E[ Y_2(1) | G = 1 ] - E[Y_2(0) | G=1]$ is nonparametrically identified.

This result is very insightful because it provides a way to solve the challenge of long-term treatment effects, utilizing two distinct and complementary data. In this sense, athey2020combining is one of the pioneering works that proposed an effective usage of experimental and observational data. However, at the same time, we note that the cross locational independence assumption (Assumption (ref)) and a form of unconfoudning conditioned on a latent variable (Assumption (ref)) may be questionable under certain contexts. The direct motivation of our paper is thus to take a closer look at both these external and internal validity assumptions and explore how we may be able to relax, interpret, or test these assumptions. This will be the central theme for Part I.

We first look at the external validity consideration, and then to internal validity. This reflects our stance towards combining data for causal inference that unless the experimental data clearly does not have external validity, it is worth trying combining data for estimating treatment effects. To formalize this, we first develop a test based on partial identification for checking external validity. Since it is based on partial identification, the power is relative low. Therefore, if the test is rejected, that should be a strong warning against data combination. If we do not reject the test, we proceed to the next step of considering the internal validity assumptions.

External validity

This section will focus on the external validity assumption (Assumption (ref)). We specifically do two things. We first show that the conditional independence assumption between $G$ and $(Y_2(1), Y_1(1))$ can be relaxed to only that of $G$ and $Y_1(1)$. Second, based on those relaxations, we consider various partial identifying assumptions specialized for this setting and their sharp identification bounds.

Dropping the external validity assumption for long-term potential outcomes

We will first show that Assumption (ref) can be relaxed to independence only concerning the short-term outcome, formalized in the following assumption and proposition. The key intuition for this relaxation is to notice that the external validity assumption is used in the identification proof of athey2020combining two times. The aggregate bias that appears without the external validity assumption on the long-term potential outcome can be shown to be of a simple form, which cancels out under the conditional independence of the short-term outcome.

assumption[Short term external validity] $G \amalg Y_1(1) ,Y_1(0) \mid X$.
proposition[Nonparametric identification of long-term ATE under weakened external validity assumption] The long-term average treatment effect in the observational population $ E[Y_2(1) -Y_2(0) | G=1]$ is nonparametrically identified replacing Assumption (ref) with Assumption (ref).
remark[Proof sketch of Proposition (ref)] We can get the following equation without any usage of the external validity assumption; \begin{align*} E[ Y_2(1) | G=1 ] = E[ E[ E[ Y_2 | Y_1 , W=1, X, G=1] | W=1, X, G=0] | G=1] + B + A \end{align*} where $B+A$ is \begin{align*} A + B =& E[ E[ Y_2(1) | X, G=1] - E[ Y_2(1) | X, G= 0 ] | G=1] \\ & \quad- E[ E[ Y_2(1)| X, G = 0] | G = 1] - E[ E[ E[ Y_2(1)| Y_1(1), X, G=1 ] | X, G = 0] | G = 1] \end{align*} This can be shown to be equal to \begin{align*} E\left[ Y_2(1)\left( 1 - \frac{ p( y_1(1) | X, G=0)}{ p( y_1(1) | X, G=1) }\right) \middle| G=1 \right] \end{align*} which is clearly equal to 0 under only the short-term outcome external validity.
remark[Intuition for the result of Proposition (ref)] This result implies that we can equally identify the long-term ATE regardless of the long-term outcome external validity, and the external validity assumption for long-term ATE is redundant. While it is difficult to generalize this result, we think the latent unconfounding assumption maintained throughout is related. The latent unconfounding assumption states that all essential information on the long-term outcome can be captured by the short-term potential outcome, implying that if external validity for the short-term outcome holds, it should also hold for the long-term. It would be an interesting direction for future work to see whether these results can be proven in other contexts.

Under this relaxation, we show how to test the external validity assumptions using partial identification methods. This approach is similar in spirit to, for example, blundell2007changes, which used it to test the gender differential wage gap in the UK. However, in their case, the problem was selection (internal validity), while we are concerned with the external validity assumption.

Partially identifying assumptions for testing external validity

An implication of relaxing the external validity assumption is that \[ F(Y_1(1) | X, G=1) = F(Y_1(1) | X, G=0).\] Rewriting from what we know, the right-hand side is nonparametrically identified because

align[align omitted — 138 chars of source]

For the left-hand side, by the law of iterated expectation and chain rule, and consistency

align[align omitted — 116 chars of source]

and hence the counterfactual quantity is only $F(Y_1(1) | W=0, X, G=1) $. A pointwise sharp bound under no additional assumption can be easily seen to be

align[align omitted — 174 chars of source]

and was indeed implemented in works like blundell2007changes, manski2009identification.

The bound above under no additional assumptions is pointwise sharp for each t. However, in practice, this will often be quite wide. On the other hand, researchers may have prior beliefs motivated by the economics of the problem and are interested in how the assumptions may translate into narrower bounds and the test results under such assumptions. The partial identification framework we provide below allows for such procedures in a systematic manner. We will consider two large classes of assumptions, the classical cross treatment($W$) assumptions and the novel cross location/design($G$) assumptions.

enumerate[(1)] • cross treatment ($W$) assumptions \begin{enumerate}[i] • Dominance conditions \begin{description} • \[\forall_{t\in\mathcal{T}}\;\forall_{x\in\mathcal{X}}\;F(Y_1(1) \leq t |X=x ,W=1, G=1 )\geq F(Y_1(1) \leq t |X=x ,W= 0, G=1)\] this is a form of negative selection into treatment, which in the case of a job training program, ashenfelter1985susing,heckman1999economics have shown how people with lower potential wage or employment probability were more likely to apply for job training. Suppose one thinks that this result holding uniformly across the support of $y_1(1)$ might be implausible but is convinced that at least for the lower 25 percent quantile of the treated, such negative selection could be true. In that case, we could consider the following. • \[\forall_{t \leq t_{0.25}}\;\forall_{x\in\mathcal{X}}\;F(Y_1(1) \leq t |X=x, W=1, G=1 )\geq F(Y_1(1) \leq t |X=x ,W= 0, G=1)\] where $t_{0.25}$ is the lower 25 percent quantile of $ F(Y_1(1) \leq t |X=x W=1, G=1 )$, which is observable. \end{description} • Exclusion conditions On the other hand, we could consider an alternative class of assumptions that assume auxiliary covariates with specific independence or inequality restrictions. Exclusion restrictions, which can be used to construct intersection bounds, are one useful assumption. It is formalized below. In the context of a job training program, if we are interested in earnings, unemployment benefits are commonly deployed as an instrumental variable that does not directly affect the outcome, as in blundell2007changes. \begin{description} • \begin{align*} &\forall_{v_1\in\mathcal{V}}\; F(Y_1(1) \leq t |V =v_1 X=x , G=1) = F(Y_1(1) \leq t |X=x , G=1) \end{align*} If one does not believe that the exclusion restriction is feasible, one could consider the monotone iv assumption, where the directionality of the iv coincides with the outcome, formalized below. • \begin{align*} &\forall_{v_1,v_2\in\mathcal{V}}\;v_1 \leq v_2 \implies F(Y_1(1) \leq t |V=v_1, X=x , G=1)\\ &\hphantom{\forall_{v_1,v_2\in\mathcal{V}}\;v_1 \leq v_2 \implies F(Y_1(1) \leq}\leq F(Y_1(1) \leq t |V =v_2, X=x , G=1) \end{align*} In our problem, the V may be a baseline covariate like test scores in the past, with lower values in such test scores may indicate lower values of the unobserved employment rate. \end{description} \end{enumerate} • cross location/design ($G$) assumptions \begin{enumerate}[i] • First order restrictions \begin{description} • \[\forall_{t\in\mathcal{T}}\;\forall_{x\in\mathcal{X}}\;F(Y_1(1) \leq t |X=x W=0, G=1 )\leq F(Y_1(1) \leq t |X=x ,W= 0, G=0)\] The stochastic dominance condition has been primarily used for the selection problem ( cross treatment $(W)$ case). However, we could equally argue for this for cross-location/design $(G)$. This could possibly justified either from a design or location perspective. For the design view, recall that in our case, $G = 0$ is experimental population, where people for $W = 0$ and $W = 1$ are homogenous. On the other hand, $G = 1$ corresponds to the case of the observational design, where negative selection as illustrated in Assumption SD take place. If the two populations themselves do not differ greatly(e.g. nearby states in the U.S.), then Assumption SDL could be plausible. On the other hand, from the location perspective, if the two groups were very heterogeneous, Assumption SDL could hold. A extreme but clear example would be, $G = 1$ corresponding to labor outcomes of a developed country, whereas $G = 0$ corresponding to that in a developing country \end{description} • Second order restrictions \begin{description} • \begin{align*} &\exists_{t_0}\;\forall_{x\in\mathcal{X}} [ \forall_{t \leq t_0}\;F(Y_1(1) \leq t |V = X=x ,W= 0, G=1) \\ &\hphantom{\exists_{t_0}\;\forall_{x\in\mathcal{X}} [ \forall_{t \leq t_0}}\qquad\leq F(Y_1(1) \leq t |V = X=x ,W= 0, G=0) \\ &\qquad\qquad\land \forall_{ t \geq t_0}\;F(Y_1(1) \leq t |V = X=x ,W= 0, G=0) \\ &\hphantom{\exists_{t_0}\;\forall_{x\in\mathcal{X}} [ \forall_{t \leq t_0}}\qquad\leq F(Y_1(1) \leq t |V = X=x ,W= 0, G=1)] \end{align*} This is differnt in nature to the previous one and discusses second order(variance) property of the two population. The basic intuition is captured in the illustration below. It is often the case that the observational data have people with covariates with a wider range, dispersed and heterogeneous. In contrast, the experimental data based on eligibility rules and the design of experiments are restricted to a more concentrated covariate value. In that case, the distribution of the observational quantile would be much wider than the experimental quantile. This difference in heterogeneity could plausibly lead to a reversal of the cumulative probability before and after some threshold $t_0$ \end{description} \end{enumerate}
figure[figure omitted — 4,170 chars of source]
remark[Relation of this framework to the literature] We briefly mention how the scheme proposed above relates to the literature on sensitivity analysis and partial identification. Firstly, in terms of assessing external validity, there is growing research in the sensitivity analysis literature that assesses the generalizability or transportability of a result in one location / research design, e.g., de2022improving. ; Nguyen et al. (2017) present a sensitivity analysis for omitted moderators, unobserved in both the experiment and the population, using linear models for expressing heterogeneity;Nie et al. (2021) introduce a non-parametric, percentile bootstrap approach to sensitivity analysis for generalization, which requires researchers to estimate a worst-case bound for the odds ratio between the misspecified and true experimental sample selection propensities; Dahabreh et al. (2019) proposed an alternative approach to adjust for bias in estimation by directly modeling the bias In practice, the most common approach first models the experimental sample inclusion probability, with the PATE then estimated using inverse probability weighted estimators (Stuart et al., 2011; Tipton, 2013; Buchanan et al., 2018). Alternative estimators focus on modeling treatment effect heterogeneity (Kern et al., 2016; Nguyen et al., 2017) or doubly robust estimation (Dahabreh et al., 2019). For the partial identification approach, however, few works use it to assess external validity. While manski2013public provides a framework of 'external assessment' using bounds, it takes a more social decision theory, minimax perspective, which is a different approach from ours. In this light, the novelty of our partial identification approach to external validity can be rephrased as follows. Consider the diagram below. While past literature has focused on the assumption that assesses the assumption on selection (across treatment groups $W$)(e.g., blundell2007changes,kreider2012identifying,andrews2019inference,lee2009training),here we also consider the cross location/design assumptions, and moreover on the intersection of those two, delivering sharp identification bounds each time. The importance of this novelty relates to the difficulty of gaining informative bounds solely from cross-treatment $W$ variation. With additional reasonable cross-locational variation assumptions, we might obtain informative bounds( leading to improved-powered tests) by taking the intersection. At the minimum, we provide a more diverse choice of partially-identifying assumptions that practitioners can choose from.
figure[figure omitted — 577 chars of source]

Uniformly sharp bounds

Having motivated some plausible assumptions, we are interested in how they translate to sharp identification bounds. Sharp identification bounds are the necessary and sufficient subset of the entire parameter space consistent with the given assumptions. In other words, it is the largest set within the parameter space that cannot be rejected by the given assumptions tamer2010. When interested in the distribution function, we would like to distinguish between pointwise and uniformly sharp bounds. Pointwise sharpness arises when we fix a particular value $t_0\in \mathcal{T}$ and observe the value a particular distribution function $F(Y_1(1)\leq t_0 | X)$ can take. However, if we see $F(Y_1(1)\leq t| X)$ as a function of t, the pointwise bound is not sharp in that it can contain distribution functions incompatible with the data. For example, it does not restrict that for any $ t_0 \le t_1$,

align[align omitted — 150 chars of source]

, which rules out some classes of distribution functions that are within the point-wise sharp bounds. See molchanov2018random,manski2003partial for more discussion. \begin{en-text} This is easiest to see in a simple counterexample as in the following: omit $\mathbf{x}$, and $G=1$ for simplicity, and let $P ( \mathbf{W} = 1 ) = \frac{2}{3} $ and let

align[align omitted — 176 chars of source]

{\color{orange}pending} Consider the distribution function

align[align omitted — 300 chars of source]

For each $t \in\mathbb{R}$, $F(t)$ lies in the tube defined by the pointwise sharp bound above. However it cannot be the CDF of $ Y_1$, because $ F(2) - F(1) = \frac{1}{9}< P (1 \leq Y_1(1) \leq 2 | W=1) P( W=1) $, directly contradicting ((ref)).

What is then the uniformly sharp identification bound in this case? The following proposition provides an answer.

propositionThe uniformly sharp identification bound for $ F(Y_1(1) | X, G=1)$ under the maintained assumptions in athey2020combining is \begin{align*} &\forall_{K\subset\mathcal{Y}}\;H_P[ Q(y | x)] =\\ &\quad \{ \tau(x) \in \mathcal{T} : \tau_{K(x)} \geq \mathrmP( Y_1(1) \in K | \mathbf{x}=x, W=1,G=1) \mathrmP(W =1 | \mathbf{x} = x, G=1)\} \end{align*} where, $\mathcal{Y}= \{0,1\}$.
proof

\end{en-text}

As the above observation suggests, it is in genreal difficult to get uniformly sharp bounds. We, therefore, rely mainly on a recent growing literature on random set theory molchanov2005theory to mechanically derive sharp bounds. The key intuition for why random set theory can be used to derive sharp identification bounds is that when a feature of a probability distribution of interest is partially identified, it is often possible to trace back the lack of point identification to the fact that either the data or the maintained assumptions yield a collection of random variables which are observationally equivalent. This collection is equal to the family of selections of a properly specified random closed set, and random set theory can be applied. Below, we provide the fundamental definitions and theorems that we will use and refer the reader to molchanov2005theory,molchanov2018random for the proofs and more in-depth explanations. In the sequel, we denote $\mathcal{Y}$ as the support of $Y_1(1)$ and standardize and normalize so that $\mathcal{Y} \subseteq [0,1]$. Also, we denote $\mathcal{K}(\mathcal{Y})$ as the family of compact sets that are subsets of $\mathcal{Y}$($\mathcal{K}$ will denote all compact subsets of $R^d$ )

definition[Random closed set] A map $\mathbf{X}$ from a probability space $( \Omega,\mathfrak{F},\mathbb{P} ) $ to the family $\mathbf{F} $ of closed subsets of $ \mathbb{R}^d $ is called a random closed set if \begin{equation} \mathbf{ X}^-(K) := \{ \omega \in \Omega : \mathbf{X} (\omega) \cap K \neq \emptyset \} \end{equation} belongs to the $\sigma $-algebra $\mathfrak{F} $ on $\Omega$ for each compact set $K$ in $\mathbb{R}^d$

The next is the capacity functional or equivalently as a containment functional that will be the key device that characterizes the sharp bounds when using the Artstein Inequality presented below.

definition[Capacity functional and containment functional] \begin{enumerate}[1.] • A functional $T_\mathbf{X}( K) : \mathcal{K} \to [0,1] $ given by \begin{equation} \mathrmT_{\mathbf{X}}( K) = \mathbb{P} ( \mathbf{X} \cap K \neq \emptyset ),\qquad (K \in \mathcal{K} ) \end{equation} is called capacity (or hitting) functional of $\mathbf{X}$. • A functional $\mathrmC_\mathbf{X}(F) : \mathbf{F} \to [0,1] $ given by \begin{equation} \mathrmC_\mathbf{X}(F) = \mathbb{P}(\mathbf{X} \subset F),\qquad (F \in \mathbf{F}) \end{equation} is called the containment functional of $\mathbf{X}$. \end{enumerate}

The next definition is the Measurable selection, and understanding what this will be in each context is usually the first step in getting sharp identification bounds.

definition[Measurable selection] For any random set $\mathbf{X}$, a (measurable) selection of $\mathbf{X}$ is a random element $\mathbf{x}$ with values in $\mathbb{R}^d$ such that $\mathbf{x}(\omega) \in \mathbf{X}( \omega)$ almost surely. We denote by $\mathrm{Sel}( \mathbf{X} )$ the set of all selections of $\mathbf{X}$.

Provided with these definitions, we can present the fundamental theorem we utilize to characterize sharp identification bounds, the Artstein Inequality. Usually, bounds that are not sharp (those that characterize the outer region) impose insufficient numbers of these Artstein Inequalities. The below theorem specifies how much is necessary and sufficient.

lemma[Artstein inequality] A probability distribution $\mu$ on $ \mathbb{R}^d$ is the distribution of a selection of a random closed set $\mathbf{X}$ in $\mathbb{R}^d$ if and only if \begin{equation} \mu(K) \leq \mathrmT(K) = \mathbb{P} \{ \mathbf{X} \cap K \neq \emptyset \} \end{equation} for all compact sets $K \subseteq \mathbb{R}^d $. Equivalently, if and only if \begin{equation} \mu(F) \geq \mathrmC(F) = \mathbb{P} (\mathbf{X} \subset F ) \end{equation} for all closed sets $ F \subset \mathbb{R}^d $. If $ \mathbf{X} $ is a compact random closed set, it suffices to check ((ref)) for compact sets $F$ only.

Given the setup above, we present the uniformly sharp bounds under the aforementioned assumptions.

First, we will provide the uniformly sharp bounds for $F(Y_1(1) = y | X, G=1)$ under the partial identifying assumption for the cross-treatment ($W$) cases.

theorem[Uniformly sharp bounds for $F(Y_1(1) = y | X, G=1)$ under cross treatment ($W$) designs] \begin{enumerate}[(i)] • Assume Assumption (SD). Absent any other information, the sharp bounds for $P(Y_1(1) = y | X, G=1)$ is \begin{align*} &H[ P( Y_1(1) = y | X, G =1 ) ] =\\ &\left\{\mu \in \Gamma \;\middle|\;\begin{array}{l} \mu ( K| X=x, G=1) \leq T_{\mathbf{Y} } ( K|X=x, G=1)\\ \mu ( [0, t]| W=1,X=x, G=1 ) \geq \mu( [0,t ]| W= 0, X=x, G=1 )\end{array} \forall_{K \in \mathcal{K}(\mathcal{Y})}\;\forall_{ x \in \mathcal{X}}\;\forall_{ t \in [0,1]} \right\}, \end{align*} where we will maintain that the random set $\mathbf{Y}$ is \[\mathbf{Y} = \begin{cases} [y_1] & \text{if}\quad W=1\\ \mathcal{Y} & \text{if}\quad W =0 \end{cases}\] throughout the paper unless otherwise mentioned. • Assume Assumption (LQD). Absent any other information, the sharp bounds for $P(Y_1(1) = y | X, G=1)$ is \begin{align*} &H[ P( Y_1(1) = y | X, G =1 ) ] = \\ &\left\{\mu \in \Gamma\middle|\begin{array}{l} \mu ( K| X=x, G=1) \leq T_{\mathbf{Y} } ( K|X=x, G=1)\\ \mu ( [0, t]| W=1,X=x, G=1 ) \geq \mu( [0,t ]| W= 0, X=x, G=1 )\end{array} \forall_{K \subset \mathcal{K}(\mathcal{Y})}\;\forall_{ x \in \mathcal{X}}\;\forall_{ t \in [0,t_{0.25}]} \right\} \end{align*} • Assume Assumption (EX). Absent any other information, the sharp bounds for $P(Y_1(1) = y | X, G=1)$ is \begin{align*} &H[ P( Y_1(1) = y | X, G =1 ) ]\\ &\scalebox{0.9}{$= \left\{ \mu \in \Gamma \;\middle|\;\begin{array}{l} \mu ( K|X=x,G=1) \\\quad\geq \operatorname*{ess\,sup}_{ v \in \mathcal{V} } P ( y \in K | W = 1, X= x , v,G=1) P ( W=1 | v,G=1)\end{array} \forall_{ K \subset \mathcal{K}(\mathcal{Y})}\;\forall_{ x \in \mathcal{X}} \right\}$} \end{align*} • Assume Assumption (MIV). Absent any other information, the sharp bounds for $P(Y_1(1) = y | X, G=1)$ is \begin{align*} &H[ P( Y_1(1) = y | X, G =1 ) ]\\ &\scalebox{0.9}{$= \left\{ \mu \in \Gamma \;\middle|\;\begin{array}{l} \mu ( K|X=x,G=1) \\\quad\geq \int \operatorname*{ess\,sup}_{ v_1 \geq v\in \mathcal{V} } P ( y \in K | X= x , v_1,G=1) P (V=v|X=x,G=1) dv\end{array} \forall_{ K \subset \mathcal{K}(\mathcal{Y})}\;\forall_{ x \in \mathcal{X}} \right\}$} \end{align*} \end{enumerate}
remark[Sketch of proof for Theorem (ref)] Recall that we assumed that $Y_1(1)$ was standardized and normalized to lie in the interval $[0,1]$, the binary outcome being a special case. Most parts of the proofs here are straightforward usage of the Artstein Inequality. The general direction for the proof is first to characterize a closed random set incorporating the assumptions and then use the Artstein Inequality on that set.
remark[Discussion of the sharp bounds in Theorem(ref) ] One may notice that the resulting sharp bounds for the Dominance conditions((i),(ii)) and the Exclusion conditions((iii),(iv)) have different shapes. Firstly, note that in (i) and (ii), we have only considered stochasticinequality restrictions,i.e. inequality relations that hold only in probability, rather than almost surely. In this case, these restrictions does not change the random set \mathbf{Y}, the selection of which must contain the data with probability one. Therefore, in (i) and (ii), the stochastic dominance assumptions simply appear as auxiliarly moment inequalities. Secondly, for (iii) and (iv), we observe there is a supremum with respect to the excluded variable $v$. The intuition is that when the independence of $v$ and $Y_1(1)$ hold, inequality restrictions should hold for any value of $v$, so the sharpest bound is attained by taking the supremum with respect to v, resulting in bounds known in the literature as intersection bounds becuase the bounds appear by taking the intersection of moment inequalities indexed by $v$.chernozhukov2013intersection,beresteanu2012partial,molchanov2018random,manski2003partial. Note that even in this case, the marginal distribution is contained in the same random set \mathbf{Y}. In Section (ref) dealing with a different problem, we will consider a Monotone Treatment Reponse assumption $Y_{it}(1) \geq Y_{it}(0) a.s. $, which changes the random set itself.
remark[Reducing the number of sufficient inequalities by utilizing Core determining class strategy] Theorem (ref) says that in order to get uniformly sharp bounds, we must consider all compact subsets of the support of Y,$\mathcal{Y}$. In the case of binary outcome, the compact subsets are merely $\{ \{0\}, \{1\}, \{0,1\}, \emptyset \}$, and easy to verify. However, when $\mathcal{Y}$ becomes larger, the number of moment inequalities increase exponentially. One promising strategy is to use the concept of core determining classto reduce the moment inequalities that are necessary and sufficient for the sharp bounds. For example, in beresteanu2012partial, they show that if $\mathcal{Y}$ is countable, it is sufficient to verify the Artstein Inequality for singleton subsets of $\mathcal{Y}$.(for \mathcal{Y}= \{0,1\}, we only need two inequalities).If $\mathcal{Y}=[0,1]$, verifying the moment inequalities for closed intervals in $[0,1]$ is sufficient. This core-determining class strategy has been succesfully applied in, e.g., galichon2009test for partially identified structural models to turn a intractable problem manageable.

Next, we will look into the cross-locational / design assumptions.

The following theorem provides the sharp bounds under the proposed assumptions. The structure of the sharp bounds closely align Dominance conditions of cross treatment ($W$).

theorem[Uniformly sharp bounds for $F(Y_1(1) = y | X, G=1)$ under cross design/location($G$) assumption] \begin{enumerate}[(i)] • Assume Assumption (SDL). Absent any further information, the sharp identification bound for $P( Y_1(1) | G=1) $ is \begin{align*} &H[ P( Y_1(1) = y | X, G =1 ) ] \\ &= \Biggl\{ \mu \in \Gamma :\begin{array}{l} \mu ( K| X=x, G=1) \leq T_{\mathbf{Y} } ( K|X=x, G=1)\\ \mu ( [0, t]| W=0,X=x, G=0 ) \leq \mu( [0,t ]| W= 0, X=x, G=1 )\end{array}\forall_{ K \subset \mathcal{K}(\mathcal{Y})}\;\forall_{ x \in \mathcal{X}} \; \forall_{ t \in [0,1]} \Biggr\} \end{align*} • Assume Assumption (HET). Absent any other condition, the sharp identification bound for $P( Y_1(1) | G=1) $ is \begin{align*} &H[ P( Y_1(1) = y | X, G =1 ) ] = \left\{ \mu \in \Gamma\middle|\begin{array}{l} \mu ( K| X=x, G=1) \leq T_{\mathbf{Y} } ( K|X=x, G=1) \\ \mu ( [0, t_1]| W=1,X=x, G=1 ) \geq \mu( [0,t_1 ]| W= 0, X=x, G=1 )\\ \mu ( [0, t_2]| W=1,X=x, G=1 ) \leq \mu( [0,t_2 ]| W= 0, X=x, G=1 )\\\forall_{ K \subset \mathcal{K}(\mathcal{Y})}\;\forall_{t_1 \in [0, t_0 ] } \forall_{t_2 \in [t_0, 1] },\forall_{ x \in \mathcal{X}}\end{array}\right\} \end{align*} \end{enumerate}
remark[Estimation consideration of sharp identification bounds] Note that all the above bounds ignore sampling uncertainty. We need to develop bounds for the bounds when for empirical investigation. In terms of inference, testing under conditional moment inequalities is an active area of research. We will deploy uniform confidence bands in the binary outcome (like unemployment rates) and turn to the framework of chernozhukov2019inference when the outcome is in general continuous and many moment inequalities are needed for uniformly sharp bounds. We discuss these issues in Section (ref).
remark[Sharp identification bounds under combination of cross treatment($W$)and location/design($G$)assumptions ] As stated in Remark (ref), one of the contributions of this paper is to consider both cross treatment($W$)and location/design($G$)assumptions. Natural interest would be in getting sharp identification bounds under those combinations. Although we do not fully present all combination of results in the main body due to limited space, we note that for the assumptions we have placed for the two regimes, they are almost immediately available. This relates to our previous remark on the shape of the bounds for cross treatment$W$( Remark (ref)) in that the stochastic dominance conditions do not directly affect the random closed set, but merely works as auxiliary moment inequalities. This implies that we can get sharp bounds under, for example, both the LQD assumption and the LSD assumption by simply adding both auxiliary inequalities. Analogous results hold for combining the intersection bounds( EX, MIV) and stochastic-dominance-type conditions(LSD, HET). Random set theory provides a simple, unified expression for sharp identification bounds, facilitating flexible partial identificaion analysis.
remark[Importance and Applicability of the framework in this section] In this section, we have provided a framework to relax and test the external validity assumption in this setting. Note that external validity is the key justification for sticking together two different datasets. Notably, our method, especially the partial identification method, can be applied in most setting, although the tightness of the bounds could vary among applications. We think that doing such a simple and universally applicable test should precede attempts to combine datasets in the future.

We will next deal with the latent unconfounding assumption, which pertains to the conditional internal validity of observational data.

Internal validity of observational data: latent unconfounding and equi-confounding bias assumption

Introducing the equi-confounding bias assumption

Unlike the previous section, we will focus on the internal validity of the observational data, i.e., under what conditions does $W$ be independent of $Y_t(w)$. As hinted in the introduction, a problem of non-nested approaches under the same data availability may be the first issue that practitioners have to deal with when having the dataset and deciding ways to establish identification. Specifically, the alternative approach for identification of the long-term ATE was recently proposed by ghassami2022combining, which posited an equi-confounding bias assumption, a form of parallel trends assumption applied to the data-combination setting. We present the conditional independence version of the original mean independence, to facilitate the comparison with latent unconfounding below.

assumption[Equiconfounding bias assumption(conditional independence version)] \begin{enumerate}[(i)] • $(Y_2(0)-Y_1(0)) \amalg G $$ (Y_2(1) - Y_1(1)) \amalg G $ \end{enumerate}

They proved the following theorem in their paper.

proposition[Nonparametric identification of LTATE under equi-confounding bias assumption ghassami2022combining] \begin{enumerate} • Replacing latent unconfounding assumption with (ref) (i) and maintaining all the other assumptions, the long-term ATT is identified, where long-term ATT is $E[Y_2(1) |W=1,G=1] - E[ Y_2(0) | W =1, G=1] $. • Replacing latent unconfounding assumption with (ref) (i) (ii), and maintaining all the other assumptions, the long-term ATE is identified. \end{enumerate}

The equiconfounding assumption can be intuitively explained that the potential growth between the short-term and long-term outcome is the same among the treated and the untreated. In the standard Difference-in-difference (DID) setup, usually, only the untreated potential outcome ((ref)(i) )is maintained to identify the ATT. An additional assumption on the equivalent growth among the treated potential outcomes((ref)(ii)) identifies the ATE.

This insightful result actually poses a dillemma upon the practioners placed with the longitudinal observational data and short-term experimental data. When only the latent unconfounding framework was available, they would have to justify the latent unconfounding assumption or give up point identification in this setting. However, now they have two different non-nested assumptions that they could rely on, and the distinction between them is not obvious. To provide one solution, we propose to focus on the selection mechanism. We use classical seletion mechanisms like ashenfelter1985susing,roy1951someas a buffer between the data(problem) and the assumption written in potential outcomes. We will see how this framework translate the difficult problem of choosing between novel assumptions into a familiar and interpretable problem in economics. For concreteness, in the main body, we model the outcome as \[\forall_{i\in[0,N]}\;\forall_{t=1,2}\;Y_{it} (0) = \alpha_i + \lambda_t + \alpha_i \lambda_t +\epsilon_{it},\; E[ \epsilon_{it} ] = 0.\] $Y_{it}(1) = Y_{it}(0) + \delta_{it}$, which generalizes the classical two-way-fixed effect model by (i) allowing for interactive fixed effect as in bai2009panel,abadie2021using (ii) allowing for arbitrary treatment effect heterogeneity while the two-way fixed effect model usually assumes constant treatment effect. we abstract from baseline covariates $X$ to ease exposition. We first consider how the two assumptions compare in estimating the long-term ATE. After this exposition, we see how the sufficient assumptions will differ when the target estimand is the long-term ATT.

Using Selection mechanism to compare internal validity assumptions for long-term ATE

Selection mechanism of Ashenfelter and Card (1985)

We first consider the selection mechanism proposed in ashenfelter1985susing.

A selection mechanism proposed by ashenfelter1985susing is

align*[align* omitted — 195 chars of source]

where $\beta \in [0,1]$ is a discount factor and $\tilde{c} = c - \lambda_1 - \beta \lambda_2$. In the original paper, they use this simple model of program participation to predict the earnings histories of the trainees of a job training program( CETA). It would be of interest to know when does $W \amalg Y_2(w) | Y_1(w)$ and $W \amalg (Y_2(w) - Y_1(w) )$ hold when $W$ is modeled as above. The following theorem shows that equi-confounding is in general unavailable while latent unconfounding may survive if we can justify myopic decision making of agents and no treatment effect heterogeneity.

theorem[Comparison for long-term ATE under selection mechanism of Ashenfelter and Card, 1985] Consider the setting specified above and the Ashenfelter-and-Card(AC) mechanism. Then Latent unconfounding for the untreated potential outcome, i.e., $W \amalg Y_2(0) | Y_1(0) $ holds if and only if $\beta =0$. Suppose we further assume that treatment effects can change between time, but are deterministic(across individuals). In that case, the latent unconfounding for the treated potential outcome, i.e., $W \amalg Y_2(1) | Y_1(1) $ holds if and only if $\beta =0$. Equi-confounding assumption does not hold in general for any $\beta \in [0,1]$
remark[Proof sketch for Theorem (ref)] Note that \[W_i = 1\{ Y_{i1}(0) + \beta Y_{i2}(0) \leq 0\} = 1\{ \alpha_i ( 1 + \lambda_1 + \beta( 1 + \lambda_2 ) ) + \lambda_1 + \beta \lambda_2 + \epsilon_{i1} + \beta \epsilon_{i2} \leq 0 \}\] When $\beta=0$, we have that $W_i =1\{Y_{i1}(0) \leq c\}$. Then $ 1\{ Y_{i1}(0) \leq c \} $ can be regarded as a constant given $ Y_{i1}(0) $. Hence it must be independent of any random variable, hence $ W_i \amalg Y_{i2}(0) | Y_{i1}(0)$ . For the treated potential outcome case, we have to show that $ W \amalg Y_{i2} (1) | Y_{i1} (1) $. Under the assumption that $(\delta_{i1} , \delta_{i2} $ is deterministic( it does not vary across individuals), then the $\sigma$-algebra conditioning on $Y_{i1}(1) $ ($Y_{i2} (1)$)is equivalent to conditioning on $Y_{i1}(0)$($Y_{i2}(0)$). So the same argument as the untreated potential outcome follows. When $ \beta \neq 0 $,$ W_i $ contains $ \epsilon_{i2} $, which cannot be conditioned on by $ Y_{i1}(0) $. Thus, in general, $W_i$ is correlated with $Y_{i2}(0) $ which also contains $\epsilon_{i2} $, A similar argument holds for the treated potential outcomes. For the equi-confounding bias case, note that $Y_{i2}(0) - Y_{i1}(0) = ( \alpha_i +1)( \lambda_2 - \lambda_1 ) + \epsilon_{i2} - \epsilon_{i1} $. This in general contains $\epsilon_{i1} $, and will be correlated with $Y_{i1}(0) $, which also contains $\epsilon_{i1}$. Similarly for the treated potential outcome difference, note that$Y_{i2}(1) - Y_{i1}(1) $=$ \delta_{i2} - \delta_{i1} + ( \alpha_i +1)( \lambda_2 - \lambda_1 ) + \epsilon_{i2} - \epsilon_{i1}$. and the containment of $\epsilon_{i1}$ remains, so it will generally not hold for any $\beta > 0$ .
remark[Discussion of Theorem (ref)] The theorem shows that when we beleive that AC type selection mechanism is appropriate, then latent unconfounding is the only alternative we have to consider, because the equi-confoudnding will not hold by the inevitable correlation induced by the common variable $\epsilon_{i2}$. To justify latent unconfounding, that requires showing that the agent makes the participation decision with full discounting and how there is no treatment effect heterogeneity. While this is arguably strong assumption, we note that (i) these assumptions have direct economic interpretation, which facilitates concrete discussion (ii) we show in Section (ref) that when the target estimand is long-term ATT we can relax this condition to only the full discounting assumption. In short, simply positing a selection mechanism has many implication for deeper analysis on the validity of assumptions.

Selection mechanism of Roy (1951)

Next, we consider the Roy selection mechanism roy1951some. The essence of the Roy selection model is that it is based on treatment effects, opposed to the untreated potential outcome in ashenfelter1985susing.

Formally, in the Roy-selection model, selection is based on a function of the treatment effects $(\delta_{i1}, \delta_{i2})$ Based on the information set of the individual, it could be a nonlinear function of this treatments which we denote as $ W= f(\delta_{i1} , \delta_{i2} ) $, for some measurable, possibly nonlinear function of the treatment effects with its codomain being $\{0,1\}$.

In this case, the comparison of the validity of the two internal validity assumptions is provided in the following theorem. It says that latent unconfounding will be generally violated while the time differencing strategy may survive

theorem[Comparison for long-term ATE under selection mechanism of Roy, 1951] Consider the setting specified above and the Roy selection mechanism. Then Latent unconfounding does not hold for $f$ that have general dependence on $(\delta_{i1}, \delta_{i2})$. Equi-confounding assumption hold for any f if (i) $(\delta_{i1}, \delta_{i2} ) \amalg (\epsilon_{i1}, \epsilon_{i2} )$ ,(ii) the time fixed effect is, in fact, invariant across time, i.e., $\lambda_1 = \lambda_2 $, and (iii)$\delta_{i2} = \delta_{i1} a.s. $
remark[Discussion of Theorem (ref)] The theorem says that for the Roy-selection mechanism, treatment effects entering both in the selection and outcome variable renders latent-unconfounding implausible. On the other hand, the time-differencing strategy of the equi-confounding assumption will survive if (i)independence between treatment effects and the time-varying unobservables in the outcome hold, time-homogeneity type conditions are placed on (ii)treatment effects and (iii)time-fixed effects. The three assumptions placed on the equi-confounding case can be intuitively interpreted as information in the short term is sufficient for inferring the long term; if the dynamic problem can be reduced to a static problem, equi-confounding will hold for Roy mechanism. To see this, note that the independence between treatment effects and the time-varying unobservables will hold if heterogeneity is sufficiently captured by αi, the time-fixed effect. Also, the two time-homogeneity condition by definition implies staticness. The proof can be done is a similar way as Theorem (ref)and will be relegated to the Appendix.
remark[Comparison of latent unconfounding and equi-confounding in light of Theorem (ref) and (ref)] In Theorem (ref) and (ref) above, we have compared the latent unconfounding and equi-confounding assumption in light of the AC and Roy mechanism. The tree graph below illustrates the overview of what we have shown in this section. In general, if we believe that the current problem fits AC mechansim better, our only choice is to try justify the assumptions for latent unconfounding, and if that seems difficult, giving up point identification may be a prudent decision. On ther other hand, our only choice is to justify the essentially 'static' problem in the Roy selection mechanism case.
remark[The role of economic theory under the current framework] Under the workflow of 1. deciding the target estimand 2. choosing the suitable selection mechanism 3. justifying the sufficient conditions for the unconfoundedness assumption, economic theory plays a key role in Step 2 and 3. Notably, for 3, we have boiled the potential outcome assumption down to conditions that can be interpretable and discussed in economic terms, like macro-economic fixed effects, discounting, treatment-effect heterogeneity. Importantly, this connects to systematic investigation on macroeconomic/policy shocks, or connect to the rich economic literature on bounded rationality(e.g., brown1981myopic,aucejo2017identification, heterogeneity and its tests( e.g., hyslop1999state,wager2018estimation,heckman1981heterogeneity),
figure[figure omitted — 472 chars of source]

Extending comparison to long-term ATT

One might note that if we are interested in the long-term ATT, equi-confounding bias assumption requires only the parallel trend on the untreated potential outcomes(Proposition (ref)). This may raise the question of how the latent unconfounding and equi-confounding assumptions compare in the case of long-term ATT estimation. Can the equi-confounding case be relaxed? To see this, however, we need identification results of long-term ATT, which is not proposed in athey2020combining or any paper we are aware of.

Thus, we will first extend the identification result, proving that almost the same nonparametric identification for ATT follows under a slightly adjusted external valididy-on-the-treated assumption and only the untreated potential outcome version of the latent unconfounding assumption formalized in the assumption below.

assumption[External validity for the the treated sub-population] Assume the following conditional independence hold. \begin{equation*} G \amalg (Y_2(1) , Y_2(0), Y_1(1), Y_1(0) ) | X , W =1 \end{equation*}

Using this, we can get the following identification result. The proof is almost the same as the long-term ATE, which we relegate to the Appendix.

corollary[Nonparametric Identification of average treatment effect on the treated (ATT) ] Assume the same assumptions as Proposition (ref), but replace Assumption (ref) with Assumption (ref) and reduce Assumption (ref) only to the untreated potential outcome. Then the Average Treatment Effect on the Treated (ATT) $E[ Y_2(1) | W =1, G=1] - E[ Y_2(0) | W-1, G=1] $ is nonparametrically identified.

Provided with this identification result, we can compare the two assumptions for the long-term ATT.

The following results can be shown, similar to Theorem (ref)and (ref).

corollary[Comparison for long-term ATT under selection mechanism of Ashenfelter and Card(AC), 1985, Corollary of Theorem (ref)]. Consider the setting specified above and the AC mechanism. Then Latent unconfounding assumption is necessary for identifying the long-term ATT holds if and only if $\beta =0$. Equi-confounding assumption still does not hold in general for any $\beta \in [0,1]$
remark[Discussion of result in Corollary (ref)] Crucially, For long-term ATT identification, we do not need to assume that treatment effects are homogeneous for the latent-unconfounding assumption to hold. On the other hand, the equi-confounding still, in general, does not hold because there is the common unobserved time-varying element $\epsilon_{i1}$ both in $W$ and $Y_2(0) - Y_1(0)$

We also have a similar result for the Roy selection case.

corollary[Comparison for long-term ATT under Roy selection mechanism Corollary of Theorem (ref)] Consider the setting specified above and the Roy selection mechanism. Then latent unconfounding assumption hold if the time-varying unobservables and time fixed effect remain static, i.e.,$ (i)\lambda_1 = \lambda_2 \, a.s.$ $(ii)\epsilon_{i1} = \epsilon_{i2} \, a.s. $ Equi-confounding assumption hold for any f if (i) $\lambda_1 = \lambda_2 \, a.s.$ (ii)$(\delta_{i1}, \delta_{i2} ) \amalg (\epsilon_{i1}, \epsilon_{i2} )$
remark[Discussion of result in Corollary (ref)] The takeaway in this case is latent unconfounding assumption could possibly hold if we argue for the time homogeneity of time fixed effects and time-varying unobservables. Moreover, the equi-confounding bias assumption also benefits from changing the target estimand from the long-term ATE to ATT. The need to only justify $W \amalg (Y_2(0) - Y_1(0) | X, G=1$ renders the very strong assumption that treatment effects are time-invariant unnecessary. We only have to argue that treatment effects are independent of time-varying unobservables and that time-fixed effects are constant.

\color{black}

remark[Comparison of long-term ATE and ATT conditions] The two corollaries above suggest that the long-term ATT is more credibly estimable than the long-term ATE by relaxing, e.g., the no treatment effect heterogeneity assumption. In the static problem, the ATE and ATT assumptions' plausibility usually coincide. We show that things change when we extend to the dynamic case, which can be an imporant direction for future work.
remark[Dynamic confounding: General takeaway from the analysis of this section] One key insight from the series of results for Section (ref) is that dynamic considerations are the crucial determinants for the assumptions to hold. In many occasions, we had to assume time-homogeneity , essentially 'static' type of assumptions for the two assumptions to hold. Systematic treatment of dynamic confounding in The literature on long-term, to our knowledge, remains underdeveloped. This motivates a deeper investigation into the economics of dynamic decision-making under uncertainty. We will examine unemployment dynamics and the role of job training programs in Part II.

Inference for moment inequalities

In this section, we will discuss the estimation for the sharp identification bounds. When there are no covariates, linear programming methods in the spirit of balke1997bounds,torgovitsky2019nonparametric will provide us with the sharp bounds necessary. However, it is usually implausible for the external validity assumption to hold without conditioning on covariates; hence we will consider inference under conditional moment inequalities methods. We have provided many types of bounds in Section (ref), and rather than to provide inference on each of them one by one, we categorize them into three classes that possess their unique challenges; (i) small number of moment inequalities but high dimensional covariates (ii) many conditional moment inequalities (iii) intersection bounds that require supremum operator. (i) will primarily correspond to the worst-case bounds of binary outcomes, which we will tackle using honest uniform confidence bands methods based on chernozhukov2014anti. (ii) will mainly correspond to the case of continuous outcomes, which the framework of chernozhukov2019inference to automatically select the one that has power, will be useful. Case (iii) corresponds to when we want to use an exclusion restriction of an instrumental that is continuous, e.g., for unemployment benefit. We use the framework of chernozhukov2013intersection in this case, and we abstract from this discussion in the main body because of its similarity to (ii) and space limitations.

Note that for each case, high-dimensional Gaussian Approximation underlies the theoretical guarantees. Our novel contribution is to refine the generic Gaussian approximation to suprema of empirical processes, Theorem 2.1 in chernozhukov2014gaussian, by utilizing the symmetrization phenomena of the Stein exchangeable pair approach for approximation.

Challenge (i): small number of moment inequalities but high dimensional covariates

For the binary outcome case, since all compact sets we need to apply is only two(by the core-determining class argument in Remark (ref)) even for the uniformly valid case, we get a simple tractable form for the uniformly sharp bound. For example, in the worst case, we can apply the confidence band approach to test our inequality. For example, when we would like to test For example, when we would like to test

align[align omitted — 174 chars of source]

rewrite it as

align[align omitted — 238 chars of source]

((ref)) ,where with a slight abuse of notation, $x := (y_1, x )$(All the capital and lower-case x below will refer to this). For $f_1(x) $ we can estimate it using a series estimator with a suitable basis (e.g., Fourier basis), and construct a band that is uniformly valid over all $ x \in \mathcal{X} $ and $P \in \mathcal{P}$, where $\mathcal{P} $ is a prespecified nonparametric distribution (H\"{o}lder class, VC-type class in the original paper's example). If there exists some $x \in \mathcal{X} $ that has a value less than zero, we can reject the null hypothesis. The issue is choosing the critical value that adjusts for the bias.

We use the following theorem in chernozhukov2014anti that justifies a critical value constructed by the Gaussian multiplier bootstrap. Their theoretical analysis justifies it for estimators with smoothing parameters chosen by Lepski's method,

Formally, let $\hat{f}_n(\cdot, l)$ be a generic estimator of $f$ with a smoothing parameter $l$, say bandwidth or resolution level, where $l$ is chosen from a candidate set $\mathcal{L}_n$.Let $\hat{l}_n=\hat{l}_n\left(X_1, \ldots, X_n\right)$ be a possibly data-dependent choice of $l$ in $\mathcal{L}_n$. Denote by $\sigma_{n, f}(x, l)$ the standard deviation of $\sqrt{n} \hat{f}_n(x, l)$, that is, \[\sigma_{n, f}(x, l):=\left(n \operatorname{Var}_f\left(\hat{f}_n(x, l)\right)\right)^{1 / 2}.\] Then we consider a confidence band of the form

align[align omitted — 245 chars of source]

We use the theorem 4.1 in chernozhukov2014anti that justify critical values decided by the multiplier bootstrap for honest uniform confidence bands\footnote{ We say a confidence set $\mathcal{C}_n$ for $f$ is honest, at level $1-\gamma$, if it satisfies

align*[align* omitted — 101 chars of source]

where $\mathcal{F}$ is the entire family of functions $f$ we wish to adapt to li1989honest. Honesty is necessary to produce practical confidence sets; it ensures that there is a known time $n$, not depending on $f$, after which the level of the confidence set is not much smaller than $1-\gamma$. }. We will deploy the four conditions in their paper that justify honest uniform confidence bands.

The first one is the estimation method. It allows many methods, like convolution kernels and sieve methods with leading examples like Wavelet, Fourier bases.

customthm{L1}[Density estimator] The density estimator $\hat{f}_n$ is either a convolution or wavelet projection kernel density estimator defined in the original paper chernozhukov2014anti. For convolution kernels, the function $K: \mathbb{R} \rightarrow \mathbb{R}$ has compact support and is of bounded variation, and is such that $\int K(s) d s=1$ and $\int s^j K(s) d x=0$ for $j=1, \ldots, r-1$. For wavelet projection kernels, the function $\phi: \mathbb{R} \rightarrow \mathbb{R}$ is either a compactly supported father wavelet of regularity $r-1$ (i.e., $\phi$ is $(r-1)$-times continuously differentiable), or a Battle-Lemarié wavelet of regularity $r-1$.

The next one is the bounds for estimation bias.

customthm{L2}[Bias bounds] There exist constants $l_0, c_3, C_3>0$ such that for every $f \in \mathcal{F} \subset \bigcup_{t \in[L, t]} \Sigma(t, L)$, there exists $t \in[\underline{t}, \bar{t}]$ with \[ c_3 2^{-l t} \leq \sup _{x \in \mathcal{X}}\left|\mathrm{E}_f\left[\hat{f}_n(x, l)\right]-f(x)\right| \leq C_3 2^{-l t}, \] for all $l \geq l_0$.

The third one state an assumption on the candidate sets of the smoothing parameter $\ell $. It is required so that the fparameter value that leads to the optimal convergence rate exists.

customthm{L3}[Candidate set] There exist constants $c_4, C_4>0$ such that for every $f \in \mathcal{F}$, there exists $l \in \mathcal{L}_n$ with \[ \left(\frac{c_4 \log n}{n}\right)^{1 /(2 t(f)+d)} \leq 2^{-l} \leq\left(\frac{C_4 \log n}{n}\right)^{1 /(2 t(f)+d)} \] for the map $t: f \mapsto t(f)$(see equation (32) of chernozhukov2014anti for details). In addition, the candidate set is $\mathcal{L}_n=$ $\left[l_{\min , n}, l_{\max , n}\right] \cap \mathbb{N}$.

Finally, a regularity assumption on the boundedness of density estimators are placed.

customthm{L4}[Density bounds] There exist constants $\delta, \underline{f}, \bar{f}>0$ such that for all $f \in \mathcal{F}$, (34) $f(x) \geq \underline{f}$ for all $x \in \mathcal{X}^\delta$ and $f(x) \leq \bar{f}$ for all $x \in \mathbb{R}^d$, where $\mathcal{X}^\delta$ is the $\delta$-enlargement of $\mathcal{X}$, that is, $\mathcal{X}^\delta=\left\{x \in \mathbb{R}^d: \inf_{y \in \mathcal{X}}|x-y| \leq \delta\right\}$.

Under the four assumptions, the following theorem of chernozhukov2014anti is justified.

theorem[Honest and adaptive confidence bands via the Gaussian multiplier-bootstrap (chernozhukov2014anti)] Suppose that Conditions L1-L4 are satisfied. In addition, suppose that there exist constants $c_5, C_5>0$ such that: (i) $2^{l_{\max , n} d}\left(\log ^4 n\right) / n \leq C_5 n^{-c_5}$, (ii) $l_{\min , n} \geq c_5 \log n$, (iii) $\gamma_n \leq C_5 n^{-c_5}$, (iv) $\left|\log \gamma_n\right| \leq C_5 \log n$, (v) $u_n^{\prime} \geq C(\mathcal{F})$ and (vi) $u_n^{\prime} \leq C_5 \log n$. Then \[ \sup _{f \in \mathcal{F}} \mathrm{P}_f\left(\sup _{x \in \mathcal{X}} \lambda\left(\mathcal{C}_n(x)\right)>C\left(1+u_n^{\prime}\right) r_n(t(f))\right) \leq C n^{-c}, \] where $\lambda(\cdot)$ denotes the Lebesgue measure on $\mathbb{R}$ and $r_n(t):=(\log n / n)^{t /(2 t+d)}$. Here the constants $c, C>0$ depend only on $c_5, C_5$, the constants that appear in Conditions L1-L4, $c_\sigma, \alpha$, and the function $K$ (when convolution kernels are used) or the father wavelet $\phi$ (when wavelet projection kernels are used).
remarkThe above theorem shows that the confidence band constructed by the method proposed in chernozhukov2014anti is asymptotically honest at a polynomial rate for the class $\mathcal{F}$. By the equivalence of confidence intervals (bands) and testing, we could test whether the external validity condition is justified at a uniform level.

In order to construct honest confidence bands, chernozhukov2014anti place high-level assumptions on the nonparametric estimation procedure (H1) - (H6), where (H1) corresponds to Gaussian approximation. To verify that (H1) is satisfied for the procedure proposed in the paper for VC-type classes, they invoke a corollary of the Gaussian approximation presented above (Theorem 2.1 in the original paper). This was specialized to VC-type classes (Corollary 2.1 in the original paper) to guarantee the existence of sequences $\epsilon_{1n} $ and $\delta_{1n}$ bounded from above by $C n^{-c} $ for some constant. We can indeed show that the approximation rate can be improved from $O(n ^{ -1/6} ) $ to $O( n^{ -1/4} ) $, which is a nontrivial extension in the high-dimensional regime.

theorem[Original Gaussian approximation to suprema of empirical processes Chernozhukov et al., 2014b 2.1] Assume the following conditions: \begin{enumerate}[({A}1)] • point wise measurability of class $\mathcal{F}$. • the integrability of the envelope F; $\exists_{ q \geq 3}:[ F \in \mathcal{L}^q(P)]$ • class $\mathcal{F}$ is pre-Gaussian \end{enumerate} Let $ Z= \sup_{f \in \mathcal{F}} \mathbb{G}_n f $. Let $\kappa >0$ be any positive constant such that $\kappa^3 \geq E[ || E_n [ | f(X_i)|^3 ] ||_{\mathcal{F}}] $Then for every $\epsilon \in (0,1] $ and $\gamma \in (0,1] $, there exists a random variable $ \tilde{Z} =^d \sup _{f \in \mathcal{F}} G_P f $ such that \begin{equation} P\{ | Z - \tilde{Z}| > K(q) \Delta _n ( \epsilon, \gamma ) \} \leq \gamma ( 1 + \delta_n (\epsilon , \gamma )) + \frac{ C \log n}{n} \end{equation} where $K(q)> 0 $ is a constant that depends only on q, and \begin{align*} \Delta_n (\epsilon, \gamma )&: = \phi_n ( \epsilon) + \gamma ^{ -1/q} \epsilon || F _{P,2}\\ &\qquad+ n^{-1/2} \gamma ^{-1/q} ||M||_q + n^{-1/2} \gamma^{-2/q} ||M||_2\\ &\qquad+ n^{-1/4} \gamma^{-1/2} ( E[ || \mathbb{G}_n ||_{ \mathcal{F}\cdot \mathcal{F}}])^{1/2} H_n^{1/2} ( \epsilon)\\ &\qquad+ n^{-1/6} \gamma^{-1/3} \kappa H_n ^{2/3} ( \epsilon)\\ \delta_n ( \epsilon, \gamma )&:= \frac{1}{4} P \{ (F/\kappa)^3 1( F/ \kappa > c \gamma ^{-1/3}n^{1/3}H_n ( \epsilon)^{-1/3}) \} \end{align*}

Note the worst rate in $\Delta_n(\epsilon, \gamma )$ is the last term with the order $n^{-1/6}$, which I show can be refined to be $n^{-1/4}$, which seems to be a large gain in the high dimensional regime. Formally, we prove the following refined Gaussian approximation to suprema of empirical processes.\footnote{There is an additional regularity condition necessary on the finiteness of moments one order above the original, which we show in the appendix.}

theorem[Refined Gaussian approximation to suprema of empirical processes] Consider the same setting as above. Then \[P\{ | Z - \tilde{Z}| > K(q) \Delta _n ( \epsilon, \gamma ) \} \leq \gamma ( 1 + \delta_n (\epsilon , \gamma )) + \frac{ C \log n}{n}\] where \begin{align*} \Delta_n (\epsilon, \gamma )&: = \phi_n ( \epsilon) + \gamma ^{ -1/q} \epsilon \|F_{P,2}\| + n^{-1/2} \gamma ^{-1/q} \|M\|_q + n^{-1/2} \gamma^{-2/q} \|M\|_2\\ &\qquad+ n^{-1/4} (\gamma^{-1/4} \kappa H_n ^{7/8} ( \epsilon) +\gamma^{-1 / 2} n^{-1 / 4}\left(E\left[\left\|\mathbb{G}_n\right\|_{\mathcal{F} \cdot \mathcal{F}}\right]\right)^{1 / 2} H_n^{1 / 2}(\epsilon))\\ \delta_n ( \epsilon, \gamma )&:= \frac{1}{4} P[ F^4 1( \frac{F}{\kappa} > c \gamma^{-1/2} n^{1/4} H_n^{- 1/8}(\epsilon) ] \end{align*}
remark[Proof Sketch of Theorem (ref)] The key source of our improved result is the refined coupling inequality of Theorem 4.1 in chernozhukov2014gaussian with notation introduce in the formal proof in our Appendix, \begin{align*} P\left(|Z-\widetilde{Z}|>2 \beta^{-1} \log p+3 \delta\right) \leq \frac{\varepsilon+C \beta \delta^{-1}\left\{B_1+\beta\left(B_2+B_3\right)\right\}}{1-\varepsilon}, \end{align*} whereas the original version is \begin{align*} P\left(|Z-\tilde{Z}|>2 \beta^{-1} \log p+3 \delta\right) \leq \frac{\epsilon+C \beta \delta^{-1}\left\{B_1+\beta^2\left(B_2+B_3\right)\right\}}{1-\epsilon} \end{align*} , which we prove by refining the remainder term of the stein interpolation \begin{align*} R=& h\left(S_n^{\prime}\right)-h\left(S_n\right)-\left(S_n^{\prime}-S_n\right)^T \nabla h\left(S_n\right) \\ &-2^{-1}\left(S_n^{\prime}-S_n\right)^T\left(\operatorname{Hess} h\left(S_n\right)\right)\left(S_n^{\prime}-S_n\right) \end{align*} where $(S_n , S_n')$ are Stein-exchangeable pairs. We take advantage of the zero skewness of the distribution of $S_n - S_n'$ that leads to higher-moment matching. We finally apply this to the discretized empirical process, where the optimal $\delta$ in our case will be \begin{align*} \delta \geq C \max \left\{\gamma^{-1 / 2} n^{-1 / 4}\left(E\left[\left\|\mathbb{G}_n\right\|_{\mathcal{F} \cdot \mathcal{F}}\right]\right)^{1 / 2} H_n^{1 / 2}(\epsilon), \gamma^{-1 / 4} n^{-1 / 4} \kappa H_n^{7 / 8}(\epsilon)\right\} \end{align*} where the worst rate is of order$n^{-1/4}$, compared to the worst case of $n^{-1/6}$ in the original \begin{align*} \delta \geq C \max \left\{\gamma^{-1 / 2} n^{-1 / 4}\left(\mathbb{E}\left[\left\|\mathbb{G}_n\right\|_{\mathcal{F} \cdot \mathcal{F}}\right]\right)^{1 / 2} H_n^{1 / 2}(\varepsilon), \gamma^{-1 / 3} n^{-1 / 6} \kappa H_n^{2 / 3}(\varepsilon)\right\}, \end{align*} Using the classical empirical process inequalities(e.g., Borel-TIS inequality),vaart1996weak and Strassen's Lemma (Lemma 4.1 of chernozhukov2014gaussian, will complete the proof.

Challenge (ii): Inference under many (conditional) moment inequalities

We follow the framework of chernozhukov2019inference. We only slightly adjust their notation so that we can remain consistent with the rest of the paper. let $\bar{X}_1, \ldots, \bar{X}_n$ be a sequence of independent and identically distributed (i.i.d.) random vectors in $\mathbb{R}^p$, where $\bar{X}_i=\left(\bar{X}_{i 1}, \ldots, \bar{X}_{i p}\right)^T$, with a common distribution denoted by $\mathcal{L}_{\bar{X}}$. For $1 \leq j \leq p$, write $\bar{\mu}_j:=\mathrm{E}\left[\bar{X}_{1 j}\right]$. We are interested in testing the null hypothesis

align[align omitted — 115 chars of source]

against the alternative

align*[align* omitted — 79 chars of source]

Our problem fits in this framework because we assume to have i.i.d. data collection, and most of our restrictions using conditional probabilities can be rewritten as conditional moment inequalities. Conditional moment inequalities following andrews2013inference, in turn, can be rewritten into an infinite number of unconditional moment inequalities from which we can choose a finite collection based on an increasing sequence which increases sufficiently fast with n. See the footnote of chernozhukov2019inference for details.

We introduce the test statistic to conduct the test ((ref)), Assume that

align[align omitted — 186 chars of source]

For $j=1, \ldots, p$, let $\widehat{\bar{\mu}}_j$ and $\widehat{\sigma}_j^2$ denote the sample mean and variance of $\bar{X}_{1 j}, \ldots, \bar{X}_{n j}$, respectively, that is

align*[align* omitted — 301 chars of source]

The test statistic we consider here is \footnote{This follows the original paper chernozhukov2019inference, see discussion for other possible test statistics therein} \[T=\max _{1 \leq j \leq p} \frac{\sqrt{n} \widehat{\bar{\mu}}_j}{\widehat{\sigma}_j}\]

There were several methods proposed by chernozhukov2019inference, and we focus on the multiplier bootstrap method since it is consistent with the choice for the confidence band approach in the previous section.

Algorithm (Multiplier bootstrap) chernozhukov2019inference

enumerate• 1. Generate independent standard normal random variables $\epsilon_1, \ldots, \epsilon_n$ independent of the data $\bar{X}_1^n=\left\{\bar{X}_1, \ldots, \bar{X}_n\right\}$. • 2. Construct the multiplier bootstrap test statistic \begin{align*} W^{M B}=\max _{1 \leq j \leq p} \frac{\sqrt{n} \mathbb{E}_n\left[\epsilon_i\left(\bar{X}_{i j}-\widehat{\bar{\mu}}_j\right)\right]}{\widehat{\sigma}_j} . \end{align*} • 3. Calculate $c^{M B}(\alpha)$ as $c^{M B}(\alpha)=$ conditional $(1-\alpha)$-quantile of $W^{M B}$ given $\bar{X}_1^n$.

They justify the validity of the test based on the multiplier bootstrap as below.

theorem[Validity of one-step Multiplier bootstrap methods Chernozhukov et al., 2019a] Let $c^B(\alpha)$ stand either for $c^{M B}(\alpha)$ or $c^{E B}(\alpha)$. Suppose that there exist constants $0<c_1<$ $1 / 2$ and $C_1>0$ such that \begin{align} \left(M_{n, 3}^3 \vee M_{n, 4}^2 \vee B_n\right)^2 \log ^{7 / 2}(p n) \leq C_1 n^{1 / 2-c_1} \end{align} Then there exist positive constants c, $C$ depending only on $c_1, C_1$ such that under $H_0$, \begin{align*} \mathrm{P}\left(T>c^B(\alpha)\right) \leq \alpha+C n^{-c} . \end{align*} In addition, if $\bar{\mu}_j=0$ for all $1 \leq j \leq p$, then \begin{align*} \left|\mathrm{P}\left(T>c^B(\alpha)\right)-\alpha\right| \leq C n^{-c} . \end{align*} Moreover, both bounds hold uniformly over all distributions $\mathcal{L}_{\bar{X}}$ satisfying ((ref)) and ((ref)).

Hence, this concludes the section for estimation consideration for moment inequalities.

\part{Combining data to estimate the ATE on the treated survivors (ATETS) under dynamic confounding: point and partial identification}

ATE on the treated survivors :motivation and nonparametric point identification

Motivation for ATETS

In this part we will focus particularly on a binary outcome of unemployment and how a job training program may affect the outcome. When employment is concerned, there has been a particular interest in the duration of employment and unemployment;van2001duration,lancaster1979econometricdocument identification and estimation of duration parameters on unemployment from observational studies ; heckman1984method discuss how to relax distributional assumptions for identification ; honore1993identificationutilize multiple spells on unemployment; recently eriksson2014unemploy discuss alternative experimental approaches to estimate duration. Then for our purpose of long-term causal inference, we would be interested in how the job training program would affect the long-term transition probability from unemployment to employment.

One might naively posit the long-term transition estimand as

equation*[equation* omitted — 70 chars of source]

and declare identification by calculating $E[ Y_2(1)| Y_1=0 , G=1]$ and $E[ Y_2(0)| Y_1 =0, G=1]$ in the treated($W=1$) and untreated($W=0$) population respectively, i.e., $E[ Y_2(1)| Y_1=0 , G=1] = E[ Y_2| Y_1=0 , W=1, G=1]$ and $E[ Y_2(0)| Y_1 =0, G=1]=E[ Y_2| Y_1 =0,W=0, G=1]$.

This approach, however overlooks dynamic selection. This is easiest to see in the best case scenario that $G=1$ is an experimental data, where $W \amalg (Y_2(1), Y_2(0), Y_1(1), Y_1(0)) |G=1 $holds. This comparability enables us to credibly estimate $E[ Y_2(1)|G=1] - E[ Y_2(0)|G=1]$ by simply estimating the $E[ Y_2(1)|G=1]$($E[ Y_2(0)|G=1]$)in the treated(untreated) population, i.e.,$E[Y_2(1)|G=1]- E[Y_2(0)|G=1]= E[Y_2|W=1,G=1] - E[Y_2| W=0, G=1]$. However, after we condition on the short term, the comparability of the treated and untreated popoulation is questionable. For example, if the job training program works reasonably well, the proportion that became employed in the treated population would likely be higher than that in the untreated population. The remaining population between the treated and untreated would be uncomparable in general, with different latent variables(like ability) that affect the likelihood of getting a job. In an observational setting, initial(static) selection compounds the dynamic selection, exacerbating the challenge in estimating dynamic treatment effects.

This observation itself is not new, and notably, the recent paper of vikstrom2018bounds have proposed a potential outcome formulation of dynamic transition probability that maintains comparability even in the long term, named The Average treatment on the treated survivors (ATETS), formally defined below.

definition[ATETS: Average treatment effect on the treated survivors] \[E[Y_2(1) | Y_1(1)=0, G=1 ] - E[ Y_2(0) | Y_1(1) =0, G=1]\]

The key point is that in $[E[Y_2(1) | Y_1(1)=0, G=1 ]] $ and $E[Y_2(0) | Y_1(1)=0, G=1 ] $, we condition on the same potential outcome $Y_1(1)$. If this quantity can be identified, it provides an interpretable treatment effect that expresses how likely the treatment will affect the long-term employment transition for the people who would not have succeeded in getting a job even if they received treatment. In fact, this treatment effect may be of particular interest in devising long-term oriented policies since if this treatment effect is estimated to be high, we should implement that treatment even if the short-term effect seems small. This would, for example, correspond to a job training program that focuses on human capital development, and the effect takes time to actually materialize in the labor outcome in the short term, but benefifical in the long term. In the famous example of California GAIN program, this would correspond to locations other than the Riverside district, which took a 'jobs first' policy; unemployers were encouraged to take any job they were offered.athey2019surrogate,hotz2005predicting.

However, we can easily see the challenge in point identifying this quantity. Especially, $E[ Y_2(0) | Y_1(1) =0, G=1]$ has two different counterfactuals in the outcome and conditioning, which makes identification seemingly impossible. Part II explores identification and estimation strategies for this challenging, but important treatment effect motivated by credible long-term causal inference free from dynamic confounding.

Nonparametric point-identification under no state-dependence

We first shortly discuss nonparametic point identification strategies. To start, we observe that the four assumptions, including the latent unconfounding assumption maintained in Part I cannot point identify the ATETS. A potential outcome variable $Y_1(1)$ in the conditioning set of $E[ Y_2(0) | Y_1(1) =0, G=1]$ cannot be identified for the observational due to initial selection. Then, is there any additional assumption that enables identification?

One possibility we note is the no-state dependence assumption introduced and elaborated in papers like heckman1981heterogeneity,heckman1984method,torgovitsky2019nonparametric. The issue in the literature centers around the interpretation of serial correlation between sequetional outcomes. Two competing interpretations are (i) there is time-varying unobservables that simultaneously affect the sequential outcome (unobserved heterogeneity), (ii) the fact that one is placed in a state of unemployment in the previous stage itself affects the possibility of unemployment this term, e.g., in the form of less opportunity for gaining social skills in unemployment (state dependence). In the latter case, where we posit no state dependence, the past potential outcome can be seen as a randomized treatment for the current potential outcome.

This assumption of no state-dependence, in potential outcomes, could be formulated as the following.

assumption[No state-dependence] The conditional independence \begin{equation} (Y_2(1), Y_2(0) ) \amalg (Y_1(1) , Y_1(0) ) | G=1 \end{equation} holds.\footnote{ To our knowledge, there are few papers that use potential outcome notations to formalize no state dependence, and arguably, the appropriate no state dependence should not be the joint independence across counterfactuals but for each treatment potential outcome, i.e., $ Y_2(w) \amalg Y_1(w) \, for w= 0,1 $. This is similar to the strong ignorability rosenbaum1983central and the weak ignorability imbens2015causal difference. In this paper, we will stick with the strong no state dependence assumption. }

Assumption (ref) states that, viewing $(Y_1(1),Y_1(0))$ as treatment, there is no direct (treatment) effect on the current outcome $(Y_2(1),Y_2(0))$. Given this assumption, we can show that combined with the assumptions in Section (ref), the ATETS reduces to the usual LTATE (long-term average treatment effect) because $E[ Y_2(1) | Y_1(1), G=1 ] = E[ Y_2(1) | G=1] $, and $ E[Y_2(0) | Y_1(0) , G=1] = E[Y_2(0) | G=1]$, and the identification strategy of athey2020combining can be used.

proposition[Nonparametric identification of ATETS under (strong) no state dependence ] Assume Assumptions (ref), (ref), (ref), (ref), and additionally Assumption (ref). Then ATETS is nonparametrically point-identified.

However, no state-dependence has known be to a strong assumption in many empirical research. Seminal works like heckman1981heterogeneity,heckman1984method,torgovitsky2019nonparametric have argued using economic theory and empirical evidence that, state dependence plays a non-trivial role, especially for unemployment. Are there other possible approaches to estimating the ATETS without imposing state dependence? In particular, we would need methods to posit some more structure on the dynamics of unemployment, and also on counterfactual states. This motivates dynamic discrete choice models on employment that incorporates behavioral structure of forward-looking optimizing agents in the form of structural parameters. We note that with this approach, identification is established only with the observational data. Still, we will show how the short-term experimental data can be useful even in this context. In Section (ref) (point identification), experimental data will be useful for model validation of structural models in line with the important work of todd2006assessing, and in Section (ref) (partial identification), it will help tighten the bounds. We first turn to point identification approach below.

Discussion and Conclusion

The above analysis concludes the theoretical framework for more reliable long-term causal inference combining data. Combining data was a very promising strategy for addressing the weaknesses of experimental data. However, the assumptions necessary for point identification were non-trivial, which motivated a deeper investigation on them.In Part I, we provided results for relaxing, understanding, and even testing for the exterrnal and internal validity assumptions, where the dynamic confounding turning out to be the source of invalidating the internal validity assumptions. In Part II, therefore, we discussed how we could credibly estimate duration paramters even under the presence of dynamic confoudnding. Here, nonparametric and structural approach to point identification had limitations. We therefore finally considered partial identification strategy that values transparency over the strength of the conclusion. In every step, we highlighted the complementary role of experiments and economic theory. These results provides a new approach to systematically tackling various new challenges that arise in our setting of interest.

Nevertheless, several important challenges still remain. For example, in a realistic scenario, the short-term outcome (e.g., children test scores) and long-term outcome of interest (e.g., labor market outcome) are qualitatively different, and inflated estimates result without considering general equilibrium effects, as in heckman1998general,blundell2003evaluating. One promising approach to this problem would be to take the mean-field approach that does not rely on strong assumptions on specific individuals decision process. Also, while we took a partial identification approach to increasing credibility, still, some may strongly prefer point estimates. In those cases, sensitivity analysis techniques specialized for long-term causal inference combining data would be a promising direction for future research.