EconBase
← Back to paper

Identification of Long-Term Treatment Effects via Temporal Links, Observational, and Experimental Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

109,253 characters · 14 sections · 126 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification of Long-Term Treatment Effects via Temporal Links, Observational, and Experimental Data

}

spacing{1.2} \begin{abstract} Recent literature proposes combining short-term experimental and long-term observational data to provide alternatives to conventional observational studies for the identification of long-term average treatment effects (LTEs). This paper re-examines the identification problem and uncovers that assumptions restricting temporal link functions -- relationships between short-term and mean long-term potential outcomes -- are central in this context. The experimental data serve to amplify the identifying power of such assumptions; absent them, the combined data are no more informative than the observational data alone. Plausible inference thus hinges on justifiable restrictions in this class. Motivated by this, I introduce two treatment response assumptions that may be defensible based on economic theory or intuition. To utilize them and facilitate future developments, I develop a novel unifying identification framework that computationally produces sharp bounds on the LTE for a general class of temporal link function restrictions and accommodates imperfect experimental compliance -- thereby also extending existing approaches. I illustrate the method by estimating the long-term effects of Head Start participation. The findings indicate that the effects on educational attainment, employment, and criminal involvement are lasting but smaller in magnitude than those established by sibling comparisons. \end{abstract}

Introduction

Identifying long-term average treatment effects (LTEs) is an important goal in economics and other fields of science. For example, researchers may be interested in the effects of childhood interventions on earnings in adulthood; the impact of early-life conditional cash transfers on employment prospects; or the long-run adverse or protective effects of vaccination. LTEs are also often of interest in private sector research gupta2019top.

Nevertheless, identifying LTEs is often challenging in practice. Long-term experimentation is frequently infeasible due to cost or institutional constraints.\footnote{Institutions supporting RCTs in development economics frequently require phase-in designs with staggered rollout of treatment to the whole sample. This limits follow-up for the control group.} Short-term experiments may be more accessible, but alone, they may not reveal outcomes of interest. As a result, researchers often rely on observational data currie2011human, hoynes2018safety. However, observational studies critically rely on identifying assumptions that may be challenging to justify.

This motivates a burgeoning strand of recent literature that seeks credible alternatives to conventional observational studies by combining (i) a long-run observational dataset with nonrandomized treatment assignment and (ii) a short-run experimental dataset in which long-run outcomes are unobserved athey2024surrogate,athey2025combining. However, follow-up work indicates that commonly-used modeling assumptions in this literature may also be challenging to justify in various economic settings, highlighting the need for alternative restrictions ghassami2022combining, van2023estimating, imbens2024long, park2024bracketing.

This paper re-examines the identification problem by first characterizing which assumptions can exploit the specific data structure. For identification of the LTE, the experimental data serve only to potentially amplify the identifying power of a well-defined class of modeling assumptions, restrictions on temporal link functions -- means of long-term potential outcomes conditional on short-term potential outcomes. Absent such restrictions, the combined data are necessarily equally informative about the LTE as the observational data alone. When restrictions on temporal link functions are imposed, however, combining data can provide additional identifying power. Therefore, plausible inference that leverages the experimental data hinges on imposing justifiable restrictions of this type.

Building on this characterization, I introduce novel modeling assumptions within this class that may be justified by economic theory or intuition. Both impose shape restrictions on temporal link functions without constraining the treatment selection in any of the datasets, and therefore constitute treatment response assumptions. The first stipulates that temporal link functions are monotonic. Intuitively, this means that, absent selection into treatment, average long-term outcomes would be monotonic in short-term ones. Monotonicity of conditional means has been widely used in other settings since it may often be defensible on economic or intuitive grounds (see e.g. manski2000monotone, manski2009identification, mogstad2018using, torgovitsky2019nonparametric). The second assumption postulates that the temporal link functions are invariant to treatment. Such invariance is implied by established mediation models, as in heckman2013understanding, garcia2020quantifying, or statistical surrogacy (prentice1989surrogate, athey2024surrogate).

Finally, the paper develops a unifying identification framework that accommodates: (i) a general class of restrictions on temporal link functions and (ii) imperfect compliance in the experiment. The framework delivers sharp bounds on the LTE computationally, by solving optimization problems in which the maintained restrictions enter as constraints. The general sharp characterization is obtained by extending arguments of beresteanu2012partial and chesher2017generalized to address a new technical challenge -- jointly bounding distribution functions and conditional means of latent variables.

The optimization-based formulation of the identified set enables direct implementation of the proposed restrictions. Moreover, it facilitates the development of new restrictions on temporal link functions by eliminating the need to derive bounds algebraically or to prove sharpness on a case-by-case basis. The framework also nests existing point-identification results and extends them to allow for imperfect compliance in the experiment, which is of substantial practical relevance. Building on shi2015simple and this characterization, I propose a consistent criterion-based estimator for the bounds.

I illustrate the method by estimating the long-term effects of Head Start participation, the largest federally funded early childhood education program in the United States. To do so, I combine data from the Head Start Impact Study, a short-term experiment, and the Child and Young Adult Supplement to the National Longitudinal Survey of Youth 1979 cohort, a longitudinal survey. I find evidence of beneficial program impacts on educational and labor-market outcomes, as well as criminal involvement in adulthood. Head Start is estimated to increase the probability of high school graduation by $1.9$ to $3.2$ percentage points (pp), and decrease the probability of grade repetition by $1.1$ to $5.3$ pp. The program is also estimated to lower the probability of idleness (neither working nor in school) by $1.5$ to $4.6$ pp and criminal involvement by $1.2$ to $4.0$ pp. The results suggest that Head Start has lasting effects, though smaller in magnitude than reported by sibling comparison studies (deming2009early). More broadly, they illustrate that the proposed assumptions can yield informative estimated bounds in applied work.

This paper is related to several strands of literature. It contributes to the recent body of work that combines long-term observational and short-term experimental data to identify long-term treatment effects (see also garcia2020quantifying, dynarski2021closing, hu2022identification, chen2023semiparametric, park2024informativeness, Aizer2024), and to the broader literature on data combination (cross2002regressions, molinari2006generalization, ridder2007econometrics, fan2014identifying, d2024partially). In doing so, it draws on and extends results from the partial identification literature based on random set theory (galichon2011set, beresteanu2012partial, molchanov2014applications, chesher2017generalized, chesher2020generalized). The paper also contributes to work on general identification frameworks that operationalize broad classes of assumptions via optimization (mogstad2018using, torgovitsky2019nonparametric, russell2021sharp, kamat2024identifying). Finally, it is related to research that combines experimental data with economic theory to identify parameters beyond what the experimental data alone reveal (todd2006assessing, attanasio2012education, todd2023best).

(ref) introduces the setting. (ref) details the role of restrictions on temporal link functions and experimental data. (ref) introduces new restrictions on temporal link functions. (ref) develops the identification framework. (ref) provides the empirical illustration. (ref) concludes. Appendix (ref) contains additional discussions; Appendix (ref) proves the main results; the Supplemental Appendix discusses estimation and proves auxiliary results.

Setting and Basic Assumptions

I formalize the problem using the standard potential outcomes model. Let $Y(d)\in\mathcal{Y}\subseteq\mathbb{R}$ and $S(d)\in{\mathcal{S}}\subseteq\mathbb{R}^{d_s}$ denote the long-term and short-term potential outcomes under some binary treatment $d\in\{0,1\}$, respectively. Denote the realized treatment by $D\in\{0,1\}$. The observed outcomes are:

align[align omitted — 108 chars of source]

Let $X\in\mathcal{X}\subseteq\mathbb{R}^{d_x}$ be a vector of observed covariates. Define the conditional long-term average treatment effect (CLTE) $\tau(x)$:

align[align omitted — 83 chars of source]

The parameter of interest can be the CLTE itself or its weighted averages, such as the average long-term treatment effect (LTE) $E[\tau(X)]$. I focus on the former for generality, noting that it is sufficient for identification of the latter when the weights are identified or given.

example{(Head Start Participation)} In the empirical illustration, $D$ denotes Head Start participation, $S(d)$ is a vector of potential cognitive test scores in childhood, and $Y(d)$ are potential outcomes in adulthood, such as high school degree status or earnings, for treatment $d$.

Observed Data

As in prior work, I maintain that the researcher observes: 1) a short-term experimental dataset; and 2) a long-term observational dataset.\footnote{This setting is increasingly common. See also garcia2020quantifying, ghassami2022combining, hu2022identification, van2023estimating, chen2023semiparametric, park2024bracketing, park2024informativeness, Aizer2024, imbens2024long, athey2025combining.} The population is partitioned into two subpopulations that are randomly sampled to generate the two datasets. Let $G\in\{O,E\}$ denote the subpopulation indicator, where $G=O$ produces the observational and $G=E$ the experimental data.

Let $Z\in\mathcal{Z}$ denote an exogenous (i.e., randomly assigned) instrument in the experimental dataset that induces individuals into treatment. The identification analysis accommodates bounded $\mathcal{Z}$ with an arbitrary number of support points. Typically, $Z\in\{0,1\}$, representing random assignment to treatment or control groups, which may differ from the realized treatment $D$. More generally, $Z$ may have more than two support points, or even satisfy $Z\in[0,1]$ as in heckman1999local. Since the binary case is predominant in practice, I adopt its terminology for expositional convenience. I refer to experiments with $P(D=Z|G=E)=1$ as having perfect compliance, and to the remaining cases as exhibiting imperfect compliance.

The short-term experimental dataset reveals $(S,D,X,Z)$, but not the long-term outcome $Y$. The long-term observational dataset reveals $(Y,S,D,X)$, but contains no exogenous instrument $Z$. Note that the data structure does not support the direct use of well-established instrumental variable methods for identification of long-term treatment effects, since the outcome of interest $Y$ is never observed together with an instrument $Z$ (e.g. as in imbens1994identification, heckman1999local and mogstad2018using).

continueexample{ex:head_start_basic} The observational dataset is the Child and Young Adult Supplement to the National Longitudinal Survey of Youth 79 Cohort (NLSY79) which reveals $(Y,S,D)$. The experimental dataset is the Head Start Impact Study (HSIS) which reveals $(S,D,Z)$, where $Z=1$ if the individual is assigned to participation in Head Start and $Z=0$ if assigned to non-participation. puma2010head explain that some individuals may have $D\neq Z$.
remarkIntroducing $Z$ in the experiment allows for (but does not require) imperfect compliance, which is of great practical relevance. In contrast, existing methods predominantly assume that $D$ is randomly assigned, which need not hold under imperfect compliance. The identification framework in (ref) nests these methods and extends them to allow this possibility.

I maintain the following assumptions throughout the paper.

myassump{RA}{(Random Assignment)} $Z\perp\mkern-9.5mu\perp (Y(d),S(d))|X,G=E$ for $d\in\{0,1\}$.
myassump{EV}{(Experimental External Validity)} $G\perp\mkern-9.5mu\perp (Y(d),S(d))|X$ for $d\in\{0,1\}$.

Assumption (ref) is standard in the program evaluation literature. It is satisfied if $Z$ in the experimental data is randomly assigned, conditional on $X$. Assumption (ref) is a key assumption in the data combination literature, linking the two datasets (ghassami2022combining, chen2023semiparametric, park2024bracketing, park2024informativeness, athey2025combining). It states that, conditional on $X$, the subpopulations generating the datasets do not differ in their counterfactual distributions for each $d$. It may be plausible if two datasets are representative of the same population, conditional on $X$. For example, in the empirical illustration, HSIS and NLSY79 are designed to be representative of the U.S. population and pertain to the same treatment. More broadly, a growing body of empirical work relies on related links across datasets in this context; see, for example, garcia2020quantifying, dynarski2021closing, hu2022identification, Aizer2024.

It is worth emphasizing what is not assumed. Maintained assumptions allow for $D\not\perp\mkern-9.5mu\perp (Y(d),S(d))|X,G=g$ for any $g\in\{O,E\}$ and $d\in\{0,1\}$. This is expected in the observational dataset, and in the experimental data when compliance is imperfect. The assumptions also do not require $P(D=1|G=g)\in(0,1)$ for any $g\in\{O,E\}$. Instead, they allow $P(D=1|G=g)\in[0,1]$, which is relevant when a treatment is available only in one dataset, typically in the experiment. This is the case with some “model” early childhood intervention programs or novel vaccines.

Under Assumption (ref), CLTE is invariant to $G$, $E[Y(1)-Y(0)|X=x,G] = E[Y(1)-Y(0)|X=x] = \tau(x)$. Henceforth, I keep conditioning on $X$ implicit. The following analysis should be understood as conditional-on-$X$; I write the parameter of interest $\tau(x)$ as:

align[align omitted — 36 chars of source]

and I continue referring to it as the LTE, with the understanding that it represents the CLTE.

remarkQuasi-experimental datasets with $Z$ satisfying Assumptions (ref) and (ref) may also serve as the experimental dataset. I continue to refer to such data as experimental, following prevailing terminology in the literature.

\noindentNotation: $\mathcal{H}(\theta)$ denotes the identified set for a parameter $\theta$. I write laws conditional on an event $\mathcal{E}$, $P(\cdot|\mathcal{E},G=g)$, as $P_g(\cdot|\mathcal{E})$ for $g\in\{O,E\}$. Whenever $P_E(\cdot|\mathcal{E}) = P_O(\cdot|\mathcal{E})$, I omit the subscript $g$. This is inherited by their features, e.g. $E_g[\cdot|\mathcal{E}] := E[\cdot|\mathcal{E},G=g]$. If necessary, I specify the random element using subscripts (e.g. $P_{S(d)}$ is the law of $S(d)$).

Role of Experimental Data

This section uncovers the role played by the experimental data in identifying $\tau$. Rather than removing the need for additional assumptions, experimental data serve to potentially amplify the identifying power of a well-defined class of modeling restrictions. (ref) leverages this insight to propose assumptions within this class that may be justifiable based on economic theory or intuition. (ref) develops a general identification framework that enables tractable implementation of various assumptions in the class. Define for $s\in\mathcal{S}$ and $d\in\{0,1\}$:

align[align omitted — 74 chars of source]

where $P_{S(d)}$ is the marginal law of $S(d)$. I refer to $m_0(s)$ and $m_1(s)$ as temporal link functions, since they “link” the short-term and long-term potential outcomes. Collect the two functions $m:=(m_0,m_1)\in\mathcal{M}$ and the marginal laws $\gamma := (\gamma_0,\gamma_1) = (P_{S(0)},P_{S(1)})$.\footnote{Let $(\Omega, \mathcal{F},P)$ denote the probability space and $\mathcal{B}$ the Borel $\sigma-$algebra. $\mathcal{M}$ is the set of Borel-measurable functions $\mu:\mathcal{S}\times\mathcal{S}\rightarrow \mathcal{Y}\times\mathcal{Y}$ such that $\mu\circ \varsigma$ is $P$-integrable for some $\mathcal{F}/\mathcal{B}(\mathcal{S}\times\mathcal{S})$-measurable function $\varsigma:\Omega\rightarrow\mathcal{S}\times\mathcal{S}$.} These objects are directly related to the parameter of interest:

align[align omitted — 176 chars of source]

Consider the class of modeling assumptions defined by the following generic restriction.

myassump{MA}{(Modeling Assumption)} $m\in\mathcal{M}^A\subseteq\mathcal{M}$ for a known or identified set $\mathcal{M}^A$.

Let $\mathcal{H}(\tau)$ be the identified set for $\tau$, that is, the set of all values of $\tau$ compatible with both datasets and Assumptions (ref), (ref), and (ref). Note that Assumption (ref) subsumes the trivial case $\mathcal{M}^A = \mathcal{M}$, corresponding to the absence of additional modeling assumptions. Denote by $\mathcal{H}^O(\tau)$ the identified set for $\tau$ under the same assumptions when only observational data are used, and by $\subsetneq$ a strict subset. The following proposition elucidates the relationship between $\mathcal{H}(\tau)$, $\mathcal{H}^O(\tau)$, and the modeling assumptions represented by $\mathcal{M}^A$.

theoremEnd[proof at the end, no link to proof]{proposition} Let Assumptions (ref), (ref) hold, and suppose $\mathcal{H}(\tau)\neq\emptyset$. If $\mathcal{H}(\tau) \subsetneq \mathcal{H}^O(\tau)$, then $\mathcal{M}^A\subsetneq\mathcal{M}$. Equivalently, if $\mathcal{M}^A=\mathcal{M}$, then $\mathcal{H}(\tau)= \mathcal{H}^O(\tau)$.
proofEndIf only Assumptions (ref), (ref) hold, by (ref) $i)$, $\mathcal{H}(P_{Y(0)},P_{Y(1)}) = \mathcal{H}^O(P_{Y(0)},P_{Y(1)})$. Then, $\mathcal{H}(\tau) = \mathcal{H}^O(\tau)$, since $\tau$ is a functional of $(P_{Y(0)},P_{Y(1)})$. Therefore, if $\mathcal{H}(\tau) \subsetneq \mathcal{H}^O(\tau)$ and (ref), (ref) hold, additional assumptions must be imposed. If $\mathcal{H}(\tau)\neq\emptyset$, these assumptions are not refuted. Henceforth, denote by $\mathcal{H}^A(\cdot)$ and $\mathcal{H}(\cdot)$ the identified sets with and without the additional assumptions. Let $\mathcal{H}^{A,O}(\cdot)$ be identified sets with additional assumptions, using only observational data. Given that $\mathcal{H}^A(\tau) $ and $\mathcal{H}^{O,A}(\tau) $ are images of $T$ over $\mathcal{H}^A(m,\gamma)$ and $\mathcal{H}^{O,A}(m,\gamma) $, respectively, the additional assumptions must directly or indirectly restrict $m$, $\gamma$ or $(m,\gamma)$. Otherwise, $\mathcal{H}^O(m,\gamma) = \mathcal{H}^{A,O}(m,\gamma) $ and $\mathcal{H}(m,\gamma) = \mathcal{H}^{A}(m,\gamma)$, so $\mathcal{H}^A(\tau) = \mathcal{H}(\tau) = \mathcal{H}^O(\tau) = \mathcal{H}^{A,O}(\tau)$ where the second equality is by (ref) $i)$. By way of contradiction, suppose that $\emptyset\neq \mathcal{H}^A(\tau)\subsetneq \mathcal{H}^{A,O}(\tau)$ and that only $\gamma$ is further restricted by the assumptions, or equivalently $\gamma\in\Gamma^A$, for some $\Gamma^A\subsetneq (\mathcal{P}^\mathcal{S})^2$. Then: \begin{align} \begin{split} \mathcal{H}^A(P_{Y(0)},P_{Y(1)},\gamma_0,\gamma_1) &=\mathcal{H}(P_{Y(0)},P_{Y(1)},\gamma_0,\gamma_1)\cap\left((\mathcal P^\mathcal{Y})^2\times\Gamma^A\right)\\ &=\mathcal{H}(P_{Y(0)},P_{Y(1)})\times\left(\mathcal{H}(\gamma_0,\gamma_1)\cap\Gamma^A\right) \end{split} \end{align} where the first equality is by the fact that only $\gamma$ is restricted by assumption, and the second is by (ref) $ii)$. It is then immediate that: \begin{align*} \mathcal{H}^A(P_{Y(0)},P_{Y(1)}) =\mathcal{H}(P_{Y(0)},P_{Y(1)}) = \mathcal{H}^O(P_{Y(0)},P_{Y(1)}) \end{align*} where the first equality is by (ref) and the definition of a projection, and the second is by (ref) $i)$. Thus, $\mathcal{H}^A(\tau) = \mathcal{H}^O(\tau)$. But since $\mathcal{H}^A(\tau)\subseteq\mathcal{H}^{A,O}(\tau)\subseteq\mathcal{H}^O(\tau)$, then it must also be that $\mathcal{H}^A(\tau)=\mathcal{H}^{A,O}(\tau) $, contradicting $\mathcal{H}^A(\tau)\subsetneq \mathcal{H}^{A,O}(\tau)$. Therefore, the additional assumptions must restrict $m$ or $(m,\gamma)$. In either case, $m\in\mathcal{M}^A\subsetneq\mathcal{M}$.

(ref) shows that it is necessary to restrict $m$ by assumption in order to gain identifying power from the experimental data. Equivalently, when the temporal link functions $m$ are left unrestricted, incorporating the experimental data provides no additional identifying power for $\tau$. Once such restrictions are imposed, the inclusion of experimental data may yield additional identifying power, enabling $\mathcal{H}(\tau) \subsetneq \mathcal{H}^O(\tau)$. The proposition thus isolates the class of assumptions that could be made more informative about $\tau$ by the addition of experimental data. In this sense, adding experimental data may amplify the identifying power of restrictions on $m$.

To gain intuition for the result, note that because $S(d)$ is revealed in the observational data whenever $D = d$, the experimental data can only provide additional information about the distribution of $S(d)$ for those individuals with $D \neq d$. For these individuals, however, the data impose no restrictions on the relationship between $S(d)$ and the mean $Y(d)$. When this relationship is left unrestricted, additional information about the distribution of $S(d)$ does not translate into additional information about the average $Y(d)$, and thus about $\tau$. In contrast, when the relationship between $S(d)$ and the average $Y(d)$ is restricted by assumption, additional information about the distribution of $S(d)$ can yield further information about the mean of $Y(d)$ and hence about $\tau$.

remarkExisting approaches typically do not impose explicit restrictions on $m$. However, their identification arguments ultimately reduce to such restrictions implied by the maintained assumptions. These point-identification results are thus nested within the identification framework here. For example, athey2025combining maintain latent unconfoundedness (LUC) -- $Y(d)\perp\mkern-9.5mu\perp D|S(d),G=O$ for $d\in\{0,1\}$, while the identification result uses its implication $m_d(s) = E_O[Y|S=s,D=d]$ for $d\in\{0,1\}$ and $s\in\mathcal{S}$. For imbens2024long, let $S_t$ be subvectors of $S$ for $t\in\{1,2,3\}$. Maintained assumptions imply the restriction used for identification: $m_d(s_3,s_2) = h(s_3,s_2,d)$ where $h$ solves $E_O[Y | S_2, S_1, D] = E_O[h(S_3, S_2, D) | S_2, S_1, D]$.

Two points are worth emphasizing. First, the addition of experimental data need not necessarily amplify the identifying power of restrictions on $m$. In particular, it is possible that $\mathcal{M}^A \subsetneq \mathcal{M}$ while $\mathcal{H}(\tau) = \mathcal{H}^O(\tau)$. Second, restrictions on $m$ may have identifying power for $\tau$ even in the absence of experimental data. Whether either case arises depends on the underlying data distribution and on the specific restriction imposed. These observations highlight that restrictions on $m$ are central for identifying $\tau$ in this setting, while the experimental data play an auxiliary role by potentially amplifying their identifying power.

(ref) in Appendix (ref) illustrates these points using the latent unconfoundedness assumption defined in (ref), and frequently invoked in this context (see, for example, hu2022identification, park2024informativeness, and Aizer2024). Specifically, the proposition shows that if the data distribution is such that $m_d(s) = E_O[Y \mid S = s, D = d]$ are constant in $s$, both the combined and observational data alone point identify $\tau$ under LUC. In this case, $\mathcal{M}^A \subsetneq \mathcal{M}$ but $\mathcal{H}(\tau) = \mathcal{H}^O(\tau)$, so the experimental data do not amplify the identifying power of LUC. The proposition further shows that for a broad class of empirically relevant data distributions, the observational-data bound $\mathcal{H}^O(\tau)$ under LUC is strictly more informative than the worst-case bounds of manski1990nonparametric that would be sharp absent LUC.

Modeling Assumptions

For the purpose of identifying $\tau$, the experimental data serve only to potentially amplify the identifying power of restrictions imposed on $m$. Consequently, plausible inference hinges on the plausibility of the restrictions within this class, despite the use of experimental data. However, commonly used assumptions implying restrictions on $m$ may be challenging to interpret or justify in economically relevant settings (see e.g. ghassami2022combining, van2023estimating, imbens2024long and park2024bracketing). Hence, I explore alternative restrictions within this class that may be defensible based on economic theory or intuition.

myassump{LIV}{(Latent Monotone Instrumental Variables)} For every $m\in\mathcal M^A$ and each $d\in\{0,1\}$, $m_d(s)$ is nondecreasing in $s$ with respect to the product order. That is, for all $s,s'\in\mathcal S$ such that $s\leq s'$ (componentwise): $ m_d(s)\leq m_d(s')$.

Assumption (ref) states that the conditional means of $Y(d)$ are nondecreasing in any individual short-term potential outcome $S(d)$.\footnote{One can also assume that $m_d(s)$ is nonincreasing for any sub-collection of elements $\{s_j\}_{j\in\mathcal{J}}$ of $s$. Results follow directly by defining $\Tilde{S}(d)$ with components $\Tilde{S}_j(d) = \mathbbm{1}[j\in\mathcal J]\left(- S_j(d)\right)+\mathbbm{1}[j\not\in\mathcal J] S_j(d)$ and observing that $E[Y(d)|\Tilde{S}(d)=s]$ satisfies LIV.} Intuitively, this would be plausible if researchers are willing to maintain that long-term outcomes would be monotonic in short-term ones, absent selection into treatment.

\setcounter{example}{1}

example{(LIV and Head Start)} Suppose first that $S(d)$ consists of a single childhood cognitive test score. Assumption (ref) then states that people with a higher potential test score $S(d)$ have a weakly higher average of potential high school degree status in adulthood $Y(d)$ than people with a lower potential test score. In other words, under an exogenously fixed treatment, high school graduation rates are nondecreasing in the childhood test score. When $S(d)$ comprises multiple test scores, the assumption asserts that the high school graduation rates are nondecreasing in each individual test score, holding the remaining ones constant.

LIV is related to the monotone instrumental variable (MIV) assumption of manski2000monotone, manski2009more. MIV posits the existence of a variable $V\in\mathcal{V}$ such that $E[Y(d)|V=v]$ is nondecreasing in $v\in\mathcal{V}$, where $V$ is observed for all individuals. The critical distinction is that the conditioning variables in Assumption (ref) are latent counterfactuals. This feature introduces additional complexity, which is addressed by the identification framework.

myassump{TI}{(Treatment Invariance)} For all $m\in\mathcal{M}^A$ and $s\in\mathcal{S}$, $m_1(s) = m_0(s)$.

The assumption states that the relationship between the potential short-term outcomes $S(d)$ and the mean of the long-term potential outcomes $Y(d)$ does not vary with the underlying treatment $d$. Intuitively, it means that the treatment affects average long-term outcomes only through short-term ones. Importantly, TI follows from theoretical models previously used in empirical research, as outlined by the following example. It is also implied by the statistical surrogacy assumption of prentice1989surrogate $Y\perp\mkern-9.5mu\perp D|S,G=E$ in the special case of perfect compliance.\footnote{See (ref) i) in Appendix (ref). An important distinction relative to work relying on the surrogacy assumption in this literature is that the comparability assumption $G\perp\mkern-9.5mu\perp Y|S$ is not imposed here (e.g. see athey2024surrogate and chen2023semiparametric). Together, comparability and surrogacy imply a selection assumption on $m$ under Assumption (ref); see (ref) vi).}

example{(TI and Head Start)} Consider the following separable model: \begin{align} Y(d) = \phi_d(S(d))+\varepsilon_d = \phi(S(d))+\varepsilon_d, \enskip \varepsilon_d\sim\varepsilon, \enskip\varepsilon_{d'}\perp\mkern-9.5mu\perp S(d), \forall d,d'\in\{0,1\} \end{align} where $S(d)$ is a vector of short-term potential outcomes, including test scores and measures of non-cognitive skills. $S(d)$ are inputs in the production function $\phi_d$ for $Y(d)$. The production function $\phi_d$ and the distributions of unobservables $\varepsilon_d$ do not depend on Head Start participation $d$. Therefore, $E[Y(d)|S(d)=s] = \phi(s)+E[\varepsilon]$ which is invariant to $d$, so TI is implied by the model. Researchers may thus use TI whenever they find the model plausible. garcia2020quantifying argue the plausibility of such a model in the context of identifying the long-term effects of an early childhood program.

Note that neither of the two assumptions imposes restrictions on selection, i.e., how individuals choose their treatment. As such, they represent treatment response assumptions, whereas related preceding work predominantly relies on selection assumptions.

Identification Framework

The preceding sections argue that restrictions on temporal link functions $m$ are central, and that experimental data can make them more informative. Here, I introduce a novel framework that enables the identification of $\tau$ under a general class of restrictions on $m$.

Define the functional $T:\mathcal{M}\times\mathcal{P}^\mathcal{S}\times\mathcal{P}^\mathcal{S}\rightarrow\Bar{\mathbb{R}}$, where $\mathcal{P}^\mathcal{S}$ is the set of distribution functions supported on $\mathcal{S}$:

align[align omitted — 147 chars of source]

Recalling that $E[Y(d)] = \int_\mathcal{S} m_d(s) d\gamma_{d}(s)$ by (ref), $T$ produces the corresponding value of $\tau$ for a given $(m,\gamma)$. One can then incorporate assumptions on $m$ to identify $\tau$ in two steps. First, find all $(m,\gamma)$ consistent with the data, Assumptions (ref), (ref), and any restriction (ref) on $m$. This yields their identified set $\mathcal{H}(m,\gamma)$. Then, the identified set $\mathcal{H}(\tau)$ is the set of all values of $\tau$ consistent with the feasible $(m,\gamma)$, or:

align[align omitted — 187 chars of source]

Identifying $(m,\gamma)$

This section shows how to represent all information about $(m,\gamma)$. $(m,\gamma)$ are features of latent random variables $Y(d)$ and $S(d)$. I exploit this fact to construct $\mathcal{H}(m,\gamma)$ for a large class of restrictions on $m$.

Because at least some potential outcomes are unobserved for each unit, the data and maintained assumptions are consistent with a set of random vectors $(S(0),S(1),Y(0),Y(1))$. Intuitively, let $\mathcal{Q}$ be the set of all such random vectors that are compatible with the data, and Assumptions (ref) and (ref). To obtain $\mathcal{H}(m,\gamma) $, it suffices to collect the corresponding pairs $(m,\gamma)$ that additionally satisfy the modeling restriction $m\in\mathcal{M}^A$. By definition:

align[align omitted — 512 chars of source]

Random set theory provides a convenient way to deliver a sharp characterization of $(m,\gamma)$ based on this intuition. It first allows one to characterize $\mathcal{Q}$ and, consequently, $\mathcal{H}(m,\gamma)$. It then yields an equivalent representation of (ref) via moment inequalities that exhaust all information contained in the data and the maintained assumptions. To formalize the argument, I introduce the necessary basic definitions specialized to finite-dimensional Euclidean spaces. I henceforth maintain that all random elements are defined on a non-atomic probability space $(\Omega, \mathcal{F},P)$.\footnote{That is, for any $A\in\mathcal{F}$ with positive measure there exists a measurable $B\subset A$ such that $0<P(B)<P(A)$.}

\noindentNotation: $B$ and $K$ denote sets. $\mathcal{K}(B)$, and $\mathcal{C}(B)$ denote the families of all compact, and closed subsets of the set $B$, respectively. $co(B)$ is the closed convex hull of the set $B$.

definitionA measurable map $\mathbf{R}:\Omega\rightarrow \mathcal{C}(\mathbb{R}^{d_R})$ is called a random (closed) set.\footnote{$\mathbf{R}$ is measurable if for every compact set $K\in\mathcal{K}(\mathbb{R}^{d_R})$: $\{\omega\in\Omega: \mathbf{R}(\omega)\cap K\neq \emptyset\}\in\mathcal{F}$. The codomain $\mathcal{C}(\mathbb{R}^{d_R})$ is equipped with the $\sigma-$algebra generated by the families of sets $\{B\in \mathcal{C}(\mathbb{R}^{d_R}):B\cap K\neq \emptyset\}$ over $K\in\mathcal{K}(\mathbb{R}^{d_R})$.}
definitionA random vector $R:\Omega\rightarrow \mathbb{R}^{d_R}$ such that $R\in\mathbf{R}$ a.s. is called a (measurable) selection of $\mathbf{R}$. $Sel(\mathbf{R})$ and $Sel^1(\mathbf{R})$ are the sets of all selections, and all integrable selections of $\mathbf{R}$, respectively.

Define the following closed random sets for $d\in\{0,1\}$:

align[align omitted — 331 chars of source]

By construction, random sets $\mathbf{Y}_d $ and $\mathbf{S}_d $ summarize all information about $Y(d)$ and $S(d)$ contained in the data, respectively. As beresteanu2012partial explain, all information in the data about $(S(0),S(1),Y(0),Y(1))$ can be represented by stating that they can be any selection of the corresponding random sets. Additional assumptions, such as Assumptions (ref) and (ref), may be imposed by further restricting the admissible selections, which are represented by the set $\mathcal{Q}$ in (ref). The following lemma formalizes $\mathcal{Q}$ and uses it to equivalently characterize $\mathcal{H}(m,\gamma)$.

theoremEnd[proof at the end, no link to proof]{lemma} Let Assumptions (ref), (ref), and (ref) hold. The identified set for $(m,\gamma)$ is: \begin{align} \begin{split} \mathcal{H}(m,\gamma) = \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}^A\times (\mathcal{P}^{\mathcal{S}})^2: \forall d\in\{0,1\}, \enskip \exists S(d)\in Sel(\mathbf{S}_d)\cap I,\\ \exists Y(d)\in Sel(\mathbf{Y}_d), \enskip \gamma_d \overset{d}{=} S(d),\enskip m_d(S(d)) = E_O[Y(d)|S(d)] a.s. \end{array}\right\}. \end{split} \end{align} where $I$ is the set of random elements $E_1\in\mathcal{S}$ such that $E_1\perp\mkern-9.5mu\perp G$ and $E_1\perp\mkern-9.5mu\perp Z|G=E$.
proofEndRecall by (ref) that $ \Tilde{Z} = \mathbbm{1}[G=E] Z+\mathbbm{1}[G=O](\sup\mathcal{Z}+1)\in \Tilde{\mathcal{Z}}$. Note that $\Tilde{Z} =Z$ when $G=E$ and $\Tilde{Z}$ equals a distinct constant when $G=O$. Therefore, Assumptions (ref) and (ref) hold if and only if $\Tilde{Z}\perp\mkern-9.5mu\perp (Y(d),S(d))$ for all $d\in\{0,1\}$. Let $\Tilde{I}$ be the set of random elements $(E_1,E_2,E_3)$ such that $(E_1,E_2,E_3)\in \mathcal{Y}\times\mathcal{S}\times\Tilde{\mathcal{Z}}$ and $(E_1,E_2)\perp\mkern-9.5mu\perp E_3$, i.e. that satisfy the two assumptions. Define the random set: \begin{align} (\mathbf{Y}_0,\mathbf{S}_0,\mathbf{Y}_1,\mathbf{S}_1) :=\begin{cases} \{(Y,S)\}\times\mathcal{Y}\times\mathcal{S}, if $(D,G)=(0,O)$\\ \mathcal{Y}\times\mathcal{S}\times\{(Y,S)\}, if $(D,G)=(1,O)$\\ \mathcal{Y}\times\{S\}\times\mathcal{Y}\times\mathcal{S}, if $(D,G)=(0,E)$\\ \mathcal{Y}\times\mathcal{S}\times\mathcal{Y}\times\{S\}, if $(D,G)=(1,E)$\\ \mathcal{Y}\times\mathcal{S}\times\mathcal{Y}\times\mathcal{S}, otherwise \end{cases}. \end{align} As in the proof of beresteanu2012partial, by definition of $(\mathbf{Y}_0,\mathbf{S}_0,\mathbf{Y}_1,\mathbf{S}_1)$, all information on $(Y(0),S(0),Y(1),S(1))$ in the observed data can be summarized by $(Y(0),S(0),Y(1),S(1))\in Sel((\mathbf{Y}_0,\mathbf{S}_0,\mathbf{Y}_1,\mathbf{S}_1))$. Also, by definition of the random set, $(\mathbf{Y}_0,\mathbf{S}_0,\mathbf{Y}_1,\mathbf{S}_1) =(\mathbf{Y}_0,\mathbf{S}_0)\times(\mathbf{Y}_1,\mathbf{S}_1)$, where for $d\in\{0,1\}$: \begin{align} (\mathbf{Y}_d,\mathbf{S}_d) =\begin{cases} \{(Y,S)\}, if $(D,G)=(d,O)$\\ \mathcal{Y}\times\{S\}, \text{ if $(D,G)=(d,E)$}\\ \mathcal{Y}\times\mathcal{S}, \text{ otherwise} \end{cases}. \end{align} By applying (ref) twice, $Sel((\mathbf{Y}_0,\mathbf{S}_0,\mathbf{Y}_1,\mathbf{S}_1)) =Sel((\mathbf{Y}_0,\mathbf{S}_0))\times Sel((\mathbf{Y}_1,\mathbf{S}_1))=Sel(\mathbf{Y}_0)\times Sel(\mathbf{S}_0)\times Sel(\mathbf{Y}_1)\times Sel(\mathbf{S}_1)$. All information about $(Y(0),S(0),Y(1),S(1))$ in the data, Assumptions (ref) (ref) and $E[|Y(d)|]<\infty$ for $d\in\{0,1\}$ can thus equivalently be expressed as $(Y(d),S(d),\Tilde{Z})\in Sel(\mathbf{Y}_d)\times Sel (\mathbf{S}_d)\times\{\Tilde{Z}\}\cap \Tilde{I}$ for $d\in\{0,1\}$. If Assumptions (ref) and (ref) hold, the intersection is non-empty. Let $\Bar{I}$ be the set of random elements $(E_1,E_2)$ such that $(E_1,E_2)\in \mathcal{S}\times\Tilde{\mathcal{Z}}$ and $E_1\perp\mkern-9.5mu\perp E_2$. The set of $(m,\gamma)$ consistent with the data and the two assumptions follows by definition as: \begin{align} \begin{split} &\mathcal{H}^{EV/RA}(m,\gamma) \\ &:= \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}\times \left(\mathcal{P}^{\mathcal{S}}\right)^2: \forall d\in\{0,1\}, \enskip \exists(\upsilon_d,\varsigma_d,\Tilde{Z})\in Sel(\mathbf{Y}_d)\times Sel (\mathbf{S}_d)\times\{\Tilde{Z}\}\cap \Tilde{I},\\ \gamma_d \overset{d}{=} \varsigma_d, \enskip m_d(\varsigma_d) = E[\upsilon_d|\varsigma_d] \text{ a.s.} \end{array} \right\}\\ &= \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}\times \left(\mathcal{P}^{\mathcal{S}}\right)^2: \forall d\in\{0,1\}, \enskip \exists(\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d, \Tilde{Z}))\cap \Bar{I},\\ \exists\upsilon_d\in Sel(\mathbf{Y}_d),\enskip (\upsilon_d,\varsigma_d)\perp\mkern-9.5mu\perp \Tilde{Z}, \enskip \gamma_d \overset{d}{=} \varsigma_d, \enskip m_d(\varsigma_d) = E[\upsilon_d|\varsigma_d] \text{ a.s.} \end{array}\right\}\\ &= \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}\times \left(\mathcal{P}^{\mathcal{S}}\right)^2: \forall d\in\{0,1\}, \enskip \exists(\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d, \Tilde{Z}))\cap \Bar{I},\\ \exists\upsilon_d\in Sel(\mathbf{Y}_d),\enskip (\upsilon_d,\varsigma_d)\perp\mkern-9.5mu\perp \Tilde{Z}, \enskip \gamma_d \overset{d}{=} \varsigma_d, \enskip m_d(\varsigma_d) = E_O[\upsilon_d|\varsigma_d] \text{ a.s.} \end{array}\right\}\\ &= \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}\times \left(\mathcal{P}^{\mathcal{S}}\right)^2: \forall d\in\{0,1\}, \enskip \exists(\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d, \Tilde{Z}))\cap \Bar{I},\\ \exists\upsilon_d\in Sel(\mathbf{Y}_d), \enskip \gamma_d \overset{d}{=} \varsigma_d, \enskip m_d(\varsigma_d) = E_O[\upsilon_d|\varsigma_d] \text{ a.s.} \end{array}\right\}\\ \end{split} \end{align} where the second equality is by (ref), and the third is by $(\upsilon_d,\varsigma_d)\perp\mkern-9.5mu\perp \Tilde{Z}$ and the fourth is by (ref). It only remains to impose Assumption (ref). Then, the identified set is: \begin{align} \begin{split} \mathcal{H}(m,\gamma) &= \mathcal{H}^{EV/RA}(m,\gamma)\cap (\mathcal{M}^A\times (\mathcal{P}^{\mathcal{S}})^2)\\ &= \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}^A\times (\mathcal{P}^{\mathcal{S}})^2:\enskip \forall d\in\{0,1\}, \enskip \exists(\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d, \Tilde{Z}))\cap \Bar{I},\\ \exists\upsilon_d\in Sel(\mathbf{Y}_d),\enskip\gamma_d \overset{d}{=} \varsigma_d, \enskip m_d(\varsigma_d) = E_O[\upsilon_d|\varsigma_d] \text{ a.s.} \end{array}\right\}\\ &= \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}^A\times (\mathcal{P}^{\mathcal{S}})^2:\enskip \forall d\in\{0,1\}, \enskip \exists\varsigma_d\in Sel(\mathbf{S}_d)\cap I,\\ \exists\upsilon_d\in Sel(\mathbf{Y}_d),\enskip\gamma_d \overset{d}{=} \varsigma_d, \enskip m_d(\varsigma_d) = E_O[\upsilon_d|\varsigma_d] \text{ a.s.} \end{array}\right\}. \end{split} \end{align} where the first equality is by observation and the second is since $\varsigma_d\in Sel(\mathbf{S}_d)\cap I$ can be equivalently stated as $(\varsigma_d,\Tilde{Z})\in Sel(\mathbf{S}_d,\Tilde{Z})\cap \Bar{I}$. Next note that for every $(m,\gamma)\in \mathcal{H}(m,\gamma)$, there exist $(\upsilon_0,\varsigma_0,\upsilon_1,\varsigma_1)$ that generate them and that are consistent with the data, modeling assumption, Assumptions (ref) and (ref). Therefore, $\mathcal{H}(m,\gamma)$ is sharp.

The lemma exploits the fact that $Y$ is observed only in the observational sample to eliminate restrictions on $(m,\gamma)$ imposed by Assumptions (ref) and (ref) that are redundant in this setting. In particular, the independence restrictions that remain relevant are $S(d)\perp\mkern-9.5mu\perp G$ and $S(d)\perp\mkern-9.5mu\perp Z|G=E$. Moreover, only observational data impose restrictions on $m$ directly, as reflected by $m_d(S(d)) = E_O[Y(d)|S(d)]$. These observations enable the following sharp characterization of $\mathcal{H}(m,\gamma)$ via moment inequalities identified by the data.

theoremEnd[proof at the end, no link to proof]{theorem} Let Assumptions (ref), (ref) and (ref) hold. Suppose that $E[|Y(d)|]<\infty$ for $d\in\{0,1\}$. The identified set for $(m,\gamma)$ is: \begin{align} \begin{split} \mathcal{H}(m,\gamma) = \left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times(\mathcal{P}^{\mathcal{S}})^2: $\forall d\in\{0,1\}$, $\forall B\in\mathcal{C}(\mathcal{S})$, \\ $\gamma_d(B)\geq \max\left(ess\sup_{Z}P_E(S\in B,D=d|Z),P_O(S\in B,D=d) \right)$,\\ $\forall u\in\{-1,1\}$: $um_d(s)\leq u\mu_d(s)\pi_{\gamma_d}(s)+ h_{co(\mathcal{Y})}(u)(1-\pi_{\gamma_d}(s))$ $ \gamma_d-$a.e. \end{array}\right\} \end{split} \end{align} where $h_{co(\mathcal{Y})}(u) := \sup_{y\in co(\mathcal{Y})} u y$, $\mu_d(s) := E_O[Y|S=s,D=d]$, and $\pi_{\gamma_d} := d P_O(S,D=d)/d\gamma_d$. If a collection of sets $\mathfrak{C}$ is a core determining class for the containment functional of $\mathbf{S}_d$, then the condition $\forall B\in\mathcal{C}(\mathcal{S})$ can be replaced with $\forall B\in\mathfrak{C}$.
proofEndRecall that $\Bar{I}$ is the set of random elements $(E_1,E_2)$ such that $(E_1,E_2)\in \mathcal{S}\times\Tilde{\mathcal{Z}}$ and $E_1\perp\mkern-9.5mu\perp E_2$. Likewise $I$ is defined by (ref) and $\Tilde{Z}$ by (ref). The proof then proceeds through a series of steps. I first show that $\mathcal{H}(m,\gamma) = \Tilde{\mathcal{H}}(m,\gamma)$ where: \begin{align} \begin{split} &\Tilde{\mathcal{H}}(m,\gamma) \\ &:=\left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times (\mathcal{P}^{\mathcal{S}})^2: \forall d\in\{0,1\}, $\exists (\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d,\Tilde{Z}))\cap \Bar{I}$, $ \gamma_d\overset{d}{=} \varsigma_d$, \\ $\forall u\in\{-1,1\}:\enskip u m_d(s)\leq u\mu_d(s)\pi_{\gamma_d}(s)+ h_{co(\mathcal{Y})}(u)(1-\pi_{\gamma_d}(s)),\text{ $\gamma_d-$a.e.}$ \end{array}\right\} \end{split} \end{align} I then show that $\Tilde{\mathcal{H}}(m,\gamma)$ is equivalent to (ref). Observing that $\varsigma_d\in Sel(\mathbf{S}_d)\cap I$ can be equivalently stated as $(\varsigma_d,\Tilde{Z})\in Sel(\mathbf{S}_d,\Tilde{Z})\cap \Bar{I}$, by (ref): \begin{align} \mathcal{H}(m,\gamma) &= \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}^A\times (\mathcal{P}^{\mathcal{S}})^2: \forall d\in\{0,1\}, \enskip \exists(\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d,\Tilde{Z}))\cap \Bar{I},\\ \exists\upsilon_d\in Sel(\mathbf{Y}_d), \enskip \gamma_d \overset{d}{=} \varsigma_d,\enskip m_d(\varsigma_d) = E_O[\upsilon_d|\varsigma_d] a.s. \end{array}\right\} \end{align} noting that $\mathcal M^A$ is such that for any $(m,\gamma)$ under consideration, each $m_d$ must be $\gamma_d$-integrable since $E[|Y(d)|]<\infty$ implies $\int |m_d|d\gamma_d = E[|E[Y(d)|S(d)]|]\leq E[|Y(d)|]<\infty$. It is immediate that then also: \begin{align} \mathcal{H}(m,\gamma) &= \left\{\begin{array}{ll} (m,\gamma)\in\mathcal{M}^A\times (\mathcal{P}^{\mathcal{S}})^2: \forall d\in\{0,1\}, \enskip \exists(\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d,\Tilde{Z}))\cap \Bar{I},\\ \exists\upsilon_d\in Sel^1(\mathbf{Y}_d), \enskip \gamma_d \overset{d}{=} \varsigma_d,\enskip m_d(\varsigma_d) = E_O[\upsilon_d|\varsigma_d] a.s. \end{array}\right\}. \end{align} $\mathcal{H}(m,\gamma) = \Tilde{\mathcal{H}}(m,\gamma)$ To show $\mathcal{H}(m,\gamma)\subseteq \Tilde{\mathcal{H}}(m,\gamma)$, fix $(m,\gamma)\in\mathcal{H}(m,\gamma)$ and $d\in\{0,1\}$. Then there exist $(\varsigma_d,\Tilde Z)\in Sel((\mathbf S_d,\Tilde Z))\cap \Bar I$ and $\upsilon_d\in Sel^1(\mathbf Y_d)$ such that $\gamma_d\overset{d}{=} \varsigma_d$ and $m_d(\varsigma_d)=E_O[\upsilon_d\mid \varsigma_d]$ a.s. Let $\sigma(\varsigma_d|G=O)$ be the sub-$\sigma$-algebra generated by $\varsigma_d$ given $\{\omega\in\Omega:G=O\}$. Let $\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]:= cl\{E_O[\upsilon_d|\varsigma_d]: \upsilon_d\in Sel^1(\mathbf{Y}_d)\}$, where the closure is taken in $L^1$ space of all $\sigma(\varsigma_d|G=O)$-measurable functions. $\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]$ exists, is a unique random set, and has at least one integrable selection (molchanov2017theory). By definition of $\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]$, it is immediate that: \begin{align} \begin{split} \exists \upsilon_d\in Sel^1(\mathbf{Y}_d): \enskip m_d(\varsigma_d)=E_O[\upsilon_d|\varsigma_d] a.s. &\Rightarrow m_d(\varsigma_d)\in \mathbb{E}_O[\mathbf{Y}_d|\varsigma_d] \text{ a.s.} \end{split} \end{align} By assumption, the probability space is non-atomic. By (ref), $P$ has no atoms over $\sigma(\varsigma_d|G=O)$ for any measurable selection $\varsigma_d$. Since $E[|Y(d)|]<\infty$ for all $d\in\{0,1\}$, $\mathbf{Y}_d$ is integrable. Thus, $\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]$ is almost surely convex and equal to $\mathbb{E}_O[co(\mathbf{Y}_d)|\varsigma_d]$ (molchanov2017theory). Therefore, $h_{\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]}(u) = h_{\mathbb{E}_O[co(\mathbf{Y}_d)|\varsigma_d]}(u)$ a.s. for all $u\in\mathbb{R}$ by definition of the support function $h$. By $\mathbb{E}_O[co(\mathbf{Y}_d)|\varsigma_d]=\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]$ and integrability of the latter, the former set is also integrable. It then follows that $h_{\mathbb{E}_O[co(\mathbf{Y}_d)|\varsigma_d]}(u) = E_O[h_{co(\mathbf{Y}_d)}(u)|\varsigma_d]$ a.s. for all $u\in\mathbb{R}$ (molchanov2017theory). Hence, recalling that $\mathbb{E}_O[co(\mathbf{Y}_d)|\varsigma_d]=\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]$, also $h_{\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]}(u) =h_{\mathbb{E}_O[co(\mathbf{Y}_d)|\varsigma_d]}(u)= E_O[h_{co(\mathbf{Y}_d)}(u)|\varsigma_d]$ a.s. for all $u\in\mathbb{R}$. Then: \begin{align} \begin{split} m_d(\varsigma_d)\in\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d] \text{ a.s.} \Leftrightarrow&\forall u\in\{-1,1\}:\enskip u m_d(\varsigma_d)\leq h_{\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]}(u) \text{ a.s.}\\ \Leftrightarrow&\forall u\in\{-1,1\}:\enskip u m_d(\varsigma_d)\leq E_O[h_{co(\mathbf{Y}_d)}(u)|\varsigma_d] \text{ a.s.} \end{split} \end{align} where the first line is by rockafellar1970convex and almost sure convexity of $\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]$, and the second is by $h_{\mathbb{E}_O[\mathbf{Y}_d|\varsigma_d]}(u) = E_O[h_{co(\mathbf{Y}_d)}(u)|\varsigma_d]$ a.s. for all $u\in\mathbb{R}$. Moreover: \begin{align} \begin{split} E_O[h_{co(\mathbf{Y}_d)}(u)|\varsigma_d] &= E_O[h_{co(\mathbf{Y}_d)}(u)|\varsigma_d,D=d]P_O(D=d|\varsigma_d)\\ &+ E_O[h_{co(\mathbf{Y}_d)}(u)|\varsigma_d,D\neq d]P_O(D\neq d|\varsigma_d)\\ &= uE_O[Y|\varsigma_d,D=d]P_O(D=d|\varsigma_d)+ h_{co(\mathcal{Y})}(u)P_O(D\neq d|\varsigma_d) \\ &= uE_O[Y|S,D=d]P_O(D=d|\varsigma_d)+ h_{co(\mathcal{Y})}(u)P_O(D\neq d|\varsigma_d) \\ &= u\mu_d(\varsigma_d)P_O(D=d|\varsigma_d)+ h_{co(\mathcal{Y})}(u)P_O(D\neq d|\varsigma_d)\\ &= u\mu_d(\varsigma_d)\pi_{\gamma_d}(\varsigma_d)+ h_{co(\mathcal{Y})}(u)(1-\pi_{\gamma_d}(\varsigma_d)) \end{split} \end{align} where the first equality is by LIE. The second follows because $co(\mathbf{Y}_d) = \{Y\}$ whenever $D=d$, $h_{\{Y\}}(u) = u Y$, and $co(\mathbf{Y}_d) = co(\mathcal{Y})$ when $D\neq d$. The third is by observing that $P_O(\varsigma_d = S|D=d)=1$ since $\varsigma_d \in Sel(\mathbf{S}_d)$ and $\mathbf{S}_d = \{S\}$ when $D=d$. The fourth is by definition of $\mu_d$ and $P_O(\varsigma_d = S|D=d)=1$. The final equality is by (ref). Then observe that: \begin{align} \begin{split} &\forall u\in\{-1,1\}:\enskip u m_d(\varsigma_d)\leq E_O[h_{co(\mathbf{Y}_d)}(u)|\varsigma_d] \text{ a.s.}\\ \Leftrightarrow&\forall u\in\{-1,1\}:\enskip u m_d(\varsigma_d)\leq u\mu_d(\varsigma_d)\pi_{\gamma_d}(\varsigma_d)+ h_{co(\mathcal{Y})}(u)(1-\pi_{\gamma_d}(\varsigma_d))\text{ a.s.}\\ \Leftrightarrow&\forall u\in\{-1,1\}:\enskip u m_d(s)\leq u\mu_d(s)\pi_{\gamma_d}(s)+ h_{co(\mathcal{Y})}(u)(1-\pi_{\gamma_d}(s))\text{ $\gamma_d-$a.e.} \end{split} \end{align} where the second line follows by (ref) and the third by $\varsigma_d\overset{d}{=} \gamma_d$. Since $\varsigma_d$ was an arbitrary selection such that $(\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d,\Tilde{Z}))\cap \Bar{I}$ and $\varsigma_d\overset{d}{=} \gamma_d$ for any $d\in\{0,1\}$, by (ref), (ref) and (ref) then $\mathcal{H}(m,\gamma)\subseteq \Tilde{\mathcal{H}}(m,\gamma)$. For the converse, for any $d\in\{0,1\}$ fix any $\varsigma_d$ such that $(\varsigma_d,\Tilde{Z})\in Sel(\mathbf{S}_d)\times\{\Tilde{Z}\}\cap \Bar{I}$, letting $\varsigma_d\overset{d}{=} \gamma_d$. Let $m\in\mathcal{M}^A$ be a pair of arbitrary $m_d$ such that for $d\in\{0,1\}$: \begin{align} \forall u\in\{-1,1\}:\enskip u m_d(s)\leq u\mu_d(s)\pi_{\gamma_d}(s)+ h_{co(\mathcal{Y})}(u)(1-\pi_{\gamma_d}(s))\text{ $\gamma_d-$a.e.} \end{align} Since $E[|Y(d)|]<\infty$ and $P(G=O)>0$, $E_O[|Y(d)|]<\infty$. Given that $\mathcal{M}^A$ must be such that $m_d$ is $\gamma_d$ integrable for any $d\in\{0,1\}$, by (ref) then for $d\in\{0,1\}$ there exist $\upsilon_d\in Sel^1(\mathbf{Y}_d)$ such that $m_d(\varsigma_d) = E_O[\upsilon_d|\varsigma_d]$ a.s. It is then immediate that $ \Tilde{\mathcal{H}}(m,\gamma)\subseteq \mathcal{H}(m,\gamma)$. \underline{$\Tilde{\mathcal{H}}(m,\gamma)$ is equivalent to (ref)} By artstein1983distributions, a distribution function characterizes a selection in $Sel((\mathbf{S}_d,\Tilde{Z}))$ if and only if: \begin{align} &\forall B\in\mathcal{C}(\mathcal{S}\times \Tilde{\mathcal{Z}}):\enskip P((S(d),\Tilde{Z})\in B) \geq P((\mathbf{S}_d,\Tilde{Z})\subseteq B) \\ \Leftrightarrow &\forall B\in\mathcal{C}(\mathcal{S}):\enskip P(S(d)\in B|\Tilde{Z}) \geq P(\mathbf{S}_d\subseteq B|\Tilde{Z}) \text{ a.s.} \end{align} where the second line follows by molchanov2018random. Now consider the containment functional $P(\mathbf{S}_d\subseteq B|\Tilde{Z})$. If $B = \mathcal{S}$, $P(\mathbf{S}_d\subseteq B|\Tilde{Z})=1$. If $B \subsetneq \mathcal{S}$, then $P(\mathbf{S}_d\subseteq B|\Tilde{Z})=P(S\in B,D=d|\Tilde{Z})$. Hence, $\exists (\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d,\Tilde{Z}))$ such that $\gamma_d\in\mathcal{P}^\mathcal{S}$ and $\gamma_d\overset{d}{=} \varsigma_d$ if and only if: \begin{align} \forall B\in\mathcal{C}(\mathcal{S}):\enskip P(\varsigma_d\in B|\Tilde{Z}) \geq P(S\in B,D=d|\Tilde{Z}) \text{ a.s.} \end{align} Since $(\varsigma_d,\Tilde{Z})\in \Bar{I}$, (ref) is equivalent to: \begin{align} \forall B\in\mathcal{C}(\mathcal{S}):\enskip P(\varsigma_d\in B)&\geq ess\sup_{\Tilde{Z}}P(S\in B,D=d|\Tilde{Z})\\ &=\max\left(ess\sup_{Z}P_E(S\in B,D=d|Z),P_O(S\in B,D=d) \right) \end{align} where the first line follows since, by definition of $\Bar{I}$, $\varsigma_d\perp\mkern-9.5mu\perp \Tilde{Z}$. The second is by definition of $\Tilde{Z}$ and $P(G=O)>0$ given that two datasets are observed. Therefore: \begin{align} \begin{split} &\text{$\exists (\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d,\Tilde{Z}))\cap \Bar{I}$ s.t. $\gamma_d\overset{d}{=} \varsigma_d$}\\ \Leftrightarrow&\forall B\in\mathcal{C}(\mathcal{S}):\enskip P(\varsigma_d\in B)\geq\max\left(ess\sup_{Z}P_E(S\in B,D=d|Z),P_O(S\in B,D=d) \right) \end{split} \end{align} By definition of $\Tilde{\mathcal{H}}(m,\gamma)$ and (ref): \begin{align} \begin{split} \Tilde{\mathcal{H}}(m,\gamma)&= \left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times (\mathcal{P}^{\mathcal{S}})^2: \forall d\in\{0,1\}, \text{$\forall B\in\mathcal{C}(\mathcal{S}),$} \\ \text{$\gamma_d(B)\geq\max\left(ess\sup_{Z}P_E(S\in B,D=d|Z),P_O(S\in B,D=d) \right)$,} \\ \text{ $\forall u\in\{-1,1\}:\enskip u m_d(s)\leq u\mu_d(s)\pi_{\gamma_d}(s)+ h_{co(\mathcal{Y})}(u)(1-\pi_{\gamma_d}(s)),\text{ $\gamma_d-$a.e.}$}\end{array}\right\} \end{split} \end{align} By its definition, if $\mathfrak{C}$ is a core determining class, (ref) is equivalent to: \begin{align} \begin{split} &\text{$\exists (\varsigma_d,\Tilde{Z})\in Sel((\mathbf{S}_d,\Tilde{Z}))\cap \Bar{I}$ s.t. $\gamma_d\overset{d}{=} \varsigma_d$}\\ \Leftrightarrow&\forall B\in\mathfrak{C}:\enskip P(\varsigma_d\in B)\geq\max\left(ess\sup_{Z}P_E(S\in B,D=d|Z),P_O(S\in B,D=d) \right) \end{split} \end{align} It is then immediate that in (ref) the condition $\forall B\in\mathcal{C}(\mathcal{S})$ can be replaced by $\forall B\in\mathfrak{C}$. The result then follows by $\mathcal{H}(m,\gamma) = \Tilde{\mathcal{H}}(m,\gamma)$.

The argument combines Artstein’s theorem (artstein1983distributions) with the conditional Aumann expectation, tools that are typically applied separately, to address a novel technical challenge: jointly bounding (i) the distribution function of a selection of a random set and (ii) a conditional mean whose conditioning $\sigma$-algebra is generated by that selection.

To build intuition for the bounds, consider conditions that are necessary for $(m,\gamma)$ to be compatible with the data and maintained assumptions. The proof shows that these are also sufficient, and hence deliver sharp bounds. First, recalling $\gamma_d:=P_{S(d)}$, it is immediate that for any closed subset $B\in\mathcal{C}(\mathcal S)$: $\gamma_d(B)\geq P(S(d)\in B,D=d)$. By Assumptions (ref) and (ref), $S(d)\perp\mkern-9.5mu\perp G$ and $S(d)\perp\mkern-9.5mu\perp Z|G=E$. It is therefore also necessary that each $\gamma_d$ satisfies:

align[align omitted — 196 chars of source]

which coincides with the restrictions on $\gamma_d$ in the theorem. This represents an intersection bound on $\gamma_d$ reflecting that $S(d)$ is observed in both datasets when $D=d$.

Second, recalling that under Assumption (ref) $m_d(S(d)) = E_O[Y(d)|S(d)]$, by the law of iterated expectations:

align*[align* omitted — 111 chars of source]

Let $\pi_{\gamma_d}(s):=P_O(D=d|S(d)=s)$ denote the latent propensity score for a given $\gamma_d\overset{d}{=} S(d)$. Since $S(d)=S$ whenever $D=d$ and $S(d)\perp\mkern-9.5mu\perp G$, $\pi_{\gamma_d}$ is the Radon-Nikodym derivative such that $\pi_{\gamma_d}(S(d))=\frac{dP_O(S,D=d)}{d\gamma_d}$. Because $E_O[Y(d)|D\neq d,S(d)]\in\mathcal{Y}$ is never observed, the usual worst-case bounds arguments imply:

align[align omitted — 283 chars of source]

which is identical to the restrictions on $m_d$ in the theorem. Accordingly, these conditions represent worst-case bounds on $m_d$ for a given $\gamma_d$. Notably, the bounds on $m_d$ may change with the particular $\gamma_d$ under consideration.

Finally, (ref) also enables the use of a core-determining class (CDC) to remove redundant constraints on $\gamma$ without loss of information. Intuitively, a CDC does so by eliminating constraints that are implied by the remaining ones.

Characterizing $\mathcal{H}(\tau)$

The sharp identified set $\mathcal{H}(\tau)$ follows as the image of the functional $T$ over the set of possible $(m,\gamma)$. The next result shows that under easily verifiable high-level conditions on $\mathcal{M}^A$, $\mathcal{H}(\tau)$ is an interval which can be characterized by solving two optimization problems.

theoremEnd[proof at the end, no link to proof]{theorem} Let Assumptions (ref), (ref), and (ref) hold. Suppose that $E[|Y(d)|]<\infty$ for $d\in\{0,1\}$ and that $\mathcal{M}^A$ is closed and convex. Then the closure of $\mathcal{H}(\tau)$ is: \begin{align} \left[\inf_{(\tilde{m},\tilde{\gamma})\in\mathcal{H}(m,\gamma)}T(\tilde{m},\tilde{\gamma}),\sup_{(\tilde{m},\tilde{\gamma})\in\mathcal{H}(m,\gamma)}T(\tilde{m},\tilde{\gamma})\right]. \end{align}
proofEndRecall $\mathcal{H}(\tau)=\{T(m,\gamma):(m,\gamma)\in\mathcal{H}(m,\gamma)\}$. The proof proceeds in three steps. Step 1: $\mathcal{H}(m,\gamma)$ is convex. Take any $(m,\gamma),(m',\gamma')\in\mathcal{H}(m,\gamma)$ and any $a\in(0,1)$. Define: \begin{align} m^a:=a m+(1-a)m',\qquad \gamma^a:=a\gamma+(1-a)\gamma'. \end{align} Since $\mathcal{M}^A$ is convex by assumption, $m^a\in\mathcal{M}^A$. $\mathcal{P}^\mathcal{S}$ is the space of probability measures and is therefore convex, so $\gamma_d^a\in\mathcal{P}^\mathcal{S}$ for each $d$. It remains to verify that set constraints on $m_d$ and $\gamma_d$ from (ref) are satisfied for $d\in\{0,1\}$ . For $\gamma_d$ constraints, note that for any $d\in\{0,1\}$ and $B\in\mathcal{C}(\mathcal{S})$: \begin{align} \gamma_d^a(B)&=a\gamma_d(B)+(1-a)\gamma'_d(B)\\ &\geq a\max\!\left(ess\sup_{Z}P_E(S\in B,D=d\mid Z),P_O(S\in B,D=d)\right)\notag\\ &\quad+(1-a)\max\!\left(ess\sup_{Z}P_E(S\in B,D=d\mid Z),P_O(S\in B,D=d)\right)\notag\\ &=\max\!\left(ess\sup_{Z}P_E(S\in B,D=d\mid Z),P_O(S\in B,D=d)\right),\notag \end{align} where the first line is by definition of $\gamma^a$, the second is by $(m,\gamma), (m',\gamma')\in \mathcal{H}(m,\gamma)$ and (ref), and the third is by observation. For the $m_d$ constraint, fix any $d\in\{0,1\}$. For any $B$ with $\gamma_d(B)=0$, $P_O(S\in B,D=d)\leq \gamma_d(B)=0$. Hence, $P_O(S\in\cdot,D=d)\ll \gamma_d $, and $P_O(S\in\cdot,D=d)\ll \gamma'_d$. Since $\gamma_d^a=a\gamma_d+(1-a)\gamma'_d$, also $P_O(S\in\cdot,D=d)\ll \gamma_d^a$. Then by the Radon-Nikodym theorem there exist measurable functions: \begin{align} \pi_{\gamma_d}:=\frac{dP_O(S\in\cdot,D=d)}{d\gamma_d},\qquad \pi_{\gamma'_d}:=\frac{dP_O(S\in\cdot,D=d)}{d\gamma'_d},\qquad \pi^a:=\frac{dP_O(S\in\cdot,D=d)}{d\gamma_d^a}. \end{align} Since $\gamma_d^a=a\gamma_d+(1-a)\gamma'_d$ with $a\in(0,1)$, $\gamma_d \ll \gamma_d^a$ and $\gamma'_d \ll \gamma_d^a$. Therefore, similarly, the following exist: \begin{align} \rho:=\frac{d\gamma_d}{d\gamma_d^a},\qquad \rho':=\frac{d\gamma'_d}{d\gamma_d^a}. \end{align} Then $\rho,\rho'\geq 0$ by nonnegativity of measures, and: \begin{align} a\rho+(1-a)\rho'=1\qquad \gamma_d^a-a.e. \end{align} by $\gamma_d^a=a\gamma_d+(1-a)\gamma'_d$ and linearity of the Radon--Nikodym derivative. By the chain rule for Radon--Nikodym derivatives, set: \begin{align} \pi^a=\pi_{\gamma_d}\rho=\pi_{\gamma'_d}\rho'\qquad \gamma_d^a-a.e. \end{align} with the convention $\pi_{\gamma_d}=0$ on the event $\{\rho=0\}$ and similarly $\pi_{\gamma'_d}=0$ on $\{\rho'=0\}$. Since $(m,\gamma)\in\mathcal{H}(m,\gamma)$, for each $u\in\{-1,1\}$: \begin{align} u m_d(s)\leq u\mu_d(s)\pi_{\gamma_d}(s)+h_{co(\mathcal{Y})}(u)\bigl(1-\pi_{\gamma_d}(s)\bigr)\qquad \gamma_d-a.e. \end{align} by (ref), and similarly with $(m',\gamma')$. Since (ref) holds $\gamma_d$-a.e. and $\gamma_d(B)=\int_B \rho(s)d\gamma_d^a(s)$ for every measurable $B$, it follows that (ref) also holds $\gamma_d^a$-a.e. on $\{\rho>0\}$. Analogously, the corresponding inequality for $(m_d',\gamma_d')$ holds $\gamma_d^a$-a.e. on $\{\rho'>0\}$. Fix $u\in\{-1,1\}$. On $\{\rho>0\}\cap\{\rho'>0\}$, by (ref) rewrite (ref) as: \begin{align} u m_d(s)\leq u\mu_d(s)\frac{\pi^a(s)}{\rho(s)}+h_{co(\mathcal{Y})}(u)\Big(1-\frac{\pi^a(s)}{\rho(s)}\Big)\qquad \gamma_d^a-a.e.\ on \{\rho>0\}, \end{align} and similarly: \begin{align} u m'_d(s)\leq u\mu_d(s)\frac{\pi^a(s)}{\rho'(s)}+h_{co(\mathcal{Y})}(u)\Big(1-\frac{\pi^a(s)}{\rho'(s)}\Big)\qquad \gamma_d^a-\text{a.e.\ on }\{\rho'>0\}. \end{align} Then define $A(s):=\frac{a}{\rho(s)}+\frac{1-a}{\rho'(s)}$ on the event $\{\rho>0\}\cap\{\rho'>0\}$ and note: \begin{align} \begin{split} u m_d^a(s) &=au m_d(s)+(1-a)u m'_d(s)\\ &\leq a\Big(u\mu_d(s)\frac{\pi^a(s)}{\rho(s)}+h_{co(\mathcal{Y})}(u)\Big(1-\frac{\pi^a(s)}{\rho(s)}\Big)\Big)\\ &\quad +(1-a)\Big(u\mu_d(s)\frac{\pi^a(s)}{\rho'(s)}+h_{co(\mathcal{Y})}(u)\Big(1-\frac{\pi^a(s)}{\rho'(s)}\Big)\Big)\\ &=u\mu_d(s)\pi^a(s) \Big(\frac{a}{\rho(s)}+\frac{1-a}{\rho'(s)}\Big) + h_{co(\mathcal{Y})}(u)\Big(1-\pi^a(s)\Big(\frac{a}{\rho(s)}+\frac{1-a}{\rho'(s)}\Big)\Big)\\ &=u\mu_d(s)\pi^a(s) A(s) + h_{co(\mathcal{Y})}(u)\bigl(1-\pi^a(s) A(s)\bigr), \text{ on }\{\rho>0\}\cap\{\rho'>0\} \end{split} \end{align} where the first line is by definition of $m^a_d$, the second is by the two inequalities above, the third is by rearrangement and factoring out $\pi^a(s)$, and the fourth is by definition of $A(s)$. Observe that: \begin{align} \begin{split} A(s)-1 &=\frac{a}{\rho(s)}+\frac{1-a}{\rho'(s)}-1\\ &=\frac{a\rho'(s)+(1-a)\rho(s)}{\rho(s)\rho'(s)}-\frac{1}{a\rho(s)+(1-a)\rho'(s)}\\ &=\frac{(a\rho'(s)+(1-a)\rho(s))(a\rho(s)+(1-a)\rho'(s))-\rho(s)\rho'(s)}{\rho(s)\rho'(s)(a\rho(s)+(1-a)\rho'(s))}\\ &=\frac{a(1-a)(\rho(s)-\rho'(s))^2}{\rho(s)\rho'(s)(a\rho(s)+(1-a)\rho'(s))}\\ &=\frac{a(1-a)(\rho(s)-\rho'(s))^2}{\rho(s)\rho'(s)}\geq 0 \qquad\text{on }\{\rho>0\}\cap\{\rho'>0\}, \end{split} \end{align} where the first line is by definition of $A(s)$, the second is by (ref), the third and fourth are by observation, and the fifth is by (ref). Note that $\mu_d(s)=E_O[Y|S=s,D=d]$ is well-defined and finite $P_O(S\in\cdot|D=d)-$a.e. because $E_O[|Y|\mathbbm{1}[D=d]]<\infty$ by assumption. Moreover, since $Y\in\mathcal{Y}$ a.s., $\mu_d(s)\in\mathrm{co}(\mathcal{Y})$ wherever it is defined and hence $u\mu_d(s)\leq h_{\mathrm{co}(\mathcal{Y})}(u),\text{ for }u\in\{-1,1\}$. Since $A(s)\geq 1$ and $u\mu_d(s)-h_{\mathrm{co}(\mathcal{Y})}(u)\leq 0$, by (ref): \begin{align} um_d^a(s)\leq h_{\mathrm{co}(\mathcal{Y})}(u)+(u\mu_d(s)- h_{\mathrm{co}(\mathcal{Y})}(u))\pi^a(s)A(s)\leq u\mu_d(s)\pi^a(s)+h_{\mathrm{co}(\mathcal{Y})}(u)(1-\pi^a(s)) \end{align} on $\{\rho>0\}\cap\{\rho'>0\}$. On $\{\rho=0\}\cup\{\rho'=0\}$, note that $\gamma_d^a(\{\rho=0\} \cap\{\rho'=0\})=0$ by (ref), so the constraint holds trivially there. On the remaining events $\{\rho>0\}\cap\{\rho'=0\}$ and $\{\rho=0\}\cap\{\rho'>0\}$, (ref) implies $\pi^a=0$ $\gamma_d^a$-a.e., so the constraint reduces to $um_d^a(s)\leq h_{\mathrm{co} (\mathcal{Y})}(u)$. Since $(m,\gamma),(m',\gamma')\in\mathcal{H} (m,\gamma)$ implies $m\in\mathcal{M}^A\subseteq\mathcal{M}$ and $m'\in\mathcal{M}^A\subseteq\mathcal{M}$, we have $m_d(s),m'_d(s) \in\mathcal{Y}\subseteq\mathrm{co}(\mathcal{Y})$ everywhere. Therefore $m_d^a(s)=am_d(s)+(1-a)m'_d(s)\in\mathrm{co}(\mathcal{Y})$ everywhere by convexity, and hence $um_d^a(s)\leq h_{\mathrm{co} (\mathcal{Y})}(u)$ holds everywhere. Combining the cases $\{\rho>0\}\cap\{\rho'>0\}$ and $\{\rho=0\}\cup\{\rho'=0\}$, for every $u\in\{-1,1\}$: \begin{align} u m_d^a(s)\leq u\mu_d(s)\pi^a(s)+h_{co(\mathcal{Y})}(u)\bigl(1-\pi^a(s)\bigr) \qquad \gamma_d^a-\text{a.e.} \end{align} Since $\pi^a=dP_O(S\in\cdot,D=d)/d\gamma_d^a=\pi_{\gamma_d^a}$, the moment constraints hold for $(m^a,\gamma^a)$. Thus $(m^a,\gamma^a)\in\mathcal{H}(m,\gamma)$ and $\mathcal{H}(m,\gamma)$ is convex. \underline{\textbf{Step 2:} $T$ is continuous on line segments over $\mathcal{H}(m,\gamma)$.} Fix arbitrary $(m^1,\gamma^1),(m^0,\gamma^0)\in\mathcal{H}(m,\gamma)$ and define for $t\in[0,1]$: \begin{align} m^t:=t m^1+(1-t)m^0,\qquad \gamma^t:=t\gamma^1+(1-t)\gamma^0. \end{align} By Step 1, $(m^t,\gamma^t)\in\mathcal{H}(m,\gamma)$ for every $t\in[0,1]$, so $T(m^t,\gamma^t)$ is well-defined and finite for every $t$ by $E[|Y(d)|]<\infty$. Fix any $t\in(0,1)$ and $d\in\{0,1\}$. Since $(m^t,\gamma^t)\in\mathcal{H}(m,\gamma)$: \begin{align} \int |m^t_d|d\gamma^t_d<\infty. \end{align} Then: \begin{align} \int |m^t_d|d\gamma^0_d\leq \frac{1}{1-t}\int |m^t_d|d\gamma^t_d<\infty, \qquad \int |m^t_d|d\gamma^1_d\leq \frac{1}{t}\int |m^t_d|d\gamma^t_d<\infty, \end{align} where the inequalities are by $\gamma^t_d=t\gamma^1_d+(1-t)\gamma^0_d\geq (1-t)\gamma^0_d$ and $\gamma^t_d\geq t\gamma^1_d$, and the finiteness is by (ref). Moreover, by $|t x+(1-t)y|\geq t|x|-(1-t)|y|$: \begin{align} |m^t_d| =|t m^1_d+(1-t)m^0_d| \geq t|m^1_d|-(1-t)|m^0_d|. \end{align} Rearranging yields $t|m^1_d|\leq |m^t_d|+(1-t)|m^0_d|$. Integrating with respect to $\gamma^0_d$ gives: \begin{align} t\int |m^1_d|d\gamma^0_d \leq \int |m^t_d|d\gamma^0_d+(1-t)\int |m^0_d|d\gamma^0_d<\infty, \end{align} where the finiteness is by (ref) and $\int |m^0_d|d\gamma^0_d<\infty$, since $(m^0,\gamma^0)\in\mathcal{H}(m,\gamma)$. Since $t>0$, $\int |m^1_d|d\gamma^0_d<\infty$. The argument for $\int |m^0_d|d\gamma^1_d<\infty$ is analogous. Hence for each $d$, the integrals $\int m^1_dd\gamma^1_d$, $\int m^1_dd\gamma^0_d$, $ \int m^0_dd\gamma^1_d$, $ \int m^0_dd\gamma^0_d$ are finite. Therefore: \begin{align} \begin{split} \int m^t_dd\gamma^t_d &=\int (t m^1_d+(1-t)m^0_d)d(t\gamma^1_d+(1-t)\gamma^0_d)\\ &=t^2\int m^1_dd\gamma^1_d +t(1-t)\int m^1_dd\gamma^0_d +t(1-t)\int m^0_dd\gamma^1_d +(1-t)^2\int m^0_dd\gamma^0_d, \end{split} \end{align} where the first line is by definition of $m^t_d$ and $\gamma^t_d$, and the second by bilinearity of the integral and the fact that individual integrals are finite. Then $f(t): = \int m^t_1d\gamma^t_1-\int m^t_0d\gamma^t_0$ is a quadratic polynomial in $t$, hence continuous on $[0,1]$. Because $f(t) = T(m^t,\gamma^t)$, $T$ is continuous on line segments. \underline{\textbf{Step 3:} $cl(\mathcal{H}(\tau))$ is a closed interval.} Step 1 implies $\mathcal{H}(m,\gamma)$ is convex, hence path-connected by line segments: for any two points $(m,\gamma),(m',\gamma')\in\mathcal{H}(m,\gamma)$, the line segment $\{(m^t,\gamma^t):t\in[0,1]\}$ lies entirely in $\mathcal{H}(m,\gamma)$. By Step 2, $T$ is continuous along every such line segment. Therefore, for any two values $\tau_0=T(m,\gamma),\tau_1=T(m',\gamma')\in\mathcal{H}(\tau)$, the set $\mathcal{H}(\tau)$ contains the continuous image $\{T(m^t,\gamma^t):t\in[0,1]\}$ of $[0,1]$ connecting $\tau_0$ and $\tau_1$. By the intermediate value theorem, this image contains the whole interval $[\min(\tau_0,\tau_1),\max(\tau_0,\tau_1)]$. Hence $\mathcal{H}(\tau)\subseteq\mathbb{R}$ is an interval. The closure of an interval in $\mathbb{R}$ is the closed interval with endpoints given by its infimum and supremum. Thus: \begin{align} cl(\mathcal{H}(\tau)) = \left[\inf_{(\tilde{m},\tilde{\gamma})\in\mathcal{H}(m,\gamma)}T(\tilde{m},\tilde{\gamma}),\; \sup_{(\tilde{m},\tilde{\gamma})\in\mathcal{H}(m,\gamma)}T(\tilde{m},\tilde{\gamma})\right]. \end{align}

Using optimization problems to characterize identified sets has become common in partial identification. Such representations typically follow from linear objectives and polyhedral, hence convex, constraint sets. (ref) requires a different argument since $T$ is bilinear -- and thus separately continuous -- and the constraint set is bilinearly constrained and therefore need not be convex. The proof shows that $\mathcal{H}(m,\gamma)$ is nevertheless convex under the assumptions of the theorem and that $T$ is continuous along line segments in $\mathcal{H}(m,\gamma)$. Then $\mathcal{H}(\tau)$ is a line-continuous image of a convex set, hence a connected set in $\mathbb{R}$, i.e., an interval.

The restrictions $\mathcal{M}^A$ enter as constraints in the optimization problems. This offers three advantages. First, it enables the implementation of LIV and TI. Second, it nests existing point-identification results and generalizes them by allowing for imperfect compliance (for examples see (ref)). Third, it provides a tool that facilitates development of new restrictions on $m$ tailored for specific empirical settings by computationally producing the corresponding sharp bounds for $\tau$. This removes the need to derive closed-form expressions and to prove sharpness for each set of assumptions.

remarkAssumptions (ref) and (ref) are representable via linear equality and inequality restrictions on $m$. Therefore, each resulting $\mathcal{M}^A$ is an intersection of (possibly infinitely many) affine subspaces and halfspaces in a linear space of functions on $\mathcal{S}$, and is thus convex. Moreover, because the restrictions are imposed pointwise (or $\gamma_d-$a.e.) and are preserved under limits, the corresponding $\mathcal{M}^A$ are closed. Furthermore, whenever $m$ is identified by the data, such as under existing assumptions in (ref), $\mathcal{M}^A$ is a singleton ($\gamma_d-$a.e.) and hence closed and convex.

Implementation and Estimation

This section discusses how the optimization problems can be used to tractably characterize and estimate $\mathcal{H}(\tau)$. The problems become finite-dimensional when $\mathcal{S}$ is a finite set, as shown by the following corollary.

theoremEnd[proof at the end, no link to proof]{corollary} Suppose that assumptions of (ref) hold and that $k:=|\mathcal{S}|<\infty$. Let $\Delta(k)$ denote the $k-$dimensional simplex. Then the closure of $\mathcal{H}(\tau) $ is: \begin{align} \left[\inf_{(\Tilde{m},\Tilde{\gamma})\in\mathcal{H}(m,\gamma)}T(\Tilde{m},\Tilde{\gamma}),\sup_{(\Tilde{m},\Tilde{\gamma})\in\mathcal{H}(m,\gamma)}T(\Tilde{m},\Tilde{\gamma})\right]. \end{align} where: \begin{align} \begin{split} \mathcal{H}(m,\gamma) =&\left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times (\Delta(k))^2: $\forall d\in\{0,1\}$, $ \forall s\in\mathcal{S}$,\\ $\gamma_d(s)\geq \max\left(ess\sup_{Z}P_E(S = s,D=d|Z),P_O(S=s,D=d) \right),$\\ $\left(m_d(s)-\inf\mathcal{Y}\right)\gamma_d(s)\geq \left(E_{O}[Y|S=s,D=d]-\inf\mathcal{Y}\right)P_{O}(S=s,D=d)$ ,\\ $\left(\sup\mathcal{Y}-m_d(s)\right)\gamma_d(s)\geq \left(\sup\mathcal{Y}-E_{O}[Y|S=s,D=d]\right)P_{O}(S=s,D=d)$ \end{array}\right\} \end{split} \end{align}
proofEndSince $\mathcal{S} = \{1,2,\hdots,k\}$, represent $\gamma_d$ as an element of the $k-$dimensional simplex $\Delta(k)$ and $m_d\in\mathcal{Y}^k$. Let $\gamma_d(s)$ and $m_d(s)$ denote the $s-$th element of the corresponding vectors. Then: \begin{align} \begin{split} \mathcal{H}(m,\gamma) &= \left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times(\Delta(k))^2: $\forall d\in\{0,1\}$, $\forall B\in\mathcal{C}(\mathcal{S})$, \\ $\gamma_d(B)\geq \max\left(ess\sup_{Z}P_E(S\in B,D=d|Z),P_O(S\in B,D=d) \right)$,\\ $\forall u\in\{-1,1\}$: $um_d(s)\leq u\mu_d(s)\pi_{\gamma_d}(s)+ h_{co(\mathcal{Y})}(u)(1-\pi_{\gamma_d}(s))$ $\gamma_d-$a.e. \end{array}\right\}\\ &= \left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times(\Delta(k))^2: $\forall d\in\{0,1\}$, $\forall B\in\mathcal{C}(\mathcal{S})$, \\ $\gamma_d(B)\geq \max\left(ess\sup_{Z}P_E(S\in B,D=d|Z),P_O(S\in B,D=d) \right)$, $\forall u\in\{-1,1\}$:\\ $um_d(s)\leq u\mu_d(s)\frac{P_O(S=s,D=d)}{\gamma_d(s)}+ h_{co(\mathcal{Y})}(u)\left(1-\frac{P_O(S=s,D=d)}{\gamma_d(s)}\right)$ $\gamma_d-$a.e. \end{array}\right\} \end{split} \end{align} where the first line is by (ref). The second is by definition of $\pi_{\gamma_d}(s)$ and $\gamma_d$ being supported on $\mathcal{S}$ with $|\mathcal{S}|<\infty$. $\mathcal{S}$ is closed by definition. Since it is finite, it is bounded. Hence, $\mathbf{S}_d$ is almost surely compact, by definition. Then, by beresteanu2012partial $\{\{s\}: s\in\mathcal{S}\}$ is a core-determining class for the containment functional of $\mathbf{S}_d$. Then: \begin{align} \begin{split} \mathcal{H}(m,\gamma) =&\left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times (\Delta(k))^2: \text{$\forall d\in\{0,1\}$, $ \forall s\in\mathcal{S}$, $\forall u\in\{-1,1\},$}\\ \text{$\gamma_d(s)\geq \max\left(ess\sup_{Z}P_E(S = s,D=d|Z),P_O(S=s,D=d) \right),$}\\ \text{$um_d(s)\leq u\mu_d(s)\frac{P_O(S=s,D=d)}{\gamma_d(s)}+ h_{co(\mathcal{Y})}(u)\left(1-\frac{P_O(S=s,D=d)}{\gamma_d(s)}\right) $ $\gamma_d-$a.e.} \end{array}\right\}\\ =&\left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times (\Delta(k))^2: \text{$\forall d\in\{0,1\}$, $ \forall s\in\mathcal{S}$, $\forall u\in\{-1,1\},$}\\ \text{$\gamma_d(s)\geq \max\left(ess\sup_{Z}P_E(S = s,D=d|Z),P_O(S=s,D=d) \right),$}\\ \text{$um_d(s)\leq uE[Y|S=s,D=d]\frac{P_O(S=s,D=d)}{\gamma_d(s)}+ h_{co(\mathcal{Y})}(u)\left(1-\frac{P_O(S=s,D=d)}{\gamma_d(s)}\right) $ $\gamma_d-$a.e.} \end{array}\right\}\\ =&\left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times (\Delta(k))^2: \text{$\forall d\in\{0,1\}$, $ \forall s\in\mathcal{S}$,}\\ \text{$\gamma_d(s)\geq \max\left(ess\sup_{Z}P_E(S = s,D=d|Z),P_O(S=s,D=d) \right),$}\\ \text{$\left(m_d(s)-\inf\mathcal{Y}\right)\gamma_d(s)\geq \left(E_{O}[Y|S=s,D=d]-\inf\mathcal{Y}\right)P_{O}(S=s,D=d)$ $\gamma_d-$a.e.},\\ \text{$\left(\sup\mathcal{Y}-m_d(s)\right)\gamma_d(s)\geq \left(\sup\mathcal{Y}-E_{O}[Y|S=s,D=d]\right)P_{O}(S=s,D=d)$ $\gamma_d-$a.e.} \end{array}\right\}\\ =&\left\{\begin{array}{l} (m,\gamma)\in\mathcal{M}^A\times (\Delta(k))^2: \text{$\forall d\in\{0,1\}$, $ \forall s\in\mathcal{S}$,}\\ \text{$\gamma_d(s)\geq \max\left(ess\sup_{Z}P_E(S = s,D=d|Z),P_O(S=s,D=d) \right),$}\\ \text{$\left(m_d(s)-\inf\mathcal{Y}\right)\gamma_d(s)\geq \left(E_{O}[Y|S=s,D=d]-\inf\mathcal{Y}\right)P_{O}(S=s,D=d)$},\\ \text{$\left(\sup\mathcal{Y}-m_d(s)\right)\gamma_d(s)\geq \left(\sup\mathcal{Y}-E_{O}[Y|S=s,D=d]\right)P_{O}(S=s,D=d)$} \end{array}\right\}. \end{split} \end{align} where the first line is by (ref) and (ref), the second line is by definition of $\mu_d(s)$, the third is by definition of $h_{co(\mathcal{Y})}(u)$ and rearrangement, and the fourth is by observation. The result then follows from (ref) and (ref).

When short-term outcomes have finite support, the infinite-dimensional problems in (ref) reduce to finite-dimensional generalized bilinear optimization programs (al1992generalized, for other examples of bilinear programs see dutz2021selection and shea2022testing). Although such problems are generally nonconvex, modern general-purpose solvers can compute globally optimal solutions via spatial branch-and-bound methods (e.g. gurobi). (ref) provides a consistent criterion-based estimator based on the corollary.

Focusing on cases where relevant variables are finitely supported or discretized has been the predominant approach in work that relies on Artstein's theorem (galichon2011set, russell2021sharp, luo2024selecting). Discretization is also common in related settings (rambachan2024program, park2024informativeness). Appendix (ref) discusses the interpretation of results when $S$ is discretized. General-purpose methods for solving bilinear infinite-dimensional programs remain an interesting open problem. In special cases, ongoing work shows that additional structure permits tractable implementation via optimal transport theory with approximation guarantees even when $\mathcal{S}$ is infinite; see (ref).

A few details are worth highlighting. First, $\mathcal{Y}$ is unrestricted. When it is unbounded, then the $m_d$ constraints involving $\mathcal{Y}$ are trivially satisfied. Second, to reduce computational complexity, the corollary uses $\{\{s\}:s\in\mathcal{S}\}$ as a CDC. (ref) illustrates the resulting reduction as a function of $|\mathcal{S}|$.\footnote{The number of constraints on $m$ imposed by the data given $\gamma$ is $4|\mathcal{S}|$; the total number depends on the modeling assumptions. The number of constraints may potentially be reduced further by adapting methods such as luo2024selecting, as $\{\{s\}:s\in\mathcal{S}\}$ is not necessarily the smallest CDC.} If $S(d)$ represents percentiles, then not using the CDC already results in a prohibitively complex constraint set.

table[table omitted — 454 chars of source]
remarkProblems $\sup/\inf_{(\Tilde{m},\Tilde{\gamma})\in\mathcal{H}(m,\gamma)} T(\Tilde{m},\Tilde{\gamma})$ become linear in certain cases. This substantially simplifies computation and occurs under either of the following two conditions. First, if $\mathcal{H}(m,\gamma)=\{m\}\times\mathcal{H}(\gamma)$, the problems reduce to linear programs over $\gamma$. For example, assumptions that point-identify $m$ independently of $\gamma$, such as LUC in (ref), yield $\mathcal{H}(m,\gamma)=\{m\}\times\mathcal{H}(\gamma)$. In this case, one may use optimal transport theory to characterize the identified set even when $\mathcal{S}$ is infinite voronin2025generalized. Second, if $\mathcal{H}(m,\gamma)=\mathcal{H}(m)\times\{\gamma\}$ and $\mathcal{M}^A$ is representable by linear constraints, the problems reduce to linear programs over $m$. This case arises under Assumptions (ref) and (ref) when the right-hand sides of the constraints for each $\gamma_d$ sum to one, as under perfect compliance or when the datasets jointly point-identify $\gamma$.

Identifying Power

Closed-form bounds, when available, can provide intuition about the sources of identifying power. Optimization-based bounds typically trade off this transparency for the ability to computationally deliver sharp bounds under a wide range of assumptions (see, e.g., mogstad2018using, torgovitsky2019nonparametric). Under Assumptions (ref) and (ref), however, one can still shed some light on the sources of identifying power by characterizing $m^{UB}$ and $m^{LB}$ that attain the upper and lower bounds on $\tau$, given a feasible $\gamma$. To simplify exposition, focus on the bounds for $E[Y(1)]$ for a fixed feasible $\gamma$, as the intuition for $E[Y(0)]$ (and hence $\tau$) is analogous.\footnote{If compliance is perfect, there is a unique feasible $\gamma$. Otherwise, there may be a set of feasible $\gamma$, and the bounds are obtained by taking the union over all feasible $\gamma$.} Based on (ref) in Appendix (ref) and letting $s'\leq s$ be the product order, under Assumption (ref):

align[align omitted — 322 chars of source]

These represent the extremal monotonic $m_1$ that are compatible with worst-case (WC) bounds on $m_1$ given $\gamma$ in (ref). Similarly, under Assumption (ref):

align[align omitted — 322 chars of source]

Under (ref), $m_1$ must satisfy the WC bounds for both $m_1$ and $m_0$. Hence the bound on it is given by the intersection of the two WC bounds at each support point $s$.

Based on these results, identifying power under Assumption (ref) comes from the extent to which the WC bounds on $m_1$ violate monotonicity. Imposing monotonicity replaces the WC lower (upper) endpoint by its left-envelope $\sup_{s'\le s}(\cdot)$ (right-envelope $\inf_{s'\ge s}(\cdot)$), which raises the lower bound (lowers the upper bound) wherever the WC endpoints are decreasing. Under Assumption (ref), identifying power comes from how small the intersection of the WC bounds for $m_1$ and $m_0$ is at each $s$. The bounds tighten when those intervals overlap little, e.g., when some $\pi_{\gamma_d}(s)$ are large, so the corresponding WC interval is narrow. Both mechanisms are illustrated for a non-pathological DGP in (ref). Finally, bounded $\mathcal{Y}$ is necessary for either assumption to yield identifying power, since otherwise WC bounds are trivial for any feasible $\gamma$.

figure[figure omitted — 1,359 chars of source]
remarkBoth Assumptions (ref) and (ref) can yield point identification of $m$, given $\gamma$. For Assumption (ref), an important case is when $D$ is constant in the observational data ($G=O$), for instance, when one treatment is not available in that dataset. For Assumption (ref), point identification arises when the WC endpoints are sufficiently non-monotonic that imposing nondecreasing $m_d$ makes it constant (since the lower and upper monotone envelopes coincide). This may happen, for example, when the WC endpoints are strictly decreasing in $s$, or when the endpoints exhibit more pronounced oscillations than in (ref).

Empirical Illustration: Long-term Effects of Head Start Participation

Head Start is the largest early childhood education program in the United States, serving approximately 730{,}000 low-income preschool-age children in 2023.\footnote{Link: \href{https://headstart.gov/program-data/article/head-start-program-facts-fiscal-year-2023}{https://headstart.gov/program-data/article/head-start-program-facts-fiscal-year-2023} (Last accessed 01/14/2025).} Established in 1965 as part of the “War on Poverty,” it was intended to help narrow gaps between disadvantaged and more advantaged children on a national scale. Its long-term effects have been studied extensively, primarily using observational designs. A common approach compares siblings within the same family who did and did not participate in Head Start (currie1995does, garces2002longer, deming2009early, bauer2016long), while others exploit variation in program funding, income-based eligibility, or rollout timing to identify local average treatment effects (ludwig2007does, carneiro2014long, bailey2021prep). kline2016evaluating instead monetize the experimental LATE of Head Start on test scores using estimates of the test-score--earnings relationship from administrative follow-up of Tennessee Project STAR participants (chetty2011does). Despite this long history, the literature has yet to reach consensus on the program’s long-term impacts (gibbs2011does, pages2020elusive), as the assumptions and generalizability of parameters identified by existing approaches remain debated (ludwig2008long, elango2015early, gonzalez2020within, garcia2020quantifying, miller2023selection). I illustrate an alternative approach by applying the proposed method to estimate long-term average treatment effects of Head Start for eligible individuals under assumptions that do not restrict selection into treatment.

Data

I combine individual-level data from the Head Start Impact Study (HSIS) and the Child and Young Adult Supplement to the National Longitudinal Survey of Youth 1979 cohort (CNLSY).

HSIS was an experimental trial of Head Start mandated by the 105th US Congress as part of the program’s 1998 reauthorization. In fall of 2002, $4{,}667$ children from nationally representative cohorts aged 3 and 4 were enrolled in the experiment. Across 383 randomly selected Head Start centers, participants were randomized either to a treatment group assigned to enroll in Head Start or to a control group barred from enrolling, yielding $Z\in\{0,1\}$. Since participants were followed only through third grade, HSIS does not contain adolescence or adulthood outcomes that may be of interest. However, it does contain childhood math and reading test scores. I follow kline2016evaluating and kamat2024identifying by pooling all children into a single cohort, then retaining only those with available scores, resulting in an experimental sample size of $n_E = 3{,}540$.

CNLSY is a biennial longitudinal survey introduced in 1986 that has tracked 11,545 children born to participants in the National Longitudinal Survey of Youth 1979 cohort (NLSY79), which, like HSIS, was designed to be nationally representative. It reveals long-term outcomes previously studied in the literature and non-randomized Head Start participation. As in deming2009early, I consider eight long-term outcomes: grade repetition, diagnosis of a learning disability, high school graduation, “idleness”, criminal involvement, teenage parenthood, self-reported health status, and average earnings. CNLSY also contains comparable childhood math and reading test scores that can be used as $S$, in conjunction with HSIS.\footnote{Fadeout, i.e. disappearance of average effects on test-scores over time, need not invalidate their use as $S$. Here, distributional effects are relevant, not only mean effects, as captured by $\gamma$ (see also bitler2014experimental). More generally, even if there are no distributional effects on $S$, $m_d$ may vary across $d$ for the scores.} Following carneiro2014long, I construct the observational sample by selecting individuals who were either eligible for Head Start or reported having participated in the program. This yields an observational sample of size $n_O=2{,}535$. (ref) provides additional details on data construction.

(ref) presents descriptive statistics for the two samples. They have comparable gender compositions and similar shares of white and non-white individuals. Moreover, both NLSY79 and HSIS were designed to be nationally representative, and both pertain to the same treatment, lending support to Assumption (ref). While lower than in the experiment, a significant proportion of CNLSY individuals participate in the program. The rate of compliance with the assigned treatment is $83.8\%$ in the experiment. Imperfect compliance and the availability of the treatment in the observational population preclude direct application of existing data-combination methods that assume either perfect compliance or that the treatment is available only in the experiment. Nevertheless, the method developed here can still be used to estimate the effects under their corresponding identifying assumptions. I illustrate this in the following section by reporting bounds under relevant restrictions following from athey2025combining and garcia2020quantifying.\footnote{Both assume perfect compliance. In garcia2020quantifying, the intervention is available only in the experiment, which is typical of “model” early childhood intervention programs such as the Carolina Abecedarian and Perry Preschool Projects, but not of large-scale programs such as Head Start.} Despite imperfect compliance, the estimated lower bounds on $\gamma_d$ are close to summing to one, suggesting that the combined data are highly informative about the short-term potential outcome distributions.\footnote{In the empirical implementation, I fix $\gamma_d$ at these bounds. This linearizes the optimization programs and substantially simplifies computation without resorting to coarse discretization; see Appendix (ref).}

table[table omitted — 2,313 chars of source]

Results

(ref) reports bound estimates using the consistent estimator proposed in (ref) under previously discussed assumptions. Worst-case bounds impose no restrictions on the temporal link functions $m$, coincide with the bounds of manski1990nonparametric, as shown in (ref), and therefore necessarily include zero. Thus, identifying even the sign of the effect requires additional assumptions. For all outcomes, each of the previously mentioned assumptions substantially reduces the width of the estimated bounds.

table[table omitted — 2,220 chars of source]

Assumption (ref) maintains that temporal link functions are monotonic in each of the two test-score components. For high school graduation and earnings, I impose weak monotonicity increasing in each potential test score. For grade repetition, learning disability diagnosis, “idleness”, crime, teen pregnancy, and poor health, I impose weak monotonicity decreasing in each potential test score. This assumption may be particularly appealing for outcomes related to education, employment, crime, and earnings. Taking high school graduation as an example, it would hold if, fixing any Head Start participation, individuals with higher math or reading scores are equally or more likely to graduate from high school than individuals with lower scores.

Estimates under Assumption (ref) reveal the sign of the effect for all but two outcomes. They further indicate that Head Start participation reduces the probability of grade repetition by at least $1.1$ percentage points (pp), idleness by at least $1.5$ pp, criminal involvement by at least $1.2$ pp, and reporting poor health by at least $2.9$ pp. The corresponding upper bounds are also quite informative. Overall, the estimates suggest that beneficial effects on grade repetition, learning disability, high school graduation, idleness, and poor health may be more modest than reported by the sibling study in deming2009early. Compared to the same study, the effect on criminal involvement found here has the expected sign, while the impact on teen pregnancy does not. Moreover, because the present data allow for longer follow-up than in deming2009early, they reveal earnings for a larger fraction of individuals. However, the sign of the effect on earnings remains unidentified.

Next, I turn to illustrative results relying on temporal link function restrictions that follow from previously proposed methods. These methods cannot be directly utilized here because of imperfect compliance and the availability of the treatment in CNLSY. However, as argued above, the proposed framework can be combined with the relevant assumptions, thereby extending their applicability. Estimates under Assumption (ref) also lead to a substantial reduction in the width of the bounds, though none exclude zero. Assumption (ref) would hold if the model proposed by garcia2020quantifying is credible. That said, HSIS does not contain all of the short-term outcomes included by the authors in their model of the Carolina Abecedarian and CARE programs, nor the medium-term experimental outcomes they use to validate their choice of short-term outcomes. The corresponding results should therefore be interpreted with caution. The final column reports estimates under Assumption LUC from athey2025combining. Since this assumption point-identifies $m$ and the $\gamma_d$ are fixed by empirical bounds, the estimated bounds for the LTE collapse to points. However, the estimated signs are not aligned with previous findings for some outcomes. As emphasized by imbens2024long, one reason for this may be that the estimates are biased by so-called long-term confounders--unobservables that relate Head Start participation, the short-term test scores, and the long-term outcome.

Conclusion

Recent literature proposes augmenting long-term observational studies with short-term experiments to provide alternatives to conventional long-term observational studies. This paper shows that such data combination is not a substitute for credible modeling assumptions. Nevertheless, it remains appealing for this purpose. Assumptions relating short-term to long-term potential outcomes may be defensible based on economic theory or intuition, and thus conducive to plausible inference. Data combination may be used to amplify the identifying power of such restrictions and thereby may yield more informative plausible inference than observational data alone.

This paper introduces two assumptions that exploit this feature of data combination. It also provides a general identification approach that enables computational derivation of bounds under new modeling assumptions, facilitating further developments. Tailor-made assumptions that are plausible in specific empirical settings are an interesting topic for future research, which may benefit from these results.

\printbibliography