Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
93,120 characters · 20 sections · 67 citation commands
The Identification Power of Combining Experimental and Observational Data for Distributional Treatment Effect Parameters
\setstretch{1.25}
\setstretch{1.7}
Researchers often have access to both experimental and observational data when evaluating public policies and medical interventions. Experimental studies, such as randomized controlled trials, are regarded as the gold standard for causal inference, whereas observational data are far more prevalent in practice and often contain rich behavioral variation. This study investigates the identification gains that result from combining these two data sources, with a focus on distributional treatment effects (DTEs).
We consider a general setting in which treatment receipt and outcomes are observed from two data sources: experimental data, where the treatment is randomly assigned, and observational data, where the treatment is self-selected. This structure arises naturally in various policy and medical contexts. In development policy, for example, the efficacy of public health interventions such as mosquito nets is evaluated through field experiments (where nets are randomly assigned) and observational data (where nets are self-purchased). In education, randomized class-size assignments (e.g., Project STAR) coexist with parental or administrative decisions regarding classroom placement. In clinical medicine, drugs may be randomly assigned in clinical trials; however, once approved, they are prescribed based on a physician’s judgment.
Experimental data are valuable because the random assignment of treatment enables identification of key causal parameters such as the average treatment effect (ATE). However, beyond the ATE, DTE parameters are also crucial for policy and medical evaluation, as well as for understanding treatment effect heterogeneity. Let $Y_1$ and $Y_0$ denote potential outcomes under treatment and no-treatment, respectively. Examples of such parameters include: (i) the fraction of individuals who benefit from treatment, $\mathbb{P}(Y_1 > Y_0)$; (ii) the distribution function of the individual treatment effect, $F_{Y_1 - Y_0}(\delta) \equiv \mathbb{P}(Y_1 - Y_0 \leq \delta)$; (iii) the ATE for disadvantaged individuals, $\mathbb{E}[Y_1 - Y_0 \mid Y_0 \leq c]$, where the subpopulation with $Y_0 \leq c$ represents individuals whose baseline outcomes fall below a threshold level $c$ (e.g., a poverty line); and (iv) the correlation between $Y_1$ and $Y_0$.
Although these parameters are highly relevant for policy and medical evaluations, they are generally not point-identified from experimental data alone because their identification requires knowledge of the joint distribution of the two potential outcomes. Experimental data identify only the marginal distributions of $Y_1$ and $Y_0$ and hence yield partial identification (e.g., heckman1997making,fan2010sharp). However, the resulting identified sets are often wide and not very informative.
This study explores whether and how combining experimental and observational data can improve the identification of distributional parameters. Specifically, we address the following questions: (i) Does combining the two data sources shrink the identified set for DTE parameters? (ii) If so, what mechanisms drive this shrinkage, and under what conditions does it occur? (iii) How can the identified set under the combined data be characterized and computed?
To address these questions, we build on a framework that incorporates both experimental and observational data. In this setting, researchers observe individuals' outcomes $Y$, treatment status $D \in \{0,1\}$, and a data source indicator $G \in \{\mathrm{exp},\mathrm{obs}\}$, indicating whether an observation comes from an experimental study ($G=\mathrm{exp}$) or an observational study ($G=\mathrm{obs}$). A key element of the framework is the latent self-selection type $S \in \{0,1\}$, which represents the treatment choice an individual would make under self-selection. This variable is observed in the observational data (where $S = D$) but unobserved in the experimental data.
We define the identified sets under experimental data alone and under the combined data, and use a copula-based approach to characterize and compare them. Under standard assumptions---random treatment assignment in the experimental data and external validity across data sources---this framework enables a systematic comparison of what can be learned from experimental data alone versus combining both data sources.
Our central theoretical results show that the identified set under the combined data is generally smaller than that under the experimental data alone, except in special cases where the latent self-selection variable $S$ is independent of the potential outcomes. Whereas the experimental data identify only the marginal distributions $F_{Y_1}$ and $F_{Y_0}$ of the potential outcomes, we show that combining the two data sources enables identification of the joint distributions $(F_{Y_1S}, F_{Y_0S})$ of each potential outcome and the self-selection type. We further show that the distribution pairs $(F_{Y_1}, F_{Y_0})$ and $(F_{Y_1S}, F_{Y_0S})$ fully characterize the identified sets under the experimental and combined data, respectively, thereby revealing a new source of identifying power arising from self-selection.
This additional identifying power arises because the latent self-selection variable $S$ encodes information about the dependence structure between the potential outcomes. For example, when selection depends on potential gains from treatment, $S$ is systematically related to $(Y_1, Y_0)$ and therefore provides information beyond the marginal distributions. As a result, knowledge of $(F_{Y_1S}, F_{Y_0S})$ imposes additional restrictions on the feasible joint distributions of $(Y_1, Y_0)$, thereby tightening the identified set.
To quantify this identification gain and derive nonparametric sharp bounds, we apply copula bound analysis that builds on and extends the work of fan2017partial, who study the identifying power of covariates, to settings with the self-selection variable $S$. For broad classes of DTE parameters represented by supermodular functions or $\varphi$-indicator functions (defined later), we derive analytical expressions for the sharp bounds and establish necessary and sufficient conditions under which combining the two data sources yields tighter identified sets. We further show that identification improves when the dependence structure between the potential outcomes varies across latent self-selection types.
To further tighten the identified sets, we also consider additional structural restrictions that are plausible in many empirical contexts. In particular, we focus on two such restrictions: positive dependence between the potential outcomes joe2014dependence,frandsen2021partial and a generalized Roy selection model. Though not required for our main results, these restrictions can substantially narrow the identified sets. To incorporate these restrictions under the combined data, we develop a linear programming approach that efficiently computes sharp bounds for a broad class of DTE parameters.
Our analysis also extends to data from doubly randomized preference trials (DRPTs) rucker1989two, long2008causal, a unique yet increasingly prevalent design. In DRPTs, individuals are randomly assigned to one of three groups: a treatment group, a control (no-treatment) group, or a self-selection group in which individuals choose between treatment and no-treatment. This design has been used in a wide range of studies across the social sciences gaines2011experimental, arceneaux2012polarized, DeBenedictis_2019, ida2024choosing and medical sciences king2005impact. An established advantage of DRPTs is that they enable identification of the ATE for each self-selection group, $\mathbb{E}[Y_1 - Y_0 \mid S = s]$ for $s \in \{0,1\}$ long2008causal. Our results highlight a novel advantage of this design: by combining the random-assignment and self-selection samples, it sharpens identification of DTE parameters.
We illustrate the empirical relevance of our approach using DRPT data from gaines2011experimental, who study the effects of negative campaign advertisements on individuals’ attitudes toward presidential candidates during the 2008 U.S. presidential election. We find that incorporating self-selection data substantially tightens the identified sets for DTE parameters such as $\mathbb{P}(Y_1 < Y_0)$, the fraction of individuals negatively affected by the advertisements. The results also indicate that the advertisements have substantial but heterogeneous effects. These findings underscore the empirical value of combining random-assignment and self-selection data to estimate DTE parameters.
This study relates to two strands of literature: (i) combining random-assignment and self-selection data for causal inference, and (ii) partial or point identification of DTE parameters.
The first strand examines the benefits of combining data from random-assignment and self-selection sources. This literature typically pursues two objectives: (1) improving the precision of estimates when experimental data already suffice for identification, and (2) identifying causal parameters that are not identifiable from experimental data alone. For the first objective, rosenman2023combining, yang2023elastic, and gui2024combining propose methods that improve the efficiency of treatment effect estimation by combining the two data sources. For the second, long2008causal show that the ATE for each self-selection group, $\mathbb{E}[Y_1-Y_0 \mid S=s]$ for $s \in \{0,1\}$, can be identified using data from a DRPT. knox2019design extend this to multiple treatment settings and derive partial identification results. In a different context, where experimental data contain secondary outcomes and observational data contain primary outcomes, athey2020combining study identification of the ATE for the primary outcome.\footnote{Related contributions include athey2025surrogate and rambachan2024program, which study data combination settings with surrogate, secondary, or missing outcomes.} Our setting differs from theirs in that both data sources contain the same type of outcome.
Our contribution to this strand is to uncover a previously unrecognized advantage of combining these two data sources: it can tighten the identified sets for DTE parameters. Unlike prior studies, which focus on point identification of new parameters or gains in estimation efficiency, we show that data combination can also improve informativeness by tightening bounds in partially identified settings.
The second strand studies the partial or point identification of DTE parameters under various assumptions and structural restrictions.\footnote{Notable contributions include heckman1997making, manski1997monotone, fan2010sharp (fan2010sharp), fan2010partial, fan2014identifying, fan2017partial, vuong2017counterfactual, kim2018identification, firpo2019partial, callaway2021bounds, frandsen2021partial, russell2021sharp, cui2025policy, and kaji2023assessing, among others.} In particular, fan2010sharp (fan2010sharp,fan2012confidence), fan2017partial, and firpo2019partial develop nonparametric bounds for the joint distribution of the potential outcomes and its functionals under minimal assumptions such as random treatment assignment. To tighten these bounds, subsequent studies have imposed additional structure, including time-dependence restrictions in panel settings callaway2021bounds and mutual stochastic monotonicity of the potential outcomes frandsen2021partial.
This study contributes to this literature by clarifying the identifying power that arises from combining experimental and observational data, a structure that has not previously been explored in this context, and by providing computable characterizations of the sharp bounds under data combination.
Finally, this study contributes to the broader causal inference literature by offering a new perspective on the role of observational data in identification. Although observational data are often viewed as redundant for identification once experimental data are available, we show that they can still sharpen identification of distributional parameters.
The remainder of the paper is organized as follows. Section (ref) introduces the data combination framework, formalizes the DTE parameters, and defines their identified sets under experimental data alone and under the combined data. Section (ref) characterizes the identified sets under each data scenario and highlights the sources of identification gains from the combined data. Section (ref) derives sharp bounds for broad classes of DTE parameters, specifically those represented by supermodular or $\varphi$-indicator functions, and establishes necessary and sufficient conditions under which data combination improves identification. Section (ref) provides numerical examples illustrating the identifying power of data combination. Section (ref) introduces additional restrictions and develops a linear programming approach for computing sharp bounds. Section (ref) presents an empirical illustration using DRPT data from gaines2011experimental. Section (ref) concludes. All proofs are provided in the appendix.
We begin by outlining the framework. Section (ref) introduces the data combination setting and fundamental assumptions. Section (ref) defines the DTE parameters and provides illustrative examples. Section (ref) defines the identified sets under experimental and combined data.
We consider observed data consisting of the quadruple $(Y, D, G, X)$, where $Y$ is the outcome; $D \in \{0,1\}$ is a binary treatment indicator; $G \in \{\mathrm{exp}, \mathrm{obs}\}$ denotes the data source; and $X$ is a vector of pre-treatment covariates with its support denoted by $\mathcal{X}$. Specifically, $G$ indicates whether an observation comes from an experimental study ($G = \mathrm{exp}$) or an observational study ($G = \mathrm{obs}$). The treatment indicator $D$ equals $1$ if the individual receives treatment and $0$ otherwise. In the experimental data ($G = \mathrm{exp}$), $D$ is randomly assigned, whereas in the observational data ($G = \mathrm{obs}$), it is determined through self-selection. Let $Y_1$ and $Y_0$ denote the potential outcomes under treatment and no-treatment, respectively, both of which are assumed to be continuous. The observed outcome is defined as $Y \equiv DY_1 + (1 - D)Y_0$.
We introduce a latent self-selection variable $S \in \{0,1\}$, defined as the treatment an individual would choose under self-selection. In the observational data ($G = \mathrm{obs}$), $S$ coincides with the observed treatment, i.e., $S = D$, whereas in the experimental data ($G = \mathrm{exp}$), $S$ is unobserved. The variable $S$ can also be interpreted as a latent preference for treatment versus no-treatment.
Let $F^{\ast}$ denote the true cumulative distribution function (CDF) of all defined variables $(Y_1, Y_0, Y, D, S, G, X)$, including both observed and unobserved ones. For a generic CDF $F$, we denote the probability and expectation under $F$ by $\mathbb{P}_{F}(\cdot)$ and $\mathbb{E}_{F}[\cdot]$, respectively. For the true CDF $F^{\ast}$, we use the shorthand $\mathbb{P}(\cdot)$ and $\mathbb{E}[\cdot]$ in place of $\mathbb{P}_{F^{\ast}}(\cdot)$ and $\mathbb{E}_{F^{\ast}}[\cdot]$.
Throughout the paper, we suppose that $F^{\ast}$ satisfies the following assumptions.
Assumption (ref) states that self-selection corresponds to the received treatment in the observational data, which is trivially satisfied by the definition of $S$. Assumption (ref) requires that treatment is randomly assigned in the experimental study, possibly conditional on $X$.
Assumption (ref) concerns the external validity of the data sources and requires that any systematic differences between the populations in the experimental and observational studies be captured by $X$. Although this assumption is automatically satisfied under a DRPT design, where $G$ is randomly assigned, it can also be plausible in empirical settings involving data combination. For example, in education research, experimental data from randomized interventions such as Project STAR are often combined with administrative or survey data from the same school system chetty2011does,athey2020combining, where differences between samples are largely driven by observable characteristics such as demographics, prior achievement, or school characteristics. Conditioning on observed covariates can therefore help make the populations comparable across data sources.
Assumption (ref) imposes an overlap condition on both the treatment status and data source.\footnote{This overlap condition can be relaxed at the cost of yielding wider identified sets under the combined data. For example, if a covariate value $x$ does not satisfy $\mathbb{P}(D = d, G = g \mid X = x) > 0$ for some $g$, then experimental and observational data cannot be combined at that value of $x$. In this case, one may still rely on either the experimental or the observational data alone to construct an identified set conditional on $x$.}
We consider a parameter of interest $\theta_{o} \in \mathbb{R}$ that has the following form:
where $\psi : \mathbb{R}^2 \to \mathbb{R}$ is a given function. With various specifications of $\psi$, this formulation encompasses a wide range of DTE parameters, as illustrated below.
See also fan2017partial for additional examples.
Note that none of the DTE parameters introduced above can be point-identified even with experimental data. Their identification generally requires knowledge of the joint distribution of the two potential outcomes, $F_{Y_1 Y_0}^{\ast}$, whereas experimental data identify only the marginal distributions, $(F_{Y_1}^{\ast}, F_{Y_0}^{\ast})$.
We can also define DTE parameters for various subpopulations: by self-selection type, $\theta_{o,s} \equiv \mathbb{E}[\psi(Y_1, Y_0) \mid S = s]$; conditional on covariates, $\theta_{o,x} \equiv \mathbb{E}[\psi(Y_1, Y_0) \mid X = x]$; and by data source, $\theta_{o,g} \equiv \mathbb{E}[\psi(Y_1, Y_0) \mid G = g]$ for $g \in \{\mathrm{exp}, \mathrm{obs}\}$. In particular, $\theta_{o,s}$ captures treatment effects for individuals who would self-select into treatment ($s=1$) or no-treatment ($s=0$), a quantity of interest in many empirical applications.
These subpopulation parameters can be incorporated into our framework with slight modifications. For example, the parameter $\theta_{o,g}$ can be expressed as $$\theta_{o,g} = \mathbb{E}\left[\psi(Y_1,Y_0)\cdot \frac{\mathbf{1}\{G=g\}}{\mathbb{P}(G=g)}\right],$$ where $\mathbb{P}(G=g)$ is point-identified from the observed data. Defining $\tilde{\psi}(y_1,y_0) = \psi(y_1,y_0)\cdot \mathbf{1}\{G=g\}/\mathbb{P}(G=g)$, we can write $\theta_{o,g} = \mathbb{E}[\tilde{\psi}(Y_1,Y_0)]$ and thus analyze it in the same way as $\theta_o$. As for $\theta_{o,s}$, Lemma (ref) in the appendix shows that $$\theta_{o,s} = \mathbb{E}[\psi(Y_1,Y_0)| D=s, G= \mathrm{obs}] = \mathbb{E}\left[\psi(Y_1,Y_0)\cdot \frac{\mathbf{1}\{D=s,G=\mathrm{obs}\}}{\mathbb{P}(D=s,G=\mathrm{obs})}\right].$$ Hence, $\theta_{o,s}$ can also be handled in the same manner as $\theta_{o}$.
In what follows, we study the partial identification of $\theta_o$ using experimental data alone and combined experimental and observational data.
We formally define the identified sets for the DTE parameter $\theta_o$ under two data scenarios: (i) experimental data alone and (ii) combined experimental and observational data. We begin with the case of experimental data alone.
Let $\mathcal{F}^{\dagger}$ denote the class of CDFs for all defined variables $(Y_1, Y_0, Y, D, S, G, X)$ that satisfy Assumptions (ref)--(ref).\footnote{Formally, $\mathcal{F}^{\dagger}$ is the set of all CDFs $F$ of $(Y_1, Y_0, Y, D, S, G, X)$ that satisfy Assumptions (ref)--(ref) with $F^{\ast}$ replaced by $F$.} We begin by defining the identified set for the joint CDF $F^{\ast}$ under experimental data as
where $F_{YDX|G}^{\ast}(\cdot,\cdot,\cdot|\mathrm{exp})$ is the distribution of the observables $(Y, D, X)$ in the experimental data, and $F_X^{\ast}$ is the marginal distribution of $X$ in the entire population. Thus, $\mathcal{F}_{\mathrm{exp}}^{\ast}$ consists of all CDFs that satisfy the maintained assumptions and are consistent with the distribution of the observed experimental data and the population distribution of $X$.
We assume that $F_X^{\ast}$ is known, as this is needed to account for potential imbalances in the covariate distributions between the experimental and observational data and to enable meaningful comparisons of the identified sets. When the two data sources are drawn from the same population (i.e., covariates are balanced), this assumption is unnecessary, since $F_X^{\ast}$ can be identified from the experimental data as $F_X^{\ast}(\cdot) = F_{X \mid G}^{\ast}(\cdot \mid \mathrm{exp})$. In other cases, covariate information for the entire population is often available from external sources such as demographic datasets. Moreover, when the parameter of interest is the covariate-conditional effect $\theta_{o,x}$ or the effect $\theta_{o,\mathrm{exp}}$ defined for the experimental population, knowledge of $F_X^{\ast}$ is not required.
Given the identified set of CDFs $\mathcal{F}_{\mathrm{exp}}^{\ast}$, the identified set for $\theta_o$ under experimental data is defined as
This set consists of all parameter values that are attainable under some distribution in $\mathcal{F}_{\mathrm{exp}}^{\ast}$. Any parameter value outside $\Theta_{I}$ is incompatible with the experimental data or the maintained assumptions.
We next consider the case in which both experimental and observational data are available. We first define the identified set for the joint CDF $F^{\ast}$ under the combined data, analogously to $\mathcal{F}_{\mathrm{exp}}^{\ast}$, as
where $F_{YDGX}^{\ast}$ is the joint distribution of the observables $(Y, D, G, X)$ in the combined data. This set consists of all CDFs that satisfy the maintained assumptions and are consistent with the observed distribution $F_{YDGX}^{\ast}$. By construction, $\mathcal{F}^{\ast} \subseteq \mathcal{F}_{\mathrm{exp}}^{\ast}$, since the condition $F_{YDGX} = F_{YDGX}^{\ast}$ in $\mathcal{F}^{\ast}$ implies both $F_{YDX|G}(\cdot,\cdot,\cdot \mid \mathrm{exp}) = F_{YDX|G}^{\ast}(\cdot,\cdot,\cdot \mid \mathrm{exp})$ and $F_{X}=F_{X}^{\ast}$.
Given the identified set of CDFs $\mathcal{F}^{\ast}$, the identified set for $\theta_o$ under the combined data is defined as
This set consists of all parameter values that are attainable under some CDF in $\mathcal{F}^{\ast}$. Any value outside $\Theta_{IC}$ contradicts the maintained assumptions or the observed combined data.
Since $\mathcal{F}^{\ast} \subseteq \mathcal{F}_{\mathrm{exp}}^{\ast}$, it follows that $\Theta_{IC} \subseteq \Theta_{I}$; that is, the identified set under the combined data is no larger than that under experimental data alone. Our primary interest, however, is whether the strict inclusion $\Theta_{IC} \subset \Theta_{I}$ holds. This is equivalent to asking whether combining the two data sources strictly tightens the identified set for $\theta_o$.
This question is nontrivial because a strict inclusion $\mathcal{F}^{\ast} \subset \mathcal{F}_{\mathrm{exp}}^{\ast}$ at the level of CDFs does not necessarily imply a strict inclusion $\Theta_{IC} \subset \Theta_{I}$ at the parameter level. We investigate this question in the following sections.
To investigate whether the strict inclusion $\Theta_{IC} \subset \Theta_{I}$ holds, the definitions of $\Theta_{I}$ and $\Theta_{IC}$ in equations ((ref)) and ((ref)) are too abstract to yield direct insight. We therefore seek a more interpretable characterization of these identified sets. To this end, we employ the concept of bivariate copulas sklar1959fonctions, which serves as a central tool in this and subsequent sections.
We begin with the case of using experimental data alone. Under randomized treatment assignment (Assumption (ref)) and external validity (Assumption (ref)), the conditional marginal distributions of the potential outcomes given $X$, $(F_{Y_1 \mid X}^{\ast}, F_{Y_0 \mid X}^{\ast})$, are identified as $F_{Y_d \mid X}^{\ast}(\cdot \mid x) = F_{Y \mid DGX}^{\ast}(\cdot \mid d, \mathrm{exp}, x)$ for $d = 0,1$. Since $F_X^{\ast}$ is assumed to be known, the joint distributions $(F_{Y_1 X}^{\ast}, F_{Y_0 X}^{\ast})$ are also identified.
Let $\mathcal{C}$ denote the class of all bivariate copula functions. For each $x \in \mathcal{X}$, let $C^{\ast}(\cdot, \cdot \mid x) \in \mathcal{C}$ denote the true conditional copula given $X=x$, which reproduces the true conditional joint distribution $F_{Y_1 Y_0 \mid X}^{\ast}(y_1, y_0 \mid x)$ from the marginals $(F_{Y_1 \mid X}^{\ast}(y_1 \mid x), F_{Y_0 \mid X}^{\ast}(y_0 \mid x))$ as \[ F_{Y_1 Y_0 \mid X}^{\ast}(y_1, y_0 \mid x) = C^{\ast}\!\left(F_{Y_1 \mid X}^{\ast}(y_1 \mid x), F_{Y_0 \mid X}^{\ast}(y_0 \mid x) \mid x\right), \] where the existence of such a copula is guaranteed by Sklar's theorem (e.g., nelsen2006introduction, nelsen2006introduction, Theorem 2.3.3).\footnote{Sklar's theorem sklar1959fonctions states that any joint distribution $F_{Y_1 Y_0}$ can be expressed as a copula of its marginals: $F_{Y_1 Y_0}(y_1, y_0) = C\bigl(F_{Y_1}(y_1), F_{Y_0}(y_0)\bigr)$ for some copula function $C$. Conversely, given any marginal distributions $F_{Y_1}$ and $F_{Y_0}$, the function $C(F_{Y_1}(y_1), F_{Y_0}(y_0))$, for any copula $C$, defines a valid bivariate distribution with those marginals.} Using the true conditional copula $C^{\ast}(\cdot, \cdot \mid x)$, the parameter $\theta_o$ can be expressed as
The true copula $C^{\ast}(\cdot, \cdot \mid x)$ is unknown. However, by allowing $C(\cdot, \cdot \mid x)$ to vary over all copula functions in $\mathcal{C}$, we obtain the identified set for $\theta_o$ based on $(F_{Y_1 X}^{\ast}, F_{Y_0 X}^{\ast})$ as
By Sklar's theorem, the collection $\bigl\{C\bigl(F_{Y_1|X}^{\ast}(\cdot|x),F_{Y_0|X}^{\ast}(\cdot|x) \big| x\bigr): C(\cdot,\cdot|x) \in \mathcal{C}\bigr\}$ coincides with the set of all conditional joint CDFs of $(Y_1,Y_0)$ given $X=x$ that share the marginals $\bigl(F_{Y_1|X}^{\ast}(\cdot|x),F_{Y_0|X}^{\ast}(\cdot|x)\bigr)$.
The following proposition shows that the identified set $\widetilde{\Theta}_{I}$, based on the distributions $(F_{Y_1 X}^{\ast}, F_{Y_0 X}^{\ast})$, coincides with the identified set $\Theta_{I}$ under experimental data.
This result formally confirms that $\Theta_{I}$ is fully characterized by the marginals $(F_{Y_1 X}^{\ast}, F_{Y_0 X}^{\ast})$.\footnote{While many studies refer to $\widetilde{\Theta}_{I}$ as the identified set under experimental data, Proposition (ref) provides the formal justification for this equivalence.}
We next seek to characterize the identified set $\Theta_{IC}$ under the combined data. Since the experimental data identify $(F_{Y_1 X}^{\ast}, F_{Y_0 X}^{\ast})$, our starting point is to examine what additional information can be gained by incorporating the observational data. The following lemma addresses this question.
Lemma (ref) shows that combining the two data sources enables identification of the joint distributions $\left(F_{Y_1 S X}^{\ast}, F_{Y_0 S X}^{\ast}\right)$ of each potential outcome, the latent self-selection type, and the covariates. The combined data therefore provide richer information than experimental data alone, which identify only $\left(F_{Y_1 X}^{\ast}, F_{Y_0 X}^{\ast}\right)$.
In particular, for $d \neq s$, $F_{Y_d \mid S X}^{\ast}(\cdot \mid s, x)$ is a counterfactual distribution and is not identifiable from either the experimental or the observational data alone. However, it becomes identifiable when the two data sources are combined, through the following decomposition:
for $d \neq s$, where $F_{Y_d \mid X}^{\ast}(\cdot \mid x)$ is identified from the experimental data, while $\mathbb{P}(S=\cdot \mid X=x)$ and $F_{Y_d \mid SX}^{\ast}(\cdot \mid d,x)$ are identified from the observational data (see the proof for details).\footnote{long2008causal show a related result: the ATE $\mathbb{E}[Y_1 - Y_0 \mid S=s]$ for each self-selection group $s \in \{0,1\}$ is identified using data from a DRPT.}
Let $C^{\ast}(\cdot,\cdot|s,x)$ denote the true conditional copula given $S=s$ and $X=x$; that is,
Then the true parameter value $\theta_o$ can be expressed as
Although the true conditional copula $C^{\ast}(\cdot, \cdot \mid s, x)$ is unknown, allowing $C(\cdot,\cdot \mid s,x)$ to vary over $\mathcal{C}$ yields the identified set for $\theta_o$ based on $(F_{Y_1 S X}^{\ast}, F_{Y_0 S X}^{\ast})$:
By Sklar's theorem, for any $C \in \mathcal{C}$, the function $C\big(F_{Y_1|SX}^{\ast}(y_1 \mid s,x), F_{Y_0|SX}^{\ast}(y_0 \mid s,x) \mid s,x\big)$ defines a valid conditional joint CDF of $(Y_1, Y_0)$ given $S=s$ and $X=x$, with marginals $F_{Y_1|SX}^{\ast}(\cdot \mid s,x)$ and $F_{Y_0|SX}^{\ast}(\cdot \mid s,x)$.
We now ask whether $\widetilde{\Theta}_{IC}$ coincides with the identified set $\Theta_{IC}$ under the combined data. This question can be rephrased as whether $\Theta_{IC}$ is fully characterized by the self-selection joint distributions $\bigl(F_{Y_1 S X}^{\ast}, F_{Y_0 S X}^{\ast}\bigr)$. This question is nontrivial because combining the two data sources might yield additional identifying information beyond these joint distributions.
The following theorem provides a central characterization result for $\Theta_{IC}$, showing that combining the two data sources yields no additional identifying information beyond $(F_{Y_1 S X}^{\ast}, F_{Y_0 S X}^{\ast})$. This result is not immediate, as it requires ruling out additional restrictions on the dependence structure.
In summary, the identified set $\Theta_I$ under experimental data is fully characterized by $(F_{Y_0 X}^{\ast}, F_{Y_1 X}^{\ast})$ (Proposition (ref)), whereas the identified set $\Theta_{IC}$ under the combined data is fully characterized by the self-selection joint distributions $(F_{Y_0 S X}^{\ast}, F_{Y_1 S X}^{\ast})$ (Theorem (ref)). This contrast highlights the additional identifying power of the combined data, which arises from the dependence between $(Y_1, Y_0)$ and the latent self-selection type $S$. The following section examines this mechanism in greater detail.
In this section, we derive sharp bounds for the DTE parameter $\theta_o$ under both experimental and combined data. Whereas fan2017partial study the identifying power of covariates for DTEs, our analysis exploits a distinct source of identifying power arising from the latent self-selection type $S$. In particular, we show that combining the two data sources leverages the dependence between $(Y_1, Y_0)$ and $S$, as characterized in Theorem (ref), to yield strictly tighter bounds.
We consider two classes of DTE parameters, according to whether $\psi$ is specified as (i) a supermodular function or (ii) a $\varphi$-indicator function. For each case, we derive closed-form expressions for the sharp bounds on $\theta_o$ under both data settings. We also establish necessary and sufficient conditions under which combining the two data sources strictly tightens the identified set, i.e., $\Theta_{IC} \subset \Theta_{I}$.
We begin with DTE parameters represented by supermodular and submodular functions.
The functions $\psi(\cdot, \cdot)$ in Examples (ref), (ref), and (ref) are either supermodular or submodular and thus fall within this framework. In particular, the function $\psi$ in Example (ref) is strict supermodular. cambanis_1976 provide many other examples of supermodular and submodular functions.
We begin by characterizing $\Theta_{I}$ using the Fréchet–Hoeffding bounds, taking fan2017partial as a benchmark. Define
where $M(u,v) \equiv \max(u + v - 1, 0)$ and $W(u,v) \equiv \min(u,v)$ denote the Fréchet–Hoeffding lower and upper bounds, respectively, for a bivariate distribution with marginals $(u,v)$.
When $\psi$ is supermodular, Theorem 3.2 of fan2017partial shows that under certain regularity conditions, the identified set $\widetilde{\Theta}_{I}$ based on $(F_{Y_1X}^{\ast},F_{Y_0X}^{\ast})$ is the interval $\left[\theta_{I}^{L},\theta_{I}^{U}\right]$, where
with $F_{Y_d|X}^{\ast,-1}(u|x) \equiv \inf\bigl\{y: F_{Y_d|X}^{\ast}(y|x) \geq u\bigr\}$ denoting the conditional quantile function of $Y_d$. Hence, by Proposition (ref), the identified set $\Theta_{I}$ under experimental data is given by $\left[\theta_{I}^{L}, \theta_{I}^{U}\right]$.
We next characterize the identified set $\Theta_{IC}$ under the combined data. Define
which extend the Fréchet–Hoeffding bounds to the self-selection setting.
Using these constructions, when $\psi$ is supermodular, the identified set $\widetilde{\Theta}_{IC}$ based on $(F_{Y_1SX}^{\ast}, F_{Y_0SX}^{\ast})$ is given by the interval $\left[\theta_{IC}^{L}, \theta_{IC}^{U}\right]$, where
with $F_{Y_d|SX}^{\ast,-1}(u|s,x) \equiv \inf\{y: F_{Y_d|SX}^{\ast}(y|s,x) \geq u\}$ denoting the conditional quantile function.
Hence, by Theorem (ref), the identified set $\Theta_{IC}$ under the combined data is equal to $\left[\theta_{IC}^{L},\theta_{IC}^{U}\right]$. We formalize these results in the following proposition.
Proposition (ref)(ii) provides a new characterization of the sharp bounds for $\theta_{o}$ under the combined data setting when $\psi$ is supermodular. Together with Lemma (ref), it yields a constructive representation of the bounds via ((ref)) and ((ref)), combined with ((ref)) and ((ref)). This representation enables direct computation of the sharp bounds and provides a basis for inference.\footnote{Pointwise valid confidence sets for $\theta_{o}$ can be constructed by extending the procedure of fan2017partial (fan2017partial, Appendix B) using nonparametric estimators of $(F_{Y_1SX}^{\ast}, F_{Y_0SX}^{\ast})$ obtained from Lemma (ref).}
The key difference between the bounds $\theta_{I}^{L}$ ($\theta_{I}^{U}$) and $\theta_{IC}^{L}$ ($\theta_{IC}^{U}$) lies in the inclusion of the latent self-selection variable $S$ in the conditioning sets. This allows $S$ to capture information about the dependence between $Y_1$ and $Y_0$ (e.g., under Roy-type selection), thereby tightening the identified set.
To compare the two identified sets, note that $\Theta_I$ is characterized by joint distributions of $(Y_1,Y_0)$ whose marginals are bounded by $F_{I}^{*,(-)}$ and $F_{I}^{*,(+)}$, whereas $\Theta_{IC}$ is characterized by the tighter bounds $F_{IC}^{*,(-)}$ and $F_{IC}^{*,(+)}$, which incorporate $S$. By Jensen's inequality,
and similarly $F_{IC}^{*,(+)}(y_1,y_0) \leq F_{I}^{*,(+)}(y_1,y_0)$. Therefore, $\theta_{I}^{L} \le \theta_{IC}^{L}$ and $\theta_{IC}^{U} \le \theta_{I}^{U}$, so incorporating $S$ weakly tightens the identified set.
For strict supermodular functions $\psi$, Theorem (ref) below establishes a necessary and sufficient condition under which $\Theta_{IC} = \Theta_{I}$; otherwise, combining the two data sources yields a strictly smaller identified set (i.e., $\Theta_{IC} \subsetneq \Theta_{I}$).
Conditions ((ref)) and ((ref)) require that, conditional on $X$, the ordering between $F_{Y_1|SX}^{\ast}(y_1|S,X)$ and $F_{Y_0|SX}^{\ast}(y_0|S,X)$ be degenerate, in the sense that it does not vary between the two self-selection types, $S=0$ and $S=1$, for $\psi_c$-almost all $(y_1,y_0)$. These conditions are highly restrictive and unlikely to hold when self-selection depends on the potential outcomes. In such cases, the theorem implies that combining experimental and observational data strictly tightens the identified set.
A notable exception arises under selection-on-observables, i.e., when $(Y_1,Y_0) \perp \!\!\! \perp S \mid X$. In this case, conditions ((ref)) and ((ref)) are satisfied, and combining the two data sources provides no additional identifying power.
We now turn to DTE parameters characterized by $\varphi$-indicator functions.
An important example in this class is the distribution function of the treatment effect, $F_{Y_1 - Y_0}^{\ast}(\delta)$ (Example (ref)), which corresponds to the choice $\varphi(y_1,y_0)=y_1-y_0$. As a special case, the fraction of positive treatment effects (Example (ref)) is given by $\mathbb{P}(Y_1 > Y_0)=1-F_{Y_1-Y_0}^{\ast}(0)$.
We first characterize the identified set $\Theta_{I}$ under experimental data, or equivalently, $\widetilde{\Theta}_{I}$ based on $(F_{Y_1X}^{\ast}, F_{Y_0X}^{\ast})$. Let $\mathcal{Y}_{1}(x)$ and $\mathcal{Y}_{0}(x)$ denote the supports of $Y_1$ and $Y_0$ given $X = x$, respectively. Define
where $\tilde{\varphi}_{y}(\delta|x) = \sup\left\{y_0 \in \mathcal{Y}_{0}(x): \varphi(y,y_{0}) < \delta\right\}$. If the set is empty, we define $\tilde{\varphi}_y(\delta|x) = -\infty$.
When $\psi$ is a $\varphi$-indicator function, fan2017partial show that the identified set $\widetilde{\Theta}_{I}$ based on $(F_{Y_1X}^{\ast}, F_{Y_0X}^{\ast})$ is given by the interval $\left[F_{I,\varphi}^{L}(\delta), F_{I,\varphi}^{U}(\delta)\right]$, where $F_{I,\varphi}^{L}(\delta) = \mathbb{E}[F_{\min,\varphi}(\delta|X)]$ and $F_{I,\varphi}^{U}(\delta) = \mathbb{E}[F_{\max,\varphi}(\delta|X)]$. Proposition (ref) then implies that the identified set $\Theta_{I}$ under experimental data is given by $\left[F_{I,\varphi}^{L}(\delta), F_{I,\varphi}^{U}(\delta)\right]$.
We next characterize the identified set $\widetilde{\Theta}_{IC}$ based on $(F_{Y_1SX}^{\ast}, F_{Y_0SX}^{\ast})$. Let $\mathcal{Y}_{1}(s,x)$ and $\mathcal{Y}_{0}(s,x)$ denote the supports of $Y_1$ and $Y_0$ given $(S,X) = (s,x)$, respectively. Then $\widetilde{\Theta}_{IC}$ is given by the interval $\left[F_{IC,\varphi}^{L}(\delta), F_{IC,\varphi}^{U}(\delta)\right]$, where $F_{IC,\varphi}^{L}(\delta) = \mathbb{E}[F_{\min,\varphi}(\delta|S,X)]$ and $F_{IC,\varphi}^{U}(\delta) = \mathbb{E}[F_{\max,\varphi}(\delta|S,X)]$, with
and $\tilde{\varphi}_{y}(\delta|s,x) = \sup\{y_0 \in \mathcal{Y}_{0}(s,x): \varphi(y,y_{0}) < \delta\}$.
Theorem (ref) then implies that the identified set $\Theta_{IC}$ under the combined data is given by $\left[F_{IC,\varphi}^{L}(\delta), F_{IC,\varphi}^{U}(\delta)\right]$. The key difference from the experimental-data case is the inclusion of the self-selection variable $S$ in (ref) and (ref). Since $S$ may encode information about the dependence between $Y_1$ and $Y_0$, its inclusion can strictly tighten the identified set.
The following proposition summarizes these characterization results.
Proposition (ref)(ii) provides a characterization of the identified set in the combined-data setting for $\varphi$-indicator functions. Combined with Lemma (ref), it also yields a constructive representation of the identified set, enabling direct computation from the combined data.
Analogously to Theorem (ref), for a $\varphi$-indicator function we can establish a necessary and sufficient condition under which $\Theta_{IC}$ is a proper subset of $\Theta_{I}$. To simplify the technical argument, the following theorem presents this result for the case in which $\mathcal{Y}_{d}(s,x)=\mathcal{Y}_{d}$ for $d=0,1$ and all $(s,x)\in\{0,1\}\times\mathcal{X}$.
The condition in Theorem (ref) requires that the locations at which the function (ref) attains its maximum and minimum be invariant to the self-selection variable $S$. This invariance condition is unlikely to hold when self-selection $S$ depends on the potential outcomes $(Y_1,Y_0)$, as outcome-dependent selection typically alters the relative ordering of the conditional distributions between $S=0$ and $S=1$. In such cases, the theorem implies that combining experimental and observational data strictly tightens the identified set.
A sufficient condition for the equivalence $\Theta_{IC} = \Theta_{I}$ in Theorem (ref) is that the self-selection variable $S$ be independent of $(Y_1,Y_0)$ conditional on $X$. Under this selection-on-observables assumption, combining the two data sources provides no additional identifying power.
To illustrate the identifying power of combining experimental and observational data, we consider a simple data-generating process (DGP) for $(Y_1, Y_0, S)$.
Suppose that the parameter of interest is $\theta_o = \mathbb{P}(Y_1 > Y_0)$. Let $\mathbb{P}(S = 1) = \mathbb{P}(S = 0) = 1/2$. Conditional on $S$, the potential outcomes $(Y_1, Y_0)$ follow independent normal distributions:
where $\mu_H > \mu_L$. This specification captures self-selection based on potential gains from treatment: individuals with $S = 1$ tend to have higher potential outcomes under treatment, whereas the reverse holds for those with $S = 0$. For concreteness, we set $(\mu_H, \mu_L, \sigma) = (1.5, -1.5, 1)$, which implies $\theta_o = 0.5$ by symmetry.
Under this DGP, the unconditional marginal distributions of $Y_1$ and $Y_0$ coincide and are given by $F_{Y_1}(y) = F_{Y_0}(y) = \frac{1}{2}\Phi(y-\mu_H) + \frac{1}{2}\Phi(y-\mu_L),$ where $\Phi(\cdot)$ denotes the standard normal CDF. Hence, under experimental data alone, Proposition (ref)(i) implies $\Theta_I = [0,1]$.
In contrast, the combined data identify the marginal distributions of $Y_1$ and $Y_0$ conditional on $S$. In this example, when $S = 1$, the distribution of $Y_1$ first-order stochastically dominates that of $Y_0$, whereas the reverse holds when $S = 0$. This additional information restricts the set of feasible joint distributions of $(Y_1, Y_0)$.
Using Proposition (ref) (ii), the identified set can be computed as
For $(\mu_H, \mu_L) = (1.5, -1.5)$, we obtain $\Theta_{IC} \approx [0.433,0.567]$. Moreover, as $\mu_H - \mu_L$ becomes large relative to $\sigma$, $\Theta_{IC}$ shrinks toward the singleton $\{0.5\}$, whereas $\Theta_I$ remains equal to $[0,1]$.
Thus, while experimental data alone yield only the trivial bounds $[0,1]$, incorporating observational data substantially tightens the identified set. This example highlights the identifying power of combining the two data sources.
In the spirit of partial identification analysis manski2003partial, additional model restrictions can, when plausible, further narrow the identified set for $\theta_o$. At the same time, such restrictions may complicate the analysis, particularly by making it difficult to derive sharp bounds analytically. In this section, we introduce two restrictions that are plausible in many empirical contexts and present a computational approach for computing the sharp bounds under these restrictions using the combined data.
The first restriction imposes a form of positive dependence between the potential outcomes, specifically the mutually left-tail decreasing (LTD) condition joe2014dependence. The potential outcomes $Y_1$ and $Y_0$ are said to be mutually LTD if they satisfy the following condition.
This assumption implies that individuals with higher potential outcomes under one treatment state tend to also have higher potential outcomes under the other state. Such an assumption is plausible in many empirical contexts. For example, in a small-class-size program, students who perform well in either small or regular classes are also likely to perform well under the alternative class size.
frandsen2021partial consider a slightly stronger dependence assumption and show that it can substantially tighten the identified set for the distribution of treatment effects. Related assumptions are also used by chetty2017fading in their empirical study of income mobility and by cui2025policy in policy learning with distributional welfare.
The second restriction concerns the self-selection mechanism for treatment.
This assumption implies that individuals who self-select into treatment are more likely to experience larger treatment effects than those who select no-treatment. We refer to this as the generalized Roy model selection assumption, since it is implied by the selection behavior in the generalized Roy model.
Because this restriction pertains to the self-selection mechanism, it cannot be exploited using experimental data alone. Thus, incorporating observational data provides an additional advantage by enabling the use of such behavioral restrictions on self-selection.
We propose a computational approach to obtain the sharp bounds for $\theta_{o}$ under the combined data, incorporating either or both Assumptions (ref) and (ref). The approach is not restricted to specific classes of objective functions $\psi$, such as supermodular or $\varphi$-indicator functions.
When Assumptions (ref) and (ref) are imposed in addition to Assumptions (ref)--(ref), the sharp lower and upper bounds for $\theta_{o}$ can be obtained by solving the following minimization and maximization problems:
The constraint in ((ref)) follows from Theorem (ref), which shows that the identified set for $\theta_{o}$ is fully characterized by $F_{Y_1SX}^{\ast}$ and $F_{Y_0SX}^{\ast}$. The constraints in ((ref)) and ((ref)) are implied by Assumptions (ref) and (ref), respectively.
Note that this optimization problem is formulated without explicitly including the treatment variable $D$ or the data source indicator $G$, since the constraint in ((ref)) already incorporates all identifying information conveyed by these variables (Theorem (ref)). This formulation reduces the computational burden by lowering the dimensionality of the optimization problem.
When all random variables are discrete, the optimization problem ((ref))--((ref)) reduces to a finite-dimensional linear program, for which efficient algorithms and solvers are available.\footnote{Many empirical applications, however, involve continuous outcomes and covariates. A common practical approach is to discretize these variables, although this may come at the cost of reduced sharpness of identification. Details of the linear programming formulation are provided in Appendix (ref).} Inference methods for bounds defined by linear programs, including those proposed by fang2023inference and cho2024simple, are applicable in this setting.
Let $\widetilde{\mathcal{F}}^{\dagger}$ denote the class of CDFs $F$ for all defined variables $(Y_1,Y_0,Y,D,S,G,X)$ that satisfy Assumptions (ref)--(ref) and (ref)--(ref), with $F^{\ast}$ replaced by $F$. The identified set for $\theta_{o}$ under the combined data and these assumptions is defined as
where $\widetilde{\mathcal{F}}^{\ast} \equiv \bigl\{F \in \widetilde{\mathcal{F}}^{\dagger}: F_{YDGX} = F_{YDGX}^{\ast}\bigr\}$.
The following proposition shows that $\Theta_{IC}^{\dagger}$ corresponds to an interval whose lower and upper bounds are given by the solutions to the minimization and maximization problems ((ref))--((ref)).
This result allows us to compute the identified set for $\theta_o$ under the additional restrictions (Assumptions (ref) and (ref)) by solving the linear program ((ref))--((ref)). Although our analysis focuses on the case in which both assumptions are imposed jointly, the sharp bounds can also be obtained by solving optimization problems of the same form when either assumption is imposed on its own.
We illustrate the proposed approach using data from the DRPT study of gaines2011experimental, conducted in Illinois as part of the 2008 Cooperative Campaign Analysis Project. The study examines the effects of negative campaign advertisements on candidate evaluations during the 2008 U.S. presidential election.
The sample consists of 483 adult Illinois residents who were randomly assigned to one of three groups: treatment ($n=118$), no-treatment ($n=129$), and self-selection ($n=236$). Individuals in the treatment group were exposed to negative campaign advertisements (e-flyers) about John McCain and Barack Obama, whereas those in the no-treatment group were not. Participants in the self-selection group were allowed to choose whether to view the advertisements, and 90 of the 236 participants opted to do so.
The outcome variable $Y$ is each respondent's feeling thermometer rating toward each candidate, measured on a scale from 0 (very unfavorable) to 100 (very favorable), with 50 indicating neutrality. The dataset includes a single categorical covariate, $X$, indicating respondents' partisanship: Republican ($n=207$), Democrat ($n=233$), and Independent ($n=43$).
Let $Y_{d,\text{McCain}}$ and $Y_{d,\text{Obama}}$ denote the potential feeling thermometer ratings toward John McCain and Barack Obama, respectively, under treatment status $d \in \{0,1\}$, where $d=1$ indicates exposure to the negative campaign advertisements. Our parameter of interest is $\mathbb{P}(Y_{1,j} < Y_{0,j})$, which represents the proportion of individuals whose evaluation of candidate $j \in \{\text{McCain}, \text{Obama}\}$ is negatively affected by the campaign material. This parameter is particularly appealing because it remains well defined and substantively meaningful even when the outcome $Y_{d,j}$ is ordinal rather than cardinal. This feature is important in our application, as feeling thermometer ratings reflect subjective assessments and hence may lack a strong cardinal interpretation (see, e.g., wilcox1989some, for discussion).\footnote{In such a context, common causal parameters, such as the ATE $\mathbb{E}[Y_{1,j}-Y_{0,j}]$, may fail to provide a meaningful interpretation.} We also estimate the distribution function of the treatment effect, $F_{Y_{1,j}-Y_{0,j}}^{\ast}(\delta)$, for each candidate.
For each candidate $j \in \{\text{McCain}, \text{Obama}\}$, we estimate the identified set for $\mathbb{P}(Y_{1,j} < Y_{0,j})$ both with and without the self-selection sample ($G=\mathrm{obs}$), and both with and without imposing Assumption (ref) (Positive Dependence). Assumption (ref) is motivated by the idea that individuals who hold a more favorable view of a candidate in the control state are likely to maintain relatively favorable views even when exposed to negative campaign advertisements, thereby inducing positive dependence between the potential outcomes.\footnote{We do not impose Assumption (ref), since self-selection into viewing negative campaign advertisements is unlikely to follow the Generalized Roy selection model.} Without Assumption (ref), the identified sets are estimated using Proposition (ref) (with $\varphi(y_1,y_0)=y_1-y_0$ and $\delta=0$), together with Lemma (ref), replacing $F_{YDGX}^{\ast}$ with its empirical distribution. When Assumption (ref) is imposed, the identified sets are estimated via linear programming using empirical distributions; see Appendix (ref) for details.
Table (ref) reports the estimated identified sets and 95% confidence intervals for $\mathbb{P}(Y_1 < Y_0)$ for each candidate, for the full population and for specified subpopulations (Democrats, Republicans, and each self-selection type $s \in \{0,1\}$), across the four combinations of (i) including versus excluding the self-selection sample ($G=\mathrm{obs}$) and (ii) imposing versus not imposing Assumption (ref).\footnote{Inference for the analytical bounds in Section (ref) is based on pointwise valid bootstrap confidence intervals constructed using plug-in estimators. For the linear-programming bounds in Section (ref), we use the perturbation bootstrap procedure of cho2024simple, which is designed for inference on functionals of set-identified parameters defined by linear moment inequalities.} The estimated identified sets based solely on the random-assignment sample ($G=\mathrm{exp}$) are wide and therefore not particularly informative. By contrast, incorporating the self-selection sample substantially tightens the identified sets, with further reductions when Assumption (ref) is imposed. The inclusion of the self-selection sample also substantially tightens the corresponding confidence intervals.
In particular, the combined use of the self-selection sample and Assumption (ref) (our baseline specification) yields substantially tighter and more informative bounds. For John McCain, the baseline results suggest that at least 40% of individuals are negatively affected by the campaign advertisements, while at least 47% appear resistant. Qualitatively similar patterns are observed for Barack Obama (see Table (ref)). These findings point to substantial yet heterogeneous effects of the negative advertisements.
Figure (ref) plots the estimated identified sets for $F_{Y_{1,j}-Y_{0,j}}^{\ast}(\delta)$, together with 95% confidence intervals, for each candidate $j \in \{\text{McCain}, \text{Obama}\}$. It compares two specifications: (i) the benchmark case, which uses only the random-assignment sample and imposes no additional restriction, and (ii) our baseline specification, which combines the self-selection sample with Assumption (ref). The figure shows that the baseline specification yields markedly tighter identified sets and confidence intervals over a wide range of $\delta$. This pattern indicates that combining the self-selection sample with Assumption (ref) substantially improves the informativeness of the treatment-effect heterogeneity analysis.
Overall, these findings illustrate the empirical value of combining random-assignment and self-selection data. They also highlight an important advantage of DRPT designs: improving the informativeness of DTE analyses through the incorporation of self-selection data.
This study investigates how combining experimental and observational data can improve the identification of DTE parameters. We show that the identified set under the combined data is fully characterized by the joint distribution of each potential outcome, the latent self-selection variable, and covariates, with latent self-selection serving as the key source of additional identifying power. For a broad class of DTE parameters represented by super(sub)modular functions and $\varphi$-indicator functions, we derive sharp bounds under the combined data. We further establish necessary and sufficient conditions under which data combination strictly shrinks the identified set, suggesting that such shrinkage arises generically unless selection-on-observables holds in the observational data. We also propose a linear programming approach for computing sharp bounds while incorporating additional structural assumptions, such as positive dependence between the potential outcomes and generalized Roy model selection. An empirical application using DRPT data on negative campaign advertisements in the U.S. presidential election illustrates the practical value of combining random-assignment and self-selection data. Although observational data are often viewed as redundant once experimental data are available, our results show that they can still provide substantial additional identifying power for distributional parameters.
\setstretch{1.4}
I thank Brian Gaines, Marc Henry, Hidehiko Ichimura, Teppei Yamamoto, and participants at various seminars and conferences for their comments and suggestions. Masaki Suzuki provided excellent research assistance. I gratefully acknowledge financial support from JSPS KAKENHI Grant (number 24K16342).
\setstretch{1.35}