EconBase
← All papers

The Identification Power of Combining Experimental and Observational Data for Distributional Treatment Effect Parameters

Shosei Sakaguchi

arXiv 17 Aug 2025 · Econometrics

arXiv:2508.12206 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This study investigates the identification power gained by combining experimental data, in which treatment is randomized, with observational data, in which treatment is self-selected, for distributional treatment effect (DTE) parameters. While experimental data identify average treatment effects, many DTE parameters, such as the distribution of individual treatment effects, are only partially identified. We examine whether and how combining these two data sources tightens the identified set for such parameters. For broad classes of DTE parameters, we derive nonparametric sharp bounds under the combined data and clarify the mechanism through which data combination improves identification relative to using experimental data alone. Our analysis highlights that self-selection in observational data is a key source of identification power. We establish necessary and sufficient conditions under which the combined data shrink the identified set, showing that such shrinkage generally occurs unless selection-on-observables holds in the observational data. We also propose a linear programming approach to compute sharp bounds that can incorporate additional structural restrictions, such as positive dependence between potential outcomes and the generalized Roy model. An empirical application using data on negative campaign advertisements in the 2008 U.S. presidential election illustrates the practical relevance of the proposed approach.

Citation extraction

40
references
85
in-text mentions
40
distinct cited
1
self-citations
13,464
main-text words

appendix boundary found by appendix_command · 58% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Cho, JoonHwan and Russell, Thomas M (2024) Simple inference on functionals of set-identified parameters defined by linear moments1.00053100%
2Fan, Yanqin and Guerre, Emmanuel and Zhu, Dongming (2017) Partial identification of functionals of the joint distribution of “potential outcomes”0.89414471%
3Frandsen, Brigham R and Lefgren, Lars J (2021) Partial identification of the distribution of treatment effects with an application to the Knowledge is Power Program (KIPP)0.81142100%
4Gaines, Brian J and Kuklinski, James H (2011) Experimental estimation of heterogeneous treatment effects related to self-selection0.81142100%
5Long, Qi and Little, Roderick J and Lin, Xihong (2008) Causal inference in hybrid intervention trials involving treatment choice0.81142100%
6Fan, Yanqin and Park, Sang Soo (2010) Sharp bounds on the distribution of treatment effects and their statistical inference0.69351100%
7Cambanis, Stamatis and Simons, Gordon and Stout, William (1976) Inequalities for E k(x,y) when the marginals are fixed0.6444250%
8Athey, Susan and Chetty, Raj and Imbens, Guido (2025) The experimental selection correction estimator: Using experiments to remove biases in observational estimates0.64422100%
9Cui, Yifan and Han, Sukjin (2025) Policy learning with distributional welfare0.64422100%
10Joe, Harry (2014) Dependence Modeling with Copulas0.64422100%

Showing the top 10 of 40 scored citations.