arXiv 17 Aug 2025 · Econometrics
arXiv:2508.12206 · PDF · DOI · OpenAlex · Extracted main text
This study investigates the identification power gained by combining experimental data, in which treatment is randomized, with observational data, in which treatment is self-selected, for distributional treatment effect (DTE) parameters. While experimental data identify average treatment effects, many DTE parameters, such as the distribution of individual treatment effects, are only partially identified. We examine whether and how combining these two data sources tightens the identified set for such parameters. For broad classes of DTE parameters, we derive nonparametric sharp bounds under the combined data and clarify the mechanism through which data combination improves identification relative to using experimental data alone. Our analysis highlights that self-selection in observational data is a key source of identification power. We establish necessary and sufficient conditions under which the combined data shrink the identified set, showing that such shrinkage generally occurs unless selection-on-observables holds in the observational data. We also propose a linear programming approach to compute sharp bounds that can incorporate additional structural restrictions, such as positive dependence between potential outcomes and the generalized Roy model. An empirical application using data on negative campaign advertisements in the 2008 U.S. presidential election illustrates the practical relevance of the proposed approach.
appendix boundary found by appendix_command · 58% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Cho, JoonHwan and Russell, Thomas M (2024) Simple inference on functionals of set-identified parameters defined by linear moments | 1.000 | 5 | 3 | 100% |
| 2 | Fan, Yanqin and Guerre, Emmanuel and Zhu, Dongming (2017) Partial identification of functionals of the joint distribution of “potential outcomes” | 0.894 | 14 | 4 | 71% |
| 3 | Frandsen, Brigham R and Lefgren, Lars J (2021) Partial identification of the distribution of treatment effects with an application to the Knowledge is Power Program (KIPP) | 0.811 | 4 | 2 | 100% |
| 4 | Gaines, Brian J and Kuklinski, James H (2011) Experimental estimation of heterogeneous treatment effects related to self-selection | 0.811 | 4 | 2 | 100% |
| 5 | Long, Qi and Little, Roderick J and Lin, Xihong (2008) Causal inference in hybrid intervention trials involving treatment choice | 0.811 | 4 | 2 | 100% |
| 6 | Fan, Yanqin and Park, Sang Soo (2010) Sharp bounds on the distribution of treatment effects and their statistical inference | 0.693 | 5 | 1 | 100% |
| 7 | Cambanis, Stamatis and Simons, Gordon and Stout, William (1976) Inequalities for E k(x,y) when the marginals are fixed | 0.644 | 4 | 2 | 50% |
| 8 | Athey, Susan and Chetty, Raj and Imbens, Guido (2025) The experimental selection correction estimator: Using experiments to remove biases in observational estimates | 0.644 | 2 | 2 | 100% |
| 9 | Cui, Yifan and Han, Sukjin (2025) Policy learning with distributional welfare | 0.644 | 2 | 2 | 100% |
| 10 | Joe, Harry (2014) Dependence Modeling with Copulas | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 40 scored citations.