Harsh Parikh, Marco Morucci, Vittorio Orlandi, Sudeepa Roy, Cynthia Rudin, Alexander Volfovsky
arXiv 4 Jul 2023 · Statistics — Methodology · 1 citations (OpenAlex)
arXiv:2307.01449 · PDF · DOI · OpenAlex · Extracted main text
Experimental and observational studies often lack validity due to untestable assumptions. We propose a double machine learning approach to combine experimental and observational studies, allowing practitioners to test for assumption violations and estimate treatment effects consistently. Our framework proposes a falsification test for external validity and ignorability under milder assumptions. We provide consistent treatment effect estimators even when one of the assumptions is violated. However, our no-free-lunch theorem highlights the necessity of accurately identifying the violated assumption for consistent treatment effect estimation. Through comparative analyses, we show our framework's superiority over existing data fusion methods. The practical utility of our approach is further exemplified by three real-world case studies, underscoring its potential for widespread application in empirical research.
appendix boundary found by appendix_command · 48% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Harsh Parikh, Cynthia Rudin, and Alexander Volfovsky (2022) Malts: Matching after learning to stretch self | 0.843 | 5 | 3 | 60% |
| 2 | Edward H Kennedy (2024) Semiparametric doubly robust targeted double machine learning: a review | 0.843 | 3 | 3 | 100% |
| 3 | Peter J Bickel, Chris AJ Klaassen, Peter J Bickel, Ya’acov Ritov, J… (1993) Efficient and adaptive estimation for semiparametric models, volume 4 | 0.811 | 4 | 2 | 100% |
| 4 | Robert J LaLonde (1986) Evaluating Econometric Evaluations of Training Programs with Experimental Data | 0.737 | 3 | 3 | 67% |
| 5 | Lili Wu and Shu Yang (2022) Integrative $ r $-learner of heterogeneous treatment effects combining experimental and observational studies | 0.737 | 3 | 2 | 100% |
| 6 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters | 0.707 | 17 | 5 | 35% |
| 7 | Max H Farrell (2015) Robust inference on average treatment effects with possibly more covariates than observations | 0.644 | 2 | 2 | 100% |
| 8 | Oliver Hines, Oliver Dukes, Karla Diaz-Ordaz, and Stijn Vansteelandt (2022) Demystifying statistical learning based on efficient influence functions | 0.644 | 2 | 2 | 100% |
| 9 | Yi Lu, Daniel O Scharfstein, Maria M Brooks, Kevin Quach, and Edward… (2019) Causal inference for comprehensive cohort studies | 0.644 | 2 | 2 | 100% |
| 10 | Matthew A Masten and Alexandre Poirier (2020) Inference on breakdown frontiers | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 53 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Data Fusion for Partial Identification of Causal Effects | 0.894 | 7 | 5 |
| 2 | Towards Generalizing Inferences from Trials to Target Populations | 0.511 | 2 | 1 |
| 3 | 2407.04448 | 0.405 | 1 | 1 |
| 4 | 2407.08602 | 0.405 | 1 | 1 |
| 5 | A Cautionary Tale on Integrating Studies with Disparate Outcome Measures for Causal Inference | 0.405 | 1 | 1 |