Justin Young, Muthoni Ngatia, Eleanor Wiske Dillon
arXiv 17 Jan 2026 · Econometrics
arXiv:2601.11845 · PDF · DOI · OpenAlex · Extracted main text
Recent developments in causal machine learning methods have made it easier to estimate flexible relationships between confounders, treatments and outcomes, making unconfoundedness assumptions in causal analysis more palatable. How successful are these approaches in recovering ground truth baselines? In this paper we analyze a new data sample including an experimental rollout of a new feature at a large technology company and a simultaneous sample of users who endogenously opted into the feature. We find that recovering ground truth causal effects is feasible -- but only with careful modeling choices. Our results build on the observational causal literature beginning with LaLonde (1986), offering best practices for more credible treatment effect estimation in modern, high-dimensional datasets.
appendix boundary found by appendix_command · 83% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | LaLonde, Robert J (1986) Evaluating the Econometric Evaluations of Training Programs with Experimental Data | 1.000 | 5 | 5 | 100% |
| 2 | Dehejia, Rajeev H. and Wahba, Sadek (1999) Causal Effects in Non-Experimental Studies: Reevaluating the Evaluation of Training Programs | 0.956 | 8 | 6 | 88% |
| 3 | Imbens, Guido W. and Xu, Yiqing (2025) LaLonde (1986) After Nearly Four Decades: Lessons Learned | 0.956 | 8 | 6 | 88% |
| 4 | Crump, Richard K. and Hotz, V. Joseph and Imbens, Guido W. and Mitni… (2009) Dealing with Limited Overlap in Estimation of Average Treatment Effects | 0.928 | 10 | 8 | 80% |
| 5 | Leo Breiman (1996) Bagging predictors | 0.928 | 4 | 3 | 100% |
| 6 | Chernozhukov, Victor and Cinelli, Carlos and Newey, Whitney and Shar… (2022) Long Story Short: Omitted Variable Bias in Causal Machine Learning | 0.843 | 5 | 3 | 60% |
| 7 | Dehejia, Rajeev H. and Wahba, Sadek (2002) Propensity Score-Matching Methods for Nonexperimental Causal Studies | 0.843 | 3 | 3 | 100% |
| 8 | Bach, Philipp and Schacht, Oliver and Chernozhukov, Victor and Klaas… (2024) Hyperparameter Tuning for Causal Inference with Double Machine Learning: A Simulation Study | 0.843 | 3 | 3 | 100% |
| 9 | Rosenbaum, Paul R. and Rubin, Donald B (1983) The Central Role of the Propensity Score in Observational Studies for Causal Effects | 0.811 | 4 | 2 | 100% |
| 10 | Chernozhukov, Victor and Hansen, Christian and Kallus, Nathan and Sp… (2024) Applied causal inference powered by ML and AI | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 40 scored citations.