Jakob Bjelac, Victor Chernozhukov, Phil-Adrian Klotz, Jannis Kueck, Theresa M. A. Schmitz
arXiv 13 Jan 2026 · Econometrics
arXiv:2601.08643 · PDF · DOI · OpenAlex · Extracted main text
In this paper, we extend the Riesz representation framework to causal inference under sample selection, where both treatment assignment and outcome observability are non-random. Formulating the problem in terms of a Riesz representer enables stable estimation and a transparent decomposition of omitted variable bias into three interpretable components: a data-identified scale factor, outcome confounding strength, and selection confounding strength. For estimation, we employ the ForestRiesz estimator, which accounts for selective outcome observability while avoiding the instability associated with direct propensity score inversion. We assess finite-sample performance through a simulation study and show that conventional double machine learning approaches can be highly sensitive to tuning parameters due to their reliance on inverse probability weighting, whereas the ForestRiesz estimator delivers more stable performance by leveraging automatic debiased machine learning. In an empirical application to the gender wage gap in the U.S., we find that our ForestRiesz approach yields larger treatment effect estimates than a standard double machine learning approach, suggesting that ignoring sample selection leads to an underestimation of the gender wage gap. Sensitivity analysis indicates that implausibly strong unobserved confounding would be required to overturn our results. Overall, our approach provides a unified, robust, and computationally attractive framework for causal inference under sample selection.
appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Michela Bia and Martin Huber and Lukáš Lafférs (2024) Double Machine Learning for Sample Selection Models | 0.941 | 12 | 5 | 83% |
| 2 | Chernozhukov, Victor and Cinelli, Carlos and Newey, Whitney and Shar… (2022) Long Story Short: Omitted Variable Bias in Causal Machine Learning self | 0.928 | 5 | 4 | 80% |
| 3 | Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters self | 0.928 | 4 | 3 | 100% |
| 4 | Chernozhukov, Victor and Newey, Whitney and Quintas-Mart\'inez, V\'i… (2022) RieszNet and ForestRiesz: Automatic Debiased Machine Learning with Neural Nets and Random Forests self | 0.874 | 5 | 2 | 100% |
| 5 | Bach, Philipp and Chernozhukov, Victor and Klaassen, Sven and Kurz,… DoubleML - Double Machine Learning in Python self | 0.737 | 3 | 3 | 67% |
| 6 | Altonji, Joseph G. and Elder, Todd E. and Taber, Christopher R (2005) Selection on Observed and Unobserved Variables: Assessing the Effectiveness of Catholic Schools | 0.644 | 2 | 2 | 100% |
| 7 | Cinelli, Carlos and Hazlett, Chad (2020) Making Sense of Sensitivity: Extending Omitted Variable Bias | 0.644 | 2 | 2 | 100% |
| 8 | Imbens, Guido W (2003) Sensitivity to Exogeneity Assumptions in Program Evaluation | 0.644 | 2 | 2 | 100% |
| 9 | Oster, Emily (2019) Unobservable Selection and Coefficient Stability: Theory and Evidence | 0.644 | 2 | 2 | 100% |
| 10 | Chernozhukov, Victor and Newey, Whitney K. and Singh, Rahul (2022) Automatic Debiased Machine Learning of Causal and Structural Effects self | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 32 scored citations.