EconBase
← All papers

Automatic debiased machine learning and sensitivity analysis for sample selection models

Jakob Bjelac, Victor Chernozhukov, Phil-Adrian Klotz, Jannis Kueck, Theresa M. A. Schmitz

arXiv 13 Jan 2026 · Econometrics

arXiv:2601.08643 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In this paper, we extend the Riesz representation framework to causal inference under sample selection, where both treatment assignment and outcome observability are non-random. Formulating the problem in terms of a Riesz representer enables stable estimation and a transparent decomposition of omitted variable bias into three interpretable components: a data-identified scale factor, outcome confounding strength, and selection confounding strength. For estimation, we employ the ForestRiesz estimator, which accounts for selective outcome observability while avoiding the instability associated with direct propensity score inversion. We assess finite-sample performance through a simulation study and show that conventional double machine learning approaches can be highly sensitive to tuning parameters due to their reliance on inverse probability weighting, whereas the ForestRiesz estimator delivers more stable performance by leveraging automatic debiased machine learning. In an empirical application to the gender wage gap in the U.S., we find that our ForestRiesz approach yields larger treatment effect estimates than a standard double machine learning approach, suggesting that ignoring sample selection leads to an underestimation of the gender wage gap. Sensitivity analysis indicates that implausibly strong unobserved confounding would be required to overturn our results. Overall, our approach provides a unified, robust, and computationally attractive framework for causal inference under sample selection.

Citation extraction

32
references
64
in-text mentions
32
distinct cited
10
self-citations
6,772
main-text words

appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Michela Bia and Martin Huber and Lukáš Lafférs (2024) Double Machine Learning for Sample Selection Models0.94112583%
2Chernozhukov, Victor and Cinelli, Carlos and Newey, Whitney and Shar… (2022) Long Story Short: Omitted Variable Bias in Causal Machine Learning self0.9285480%
3Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters self0.92843100%
4Chernozhukov, Victor and Newey, Whitney and Quintas-Mart\'inez, V\'i… (2022) RieszNet and ForestRiesz: Automatic Debiased Machine Learning with Neural Nets and Random Forests self0.87452100%
5Bach, Philipp and Chernozhukov, Victor and Klaassen, Sven and Kurz,… DoubleML - Double Machine Learning in Python self0.7373367%
6Altonji, Joseph G. and Elder, Todd E. and Taber, Christopher R (2005) Selection on Observed and Unobserved Variables: Assessing the Effectiveness of Catholic Schools0.64422100%
7Cinelli, Carlos and Hazlett, Chad (2020) Making Sense of Sensitivity: Extending Omitted Variable Bias0.64422100%
8Imbens, Guido W (2003) Sensitivity to Exogeneity Assumptions in Program Evaluation0.64422100%
9Oster, Emily (2019) Unobservable Selection and Coefficient Stability: Theory and Evidence0.64422100%
10Chernozhukov, Victor and Newey, Whitney K. and Singh, Rahul (2022) Automatic Debiased Machine Learning of Causal and Structural Effects self0.58531100%

Showing the top 10 of 32 scored citations.