EconBase
← All papers

Double machine learning for sample selection models

Michela Bia, Martin Huber, Lukáš Lafférs

arXiv 30 Nov 2020 · Econometrics · publishedJournal of Business and Economic Statistics (2023) · 34 citations (OpenAlex)

arXiv:2012.00745 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper considers the evaluation of discretely distributed treatments when outcomes are only observed for a subpopulation due to sample selection or outcome attrition. For identification, we combine a selection-on-observables assumption for treatment assignment with either selection-on-observables or instrumental variable assumptions concerning the outcome attrition/sample selection process. We also consider dynamic confounding, meaning that covariates that jointly affect sample selection and the outcome may (at least partly) be influenced by the treatment. To control in a data-driven way for a potentially high dimensional set of pre- and/or post-treatment covariates, we adapt the double machine learning framework for treatment evaluation to sample selection problems. We make use of (a) Neyman-orthogonal, doubly robust, and efficient score functions, which imply the robustness of treatment effect estimation to moderate regularization biases in the machine learning-based estimation of the outcome, treatment, or sample selection models and (b) sample splitting (or cross-fitting) to prevent overfitting bias. We demonstrate that the proposed estimators are asymptotically normal and root-n consistent under specific regularity conditions concerning the machine learners and investigate their finite sample properties in a simulation study. We also apply our proposed methodology to the Job Corps data for evaluating the effect of training on hourly wages which are only observed conditional on employment. The estimator is available in the causalweight package for the statistical software R.

Citation extraction

63
references
91
in-text mentions
63
distinct cited
4
self-citations
24,326
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, Newey, and Robins (2018) Double/debiased machine learning for treatment and structural parameters1.000103100%
2Levy (2019) Tutorial: Deriving The Efficient Influence Curve for Large Models0.81142100%
3Huber (2012) Identification of average treatment effects in social experiments under alternative forms of attrition self0.73732100%
4Huber (2014) Treatment evaluation in the presence of sample selection self0.73732100%
5Neyman (1959) Optimal asymptotic tests of composite statistical hypotheses0.73732100%
6Bodory and Huber (2018) The causalweight package for causal inference in R0.64422100%
7Das, Newey, and Vella (2003) Nonparametric Estimation of Sample Selection Models0.64422100%
8Newey (2007) Nonparametric continuous/discrete choice models0.64422100%
9Robins, Rotnitzky, and Zhao (1994) Estimation of Regression Coefficients When Some Regressors Are not Always Observed0.64422100%
10Robins, Rotnitzky, and Zhao (1995) Analysis of Semiparametric Regression Models for Repeated Outcomes in the Presence of Missing Data0.64422100%

Showing the top 10 of 63 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Generalized Kernel Ridge Regression for Causal Inference with Missing-at-Random Sample Selection0.977154
2Automatic debiased machine learning and sensitivity analysis for sample selection models0.941125
32406.138260.84333
42411.094520.51121
5Semiparametric Estimation of Long-Term Treatment Effects$^*$0.40511
6a framework for generalization and transportation of causal estimates under covariate shift0.40511
7Double Machine Learning for Static Panel Models with Fixed Effects0.40511
8Locally robust semiparametric estimation of sample selection models without exclusion restrictions0.40511
9An Introduction to Double/Debiased Machine Learning0.40511
10xtdml: Double Machine Learning Estimation to Static Panel Data Models with Fixed Effects in R0.40511