EconBase
← All papers

Hyperparameter Tuning for Causal Inference with Double Machine Learning: A Simulation Study

Philipp Bach, Oliver Schacht, Victor Chernozhukov, Sven Klaassen, Martin Spindler

arXiv 7 Feb 2024 · Econometrics · 6 citations (OpenAlex)

arXiv:2402.04674 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Proper hyperparameter tuning is essential for achieving optimal performance of modern machine learning (ML) methods in predictive tasks. While there is an extensive literature on tuning ML learners for prediction, there is only little guidance available on tuning ML learners for causal machine learning and how to select among different ML learners. In this paper, we empirically assess the relationship between the predictive performance of ML methods and the resulting causal estimation based on the Double Machine Learning (DML) approach by Chernozhukov et al. (2018). DML relies on estimating so-called nuisance parameters by treating them as supervised learning problems and using them as plug-in estimates to solve for the (causal) parameter. We conduct an extensive simulation study using data from the 2019 Atlantic Causal Inference Conference Data Challenge. We provide empirical insights on the role of hyperparameter tuning and other practical decisions for causal estimation with DML. First, we assess the importance of data splitting schemes for tuning ML learners within Double Machine Learning. Second, we investigate how the choice of ML methods and hyperparameters, including recent AutoML frameworks, impacts the estimation performance for a causal parameter of interest. Third, we assess to what extent the choice of a particular causal model, as characterized by incorporated parametric assumptions, can be based on predictive performance metrics.

Citation extraction

29
references
43
in-text mentions
29
distinct cited
7
self-citations
6,437
main-text words

appendix boundary found by appendix_command · 65% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters self0.8746467%
2Alexandre Belloni, Victor Chernozhukov, and Christian Hansen (2013) Inference on Treatment Effects after Selection among High-Dimensional Controls† self0.84333100%
3Victor Chernozhukov, Whitney K. Newey, Victor Quintas-Martinez, and… (2022) Riesznet and forestriesz: Automatic debiased machine learning with neural nets and random forests, 2022 self0.7373367%
4Michael C Knaus (2022) Double machine learning-based programme evaluation under unconfoundedness0.64422100%
5Chi Wang and Qingyun Wu (1911) FLO: fast and lightweight hyperparameter optimization for automl0.64422100%
6Tianqi Chen and Carlos Guestrin (2016) XGBoost: A scalable tree boosting system0.64422100%
7F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. G… (2011) Scikit-learn: Machine learning in Python0.51121100%
8Claudia Shi, David M. Blei, and Victor Veitch (2019) Adapting neural networks for the estimation of treatment effects, 20190.51121100%
9James M Robins, Andrea Rotnitzky, and Lue Ping Zhao (1994) Estimation of regression coefficients when some regressors are not always observed0.40511100%
10Ezequiel Smucler, Andrea Rotnitzky, and James M Robins (2019) A unifying approach for doubly-robust $_1$ regularized estimation of causal contrasts0.40511100%

Showing the top 10 of 29 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Reevaluating Causal Estimation Methods with Data from a Product Release0.84333
2DoubleMLDeep: Estimation of Causal Effects with Multimodal Data0.40511
3Structure-agnostic Optimality of Doubly Robust Learning for Treatment Effect Estimation0.40511
4Semiparametric inference for impulse response functions using double/debiased machine learning0.40511
5An Introduction to Double/Debiased Machine Learning0.40511
6A Unifying Framework for Robust and Efficient Inference with Unstructured Data0.40511
7Bayesian Double Machine Learning for Causal Inference0.40511
8xtdml: Double Machine Learning Estimation to Static Panel Data Models with Fixed Effects in R0.40511
9Sharp Structure-Agnostic Lower Bounds for General Linear Functional Estimation0.40511
10Automatic debiased machine learning and sensitivity analysis for sample selection models0.40511