EconBase
← All papers

ddml: Double/debiased machine learning in Stata

Achim Ahrens, Christian B. Hansen, Mark E. Schaffer, Thomas Wiemann

arXiv 23 Jan 2023 · Econometrics · publishedThe Stata Journal Promoting communications on statistics and Stata (2024) · 53 citations (OpenAlex)

arXiv:2301.09397 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We introduce the package ddml for Double/Debiased Machine Learning (DDML) in Stata. Estimators of causal parameters for five different econometric models are supported, allowing for flexible estimation of causal effects of endogenous variables in settings with unknown functional forms and/or many exogenous variables. ddml is compatible with many existing supervised machine learning programs in Stata. We recommend using DDML in combination with stacking estimation which combines multiple machine learners into a final predictor. We provide Monte Carlo evidence to support our recommendation.

Citation extraction

45
references
88
in-text mentions
45
distinct cited
2
self-citations
16,514
main-text words

appendix boundary found by appendix_command · 79% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W… (2018) Double/debiased machine learning for treatment and structural parameters1.000135100%
2width30.25006ptheight2.62222ptdepth-2.25222pt (2023) pystacked: Stacking generalization and machine learning in Stata0.92844100%
3Belloni, A., D. Chen, V. Chernozhukov, and C. Hansen (2012) Sparse Models and Methods for Optimal Instruments With an Application to Eminent Domain0.92843100%
4Ahrens, A., C. B. Hansen, and M. E. Schaffer (2018) PDSLASSO: Stata module for post-selection and post-regularization OLS or IV estimation and inference self0.81142100%
5Ahrens, A., C. B. Hansen, M. E. Schaffer, and T. Wiemann (2024) Model Averaging and Double Machine Learning. https://arxiv.org/abs/2401.01645 self0.81142100%
6width30.25006ptheight2.62222ptdepth-2.25222pt (2015) Valid Post-Selection and Post-Regularization Inference: An Elementary, General Approach0.64441100%
7width30.25006ptheight2.62222ptdepth-2.25222pt (2020) lassopack: Model selection and prediction with regularized regression in Stata0.64422100%
8Athey, S., J. Tibshirani, and S. Wager (2019) Generalized random forests0.64422100%
9Belloni, A., V. Chernozhukov, I. Fernández-Val, and C. Hansen (2017) Program Evaluation and Causal Inference With High-Dimensional Data0.64422100%
10Breiman, L (1996) Stacked regressions0.64422100%

Showing the top 10 of 45 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Model Averaging and Double Machine Learning0.84343
2On the Asymptotic Properties of Debiased Machine Learning Estimators0.84333
3Nonparametric rich covariateswithout saturation0.51121
4Reproducible Aggregation of Sample-Split Statistics$^*$0.40511
5An Introduction to Double/Debiased Machine Learning0.40511