Achim Ahrens, Christian B. Hansen, Mark E. Schaffer, Thomas Wiemann
arXiv 23 Jan 2023 · Econometrics · publishedThe Stata Journal Promoting communications on statistics and Stata (2024) · 53 citations (OpenAlex)
arXiv:2301.09397 · PDF · DOI · OpenAlex · Extracted main text
We introduce the package ddml for Double/Debiased Machine Learning (DDML) in Stata. Estimators of causal parameters for five different econometric models are supported, allowing for flexible estimation of causal effects of endogenous variables in settings with unknown functional forms and/or many exogenous variables. ddml is compatible with many existing supervised machine learning programs in Stata. We recommend using DDML in combination with stacking estimation which combines multiple machine learners into a final predictor. We provide Monte Carlo evidence to support our recommendation.
appendix boundary found by appendix_command · 79% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W… (2018) Double/debiased machine learning for treatment and structural parameters | 1.000 | 13 | 5 | 100% |
| 2 | width30.25006ptheight2.62222ptdepth-2.25222pt (2023) pystacked: Stacking generalization and machine learning in Stata | 0.928 | 4 | 4 | 100% |
| 3 | Belloni, A., D. Chen, V. Chernozhukov, and C. Hansen (2012) Sparse Models and Methods for Optimal Instruments With an Application to Eminent Domain | 0.928 | 4 | 3 | 100% |
| 4 | Ahrens, A., C. B. Hansen, and M. E. Schaffer (2018) PDSLASSO: Stata module for post-selection and post-regularization OLS or IV estimation and inference self | 0.811 | 4 | 2 | 100% |
| 5 | Ahrens, A., C. B. Hansen, M. E. Schaffer, and T. Wiemann (2024) Model Averaging and Double Machine Learning. https://arxiv.org/abs/2401.01645 self | 0.811 | 4 | 2 | 100% |
| 6 | width30.25006ptheight2.62222ptdepth-2.25222pt (2015) Valid Post-Selection and Post-Regularization Inference: An Elementary, General Approach | 0.644 | 4 | 1 | 100% |
| 7 | width30.25006ptheight2.62222ptdepth-2.25222pt (2020) lassopack: Model selection and prediction with regularized regression in Stata | 0.644 | 2 | 2 | 100% |
| 8 | Athey, S., J. Tibshirani, and S. Wager (2019) Generalized random forests | 0.644 | 2 | 2 | 100% |
| 9 | Belloni, A., V. Chernozhukov, I. Fernández-Val, and C. Hansen (2017) Program Evaluation and Causal Inference With High-Dimensional Data | 0.644 | 2 | 2 | 100% |
| 10 | Breiman, L (1996) Stacked regressions | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 45 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Model Averaging and Double Machine Learning | 0.843 | 4 | 3 |
| 2 | On the Asymptotic Properties of Debiased Machine Learning Estimators | 0.843 | 3 | 3 |
| 3 | Nonparametric rich covariateswithout saturation | 0.511 | 2 | 1 |
| 4 | Reproducible Aggregation of Sample-Split Statistics$^*$ | 0.405 | 1 | 1 |
| 5 | An Introduction to Double/Debiased Machine Learning | 0.405 | 1 | 1 |