Yiyan Huang, Cheuk Hang Leung, Qi Wu, Xing Yan
arXiv 22 Mar 2021 · Statistics — Machine Learning
arXiv:2103.11869 · PDF · DOI · OpenAlex · Extracted main text
Causal learning is the key to obtaining stable predictions and answering what if problems in decision-makings. In causal learning, it is central to seek methods to estimate the average treatment effect (ATE) from observational data. The Double/Debiased Machine Learning (DML) is one of the prevalent methods to estimate ATE. However, the DML estimators can suffer from an error-compounding issue and even give extreme estimates when the propensity scores are close to 0 or 1. Previous studies have overcome this issue through some empirical tricks such as propensity score trimming, yet none of the existing works solves it from a theoretical standpoint. In this paper, we propose a Robust Causal Learning (RCL) method to offset the deficiencies of DML estimators. Theoretically, the RCL estimators i) satisfy the (higher-order) orthogonal condition and are as consistent and doubly robust as the DML estimators, and ii) get rid of the error-compounding issue. Empirically, the comprehensive experiments show that: i) the RCL estimators give more stable estimations of the causal parameters than DML; ii) the RCL estimators outperform traditional estimators and their variants when applying different machine learning models on both simulation and benchmark datasets, and a mimic consumer credit dataset generated by WGAN.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W.… (2018) Double/debiased machine learning for treatment and structural parameters | 1.000 | 13 | 4 | 100% |
| 2 | L. Mackey, V. Syrgkanis, and I. Zadik, “Orthogonal machine learning:… (2018) Orthogonal machine learning: Power and limitations | 1.000 | 10 | 5 | 100% |
| 3 | J. L. Hill, “Bayesian nonparametric modeling for causal inference,”… (2011) Bayesian nonparametric modeling for causal inference | 0.811 | 4 | 2 | 100% |
| 4 | U. Shalit, F. D. Johansson, and D. Sontag, “Estimating individual tr… (2017) Estimating individual treatment effect: generalization bounds and algorithms | 0.811 | 4 | 2 | 100% |
| 5 | C. Shi, D. Blei, and V. Veitch, “Adapting neural networks for the es… (2019) Adapting neural networks for the estimation of treatment effects | 0.811 | 4 | 2 | 100% |
| 6 | J. Yoon, J. Jordon, and M. Van Der Schaar, “Ganite: Estimation of in… (2018) Ganite: Estimation of individualized treatment effects using generative adversarial nets | 0.737 | 3 | 2 | 100% |
| 7 | A. Linden, S. D. Uysal, A. Ryan, and J. L. Adams, “Estimating causal… (2016) Estimating causal effects for multivalued treatments: a comparison of approaches | 0.644 | 2 | 2 | 100% |
| 8 | C. Louizos, U. Shalit, J. M. Mooij, D. Sontag, R. Zemel, and M. Well… (2017) Causal effect inference with deep latent-variable models | 0.644 | 2 | 2 | 100% |
| 9 | J. Robins, L. Li, E. Tchetgen, A. van der Vaart et al., “Higher orde… (2008) Higher order influence functions and minimax estimation of nonlinear functionals | 0.644 | 2 | 2 | 100% |
| 10 | S. Athey, G. W. Imbens, J. Metzger, and E. Munro, “Using wasserstein… (2021) Using wasserstein generative adversarial networks for the design of monte carlo simulations | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 48 scored citations.