Fengshi Niu, Harsha Nori, Brian Quistorff, Rich Caruana, Donald Ngwe, Aadharsh Kannan
arXiv 22 Feb 2022 · Statistics — Machine Learning · 5 citations (OpenAlex)
arXiv:2202.11043 · PDF · DOI · OpenAlex · Extracted main text
Estimating heterogeneous treatment effects in domains such as healthcare or social science often involves sensitive data where protecting privacy is important. We introduce a general meta-algorithm for estimating conditional average treatment effects (CATE) with differential privacy (DP) guarantees. Our meta-algorithm can work with simple, single-stage CATE estimators such as S-learner and more complex multi-stage estimators such as DR and R-learner. We perform a tight privacy analysis by taking advantage of sample splitting in our meta-algorithm and the parallel composition property of differential privacy. In this paper, we implement our approach using DP-EBMs as the base learner. DP-EBMs are interpretable, high-accuracy models with privacy guarantees, which allow us to directly observe the impact of DP noise on the learned causal model. Our experiments show that multi-stage CATE estimators incur larger accuracy loss than single-stage CATE or ATE estimators and that most of the accuracy loss from differential privacy is due to an increase in variance, not biased estimates of treatment effects.
appendix boundary found by appendix_titled_section at “Appendix: Proof of Theorem \ref{thm: DP guarantee}” · 93% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | X Nie and S Wager (2020) Quasi-oracle estimation of heterogeneous treatment effects | 1.000 | 7 | 3 | 100% |
| 2 | Edward H Kennedy (2020) Optimal doubly robust estimation of heterogeneous causal effects | 1.000 | 5 | 3 | 100% |
| 3 | Harsha Nori, Rich Caruana, Zhiqi Bu, Judy Hanwen Shen, and Janardhan… (2021) Accuracy, interpretability, and differential privacy via explainable boosting self | 0.928 | 4 | 3 | 100% |
| 4 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters | 0.737 | 3 | 2 | 100% |
| 5 | Jinshuo Dong, Aaron Roth, and Weijie J Su (2019) Gaussian differential privacy | 0.737 | 3 | 2 | 100% |
| 6 | Harsha Nori, Samuel Jenkins, Paul Koch, and Rich Caruana (2019) Interpretml: A unified framework for machine learning interpretability self | 0.737 | 3 | 2 | 100% |
| 7 | Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith (2006) Calibrating noise to sensitivity in private data analysis | 0.644 | 2 | 2 | 100% |
| 8 | Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Yu (2019) Metalearners for estimating heterogeneous treatment effects using machine learning | 0.644 | 2 | 2 | 100% |
| 9 | James M. Robins, Andrea Rotnitzky, and Lue Ping Zhao (1994) Estimation of regression coefficients when some regressors are not always observed | 0.644 | 2 | 2 | 100% |
| 10 | Mark J Van der Laan and Sherri Rose (2011) Targeted learning: Causal inference for observational and experimental data | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 35 scored citations.