Sushant More, Priya Kotwal, Sujith Chappidi, Dinesh Mandalapu, Chris Khawand
arXiv 3 Sep 2024 · Machine Learning · 1 citations (OpenAlex)
arXiv:2409.02332 · PDF · DOI · OpenAlex · Extracted main text
Causal Impact (CI) of customer actions are broadly used across the industry to inform both short- and long-term investment decisions of various types. In this paper, we apply the double machine learning (DML) methodology to estimate the CI values across 100s of customer actions of business interest and 100s of millions of customers. We operationalize DML through a causal ML library based on Spark with a flexible, JSON-driven model configuration approach to estimate CI at scale (i.e., across hundred of actions and millions of customers). We outline the DML methodology and implementation, and associated benefits over the traditional potential outcomes based CI model. We show population-level as well as customer-level CI values along with confidence intervals. The validation metrics show a 2.2% gain over the baseline methods and a 2.5X gain in the computational time. Our contribution is to advance the scalable application of CI, while also providing an interface that allows faster experimentation, cross-platform support, ability to onboard new use cases, and improves accessibility of underlying code for partner teams.
appendix boundary found by appendix_command · 89% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Pages C1–C68, https://doi.org/10.1111/ectj.12097doi.org/10.1111/ectj.12097 | 0.737 | 3 | 2 | 100% |
| 2 | A. Abadie and G. Imbens, On the failure of bootstrap for matching es… (2008) 1537-1157 | 0.405 | 1 | 1 | 100% |
| 3 | P. C. Austin and E. A. Stuart, Moving towards Best Practice When usi… (2015) | 0.405 | 1 | 1 | 100% |
| 4 | Chernozhukov, V., Goldman, M., Semenova, V., and Taddy, M (2017) Orthogonal Machine Learning for Demand Estimation: High Dimensional Causal Inference in Dynamic Panels | 0.405 | 1 | 1 | 100% |
| 5 | Chernozhukov, Victor, Mert Demirer, Esther Duflo, and Ivan Fernandez… (2022) Generic machine learning inference on heterogenous treatment effects in randomized experiments | 0.405 | 1 | 1 | 100% |
| 6 | Holland, Paul W. Statistics and Causal Inference (1986) J. Amer. Statist. Assoc | 0.405 | 1 | 1 | 100% |
| 7 | Huber, Peter J. The behavior of maximum likelihood estimates under n… (1967) Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability | 0.405 | 1 | 1 | 100% |
| 8 | Sekhon, Jasjeet. The Neyman–Rubin Model of Causal Inference and Esti… (2007) The Oxford Handbook of Political Methodology | 0.405 | 1 | 1 | 100% |
| 9 | Knaus, M. C., Lechner, M., & Strittmatter, A (2021) Machine learning estimation of heterogeneous causal effects: Empirical monte carlo evidence | 0.405 | 1 | 1 | 100% |
| 10 | Neyman, Jerzy. Sur les applications de la theorie des probabilites a… (1923) Excerpts reprinted in English | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 14 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Valuing an Engagement Surface using a Large Scale Dynamic Causal Model | 0.511 | 2 | 1 |