arXiv 21 Jan 2026 · Statistics — Machine Learning
arXiv:2601.15360 · PDF · DOI · OpenAlex · Extracted main text
Estimating Heterogeneous Treatment Effects (HTE) in industrial applications such as AdTech and healthcare presents a dual challenge: extreme class imbalance and heavy-tailed outcome distributions. While the X-Learner framework effectively addresses imbalance through cross-imputation, we demonstrate that it is fundamentally vulnerable to "Outlier Smearing" when reliant on Mean Squared Error (MSE) minimization. In this failure mode, the bias from a few extreme observations ("whales") in the minority group is propagated to the entire majority group during the imputation step, corrupting the estimated treatment effect structure. To resolve this, we propose the Robust X-Learner (RX-Learner). This framework integrates a redescending γ-divergence objective -- structurally equivalent to the Welsch loss under Gaussian assumptions -- into the gradient boosting machinery. We further stabilize the non-convex optimization using a Proxy Hessian strategy grounded in Majorization-Minimization (MM) principles. Empirical evaluation on a semi-synthetic Criteo Uplift dataset demonstrates that the RX-Learner reduces the Precision in Estimation of Heterogeneous Effect (PEHE) metric by 98.6% compared to the standard X-Learner, effectively decoupling the stable "Core" population from the volatile "Periphery".
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Künzel, S. R., Sekhon, J. S., Bickel, P. J., & Yu, B (2019) Metalearners for estimating heterogeneous treatment effects using machine learning | 0.928 | 4 | 3 | 100% |
| 2 | Diemert, E., Betlei, A., Renaudin, C., & Aigouy-Girard, M. R (2018) A large scale benchmark for uplift modeling | 0.644 | 2 | 2 | 100% |
| 3 | Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters | 0.644 | 2 | 2 | 100% |
| 4 | Fujisawa, H., & Eguchi, S (2008) Robust parameter estimation with a small bias against heavy contamination | 0.644 | 2 | 2 | 100% |
| 5 | Holland, P. W., & Welsch, R. E (1977) Robust regression using iteratively reweighted least-squares | 0.644 | 2 | 2 | 100% |
| 6 | Hunter, D. R., & Lange, K (2004) A tutorial on MM algorithms | 0.644 | 2 | 2 | 100% |
| 7 | Nie, X., & Wager, S (2021) Quasi-oracle estimation of heterogeneous treatment effects | 0.644 | 2 | 2 | 100% |
| 8 | Athey, S., & Imbens, G. W (2016) Recursive partitioning for heterogeneous causal effects | 0.585 | 3 | 1 | 100% |
| 9 | Geman, D., & Reynolds, G (1992) Constrained restoration and the recovery of discontinuities | 0.405 | 1 | 1 | 100% |
| 10 | Hampel, F. R., Ronchetti, E. M., Rousseeuw, P. J., & Stahel, W. A (1986) Robust statistics: The approach based on influence functions | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 16 scored citations.