Joel Persson, Jurriën Bakker, Dennis Bohle, Stefan Feuerriegel, Florian von Wangenheim
arXiv 23 Feb 2026 · Statistics — Methodology
arXiv:2602.20383 · PDF · DOI · OpenAlex · Extracted main text
Heterogeneous treatment effects (HTEs) are increasingly estimated using machine learning models that produce highly personalized predictions of treatment effects. In practice, however, predicted treatment effects are rarely interpreted, reported, or audited at the individual level but, instead, are often aggregated to broader subgroups, such as demographic segments, risk strata, or markets. We show that such aggregation can induce systematic bias of the group-level causal effect: even when models for predicting the individual-level conditional average treatment effect (CATE) are correctly specified and trained on data from randomized experiments, aggregating the predicted CATEs up to the group level does not, in general, recover the corresponding group average treatment effect (GATE). We develop a unified statistical framework to detect and mitigate this form of group bias in randomized experiments. We first define group bias as the discrepancy between the model-implied and experimentally identified GATEs, derive an asymptotically normal estimator, and then provide a simple-to-implement statistical test. For mitigation, we propose a shrinkage-based bias-correction, and show that the theoretically optimal and empirically feasible solutions have closed-form expressions. The framework is fully general, imposes minimal assumptions, and only requires computing sample moments. We analyze the economic implications of mitigating detected group bias for profit-maximizing personalized targeting, thereby characterizing when bias correction alters targeting decisions and profits, and the trade-offs involved. Applications to large-scale experimental data at major digital platforms validate our theoretical results and demonstrate empirical performance.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Greenland, Sander and Pearl, Judea and Robins, James M (1999) Confounding and Collapsibility in Causal Inference | 1.000 | 6 | 3 | 100% |
| 2 | Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/Debiased Machine Learning for Treatment and Structural Parameters | 1.000 | 5 | 4 | 100% |
| 3 | Leng, Yan and Dimmery, Drew (2024) Calibration of Heterogeneous Treatment Effects in Randomized Experiments | 1.000 | 5 | 4 | 100% |
| 4 | Colnet, Bénédicte and Josse, Julie and Varoquaux, Gaël and Scornet,… (2023) Risk Ratio, Odds Ratio, Risk Difference ... Which Causal Measure is Easier to Generalize? | 0.928 | 4 | 3 | 100% |
| 5 | Wager, Stefan and Athey, Susan (2018) Estimation and Inference of Heterogeneous Treatment Effects using Random Forests | 0.843 | 3 | 3 | 100% |
| 6 | Lemmens, Aurélie and Roos, Jason and Gabel, Sebastian and Ascarza, E… (2025) Personalization and Targeting: How to Experiment, Learn & Optimize | 0.811 | 4 | 2 | 100% |
| 7 | Corbett-Davies, Sam and Pierson, Emma and Feller, Avi and Goel, Shar… (2017) Algorithmic Decision Making and the Cost of Fairness | 0.737 | 3 | 2 | 100% |
| 8 | Diemert, Eustache and Betlei, Artem and Renaudin, Christophe and Mas… (2018) A Large Scale Benchmark for Uplift Modeling | 0.737 | 3 | 2 | 100% |
| 9 | Goldenberg, Dmitri and Albert, Javier and Bernardi, Lucas and Esteve… (2020) Free Lunch! Retrospective Uplift Modeling for Dynamic Promotions Recommendation within ROI Constraints | 0.737 | 3 | 2 | 100% |
| 10 | Gordon, Brett R and Moakler, Robert and Zettelmeyer, Florian (2023) Close Enough? A Large-Scale Exploration of Non-experimental Approaches to Advertising Measurement | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 63 scored citations.