Aldo Gael Carranza, Susan Athey
arXiv 21 May 2023 · Machine Learning
arXiv:2305.12407 · PDF · DOI · OpenAlex · Extracted main text
We consider the problem of learning personalized decision policies from observational bandit feedback data across multiple heterogeneous data sources. In our approach, we introduce a novel regret analysis that establishes finite-sample upper bounds on distinguishing notions of global regret for all data sources on aggregate and of local regret for any given data source. We characterize these regret bounds by expressions of source heterogeneity and distribution shift. Moreover, we examine the practical considerations of this problem in the federated setting where a central server aims to train a policy on data distributed across the heterogeneous sources without collecting any of their raw data. We present a policy learning algorithm amenable to federation based on the aggregation of local policies trained with doubly robust offline policy evaluation strategies. Our analysis and supporting experimental results provide insights into tradeoffs in the participation of heterogeneous data sources in offline policy learning.
appendix boundary found by appendix_command · 26% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Agarwal, A., Basu, S., Schnabel, T., and Joachims, T (2017) Effective evaluation using logged bandit feedback from multiple loggers | 0.843 | 4 | 4 | 75% |
| 2 | Swaminathan, A. and Joachims, T (2015) Batch learning from logged bandit feedback through counterfactual risk minimization | 0.843 | 3 | 3 | 100% |
| 3 | Athey, S. and Wager, S (2021) Policy learning with observational data self | 0.830 | 7 | 7 | 57% |
| 4 | Zhou, Z., Athey, S., and Wager, S (2023) Offline multi-action policy learning: Generalization and optimization self | 0.822 | 9 | 7 | 56% |
| 5 | Jin, Y., Ren, Z., Yang, Z., and Wang, Z (2022) Policy learning" without”overlap: Pessimism and generalized empirical bernstein's inequality | 0.737 | 3 | 3 | 67% |
| 6 | Bietti, A., Agarwal, A., and Langford, J (2021) A contextual bandit bake-off | 0.644 | 2 | 2 | 100% |
| 7 | Hong, J., Kveton, B., Zaheer, M., Katariya, S., and Ghavamzadeh, M (2023) Multi-task off-policy learning from bandit feedback | 0.644 | 2 | 2 | 100% |
| 8 | Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhag… (2021) Advances and open problems in federated learning | 0.644 | 2 | 2 | 100% |
| 9 | Kallus, N., Saito, Y., and Uehara, M (2021) Optimal off-policy evaluation from multiple logging policies | 0.644 | 2 | 2 | 100% |
| 10 | Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 43 scored citations.