Ruohan Zhan, Zhimei Ren, Susan Athey, Zhengyuan Zhou
arXiv 5 May 2021 · Statistics — Machine Learning · publishedManagement Science (2023) · 2 citations (OpenAlex)
arXiv:2105.02344 · PDF · DOI · OpenAlex · Extracted main text
Learning optimal policies from historical data enables personalization in a wide variety of applications including healthcare, digital recommendations, and online education. The growing policy learning literature focuses on settings where the data collection rule stays fixed throughout the experiment. However, adaptive data collection is becoming more common in practice, from two primary sources: 1) data collected from adaptive experiments that are designed to improve inferential efficiency; 2) data collected from production systems that progressively evolve an operational policy to improve performance over time (e.g. contextual bandits). Yet adaptivity complicates the optimal policy identification ex post, since samples are dependent, and each treatment may not receive enough observations for each type of individual. In this paper, we make initial research inquiries into addressing the challenges of learning the optimal policy with adaptively collected data. We propose an algorithm based on generalized augmented inverse propensity weighted (AIPW) estimators, which non-uniformly reweight the elements of a standard AIPW estimator to control worst-case estimation variance. We establish a finite-sample regret upper bound for our algorithm and complement it with a regret lower bound that quantifies the fundamental difficulty of policy learning with adaptive data. When equipped with the best weighting scheme, our algorithm achieves minimax rate optimal regret guarantees even with diminishing exploration. Finally, we demonstrate our algorithm's effectiveness using both synthetic data and public benchmark datasets.
appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Zhou, Z., Athey, S., and Wager, S (2022) Offline multi-action policy learning: Generalization and optimization self | 1.000 | 10 | 5 | 100% |
| 2 | Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments self | 1.000 | 8 | 3 | 100% |
| 3 | Athey, S. and Wager, S (2021) Policy learning with observational data self | 1.000 | 5 | 3 | 100% |
| 4 | Rakhlin, A., Sridharan, K., and Tewari, A (2015) Sequential complexities and uniform martingale laws of large numbers | 0.874 | 9 | 2 | 100% |
| 5 | Zhan, R., Hadad, V., Hirshberg, D. A., and Athey, S (2021) Off-policy evaluation via adaptive weighting with data from contextual bandits self | 0.874 | 7 | 2 | 100% |
| 6 | Luedtke, A. R. and van der Laan, M. J (2016) Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy | 0.874 | 5 | 2 | 100% |
| 7 | Agrawal, S. and Goyal, N (2013) Thompson sampling for contextual bandits with linear payoffs | 0.737 | 3 | 2 | 100% |
| 8 | Dudḱ, M., Langford, J., and Li, L (2011) Doubly robust policy evaluation and learning | 0.737 | 3 | 2 | 100% |
| 9 | Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.644 | 4 | 1 | 100% |
| 10 | Dimakopoulou, M., Zhou, Z., Athey, S., and Imbens, G (2017) Estimation considerations in contextual bandits self | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 83 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.