EconBase
← All papers

Policy Learning with Adaptively Collected Data

Ruohan Zhan, Zhimei Ren, Susan Athey, Zhengyuan Zhou

arXiv 5 May 2021 · Statistics — Machine Learning · publishedManagement Science (2023) · 2 citations (OpenAlex)

arXiv:2105.02344 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Learning optimal policies from historical data enables personalization in a wide variety of applications including healthcare, digital recommendations, and online education. The growing policy learning literature focuses on settings where the data collection rule stays fixed throughout the experiment. However, adaptive data collection is becoming more common in practice, from two primary sources: 1) data collected from adaptive experiments that are designed to improve inferential efficiency; 2) data collected from production systems that progressively evolve an operational policy to improve performance over time (e.g. contextual bandits). Yet adaptivity complicates the optimal policy identification ex post, since samples are dependent, and each treatment may not receive enough observations for each type of individual. In this paper, we make initial research inquiries into addressing the challenges of learning the optimal policy with adaptively collected data. We propose an algorithm based on generalized augmented inverse propensity weighted (AIPW) estimators, which non-uniformly reweight the elements of a standard AIPW estimator to control worst-case estimation variance. We establish a finite-sample regret upper bound for our algorithm and complement it with a regret lower bound that quantifies the fundamental difficulty of policy learning with adaptive data. When equipped with the best weighting scheme, our algorithm achieves minimax rate optimal regret guarantees even with diminishing exploration. Finally, we demonstrate our algorithm's effectiveness using both synthetic data and public benchmark datasets.

Citation extraction

83
references
145
in-text mentions
83
distinct cited
7
self-citations
15,849
main-text words

appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Zhou, Z., Athey, S., and Wager, S (2022) Offline multi-action policy learning: Generalization and optimization self1.000105100%
2Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments self1.00083100%
3Athey, S. and Wager, S (2021) Policy learning with observational data self1.00053100%
4Rakhlin, A., Sridharan, K., and Tewari, A (2015) Sequential complexities and uniform martingale laws of large numbers0.87492100%
5Zhan, R., Hadad, V., Hirshberg, D. A., and Athey, S (2021) Off-policy evaluation via adaptive weighting with data from contextual bandits self0.87472100%
6Luedtke, A. R. and van der Laan, M. J (2016) Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy0.87452100%
7Agrawal, S. and Goyal, N (2013) Thompson sampling for contextual bandits with linear payoffs0.73732100%
8Dudḱ, M., Langford, J., and Li, L (2011) Doubly robust policy evaluation and learning0.73732100%
9Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.64441100%
10Dimakopoulou, M., Zhou, Z., Athey, S., and Imbens, G (2017) Estimation considerations in contextual bandits self0.64422100%

Showing the top 10 of 83 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Policy Learning with Observational Data : The Case of Hepatitis C Treatment for HIV/HCV Co-Infected Patients0.64422
2Safe Policy Learning under Regression Discontinuity Designs with Multiple Cutoffs0.51121
3Double/Debiased Machine Learning for Dynamic Treatment Effects via $g$-Estimation0.40511
4Contextual Bandits in a Survey Experiment on Charitable Giving: Within-Experiment Outcomes versus Policy Learning0.40511
5Federated Offline Policy Learning0.40511
6Personalized Policy Learning through Discrete Experimentation: Theory and Empirical Evidence0.40511
7Benefits and Costs of Adaptive Sampling0.40511
8Wasserstein Policy Learning for Distributional Outcomes0.40511
9Estimating Causal Effects from Data Generated by Stochastic Algorithms0.40511