EconBase
← All papers

Estimation Considerations in Contextual Bandits

Maria Dimakopoulou, Zhengyuan Zhou, Susan Athey, Guido Imbens

arXiv 19 Nov 2017 · Statistics — Machine Learning · 27 citations (OpenAlex)

arXiv:1711.07077 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Contextual bandit algorithms are sensitive to the estimation method of the outcome model as well as the exploration method used, particularly in the presence of rich heterogeneity or complex outcome models, which can lead to difficult estimation problems along the path of learning. We study a consideration for the exploration vs. exploitation framework that does not arise in multi-armed bandits but is crucial in contextual bandits; the way exploration and exploitation is conducted in the present affects the bias and variance in the potential outcome model estimation in subsequent stages of learning. We develop parametric and non-parametric contextual bandits that integrate balancing methods from the causal inference literature in their estimation to make it less prone to problems of estimation bias. We provide the first regret bound analyses for contextual bandits with balancing in the domain of linear contextual bandits that match the state of the art regret bounds. We demonstrate the strong practical advantage of balanced contextual bandits on a large number of supervised learning datasets and on a synthetic example that simulates model mis-specification and prejudice in the initial training data. Additionally, we develop contextual bandits with simpler assignment policies by leveraging sparse model estimation methods from the econometrics literature and demonstrate empirically that in the early stages they can improve the rate of learning and decrease regret.

Citation extraction

66
references
122
in-text mentions
66
distinct cited
3
self-citations
11,731
main-text words

appendix boundary found by appendix_command · 67% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1L. Li, W. Chu, J. Langford, and R. Schapire (2010) A contextual-bandit approach to personalized news article recommendation0.9507386%
2S. Agrawal and N. Goyal (2013) Thompson sampling for contextual bandits with linear payoffs0.9209478%
3A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire (2014) Taming the monster: A fast and simple algorithm for contextual bandits0.8947371%
4O. Chapelle and L. Li (2011) An empirical evaluation of thompson sampling0.87452100%
5M. Dudik, J. Langford, and L. Li (2011) Doubly robust policy evaluation and learning0.73732100%
6N. Kallus (2017) Balanced policy evaluation and learning0.73732100%
7H. Lei, A. Tewari, and S. Murphy (2017) An actor-critic contextual bandit algorithm for personalized mobile health interventions0.73732100%
8Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire (2011) Contextual bandits with linear payoff functions0.69310250%
9S. Athey, G. Imbens, and S. Wager (2017) Approximate residual balancing0.64422100%
10S. Athey, J. Tibshirani, and S. Wager (2017) Generalized random forests0.64422100%

Showing the top 10 of 66 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Policy Learning with Adaptively Collected Data0.64422
2Offline Multi-Action Policy Learning: Generalization and Optimization0.40511
3Audits as Evidence: Experiments, Ensembles, and Enforcement0.40511
4Online Causal Inference for Advertising in Real-Time Bidding Auctions0.40511
5Online Multi-Armed Bandits with Adaptive Inference0.40511
6Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits0.40511
7Reinforcing RCTs with Multiple Priors while Learning about External Validity0.40511
8Bandit Algorithms for Policy Learning: Methods, Implementation, and Welfare-performance0.40511
9Estimating Causal Effects from Data Generated by Stochastic Algorithms0.40511