EconBase
← All papers

Debiased Off-Policy Evaluation for Recommendation Systems

Yusuke Narita, Shota Yasui, Kohei Yata

arXiv 20 Feb 2020 · Machine Learning

arXiv:2002.08536 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Efficient methods to evaluate new algorithms are critical for improving interactive bandit and reinforcement learning systems such as recommendation systems. A/B tests are reliable, but are time- and money-consuming, and entail a risk of failure. In this paper, we develop an alternative method, which predicts the performance of algorithms given historical data that may have been generated by a different algorithm. Our estimator has the property that its prediction converges in probability to the true performance of a counterfactual algorithm at a rate of $\sqrt{N}$, as the sample size $N$ increases. We also show a correct way to estimate the variance of our prediction, thus allowing the analyst to quantify the uncertainty in the prediction. These properties hold even when the analyst does not know which among a large number of potentially important state variables are actually important. We validate our method by a simulation experiment about reinforcement learning. We finally apply it to improve advertisement design by a major advertisement company. We find that our method produces smaller mean squared errors than state-of-the-art methods.

Citation extraction

24
references
59
in-text mentions
24
distinct cited
1
self-citations
7,911
main-text words

appendix boundary found by appendix_command · 73% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Nan Jiang and Lihong Li (2016) Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In Proceedings of the 33rd International Conference on Mac…1.00083100%
2Philip Thomas and Emma Brunskill (2016) Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning. In Proceedings of the 33rd International Conference on M…1.00073100%
3Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters0.9098475%
4Miroslav Dudḱ, Dumitru Erhan, John Langford, and Lihong Li (2014) Doubly Robust Policy Evaluation and Optimization0.8434375%
5Nathan Kallus and Masatoshi Uehara (2020) Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes, In Proceedings of the 37th Inter…0.84333100%
6Yusuke Narita, Shota Yasui, and Kohei Yata (2019) Efficient Counterfactual Learning from Bandit Feedback, In Proceedings of the 33rd AAAI Conference on Artificial Intelligence self0.7373367%
7Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh (2018) More Robust Doubly Robust Off-policy Evaluation. In Proceedings of the 35th International Conference on Machine Learning. 1447–1…0.64422100%
8Alex Strehl, John Langford, Lihong Li, and Sham M Kakade (2010) Learning from Logged Implicit Exploration Data, In Advances in Neural Information Processing Systems 230.64422100%
9Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, Joh… (2016) OpenAI Gym0.64422100%
10Yao Liu, Omer Gottesman, Aniruddh Raghu, Matthieu Komorowski, Aldo A… (2018) Representation balancing mdps for off-policy policy evaluation. In Advances in Neural Information Processing Systems 31. 2644–26530.64422100%

Showing the top 10 of 24 scored citations.