Yusuke Narita, Shota Yasui, Kohei Yata
arXiv 20 Feb 2020 · Machine Learning
arXiv:2002.08536 · PDF · DOI · OpenAlex · Extracted main text
Efficient methods to evaluate new algorithms are critical for improving interactive bandit and reinforcement learning systems such as recommendation systems. A/B tests are reliable, but are time- and money-consuming, and entail a risk of failure. In this paper, we develop an alternative method, which predicts the performance of algorithms given historical data that may have been generated by a different algorithm. Our estimator has the property that its prediction converges in probability to the true performance of a counterfactual algorithm at a rate of $\sqrt{N}$, as the sample size $N$ increases. We also show a correct way to estimate the variance of our prediction, thus allowing the analyst to quantify the uncertainty in the prediction. These properties hold even when the analyst does not know which among a large number of potentially important state variables are actually important. We validate our method by a simulation experiment about reinforcement learning. We finally apply it to improve advertisement design by a major advertisement company. We find that our method produces smaller mean squared errors than state-of-the-art methods.
appendix boundary found by appendix_command · 73% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Nan Jiang and Lihong Li (2016) Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In Proceedings of the 33rd International Conference on Mac… | 1.000 | 8 | 3 | 100% |
| 2 | Philip Thomas and Emma Brunskill (2016) Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning. In Proceedings of the 33rd International Conference on M… | 1.000 | 7 | 3 | 100% |
| 3 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters | 0.909 | 8 | 4 | 75% |
| 4 | Miroslav Dudḱ, Dumitru Erhan, John Langford, and Lihong Li (2014) Doubly Robust Policy Evaluation and Optimization | 0.843 | 4 | 3 | 75% |
| 5 | Nathan Kallus and Masatoshi Uehara (2020) Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes, In Proceedings of the 37th Inter… | 0.843 | 3 | 3 | 100% |
| 6 | Yusuke Narita, Shota Yasui, and Kohei Yata (2019) Efficient Counterfactual Learning from Bandit Feedback, In Proceedings of the 33rd AAAI Conference on Artificial Intelligence self | 0.737 | 3 | 3 | 67% |
| 7 | Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh (2018) More Robust Doubly Robust Off-policy Evaluation. In Proceedings of the 35th International Conference on Machine Learning. 1447–1… | 0.644 | 2 | 2 | 100% |
| 8 | Alex Strehl, John Langford, Lihong Li, and Sham M Kakade (2010) Learning from Logged Implicit Exploration Data, In Advances in Neural Information Processing Systems 23 | 0.644 | 2 | 2 | 100% |
| 9 | Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, Joh… (2016) OpenAI Gym | 0.644 | 2 | 2 | 100% |
| 10 | Yao Liu, Omer Gottesman, Aniruddh Raghu, Matthieu Komorowski, Aldo A… (2018) Representation balancing mdps for off-policy policy evaluation. In Advances in Neural Information Processing Systems 31. 2644–2653 | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 24 scored citations.