arXiv 4 Jul 2020 · Machine Learning · 2 citations (OpenAlex)
arXiv:2007.02141 · PDF · DOI · OpenAlex · Extracted main text
Off-policy evaluation (OPE) is the problem of evaluating new policies using historical data obtained from a different policy. In the recent OPE context, most studies have focused on single-player cases, and not on multi-player cases. In this study, we propose OPE estimators constructed by the doubly robust and double reinforcement learning estimators in two-player zero-sum Markov games. The proposed estimators project exploitability that is often used as a metric for determining how close a policy profile (i.e., a tuple of policies) is to a Nash equilibrium in two-player zero-sum games. We prove the exploitability estimation error bounds for the proposed estimators. We then propose the methods to find the best candidate policy profile by selecting the policy profile that minimizes the estimated exploitability from a given policy profile class. We prove the regret bounds of the policy profiles selected by our methods. Finally, we demonstrate the effectiveness and performance of the proposed estimators through experiments.
appendix boundary found by appendix_command · 34% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Nan Jiang and Lihong Li (2016) Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In ICML. 652–661 | 1.000 | 6 | 3 | 100% |
| 2 | Nathan Kallus and Masatoshi Uehara (2019) Double reinforcement learning for efficient off-policy evaluation in markov decision processes | 0.946 | 13 | 6 | 85% |
| 3 | Zhengyuan Zhou, Susan Athey, and Stefan Wager (2018) Offline multi-action policy learning: Generalization and optimization | 0.874 | 12 | 5 | 67% |
| 4 | Masahiro Kato, Masatoshi Uehara, and Shota Yasui (2020) Off-Policy Evaluation and Learning for External Validity under a Covariate Shift | 0.737 | 3 | 2 | 100% |
| 5 | Michael L Littman (1994) Markov games as a framework for multi-agent reinforcement learning | 0.737 | 3 | 2 | 100% |
| 6 | Susan Athey and Stefan Wager (2017) Efficient policy learning | 0.644 | 2 | 2 | 100% |
| 7 | Nathan Kallus and Masatoshi Uehara (2019) Efficiently breaking the curse of horizon: Double reinforcement learning in infinite-horizon processes | 0.644 | 2 | 2 | 100% |
| 8 | Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.644 | 2 | 2 | 100% |
| 9 | Adith Swaminathan and Thorsten Joachims (2015) Batch learning from logged bandit feedback through counterfactual risk minimization | 0.644 | 2 | 2 | 100% |
| 10 | Philip Thomas and Emma Brunskill (2016) Data-efficient off-policy policy evaluation for reinforcement learning. In ICML. 2139–2148 | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 58 scored citations.