EconBase
← All papers

Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games

Kenshi Abe, Yusuke Kaneko

arXiv 4 Jul 2020 · Machine Learning · 2 citations (OpenAlex)

arXiv:2007.02141 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Off-policy evaluation (OPE) is the problem of evaluating new policies using historical data obtained from a different policy. In the recent OPE context, most studies have focused on single-player cases, and not on multi-player cases. In this study, we propose OPE estimators constructed by the doubly robust and double reinforcement learning estimators in two-player zero-sum Markov games. The proposed estimators project exploitability that is often used as a metric for determining how close a policy profile (i.e., a tuple of policies) is to a Nash equilibrium in two-player zero-sum games. We prove the exploitability estimation error bounds for the proposed estimators. We then propose the methods to find the best candidate policy profile by selecting the policy profile that minimizes the estimated exploitability from a given policy profile class. We prove the regret bounds of the policy profiles selected by our methods. Finally, we demonstrate the effectiveness and performance of the proposed estimators through experiments.

Citation extraction

57
references
98
in-text mentions
58
distinct cited
0
self-citations
9,130
main-text words

appendix boundary found by appendix_command · 34% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Nan Jiang and Lihong Li (2016) Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In ICML. 652–6611.00063100%
2Nathan Kallus and Masatoshi Uehara (2019) Double reinforcement learning for efficient off-policy evaluation in markov decision processes0.94613685%
3Zhengyuan Zhou, Susan Athey, and Stefan Wager (2018) Offline multi-action policy learning: Generalization and optimization0.87412567%
4Masahiro Kato, Masatoshi Uehara, and Shota Yasui (2020) Off-Policy Evaluation and Learning for External Validity under a Covariate Shift0.73732100%
5Michael L Littman (1994) Markov games as a framework for multi-agent reinforcement learning0.73732100%
6Susan Athey and Stefan Wager (2017) Efficient policy learning0.64422100%
7Nathan Kallus and Masatoshi Uehara (2019) Efficiently breaking the curse of horizon: Double reinforcement learning in infinite-horizon processes0.64422100%
8Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.64422100%
9Adith Swaminathan and Thorsten Joachims (2015) Batch learning from logged bandit feedback through counterfactual risk minimization0.64422100%
10Philip Thomas and Emma Brunskill (2016) Data-efficient off-policy policy evaluation for reinforcement learning. In ICML. 2139–21480.64422100%

Showing the top 10 of 58 scored citations.