EconBase
← All papers

A Practical Guide of Off-Policy Evaluation for Bandit Problems

Masahiro Kato, Kenshi Abe, Kaito Ariu, Shota Yasui

arXiv 23 Oct 2020 · Machine Learning · 4 citations (OpenAlex)

arXiv:2010.12470 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Off-policy evaluation (OPE) is the problem of estimating the value of a target policy from samples obtained via different policies. Recently, applying OPE methods for bandit problems has garnered attention. For the theoretical guarantees of an estimator of the policy value, the OPE methods require various conditions on the target policy and policy used for generating the samples. However, existing studies did not carefully discuss the practical situation where such conditions hold, and the gap between them remains. This paper aims to show new results for bridging the gap. Based on the properties of the evaluation policy, we categorize OPE situations. Then, among practical applications, we mainly discuss the best policy selection. For the situation, we propose a meta-algorithm based on existing OPE estimators. We investigate the proposed concepts using synthetic and open real-world datasets in experiments.

Citation extraction

38
references
83
in-text mentions
38
distinct cited
4
self-citations
7,594
main-text words

appendix boundary found by appendix_command · 38% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Narita, Y., Yasui, S., and Yata, K (2019) Efficient counterfactual learning from bandit feedback self0.84310660%
2Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters0.8435460%
3Saito, Y., Aihara, S., Matsutani, M., and Narita, Y (2020) A large-scale open dataset for bandit algorithms0.8435460%
4Kallus, N. and Uehara, M (2019) Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning0.84333100%
5Kato, M., Uehara, M., and Yasui, S (2002) Off-policy evaluation and learning for external validity under a covariate shift self0.73732100%
6Li, L., Chu, W., Langford, J., and Wang, X (2011) Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms0.6443267%
7Dudḱ, M., Langford, J., and Li, L (2011) Doubly Robust Policy Evaluation and Learning0.64422100%
8Kallus, N. and Uehara, M (2020) Efficient evaluation of natural stochastic policies in offline reinforcement learning0.64422100%
9Kato, M., Ishihara, T., Honda, J., and Narita, Y (2002) Adaptive experimental design for efficient treatment effect estimation: Randomized allocation via contextual bandit algorithm self0.58510520%
10Hamilton, J (1994) Time series analysis0.5237414%

Showing the top 10 of 38 scored citations.