Yusuke Narita, Kyohei Okumura, Akihiro Shimizu, Kohei Yata
arXiv 4 Dec 2022 · Machine Learning · publishedProceedings of the AAAI Conference on Artificial Intelligence (2023) · 1 citations (OpenAlex)
arXiv:2212.01925 · PDF · DOI · OpenAlex · Extracted main text
Off-policy evaluation (OPE) attempts to predict the performance of counterfactual policies using log data from a different policy. We extend its applicability by developing an OPE method for a class of both full support and deficient support logging policies in contextual-bandit settings. This class includes deterministic bandit (such as Upper Confidence Bound) as well as deterministic decision-making based on supervised and unsupervised learning. We prove that our method's prediction converges in probability to the true performance of a counterfactual policy as the sample size increases. We validate our method with experiments on partly and entirely deterministic logging policies. Finally, we apply it to evaluate coupon targeting policies by a major online platform and show how to improve the existing policy.
appendix boundary found by appendix_command · 35% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Narita, Y.; and Yata, K (2022) Algorithm is Experiment: Machine Learning, Market Design, and Policy Eligibility Rules self | 0.737 | 3 | 2 | 100% |
| 2 | Dudḱ, M.; Erhan, D.; Langford, J.; and Li, L (2014) Doubly robust policy evaluation and optimization | 0.585 | 3 | 1 | 100% |
| 3 | Strehl, A.; Langford, J.; Li, L.; and Kakade, S. M (2010) Learning from logged implicit exploration data | 0.585 | 3 | 1 | 100% |
| 4 | Farajtabar, M.; Chow, Y.; and Ghavamzadeh, M (2018) More robust doubly robust off-policy evaluation | 0.511 | 2 | 1 | 100% |
| 5 | Swaminathan, A.; and Joachims, T (2015) The self-normalized estimator for counterfactual learning | 0.511 | 2 | 1 | 100% |
| 6 | Precup, D (2000) Eligibility traces for off-policy policy evaluation | 0.511 | 2 | 1 | 100% |
| 7 | Su, Y.; Dimakopoulou, M.; Krishnamurthy, A.; and Dudik, M (2020) Doubly robust off-policy evaluation with shrinkage | 0.511 | 2 | 1 | 100% |
| Irpan2019OffPolicyEV | unmatched citation key Irpan2019OffPolicyEV | 0.405 | 1 | 1 | 100% |
| Jiang16 | unmatched citation key Jiang16 | 0.405 | 1 | 1 | 100% |
| 10 | Kuzborskij, I.; Vernade, C.; Gyorgy, A.; and Szepesvari, C (2021) Confident Off-Policy Evaluation and Selection through Self-Normalized Importance Weighting | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 37 scored citations. 2 of these could not be matched to a bibliography entry, so only the citation key is shown.