EconBase
← All papers

Counterfactual Learning with General Data-generating Policies

Yusuke Narita, Kyohei Okumura, Akihiro Shimizu, Kohei Yata

arXiv 4 Dec 2022 · Machine Learning · publishedProceedings of the AAAI Conference on Artificial Intelligence (2023) · 1 citations (OpenAlex)

arXiv:2212.01925 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Off-policy evaluation (OPE) attempts to predict the performance of counterfactual policies using log data from a different policy. We extend its applicability by developing an OPE method for a class of both full support and deficient support logging policies in contextual-bandit settings. This class includes deterministic bandit (such as Upper Confidence Bound) as well as deterministic decision-making based on supervised and unsupervised learning. We prove that our method's prediction converges in probability to the true performance of a counterfactual policy as the sample size increases. We validate our method with experiments on partly and entirely deterministic logging policies. Finally, we apply it to evaluate coupon targeting policies by a major online platform and show how to improve the existing policy.

Citation extraction

19
references
58
in-text mentions
37
distinct cited
1
self-citations
8,243
main-text words

appendix boundary found by appendix_command · 35% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Narita, Y.; and Yata, K (2022) Algorithm is Experiment: Machine Learning, Market Design, and Policy Eligibility Rules self0.73732100%
2Dudḱ, M.; Erhan, D.; Langford, J.; and Li, L (2014) Doubly robust policy evaluation and optimization0.58531100%
3Strehl, A.; Langford, J.; Li, L.; and Kakade, S. M (2010) Learning from logged implicit exploration data0.58531100%
4Farajtabar, M.; Chow, Y.; and Ghavamzadeh, M (2018) More robust doubly robust off-policy evaluation0.51121100%
5Swaminathan, A.; and Joachims, T (2015) The self-normalized estimator for counterfactual learning0.51121100%
6Precup, D (2000) Eligibility traces for off-policy policy evaluation0.51121100%
7Su, Y.; Dimakopoulou, M.; Krishnamurthy, A.; and Dudik, M (2020) Doubly robust off-policy evaluation with shrinkage0.51121100%
Irpan2019OffPolicyEVunmatched citation key Irpan2019OffPolicyEV0.40511100%
Jiang16unmatched citation key Jiang160.40511100%
10Kuzborskij, I.; Vernade, C.; Gyorgy, A.; and Szepesvari, C (2021) Confident Off-Policy Evaluation and Selection through Self-Normalized Importance Weighting0.40511100%

Showing the top 10 of 37 scored citations. 2 of these could not be matched to a bibliography entry, so only the citation key is shown.