EconBase
← All papers

Off-Policy Evaluation of Bandit Algorithm from Dependent Samples under Batch Update Policy

Masahiro Kato, Yusuke Kaneko

arXiv 23 Oct 2020 · Machine Learning · 2 citations (OpenAlex)

arXiv:2010.13554 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The goal of off-policy evaluation (OPE) is to evaluate a new policy using historical data obtained via a behavior policy. However, because the contextual bandit algorithm updates the policy based on past observations, the samples are not independent and identically distributed (i.i.d.). This paper tackles this problem by constructing an estimator from a martingale difference sequence (MDS) for the dependent samples. In the data-generating process, we do not assume the convergence of the policy, but the policy uses the same conditional probability of choosing an action during a certain period. Then, we derive an asymptotically normal estimator of the value of an evaluation policy. As another advantage of our method, the batch-based approach simultaneously solves the deficient support problem. Using benchmark and real-world datasets, we experimentally confirm the effectiveness of the proposed method.

Citation extraction

36
references
85
in-text mentions
36
distinct cited
2
self-citations
7,835
main-text words

appendix boundary found by appendix_command · 39% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1van der Laan, M. J. and Lendle, S. D (2014) Online targeted learning1.000103100%
2Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2019) Confidence intervals for policy evaluation in adaptive experiments1.00093100%
3Kato, M., Ishihara, T., Honda, J., and Narita, Y (2002) Adaptive experimental design for efficient treatment effect estimation: Randomized allocation via contextual bandit algorithm self0.87472100%
4van der Laan, M. J (2008) The construction and analysis of adaptive group sequential designs0.87452100%
5Luedtke, A. R. and van der Laan, M. J (2016) Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy0.87452100%
6Narita, Y., Yasui, S., and Yata, K (2019) Efficient counterfactual learning from bandit feedback0.87452100%
7Hahn, J., Hirano, K., and Karlan, D (2011) Adaptive experimental design using the propensity score0.81142100%
8Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters0.73732100%
9Yang, Y. and Zhu, D (2002) Randomized allocation with nonparametric estimation for a multi-armed bandit problem with covariates0.73732100%
10Zheng, W. and van der Laan, M. J (2011) Cross-validated targeted minimum-loss-based estimation0.64422100%

Showing the top 10 of 36 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Adaptive Doubly Robust Estimator from Non-stationary Logging Policy under a Convergence of Average Probability0.40511