EconBase
← All papers

Zero-Inflated Bandits

Haoyu Wei, Runzhe Wan, Lei Shi, Rui Song

arXiv 25 Dec 2023 · Statistics — Machine Learning

arXiv:2312.15595 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Many real-world bandit applications are characterized by sparse rewards, which can significantly hinder learning efficiency. Leveraging problem-specific structures for careful distribution modeling is recognized as essential for improving estimation efficiency in statistics. However, this approach remains under-explored in the context of bandits. To address this gap, we initiate the study of zero-inflated bandits, where the reward is modeled using a classic semi-parametric distribution known as the zero-inflated distribution. We develop algorithms based on the Upper Confidence Bound and Thompson Sampling frameworks for this specific structure. The superior empirical performance of these methods is demonstrated through extensive numerical studies.

Citation extraction

88
references
178
in-text mentions
89
distinct cited
8
self-citations
11,491
main-text words

appendix boundary found by appendix_titled_section at “Supplement to Simulation” · 30% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Wu, S., Wang, C.-H., Li, Y., and Cheng, G (2022) Residual bootstrap exploration for stochastic linear bandit0.9416583%
2Lattimore, T. and Szepesvári, C (2020) Bandit algorithms0.87412767%
3Zhang, H. and Chen, S. X (2020) Concentration inequalities for statistical inference0.8434375%
4Agrawal, S. and Goyal, N (2013) Thompson sampling for contextual bandits with linear payoffs0.8435460%
5Zhang, H. and Wei, H (2022) Sharper sub-weibull concentrations self0.8435360%
6Bubeck, S., Cesa-Bianchi, N., and Lugosi, G (2013) Bandits with heavy tail0.7948550%
7Li, L., Lu, Y., and Zhou, D (2017) Provably optimal algorithms for generalized linear contextual bandits0.7946550%
8Jin, T., Xu, P., Shi, J., Xiao, X., and Gu, Q (2021) Mots: Minimax optimal thompson sampling0.73715640%
9Shi, Z., Kuruoglu, E. E., and Wei, X (2022) Thompson sampling on asymmetric $$-stable bandits0.7374350%
10Hao, B., Abbasi Yadkori, Y., Wen, Z., and Cheng, G (2019) Bootstrapping upper confidence bound0.7373367%

Showing the top 10 of 89 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
12412.022511.00053