Haoyu Wei, Runzhe Wan, Lei Shi, Rui Song
arXiv 25 Dec 2023 · Statistics — Machine Learning
arXiv:2312.15595 · PDF · DOI · OpenAlex · Extracted main text
Many real-world bandit applications are characterized by sparse rewards, which can significantly hinder learning efficiency. Leveraging problem-specific structures for careful distribution modeling is recognized as essential for improving estimation efficiency in statistics. However, this approach remains under-explored in the context of bandits. To address this gap, we initiate the study of zero-inflated bandits, where the reward is modeled using a classic semi-parametric distribution known as the zero-inflated distribution. We develop algorithms based on the Upper Confidence Bound and Thompson Sampling frameworks for this specific structure. The superior empirical performance of these methods is demonstrated through extensive numerical studies.
appendix boundary found by appendix_titled_section at “Supplement to Simulation” · 30% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Wu, S., Wang, C.-H., Li, Y., and Cheng, G (2022) Residual bootstrap exploration for stochastic linear bandit | 0.941 | 6 | 5 | 83% |
| 2 | Lattimore, T. and Szepesvári, C (2020) Bandit algorithms | 0.874 | 12 | 7 | 67% |
| 3 | Zhang, H. and Chen, S. X (2020) Concentration inequalities for statistical inference | 0.843 | 4 | 3 | 75% |
| 4 | Agrawal, S. and Goyal, N (2013) Thompson sampling for contextual bandits with linear payoffs | 0.843 | 5 | 4 | 60% |
| 5 | Zhang, H. and Wei, H (2022) Sharper sub-weibull concentrations self | 0.843 | 5 | 3 | 60% |
| 6 | Bubeck, S., Cesa-Bianchi, N., and Lugosi, G (2013) Bandits with heavy tail | 0.794 | 8 | 5 | 50% |
| 7 | Li, L., Lu, Y., and Zhou, D (2017) Provably optimal algorithms for generalized linear contextual bandits | 0.794 | 6 | 5 | 50% |
| 8 | Jin, T., Xu, P., Shi, J., Xiao, X., and Gu, Q (2021) Mots: Minimax optimal thompson sampling | 0.737 | 15 | 6 | 40% |
| 9 | Shi, Z., Kuruoglu, E. E., and Wei, X (2022) Thompson sampling on asymmetric $$-stable bandits | 0.737 | 4 | 3 | 50% |
| 10 | Hao, B., Abbasi Yadkori, Y., Wen, Z., and Cheng, G (2019) Bootstrapping upper confidence bound | 0.737 | 3 | 3 | 67% |
Showing the top 10 of 89 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | 2412.02251 | 1.000 | 5 | 3 |