Pengjie Zhou, Haoyu Wei, Huiming Zhang
arXiv 3 Dec 2024 · Statistics — Machine Learning · publishedMathematics (2025) · 10 citations (OpenAlex)
arXiv:2412.02251 · PDF · DOI · OpenAlex · Extracted main text
Reinforcement Learning (RL) is a widely researched area in artificial intelligence that focuses on teaching agents decision-making through interactions with their environment. A key subset includes stochastic multi-armed bandit (MAB) and continuum-armed bandit (SCAB) problems, which model sequential decision-making under uncertainty. This review outlines the foundational models and assumptions of bandit problems, explores non-asymptotic theoretical tools like concentration inequalities and minimax regret bounds, and compares frequentist and Bayesian algorithms for managing exploration-exploitation trade-offs. Additionally, we explore K-armed contextual bandits and SCAB, focusing on their methodologies and regret analyses. We also examine the connections between SCAB problems and functional data analysis. Finally, we highlight recent advances and ongoing challenges in the field.
appendix boundary found by appendix_command · 79% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Lattimore, T.; Szepesvári, C (2020) Bandit Algorithms; Cambridge University Press: Cambridge, UK, 2020 | 1.000 | 19 | 4 | 100% |
| 2 | Zhang, H.; Wei, H.; Cheng, G (2023) Tight non-asymptotic inference via sub-Gaussian intrinsic moment norm self | 1.000 | 7 | 3 | 100% |
| 3 | Lu, Y.; Xu, Z.; Tewari, A (2024) Bandit algorithms for precision medicine | 1.000 | 5 | 3 | 100% |
| 4 | Wei, H.; Wan, R.; Shi, L.; Song, R (2023) Zero-Inflated Bandits self | 1.000 | 5 | 3 | 100% |
| 5 | Li, L (2019) A perspective on off-policy evaluation in reinforcement learning | 0.737 | 3 | 2 | 100% |
| 6 | Ren, H.; Zhang, C.H (2024) On Lai's Upper Confidence Bound in Multi-Armed Bandits | 0.737 | 3 | 2 | 100% |
| 7 | Srinivas, N.; Krause, A.; Kakade, S.; Seeger, M (2010) Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design | 0.693 | 6 | 1 | 100% |
| 8 | Burtini, G.; Loeppky, J.; Lawrence, R (2015) A survey of online experiment design with the stochastic multi-armed bandit | 0.644 | 2 | 2 | 100% |
| 9 | Cai, T.T.; Pu, H (2022) Stochastic continuum-armed bandits with additive models: Minimax regrets and adaptive algorithm | 0.644 | 2 | 2 | 100% |
| 10 | Elena, G.; Milos, K.; Eugene, I (2021) Survey of multiarmed bandit algorithms applied to recommendation systems | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 153 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Zero-Inflated Bandits | 0.405 | 1 | 1 |