EconBase
← All papers

Selective Reviews of Bandit Problems in AI via a Statistical View

Pengjie Zhou, Haoyu Wei, Huiming Zhang

arXiv 3 Dec 2024 · Statistics — Machine Learning · publishedMathematics (2025) · 10 citations (OpenAlex)

arXiv:2412.02251 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Reinforcement Learning (RL) is a widely researched area in artificial intelligence that focuses on teaching agents decision-making through interactions with their environment. A key subset includes stochastic multi-armed bandit (MAB) and continuum-armed bandit (SCAB) problems, which model sequential decision-making under uncertainty. This review outlines the foundational models and assumptions of bandit problems, explores non-asymptotic theoretical tools like concentration inequalities and minimax regret bounds, and compares frequentist and Bayesian algorithms for managing exploration-exploitation trade-offs. Additionally, we explore K-armed contextual bandits and SCAB, focusing on their methodologies and regret analyses. We also examine the connections between SCAB problems and functional data analysis. Finally, we highlight recent advances and ongoing challenges in the field.

Citation extraction

153
references
225
in-text mentions
153
distinct cited
5
self-citations
25,095
main-text words

appendix boundary found by appendix_command · 79% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lattimore, T.; Szepesvári, C (2020) Bandit Algorithms; Cambridge University Press: Cambridge, UK, 20201.000194100%
2Zhang, H.; Wei, H.; Cheng, G (2023) Tight non-asymptotic inference via sub-Gaussian intrinsic moment norm self1.00073100%
3Lu, Y.; Xu, Z.; Tewari, A (2024) Bandit algorithms for precision medicine1.00053100%
4Wei, H.; Wan, R.; Shi, L.; Song, R (2023) Zero-Inflated Bandits self1.00053100%
5Li, L (2019) A perspective on off-policy evaluation in reinforcement learning0.73732100%
6Ren, H.; Zhang, C.H (2024) On Lai's Upper Confidence Bound in Multi-Armed Bandits0.73732100%
7Srinivas, N.; Krause, A.; Kakade, S.; Seeger, M (2010) Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design0.69361100%
8Burtini, G.; Loeppky, J.; Lawrence, R (2015) A survey of online experiment design with the stochastic multi-armed bandit0.64422100%
9Cai, T.T.; Pu, H (2022) Stochastic continuum-armed bandits with additive models: Minimax regrets and adaptive algorithm0.64422100%
10Elena, G.; Milos, K.; Eugene, I (2021) Survey of multiarmed bandit algorithms applied to recommendation systems0.64422100%

Showing the top 10 of 153 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Zero-Inflated Bandits0.40511