arXiv 30 Jun 2025 · Econometrics
arXiv:2506.24007 · PDF · DOI · OpenAlex · Extracted main text
This study investigates minimax and Bayes optimal strategies in fixed-budget best-arm identification. We consider an adaptive procedure consisting of a sampling phase followed by a recommendation phase, and we design an adaptive experiment within this framework to efficiently identify the best arm, defined as the one with the highest expected outcome. In our proposed strategy, the sampling phase consists of two stages. The first stage is a pilot phase, in which we allocate each arm uniformly in equal proportions to eliminate clearly suboptimal arms and estimate outcome variances. In the second stage, arms are allocated in proportion to the variances estimated during the first stage. After the sampling phase, the procedure enters the recommendation phase, where we select the arm with the highest sample mean as our estimate of the best arm. We prove that this single strategy is simultaneously asymptotically minimax and Bayes optimal for the simple regret, with upper bounds that coincide exactly with our lower bounds, including the constant terms.
appendix boundary found by appendix_command · 46% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Junpei Komiyama, Kaito Ariu, Masahiro Kato, and Chao Qin (2023) Rate-optimal bayesian simple regret in best arm identification self | 0.950 | 7 | 4 | 86% |
| 2 | Aurélien Garivier and Emilie Kaufmann (2016) Optimal best arm identification with fixed confidence | 0.928 | 5 | 4 | 80% |
| 3 | Jean-Yves Audibert, Sébastien Bubeck, and Remi Munos (2010) Best arm identification in multi-armed bandits | 0.843 | 4 | 3 | 75% |
| 4 | Tze Leung Lai and Herbert Robbins (1985) Asymptotically efficient adaptive allocation rules | 0.843 | 4 | 3 | 75% |
| 5 | Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier (2016) On the complexity of best-arm identification in multi-armed bandit models | 0.794 | 6 | 5 | 50% |
| 6 | Sébastien Bubeck, Rémi Munos, and Gilles Stoltz (2011) Pure exploration in finitely-armed and continuous-armed bandits | 0.737 | 5 | 2 | 60% |
| 7 | Kaito Ariu, Masahiro Kato, Junpei Komiyama, Kenichiro McAlinn, and C… (2021) Policy choice and best arm identification: Asymptotic analysis of exploration sampling, 2021 self | 0.737 | 4 | 2 | 75% |
| 8 | Chun-Hung Chen, Jianwu Lin, Enver Yücesan, and Stephen E. Chick (2000) Simulation budget allocation for further enhancing the efficiency of ordinal optimization | 0.737 | 3 | 3 | 67% |
| 9 | Rémy Degenne (2023) On the existence of a complexity in fixed budget bandit identification | 0.644 | 4 | 2 | 50% |
| 10 | Emilie Kaufmann (2020) Contributions to the Optimal Solution of Several Bandits Problems | 0.644 | 4 | 2 | 50% |
Showing the top 10 of 36 scored citations.