Yu-Shiou Willy Lin, Dae Woong Ham, Iavor Bojinov
arXiv 27 Apr 2026 · Statistics — Methodology
arXiv:2604.24652 · PDF · DOI · OpenAlex · Extracted main text
Multi-armed bandits are widely used for sequential experimentation in clinical trials, recommendation systems, and online platforms. While regret minimization and valid inference from adaptively collected data have each been studied extensively, a basic question remains: when does adaptivity improve estimation precision relative to uniform designs, and how should inference be balanced against the online cost of experimentation? We first study arm-level mean estimation under mean-squared-error (MSE) objectives. We characterize when an adaptive Neyman allocation, which allocates samples according to arm variance, yields strict MSE improvements over uniform sampling. When there is variance heterogeneity across arms, these improvements arise at modest sample sizes, clarifying that adaptivity can be preferable for inference not only asymptotically, but also in many practical finite-sample settings. We then study a joint inference-regret objective that accounts for the cost of assigning units to inferior arms during experimentation. We propose the Static-Allocation Rate Policy (SARP) and Neyman-Adaptive Rate Policy (NARP), which interpolates between inference- and regret-oriented policies by adjusting exploration to the local structure of the instance. We show that SARP and NARP converge to the complete-information benchmark at the optimal rate as the sampling budget grows. Our proposed policies are practically attractive as it linearly interpolates between any standard regret-minimizing algorithm and inference-targeting adaptive policies. Yet we show it still enjoys the oracle-based asymptotic optimal rate. Simulations support the theory by demonstrating improved precision over uniform allocation while controlling performance loss across a range of instances.
appendix boundary found by appendix_command · 51% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Akram Erraqabi and Alessandro Lazaric and Michal Valko and Emma Brun… (2017) Trading off Rewards and Errors in Multi-Armed Bandits | 1.000 | 5 | 3 | 100% |
| 2 | Jinglong Zhao (2023) Adaptive Neyman Allocation | 0.811 | 4 | 2 | 100% |
| 3 | András Antos and Varun Grover and Csaba Szepesvári (2010) Active learning in heteroscedastic noise | 0.737 | 3 | 2 | 100% |
| 4 | Simchi-Levi, David and Wang, Chonghuan (2025) Multi-armed Bandit Experimental Design: Online Decision-Making and Adaptive Inference | 0.737 | 3 | 2 | 100% |
| 5 | Carpentier, Alexandra and Lazaric, Alessandro and Ghavamzadeh, Moham… (2011) Upper-Confidence-Bound Algorithms for Active Learning in Multi-armed Bandits | 0.644 | 2 | 2 | 100% |
| 6 | Dai, Jessica and Gradu, Paula and Harshaw, Christopher (2023) Clip-ogd: An experimental design for adaptive neyman allocation in sequential experiments | 0.644 | 2 | 2 | 100% |
| 7 | Athey, Susan and Bickel, Peter J and Chen, Aiyou and Imbens, Guido W… (2023) Semi-parametric estimation of treatment effects in randomised experiments | 0.405 | 1 | 1 | 100% |
| 8 | Kari Lock Morgan and Donald B. Rubin (2012) Rerandomization to improve covariate balance in experiments | 0.405 | 1 | 1 | 100% |
| 9 | Chernoff, Herman (1992) Sequential Design of Experiments | 0.405 | 1 | 1 | 100% |
| 10 | T.L Lai and Herbert Robbins (1985) Asymptotically efficient adaptive allocation rules | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 33 scored citations.