EconBase
← All papers

Risk and optimal policies in bandit experiments

Karun Adusumilli

arXiv 13 Dec 2021 · Econometrics · publishedEconometrica (2025) · 2 citations (OpenAlex)

arXiv:2112.06363 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We provide a decision theoretic analysis of bandit experiments under local asymptotics. Working within the framework of diffusion processes, we define suitable notions of asymptotic Bayes and minimax risk for these experiments. For normally distributed rewards, the minimal Bayes risk can be characterized as the solution to a second-order partial differential equation (PDE). Using a limit of experiments approach, we show that this PDE characterization also holds asymptotically under both parametric and non-parametric distributions of the rewards. The approach further describes the state variables it is asymptotically sufficient to restrict attention to, and thereby suggests a practical strategy for dimension reduction. The PDEs characterizing minimal Bayes risk can be solved efficiently using sparse matrix routines or Monte-Carlo methods. We derive the optimal Bayes and minimax policies from their numerical solutions. These optimal policies substantially dominate existing methods such as Thompson sampling; the risk of the latter is often twice as high.

Citation extraction

33
references
88
in-text mentions
33
distinct cited
0
self-citations
13,471
main-text words

appendix boundary found by appendix_command · 40% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1T. Lattimore and C. Szepesvári, Bandit algorithms. 1em plus 0.5em mi… (2020)1.00053100%
2X. Kuang and S. Wager, “Weak signal asymptotics for sequentially ran… (2024) Weak signal asymptotics for sequentially randomized experiments0.87462100%
3M. Kasy and A. Sautmann, “Adaptive treatment assignment in experimen… (2021) Adaptive treatment assignment in experiments for policy choice0.8434375%
4L. Fan and P. W. Glynn, “Diffusion approximations for thompson sampl… (2021) Diffusion approximations for thompson sampling0.81142100%
5L. Le Cam and G. L. Yang, Asymptotics in Statistics: Some basic conc… (2000)0.7375340%
6M. G. Crandall, H. Ishii, and P.-L. Lions, “User's guide to viscosit… (1992) User's guide to viscosity solutions of second order partial differential equations0.7374350%
7T. L. Lai, “Adaptive treatment allocation and the multi-armed bandit… (1987) Adaptive treatment allocation and the multi-armed bandit problem0.73732100%
8A. W. Van der Vaart, Asymptotic statistics. 1em plus 0.5em minus 0.4… (2000)0.66910330%
9K. Hirano and J. R. Porter, “Asymptotics for statistical treatment r… (2009) Asymptotics for statistical treatment rules0.64422100%
10G. Barles and P. E. Souganidis, “Convergence of approximation scheme… (1991) Convergence of approximation schemes for fully nonlinear second order equations0.5854325%

Showing the top 10 of 33 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Continuous time asymptotic representations for adaptive experiments1.00085
2How to sample and when to stop sampling: The generalized Wald problem and minimax policies0.82294
3Asymptotic Representations for Sequential Decisions, Adaptive Experiments, and Batched Bandits0.58531
4Best Arm Identification with Contextual Information under a Small Gap0.40511
5A Primer on the Analysis of Randomized Experiments and a Survey of some Recent Advances0.40511
6Bandit Algorithms for Policy Learning: Methods, Implementation, and Welfare-performance0.40511
7Optimizing Returns from Experimentation Programs0.40511
8Valid Post-Contextual Bandit Inference0.40511
9Dynamic Decision-Making under Model Misspecification0.40511
10Sequential Decision Problems with Missing Feedback0.40511