Ramon van den Akker, Bas J. M. Werker, Bo Zhou
arXiv 20 May 2025 · Econometrics
arXiv:2505.13897 · PDF · DOI · OpenAlex · Extracted main text
We establish an asymptotic framework for the statistical analysis of the stochastic contextual multi-armed bandit problem (CMAB), which is widely employed in adaptively randomized experiments across various fields. While algorithms for maximizing rewards or, equivalently, minimizing regret have received considerable attention, our focus centers on statistical inference with adaptively collected data under the CMAB model. To this end we derive the limit experiment (in the Hajek-Le Cam sense). This limit experiment is highly nonstandard and, applying Girsanov's theorem, we obtain a structural representation in terms of stochastic differential equations. This structural representation, and a general weak convergence result we develop, allow us to obtain the asymptotic distribution of statistics for the CMAB problem. In particular, we obtain the asymptotic distributions for the classical t-test (non-Gaussian), Adaptively Weighted tests, and Inverse Propensity Weighted tests (non-Gaussian). We show that, when comparing both arms, validity of these tests requires the sampling scheme to be translation invariant in a way we make precise. We propose translation-invariant versions of Thompson, tempered greedy, and tempered Upper Confidence Bound sampling. Simulation results corroborate our asymptotic analysis.
appendix boundary found by appendix_command · 80% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Van der Vaart, A. W (2000) Asymptotic statistics | 1.000 | 7 | 3 | 100% |
| 2 | Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments | 1.000 | 5 | 3 | 100% |
| 3 | Zhang, K., Janson, L., and Murphy, S (2021) Statistical inference with m-estimators on adaptively collected data | 1.000 | 5 | 3 | 100% |
| 4 | Kuang, X. and Wager, S (2024) Weak signal asymptotics for sequentially randomized experiments | 0.909 | 8 | 3 | 75% |
| 5 | Deshpande, Y., Mackey, L., Syrgkanis, V., and Taddy, M (2018) Accurate inference for adaptive linear models, in | 0.843 | 3 | 3 | 100% |
| 6 | Jeganathan, P (1995) Some aspects of asymptotic theory with applications to time series models | 0.737 | 3 | 2 | 100% |
| 7 | Fan, L. and Glynn, P. W (2025) Diffusion approximations for thompson sampling | 0.737 | 3 | 2 | 100% |
| 8 | Lai, T. L. and Robbins, H (1985) Asymptotically efficient adaptive allocation rules | 0.644 | 3 | 2 | 67% |
| 9 | Zhang, K., Janson, L., and Murphy, S (2020) Inference for batched bandits | 0.644 | 2 | 2 | 100% |
| 10 | Hallin, M., van den Akker, R., and Werker, B. J (2015) On quadratic expansions of log-likelihoods and a general asymptotic linearity result, in | 0.511 | 3 | 2 | 33% |
Showing the top 10 of 32 scored citations.