EconBase
← All papers

Valid Post-Contextual Bandit Inference

Ramon van den Akker, Bas J. M. Werker, Bo Zhou

arXiv 20 May 2025 · Econometrics

arXiv:2505.13897 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We establish an asymptotic framework for the statistical analysis of the stochastic contextual multi-armed bandit problem (CMAB), which is widely employed in adaptively randomized experiments across various fields. While algorithms for maximizing rewards or, equivalently, minimizing regret have received considerable attention, our focus centers on statistical inference with adaptively collected data under the CMAB model. To this end we derive the limit experiment (in the Hajek-Le Cam sense). This limit experiment is highly nonstandard and, applying Girsanov's theorem, we obtain a structural representation in terms of stochastic differential equations. This structural representation, and a general weak convergence result we develop, allow us to obtain the asymptotic distribution of statistics for the CMAB problem. In particular, we obtain the asymptotic distributions for the classical t-test (non-Gaussian), Adaptively Weighted tests, and Inverse Propensity Weighted tests (non-Gaussian). We show that, when comparing both arms, validity of these tests requires the sampling scheme to be translation invariant in a way we make precise. We propose translation-invariant versions of Thompson, tempered greedy, and tempered Upper Confidence Bound sampling. Simulation results corroborate our asymptotic analysis.

Citation extraction

32
references
68
in-text mentions
32
distinct cited
0
self-citations
21,604
main-text words

appendix boundary found by appendix_command · 80% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Van der Vaart, A. W (2000) Asymptotic statistics1.00073100%
2Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments1.00053100%
3Zhang, K., Janson, L., and Murphy, S (2021) Statistical inference with m-estimators on adaptively collected data1.00053100%
4Kuang, X. and Wager, S (2024) Weak signal asymptotics for sequentially randomized experiments0.9098375%
5Deshpande, Y., Mackey, L., Syrgkanis, V., and Taddy, M (2018) Accurate inference for adaptive linear models, in0.84333100%
6Jeganathan, P (1995) Some aspects of asymptotic theory with applications to time series models0.73732100%
7Fan, L. and Glynn, P. W (2025) Diffusion approximations for thompson sampling0.73732100%
8Lai, T. L. and Robbins, H (1985) Asymptotically efficient adaptive allocation rules0.6443267%
9Zhang, K., Janson, L., and Murphy, S (2020) Inference for batched bandits0.64422100%
10Hallin, M., van den Akker, R., and Werker, B. J (2015) On quadratic expansions of log-likelihoods and a general asymptotic linearity result, in0.5113233%

Showing the top 10 of 32 scored citations.