EconBase
← All papers

Demistifying Inference after Adaptive Experiments

Aurélien Bibaut, Nathan Kallus

arXiv 2 May 2024 · Statistics — Methodology

arXiv:2405.01281 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Adaptive experiments such as multi-arm bandits adapt the treatment-allocation policy and/or the decision to stop the experiment to the data observed so far. This has the potential to improve outcomes for study participants within the experiment, to improve the chance of identifying best treatments after the experiment, and to avoid wasting data. Seen as an experiment (rather than just a continually optimizing system) it is still desirable to draw statistical inferences with frequentist guarantees. The concentration inequalities and union bounds that generally underlie adaptive experimentation algorithms can yield overly conservative inferences, but at the same time the asymptotic normality we would usually appeal to in non-adaptive settings can be imperiled by adaptivity. In this article we aim to explain why, how, and when adaptivity is in fact an issue for inference and, when it is, understand the various ways to fix it: reweighting to stabilize variances and recover asymptotic normality, always-valid inference based on joint normality of an asymptotic limiting sequence, and characterizing and inverting the non-normal distributions induced by adaptivity.

Citation extraction

27
references
60
in-text mentions
27
distinct cited
5
self-citations
9,381
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Aurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz,… (2021) Post-contextual-bandit inference self0.87482100%
2Keisuke Hirano and Jack R Porter (2023) Asymptotic representations for sequential decisions, adaptive experiments, and batched bandits0.69381100%
3Karun Adusumilli (2023) Optimal tests following sequential experiments0.69371100%
4Aurelien Bibaut, Nathan Kallus, and Michael Lindon (2022) Near-optimal non-parametric sequential tests and confidence sequences with possibly dependent observations self0.69361100%
5Peter Hall and Christopher C Heyde (2014) Martingale limit theory and its application0.64441100%
6Aurélien Bibaut, Nathan Kallus, Maria Dimakopoulou, Antoine Chambaz,… (2021) Risk minimization from adaptively collected data: Guarantees for supervised and policy learning self0.51121100%
7Vitor Hadad, David A Hirshberg, Ruohan Zhan, Stefan Wager, and Susan… (2021) Confidence intervals for policy evaluation in adaptive experiments0.51121100%
8Alexander R Luedtke and Mark J Van Der Laan (2016) Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy0.51121100%
9LJ Wei, Robert T Smythe, DY Lin, and TS Park (1990) Statistical inference with data-dependent treatment allocation rules0.51121100%
10Ruohan Zhan, Vitor Hadad, David A Hirshberg, and Susan Athey (2021) Off-policy evaluation via adaptive weighting with data from contextual bandits0.51121100%

Showing the top 10 of 27 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Optimal Conditional Inference in Adaptive Experiments0.40511