Aurélien Bibaut, Nathan Kallus
arXiv 2 May 2024 · Statistics — Methodology
arXiv:2405.01281 · PDF · DOI · OpenAlex · Extracted main text
Adaptive experiments such as multi-arm bandits adapt the treatment-allocation policy and/or the decision to stop the experiment to the data observed so far. This has the potential to improve outcomes for study participants within the experiment, to improve the chance of identifying best treatments after the experiment, and to avoid wasting data. Seen as an experiment (rather than just a continually optimizing system) it is still desirable to draw statistical inferences with frequentist guarantees. The concentration inequalities and union bounds that generally underlie adaptive experimentation algorithms can yield overly conservative inferences, but at the same time the asymptotic normality we would usually appeal to in non-adaptive settings can be imperiled by adaptivity. In this article we aim to explain why, how, and when adaptivity is in fact an issue for inference and, when it is, understand the various ways to fix it: reweighting to stabilize variances and recover asymptotic normality, always-valid inference based on joint normality of an asymptotic limiting sequence, and characterizing and inverting the non-normal distributions induced by adaptivity.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Aurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz,… (2021) Post-contextual-bandit inference self | 0.874 | 8 | 2 | 100% |
| 2 | Keisuke Hirano and Jack R Porter (2023) Asymptotic representations for sequential decisions, adaptive experiments, and batched bandits | 0.693 | 8 | 1 | 100% |
| 3 | Karun Adusumilli (2023) Optimal tests following sequential experiments | 0.693 | 7 | 1 | 100% |
| 4 | Aurelien Bibaut, Nathan Kallus, and Michael Lindon (2022) Near-optimal non-parametric sequential tests and confidence sequences with possibly dependent observations self | 0.693 | 6 | 1 | 100% |
| 5 | Peter Hall and Christopher C Heyde (2014) Martingale limit theory and its application | 0.644 | 4 | 1 | 100% |
| 6 | Aurélien Bibaut, Nathan Kallus, Maria Dimakopoulou, Antoine Chambaz,… (2021) Risk minimization from adaptively collected data: Guarantees for supervised and policy learning self | 0.511 | 2 | 1 | 100% |
| 7 | Vitor Hadad, David A Hirshberg, Ruohan Zhan, Stefan Wager, and Susan… (2021) Confidence intervals for policy evaluation in adaptive experiments | 0.511 | 2 | 1 | 100% |
| 8 | Alexander R Luedtke and Mark J Van Der Laan (2016) Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy | 0.511 | 2 | 1 | 100% |
| 9 | LJ Wei, Robert T Smythe, DY Lin, and TS Park (1990) Statistical inference with data-dependent treatment allocation rules | 0.511 | 2 | 1 | 100% |
| 10 | Ruohan Zhan, Vitor Hadad, David A Hirshberg, and Susan Athey (2021) Off-policy evaluation via adaptive weighting with data from contextual bandits | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 27 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Optimal Conditional Inference in Adaptive Experiments | 0.405 | 1 | 1 |