arXiv 17 Feb 2026 · Statistics — Methodology
arXiv:2602.15559 · PDF · DOI · OpenAlex · Extracted main text
Adaptive randomized experiments update treatment probabilities as data accrue, but still require an end-of-study interval for the average treatment effect (ATE) at a prespecified horizon. Under adaptive assignment, propensities can keep changing, so the predictable quadratic variation of AIPW/DML score increments may remain random. When no deterministic variance limit exists, Wald statistics normalized by a single long-run variance target can be conditionally miscalibrated given the realized variance regime. We assume no interference, sequential randomization, i.i.d. arrivals, and executed overlap on a prespecified scored set, and we require two auditable pipeline conditions: the platform logs the executed randomization probability for each unit, and the nuisance regressions used to score unit $t$ are constructed predictably from past data only. These conditions make the centered AIPW/DML scores an exact martingale difference sequence. Using self-normalized martingale limit theory, we show that the Studentized statistic, with variance estimated by realized quadratic variation, is asymptotically N(0,1) at the prespecified horizon, even without variance stabilization. Simulations validate the theory and highlight when standard fixed-variance Wald reporting fails.
appendix boundary found by appendix_command · 75% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Cook, T., Mishler, A., & Ramdas, A (2024) March) | 0.693 | 6 | 1 | 100% |
| 2 | Kato, M., Ishihara, T., Honda, J., & Narita, Y (2020) Efficient adaptive experimental design for average treatment effect estimation | 0.644 | 4 | 1 | 100% |
| 3 | Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., & Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments | 0.585 | 3 | 1 | 100% |
| 4 | Kato, M., McAlinn, K., and Yasui, S (2021) The adaptive doubly robust estimator and a paradox concerning logging policy | 0.585 | 3 | 1 | 100% |
| 5 | Sengupta, S., Khamaru, K., Ghosh, S., & Dasgupta, T (2025) Design Stability in Adaptive Experiments: Implications for Treatment Effect Estimation | 0.585 | 3 | 1 | 100% |
| 6 | Zhan, R., Hadad, V., Hirshberg, D. A., & Athey, S (2021) August) | 0.585 | 3 | 1 | 100% |
| 7 | Li, H. H., & Owen, A. B (2024) Double machine learning and design in batch adaptive experiments | 0.511 | 2 | 1 | 100% |
| 8 | Bibaut, A., Dimakopoulou, M., Kallus, N., Chambaz, A., & van Der Laa… (2021) Post-contextual-bandit inference | 0.511 | 2 | 1 | 100% |
| 9 | Howard, S. R., Ramdas, A., McAuliffe, J., & Sekhon, J (2021) Time-uniform, nonparametric, nonasymptotic confidence sequences | 0.511 | 2 | 1 | 100% |
| 10 | Waudby-Smith, I., Arbour, D., Sinha, R., Kennedy, E. H., & Ramdas, A (2024) Time-uniform central limit theory and asymptotic confidence sequences | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 33 scored citations.