Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
19,038 characters · 9 sections · 16 citation commands
Inference for Batched Adaptive Experiments
Adaptive experiments have become increasingly common because they allow for earning while learning. Such designs have been applied, for example, by Kasy2021, Caria2023, Offer2021, Avivi2021, Tabord2023, Hoffmann2023, Gaul2025. They combine exploration and exploitation by updating treatment probabilities based on accumulated evidence. However, the dependence of assignment on past outcomes breaks the usual assumptions of random sampling and independent treatment assignment, complicating statistical inference. This is particularly problematic when there is no clear difference between outcomes under different treatments. For example, usual confidence intervals and bootstrap methods may overreject nullhypotheses. Hadad2021 use large number-of-trials asymptotics to construct generally valid confidence intervals. Zhangetal2020 note that typically the number of trials is limited but treatment assignment is adapted after each batch of observations arrives. For this important case, they derive valid frequentist inference procedures using large batch size asymptotics under homoskedasticity. This note extends their argument to the more general and empirically relevant case of heteroskedastic outcomes, deriving the corresponding BOLS (batched OLS) test statistic and explores its asymptotic distribution. The heteroskedastic case is relevant because researchers usually design experiments such that not only the outcome means but also their variances differ by treatment arm. Often the outcome is binary (success/failure), which results in heteroskedasticity by construction. The results are in line with Hirano2025 who show that any limit distribution generated by joint choices of adaptive assignment rules and statistics can be represented within a unified Gaussian limit experiment.
Let periods be indexed by $t = 1,\dots,T$. In period $t$, there are $N_{1,t}$ treated and $N_{0,t}$ control units, with $n_t=N_{1,t}+N_{0,t}$ and treatment share $\pi_t = N_{1,t}/n_t$, so $\pi_t(1-\pi_t)$ is a measure of balance. Let the per-period difference in sample means be
For each treatment arm $a \in \{0,1\}$, outcomes of individuals $i=1,2,...N_{a,t}$ satisfy $Y_{i,t}(a) = \mu_{a,t} + \varepsilon_{i,t}(a)$. Within period $t$, the treated and control sample means are independent with possibly different variances. Thus, $\varepsilon_{i,t}(1)$ is independent of $\varepsilon_{j,t}(0)$ for all $i,j$ with $\varepsilon_{i,t}(a) \stackrel{i.i.d.}{\sim} (0,\sigma^2_{a,t})$ and \[ \operatorname{Var}(\bar Y_{1,t})=\frac{\sigma_{1,t}^2}{N_{1,t}}, \qquad \operatorname{Var}(\bar Y_{0,t})=\frac{\sigma_{0,t}^2}{N_{0,t}}. \] Hence, the variance of the period difference is
In adaptive experiments, the selection probability $\pi_t$ is random because it depends on the realized history. Thus, the variance of the OLS estimator across periods depends on the selection probability which may result in asymptotic non-normality. Intuitively, if outcomes under two treatments are hard to distinguish, either treatment might get assigned more observations in repeated samples, and consequently the selection probability does not concentrate. Zhangetal2020 show that the selection probability is fixed, when conditioning on the history up to a given batch, and that the batchwise OLS, scaled by the selection probability, is asymptotically normal. We construct an estimator across periods scaled by the inverse standard error, such that each period’s standardized mean difference has the same influence. Let \[ w_t = \frac{1}{\sqrt{v_t}}, \qquad S = \sum_{s=1}^T w_s. \] Define the weighted average effect estimate
By construction, $w_t^2 v_t = 1$, hence the conditional variance given the realized weights is
Therefore,
In practice, the arm- and period-specific variances are unknown. Let $\hat\sigma_{a,t}^2$ be consistent estimators for $a\in\{0,1\}$ and the feasible test statistic $\hat{Z}_{het}$.
We compare three test statistics: the heteroskedasticity-robust OLS statistic, the BOLS statistic derived under homoskedasticity, and our heteroskedasticity-robust BOLS statistic. We report rejection rates, i.e., the proportion of simulated samples in which the null hypothesis is rejected. Under $H_{0}:\,\Delta = 0$, this quantity measures the empirical size of the test (nominal 5%). Data are generated using two common adaptive sampling algorithms, $\varepsilon$-Greedy and Bernoulli Thompson Sampling in two-arm settings.
\paragraph{Simulation I: Non-normality of OLS and homoskedastic BOLS}
Figure (ref) shows the empirical distributions of the test statistics with 25 batches each with a size of 500 observations. For $\varepsilon$-Greedy (left panels), both arms have Gaussian outcomes (think of log incomes) with mean 1, with variances $(1^2, 1^2)$ in panel a) and $(4^2,1^2)$ in panel c). The exploration rate is $\varepsilon=0.2$. Each design is repeated 100{,}000 times. Consistent with Zhangetal2020, panel a) shows that OLS under zero-margin is non-normally distributed and homoskedastic BOLS recovers normality. But in the same setting under heteroskedasticity (panel c), both OLS and the homoskedastic BOLS statistic deviate markedly from normality in the zero-margin case. The homoskedastic BOLS statistic severely overrejects (17% instead of 5%). OLS yields approximately correct rejection rates but exhibits non-normal behavior. In contrast, our heteroskedasticity-robust BOLS statistic closely matches the standard normal distribution and delivers correct 5% rejection rates in both cases.
For Thompson Sampling (right panels), we set Bernoulli success probabilities to $(p_1,p_2)=(0.5,0.5)$ in panel b) and to $(p_1,p_2)=(0.7,0.4)$ in panel d). Panel b) confirms that OLS is non-normally distributed and the homoskedastic BOLS test statistics fixes the problem. But under mechanical heteroskedasticity in the non-zero margin case (panel d) the homoskedastic BOLS statistic overrejects ($\approx6\%$) because it ignores heteroskedasticity. As the success probabilities differ substantially, both OLS and our heteroskedasticity-robust BOLS statistic closely match the standard normal distribution and deliver correct 5% rejection rates.
\paragraph{Simulation II: Rejection Rates at Small and Large Margins in Small and Large Samples}
To study behavior in smaller samples, we vary the number of batches (10–100) and batch sizes (10–100). Each configuration is repeated 10{,}000 times. Figures (ref) and (ref) report rejection rates of a 5% significance level test of $H_0:\Delta=0$.
For $\varepsilon$-Greedy (Figure (ref)), the homoskedastic BOLS statistic is unreliable in all settings: it overrejects sharply at $\Delta=0$ and underrejects for positive margins. The heteroskedasticity-robust OLS statistic performs reasonably well and improves as the margin grows. Across all designs, our heteroskedasticity-robust BOLS statistic maintains rejection rates close to 5%.
For Thompson Sampling (Figure (ref)), the zero-margin case exhibits no heteroskedasticity, so the heteroskedastic and the homoskedastic BOLS statistic perform well with rejection rates near 5%. The heteroskedasticity-robust OLS overrejects somewhat. When margins increase, inducing heteroskedasticity, the homoskedastic BOLS statistic begins to overreject, while both robust statistics remain close to nominal size.
Adaptive experiments have made selection-weighted inference increasingly relevant in sequential settings where treatment assignment depends on past outcomes. The BOLS selection-weighted statistic provides a simple, asymptotically valid procedure for inference under heteroskedasticity, extending previous results derived for the homoskedastic case.