EconBase
← Back to paper

Efficient Adaptive Experimental Design for Average Treatment Effect Estimation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

88,923 characters · 28 sections · 111 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Efficient Adaptive Experimental Design for Average Treatment Effect Estimation

abstractWe study how to efficiently estimate average treatment effects (ATEs) using adaptive experiments. In adaptive experiments, experimenters sequentially assign treatments to experimental units while updating treatment assignment probabilities based on past data. We start by defining the efficient treatment-assignment probability, which minimizes the semiparametric efficiency bound for ATE estimation. Our proposed experimental design estimates and uses the efficient treatment-assignment probability to assign treatments. At the end of the proposed design, the experimenter estimates the ATE using a newly proposed Adaptive Augmented Inverse Probability Weighting (A2IPW) estimator. We show that the asymptotic variance of the A2IPW estimator using data from the proposed design achieves the minimized semiparametric efficiency bound. We also analyze the estimator's finite-sample properties and develop nonparametric and nonasymptotic confidence intervals that are valid at any round of the proposed design. These anytime valid confidence intervals allow us to conduct rate-optimal sequential hypothesis testing, allowing for early stopping and reducing necessary sample size. \footnote{This paper was first made public in February 13, 2020 at \url{https://arxiv.org/abs/2002.05308v1}, and was presented at the NeurIPS 2020 Workshop on Causal Discovery & Causality-Inspired Machine Learning, as well as workshops at the University of Tokyo and Keio University. We thank the seminar participants for their valuable comments and feedback. We would also like to express our deep gratitude to Professor Hidehiko Ichimura for his insightful guidance and encouragement.}

Introduction

Adaptive experiments are increasingly common in the social sciences, the tech industry, and medicine. In adaptive experiments, experimenters sequentially assign treatments to experimental units while updating treatment assignment probabilities based on past data. Compared to the non-adaptive randomized control trial (RCT), adaptive designs often allow experimenters to more efficiently or quickly detect causal effects, thus exposing fewer experimental units to costly or harmful treatments. This merit has led organizations such as the US Food and Drug Administration to recommend adaptive designs fda. Adaptive experiments also produce social and economic applications and spark theoretical interest.

This paper studies how to design an adaptive experiment for efficient estimation of the average effects of treatment (ATE) and hypothesis testing. Let $Y(1), Y(0) \in \mathcal{Y}$ be potential outcomes of treatment $1$ and control $0$, respectively, where $\mathcal{Y} \subset \mathbb{R}$ is a bounded outcome space (see Assumption (ref)). Let $X \in \mathcal{X}$ be covariates, where $\mathcal{X}$ represents a space of covariates. The random variables $(X, Y(1), Y(0))$ jointly follow an unknown distribution $P_0 \in \mathcal{P}$, where $\mathcal{P}$ is the set of the distributions over $(X, Y(1), Y(0))$. We are interested in the estimation of average treatment effect (ATE), defined as

align*[align* omitted — 73 chars of source]

where $\mathbb{E}[Y(a)]$ denotes the mean potential outcome for each treatment $a\in\{1, 0\}$. The experiment involves $T \in \mathbb{N}$ experimental units, who are assigned to the treatment ($1$) or the control ($0$). For each $t \in [T]$, let $(X_t, Y_t(1), Y_t(0))$ be an i.i.d. draw of $(X, Y(1), Y(0))$ following the distribution $P_0$.

We propose the following adaptive experiment consisting of (1) a treatment-assignment phase and (2) an ATE-estimation phase using a novel estimator. :

itemize• Step 1. Treatment-assignment phase: \begin{itemize} • In each round $t \in [T] = {1, 2, \dots, T}$, an experimental unit with covariate $X_t \in \mathcal{X}$ visits the experimenter; • The experimenter assigns treatment $A_t \in \{1, 0\}$ with probability $\pi_t(a \mid X_t, \mathcal{H}_{t-1})$, based on the covariate $X_t$ and past observations \[\mathcal{H}_{t-1} \coloneqq \{X_1, A_1, Y_1, X_2, \dots, Y_{t-2}, X_{t-1}, A_{t-1}, Y_{t-1}\},\] where $Y_t = \mathbbm{1}[A_t = 1]Y_t(1) + \mathbbm{1}[A_t = 0]Y_t(0)$ is the observed outcome; • After treatment assignment, the experimenter observes the outcome $Y_t \in \mathbb{R}$; \end{itemize} • Step 2. ATE-estimation phase: \begin{itemize} • We estimate ATE $\theta_0$ using observations \[\mathcal{H}_{T} = \{(X_i, A_i, Y_i)\}^T_{i=1}.\] \end{itemize}

The treatment-assignment probability can be updated after each round based on the observations collected up to that point. Our method is also applicable to batch settings, where updates occur only in specified rounds. The treatment-assignment probability $\pi_t$ is usually called a propensity score in observational studies.

Note that from our assumption that $(X_t, Y_t(1), Y_t(0))$ is i.i.d. over $t\in[T]$, the Stable Unit Treatment Value Assumption (SUTVA) holds imbens_rubin_2015. Furthermore, unconfoundedness also holds from the construction of the treatment assignment probability $\pi_t(a\mid X_t, \mathcal{H}_{t-1})$; that is, outcomes $(Y_t(1), Y_t(0))$ and treatment $A_t$ are conditionally independent given $X_t$ and $\mathcal{H}_{t-1}$.

In addition to ATE estimation, we also analyze hypothesis testing about $\theta_0$ with null and alternative hypotheses defined for some $\mu \in \mathbb{R}$ as

align[align omitted — 85 chars of source]

We begin by investigating the semiparametric efficiency bound for ATE estimators. Following the approach of Hahn2011, we minimize the semiparametric efficiency bound with respect to treatment-assignment probabilities and define the minimizer as the efficient treatment-assignment probability. This efficient treatment-assignment probability is expressed as the ratio of the covariate-conditional standard deviations of the potential outcomes. This treatment-assignment probability is a variant of the one proposed in Neyman1934OnTT, which recently has been called the Neyman allocation.

Step (1) of our adaptive experiment sequentially estimates these conditional standard deviations, calculates the efficient treatment-assignment probability, and assigns treatment based on this estimate. To implement Step (2) of efficient ATE estimation, we introduce and use an ATE estimator, which we call the Adaptive Augmented Inverse Probability Weighting (A2IPW) estimator, which is a variant of the Augmented Inverse Probability Weighting (AIPW) estimator designed for adaptive experiments BangRobins2005.

We analyze both the infinite-sample and finite-sample properties of the A2IPW estimator. In the infinite-sample analysis, we demonstrate its consistency and asymptotic normality, showing that its asymptotic variance reaches the minimized semiparametric efficiency bound.

We then study hypothesis testing under two frameworks: single-stage testing and sequential testing. In the single-stage approach, we perform standard hypothesis testing by constructing confidence intervals with a fixed sample size to decide whether to reject the null hypothesis. In the sequential testing approach, the sample size is not fixed; instead, we continue collecting data until a decision can be made with a predetermined Type I error probability. Sequential testing has the potential to reduce the sample size by stopping the adaptive experiment early.

We propose a sequential testing procedure based on the finite-sample analysis of our estimator. Specifically, we derive a confidence interval that is nonparametric and non-asymptotic; it does not rely on a distributional assumption and an asymptotic approximation. We derive our confidence interval based on the Law of the Iterated Logarithm Balsubramani2016,Howard2020TimeuniformNN. In addition, our confidence intervals are Bernstein-type and use information about the variance of potential outcomes. As a result, our sequential testing with LIL-type anytime valid confidence intervals is rate-optimal for stopping time and effectively reduces the sample size Jamieson2014. In particular, our confidence intervals are narrower than other confidence intervals, such as those based on Hoeffding's inequality, which rely solely on the boundedness of outcomes.

\color{black}

Related Work

This study contributes to the growing work on adaptive experimental design for efficient estimation and inference of treatment effects. Important problems include how to design treatment assignment probabilities Hahn2011 and how to make statistical decisions Manski2000. This paper addresses those problems by designing an adaptive experiment for efficiently estimating the ATE with associated hypothesis testing and decision-making methods.

Compared to existing studies such as Hahn2011, our adaptive experiment offers the following advantages:

description• Our proposed design does not require dividing experimental units into discrete prespecified batches (though our design can also be used in such batch settings). Without prefixing the sample size for batches, our approach allows for the sequential construction of the optimal treatment assignment. • Our experiment does not require the Donsker condition for the estimators of the nuisance parameters (i.e., the conditional expected outcome and the efficient treatment-assignment probability). Instead, we impose convergence rate conditions for the estimators, similar to double machine learning in ChernozhukovVictor2018Dmlf. This flexibility allows us to use a variety of machine learning estimators for estimating nuisance parameters. • We do not require specific assumptions (such as discrete support) on the covariate distribution as long as the convergence rate conditions are satisfied.

Furthermore, our study examines the finite sample properties of ATE estimation and the sequential testing method.

Kato2021adr complements this work by highlighting that the proposed A2IPW estimator is a variant of double machine learning. They generalize the A2IPW estimator into the Adaptive Doubly Robust (ADR) estimator, which enables the estimation of the treatment-assignment probability. Their findings indicate that empirical performance can be improved by replacing the treatment-assignment probability with its estimator, even when the true value of the treatment-assignment probability is known. For a detailed discussion of double machine learning in adaptive experiments, see their paper. kato2021adaptivedoublyrobustestimator further extends the ADR estimator for the case where the average of the treatment-assignment probability converges to a constant, even if the probability itself does not converge.

Our method and the framework for adaptive experimental design for ATE estimation have been extended in various directions. Some works have relaxed our assumptions cook2023semiparametric, Waudby-Smith2024, and others have adapted our proposed estimator for cases with unknown treatment-assignment probabilities Kato2021adr, li2023double. Vikas2024 refines the asymptotic optimality in this problem. gupta2021efficient and chandak2024adaptive address endogeneity problems with instrumental variables, while li2024privacy explore privacy-preserving aspects. Simchi-Levi2023 investigates the setting under nonstationarity. zrnic2024active, kato2024active and ao2024predictionguidedactiveexperiments introduce the idea of active learning for this problem setting.

The framework of Hahn2011 is called a stratified experiment, where experimental units are divided into several strata based on their covariates Bugni2018,Bugni2019. Concurrently with our work, Meehan2022 proposes a stratification method based on a tree-based algorithm within a two-stage experimental framework to relax the assumption of discrete support in Hahn2011. In contrast, our algorithm does not depend on specific models or algorithms for determining treatment-assignment probabilities or for estimating the ATE. Instead, our method incorporates double machine learning techniques into our experimental design ChernozhukovVictor2018Dmlf, allowing for a wide range of traditional and modern machine learning estimators. Furthermore, our method is applicable to various settings of adaptive experimental design, including two-stage, multi-stage, and sequential experiments.

Furthermore, after the initial public draft of this paper Kato2020adaptive, several related studies have emerged. Kallus2021, bai2025efficiencyfinelystratifiedexperiments, and rafi2023efficientsemiparametricestimationaverage discuss efficiency bounds or efficient experiments under the stratification setting. Armstrong2022 and hirano2023asymptotic investigate asymptotically optimal treatment rules in adaptive experiments. Furthermore, Cai2024 and zhao2023adaptive investigate the Neyman allocation from perspectives different from ours.

We derive the asymptotic distribution of our A2IPW estimator using martingale theory. Notably, our asymptotic normality result does not require the Donsker condition for the nuisance parameter estimator. This approach is similar in spirit to sample-splitting methods used in the semiparametric analysis, such as double machine learning klaassen1987,ZhengWenjing2011CTME,ChernozhukovVictor2018Dmlf. hadad2019 also independently proposes a closely related estimator, including ATE estimation, for bandit problems, focusing on cases where the treatment-assignment probability approaches zero at a certain rate with respect to $t$.

Efficient estimation with adaptive experiments is closely related to the Best Arm Identification (BAI) problem in multi-armed bandit (MAB) settings Bubeck2009, Kasy2021. Neyman allocation is known to be optimal in BAI problems under certain conditions, such as Gaussian outcomes when variances are known chen2000, glynn2004large, Kaufman2016complexity. When variances are unknown, our proposed A2IPW strategy is still optimal in BAI as the ATE approaches zero adusumilli2022minimax,kato2025generalizedneymanallocationlocally. adusumilli2022minimax proves that the Neyman allocation is minimax optimal for the BAI problem. Armstrong2022 and Adusumilli2021risk study asymptotic treatment rules in adaptive experiments. In the setting of BAI, kato2025generalizedneymanallocationlocally generalizes the Neyman allocation for the multi-armed case. In BAI problems with covariates, researchers investigate identifying the best treatment arm based on expected outcomes marginalized over the covariate distribution or the conditional on covariates Russac2021, kato2021role, simchi2024experimentation, kato2024adaptivepolicylearning. SimchiLevi2023 and Caria2023 integrate the statistical inference problem with the regret minimization problem in MAB.

This study investigates the finite-sample property of the AIPW estimator in adaptive experiments. Our non-asymptotic error analysis is based on the law of the iterated logarithm Darling1967,Howard2020TimeuniformNN. The LIL plays an important role in finite-sample analysis and sequential testing since it is known to return tighter confidence intervals. Balsubramani2016 propose nonparametric sequential testing using the LIL, and we apply their results to adaptive ATE estimation with the A2IPW estimator. Smith2024 and Cai2024 also address the finite-sample analysis.

Organization

This study is organized as follows. In Section (ref), we introduce the data-generating process and discuss the semiparametric efficiency bound. In Section (ref), we design an adaptive experiment for efficient ATE estimation and powerful hypothesis testing, and we also propose the A2IPW estimator. Next, in Section (ref), we present the theoretical properties of our A2IPW estimator, focusing on its asymptotic normality, efficiency, and non-asymptotic results. Notably, its asymptotic variance aligns with the semiparametric efficiency bound. In Section (ref), we examine hypothesis testing under our adaptive experimental framework. Finally, in Section (ref), we assess the empirical performance of the proposed method using both synthetic and semi-synthetic data. All Appendices are provided in the online supplementary materials.

Semiparametric Efficiency Bound and Efficient Assignment Probability

Semiparametric Efficiency Bound in Adaptive Experimental Design

This section provides a lower bound for the asymptotic variance of regular estimators of the ATE in adaptive experiments, following the arguments in Hahn2011. Specifically, we focus on the semiparametric lower bound, which establishes a theoretical limit for the asymptotic variances of regular ATE estimators under semiparametric models.\footnote{The asymptotic variance can also be interpreted as the asymptotic mean squared error when the ATE estimator is asymptotically normal. Consequently, the semiparametric lower bound serves as a lower bound for the estimation error.}

Consider i.i.d. observations $\{(X_i, A_i, Y_i)\}_{i=1}^n$ generated from a distribution $P_0$ with a treatment-assignment probability $\pi_0(a \mid X_i)$. This treatment-assignment probability can be optimized for ATE estimation; thus, we refer to an algorithm with such a treatment-assignment probability as an oracle algorithm. In this case, from Theorem 1 of Hahn1998, the semiparametric efficiency bound is given as follows:

proposition[Semiparametric efficiency bound of ATE estimators. Based on Theorem 1 of Hahn1998.] Suppose that the same regularity conditions assumed in Theorem 1 of Hahn1998 hold. Under an oracle algorithm with treatment-assignment probability $\pi_0$, the asymptotic variance of regular ATE estimators is lower bounded by \begin{align} V(\pi_0) \coloneqq \mathbb{E}_{P_0}\left[\frac{\sigma^2_0\big(1, X\big)}{\pi_0(1 \mid X)} + \frac{\sigma^2_0\big(0, X\big)}{\pi_0(0 \mid X)} + \Big(\theta_0(X) - \theta_0\Big)^2\right], \end{align} where $\sigma^2_0\big(a, X\big)$ is the conditional variance of $Y(a)$ given $X$ for $a\in\{1, 0\}$.

This proposition corresponds to the case where the oracle treatment-assignment probability $\pi_0$ is known in advance, eliminating the need for estimation during the adaptive experiment. In such a scenario, treatments are assigned directly using the oracle treatment-assignment probability. This static oracle algorithm serves as a benchmark in the study of adaptive experimental designs Hahn2011.

If we restrict the algorithm to those where $\pi_t \xrightarrow{\mathrm{p}} \pi_0$ as $t \to \infty$, the result can be extended to non-i.i.d. observations using the martingale central limit theorem, as demonstrated in the derivation of asymptotic normality. Various extensions of lower bounds also have been proposed li2023double,rafi2023efficientsemiparametricestimationaverage.

Efficient Treatment-assignment Probability

In the semiparametric efficiency bound ((ref)), decision-makers can select $\pi_0$ to minimize the asymptotic variance. Denote the efficient treatment-assignment probability by \[ \pi^* \coloneqq \operatorname*{arg\,min}_{\pi_0 \in \Pi} V(\pi_0). \] The minimization problem has a closed-form solution, as shown below:

proposition[Efficient treatment-assignment probability] The efficient treatment-assignment probability $\pi^*$ is: \begin{align*} \pi^*(a \mid x) = \frac{\sqrt{\sigma^2_0(a, x)}}{\sqrt{\sigma^2_0(1, x)} + \sqrt{\sigma^2_0(0, x)}}, \quad \forall a \in \{1, 0\},\ \forall x \in \mathcal{X}. \end{align*}

The proof is presented in Appendix (ref).

Intuitively, conditional on $x$, the asymptotic variance can be minimized by assigning the treatment with a higher variance of the potential outcome. This treatment-assignment probability is recently referred to as the Neyman allocation Neyman1934OnTT and has been investigated in various studies on experimental design chen2000,glynn2004large,Laan2008TheCA,Hahn2011,Kaufman2016complexity,Meehan2022.

Semiparametric Efficient Adaptive Experiment

In this section, we design an adaptive experiment that minimizes the semiparametric efficiency bound and an ATE estimator whose asymptotic variance hits the minimized semiparametric efficiency bound. As explained in the Introduction, our experiment consists of two steps:

itemize• Step (1). Treatment-assignment phase: In each round $t\in[T]$, we estimate the efficient treatment-assignment probability $\pi^*(a\mid x)$ and assign a treatment based on the estimated efficient treatment-assignment probability. • Step (2). ATE-estimation phase: At the end of the experiment, we estimate the ATE using our proposed A2IPW estimator.

The pseudo-code is provided in Algorithm (ref). In the following subsections, we explain the details of our experimental design.

Step (1): Treatment-Assignment Phase

We assign treatments in each round $t \in [T]$ to gather data. Although assigning treatments with probability $\pi^*$ minimizes the semiparametric efficiency bound, it is infeasible since we do not know the conditional variance $\sigma^2_0(a, x)$. To overcome this challenge, in each round $t$, we estimate the conditional variance $\sigma^2_0(a, x)$, estimate the efficient treatment-assignment probability $\pi^*$ using the estimator of $\sigma^2_0(a, x)$, and assign a treatment based on the estimated efficient treatment-assignment probability.

Let $T_0$ ($2 \leq T_0 \leq T$) be the number of initialization rounds, which is a constant independent of $T$. In the initialization rounds $t = 1, 2, \dots, T_0$, we assign treatment $A_t = 1$ if $t$ is odd and $A_t = 0$ if $t$ is even; for example, if $T_0 = 6$, $(A_1, A_2, A_3, A_4, A_5, A_6) = (1, 0, 1, 0, 1, 0)$. We set $\pi_t(1\mid X_t, \mathcal{H}_{t-1}) = 1/2$ for all $t = 1, 2, \dots, T_0$.

In each round $t \in \{T_0 + 1, T_0 + 2, \dots, T\}$, we construct a consistent estimator $\widehat{\sigma}^2_t(a, x)$ of $\sigma^2_0(a, x)$ such that $\widehat{\sigma}^2_t(a, x) \in (0, \infty)$ for all $a \in \{1, 0\}$ and $x \in \mathcal{X}$, and $\widehat{\sigma}^2_t(a, x)$ is constructed only by using $\mathcal{H}_{t-1}$. The reason we use only $\mathcal{H}_{t-1}$ is to construct an ATE estimator whose scores consist of a martingale difference sequence, as shown in the next subsection. Under this property, we can apply the martingale central limit theorem and martingale concentration inequality to analyze the asymptotic and non-asymptotic behaviors of the ATE estimator.

To estimate $\sigma^2_0(a, X_t)$, we propose estimating $f_0(a, X_t) = \mathbb{E}[Y_t(a) \mid X_t]$ and $e_0(a, X_t) = \mathbb{E}[Y_t^2(a) \mid X_t]$ using nonparametric models based on observations $\mathcal{H}_{t-1}$ up to round $t$. Let $\widehat{f}_t(a, X_t)$ and $\widehat{e}_t(a, X_t)$ denote such estimators. In MAB problems, several nonparametric estimators, such as $K$-nearest neighbor regression and Nadaraya–Watson kernel regression, have been shown to be consistent yang2002,Qian2016. For example, given a bandwidth $h_T > 0$ and a kernel function $K:\mathcal{X} \to \mathbb{R}$, a Nadaraya-Watson estimator of $f_0(a, X_t)$ is defined as $\widehat{f}_t(a, X_t) = \frac{1}{\frac{1}{t - 1}\sum^{t-1}_{s=1}\mathbbm{1}[A_s = a]K((X_s - X_t)/h_t)}\frac{1}{t - 1}\sum^{t-1}_{s=1}Y_s\mathbbm{1}[A_s = a]K((X_s - X_t)/h_t)$. We can also estimate $e_0(a, X_t)$. By appropriately obtaining samples, we can also employ random forests WagerAthey2018 and neural networks as nonparametric estimators SchmidtHieber2020,Farrell2021.

We then estimate $\sigma^2_0(a, X_t)$ as follows: \[ \widehat{\sigma}^2_t =

cases\widehat{e}_t(a, X_t) - \widehat{f}^2_t(a, X_t) & if \widehat{e}_t(a, X_t) - \widehat{f}^2_t(a, X_t) > 0, \\ \varepsilon & otherwise,

\] where $\varepsilon > 0$ is a small positive constant introduced to ensure that $\widehat{\sigma}^2_t$ remains non-negative. Note that when $\sigma^2_0(a, X_t) > 0$, the term $\varepsilon$ becomes unnecessary as $t$ grows large.

We assign treatment $A_t$ with probability $\pi_t(A_t \mid X_t, \mathcal{H}_{t-1})$, defined as \[ \pi_t(a \mid X_t, \mathcal{H}_{t-1}) = \frac{\sqrt{\widehat{\sigma}^2_t(a, X_t)}}{\sqrt{\widehat{\sigma}^2_t(1, X_t)} + \sqrt{\widehat{\sigma}^2_t(0, X_t)}} \quad \forall a \in \{1, 0\}, \] Note that our experiment can be used in a batch setting, where we update $\pi_t(a \mid x, \mathcal{H}_{t-1})$ only at certain rounds $T_1, T_2, \dots \in \{1, \dots, T\}$. We require that $\pi_t(a \mid x, \mathcal{H}_{t-1}) \to \pi^*(a \mid x)$ for each $x \in \mathcal{X}$ as $t \to \infty$.\footnote{As long as this condition is satisfied, we do not need to sequentially update $\pi_t(a \mid x, \mathcal{H}_{t-1})$. This implies that we can keep $\pi_t(a \mid x, \mathcal{H}_{t-1})$ constant for several rounds and update $\pi_t(a \mid x, \mathcal{H}_{t-1})$ in specific rounds. For example, we can consider a two-stage design similar to Hahn2011. In this case, we update $\pi_t(a \mid x, \mathcal{H}_{t-1})$ only at $T_1$. Assume that $T_1 = r T$, where $r \in (0, 1)$ is a constant independent of $T$. In rounds $1, 2, \dots, T_1$, we assign treatment $a \in \{1, 0\}$ with probability $1/2$, where $\pi_t(a \mid x, \mathcal{H}_{t-1}) = 1/2$ for all $x \in \mathcal{X}$. Afterward, we update $\pi_t(a \mid x, \mathcal{H}_{t-1})$ by estimating $\pi^*(a \mid x)$. If $\pi_t(a \mid x, \mathcal{H}_{t-1}) \to \pi^*(a \mid x)$ as $t \to \infty$ holds for all $x \in \mathcal{X}$, we can prove the same asymptotic optimality of our experimental design. To verify that $\pi_t(a \mid x, \mathcal{H}_{t-1}) \to \pi^*(a \mid x)$ as $t \to \infty$, it is sufficient to check $\pi_{T_1}(a \mid x, \mathcal{H}_{t-1}) \to \pi^*(a \mid x)$ as $T_1 \to \infty$ ($T \to \infty$).}

Step (2): ATE-Estimation Phase

At the end of the experiment, we construct an ATE estimator that is asymptotically normal with an asymptotic variance, achieving the semiparametric lower bound ((ref)). In adaptive experiments, due to the changing assignment probabilities, dependencies among samples can complicate the estimation process. To address this dependency problem, we propose the A2IPW estimator: \[ \widehat{\theta}^{\mathrm{A2IPW}}_T = \frac{1}{T}\sum^T_{t=1} \Psi_t, \]

align*[align* omitted — 429 chars of source]

and $\widehat{f}_t(a, x)$ is an estimator of $f_0(a, x)$, constructed from $\mathcal{H}_t$. As stated in Theorem (ref), the asymptotic optimality of our proposed ATE estimator holds with any consistent estimator for $f_0(a, x)$, due to the unbiasedness of $\widehat{\theta}^{\mathrm{A2IPW}}_T$ for $\theta_0$. This point is also discussed in Section 4 of Kato2021adr, our follow-up study. Additionally, consistency holds even if the estimator of $f_0(a, x)$ is inconsistent, as stated in Corollary (ref). Here, $\Psi_t$ is the semiparametric efficient score for ATE estimators. Regular estimators with scores $\Psi_t$ achieve the smallest asymptotic variance within the class of such estimators.

For $z_t = \Psi_t - \theta_0$, the sequence $\{z_t\}_{t=1}^T$ forms a martingale difference sequence, which means that $\mathbb{E}[z_t \mid \mathcal{H}_{t-1}] = 0$. Using this property, we will derive the theoretical results for $\widehat{\theta}^{\mathrm{A2IPW}}_T$. This construction shares a similar motivation to sample-splitting techniques in semiparametric inference klaassen1987, including double machine learning ChernozhukovVictor2018Dmlf.

Stabilizations and Extensions

While not required to obtain the asymptotic properties, here we introduce stabilization techniques that contribute to the finite-sample stabilization of the designed experiment. The above sections show that our designed experiment is asymptotically efficient in the sense that the asymptotic variance of the ATE estimator aligns with the semiparametric efficiency bound. However, such asymptotic optimality does not necessarily guarantee accurate ATE estimation in finite samples.

The ADR estimator. Kato2021adr reports that replacing the true $\pi_t$ with its estimate can paradoxically improve performance. This is because the original $\pi_t$ may take values close to zero, causing the inverse of $\pi_t$ to become large and making the A2IPW estimator unstable. By replacing $\pi_t$ with its estimate, even when the true value of $\pi_t$ is known, the A2IPW estimator can be stabilized.\footnote{Note that, unlike the classical problem regarding the use of an estimated propensity score in the IPW estimators, the asymptotic properties remain unchanged between the cases where the true $\pi_t$ is used and where $\pi_t$ is estimated when we use the AIPW estimator hirano03,Henmi2004paradox. }

The ADR estimator is defined as follows:

align*[align* omitted — 407 chars of source]

where $\widehat{g}_t(a \mid X_t)$ is an estimator of $\pi_t(a \mid X_t, \mathcal{H}_{t-1})$ constructed from the past observations $\{(X_s, A_s, Y_s)\}_{s=1}^{t-1}$. Although this estimator is no longer unbiased, asymptotic normality holds under convergence rate conditions for $\widehat{f}_t$ and $\widehat{\pi}_t$, as well as double machine learning techniques. The theorem regarding its asymptotic normality is introduced in Proposition (ref).

Kato2021adr refer to the sample splitting used in both the A2IPW and ADR estimators as adaptive fitting, where only past observations up to time $t$ are used to obtain the plug-in estimators for each $t$. Figure (ref) illustrates the difference between cross-fitting as described in ChernozhukovVictor2018Dmlf and our adaptive fitting approach.

figure[figure omitted — 355 chars of source]

Stabilization techniques. To stabilize the finite-sample behavior, we can further introduce certain elements into our experiment. These elements are designed not to affect the asymptotic behavior, meaning their influence vanishes as $t \to \infty$.

description• We define the treatment-assignment probability as \begin{align*} \pi_t(1 \mid x, \mathcal{H}_{t-1}) &= \gamma_t\frac{1}{2} + (1-\gamma_t)\frac{\sqrt{\widehat{\sigma}^2_{t-1}(1, x)}}{\sqrt{\widehat{\sigma}^2_{t-1}(1, x)} + \sqrt{\widehat{\sigma}^2_{t-1}(0, x)}},\\ \pi_t(0 \mid x, \mathcal{H}_{t-1}) &= 1 - \pi_t(1 \mid x, \mathcal{H}_{t-1}), \end{align*} where $\gamma_t = O(1/\sqrt{t})$; • As an alternative estimator, we propose the mixed A2IPW (MA2IPW) estimator, defined as $\widehat{\theta}^{\mathrm{MA2IPW}}_T = \zeta_T \widehat{\theta}^{\mathrm{IPW}}_T + (1-\zeta_T)\widehat{\theta}^{\mathrm{A2IPW}}_T$, where $\widehat{\theta}^{\mathrm{IPW}}_T$ is the IPW estimator defined as $ \widehat{\theta}^{\mathrm{IPW}}_T = \frac{1}{T}\sum^T_{t=1} \left(\frac{\mathbbm{1}[A_t=1]Y_t}{\pi_t(1\mid X_t, \mathcal{H}_{t-1})} - \frac{\mathbbm{1}[A_t=0]Y_t}{\pi_t(0\mid X_t, \mathcal{H}_{t-1})}\right)$ and $\zeta_T = o(1/\sqrt{T})$.

Note that the IPW estimator is a special case of the A2IPW estimator with $\widehat{f}_{t- 1}(x) = 0$.

Technique (a) aims to stabilize the treatment assignment probability. If $\pi_t$ fluctuates significantly, it may unstabilize of the A2IPW estimator for the ATE. Furthermore, when the value of $\pi_t$ is too close to zero, its inverse is included in the elements averaged by A2IPW, potentially causing those elements to become extremely large. Technique (a) prevents such cases.\footnote{Performance can be further improved by replacing $\pi_t$ with its estimator constructed from past observations $\{(X_s, A_s, Y_s)\}_{s=1}^{t-1}$ and $X_t$, as noted by Kato2021adr, a subsequent study to our study. Kato2021adr observes that the A2IPW estimator with the true $\pi_t$ incurs a larger mean squared error than when using an estimated $\pi_t$. This is because when the true $\pi_t$ fluctuates significantly, the A2IPW estimator also becomes unstable. However, Kato2021adr finds that replacing the volatile $\pi_t$ with a more stable estimator helps stabilize the behavior of the A2IPW estimator. The A2IPW estimator with an estimated $\pi_t$ is referred to as the Adaptive Doubly Robust (ADR) estimator. Although we do not focus on this type of stabilization in this study, it is a promising approach. We compare our estimator with the ADR estimator in our simulation studies. cook2023semiparametric also develops stabilization techniques based on Waudby-Smith2024.}

Technique (b) controls the estimator's behavior by avoiding situations where $\widehat{f}_{t-1}$ takes unpredictable values in the early stages. Since the nonparametric convergence rate is generally slower than $1/\sqrt{t}$, the convergence rate of $\pi_t$ to $\pi^*$ does not exceed $O(1/\sqrt{t})$. Therefore, $\gamma_t = O(1/\sqrt{t})$ does not asymptotically affect the convergence rate of the treatment-assignment probability. Similarly, the asymptotic distribution of $\widehat{\theta}^{\mathrm{MA2IPW}}_T$ is asymptotically equivalent to $\widehat{\theta}^{\mathrm{A2IPW}}_T$ because it holds that $ \sqrt{T} \widehat{\theta}^{\mathrm{MA2IPW}}_T = \sqrt{T}\big(\zeta_T \widehat{\theta}^{\mathrm{IPW}}_T + (1-\zeta_T)\widehat{\theta}^{\mathrm{A2IPW}}_T\big)=\sqrt{T} \widehat{\theta}^{\mathrm{A2IPW}}_T + o(1) $ as $T \to \infty$.

There are additional stabilization techniques. For example, cook2023semiparametric also develops stabilization techniques based on Waudby-Smith2024.

algorithm[algorithm omitted — 1,331 chars of source]

Theoretical Results about Treatment Effect Estimation

This section provides theoretical results on the A2IPW estimator. We present its asymptotic distribution, a regret bound, and a non-asymptotic confidence bound for the A2IPW estimator.

Consistency and Asymptotic Normality of the A2IPW Estimator

We first show the asymptotic normality of the A2IPW estimator $\widehat{\theta}^{\mathrm{A2IPW}}_T$. Before showing the asymptotic normality, we make the following assumption.

assumption[Boundedness] There exists an absolute constant $C$ such that $|Y_t(a)| \leq C$ holds for $a\in\{1, 0\}$.

The following theorem states the asymptotic normality.

theorem[Asymptotic distribution of the A2IPW estimator] Suppose that Assumption (ref) holds, and \begin{description} • point convergence in probability of $\widehat{f}_{t-1}$ and $\pi_t$, i.e., for all $x\in\mathcal{X}$ and $a\in\{0,1\}$, \begin{align*} &\widehat{f}_{t-1}(a, x)-f_0(a, x)\xrightarrow{\mathrm{p}}0\quad \mathrm{and}\quad \pi_t(a\mid x, \mathcal{H}_{t-1})-\widetilde{\pi}(a\mid x)\xrightarrow{\mathrm{p}}0, \end{align*} where $\widetilde{\pi} \in \Pi$; • there exists a constant $C_f$ such that $|\widehat{f}_{t-1}| \leq C_f$. \end{description} Then, the A2IPW estimator is asymptotically normal: $$\sqrt{T}\left(\widehat{\theta}^{\mathrm{A2IPW}}_T-\theta_0\right)\xrightarrow{d}\mathcal{N}\left(0, V\right),$$ where $$V \coloneqq \mathbb{E}\left[\frac{\widetilde{\sigma}^2\big(1, X_t\big)}{\widetilde{\pi}(1\mid X_t)} + \frac{\widetilde{\sigma}^2\big(0, X_t\big)}{\widetilde{\pi}(0\mid X_t)} + \Big(f_0(1, X_t) - f_0(0, X_t) - \theta_0\Big)^2\right].$$

The asymptotic variance aligns with the semiparametric efficiency bound derived under the treatment-assignment probability $\widetilde{\pi}$. Note that we do not have to impose the Donsker condition, similar to cross-fitting klaassen1987,ZhengWenjing2011CTME,ChernozhukovVictor2018Dmlf. Here, we do not impose the convergence rate of $\widehat{f}_{t-1}$ owing to the unbiasedness of the A2IPW estimator $\widehat{\theta}^{\mathrm{A2IPW}}_T$ for the ATE $\theta_0$.

Consistency holds under a weaker assumption, i.e., even if the treatment-assignment probability $\pi_t$ does not converge. We omit the proof because it follows from the boundedness of $z_t$ and the weak law of large numbers for a martingale difference sequence (Proposition (ref) in Appendix (ref)).

corollary[Consistency of the A2IPW estimator] Suppose that there exists a constant $C_f$ such that $|\widehat{f}_{t-1}| \leq C_f$. Then, under Assumption (ref), $\widehat{\theta}^{\mathrm{A2IPW}}_T\xrightarrow{\mathrm{p}} \theta_0$ holds as $T\to\infty$.

Note that Corollary (ref) holds even if $\widehat{f}_t$ is inconsistent. Therefore, compared to Theorem (ref), Corollary (ref) holds with a weaker assumption.

We also present the theorem about the asymptotic normality of the ADR estimator from Kato2021adr, which is a follow-up study that investigates the A2IPW estimator and generalizes it as the ADR estimator.

proposition[Asymptotic distribution of the ADR estimator. From Theorem 1 in Kato2021adr.] Suppose that Assumption (ref) holds, and \begin{description} • For all $x\in\mathcal{X}$ and $a\in\{0,1\}$, there exist $p, q > 0$ such that $p+q = 1/2$, it holds that $|\widehat{g}_{t-1}(a\mid x) -\widetilde{\pi}(a\mid x)|=\mathrm{o}_{p}(t^{-p})$, and $|\widehat{f}_{t-1}(a, x)-f_0(a, x)|=\mathrm{o}_{p}(t^{-q})$ where $\widetilde{\pi} \in \Pi$; • There exists a constant $C_f$ such that $|\widehat{f}_{t-1}| \leq C_f$. \end{description} Then, the ADR estimator is consitent and asymptotically normal: $$\sqrt{T}\left(\widehat{\theta}^{\mathrm{ADR}}_T-\theta_0\right)\xrightarrow{d}\mathcal{N}\left(0, V\right).$$

Regret Bound of the A2IPW Estimator

In addition to the above asymptotic analysis, we introduce the finite-sample regret framework often used in the literature on the MAB problem. We define regret based on the MSE. We define the optimal experiment $\Pi^{\mathrm{OPT}}$ as an experiment that chooses a treatment with the probability $\pi^*$ defined in Proposition (ref), and an estimator $\widehat{\theta}^\mathrm{OPT}_T$ with oracle $f_0$ as

multline*[multline* omitted — 263 chars of source]

For any experiment $\Pi$ adapted by the experimenter, we define the regret of $\Pi$ as \[{\tt regret} = \mathbb{E}_{\Pi}\left[\left(\theta_0 - \widehat{\theta}^{\mathrm{A2IPW}}_T \right)^2\right] -\mathbb{E}_{\Pi^{\mathrm{OPT}}}\left[\left(\theta_0-\widehat{\theta}^{\mathrm{OPT}}_T\right)^2\right],\] where the expectations are taken over each experiment. The following theorem provides an upper bound on the regret.

theorem[Regret Bound of A2IPW] Suppose that there exists a constant $C_f$ such that $|\widehat{f}_{t-1}| \leq C_f$. Then, under Assumption (ref), there exist constants $C > 0$ and $T_0$ such that for all $T > T_0$, it holds that \begin{align*} {\tt regret} &\leq \frac{C}{T^2}\sum_{a\in\{1,0\}}\sum^T_{t=1}\Bigg(\mathbb{E}\left[\Big| \sqrt{\pi^*(a\mid X_t)} - \sqrt{\pi_t(a\mid X_t, \mathcal{H}_{t-1})} \Big|\right]\\ &\ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ + \mathbb{E}\left[\Big| f_0(a, X_t) - \widehat{f}_{t-1}(a, X_t) \Big|\right]\Bigg), \end{align*} where the expectation is taken over the random variables including $\mathcal{H}_{t-1}$.

The proof is shown in Appendix (ref). This result tells us that regret is bounded by $o(1/T)$ under the consistencies of $\pi_t$ and $\hat f_t$, since under the consistencies and uniform integrability, as $T \to \infty$, it holds that

align*[align* omitted — 346 chars of source]

By contrast, if we use a constant value for $\pi_t$, regret is $O(1/T)$. The regret bound for finite samples can also be obtained by substituting the finite sample bounds of \[\mathbb{E}\left[\Big| \sqrt{\pi^*(a\mid X_t)} - \sqrt{\pi_t(a\mid X_t, \mathcal{H}_{t-1})} \Big|\right]\] and $\mathbb{E}\left[\Big| f_0(a, X_t) - \widehat{f}_{t-1}(a, X_t) \Big|\right]$. We can bound $\widehat{f}_{t-1}(a, X_t)$ and $\sqrt{\pi_t(a\mid X_t, \mathcal{H}_{t-1})}$ by the same argument as existing work on the MAB problem such as yang2002.

Any Time Confidence Interval

The asymptotic normality shown in the previous section holds for large fixed $T$. In this section, we consider a confidence interval that is valid for any $t \in [T]$. This type of anytime confidence interval guarantees a finite sample estimation error and plays an important role in sequential hypothesis testing.

Among the various candidates for constructing confidence intervals, we employ concentration inequalities based on the LIL. The LIL is originally derived as an asymptotic property of independent random variables by Khintchine1924 and Kolmogoroff1929. Following their methods, several works have derived an asymptotic LIL for a martingale difference sequence under some regularity conditions Stout1970,Fisher1992. Balsubramani2016 derived a non-asymptotic LIL-based concentration inequality for sequential testing. The reason for using the LIL-based concentration inequality is that sequential testing with the LIL-based confidence sequence requires a smaller sample size needed to identify the parameter of interest since the confidence intervals depend on the distributional information and are said to be tight Jamieson2014, as explained later. Due to the tightness of the inequality, LIL-based concentration inequalities have been widely accepted in sequential testing Balsubramani2016 and in the best arm identification in the Multi-Armed Bandit (MAB) problem Jamieson2014,NIPS2018_7624.

We construct the confidence sequence $\big\{q_t\big\}_{t\in\mathbb{N}}$ based on the LIL-based concentration inequality for the A2IPW estimator as follows.

theorem[Concentration Inequality of the A2IPW Estimator] Suppose that the null hypothesis is correct; that is, $\mu = \theta_0$ and $z_t = \Psi_t - \theta_0$. Let $C > 0$ and $C_z > 0$ be constants independent of $t$ and $T$ such that $|z_t| \leq C$ and $|(z_t - z_{t-1})^2 - \mathbb{E}[(z_t - z_{t-1})^2| \mathcal{H}_{t-1}]| \leq C_z$ hold. For any $\delta$, with probability $\geq 1-\delta$, for all $t\geq T_0$ simultaneously, \begin{align*} &\left|\sum^t_{i=1}z_i\right| = t\left|\widehat{\theta}^{\mathrm{A2IPW}}_t-\theta_0\right| \leq \frac{2C}{e^2}\left(C_0(\delta) + \sqrt{2C_1\widehat{V}^*_t\left(\log\log\widehat{V}^*_t + \log\left(\frac{4}{\delta}\right)\right)}\right), \end{align*} where $\widehat{V}^*_t = C_f\left(\frac{e^4}{4C^2}\sum^t_{i=1}z^2_i + \frac{2C_0(\delta)C_z}{e^2}\right)$, $C_0(\delta)=3(e-2) + 2\sqrt{\frac{173}{2(e-2)}} \log\left(\frac{4}{\delta}\right)$, $C_1 = 6(e-2)$ and $C_f$ is an absolute constant.

The proof is provided in Appendix (ref). This result, derived by applying the findings of Balsubramani2014SharpFI, not only shows an anytime confidence interval but also establishes a finite-sample estimation error bound in estimating $\theta_0$.

We obtain confidence sequences, $\big\{q_t\big\}^T_{t=1}$, with the Type I error at $\alpha$ from the results of Theorem (ref) and Balsubramani2016 as $$ q_t \propto \log \left(\frac{1}{\alpha}\right) + \sqrt{2\sum^t_{i=1}z^2_i\left(\log\frac{\log \sum^t_{i=1}z^2_i}{\alpha}\right)}. $$ Balsubramani2016 proposes using constant $1.1$ to specify $q_t$, namely, $$ q_t = 1.1\left(\log \left(\frac{1}{\alpha}\right) + \sqrt{2\sum^t_{i=1}z^2_i\left(\log\frac{\log \sum^t_{i=1}z^2_i}{\alpha}\right)}\right). $$ This choice is motivated by the asymptotic property of the LIL such that $$ \limsup_{t\to\infty}\frac{\left|t\widehat{\theta}^{\mathrm{A2IPW}}_t - t\theta_0\right|}{\sqrt{2\widetilde{V}^*_t\left(\log\log\widetilde{V}^*_t\right)}} = 1 $$ with probability $1$ for sufficiently large samples Stout1970,Balsubramani2016, where $\tilde{V}^2_t = \sum^t_{i=1}\mathbb{E}[z^2_i\mid \mathcal{H}_{i-1}]$, as well as the empirical results of Balsubramani2016.

The confidence interval tightly depends on the underlying distribution through the variances, in contrast to confidence intervals that rely on less information, such as those based on Hoeffding's inequality, which only uses the boundedness of the outcomes. Additionally, the tightness of the confidence intervals is also guaranteed by the lower bounds for sequential testing, as discussed in Jamieson2014. It is known that the $O(\sqrt{t^{-1}\log\log t})$ asymptotic rate of the confidence intervals aligns with the lower bound implied by the LIL Farrell1964. Non-asymptotic bounds of this form are referred to as finite LIL bounds Howard2020TimeuniformNN.

Hypothesis Testing

This section studies hypothesis testing about the ATE. We begin by formulating the hypothesis testing framework, introduce the testing procedures, and conclude by presenting the theoretical properties. We demonstrate how to compute the required sample size for hypothesis testing. The pseudo-code for our experimental design incorporating hypothesis testing is provided in Algorithm (ref), which encompasses Algorithm (ref).

algorithm[algorithm omitted — 2,125 chars of source]

Hypothesis Testing in Adaptive Experiments

The experimenter aims to decide whether to reject the null hypothesis $H_0$ in ((ref)) while maximizing the power and controlling the Type I error. In adaptive experiments, hypothesis testing can be framed in two ways: single-stage testing and sequential hypothesis testing. In single-stage testing, the test is performed only at the end of the experiment ($t = T$). In sequential hypothesis testing, the test is conducted sequentially at each stage of the experiment, where the sample size is treated as a stopping time (a random variable). Sequential testing is expected to reduce the required sample size by allowing the experiment to stop earlier.

For each $t \in [T]$, let $\widehat{\theta}_T$ be an ATE estimator constructed using $\mathcal{H}_T$. In single-stage testing, we fix a threshold $p_T \in \mathbb{R}^+$ before gathering data via the experiment. At the end of the experiment, we reject the null hypothesis if: \[ T\left|\widehat{\theta}_T - \mu\right| > p_T. \] We can conduct the most powerful test in single-stage testing by using the $t$-test with our efficient ATE estimator, which is asymptotically normal.

In sequential testing, we define thresholds $q_t \in \mathbb{R}^+$ for each $t \in [T]$. At each time $t \in [T]$, we reject the null hypothesis if: \[ t\left|\widehat{\theta}_t - \mu\right| \geq q_t. \]

The difference between single-stage and sequential testing is illustrated in Figure (ref).

figure*[figure* omitted — 777 chars of source]

In single-stage testing, it is natural to analyze the power of the test. In contrast, in sequential testing, we focus on the expected sample size (stopping time). Both are related to the minimum required sample size under Type I error control.

Controlling the Type I and Type II errors. Recall that our null and alternative hypotheses are $H_0: \theta_0 = \mu$ and $H_1: \theta_0 \neq \mu$, respectively. Let $\mathbb{P}_{H_0}$ and $\mathbb{P}_{H_1}$ represent the probabilities when the null and alternative hypotheses are correct, respectively. When $\mathbb{P}_{H_0}\big(\text{reject}\ H_0\big)\leq \alpha$, we say that we control the Type I error at level $\alpha$. Similarly, when $\mathbb{P}_{H_1}\big(\text{reject}\ H_0\big)\geq 1-\beta$, we say that we control the Type II error, where $\beta$ is also referred to as the power of the test.

We set $p_T$ and $q_t$ to control both the Type I and Type II errors. We use asymptotic normality to construct $p_T$ and the LIL-based concentration inequality to construct $\{q_t\}_{t=1}^T$. Controlling errors in sequential testing is more complex than in single-stage testing. If we naively apply standard single-stage testing at each $t$ sequentially, the probability of a Type I error increases due to the multiple testing problem Balsubramani2016.\footnote{By contrast, the probability of a Type II error does not increase in sequential testing Balsubramani2016, though there are methods to control the Type II error more precisely NIPS2018_7624.} A common approach to this problem is to apply multiple testing corrections, such as the Bonferroni (BF) or Benjamini–Hochberg procedures. However, these methods tend to be overly conservative, resulting in suboptimal outcomes when conducting many tests. To avoid this issue in sequential hypothesis testing, we employ the non-asymptotic anytime confidence interval derived in Section (ref), which holds for any time $t$ 1512.04922,Howard2020TimeuniformNN.

Sample size and stopping time. We are interested in determining the sample size required to reject the null hypothesis while controlling the Type II error at level $\beta$, assuming the alternative hypothesis $H_1$ is true.

To control the Type II error, we introduce a parameter $\Delta > 0$, commonly referred to as the effect size in hypothesis testing literature. We redefine the alternative hypothesis as $H_1(\Delta): |\theta_0 - \mu| > \Delta$, where $\mathbb{P}_{H_1(\Delta)}$ represents the probability when the alternative hypothesis $H_1(\Delta)$ is correct. Let $R_n$ denote the rejection region for controlling the Type II error at level $\beta$ given $n$ observations. In other words, when $\widehat{\theta}^{\mathrm{A2IPW}}_n \in R_n$ and the alternative hypothesis $H_1$ is true, the null hypothesis is rejected with a probability of at least $1 - \beta$. For $\Delta$ and $\beta$, the minimum sample size required to control the Type II error at $\beta$ is defined as: \[ n^*_{\beta}(\Delta) = \min\left\{n: \mathbb{P}_{H_1(\Delta)}\left(\widehat{\theta}^{\mathrm{A2IPW}}_n \in R_n\right) \geq 1 - \beta\right\}. \]

In single-stage testing, we can compute the sample size $n^*_{\beta}$ by using the asymptotic distribution of $\widehat{\theta}^{\mathrm{A2IPW}}_T$. See Section (ref). Note that to compute $n^*_{\beta}(\Delta)$, we need to know the conditional variance $\sigma^2_0(a, x)$ to calculate $V$ in Theorem (ref). In practice, conjectured values or upper bounds of the conditional variance or $V$ can be used. It is important to note that as the conditional variance or $V$ increases, the required sample size also increases.

In sequential testing, the sample size corresponds to the stopping time when the algorithm stops after rejecting the null hypothesis. Letting $\tau$ denote the stopping time, we evaluate the expected value of $\tau$, which is also referred to as the sample complexity. In Theorem (ref), we show that the sequential test is essentially as powerful as a batch test with a sample size of $T$.

Implementation

Let $\alpha \in (0, 1)$ be the target Type-I error, and the experimenter aims to perform hypothesis testing without a Type-I error exceeding $\alpha$.

Single-stage testing. When our interest lies in single-stage testing (at the end of the experiment), we utilize the asymptotic normality of $\widehat{\theta}^{\mathrm{A2IPW}}_T$ in Theorem (ref): \[ \sqrt{T}\left(\widehat{\theta}^{\mathrm{A2IPW}}_T - \theta_0\right) \xrightarrow{\mathrm{d}} \mathcal{N}\big(0, V\big). \] In this case, we apply the (asymptotic) Student's $t$-test using the $t$-statistic $\frac{\widehat{\theta}^{\mathrm{A2IPW}}_T - \mu}{\sqrt{\widehat{V}/T}}$, where $\widehat{V}$ is a consistent estimator of $V$.

If the null hypothesis (i.e., $\theta_0=0$) is true, the $t$-statistic asymptotically follows the standard normal distribution. Based on these results, the test rejects the null hypothesis when \[ T\left|\widehat{\theta}^{\mathrm{A2IPW}}_T - \mu\right| > \sqrt{T\widehat{V}}z_{1-\alpha/2} := p_T, \] where $z_\alpha$ is the $\alpha$ quantile of the standard normal distribution. When the sample size $T$ is large, the Type I error is controlled as \[ \mathbb{P}_{H_0}\left(T\left|\widehat{\theta}^{\mathrm{A2IPW}}_T - \mu \right| > p_T\right) \leq \alpha. \]

Sequential testing. In sequential testing, we construct a confidence interval using a LIL-based concentration inequality, as shown in Theorem (ref). Based on this result, we define the confidence sequences $\{q_t\}_{t \in [T]}$ as \[ q_t \coloneqq 1.1\left(\log \left(\frac{1}{\alpha}\right) + \sqrt{2\sum^t_{i=1}z^2_i\left(\log\frac{\log \sum^t_{i=1}z^2_i}{\alpha}\right)}\right), \] where $z_t \coloneqq z_t(\mu) \coloneqq \Psi_t - \mu$.

Sample Size Computation

In this subsection, we compute the sample size needed to control the Type I error at $\alpha$ while achieving power $\beta$. In single-stage testing, we calculate the required minimum sample size $T$. In sequential testing, we compute the expected stopping time $\mathbb{E}[\tau]$, where $\tau$ is the stopping time when the null hypothesis is rejected.

Minimum Sample Size under the Optimal Experiment

First, we derive the required minimum sample size for single-stage testing. Theorem (ref) shows that $ \sqrt{T}\left(\widehat{\theta}_T - \theta_0\right) \xrightarrow{\mathrm{d}} \mathcal{N}(0, V)$. When the null hypothesis is true ($\theta_0 = \mu$), \[ \frac{\sqrt{T}\left(\widehat{\theta}^{\mathrm{A2IPW}}_T - \mu\right)}{\sqrt{V}} \xrightarrow{\mathrm{d}}_{H_0} \mathcal{N}(0, 1). \] Based on these results, with sufficient samples and knowledge of $\sigma^2_0$, we reject the null hypothesis when \[ \left| \sqrt{T}\left(\widehat{\theta}^{\mathrm{A2IPW}}_T - \mu\right) \right| > \sqrt{V} z_{1-\alpha/2}, \] where $z_{1-\alpha/2}$ is the $1 - \alpha/2$ quantile of the standard normal distribution. As explained in Section (ref), the Type I error is controlled at $\alpha$.

We now compute the smallest sample size $n^{\mathrm{OPT}*}_{\beta}(\Delta)$ required to achieve power $\beta$. The asymptotic power is given as

align*[align* omitted — 631 chars of source]

For $T \geq \frac{\sigma^2}{\Delta^2}\big(z_{1-\alpha/2} - z_{\beta}\big)^2$, the asymptotic power becomes at least $\beta$. Therefore, to achieve power $\beta$, the required sample size is: \[ n^{\mathrm{OPT}*}_{\beta}(\Delta) = \frac{\mathbb{E}\left[\frac{\sigma^2_0(1, X_t)}{\pi^*(1 \mid X_t)} + \frac{\sigma^2_0(0, X_t)}{\pi^*(0 \mid X_t)} + \left(f_0(1, X_t) - f_0(0, X_t) - \theta_0\right)^2\right]}{\Delta^2}\big(z_{1-\alpha/2} - z_{\beta}\big)^2. \]

Expected Sample Size in Sequential Testing

In this section, we calculate the upper bound of the expected stopping time $\tau$. In sequential testing using an LIL-based concentration inequality, we propose an algorithm that rejects the null hypothesis when \[ \left|t \widehat{\theta}^{\mathrm{A2IPW}}_t - t \mu\right| > 1.1\left(\log \left(\frac{1}{\alpha}\right) + \sqrt{2\sum^t_{i=1}z^2_i\left(\log\frac{\log \sum^t_{i=1}z^2_i}{\alpha}\right)}\right) = q_t. \] Let $\tau$ be the stopping time of the sequential test, i.e., $\tau = \min \left\{t: \left|t \widehat{\theta}^{\mathrm{A2IPW}}_t - t \mu\right| > q_t\right\}$. When $t = \tau$, the null hypothesis is rejected.

We show that as time progresses, the probability that the sequential test does not reject the hypothesis becomes small. We bound $\mathbb{P}_{H_1}(\tau > \widetilde{t})$ for sufficiently large $\widetilde{t}$ such that $\widetilde{t} \Delta \gg 1.1\left(\log \left(\frac{1}{\alpha}\right) + \sqrt{2C^2 \widetilde{t}\left(\log\frac{\log C^2 \widetilde{t}}{\alpha}\right)}\right)$. First, we consider the probability of $\tau \geq \widetilde{t}$ for a stopping time $\tau$. The proof is shown in Appendix (ref).

lemmaWhen the alternative hypothesis is true, $\tau > \widetilde{t}$ occurs with probability \begin{align*} &\mathbb{P}_{H_1}(\tau > \widetilde{t})=O\left( \exp\left(-\frac{\widetilde{t}\Delta^2}{8C^2}\right)\right). \end{align*}

With this lemma, we prove the following theorem. The proof is in Appendix (ref).

theorem[Expected sample size in sequential testing] When the alternative hypothesis is true, if \[n^{\mathrm{OPT}*}_{\beta}(\Delta)\Delta \gg 1.1\left(\log \left(\frac{1}{\alpha}\right) + \sqrt{2C^2 n^{\mathrm{OPT}*}_{\beta}(\Delta)\left(\log\frac{\log C^2 n^{\mathrm{OPT}*}_{\beta}(\Delta)}{\alpha}\right)}\right),\] then the expected sample size in sequential testing satisfies \begin{align*} \mathbb{E}_{H_1}[\tau]=\left(1+\frac{8C^2}{V\big(z_{1-\alpha/2}-z_{\beta}\big)^2}\mathbb{P}_{H_1}(\tau > n^{\mathrm{OPT}*}_{\beta}(\Delta))\right)n^{\mathrm{OPT}*}_{\beta}(\Delta). \end{align*}

This result implies that given $\alpha$ and $\beta$, $\mathbb{E}_{H_1}[\tau]$ is approximately equal to $n^{\mathrm{OPT}*}_{\beta}(\Delta)$, multiplied by a constant term independent of $\Delta$. As $\Delta$ approaches zero, both $\mathbb{E}_{H_1}[\tau]$ and $n^{\mathrm{OPT}*}_{\beta}(\Delta)$ approach infinity. From this result, we find that the expected stopping time $\mathbb{E}_{H_1}[\tau]$ in sequential testing grows proportionally to the sample size $n^{\mathrm{OPT}*}_{\beta}(\Delta)$ in single-stage testing (\(\mathbb{E}_{H_1}[\tau] = (1 + O(1))n^{\mathrm{OPT}*}_{\beta}(\Delta)\) as \(\Delta \to 0\)). Hence, for sufficiently small $\Delta$, we can consider that $\mathbb{E}_{H_1}[\tau]$ becomes close to $n^{\mathrm{OPT}*}_{\beta}(\Delta)$.

This result suggests that sequential testing has the potential to stop an experiment earlier than single-stage testing since the expected sample size is nearly identical to the (non-random) oracle sample size of single-stage testing, even though we do not know $n^{\mathrm{OPT}*}_{\beta}(\Delta)$ in advance of the experiment. That is, our sequential testing only uses the (unknown) minimum sample size in expectation.

Here, we emphasize that the oracle sample size $n^{\mathrm{OPT}*}_{\beta}(\Delta)$ is unknown because computing it requires the efficiency bound, which depends on the true expected conditional outcomes and the conditional variances. In single-stage testing, we cannot change the sample size during an experiment, as doing so is considered a violation of standard experimental design principles. Sequential testing, on the other hand, allows us to conduct a nearly optimal adaptive experiment without knowing $n^{\mathrm{OPT}*}_{\beta}(\Delta)$. Thus, sequential testing effectively reduces the sample size.

Note that the condition \[n^{\mathrm{OPT}*}_{\beta}(\Delta)\Delta \gg 1.1\left(\log \left(\frac{1}{\alpha}\right) + \sqrt{2C^2 n^{\mathrm{OPT}*}_{\beta}(\Delta)\left(\log\frac{\log C^2 n^{\mathrm{OPT}*}_{\beta}(\Delta)}{\alpha}\right)}\right)\] holds when $\beta$ is sufficiently close to 0. This theorem leads to the following corollary.

corollarySuppose that \[n^{\mathrm{OPT}*}_{\beta}(\Delta)\Delta \gg 1.1\left(\log \left(\frac{1}{\alpha}\right) + \sqrt{2C^2 n^{\mathrm{OPT}*}_{\beta}(\Delta)\left(\log\frac{\log C^2 n^{\mathrm{OPT}*}_{\beta}(\Delta)}{\alpha}\right)}\right)\] and $\pi_t = \pi^*$. Under $H_1$, for a sufficiently large sample size, the expected stopping time for the sequential test using $q_t$ is proportional to $n^{\mathrm{OPT}*}_{\beta}(\Delta)$.

Minimum Sample Size and Early Stopping

For a user-defined treatment assignment probability $\pi_t$, if $\pi_t(a\mid x)\xrightarrow{\mathrm{p}}\pi^*(a\mid x)$ holds for all $a, x$, the asymptotic variance is the same as $\widetilde{\sigma}^2$ from Theorem (ref). Therefore, when $\pi_t(a\mid x)\xrightarrow{\mathrm{p}}\pi^*(a\mid x)$ holds for all $a, x$, the minimum sample size required for hypothesis testing is also $n^{\mathrm{OPT}*}_{\beta}(\Delta)$. By applying the same method as in the previous section, we can verify that the expected stopping time for sequential testing under a user-defined treatment assignment probability $\pi_t$ using $q_t$ is proportional to $n^{\mathrm{OPT}*}_{\beta}(\Delta)$.

Summary

We introduced two approaches for hypothesis testing: single-stage testing and sequential testing. Single-stage testing employs a fixed, non-random sample size determined prior to the experiment, while sequential testing continues the experiment until a predefined stopping criterion is satisfied. Single-stage testing is justified by the asymptotic normality of the A2IPW estimator, which allows us to compute both the statistical power and the required sample size. Sequential testing, in contrast, is designed for finite-sample analysis and allows early stopping.

In contrast to single-stage testing with a fixed sample size, sequential testing has the potential to reduce sample size by terminating the experiment early. For example, if the null hypothesis assumes zero ATE but the actual ATE is significantly large, sequential testing may be able to reject the null and finish the experiment in an early round. On the other hand, in cases where the null hypothesis is not easily rejected, sequential testing may require larger sample sizes. Even in such cases, if the true ATE is sufficiently small and the sample size required for single-stage testing is large, the expected sample size for sequential testing is approximately equal to the fixed sample size used in single-stage testing. \color{black}

table[table omitted — 2,644 chars of source]
table[table omitted — 2,612 chars of source]
table*[table* omitted — 1,948 chars of source]

Simulation Studies

In this section, we evaluate the effectiveness of the proposed algorithm through experimental comparisons. The proposed method using the A2IPW estimator is compared against several alternative approaches, including the MA2IPW estimator, the IPW estimator, a randomized controlled trial (RCT) with a fixed treatment assignment probability of \( \pi_t(1\mid X_t, \mathcal{H}_{t-1}) = \pi_t(0\mid X_t, \mathcal{H}_{t-1}) = 0.5 \) for all \( t \), an oracle estimator \( \hat{\theta}^\mathrm{OPT}_T \) that operates under the optimal treatment-assignment probability, and a direct method (DM) estimator defined as $\frac{1}{T} \sum^T_{t=1} \left( \widehat{f}_t(1, X_t) - \widehat{f}_t(0, X_t) \right)$.

To estimate the treatment-assignment probability and expected outcomes, we consider two cases using different nonparametric estimators: the Nadaraya–Watson (NW) estimator and the \( K \)-nearest neighbor (K-nn) estimator. For the MA2IPW estimator, we set the parameter as \( \zeta = t^{-1/1.5} \).

In Appendix (ref), we show simulation studies in which we compare our method using the A2IPW and the ADR estimator with the stratification tree method proposed in Meehan2022.

Setting

We conduct simulation studies using synthetic and semi-synthetic datasets. In each dataset, we perform the following three types of hypothesis testing:

itemize• Single-stage testing using a $T$-test. • Sequential testing with Bonferroni (BF) correction. • Sequential testing based on an adaptive confidence sequence derived from the LIL-based concentration inequality.

For all settings, the null and alternative hypotheses are given by \[ \mathcal{H}_0: \theta_0 = 0, \quad \mathcal{H}_1: \theta_0 \neq 0. \] For the standard hypothesis testing, we construct confidence intervals using $T$-statistics derived from the asymptotic distribution in Theorem (ref). The sequential testing with BF correction is conducted at $t=150, 250, 350, 450$. For the LIL-based sequential testing, confidence intervals are constructed using $q_t$ as shown in Theorem (ref).

Simulation Studies with Synthetic Dataset

We first conduct experiments using synthetic datasets to evaluate the proposed method. At each round $t$, a covariate vector $X_t \in \mathbb{R}^5$ is generated as \[ X_t = (X_{t1}, X_{t2}, X_{t3}, X_{t4}, X_{t5})^\top, \quad X_{tk} \sim \mathcal{N}(0, 1) \text{ for } k=1,2,3,4,5. \] The potential outcome model is given by \[ Y_t(d) = \mu_d + \sum^5_{k=1} X_{tk} + e_{td}, \] where $\mu_d$ is a constant and $e_{td}$ follows a normal distribution with standard deviation $\sigma_d$. The expectation of the potential outcome is $\mathbb{E}[Y_t(d)] = \mu_d$.

We generate four datasets, each containing $500$ units, under different settings for $\mu_d$ and $\sigma_d$:

itemize• Dataset 1: $\mu_1=0.8$, $\mu_0=0.3$, $\sigma_1=0.8$, $\sigma_0=0.3$. • Dataset 2: $\mu_1=0.5$, $\mu_0=0.5$, $\sigma_1=0.8$, $\sigma_0=0.3$. • Dataset 3: $\mu_1=0.8$, $\mu_0=0.3$, $\sigma_1=0.6$, $\sigma_0=0.4$. • Dataset 4: $\mu_1=0.5$, $\mu_0=0.5$, $\sigma_1=0.6$, $\sigma_0=0.4$.

For each setting, we conduct $1000$ independent trials. The results are summarized in Tables (ref), (ref), and (ref). We report the mean squared error (MSE) between $\theta$ and $\hat{\theta}$, the standard deviation of the MSE (STD), and the rejection rates of hypothesis testing based on $T$-statistics at the $150$th (mid) and $300$th (final) rounds. Additionally, we present the stopping times for the LIL-based algorithm and the multiple testing with BF correction. If the null hypothesis is not rejected in sequential testing, the stopping time is set to $500$.

Across various datasets, the proposed algorithm achieved lower MSE compared to other methods. The DM estimator tends to reject the null hypothesis with small samples in Dataset 1 but exhibited high Type II error in Dataset 2.

Simulation Studies with Semi-Synthetic Data

We also evaluated the proposed algorithm using semi-synthetic datasets constructed from the Infant Health and Development Program (IHDP). The IHDP dataset consists of simulated outcomes and covariates based on a real study, following the simulation setting proposed by doi:10.1198/jcgs.2010.08162. The dataset contains $747$ units with $6$ continuous and $19$ binary covariates, with outcomes generated artificially.

doi:10.1198/jcgs.2010.08162 considers two response surfaces:

itemize• Response Surface A: \begin{align*} &Y_t(0) \sim \mathcal{N}(X^\top_{t}\bm{\beta}_A, 1),\\ &Y_t(1) \sim \mathcal{N}(X^\top_{t}\bm{\beta}_A+4, 1), \end{align*} where elements of $\bm{\beta}_A \in \mathbb{R}^{25}$ are randomly sampled from $\{0, 1, 2, 3, 4\}$ with probabilities $(0.5, 0.2, 0.15, 0.1, 0.05)$. • Response Surface B: \begin{align*} &Y_t(0) \sim \mathcal{N}(\exp((X_{t} + W)^\top\bm{\beta}_B), 1),\\ &Y_t(1) \sim \mathcal{N}(X^\top_{t}\bm{\beta}_B - q, 1), \end{align*} where $\bm{W}$ is an offset matrix with all elements equal to $0.5$, $q$ is a constant ensuring an average treatment effect of $4$, and elements of $\bm{\beta}_B$ are randomly sampled from $\{0, 0.1, 0.2, 0.3, 0.4\}$ with probabilities $(0.6, 0.1, 0.1, 0.1, 0.1)$.

For experiments, we randomly select $500$ units from the dataset. The results, summarized in Tables (ref) and (ref), include MSE, the standard deviation of MSE (STD), rejection rates at the $150$th and $300$th periods, and stopping times for the LIL-based and BF correction-based sequential testing. If the hypothesis is not rejected in sequential testing, the stopping time is set to $500$.

Results

The experimental results in Tables (ref)--(ref) (synthetic data) and Tables (ref)--(ref) (semi-synthetic IHDP data) reveal several notable patterns. In all datasets, the proposed adaptive algorithm using the A2IPW estimator, particularly with the Nadaraya--Watson kernel, consistently yields lower mean squared error than the baseline RCT and DM estimators. This performance gap becomes more pronounced at larger sample sizes, indicating that adaptively refining treatment-assignment probabilities based on accumulated data improves estimation accuracy.

Another observation concerns the DM estimator, which sometimes rejects the null hypothesis more readily when the true effect is clearly different from zero (as in Dataset 1). However, in scenarios where the true effect is close to zero (Dataset 2), it can fail to reject the null and thus exhibit higher Type II error. This pattern underscores that DM methods are sensitive to both sample size and the true effect magnitude and may be less robust when the treatment effect is marginal.

Standard RCT designs with constant assignment probabilities maintain unbiasedness but often show higher mean squared error relative to A2IPW. The adaptive nature of A2IPW allows it to focus allocations more efficiently, leading to more precise estimates of the treatment effect. The oracle (optimal) estimator, which knows the true assignment probabilities in advance, outperforms all other methods in terms of mean squared error and is included only to demonstrate the theoretical upper bound of performance.

The sequential testing procedures exhibit distinct behaviors. The LIL-based approach is typically conservative in practice and requires larger sample sizes before rejecting the null hypothesis, while the Bonferroni-based correction often stops earlier but can inflate Type I error. For instance, Tables (ref) and (ref) show cases where the Bonferroni-based method rejects the null more frequently, even when the underlying effect is subtle. Standard hypothesis testing (using a fixed sample size and T-statistics) avoids the complexity of sequential testing but does not allow the possibility of early termination.

Overall, the results suggest that the A2IPW-based adaptive design achieves lower estimation error and maintains favorable operating characteristics in both standard and sequential testing frameworks. Whether to use a sequential testing procedure depends on factors such as how rapidly decisions must be reached, the acceptable risk of false positives, and whether the total sample size can be determined in advance.

Conclusion

In this study, we designed an adaptive experimental framework to efficiently estimate the ATE and conduct hypothesis testing. We began by reviewing the semiparametric efficiency bound, which characterizes the fundamental limits of estimation efficiency as a function of the treatment-assignment probability. We then defined the efficient treatment-assignment probability as the minimizer of the semiparametric efficiency bound and leveraged this result to develop an optimal adaptive experimental design.

Our proposed method consists of two key phases: the treatment-assignment phase and the ATE-estimation phase. In the treatment-assignment phase, treatments are adaptively assigned based on an estimate of the efficient treatment-assignment probability. In the ATE estimation phase, we estimate the ATE using the proposed A2IPW estimator, which is constructed from the data collected in the treatment-assignment phase. We demonstrated that this estimator achieves asymptotic optimality by proving that its asymptotic variance matches the semiparametric efficiency bound. This optimality also ensures smaller sample sizes in hypothesis testing, improving the efficiency of experimental design.

In addition to establishing asymptotic optimality, we derived both asymptotic and non-asymptotic confidence intervals for the A2IPW estimator. The non-asymptotic bounds provide finite-sample guarantees, which are particularly useful in practical applications where sample sizes are limited. These confidence intervals enable rigorous inference while maintaining a tight dependence on the underlying data distribution.

Furthermore, we developed a hypothesis-testing framework tailored to our adaptive experimental design. We introduced two approaches: single-stage testing, which relies on a fixed sample size and asymptotic normality, and sequential testing, which dynamically determines sample size based on intermediate test results. We analyzed the theoretical properties of both approaches and highlighted scenarios where sequential testing can substantially reduce the required sample size while maintaining statistical rigor.

Our study contributes to the broader literature on adaptive experimental design by providing a theoretically grounded and practically implementable methodology for efficient ATE estimation and inference. Future research directions include extending our framework to accommodate more complex settings, such as network interference viviano2022experimentaldesignnetworkinterference, clustered experimental designs viviano2025causalclusteringdesigncluster, and heterogeneous treatment effects kato2024adaptivepolicylearning. Further exploration of optimality guarantees in finite-sample regimes and their connections to best-arm identification remains an important avenue for research Kasy2021,KOCK2023624. Additionally, incorporating reinforcement learning techniques into the treatment-assignment phase may enhance adaptability and extend the applicability of our approach to more dynamic experimental settings NathanUehara2019,adusumilli2024dynamicallyoptimaltreatmentallocation,sakaguchi2024policylearningoptimaldynamic.

In summary, this study provides a comprehensive methodological framework for designing and analyzing adaptive experiments, ensuring both statistical efficiency and practical applicability in treatment-effect estimation and hypothesis testing.

\onecolumn