EconBase
← Back to paper

Fixed-Horizon Self-Normalized Inference for Adaptive Experiments via Martingale AIPW/DML with Logged Propensities

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

88,301 characters · 0 sections · 30 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Fixed-Horizon Self-Normalized Inference for Adaptive Experiments via Martingale AIPW/DML with Logged Propensities

\begingroup \thispagestyle{plain}

{\color{TitleNavy}\rule{\textwidth}{1.15pt}}

{\color{TitleNavy}Fixed-Horizon Self-Normalized Inference for Adaptive Experiments via Martingale AIPW/DML \\ with Logged Propensities}

{ Gabriel Saco} {\normalsizeUniversidad del Pac\'ifico}

{\smallORCID: \href{https://orcid.org/0009-0009-8751-4154}{0009-0009-8751-4154}} {\smallReplication code: \href{https://github.com/gsaco/martingale-aipw-dml}{\nolinkurl{https://github.com/gsaco/martingale-aipw-dml}}}

{\color{TitleGold}\rule{0.72\textwidth}{0.9pt}} \endgroup

center[center omitted — 1,473 chars of source]

\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Introduction}

Adaptive randomized experiments---including response-adaptive clinical trials, contextual bandits, and large-scale platform experimentation systems---update assignment probabilities as data accrue to balance learning and deployment kasySautmann2021. In practice, however, many platforms still require conventional end-of-study reporting for classical causal estimands such as the superpopulation average treatment effect (ATE) computed once at a prespecified horizon. Hereafter, we denote by $\pi_t$ the executed assignment probability used to randomize unit $t$ after applying any platform guardrails, as recorded in the experiment log. In this paper, we study fixed-horizon Wald inference for the standard logged-propensity AIPW/DML estimator, the sample average of doubly robust pseudo-outcomes scored using these logged propensities. Under adaptive assignment, the propensity process $\{\pi_t\}$ is itself data-dependent, so the predictable quadratic variation of the AIPW/DML score increments can remain replication-random and need not converge to a single deterministic long-run variance target.

Some end-of-study Wald arguments for AIPW/A2IPW (and related DML estimators) under adaptivity proceed via Slutsky steps that rely on a deterministic variance target for the predictable quadratic variation, often enforced through stabilization-type or design-stability conditions on assignment probabilities and/or average conditional variances hadad2021confidence,zhan2021off,katoIshiharaHondaNarita2020,cook2024semiparametric,LiOwen2024,sengupta2025designstability. On modern platforms, however, the policy can keep reacting to noisy intermediate estimates. Clipping and guardrails can activate intermittently and batch updates can induce regime switches. When this happens, a Wald statistic normalized using a single deterministic variance target can be systematically miscalibrated conditional on the realized variance regime, even if marginal coverage appears close to nominal.

We treat the adaptive assignment policy as given but assume that the platform logs the executed propensity used to randomize each unit and that nuisance regressions used for AIPW/DML scoring are fit predictably using only past data. Related work emphasizes that careful use of the logging policy and past-only fitting is central for post-adaptive inference bibaut2021post,katoYasuiMcAlinn2021,cook2024semiparametric. Under these auditable conditions, the centered score increments form an exact martingale difference sequence, and we obtain fixed-horizon Wald inference by studentizing with realized quadratic variation along the realized propensity path. This yields asymptotic $\mathcal{N}(0,1)$ calibration without requiring the predictable quadratic variation to converge to a deterministic long-run variance target.\footnote{We do not claim anytime-valid or optional-stopping guarantees.}

\noindentContributions.

itemize• Auditable martingale scoring. We formalize a logging/predictability contract, logged executed propensities and predictable nuisance fitting, under which centered AIPW/DML score increments form an exact martingale difference sequence (Lemma (ref)). • Fixed-horizon self-normalized Wald inference. We prove that the usual Studentized statistic, with variance estimated by realized quadratic variation, is asymptotically $\mathcal{N}(0,1)$ at a prespecified horizon even when no deterministic long-run variance limit exists (Theorem (ref)). • Feasible studentization. We show that the standard plug-in studentizer used in practice consistently estimates realized quadratic variation, so the feasible Wald interval inherits the same fixed-horizon validity (Proposition (ref)). • Oracle benchmarking and nuisance-learning effects. We provide a conditional second-moment decomposition yielding an oracle precision benchmark and isolate a nonnegative augmentation term capturing variance inflation from nuisance error (Proposition (ref)). Under weighted $L^2$ convergence, the feasible statistic is asymptotically oracle-equivalent (Theorem (ref)).

The subsequent sections are structured as follows. Section (ref) reviews related work. Section (ref) states the model and assumptions. Section (ref) presents the estimator and auditable implementation details. Section (ref) develops the main theoretical results and Section (ref) reports simulations. The appendices collect supporting limit-theory background, additional results and proofs, and an operational logging protocol.

\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Related Work}

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Inference after adaptive data collection} A growing literature studies inference with adaptively collected data, where observations are generated under an evolving information set and classical i.i.d.\ arguments do not directly apply. Early econometric work by hahnHiranoKarlan2011 highlighted how propensity information can be leveraged for inference in sequential designs. More recent general frameworks derive asymptotic representations for sequential decisions and adaptive experiments under broad conditions hiranoPorter2023asymptotics. In the contextual-bandit and adaptive-experiment literature, fixed-horizon inference has also been developed via batched OLS/batchwise studentization arguments zhang2020large. Our focus is narrower but operationally central for experimentation platforms: fixed-horizon ATE reporting with logged executed propensities and predictable AIPW/DML scoring. In particular, we target settings where the predictable quadratic variation remains random across replications, so that deterministic-variance normalizations can be conditionally miscalibrated. General background on response-adaptive randomization and bandit-style designs can be found in rosenberger2015randomization and villar2015bandit.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Stabilization, adaptive weighting and batching} In adaptive experimentation and off-policy evaluation, evolving propensities can create heavy tails and regime-dependent uncertainty, motivating variance-control strategies. In policy evaluation, hadad2021confidence and zhan2021off develop adaptive weighting schemes for augmented IPW/DR scores to obtain asymptotically normal $t$-statistics. For post-contextual-bandit inference, bibaut2021post proposes stabilized doubly robust constructions that estimate conditional scale components using only past data. Another route restricts the data-collection design---for example through batching---to recover classical CLTs for bandit estimators zhang2020large. In contrast, we keep the standard logged-propensity AIPW/DML estimator unchanged and do not require batching or adaptive reweighting. Instead, we rely on a martingale score representation and studentization by realized quadratic variation.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Design stability and deterministic variance-limit CLTs} A further line of work derives fixed-horizon CLTs for IPW/AIPW-type ATE estimators under explicit design stability conditions ensuring that inverse-propensity averages and/or average conditional variances converge to nonrandom limits, yielding conventional deterministic asymptotic variances for Wald reporting sengupta2025designstability. Related perspectives arise when one engineers assignment rules (or batchwise designs) to target efficiency or precision for A2IPW/DML-style estimators katoIshiharaHondaNarita2020,LiOwen2024,cook2024semiparametric. More generally, the adaptive-experiment literature emphasizes subtleties around what “the logging policy” means operationally when platforms implement clipping, guardrails, or algorithmic randomness katoYasuiMcAlinn2021. Our contribution is complementary: we assume the executed propensity actually used to randomize each unit is logged and we avoid requiring convergence of the predictable quadratic variation to a deterministic limit by using realized quadratic variation as the normalizer.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Self-normalized martingale theory for quadratic-variation studentization} Our fixed-horizon Wald statistic is a self-normalized martingale functional. Classic references for martingale CLTs and self-normalized processes include hallHeyde1980 and delapena2009self, as well as the survey shao2013survey. Modern probability theory provides refined asymptotic and nonasymptotic control for self-normalized martingales, including Berry--Esseen bounds fanShao2017, Cram{\'e}r-type moderate deviations fanGramaLiuShao2019, and concentration inequalities bercuTouati2019. We apply this theory to AIPW/DML score increments under adaptive assignment, using realized quadratic variation to obtain a fixed-horizon $\mathcal{N}(0,1)$ approximation without a deterministic variance limit.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Anytime-valid and time-uniform alternatives} The results in this paper are prespecified-horizon and do not provide optional-stopping guarantees. When inference must remain valid under continuous monitoring or data-dependent stopping, time-uniform methods based on test supermartingales and confidence sequences are appropriate howard2021confidence,waudby2024timeuniform. Related time-uniform tools also appear alongside fixed-time inference in adaptive-experiment work cook2024semiparametric,katoIshiharaHondaNarita2020,waudbySmithRamdas2024betting. We focus instead on conventional fixed-horizon reporting with realized quadratic-variation studentization (see Remark (ref) for additional discussion). \@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Model and Assumptions}

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Propensities and Logged Assignment Rule}

Consider a stream of experimental units indexed by $t=1,\dots,n$. For each unit, covariates $X_t$ are observed before assignment. The platform computes an assignment probability $\pi_t\in(0,1)$. Then, randomizes $A_t\in\{0,1\}$ using $\pi_t$. Finally, observes an outcome $Y_t$.

We observe and store $Z_t:=(X_t,A_t,Y_t,\pi_t)$ in time order. Let

equation[equation omitted — 113 chars of source]

be the data filtration. Let

equation[equation omitted — 81 chars of source]

be the $\sigma$-field immediately before randomization at time $t$, after the platform has computed the executed assignment probability $\pi_t$. The policy may be arbitrarily adaptive or randomized. The fixed-horizon validity results below require only that the realized executed propensity used to randomize $A_t$ is recorded in the log (Assumption (ref)).

definition[Adaptive logged executed propensity] At time $t$, after observing $(\mathcal{F}_{t-1},X_t)$ and any algorithmic randomness used to form the propensity, the platform records the executed assignment probability $\pi_t\in(0,1)$ and draws treatment according to \[ A_t \mid \mathcal{G}_t \sim \mathrm{Bernoulli}(\pi_t), \] where $\mathcal{G}_t$ is the pre-treatment $\sigma$-field in (ref). If $\pi_t$ is deterministic given $(\mathcal{F}_{t-1},X_t)$, then $\pi_t=\mathbb{P}(A_t=1\mid \mathcal{F}_{t-1},X_t)$. If $\pi_t$ is produced using additional exogenous randomness, then $\mathbb{P}(A_t=1\mid \mathcal{F}_{t-1},X_t)=\mathbb{E}[\pi_t\mid \mathcal{F}_{t-1},X_t]$ while still $\mathbb{P}(A_t=1\mid \mathcal{G}_t)=\pi_t$. We treat any such platform-side randomness as exogenous, i.e., independent of $(Y_t(0),Y_t(1))$ conditional on $(\mathcal{F}_{t-1},X_t)$.
assumption[Logging integrity] For each $t$, the logged propensity $\pi_t$ is $\mathcal{G}_t$-measurable, takes values in $(0,1)$, and equals the probability passed to the randomization device that generated $A_t$, i.e., \[ \mathbb{P}(A_t=1\mid \mathcal{G}_t)=\pi_t \quad \text{a.s.} \] Equivalently, $A_t\mid \mathcal{G}_t\sim \mathrm{Bernoulli}(\pi_t)$ with the logged $\pi_t$.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Audit diagnostics and data contract}

remark[Auditing logged propensities] Assumption (ref) requires that, conditional on the platform history $\mathcal{G}_t$ (including the logged probability $\pi_t$), the treatment assignment satisfies $A_t\mid \mathcal{G}_t \sim \mathrm{Bernoulli}(\pi_t)$. This is a design-stage property. It holds only if the log records the probability that was actually passed to the randomization device. A basic diagnostic is calibration of $A_t$ against $\pi_t$. For any bin $B\subset(0,1)$ with many observations, let \[ N_B:=\sum_{t=1}^n \mathbbm{1}\{\pi_t\in B\},\qquad \overline A_B:=\frac{1}{N_B}\sum_{t:\pi_t\in B} A_t,\qquad \overline \pi_B:=\frac{1}{N_B}\sum_{t:\pi_t\in B} \pi_t. \] In practice, restrict attention to bins with $N_B\ge N_{\min}$ for a user-chosen $N_{\min}\ge 1$ (or merge empty/small bins). Under correct logging, $\overline A_B$ should be close to $\overline \pi_B$ up to martingale sampling variability; under adaptivity, dependence can matter, so formal testing can be based on martingale concentration/self-normalized methods for the MDS $(A_t-\pi_t)$. These checks can detect severe mis-logging or implementation bugs, but they cannot certify correct logging or validate the causal assumptions in Assumptions (ref)--(ref).
remark[Timing and measurability] The experiment unfolds in the order (1) observe $X_t$; (2) the platform computes and logs an executed propensity $\pi_t$ as a function of $(\mathcal{F}_{t-1},X_t)$; (3) treatment is drawn according to $A_t\mid \mathcal{G}_t\sim\mathrm{Bernoulli}(\pi_t)$; and (4) the outcome $Y_t$ is revealed and logged. In particular, $\pi_t$ is $\mathcal{G}_t$-measurable, while $(A_t,Y_t)$ are $\mathcal{F}_t$-measurable. This timing underlies the predictability requirement in Section (ref) and the data-contract protocol in Appendix (ref).
table[table omitted — 872 chars of source]
table[table omitted — 1,441 chars of source]

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Potential Outcomes and Target Parameter}

We adopt the potential outcomes framework with the usual no interference and consistency conventions.

assumption[Consistency and no interference; SUTVA] For each unit $t$, there are well-defined potential outcomes $(Y_t(0),Y_t(1))$ that depend only on the assignment of unit $t$ itself rubin1980comment. The observed outcome satisfies \[ Y_t = Y_t(A_t) = A_t Y_t(1) + (1-A_t)Y_t(0)\qquad\text{a.s.} \]

The target parameter is the average treatment effect (ATE) in the superpopulation:

equation[equation omitted — 76 chars of source]

where $(X,Y(0),Y(1))$ denotes a generic draw from the superpopulation described in Assumption (ref). Assumption (ref) is standard. It rules out spillovers or network effects, which require separate methods.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Assumptions}

We maintain the following assumptions throughout.

assumption[Superpopulation arrivals; no selection] The sequence $\{(X_t,Y_t(0),Y_t(1))\}_{t=1}^n$ is independent across $t$ and identically distributed with generic draw $(X,Y(0),Y(1))$. In particular, the adaptive policy affects only treatment assignment, not which units/covariates arrive.

We write all expectations with respect to this superpopulation draw, so $\theta_0=\mathbb{E}[Y(1)-Y(0)]$ and $\tau(x):=\mathbb{E}[Y(1)-Y(0)\mid X=x]$ satisfy $\theta_0=\mathbb{E}[\tau(X)]$. As noted in Remark (ref), the proofs only require the conditional mean-stationarity condition $\mathbb{E}[\tau(X_t)\mid \mathcal{F}_{t-1}]=\theta_0$.

assumption[Sequential randomization/no anticipation] $A_t \mathrel{\perp\!\!\!\perp} (Y_t(0),Y_t(1))\mid \mathcal{G}_t$, where $\mathcal{G}_t$ is the pre-treatment $\sigma$-field in (ref). Equivalently, conditional on the pre-treatment information (including the realized executed propensity $\pi_t$), the treatment draw is independent of the potential outcomes.
assumption[Executed overlap on scored units] There exists $\varepsilon>0$ such that for every time index whose score is included in the final estimator (i.e., $t\in\mathcal{T}$ in Section (ref)), \begin{equation} \varepsilon \le \pi_t \le 1-\varepsilon \qquad a.s. \end{equation} A sufficient operational mechanism is to enforce clipping of the executed propensity before randomization on all scored units.
assumption[Moment conditions] For each $a\in\{0,1\}$, $\mathbb{E}[|Y(a)|^4]<\infty$.
assumption[Nuisance stability] There exists a finite constant $C<\infty$ such that for each $a\in\{0,1\}$, \[ \sup_{n\ge 1}\sup_{t\in\mathcal{T}} \mathbb{E}[|\widehat m_{t-1,a}(X_t)|^4]\le C. \] In unbounded-outcome settings, a sufficient alternative is to truncate/clamp $\widehat m_{t-1,a}(X_t)$ at a large threshold; see Remark (ref).
remark[How to enforce Assumption (ref)] If outcomes are known to be bounded, say $Y_t\in[-B,B]$, we can enforce Assumption (ref) by clipping $\widehat m_{t-1,a}(x)$ to $[-B,B]$ (or slightly wider) for each $a$. For unbounded outcomes, one may instead truncate at a large threshold or use robust regression; this is a technical device to control moments and does not change the target estimand.
remark[Weaker moment conditions are possible but are not needed for this paper] Self-normalized martingale CLTs only require finite $(2+\delta)$ moments of the score increments. We impose fourth moments to keep the quadratic-variation equivalence proof fully transparent.
remark[Design-stage, analysis-stage and auditable conditions] Our validity results combine (i) substantive causal/model assumptions and (ii) pipeline conditions that can be checked from logs and stored artifacts. \begin{enumerate} • Substantive assumptions. SUTVA/no interference (Assumption (ref)) and sequential randomization with a stable target estimand (Assumption (ref)), together with i.i.d.\ arrivals/no selection (Assumption (ref)) or an explicit mean-stationarity alternative (Remark (ref)). • Auditable design-time requirements. Logged executed propensities (Assumption (ref)) and executed overlap on the scored units (Assumption (ref)), which must be enforced at assignment time if needed (e.g., by clipping $\pi_t$). • Auditable analysis-time requirements. A prespecified (deterministic) scored set $\mathcal{T}$ (Assumption (ref)) and predictable nuisance construction (Assumption (ref)), which rule out “peeking” when fitting the regressions used to score each unit. • Regularity conditions. Moment and stability bounds (Assumptions (ref) and (ref)) and variance growth (Assumption (ref)) needed to apply martingale limit theory. \end{enumerate} Finally, all confidence intervals are fixed-horizon: they should be reported only at the prespecified sample size and are generally invalid under optional stopping (Remark (ref) and Appendix (ref)).

\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Predictable AIPW/DML Estimation}

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{The Doubly Robust Score}

Let $m_a^\star(x):=\mathbb{E}[Y(a)\mid X=x]$ and write $m^\star:=(m_0^\star,m_1^\star)$. For any candidate $m=(m_0,m_1)$ and any $\pi\in(0,1)$, define the usual augmented inverse-propensity weighted (AIPW) / doubly robust (DR) score robinsRotnitzkyZhao1994,bangRobins2005 (pseudo-outcome)

equation[equation omitted — 151 chars of source]

When evaluating at the logged executed propensity, we write $\phi_t(m):=\phi_t(m,\pi_t)$ for brevity. The quantity $\phi_t(m,\pi_t)$ is observable given $(X_t,A_t,Y_t,\pi_t)$ and a supplied regression pair $m$.

For the oracle regression $m^\star$, one can decompose

equation[equation omitted — 169 chars of source]

The centered increment $\phi_t(m^\star)-\theta_0$ has mean zero and is the object governed by the martingale CLT.

remark[Terminology] The estimating function $\phi_t(m,\pi_t)$ in (ref) has the classical doubly robust/AIPW form. When evaluated at the true regression functions $m^\star$ (and the executed propensities), the centered score $\phi_t(m^\star,\pi_t)-\theta_0$ coincides with the efficient influence function for the ATE in the corresponding i.i.d.\ model with known propensity (see, e.g., hahn1998role,tsiatis2006semiparametric). For general $m$, it is simply an estimating function with the same doubly robust form. In off-policy evaluation and contextual-bandit settings, essentially the same form appears as a doubly robust score for value estimation dudikLangfordLi2011,dudikErhanLangfordLi2014.
remark[Multi-arm extensions] The results extend directly to $K>2$ arms by replacing the scalar propensity with a probability vector and using the corresponding multivariate AIPW score. For simplicity, we present the binary case.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Forward Cross-Fitting}

Standard cross-fitting partitions data into folds and estimates nuisances on held-out folds. Under adaptivity, we must respect time: nuisance estimates used at time $t$ must be constructed from data strictly before $t$.

definition[Forward cross-fitting] Partition indices $\{1, \ldots, n\}$ into $K$ contiguous blocks $I_1, \ldots, I_K$. The partition is deterministic and fixed prior to data collection. For each block $k \ge 2$: \begin{enumerate}[leftmargin=2em, label=\arabic*.] • Estimate nuisance functions $(\widehat m^{(-k)}_0, \widehat m^{(-k)}_1)$ using only data from blocks $I_1, \ldots, I_{k-1}$. • For each $t \in I_k$, evaluate the score using these estimates and the logged propensity $\pi_t$. \end{enumerate} For $t \in I_1$, use a pilot estimate or exclude $I_1$ from inference.
remark[Common pitfall] A common pitfall is to use i.i.d.-style sample splitting or cross-fitting patterns that inadvertently allow information from unit $t$ (or future units) to enter the nuisance regression used to score unit $t$. In adaptive experiments, such “leakage” violates Assumption (ref) and can break the martingale difference property in Lemma (ref). As a result, the Studentized Wald statistic built from leaky scores generally lacks the fixed-horizon guarantee of Theorem (ref): it may still work in some finite-sample designs, but its validity is no longer ensured by the martingale argument. Appendix (ref) gives an auditable implementation pattern (forward cross-fitting) that enforces predictability.
assumption[Predictable nuisance construction] For each scored time $t\in\mathcal{T}$ and treatment arm $a\in\{0,1\}$, the regression function $\widehat m_{t-1,a}$ used in the score for unit $t$ is $\mathcal{F}_{t-1}$-measurable. Equivalently, conditional on the past $\mathcal{F}_{t-1}$, the function $\widehat m_{t-1,a}$ does not depend on $(A_t,Y_t)$ or on any future data.

Assumption (ref) is the formal predictability requirement (“no-peeking”, i.e., no use of contemporaneous or future outcomes) that makes the scored pseudo-outcomes a martingale difference sequence, and predictable fits are not based on non-predictable (“leaky”) scores (see Lemma (ref) below). The measurability here refers to the fitted function object $\widehat m_{t-1,a}$; it may then be evaluated at the current covariate $X_t$ to form $\widehat m_{t-1,a}(X_t)$. It rules out i.i.d.-style cross-fitting schemes that, even indirectly, use contemporaneous or future outcomes when constructing the nuisance for time $t$. A simple way to enforce predictability is forward cross-fitting (Definition (ref)): partition the time axis into blocks $I_1,\dots,I_K$, fit each nuisance model once per block using only data from $\cup_{j<k} I_j$, and hold the fitted model fixed while scoring units in $I_k$. To make this requirement auditable, an implementation should persist (i) the training indices used for each fit, (ii) fold assignments if internal cross-validation is used, and (iii) all learner randomness.

lemma[A sufficient operational condition for predictability] Under forward cross-fitting (Definition (ref)), if for each block $I_k$ the analyst fits $(\widehat m_0^{(-k)},\widehat m_1^{(-k)})$ using only data from blocks $I_1,\dots,I_{k-1}$ and then reuses these fitted objects unchanged for all $t\in I_k$, then Assumption (ref) holds.
proofFor $t\in I_k$, the fitted objects depend only on $\sigma(Z_s: s\in I_1\cup\cdots\cup I_{k-1})\subseteq \mathcal{F}_{t-1}$ and are therefore $\mathcal{F}_{t-1}$-measurable.
remark[Connection to A2IPW terminology] Our $\widehat{\phi}_t$ is the same AIPW/DR pseudo-outcome used in the adaptive-experiment literature under the name A2IPW (adaptive AIPW). At time $t$ the score is computed using outcome regressions fitted on past data only, together with the logged executed propensity used to randomize $A_t$ katoIshiharaHondaNarita2020,cook2024semiparametric. The present paper emphasizes that this predictability requirement is not merely a convenience: it is the condition that yields an exact martingale difference sequence and enables fixed-horizon Studentized inference without any stabilization of propensities or conditional variances.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{The Estimator}

Let $\mathcal{T}\subseteq\{1,\dots,n\}$ denote the set of indices for which the nuisance used at time $t$ is predictable (i.e., $\widehat m_{t-1}$ is $\mathcal{F}_{t-1}$-measurable). Under forward cross-fitting with a burn-in block $I_1$, one typically takes $\mathcal{T}:=\{1,\dots,n\}\setminus I_1$. Let $n_{\mathrm{eff}}:=\sum_{t=1}^n \mathbbm{1}\{t\in\mathcal{T}\}=|\mathcal{T}|$. Throughout the asymptotic theory we treat the scored index set $\mathcal{T}$ as deterministic, as is the case under the forward-block construction in Definition (ref).

assumption[Deterministic scored set] The scored index set $\mathcal{T}\subseteq\{1,\dots,n\}$ is fixed prior to data collection (equivalently, $T_t:=\mathbbm{1}\{t\in\mathcal{T}\}$ is nonrandom for each $t$).
remark[Predictable scored sets] If $T_t:=\mathbbm{1}\{t\in\mathcal{T}\}$ is allowed to be $\mathcal{F}_{t-1}$-measurable, the same proof strategy applies to $\tilde\xi_t:=T_t(\widehat\phi_t-\theta_0)$. In this case, interpret $n_{\mathrm{eff}}:=\sum_{t=1}^n T_t$ and define $Q_{\mathcal{T}}$ and $V_{\mathcal{T}}^2$ using $\tilde\xi_t$. We omit further details of the predictable-$\mathcal{T}$ extension to keep the note focused.

\paragraph{Estimator and studentizing factor.} The cross-fitted pseudo-outcome is

equation[equation omitted — 180 chars of source]

where $\widehat m_{t-1}$ is the predictable nuisance estimate.

equation[equation omitted — 182 chars of source]
remark[The studentizer as realized quadratic variation] Let $n_{\mathrm{eff}}:=|\mathcal{T}|$ and recall $\widehat V$ in (ref) and $Q_{\mathcal{T}}$ in (ref). A direct calculation gives \begin{equation} (n_{\mathrm{eff}}-1)\widehat V = Q_{\mathcal{T}} - n_{\mathrm{eff}}(\widehat\theta-\theta_0)^2. \end{equation} Thus $\widehat V$ is essentially the realized quadratic variation $Q_{\mathcal{T}}$ up to the negligible centering term $n_{\mathrm{eff}}(\widehat\theta-\theta_0)^2/(n_{\mathrm{eff}}-1)$. Proposition (ref) records a minimal condition under which $(n_{\mathrm{eff}}-1)\widehat V/Q_{\mathcal{T}}\to_p 1$.
proposition[Feasibility of the studentizer] Let $\xi_t:=\widehat\phi_t-\theta_0$ and $S_{\mathcal{T}}:=\sum_{t\in\mathcal{T}}\xi_t$, so that $\widehat\theta-\theta_0=S_{\mathcal{T}}/n_{\mathrm{eff}}$ and $Q_{\mathcal{T}}=\sum_{t\in\mathcal{T}}\xi_t^2$. If $n_{\mathrm{eff}}\to\infty$ and $S_{\mathcal{T}}/\sqrt{Q_{\mathcal{T}}}=O_p(1)$, then \[ \frac{(n_{\mathrm{eff}}-1)\widehat V}{Q_{\mathcal{T}}}\to_p 1. \] In particular, the condition $S_{\mathcal{T}}/\sqrt{Q_{\mathcal{T}}}=O_p(1)$ holds whenever $S_{\mathcal{T}}/\sqrt{Q_{\mathcal{T}}}\Rightarrow \mathcal{N}(0,1)$.
proofBy (ref), \[ \frac{(n_{\mathrm{eff}}-1)\widehat V}{Q_{\mathcal{T}}} = 1-\frac{1}{n_{\mathrm{eff}}}\Big(\frac{S_{\mathcal{T}}}{\sqrt{Q_{\mathcal{T}}}}\Big)^2, \] which converges to $1$ in probability under the stated conditions.

\paragraph{Score sum notation.} Define the centered increments $\xi_t:=\widehat\phi_t-\theta_0$ and the score sum $S_{\mathcal{T}}:=\sum_{t\in\mathcal{T}}\xi_t=n_{\mathrm{eff}}(\widehat\theta-\theta_0)$.

equation[equation omitted — 92 chars of source]
equation[equation omitted — 122 chars of source]
remark[Finite-$n$ reporting convention] For small $n_{\mathrm{eff}}$, it is common as a pragmatic finite-sample convention to also report a $t$-critical variant using $t_{0.975, n_{\mathrm{eff}}-1}$, namely $CI_t:=\widehat\theta\pm t_{0.975, n_{\mathrm{eff}}-1}\sqrt{\widehat V/n_{\mathrm{eff}}}$, where $\widehat V$ is the sample variance in (ref). This is a reporting convention; we do not claim finite-sample Student-$t$ validity, and it is not covered by Theorem (ref).

\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Main Theoretical Results}

remark[Triangular-array and uniformity convention] All objects may depend on the horizon $n$ (e.g.\ forward blocks and nuisance fits). We suppress the index $n$ and assume overlap and moment constants are uniform over $n$.

This section isolates the logic behind predictable AIPW/DML in adaptive experiments. First, under logged executed propensities and predictable nuisance fits, the centered score is an exact martingale difference, yielding finite-sample conditional unbiasedness (Proposition (ref) and Lemma (ref)). Second, fixed-horizon inference follows from a self-normalized martingale CLT applied to the martingale sum, with feasibility provided by a sample-variance approximation (Theorem (ref)). Table (ref) in Appendix (ref) maps each assumption to each statement, and Appendix (ref) states the martingale limit theorem invoked in the proofs.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{The Martingale Structure}

The central insight of our analysis is that forward cross-fitting induces a martingale difference structure. We first establish an identification result for the pseudo-outcome.

proposition[Identification, including predictable random nuisances] Let $m_{t-1}=(m_{t-1,0},m_{t-1,1})$ be any (possibly random) pair of regression functions that is $\mathcal{F}_{t-1}$-measurable. Under Assumptions (ref), (ref), and (ref), for each $t$, \begin{equation} \mathbb{E}\!\big[\phi_t(m_{t-1})\mid \mathcal{G}_t\big]=\tau(X_t). \end{equation} Consequently, $\mathbb{E}[\phi_t(m_{t-1})]=\theta_0$.
proofFix $t$ and condition on $\mathcal{G}_t=\sigma(\mathcal{F}_{t-1},X_t,\pi_t)$. Since $m_{t-1}$ is $\mathcal{F}_{t-1}$-measurable and $\mathcal{F}_{t-1}\subseteq \mathcal{G}_t$, the function $m_{t-1}$ is fixed under this conditioning. By sequential randomization (Assumption (ref)), \[ \mathbb{E}[Y_t \mid \mathcal{G}_t, A_t=1]=\mathbb{E}[Y_t(1)\mid \mathcal{G}_t],\qquad \mathbb{E}[Y_t \mid \mathcal{G}_t, A_t=0]=\mathbb{E}[Y_t(0)\mid \mathcal{G}_t]. \] By i.i.d.\ arrivals (Assumption (ref)) and the exogeneity condition in Definition (ref), $(Y_t(0),Y_t(1))\mathrel{\perp\!\!\!\perp} (\mathcal{F}_{t-1},\pi_t)\mid X_t$ and thus $\mathbb{E}[Y_t(a)\mid X_t,\mathcal{F}_{t-1},\pi_t]=\mathbb{E}[Y_t(a)\mid X_t]=m_a^\star(X_t)$ for $a\in\{0,1\}$. Using $\mathbb{E}[A_t\mid \mathcal{G}_t]=\pi_t$ and the tower property, \begin{align*} \mathbb{E}\!\left[\frac{A_t(Y_t-m_{t-1,1}(X_t))}{\pi_t}\,\Big|\,\mathcal{G}_t\right] &= \mathbb{E}\!\left[\frac{A_t}{\pi_t}\,\mathbb{E}[Y_t-m_{t-1,1}(X_t)\mid \mathcal{G}_t,A_t]\,\Big|\,\mathcal{G}_t\right]\\ &= \mathbb{E}\!\left[\frac{A_t}{\pi_t}\,(m_1^\star(X_t)-m_{t-1,1}(X_t))\,\Big|\,\mathcal{G}_t\right]\\ &= m_1^\star(X_t)-m_{t-1,1}(X_t), \end{align*} and similarly \[ \mathbb{E}\!\left[\frac{(1-A_t)(Y_t-m_{t-1,0}(X_t))}{1-\pi_t}\,\Big|\,\mathcal{G}_t\right] = m_0^\star(X_t)-m_{t-1,0}(X_t). \] Plugging into the definition of $\phi_t(m_{t-1})$ in (ref) yields \begin{align*} \mathbb{E}[\phi_t(m_{t-1})\mid \mathcal{G}_t] &= (m_{t-1,1}(X_t)-m_{t-1,0}(X_t)) +(m_1^\star(X_t)-m_{t-1,1}(X_t))\\ &\quad -(m_0^\star(X_t)-m_{t-1,0}(X_t))\\ &= m_1^\star(X_t)-m_0^\star(X_t) =\tau(X_t), \end{align*} which proves (ref). Taking unconditional expectations gives $\mathbb{E}[\phi_t(m_{t-1})]=\mathbb{E}[\tau(X_t)]=\theta_0$.

The key consequence is that the pseudo-outcome, when evaluated with any predictable nuisance estimates, remains conditionally unbiased.

lemma[Martingale difference structure] Suppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Let $\{\widehat m_{t-1}\}$ satisfy Assumption (ref). For each scored index $t\in\mathcal{T}$, define $\widehat\phi_t:=\phi_t(\widehat m_{t-1})$ and $\xi_t:=\widehat\phi_t-\theta_0$. Then \begin{equation} \mathbb{E}[\widehat\phi_t \mid \mathcal{F}_{t-1}] = \theta_0, \end{equation} and $\{\xi_t,\mathcal{F}_t\}$ is a martingale difference sequence over the scored indices.
proof(A finite first moment suffices here; this follows from Assumption (ref). Higher-moment bounds used later are provided by Lemma (ref).) Using iterated expectations and Proposition (ref), \[ \mathbb{E}[\phi_t(\widehat m_{t-1})\mid \mathcal{F}_{t-1}] = \mathbb{E}\!\big[\mathbb{E}[\phi_t(\widehat m_{t-1})\mid \mathcal{G}_t]\mid \mathcal{F}_{t-1}\big] = \mathbb{E}[\tau(X_t)\mid \mathcal{F}_{t-1}]. \] By Assumption (ref), $X_t$ is independent of $\mathcal{F}_{t-1}$, so $\mathbb{E}[\tau(X_t) \mid \mathcal{F}_{t-1}] = \mathbb{E}[\tau(X)] = \theta_0$.
corollary[Finite-sample unbiasedness] Under Assumptions (ref), (ref), (ref), (ref), (ref), (ref), and predictable nuisances as in Assumption (ref), we have $\mathbb{E}[\widehat{\theta}]=\theta_0$.
proofBy linearity and Lemma (ref), \[ \mathbb{E}[\widehat\theta] = \frac{1}{n_{\mathrm{eff}}}\sum_{t\in\mathcal{T}} \mathbb{E}[\phi_t(\widehat m_{t-1})] = \frac{1}{n_{\mathrm{eff}}}\sum_{t\in\mathcal{T}} \theta_0 = \theta_0. \]
remark[Interpretation] Lemma (ref) is the conceptual replacement for fold-wise independence in i.i.d.\ DML. Rather than requiring approximate independence between nuisance estimation and score evaluation, we exploit exact conditional unbiasedness given the past. This martingale structure holds for any quality of nuisance estimates, provided they are predictable.
lemma[Moment bounds] Suppose Assumptions (ref), (ref), and (ref) hold. Then there exists a constant $C<\infty$ such that \[ \sup_{n\ge 1}\sup_{t\in\mathcal{T}} \mathbb{E}\big[|\widehat\phi_t|^4\big]\le C. \]
proofFix $t\in\mathcal{T}$. By overlap, $\pi_t\in[\varepsilon,1-\varepsilon]$ almost surely, so the inverse-propensity weights are uniformly bounded on scored indices. Applying $(a+b+c+d)^4\le 4^3(a^4+b^4+c^4+d^4)$ to the score formula (ref) (and hence to (ref)) and using Assumptions (ref) and (ref) yields the claimed uniform bound.
remark[Weaker than i.i.d.] Lemma (ref) uses Assumption (ref) only through the implication $\mathbb{E}[\tau(X_t)\mid \mathcal{F}_{t-1}]=\theta_0$. Thus the MDS argument extends to any arrival process satisfying this mean-stationarity condition.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{The Variance Structure}

Before stating the CLT, we analyze the variance structure. Define the realized and predictable quadratic variations:

equation[equation omitted — 192 chars of source]
proposition[Conditional second-moment decomposition and oracle benchmark] Suppose Assumptions (ref), (ref), and (ref) hold. Let $\sigma_a^2(X_t):=\operatorname{Var}(Y_t(a)\mid X_t)$ and define the regression bias terms $b_a(x):=m_a(x)-m_a^\star(x)$. Then for each $t\in\mathcal{T}$, \begin{equation} \mathbb{E}[(\phi_t(m)-\theta_0)^2 \mid \mathcal{G}_t] = (\tau(X_t)-\theta_0)^2 +\frac{\sigma_1^2(X_t)}{\pi_t} +\frac{\sigma_0^2(X_t)}{1-\pi_t} +\pi_t(1-\pi_t)\left(\frac{b_1(X_t)}{\pi_t}+\frac{b_0(X_t)}{1-\pi_t}\right)^2. \end{equation} Moreover, since $\mathbb{E}[\phi_t(m)\mid \mathcal{G}_t]=\tau(X_t)$ for any $m$ that is $\mathcal{F}_{t-1}$-measurable, the last three terms in (ref) equal $\operatorname{Var}(\phi_t(m)\mid \mathcal{G}_t)$. In particular, relative to the oracle score $\phi_t(m^\star)$ (for which $b_0=b_1\equiv 0$), using a misspecified $m$ adds a nonnegative augmentation term to the conditional second moment and conditional variance.
proofDefine $\varepsilon_{t,a}:=Y_t(a)-m_a^\star(X_t)$, so that $\mathbb{E}[\varepsilon_{t,a}\mid X_t]=0$ and $\mathbb{E}[\varepsilon_{t,a}^2\mid X_t]=\sigma_a^2(X_t)$. A direct expansion using sequential randomization yields \[ \phi_t(m)-\tau(X_t) = \frac{A_t}{\pi_t}\varepsilon_{t,1}-\frac{1-A_t}{1-\pi_t}\varepsilon_{t,0} -\frac{A_t-\pi_t}{\pi_t}b_1(X_t)-\frac{A_t-\pi_t}{1-\pi_t}b_0(X_t), \] from which (ref) follows by taking conditional second moments given $\mathcal{G}_t$ and adding $(\tau(X_t)-\theta_0)^2$. Finally, since $\mathbb{E}[\phi_t(m)\mid \mathcal{G}_t]=\tau(X_t)$ for any $m$ that is $\mathcal{F}_{t-1}$-measurable (in particular, any predictable learner), the last three terms in (ref) are all centered at the oracle benchmark $\tau(X_t)$, and the proof is complete.
remark[Oracle case and variance inflation are immediate] Setting $m=m^\star$ in Proposition (ref) removes the nonnegative augmentation term and yields the oracle conditional second moment. The factors $1/\pi_t$ and $1/(1-\pi_t)$ make explicit how extreme propensities inflate uncertainty, motivating overlap enforcement at the design stage.
remark[Why learn $m$ without rates?] Theorem (ref) yields valid prespecified-horizon coverage under predictability and logged propensities without requiring that $\widehat m_{t-1,a}$ converges to $m_a^\star$ at any particular rate. Learning $m$ is therefore not needed for validity; it is needed for precision. Proposition (ref) shows that the oracle score $\phi_t(m^\star)$ removes a nonnegative augmentation term from the conditional second moment, providing a natural benchmark for how nuisance quality affects the size of the studentizer. The oracle regression $m^\star$ minimizes the conditional variance within the AIPW score family (Proposition (ref)), and the augmentation term shrinks as $m$ approaches $m^\star$. This does not imply monotone tightening relative to an arbitrary baseline across learning iterations. There is no monotonic guarantee relative to an arbitrary baseline. A poorly chosen nuisance can increase or decrease variance compared to another misspecified choice.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Studentized Asymptotic Normality}

We now state our main inferential result. The key ingredient is that the realized quadratic variation $Q_{\mathcal{T}}$ and its predictable counterpart $V_{\mathcal{T}}^2$ are asymptotically equivalent under mild conditions, enabling studentization with the feasible sample variance $\widehat V$.

\paragraph{Roadmap.} Lemma (ref) identifies the scored sum $S_{\mathcal{T}}$ as a martingale sum with predictable quadratic variation $V_{\mathcal{T}}^2$ and realized quadratic variation $Q_{\mathcal{T}}$; Assumption (ref) ensures that the scored indicator $\mathbbm{1}\{t\in\mathcal{T}\}$ is deterministic (hence predictable). To apply the self-normalized martingale CLT in Theorem (ref), we verify: (i) $V_{\mathcal{T}}^2\to\infty$ (Assumption (ref)); (ii) $Q_{\mathcal{T}}/V_{\mathcal{T}}^2\to_p 1$ via a standard martingale variance-ratio argument, using the uniform fourth-moment bound from Lemma (ref); and (iii) a Lyapunov condition, again by Lemma (ref) together with Assumption (ref). Finally, Proposition (ref) justifies replacing $Q_{\mathcal{T}}$ by $(n_{\mathrm{eff}}-1)\widehat V$ in the Studentized statistic.

assumption[Variance growth / nondegeneracy] As $n\to\infty$, $n_{\mathrm{eff}}:=|\mathcal{T}|\to\infty$ and there exists $v_->0$ such that \[ \mathbb{P}\left(V_{\mathcal{T}}^2 \ge v_- n_{\mathrm{eff}}\right)\to 1. \] This is a nondegeneracy lower bound ensuring the predictable quadratic variation grows at least linearly in $n_{\mathrm{eff}}$; it does not impose variance stabilization.

A simple primitive sufficient condition for Assumption (ref) is that outcome noise is nondegenerate and overlap holds under the standing causal/sequential assumptions of Section (ref). The following corollary records a sufficient condition; see Appendix (ref) for the proof.

corollary[A sufficient condition for variance growth] Suppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold, where Assumption (ref) (Appendix) is the nondegeneracy condition $\mathbb{E}[\sigma_1^2(X)+\sigma_0^2(X)]\ge \underline{\sigma}^2>0$. Then Assumption (ref) holds.
remark[Nondegeneracy and numerical guardrails] Assumption (ref) implies the quadratic variation grows, so $\widehat V$ is bounded away from zero with high probability asymptotically. In finite samples, if $\widehat V=0$ (e.g.\ constant outcomes), inference is uninformative; report this as a design/measurement degeneracy.
theorem[Studentized fixed-horizon inference without variance stabilization] Suppose Assumptions (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref), and (ref) hold. Let $n_{\mathrm{eff}}:=|\mathcal{T}|$ and recall $S_{\mathcal{T}}=\sum_{t\in\mathcal{T}}(\widehat\phi_t-\theta_0)$ and $Q_{\mathcal{T}}=\sum_{t\in\mathcal{T}}(\widehat\phi_t-\theta_0)^2$. Then \begin{equation} \frac{S_{\mathcal{T}}}{\sqrt{Q_{\mathcal{T}}}}\Rightarrow \mathcal{N}(0,1), \end{equation} and the practical Studentized statistic satisfies \begin{equation} \frac{\sqrt{n_{\mathrm{eff}}}(\widehat\theta-\theta_0)}{\widehat V^{1/2}}\Rightarrow \mathcal{N}(0,1), \end{equation} where $\widehat V$ is defined in (ref). Consequently, the Wald interval (ref) has asymptotically correct coverage at the prespecified horizon, without requiring variance stabilization (i.e., $V_{\mathcal{T}}^2/n_{\mathrm{eff}}\to V$ deterministic; Assumption (ref)).
proofLet $\xi_t:=\mathbbm{1}\{t\in\mathcal{T}\}(\widehat\phi_t-\theta_0)$ and $S_{\mathcal{T}}:=\sum_{t\in\mathcal{T}}\xi_t$. Under Assumptions (ref), (ref), (ref), (ref), and (ref), Lemma (ref) implies that $(\xi_t,\mathcal{F}_t)$ is a martingale difference array. Assumption (ref) gives $V_{\mathcal{T}}^2:=\sum_{t\in\mathcal{T}}\mathbb{E}[\xi_t^2\mid \mathcal{F}_{t-1}]\to\infty$. Lemma (ref) (which uses Assumptions (ref), (ref), and (ref)) provides a uniform fourth-moment bound for the scored increments; together with Assumption (ref), a standard martingale variance-ratio calculation (see Appendix (ref)) verifies the variance-ratio and Lyapunov conditions required by Theorem (ref). Therefore Theorem (ref) yields $S_{\mathcal{T}}/\sqrt{Q_{\mathcal{T}}}\Rightarrow \mathcal{N}(0,1)$, i.e., (ref). Finally, Proposition (ref) implies $(n_{\mathrm{eff}}-1)\widehat V/Q_{\mathcal{T}}\to_p 1$, and Slutsky's theorem yields (ref).
remark[Validity vs.\ precision and consistency] The studentized CLT in Theorem (ref) does not require $\widehat m_{t-1}$ to converge to $m^\star$ for validity. Predictability (Assumption (ref)) together with overlap and moment/stability conditions (Assumptions (ref)--(ref)) and variance growth (Assumption (ref)) are sufficient. Nuisance learning is nonetheless valuable: by Proposition (ref) the oracle regression $m^\star$ minimizes the conditional variance within the AIPW score family (Proposition (ref)), and the augmentation term shrinks as $m$ approaches $m^\star$; this does not imply monotone tightening relative to an arbitrary baseline across learning iterations. In particular, $\widehat\theta$ is consistent whenever $n_{\mathrm{eff}}\to\infty$, and $\widehat V$ is a feasible studentizer even when $V_{\mathcal{T}}^2/n_{\mathrm{eff}}$ does not stabilize.
remark[Relation to stabilization-based and anytime-valid approaches] Some fixed-horizon CLTs in the adaptive-experiment literature normalize by a deterministic asymptotic variance, imposing stabilization/design-stability conditions ensuring that an average conditional variance (or, equivalently, $V_{\mathcal{T}}^2/n_{\mathrm{eff}}$) converges to a deterministic limit hadad2021confidence,zhan2021off,katoYasuiMcAlinn2021,cook2024semiparametric,sengupta2025designstability,zenati2025kernel. Such assumptions are appropriate when the deployed policy stabilizes or when one explicitly engineers stability. Theorem (ref) takes a different route. We normalize by the realized quadratic variation, so validity does not require a deterministic variance limit. This directly accommodates regimes where propensities oscillate, converge to random limits, or otherwise fail to stabilize, while still delivering a conventional fixed-horizon Wald interval. Finally, Theorem (ref) is a prespecified-horizon result. If the experiment is continuously monitored and potentially stopped early based on the data, time-uniform methods such as confidence sequences and time-uniform CLTs are needed howard2021confidence,waudbySmithRamdas2024betting,waudby2024timeuniform. Because our score increments form an exact MDS, the setup is compatible with standard time-uniform martingale methods, but we do not develop anytime-valid inference here.

Appendix (ref) collects three extensions not needed for the main methods-note claim: (i) a non-studentized CLT under variance stabilization; (ii) oracle equivalence under $L^2$ nuisance consistency; and (iii) a primitive sufficient condition for variance growth.

\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Simulation Study}

We compare fixed-horizon 95% Wald intervals of the form (ref) under different normalizations and nuisance-fitting regimes, depending on the design:

enumerate[leftmargin=2em,label=(CI\arabic*)] • SN (self-normalized): the proposed studentized Wald interval (ref). • Fixed-$V$ (stabilization-style baseline): a Wald interval that replaces the (potentially random) long-run variance by a fixed constant $V_{\mathrm{fix}}$, \[ \widehat\theta \pm z_{1-\alpha/2}\sqrt{V_{\mathrm{fix}}/n_{\mathrm{eff}}}, \] thereby behaving as if Assumption (ref) held with deterministic limit. We specify $V_{\mathrm{fix}}$ explicitly in each design. • Leaky baselines (Designs C1 and D): nuisance fits that violate predictability (Assumption (ref)) by using contemporaneous or future outcomes when scoring unit $t$. When computed, the score is still (ref) and the CI is (ref), but the fixed-horizon guarantee of Theorem (ref) does not apply. • Regime-aware Fixed-$V$ (Design A only): an infeasible regime-aware fixed-variance benchmark that plugs in the correct regime-specific $V_{\mathrm{fix}}$ based on the burn-in sign.

For each design we report: (i) two-sided 95% coverage; (ii) average CI length; and (iii) where informative, mean bias and/or null rejection rates. Monte Carlo standard errors for coverage are {\,($\sqrt{\hat p(1-\hat p)/R}$)}. Tables report $n$ and $n_{\mathrm{eff}}$ explicitly. Unless noted, we use $R=1000$ replications and $n\in\{250,500,1000,2000,5000\}$.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design A: random long-run variance}

During burn-in, $\pi_t=0.5$ for $t\le n_0=50$, and we compute the burn-in estimate \[ \widehat\tau_{\mathrm{burn}} :=\frac{1}{n_0}\sum_{t=1}^{n_0}\left(\frac{A_tY_t}{\pi_t}-\frac{(1-A_t)Y_t}{1-\pi_t}\right) \] using only burn-in units. After burn-in, if $\widehat\tau_{\mathrm{burn}}\ge 0$ then $\pi_t=0.8$, else $\pi_t=0.2$. Outcomes: $Y(0)\sim\mathcal{N}(0,1)$ and $Y(1)=\varepsilon_1$ with $\varepsilon_1\sim\mathcal{N}(0,9)$. There are no covariates; with $m_0\equiv m_1\equiv 0$, the AIPW score (ref) reduces to the usual inverse-propensity weighted (IPW) score.

On $\mathcal{T}=\{n_0+1,\dots,n\}$ we have $n_{\mathrm{eff}}=n-n_0$ and $\pi_t$ is constant after burn-in, so the oracle conditional variance converges to a random limit: \[ \frac{V_{\mathcal{T}}^2}{n_{\mathrm{eff}}}\ \to\

cases9/0.8 + 1/0.2 = 16.25, & \widehat\tau_{\mathrm{burn}}\ge 0,\\ 9/0.2 + 1/0.8 = 46.25, & \widehat\tau_{\mathrm{burn}}<0.

\] Thus Assumption (ref) fails (no deterministic variance limit). The SN CI remains valid under Theorem (ref) because it self-normalizes by realized quadratic variation. SN uses (ref). Fixed-$V$ sets $V_{\mathrm{fix}}:=31.25$ (the unconditional mean of the two variance limits above). Regime-aware Fixed-$V$ plugs in $V_{\mathrm{fix}}=16.25$ or $46.25$ depending on the realized burn-in sign.

table[table omitted — 1,915 chars of source]

Table (ref) reports marginal coverage and average length for Design A as $n$ increases. In this design, the studentized SN interval remains close to the nominal $95\%$ level across horizons, consistent with Theorem (ref). The fixed-$V$ interval normalizes by a single deterministic target $V$ that averages over the two post--burn-in propensity regimes. As a consequence, its marginal coverage can be close to nominal while still exhibiting regime-dependent over/under coverage (Table (ref)). The regime-aware fixed-$V$ benchmark, which plugs in the correct regime-specific variance constant, illustrates that classical normalization works once one conditions on (and correctly accounts for) the realized variance regime. The SN interval achieves this conditioning automatically via studentization.

table[table omitted — 773 chars of source]

Design A illustrates the basic phenomenon targeted by Theorem (ref): the variance proxy $V_{\mathcal{T}}^2/n_{\mathrm{eff}}$ does not converge to a single deterministic constant because it depends on the realized post--burn-in propensity regime. Table (ref) shows that a fixed-$V$ normalization can be conservative when the realized regime has low variance ($\pi_t=0.8$) and anti-conservative when the realized regime has high variance ($\pi_t=0.2$). In contrast, the SN interval studentizes by the realized quadratic variation and maintains near-nominal conditional coverage in both regimes.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design B: stabilization benchmark}

Design B provides a benchmark setting in which the conditional variance stabilizes. We adopt the outcome model from Design A with $\tau=0$, $Y(0)\sim\mathcal{N}(0,1)$, and $Y(1)=\varepsilon_1$ where $\varepsilon_1\sim\mathcal{N}(0,9)$. Treatment is assigned with a constant executed propensity $\pi_t\equiv 0.6$ for all $t\le n$, so there is no burn-in and $n_{\mathrm{eff}}=n$. Because there are no covariates, the AIPW score coincides with the IPW score. In this setting, Assumption (ref) holds with a deterministic long-run variance, \[ V_{\mathrm{fix}} \;=\; \frac{9}{0.6} + \frac{1}{0.4} \;=\; 17.5, \] and therefore the SN and Fixed-$V$ intervals are asymptotically equivalent.

table[table omitted — 1,199 chars of source]

Table (ref) shows that when the variance stabilizes, SN and Fixed-$V$ give nearly identical coverage and length. Differences are within Monte Carlo error, consistent with the deterministic $V_{\mathrm{fix}}$.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design C1: leakage stress test}

Design C1 is a stress test for Assumption (ref) and Remark (ref). Covariates are $X_t\in\mathbb{R}^{p}$ with $p=20$ and i.i.d.\ $X_t\sim\mathcal{N}(0,I_p)$. Potential outcomes follow \[ Y_t(0)=m_0(X_t)+\varepsilon_{t0},\qquad Y_t(1)=m_0(X_t)+\tau(X_t)+\varepsilon_{t1}, \] where $m_0(x)$ is nonlinear and $\tau(x)=\tau_0+\delta\sin(x_1)$ so the superpopulation ATE equals $\theta_0=\mathbb{E}[\tau(X_t)]=\tau_0$. We report calibration at $\tau_0=0$ with $\delta=0.2$; errors are independent $t_{30}$ draws rescaled to unit variance (finite moments, but heavier tails than Gaussian).

To induce adaptive feedback, we split $\{1,\ldots,n\}$ into $K=5$ contiguous blocks and use the first block as burn-in with $\pi_t\equiv 0.5$. On subsequent blocks, the executed propensity is \[ \pi_t=\varepsilon+(1-2\varepsilon)\mathrm{expit}\big(\lambda\,\widehat\tau_{t-1}(X_t)\big),\qquad \widehat\tau_{t-1}(x)=\widehat m_{t-1,1}(x)-\widehat m_{t-1,0}(x), \] with $\varepsilon=0.1$ and $\lambda=2.5$. The predictable nuisance $\widehat m_{t-1,a}$ is fit by ridge regression on a low-dimensional feature map using only prior blocks and held fixed within the current block (forward fitting). The leaky baseline fits a richer ridge model once on the full sample, violating predictability.

table[table omitted — 2,237 chars of source]

Table (ref) makes the predictability requirement operational: the predictable SN-AIPW procedure remains near nominal, while the leaky full-sample nuisance substantially undercovers and over-rejects at smaller horizons, consistent with the warning in Remark (ref).

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design C2: nuisance quality}

Design C2 isolates how nuisance quality affects precision without confounding from adaptivity. Covariates $X_t\in\mathbb{R}^5$ follow a correlated normal with AR(1) correlation $\rho=0.5$; treatment is assigned with constant propensity $\pi_t\equiv 0.5$. Potential outcomes are linear with homoskedastic noise: \[ Y_t(0)=X_{t1}+X_{t2}+\varepsilon_{t0},\qquad Y_t(1)=X_{t1}+X_{t2}+\tau+\varepsilon_{t1}, \qquad \varepsilon_{t0},\varepsilon_{t1}\stackrel{\text{i.i.d.}}{\sim}\mathcal{N}(0,1), \] so the oracle regression removes the augmentation term in Proposition (ref). We use $K=10$ forward blocks (first block omitted) and report results at $n=5000$ under $\tau=0$.

table[table omitted — 1,224 chars of source]

Table (ref) shows the oracle benchmarking message: coverage remains near nominal across nuisance choices, while interval length (and $\widehat V$) inflates as the nuisance becomes more misspecified, consistent with Proposition (ref) and the “validity vs precision” discussion following Theorem (ref).

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design D: adaptive policy and logging integrity}

Design D is a simple contextual-bandit-style adaptive experiment that highlights the design-time role of executed propensity logging (Assumption (ref)). Covariates are $X_t\in\mathbb{R}^{10}$ with correlated normal $X_t\sim\mathcal{N}(0,\Sigma)$ where $\Sigma_{ij}=\rho^{|i-j|}$ and $\rho=0.3$. The baseline outcome, heterogeneous treatment effect, and heteroskedastic noise are \[ m_0(x)=0.8x_1+0.5x_2^2-0.5\cos(x_3)+0.25x_4,\qquad \tau(x)=0.5x_1+0.5\sin(x_2)+0.25\mathbbm{1}\{x_3>0\}-0.25x_4x_5, \] \[ \sigma(x)=1+0.5|x_1|,\qquad Y_t(a)=m_a(X_t)+\varepsilon_{ta},\ \ \varepsilon_{ta}\mid X_t\sim \mathcal{N}(0,\sigma(X_t)^2), \] so $\theta_0=\mathbb{E}[\tau(X_t)]$ is the ATE. We use burn-in $n_0=100$ with $\pi_t\equiv 0.5$ and then update in blocks of size 100: on each block we fit arm-specific linear regressions on past data only and choose the executed propensity by either an $\varepsilon$-greedy rule or a softmax rule, clipping to $[0.05,0.95]$ to enforce overlap.

We compare: (i) SN-AIPW with predictable (past-only) nuisance fits; (ii) Naive-iid-DML using leaky 5-fold cross-fitting that ignores time ordering; (iii) SN-IPW with $m\equiv 0$; (iv) SN-Oracle using the true $m_0,m_1$; and (v) SN-IPW-Assume0p5, an analysis-time mis-logging baseline that computes the score using $\pi_t\equiv 0.5$ rather than the logged executed propensity $\pi_t$.

table[table omitted — 2,818 chars of source]
table[table omitted — 2,604 chars of source]

Tables (ref)--(ref) show two complementary messages: with the logged executed propensities and predictable nuisance fitting, the studentized intervals are close to nominal under both adaptive policies; in contrast, the mis-logged “Assume0p5” baseline is severely biased and has catastrophic undercoverage, illustrating why Assumption (ref) is design-critical.

\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Variance-ratio variability}

figure[figure omitted — 624 chars of source]

Figure (ref) shows two distinct variance modes for $V_{\mathcal{T}}^2/n_{\mathrm{eff}}$ induced by the post--burn-in propensity regime in Design A. This bimodality explains the regime-dependent length and coverage patterns in Table (ref): intervals that plug in a single deterministic variance target are too short in the high-variance regime and too long in the low-variance regime, even if their marginal coverage can be closer to nominal by averaging over regimes. The SN interval uses the realized quadratic variation (equivalently $\widehat V$) within each replication and therefore adapts to the realized variance regime by construction.

\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Conclusion}

Modern adaptive experiments often retain a classical reporting requirement: a single end-of-study confidence interval for the superpopulation ATE $\theta_0$ at a prespecified horizon. The difficulty is not estimation per se, but calibration. When assignment probabilities evolve, the predictable quadratic variation of AIPW/DML score increments can remain replication-random, so a Wald statistic normalized by a deterministic variance target can be well behaved marginally yet miscalibrated conditional on the realized propensity regime. Our analysis keeps the usual logged-propensity AIPW/DML estimator fixed and instead enforces a martingale representation through an auditable scoring contract. The contract is minimal but concrete: the experiment log must record, unit by unit, the executed propensity actually passed to the randomization device (Assumption (ref)), and the nuisance regressions used to score unit $t$ must be measurable with respect to the past history (Assumption (ref)). Forward cross-fitting provides an implementation pattern that satisfies this predictability requirement while allowing flexible learning.

Under the standard causal model, i.i.d.\ arrivals, and executed overlap on a prespecified scored set, these pipeline conditions imply that centered AIPW/DML score increments form an exact martingale difference sequence (Lemma (ref)). This step replaces the fold-wise independence logic common in i.i.d.\ DML with exact conditional unbiasedness given the past. With the martingale structure in place, inference follows by studentization along the realized experiment path. Theorem (ref) applies self-normalized martingale limit theory to show that the Studentized score sum, normalized by realized quadratic variation, converges to $\mathcal{N}(0,1)$ at the prespecified horizon even when the variance does not stabilize to any deterministic limit. Proposition (ref) justifies the practical studentizer: the sample-variance plug-in used in standard reporting is asymptotically equivalent to realized quadratic variation under the stated conditions.

Nuisance learning enters only through precision. Proposition (ref) decomposes the conditional second moment, yielding an oracle benchmark and isolating a nonnegative augmentation term attributable to regression error; learning better outcome regressions reduces this term and shrinks the realized quadratic variation. Consistency is not needed for fixed-horizon validity, but under $L^2$ convergence the feasible procedure is asymptotically equivalent to the oracle AIPW/DML score (Theorem (ref)).

The simulation designs in Section (ref) make the operational messages tangible: deterministic-variance normalizations can under- and over-cover across realized variance regimes, predictable scoring avoids leakage-induced failures, and analysis based on mis-logged propensities can be severely biased. These examples motivate the logging-and-fitting protocol and formalized as a data-contract checklist in Appendix (ref). In practice, this means persisting the executed propensity and randomization metadata for each unit, enforcing overlap at assignment time on the scored set, recording the exact training indices and randomness used for each nuisance fit, and producing intervals only at the prespecified horizon so that “peeking” and optional stopping are excluded by design.

The contract is auditable, but it is not a substitute for causal assumptions. Logged propensities and predictable fitting can be checked from stored artifacts, whereas no interference, sequential randomization, and the no-selection/i.i.d.\ arrival condition (or its mean-stationarity alternative) are substantive and not fully testable from logs (Remark (ref)). Likewise, our conclusions are strictly fixed-horizon: the intervals are not anytime-valid and should not be used under continuous monitoring or data-dependent stopping. When time-uniform guarantees are required, confidence sequences and related martingale methods are the appropriate alternatives.

Several extensions are immediate within the present framework. One is to develop time-uniform analogues that leverage the same martingale score construction while changing only the inferential target. A second is to formalize predictable scored-set rules (Remark (ref)) and outcome-delay settings within the same audit contract. Further work could relax the fourth-moment requirement to the $(2+\delta)$-moment regime highlighted in Remark (ref), and broaden the arrival model beyond i.i.d.\ using the mean-stationarity condition noted in Remark (ref).