Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,301 characters · 0 sections · 30 citation commands
Fixed-Horizon Self-Normalized Inference for Adaptive Experiments via Martingale AIPW/DML with Logged Propensities
\begingroup \thispagestyle{plain}
{\color{TitleNavy}\rule{\textwidth}{1.15pt}}
{\color{TitleNavy}Fixed-Horizon Self-Normalized Inference for Adaptive Experiments via Martingale AIPW/DML \\ with Logged Propensities}
{ Gabriel Saco} {\normalsizeUniversidad del Pac\'ifico}
{\smallORCID: \href{https://orcid.org/0009-0009-8751-4154}{0009-0009-8751-4154}} {\smallReplication code: \href{https://github.com/gsaco/martingale-aipw-dml}{\nolinkurl{https://github.com/gsaco/martingale-aipw-dml}}}
{\color{TitleGold}\rule{0.72\textwidth}{0.9pt}} \endgroup
\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Introduction}
Adaptive randomized experiments---including response-adaptive clinical trials, contextual bandits, and large-scale platform experimentation systems---update assignment probabilities as data accrue to balance learning and deployment kasySautmann2021. In practice, however, many platforms still require conventional end-of-study reporting for classical causal estimands such as the superpopulation average treatment effect (ATE) computed once at a prespecified horizon. Hereafter, we denote by $\pi_t$ the executed assignment probability used to randomize unit $t$ after applying any platform guardrails, as recorded in the experiment log. In this paper, we study fixed-horizon Wald inference for the standard logged-propensity AIPW/DML estimator, the sample average of doubly robust pseudo-outcomes scored using these logged propensities. Under adaptive assignment, the propensity process $\{\pi_t\}$ is itself data-dependent, so the predictable quadratic variation of the AIPW/DML score increments can remain replication-random and need not converge to a single deterministic long-run variance target.
Some end-of-study Wald arguments for AIPW/A2IPW (and related DML estimators) under adaptivity proceed via Slutsky steps that rely on a deterministic variance target for the predictable quadratic variation, often enforced through stabilization-type or design-stability conditions on assignment probabilities and/or average conditional variances hadad2021confidence,zhan2021off,katoIshiharaHondaNarita2020,cook2024semiparametric,LiOwen2024,sengupta2025designstability. On modern platforms, however, the policy can keep reacting to noisy intermediate estimates. Clipping and guardrails can activate intermittently and batch updates can induce regime switches. When this happens, a Wald statistic normalized using a single deterministic variance target can be systematically miscalibrated conditional on the realized variance regime, even if marginal coverage appears close to nominal.
We treat the adaptive assignment policy as given but assume that the platform logs the executed propensity used to randomize each unit and that nuisance regressions used for AIPW/DML scoring are fit predictably using only past data. Related work emphasizes that careful use of the logging policy and past-only fitting is central for post-adaptive inference bibaut2021post,katoYasuiMcAlinn2021,cook2024semiparametric. Under these auditable conditions, the centered score increments form an exact martingale difference sequence, and we obtain fixed-horizon Wald inference by studentizing with realized quadratic variation along the realized propensity path. This yields asymptotic $\mathcal{N}(0,1)$ calibration without requiring the predictable quadratic variation to converge to a deterministic long-run variance target.\footnote{We do not claim anytime-valid or optional-stopping guarantees.}
\noindentContributions.
The subsequent sections are structured as follows. Section (ref) reviews related work. Section (ref) states the model and assumptions. Section (ref) presents the estimator and auditable implementation details. Section (ref) develops the main theoretical results and Section (ref) reports simulations. The appendices collect supporting limit-theory background, additional results and proofs, and an operational logging protocol.
\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Related Work}
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Inference after adaptive data collection} A growing literature studies inference with adaptively collected data, where observations are generated under an evolving information set and classical i.i.d.\ arguments do not directly apply. Early econometric work by hahnHiranoKarlan2011 highlighted how propensity information can be leveraged for inference in sequential designs. More recent general frameworks derive asymptotic representations for sequential decisions and adaptive experiments under broad conditions hiranoPorter2023asymptotics. In the contextual-bandit and adaptive-experiment literature, fixed-horizon inference has also been developed via batched OLS/batchwise studentization arguments zhang2020large. Our focus is narrower but operationally central for experimentation platforms: fixed-horizon ATE reporting with logged executed propensities and predictable AIPW/DML scoring. In particular, we target settings where the predictable quadratic variation remains random across replications, so that deterministic-variance normalizations can be conditionally miscalibrated. General background on response-adaptive randomization and bandit-style designs can be found in rosenberger2015randomization and villar2015bandit.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Stabilization, adaptive weighting and batching} In adaptive experimentation and off-policy evaluation, evolving propensities can create heavy tails and regime-dependent uncertainty, motivating variance-control strategies. In policy evaluation, hadad2021confidence and zhan2021off develop adaptive weighting schemes for augmented IPW/DR scores to obtain asymptotically normal $t$-statistics. For post-contextual-bandit inference, bibaut2021post proposes stabilized doubly robust constructions that estimate conditional scale components using only past data. Another route restricts the data-collection design---for example through batching---to recover classical CLTs for bandit estimators zhang2020large. In contrast, we keep the standard logged-propensity AIPW/DML estimator unchanged and do not require batching or adaptive reweighting. Instead, we rely on a martingale score representation and studentization by realized quadratic variation.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Design stability and deterministic variance-limit CLTs} A further line of work derives fixed-horizon CLTs for IPW/AIPW-type ATE estimators under explicit design stability conditions ensuring that inverse-propensity averages and/or average conditional variances converge to nonrandom limits, yielding conventional deterministic asymptotic variances for Wald reporting sengupta2025designstability. Related perspectives arise when one engineers assignment rules (or batchwise designs) to target efficiency or precision for A2IPW/DML-style estimators katoIshiharaHondaNarita2020,LiOwen2024,cook2024semiparametric. More generally, the adaptive-experiment literature emphasizes subtleties around what “the logging policy” means operationally when platforms implement clipping, guardrails, or algorithmic randomness katoYasuiMcAlinn2021. Our contribution is complementary: we assume the executed propensity actually used to randomize each unit is logged and we avoid requiring convergence of the predictable quadratic variation to a deterministic limit by using realized quadratic variation as the normalizer.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Self-normalized martingale theory for quadratic-variation studentization} Our fixed-horizon Wald statistic is a self-normalized martingale functional. Classic references for martingale CLTs and self-normalized processes include hallHeyde1980 and delapena2009self, as well as the survey shao2013survey. Modern probability theory provides refined asymptotic and nonasymptotic control for self-normalized martingales, including Berry--Esseen bounds fanShao2017, Cram{\'e}r-type moderate deviations fanGramaLiuShao2019, and concentration inequalities bercuTouati2019. We apply this theory to AIPW/DML score increments under adaptive assignment, using realized quadratic variation to obtain a fixed-horizon $\mathcal{N}(0,1)$ approximation without a deterministic variance limit.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}*{Anytime-valid and time-uniform alternatives} The results in this paper are prespecified-horizon and do not provide optional-stopping guarantees. When inference must remain valid under continuous monitoring or data-dependent stopping, time-uniform methods based on test supermartingales and confidence sequences are appropriate howard2021confidence,waudby2024timeuniform. Related time-uniform tools also appear alongside fixed-time inference in adaptive-experiment work cook2024semiparametric,katoIshiharaHondaNarita2020,waudbySmithRamdas2024betting. We focus instead on conventional fixed-horizon reporting with realized quadratic-variation studentization (see Remark (ref) for additional discussion). \@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Model and Assumptions}
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Propensities and Logged Assignment Rule}
Consider a stream of experimental units indexed by $t=1,\dots,n$. For each unit, covariates $X_t$ are observed before assignment. The platform computes an assignment probability $\pi_t\in(0,1)$. Then, randomizes $A_t\in\{0,1\}$ using $\pi_t$. Finally, observes an outcome $Y_t$.
We observe and store $Z_t:=(X_t,A_t,Y_t,\pi_t)$ in time order. Let
be the data filtration. Let
be the $\sigma$-field immediately before randomization at time $t$, after the platform has computed the executed assignment probability $\pi_t$. The policy may be arbitrarily adaptive or randomized. The fixed-horizon validity results below require only that the realized executed propensity used to randomize $A_t$ is recorded in the log (Assumption (ref)).
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Audit diagnostics and data contract}
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Potential Outcomes and Target Parameter}
We adopt the potential outcomes framework with the usual no interference and consistency conventions.
The target parameter is the average treatment effect (ATE) in the superpopulation:
where $(X,Y(0),Y(1))$ denotes a generic draw from the superpopulation described in Assumption (ref). Assumption (ref) is standard. It rules out spillovers or network effects, which require separate methods.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Assumptions}
We maintain the following assumptions throughout.
We write all expectations with respect to this superpopulation draw, so $\theta_0=\mathbb{E}[Y(1)-Y(0)]$ and $\tau(x):=\mathbb{E}[Y(1)-Y(0)\mid X=x]$ satisfy $\theta_0=\mathbb{E}[\tau(X)]$. As noted in Remark (ref), the proofs only require the conditional mean-stationarity condition $\mathbb{E}[\tau(X_t)\mid \mathcal{F}_{t-1}]=\theta_0$.
\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Predictable AIPW/DML Estimation}
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{The Doubly Robust Score}
Let $m_a^\star(x):=\mathbb{E}[Y(a)\mid X=x]$ and write $m^\star:=(m_0^\star,m_1^\star)$. For any candidate $m=(m_0,m_1)$ and any $\pi\in(0,1)$, define the usual augmented inverse-propensity weighted (AIPW) / doubly robust (DR) score robinsRotnitzkyZhao1994,bangRobins2005 (pseudo-outcome)
When evaluating at the logged executed propensity, we write $\phi_t(m):=\phi_t(m,\pi_t)$ for brevity. The quantity $\phi_t(m,\pi_t)$ is observable given $(X_t,A_t,Y_t,\pi_t)$ and a supplied regression pair $m$.
For the oracle regression $m^\star$, one can decompose
The centered increment $\phi_t(m^\star)-\theta_0$ has mean zero and is the object governed by the martingale CLT.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Forward Cross-Fitting}
Standard cross-fitting partitions data into folds and estimates nuisances on held-out folds. Under adaptivity, we must respect time: nuisance estimates used at time $t$ must be constructed from data strictly before $t$.
Assumption (ref) is the formal predictability requirement (“no-peeking”, i.e., no use of contemporaneous or future outcomes) that makes the scored pseudo-outcomes a martingale difference sequence, and predictable fits are not based on non-predictable (“leaky”) scores (see Lemma (ref) below). The measurability here refers to the fitted function object $\widehat m_{t-1,a}$; it may then be evaluated at the current covariate $X_t$ to form $\widehat m_{t-1,a}(X_t)$. It rules out i.i.d.-style cross-fitting schemes that, even indirectly, use contemporaneous or future outcomes when constructing the nuisance for time $t$. A simple way to enforce predictability is forward cross-fitting (Definition (ref)): partition the time axis into blocks $I_1,\dots,I_K$, fit each nuisance model once per block using only data from $\cup_{j<k} I_j$, and hold the fitted model fixed while scoring units in $I_k$. To make this requirement auditable, an implementation should persist (i) the training indices used for each fit, (ii) fold assignments if internal cross-validation is used, and (iii) all learner randomness.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{The Estimator}
Let $\mathcal{T}\subseteq\{1,\dots,n\}$ denote the set of indices for which the nuisance used at time $t$ is predictable (i.e., $\widehat m_{t-1}$ is $\mathcal{F}_{t-1}$-measurable). Under forward cross-fitting with a burn-in block $I_1$, one typically takes $\mathcal{T}:=\{1,\dots,n\}\setminus I_1$. Let $n_{\mathrm{eff}}:=\sum_{t=1}^n \mathbbm{1}\{t\in\mathcal{T}\}=|\mathcal{T}|$. Throughout the asymptotic theory we treat the scored index set $\mathcal{T}$ as deterministic, as is the case under the forward-block construction in Definition (ref).
\paragraph{Estimator and studentizing factor.} The cross-fitted pseudo-outcome is
where $\widehat m_{t-1}$ is the predictable nuisance estimate.
\paragraph{Score sum notation.} Define the centered increments $\xi_t:=\widehat\phi_t-\theta_0$ and the score sum $S_{\mathcal{T}}:=\sum_{t\in\mathcal{T}}\xi_t=n_{\mathrm{eff}}(\widehat\theta-\theta_0)$.
\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Main Theoretical Results}
This section isolates the logic behind predictable AIPW/DML in adaptive experiments. First, under logged executed propensities and predictable nuisance fits, the centered score is an exact martingale difference, yielding finite-sample conditional unbiasedness (Proposition (ref) and Lemma (ref)). Second, fixed-horizon inference follows from a self-normalized martingale CLT applied to the martingale sum, with feasibility provided by a sample-variance approximation (Theorem (ref)). Table (ref) in Appendix (ref) maps each assumption to each statement, and Appendix (ref) states the martingale limit theorem invoked in the proofs.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{The Martingale Structure}
The central insight of our analysis is that forward cross-fitting induces a martingale difference structure. We first establish an identification result for the pseudo-outcome.
The key consequence is that the pseudo-outcome, when evaluated with any predictable nuisance estimates, remains conditionally unbiased.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{The Variance Structure}
Before stating the CLT, we analyze the variance structure. Define the realized and predictable quadratic variations:
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Studentized Asymptotic Normality}
We now state our main inferential result. The key ingredient is that the realized quadratic variation $Q_{\mathcal{T}}$ and its predictable counterpart $V_{\mathcal{T}}^2$ are asymptotically equivalent under mild conditions, enabling studentization with the feasible sample variance $\widehat V$.
\paragraph{Roadmap.} Lemma (ref) identifies the scored sum $S_{\mathcal{T}}$ as a martingale sum with predictable quadratic variation $V_{\mathcal{T}}^2$ and realized quadratic variation $Q_{\mathcal{T}}$; Assumption (ref) ensures that the scored indicator $\mathbbm{1}\{t\in\mathcal{T}\}$ is deterministic (hence predictable). To apply the self-normalized martingale CLT in Theorem (ref), we verify: (i) $V_{\mathcal{T}}^2\to\infty$ (Assumption (ref)); (ii) $Q_{\mathcal{T}}/V_{\mathcal{T}}^2\to_p 1$ via a standard martingale variance-ratio argument, using the uniform fourth-moment bound from Lemma (ref); and (iii) a Lyapunov condition, again by Lemma (ref) together with Assumption (ref). Finally, Proposition (ref) justifies replacing $Q_{\mathcal{T}}$ by $(n_{\mathrm{eff}}-1)\widehat V$ in the Studentized statistic.
A simple primitive sufficient condition for Assumption (ref) is that outcome noise is nondegenerate and overlap holds under the standing causal/sequential assumptions of Section (ref). The following corollary records a sufficient condition; see Appendix (ref) for the proof.
Appendix (ref) collects three extensions not needed for the main methods-note claim: (i) a non-studentized CLT under variance stabilization; (ii) oracle equivalence under $L^2$ nuisance consistency; and (iii) a primitive sufficient condition for variance growth.
\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Simulation Study}
We compare fixed-horizon 95% Wald intervals of the form (ref) under different normalizations and nuisance-fitting regimes, depending on the design:
For each design we report: (i) two-sided 95% coverage; (ii) average CI length; and (iii) where informative, mean bias and/or null rejection rates. Monte Carlo standard errors for coverage are {\,($\sqrt{\hat p(1-\hat p)/R}$)}. Tables report $n$ and $n_{\mathrm{eff}}$ explicitly. Unless noted, we use $R=1000$ replications and $n\in\{250,500,1000,2000,5000\}$.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design A: random long-run variance}
During burn-in, $\pi_t=0.5$ for $t\le n_0=50$, and we compute the burn-in estimate \[ \widehat\tau_{\mathrm{burn}} :=\frac{1}{n_0}\sum_{t=1}^{n_0}\left(\frac{A_tY_t}{\pi_t}-\frac{(1-A_t)Y_t}{1-\pi_t}\right) \] using only burn-in units. After burn-in, if $\widehat\tau_{\mathrm{burn}}\ge 0$ then $\pi_t=0.8$, else $\pi_t=0.2$. Outcomes: $Y(0)\sim\mathcal{N}(0,1)$ and $Y(1)=\varepsilon_1$ with $\varepsilon_1\sim\mathcal{N}(0,9)$. There are no covariates; with $m_0\equiv m_1\equiv 0$, the AIPW score (ref) reduces to the usual inverse-propensity weighted (IPW) score.
On $\mathcal{T}=\{n_0+1,\dots,n\}$ we have $n_{\mathrm{eff}}=n-n_0$ and $\pi_t$ is constant after burn-in, so the oracle conditional variance converges to a random limit: \[ \frac{V_{\mathcal{T}}^2}{n_{\mathrm{eff}}}\ \to\
\] Thus Assumption (ref) fails (no deterministic variance limit). The SN CI remains valid under Theorem (ref) because it self-normalizes by realized quadratic variation. SN uses (ref). Fixed-$V$ sets $V_{\mathrm{fix}}:=31.25$ (the unconditional mean of the two variance limits above). Regime-aware Fixed-$V$ plugs in $V_{\mathrm{fix}}=16.25$ or $46.25$ depending on the realized burn-in sign.
Table (ref) reports marginal coverage and average length for Design A as $n$ increases. In this design, the studentized SN interval remains close to the nominal $95\%$ level across horizons, consistent with Theorem (ref). The fixed-$V$ interval normalizes by a single deterministic target $V$ that averages over the two post--burn-in propensity regimes. As a consequence, its marginal coverage can be close to nominal while still exhibiting regime-dependent over/under coverage (Table (ref)). The regime-aware fixed-$V$ benchmark, which plugs in the correct regime-specific variance constant, illustrates that classical normalization works once one conditions on (and correctly accounts for) the realized variance regime. The SN interval achieves this conditioning automatically via studentization.
Design A illustrates the basic phenomenon targeted by Theorem (ref): the variance proxy $V_{\mathcal{T}}^2/n_{\mathrm{eff}}$ does not converge to a single deterministic constant because it depends on the realized post--burn-in propensity regime. Table (ref) shows that a fixed-$V$ normalization can be conservative when the realized regime has low variance ($\pi_t=0.8$) and anti-conservative when the realized regime has high variance ($\pi_t=0.2$). In contrast, the SN interval studentizes by the realized quadratic variation and maintains near-nominal conditional coverage in both regimes.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design B: stabilization benchmark}
Design B provides a benchmark setting in which the conditional variance stabilizes. We adopt the outcome model from Design A with $\tau=0$, $Y(0)\sim\mathcal{N}(0,1)$, and $Y(1)=\varepsilon_1$ where $\varepsilon_1\sim\mathcal{N}(0,9)$. Treatment is assigned with a constant executed propensity $\pi_t\equiv 0.6$ for all $t\le n$, so there is no burn-in and $n_{\mathrm{eff}}=n$. Because there are no covariates, the AIPW score coincides with the IPW score. In this setting, Assumption (ref) holds with a deterministic long-run variance, \[ V_{\mathrm{fix}} \;=\; \frac{9}{0.6} + \frac{1}{0.4} \;=\; 17.5, \] and therefore the SN and Fixed-$V$ intervals are asymptotically equivalent.
Table (ref) shows that when the variance stabilizes, SN and Fixed-$V$ give nearly identical coverage and length. Differences are within Monte Carlo error, consistent with the deterministic $V_{\mathrm{fix}}$.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design C1: leakage stress test}
Design C1 is a stress test for Assumption (ref) and Remark (ref). Covariates are $X_t\in\mathbb{R}^{p}$ with $p=20$ and i.i.d.\ $X_t\sim\mathcal{N}(0,I_p)$. Potential outcomes follow \[ Y_t(0)=m_0(X_t)+\varepsilon_{t0},\qquad Y_t(1)=m_0(X_t)+\tau(X_t)+\varepsilon_{t1}, \] where $m_0(x)$ is nonlinear and $\tau(x)=\tau_0+\delta\sin(x_1)$ so the superpopulation ATE equals $\theta_0=\mathbb{E}[\tau(X_t)]=\tau_0$. We report calibration at $\tau_0=0$ with $\delta=0.2$; errors are independent $t_{30}$ draws rescaled to unit variance (finite moments, but heavier tails than Gaussian).
To induce adaptive feedback, we split $\{1,\ldots,n\}$ into $K=5$ contiguous blocks and use the first block as burn-in with $\pi_t\equiv 0.5$. On subsequent blocks, the executed propensity is \[ \pi_t=\varepsilon+(1-2\varepsilon)\mathrm{expit}\big(\lambda\,\widehat\tau_{t-1}(X_t)\big),\qquad \widehat\tau_{t-1}(x)=\widehat m_{t-1,1}(x)-\widehat m_{t-1,0}(x), \] with $\varepsilon=0.1$ and $\lambda=2.5$. The predictable nuisance $\widehat m_{t-1,a}$ is fit by ridge regression on a low-dimensional feature map using only prior blocks and held fixed within the current block (forward fitting). The leaky baseline fits a richer ridge model once on the full sample, violating predictability.
Table (ref) makes the predictability requirement operational: the predictable SN-AIPW procedure remains near nominal, while the leaky full-sample nuisance substantially undercovers and over-rejects at smaller horizons, consistent with the warning in Remark (ref).
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design C2: nuisance quality}
Design C2 isolates how nuisance quality affects precision without confounding from adaptivity. Covariates $X_t\in\mathbb{R}^5$ follow a correlated normal with AR(1) correlation $\rho=0.5$; treatment is assigned with constant propensity $\pi_t\equiv 0.5$. Potential outcomes are linear with homoskedastic noise: \[ Y_t(0)=X_{t1}+X_{t2}+\varepsilon_{t0},\qquad Y_t(1)=X_{t1}+X_{t2}+\tau+\varepsilon_{t1}, \qquad \varepsilon_{t0},\varepsilon_{t1}\stackrel{\text{i.i.d.}}{\sim}\mathcal{N}(0,1), \] so the oracle regression removes the augmentation term in Proposition (ref). We use $K=10$ forward blocks (first block omitted) and report results at $n=5000$ under $\tau=0$.
Table (ref) shows the oracle benchmarking message: coverage remains near nominal across nuisance choices, while interval length (and $\widehat V$) inflates as the nuisance becomes more misspecified, consistent with Proposition (ref) and the “validity vs precision” discussion following Theorem (ref).
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Design D: adaptive policy and logging integrity}
Design D is a simple contextual-bandit-style adaptive experiment that highlights the design-time role of executed propensity logging (Assumption (ref)). Covariates are $X_t\in\mathbb{R}^{10}$ with correlated normal $X_t\sim\mathcal{N}(0,\Sigma)$ where $\Sigma_{ij}=\rho^{|i-j|}$ and $\rho=0.3$. The baseline outcome, heterogeneous treatment effect, and heteroskedastic noise are \[ m_0(x)=0.8x_1+0.5x_2^2-0.5\cos(x_3)+0.25x_4,\qquad \tau(x)=0.5x_1+0.5\sin(x_2)+0.25\mathbbm{1}\{x_3>0\}-0.25x_4x_5, \] \[ \sigma(x)=1+0.5|x_1|,\qquad Y_t(a)=m_a(X_t)+\varepsilon_{ta},\ \ \varepsilon_{ta}\mid X_t\sim \mathcal{N}(0,\sigma(X_t)^2), \] so $\theta_0=\mathbb{E}[\tau(X_t)]$ is the ATE. We use burn-in $n_0=100$ with $\pi_t\equiv 0.5$ and then update in blocks of size 100: on each block we fit arm-specific linear regressions on past data only and choose the executed propensity by either an $\varepsilon$-greedy rule or a softmax rule, clipping to $[0.05,0.95]$ to enforce overlap.
We compare: (i) SN-AIPW with predictable (past-only) nuisance fits; (ii) Naive-iid-DML using leaky 5-fold cross-fitting that ignores time ordering; (iii) SN-IPW with $m\equiv 0$; (iv) SN-Oracle using the true $m_0,m_1$; and (v) SN-IPW-Assume0p5, an analysis-time mis-logging baseline that computes the score using $\pi_t\equiv 0.5$ rather than the logged executed propensity $\pi_t$.
Tables (ref)--(ref) show two complementary messages: with the logged executed propensities and predictable nuisance fitting, the studentized intervals are close to nominal under both adaptive policies; in contrast, the mis-logged “Assume0p5” baseline is severely biased and has catastrophic undercoverage, illustrating why Assumption (ref) is design-critical.
\@startsection{subsection}{2}{\z@} {-3.0ex \@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Variance-ratio variability}
Figure (ref) shows two distinct variance modes for $V_{\mathcal{T}}^2/n_{\mathrm{eff}}$ induced by the post--burn-in propensity regime in Design A. This bimodality explains the regime-dependent length and coverage patterns in Table (ref): intervals that plug in a single deterministic variance target are too short in the high-variance regime and too long in the low-variance regime, even if their marginal coverage can be closer to nominal by averaging over regimes. The SN interval uses the realized quadratic variation (equivalently $\widehat V$) within each replication and therefore adapts to the realized variance regime by construction.
\@startsection{section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus .2ex} {\normalfont\color{TitleNavy}}{Conclusion}
Modern adaptive experiments often retain a classical reporting requirement: a single end-of-study confidence interval for the superpopulation ATE $\theta_0$ at a prespecified horizon. The difficulty is not estimation per se, but calibration. When assignment probabilities evolve, the predictable quadratic variation of AIPW/DML score increments can remain replication-random, so a Wald statistic normalized by a deterministic variance target can be well behaved marginally yet miscalibrated conditional on the realized propensity regime. Our analysis keeps the usual logged-propensity AIPW/DML estimator fixed and instead enforces a martingale representation through an auditable scoring contract. The contract is minimal but concrete: the experiment log must record, unit by unit, the executed propensity actually passed to the randomization device (Assumption (ref)), and the nuisance regressions used to score unit $t$ must be measurable with respect to the past history (Assumption (ref)). Forward cross-fitting provides an implementation pattern that satisfies this predictability requirement while allowing flexible learning.
Under the standard causal model, i.i.d.\ arrivals, and executed overlap on a prespecified scored set, these pipeline conditions imply that centered AIPW/DML score increments form an exact martingale difference sequence (Lemma (ref)). This step replaces the fold-wise independence logic common in i.i.d.\ DML with exact conditional unbiasedness given the past. With the martingale structure in place, inference follows by studentization along the realized experiment path. Theorem (ref) applies self-normalized martingale limit theory to show that the Studentized score sum, normalized by realized quadratic variation, converges to $\mathcal{N}(0,1)$ at the prespecified horizon even when the variance does not stabilize to any deterministic limit. Proposition (ref) justifies the practical studentizer: the sample-variance plug-in used in standard reporting is asymptotically equivalent to realized quadratic variation under the stated conditions.
Nuisance learning enters only through precision. Proposition (ref) decomposes the conditional second moment, yielding an oracle benchmark and isolating a nonnegative augmentation term attributable to regression error; learning better outcome regressions reduces this term and shrinks the realized quadratic variation. Consistency is not needed for fixed-horizon validity, but under $L^2$ convergence the feasible procedure is asymptotically equivalent to the oracle AIPW/DML score (Theorem (ref)).
The simulation designs in Section (ref) make the operational messages tangible: deterministic-variance normalizations can under- and over-cover across realized variance regimes, predictable scoring avoids leakage-induced failures, and analysis based on mis-logged propensities can be severely biased. These examples motivate the logging-and-fitting protocol and formalized as a data-contract checklist in Appendix (ref). In practice, this means persisting the executed propensity and randomization metadata for each unit, enforcing overlap at assignment time on the scored set, recording the exact training indices and randomness used for each nuisance fit, and producing intervals only at the prespecified horizon so that “peeking” and optional stopping are excluded by design.
The contract is auditable, but it is not a substitute for causal assumptions. Logged propensities and predictable fitting can be checked from stored artifacts, whereas no interference, sequential randomization, and the no-selection/i.i.d.\ arrival condition (or its mean-stationarity alternative) are substantive and not fully testable from logs (Remark (ref)). Likewise, our conclusions are strictly fixed-horizon: the intervals are not anytime-valid and should not be used under continuous monitoring or data-dependent stopping. When time-uniform guarantees are required, confidence sequences and related martingale methods are the appropriate alternatives.
Several extensions are immediate within the present framework. One is to develop time-uniform analogues that leverage the same martingale score construction while changing only the inferential target. A second is to formalize predictable scored-set rules (Remark (ref)) and outcome-delay settings within the same audit contract. Further work could relax the fourth-moment requirement to the $(2+\delta)$-moment regime highlighted in Remark (ref), and broaden the arrival model beyond i.i.d.\ using the mean-stationarity condition noted in Remark (ref).