EconBase
← Back to paper

Randomization Tests in Switchback Experiments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

85,145 characters · 18 sections · 37 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Randomization Tests in Switchback Experiments

\setstretch{1.5}

abstractSwitchback experiments—alternating treatment and control over time—are widely used when unit-level randomization is infeasible, outcomes are aggregated, or user interference is unavoidable. In practice, experimentation must support fast product cycles, so teams often run studies for limited durations and make decisions with modest samples. At the same time, outcomes in these time-indexed settings exhibit serial dependence, seasonality, and occasional heavy-tailed shocks, and temporal interference (carryover or anticipation) can render standard asymptotics and naive randomization tests unreliable. In this paper, we develop a randomization-test framework that delivers finite-sample valid, distribution-free p-values for several null hypotheses of interest using only the known assignment mechanism, without parametric assumptions on the outcome process. For causal effects of interests, we impose two primitive conditions—non-anticipation and a finite carryover horizon \(m\)—and construct conditional randomization tests (CRTs) based on an ex ante pooling of design blocks into “sections,” which yields a tractable conditional assignment law and ensures imputability of focal outcomes. We provide diagnostics for learning the carryover window and assessing non-anticipation, and we introduce studentized CRTs for a session-wise weak null that accommodates within-session seasonality with asymptotic validity. Power approximations under distributed-lag effects with AR(1) noise guide design and analysis choices, and simulations demonstrate favorable size and power relative to common alternatives. Our framework extends naturally to other time-indexed designs.

KEYWORDS: Causal inference; Conditional Randomization Test; Switchback Experiments.

\hypersetup{pageanchor=false} \thispagestyle{empty} \hypersetup{pageanchor=true} \setcounter{page}{1}

\setstretch{1.9}

Introduction

Randomized experiments play a central role in guiding decision making in online and operational systems such as marketplaces, transportation platforms, and digital advertising whymarketplace,Blake2015,Gordon2019,kohavi2020trustworthy,Wager2021,Bojinov2022Online,johari2021,li2022,Christensen2023,Tang2024,Masoero2026. Yet in many of these environments, unit-level randomization is infeasible, outcomes are observed only in aggregate, or interference across users is unavoidable. These constraints motivate switchback experiments, in which a platform alternates between treatment and control over time rather than across individuals Bojinov2023,jiang2025principledanalysiscrossoverdesigns,lin2025unifyingregressionbaseddesignbasedcausal.

Reliable inference in these time-indexed experiments is both practically important and methodologically challenging. Platforms often need to make decisions quickly—sometimes within days or weeks—which limits the available sample size kohavi2020trustworthy,gupta2019top. At the same time, outcomes form a time series with serial dependence, seasonal patterns, and occasionally heavy-tailed shocks, so normal approximations can be poorly calibrated at realistic horizons; for example, Bojinov2023 document settings in which calibration may require on the order of $T=1200$ periods. Inference is further complicated by temporal interference: outcomes may depend on recent treatments (carryover) or even future scheduled assignments (anticipation), making many null hypotheses non-sharp and invalidating naive Fisher randomization tests Athey2018,Basse2019,zhong2024unconditional.

To address these challenges, this paper develops a randomization-test framework for several null hypotheses of interest in switchback experiments. Our approach yields finite-sample valid, distribution-free $p$-values using only the known assignment mechanism, without imposing parametric assumptions on the outcome process. The goal is to make randomization inference operationally reliable in time-indexed experiments conducted over realistic, and often short, horizons.

To conduct inference for the total treatment effect—defined as the effect of sustained exposure to the intervention—we require only two standard assumptions on potential outcomes: non-anticipation, which rules out dependence on future assignments, and a finite carryover horizon, which allows outcomes to depend on recent treatment history but not on distant past assignments. Within regular switchback experiments in the sense of Bojinov2023, we construct conditional randomization tests that remain finite-sample valid and can be implemented via simple Monte Carlo resampling Athey2018,Basse2019,puelz2021,basse2024,liu2025randomizationinferencetwosidedmarket. Our approach pools design blocks into candidate sections in advance of observing the assignment path and conditions on those that are constant under the realized schedule, yielding a tractable randomization law and a valid reference distribution for outcomes not contaminated by carryover.

To mitigate the concern of potential violation of the assumptions, we also propose two complementary diagnostics for the temporal structure of interference. First, we develop a family of tests for the null hypothesis of at most $m$-period carryover effects and exploit the nested structure of these nulls to obtain a sequential procedure that controls the family-wise error rate while delivering an interpretable estimate of treatment memory length. Second, to assess the non-anticipation restriction, we adapt the unconditional Pairwise Imputation--based Randomization Test (PIRT) of zhong2024unconditional to switchback schedules, using a notion of imputable times where potential outcomes can be paired across schedules under the anticipation null.

In light of the practical interest for testing average effect, we also address weak-null inference, where the target is a mean-zero restriction on treatment effects in the post--burn-in regime. We introduce a session-wise weak null that requires mean-zero effects at each within-session position among the focal periods (e.g., Monday vs.\ Tuesday), accommodating within-session seasonality in practice. We construct studentized CRTs for this null and establish asymptotic validity under mild moment and stabilization conditions Ding2017,Wu2021,ZHAO2021278.

Finally, we study power under a superpopulation model with distributed-lag treatment effects and AR(1) noise. The resulting approximations clarify how block length, burn-in, and predetermined pooling jointly determine the signal-to-noise ratio of the studentized CRTs, providing practical guidance for design and analysis choices. Monte Carlo simulations corroborate these insights: the proposed CRT maintains near-nominal size under both Gaussian and heavy-tailed shocks, while Fisher tests under a misspecified sharp null over-reject in heavy tails and Horvitz--Thompson asymptotic tests are often conservative at realistic horizons; in terms of power, the CRT matches asymptotic tests under Gaussian noise and dominates size-correct alternatives under heavy-tailed disturbances. For diagnostics, the carryover CRT achieves accurate size with power increasing in effect magnitude and sample size, and the PIRT-based non-anticipation test controls size and delivers meaningful power; together with our theoretical approximations, these findings highlight block length, burn-in, and section pooling as first-order design levers for sensitivity in time-indexed experiments.

This paper contributes to the growing literature on switchback and other time‑indexed experimental designs. Bojinov2023 study regular switchback designs and develop largely asymptotic inference procedures. Related design‑based analyses for experiments with temporal structure include jiang2025principledanalysiscrossoverdesigns and lin2025unifyingregressionbaseddesignbasedcausal. By contrast, we adopt a Fisherian randomization perspective and construct conditioning‑based tests that deliver finite‑sample exact, distribution‑free p-values for standard switchback schedules.

The remainder of the paper is organized as follows. Section (ref) describes the setup and notation. Section (ref) develops finite-sample valid CRTs for total treatment effects, carryover length and non-anticipation. Section (ref) develops studentized CRTs for session-wise weak null hypotheses and provides asymptotic validity results. Section (ref) derives power approximations under a superpopulation model. Section (ref) illustrates the simulation results. Section (ref) concludes.

Setup and Notation

Switchback experiments and regular block randomization

We observe a single outcome time series over $T$ measurement periods. Let $[T]:=\{1,2,\dots,T\}$ index time, and let \[ \mathbf{W}:=(W_1,\dots,W_T)\in\{0,1\}^T \] denote the treatment assignment path, where $W_t=1$ indicates treatment is “on” during period $t$ and $W_t=0$ indicates treatment is “off.” For any candidate assignment path $\mathbf{w}=(w_1,\dots,w_T)\in\{0,1\}^T$, let $Y_t(\mathbf{w})\in\mathbb{R}$ be the potential outcome at time $t$ that would be observed if the assignment path were $\mathbf{w}$. The observed outcome satisfies the consistency relationship \[ Y_t^{\mathrm{obs}}=Y_t(\mathbf{W}^{\mathrm{obs}}),\qquad t\in[T], \] where $\mathbf{W}^{\mathrm{obs}}$ denotes the realized assignment.

Throughout, for integers $1\le a\le b\le T$, we write $\mathbf{w}_{a:b}:=(w_a,\dots,w_b)$ for a subvector, and we use $\mathbf{1}_d$ and $\mathbf{0}_d$ to denote length-$d$ all-ones and all-zeros vectors, respectively.

A switchback experiment is a time-based randomized experiment in which the treatment is held constant over contiguous blocks of time and may switch only at pre-specified times. In this paper, we work with the class of regular switchback experiments of Bojinov2023. Following Bojinov2023, let \[ \mathbb{T}=\{t_0=1<t_1<\cdots<t_K\}\subseteq[T], \qquad t_{K+1}:=T+1, \] be a deterministic set of switch times. These induce $K+1$ design blocks \[ B_k:=\{t: t_k\le t\le t_{k+1}-1\}, \qquad k=0,1,\dots,K. \] Let \[ \mathbb{Q}=(q_0,\dots,q_K)\in(0,1)^{K+1} \] denote the block-level treatment probabilities.

definition[Regular switchback experiment] Fix $(\mathbb{T},\mathbb{Q})$ as above. A regular switchback experiment draws independent block-level assignments $\{W^{(k)}\}_{k=0}^K$ with \[ W^{(k)}\sim \mathrm{Bernoulli}(q_k), \qquad k=0,1,\dots,K, \] and then sets \[ W_t:=W^{(k)}\quad\text{for all } t\in B_k. \] Equivalently, $\mathbf{W}$ is blockwise constant on each $B_k$, with independent block labels.

Definition (ref) covers common switchback implementations in which time is partitioned into intervals (e.g., hours, days, or weeks), and each interval is independently assigned treatment with a possibly time-varying probability. We will use

equation[equation omitted — 137 chars of source]

to denote the set of assignment paths that respect the design blocks.

Assumptions on temporal interference

Switchback experiments are typically used when outcomes may depend on recent treatment history, so we allow for temporal interference across periods through two primitive restrictions on potential outcomes.

assumption[Non-anticipating potential outcomes] For any $t\in[T]$, any prefix $\mathbf{w}_{1:t}\in\{0,1\}^t$, and any two continuation paths $\mathbf{w}'_{t+1:T},\mathbf{w}''_{t+1:T}\in\{0,1\}^{T-t}$, \[ Y_t(\mathbf{w}_{1:t},\mathbf{w}'_{t+1:T}) = Y_t(\mathbf{w}_{1:t},\mathbf{w}''_{t+1:T}). \]

Assumption (ref) rules out dependence of $Y_t$ on future treatment assignments. It is a timing restriction: potential outcomes at time $t$ are allowed to depend arbitrarily on the history up to $t$, but not on $\{w_{t+1},\dots,w_T\}$.

assumption[$m$-carryover effects] There exists a fixed and given $m\in\{0,1,\dots,T-1\}$ such that for any $t\in\{m+1,m+2,\dots,T\}$, any $\mathbf{w}_{t-m:t}\in\{0,1\}^{m+1}$, and any two histories $\mathbf{w}'_{1:t-m-1},\mathbf{w}''_{1:t-m-1}\in\{0,1\}^{t-m-1}$, \[ Y_t(\mathbf{w}'_{1:t-m-1},\mathbf{w}_{t-m:T}) = Y_t(\mathbf{w}''_{1:t-m-1},\mathbf{w}_{t-m:T}). \]

Assumption (ref) bounds the “memory” of the treatment: outcomes at time $t$ may depend on the most recent $m+1$ treatment indicators $(w_{t-m},\dots,w_t)$ but not on assignments more than $m$ periods in the past.

Together, Assumptions (ref) and (ref) imply a local dependence structure: for all $t\ge m+1$, the potential outcome $Y_t(\mathbf{w})$ depends on $\mathbf{w}$ only through the length-$(m+1)$ window $\mathbf{w}_{t-m:t}$. In particular, if two assignment paths $\mathbf{w}$ and $\mathbf{w}'$ satisfy $\mathbf{w}_{t-m:t}=\mathbf{w}'_{t-m:t}$, then $Y_t(\mathbf{w})=Y_t(\mathbf{w}')$ for all $t\ge m+1$.

remark[Interpretation of $m$] The parameter $m$ is the maximal carryover horizon: if treatment switches at time $s$, then its effect on outcomes may persist through time $s+m$, but it does not affect $Y_t$ for $t>s+m$ via channels other than the contemporaneous and recent treatment indicators. The case $m=0$ corresponds to no carryover beyond the current period.

Main Results

Overview of Conditional Randomization Tests

The classical Fisher Randomization Test (FRT) simulates the randomization distribution of a test statistic under the known assignment mechanism and compares it to the observed statistic imbens2015causal. This procedure is finite-sample valid for sharp null hypotheses, i.e., nulls that allow the analyst to impute all missing potential outcomes under every assignment in the design support. In switchback experiments with temporal interference, many hypotheses of interest are not sharp over the full assignment space. For example, a null that only links potential outcomes under two particular assignment paths (e.g., all-treated versus all-control) does not determine potential outcomes under intermediate switching paths. As a result, a naive FRT that resamples the full assignment path generally cannot be implemented by imputation.

Conditional randomization tests (CRTs) address this problem by restricting the resampling procedure to a conditioning event that renders the null hypothesis sharp on a subset of units and assignments. In the terminology of Basse2019, a conditioning event takes the form \[ \mathcal{C}=(\mathcal{U},\mathcal{W}), \] where $\mathcal{U}$ is a subset of units (here, time periods) and $\mathcal{W}$ is a subset of assignments. In the literature, $\mathcal{U}$ and $\mathcal{W}$ are commonly referred to as the focal units and the focal assignments, respectively. The analyst specifies a conditioning mechanism $p(\mathcal{C}\mid \mathbf{W})$, which may be random and may depend on the realized assignment $\mathbf{W}$. The CRT then samples assignments from the conditional law

equation[equation omitted — 123 chars of source]

and computes a test statistic using only outcomes on $\mathcal{U}$. Finite-sample validity follows when the test statistic is imputable on $\mathcal{C}$ under the null, meaning that its value under any $\mathbf{w}\in\mathcal{W}$ can be computed from the observed outcomes given the null restriction Basse2019.

Two general recipes are commonly used to construct conditional randomization tests. First, the framework of Basse2019 allows for non-degenerate conditioning mechanisms, in which $p(\mathcal{C}\mid \mathbf{W})$ is genuinely randomized. A valid CRT then draws a single conditioning event $\mathcal{C}\sim p(\cdot\mid \mathbf{W}^{\mathrm{obs}})$ and samples assignments from the induced conditional distribution (ref). The main drawback is computational: evaluating and sampling from (ref) can be expensive when the conditional randomization space is large or lacks exploitable structure. Second, puelz2021 study CRTs under degenerate conditioning mechanisms of the form \[ p(\mathcal{C}\mid \mathbf{W}) = \mathbf{1}\{\mathbf{W}\in\mathcal{W}\}, \] for some (possibly data-dependent) restricted randomization space $\mathcal{W}$. In this regime, validity hinges on an invariance requirement: the conditioning event used for resampling must not change as we vary assignments within the restricted space. Equivalently, if we view the conditioning event as a mapping $\mathcal{C}(\mathbf{w})$ generated from an assignment $\mathbf{w}$, then invariance requires

equation[equation omitted — 154 chars of source]

so that the conditioning event is constant across all focal assignments considered by the test.

Our switchback procedures fall into this second class. We employ a degenerate conditioning mechanism---$\mathcal{C}$ is a deterministic function of the realized assignment path---and our main technical task is to construct $\mathcal{C}$ so that (i) the null becomes imputable on the selected focal periods $\mathcal{U}$ and (ii) the invariance condition (ref) holds for the resulting restricted assignment set. In particular, our conditioning events are built from predetermined time partitions (fixed ex ante), so that resampling treatment labels within focal assignments cannot change the event itself. This yields a tractable conditional assignment law and permits efficient Monte Carlo CRT implementation while maintaining finite-sample validity under the nulls considered below.

Testing total treatment effects

We begin with randomization tests for the partially sharp null hypothesis of no total treatment effect in a switchback experiment:

equation[equation omitted — 124 chars of source]

where $\mathbf{1}_T$ and $\mathbf{0}_T$ denote the constant all-ones and all-zeros assignment paths. Under Assumptions (ref)--(ref), (ref) is equivalent to $Y_t(\mathbf{1}_{m+1})=Y_t(\mathbf{0}_{m+1})$ for all $t\ge m+1$ because $Y_t(\cdot)$ depends on $\mathbf{w}$ only through the last $m+1$ entries.

The key challenge is that, under carryover, (ref) is not directly imputable for all periods: changing the assignment at time $t$ can affect outcomes at times $t+1,\dots,t+m$. To construct a conditioning event that isolates imputable outcomes, we introduce predetermined time sections. Formally, a section is a contiguous time interval $[s,e]$ whose endpoints are fixed ex ante and that is formed by merging consecutive design blocks. We will restrict attention to periods $t$ whose relevant treatment history $(W_{t-m},\dots,W_t)$ lies entirely within a section that is constant under the realized assignment.

Fix a predetermined family of disjoint sections \[ \mathcal{S}^{\text{pre}} =\{[s_1,e_1],\dots,[s_J,e_J]\}, \qquad 1\le s_1\le e_1 < s_2\le \cdots < s_J\le e_J\le T, \] constructed ex ante (independent of $\mathbf{W}$). Each section is obtained by pooling consecutive design blocks, and must satisfy:

enumerate• (Merged from design blocks) For each $j$ there exist $0\le a_j\le b_j\le K$ such that \[ [s_j,e_j]=\{t: t_{a_j}\le t\le t_{b_j+1}-1\}. \] • (Length constraint) $e_j-s_j\ge m$ (equivalently, $|[s_j,e_j]|\ge m+1$).

In simulation, $\mathcal{S}^{\text{pre}}$ is constructed deterministically from the randomization design and the pre-specified carryover horizon $m$. Specifically, we pool consecutive design blocks greedily until the pooled length is at least $m+1$, then repeat this procedure on the remaining blocks. This produces a disjoint cover of $[T]$ and fixes section boundaries independently of the realized assignment path and all outcomes.\footnote{Any deterministic rule depending only on the design and $m$ would be valid. We adopt greedy pooling because it guarantees admissible section length while minimizing pooling.}

Given a realized assignment $\mathbf{W}$, let $\mathcal{S}(\mathbf{W})\subseteq\mathcal{S}^{\text{pre}}$ denote the subcollection of predetermined sections that are constant under $\mathbf{W}$: \[ \mathcal{S}(\mathbf{W}) = \Bigl\{[s_j,e_j]\in\mathcal{S}^{\text{pre}}:\ W_t=W_{t'}\ \text{for all }t,t'\in[s_j,e_j]\Bigr\}. \] Because $\mathcal{S}^{\text{pre}}$ is fixed ex ante, the map $\mathbf{W}\mapsto \mathcal{S}(\mathbf{W})$ is fully determined by the realized assignment and involves no analyst discretion.

Within each constant section, only the last $e-s-m+1$ periods are free of carryover contamination from outside the section. Accordingly, define the set of focal units

equation[equation omitted — 133 chars of source]

If $\mathcal{S}(\mathbf{W})=\varnothing$, then $\mathcal{U}(\mathbf{W})=\varnothing$ and no focal units are available.

In a CRT, focal units are the units on which the test statistic is computed. The availability of focal units depends on the design, the carryover horizon $m$, and the realized assignment path. Short sections relative to $m$ may yield $\mathcal{U}(\mathbf{W})=\varnothing$ with non-negligible probability. To guide implementation, practitioners may evaluate ex ante the expected number of focal units under the randomization design, see section (ref) for a detailed discussion.

Next, we describe how we construct the focal assignment space, over which we perform Monte Carlo resampling of treatment assignments. Conditioning on the realized constant sections $\mathcal{S}(\mathbf{W}^{\mathrm{obs}})$, we restrict resampling to assignments that (i) remain blockwise constant, (ii) are constant within each realized constant section, and (iii) match the observed assignment outside those sections:

equation[equation omitted — 464 chars of source]

This construction is tailored to imputability: under $H_0$ and Assumptions (ref)--(ref), the outcomes $\{Y_t^{\mathrm{obs}}: t\in\mathcal{U}(\mathbf{W}^{\mathrm{obs}})\}$ are unaffected by changes to $\mathbf{w}$ within $\mathcal{W}(\mathcal{S}(\mathbf{W}^{\mathrm{obs}}))$ outside the focal histories.

Let $T(\{Y_t\}_{t\in\mathcal{U}},\mathbf{w})$ be any test statistic computed from the focal outcomes and the assignment path.\footnote{A common choice is a weighted difference-in-means comparing focal outcomes in treated versus control periods, possibly with inverse-probability weights if $\{q_k\}$ vary across blocks.} Algorithm (ref) describes a Monte Carlo CRT that samples from the conditional assignment law on $\mathcal{W}(\mathcal{S}(\mathbf{W}^{\mathrm{obs}}))$ and computes a one-sided $p$-value.

algorithm[algorithm omitted — 1,425 chars of source]

The next lemma characterizes the conditional law of a merged section under independent block Bernoulli assignment.

lemma[Conditional law on a constant merged section] Suppose blocks $W^{(k)}$ are independent with $W^{(k)} \sim \mathrm{Bern}(q_k)$. For a merged section covering blocks $a,\ldots,b$, conditional on the event $\{W^{(a)} = \cdots = W^{(b)}\}$, the common value is Bernoulli with probability \[ p=\frac{\prod_{k=a}^{b} q_k}{\prod_{k=a}^{b} q_k + \prod_{k=a}^{b}(1-q_k)}. \] Moreover, across disjoint sections these conditional draws are independent.
theorem[Finite-sample validity] Suppose the assignment follows Definition (ref) and Assumptions (ref) and (ref) hold. Consider testing the partially sharp null (ref). Let $\mathcal{C}(\mathbf{W}) := (\mathcal{S}(\mathbf{W}),\mathcal{U}(\mathbf{W}))$ denote the conditioning event. Let $\hat p$ be the Monte Carlo $p$-value returned by Algorithm (ref). Then, under $H_0$, \[ \mathbb{P}\!\left(\hat p \le \alpha \,\middle|\, \mathcal{C}(\mathbf{W})=\mathcal{C}(\mathbf{W}^{\mathrm{obs}})\right) \le \alpha \quad \text{for all } \alpha\in(0,1). \]

Theorem (ref) is a finite-population statement: potential outcomes are treated as fixed, and randomness enters only through the switchback assignment mechanism and the Monte Carlo resampling. The result follows from (i) imputability of focal outcomes under $H^{tot}_0$ given the conditioning event and (ii) correct simulation of the conditional assignment law via Lemma (ref) Basse2019,Athey2018.

remark[Predetermined sections and invariance] Theorem (ref) relies on the invariance requirement for degenerate CRTs puelz2021: the conditioning event must be the same for every focal assignment considered by the test. This is why we fix a predetermined collection of disjoint candidate sections $\mathcal S^{\mathrm{pre}}$ independent of $\mathbf W$, and then define $\mathcal S(\mathbf W)$ only by selecting those predetermined sections that are constant under $\mathbf W$. Because section boundaries are fixed ex ante, any $\mathbf w\in\mathcal{W}(\mathcal{S}(\mathbf W^{\mathrm{obs}}))$ preserves the same selected sections (and hence the same focal set), so $\mathcal C(\mathbf w)=\mathcal C(\mathbf W^{\mathrm{obs}})$. If sections were defined adaptively from $\mathbf W$ (e.g., via maximal runs of all-ones and all-zeros paths), resampling could create longer consecutive all-ones/all-zeros paths, changing $\mathcal C$ and breaking invariance.
remark[Permutation implementations under homogeneous section probabilities] Algorithm (ref) samples section labels using their conditional treatment probabilities $\{p_j\}$. If the design is homogeneous in the sense that $q_k\equiv q$ and every realized section in $\mathcal{S}(\mathbf{W}^{\mathrm{obs}})$ merges the same number of design blocks (equivalently, has the same length in block units), then $p_j\equiv p$ for all selected sections. In this case the selected section labels are exchangeable, and one may implement the CRT via a simple permutation: condition additionally on the treated count $\sum_{j\in\mathcal{J}} Z_j$ and permute the observed labels across the selected sections. When the section probabilities are not all equal, exchangeability fails. A convenient exact alternative is stratified permutation: partition the selected sections into strata with common $p_j$, condition on the treated count within each stratum, and permute labels only within strata. This retains the simplicity of permutations while respecting heterogeneous assignment probabilities.

Testing the carryover horizon

The CRT in Section (ref) assumes a known carryover horizon $m$. In applications, however, $m$ is rarely known a priori, and it is often important to assess whether the carryover window used to define focal outcomes is plausibly long enough. This section develops randomization tests for the null hypothesis that carryover effects vanish after $m$ periods.\footnote{Bojinov2023 also propose a test for identifying the order of carryover effects. Their approach, however, requires a specialized design implemented across different units—thus necessitating changes to the experimental protocol—and its validity is primarily asymptotic. In contrast, our method applies directly to existing experimental designs and yields finite‑sample validity, which is particularly important when sample sizes are modest.}

definition[Null of $m$-carryover effects $H_0^m$] Fix $m \in \{0,1,\dots,T-1\}$ and assume no-anticipation as in Assumption (ref). The null hypothesis of at most $m$-period carryover effects, denoted $H_0^m$, states that for any $t \in \{m+1,\dots,T\}$, any continuation path $\mathbf{w}_{t-m:t}\in\{0,1\}^{m+1}$, and any two histories $\mathbf{w}_{1:t-m-1}^\prime,\mathbf{w}_{1:t-m-1}^{\prime\prime}\in\{0,1\}^{t-m-1}$, \[ Y_t(\mathbf{w}_{1:t-m-1}^\prime,\mathbf{w}_{t-m:t}) = Y_t(\mathbf{w}_{1:t-m-1}^{\prime\prime},\mathbf{w}_{t-m:t}). \]

Definition (ref) makes precise the intuition that, under $H_0^m$, outcomes at time $t$ are unaffected by assignments more than $m$ periods in the past.

We reuse the predetermined section family $\mathcal{S}^{\text{pre}}=\{[s_j,e_j]\}_{j=1}^J$ from Section (ref) and index sections in increasing time order. To decouple the randomized labels from the outcomes used for testing, we employ an “alternating holdout” device: we hold out every other predetermined section and use the remaining sections as focal outcome sections.

Define the focal sections as $\{[s_{2i},e_{2i}]\}_{i=1}^{\lfloor J/2\rfloor}$. Within each focal section, we keep exactly those times whose last $m$ lags remain in the same section:

equation[equation omitted — 136 chars of source]

Under $H_0^m$, outcomes $\{Y_t(\mathbf{w}): t\in\mathcal{U}^{(m)}\}$ depend only on assignments inside their own focal section, and hence they are invariant to changes in assignment outside that section.

lemma[Local dependence of $Y_t$ under $H_0^m$] If $t \in [s_{2i}+m,\, e_{2i}]$, then under $H_0^m$ the quantity $Y_t(\mathbf{w})$ depends only on the assignments $\mathbf{w}_{t-m:t}$, and in particular \[ \mathbf{w}_{t-m:t} \subseteq [s_{2i},\, e_{2i}]. \] Therefore, changing assignments outside $[s_{2i},\, e_{2i}]$ cannot change $Y_t$.

Let $\mathcal{W}_0$ be the blockwise-constant assignment set in (ref). Conditioning on $\mathcal{S}^{\text{pre}}$ (which is predetermined), we define the randomization space by keeping the assignment fixed on focal sections while allowing the remaining blocks to vary according to the design:

equation[equation omitted — 199 chars of source]

A natural test statistic aggregates focal outcomes section-by-section and treats the treatment levels of the preceding (non-focal) sections as randomized labels. For each even-numbered section $2i$, define the focal mean \[ \bar Y_{2i}^{\mathrm{obs}} = \frac{1}{|\mathcal{U}_{2i}^{(m)}|} \sum_{t=s_{2i}+m}^{e_{2i}} Y_t^{\mathrm{obs}}, \qquad \mathcal{U}_{2i}^{(m)}=\{t : s_{2i}+m \le t \le e_{2i}\}, \] which averages observed outcomes over times whose last $m$ lags remain within the focal section. For the preceding non-focal section $2i-1$, define its section label as the treatment assignment at its final time point, \[ Z_{2i-1} := W_{e_{2i-1}}. \] Because treatment is blockwise constant within each section, $Z_{2i-1}$ indexes the common treatment level applied throughout the entire preceding section $[s_{2i-1},e_{2i-1}]$.

When the assignment probabilities are known from the design, we form an inverse-probability-weighted contrast across adjacent section pairs: \[ T_m(W) = \frac{1}{\lfloor J/2 \rfloor} \sum_{i=1}^{\lfloor J/2 \rfloor} \left( \frac{Z_{2i-1}\,\bar Y_{2i}^{\mathrm{obs}}}{\Pr(Z_{2i-1}=1)} - \frac{(1-Z_{2i-1})\,\bar Y_{2i}^{\mathrm{obs}}}{1-\Pr(Z_{2i-1}=1)} \right). \]

Under $H_0^m$, focal outcomes depend only on assignments within their own section, and hence are independent of the treatment level in the preceding section under the randomization distribution. The statistic $T_m(W)$ therefore tests for residual dependence of focal outcomes on treatment exposure in adjacent prior sections. The corresponding randomization distribution is obtained by resampling the non-focal section labels in $\mathcal{W}^{(m)}$ while holding the focal outcomes $\{Y_t^{\mathrm{obs}} : t \in \mathcal{U}^{(m)}\}$ fixed. \footnote{Any statistic measurable with respect to focal outcomes and the randomized labels, and imputable under Lemma (ref), yields a valid conditioning-based randomization test.}

We now make explicit how to obtain the $p$-value for testing a given $H_0^m$. Fix $m$ and treat the focal set $\mathcal{U}^{(m)}$ in (ref) and the conditional assignment space $\mathcal{W}^{(m)}$ in (ref) as the conditioning event. Under $H_0^m$ and Lemma (ref), the focal outcomes $\{Y_t^{\mathrm{obs}}:t\in\mathcal{U}^{(m)}\}$ are invariant to re-randomizing assignments on the non-focal sections, so any statistic that depends on the observed outcomes only through $\mathcal{U}^{(m)}$ and on assignments only through the randomized labels is imputable. Algorithm (ref) gives a Monte Carlo CRT that samples from the conditional assignment law on $\mathcal{W}^{(m)}$ and returns a one-sided $p$-value.

algorithm[algorithm omitted — 1,080 chars of source]

The validity of Algorithm (ref) follows directly from the same argument as in Theorem (ref). In particular, for each fixed \(m\), it yields a valid level-\(\alpha\) test of \(H_0^m\). In many applications, however, \(m\) is unknown, and the objective is to identify a plausible minimal horizon \(\bar m\) beyond which carryover effects vanish. Because the hypotheses $\{H_0^m\}_{m=0}^M$ are nested (Proposition (ref)), false nulls must occur before true nulls. This monotonic structure allows us to run the tests sequentially using the $p$-values $\{\widehat p_m\}$ from Algorithm (ref) and stop at the first non-rejection, as formalized in Algorithm (ref). Theorem (ref) then guarantees FWER control (Definition (ref)) without multiplicity corrections.

proposition[Nestedness] Suppose there exists $\bar m \in \{0,\dots,M\}$ such that $H_0^{\bar m}$ is true. Then $H_0^{m}$ is true for all $m \ge \bar m$.
definition[FWER under nested nulls] Let $\varphi = (\varphi_0,\dots,\varphi_M)$ be a multiple testing rule, where $\varphi_m(\mathbf{W}^{\mathrm{obs}})=1$ denotes rejection of $H_0^m$. Let $\bar m$ denote the smallest index such that $H_0^{\bar m}$ is true. The family-wise error rate (FWER) is \[ \mathrm{FWER} = \Pr\bigl(\exists\, m \ge \bar m \text{ such that } \varphi_m(\mathbf{W}^{\mathrm{obs}})=1 \bigr), \] the probability of rejecting at least one true null hypothesis.
algorithm[algorithm omitted — 675 chars of source]
theoremAlgorithm (ref) controls the family-wise error rate at level $\alpha$.

Algorithm (ref) effectively provides a lower bound (or conservative estimate) for the carryover horizon $m$, rather than point identification. That being said, it remains useful in practice. For example, Bojinov2023 note that when $m$ is unknown, it is generally preferable to choose a value slightly larger than the true $m$ rather than substantially smaller. This suggests that even a conservative estimate or lower bound on $m$ can be informative for experimental design and inference, since underestimating the carryover horizon can lead to more serious distortions than modest overestimation.

Testing non-anticipation

Sections (ref)--(ref) treat non-anticipation as a maintained assumption. Here we construct a finite-sample valid test of this restriction. Because anticipation concerns dependence on future assignments, constructing convenient conditioning events for conditional randomization tests is difficult without strong structural restrictions. We therefore employ an unconditional randomization test based on the Pairwise Imputation--based Randomization Test (PIRT) of zhong2024unconditional, adapted to switchback schedules.

\paragraph{Imputable times.}

Under Assumption (ref), potential outcomes at time $t$ depend only on the assignment prefix $\mathbf w_{1:t}$. Thus if two schedules agree through time $t$, the corresponding potential outcomes must coincide.

definition[Imputable times] For schedules $\mathbf w,\mathbf w'\in\{0,1\}^T$, define \[ \mathbb I(\mathbf w,\mathbf w') = \bigl\{t:\ \mathbf w_{1:t}=\mathbf w'_{1:t}\bigr\}. \]

Times in $\mathbb I(\mathbf w,\mathbf w')$ are precisely those at which the non-anticipation null implies equality of potential outcomes across the two schedules and hence support imputation-based inference.

\paragraph{Prefix-preserving randomization.}

Independently drawn switchback schedules typically diverge quickly, producing few imputable times. To ensure adequate overlap, we restrict randomization to schedules that share an initial prefix with the observed assignment.

Fix $L\in\{0,\ldots,T\}$. Given $\mathbf W^{\mathrm{obs}}$, alternative schedules $\mathbf W^\ast$ are drawn from the design distribution subject to \[ \mathbf W^\ast_{1:L}=\mathbf W^{\mathrm{obs}}_{1:L}. \] Assignments after time $L$ follow the original design mechanism. This guarantees $\{1,\ldots,L\}\subseteq \mathbb I(\mathbf W^{\mathrm{obs}},\mathbf W^\ast)$ for every draw.

\paragraph{Centered signed-score statistic.}

Let $s_t(\mathbf w)=2w_{t+1}-1$ for $t<T$, and define \[ I(\mathbf w',\mathbf w)=\mathbb I(\mathbf w',\mathbf w)\cap\{1,\ldots,T-1\}. \] Given observed outcomes $\mathbf Y^{\mathrm{obs}}=\mathbf Y(\mathbf W^{\mathrm{obs}})$, define

equation[equation omitted — 282 chars of source]

where $\bar Y_A^{\mathrm{obs}}=|A|^{-1}\sum_{t\in A}Y_t^{\mathrm{obs}}$, and $T_{\mathrm{NA}}=+\infty$ if $I(\mathbf w',\mathbf w)=\varnothing$.

This statistic depends only on outcomes at imputable times and is therefore pairwise imputable under the non-anticipation null. Centering removes sensitivity to the overall level of outcomes on the imputable set and ensures the statistic is driven by association between outcomes and the next-period assignment sign. Under anticipation alternatives where outcomes increase with future treatment, this association tends to be positive, yielding power while preserving finite-sample validity.

\paragraph{PIRT procedure.}

Let $\mathbf W^\ast$ be drawn from the prefix-preserving design. Define the pairwise statistic \[ A=T_{\mathrm{NA}}(\mathbf Y^{\mathrm{obs}},\mathbf W^\ast,\mathbf W^{\mathrm{obs}}), \qquad B=T_{\mathrm{NA}}(\mathbf Y^{\mathrm{obs}},\mathbf W^{\mathrm{obs}},\mathbf W^\ast). \]

The PIRT $p$-value is \[ p = \mathbb P\!\left(A\ge B\right), \] with probability taken over the prefix-preserving randomization distribution. Two-sided variants use $|A|\ge|B|$.

\paragraph{Finite-sample validity.}

Under non-anticipation, observed outcomes coincide with the corresponding potential outcomes at all imputable times for every pair $(\mathbf W^{\mathrm{obs}},\mathbf W^\ast)$. Because the statistic depends only on imputable outcomes and the randomization distribution is known, the resulting test controls size exactly in finite samples under the conditional design. See zhong2024unconditional for general validity results.

Testing Weak-null Hypotheses

From this section onward, we specialize to switchback designs in which the treatment is randomized over equal-length, naturally defined design blocks (e.g., days or weeks); we refer to these blocks as sessions\footnote{We use blocks to denote the (possibly unequal-length) design intervals induced by the switch times. In contrast, sessions are equal-length, naturally defined blocks (used only in this section), while sections (or pooled sections) are predetermined unions of consecutive blocks used for conditioning in our CRTs.}. This restriction is motivated by common practice and by interpretability: the weak-null hypotheses we consider are formulated position-by-position within a session (e.g., Monday vs.\ Tuesday), which is meaningful only when the design blocks correspond to a stable time unit.

The goal of this section is to clarify which weak (mean-zero) null hypotheses are testable using randomization-based inference in switchback experiments with carryover. Because weak nulls are not sharp, finite-sample exact CRTs are generally unavailable; instead, we use studentized CRTs to obtain asymptotic size control chung2013,Ding2017,DiCiccio&Romano2017,ZHAO2021278,Wu2021. We show that studentized CRTs can test: (i) a joint weak null requiring mean-zero effects at each within-session focal position, (ii) a weak null requiring mean-zero average effects over the focal periods at the end of each session, and (iii) position-wise nulls for any individual focal position. We also highlight an important limitation: without further structure, we generally cannot test a global weak null that averages treatment effects over the entire time series.

Weak-null hypotheses under fixed-length sessions

Fix a session length $L$ and a carryover length $m$ with $L>m$, and assume for simplicity that $T=JL$ for some integer $J\ge 1$. Sessions are deterministic intervals \[ [s_j,e_j]:=\{(j-1)L+1,\dots,jL\}, \qquad j=1,\dots,J, \] and the switchback design assigns a constant treatment label $Z_j\in\{0,1\}$ to each session $j$ with known probability $p_j:=\Pr(Z_j=1)\in(0,1)$. Under the maintained non-anticipation and $m$-carryover restrictions, periods near the start of a session may depend on the previous session’s assignment. Accordingly, we focus on the within-session focal window at the end of each session, \[ \mathcal U_j := \{t\in[T]: s_j+m \le t \le e_j\}, \qquad n:=|\mathcal U_j|=L-m, \] and the overall focal set $\mathcal U:=\cup_{j=1}^J \mathcal U_j$. Index focal times in session $j$ by $\ell=1,\dots,n$ via $t_{j,\ell}:=s_j+m+\ell-1$.

For $z\in\{0,1\}$, define the focal potential outcome at within-session position $\ell$ under the constant path $\mathbf z_T$ by \[ Y_{j,\ell}(z):=Y_{t_{j,\ell}}(\mathbf{z}_T),\qquad z\in\{0,1\}. \] Under $m$-carryover and $L>m$, $Y_{j,\ell}(z)$ can be interpreted as the outcome at the $\ell$th focal position in session $j$ when the session is assigned $z$ (the relevant $(m+1)$-period treatment history lies entirely within the session’s constant segment).

Define the cross-session average effect at within-session focal position $\ell$ as \[ \tau_\ell := \frac{1}{J}\sum_{j=1}^J\{Y_{j,\ell}(1)-Y_{j,\ell}(0)\}, \qquad \ell=1,\dots,n. \] We consider three empirically relevant weak-null hypotheses:

\paragraph{(1) Joint session-wise weak null (strong).} The strongest focal weak null requires mean-zero effects at each within-session focal position:

equation[equation omitted — 112 chars of source]

This null is attractive in practice when within-session seasonality is important (e.g., day-of-week effects), because it rules out cancellation across positions.

\paragraph{(2) Focal-average weak null (weaker).} Let $\bar Y_j(z):=n^{-1}\sum_{\ell=1}^n Y_{j,\ell}(z)$ denote the session-level mean potential outcome over the focal window, and define the corresponding focal-average effect \[ \tau_{\mathcal U} := \frac{1}{J}\sum_{j=1}^J\{\bar Y_j(1)-\bar Y_j(0)\}. \] We will also test the weaker null

equation[equation omitted — 85 chars of source]

Since $\tau_{\mathcal U}=n^{-1}\sum_{\ell=1}^n \tau_\ell$, the strong joint null $H_0^{\mathrm{sw}}$ implies $H_0^{\mathcal U}$, but not conversely. The null $H_0^{\mathcal U}$ is directly aligned with the focal set $\mathcal U$ used by our CRTs and is therefore a natural target for studentized randomization inference.

\paragraph{(3) Position-wise weak nulls.} For any focal position $\ell\in\{1,\dots,n\}$, we can test \[ H_{0,\ell}:\quad \tau_\ell=0, \] and, more generally, test subsets of positions by restricting attention to $\{\tau_\ell:\ell\in\mathcal L\}$ for any $\mathcal L\subseteq\{1,\dots,n\}$.

\paragraph{Why the global weak null is generally not testable.} A common estimand under constant paths is the global post--burn-in average effect \[ \bar Y_m(\mathbf{w}) := \frac{1}{T-m}\sum_{t=m+1}^T Y_t(\mathbf{w}), \qquad \tau_m^{\mathrm{glob}}:=\bar Y_m(\mathbf{1}_T)-\bar Y_m(\mathbf{0}_T), \] and the corresponding global null $H_0^{\mathrm{glob}}:\tau_m^{\mathrm{glob}}=0$. In switchback designs with carryover, however, valid randomization-based inference is naturally tied to focal periods whose relevant treatment histories are contained within constant segments. A studentized CRT built on the focal set $\mathcal U$ targets $\tau_{\mathcal U}$; under the global null $\tau_m^{\mathrm{glob}}=0$, the focal-average effect $\tau_{\mathcal U}$ need not be zero (e.g., under within-session heterogeneity or sign-reversing patterns), so the randomization distribution of a focal statistic is not generally centered at zero and uniform validity cannot be guaranteed without additional structure linking focal and non-focal periods. That said, when the outcome process is approximately stationary across within-session positions and $m$ is small relative to $L$ (so $n=L-m$ is close to $L$), $\tau_{\mathcal U}$ can be close to the global average effect, making $H_0^{\mathcal U}$ a practically informative proxy in many applications.

A studentized CRT for the focal-average null

Let $Z_j^{\mathrm{obs}}\in\{0,1\}$ denote the realized session assignment and $p_j:=\Pr(Z_j=1)$ its known design probability. Define the observed session-level focal mean \[ \bar Y_j^{\mathrm{obs}}:=\frac{1}{n}\sum_{t\in\mathcal U_j}Y_t^{\mathrm{obs}}. \] Under the fixed-length session design and $m$-carryover, $\bar Y_j^{\mathrm{obs}}=\bar Y_j(Z_j^{\mathrm{obs}})$ for each $j$, so $\{\bar Y_j(1),\bar Y_j(0)\}_{j=1}^J$ are well-defined session-level potential outcomes for the focal window. We estimate the focal-average effect $\tau_{\mathcal U}$ using the Horvitz--Thompson (HT) estimator

equation[equation omitted — 222 chars of source]

and studentize using the conservative upper-bound form

equation[equation omitted — 293 chars of source]

The studentized CRT compares $T_{\mathrm{stud}}^{\mathrm{obs}}$ to its randomization distribution obtained by resampling $Z_1^\ast,\dots,Z_J^\ast$ independently with $Z_j^\ast\sim\mathrm{Bernoulli}(p_j)$ and recomputing $T_{\mathrm{stud}}^\ast$ while holding $\{\bar Y_j^{\mathrm{obs}}\}$ fixed. The resulting one-sided Monte Carlo $p$-value is denoted $\hat p_{\mathrm w}$. (Here, the session partition and focal set are deterministic, so “conditioning” reduces to conditioning on a deterministic event.) This procedure can be viewed as the fixed-length-session analogue of the CRT in Section (ref), with studentization added to handle weak nulls.

assumption[Weak-null asymptotics under fixed-length sessions] The carryover length $m$ and session length $L$ are fixed with $L>m$, and $T=JL\to\infty$ so that $J\to\infty$. There exists $\underline p\in(0,1/2)$ such that $\underline p \le p_j \le 1-\underline p$ for all $j$ and all $T$. Let $\bar Y_j(1)$ and $\bar Y_j(0)$ denote the session-level focal mean potential outcomes above and define $M_j:=\max\{|\bar Y_j(1)|,|\bar Y_j(0)|\}$. Assume: \begin{enumerate} • (Uniform fourth-moment bound) There exists $C<\infty$ such that $\sup_{J\ge 1} \frac{1}{J}\sum_{j=1}^J M_j^4 \le C$. • (Stabilization) The empirical measures $\nu_J := \frac{1}{J}\sum_{j=1}^J \delta_{(p_j,\bar Y_j(1),\bar Y_j(0))}$ converge weakly to some probability measure $\nu$ on $[\underline p,1-\underline p]\times\mathbb{R}^2$. • (Nondegeneracy) If $(P,Y_1,Y_0)\sim\nu$, then \[ \sigma_{\mathrm S}^2 := \mathbb{E}_\nu\!\left[ \left( \sqrt{\frac{1-P}{P}}\,Y_1 + \sqrt{\frac{P}{1-P}}\,Y_0 \right)^2 \right] >0. \] \end{enumerate}

Assumption (ref)(ii) is a standard triangular-array stabilization condition in randomization CLTs: it requires that as $J\to\infty$, the empirical distribution of session-level potential outcomes (and assignment probabilities) converges to a stable limiting “population.” This assumption is weaker than independence or stationarity of the outcome process; it formalizes the idea that the experimental environment does not drift arbitrarily as more sessions are observed. It rules out, for example, systematic time trends in treatment effects or assignment probabilities, progressively heavier-tailed session means, or structural regime changes that make early and late sessions incomparable.

theorem[Asymptotic validity for the focal-average weak null] Suppose Definition (ref) holds with equal-length sessions of length $L>m$, and Assumptions (ref) and (ref) hold. If Assumption (ref) holds, then under $H_0^{\mathcal U}$ in (ref), \[ \limsup_{T\to\infty} \mathbb{P}\!\left(\hat p_{\mathrm w} \le \alpha \right) \le \alpha \qquad \text{for all } \alpha\in(0,1), \] where $\hat p_{\mathrm w}$ is the Monte Carlo $p$-value computed from the randomization distribution of $T_{\mathrm{stud}}$ in (ref). In particular, the same test is asymptotically valid under the stronger joint null $H_0^{\mathrm{sw}}$ in (ref).

Position-wise and joint tests across within-session positions

The statistic $T_{\mathrm{stud}}$ targets the focal-average effect $\tau_{\mathcal U}$ and can therefore have low power against alternatives where effects vary across within-session positions and cancel in the average (e.g., sign-reversing patterns). To probe heterogeneity and to test the strong joint null $H_0^{\mathrm{sw}}$ more directly, it is natural to use position-wise and joint tests.

\paragraph{Position-wise tests.} Fix $\ell\in\{1,\dots,n\}$ and let $Y_{j,\ell}^{\mathrm{obs}}:=Y_{j,\ell}(Z_j^{\mathrm{obs}})$ denote the observed outcome at focal position $\ell$ in session $j$. Consider the HT estimator \[ \hat\tau_{\ell,\mathrm{HT}} := \frac{1}{J}\sum_{j=1}^J \left( \frac{Z_j^{\mathrm{obs}}\,Y_{j,\ell}^{\mathrm{obs}}}{p_j} - \frac{(1-Z_j^{\mathrm{obs}})\,Y_{j,\ell}^{\mathrm{obs}}}{1-p_j} \right), \] and the upper-bound variance estimator \[ \widehat V_{\ell,\mathrm{up}} := \frac{1}{J^2}\sum_{j=1}^J (Y_{j,\ell}^{\mathrm{obs}})^2 \left( \frac{Z_j^{\mathrm{obs}}}{p_j^2} + \frac{1-Z_j^{\mathrm{obs}}}{(1-p_j)^2} \right), \qquad T_{\ell}:=\frac{\hat\tau_{\ell,\mathrm{HT}}}{\sqrt{\widehat V_{\ell,\mathrm{up}}}}. \] A studentized CRT for the position-wise null $H_{0,\ell}:\tau_\ell=0$ is obtained by the same resampling scheme: resample $Z_1^\ast,\dots,Z_J^\ast$ independently with $Z_j^\ast\sim\mathrm{Bernoulli}(p_j)$, hold $\{Y_{j,\ell}^{\mathrm{obs}}\}_{j=1}^J$ fixed, and recompute $T_\ell^\ast$. The resulting Monte Carlo $p$-value is asymptotically valid under the same conditions as Theorem (ref). The same construction applies to any subset of focal positions by restricting $\ell$ to a set $\mathcal L$.

\paragraph{Joint tests via quadratic-form ($F$-type) statistics.} To test the strong joint null $H_0^{\mathrm{sw}}$ (or, more generally, $\tau_\ell=0$ for all $\ell\in\mathcal L$ for a subset $\mathcal L$), one can combine the vector of position-wise effect estimators using an $F$-type quadratic form. Let $\mathbf Y_j^{\mathrm{obs}}:=(Y_{j,1}^{\mathrm{obs}},\dots,Y_{j,n}^{\mathrm{obs}})^\top$ and define the vector HT estimator $\widehat{\boldsymbol\tau}_{\mathrm{HT}}:=(\hat\tau_{1,\mathrm{HT}},\dots,\hat\tau_{n,\mathrm{HT}})^\top$. A natural matrix analogue of (ref) is \[ \widehat\Sigma_{\mathrm{up}} := \frac{1}{J^2}\sum_{j=1}^J \mathbf Y_j^{\mathrm{obs}}(\mathbf Y_j^{\mathrm{obs}})^\top \left( \frac{Z_j^{\mathrm{obs}}}{p_j^2} + \frac{1-Z_j^{\mathrm{obs}}}{(1-p_j)^2} \right), \] and an omnibus statistic is \[ T_F := \widehat{\boldsymbol\tau}_{\mathrm{HT}}^\top \,\widehat\Sigma_{\mathrm{up}}^{-1}\,\widehat{\boldsymbol\tau}_{\mathrm{HT}}, \] with the convention that a generalized inverse may be used when needed (e.g., when $n$ is large relative to $J$). A randomization $p$-value is obtained by recomputing $T_F^\ast$ under each resampled assignment vector $\mathbf Z^\ast=(Z_1^\ast,\dots,Z_J^\ast)$. Such quadratic-form tests are standard in randomization inference for multivariate outcomes; see, e.g., Dasgupta2017. These joint tests are typically much more sensitive than focal averages to structured or sign-reversing alternatives.

Formal asymptotic validity results for the position-wise and joint studentized CRTs described above are stated and proved in Appendix (ref).

Power analysis under a superpopulation model

Sections (ref) and (ref) establish validity of our randomization tests under a finite-population framework. This section studies power under a superpopulation model for the potential-outcome time series. The goal is to obtain interpretable approximations that clarify how block length and predetermined pooling enter the signal-to-noise ratio of the studentized CRTs.

Similarly to Section (ref), we consider a simple setup where the design blocks have a fixed length and are longer than the carryover memory length $m$.\footnote{In this section, we use the term “block” for fixed-length design intervals rather than “sessions,” since blocks need not correspond to naturally occurring time units (e.g., days or weeks) to support meaningful weak null hypotheses as in Section (ref).} Formally, fix an experiment length $T$ and a block length $L\ge 1$ such that $T=ML$ for some integer $M$. Define design blocks \[ B_k := \{(k-1)L+1,\dots,kL\},\qquad k=1,\dots,M. \] Assume a regular switchback design with constant assignment probability $q\in(0,1)$:

equation[equation omitted — 137 chars of source]

Fix a pooling size $r\ge 1$ and assume $r\mid M$. Set $J:=M/r$ and define predetermined pooled sections \[ [s_j,e_j] := \bigcup_{k=(j-1)r+1}^{jr} B_k = \{(j-1)rL+1,\dots,jrL\}, \qquad j=1,\dots,J, \] each of length $rL$. Let $E_j$ denote the event that pooled section $j$ is blockwise constant: \[ E_j=\{W^{((j-1)r+1)}=\cdots=W^{(jr)}\}. \] On $E_j$, let $Z_j$ denote the common block value.

lemma[Conditional pooled-section law under predetermined pooling] Under (ref), for each $j$, \[ \mathbb{P}(Z_j=1\mid E_j) = p(r;q) := \frac{q^r}{q^r+(1-q)^r}. \] Moreover, across disjoint sections the pairs $\{(E_j,Z_j)\}$ are independent, and hence conditional on $\{E_j: j\in\mathcal{J}(\mathbf{W})\}$ the labels $\{Z_j: j\in\mathcal{J}(\mathbf{W})\}$ are independent $\mathrm{Bernoulli}(p(r;q))$.

Fix a burn-in parameter $m\ge 0$ and assume $rL>m$. Each pooled section contributes

equation[equation omitted — 39 chars of source]

focal time points (the last $n$ periods in the pooled section).

We posit the distributed-lag model

equation[equation omitted — 134 chars of source]

where $m_0\ge 0$ is the true carryover horizon, $\mu\in\mathbb{R}$ is a baseline level, and $\{\varepsilon_t\}$ is a stationary mean-zero process independent of the assignment mechanism. For closed-form expressions, assume AR(1) errors:

equation[equation omitted — 140 chars of source]

with $\sigma_\varepsilon^2:=\sigma_u^2/(1-\rho^2)$.

Under constant paths, the total long-run effect is \[ \tau_{\mathrm{tot}} := \mathbb{E}\!\left[Y_t(\mathbf{1}_T)-Y_t(\mathbf{0}_T)\right] = \sum_{\ell=0}^{m_0}\beta_\ell. \] The carryover null $H_0^m$ corresponds to $\beta_\ell=0$ for all $\ell>m$.

remark[Asymptotic regime] Throughout, we consider $M\to\infty$ with $(L,r,m,m_0,q)$ fixed, so $T=ML\to\infty$ and $J=M/r\to\infty$. For the total-effect test, the number of usable (constant) pooled sections $J_{\mathrm{tot}}:=|\mathcal J(\mathbf W)|$ is random, where $\mathcal J(\mathbf W):=\{j:\,E_j\}$; under (ref), $J_{\mathrm{tot}}/J\to \pi_r(q)$ with $\pi_r(q)=q^r+(1-q)^r$, and hence $J_{\mathrm{tot}}\to\infty$ with high probability. In the proof, we first derive a finite-population CLT conditional on the realized potential outcomes, then use the DGP only to approximate the resulting random variance functionals and obtain an unconditional normal approximation. This is not a superpopulation CLT: the estimand remains sample-dependent, and the DGP is invoked solely to make variance and design/power trade-offs transparent—not to redefine the target of inference.

\paragraph{Total-effect statistic.} Conditional on the realized set of usable pooled sections, let $J_{\mathrm{tot}}=|\mathcal{J}(\mathbf{W})|$ denote their number and write $Z_j\in\{0,1\}$ for the (common) pooled-section assignment label on each $j\in\mathcal J(\mathbf W)$. Let $\bar Y_j^{\mathrm{obs}}$ denote the mean outcome over the $n$ focal periods in pooled section $j$.

With constant $p=p(r;q)$ and equal $n$, the pooled-section HT estimator equals

equation[equation omitted — 221 chars of source]

and we studentize using

equation[equation omitted — 299 chars of source]

\paragraph{$m$-carryover statistic (paired predetermined pooled sections).} Following the carryover test in Section (ref), we use an alternating holdout: even pooled sections provide outcomes, and the immediately preceding odd pooled sections provide randomized labels. Let $J_e:=\lfloor J/2\rfloor$ and consider the $J_e$ adjacent pairs $\big([s_{2j-1},e_{2j-1}],[s_{2j},e_{2j}]\big)$ for $j=1,\dots,J_e$. For each even pooled section $2j$, define the focal mean \[ \bar Y_{2j}^{\mathrm{obs}} := \frac{1}{n}\sum_{t=s_{2j}+m}^{e_{2j}} Y_t^{\mathrm{obs}}, \qquad n:=rL-m. \]

Define the randomized label from the preceding odd pooled section as the assignment on its last block: \[ Z_{2j-1} := W^{((2j-1)r)} \in\{0,1\}. \] Under (ref), $\Pr(Z_{2j-1}=1)=q$ and $\{Z_{2j-1}\}_{j=1}^{J_e}$ are independent.

We estimate the tail carryover signal using the HT contrast

equation[equation omitted — 203 chars of source]

and studentize with the corresponding upper-bound form

equation[equation omitted — 280 chars of source]

We first record the variance of the average of $n$ consecutive AR(1) errors:

equation[equation omitted — 220 chars of source]
proposition[Power for the total-effect test] Assume (ref), the predetermined pooled-section construction, and the DGP (ref)--(ref). Assume the analyst burn-in satisfies $m\ge m_0$ so that focal means are uncontaminated by carryover. Let $\varphi_{\mathrm{tot}}$ be the one-sided level-$\alpha$ CRT that rejects for large $T_{\mathrm{tot}}$ in (ref). Then, conditional on $J_{\mathrm{tot}}$, \begin{equation} \mathbb{P}\!\left(\varphi_{\mathrm{tot}}=1 \,\middle|\, J_{\mathrm{tot}}\right) \;\approx\; 1-\Phi\!\left(\frac{z_{1-\alpha}-\mu_{\mathrm{tot}}(J_{\mathrm{tot}})}{\sigma_{\mathrm{tot}}}\right), \end{equation} where $p=p(r;q)$, $n=rL-m$, and \begin{equation} \mu_{\mathrm{tot}}(J_{\mathrm{tot}}) := \frac{\tau_{\mathrm{tot}}\sqrt{J_{\mathrm{tot}}}} {\sqrt{\ \frac{\sigma_{\bar\varepsilon}^2(n,\rho)}{p(1-p)}+\frac{(\mu+\tau_{\mathrm{tot}})^2}{p}+\frac{\mu^2}{1-p}\ }}, \qquad \sigma_{\mathrm{tot}}^2 := \frac{\sigma_{\bar\varepsilon}^2(n,\rho)+\{\mu+(1-p)\tau_{\mathrm{tot}}\}^2} {\sigma_{\bar\varepsilon}^2(n,\rho)+p\mu^2+(1-p)(\mu+\tau_{\mathrm{tot}})^2}. \end{equation} Moreover, $0<\sigma_{\mathrm{tot}}^2\le 1$ and $\sigma_{\mathrm{tot}}^2\to 1$ under local alternatives $\tau_{\mathrm{tot}}=o(1)$.
proposition[Power for the $m$-carryover test] Assume (ref), the predetermined pooled-section construction, and the DGP (ref)--(ref). Assume $m_0>m$ and \begin{equation} m_0\le rL \end{equation} so that, after the $m$-period burn-in, any lag leaving an even pooled section can reach only into the immediately preceding odd pooled section. For the carryover test statistic (ref)--(ref), only the assignment on the last design block of the preceding odd pooled section is used as the randomized label. Define the corresponding tail-signal functional \begin{equation} \delta_m(n) := \frac{1}{n}\sum_{\ell=m+1}^{m_0}\min\{L,\ell-m\}\,\beta_\ell, \qquad n=rL-m, \end{equation} and let $\varphi_m$ be the one-sided level-$\alpha$ CRT that rejects for large $T_m$ in (ref). Then, as $J\to\infty$, \begin{equation} \mathbb{P}\!\left(\varphi_m=1 \right) \;\approx\; 1-\Phi\!\left(\frac{z_{1-\alpha}-\mu_{m}(J)}{\sigma_{m}}\right), \end{equation} where $J_e=\lfloor J/2\rfloor$ and \begin{equation} \mu_{m}(J) := \frac{\delta_m(n)\sqrt{J_e}}{\sqrt{v_{\mathrm{up}}^{(m)}}}, \qquad \sigma_{m}^2 := \frac{\sigma_{\delta}^2}{v_{\mathrm{up}}^{(m)}}, \end{equation} with $\pi_1=q$, $\pi_0=1-q$, and \begin{align} v_{\mathrm{up}}^{(m)} &:= \frac{\mathbb{E}[\bar Y(1)^2]}{\pi_1}+\frac{\mathbb{E}[\bar Y(0)^2]}{\pi_0}, \\ \sigma_{\delta}^2 &:= \mathbb E\!\left[\left(\frac{1}{\pi_1}-1\right)\bar Y(1)^2+\left(\frac{1}{\pi_0}-1\right)\bar Y(0)^2+2\,\bar Y(1)\bar Y(0)\right]. \end{align} Here $\bar Y(1)$ and $\bar Y(0)$ denote the even-section focal mean under the counterfactual intervention that the last design block of the preceding odd pooled section is set to $1$ versus $0$, respectively (with all other block assignments and the error process following their original laws).

Pooling more blocks (larger $r$) increases the within-section focal sample size $n=rL-m$ and reduces $\sigma_{\bar\varepsilon}^2(n,\rho)$ via within-section averaging, but it also reduces the number of pooled sections $J=M/r$ (and hence the number of odd--even pairs $J_e=\lfloor J/2\rfloor$ used by the carryover diagnostic). For the total-effect test, usable pooled sections must be constant; the selection probability is $\pi_r(q)=q^r+(1-q)^r$, so $J_{\mathrm{tot}}\sim\mathrm{Binomial}(J,\pi_r(q))$ and $\mathbb E[J_{\mathrm{tot}}]=J\pi_r(q)$. For the carryover test based on (ref)--(ref), pooling affects power primarily through the tail-signal functional $\delta_m(n)$ in (ref) and through the variance of the even-section focal mean: each lag coefficient $\beta_\ell$ with $\ell>m$ contributes to $\delta_m(n)$ only through those focal times whose lag-$\ell$ exposure falls in the last block of the preceding odd pooled section, which yields weights capped at $L$ and a $1/n$ dilution from averaging over $n$ focal periods. Propositions (ref) and (ref) quantify how these effects (signal dilution, variance reduction from larger $n$, and fewer pairs $J_e$ when $r$ grows) jointly determine the resulting noncentrality parameter.

Given pilot data (or pre-experiment time series), one can estimate the noise parameters in (ref) (e.g., $\rho$ and $\sigma_\varepsilon^2$, hence $\sigma_{\bar\varepsilon}^2(n,\rho)$) and specify a plausible effect scale (e.g., $\tau_{\mathrm{tot}}$ for the total-effect test or a tail profile $\{\beta_\ell:\ell>m\}$ for the carryover test). For each candidate $r$, compute $n=rL-m$ and $\pi_r(q)$, approximate the distribution of $J_{\mathrm{tot}}\sim\mathrm{Binomial}(J,\pi_r(q))$, and plug these into (ref)--(ref) (with $J_{\mathrm{tot}}$ replaced by $\mathbb E[J_{\mathrm{tot}}]$ or averaged over simulated draws of $J_{\mathrm{tot}}$). A practical constraint is to avoid overly aggressive pooling that yields too few usable sections (or too few informative odd sections), e.g., by requiring $\mathbb E[J_{\mathrm{tot}}]$ and $J_e$ to exceed minimal thresholds for stable studentization and accurate Monte Carlo calibration.

In Propositions (ref)--(ref), experimental design enters power through the inverse-probability factors $1/p$ and $1/(1-p)$ (and, under centered outcomes and local alternatives, essentially through $p(1-p)$ in the noise term), so highly imbalanced designs can substantially inflate the studentization term and reduce the resulting noncentrality parameter. When outcomes are approximately centered (so $\mu\approx 0$) and effect sizes are modest, a balanced assignment $p\approx 1/2$ is therefore typically near-optimal. More generally, practitioners can plug pilot estimates into the power approximations (ref)--(ref) and select $p$ (or, in the pooled-section setup, the underlying block-level fraction $q$, which determines $p=p(r;q)$ and $\mathbb E[J_{\mathrm{tot}}]=J\pi_r(q)$) by numerically maximizing the resulting power proxy subject to operational constraints (e.g., minimum allocation rules).

Simulation

This section evaluates the finite-sample size control and power properties of the proposed testing procedures. We study three classes of hypotheses: (i) the null of no total treatment effect, (ii) the $m$-carryover restriction, and (iii) the no-anticipation restriction. The simulation design is intended to assess both validity and power in time-series environments characterized by temporal dependence and heavy-tailed disturbances.

In this simulation study, we adopt a potential-outcomes framework similar to Bojinov2023, but allow for a richer stochastic structure in the disturbance term in order to stress-test finite-sample validity. Potential outcomes are generated according to

equation[equation omitted — 220 chars of source]

Here $\mu$ is a constant baseline level, $\alpha_t$ is a deterministic time effect, $\delta^{(t+1-s)}$ are non-stochastic carryover coefficients governing the effect of treatment assignments at different lags, $w_s\in\{0,1\}$ denotes treatment assignment at time $s$, and $\varepsilon_t$ is a stochastic disturbance. The indicator term allows the scale of the disturbance to depend on whether treatment has remained constant over the most recent $m+1$ periods, thereby generating treatment-history–dependent heteroskedasticity. Different configurations of the carryover coefficients correspond to different null hypotheses and alternative effect structures.

Throughout, we set $\mu=0$ and $\alpha_t=\log t$. The number of periods is \( T \in \{60,120,180,240,300\}, \) representing experimentally realistic horizons at which asymptotic approximations may be unreliable. We consider two disturbance specifications:

enumerate$\varepsilon_t \overset{i.i.d.}{\sim} \mathcal{N}(0,1)$, • $\varepsilon_t \overset{i.i.d.}{\sim} t(1)$, representing heavy-tailed shocks.

We fix the treatment assignment mechanism to the optimal regular switchback design of Bojinov2023. Switch points occur at times \[ \{1, 2m+1, 3m+1, \ldots,(n-2)m+1\}, \] with switching probability $1/2$ at each decision point. Holding this assignment rule fixed across all configurations allows us to focus on the finite-sample behavior of the inference procedures rather than design choice.

In all simulation experiments, we implement the pre-determined sectioning rule described in Section (ref). Concretely, given the design blocks $\{B_k\}_{k=1}^K$ and carryover horizon $m$, we pool consecutive blocks greedily until the pooled length is at least $m+1$, producing a disjoint cover of $[T]$.

Null of no total treatment effect

To evaluate size and power under $H^{tot}_0$, we fix the carryover horizon at $m=2$ and assume it is correctly specified in the analysis as in Bojinov2023. We set $\delta^{(p)}=0$ for $p\notin\{1,2,3\}$ and impose \[ \delta^{(1)}=\delta^{(2)}=\delta^{(3)}=\delta, \qquad \delta\in\{0,1,2,3\}. \]

When $\delta=0$, the null (ref) holds and empirical rejection frequencies measure test size. When $\delta>0$, rejection frequencies measure power. Unlike the sharp-null configuration considered in Bojinov2023, our specification does not imply a sharp null when $\delta=0$, allowing a direct comparison between Fisher randomization tests and conditional randomization tests under a non-sharp null.

For each parameter configuration, we generate one treatment assignment path and compute observed outcomes from (ref). We then apply four testing procedures:

enumerate• the proposed conditional randomization test (CRT), • a Fisher randomization test imposed under a misspecified sharp null, • the unconditional PIRT procedure of zhong2024unconditional with rejection threshold $\alpha$, • the Horvitz--Thompson asymptotic test of Bojinov2023.

For asymptotic inference, we use the conservative variance upper bound proposed in Bojinov2023. All tests are conducted at nominal level $\alpha=0.05$ using the Horvitz--Thompson estimator. Each configuration is replicated 1,000 times, and rejection rates are reported as Monte Carlo averages.

figure[figure omitted — 782 chars of source]

Figure (ref) reports rejection frequencies across sample sizes and error distributions. The CRT maintains size control across all configurations. In contrast, the Fisher randomization test exhibits substantial over-rejection under heavy-tailed disturbances, reflecting misspecification of the sharp null when treatment effects are not fully imputable. This pattern is consistent with the theoretical results of Athey2018. Both the unadjusted PIRT at level $\alpha$ and the Horvitz--Thompson asymptotic test display conservative behavior in several configurations.

In terms of power, the CRT performs competitively across all scenarios. Under Gaussian disturbances, the CRT, Fisher test, and asymptotic test exhibit similar power, while PIRT is somewhat less powerful. Under heavy-tailed disturbances, the CRT achieves the highest power among procedures that maintain valid size. Although the Fisher test may appear more powerful in some configurations, this reflects size distortion rather than genuine power gains.

Overall, the results indicate that the proposed CRT achieves reliable size control while maintaining strong power, and remains robust to heavy-tailed shocks and finite-sample dependence.

Tests for $m$-carryover and anticipation

This section evaluates the finite-sample size and power of the two diagnostic procedures proposed above. We study the $m$-carryover restriction using the conditional randomization test (CRT), and the non-anticipation restriction using the PIRT procedure.

\paragraph{Carryover diagnostic.}

To assess the $m$-carryover test, we fix the true carryover horizon at $m=2$ and evaluate the null that treatment effects vanish beyond this horizon. The data-generating process introduces carryover effects strictly beyond the tested horizon by setting \[ \delta^{(p)}=0 \quad \text{for } p\notin\{4,5,6\}, \qquad \delta^{(4)}=\delta^{(5)}=\delta^{(6)}=\delta, \] with $\delta\in\{0,1,2,3\}$. Thus $\delta=0$ corresponds to the null of no carryover beyond $m=2$, while $\delta>0$ generates violations of the null. Inference is conducted using the proposed CRT. Empirical rejection frequencies therefore measure size when $\delta=0$ and power when $\delta>0$.\footnote{For both carryover and anticipation diagnostics we use the same outcome model. We do not compare directly with the method of Bojinov2023 (Section 4.4), since their approach requires a different experimental design tailored to each hypothesis.}

\paragraph{Anticipation diagnostic.}

To evaluate the non-anticipation test, we introduce dependence on future treatment assignments by setting \[ \delta^{(p)}=0 \quad \text{for } p\notin\{-2,-1,0\}, \qquad \delta^{(-2)}=\delta^{(-1)}=\delta^{(0)}=\delta, \] again with $\delta\in\{0,1,2,3\}$. The null of non-anticipation holds when $\delta=0$ and is violated when $\delta>0$. Inference is conducted using the proposed PIRT procedure with prefix-preserving randomization and holdout length $L=2m+1=5$.

figure[figure omitted — 478 chars of source]

Figure (ref) reports rejection frequencies across sample sizes and error distributions for the carryover diagnostic. The CRT maintains accurate size control across all configurations. Power increases with the magnitude of the violation and with the sample size, reaching approximately 60% in the largest designs under Gaussian disturbances.

figure[figure omitted — 479 chars of source]

Figure (ref) reports rejection frequencies for the non-anticipation diagnostic using the PIRT procedure. Size control remains close to nominal across designs. Power increases with the magnitude of anticipation effects but does not vary monotonically with the total horizon, because the effective sample size is driven by the number of imputable times—i.e., stochastic overlaps of prefixes—rather than by $T$ per se. Unlike carryover testing, where focal sets are deterministic given the design, anticipation testing inherits additional finite-sample variability from this random overlap structure, so the count of imputable times can remain limited even as $T$ grows, leading to irregular power patterns across realizations.

Overall, the results indicate that the proposed procedures provide reliable finite-sample size control and nontrivial power in realistic time-series environments. A more detailed theoretical analysis of power for PIRT-based diagnostics in dependent time-series settings is left for future work.

Conclusion

We study inference for switchback experiments run over realistic operational horizons, where serial dependence, heavy-tailed shocks, and temporal interference can undermine classical approximations. Our contribution is a randomization-test framework that yields finite-sample valid, distribution-free $p$-values using only the known assignment mechanism and two primitive conditions on potential outcomes: non-anticipation and a finite carryover horizon. We develop randomization tests for total effects alongside diagnostics for carryover length and non-anticipation. For average-effect targets, we develop studentized CRTs that accommodate within-session seasonality and establish asymptotic validity under mild conditions.

We complement the theory with power approximations under distributed-lag effects with AR(1) noise, which highlight block length, burn-in, and section pooling as first-order design levers for sensitivity. Monte Carlo evidence is consistent with these insights, indicating reliable finite-sample performance of the proposed tests across realistic settings. A fuller treatment of power optimization and simulation-based design tuning remains an important direction for future work.

Moreover, our testing framework naturally extends beyond the canonical switchback setting and applies directly to a broad class of experimental designs with a time dimension—including time-series experiments, cross-over designs, and regular balanced switchback designs—thereby providing a unified approach to exact inference in temporally structured experiments. Further advances in design optimization, richer temporal interference models, and scalable implementations within platform experimentation systems offer promising avenues for continued research.