Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
131,143 characters · 33 sections · 54 citation commands
Real-time Program Evaluation using Anytime-valid Rank Tests*
\begingroup \footnotetext{Corresponding author: Sam van Meer, [email removed]. We thank Alberto Abadie, Dick van Dijk, Stan Koobs, Bernhard van der Sluis, and Wendun Wang for their valuable feedback. We are also grateful to participants at IAAE 2024, EAYE 2024, Camp Econometrics 2024, and the ISEO Summer School 2024 for their helpful comments.}
\endgroup
We sequentially observe outcomes $Y^T = (Y_1, Y_2, \dots, Y_{T_0}, \dots, Y_{T})$, of some individual or unit that undergoes a treatment at time $T_0$, up until some possibly infinite final time $T$. Starting from the treatment time $T_0$ onwards, we are interested in testing the hypothesis that the treatment has no effect on our outcome $Y^T$.
We monitor this by comparing the post-treatment outcomes to the pre-treatment outcomes. To make these comparable, we assume they are appropriately transformed to be exchangeable in absence of a treatment effect. Depending on the context, this means the treatment outcomes may be directly observed, or they may be estimates obtained from some counterfactual mean method such as difference-in-differences or synthetic control, as we demonstrate later. We now formulate the hypothesis of no treatment effect $H_0$ as
Here, exchangeability means that for any finite $t\leq T$, the distribution of data $Y_1, \dots Y_{t}$ is invariant under re-ordering of its observations. The advantage of this abstract formulation of the notion of `no treatment effect' is that it can not only be used to capture changes in the mean, but also other changes in the distribution.
This problem is well-studied, but only for situations where the number of observations $T$ is pre-specified or, equivalently, independently specified. For example, chernozhukov2021exact and abadie2021synthetic propose a permutation $p$-value $p_T$ that is valid for every pre-specified $T > T_0$:
This guarantees that at every point in time $T$, the probability that the $p$-value $p_T$ is below $\alpha$ is at most $\alpha$.
When evaluating the impact of a treatment, the requirement to pre-specify the number of observations $T$ is restrictive. Indeed, if we set $T$ too large, then we incur the cost of collecting too much data, and an opportunity cost due to the delay. On the other hand, setting $T$ too small may lead to an underpowered test, which comes with the opportunity cost associated to not discovering an effective treatment.
To prevent collecting too much data, it may be tempting to reject the hypothesis $H_0$ at the first time $T$ at which $p_T \leq \alpha$. However, for traditional tests, such real-time monitoring is prohibited, as this makes the number of observations $T$ data-dependent while the Type-I error guarantee in (ref) only holds for data-independent $T$. The probability that a standard $p$-value ever dips below $\alpha$ is typically much larger than $\alpha$.
The primary contribution of this paper is to develop anytime-valid $p$-values $\mathfrak{p}_T$ for real-time program evaluation. In contrast to a regular $p$-value as in (ref), such an anytime-valid $p$-value is valid at arbitrary data-dependent times $T$. In particular, the probability under $H_0$ that $\mathfrak{p}_T$ ever dips below $\alpha$ is bounded by $\alpha$:
A consequence is that we need not pre-specify a number of observations, but we can test at every moment in time without invalidating the Type I error guarantee.
We tailor our anytime-valid $p$-values to the key feature of this problem: the fact that at the moment we start testing, we already have an entire batch of $T_0$ pre-treatment observations available to us. If we order the pre-treatment observations by magnitude, the exchangeability hypothesis $H_0$ implies that the first incoming post-treatment observation is equally likely to land in any position between these pre-treatment observations. We exploit this to design two frameworks, centered around two alternative hypotheses, for which we construct our anytime-valid $p$-values. \\
Post-exchangeable alternative \\ In the first framework, we construct our $p$-values by keeping track of how the newly incoming post-treatment observations rank among the pre-treatment observations. As this framework does not depend on the ranking of the post-treatment data among themselves, this can be translated into an alternative hypothesis in which the post-treatment data are exchangeable:
An interpretation of this alternative is that after time $T_0$, the treatment shifts the data generating process to some new stable equilibrium from which the post-treatment data are drawn. \\
Generic alternative \\ In our second framework, we do not just consider how the post-treatment data ranks among the pre-treatment data, but also how the post-treatment data ranks among itself. This translates into the `generic' alternative hypothesis that the treatment outcomes are simply not exchangeable:
This lends itself better, for example, to situations in which the underlying treatment effect fluctuates or continues to grow over time. Note that this hypothesis nests the post-exchangeable alternative defined before. Therefore, tests against this generic alternative are generally expected to be less powerful against the narrower post-exchangeable alternative.\\
For both these alternatives, we derive a flexible way to construct an anytime-valid $p$-value. In particular, our construction works for any choice of non-negative test statistic, that may arbitrarily depend on the previously observed ranks or on external information. If the data generating process under the alternative is known, we show how to exploit this to derive a $p$-value with a type of optimal shrinkage rate grunwald2024safe. As an example, we derive the optimal anytime-valid procedure under a Gaussian alternative. If the data generating process under the alternative is not known, we derive an adaptive test statistic that learns the alternative during the sequential testing procedure fedorova2012plug.
In addition, we demonstrate how our methodology can be applied in an interactive fixed-effects model. There, both difference-in-differences and synthetic control produce exchangeable treatment estimates in absence of a treatment effect under standard assumptions abadie2021synthetic, chernozhukov2021exact. We use this setting to illustrate our methods in a simulation study, and demonstrate that our methods control size over time when the treatment estimators are (approximately) exchangeable under the null. If exchangeability is strongly violated due to presence of serial dependence, we show that size distortions can be mitigated by imposing a block structure, as proposed in the fixed-$T$ setting by chernozhukov2021exact. In terms of power, we find that the cost of performing anytime-valid inference is limited, especially when the expected effect size is accurately specified. We highlight the benefit of our real-time anytime-valid inference in a discounted utility model, where we find our anytime-valid tests to be preferred over the fixed-$T$ test for a wide range of intertemporal discount factors.
In this section, we briefly illustrate how counterfactual and synthetic control (CSC) methods chernozhukov2021exact can give rise to exchangeable treatment estimates, so that our methods are applicable.
We consider a panel with $J + 1$ units over time, where only the first unit $i = 1$ is treated after time $T_0$, and the remaining $J$ units are never treated. The observed outcomes of interest $Y^{\text{obs}}_{it}$ are modelled through the potential outcomes framework:
where $Y_{it}(1)$ is the outcome at time $t$ of unit $i$ if it is treated, and $Y_{it}(0)$ if not. Assuming the data is real-valued, we are interested in the causal effect of treatment on the treated unit, which is defined as $\tau_t=Y_{1t}(1)-Y_{1t}(0)$.
Unfortunately, by construction, we only observe $Y_{1t}(0)$ up until time $T_0$ and only observe $Y_{1t}(1)$ beyond time $T_0$: we never observe both $Y_{1t}(1)$ and $Y_{1t}(0)$ at any particular time $t$. As a consequence, $\tau_{t}$ cannot simply be computed, but needs to be inferred. CSC methods estimate $\tau_t$ by constructing an approximation $\widehat{Y}_{1t}(0)$ of the counterfactual outcome $Y_{1t}(0)$ of the treated unit, had it not been treated. This approximation is then used to construct the treatment estimator $\widehat{\tau}_t = Y_{1t}^{\text{obs}} - \widehat{Y}_{1t}(0)$.
To construct the counterfactual approximation, it is common to use the observed outcomes $Y^{\text{obs}}_{it}$ of the untreated units. The Synthetic Control Method abadie2003economic, then constructs the approximation of the counterfactual $\widehat{Y}_{1t}(0)$ as a linear combination of the untreated outcomes, $\widehat{Y}_{1t}(0)=\sum_{i=2}^{J+1}w_i Y^{\text{obs}}_{it}$ for some appropriate weights $w_i$.
Let $u_t$ denote the approximation error of this proxy, such that $\widehat{\tau}_t=\tau_t+u_t$. If $u_1, u_2, ...$ are exchangeable under $H_0$, then we can recast the null-hypothesis $H_0: \tau_{T_0+1}=\tau_{T_0+2}=...=0$ into $H_0: \{\widehat{\tau}_t\}_{t>0}$ is exchangeable chernozhukov2021exact. Hence, sequential testing for a treatment effect can be done by sequentially testing for exchangeability of the CSC estimator. Later, we show how the interactive fixed effects model by bai2009panel can give rise to these exchangeable approximation errors.
As mentioned, our work is related to that of chernozhukov2021exact, which studies inference for counterfactual mean models. Their work is also based on exchangeability over time, and relies on the same assumptions. However, they only consider testing at a fixed number of observations $T$, whereas our methods permit real-time inference with an adaptive number of observations.
Our paper also relates to the new synthetic design literature, which applies the synthetic control method in experimental setting abadie2021synthetic,doudchenko2019designing,doudchenko2021synthetic. Compared to traditional A/B testing, synthetic control has the advantage that it can account for regional market effects and spillovers. A key focus of this literature is on experimental design, particularly in optimizing the selection of treatment units. Since our methods allow for endogenous treatment selection, they can be integrated with such experimental designs to enable sequential inference in the context of synthetic design.
There already exists some work on sequential causal inference under more restrictive assumptions. maharaj2023anytime study anytime-valid inference for A/B testing, but they assume a multiple-treated-units framework in which the outcome of a new treated unit is observed with every observation. In contrast, our methodology also works for a single treated unit that undergoes one treatment. Similarly, ham2022design develop asymptotically anytime-valid confidence intervals for causal inference. Their methods work in the time-series/panel setting, but require that treatment is randomly assigned at every time point. Our method is different from theirs in that we allow for endogenous treatment selection, and do not require new treatment assignment at every time point. Also, our proposed methods do not rely on asymptotic approximations but have finite sample size guarantees.
Rank-based tests for sequential exchangeability in an abstract context can be traced back to at least vovk2003testing. However, to the best of our knowledge, they have not been previously applied in treatment evaluation and have generally been surprisingly underappreciated in statistical inference. Other recent applications include testing independence henzi2024rank and outlier detection bates2023testing. Moreover, koning2024post studies how this can be generalized beyond testing exchangeability to other forms of invariance.
Compared to existing work on rank-based anytime-valid testing of exchangeability, our work has two unique features. First, we have an entire batch of pre-treatment observations available to us when we start testing. Second, we exploit the pre-/post-treatment structure in an alternative hypothesis under which we only consider the ranks of the post-treatment data among these pre-treatment observations. Interestingly, fischer2024sequential study a related approach in a completely unrelated context: reducing the number of Monte Carlo simulations in a Monte Carlo-based hypothesis test. Abstractly speaking, their approach can be interpreted as the special case of our method, in which only a single pre-treatment period is considered.
To simplify the discussion, we consider testing a simple null hypothesis that contains a single distribution $\mathbb{P}$ throughout this section. While our null hypothesis of exchangeability is highly composite, we show that it can be reduced to a simple hypothesis in Section (ref).
The construction of our anytime-valid sequential tests relies on a recently introduced measure of evidence called the $e$-value shafer2011test, grunwald2024safe,howard2021time,vovk2021values, ramdas2023game,grunwald2024safe,koning2024continuous. An $e$-value is typically defined as a non-negative random variable $e$ with expectation bounded by 1 under the null hypothesis:
The definition of the $e$-value does not immediately convey its usefulness in a testing context, but a simple application of Markov's inequality shows that we can construct a valid test from an $e$-value. In particular, a test that rejects when $e \geq 1/\alpha$ is valid at level $\alpha > 0$:
Moreover, $e$-values inherit simple merging rules from the properties of the expectation: the average of two $e$-values and the product of independent $e$-values is also an $e$-value. A generalization of this product-merging property combined with the possibility to convert an $e$-value into a test forms the basis for a powerful machinery for constructing anytime-valid sequential tests ramdas2020admissible. We explain the key parts of this machinery in this section.
In the context of sequential testing, we must carefully describe the available information at every moment in time. This means that our sample space is not just equipped with some sigma algebra of events, but with an entire filtration $\{\mathcal{F}_t\}_{t \geq 0}$ that describes the available information $\mathcal{F}_t$ at time $t \geq 0$. With such a filtration, we can speak of a sequential $e$-value $e_t$ at time $t$, if it is an $e$-value with respect to the available information $\mathcal{F}_{t-1}$. Without loss of generality, we set $e_0 = 1$ out of convenience.
In order to construct an anytime-valid sequential test out of these sequential $e$-values, we start by stringing them together through multiplication. In particular, the running product $W_t = \prod_{s = 0}^t e_s$ of a sequence of sequential $e$-values $e_0, e_1\dots$, where $e_0=1$, constitutes a non-negative supermartingale that starts at 1. Such martingales are also referred to as test martingales, for their applications in testing shafer2011test.
Recall that our aim is to construct an anytime-valid $p$-value $\mathfrak{p}_t$, as defined in ((ref)). It turns out that this can be accomplished by simply setting the $p$-value equal to the reciprocal of a test martingale: $\mathfrak{p}_t:=\frac{1}{W_t}$. Indeed, this follows from an application of Ville's inequality:
under the null hypothesis ville1939etude,grunwald2024safe. In words, this means that for the $p$-value $\mathfrak{p}_t$, the probability that it ever dips below $\alpha$ is bounded by $\alpha$. As such $p$-values are actually stochastic processes, they are sometimes also referred to as $p$-processes.
Ville's inequality, the inequality in (ref), can be viewed as a sequential generalization of Markov's inequality for martingales. If we interpret Markov's inequality as stating that the odds to double our money in a fair bet are no more than 50/50, then Ville's inequality states that the odds to ever double our money in a sequence of fair bets is no more than 50/50.
A testing interpretation of Ville's inequality is that the probability our test martingale $W_t$ ever exceeds the critical value $1/\alpha$ is bounded by $\alpha$. This is true in particular for the first data-dependent time $\inf\{t > 0 \mid W_t \geq 1/\alpha\}$ at which the test martingale exceeds the threshold, but also for any other possibly data-dependent time. A consequence is that we can continuously evaluate whether our test martingale exceeds $1/\alpha$, which allows for early rejection if the evidence is strong or continuation if the evidence remains weak.
It remains to discuss how to construct a good anytime-valid $p$-value. Ideally, we would like a $p$-value that shrinks `as quickly as possible' when the alternative hypothesis is true. As our anytime-valid $p$-value is the reciprocal of a test martingale, this is equivalent to finding a test martingale that grows as quickly as possible.
Before we can even define what it means for a test martingale to grow as quickly as possible, we must first define when it should grow as quickly as possible. To do this, let us assume for now that if the alternative hypothesis is true, then the data are sampled from the distribution $\mathbb{Q}$. This means our test martingale should grow as quickly as possible if the data comes from $\mathbb{Q}$.
As our test martingale is the running product of sequential $e$-values, the problem can be decomposed into choosing sequential $e$-values that lead to a high growth rate. Let $\mathbb{Q}_t$ denote the conditional distribution given the available information $\mathcal{F}_{t-1}$ at time $t -1$, which is the moment at which we must choose the sequential $e$-value that we will use at time $t$.
A naive strategy would be to choose the sequential $e$-value that has the largest expectation under $\mathbb{Q}_t$. Unfortunately, this sequential $e$-value is usually zero with a substantial positive probability. This is problematic: if we encounter a single zero along the way, this will multiply the current value of our martingale by zero, which means it will equal zero and can never grow again. For this reason, the anytime-validity literature typically considers different targets.
A popular target is to maximize the expected logarithm $\mathbb{E}^{\mathbb{Q}_t}[\log e_t]$, which is equivalent to maximizing the geometric expectation. The motivation for this target comes from the i.i.d. setting, where this maximizes the long-run growth rate kelly1956new. Indeed, if $e_1, \dots e_t$ are i.i.d., then a combination of the strong law of large numbers and the continuous mapping theorem tells us that this maximizes the average asymptotic growth rate:
Remarkably, for a simple null hypothesis with distribution $\mathbb{P}$ and alternative $\mathbb{Q}$, where $\mathbb{P}\gg \mathbb{Q}$, this `log-optimal' sequential $e$-value simply equals the conditional likelihood ratio kelly1956new, koolen2022log, grunwald2024safe, larsson2024numeraire
This means that the growth-rate maximizing test martingale is the product of these conditional likelihood ratios.
In our application, the alternative hypothesis $H_1$ is typically composite, as it is often hard to know beforehand how effective a treatment will be. This means that we do not know the precise distribution $\mathbb{Q}$ from which the data are sampled when the alternative hypothesis is true. Fortunately, we can use the sequential nature of the setting to adaptively learn about this distribution while we are testing, and update our choice of sequential $e$-values accordingly.
One such approach is the plug-in approach, where we plug our current estimator of $\mathbb{Q}$, given the available information, in for $\mathbb{Q}$ in the target. Another approach is the method of mixtures, where our sequential $e$-value at time $t$ is an adaptive mixture over multiple candidate sequential $e$-values at time $t$ ramdas2023game, ramdas2020admissible. Here, an adaptive mixture means that we can update the mixing weights based on the historical performances of each of these $e$-values.
In case we consider $k \in \mathbb{N}$ candidate sequential $e$-values, there exists a weighting scheme for these $e$-values with an attractive minimax regret property, as demonstrated in Appendix (ref). This adaptive weighting scheme, when applied to the candidate $e$-values, reduces to taking a simple average over their respective martingales. The regret bound states that the average growth of this mixture martingale at time $t$ is at most $\frac{\log k}{t}$ away from that of the $e$-value candidate with highest average growth.
In this section, we present our main contribution: the construction of anytime-valid $p$-values for program evaluation. In Section (ref), we showed that such an anytime-valid $p$-value $\mathfrak{p}_t$ can be constructed as the reciprocal $\mathfrak{p}_t = 1/W_t$ of a test martingale $W_t$. In turn, a test martingale is formed as the running product $W_t = \prod_{s = 1}^t e_s$ of sequential $e$-values $e_1, \dots, e_t$. Therefore, we focus in this section on the construction of sequential $e$-values for program evaluation, as they are the building blocks for anytime-valid $p$-values.
Recall that we sequentially observe the treatment estimates $Y_1, Y_2, \dots, Y_{T_0}, \dots$, where $Y_1, \dots Y_{T_0}$ are considered untreated, and $Y_{T_0 + 1}, Y_{T_0 + 2}, \dots$ treated. We assume throughout that these treatment estimates have been appropriately de-meaned or otherwise transformed so that they are exchangeable in absence of a treatment effect. This means we can reject the hypothesis of no treatment effect, if we reject the hypothesis of exchangeability:
In Section (ref), we detail in an interactive fixed effects model how assumptions on the data generating process for difference-in-differences and synthetic control give rise to such exchangeable treatment estimates.
Moreover, recall our two alternatives: the more specific alternative $H_1^\textnormal{post-exch.}$, where post-treatment data are assumed to be exchangeable among each other, but not exchangeable with the pre-treatment data, and the broader alternative $H_1^\textnormal{generic}$, where the post-treatment data are not exchangeable with the pre-treatment data or among each other.
Interestingly, there exist no non-trivial test martingales for exchangeability under the natural filtration: the filtration under which we have full information about the original data $Y_1, Y_2, \dots, Y_T$ vovk2021testing, ramdas2022testing. This means it is necessary to consider a less informative `coarsened' filtration. The practical consequence is that in exchange for having a powerful test martingale, we are only allowed to base our decision to stop or continue testing on some statistic of the data; not the underlying data itself.
This may superficially seem paradoxical: if we only base our decision to stop or continue on a statistic of the data, then we must surely lose power? This paradox is resolved by realizing the coarsening of the filtration simultaneously reduces the collection of data-dependent (`stopping') times under which our method is required to be anytime-valid. This is aptly named `the power of forgetting' by vovk2023power.
Another interpretation is that we may only be interested in alternatives that depend on this statistic, so that the precise values of the underlying data are uninteresting to us. This means that we consider a narrower alternative hypothesis towards which we can direct our power. In this interpretation, the coarsening of the filtration is inconsequential: it discards irrelevant information.
For exchangeability, the natural\footnote{Ranks are natural because, if there are no ties, they are in bijection with the group of permutations that underlies exchangeability koning2024post. This means that testing uniformity of the ranks can be interpreted as testing uniformity on the group of permutations.} reduction of the filtration is obtained by converting the data into their sequential ranks: the rank $R_t$ of the new observation $Y_t$ among the previous observations $Y_1, \dots, Y_{t-1}$. We denote the filtration generated by the sequential ranks by $\mathcal{G} = \{\mathcal{G}_t\}_{t\geq0}$. The consequence of this reduction is that we can obtain powerful test martingales, at the cost that our inference can only depend on the sequential ranks.
A key property of the sequential ranks is that under exchangeability, the previous ranks are completely uninformative about the next rank. This means that under this reduced filtration, our null hypothesis can be recast as sequentially testing whether the sequential ranks are uniformly distributed vovk2003testing:
for every $t > T_0$.
This reduces the highly composite hypothesis of exchangeability to a simple hypothesis on the ranks. Under this coarsened filtration, the generic alternative hypothesis is simply the negation of the null:
which means that the alternative remains composite. We handle this in Section (ref).
In Theorem (ref), we define the building block for our anytime-valid $p$-value for this hypothesis: a sequential $e$-value $e_t$. This sequential $e$-value offers great flexibility, as it can be defined for an arbitrary non-negative test statistic $S_t$, that may depend on the past sequential ranks $\mathcal{G}_{t-1}$ and on external randomization. An interpretation of this $e$-value is that it quantifies how the current value of the statistic $S_t(R_t)$ compares to the average among its other possible realizations.
Moreover, Theorem (ref) also shows that if the alternative of the $t$'th sequential rank conditional on $\mathcal{G}_{t-1}$ is simple, then log-optimality (see Section (ref)) is attained by choosing $S_t$ proportional to the conditional density of this alternative. All proofs are presented in Appendix (ref). In Section (ref), we appropriate choices for composite alternatives.
Recall that the post-exchangeable alternative posits that the entire sequence of observations $Y_1, Y_2, \dots, Y_{T_0}, \dots, Y_T$ is not exchangeable, but the post-treatment observations $Y_{T_0+1}, Y_{T_0+2}, \dots$ are. For this alternative, we are unconcerned with the ranks of the post-treatment observations $Y_{T_0 +1}, \dots Y_{T}$ amongst themselves: we only care about how the post-treatment observations rank among the pre-treatment observations $Y_1, \dots Y_{T_0}$. As a consequence, we can reduce the entire problem to a question concerning what we call the reduced sequential ranks.
We call these ranks reduced, because they are less informative than the sequential ranks and take value on the reduced space $\{1, \dots, T_0 + 1\} \subseteq \{1, \dots, t\}$, $t > T_0$. As we are only interested in these reduced sequential ranks here, we may freely coarsen the filtration to the one induced by these reduced ranks $\{\widetilde{\mathcal{G}}_t\}_{t \geq 0}$.
Under exchangeability, the reduced sequential ranks are generally not independent nor uniform like the conventional sequential ranks in (ref). Instead, they follow a conditional categorical distribution over $T_0 + 1$ options,
where the probability $q_t^{i}$ that the $t$'th reduced sequential rank $\widetilde{R}_t$ equals $i \in \{1, \dots, T_0 + 1\}$ is proportional to one plus the number of previous sequential ranks that equal $i$:
This distribution is uniform on $\{1, \dots, T_0 + 1\}$ at time $t = T_0 + 1$, but generally different at later times, depending on where the previous post-treatment observations rank among the pre-treatment observations. As this may be counterintuitive, we provide intuition for this null distribution in Section (ref).
For the reduced sequential ranks, the post-exchangeable alternative hypothesis is again composite, and simply states that the reduced sequential ranks do not follow this distribution
In Theorem (ref), we present a sequential $e$-value (ref) for the reduced sequential ranks that is similar to the one for the sequential ranks from Section (ref), but now rescaled using the previously observed reduced ranks. The proof of this theorem is similar to that of Theorem (ref) and is presented in Appendix (ref).
In Figure (ref), we illustrate the difference between the null distributions of the sequential ranks and reduced sequential ranks. It pictures the null distribution as a histogram of both the rank and reduced rank of the $t = 8$th observation $Y_8$, having observed the same $T_0 = 4$ pre-treatment observations and 3 post-treatment observations. In Panel (ref), we see that the exchangeability means that $Y_8$ is equally likely to fall in any slot between pre- or post-treatment observations, and so its rank $R_8$ is uniform on $\{1, \dots, 8\}$. In Panel (ref), we picture the null distribution of the reduced sequential rank, which discards the position of $Y_8$ among the post-treatment observations and only conveys its position among the pre-treatment observations.
Comparing the two panels, we see that the mass of the three slots between the first two pre-treatment observations in Panel (ref) is 3 out of 8, which exactly corresponds to the mass at slot 2 between the first two pre-treatment observations in Panel (ref). More generally, the mass of the null distribution of the reduced rank at a slot between two pre-treatment observations equals one plus the number of post-treatment observations that previously fell into the same slot.
We now consider the construction of good test statistics $S_t$ for the alternatives \( H_1^\text{generic} \) and \( H_1^\text{post-exch.} \). Recall from Theorems (ref) and (ref) that the log-optimal test statistic is proportional to the conditional density of the true data generating distribution under the alternative. Unfortunately, this true data generating distribution is unknown as \( H_1^\text{generic} \) and \( H_1^\text{post-exch.} \) are composite, so that we cannot choose the log-optimal test statistic as such. To address this problem, we consider two approaches: sequentially learning the distribution under the alternative through a plug-in method, and transforming a (possibly composite) alternative on the outcomes to a single implied distribution on the ranks.
The idea behind the plug-in method is that if the distribution of the post-treatment ranks is unknown, then perhaps we can sequentially learn it due to our sequential setup. The idea is to choose the test statistic $S_t$ to be proportional to (a smoothed version of) the empirical density of the post-treatment ranks that we have observed so-far. If the distribution of the post-treatment ranks is stable over time, then after sufficiently many post-treatment draws we should start to learn this distribution.
The plug-in method we describe here is inspired by the work of fedorova2012plug, who use it for testing exchangeability abstractly. We modify their approach to accommodate the pre- and post-treatment structure of our problem. Moreover, we enhance the plug-in approach for testing against the narrower alternative $H_{1}^\text{post-exch.}$.
\paragraph*{Rank based test} For $H_1^\text{generic}$, we aim to approximate the conditional distribution of the sequential ranks $R_t$, based on $\mathcal{G}_{t-1}$. As the support of the ranks $\{1,\dots,t\}$ is discrete and increases with $t$, we rescale and smooth the sequential ranks before estimating the kernel density.\footnote{Adding independent noise for kernel estimation of discrete data is also called jittering, which is shown to work well in finite samples nagler2018asymptotic.} Using independent noise $u_t\sim U(0,1)$ for smoothing, we define the smoothed rescaled ranks as
where $V_t$ has support on $[0,1]$ and is independently uniformly distributed under $H_0^\text{rank}$.
We estimate the distribution of $V_t$ with a continuous kernel density. Similar to fedorova2012plug, we apply a standard kernel trick to improve the accuracy of kernel estimates near the boundaries of \([0,1]\). Specifically, we estimate the kernel density using an extended sample, defined as $\mathcal{V}_{t-1}=\cup_{s=T_0+1}^{t-1}\{-V_s,V_s,2-V_s\}$. Let $\widehat{f}_t$ denote the resulting kernel density estimate at time $t$, that is suitably normalized to integrate to 1 on the domain $[0, 1]$ of $V_t$. This then yields the plug-in test statistic
which can be interpreted as the estimated kernel density of the sequential ranks, based on the information in $\mathcal{G}_{t-1}$. Since $\widehat{f}_t$ is estimated based on only $R_s,u_s$ for $s<t$, we have that $S_t\perp R_t|\mathcal{G}_{t-1}$. Also, the test statistic is non-negative. Hence $S_t$ adheres to all conditions of Theorem (ref) and can be used for anytime-valid inference.\footnote{In this particular case, the denominator in ((ref)) can be set equal to 1 to decrease variation, while still maintaining validity.}
The log-optimality results imply that if the kernel is a good estimator of the true distribution, then we expect our test martingale to grow quickly. fedorova2012plug show that this holds asymptotically if the underlying distribution of sequential ranks is stable, i.e, the empirical distribution of $\{V_t\}_{t\geq0}$ converges to some distribution with a continuous density. Moreover, they show that for a consistent kernel estimator, the asymptotic log growth of the martingale corresponding to $S_t^{\text{plug-in}}$ dominates every martingale based on a test statistic $S_t$ that is constant over time. \paragraph*{Reduced rank based test} We now propose a specialized improvement of the plug-in method that is more powerful for the post-exchangeable alternative. This improvement relies on the fact that the reduced sequential ranks disregard any ordering of post-treatment observations and only retains their position relative to the pre-treatment observations. As these post-treatment observations are exchangeable under $H_1^\text{post-exch.}$, the fact that the reduced ranks disregard this information can improve the estimation.
We consider the same kernel density estimator $\widehat{f}_t$, and its CDF $\widehat{F}_t$ as the ones we used for the rank based test. We estimate the discrete distribution of $\widetilde{R}_{t}$ by integrating $\widehat{f}_t$ over the regions corresponding to every realization of $\widetilde{R}_t$. Some algebraic manipulation shows that this can be written as the difference of the kernel CDF evaluated at points $\sum_{i=1}^{\widetilde{R}_t-1}q_t^{i}$ and $\sum_{i=1}^{\widetilde{R}_t}q_t^{i}$,
The conditions for anytime-validity as stated in Theorem (ref) hold, as $\widetilde{S}_t^\text{plug-in}$ is non-negative, and independent of $\widetilde{R}_t$ given $\mathcal{G}_{t-1}$. The following theorem shows that, under $H_1^{\text{post-exch.}}$, an anytime-valid test based on this $\widetilde{S}^\text{plug-in}$ statistic weakly dominates the sequential $e$-value based on $S^\text{plug-in}$ in terms of the log-target (proof in Appendix (ref)). Moreover, in simulations we see that it performs substantially better in practice, if the alternative is indeed post-exchangeable.
Even if we would like to specify a simple distribution on the (reduced) sequential ranks, it can be hard to express such a distribution. Indeed, we may only have an idea about the distribution of the outcomes $Y_1,\dots,Y_T$. Luckily, we can simply take such a distribution on the outcome space and recover the distribution it induces on the (reduced) ranks. Moreover, many distributions on the outcome space may have the same induced distribution on the ranks.
In this section, we illustrate this for a Gaussian alternative. We find that Gaussian distributions with the same effect size / signal-to-noise ratio all have the same implied distribution on the sequential ranks. As a consequence we do not need to specify both the mean and the variance: we only need to choose an effect size.
While it may feel as a burden that one has to specify such an effect size, note that this is usually also required in the fixed-$T$ setting: either explicitly to plan the number of to-be-gathered observations, or implicitly by gauging whether the number of available observations would be sufficient to find an effect.
\paragraph*{Rank based test} Let us consider a simple setting for $\mathbb{Q}$, where the outcomes (pre- and post-treatment) are independently\footnote{The more general exchangeable setting with a correlation coefficient of $\rho$ can also be easily accommodated by dividing $Y_t$ by $\sqrt{(1-\rho^2)}$.} normally distributed, $Y_t \sim \mathcal{N}(\mu_t, \sigma^2)$, with some mean $\mu_t$ and variance $\sigma^2$. Suppose the impact of the treatment is a change in the mean of this normal distribution. In particular, $\mu_t = \mu_{\text{pre}}$ for $t \leq T_0$ and $\mu_t=\mu_\text{t,post}$, otherwise.
By Theorem (ref), the log-optimal sequential $e$-value is obtained for $S_t\propto f_t$ where
for all $t>T_0$.
These densities do not permit a simple expression, but they can be easily simulated through Monte Carlo draws of the Gaussian distribution $Y_t\sim \mathcal{N}(\frac{\mu_t-\mu_{\text{pre}}}{\sigma},1)$, where we use that the distribution of the ranks is invariant to univariate shifts in the mean or variance.
A practical disadvantage is that we must specify a compete path for $\mu_t$, $t > T_0$ under the alternative, which may be difficult. The post-exchangeable alternative is exactly the setting where this is not necessary, and we simply fix $\mu_t$ over time.
\paragraph*{Reduced rank based test} Under post-exchangeability, the post-treatment observations are exchangeable and so identically distributed. As a consequence, the mean is the same for the entire post-treatment period: $\mu_{t,post}=\mu_{\text{post}}$, for all $t > T_0$.
Choosing the test statistic $\widetilde{S}_t=\widetilde{f}_t$, where $\tilde{f}_t$ is the distribution of $\widetilde{R}_t|\mathcal{G}_{t-1}$, ensures log-optimality through Theorem (ref). An efficient way to approximate this density for the Gaussian setting, based on the expected effect size $\frac{\mu_\text{post}-\mu_\text{pre}}{\sigma}$, is provided in Appendix (ref). The calculation of this $\tilde{f}_t$ is easier than that of $f_t$ as each $\tilde{f}_t$ only requires draws from a $T_0$-dimensional standard normal distribution, instead of a $t$-dimensional distribution.
In practice, this effect size $ \frac{\mu_\text{post} - \mu_\text{pre}}{\sigma}$ is unknown, but it is possible to average over multiple candidates using the adaptive mixture method described in Appendix (ref). An appealing property of this weighted averaging is that its cumulative log-growth converges to that of the best-performing candidate in hindsight.
This model averaging approach can be viewed as a computationally tractable, discrete approximation of the mixture method for learning alternatives in the anytime-valid inference literature. While we focus on a finite set of candidates with uniform weighting for simplicity and efficiency, one could also incorporate informative priors or adopt continuous mixture models for greater flexibility.
In this section, we illustrate how testing for a treatment effect in the interactive fixed effects (IFE) model by bai2009panel can be recast to testing for exchangeability. We first present the IFE, and then state conditions under which difference-in-differences (DiD) and the synthetic control method (SCM) generate exchangeable treatment estimates. The exchangeability conditions are identical to the conditions imposed by abadie2021synthetic and chernozhukov2021exact. If these are not strictly met—for instance, due to serial correlation in treatment estimates—we adopt a block structure, as recommended by chernozhukov2018exact.
The IFE model is a common way to model treatment effects for the synthetic control estimator as it can capture violations of the parallel trend assumption. Consider the following panel for the outcomes of $N+1$ units, $i = 1, \dots, N+1$, of which only the first unit $i = 1$ is treated after time $T_0$,
Here, $I[\cdot]$ is the indicator function, $\bm{\mu}_i$ and $\bm{\lambda}_t$ denote the $r$-dimensional factor loadings and time-varying factors, $\tau_t$ is the treatment effect at time $t$, and $\varepsilon_{it}$ is mean zero noise. The $\ell$-dimensional covariates of unit $i$ are denoted with $\bm{Z}_i$ and its corresponding time-varying parameter is denoted with $\bm{\theta}_t$. This specification allows units to have different responses to time trends and thereby violate the parallel trends assumption.
We follow abadie2021synthetic by splitting the pre-treatment observations into a set of blank periods $\mathcal{B}=\{1,...,T_{\mathcal{B}}\}$ and a set of training periods used for estimation $\mathcal{E}=\{1,...,T_0\}\setminus\mathcal{B}$. The training periods are used for estimating the pre-treatment means (DiD) and weights (SCM), and the blank periods are compared to the post-treatment observations. The fact that we do not use these blank periods for estimating weights, helps ensuring exchangeability of the blank- and post-treatment periods under the null.
We are interested in conducting inference on the treatment effect $\tau_t$. Specifically, we want to test that there is no treatment effect,
In the following two sections, we show under which conditions DiD and SCM generate treatment estimates $\widehat{\tau}_t$, such that $H_0^\text{treatment}$ can be reduced to the hypothesis studied in Sections (ref) and (ref):
Suppose all units respond equally to changes in the time-varying factors $\bm{\lambda}_t$. In the DiD literature, this is also known as the parallel trends assumption. This corresponds to the IFE model with no covariates ($\ell=0$) and, without loss of generality,
The goal of DiD is to estimate treatment effects for the blank pre-treatment periods $\mathcal{B}$ and the post-treatment observations $\{T_0+1,...,T\}$. The DiD estimator does this by filtering out the unit- and time-fixed effects by double differencing,
where $\bar{Y}_t$ is the mean outcome of control units at time $t$, $\bar{Y}_1$ is the average outcome of unit 1 over training periods $\mathcal{E}$, and $\bar{Y}$ is the average outcome of all control units over training period $\mathcal{E}$.
For the IFE model under the parallel trends assumption, the treatment estimator equals
where $\bar{\varepsilon}_t$, $\bar{\varepsilon}_1$, and $\bar{\varepsilon}$ are the respective error terms of $\bar{Y}_t$, $\bar{Y}_1$, and $\bar{Y}$. The following proposition states that exchangeability of $\bm{\varepsilon_t}$ implies exchangeability of this DiD estimator under $H_0$.
This proposition is a special case of Proposition (ref), which is presented next in Section (ref).
In the more general case where the parallel trends assumption is violated, one can use the synthetic control method abadie2003economic. This technique attempts to filter out the interactive fixed effects and time-varying parameter $\theta_t$ by constructing a synthetic control unit that approximates the factor loadings and covariates of the treated unit. Specifically, it first estimates an $N$-dimensional vector of linear weights $\widehat{\bm{w}}=[\widehat{w}_2,...,\widehat{w}_{N+1}]$ based on the training periods $\mathcal{E}$, and then estimates the synthetic outcomes as $\sum_{i=2}^{N+1} \widehat{w}_iY_{it}$. The treatment effect estimator is then the difference between the outcome of the treated unit and the synthetic outcome,
The SCM weights $\widehat{\bm{w}}$ are usually constructed as follows. Let $\bm{X}_1$ be a $K$-dimensionional vector that contains $K$ characteristics based on the pre-treatment data in $\mathcal{T}$, and collect the same characteristics for the $N$ control units in the $K\times N$ matrix $\bm{X}_0$. For a given non-negative diagonal matrix $\bm{V}$ that determines the importance of these characteristics\footnote{See abadie2010synthetic for a data-driven approach to determine $\bm{V}$.}, $\widehat{\bm{w}}$ minimizes the weighted Euclidean distance between the pre-treatment characteristics of the treated and the synthetic control unit,
where $\Delta_N$ denotes the $N$-dimensional simplex.
To show how a properly chosen $\widehat{\bm{w}}$ can lead to an (approximately) exchangeable treatment estimator, we decompose the treatment effect estimator presented in (ref),
It is easy to verify whether the covariates are accurately reconstructed by the synthetic control unit ($\bm{Z}_1 \approx \sum_{i=2}^{N+1}\bm{Z}_i\widehat{w}_i$) as the covariates are observable. This is more difficult for the unobserved factor loadings $\bm{\mu}_i$. abadie2010synthetic show however, that under mild conditions, when the $K$ characteristics include the pre-treatment outcomes, the weights $\widehat{\bm{w}}^{(T_{\mathcal{E}})}$ asymptotically reconstruct the factor loadings of the treated unit:
Unfortunately, in finite samples, the SCM does not exactly reconstruct the factor loadings and hence we require stronger assumptions for exchangeability than just the exchangeable error term assumption of DiD. The following result restates Theorem 2 of abadie2021synthetic for the conditions on exchangeability of the SCM estimator.
In some cases, size distortions due to violations of the exchangeability conditions can be reduced. If the treatment estimators are expected to be serially dependent, imposing a block-structure on the treatment estimators can reduce some of this dependence chernozhukov2018exact. This is done by partitioning the blank periods and post-treatment periods into blocks of $B$ observations (assuming for simplicity that the number of blank periods $T_{\mathcal{B}}$ is divisible by $B$). For each block, we then take the block-wise means of the treatment estimators and perform the procedure with these block-wise treatment estimators. It is important to note that these blocks also reduce our collection of stopping times; instead of being able to reject after every single observation, we now can only evaluate our test statistic after every $B$ observations.
We start by illustrating the potential benefits of anytime-valid inference in a stylized setting for the DiD estimator. A less stylized simulation setup is presented in Section (ref). In our first setting, we assume the conditions of Proposition (ref) hold, such that the DiD estimator produces exchangeable treatment estimators under $H_0$. We compare the anytime-valid tests with the fixed-$T$ test by abadie2021synthetic and show how anytime-valid tests are preferred for a wide range of intertemporal preference profiles.
More specifically, we draw from the IFE model in (ref), excluding covariates ($\ell=0$) or fixed effects ($r=0$). We set $\bm{\varepsilon}_t \overset{\text{i.i.d}}{\sim} N(0,I)$. The anytime-valid tests are based on the reduced ranks, and we use the plug-in and Gaussian statistics as outlined in Sections (ref) and (ref), initializing the plug-in method at $t = T_0 + 1$ with the optimal test statistic of the Gaussian alternative. For now, we correctly specify the effect size of this Gaussian alternative. These anytime-valid tests are compared to the one-sided fixed-$T$ test by abadie2021synthetic (see Appendix (ref)). As the exchangeability conditions hold exactly, we use a block size of $B=1$.
We compare the rejection rates of the two anytime-valid tests to that of the not anytime-valid fixed-$T$ test. Clearly, these are not direct competitors as the fixed-$T$ test is not anytime-valid, but this comparison allows us to study how much power we lose in exchange for the anytime-validity.
We consider two variations of the fixed-$T$ test: the single fixed-$T$ test where we correctly evaluate the fixed-$T$ $p$-value at a single time $T$, and the repeated fixed-$T$ test where we naively apply the fixed-$T$ methodology sequentially by rejecting as soon as the $p$-value dips below $\alpha$. The former controls the Type I error but can only be executed once at $t=T$, whereas the latter naively applies the fixed-$T$ methodology at every time $t$ but does not control size.
Figure (ref) illustrates the rejection rates of these tests over time ($\alpha = 0.05$) under $H_0$ and $H_1: \tau_t = 1.5$ for $t > T_0$. The first thing we see in Figure (ref) is that repeatedly conducting the fixed-$T$ test leads to a large size distortion, reaching up to 20%. As shown by the `single fixed-$T$' line, applying this test only once ensures size control. Both anytime-valid tests maintain proper size control throughout the post-treatment period. We extended this simulation for up to 1000 post-treatment observations, and their size remained below $\alpha = 0.05$.
The rejection rates under the alternative are shown in Figure (ref). The power of the repeated fixed-$T$ test should be ignored, as it comes at the costs of high size distortion. The power of the single fixed-$T$ method is next highest, but this power can of course only be feasibly attained if we had pre-specified the number of observations to be equal to $T$. At a given time $T$, we see that the anytime-valid methods pay a price in terms of power. In exchange, a single anytime-valid method attains the displayed power at every time: not just at a single pre-specified time $T$. We highlight this difference in Section (ref).
Among the anytime-valid tests, the test with Gaussian alternative demonstrates greater power compared to the plug-in method. This is expected, as it is log-optimal in this setting, whereas the plug-in method must learn the distribution. Both anytime-valid tests converge to full power as the number of post-treatment observations grows.
For a fair comparison between the anytime-valid tests and the fixed-$T$ test, we now compare our tests to a single fixed-$T$ test. We find that our anytime-valid tests can benefit greatly from being able to reject early or continue testing. Figure (ref) presents the rejection rates over time for the anytime-valid test with Gaussian alternative and that of a fixed-$T$ test performed after either 3 (Figure (ref)) or 6 (Figure (ref)) post-treatment observations. We highlight differences in rejection rates by shading regions where the anytime-valid test is preferred in dark grey, and the other regions in light grey.
In Figure (ref), we study a fixed-$T$ test that is executed after 3 post-treatment observations: $T = T_0 + 3$. We see that most of the advantage of the anytime-valid method stems from rejections occurring after $T$, but it can also benefit from rejecting early (before $T$). The fixed-$T$ test has higher rejection rates only within a narrow band of the number of post-treatment observations.
In Figure (ref), we make the same comparison for a fixed-$T$ test that rejects after 6 post-treatment observations. Most of the benefit of the anytime-valid test now comes from the ability to reject earlier than the fixed-$T$ test, though some advantage remains for rejections occurring later.
Comparing the volume of the dark gray and light gray regions in Figure (ref) suggests that this anytime-valid test offers some intertemporal advantage over the fixed-$T$ tests. We study this in more detail in the following section.
In practice, one may have intertemporal preferences due to, the (opportunity) costs involved with conducting an experiment. To capture these preferences over time, we consider a simple discounted version of power.
Suppose $H_0$ is false, and a decision-maker derives utility from knowing this at $t > T_0$. Consider a normalized discounted utility model, where the expected utility $U_S$ of test $S$ is a discounted sum of its rejection probabilities:
for some scalar discount factor $\delta \in (0,1]$. When $\delta=1$, the difference in utility between the anytime-valid and fixed-$T$ is simply the difference in volume of the dark grey area and the light grey area in Figure (ref). A smaller $\delta$ implies that the volume of earlier shaded regions is weighted more heavily. One economic interpretation of $\delta$ is that it represents the (opportunity) costs of continuing the study. A smaller $\delta$ suggests that future rejections carry higher costs, making them less important for the decision-maker.
As seen in Section (ref), the intertemporal advantage of an anytime-valid test over a fixed-$T$ test varies with its rejection time $T$. We now demonstrate that for a wide range of discount factors $\delta$, the anytime-valid tests are preferred over almost every possible fixed-$T$ test. In practice, the $T$ for which the fixed-$T$ test attains the highest discounted utility is unknown a priori, further supporting the use of anytime-valid tests.
In Figure (ref), we compare the utility of different fixed-$T$ tests with the utility of the anytime-valid tests. Each $x$-axis tick corresponds to a fixed-$T$-test executed only at that specific time $T$. We shade the regions of discount factors $\delta$ where fixed-$T$ tests performed at time $T$ are preferred over the anytime-valid tests. Figure (ref) shows that for $\delta \geq 0.8$, the anytime-valid test with Gaussian alternative dominates all fixed-$T$ tests. Hence, for a (mildly) patient decision-maker, there does not exist any $T$ at which it would have been better to perform a fixed-$T$ test. For low discount factors, the fixed-$T$ test with few post-treatment observations is preferred.
In contrast, Figure (ref) shows that the plug-in test is outperformed by a substantial range of fixed-$T$ tests for all values of $\delta$. However, as opposed to the fixed-$T$ test, this plug-in method requires no prior knowledge on the effect size. This suggests that in this setting, the plug-in method is only beneficial when there is high uncertainty on the actual effect size.
In Section (ref), we considered a stylised setting, in which the treatment estimates were exactly exchangeable. In this setting, we illustrate our methodology on SCM and simultaneously demonstrate the robustness under possible misspecification. We show how size distortion, caused by violations of the conditions in Proposition (ref), can be effectively reduced by employing a block structure. In addition, we consider a setting where the anytime-valid test with Gaussian alternative is based on the incorrect effect size specification. There, we show how adaptive mixtures across multiple alternatives can enhance power. Lastly, we also consider a setting with a dynamic treatment effect, which has a varying strength over time.
In this setup, we consider 3 time-varying factors ($r=3$) and generate the time-varying factors and errors through a VAR(1) process defined as follows:
where $\rho_\lambda, \rho_\varepsilon \in (-1,1)$ are the autoregressive parameters of $\bm{\lambda}_t$ and $\bm{\varepsilon}_t$, and we set $\sigma^2=1$. Proposition (ref) implies that the tests are exactly valid when $\rho_\lambda = \rho_\varepsilon = 0$. Since the rejection rates over time and intertemporal preferences in this exactly valid case are similar to the DiD setting (see Appendix (ref)), we immediately turn to the effect of non-exchangeability on the size.
We examine the size distortion of three tests: a fixed-$T$ test carried out after 12 post-treatment observations, and two reduced-rank anytime-valid tests based on 30 post-treatment observations. The sample of the anytime-valid tests is deliberately set larger as its size control should also hold after $T=12$.
Table (ref) presents the size of these tests for different sample dimensions and values of $\rho_\lambda$ and $\rho_\varepsilon$, imposing a block-structure ($B=3$). We also consider the case where $B=1$ in (ref), where we find settings with high size distortions, but these are similar across fixed-$T$ and anytime-valid tests\footnote{As we impose a block structure of $B$, we rescale the pre- and post-treatment sample lengths with $B$ to ensure a fair comparison between $B=1$ and $B=3$.}.
As expected, when the autoregressive components are equal to 0 ($\rho_\lambda = \rho_\varepsilon = 0$), all tests maintain size control across all sample dimensions. The anytime-valid tests are slightly undersized for this finite horizon. This is not surprising, as they are valid for any final horizon $T$. When the autoregressive component of the time-varying factors $\rho_\lambda$ increases, we observe larger sizes for all three tests. These distortions are largest when $\rho_\lambda$ is large and the pre-treatment sample size is small ($T_0 = 20$). In these small samples, the SCM estimator does not adequately reconstruct the factor loadings of the treated unit, leading to non-exchangeability of the SCM estimates. When $T_0$ is large, most tests exhibit relatively low size distortion. This especially holds for the Gaussian anytime-valid test (see column 6 of Table (ref))
When the idiosyncratic serial correlation ($\rho_\varepsilon$) grows, the size increases, regardless of the dimension of the pre-treatment sample. This is because the SCM only filters out variation that is shared across units, and these idiosyncratic shocks are unit-specific by construction. In practice, one often assumes that the idiosyncratic components $\varepsilon_{it}$ have some common variation across units and their serial dependence can therefore be captured by the unobserved time-varying factors.
Overall, the size distortions observed for the fixed-$T$ test and the anytime-valid tests do not differ substantially. However, between the two anytime-valid tests, the plug-in method appears less robust than the Gaussian alternative. An explanation for this is that the flexibility of this plug-in method also makes it more sensitive to violations of exchangeability. As the plug-in method adaptively learns the distribution of the ranks, it gradually learns in what way the exchangeability is violated and attempts to exploit this.
{3.5pt}
In the previous simulations, we used the true effect size for the Gaussian anytime-valid method. We now study the sensitivity of the power to a possible misspecification of the effect size. The key finding is that the power can vary quite strongly if the effect size is misspecified, and that this can be overcome by taking a mixture over different plausible effect sizes.
Using the same IFE setup with $\rho_\lambda=\rho_\varepsilon=0$, Figure (ref) shows the power of the anytime-valid test with Gaussian alternative when the true effect size is off by a factor $c > 0$. For $c > 1$, this means that we test against an effect size that is larger than the true effect size. For $c < 1$, we test against an effect size that is smaller than the true one.
We find that if $c > 1$, the rejection rate is higher at the beginning, but it grows less quickly later on. The tests with smaller $c$ start out less powerful, but grow more powerful later on. This illustrates that different $c$ correspond to different patience preferences.\footnote{koning2024continuous discusses how in a Gaussian location model, these different effect sizes correspond to optimal $e$-values for different power targets.}
In practice, the actual effect size is often unknown. Let $C$ denote a countable set of candidate $c$'s. Proposition (ref) in Appendix (ref) implies that if $1\in C$, then the average over the martingales implied by $C$ has asymptotic log-optimal growth. Since this average is actually an adaptive weighted average over $e$-values, we call this strategy the adaptive mixture. Another possibility is to take a uniform average over all the candidate $e$-values and constructing a martingale through their running product. We call this strategy the average mixture.
We compare our mixture strategies, with $C=\{0.25,0.5,1,2,4\}$, to the log-optimal rejection rates in Figure (ref). Both weighting strategies are less powerful than the log-optimal strategy, which is explained by the inclusion of the sub-optimal martingales. After $T=10$, the adaptive mixture starts outperforming the average mixture and its power approaches that of the log-optimal strategy.
In practice, the treatment effect may be dynamic, with which we mean that the treatment effect may change over time. In this section, we consider a setting where the treatment effect grows linearly over time $\tau_t=1 +\frac{t-T_0}{15}$. We compare the fixed-$T$ test, plug-in method with sequential ranks, and an adaptive mixture (Section (ref)) over reduced rank anytime-valid Gaussian tests with $C=\{0.25,0.5,1,2,4\}$. In short, the results show that although the anytime-valid methods remain effective, their advantage over the fixed-$T$ test is slightly reduced.
Figure (ref) presents the intertemporal preference plots for both of the anytime-valid tests. The anytime-valid test with Gaussian alternative is still preferred over almost all fixed-$T$ tests for discount factors $\delta>0.9$. This is slightly weaker than the constant treatment effect, where this held for all fixed-$T$ tests. The plug-in test statistic seems to perform less well. It is dominated by a large set of fixed-$T$ tests for every discount factor $\delta$. An explanation for this is that the distribution of the ranks is not stable, which makes it difficult to effectively learn the distribution of the ranks. One important note is that the inherent uncertainty regarding the dynamics of the treatment effect simultaneously make it more difficult to specify at which $T$ to test for a fixed-$T$ test. Hence, the increased range for which the fixed-$T$ is preferred over these anytime-valid tests compared to the instantanuous setting is partially offset by the additional insecurity in specifying $T$.
This paper introduces anytime-valid testing for program evaluation. We construct anytime-valid $p$-values and derive conditions under which these $p$-values have optimal shrinkage rates. In the case of composite alternatives, we propose two strategies: either learning the most promising alternative, or transforming a composite alternative on the outcomes to a simple alternative on the ranks. We show how these methods are often preferred over a standard test that requires a pre-specified number of observations $T$ for a simple intertemporal discounted notion of power.
In future work, we plan to more extensively study the setting with a dynamic treatment effect. Such settings are challenging in the fixed-$T$ setting because the correct value for $T$ is harder to find, and challenging in the sequential setting as it makes the alternative hard to learn / pre-specify. A fruitful idea may be to design techniques that are tailored to sequentially learning dynamic treatment effects.
Finally, we note that there are multiple techniques to further increase power based on test martingales. For a given significance level $\alpha$, fischer2024improving show that the power of an anytime-valid test can be slightly, but almost surely, improved by avoiding `overshooting'. This is done by capping the maximum of our test statistics such that the martingale never exceeds $\frac{1}{\alpha}$, thereby strictly increasing the density of the $e$-value in other regions. In addition, if after a certain time $T$, one wishes to terminate the experiment, randomized Markov's inequality can be used to have an almost surely lower critical value at time $T$ ramdas2023randomized.