EconBase
← Back to paper

Limitations of Randomization Tests in Finite Samples

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

25,139 characters · 9 sections · 7 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Limitations of Randomization Tests in Finite Samples

\singlespacing

\abstract{ Randomization tests deliver exact finite-sample Type 1 error control when the null satisfies the randomization hypothesis. In practice, achieving these guarantees often requires stronger conditions than the null hypothesis of primary interest. For example, sign-change tests of mean zero require symmetry and need not control finite-sample size for non-symmetric mean-zero distributions. We investigate whether the mismatch between the null and the invariance conditions required for exactness reflects the use of particular transformations or a more fundamental limitation. We provide a simple necessary and sufficient condition for a null hypothesis to admit a randomization test. Applying this framework to one-sample problems, we characterize the nulls that admit randomization tests on finite supports and derive impossibility results on continuous supports. In particular, we show that several common nulls, including mean zero, do not admit randomization tests. We further show that, among one-sample tests using linear group actions, the admissible nulls are limited to subsets of symmetric or Gaussian distributions. These results confirm that the absence of exact finite-sample validity is inherent for many commonly studied nulls and that practitioners using existing tests are not foregoing feasible exact alternatives. }

\onehalfspacing

Introduction

Consider the problem of testing a null hypothesis given a finite sample of data. If this null satisfies a group invariance property referred to as the “randomization hypothesis,” then one can construct randomization tests that obtain exact finite-sample Type 1 error control. Intuitively, this property requires that, under the null, certain transformations of the data leave its distribution unchanged. In part due to the appeal of finite-sample validity, there is a large methodological literature on randomization tests and a similarly large applied literature using these methods ritzwoller2024randomization.

In practice, randomization tests often rely on invariance conditions that are stronger than the null hypothesis of primary interest. This is a problem because it re-introduces an inability to control Type 1 error: the test can over-reject when the data is drawn from a distribution that satisfies the hypothesis of primary interest but not the stronger conditions. For example, a commonly used randomization test for the null of mean zero with i.i.d. observations is based on sign changes. This test controls finite-sample size only under symmetry and may over-reject for non-symmetric mean-zero distributions.

This paper investigates whether the failure to control Type 1 error for certain null hypotheses (e.g. the null of mean zero) is due to the specific randomization tests used in practice (e.g. the use of sign changes) or from a more fundamental limitation. To formalize this distinction, we develop a framework that characterizes which null hypotheses admit randomization tests. We provide a simple necessary and sufficient condition for a null to satisfy the randomization hypothesis that avoids group-theoretic structure.

We then apply this framework to one-sample problems. We show that certain null hypotheses---such as the null of mean zero---do not admit randomization tests that achieve exact finite-sample validity. These results relate to classical non-existence results for nonparametric inference bahadur1956nonexistence,lehmann1990pointwise,romano2004non. By focusing on randomization tests, we derive conditions for the existence and non-existence of exact finite-sample procedures across a range of null hypotheses and distributional classes. These results imply that existing applications of randomization tests, such as sign-change tests, are not overlooking alternative procedures that would achieve exact finite-sample validity.

We also provide guidance for constructing randomization tests when standard approaches are not applicable. On finite supports, we give an explicit characterization of nulls that admit randomization tests. On continuous supports, we study tests based on linear group actions. We show that, in one-sample settings, such tests are very limited: the admissible nulls correspond to subsets of symmetric or Gaussian distributions.

As part of identifying whether a null admits a randomization test, we also show how our framework can be used to construct these tests. We do so by providing explicit examples --- including a construction of a normality test that makes use of rotation symmetries --- that may be of theoretic or applied interest.

Taken together, these results yield a positive implication for applied work: the use of standard randomization tests does not overlook alternative procedures that achieve exact finite-sample validity for many commonly studied nulls. They also affirm and motivate a focus on asymptotic properties of randomization tests (see, e.g., romano1989bootstrap,romano1990behavior,diciccio2017robust,canay2017randomization,lei2021assumption,pouliot2024ttest,bai2024inference). Finally, our results suggest that one-sample randomization tests that go beyond symmetry or normality must rely on non-linear transformations.

The remainder of the paper is organized as follows. Section (ref) reviews the randomization test procedure and develops our framework. Section (ref) studies one-sample tests with finite support. Section (ref) considers continuous supports. Section (ref) concludes.

The randomization test

Construction of the randomization test

The construction follows lehmannromano1986; see zhang2023randomization for a recent formulation. Let $X = (X_1,\dots,X_n)$ be data taking values in $\ensuremath{\mathcal{X}}$ with distribution $P \in \Omega$. Consider the problem of testing $H_0 : P \in \Omega_0$, and suppose that there is a group $\ensuremath{\mathbb{G}}$ of measurable transformations $g:\ensuremath{\mathcal{X}}\to\ensuremath{\mathcal{X}}$ such that $gX \stackrel{d}{=} X$ whenever $X \sim P \in \Omega_0$. For ease of exposition, we assume $\ensuremath{\mathbb{G}}$ is finite with size $M$.

Let $T:\ensuremath{\mathcal{X}}\to\mathbb{R}$ be any test statistic used to test the null. For every $x \in \ensuremath{\mathcal{X}}$, let $T^{(1)}(x;\ensuremath{\mathbb{G}}) \leq \dots \leq T^{(M)}(x;\ensuremath{\mathbb{G}})$ be the ordered values of $T(gx)$ as $g$ ranges over $\ensuremath{\mathbb{G}}$. Given nominal level $\alpha \in (0,1)$, let $k \equiv M - \lfloor M\alpha \rfloor,$ where $\lfloor M\alpha \rfloor$ denotes the largest integer less than or equal to $M\alpha$. Let $M^{+}(x)$ and $M^{0}(x)$ be the number of values $T^{(j)}(x)$ that are greater than $T^{(k)}(x)$ and equal to $T^{(k)}(x)$, respectively. Set

equation*[equation* omitted — 65 chars of source]

Define the randomization test as

equation[equation omitted — 316 chars of source]

The following Theorem shows that the randomization test has nominal level $\alpha$. The proof we present is slightly different than the proof in lehmannromano1986, in part due to our focus on the role of the group structure $\ensuremath{\mathbb{G}}$.

theoremSuppose that for any $g \in \ensuremath{\mathbb{G}}$, $gX \stackrel{d}{=} X$ whenever $X \sim P \in \Omega_0$. Then $\operatorname{\mathbb{E}}_P[\phi(X;\ensuremath{\mathbb{G}})] = \alpha$ $\forall \; P \in \Omega_0$.
proof[Proof of Theorem (ref)] Consider any $X \sim P \in \Omega_0.$ Then \begin{align*} \operatorname{\mathbb{E}}_P[\phi(X;\ensuremath{\mathbb{G}})] = \frac{1}{M} \operatorname{\mathbb{E}}_P \left[ \sum_{g \in \ensuremath{\mathbb{G}}} \phi(X;\ensuremath{\mathbb{G}})\right] = \frac{1}{M} \operatorname{\mathbb{E}}_P \left[ \sum_{g \in \ensuremath{\mathbb{G}}} \phi(gX;g\ensuremath{\mathbb{G}} )\right], \end{align*} where the last equality used the form of $\phi$ as in (ref) and that $X \stackrel{d}{=} gX$ for all $g \in \ensuremath{\mathbb{G}}$. By construction, $\frac{1}{M} \sum_{g \in \ensuremath{\mathbb{G}}} \phi(gx;\ensuremath{\mathbb{G}}) = \alpha$ for any $x \in \ensuremath{\mathcal{X}}$, so that $\frac{1}{M} \operatorname{\mathbb{E}}_P\left[ \sum_{g \in \ensuremath{\mathbb{G}}} \phi(gX;\ensuremath{\mathbb{G}})\right] = \alpha. $ To complete the proof, it thus suffices to show that \begin{equation} \operatorname{\mathbb{E}}_P\left[ \sum_{g \in \ensuremath{\mathbb{G}}} \phi(g X;\ensuremath{\mathbb{G}})\right] = \operatorname{\mathbb{E}}_P \left[ \sum_{g \in \ensuremath{\mathbb{G}}} \phi(gX;g \ensuremath{\mathbb{G}})\right]. \end{equation} This follows from the group property that $g \ensuremath{\mathbb{G}} = \ensuremath{\mathbb{G}}$ for all $g \in \ensuremath{\mathbb{G}}$, and thus $\phi(gx;\ensuremath{\mathbb{G}}) = \phi(gx;g\ensuremath{\mathbb{G}})$ for all $x$ and $g \in \ensuremath{\mathbb{G}}$, so that (ref) holds.
remarkSuppose that we did not require that the collection of invertible transformations $\ensuremath{\mathbb{G}}$ is a group. For (ref) to hold generally, we require that $\ensuremath{\mathbb{G}} g= \ensuremath{\mathbb{G}}$ for all $g\in \ensuremath{\mathbb{G}}$. A basic property of groups is that if the set $\ensuremath{\mathbb{G}}$ is a collection of invertible transformations, then $\ensuremath{\mathbb{G}} g= \ensuremath{\mathbb{G}} $ for all $g \in \ensuremath{\mathbb{G}}$ if and only if $\ensuremath{\mathbb{G}}$ is a group (with composition as the operation). In this sense, the requirement that $\ensuremath{\mathbb{G}}$ is a group is both necessary and sufficient.

Framework

Having established the construction, we now develop a simple framework to examine which null hypotheses admit a group that is invariant to all distributions in the null.

definition[Randomization hypothesis] We say $\Omega_0$ satisfies the randomization hypothesis if there exists a group $\ensuremath{\mathbb{G}}$ of measurable transformations $g:\ensuremath{\mathcal{X}}\to\ensuremath{\mathcal{X}}$ such that \begin{equation} \Omega_0 \subseteq \Omega_{\ensuremath{\mathbb{G}}} \subset \Omega, \end{equation} where $\Omega_{\ensuremath{\mathbb{G}}} \equiv \{P \in \Omega \; : \; g X \stackrel{d}{=} X \; \forall \; g \in \ensuremath{\mathbb{G}} \ , \ X \sim P\}.$
remarkThe set inclusions in (ref) enforce that $\Omega_0$ is invariant to $\ensuremath{\mathbb{G}}$ while also ensuring that $\ensuremath{\mathbb{G}}$ is not invariant to all distributions in $\Omega$. The latter condition rules out trivial groups (e.g. the identity) that are invariant to all distributions.

Our analysis will focus on characterizing which null hypotheses do and do not satisfy the randomization hypothesis. For those that do, we will be interested in constructing a group to implement the test. The following result provides a tool to answer both of these questions in a way that avoids group-theoretic concepts.

proposition$\Omega_0$ satisfies the randomization hypothesis if and only if there exists a bijective measurable function $f: \ensuremath{\mathcal{X}} \to \ensuremath{\mathcal{X}}$ such that \begin{equation} \Omega_0 \subseteq \Omega_f \subset \Omega, \end{equation} where $\Omega_f \equiv \big\{P \in \Omega \; : \; f X \stackrel{d}{=} X, X \sim P\big\}.$ When this holds, a randomization test can be constructed by using the group generated by $f$.

All omitted proofs can be found in Supplementary Appendix (ref). The proof of the forward direction is immediate and the proof of the reverse direction follows from showing that if $f$ is invariant to all distributions in $\Omega_0$, then so is $f^{-1}$. We can then show that $\Omega_0$ satisfies the randomization hypothesis using the group generated by $f$.

Proposition (ref) shows that to study which null hypotheses satisfy the randomization hypothesis, it suffices to consider individual bijective functions. Given such a function satisfying (ref), a randomization test can be implemented using the group it generates.

In the remainder of the paper, we study one-sample problems with $X = (X_1,\dots,X_n)$ with $X_i$ i.i.d. from $P \in \Omega$, supported on a subset of the real line. Here, $\Omega$ denotes the set of distributions for $X_i$.

One-sample tests on a finite support

In this section, we suppose $X_i$ are drawn i.i.d. from $P$ supported on a finite set $\mathcal{A} \equiv \{\alpha_1,\dots,\alpha_K\} \subset \mathbb{R}$. We represent $P$ by its probability mass function $p: \mathcal{A} \to [0,1]$ and define

equation[equation omitted — 136 chars of source]

By Proposition (ref), a null hypothesis $\Omega_0 \subset \Omega_{\text{discrete}}$ satisfies the randomization hypothesis if and only if there exists a bijection $f: \mathcal{A}^n \to \mathcal{A}^n$ such that $\Omega_0 \subseteq \Omega_f \subset \Omega_{\text{discrete}}$. On a finite support, we can explicitly write $\Omega_f$ as

equation[equation omitted — 182 chars of source]

Each of the individual sets in this intersection corresponds to the invariances of simpler functions that equal $f$ for a single value and otherwise equal the identity. This provides an explicit characterization of Proposition (ref) for one-sample tests over finite supports.

proposition$\Omega_0 \subset \Omega_{\text{discrete}}$ satisfies the randomization hypothesis if and only if there exists $x,y \in \mathcal{A}^n$ such that $y$ is not a permutation of $x$ and such that $\prod_{j=1}^n p(x_j) = \prod_{j=1}^n p(y_j)$ for all $p \in \Omega_0$.

Proposition (ref) is useful for two reasons. First, it can be used to identify bijective functions that satisfy the invariance. We can then construct randomization tests by taking the groups generated by the functions. To illustrate this, Example (ref) of Supplementary Appendix (ref) uses the result to construct the usual test of symmetry around zero and to construct a test of two points having equal mass.

Second, Proposition (ref) shows that null hypotheses satisfying the randomization hypothesis are characterized by equal probability mass assigned to distinct configurations (accounting for multiplicities). Null hypotheses not defined by such restrictions therefore do not satisfy the randomization hypothesis. In what follows, we apply Proposition (ref) to show that null hypotheses defined by $k$-th moments (including the mean) and quantiles do not satisfy the randomization hypothesis. These results illustrate how the proposition can be used to determine whether a null hypothesis admits a randomization test.

propositionIf $|\mathcal{A}| > 4$, then for any positive integer $t$ and $\beta \in (\alpha^{t}_1,\alpha^{t}_K)$, $\Omega_0 = \{P \in \Omega_{\text{discrete}} : \operatorname{\mathbb{E}}_P[X_i^{t}] = \beta\}$ does not satisfy the randomization hypothesis.
propositionIf $|\mathcal{A}| > 2$, then for any $p \in (0,1)$ and $q \in (\alpha_1,\alpha_K)$, $\Omega_0 = \{P \in \Omega_{\text{discrete}} : \operatorname{\mathbb{P}}[X_i \leq q] = p\}$ does not satisfy the randomization hypothesis.

The proofs of Propositions (ref) and (ref) proceed by showing that for any $x,y \in \mathcal{A}^n$ that are not equal up to permutation, one can construct a distribution in the null that assigns different probabilities to $x$ and $y$. The result then follows from Proposition (ref).

remarkThe cardinality requirements in Propositions (ref) and (ref) exclude edge cases. For example, when $|\ensuremath{\mathcal{A}}| =3$ and $\mathcal{A} = \{-L,0,L\}$, the null of mean zero coincides with symmetry, which satisfies the randomization hypothesis via sign changes. The proof of Proposition (ref) provides an analogous edge case for $|\ensuremath{\mathcal{A}}| = 4$. For Proposition (ref), the result continues to hold when $|\mathcal{A}| = 2$ except at $p = 1/2$, where sign changes yield a valid randomization test.

One-sample tests on a continuous support

In this section, we suppose $X_i$ is drawn i.i.d. from $P$ supported on the real line. We first characterize all null hypotheses that are invariant to linear group actions. We then consider arbitrary groups and develop a necessary condition for a null hypothesis to satisfy the randomization hypothesis.

Groups of linear transformations

A group of linear transformations $\ensuremath{\mathbb{G}}$ is one in which each element $g$ corresponds to a linear map $A_g:\mathbb{R}^n \to \mathbb{R}^n$. The following result characterizes the distributions invariant to such groups.

propositionLet $P \in \Omega_{\text{cts}}$, where $\Omega_{\text{cts}}$ is a subset of distributions supported on the real line. Let $\ensuremath{\mathbb{G}}$ be a group of linear transformations acting on $\mathbb{R}^n$. Then exactly one of the following is true: 1. $\Omega_{\ensuremath{\mathbb{G}}} = \Omega_{\text{cts}}$, 2. $\Omega_{\ensuremath{\mathbb{G}}} = \{P \in \Omega_{\text{cts}} : \text{$P$ symmetric about $0$}\}$, 3. $\Omega_{\ensuremath{\mathbb{G}}} \subseteq \{N(\mu,\sigma^2) \in \Omega_{\text{cts}} : \mu \in \mathbb{R},\sigma^2 >0 \}$, or 4. $\Omega_{\ensuremath{\mathbb{G}}} = \emptyset$.

The proof proceeds by writing $gX = AX$ for some matrix $A$. If $A$ is diagonal, invariance requires each coordinate to be symmetric about zero. If $A$ is not diagonal, then at least one component of $AX$ is a non-trivial linear combination of multiple coordinates of $X$. Since both $AX$ and $X$ consist of i.i.d. components, the Darmois-Skitovich theorem implies that $X_i$ must be Gaussian.

Proposition (ref) shows that, when restricted to linear group actions, one-sample randomization tests extend permutation-based invariance in limited ways. In particular, they can test membership in the class of distributions symmetric about zero or in subsets of the Gaussian family. The former can be implemented using sign changes. Example (ref) in Supplementary Appendix (ref) illustrates the latter.

Arbitrary groups

We now allow $\ensuremath{\mathbb{G}}$ to be arbitrary. The classical result of bahadur1956nonexistence shows that sufficiently rich classes of distributions cannot be tested in a nonparametric setting. In our context, a related conclusion holds under a local version of this richness condition.

Let $\Omega_{\text{cts}}(\mathcal{B})$ be the set of distributions supported on $\mathcal{B}\subseteq \mathbb{R}$. Informally, a null is locally-dense if it can locally replicate any density on finitely many small intervals (up to scale).

definition[Locally-dense] A set $\Omega_0 \subset \Omega_{\text{cts}}(\mathcal{B})$ is locally-dense if there exists $L:\mathbb{N} \to \mathbb{R}_+$ such that for any collection of intervals $I_1,\dots,I_m$ with $|I_i| < L(m)$ and any density $p$ supported on $\bigcup_{i=1}^m I_i$, there exist $H \in \Omega_0$ with density $h$ and $\alpha \in (0,1]$ such that \[ h(x) = \alpha p(x) \quad \text{for all } x \in \bigcup_{i=1}^m I_i. \]

The following result shows that locally-dense null hypotheses do not satisfy the randomization hypothesis.

propositionIf $\Omega_0$ is locally-dense, then $\Omega_0$ does not satisfy the randomization hypothesis.

The proof uses Proposition (ref) and shows that for any function $f$ with $\Omega_0 \subseteq \Omega_f$, local denseness implies $\Omega_f = \Omega$. Intuitively, local denseness ensures that the null can approximate any density on finitely many regions, allowing the invariance to extend to all distributions.

We now use Proposition (ref) to show that the same negative results obtained in Section (ref) also hold when considering continuous distributions or when considering continuous distributions over a fixed interval.

propositionFor any positive integer $t$ and $\beta \in \mathbb{R}$, $\Omega_0 = \{P \in \Omega_{\text{cts}}(\mathbb{R}) : \operatorname{\mathbb{E}}_P[X_i^{t}] = \beta\}$ is locally-dense and thus does not satisfy the randomization hypothesis.

To illustrate the content of this result, consider the null $\operatorname{\mathbb{E}}_P[X_i] = \beta$. Let $I_1,\dots,I_m$ be intervals with width bounded by $\delta > 0$. For any density $p$ supported on these intervals with mean $\gamma_p$ and any $\alpha \in (0,1]$, construct a distribution $H$ with density $\alpha p(x)$ on these intervals. Then $H$ has mean \[ \alpha \gamma_p + (1-\alpha)\operatorname{\mathbb{E}}_H[X_i \mid X_i \notin \cup_j I_j]. \] By placing the remaining mass on two intervals above and below $\beta$, one can choose $\alpha$ and the remaining mass so that the mean equals $\beta$. Thus the null is locally-dense. The following results all make use of similar constructions.

propositionFor any $p \in (0,1)$ and $q \in \operatorname{\mathbb{R}}$, $\Omega_0 = \{P \in \Omega_{\text{cts}}(\mathbb{R}): \operatorname{\mathbb{P}}[X_i \leq q] = p\}$ is locally-dense and thus does not satisfy the randomization hypothesis.
remarkWe show in Supplementary Appendix (ref) that Proposition (ref) extends to $\Omega_{\text{cts}}([b_0,b_1])$ for interval $[b_0,b_1]$ and $\beta$ in the support. The same extension holds for Proposition (ref) for any $q \in (b_0,b_1)$. We also show that the null hypothesis of fixed variance is locally-dense and therefore does not satisfy the randomization hypothesis.

Conclusion

Randomization tests yield exact finite-sample Type 1 error control when the null satisfies the randomization hypothesis. In practice, randomization tests often rely on invariance conditions that are stronger than the null hypothesis of primary interest. This paper developed a framework for characterizing which null hypotheses admit randomization tests with exact finite-sample validity.

Applying our framework to one-sample problems revealed impossibilities and limited possibilities of uses of finite-sample randomization tests. Many standard null hypotheses, including the null of mean zero, do not admit randomization tests with exact finite-sample validity. We further show that, when restricted to linear group actions, one-sample randomization tests can only test membership in the class of symmetric distributions or in subsets of the Gaussian family. These results imply that existing procedures are not overlooking alternative randomization tests that would achieve finite-sample validity.

While testing beyond symmetry or normality requires non-linear transformations, we show that even such transformations cannot be used to test many commonly studied null hypotheses. More broadly, the framework developed here delineates the limits of finite-sample randomization tests and provides constructive guidance for identifying when nulls admit such tests and how to construct them when they exist. Overall, our findings indicate that current practice is not omitting feasible exact procedures and further motivate a focus on asymptotic validity.

\addcontentsline{toc}{section}{References} {\singlespacing}

\setcounter{page}{1}

\addtocontents{toc}{\setcounter{tocdepth}{1}}