EconBase
← Back to paper

Some Finite Sample Properties of the Sign Test

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

18,278 characters · 9 sections · 18 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

1 Some Finite Sample Properties of the Sign Test

abstract\setstretch{1} This paper contains two finite-sample results concerning the sign test. First, we show that the sign-test is unbiased with independent, non-identically distributed data for both one-sided and two-sided hypotheses. The proof for the two-sided case is based on a novel argument that relates the derivatives of the power function to a regular bipartite graph. Unbiasedness then follows from the existence of perfect matchings on such graphs. Second, we provide a simple theoretical counterexample to show that the sign test over-rejects when the data exhibits correlation. Our results are useful for understanding the properties of approximate randomization tests in settings with few clusters.

Introduction

In recent years, randomization tests have become increasingly popular as a method for inference, due in no small part to the fact that they control size exactly even in finite sample. For example, ai2017 and young2019 advocate their in policy evaluation for this reason. Exactness has also been useful in settings with a small number of “effective observations". Here, tests are set up to be asymptotically equivalent to randomization tests with a finite number of observations that are independently but not identically distributed. Exactness of randomization tests then yields size control. Such as approach has been fruitful in for inference for cluster-dependent data with few clusters crs2017, css2019 and specification tests for RDD involving few units near the cut-off (canay2018approximate, bugni2021testing), among others.

In the above examples, randomization tests with a fixed number of observations are often used directly or indirectly for inference. It is therefore useful to understand their finite sample properties. hoeffding1952 showed that randomization tests control size in this setting, provided that the randomization hypothesis holds -- that is, that the data satisfies certain symmetries. The power of randomization tests is less well documented, particularly when observations are not identically distributed. It is challenging to analyze the power of randomization tests in finite sample due to their combinatorial nature. However, allowing for non-identical observations is important. In crs2017, each observation corresponds to a cluster, so accommodating non-identical observations is necessary to allow for non-identical clusters, a keystone of the clustering literature. Additionally, it is useful to assess the sensitivity of randomization tests to violations of the randomization hypothesis. In the clustering context, researchers rarely know the cluster structure ex ante. Misspecify the cluster structure may lead to (approximate) randomization tests with dependence between observations. Understanding how randomization tests behave under violations of independence is thus of relevance to applied researchers.

In this paper, we attempt to make headway on the above issues by focusing on the sign test, which is a test for the median of a collection of random variables. If these random variables are symmetrically distributed, it is also a test for the mean. In addition to being an especially tractable randomization test, the sign test is also relevant given its use in financial econometrics to test for abnormal returns in event studies (corrado1992specification), in RDD to test the continuity of the running variable (bugni2021testing), or to test for the level of clustering in computing clustered standard error (cai2023modified), among other applications.

We document two properties of the sign test in finite samples. Our first result shows that when observations are independent but not identically distributed, the sign test is unbiased against one and two-sided alternatives. Our proof for the two sided case is based on a novel argument that relates the derivative of the power function to a regular bipartite graph. Hall's Marriage Theorem, which implies the existence of a perfect matching for regular bipartite graphs, then yields unbiasedness. Our second result shows that sign tests with correlated normal random variables strictly over-rejects. This provides a clear counterexample, showing that it is important to account for correlation within units when conducting sign tests and randomization tests more generally.

This paper complements the literature on the properties of the sign test. It is well-known since hoeffding1952 that the sign test for the median has exact size in finite sample, even when observations non-identically distributed. On the issue of finite sample power, under the assumption that observations are identically distributed, the sign test is known to be the uniformly most powerful unbiased test for one- and two-sided hypotheses about the median (see Section 4.9 and Theorem 4.4.1 in lr2005). dixon1953power computes the power of the sign test under alternatives which are identically, normally distributed. However, to our knowledge, there does not exist power results once we allow for observations that are not identically distributed. In particular, it is not clear how the standard approach generalizes to this case (see Remark (ref)). Our paper provides a novel argument to show unbiasedness. On the issue of dependence, gastwirth1971effect shows that the sign test over-rejects asymptotically when observations exhibit first order auto-correlation. Our counterexample, based on equicorrelated normal random variables, demonstrates the same phenomenon simply in the finite sample setting.

The rest of this paper is organized as follows. Section 2 presents notation and describes the sign test. Section 3 contains our main results. Section 4 concludes. All proofs are contained in the appendix.

Set up and Notation

Let $\mu = \text{Med}(X_1) = ... = \text{Med}(X_q)$.

Two-Sided Test

We want to test $$H_0: \mu = \mu_0 \quad \text{against} \quad H_A: \mu \neq \mu_0~.$$ First define $Y = \left(X_1 - \mu_0, ..., X_q - \mu_0\right)'$. Next, define the two-sided test statistic

equation[equation omitted — 127 chars of source]

Next, denote by $\mathbf{G}$ the set of sign changes. $\mathbf{G}$ can be identified with the set of $g \in \{-1, 1\}^{q}$ so that $gY = \left(g_1Y_1, ... ,g_qY_q\right)'$. The two-sided sign test $\phi_{2,n}$ is the randomization test that rejects the null hypothesis when $T(Y)$ takes on extreme value relative to $T(gY)$. It proceeds as follows.

Define $M = |\mathbf{G}|$ and let:

equation*[equation* omitted — 66 chars of source]

be the ordered values of $T(gY)$ as $g$ varies in $\mathbf{G}$. For a fixed nominal level $\alpha$, let $k$ be defined as

equation*[equation* omitted — 49 chars of source]

where $\lceil x \rceil$ denotes the smallest integer greater or equal to $x$. In addition, define:

align*[align* omitted — 128 chars of source]

and set

equation*[equation* omitted — 57 chars of source]

We can then define the two-sided sign test as:

align[align omitted — 176 chars of source]

In words, this test rejects the null hypothesis with certainty when $T(Y) > T^{(k)}(Y)$. When $T(Y) = T^{(k)}(Y)$, it rejects the null hypothesis with probability $a(Y)$. The test does not reject when $T(Y) < T^{(k)}(Y)$. Suppose $\{Y_i\}_{i=1}^q$ do not have point mass at $\mu$. Then under the null hypothesis, $\text{sgn}\{Y_i\}$ is a Rademacher random variable and $$\text{Unif}\left(T^{(1)}(Y), T^{(2)}(Y), ..., T^{(M)}(Y)\right) \, \bigg\lvert \, Y \sim \text{Binomial}\left(q, \frac{1}{2}\right)~.$$ As such, the critical value for our test can equivalently be written in terms of the quantiles of the $\text{Binomial}(q, \frac{1}{2})$ distribution.

One-Sided Test

Suppose we are instead interested in the hypothesis: $$H_0: \mu = \mu_0 \quad \text{against} \quad H_A: \mu \geq \mu_0~.$$ We perform the one-sided sign test with the test statistic:

equation[equation omitted — 69 chars of source]

The rest of the testing procedure is identical to that of the two-sided test, except with $S(Y)$ taking the place of $T(Y)$. Let the corresponding test by $\phi_{1,n}$.

In the following sections, $E_\mu[\,\cdot\,]$ denotes expectation evaluated with respect to the data-generating process with median $\mu$.

Results

Unbiasedness of the Sign Test

In this subsection, we show that the sign test is unbiased against two-sided alternatives in finite sample. Our result implies that the approximate sign test is asymptotically unbiased.

assumptionLet $\{Y_i\}_{i=1}^q$ be independent, not necessarily identical random variables with median $\mu$.

It is well known that the sign test (and randomization tests more generally) has exact $p$-values in finite sample:

theorem[hoeffding1952] Given assumption (ref), suppose $\mu = \mu_0$. Then $E_{\mu_0}[\phi_n(Y)] = \alpha$.

It is also known that the sign test against one-sided alternatives is unbiased even with random variables that are not identically distributed (see Lemma 5.9.1 in lr2005 and also Appendix (ref)). Compared to one-sided alternatives, tests against two-sided alternatives are arguably more relevant, given the state of applied econometrics today. Our unbiasedness result for two-sided alternatives states that:

propositionGiven Assumption (ref), suppose $\mu \neq \mu_0$. Then, \begin{itemize} • $E_{\mu}[\phi_{2,n}(Y)] \geq \alpha$. • If $\alpha < 1$ and at least one of $\{Y_i\}_{i=1}^n$ is continuously distributed around $\mu$, the inequality in (a) is strict. • Furthermore, as $|\mu - \mu_0| \to \infty$, $E_\mu[\phi_{2,n}(Y)] \to 1$. \end{itemize} The same results hold for $\phi_{1,n}$ if $\mu > \mu_0$.

To show that $E_{\mu}[\phi_{t,n}(Y)] > \alpha$, we need to avoid the case in which $P_\mu(Y_i \leq \mu_0) = P_\mu(Y_i > \mu_0)$ for all $i$. This might occur if for instance all the random variables have zero density in a neighbourhood of $\mu$. This definitely cannot happen for continuous random variable, as is the case with approximate sign tests. Note that power cannot be 1 for any finite parameter value. This is because we are conducting inference using finite number of “effective observations".

Unbiasedness in the one-sided test follows straightforwardly from monotonicity of the test statistic in $\mu_0$. Our argument is similar in spirit to Lemma 5.9.1 in lr2005, which concerns the permutation test for two-sample comparison of means. Our argument for the two-sided test is based takes a novel approach that may be of interest. It relates derivatives of the power function to a regular bipartite graph. On this graph, each node represents a term in the derivative and an edge represents an inequality relationship. A perfect matching, which exists by Hall's Marriage Theorem, then allows us to sign the derivative.

remarkWhen observations are identically distributed, the approach of lr2005 to prove unbiasedness proceeds by showing that the distribution of signs under is a member of the one-parameter exponential family. The sign test can then be shown to be the uniformly most powerful test, which in turn implies unbiasedness. This argument is presented in Section 4.9, drawing on Theorems 4.4.1 and 3.7.1. When observations are allowed to be non-identically distributed, the distribution of signs under the null hypothesis is not a member of the one-parameter exponential family. In particular, under a given alternative hypothesis, each random variable $\text{sgn}\{Y_i\}$ will be $1$ with a different probability $p_i$. It is therefore not clear how to adapt this approach to the non-identical distributions.

The corollary below states that an approximate randomization test is unbiased as well. It follows immediately by the continuous mapping theorem.

corollarySuppose we have variables $S_n \to \text{N}(\mu\cdot \iota, D)$, where $D$ is a $q \times q$ diagonal matrix and $\iota$ is the $q$-vector of 1's. Then for $t \in \{1,2\}$, \begin{itemize} • If $\alpha < 1$, $\lim_{n \to \infty }E[\phi_{t,n}(S_n)] > \alpha$ • As $|\mu - \mu_0| \to \infty$, $\lim_{n \to \infty} E_\mu[\phi_{t,n}(S_n)] \to 1$ \end{itemize}

Size Distortion under Equicorrelation

Here, we study size distortion of the sign test with normal, equicorrelated random variables that are wrongly assumed to be independent. This is a counterexample showing that over-rejection could be a consequence when practitioners using approximate sign test gets the cluster structure wrong.

assumptionLet $Y \sim \text{N}(\mu, (1-\rho)\mathbf{I}_q + \rho \cdot \iota \iota^T)$ be equicorrelated normal random variables.

Consider again test (ref). We show that under $H_0$, when $\mu = \mu_0$, the test will over-reject:

propositionGiven Assumption (ref), suppose that $\mu = \mu_0$, $\rho > 0$, $q > 1$ and $\alpha < 0.25$. Then $$E[\phi_{2,n}(Y)] > \alpha~.$$ Furthermore, as $\rho \to 1$, $E[\phi_n(Y)] \to 1$.

In other words, the sign test could mistake correlation for violation of the null. Intuitively, the sign test checks for balance of the signs around the mean. If the variables are positively correlated, there will be an unusually large number of random variables with the same sign relative to independence, so that the test wrongly rejects. Our result is stated for the two-sided test, though the analogous result for the one-sided test follows straightforwardly. Our counterexample complements a large literature showing that the usual $t$- and Wald tests can experience significant size distortion when dependence is ignored (see for instance bdm2004, cgm2008).

To further understand size distortion under equicorrelation, we compute asymptotic power numerically. This is possible because the size (in particular equation (ref) in appendix A.2) can be expressed as a scalar integral that can be quickly and accurately computed using Gauss-Hermite quadratures. We evaluate the power of the test at $5\%$ and $10\%$ levels of significance using 1000 nodes. The results are presented in figures (ref) and (ref) respectively.

For the test at 5%-level, we see that for the $\rho$'s considered, size starts at $5\%$ when the number of sub-clusters is 1. It then monotonically increases as we increase the number of sub-clusters. Furthermore, fixing the number of sub-clusters, the size of the test increases also increases monotonically in the correlation across the sub-clusters. The test at 10%-level yields similar results. All in all, the numerical results concord with that of Proposition (ref).

figure[figure omitted — 432 chars of source]
figure[figure omitted — 435 chars of source]

Size Distortion under Minimal Correlation

Equicorrelation appears to be a strong assumption, since all of $Y$ is assume to be correlated. Suppose only the first two entries of $Y$ are correlated. In turns out that size distortion arises as well.

assumptionLet $Y \sim \text{N}(\mu, \Sigma)$, where \begin{equation} \Sigma = \begin{pmatrix} 1 & \rho & 0 & \cdots & 0 \\ \rho& 1 & 0 & \cdots & 0 \\ 0 & 0 & 1 & \cdots & 0 \\ \vdots& \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & 0 & \cdots & 1 \\ \end{pmatrix} . \end{equation}

Our earlier result also applies to this case:

propositionGiven Assumption (ref), suppose that $\mu = \mu_0$, $\rho > 0$, $q > 1$ and $\alpha < 0.25$. Then $$E[\phi_{2,n}(Y)] > \alpha~.$$

Conclusion

We document two finite sample properties of the sign test. We show that it is unbiased against one- and two-sided alternatives, even with non-identically distributed observations, but can experience size distortion if the samples are positively correlated. In the context of approximate sign tests under fixed cluster asymptotics, the latter result shows that misspecifying the cluster structure will lead to over-rejection of the null. Our result can be a useful first step towards a more general theory for the finite sample properties of randomization tests.

\\