EconBase
← Back to paper

Inference with a single treated cluster

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

186,533 characters · 5 sections · 48 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference with a single treated cluster

\address{Department of Economics, University of Michigan, 611 Tappan Ave, Ann Arbor, MI 48109, USA. Tel.: +1 (734) 764-2355. Fax: +1 (734) 764-2769} \email{[email removed]} \urladdr{\href{https://umich.edu/ hagem}{umich.edu/ hagem}}

abstractI introduce a generic method for inference about a scalar parameter in research designs with a finite number of heterogeneous clusters where only a single cluster received treatment. This situation is commonplace in difference-in-differences estimation but the test developed here applies more generally. I show that the test controls size and has power under asymptotics where the number of observations within each cluster is large but the number of clusters is fixed. The test combines weighted, approximately Gaussian parameter estimates with a rearrangement procedure to obtain its critical values. The weights needed for most empirically relevant situations are tabulated in the paper. Calculation of the critical values is computationally simple and does not require simulation or resampling. The rearrangement test is highly robust to situations where some clusters are much more variable than others. Examples and an empirical application are provided. \vskip 1em JEL classification: C01, C22, C32\\ Keywords: cluster-robust inference, difference in differences, two-way fixed effects, clustered data, dependence, heterogeneity

Introduction

Inference about the average effect of a binary treatment or policy intervention is often much more challenging than its estimation. For example, calculating a difference-in-differences estimate can be as simple as comparing the difference in average outcomes of individuals in a group before and after an intervention to the same differences in unaffected groups. The main challenge for inference is that individuals within each of these groups likely depend on one another in unobservable ways. Taking this dependence into account generally requires knowledge of an explicit ordering of the dependence structure within each group. While time-dependent data have a natural ordering, it may be difficult or impossible to credibly order cross-sectionally dependent data within states or villages. Researchers commonly try to sidestep this problem by splitting large groups into smaller clusters that are presumed to be independent in order to have access to standard inferential procedures based on cluster-robust standard errors or the bootstrap. Splitting states, villages, or other large groups into smaller clusters is often difficult to justify but necessary for most of the available inferential procedures because they achieve consistency by requiring the number of clusters to go to infinity. If a procedure is valid with a fixed number of clusters, it typically requires at least two treated clusters unless strong homogeneity conditions are satisfied. Numerical evidence by bertrandetal2004, mackinnonwebb2014, and others suggests that ignoring dependence and heterogeneity may lead to heavily distorted inference in empirically relevant situations. In both cases, the actual size of the test can exceed its nominal level by several orders of magnitude, i.e., nonexistent effects are far too likely to show up as highly significant.

In this paper, I introduce an asymptotically valid method for inference with a single treated cluster that allows for heterogeneity of unknown form. The number of observations within each cluster is presumed to be large but the total number of clusters is fixed. The method, which I refer to as a rearrangement test, applies to standard difference-in-differences estimation and other settings where treatment occurs in a single cluster and the treatment effect is identified by between-cluster comparisons. The key theoretical insight for the rearrangement test is that a mild restriction on some but not all of the heterogeneity in two samples of independent normal variables allows testing the equality of their means even if one sample consists of only a single observation. I prove that this is possible for empirically relevant levels of significance if the other sample consists of at least ten observations. The rearrangement test compares the data to a reordered version of itself after attaching a special weight to the sample with a single observation. The weights needed for most standard situations are tabulated in the paper and calculating additional weights is computationally simple. I also show that the weights remain approximately valid if the two samples of independent heterogeneous normal variables arise as a distributional limit. I exploit this result in the context of cluster-robust inference by constructing asymptotically normal cluster-level statistics to which the rearrangement test can be applied. The resulting test is consistent against all fixed alternatives to the null, powerful against $1/\sqrt{n}$ local alternatives, and does not require simulation or resampling.

Inference based on cluster-level estimates goes back at least to famamacbeth1973. Their approach is generalized and formally justified by ibragimovmueller2010, ibragimovmueller2016, who construct $t$ statistics from cluster-level estimates and show that these statistics can be compared to Student $t$ critical values. canayetal2014 obtain null distributions by permuting the signs of cluster-level statistics under symmetry assumptions. hagemann2019b permutes cluster-level statistics directly but adjusts inference to control for the potential lack of exchangeability. All of these methods allow for a fixed number of large and heterogeneous clusters but require several treated clusters. The rearrangement test complements these methods because it relies on the same type of high-level condition on the cluster-level statistics but is valid with a single treated cluster. Other methods that are valid with a fixed number of clusters are the tests of besteretal2014 and a cluster-robust version of the wild bootstrap cameronetal2008, djogbenouetal2019 analyzed by canayetal2018. However, these papers rely on strong homogeneity conditions across clusters that are not needed here.

Several approaches for inference have been developed specifically for difference-in-differences estimation. conleytaber2011 provide a method that is valid with a single treated cluster and infinitely many control clusters under strong independence and homogeneity conditions that justify an exchangeability argument. fermanpinto2019 extend this approach to situations where the form of heteroskedasticity is known exactly. Another extension by ferman2020 allows for spatial correlation while maintaining conleytaber2011's exchangeability condition. The rearrangement test differs from these methods because it is not limited to models estimated by difference in differences, does not rely on exchangeability conditions, and allows for completely unknown forms of heterogeneity. Other approaches due to mackinnonwebb2019b, mackinnonwebb2019a use randomization (permutation) inference for difference-in-differences estimation and other models with few treated clusters. They test “sharp” fisher1935 nulls under randomization hypotheses and asymptotics where the number of clusters is eventually infinite. In contrast, the present paper is able to test conventional nulls in a setting with finitely many clusters.

The remainder of the paper is organized as follows: Section (ref) proves several new results on normal random vectors with independent, heterogeneous entries after a specific transformation and introduces the rearrangement test. Section (ref) establishes the asymptotic validity of the test in the presence of finitely many heterogeneous clusters when only one cluster received treatment and discusses several examples. Section (ref) illustrates the finite sample behavior of the new test in simulations and in data used by garthwaiteetal2014, who analyze the effects of a large-scale disruption of public health insurance in Tennessee. Section (ref) concludes. The appendix contains auxiliary results and proofs.

I will use the following notation. $1\{A\}$ is an indicator function that equals one if $A$ is true and equals zero otherwise. Limits are as $n\to\infty$ unless noted otherwise and $\leadsto$ denotes convergence in distribution.

Inference with heterogenous normal variables

In this section, I construct a test for the equality of means of two samples of independent heterogeneous normal variables where one sample consists of only a single observation. The other sample has finitely many observations. I show that the test has power while controlling size (Theorem (ref)) and remains approximately valid if this two-sample problem characterizes the large sample distribution of a random vector of interest (Proposition (ref)).

Consider $q$ independent variables $X_{0,1},\dots, X_{0,q}$ with $X_{0,k}\sim N(\mu_0,\sigma_k^2)$ for $1\leqslant k\leqslant q$. Independently, there is an additional variable $X_1 \sim N(\mu_1, \sigma^2)$. I interpret this as a two-sample problem with “control” sample $X_{0,1},\dots, X_{0,q}$ and “treatment” sample $X_1$, although all of the following still applies if these roles are reversed. The objective is to test the null hypothesis of equality of means,

equation*[equation* omitted — 42 chars of source]

without knowledge of $\mu_0, \sigma, \sigma_1,\dots,\sigma_q$ and without assuming that these quantities can be consistently estimated. I account for the uncertainty about $\mu_0$ by recentering the data $X = (X_1, X_{0,1},\dots, X_{0,q})$ with $\bar{X}_0 = q^{-1}\sum_{k=1}^q X_{0,k}$ to define

equation[equation omitted — 141 chars of source]

for some known weight $w\in(0,1)$ that will be chosen shortly. If $X_1-\bar{X}_0 >0$, the $1+w$ increases $X_1-\bar{X}_0$ and $1-w$ decreases $X_1-\bar{X}_0$. If $X_1-\bar{X}_0 < 0$, these effects are reversed. The idea underlying the test is that if the decreased version of $X_1-\bar{X}_0$ is still large in comparison to $X_{0,1} - \bar{X}_0,\dots, X_{0,q} - \bar{X}_0$, then this size difference is unlikely to be only due to heterogeneity in $\sigma^2, \sigma_1^2,\dots,\sigma_q^2$ but provides evidence that $\mu_1$ and $\mu_0$ are in fact not equal. I show below that $w$ gives precise probabilistic control over this comparison. In particular, choosing $w$ appropriately allows me to construct a test whose size can be bounded at a predetermined significance level.

Before defining the test statistic, I first introduce some notation. For a given vector $s\in\mathbb{R}^d$, let $s_{(1)}\leqslant\cdots\leqslant s_{(d)}$ be the ordered entries of $s$. Denote by $s\mapsto s^\triangledown = (s_{(d)},\dots, s_{(1)})$ the operation of rearranging the components of $s$ from largest to smallest. The test uses $S(X,w)$ and its rearranged version $S(X,w)^\triangledown$ in the difference-of-means statistic

equation[equation omitted — 128 chars of source]

to define the test function

equation[equation omitted — 123 chars of source]

The test, which I refer to as rearrangement test, rejects if $\varphi (X,w) = 1$ and does not reject otherwise. As stated, the test is against the alternative of a positive treatment effect, $H_1\colon \mu_1 > \mu_0$. For a test against $H_1\colon \mu_1 < \mu_0$, simply use $\varphi (-X, w)$. These alternatives can be combined to provide a two-sided test. I describe the exact implementation below equation (ref) ahead. Also note that the first difference of means in (ref) simplifies to $T(S(X,w))=X_1-\bar{X}_0$ but $T(S(X,w)^\triangledown)$ is in general a complicated function of $w$.

Intuitively, the rearrangement test can be interpreted as a permutation test that treats $S = S(X,w)$ as if it were the data and uses the second largest permutation statistic of $T(S)$ as critical value $c$. If $T(S) > c$, then the only possibility left is that $T(S)$ equals its largest permutation statistic. For the difference of means $T(S)$, that statistic must be $T(S^\triangledown)$ and therefore $T(S) > c$ is equivalent to $\varphi (X,w) = 1$. Because $S$ is being permuted and not $X$, this also explains why it is sensible to write $T(S(X,w))$ instead of $X_1-\bar{X}_0$ in the definition of the test function (ref). A classical permutation test would then use an exchangeability condition on $S$ to determine the size of the test. Even though the $S$ constructed here is far from exchangeable, I will show that this test has power while controlling size at a predetermined level. Instead of relying on exchangeability, the results here depend on the joint normality of $X$ combined with the location and scale invariance property $\varphi (X,w) = \varphi ((X-\mu_0 1_{q+1})/\sigma,w)$, where $1_{q+1}$ is a ($q+1$)-vector of ones. The location invariance is forced by the recentering of $X$ with $\bar{X}_0$ and effectively removes $\mu_0$ from the list of nuisance quantities. The scale invariance is ensured by the specific choices of $T$ and $\varphi$. It reduces the dimensionless unknowns $\sigma, \sigma_1,\dots,\sigma_q$ to the more tractable ratios $\sigma_1/\sigma,\dots,\sigma_q/\sigma$.

I start with the analysis of size and power, and connect these results with the situation where $X = (X_1, X_{0,1},\dots, X_{0,q})$ is an asymptotic approximation later on. I assume that the variances $\sigma_k^2$ of the $X_{0,k}$, $1\leqslant k\leqslant q$, are bounded away from zero by some $\text{\b{$\sigma$}}^2>0$ for all but one $k$. This avoids a trivial and in practice easily recognizable situation where some of the $X_{0,k}$ are exactly equal. I also restrict the variance $\sigma^2$ of $X_1$ to be bounded above by some $\bar{\sigma}^2<\infty$ because letting $\sigma\to\infty$ in $\varphi(X,w)$ would have the same effect as setting all $\sigma_k^2$ equal to zero. Under the null hypothesis, the distribution of $\varphi (X,w)$ is then determined by the unknown value of \[\lambda \in \Lambda \coloneqq \{ (\mu_0, \sigma, \sigma_1,\dots, \sigma_q) \in\mathbb{R}\times (0,\infty)^{q+1} : \sigma \leqslant \bar{\sigma} \text{ and } \sigma_k \geqslant \text{\b{$\sigma$}} \text{ for all $k$ but one} \}.\] Under the alternative, the distribution of $\varphi (X,w)$ also depends on the treatment effect $\delta = \mu_1 - \mu_0$. I write ${\mathord \mathrm{E}}_{\lambda, \delta}$ and ${\mathord P}_{\lambda, \delta}$ to emphasize this dependence but occasionally drop subscripts to prevent clutter.

My strategy is to first bound the null rejection probability ${\mathord \mathrm{E}}_{\lambda, 0} \varphi (X,w)$ uniformly in $\lambda\in\Lambda$ by a smooth function of the weight $w$. I can then find a $w$ to make the bound exactly equal to the desired significance level to guarantee size control. The bound is also a function of the number of control observations $q$ and the maximal relative heterogeneity $\varrho = \bar{\sigma}/\text{\b{$\sigma$}}$ of treated and untreated observations. The parameter $\varrho$ is user chosen and has a simple interpretation: it restricts how much more variable $X_1$ can be relative to the $X_{0,k}$ when one of the $\sigma_k$ equals zero and the remaining $\sigma_k$ are all equal to the lower limit $\text{\b{$\sigma$}}$. This is the worst-case scenario for the test because $X_1$ is then likely to be very large on accident in comparison to the $X_{0,k}$. In that scenario, a $\varrho$ of 5 simply means that the variance of $X_1$ can be up to $5^2 = 25$ times larger than the variances of all but one of the $X_{0,k}$ and “infinitely more variable” than the remaining $X_{0,k}$. There are no restrictions on how much less variable $X_1$ can be than $X_{0,1},\dots,X_{0,q}$ and, in particular, $\bar{\sigma}/\text{\b{$\sigma$}}$ can be less than one.

The following theorem is the main theoretical result of the paper. It establishes the existence of a size bound that is valid for a fixed number of control observations $q$ and fully accounts for the uncertainty about the parameters in $\Lambda$. The theorem also shows that the test has power against the alternative $H_1\colon \mu_1 > \mu_0$. Results in the other direction follow by considering ${\mathord \mathrm{E}}_{\lambda,-\delta}\varphi(-X, w)$ instead of ${\mathord \mathrm{E}}_{\lambda,\delta}\varphi(X, w)$. The discussion immediately below focuses on the implications of the theorem. I address some of its technical aspects towards the end of this section. Let $\Phi$ and $\phi$ denote the normal distribution and density functions, respectively.

theorem[Size and power] Let $X_1, X_{0,1,}\dots, X_{0,q}$ be independent with $X_1\sim N(\mu_0 + \delta, \sigma^2)$ and $X_{0,k}\sim N(\mu_0, \sigma_k^2)$ for $1\leqslant k\leqslant q$. If $\delta = 0$, then for all $w\in (0,1)$, \begin{align} \sup_{\lambda \in \Lambda} {\mathord \mathrm{E}}_{\lambda, 0} \varphi (X, w) \leqslant \xi_q(w,\varrho) \coloneqq \frac{1}{2^{q+1}} + &\int_{0}^{\infty} \Phi\bigl((1-w) \varrho y\bigr)^{q-1} \phi(y) dy \\ & + \min_{t > 0} \biggl ( \Phi\Bigl(\sqrt{q-1} w t \Bigr)^{q-1} + 2\Phi(- qt ) \biggr). \nonumber \end{align} Furthermore, for every $\lambda \in \Lambda$ and $w\in (0,1)$, we have $\lim_{\delta \to \infty}{\mathord \mathrm{E}}_{\lambda,\delta} \varphi (X, w)= 1$ and $\lim_{\delta \to \infty}{\mathord \mathrm{E}}_{\lambda, \delta} \varphi (X, 1)= 0$.

The theorem implies that the rearrangement test controls size, i.e., \[ \sup_{\lambda \in \Lambda} {\mathord \mathrm{E}}_{\lambda, 0} \varphi (X, w) \leqslant \alpha, \] whenever $q$, $w$, and $\varrho$ are such that $\xi_q(w,\varrho) \leqslant \alpha$ for the desired significance level $\alpha$. The bound $\xi_q(w,\varrho)$ has several properties that make this possible. In particular, it is monotonically increasing in $\varrho$ and decreasing in $q$. The reason for the monotonicity is that if $X_1$ can be more variable than $X_{0,1},\dots,X_{0,q}$, then the burden of proof to show “$\mu_1 > \mu_0$” as opposed to “$\mu_1 = \mu_0$ with a large realization of $X_1$” becomes necessarily higher. A large $q$ can ameliorate this effect somewhat because it removes uncertainty about $\mu_0$. The bound also tends to be decreasing in $w \in [0,1]$ because the integral generally dominates the other components, but can increase slightly in some situations.

figure[figure omitted — 85,078 chars of source]

This is illustrated in Figure (ref), where $w\mapsto \xi_q(w,\varrho)$ (solid lines) is essentially decreasing over the entire domain except for $\varrho=2$ and $w\geqslant .85$. Most importantly, it can be seen that $w\mapsto \xi_q(w,\varrho)$ decreases enough to dip below the desired significance level $\alpha = .05$ (dashed line) for all values of $\varrho$. As $q$ increases (not shown), $w\mapsto \xi_q(w,\varrho)$ is pushed towards zero but the shape of the function does not change meaningfully with $q$. The $w$ at which $\xi_q(w,\varrho)=\alpha$ is generally unique for most empirically relevant $\alpha$ and does not exist in some extreme situations. This can be seen in Figure (ref), where $w\mapsto \xi_q(w,\varrho)$ crosses $\alpha = .05$ only once for each $\varrho$ but, for example, $\xi_q(w,\varrho) = .6$ is never attained.

Theorem (ref) also provides information about the interplay between $w$ and the test under the alternative. In particular, it shows that the rearrangement test has power against $H_1 : \mu_1 > \mu_0$ for every $w\in (0,1)$ but the power declines sharply at $w=1$. I therefore explore the behavior of the test with $w$ near $1$ further in the following result. It provides a lower bound on the power of the test for fixed $\delta$.

proposition[Lower bound on power] Let $X_1, X_{0,1},\dots, X_{0,q}$ be independent with $X_1\sim N(\mu_0 + \delta, \sigma^2)$ and $X_{0,k}\sim N(\mu_0, \sigma_k^2)$ for $1\leqslant k\leqslant q$. For every $w\in (0,1)$, $\sigma, \sigma_1,\dots, \sigma_q > 0$, and $\delta > 0$, \[ \inf_{\mu_0\in\mathbb{R}}{\mathord \mathrm{E}}_{\lambda, \delta} \varphi (X, w) \geqslant 2^q\sup_{t \geqslant 0}\Phi\biggl(\frac{\delta}{\sigma} - \frac{1+w}{1-w}t\biggr)\prod_{k=1}^q\Biggl(\Phi\biggl(\frac{\sigma}{\sigma_k}t\biggr) - 0.5\Biggr) \] The supremum is attained on $t\in (0,\infty)$. The right-hand side is strictly positive and converges to $1$ as $\delta \to \infty$.

The bound shows that the test exhibits a standard relationship between the signal $\delta$ and the noise components $\sigma_1,\dots,\sigma_q$. Power is low if the signal relative to $\sigma$ is weak or the noise in the control group relative to $\sigma$ is strong. The latter relationship is in contrast to Theorem (ref), where small $\sigma_k$ relative to $\sigma$ were problematic. In addition, the bound also clarifies that $w$ dampens $\delta$ through the function $w\mapsto (1+w)/(1-w)$, which is arbitrarily large for $w$ sufficiently close to $1$. A $w$ very close to $1$ can therefore drown out a large treatment effect even if the noise coming from the control observations is mild. (The role of the supremum is simply to find the best possible balance for a given set of parameters.) It is also worth noting that the bound is tight enough to converge to $1$ as $\delta \to \infty$ and to $0$ as $w \to 1$.

Because the $w$ that satisfies $\xi_q(w,\varrho) = \alpha$ is not necessarily unique and because Proposition (ref) suggests that power against the alternative $H_1 : \mu_1 > \mu_0$ for $w$ near one can be low, it is sensible to choose the smallest feasible $w$, denoted by

equation[equation omitted — 113 chars of source]

in the definition of the rearrangement test function for a test of size $\alpha$,

equation[equation omitted — 118 chars of source]

The test $\varphi_\alpha$ also depends on $\varrho$ but this is suppressed here to prevent clutter. Table (ref) lists values of $w_q(\alpha,\varrho)$ for common choices of $\alpha$ as a function of $\varrho$ and $q$. They guarantee

equation[equation omitted — 133 chars of source]

The list is not exhaustive and additional values can be easily calculated by numerical integration. An R command that performs the calculations can be found at https://hgmn.github.io/rea.

table[table omitted — 4,284 chars of source]

Table (ref) shows that the rearrangement test is available in a wide variety of situations depending on the desired significance level and tolerance for heterogeneity. For instance, a test with a 10% significance level is already available with $q=10$ control observations. A 5% level test becomes available at $q=15$, a 1% level test at $q=20$, and for $q\geqslant 25$ there are essentially no restrictions to the level and underlying heterogeneity. This provides two avenues for implementation:

enumerate• Choose a desired maximal degree of heterogeneity $\varrho$ and make test decisions based on this choice. • Determine at which degree of maximal heterogeneity the null hypothesis can no longer be rejected.

The first option is similar in spirit to the ubiquitous staigerstock1997 rule of thumb for weak instruments, where an $F$ statistic larger than 10 corresponds to a tolerance for an at most 10% bias (as defined in stockyogo2001) in the instrumental variables estimator relative to least squares. The second option takes the form of a “robustness check.” It has a meaningful interpretation because a result that is robust to a tenfold larger standard deviation in the treated observation relative to the control sample is more credible than a result that only survives a twofold difference in standard deviation. This second option leaves it up to the reader to decide whether the results are convincing.

The test decision itself is simple. Choose $w = w_q(\alpha,\varrho)$ from Table (ref) for a given number of control observations $q$, desired significance level $\alpha$, and maximal tolerance for heterogeneity, e.g., $\varrho = 2$. For this $w$, compute $S = S(X, w)$ as in (ref) and reorder the entries of $S$ from largest to smallest to obtain $S^\triangledown$. For an $\alpha$-level test of $\mu_1 = \mu_0$, reject in favor of $\mu_1 > \mu_0$ if $T(S) = T(S^\triangledown)$ as defined in (ref). For a one-sided test with level $\alpha$ against $\mu_1 < \mu_0$, reject if $T(-S) = T((-S)^\triangledown)$. For a two-sided test with level $2\alpha$, reject in favor of $\mu_1 \neq \mu_0$ if either $T(S) = T(S^\triangledown)$ or $T(-S) = T((-S)^\triangledown)$. The “robustness check” increases $\varrho$ until the null hypothesis can no longer be rejected against the desired alternative. The test decision is monotonic in $\varrho$, i.e., if $\varrho' > \varrho$ lead to the same test decision, then the decision does not change for any value between $\varrho$ and $\varrho'$. An R command that implements the test and the robustness check for any choice of $\varrho$ is available at \href{https://hgmn.github.io/rea}{https://hgmn.github.io/rea}.

I now turn to a discussion of some technical aspects of the size bound $\xi_q(w,\varrho)$ that forms the theoretical underpinning for the rearrangement test. The bound, defined in (ref), has three components with simple interpretations: The $1/2^{q+1}$ removes an unlikely event ($X_1 < \mu_0$, $X_{0,1} < \mu_0, \dots, X_{0,q} < \mu_0$ at the same time) from consideration. This forces a monotonicity property over the complement of this event and allows tightly bounding an oracle version of the problem where $\mu_0$ replaces $\bar{X}_0$ in (ref). This bound is the integral in (ref). The minimization problem then adjusts for the fact that the data are centered by $\bar{X}_0$ instead of the unknown $\mu_0$. The minimizer does not have closed form but is easily found numerically.\footnote{In particular, at $t = 1/q$, $\Phi(\sqrt{q-1} w t )^{q-1} + 2\Phi(- qt ) < \Phi(1/\sqrt{q})^{q-1} + 2\Phi(- 1) < 1$ for $q > 2$. Because $\Phi(\sqrt{q-1} w t )^{q-1} + 2\Phi(- qt ) \geqslant 1$ at $t\in \{0,\infty \}$, the minimization problem always has an interior solution. This also implies that the bound as a whole is a smooth function of $w$ and $\varrho$.} Taken together, $\xi_q(w_q(\alpha, \varrho),\varrho)$ can therefore be roughly viewed as a tight bound for a high-probability event plus two small adjustments. I use Table (ref) to illustrate the relative size of these adjustments. In the table, empty cells correspond to situations where there is either no $w$ such that $\xi_q(w, \varrho) = \alpha$ or more than $\alpha/2$ of $\xi_q(w_q(\alpha, \varrho),\varrho)$ is taken up by the non-tight parts of the bound. Cells in italics are settings where between $\alpha/2$ and $\alpha/10$ of the bound are taken up by the non-tight parts. The lack of tightness in the remaining cells is less than $\alpha/10$. For these cells $\sup_{\lambda \in \Lambda} {\mathord \mathrm{E}}_{\lambda,0} \varphi_\alpha (X)$ approximately equals $\alpha$. As the table shows, $\xi_q(w_q(\alpha, \varrho),\varrho)$ is an essentially tight bound for $\sup_{\lambda \in \Lambda} {\mathord \mathrm{E}}_{\lambda,0} \varphi_\alpha (X)$ for $q\geqslant 30$. The bound is also nearly tight for values of $q$ as small as 15 as long as $\varrho$ is not too large. I return to a discussion of this aspect of the rearrangement test in Example (ref) (ahead), where I illustrate the size of the test numerically.

Finally, before concluding this section, I show that the rearrangement test remains approximately valid for random vectors $X_n$ converging in distribution to the random vector $X = (X_1, X_{0,1},\dots, X_{0,q})$ described in Theorem (ref). The reason is that ${\mathord \mathrm{E}} \varphi(X_n, w)$ and ${\mathord \mathrm{E}} \varphi(X, w)$ eventually coincide whenever $X$ has independent entries and a smoothly distributed first entry. The $X$ in Theorem (ref) easily satisfies these conditions, which makes $\varphi_\alpha(X_n)$ asymptotically an $\alpha$-level test.

proposition[Large sample approximation] Let $X_1, X_{0,1},\dots, X_{0,q}$ be independent and let $X_1$ have a continuous distribution. If $X_n\leadsto X$, then ${\mathord \mathrm{E}} \varphi (X_n, w) \to {\mathord \mathrm{E}}\varphi (X, w)$ for every $w\in(0,1)$.

I use Theorem (ref) and Proposition (ref) in the next section to construct a simple method for inference with a single treated cluster. Section (ref) shows how the rearrangement test performs in Monte Carlo experiments.

Inference with a single treated cluster

In this section, I use a single high-level condition to extend the rearrangement test introduced in the previous section to a test about a scalar parameter in research designs with a finite number of large, heterogeneous clusters where only a single cluster received treatment. I then outline how these results can be applied in empirical practice.

Suppose data from $q+1$ large clusters (e.g., states, industries, or villages observed over one or more time periods) are available. Data are dependent within clusters but independent across clusters. The exact form of dependence is unknown and not presumed to be estimable. An intervention took place during which one cluster received treatment and and $q$ clusters did not. The quantity of interest is a treatment effect or an object related to it that can be represented by a scalar parameter $\delta$. Because the entire cluster was treated, this parameter is only identified up to a location shift $\theta_0$ within the treated cluster and therefore only the left-hand side of \[ \theta_1 = \theta_0 + \delta \] can be identified from this cluster. If the treated cluster would have behaved similarly to the untreated clusters in the absence of an intervention, then $\theta_0$ can be identified from each untreated cluster. Pairwise comparison then identifies $\delta$.

The identification strategy outlined in the preceding paragraph is the idea behind differences in differences---arguably the most popular identification strategy in modern empirical research---and a variety of other models. The goal of this section is to use the rearrangement test to provide a generic method for testing the hypothesis \[H_0\colon \delta = 0, \] or, equivalently, $H_0\colon \theta_1 = \theta_0$. I achieve this by obtaining an estimate $\hat{\theta}_{1}$ of $\theta_1$ and estimates \smash{$\hat{\theta}_{0,1},\dots, \hat{\theta}_{0,q}$} of $\theta_0$ so that \[ \hat{\theta}_n = (\hat{\theta}_{1}, \hat{\theta}_{0,1} \dots, \hat{\theta}_{0,q}) \] is approximately a vector of independent but potentially heterogeneous normal variables that can be used as if it were the data vector $X$ from Section (ref).

The following example explains how to construct $\hat{\theta}_n$ in a simple situation. I discuss construction of $\hat{\theta}_n$ for difference in differences towards the end of this section.

example[Regression with cluster-level treatment] Consider a linear regression model \begin{equation*} Y_{i,k} = \theta_0 + \delta D_{k} + \beta_k' X_{i,k} + U_{i,k}, \end{equation*} where $i$ indexes individuals within cluster $k$. There are $q+1$ clusters and individuals in cluster $k = q+1$ received treatment ($D_k=1$) but those in $1\leqslant k\leqslant q$ did not ($D_k=0$). The parameter of interest $\delta$ on the treatment indicator $D_k$ can be interpreted as an average treatment effect under suitable conditions. See, e.g., sloczynski2018, sloczynski2020 and references therein for a precise discussion. The regression may also include covariates $X_{i,k}$ that vary within each cluster and have coefficients $\beta_k$ that may vary across clusters. The condition ${\mathord \mathrm{E}}(U_{i,k}\mid D_k, X_{i,k}) = 0$ identifies $\theta_1 = \theta_0 + \delta$ within the treated cluster and $\theta_0$ within the untreated clusters. The preceding display can then be written as \begin{align*} Y_{i,k} = \begin{cases} \theta_0 + \beta_k' X_{i,k} + U_{i,k}, &1\leqslant k\leqslant q,\\ \theta_1 + \beta_k' X_{i,k} + U_{i,k}, &k = q + 1. \end{cases} \end{align*} View these as $q+1$ separate regressions and use the least squares estimates of the constants $\theta_1$ and $\theta_0$ as the vector $\hat{\theta}_n = (\hat{\theta}_{1}, \hat{\theta}_{0,1} \dots, \hat{\theta}_{0,q})$ described above. \leavevmode\unskip\penalty9999 \hbox\nobreak \quad\hbox{\ensuremath{\square}}

I will now show that the cluster-level statistics $\hat{\theta}_n$ can be used together with the results in the previous section to perform a consistent test as the sample size $n$ grows large. The test is not limited to parameters estimated by least squares. Instead, consistency relies on the condition that a centered and scaled version of some estimate $\hat{\theta}_n$ converges to a $(q+1)$-dimensional normal distribution,

equation[equation omitted — 407 chars of source]

where $\mathchoice {\raisebox{.0em}{ $\overset{\theta}{\leadsto}$ }} {\raisebox{-.15em}{ $\overset{\raisebox{-.25em}{\scriptsize$\theta$}}{\leadsto}$ }} {} {}$ denotes weak convergence under $\theta = (\theta_1,\theta_0)$. For fixed $\theta$, the display can be interpreted as $\sqrt{n}(\hat{\theta}_{1} - \theta_1, \dots,\hat{\theta}_{0,1} - \theta_0, \dots, \hat{\theta}_{0,q} - \theta_0)\leadsto N(0, \operatorname*{diag}(\sigma, \sigma_1, \dots, \sigma_q))$ to include the case that one of the $\sigma_1,\dots, \sigma_q$ may be zero as in Theorem (ref).

A key feature of condition (ref) is that the $\sigma$ and $\sigma_1,\dots, \sigma_q$ are not assumed to be known or estimable by the researcher. This is important for applications because consistent variance estimation generally requires knowledge of an explicit ordering of the dependence structure within each cluster. While time-dependent data are automatically ordered, it may be difficult or impossible to infer or credibly assume an ordering of the data within states or villages. In contrast, (ref) can be established under weak (short-range) dependence conditions that only require existence of a potentially unknown ordering for which the dependence of more distant units decays sufficiently fast. machkouriaetal2013 present convenient moment bounds and limit theorems for this situation. For more results in this direction, see also besteretal2014 and references therein. In general, the convergence in (ref) also implicitly requires the number of observations in all clusters to grow with the sample size $n$. However, the clusters are not required to have similar or even identical sizes. Another noteworthy feature of condition (ref) is the diagonal covariance matrix of the limiting distribution. It is the only independence condition that is imposed on the clusters.

I now show that under the joint convergence (ref), a rearrangement test that uses $\hat{\theta}_n$ is asymptotically of level $\alpha$ with a single treated cluster and a fixed number of control clusters. The test $\varphi_\alpha(\hat{\theta}_n)$, as defined in (ref), has power against all fixed alternatives $\theta_1 = \theta_0 + \delta$ with $\delta > 0$ and local alternatives $\theta_1 = \theta_0 + \delta/\sqrt{n}$ converging to the null. In the latter situation, $\theta_0$ is fixed and $\theta = (\theta_0 + \delta/\sqrt{n}, \theta_0)$ implicitly depends on $n$. The convergence in (ref) is then a statement about an entire sequence $(\theta_0 + \delta/\sqrt{n}, \theta_0)$ instead of a single point. Results for alternatives with $\delta < 0$ follow from the same result by considering $\varphi_\alpha(-\hat{\theta}_n)$. These tests can be combined into a two-sided test that has power against fixed and local alternatives from either direction. Algorithm (ref) at the end of this section shows how this can be implemented.

theorem[Consistency and local power] Suppose (ref) holds with $\sigma^2 >0$ and at most one $\sigma_k = 0$. If $\theta_1 = \theta_0$, then \[ \lim_{n\to\infty}{\mathord \mathrm{E}} \varphi_\alpha(\hat{\theta}_n) \leqslant \alpha,\qquad\text{every $\alpha,\varrho$ with $0 < w(\alpha,\varrho) < 1$,} \] and if $\theta_1 > \theta_0$, then ${\mathord \mathrm{E}} \varphi_\alpha(\hat{\theta}_n)\to 1$. If (ref) holds with $\theta = (\theta_0 + \delta/\sqrt{n},\theta_0)$ and the $\sigma, \sigma_1,\dots, \sigma_{q}$ are continuous and positive at $\theta_0$, then \begin{equation*} \lim_{n\to\infty}{\mathord \mathrm{E}} \varphi_\alpha(\hat{\theta}_n) \geqslant 2^q\sup_{t \geqslant 0}\Phi\Biggl(\biggl(\frac{\delta}{\sigma(\theta_0)} - \frac{1+w_q(\alpha,\varrho)}{1-w_q(\alpha,\varrho)}t\biggr)\Biggr)\prod_{k=1}^q\Biggl(\Phi\biggl(\frac{\sigma(\theta_0)}{\sigma_k(\theta_0)}t\biggr) - 0.5\Biggr)>0. \end{equation*}
remarks(i) Because $\varphi_\alpha(\hat{\theta}_n) = 1$ if and only if $\varphi_\alpha(a(\hat{\theta}_n - \theta_0 1_{q+1})) = 1$, where $a> 0$ and $1_{q+1}$ is a $(q+1)$-vector of ones, the $\sqrt{n}$-rate in (ref) and in the theorem can be replaced by any other rate as long as the asymptotic normal distribution in (ref) is still attained. Several semiparametric or nonstandard estimators are therefore covered by the theorem. (ii) It is sometimes of interest in applications to test the null hypothesis $H_0\colon \theta_1 = \theta_0 + \gamma$ for a given $\gamma$. In that case, define $\Gamma = (\gamma 1\{k=1 \})_{1\leqslant k\leqslant q+1}$ and reject if $\varphi_\alpha(\hat{\theta}_n - \Gamma) = 1$. Replace $\theta_0$ by $\theta_0 + \gamma$ in Theorem (ref) and use part (i) of this remark to see that this leads to a consistent test. \leavevmode\unskip\penalty9999 \hbox\nobreak \quad\hbox{\ensuremath{\square}}

I now discuss how the high-level condition (ref) can be verified in an application. The specific example I use is difference-in-differences estimation but the arguments presented here apply more broadly. See also canayetal2014 and hagemann2019b for similar types of arguments in other models. For simplicity, I focus on (ref) under the null hypothesis $H_0 : \theta_1 = \theta_0$.

example[Difference in differences] Consider the panel model \begin{equation} Y_{i,t,k} = \theta_0 I_t + \delta I_t D_{k} + \beta_k' X_{i,t,k} + \zeta_{i,k}+ U_{i,t,k}, \end{equation} where $i$ indexes individuals $i$ in unit $k\in\{1,\dots, q+1\}$ at time $t\in\{0,1\}$. Treatment occurred between periods $0$ and $1$. Right-hand side variables are a post-intervention indicator $I_t=1\{t= 1\}$, a treatment indicator $D_k$ that equals $1$ if unit $k$ ever received treatment, individual fixed effects $\zeta_{i,k}$, and other covariates $X_{i,t,k}$ that for every $k$ vary at least before or after the intervention. The collection of pre and post intervention data from unit $k$ forms the $k$-th cluster. Let $n_k$ be the number of individuals in cluster $k$ so that $n = 2\sum_{k=1}^{q+1} n_k$ is the total sample size. View each cluster as a separate regression and rewrite (ref) in first differences as \begin{equation*} \Delta Y_{i,k} = \begin{cases} \theta_0 + \beta_k' \Delta X_{i,k} + \Delta U_{i,k}, &1\leqslant k\leqslant q, \\ \theta_1 + \beta_k' \Delta X_{i,k} + \Delta U_{i,k}, &k = q + 1, \end{cases} \end{equation*} where $\Delta Y_{i,k} = Y_{i,1,k} - Y_{i,0,k}$ and so on. Provided ${\mathord \mathrm{E}} (\Delta U_{i,k}\mid \Delta X_{i,k}) = 0$, the data identify $\theta_1 = \theta_0 + \delta$ in a treated cluster and $\theta_0$ in an untreated cluster. The least squares estimates $\hat{\theta}_1$ and $\hat{\theta}_{0,k}$ of the parameters $\theta_1$ and $\theta_0$ are suitable cluster-level estimates if $\hat{\theta}_n = (\hat{\theta}_{1}, \hat{\theta}_{0,1},\dots, \hat{\theta}_{n,q})$ satisfies condition (ref). In the absence of covariates (i.e., $\beta_k\equiv 0$), the centered and scaled least squares estimate in a control cluster under $H_0$ can be expressed as \[ \sqrt{n}(\hat{\theta}_{0,k} - \theta_0) = \biggl(\frac{n}{n_{k}}\biggr)^{1/2} n_{k}^{-1/2}\sum_{i=1}^{n_k}\Delta U_{i,k}. \] The same is true for $\sqrt{n}(\hat{\theta}_{1} - \theta_0)$ with $k=q+1$ on the right-hand side of the display. If the number of individuals per cluster is large in the sense that $n/n_k \to c_k \in (0,\infty)$ for $1\leqslant k\leqslant q+1$, then condition (ref) already holds if $n^{-1/2}(\sum_{i=1}^{n_k}U_{i,0,k},$ $\sum_{i=1}^{n_k}U_{i,1,k})$ is independent across $1\leqslant k\leqslant q+1$ and has a non-degenerate normal limiting distribution for each $k$. The latter condition can be ensured with a central limit theorem for spatially dependent data. See, e.g., jenischprucha2009 and machkouriaetal2013 for appropriate results. If the number of individuals per cluster is small, then Theorem (ref) implies that the rearrangement test can still be applied under the assumption that $((U_{i,0,k})^T_{1\leqslant i\leqslant n_k}, (U_{i,1,k})^T_{1\leqslant i\leqslant n_k})$ is multivariate normal for $1\leqslant k\leqslant q+1$. This last condition may be strong but serves to illustrate that $\hat{\theta}_1$ and $\hat{\theta}_{0,k}$ need not even be consistent for the test to be valid. Now consider pooled cross sections with $n_k$ individuals in period $0$, $m_k$ individuals in period $1$, and $\zeta_{i,k}\equiv \zeta_k$. The calculations in the preceding paragraph still apply with minor modifications. For period $1$, $n_k$ has to be replaced by $m_k$. The analysis is no longer in first differences but the underlying conditions are essentially identical as long as $n/n_k \to c_k\in (0,\infty)$ and $n/m_k \to c_k' \in (0,\infty)$ for $1\leqslant k\leqslant q+1$, where $n$ is the total sample size. If the number of individuals available post intervention $m = \sum_{k=1}^{q+1} m_k$ is relatively small in the sense that $m/n_k \to 0$ and $m/m_k \to c_k' \in (0,\infty)$, the scale invariance discussed in the remarks below Theorem (ref) allows replacement of the $\sqrt{n}$ in (ref) by $\sqrt{m}$. Then (ref) holds if $n_k^{-1/2}\sum_{t=1}^{n_k}U_{i,0,k} = O_P(1)$ and $m_k^{-1/2}\sum_{t=1}^{m_k}U_{i,1,k}$ obeys a central limit theorem for $1\leqslant k\leqslant q + 1$. The same argument applies with the roles of $n_k$ and $m_k$ reversed if relatively few individuals are available pre intervention. The calculations in the preceding two paragraphs can be generalized to include covariates and additional time periods at the expense of more involved notation and non-singularity conditions. The same types of arguments also apply if each cluster consists of one or few units over many time periods, although the conditions for time dependence are generally less involved. See dedeckeretal2007 for a comprehensive overview. These remarks and the calculations in this example also apply to the regression model in Example (ref). \leavevmode\unskip\penalty9999 \hbox\nobreak \quad\hbox{\ensuremath{\square}}
remark[Nonlinear models] The methodology presented here also includes nonlinear models because the parameter $\delta$ does not need to be interpretable by itself. For example, suppose the model in Example (ref) is the latent model in a binary choice framework with symmetric link function $F$ and $\beta_k \equiv \beta$. Then $F(\theta_0 + \delta + \beta' x) - F(\theta_0 + \beta' x)$ for some $x$ may be the treatment effect of interest but $H_0\colon \delta = 0$ still determines whether the treatment effect is zero or not. Estimates of $\theta_0$ and $\theta_1 = \theta_0 + \delta$ from these models typically do not have closed form in the presence of covariates but generally have asymptotic linear representations to which the same types of arguments as in Example (ref) can be applied. \leavevmode\unskip\penalty9999 \hbox\nobreak \quad\hbox{\ensuremath{\square}}

Before concluding this section, I present a brief summary of how the rearrangement test can be implemented in practice. By Theorem (ref), the following procedure provides an asymptotically $\alpha$-level test in the presence of a finite number of large clusters when only a single cluster received treatment. The test is computationally simple and does not require simulation or resampling, can be two-sided or one-sided in either direction, is able to detect all fixed alternatives, and is powerful against $1/\sqrt{n}$-local alternatives. Recall that $\varrho$ here measures how much more variable the estimate from the treated cluster $\smash{\hat{\theta}_1}$ can be relative to the second-least variable control cluster estimate $\hat{\theta}_{0,k}$. A $\varrho$ of $5$ means that the (asymptotic) variance of $\smash{\hat{\theta}_1}$ can be up to $5^2 = 25$ times larger. There is no restriction on how much less variable $\smash{\hat{\theta}_1}$ can be than any of the other estimates and $\smash{\hat{\theta}_1}$ can be infinitely more variable than the least variable control cluster. (See also the discussion above Theorem (ref).)

algorithm[algorithm omitted — 1,601 chars of source]

This test can also be used as a “robustness check” if inference was originally performed with a method designed for a finer level of clustering, e.g., at the county level instead of the state level. In that case Algorithm (ref) can illustrate how well the results of the original test hold up if there is dependence across counties. As I point out in Section (ref), one could start at $\varrho = 0$ or $\varrho = 1$ and increase $\varrho$ until the null hypothesis can no longer be rejected. This is informative because a result that holds up to a potentially $\varrho^2 = 25$ times larger variance is more credible than a result that only holds if $\varrho^2 = 1$, i.e., if $\hat{\theta}_1$ cannot be more variable than all but one $\hat{\theta}_{0,k}$. If the rearrangement test is used in difference-in-differences models in conjunction with the popular conleytaber2011 test, it is important to note that $\varrho^2 = 1$ still allows for substantial heterogeneity whereas the \citetalias{conleytaber2011} test presumes full homogeneity across clusters.

An R command that implements Algorithm (ref) and the robustness check for any choice of $\varrho$ is available at \href{https://hgmn.github.io/rea}{https://hgmn.github.io/rea}. The next section shows how the rearrangement test performs in simulations and an application.

Numerical results

This section explores the finite-sample behavior of the rearrangement test in two experiments. Example (ref) compares the rearrangement test to the widely used conleytaber2011 test in the two-way fixed effects model with clusters. Example (ref) applies the rearrangement test as a robustness check for the results of garthwaiteetal2014. The discussion focuses on one-sided tests to the right but the results apply more generally.

example[Two-way fixed effects; conleytaber2011] This example uses a Monte Carlo experiment to compare rearrangement to the conleytaber2011 (conleytaber2011) test. The \citetalias{conleytaber2011} test is designed specifically for difference in differences and applies to models with a single treated cluster. Following conleytaber2011, the data are generated from the two-way fixed effects model \begin{equation} Y_{t,k} = \delta I_t D_k + \eta_t + \zeta_k + U_{t, k}, \end{equation} where $I_t$ is a post-intervention indicator, $D_k$ is a treatment indicator, and $\eta_t$ and $\zeta_k$ are time and cluster fixed effects, respectively. The error term satisfies \begin{equation} U_{t,k} = \gamma U_{t-1,k} + \sigma^{1\{k=q+1\}} V_{t, k},\end{equation} where the $V_{t, k}$ are iid copies of a standard normal variable and $k=q+1$ is the one cluster that received treatment. The model uses $\eta_t \equiv 0 \equiv \zeta_k$, ten time periods with four post-intervention periods, and, unless stated otherwise, $\gamma = .5$ and $\delta = 0$. I do not consider all of conleytaber2011's variations of their model and, to focus on the simplest possible situation, I do not include covariates. I expand upon their analysis by investigating smaller numbers of control clusters $q$ and values of $\sigma$ other than one. In the latter situation, the \citetalias{conleytaber2011} test can be expected to fail because it relies heavily on homogeneity of all clusters in absence of an intervention. The \citetalias{conleytaber2011} test can be restored (as $q\to \infty$) if the exact form of heterogeneity is known fermanpinto2019, ferman2020 but this is not assumed here. The \citetalias{conleytaber2011} test with one treated cluster can be computed as follows: (1) Regress the outcome on $I_t D_k$, time and cluster fixed effects, and other covariates (if available). Denote the coefficient on $I_t D_k$ by $\hat\delta$. (2) Split the residuals by cluster and run, for each of the $q$ control clusters separately, regressions of the residuals on a constant and $I_t$. (3) Compute the $1-\alpha$ empirical quantile of the $q$ coefficients on $I_t$. Reject $H_0\colon \delta = 0$ if $\hat\delta$ is larger than that quantile. The rearrangement test can be computed similarly from $q+1$ separate artificial regressions of $Y_{t,k}$ on a constant and the post-intervention indicator $I_t$, \begin{align*} Y_{t, k} &= \zeta + \theta_0 I_t + \mathrm{error}_{t,k}, \qquad 1\leqslant k\leqslant q,\\ Y_{t, k} &= \zeta + \theta_1 I_t + \mathrm{error}_{t,k}, \qquad k = q+1, \end{align*} where $\zeta$ is the intercept in each regression. The coefficients on the post-intervention indicator can be expressed as $\theta_0 = \bar{\eta}_+ - \bar{\eta}_-$ and $\theta_1 = \delta + \bar{\eta}_+ - \bar{\eta}_-$, where $\bar{\eta}_-$ and $\bar{\eta}_+$ are time averages of $\eta_t$ pre and post intervention, respectively. Because $\delta = \theta_1 - \theta_0$, I apply the rearrangement test to the least squares estimates $\hat{\theta}_{0,1}, \dots, \hat{\theta}_{0,q}$ and $\hat{\theta}_1$ of $\theta_0$ and $\theta_1$, respectively. I view (ref) as coming from individual-level data aggregated to the cluster level with a fixed number of time periods. The estimates $\hat{\theta}_1, \hat{\theta}_{0,1}, \dots, \hat{\theta}_{0,q}$ should therefore be approximately normal for the rearrangement test to apply. To test deviations from this assumption in finite samples, I also consider a situation where the innovations $V_{t,k}$ in (ref) are $\chi^2_2/2$ variables centered at zero. These innovations are asymmetric but still have unit variance. \begin{figure} \resizebox{\textwidth}{!}{ \begin{tikzpicture}[x=1pt,y=1pt] \definecolor{fillColor}{RGB}{255,255,255} \path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (433.62,216.81); \begin{scope} \path[clip] ( 48.00, 42.00) rectangle (204.81,192.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 53.81, 75.89) -- ( 58.65, 79.38) -- ( 63.49, 82.91) -- ( 68.33, 86.27) -- ( 73.17, 91.39) -- ( 78.01, 93.39) -- ( 82.85, 98.46) -- ( 87.69,102.42) -- ( 92.53,105.77) -- ( 97.37,108.89) -- (102.21,116.29) -- (107.05,118.38) -- (111.89,123.18) -- (116.73,122.29) -- (121.57,126.53) -- (126.41,132.95) -- (131.24,135.14) -- (136.08,136.49) -- (140.92,137.09) -- (145.76,145.47) -- (150.60,144.68) -- (155.44,148.36) -- (160.28,149.06) -- (165.12,154.55) -- (169.96,157.48) -- (174.80,156.13) -- (179.64,158.13) -- (184.48,163.63) -- (189.32,164.74) -- (194.16,167.77) -- (199.00,171.03); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (433.62,216.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- (199.00, 42.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- ( 53.81, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (102.21, 42.00) -- (102.21, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (150.60, 42.00) -- (150.60, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (199.00, 42.00) -- (199.00, 36.00); \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 53.81, 20.40) {1.0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (102.21, 20.40) {1.5}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (150.60, 20.40) {2.0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (199.00, 20.40) {2.5}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 48.00,187.22); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 42.00, 47.59); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 70.86) -- ( 42.00, 70.86); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 94.13) -- ( 42.00, 94.13); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,117.41) -- ( 42.00,117.41); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,140.68) -- ( 42.00,140.68); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,163.95) -- ( 42.00,163.95); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,187.22) -- ( 42.00,187.22); \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 47.59) {0.00}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 94.13) {0.10}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,140.68) {0.20}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,187.22) {0.30}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 42.00) -- (204.81, 42.00) -- (204.81,192.81) -- ( 48.00,192.81) -- ( 48.00, 42.00); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (216.81,216.81); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.20] at (126.41,200.68) {Conley-Taber (size, $\delta = 0$)}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (126.41, 2.40) {$\sigma$}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 15.60,117.41) {Rejection frequency}; \end{scope} \begin{scope} \path[clip] ( 48.00, 42.00) rectangle (204.81,192.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] ( 53.81, 74.30) -- ( 58.65, 80.63) -- ( 63.49, 83.33) -- ( 68.33, 86.64) -- ( 73.17, 90.83) -- ( 78.01, 94.04) -- ( 82.85, 99.53) -- ( 87.69,103.12) -- ( 92.53,104.98) -- ( 97.37,111.77) -- (102.21,110.33) -- (107.05,113.50) -- (111.89,120.48) -- (116.73,123.97) -- (121.57,122.34) -- (126.41,128.16) -- (131.24,128.76) -- (136.08,134.49) -- (140.92,137.14) -- (145.76,140.40) -- (150.60,141.24) -- (155.44,142.40) -- (160.28,148.54) -- (165.12,149.01) -- (169.96,151.06) -- (174.80,150.87) -- (179.64,153.62) -- (184.48,157.30) -- (189.32,158.55) -- (194.16,160.97) -- (199.00,163.53); \path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] ( 53.81, 75.19) -- ( 58.65, 78.45) -- ( 63.49, 78.82) -- ( 68.33, 84.12) -- ( 73.17, 89.15) -- ( 78.01, 93.81) -- ( 82.85, 94.78) -- ( 87.69, 98.97) -- ( 92.53,104.28) -- ( 97.37,107.21) -- (102.21,109.17) -- (107.05,113.40) -- (111.89,114.61) -- (116.73,120.48) -- (121.57,118.66) -- (126.41,125.32) -- (131.24,126.57) -- (136.08,128.02) -- (140.92,130.90) -- (145.76,139.00) -- (150.60,138.91) -- (155.44,142.12) -- (160.28,144.45) -- (165.12,145.84) -- (169.96,146.45) -- (174.80,149.66) -- (179.64,157.81) -- (184.48,154.92) -- (189.32,156.46) -- (194.16,159.67) -- (199.00,162.93); \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 48.00, 70.86) -- (204.81, 70.86); \definecolor{fillColor}{RGB}{255,255,255} \path[fill=fillColor] ( 48.97,191.88) rectangle (128.62,134.88); \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 57.52,180.48) -- ( 74.62,180.48); \path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] ( 57.52,169.08) -- ( 74.62,169.08); \path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] ( 57.52,157.68) -- ( 74.62,157.68); \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 57.52,146.28) -- ( 74.62,146.28); \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,177.21) {$q=50, N$}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,165.81) {$q=15, N$}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,154.41) {$q=50, \chi^2_2$}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,143.01) {5% level}; \end{scope} \begin{scope} \path[clip] (264.81, 42.00) rectangle (421.62,192.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (270.62, 48.33) -- (275.46, 48.61) -- (280.30, 48.89) -- (285.14, 49.31) -- (289.98, 49.73) -- (294.82, 50.61) -- (299.66, 50.61) -- (304.50, 52.75) -- (309.34, 52.71) -- (314.18, 54.06) -- (319.02, 54.89) -- (323.86, 56.85) -- (328.70, 57.08) -- (333.54, 59.08) -- (338.38, 60.01) -- (343.22, 61.88) -- (348.05, 64.20) -- (352.89, 64.57) -- (357.73, 65.97) -- (362.57, 69.23) -- (367.41, 69.37) -- (372.25, 71.32) -- (377.09, 73.00) -- (381.93, 76.49) -- (386.77, 77.75) -- (391.61, 80.03) -- (396.45, 79.66) -- (401.29, 83.66) -- (406.13, 85.71) -- (410.97, 86.78) -- (415.81, 90.45); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (433.62,216.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (415.81, 42.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (270.62, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (319.02, 42.00) -- (319.02, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (367.41, 42.00) -- (367.41, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (415.81, 42.00) -- (415.81, 36.00); \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (270.62, 20.40) {1.0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (319.02, 20.40) {1.5}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (367.41, 20.40) {2.0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (415.81, 20.40) {2.5}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (264.81,187.22); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (258.81, 47.59); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 70.86) -- (258.81, 70.86); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 94.13) -- (258.81, 94.13); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,117.41) -- (258.81,117.41); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,140.68) -- (258.81,140.68); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,163.95) -- (258.81,163.95); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,187.22) -- (258.81,187.22); \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 47.59) {0.00}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 94.13) {0.10}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,140.68) {0.20}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,187.22) {0.30}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 42.00) -- (421.62, 42.00) -- (421.62,192.81) -- (264.81,192.81) -- (264.81, 42.00); \end{scope} \begin{scope} \path[clip] (216.81, 0.00) rectangle (433.62,216.81); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.20] at (343.21,200.68) {Rearrangement (size, $\delta = 0$)}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (343.21, 2.40) {$\sigma$}; \end{scope} \begin{scope} \path[clip] (264.81, 42.00) rectangle (421.62,192.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (270.62, 48.47) -- (275.46, 48.47) -- (280.30, 48.89) -- (285.14, 48.80) -- (289.98, 49.45) -- (294.82, 49.49) -- (299.66, 50.38) -- (304.50, 52.01) -- (309.34, 51.73) -- (314.18, 53.17) -- (319.02, 52.98) -- (323.86, 53.68) -- (328.70, 55.45) -- (333.54, 56.71) -- (338.38, 55.87) -- (343.22, 58.62) -- (348.05, 59.08) -- (352.89, 60.34) -- (357.73, 61.27) -- (362.57, 63.83) -- (367.41, 64.44) -- (372.25, 65.88) -- (377.09, 66.16) -- (381.93, 67.79) -- (386.77, 68.76) -- (391.61, 71.23) -- (396.45, 73.70) -- (401.29, 74.40) -- (406.13, 74.54) -- (410.97, 76.58) -- (415.81, 79.84); \path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (270.62, 48.80) -- (275.46, 48.98) -- (280.30, 50.10) -- (285.14, 50.38) -- (289.98, 50.61) -- (294.82, 51.40) -- (299.66, 51.17) -- (304.50, 52.75) -- (309.34, 53.31) -- (314.18, 54.38) -- (319.02, 55.54) -- (323.86, 56.89) -- (328.70, 56.80) -- (333.54, 58.90) -- (338.38, 59.22) -- (343.22, 60.20) -- (348.05, 62.71) -- (352.89, 64.16) -- (357.73, 64.53) -- (362.57, 67.74) -- (367.41, 67.60) -- (372.25, 70.44) -- (377.09, 70.77) -- (381.93, 70.81) -- (386.77, 71.28) -- (391.61, 75.19) -- (396.45, 78.21) -- (401.29, 79.00) -- (406.13, 79.10) -- (410.97, 80.35) -- (415.81, 83.71); \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (264.81, 70.86) -- (421.62, 70.86); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (367.41, 42.00) -- (367.41,192.81); \end{scope} \end{tikzpicture} } \caption{Rejection frequencies of a true null as a function of the heterogeneity $\sigma$ for the \citetalias{conleytaber2011} test (left) and the rearrangement test (right) with (i) $q=50$ control clusters and normal errors (solid lines), (ii) $q=15$ and normal errors (long-dashed), and (iii) $q=50$ and chi-squared errors (dotted). The short-dashed line equals $.05$. The rearrangement test uses $\varrho = 2$ (vertical line).} \end{figure} Figure (ref) shows the rejection frequencies of a true null hypothesis $H_0 \colon \delta = 0$ as a function of $\sigma \in \{1, 1.05, 1.1, \dots, 2.5\} $ for the two tests at the 5% level (short-dashed lines). The assumptions of the \citetalias{conleytaber2011} test (left) hold as $q\to \infty$ when $\sigma = 1$ but are violated at any sample size as soon as $\sigma > 1$. The rearrangement test (right) here uses $\varrho = 2$ (vertical line). The assumptions of the rearrangement test are violated as soon as $\sigma > 2$. The figure shows rejection rates in 10,000 Monte Carlo experiments for each horizontal coordinate with (i) $q=50$ control clusters (solid lines), (ii) $q=15$ (long-dashed), and (iii) $q=50$ but the $V_{t,k}$ are iid copies of a $(\chi_2^2-2)/2$ variable (dotted). Both methods were faced with the same data. As can be seen, the \citetalias{conleytaber2011} test over-rejected slightly at $\sigma=1$ but quickly became unusable as $\sigma$ increased. It exceeded a 10% rejection rate at about $\sigma = 1.25$. At $\sigma = 2.5$, the \citetalias{conleytaber2011} test falsely discovered a nonzero effect in about 25% of all cases. In contrast, the rearrangement test was able to reject at or below the nominal level of the test as long as $\sigma \leqslant \varrho$. For $\sigma > \varrho$, the rearrangement test eventually started to over-reject. It performed worst at $\sigma = 2.5$, where it rejected in 6.9-9.2% of all cases. I also conducted a large number of additional experiments under the null. I considered (not shown) other distributions for $V_{t,k}$ and other values of the AR(1) coefficient $\gamma$, the number of time periods, the number of post-intervention periods, and the number of control clusters. However, I found that these changes had little impact on the results in the preceding paragraph. The \citetalias{conleytaber2011} test performed well when there was no heterogeneity but over-rejected wildly otherwise. More results in this direction can be found in canayetal2014, who come to the same conclusion in their experiments. The rearrangement test continued to be highly robust to heterogeneity as long as $\varrho$ was not chosen to be much too small. \begin{figure} \resizebox{\textwidth}{!}{ \begin{tikzpicture}[x=1pt,y=1pt] \definecolor{fillColor}{RGB}{255,255,255} \path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (433.62,216.81); \begin{scope} \path[clip] ( 48.00, 42.00) rectangle (204.81,192.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 53.81, 67.66) -- ( 58.65, 68.50) -- ( 63.49, 69.09) -- ( 68.33, 71.38) -- ( 73.17, 73.23) -- ( 78.01, 72.76) -- ( 82.85, 74.92) -- ( 87.69, 76.94) -- ( 92.53, 77.82) -- ( 97.37, 78.67) -- (102.21, 80.56) -- (107.05, 82.08) -- (111.89, 82.67) -- (116.73, 81.54) -- (121.57, 84.90) -- (126.41, 86.04) -- (131.24, 87.30) -- (136.08, 87.68) -- (140.92, 88.03) -- (145.76, 89.62) -- (150.60, 90.56) -- (155.44, 90.68) -- (160.28, 91.75) -- (165.12, 93.86) -- (169.96, 94.02) -- (174.80, 94.35) -- (179.64, 94.77) -- (184.48, 95.97) -- (189.32, 96.41) -- (194.16, 98.12) -- (199.00, 99.06); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (433.62,216.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- (199.00, 42.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- ( 53.81, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (102.21, 42.00) -- (102.21, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (150.60, 42.00) -- (150.60, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (199.00, 42.00) -- (199.00, 36.00); \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 53.81, 20.40) {1.0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (102.21, 20.40) {1.5}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (150.60, 20.40) {2.0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (199.00, 20.40) {2.5}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 48.00,187.22); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 42.00, 47.59); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 82.50) -- ( 42.00, 82.50); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,117.41) -- ( 42.00,117.41); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,152.31) -- ( 42.00,152.31); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,187.22) -- ( 42.00,187.22); \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 47.59) {0.0}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 82.50) {0.2}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,117.41) {0.4}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,152.31) {0.6}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,187.22) {0.8}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 42.00) -- (204.81, 42.00) -- (204.81,192.81) -- ( 48.00,192.81) -- ( 48.00, 42.00); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (216.81,216.81); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.20] at (126.41,200.68) {Rearrangement (power, $\delta = 2$)}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (126.41, 2.40) {$\sigma$}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 15.60,117.41) {Rejection frequency}; \end{scope} \begin{scope} \path[clip] ( 48.00, 42.00) rectangle (204.81,192.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] ( 53.81, 61.46) -- ( 58.65, 63.31) -- ( 63.49, 63.77) -- ( 68.33, 65.53) -- ( 73.17, 66.24) -- ( 78.01, 66.28) -- ( 82.85, 67.71) -- ( 87.69, 69.04) -- ( 92.53, 69.72) -- ( 97.37, 72.02) -- (102.21, 71.17) -- (107.05, 72.51) -- (111.89, 73.49) -- (116.73, 75.50) -- (121.57, 74.50) -- (126.41, 76.46) -- (131.24, 76.46) -- (136.08, 78.64) -- (140.92, 78.71) -- (145.76, 79.82) -- (150.60, 80.35) -- (155.44, 80.10) -- (160.28, 82.81) -- (165.12, 82.70) -- (169.96, 83.30) -- (174.80, 83.37) -- (179.64, 84.26) -- (184.48, 85.39) -- (189.32, 85.67) -- (194.16, 87.07) -- (199.00, 87.35); \path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] ( 53.81, 59.70) -- ( 58.65, 59.94) -- ( 63.49, 61.03) -- ( 68.33, 61.60) -- ( 73.17, 62.39) -- ( 78.01, 64.88) -- ( 82.85, 65.02) -- ( 87.69, 66.14) -- ( 92.53, 68.08) -- ( 97.37, 68.04) -- (102.21, 68.79) -- (107.05, 70.64) -- (111.89, 71.36) -- (116.73, 72.79) -- (121.57, 72.00) -- (126.41, 73.38) -- (131.24, 75.02) -- (136.08, 76.04) -- (140.92, 76.28) -- (145.76, 78.50) -- (150.60, 78.24) -- (155.44, 79.58) -- (160.28, 80.37) -- (165.12, 79.82) -- (169.96, 81.43) -- (174.80, 82.32) -- (179.64, 84.55) -- (184.48, 84.35) -- (189.32, 84.29) -- (194.16, 85.60) -- (199.00, 85.99); \definecolor{drawColor}{RGB}{169,169,169} \path[draw=drawColor,line width= 1.6pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 53.81,102.18) -- ( 58.65,103.41) -- ( 63.49,103.48) -- ( 68.33,105.83) -- ( 73.17,105.80) -- ( 78.01,105.20) -- ( 82.85,107.18) -- ( 87.69,109.29) -- ( 92.53,109.32) -- ( 97.37,109.93) -- (102.21,110.11) -- (107.05,110.44) -- (111.89,111.03) -- (116.73,109.97) -- (121.57,111.99) -- (126.41,113.48) -- (131.24,113.83) -- (136.08,114.93) -- (140.92,113.16) -- (145.76,115.69) -- (150.60,115.76) -- (155.44,114.44) -- (160.28,116.25) -- (165.12,116.90) -- (169.96,117.51) -- (174.80,117.42) -- (179.64,116.93) -- (184.48,118.52) -- (189.32,118.61) -- (194.16,119.90) -- (199.00,120.55); \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 53.81, 54.15) -- ( 58.65, 54.85) -- ( 63.49, 56.02) -- ( 68.33, 57.20) -- ( 73.17, 57.53) -- ( 78.01, 58.74) -- ( 82.85, 59.91) -- ( 87.69, 61.72) -- ( 92.53, 62.42) -- ( 97.37, 63.66) -- (102.21, 64.83) -- (107.05, 65.16) -- (111.89, 67.13) -- (116.73, 67.05) -- (121.57, 69.49) -- (126.41, 69.44) -- (131.24, 71.29) -- (136.08, 71.95) -- (140.92, 72.56) -- (145.76, 73.79) -- (150.60, 75.08) -- (155.44, 76.33) -- (160.28, 76.70) -- (165.12, 78.29) -- (169.96, 78.20) -- (174.80, 79.27) -- (179.64, 81.52) -- (184.48, 81.66) -- (189.32, 81.90) -- (194.16, 83.33) -- (199.00, 83.77); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 48.00, 56.31) -- (204.81, 56.31); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (150.60, 42.00) -- (150.60,192.81); \definecolor{fillColor}{RGB}{255,255,255} \path[fill=fillColor] ( 48.97,192.46) rectangle (158.34,124.06); \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 57.52,181.06) -- ( 74.62,181.06); \path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] ( 57.52,169.66) -- ( 74.62,169.66); \definecolor{drawColor}{RGB}{169,169,169} \path[draw=drawColor,line width= 1.6pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 57.52,158.26) -- ( 74.62,158.26); \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 57.52,146.86) -- ( 74.62,146.86); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] ( 57.52,135.46) -- ( 74.62,135.46); \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,177.79) {$q=50, N, \gamma = .5$}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,166.39) {$q = 15, N, \gamma = .5$}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,154.99) {$q=50, N, \gamma = .1$}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,143.59) {$q=50, N, \gamma = .9$}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.95] at ( 83.17,132.19) {$q = 50, \chi^2_2, \gamma = .5$}; \end{scope} \begin{scope} \path[clip] (264.81, 42.00) rectangle (421.62,192.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (270.62,111.51) -- (275.46,112.34) -- (280.30,112.48) -- (285.14,112.97) -- (289.98,113.90) -- (294.82,112.92) -- (299.66,115.29) -- (304.50,116.71) -- (309.34,116.51) -- (314.18,116.92) -- (319.02,116.83) -- (323.86,116.93) -- (328.70,117.27) -- (333.54,116.48) -- (338.38,117.61) -- (343.22,118.78) -- (348.05,119.78) -- (352.89,119.90) -- (357.73,119.13) -- (362.57,120.29) -- (367.41,121.63) -- (372.25,121.00) -- (377.09,120.67) -- (381.93,121.66) -- (386.77,122.47) -- (391.61,121.28) -- (396.45,121.21) -- (401.29,122.61) -- (406.13,123.64) -- (410.97,123.71) -- (415.81,124.00); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (433.62,216.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (415.81, 42.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (270.62, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (319.02, 42.00) -- (319.02, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (367.41, 42.00) -- (367.41, 36.00); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (415.81, 42.00) -- (415.81, 36.00); \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (270.62, 20.40) {1.0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (319.02, 20.40) {1.5}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (367.41, 20.40) {2.0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (415.81, 20.40) {2.5}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (264.81,187.22); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (258.81, 47.59); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 82.50) -- (258.81, 82.50); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,117.41) -- (258.81,117.41); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,152.31) -- (258.81,152.31); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,187.22) -- (258.81,187.22); \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 47.59) {0.0}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 82.50) {0.2}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,117.41) {0.4}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,152.31) {0.6}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,187.22) {0.8}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 42.00) -- (421.62, 42.00) -- (421.62,192.81) -- (264.81,192.81) -- (264.81, 42.00); \end{scope} \begin{scope} \path[clip] (216.81, 0.00) rectangle (433.62,216.81); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.20] at (343.21,200.68) {Rearrangement (power, $\delta = 3$)}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (343.21, 2.40) {$\sigma$}; \end{scope} \begin{scope} \path[clip] (264.81, 42.00) rectangle (421.62,192.81); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (270.62, 91.21) -- (275.46, 93.18) -- (280.30, 92.93) -- (285.14, 95.29) -- (289.98, 95.24) -- (294.82, 95.29) -- (299.66, 95.81) -- (304.50, 97.45) -- (309.34, 97.10) -- (314.18,100.53) -- (319.02, 99.04) -- (323.86, 99.88) -- (328.70,101.36) -- (333.54,102.22) -- (338.38,100.51) -- (343.22,102.69) -- (348.05,103.41) -- (352.89,104.26) -- (357.73,105.29) -- (362.57,105.20) -- (367.41,106.29) -- (372.25,104.75) -- (377.09,106.91) -- (381.93,106.20) -- (386.77,106.39) -- (391.61,106.93) -- (396.45,106.46) -- (401.29,108.94) -- (406.13,109.18) -- (410.97,108.61) -- (415.81,109.38); \path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (270.62, 86.67) -- (275.46, 85.69) -- (280.30, 86.88) -- (285.14, 87.28) -- (289.98, 87.98) -- (294.82, 91.14) -- (299.66, 89.93) -- (304.50, 91.89) -- (309.34, 94.38) -- (314.18, 93.42) -- (319.02, 95.29) -- (323.86, 95.17) -- (328.70, 96.74) -- (333.54, 97.61) -- (338.38, 95.95) -- (343.22, 98.24) -- (348.05, 98.48) -- (352.89,100.58) -- (357.73, 99.93) -- (362.57,101.99) -- (367.41,101.82) -- (372.25,102.46) -- (377.09,104.17) -- (381.93,103.35) -- (386.77,102.76) -- (391.61,104.56) -- (396.45,106.86) -- (401.29,106.08) -- (406.13,105.15) -- (410.97,106.58) -- (415.81,106.67); \definecolor{drawColor}{RGB}{169,169,169} \path[draw=drawColor,line width= 1.6pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (270.62,180.87) -- (275.46,178.51) -- (280.30,177.01) -- (285.14,175.27) -- (289.98,174.52) -- (294.82,172.39) -- (299.66,172.58) -- (304.50,171.81) -- (309.34,171.22) -- (314.18,170.43) -- (319.02,169.18) -- (323.86,167.40) -- (328.70,167.94) -- (333.54,165.23) -- (338.38,165.53) -- (343.22,165.14) -- (348.05,163.54) -- (352.89,164.08) -- (357.73,161.55) -- (362.57,162.60) -- (367.41,161.25) -- (372.25,161.09) -- (377.09,160.80) -- (381.93,159.12) -- (386.77,159.16) -- (391.61,158.95) -- (396.45,158.98) -- (401.29,157.53) -- (406.13,157.97) -- (410.97,157.52) -- (415.81,158.35); \path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (270.62, 68.15) -- (275.46, 68.74) -- (280.30, 70.61) -- (285.14, 71.17) -- (289.98, 72.44) -- (294.82, 74.01) -- (299.66, 74.66) -- (304.50, 77.45) -- (309.34, 78.55) -- (314.18, 79.49) -- (319.02, 81.17) -- (323.86, 81.26) -- (328.70, 82.36) -- (333.54, 82.18) -- (338.38, 84.75) -- (343.22, 84.96) -- (348.05, 87.14) -- (352.89, 87.61) -- (357.73, 89.39) -- (362.57, 89.55) -- (367.41, 91.19) -- (372.25, 91.43) -- (377.09, 92.11) -- (381.93, 93.47) -- (386.77, 93.33) -- (391.61, 94.00) -- (396.45, 96.08) -- (401.29, 96.49) -- (406.13, 96.51) -- (410.97, 97.61) -- (415.81, 98.27); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (264.81, 56.31) -- (421.62, 56.31); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (367.41, 42.00) -- (367.41,192.81); \end{scope} \end{tikzpicture} } \caption{Rejection frequencies of the rearrangement test ($\varrho = 2$) under the alternative as a function of the heterogeneity $\sigma$ at $\delta = 2$ (left) and $\delta = 3$ (right) with (i) and (ii) as in Figure (ref), (iii) is (i) with weak time dependence $\gamma=.1$ (short-dashed grey), (iv) is (i) with strong time dependence $\gamma=.9$ (solid grey) (v) is (i) with chi-squared errors (dotted). The short-dashed line equals $.05$.} \end{figure} I now turn to the performance of the rearrangement test under the alternative. The behavior of the \citetalias{conleytaber2011} test under the alternative is not discussed due to its massive size distortion. I consider the same models as before together with some variations mentioned in the preceding paragraph but use nonzero $\delta$. Figure (ref) shows the results with $\delta = 2$ (left) and $\delta = 3$ (right). The base model is again model (i) with $q=50$ control clusters, standard normal $V_{t,k}$, and time dependence set to $\gamma = .5$ (solid lines). The other models deviate from (i) in the following ways: (ii) uses $q=15$ (long-dashed), (iii) lowers the time dependence to $\gamma = .1$ (short-dashed grey), (iv) increases the time dependence to $\gamma = .9$ (solid grey), and (v) changes the innovations to $(\chi_2^2-2)/2$ (dotted). As can be seen, having to guard against near arbitrary heterogeneity of unknown form made it difficult to detect a relatively small treatment effect (left) when the number of control clusters was low, the distribution of the innovations was non-normal, or the treatment effect was obfuscated by strong time dependence. However, the rearrangement test reliably detected smaller treatment effects when the time dependence was relatively weak. Increasing the treatment effect (right) improved detection rates substantially and uniformly across models, with strong time dependence again being the most challenging situation. The rearrangement test now had considerable power even when only 15 control clusters were available, the innovations were asymmetric, or the time dependence was not extreme. Power was very high when there was little time dependence. Figures (ref) and (ref) also illustrate two noteworthy aspects of the rearrangement test: (1) The inequality the rearrangement is based on is nearly tight (as discussed below equation (ref)) in the sense that it cannot be meaningfully be improved upon unless $q$ is very small. This can be seen in the right panel of Figure (ref), where the rejection rate of the test was essentially at or slightly below nominal level when $\sigma = \varrho$. (2) Rejection rates under the null hypothesis increase with $\sigma$ but this does not necessarily translate into increased rejection rates under the alternative for large $\sigma$. This is seen in the right panel of Figure (ref), where the power decreases with $\sigma$ in the presence of weak time dependence ($\gamma = .1$). \leavevmode\unskip\penalty9999 \hbox\nobreak \quad\hbox{\ensuremath{\square}}
example[Health insurance and labor supply; garthwaiteetal2014] In this example, I use the rearrangement test to reanalyze the results of garthwaiteetal2014. They use a difference-in-differences design to study the effects of a large-scale disruption of public heath insurance on labor supply. Their design exploits that in 2005 approximately 170,000 adults in Tennessee (roughly 4% of the state's non-elderly, adult population) abruptly lost access to TennCare, the state's public health insurance system. garthwaiteetal2014\ use data from the 2001-2008 March Current Population Survey to determine health insurance and work status for the years 2000-2007. The comparison groups for Tennessee are the 16 other Southern states\footnote{The Southern states are Alabama, Arkansas, Delaware, the District of Columbia, Florida, Georgia, Kentucky, Louisiana, Maryland, Mississippi, North Carolina, Oklahoma, Tennessee, Texas, Virginia, South Carolina, and West Virginia.} defined by the U.S.\ Census Bureau. The main treatment effect in garthwaiteetal2014\ can be estimated as $\delta$ in \[ Y_{t, k} = \theta_0 I_t + \delta I_t D_k + \zeta_k + U_{t,k}, \] where $Y_{t,k}$ is a state-by-year mean of an outcome of interest for state $k$ in year $t$, $I_t = 1\{t\geqslant 2006\}$ is a post-intervention indicator, and $D_k$ equals one for an observation from Tennessee and equals zero otherwise. There are $17\times 8 = 136$ state-by-year means in total. garthwaiteetal2014\ estimate the model in the preceding display by least squares and conduct inference about $\delta$ with bootstrap standard errors that are compared to Student $t$ critical values with 16 degrees of freedom. Their preferred bootstrap first draws states with replacement and then draws individuals within those states with replacement. This type of inference accounts for autocorrelation within individuals over time but generally requires the number of clusters to be infinite for the asymptotics. This bootstrap also does not account for potential dependence within states. \begin{table}\caption{Effects of TennCare disenrollment in garthwaiteetal2014 with their auto-correlation robust bootstrap standard errors (top) and the largest $\varrho^2$ at which a rearrangement test robust to arbitrary correlation within states and over time still detects an effect (bottom).} { \scalebox{.92}{ \begin{tabular}{cp{0cm}cccccc} \hline & & (1) & (2) & (3) & (4) & (5) & (6) \\ \cline{3-8} & & & &Employed &Employed &Employed &Employed\\ & &Has public & &working & working & working & working\\ & &health & &$<$20 hours&$\geqslant$20 hours&20-35 hours &$\geqslant$35 hours \\ & &insurance &Employed &per week &per week &per week &per week \\ \cline{3-8} $\hat{\delta}$ &&$-$0.046\phantom{$-$} &0.025 &$-$0.001\phantom{$-$} &0.026 &0.001 &0.025 \\ s.e.&&(0.010) &(0.011) &(0.004) &(0.010) &(0.007) &(0.011)\\ $p$-val.&&[0.000] &[0.019] &[0.621] &[0.011] &[0.453] &[0.020] \\ \\ & &\multicolumn{6}{c}{Rearrangement test: largest $\varrho^2$ at which $H_0\colon \delta = 0$ is rejected}\\ $\alpha$ & &\multicolumn{6}{c}{(“$\times$” indicates that $H_0\colon \delta = 0$ cannot be rejected for any $\varrho \geqslant 0$)} \\ \cline{1-1}\cline{3-8} $.10$ & &5.434 &1.793 &$\times$ &2.208 &$\times$ &$\times$ \\ $.05$ & &2.914 &0.972 &$\times$ &1.195 &$\times$ &$\times$ \\ \hline \end{tabular} }} \end{table} I replicate the findings of garthwaiteetal2014 in the top panel of Table (ref). They estimate the causal effect of the TennCare disenrollment on the probability of (1) having public health insurance, (2) being employed, and (3)-(6) being employed for a certain number of hours per week. I show their bootstrap standard errors in parentheses but report one-sided $p$-values in brackets instead of their two-sided $p$-values. In (1) the alternative is a negative effect, for (2)-(6) the alternative is positive. garthwaiteetal2014\ find a highly significant 4.6 percentage point decrease for (1) and mostly significant positive effects for (2)-(6). They document an approximately 2.5 percentage point increase in employment and find the same effect if the outcome is restricted to individuals working more than 20 hours or more than 35 hours a week. All three effects are significant at the 5% level. The inference in garthwaiteetal2014\ shows no significant effect for individuals working less than 20 hours or 20-35 hours. I now apply the rearrangement test as a robustness check. I view each state over time as a single cluster and run 17 separate least squares regressions of the form \begin{align*} Y_{t, k} &= \theta_0 I_t + \zeta_k + U_{t,k}, \qquad 1\leqslant k\leqslant 16,\\ Y_{t, k} &= \theta_1 I_t + \zeta_k + U_{t,k}, \qquad k = 17, \end{align*} to obtain $\hat{\theta}_{0,k}$ ($1\leqslant k\leqslant 16$) from each of the Southern states except Tennessee and $\hat{\theta}_1$ from Tennessee ($k=17$). Note that the $\zeta_k$ are now the constant terms in each regression. To perform the robustness check, I start with $\varrho = 0$ and increase $\varrho$ by $.001$ in Algorithm (ref) as long as the null hypothesis $H_0\colon\delta = 0$ is still rejected. The bottom panel of Table (ref) shows the largest feasible value of $\varrho^2$ for outcomes (1)-(6). At the 10% level, the result in (1) survives an up to 5.4 times larger variance in the estimate from Tennessee relative to the second-least variable control cluster estimate. The result in (2) holds if Tennessee has a 1.8 times larger variance and (4) holds even with an up to 2.2 times larger variance. At the 5% level, these three results remain valid with smaller $\varrho^2$ but the result in (2) only survives if the estimate from Tennessee is at most slightly less variable than the second-least variable control cluster estimate. The results in (3) and (5) confirm findings in garthwaiteetal2014 in that they are not significant at any level and for any value of $\varrho$. A noteworthy situation occurs in (6), where the rearrangement test disagrees sharply with the significant effect found by garthwaiteetal2014. The rearrangement test finds no effect at any significance level and for any $\varrho$. In contrast, the effects in (2) and (6) are not only essentially identical but also have identical standard errors. (The $p$-values differ slightly because of rounding.) This also illustrates that the rearrangement test differs fundamentally from inference based on $t$ statistics and resampling. In sum, the rearrangement test robustly confirms---with one exception---the results of garthwaiteetal2014. There is statistical evidence of increased employment concentrated among individuals working at least 20 hours per week even if one accounts for arbitrary dependence within states and over time. The results hold up to substantial heterogeneity across clusters even if the number of clusters is treated a fixed for the analysis. It is also worth noting that $\varrho$ only restricts heterogeneity in one direction. All of the results presented here are robust to arbitrary heterogeneity in any other direction and to Tennessee being infinitely more variable than the least variable control cluster. \leavevmode\unskip\penalty9999 \hbox\nobreak \quad\hbox{\ensuremath{\square}}

Conclusion

I introduce a generic method for inference about a scalar parameter in research designs with a finite number of large, heterogeneous clusters where only a single cluster received treatment. This situation is commonplace in difference-in-differences estimation but the test developed here applies more generally. I show that the test asymptotically controls size and has power in a setting where the number of observations within each cluster is large but the number of clusters is fixed. The test combines independent, approximately Gaussian parameter estimates from each cluster with a weighting scheme and a rearrangement procedure to obtain its critical values. The weights needed for most empirically relevant situations are tabulated in the paper. The critical values are computationally simple and do not require simulation or resampling. The test is highly robust to situations where some clusters are much more variable than others. Examples and an empirical application are provided.