Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
66,639 characters · 14 sections · 44 citation commands
A Modified Randomization Test for the Level of Clustering
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Linear Regression, Clustered Standard Errors, Small-Cluster Asymptotics
\spacingset{1.45}
Consider the following regression:
where a researcher wants to perform inference on $\beta$. If the researcher is concerned about correlation between $U_{i}$ and $U_{i'}$, it is frequently helpful to group observations into independent clusters. These independent clusters can then be used to construct cluster-robust covariance estimators (CCE) as in lz1986, or for approximate randomization tests as in crs2017 and ccks2021.
However, these procedures require the assignment of units to clusters be known ex ante. In practice, researchers often have some freedom in choosing the level at which to cluster their standard errors. For example, those working with the American Community Survey (ACS) can cluster their data either at the individual, county or state level. Alternatively, those working with firm data from COMPUSTAT have the option to cluster firms at either the 4-digit, 3-digit or 2-digit Standard Industrial Classification (SIC) level.
Clustering at the correct level is important for valid inference. A large body of simulation evidence shows that ignoring cluster dependence -- in other words, clustering at too fine a level -- leads to type I errors that exceed the nominal error by as much as 10 times (bdm2004; cgm2008). On the other hand, clustering at excessively coarse levels can also lead to problems. For one, coarse clusters tend to be few in number. It is well-known that confidence intervals based on the cluster-robust standard errors tend to under-cover when the number of clusters is small (see mhe2008 for instance), leading to poor size control. In the absence of under-coverage issues, unnecessarily coarse levels of clustering can also lead to tests with poor power since the researcher assumes less information than they actually have. aaiw17 demonstrate via simulations, in a many-cluster setting, that CCEs based on coarse clusters can be too large. They also provide theoretical results in this vein, though they do so in the context of their “design-based" asymptotics that differ from those traditionally used to analyze clustered standard errors. Nonetheless, the problems with tests based on excessively coarse-clustering arise even with few clusters -- the setting of interest for our paper. We present a simple simulation to demonstrate these issues in Appendix (ref).
Given the above considerations, a researcher may choose to cluster at a fine level (e.g. individual or county) even when a coarse level of clustering (e.g. state), which is known to be valid, is also available. Nonetheless, they may be unsure if the fine level is appropriate. That is, whether observations across the fine clusters are approximately independent.
To help researchers assess the validity of their chosen clusters, we propose a modified randomization test that can be used as a robustness check for a given clustering specification. Our test requires large (fine) sub-clusters, but is justified under asymptotics that take the number of (coarse) clusters and (fine) sub-clusters as fixed. Inference is difficult in this setting because scores are not independent across sub-clusters even asymptotically, as we will explain in Section (ref). Randomization tests, which typically require some type of asymptotic independence, thus cannot be directly applied. We get around this problem by searching for worst-case values of the unobserved parameters to guard against over-rejection. We describe a simple method to search for this value, so that the computational complexity of the test is of the same order as the number of sub-clusters. This is reasonable since our test is targeted towards applications with few sub-clusters. Our test has no power against negative correlation. However, ignoring negative correlation leads to variance estimators that are too large, and is thus less of an issue if the researcher is concerned about size control when performing inference on $\beta$.
To our knowledge, there are two other tests for the level of clustering. mnw2020 proposes a test based on having large number of coarse clusters, relying on the wild bootstrap to improve finite sample performance. Meanwhile, im2016 proposes a test for the case when there are many sub-clusters. Our test, which takes the number of clusters and sub-clusters to be fixed, handles a more challenging situation, though this comes at the cost of being conservative, especially in settings with homogeneous clusters. However, as our simulations in Section (ref) show, it has competitive power given heterogeneous clusters -- a setting that could be relevant for empirical work. Indeed, our test detects correlation in the clusters chosen by gllqsx2019, demonstrating its potential usefulness in applied work (see Section (ref)). Finally, we note that the test of im2016 also has no power against negative correlation, although that of mnw2020 does not share this limitation.
aaiw17 takes a different approach to this issue. They argue for a “design-based" perspective on clustering, requiring researchers to determine ex ante the uncertainty that they face in either sampling or treatment assignment. For example, if the researcher believes that in their specific context, treatment assignment occurs at the sub-cluster level, then sub-clusters should be used for computing standard errors, regardless of whether or not residuals are correlated across the sub-clusters. While insightful, this approach requires researchers to answer an alternative question on which there is equally little theoretically guidance. We therefore develop our method under the “model-based" framework, in which the researcher has in mind some data-generating process that entails dependent clusters.
The remainder of this paper is organized as follows. Section (ref) describes our proposed test. Section (ref) presents Monte Carlo simulations. Section (ref) demonstrates an application to gllqsx2019. Section (ref) concludes. Proofs are collected in Appendix (ref).
In the following, we assume that the researcher has conducted inference on $\beta \in \mathbf{R}$, and seeks a robustness check for the level of clustering used for said inference. As will become clear in Section (ref), using a scalar $\beta$ yields computational advantages, though the test can be feasibly computed for moderate dimensions of $\beta$. For this reason and for ease of exposition we limit our discussion to the scalar case.
Consider the linear regression:
where $\beta \in \mathbf{R}$ is the parameter of interest and $\gamma \in \mathbf{R}^d$ is a nuisance parameter. Suppose there are $r$ clusters, indexed by $k \in \mathcal{K}$. Within each cluster $k$, there are $q_k$ sub-clusters, indexed by $j \in \mathcal{J}_k$. Within each sub-cluster $j$, there are $n_j$ individuals indexed by $i \in \mathcal{I}_j$. Let $\mathcal{J} = \bigcup_{k \in \mathcal{K}} \mathcal{J}_k$ and $\mathcal{I} = \bigcup_{j \in \mathcal{J}} \mathcal{I}_j$. Further, let $n = \sum_{j \in \mathcal{J}} n_j$ and $q = \sum_{k \in \mathcal{K}} q_k = |\mathcal{J}|$. We also write $i \in \mathcal{I}_k$ when $i \in \mathcal{I}_j$ and $j \in \mathcal{J}_k$. In the following, we suppress dependence on $j$ and $k$ whenever this does not cause confusion.
Suppose that within a sub-cluster, $W$ has full rank. Then $\hat{\Pi}_j$ can be chosen as the sub-cluster level OLS estimator of $X$ on $W$. Otherwise, we can just drop variables until we obtain a linearly independent subset $\tilde{W}$. The entries of $\hat{\Pi}_j$ corresponding to the dropped variables can then be set to $0$ while the remaining entries are chosen to be the corresponding coefficients from the sub-cluster level regression of $X$ on $\tilde{W}$. Alternatively, if the researcher is willing to assume that $\Pi_j$ is identical across clusters, $\hat{\Pi}_j$ can also be obtained from the full sample regression of $X$ on $W$. Now define
where $\hat{U}_i$ is the full-sample OLS residual using equation ((ref)). Suppose we know that clusters are independent, so that $E[Z_{i}Z_{i'}] = 0$ when $i \in \mathcal{I}_k, i' \in \mathcal{I}_{k'}$ and $k \neq k'$. Under this assumption, we test the null hypothesis that sub-clusters are uncorrelated:
against the alternative hypothesis that there exists sub-clusters within at least one cluster that exhibit correlation:
Note that changing the choice of $X_i$ and $W_i$ corresponds to testing different null hypotheses and could lead to differing outcomes. If a researcher wants to test the level of clustering used for inference on $\beta$, $X_i$ should be projected onto $W_i$. Similarly, if inference was conducted on $\gamma$, then $W_i$ should take the place of $X_i$ in equation ((ref)).
We further assume the following:
In other words, we assume that the errors are weakly correlated within each sub-cluster $j$. Imposing weak dependence within a (sub-)cluster is not an uncommon assumption (see for instance the discussion in crs2017 and bch2011). We note that under $H_0$, $\Omega$ is a diagonal matrix. On the other hand, under the alternative, it has a block diagonal structure due to correlation between sub-clusters.
In this subsection, we define the test statistic and explain the need to search over the worst case critical value. Before doing so, we first consider the infeasible test in which the true parameters -- $\beta$, $\gamma$ and $\Pi$ as defined in equations ((ref)) and ((ref)) -- are observed. Readers who are only interested in the details of implementation can skip to the end of Section (ref).
Suppose we know $\beta$, $\gamma$ and $\Pi$. Given $Y_i$ and $X_i$, we can back out $U_i$ and construct the vector $S_n^*$, whose $j^\text{th}$ entry is
Given $S^*_{n}$, we can then define the infeasible test statistic:
The inner sum is the net number of positive ${S}^*_{n,j}$ within each cluster $k$. Intuitively, if the sub-clusters are independent, the net number of positive $S^*_{n,j}$ should be close to $0$. Conversely, if they are positively correlated, this number will be large in absolute value, since many sub-clusters will have $S_{n,j}^*$ of the same sign. On the other hand, if they are negatively correlated, this number will be more concentrated around $0$ than in the independent case. As will become clear below, our test interprets large absolute values of $T(S_n^*)$ as violation of the null. For this reason, we it will not have power against negative correlation.
Next, denote by $\mathbf{G}$ the set of sign changes. $\mathbf{G}$ can be identified with the set of $g \in \{-1, 1\}^{q}$ so that
Now let $p^*(S_n^*)$ be the proportion of $T\left(g{S}^*_n\right)$ that are no smaller than $T\left({S}^*_n\right)$:
The test rejects the null hypothesis when $p(S_n^*)$ is small -- that is, when $T\left({S}^*_n\right)$ is extreme relative to $T\left(g{S}^*_n\right)$:
The intuition for the randomization test is as follows. Since $S^*_{n,j}$ involves only units within the same sub-cluster, under the null hypothesis, $S_n^*$ converges to a mean-zero normal distribution with independent components. Independence, together with symmetry of normal random variables about their means, implies that for any $g \in \mathbf{G}$, $gS^*_n$, has the same distribution as $S_n^*$. Hence, the randomization distribution $\{T(gS_n^*)\}_{g\in \mathbf{G}}$ is in fact the distribution of $T(S_n^*)$ conditional on the values of $|S_n^*|$, where $|\cdot|$ is applied component-wise. Rejecting the null hypothesis when we observe values of $T(S_n^*)$ that are extreme relative to $\{T(gS_n^*)\}_{g\in \mathbf{G}}$ therefore leads to a test with the correct size.
Note that the randomization test defined above is non-randomized. Randomization tests can also employ a randomized rejection rule for the situation when
Using a randomized rejection rule, we have that provided the necessary symmetry properties hold in finite sample, the randomization test will have size equal to $\alpha$ exactly. The test defined in equation ((ref)) is conservative since it never rejects when the above situation occurs. However, we present the deterministic version since the test that we propose is based on it.
Tests based on $S_n^*$ are infeasible since $\beta$, $\gamma$ and the $\Pi_j$'s are unknown. Suppose we simply replaced $Z_i$ with $\hat{Z}_i$ and performed the randomization test with the estimated scores. It turns out that this procedure is incorrect. To see this, let $\tilde{S}_n$ be $S_n^*$ but with $\hat{Z}_i$ replacing $Z_i$. Then we can write each component of $\tilde{S}_n$ as:
In the above equation, $S_{n,j}^*$ is the part that is informative about cluster structure. However, each component now has an additional nuisance term $A_j$ that does not go away under asymptotics that take the number of sub-clusters to be fixed. Because $\hat{\beta} - \beta$ is common across the $A_j$'s, it induces correlation across $\tilde{S}_{n,j}$ even when the $S_{n,j}^*$'s are independent, leading potentially to over-rejection. Addressing this complication which does not arise in frameworks taking $q \to \infty$ results in the conservativeness of our test.
If we knew $\hat{\beta} - \beta$, we could back out ${S}^*_{n,j}$ for the randomization test using equation ((ref)). Since that is not possible, we propose to search over values of $\hat{\beta} - \beta$ to ensure that the test controls size when the unobserved term takes on extreme values.
For a given $\lambda \in \mathbf{R}$, let $\hat{S}_n({\lambda})$ be $q\times 1$ vector whose $j^\text{th}$ entry is the following term:
Note that $\hat{S}_{n,j}(\hat{\beta} - \beta) = S^*_{n,j} + o_p(1)$. Define:
For a given $\lambda$, this is just the test statistic in equation ((ref)) but with $\hat{S}_{n,j}(\lambda)$ taking the place of $S_{n,j}^*$. As before, we denote by $\mathbf{G}$ the set of sign changes and write:
Now let $p(\hat{S}_n(\lambda))$ be the proportion of $T\left(g\hat{S}_n(\lambda)\right)$ that takes on extreme values relative to $T\left(\hat{S}_n(\lambda)\right)$:
We can then define the randomization test as:
We can then prove the following result:
The test is a two-stage process. In the first stage, it searches for the value of $\lambda$ that leads to the largest $p$-value. In the second stage, the test rejects if this worst-case $p$-value is still smaller than the desired level of significance $\alpha$. Since the worst-case $p$-value bounds the true $p$-value from above, the rejection rule based on the worst-case $p$-value must be conservative.
As the Monte Carlo simulations in Section (ref) shows, the test has size that could be much smaller than $\alpha$ under the null hypothesis. However, the same simulations also show that the test has reasonable power under the alternative hypothesis, particularly in settings where clusters are heterogeneous in their variances. The potential usefulness of our test is further seen in the empirical application (Section (ref)), where it detects dependence in the clusters chosen by gllqsx2019.
In this subsection we describe an efficient way of searching for $\lambda \in \mathbf{R}$. This search is simplified by the fact that $p(\hat{S}_n(\lambda))$ depends only on the sign of $\hat{S}_{n,j}(\lambda)$'s. As such, to find $\sup_{\lambda \in \mathbf{R}}$, we only need to search over sign combinations of $\hat{S}_{n,j}$. When $\beta$ is scalar, the search can be completed in $O(q)$ time. This is reasonable since the test is designed for use when $q$ is small.
Suppose for now that $\sum_{i \in \mathcal{I}_j} \left(X_i - W_i'\hat{\Pi}_j\right)^2 > 0$ for all $j \in \mathcal{J}$. Define:
Then, $ \hat{S}_{n,j}({\lambda}) \geq 0 \Leftrightarrow R_j + \lambda \geq 0~. $ Sort the values of $R_j$'s so that $ R^{(1)} \geq R^{(2)} \geq ... \geq R^{(q)}~. $ We must have that $R^{(j)} +\lambda \geq 0 \Rightarrow R^{(j')} +\lambda \geq 0$ for all $j' \leq j$. Let $\hat{S}_{n, (1)}(\lambda), ..., \hat{S}_{n, (q)}(\lambda)$ denote the values of $\hat{S}_{n,j}(\lambda)$ corresponding to $R^{(1)}, ..., R^{(q)}$. Therefore, we only need to consider sequences of the form
for some cut-off $j$. Since the $p$-value, as defined in equation ((ref)), depends only on the sign of $\hat{S}_n$, we can compute it using $\check{S}_n$ in the place of $\hat{S}_{n,j}(\lambda)$:
Here, we see that even when we are searching over the worst case $\lambda$, we are only allowed to choose the cut-off point at which the signs change. We can therefore complete the search with no more than $q$ randomization tests. Assuming that the time it takes for each test is $O(1)$, the procedure takes $O(q)$ time. The restriction that $\check{S}_{n, (j)} \geq \check{S}_{n, (j')}$ for all $j \leq j'$ also gives the test power. If all combinations of signs for the $S_{n,j}$'s were allowed, the test will always return a $p$-value of 1 and will have no power.
Finally, suppose there are sub-clusters such that $\sum_{i \in \mathcal{I}_j} \left(X_i - W_i'\hat{\Pi}_j\right)^2 = 0$. We can repeat the above procedure excluding these sub-clusters. In the final step, we set $\check{S}_{n,j}$ corresponding to these clusters to 0. Hence,
We summarise the implementation procedure in Algorithm (ref).
To our knowledge, two other tests have been proposed for the level of clustering. They take either the number of sub-clusters in each cluster to infinity or the number of clusters to infinity. We assume both to be fixed. For ease of exposition, we restrict our discussion of these tests to the univariate case.
im2016 (IM hereafter) adopts an asymptotic framework that takes $q_k \to \infty$ for all $k \in \mathcal{K}$. Consider estimating a regression coefficient cluster-by-cluster. Let $\hat{\beta}_k$ denote coefficients estimated using only cluster $k$. The IM test is based on the asymptotic distribution of an estimator for the variance of $\frac{1}{r} \sum_{k = 1}^r\hat{\beta}_k$. Let this variance be denoted by $V$ and let $\hat{\Omega}^\text{CCE}_k$ be the cluster-robust variance estimator for $\hat{\beta}_k$, where the clustering is done at the sub-cluster level using $j \in \mathcal{J}_k$. Under the null hypothesis, $\hat{\Omega}^\text{CCE}_k$ consistently estimates the variance of each $\hat{\beta}_k$.
Under either the null or the alternative, but maintaining the assumption that coarse clusters are independent, consider estimating $V$ by:
IM show that under the null, $\hat{V} \overset{d}{\to} V^W$, where $V^W = \frac{1}{r-1} \sum_{k = 1}^r ({W}_k - \bar{W})^2$ and
The IM test constructs a reference distribution $\hat{V}^W$ by drawing $W$ from $$N(0, \text{diag}(\hat{\Omega}^\text{CCE}_1, \hat{\Omega}^\text{CCE}_2, ..., \hat{\Omega}^\text{CCE}_r))$$ and seeing if $\hat{V}$ is larger than the $\left( 1-{\alpha}\right)^\text{th}$ quantile of $\hat{V}^W$.
There are two limitations to the IM test that our test does not share. Firstly, they require the regression to be estimated cluster-by-cluster. This would be infeasible in, for example, differences-in-differences set ups where treatment varies at the cluster level. Secondly, since their asymptotics take $q_k \to \infty$, we expect the test to have poor properties when $q_k$ is small. Instead, our test is expected to have good properties even when $q_k$ is small as long as $n_j$ is large. These benefits come at a cost. We expect our test to perform worse if observations within sub-clusters are highly correlated, whereas the IM test allows unrestricted covariance within sub-clusters. Our test is also conservative under the null hypothesis. We note also that neither test has power against negative correlations. This is because both tests use test statistics that take on large value relative to their reference distributions only when there is positive correlation.
mnw2020 (MNW hereafter) considers an asymptotic framework that takes $r \to \infty$. In the same spirit as IM, the MNW test is a Hausman-type test based on the variance of regression coefficients. Consider the full sample regression coefficient $\hat{\beta}$. Under the null hypothesis, the (full-sample) cluster-robust covariance estimator at the sub-cluster level, denoted, $\hat{\Omega}^\text{CCE}_J$, is consistent for the asymptotic variance-covariance matrix.
Under either the null or the alternative, but maintaining the assumption that coarse clusters are independent, the (full-sample) cluster-robust covariance estimator at the cluster level, denoted, $\hat{\Omega}^\text{CCE}_K$, is consistent for the asymptotic variance-covariance matrix. Under the null hypothesis, the authors show that their test statistic converges to a standard normal distribution: $ \frac{\hat{\Omega}^\text{CCE}_K - \hat{\Omega}^\text{CCE}_J}{\hat{V}^{MNW}} \overset{d}{\to} N(0,1) $ for an appropriately defined $\hat{V}^{MNW}$.
It is well known that the cluster-robust covariance estimator can be severely biased when $r$ is small. In order to deal with such situations, the authors propose to conduct the test using wild (sub-)cluster bootstrap. They prove the consistency of this approach in their large-$r$ framework, showing power even against alternatives with negative correlations.
Compared to the MNW test, our test is theoretically justified when both $r$ and $q$ are small, provided that $n_j$'s are large. Our test could therefore be preferable in such applications since it is presently not known if the MNW test remains valid once we take $r$ and $q$ to be fixed. However, as with the IM test, the MNW test allows unrestricted covariance within sub-clusters, whereas our test is expected to have poor performance if observations within sub-clusters are highly correlated. Our test is also conservative relative to the MNW test. On the other hand, simulation evidence suggests that it has comparable performance with the MNW test when clusters have differing variances (see Section (ref)).
In this section, we examine the finite sample performance of our worst-case randomization test (WCR) together with the IM and bootstrap version of the MNW tests via Monte Carlo simulations. We also study the performance of the na\"ive randomization test (NR) as described in Section (ref). We consider two data generating processes described below.
Model 1: Model 1 is defined by the following:
In particular, we set $X_{t,j,k} = \beta = 1$ and $\phi = 0.25$. Errors are correlated within a sub-cluster, according to an $AR(1)$ process, with autocorrelation coefficient $\phi$. $\rho$ captures the importance of cluster level shock. Since $\frac{1}{\sqrt{1 - \phi^2}} U_{t,j,k}$ has unit variance, $\rho$ is exactly the relative variance of cluster- to sub-cluster-level shocks. $\sigma_{j,k}$ controls the variance of the unobserved term in each cluster $k$. Here in Section (ref), we set $\sigma_{j,k} = 1$ for all $j \in \mathcal{J}, k \in \mathcal{K}$. In Section (ref), we explore the consequences of cluster heterogeneity by varying $\sigma_{j,k}$.
Model 2: This is the model used in the simulations of mnw2020, with the constant omitted. Let $m_k = \sum_{j \in \mathcal{J}_k} n_j$ be the total number of observations in cluster $k$. Let $U_k$ be the $m_k \times 1$ vector of $U_{t,j,k}$ for all observations in cluster $k$. Then
where $\xi_k$ is a $10 \times 1$ vector distributed as:
and $W_\xi$ is the $m_k \times 10$ loading matrix with the $(i,j)^\text{th}$ entry $\mathbf{1}\left\{ j = \lfloor (i-1)10/m_k \rfloor + 1 \right\}$. Under this model, $\frac{1}{10}$ of the observations in each cluster are correlated because they depend directly on the same $\xi_{k, l}$. In addition, there is correlation between the $\xi_{k,l}$'s since it is generated according to an AR(1) process. Observations are then ordered so that every sub-cluster contains the same number of observations that depend on each $\xi_{k,l}$. Finally, $\beta = (1,1)'$ and the two covariates are independent and generated in the same way as $U$. This model features more complex correlations between and within the sub-clusters. Clusters are independent and identically distributed. As in Section 5.2 of mnw2020, we set $\phi = 0.5$. $\rho$ here is directly comparable to $w_\xi$ in their simulations.
For our simulations, we perform the test at the 5% level. 1,000 Monte Carlo simulations were drawn for each combination of the parameters. The non-standard reference distribution in IM is evaluated using 1,000 Monte Carlo draws. Wild bootstrap in MNW is evaluated using $399$ draws as in their simulations.
To understand the size and power of each of our tests in scenarios with few clusters and few sub-clusters, we consider equal-sized clusters and sub-clusters, with $r \in \{4, 8, 12\}$, $q_k \in \{4, 8, 12\}$ and $n_j \in \{25, 50, 100\}$. We consider $\rho \in \{0, 0.5\}$.
Table (ref) presents results under the null hypothesis ($\rho = 0$). Across the two models, we see that regardless of $r$, the IM test performs poorly when $q_k$ is small. With $q_k = 4$, type I error is between 15% and 20%. By $q_k = 12$, however, the size is between 6-7%. Comparatively, our test, which is highly conservative, has type I error less than 2% across all values of $q_k$. The MNW and NR tests perform well across the board. Table (ref) presents results under the alternative $\rho = 0.5$. Relative to the IM and MNW tests, our test has power that is consistently lower. In particular, our test does poorly when $q_k$ is small. This is the weakness of the worst-case approach.
Figure (ref) presents power of the tests for $r = 8$, $q_k = 8$, $n_j = 100$ as we vary $\rho$ from $0$ to $2$ in model 1 and $0$ to $1$ in model 2. Across the two models, we see that the IM and MNW tests have greater power than our test. However, as $\rho$ increases, our test quickly catches up in power.
The previous section suggests that our test has poor performance compared to all other tests, including NR. However, a different picture emerges once we allow clusters and sub-clusters to be heterogeneous in their variances.
We first consider what happens when clusters are heterogeneous. Specifically, we return to model 1 but with $\sigma_{j,1} \in \{5,10,15\}$. That is, when all sub-clusters in cluster 1 are much noisier than the rest. Figure (ref) plots power curves with $r = 8, q_k = 12, n_j = 100$ for $\sigma_{j,1} \in \{5,10,15\}$. These curves are directly comparable with Figure (ref). Starting from the within test comparison, we see that the performance of our test is unaffected by $\sigma_{j,1}$. However, power of IM and MNW quickly degrade as $\sigma_{j,1}$ increases. Turning to the across test comparison, we see that the tests perform similarly when $\sigma_{j,1} = 5$. As $\sigma_{j,1}$ increases to 10, our test starts to have more power than the IM and MNW tests for $\rho \geq 1$. The across test comparison also shows how the NR test fails to control size. In particular, when $\sigma_{j,1}$, an NR test with nominal size 5% could wrongly reject over 40% of the time.
We see the same patterns when sub-clusters are heterogeneous. Consider again model 1 but with $\sigma_{1,k} \in \{5,10,15\}$. That is, when the first sub-cluster in each cluster is much noisier than the rest. Figure (ref) presents the results. Again, our test is not affected by changing $\sigma_{1,k}$. The power of the IM test falls by a large extent as $\sigma_{1,k}$ increases. The MNW test is also negatively affected by $\sigma_{1,k}$, though less so than the IM test.
All in all, the simulation evidence suggests that our test manages to maintain type I error below $\alpha$ when $q$ is small, whereas the IM and NR tests may see size distortion in such a setting. The cost of size control in a fixed $q$ setting is that the procedure is very conservative. This conservativeness limits the power of our test. However, the performance of our test is less sensitive to heterogeneous variances within and across clusters, such that it could be more powerful than the IM and MNW tests when some clusters or sub-clusters are much noisier than others. Hence, our test is suited for applications with small $q$ and heterogenous clusters. Indeed, as we will see in the next section, our test detects dependence in the clusters of gllqsx2019, demonstrating its potential relevance for empirical work.
In recent years, the poor performance of American students in assessment tests such as the Programme for International Student Assessment (PISA) has raised concerns among policymakers. gllqsx2019 argues that the testing gap reflects, among other things, the low effort that American students put in on tests, especially when compared to their higher scoring counterparts in other countries.
The authors test their hypothesis by a randomized controlled experiment in which students were rewarded with cash for correct answers in a 25-question test. Those assigned to the treatment group were offered roughly \$1 USD per correct answer, while the control group received no payment. Students were informed right before the test started to prevent them from changing their effort in test preparation. The experiments were conducted at 4 schools in Shanghai and 2 schools in the US. Due to logistical reasons, the authors randomized treatment at the class level for some schools and individual level for others.
Various regression analyses were conducted to study the effect of treatment on test-taking effort and test performance. Panel A in Table 3 examines whether monetary incentive increased the probability that students attempt a given question -- a proxy for effort. It does so by estimating the following equation:
Here, the unit of analysis is a question and $Y_{qi}$ is an indicator for whether student $i$ attempted question $q$. $Z_i$ is the treatment indicator and $W_i$ is a vector of control variables, which include terms such as gender, ethnicity as well as question number fixed effects. We focus on Column 1 in Panel A, which looks at US students' responses to all 25 questions in the test, and Column 4, which looks at Shanghai students' responses to the same test.
The authors present their linear regression estimate of $\beta$, together with standard errors clustered at the level of randomization. However, other levels of clustering are plausible:
We will refer to these levels of clustering by their initials hereafter. More information on the sizes of clusters can be found in appendix (ref).
While the authors chose to cluster their standard errors by $G$, it seems reasonable to be concerned about correlation across individuals within the same school or among those who took the test in the same year. If these clusters were not independent, $t$-tests using the presented standard errors could lead to the wrong conclusions.
Table (ref) presents the OLS estimates from gllqsx2019 as well as the $p$-values that would be obtained from testing the null hypothesis that $\beta = 0$ using several methods. Specifically, we consider the wild cluster bootstrap (cgm2008), approximate randomization tests (crs2017) and the t-distribution based procedure of ibragimov2010t, denoted IM2010. We perform these tests using the various plausible levels of clustering. For Column 1, we consider the increasingly coarse levels of clustering $G$, $ST$ and $S$. For the US, there are no schools sampled over multiple years, so $SY$ is the same as $S$. For Column 4, we consider the increasingly coarse levels of clustering $G$, $SY$ and $S$. In Shanghai schools, students are not separated by track, so $ST$ is the same as $S$.
Turning to the results, for column 1, we see that CCE SE's decrease as we move to increasingly coarse levels of clustering. Correspondingly, $p$-values from CCE-based $t$-tests decrease as we coarsen the clusters. Such a pattern is typically interpreted as arising from the downward bias of CCEs with few clusters (mhe2008), so that these $p$-values would be considered unreliable. Faced with downward bias, practitioners commonly turn to the wild cluster bootstrap. With this method, the $p$-values increase as we coarsen the clusters. While clustering at $G$ and $ST$ may lead one to conclude that there is strong evidence that $\beta \neq 0$, the $p$-value at $S$ suggests the absence of strong evidence. The same phenomenon arises with approximate randomization tests: at $ST$ there appears to be strong evidence that $\beta \neq 0$. At $S$, this is no longer true. With IM2010, the test does not reject in either case. We note that ART and IM2010 cannot be applied with $G$ as the chosen level of clustering, since both methods require $\beta$ to be estimated cluster-by-cluster. The results for column 4 are qualitatively similar. At $G$, CCE-based $t$-test and the wild cluster bootstrap find strong evidence that $\beta \neq 0$. This conclusion is overturned once we cluster at either $SY$ or $S$.
To assess the validity of the above specifications, we apply our WCR test, the IM test and the MNW tests. Table (ref) presents the resulting $p$-values. The notation $G \to S$ means that the null hypothesis involves sub-clusters $G$ and coarse clusters $S$. For Column 1, clustering at $G$ appears to be appropriate, as all 3 tests fail to reject the null hypotheses $G \to ST$ and $G \to S$. For Column 4, all 3 tests find strong evidence that sub-clusters $G$ are inappropriate. The WCR test has higher $p$-values than the IM and MNW tests, likely due to its lower power. Nonetheless, they are close to 5%. The WCR and MNW tests do not reject the null hypothesis for $SY \to S$, whereas the IM test does. Given the that there are at most 2 school$\times$year per school, the IM test is likely to over-reject. As such, we consider the conclusion of the WCR and MNW test to be more reliable in this instance. Thus, results based on clustering at $SY$ are plausible.
All in all, we see that settings with varying numbers of clusters and sub-clusters arise in empirical work. Our test, designed for applications with few clusters and sub-clusters is relevant and appears to work well in practical settings.
We propose to test for the level of clustering in a regression by means of a modified randomization test. We show that the test controls size even when the number of clusters and sub-clusters are small, provided that the size of sub-clusters are relatively large. This is a challenging situation not accommodated by existing tests. To ensure size control, our procedure may be conservative when clusters are homogeneous. However, in settings with heterogeneous clusters, it has power that is comparable with other tests. As such, our test can be useful when the researcher faces an application with few sub-clusters, particularly when these clusters are likely to be heterogeneous. Finally, we note that the test is easy to implement and could serve as a helpful robustness check to researchers working with clustered data. An R package is available from the author's website.