EconBase
← Back to paper

Panel Data with Unknown Clusters

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

56,539 characters · 18 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

1 Panel Data with Unknown Clusters

abstract\setstretch{1} Clustered standard errors and approximate randomization tests are popular inference methods that allow for dependence within observations. However, they require researchers to know the cluster structure ex ante. We propose a procedure to help researchers discover clusters in panel data. Our method is based on thresholding an estimated long-run variance-covariance matrix and requires the panel to be large in the time dimension, but imposes no lower bound on the number of units. We show that our procedure recovers the true clusters with high probability with no assumptions on the cluster structure. The estimated clusters are independently of interest, but they can also be used in the approximate randomization tests or with conventional cluster-robust covariance estimators. The resulting procedures control size and have good power.

Introduction

Consider the following regression with panel data:

equation*[equation* omitted — 82 chars of source]

where a researcher wants to conduct inference on $\beta$. If the researcher is concerned about correlations between $U_{it}$ and $U_{i't'}$, it is frequently helpful to group units into independent clusters. These independent clusters can then be used to construct cluster-robust covariance estimators (CCE) as in lz1986, or for approximate randomization tests as in crs2017 and ccks2021.

However, cluster assignments are rarely known ex ante. In many contexts, multiple levels of clustering are plausible. For example, with the American Community Survey, researchers have the choice of clustering at the individual, county or state level and the appropriate level of clustering is not always obvious. In other situations, researchers may not be able to identify their desired level of clustering. For example, in the presence of peer effects, researchers might like to cluster their observations along friend groups since this is the level at which spillover occurs. However, unless researchers observe friendship networks, it would not be possible to cluster at this chosen level.

Clustering at the correct level is important for inference. It is well known that ignoring cluster dependence -– in other words, clustering at too fine a level -- leads to tests with excessive type I errors. bdm2004 and cgm2008, for example, find nominal size find that neglecting dependence can lead to type I that exceeds their nominal size by as much as 10 times. Conversely, excessively coarse levels of clustering bring their own problems. Firstly, coarse clusters tend to be few in numbers. A large body of work show that confidence intervals based on the cluster-robust standard errors tend to under-cover when the number of clusters is small (see mhe2008 and cgm2008 among others), leading to poor size control. Secondly, setting aside under-coverage issues, excessively coarse levels of clustering cause tests to have poor power (aaiw17), since the researcher assumes less information than they actually have. Despite the importance of these issues, there has been limited theoretical guidance on choosing the appropriate level of clustering.

We propose a procedure to help researchers discover clusters in the panel data. Our method is based on thresholding an estimated long-run variance-covariance matrix and requires the panel to be large in the time dimension, but imposes no lower bound on the number of units. We show that the procedure recovers the true clusters with high probability with few assumptions on the cluster structure. We believe the estimated clusters can be of independent interest to researchers. However, they can also directly be used in the approximate randomization tests of crs2017 or in tests based on cluster-robust covariance estimators. We show that doing so leads to tests which control size in asymptotic frameworks that takes the number of clusters to be fixed or growing to infinity.

Our paper is similar in spirit to tests for level of clustering, which aims to provide a robustness checks for researchers who have already chosen a given level of clustering. However, these methods are designed to test between two specified nested levels of clustering and are not suitable for discovering the cluster structure. Our paper also relates to the extensive literature on panel data with interactive fixed effects, which effectively assumes that cluster dependence takes factor structure. We are unable to accommodate the type of endogeneity allowed in this literature, but we allow for richer patterns of dependence while maintaining ease of computation.

Our paper is most similar to bcl2020, which provides a method for inference in panel data when clusters are unknown, but when correlation across units are sparse. Both our methods are based on thresholding a long-run variance-covariance matrix. The method of bcl2020 even when in the absence of cluster structures, since they rely on sparsity in the dependence structure between units. In the cluster setting, their sparsity assumption translates into many-cluster asymptotics, which may not be realistic in all applications. Unlike their method, ours is able to accommodate small-cluster asymptotics. Furthermore, we provide a novel cluster-recovery result. Simulations also suggest that our method leads to tests which are more powerful. Finally, aaiw17 advocates a design-based approach on the issue of clustering, arguing that dependence between units is neither necessary nor sufficient for clustering standard errors. We argue that our method is useful even for researchers who takes such a perspective. We expand on the aforementioned points in section (ref).

The rest of this paper is organized as follows. Section (ref) presents our model and assumptions. The proposed method as well as our theoretical results are contained in section (ref). Section (ref) relates our paper to existing literature. Section (ref) presents results from Monte Carlo simulation. Section (ref) concludes. Proofs are contained in the appendices.

Model and Assumptions

In this section we discuss the assumptions which are needed for cluster recovery and inference. The most onerous assumption of the method is that it requires panel data that is large in the time dimension. Beyond that, we require the covariates and error terms to have tails that are sufficiently thin, and also for dependency across time to decay quickly enough. Our assumptions are generally standard in the panel data literature.

We work with the usual linear model:

assumption[Model] Consider the model \begin{equation*} Y_{i,t} = X_{i,t}'\beta + U_{i,t} \quad , \quad E[X_{i,t}U_{i,t}] = 0 \end{equation*} Suppose also that for all $N, T \in \mathbb{N}$, there exists $\underline{\lambda} \in \mathbb{R}$ so that \begin{equation*} \lambda_min \left( \frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T E\left[X_{i,t}X_{i,t}'\right]\right) \geq \lambda \end{equation*} where $\lambda_\text{min}(M)$ is the smallest singular value of the matrix $M$.

The lower bound on the singular values of the expected Gram matrix is a common alternative for the full rank assumption when working with independent but not identical clusters.

Our method learns cluster structure from what we will call the long-run correlation matrix. For this matrix to be estimable, we assume strong mixing and stationarity at the unit level.

definitionDefine the $\alpha$-mixing coefficient (for stationary random variables) as \begin{equation*} \alpha(h) = \sup_{A \in \mathcal{F}^0_{-\infty} , B \in \mathcal{F}_h^\infty } \lvert P(A)P(B) - P(A \cup B) \rvert \end{equation*} where $\left(\mathcal{F}^0_{-\infty} , \mathcal{F}_h^\infty \right)$ are the $\sigma$-algebras generated by $\{\mathbb{X}_t, \mathbb{U}_t\}^0_{t=-\infty}$ and $\{\mathbb{X}_t, \mathbb{U}_t\}^\infty_{t=h}$ respectively.

Our strong mixing assumption is that:

assumption[Strong Mixing] Suppose $\{\mathbb{X}_t, \mathbb{U}_t\}$ are stationary, and there exists $C_1 > 0$ and $\kappa > 0$ such that $\alpha(h) \leq \exp(-C_1h^\kappa)$.

In order to achieve a sufficiently fast rate of convergence, we also need to control higher moments of the scores. In particular, we assume that they decrease fast enough to satisfy Bernstein's condition:

assumption[Bernstein's Condition] Suppose there exists $C_2 > 0$ and $M$ such that for all $h \in \mathbb{Z}_+$ and $i, j \in \mathbb{N}$, we have that for all $k \geq 1$, \begin{equation*} E\left[\lvert W_{i,t}W_{j,t-h} \rvert^k \right] \leq C_2^{k-2}k!E[W_{i,t}^2W_{j,t-h}^2] , \end{equation*} where $W_{it} \in \left\{X^{(1)}_{i,t}U_{i,t}, ..., X^{(p)}_{i,t}U_{i,t} \right\}$ and $W_{j,t-h} \in \left\{X^{(1)}_{j,t-h}U_{j,t-h}, ..., X^{(p)}_{j,t-h}U_{j,t-h} \right\}$.

Assumptions (ref) and (ref) -- or their analogues -- are frequently seen in the panel data literature. See for instance bcl2020 or bm2015.

Our next assumption restricts heterogeneity across individuals:

assumption[Uniformity Conditions] Suppose there exists $M_2$ and $M_k$ for some $k \geq 4$ such that for all $h \in \mathbb{Z}_+$ and $i, j \in \mathbb{N}$ \begin{align*} E[W_{i,t}^2W_{j,t-h}^2] \leq M_2 \quad and \quad E[W_{i,t}^kW_{j,t-h}^k] \leq M_k \end{align*} for all $i$, $j$, $h$ , $W_{it} \in \left\{X^{(1)}_{i,t}U_{i,t}, ..., X^{(p)}_{i,t}U_{i,t} \right\}$ and $W_{j,t-h} \in \left\{X^{(1)}_{j,t-h}U_{j,t-h}, ..., X^{(p)}_{j,t-h}U_{j,t-h} \right\}$.

In particular, we require that the variance and covariance across individuals to be bounded. As such, while individuals can be different, we cannot have a few individuals dominating the regression estimates.

The next assumption concerns the existence of a cluster structure.

definitionDefine: \begin{align*} \sigma_{i,j}^{a,b} = E[X_{i,t}^{(a)}U_{i,t}X_{j,t}^{(b)}U_{j,t}] + \sum_{h=1}^{\infty} E[X_{i,t}^{(a)}{U}_{i,t}X_{j,t-h}^{(b)}{U}_{j,t-h}] + E[X_{i,t-h}^{(a)}{U}_{i,t-h}X_{j,t}^{(b)}{U}_{j,t}] . \end{align*} where $X_{i,t}^{(a)}$ refers to the $a^{\text{th}}$ component of the vector $X_{i,t}$. Let the long-run covariance between two individuals be denoted \begin{align*} \sigma_{i,j} = \sum_{a = 1}^p \sum_{b = 1}^p \left\lvert \sigma_{i,j}^{a,b} \right\rvert . \end{align*} Define $\Sigma$ to be the $N\times N$ matrix with $\sigma_{i,j}/\sqrt{\sigma_{i,i} \sigma_{j,j}}$ in its $(i,j)^\text{th}$ entry.
assumption[Cluster Structure] Suppose that each individual $i$ belongs to one of $q$ clusters, where $q$ is unknown. Let $g(i): [N] \to [q]$ denote the function that maps $i$ to its cluster. Further, suppose that \begin{enumerate} • $\sigma_{ij} = 0$ if $i$ and $j$ belong to different cluster • For each $i$, $|g(i)| > 1$ implies that there exists at least one $j \neq i$ s.t. $g(j) = g(i)$ and $\sigma_{ij} \neq 0$. • There exists $0 < \underline{\sigma} \leq \overline{\sigma} < \infty$ such that for all $i,j \in \mathbb{N}$, $\sigma_{i,j} \neq 0 \Rightarrow \underline{\sigma} \leq\sigma_{ij} \leq \overline{\sigma}$. \end{enumerate}

In other words, when the individuals are sorted according to their clusters, we have that:

equation*[equation* omitted — 190 chars of source]

where each $\Sigma_g$ is either a scalar, or has a non-zero entry in every row.

The above assumption does not restrict cluster structure since we make no assumption on $q$. In particular, any pattern of correlation between units is permitted if we set $q = 1$. As will become clear in section (ref), we do not need any assumption on $q$ to recover the cluster structure. However, we will introduce assumptions on $q$ for inference.

Where the assumption has bite is in restricting $\sigma_{i,j}$. The lower bound is similar to the beta-min condition seen in the LASSO literature. It rules out covariance terms that are arbitrarily close to $0$, which would be difficult to estimate.

In addition, we require the $\sigma_{i,j}$ to be bounded above. This ensures that the estimate of the long-run correlation matrix $\Sigma$ is well-behaved. However, we can also threshold a covariance matrix rather than a correlation matrix, in which case the upper bound would be unnecessary.

Lastly, we note that our goal is only to recover the cluster structure up to relabeling. To be precise, we define the following:

definitionWe say that cluster structure $g$ is equivalent to $\tilde{g}$ if there exists permutation $\pi$ such that $g([N]) = \pi \tilde{g}([N])$. We write $g \cong \tilde{g}$.

In other words, we seek $\hat{g}$ so that $\hat{g} \cong g$ with high probability.

The Proposed Method

In this section we discuss our proposed method for recovering the clusters and for performing inference. The method involves $2$ tuning parameters. We provide heuristics for choosing these parameters in section (ref).

Cluster Recovery

We first define our estimator for the long-run correlation matrix. The estimator is based on the heteroskedasticity and autocorrelation consistent (HAC) estimator of nw1994 and involves a bandwidth (tuning) parameter.

definition[Bartlett Kernel] We define the Bartlett Kernel to be $\omega: \mathbb{Z}_+ \times \mathbb{Z}_+ \to \mathbb{R}$, \begin{align*} \omega(h, L) = \begin{cases} \frac{L - h}{L} & if h \leq L \\ 0 & otherwise. \end{cases} \end{align*}
definitionGiven the bandwidth parameter $L \in \mathbb{Z}_+$, define: \begin{align} \begin{aligned} \hat{\sigma}^{a,b}_{i,j} & = \frac{1}{T}\sum_{t=1}^{T} X_{i,t}^{(a)}\hat{U}_{i,t}X_{j,t}^{(b)}\hat{U}_{j,t} \\ & + \sum_{h = 1}^L \frac{\omega(h, L)}{T-h} \sum_{t=h+1}^{T} X_{i,t}^{(a)}\hat{U}_{i,t}X_{j,t-h}^{(b)}\hat{U}_{j,t-h} \\ & + \sum_{h = 1}^L \frac{\omega(h, L)}{T-h} X_{i,t-h}^{(a)}\hat{U}_{i,t-h} X_{j,t}^{(b)}\hat{U}_{j,t-h} \end{aligned} \end{align} and let \begin{equation*} \hat{\sigma}_{i,j} = \sum_{a = 1}^{p} \sum_{b = 1}^p \left\lvert \hat{\sigma}_{i,j}^{a,b} \right\rvert \end{equation*} Let the estimator of $\Sigma$, denoted $\hat{\Sigma}$, be the $N\times N$ matrix with $\hat{\sigma}_{i,j}/\sqrt{ \hat{\sigma}_{i,i} \hat{\sigma}_{j,j} }$ in its $(i,j)^\text{th}$ entry.

We propose to estimate the cluster structure using algorithm (ref). In words, the algorithm estimates $\Sigma$ entry-by-entry using a HAC-based estimator. Given $\hat{\Sigma}$, we set correlations that are smaller than $T^{\eta - 1/2}$ to $0$, because these links are likely spurious. Using the remaining links, we can then group individuals who are correlated into the same cluster.

algorithm[algorithm omitted — 539 chars of source]
remarkWe state our result for the Bartlett kernel. Our obtained rates of convergence rely on the fact that partial sums of the Bartlett kernel is bounded by $\log_2(T)$. Other kernels can be used but they might lead to different rates of convergence.

Inference

Using the estimated clusters, there are two possible methods for performing inference, depending on the assumption that is made on $q$. When $q$ is small, the approximate randomization tests of crs2017 controls size. It's main drawback is that $\hat{\beta}$ needs to be estimable cluster-by-cluster. On the other hand, if $q$ is large, inference using the conventional clustered covariance estimator errors will have good properties, though this method is not compatible with small $q$.

Approximate Randomization Tests

Approximate Randomization Tests (ARTs) are appropriate when there are few clusters. In this subsection we explain how to test linear hypotheses using ARTs. For ease of exposition, we discuss the case with a single linear restriction, though the test also accommodates tests of multiple linear restrictions. The hypothesis of interest is:

equation[equation omitted — 120 chars of source]

The test is based on cluster-by-cluster estimates of $\beta$, which we denote $\{\hat{\beta}_j\}_{j=1}^{\hat{q}}$. Define:

equation*[equation* omitted — 117 chars of source]

Our test statistic is:

equation*[equation* omitted — 287 chars of source]

This is the familiar $t$-statistic that takes $S_T$ as the “raw data". Intuitively, $R(S_T)$ is large if the $r\hat{\beta}_j$'s are far from $\lambda$ and small otherwise.

Next, denote by $\mathbf{G}$ the group of $\hat{q} \times 1$ sign changes. $\mathbf{G}$ can be identified with the set of $g \in \{-1, 1\}^{\hat{q}}$ so that:

equation*[equation* omitted — 166 chars of source]

As we elaborate in section (ref), the basis for our randomization test is that asymptotically, $gS_T$ has the same distribution as $S_T$. Hence, the set of $R(gS_T)$ provides a valid reference distribution for the $R(S_T)$. The test therefore rejects the null hypothesis when $R(S_T)$ takes on extreme values relative to $R(gS_T)$. It proceeds as follows.

Define $M = |\mathbf{G}| = 2^{\hat{q}}$ and let:

equation*[equation* omitted — 72 chars of source]

be the ordered values of $R(gS_T)$ as $g$ varies in $\mathbf{G}$. For a fixed nominal level $\alpha$, let $k$ be defined as

equation*[equation* omitted — 49 chars of source]

where $\lceil x \rceil$ denotes the smallest integer greater or equal to $x$. In addition, define:

equation[equation omitted — 192 chars of source]

and set

equation[equation omitted — 82 chars of source]

We can then define the randomization test as:

align[align omitted — 185 chars of source]

In words, this test rejects the null hypothesis with certainty when $R(S_T) > R^{(k)}(S_T)$. When $R(S_T) = R^{(k)}(S_T)$, it rejects the null hypothesis with probability $a(S_T)$. The test does not reject when $R(S_T) < R^{(k)}(S_T)$.

remarkAs with standard randomization tests, $|\mathbf{G}|$ may sometimes be too large so that computation of $\{R(gS_T)\}_{g \in \mathbf{G}}$ becomes onerous. In these instances, it is possible to replace $\{R(gS_T)\}_{g \in \mathbf{G}}$ with a stochastic approximation. Formally, for a given $B \in \mathbb{Z}_+$, let \begin{equation*} \hat{\mathbf{G}} = \left\{ g^1, ... ,g^B \right\} \end{equation*} where $g^1$ is the identity transformation and $g^2, ..., g^B$ are independent draws from Uniform($\mathbf{G}$). Using $\hat{\mathbf{G}}$ instead of ${\mathbf{G}}$ in equation ((ref)) does not affect validity of our results in section (ref).
remarkOur test is possibly randomized. A deterministic but conservative version of the test can be implemented by rejecting the null hypothesis if and only if $R(S_T) > R^{(k)}(S_T)$. In the spirit of restricting researcher degree-of-freedom, this is also the version of the test that we implement in the Monte Carlo simulations of section (ref). With appropriately chosen tuning parameters, the test has good power.

In sum, we propose to conduct inference by treating the estimated clusters as the true clusters and applying the usual ART. We summarize the implementation procedure in the algorithm below:

algorithm[algorithm omitted — 1,436 chars of source]

Inference with Clustered Covariance Estimators

When there are many clusters, we can simply perform tests based on the usual clustered covariance estimators, using the estimated clusters in place of the unknown true clusters. We briefly review the method below.

Using the estimated clusters, the clustered covariance estimator is defined as:

equation[equation omitted — 192 chars of source]

where $\mathbf{X}_j$ is the $n_j \times p$ matrix formed by stacking row-wise the covariates of all individuals whose estimated cluster is $j$. That is, the $l^{\text{th}}$ row of $\mathbf{X}_j$ is $X_i'$ for some $i$ for whom $g(i) = j$. $\hat{\mathbf{\varepsilon}}_j$ is analogously defined. In other words, $\hat{V}^{CCE}$ is the usual clustered covariance estimator but constructed taking the estimated clusters as true clusters.

To test the hypothesis in equation (ref), we form the $t$-statistic:

equation*[equation* omitted — 106 chars of source]

The test is then:

align[align omitted — 164 chars of source]

Theoretical Results

Our main result shows that clusters are exactly recovered with high probability:

theorem[Cluster Recovery] Given assumptions (ref), (ref), (ref), (ref) and (ref), suppose that $L \to \infty$, $L = o(\sqrt{T})$ and $N = O(T^c)$. Then as $T \to \infty$, \begin{equation*} P(\hat{g} \cong g) \to 1 . \end{equation*}

Note that the theorem does not place any lower bound on the $N$. In particular, it is allowed to be fixed. The method also does not place any restrictions on $q$. Our cluster recovery method is therefore consistent with both large and small cluster asymptotics.

Given the “oracle property" described in the above theorem, it is immediate that approximate randomization tests and CCE-based tests that use the estimated clusters are valid. This is because with high probability, using the estimated clusters is equivalent to using the true clusters. This leads to the following corollaries:

corollary[ART-based Test] Given assumptions (ref), (ref), (ref), (ref) and (ref), suppose that $L \to \infty$, $L = o(\sqrt{T})$, $N = O(T^c)$ and that $q$ is fixed as $T \to \infty$. Suppose further that for each $j \in [q]$, there exists $\underline{\lambda} > 0$ so that \begin{equation*} \lambda_min \left( \frac{1}{n_j T}\sum_{i \in I_j}^N\sum_{t=1}^T E\left[X_{i,t}X_{i,t}'\right]\right) \geq \lambda \quad for all T \in \mathbb{N} . \end{equation*} where $n_j = \lvert\{i : g(i) = j\}\rvert$. Then, \begin{equation*} \underset{T \to \infty}{\lim \sup} \,\, E\left[\phi^{ART}_T\right] \leq \alpha . \end{equation*}
corollary[CCE-based Test] Given assumptions 1, 2, 3, 4 and 5, suppose that $L \to \infty$, $L = o(\sqrt{T})$, $N = O(T^c)$ and that $q \to \infty$ is fixed as $T \to \infty$. Suppose further that for some $2 \leq r < \infty$, \begin{equation*} T \cdot \frac{\left(\sum_{j = 1}^q |n_j|^r \right)^{2/r}}{N} \leq C < \infty \quad , \quad \max_{j \in [q]} T \cdot \frac{|n_j|^2}{N} \to 0 \end{equation*} Then, \begin{equation*} \hat{V}_{CCE}^{-1/2}\left(\hat{\beta} - \beta \right) \overset{d}{\to} N(0, I) \end{equation*} where $n_j = \lvert\{i : g(i) = j\}\rvert$ and $\hat{V}_{CCE}$, as displayed in equation (ref), is the usual CCE clustered using $\hat{g}$. Furthermore, \begin{equation*} \underset{T \to \infty}{\lim \sup} \,\, E\left[\phi_T^{CCE}\right] = \alpha \end{equation*}

The corollaries above follow from imposing the additional assumptions which are necessary for ART and CCE to yield valid tests. In corollary (ref), we assume that $q$ is fixed. The additional condition ensures that $\hat{\beta}$ can be estimated cluster-by-cluster. Size control then follows immediately from crs2017.

Similarly, in corollary (ref), we take $q \to \infty$. The remaining assumptions, taken from hl2019, control the amount of heterogeneity across clusters and ensures that CCE with the true clusters yield valid inference.

Choice of Tuning Parameters

Our cluster recovery method involves two tuning parameters. We propose to choose them by cross-validation. First, note that $T^{\eta - 1/2} \in [0, 1]$. Equivalently, we search the unit interval for the optimal value of $\tilde{\eta} = T^{\eta - 1/2}$. For $L$, the relevant range over which to search is $\left[0, \sqrt{T}\right]$. Our proposed cross-validation procedure, adapted from that of bcl2020 and bl2008, is as follows.

First, divide the $T$ time periods into $P = [\log T]$ continuous blocks each of size $[T/\log T] \pm 1$. Number the blocks $1$ to $P$. Fix $L$. For $p \in [P]$, using only the observations in block $p$, compute $\hat{\Sigma}$ as in definition (ref) using windows of size up to $L$.

Denote this estimate $\hat{\Sigma}_p(L)$. We write $\hat{s}_{p, ij}(L)$ to denote the $(i,j)^{\text{th}}$ entry of $\hat{\Sigma}_p(L)$. For a given $\tilde{\eta}$, let $\tilde{\Sigma}_p(L, \tilde{\eta})$ be $\hat{\Sigma}_p(L)$ thresholded at $\tilde{\eta}$. In other words,

align*[align* omitted — 203 chars of source]

The cross-validation objective function is:

equation*[equation* omitted — 161 chars of source]

The cross-validation value of $L$ and $\tilde{\eta}$ are then:

equation*[equation* omitted — 171 chars of source]

Intuitively, our cross-validation procedure relies on stationarity. If the correlation between units are stable over time, the properly thresholded estimator for some block $p$ should be close to the unthresholded estimators obtained from all other blocks.

Although the theoretical results in section (ref) do not allow for data-dependent choices of $\eta$, the simulations in section (ref) suggest that cross-validation works well in practice.

remarkIn the objective function above, we cannot use the thresholded estimator for both ${p}$ and $p'$. Such an objective function would be minimized by $\tilde{\eta} = 1$, since eliminating all entries other than the diagonal will lead to a cross-validation error of $0$.

Relation to Existing Literature

In this subsection, we discuss the work most closely related to ours bcl2020 (section (ref)), as well as the literatures on tests for level of clustering (section (ref)) and panel data with interactive fixed effects (section (ref)). Finally, we discuss how our work can be useful even when researchers take the “design-based" approach of aaiw17 (section (ref)).

bcl2020

Our paper is most closely related to bcl2020 (henceforth BCL), which to our knowledge is the only other paper that is explicitly concerned with inference in settings with unknown clusters. The starting point of their method is the observation that dependence across units affect the variance $\hat{\beta}$ only through:

align*[align* omitted — 403 chars of source]

where

equation*[equation* omitted — 276 chars of source]

Consider the Newey-West estimator of $V_{i,j}$:

equation*[equation* omitted — 272 chars of source]

Since they are consistent for $V_{i,j}$, it seems reasonable to construct the estimator:

equation*[equation* omitted — 66 chars of source]

However, it turns out that when $N$ is large, such an estimator “accumulates a large number of cross-sectional estimation noises[sic]" (BCL, pg 3). Conventional clustered covariance estimator overcomes this problem by setting $V_{i,j}$ for which $i,j$ are in different clusters to $0$, so that a large number of terms do not need to be estimated.

In the absence of this information, BCL assumes that the set of $V_{i,j}$'s are sparse -- that is, that they are mostly $0$ -- and propose to use a thresholding method to identify the entries which are $0$. In effect, they employ the estimator:

equation*[equation* omitted — 132 chars of source]

which can then be used to construct an estimator for the variance-covariance matrix of $\hat{\beta} - \beta$. The authors show that their estimator is consistent and leads to tests which are valid.

Our method is similar to BCL since we also threshold an long-run correlation matrix. In fact, $\hat{\sigma}_{i,j} = \iota' \hat{V}_{i,j}\iota$. The advantage of BCL over our method is that they do not require the existence of clusters for inference. In particular, they could accommodate a dependency structure in which for any two units, one can find a chain of units along which pairwise correlation is not $0$. Our method would have poor power in such a situation since there is only one cluster.

However, in a cluster context, BCL's sparsity assumption implies $q \to \infty$. Using ART with the recovered clusters, we are able to accommodate a fixed $q$ setting. This means, for example, that we allow a given unit to be correlated with $O(N)$ other units, a setting for which BCL is unsuitable. In addition, we state a cluster recovery result which is novel. This is not found in BCL, although this is because they do not assume the existence of a cluster-structure.

Furthermore, we relax two potentially restrictive assumptions in BCL. First, unlike BCL, we do not require $N \to \infty$. Note that both methods require $T \to \infty$ and $N = O(T^c)$ for some $c$.\footnote{In fact, BCL allows $N$ to grow at an exponential rate in relation to $T$. Our method does as well, though for ease of exposition, we have opted to state our result in terms of arbitrarily large polynomial rate growth.} Secondly, we do not impose any lower bound on the $\alpha$-mixing coefficients across time. Such a lower bound excludes the setting in which observations are independent over time. This is unnatural since we expect that the absence of dependence would make the inference problem easier.

Tests for Level of Clustering

Our concern with the fact that clusters are typically unknown ex ante is shared by the literature on tests for level of clustering (im2016, mnw2020 and cai--testforclustering). These tests assume that researchers are able to identify clusters which are independent and use this information to test whether a conjectured finer level of clustering is valid.

Researchers could plausibly use a sequence of these tests to discover the level of clustering by testing the validity of increasingly fine sub-clusters. However, such a method is unsuitable for use in discovering clusters. First, these tests would have to be adjusted for multiple testing, especially if the goal is inference. No sophisticated method of adjustment is available and Bonferroni corrections will lead to tests which are too conservative. Secondly, these methods require the conjectured clustering to be nested within the valid clustering that is known. This is restrictive.

Lastly and most importantly, researchers may not know of any independent clusters. With these methods, assuming that there is one cluster containing all the observations is not an option. This is because im2016 and mnw2020 require at least two clusters to be computable, while cai--testforclustering has trivial power when there is only one cluster.

Our method and BCL therefore complement this literature by providing methods that choose the level of clustering directly, without requiring input from the researcher beyond specifying the linear model.

Panel Regression with Interactive Fixed Effects

Our paper is also related to a large literature on panel regression models with interactive fixed effects (see for instance bai2009, mw2015, bm2015 among many more.) Relative to our model, these papers further assume that the error term has a factor structure:

align*[align* omitted — 77 chars of source]

In the above equation, $E_{i,t}$ is white noise, although $f_{t,r}$ maybe correlated with $X_{i,t}$. If $f_{t,r}$ are exogenous to $X_{i,t}$, the above model would be nested in ours, which accommodates richer patterns of correlation than the factor structure.

Allowing $f_{t,r}$ to be endogenous necessitates estimation procedures that are either computationally difficult, or require further assumptions on the $X_{i,t}$'s. Our procedure is unable to accommodate endogenous $f_{t,r}$, but it is computationally simple without these additional requirements.

Design-Based Approach

aaiw17 (henceforth AAIW) advocate a “design-based" perspective on the issue of clustering. They are concerned with two types of design issues. Clustering is a sampling design issue when samples are drawn according to a two-stage process, in which the first step entails sampling clusters and individuals are drawn only in the second stage. Because clusters could differ systematically and not all clusters are observed, researchers wanting to learn the average effect over all clusters have to take into account the uncertainty introduced by this clustered sampling procedure. Conversely, clustering is an experimental design issue when clusters of units, rather than individual ones, are assigned to a treatment by the assignment mechanism. Within each cluster, we observe only one of two potential outcomes. Because clusters could differ systematically in their potential outcomes, researchers who want to learn an average effect must again cluster their standard errors. AAIW then call on researcher to be explicit in their beliefs about the sampling and treatment assignment process, and use these beliefs to guide their clustering decisions.

For a concrete example, suppose we are interested in studying the returns to education in the US. In other words, we are interested in the average effect of an additional year of education on income for the average individual in the US. Suppose we only observe individuals in 25 randomly chosen states, but that education is randomly assigned to all individual within this sample. By AAIW, the researcher should cluster at the state level due to clustering in the sampling process. Suppose instead that we observe individuals in all states, but that all individuals in a given state were randomly assigned the same years of education. Then, because of clustering in experimental design, standard errors should again be clustered at the state level. Finally, suppose that we observe all individuals, and that years of education was randomly assigned to each individual. Then there is no clustering arising from sample design or experimental design. There is no need to cluster standard errors, even though individuals in the same state may have scores $X_{i,t}U_{i,t}$ that are correlated (provided of course that these unobservables are not correlated with years of education). As such, researchers should cluster their standard errors at the state only if they believe that their sample is drawn from a subset of states, or if they believe that education was randomly assigned at the state level.

At first sight, the design-based approach contradicts the conventional model-based approach to clustering: the latter demands that researchers cluster their standard errors when there is correlation in the scores of the clusters. This contradiction can be resolved once we realise that the design-based approach is concerned with a different estimand: the conditional treatment effect. Appendix (ref) provides a simple illustration of this point.

The approach of AAIW is theoretically insightful, but not necessarily helpful for a researcher deciding how to cluster. This is because they require researchers to know the sampling or treatment assignment process, which researchers may not know. Our procedure can still be useful in such a scenario. For example, suppose the treatment variable, $D_{i,t}$, is $\alpha$-mixing over time. Then, applying our cluster recovery method with $D_{i,t}$ replacing $X_{i,t}\hat{U}_{i,t}$ will yield clusters across which treatment assignment is asymptotically independent. More generally, applying our method in their framework will still yield clusters that are valid -- that is, clusters with observables and unobservables that are asymptotically independent. The trade-off is that these clusters might be too coarse, leading to relatively lower power.

Hence, our method is compatible with the design-based approach to clustering. Despite the potential power loss, researchers might find our data-driven method useful, especially when they have little information or are unwilling to make assumptions on the true sampling and treatment assignment mechanisms.

Monte Carlo Simulations

In this section we study the finite sample performance of our method via Monte Carlo simulations. Section (ref), evaluates the method's ability to reliably recover clusters from the data. Section (ref) considers the size and power of tests based on the recovered clusters.

We employ the following data generating process:

align*[align* omitted — 257 chars of source]

where $g(i)$ denotes the cluster to which each unit $i$ belongs. Specifically, we set $\rho = \phi = 0.2$ and $\beta = 1$. We study the performance of the test as the following parameters vary: $q \in \{5, 10, 25\}$, $n \in \{50, 200\}$, $t \in \{100, 200\}$. We select the tuning parameters $\eta$ and $L$ by the cross-validation procedure of section (ref). Our results show that the method is effective in recovering the true clusters and leads to tests which are both valid and powerful.

Cluster Recovery

Table (ref) presents results on cluster recovery. We assess the quality of the estimated clusters on purity: the largest number of individuals in each estimated cluster that truly belong to the same cluster. We consider both the minimum purity across the estimated clusters as well as average purity. Our second criterion is the number of estimated clusters $\hat{q}$. The method achieves perfect recovery if and only if minimum and average cluster purity are both equal to $1$ and $\hat{q} = q$.

We see that the method reliably recovers clusters in the data. Both minimum and average purity of the estimated clusters are high. For $q = 5$, $\hat{q}$ is close to $q$, suggesting that the clusters are very high quality. For $q = 10$ and $q = 25$, we still have high cluster purity. However, the $\hat{q}$ can be very large when $n =100$. This means that the estimated clusters are splintered version of the true clusters. By the time $n = 200$, however, the problem mostly goes away. $q = 25$ is the most challenging for cluster recovery. This is unsurprising since with $25$ clusters, the long-run covariance matrix is mostly noise, making cluster recovery a challenging task. Nonetheless, it performs reasonably well.

table[table omitted — 1,618 chars of source]

Test Performance

Table (ref) presents type I error of our test as well as that of bcl2020 for $\alpha = 10\%$. We see that the ART version of our test has good size control when $q = 5$. When $q = 10$, the test initially over-rejects when $t = 100$. However, type I error is close to $10\%$ once $t = 200$. For $q = 5, 10$, the test performs very similarly to the oracle version of ART. This is unsurprising since the clusters are very well estimated at these values. When $q = 25$, ART is over-rejects at close to 20%. The fact that clusters are relatively poorly estimated affects size control here, though we see that as $T$ grows, the problem quickly goes away. In contrast, the CCE version of our method does well when $q = 25$. It has poor size control when $q = 5$, due to the well known fact that CCE's are downward biased when $q$ is small. However, as $q$ grows, the CCE method controls size well. In particular, it handles the situation with $q = 25$ better than the ART method. This suggests that the CCE method is less sensitive to clusters that are wrongly estimated. In general, our simulation results concord with the theoretical results in section (ref). Turning to BCL, we see that across all parameter values, the method is fairly conservative, with size below 1%.

table[table omitted — 4,613 chars of source]

Table (ref) presents the power of our test as well as that of bcl2020 for $\alpha = 10\%$ when $\beta_0 = 0.95$. Across the parameter values considered, both the ART and CCE version of our procedure has least 10% more power than BCL, often much more. The difference is less stark once we turn to table (ref). Here, the null hypothesis is so clearly wrong that both methods reject it at very high rates.

table[table omitted — 2,283 chars of source]

In summary, our simulation result suggests that our cluster recovery method finds situations with large $q$ more challenging, but generally performs well. Furthermore, inference method that combined our recovered clusters with either ART or CCE are able to control size well, and has much more power than BCL. Our method is hence useful for inference in panel data with unknown clusters.

Conclusion

We provide a method for inference in panel data with unknown clusters. We propose a procedure to help researchers discover clusters in panel data. Our method is based on thresholding an estimated long-run variance-covariance matrix and requires the panel to be large in the time dimension, but imposes no lower bound on the number of units. We provide a novel theoretical result showing exact cluster recovery with high probability. Furthermore, the recovered clusters can be combined with either approximate randomization tests or tests based on clustered covariance estimators to yield valid inference. The test based on approximate randomization test controls size even when the number of clusters is small, a setting that is not currently handled by existing papers. Simulation results show that our method has more power than existing methods, making it a useful addition to the toolbox of applied economists.