Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
33,444 characters · 7 sections · 55 citation commands
Asymptotic Theory for Two-Way Clustering
\pagenumbering{arabic}
Clustering standard errors on multiple dimensions is common and attractive in applied econometrics because it allows observations to be dependent whenever they share a cluster on any dimension. Though more broadly applicable, a common instance of two-way clustering is in linear regressions, where a researcher wants to do inference on the coefficient of interest when the residual is two-way clustered. The variance estimator proposed by cameron2011robust (henceforth CGM) has thus been widely applied to contexts with such two-way dependence.\footnote{CGM has 3886 citations on Google Scholar at the time of writing.} For instance, nunn2011slave clustered on ethnic group and district when studying the effect of slave trade on trust; michalopoulos2013pre clustered on country and ethnolinguistic family when studying the effect of pre-colonial institutions on development; jackson2018test clustered on teacher and student when studying the effect of the teacher on students' skill; neumark2019harder clustered on resume and job ad when studying the effect of age on getting a call-back. The existing justification for the asymptotic validity of the CGM estimator and other inference procedures in two-way clustering (e.g., mackinnon2021wild; davezies2021empirical; menzel2021bootstrap) relies on separate exchangeability, which implies homogeneity of clusters, a restriction that is not required in one-way clustering. This paper provides sufficient general conditions for valid inference in two-way clustering by proving that, even with cluster heterogeneity, a central limit theorem holds, and the CGM variance estimator is consistent.
An environment with two-way clustering permits dependence whenever observations share at least one cluster. To fix ideas, consider jackson2018test: observations of the same student or of the same teacher are plausibly correlated, but two observations of different students and different teachers are assumed to be independent.\footnote{This setting permits more general dependence structures than one-way clustering. If there is one-way clustering by student, then two observations from different students are automatically independent. In two-way clustering, two observations from different students are not necessarily independent because they may share the same teacher.} The CGM variance estimator accommodates such dependence, and a subsequent literature provided a theoretical basis for its validity: mackinnon2021wild obtained sufficient conditions for validity of the CGM estimator in regression models; davezies2021empirical obtained analogous results for empirical processes. menzel2021bootstrap also showed the validity of a bootstrap procedure for two-way clustering that is robust to asymptotic non-normalities.\footnote{menzel2021bootstrap pointed out that a purely interactive data-generating processes unique to two-way dependence has an asymptotic distribution that is not normal. Section (ref) will consider this process and show how the assumptions of this paper rule it out.} The theoretical basis for inference thus far relies on separate exchangeability, the assumption that random variables are exchangeable on either clustering dimension, though not necessarily both.
As noted by mackinnon2021wild, however, separate exchangeability implies identical marginal distributions. Separate exchangeability in the student-teacher example thus implies the random variables for all students must be drawn from the same distribution, including students of different cohorts over time. As wooldridge2010econometric notes in the discussion of pooled data in his graduate textbook, distributions of variables tend to change over time, so the identical distribution assumption is not usually valid. In other examples, separate exchangeability implies that countries michalopoulos2013pre and jobs neumark2019harder are identically distributed. Applied researchers surely would want size to be controlled in such heterogeneous environments, but the existing theories that rely on separate exchangeability do not imply this result. Further, in linear regressions with regressor $X_i$ and residual $u_i$, asymptotic theory is applied to $X_i u_i$. Separate exchangeability of the product implies that the regressors must also be separately exchangeable, which is not plausible when the regressors include a time trend, say.
In contrast, existing asymptotic theory on one-way clustering (e.g., hansen2019asymptotic; djogbenou2019asymptotic) allows the distribution of the random variable to be heterogeneous over clusters. Since the only available conditions for the validity of two-way clustering require separate exchangeability, the literature lacks conditions for two-way clustering that generalize one-way clustering and permit heterogeneity over clusters. This paper fills the gap, and thus justifies two-way clustering as a more robust version of one-way clustering. In particular, when the conditions of the central limit theorem hold, the variance estimator for two-way clustering asymptotically converges to the true variance for one-way clustering when the researcher clusters on more dimensions than required.
The main result is a central limit theorem for two-way clustering with heterogeneous cluster sizes and distributions. This result is proven using Stein's method. It adapts the strategy from ross2011fundamentals Theorem 3.6: I first derive an upper bound on the distance between the distribution of a pivotal statistic and the standard normal, then show that this distance converges to zero asymptotically. This proof strategy hence yields intermediate results on non-asymptotic Berry-Esseen type bounds that provide worst-case bounds on the quality of approximation between the pivotal statistic and the standard normal, which may be of independent interest. I apply the theorem to a simple setting of a linear regression, but it is more broadly applicable to many other econometric procedures that exhibit a similar clustering structure.
This paper contributes to the literature on multi-way clustering and Stein's method. This paper differs from the existing literature on multi-way clustering (e.g., mackinnon2021wild; davezies2021empirical; menzel2021bootstrap; chiang2022standard; chiang2023using) in that it does not rely on separate exchangeability. Stein's method has been applied to other contexts such as two-way fixed effects verdier2020estimation, and spillover effects (e.g., chin2018central, leung2022rate and braun2023estimation). Unlike the aforementioned papers, this paper speaks directly to multi-way clustering, and it makes a modification to the proof of ross2011fundamentals Theorem 3.6 to obtain the result instead of applying the theorem directly.
Consider a setup with two-way clustering on dimensions $G$ and $H$ for random vectors $\{ W_i \}_{i=1}^n$, where $W_i := (W_{i1}, W_{i2}, \cdots, W_{iK})' \in \mathbb{R}^K$ and $i=1,\dots,n$ is the unit of observation. For example, $G$ could denote states and $H$ could denote industries. Clustering in more than two dimensions is possible, and derivations are entirely analogous. This section establishes a central limit theorem (CLT) for $\sum_i W_i$, as $n \rightarrow \infty$. Here and in the following, sums are over (subsets of) $\{ 1, 2, \dots, n \}$. For $C \in \{G ,H \}$, let $\mathcal{N}^C_c$ denote the set of observations in cluster $c$ on dimension $C$ --- this setup partitions the sample on the $C$ dimension.
Let $g(i)$ and $h(i)$ denote the cluster that observation $i$ belongs to on the $G$ and $H$ dimensions respectively. These cluster identities are nonstochastic and observed. Let $N^C_c := |\mathcal{N}^C_c|$ denote the cluster size for $C \in \{G, H \}$ and $N_{gh} := |\mathcal{N}^G_g \cap \mathcal{N}^H_h|$. These cluster sizes are allowed to be heterogeneous in a way that will be formalized in the assumptions below. $W_i$ is assumed to be independent of the joint distribution of $\{ W_j \}$ for $j \notin \mathcal{N}^G_{g(i)} \cup \mathcal{N}^H_{h(i)} =: \mathcal{N}_i$, i.e., when $i$ and $j$ do not share a cluster on either dimension. Hence, $\mathcal{N}_i$ is the set of observations that are arbitrarily dependent with $i$. This environment is stated as Assumption (ref).
While the dependence structure is implicitly described in the setup of many clustering papers (e.g., hansen2019asymptotic; menzel2021bootstrap), Assumption (ref) makes the dependence structure explicit. Assumption (ref)(a) is a dissociation assumption similar to definition 3.5 of ross2011fundamentals required to apply Stein's method. Assumption (ref)(b) is required because, for a scalar $W_i$, a crucial step of the proof requires $E[W_i W_j W_k W_l] = E[W_i W_k] E[W_j W_l]$ when $j,l$ do not share any cluster with $i,k$. Even when $W_i \!\perp\!\!\!\perp (W_j, W_l)$ and $W_k \!\perp\!\!\!\perp (W_j, W_l)$, we cannot conclude that $E[W_i W_j W_k W_l] = E[W_i W_k] E[W_j W_l]$ in general, because independence of marginal distributions does not imply independence of the joint distribution. Assumption (ref)(b) hence makes an assumption on the joint distribution. It can alternatively be stated as $(W_i, W_k) \!\perp\!\!\!\perp (W_j, W_l)$, which is stronger but more interpretable than the zero-covariance assumption. I further discuss the relationship between Assumption (ref) and the existing literature in Section (ref).
Assumption (ref) is agnostic about the dependence structure when $W_i$ and $W_j$ share at least one cluster. It also allows the data generating process to be arbitrarily heterogeneous across different clusters, mimicking the heterogeneity permitted in one-way clustering (e.g., hansen2019asymptotic). Since one-way clustering is a special case of two-way clustering where the $H$ cluster consists of single observations, the result here generalizes the existing results in one-way clustering. In contrast, the existing literature on two-way clustering assumes separate exchangeability that additionally imposes identical distribution over clusters, so it does not generalize the results on one-way clustering.
For positive definite matrix $Q$, let $\lambda_{\operatornamewithlimits{min}} (Q)$ denote the smallest eigenvalue of $Q$. Then, let $Q_n := Var \left( \sum_i W_i \right)$ denote the variance of the sum and $\lambda_n := \lambda_{\operatornamewithlimits{min}}(Q_n)$ denote its smallest eigenvalue. For example, when $K=1$, $W_i$ is a scalar and $\lambda_n = Q_n = Var(\sum_i W_i)$. $K_0$ is used throughout the paper to denote an arbitrary constant.
Assumption (ref)(a) requires the fourth moment to be bounded, which is stronger than the moment condition in one-way clustering.\footnote{See Equation (7) of hansen2019asymptotic for the condition in one-way clustering.} The proof in one-way clustering usually verifies a Lindeberg condition because blocks of observations are independent of each other. With two-way dependence, we no longer have independent blocks because each cluster can have observations that are dependent on observations from a different cluster when these observations share a cluster on a different dimension. Hence, a different proof strategy is required. The proof in this paper uses Stein's method, which requires stronger moment restrictions, but provides a non-asymptotic bound on the approximation error --- details are in Subsection (ref). By using this strategy, a bounded fourth moment is required.
Assumption (ref)(b) requires the contribution of the largest cluster to be small relative to the total variance. This condition mimics the sparsity condition in the networks literature (e.g., graham2020sparse). Intuitively, this condition is required so that the removal of a cluster does not change the variance substantively. This assumption allows the ratio of any two cluster sizes to diverge to infinity. It is identical to Equation (12) of hansen2019asymptotic for one-way clustering. Assumption (ref)(b) also rules out having components that are perfectly correlated: if the components of the vector were perfectly correlated (i.e., $\mu^\prime W_i =0$ for some $\mu \ne (0,\dots,0)^\prime$), then $\lambda_n =0$. If cluster sizes are uniformly bounded, and $\lambda_n \rightarrow \infty$, then Assumption (ref)(b) is satisfied.\footnote{Assumption (ref)(b) is hence a more general version of sparsity than having the size of the dependency neighborhood (i.e., the number of observations plausibly correlated with some observation $i$) being bounded above. The conditions are also comparable with verdier2020estimation in the two-way fixed effects literature: when the neighborhood size is bounded, $\lambda_n \asymp n$, which matches his assumption 2(c).}
Assumption (ref)(c) is a summability condition that requires $\lambda_n$ not to be too small, and requires $\lambda_n$ to be the same order as $\sum_c (N^C_c)^2$, i.e., $\lambda_n \asymp \sum_c (N^C_c)^2$, $C \in \{ G,H \}$.\footnote{For sequences $a_n$ and $b_n$, $a_n \asymp b_n$ if and only if there exists $K_0 <\infty$ such that $a_n/b_n, b_n/a_n \in [-K_0, K_0]$ for all elements in the sequence.} With strictly positive covariance within clusters, $\lambda_n \asymp \sum_c (N^C_c)^2$ is satisfied. However, if the researcher were conservative and clustered on $C$ when the data is indeed iid, then $\lambda_n \asymp n$, which then requires $\sum_c (N^C_c)^2 \asymp n$ for the condition to hold --- this condition holds when the cluster sizes are not too large. The assumption that $(1/\lambda_n) \sum_c (N_c^C)^2 \leq K_0$ matches Equation (11) of hansen2019asymptotic.
The theorem tells us that, under the aforementioned conditions, $Q_n^{-1/2} \sum_i (W_i - E[W_i])$ is asymptotically standard normal and the plug-in variance estimator following CGM is consistent for two-way clustering. One-way clustering is a special case of this theorem when one dimension is weakly nested within the other: examples include $G=H$ so both dimensions are identical, and clustering by county and state (as counties are nested in states). A sufficient condition for consistent variance estimation is $E[W_i] = 0$, similar to Theorem 3 of hansen2019asymptotic. This assumption is sufficient in many applications: for example, linear regressions considered in Section (ref) are identified by requiring the expectation of the residual term to be zero.
The following two subsections discuss technicalities on the dependence structure and the proof sketch. A general-interest audience may wish to proceed immediately to Section (ref).
To compare the setup used in Assumption (ref) to the existing literature, I carefully define a few terms used in menzel2021bootstrap, whose setup uses a dissociated separately exchangeable array. Let $Y_{gh}$ denote an infinite array of observations in cluster $g$ on the $G$ dimension and cluster $h$ on the $H$ dimension. $Y_{gh}$ is a separately exchangeable array if, for any integers $\tilde{G}, \tilde{H}$ and permutations $\pi_1: \{ 1,\dots,\tilde{G} \} \rightarrow \{ 1,\dots,\tilde{G} \}$ and $\pi_2: \{ 1,\dots,\tilde{H} \} \rightarrow \{ 1,\dots,\tilde{H} \}$, we have:
where $\stackrel{d}{=}$ denotes equality in distribution. Such an array is dissociated if, for any $G_0, H_0 \geq 1$, $(Y_{gh})_{g=1,h=1}^{g=G_0,h=H_0}$ is independent of $(Y_{gh})_{g>G_0,h>H_0}$. Dissociation is how the existing literature formally incorporates the multi-way clustering structure. Separate exchangeability implies that the cluster indices are not meaningful, and it is stronger than having identical distributions across clusters. This environment is a special case of Assumption (ref), as the following proposition claims.
Assumption (ref) on independence and zero covariance is also arguably more transparent and interpretable for an economics audience than formal generalizations of separate exchangeability, e.g., relative exchangeability in crane2018relatively that requires defining a signature, arity, structure, and embedding.
The proof of Theorem (ref) proceeds by first proving a CLT for a scalar random variable, then applying the Cramer-Wold device to obtain the multivariate CLT. The scalar CLT is proven using Stein's method. I adapt the proof strategy from ross2011fundamentals Theorem 3.6 to obtain an upper bound on the Wasserstein distance between a pivotal statistic and the standard normal random variable. By exploiting the two-way clustering structure, the upper bound on the distance can be shown to converge to zero. All details are in Appendix (ref).
For ease of exposition, consider a simpler environment where $K=1$, and $E[W_i]=0$. Let $\sigma_n^2 := Q_n$, $R = \sum_i W_i/ \sigma_n$, and $Z \sim N(0,1)$. Lemma (ref) in Appendix (ref) provides an explicit bound on the Wasserstein distance between $R$ and $Z$. With $d_W(.)$ denoting the Wasserstein distance, and $d_K(.)$ denoting the Kolmogorov distance, Proposition 1.2 from ross2011fundamentals implies that $d_K(R,Z) \leq (2/\pi)^{1/4} \sqrt{d_W(R,Z)}$. The Kolmogorov distance is the maximal distance between two CDF's, so it is informative of the maximum distance between the distribution of the pivotal statistic and the standard normal. Then, by using Assumption (ref) to adapt the proof of Theorem 3.6 in ross2011fundamentals,
Observe that this intermediate result is informative of the quality of the normal approximation. This bound on the Wasserstein distance (and hence the Kolmogorov distance) is non-asymptotic, and of the Berry-Esseen type, thereby giving a worst-case bound on the distance between the pivotal statistic and the standard normal.
At this point, my proof departs from the proofs in the existing statistical literature that employ Stein's method (e.g., chen2004normal). Let $N_i := |\mathcal{N}_i|$. H\"{o}lder's inequality is employed on objects such as $\sum_i \sum_{j,k \in \mathcal{N}_i} E[|W_i| W_j W_k]$. The existing literature uses the $L^1$ norm of moments $E[|W_i|^3]$ and the $L^\infty$ norm of $N_i$, resulting in $\left( \operatornamewithlimits{max}_m N_m \right)^2 \sum_i E[|W_i|^3]$. In contrast, my proof uses the $L^\infty$ norm of $E[|W_i|^3]$ and the $L^1$ norm of $N_i$, resulting in $\operatornamewithlimits{max}_m E[|W_m|^3] \sum_i N_i^2$. Hence,
Since $\operatornamewithlimits{max}_m E[|W_m|^3]$ is bounded by Assumption (ref)(a), it suffices to show that $\sum_i N_i^2 / \sigma_n^3 \rightarrow 0$. Due to Assumption (ref)(a), $N_i \leq N^G_{g(i)} + N^H_{h(i)}$, so
Since $\lambda_n = \sigma_n^2$ when $K=1$, $\operatornamewithlimits{max}_{g,h} (N^G_g + N^H_h) /\sigma_n \rightarrow 0$ by Assumption (ref)(b) and the final term $\left( \sum_g (N^G_g)^2 + \sum_h (N^H_h)^2 \right) /\sigma_n^2$ is bounded by Assumption (ref)(c). Hence, the term is $o(1)$.
A similar argument is made for the fourth moment that features in $Var\left( \sum_i \sum_{j \in \mathcal{N}_i} W_i W_j \right)$. To complete the proof for variance estimation, observe that since the fourth moments exist, the consistency of the plug-in variance estimator can be proven by using Chebyshev's inequality and the existing intermediate results.
This section applies Theorem (ref) to linear regressions, showing that using the normal approximation with the CGM variance estimator is valid. Consider a linear model where the scalar outcome $Y_i$ is generated by
with $X_i \in \mathbb{R}^K$. We are interested in estimating $\beta$. Suppose $E[X_i u_i] =0$ for all $i$, and $(X_i^\prime, u_i)$ is allowed to be two-way clustered. The standard OLS estimator is
This object is assumed to be well-defined in that $\sum_i X_i X_i'$ is invertible. Define $S_n := \sum_i E[X_i X_i']$ and $Q_n := Var \left( \sum_i X_i u_i \right)$, and denote their sample analogs as $\hat{S}_n = \sum_i X_i X_i'$ and $\hat{Q}_n := \sum_i \sum_{j \in \mathcal{N}_i} \hat{u}_i \hat{u}_j X_i X_j'$. Let the smallest eigenvalue of $Q_n$ be $\lambda_n := \lambda_{\operatornamewithlimits{min}}(Q_n)$. The asymptotic variance of $\hat{\beta}$ and its sample analog are $V(\hat{\beta}) := S_n^{-1} Q_n S_n^{-1}$ and $ \hat{V}(\hat{\beta}) := \hat{S}_n^{-1} \hat{Q}_n \hat{S}_n^{-1}$ respectively.
Assumption (ref) provides sufficient conditions for the estimator $\hat{\beta}$ to be asymptotically normal and for the CGM variance estimator to be consistent. The conditions mimic Assumption (ref) so that Theorem (ref) is applicable to the random vector $X_i u_i$. The new condition is a weak regularity condition that $\lambda_{\operatornamewithlimits{min}} \left( S_n /n \right) \geq K_1 > 0$, mimicking the rank condition in OLS.
Proposition (ref) is useful for performing F tests on a subvector of $\beta$. The proof of Proposition (ref) proceeds by applying Theorem (ref) to $\sum_i X_i u_i$, then showing that $S_n^{-1} \hat{S}_n \xrightarrow{p} I_K$, which uses the rank condition of Assumption (ref)(e). It then remains to show that the remainder terms are asymptotically negligible.
The practitioner's takeaway from Proposition (ref) is that the existing CGM variance estimator can be used for valid inference with two-way clustering. The result provides the formal theoretical guarantee for using the estimator, under conditions that permit heterogeneity across clusters.
Besides the application mentioned, Theorem (ref) also has implications on the conditions required for valid inference when the random variable is two-way clustered in many other econometric models, including design-based settings and instrumental variables models. This theory is especially relevant for design-based settings where the researcher conditions on potential outcomes, so the random variable cannot be separately exchangeable by construction --- see yap2023design, for instance. Inference for estimators based on moment conditions can be done by straightforward application of Theorem (ref) as in linear regression. Practically, this paper has shown that the popular CGM estimator is robust in an environment without separate exchangeability, but practitioners should exercise caution when applying bootstrap methods to environments that are not separately exchangeable. While the results are presented for two-way clustering, they can be easily extended to clustering on three or more dimensions.