EconBase
← Back to paper

Design-Based Multi-Way Clustering

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

31,827 characters · 8 sections · 25 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Design-Based Multi-Way Clustering

\pagenumbering{arabic}

abstractThis paper extends the design-based framework to settings with multi-way cluster dependence, and shows how multi-way clustering can be justified when clustered assignment and clustered sampling occurs on different dimensions, or when either sampling or assignment is multi-way clustered. Unlike one-way clustering, the plug-in variance estimator in multi-way clustering is no longer conservative, so valid inference either requires an assumption on the correlation of treatment effects or a more conservative variance estimator. Simulations suggest that the plug-in variance estimator is usually robust, and the conservative variance estimator is often too conservative.

Introduction

There have been thousands of papers that adjust standard errors for clustering in linear regression. To address the issue of when such clustering is appropriate, abadie2023should extended the design-based framework of abadie2020sampling to show how clustering on a particular dimension, such as state, is justified when there is clustered sampling or assignment on that dimension. However, their framework and theory apply only to one-way clustering, which leaves an open question on how papers that clustered on multiple dimensions, such those applying the plug-in variance of cameron2011robust, henceforth CGM, can be justified. There are situations where sampling and assignment occur on different clustering dimensions (e.g., gruber1995health). There are also sampling (e.g., hersch1998compensating) or assignment (e.g., nunn2011slave) mechanisms that are multi-way clustered. Then, there is an open question of whether the results of abadie2023should generalize to multi-way clustering settings. This paper fills the gap.

Until recently, accounting for the large sample behavior of design-based settings with multi-way clustering has been a difficult problem. Asymptotic theory for variables that have multi-dimensional dependence has thus far relied on separate exchangeability (e.g., davezies2018asymptotic). Separate exchangeability implies that the marginal distributions of clusters are exchangeable. mackinnon2021wild However, by construction, separate exchangeability is violated in a design-based framework because the residual $u_i = W_i u_i(1) + (1-W_i) u_i(0)$ depends on nonstochastic potential residual $u_i(w)$ and treatment $W_i$. Hence, even if the treatments $W_i$'s were identically distributed over clusters, the marginal distribution of $u_i$ cannot be the same as the $W_i$'s are weighted differently. Hence, limit theory that accommodates heterogeneity of clusters and observations is required. By building on the central limit theorem in yap2023general, this paper obtains results on large-sample behavior of standard estimators in this environment.

This paper shows that while most results in the one-way design-based setup generalize to multi-way clustering, there are some nuances in multi-way clustering. As we would expect, when sampling and assignment occur on different clustering dimensions, it is necessary to cluster on both dimensions to obtain valid inference. When sampling is two-way clustered, it is necessary to cluster on both dimensions for valid inference. When assignment is one-way clustered, abadie2023should showed that the standard liang1986longitudinal plug-in variance estimator is conservative, even though it is still necessary to cluster in some way for valid inference. Similarly, in several data-generating processes with multi-way clustered assignment, the CGM plug-in estimator is conservative. However, unlike one-way clustering, the CGM estimator is no longer always conservative. In fact, it is possible to construct a data-generating process that makes the CGM variance estimator anticonservative when there is multiway sampling.

In response to the anticonservativeness of the CGM estimator in design-based settings, there are two approaches that empirical researchers may take. The first approach is to make an assumption on how the individual treatment effects are correlated within the same cluster: CGM is conservative when the correlation is positive, which is reasonable in most applications. The second approach is to remain agnostic and to use CGM2, a more conservative version of the CGM variance estimator proposed by davezies2018asymptotic. Since the simulations show that CGM2 is often unnecessarily conservative, making an assumption on the correlation is usually the more reasonable approach.

Setting

We have a sequence of populations indexed by $k$, each with $n_k$ units. Each unit is indexed by $i = 1, \cdots , n_k$, partitioned into clusters on two dimensions $G,H$, indexed by $g$ and $h$. Let $m=(g,h)$ denote the intersection of two cluster indices that are nonempty. Let $g_{ki}, h_{ki}$ denote the cluster that unit $i$ belongs to on the respective dimensions, and $m_{ki}$ its intersection of clusters. For treatment variable $W \in \{ 0,1 \}$, we have the potential outcome is denoted $y_{ki} (w)$ that is nonstochastic.

We are interested in the population average treatment effect (ATE):

align*[align* omitted — 81 chars of source]

With $\alpha_k := (1/n_k)\sum_{i=1}^{n_k} y_{ki}(0)$ and residuals $U_{ki} := Y_{ki} - \alpha_k - \tau_k W_{ki}$, the potential residuals are denoted:

align*[align* omitted — 103 chars of source]

Let $R_{ki}$ denote the sampling indicator, so observation $i$ in population $k$ is observed when $R_{ki}=1$. Similarly, $W_{ki}$ denotes the treatment indicator. Hence, for $R_{ki}=1$, we observe the triple $\{ Y_{ki}, W_{ki}, m_{ki} \}$, with $Y_{ki}= W_{ki} y_{ki}(1) + (1-W_{ki})y_{ki}(0)$. The high-level assumption on assignment and sampling is that $(R_{ki},W_{ki}) \perp \!\!\! \perp (R_{kj},W_{kj})$ whenever they do not share any cluster, and that sampling and assignments are independent, as stated in Assumption (ref).

assumption$(R_{ki},W_{ki}) \perp \!\!\! \perp (R_{kj},W_{kj})$ if $g_{ki} \ne g_{kj}$ and $h_{ki} \ne h_{kj}$. Further, $\{ R_{ki} \}_{i=1}^{n_k} \perp \!\!\! \perp \{ W_{ki} \}_{i=1}^{n_k}$. For all $i$, there exists some $b_{k1}$ and $b_{k0}$ such that $E\left[R_{ki}W_{ki}\right] = b _{k1}$ and $E\left[R_{ki}\left(1-W_{ki}\right)\right] =b_{k0}$.

$b_{k1}$ is the probability that an individual is observed and treated; $b_{k0}$ is the expected probability that an individual is observed and untreated.

The above assumption can be generated from several design-based settings. One possibility is to adapt the data-generating process from abadie2023should for one-way clustering, just that the sampling and assignment processes are allowed to occur on different clustering dimensions. This environment strictly generalizes their setting: their setting is a special case where both sampling and assignment occur on the same dimension, and everyone is in their own cluster on the other dimension. Another possibility is to have multi-way clustering occur on either the assignment or sampling dimension, and independence over units on the other dimension. I detail both possibilities in the next two subsections, and provide empirical examples.

Sampling and Assignment on Different Dimensions

Without loss of generality, suppose there is clustered sampling on the $G$ dimension. Here, every $G$ cluster is independently sampled with probability $q_k$. If a cluster is sampled, then units within the cluster are sampled independently with probability $p_k$. Clustered assignment can occur on the $H$ dimension. Every cluster on the $H$ dimension has an assignment probability $B_{kh}$, drawn from a distribution with mean $\mu_k$ and $\sigma_k^2$. Then, for a given cluster $h$, units within that cluster are independently assigned treatment with probability $B_{kh}$.

An empirical example is gruber1995health, who study the effect of health insurance on retirement. We may want to cluster by household and state-year cell. They use the CPS data which has a sample that is negligibly small relative to the superpopulation of households, and each household reveals data on employment in previous years (where they could have moved across different states). It then seems reasonable to believe that sampling occurs on the household dimension. Assignment occurs on the state dimension because insurance is affected by state policy. Since the same incumbent party has differing policies over states, and the same state has differing policies when the party differs over time, assignment to health insurance varies by state clusters.

Multiway Clustering on Assignment

This subsection explains a possible mechanism for multiway clustering on the assignment dimension. Since the setup for multiway sampling is analogous, its exposition is omitted for brevity.

Data is generated by independently drawing $A_{kg}\in[0,1]$, $B_{kh}\in[0,1]$ and $e_{ki}\sim U[0,1]$, with $W_{ki}=1\left\{ e_{i}<A_{g(i)}B_{h(i)}\right\}$. The random variables $A_{kg}$ and $B_{kh}$ have means $\mu_{Ak}, \mu_{Bk}$ and variances $\sigma_{Ak}^2, \sigma_{Bk}^2$ respectively. This process nests several cases. If assignment is one-way clustered, then we can simply set $A_{kg}=1$. A special case is where assignment occurs at the intersection level, and we need both dimensions $G$ and $H$ to be assigned treatment for the unit to be treated. Then, $A_{kg}, B_{kh} \in \{0,1 \}$. This mechanism also nests the assignment mechanism where a unit is treated whenever either of its clustering dimensions is treated. To see this, observe that the unit is untreated only if both its clusters are untreated, so we can simply switch the labels of treatment and non-treatment.

As an example of multi-way assignment, nunn2011slave studied the effect of slave trade on trust today. They clustered by ethnic groups and district. The unit of observation is an individual. Both these dimensions affect assignment because people of similar ethnicities have similar history of slave trade. District is relevant since people in the same location are more likely hit by slave trade.

An example of multi-way sampling is from hersch1998compensating, who was interested in the effect of injury risk on wages. She clustered by industry and occupation. When sampling from a large number of industries and occupations, we would observe an individual only when both his industry and his occupation were sampled.

Least Squares Estimator and Variance

Let $N_k := \sum_{i=1}^{n_k} R_{ki}$ denote the number of sampled units. Further, $N_{k1} := \sum_{i=1}^{n_k} R_{ki} W_{ki}$ and $N_{k0} := \sum_{i=1}^{n_k} R_{ki} (1-W_{ki})$. The least squares estimator is:

align*[align* omitted — 149 chars of source]

The rest of this section first shows how the estimator is asymptotically normal, and derives an expression for its asymptotic variance. Then, it presents the asymptotic limits of various commonly-used variance estimators.

Large Sample Properties of OLS Estimator

To state the main result, I first define a few terms. Let $\mathcal{N}_i$ denote the neighborhood of $i$, which is the set of observations that are plausibly correlated with $i$. In the multi-way clustering context, $\mathcal{N}_i := \{ j: g_{ki}=g_{kj} \} \cup \{ j: h_{ki}=h_{kj} \}$. For $C \in \{G ,H \}$, let $\mathcal{N}^C_c$ denote the set of observations in cluster $c$, and $N^C_c := |\mathcal{N}^C_c|$ denote the number of observations in cluster $c$ on the $C$ dimension. Further, define $\xi_{ki}$ as the demeaned residual for individual $i$ that features in the variance of $\hat{\tau}_k$:

align*[align* omitted — 157 chars of source]
assumptionLet $Q_{ki}$ denote $R_{ki} W_{ki}$, $R_{ki} (1-W_{ki})$ or $\xi_{ki}$. For some positive $K_0<\infty$, $C\in \{ G,H \}$ and $\lambda_k = Var(\sum_{i=1}^{n_k} Q_{ki})$, the following hold: \begin{enumerate} • $|y_{ki}(w)| \leq K_0$. • $\frac{1}{\lambda_k} \operatornamewithlimits{max}_c (N^C_c)^2 \rightarrow 0$. • $\frac{1}{\lambda_k} \sum_c (N^C_c)^2 \leq K_0$. \end{enumerate}

These regularity conditions are required in the multi-way clustering context so that the central limit theorem (CLT) from yap2023general can be applied. These regularity conditions require bounded potential outcomes, the contribution of the largest cluster to the total variance be small, and a summability condition.

theoremUnder (ref) and (ref), \begin{equation} \frac{\sqrt{N_k} (\hat{\tau}_k - \tau_k)}{\sqrt{v_k}} \xrightarrow{d} N(0,1) \end{equation} where \begin{align*} v_k &:= \frac{N_k}{n_k^2} \sum_{i=1}^{n_k} \sum_{j \in \mathcal{N}_i} E[\xi_{ki} \xi_{kj}] \end{align*} and \begin{align*} E\left[\xi_{ki}\xi_{kj}\right] & =\left(\frac{1}{b_{k1}^{2}}E\left[R_{ki}R_{kj}\right]E\left[W_{ki}W_{kj}\right]-1\right)u_{ki}(1)u_{kj}(1) \\ &\qquad +\left(\frac{1}{b_{k0}^{2}}E\left[R_{ki}R_{kj}\right]E\left[\left(1-W_{ki}\right)\left(1-W_{kj}\right)\right]-1\right)u_{ki}(0)u_{kj}(0)\\ & \qquad-\left(\frac{1}{b_{k1}b_{k0}}E\left[R_{ki}R_{kj}\right]E\left[W_{ki}\left(1-W_{kj}\right)\right]-1\right)u_{ki}(1)u_{kj}(0) \\ &\qquad-\left(\frac{1}{b_{k1}b_{k0}}E\left[R_{ki}R_{kj}\right]E\left[\left(1-W_{ki}\right)W_{kj}\right]-1\right)u_{ki}(0)u_{kj}(1) \end{align*}

In (ref), the variance expression is a function of expectations of $R$ and $W$ cross products across different units. These cross-products have different expressions, depending on the environment that generated the data, as stated in Lemmas (ref) and (ref) below.

lemmaSuppose (ref) holds, there is clustered sampling on dimension $G$, and clustered assignment on dimension $H$. Then, $g_{ki} \ne g_{kj}$ implies $E\left[R_{ki}R_{kj}\right] = p_{k}^{2}q_{k}^{2}$ and $h_{ki} \ne h_{kj}$ implies $E\left[W_{ki}W_{kj}\right]=\mu^{2}$. If $g_{ki}=g_{kj}$, $E\left[R_{ki}R_{kj}\right]=q_{k}p_{k}^{2}$. If $h_{ki}=h_{kj}$, $E\left[W_{ki}W_{kj}\right]= \sigma^{2}+\mu^{2}$, $E\left[W_{ki}\left(1-W_{kj}\right)\right] = \mu\left(1-\mu\right)-\sigma^{2}$, and $E\left[\left(1-W_{ki}\right)\left(1-W_{kj}\right)\right] =\left(1-\mu\right)^{2}+\sigma^{2}$.
lemmaSuppose (ref) holds and there is multiway clustering in assignment. Then, if $m_{ki}=m_{kj}$, then $E\left[W_{ki}W_{kj}\right]=\left(\mu_{A}^{2}+\sigma_{A}^{2}\right)\left(\mu_{B}^{2}+\sigma_{B}^{2}\right)$. If $g_{ki}=g_{kj},h_{ki} \ne h_{kj}$, then $E\left[W_{ki}W_{kj}\right]=\left(\mu_{A}^{2}+\sigma_{A}^{2}\right)\mu_{B}^{2}$. If $g_{ki}\ne g_{kj}$ and $h_{ki} \ne h_{kj}$, then $E\left[W_{ki}W_{kj}\right]=\mu_{A}^{2}\mu_{B}^{2}$.

Multi-way sampling can be treats $R$ analogously to $W$ in Lemma (ref).

Variance Estimators

The true variance as stated above is a function of potential residuals, so a direct plug-in is infeasible. Variance estimators hence use a feasible analog. Define:

align*[align* omitted — 288 chars of source]

Observe that $\xi_{ki} = \eta_{ki} - E[\eta_{ki}]$, and $E[\eta_{ki}] = u_{ki}(1)-u_{ki}(0)$. The common variance estimators include the heteroskedasticity-robust variance estimator attributed to eicker1967limit, huber1967under and white1980heteroskedasticity (EHW), and the one-way cluster robust variance estimator attributed to liang1986longitudinal (LZ). The CGM estimator cameron2011robust is the sum of LZ estimators on the two different dimensions minus the LZ estimator of the intersection to avoid double-counting terms. The CGM2 estimator davezies2018asymptotic uses the sum without subtracting the intersection so terms are double-counted. To be precise,

align*[align* omitted — 620 chars of source]

where, $\hat{V}_{LZG}$ and $\hat{V}_{LZH}$ denote the LZ estimator on the G and H dimensions respectively, while $\hat{V}_{LZM}$ is the LZ estimator treating each intersection $m$ as a single cluster.

assumptionFor $C, C^{\prime} \in \{ G,H \}$, let $\lambda^C_k := \sum_{i=1}^{n_k} \sum_{j \in \mathcal{N}^{C}_{c_{ki}}} E[\eta_{ki} \eta_{kj}]$. Then, $\frac{1}{\lambda^C_k} \operatornamewithlimits{max}_c (N^{C^\prime}_c)^2 \rightarrow 0$ and $\frac{1}{\lambda^C_k} \sum_c (N^{C^\prime}_c)^2 \leq K_0$.

The conditions in Assumption (ref) are required so that the asymptotic error incurred by using $\hat{V}$ relative to the true $V$ converges to zero. Since the strategy for showing such convergence is similar to yap2023general, an analogous summability condition and a condition on the largest cluster having a negligible contribution to the variance are required.

theoremUnder (ref), (ref) and (ref), for $LZC \in \{ LZG,LZH \}$, \begin{align*} \frac{\hat{V}_{EHW}}{V_k^{EHW}} \xrightarrow{p} 1 \qquad \frac{\hat{V}_{LZC}}{V_k^{LZC}} \xrightarrow{p} 1 \qquad \frac{\hat{V}_{CGM}}{V_k^{CGM}} \xrightarrow{p} 1 \qquad \frac{\hat{V}_{CGM2}}{V_k^{CGM2}} \xrightarrow{p} 1 \end{align*} where \begin{align*} V_k^{EHW} &:= \frac{1}{N_k} \sum_{i=1}^{n_k} E[\eta_{ki}^2] \\ V_k^{LZC} &:= \frac{1}{N_k} \sum_{i=1}^{n_k} \sum_{j \in \mathcal{N}^C_{c_{ki}}} E[\eta_{ki} \eta_{kj}] \\ V_k^{CGM} &:= \frac{1}{N_k} \sum_{i=1}^{n_k} \sum_{j \in \mathcal{N}_i} E[\eta_{ki} \eta_{kj}] \\ V_k^{CGM2} &:= \frac{1}{N_k} \sum_{i=1}^{n_k} \sum_{j \in \mathcal{N}_i} E[\eta_{ki} \eta_{kj}] + \frac{1}{N_k} \sum_{i=1}^{n_k} \sum_{j \in \mathcal{N}^M_{m_{ki}}} E[\eta_{ki} \eta_{kj}] \\ E[\eta_{ki} \eta_{kj}] &= E[\xi_{ki} \xi_{kj}] + E[\eta_{ki}] E[\eta_{kj}] \end{align*}

The result allows us to observe differences in variance estimators in Corollary (ref).

corollaryThe differences between the limit of the variance estimators and the true variance are as follows: \begin{align*} V_{k}^{CGM}-v_{k} & =\frac{N_{k}}{n_{k}^{2}}\sum_{i}\sum_{j\in\mathcal{N}_{i}}E\left[\eta_{ki}\right]E\left[\eta_{kj}\right] \\ V_{k}^{EHW}-v_{k} & =\frac{N_{k}}{n_{k}^{2}}\sum_{i}E\left[\eta_{ki}\right]^{2}-\frac{N_{k}}{n_{k}^{2}}\sum_{i}\sum_{j\in\mathcal{N}_{i}\backslash\left\{ i\right\} }E\left[\xi_{ki}\xi_{kj}\right] \\ V_{k}^{LZG}-v_{k} & =-\frac{N_{k}}{n_{k}^{2}}\sum_{i}\sum_{j\in\mathcal{N}_{i}\backslash\mathcal{\mathcal{N}}_{g(i)}^{G}}E\left[\xi_{ki}\xi_{kj}\right]+\frac{N_{k}}{n_{k}^{2}}\sum_{i}\sum_{j\in\mathcal{\mathcal{N}}_{g(i)}^{G}}E\left[\eta_{ki}\right]E\left[\eta_{kj}\right] \\ V_{k}^{CGM2}-v_{k} & =\frac{N_{k}}{n_{k}^{2}}\sum_{i}\sum_{j\in\mathcal{\mathcal{N}}_{m(i)}^{M}}E\left[\xi_{ki}\xi_{kj}\right]+\frac{N_{k}}{n_{k}^{2}}\sum_{i}\sum_{j\in\mathcal{\mathcal{N}}_{h(i)}^{H}}E\left[\eta_{ki}\right]E\left[\eta_{kj}\right]+\frac{N_{k}}{n_{k}^{2}}\sum_{i}\sum_{j\in\mathcal{\mathcal{N}}_{g(i)}^{G}}E\left[\eta_{ki}\right]E\left[\eta_{kj}\right] \end{align*}

In general, we cannot guarantee that $V_k^{CGM}-v_k$ is positive. The simulation presents one example where this difference is negative. Nonetheless, observe that, since $u_{ki}(1) - u_{ki}(0) = E[\eta_{ki}]$,

align*[align* omitted — 172 chars of source]

Hence, $V_{k}^{CGM}-v_{k}$ will be positive whenever $\sum_{j \in \mathcal{N}_i} (\tau_{ki} - \tau_k) (\tau_{kj} - \tau_k) \geq 0$ i.e., when the sum of the correlation in the treatment effects for observations that are plausibly correlated is positive. This assumption is implied by the treatment effects of all units who are plausibly correlated having a positive correlation (i.e., $(\tau_{ki} - \tau_k) (\tau_{kj} - \tau_k) \geq 0$ for all $j \in \mathcal{N}_i$). However, the requirement for $V_k^{CGM} - v_k$ is weaker than this uniform assumption on the correlation, because we can have some negative correlations --- we merely require their overall sum to be weakly positive. The assumption that $\sum_{j \in \mathcal{N}_i} (\tau_{ki} - \tau_k) (\tau_{kj} - \tau_k) \geq 0$ is reasonable in many empirical situations, as we would usually expect treatment effects of neighbors to be positively rather than negatively correlated. However, this assumption cannot be written in terms of $R_{ki}$ and $W_{ki}$ that our design-based setting has restrictions on. Hence, nothing in the design-based setting necessarily implies that CGM is conservative.

The asymptotic underestimation of the EHW estimator comes from failing to account for correlations in $\xi$, but it may be compensated by the nonzero mean of $\eta$. A similar interpretation is obtained in that underestimation comes from the failure to account for correlations in H that are not part of the g cluster. This underestimation may be offset by $\sum_{i}\sum_{j\in\mathcal{\mathcal{N}}_{g(i)}^{G}}E\left[\eta_{ki}\right]E\left[\eta_{kj}\right] \geq 0$, which is weakly positive because the object can be written as a sum of outer products.

Finally, it can be shown that CGM2 is conservative.

corollary$V_{k}^{CGM2}-v_{k} \geq 0$.

The CGM2 difference above is guaranteed to be positive, because it can be written as an additive function of one-way clustered objects. Its conservativeness arises from double-counting the intersection of clusters, and the nonzero mean of $\eta$.

Consequently, in contrast to the plug-in (LZ) estimator in one-way clustering being conservative, the plug-in (CGM) estimator in two-way clustering is no longer conservative. This leaves researchers with two options: they can either assume that $\sum_{j \in \mathcal{N}_i} (\tau_{ki} - \tau_k) (\tau_{kj} - \tau_k) \geq 0$, or they can use CGM2.

Simulations

This section describes and presents results for simulations to illustrate how the theoretical results perform in practice. The simulations use the following procedure:

enumerate• Construct a population of size $n_k$. Each unit in the population is a tuple of $(g(i),h(i),Y_{ki}(1),Y_{ki}(0))$. • For every simulation $s$: \begin{enumerate} • Draw random variables $(R_{ki},W_{ki})$ for every unit and hence construct a sample of observations with $R_{ki}=1$. For every unit in the sample, $(g(i),h(i),W_{ki},Y_{ki})$ is observed. • Calculate $\hat{\tau}$ from the regression. • Calculate the variance estimates for various procedures and construct a confidence interval to determine if the CI covers $\tau$. \end{enumerate} • Aggregate the coverage over the simulations.

Then, every design modifies some part of the procedure. The following parts can be modified:

enumerate• How clusters $(g,h)$ are assigned in the population. \begin{enumerate} • Balanced: Use $n_k=1e6$, with $1000$ $G$ clusters and $1000$ $H$ clusters, so there is one observation for every intersection. • Staircase: Let $k_{1}$ denote an odd number, and $(g,h)$ describe a cluster intersection that is in cluster $g$ on the $G$ dimension and in cluster $h$ on the $H$ dimension. In the population, individuals can only belong to cluster intersections of the form $(k_{1},k_{1})$, $(k_{1},k\pm1)$ and $(k_{1}\pm1,k_{1})$. The mass of the population in $(k_{1},k_{1})$ is four times larger than cluster intersections of the other four forms, which are of equal mass. \end{enumerate} • How potential outcomes and hence $\tau_{ki}$ are constructed in the population. Use $u_i \sim N(0,0.1)$ so that $Y_{ki} = \tau_{ki}W_{ki} + u_i$. \begin{enumerate} • Gvar: $\tau_{ki}=\tau_{g(i)} + \tau_{h(i)}$, where $\tau_g=\pm2$ with equal probability and $\tau_h=\pm1/2$ with equal probability. • Hvar: $\tau_{ki}=\tau_{g(i)} + \tau_{h(i)}$, where $\tau_h=\pm2$ with equal probability and $\tau_g=\pm1/2$ with equal probability. • same: $\tau_{ki}=\tau_{g(i)} + \tau_{h(i)}$, where $\tau_g=\pm1$ with equal probability and $\tau_h=\pm1$ with equal probability. • constant: $\tau_{ki}=1$. • oddeven: $\tau_{ki}=1$ if $g(i)$ and $h(i)$ are both odd, and $-1$ otherwise. \end{enumerate} • How $R_{ki}$'s are generated for every sample. \begin{enumerate} • $p_k$ individual sampling probability • $q_k$ cluster sampling probability for $G$ • Multiway Sampling: The intersection $(g,h)$ is sampled when $A^{sam}_{kg}=1$ and $B^{sam}_{kh}=1$, each independently drawn from $Be(0.25)$. Then, $p_k=0.25$. \end{enumerate} • How $W_{ki}$'s are generated for every sample. \begin{enumerate} • Hway: Assignment probability $B_{kh}$ for every cluster is drawn from $U[0,1]$. $W_{ki} = 1$ with probability $B_{kh(i)}$. • none: $W_{ki}=1$ with probability 1/2 so $W_{ki}$ is iid. • AND: $A_{kg} \in\{0,1\}$ and $B_{kh}\in\{0,1\}$ are drawn independently, and each takes the value 1 with probability $1/\sqrt{2}$. Then, $W_{ki}=A_{kg(i)}B_{kh(i)}$. \end{enumerate}

The staircase design for constructing the clusters is artificially designed so that $V_{CGM,k}-v_k\leq0$. Details are in Appendix (ref). I use $nsim=5000$ simulation draws and report eight informative designs.

table[table omitted — 2,001 chars of source]

Table (ref) reports the results for various designs (D). D1 shows that there is massive over-coverage in two-way assignment. However, by virtue of clustered assignment, if we used the EHW standard errors, we would not have correct coverage. D2 shows how $\tau$ matters. Even though there is multiway assignment, because there is more variation in $\tau$ on the $H$ dimension, it suffices to cluster on $H$. D3 shows that when there is multiway sampling, we need to cluster on both dimensions to have valid confidence intervals --- doing so on just one dimension will under-cover. D4 reflects some necessity for considering sampling and assignment. When sampling and assignment occur on different clustering dimensions, if we had just clustered on either one, we would not have 0.95 coverage, so we need to cluster on both.

D5 is an example when CGM yields correct coverage, namely, when $\tau$ is constant. D6 shows that when there is clustered sampling on $G$ but no clustered assignment, it suffices to cluster on $G$. If we had clustered on $H$ too, the variance would have been twice as large. D7 shows that there is no need to cluster on $G$ when there is clustered assignment only on $H$. The variance is four times as large when clustering on the unnecessary dimension. D8 provides the counterexample where CGM under-covers when there is clustered sampling. The details of this construction is given in Appendix (ref).

Overall, the simulations illustrate the theoretical results that sampling and assignment matter. We can get exact coverage with multiway sampling or when treatment effects are constant. CGM can under-cover in multi-way sampling, while CGM2 is always conservative. In fact, the conservativeness of CGM2 can be quite substantial: CGM2 variance is five times what is required for valid inference in D7, for instance. While D8 shows how CGM may be anticonservative, it is usually possible to rule out such designs by institutional details or some economic model.

Design-based inference fundamentally relies on the institutional knowledge of how sampling and assignment occur. This knowledge then informs what dimensions the researcher should cluster on, if at all. When moving from one-way clustering to multi-way clustering, this paper has shown how such demands on institutional knowledge still apply. To use the plug-in variance estimator in multi-way settings, however, researchers need to make an additional assumption on the correlation of treatment effects within clusters. Without such an assumption, valid inference may require a variance estimator that is unnecessarily conservative.