Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
54,849 characters · 12 sections · 50 citation commands
Inference in Difference-in-Differences: How Much Should We Trust in Independent Clusters?
\def\spacingset#1{ {#1}} \spacingset{1}
\newsavebox{\tablebox} \newlength{\tableboxwidth}
We analyze the challenges for inference in difference-in-differences (DID) when there is spatial correlation. We present novel theoretical insights and empirical evidence on the settings in which ignoring spatial correlation should lead to more or less distortions in DID applications. We show that details such as the time frame used in the estimation, the choice of the treated and control groups, and the choice of the estimator, are key determinants of distortions due to spatial correlation. We also analyze the feasibility and trade-offs involved in a series of alternatives to take spatial correlation into account. Given that, we provide relevant recommendations for applied researchers on how to mitigate and assess the possibility of inference distortions due to spatial correlation.
\
{\it Keywords:} spatial correlation; clustered standard errors; linear factor model
\
{\it JEL Codes:} C12; C21; C23; C33
\spacingset{1.45}
\doublespace
Difference-in-Differences (DID) is one of the most widely used methods for identification of causal effects in social sciences. However, inference in DID can be complicated by both serial and spatial correlations. \Copy{Cluster}{ After an influential paper by Bertrand04howmuch, showing that serial correlation can lead to severe over-rejection in DID applications, most papers applying DID use inference methods that are robust to arbitrary forms of serial correlation. A common alternative in this case is to rely on cluster robust variance estimator (CRVE) at the unit level, which allows for arbitrary serial correlation, but generally relies on the assumption that these unit-level clusters are independent. In most cases, DID papers do not take the possibility of spatial correlation across these clusters into account.\footnote{In case we have, for example, individual-level data and a state-level policy, clustering at the state level would allow for arbitrary correlation between individuals in the same state. Since clustering at the state level takes within-state correlation into account, we focus on the possibility of across state spatial correlations. } }
While there are some alternatives for inference in the presence of spatial correlation, they generally require knowledge about the relevant distance metric, impose assumptions on the serial correlation, and/or rely on more data (such as a large number of periods).\footnote{For example, CONLEY19991, KIM201385, CT (in their online appendix A.3), BESTER2011137, and Muller1,Muller2 rely on distance measures across units. Other papers exploit the time dimension to perform inference in the presence of spatially correlated shocks. However, these methods rely on a large number of periods (for example, VOGELSANG2012303, FP (Section 4) and Chernozhukov). ferman2020inference considers a setting in which spatial dependence is unknown, and the number of pre-treatment periods is fixed. However, his conclusions rely on a strong mixing condition for the spatial correlation, and are only valid for settings with few treated and many control units. } We discuss different settings in which such alternatives are unfeasible in DID applications. As we illustrate below, this may be the case when the relevant source of spatial correlation is unknown by the applied researcher. Moreover, even when the relevant distance metric is known, there might not be enough variation in the data to estimate the spatial correlation. We also show that, even when feasible, some alternatives might have important limitations that have not been previously considered in the literature.\footnote{For example, while the common correlated effects estimator proposed by Pesaran may provide an interesting alternative in some settings, we discuss some limitations in Appendix (ref). We also discuss in Appendix (ref) the use of the inference method proposed by Muller1,Muller2, for settings in which the distance metric is known, but there is limited variation on the distance metric in the data. }
Given that correcting for spatial correlation is not always feasible, we consider the consequences of ignoring spatial correlation in DID applications. We analyze a setting in which the spatial correlation follows a linear factor model, which allows for a rich variety of spatial correlation structures. We show in Section (ref) that, in a setting with no variation in treatment timing, inference ignoring spatial correlation becomes more problematic when both (i) the variance of the difference between the pre- and post-treatment averages of common shocks is large relative to the variance of the same difference for the idiosyncratic shocks, and (ii) the distribution of factor loadings has different expected values for treated and control units. When at least one of these conditions does not hold, the time and/or unit fixed effects would absorb most of the relevant spatial correlation. This provides novel insights on the settings in which spatial correlation should lead to more or less distortions for inference in DID applications. In particular, we show that details such as the time frame used in the estimation, the choice of the treated and control groups, and the choice of the estimator, are key determinants of distortions due to spatial correlation.
We then present in Section (ref) two sets of simulations, based on the American Community Survey (ACS) and the Current Population Survey (CPS). In the first one, we illustrate a setting in which the source of spatial correlation is unknown to the applied researcher. In the second one, the source of spatial correlation is known, but there is not enough variation in the data to take that into account. In both cases, ignoring spatial correlation does not significantly affect inference when the time frame is short, but can lead to relevant size distortions when the time frame is long. Based on our theoretical results, this is consistent with common shocks also being more serially correlated relative to the idiosyncratic shocks. In the second setting, it is also possible to ameliorate the spatial correlation problems by considering treated and control groups that are more alike. This is also consistent with our theoretical results. Section (ref) concludes with recommendations for applied researchers.
We start presenting in Section (ref) a general DID model in which we discuss the consequences of ignoring spatial correlation for inference. In Section (ref), we discuss existing alternatives to take spatial correlation into account, and explain the reasons why correcting for spatial correlation may be unfeasible in some settings. In Section (ref), we impose more structure on the spatial correlation, so that we can provide further insights on the settings in which we should expect spatial correlation to lead to more or less distortions for inference. Throughout, we consider the case in which a parallel trends assumption remains valid, so that the inclusion of spatial correlation does not imply that the DID model is misspecified.
We start considering a standard model for the potential outcomes. Let $Y_{jt}(0)$ ($Y_{jt}(1)$) be the potential outcome of unit $j$ at time $t$ when this unit is untreated (treated) at this period. We consider first that potential outcomes are given by
where $\theta_j$ and $\gamma_t$ are, respectively, unit- and time-invariant unobserved variables, while $\eta_{jt}$ represents unobserved variables that may vary at both dimensions. We do not impose any restriction on the serial and spatial correlations of $\eta_{jt}$, so this is a very general model; $\alpha_{jt}$ is the (possibly heterogeneous) treatment effect on unit $j$ at time $t$.\footnote{All results remain unchanged in case we consider a setting with individual-level observations $i$ within units $j$, $Y_{ijt}$, and we consider clustering at the unit level. Within-unit spatial correlation can be taken into account by clustering at the unit level, so we are mainly concerned about the possibility of across-units spatial correlation. }
Equation (ref) leads to a standard DID model for $Y_{jt} = d_{jt}Y_{jt}(1) + (1-d_{jt}) Y_{jt}(0)$, given by
where $\bar \alpha$ is defined as the two-way fixed effect (TWFE) estimand, $ \widetilde \eta_{jt} = \eta_{jt} + (\alpha_{jt} - \bar \alpha) d_{jt}$, and $d_{jt}$ is an indicator variable equal to one if unit $j$ is treated at time $t$, and zero otherwise.
Consider a simpler case in which $d_{jt}$ changes to 1 for all treated units starting after date $t^\ast$, and define a dummy variable $D_j$ equal to one if unit $j$ is treated. There are $N_1$ treated units, $N_0$ control units, and $T$ time periods. Let $\mathcal{I}_1$ ($\mathcal{I}_0$) be the set of treated (control) units, while $\mathcal{T}_1$ ($\mathcal{T}_0$) be the set of post- (pre-) treatment periods. For a generic variable $A_t$, define $\nabla A = \frac{1}{T-t^\ast} \sum_{t \in \mathcal{T}_1} A_{t} - \frac{1}{t^\ast} \sum_{t \in \mathcal{T}_0} A_{t}$. In particular, we consider $W_j = \nabla \widetilde \eta_{j}$, which is the post-pre difference in average errors for each unit $j$.
In this section, we consider a repeated sampling framework over the distribution of $\{W_j \}_{j \in \mathcal{I}_0 \cup \mathcal{I}_1 }$, conditional on $\mathbf{D} = \mathbf{d}$, where $\mathbf{D} = (D_1,...,D_N)$. For example, we can think of $W_j$ as a linear combination of economic or weather shocks that may affect unit $j$, and we analyze the distribution of the DID estimator over the distribution of those shocks. In this case, we have that $\bar \alpha = \mathbb{E} [ \frac{1}{N_1} \frac{1}{T-t^\ast}\sum_{j \in \mathcal{I}_1} \sum_{t \in \mathcal{T}_1}\alpha_{jt} | \mathbf{D} = \mathbf{d} ]$.
In this setting, the DID estimator is the same as the TWFE estimator, which is given by
\Copy{W}{If we have $\mathbb{E}[W_j | \mathbf{D} = \mathbf{d}] =0$ for all $j$, then the DID estimator $\hat \alpha$ will be unbiased for $\bar\alpha$, regardless of the assumptions on the serial and spatial correlations of $\widetilde \eta_{jt}$. } However, inference is only possible if we impose assumptions on either the serial or the spatial correlation of $\widetilde \eta_{jt}$. Most commonly, inference methods for DID do not impose restrictions on the serial correlation of $\widetilde \eta_{jt}$, but assumes that $\widetilde \eta_{jt}$ are independent across $j$.\footnote{See, for example, Arellano, Bertrand04howmuch, cameron2008bootstrap, IZA, CT, FP, Canay, and MW. } A common alternative in this case is to rely on CRVE at the unit level which, assuming independence across $j$, is valid when both $N_1$ and $N_0$ are large.
Now consider a setting in which treatment allocation is such that units that are exposed to similar shocks are also more likely to be allocated into the same treatment status. Then, once we condition on $\mathbf{D} = \mathbf{d}$, we should expect a strong correlation between $W_j$ and $W_{j'}$ if $j$ and $j'$ received the same treatment allocation. In such cases, not taking such spatial correlation into account can lead to over-rejection. The intuition is the following. Imagine there is an unobserved variable in $W_j$ that equally affects all treated units, but does not affect the control units.\footnote{We assume that the expected value of this variable is equal to zero conditional on $\mathbf{D}= \mathbf{d}$, so the presence of such correlated shock does not affect the identification assumption of the DID model.} If the null $H_0: \bar \alpha=0$ is true, then $\hat \alpha = \frac{1}{N_1} \sum_{j \in \mathcal{I}_1} W_j - \frac{1}{N_0} \sum_{j \in \mathcal{I}_0} W_j$. Therefore, under the null, finding a “large” value for $\hat \alpha$ would only be possible if many of those $W_j$ for $j \in \mathcal{I}_1$ were positive, and/or many of those $W_j$ for $j \in \mathcal{I}_0$ are negative. We would consider that this event has a lower probability than the true one if we (mistakenly) assume that $W_j$ are independent, leading to over-rejection.
There are alternatives for inference when we relax the assumption that clusters are independent. However, such alternatives generally assume that there is a distance metric across units, impose assumptions on the serial correlation, and/or rely on more data, (such as a large number of periods).\footnote{See Footnote (ref).}
We focus on settings in which alternatives to take spatial correlation into account may be unfeasible. For example, it may be that the relevant source of spatial correlation is unknown to the econometrician. While it is natural to think about spatial correlation considering geographical distances, the relevant spatial correlation may arise from other sources. For example, we consider in Section (ref) simulations in which PUMA's with some specific industry compositions are more likely to receive treatment. In such case, even if we assume conditions such that the DID estimator is unbiased, ignoring spatial correlation from unobserved shocks that are related to industry composition might generate relevant size distortions. Moreover, attempts to correct for that considering that the relevant distance metric is geographical would generally not solve the problem.
There are alternatives that take spatial correlation into account even when the source of spatial correlation is unknown, exploiting the time series of the data.\footnote{See Footnote (ref). Also, considering the use of two-way cluster at the unit and time dimensions would not provide a valid solution in this setting, even if both $N$ and $T$ are large, because it would not take into account the correlation between $\eta_{jt}$ and $\eta_{j't'}$, for $j \neq j'$ and $t \neq t'$. See multi_way, THOMPSON20111, Davezies, menzel, and Webb_multi for recent developments on multi-way clustering. } However, such alternatives generally require a large number of periods, while, as Roth points out, settings with short time series are prevalent in DID applications. One exception that may take spatial correlation into account, even when the source of spatial correlation is unknown and $T$ is finite, is the common correlated effects (CCE) estimator, proposed by Pesaran. However, we show in Appendix (ref) that there are some limitations and trade-offs involved in using this alternative in our setting. First, it requires variation in treatment timing, so it would not be an option in common settings in which all treated units start treatment at the same time. Also, the CCE estimator imposes restrictions on the dimension of the common shocks.\footnote{We present simulations in Appendix (ref) in which the spatial correlation comes from a linear factor model, as we consider in Section (ref). Inference for CCE leads to large over-rejections when the dimension of the linear factor model is greater than two.} Moreover, in some cases there is a loss in precision relative to considering the TWFE estimator. Finally, if treatment effects are heterogenous, we show that the CCE estimand may be negative even when treatment effects are always positive.\footnote{This problem has been documented in the DID literature for the TWFE estimator Bacon,chaisemartin2018twoway, but not for the CCE estimator. We also show that standard solutions to this problem are unfeasible when we consider the CCE estimator.}
Moreover, even if the source of spatial correlation is known, it might be unfeasible to take the spatial correlation into account. This may happen when we do not have enough variation in the data to estimate the relevant spatial correlations. As an example, suppose we have data on students' test scores for grades one and two, and we have a policy that affected only second graders in the post-treatment periods. In this case, there might be common shocks that differentially affect different grades, which might generate relevant spatial correlation for the DID estimator. However, it would unfeasible to cluster at the grade level, or to consider alternative spatial correlation-robust methods, with only two grades.
Finally, we note that a commonly-used rule-of-thumb is to consider CRVE “at the level of the treatment assignment”.\footnote{See, for example, NBERw24003 and MW2020.} Consider, for example, a setting in which we analyze a state-level policy, and we have county-level data. In this case, clustering at the county level would generally lead to over-rejection. In contrast, if we have treatment completely randomly assigned at the state level, then CRVE at the state level would be valid if we have a large number of treated and control states, as IK and NBERw24003 show considering a design-based approach for inference.\footnote{Following IK, we consider that treatment is “completely randomly assigned at the state level” if all possible treatment allocations subject to the constraints on the number of treated states have the same probability.} While we focus in the main text on a setting in which potential outcomes are stochastic, we also consider in Appendix (ref) a design-based approach for inference.
More generally, however, clustering at the level of the treatment assignment may not solve the problem in case we have more complex treatment assignments. For example, consider we have two regions, and treatment is assigned based on a two-stage randomization. First, we have that the proportions of treated states in regions A and B are either $(70\%,30\%)$ or $(30\%,70\%)$, with equal probabilities. Then, states within each region are randomly allocated into treatment according to those proportions. In this case, the DID estimator is unbiased rambachan2020designbased. In such setting, we have that counties within the same state have the same treatment status. Moreover, in each region we would have both treated and control counties. Therefore, an applied researcher looking at the data might say that “treatment is allocated at the state level.” However, clustering at the state level would not generally be valid in this case.
While an alternative in this case would be to cluster at a higher level (in this case, regions), the econometrician may be unaware of this more complex assignment design, and/or not have information to construct the relevant cluster level. Moreover, even if this information is available, when we consider clustering at higher levels, we may end up with very few clusters to estimate the standard errors. While there are alternatives that work in settings with few clusters,\footnote{See, for example, cameron2015practitioner, Ibragimov, Canay, Hagemann. } such alternatives generally do not work well in the limit when we end up with only two or three clusters.
In order to provide further insights on the implications of spatial correlation, we impose more structure on the errors. We assume that potential outcomes follow a linear factor model
where $\lambda_t$ is an $(1 \times F)$ vector of common shocks, while $\mu_j$ is an $(F \times 1)$ vector of factor loadings determining how unit $j$ is affected by $\lambda_t$. While $\theta_j$ and $\gamma_t$ could have been included as components of $\mu_j$ and $\lambda_t$, we consider them separately to highlight that we can still have time-invariant and unit-invariant shocks as in standard DID model, so what we add is the possibility of other spatially correlated shocks that are not time- nor unit-invariant, which are captured by $\lambda_t \mu_j$.\footnote{We discuss in Appendix (ref) the possibilities of using alternative estimators designed for panel data settings with an error structure following a linear factor model. }
Such structure allows for a rich variety of spatial correlation structures. We can consider, for example, the case in which spatial correlation comes from counties with similar industry compositions having correlated errors. In this case, we would have $F$ industries, and vector $\mu_j$ would represent the exposure of county $j$ to each of these industries, while $\lambda_t$ would represent industry shocks. {We can also consider the case of $N$ municipalities divided into $F$ states, where there are relevant state-level shocks. In this case, if municipality $j$ belongs to state $f$, we could model that by setting the $f-$th entry of $\mu_j$ equal to one and zero otherwise.}\footnote{This simple structure would not allow for arbitrary spatial correlation within states, as it considers a common state-level shock. We would be able to consider more complex within-state correlations by increasing the dimension of the $\lambda_t$. } As another example, this structure can encompass the common notion that spatial correlation depends on geographical distances. Finally, note that the spatial correlation structure may involve different notions of distance (for example, depending on both geographical position and industry composition distances, as considered in the simulations in Section (ref)).\footnote{While we focus in the case in which the dimension $F$ is fixed, we consider in Appendix (ref) a setting in which the dimension $F$ may increase with $N$. }
We continue to consider that treated units start treatment after $t^\ast$, and let $D_j =1$ if unit $j$ is treated, and $0$ otherwise. But now we consider the distribution of the DID estimator based on a repeated sampling framework over the distributions of $D_j$, $\lambda_t$, $\mu_j$, $\epsilon_{jt}$ and $\alpha_{jt}$.
Assumption (ref) implies that all spatial correlation is captured by this linear factor structure, so that the idiosyncratic shocks $\epsilon_{jt}$ are independent across $j$. We do allow, however, for arbitrary serial correlation in both $\epsilon_{jt}$ and $\lambda_t$. We also assume for simplicity that treatment effects $\alpha_{jt}$ are independent across $j$.\footnote{See Footnote (ref) for the consequences of relaxing this assumption.} We do not need to impose assumptions on $\theta_j$ and $\gamma_t$.
{Since we do not restrict the dependence between $\mu_j$ and $D_j$, this sampling scheme can encompass settings in which we have relevant spatial correlation in the treatment assignment mechanism. For example, consider the case in which $\mu_j$ represents exposure to specific industry shocks. In this case, we may have that the probability of being assigned to treatment is larger for units that, for example, are more exposed to a specific industry, generating relevant spatial correlation. Likewise, if we think about the factors as representing geographical locations, then this formulation would allow for spatial correlation due to geographical distance. }
The TWFE estimand in this case is given by $\alpha \equiv \mathbb{E}[\frac{1}{T - t^\ast} \sum_{t \in \mathcal{T}_1} \alpha_{jt} | D_j =1]$, which we can think of as the population average treatment effects on the treated. If we let $\mu^e = \mathbb{E}[ \mu_j]$, and $\mu_w^e = \mathbb{E}[\mu_j | D_j=w]$, for $w \in \{ 0,1\}$, then
where, with some abuse of notation, $\nabla \alpha_j$ is the post-treatment average of $\alpha_{jt}$ across $t$.
We consider a setting in which the linear factor structure does not affect the counterfactual trends, so the DID model is not misspecified. We impose the following assumption, which implies a standard parallel trends assumption $\mathbb{E}[\nabla Y_j(0) | D_j = 1] = \mathbb{E}[\nabla Y_j(0) | D_j = 0]$.\footnote{This assumption is implied by the assumption of parallel trends for all periods. We can extend our results to consider alternative parallel trends assumptions Marcus.}
The first part of Assumption (ref) states that idiosyncratic errors are uncorrelated with treatment assignment. The second part implies that factor structure does not affect the expected value of the DID estimator. Note that $\mathbb{E}[\nabla \lambda] (\mu_1^e - \mu_0^e) = \sum_{f=1}^F \mathbb{E}[\nabla \lambda(f) ] (\mu_1^e(f) - \mu_0^e(f))$, where $v(f)$ is the $f-$th coordinate of vector $v$. If we do not take into account knife-edge cases in which elements of this sum cancel out, Assumption (ref) implies that, for each $f=1,...,F$, either one of two conditions hold. First, it may be that $\mathbb{E}[\bar \lambda_{\mbox{\tiny post}}(f)] = \mathbb{E}[ \bar \lambda_{\mbox{\tiny pre}}(f)]$, so the first moment of the distribution of the common factor $f$ is stable in the pre- and post-treatment periods. In this case, even if treated and control units are differentially affected by this common factor, this would not generate bias on the DID estimator over the distribution of $\lambda_t(f)$. Alternatively, it may be that $ \mu^e_1(f) = \mu^e_0(f)$. In this case, even if the expected value of $\lambda_t(f)$ differs in the pre- and post-treatment periods, this common factor does not systematically affect treated units differently relative to control units, so this would not generate bias for the DID estimator over the distribution of $\mu_j(f)$. Since we also have $\mathbb{E}[\nabla \alpha_j | D_j=1] = \alpha$, Assumption (ref) implies that $\hat \alpha$ is unbiased.
Overall, we can think that there are unit- and/or time-invariant unobserved variables that may be arbitrarily correlated with treatment assignment, but the other common shocks are not correlated with treatment assignment once we condition on these fixed effects.
In order to derive the asymptotic distribution of the DID estimator in this setting, we consider a local-to-0 approximation in which the variance of $\nabla \lambda (\mu^e_1 - \mu^e_0)$ drifts to zero. This way, we can rely on an asymptotic theory to approximate settings in which the ratio between the variance of the common shocks and the variance of the average of the idiosyncratic shocks assumes any value in $[0,\infty)$. Therefore, we can consider approximations to settings in which common shocks have negligible, moderate, or large relevance relative to the sampling variation.\footnote{If we do not consider a local asymptotics, this ratio would diverge when $N \rightarrow \infty$, and this would not provide reasonable approximations to many relevant applications. Roth considers a similar assumption. We consider in Appendix (ref) the case in which $var(\nabla \lambda (\mu^e_1 - \mu^e_0)$ does not drift to zero.}
We present details of the proof in Appendix (ref). While $\hat \alpha$ is unbiased despite the spatial correlation, Proposition (ref) shows that $\hat \alpha$ may not be asymptotically normal if $\nabla \xi (\mu^e_1 - \mu^e_0)$ is not normally distributed. As a consequence, we may have distortions for inference based on a $t$-statistic for two reasons. First, the asymptotic distribution of the $t$-statistic, under the null, will have a variance greater than one. Second, the asymptotic distribution of the $t$-statistic, under the null, may not be normal.
If we assume that $\nabla \xi(\mu_1^e - \mu_0^e)$ is normally distributed, then the $t$-statistic based on CRVE, under the null, would be asymptotically normal with mean zero, but its variance would be greater than one if by $\Lambda_\lambda \equiv (\mu^e_1 - \mu^e_0)' \Omega (\mu^e_1 - \mu^e_0)>0$.
Proposition (ref) and Corollary (ref) make it clear that, when $\Lambda_\lambda >0$, ignoring spatial correlation in this setting leads to over-rejection for two-sided $t$-tests.\footnote{Even if $V$ is not normal, $Z \sim N(0,1)$ and $Z \perp V$ implies that, for any $c \in \mathbb{R}$, $Pr(|Z+V|>c) = \int Pr(|Z+v|>c)dF_v(v) > \int Pr(|Z|>c)dF_v(v) = Pr(|Z|>c)$. This inequality follows from $Pr(|Z + v | > c) > Pr(|Z| > c)$ for any $v \neq 0$, and $Pr(V = 0) < 1$.} Moreover, over-rejection will be larger when $\Lambda_\lambda$ is larger relative to $\Lambda_\epsilon \equiv \frac{1}{c} \sigma_\epsilon^2(1) + \frac{1}{1-c} \sigma_\epsilon^2(0)$.\footnote{Under the assumption that $\nabla \alpha_j$ is iid, a larger treatment effect heterogeneity ($var(\nabla \alpha_j | D_j =1)$) leads to larger $\sigma_\epsilon^2(1)$, which in turn implies smaller underestimations by the CRVE. This happens because the treatment effect heterogeneity in this case is captured by the CRVE. However, we should expect the opposite in case there is strong spatial correlation in $\nabla \alpha_j$. Overall, whether larger treatment effects heterogeneity leads to more or less underestimation by the CRVE depends on the degree of spatial correlation in the treatment effects heterogeneity. } Importantly, implementation details, such as the time frame used in the estimation and the choice of the control group will affect the relative magnitude between $\Lambda_\lambda$ and $\Lambda_\epsilon$.
Choice of the time frame: if the spatially correlated shocks are also more serially correlated than the idiosyncratic shocks, then considering shorter time frames around the treatment would lead to less size distortions. More specifically, assume that $\xi_t (\mu_1^e - \mu_0^e)$ follows an AR(1) process with serial correlation $\rho_\xi$, while $\epsilon_{jt}$, conditional on either $D_j = 0$ or $D_j = 1$, follows an AR(1) process with serial correlation $\rho_\epsilon$. Consider a DID estimator using $T/2$ periods before and $T/2$ periods after the treatment. As we show in Appendix (ref), if $0 \leq \rho_\epsilon < \rho_\xi < 1$, then inference distortions based on CRVE are increasing in $T$.\footnote{In Appendix (ref), we derive the formula for $\phi(\rho_\xi,\rho_\epsilon,T) = \Lambda_\lambda / \Lambda_\epsilon$, in a setting in which $\xi_t$ and $\epsilon_{j,t}$ are AR(1), and we have a DID estimator with $T$ periods. We show numerically that $\phi(\rho_\xi,\rho_\epsilon,T)$ is increasing in $T$ when $ 0 \leq \rho_\epsilon < \rho_\xi <1 $ for all reasonable values of $T$. }
An important caveat is that considering different time frames implies that the DID estimand may change. More specifically, the estimand would be given by $\mathbb{E}[\frac{1}{\tilde t} \sum_{t \in \widetilde{\mathcal{T}}} \alpha_{jt} | D_j =1]$, where $\widetilde{\mathcal{T}}$ is the set of post-treatment time periods used to construct the DID estimator. Therefore, if we consider shorter time frames, then we would only estimate short-term effects of the policy. Note also that changing the time frame also implies that we would consider a modified Assumption (ref), which can be valid considering only the time periods used for the estimation.
Likewise, consider a dynamic DID, in which one uses a base period (for example, $t^\ast$) and a period $t^\ast + \tau$ for varying $\tau$, where $\tau < 0$ provides evidence on the parallel trends assumptions, while $\tau > 0$ provides estimates for the effect $\tau$ periods after the treatment. Our results also imply that the degree of size distortions when spatial correlation is ignored can vary substantially for different values of $\tau$. In particular, under the assumption $0 \leq \rho_\epsilon < \rho_\xi < 1$, we should expect more size distortions when $|\tau|$ increases (more details in Appendix (ref)).
Choice of treated and control units: Proposition (ref) and Corollary (ref) also imply that size distortions would be lower when $\mu_1^e \approx \mu_0^e$. In such cases, the time fixed effects would absorb most of the spatial correlation, and inference based on CRVE at the unit level would lead to less distortions.\footnote{This is related to the idea of using state-border DID, as considered by Dube. }
In settings in which the nature of the spatial correlation is unknown, it would not be possible to select the treated/control group taking that into account. However, in some settings this conclusion may be useful in practice. For example, consider a setting in which we observe students from grades one to four, and consider a treatment that starts in the post-treatment periods for grades three and four. In this case, if closer grades are more similarly exposed to the common shocks relative to more distant grades, then we should expect smaller size distortions if we consider a DID estimator comparing students from grades two and three, relative to a DID estimator using the full sample. In Section (ref), we present simulations based on the CPS that corroborate this conclusion. An important caveat is that, if treatment effects are heterogenous, then this approach might change the DID estimand. In this example, we would estimate the effects for third graders (instead of an average effect for third and fourth graders).
We consider two sets of simulations, one that mimics a setting in which the relevant source of spatial correlation is unknown by the econometrician, and another one in which it is known, but there is not enough variation to take that into account.
We first consider MC simulations in which we estimate a spatial correlation structure based on the ACS ipums. We aggregate the data at the Public Use Microdata Area (PUMA) $\times$ year level, considering 2005 to 2019. Following Bertrand04howmuch, we restrict the sample to women between the ages 25 and 50, and focus on log wages as the outcome variable. We consider simulations in which we fix the total number of years, $T$, and treatment starts in the middle of the time frame. For a given $T \in \{2,3,...,15 \}$, we estimate a covariance matrix in which $cov(W_j,W_k)$ may depend on whether PUMA's $j$ and $k$ are in the same state, and/or on whether they have similar industry compositions. We present details on how the DGP is constructed in Appendix (ref).
For a given $T$, we simulate a Gaussian model with the estimated covariance structure for such $T$.\footnote{Note that the covariance structure estimated from the ACS is identified even when errors are not normal. In particular, if the estimated covariance structure is such that $var(Z + V) \approx 1$, then $var(V) \approx 0$ (where $Z$ and $V$ are defined in Proposition (ref)). Therefore, we should expect the same patterns regarding settings in which spatial correlation does not lead to large distortions if we did not assume a Gaussian DGP. Consistent with that, a previous version of this paper Ferman_OLD presented simulations in a design-based approach, in which we did not impose that the potential outcomes are normal, and found similar results. Moreover, the simulations in Section (ref) do not assume normality, and find similar results. } We consider two alternative treatment assignment mechanisms. In the first one, we consider PUMA's completely randomly assigned. As presented in Figure (ref).A, even if we consider CRVE at the PUMA level, rejection rates are close to 5% regardless of the spatial correlation in the DGP. This is consistent with the conclusions from IK.
In the second assignment mechanism, we consider a setting in which PUMA's with industry compositions that are more concentrated in manufacturing have higher probability of receiving treatment. Therefore, in this case $\mu_1^e \neq \mu_0^e$ for the common shocks related to industry composition. Still, since $\mathbb{E}[W_j | D_j]=0$ for all $j$, the DID estimator is unbiased.\footnote{We provide evidence that this is a reasonable assumption in this setting in Appendix (ref). }
It is conceivable that an applied researcher might be unaware about (or may not have information on) such industry-level shocks. Therefore, we consider first inference based on CRVE at the PUMA level. In this case, rejection rates are relatively close to 5% when $T$ is small (for example, at 7% when $T=2$), but over-rejection becomes more problematic when $T$ increases, reaching 19% when $T=15$ (Figure (ref).B). Considering the results from Section (ref), this is consistent with industry-level shocks being more serially correlated relative to the idiosyncratic shocks.\footnote{We recall that the parameters of the spatial correlation in this DGP were estimated based on a real (and widely used by applied researchers) dataset. If it were the case that the data is such that the idiosyncratic shocks are relatively more serially correlated, then we should expect the reverse pattern in terms of size distortions in Figure (ref).B. } More generally, this example illustrates that the relevance of (ignored) spatial correlation depends crucially on the time frame considered in the application. We present in Appendix Figure (ref) results for a dynamic DID specification. We similarly find that size distortions are minor when we consider estimation of shorter-term effects, but become more relevant when we consider longer-term effects.
Now consider that the applied researcher attempts to correct for spatial correlation, but considers a geographical distance as the relevant distance metric. We consider a wild cluster bootstrap (WCB) at the state level.\footnote{Results with CRVE at the state level are similar, but with slightly larger rejection rates due to some large state clusters.} Over-rejection becomes slightly smaller in most cases, but we still find relevant over-rejection when $T $ is large. The reason is that the state-level cluster captures some of the industry-level shocks, because some states have a relatively higher concentration of PUMA's more exposed to manufacturing. However, we still have relevant over-rejection when $T$ is large, because there are relevant across-state correlations that are not taken into account. We also consider a border DID approach. Again, this approach slightly ameliorates the inference problem, but does not completely solve it.
Finally, not surprisingly, if the applied researcher had complete knowledge that the relevant spatial correlation came from such industry shocks, then clustering at the industry-group level would be valid regardless of $T$.
We now present simulations using the CPS data from 1990 to 2018, still considering log wages for women between the ages of 25 and 50. For each simulation, in addition to selecting a time frame with $T \in \{2,3,...,10 \}$, we also select an age frame with $\delta_{\mbox{\tiny age}} \in \{2,3,...,10 \}$. We construct a DGP based on this dataset in which we allow for individuals of similar ages to be more spatially correlated, in addition to allowing for within-state correlation and for serial correlation. We present in details how this DGP is constructed in Appendix (ref). Given $T$ and $\delta_{\mbox{\tiny age}}$, we consider simulations in which treatment starts in the second half of the years for individuals above median in terms of age. In this setting, we expect relevant spatial correlation if individuals of closer ages (whether or not they are in the same state) are likely to be affected by similar shocks. We consider DID regressions including state $\times$ age-group fixed effects, time fixed effects, and the DID dummy. In those simulations, the DID estimator is unbiased, and the null hypothesis is true.\footnote{We present evidence in Appendix (ref) that it is reasonable to assume parallel trends in these simulations.}
Table (ref) presents rejection rates when inference is based on CRVE at the state level. There is large over-rejection (with rejection rates up to 36%) when both the time and the age frames are large. In contrast, there is not much over-rejection when $T$ is small, regardless of the age frame. This is again consistent with spatially correlated shocks being relatively more serially correlated relative to the idiosyncratic shocks. More interesting, even when $T$ is large, there is not much over-rejection when we keep the age frame small. This is consistent with the theoretical results that the distortions are mitigated when treated and control groups are more similar. Differently from the setting considered in Section (ref), in this setting it would be possible to select treated and control groups in such a way.
In Appendix (ref), we show that spatial-correlation robust standard errors do not work well in these simulations, given that we have little variation in the age groups, even when $\delta_{\mbox{\tiny age}}=10$.
Spatial correlation can lead to substantial over-rejection. Whenever feasible, applied researchers should consider methods that take that into account, as the ones discussed in Section (ref). However, there are common settings in which such solutions are unfeasible. Also, even when feasible, some alternatives may involve relevant trade-offs. For example, with variation in treatment timing, the CCE estimator may take spatial correlation into account, but there are some limitations and trade-offs in considering this alternative (see details in Appendix (ref)). Therefore, even if an applied researcher decides to use the CCE estimator, it might be valuable to also consider DID alternatives as a robustness check.\footnote{For example, it might be interesting to consider a DID estimator that deals with the problem of aggregating heterogeneous treatment effects when there is variation in treatment timing, which is a potential problem for the CCE estimator.}
Given that, the results we present in this paper provide guidelines on how applied researchers could proceed in empirical applications to mitigate and assess the relevance of spatial correlation when relying on inference methods that assume independent errors in the cross-section (such as CRVE at the unit level).
Consider a setting with more than one pre- and post-treatment periods. In this case, a longer time series would imply larger over-rejection if common factors exhibit stronger serial correlation relative to the idiosyncratic shocks. The simulations from Section (ref) provide evidence that this is the case for the ACS and CPS datasets. One robustness check in this case is to consider a specification restricting the sample to a few periods before and a few periods after the treatment. In this case, the unit fixed effects would absorb more of these common shocks, making inference assuming independent units more reliable. An important caveat is that, in this case, the DID estimand would provide the short-term effect of the policy. As discussed in Remark (ref), in this setting it would also possible to use pre-treatment data to check whether inference based on short differences is indeed reliable.
Relatedly, we also show that, in dynamic DID specifications, size distortions may vary substantially, depending on the time horizon that we analyze. In particular, if common factors exhibit stronger serial correlation relative to the idiosyncratic shocks, then we should expect relatively larger inference distortions for longer-term effects.
Another alternative to mitigate the spatial correlation problem is to make treated and control units as similar as possible. This alternative is unfeasible if the source of spatial correlation is unknown. However, as illustrated in the simulations in Section (ref), this can be a valid alternative in case there is information about the source of spatial correlation, but spatial correlation-robust standard errors do not work well.
Overall, this paper analyzes the challenges for inference in DID when there is spatial correlation. We present a series of novel insights and empirical evidence on the settings in which ignoring spatial correlation should lead to more or less distortions in DID applications. We show that details such as the time frame used in the estimation, the choice of the treated and control groups, and the choice of the estimator, are key determinants of distortions due to spatial correlation. We also analyze in detail the feasibility and trade-offs involved in a series of alternatives to take spatial correlation into account. Given that, we provide relevant recommendations for applied researchers on how to mitigate and assess the possibility of inference distortions due to spatial correlation.
\singlespace