Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
42,380 characters · 11 sections · 47 citation commands
ON THE USE OF DESIGN-BASED SIMULATIONS
Keywords: design-based simulations, inference, shift-share designs, spatial correlation
JEL Codes: C12; C21
Simulation-based analyses play a central role in econometric research. They are routinely used to study the finite-sample properties of estimators and inference procedures, and to illustrate potential failures of commonly used methods. An important class of such exercises consists of what we refer to as design-based simulations, in which realized outcomes are treated as fixed and uncertainty in the simulations is generated by resampling treatment assignment, shocks, or other components of the research design.
Design-based simulations have been used in methodological papers to illustrate inference problems in prominent empirical settings. For example, Bertrand04howmuch simulate placebo laws using state-level data on female wages from the Current Population Survey to illustrate severe size distortions of conventional inference methods in difference-in-differences (DiD) designs caused by serial correlation. More recently, design-based simulations have played a central role in the analysis of inference methods in shift-share designs. Adao and Borusyak use simulations based on resampling sectoral shocks to illustrate how spatial correlation can invalidate standard inference procedures and to analyze the properties of alternative inference methods in this setting. Related simulation-based approaches are also considered by Chaisemartin_cluster in the context of small-strata randomized experiments. Beyond their use as illustrative tools in methodological work, design-based simulations can also be employed by applied researchers as application-specific diagnostics. In this context, researchers may use design-based simulations tailored to their empirical design to assess whether a given inference method is likely to perform well in their specific setting.
In design-based simulations, researchers usually define a data-generating process (DGP) based on an empirical application or a specific dataset, in which potential outcomes are held fixed and the null hypothesis is satisfied by construction. Interpreting the results from such simulations, however, can be subtle. Whether the resulting rejection frequencies are informative about validity of an inference method in the original empirical setting depends critically on how the DGP used in the design-based simulations relates to the true DGP along dimensions relevant for inference. As a result, design-based simulations may be highly informative in some settings, but potentially misleading in others. Moreover, design-based simulations have often been used to illustrate issues related to the distribution of the errors (for example, serial or spatial correlation), which at first glance appears at odds with the fact that potential outcomes---and therefore errors---are held fixed in these simulations.
This paper studies the properties of design-based simulations, with a particular focus on shift-share designs, where such simulations have been especially influential in recent methodological debates. We show that when treatment effects are present, standard design-based simulations may mechanically confound treatment effects with error dependence, overstating the over-rejection of inference methods that do not allow for spatial correlation. We then discuss two alternative design-based simulation designs that circumvent these problems. Our analysis also clarifies how design-based simulations, despite holding errors fixed, can nonetheless be informative about characteristics of the error distribution. We also show that the alternative design-based simulations we propose can be more informative about distortions due to spatial correlation than tests based on the statistical significance of regressions using pre-treatment outcomes.
Finally, we analyze the use of design-based simulations in three shift-share design empirical applications. Our results reinforce the central insight of Adao that standard inference methods should be used with caution in shift-share applications, given the possibility of spatial correlation. At the same time, we show that commonly used design-based simulations tend to overstate the magnitude of size distortions attributable to spatial correlation, because they confound true treatment effects with error dependence. As a result, applied researchers who rely on standard design-based simulations in their own empirical applications may be led to misleading conclusions about the relative performance of alternative inference methods. We illustrate how the alternative simulation designs we propose avoid these issues and provide more informative guidance for inference in shift-share designs.
We consider first a simple setting of randomized experiments, in which we can illustrate the fundamental problem that generally prevents design-based simulations from recovering the true DGP in an empirical application, even when the distribution of treatment allocation is known.
Suppose we observe $i=1,...,N$ individuals whose potential outcomes are given by $Y_i(0)$ and $Y_i(1)$, and the only source of uncertainty comes from the treatment allocation $\mathbf{T} = (T_1,...,T_N)$, which is assumed to have a known distribution Finite_pop,NBERw24003,fisher,Neyman1990. The target parameter is the sample average treatment effect (SATE), $\frac{1}{N} \sum_{i=1}^N (Y_i(1) - Y_i(0))$.\footnote{Finite_pop and NBERw24003 also consider the possibility of sampling uncertainty. In this case, an alternative target parameter would be the population average treatment effect. For simplicity, we abstract from that.} In this case, inference using robust standard errors is asymptotically valid for the SATE in this sampling framework when $N \rightarrow \infty$ (although in some cases it may be conservative). Still, it might be that an econometrician or an applied researcher wants to study the properties of such inference method in a specific empirical application with a fixed $N$.
In this setting, the true DGP is characterized by $\{Y_i(0),Y_i(1) \}_{i=1}^N$ and the distribution of $\mathbf{T}$. Therefore, if we knew $Y_i(0)$ and $Y_i(1)$ for all $i$, then knowledge of the distribution of $\mathbf{T}$ would provide all information needed to recover the true probability of rejection for a test based on, for example, robust standard errors. More specifically, note that the distribution of the t-statistic with robust standard errors will be a function of $\{Y_i(0),Y_i(1)\}_{i=1}^N$ and $\mathbf{T}$, where the only stochastic component in this design-based setting is $\mathbf{T}$. Therefore, we would be able to calculate the exact probability that the t-statistic would be greater than a given critical value and, thus, the probability of rejection.\footnote{Since the true SATE may differ from zero, calculating the size of the test in this hypothetical scenario would require testing the null $H_0: SATE = \frac{1}{N}\sum_{i=1}^N \left(Y_i(1)-Y_i(0)\right)$, that is, the true SATE. Of course, since $\frac{1}{N}\sum_{i=1}^N \left(Y_i(1)-Y_i(0)\right)$ is generally unknown, this is not feasible in practice.}
However, we do not know $Y_i(0)$ and $Y_i(1)$ for all $i$. In this case, a commonly used way to construct design-based simulations in this setting is to hold the vector of realized outcomes fixed ($\mathbf{Y}$), and consider different allocations of $\mathbf{T}$ coming from the known distribution of $\mathbf{T}$. Then, for each allocation of $\mathbf{T}$ we estimate the parameter of interest and test the null using the inference method being analyzed. This design-based simulation implicitly considers a DGP in which potential outcomes are given by $\{ \widetilde Y_i(0),\widetilde Y_i(1) \}_{i=1}^N$, where $\widetilde Y_i(1) = \widetilde Y_i(0) = Y_i$, and the true distribution of $\mathbf{T}$. Therefore, in this DGP the null hypothesis that the SATE equals zero is true, because $\frac{1}{N} \sum_{i=1}^N (\widetilde Y_i(1) - \widetilde Y_i(0))=0$, and the design-based simulations gives us the size of this test if potential outcomes were given by $\{ \widetilde Y_i(0),\widetilde Y_i(1) \}_{i=1}^N$.
The fundamental challenge with this approach is that the true DGP would not be given by $\{ \widetilde Y_i(0),\widetilde Y_i(1) \}_{i=1}^N$ if $Y_i(1) \neq Y_i(0)$ for some $i$. Hence, even if we know the distribution of $\mathbf{T}$, a design-based simulation holding $\mathbf{Y}$ fixed would not necessarily recover the true finite-sample rejection rate of a given inference method or the true coverage for a confidence interval. It is therefore crucial to understand how the DGP underlying such design-based simulations may differ from the true DGP along dimensions that are relevant for finite-sample rejection rates of the inference method under consideration. For example, suppose that $\{Y_i(0)\}_{i=1}^N$ and $\{Y_i(1)\}_{i=1}^N$ are symmetric and single-peaked in their empirical distribution, with $Y_i(1)=Y_i(0)+\beta$ for all $i$ and some constant $\beta>0$. In this case, a design-based simulation that fixes $\mathbf Y$ and sets $\widetilde Y_i(1)=\widetilde Y_i(0)=Y_i$ would instead consider a DGP in which the distributions $\{\widetilde{Y}_i(0)\}_{i=1}^N$ and $\{\widetilde{Y}_i(1)\}_{i=1}^N$ may be dual peaked, thus differing fundamentally from the true DGP.
In Section (ref), we analyze in detail how the fact that design-based simulations do not recover exactly the true DGP implies an important challenge for the use of design-based simulations in shift-share designs.
We consider now the use of design-based simulations in shift-share design settings. Shift-share designs are specifications that study the impact of a set of shocks on units differentially exposed to them. Consider a setting with $i=1,...,N$ regions that are subject to $f=1,...,F$ aggregate shocks $\mathcal{X}_f$. The shift-share variable is given by ${x}_i = \sum_{f=1}^F w_{if}\mathcal{X}_f$, where the shares $w_{if}$ reflect how shock $\mathcal{X}_f$ affects unit $i$. The regression model is then given by
where $\beta$ reflects the effect of $x_i$ on $y_i$. Let $\mathbf{y}$ be the $N \times 1$ vector of outcomes, and $\Omega$ be the $N\times F$ matrix with all shares for all regions.
A key insight from Adao is that, when testing the null $H_0: \beta=0$, an important challenge in this setting is that the error term $\epsilon_i$ may be spatially correlated. In particular, if regions with similar shares tend to have correlated errors, this spatial correlation leads to over-rejection if one relies on robust or cluster-robust standard errors. In order to illustrate this issue, Adao consider design-based simulations in which they hold $\mathbf{y}$ and $\Omega$ fixed, and draw realizations of $\{\tilde{\mathcal{X}}_f\}_{f=1}^F$ for a given distribution for the shocks. For each realization of the shocks in these simulations, they estimate the shift-share estimator from $y_i = \gamma_0 + \gamma \tilde x_i^b + \tilde \epsilon_i$, where $\tilde x_i^b = \sum_{f=1}^F w_{if}\tilde{\mathcal{X}}^b_f$, and test the null $\gamma=0$ using robust or cluster-robust standard errors. In these simulations, they find rejection rates much higher than the nominal size of the tests.
Adao and Borusyak propose interesting alternatives for inference taking into account the possibility of spatial correlation. They consider a design-based setting, where shocks $\mathcal{X}_f$ are stochastic, while potential outcomes and shares are conditioned on.\footnote{See GP for an alternative sampling framework for shift-share designs.} A simplified version of their potential outcomes model is given by $y_i(x) = y_i(0) + \beta x$, where we assume for simplicity linearity and treatment effect homogeneity across shocks. In this setting, the identification assumption is that $\mathbb{E}[\mathcal{X}_f | \mathcal{L}]=c$ for a constant $c$, where $\mathcal{L}$ includes potential outcomes and shares. The inference methods they developed are asymptotically valid in this sampling framework when (i) shocks are independent, (ii) the number of shocks goes to infinity, and (iii) the size of each shock becomes asymptotically negligible.\footnote{They also allow for clusters of shocks. In this case, the assumption is that shocks from different clusters are independent, the number of clusters of shocks goes to infinity, and the size of each cluster of shocks becomes asymptotically negligible.} The fact that they developed their theory in this framework does not necessarily mean that the focus in shift-share designs should be on inference conditional on $\mathcal{L}$. As discussed by Adao, the idea of conditioning on potential outcomes is so that they “can allow for any correlation structure of the regression residuals across regions.”
Consider now the shift-share design setting from Section (ref). We formalize what a design-based simulation holding $\mathbf{y}$ and $\Omega$ fixed would recover in this setting, depending on the true treatment effect $\beta$ and on the presence or absence of spatial correlation in the empirical application used to construct the DGP used in the simulations.
To this end, consider a simpler version of the shift-share design setting from Section (ref) in which observations $i=1,...,N$ are partitioned into equally-sized groups $\Lambda_1,...,\Lambda_F$, with $w_{if} = 1$ if $i \in \Lambda_f$, and $w_{if} = 0$ otherwise. Assume also that $\mathcal{X}_f \in \{0,1\}$ and $ \sum_{f=1}^F \mathcal{X}_f =F/2$. Potential outcomes are given by $y_i(0) = \beta_0 + \epsilon_i$ and $y_i(1) = \beta + y_i(0)$.\footnote{Given the structure of this shift-share design example, we only need to define the potential outcomes $y_i(x)$ for $x \in \{0,1\}$.} This particular shift-share design setting can be seen as a randomized experiment in which treatment is assigned at the group $\Lambda_f$ level. In this particular shift-share design setting, the standard error proposed by Adao would coincide (up to a degrees-of-freedom correction) with standard errors clustered at the group level. Our conclusions in this section can also be extrapolated for DiD designs, which was the setting analyzed by Bertrand04howmuch.
In this setting, we know that inference based on robust standard errors would be asymptotically valid when $N \rightarrow \infty$ if errors are independent across $i$. However, it would lead to over-rejection in case errors within partitions are positive correlated. Therefore, it is natural to think that one could run design-based simulations assessing the rejection rates when inference is based on robust standard errors, in order to assess whether spatial correlation leads to relevant size distortions in specific empirical applications. We analyze whether this is actually so.
Consider design-based simulations with random draws of $\widetilde{\mathcal{X}}_f \in \{0,1\}$ such that $\sum_{f=1}^F \widetilde{\mathcal{X}}_f =F/2$, while holding $\mathbf{y}$ fixed. More specifically, for each draw of $\widetilde{\mathcal{X}}_f \in \{0,1\}$, we run the regression $y_i = \gamma_0 + \gamma \tilde x_i + \tilde \epsilon_i$, where $\tilde x_i = \sum_{f=1}^F w_{if} \tilde{\mathcal{X}_f}$, yielding $\hat \gamma^b$. Then we test the null $\gamma=0$ using robust standard errors at a significance level $\alpha$, and calculate the rejection rate in these simulations. The DGP in the simulations sets potential outcomes as $\tilde y_i(0) = \tilde y_i(1) = y_i$ (which are fixed, given the sampling framework of the simulations) and the distribution for $\widetilde{\mathcal{X}}_f $ described above. Therefore, the null hypothesis $H_0:\gamma=0$ is true in this DGP.
Uncertainty in these simulations comes only from realizations of $ \widetilde{\mathcal{X}}_f$. Let $\mathbb{E}^\ast[.| \mathbf{y}]$ and $\mathbb{V}^\ast(. | \mathbf{y})$ denote the expectation and variance operators with respect to this measure, conditional on $\mathbf{y}$. From Lemma 5 from IK, $\mathbb{E}^\ast [\hat \gamma^b | \mathbf{y} ] =0$, so the estimator $\hat \gamma^b$ is unbiased. Let $\mathbb{V}^\ast_{\mbox{\tiny true}} \equiv \mathbb{V}^\ast\left(\hat \gamma^b | \mathbf{y} \right)$ be the true variance of $\hat \gamma^b$ in these design-based simulations. Note that this is a number for a fixed $\mathbf{y}$, and a random variable depending on the errors $\epsilon_i$ when $\mathbf{y}$ is treated as a random variable. Also, let $\mathbb{V}^\ast_{\mbox{\tiny robust}}$ be the true variance in case treatment were assigned at the individual level in the design-based simulation. This is what the robust standard errors would asymptotically recover in these simulations when $F \rightarrow \infty$.
Therefore, the ratio ${\mathbb{V}^\ast_{\mbox{\tiny robust}}}/{\mathbb{V}^\ast_{\mbox{\tiny true}}}$ is crucial to understand whether the design-based simulations would indicate or not relevant over-rejections when we assess inference using robust standard errors. Considering this framework with $ \mathbf{y}$ fixed, if ${\mathbb{V}^\ast_{\mbox{\tiny robust}}}/{\mathbb{V}^\ast_{\mbox{\tiny true}}}$ converges to a value smaller than one in a sequence with $F \rightarrow \infty$, this means that the robust variance estimator would tend to underestimate the true variance in the simulations, generating a rejection rate larger than $\alpha$. In contrast, if ${\mathbb{V}^\ast_{\mbox{\tiny robust}}}/{\mathbb{V}^\ast_{\mbox{\tiny true}}}$ converges to one in a sequence with $F \rightarrow \infty$, then the robust variance calculated in the simulations would adequately estimate the true variance in the simulations, so we should expect rejection rates in the design-based simulations close to $\alpha$.
In this setting, we have the following result.
First, consider a setting with no spatial correlation, so $\rho =0$. This is a setting in which robust standard errors would be asymptotically valid when $F \rightarrow \infty$. However, Proposition (ref) shows that, with probability one, ${\mathbb{V}^\ast_{\mbox{\tiny robust}}}/{\mathbb{V}^\ast_{\mbox{\tiny true}}} \rightarrow k<1$, whenever the true DGP exhibits a true treatment effect ($\beta \neq 0$) and $m>1$. Therefore, the design-based assessment would tend to be larger than $\alpha$, (incorrectly) suggesting that robust standard errors are invalid due to spatial correlation. The intuition is that the true treatment effect $\beta$ is confounded with spatial correlation that affects the $m>1$ units within the same partition in a similar way in the DGP used in the simulations. The main problem is that the DGP used in the simulations would differ from the true DGP in a way that is relevant in determining the properties of a specific inference method.
Consider now a setting in which we should expect $\beta=0$. In this case, with probability one, ${\mathbb{V}^\ast_{\mbox{\tiny robust}}}/{\mathbb{V}^\ast_{\mbox{\tiny true}}} \rightarrow 1$ when $\rho =0$. Therefore, if we are in a setting with no spatial correlation (and $F$ is large), rejection rates in this design-based simulation would be close to $\alpha$, (correctly) suggesting that robust standard errors would be asymptotically valid. In contrast, if we are in a setting with $\rho >0$, then the rejection rates in the design-based simulations would be larger than $\alpha$, (correctly) suggesting that robust standard errors are invalid due to spatial correlation. Therefore, even though these simulations hold $\mathbf{y}$ fixed and consider only covariates as stochastic, this formalization also clarifies that, considering a setting with no treatment effect ($\beta=0$), these simulations are informative about potential problems in the errors, such as spatial correlation. A direct consequence is that, if we want to assess the prevalence of over-rejection due to spatial correlation in a specific setting, then one possibility would be to construct simulations based on a dataset in which we expect the true treatment effect to be zero. If we are analyzing a specific shift-share design empirical application, then one alternative would be to consider simulations using pre-treatment outcomes, so that we should expect similar patterns in terms of spatial correlation of the potential outcomes, but no treatment effect.
In case pre-treatment data is not available, another alternative to assess the relevance of spatial correlation when $\beta \neq 0$ is to consider $\boldsymbol{\epsilon}$-fixed (rather than $\mathbf{y}$-fixed) design-based simulations. This is how Borusyak construct their simulations in their Appendix A.11. In this case, we consider a DGP in which we fix potential outcomes as $\dot y_i(0) = \dot y_i(1) = y_i - \hat \beta x_i$. If potential outcomes in the true model are given by $y_i(x) = y_i +\beta x$ and $\hat \beta$ converges almost surely to $\beta$ when $F\rightarrow \infty$, then we would have simulations with a structure for the potential outcomes that is more similar to the true structure in the application.\footnote{In the stylized setting considered in this setting, the shift-share/OLS estimator $\hat{\beta}$ converges almost surely to $\beta$ as $F \rightarrow \infty$, even allowing for within-group error correlation. Our argument, however, applies to any estimator $\hat{\beta}$ satisfying $\hat{\beta} \xrightarrow{a.s.} \beta$ under the asymptotic sequence considered, and we therefore maintain this as a high-level assumption.} Therefore, the $\boldsymbol{\epsilon}$-fixed design-based simulations are informative about the presence of spatial correlation (when $F \rightarrow \infty$), even when we have a true treatment effect $\beta \neq 0$. We formalize this result in Appendix (ref). Interestingly, if treatment effects are heterogeneous or non-linear, then treatment effects would still generate spatially correlated errors in the design-based DGP, even after subtracting $\hat \beta x_i$ to construct the potential outcomes for the simulations. Therefore, these simulations would still be informative about problems with robust standard errors in shift-share design regressions due to heterogeneous or non-linear treatment effects.
Appendix (ref) presents simulations illustrating the use of design-based simulations to detect spatial correlation. We consider DGPs that follow the structure of the shift-share design setting analyzed in Section (ref). The simulations confirm the novel insights presented in Section (ref): (i) that $\mathbf{y}$-fixed design-based simulations may incorrectly indicate spatial correlation when the true treatment effect is different from zero; (ii) that $\boldsymbol{\epsilon}$-fixed design-based simulations circumvent this problem; (iii) that $\mathbf{y}$-fixed design-based simulations can be informative about spatial correlation when $\beta = 0$ (for example, when simulations are based on a placebo specification); and (iv) that $\boldsymbol{\epsilon}$-fixed design-based simulations are informative about inference problems when treatment effects are heterogeneous.
Interestingly, these simulations also highlight an additional insight about the use of design-based simulations. While testing significance in a pre-treatment outcome specification can be informative not only about the possibility of pre-trends but also about spatial correlation Ferman_DID, design-based simulations based on pre-treatment outcomes can be even more informative about spatial correlation problems. The intuition is the following. Consider the stylized model from Section (ref), and assume there are group-level iid shocks that take values $-\omega$ or $+\omega$ with equal probabilities. Let $\tau_1$ ($\tau_0$) denote the proportion of positive shocks among treated (control) groups. Ignoring idiosyncratic shocks, the pre-treatment outcome regression rejects the null when $|\tau_1-\tau_0|$ is sufficiently large, as this generates a large point estimate. Thus, rejection requires not only variability in the signs of the shocks, but also an unbalanced allocation of those shocks between treated and control groups. In contrast, design-based simulations analyze rejection rates across all possible treatment assignments rather than only the realized one. As a result, the simulations can flag spatial correlation problems whenever there is sufficient variation in the shocks—so that some treatment assignments would generate large values of $|\tau_1-\tau_0|$—even if $|\tau_1-\tau_0|$ is small for the realized treatment assignment.
We illustrate the use of design-based simulations in three shift-share design empirical applications, based on Autor, Dix, and Acemoglu.
We consider first the use design-based simulations to assess the performance of cluster-robust standard errors in these applications. We know that cluster-robust standard errors might be problematic in this setting if (i) there is spatial correlation in the errors (beyond the clusters), and/or (ii) there are few clusters. Ferman_assessment show that, in the absence of spatial correlation, asymptotic approximations for cluster-robust standard errors are relatively less reliable for the weighted OLS specifications in these applications. Since the focus in this section is to understand how different design-based simulations may be used to detect spatial-correlation problems, we therefore focus on the unweighted specifications. We emphasize, though, that those simulations can potentially be used to detect both of these problems.\footnote{In order to distinguish between these two problems, one could also run simulations with both errors and shocks stochastic. In this case, we would consider a DGP in which errors are independent, so these simulations would be informative about the first problem, but not about the second one.}
We consider first design-based simulations with $\mathbf{y}$ fixed and stochastic shocks. We find large rejection rates in columns 1, 3 and 5 of Table (ref), ranging from 34% to 70%. However, as discussed in Section (ref), a true treatment effect of ${x}_i$ may be confounded with spatially-correlated errors in these simulations. Therefore, we could find large rejection rates in these simulations, even when errors are not spatially correlated.
We consider the two alternatives discussed in Section (ref) to assess whether spatial correlation is a relevant concern for cluster-robust standard errors in these applications. All results are presented in Table (ref). For the application in Autor, our results reinforce the conclusions of Adao that spatial correlation leads to substantial over-rejection in this setting. At the same time, the over-rejection we find with these alternative design-based simulations is slightly less severe than that obtained with $\mathbf{y}$-fixed design-based simulations. This is consistent with the theoretical results presented in Section (ref), where we show that $\mathbf{y}$-fixed design-based simulations tend to overstate the magnitude of spatial correlation when there is a true treatment effect.
For Acemoglu, the $\boldsymbol{\epsilon}$-fixed design-based simulations remains large, but the design-based simulations using a placebo outcome is close to 5%. This provides evidence that inference based on cluster-robust standard errors might be reasonable for testing a sharp null of no effect whatsoever, but that we may have relevant heterogeneous treatment effects. As we discuss in Appendix (ref), the simulations using the pre-treatment outcome would have a relatively high probability of flagging problems due to spatial correlation in this application.
Finally, for Dix, both alternative simulations lead to rejection rates smaller than 5%, providing some indication that spatial correlation is not a problem in this application. An important caveat, however, is that those simulations would have a relatively lower probability of flagging a problem in case there is relevant spatial correlation in this application (details in Appendix (ref)). An alternative in cases like that may be to run simulations for a number of pre-treatment outcomes, if available.
We can also use design-based simulations to assess whether the new inference methods proposed for shift-share designs provide good asymptotic approximations in these settings, given the structure of shares. We find evidence that these inference methods work well in Autor. However, for the other two applications the design-based simulations suggest that these inference methods can lead to large distortions, with rejection rates up to 57% (Table (ref), Panel C). This is consistent with these applications having a smaller number of sectors (details in Table (ref)). Note that it is not a problem to consider $\mathbf{y}$-fixed design-based simulations in this case, because we are evaluating inference methods that allow for spatial correlation. Therefore, even if a true treatment effect is confounded with spatial correlation in the simulations, this would not be a problem in this case.
The results in this section illustrate that it is not trivial to determine which inference methods are more reliable in specific shift-share design applications, and that design-based simulations may be used by applied researchers to inform on that.
If we have evidence that the methods from Adao/Borusyak provide accurate asymptotic approximations in a given application, then they should be preferred, as they impose fewer restrictions on the spatial correlation. Therefore, our recommendation is to start with design-based simulations to assess whether inference based on their methods performs well in the application at hand, as we illustrate in Section (ref). This is the case for the application from Autor.
In case the design-based simulations suggest these new inference methods are unreliable, then the next step should be to use one of the alternative design-based simulations we propose (design-based simulations on a placebo outcome and/or $\boldsymbol{\epsilon}$-fixed design-based simulations) to evaluate whether inference based on cluster-robust standard errors would be reliable. These design-based simulations would be informative about two potential problems with cluster-robust standard errors in this setting: (i) in case the number of clusters is not large enough, and (ii) in case it is unreasonable to assume that there is no relevant spatial correlation. In our empirical illustrations, we find evidence that inference based on cluster-robust standard errors is relatively more reliable than inference based on these new methods for the applications from Acemoglu (for testing a sharp null), and from Dix (with the caveat that the design-based simulations have lower probabilities of detecting problems in this application).
The placebo specifications in these two applications provide further evidence on the conclusion that, in these applications, cluster-robust standard errors are relatively more reliable: in both cases, we would reject the null for their placebo specifications if we consider Adao standard errors, while we do not reject these nulls if we use cluster-robust standard errors (Appendix Table (ref), columns 4 and 6). Importantly, if applied researchers consider standard (fixed-$\mathbf{y}$) design-based simulations resampling shocks in these last two applications (instead of the alternatives we propose), they might incorrectly conclude that cluster-robust standard errors are less reliable than the new inference methods in these two applications.
In case we have evidence that none of these methods are reliable, then one would have to consider other alternatives. For example, Borusyak2 and Alvarez consider the use of randomization inference tests for shift-share designs, and show conditions in which these tests are exact in finite samples even when we allow for unrestricted spatial correlation. However, these tests rely on relatively strong assumptions on the shock assignment mechanism (such as correct specification of the distribution of shocks or exchangeability) for this finite-sample validity. Applied researcher would then have to analyze whether those are reasonable assumptions in their settings.
Design-based simulations are widely used in both methodological and applied work to study the finite-sample properties of inference procedures and to evaluate their performance in specific empirical applications. This paper shows that the interpretation of such simulations requires care. Because they construct a DGP by holding realized outcomes fixed, design-based simulations may fail to replicate key features of the true DGP—particularly when treatment effects are present or heterogeneous. In shift-share designs, we show that standard $\mathbf{y}$-fixed simulations can confound treatment effects with spatial correlation, potentially overstating inference distortions.
At the same time, design-based simulations can be informative when properly constructed. We clarify the conditions under which they reveal meaningful features of the error structure and propose alternative simulation designs that better isolate the relevant sources of distortion. Our empirical illustrations demonstrate that the choice of simulation design can materially affect conclusions about the reliability of competing inference methods.
Overall, the usefulness of design-based simulations depends on how closely the simulated DGP aligns with the true one along dimensions that matter for inference. Careful construction and interpretation are therefore essential.
\singlespace