EconBase
← Back to paper

Randomization Inference Tests for Shift-Share Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

32,680 characters · 7 sections · 16 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Randomization Inference Tests for Shift-Share Designs

\def\spacingset#1{ {#1}} \spacingset{1}

\newsavebox{\tablebox} \newlength{\tableboxwidth}

center[center omitted — 75 chars of source]

We consider the problem of inference in shift-share research designs. The choice between existing approaches that allow for unrestricted spatial correlation involves tradeoffs, varying in terms of their validity when there are relatively few or concentrated shocks, and in terms of the assumptions on the shock assignment process and treatment effects heterogeneity. We propose alternative randomization inference methods that combine the advantages of different approaches. These methods are valid in finite samples under relatively stronger assumptions, while asymptotically valid under weaker assumptions.

\

{\it Keywords:} shift-share designs; inference; spatial correlation;

\

{\it JEL Codes: C18,C21,C26.}

\spacingset{1.45}

\onehalfspacing

Introduction

Shift-share research designs consider instrumental variables that are constructed based on a common set of shocks that differentially affect different regions, depending on their exposure to those shocks. Prominent examples of papers that used this methodology include RePEc:upj:ubooks:wbsle, RePEc:bin:bpeajo:v:23:y:1992:i:1992-1:p:1-76, RePEc:ucp:jlabec:v:19:y:2001:i:1:p:22-64, and Autor.

AKM (henceforth, AKM) show that usual standard error formulas may substantially over-state the true variability of the shift-share estimator, if regions with similar exposures to the sector-level shocks also have correlated errors. AKM and BHJ (henceforth, BHJ) propose alternative estimators for the asymptotic variance of the shift-share estimator that are valid under arbitrary cross-regional correlation in the regression residuals. These methods rely on an asymptotic theory in which we have a large number of sectors, and the relevance of each sector becomes asymptotically negligible. While these methods provide reliable inference in many applications, they may lead to large over-rejection when such asymptotic theory does not provide a reasonable approximation to the empirical setting Ferman_assessment. BH (henceforth, BH) propose another alternative, based on the ideas of randomization inference (RI), that is valid even in finite samples. However, their approach relies on {assumptions on the shock assignment mechanism (such as, for example, knowledge of the distribution of the shocks, or that shocks are iid), which may be relatively harder to justify in some settings}. Given the advantages and disadvantages of each approach, BH state that “the choice between RI and asymptotic approaches involves tradeoffs.”

In this paper, we consider alternative inference methods based on RI that combine the advantages of the asymptotic methods proposed by AKM and BHJ, and of the RI method proposed by BH. The inference methods we propose are valid in finite samples under relatively stronger assumptions, including homogeneous treatment effects, and correct specification of the distribution of the shocks up to a scale parameter. We also consider alternatives that rely on the assumptions that the distribution of shocks is symmetric around a known mean or that shocks are iid, instead of assuming correct specification of the distribution of the shocks. At the same time, these inference methods are also asymptotically valid when the number of sectors increases under weaker assumptions on the treatment effects heterogeneity, and even when the distribution of the shocks is misspecified.\footnote{The RI tests we consider will be asymptotically conservative whenever the conditions stated by AKM in their Appendix A.1.6 hold. Those conditions limit the correlation between treatment effect heterogeneity and exposure weights.} Therefore, we provide inference methods for shift-share designs that are valid under relatively stronger assumptions in finite samples, but that we can relax those assumptions once the number of sectors increases. This eliminates the tradeoffs between RI and asymptotic approaches mentioned by BH.

Our approaches build on a large literature that studies the use of RI methods in other settings, and considers RI with studentized test statistics that are valid under stronger assumptions (or, alternatively, for inference on sharper null hypotheses) in finite samples, and also asymptotically valid under weaker assumptions (or, alternatively, for inference on less stringent null hypotheses). See, for example, Janssen1997, Chapter 15 of Lehmann2005, Chung2013, Bugni2018, Wu2020, Ferman_matching.

Main results

Setting

For an outcome of interest $Y$, we consider the structural model

equation[equation omitted — 60 chars of source]

where $x$ denotes a treatment of interest, and $\epsilon$ are the remaining determinants of $Y$. We consider for simplicity the case without a constant and without other covariates. However, all our results remain valid for a more general setting. For a sample of $N$ units, observed outcomes are given by

equation[equation omitted — 63 chars of source]

where $\{(X_i, \epsilon_i)\}_{i=1}^N$ are random variables. We have access to a shift-share instrument constructed as

equation[equation omitted — 45 chars of source]

where $\boldsymbol{s}_i \in \mathbb{R}^J_+$ are a set of exposures of $i$ and $g \in \mathbb{R}^J$ are a set of sector shocks. Let $\boldsymbol{S} = [\boldsymbol{s}_1,\boldsymbol{s}_2,\ldots \boldsymbol{s}_N]'$ and $\boldsymbol{\epsilon} = (\epsilon_1,\ldots, \epsilon_N)'$. We consider the following assumption on the distribution of the shocks, which is standard in the shift-share design literature (see AKM and BHJ).

assumption[Shock exogeneity] $\mathbb{E}[g|\boldsymbol{S},\boldsymbol{\epsilon}] = 0$.

Assumption (ref) imposes that shocks are mean-independent from unobserved determinants, conditional on exposures, with a common mean. More generally, we could have assumed that there exists $\mu \in \mathbb{R}$ such that $\mathbb{E}[g_j|\boldsymbol{S},\boldsymbol{\epsilon}] = \mu$ for every $j=1,\ldots, J$. In this case, the researcher may conduct inference by working with statistics that depend on demeaned shocks $\tilde{g}_j = g_j -\bar{g}$ (e.g. the shift-share estimator constructed with demeaned shocks). This is the solution proposed by BH to inference in linear shift-share designs where the sum of exposures may be uneven across units.\footnote{Our results remain valid in such setting. Specifically, valid finite sample inference under correct specification of the shock assignment mechanism (Proposition (ref)) would solely require that shocks are correctly specified up to a common location shift. Another alternative would be to control for the sum of the exposures, as proposed by BHJ.}

In our setting, the shift-share estimator is given by

equation[equation omitted — 260 chars of source]

where $\boldsymbol{X} = (X_1,\ldots, X_N)'$, $\boldsymbol{Y} = (Y_1,\ldots, Y_N)'$ and $\boldsymbol{Z} = (Z_1,\ldots, Z_N)'$. All our results remain valid if we consider the reduced-form case, in which case we set $X_i = Z_i$. Also, some of the results we present in Section (ref) are only valid for the reduced-form case.

AKM and BHJ show that, when shocks are independent and the importance of each sector becomes asymptotically negligible, then, as $J, N \to \infty$,

equation[equation omitted — 143 chars of source]

for every $c \in \mathbb{R}$, where $V_{SS}=\frac{\left(\sum_{i=1}^N \epsilon_i\boldsymbol{s}_i\right)' \mathbb{V}[g|\boldsymbol{S},\boldsymbol{\epsilon}]\left(\sum_{i=1}^N \epsilon_i\boldsymbol{s}_i\right)}{(\sum_{i=1}^N Z_i X_i)^2}$. This representation motivates the variance estimators proposed by AKM and BHJ. While their approaches provide reliable inference in empirical applications with many sectors, we may have relevant size distortions when there are few or concentrated sectors.

Randomization inference in shift-share designs

Given that the inference methods proposed by AKM and BHJ may not work well in some applications when there are few or concentrated sectors, we consider the use of RI in this setting. BH propose randomization-based inference in shock-based designs for a more general setting in which we have non-random exposure to exogenous shocks, where shift-share designs would be a particular example. Differently from BH, by focusing on shift-share design applications we are able to consider RI tests that are valid under relatively stronger assumptions in finite samples (similar to the approach proposed by BH), but that are also valid under weaker assumptions when the number of sectors increases. The main reason is that, in the setting we consider, the shift-share estimator has well-stablished asymptotic results (AKM, BHJ), so we are able to consider a studentized test statistic. In contrast, many of the settings considered by BH do not have well-stablished asymptotic results.

Finite-sample results

Suppose our goal is to test the null that $\beta = b$ against either a unilateral or bilateral alternative. Let $\hat{T} = \bar{T}(g, \boldsymbol{S}, \boldsymbol{X}, \boldsymbol{Y})$ be a test statistic, where large values of $\hat{T}$ constitute evidence against the null. We assume the following condition on the test statistic.

assumption{Under the null}, the map $\bar{T}$ satisfies $\bar{T}(g, \boldsymbol{S}, \boldsymbol{Y},\boldsymbol{X}) = T(g, \boldsymbol{S}, \boldsymbol{Y} - \boldsymbol{X}b)$ for some other map $T$.

That is, under the null, the test-statistic depends on $\boldsymbol{Y}$ and $\boldsymbol{X}$ solely through the “null-imposed residuals” $\boldsymbol{e}_b = \boldsymbol{Y} - \boldsymbol{X}b$.

Suppose now that the researcher has a guess on the shock assignment mechanism, i.e on the conditional probabilities $\mathbb{P}[g \leq v|\boldsymbol{S},\boldsymbol{\epsilon}]$ for each $v \in \mathbb{R}^J$. Let us denote such guess by a conditional distribution function $\mathbb{H}(\cdot|\boldsymbol{s},\boldsymbol{e})$ which specifies, for each $(\boldsymbol{s},\boldsymbol{e})$ in the support of $(\boldsymbol{S},\boldsymbol{\epsilon})$, a cummulative distribution function $\mathbb{H}(\cdot|\boldsymbol{s},\boldsymbol{e})$ on $\mathbb{R}^J$. In this case, the researcher is able to compute critical values by analysing the quantiles of

equation[equation omitted — 187 chars of source]

Such quantity is easily estimable by simulation. Indeed, if we are able to draw $L$ independent draws $g^*_l$, $l=1,\ldots, L$, from $\mathbb{H}[\cdot |\boldsymbol{S},\boldsymbol{e}_b]$, then $H_b(c|\boldsymbol{S},\boldsymbol{X},\boldsymbol{Y})$ may be estimated as

equation[equation omitted — 185 chars of source]

Clearly, if the shock-assignment process is correctly specified, then the procedure above provides valid inference.

propositionSuppose that Assumption (ref) holds, and that $\mathbb{P}[g \leq \cdot|\boldsymbol{S},\boldsymbol{\epsilon}] = \mathbb{H}[\cdot|\boldsymbol{S},\boldsymbol{\epsilon}]$. Then, under the null $\beta = b$, $$ H_b(c|\boldsymbol{S},\boldsymbol{X},\boldsymbol{Y}) = \mathbb{P}[\hat{T}\leq c|\boldsymbol{S},\boldsymbol{\epsilon}].$$ Consequently, a test that rejects the null if $\hat{T}$ exceeds the $1-\alpha$ quantile of $H_b(\cdot|\boldsymbol{S},\boldsymbol{X},\boldsymbol{Y})$ is (conditional on $(\boldsymbol{S},\boldsymbol{X},\boldsymbol{Y})$) level $\alpha$, where $\alpha \in (0,1)$. Similarly, a test that rejects the null if $\hat{T}$ exceeds the $1 +\frac{2}{L+1} -\frac{\lceil\alpha (L+1)\rceil}{L+1}$ quantile of $\hat{H}_b$ is conditionally level $\alpha$.

The result presented above is similar to Proposition S.3 from BH, with a couple of minor differences. First, we allow the shock assignment mechanism to depend on the non-observables. This way, we allow for the distribution of $S$ to depend on $\boldsymbol{\epsilon}$ (though in practice we expect that applied researchers would rarely choose a $\mathbb{H}[\cdot|\boldsymbol{S},\boldsymbol{e}_b]$ that depends on $\boldsymbol{e}_b$). Also, we provide level guarantees with a finite number of simulations. We present details of the proof in Appendix (ref).

In their paper, BH consider basing inference on the test statistic

equation[equation omitted — 153 chars of source]

which depends on $(\boldsymbol{X},\boldsymbol{Y})$ solely through $\boldsymbol{e}_b$, so Assumption (ref) holds.

In contrast, we consider the test-statistic

equation[equation omitted — 140 chars of source]

where $\hat{V}_{N,b}$ is the null-imposed variance estimator from {AKM and BHJ,\footnote{In a setting without an intercept or controls, BHJ show their variance estimator collapses to AKM's.}}

$$\hat{V}_{N,b} = \frac{\sum_{j=1}^J\left(\sum_{i=1}^N (Y_i - bX_i)\boldsymbol{s}_{ij}\right)^2 g_j^2}{(\sum_{i=1}^N Z_i X_i)^2}.$$

Observe that, under the null, the above map satisfies

equation[equation omitted — 268 chars of source]

so this test statistic also satisfies Assumption (ref). This corresponds to a rescaled version of the test statistic in BH.

Inference may be conducted as follows:

\paragraph{Algorithm 1}

enumerate• For simulations $l=1,\ldots,L$: \begin{enumerate} • Draw $g^*_l \sim \mathbb{H}[\cdot|\boldsymbol{S},\boldsymbol{e}_b]$. • Construct simulated instruments, $Z^*_{il} = \boldsymbol{s}_i'g^*_l$. • Run the shift-share IV estimator and null-imposed standard errors using the original data $Y_i$ and $X_i$ with artificial instruments $Z_{il}^*$. Construct the $t$-test based on the obtained values. \end{enumerate} • Reject the null if the observed test statistic is at the tails of the simulated distribution.

We adopt null-imposed standard errors in the test statistic $\hat T_1$, because otherwise it would depend on the endogenous regressor $X_i$, and we do not model assignment of these. If, however, $X_i = Z_i$, such problem disappears, and we may consider the test statistic

equation[equation omitted — 127 chars of source]

where $\hat{V}_{F} = \frac{\sum_{j=1}^J\left(\sum_{i=1}^N (Y_i - \hat{\beta}_{SS}X_i)\boldsymbol{s}_{ij}\right)^2 g_j^2}{(\sum_{i=1}^N X_i^2)^2}$ is the AKM or BHJ standard errors without imposing the null. Under the null, such statistic may be written as

eqnarray[eqnarray omitted — 399 chars of source]

Therefore, when we are in the case in which $X_i = Z_i$, we have that this test statistic satisfies Assumption (ref). Note, however, that this assumption would not be satisfied for this test statistic if we considered the case in which $X_i \neq Z_i$.

In this case, inference may be conducted as follows:

\paragraph{Algorithm 2}

enumerate• For a given $b$, compute $\boldsymbol{e}_b = \mathbf{Y} - \mathbf{X} b$. • For simulations $l=1,\ldots,L$: \begin{enumerate} • Draw $g^*_l \sim \mathbb{H}[\cdot|\boldsymbol{S},\boldsymbol{e}_b]$. • Construct data $\boldsymbol{Y}_l^* = b \boldsymbol{S} g^*_l + \boldsymbol{e}_b$. • Run the shift-share regression and shock-robust standard errors using the artificial data $\boldsymbol{Y}_l^*$, $g_l^*$ and $\boldsymbol{S}$. Construct the $t$-test based on the obtained values. \end{enumerate} • Reject the null if the observed test statistic is at the tails of the simulated distribution.

Following Proposition (ref), this inference procedure would be valid for settings in which $X_i = Z_i$.

remark[Scale-invariance of $\bar{T}_1$ and $\bar{T}_2$] \normalfont We observe that, when inference is based on the test-statistics $\hat{T}_1$ or $\hat{T}_2$, the requirement in Proposition (ref) may be weakened to: the distribution of shocks $g$ is correctly specified, up to multiplication of $g$ by a positive scalar. Indeed, test statistics $\hat{T}_1$ and $\hat{T}_2$ are invariant to multiplication of the shocks by a common positive constant. This contrasts with the test statistic $\hat{T}_0$, which requires the researcher to correctly specify the scale of shocks.
remark[Group transformations] \normalfont Instead of assuming that the shock-assignment mechanism is known, an alternative would be to consider a group $\boldsymbol{H}$ of transformations on $\mathbb{R}^J$ such that, under the null, for any $h \in \boldsymbol{H}$, $\bar{T}(h(g), \boldsymbol{S}, \boldsymbol{\epsilon})|\boldsymbol{S},\boldsymbol{\epsilon} \overset{d}{=} \bar{T}(g, \boldsymbol{S}, \boldsymbol{\epsilon})|\boldsymbol{S},\boldsymbol{\epsilon}$. In these settings, it follows from well-established results on randomization tests Lehmann2005 that the procedure described in Proposition (ref) remains valid if simulated shocks are constructed as $g^* = \boldsymbol{h}(g)$, where $\boldsymbol{h} \sim \operatorname{Uniform}(\boldsymbol{H})$, independently from the data. For example, if, conditional on $(\boldsymbol{S}, \boldsymbol{\epsilon})$, shocks were assumed independently drawn from symmetric distributions with known common symmetry point $m$, then one could take the group of transformations to be recentred sign changes, i.e. $h(g) = \kappa \odot (g - m \cdot \iota_J) + m \cdot \iota_J$ for $\kappa \in \{-1,1\}^J$, where $\odot$ denotes entry-by-entry multiplication and $\iota_J$ is a $J$ dimensional vector of ones.\footnote{ If the symmetry point were estimated (for example, by using the sample mean as an estimator of $m$), then the simulation procedure would no longer retain finite sample validity. In this case, conservative inference could be conducted by computing p-values under different choices of $m$, as $m$ varies over a valid confidence set, and then taking the supremum and adding one minus the confidence of the confidence set to it Berger1994. See Proposition S6 in BH for details.} {BH consider this kind of simulations in their Appendix D4.} Similarly, if, conditional on $(\boldsymbol{S}, \boldsymbol{\epsilon})$, shocks were assumed to be iid, then one could take the group to be the set of permutations of a $J$-dimensional vector, as also discussed by BH.

Asymptotic results

In this section, we consider the asymptotic properties of the simulation-based approach. We consider the properties of the tests in a framework where the number of sectors, $J$, is large. The number of units, $N$, is (implicitly) indexed by $J$, and is also allowed to grow.

We adopt a finite population perspective and allow for treatment effects to vary by unit. Formally, for a given $J \in \mathbb{N}$, potential outcomes are given by

equation[equation omitted — 87 chars of source]

and observed outcomes are given by

equation[equation omitted — 99 chars of source]

whereas the instrument is given by $Z_{i,J} = \boldsymbol{s}_{i,J}'g_{J}$. We will treat the $\epsilon_{i,J}$, $\beta_{i,J}$ and $\boldsymbol{s}_{i,J}$ as nonrandom throughout -- the only source of randomness stems from the assignment of $g_J$ and the treatment $X_{i,J}$. In other words, we follow a “design-based” approach. In this setting, the shift-share identification assumption is written as follows.

assumption[Shock-exogeneity] $\mathbb{E}[g_J] = 0$.

We rewrite the outcome model as

equation[equation omitted — 73 chars of source]

where

equation[equation omitted — 129 chars of source]

and

equation[equation omitted — 61 chars of source]

We consider the goal of the researcher to be to conduct inference on $\beta_J$, an affine combination of individual treatment effects. Specifically, she would like to test the null that $\beta_J = b_J$, and for that she uses one of the procedures described in the previous section.

Following AKM and BHJ, we put $v_J = \sum_{j=1}^J|\sum_{i=1}^N {s}_{i,j,J}|^2$ and assume $v_J \to \infty$. In the next proposition, we provide conditions for (conditional) asymptotic normality of $T^*_1 = T_1(g^*_J, \boldsymbol{S}, \boldsymbol{Y} - b_J \boldsymbol{X})$, the test statistic ${T}_1$ constructed under simulated shocks $g^*_J$.

proposition[Asymptotic normality of $T_1^*$ statistic] Assume that, conditional on $g_J$ and $\boldsymbol{X}$, simulated shocks $g^*_{j,J}$ are drawn independently across $j$ from distributions (not necessarily identical) satisfying: \begin{enumerate} • $\frac{1}{\sqrt{v_J}}\sum_{j=1}^J \left( \sum_{i=1}^N [\epsilon_{i,J}+\eta_{i,J} + (\beta_J-b_J)X_{i,J}]\boldsymbol{s}_{i,j,J}\right)\mathbb{E}[g_{j,J}^*|g_J,\boldsymbol{X}]=o_p(1)$; • $\sum_{j=1}^J\frac{1}{v_J}\left( \sum_{i=1}^N [\epsilon_{i,J}+\eta_{i,J} + (\beta_J-b_J)X_{i,J}]\boldsymbol{s}_{i,j,J}\right)^2\mathbb{E}[g_{j,J}^{*2}|g_J,\boldsymbol{X}] \overset{p}{\to} \sigma^2_* $, where $\sigma^2_* > 0$; and • $\sum_{j=1}^J\frac{1}{v_J^2}\left( \sum_{i=1}^N [\epsilon_{i,J}+\eta_{i,J} + (\beta_J-b_J)X_{i,J}]\boldsymbol{s}_{i,j,J}\right)^4\mathbb{E}[g_{j,J}^{*4}|g_J,\boldsymbol{X}] \overset{p}{\to} 0 $. \end{enumerate} Then, the simulated distribution converges in distribution to a standard normal, in probability, i.e. $ \mathbb{P}[T_1^* \leq c|\boldsymbol{X},g_J]\overset{p}{\to} \Phi(c)$, for every $c \in \mathbb{R}$.

We present details of the proof in Appendix (ref). Proposition (ref) provides high-level conditions for conditional asymptotic normality of the simulated statistic. In Appendix (ref), we show that these conditions are satisfied for three examples of simulation distributions: (i) when we sample with replacement from the (recentered) empirical distribution of shocks; (ii) when we consider shocks iid $N(0,1)$, independently from $\mathbf{X}$ and $g_J$, and; (iii) when we consider sign-changes of observed shocks.\footnote{{In the Appendix, we consider sign changes without recentering shocks ($m=0$). We note, however, that convergence would hold for any choice of recentering parameter $m$, including the case in which it is misspecified, and the case in which $m$ is replaced by an estimator such as the sample mean of shocks; {provided we work with the shift-share estimator that uses demeaned shocks (as per footnote 2)}.}} The crucial point is that we consider a studentized test statistic, so the simulated test statistic is asymptotically $N(0,1)$. In contrast, if we considered alternative test statistics, such as $\hat T_0$, then we would not reach this conclusion. Studentizing the test statistic using robust standard errors would also generally not work.

As a byproduct of Proposition (ref), whenever inference based on $\hat{T}_1$ and normal critical values provides asymptotically conservative inference, the simulation-based approach will also lead to asymptotically conservative inference. We summarize this fact in the corollary below.

corollarySuppose that, under the null $ \beta_J = b_J$, there exists $ v\geq 1$ such that, for every $c \in \mathbb{R}$, $\mathbb{P}[\hat{T}_1 \leq c] \to \Phi(v\cdot c)$. Assume that the conditions in Proposition (ref) hold under the null. Then the simulation-based approach to inference of Algorithm 1 will be asymptotically conservative, in the sense that, under the null, the probability of rejecting the null converges to a number smaller than the nominal significance level.

When there is no treatment effect heterogeneity, it follows that, under the conditions in AKM and BHJ, $v=1$. These conditions include that shocks are independent, that the number of sectors increase, and that the relevance of sectors are asymptotically negligible (AKM and BHJ consider alternatives that relax the assumption that shocks are independent, and we discuss that in Remark (ref)). In this case, inference based on the simulation approach is asymptotically size $\alpha$. More generally, when there is treatment effect heterogeneity, AKM provide sufficient conditions for inference based on $\hat{T}_1$ and normal critical-values being conservative (i.e. $v \geq 1$). These conditions limit the correlation between treatment effect heterogeneity and exposure weights. In this case, our simulation-based approach will also lead to conservative inference. Notice that, in contrast to our finite sample results, which require homogeneous treatment effects, asymptotically our method may be able to provide conservative inference under treatment effect heterogeneity.

remark\normalfont We note that the statement of Proposition (ref) does not require the null to be true. Specifically, if the conditions in Proposition (ref) can be shown to be valid under a given sequence of alternatives,\footnote{See Appendix (ref) for sufficient conditions in our three examples of simulation distributions.} then it follows that the distribution of the simulated statistic converges to a standard normal along such sequence. In this case, the power of the null-imposed t-test and our simulation-based approach coincide asymptotically along this sequence.
remark\normalfont Suppose that instead of assuming that shocks are independent, we consider that we have clusters of shocks that are independent, but that there may be correlation between shocks within the same cluster. In this case, our results from Proposition (ref) and Corollary (ref) should remain valid if we studentized the test statistic using AKM and BHJ standard errors with clusters of shocks (with the null imposed), provided the number of clusters is large. We may also consider using a distribution for the simulated shocks that allows for correlation within clusters.

Next, we analyze the test statistic $\hat T_2$. In this case, since the shift-share estimator is being recomputed across samples and then used in the calculation of the standard error, we need to ensure that $\hat{\beta}_{SS}^*$, the simulated shift-share estimator from Algorithm 2, is consistent at a given rate. In addition to the assumptions in Proposition (ref), we require a “strong simulated shock” assumption that ensures that the variance of the simulated shift-share regressor does not vanish asymptotically; as well as conditions that ensure the estimation error of the standard error vanishes. We state these requirements in the proposition below:

propositionSuppose, in addition to the assumptions in Proposition (ref), that: (i) $\frac{1}{N} \sum_{i=1}^N (\boldsymbol{s}_{i,J}'g_J^*)^2 \overset{p}{\to} \pi^* > 0$. Then $\frac{N}{\sqrt{v_J}}(\hat{\beta}^*_{SS} -b_J) = O_P(1)$. Moreover, if we assume that: \begin{enumerate} • \begin{equation*} \begin{aligned} \frac{1}{v_J} \sum_{j=1}^J \Bigg(\sum_{i=1}^N \sum_{l=1}^N s_{i,j,J}s_{l,j,J}[(\boldsymbol{s}_l'g^*_{J})(\epsilon_{i,J}+\eta_{i,J} + (\beta_J-b_J)X_{i,J}) + \\ (\boldsymbol{s}_i'g^*_{J})(\epsilon_{l,J}+\eta_{l,J}+ (\beta_J-b_J)X_{i,J})]\Bigg) g_{j,J}^{*2} = o_p\left(\frac{N}{\sqrt{v_J}}\right); and \end{aligned} \end{equation*} • \begin{equation*} \frac{1}{v_J} \sum_{j=1}^J \Bigg(\sum_{i=1}^N \sum_{l=1}^N [ (\boldsymbol{s}_i'g^*_{J}) (\boldsymbol{s}_l'g^*_{J})s_{i,j,J}s_{l,j,J} ]\Bigg) g_{j,J}^{*2} = o_p\left(\frac{N^2}{{v}_J}\right), \end{equation*} \end{enumerate} we may then conclude that $\mathbb{P}[T_2^* \leq c|\boldsymbol{X},g_J] \overset{p}{\to} \Phi(c)$ for every $c \in \mathbb{R}$.

We present details of the proof in Appendix (ref). In Appendix (ref), we discuss assumptions (i)-(iii) of the proposition in the context of our three examples of simulation distributions.

corollarySuppose that, under the null $\beta_J=b_J$, there exists $ v\geq 1$ such that, for every $c \in \mathbb{R}$, $\mathbb{P}[\hat{T}_2 \leq c] \to \Phi(v\cdot c)$. Assume that the conditions in Proposition (ref) hold under the null. Then the simulation-based approach to inference provided by Algorithm 2 will be asymptotically conservative, in the sense that, under the null, the probability of rejecting the null converges to a number smaller than the nominal significance level.
remark\normalfont Remarks (ref) and (ref) also apply to Proposition (ref) and Corollary (ref) when we consider the test statistic $\hat T_2$.

Conclusions

We consider the problem of inference in shift-share research designs. There are two main existing approaches that allow for unrestricted spatial correlation. The RI approach is valid even with relatively few or concentrated shocks, but relies on relatively strong assumptions on the shock assignment process and on treatment effect heterogeneity. In contrast, the asymptotic approach relies on weaker assumptions on the shock assignment process and on treatment effect heterogeneity, but asymptotic approximations may be inaccurate in some applications.

We propose alternative RI methods that combine the advantages of both approaches. More specifically, the inference methods we propose are exact under relatively strong assumptions, and also asymptotically valid under weaker assumptions. The latter is achieved through studentization, which ensures convergence of the simulated distribution of the test-statistic to a standard normal under mild regularity conditions.

\singlespace