Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
32,680 characters · 7 sections · 16 citation commands
Randomization Inference Tests for Shift-Share Designs
\def\spacingset#1{ {#1}} \spacingset{1}
\newsavebox{\tablebox} \newlength{\tableboxwidth}
We consider the problem of inference in shift-share research designs. The choice between existing approaches that allow for unrestricted spatial correlation involves tradeoffs, varying in terms of their validity when there are relatively few or concentrated shocks, and in terms of the assumptions on the shock assignment process and treatment effects heterogeneity. We propose alternative randomization inference methods that combine the advantages of different approaches. These methods are valid in finite samples under relatively stronger assumptions, while asymptotically valid under weaker assumptions.
\
{\it Keywords:} shift-share designs; inference; spatial correlation;
\
{\it JEL Codes: C18,C21,C26.}
\spacingset{1.45}
\onehalfspacing
Shift-share research designs consider instrumental variables that are constructed based on a common set of shocks that differentially affect different regions, depending on their exposure to those shocks. Prominent examples of papers that used this methodology include RePEc:upj:ubooks:wbsle, RePEc:bin:bpeajo:v:23:y:1992:i:1992-1:p:1-76, RePEc:ucp:jlabec:v:19:y:2001:i:1:p:22-64, and Autor.
AKM (henceforth, AKM) show that usual standard error formulas may substantially over-state the true variability of the shift-share estimator, if regions with similar exposures to the sector-level shocks also have correlated errors. AKM and BHJ (henceforth, BHJ) propose alternative estimators for the asymptotic variance of the shift-share estimator that are valid under arbitrary cross-regional correlation in the regression residuals. These methods rely on an asymptotic theory in which we have a large number of sectors, and the relevance of each sector becomes asymptotically negligible. While these methods provide reliable inference in many applications, they may lead to large over-rejection when such asymptotic theory does not provide a reasonable approximation to the empirical setting Ferman_assessment. BH (henceforth, BH) propose another alternative, based on the ideas of randomization inference (RI), that is valid even in finite samples. However, their approach relies on {assumptions on the shock assignment mechanism (such as, for example, knowledge of the distribution of the shocks, or that shocks are iid), which may be relatively harder to justify in some settings}. Given the advantages and disadvantages of each approach, BH state that “the choice between RI and asymptotic approaches involves tradeoffs.”
In this paper, we consider alternative inference methods based on RI that combine the advantages of the asymptotic methods proposed by AKM and BHJ, and of the RI method proposed by BH. The inference methods we propose are valid in finite samples under relatively stronger assumptions, including homogeneous treatment effects, and correct specification of the distribution of the shocks up to a scale parameter. We also consider alternatives that rely on the assumptions that the distribution of shocks is symmetric around a known mean or that shocks are iid, instead of assuming correct specification of the distribution of the shocks. At the same time, these inference methods are also asymptotically valid when the number of sectors increases under weaker assumptions on the treatment effects heterogeneity, and even when the distribution of the shocks is misspecified.\footnote{The RI tests we consider will be asymptotically conservative whenever the conditions stated by AKM in their Appendix A.1.6 hold. Those conditions limit the correlation between treatment effect heterogeneity and exposure weights.} Therefore, we provide inference methods for shift-share designs that are valid under relatively stronger assumptions in finite samples, but that we can relax those assumptions once the number of sectors increases. This eliminates the tradeoffs between RI and asymptotic approaches mentioned by BH.
Our approaches build on a large literature that studies the use of RI methods in other settings, and considers RI with studentized test statistics that are valid under stronger assumptions (or, alternatively, for inference on sharper null hypotheses) in finite samples, and also asymptotically valid under weaker assumptions (or, alternatively, for inference on less stringent null hypotheses). See, for example, Janssen1997, Chapter 15 of Lehmann2005, Chung2013, Bugni2018, Wu2020, Ferman_matching.
For an outcome of interest $Y$, we consider the structural model
where $x$ denotes a treatment of interest, and $\epsilon$ are the remaining determinants of $Y$. We consider for simplicity the case without a constant and without other covariates. However, all our results remain valid for a more general setting. For a sample of $N$ units, observed outcomes are given by
where $\{(X_i, \epsilon_i)\}_{i=1}^N$ are random variables. We have access to a shift-share instrument constructed as
where $\boldsymbol{s}_i \in \mathbb{R}^J_+$ are a set of exposures of $i$ and $g \in \mathbb{R}^J$ are a set of sector shocks. Let $\boldsymbol{S} = [\boldsymbol{s}_1,\boldsymbol{s}_2,\ldots \boldsymbol{s}_N]'$ and $\boldsymbol{\epsilon} = (\epsilon_1,\ldots, \epsilon_N)'$. We consider the following assumption on the distribution of the shocks, which is standard in the shift-share design literature (see AKM and BHJ).
Assumption (ref) imposes that shocks are mean-independent from unobserved determinants, conditional on exposures, with a common mean. More generally, we could have assumed that there exists $\mu \in \mathbb{R}$ such that $\mathbb{E}[g_j|\boldsymbol{S},\boldsymbol{\epsilon}] = \mu$ for every $j=1,\ldots, J$. In this case, the researcher may conduct inference by working with statistics that depend on demeaned shocks $\tilde{g}_j = g_j -\bar{g}$ (e.g. the shift-share estimator constructed with demeaned shocks). This is the solution proposed by BH to inference in linear shift-share designs where the sum of exposures may be uneven across units.\footnote{Our results remain valid in such setting. Specifically, valid finite sample inference under correct specification of the shock assignment mechanism (Proposition (ref)) would solely require that shocks are correctly specified up to a common location shift. Another alternative would be to control for the sum of the exposures, as proposed by BHJ.}
In our setting, the shift-share estimator is given by
where $\boldsymbol{X} = (X_1,\ldots, X_N)'$, $\boldsymbol{Y} = (Y_1,\ldots, Y_N)'$ and $\boldsymbol{Z} = (Z_1,\ldots, Z_N)'$. All our results remain valid if we consider the reduced-form case, in which case we set $X_i = Z_i$. Also, some of the results we present in Section (ref) are only valid for the reduced-form case.
AKM and BHJ show that, when shocks are independent and the importance of each sector becomes asymptotically negligible, then, as $J, N \to \infty$,
for every $c \in \mathbb{R}$, where $V_{SS}=\frac{\left(\sum_{i=1}^N \epsilon_i\boldsymbol{s}_i\right)' \mathbb{V}[g|\boldsymbol{S},\boldsymbol{\epsilon}]\left(\sum_{i=1}^N \epsilon_i\boldsymbol{s}_i\right)}{(\sum_{i=1}^N Z_i X_i)^2}$. This representation motivates the variance estimators proposed by AKM and BHJ. While their approaches provide reliable inference in empirical applications with many sectors, we may have relevant size distortions when there are few or concentrated sectors.
Given that the inference methods proposed by AKM and BHJ may not work well in some applications when there are few or concentrated sectors, we consider the use of RI in this setting. BH propose randomization-based inference in shock-based designs for a more general setting in which we have non-random exposure to exogenous shocks, where shift-share designs would be a particular example. Differently from BH, by focusing on shift-share design applications we are able to consider RI tests that are valid under relatively stronger assumptions in finite samples (similar to the approach proposed by BH), but that are also valid under weaker assumptions when the number of sectors increases. The main reason is that, in the setting we consider, the shift-share estimator has well-stablished asymptotic results (AKM, BHJ), so we are able to consider a studentized test statistic. In contrast, many of the settings considered by BH do not have well-stablished asymptotic results.
Suppose our goal is to test the null that $\beta = b$ against either a unilateral or bilateral alternative. Let $\hat{T} = \bar{T}(g, \boldsymbol{S}, \boldsymbol{X}, \boldsymbol{Y})$ be a test statistic, where large values of $\hat{T}$ constitute evidence against the null. We assume the following condition on the test statistic.
That is, under the null, the test-statistic depends on $\boldsymbol{Y}$ and $\boldsymbol{X}$ solely through the “null-imposed residuals” $\boldsymbol{e}_b = \boldsymbol{Y} - \boldsymbol{X}b$.
Suppose now that the researcher has a guess on the shock assignment mechanism, i.e on the conditional probabilities $\mathbb{P}[g \leq v|\boldsymbol{S},\boldsymbol{\epsilon}]$ for each $v \in \mathbb{R}^J$. Let us denote such guess by a conditional distribution function $\mathbb{H}(\cdot|\boldsymbol{s},\boldsymbol{e})$ which specifies, for each $(\boldsymbol{s},\boldsymbol{e})$ in the support of $(\boldsymbol{S},\boldsymbol{\epsilon})$, a cummulative distribution function $\mathbb{H}(\cdot|\boldsymbol{s},\boldsymbol{e})$ on $\mathbb{R}^J$. In this case, the researcher is able to compute critical values by analysing the quantiles of
Such quantity is easily estimable by simulation. Indeed, if we are able to draw $L$ independent draws $g^*_l$, $l=1,\ldots, L$, from $\mathbb{H}[\cdot |\boldsymbol{S},\boldsymbol{e}_b]$, then $H_b(c|\boldsymbol{S},\boldsymbol{X},\boldsymbol{Y})$ may be estimated as
Clearly, if the shock-assignment process is correctly specified, then the procedure above provides valid inference.
The result presented above is similar to Proposition S.3 from BH, with a couple of minor differences. First, we allow the shock assignment mechanism to depend on the non-observables. This way, we allow for the distribution of $S$ to depend on $\boldsymbol{\epsilon}$ (though in practice we expect that applied researchers would rarely choose a $\mathbb{H}[\cdot|\boldsymbol{S},\boldsymbol{e}_b]$ that depends on $\boldsymbol{e}_b$). Also, we provide level guarantees with a finite number of simulations. We present details of the proof in Appendix (ref).
In their paper, BH consider basing inference on the test statistic
which depends on $(\boldsymbol{X},\boldsymbol{Y})$ solely through $\boldsymbol{e}_b$, so Assumption (ref) holds.
In contrast, we consider the test-statistic
where $\hat{V}_{N,b}$ is the null-imposed variance estimator from {AKM and BHJ,\footnote{In a setting without an intercept or controls, BHJ show their variance estimator collapses to AKM's.}}
$$\hat{V}_{N,b} = \frac{\sum_{j=1}^J\left(\sum_{i=1}^N (Y_i - bX_i)\boldsymbol{s}_{ij}\right)^2 g_j^2}{(\sum_{i=1}^N Z_i X_i)^2}.$$
Observe that, under the null, the above map satisfies
so this test statistic also satisfies Assumption (ref). This corresponds to a rescaled version of the test statistic in BH.
Inference may be conducted as follows:
\paragraph{Algorithm 1}
We adopt null-imposed standard errors in the test statistic $\hat T_1$, because otherwise it would depend on the endogenous regressor $X_i$, and we do not model assignment of these. If, however, $X_i = Z_i$, such problem disappears, and we may consider the test statistic
where $\hat{V}_{F} = \frac{\sum_{j=1}^J\left(\sum_{i=1}^N (Y_i - \hat{\beta}_{SS}X_i)\boldsymbol{s}_{ij}\right)^2 g_j^2}{(\sum_{i=1}^N X_i^2)^2}$ is the AKM or BHJ standard errors without imposing the null. Under the null, such statistic may be written as
Therefore, when we are in the case in which $X_i = Z_i$, we have that this test statistic satisfies Assumption (ref). Note, however, that this assumption would not be satisfied for this test statistic if we considered the case in which $X_i \neq Z_i$.
In this case, inference may be conducted as follows:
\paragraph{Algorithm 2}
Following Proposition (ref), this inference procedure would be valid for settings in which $X_i = Z_i$.
In this section, we consider the asymptotic properties of the simulation-based approach. We consider the properties of the tests in a framework where the number of sectors, $J$, is large. The number of units, $N$, is (implicitly) indexed by $J$, and is also allowed to grow.
We adopt a finite population perspective and allow for treatment effects to vary by unit. Formally, for a given $J \in \mathbb{N}$, potential outcomes are given by
and observed outcomes are given by
whereas the instrument is given by $Z_{i,J} = \boldsymbol{s}_{i,J}'g_{J}$. We will treat the $\epsilon_{i,J}$, $\beta_{i,J}$ and $\boldsymbol{s}_{i,J}$ as nonrandom throughout -- the only source of randomness stems from the assignment of $g_J$ and the treatment $X_{i,J}$. In other words, we follow a “design-based” approach. In this setting, the shift-share identification assumption is written as follows.
We rewrite the outcome model as
where
and
We consider the goal of the researcher to be to conduct inference on $\beta_J$, an affine combination of individual treatment effects. Specifically, she would like to test the null that $\beta_J = b_J$, and for that she uses one of the procedures described in the previous section.
Following AKM and BHJ, we put $v_J = \sum_{j=1}^J|\sum_{i=1}^N {s}_{i,j,J}|^2$ and assume $v_J \to \infty$. In the next proposition, we provide conditions for (conditional) asymptotic normality of $T^*_1 = T_1(g^*_J, \boldsymbol{S}, \boldsymbol{Y} - b_J \boldsymbol{X})$, the test statistic ${T}_1$ constructed under simulated shocks $g^*_J$.
We present details of the proof in Appendix (ref). Proposition (ref) provides high-level conditions for conditional asymptotic normality of the simulated statistic. In Appendix (ref), we show that these conditions are satisfied for three examples of simulation distributions: (i) when we sample with replacement from the (recentered) empirical distribution of shocks; (ii) when we consider shocks iid $N(0,1)$, independently from $\mathbf{X}$ and $g_J$, and; (iii) when we consider sign-changes of observed shocks.\footnote{{In the Appendix, we consider sign changes without recentering shocks ($m=0$). We note, however, that convergence would hold for any choice of recentering parameter $m$, including the case in which it is misspecified, and the case in which $m$ is replaced by an estimator such as the sample mean of shocks; {provided we work with the shift-share estimator that uses demeaned shocks (as per footnote 2)}.}} The crucial point is that we consider a studentized test statistic, so the simulated test statistic is asymptotically $N(0,1)$. In contrast, if we considered alternative test statistics, such as $\hat T_0$, then we would not reach this conclusion. Studentizing the test statistic using robust standard errors would also generally not work.
As a byproduct of Proposition (ref), whenever inference based on $\hat{T}_1$ and normal critical values provides asymptotically conservative inference, the simulation-based approach will also lead to asymptotically conservative inference. We summarize this fact in the corollary below.
When there is no treatment effect heterogeneity, it follows that, under the conditions in AKM and BHJ, $v=1$. These conditions include that shocks are independent, that the number of sectors increase, and that the relevance of sectors are asymptotically negligible (AKM and BHJ consider alternatives that relax the assumption that shocks are independent, and we discuss that in Remark (ref)). In this case, inference based on the simulation approach is asymptotically size $\alpha$. More generally, when there is treatment effect heterogeneity, AKM provide sufficient conditions for inference based on $\hat{T}_1$ and normal critical-values being conservative (i.e. $v \geq 1$). These conditions limit the correlation between treatment effect heterogeneity and exposure weights. In this case, our simulation-based approach will also lead to conservative inference. Notice that, in contrast to our finite sample results, which require homogeneous treatment effects, asymptotically our method may be able to provide conservative inference under treatment effect heterogeneity.
Next, we analyze the test statistic $\hat T_2$. In this case, since the shift-share estimator is being recomputed across samples and then used in the calculation of the standard error, we need to ensure that $\hat{\beta}_{SS}^*$, the simulated shift-share estimator from Algorithm 2, is consistent at a given rate. In addition to the assumptions in Proposition (ref), we require a “strong simulated shock” assumption that ensures that the variance of the simulated shift-share regressor does not vanish asymptotically; as well as conditions that ensure the estimation error of the standard error vanishes. We state these requirements in the proposition below:
We present details of the proof in Appendix (ref). In Appendix (ref), we discuss assumptions (i)-(iii) of the proposition in the context of our three examples of simulation distributions.
We consider the problem of inference in shift-share research designs. There are two main existing approaches that allow for unrestricted spatial correlation. The RI approach is valid even with relatively few or concentrated shocks, but relies on relatively strong assumptions on the shock assignment process and on treatment effect heterogeneity. In contrast, the asymptotic approach relies on weaker assumptions on the shock assignment process and on treatment effect heterogeneity, but asymptotic approximations may be inaccurate in some applications.
We propose alternative RI methods that combine the advantages of both approaches. More specifically, the inference methods we propose are exact under relatively strong assumptions, and also asymptotically valid under weaker assumptions. The latter is achieved through studentization, which ensures convergence of the simulated distribution of the test-statistic to a standard normal under mild regularity conditions.
\singlespace