Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
176,743 characters · 29 sections · 50 citation commands
SHIFT-SHARE DESIGNS: THEORY AND INFERENCE
\thispagestyle{empty} \setcounter{page}{0}
We study how to perform inference in shift-share designs: regression specifications in which one studies the impact of a set of shocks, or “shifters”, on units differentially exposed to them, with the exposure measured by a set of weights, or “shares”. Specifically, shift-share regressions have the form
For example, in an investigation of the impact of sectoral demand shifters on regional employment changes, $Y_{i}$ is the change in employment in region $i$, the shifter $\ensuremath\mathcal{X}_{s}$ is a measure of the change in demand for the good produced by sector $s$, and the share $w_{is}$ may be measured as the initial share of region $i$'s employment in sector $s$. Other observed characteristics of region $i$ are captured by the vector $Z_{i}$, which includes the intercept, and $\epsilon_{i}$ is the regression residual. Shift-share specifications are increasingly common in many contexts (see, e.g., bartik1991benefits, blanchard1992regional, card2001immigrant, or \citet*{autordornhanson2013}). However, their formal properties are relatively understudied.
Our starting point is the observation that usual standard error formulas may substantially understate the true variability of OLS estimators of $\beta$ in (ref). We illustrate the importance of this issue through a placebo exercise. As outcomes, we use 2000--2007 changes in employment rates and average wages for 722 Commuting Zones in the United States. We build a shift-share regressor by combining actual sectoral employment shares in 1990 with randomly drawn sector-level shifters for 396 4-digit SIC manufacturing sectors. The placebo samples thus differ exclusively in the randomly drawn sectoral shifters. For each sample, we compute the OLS estimate of $\beta$ in (ref) and test if its true value is zero. Since the shifters are randomly generated, their true effect is indeed zero. Valid 5% significance level tests should therefore reject the null of no effect in at most 5% of the placebo samples. We find, however, that usual standard errors---clustering on state as well as heteroskedasticity-robust errors---are much smaller than the standard deviation of the OLS estimator and, as a result, lead to severe overrejection. Depending on the labor market outcome used, the rejection rate for 5% level tests can be as high as 55% for heteroskedasticity-robust standard errors and 45% for standard errors clustered on state, and it is never below 16%.
To explain the source of this overrejection problem, we introduce a stylized economic model featuring multiple regions, each of which produces output in multiple sectors. The key ingredients of our model are a sector- and region-specific labor demand and a regional labor supply. We assume that labor demand in each sector-region pair has a sector-specific elasticity with respect to wages and an intercept that aggregates several sector-specific components (e.g.\ sectoral productivities and demand shifters for the corresponding sectoral good). Labor supply in each region is upward-sloping and has a region-specific intercept that may aggregate group-specific labor supply shifters (e.g.\ push factors that raise immigration from different countries of origin). Up to a first-order approximation, the impact of sector-level shocks on labor market outcomes takes the form of a shift-share specification similar to that in (ref).
A key insight of our model is that the regression residual $\epsilon_{i}$ in (ref) will generally account for shift-share components that aggregate all unobserved sector-level shocks using the same shares $w_{is}$ that enter the construction of the regressor $X_{i}$, as well as shift-share components that aggregate unobserved group-specific labor supply shifters using exposures $\tilde{w}_{ig}$ of region $i$ to group-$g$ specific shocks. Thus, the residual may incorporate multiple shift-share terms with shares correlated with those defining the shift-share regressor $X_{i}$. Consequently, whenever two regions have similar shares, they will not only have similar exposure to the shifters $\ensuremath\mathcal{X}_{s}$, but will also tend to have similar values of the residuals $\epsilon_{i}$. While traditional inference methods allow for some forms of dependence between the residuals, such as spatial dependence within a state, they do not directly address the possible dependence between residuals generated by unobserved shift-share components. This is why, in our placebo exercise, traditional inference methods underestimate the variance of the OLS estimator of $\beta$, creating the overrejection problem.
We then establish the large-sample properties of the OLS estimator of $\beta$ in (ref) under repeated sampling of the shifters $\ensuremath\mathcal{X}_s$, conditioning on the realized shares $w_{is}$, controls $Z_{i}$, and residuals $\epsilon_i$. This sampling approach is motivated by our economic model: we are interested in what would have happened to outcomes if the sector-level shocks $\ensuremath\mathcal{X}_s$ had taken different values, holding everything else constant. Our framework allows for heterogeneous effects of the shifters: one unit increase in $\ensuremath\mathcal{X}_{s}$ causes the outcome in region $i$ to increase by $w_{is}\beta_{is}$, where $\beta_{is}$ is an unknown parameter.
Our key assumption is that, conditional on the controls and the shares, the shifters are as good as randomly assigned and independent across sectors. An advantage of this assumption is that it allows us to do inference conditionally on $\epsilon_{i}$; as a result, we can allow for any correlation structure of the regression residuals across regions.\footnote{This is similar to the insight in barrios2012clustering, who consider cross-section regressions estimated at an individual level when the variable of interest varies only across groups of individuals. They show that, as long as the regressor of interest is as good as randomly assigned and independent across the groups, standard errors clustered on groups are valid under any correlation structure of the residuals.} In contrast, if, instead of assuming independence of the shifters across sectors, we modeled the correlation structure in the residual, as in the spatial econometrics literature conley_gmm_1999 or in the interactive fixed effects literature bai2009,gobillonmagnac2016, the resulting inference would be sensitive to the validity of the modeling assumptions. We show that the regression estimand $\beta$ in (ref) corresponds to a weighted average of the heterogeneous parameters $\beta_{is}$ and derive novel confidence intervals that are valid in samples with many regions and sectors. We also derive an analogous formula when $X_{i}$ is used as an instrument in an instrumental variables regression, which follows directly from the fact that the associated first-stage and reduced-form regressions take the form in (ref).
To gain intuition for our formula, it is useful to consider the special case in which each region is fully specialized in one sector (i.e.\ for every $i$, $w_{is}=1$ for some sector $s$). In this case, our procedure is identical to using the usual clustered standard error formula, but with clusters defined as groups of regions specialized in the same sector. This is in line with the rule of thumb that one should “cluster” at the level of variation of the regressor of interest. In the general case, our standard error formula essentially forms sectoral clusters, the variance of which depends on the variance of a weighted sum of the regression residuals $\epsilon_{i}$, with weights that correspond to the shares $w_{is}$.
We extend our baseline results in three ways. We provide versions of our standard errors that only require the shifters to be independent across “clusters” of sectors, allowing for arbitrary correlation among sectors belonging to the same “cluster.” We also show how to apply our framework to panel data settings in which we have multiple observations of each region over time. Finally, we cover applications in which the shifter is unobserved, but can be estimated using observable local shocks.
We illustrate the finite-sample properties of our novel inference procedure in the same placebo exercise that we use to show the bias of the usual standard error formulas. Our new formulas give a good approximation to the variability of the OLS estimator across the placebo samples; consequently, they yield rejection rates that are close to the nominal significance level. As predicted by the theory, our standard error formula remains accurate under alternative distributions of both the shifters and the regression residuals. When the number of sectors is small or there is a sector that is significantly larger than the rest, our method overrejects, although the overrejection is milder in comparison with the usual standard error formulas. If the shifters are not independent across sectors, we show that it is important to properly account for their correlation structure.
In the final part of the paper, we illustrate the implications of our new inference procedure for two popular applications of shift-share regressions. First, we study the effect of changes in sector-level Chinese import competition on labor market outcomes across U.S. Commuting Zones, as in \citet*{autordornhanson2013}. Second, we use changes in sector-level national employment to estimate the regional inverse labor supply elasticity, as in bartik1991benefits.\footnote{Additionally, in
, we use changes in the stock of immigrants from various origin countries to investigate the impact of immigration on employment and wages, following altonji1991effects and card2001immigrant.} Our new confidence intervals for the effects of Chinese competition on local labor markets increase by 23%--66% relative to those implied by state-clustered or heteroskedasticity-robust standard errors, although these effects remain statistically significant. In contrast, our confidence intervals for the inverse labor supply elasticity estimated using the procedure in bartik1991benefits are very similar to those constructed using standard approaches.
Shift-share designs have been applied to estimate the effect of a wide range of shocks. For example, in seminal papers, bartik1991benefits and blanchard1992regional use shift-share designs to analyze the impact on local labor markets of shifters measured as changes in national sectoral employment. More recently, shift-share strategies have been applied to investigate the local labor market impact of various shocks, including international trade competition \citep*{topalova2007trade,topalova2010factor, kovak2013regional,autordornhanson2013, dixcarneirokovak2017, pierce2017trade}, credit supply \citep*{Greenstone2015credit}, technological change acemoglu2017robots,acemoglurestrepo2018, and industry reallocation chodorow2018reallocation. Shift-share regressors have been used as well to estimate the impact of immigration on labor markets, as in card2001immigrant and many other papers following his approach; see reviews in lewisperi2015 and \citet*{dustmann2016different}. Furthermore, recent papers use shift-share strategies to estimate how firms respond to changes in outsourcing costs and foreign demand hummels2014wage, aghion2018impact.\footnote{Shift-share regressors have also been used to study the impact of sectoral shocks on political preferences autor2017elections, che2017tradechinaelections,colantonestanig2018, marriage patterns \citep*{autordornhanson2018}, crime levels \citep*{dixcarneiro2017crime}, and innovation acemoglulinn2004,autordornhanson2017innovation. In addition to using shift-share designs to estimate the overall impact of a shifter of interest, other work has used them as part of a more general structural estimation approach; see diamond2016, adao2016, \citet*{galle2017slicing}, bursteinhansontianvogel2018, bartelme2018tradecosts. baum2015causal review additional applications in the context of urban economics.}
Our paper is related to two other papers studying the statistical properties of shift-share instrumental variables. First, \citet*{goldsmith2018bartik} consider using the full vector of shares $(w_{i1}, \dotsc, w_{iS})$ as an instrument for endogenous treatment. They conclude that this approach requires the entire vector of shares to be as good as randomly assigned conditional on the shifters. Second, \citet*{borusyakhulljaravel2018shiftshare}, focusing on the use of a shift-share regressor as an instrument, show it is a valid instrument if the set of shifters is as good as randomly assigned conditional on the shares, and discuss consistency of the instrumental variables estimator in this context. We follow \citet*{borusyakhulljaravel2018shiftshare} by modeling the shifters as randomly assigned, since this approach follows naturally from our economic model. Using this assumption, we point out the potential bias of standard inference procedures when applied to shift-share designs, and provide a novel inference procedure that is valid in this context.
While our paper focuses on the statistical properties of the OLS estimator of $\beta$ in (ref), there exists a prior literature that has focused on studying the validity of different economic interpretations that one may attach to the estimand $\beta$. For example, this prior literature has studied how this interpretation may be affected by the presence of cross-regional general equilibrium effects \citep*{beraja2016aggregate,adao2018spatial}, slow adjustment of labor market outcomes to the shifters $\ensuremath\mathcal{X}_{s}$ \citep*{jaeger2018shift}, and heterogeneous effects of the shifters across sectors and regions \citep*{montereddingrossi2018}.
The rest of this paper is organized as follows. (ref) presents a placebo exercise illustrating the properties of the usual inference procedures. (ref) introduces a stylized economic model and maps its implications into a potential outcome framework. (ref) establishes the asymptotic properties of the OLS estimator of $\beta$ in (ref), as well as the properties of an instrumental variables estimator that uses a shift-share variable as an instrument. (ref) discusses extensions of our baseline framework. (ref) examines the performance of our novel inference procedures in a series of placebo exercises. (ref) revisits two prior applications of shift-share designs, and (ref) concludes. Proofs and additional results are collected in an Online Appendix.
In this section, we implement a placebo exercise to evaluate the finite-sample performance of the two inference methods most commonly applied in shift-share regression designs: (a) Eicker-Hubert-White---or heteroskedasticity-robust---standard errors, and (b) standard errors clustered on groups of regions geographically close to each other. In our placebo, we regress observed changes in U.S. regional labor market outcomes on a shift-share regressor that is constructed by combining actual data on initial sectoral employment shares for each region with randomly generated sector-level shocks. We describe the setup in (ref) and discuss the results in (ref).
We generate $30,000$ placebo samples indexed by $m$. Each of them contains $N = 722$ regions and $S = 396$ sectors. We identify each region $i$ with a U.S.\ Commuting Zone (CZ) and each sector $s$ with a 4-digit SIC manufacturing industry.
Using the notation from (ref), the shares $\{w_{is}\}_{i=1,s=1}^{N, S}$, and the outcomes $\{Y_{i}\}_{i=1}^{N}$ are identical in each placebo sample. The shares correspond to employment shares in 1990, and the outcomes correspond to changes in employment rates and average wages for different subsets of the population between 2000 and 2007. Our source of data on employment shares is the County Business Patterns, and our measures of changes in employment rates and average wages are based on data from the Census Integrated Public Use Micro Samples in 2000 and the American Community Survey for 2006 through 2008. Given these data sources, we construct our variables following the procedure described in the Online Appendix of \citet*{autordornhanson2013}.
The placebo samples differ exclusively in the shifters $\{\ensuremath\mathcal{X}_{s}^{m}\}_{s=1}^{\ensuremath{N}}$, which are drawn i.i.d.\ from a normal distribution with zero mean and variance equal to five in each placebo sample $m$. Since the shifters are independent of both the outcomes and the shares, the parameter $\beta$ is zero; this is true irrespective of the dependence structure between the outcomes and the shares.
For each placebo sample $m$, given the observed outcome $Y_{i}$, the generated shift-share regressor $X^{m}_{i}$ and a vector of controls $Z_{i}$ including only an intercept, we compute the OLS estimate of $\beta$, the heteroskedasticity-robust standard error (which we label Robust), and the standard error that clusters CZs in the same state (labeled Cluster).
(ref) presents the median and standard deviation of the empirical distribution of the OLS estimates of $\beta$ across the 30,000 placebo samples, along with the median standard error estimates, and rejection rates for 5% significance level tests of the null hypothesis $H_{0}\colon \beta=0$. We present these statistics for several outcome variables, which are listed in the leftmost column.
Column (1) of (ref) shows that, up to simulation error, the average of the OLS estimates is zero for all outcomes. Column (2) reports the standard deviation of the estimated coefficients. This dispersion is the target of the estimators of the standard error of the OLS estimator.\footnote{
reports the empirical distribution of the OLS estimates when the dependent variable is the change in each CZ's employment rate. Its distribution resembles a normal distribution centered around $\beta=0$.} Columns (3) and (4) report the median standard error estimates for the Robust and Cluster procedures, respectively, and show that both standard error estimators are downward biased. On average across all outcomes, the median magnitudes of the heteroskedasticity-robust and state-clustered standard errors are, respectively, 55% and 46% lower than the standard deviation.
The downward bias in the Robust and Cluster standard errors translates into a severe overrejection of the null hypothesis $H_{0}\colon\beta=0$. Since the true value of $\beta$ equals $0$ by construction, a correctly behaved test with significance level 5% should have a 5% rejection rate. Columns (5) and (6) in (ref) show that traditional standard error estimators yield much higher rejection rates. For example, when the outcome variable is the CZ's employment rate, the rejection rate is 48.5% and 38.1% when Robust and Cluster standard errors are used, respectively. These rejection rates are very similar when the dependent variable is instead the change in the average log weekly wage.
These results are quantitatively important. To see this, consider the following thought-experiment. Suppose we were to provide the $30,000$ simulated samples to $30,000$ researchers without disclosing the origin of the data to them. Instead, we would tell them that the shifters correspond to changes in a sectoral shock of interest---for instance, trade flows, tariffs, or national employment. If the researchers set out to test the null that the impact of this shock is zero using standard inference procedures at a 5% significance level, then over a third of them would conclude that our computer generated shocks had a statistically significant effect on the evolution of employment rates between 2000 and 2007.
The following remark summarizes the results of our placebo exercise.
To understand the source of this overrejection problem, note that the standard error estimators reported in (ref) assume that the regression residuals are either independent across all regions (for Robust), or between geographically defined groups of regions (for Cluster). Given that shift-share regressors are correlated across regions with similar employment shares $\{w_{is}\}_{s=1}^{S}$, these methods generally lead to a downward bias in the standard error estimate whenever regions with similar employment shares $\{w_{is}\}_{s=1}^{S}$ also have similar regression residuals. In the next section, we show how such correlations between regression residuals may arise.
This section presents a stylized economic model mapping labor demand and labor supply shocks to labor market outcomes for a set of regional economies. The aim of the model is twofold. First, it illustrates the economic mechanisms behind the overrejection problem documented in (ref). Second, it provides guidance on how to estimate: (i) the impact of sector-specific labor demand shifters on regional labor market outcomes; and (ii) the regional inverse labor supply elasticity. We describe the model fundamentals in (ref), discuss its main implications in (ref), and map these implications to a potential outcome framework in (ref).
We consider an economy with multiple sectors $s=1, \dotsc, S$ and multiple regions $i=1, \dotsc, N$. We assume that the labor demand in sector $s$ and region $i$, $L_{is}$, is given by
where $\omega_i$ is the wage rate in region $i$, $\sigma_s$ is the labor demand elasticity in sector $s$, and $D_{is}$ is a region- and sector-specific labor demand shifter. This shifter may account for multiple sectoral components. Specifically, we decompose $D_{is}$ into a sectoral shifter of interest $\chi_{s}$, other shifters that vary by sector $\mu_s$, and a residual region- and sector-specific shifter $\eta_{is}$:
We assume that the labor supply in region $i$ is given by
where $\phi$ is the labor supply elasticity, and $v_i$ is a region-specific labor supply shifter. We allow this shifter to have a shift-share structure that yields region-specific aggregates of group-specific labor supply shocks. In particular, indexing labor groups by $g=1, \dots, G$, we decompose
where $\nu_g $ is a group-specific labor supply shifter, $\tilde{w}_{ig}$ measures the exposure of region $i$ to group $g$ labor supply shifter, and $\nu_i$ captures region-specific factors affecting labor supply. The variable $\nu_g$ captures factors that affect the supply of labor of group $g$ in all regions in the population of interest. Workers may be classified into groups according to their education level, gender, or country of origin.
We assume that workers cannot move across regions but are freely mobile across sectors. Thus, labor markets clear if
We assume that, in each period, the model described by (ref) characterizes the labor market equilibrium in every region, and that, across periods, changes in the labor market outcomes $\{\omega_{i}, L_{i}\}_{i=1}^{N}$ are due to changes in either the labor demand shifters, $\{\chi_s, \mu_{s}\}_{s=1}^S$ and $\{\eta_{is}\}_{i=1,s=1}^{N, S}$, or the labor supply shifters, $\{\nu_g\}_{g=1}^{G}$ and $\{\nu_{i}\}_{i=1}^{N}$.
We use $\hat{z} = \log(z^{t}/z^{0})$ to denote log-changes in a variable $z$ between a period $t=0$ and some other period $t$. We assume that the realized changes between any two periods in all labor demand and supply shifters are draws from a joint distribution $F(\cdot)$:
Up to a first-order approximation around the initial equilibrium, (ref) imply that the changes in employment and wages in region $i$ are given by
where $l_{is}^{0}=L_{is}^{0}/L_{i}^{0}$ is the initial employment share of sector $s$ in region $i$, $\lambda_i =\phi \left[\phi + \sum_{s=1}^{S} l_{is}^0 \sigma_s\right]^{-1}$, and $\theta_{is} = \rho_s \lambda_i$.
Consider first the model's implications for the impact on regional labor market outcomes of changes in sector-specific labor demand. We focus here on the impact of the demand shocks $\{\hat{\chi}_{s}, \hat{\mu}_{s}\}_{s=1}^{S}$ on the change in the employment rate $\hat{L}_i$; however, given the symmetry between (ref), the model's implications for the impact of these shocks on the change in the wage level $\hat{\omega}_i$ are analogous.
According to (ref), the change in the employment rate in region $i$ depends on two shift-share components that aggregate the impact of the sector-specific labor demand shocks. In both components, the “share” term is the initial employment share $l_{is}^{0}$; the “shift” term corresponds in each of them to one of the two sector-specific labor demand shocks, $\hat{\chi}_{s}$ or $\hat{\mu}_s$. Furthermore, $\hat{L}_i$ also depends on additional shift-share terms that aggregate the impact of group-specific labor supply shocks. In this case, the “share” term is the region's exposure to each group-specific shock, $\tilde{w}_{ig}$. Conditional on a sector $s$ and a labor group $g$, the shares $\{l_{is}^0\}_{i=1}^{N}$ and $\{\tilde{w}_{ig}\}_{i=1}^{N}$ may be correlated. Settings in which the outcome of interest depends on multiple shift-share terms with potentially correlated shares is central to understanding the placebo results presented in (ref).
Another implication of (ref) is that, even conditional on the initial employment share $l_{is}^0$, the impact of sectoral labor demand shocks on regional employment may be heterogeneous across sectors and regions; e.g., the impact of $\hat{\chi}_{s}$ on $\hat{L}_i$ depends not only on $l_{is}^0$ but also on $\theta_{is}$, which may vary across $i$ and $s$. While datasets usually contain information on the initial employment shares for every sector and region $\{l_{is}^0\}_{i=1,s=1}^{N, S}$, the parameters $\{\theta_{is}\}_{i=1,s=1}^{N, S}$ are not generally known.
We summarize the discussion in the last two paragraphs in the following remark:
show that there are multiple microfoundations consistent with the insights summarized in (ref). Alternative microfoundations may differ in the mapping between the labor demand and supply elasticities, $\sigma_{s}$ and $\phi$, and structural parameters, or in the interpretation of the different terms entering the labor demand shifter $D_{is}$ in (ref).\footnote{In
, we derive (ref) from a multisector gravity model with endogenous labor supply that follows closely that in \citet*{adao2018spatial}. In
, we show that (ref) is consistent with a jones1971specific model featuring sector-specific production inputs, as in kovak2013regional. In
, we show that it is also consistent with a roy1951distribution model featuring workers with heterogeneous preferences for employment across sectors, as in \citet*{galle2017slicing}, lee2017inequality and \citet*{burstein2018inequality}.} In addition,
shows that similar insights arise in a model that allows for migration across regions. In this case, the change in regional employment depends not only on the region's own shift-share terms included in (ref), but also on a component, common to all regions, that combines the shift-share terms corresponding to all $N$ regions. In this environment, $l_{is}^0 \theta_{is}$ is the partial effect of the shifter $\hat{\chi}_{s}$ on $\hat{L}_i$ conditional on a fixed effect that absorbs cross-regional spillovers created by migration.
Turning to the estimation of the inverse labor supply elasticity, (ref) imply that
It follows from (ref) that the change in region $i$'s employment rate, $\hat{L}_{i}$, also depends on the term $\sum_{g=1}^G \tilde{w}_{ig} \hat{\nu}_g + \hat{\nu}_i$. Thus, the two terms on the right-hand side of (ref) are correlated with each other, creating an endogeneity problem. The instrumental variables solution to this problem relies on the observation that using (ref), one can write the inverse labor supply elasticity as the ratio of the impact of a sector-specific labor demand shock (e.g. $\hat{\chi}_{s}$) on wages to that on employment:
In (ref), we use the model described here to provide an economic interpretation for the econometric assumptions we impose when discussing identification and estimation in shift-share designs. These assumptions imply restrictions on the distribution of labor supply and demand shocks $F(\cdot)$ introduced in (ref). In (ref), we return to this economic model when interpreting empirical estimates of the impact of sector-specific labor demand shifters on regional labor market outcomes ((ref)); and the regional inverse labor supply elasticity ((ref)).
We build on the results in (ref) to propose a general framework for the estimation of the impact of shifters on outcomes measured at a different unit of observation. For concreteness, we refer to the level at which shifters vary as sectors and to the level at which the outcome varies as regions.
To make precise what we mean by “the effect of shifters on an outcome”, we use the potential outcomes notation, writing $Y_{i}(\ensuremath\mathcal{x}_{1}, \dotsc, x_{S})$ to denote the potential (counterfactual) outcome that would occur in region $i$ if the shocks to the $S$ sectors were exogenously set to $\{\ensuremath\mathcal{x}_{s}\}_{s=1}^{S}$. Consistently with (ref), we assume that the potential outcomes are linear in the shocks,
and $Y_{i}(0)=Y_{i}(0,\dotsc,0)$ denotes the potential outcome in region $i$ when all shocks $\{\ensuremath\mathcal{x}_{s}\}_{s=1}^{S}$ are set to zero. Thus, increasing $\ensuremath\mathcal{x}_{s}$ by one unit, holding the shocks to the other sectors constant, leads to an increase in region $i$'s outcome of $w_{is}\beta_{is}$ units. This is the treatment effect of $\ensuremath\mathcal{x}_{s}$ on $Y_{i}(\ensuremath\mathcal{x}_{1}, \dotsc, \ensuremath\mathcal{x}_{S})$. The actual (observed) outcome is given by $Y_{i}=Y_{i}(\ensuremath\mathcal{X}_{1}, \dotsc, \ensuremath\mathcal{X}_{S})$, which depends on the realization of the shifters, $(\ensuremath\mathcal{X}_{1}, \dotsc, \ensuremath\mathcal{X}_{S})$.
If the shifters of interest are the sectoral labor demand shocks $\{\hat{\chi}_{s}\}_{s=1}^{S}$, and the outcome of interest is the employment change $\hat{L}_{i}$, we can map (ref) into (ref) by defining
Observe that $Y_{i}(0)$ aggregates all shifters other than the sectoral shocks of interest $\{\hat{\chi}_{s}\}_{s=1}^{S}$.\footnote{Given the mapping in (ref), the expression in (ref) captures the first-order impact of the labor demand shocks $\{\hat{\chi}_{s}\}_{s=1}^{S}$ on changes in the employment rate. We focus on this first-order impact because it helps connecting our analysis to linear specifications used extensively in the shift-share literature. See
for a discussion of the approximation error arising from the linear specification imposed in (ref).}
We are interested in the properties of the OLS estimator $\hat{\beta}$ of the coefficient on the shift-share regressor $X_{i}=\sum_{s=1}^{S}w_{is}\ensuremath\mathcal{X}_{s}$ in a regression of $Y_{i}$ onto $X_{i}$.\footnote{We assume for now that the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$ are directly observable. In (ref), we consider the case in which we only observe noisy estimates of these shifters.} To focus on the key conceptual issues, we abstract away from any additional covariates or controls for now, and assume that $\ensuremath\mathcal{X}_{s}$ and $Y_{i}$ have been demeaned, so that we can omit the intercept in a regression of $Y_{i}$ on $X_{i}$ (see (ref) for the case with controls). In this simplified setting, the OLS estimator of the coefficient on $X_{i}$ is given by
and we can write the regression equation as
The definition of the estimand $\beta$ in (ref) and the properties of the estimator $\hat{\beta}$ will depend on: (a) what is the population of interest; and (b) how we think about repeated sampling. For (a), we define the population of interest to be the observed set of $N$ regions, as opposed to focusing on a large superpopulation of regions from which the $N$ observed regions are drawn. Consequently, we are interested in the parameters $\{\beta_{is}\}_{i=1,s=1}^{N, S}$ and the treatment effects $\{w_{is}\beta_{is}\}_{i=1,s=1}^{N, S}$ themselves, rather than the distributions from which they are drawn, which would be the case if we were interested in a superpopulation of regions.\footnote{Treating the set of observed regions as the population of interest is common in applications of the shift-share approach. For example, the abstract of \citet*{autordornhanson2013} reads: “We analyze the effect of rising Chinese import competition between 1990 and 2007 on U.S. local labor markets”. Similarly, the abstract of dixcarneirokovak2017 reads: “We study the evolution of trade liberalization's effects on Brazilian local labor markets” (emphases added).} For (b), given our interest on estimating the ceteris paribus impact of a specific set of shocks $(\ensuremath\mathcal{X}_{1}, \dotsc, \ensuremath\mathcal{X}_{S})$, we consider repeated sampling of these shocks, while holding the shares $\{w_{is}\}_{i=1,s=1}^{N, S}$, the parameters $\{\beta_{is}\}_{i=1,s=1}^{N, S}$, and the potential outcomes $\{Y_{i}(0)\}_{i=1}^{N}$ fixed.
Given these assumptions, the estimand $\beta$ is defined as the population analog of (ref) under repeated sampling of the shocks $\ensuremath\mathcal{X}_{s}$,
and, given (ref), the regression error $\epsilon_{i}$ is then defined as the residual
Thus, the statistical properties of the regression residual $\epsilon_{i}$ depend on the properties of the potential outcome $Y_{i}(0)$, the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$, the shares $\{w_{is}\}_{s=1}^{S}$, and the difference between the parameters $\{\beta_{is}\}_{s=1}^{S}$ and the estimand $\beta$. Importantly, as illustrated in (ref), the potential outcome $Y_{i}(0)$ will generally incorporate terms that have a shift-share structure with shares that are either identical to (e.g.\ the term $\sum_{s=1}^{S} w_{is}\hat{\mu}_s$) or different from but potentially correlated with (e.g.\ the term $\sum_{g=1}^G \tilde{w}_{ig} \hat{\nu}_g$) the shares $\{w_{is}\}_{s=1}^{S}$ that define the shift-share regressor $X_{i}$. It then follows from (ref) that the residuals $\epsilon_{i}$ and $\epsilon_{i'}$ will generally be correlated for any pair of regions $i$ and $i'$ with similar values of the shift-share regressor.
We summarize this discussion in the following remark.
(ref) has important implications for estimating the sampling variability of $\hat{\beta}$. In particular, traditional inference procedures do not account for correlation in $\epsilon_i$ among regions with similar shares and, therefore, tend to underestimate the variability of $\hat{\beta}$. As we formalize in the next section, this is the main reason for the overrejection problem described in (ref).
In this section, we formulate the statistical assumptions that we impose on the data generating process (DGP), use them to derive asymptotic results, and provide an economic interpretation of these assumptions using the model introduced in (ref). In (ref), we consider the case in which there is a single shift-share regressor and no controls. We account for controls in (ref). In (ref), we consider using the shift-share variable as an instrument for a regional treatment variable. All proofs and technical details are collected in
.
We follow the notation from (ref) by writing sector-level variables (such as the shifter $\ensuremath\mathcal{X}_{s}$) in script font style and region-level aggregates (such as $X_{i}$) in normal style. We use standard matrix and vector notation. In particular, for a (column) $L$-vector $A_{i}$ that varies at the regional level, $A$ denotes the $\ensuremath{N}\times L$ matrix with the $i$th row given by $A_{i}'$. For an $L$-vector $\mathcal{A}_{s}$ that varies at the sectoral level, $\mathcal{A}$ denotes the $S\times L$ matrix with the $s$th row given by $\mathcal{A}_{s}'$. If $L=1$, then ${A}$ and $\mathcal{A}$ are an $\ensuremath{N}$-vector and an $S$-vector, respectively. Let $W$ denote the $\ensuremath{N}\times S$ matrix of shares, so that its $(i, s)$ element is given by $w_{is}$, and let $B$ denote the $\ensuremath{N}\times S$ matrix with $(i, s)$ element given by $\beta_{is}$.
We focus here on the statistical properties of the OLS estimator $\hat{\beta}$ defined in (ref).
We consider large-sample properties of $\hat{\beta}$ as the number sectors goes to infinity, $S\to\infty$. The assumptions below imply that $N\to\infty$ as $S\to\infty$. To assess how large $S$ needs to be in order that these asymptotics provide a good approximation to the finite sample distribution of $\hat{\beta}$, we conduct a series of placebo simulations in (ref). We describe here the main substantive assumptions, and collect technical regularity conditions in
. As in (ref), let $\mathcal{F}_{0}=(Y(0), B, W)$.
(ref) requires that the potential outcomes are linear in the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$. As discussed in (ref), one can generate such linear specification from a first-order approximation of the impact of the shifters $(\ensuremath\mathcal{X}_{1}, \dotsc, \ensuremath\mathcal{X}_{S})$ on the outcome $Y_{i}$. This approximation may be subject to error. In
, we generalize (ref) to allow for a linearization error and derive restrictions on this error under which our inference procedures remain valid.
(ref) imposes that the sectoral shifters $\ensuremath\mathcal{X}$ are mean independent of the shares $W$, potential outcomes $Y(0)$, and parameters $B$; the assumption that the shifters are mean zero is a normalization to allow us to drop the intercept; we relax it in (ref). This random assignment assumption is a key assumption for identifying the causal impact of a shift-share covariate; a version of this assumption has been previously proposed by \citet*{borusyakhulljaravel2018shiftshare}.
If we are interested in studying the effect of labor demand shifters in the context of the model in (ref) (i.e. $\ensuremath\mathcal{X}_{s}=\hat{\chi}_{s}$), (ref) will hold if the shifters $\{\hat{\chi}_{s}\}_{s=1}^{S}$ are mean independent of the other labor demand shifters, $\{\hat{\mu}_{s}\}_{s=1}^{S}$ and $\{\hat{\eta}_{is}\}_{i=1,s=1}^{N, S}$, and of the labor supply shifters, $\{\hat{\nu}_{g}\}_{g=1}^{G}$ and $\{\hat{\nu}_{i}\}_{i=1}^{N}$. The plausibility of this restriction depends on the specific empirical application. For example, if all $N$ regions in the sample are regions within a small open economy, $\hat{\chi}_{s}$ denotes changes in international prices in sector $s$, and $\hat{\mu}_{s}$ denotes changes in the tariffs that this small open economy charges on its sector $s$ imports; then, (ref) requires these changes in tariffs to be independent of the changes in tariffs in any country that is large enough for their tariff changes to affect international prices (see
for additional details).
(ref) requires the shifters to be independent. It adapts to our setting the assumption underlying randomization-style inference in randomized controlled trials that the treatment assignment is independent across entities imbens_causal_2015. An independence or a weak dependence assumption of this type is generally necessary in order to do inference.\footnote{For example, for inference on average treatment effects, which is commonly the goal when running a regression, one typically assumes that the sample is a random sample from the population of interest and, thus, that the treatment variable is independent across the individuals in the sample.} One could alternatively impose assumptions on the correlation structure of the regression residuals, either by imposing a particular structure on them, as in the literature on interactive fixed effects gobillonmagnac2016, or by imposing a distance metric on the observations, as in the spatial econometrics literature conley_gmm_1999. However, as the economic model in (ref) shows, the structure of the residuals may be very complex. The residuals may include potentially correlated region-specific terms as well as several shift-share terms, which may or may not use the same shares as the covariate of interest $X_{i}$. It is thus difficult to conceptualize which exact restriction on their joint distribution one should impose.
By instead imposing restrictions on the distribution of the vector of shifters $(\ensuremath\mathcal{X}_{1}, \dotsc, \ensuremath\mathcal{X}_{S})$ conditional on $\mathcal{F}_{0}=(Y(0), B, W)$, (ref) ensures that the standard errors we derive remain valid under any dependence structure between the shares $w_{is}$ across sectors and regions, and under any correlation structure of the potential outcomes $Y_{i}(0)$ or, equivalently, of the regression errors $\epsilon_{i}$, across regions.\footnote{Since our inference is valid conditional on $\{\epsilon_{i}\}_{i=1}^{N}$, it accounts for any correlation structure they may have, including spatial, or, in applications with multiple periods, temporal correlations. See (ref) for settings with multiple periods.} We thus do not have to worry about correctly specifying this correlation structure, as one would under the alternative approaches mentioned above. Our approach allows (but does not require) the residual to have a shift-share structure; it similarly allows all $\{w_{is}\}_{i=1,s=1}^{N, S}$ to be equilibrium objects responding to the same economic shocks, and thus be correlated across regions and sectors.\footnote{This conceptualization of all the shares $w_{is}$ as equilibrium objects that respond (at least partly) to the same set of shocks is consistent with the model in (ref). As shown in (ref), each share $w_{is}$ corresponds to the share of workers in region $i$ employed in sector $s$ in an initial equilibrium, $l^{0}_{is}$. Furthermore, each of these initial employment shares will be a function of the same sector-specific demand shocks and group-specific labor supply shocks; consequently $l^{0}_{is}$ will generally be correlated with $l^{0}_{i's'}$ even for $i\neq i'$ and $s\neq s'$.} In (ref), we relax (ref) and allow for a non-zero correlation in the shifters $(\ensuremath\mathcal{X}_{1}, \dotsc, \ensuremath\mathcal{X}_{S})$ within clusters of sectors; we only require that the shifters are independent across the clusters. Additionally, in the context of the empirical application in (ref), we discuss how to perform inference in a setting in which all shifters of interest are generated by a common shock that has heterogeneous effects across sectors.
In the economic model in (ref), if $\ensuremath\mathcal{X}_{s}=\hat{\chi}_{s}$ and we interpret these shocks as, for example, sector-specific productivity shocks, (ref) requires that there is no common component driving the changes in sectoral productivities. Our approach does not require the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$ to be identically distributed; we allow, for example, the variance of the shock to differ across sectors.
(ref) are our main regularity conditions.\footnote{In the context of a shift-share instrumental variables regression, \citet*{goldsmith2018bartik} discuss similar conditions stated in terms of Rotemberg weights. This is convenient under the baseline assumption considered in \citet*{goldsmith2018bartik} that the vector of shares $(w_{i1}, \dotsc, w_{iS})$ is exogenous, because the Rotemberg weights determine the asymptotic bias of the estimator under local failures of this exogeneity condition. Since we do not assume exogeneity of the shares, this interpretation is not available under our setup.} (ref) is needed for consistency: it requires that the size of each sector, $n_{s}$, is asymptotically negligible. This assumption is analogous to the standard consistency condition in the clustering literature that the largest cluster be asymptotically negligible. To see the connection, consider the special case with “concentrated sectors”, in which each region $i$ specializes in one sector $s(i)$; i.e. $w_{is}=1$ if $s=s(i)$ and $w_{is}=0$ otherwise, and $n_{s}$ is thus the number of regions that specialize in sector $s$. In this case, $X_{i}=\ensuremath\mathcal{X}_{s(i)}$, so that, if (ref) holds, $\hat{\beta}$ is equivalent to an OLS estimator in a randomized controlled trial in which the treatment varies at a cluster level; here the $s$th cluster consists of regions that specialize in sector $s$. The condition $\max_{s}n_{s}/\sum_{t=1}^{S}n_{t}\to 0$ then reduces to the assumption that the largest cluster be asymptotically negligible. (ref) is needed for asymptotic normality---it ensures that the Lindeberg condition holds. It strengthens (ref) slightly by requiring that the contribution of each sector to the asymptotic variance is asymptotically negligible; otherwise the estimator will not generally be asymptotically normal, even if it is consistent.
In terms of the economic model introduced in (ref), (ref) require that no sector dominates the rest in terms of initial employment at the national level; i.e. $\sum_{i=1}^{N}l^{0}_{is}$ is not too large for any sector. (ref) shows that this assumption is reasonable for the U.S. if the $S$ sectors used to construct the treatment of interest $X_{i}$ correspond to the 396 4-digit manufacturing sectors (see (ref)). In (ref), we illustrate the consequences of the failure of this assumption due to the inclusion of a large aggregate sector, the non-manufacturing sector, in $X_{i}$.
We now establish that the OLS estimator in (ref) is consistent and asymptotically normal.
This \namecref{theorem:consistency-noZ} gives two results. First, it shows that the estimand $\beta$ in (ref) can be expressed as a weighted average of the region- and sector-specific parameters $\{\beta_{is}\}_{i=1,s=1}^{N, S}$, with the weight $\pi_{is}$ increasing in the share $w_{is}$ and in the conditional variance of the shifter $\operatorname{var}(\ensuremath\mathcal{X}_{s}\mid \mathcal{F}_{0})$. Second, it states that the OLS estimator $\hat{\beta}$ converges to this estimand as $S\to\infty$. The special case with concentrated sectors is again useful in interpreting (ref). In this case, $\sum_{s=1}^{S}\pi_{is}\beta_{is}=\operatorname{var}(\ensuremath\mathcal{X}_{s(i)}\mid \mathcal{F}_{0})\beta_{is(i)}$ and, therefore, the first result in (ref) reduces to the standard result from the randomized controlled trials literature with cluster-level randomization (with each “cluster” defined as all regions specialized in the same sector) that the weights are proportional to the variance of the shock.
The estimand $\beta$ does not in general equal a weighted average of the heterogeneous treatment effects. As discussed in (ref), the effect on the outcome in region $i$ of increasing the value of the sector $s$ shock in one unit is equal to $w_{is}\beta_{is}$; weighting this effect using a set of region- and sector-specific weights $\{\xi_{is}\}_{i=1,s=1}^{N, S}$, yields the weighted average treatment effect
Alternatively, the total effect of increasing the shifters simultaneously in every sector by one unit is $\sum_{s=1}^{S}w_{is}\beta_{is}$; weighting it using a set of region-specific weights $\{\zeta_{i}\}_{i=1}^{\ensuremath{N}}$ yields the weighted total treatment effect $\tau^{T}_{\zeta}=\sum_{i=1}^{\ensuremath{N}}\zeta_{i}\sum_{s=1}^{S}w_{is}\beta_{is}/\sum_{i=1}^{\ensuremath{N}}\zeta_{i}$. If $\beta_{is}$ is constant across $i$ and $s$, then $\beta=\tau^{T}_{\zeta}$, provided $\sum_{s=1}^{S}w_{is}=1$ in every region $i$; otherwise, we can consistently estimate $\tau^{T}_{\zeta}$ by $\hat{\beta}\cdot \sum_{i=1}^{\ensuremath{N}}\zeta_{i}\sum_{s=1}^{S}w_{is}/\sum_{i=1}^{\ensuremath{N}}\zeta_{i}$. Similarly, if $\beta_{is}$ is constant across $i$ and $s$, $\tau_{\xi}$ is consistently estimated by $\hat{\beta}\cdot\sum_{i=1}^{N}\sum_{s=1}^{S}\xi_{is}w_{is} / \sum_{i=1}^{N}\sum_{s=1}^{S}\xi_{is}$. On the other hand, if $\beta_{is}$ varies across regions and sectors, then it is not clear in general how to exploit knowledge of the estimand $\beta$ defined in (ref) to learn something about $\tau_{\xi}$ or $\tau^{T}_{\zeta}$. A special case in which it is possible to consistently estimate $\tau_{\xi}$ even if $\beta_{is}$ varies across $i$ or $s$ arises when $\ensuremath\mathcal{X}_{s}$ is homoskedastic, $\operatorname{var}(\ensuremath\mathcal{X}_{s}\mid \mathcal{F}_{0})=\sigma^{2}$, and $\xi_{is}=w_{is}$; in this case, a consistent estimate of $\tau_{\xi}$ is given by $\hat{\beta}\sum_{i=1}^{\ensuremath{N}}\sum_{s=1}^{S}w_{is}^{2}/\sum_{i=1}^{\ensuremath{N}}\sum_{s=1}^{S}w_{is}$.\footnote{In general, one can consistently estimate $\tau_{\xi}$ or $\tau^{T}_{\zeta}$ by imposing a mapping between $\beta_{is}$ and structural parameters, and obtaining consistent estimates of these structural parameters. However, since this mapping will vary across models, the consistency of such estimator will not be robust to alternative modeling assumptions, even if all these assumptions predict an equilibrium relationship like that in (ref); e.g.\ see
for examples of this mapping in different models.}
This proposition shows that $\hat{\beta}$ is asymptotically normal, with a rate of convergence equal to $\ensuremath{N}(\sum_{s=1}^{S}n_{s}^{2})^{-1/2}$. If all sector sizes $n_{s}$ are of the order $\ensuremath{N}/S$, the rate of convergence equals $\sqrt{S}$. However, if the sizes are unequal, the rate may be slower.
According to (ref), the asymptotic variance formula has the usual “sandwich” form. Since $X_{i}$ is observed, to construct a consistent standard error estimate, it suffices to construct a consistent estimate of $\mathcal{V}_{\ensuremath{N}}$, the middle part of the sandwich. To motivate our standard error formula, suppose that $\beta_{is}$ is constant across $i$ and $s$, $\beta_{is}=\beta$. Then it follows from (ref) and (ref) that
Replacing $\operatorname{var}(\ensuremath\mathcal{X}_{s}\mid \mathcal{F}_{0})$ by $\ensuremath\mathcal{X}_{s}^{2}$, and $\epsilon_{i}$ by the regression residual $\hat{\epsilon}_{i}=Y_{i}-X_{i}\hat{\beta}$, we obtain the estimate
When $\beta_{is}=\beta$, we show formally that this variance estimate leads to valid inference under regularity conditions in (ref). In
we show that this variance estimate remains valid under heterogeneous $\beta_{is}$ under further regularity conditions.
To gain intuition for the variance estimate in (ref), consider the case with concentrated sectors. Then the numerator in (ref) becomes $\sum_{s=1}^{S}\ensuremath\mathcal{X}_{s}^{2}\hat{R}_{s}^{2}=\sum_{s=1}^{S}(\sum_{i=1}^{\ensuremath{N}}\1{s(i)=s}X_{i}\hat{\epsilon}_{i})^{2}$, so that (ref) reduces to the cluster-robust variance estimate that clusters on the sector that each region is specialized. This is consistent with the rule of thumb that one should “cluster” at the level of variation of the regressor of interest. More generally, the variance estimate essentially forms sectoral clusters with variance that depends on the variance of $\hat{R}_{s}$, a weighted sum of the regression residuals $\{\hat{\epsilon}_{i}\}_{i=1}^{N}$, with weights that correspond to the shares $\{w_{is}\}_{i=1}^{N}$. An important advantage of $\widehat{V}_{AKM}(\hat{\beta})$ is that it allows for an arbitrary structure of cross-regional correlation in residuals:
To understand the source of the overrejection problem discussed in (ref), let us compare the variance estimate $\widehat{V}_{AKM}(\beta)$ with the cluster-robust variance estimate when the residuals $\hat{\epsilon}_{i}$ are computed at the true $\beta$ (so that $\hat{\epsilon}_{i}=\epsilon_{i}$). These variance estimates differ in the middle sandwich, with the cluster-robust estimate replacing $\hat{\mathcal{V}}_{AKM}({\beta})$ in (ref) with $\hat{\mathcal{V}}_{CL}({\beta})=\sum_{i=1}^{N}\sum_{j=1}^{N}\1{c(i)=c(j)}X_{i}X_{j}{\epsilon}_{i}{\epsilon}_{j}$, where $c(i)$ denotes the cluster that region $i$ belongs to (the comparison with heteroskedasticity-robust standard errors obtains as a special case if $c(i)=i$, so that each region belongs to its own cluster). Assuming for simplicity that the conditional variance of $\ensuremath\mathcal{X}_{s}$ does not depend on $Y(0)$, it follows by simple algebra that the expectation of the difference between these terms is given by
This expression is non-negative so long as the correlation between the residuals is non-negative. The magnitude of the difference will be large if regions located in different clusters (so that $c(i)\neq c(j)$) that have similar shares (i.e.\ large values of $\sum_{s=1}^{S}w_{is}w_{js}$) also tend to have similar residuals (i.e.\ large values of $E[\epsilon_{i}\epsilon_{j}\mid W]$). For illustration, consider a simplified version of the model described in (ref) in which: (a) $\sigma_{s}\geq 0$ for all $s$ and $\phi\geq 0$, so that $0\leq \lambda_{i}\leq 1$; (b) region-specific labor demand and supply shocks $\{\hat{\eta}_{is}\}_{s=1}^{S}$ and $\hat{\nu}_{i}$ are independent across regions; and (c) all labor demand and supply shocks are independent of each other. Then, it follows from (ref) that, for any $i\neq j$,
which by the law of iterated expectations implies that $E[\hat{\mathcal{V}}_{AKM}(\beta)-\hat{\mathcal{V}}_{CL}(\beta)\mid W]\geq 0$. This expression illustrates that regions with similar shares will tend to have similar residuals in two cases. First, if the variance of the unobserved shifter $\hat{\mu}_{s}$ is large, so that $E[\hat{\mu}_{s}^{2}\mid W, \tilde{W}]$ is large. In other words, standard inference methods lead to overrejection if the residual contains important shift-share terms that affect the outcome of interest through the same shares $\{w_{is}\}_{s=1}^{S}$ as those defining the covariate of interest $X_{i}$. Second, if the variance of the unobserved shifter $\hat{\nu}_{g}$ is large, so that $E[\hat{\nu}_{g}^{2} \mid W, \tilde{W}]$ is large, and the shares $\tilde{w}_{ig}$ through which these shifters affect the outcome variable have a correlation structure that is similar to that of $w_{is}$ (so that $\sum_{g=1}^{G}\tilde{w}_{ig}\tilde{w}_{jg}$ is large whenever $\sum_{s=1}^{S}w_{is}w_{js}$ is large). Thus, standard inference methods may overreject even when the unobserved shifters contained in the residual vary along a different dimension than the shift-share covariate of interest.
We now study the properties of the OLS estimator $\hat{\beta}$ of the coefficient on $X_{i}$ in a regression of $Y_{i}$ onto $X_{i}$ and a $K$-vector of controls $Z_{i}$. To this end, let $Z$ denote the $\ensuremath{N}\times K$ matrix with $i$-th row given by $Z_{i}'=(Z_{i1}, \dotsc, Z_{iK})$, and let $\ddot{X}=X-Z(Z'Z)^{-1}Z'X$ denote an $N$-vector with $i$-th element equal to the regressor $X_{i}$ with the controls $Z_{i}$ partialled out (i.e.\ the residual from regressing $X_{i}$ onto $Z_{i}$). Then, by the Frisch–Waugh–Lovell theorem, $\hat{\beta}$ can be written as
The controls may play two roles. First, they may be included to increase the precision of $\hat{\beta}$. Second, and more importantly, they may be included because one may worry that the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$ are correlated with the potential outcomes $\{Y_{i}(0)\}_{i=1}^{N}$, violating (ref). To formalize how $Z_{i}$, a regional variable, may be a control variable for the shifters, which vary at a sectoral level, we project $Z_{i}$ onto the sectoral space using the same shares as those defining the shift-share regressor $X_{i}$,
We think of $\{\ensuremath\mathcal{Z}_{s}\}_{s=1}^{S}$ as latent sector-level shocks that may have an independent effect on the outcome $Y$ and may also be correlated with the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$, with $U_{i}$, the residual in this projection, mean-independent of the shifters. If the $k$th control $Z_{ik}$ is included for precision, then the sector-level shocks $\{\ensuremath\mathcal{Z}_{sk}\}_{s=1}^{S}$ and, thus, $Z_{ik}$, are uncorrelated with $X_{i}$. If $Z_{ik}$ is included because one worries that otherwise $X_{i}$ may not be as good as randomly assigned, we interpret $Z_{ik}$ as a proxy for the confounding sector-level shocks $\{\ensuremath\mathcal{Z}_{sk}\}_{s=1}^{S}$, and think of $U_{ik}$ as a measurement error in this proxy.
To make this concrete, consider the model in (ref), with the equivalences in (ref). Then we may include $Z_{ik}=\sum_{s=1}^{S}l^{0}_{is}\hat{\mu}_{s}$ as a control. Here the measurement error in (ref) is zero, and $\ensuremath\mathcal{Z}_{sk}=\hat{\mu}_{s}$. If the shifters $\{\hat{\chi}_{s}\}_{s=1}^{S}$ are correlated with the demand shocks $\{\hat{\mu}_{s}\}_{s=1}^{S}$, then not including this control will generate omitted variable bias. Alternatively, we may include $Z_{ik}=\sum_{s=1}^{S}w_{is} \hat{\eta}_{is}$ as a control. Here $\ensuremath\mathcal{Z}_{sk}=0$, and $U_{ik}=Z_{ik}$ is a regional aggregation of idiosyncratic region- and sector-specific labor-demand shocks that are independent of $\ensuremath\mathcal{X}_{s}$. In this case, if the shifters $\{\hat{\chi}_{s}\}_{s=1}^{S}$ are independent of the demand shocks $\{\eta_{is}\}_{i=1,s=1}^{N, S}$, then including the control will help increase the precision of $\hat{\beta}$, but it is not necessary for consistency.
For clarity of exposition, we focus here on the main substantive assumptions and relegate technical regularity conditions to
. Let $\mathcal{F}_{0}=(Y(0), W, B, \ensuremath\mathcal{Z}, U)$; without controls, this set of variables reduces to $(Y(0), B, W)$, as in (ref). Here, $\ensuremath\mathcal{Z}$ denotes the $S\times K$ matrix with $s$th row given by $\ensuremath\mathcal{Z}_{s}'$, and $U$ denotes the $N\times K$ matrix with $i$-th element given by $U_{i}'$.
We maintain (ref) with $\mathcal{F}_{0}=(Y(0), W, B, \ensuremath\mathcal{Z}, U)$. The inclusion of controls allows us to weaken (ref) and instead impose the following identification assumption:
(ref) weakens (ref) by only requiring the shifters to be as good as randomly assigned conditional on $\ensuremath\mathcal{Z}$, in the sense that (ref) holds. To interpret this restriction, consider a projection of the regional potential outcomes onto the sectoral space. For simplicity, consider the case with constant effects, $\beta_{is}=\beta$ for all $i$ and $s$, and project $Y_{i}(0)$ onto the shares $(w_{i1}, \dotsc, w_{iS})$, so that we may write $Y_{i}(0)=\sum_{s=1}^{S}w_{is}\mathcal{Y}_{s}(0)+\kappa_{i}$. Then, (ref) holds if (i) $\mathcal{Y}_{s}(0)$ is spanned by the vector of controls $\ensuremath\mathcal{Z}_{s}$; and (ii) $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$ is mean-independent of the projection residuals $\{\kappa_{i}\}_{i=1}^{N}$.
As an example, consider again the model in (ref), with the outcomes $Y_{i}$ generated by (ref). Then (ref) holds, for example, if we set $\mathcal{Z}_{s}=\mathcal{Y}_{s}(0)=\hat{\mu}_{s}$ and if, conditional on the sector-specific labor demand shocks $\{\hat{\mu}_{s}\}_{s=1}^{S}$, the shifters of interest $\{\hat{\chi}_{s}\}_{s=1}^{S}$ are mean independent of the sector- and region-specific labor demand shocks $\{\hat{\eta}_{is}\}_{i=1,s=1}^{N, S}$ and of the labor supply shocks $\{\hat{\nu}_{g}\}_{g=1}^{G}$ and $\{\hat{\nu}_{i}\}_{i=1}^{N}$. Suppose, for instance, the shocks of interest $\{\hat{\chi}_{s}\}_{s=1}^{S}$ are changes in tariffs kovak2013regional and that other potential labor demand shocks are those induced by automation and robots acemoglu2017robots. Splitting the impact of automation into nationwide sector-specific effects, as captured by $\{\hat{\mu}_{s}\}_{s=1}^{S}$, and sector- and region-specific deviations from the nationwide effects, as captured by $\{\hat{\eta}_{is}\}_{i=1,s=1}^{N, S}$, (ref) allows the political entity responsible for setting the tariffs to do so influenced by the nationwide sector-specific effects of automation, but not by any region-specific deviation from those national effects. In contrast, (ref) would require that the tariffs are also independent of the nationwide effects of automation.
Under (ref), one generally needs to include the controls non-parametrically; by imposing (ref), we ensure that it suffices to include the controls as additional covariates in a linear regression. If the shifters $\ensuremath\mathcal{X}_{s}$ are not mean zero (in the sense that the regression intercept on the right-hand side of (ref) is non-zero), (ref) requires that we include a constant $\ensuremath\mathcal{Z}_{sk}=1$ as one of the controls. If the shares sum to one, $\sum_{s=1}^{S}w_{is}=1$, this amounts to including an intercept $Z_{ik}=1$ as a control in the regression. Importantly, if the shares do not sum to one, this amounts to including $\sum_{s=1}^{S}w_{is}$ as a control \citep*[see][for a more extensive discussion of this point]{borusyakhulljaravel2018shiftshare}. For instance, if the shares $w_{is}$ correspond to labor shares in different manufacturing sectors, one needs to include the size of the manufacturing sector $\sum_{s=1}^{S}w_{is}$ in each region as a control.
Given (ref), if we observed $\{\ensuremath\mathcal{Z}_{s}\}_{s=1}^{S}$ directly, we could include the vector $Z_{i}^{*}=\sum_{s=1}^{S}w_{is}\ensuremath\mathcal{Z}_{s}$ directly as control. However, the definition of each regional control $Z_{i}$ in (ref) allows for $Z_{i}^{*}$ to be observed with measurement error $U_{i}$. If $\gamma_{k}=0$, such as when $Z_{ik}$ is included for precision, then this measurement error in $Z_{ik}^{*}$ does not matter; if $\gamma_{k}\neq0$, this measurement error will in general induce a bias in $\hat{\beta}$. This is analogous to the classic linear regression result that measurement error in a control variable generally leads to a bias in the estimate of the coefficient on the variable of interest. (ref) ensures that any such bias disappears in large samples by imposing that the variance of the measurement error for controls that matter (i.e.\ those with $\gamma_{k}\neq 0$) converges to zero as $S\to\infty$. This ensures consistency of $\hat{\beta}$. For asymptotic normality, we need to strengthen this condition in (ref) by requiring that the variance of the measurement error converges to zero sufficiently fast. (ref) holds, for instance, if $U_{i}=S^{-1}\sum_{s=1}^{S}\psi_{is}$, where $\psi_{is}$ is an idiosyncratic measurement error that is independent across $s$. In intuitive terms, this condition guarantees that $Z_{i}$ is a sufficiently good proxy for the confounding latent shocks $\{\ensuremath\mathcal{Z}_{s}\}_{s=1}^{S}$.
The following result generalizes (ref):
The only difference in the characterization of the probability limit relative to (ref) is that the weights $\pi_{is}$ now reflect the variance of $\ensuremath\mathcal{X}_{s}$ that also conditions on the controls.
To state the asymptotic normality result, define $\delta=E[Z'Z]^{-1}E[Z'(Y-X\beta)]$, so that we can define the regression residual in (ref) as $\epsilon_{i}=Y_{i}-X_{i}\beta-Z_{i}'\delta$.
Relative to (ref), the main difference is that $X_{i}$ in the definition of $\mathcal{V}_{\ensuremath{N}}$ is replaced by $X_{i}-Z_{i}'\gamma$, and that $X_{i}$ is replaced by $\ddot{X}_{i}$ in the outer part of the “sandwich.” To motivate our standard error formula, suppose that $\beta_{is}=\beta$ for all $i$ and $s$. Under $\beta_{is}=\beta$, it follows from (ref) and (ref) that
A plug-in estimate of $R_{s}$ can be constructed by replacing $\epsilon_{i}$ with the estimated regression residuals $\hat{\epsilon}_{i}=Y_{i}-X_{i}\hat{\beta}-Z_{i}\hat{\delta}$, where $\hat{\delta}=(Z'Z)^{-1}Z'(Y-X\hat{\beta})$ is an OLS estimate of $\delta$. We can estimate the variance $\operatorname{var}(\tilde{\ensuremath\mathcal{X}}_{s}\mid \mathcal{F}_{0})$ by $\widehat{\ensuremath\mathcal{X}}^{2}$, where
projects the estimate $\ddot{X}$ of $X-Z'\gamma$ onto the sectoral space by regressing it onto the shares $W$. To carry out the regression in (ref), $W$ must be full rank; this requires that there are more regions than sectors, $\ensuremath{N}\geq S$. These steps lead to the standard error estimate
The next remark summarizes the steps needed for the construction of the standard error $\widehat{se}(\hat{\beta})$:
To gain intuition for the procedure in (ref), it is useful to consider again the case with concentrated sectors. Suppose that $U_{i}=0$ for all $i$, so that the regression of $Y_{i}$ onto $X_{i}$ and $Z_{i}$ is identical to the regression of $Y_{i}$ onto $\ensuremath\mathcal{X}_{s(i)}$ and $\ensuremath\mathcal{Z}_{s(i)}$. Then the standard error formula in (ref) reduces to the usual cluster-robust standard error, with clustering on $s(i)$.
The cluster-robust standard error is generally biased due to estimation noise in estimating ${\epsilon}_{i}$, which can lead to undercoverage, especially in cases with few clusters (see cameron_practitioners_2014 for a survey). Since the standard error in (ref) can be viewed as generalizing the cluster-robust formula, similar concerns arise in our setting. We thus consider a modification $\widehat{se}_{\beta_{0}}(\hat{\beta})$ of $\widehat{se}(\hat{\beta})$ that imposes the null hypothesis when estimating the regression residuals to reduce the estimation noise in estimating ${\epsilon}_{i}$.\footnote{Alternatively, one could construct a bias-corrected variance estimate; see, for example, BeMc02 for an example of this approach in the context of cluster-robust inference.} To calculate the standard error $\widehat{se}_{\beta_{0}}(\hat{\beta})$ for testing the hypothesis $H_{0}\colon \beta=\beta_{0}$ against a two-sided alternative at significance level $\alpha$, one replaces $\hat{\epsilon}_{i}$ with $\hat{\epsilon}_{\beta_{0}, i}$, the residual from regressing $Y_{i}-X_{i}\beta_{0}$ onto $Z_{i}$ ($\hat{\epsilon}_{\beta_{0}, i}$ is an estimate of the residuals with the null imposed). The null is rejected if the absolute value of the $t$-statistic $(\hat{\beta}-\beta_{0})/\widehat{se}_{\beta_{0}}(\hat{\beta})$ exceeds $z_{1-\alpha/2}$, the $1-\alpha/2$ quantile of a standard normal distribution (1.96 for $\alpha=0.05$). To construct a confidence interval (CI) with coverage $1-\alpha$, one collects all hypotheses $\beta_{0}$ that are not rejected. The endpoints of this CI are a solution to a quadratic equation, and are thus available in closed form---one does not have to numerically search for all the hypotheses that are not rejected. The next remark summarizes this procedure.
Since in both $\hat{\epsilon}_{i}$ and $\hat{\epsilon}_{\beta_{0}, i}$ are consistent estimates of the residuals, this proposition shows that the procedures in (ref) both yield asymptotically valid confidence intervals. The additional assumptions of (ref) ensure that the estimation error in $\widehat{\ensuremath\mathcal{X}}_{s}$ that arises from having to back out the sector-level shocks $\ensuremath\mathcal{Z}_{s}$ from the controls $Z_{i}$ is not too large. If the sectors are concentrated, then $((W'W)^{-1}W')_{si}=\1{s(i)=s}/n_{s}$, so that $\max_{s}\sum_{i=1}^{N}\abs{((W'W)^{-1}W')_{si}}=1$, and the assumption always holds. We show in
that the procedures in (ref) continue to yield valid inference if $\beta_{is}$ is heterogeneous across regions and sectors, as long as further regularity conditions hold.
Although both standard errors $\widehat{se}_{\beta_{0}}(\hat{\beta})$ and $\widehat{se}(\hat{\beta})$ are consistent (and one could further show that the resulting confidence intervals are asymptotically equivalent), they will in general differ in finite samples. In particular, it can be seen from (ref) that the confidence interval with the null imposed is not symmetric around $\hat{\beta}$, but its center is shifted by $A$.\footnote{This is analogous to the differences in likelihood models between confidence intervals based on the Lagrange multiplier test (which imposes the null and is not symmetric around the maximum likelihood estimate) and the Wald test (which does not impose the null and yields the usual confidence interval).} As we show in (ref), this recentering tends to improve the finite-sample coverage properties of the confidence interval. On the other hand, the confidence interval described in (ref) tends to be longer on average than that in (ref).
We now turn to the problem of estimating the effect of a regional treatment variable $Y_{2i}$ on a regional outcome $Y_{1i}$ using the shift-share variable $X_{i}=\sum_{s=1}^{S}w_{is}\ensuremath\mathcal{X}_{s}$ as an instrumental variable (IV). To set up the problem precisely, we again use the potential outcome framework. In particular, we assume that
where $\alpha$, our parameter of interest, measures the causal effect of $Y_{2i}$ onto $Y_{1i}$. We assume for simplicity that this causal effect is linear and constant across regions.\footnote{If we weaken the assumption of constant treatment effects and instead assume $Y_{1i}(y_{2})=Y_{1i}(0)+y_{2}\alpha_{i}$, then it follows by a mild extension of the results in
that our methods would deliver inference on the estimand $\sum_{i=1}^{\ensuremath{N}}\pi_{i}\alpha_{i}/\sum_{i=1}^{\ensuremath{N}}\pi_{i}$, with $\pi_{i}=\sum_{s=1}^{S}w_{is}^{2}\operatorname{var}(\ensuremath\mathcal{X}_{s}\mid\mathcal{F}_{0})\beta_{is}$, where $\mathcal{F}_{0}=(\ensuremath\mathcal{Z}, U, Y_{1}(0), Y_{2}(0), B, \alpha, W)$, and $\beta_{is}$ is defined in (ref).} In analogy with (ref), we denote the region-$i$ treatment level that would occur if the region received shocks $(\ensuremath\mathcal{x}_{1}, \dotsc, \ensuremath\mathcal{x}_{S})$ as
The observed outcome and treatment variables are given by $Y_{1i}=Y_{1i}(Y_{2i})$ and $Y_{2i}=Y_{2i}(\ensuremath\mathcal{X}_{1}, \dotsc, \ensuremath\mathcal{X}_{S})$, respectively.
The framework in (ref) maps directly to the problem of estimating the regional inverse labor supply elasticity. In particular, in the context of the model in (ref), (ref) map directly into (ref) if we define
and $Y_{2i}(0)$ is given by the expression for $Y_{i}(0)$ in (ref).\footnote{In some applications of shift-share IVs, the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$ are unobserved and have to be estimated. We assume here that $\ensuremath\mathcal{X}_{s}$ is directly measurable for every sector $s$, and study the case with estimated shifters in (ref).} As this mapping illustrates, the potential outcome $Y_{1i}(0)$ will generally have a shift-share structure, with the shifters being group-specific labor supply shocks (e.g.\ growth in the number of workers by education group). Consequently, the regression residual in the structural equation will generally have a shift-share structure. Similarly, as (ref) illustrates, the potential outcome $Y_{2i}(0)$ will also generally include several shift-share components, with the shifters being either sector-specific labor demand shocks or the same group-specific labor supply shocks appearing in $Y_{1i}(0)$. Thus, the regression residual in the first-stage regression of $Y_{2i}$ onto $X_{i}$ will also generally have a shift-share structure.
Our estimate of $\alpha$ is given by an IV regression of $Y_{1i}$ onto $Y_{2i}$ and a $K$-vector of controls $Z_{i}$, with $X_{i}$ used as an instrument for $Y_{2i}$. This IV estimate can be written as
where, as in (ref), $\ddot{X}_{i}$ denotes the residual from regressing $X_{i}$ onto $Z_{i}$.
(ref) is a generalization of (ref). Let $\mathcal{F}_{0}=(\ensuremath\mathcal{Z}, U, Y_{1}(0), Y_{2}(0), B, W)$.
(ref) adapts the standard instrument exogeneity condition (see, e.g., Condition 1 in imbens_identification_1994) to our setting. Our approach follows \citet*{borusyakhulljaravel2018shiftshare}, who impose a similar identification condition. To illustrate the restrictions that (ref) may impose, consider again the problem of estimating the inverse labor supply elasticity within the context of the model in (ref), with the mapping between this model and the potential outcomes in (ref) given in (ref). If the controls $\{\ensuremath\mathcal{Z}_{s}\}_{s=1}^{S}$ correspond to the shocks $\{\hat{\mu}_{s}\}_{s=1}^{S}$, then (ref) requires that, conditional on $\{\hat{\mu}_{s}\}_{s=1}^{S}$, the labor demand shocks $\{\hat{\chi}_s\}_{s=1}^{S}$ used to construct our IV are mean-independent of the idiosyncratic labor demand shocks $\{\hat{\eta}_{is}\}_{i=1,s=1}^{N, S}$ and of the labor supply shifters $\{\hat{\nu}_{i}\}_{i=1}^{N}$ and $\{\hat{\nu}_{g}\}_{g=1}^{G}$.\footnote{If, instead of (ref), we defined the first stage as simply the projection of $Y_{2i}$ onto the shift-share instrument, we could further relax this condition and only require $\{\hat{\chi}_s\}_{s=1}^{S}$ to be mean-independent of the labor supply shifters. An advantage of the current setup is that it allows us to derive primitive conditions for the consistency of the estimates of the first-stage regression and, thus, of the IV estimator.} For example, if $\{\hat{\chi}_s\}_{s=1}^{S}$ are sectoral productivity shocks, then these productivity shocks need to be independent of shocks to individuals' willingness to work in different groups and regions. (ref) requires that the coefficient on the instrument in the first-stage equation, which can be written as $\beta=\sum_{i=1}^{N}\sum_{s=1}^{S}w^{2}_{is}\operatorname{var}(\ensuremath\mathcal{X}_{s}\mid \mathcal{F}_{0})\beta_{is}/\sum_{i=1}^{N}\sum_{s=1}^{S}w^{2}_{is}\operatorname{var}(\ensuremath\mathcal{X}_{s}\mid \mathcal{F}_{0})$, is non-zero---this is the standard IV relevance assumption. For consistency and inference, in an analogy to the OLS case, we assume that (ref) holds with $\mathcal{F}_{0}=(\ensuremath\mathcal{Z}, U, Y_{1}(0), Y_{2}(0), B, W)$.
In a recent paper, \citet*{goldsmith2018bartik} explore a different approach to identification and inference on the treatment effect $\alpha$. Focusing here for simplicity on the case without controls, in place of (ref), they assume that the shares $\{w_{is}\}_{s=1}^{S}$ are as good as randomly assigned conditional on the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$; so that they are mean-independent of the potential outcomes $Y_{1}(0)$ and $Y_{2}(0)$ conditional on $\ensuremath\mathcal{X}$. As \citet*{goldsmith2018bartik} show, under this alternative assumption, one can replace the shift-share instrument $X_{i} = \sum_{s=1}^{S} w_{is} \ensuremath\mathcal{X}_s$ by the full vector of shares $(w_{i1}, \dotsc, w_{iS})$ in the first-stage equation. For estimation and inference, this alternative approach requires that, conditionally on the shifters, either the shares $(w_{i1}, \dotsc, w_{iS})$ or else the structural residuals be independent across regions or clusters of regions.
For estimating the inverse labor supply elasticity in the context of the model in (ref), (ref) illustrates that this alternative identification assumption requires that, conditional on $\{\hat{\chi}_{s}\}_{s=1}^{S}$, the region-specific employment shares in the initial equilibrium $\{l^{0}_{is}\}_{s=1}^{S}$ are mean-independent of both the region-specific exposure shares $\{\tilde{w}_{ig}\}_{g=1}^{G}$, and the region-specific labor supply shock $\nu_{i}$. This assumption is violated if regions more exposed to labor demand shocks in a sector $s$ (e.g.\ to changes in tariffs in the food sector) are also more exposed to labor supply shocks affecting workers of a group $g$ monras2018Mexico.\footnote{To allow for a shift-share component in the structural residual, \citet*{goldsmith2018bartik} view the shares $(w_{i1}, \dots, w_{iS})$ as “invalid” instruments, since, in this case, $E[\epsilon_{i}w_{{is}}\mid \ensuremath\mathcal{X}]\neq 0$, where $\epsilon_{i}$ denotes the structural error. \citet*{goldsmith2018bartik} show that if these shares are used to construct a single shift-share instrument $X_{i}$, the bias in the IV estimator coming from the correlation between any $w_{is}$ and the structural residual averages out under certain conditions as $S\to\infty$, as in the many invalid instrument setting studied in kolesar15invalid. Under the current setup, in contrast, (ref) implies that $X_{i}$ is a valid instrument for any fixed $S$. Leveraging exogeneity of $\ensuremath\mathcal{X}_{s}$ is a key difference between our approach and that in kolesar15invalid and \citet*{goldsmith2018bartik}. It allows us to do inference without imposing a particular correlation structure on the residuals $\epsilon_{i}$, and it allows us to achieve identification without requiring $S\to\infty$; the latter is only needed for consistency and inference.}
In terms of inference, since the structural residuals will not be independent across regions unless they contain no shift-share component (which, according to the economic model in (ref), is unlikely), the approach in \citet*{goldsmith2018bartik} generally requires that the shares are independent across (clusters of) regions. This assumption is, from the perspective of the model in (ref), conceptually very different from assuming independence of the shifters $\ensuremath\mathcal{X}_{s}$ across sectors. Since the shifters $\ensuremath\mathcal{X}_{s}=\hat{\chi}_{s}$ are exogenous, the latter only involves assumptions on model fundamentals by restricting the distribution in (ref). In contrast, each share $w_{is}=l_{is}^{0}$ corresponds to the employment allocation across sectors in a region $i$ in an initial equilibrium, so that the former involves imposing restrictions on an endogenous outcome of the model. Furthermore, since all the shares $\{w_{is}\}_{i=1,s=1}^{N, S}$ depend on the same set of sector-specific labor demand shifters $\{(\chi_{s}, \mu_{s})\}_{s=1}^{S}$, they will generally be correlated across regions.\footnote{For instance, if $\sigma_{s}=\sigma$ for all $s$, then $l_{is}^0 = D_{is}^0 / (\sum_{t=1}^{S}D_{it}^0)$, where $D^{0}_{is}$ is the labor demand shifter of sector $s$ in region $i$ in the initial equilibrium. According to (ref), for any $s$, all shifters $\{D^{0}_{is}\}_{i=1}^{N}$ depend on the same sector-level demand shocks, $\{(\chi_{s}, \mu_{s})\}_{s=1}^{S}$ and, thus, the labor shares $l_{is}^0$ will generally be correlated across all regions for any given sector.}
Which identification and inference approach is more attractive depends on the context of each particular empirical application. While the economic model in (ref) motivates the approach we pursue here, this does not mean that our approach is generally more attractive. In other empirical applications (e.g.\ when the shares are exogenous variables from the perspective of an economic framework), the approach of \citet*{goldsmith2018bartik} may be more appropriate.
It follows by adapting the arguments in the proof of (ref) that, if (ref) holds, and (ref) holds with $\mathcal{F}_{0}=(\ensuremath\mathcal{Z}, U, Y_{1}(0), Y_{2}(0), B, W)$, then, under mild technical regularity conditions (see
for details and proof),
where $\epsilon_{i}=Y_{1i}-Y_{2i}\alpha-Z_{i}'\delta$ is the residual in the structural equation, with $\delta=E[Z'Z]^{-1}E[Z'(Y_{1}-Y_{2}\alpha)]$. This suggests the standard error estimate
where $\widehat{\ensuremath\mathcal{X}}_{s}$ is constructed as in (ref), $\hat{\epsilon}=Y_{1}-Y_{2}\hat{\alpha}-Z'(Z'Z)^{-1}Z'(Y_{1}-Y_{2}\hat{\alpha})$ is the estimated residual of the structural equation, and $\hat{\beta}=\sum_{i=1}^{N}\ddot{X}_{i}Y_{2i}/\sum_{i=1}^{N}\ddot{X}_{i}^{2}$ is the first-stage coefficient.
The difference between the IV standard error formula in (ref) and the OLS version in (ref) is analogous to the difference between IV standard errors and OLS heteroskedasticity-robust standard errors for the corresponding reduced-form specification: the residual $\hat{\epsilon}_{i}$ corresponds to the residual in the structural equation, and the denominator is scaled by the first-stage coefficient. To obtain the IV analog of the standard error estimator under the null $H_{0}\colon \alpha=\alpha_{0}$, we use the formula in (ref) except that, instead of $\hat{\epsilon}_{i}$, we use the structural residual computed under the null, $\hat{\epsilon}_{\alpha_{0}}=(I-Z'(Z'Z)^{-1}Z')(Y_{1}-Y_{2}\alpha_{0})$. The resulting confidence interval is a generalization of the andersonrubin1949 confidence interval (which assumes that the structural errors are independent). For this reason, this confidence interval will remain valid even if the shift-share instrument is weak.
We now discuss three extensions to the basic setup. In (ref), we relax the assumption that the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$ are independent, allowing them to be correlated within clusters of sectors. (ref) generalizes our results to settings in which we have multiple observations for each region. (ref) considers the case in which the shifters are not directly observed, and have to be estimated.
Suppose that the sectors can be grouped into larger units, which we refer to as “clusters”, with $c(s)\in\{1,\dotsc, C\}$ denoting the cluster that sector $s$ belongs to; e.g., if each $s$ corresponds to a four-digit industry code, $c(s)$ may correspond to a three-digit code. With this structure, we replace (ref) with the weaker assumption that, conditional on $\mathcal{F}_{0}$, the shocks $\ensuremath\mathcal{X}_{s}$ and $\ensuremath\mathcal{X}_{k}$ are independent if $c(s)\neq c(k)$, and we replace (ref) with the assumption that, as $C\to\infty$, the largest cluster makes an asymptotically negligible contribution to the asymptotic variance; i.e. $\max_{c}\tilde{n}_{c}^{2}/\sum_{d=1}^{C}\tilde{n}_{d}^{2}\to 0$, where $\tilde{n}_{c}=\sum_{s=1}^{S}\1{c(s)=c}n_{s}$ is the total share of cluster $c$.
Under this setup, by generalizing the arguments in (ref), one can show that, as $C\to\infty$,
and, assuming that $\beta_{is}=\beta$ for every region and sector, the term $\mathcal{V}_{\ensuremath{N}}$ is now given by
As a result, we replace the standard error estimate in (ref) with a version that clusters $\widehat{\ensuremath\mathcal{X}}_{s}\hat{R}_{s}$,
where $\widehat{\ensuremath\mathcal{X}}_{s}$ is defined as in (ref). Confidence intervals with the null imposed can be constructed as in (ref), replacing $\hat{\epsilon}_{i}$ with $\hat{\epsilon}_{\beta_{0}, i}$ in (ref). In the IV setting considered in (ref), the standard error for $\hat{\alpha}$ is analogous to that in (ref), except that $\hat{\epsilon}_{i}$ denotes the residual in the structural equation, and we divide the expression by the absolute value of the first-stage coefficient, $\sum_{i=1}^{N}\ddot{X}_{i}Y_{2i}/\sum_{i=1}^{N}\ddot{X}_{i}^{2}$.
Consider a setting with $j=1, \dotsc, J$ regions, $k=1, \dotsc, K$ sectors, and $t=1, \dotsc, T$ periods. For each period $t$, we have data on shifters $\{\ensuremath\mathcal{X}_{kt}\}_{k=1}^{K}$, outcomes $\{Y_{jt}\}_{j=1}^{J}$, and shares $\{w_{jkt}\}_{j=1,k=1}^{J, K}$. This setup maps into the potential outcome framework in (ref) if we identify a “sector” with a sector-period pair $s=(k, t)$, and a “region” with a region-period pair $i=(j, t)$, so that we can index outcomes and shifters as $Y_{i}=Y_{jt}$ and $\ensuremath\mathcal{X}_{s}=\ensuremath\mathcal{X}_{kt}$, with the shares given by
If the shifters $\ensuremath\mathcal{X}_{kt}$ are independent across time and sectors, (ref) immediately give the large-sample distribution of the OLS estimator. In general, however, it will be important to allow the shifters $\ensuremath\mathcal{X}_{kt}$ to be correlated across time within each sector $k$. In this case, one can use the clustered standard error derived in (ref) by grouping observations over time for each sector $k$ into a common cluster, so that $c(k, t)=c(k', t')$ if $k=k'$. We can then apply the formula in (ref) to allow for any arbitrary time-series correlation in the sector-level shocks $\ensuremath\mathcal{X}_{kt}$ for any given sector $k$. Regardless of whether the sector-period pairs $(k, t)$ are clustered, as discussed in (ref), our standard error formulas allow for arbitrary dependence patterns in the regression residuals---in particular, they account for potential serial dependence in the regression residuals.
If the shift-share regressor is used as an IV in a regression of an outcome $Y_{1jt}$ onto a treatment $Y_{2jt}$, the mapping to (ref) is analogous, and one can use an IV version of the formula in (ref) for inference.
We now consider a setting in which the sectoral shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$ that define the shift-share IV studied in (ref) are not directly observed. We follow the setup in (ref) but assume that, instead of observing $\ensuremath\mathcal{X}_{s}$ directly, we only observe a noisy measure of it,
for each sector-region pair. We consider IV regressions that use two different estimates of $X_{i}=\sum_{s=1}^{S}w_{is}\ensuremath\mathcal{X}_{s}$. First, an estimate that replaces $\ensuremath\mathcal{X}_{s}$ with an estimate $\hat{\ensuremath\mathcal{X}}_{s}=\sum_{i=1}^{N}\check{w}_{is}X_{is}/\check{n}_{s}$, where $\check{n}_{s}=\sum_{i=1}^{N}\check{w}_{is}$ and the weights $\check{w}_{is}$ are not necessarily related to $w_{is}$. The resulting estimate of $X_{i}$ is
and it yields the IV estimate $\tilde{\alpha}=\ddot{\hat{X}}'Y_{1}/\ddot{\hat{X}}'Y_{2}$, where $\ddot{\hat{X}}=\hat{X}-Z(Z'Z)^{-1}Z'\hat{X}$ is the residual from regressing $\hat{X}_{i}$ onto $Z_{i}$. Second, we consider the leave-one-out estimator
where $\hat{\ensuremath\mathcal{X}}_{s, -i}=\sum_{j=1}^{N}\1{j\neq i}\check{w}_{js}X_{js}/\check{n}_{s, -i}$ is an estimate of $\ensuremath\mathcal{X}_{s}$ that excludes region $i$. A version of this estimator has been used in autorduggan2003. This leave-one-out estimator of the shift-share instrument $X_{i}$ yields the IV estimate $\hat{\alpha}_{-}=\ddot{\hat{X}}_{-}'Y_{1}/\ddot{\hat{X}}_{-}'Y_{2}$, where $\ddot{\hat{X}}_{-}=\hat{X}_{-}-Z'(Z'Z)^{-1}Z'\hat{X}_{-}$.
While we assume that $\ensuremath\mathcal{X}_{s}$ satisfies the exogeneity restriction in (ref) for every $s$, we allow the measurement errors $\psi_{i}=(\psi_{i1}, \dotsc, \psi_{iS})'$ to be potentially correlated with the potential outcomes $Y_{1i}(0)$ and $Y_{2i}(0)$ in the same region $i$. We assume, however, that $\psi_{i}$ is independent of the errors $\psi_{j}$ and of the potential outcomes $Y_{1j}(0)$ and $Y_{2j}(0)$ for any region $j\neq i$ (see
for a formal statement). In
, we use the model in (ref) to discuss these assumptions in the context of estimating the inverse labor supply elasticity.\footnote{Specifically, we show in
that, if $X_{is}$ corresponds to employment growth rates, then $\psi_{i}$ will generally not be independent of $(\psi_{j}, Y_{1j}(0), Y_{2j}(0))$ in others regions $j\neq i$, unless one makes restrictive assumptions about the demand elasticities $\sigma_{s}$, such as $\sigma_{s}=0$. We also construct alternative shift-share IVs that satisfy this independence assumption under weaker restrictions on $\sigma_{s}$, but require adjusting the shifter used in estimation.}
The potential correlation between $\psi_{i}$ and the potential outcomes in region $i$ implies that the estimation error in $\hat{X}_{i}$, which is a function on $\psi_{i}$, may be correlated with the residual in the structural equation. Thus, including the $i$th observation in the construction of $\hat{X}_{i}$ induces an own-observation bias in the IV estimator $\tilde{\alpha}$ of $\alpha$. See \citet*{goldsmith2018bartik} and \citet*{borusyakhulljaravel2018shiftshare} for a discussion. This bias is analogous to the bias of the two-stage least squares estimator in settings with many instruments (e.g.\ \citealp*{bekker_alternative_1994,angrist_jackknife_1999}), such as when one uses group indicators as instruments.\footnote{See, e.g., \citet*{maestas2013,DoSo15,AiDo15}, or Silver16.} We show in
that the magnitude of the bias is of the order $\frac{1}{\ensuremath{N}} \sum_{i=1}^{N}\sum_{s=1}^{S}\frac{w_{is}\check{w}_{is}}{\check{n}_{s}}\leq S/\ensuremath{N}$, so that consistency of $\tilde{\alpha}$ generally requires the number of sectors to grow more slowly than the number of regions. Furthermore, to ensure that the asymptotic bias in $\tilde{\alpha}$ does not induce undercoverage of the resulting confidence intervals, one generally requires $S^{3/2}/N\to 0$.
The estimator $\hat{\alpha}_{-}$, which can be thought of as a shift-share analog of the jackknife IV estimator studied in \citet*{angrist_jackknife_1999}, remains consistent, as shown in \citet*{borusyakhulljaravel2018shiftshare} and in
. We also show in this appendix that, under regularity conditions, its asymptotic distribution is given by
with $\mathcal{V}_{\ensuremath{N}}$ defined as in (ref), and
The term $\mathcal{W}_{\ensuremath{N}}$ accounts for the additional uncertainty stemming from the fact that the shift-share IV is estimated. It is analogous to the many-instrument term in the jackknife IV estimator under many instrument asymptotics (see chao_asymptotic_2012). Using simulations, we show in
several designs in which, while correcting for the own-observation bias by using $\hat{\alpha}_{-}$ instead of $\tilde{\alpha}$ is quantitatively important, accounting for the additional variance term $\mathcal{W}_{N}$ is less important.
In (ref), we revisit the placebo exercise in (ref) to examine the finite-sample properties of the inference procedures described in (ref). In (ref), we show that our baseline placebo results are robust to several changes in the placebo design.
We first consider the performance of the standard error estimator in (ref) (which we label AKM), and the standard error and confidence interval in (ref) (with label AKM0) in the baseline placebo design described in (ref).\footnote{We fix the matrix $Z$ to be a column of ones when implementing the formulas in (ref).}$^{,}$\footnote{In
, we explore the sensitivity of our results to using counties (instead of CZs) as the regional unit of analysis, and occupations (instead of sectors) as the unit at which the shifter is defined.}
For the AKM and AKM0 inference procedures, (ref) presents median standard error estimates and rejection rates for 5% significance level tests of the null hypothesis $H_{0}\colon\beta=0$. In the case of AKM0, since the standard error depends on the null being tested, the table reports the median “effective standard error”, defined as the length of the 95% confidence interval divided by $2\times 1.96$.
The results in (ref) show that the inference procedures introduced in (ref) perform well. The median AKM standard error is slightly lower than the standard deviation of $\hat{\beta}$, by about 5% on average across all outcomes. The median AKM0 effective standard error is slightly larger than the standard deviation of $\hat{\beta}$, by about 11% on average. The implied rejection rates are close to the 5% nominal rate: the AKM procedure has rejection rates between 7.5% and 9.1% and the AKM0 rejection rates are always between 4.3% and 4.5%. As discussed in (ref), the AKM and AKM0 confidence intervals are asymptotically equivalent. The differences in rejection rates between the \textit{AKM} and \textit{AKM0} inference procedures are thus due to differences in finite-sample performance. As noted in other contexts (see, e.g., lazarus2018har), imposing the null can lead to improved finite-sample size control. The better size control of the \textit{AKM0} procedure is consistent with these results.
In (ref), we show theoretically that the AKM and AKM0 inference procedures are valid in large samples only if: (a) the number of sectors goes to infinity; (b) all sectors are asymptotically “small”; (c) the sectoral shocks are independent across sectors. Given these conditions, these inference procedures remain valid under (d) any distribution of the sectoral shifters; and (e) arbitrary correlation structure of the regression residuals. In this section, we evaluate the sensitivity of these inference procedures to requirements (a) to (c) above, and illustrate points (d) and (e) by documenting the robustness of these procedures to alternative distributions of the shifters and the residuals. In all cases, we also report Robust and Cluster standard errors estimates and rejection rates. We focus on the change in the share of working-age population employed as the outcome variable of interest.
We first evaluate how the performance of different inference procedures depends on the number of sectors. Panel A of (ref) shows that the overrejection problem affecting standard inference procedures worsens when the number of sectors decreases: the rejection rates of 5% significance level tests based on Robust and Cluster standard errors reach 70.6% and 56.1%, respectively, when we construct the shift-share covariate using 20 2-digit SIC sectors (instead of the 396 4-digit SIC sectors we use in the baseline placebo). In line with the findings of the literature on clustered standard errors with few clusters, the rejection rates of hypothesis tests that rely on AKM standard errors also increase to 12%, but rejection rates for hypothesis tests that apply the AKM0 inference procedure remain very close to the nominal 5% significance level.
Panels B to D of (ref) examine the robustness of the results in (ref) to alternative distributions of the shifters. In Panel B, as in our baseline placebo exercise, the shifters are drawn i.i.d.\ from a normal distribution, but we change the variance to both a lower ($\sigma^{2}=0.5$) and a higher value ($\sigma^{2}=10$) than in the baseline ($\sigma^{2}=5$). In Panel C, we draw the shifters from a log-normal distribution re-centered to have mean zero and scaled to have the same variance as in the baseline. Panel D investigates the robustness of our results to heteroskedasticity in the sector-level shocks. We set variance of the shock in each sector $s$, to $\sigma^{2}_{s}=5 + \lambda(n_{s} - S/N)$. Thus, the cross-sectional average of the variance of the sector-level shocks is the same as in the baseline (which corresponds to setting $\lambda=0$), but this variance now varies across sectors. Comparison of the results in Panels B to D of (ref) to those in (ref) suggests that our baseline results are not sensitive to specific details of the distribution of sector-level shifters. This is consistent with the claim (d) above.
Panels E and F of (ref) explore the robustness of our baseline results to different patterns of correlation in the regression residuals. In the baseline placebo, since $\beta=0$, the regression residuals inherit the correlation patterns in the outcome variable. Here, we modify these patterns by adding a random shock $\eta^{m}_{i}$ in each placebo sample $m$ to the outcome $Y_{i}$. Panel E explores the impact of increasing the correlation between the regression residuals of CZs that belong to the same state. Specifically, we generate a random variable $\tilde{\eta}_{k}^m$ for each state $k$ and simulation $m$ such that $\tilde{\eta}_{k}^m \sim \mathcal{N}(0,6)$. We then set $\eta_{i}^m = \tilde{\eta}_{k(i)}^m$ where $k(i)$ is the state of CZ $i$. Since we have now increased the relative importance of the correlation pattern accounted for by Cluster standard errors, the resulting overrejection decreases from 38.3% to 30.4%. In line with claim (e) above, the rejection rates of the AKM and AKM0 inference procedures are not affected. In Panel F, we evaluate the robustness of our results to adding a shock to the non-manufacturing sector that is included in the regression residual. Specifically, in each simulation $m$, we set $\eta_i^m = (1-\sum_{s=1}^{S}w_{is})\hat{\eta}^m_{S}$ with $\hat{\eta}^m_{S} \sim \mathcal{N}(0,5)$, where $\sum_{s=1}^{S}w_{is}$ is the 1990 aggregate employment share of the 396 4-digit SIC manufacturing sectors included in the definition of the shift-share regressor of interest. The results in Panel F of (ref) show that adding this component to the regression residual does not affect the rejection rates.
Lastly, Panel G in (ref) explores the consequences of adding the non-manufacturing sector to the shift-share regressor. In Panel F, the shock to the non-manufacturing sector is part of the regression residual; in Panel G, we use this shock, in combination with the shocks to all manufacturing sectors, to construct the shift-share regressor. Across CZs, the average initial employment share in the non-manufacturing sector is 77.5%; i.e. $N^{-1}\sum_{i=1}^{N}(1-\sum_{s=1}^{S}w_{is})=77.5\%$. Including such a large sector in the shift-share regressor violates (ref). As a result, the AKM and AKM0 inference procedures overreject severely; standard inference procedures fare even worse, with rejection rates reaching up to 92%. The results in Panels F and G suggest that, provided that the shifters are independent across sectors, it is better to exclude large sectors from the shift-share regressor of interest, and thus let the shocks associated with them enter the regression residual. One should, however, bear in mind that, if $\beta_{is}$ in (ref) varies across sectors, excluding large sectors from the shift-share regressor will change the estimand $\beta$ (see (ref)).
In the placebo simulations described in (ref), we have drawn the shifters independently from a mean-zero distribution. In (ref), we allow for non-zero correlation in the shifters within “clusters” of sectors.\footnote{In
, we study the impact of drawing the shifters from a distribution with non-zero mean. We show that, in line with the discussion in (ref), it is important to control for the region-specific sum of shares $\sum_{s=1}^{S}w_{is}$.} Specifically, we report results from placebo exercises in which the shifters are drawn from the joint distribution $(\ensuremath\mathcal{X}_{1}^m, \dotsc, \ensuremath\mathcal{X}_{S}^m) \sim \mathcal{N}\left(0,\Sigma\right)$, where $\Sigma$ is an $S\times S$ covariance matrix with elements $\Sigma_{sk}=(1-\rho)\sigma\1{s=k}+\rho\sigma\1{c(s)=c(k)}$ and $c(s)$ indicates the “cluster” that industry $s$ belongs to. In panels A, B, and C, these clusters correspond to the 3-, 2-, and 1-digit SIC sector that the 4-digit SIC sector $s$ belongs to, respectively.
\afterpage{
}
Panel A of (ref) shows that introducing correlation within 3-digit SIC sectors has a moderate effect on the rejection rates of both the traditional methods and versions of the AKM and AKM0 methods that assume that the sectoral shocks are independent. Rejection rates close to 5% are obtained with versions of the AKM and AKM0 inference procedures that cluster the shifters at a 2-digit SIC level (see (ref)). As shown in Panel B, the overrejection problem affecting both traditional inference procedures and versions of the AKM and AKM0 procedures that assume independence of shifters is more severe when the shifters are correlated at the 2-digit level. However, the last two columns show that, in this case, the versions of \textit{AKM} and \textit{AKM0} that cluster the sectoral shocks at the 2-digit level achieve rejection rates close to the nominal level. Finally, Panel C shows that the overrejection problem is much more severe in the presence of high correlation in shifters within the two 1-digit aggregate sectors, and this problem is not solved by clustering at the 2-digit level.
The last panel in (ref) illustrates the inferential problems that arise in empirical applications of shift-share designs when all shifters are correlated with each other. Such correlations also arise, for example, when all shifters are generated (at least in part) by a common shock with potentially heterogeneous effects across sectors.\footnote{There is an extensive empirical literature documenting the importance of common factors driving changes in sector-specific variables such as sectoral industrial production, employment and value added \citep*[see, e.g.,][]{altonji_comovement_1990,shea_2002,foerster_aggshocks_2011}.} As simulations presented in
illustrate, if there is a common component affecting all shifters, it is important to first estimate this common component and to control for it in the shift-share regression of interest. Otherwise, hypothesis tests based on standard inference procedures as well as on the AKM and AKM0 inference procedures may suffer from an overrejection problem.
We summarize the conclusions from (ref) in the following remark.
In
we present results from additional placebo simulations in which we investigate the consequences of: (a) the violation of the assumption that the shifters of interest are as good as randomly assigned; (b) the presence of serial correlation in both the shifters of interest and the regression residuals, in panel data settings; (c) the true potential outcome function being nonlinear, implying that the linearly additive potential outcome framework in (ref) is misspecified; (d) the presence in the regression residuals of shift-share components with shares correlated in different degrees with those entering the shift-share covariate of interest; and, (e) the presence of treatment heterogeneity across regions and sectors.
We now apply the AKM and AKM0 inference procedures to two empirical applications. First, the effect of Chinese competition on U.S. local labor markets, as in \citet*{autordornhanson2013}. Second, the estimation of the local inverse elasticity of labor supply, as in bartik1991benefits. Additionally, in
, we apply the AKM and AKM0 inference procedures to the study of the impact of immigration on labor market outcomes of U.S. natives.
\citet*[henceforth ADH]{autordornhanson2013}, explore the impact of exports from China on labor market outcomes across U.S. CZs. Specifically, ADH present IV estimates for a specification that fits within the panel data setting described in (ref), with each region $j=1,\dots,722$ denoting a CZ, each sector $k=1,\dots,396$ denoting a 4-digit SIC industry, and each period $t=1,2$ denoting either 1990--2000 changes or 2000--2007 changes. As in (ref), we index here the intersection of a region $j$ and a period $t$ by $i$, and the intersection of a sector $k$ and a period $t$ by $s$. In ADH, the outcome $Y_{1i}$ is a ten-year equivalent change in a labor-market outcome, the endogenous treatment is $Y_{2i} = \sum_{s=1}^{S} \bar{w}_{is} \ensuremath\mathcal{X}_s^{US}$, where $\ensuremath\mathcal{X}_{s}^{US}$ is the change in U.S. imports from China normalized by the start-of-period total U.S. employment in the sector, and $\bar{w}_{is}$ is the start-of-period employment share of a sector in a CZ\@. ADH use the shift-share IV $X_{i} = \sum_{s=1}^{S} w_{is} \ensuremath\mathcal{X}_s$, where $\ensuremath\mathcal{X}_s$ denotes imports from China by high-income countries other than the U.S. normalized by a ten-year-lag of the start-of-period total U.S.\ employment in the sector, and $w_{is}$ is the ten-year-lag of the employment share $\bar{w}_{is}$. To measure these variables, we use the data sources described in (ref). In all regression specifications, we include a vector of controls $Z_{i}$ corresponding to the largest set of controls used in ADH.\footnote{See column (6) of Table 3 in ADH\@. The vector $Z_{i}$ aims to control for labor supply shocks and labor demand shocks other than the changes in imports from China, and it includes the start-of-period percentage of employment in manufacturing. The discussion in (ref) implies that one should instead control for the ten-year-lagged of the start-of-period employment share in manufacturing, to match the shares that enter the definition of the shift-share IV\@. However, to facilitate the comparison with the original results in ADH, we use their vector of controls. As shown in \citet*{borusyakhulljaravel2018shiftshare}, controlling for the ten-year-lagged manufacturing employment shares does not substantively affect the estimates.}
(ref) reports 95% CIs computed using different methodologies for the specifications in Tables 5 to 7 in ADH\@. Panels A, B, and C present the IV, reduced-form and first-stage estimates, respectively. Following autor2014trade, the AKM and AKM0 CIs cluster the shifters $\{\ensuremath\mathcal{X}_{s}\}_{s=1}^{S}$ by 3-digit SIC industry; thus, the AKM and AKM0 CIs we report are robust to serial correlation in the shifters as well as to cross-sectoral correlation in the shifters within 3-digit SIC industries.
report AKM and AKM0 CIs for alternative definitions of clusters.
In
, we present placebo simulations that depart from our baseline placebo design in ways that explore specific features of the empirical setting studied in this section. In
, we draw the shifters from the empirical distribution of shifters used to construct the ADH IV (instead of drawing them from a normal distribution); the resulting rejection rates are very similar to those in the baseline simulation. In
, we draw shifters that have a common component with factor structure; since the resulting correlation structure cannot be captured by clustering, we show that it is important in this case to include an estimate of the common factor component as an additional control.\footnote{For placebo simulation evidence under our baseline assumption that the shifters are independent across 3-digit clusters, using data for outcomes $Y_{1i}$ and shares $w_{is}$ identical to that used in this section, see
.}
In (ref), state-clustered CIs are very similar to the heteroskedasticity-robust ones. In contrast, our proposed CIs are wider than those implied by state-clustered standard errors. For the IV estimates reported in Panel A, the average increase across all outcomes in the length of the 95% CI is 24% with the AKM procedure and 65% with the AKM0 procedure. When the outcome is the change in the manufacturing employment rate, the length of the 95% CI increases by 26% with the AKM procedure and by 65% with the AKM0 procedure. In light of the lack of impact of state-clustering on the 95% CI, the wider intervals implied by our inference procedures indicate that cross-region residual correlation is driven by similarity in sectoral compositions rather than by geographic proximity.
Panel B of (ref) reports CIs for the reduced-form specification. In this case, the increase in the CI length is slightly larger than for the IV estimates: across outcomes, it increases on average by 54% for AKM and 130% for AKM0. The smaller relative increase in the CI length for the IV estimate relative to its increase for the reduced-form estimate is a consequence of the fact that all inference procedures yield similar CIs for the first-stage estimate, as reported in Panel C.
As discussed in (ref), the differences between AKM (or AKM0) CIs and state-clustered CIs are related to the importance of shift-share components in the regression residual. The results in Panel C suggest that, once we account for changes in sectoral imports from China to other high-income countries, there is not much sectoral variation left in the first-stage regression residual; i.e., there are no other sectoral variables that are important to explain changes in sectoral imports from China to the U.S.\footnote{This is analogous to what we would observe in a regression in which the regressor of interest varies at the state level, and we control for all state-specific covariates affecting the outcome variable: state-clustered standard errors would be similar to heteroskedasticity-robust standard errors, since there is little within-state correlation left in the residuals.} To investigate this claim,
reports the rejection rates implied by a placebo exercise designed to match the first-stage specification reported in Panel C of (ref). The placebo results show that, while traditional methods still suffer from severe overrejection when no controls are included, the overrejection is attenuated once we include as controls the shift-share IV and the control vector $Z_{i}$ we use in (ref), indicating that these variables soak up much of the cross-CZ correlation in the treatment variable used in ADH\@.
Overall, (ref) shows that, despite the wider confidence intervals obtained with our procedures, the qualitative conclusions in ADH remain valid at usual significance levels. However, the increased width of the 95% CI shows that the uncertainty regarding the magnitude of the impact of Chinese import exposure on U.S. labor markets is greater than that implied by usual inference procedures. In particular, the AKM0 CI is much wider than that based on state-clustered standard errors; furthermore due to its asymmetry around the point estimate, using the AKM0 CI, we cannot rule out impacts of the China shock that are two to three times larger than the point estimates of these effects.\footnote{It follows from (ref) (see the expression for the quantity $A$) that the asymmetry in the AKM0 CI comes from the correlation between the regression residuals $\hat{R}_{s}$ and the shifters cubed. In large samples, this correlation is zero and the AKM and AM0 CIs are asymptotically equivalent. The differences between both CIs in (ref) thus reflect differences in their finite-sample properties. This notwithstanding, the placebo exercise presented in
shows that both inference procedures yield close to correct rejection rates in a sample analogous to that used in ADH.}
In our second application, we estimate the inverse labor supply elasticity. Specifically, using the notation of (ref), we estimate the parameter $\tilde{\phi}$ in the equation
where $\hat{L}_{i}$ denotes the log change in the employment rate in CZ $i$, $\hat{\omega}_{i}$ denotes the log change in wages, $Z_{i}$ is a vector of controls, and $\epsilon_{i}$ is a regression residual. We use the same sample, data sources, and vector of controls $Z_{i}$ as in (ref).\footnote{
investigates the robustness of our results to alternative sets of controls.}
The model in (ref) has implications for the properties of different strategies for estimating the inverse labor supply elasticity $\tilde{\phi}$. By (ref), the residual $\epsilon_{i}$ in (ref) accounts for changes in labor supply shocks, $\sum_{g=1}^G \tilde{w}_{ig} \hat{\nu}_g + \hat{\nu}_i$, not controlled for by the vector $Z_{i}$. Second, it follows from (ref) that, up to a first-order approximation around an initial equilibrium, changes in regional employment rates, $\hat{L}_{i}$, can be written as a function of both shift-share aggregators of sectoral labor demand shocks and the same labor supply shocks potentially entering $\epsilon_{i}$ in (ref), $\sum_{g=1}^G \tilde{w}_{ig} \hat{\nu}_g + \hat{\nu}_i$. Thus, $\hat{L}_i$ and $\epsilon_{i}$ will generally be correlated and the OLS estimator of $\tilde{\phi}$ in (ref) will be biased. However, as discussed in (ref), the model in (ref) also implies that we can instrument for $\hat{L}_i$ using shift-share aggregators of sectoral labor demand shocks that are independent of the unobserved labor supply shocks (see
for more details).
In this section, we use three different shift-share IVs to estimate $\tilde{\phi}$ in (ref). For each of them, (ref) presents the reduced-form, first-stage and 2SLS estimates. First, in Panel A, we use the instrumental variable in bartik1991benefits; i.e. $\hat{X}_{i}=\sum_{i=1}^{N}w_{is}\hat{L}_{s}$, where $\hat{L}_{s}$ denotes the nation-wide employment growth in sector $s$. Second, in Panel B, we use the leave-one-out version of this instrument; i.e. $\hat{X}_{i}=\sum_{i=1}^{N}w_{is}\hat{L}_{s, -i}$, where $\hat{L}_{s, -i}$ denotes the employment growth in sector $s$ over all CZs excluding CZ $i$.\footnote{The leave-one-out version of the instrument in bartik1991benefits was originally proposed by autorduggan2003. In Online Appendix
, we clarify the assumptions under which the model in (ref) is consistent with the validity of the leave-one-version of the Bartik IV\@.
presents placebo exercises attesting that the AKM and AKM0 CIs reported in this section have appropriate coverage in the context of this empirical application.} Third, in Panel C, we use the IV used in \citet*{autordornhanson2013}, which we denote as ADH IV and describe in detail in (ref).\footnote{The effect of these IVs on the changes in the employment rate may be heterogeneous across regions and sectors (see (ref)). This does not affect the validity of our inference procedures since, as discussed in (ref), we allow for heterogeneous effects in the first-stage regression.} As in (ref), we report versions of the AKM and AKM0 CIs with shifters clustered at the 3-digit SIC industry for all periods.
Column (3) of (ref) shows that the estimates of the inverse labor supply elasticity are similar no matter which IV we use: 0.80 when using the original Bartik IV, 0.82 when using the leave-one-out version of this estimator, and 0.67 when using the ADH IV\@.\footnote{One explanation for the similarity between the leave-one-out and the original Bartik IV is that, as discussed in (ref), the bias of the original Bartik IV is of the order $\frac{1}{\ensuremath{N}} \sum_{i=1}^{N}\sum_{s=1}^{S}\frac{w_{is}\check{w}_{is}}{\check{n}_{s}}$. This quantity equals $0.004$ in this application, indicating that the own-observation bias is likely to be small.} In both Panel A and Panel B, the AKM and AKM0 CIs are very similar to the state-clustered CI\@. In Panel C, the AKM and AKM0 CIs are only moderately wider than those obtained with state-clustered standard errors.
Columns (1) and (2) of (ref) show the first-stage and reduced-form estimates, respectively. In Panel A and Panel B, the AKM0 CIs are similar to the state-clustered CIs; in contrast, in Panel C, the first-stage and reduced-form AKM0 CIs more twice as wide, and more than three times as wide as the state-clustered CI, respectively. Thus, the first-stage and reduced-form AKM and AKM0 CIs differ more from the state-clustered CI when the ADH IV is used than when the Bartik IV is used. A possible explanation for this finding is that the shift-share component of the first-stage and reduced-form regression residuals is much smaller in the latter than in the former case. The Bartik IV absorbs the bulk of the shift-share covariates that affect the change in the employment rate and wages across CZs. In contrast, the ADH IV is just one of the possibly various shift-share terms affecting the change in the outcome and endogenous treatment of interest. With the remaining shift-share entering the regression residual, it becomes quantitatively important to use our inference procedures to obtain CIs with the right coverage.
This paper studies inference in shift-share designs. We show that standard economic models predict that changes in regional outcomes depend on observed and unobserved sector-level shocks through several shift-share terms. Our model thus implies that the residual in shift-share regressions is likely to be correlated across regions with similar sectoral composition, independently of their geographic location, due to the presence of unobserved shift-share terms. Such correlations are not accounted for by inference procedures typically used in shift-share regressions, such as when standard errors are clustered on geographic units. To illustrate the importance of this shortcoming, we conduct a placebo exercise in which we study the effect of randomly generated sector-level shocks on actual changes in labor market outcomes across CZs in the United States. We find that traditional inference procedures severely overreject the null hypothesis of no effect. We derive two novel inference procedures that yield correct rejection rates.
It has become standard practice to report cluster-robust standard errors in regression analysis whenever the variable of interest varies at a more aggregate level than the unit of observation. This practice guards against potential correlation in the residuals that arises whenever these residuals contain unobserved shocks that also vary at the same level as the variable of interest. In the same way, we recommend that researchers report confidence intervals in shift-share designs that allow for a shift-share structure in the residuals, such as one of the two confidence intervals that we propose.
University of Chicago Booth School of Business\\ Princeton University\\ Princeton University