Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,821 characters · 24 sections · 12 citation commands
Relaxing the Exclusion Restriction in Shift-Share Instrumental Variable Estimation
\pagenumbering{gobble}
Keywords: Causal inference, Invalid instruments, Lasso, Shift-Share instrument \\ JEL classification: C36, C52, F22, F66
\pagenumbering{arabic}
The shift-share instrument is often used in applied economics to obtain estimates of causal effects. The numerous applications have spanned three decades, beginning with \citeauthor*{Bartik1991Who}'s\ (Bartik1991Who) seminal paper, and included several fields, such as migration, labor, international economics and many other. A shift-share strategy exploits shares at an earlier point in time and current aggregate-level changes to create an instrumental variable (IV). For example, the share of migrants from a certain origin country is interacted with the inflow from that country. The methods proposed in this paper are not restricted to the examples mentioned here, but can be applied to a wide range of studies in which shift-share instruments are used.
The key assumptions and properties of shift-share instruments have been investigated only in very recent work. Centrally, in an important setting the exclusion restriction must hold for all initial shares, for the estimator to be consistent. In practice, a potentially large number of shares cannot have a direct effect on the outcome. This assumption is very strict, because it requires the researcher to have perfect structural knowledge about all shares, which is typically unavailable. The natural question to ask, therefore is: Is consistent estimation still possible when the exclusion restriction is violated for some but not all shares?
In this paper, my main contribution is to show how consistent shift-share estimation is possible, when not all shares fulfil the exclusion restriction. This paper is a practitioner's guide on how to select invalid shares in the shift-share setting, using two methods developed in statistical learning. I also extend one of the methods that I present to allow for multiple endogenous regressors. This was not possible so far and is a further contribution of the paper.
The proposed methods go beyond existing econometric diagnostics typically applied in this setting. Rotemberg weights, proposed by \citet*{Goldsmith-Pinkham2020Bartik}, report the sensitivity to misspecification of the different shares. These weights often fail to provide clear-cut guidance as to which shares should finally be included in the construction of the instrument, because they tell the researcher how large the relative bias of the entire shift-share estimator stemming from the bias of a single industry is. Instead, with the methods proposed in this paper the researcher obtains an estimate for the identity of valid and invalid instruments and a consistent estimate, adjusted from the absolute bias. This bears the advantage that valid instruments do not need to be discarded just because their potential invalidity could lead to bias.
Applying the new methods to the estimation of the effects of immigration and of Chinese import exposure on the US labor market illustrates that the shares selected as invalid in these applications are consistent with those discussed as problematic in the literature. So far, no way to locate invalid shares has been proposed. This paper fills this gap and provides a principled approach to share selection. To make the methods more widely accessible to practitioners, I also provide simple-to-use Stata-programs.
I begin by presenting the shift-share IV and its key identifying assumption, the exclusion restriction. In the shift-share approach, there are multiple class-specific shares and shifts which are interacted to produce the final instrument. \citet*{Goldsmith-Pinkham2020Bartik} show that if the exclusion restriction holds for each class-specific share, the IV estimator is consistent. That is, shares should not be directly correlated with the outcome variable through unobservable shocks or longterm effects. The exclusion restriction from the shares perspective is a sufficient condition for the consistency of the shift-share IV estimator. Instruments which fulfil the exclusion restriction are called valid, while those that do not fulfil it are called invalid. This definition of validity assumes that all instruments are related to the treatment. This exclusion restriction is very restrictive because it must hold for all classes. While the general idea behind the shift-share IV is credible, typically structural relationships between instruments and outcome variables are difficult to exclude for each single class.
In Section (ref), I show how to obtain consistent estimators, when many shares are invalid. To achieve consistency, invalid shares are selected using two methods: the adaptive Least absolute shrinkage and selection operator (AL) by \citet*{Windmeijer2019Use} and the Confidence Interval Method (CIM) by \citet*{Windmeijer2020Confidence}. The two methods have been developed primarily for the use in Mendelian randomization, which is the application of instrumental variable estimation in genetic epidemiology. In these applications, genetic markers are used as IVs when estimating the effect of an exposure on a health outcome.
The key advantage of the leveraged methods is that they consistently select shares, which violate the exclusion restriction. When the majority of instruments is valid, both metods have so-called oracle properties and when the largest group of shares fulfills the exclusion restriction the CIM has oracle properties. This means that asymptotically the post-selection estimators perform as well as if the researcher knew the identity of invalid IVs. Intuitively, these estimators both exploit the fact that just-identified estimates, which use only one valid IV at a time, converge against the same value. None of the existing methods allow for multiple endogenous regressors. I therefore propose a simple extension of the AL. Here, the exclusion restriction becomes stricter with increasing number of endogenous regressors.
To show the implications and generality of the presented methods in practice, in Sections (ref) and (ref), I apply them to two empirical examples. I apply the methods to the estimation of the effect of immigration on wages in the US as in \citet*{Altonji1991effects} and \citet*{Card2001Immigrant} and of Chinese import exposure on manufacturing employment, following \citet*{Autor2013china}. These two empirical examples are representative for a long series of applications in international and migration economics, which rely heavily on shift-share IVs.
The use of the new methods suggests a lower effect of immigration on wages in the US. Using data from the US, the coefficients for the estimates that do not account for endogenous shares are positive. Using the estimators which adjust for shares selected as invalid, the estimates become smaller, often even change sign and retain statistical significance, when they were significant in the standard estimations. For example the effects on high-skilled wages change from 0.52 to -0.53 ($p < 0.01$). Among the selected there are many countries which were suspected of invalidity in the literature, such as the Philippines \citep*{Card2009Immigration}. When using a model with lagged immigration, the standard shift-share analysis produces estimates which point in different directions than expected. When using the proposed extension, the coefficients get switched in the expected direction again. Overall, the proposed methods seem to induce a large qualitative difference and should be used as a robustness check in migration settings.
When estimating the effect of import competition on employment, a large number of instruments is selected as invalid in some settings. It is noteworthy that many of the industry classes which have been discussed as problematic in \citet*{Autor2013china} and are likely to affect estimates \citep*{Goldsmith-Pinkham2020Bartik} are selected as invalid.
These two applications illustrate the value of the proposed methods. The identity of chosen shares is consistent with economic intuition, because many of the origin countries and industries chosen as invalid are also discussed as being potentially problematic in the literature. This speaks for the plausibility of the outcomes of the new methods. Also, many shares which are similar to the discussed groups but were not specifically pointed out as being problematic in the literature, have been chosen by the new methods. Results can change qualitatively, when using the adjusted estimators. This shows that the methods are complementary to economic intuition and can help in designing an appropriate shift-share instrument.
In an online Appendix I show an extension of the method to the multiple regressor case, a summary and illustrations of the methods as well as simulations confirming that in applications with shares as IVs with increasing sample size the performance approaches that of the estimator which uses only valid shares. This holds already for relatively small sample sizes. Second, allowing for weak instruments and stronger direct effects of the instruments on the outcome does not change the fact that the estimators converge to oracle performance quickly. In a third set of simulations, I test the performance of my extension to multiple endogenous regressors. The simulations confirm that with increasing number of regressors, the allowed number of invalid IVs becomes lower.
This paper relates to two strands of the literature. Recent work has brought forward two ways of motivating shift-share research designs: one which is justified by quasi-random shocks and one which stresses the exclusion restriction for shares only. \citet*{Borusyak2020Quasi} show consistency of the shift-share estimator when shocks are quasi-randomly assigned, conditionally on the shares. Also in this setting, \citet*{Adao2019Shift} discuss issues in inference. Unlike these two papers, \citet*{Goldsmith-Pinkham2020Bartik} put forward an interpretation of shift-share designs, according to which the shares express differential exposure to common shocks. The identification of the causal effect relies on the exclusion restriction for shares. Which setting is appropriate depends on the economic question at hand.
This paper mainly relates to the exogenous share setting. Still, in situations in which it comes natural to think of the instrument from the perspective of random shifts, when many class-specific shifts are available these can be used to create multiple shift-share IVs or can be used as IVs directly. The selection methods then select among them instead of among the shares. For example, when imports are available for several high-income countries, in the \citet*{Autor2013china} example, the shifts can be used separately in a regression where the observations are industry-specific.
The approach that I propose is helpful when both of these approaches fail. When the shifts are not random, consistency as derived by \citet*{Borusyak2020Quasi} does not follow. When the identifying assumption is motivated through the shares, but some of the class-specific shares are subject to criticism, but the general motivation of exogeneity still stands, the methods proposed here can offer interesting insights.
This paper also relates to the literature that proposes the use of machine learning methods for causal inference \citep*[e.g.][]{Athey2019Machine}. Shift-share IV estimation does not resemble a high-dimensional problem prima facie, because there is only one instrument. Arguably however, all of the shares can be used as separate instruments and the need to select invalid shares substantially increases the complexity of the problem, making the use of machine learning methods appropriate. \citet*{Mullainathan2017Machine} have pointed out that in the context of economic research, machine learning methods lend themselves mostly to predictive tasks and less to causal inference. This paper provides a remedy for a commonly seen endogeneity problem, which threatens the reliability of causal inference in a wide range of economic studies.
In this section, I present the shift-share setup and the exclusion restriction in terms of shares. I show under which conditions the exclusion restriction is fulfilled and show a setting in which it is plausible that some shares are valid and some invalid. A discussion of indications that a some valid - some invalid setting applies concludes this section.
First, consider a linear model with a constant treatment effect $\beta$:
where $l$ indicates the location and $t$ the time period. A discussion of the constant treatment effect assumption can be found later in the text. The outcome variable is denoted by $y_{lt}$, $x_{lt}$ is the treatment, $u_{lt}$ is an idiosyncratic error term with $Cov(u_{lt}, x_{lt})=0$ and $\theta_{lt}$ denotes unobservable shocks which might be correlated with the treatment, i.e. $Cov(\theta_{lt}, x_{lt}) \neq 0$. For example, the outcome variable is employment growth in a certain region and year, the independent variable is growth of the immigrant share and the unobserved shocks $\theta_{lt}$ are labor demand shocks which might be correlated with the growth of the immigrant share. I abstract from covariates for ease of exposition.
In model (ref), assume the treatment variable has the structure $$x_{lt} \equiv \sum_{j=1}^J z_{jlt} \cdot g_{jlt}\text{,}$$ where $j$ indicates a class (e.g. the industry or the origin country of migrants), $z_{jlt}$ is the class-specific share in a certain region and $g_{jlt}$ is the region-specific growth-rate (or shift) of that class at time t. For example, $z_{Mexico,CA,2020}$ is the share of Mexicans in California in 2020 and $g_{Mexico,CA,2020}$ is the inflow of migrants from Mexico to California in 2020. These shifts and shares are available for $J$ classes, i.e. origin countries, in the migration example.
In many settings, $x_{lt}$ can be subject to endogeneity problems such as correlation with unobserved shocks and reverse causality. In this model, the regressor is endogenous when $Cov(x_{lt}, \theta_{lt}) \neq 0$. In the migration context, Mexican migrants may have chosen to settle down in California precisely because of the high wages at destination. Part of the correlation that is measured with ordinary least-squares regressions would thus be due to migrant selection into regions.
To circumvent this problem, a shift-share approach replaces components of the treatment variable by shares and shifts which are presumably unrelated with changes of the outcome variable. For example the share of Mexicans in California relative to Mexicans in the US is replaced with the same share, at a certain base period $t^0$ earlier in time (say 1990), while the growth rate of Mexican immigrants in California is replaced by its equivalent at the national level. The resulting shift-share IV is
where $g_{jt}$ is the national growth rate of industry $j$ (i.e. the shift) at time $t$ and $s_{lt}$ is then used to instrument for $x_{lt}$.
The exclusion restriction is the key identifying assumption for any instrumental variable approach. In this setting, the exclusion restriction is stated in terms of shares. This is the setting proposed by \citet*[GSS,][]{Goldsmith-Pinkham2020Bartik}. To show fulfilledness, violation and partial violation of the exclusion restriction, I set up a simple model. The structural equation is augmented by the shares $z_{jlt^0}$, with coefficients $\alpha_j$ which model the direct effects on the outcome. This is the definition of validity found in \citet*{Kang2016Instrumental}.
The model becomes
Equation (ref) denotes the first stage. Relevance is given when $\gamma_s \neq 0$. When all shares are to be used as instruments, separately, relevance is given when $\gamma_j\neq0$ for all $j$ in equation (ref). This paper focuses on the exclusion restriction. Relevance is plausible because the underlying idea of this instrument is that immigrants settle in regions where they find communities of earlier migrants from their same country of origin, for example because they rejoin family members or there is a network of their country of origin which eases their arrival. This is why the shift-share instrument has also been called “network”, “enclave” or “past settlement instrument”. The higher probability to settle in regions in which communities of their same origin country can be found creates a correlation between past and present settlement, and the instrument is relevant.
Shares might fail validity because they have a direct effect on the outcome, as measured by $\alpha_j$ but they might also need to be discarded because they are related to the outcome through unobservable shocks, $\theta_{lt}$. To show this, I allow for a non-zero correlation between current and past unobservable shocks:
Now, assume that the past unobservable shocks can be written as
Then, the structural equation becomes
where $\xi_{lt} = \rho\epsilon_{lt} + \nu_{lt} + u_{lt}$. In order for the initial shares to be valid instruments, they should not be directly related with the outcome ($\alpha_j=0$) and they should not be related to the initial shocks ($\phi_j=0$) or there should be no serial correlation between initial and current shocks ($\rho=0$).
Next, I summarize the above and introduce the definition of share validity in the context of this model to state the exclusion restriction more easily.
In the migration example, the first part of validity means that there is no adjustment through other factors of production. The second part means that unobserved shocks are not related to initial shares and/or these shocks are not correlated over time. The strict exclusion restriction can now be stated as
Under the strict exclusion restriction and relevance, the shift-share IV estimator is consistent (Proposition 2 in GSS). Note that when Assumption (ref) is fulfilled, the shifts do not play a role for the validity of the instrument.
Another way to achieve consistency of the estimator is relying on random shifts. This is the setting in \citet*{Borusyak2020Quasi} and \citet*{Adao2019Shift}. Which of the settings should be considered is dependent on the application. Still, the methods proposed here are also applicable to the random shocks setting of \citet*{Borusyak2020Quasi}, when there are multiple shifts. The following sections discuss how this can be achieved.
Applications in labor and migration economics often are related to share exogeneity, because they stress that past shares are not directly related with the outcome of interest and are hence valid. Twenty-one examples for this are listed in Table (ref) in the Appendix. This list is not exhaustive. In the mentioned papers, the reader can find explicit statements that share exogeneity motivates the validity of the shift-share IV strategy.
The definition of validity in the preceding section makes it clear that violations of the exclusion restriction can come from two different sources. First, a non-zero $\alpha_j$ invalidates the shares. In the migration setting, \citet*{Jaeger2020Shift} warn that there might be direct effects through general equilibrium adjustments. The concern is that the economy reacts dynamically to migrant inflows. If this is the case, there is a direct correlation between instrument and outcomes, through native labor, capital and other general equilibrium adjustment channels, invalidating the instrument. One way that this might apply is illustrated by \citet*{Borjas2003Labor}: if migrants choose to move to regions with persistently high wages and native workers choose to migrate in response to the immigration of foreign workers, then the effect of immigration is positively biased.
Second, when unobserved shocks today and at the initial period $t^0$ are correlated ($\rho \neq 0$) and the initial shocks are related with initial shares ($\phi_j \neq 0$), this induces a non-zero correlation between instrument and error term. A violation is plausible, because serial correlation of unobservables is typically discussed in the literature \citep*[see Table (ref) and ][]{Jaeger2020Shift} and initial migrants might well have been attracted by economic conditions. In principle, the bias could go in both directions because migrants might endogenously select into regions with higher wages, or into regions with lower growth potential.
The exclusion restriction is strict in the sense that it must hold for all $J$ shares. What looks like a single exclusion restriction in a just-identified model is in fact a set of $J$ exclusion restrictions. Therefore, the researcher needs to feel comfortable defending the exclusion restriction for Mexicans, Cubans, Canadians, Indians and all origin countries used when constructing the IV.
In practice, it is very difficult, if not impossible to credibly uphold the strict exclusion restriction. While building an intuition about which shares are valid might be feasible, arguing that none of them had a long-term effect or was correlated with initial shocks is very restrictive. Thinking about which factors determined migrant settlement at an initial point in time makes it clear how difficult and hypothetical such an argument is destined to be. Institutional knowledge about which origin country group was mostly drawn into cities which were experiencing a boom at the time of settlement is typically unavailable. This holds true especially in settings in which a large number of countries of origin is used. Such detailed knowledge about the structural mechanisms at work is only available for very few countries, if any.
Until now, there have not been attempts to make shift-share designs robust to violations of the exclusion restriction in Assumption (ref). GSS propose computing sensitivity-to-misspecification (Rotemberg) weights, which indicate by what percentage the bias of the shift-share IV estimator changes if the bias from a certain share increases by one percent. The authors point out that one should argue prudently for the validity of shares associated with large weights. While these weights indicate the relative importance with which an individual invalid share contributes to the bias of the estimator, the latter can still be considerable in absolute terms, even if only shares associated with low weights violate the exclusion restriction. Therefore, it does not suffice to argue for the validity of the shares associated with the largest weights to make a case for a low bias in absolute terms.
These reasons for violations of the strict exclusion restriction indicate that in many settings it can at best be hoped that some but not all shares are valid. The general share validity setup as stated by GSS might be credible, but not for all shares.
The migration example applies to a setting with partial violation of the exclusion restriction for the following reasons. First, when $\rho\neq0$, some migrant groups might be related with labor demand shocks at the base year ($\phi_j\neq0$), while others are not. The absence of correlation with unobservable shocks is credible for some shares, because only some origin country groups might have migrated mainly because of economic reasons. This is in line with Jaeger2007Green, who finds that migrants with employment visa where most responsive to economic conditions in their location choice. If the visa composition varies by origin groups, then some shares might have been driven mostly by factors orthogonal to economic conditions. Jaeger2007Green also finds that in the beginning of the 1970s the share of employment-based visa was low. The increase of employment visa over the decades implies that origin country groups in a later base period are more likely to be invalid.
Second, there are multiple sets of shares, which vary by base year. Some base years are correlated with the current shocks. Then for some years, $\rho_{0}\neq0$, while for others $\rho_{-1}=0$, when the correlation breaks after a few decades. Third, some origin country groups might have had long-term effects on wages, while the effects of others have worn off quickly. This might be the case when origin country shares which consisted mostly of people with family visas did not affect other factors of production in the long-term.
In applications, the discussion of single shares as potentially problematic indicates that the researchers think of a setting in which some shares are valid, while others are invalid. Another telltale sign of such a setting is when researchers report Rotemberg weights and exclude the shares with the highest weights as a robustness exercise. The questions in the application of this diagnostic are: “By how much does the bias of the estimator change, if a certain share is invalid? What happens if we assume that the most influential shares are invalid and exclude them from the estimation?” These questions imply that it is feasible that some shares are valid, while others are invalid.
In the literature, there is also evidence that such a some valid - some invalid setting is indeed the case. Tabellini2020Gifts raises the concern that specific origin country shares violate the exclusion restriction because Italian or Irish migrants could have chosen their city of location endogenously, based on the possibility to influence the local economy and politics. \citet*{Hunt2017Impact} and \citet*{Wozniak2012Timing} use adapted versions of the shift-share instrument where certain origin countries are excluded from the construction.
Arguably, the random shocks setting can be used when the strict exclusion restriction fails, but when the shocks are not numerous and random, which is often the case, this alternative approach is of little help. This offers a further setting where the some valid - some invalid IV setting applies. Several shift-share IVs can be constructed by using various push factors of emigration as shifts, such as economic, conflict- or civil liberties related variables. One might argue that economic variables are most likely to be related across countries, while political variables at origin are more likely to be unrelated with the local economic outcomes at destination. This setting is not based on the validity of shares and illustrates that the methods are in fact more widely applicable, also to the \citet*{Borusyak2020Quasi} context.
In this section, I introduce how to obtain modified estimators which are robust to invalid shares. I present the general procedure, the leveraged methods and extensions of these methods.
The idea of the procedure is to preselect valid shares beforehand with methods that will be presented in the following. I first introduce some notation. Let $\mathbf{Z}_\mathcal{V}$ be the matrix of valid IVs with $\mathcal{V} = \{j: \alpha_j = 0\}$ the set of valid IVs and $\hat{\mathcal{V}}$ the set of IVs selected as valid. Let $\mathbf{Z}_\mathcal{I}$ be the matrix of invalid IVs with $\mathcal{I} = \{j: \alpha_j \neq 0\}$ the set of invalid IVs and $\hat{\mathcal{I}}$ the set of IVs selected as invalid. Further, $|\mathcal{V}|$ is the number of valid and $|\mathcal{I}|$ is the number of invalid IVs.
In short, the procedure works as follows:
It is important that the shares selected as invalid are controlled for. The invalid shares can only be omitted from the regression, if they are uncorrelated with the valid shares. However, this is unlikely to be the case in practice. Consistency of the proposed method follows directly if the selection methods used in the first step consistently select the invalid shares.
When validity is plausible only with random shifts and there are multiple shifts, one can also apply an industry-level regression with multiple shifts and select shifts instead of shares, analogously to above. Disregarding which source of validity is put emphasis on, the preselection of variables starts with theoretical arguments. The set of shares (or shifts) selected is hence the intersection of the shares considered to be valid by the researcher and the algorithm.
In this section, I start with the critical assumptions needed for identification in the adaptive Least absolute shrinkage and selection operator (AL) by \citet*{Windmeijer2019Use} and the Confidence Interval Method (CIM) by \citet*{Windmeijer2020Confidence}, that I will use in this paper. Descriptions of the methods can be found in appendix (ref) and in the original papers.
The properties of the methods that will be leveraged to improve shift-share estimation are the so-called “oracle properties”. Oracle properties mean consistent selection of invalid IVs and convergence in distribution to the ideal (oracle) estimator that uses the model under perfect knowledge about the identity of invalid IVs. In Appendix (ref) I describe the oracle estimator more closely.
In other words, if an estimator has oracle properties, it works as well as if one knew the true identity of invalid IVs. The AL has oracle properties when the majority of IVs is valid. All of the IVs also need to be relevant, as noted in equation (ref).
The Confidence Interval Method has oracle properties when the largest group of IVs is valid. The plurality condition in \citet*{Windmeijer2020Confidence} states that the group of valid IVs is larger than any other group. A group is defined as a set of IVs associated with an estimate which asymptotically deviates from the true $\beta$ by the same constant $c = \frac{\alpha_j}{\gamma_j}$. For the valid group, $c$ is zero. Formally, the plurality exclusion restriction is
To compare these two assumptions, consider the following example: there are five IVs. The true effect is $\beta=1$. For three of these IVs: $\alpha_j=0$ and hence the three IV-specific estimands are $\beta_j = \beta$, while the remaining estimands are $\beta + \frac{\alpha_j}{\gamma_j}$ with $\alpha_j\neq0$. In this example: $\beta_1=\beta_2=\beta_3=1$ while e.g. $\beta_4=4$ and $\beta_5=5$. Clearly, the majority assumption is fulfilled. The plurality assumption is also fulfilled, because the largest group of IVs is valid. When only two IVs are valid and the third now has an estimand which is $\beta_3 = 3$, the majority is violated, because only 2/5 IVs are valid, but the plurality is still fulfilled, because there is one valid group of two IVs and three singleton groups. Therefore, the plurality assumption can still hold even when the majority is violated. Next, I discuss the choice of methods. I introduce how the two methods work in Appendices (ref) and (ref).
The procedure builds on two methods from an emerging literature that investigates IV estimation in presence of invalid IVs. The proposed methods are the only ones which combine the following four benefits.
First, they are computationally feasible. \citet*{Andrews1999Consistent} requires to search over all possible models, which is computationally infeasible when the number of IVs is moderately large. Second, they do not require a priori knowledge about an initial set of valid IVs. \citet*{Caner2018Adaptive} also allow for invalid IVs when a set of valid IVs is known a priori. Third, the methods do not need assumptions on the correlation of first-stage and structural parameters. \citet*{Kolesar2015Identification} assume that first stage and direct effects are uncorrelated, but in applications, this assumption is rather strict. Finally, the direct effect of invalid IVs on the outcome need not be close to zero. This needs to be the case in \citet*{Conley2012Plausibly}, where additionally prior knowledge on possible values of $\alpha$ is needed. The methods used in this paper allow for arbitrarily strong direct effects. In fact, their performance even improves when the direct effects are large.
In the following, I apply the methods to two real-world examples. I first reproduce the original estimates by using the standard shift-share IV which uses all shares, irrespectively of their validity. I then compare this regression with the result of the adjusted estimators, using AL and the confidence interval method. In the Appendix (Section (ref)) I also apply the methods in a Monte Carlo simulation that illustrates how the methods work with weaker IVs and strong violations.
The first empirical application is the estimation of the effect of immigration on wages in the United States. \citet*{Basso2015Association} estimate the linear model\footnote{ I choose to use this paper as a reference even though it is unpublished, for the following reasons: the number of locations is large, which is helpful, because the methods I use make asymptotic arguments. Many of the published papers for which data is available have observation numbers which are low. For example, Card2009Immigration, which GSS use as an illustration, uses only 124 city-observations. }
with three time periods $t$ (1990, 2000, 2010) and 722 commuting zones $l$. On the left hand side, $\Delta y_{lt}$ stands for the three dependent variables used in separate regressions: the change in log weekly wages, and the change in log weekly wages of high- and low-skilled workers. On the right-hand-side, $\Delta immi_{lt}$ is the change in share of immigrants in total employment and $\beta$ is the coefficient of interest. Decade fixed-effects are denoted by $\psi_{t}$ and $\varepsilon_{lt}$ is the error term. Commuting zone fixed effects are accounted for by first-differencing.
As discussed before, estimating by OLS does not account for migrant sorting into regions: migrants might select into more prosperous or declining regions, creating a correlation between migrant location and outcome, which cannot be accounted to the impact of immigration. To tackle this problem, a shift-share IV, which uses origin-specific migrant shares in 1970 and changes in migrant populations, is used. The shift-share IV is
where $z_{lc,1970}$ is the share of immigrants from country $c$ in a location $l$ at base period 1970, and 19 origin country groups are used. The change of immigrants from country $c$ is denoted by $g_{ct}$. All of the origin countries are assumed to plausibly fulfill the exclusion restriction a priori.
The SSIV estimates are in column 1 of (ref). The coefficient estimate of the change in log weekly wages of the natives is 0.09, but it is insignificant. The 2SLS estimate is 0.479 ($p < 0.05$) and the LIML estimate is 0.568 ($p > 0.1$). For the change in log weekly wages of the high-skilled, the coefficients are 0.35 for SSIV, 0.519 for 2SLS and 0.672 for LIML. The coefficient is significant for 2SLS. For low-skilled workers, the effects on wages are negative at -0.66 and statistically significant for the shift-share analysis, and they are close to 0.1 for 2SLS and LIML. Overall, the standard estimates suggest positive effects on high-skilled and null or negative effects on low-skilled wages. The first-stage F-statistics are at 171.3 for the overidentified models and at 21.8 for the SSIV.
The possible violations in the migration context have been discussed in section (ref). However, it is unclear whether these concerns really applied to some origin country groups and if yes to which. The remainder of this section shows the shares selected as invalid and how large the adjusted estimate is.
\afterpage{ \newgeometry{left=0.6in, right=0.6in, top=1in, bottom=1in}
\restoregeometry }
\afterpage{
}
The results of applying AL and CIM on the immigration example are in Table (ref). Each panel presents the results for one of the outcomes. A list of selected countries can be found in table (ref). Overall, with the new methods, the coefficients of immigration decrease and often switch sign. The decrease in coefficients tends to be stronger when CIM is used in the selection step.
When choosing the significance level of $0.1/ln(N)$ (0.01375) in the downward testing procedure as proposed in WLHB, no country is chosen as invalid. Thus, the adjusted estimators are identical to the original ones (col. 1). \citet*{Bowsher2002testing} shows that the use of many IVs leads to low power when using the HS test. Larger significance levels of the HS test are more conservative, which is the inverse logic as with conventional tests of coefficient significance \citep*{Roodman2009note}. Hence, a more conservative strategy would be to set the threshold to a more conventional level, for example to $0.05$ or $0.1$. Increasing this threshold in the testing procedure leads to the selection of a few countries for low-skilled wages (Panel C), but does not change the results qualitatively.
One might be concerned that the IVs are weak and hence the HS test is unreliable. To address this concern, I also use the Anderson-Rubin test in the downward selection procedure. As a threshold I use $0.1/ln(N)$, as originally proposed in WFDS. Now, both methods select many IVs. Often, more than half of shares are selected as invalid. If a majority is invalid, AL in fact does not have oracle properties. I therefore rely on the CIM. The preferred analyses are hence those in the last column of table (ref) for 2SLS and LIML. All estimates decrease strongly, and mostly become negative.
For overall weekly wages, the coefficients become negative and are statistically insignificant. For wages of the high-skilled, the estimates from overidentified models become negative and statistically significant when selecting via CIM, which is in stark contrast with the original results. Interestingly, the absolute size of coefficients is very similar to the original ones, with the difference that they changed direction. For wages of the low-skilled, now 15 countries are chosen as invalid by CIM. The coefficients of 2SLS and LIML are negative but none of them are significant. This might be because the F-statistic becomes low. Still, for most other analyses the F-statistics are still reasonably high.
It is reassuring to see that the differences between 2SLS and LIML estimators are smaller with the corrected as compared to the standard estimators. The remaining differences in the adjusted shift-share (SSIV), 2SLS and LIML estimates may stem from different reasons. First, LIML approximately eliminates the finite sample bias that is due to weak instruments when using 2SLS. Second, the weighting scheme of each just-identified IV estimate differs across SSIV and 2SLS. The weights shown in GSS are dependent on the shift variable. The weighting implicit in the 2SLS estimator which uses only shares, does not take shifts into account. Therefore, different results may also arise because different methods estimate different weighted combinations of just-identified estimates.
}
}
{ \def\sym#1{\ifmmode^{#1}\else\(^{#1}\)\fi}
}
\floatfoot{ Note: This table reports estimates of $\beta$ in equation (ref). $N= 538$ in panels A and B and $N= 526$ in panels C and D. Standard errors (in parentheses) are clustered by commuting zone. First-stage F-statistics are reported. Observations are weighted by beginning-of-period population. Outcome variables are listed in the panel heads. In the first row of each panel, shift-share IV results are reported. In the second and third rows, the full vector of shares is used for 2SLS and LIML. In the last rows, the number of countries chosen as invalid and the thresholds used in the HS procedure are reported. In column (1), all shares are assumed to be valid. Column heads of columns 2 to 5 denote which method has been used for selection.} \end{table} \end{center} \end{landscape} } \fi
Taking into account the critique of \citet*{Jaeger2020Shift}, who argue that using a single regressor compresses the long- and short-term effects, means to include lagged migration. The equation now becomes
where $\beta_c$ is the coefficient of interest for the contemporaneous impact and $\beta_l$ denotes the coefficient of interest for the lagged impact. I include an additional shift-share IV, now using 1980 as a base period. When using 2SLS or LIML, the number of shares increases to 38 (19 per base year).
The results are shown in Table (ref). The standard estimates always suggest positive effects in the short and negative effects in the long run, across estimators and outcomes. This is exactly the opposite of what \citet*{Jaeger2020Shift} expect: partial equilibrium effects should be negative and general equilibrium adjustments are expected to offset these negative effects. These unexpected coefficient estimates might be due to the same endogeneity problems as before: both base years of the number of foreign borns might be directly correlated with the endogenous treatment.
Using the extension of the adaptive Lasso presented in Appendices (ref) and (ref), with the Hansen-Sargan downward testing procedure, only for wages of the high-skilled the UK and Ireland are selected as invalid once and the adjusted estimates still have the same signs.
When using the Anderson-Rubin test instead, nine to eleven countries are selected for each outcome. The countries selected have a large overlap. Most variables selected come from the year 1980. I focus on 2SLS and LIML results, because the Cragg-Donald statistic of the SSIV is very low. The estimates now have the expected sign: the coefficient of contemporaneous immigration is negative and that of lagged immigration is positive.
}
}
{ \def\sym#1{\ifmmode^{#1}\else\(^{#1}\)\fi}
}
\floatfoot{ Note: This table reports estimates of $\beta$ in equation (ref). $N= 2166$ (722 CZ $\times$ 3). Standard errors (in parentheses) are clustered by commuting zone. First-stage F-statistics are reported. Observations are weighted by beginning-of-period population. Outcome variables are listed in the panel heads. In the first row of each panel, shift-share IV results are reported. In the second and third rows, the full vector of shares is used for 2SLS and LIML. In the last rows, the number of countries chosen as invalid and the thresholds used in the HS procedure are reported. In column (1), all shares are assumed to be valid. Column heads of columns 2 to 5 denote which method has been used for selection.} \end{center} \end{table} \restoregeometry } \fi
One might fundamentally question the share exogeneity interpretation of the shift-share design in the migration setting. In principle all shares could directly affect wages. If this is the case, the shift-share IV can be motivated via random shifts as in \citet*{Borusyak2020Quasi}. This offers an alternative starting point for the selection methods. This shows that the new methods are not restricted to the exogenous share world of \citet*{Goldsmith-Pinkham2020Bartik}.
If validity was still a concern for all shares, one additional way to check for robustness of the results is to motivate the exclusion restriction through quasi-random shifts and to use country-of-origin specific push factors related to war, civil liberties or natural disasters. One example for such an approach is \citet*{Llull2017effect}. The selection methods can then be used with the different shifts in an overidentified model. There would be reason to believe that some instruments are valid while others are invalid. Some shifts are related to war, others to politics, again others to other country-of-origin factors. Some might be correlated with unobservable shocks which drive wages at destination, for others it is difficult to think of a reason why that would be the case.
If multiple shifts are available, multiple shift-share instruments can be generated and used in an over-identified model. The SSIV constructed with shocks that fulfill the conditions in \citet*{Borusyak2020Quasi} can then be selected. Alternatively, one could also directly use the class-level regression, and use the shocks as IVs directly. The latter approach will be used in the international trade application. \citet*{Borusyak2020Quasi} show consistency of the IV-estimator, taking into account that the data is non-iid. This does not pose a challenge for the selection methods, because the key assumption is that a large-enough group of IV-specific estimators is consistent, regardless of how that consistency is established.
Moreover, selection of a particular group does not necessarily mean that the other groups are invalid instruments. Different shocks might produce heterogeneous effects. The migration inflow due to war might be different from that due to a decrease in civil liberties, which is more likely to induce migration of the elite.
To illustrate this approach, I used eleven shifts and produced eleven shift-share instruments, still using the 1970 country shares.\footnote{The shifts and their sources are as follows: migration (as in the preceding subsections), battle-related deaths, onesided violence and nonstate violence (Uppsala Conflict Data program, www.ucdp.uu.se), population (World Development Indicators), Civil Liberties, Political Rights, Freedom House Status House2020Freedom, Polity Score (Polity V project), Press Freedom Status and Press Freedom Score House2017Freedom.} I directly estimate the dynamic model suggested by Jaeger2020Shift, including lagged immigration. My findings can be found in table (ref). In brief, the main results stay the same: with the HS-testing procedure, only few IVs are selected as invalid, while with the Anderson-Rubin test, more IVs are selected. With the AR testing procedure, all estimates turn negative but insignificant. This could be due to a loss in relevance, as the first-stage F-statistic becomes low. The shifts selected as invalid can be found in table (ref). The variables that are selected most often are the IVs constructed with battle-related deaths, with the political freedom indicator (every analysis) and with the Press Freedom Score (five times). Since the last two express similar things, it makes sense that they constitute a group.
There are five key takeaways from the application of AL and CIM to the estimation of the effect of immigration on wages. First, the results from adjusted estimators suggest a strong positive bias of standard estimates. This is in line with most of the literature, that expects an upward bias. This doesn't seem to be due to weaker instruments after selection, because the first-stage statistics are still reasonably high and the use of the LIML estimator, which has better finite-sample properties in presence of weak IVs suggests the same direction of the bias.
Second, the selection of shares is consistent with economic intuition. The selection of Central and Eastern Europe (including Russia) in almost all analyses can be explained by the emigration from the Soviet Union in the 1970s and the Post-Soviet countries in the 1990s. The emigrants predominantly chose coastal cities which had large country-of-origin communities, but also cities which had experienced lasting prosperity. The conditions which have made these places attractive might be correlated over time. This makes a violation of the exclusion restriction likely. The share of migrants from the UK and Ireland, which has been picked by Tabellini2020Gifts as an example for possibly invalid shares is chosen nine times. When shares from multiple base years are used, mostly IVs with base year 1980 are selected. This is consistent with more job-related visa in 1980 as compared to more family-related visa 1970 and is in line with the common practice in the literature of choosing longer lags to break eventual correlation between shares and current unobservable shocks.
Third, the application shows the added value of the methods to existing econometric tools. The Rotemberg weights proposed by GSS help understand which share's invalidity is most likely to bias results, but it does not tell the researcher whether this bias is large in absolute terms and it does not lend guidance on which country should be excluded effectively. A few of the countries flagged as potentially problematic by high weights have been selected. The Philippines have received the highest sensitivity-to-misspecification weight in GSS. Indeed, they have been selected seven times by AL and CIM, and adjusting for them results in large qualitative changes of the coefficients.
However, if the Rotemberg weights for some origin countries are low, their invalidity could still contribute to a large part of the inconsistency of estimators. Notably, many country groups which are not worrisome according to the top-5 Rotemberg weights, such as Central and Eastern Europe have been chosen as invalid, while some that have high weights have not been selected. This shows how the new methods can guide the selection of shares beyond the discretion of researchers.
Fourth, the fraction of shares selected as invalid can be high. A maximum of 15 out of 19, are selected as invalid suggesting that the majority assumption is likely to be violated. This suggests that the adaptive Lasso can not consistently select valid IVs in the migration setting with one regressor. Also, there is a large overlap of selected countries as invalid, by variables and methods used. This is reassuring in that it confirms that the share selection is not erratic.
Fifth, when including lagged immigration, the coefficient estimates have the expected sign, only with the proposed AL adjustment. The origin-country variables selected are mostly those from the year 1980. This is consistent with \citet*{Jaeger2020Shift}, who worry that spatial adjustments might take around ten years\footnote{“Research on regional evolutions in the U.S. concludes, however, that spatial adjustments can take around a decade or more.” (p. 10)}. In my analysis, I also use data from 1990. Hence, my analysis confirms \citeauthor*{Jaeger2020Shift}'s\ (Jaeger2020Shift) result that using contemporaneous and lagged immigration can help uncover the effects of immigration. It also confirms the common practice of taking longer lags of the country-of-origin distribution to plausibly fulfill the exclusion restriction.
\citet*[][ADH]{Autor2013china} study the impact of Chinese imports on employment in manufacturing in the US. The regression equation is
where the left-hand side is decadal change in manufacturing employment in commuting zone $l$, $\beta_1$ is the coefficient of interest and $\Delta IPW_{ult}$ is import exposure, defined as $\Sigma_{j} z_{ljt} g_{jt}$. Here, $z_{ljt}$ are the shares of workers in commuting zone $l$ employed in industry $j$ at time $t$ and $g_{jt}$ measures the growth of imports from China in industry $j$. This regression is estimated in first-differences to exclude commuting-zone fixed effects and augmented by a time dummy and a set of commuting-zone-level controls. The time period used ranges from 1990 to 2007 and there are 397 industry shares, indexed by four-digit SIC codes.
The endogeneity issue that affects this analysis is that both employment and imports might be correlated with unobserved shocks to US demand. To address this problem, a shift-share instrument is used, which replaces the share of workers with the same share ten years earlier and uses import exposure of other high-income countries rather than the US. ADH find a coefficient of -0.596. I report the same coefficient for the original estimate in row 1, column 1 (1,1) of table (ref). When using all shares separately in a 2SLS estimation, a lower coefficient of -0.183 is found (2,1). The same model is also estimated by LIML (3,1). These are the baseline coefficients to which the adjusted estimation results will be compared.
One might understand the analysis from the viewpoint of GSS in the framework of a pooled exposure research design, in which employment shares capture local exposure to common import shocks. My results show which industry shares one should worry about if one chooses to rely on share validity.
ADH discuss the possible invalidity of three specific industries: the computer industry, construction materials as well as apparel, footwear and textiles. GSS show that electronic computers display the highest sensitivity-to-misspecification weight, making the validity of this specific share especially important.
\afterpage{
}
The results of the AL-adjusted IV estimators are presented in table (ref). Using AL, the coefficients change by little. With the default threshold of the over-identification test at $0.01375$ ($0.1/ln(N)$) as in WFDS, the test does not reject the Null hypothesis, all shares can be used for the construction of the shift-share IV and all coefficients are identical to the original estimates (column 2 of table (ref)). To account for the problem of too many instruments in the HS-test, I set the threshold to $0.05$. Now, only one industry is selected. When excluding this industry from the construction of the instrument in column 3, the estimate is virtually unaltered.
When applying CIM, the industry chosen by AL and seven additional industries are selected as invalid. The estimates becomes larger in absolute terms but the confidence interval still includes the original estimate. Hence, the application is also robust to omitting shares chosen as invalid.
When estimating the post-selection model by LIML, the estimates are very different from 2SLS. This might indicate that many IVs are weak and 2SLS is therefore biased. In order to adjust for this, I use the Anderson-Rubin test in the downward testing procedure instead of the HS-test. When using adaptive Lasso with the AR-test, the method now selects 63 shares as invalid. When selecting with CIM and the AR-test (column 7 of table (ref)), even 128 shares are chosen as invalid. The estimate for SSIV now moves to -0.92, but the 95% significance interval still includes the original estimate. The 2SLS estimate becomes positive, with a coefficient of 0.12, while the coefficient of LIML is positive and large.
The industries chosen as invalid are listed in table (ref) of the supplementary material. These industries concord with those discussed in ADH. The first industry labeled as problematic was the computer industry. Industries belonging to electronic and computer equipment (SIC35 and 36) are among those chosen most often, constituting up to 29 percent of the shares selected as invalid. The second industry class that is discussed in ADH is related to construction. Up to 16 percent of selected shares come from industries that are associated with construction (32, 33, 34). The third industry discussed in ADH is apparel, footwear and textiles. Also, up to 16 percent of shares selected comes from these industries (22, 23, 31).
The analysis offers additional information beyond the sensitivity to misspecification illustrated by Rotemberg weights. Games, Toys and Children Vehicles as well as Household Audio and Video Equipment have obtained the second- and third-largest Rotemberg weights in GSS, and they have been selected by AL and CIM. The industry with fourth-largest Rotemberg has been selected by CIM and the one with the fifth-largest weight has been selected by both methods. However, the SIC-4 industry with the largest weight has not been selected, while numerous industries from the SIC-2 industry related to it and SIC-codes from the food sector have been selected. This shows how the proposed methods can guide share selection in expected ways but it can also inspire to think about the possible endogeneity of some other industries.
Overall, if one believes that some shares are valid and some invalid the selection procedures single out industries which are also in harmony with the ones discussed by ADH. The results are relatively robust to the use of the new methods.
If share exogeneity is not credible, one can also understand exogeneity of the shift-share instrument from a random shift perspective, as in \citet*{Borusyak2020Quasi}, who run an equivalent industry-level regression which uses the shift-variable as instrument. In table C4 of their paper, they use an overidentified model with all eight shifts from other high-income countries instead of the aggregated shift.\footnote{In this analysis, the authors add lagged sum of shares for each period as control variables. I follow this modeling choice to keep results comparable.} The common concern is usually that imports of high-income countries are correlated with unobservable shocks. The estimates for these estimations lie at roughly -0.24 for both 2SLS and LIML. It is reassuring to see that the coefficients of the two methods coincide.
I reproduce the results from this table in unreported estimations. The HS- and AR-tests do not reject at any conventional significance level, and therefore the selection algorithms select all eight shifts as valid. This robustness to the use of the new methods illustrates how in this example identification should be thought of in terms of shifts. This is in line with \citeauthor*{Borusyak2020Quasi}'s\ (Borusyak2020Quasi) and \citeauthor*{Adao2019Shift}'s\ (Adao2019Shift) interpretation of the exclusion restriction as shock exogeneity.
Researchers can also leverage a larger set of import shifts from even more countries in which they are confident that most of the shifts are valid and select shifts via AL and CIM. This example hence shows that even in settings where the exogenous share interpretation is controversial the two methods can be helpful.
This paper proposes adjusted shift-share IV estimators which require that only a majority or plurality of shares is valid. New statistical methods are used to select invalid shares. The STATA-programs in the supplementary material offer a simple way to apply the proposed methods.
In the migration setting, many shares are chosen as invalid and the adjusted estimates are much lower than the original ones, suggesting negative effects of immigration on wages. When including lagged migration, with the adjustments the coefficients have the expected signs. In this setting the proposed methods can be helpful for retrieving a causal estimate. In the China shock example the results are mostly robust to the use of the new methods. The results are also robust to the use of the new methods when the exclusion restriction is motivated through random shifts. In simulations I show that even in settings with weak instruments the estimators can continue to perform well. Severe violations of the exclusion restriction even improve the performance of the estimators in small-sample settings.
In the appendices I provide detailed descriptions of the employed methods, briefly discuss the implications of weak IVs and heterogeneous effects, provide an extension to multiple endogenous regressors and additional simulations.
The methods are complementary to the recent literature on shift-share IVs. Before using them, it is important to think carefully about which source of validity is most feasible. The methods can be most helpful when researchers think about the shift-share instrument from the perspective of valid shares and whenever some shares are suspected to be directly correlated with the outcome variable. If there are many class-specific shocks, for example for multiple high-income countries, there is also scope for applying the methods in the quasi-random shocks setting. When doing so in this paper, my conclusions do not change, qualitatively.
I conclude with two shortcomings of the methods. First, the original methods only allow for one endogenous regressor. In the Appendix, I developed an extension of AL to multiple endogenous regressors which calls for stricter qualified majority assumptions. These new assumptions are confirmed in simulations. Further improvements would be to develop methods which can be readily extended to the multiple endogenous regressor case without making the exclusion restriction stricter.
Second, the validity of all shares might be a concern. Given that validity of shares relies on similar arguments, it is possible that they are all inconsistent in similar ways. In this case, consistent selection can not be guaranteed. In fact, even though the majority and plurality assumptions are considerable relaxations of the strict exclusion restriction, they are still strict. Importantly, researchers should find a set of variables whose validity can be credibly defended from a theoretical point of view. The methods proposed here complement thorough theoretical considerations and do not replace them; a convincing justification of the exclusion restriction is still imperative for the new estimators.
I'd like to thank Kirill Borusyak, David Dorn, Ben Elsner, Helmut Farbmacher, Paul Goldsmith-Pinkham, Chirok Han, Ines Helm, Stephan Huber, Peter Hull, Xiaoran Liang, Jan Stuhler, Frank Windmeijer, Joachim Winter and seminar participants at LMU Munich, TU Munich, DAGStat 2019, IAAEU, Regensburg, the Ammersee Workshop, EALE 2019 and the EALE/SOLE/AASLE World Meeting for helpful comments and discussions. I also thank Gaetano Basso and Giovanni Peri for sharing their data and code. I acknowledge funding through the International Doctoral Program “Evidence-Based Economics” of the Elite Network of Bavaria.
\newgeometry{left=1in, top=1in,right=1in, bottom=1in}