Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,523 characters · 22 sections · 83 citation commands
Nonlinear Treatment Effects in Shift-Share Designs
\setstretch{1}
\newsavebox{\tablebox} \newlength{\tableboxwidth}
We analyze heterogenous, nonlinear treatment effects in shift-share designs with exogenous shares. We employ a triangular model and correct for treatment endogeneity using a control function. Our tools identify four target parameters. Two of them capture the observable heterogeneity of treatment effects, while one summarizes this heterogeneity in a single measure. The last parameter analyzes counterfactual, policy-relevant treatment assignment mechanisms. We propose flexible parametric estimators for these parameters and apply them to reevaluate the impact of Chinese imports on U.S. manufacturing employment. Our results highlight substantial treatment effect heterogeneity, which is not captured by commonly used shift-share tools.
\
Keywords: Nonseparable models, Control Variables, Policy Effect, Shift-Share Instruments, Globalization, Employment.
\
JEL Codes: C14, C31, F14, J23.
\doublespacing
Shift-share designs have become a widely used tool in many fields, including trade autor2013, political economy Dippel2021,Campante2023, development economics topalova2010, and labor economics Acemoglu2020. The empirical importance of this tool motivated many methodological articles to discuss its identification and inferential aspects adao2019,goldsmith2020,borusyak2021. However, those papers rely on linearity or homogeneity assumptions, which may be potentially strong in many applications hahn2024. For instance, in the China shock setting autor2013, opening trade relations with China could have a large effect, whereas increasing import exposure when trade relations with China are already well-established may have little effect. Moreover, allowing for nonlinear heterogeneity in the context of two-stage least squares may lead to negative weighting problems, affecting the causal interpretation of the estimand heckman2006,blandhol2022,słoczyński2024,alvarez2024,hahn2024.
To overcome these two limitations, we propose a nonparametric model for shift-share designs. We adapt the triangular equation model proposed by imbens2009, allowing the outcome to be a nonseparable structural function of the treatment and a (possibly infinitely dimensional) shock, and the treatment to be a nonseparable structural function of the instrument and a scalar idiosyncratic term. The instrument combines common shocks (shifts) and individual measures of exposure to those shocks (shares). We leverage variation in the exogenous shares to construct a variable that controls for the endogenous part of the outcome shock.
By doing so, we identify both Average and Policy Effects. For Average Effects, we nonparametrically identify the Local Average Response altonji2005, the Average Derivative imbens2009, and the Average Structural Function blundell2003. First, the Local Average Response (LAR) function is the effect of the treatment on the outcome for a given value of the treatment, capturing the observed heterogeneity of our treatment effects. Second, the Average Derivative (AD) summarizes this heterogeneity into a single parameter that does not suffer from negative weighting issues. Third, the Average Structural Function (ASF) represents the average outcome at a fixed treatment value, capturing the observed heterogeneity in the levels of the outcome variable. Lastly, the Policy Effect (PE) captures the effect of a given treatment policy on the outcome variable.\footnote{An important limitation of shift-share designs, which also applies to our framework, is that they generally do not allow the identification of in-level effects, only in-changes. This drawback stems from the nature of the identifying variation, which relies on differential changes in exposure to common shocks.}
The Policy Effect is, potentially, a very interesting parameter to explore in Shift-Share applications. For example, in the China shock application, we may be interested in understanding how a new tariff policy that reduces imports from China affects the U.S. labor market. Interestingly, this parameter enables us to study counterfactual policies that have not yet been implemented.
In addition to identifying different parameters of interest, we provide a brief discussion of the two-stage least squares estimand (2SLS). By connecting this estimand to our nonlinear setting, we show that the 2SLS estimand may lack a causal interpretation without strong assumptions regarding the instrument-treatment relationship. Moreover, even if those assumptions are met, the 2SLS estimand captures a hard-to-interpret convex combination of local effects.
Since estimating those parameters nonparametrically is complicated by the large number of instruments (i.e., the size of the vector of shares), we also identify them using a semiparametric method. This semiparametric control function approach adapts the method proposed by chernozhukov2020 to the shift-share setting. It also provides a semiparametric estimator of the control function and flexibly parametric estimators of the target parameters. Moreover, we propose a uniform inference procedure using the weighted bootstrap.
Lastly, we reevaluate the impact of increasing Chinese imports on the growth of manufacturing employment in the U.S. between 1990-2000 and 2000-2007. To do so, we use commuting zone data from autor2013 and flexibly estimate our four target parameters. We also compare them against estimates based on a linear 2SLS estimator, finding that allowing for nonlinear heterogeneous treatment effects is fundamental to understanding the China shock’s impact on the U.S. economy.
Our Average Derivative estimates are small and do not reject the null hypothesis of zero average effects. In contrast, the 2SLS estimates are negative and significant. We find that these differences are explained by the negative weights within the 2SLS estimand in our dataset.
Our estimates of the Local Average Response function find strong evidence of nonlinear effects between 2000 and 2007. In particular, for regions that faced lower exposure to growth in Chinese imports, a marginal increase in exposure to growth in Chinese imports leads to a greater intertemporal difference in manufacturing employment. In other words, for lower values of exposure to Chinese imports, employment in manufacturing decreases less than it would without the marginal increase in exposure.
Our estimates of the Average Structural Function also find evidence of nonlinear effects between 2000 and 2007. They suggest that, for lower values of growth in exposure to Chinese imports, the reduction in the manufacturing employment rate was smaller than for those that faced higher values of growth in exposure to Chinese imports. However, in this case, the results of our nonlinear estimator are similar to the estimates based on linear 2SLS regressions.
For the Policy Effect, we evaluate the effects of counterfactually increasing U.S. import tariffs during the analyzed period. To do so, we connect the exposure to Chinese imports with import tariffs by using the elasticity of substitution estimated by fajgelbaum2019. We find that larger tariffs would not be sufficient to compensate for the loss of manufacturing employment caused by the increased exposure to Chinese imports. Importantly, the 2SLS estimates for the policy effect are larger and fall outside the confidence band of our nonlinear method. This result highlights how the linearity imposed by the 2SLS specification may lead researchers to overestimate the effect of tariffs. Consequently, allowing for nonlinear heterogeneous treatment effects is essential to our understanding of the impact of Chinese imports on the U.S. economy.
Related Literature. This article contributes to three distinct strands of literature. Concerning its contribution to shift-share designs, there exists a growing literature on identification and estimation bartik1991,goldsmith2020,borusyak2021,dechaisemartin2022,hahn2024 and inference adao2019,alvarez2022. Our contribution is focused on identification. While these articles clarify the identifying variation in shift-share designs, they rely on assumptions of linearity or homogeneity.
We depart from this framework by allowing for heterogeneous and nonlinear effects in a nonparametric setting. By doing so, we broaden the applicability of shift-share designs to contexts where effect heterogeneity is empirically and theoretically plausible. Such an extension is methodologically relevant because hahn2024 show that causally interpreting the 2SLS estimand in shift-share designs implies implausibly strong necessary conditions.
Concerning its contribution to control function approaches, our work is inserted in the literature about identifying treatment effect parameters when the structural functions follow a triangular nonseparable model chesher2003,imbens2009,blundell2014. We adapt the triangular model proposed by imbens2009 and the semiparametric tools proposed by chernozhukov2020 to the shift-share setting. These strategies complement existing work on control function models, providing a practical solution for applied researchers using shift-share instruments.
Lastly, we speak directly to the empirical literature on the China shock and its labor market consequences autor2013,costa2016. Our paper revisits the influential work of autor2013 through a nonparametric lens and shows that the standard 2SLS estimates may obscure meaningful heterogeneity in the impact of Chinese import exposure. By doing so, we contribute to a more nuanced understanding of the consequences of trade shocks and the limits of protectionist policy responses.
Paper Organization. This paper is organized as follows. Section (ref) describes the econometric framework, presents the target parameters, and discusses our identifying assumptions. Section (ref) presents our identification results and analyzes the 2SLS estimand in a nonlinear setting. Section (ref) explains our estimation and inferential procedures. Section (ref) discusses our empirical results with data from autor2013, while Section (ref) concludes.
In this section, we explain our econometric framework. Section (ref) starts by describing our model’s setting. Then, Section (ref) defines our target parameters. Lastly, Section (ref) states the model’s assumptions. In all these sections, we use our empirical application as an example to provide intuition.
We analyze the following triangular system with nonparametric, nonseparable equations:
where $Y$ is an outcome variable (e.g., growth in the manufacturing employment rate in a commuting zone in the U.S.) with support in $\mathcal{Y} \subseteq \mathbb{R}$, $X$ is a continuous treatment variable (e.g., temporal change in exposure to Chinese imports in a commuting zone) with support in $\mathcal{X} \subseteq \mathbb{R}$, $D$ is a vector of covariates (e.g., commuting zone's demographic characteristics) with support in $\mathcal{D} \subseteq \mathbb{R}^{dim(D)}$, $Z$ is a shift-share variable (e.g., $Z \coloneqq S \cdot W$) with support in $\mathcal{Z} \subseteq \mathbb{R}$, $S$ is a vector of shifts (e.g., sectoral growth of Chinese imports in high-income countries that are not the U.S.) with support in $\mathcal{S} \subseteq \mathbb{R}^J$, and $W$ is a vector of shares (e.g., sectoral employment shares in a commuting zone) with support in $\mathcal{W} \subseteq \mathbb{R}^J$. The natural number $J$ is the dimension of the vector of shifts and the vector of shares.
This model also has two latent variables. First, $\varepsilon$ is the outcome-equation error with a (possibly) infinite-dimensional support $\mathcal{E}_2$. In our empirical application, it can be interpreted as local policies that impact the growth of the manufacturing employment rate. Second, $\eta$ is a scalar unobserved variable in the treatment assignment equation with support $\mathcal{E}_1 \subseteq \mathbb{R}$. In our empirical application, it can be interpreted as local shocks in the demand for Chinese imports that are not captured by shocks in other high-income countries.
Moreover, this model has three key functions. First, $h_2: \mathcal{X} \times \mathcal{D} \times \mathcal{E}_2 \rightarrow \mathcal{Y}$ is the outcome function. Second, $h_1: \mathcal{Z} \times \mathcal{D} \times \mathcal{E}_1 \rightarrow \mathcal{X}$ is the treatment assignment function.\footnote{Our model allows for simultaneity between X and Y under additional assumptions. blundell2014 discuss which structural assumptions are necessary to write a simultaneous model as a triangular model.} Third, $h_0: \mathcal{S} \times \mathcal{W} \rightarrow \mathcal{Z}$ is the shift-share function, which is chosen by the researcher. As we see further in this paper, there is no need for the researcher to specify a function $h_0$, since she will only need the vector $W$ as the instrument.
Although the structural equations above are presented in levels for generality and clarity, these variables are typically expressed in first differences in most empirical applications of shift-share designs. Our framework is flexible enough to accommodate this specification. In particular, when the empirical setting identifies causal effects from changes in exposure to aggregate shocks—rather than from levels—our model can be reinterpreted with $Y$, $X$, and $Z$ representing temporal changes rather than levels. For instance, in autor2013, the identifying variation arises from differential trends across commuting zones, rather than static levels of exposure. For this reason, the outcome variable is the ten-year change in manufacturing employment, the treatment variable is the ten-year change in exposure to Chinese imports, and the structural functions $h_1$ and $h_2$ are defined directly for differenced variables. Consequently, we can only identify the effect of differential trade shocks on the growth rate of employment in each commuting zone. We cannot identify the effect of trade on employment levels.
In this section, we define our four parameters of interest. All objects are defined conditioning in $S = s$. In our empirical application, we treat each time period as a separate dataset. Consequently, it is as if we observed a single draw of the distribution of $S$, implying that conditioning on $S = s$ is basically conditioning on the available population.
The first target parameter is the Local Average Response (LAR) function, studied in altonji2005. It is defined as
It summarizes the marginal effect of $x$ on $Y$ over the population of $D$ and $\varepsilon$ for a given value of $X = x$. It captures the observable heterogeneity from the model and can be interpreted as a generalization of the Conditional Average Treatment Effect (CATE). In our empirical application, it captures the effect of marginally increasing the change in exposure to Chinese imports on the temporal change in the manufacturing employment rate.
The second target parameter is a single summary measure of the marginal effect of $x$ on $Y$: the Average Derivative (AD). It is studied by imbens2009 and is defined as
The Average Derivative integrates the Local Average Response function over the population of $X$ and summarizes the observable heterogeneity into a single object.\footnote{In Section (ref), we will relate the LAR and the AD to the 2SLS estimand.} It can be interpreted as a generalization of the Average Treatment Effect (ATE). In our empirical application, it captures the average effect of marginally increasing the change in exposure to Chinese imports on the temporal change in the manufacturing employment rate.
The third target parameter is the Average Structural Function (ASF), studied by blundell2003. It is defined as
Similar to the Local Average Response function, this parameter captures the observable heterogeneity of the model, but at the level of the outcome variable. In our empirical application, it captures how the percentage point change in manufacturing employment rate differs, on average, for different values of change in the Chinese import exposure. In other words, the Average Structural Function describes the average intertemporal change in manufacturing employment rate for a commuting zone that faced a growth in Chinese import exposure of $X = x$.
Our fourth target parameter is the Policy Effect, studied in imbens2009. It is defined as
where $\ell:\mathcal{X} \rightarrow \mathcal{X}$ is policy function chosen by the researcher. This parameter captures the average effect of introducing policy $l$ on the outcome $Y$. In our empirical application, the researcher could be interested in the effects of a policy $\ell$ that imposes an upper bound on exposure to Chinese imports, $X$. It could be through an import restriction on some specific sector (e.g., a policy that bans imports of cars from China) or aggregated in terms of exposure to all Chinese imports. This parameter is interesting, as it allows the researcher to capture the causal effects of policies that have never been introduced in real life. In Section (ref), we provide a detailed discuss about this parameter.
Lastly, note that, when the model is implemented using first-differenced variables (as is common in shift-share applications), the target parameters should be interpreted as marginal or average effects in changes, rather than in levels. For example, in our empirical application, the LAR, AD, ASF, and Policy Effect capture how increases in import exposure affect the decline in manufacturing employment over time, rather than the level of employment itself. This distinction is crucial for interpreting the results appropriately.
To identify the parameters described in the last section, we impose five assumptions. The first two assumptions allow us to identify a control function that will be used to identify all target parameters. Then, when we impose our third assumption, we can identify the Average Structural Function. Lastly, the addition of the fourth assumption allows us to identity the Local Average Response and the Average Derivative, while the addition of the fifth assumption allows us to identify policy effects.
Our primary assumption imposes the exogeneity of the vector of shares. It is closely connected to the identification assumption used by goldsmith2020 and hahn2024. Formally, it imposes the following restriction on our data-generating process.
Assumption (ref) says that the vector of shares, $W$, is independent of the treatment-assignment and outcome-equation errors, $\eta$ and $\varepsilon$, given the vector of shifts, $S$, and the vector of covariates, $D$.\footnote{An alternative identification strategy would impose exogenous shifts as done by borusyak2021. Challengingly, the shifts are the same for every region, implying that we cannot find exogenous variation at the regional level in a nonparametric setting. To circumvent this issue, borusyak2021 rely on linearity restrictions to derive an equivalence result between a region-level model and an industry-level model. Using the latter model, a researcher can explore “shift” variation across industries to identify the linear effect of interest. However, such an equivalence result is not trivial in a nonlinear setting such as ours. For this reason, shift-share designs with exogenous shifts are outside the scope of this paper.} In our empirical application, it imposes that employment in manufacturing would have trended similarly for regions that were more vs. less exposed to a possible shock in the previous period if there were no changes in exposure to Chinese imports.\footnote{goldsmith2020 and hahn2024 find that, when combined in a Bartik instrument, share variation does not seem to be exogenous in our China shock application. However, both groups of authors argue that it is still possible to analyze this empirical setting by directly using the shares as instruments without combining them into a Bartik instrument. Importantly, our identifying assumption adopts this approach of directly using the shares as instruments.} As noted by goldsmith2020, the Exogenous Shares assumption could be interpreted as a set of parallel trend conditions when the outcome is measured in changes.
Our second assumption imposes monotonicity of the treatment-assignment error, $\eta$.
Assumption (ref) is a generalization of the common IV monotonicity assumption imbens2009. In our setting, it requires that the function $h_1$ is strictly increasing or strictly decreasing in the unobserved variable of the treatment assignment equation, $\eta$. Combined with the fact that this treatment-assignment error is a scalar, it allows us to invert the function $h_1$ with respect to $\eta$. Similarly to the work of imbens2009, this step is essential to identifying the control function in our setting.
Our third assumption is necessary to connect our model with the shift-share structure present in our empirical application and restricts the number of observed shocks.
Assumption (ref) says that the vector of shifts $S$ is common across the entire population. Consequently, conditioning on $S$ is equivalent to condition on the observed population. To the best of our knowledge, all shift-share applications assume that there is only one common vector of shifts. For example, in our empirical application, the shift is a vector of changes in imports from China to high-income countries. Each entry of the vector corresponds to an industry sector, but the vector is common to all regions.
To state our next two assumptions, we need to define the following variable:
The random variable $V$ is based entirely on observable variables and, later, works as our control function. This result is shown in Proposition (ref).
Our fourth assumption imposes three regularity conditions and is connected to the assumptions in Theorem 6 by imbens2009.
Assumption (ref) is necessary to identify the Local Average Response (LAR) function and the Average Derivative (AD). Condition (ref).(ref) requires that the outcome function, $h_2$, to be continuously differentiable in the treatment variable, $x$, since this derivative appears in the definitions of the LAR function and the AD. Condition (ref).(ref) allows us to identify the LAR function for the entire support of $X$ and, then, integrate the LAR function over the distribution of $X$ to identify the Average Derivative. Finally, Condition (ref).(ref) is the weakest possible restriction that allows us to change the order of the derivative and the integral.
Lastly, our fifth assumption is a common support assumption.
Assumption (ref) imposes that the support of the treatment variable after the policy $\ell$ is imposed is contained in the observed support of the treatment variable. When combined with Assumptions (ref)-(ref), Assumption (ref) is sufficient to identify the Policy Effect.
In this section, we present the identification results for the target parameters listed in the previous section. Section (ref) provides nonparametric identification results, while Section (ref) relates our target parameters to the 2SLS estimand. Lastly, (ref) presents semiparametric identification results that connect directly with our proposed estimation and inference procedures in Section (ref).
Similarly to imbens2009, we adopt the control function approach to identification and estimation. We begin by identifying the control function variable. Then, we identify the average structural function, the local average response function, and the average derivative. Lastly, we identify the policy effect.
Our first proposition identifies our control function variable.
Proof. See Appendix (ref).
The intuition behind Proposition (ref) is that we want to clean out the unobserved endogenous variation of $X$ that is driven by $\eta$. To do that, we construct a proxy for $\eta$. When controlling for this proxy and the other observed variables ($D$ and $S$), we isolate the exogenous variation in $X$ that is driven by $W$. Consequently, we can identify its effects on $Y$. This reasoning is formalized by the last statement in Proposition (ref).
Before identifying our target parameters, we define the following function:
where the second equality follows from Equation (ref) and the third equality follows from Proposition (ref). Defining the function $m(x,d,v)$ simplifies our notation significantly because it appears in most of our identification results. Note that the function $m(x,d,v)$ is defined using observable variables only.
Our second proposition identifies the Average Structural Function (Equation (ref)).
Proof. See Appendix (ref).
Proposition (ref) states that the conditional expectation of the structural function $h_2$ evaluated at point $x \in \mathcal{X}$ (i.e., $\mu(x)$) is captured by the conditional expectation of the function $m$, defined in Equation (ref). This result is closely related to blundell2003. However, we impose a slightly different exogeneity assumption.\footnote{In blundell2003, they use the conditional independence assumption $\varepsilon\mid X,Z \sim \varepsilon\mid X, \eta$.}
Our third proposition identifies the Local Average Response function (Equation (ref)) and the Average Derivative (Equation (ref)).
Proof. See Appendix (ref).
The results in Proposition (ref) state that the conditional expectation of the derivative of the structural function $h_2$ in $x$ is captured by the conditional expectation of the derivative of the function $m$ in $x$.
Our fourth proposition identifies the Policy Effect (Equation (ref)).
Proof. See Appendix (ref).
Proposition (ref) says that the conditional expectation of the structural function $h_2$ when applying the policy transformation $\ell$ in the random variable $X$ is captured by the conditional expectation of the function $m$ when applying the same policy transformation.
Importantly, the expectations in Proposition (ref) are properly defined only when Assumption (ref) holds. This common support assumption is potentially a strong restriction, limiting our choice of policy functions in a fully nonparametric setting. To avoid this type of restriction in our empirical application, we use semiparametric assumptions as explained in Section (ref).
In this section, we discuss the causal interpretation of the Two-Stage Least Squares (2SLS) estimand in our model and compare this estimand against our target parameters. For brevity, we omit the extra covariates $D$ from the model. In this case, the 2SLS estimand is given by
Our fifth proposition connects the 2SLS estimand with the structural functions in Equations (ref) and (ref).
Proof. See Appendix (ref).
Proposition (ref) states that the 2SLS estimand identifies an average of the derivative of the outcome function (Equation (ref)) weighted by $\lambda(z,\eta)$.\footnote{This result adapts the result in angrist2000 for the case of a triangular model. It is also related to results derived by adao2019, borusyak2021, dechaisemartin2022, and hahn2024. Most of these authors analyze how to interpret the 2SLS in a linear, heterogeneous shift-share model, while we use a non-linear shift-share model. Importantly, borusyak2021 analyze a partially linear model, but they impose that our function $h_{0}$ has an inner product structure while we left it unrestricted (Equation (ref)).} This estimand identifies a convex combination of causal effects when $\lambda(z,\eta) \geq 0 $ for all $(z,\eta) \in \mathcal{Z}\times \mathcal{E}_1$. These weights are proportional to the first-stage effect, $\partial h_1(z,\eta)/\partial z$, and they will be nonnegative if, for all $(z,\eta) \in \mathcal{Z}\times \mathcal{E}_1$, either $\partial h_1(z,\eta)/\partial z \geq 0$ or $\partial h_1(z,\eta)/\partial z \leq 0$.\footnote{We identify the first-stage effect in Appendix (ref).}\textsuperscript{,}\footnote{hahn2024 derive necessary conditions for the 2SLS estimand to be a positively weighted average of causal effects in a linear heterogeneous treatment effects model under either the exogenous shares assumption or the exogenous shifts assumption. They argue that these necessary conditions are implausible in many empirical contexts. Consequently, analyzing nonlinear heterogeneous models like ours is methodologically relevant.} If the first-stage effect function has different signs for a positive mass of points in the support of $Z$ and $\eta$, then the 2SLS estimand faces negative weighting problems and is not weakly causal according to blandhol2022.
Even without negative weighting problems, the 2SLS estimand lacks a straightforward causal interpretation. Note that, instead of using the distribution of observable and unobservable variables like the LAR function and the AD parameter (Equations (ref) and (ref)), the 2SLS estimand places more weight on points where $\partial h_1(z,\eta)/\partial z$ is greater. Moreover, the weights used by the 2SLS estimand are not connected with policies motivated by economic theory, such as our policy effect (Equation (ref)).
Although the nonparametric identification results are valid for a wide class of structural functions, they have two main drawbacks. First, when connecting them to nonparametric estimators, we must estimate a complex function with many covariates (e.g., there are 397 industry sectors in our empirical application). Consequently, these estimators would suffer greatly from the curse of dimensionality. Second, the common support assumption significantly limits our choice of policy functions. To avoid these issues, this section describes a semiparametric identification strategy based on the methods proposed by chernozhukov2020.
Before stating the required assumptions for semiparametric identification, we must introduce some notation. Let $q_A$ be a vector of transformations, such as powers, splines, and interactions, referring to a random variable $A$.\footnote{Chen2007 provides a detailed review of sieve estimators, explaining power, spline and other series that may be used in our vector of transformations.} Now, define $$K_1(W,D) := q_W(W) \otimes q_D(D) \qquad \text{and} \qquad K_2(X,D,V):= q_X(X) \otimes q_D(D) \otimes q_V(V)$$ where $\otimes$ denotes the Kronecker product. Moreover, let $Q_{A}\left(\left. \tau \right\vert B\right)$ denote the $\tau$-th quantile of variable $A$ conditional on variable $B$.
Our first semiparametric assumption restricts the functional form of our structural functions (Equations (ref) and (ref)).
Assumption (ref) imposes that our structural functions follow a semiparametric quantile regression model.\footnote{We chose a Quantile Regression model, but one could opt for other semiparametric models. For example, chernozhukov2020 also derives identification results for Distribution Regression models.} According to chernozhukov2020, the quantile regression model is valid when the structural functions (Equations (ref) and (ref)) follow a restricted random coefficient model or a heteroskedastic normal system of equations.
Our second semiparametric assumption is a rank condition.
Assumption (ref) guarantees that the vector $\pi_2(U)$ is unique, as shown by chernozhukov2020.
Now, we briefly discuss our semiparametric identification strategy. Note that, under Assumptions (ref) and (ref), Equation (ref) implies that
where $\pi_{2} \coloneqq \int_{0}^{1} \pi_{2}\left(u\right) \, du$ according to chernozhukov2020. This result, when combined with Propositions (ref) and (ref), implies that we can semiparametrically identify the Average Structural Function and the Policy Effect as
when we add Assumptions (ref)-(ref) only. Importantly, we do not need the common support assumption to semiparametrically identify the Policy Effect because the Quantile Regression restrictions allow us to extrapolate outside the support of the treatment variable.
Lastly, we combine Equation (ref) with Proposition (ref) to semiparametrically identify the Local Average Response function and the Average Derivative as
when we add Assumption (ref).
Note that Equations (ref)-(ref) are key results to understand the estimation method proposed in Section (ref).
In this section, we propose a three-step estimation process to estimate the parameters of interest. Section (ref) uses a semiparametric estimator to estimate the control function, $V$, based on chernozhukov2020. Section (ref) proposes a flexibly parametric procedure to estimate the function $m(x,d,v)$ and its derivative, while Section (ref) uses the objects estimated in Section (ref) to estimate the parameters of interest. Lastly, Section (ref) describes a simple-to-implement estimation algorithm with a bootstrap inference procedure.
Below, we assume that we observe a sample $\left\lbrace Y_{i}, X_{i}, W_{i}, D_{i}\right\rbrace_{i = 1}^{N}$ of size $N \in \mathbb{N}$. Moreover, our sample is exposed to a single common shift shock $S_{i} = \Tilde{s}$ for all $i \in \left\lbrace 1, \ldots, N \right\rbrace$. Consequently, conditioning on the vector of shifts is equivalent to conditioning on our dataset.
Our first step is to estimate the values of the control function, $V_i = F_{X \mid W,D,S}(X_i \mid W_i, D_i, \Tilde{s})$, for $i \in \{1,\dots,N\}$. Following in chernozhukov2020, we estimate this distribution in a trimmed support $\overline{\mathcal{X}}$ to avoid far tails. We use bars to denote trimmed supports with respect to $X$, such as $\overline{\mathcal{X}\mathcal{W}\mathcal{D}} := \{(x,w,d) \in \mathcal{X}\times \mathcal{W}\times \mathcal{D}: x \in \overline{\mathcal{X}}\}$. Moreover, we denote the usual check function by $\rho_v(a) = (v - \mathds{1}\{a < 0\})\cdot a$.
Now, we can estimate the first stage as
for a small constant $\epsilon > 0$. chernozhukov2020 adjusts the limits of the integral in Equation (ref) to avoid tail estimation of quantiles.
Given Equation (ref), we estimate the control function variable as $$\hat{V}_i = \hat{F}_{X\mid W,D,S}(X_i \mid W_i, D_i, \Tilde{s}).$$
Here, we provide a flexibly parametric procedure to estimate the function $m(x,d,v)$ (Equation (ref)) and its derivative with respect to $x$. To estimate the function $m(x,d,v)$, we perform a OLS regression of $Y_i$ on $K_2(X_i,D_i,\hat{V}_i)$. From this OLS regression, we have that
where $\hat{\pi}_2$ is the vector of estimated parameters from the OLS regression.
Treating the dimension of $K_2(x,d,v)$ as fixed, we know the functional form of $\hat{m}(x,d,v)$. Then, we can take the derivative of $K_2(x,d,v)$ with respect to $x$ in order to estimate the derivative of the function $m(x,d,v)$. To simplify notation, let $m_x(x,d,v) := \partial m(x,d,v)/\partial x$. Therefore, the estimator for $m_x(x,d,v)$ is
For the final parametric step, we use the estimated functions from Section (ref) to construct estimators for the target parameters.
We start by constructing an estimator for the Average Structural Function (ASF):
In Equation (ref), the ASF estimator is a function of $x$, integrating over observed values of $D_i$ and $\hat{V}_i$.
Next, we construct an estimator for the Policy Effect of a given policy $\ell(\cdot)$:
Furthermore, using the estimator for the derivative of the function $m(x,d,v)$, we can construct an estimator for the Average Derivative (AD):
Lastly, to estimate the Local Average Response (LAR) function, we perform an OLS regression of $\hat{m}_x(X_i,D_i,\hat{V}_i)$ on $q_X(X_i)$. Then, our estimator for the LAR is
where $\hat{\pi}_X$ is the vector of estimated parameters from the OLS regression of $\hat{m}_x(X_i,D_i,\hat{V}_i)$ on $q_X(X_i)$.
Here, we provide an algorithm for the three-stage estimation procedure and an algorithm to perform uniform inference for the target parameters using the weighted bootstrap to estimate the standard errors.
Algorithm (ref) provides the three-stage estimation procedure. The first stage is identical to the first stage of the Quantile Regression specification proposed by chernozhukov2020. The later stages are based on procedures described in Sections (ref) and (ref). In the empirical application of chernozhukov2020, they find that their estimates are not sensitive to values of $M_1$ and $\epsilon$. Similarly, we also perform a robustness analysis with respect to those parameters in our empirical section.
Next, we present the inference procedure in Algorithm (ref). This procedure is based on the uniform inference procedure proposed by chernozhukov2020. We begin by performing a weighted bootstrap using the standard exponential distribution.\footnote{One could use any random variable satisfying $e \geq 0$, $\mathbb{E}[e] = 1$, $\mathrm{Var}(e) = 1$, and $\mathbb{E}\vert e \vert^{2+\delta} < \infty$ for some $\delta > 0$. See Assumption 3 in chernozhukov2020 for more details.} Next, we compute the standard errors using the interquartile range function. Then, for each bootstrap iteration, we compute the maximal $t$-statistics, so we can finally form $(1-\alpha)$-confidence bands in the last step.
In this section, we estimate the effects of the time evolution of Chinese imports on the temporal change of manufacturing employment in the United States using data previously analyzed by autor2013. In Section (ref), we provide the specification of our model. Section (ref) provides the results for both the Local Average Response (LAR) and the Average Derivative (AD), while Section (ref) provides the results for the Average Structural Function (ASF). Lastly, Section (ref) provides the results for the Policy Effect.
When analyzing the same empirical application as ours, hahn2024 reject the null hypothesis of constant and linear effects in this setting. They find statistical evidence to reject this hypothesis using either the “exogenous shares” identification approach goldsmith2020 or the “exogenous shift” identification approach borusyak2021. Their results highlight the importance of adopting a nonlinear heterogeneous treatment effect model such as the one we use in the following sections.
Moreover, goldsmith2020 and hahn2024 argue that, when exploring share variation to analyze the China shock application, any researcher should directly use the shares as instruments without combining them into a Bartik instrument. Our identification strategy and estimation algorithm follow exactly this approach.
autor2013 studies the effects of the time evolution of Chinese imports on the temporal change of manufacturing employment in the United States. Our main specification relies on the specification in Column (6) in Table 3 of autor2013, which includes the full set of covariates:
where $Y_{it}$ is the percentage point change in manufacturing employment rate for location $i$ and period $t$, $X_{it}$ is the change in Chinese import exposure in the United States per worker in a region, and $Z_{it}$ is the change in Chinese import exposure of other high-income countries.\footnote{In Appendix (ref), we plot the empirical distribution of $X_{it}$, and we plot a map of $X_{it}$ by commuting zone in the US.} Note that both $X_{it}$ and $Z_{it}$ are constructed as shift-share variables. $X_{it}$ is a linear combination of a normalized measure of the growth of imports from China to the US in industry $j$, $S_{jt}^{US}$, weighted by the contemporaneous start-of-period sector share of industry $j$ in each commuting zone $i$, $W_{ijt}$. Similarly, $Z_{it}$ is a linear combination of a normalized measure of the growth of imports from China to other high-income countries, $S_{jt}^{HI}$, weighted by the lagged sector share of industry $j$ in each commuting zone $i$, $W_{ijt-1}$. Finally, $D_{it}$ is a set of fifteen covariates.\footnote{The set of covariates includes start-of-decade labor force and demographic composition variables.} Importantly, when we connect this application to Assumption (ref), we have that $(\varepsilon_{it},\eta_{it}) \protect\mathpalette{\protect\independenT}{\perp} W_{i,t-1} \mid D_{it}, S_{it}^{HI}$.
In the data, we have a total of 722 locations ($N = 722$), 397 industries ($J = 397$), and two time periods: 1990-2000 and 2000-2007. autor2013 include a dummy for the second period as a covariate. However, including this indicator variable and using both time periods simultaneously would violate Assumption (ref). To avoid this issue, we split the sample into two disjoint datasets, one for each period, and estimate the parameters of interest separately.
Lastly, our estimator and inference procedures require choosing (i) sieves for shares, covariates, treatment variable, and control variable, (ii) trimming parameters, (iii) the number of bootstrap repetitions, and (iv) the confidence level. We present our choices below.
First Stage Specification. We choose $\epsilon = 0.01$, and $M_1 = 599$. We choose a linear specification for $K_1(W,D)$ because of the large number of shares and covariates.
Second Stage Specification. We choose a B-Spline bases for $q_X(X)$ and $q_V(\hat{V})$ with degree 3 and 4 knots. We interact those two basis as $q_X(X) \otimes q_V(\hat{V})$, and add $D$ linearly to construct $K_2(X,D,\hat{V})$. Consequently, the derivative of $K_2(x,d,v)$ with respect to $x$ will not depend on $d$.
Inference. We choose $B = 199$ for the weighted bootstrap procedure and $\alpha = 0.1$ to construct 90%-confidence bands.
In this section, we show the results for the AD and LAR function (Equations (ref) and (ref)). Here, the LAR captures the effect of marginally increasing the change in exposure to Chinese imports on the temporal change in the manufacturing employment rate. In other words, we compare the percentage point change in the manufacturing employment rate when $X = x$ against its counterfactual when $X = x + dx$, where $dx \rightarrow 0$. The AD should be interpreted as the average of these effects.
We start by comparing the AD results with the estimates obtained from 2SLS specifications. The first 2SLS specification is the one used by autor2013 and is given by
where $Period_t$ is a dummy that equals 1 when the period $t$ is equal to 2000-2007, and 0 otherwise. This specification is reported in Column (3) in Table (ref). We also report estimates for a non-pooled 2SLS specification:
We estimate this specification for both periods separately.
Table (ref) reports the results for the AD and the 2SLS estimates. For both periods, the AD estimates are positive, but not statistically significant. In other words, we do not reject the null that, on average, marginally increasing the exposure to growth in Chinese imports in a region will not affect the change in employment in manufacturing. This finding contradicts the results obtained by the pooled 2SLS method used by autor2013, as this specification yields a negative and statistically significant estimate of -0.303.\footnote{autor2013 estimate an effect of -0.596, but their specification includes regression weights accounting for the start-of-period commuting zones' shares of the national population. For simplicity, we do not use these weights in any estimates.} These differences can be interpreted using the results in Section (ref). According to Proposition (ref), the derivatives of $h_1$ with respect to $z$ are likely larger in magnitude in points where the derivative of $h_2$ with respect to $x$ is more negative.
When comparing the AD estimates with the 2SLS estimates for each period, the difference between the estimates gets smaller. For the period of 1990-2000, the 2SLS estimate is not statistically significant. However, for the period of 2000-2007, the 2SLS estimate is negative and statistically significant. Moreover, note that the pooled 2SLS estimate, reported in Column (3), is not a linear combination of the separate 2SLS estimates. These results suggests that the 2SLS specifications may suffer from negative weighting problems.
To better understand the weighting problems in the 2SLS specifications, we estimate the first-stage effects, which are a key component of the results in Proposition (ref). This derivative is identified in Proposition (ref) in Appendix (ref) and is estimated similarly to the methods described in Section (ref). First, we follow the semiparametric estimation procedure in chernozhukov2020 to estimate $\hat{\pi}_1(V)$, as in Section (ref). Then, we choose the spline basis $K_1(Z,D)$. Here, we follow the main specification in Section (ref), with the basis for $Z$ as a spline of degree 3 and 4 knots, and the vector $D$ entering linearly. Consequently, the derivative of $K_1$ with respect to $Z$ does not depend on $D$.
Figure (ref) shows the estimates of the first-stage effects associated with 2SLS regressions in the first two columns of Table (ref). Importantly, the estimates of the $h_{1}$ function change sign depending on the value of $z$ and $v$. Combining these with Proposition (ref), we have evidence that we cannot interpret the 2SLS estimates as weakly causal.\footnote{A similar conclusion is reached by hahn2024 using a different testing procedure.} Consequently, taking nonlinear treatment effects into consideration and focusing on the target parameters presented in Section (ref) are fundamental to understanding the impact of the China shock on the U.S. economy.
To deepen our understanding of these nonlinear treatment effects, Figure (ref) reports the LAR estimates for the periods of 1990-2000 and 2000-2007. The estimates in Panel (a) indicate that the LAR function for the period 1990-2000 appears to be linear, as we can fit constant functions inside its confidence bands. In particular, a constant null effect is not rejected. On the other hand, the estimates in Panel (b) indicate that the LAR function for the period 2000-2007 is nonlinear. In particular, we cannot fit any constant function inside its confidence bands, since the maximum lower bound is greater than the minimum upper bound.\footnote{We note that the confidence bands are very sensitive to the chosen specification, specifically to the chosen number of knots. Appendix (ref) plots five alternative specifications of our estimator. When we lower the number of knots, we find tighter confidence bands. In particular, the positive values of the LAR function for the period of 2000-2007 become statistically significant when we choose 3 knots, no matter the chosen degree of the splines.}
The point estimates in Panel (b) in Figure (ref) indicate that the effects for the period of 2000-2007 are positive for lower values of $X$, and they get negative for higher values. Specifically, for regions that faced lower exposure to growth in Chinese imports, a marginal increase in exposure to growth in Chinese imports leads to a greater intertemporal difference in manufacturing employment. In other words, for lower values of exposure to Chinese imports, employment in manufacturing decreases less than it would without the marginal increase in exposure. This type of nonlinearity underscores the importance of focusing on the target parameters presented in Section (ref) to understand the China shock's impact on the U.S. economy.
In this section, we present the results for the ASF (Equation (ref)). Since $Y$ and $X$ are defined as first differences, the ASF captures the relationship between a change in Chinese import exposure and the percentage point change in the manufacturing employment rate. Therefore, the ASF provides us with the average intertemporal change in the manufacturing employment rate for a commuting zone that experienced a growth in Chinese import exposure of $X = x$.
To illustrate the difference between nonlinear and linear models, we compare our specification with a linear estimator of the ASF. This linear estimator captures the ASF function if the true model (i.e., functions $h_{2}$ and $h_{1}$ in Equations (ref) and (ref)) is given by the linear model associated with the non-pooled 2SLS regressions in Equation (ref). The linear ASF estimator is given by
Figure (ref) plots the ASF estimates for our nonlinear specification (Section (ref)) and the linear specification in Equation (ref). Moreover, we report uniform 90%-confidence bands for our specification.
Panel (a) plots the ASF estimates for the period of 1990-2000. The estimates provided by our specification suggest that, on average, a change in exposure to Chinese imports is associated with a reduction in the manufacturing employment rate that lies between 0 and 2 p.p. We observe little heterogeneity for that period, allowing us to fit constant functions within the confidence bands. The red line indicates the ASF produced by the 2SLS estimator in Equation (ref). It also fits inside the confidence bands, indicating that the 2SLS specification in Equation (ref) is flexible enough to estimate the ASF in this case. The point estimates for our specification also do not differ meaningfully from those produced by the 2SLS estimator.
Panel (b) in Figure (ref) plots the ASF estimates for the period of 2000-2007. The ASF estimates suggest that, for lower values of growth in exposure to Chinese imports, the reduction in the manufacturing employment rate was smaller than for those that faced higher values of growth in exposure to Chinese imports. Similarly to the period 1990-2000, the red line, which indicates the estimates produced by the 2SLS estimator in Equation (ref), falls within the confidence bands for the period 2000-2007. On the other hand, the point estimates for our specification indicate that the reduction in the manufacturing employment rate was more negative than the estimates for the 2SLS specification for higher values of growth in exposure to Chinese imports.
In this section, we discuss the results for the Policy Effects (Equation (ref)). We are interested in the effects of changes in import tariffs. Section (ref) explains how we construct the policy of interest (function $\ell$ in Equation (ref)) as a function of the counterfactual increase in import tariffs, while Section (ref) interprets the estimated policy effects.
To connect the changes in import tariffs with changes in our measure of exposure to growth in Chinese imports, we use the results discussed by fajgelbaum2019. They estimate the elasticity of substitution between domestic goods and imports within a sector, $\hat{\kappa}$, of 1.19. They state the following relation between these variables:
where $M_{jt}$ is imports of products from sector $j$ in period $t$, $D_{jt}$ is expenditure of domestically produced goods, and $\phi_{jt}$ is the import tariff. Then, $\Delta \log (M_{jt}/D_{jt})$ is the log change in the ratio of imports from China to US in sector $j$ and period $t$ over domestically produced goods, driven by the log change in import tariffs $\Delta \log (1 + \phi_{jt})$. From Equation (ref), we can derive the following relation:
implying that we can construct a counterfactual $\frac{\Tilde{M}_{jt}}{\Tilde{D}_{jt}}$ from an increase of $\Delta\phi_j$ in the actual import tariff:
We perform a partial equilibrium exercise, in which we assume that consumers do not change their consumption of domestically produced goods, i.e., $D_{jt} = \Tilde{D}_{jt}$. Combining Equations (ref) and (ref), we have that
where we define $\Tilde{\phi}_j$ as the percentage increase in import tariff related to the actual import tariff. By isolating $\Tilde{M}_{jt}$ in Equation (ref), we find that
Moreover, according to the definition of $S_{jt}^{US}$ used by autor2013, we can construct a counterfactual $\Tilde{S}_{jt}^{US}$ using Equation (ref) as
where $L_{jt}$ is the share of the workforce working in sector $j$.
Furthermore, we are only interested in estimating the effects of a common increase in tariffs for all industries. For this reason, we impose that $\Tilde{\phi}_j = \Tilde{\phi}$ for all $j \in \{1,\dots,J\}$. Consequently, Equation (ref) and the estimated elasticity in fajgelbaum2019 imply that our policy function is given by
Using Equation (ref) as our policy function, our Policy Effect (Equation (ref)) captures the effect of a change in the actual exposure to growth in Chinese imports, driven by the percentage increase in the actual import tariff, on manufacturing employment at the end of the period. Consequently, we interpret this effect as a percentage point change in manufacturing employment rate at the end of the period driven by a percentage increase of $\Tilde{\phi}$ in the import tariff.\footnote{The factual world in Equation (ref), $\mathbb{E}[Y\mid S = \tilde{s}]$, contains the impacts of the actual growth in Chinese imports on the time change in the manufacturing employment rate. The counterfactual world in Equation (ref), $E[h_2(\ell(X),D, \varepsilon)\mid S = \tilde{s}]$, contains the impact of changing the import tariffs through the term $\tilde{\phi}$ and the impacts of the actual growth in Chinese imports through the term $M_{jt} - M_{jt-1}$. Consequently, their difference---the policy effect---captures the effect of increasing import tariffs on the temporal change in the manufacturing employment rate.}
To illustrate the difference between nonlinear and linear models, we compare our specification with a linear estimator of the policy effect. This linear estimator captures the policy effect if the true model (i.e., functions $h_{2}$ and $h_{1}$ in Equations (ref) and (ref)) is given by the linear model associated with the non-pooled 2SLS regressions in Equation (ref). The linear policy effect estimator is given by
We estimate the effects of the Policy Effect for values of $\Tilde{\phi}$ between 0.01 and 0.3, i.e., from an increase of 1% to an increase of 30% in import tariffs. Figure (ref) plots the estimates for the periods 1990-2000 and 2000-2007. In both periods, the estimated effects are not statistically significant.
Panel (a) in Figure (ref) plots the estimates of the Policy Effect for the period of 1990-2000. The point estimates are negative and very close to zero, while the confidence band becomes wider for higher values of the increase in import tariffs. Those estimates suggest that increasing the import tariffs for the period of 1990-2000 would not affect the manufacturing employment rate at the end of the period. These results are consistent with the results in Sections (ref) and (ref).\footnote{We note that the Policy Effect is sensitive to the chosen specification. Unlike our other objects of interest, not only the confidence bands, but also the point estimates of the Policy Effect are very sensitive to the choice of tuning parameters (e.g., the specification with degree of 3 and 3 knots). Appendix (ref) plots the Policy Effects for five alternative specifications.}
Panel (b) in Figure (ref) plots the estimates of the Policy Effect for the period of 2000-2007. Although nonsignificant, the point estimates are now positive, and the area of the confidence band is more concentrated in the positive region of the plot. One interesting finding that our point estimates reveal is that the policy effect, as a function of the tariff increase, appears to become constant after a certain level of import tariffs is reached. For example, a 15% increase in import tariffs has a similar effect to a 30% increase.
Importantly, the 2SLS estimates for the policy effect fall outside the confidence band of our nonlinear method. This result highlights how the linearity imposed by the 2SLS specification overlooks the heterogeneity found by our methods and may lead researchers to overestimate the effect of tariffs. Consequently, allowing for nonlinear heterogeneous treatment effects is fundamental to understanding the China shock's impact on the U.S. economy.
In this paper, we analyze nonlinear treatment effects in shift-share designs when shares are exogenous. To do so, we use a triangular model, similarly to imbens2009, and identify four parameters of interest using a control function approach. Moreover, we propose a flexibly parametric estimation procedure to estimate these parameters. In this section, we discuss the contexts in which our proposed methodology can be applied and further develop our empirical discussion.
Our methodology can be applied to any empirical problem that fits a shift-share design with exogenous shares. In our empirical application, we focus on the effects of the China shock on changes in manufacturing employment in the U.S., using data from autor2013. Our results highlight substantial treatment effect heterogeneity, which commonly used tools in shift-share applications are not designed to capture.
Regarding its empirical contribution, our work offers novel tools that answer policy-relevant questions in shift-share settings. For instance, we assess the impact of increasing import tariffs to compensate for the China shock, contributing to the literature on the effects of protectionist policies. Moreover, researchers can use our methodology to investigate the effects of different treatment assignment policies.
\singlespace