Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
93,441 characters · 21 sections · 40 citation commands
Heterogeneous Effects of Endogenous Treatments with Interference and Spillovers in a Large Network
\addcontentsline{toc}{section}{Abstract} { \centerline{Abstract}{6.8mm} This paper studies the identification and estimation of heterogeneous effects of an endogenous treatment under interference and spillovers in a large single-network setting. We model endogenous treatment selection as an equilibrium outcome that explicitly accounts for spillovers and derive conditions guaranteeing the existence and uniqueness of this equilibrium. We then identify heterogeneous marginal exposure effects (MEEs), which may vary with both the treatment status of neighboring nodes and unobserved heterogeneity. We develop estimation strategies and establish their large-sample properties. Equipped with these tools, we analyze the heterogeneous effects of import competition on U.S.\ local labor markets in the presence of interference and spillovers. We find negative MEEs, consistent with the existing literature. However, these effects are amplified by spillovers in the presence of treated neighbors and among localities that tend to select into lower levels of import competition. These additional empirical findings are novel and would not be credibly obtainable without the econometric framework proposed in this paper.
Keywords: causal inference, heterogeneity, interference, spillover, network
JEL classification: C31, C35, C36 }
The recent literature has seen a growing number of studies on causal effects in large single-network settings Cai3:15,Hu3:22,Li2:22,Leung:22ECTA,leung2022graph,Hoshino2:23_JSAS,Vazquez-Bare:23,Gao2:25. These studies examine how to learn the effects of exogenous and endogenous treatments in the presence of interference and spillovers. By contrast, much less is known about how to identify and estimate heterogeneous treatment effects in this environment, disentangling the heterogeneous effects of an individual’s own treatment from those of neighbors’ treatments.
Causal inference in this sophisticated environment requires (i) a model that characterizes individuals’ treatment decisions in equilibrium and (ii) a model that describes how treatments affect outcomes in the presence of spillovers. Both treatment decisions and outcomes generally depend on the behavior of others. For example, when assessing the effects of import competition on local labor markets treated as nodes, the labor market response in one region (node) may influence neighboring regions (nodes) through labor commuting/migration. Moreover, local labor market composition, which determines the extent of exposure to trade shocks, may itself depend on conditions in neighboring regions. Hence, a credible model must accommodate interference in both treatment assignment and outcome determination.
We begin by modeling joint treatment decisions under incomplete information and derive conditions for the uniqueness of the rational expectations equilibrium, building on the literature on discrete games on networks Lee3:14,Xu:18,Lin2:24. Uniqueness of equilibrium enables point identification of the treatment decision parameters and, more importantly, the propensity score. Leveraging this property, we then identify the treatment effects by adapting identification strategies from the causal inference literature. Following the network causal inference literature Leung:22ECTA,Hoshino2:23_JSAS, we take advantage of an exposure mapping, a low-dimensional statistic that summarizes interference effects. Under this assumption, we show that heterogeneous causal effects are identified without imposing additional functional form restrictions on the outcome model.
We estimate the model using the generalized method of moments (GMM). To study the asymptotic properties of GMM under network dependence, we adopt the $\psi$-dependence framework developed by Kojevnikov3:21, which establishes pointwise laws of large numbers, central limit theorems, and variance estimators that are robust to network dependence. Due to the nonlinear nature of our model, however, stronger conditions are required for valid statistical inference. In particular, we rely on the uniform law of large numbers Sasaki:25 in the $\psi$-dependence framework of KMS. Building on these, we show that the GMM estimator is consistent and converges stably to a mixed normal distribution. Moreover, we demonstrate that standard inference procedures remain valid even when the asymptotic variance-covariance matrix is random. Related results can be found in Andrews:05, Kuersteiner2:13, and Kuersteiner2:20ECMA. Because our asymptotic analysis is conducted conditionally on the common shocks that generate the network, the resulting estimator enjoys desirable asymptotic properties regardless of the underlying network formation model.
We illustrate our method using the Autor3:13 dataset on Chinese import competition and manufacturing employment shares across U.S. commuting zones (CZs) as nodes. Treating each CZ as an independent open economy, their two-stage least squares (2SLS) analysis documents a large negative effect of import competition on manufacturing employment. However, because CZs are geographically linked, interference may arise both in the extent of exposure to trade shocks and in labor market responses. We model sectoral composition choices as outcomes of a rational expectations equilibrium on a spatial network. The results indicate strategic substitution in sectoral choices, with regions tilting away from sectors facing greater competition from Chinese imports. We then examine the direct MEEs of import competition on local labor markets. We find substantial heterogeneity across CZs depending on network exposure: the exposure effects differ markedly between CZs with and without treated neighbors, as well as between CZs experiencing above- versus below-median import competition. These findings highlight the importance of accounting for both network-induced interference and heterogeneity when evaluating treatment effects.
This paper contributes to the literature on network causal inference that employs the $\psi$-dependence framework for statistical inference Leung:22ECTA,Hoshino2:23_JSAS,Gao2:25. Our approach differs from these studies in two respects. First, we focus on marginal responses that allow exposure effects to vary with unobserved heterogeneity. Second, we adopt a conditional $\psi$-dependence framework by conditioning on the common shocks that jointly generate the covariates and the network structure. As a result, the limiting distribution of our GMM estimator is mixed normal, with a random variance-covariance matrix that reflects the randomness induced by these common shocks.
This paper is also related to the literature on endogenous treatment with interference confined within groups or markets: Hoshino2:23JoE study how the delinquency of an opposite-gender best friend affects a student’s GPA; Balat2:23 examine how competition among airlines influences air pollution across U.S.\ cities. In these settings, interference is localized, which allows the authors to identify outcomes under all possible treatment configurations and to conduct the corresponding comparisons. In contrast, we consider a setting in which all individuals are connected through a large single network.
The remainder of the paper is organized as follows. Section (ref) presents the model setup. Section (ref) establishes identification. Section (ref) proposes the estimators, and Section (ref) analyzes their large-sample properties. Section (ref) evaluates the finite-sample performance of the proposed methods. Section (ref) presents an empirical application, and Section (ref) concludes. All mathematical details are relegated to the appendix.
Our observed data consist of $n$ observations $\{(Y_i,X_i,D_i,Z_i)\}_{i=1}^n$, where $Y_i$ denotes the outcome, $X_i$ the vector of covariates, $D_i$ the binary treatment status, and $Z_i$ the vector of exogenous variables. We allow $Z_i$ to include $X_i$ as a subvector, along with additional excluded exogenous variables that serve as instruments. In addition to these variables, we observe the network structure linking the $n$ individuals as nodes, represented by an adjacency matrix $A_n$, where $A_{ij}=1$ indicates that nodes $i$ and $j$ are connected and $A_{ij}=0$ otherwise. Wherever there is little risk of confusion, we suppress the subscript $n$ and simply write the adjacency matrix as $A$. We rule out self-spillover effects by assuming $A_{ii}=0$ for all $i$.
In a network setting, it is natural to allow an individual’s potential outcomes to depend on the treatment status of other individuals. Let $\mathbf{D}=(D_i)_{i=1}^n$ denote the collection of treatment assignments, and let $Y_i(\mathbf{D})$ be the potential outcome for node $i$ under the joint treatment status $\mathbf{D}$. Given the treatment status $D_i$ of node $i$ and the collection $\mathbf{D}_{-i}=(D_1,\ldots,D_{i-1},D_{i+1},\ldots,D_n)$ of treatment assignments excluding $D_i$, we model the outcome-generating process as
and $\epsilon_i^{(D_i,\mathbf{D}_{-i})}$ denotes the productivity shock for node $i$ under the treatment configuration $(D_i,\mathbf{D}_{-i})$. This shock is unobserved by the other nodes, as well as by the econometrician.
The $n$ nodes make simultaneous treatment decisions in an environment of incomplete information. The type of node $i$ is defined by the pair $(Z_i,\nu_i)$, where $Z_i$ is publicly observed by all nodes, while $\nu_i$ is private information known only to node $i$. The network structure $A$ is also publicly observed. Let $n_i=\sum_{j} A_{ij}$ denote the number of neighbors of node $i$.
Given the treatment decisions of other nodes $\mathbf{D}_{-i}$, node $i$ receives the payoff
if it chooses to take the treatment ($D_i=1$), and zero otherwise ($D_i=0$). Here, $\beta_D$ denotes the vector of coefficients associated with $Z_i$, and $\lambda$ captures the spillover effect from treated neighbors. Let $\bm{\beta}_1 = (\beta_D',\lambda)'$ for short hand.
Because node $i$ does not observe other nodes’ private information $\nu_{-i}$, it must base its decision on beliefs about others’ treatment choices. Accordingly, the best response of node $i$ is given by
We introduce the shorthand notation $P_i(\mathbf{Z},A) = \Pr(D_i^* = 1 \mid \mathbf{Z}, A)$ for the propensity score. The rational-expectations (Bayesian Nash) equilibrium choice probabilities are given by the vector $\mathbf{P}(\mathbf{Z},A) = \big(P_1(\mathbf{Z},A), \ldots, P_n(\mathbf{Z},A)\big)$, which solves the following system of equations:
The existence of a solution to (ref) follows from Brouwer’s fixed-point theorem. Moreover, uniqueness of the equilibrium can be established via a contraction mapping argument. To formally establish these properties, we impose the following assumption.
Assumption (ref)(i) is commonly imposed in the literature on games with incomplete information Brock2:02,Bajari4:10,Xu:18,Lin2:24. Assumption (ref)(ii) extends the standard instrument exogeneity condition to a network setting. Under this assumption, we can write $F_{\nu \mid \mathbf{X}}$ in place of $F_{\nu \mid \mathbf{Z}}$. Regarding Assumption (ref)(iii), with an abuse of notation, we rewrite the system of equations in (ref) as $\mathbf{P} = \varphi(\mathbf{P})$. Lee3:14 show that when the selection peer effect $\lambda$ is moderate, the norm of the Jacobian matrix of $\varphi$ is strictly less than one. This contraction property guarantees the existence and uniqueness of the equilibrium. We formally summarize this discussion in the following lemma.
See Appendix (ref) for proof.
Thanks to the uniqueness property established in Lemma (ref), we can directly identify the equilibrium choice probability as $ P_i(\mathbf{Z},A) = \Pr(D_i = 1 \mid \mathbf{Z}, A). $ We now proceed to identify the model parameters. We refer to the treatment choice equation as the first stage, and the outcome-generating process as the second stage.
Given the uniqueness property already established, the first stage reduces to a standard logistic regression in the network framework, similar to that studied in the literature Xu:18. In light of Assumption (ref)(i), the equilibrium propensity score can be written as
Applying the logistic inversion yields
With the shorthand notation $ \phi_i(\mathbf{Z},A) = \left( Z_i', \; n_i^{-1} \sum_{j} A_{ij} P_j(\mathbf{Z},A) \right)', $ consider the following matrix invertibility assumption.
This assumption is the standard rank condition and corresponds to Assumption 4 of Xu:18 and Assumption 7 of Lin2:24, among others. First, the uniqueness of equilibrium implies that the quantity $P_i(\mathbf{Z},A)$ is identified. Since $\phi_i(\mathbf{Z},A)$ depends only on the first-stage controls $Z_i$ and the propensity scores $\{P_j(\mathbf{Z},A)\}_{j=1}^n$, this object is also identified. Moreover, the right-hand side of (ref) is linear in the parameter vector $\bm{\beta}_1=(\beta_D',\lambda)'$. Consequently, the invertibility of the matrix implies identification of the first-stage structural parameters, as formally stated in the following theorem.
See Appendix (ref) for proof.
Let $V_i = F_{\nu \mid \mathbf{X}}(\nu_i)$. We note three properties of this variable. First, under Assumption (ref), the random variable $V_i$ follows a uniform distribution on $[0,1]$ and is independent of $V_{-i}$. Second, since $V_i \mid \mathbf{X} \sim \mathrm{Uniform}(0,1)$ does not depend on $\mathbf{X}$, it follows that the vector $\mathbf{V}$ is independent of $\mathbf{X}$. Finally, under this notation, $D_i = 1$ if and only if $P_i(\mathbf{Z},A) > V_i$.
We now proceed with the identification of the second-stage parameters. In what follows, we maintain the following assumption.
Assumption (ref)(i) extends Assumption (ref)(ii) from Section (ref). It requires that the instrument $\mathbf{Z}$ be exogenous in both the first and second stages, conditional on $\mathbf{X}$. In general, the presence of multiple equilibria may cause the equilibrium propensity score to vary in a non-smooth manner, in which case the existence of a continuous player-specific instrumental variable—say, $\mathbf{Z}_{1} = (Z_{11},\ldots,Z_{n1})$—is not sufficient to satisfy Assumption (ref)(ii). However, once uniqueness of equilibrium is established under the logistic error specification (Lemma (ref)), the implicit function theorem applies. As a result, the existence of at least one continuous coordinate in the player-specific instrument $Z_i$ is sufficient for Assumption (ref)(ii).\footnote{As $Z_{i1}$ enters $P_i(\mathbf{Z},A)$ through the linear index $Z_i'\beta_D$, and since our goal is to establish continuity of $\mathbf{P}$ with respect to this player-specific continuous coordinate $\mathbf{Z}_1$, we may, without loss of generality, restrict attention to the case in which $X_i$ is empty and $Z_i=Z_{i1}$. First, rewrite the system of equations in (ref) as $\mathbf{P}(\mathbf{Z},A)=\varphi(\mathbf{P}(\mathbf{Z},A),\mathbf{Z})$. Define $\varPhi(\mathbf{P}(\mathbf{Z},A),\mathbf{Z})=\mathbf{P}(\mathbf{Z},A)-\varphi(\mathbf{P}(\mathbf{Z},A),\mathbf{Z})$. Fix $\mathbf{Z}=\mathbf{z}$ and let $\mathbf{P}(\mathbf{z})=\mathbf{p}$. It is immediate that $\varPhi(\mathbf{p},\mathbf{z})$ is continuously differentiable with respect to each component of $\mathbf{p}$ and $\mathbf{z}$. The equilibrium condition implies that $\varPhi(\mathbf{p},\mathbf{z})=0$. It therefore remains to verify that $\nabla_{\mathbf{p}}\varPhi(\mathbf{p},\mathbf{z})$ is invertible. Let $F_i=F_{\nu|\mathbf{X}}(z_i\beta_D+\lambda\cdot n_i^{-1} A_{ik}p_k)$. Noting that $F_i'(\nu)=F_i(\nu)\cdot(1-F_i(\nu))$, we obtain $ {\partial F_i} / {\partial p_j} = F_i\cdot(1-F_i)\cdot \lambda\cdot n_i^{-1} A_{ij}. $ We can then decompose $\nabla_{\mathbf{p}}\varPhi(\mathbf{p},\mathbf{z})$ as $\nabla_{\mathbf{p}}\varPhi(\mathbf{p},\mathbf{z})=\nabla_{\mathbf{p}}\mathbf{p}-\nabla_{\mathbf{p}}\varphi(\mathbf{p},\mathbf{z})=I-\lambda D_\mathbf{p} \widetilde{A}$, where $D_{\mathbf{p}}=\text{diag}\{F_1\cdot(1-F_1),\dots,F_n\cdot(1-F_n)\}$ and $\widetilde{A}$ denotes the row-normalized adjacency matrix. Consider the matrix infinity norm $\lVert\cdot\rVert_\infty$. We have $\lVert \lambda D_\mathbf{p} \widetilde{A} \rVert_\infty\leq |\lambda| \lVert D_\mathbf{p} \rVert_\infty \lVert \widetilde{A}\rVert_\infty=|\lambda|\cdot 1/4<1$. Hence, $I-\lambda D_\mathbf{p}\widetilde{A}$ is invertible. By the implicit function theorem, there exists a unique and continuously differentiable function $g^*$ such that $\mathbf{p}=g^*(\mathbf{z})$. Finally, by uniqueness of equilibrium, $\mathbf{p}$ is a reduced-form function of $\mathbf{z}$, and therefore $\mathbf{P}(\mathbf{Z})=g^*(\mathbf{Z})$ is continuously differentiable in $\mathbf{Z}$. } This feature is common in the market entry literature, where the distance between a firm’s headquarters and a given market serves as a cost shifter that affects the firm’s own entry decision but does not directly influence its competitors’ profits, except through the firm’s entry decision.
In a single large-network setting, an individual can in principle have $2^n$ distinct potential outcome distributions, corresponding to all possible treatment configurations in the network. As a result, it is impossible to identify all parameters without imposing additional restrictions. Following the network causal inference literature Cai3:15,Leung:22ECTA,Hoshino2:23_JSAS, we introduce a low-dimensional summary statistic, referred to as the exposure mapping $T(i,\mathbf{D},A)$. Whenever there is little risk of confusion about the underlying network, we use the simplified notation $T_i(D_i,\mathbf{D}_{-i})$. Introducing this statistic amounts to assuming that
In other words, all interference effects influence potential outcomes solely through the exposure mapping. Consider the following examples commonly considered in the literature.\footnote{In the literature on strategic interactions outside the network context, interference is often assumed to operate through an aggregate state in empirical applications Menzel:2015. For example, Berry:1992 study how the number of entrants affects firms’ profits and, in turn, their entry decisions in airline markets. Similarly, Brock2:01 analyze how an individual’s binary choice depends on the proportion of members in a reference group who take a given action.}
Let $l_A(i,j)$ denote the path distance between nodes $i$ and $j$ on the network represented by $A$. In addition, for $K>0$, let $N_A(i,K)=\{j:l_A(i,j)\leq K\}$ denote the $K$-neighborhood of node $i$. We also introduce the subnetwork $A_{N_A(i,K)}=(a_{kl}:k,l\in N_A(i,K))$ and subvector $\mathbf{D}_{N_A(i,K)}=(D_{k}:k\in N_A(i,K))$. Withe this notation, consider the following assumption.
In what follows, we denote the $K$-neighborhood $N_A(i,K)$ by $N_i$ for brevity. We also write $\mathbf{V}_{N_i}=(V_j)_{j\in N_i}$ and $\mathbf{p}_{N_i}=(p_j)_{j\in N_i}$ for the corresponding subvectors of random variables and their realizations, respectively. Letting $\partial_{\mathbf{p}_{N_i}}:=\partial^{|N_i|}/\big(\prod_{j\in N_i}\partial p_j\big)$, we define the observationally identifiable quantity
where $\mathbf{l}_i'=\{j\in N_i\mid d_j=0\}$ denotes the set of local nodes whose treatment status equals zero. The definition of this quantity depends on the local network structure $N_i$. However, conditional on a given local network structure and the exposure mapping specified, the denominator is fully determined and involves no additional uncertainty.
The object in (ref) can be shown to identify the marginal exposure response (MER), defined as
in the spirit of Hoshino2:23JoE, but extended to a network setting. This result is formally stated in the following theorem.
See Appendix (ref) for proof.
As in general causal inference frameworks, an individual’s potential outcomes $Y_i(d,\mathbf{d}_{-i})$ cannot be identified per se. Theorem (ref) shows that their conditional mean (ref), given the information generated by observed characteristics $\mathbf{X}$ and unobserved heterogeneity $\mathbf{V}_{N_i}$, can be identified by leveraging the exposure mapping $T_i$ as a low-dimensional summary statistic.
While Theorem (ref) establishes identification of the MER under general conditions, this generality poses substantial challenges for practical estimation. The difficulty is that the estimand $\psi(t,\mathbf{x},\mathbf{p}_{N_i},N_i)$ depends on the local neighborhood $N_i$ and the vector $\mathbf{p}_{N_i}$, among other objects. In other words, it depends on both the local network structure and a vector of continuous random variables whose dimension depends on it. This feature severely limits the effective (local) sample size when the identifying expression in Theorem (ref) is used directly for estimation. In light of this practical challenge, the present subsection introduces additional model restrictions that enable dimension reduction and thereby facilitate estimation.
Reduce the dimension of the MER in (ref) as
We will show that (ref) can be identified by the following modification of (ref):
where $N_i^\circ=N_i\backslash \{i\}$ and $\Tilde{\mathbf{l}}' = \{j \in N_i^\circ|d_j=0\}$. Furthermore, it will be shown that this estimand (ref) can be significantly simplified under the following set of restrictions.
Assumption (ref)(i) is satisfied under the anonymous interference assumption in Li2:22, Example (ref) as in Cai3:15, and Example (ref). In general, the exposure mapping may depend on treatment status in an arbitrary manner. Under Assumption (ref)(i), however, we can clearly separate the role of an individual’s own treatment status $D_i$ from that of neighbors’ treatment assignments, thereby simplifying the latter. We note that the function $f_i$ remains quite general, as it may be multidimensional even under Assumption (ref)(i). Assumption (ref)(ii) imposes an additively separable structure between the observed covariates $\mathbf{X}_i$ and the unobservables $\mathbf{V}_i$. Even with this restriction, however, the exposure indices $T_i(\mathbf{d})$ may still affect potential outcomes in an arbitrary manner.
The following corollary to Theorem (ref) formally establishes the dimension reduction and simplification for second-stage identification.
See Appendix (ref) for proof.
Corollary (ref) implies that $\overline{\psi}(t,\mathbf{x},p,N_i)$ no longer depends on $N_i$. In light of this result, we suppress this argument and simply write $\overline{\psi}(t,\mathbf{x},p)$ whenever this corollary is applied. Define the marginal exposure effect (MEE) of changing $t$ to $t'$ by
For instance, the MEE with $t=(0,0)$ and $t'=(1,0)$ can be interpreted as the direct effect of own treatment in the absence of treated neighbors. The MEE with $t=(0,0)$ and $t'=(0,1)$ can be interpreted as the effect of exposure to treated neighbors on untreated individuals.
The following section proposes estimation strategies for these exposure effects.
Following the identification steps in Section (ref), the first stage reduces to a logit discrete choice model on the network. Under Assumption (ref)(i), we may employ the log-likelihood
Note that $P_i(\mathbf{Z},A)$ is not directly observed in the data. Under Assumption (ref)(i) and the treatment selection model, we can compute the propensity scores $P_i(\mathbf{Z},A)$ by solving for the fixed point of the following system of equations:
Moreover, given the functional form of $P_i(\mathbf{Z},A)$, we can compute the derivatives of $\mathbf{P}$ with respect to the parameters as follows:
where $\mathbf{Z}_k=(Z_{ik})_{i=1}^n$, with $Z_{i0}=1$ for all $i$, corresponding to the intercept term. \footnote{In implementation, we clip $P_i$ at $10^{-5}$ from both ends. Note that Assumption (ref) (iii) with $|\lambda|<4$ is sufficient to ensure invertibility of the relevant matrix. To see this, let $D=\text{diag}\!\left(\frac{1}{P_i(1-P_i)}\right)$. Factoring out $D^{-1}$, it suffices to examine the invertibility of $I-\lambda D^{-1}A$. This matrix is invertible if the product of $\lambda$ and the spectral radius of $D^{-1}A$ lies strictly between $-1$ and $1$. Moreover, the spectral radius of $D^{-1}A$ is bounded above by the maximum row sum, which equals $\max_i P_i(1-P_i)\le 1/4$. Hence, $|\lambda|<4$ is a universally sufficient condition.}
Given the derivatives with respect to the parameters, we obtain the first-order conditions for the maximum likelihood estimator (MLE) as
These first-stage score functions are subsequently used as moments in our Generalized Method of Moments (GMM) estimation.
For the second stage, we consider a concrete specification following Section (ref). Assumption (ref) imposes only an additively separable structure. We parametrize the functions $\mu_X^{(t)}$ and $\mu_\epsilon^{(t)}$ and adopt a linear-in-parameters specification that is nested within Assumption (ref).
Under Assumption (ref), the estimand from Corollary (ref) becomes
The parametric assumption further implies that $\overline{\psi}$ now depends only on $x$, rather than on the covariates of the entire sample $\mathbf{x}$.
We next discuss how to estimate the coefficients $(\beta_X^{(t)},\beta_\epsilon^{(t)})$. In what follows, we focus on the case $t_1=1$; the case $t_1=0$ follows analogously. Assumptions (ref) and (ref)(i) then imply the following linear equation
However, this approach is infeasible because $V_i$ is unobserved. Alternatively, consider the following linear estimating equation.
Taking the conditional expectation on the both sides and noting $t_1=1$, we have
where the second equality follows from Assumptions (ref), (ref) (i), and (ref) (ii), and the last equality is due to $E[e_i^{(t)}|\mathbf{X},\mathbf{V}_{N_i}]=0$ under Assumption (ref) (ii) and $E[V_i|V_i < P_i]=P_i/2$ for the uniform random variable.
This observation motivates the feasible linear estimating equation
where $E[\Tilde{e}_i^{(t)}|X_i,P_i]=0$ by construction. Thus, as the first-order conditions of the resulting least-squares criterion, we can incorporate the following moments into our GMM estimation:
For the case $t_1=0$, an analogous argument applies by replacing $E[V_i \mid V_i < P_i]$ with $E[V_i \mid V_i \ge P_i]$, which yields $1+P_i$ in place of $P_i$.
Since our objective is to learn about the effect of changing the exposure mapping from $t$ to $t'$, we need to estimate the parameters corresponding to both $t$ and $t'$. Thus, we use the following moment function
where $\bm{\beta} = (\beta_D,\lambda,\beta_X^{(t)},\beta_p^{(t)},\beta_X^{(t')},\beta_p^{(t')})$. Its dimension, $\dim(\bm{g})=(\dim(Z)+1)+1+2(\dim(X)+1)$, corresponds to the number of unknown parameters in $\bm{\beta}$, yielding just identification in the current specification. With this said, the general GMM estimator takes the form of
where $\widehat{\bm{g}}_n(\bm{\beta})=n^{-1}\sum_i^n \bm{g}((Y_i,X_i,D_i,T_i,\mathbf{Z}),\bm{\beta})$, and $\widehat{\Xi}_n$ is a sequence of positive definite weight matrices.
In practice, it is desirable to implement a two-step GMM procedure to improve efficiency in general over-identified settings. In the two-step GMM, one begins by obtaining the first-step estimate
We take advantage of the network HAC estimator proposed by KMS to construct an optimal weighting matrix.\footnote{To prevent the HAC estimator from being non–positive semi-definite, we apply an eigenvalue (diagonal) decomposition to the variance–covariance matrix and replace any eigenvalues smaller than $10^{-10}$ with this threshold. In our simulations, a small number of iterations required this adjustment when the sample size was 250. However, once the sample size increased to 1000, no iterations required such adjustments, suggesting that the variance estimator converges to a strictly positive-definite matrix.} With the shorthand notation $\bm{g}_i(\widehat{\bm{\beta}}^{(1)})=\bm{g}((Y_i,X_i,D_i,T_i,\mathbf{Z}),\widehat{\bm{\beta}}^{(1)})$, the network HAC estimator of the variance of $\bm{g}$ is given by
$\omega_n(s)=\omega(s/b_n)$, $\omega$ is a pre-specified kernel function, and $b_n$ denotes the bandwidth. We then obtain the two-step GMM estimator
With the parameter estimate $(\widehat{\beta}_X^{(t)},\widehat{\beta}_\epsilon^{(t)})$, we in turn estimate the MER by the plug-in
The marginal exposure effect of changing the exposure mapping from $t$ to $t'$ is estimated by
Consider a triangular array of random variables, where $\mathcal{N}_n=\{1,\dots,n\}$ denotes the index set in the $n$th row. Accordingly, we extend the notation $W_i$ to $W_{n,i}$ to emphasize that the random variables belong to the $n$th row of the triangular array. Boldface notation $\mathbf{W}_n=(W_{n,i})_{i=1}^n$ is used to denote the entire vector of random variables in the $n$th row.
Samples are drawn from a probability space $(\Omega,\mathcal{F},P)$. For each $n$, let $\mathcal{C}_n\subseteq\mathcal{F}$ denote a $\sigma$-subalgebra representing common shocks. In particular, these common shocks determine both the adjacency matrix $A_n$ and the covariates $\{X_{n,i},Z_{n,i}\}_{i=1}^n$. Moreover, there may exist a $\sigma$-algebra $\mathcal{C}$ representing common factors that is contained in every $\mathcal{C}_n$, that is, $\mathcal{C}\subset\bigcap_{n\ge 1}\mathcal{C}_n$.
We first fix some basic notation. Define a metric \(d_h\) on $\mathbb{R}^{h \times k}$ by \[ d_h(\mathbf{W}, \widetilde{\mathbf{W}}) = \sum_{r=1}^h \lVert w_r - \widetilde{w}_r \rVert_2 \] for \(\mathbf{W}, \widetilde{\mathbf{W}} \in \mathbb{R}^{h \times k}\) written as \(\mathbf{W} = (w_1, \dots, w_k)\) and \(\widetilde{\mathbf{W}} = (\widetilde{w}_1, \dots, \widetilde{w}_k)\), where \(\lVert \cdot \rVert_2\) denotes the Euclidean norm. Also define its conditional \(L^p\) norm $\lVert \cdot \rVert_{\mathcal{C}_n, L^p}$ with respect to \(\mathcal{C}_n\) by \[ \lVert W_i \rVert_{\mathcal{C}_n, L^p} = \left( E \left[ \sum_{r=1}^k |W_{ir}|^p \,\middle|\, \mathcal{C}_n \right] \right)^{1/p} \] for a vector \(W_i = (W_{i1}, \dots, W_{ik}) \in \mathbb{R}^k\). Its unconditional counterpart is denoted by \(\lVert \cdot \rVert_{L^p}\).
Let \(\mathcal{P}_n(h, h', s)\) denote the collection of all pairs of subsets of nodes of sizes \(h\) and \(h'\) such that the path distance between any two nodes from different subsets is at least \(s\). Define \[ \mathcal{L}_{w,a} = \left\{ f : \mathbb{R}^{w \times a} \to \mathbb{R} \;:\; \lVert f \rVert_{\infty} < \infty,\; \mathrm{Lip}(f) < \infty \right\} \] as the set of bounded Lipschitz functions, where \(\lVert f \rVert_{\infty} = \sup_{x} |f(x)|\) and \(\mathrm{Lip}(f)\) denotes the Lipschitz constant of \(f\). Following KMS, we now formally define the notion of the conditional \(\psi\)-dependence.
The central idea of the \(\psi\)-dependence is that the influence of distant observations on local observations diminishes as the path distance between them increases. KMS develop a fundamental asymptotic theory under this notion of network dependence, of which we can take advantage. Thus, as a first step in our analysis, we characterize the \(\psi\)-dependence properties of the individual moment conditions for our network model. Conditioning on \(\mathcal{C}_n\), the moment condition for sample \(i\) takes the treatment status \(D_{n,i}\), the exposure mapping \(T_{n,i}\), and the outcome \(Y_{n,i}\) as its arguments. Accordingly, we will establish the \(\psi\)-dependence of $\{W_{n,i}\}$ conditional on \(\mathcal{C}_n\), where $ W_{n,i} = (Y_{n,i},\, T_{n,i},\, D_{n,i}). $
Note that $\{D_{n,i}\}_{i=1}^n$ is a transformation of $\mathbf{Z}_n=\{Z_{n,i}\}_{i=1}^n$ and $\bm{\nu}_n=\{\nu_{n,i}\}_{i=1}^n$, as $ D_{n,i}=\sigma_{n,i}^D(\bm{\nu}_n,\mathbf{Z}_n). $ Conditional on $\mathcal{C}_n$, $D_{n,i}$ can be viewed as a function of $\bm{\nu}_n$. Thus, we suppress the argument of $\mathbf{Z}_n$ and simply write $D_{n,i}=\sigma_{n,i}^D(\bm{\nu}_n)$ for brevity. We also define $ D_{n,i}^{(s)}=\sigma_{n,i}^D(\bm{\nu}_n^{(s,i)}), $ where $\bm{\nu}_n^{(s,i)} = \left( \mathbbm{1}\{j\in N_{A_n}(i,s)\}\cdot \nu_{n,j} \right)_{j=1}^n$ replaces the distant shock with zero. In the following proposition, this quantity will be shown to play a role in characterizing the $\psi$-dependence of $\{D_{n,i}\}$.
Similarly, the outcomes \(\{Y_{n,i}\}_{i=1}^n\) are transformations of \((\mathbf{D}_n, \mathbf{X}_n, \mathbf{Z}_n, \bm{\nu}_n, \{\bm{e}_n^{(t)}\}_{t \in \mathcal{T}})\), where the dependence on \(\mathbf{D}_n\) arises through the exposure mapping. Suppressing \(\mathbf{X}_n\) and \(\mathbf{Z}_n\), we denote this as $ Y_{n,i} = \sigma_{n,i}^Y(\mathbf{D}_n, \nu_{n,i}, \{ e_{n,i}^{(t)} \}_{t \in \mathcal{T}}). $ We note that while \(\mathbf{D}_n\) depends on the entire vector \(\bm{\nu}_n\), only \(\mathbf{D}_n\), \(\nu_{n,i}\), and \(\{ e_{n,i}^{(t)} \}_{t \in \mathcal{T}}\) directly enter \(Y_{n,i}\) in the potential outcome, either as part of the regressor \(V_i = F_{\nu|\mathbf{X}}(\nu_{n,i})\) or as error terms. We similarly define $ Y_{n,i}^{(s)} = \sigma_{n,i}^Y(\mathbf{D}_n^{(s,i)}, \nu_{n,i}, \{ e_{n,i}^{(t)} \}_{t \in \mathcal{T}}), $ with $\mathbf{D}_n^{(s,i)} = \left( \mathbbm{1}\{j\in N_{A_n}(i,s)\}\cdot D_{n,j} \right)_{j=1}^n$. Although \(\nu_{n,i}\) and \(\{ \epsilon_{n,i}^{(t)} \}\) are i.i.d. random variables conditional on \(\mathcal{C}_n\), the treatment assignments \(\{ D_{n,i} \}\) are \(\psi\)-dependent. Thus, we cannot directly apply the previous result. Nevertheless, we can still show that \(\{Y_{n,i}\}\) is conditionally \(\psi\)-dependent. We state the result in the following theorem.
We now move on to characterize the conditional \(\psi\)-dependence of the exposure mapping \(\{T_{n,i}\}\). There are two distinct features to note. First, unlike the outcome variables, \(T_{n,i}\) depends only on \(\mathbf{D}_n\). Second, \(T_{n,i}\) is a vector-valued function. Bearing these two features in mind, we write
to emphasize its vector-valued nature. Similarly, we define
We omit the proof of this proposition, as it follows the same structure as the proof of Proposition (ref). Finally, we are ready to establish the conditional \(\psi\)-dependence of \(\{W_{n,i}\}\) with \(W_{n,i} = (Y_{n,i}, T_{n,i}, D_{n,i})\).
Observe that \(W_{n,i}\) is a vector formed by stacking \(Y_{n,i}\), \(T_{n,i}\), and \(D_{n,i}\), with the first two components constructed from the last. It therefore follows that the conditional \(\psi\)-dependence of \(\{W_{n,i}\}\) inherits both the functional form and the \(\theta\)-sequence from those of \(\{Y_{n,i}\}\) and \(\{T_{n,i}\}\). As such, the result immediately follows from propositions (ref), (ref), and (ref).
We now establish the asymptotic properties of our estimator and discuss their implications for statistical inference. We extend the analysis of Sasaki:25 to derive large-sample properties conditional on the common sub-$\sigma$-field $\mathcal{C}$. The key advantage of this conditional framework is that it allows the underlying network-generating process to be non-deterministic. As shown by KMS, this approach accommodates a broad class of models, including network formation models, random fields on graphs, and conditional dependency graphs.
We are going to use the following version of the definition of stable convergence, as given by Daley2:03.
There exist several equivalent formulations of stable convergence, as noted by Daley2:03 and Häusler2:15. We summarize these equivalences in Proposition (ref) in Appendix (ref). A special case arises when $W$ is $\mathcal{C}$-measurable, in which case stable convergence is equivalent to convergence in probability Häusler2:15. Since the true parameter vector $\bm{\beta}_0$ is treated as fixed, this corollary implies the usual convergence-in-probability result for consistency.
We begin by establishing the consistency of the GMM estimator. Let $N_{A_n}^{\partial}(i,s)$ denote the set of nodes that are exactly $s$ path-lengths away from node $i$, and define
With this notation, we impose the following assumptions.
Parts (i) and (ii) of Assumption (ref) correspond to Assumptions 3.1 and 3.2 of KMS, respectively. These conditions are used to establish the pointwise law of large numbers. Notably, Assumption (ref)(ii) characterizes the trade-off between network density and cross-sample dependence. A standard requirement for the consistency of nonlinear GMM estimators is the uniform convergence of the GMM objective function Newey2:94, Vaart:98. Accordingly, we invoke Assumptions (ref)(iii)--(v) in addition to construct a $\delta$-net and extend the pointwise law of large numbers to a uniform law of large numbers under network interference similarly to Sasaki:25.
Recall from (ref) that our sample objective function is given by \[ \widehat{Q}_n(\bm{\beta}) = \widehat{\bm{g}}_n(\bm{\beta})'\widehat{\Xi}_n\widehat{\bm{g}}_n(\bm{\beta}). \] Since our model explicitly depends on the network structure, we consider a sequence of objective functions that are $\mathcal{C}_n$-measurable. The corresponding population objective function $Q_n$ is defined by
where $\bm{g}_n(\bm{\beta}) = E[\bm{g}((W_{n,i},X_{n,i},\mathbf{Z}), \bm{\beta}) \mid \mathcal{C}_n]$, and $\Xi_n$ denotes a random matrix that is $\mathcal{C}_n$-measurable, has finite elements, and is positive definite almost surely.
As in standard nonlinear GMM or M estimation frameworks, a crucial requirement is the uniform law of large numbers (ULLN) for the objective function Newey2:94, Vaart:98. Following the conditional framework of KMS, we establish the conditional ULLN for the objective function by leveraging the conditional moment results of Sasaki:25. Lemma (ref) in the appendix states this result formally. Specifically, for all $\epsilon > 0$,
From this almost sure convergence, it follows that \[ \sup_{\bm{\beta} \in \Theta} \big| \widehat{Q}_n(\bm{\beta}) - Q_n(\bm{\beta}) \big| \overset{P}{\to} 0. \]
Define the true parameter vector as the solution to the sequential conditional moment conditions in the sense that
Here, recall that the identification (i.e., the “only if” implication) is justified on economic grounds in Section (ref). Let $\lVert \cdot \rVert_F$ denote the Frobenius norm, defined by $\lVert B \rVert_F = \sqrt{\mathrm{trace}(B'B)}$ for any real matrix $B$. The following theorem establishes that the GMM estimator converges in probability to the true parameter value.
Next, we are going to establish the asymptotic normality of the GMM estimator. By Assumption (ref) (iv) and the conditional $\psi$-dependence of $\{ W_{n,i} \}$, Corollary (ref) in Appendix (ref) shows that $c'\bm{g}((W_{n,i},X_{n,i},\mathbf{Z}), \bm{\beta}_0)$ is also conditionally $\psi$-dependent for any vector $c \in \mathbbm{R}^{\mathrm{\dim}(g)}$, with functional $\psi^{cg}$ and coefficients $\{ \theta_{n,s}^{cg} \}$, defined in Corollary (ref). Let \[ S_n = \sum_{i \in \mathcal{N}_n} \bm{g}((W_{n,i},X_{n,i},\mathbf{Z}), \bm{\beta}_0) \quad \text{and} \quad \Omega_{g,n} = \mathrm{Var}(S_n \mid \mathcal{C}_n). \] Given $c \in \mathbbm{R}^{\mathrm{dim}(g)}$, define $\sigma_n^2(c) = \mathrm{Var}(c'S_n \mid \mathcal{C}_n) = c'\Omega_{g,n}c$. For notational convenience, we also define
We now impose the following conditions to establish asymptotic normality.
By Assumptions (ref) (i)–(ii) and the conditional $\psi$-dependence of $c'\bm{g}((W_{n,i},X_{n,i},\mathbf{Z}), \bm{\beta}_0)$, we extend the conditional central limit theorem in Theorem 3.2 of KMS to the multivariate case of $\bm{g}$ via the Cramér–Wold theorem. Assumptions (ref) (iii)--(v) impose standard regularity conditions for nonlinear GMM estimators that are widely adopted in the literature Newey2:94.
The following theorem establishes the conditional asymptotic normality for $\widehat\beta$.
Conditional convergence in distribution to normal distributions generally implies a non-normal unconditional limiting distribution. At first glance, this seems to preclude standard normal-based inference; nevertheless, valid statistical inference remains feasible. Let $\widehat{\Omega}_{\beta}$ denote a consistent estimator of the limiting covariance matrix $\Omega_{\beta,0}$. We consider testing the $k$-dimensional null hypothesis $H_0\!: r(\bm{\beta}_0) = 0$ against the alternative $H_1\!: r(\bm{\beta}_0) \neq 0$, where $r(\cdot)$ is continuously differentiable in a neighborhood of $\bm{\beta}_0$. Let $R(\bm{\beta}) = \nabla_{\bm{\beta}} r(\bm{\beta})$ be a $\mathrm{dim}(g) \times k$ matrix with full column rank, where $k \leq \mathrm{dim}(g)$. The Wald statistic is then given by
To operationalize the result of Theorem (ref), the following theorem establishes the asymptotic validity of the Wald test, even when the variance–covariance matrix is random due to the conditioning on $\mathcal{C}$.
This theorem is readily applicable in statistical inference in practice.
This section investigates the finite-sample performance of our proposed method through simulation studies.
We consider two types of network structures in our simulation study. The first is the ring network, in which each node is connected to its two immediate neighbors. The second is the random geometric graph (RGG), constructed following Leung:22ECTA.
For the RGG model, a network is generated as follows. Let $\xi_i=(\xi_{1i},\xi_{2i})$, $i=1,\ldots,n$, be independent random variables drawn from $\mathrm{Uniform}([0,1]^2)$. Nodes $i$ and $j$ are connected if $\lVert \xi_i - \xi_j \rVert_2 \le r_n$, where $r_n=\left(\frac{\kappa}{\pi n}\right)^{1/2}$ and $\kappa$ denotes the expected degree. To ensure realism, we set $\kappa$ equal to the average degree of the empirical network used in our application, which is approximately $5.63$.
The exposure mapping is specified by $ T_i(D_i,\mathbf{D}_{-i})=\left(D_i,\mathbbm{1}\left\{\sum_{j\neq i}A_{ij}D_j>0\right\}\right), $ so that $t\in\mathcal{T}=\{(1,1),(1,0),(0,1),(0,0)\}$. For the treatment decision and outcome models, we largely follow the setup of Hoshino2:23JoE, adapted to our framework:
The covariates and instrumental variables are drawn from standard normal distributions, \( X_i, Z_i \sim N(0,1) \). The parameter values for the outcome model are set as \( \beta_Y^{(1,1)} = \beta_Y^{(1,0)} = \beta_Y^{(0,1)} = (2,1) \) and \( \beta_Y^{(0,0)} = (1,2) \), where \( \beta_Y^{(t)} = (\beta_{X0}^{(t)}, \beta_{X1}^{(t)})' \). For the selection model, the parameters are set to \( (\beta_{D0}, \beta_{D1}) = (-1, 2) \) and \( \lambda = 1 \). The unobservable is generated as
where $\beta^{(1,1)}_{p}=1.5,\beta^{(1,0)}_{p}=\beta^{(0,1)}_{p}=1,\beta^{(0,0)}_{p}=0.5$, and $e_i\overset{i.i.d}{\sim}N(0,1)$.
For estimation, we implement our two-step GMM estimator using the HAC variance estimation procedure described between Equations (ref) and (ref).\footnote{For the just-identified case, a two-step estimation procedure is not necessary. Nevertheless, we implement the two-step method for the sake of generality in our code, so that it can readily accommodate more general over-identified settings as extensions.} A crucial component of this estimator is the choice of the bandwidth and the kernel function. Following KMS, we set the bandwidth as
and use the Parzen kernel:
Here, \(\text{avg.deg}\) denotes the average degree of nodes in the given network. We fix \(\epsilon = 0.05\) and consider \( c \in \{ 1.7, 1.8, 1.9, 2.0, 2.1, 2.2 \} \) as in KMS. We report results with \( c = 2.0 \), as the conclusions are qualitatively similar across the other values of \( c \).
For each simulation design, we consider sample sizes $n=250$ and $n=1000$ to assess the convergence rate. In particular, root-$n$ convergence is evaluated by comparing the ratio of the standard deviations of the estimates across these two sample sizes to two. Each experiment consists of 1,000 Monte Carlo replications. In each replication, we compute the estimate $\widehat{\bm{\beta}}$, its standard error, and construct the $95\%$ asymptotic confidence interval as $\widehat{\bm{\beta}} \pm \operatorname{diag}\!\left(\widehat{\Omega}_\beta / n\right)^{1/2}$, which we use to evaluate coverage probabilities. We additionally compute the corresponding marginal exposure responses (MERs), $\widehat{\overline{\psi}}(t,x,p)$, fixing $x=1$ and considering $p\in\{0.2,0.5,0.8\}$.
The simulation results for $\bm{\beta}$ under the ring network design are reported in Table (ref). We make the following observations. First, the bias when $n=1000$ is generally smaller than when $n=250$, reflecting improved estimation accuracy with larger sample sizes. Second, the standard deviations for $n=1000$ are approximately half the size of those for $n=250$, indicating a $\sqrt{n}$ convergence rate, which is consistent with our theoretical results. Finally, the Monte Carlo coverage frequencies of the $95\%$ confidence intervals are close to the nominal level, particularly when $n=1000$.
Table (ref) reports the simulation results for the MER estimator $\widehat{\overline{\psi}}$ under the ring network design. Similar to the results for $\bm{\beta}$, the MER estimator also exhibits a $\sqrt{n}$ convergence rate, and its coverage accuracy improves as the sample size increases.
One potential concern the reader may have is that the favorable results reported above may be driven by the simplicity of the ring network structure. Tables (ref) and (ref) report the results for networks generated by the RGG model. Even though the network structure is now more irregular, the estimator continues to exhibit a $\sqrt{n}$ convergence rate, and our inference procedure achieves coverage probabilities close to the nominal $95\%$ level when $n=1000$.
The impact of the global economy on local economies has long been of great interest to economists. In particular, following China's accession to the WTO in 2001, the effect of Chinese exports on U.S.\ local labor markets has received considerable attention. A central concern in using observational data for such empirical analysis is the potential endogeneity of local exposure to imports from China. In their notable study, Autor3:13 exploit supply-side factors in international trade to address the endogeneity of local import competition and provide evidence of the effects of such exposure. Under the assumption that the U.S.\ local labor markets are presumably not directly linked to China's supply-side shocks, their analysis provides convincing empirical evidence.
Autor3:13 compile a dataset of 722 commuting zones (CZs) in the mainland United States as units of observation representing local labor markets. They construct a first-differenced sample from the panel of three decadal periods to capture the growth between 1991–2000 and 2000–2007.\footnote{In their original work, they use import volumes for the years 1991, 2000, and 2007 from the UN Comtrade Database and scale the first-differenced values by 10/9 and 10/7 to obtain decadal figures.} The key regressor is the change in Chinese import exposure per worker in a CZ, and the outcome of interest is the decadal change in the manufacturing employment share of the working-age population. We illustrate the CZs on a map in Figure (ref).
While the existing literature assumes that the effects of import competition on local labor markets are confined within each commuting zone (CZ), in reality, spillovers across CZs are likely. Moreover, the extent of local import competition may not only be endogenously determined but also interdependent with that of neighboring CZs, for instance, due to economies of scale in import costs within adjacent geographic areas. In other words, the degree of exposure is determined as an equilibrium choice under spatial interactions. To address these features, we employ a model that accommodates spillovers and interference, while also allowing for endogeneity and heterogeneity.
Incorporating the network structure, we consider the following specification:
The outcome variable $Y_i$ denotes the decadal change in the manufacturing employment share of the working-age population in the $i$-th CZ. The treatment variable $D_i$ represents the per-worker exposure to the change in Chinese imports in the $i$-th CZ. As instruments, following Autor3:13, we use the composition and growth of Chinese imports in eight other developed countries. The idea is to exploit variation in supply-side factors within China to isolate the corresponding exogenous variation in the treatment variable. For the control variables $X_i$, we include the period dummy and the share of manufacturing in the start-of-period employment.
The exposure mapping is specified as \[ T_i(D_i,\mathbf{D}_{-i}) = \left( D_i, \, \mathbbm{1}\!\left\{\sum_{j\neq i} A_{ij} D_j \geq 1 \right\} \right). \] As discussed in Example (ref), this exposure mapping allows us to capture heterogeneity in the treatment effect by accounting for whether a unit is exposed to a treated neighbor. We focus on the direct MEE while controlling for the treatment status of neighbors as
The empirical model (ref) allows the MEE (ref) to be heterogeneous depending on whether there are treated neighbors. Specifically, the difference between $\theta_{1}$ and $\theta_{0}$ captures the heterogeneity of the direct effect depending on the presence of a treated neighbor. Moreover, the model (ref) also accommodates unobserved heterogeneity in the MEE (ref). In particular, the difference $\beta_p^{(t')}-\beta_p^{(t)}$ captures such heterogeneity, which itself may vary depending on whether a neighbor is treated. While this empirical model differs from the specification in (ref) used for our simulation study, it is designed to capture these empirically relevant features for the current application.
Before proceeding with the estimation of the MEEs, we first present the routine overlap check. Figure (ref) displays the histograms of the estimated propensity scores. Unlike conventional models, however, our setting involves multi-dimensional treatments due to the coexistence of direct and neighbor effects. The left panel of the figure shows the histograms for $t=(0,0)$ and $t=(1,0)$, while the right panel shows those for $t=(0,1)$ and $t=(1,1)$. In both cases, we observe substantial overlap in this data set.
Figure (ref) displays the estimated MEE values along with their pointwise 95% confidence intervals. Before discussing the results, it is worth noting that 83.59% of the CZs have at least one treated neighbor.\footnote{ More specifically, 72.85% and 94.32% of the first and second first-differenced samples, respectively, have treated neighbors. } First, in this light, the MEEs are significantly negative for the vast majority of CZs. This finding is consistent with earlier results in the literature and aligns with the idea that stronger import competition from China displaces manufacturing workers into other sectors. Second, we observe that the MEE without a treated neighbor is always smaller (in absolute value) than the MEE with a treated neighboring zone, suggesting that spillovers amplify the trade shock in local labor markets. Finally, the downward-sloping pattern of the MEEs indicates that CZs with higher values of $V$ face more severe trade effects. Note that CZs with larger values of $V$ are those less likely to be exposed to import competition. It is, therefore, natural that the CZs that suffer more from import competition are also those that tend to select out of that treatment status.
This paper develops an econometric framework for identifying and estimating heterogeneous effects of an endogenous treatment in the presence of network interference and spillovers. We characterize endogenous treatment choices as outcomes of a rational‐expectations equilibrium that incorporates spillover considerations, and we provide conditions ensuring that such an equilibrium exists and is unique. Building on this structure, we obtain identification of heterogeneous marginal exposure effects (MEEs), effects that vary with both neighbors’ treatment statuses and with unobserved heterogeneity, and we establish estimation procedures together with their asymptotic properties.
Applying our methods to the impact of import competition on U.S. local labor markets, we document negative MEEs in line with prior findings. Importantly, we show that these negative effects intensify when neighboring regions are also exposed to treatment and for areas that tend to self‐select into lower levels of import competition. These results, which reveal nuanced spillover‐driven heterogeneity, would not have been accessible without the econometric tools developed in this paper.