Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
56,851 characters · 13 sections · 26 citation commands
Causal Interpretation of Linear Social Interaction Models with Endogenous Networks
Many studies have estimated treatment spillover effects using an experimental approach that randomly assigns individuals to treatments to exploit the exogeneity of their peers' treatment status (e.g., oster2012determinants, cai2015social, paluck2016changing, angelucci2019incentives). In most of these studies, although the treatment assignments are randomized, they often ignore the endogenous nature of existing networks, which can hinder the correct identification of the causal effects of treatment spillovers.
For example, paluck2016changing investigated the impact of an anti-conflict intervention program on adolescents' attitudes and their propagation through friendship networks. In friendship networks, non-combative students are more likely to be friends with students of the same mindset (i.e., homophily). Therefore, even if we can observe a group of students engaging in conflict-mitigation activities together, determining whether this reflects the causal spillover effect from the program's participant friends or a mere correlation between their personalities (or both) is not straightforward. In this example, although the treatments were completely randomized, whose treatment status matters to whom remained endogenously determined.\footnote{ In other words, actual treatment exposure consists of a product of exogenous and endogenous factors. borusyak2020non studied this situation in a more general framework than ours. }
One possible approach to circumvent the network endogeneity problem is to randomly assign peers, as in sacerdote2001peer, zimmerman2003peer, guryan2009peer, and booij2017ability. A clear limitation of this approach is that it is not feasible in most situations. Meanwhile, several studies consider specific regression models of social interactions to deal with the endogeneity of networks (e.g., goldsmith2013social, hsieh2016social, johnsson2021estimation, jochmans2022peer). Their results are not directly applicable to the potential outcome framework with heterogeneous treatment effects, and simply extending them to a nonparametric causal model is often challenging due to the curse of dimensionality. Therefore, considering empirical feasibility, it would be preferable to focus on a simple regression approach based on existing networks. Although the causal interpretation of linear regression models under network interference has been studied, for example, in baird2018optimal, DITRAGLIA20231589, and vazquez2022identification,vazquez2023causal, none of these accounts for network endogeneity. The purpose of our study is to bridge this gap.
This paper is organized as follows. In Section (ref), to clarify the problem and find possible remedies, we consider a toy model in which treatments are randomly assigned and each unit has exactly one potential interacting partner. We assume that whether one's outcome is influenced by his partner or not is determined endogenously. We show that when the network endogeneity is overlooked, although the ordinary least squares (OLS) estimand can correctly capture the average direct effect from one's own treatment, it is biased for the average spillover effect due to the correlation between potential outcomes and network connectivity. To account for the endogeneity issue, we use an instrumental variable (IV) approach. We employ the partner's treatment assignment as an IV, which is a valid IV for spillover exposure. Based on this IV, we prove that a two-stage least squares (2SLS) regression has a local average treatment effect (LATE) type causal interpretation. Furthermore, as a more efficient alternative to 2SLS, we propose a weighted least squares (WLS) method with the same LATE interpretation.
Section (ref) extends the discussion in Section (ref) to a general model that allows each unit to have multiple peers. Although potential heterogeneity in the network structure complicates the analysis, we can confirm that essentially the same results as in the toy model hold. We also show that in this general model, the OLS estimand suffers not only from the endogeneity bias but also from a negative-weights problem.
In general, statistical inference for fully heterogenous treatment effect models under non-identical data structures is a challenging task. In Section (ref), we consider several empirically tractable alternatives to perform statistical inference, including a randomization test. In Section (ref), we revisit the data in paluck2016changing and demonstrate that the spillover effect indeed exists even after controlling for the network endogeneity. In the supplementary appendix, we present proofs of the technical results and results of some Monte Carlo simulations.
Suppose there are two non-overlapping samples: the focal sample $\mathcal{I}$ and the partner sample $\mathcal{J}$. Only the focal sample is used to estimate causal effects. Each focal unit $i \in \mathcal{I}$ has exactly one potential partner $j(i) \in \mathcal{J}$, and $j(i)$ cannot be a partner of the other focal units. Suppose that $j(i)$ does not have to always interact with $i$. For each pair $(i, j(i))$, let $A_{j(i)}$ denote a dummy variable indicating whether $j(i)$'s treatment affects $i$'s outcome. The outcome and treatment of interest are denoted by $Y \in \mathbb{R}$ and $D \in \{0,1\}$, respectively. For example, this situation occurs when $(i, j(i))$ represents a couple, $A_{j(i)}$ indicates whether they are living together, and $D_{j(i)}$ indicates their partner's lifestyle, such as diet (e.g., vegetarian), with $Y$ being a health outcome. We define the effective treatment spillover variable as $S_i \coloneqq A_{j(i)}D_{j(i)}$.
Let $Y_i(d,s)$ be the potential outcome when $D_i = d$ and $S_i = s$. The observed outcome is written as
To focus solely on the endogeneity issue caused by link connectivity $A_{j(i)}$, we rule out cases of self-selected treatments. Specifically, we consider the following Bernoulli-type experimental setting:
If we allow for the dependence among the treatment assignments, it causes a significant difficulty in interpreting our estimands.\footnote{When the treatment assignments are dependent, even the average direct effect from one's own treatment may involve an identification issue (cf. Theorem 2 in vazquez2022identification).} For this reason, studies in the literature of network interference often focus on the Bernoulli design (e.g., hu2022average; li2022random). As we will demonstrate later, even in this ideal experimental setting, the spillover effects cannot be correctly identified in the presence of network endogeneity (not to mention in other experimental designs).
We also assume that the data are independent and identically distributed (IID).
We allow $A_{j(i)}$ to correlate with $Y_i(d,s)$, which is the source of endogeneity of concern. In the example above, it is natural to imagine that $A_{j(i)}=1$ is more likely if $i$ and $j(i)$ share similar lifestyle preferences, which suggests that $A_{j(i)}$ and $Y_i(d,s)$ are dependent.
We first discuss the selection bias in the OLS regression. Suppose that a researcher estimates a linear regression model of $Y_i$ on $(D_i, S_i)$ using the OLS estimator. Define
This characterization implicitly relies on the IID assumption. Note that since $D_i$ and $S_i$ are independent, we do not consider including a cross-term between them.
Here, we define the "direct" and "spillover" treatment effects as follows:
We summarize the causal interpretations of $(\beta_d^{ols}, \beta_s^{ols})$ as follows:
As shown in result (i), $\beta_d^{ols}$ is the weighted average of the average direct treatment effects for $A_{j(i)}= 1$ and $A_{j(i)}= 0$, which has a clear causal interpretation. Hence, ignoring the endogeneity in $A_{j(i)}$ is harmless for identifying the average direct effect. By contrast, result (ii) shows that the spillover effect parameter $\beta_s^{ols}$ is a summation of the average spillover effect for $A_{j(i)}= 1$ and the selection bias term. This bias term originates from the correlation between $Y_i(d,0)$ and $A_{j(i)}$.
To circumvent the selection bias in the OLS, we use the IV method. For the choice of the IV for $S_i$, one obvious candidate is $D_{j(i)}$, which is completely exogenous according to the experimental design and is a determinant of $S_i$. More importantly, in practical terms, it is not necessary to search for other IV candidates. Then, based on this IV, we investigate the causal interpretation of the 2SLS estimand $(\beta_0^{2sls}, \beta_d^{2sls}, \beta_s^{2sls})$, which we define as
where $\mathbb{L}(S_i \mid D_i, D_{j(i)}) = \gamma_0 + \gamma_d D_i + \gamma_s D_{j(i)}$ is the linear projection of $S_i$ onto $(D_i, D_{j(i)})$ obtained from the first-stage regression $(\gamma_0, \gamma_d, \gamma_s) = \operatorname*{argmin}_{a_0, a_d, a_s} \mathbb{E}[(S_i - a_0 - a_d D_i - a_s D_{j(i)})^2]$. Because $D_i$ and $S_i$ are independent, instrumenting for $S_i$ does not alter the interpretation of $\beta_d^{2sls}$; that is, $\beta_d^{2sls} = \beta_d^{ols}$ holds. For a causal interpretation of $\beta_s^{2sls}$, we obtain the following result:
It is known that we can improve the estimation efficiency of the LATE parameter by weighting each observation according to its compliance probability (e.g., joffe2003weighting, coussens2021improving). The same discussion applies here. Moreover, because we can precisely identify each unit's compliance status (i.e., $A_{j(i)}$), the resulting estimator is reduced to a simple least squares regression of $Y_i$ on $(D_i, D_{j(i)})$ for those satisfying $A_{j(i)} = 1$. Then, we define the following weighted least squares (WLS) estimand:
Note that if every $i$ is affected by his/her partner, then the OLS and WLS coincide. The next proposition provides a causal interpretation of the WLS estimand.
Because of the sample selection, the interpretation of the direct effect $\beta_d^{wls}$ is slightly different from that for $\beta_d^{ols}$ and $\beta_d^{2sls}$. The average direct effect for the non-compliers $\mathbb{E}[\delta_i(0) \mid A_{j(i)}= 0]$ is not incorporated in $\beta_d^{wls}$. Meanwhile, the interpretation of $\beta_s^{wls}$ is the same as that of $\beta_s^{2sls}$.
\paragraph{Numerical example} Here, we briefly demonstrate the severity of the selection bias. The simulation setup is as follows: $D_i, D_{j(i)} \sim \text{Bernoulli}(0.5)$, $A_{j(i)} = \bm{1}\{e^{(U_i + U_{j(i)})}/[1 + e^{(U_i + U_{j(i)})}] > 0.5\}$, where $U_i, U_{j(i)} \sim \text{Uniform}(-1,1)$, and $Y_i = \xi_i + U_i$, where $\xi_i \sim N(1,1)$. The sample size is $n = 1000$. By construction, both the direct treatment effect and spillover effect do not exist. The results of OLS regression, 2SLS, and WLS are summarized in the next table.
The table shows that in the presence of network endogeneity, the simple OLS regression incorrectly detects the spillover effect, even though the treatments are completely randomly assigned. In contrast, the 2SLS and WLS estimators correctly evaluate the spillover effect. The replication R code for this experiment is provided in Appendix (ref).
\paragraph{Efficiency comparison}
In Appendix (ref), under standard moment conditions, we show that the sample analog of the 2SLS estimand and that of WLS are both asymptotically normal in the following sense:
where $n = |\mathcal{I}|$, $\sigma^2_{\varepsilon}(d) \coloneqq \mathbb{E}[\varepsilon_i^2 \mid D_{j(i)} = d]$, $\varepsilon_i \coloneqq Y_i - \beta_0^{2sls} - \beta_d^{2sls} D_i - \beta_s^{2sls} S_i$, $\sigma^2_{\epsilon, 1}(d) \coloneqq \mathbb{E}[\epsilon_{1, i}^2 \mid D_{j(i)} = d, A_{j(i)} = 1]$, and $\epsilon_{1, i} \coloneqq Y_i - \beta_0^{wls} - \beta_d^{wls} D_i - \beta_s^{wls} D_{j(i)}$. This is a common result for the 2SLS estimator: the asymptotic variance is inversely proportional to the “square” of the compliance probability. On the other hand, the asymptotic variance of the WLS estimator is inversely proportional only to $p_1^A$. In particular, for the case of homoscedasticity such that $\sigma^2_{\varepsilon}(d) = \sigma^2_{\epsilon, 1}(d)$, the WLS estimator is more efficient than the 2SLS estimator exactly by a factor of $p_1^A$.
In this section, we generalize the above discussion to models in which interactions can occur among more than two individuals. For each $i \in \mathcal{I}$, let $\mathcal{P}_i \subseteq \mathcal{J}$ be a group of potential peers, such as family members, classmates, and local neighborhoods, depending on the context. The size of $\mathcal{P}_i$ is denoted by $n_i \coloneqq |\mathcal{P}_i|$ and may vary across $i$ and $n_i \ge 1$ for all $i \in \mathcal{I}$. We assume that $\mathcal{P}_i$ and $\mathcal{P}_{i'}$ are disjoint for any $i \neq i'$.
It is important to note that in the following analysis, $\mathcal{P}_i$ is treated as fixed for each $i$ (which may or may not be endogenous in reality). Thus, the subsequent discussion should be viewed as being conditional on $\mathcal{P}_i$'s. If we allow them to be random, then we will need to take into account the correlation between the potential outcomes and the composition and size of $\mathcal{P}_i$, which significantly complicates the analysis.\footnote{ Insofar as the analysis is based on existing networks, this level of network endogeneity may not be addressed without additional information, such as the availability of individual covariates to instrument $\mathcal{P}_i$. For example, assuming that $\mathcal{P}_i$ is exogenously constructed once conditioned on each school, then we can consider segmenting the data by school to control for the endogeneity. Otherwise, if we want to obtain “unconditional” spillover effects, we might need to resort to an experimental approach that intervenes in the network structure itself. }
Denoting the elements of $\mathcal{P}_i$ as $\mathcal{P}_i = \{1(i), \ldots, n_i(i)\}$, the link connections are characterized by $\bm{A}_{\mathcal{P}_i} = (A_{1(i)}, \ldots, A_{n_i(i)})^\top$, whose support is $\mathcal{A}_i \coloneqq \{0, 1\}^{n_i}$, where $A_{j(i)} = 1$ means that the treatment of $j$-th peer affects $i$'s outcome. The peer treatments are denoted by $\bm{D}_{\mathcal{P}_i} \coloneqq (D_{1(i)}, \ldots, D_{n_i(i)})^\top$. Then, the number of treated effective peers can be written as $R_i \coloneqq \bm{A}_{\mathcal{P}_i}^\top \bm{D}_{\mathcal{P}_i}$, which ranges over $\mathcal{R}_i \coloneqq \{0,1, \ldots, n_i\}$.
We assume that $R_i$ contains sufficient information on the treatment spillover effects in the sense that
where $Y_i(d,r)$ denotes the potential outcome when $D_i = d$ and $R_i = r$. This implicitly imposes anonymity and homogeneity in the treatment spillover mechanism, which is a standard assumption in the literature. It is worth noting that $R_i$ becomes an exogenous variable when conditioned on $i$'s degree: $\overline A_i \coloneqq \sum_{j = 1}^{n_i} A_{j(i)}$. That is, we have $\mathbb{E}[Y_i \mid R_i = r, D_i = d, \overline A_i = \overline a] = \mathbb{E}[Y_i(d,r) \mid \overline A_i = \overline a]$, implying that we can identify the average direct and spillover effects conditional on $\overline A_i = \overline a$ by a nonparametric regression of $Y_i$ on $(R_i, D_i, \overline A_i)$ (cf. leung2020treatment). However, since all these regressors are discrete, performing this nonparametric regression is impractical due to the curse of dimensionality. Instead, in practice, researchers often employ a linear social interaction model as follows:
where $M_i: \mathcal{R}_i \to \mathbb{R}$ is a known non-decreasing transformation of $R_i$. Common choices for $M_i$ would be $M_i(r) = r$, $M_i(r) = r/n_i$, and $M_i(r) = \bm{1}\{r > 0\}$.\footnote{ These three $M_i$ transformations are deterministic. Another empirically common choice is $M_i(r) = r/\overline A_i$. Since $\overline A_i$ is random and may be correlated with the potential outcomes, this case needs to be discussed separately from the others. Fortunately, it turns out that our causal interpretations for the 2SLS and WLS estimands remain valid for this $M_i$ too -- see the proofs of Theorems (ref) and (ref). } Note again that since $D_i$ and $M_i$ are independent, a cross-term between them is not included in the model.
To facilitate the analysis, we assume an experimental setup similar to that considered above.
We first characterize the selection bias in the OLS estimand caused by network endogeneity. The parameters of interest are as follows:
In contrast to the previous case, because the data may be non-identically distributed owing to the heterogeneity in $\mathcal{P}_i$, the target parameters essentially depend on the specific composition of $\mathcal{I}$ (but we suppress the dependence for notational simplicity). Note that the estimator does not use the information of $\mathcal{P}_i$, and thus it can be unobserved to us as long as $R_i$ is known. The same is true for the 2SLS and WLS.
Define
Here, $\tau_i^0(d, r)$ measures the spillover effect, using $Y_i(d, 0)$ as the baseline.
From Theorem (ref), $\beta_d^{ols}$ can be expressed as a weighted average of the conditional average direct effects. Thus, it does not lose causal interpretability even if the network endogeneity is ignored, similar to Proposition (ref). On the other hand, $\beta_s^{ols}$ includes the causal effect term $\sum_{(r, \vec{a}) \in \mathcal{R}_i \times \mathcal{A}_i} \pi_i(r, \vec{a}) \mathbb{E}[\overline \tau_i^0(r) \mid \bm{A}_{\mathcal{P}_i} = \vec{a}]$ and the selection bias term $\sum_{(r, \vec{a}) \in \mathcal{R}_i \times \mathcal{A}_i} \pi_i(r, \vec{a})\mathbb{E}[\overline Y_i(0) \mid \bm{A}_{\mathcal{P}_i} = \vec{a}]$. The bias comes from the correlation between $Y_i(d,0)$ and $\bm{A}_{\mathcal{P}_i}$. Moreover, even if the selection bias happens to be zero, the causal effect term is not purely causal because some weights $\{\pi_i(r, \vec{a})\}$ can be negative: indeed, we can easily see that $\frac{1}{n} \operatorname*{\sum_{\it i \in \mathcal{I}}} \sum_{(r, \vec{a}) \in \mathcal{R}_i \times \mathcal{A}_i} \pi_i(r, \vec{a}) = 0$. This is a tricky problem that does not arise in models with binary spillover exposure, which was originally pointed out by vazquez2022identification.
To perform the 2SLS estimation, a natural IV candidate for $R_i$ would be the summation of the peers' treatments $\sum_{j = 1}^{n_i} D_{j(i)}$. However, this choice of IV is not favorable in terms of obtaining a clear causal interpretation because $R_i$ is not directly a function of it. Another possibility is to use $\bm{D}_{\mathcal{P}_i}$ as separate IVs. With this approach, we could define various compliance types according to the connectivity of $A_{j(i)}$'s, such as $D_{1(i)}$-complier, $D_{2(i)}$-complier, etc, as in mogstad2021causal. Note however that even for a two-IV case, there are four different compliance patterns (see Table 2 in mogstad2021causal). More importantly, we allow each unit to have a heterogeneous number of peers. This means that for some units, some $D_{j(i)}$'s may be non-existent, which further complicates the interpretation and raises practical feasibility concerns. In addition, the peers are neither ordered nor labelled in general.
Considering these points, in the following, we focus on the case where a "single" element of $\bm{D}_{\mathcal{P}_i}$ is used as the IV for $R_i$. Without loss of generality, suppose that the IV is the first element of $\bm{D}_{\mathcal{P}_i}$, $D_{1(i)}$. Then, the potential treatment when $D_{1(i)} = d$ can be written as
The observed treatment is $R_i = D_{1(i)} R_i(1) + (1 - D_{1(i)}) R_i(0)$. By definition, the monotonicity condition $R_i(1) \ge R_i(0)$ holds trivially, and the inequality is strict if and only if $A_{1(i)} = 1$.\footnote{ As pointed out by a referee, as the size of potential peers $n_i$ increases, the contribution of each peer member decreases, which may lead to a weak IV problem. Therefore, when $n_i$ is relatively large, it might be better to reconsider the possibility of using multiple IVs. Also, the strength of $D_{1(i)}$'s as IVs is maximised when $A_{1(i)} =1$ for all $i$ (which is the source of the efficiency improvement in the WLS estimator). In this sense, since the choice of “first peer” is completely arbitrary, one should choose a person who is actually a friend of $i$ as $1(i)$. }
With this IV, the population 2SLS estimand $(\beta_0^{2sls}, \beta_d^{2sls}, \beta_s^{2sls})$ is defined as
where $\mathbb{L}(M_i \mid D_i, D_{1(i)}) \coloneqq \gamma_0 + \gamma_d D_i + \gamma_s D_{1(i)}$, and $(\gamma_0, \gamma_d, \gamma_s) = \operatorname*{argmin}_{a_0, a_d, a_s} \operatorname*{\sum_{\it i \in \mathcal{I}}} \mathbb{E}[(M_i - a_0 - a_d D_i - a_s D_{1(i)})^2]$. As in the pair-interaction model, as $D_i$ is independent of $M_i$, $\beta_d^{ols} = \beta_d^{2sls}$ holds. Thus, the same characterization as in Theorem (ref)(i) applies to $\beta_d^{2sls}$.
Let
The next theorem presents a causal interpretation of $\beta_s^{2sls}$.
From the same argument as before, it is possible to improve the efficiency of the 2SLS estimator by appropriately weighting the data. In this case, the compliers are those with a link connection with the first peer. Thus, the WLS estimand $(\beta_0^{wls}, \beta_d^{wls}, \beta_s^{wls})$ is defined as
where $\mathbb{L}_1(M_i \mid D_i, D_{1(i)}) \coloneqq \gamma_{0,1} + \gamma_{d,1} D_i + \gamma_{s,1} D_{1(i)}$, and $(\gamma_{0,1}, \gamma_{d,1}, \gamma_{s,1}) = \operatorname*{argmin}_{a_0, a_d, a_s} \operatorname*{\sum_{\it i \in \mathcal{I}}} \mathbb{E}[A_{1(i)}(M_i - a_0 - a_d D_i - a_s D_{1(i)})^2]$. Notice that if, for example, $1(i)$ represents $i$'s best friend and everyone has his/her best friend, the WLS is identical to the 2SLS.
As shown in Theorem (ref)(i), the causal interpretation of $\beta_d^{wls}$ is similar to that of $\beta_d^{ols}$ and $\beta_d^{2sls}$, except that it is conditioned on $A_{1(i)} = 1$. Result (ii) shows that the WLS estimand $\beta_s^{wls}$ and the 2SLS estimand $\beta_s^{2sls}$ have a completely identical characterization; thus, the same LATE interpretation as in Remark (ref) applies.
This subsection briefly discusses the asymptotic distributions of the 2SLS and WLS estimators. The sample analog of $\beta_s^{2sls}$ can be obtained by $\widehat \beta_s^{2sls} = \widetilde{\bm{D}}_{1,n}^\top \bm{Y}_n/ \widetilde{\bm{D}}_{1,n}^\top \bm{M}_n$, where $\bm{Y}_n = (Y_1, \ldots, Y_n)^\top$, $ \widetilde{\bm{D}}_{1,n} \coloneqq \bm{D}_{1,n} - \bm{D}_{c,n}(\bm{D}_{c,n}^\top \bm{D}_{c,n})^{-1}\bm{D}_{c,n}^\top \bm{D}_{1,n}$, $D_{c,i} = (1, D_i)^\top$, $\bm{D}_{c,n} = (D_{c,1}, \ldots, D_{c,n})^\top$, $\bm{D}_{1,n} = (D_{1(1)}, \ldots, D_{1(n)})^\top$, and $\bm{M}_n = (M_1, \ldots, M_n)^\top$. Furthermore, the population 2SLS residual can be written as $\varepsilon_i \coloneqq Y_i - \beta_0^{2sls} - \beta_d^{2sls} D_i - \beta_s^{2sls} M_i$. Note that the population residuals may have unknown heterogeneous means caused by heterogeneity in the distribution of potential outcomes and network structure. Hence, we introduce additional assumptions on the data to facilitate the derivation of the limiting distribution. Specifically, suppose that $\mathbb{E}[\varepsilon_i] = \mathbb{E}[D_i \varepsilon_i] = \mathbb{E}[D_{1(i)} \varepsilon_i] = 0$. Then,
where $\sigma^2_{\varepsilon,i}(d) \coloneqq \mathbb{E}[\varepsilon_i^2 \mid D_{1(i)} = d]$. Noting the equality $\mathbb{E}[M_i^1 - M_i^0] = \mathbb{E}[M_i^1 - M_i^0 \mid A_{1(i)} = 1] \Pr(A_{1(i)} = 1)$, we can observe that the asymptotic variance is inversely related to the square of the compliance probability.
We now turn to the asymptotic distribution of the WLS estimator: $\widehat \beta_s^{wls} = \widetilde{\bm{D}}_{1,n,A}^\top \bm{Y}_n/ \widetilde{\bm{D}}_{1,n,A}^\top \bm{M}_n$, where $\mathbb{I}_{n,A} \coloneqq \text{diag}(A_{1(1)}, \ldots, A_{1(n)})$, and $\widetilde{\bm{D}}_{1,n,A} \coloneqq \mathbb{I}_{n,A} \bm{D}_{1,n} - \mathbb{I}_{n,A} \bm{D}_{c,n}(\bm{D}_{c,n}^\top \mathbb{I}_{n,A} \bm{D}_{c,n})^{-1} \bm{D}_{c,n}^\top \mathbb{I}_{n,A} \bm{D}_{1,n}$. The population residual term is $\epsilon_i \coloneqq A_{1(i)} \epsilon_{1, i}$, where $\epsilon_{1, i} \coloneqq Y_i - \beta_0^{wls} - \beta_d^{wls} D_i - \beta_s^{wls} M_i$. Similarly, we assume $\mathbb{E}[\epsilon_i] = \mathbb{E}[D_i \epsilon_i] = \mathbb{E}[D_{1(i)} \epsilon_i] = 0$. Then, we have
where $\sigma^2_{\epsilon, 1, i}(d) \coloneqq \mathbb{E}[\epsilon_{1,i}^2 \mid D_{1(i)} = d, A_{1(i)} = 1]$. We can see that the asymptotic variances of the 2SLS and WLS estimators have the same denominator term, while the numerator of the WLS is weighted by the compliance probability $\Pr(A_{1(i)} = 1)$. Thus, if $\sigma^2_{\varepsilon,i}(d)$ and $\sigma^2_{\varepsilon,1,i}(d)$ take close values, the spillover-effect parameter can be estimated more efficiently using the WLS estimator.
For derivation of (ref) and (ref), see Appendix (ref).
The methods presented in the preceding sections may be difficult to apply directly in practice because of some restrictive assumptions regarding the data structure. In addition, we need to deal with the heterogeneity in the distribution of population residuals for statistical inference. Considering these points, we present several empirically tractable inference procedures by introducing additional constraints that may or may not be plausible depending on the data at hand.
To utilise the asymptotic normality results shown above, we require the population residuals to have zero mean uniformly. One possible way to enforce this property is to limit our attention to a subset in which both the potential outcomes and network structures are IID. The units in this subset must have potential peers of the same size. A similar approach is considered in vazquez2022identification. If we can find such a subset, say $\mathcal{I}'$, then the population 2SLS estimand can be defined as
where $\mathbb{L}(M_i \mid D_i, D_{1(i)}) \coloneqq \gamma_0 + \gamma_d D_i + \gamma_s D_{1(i)}$, and $(\gamma_0, \gamma_d, \gamma_s) = \operatorname*{argmin}_{a_0, a_d, a_s} \mathbb{E}[(M_i - a_0 - a_d D_i - a_s D_{1(i)})^2 \mid i \in \mathcal{I}']$. For simplicity, the dependence of the parameters on the choice of $\mathcal{I}'$ is suppressed. Then, for each $i \in \mathcal{I}'$, we can show that $\mathbb{E}[\varepsilon_i \mid i \in \mathcal{I}'] = \mathbb{E}[D_i \varepsilon_i \mid i \in \mathcal{I}'] = \mathbb{E}[D_{1(i)} \varepsilon_i \mid i \in \mathcal{I}'] = 0$ holds for the population residual $\varepsilon_i = Y_i - \beta_0^{2sls} - D_i \beta_d^{2sls} - M_i \beta_s^{2sls}$. Moreover, the asymptotic distribution of the 2SLS estimator is simplified as
where $n' \coloneqq |\mathcal{I}'|$, $\sigma^2_\varepsilon (d) \coloneqq \mathbb{E} [\varepsilon_i^2 \mid D_{1(i)} = d, i \in \mathcal{I}']$, and $q_d^A \coloneqq \Pr(A_{1(i)} = 1 \mid i \in \mathcal{I}')$. A similar result is obtained for the WLS estimator:
where $\sigma^2_{\epsilon,1} (d) \coloneqq \mathbb{E} [\epsilon_{1,i}^2 \mid D_{1(i)} = d, A_{1(i)} = 1, i \in \mathcal{I}']$. See Appendix (ref) for the derivation of these results.
Another possible approach, as in many empirical studies, is to assume a homogeneous treatment effect model:
Here, we implicitly require that the specification of the exposure mapping $M_i(\cdot)$ be correct (otherwise, the misspecification error may correlate with IV). There are many advantages of this approach: we do not need to account for the heterogeneity in the data structure, the treatment assignment does not have to follow a simple Bernoulli design, we can incorporate $\mathcal{J}$ in the estimation as well, the IV does not have to be binary, etc. One clear limitation is that it is often difficult to assume treatment homogeneity in practical applications.
If one simply wants to test for the presence of spillover effects within the data at hand, the Fisher's randomization test is a promising alternative. For example, consider the following null hypothesis:
If $\mathbb{H}_0$ is true, the IV has no impact on the outcome ($\beta_s^{2sls} = \beta_s^{wls} = 0$) so that we can impute the values of all $\{Y_i(D_i, R_i(d))\}_{d \in \{0,1\}}$ as $Y_i(D_i, R_i)$ ($=Y_i$). Thus, we can consider a conditional randomization test in which $\bm{D}_{1,n} = (D_{1(1)}, \ldots, D_{1(n)})^\top$ are randomized and everything else is fixed. Specifically, letting $T(\bm{D}_{1,n})$ be a predetermined test statistic, we can approximate the $p$-value of the statistic in the following manner: (Step 1) Compute $T(\bm{D}_{1,n})$; (Step 2) Draw $\bm{d}_{1,n}^{(b)}$ independently from the appropriate conditional distribution of $\bm{D}_{1,n}$ and compute $T(\bm{d}_{1,n}^{(b)})$ for $b = 1, \ldots , B$; and (Step 3) Compute $\widehat p_B \coloneqq B^{-1} \sum_{b=1}^B \bm{1}\{T(\bm{d}_{1,n}^{(b)}) \ge T(\bm{D}_{1,n})\}$.
Two obvious candidates for the test statistic $T(\bm{D}_{1,n})$ are $\widehat \beta_s^{2sls}$ and $\widehat \beta_s^{wls}$ possibly with some normalization. For other options, we can consider the intention-to-treat (ITT) statistic and ITT for compliers (ITTC):
Similar randomization tests can be found in the literature (e.g., rogowski2012estimating, forastiere2018posterior, kang2018inference). Note that to perform the randomization test, the experimental design does not have to be a Bernoulli design as long as it is known (although the above causal interpretations may not hold).
In this section, we revisit the data from paluck2016changing, who conducted a large social experiment on anti-conflict intervention programs in American middle schools. Half of these schools were randomly selected to host the programs. Within each selected school, a group of students ({\it seed-eligible students}) were selected and half of them ({\it seed students}) were randomly invited to join the meeting program. The students' social networks were measured by asking them to nominate up to 10 closest friends in their school.
In our analysis, the treatment variable $D_i$ indicates whether student $i$ was a seed student.\footnote{ Participation in the anti-conflict intervention program was not mandatory for seed students. Therefore, what we are estimating here is the effect of being nominated as a seed student. } We assume identity mapping for $M_i(\cdot)$; that is, $M_i = R_i$. For the IV of $M_i$, we use $D_{1(i)}$, where $1(i)$ denotes the closest seed-eligible friend to $i$. Outcome $Y$ is a dummy indicator for wearing the program wristband given as a reward for engaging in conflict-mitigating behaviors.
Under this setup, we consider the following two samples: (1) a sample of students in treatment schools who have at least one seed-eligible friend, and (2) a subsample of (1) constructed so that the friend networks of the units are non-overlapping. We perform OLS and 2SLS estimations on these samples. Because every unit has a seed-eligible friend (i.e., $A_{1(i)} = 1$ for all $i$), the 2SLS and WLS estimators are equivalent. In sample (1), we do not rule out the possibility that the units in this sample share the same friends. While this is not consistent with the theoretical analysis part, we examine this sample for robustness checks. Sample (2) is the one that conforms with the theoretical framework considered above. Note that the formation of potential peers (i.e., the set of general schoolmates) is treated as a given factor, and thus any school-level endogeneity may not be isolated from the estimates (see also the discussion in footnote 3).
The results are summarized in Table (ref). The spillover effect is significant in all four cases, irrespective of the estimation method; having one additional treated friend increases the probability of receiving the reward by roughly 5% or so. As a robustness check, we also perform the randomization test proposed in Subsection (ref) using sample (2). The $p$-values obtained from the 2SLS and ITT statistics are 0.010 and 0.016, respectively. From these results, it would be safe to conclude that there exist spillover effects of the program.