EconBase
← Back to paper

Causal Inference with Noncompliance and Unknown Interference

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

88,494 characters · 18 sections · 66 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Causal Inference with Noncompliance and Unknown Interference

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

\if01 {

center[center omitted — 89 chars of source]

} \fi

abstractWe consider a causal inference model in which individuals interact in a social network and they may not comply with the assigned treatments. In particular, we suppose that the form of network interference is unknown to researchers. To estimate meaningful causal parameters in this situation, we introduce a new concept of exposure mapping, which summarizes potentially complicated spillover effects into a fixed dimensional statistic of instrumental variables. We investigate identification conditions for the intention-to-treat effects and the average treatment effects for compliers, while explicitly considering the possibility of misspecification of exposure mapping. Based on our identification results, we develop nonparametric estimation procedures via inverse probability weighting. Their asymptotic properties, including consistency and asymptotic normality, are investigated using an approximate neighborhood interference framework. For an empirical illustration, we apply our method to experimental data on the anti-conflict intervention school program. The proposed methods are readily available with the companion { R} package \href{https://tkhdyanagi.github.io/latenetwork/}{{ latenetwork}}.

{\it Keywords:} exposure mapping, instrumental variables, local average treatment effect, network interference, spillover effects.

\spacingset{1.9}

Introduction

Estimating causal effects under cross-unit interference has become increasingly important in various fields. When individuals interact with each other, using the conventional potential outcome framework of rubin1980discussion based on the stable unit treatment value assumption (SUTVA) is inappropriate. To address the potential interference, there has been a rapidly growing number of studies that attempt to mitigate SUTVA by replacing it with some weaker restrictions.

A common approach to dealing with interference is to assume the existence of a low-dimensional exposure mapping which serves as a sufficient statistic of spillover effects in that others' treatments affect one's outcomes only through this function (e.g., hong2006evaluating; hudgens2008toward; manski2013identification; aronow2017estimating; li2019randomization; forastiere2021identification; li2022random). Some frequently-used forms of exposure mapping include, for example, simply extracting the neighbors' treatments from the entire treatment vector or calculating the proportion of treated neighbors. The exposure mapping is useful for summarizing potentially complicated spillover effects, but there is an inherent difficulty in how to choose the “right” functional form. Thus, some recent studies investigate under what conditions one can estimate meaningful causal parameters even under unknown interference (savje2021average; leung2022causal; savje2023causal).

Including the aforementioned studies, much of the research on causal inference with interference assumes the availability of experimental data where the individuals fully comply with their assigned treatments. However, this should be restrictive in many applications (e.g., miguel2004worms; dupas2014short; zelizer2019position). As a real example, consider the experiment on social norms and behaviors of adolescents conducted by paluck2016changing. They randomly selected students to participate in the anti-conflict intervention program where the participants were encouraged to take on leadership roles to reduce conflicts in school. The authors were interested in assessing the effectiveness of the intervention against one's own behavior, as well as whether the participants influence their peers through their friendship network. Unfortunately, a certain proportion of the selected students did not join the intervention program, which led the authors to compromise with an intention-to-treat (ITT) analysis.

Although the coexistence of spillovers and noncompliance should be prevalent in empirical applications, only a few studies have explicitly tackled this issue. sobel2006randomized shows that the conventional methods, such as the two-stage least squares estimator, may not admit causal interpretations when ignoring spillover effects. While there are studies that deal with both spillovers and noncompliance using an instrumental variable (IV) method by extending the local average treatment effect (LATE) framework of imbens1994identification (e.g., kang2016peer; kang2018spillover; imai2020causal; ditraglia2023identifying; vazquez2023causal), they rule out interactions on a large network and, more importantly, do not explicitly consider the misspecification of exposure mappings.

Taken these points together, it should be of primary importance to understand what causal parameters we can identify (if any) and how to perform statistical inference on them under the possibility of noncompliance and network interference of unknown form, which is the objective of this study. We consider a model in which individuals are connected through a single large network and they may self-select their treatment status. To account for the noncompliance issue and network interference, we employ the IV method and introduce a new concept of exposure mapping, which we call instrumental exposure mapping (IEM). The IEM is similar to the conventional exposure mapping in that it is a function summarizing the spillover effects into low-dimensional variables, but it differs in that it is a function of IVs.

We begin by considering the ITT analysis, wherein the estimands of interest are the average direct effect (ADE) and average indirect effect (AIE) of the IV on the outcome and those on the treatment choice. We show that these estimands have clear causal interpretation even with a misspecified IEM. Next, we focus on identifying the average direct and indirect effects for compliers who comply with their assigned treatments, which we call the local average direct effect (LADE) and local average indirect effect (LAIE), respectively. Under certain conditions, these LATE-type parameters capture the direct and indirect effects of the treatment receipt on the outcome for compilers, and thus, they should be more interpretable and policy-relevant than the simple ITT parameters. The technical difficulty in identifying these LATE-type parameters is that the standard identification argument in the LATE literature cannot be directly applied if no additional restrictions on the interaction structure are given. To address this problem, we extend the restricted interference assumption in imai2020causal to our situation. It is shown that the LADE and LAIE parameters are identifiable from Wald-type estimands under certain restricted interference assumptions. Importantly, these identification results imply that the ITT analysis disregarding the noncompliance may underestimate the direct and indirect treatment effects.

We propose nonparametric estimation procedures via inverse probability weighting. Our estimators are easy to implement, but their statistical properties are non-trivial because of unknown interference. We impose two key assumptions to show that our estimators are consistent and asymptotically normal. The first is the approximate neighborhood interference (ANI) assumption of leung2022causal, which is suitable for many empirical situations where spillover effects from distant units are weaker than those from close ones. The second assumption is that the network structure is “sparse” such that the number of link connections is sufficiently small for each unit.

\paragraph{Related literature}

Our identification results for the ADE parameters build particularly on imai2020causal, who study the identification of average causal effects for compliers in two-stage randomized experiments under noncompliance. A crucial assumption underlying their model is that interference is restricted within disjoint groups (i.e., partial interference). More importantly, they a priori assume a stratified interference mechanism in which spillovers are determined only through the number of treatment assignments within each cluster. In contrast, this paper addresses the network interference of unknown form leveraging a potentially misspecified IEM. See Remark (ref) for further comparison.

Our definitions and identification arguments for the AIE parameters extend hu2022average to the case of noncompliance. As in their paper, we define the AIE parameters based on the interference graph (cf. aronow2021spillover). Compared to hu2022average, we explicitly consider a potential misspecification of interference set -- the set of units who are affected by a focal unit. Specifically, we allow for the possibility that the effects originated from a focal unit spillover beyond its interference set.

Another closely related study is leung2022causal. He proposes an ANI model in experimental situations with perfect compliance and develops inferential methods for average treatment effects while explicitly allowing for the misspecification of exposure mapping. The major distinction between leung2022causal and ours is that not only the spillover effects of treatments on the outcome but also that of others' IVs on own treatment choice are considered in our study.

\paragraph{Paper organization}

Section (ref) presents our model setup. Sections (ref) and (ref) provide the identification and estimation results, respectively. Section (ref) reports the numerical results. The companion { R} package \href{https://tkhdyanagi.github.io/latenetwork/}{{ latenetwork}} is available from the authors' websites.

Model

Consider a finite population of $n \in \mathbb{N}$ units $N_n \coloneqq \{1, 2, \dots, n\}$. The units form an undirected network represented by the $n \times n$ symmetric adjacency matrix $\bm{A} = (A_{ij})_{i,j \in N_n}$, where $A_{ij} \in \{ 0, 1 \}$ indicates whether or not $i$ and $j$ are connected. We assume that there are no self-links so that $A_{ii} = 0$ for all $i \in N_n$. Denote the set of possible adjacency matrices of $n$ units as $\mathcal{A}_n$.

In a later section, we study asymptotic theory under the condition that the network size $n$ grows to infinity. This means that we consider a sequence of networks $\{ \bm{A}_m \}$ for $m = 1, 2, \dots$. The observed adjacency matrix $\bm{A}$ with no subscript is regarded as an $n$-th element of the sequence. Generally, networks $\bm{A}_{m_1}$ and $\bm{A}_{m_2}$ are completely unrelated, such that the members of the former and latter do not overlap. Meanwhile, it is possible to create a new network $\bm{A}_{m_3}$ ($m_3 = m_1 + m_2$) as a block diagonal matrix with the diagonal submatrices $\bm{A}_{m_1}$ and $\bm{A}_{m_2}$. Furthermore, the distributions of variables such as treatment assignments may be specific to each network in general; that is, they form a triangular array defined along with the network sequence. However, for notational simplicity, we suppress the dependence of variables on the network.

Let $Y_i \in \mathbb{R}$ be an observed outcome variable and $D_i \in \{ 0, 1 \}$ an indicator of the treatment receipt for $i$. In observational studies or randomized experiments with possible noncompliance, individuals may self-select their treatment status and the existing methods under perfect compliance may not be applicable. To address this problem, suppose that there is a binary IV, $Z_i \in \{ 0, 1 \}$. In an experimental setup, $Z_i$ is typically an indicator of initial treatment recommendation for $i$. Denote the $n$-dimensional vector of realized treatments as $\bm{D} = (D_i)_{i \in N_n}$, and similarly let $\bm{Z} = (Z_i)_{i \in N_n}$. We write the support of $\bm{D}$ and that of $\bm{Z}$ as $\mathcal{D}_n = \{ 0, 1 \}^n$ and $\mathcal{Z}_n = \{ 0, 1 \}^n$, respectively. For each $\bm{d} \in \mathcal{D}_n$ and $\bm{z} \in \mathcal{Z}_n$, we denote the potential outcome of unit $i$ when $\bm{D} = \bm{d}$ and $\bm{Z} = \bm{z}$ as $Y_i(\bm{d}, \bm{z})$. Similarly, the potential treatment status given $\bm{Z} = \bm{z}$ is written as $D_i(\bm{z})$. Let $\bm{D}(\bm{z}) = (D_i(\bm{z}))_{i \in N_n}$ be the $n$-dimensional vector of potential treatments. By construction, we have $Y_i = Y_i(\bm{D}, \bm{Z})$, $D_i = D_i(\bm{Z})$, and $\bm{D} = \bm{D}(\bm{Z})$. Hence, we can further write $y_i(\bm{z}) = Y_i(\bm{D}(\bm{z}), \bm{z})$ for some function $y_i: \mathcal{Z}_n \to \mathbb{R}$, and we have $Y_i = y_i(\bm{Z}) = \sum_{\bm{z} \in \mathcal{Z}_n} \bm{1}\{ \bm{Z} = \bm{z} \} y_i(\bm{z})$. Denoting $\bm{z}_{-i} = (z_k)_{k \neq i}$, we write the potential outcome of unit $i$ given $Z_j = z_j$ and $\bm{Z}_{-j} = \bm{z}_{-j}$ as $y_i(Z_j = z_j, \bm{Z}_{-j} = \bm{z}_{-j})$, where $i$ may differ from $j$. When $i = j$, we use both $y_i(z_i, \bm{z}_{-i})$ and $y_i(\bm{z})$ interchangeably depending on the situation. The same notation applies to the other functions of $\bm{z}$.

Because we can observe only one realization from $(y_i(\bm{z}), D_i(\bm{z}))_{\bm{z} \in \mathcal{Z}_n}$ for each unit, it is generally impossible to define identifiable causal estimands without introducing some restrictions. Here, we consider a pre-specified function $T: N_n \times \{0,1\}^{n-1} \times \mathcal{A}_n \to \mathcal{T}$, where $\mathcal{T} \subset \mathbb{R}^{\dim(T)}$ does not depend on $i$ and $n$, and $\dim(T)$ is a fixed positive integer.\footnote{ In the literature, forastiere2021identification consider an exposure mapping whose range may be heterogeneous across $i$ and $n$. Although our results would hold with minor modifications even in the presence of heterogeneity in exposure mappings, we consider this aspect beyond this study's scope because such a generalization substantially complicates asymptotic theory. Nevertheless, the common range assumption should not be too restrictive in practice, considering that in our framework, researchers can arbitrarily specify the form of IEM. } We call the function $T$ the {\it instrumental exposure mapping} (IEM) and its realization $T_i = T(i, \bm{Z}_{-i}, \bm{A})$ the {\it instrumental exposure}. When no confusion exists, we suppress the dependence of $T$ on $\bm{A}$, that is, $T(i, \bm{z}_{-i}) = T(i, \bm{z}_{-i}, \bm{A})$. Without loss of generality, we may suppose that the functional form of $T$ does not depend on $i$'s own $Z_i$.\footnote{ Even when the level of exposure changes depending on unit's own IV, we can consider the exposures when $Z_i = 1$ and $Z_i = 0$ separately, say $T_1(i, \bm{z}_{-i})$ and $T_0(i, \bm{z}_{-i})$, respectively, and re-define $T(i, \bm{z}_{-i}) \coloneqq (T_0(i, \bm{z}_{-i}), T_1(i, \bm{z}_{-i}))$. We thank the referee for suggesting this point. } We say that the IEM is {\it correctly specified} on $\bm{A}$ if

align[align omitted — 204 chars of source]

for all $i \in N_n$, $z_i \in \{ 0, 1 \}$, and $\bm{z}_{-i}, \bm{z}_{-i}' \in \{ 0, 1 \}^{n-1}$.\footnote{ Note that the definition in (ref) does not imply the uniqueness of correct IEM. For example, if having at least one treatment-eligible neighbor is a correct IEM for $i$ (i.e., $T_i^\text{max} = \max\{Z_j: A_{ij} = 1\}$), so is $i$'s neighborhood average $T_i^\text{ave} = \sum_{j \neq i} A_{ij} Z_j/\sum_{j \neq i} A_{ij}$, because $T_i^\text{ave} = 0$ and $T_i^\text{max} = 0$ are equivalent and $T_i^\text{ave} = t$ for any $t > 0$ is only a special case of $T_i^\text{max} = 1$. That is, $T_i^\text{ave}$ is a “finer” exposure than $T_i^\text{max}$. Based on such a hierarchical structure of different exposure mappings, hoshino2023randomization propose a randomization test for the specification of exposure mappings. If any spillover effects do not exist in the first place, any IEM is correct. }

If the IEM is correctly specified, it serves as a fixed dimensional sufficient statistic that summarizes potentially high-dimensional spillover effects. That is, the potential treatment status and the potential outcome of unit $i$ can be fully characterized by $i$'s own IV $Z_i$ and her instrumental exposure $T_i$, and there exist functions $\widetilde d_i: \{ 0, 1 \} \times \mathcal{T} \to \{ 0, 1 \}$ and $\widetilde y_i: \{ 0, 1 \} \times \mathcal{T} \to \mathbb{R}$ satisfying $\widetilde d_i(z_i, T(i, \bm{z}_{-i})) = D_i(z_i, \bm{z}_{-i})$ and $\widetilde y_i(z_i, T(i, \bm{z}_{-i})) = y_i(z_i, \bm{z}_{-i})$ for any $z_i \in \{ 0, 1 \}$ and $\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}$. Then, $\widetilde y_i(z, t)$ and $\widetilde d_i(z, t)$ represent the potential outcome and the potential treatment status, respectively, given $Z_i = z$ and $T_i = t$. In this way, a properly specified IEM alleviates the complexity of handling general spillover effects and greatly simplifies the estimation of causal parameters under interference. However, in reality, the user-specified IEM is generally incorrect. In this case, $\widetilde y_i(z, t)$ and $\widetilde d_i(z, t)$ are no longer well-defined.

Throughout the paper, following the recent literature on causal inference with interference, we focus on a design-based uncertainty framework where the randomness comes only from $\bm{Z}$. That is, we treat the potential outcomes, potential treatments, and adjacency matrix as non-stochastic components. The design-based approach is suitable for randomized experiments where researchers can design the random assignment mechanism for treatment eligibility or initial recommendations, as in dupas2014short and paluck2016changing. Even in observational studies, the design-based approach should be relevant when we can observe the entire population or most of the finite population (cf. abadie2020sampling). It should be noted that we can also view our framework as a random design on which everything other than $\bm{Z}$ is conditioned.

We assume that the IV can affect the outcome only through the treatment.

assumption[Exclusion restriction] $Y_i(\bm{d}, \bm{z}) = Y_i(\bm{d}, \bm{z}')$ for all $i \in N_n$, $\bm{d} \in \mathcal{D}_n$, and $\bm{z}, \bm{z}' \in \mathcal{Z}_n$.

Under Assumption (ref), we can reduce the potential outcome when $\bm{D} = \bm{d}$ to $Y_i(\bm{d}) = Y_i(\bm{d}, \bm{z})$, and we have $y_i(\bm{z}) = Y_i(\bm{D}(\bm{z}))$. Note that this assumption is not essentially necessary in terms of the ITT analysis, but it can greatly improve the causal interpretation of our parameters.

We provide two specific examples that can be effectively analyzed within our model.

exampleSuppose that the outcome is generated as $Y_i = \beta_{0i} + \beta_1 D_i + \beta_2 \cdot \bm{1}\{ \sum_{j \neq i} A_{ij} D_j > c \}$, where $\beta_{0i}$ is an idiosyncratic intercept term, $\beta_1$ and $\beta_2$ indicate the direct and spillover effects, respectively, and $c$ is a given threshold. Assume that the treatment status of each unit is determined only by her own IV: $D_i(z_i) = D_i(z_i, \bm{z}_{-i})$. Then, the potential outcome when $\bm{Z} = \bm{z}$ is $y_i(\bm{z}) = \beta_{0i} + \beta_1 D_i(z_i) + \beta_2 \cdot \bm{1} \{ \sum_{j \neq i} A_{ij} D_j(z_j) > c \}$. A correctly specified IEM is, for example, $T(i, \bm{Z}_{-i}) = \bm{1}\{ \sum_{j \neq i} A_{ij} D_j(Z_j) > c \}$ with $\mathcal{T} = \{ 0, 1 \}$. In the literature, this type of exposure mapping is used, for example, in hong2006evaluating and leung2022causal.
exampleSuppose that Assumption (ref) holds and that no interference exists in the outcome. For the treatment choice, consider the latent index model $D_i = \bm{1}\{ \gamma_{0i} + \gamma_{1i} Z_i + \gamma_{2i} \cdot \bm{1} \{ \sum_{j \neq i} A_{ij} Z_j > c \} > 0 \}$, where $\gamma_{0i}$ is the preference heterogeneity for the treatment, and $\gamma_{1i}$ and $\gamma_{2i}$ respectively capture the direct and spillover effects of the IV. In this situation, the potential outcome when $\bm{Z} = \bm{z}$ is given by $y_i(\bm{z}) = \beta_{0i} + \beta_{1i} D_i(\bm{z})$. As such, the model is a simple binary treatment model with potentially many IVs. It is straightforward to find that we can estimate a LATE-type parameter using the two-stage least squares method under a monotonicity condition between $D_i$ and $Z_i$ (e.g., $\gamma_{1i} \ge 0$ for all $i$), ignoring the spillover effect in the treatment choice model. If we set $T(i, \bm{Z}_{-i}) = \bm{1}\{ \sum_{j \neq i} A_{ij} Z_j > c \}$, this is clearly a correct IEM.

Identification

To begin with, we introduce the following assumption:

assumption[Independence] $\{Z_i\}_{i \in N_n}$ are mutually independent.

This assumption would be reasonable for many empirical situations. For example, the assumption is satisfied in a randomized experiment where the treatment eligibility is independently assigned to each unit with a given probability. Another example is an observational study in which the IV of each unit is determined independently from the other units.

Assumption (ref) will be used for both the identification analysis and asymptotic investigations. If the IVs are correlated with each other, this causes non-trivial difficulties in proceeding the subsequent analysis. For the same reason, several studies in the literature of network interference predominantly focus on Bernoulli experiments (e.g., hu2022average; li2022random). Investigating more general treatment assignment mechanisms is left for future research.

Intention-to-treat estimands and causal parameters of interest

We define the ITT estimands and the causal parameters of interest. To this end, consider a non-random sub-population $S_n \subseteq N_n$. Throughout the paper, we assume $S_n$ to be non-empty, and consider estimating causal parameters specific to this sub-population. For an example of $S_n$, let $S_n(\delta)$ be the set of units whose degrees are $\delta$: $S_n(\delta) = \{ i \in N_n: \sum_{j \neq i} A_{ij} = \delta \}$. In this case, we can examine whether the causal impacts vary across individuals with different centrality by comparing the estimates obtained with different $\delta$'s. See Remark (ref) for further discussion.

To define the ADE estimands, let $\mu_i^Y(z, t) \coloneqq \operatorname*{\mathbb{E}}[ Y_i \mid Z_i = z, T_i = t]$ for $z \in \{ 0, 1 \}$ and $t \in \mathcal{T}$. Here, the expectation is taken with respect to the distribution of $\bm{Z}$; that is, $\mu_i^Y(z, t) = \sum_{\bm{z} \in \mathcal{Z}_n} \Pr[\bm{Z} = \bm{z} \mid Z_i = z, T_i = t] y_i(\bm{z})$. Due to the heterogeneity in the potential outcomes, generally we cannot obtain a consistent estimator of $\mu_i^Y(z, t)$ in the design-based approach. Denoting $\bar \mu_{S_n}^Y(z, t) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \mu_i^Y(z, t)$, where $|S_n|$ is the cardinality of $S_n$, the ADE of the IV is defined by $\mathrm{ADEY}_{S_n}(t) \coloneqq \bar \mu_{S_n}^Y(1, t) - \bar \mu_{S_n}^Y(0, t)$. Similarly, we define $\mu_i^D(z, t) \coloneqq \operatorname*{\mathbb{E}}[ D_i \mid Z_i = z, T_i = t]$, $\bar \mu_{S_n}^D(z, t) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \mu_i^D(z, t)$, and $\mathrm{ADED}_{S_n}(t) \coloneqq \bar \mu_{S_n}^D(1, t) - \bar \mu_{S_n}^D(0, t)$.

Next, we turn to the AIE estimands. Let $\ell_{\bm{A}}(i, j)$ denote the path distance (defined on the whole population $N_n$) between units $i$ and $j$.\footnote{ The {\it path distance} between $i$ and $j$ is the length of the shortest sequence of neighboring edges connecting them. As convention, we set $\ell_{\bm{A}}(i, j) = \infty$ when no path exists between $i$ and $j$ in $\bm{A}$ and 0 if $i = j$. } Suppose that each $i$'s instrumental exposure $T_i$ depends only on unit $j$'s such that $1 \le \ell_{\bm{A}}(i, j) \le K$ with some constant $K \ge 1$ (see Assumption (ref)). Given this, we define the interference graph $\bm{E} = ( E_{ij} )_{i,j \in N_n}$, where $E_{ij} \coloneqq \bm{1}\{1 \le \ell_{\bm{A}}(i, j) \le K\}$. For each $i \in S_n$, let $\mathcal{E}_i \coloneqq \{ j \in N_n: E_{ij} = 1 \}$, which we call $i$'s interference set. Note that $i \notin \mathcal{E}_i$ because $\ell_{\bm{A}}(i, i) = 0$. Write $\mu_{ji}^Y(z) \coloneqq \operatorname*{\mathbb{E}}[ Y_j \mid Z_i = z ]$ and $\bar \mu_{S_n}^{Y}(z ; \mathcal{E}) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \sum_{j \in \mathcal{E}_i} \mu_{ji}^Y(z)$. Then, we define $\mathrm{AIEY}_{S_n} \coloneqq \bar \mu_{S_n}^Y(1; \mathcal{E}) - \bar \mu_{S_n}^Y(0; \mathcal{E})$. Similarly, we define $\mu_{ji}^D(z) \coloneqq \operatorname*{\mathbb{E}}[ D_j \mid Z_i = z ]$, $\bar \mu_{S_n}^D(z; \mathcal{E}) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \sum_{j \in \mathcal{E}_i} \mu_{ji}^D(z)$, and $\mathrm{AIED}_{S_n} \coloneqq \bar \mu_{S_n}^D(1; \mathcal{E}) - \bar \mu_{S_n}^D(0; \mathcal{E})$.

We will show later that these ITT estimands have clear causal interpretation as a certain weighted average of the effect of the IV on the outcome and as that on the treatment receipt. However, as will be highlighted, the ITT parameters may underestimate the effects of the treatment. To depart from the ITT analysis, we extend the notions of {\it compliers} and LATE to our setting. For each given $\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}$, let $\mathcal{C}_i(\bm{z}_{-i}) \coloneqq \mathbf{1}\{ D_i(1, \bm{z}_{-i}) = 1, D_i(0, \bm{z}_{-i}) = 0 \}$ indicate a complier who takes the treatment only when $Z_i = 1$ conditional on $\bm{Z}_{-i} = \bm{z}_{-i}$. Notably, the compliance status may change with the others' IV values. Denote the realized compliance status as $\mathcal{C}_i = \mathcal{C}_i(\bm{Z}_{-i})$. Letting $\pi_i(\bm{z}_{-i}, t) \coloneqq \Pr[ \bm{Z}_{-i} = \bm{z}_{-i} \mid T_i = t ]$, the expected compliance status conditional on $T_i = t$ is $\operatorname*{\mathbb{E}}[ \mathcal{C}_i \mid T_i = t] = \sum_{\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}} \mathcal{C}_i(\bm{z}_{-i}) \pi_i(\bm{z}_{-i}, t)$. Then, the LADE is defined by the weighted average of $y_i(1, \bm{z}_{-i}) - y_i(0, \bm{z}_{-i})$ over the compliers:

align*[align* omitted — 294 chars of source]

Similarly, noting that $\operatorname*{\mathbb{E}}[ \mathcal{C}_i ] = \sum_{\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}} \mathcal{C}_i(\bm{z}_{-i}) \pi_i(\bm{z}_{-i})$ with $\pi_i(\bm{z}_{-i}) \coloneqq \Pr[ \bm{Z}_{-i} = \bm{z}_{-i} ]$, the LAIE is defined by the weighted average of $\sum_{j \in \mathcal{E}_i} \{ y_j(Z_i = 1, \bm{Z}_{-i} = \bm{z}_{-i}) - y_j(Z_i = 0, \bm{Z}_{-i} = \bm{z}_{-i}) \}$ over the compliers:

align*[align* omitted — 337 chars of source]
remark[ADE parameters with a constant IEM] The identification and estimation results presented below are still valid even when $T(i, \bm{z}_{-i})$ is constant for all $i$ and $\bm{z}_{-i}$. In this case, $\mathrm{ADEY}_{S_n}(t)$ reduces to $\mathrm{ADEY}_{S_n} \coloneqq \bar \mu_{S_n}^Y(1) - \bar \mu_{S_n}^Y(0)$, where $\bar \mu_{S_n}^{Y}(z) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \mu_i^Y(z)$ with $\mu_i^Y(z) \coloneqq \operatorname*{\mathbb{E}}[ Y_i \mid Z_i = z ]$. Analogously, $\mathrm{ADED}_{S_n}(t) = \mathrm{ADED}_{S_n} \coloneqq \bar \mu_{S_n}^D(1) - \bar \mu_{S_n}^D(0)$, where $\bar \mu_{S_n}^{D}(z) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \mu_i^D(z)$, and $\mu_i^D(z) \coloneqq \operatorname*{\mathbb{E}}[ D_i \mid Z_i = z ]$. Further, $\mathrm{LADE}_{S_n}(t)$ reduces to \begin{align*} \mathrm{LADE}_{S_n} \coloneqq \sum_{i \in S_n} \sum_{\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}} \{ y_i(1, \bm{z}_{-i}) - y_i(0, \bm{z}_{-i}) \} \frac{ \mathcal{C}_i(\bm{z}_{-i}) \pi_i(\bm{z}_{-i}) }{ \sum_{i \in S_n} \operatorname*{\mathbb{E}}[ \mathcal{C}_i ] }. \end{align*} These parameters are natural extensions of those in hu2022average to the case of noncompliance. Compared to the ADE parameters discussed above, the parameters defined here do not depend on the instrumental exposure explicitly. Thus, the latter parameters would be easier to interpret than the former with incorrect IEM. Meanwhile, if a correct IEM is used, the former parameters are more informative for understanding the treatment effect heterogeneity for the exposure level.

Average direct effects

Intention-to-treat analysis

The following proposition presents the causal interpretation of $\mathrm{ADEY}_{S_n}(t)$ and $\mathrm{ADED}_{S_n}(t)$.

propositionUnder Assumption (ref), \begin{align*} \mathrm{ADEY}_{S_n}(t) & = \frac{1}{|S_n|} \sum_{i \in S_n} \sum_{\bm{z}_{-i} \in \{0, 1\}^{n-1}} \{ y_i(1, \bm{z}_{-i}) - y_i(0, \bm{z}_{-i}) \} \pi_i(\bm{z}_{-i}, t), \\ \mathrm{ADED}_{S_n}(t) & = \frac{1}{|S_n|} \sum_{i \in S_n} \sum_{\bm{z}_{-i} \in \{0, 1\}^{n-1}} \{ D_i(1, \bm{z}_{-i}) - D_i(0, \bm{z}_{-i}) \} \pi_i(\bm{z}_{-i}, t). \end{align*}

Proposition (ref) shows that $\mathrm{ADEY}_{S_n}(t)$ and $\mathrm{ADED}_{S_n}(t)$ have clear causal interpretation as the weighted average of $y_i(1, \bm{z}_{-i}) - y_i(0, \bm{z}_{-i})$ and as that of $D_i(1, \bm{z}_{-i}) - D_i(0, \bm{z}_{-i})$, respectively. If we additionally impose Assumption (ref), the result for $\mathrm{ADEY}_{S_n}(t)$ can be further well interpreted. To see this, let $\bm{D}_{-i}(Z_i = z_i, \bm{Z}_{-i} = \bm{z}_{-i}) = (D_k(Z_i = z_i, \bm{Z}_{-i} = \bm{z}_{-i}))_{k \neq i}$. By the definition of $y_i(z_i, \bm{z}_{-i})$ and Assumption (ref), we can observe that

align*[align* omitted — 526 chars of source]

That is, $y_i(1, \bm{z}_{-i}) - y_i(0, \bm{z}_{-i})$ comprises of the direct effect of changing $i$'s own treatment status from $D_i(0, \bm{z}_{-i})$ to $D_i(1, \bm{z}_{-i})$ and the spillover effect by changing the others' treatments from $\bm{D}_{-i}(Z_i = 0, \bm{Z}_{-i} = \bm{z}_{-i})$ to $\bm{D}_{-i}(Z_i = 1, \bm{Z}_{-i} = \bm{z}_{-i})$. Hence, Proposition (ref) can be read as that $\mathrm{ADEY}_{S_n}(t)$ consists of the sum of the average direct effect from own IV and the average spillover effect caused by changing the unit's own IV.

Proposition (ref) is also useful for interpreting the ADEs obtained with different IEMs. For example, suppose that two IEMs $T$ and $T'$ generate the same ADEY value at $t$ and at $t'$, respectively:

align*[align* omitted — 291 chars of source]

Thus, if this equality holds, it is a strong indication that $y_i(1, \bm{z}_{-i}) - y_i(0, \bm{z}_{-i})$ is homogeneous with respect to $\bm{z}_{-i}$ for all individuals. Notably, if it is indeed homogeneous, the above equality must hold for any combination of IEMs, which is testable from the data.

Local average direct effect

The following two conditions are analogous to the IV relevance condition and the monotonicity condition for the standard LATE estimation without interference.

assumption[Relevance 1] $|S_n|^{-1} \sum_{i \in S_n} \operatorname*{\mathbb{E}}[ \mathcal{C}_i \mid T_i = t ] \ge c$ for a constant $c > 0$.
assumption[Monotonicity 1] $D_i(1, \bm{z}_{-i}) \ge D_i(0, \bm{z}_{-i})$ for all $i \in S_n$ and $\bm{z}_{-i} \in \{0, 1\}^{n - 1}$ such that $\pi_i(\bm{z}_{-i}, t) > 0$.

Assumption (ref) states that there is a non-negligible proportion of units among those with $T_i = t$ whose treatment status is positively affected by the IV. The assumption is necessary to well-define the LADE. Assumption (ref) requires that there do not exist {\it defiers}, whose treatment status is negatively affected by the IV. This assumption limits the heterogeneity in treatment choice in that the response to the IV must be uniform. For instance, the treatment choice equation in Example (ref) satisfies the condition if $\gamma_{1i} \ge 0$ for all $i$. Note that these assumptions do not have to be fulfilled uniformly in $t \in \mathcal{T}$ as the LADE of interest is conditioned on $T_i = t$ at a given $t$.

Under Assumption (ref), conditional on $\bm{Z}_{-i}$, each individual $i$ can be classified into one of the following three latent types: {\it compliers}, {\it always takers} (those who always take the treatment), and {\it never takers} (those who never take the treatment). More precisely, we define the indicator for always takers as $\mathcal{AT}_i = \mathcal{AT}_i(\bm{Z}_{-i}) \coloneqq \bm{1}\{ D_i(1, \bm{Z}_{-i}) = D_i(0, \bm{Z}_{-i}) = 1 \}$. Similarly, the indicator for never takers is $\mathcal{NT}_i = \mathcal{NT}_i(\bm{Z}_{-i}) \coloneqq \bm{1}\{ D_i(1, \bm{Z}_{-i}) = D_i(0, \bm{Z}_{-i}) = 0 \}$. Under Assumption (ref), only one of $\mathcal{C}_i$, $\mathcal{AT}_i$, and $\mathcal{NT}_i$ equals one for each $i \in S_n$.\footnote{ While $\mathcal{C}_i$, $\mathcal{AT}_i$, and $\mathcal{NT}_i$ are determined from $D_i(1, \bm{Z}_{-i})$ and $D_i(0, \bm{Z}_{-i})$ conditional on $\bm{Z}_{-i}$, it is possible to more finely classify the units by incorporating other potential treatment responses $\{ D_i(\bm{z}) \}_{\bm{z} \in \mathcal{Z}_n}$. As such an example, vazquez2023causal classifies the units into always takers, never takers, compliers, social-interaction compliers, and group compliers. See Appendix F for further discussion. }

Unlike conventional identification results without interference, the set of the exclusion restriction, relevance condition, and monotonicity condition does not suffice to identify the LADE. This is because we need to account for two potential interference channels at the same time: one is the spillover effect of the IV on the treatment receipt, and the other is the spillover effect of the treatment on the outcome. As in the conventional method, we use the variation in the instrumental value to identify the LADE, but in the present situation, the effect of shifting the IV can be amplified in two steps by the two different spillovers. Therefore, to facilitate the identification of the LADE, some additional restriction on the interference structure is needed. In this study, similar to imai2020causal, we require the potential outcome $y_i(z_i, \bm{z}_{-i})$ of noncompliers to be insensitive to their own instrumental value $z_i$.

assumption[Restricted interference 1] For all $i \in S_n$ and $\bm{z}_{-i} \in \{ 0, 1 \}^{n - 1}$ such that $\pi_i(\bm{z}_{-i}, t) > 0$, $y_i(1, \bm{z}_{-i}) = y_i(0, \bm{z}_{-i})$ holds whenever $D_i(1, \bm{z}_{-i}) = D_i(0, \bm{z}_{-i})$.

Here, we provide three empirically relevant sufficient conditions for this assumption. The first condition is no spillovers between the IV and treatment choice:

align[align omitted — 181 chars of source]

This corresponds to the {\it personalized encouragement} assumption of kang2016peer, which states that an incentive to take the treatment must be personalized to everyone. Under this condition, we can define the potential treatment status as $D_i(z_i) = D_i(z_i, \bm{z}_{-i})$. Then, the potential outcome satisfies $y_i(z_i, \bm{z}_{-i}) = Y_i(D_i(z_i), (D_k(z_k))_{k \neq i})$, implying Assumption (ref).

The second situation in which Assumption (ref) holds is when there is no treatment spillover effect on the outcome; that is,

align[align omitted — 181 chars of source]

Then, we may write the potential outcome given $D_i = d_i$ as $Y_i(d_i)$. It is easy to see that (ref) implies Assumption (ref).

The third sufficient condition for Assumption (ref) is that the IV of any noncomplier does not affect the treatment status of all other units; specifically, for any $\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}$,

align[align omitted — 192 chars of source]

If this condition holds, the potential outcome of unit $i$ with $D_i(1, \bm{z}_{-i}) = D_i(0, \bm{z}_{-i})$ satisfies $y_i(1, \bm{z}_{-i}) = y_i(0, \bm{z}_{-i})$, which implies Assumption (ref).

Although it is difficult to directly verify conditions (ref) -- (ref) from data alone, they have different testable implications (see Appendix F). There would also be cases where the experimental design suggests which is more likely to hold than the others. For example, consider the empirical analysis of the anti-conflict intervention program, where $Y_i$ is an outcome variable relating to anti-conflict norms and behaviors, $D_i$ is an indicator for participating in the program, and $Z_i$ indicates whether receiving an invitation for the program. In this experiment, the participation in the program was not compulsory and the students without an invitation were not able to attend. Thus, there are only compliers and never takers in this empirical example, and joining in the intervention program means that the student is a complier. It is plausible to imagine that the never takers were unlikely to affect the participation of others, as the never takers would be those who were not interested in the program. Thus, the sufficient condition in (ref) would be met here.

The causal interpretation of the LADE can be different depending on which sufficient condition the researcher considers for Assumption (ref) to hold. To see this, when $i$ is a complier, we have

align*[align* omitted — 432 chars of source]

Here, under (ref) or (ref), the second line vanishes. Therefore, the LADE purely captures the direct treatment effect for the compliers. On the contrary, (ref) admits that the IV of a complier may affect the treatment status of others. In this case, the second line is generally nonzero, and the LADE contains the average spillover effect for the compliers as well.

The following theorem shows that the LADE is identifiable from a Wald-type estimand.

theoremUnder Assumptions (ref) -- (ref), $\mathrm{LADE}_{S_n}(t) = \mathrm{ADEY}_{S_n}(t) / \mathrm{ADED}_{S_n}(t)$.
remark[Failure of Assumption (ref)] The Wald-type estimand does not generally admit valid causal interpretation without Assumption (ref) due to potential spillovers through noncompliers (see Appendix A.3 for details). The assumption fails if a noncomplier's IV affects the treatment choices of other units, which further influence on her own outcome. In our empirical setting, this occurs when students who are not interested in the anti-conflict campaign discourage other students from participating in the program, and this further reinforces their negative attitudes toward the school climate.
remark[imai2020causal] The IV relevance condition, monotonicity condition, and restricted interference in Assumptions (ref) -- (ref) are essentially the same as those in imai2020causal. Consequently, Theorem (ref) follows from nearly the same argument as in Theorem 2(1) of imai2020causal. However, the two papers have two notable differences in terms of the identification of ADEs. First, we impose the mutual independence of IVs in Assumption (ref), while imai2020causal focus on a two-stage randomized experiment. Second and more importantly, the two papers consider different network structures and asymptotic frameworks to establish the estimability for the target parameters. Our paper achieves this by showing that $\mathrm{ADEY}_{S_n}(t)$ and $\mathrm{ADED}_{S_n}(t)$ can be consistently estimated under the ANI framework on a large network and certain restrictions on the denseness of the network (see Section (ref)). By contrast, imai2020causal prove the consistency of their estimators in the setting of partial interference assuming that both the cluster size and number of clusters grow to infinity.

Average indirect effects

Intention-to-treat analysis

The following proposition presents the causal interpretation of $\mathrm{AIEY}_{S_n}$ and $\mathrm{AIED}_{S_n}$.

propositionUnder Assumption (ref), \begin{align*} \mathrm{AIEY}_{S_n} & = \frac{1}{|S_n|} \sum_{i \in S_n} \sum_{\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}} \sum_{j \in \mathcal{E}_i} \{ y_j(Z_i = 1, \bm{Z}_{-i} = \bm{z}_{-i}) - y_j(Z_i = 0, \bm{Z}_{-i} = \bm{z}_{-i}) \} \pi_i(\bm{z}_{-i}), \\ \mathrm{AIED}_{S_n} & = \frac{1}{|S_n|} \sum_{i \in S_n} \sum_{\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}} \sum_{j \in \mathcal{E}_i} \{ D_j(Z_i = 1, \bm{Z}_{-i} = \bm{z}_{-i}) - D_j(Z_i = 0, \bm{Z}_{-i} = \bm{z}_{-i}) \} \pi_i(\bm{z}_{-i}). \end{align*}

In view of Proposition (ref), these estimands measure the weighted average of the effect of $i$'s IV on the sum of $j$'s ($j \in \mathcal{E}_i$) outcomes and that on the sum of $j$'s treatments with the weight equal to $\pi_i(\bm{z}_{-i})$. Furthermore, under Assumption (ref), we can see that $\mathrm{AIEY}_{S_n}$ captures both the effect of changing $i$'s treatment through $i$'s IV on the sum of $j$'s outcomes and that of changing others' treatments (including $j$'s treatment) . This result immediately follows from the same arguments as in the previous subsection, once we notice that

align[align omitted — 160 chars of source]

While these ITT estimands help infer the spillover effects even with a misspecified IEM, whether they can correctly capture the spillover effects from distant units depends on the correctness of the selected IEM. To see this, let $T_i^*$ denote the true instrumental exposure defined by $j$'s such that $1 \le \ell_{\bm{A}}(i, j) \le K^*$ with some $K^*$. We write $i$'s interference set based on $T_i^*$ as $\mathcal{E}_i^* \coloneqq \{ j \in N_n : E_{ij}^* = 1 \}$. If $K < K^*$ so that $\mathcal{E}_i \subset \mathcal{E}_i^*$ for some $i \in S_n$, there are some spillover effects that are not accounted for by the AIEY under $K$. Meanwhile, if $K \ge K^*$ so that $\mathcal{E}_i \supseteq \mathcal{E}_i^*$ for all $i \in S_n$, the AIEY defined by $K$ coincides with the one defined by $K^*$. This result immediately follows from Proposition (ref) and the fact that $y_j(Z_i = 1, \bm{Z}_{-i} = \bm{z}_{-i}) = y_j(Z_i = 0, \bm{Z}_{-i} = \bm{z}_{-i})$ for all $j \notin \mathcal{E}_i^*$. Considering this, one might expect that choosing a large $K$ is practically desirable. However, a large $K$ implies a strong network dependence between distant units, and might deteriorate the estimation precision (see Lemma B.7 for a related result). Studying how to balance such a trade-off is beyond the scope of this paper.

Local average indirect effect

We introduce the following three assumptions, which are parallel to Assumptions (ref), (ref), and (ref), respectively.

assumption[Relevance 2] $|S_n|^{-1} \sum_{i \in S_n} \operatorname*{\mathbb{E}}[ \mathcal{C}_i ] \ge c$ for a constant $c > 0$.
assumption[Monotonicity 2] $D_i(1, \bm{z}_{-i}) \ge D_i(0, \bm{z}_{-i})$ for all $i \in S_n$ and $\bm{z}_{-i} \in \{0, 1\}^{n - 1}$ such that $\pi_i(\bm{z}_{-i}) > 0$.
assumption[Restricted interference 2] For all $i \in S_n$, $j \in \mathcal{E}_i$, and $\bm{z}_{-i} \in \{ 0, 1 \}^{n-1}$ such that $\pi_i(\bm{z}_{-i}) > 0$, $y_j(Z_i = 1, \bm{Z}_{-i} = \bm{z}_{-i}) = y_j(Z_i = 0, \bm{Z}_{-i} = \bm{z}_{-i})$ holds whenever $D_i(1, \bm{z}_{-i}) = D_i(0, \bm{z}_{-i})$.

As in Assumption (ref), Assumption (ref) restricts the interference structure, but they are different in that the latter limits (not $i$'s own but) $j$'s outcome value when $i$ is a noncomplier. In the same manner as in the previous subsection, we can see that (ref) or (ref) fulfills Assumption (ref). Meanwhile, (ref) is not sufficient for Assumption (ref). The causal interpretation of the LAIE varies with which sufficient condition holds for Assumption (ref). Specifically, for a complier $i$, (ref) leads to the following decomposition:

align*[align* omitted — 561 chars of source]

Under (ref), the second line vanishes, and the LAIE captures the average effect of complier $i$'s treatment on the sum of $j$'s outcomes. By contrast, the second line remains in the case of (ref), and the LAIE recovers the sum of the average effect of complier $i$'s treatment on the others' outcomes and that of others' treatments caused by changing $i$'s IV.

The next theorem shows that the LAIE is identifiable from a Wald-type estimand.

theoremUnder Assumptions (ref) and (ref) -- (ref), $\mathrm{LAIE}_{S_n} = \mathrm{AIEY}_{S_n} / \mathrm{ADED}_{S_n}$.
remark[Another Wald-type estimand] We can naturally think of another Wald-type estimand $\mathrm{AIEY}_{S_n} / \mathrm{AIED}_{S_n}$. A causal interpretation of this parameter can be derived but with a set of assumptions that appears somewhat restrictive in practice. See Appendix F for further discussion.

Estimation and Asymptotic Theory

Estimators

We consider the following data generating process (DGP):

assumption[DGP] (i) Assumption (ref) holds. (ii) $\{ (Z_i, T_i) \}_{i \in S_n}$ (resp. $\{ Z_i \}_{i \in S_n}$) are identically distributed across $i \in S_n$ for estimating the parameters conditioned on $(Z_i, T_i) = (z, t)$ (resp. $Z_i = z$). (iii) $\mathcal{T}$ is a finite subset of $\mathbb{R}^{\dim(T)}$. (iv) $|S_n| \to \infty$.

We reintroduce Assumption (ref) here for the sake of self-containedness of this section. The identical distribution of $\{ (Z_i, T_i) \}_{i \in S_n}$ in Assumption (ref)(ii) can be justified by appropriately choosing IEM $T$ and sub-population $S_n$. For example, this assumption holds when $T_i = \bm{1}\{ \sum_{j \neq i} A_{ij} Z_j > c \}$ and $S_n = \{ i \in N_n: \sum_{j \neq i} A_{ij} = \delta \}$ for some $c$ and $\delta$, provided that $\{ Z_i \}_{i \in N_n}$ are IID. We require this assumption to construct a consistent estimator of $\Pr[Z_i = z, T_i = t]$ for those $i \in S_n$ and prove a weak dependence property of some variables in the estimation of the ADE parameters. Similarly, the identical distribution of $\{ Z_i \}_{i \in S_n}$ in Assumption (ref)(ii) is used to consistently estimate $\Pr[Z_i = z]$ for those $i \in S_n$ and to analyze the dependency structure in the AIE estimation. Assumption (ref)(iii) is for simplicity.

Under Assumption (ref), for all $i \in S_n$, we can write $p_{S_n}(z, t) \coloneqq \Pr[Z_i = z, T_i = t]$ and $p_{S_n}(z) \coloneqq \Pr[Z_i = z]$. We estimate these by $\widehat p_{S_n}(z, t) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \bm{1}\{ Z_i = z, T_i = t \}$ and $\widehat p_{S_n}(z) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \bm{1}\{ Z_i = z \}$.\footnote{ If one knows the experimental design completely, $p_{S_n}(z, t)$ and $p_{S_n}(z)$ can be exactly computed without estimation, as in leung2022causal. In this case, we can achieve unbiased estimation of $\bar \mu_{S_n}^Y(z, t)$ and $\bar \mu_{S_n}^Y(z; \mathcal{E})$. In this study, for generality, we investigate the case where they need to be estimated. } Then, $\bar \mu_{S_n}^Y(z, t)$ and $\bar \mu_{S_n}^Y(z; \mathcal{E})$ can be estimated by

align*[align* omitted — 338 chars of source]

With these estimators, we define $\widehat{\mathrm{ADEY}}_{S_n}(t) \coloneqq \widehat \mu_{S_n}^Y(1, t) - \widehat \mu_{S_n}^Y(0, t)$ and $\widehat{\mathrm{AIEY}}_{S_n} \coloneqq \widehat \mu_{S_n}^Y(1; \mathcal{E}) - \widehat \mu_{S_n}^Y(0; \mathcal{E})$. Similarly, we can obtain the estimators for the treatment receipt, namely, $\widehat \mu_{S_n}^D(z, t)$, $\widehat \mu_{S_n}^D(z)$, $\widehat \mu_{S_n}^{D}(z; \mathcal{E})$, $\widehat{\mathrm{ADED}}_{S_n}(t)$, $\widehat{\mathrm{ADED}}_{S_n}$, and $\widehat{\mathrm{AIED}}_{S_n}$. Finally, we have $\widehat{\mathrm{LADE}}_{S_n}(t) \coloneqq \widehat{\mathrm{ADEY}}_{S_n}(t) / \widehat{\mathrm{ADED}}_{S_n}(t)$ and $\widehat{\mathrm{LAIE}}_{S_n} \coloneqq \widehat{\mathrm{AIEY}}_{S_n} / \widehat{\mathrm{ADED}}_{S_n}$.

remark[Choice of the subpopulation] To estimate the parameters conditioned on $(Z_i, T_i) = (z, t)$, we need to construct $S_n$ appropriately to satisfy Assumption (ref)(ii). When all units in $N_n$ have the same degree (e.g., a ring network where every unit connects only to the two adjacent units), we can set $S_n = N_n$ (if $\{ Z_i \}_{i \in N_n}$ are IID). For a more general network, the assumption requires us to select $S_n$ as a proper subset of $N_n$ to ensure the homogeneity of degrees over $S_n$.

Asymptotic properties

Average direct effects

We impose the following conditions.

assumption[Bounded outcome] There exists a constant $\bar y$ such that $|y_i(\bm{z})| \le \bar y < \infty$ for all $i \in S_n$ and $\bm{z} \in \mathcal{Z}_n$.
assumption[Overlap] There exist constants $\underline{p}, \bar p \in (0, 1)$ such that $p_{S_n}(z, t) \in [\underline{p}, \bar{p}]$ and $p_{S_n}(z) \in [\underline{p}, \bar{p}]$ for all $z \in \{ 0, 1 \}$ and a given $t \in \mathcal{T}$.

Assumption (ref) depends on the specification of IEM $T$, choice of sub-population $S_n$, distribution of $\bm{Z}$, and structure of network $\bm{A}$. For example, the assumption is fulfilled for each $t \in \{ 0, 1 \}$ if $T_i = \bm{1} \{ \sum_{j \neq i} A_{ij} Z_j > 0 \}$, $S_n = \{ i \in N_n: \sum_{j \neq i} A_{ij} = \delta \}$ for some $\delta \ge 1$ is non-empty, and $Z_i \stackrel{\mathrm{IID}}{\sim} \mathrm{Bernoulli}(q)$ for all $i \in N_n$ and some fixed $q \in (0, 1)$.

For a non-negative integer $s \ge 0$, let $N_{\bm{A}}(i, s) \coloneqq \{ j \in N_n: \ell_{\bm{A}}(i, j) \le s \}$ be the set of units within $s$ distance from unit $i$; namely, unit $i$'s $s$-neighborhood. Note that $i \in N_{\bm{A}}(i, s)$ for all $s \ge 0$. We write the sub-vector of $\bm{z} \in \mathcal{Z}_n$ restricted on $N_{\bm{A}}(i, s)$ as $\bm{z}_{N_{\bm{A}}(i, s)} \coloneqq (z_j)_{j \in N_{\bm{A}}(i, s)}$. Similarly, let $\bm{A}_{N_{\bm{A}}(i, s)} = (A_{kl})_{k,l \in N_{\bm{A}}(i, s)}$ denote the sub-matrix of $\bm{A}$ restricted on $N_{\bm{A}}(i, s)$.

assumption[IEM] There exists a known positive integer $K \in \mathbb{N}$ such that, for all $i \in S_n$, $\bm{A}, \bm{A}' \in \mathcal{A}_n$, and $\bm{z}, \bm{z}' \in \mathcal{Z}_n$, \begin{align*} $N_{\bm{A}}(i, K) = N_{\bm{A}'}(i, K)$, $\bm{A}_{N_{\bm{A}}(i, K)} = \bm{A}'_{N_{\bm{A}'}(i, K)}$, and $\bm{z}_{N_{\bm{A}}(i, K)} = \bm{z}'_{N_{\bm{A}'}(i, K)}$ \; \Longrightarrow \; T(i, \bm{z}, \bm{A}) = T(i, \bm{z}', \bm{A}'). \end{align*}

The assumption states that the instrumental exposure of each unit depends only on the unit's own $K$-neighborhood. This would be a mild requirement that most practical IEMs should satisfy.

We introduce the concept of ANI, which originates from leung2022causal. Let $N_{\bm{A}}^c(i,s) \coloneqq N_n \setminus N_{\bm{A}}(i,s)$ denote the set of units who are more than distance $s$ away from $i$. Writing $\bm{Z}'$ as an independent copy of $\bm{Z}$, define $\bm{Z}_i^{(s)} \coloneqq (\bm{Z}_{N_{\bm{A}}(i,s)}, \bm{Z}'_{N_{\bm{A}}^c(i,s)})$ by combining the sub-vector of $\bm{Z}$ on $N_{\bm{A}}(i,s)$ and that of $\bm{Z}'$ on $N_{\bm{A}}^c(i,s)$. Let $\theta_{n,s}^{\mathrm{ADE}} \coloneqq \max \{ \max_{i \in S_n} \operatorname*{\mathbb{E}} | y_i(\bm{Z}) - y_i(\bm{Z}_i^{(s)}) |, \; \max_{i \in S_n} \operatorname*{\mathbb{E}} | D_i(\bm{Z}) - D_i(\bm{Z}_i^{(s)}) | \}$. This measures the intensity of interference with units that are at least $s$ distance away. By Assumption (ref), $\theta_{n,s}^{\mathrm{ADE}}$ is uniformly bounded in $n$ and $s$.

assumption[ANI 1] $\sup_{n \in \mathbb{N}} \theta_{n,s}^{\mathrm{ADE}} \to 0$ as $s \to \infty$.

The ANI assumption says that spillover effects from units that are sufficiently far away should be sufficiently small. In particular, those not connected with $i$ do not affect the outcome and the treatment response of $i$. Thus, the ANI would be reasonable in practical situations where only nearby people have major impacts on one's behavior. In other words, the assumption precludes situations such as herding behavior or social pressure, whereby one may be influenced by arbitrary unknown others, irrespective of actual connections. Note that the ANI is much weaker than the commonly used clustered interference assumption that requires $\theta_{n,L}^{\mathrm{ADE}}=0$ for some $L$.

Let $S_{\bm{A}}^{\partial}(i, s) \coloneqq \{ j \in S_n: \ell_{\bm{A}}(i, j) = s \}$ be the subset of $S_n$ that are exactly at distance $s$ from unit $i \in S_n$. We denote its $k$-th sample moment as $M_{S_n}^{\partial}(s; k) \coloneqq |S_n|^{-1} \sum_{i \in S_n} | S_{\bm{A}}^{\partial}(i, s)|^k$, which measures the denseness of $\bm{A}$ restricted on $S_n$. When $k = 1$, we write $M_{S_n}^{\partial}(s) = M_{S_n}^{\partial}(s; 1)$. Further, letting $\lfloor \cdot \rfloor$ indicate the floor function, define

align[align omitted — 227 chars of source]
assumption[Weak dependence 1] (i) $\max_{1 \le s \le 2K} M_{S_n}^{\partial}(s) = O(1)$, where $K$ is as given in Assumption (ref). (ii) $|S_n|^{-1} \sum_{s = 1}^{n - 1} M_{S_n}^{\partial}(s) \widetilde \theta_{n,s}^{\mathrm{ADE}} = o(1)$.

Assumption (ref)(i) rules out that there are a non-negligible proportion of units whose $2K$ neighborhoods may grow to infinity as $n$ increases. If one assumes that each individual can hold only a limited number of interacting partners, Assumption (ref)(i) is satisfied with $M_{S_n}^{\partial}(s; k) < \infty$ for all $s,k < \infty$. However, for example, it is violated if the network is a complete graph. Assumption (ref)(ii) is analogous to Assumption 5 of leung2022causal and Assumption 3.2 of kojevnikov2021limit. This assumption restricts the rate of convergence of $\widetilde \theta_{n,s}^{\mathrm{ADE}}$. For example, in the case of a ring network, we can see that $M_{S_n}^{\partial}(s) \le 2$ for all $s$, and Assumption (ref)(ii) is reduced to the condition $|S_n|^{-1} \sum_{s = 1}^{n - 1} \widetilde \theta_{n, s}^{\mathrm{ADE}} = o(1)$.

The following theorem establishes the consistency of the ADE estimators.

theoremSuppose that Assumptions (ref) -- (ref) hold. Then, we have (i) $\widehat{\mathrm{ADEY}}_{S_n}(t) - \mathrm{ADEY}_{S_n}(t) \stackrel{p}{\to} 0$ and (ii) $\widehat{\mathrm{ADED}}_{S_n}(t) - \mathrm{ADED}_{S_n}(t) \stackrel{p}{\to} 0$. Additionally, if Assumptions (ref) -- (ref) hold, we have (iii) $\widehat{\mathrm{LADE}}_{S_n}(t) - \mathrm{LADE}_{S_n}(t) \stackrel{p}{\to} 0$.
remark[Rate of convergence] The convergence rates of the proposed estimators are determined by the convergence rate given in Assumption (ref)(ii); see Lemma B.2. In particular, $\sqrt{|S_n|}$-consistency can be achieved if Assumption (ref)(ii) is strengthened to $\sum_{s = 1}^{n - 1} M_{S_n}^{\partial}(s) \widetilde \theta_{n,s}^{\mathrm{ADE}} = O(1)$.

Next, we investigate the asymptotic distributions of the ADE estimators. For each $i \in S_n$, let

align*[align* omitted — 493 chars of source]

where

align[align omitted — 466 chars of source]

For notational simplicity, we suppress the dependence of $V$'s and $W$'s on the IEM value $t$; the same notation applies to other variables introduced below. In the proof of the theorem presented below, we will show that the asymptotic distribution of $\widehat{\mathrm{ADEY}}_{S_n}(t) - \mathrm{ADEY}_{S_n}(t)$ can be obtained by that of $|S_n|^{-1} \sum_{i \in S_n}(V_i^{\mathrm{ADEY}} - \operatorname*{\mathbb{E}} [V_i^{\mathrm{ADEY}}])$. Similar results hold for the other cases. Let $(\sigma_{S_n}^\mathrm{ADEY})^2 \coloneqq \operatorname*{\mathrm{Var}}[ |S_n|^{-1/2} \sum_{i \in S_n} V_i^{\mathrm{ADEY}} ]$, and similarly we define $( \sigma_{S_n}^\mathrm{ADED} )^2$ and $( \sigma_{S_n}^\mathrm{LADE} )^2$.

To derive the asymptotic distributions, we employ the central limit theorem (CLT) for $\psi$-weakly dependent processes in kojevnikov2021limit (see Definition B.1). Under Assumptions (ref) -- (ref), for each $V_i = V_i^{\mathrm{ADEY}}$, $V_i^{\mathrm{ADED}}$, and $V_i^{\mathrm{LADE}}$, we show that $\{ V_i \}_{i \in S_n}$ is a $\psi$-weakly dependent process with the dependence coefficients $\{ \widetilde \theta_{n,s}^{\mathrm{ADE}} \}_{s \ge 0}$. Then, we can apply their CLT to our context with additional restrictions on the network structure. Let $S_{\bm{A}}(i, s) \coloneqq \{ j \in S_n : \ell_{\bm{A}}(i, j) \le s \}$ and $\Delta_{S_n}(s, m; k) \coloneqq |S_n|^{-1} \sum_{i \in S_n} \max_{j \in S^\partial_{\bm{A}}(i, s)} |S_{\bm{A}}(i, m) \setminus S_{\bm{A}}(j, s - 1)|^k$, where $S_{\bm{A}}(j, s - 1) = \varnothing$ if $s = 0$. This is the $k$-th sample moment of the maximum number (over $j$'s at distance $s$ from $i$) of units who are within distance $m$ from $i$ but at least distance $s$ apart from $j$. Note that $\Delta_{S_n}(s, m; k)$ increases as $m$ becomes larger, but at the same time decreases fast to zero as $s$ grows because $S_{\bm{A}}(j, s - 1)$ tends to become large quickly; for example, if all units have approximately $L$ links, $|S_{\bm{A}}(j, s - 1)| = O(L^{s-1})$. In addition, we define $c_{S_n}(s, m; k) \coloneqq \inf_{\alpha > 1} [\Delta_{S_n}(s, m; k \alpha)]^{\frac{1}{\alpha}} [ M_{S_n}^{\partial} (s; \alpha/(\alpha - 1) ) ]^{1 - \frac{1}{\alpha}}$. This quantity measures the denseness of the network, which plays an important role in establishing the CLT.

assumption[Weak dependence 2] For each $\sigma_{S_n} = \sigma_{S_n}^\mathrm{ADEY}$, $\sigma_{S_n}^\mathrm{ADED}$, and $\sigma_{S_n}^\mathrm{LADE}$, there exist some positive sequence $m_n \to \infty$ and a constant $0 < \varepsilon < 1$ such that for each $k \in \{ 1, 2 \}$, (i) $|S_n|^{-k/2} \sigma_{S_n}^{-(2 + k)} \sum_{s=0}^{n-1} c_{S_n}(s, m_n; k) (\widetilde \theta_{n,s}^{\mathrm{ADE}})^{1-\varepsilon} \to 0$ and (ii) $|S_n|^{k/2}\sigma_{S_n}^{-k} (\widetilde \theta_{n,m_n}^{\mathrm{ADE}})^{1 - \varepsilon} \to 0$.

This corresponds to Assumption 3.4 of kojevnikov2021limit.\footnote{ Note that Assumption (ref) is weaker than Assumption 3.4 of kojevnikov2021limit. This comes from the following two facts. First, the $\psi$-weak dependent processes considered here are uniformly bounded by Assumptions (ref) and (ref), while kojevnikov2021limit only assume the existence of $4 + \varepsilon$ moments of them. Second, they consider a more general form of $\psi$-function than ours. See Assumption 2.1 of their paper and Lemma B.3. } Note that it restricts not only the network structure but also our choice of sub-population $S_n$. In particular, when there exist constants $\underline{C}$ and $\bar C$ such that $0 < \underline{C} \le \sigma_{S_n} \le \bar C < \infty$ for all sufficiently large $n$, Assumption (ref) can be reduced to (i) $|S_n|^{-k/2} \sum_{s=0}^{n-1} c_{S_n}(s, m_n; k) (\widetilde \theta_{n, s}^{\mathrm{ADE}})^{1 - \varepsilon} \to 0$ and (ii) $|S_n|^{k/2} (\widetilde \theta_{n, m_n}^{\mathrm{ADE}})^{1 - \varepsilon} \to 0$.

The following theorem shows that the ADE estimators are asymptotically normal.

theoremSuppose that Assumptions (ref) -- (ref) hold. Then, we have \begin{align*} \begin{array}{cl} (i) & \frac{\sqrt{|S_n|} \left( \widehat{\mathrm{ADEY}}_{S_n}(t) - \mathrm{ADEY}_{S_n}(t) \right) }{\sigma_{S_n}^{\mathrm{ADEY}}} \stackrel{d}{\to} \mathrm{Normal}(0, 1)\\ (ii) & \frac{\sqrt{|S_n|} \left( \widehat{\mathrm{ADED}}_{S_n}(t) - \mathrm{ADED}_{S_n}(t) \right) }{\sigma_{S_n}^{\mathrm{ADED}}} \stackrel{d}{\to} \mathrm{Normal}(0, 1) \end{array} \end{align*} provided that $(\sigma_{S_n}^{\mathrm{ADEY}})^{-1} = O(1)$ and $(\sigma_{S_n}^{\mathrm{ADED}})^{-1} = O(1)$. Additionally, if Assumptions (ref) -- (ref) hold, we have \begin{align*} \begin{array}{cl} (iii) & \frac{\sqrt{|S_n|} \left( \widehat{\mathrm{LADE}}_{S_n}(t) - \mathrm{LADE}_{S_n}(t) \right) }{\sigma_{S_n}^{\mathrm{LADE}}} \stackrel{d}{\to} \mathrm{Normal}(0, 1), \end{array} \end{align*} provided that \begin{align} \frac{1}{\sigma_{S_n}^{\mathrm{LADE}}} = O(1), \qquad \frac{\sigma_{S_n}^{\mathrm{ADEY}} \sigma_{S_n}^{\mathrm{ADED}}}{\sqrt{|S_n|} \sigma_{S_n}^{\mathrm{LADE}}} = o(1), \qquad \frac{(\sigma_{S_n}^{\mathrm{ADED}})^2}{\sqrt{|S_n|} \sigma_{S_n}^{\mathrm{LADE}}} = o(1). \end{align}

The conditions in (ref) are fairly mild, which are satisfied especially with $\sqrt{|S_n|}$-consistency.

In Appendix C, we consider inference methods based on network HAC estimation and a wild bootstrap approach. It is shown that the HAC estimators have asymptotic biases due to the fact that we cannot estimate heterogeneous means in the asymptotic variances in Theorem (ref). This is a well-known issue in the design-based uncertainty framework (cf. imbens2015causal).

Average indirect effects

Next, we focus on the estimators of $\mathrm{AIEY}_{S_n}$, $\mathrm{ADED}_{S_n}$, and $\mathrm{LAIE}_{S_n}$; the discussion of $\mathrm{AIED}_{S_n}$ is analogous. The asymptotic properties of these estimators can be derived in the same way as in the previous subsection, but we impose an additional condition on the denseness of the network and an ANI condition slightly different from Assumption (ref). Let $\theta_{n,s}^{\mathrm{AIE}} \coloneqq \max \{ \max_{i \in S_n} \max_{j \in \mathcal{E}_i} \operatorname*{\mathbb{E}} | y_j(\bm{Z}) - y_j(\bm{Z}_i^{(s)}) |, \;\; \max_{i \in S_n} \operatorname*{\mathbb{E}} | D_i(\bm{Z}) - D_i(\bm{Z}_i^{(s)}) | \}$. Here, $\operatorname*{\mathbb{E}} | y_j(\bm{Z}) - y_j(\bm{Z}_i^{(s)}) |$ measures to what extent the outcome of $j \in \mathcal{E}_i$ is affected by the IVs of units that are apart from $i$ more than $s$ distance. Define $\widetilde \theta_{n,s}^{\mathrm{AIE}}$ in the same way as in (ref).

assumption[ANI 2] $\sup_{n \in \mathbb{N}} \theta_{n,s}^{\mathrm{AIE}} \to 0$ as $s \to \infty$.
assumption[Weak dependence 3] Assumption (ref)(i) -- (ii) hold when $\widetilde \theta_{n,s}^{\mathrm{ADE}}$ is replaced by $\widetilde \theta_{n,s}^{\mathrm{AIE}}$. Additionally, (iii) $\max_{i \in S_n} |\mathcal{E}_i| = O(1)$.

Similar to Assumption (ref), Assumption (ref) restricts the “sparseness” of the network in several ways. For example, Assumption (ref)(i) and (iii) are fulfilled if the degree in the network $\bm{A}$ is uniformly bounded in $i \in N_n$ and $n \in \mathbb{N}$. Moreover, Assumption (ref)(ii) is also satisfied if $|S_n|^{-1} \sum_{s = 1}^{n-1} \widetilde \theta_{n,s}^{\mathrm{AIE}} = o(1)$ additionally holds. However, Assumption (ref) rules out, for example, a small-world network, where any pair of units are connected within a short distance.

Given these assumptions, we can prove the following consistency results.

theoremSuppose that Assumptions (ref) -- (ref) and (ref) -- (ref) hold. Then, we have (i) $\widehat{\mathrm{AIEY}}_{S_n} - \mathrm{AIEY}_{S_n} \stackrel{p}{\to} 0$ and (ii) $\widehat{\mathrm{ADED}}_{S_n} - \mathrm{ADED}_{S_n} \stackrel{p}{\to} 0$. Additionally, if Assumptions (ref) -- (ref) hold, we have (iii) $\widehat{\mathrm{LAIE}}_{S_n} - \mathrm{LAIE}_{S_n} \stackrel{p}{\to} 0$.
remark[Denseness] Assumption (ref)(iii) requires that the size of the interference set is uniformly bounded in $i \in S_n$ and $n \in \mathbb{N}$, which plays an essential role in the proof of Theorem (ref). If the network at hand is denser than that considered in Assumption (ref), $\mathcal{E}_i$ may grow with the sample size and our AIE estimators may not achieve the consistency; see Proposition 5 of li2022random for a related result. Hence, we should be cautious about the denseness of the network especially when estimating the AIE parameters.

To discuss the asymptotic normality results, for each $i \in S_n$, define

align*[align* omitted — 665 chars of source]

where

align[align omitted — 507 chars of source]

Let $(\sigma_{S_n}^\mathrm{AIEY})^2 \coloneqq \operatorname*{\mathrm{Var}}[ |S_n|^{-1/2} \sum_{i \in S_n} V_{\mathcal{E}_i}^{\mathrm{AIEY}} ]$, and define $( \sigma_{S_n}^\mathrm{LAIE} )^2$ analogously. With an abuse of notation, let us denote $( \sigma_{S_n}^\mathrm{ADED} )^2 \coloneqq \operatorname*{\mathrm{Var}}[ |S_n|^{-1/2} \sum_{i \in S_n} V_{\mathcal{E}_i}^{\mathrm{ADED}} ]$.

assumption[Weak dependence 4] For each $\sigma_{S_n} = \sigma_{S_n}^\mathrm{AIEY}$, $\sigma_{S_n}^\mathrm{ADED}$, and $\sigma_{S_n}^\mathrm{LAIE}$, Assumption (ref) holds when $\widetilde \theta_{n,s}^{\mathrm{ADE}}$ is replaced by $\widetilde \theta_{n,s}^{\mathrm{AIE}}$.

The following theorem presents the asymptotic normality results.

theoremSuppose that Assumptions (ref) -- (ref) and (ref) -- (ref) hold. Then, we have \begin{align*} \begin{array}{cl} (i) & \frac{\sqrt{|S_n|} \left( \widehat{\mathrm{AIEY}}_{S_n} - \mathrm{AIEY}_{S_n} \right) }{\sigma_{S_n}^{\mathrm{AIEY}}} \stackrel{d}{\to} \mathrm{Normal}(0, 1)\\ (ii) & \frac{\sqrt{|S_n|} \left( \widehat{\mathrm{ADED}}_{S_n} - \mathrm{ADED}_{S_n} \right) }{\sigma_{S_n}^{\mathrm{ADED}}} \stackrel{d}{\to} \mathrm{Normal}(0, 1) \end{array} \end{align*} provided that $(\sigma_{S_n}^{\mathrm{AIEY}})^{-1} = O(1)$ and $(\sigma_{S_n}^{\mathrm{ADED}})^{-1} = O(1)$. Additionally, if Assumptions (ref) -- (ref) hold, we have \begin{align*} \begin{array}{cl} (iii) & \frac{\sqrt{|S_n|} \left( \widehat{\mathrm{LAIE}}_{S_n} - \mathrm{LAIE}_{S_n} \right) }{\sigma_{S_n}^{\mathrm{LAIE}}} \stackrel{d}{\to} \mathrm{Normal}(0, 1), \end{array} \end{align*} provided that the conditions in (ref) hold when $\sigma_{S_n}^{\mathrm{ADEY}}$ and $\sigma_{S_n}^{\mathrm{LADE}}$ are replaced by $\sigma_{S_n}^{\mathrm{AIEY}}$ and $\sigma_{S_n}^{\mathrm{LAIE}}$, respectively.

Numerical Illustrations

Monte Carlo simulation

We investigate the finite sample properties of our methods using a set of Monte Carlo experiments. We conduct these experiments based on an artificial ring-shaped network and on a real students' friendship network separately. To save space, the detailed experimental setups and simulation results are summarized in Appendix G.

The main findings are as follows: First, regardless of whether the IEM is correctly- or mis-specified, our estimators work satisfactorily well with sufficiently small biases. Second, the RMSE values for the LADE can be significantly large in some cases. This is because in some situations, the estimated ADED is nearly zero, resulting in extremely large LADE estimates. This phenomenon is not persistent in other DGPs where the ADED is sufficiently away from zero. Third, overall, we can observe that the empirical coverage rates are reasonably close to the nominal 95% level for both HAC and bootstrap estimators in the experiments based on the artificial network. For the experiments based on the real network, the coverage rates tend to be slightly below the nominal level. Considering that these confidence intervals contain non-negligible asymptotic biases, the above results should be attributable to this bias to some extent.

Empirical illustration

We apply the proposed methods to the data from paluck2016changing's (paluck2016changing) field experiment on anti-conflict intervention programs at American middle schools. During the 2012-2013 school year, the research team organized intervention meetings to help students identify common conflict behaviors in their schools and instruct them on behavioral strategies to mitigate conflicts.

The data include $n = 24,471$ students in 56 public middle schools in the state of New Jersey. A group of students (called {\it seed-eligible students}) were non-randomly selected by the research team, and half of these students (called {\it seed students} or {\it treatment-eligible students}) were randomly invited to join the program.

The students' social networks were measured by asking them to nominate up to 10 students in their school with whom they had spent time in person or online in the past few weeks. We construct a symmetric adjacency matrix $\bm{A}$ by treating the pair of students as friends if either student nominated the other, as in aronow2017estimating.

In our analysis, $Z_i \in \{ 0, 1 \}$ indicates whether student $i$ received an invitation to the program (i.e., whether student $i$ was a seed student), and $D_i \in \{ 0, 1 \}$ represents the participation in the intervention program (i.e., whether student $i$ attended at least one meeting). Let $Y_i \in \{ 0, 1 \}$ be an indicator for the wearing of a program wristband given by the treated students as a reward to students for engaging in conflict-mitigating behaviors. This is regarded as a proxy variable of student's willingness to endorse anti-conflict norms and behaviors, and the same outcome variable is used in aronow2017estimating and leung2022causal. We consider the following two IEMs: $T_{1i} = \bm{1}\{ \sum_{j \neq i} A_{ij} Z_j > 0 \}$ and $T_{2i} = \bm{1}\{ \sum_{j \neq i} A_{ij} D_j > 0 \}$, respectively labeled as “IEM1” and “IEM2”. For the estimation of ADEs conditional on $T_i = t$, in view of Assumption (ref)(ii), we focus on the sub-populations $S_n(\delta) = \{ i \in N_n: i$ is seed-eligible and has $\delta$ seed-eligible friend(s) $\}$ for $\delta \in \{ 1, 2, 3 \}$. Meanwhile, when estimating the AIE parameters, we consider the sub-population $S_n^{\ge 1} = \{ i \in N_n: i$ is seed-eligible and has at least one friend $\}$.

Panels (a) and (b) of Table (ref) present the ITT estimates for the ADE with the standard errors based on the network HAC estimation using bandwidth $b_n = 2$.\footnote{ Almost the same HAC estimates were obtained for other bandwidths $b_n \in \{1, 3\}$. In addition, the wild bootstrap produced similar standard errors to those reported here. To save space, we omit those results. } Receiving an invitation has a statistically significant positive effect on the probability of wearing a wristband, which is consistent with previous findings (e.g., aronow2017estimating; leung2022causal). For example, the estimate of $\mathrm{ADEY}_{S_n(1)}(0)$ for IEM1 indicates that receiving an invitation leads to about a nine percentage point increase in the probability of wearing a wristband for the seed-eligible students whose seed-eligible friend is not treatment-eligible. Similarly, the ADED estimates indicate positive effects of receiving an invitation on the probability of participation, which supports the IV relevance condition in Assumption (ref).

The LADE estimates and their standard errors based on the HAC estimation are also reported in panels (a) and (b) of Table (ref). For example, the estimate of $\mathrm{LADE}_{S_n(1)}(1)$ based on IEM1 indicates an eighteen percentage point increase in the probability of wearing a wristband for the seed-eligible students who have a treatment-eligible friend. Importantly, the LADE estimates tend to be larger than the corresponding ITT estimates, implying that the ITT analysis might underestimate the effect of the anti-conflict intervention program.

Nonetheless, we should be cautious in interpreting the LADE estimates because the interpretation of LADE essentially depends on which sufficient condition we consider for Assumption (ref). As discussed in Section (ref), the sufficient condition in (ref) would be met here, and the LADE aggregates the direct effect from participating in the intervention program and the spillover effect from the student's own treatment eligibility. However, since one's treatment eligibility seems to have little impact on the others' treatment choice as observed below, the LADE should mainly account for the direct effect of the intervention program.

Panel (c) of Table (ref) presents the AIE estimates when we set $K = 1$ in line with IEM1 and IEM2. The estimates of $\mathrm{AIEY}_{S_n}$ and $\mathrm{LAIE}_{S_n}$ are positive and statistically significant, indicating substantial spillover effects of one's treatment eligibility and treatment take-up on others' wristband wearing. By contrast, the estimated $\mathrm{AIED}_{S_n}$ provides no strong evidence on such spillovers between treatment eligibility and treatment decision.

table[table omitted — 4,419 chars of source]

\if11 {

center[center omitted — 44 chars of source]

The authors are grateful to the co-editor, the associate editors, three anonymous referees, and Ryo Okui for their beneficial comments and suggestions. This work was supported by JSPS KAKENHI Grant Numbers 19H01473 and 20K01597. The data set used in this study is available through the Inter-university Consortium for Political and Social Research (paluck2020changing). } \fi

\if01 \fi

center[center omitted — 47 chars of source]

The authors report there are no competing interests to declare.

center[center omitted — 49 chars of source]
description• The proofs of all technical results and the other supplementary results (PDF file) • The ACC form and the codes to reproduce the computational results (ZIP file)

\setcounter{page}{1} \setcounter{table}{0}

center[center omitted — 357 chars of source]
abstractAppendices (ref) and (ref) provide the proofs of all technical results in the main text. In Appendix (ref), we develop statistical inference methods based on network HAC estimation and a wild bootstrap approach. We consider the identification and estimation of average overall effects and average spillover effects in Appendices (ref) and (ref), respectively. Appendix (ref) contains the additional discussion of the identification analysis developed in Section (ref). In Appendix (ref), we report the results of Monte Carlo experiments. Appendix (ref) presents the additional empirical results.

\spacingset{1.9}