EconBase
← Back to paper

Endogenous Interference in Randomized Experiments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

127,624 characters · 20 sections · 126 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Endogenous Interference in Randomized Experiments

\onehalfspacing

\thispagestyle{plain}

\pagenumbering{arabic}

spacing{1.3} \begin{abstract} This paper investigates the identification and inference of treatment effects in randomized controlled trials with social interactions. Two key network features characterize the setting and introduce endogeneity: (1) latent variables may affect both network formation and outcomes, and (2) the intervention may alter network structure, mediating treatment effects. I make three contributions. First, I define parameters within a post-treatment network framework, distinguishing direct effects of treatment from indirect effects mediated through changes in network structure. I provide a causal interpretation of the coefficients in a linear outcome model. For estimation and inference, I focus on a specific form of peer effects, represented by the fraction of treated friends. Second, in the absence of endogeneity, I establish the consistency and asymptotic normality of ordinary least squares estimators. Third, if endogeneity is present, I propose addressing it through shift-share instrumental variables, demonstrating the consistency and asymptotic normality of instrumental variable estimators in relatively sparse networks. For denser networks, I propose a denoised estimator based on eigendecomposition to restore consistency. Finally, I revisit Prina2015 as an empirical illustration, demonstrating that treatment can influence outcomes both directly and through network structure changes. \end{abstract}

KEYWORDS: Causal inference, endogeneity, interference, mediation analysis, peer effect, random graph, shift-share instrument variable.

Introduction

Peer effects have been extensively studied in the economics literature. However, identifying these effects can be challenging without randomized experiments. Recent research has integrated peer effects with randomized controlled trials (RCTs) to improve their identification across various fields, including education, microfinance, public health, agriculture, and social psychology.\footnote{See, for example, Sacerdote2001a, MiguelKremer2004, Sobel2006a, BanerjeeChandrasekharDuflo2013, CaiJanvrySadoulet2015 and PaluckShepherdAronow2016} These studies go beyond direct effects, exploring how interventions spread through social interaction, potentially amplifying or dampening their impact. This introduces a phenomenon known as “interference.”\footnote{The term “interference” originates from the assumption of no interference between individuals in the potential outcomes framework for causal inference Cox1958, Rubin1980. It initially referred to how spillover effects disrupted or “interfered with” the standard comparison between treated and untreated groups, complicating the estimation of direct treatment effects. However, the concept of interference has evolved, and in many contexts, it now encompasses situations where spillover effects are not just nuisances but are themselves of primary interest.} The literature on spillover effects often assumes exogenous networks, where no latent variables influence both network formation and outcomes, and the network remains unchanged after the intervention. However, empirical evidence suggests that networks can be endogenous even in randomized experiments, as latent variables may influence both network formation and outcomes, and treatment can alter the network structure. Motivated by these concerns, this paper investigates the identification and inference of causal effects in RCTs with interference, accounting for both sources of endogeneity.

Many empirical studies suggest that network structures are influenced by latent variables and evolve in response to interventions. For example, Prina2015 conducted an RCT offering savings accounts to villagers in Nepal and found significant network changes before and after the intervention. Similarly, BanerjeeBrezaChandrasekhar2023 found that introducing formal financial institutions in Indian villages reduced social connections, as access to formal institutions diminished the need for social ties, demonstrated through both an observational study and an RCT. BarnhardtFieldPande2017 analyzed a housing lottery program designed to improve housing situations but found that it increased isolation from family and exacerbated financial insecurity. Likewise, CarrellSacerdoteWest2013 conducted a group formation experiment to improve low-ability students' performance, but the intervention backfired as low-ability students formed stronger bonds with similar peers, worsening their outcomes.

In empirical studies, researchers often rely on pre-intervention network data to fit regressions or control for unobserved confounders CarterLaajajYang2021a. However, this approach can be problematic. When post-intervention networks ultimately influence outcomes, relying on pre-intervention networks introduces measurement error. Conversely, using post-intervention networks, which are shaped by both treatment and latent variables, introduces endogeneity issues, even in randomized experiments. Specifically, treatment-induced changes in the network make it challenging to disentangle the direct effects of treatment from the indirect effects mediated by changes in network structure. Separating these effects is crucial for understanding the mechanisms of the intervention and designing effective policies. Furthermore, latent variables complicate the identification and estimation of causal effects, potentially leading to biased and inconsistent estimates of causal effects. This paper demonstrates a novel approach to utilizing panel network data for separating and consistently estimating these effects by fitting the regression with post-intervention network data while constructing instrumental variables from pre-intervention network data.

To account for network evolution induced by the intervention, I apply the mediation analysis framework Pearl2001, HeckmanPinto2015 to distinguish the direct effects of treatment from the indirect effects mediated by changes in network structure. Researchers often use {effective treatment}, also known as {exposure mapping}, for dimensionality reduction Manski2013, AronowSamii2017. This low-dimensional statistic captures interference patterns by mapping units, treatment assignments, and network structures to the exposures each unit receives. In my context, since the network is a post-treatment variable, I treat the exposure mapping as a mediator transmitting the indirect effect of the treatment on the outcome. I define the causal parameters within this framework, distinguishing between the direct effect of the treatment on the outcome and the indirect effect mediated by changes in the network. I assume a linear outcome model that incorporates both the treatment and the network mediator while allowing for additive and flexible forms of unobserved confounding. In this model, the coefficients have clear causal interpretations, and the framework can be applied to any mediator.

For estimation and inference, I focus on a specific mediator defined as the fraction of treated friends. This approach aligns with the anonymous interference assumption from HudgensHalloran2008a and Manski2013. Consequently, the linear model considered in this paper represents a special case of the linear-in-means (LIM) model Manski1993a, BlumeBrockDurlauf2015a, focusing on contextual peer effects, which captures the mean impact of friends' treatments. Within this LIM framework, the endogeneity issue arises only when both conditions are met: the network depends on the treatment but is not mean-independent, and an unobserved confounder is present. Although this paper focuses on a specific form of peer effects, the discussion of the SSIV relevance condition offers valuable insights into other types of peer effects.

When no endogeneity is present, I propose using OLS estimation with post-intervention network data. I demonstrate the consistency and asymptotic normality of the OLS estimators across different levels of network sparsity. With the fraction mediator, I handle the diminishing variation of the regressor in denser networks and find that treatment-induced network changes can increase sample variation in the fraction and enhance the convergence rate. Additionally, I demonstrate that the standard heteroskedasticity-consistent variance estimator ensures valid inference, even with dependent regressors and nonstandard convergence rates in estimation.

When unobserved confounders are present, I address endogeneity using shift-share (or “Bartik”) instrumental variables (SSIV) Bartik1991. The SSIV mitigates endogeneity by combining a set of shocks using exposure share weights. This paper extends the shift-share design to the network setting by constructing the IV as a combination of the random shock (others' treatment assignments) and non-exogenous exposure (pre-intervention network structure), following BorusyakHull2023. I provide a thorough examination of the conditions required for a valid IV, with particular emphasis on the relevance condition across varying levels of sparsity. I demonstrate that the relevance condition for SSIV fails as networks become denser, as increased overlap in friendships reduces variation in SSIV across units, which limits its effectiveness as an instrument for the endogenous network mediator. Specifically, I show that IV estimators using SSIV are consistent in relatively sparse networks. In denser networks, where IV estimators using SSIV may be inconsistent, I adapt the approach of LiWager2022 to propose an eigendecomposition-based denoised SSIV estimator for consistent estimation. Furthermore, I establish the asymptotic normality of the IV estimators using both SSIV and modified SSIV and provide consistent variance estimators that account for unit dependencies introduced by the SSIV, ensuring valid inference.

I present simulation evidence supporting my theoretical results for both OLS and IV estimators using SSIV and modified SSIV. These results are derived in an asymptotic framework where networks are modeled as random graphs with varying sparsity rates. The estimation results across various sample sizes and sparsity levels confirm the asymptotic results of my theoretical analysis.

For empirical illustration, I revisit Prina2015, which offered access to formal savings accounts to a random sample of female household heads in 19 villages in Nepal. My results demonstrate that IV estimation effectively addresses the endogeneity issue arising from unobserved confounders. Results from IV estimation and the OLS estimates in Prina2015 can differ in both sign and significance when the indirect effect is significant. Additionally, I find that the intervention influences various outcomes through multiple channels: some are directly impacted by the treatment, while others operate through changes in the network mediator, which captures patterns of interference.

Related Literature

First, it builds on the peer effects literature, which explores how individuals' outcomes are shaped by their peers' behaviors, actions, or characteristics. The standard empirical framework for peer effects is the LIM model, which assumes that agents are influenced by the average actions of their peers. The seminal work of Manski1993a formalizes the reflection problem, highlighting the difficulty of disentangling endogenous peer effects from correlated and contextual influences within the LIM framework. BramoulleDjebbariFortin2009b extend this line of inquiry by characterizing the identification conditions for peer effects in network settings, providing crucial insights for empirical applications. In a critical review of the literature, Angrist2014 scrutinizes various econometric methods and empirical studies on peer effects, offering sharp critiques of existing approaches and underscoring key challenges in this field.

The literature outlines four broad strategies for identifying peer effects while accounting for correlated effects: random peers, random shocks, structural endogeneity, and panel data BramoulleDjebbariFortin2020. Researchers estimate causal peer effects by studying contexts with randomly assigned peers, through natural or artificial experiments Sacerdote2001a, CarrellSacerdoteWest2013, DeGiorgiPellizzariRedaelli2010a. Identification relies on the fact that random assignment ensures an agent's characteristics, both observed and unobserved, are uncorrelated with those of her peers. When peers are not random, researchers seek to identify peer effects using other sources of exogenous variation, such as randomized interventions or quasi-random experiments DieyeDjebbariBarrera-Osorio2014, NicolettiSalvanesTominey2018, DeGiorgiFrederiksenPistaferri2020, ArduiniPatacchiniRainone2020. In the absence of a clear source of exogenous variation, researchers have developed structural frameworks to address correlated effects and identify peer effects in networks Goldsmith-PinkhamImbens2013b, Graham2015, Graham2017, HsiehLee2016a, HsiehLeeBoucher2019, JohnssonMoon2019, employing methods such as Bayesian approaches and control function techniques. Combined with panel data, the inclusion of individual fixed effects enables researchers to control for agents' time-invariant unobserved characteristics, helping to address issues of correlated effects ArcidiaconoFosterGoodpaster2012, DeGiorgiFrederiksenPistaferri2020, ComolaPrina2021, dePaulaRasulSouza2023.

This paper contributes to the peer effects literature by leveraging randomized treatment and panel network data in a novel way for identification and estimation, accounting for treatment-induced network changes. Specifically, I combine randomized treatments and pre-intervention networks to construct SSIV, instrumenting endogenous variables from the post-intervention network. This approach is conceptually similar to the use of lagged variables as instruments in panel data models, as developed by ArellanoBond1991 and later expanded by BunSarafidis2015. By incorporating pre-intervention networks into the SSIV framework, I aim to address potential endogeneity issues while capturing the dynamic evolution of network structures over time. The most related work is ComolaPrina2021, who study interventions that impact network structure. They allow for endogenous peer effects but assume conditional exogeneity of pre- and post-intervention networks, addressing different endogeneity sources. They adapt “lagged” partner characteristics as instruments, focusing on identification conditions rather than asymptotic properties. Other recent studies explore dynamic networks, often focusing on edge-level variables. Auerbach2022 develops a test to assess how treatments like social programs or trade shocks impact network link formation, while AuerbachCai2023 examine social disruption by analyzing the formation and dissolution of network connections in response to policy using a random or quasi-random assignment framework.

Second, I contribute to the shift-share IV literature by demonstrating the relevance condition of SSIV across a broad range of network sparsity regimes. Shift-share specifications are increasingly common in many contexts, including labor, public, development, macroeconomics, international trade, and finance; see Card2009, AutorDornHanson2013, NakamuraSteinsson2014, BourveauSheZaldokas2020 and Breuer2022. There are two main approaches to identification in the SSIV literature. Bartik1991 and Goldsmith-PinkhamSorkinSwift2020 suggest identification based on the exogeneity of the exposure shares. In contrast, BorusyakHullJaravel2022 suggest identification through the quasi-random assignment of shocks, which allows for endogenous exposure shares, as do AdaoKolesarMorales2019 and BorusyakHull2023. AdaoKolesarMorales2019 investigate inference in shift-share regression designs and develop new results for the inference that remain valid even in the presence of arbitrary cross-regional correlation in the regression residuals. Their findings suggest that cluster-robust standard errors, commonly reported in such settings, may lead to the overrejection of the null hypothesis. BorusyakHull2023 extend the SSIV idea to the nonlinear case, where multiple sources of variation are combined according to a known formula. However, these papers impose a key condition for the relevance of SSIV to hold. Specifically, most observations are primarily exposed to a small number of shocks influencing treatment. In this paper, I characterize the regimes in which the relevance condition holds, ensuring that IV estimators with SSIV lead to consistent estimation. I also establish a connection to the work of LiWager2022, noting that their estimator for indirect effects fundamentally employs the SSIV approach to address the endogeneity issue of unobserved confounding.

Third, I contribute to the interference literature by accounting for stochastic and treatment-induced network. The concept of “interference” challenges the “stable unit treatment value assumption” (SUTVA), a foundational principle of classical causal inference that assumes no interference between units Cox1958, Rubin1980, ImbensRubin2015. The existing literature on estimating treatment effects under interference primarily follows a design-based approach HudgensHalloran2008a, AronowSamii2017, AbadieAtheyImbens2020, Leung2022, GaoDing2023. It makes no assumptions about outcome models and network formation, and inference is based on random treatment assignment. Given the stochastic nature of networks, it is natural to question how this affects inference. Recent studies have started exploring the impact of network stochasticity by modeling network graphs as realizations from an (unknown) graphon. For example, Leung2020 analyzes nonparametric and regression estimators for treatment and spillover effects in sparse networks, while LiWager2022 examines the asymptotics of treatment effect estimation under network interference, allowing for arbitrary dependencies on unobserved latent variables. However, these studies do not consider how treatment-induced changes in network structure further influence causal effects. Many studies on interference assume partial interference, with units divided into clusters where interference occurs only within each cluster Sobel2006a, HudgensHalloran2008a, TchetgenVanderWeele2012, LiuHudgensSaul2019. In contrast, this paper examines a single large network, allowing for arbitrary interference patterns.

Our paper also relates to the literature on mediation analysis VanderweeleHongJones2013, Pearl2001, HeckmanPinto2015, ChengGuoLiu2022 by treating exposure mapping as a mediator, accounting for post-intervention network changes. Previous studies propose regression estimators for latent mediation but often lack rigorous asymptotic theory or causal interpretation CheJinZhang2021, LiuJinZhang2021, DiMariaAbbruzzoLovison2022. HayesFredricksonLevin2023 consider latent network positions while excluding interference, while Sweet2018, SweetAdhikari2020, and GuhaRodriguez2021 treat entire networks as mediators.

\paragraph*{Organization of the paper} In Section (ref), I introduce the framework, define the parameters of interest, and provide the identification results. In Section (ref), I discuss the endogeneity issue associated with the fraction mediator and the common practice of using pre-intervention network data to mitigate it. In Section (ref), I explore estimation across various cases. I first present the asymptotic properties of OLS estimators in the absence of endogeneity. I then account for unobserved confounders and employ SSIV for estimation, analyzing its asymptotic properties. Additionally, I propose a modification to SSIV for cases where the network gets denser. Section (ref) presents the results of the Monte Carlo simulations. Section (ref) illustrates the results in an empirical application based on the RCT in Prina2015. Section (ref) offers the concluding remarks. The appendix of the paper collects all of the proofs, as well as some intermediate results.

\paragraph*{Notation} I use $O(), O_{\mathbb{P}}(), o_{\mathbb{P}}(), \asymp, \succ, \prec, \succcurlyeq, \preccurlyeq$ in the following sense: $a_n = O(b_n)$ if $|a_n| \leq C b_n$ for $n$ large enough; $X_n = O_{\mathbb{P}}(b_n)$, if for any $\delta > 0$, there exists $M, N > 0$, s.t. $\mathbb{P}[|X_n| \geq Mb_n] \leq \delta$ for any $n > N$; $X_n = o_{\mathbb{P}}(b_n)$, if $\lim \mathbb{P}[|X_n| \geq \varepsilon b_n] \rightarrow 0$ for any $\varepsilon > 0$; $a_n \asymp b_n$ if there exists $k_1, k_2>0$ and $n_0$, s.t. for all $n > n_0$, $k_1 a_n \leq b_n \leq k_2 a_n$; $a_n \succ b_n$ if $\lim a_n/b_n = \infty$; $a_n \prec b_n$ if $\lim a_n/b_n = 0$; $a_n \succcurlyeq b_n$ if $b_n = O(a_n)$; $a_n \preccurlyeq b_n$ if $a_n = O(b_n)$. Let $1_{n-1}$ represent a vector of ones with $n - 1$ elements, and $0_{n-1}$ represent a vector of zeros with $n - 1$ elements. Let $\| \cdot \|_{\textup{op}}$ denote the operator norm. Define $E_{w_i}(\cdot)$ as the expectation taken over the marginal distribution of $w_i$. Let $X_{-i}$ represent the set $\{X_j\}_{j \neq i}$. The abbreviation “i.i.d.” stands for “independent and identically distributed.”

Setup

Framework and Notation

I consider an RCT with $n$ participants. For each participant $i\in\{1,\dots,n\}$, let $Y_i \in \mathbb{R}$ denote the observed outcome of interest, $T_i \in \{0,1\}$ denote the treatment assignment where $T_i \overset{}{} \sim \text{Bernoulli}(\pi)$ for some $0<\pi<1$, and $w_i$ denote the unobserved covariates. I consider a single large network, allowing for an arbitrary form of interference within it. The network structure is represented by an adjacency matrix $A = \{A_{ij}\}_{i,j=1}^n$, where the $(i,j)$th entry $A_{ij} \in \{0,1\}$ indicates whether units $i$ and $j$ are connected. The adjacency matrix is assumed to be undirected (symmetric), unweighted (binary values), and has no self-links ($A_{ii} = 0$). I differentiate between two network structures: $A^\text{pre}$, which represents the network observed before the intervention, and $A^\text{post}$, which corresponds to the network observed after the intervention. I assume the researchers observe both $A^\text{pre}$ and $A^\text{post}$.

Figure (ref) illustrates the causal mechanism considered in this paper.\footnote{I use the graph for illustration without using formal graphical language.} I assume that the post-intervention network $A^{\text{post}}$, influenced by the treatment vector, plays a critical role in determining the outcome. Consequently, the exposure mapping, also known as the effective treatment, which maps $\{T_i\}_{i=1}^n$ and $A^{\text{post}}$ to some low-dimensional statistics, serves as a mediator through which the treatment indirectly affects the outcome. I denote this mediator as $M_i$. Specifically, I assume that $T_i$ has both a direct effect on $Y_i$ and an indirect effect through changing network mediator $M_i$. The network $A^{\text{post}}$ is shaped by both the treatments and the latent variable $w_i$, which may also influence the outcome. As a result, $w_i$ acts as a confounder between the mediator and the outcome. The treatments of other units $T_{-i}$ affect $Y_i$ exclusively through the network mediator $M_i$.

figure[figure omitted — 1,144 chars of source]

Mediators can take various forms. For example:

equation[equation omitted — 116 chars of source]

which measures the fraction of treated friends after the intervention. This mediator depends solely on the number of (treated) friends, regardless of their identity. It is a specific form of the anonymous interference assumption proposed by HudgensHalloran2008a, also referred to as the anonymous interactions assumption by Manski2013. Other examples of anonymous interference include:

enumerate[(1)] • $M_i = \sum_{j=1}^n A_{ij}^{\text{post}} T_j$, which measures the total number of treated friends after the intervention; • $M_i = 1\{ \sum_{j=1}^n A_{ij}^{\text{post}} T_j > 0\}$, which measures whether there is at least one treated friend after the intervention.

Potential Value Notation

To define the causal parameters of interest, I introduce the potential value notations for the mediator and the outcome, following the mediation analysis literature RobinsGreenland1992, Pearl2001, HeckmanPinto2015. First, I consider a hypothetical intervention on the treatment vector $t$ and define the potential values of the mediators for unit $i$ corresponding to this intervention as:

equation*[equation* omitted — 58 chars of source]

Note that potential mediator $M_i(t)$ depends on the treatment vector $t$ in two ways: it is a direct function of $t$ and the network structure $A^{\text{post}}$, which implicitly depends on $t$ as well. I then consider a hypothetical intervention on both $t_i$ and $m_i$. Define the potential outcomes corresponding to the interventions on $t_i$ and $m_i = M_i(t_i, t_{-i})$ for unit $i$ as:

equation*[equation* omitted — 86 chars of source]

where $\mathcal{M}$ contains all possible values of $m_i$. This notation reflects that the potential outcome $Y_i(t_i, m_i)$ depends not only on $t_i$, the treatment assignment of unit $i$, but also on the network changes induced by the intervention, as captured by the function $M_i$. Additionally, it indicates that the treatment assignments of others influence the potential outcome of individual $i$ only indirectly through the mediator $M_i$. I also define the nested potential outcomes corresponding to an intervention on $t_i$ and $m_i = M_i(t_i', t_{-i})$ as:

equation*[equation* omitted — 123 chars of source]

The notation $Y_i(t_i, M_i(t_i', t_{-i}))$ represents the hypothetical outcome when the treatment is set to level $t_i$ and the mediator is set to its potential value $M_i(t_i', t_{-i})$, corresponding to the treatments $t_i'$ for unit $i$ and $t_{-i}$ for the other units. I allow $t_i$ and $t_i'$ to differ to define counterfactual outcomes where either the treatment level or the mediator level changes, while the other remains fixed. This approach enables the separation of the two channels through which the treatment affects the outcome, either directly or indirectly. The observed values of the mediator $M_i$ and the outcome $Y_i$ are related to their potential values as follows: \[ M_i = M_i(t_i, t_{-i}) \text{ if } T_i = t_i \text{ and } T_{-i} = t_{-i}, \] and \[ Y_i = Y_i(t_i, m_i) \text{ if } T_i = t_i \text{ and } M_i = m_i. \]

remarkThe interference literature typically defines potential outcomes as a function of the treatment vector HudgensHalloran2008a: \begin{equation*} \left\{ Y_i(t_i, t_{-i}): t_i \in \{0,1\}; t_{-i} \in \{0,1\}^{n-1} \right\}, \end{equation*} which does not account for how an individual's treatment affects their exposure to others in the network, thereby indirectly influencing the outcome. Specifically, it cannot capture the scenario where the treatment or mediator levels change while the other remains fixed, i.e., when $t_i \ne t_i'$. The expression $Y_i(t_i, M_i(t_i', t_{-i}))$ nests $Y_i(t_i, t_{-i})$ as a special case when $t_i = t_i'$ and $M_i = t_{-i}$. In the causal graph shown in Figure (ref), the arrow from $T_i$ to $M_i$ would be omitted if the treatment does not affect the network.

Data-Generating Process for Networks

I specify that the relationship between unit $i$ and $j$ before and after the intervention is determined by the following random graph models:

align[align omitted — 279 chars of source]

where $\{w_i\}_{i=1}^n$ are the i.i.d. latent variables and $\{\eta_{ij}\}_{i,j=1}^n$ is a symmetric matrix of unobserved scalar disturbances with upper diagonal entries that are mutually independent. I allow the sparsity parameters, $q_n^\text{pre}$ and $q_n^\text{post}$, to differ after the intervention for generality. The pre-intervention graphon, $g^\text{pre}$, depends on the latent variable of the pairs, $w_i$ and $w_j$, while the post-intervention graphon, $g^\text{post}$, also incorporates the treatment assignments $T_i$ and $T_j$ in forming links. Since the idiosyncratic term $\eta_{ij}$ and the latent variable $w_i$ are not time-varying, the framework to account for cases where the network remains unchanged after the intervention, with $g^\text{pre} = g^\text{post}$.\footnote{The assumption that $\eta_{ij}$ is not time-varying can be relaxed. Specifically, $\eta_{ij}$ in the pre- and post-intervention networks can be treated as independent draws from the same underlying distribution or as draws from correlated distributions. The conjecture is that the SSIV constructed based on $A^{\text{pre}}$ may be weaker, as additional exogenous shocks reduce its relevance.} Changes in the network are primarily attributed to the intervention, while exogenous shocks unrelated to the intervention are reflected in the difference between $g^\text{pre}$ and $g^\text{post}$. Thus, the network formation described in (ref) and (ref) captures the features that network formation is driven by latent variables and may evolve over time, with changes induced by the intervention.

remarkThe latent variable $w_i$ may be either a scalar or a vector, with the key requirement being a common or correlated component that influences both the pre- and post-intervention networks. This relationship ensures that the pre-intervention network is predictive of the post-intervention network, allowing it to serve as an instrument.

The distribution of \(\eta_{ij}\) is not separately identified from the graphons and sparsity parameters, so it is typically normalized to follow a standard uniform distribution. As a result, $q_n^{\diamond} g^{\diamond}$ represents the probability that a given pair is connected, for $\diamond \in \{\textup{pre}, \textup{post}\}$. The graphon-based network model in (ref) and (ref), which takes the form of an inhomogeneous Erdős-Rényi graph, is motivated by Aldous-Hoover Theorem on exchangeable arrays Aldous1981, LovaszSzegedy2006,BickelChen2009 and has recently gained considerable attention in econometrics GaoLuZhou2015, ZhangLevinaZhu2017, Graham2020b, PariseOzdaglar2023. Recent studies have also employed this model to account for randomness in network formation; see Auerbach2022a, Cai2022, and LiWager2022. The sparsity parameters $q_n^{\diamond}$, for $\diamond \in \{\textup{pre}, \textup{post}\}$, serve as theoretical tools, (possibly) driving the probability of any pair being connected to zero as $n \to \infty $. The goal of this paper is to derive theoretical results across a broad range of sparsity rates, encompassing the following three cases:

enumerate[(a)] • $q_n^{\diamond} \asymp n^{-1}$ (bounded degree graph): the total number of friends for each unit remains bounded and not vanishing in expectation; • $\lim_{n\to\infty} q_n^{\diamond} = 0$ and $\lim_{n\to\infty} n q_n^{\diamond} = \infty$ (sparse graph): the probability of any pair forming a connection decreases as the sample size increases, while the total number of friends for each unit would increase in expectation; • $q_n^{\diamond} \asymp 1$ (dense graph): the total number of friends for each unit is approximately of order $n$ in expectation, growing at the same rate as the sample size.

I summarize the assumption on the network models below.

assumptionThe latent variables $\{w_i\}_{i=1}^n$ are i.i.d. and independent of $\{\eta_{ij}\}_{i,j=1}^n$, where $\eta_{ij} \overset{\text{i.i.d.}}{\sim } U[0,1]$ for $j>i$ and $\eta_{ij} = \eta_{ji}$. The networks are randomly generated according to equations (ref) and (ref), where $g^\text{pre}$ is a symmetric measurable function of $w_i$ and $w_j$, and $g^\text{post}$ is a symmetric measurable function of $(T_i, w_i)$ and $(T_j, w_j)$, both mapping into $[0,1]$. The sparsity parameters satisfy $n^{-1} \preccurlyeq q_n^{\diamond} \preccurlyeq 1$, for $\diamond \in \{\textup{pre}, \textup{post}\}$.

Parameters of Interest and Identification

Now I define the causal parameters of interest in my setting.

definition\begin{enumerate}[(1)] For $t\in \{0,1\}$, define • the average total effect $\textup{ToE}$ of unit $i$'s treatment as \begin{equation} ToE = E\left[Y_i\left(1, M_i\left(1, T_{-i}\right) \right) - Y_i\left(0, M_i\left(0, T_{-i}\right) \right)\right]; \end{equation} • the average direct effect $\textup{DE}(t)$ of unit $i$'s treatment as \begin{equation} DE(t) = E\left[Y_i\left(1, M_i\left(t, T_{-i}\right) \right)-Y_i\left(0, M_i\left(t, T_{-i}\right) \right)\right]; \end{equation} • the average indirect effect $\textup{IE}(t)$ of unit $i$'s treatment as \begin{equation} IE(t) = E\left[Y_i\left(t, M_i\left(1, T_{-i}\right) \right) - Y_i\left(t, M_i\left(0, T_{-i}\right) \right)\right]; \end{equation} • the average \textit{spillover effect} $\textup{SE}(t)$ as \begin{equation} \textup{SE}(t) = E\left[Y_i\left(t, M_i\left(t_i = t, t_{-i} = 1_{n-1} \right) \right) - Y_i\left(t, M_i\left(t_i = t, t_{-i} = 0_{n-1} \right) \right)\right]. \end{equation} \end{enumerate}

The average total effect $\textup{ToE}$ measures the average overall effect of changing the treatment of unit $i$ on that unit's outcome. I then decompose the $\textup{ToE}$ into two channels: one without changes in the mediator and the other through changes in the mediator. The average direct effect $\textup{DE}(t)$ captures the former, measuring the average effect of changing the treatment of unit $i$, while keeping the mediator at the level it would have taken if $T_i$ had been set to $t$. The average indirect effect $\textup{IE}(t)$ captures the latter, measuring the average effect of changing the treatment of unit $i$ solely through its impact on the mediator while keeping the intervention on unit $i$ at level $t$. The average spillover effect $\textup{SE}(t)$ measures the effect on unit $i$'s outcome of assigning everyone else to the treatment group versus the control group, while holding unit $i$'s treatment status fixed at level $t$.\footnote{The $\textup{SE}(t)$ in (ref) is one way to define spillover effects. Alternatively, spillover effects can be defined as the change in unit $i$'s outcome resulting from differences between any two distinct intervention programs through variations in the treatment of others.} In defining these causal effects, the expectation is taken over all sources of randomness, including potential outcomes and the treatment assignments of others.

remarkThe existing literature on interference typically defines the direct and indirect effects of a binary treatment, based on the potential outcome notation in Remark (ref). For example, HuLiWager2022b defined the average direct effect of a binary treatment as: \[ \tau_{\textsc{ade}} = \frac{1}{n} \sum_{i=1}^n {E} \left[ Y_i(t_i = 1, T_{-i}) - Y_i(t_i = 0, T_{-i}) \right], \] and the average indirect effect as: \begin{equation} \tau_{aie} = \frac{1}{n} \sum_{i=1}^n \sum_{j \neq i} {E}\left[ Y_j(t_i = 1, T_{-i}) - Y_j(t_i = 0, T_{-i}) \right]. \end{equation} The estimand $\tau_{\textsc{ade}}$ corresponds to the $\textup{ToE}$ in (ref). To better understand the intervention mechanism, I further decompose the effect of $T_i$ on $Y_i$ into the $\textup{DE}(t)$ in (ref) and the $\textup{IE}(t)$ in (ref), which operates through changes in the mediator.

The mediation formula relies on the following assumption.

assumptionFor all $i=1,\cdots,n$, \begin{enumerate}[{(a)}] • $ (T_i, T_{-i}) \perp \!\!\! \perp Y_{i}\left(t_i, m_i \right)$ for all $t_i \in \{0,1\}$ and $m_i \in \mathcal{M}$; • $ (T_i, T_{-i}) \perp \!\!\! \perp M_i(t_i, t_{-i}) $ for all $t_i \in \{0,1\}$ and $t_{-i} \in \{0,1\}^{n-1}$; • $M_i \perp \!\!\! \perp Y_i(t_i, m_i) \mid w_i$ for all $t_i \in \{0,1\}$ and $m_i \in \mathcal{M}$; • $M_i(t_i', t_{-i}) \perp \!\!\! \perp Y_{i}\left(t_i, m_i \right) \mid w_i$ for all $t_i\in \{0,1\}$, $t_i'\in \{0,1\}$, $t_{-i} \in \{0,1\}^{n-1}$ and $m_i \in \mathcal{M}$. \end{enumerate}

Assumption (ref) is standard in the literature of mediation analysis Pearl2001. Assumptions (ref)(a)-(b) assume no treatment-outcome confounding and no treatment-mediator confounding, respectively. Assumptions (ref)(a)-(b) hold under experiments with randomized treatment. Assumption (ref)(c) assumes that the variable $w_i$ captures all the confounding factors between the mediator and the outcome. Assumption (ref)(d) assumes the cross-world independence between the potential outcomes and potential mediators. However, Assumption (ref)(d) can never be validated, as it is impossible to observe both $M_i(t_i', t_{-i})$ and $Y_{i}\left(t_i, m_i \right)$ in any experiment if $t_i \ne t_i'$.

Theorem (ref) expresses the parameters of interest in terms of the observed data up to the unknown distribution of $w_i$.

theoremUnder Assumption (ref), \begin{enumerate}[(1)] • $\textup{ToE} = E(Y_i \mid T_i=1) - E(Y_i \mid T_i=0)$; • $\textup{DE}(t) = E_{w_i} \left\{ E\left[ E\left(Y_i \mid M_i, T_i = 1, w_i \right) - E\left(Y_i \mid M_i, T_i = 0, w_i \right) \mid T_i = t, w_i \right] \right\}$; • $\textup{IE}(t) = E_{w_i} \left\{ \begin{array}{c} E\left[ E\left(Y_i \mid M_i, T_i = t, w_i \right) \mid T_i = 1, w_i \right] \\ - E\left[ E\left(Y_i \mid M_i, T_i = t, w_i \right) \mid T_i = 0, w_i \right] \end{array}\right\} $; • $\textup{SE}(t) = E_{w_i} \left\{ \begin{array}{c} E\left[ E\left(Y_i \mid M_i, T_i = t, w_i \right) \mid T_i = t, T_{-i} = 1_{n-1} , w_i \right] \\ - E\left[ E\left(Y_i \mid M_i, T_i = t, w_i \right) \mid T_i = t, T_{-i} = 0_{n-1}, w_i \right] \end{array} \right\}$. \end{enumerate}

If the parameter of interest is solely $\textup{ToE}$, it can be identified, and consistent estimators can be obtained by regressing the outcome on the treatment indicator with an intercept, or equivalently, by using the difference-in-means estimator, provided the treatment is randomly assigned. However, if the focus is on distinguishing between the direct and indirect effects, or when the spillover effect is of interest, these causal effects are identified only up to the unknown distribution of $w_i$.

Following the classical Baron-Kenny method BaronKenny1986, I assume that the potential outcome is linear in the treatment and mediator, and additive with respect to any form of $w_i$, as stated in Assumption (ref), which implies Assumptions (ref)(c)-(d). However, I relax the assumption that the mediator $M_i$ is a linear function of $T_i$ and $w_i$.

assumptionAssume the following (partially) linear model for the potential outcome: \begin{equation} Y_i(t_i, m_i) = \beta_0 + \beta_1 t_i + \beta_2 m_i + \lambda(w_i) + \varepsilon_i, \end{equation} where $\lambda(\cdot)$ is an unknown measurable function with $E(\lambda(w_i)) = 0$. Additionally, assume that $\{\varepsilon_i\}_{i=1}^n$ are i.i.d., $E(\varepsilon_i^2) < \infty$, and $E(\varepsilon_i \mid T_i, T_{-i}, A^{\text{post}}) = 0$.

Under Assumption (ref), the observed outcome also follows the linear model: \[ Y_i = \beta_0 + \beta_1 T_i + \beta_2 M_i + \lambda(w_i) + \varepsilon_i. \] Let $u_i$ denote the error term in the outcome model: $u_i = \lambda(w_i) + \varepsilon_i$.

Although the framework in Section (ref) applies to general forms of mediators, Sections (ref) and (ref) focus on the estimation and inference with the mediator specified in (ref). I follow the convention of $0/0=0$. I focus on $M_i$ in (ref) for several reasons. First, different forms of mediators exhibit varying levels of dependency and require different proofs. The magnitude of $M_i$ in (ref) is normalized and remains bounded as networks become denser and sample sizes increase, ensuring the outcome does not diverge to infinity under a constant treatment effect. Second, with this fraction as the mediator, the linear outcome model introduced in Assumption (ref) becomes a special case of the linear-in-means model Manski2013. Incorporating endogenous peer effects is also of interest and is left for future work.

remarkAuerbach2022 investigates the identification and estimation of the following partially linear model, with the parameters of interest being $\beta$ and $\lambda(w_i)$: \begin{align} Y_i &= \beta^\top X_i + \lambda(w_i) + \varepsilon_i, \end{align} where $w_i$ is some unobserved latent variable that also drives the network formation: \[ A_{ij} = {1}\left\{\eta_{ij} \leq g(w_i, w_j)\right\} {1}\{j \neq i\}. \] Auerbach2022a focuses on the dense network, where the sparsity parameter equals $1$. The regressor $X_i$ in (ref) corresponds to $(1, T_i, M_i)$ in my setting. Auerbach2022a uses a matching approach based on network data $A_{ij}$ to identify and estimate $\beta$ and $\lambda(w_i)$. However, this method relies on dense networks and sufficient variation in the regressors after controlling for the latent variable. In my context, which involves exogenous shocks from the randomized treatment and does not prioritize identifying the latent variable part $\lambda(w_i)$, I tackle the endogeneity issue from the unobserved confounder using the IV method with SSIV.

Corollary (ref) simplifies the statement of Theorem (ref) under Assumption (ref), and provides the causal interpretation of the coefficients, which is analogous to the Baron-Kenny formulas for mediation BaronKenny1986.

corollaryUnder Assumption (ref), then \begin{enumerate}[(1)] • $\textup{DE}(1) = \textup{DE}(0) = \textup{DE} = \beta_1$; • $\textup{IE}(1) = \textup{IE}(0) = \textup{IE} = \beta_2 \cdot \left\{ E(M_i \mid T_i=1) - E(M_i \mid T_i=0) \right\}$; • $\textup{ToE} = \beta_1 + \beta_2 \cdot \left\{ E(M_i \mid T_i=1) - E(M_i \mid T_i=0)\right\}$; • $\textup{SE}(t) = \beta_2 \cdot \left\{ E(M_i \mid T_i=t, T_{-i} = 1_{n-1}) - E(M_i \mid T_i=t, T_{-i} = 0_{n-1})\right\}$. \end{enumerate}

Under Assumption (ref), where the treatment $T_i$ does not interact with the mediator $M_i$, $\textup{DE}(t)$ and $\textup{IE}(t)$ do not depend on the value of $t$. Furthermore, Corollary (ref) shows that $\beta_1$ captures the direct effect of the treatment, while the indirect and spillover effects depend on $\beta_2$, scaled by the magnitude of changes in the mediator in response to changes in $T_i$ and $T_{-i}$. When the post-intervention network $A^{\text{post}}$ is uncorrelated with the treatment, the mediator does not respond to changes in $T_i$, resulting in a zero indirect effect and making the total effect equivalent to the direct effect. However, the spillover effect would still be present. Additionally, with $M_i$ in (ref), $ E(M_i \mid T_i=t, T_{-i} = 1_{n-1}) - E(M_i \mid T_i=t, T_{-i} = 0_{n-1}) = 1$, indicating that $\beta_2$ captures the spillover effect $\textup{SE}(t)$ defined in (ref). Corollary (ref) justifies the emphasis on the estimation and inference of $\beta_1$ and $\beta_2$ in this paper.

To summarize, I account for two key features of networks when estimating $\beta$'s:

enumerate[(1)] • There exists some latent variable $w_i$ that influences both network formation, and consequently the mediator, as well as the outcome of interest. • Treatment assignments affect how units are exposed to others in the network, with the exposure mapping serving as a mediator for the treatment’s effect on the outcome.

As a result, even in a randomized setting, if the causal effects of interest involve the network, such as the $\textup{IE}(t)$ in (ref) and the $\textup{SE}(t)$ in (ref), endogeneity becomes a concern, potentially biasing the estimation of $\beta_2$. Given the correlation between $T_i$ and $M_i$, this bias in estimating $\beta_2$ can further affect the estimation of $\beta_1$.

Motivating Examples

I present three empirical examples illustrating how social network structures can endogenously evolve in response to interventions, potentially undermining their intended effects.

examplePrina2015 conducted an RCT to assess the impact of offering access to formal savings on households' financial situations in 19 villages in Nepal with 915 households. Half of the female household heads were randomly offered the savings accounts, while the other half were not. The financial links were measured by asking survey questions such as “who did you exchange loans or gifts with?” The network experienced considerable reshuffling, despite the total number of links remaining nearly constant (328 at baseline, 329 at endline). Of these, only 73 links persisted from baseline, while 255 links were broken after the intervention, and 256 new links were formed by the endline. Importantly, the distribution of treat-treat, control-control, and treat-control pairs among these link types was imbalanced. Of the persisting links, 26 were treat-treat pairs, 17 were control-control pairs, and 30 were treat-control pairs. Among the broken links, 68 were treat-treat, 74 were control-control, and 113 were treat-control. In contrast, among the newly formed links, 78 were treat–treat, 53 were control-control, and 125 were treat-control. I will revisit this study in Section (ref) as an empirical illustration.
exampleBanerjeeBrezaChandrasekhar2023 investigated the effect of introducing formal financial institutions on informal lending practices and social networks through studies in two distinct settings. In their first study, they analyzed the non-random introduction of microfinance in 43 out of 75 villages in Karnataka, India. They found that social networks contracted more in villages where microfinance was introduced. To validate these findings, they conducted an RCT in Hyderabad, India, where 52 out of 104 neighborhoods were randomly selected for microfinance introduction. Similar to the first setting, these neighborhoods also experienced a reduction in social connections, even among households initially unlikely to borrow. Together, these studies suggest that access to microfinance reduces the incentive to maintain and form social links, affecting both borrowers and non-borrowers alike.
exampleCarrellSacerdoteWest2013 examined the impact of group randomization on academic performance at the United States Air Force Academy, with the hypothesis that low-ability students would benefit from exposure to high-ability peers. Incoming freshmen were randomly assigned to treatment or control groups. The control group followed the usual distribution of abilities, while the treatment group paired low-ability students with high-ability peers and placed middle-ability students in separate squadrons. Contrary to expectations, the intervention had a significantly negative effect on the academic performance of low-ability students. While several factors could explain this outcome, the study found that the endogenous sorting of roommates, study partners, and friends evolved differently in the treatment group compared to the control group. Specifically, low-ability students in the treatment group saw a notable increase in the number of low-ability study partners and friends.

Examples (ref)-(ref) also highlight concerns about latent variables that influence network formation. Example (ref) demonstrates that transfers were linked to treatment households having more assets and greater financial inclusion at baseline, suggesting that households with more resources and higher socioeconomic status may also increase transfers to others. Example (ref) suggests that access to microfinance reduces borrowers' incentives to maintain connections, leading even non-participants to scale back efforts to form links. This is partly because individuals connected to potential borrowers see diminished returns from these relationships. Example (ref) indicates that the increase in low-ability study partners and friends was not merely due to the higher proportion of low-ability peers in the treatment group but also reflected a pattern of homophily, as students gravitated toward others with similar abilities.

OLS estimation

In this section, I begin with discussing the properties of $M_i$ and the conditions under which endogeneity arises in Section (ref). In Section (ref), I discuss the bias resulting from the common practice in empirical studies of collecting and using only pre-intervention network data. When no endogeneity arises, I explore the asymptotic properties of OLS estimators in Section (ref).

Property of $M_i$

Combined with the form of mediator in (ref), the outcome model of interest is \[ Y_i = \beta_0 + \beta_1 T_i + \beta_2 \frac{\sum_{j=1}^n A_{ij}^\text{post} T_j }{\sum_{j=1}^n A_{ij}^\text{post} } + \lambda(w_i) + \varepsilon_i. \] In addition to assuming anonymous interference, this model assumes “no second- or higher-order effects,” meaning the treatment of units beyond direct connections has no impact.

The properties of the estimator depend critically on the characteristics of $M_i$. The regressor $M_i$ is dependent across units, and its fractional form raises natural concerns about whether it provides sufficient variation, potentially leading to (near) multicollinearity issues. To better understand the properties of $M_i$, I decompose it as $M_i = \xi_i + r^*_i$, where $\xi_i$ is defined as the ratio of the conditional expectations of the numerator and denominator of $M_i$ in (ref):

equation[equation omitted — 190 chars of source]

and $r^*_i$ is some remainder term. Note that $\xi_i$ represents the conditional probability that a neighbor of individual $i$ (i.e., a connected individual $j$) is treated, given $i$'s own characteristics $w_i$ and treatment status $T_i$. Throughout the paper, I present the results based on the following two cases, depending on whether $\xi_i$ is degenerate or not:

enumerate[(a)] • $\xi_i$ is an i.i.d. variable across $i$ with constant variance, implying $\operatorname{Var}(\xi_i)>0$;\footnote{The term $\xi_i$ does not depend on $n$ because the sparsity rate $q_n^\text{post}$ cancels out between the numerator and denominator.} • $\xi_i = \pi$, implying $\operatorname{Var}(\xi_i)=0$.

Case (b) occurs when $A_{ij}^\text{post}$ is mean independent of $T_j$ given $T_i$ and $ w_i$, i.e., $E(T_j \mid A_{ij}^\text{post}, T_i, w_i) = E(T_j \mid T_i, w_i) = \pi$. This case encompasses scenarios where the network remains unchanged after the intervention (i.e., $A^{\text{pre}} = A^{\text{post}}$), undergoes exogenous changes unrelated to the treatment, or when $A_{ij}^{\text{post}}$ depends on $T_j$ but is mean-independent. It has two important implications for discussion. First, in Case (b), no endogeneity issue arises, even with the presence of an unobserved confounder: \[ E\left[ M_i u_i \right] = E\left[ \frac{\sum_{j=1}^n A_{ij}^{\text{post}} T_j}{\sum_{j=1}^n A_{ij}^{\text{post}}} u_i \right] = \pi E\left[ \frac{\sum_{j=1}^n A_{ij}^{\text{post}}}{\sum_{j=1}^n A_{ij}^{\text{post}}} u_i \right] = 0. \] The equality holds because $A_{ij}^\text{post}$ is conditionally mean independent of $T_j$, and Assumption (ref) implies that $A_{ik}^\text{post}$ is independent of $T_j$ for any $k \neq j$. This property, resulting from the row normalization of the mediator, is also documented by DieyeDjebbariBarrera-Osorio2014 and applied in the construction of the normalized SSIV Card2009, AutorDornHanson2013, PeriShihSparber2015, AdaoKolesarMorales2019, Goldsmith-PinkhamSorkinSwift2020, BorusyakHullJaravel2022. Second, the consistency condition and convergence rate of OLS estimators depend on whether $\xi_i$ is degenerate. Treatment-induced network changes introduce greater variation in the fraction mediator, potentially improving the convergence rate. By comparing the results of Cases (a) and (b), I evaluate how the network's response to the intervention influences inference.

Bias from OLS estimation with pre-intervention network

In empirical studies, researchers often assume the network remains unchanged and use only pre-intervention data for estimation. Alternatively, even when acknowledging possible network changes, they avoid post-intervention data due to endogeneity concerns and rely on pre-intervention data to control for unobserved confounders CarterLaajajYang2021a. However, using $A^\text{pre}$ instead of $A^\text{post}$ in the regression model introduces omitted variable bias (OVB) and impedes the identification and consistent estimation of direct, indirect, and spillover effects.

Define ${M}^\text{pre}_i$ as the mediator calculated using $A^\text{pre}$, i.e., ${M}^\text{pre}_i = \sum_{j=1}^n A_{ij}^\text{pre} T_j / \sum_{j=1}^n A_{ij}^\text{pre}$. Researchers typically estimate a linear regression with the following regressors: \[ {X}^\text{pre}_i = (1, T_i, {M}^\text{pre}_i). \] Let $\beta^\text{pre}$ denote the vector of population coefficients from the above OLS fit. By the argument of OVB,

align*[align* omitted — 254 chars of source]

This implies that if the mediator $M_i$ is uncorrelated with the treatment $T_i$, there is no OVB in $\beta_1^\text{pre}$. However, when a correlation exists between $T_i$ and $M_i$, $\beta_1^\text{pre}$ captures the total effect of $T_i$ on $Y_i$, encompassing both the direct effect and the indirect effect through changes in the mediator. The coefficient $\beta_2^\text{pre}$ represents the attenuated indirect effect, which remains uncontaminated by $\beta_1$. If researchers are only interested in the total effect $\textup{ToE}$, fitting an OLS regression with ${X}^\text{pre}_i$ using the pre-intervention network can still yield the desired result. However, this approach cannot distinguish between direct and indirect effect channels and fails to recover the spillover effect.

Asymptotic properties of OLS estimators

As discussed in Section (ref), no endogeneity arises when either $\operatorname{Var}\left( \xi_i \right) > 0$ with $\lambda(w_i) = 0$, or $\operatorname{Var}\left( \xi_i \right) = 0$. In such cases, I propose using OLS estimation with the regressor vector $X_i$ to estimate $\beta$: \[ X_i = (1, T_i, M_i). \] Let $\hat{\beta}^{\textsc{ols}}$ denote the vector of coefficients obtained from the above OLS fit.

theoremUnder Assumptions (ref) and (ref), \begin{enumerate}[(a)] • if $\operatorname{Var}\left( \xi_i \right) > 0$ and $\lambda(w_i)=0$, then $\hat{\beta}^{\textsc{ols}} - \beta = O_{\mathbb{P}}(n^{-1/2})$; • if $\operatorname{Var}\left( \xi_i \right) = 0$, then $\hat{\beta}^{\textsc{ols}}_0 - \beta_0 = O_{\mathbb{P}}\left(\sqrt{q_n^\text{post}}\right)$, $\hat{\beta}^{\textsc{ols}}_1 - \beta_1 = O_{\mathbb{P}}(n^{-1/2})$ and $\hat{\beta}^{\textsc{ols}}_2 - \beta_2 = O_{\mathbb{P}}\left(\sqrt{q_n^\text{post}}\right)$. \end{enumerate}

Theorem (ref) establishes the consistency conditions of $\hat{\beta}^{\textsc{ols}}$ in the absence of endogeneity under Cases (a) and (b), respectively. Under Case (a), the convergence rate is the standard $\sqrt{n}$ and does not depend on the sparsity rate $q_n^\text{post}$. The intuition is that, by the decomposition, $\xi_i$ represents the key term in $M_i$. When $\xi_i$ is non-degenerate and thus i.i.d., the estimation reduces to the standard case with i.i.d. data. Under Case (b), when $\xi_i$ is degenerate, the variation of $M_i$ comes from the remainder term $r_i^*$. As the network becomes denser, the dependence among units increases, and the variation of $M_i$ across units decreases, which slows the convergence rate. This explains why $\hat{\beta}^{\textsc{ols}}_0$ and $\hat{\beta}^{\textsc{ols}}_2$ are consistent only when the network is not dense (i.e., $q_n^\text{post} \prec 1$), with a (possibly) slower convergence rate than the usual rate $\sqrt{n}$, scaled by $\sqrt{ n q_n^\text{post}}$. The convergence rate of $\hat{\beta}^{\textsc{ols}}_1$ remains the standard $\sqrt{n}$, since $T_i$ is uncorrelated with the other regressor $M_i$ and thus unaffected by the dependency of $M_i$ across $i$, as in the i.i.d. setting.

Next, I study the asymptotic distribution of $\hat{\beta}^{\textsc{ols}}$, focusing on the regimes where it is consistent. Define $\hat{u}_i^\textsc{ols}$ as the residual from the above OLS fit and stack $X_i$ to form the $n\times 3$ design matrix $X$. The variance estimator $\hat{V}^{\textsc{ols}}$ is the standard heteroskedasticity-consistent (HC) variance estimator, defined as:

equation[equation omitted — 129 chars of source]

where the middle term, $\hat{V}^{\textsc{ols}}_{\text{num}}$, is given by \[ \hat{V}^{\textsc{ols}}_{\text{num}} = X^\top \text{diag}\left\{(\hat{u}_i^{\textsc{ols}})^2, i = 1, \dots, n\right\} X. \] Theorem (ref) focuses on the regimes of $q_n^{\text{post}}$ under which the OLS estimators are consistent.

theoremSuppose either $\operatorname{Var}\left( \xi_i \right) > 0$ with $\lambda(w_i)=0$ or $\operatorname{Var}\left( \xi_i \right) = 0$ with $q_n^{\text{post}} \prec 1$. Under Assumptions (ref) and (ref), then \[ \left( \hat{V}^{\textsc{ols}} \right)^{-1/2} \left(\hat{\beta}^{\textsc{ols}} - \beta \right) \stackrel{d}{\rightarrow} \mathcal{N}(0, I_3). \]

Theorem (ref) establishes the asymptotic normality of $\hat{\beta}^{\textsc{ols}}$ and shows that $\hat{V}^{\textsc{ols}}$ closely approximates the asymptotic variance. There are two key implications. First, despite the dependence of the regressor $M_i$ across units, the usual HC variance estimator $\hat{V}^{\textsc{ols}}$ in (ref) remains valid since the error term $u_i$ is assumed to be i.i.d. Second, the normal approximation exhibits a self-normalization property, ensuring that inference based on standard $t$-tests is reliable. Although the convergence rate of $\hat{\beta}^{\textsc{ols}}$ (potentially) depends on the sparsity parameter $q_n^{\text{post}}$, its asymptotic normality is unaffected by the sparsity level, as the varying order of $\hat{\beta}^{\textsc{ols}}$ is “canceled out” by the order of the variance component. This result is crucial because the sparsity parameter $q_n^{\text{post}}$ is generally not identified BickelChenLevina2011. A similar self-normalization property is observed by Cai2022 in the context of network centrality regression and by HansenLee2019 in the context of cluster-dependent data.

IV estimation

In this section, I examine the IV estimation to address the potential endogeneity issue when $\lambda(w_i)\ne 0$. I study the asymptotic properties of IV estimators using SSIV in Section (ref), showing that they become inconsistent as networks grow denser. To address this, in Section (ref), I propose a modification to SSIV to restore consistency in denser networks.

Asymptotic properties of SSIV estimators

Endogeneity issue (potentially) arises when $\lambda(w_i) \ne 0$. To address this, one needs instruments that are uncorrelated with $u_i$ and shifts $M_i$ enough to identify $\beta_2$. I tackle this issue using the shift-share instrumental variable approach. BorusyakHull2023 proposed a general approach for constructing SSIVs by combining exogenous shocks with non-exogenous exposure through a known formula, adjusting for expected treatment. Applying this approach to my setting, I derive the following linear SSIV for $M_i$: \[ Z_i^{\textsc{ssiv}} = \sum_{j=1}^n A_{ij}^\text{pre} (T_j - \pi), \] which leverages the treatment assignment of others, $T_j$, as the exogenous shock, weighted by pre-intervention network information as the share.

To estimate $\beta$, I propose using IV estimation with the IV vector ${Z}_i$:

equation*[equation* omitted — 91 chars of source]

Stack ${Z}_i$ to obtain the $n\times 3$ matrix ${Z}$. Let $\hat{\beta}^{\textsc{iv}}$ denote the vector of the coefficients obtained from the IV fits with ${Z}_i$.

I assess the validity of SSIV, and thus the identification of $\beta$, by verifying two conditions: exogeneity and relevance. The exogeneity of SSIV holds by construction, due to the random assignment of treatments and the centering of treatment around the assignment probability:

align*[align* omitted — 193 chars of source]

The question now turns to the relevance of the SSIV. BorusyakHull2023 assume weak mutual dependence of the SSIV to ensure the convergence of the sample first stage, which, in my model, roughly implies that the network is not too dense. I thoroughly analyze the relevance condition in my context and characterize the sparsity regimes under which the SSIV provides consistent estimators, as stated in Theorem (ref). In line with BorusyakHull2023, I demonstrate that consistency is achieved when both pre- and post-treatment networks are relatively sparse.

theoremUnder Assumptions (ref) and (ref), \begin{enumerate}[(a)] • if $\operatorname{Var}\left( \xi_i \right) > 0$ with $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec n^{-1/2}$, then \[ \hat{\beta}^{\textsc{iv}} - \beta = O_{\mathbb{P}}\left( \sqrt{n} \max\{q_n^\text{pre}, q_n^\text{post}\} \right); \] • if $\operatorname{Var}\left( \xi_i \right) = 0$ with $ \max\{q_n^\text{pre}, q_n^\text{post} \} \prec \sqrt{n} q_n^\text{post}$, then \begin{align*} \hat{\beta}^{iv}_1 - {\beta}_1 = & O_{\mathbb{P}}\left( \frac{1}{\sqrt{n}} \max\left\{ \frac{ \max\{q_n^pre, q_n^post \} }{ \sqrt{ q_n^post} }, 1 \right\} \right). \end{align*} Moreover, with $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec n^{-1/2}$, then \begin{align*} \hat{\beta}^{iv}_0 - {\beta}_0 = & O_{\mathbb{P}}\left( \sqrt{n} \max\{q_n^pre, q_n^\text{post}\} \right) \text{ and } \hat{\beta}^{\textsc{iv}}_2 - {\beta}_2 = O_{\mathbb{P}}\left( \sqrt{n} \max\{q_n^\text{pre}, q_n^\text{post}\} \right). \end{align*} \end{enumerate}

Theorem (ref) specifies the regimes under which $\hat{\beta}^{\textsc{iv}}$ is consistent for Cases (a) and (b). The consistency regime and convergence rates of $\hat{\beta}^{\textsc{iv}}_0$ and $\hat{\beta}^{\textsc{iv}}_2$ are invariant across these cases, both depending on $q_n^\text{pre}$ and $q_n^\text{post}$, and are (possibly) slower than the usual rate $\sqrt{n}$, scaled by the degree in the denser network, $n \max\{q_n^\text{pre}, q_n^\text{post}\}$. In Case (b), $\hat{\beta}^{\textsc{iv}}_1$ converges at a (possibly) faster rate compared to Case (a) and is consistent under a less restrictive condition. This observation aligns with Theorem (ref), which shows that, in Case (b), the estimation of $\beta_1$ is less affected by the dependency of $M_i$ across $i$ and exhibits a faster convergence rate than the estimator of $\beta_2$.

Theorem (ref) demonstrates that the consistency of the SSIV estimators breaks down when networks are relatively dense, i.e., $\max\{q_n^\text{pre}, q_n^\text{post}\} \succcurlyeq n^{-1/2}$. Here is the intuition: the SSIV method leverages variation in the number of treated friends for each individual, driven by random fluctuations in treatment assignment. Incorporating network information is essential, as the endogenous variable $M_i$ depends on the network structure. However, as the network becomes denser, the variation in the number of treated friends across individuals decreases. Consider an extreme case where all individuals are fully connected before the intervention. In this scenario, the SSIV reduces to $Z_i^{\textsc{ssiv}} = \sum_{j \neq i}(T_j - \pi)$, where the only distinction across units is whether they are treated or not.

remarkA commonly used instrument in the peer effects literature is the “peer-of-peer” IV BramoulleDjebbariFortin2009b, DeGiorgiFrederiksenPistaferri2020. When peers of peers are not direct peers, their characteristics influence individual outcomes only through their effect on peers' outcomes, providing valid instruments. Identification also requires a non-overlapping peer network, consistent with the findings on SSIV: the network must not be too dense and should exhibit sufficient structural variation.
remarkIn Appendix (ref), I discuss the performance of the IV estimators using the normalized SSIV suggested in BorusyakHullJaravel2022: $\sum_{j=1}^n A_{ij}^\text{pre}T_j / \sum_{j=1}^n A_{ij}^\text{pre}$. Instead of centering the treatment around the assignment probability, this approach normalizes the total number of treated friends to guarantee the exogeneity. However, as the network becomes denser, this ratio converges to the assignment probability as the sample size increases. This convergence can lead to a weak IV problem if the assignment probability is constant. See the consistency condition of the IV estimators using this normalized SSIV in Theorem (ref).
remarkThe BLP instrument is widely used in empirical industrial organization to address endogeneity in demand estimation by leveraging the characteristics of competing products as instruments BerryLevinsohnPakes1995, Armstrong2016. Both the BLP instrument and the shift-share instrument share the fundamental idea of leveraging variation from external, exogenous sources to identify causal effects. However, both the BLP instrument and the normalized SSIV face challenges in maintaining identifying power in limiting cases. The BLP instrument has limited identifying power asymptotically in large markets, as market shares approach zero and the dependence of equilibrium markups on other products' characteristics diminishes. Similarly, the normalized SSIV, $\sum_{j=1}^n A_{ij}^\text{pre} T_j / \sum_{j=1}^n A_{ij}^\text{pre}$, becomes degenerate in denser networks, weakening its relevance.

Corollary (ref) simplifies the results in Theorem (ref) to the special case where $q_n^\text{pre} \preccurlyeq q_n^\text{post}$.

corollarySuppose $q_n^\text{pre} \preccurlyeq q_n^\text{post}$. Under Assumptions (ref) and (ref), \begin{enumerate}[(a)] • if $\operatorname{Var}\left( \xi_i \right) > 0$ with $q_n^\text{post} \prec n^{-1/2}$, then $\hat{\beta}^{\textsc{iv}} - \beta = O_{\mathbb{P}}(\sqrt{n} q_n^\text{post})$; • if $\operatorname{Var}\left( \xi_i \right) = 0$ it holds that $\hat{\beta}^{\textsc{iv}}_1 - {\beta}_1 = O_{\mathbb{P}}\left( \frac{1}{\sqrt{n}} \right)$. Moreover, with $q_n^\text{post} \prec n^{-1/2}$, then $\hat{\beta}^{\textsc{iv}}_0 - {\beta}_0 = O_{\mathbb{P}}\left( \sqrt{n} q_n^\text{post} \right)$ and $\hat{\beta}^{\textsc{iv}}_2 - {\beta}_2 = O_{\mathbb{P}}\left( \sqrt{n} q_n^\text{post} \right)$. \end{enumerate}

Corollary (ref) shows that in both Cases (a) and (b), $\hat{\beta}_0^{\textsc{iv}}$ and $\hat{\beta}_2^{\textsc{iv}}$ exhibit the same convergence rate, which depends on $q_n^\text{post}$ and is (possibly) slower than the usual rate $\sqrt{n}$. In Case (a), the convergence rate of $\hat{\beta}_1^{\textsc{iv}}$ matches that of $\hat{\beta}_0^{\textsc{iv}}$ and $\hat{\beta}_2^{\textsc{iv}}$. However, in Case (b), where $A^{\text{post}}$ is conditionally mean-independent of the treatment, $\hat{\beta}^{\textsc{iv}}_1$ is less affected by the dependence of $M_i$ across individuals and retains the usual rate $\sqrt{n}$.

Now I show the asymptotic normality of the IV estimators $\hat{\beta}^{\textsc{iv}}$. I focus on the regimes where the IV estimators are consistent. Define $V^{\textsc{iv}}_{\text{num}} = \operatorname{Var}\left( \sum_{i=1}^n {Z}_i u_i \right)$, the variance of the numerator of the centered estimator $\hat{\beta}^{\textsc{iv}} - \beta$. It can be shown that

align[align omitted — 509 chars of source]

The $(2,3)$, $(3,2)$ and $(3,3)$ elements of $ V^{\textsc{iv}}_{\text{num}}$ in (ref) account for the dependence of ${Z}_i^{\textsc{ssiv}}$ across $i$, even though $u_i$ is i.i.d.. Let $\hat{u}_i^{\textsc{iv}}$ denote the residual from the IV fits with ${Z}_i$. Define $\hat{V}^{\textsc{iv}}_{\text{num}}$ as the plug-in estimator of $V^{\textsc{iv}}_{\text{num}}$:

align[align omitted — 673 chars of source]

Define $\hat{V}^{\textsc{iv}} = ({Z}^\top X)^{-1} \hat{V}_{\text{num}}^{\textsc{iv}} (X^\top {Z})^{-1}$ as the estimator of the asymptotic variance of $\hat{\beta}^{\textsc{iv}}$. I demonstrate the asymptotic normality of $\hat{\beta}^{\textsc{iv}}$, specifically focusing on the regimes where the IV estimators are consistent.

theoremUnder Assumptions (ref) and (ref), and with $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec n^{-1/2}$, then \[ \left(\hat{V}^{\textsc{iv}}\right)^{-1/2} \left( \hat{\beta}^{\textsc{iv}} - \beta \right) \stackrel{d}{\rightarrow} \mathcal{N}(0, I_3). \]

Theorem (ref) establishes the asymptotic normality of the IV estimators $\hat{\beta}^{\textsc{iv}}$ using SSIV, along with a consistent variance estimator, thereby validating inference results based on the usual $t$-test. Similar to Theorem (ref), it exhibits the property of self-normalization, which does not require to know the sparsity parameters.

remarkThe usual HC variance estimator of IV estimators is given by $\hat{V}^{{\textsc{iv}},\text{hc}} = ({Z}^\top X)^{-1} \hat{V}_{\text{num}}^{{\textsc{iv}},\text{hc}} (X^\top {Z})^{-1}$ where \begin{align*} \hat{V}^{{iv},hc}_{num} = \begin{pmatrix} \sum\limits_{i=1}^n (\hat{u}_i^{iv})^2 & \sum\limits_{i=1}^n T_i (\hat{u}_i^{iv})^2 & \sum\limits_{i=1}^n (\hat{u}_i^{iv})^2 \sum\limits_{j=1}^n A_{ij}^\text{pre} (T_j - \pi) \\ \sum\limits_{i=1}^n T_i (\hat{u}_i^{\textsc{iv}})^2 & \sum\limits_{i=1}^n T_i (\hat{u}_i^{\textsc{iv}})^2 & \sum\limits_{i=1}^n T_i (\hat{u}_i^{\textsc{iv}})^2 \sum\limits_{j=1}^n A_{ij}^\text{pre} (T_j - \pi) \\ \sum\limits_{i=1}^n (\hat{u}_i^{\textsc{iv}})^2 \sum\limits_{j=1}^n A_{ij}^\text{pre} (T_j - \pi) & \sum\limits_{i=1}^n T_i (\hat{u}_i^{\textsc{iv}})^2 \sum\limits_{j=1}^n A_{ij}^\text{pre} (T_j - \pi) & \sum\limits_{i=1}^n (\hat{u}_i^{\textsc{iv}})^2 \left( \sum\limits_{j=1}^n A_{ij}^\text{pre} (T_j - \pi) \right)^2 \end{pmatrix}. \end{align*} After appropriately scaling the $(3,1)$, $(1,3)$, $(3,2)$, and $(2,3)$ elements of $\hat{V}^{{\textsc{iv}},\text{hc}}_{\text{num}}$, these terms converge to zero. Thus, the primary differences between $\hat{V}^{\textsc{iv},\text{hc}}_{\text{num}}$ and $\hat{V}^{\textsc{iv}}_{\text{num}}$ lie in the $(2,3)$, $(3,2)$, and $(3,3)$ terms. The variance estimator $\hat{V}^{\textsc{iv},\text{hc}}_{\text{num}}$, designed for i.i.d. data, does not account for dependencies across units induced by the SSIV $Z_i^{\textsc{ssiv}}$. In contrast, $\hat{V}^{\textsc{iv}}_{\text{num}}$ captures these dependencies, leading to more accurate variance estimation. This issue is similar to that documented by AdaoKolesarMorales2019, who study inference in shift-share regression designs, where regional outcomes are regressed on a weighted average of sectoral shocks using regional sector shares as weights. They show that over-rejection arises when using cluster-robust standard errors because regression residuals are correlated across regions with similar sectoral shares, regardless of their geographic location.

Modification of SSIV

I begin this subsection with a more formal explanation of the failure of relevance, offering insights into how to modify the SSIV. To simplify the argument, consider the special case where $q_n^\text{pre} = q_n^\text{post} = q_n$. The relevance of SSIV is measured by $n^{-1}\sum_{i=1}^n M_i Z_i^{\textsc{ssiv}}$.\footnote{For the sake of a heuristic argument, I focus on $\frac{1}{n}\sum_{i=1}^n M_i Z_i^{\textsc{ssiv}}$ instead of the sample covariance $\frac{1}{n}\sum_{i=1}^n M_i Z_i^{\textsc{ssiv}} - \frac{1}{n}\sum_{i=1}^n M_i \frac{1}{n}\sum_{i=1}^n Z_i^{\textsc{ssiv}}$. Since the SSIV $Z_i^{\textsc{ssiv}}$ is mean zero, this simplification does not change the essence of the argument.} Recall the decomposition that $M_i = \xi_i + r^*_i$, where $\xi_i$ is defined in (ref) and $r^*_i$ is the remainder term, capturing information from all other units $j\ne i$. The relevance of SSIV is then given by: \[ \frac{1}{n}\sum_{i=1}^n M_i Z_i^{\textsc{ssiv}} = \frac{1}{n}\sum_{i=1}^n \xi_i Z_i^{\textsc{ssiv}} + \frac{1}{n}\sum_{i=1}^n r^*_i Z_i^{\textsc{ssiv}}. \] By definition, $\xi_i$ is a function of $T_i$ and $w_i$ and is uncorrelated with the treatments $T_j$ of all other units $j \neq i$, and thus uncorrelated with $Z_i^{\textsc{ssiv}}$. Consequently, the first term, $n^{-1} \sum_{i=1}^n \xi_i Z_i^{\textsc{ssiv}}$, has mean zero but a variance of order $O(n q_n^2)$, referred to as the noise term. The variance increases with the sparsity parameter as the dependence of $Z_i$ across units strengthens in denser networks. Thus, this term does not contribute to the signal in the relevance measure but instead adds noise. In contrast, the second term, $n^{-1} \sum_{i=1}^n r_i^* Z_i^{\textsc{ssiv}}$, where $r_i^*$ contains information from all other units $j$, contributes to the signal in the relevance measure and is referred to as the signal term. It has a non-zero mean and remains of constant scale. In other words, in $n^{-1} \sum_{i=1}^n M_i Z_i^{\textsc{ssiv}}$, the first term contributes only noise, while the second term carries the signal for the relevance measure. As networks become denser, the noise term increasingly dominates the signal term, preventing the first stage from converging to a nonzero constant and resulting in an inconsistent estimator. Therefore, restoring the consistency of the IV estimators is feasible if the noise term can be effectively reduced while preserving the signal term.

remarkThere is a remarkable connection between SSIV and the estimator of the indirect effect in LiWager2022, who study asymptotics for treatment effect estimation under network interference, with the network randomly drawn from a graphon. They consider anonymous interference, where potential outcomes depend on the fraction of treated friends, not their identities. Unlike my setting, they assume that the network remains unchanged following the intervention. The potential outcome for individual $i$ is $Y_i(t_i, t_{-i}) = f_i\left(t_i, \sum_{j=1}^n A_{ij} t_j/ \sum_{j =1}^n A_{ij} \right)$, where $f_i \in \mathcal{F}$ allows arbitrary dependence on the latent variable $w_i$. They propose an estimator for the indirect effect defined in (ref), \[ \hat{\tau}_{\text{IND}}^{\text{U}} = \frac{1}{n}\sum_{i=1}^n Y_i \left( \frac{\sum_{j\ne i} A_{ij} T_j }{\pi} - \frac{\sum_{j\ne i}A_{ij} (1-T_j)}{1-\pi} \right), \] which is derived as the difference between the Horvitz--Thompson estimators for direct and total effects, and thus is unbiased. Rewriting this estimator gives \[ \hat{\tau}_{\text{IND}}^{\text{U}} = \frac{1}{\pi(1-\pi)} \frac{1}{n} \sum_{i=1}^n Y_i \left( \sum_{j\ne i}A_{ij} (T_j - \pi) \right), \] which implicitly uses SSIV to address endogeneity.

Here, I explain how to reduce the noise term, drawing inspiration from the PC balancing idea in LiWager2022. Ideally, the term $\frac{1}{n} \sum_{i=1}^n \xi_i Z_i^{\textsc{ssiv}}$ could be eliminated by adjusting the SSIV to be the residual from the projection of $Z_i^{\textsc{ssiv}}$ onto $\xi_i$, leveraging the orthogonality property. However, since $\xi_i$ is unknown, the noise term $\frac{1}{n} \sum_{i=1}^n \xi_i Z_i$ can instead be reduced by projecting $w_i$ out of $Z_i$, thereby decreasing the noise while preserving the relevance signal from other units $j$. Although $w_i$ is unobserved, its information can be extracted through spectral analysis of the graphon.

To be more specific, let $G_n^\text{pre}$ denote the graphon matrix, where the $(i,j)$ element is given by $G^\text{pre}_{n,ij} = g^\text{pre}(w_i, w_j)$. By performing eigenvalue decomposition of $q_n^\text{pre} G_n^\text{pre}$, I express it as $q_n^\text{pre} G_n^\text{pre} = \sum_{k=1}^n \lambda_k^* \psi_k^* \psi_k^{*\top}$, where $\lambda_k^*$ represents the $k$-th largest eigenvalue of $q_n^\text{pre} G_n^\text{pre}$. I assume that the graphon is approximately low-rank, as formalized in Assumption (ref), meaning it can be well-approximated by the leading $r$ terms: $q_n^\text{pre} \tilde{G}^\text{pre}_n = \sum_{k=1}^r \lambda_k^* \psi_k^* \psi_k^{*\top}$. I project $Z_i$ onto the first $r$ eigenvectors, $\{{\psi}_k^*(w_i)\}_{k=1}^r$, with the coefficients:

align*[align* omitted — 109 chars of source]

which follows from the orthogonality of eigenvectors. The residual from this projection yields the “oracle version” of the modified SSIV: ${Z}_i^{\textsc{de}*} = Z_i^{\textsc{ssiv}} - \sum_{k=1}^r {\gamma}^*_k {\psi}_k^*(w_i)$. For a feasible estimatior, I approximate ${\psi}_k^*(w_i)$ using the eigenvector of pre-intervention adjacency matrix, $A^{\text{pre}} = \sum_{k=1}^n \hat{\lambda}_k \hat{\psi}_k \hat{\psi}_k^\top$. Therefore, the modified SSIV, denoted as ${Z}_i^{\textsc{de}}$ (for “denoised”), is given by: \[ {Z}_i^{\textsc{de}} = Z_i^{\textsc{ssiv}} - \sum_{k=1}^r \hat{\gamma}_k \hat{\psi}_k(w_i) \] where $\hat{\gamma}_k = \sum_{i=1}^n \hat{\psi}_k(w_i) Z_i^{\textsc{ssiv}}$. By the properties of projection, the modified SSIV ${Z}_i^{\textsc{de}}$ is orthogonal to the space spanned by $\{\hat{\psi}_k(w_i)\}_{k=1}^r$, which encodes information about $w_i$. This effectively reduces noise without compromising the signal.

In the network literature, particularly in settings involving inference with eigenvectors Cai2022, LeLi2022, LiWager2022, it is widely assumed that the graphon is low-rank in terms of its eigenfunctions, such that

align[align omitted — 92 chars of source]

where $\mathbb{E}\left[\psi_k^2(w_i)\right]=1$ and $\mathbb{E}\left[\psi_k(w_i) \psi_l(w_i)\right]=0 \text { for } k \neq l$. This assumption is satisfied by many well-known network models, including the stochastic block model HollandLaskeyLeinhardt1983a and the random dot product graph YoungScheinerman2007, AthreyaFishkindTang2017. Another popular model is the latent space model HoffRafteryHandcock2002, which falls within the class of inhomogeneous Erdős--Rényi models, including the homophily model and the beta model. The latent space model and the more general graphon models typically do not impose a low-rank structure as in (ref), but instead rely on certain smoothness conditions for the graphon function. This paper assumes that the graphon $g^\text{pre}$ can be well-approximated by a low-rank representation based on its eigenvalues, as outlined in Assumption (ref).

assumptionThere exists some constant $r$ $(r \prec n)$ such that \begin{enumerate}[(a)] • $\min\limits_{k \in \{1,\cdots, r-1\}}(\lambda_k -\lambda_{k+1}) \asymp n q_n^\text{pre}$; • $\left\|\sum_{k=r+1}^n \lambda_k^* \psi_k^* \psi_k^{*\top} \right\|_{\textup{op}}=O_{\mathbb{P}}( q_n^\text{pre} )$. \end{enumerate}

Assumption (ref)(a) requires the minimum eigen-gap, i.e., the spacing between an eigenvalue and the rest of the spectrum, to be sufficiently large. Several papers propose eigenvector estimators that are robust to small eigen-gaps. For example, ChengWeiChen2021 tackle eigenvector estimation for low-rank matrices with small eigen-gaps and noisy observations, and LiCaiPoor2022 focus on estimating linear functionals of unknown eigenvectors under similarly tight eigen-gap conditions. While these robust estimators could potentially extend my methods to such settings, doing so is beyond the scope of this paper. Assumption (ref)(b) is analogous to the sparsity assumption in high-dimensional analysis, which assumes that only a small subset of features (variables) significantly contribute to the outcome. This assumption enables more efficient estimation and prediction in models with a large number of features, possibly exceeding the number of observations Tibshirani1996.

To estimate $\beta$'s, I propose the IV fits with the modified IV vector $\tilde{Z}_i^{\textsc{de}}$: \[ \tilde{Z}_i^{\textsc{de}} = \left( 1, T_i, {Z}_i^{\textsc{de}} \right). \] Stack $\tilde{Z}_i^{\textsc{de}}$ to form the $n\times 3$ matrix $\tilde{Z}^{\textsc{de}}$. Let $\hat{\beta}^{\textsc{de}}$ denote the vector of the coefficients obtained from the above IV fit.

theoremSuppose $q_n^\text{pre} \succ {\frac{\log(n)}{\log(\log(n))}} / n$. Under Assumptions (ref), (ref) and (ref), \begin{enumerate}[(a)] • if $\operatorname{Var}\left( \xi_i \right) > 0$ with $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec \sqrt{q_n^\text{pre}}$, then \begin{align*} \hat{\beta}^{de} - {\beta} = O_{\mathbb{P}}\left( \frac{ \max\{q_n^pre, q_n^post \} }{ \sqrt{q_n^pre }} \right); \end{align*} • if $\operatorname{Var}\left( \xi_i \right) = 0$, it holds that $\hat{\beta}^{\textsc{de}}_1 - {\beta}_1 = O_{\mathbb{P}}\left( \frac{ 1 }{ \sqrt{n} } \right)$. Moreover, with $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec \sqrt{q_n^\text{pre}}$, then \begin{align*} \hat{\beta}^{de}_0 - {\beta}_0 = O_{\mathbb{P}}\left( \frac{ \max\{q_n^pre, q_n^\text{post} \} }{ \sqrt{q_n^\text{pre} }} \right) \text{ and } \hat{\beta}^{\textsc{de}}_2 - {\beta}_2 = O_{\mathbb{P}}\left( \frac{ \max\{q_n^\text{pre}, q_n^\text{post} \} }{ \sqrt{q_n^\text{pre} }} \right). \end{align*} \end{enumerate}

To make the modified SSIV work, it requires precise estimation of the eigenvectors, which becomes problematic when the network is too sparse AltDucatezKnowles2021, Benaych-GeorgesBordenaveKnowles2019, Benaych-GeorgesBordenaveKnowles2020. Despite this added requirement, ${\frac{\log(n)}{\log(\log(n))}} / n$ is much sparser than the threshold at which SSIV fails, $n^{-1/2}$. As a result, the consistency regime of the modified SSIV overlaps with that of SSIV, while also extending to the regime where SSIV becomes inconsistent.

remarkTo clarify the condition $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec \sqrt{q_n^\text{pre}}$, I discuss it in two cases: \begin{enumerate}[(1)] • if $q_n^\text{pre} \prec q_n^\text{post}$, then $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec \sqrt{q_n^\text{pre}}$ holds under $\sqrt{q_n^\text{pre}} \succ q_n^\text{post}$, saying that the pre-intervention network can not be too sparse compared with the post-intervention network; • if $q_n^\text{pre} \succcurlyeq q_n^\text{post}$, then $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec \sqrt{q_n^\text{pre}}$ holds under $q_n^\text{pre} \prec 1$. \end{enumerate} The modified SSIV cannot be extended to the case where both pre- and post-intervention networks are dense. In such a case, one potential solution is to use the matching method from Auerbach2022a if some regularity assumption holds, e.g., $M_i$ has sufficient variation when controlling for $w_i$.

Corollary (ref) simplifies the results in Theorem (ref) under the special case $q_n^\text{post} \succcurlyeq q_n^\text{pre}$.

corollarySuppose $q_n^\text{post} \succcurlyeq q_n^\text{pre} \succ {\frac{\log(n)}{\log(\log(n))}} / n$. Under Assumptions (ref), (ref) and (ref), then \begin{enumerate}[(a)] • if $\operatorname{Var}\left( \xi_i \right) > 0$ with $\sqrt{q_n^\text{pre}} \succ q_n^\text{post}$, then $\hat{\beta}^{\textsc{de}} - \beta = O_{\mathbb{P}}\left( \frac{q_n^\text{post}}{\sqrt{q_n^\text{pre}}} \right)$; • if $\operatorname{Var}\left( \xi_i \right) = 0$, it holds that $\hat{\beta}^{\textsc{de}}_1 - {\beta}_1 = O_{\mathbb{P}}\left( \frac{ 1 }{ \sqrt{n} } \right)$. Moreover, with $\sqrt{q_n^\text{pre}} \succ q_n^\text{post}$, then $\hat{\beta}^{\textsc{de}}_0 - {\beta}_0 = O_{\mathbb{P}}\left( \frac{q_n^\text{post}}{\sqrt{q_n^\text{pre}}} \right)$ and $\hat{\beta}^{\textsc{de}}_2 - \beta_2 = O_{\mathbb{P}}\left( \frac{q_n^\text{post}}{\sqrt{q_n^\text{pre}}} \right)$. \end{enumerate}

Aligning with Corollary (ref), Corollary (ref) demonstrates that under Case (a), the convergence rate of $\hat{\beta}$ depends on the sparse parameters and is (possibly) slower than the usual rate $\sqrt{n}$, but faster than the IV estimators using SSIV, which converge at the rate $1/(\sqrt{n} q_n^\text{post})$. In Case (b), I also find that when the post-intervention network $A^{\text{post}}$ is conditionally mean independent of the treatment, $\hat{\beta}^{\textsc{de}}_1$ is less affected by the dependency of $M_i$ across $i$ and retains the usual rate $\sqrt{n}$. Meanwhile, the convergence rates of $\hat{\beta}^{\textsc{de}}_0$ and $\hat{\beta}^{\textsc{de}}_2$ are (possibly) slower than $\sqrt{n}$.

remarkOne takeaway from this subsection is that the discussion on the relevance condition of SSIV sheds light on other forms of mediators. Specifically, if a mediator can be decomposed into two parts, $M_i = \xi_i + r_i^*$, where $\xi_i$ contains only unit $i$'s own information, then $\xi_i$ introduces noise that potentially weakens relevance. The proposed denoising procedure has the potential to restore consistency.

Now I show the asymptotic normality of $\hat{\beta}^{\textsc{de}}$. Define $\mu_k^{u} = \sum_{i=1}^n u_i \psi_k^*(w_i)$ and $\eta_i^{u} = u_i - \sum_{k=1}^r \mu_k^{u} \psi_k^*(w_i)$. Define $V^{\textsc{de}}_{\text{num}} = \operatorname{Var}\left( \sum_{i=1}^n \tilde{Z}^{\textsc{de}}_i u_i \right)$, the variance of the numerator of the estimation bias $(\hat{\beta}^{\textsc{de}} - \beta)$. We can show

align[align omitted — 355 chars of source]

By comparing ${V}^{\textsc{iv}}_{\text{num}}$ in (ref) with ${V}^{\textsc{de}}_{\text{num}}$ in (ref), I observe that the $(3,2)$ and $(2,3)$ terms in (ref) are zero, whereas they are non-zero in (ref). Additionally, the $(3,3)$ term in (ref) accounts for the covariance component, while the $(3,3)$ term in (ref) accounts only for the diagonal component. This modification effectively reduces noise by mitigating the dependency of the IV across units due to information from $w_i$.

Let $\hat{u}_i^{\textsc{de}}$ denote the residual from the IV fits with $\tilde{Z}_i^{\textsc{de}}$. Define $\hat{\mu}_k^{u} = \sum_{i=1}^n \hat{u}_i^{\textsc{de}} \hat{\psi}_{ki}$ and $\hat{\eta}^{u}_i = \hat{u}_i^{\textsc{de}} - \sum_{k=1}^r \hat{\mu}_k^u \hat{\psi}_{ki} $. Define the plug-in estimator of ${V}_{\text{num}}^{\textsc{de}}$ as

align*[align* omitted — 414 chars of source]

Define $\hat{V}^{\textsc{de}} = ((\tilde{Z}^{\textsc{de}})^\top X)^{-1} \hat{V}_{\text{num}}^{\textsc{de}} (X^\top \tilde{Z}^{\textsc{de}})^{-1}$ as the variance estimator of $\hat{\beta}^{\textsc{de}}$.

Theorem (ref) focuses on the sparsity regimes under which the IV estimators $\hat{\beta}^{\textsc{de}}$ are consistent.

theoremSuppose $q_n^\text{pre} \succ {\frac{\log(n)}{\log(\log(n))}} / n$. Under Assumptions (ref), (ref) and (ref), and with $\max\{q_n^\text{pre}, q_n^\text{post}\} \prec \sqrt{q_n^\text{pre}}$, then \[ \left(\hat{V}^{\textsc{de}}\right)^{-1/2} \left( \hat{\beta}^{\textsc{de}} - \beta \right) \stackrel{d}{\rightarrow} \mathcal{N}(0, I_3). \]

Theorem (ref) demonstrates the asymptotic normality of the IV estimators $\hat{\beta}^{\textsc{de}}$ using the modified SSIV, and the consistency of the variance estimator, which together validates the inference results based on the usual $t$ test. While comparing the asymptotic variances of $\hat{\beta}^{\textsc{iv}}$ and $\hat{\beta}^{\textsc{de}}$ is not straightforward, simulation results in Section (ref) indicate that, in the regimes where both SSIV and modified SSIV yield consistent estimators, using the modified SSIV does not increase the variance.

Monte Carlo Simulation

In this section, I provide simulation evidence to support my theory. I draw treatment indicator $T_i \overset{\text{i.i.d.}}{\sim} \text{Bern}(0.5)$. I consider the following four designs for network generation: \paragraph{Design 1: Rank-3 Stochastic Block Model}

itemize$g^\text{pre}(w_i, w_j) = \begin{cases} 3/5 & \text { if } \Phi(w_i) \leq \frac13 \text { and } \Phi(w_j) \le \frac13; \\ 1/3 & \text { if } \frac13 < \Phi(w_i) \leq \frac23 \text { and } \frac13 < \Phi(w_j) \leq \frac23; \\ 1/2 & \text { if } \frac23 < \Phi(w_i) \leq 1 \text { and } \frac23 < \Phi(w_j) \leq 1; \\ 1/5 & \text { otherwise}. \end{cases}$$g^\text{post}(w_i, w_j, T_i, T_j) = \begin{cases} 3/5 & \text { if } \Phi(w_i(1-T_i)) \leq \frac13 \text { and } \Phi(w_j(1-T_j)) \leq \frac13; \\ 1/3 & \text { if } \frac13 < \Phi(w_i(1-T_i)) \leq \frac23 \text { and } \frac13 < \Phi(w_j(1-T_j)) \leq \frac23; \\ 1/2 & \text { if } \frac23 <\Phi(w_i(1-T_i)) \le 1 \text { and } \frac23 <\Phi(w_j(1-T_j)) \le 1; \\ 1/5 & \text { otherwise}. \end{cases}$

\paragraph{Design 2: Homophily Model}

itemize$g^\text{pre}(w_i, w_j) = 1-(\Phi(w_i)-\Phi(w_j))^2$. • $g^\text{post}(w_i, w_j, T_j, T_j) = 1- \left(\Phi(w_i)(1-T_i) -\Phi(w_j)(1-T_j) \right)^2$.

\paragraph{Design 3: Beta Model}

itemize$g^\text{pre}(w_i, w_j) = \frac{\exp(\Phi(w_i) + \Phi(w_j))}{1+ \exp(\Phi(w_i) + \Phi(w_j))}$. • $g^\text{post}(w_i, w_j, T_j, T_j) = \frac{\exp(\Phi(w_i) + \Phi(w_j) + T_i + T_j + T_iT_j)}{1+ \exp(\Phi(w_i) + \Phi(w_j) + T_i + T_j + T_iT_j)}$.

\paragraph{Design 4: Homophily Model}

itemize$g^\text{pre}(w_i, w_j) = 1- \left(\Phi(w_i)-\Phi(w_j)\right)^2$. • $g^\text{post}(w_i, w_j, T_i, T_j) = 1- \left(\Phi(w_i(1-T_i))-\Phi(w_j (1-T_j))\right)^2 $.

For all four models, I set the diagonal elements of pre- and post-intervention networks to zero, i.e., $g^\text{pre}(w_i, w_i) = 0$ and $g^\text{post}(w_i, w_i, T_i, T_i) = 0$. The ranks $r$ for these four designs are $3$, $2$, $2$ and $2$, respectively.

I explore the sparsity regime where $q_n^\text{pre} = q_n^\text{post} = q_n$, considering different levels of sparsity. The adjacency matrices for the pre- and post-intervention networks are generated as follows:

align*[align* omitted — 194 chars of source]

where $w_i \overset{\text{i.i.d.}}{\sim } \mathcal{N}(0,1)$ and $\eta_{ij} \overset{\text{i.i.d.}}{\sim } U[0,1]$ for $i<j$ and $\eta_{ij} = \eta_{ji}$. Each simulated dataset has $n\in \{200, 800\}$ with $5,000$ replications.

Table (ref) reports the within-sample mean and standard error of the mediator $M_i$ for varying levels of sparsity and sample sizes over four designs. The results for $q_n = n^{-1}$ show that Designs 1--4 maintain stable standard errors as the sample size increases. For $q_n = n^{-1/2}$, the standard errors for Designs 1--4 decrease with sample size, with Designs 3 and 4 shrinking at a faster rate. Finally, for $q_n = 1$, Designs 1 and 2 maintain stable standard errors as the sample size grows, while Designs 3 and 4 exhibit a reduction in standard errors by half when the sample size increases from 200 to 800. Moreover, for Designs 3 and 4, when the network is denser, the standard errors of $M_i$ decrease, suggesting insufficient variation within the sample and a potential issue of near collinearity when fitting linear regression. Consequently, I use Designs 1 and 2 to represent Case (a) with non-degenerate $\xi_i$, and Designs 3 and 4 to represent Case (a) with constant $\xi_i$. Notably, both Designs 2 and 4 are homophily models, but how treatments enter the network model differs, consequently influencing the properties of $M_i$ and the performance of the estimators.

table[table omitted — 1,023 chars of source]

Section (ref) presents the OLS estimation results without unobserved confounders, while Section (ref) presents the IV estimation results accounting for unobserved confounders. I omit the results of $\beta_0$ for brevity.

OLS estimation

I generate the outcomes as follows: \[ Y_i = \beta_0+ \beta_1 T_i + \beta_2 \frac{\sum_{j=1}^n A_{ij}^\text{post}T_j}{\sum_{j=1}^n A_{ij}^\text{post}} +\varepsilon_i \] with $(\beta_0, \beta_1, \beta_2) = (1,1,0.5)$, $\varepsilon_i \overset{\text{i.i.d.}}{\sim} U[-1,1]$ and $\varepsilon_i \perp \!\!\! \perp w_i$. Thus, there is no endogeneity issue.

Table (ref) presents the results for $q_n = n^{-1/2}$. The results for $q_n = n^{-2/3}$ (sparser networks) and $q_n = n^{-1/3}$ (denser networks) follow similar patterns and are therefore omitted. The top panel of Table (ref) reports the results for $n = 200$, while the bottom panel shows the results for $n = 800$. I report the mean and the standard deviation of the OLS estimators across simulation draws under “$\hat{\beta}^{\textsc{ols}}$” and “std($\hat{\beta}^{\textsc{ols}}$)”, using $A^\text{post}$ and $A^\text{pre}$ to construct the fraction of treated friends, respectively, in order to approximate the true value and asymptotic standard deviation of the estimators. For the OLS fits using $A^\text{post}$, I also report the heteroskedastic consistent standard errors under “s.e.”, and the corresponding coverage of 95% Confidence Intervals (CIs) under “coverage.” For the OLS fits using $A^\text{pre}$, I report the coverage of 95% CIs based on the corresponding heteroskedastic consistent standard error under “coverage.”

In line with Theorem (ref), the estimators $\hat{\beta}^{\textsc{ols}}$ are consistent, with estimators for all four designs concentrated around the true values. As the sample size increases from 200 to 800, the standard deviations of $\hat{\beta}_1^{\textsc{ols}}$ across all models decrease by approximately half, consistent with the expected $\sqrt{n}$ convergence rate. For $\hat{\beta}_2^{\textsc{ols}}$, the standard deviations in Designs 1 and 2 also decrease by half, whereas those in Designs 3 and 4 exhibit a slower convergence rate due to smaller variation in $M_i$ within the sample. In line with Theorem (ref), the coverage of the 95% CIs using the heteroskedasticity-consistent standard errors is close to the target level.

However, if I mistakenly use $A^{\text{pre}}$ to calculate the mediator variable $M_i$, the coefficients may be inaccurate, and the coverage can become arbitrarily poor. It is worth noting that, for Designs 3 and 4, where the post-intervention networks have a weaker dependency on the treatment, the estimators are not as far off, and the coverage of the CIs is better than in Designs 1 and 2.

table[table omitted — 2,053 chars of source]

IV estimation with SSIV

I generate the outcome as follows: \[ Y_i = \beta_0+ \beta_1 T_i + \beta_2 \frac{\sum_{j=1}^n A_{ij}^\text{post}T_j}{\sum_{j=1}^n A_{ij}^\text{post}} + (w_i + \varepsilon_i)/2 \] with $(\beta_0, \beta_1, \beta_2) = (1,1,0.5)$, $\varepsilon_i \overset{\text{i.i.d.}}{\sim} U[-1,1]$ and $\varepsilon_i \perp \!\!\! \perp w_i$. Hence, the error term is $u_i = (w_i + \varepsilon_i)/2$.

Tables (ref) and (ref) present the estimation results of IV fits using SSIV and modified SSIV with $q_n = {\frac{\log(n)}{\log(\log(n))}} / n$ and $q_n = n^{-1/5}$, respectively. For each table, the top panel reports the results for $n = 200$, while the bottom panel shows the results for $n = 800$. The left panel presents the mean and standard deviation of the IV estimators using SSIV, along with our estimated standard errors and the coverage of 95% CIs across simulations, reported as “$\hat{\beta}^{\textsc{iv}}$”, “std($\hat{\beta}^{\textsc{iv}})$”, “s.e.”, and “coverage”. The right panel provides the corresponding statistics for the IV estimators using modified SSIV, reported as “$\hat{\beta}^{\textsc{de}}$”, “std($\hat{\beta}^{\textsc{de}})$”, “s.e.”, and “coverage”.

I begin with Table (ref), where the network is relatively sparse. In this regime, the network is relatively sparse, and Theorem (ref) asserts that the SSIV estimators $\hat{\beta}^{\textsc{iv}}$ are consistent. I observe that the estimators are concentrated around the true values. As the sample size increases from $200$ to $800$, the standard deviations of $\hat{\beta}_1^{\textsc{iv}}$ decrease by half, indicating a convergence rate of $\sqrt{n}$. However, the standard deviation of $\hat{\beta}_2^{\textsc{iv}}$ shows a slower convergence rate. In line with Theorem (ref), the coverage of the 95% CIs using our variance estimate is close to the target level.

Table (ref) presents the results for the IV fits with $q_n = n^{-1/5}$, denser than the threshold where SSIV remains valid. In Table (ref), I also report the results with $n = 1600$. In line with Theorem (ref), I observe that $\hat{\beta}^{\textsc{iv}}$ is no longer consistent for Designs 1 and 2; the estimators can be arbitrarily off, and the standard deviation does not shrink as the sample size increases. For Designs 3 and 4, $\hat{\beta}^{\textsc{iv}}_1$ remains consistent with convergence rate $\sqrt{n}$, while $\hat{\beta}^{\textsc{iv}}_2$ is not. By applying the modified SSIV, I effectively project out some noise and reduce the estimation error. The IV estimator $\hat{\beta}^{\textsc{de}}_1$ is consistent with a convergence rate of $\sqrt{n}$, while $\hat{\beta}^{\textsc{de}}_2$ is consistent with a convergence rate of $1/\sqrt{q_n}$; the coverage of the 95% confidence intervals for both estimators is approximately at the target level.

table[table omitted — 2,180 chars of source]
table[table omitted — 2,896 chars of source]

Empirical Application

In this section, I revisit Prina2015, which offered access to formal savings accounts to a random sample of poor households in 19 villages in Nepal.\footnote{Prina2015 conducted the RCT and collected network data but did not perform any network analysis. ComolaPrina2021 studied the treatment effects while accounting for network changes and revisited Prina2015 as an empirical illustration. ComolaPrina2023 also revisited the RCT from Prina2015 and focused on dyadic-level regressions to assess the treatment's impact on financial flows. I integrate the panel network data released by ComolaPrina2023 with the outcome variables from Prina2015. The treatment variable is available in both datasets.} I apply the method in this paper to make estimations and inference on the causal effects of interest.

I now provide a more detailed description of the empirical setting of the RCT in Prina2015, expanding on Example (ref). Before the introduction of savings accounts, a baseline survey was conducted in May 2010 across 19 villages in Pokhara. All households with a female head aged 18-55 were surveyed. Following the baseline survey, half of the female household heads were randomly assigned to the treatment group through a public lottery between late May and early June 2010, offering them savings accounts at the local bank. The remaining half, in the control group, did not receive this offer. An endline survey was conducted in June 2011, after the intervention. This study analyzes a sample of 915 households that participated in both survey waves. Both surveys collected data on household socioeconomic characteristics and informal financial transactions.

Following ComolaPrina2023, I define financial links as transfers--loans or gifts, either sent or received--between households, based on survey questions such as, “With whom did you exchange loans or gifts?” The resulting networks, both pre- and post-intervention, were sparse, with small groups and minimal clustering. Despite the total number of links remaining stable (328 at baseline, 329 at endline), the network underwent significant reshuffling. As shown in Example (ref), 255 links were broken, and 256 new links formed by the endline. ComolaPrina2023 reported high uptake and usage of the savings accounts, with over 84% of treated households opening accounts and depositing about 8% of their baseline weekly income nearly every week during the first year.

Table (ref) presents the results of OLS estimates and IV estimators using both SSIV and modified SSIV, across various outcomes, with corresponding standard errors shown in parentheses. I also report the point estimates using the normalized SSIV, $\frac{\sum_{j=1}^n A_{ij}^\text{pre}T_j}{\sum_{j=1}^n A_{ij}^\text{pre}}$, without standard errors, as I do not conduct an asymptotic distribution analysis. The top panel reports the results of household expenditures in the 30 days before the endline survey for different categories, measured by Nepalese rupee. The expenditure categories include health, education, festivals and ceremonies, meat and fish, fish, and other expenditures.\footnote{Other expenditures include clothes and footwear, personal care items, house cleaning articles, house maintenance, and bus and taxi fares.} The bottom panel focuses on education-related expenditures, including school fees, textbooks, uniforms, and school supplies (e.g., pens and pencils). For these expenditures, the sample is restricted to households with school-age children (ages 6--16). For brevity, I omit the results for $\beta_0$ across the various estimation specifications.

The discrepancy between the OLS and the IV estimators using SSIV, in either the sign of the coefficients or their significance, highlights the bias caused by endogeneity due to unobserved confounders. The IV results using SSIV show that the treatment affects various outcomes through different channels. In the “Fish” expenditure column, the direct effect of access to savings accounts on fish consumption is not statistically significant, while the indirect effect, mediated through the fraction of treated friends, is significantly positive. A $0.1$ increase in the fraction of friends with access to savings accounts leads to an increase in fish consumption by Rs. $25.29$. Since the fraction of friends with access to savings accounts is $0.062$ higher for the treatment group compared to the control group, this implies an additional Rs. $15.63$ in fish consumption for the treatment group. The spillover effects, from assigning others to the treated group versus the control group, result in an increase in fish consumption by Rs. $252.91$.

For total education expenditures, as well as expenditures on textbooks and school supplies, the direct effects of access to savings accounts are positive and significant, while the indirect effects through fraction of treated friends are not significant. Access to free savings accounts leads to an increase of Rs. 574.43, Rs. 223.26, and Rs. 104.47 in total education expenditure, school textbooks, and school supplies, respectively, compared to control units, holding the fraction of treated friends constant. Our methods effectively disentangle the direct treatment effects from the indirect effects mediated through networks, providing a clearer understanding of the mechanisms by which the intervention influences the outcomes of interest. Moreover, for outcomes where $\hat{\beta}_2^{\textsc{iv}}$ is not significant, the estimates $\hat{\beta}_1^{\textsc{ols}}$ and $\hat{\beta}_1^{\textsc{iv}}$ are closely aligned, as the treatment is randomly assigned and omitted variable bias is not a concern in these cases.

Furthermore, the point estimators from IV fits using SSIV and normalized SSIV are closely aligned, demonstrating the effectiveness of both versions of SSIV. Despite the relative sparsity of the pre-intervention network, the results from the modified SSIV show that the point estimators are closely aligned with those from using the SSIV and normalized SSIV. This coherence further suggests that the modification based on the estimated eigenvectors does not introduce excessive noise into the estimation.

For the sake of comparison, Table (ref) also includes the corresponding results from Prina2015.\footnote{Prina2015 did not analyze the expenditure on “Fish”, so I leave the result blank.} Prina2015 estimated the average effect of being assigned to the treatment group on each outcome variable, with explanatory variables including the treatment, the baseline value of the outcome variable, and a vector of baseline characteristics.\footnote{Baseline characteristics include age, years of education, marital status of the account holder; number of household members; baseline household income; and three dummies for the main source of household income.} The treatment indicator's coefficient is referred to as the ITT due to the presence of non-compliance. Similar to the OLS results in the top panel, for outcomes where $\hat{\beta}_2^{\textsc{iv}}$ in the IV fits is not significant, the ITT estimators reported by Prina2015 and the $\hat{\beta}_1^{\textsc{iv}}$ estimators from the IV fits are closely aligned. This alignment suggests that unobserved confounders do not bias the estimation of the (direct) treatment effect in these cases, as the treatment is randomly assigned.

table[table omitted — 3,651 chars of source]

Conclusion

This paper investigates the identification and inference of treatment effects in RCTs with network interference, focusing on two key aspects: (1) unobserved confounders affecting both network formation and outcomes, and (2) network changes induced by the intervention. The framework demonstrates that treatment affects outcomes in two distinct ways: directly and indirectly through changes in the network mediator. Disentangling these channels offers deeper insights into the intervention's mechanisms and informs more effective policy design. This paper presents methods for the estimation and inference of causal effects in the presence of network interference, addressing endogeneity concerns. In the absence of endogeneity, I recommend OLS estimation using post-intervention network data. When unobserved confounders are present, I recommend IV estimation with SSIV for relatively sparse networks and propose a modification to SSIV for denser networks. For all estimators, I provide consistent variance estimators for normal approximation, ensuring valid inference. More generally, this paper highlights a concern regarding the relevance condition of SSIV as the network becomes denser. While I focus on the mean impact of peers, the discussion of the SSIV relevance condition and denoising modification also provides insights applicable to other forms of mediators.

This paper opens several avenues for future research. A natural extension would be to incorporate endogenous and correlated peer effects, along with covariates, into the linear model. An expanding body of research highlights the importance of measuring and accounting for network changes when evaluating the welfare impacts of policies CarrellSacerdoteWest2013, Jackson2021, BanerjeeBrezaChandrasekhar2023. While the literature extensively examines optimal treatment assignment under interference BairdBohrenMcIntosh2018b, CaiPouget-AbadieAiroldi2022, Viviano2023, Viviano2023a, the challenge of designing effective policies that account for endogenous network evolution remains unresolved. Furthermore, my methods rely on detailed panel network data, but relaxing this requirement to single-period or aggregated network data would be desirable AlidaeeAuerbachLeung2020, BrezaChandrasekharMcCormick2020. I leave these questions for future research.

appendix