Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Design-Based and Network Sampling-Based Uncertainties in Network Experiments
\address{Department of Economics, University of Wisconsin, Madison. 1180 Observatory Drive, Madison, WI 53706-1393, USA.}
\email{[email removed]}
\address{Department of Economics, University of Wisconsin, Madison. 1180 Observatory Drive, Madison, WI 53706-1393, USA.}
\email{[email removed]}
abstractOrdinary least squares (OLS) estimators are widely used in network experiments to estimate spillover effects. We study the causal interpretation of, and inference for the OLS estimator under both design-based uncertainty from random treatment assignment and sampling-based uncertainty in network links. We show that correlations among regressors that capture the exposure to neighbors' treatments can induce contamination bias, preventing OLS from aggregating heterogeneous spillover effects for a clear causal interpretation. We derive the OLS estimator's asymptotic distribution and propose a network-robust variance estimator. Simulations and an empirical application demonstrate that contamination bias can be substantial, leading to inflated spillover estimates.
{\it Keywords: Network Sampling, Design-based Inference, Network Experiments, Spillover Effects, Potential Outcomes}
JEL classification codes: C13, C21
bibunit[apecon]
\section{Introduction}
Network experiments, or randomized controlled trials (RCTs) on networks, have become increasingly common in applied economics (e.g., cai2015social; dizon2020; carter2021subsidies; fernando2021; Beaman2021-hw). A central objective of these experiments is to estimate the “spillover effect” of policy interventions as they propagate through networks. For example, cai2015social estimate spillover effects from randomly assigned information sessions on rice farmers' decisions to purchase a weather insurance product in Chinese villages. In this paper, we develop a comprehensive theoretical framework for ordinary least squares (OLS) estimators in network experiments, explicitly accounting for both design-based uncertainty, arising from randomness in treatment assignment, and sampling-based uncertainty, arising from randomness in sampling units and network links. Our theory is motivated by two key gaps between empirical practice in applied work and existing econometric theory.
The first gap lies in the choice of estimator.
In applications, researchers predominantly use OLS estimators to estimate spillover effects, employing exposure mappings that summarize treatment status and network structure. In our survey of 29 papers analyzing network experiments, published in the “top 5” economics journals and two leading field journals, all of the studies report using the OLS estimator, while only two papers use propensity score-based estimators.\footnote{Specifically, we considered papers published from April 2010 through April 2025 in the following journals: American Economic Review, Econometrica, Quarterly Journal of Economics, Journal of Political Economy, Review of Economic Studies, American Economic Journal: Applied Economics, and Journal of Development Economics. We searched for articles that listed “networks” and either “field experiments” or “randomized trial” as keywords on the Web of Science platform. This search resulted in 52 papers, of which 29 conducted network experiments and are mentioned in the text. These papers are referenced in (ref).} This pattern stands in contrast to the theoretical literature on inference in network experiments (e.g., aronow2017estimating; leung2022causal; gao2023causal), which provides inference results for inverse probability weighting (IPW) estimators that directly estimate average spillover effects.
The other gap is due to ignoring a source of randomness.
In many applied cases, researchers need to collect network information through surveys. This collection process can introduce an extra layer of uncertainty beyond design-based uncertainty.
Moreover, the collected network may only partially capture the true network governing the propagation mechanism. By contrast, the theoretical literature on causal inference in network experiments typically abstracts away from sampling-based uncertainty, assuming that the data correspond to the entire population and that the observed network is complete.
To address these gaps, we make three contributions. First, we develop a novel framework that jointly incorporates design-based randomness in treatment assignment and sampling-based randomness in network links. Our framework considers a finite population of $n$ units, from which units are randomly sampled and treatments are assigned. We explicitly model the network sampling process, focusing on two common sampling methods: (i) induced subgraph sampling, where each sampled unit reports friends within the sample, and (ii) star sampling, where each sampled unit reports friends from the entire population. In this setup, unlike in non-network experiments, sampling-based uncertainty arises from two sources: (i) which units are sampled, and (ii) which links are observed.
We consider potential outcomes that depend on the entire treatment vector, thus violating the Stable Unit Treatment Value Assumption (SUTVA). To address the resulting dimensionality problem, we assume that the potential outcomes are linear in an exposure mapping, a set of user-specified sufficient statistics summarizing treatment status and network structure. For example, a common exposure mapping includes the fraction of one's friends who are treated. Importantly, we do not assume that the user-specified exposure mapping is correctly specified; it may differ from the true exposure mapping in both functional form and dimension. This flexibility also allows us to incorporate censored network links in a unified way.
As our second contribution, we investigate whether the estimands associated with the OLS estimator can be interpreted as causal spillover effects.
We distinguish between two causal targets: a population-level estimand and a sample-level estimand. The population-level estimand is defined as the weighted average of the treatment effect vector across the entire population, including those who are not sampled, with complete network information. On the other hand, the sample-level estimand is defined as the sample average of the treatment effect vector across the sampled units, with the sampled network information.
We show that both types of estimands can be contaminated: each element of the multi-dimensional estimands may reflect causal effects from other elements of the exposure mapping vector.
With heterogeneous treatment effects, correlations among elements in the exposure mapping vector (e.g., the proportion of treated friends and the proportion of friends' treated friends) blur the distinction between the true causal effects in one element and those in another. Although the population-level causal estimand can be free from contamination if the exposure mapping is defined such that there is no correlation among its elements, the sample-level causal estimand can still be subject to contamination, and thus lacks causal interpretability due to network sampling. Missing links can create undesirable correlations between the observed and true exposure mapping across different elements. As a result, the two estimands can remain distinct even in large samples unless the exposure mapping is correctly specified and the network links for the neighborhood are completely sampled.
In our third contribution, we derive asymptotic theory for the OLS estimator and find conditions under which the OLS estimator approximates the estimands.
We show that the OLS estimator is consistent for the sample-level causal estimand, conditionally or unconditionally on the sampling uncertainty. However, because the sample-level causal estimand generally lacks causal interpretability, results from OLS estimation should be interpreted with caution.
If the exposure mapping is correctly specified and there is no potential correlation between the true and observed exposure mappings, the sample-level causal estimand is consistent for the population-level causal estimand; thus, we can guarantee a clear interpretation of the OLS estimator.
We further derive the estimator's asymptotic distribution and provide a conservative network heteroskedasticity and autocorrelation consistent (HAC) variance estimator.
This paper contributes to the literature on design-based inference in network experiments (aronow2017estimating; leung2022causal; gao2023causal; hoshino2024). Previous works have primarily focused on design-based uncertainty, where treatment assignment is the only source of randomness and complete network information is assumed to be available without sampling uncertainty. Additionally, these works have mainly considered IPW estimators, which allow for direct estimation of causal spillover effects, while the OLS estimator has received less attention. To focus on IPW estimators, these works typically assume that the exposure mapping takes discrete values, such as an indicator of whether a unit has at least one treated friend.\footnote{gao2023causal discuss the potential application of IPW-based estimators to continuous exposure mappings.}
In contrast, this paper considers both design-based and sampling-based uncertainties with an explicit network collection process, and focuses on the OLS estimator with exposure mappings as regressors, which is widely used in empirical applications and allows for continuous exposure mappings.
This paper also relates to the literature on simultaneous design-based and sampling-based inference (see abadie2020sampling; xu2022design; abadie2023should; viviano2024policy). Our framework extends the approach of abadie2020sampling to the network setting by allowing for both design-based and sampling-based uncertainties in network experiments, and by focusing on both population-level and sample-level estimands. We differ from abadie2020sampling in several important respects. First, we explicitly model network sampling, where the observed network may be only partially observed. Second, we study the OLS estimator with exposure mappings as regressors, which induces dependence among outcomes and between regressors and sampling indicators, features not present in their analysis. Third, we provide an element-wise causal interpretation of the estimands and the OLS estimator, which is not addressed in their work.
Relatedly, viviano2024policy also considers both design-based and sampling-based uncertainties, including uncertainty arising from network sampling. However, while his approach assumes that all relevant network information for computing the true exposure mapping is observed, our framework allows for the possibility that some relevant network information is unobserved due to sampling uncertainty. Additionally, while viviano2024policy focuses on a sample-level estimand that maximizes a welfare measure, our study is concerned with inference for both population-level and sample-level causal estimands, emphasizing the potential divergence between the two.
This paper is also related to the literature studying the impact of network data collection on parameters of interest (chandrasekhar2016econometrics; griffith2022name; lewbel2023ignoring; hsieh2024). While these papers share a similar motivation in that the network sampling process can affect the estimation of spillover effects, they primarily focus on the potential bias of estimators with respect to homogeneous parameters due to network sampling. In contrast, this paper focuses on the causal interpretability of the OLS estimator with heterogeneous spillover effects. This distinction is important because attenuation bias, as highlighted for example in chandrasekhar2016econometrics, does not necessarily hinder learning about spillover effects if the estimator preserves the sign of the underlying effects. However, we show that the OLS estimator with exposure mappings may not preserve the sign of the true spillover effects due to contamination bias, potentially leading to misleading conclusions.
More broadly, this paper contributes to the literature on the causal interpretability of estimators in linear regressions with heterogeneous treatment effects (angrist1998; goldsmith2022contamination; borusyak2024negative). In particular, goldsmith2022contamination show that the OLS estimator with multi-dimensional treatment indicators can be contaminated in the presence of heterogeneous treatment effects, which aligns with our findings in (ref). There are two important differences. First, we consider a finite population model, whereas goldsmith2022contamination focus on an infinite population model, making it nontrivial to extend their results to our setting. Second, we allow for general exposure mappings as regressors, while goldsmith2022contamination restrict attention to mutually exclusive multi-dimensional treatment indicators. In our context, contamination bias arises from overlaps in the treatment status across elements of the exposure mapping, whereas such overlaps are not possible in the non-network setup of goldsmith2022contamination.
The remainder of this paper is organized as follows.
(ref) introduces the framework for network sampling, the model, and assumptions.
(ref) presents the main results, including a causal interpretation and asymptotic theory.
(ref) proposes a network heteroskedasticity and autocorrelation consistent (HAC) estimator for the standard errors.
(ref) provides a simulation study to illustrate the finite sample properties of the proposed estimator.
(ref) applies the proposed method to a real-world dataset.
(ref) concludes the paper and provides a flowchart ((ref)) outlining recommended steps for conducting inference in network experiments using the OLS estimator.
(ref) discusses how to estimate the nuisance parameters consistently,
(ref) contains technical lemmas, (ref) contains proofs, (ref) presents additional simulation results, and
(ref) lists the papers included in the survey of network experiment research presented in the Introduction.
\section{Model}
In this section, we first outline our framework for modeling network experiments. We then introduce the estimands of interest, which are defined both for the entire population and for the sampled group, as well as the OLS estimator used to estimate these estimands.
\subsection{Population} As in abadie2020sampling, we consider a sequence of finite populations.
There are finitely many units ($n<\infty$) in the population, denoted by $\mathcal{N}_{n}=\{1,...,n\}$. These units are connected through the network represented by an adjacency matrix $\boldsymbol{A}_{n}=[A_{n,i,j}]_{i,j\in\mathcal{N}_{n}}\in\{0,1\}^{n\times n}$.
We define $A_{n,i,j}=1$ if there is a network link between units $i$ and $j$, and $A_{n,i,j}=0$ otherwise.
We assume that the network is undirected $(A_{n,i,j}=A_{n,j,i})$ and has no self-loops ($A_{n,i,i}=0$). Each unit $i$ is characterized by a vector of covariates $Z_{n,i}\in \mathcal{Z}_{n}\subset\mathbb{R}^{d_{Z}}$, potential outcomes $Y_{n,i}^{*}(\cdot)\in\mathcal{Y}_{n}\subset\mathbb{R}$ that depend on the entire vector of binary treatments $\boldsymbol{D}_{n}=[D_{n,i}]_{i\in\mathcal{N}_{n}}\in \{0,1\}^{n}$. We consider the setup where the researcher assigns treatments only to the sampled units, but spillovers to non-sampled units are allowed. The covariates $Z_{n,i}$ include both network information (e.g., number of $i$'s neighbors, degree: $\text{deg}_{n,i}=\sum_{j\neq i}A_{n,i,j}$) and individual information (e.g., $i$'s age). Also, the potential outcomes may violate the Stable Unit Treatment Value Assumption (SUTVA) by allowing for others' treatment status as inputs.
\subsection{Sampling} From a finite population of $n$ units, we draw a sample of $N=\sum_{i=1}^{n}R_{n, i}$ units (hence $n\geq N$), where $R_{n, i}\in\{0,1\}$ is the sampling indicator for the $i$-th unit: $R_{n,i}=1$ if $i$ is in the sample and otherwise $R_{n,i}=0$.
\begin{comment}
We assume that treatments $\boldsymbol{D}_{n}$ are assigned in a way that they are jointly independent and independent from $\boldsymbol{R}_{n}$.
\end{comment}
Given the sampling indicator vector $\boldsymbol{R}_{n}$, partial elements of the true network $\boldsymbol{A}_{n}$ are sampled. We denote the sampled network, given the sampling indicator vector $\boldsymbol{R}_{n}$, as $\widetilde{\boldsymbol{A}}_{n}(\boldsymbol{R}_{n})$.
When the dependence on $\boldsymbol{R}_{n}$ is clear from context, we simply write $\widetilde{\boldsymbol{A}}_{n}$.
The sampled adjacency matrix $\widetilde{\boldsymbol{A}}_{n}=[\widetilde{A}_{n,i,j}]_{i,j\in\mathcal{N}_{n}}\in\{0,1\}^{n\times n}$ has $(i,j)$-element $\widetilde{A}_{n,i,j}$, which equals one if there is a true network link between units $i$ and $j$ and the link is sampled, and zero otherwise.
In this paper, we focus on two canonical network sampling methods: (i) induced subgraph sampling, and (ii) star sampling.
In the induced subgraph sampling case, we sample $\widetilde{\boldsymbol{A}}_{n}= \boldsymbol{R}_{n}\boldsymbol{R}_{n}'\odot \boldsymbol{A}_{n}$ where $\odot$ is the element-wise product and the $(i,j)$-element of $\widetilde{\boldsymbol{A}}_{n}$, $\widetilde{A}_{n,i,j}=R_{n,i}R_{n,j}A_{n,i,j}$ represents a network link between the units $i$ and $j$, which is sampled if both units are sampled.
In the star sampling case, we sample $\widetilde{\boldsymbol{A}}_{n}=\left(\boldsymbol{1}_{n}\boldsymbol{1}_{n}'-(\boldsymbol{1}_{n}-\boldsymbol{R}_{n})(\boldsymbol{1}_{n}-\boldsymbol{R}_{n})'\right)\odot \boldsymbol{A}_{n}$, where $\widetilde{A}_{n,i,j}=\max\{R_{n,i},R_{n,j}\}A_{n,i,j}$ represents a network link between the units $i$ and $j$, which is sampled if at least one of the two units is sampled.
Sampled networks under induced subgraph and star sampling are illustrated in (ref).
In the figure, the sampled units are in blue, and the non-sampled units are in light gray. The sampled links are shown as solid black lines, and the non-sampled links as dashed gray lines.
In practice, if the researcher asks the sampled units to list their friends from the list of sampled units, the induced subgraph sampling network is sampled (e.g., Conley2010,dizon2020; carter2021subsidies). If the researcher asks the sampled units to list their friends from the population, the star sampling network is sampled (e.g., banerjee2013diffusion; cai2015social; Beaman2021-hw).
See Section 5.3 of kolaczyk2014statistical for further examples of network sampling.
\begin{figure}[ht]
\caption{Comparison of induced subgraph sampling (left) and star sampling (right).}
\begin{threeparttable}
\begingroup
\begin{subfigure}{0.49\textwidth}
\caption{Induced subgraph sampling}
\end{subfigure}
\begin{subfigure}{0.49\textwidth}
\caption{Star sampling}
\end{subfigure}
\begin{tablenotes}
• {\it Note:} Blue nodes indicate sampled units, while light gray nodes denote non-sampled units. Solid black links are observable to the researcher; dashed gray links are unobserved.
\end{tablenotes}
\endgroup
\end{threeparttable}
\end{figure}
We denote the observed covariates by $\widetilde{Z}_{n,i}$, which may differ from $Z_{n,i}$ due to network sampling. For example, if $Z_{n,i}$ includes $i$'s degree, then $\widetilde{Z}_{n,i}$ contains $i$'s degree computed from the sampled network $\widetilde{\boldsymbol{A}}_{n}$: $\widetilde{\text{deg}}_{n,i}=\sum_{j\neq i}\widetilde{A}_{n,i,j}$.
Note that we allow both $Z_{n,i}$ and $\widetilde{Z}_{n,i}$ to depend on $\boldsymbol{R}_n$.
Throughout the paper, we maintain the following assumption regarding the sampling process and the assignment mechanism.
\begin{assumption}
(i) Random sampling:
\begin{align*}
R_{n,i}\sim Bernoulli(\rho_{n}) i.i.d.,
\end{align*}
where $\rho_{n}\in(0,1]$ is a sequence of sampling probabilities such that $\rho_{n}\to\rho\in(0,1]$.
(ii) Network sampling:
Given a fixed entire network sequence $\boldsymbol{A}_{n}\in\{0,1\}^{n\times n}$, the $(i,j)$-element of sampled network $\widetilde{\boldsymbol{A}}_{n}$ is generated by the induced subgraph sampling
$\widetilde{A}_{n,i,j}=R_{n,i}R_{n,j}A_{n,i,j}$ or the star sampling $\widetilde{A}_{n,i,j}=\max\{R_{n,i},R_{n,j}\}A_{n,i,j}$.
(iii) Treatment assignment mechanism: Let $\boldsymbol{R}_{n,-i}$ denote the vector $\boldsymbol{R}_n$ excluding the $i$-th element, $R_{n,i}$.
The assignment mechanism $D_{n,i}$ is independent of $\boldsymbol{R}_{n,-i}$ and drawn independently (but not necessarily identically) from a known distribution. The distribution of $D_{n,i}$ is degenerate at $0$ if and only if $R_{n,i}=0$.
\end{assumption}
(ref) (iii) implies $D_{n,i}=0$ if $R_{n,i}=0$, which means we treat only the sampled units.
The simplest example is
\begin{equation}
D_{n,i}\sim \text{Bernoulli}(R_{n,i}p_{n,i}) \text{ independently}.
\end{equation}
While we use the Bernoulli assignment in (ref) for all illustrations in the paper, our theoretical results accommodate more general assignment mechanisms, as specified in (ref) (iii).
Since (ref) (iii) does not require the identical draws, $p_{n,i}$ could depend on $\boldsymbol{A}_{n}$, $Z_{n,i}$ or other observed characteristics of unit $i$. We can equivalently write (ref) (iii) as $D_{n,i}=R_{n,i}D^*_{n,i}$, where $D^*_{n,i}$ is defined as the latent treatment indicator generated by $D^*_{n,i}\sim \text{Bernoulli}(p_{n,i})$ (or more general distribution satisfying the assumption) independently. Note also that the treatment assignment mechanism is known to the researcher, which is satisfied in a randomized controlled trial and commonly assumed in the design-based inference literature.
\begin{remark}
(ref) (i) rules out cluster sampling and multi-wave network sampling, because in such designs $R_{n,i}$ may depend on $R_{n,j}$ for some $j\neq i$ through the cluster or the network, respectively.
(ref) (ii) prohibits censoring of $\widetilde{\boldsymbol{A}}_n$; when censoring occurs, it can be treated as misspecification of the exposure mapping (see (ref)).
(ref) (iii) excludes complex assignment schemes, such as matched-pair or blocked randomization.
\end{remark}
\subsection{Potential Outcome}
As discussed above, each unit's potential outcome $Y^*_{n,i}(\cdot)$ is a function of the full treatment vector $\boldsymbol{D}_{n}$. By (ref) (iii), we can write $\boldsymbol{D}_{n}=\boldsymbol{R}_n\odot\boldsymbol{D}_{n}^*$, where $\boldsymbol{D}^*_{n}=[D^*_{n,i}]_{i\in\mathcal{N}_n}$. Following the literature (e.g., aronow2017estimating), we assume that there is an exposure mapping $T_{n,i}\in\mathcal{T}_n\subset\mathbb{R}^{d_T}$ that essentially determines $i$'s potential outcome by summarizing the network structure and the treatment status vector. See (ref) below for a detailed definition and discussion on the exposure mapping.
We consider a linear potential outcome model, so that for each $t\in\mathcal{T}_{n}$, $Y_{n,i}^{*}(t)$ is defined as follows.
\begin{assumption}
For all $t\in\mathcal{T}_{n}$,
\begin{align*}
Y_{n,i}^{*}(t)=t'\theta_{n,i} + \nu_{n,i},
\end{align*}
where $\theta_{n,i}$ and $\nu_{n,i}$ are non-stochastic.
\end{assumption}
Note that $\theta_{n,i}$ is a vector of heterogeneous treatment effects. Each element $\theta_{n,i,(k)}$ represents the marginal effect of the $k$-th component of the exposure mapping $T_{n,i}$ on the potential outcome $Y_{n,i}^{*}(t)$. For example, if the $k$-th component of $T_{n,i}$ is the share of treated friends, then $\theta_{n,i,(k)}$ captures the causal spillover effect from $i$'s treated friends on $i$'s outcome. Since $\theta_{n,i}$ is non-stochastic, it may depend on the population network $\boldsymbol{A}_{n}$, allowing for heterogeneity based on network structure and unit $i$'s position.
Although a linear model may seem restrictive, when $|\mathcal{T}_{n}|$ is finite (e.g., $\mathcal{T}_{n}=\{0,1\}^2$), this assumption is without loss of generality as discussed in abadie2020sampling.
The realized outcome is $Y_{n,i}=Y^*_{n,i}(T_{n,i})$. Thus, the outcome depends on $\boldsymbol{D}_n$ only via the exposure mapping $T_{n,i}$.
\subsection{Exposure Mapping}
Let the true exposure mapping be $T_{n,i}=g(i,\boldsymbol{D}_{n},\boldsymbol{A}_{n})\in\mathcal{T}_{n}\subset\mathbb{R}^{d_T}$, where $g:\mathcal{N}_{n}\times \{0,1\}^{n} \times \{0,1\}^{n\times n} \to \mathcal{T}_{n}$ is a function that generates the true exposure mapping for each unit. Specifically, for unit $i$, it takes (i) $i$'s index, (ii) the treatment vector $\boldsymbol{D}_{n}$, and (iii) the true network $\boldsymbol{A}_n$ as inputs, and returns a lower-dimensional vector of summary statistics for the outcome. For example, applied researchers use the presence of $i$'s treated friends and the share of $i$'s treated friends as the exposure mapping.
This paper allows the researcher to misspecify the functional form of $g$.
For example, if the researcher uses the presence of treated friends as the exposure mapping, while the true potential outcome is linear in the share of treated friends, then the exposure mapping is misspecified.
We denote this misspecified exposure mapping function by $\widetilde{g}_n:\mathcal{N}_{n}\times \{0,1\}^{n}\times \{0,1\}^{n\times n}\to \widetilde{\mathcal{T}}_{n}$, where $\widetilde{\mathcal{T}}_{n}\in\mathbb{R}^{d_{\widetilde{T}}}$. Note that the dimensions $d_T$ and $d_{\widetilde{T}}$ may differ.
The functional form $\widetilde{g}_n$ could depend on the sample size $n$ (as in (ref)), but for notational simplicity, we omit the subscript $n$.
We assume that dimensions $d_T$ and $d_{\widetilde{T}}$ are constants independent of $n$.
If $\widetilde{g}=g$, then the observed exposure mapping $\widetilde{T}_{n,i}$ can be written as $\widetilde{T}_{n,i}=g(i,\boldsymbol{D}_{n},\widetilde{\boldsymbol{A}}_{n})$. That is, the only difference between the true exposure mapping and the observed exposure mapping is the network input, between $\boldsymbol{A}_{n}$ and $\widetilde{\boldsymbol{A}}_{n}$.
More generally, if the researcher misspecifies $g$ as $\widetilde{g}$, then the observed exposure mapping is $\widetilde{T}_{n,i}=\widetilde{g}(i,\boldsymbol{D}_{n},\widetilde{\boldsymbol{A}}_{n})$.
In this case, the dimensions $d_T$ and $ d_{\widetilde{T}}$ may differ.
Below, we provide some examples of exposure mappings.
\begin{example}
Suppose that the true exposure mapping is $i$'s own treatment indicator:
\begin{align*}
T_{n,i}=g(i,\boldsymbol{D}_{n},\boldsymbol{A}_{n})&=D_{n,i}=R_{n,i}D^*_{n,i}.
\end{align*}
Note that the exposure mapping does not depend on the network information, and as long as the researcher correctly specifies the exposure mapping $g=\widetilde{g}$, we have $T_{n,i}=\widetilde{T}_{n,i}$ for all $i\in\mathcal{N}_n$.
\end{example}
\begin{example}
Suppose that the true exposure mapping is an indicator of the existence of at least one treated friend:
\begin{align*}
T_{n,i}=g(i,\boldsymbol{D}_{n},\boldsymbol{A}_{n})&=\mathds{1}\left\{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}>0\right\},
\end{align*}
and the researcher correctly specifies the exposure mapping as $\widetilde{T}_{n,i}=g(i,\boldsymbol{D}_{n},\widetilde{\boldsymbol{A}}_{n})$. Thus, for the induced subgraph sampling case ($\widetilde{A}_{n,i,j}=R_{n,i}R_{n,j}A_{n,i,j}$),
\begin{align*}
\widetilde{T}_{n,i}&=\mathds{1}\left\{\sum_{j\neq i}R_{n,i}R_{n,j}A_{n,i,j}R_{n,j}D^*_{n,j}>0\right\}=\mathds{1}\left\{R_{n,i}\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}>0\right\}.
\end{align*}
Thus, when $R_{n,i}=1$, we have $T_{n,i}=\widetilde{T}_{n,i}$.
For the star sampling case ($\widetilde{A}_{n,i,j}=\max\{R_{n,i},R_{n,j}\}A_{n,i,j}$),
\begin{align*}
\widetilde{T}_{n,i}&=\mathds{1}\left\{\sum_{j\neq i}\max\{R_{n,i},R_{n,j}\}A_{n,i,j}R_{n,j}D^*_{n,j}>0\right\}=\mathds{1}\left\{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}>0\right\},
\end{align*}
and we have $T_{n,i}=\widetilde{T}_{n,i}$ for all $i\in\mathcal{N}_n$.
\end{example}
Although the two preceding examples correctly specify the exposure mapping, the subsequent example fails to do so.
\begin{example}
Suppose that the true exposure mapping is a vector of a direct treatment, a spillover treatment through a fraction of treated peers, and their interaction term:
\begin{align*}
T_{n,i}=g(i,\boldsymbol{D}_{n},\boldsymbol{A}_{n})&=\left(R_{n,i}D^*_{n,i},\frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}A_{n,i,j}},R_{n,i}D^*_{n,i}\times \frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}A_{n,i,j}}\right).
\end{align*}
By convention, we usually set $\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}/\sum_{j\neq i}A_{n,i,j}=0$ if $\sum_{j\neq i}A_{n,i,j}=0$ to negate the spillover effect.
Suppose that the researcher misspecifies $\widetilde{g}$ as
\begin{align*}
\widetilde{T}_{n,i}=\widetilde{g}(i,\boldsymbol{D}_{n},\widetilde{\boldsymbol{A}}_{n})&=\left(R_{n,i}D^*_{n,i},\mathds{1}\left\{\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}>0\right\}\right).
\end{align*}
In this specification, it is evident that $g \neq \widetilde{g}$ because $d_{T} > d_{\widetilde{T}}$.
The misspecified $\widetilde{g}$ accounts only for the direct effect and the spillover effect represented by an indicator of the presence of at least one treated friend. Consequently, not only do the dimensions differ, but the structures of the variables capturing spillover effects are also distinct.
\end{example}
\subsection{Censored Network}
We can also treat censoring on a sampled network as arising from a misspecified exposure mapping as $\widetilde{g}$ can specify which links in a sampled network $\widetilde{\boldsymbol{A}}$ to be used to compute the exposure mapping. This is empirically relevant as in practice, some studies impose a cap on the number of links each sampled unit can report, leading to a discrepancy between the sampled and censored networks. For example, in cai2015social, each sampled unit was asked to report up to five closest friends, which potentially introduces censoring in the observed network. See also griffith2022name for further examples and a detailed discussion of censoring in network data collection.
The following example illustrates how censoring can be framed as a misspecified exposure mapping:
\begin{example}
Let $g$ be the same as in (ref). Suppose that the researcher misspecifies $\widetilde{g}$ due to the censoring as
\begin{align*}
\widetilde{T}_{n,i}&=\widetilde{g}(i,\boldsymbol{D}_{n},\widetilde{\boldsymbol{A}}_{n})=g(i,\boldsymbol{D}_{n},\boldsymbol{C}_n(\widetilde{\boldsymbol{A}}_{n})\odot\widetilde{\boldsymbol{A}}_{n})=\mathds{1}\left\{\sum_{j\neq i}C_{n,i,j}(\widetilde{\boldsymbol{A}}_{n})\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}>0\right\},
\end{align*}
where $\boldsymbol{C}_n(\widetilde{\boldsymbol{A}}_{n})$ is the censoring indicator matrix whose $(i,j)$-element is $C_{n,i,j}(\widetilde{\boldsymbol{A}}_{n})\in\{0,1\}$, a binary variable that indicates whether unit $j$ is censored from $i$'s perspective. The censoring indicator can be a random variable, as we allow it to be an unknown function of the sampled network $\widetilde{\boldsymbol{A}}_{n}$. For example, $C_{n,i,j}=1$ when unit $i$ (or $j$) with $R_{n,i}=1$ (or $R_{n,j}=1$) is asked to list their five closest friends and $j$ (or $i$) is one of them.\footnote{We can define $C_{n,i,i}(\widetilde{\boldsymbol{A}}_{n})$ arbitrarily because $A_{n,i,i}=0$.} In this example, $g\neq \widetilde{g}$ in general and misspecification occurs due to the censoring.
\end{example}
We distinguish between the sampled network $\widetilde{\boldsymbol{A}}_{n}$ and the censored network $\boldsymbol{C}_n(\widetilde{\boldsymbol{A}}_{n})\odot\widetilde{\boldsymbol{A}}_{n}$, and the discrepancy is framed as the misspecification of the exposure mapping.
This framework is useful for separating the sampling effect from the censoring. In the extreme case with $\rho_n=1$, we sample the entire network $\widetilde{\boldsymbol{A}}_{n}=\boldsymbol{A}_{n}$, but the censoring still matters as we observe $\boldsymbol{C}_n(\boldsymbol{A}_{n})\odot\boldsymbol{A}_{n}$.
For convenience, we will omit the notational dependence of $\boldsymbol{C}_n$ on $\widetilde{\boldsymbol{A}}_{n}$.
The dependence of $\boldsymbol{C}_n$ on $\widetilde{\boldsymbol{A}}_{n}$ is motivated as follows.
In practice, the censored induced subgraph sampling network is observed if the researcher asks the sampled unit to list a fixed number of closest friends \textit{from the sampled friends}. Thus, it usually depends on $[\widetilde{A}_{n,i,j}]_{j\in\mathcal{N}_n}$.
The censored star sampling network is observed if the researcher asks $i$ with $R_{n,i}=1$ to list a fixed number of closest friends \textit{from their friends in population} $[A_{n,i,j}]_{j\in\mathcal{N}_n}$. Since $\widetilde{A}_{n,i,j}=A_{n,i,j}$ holds for $R_{n,i}=1$ for the star sampling network, the censoring depends on $[\widetilde{A}_{n,i,j}]_{j\in\mathcal{N}_n}$. We also allow the arbitrary dependence of $\boldsymbol{C}_n$ on other deterministic variables, such as individuals' preferences regarding their friends, which is a benefit of the design-based framework.
\subsection{Estimands and Estimator}
To facilitate the introduction of our estimands and OLS estimator, we first transform the exposure mappings.
Recall that the exposure mappings $T_{n,i}$ and $\widetilde{T}_{n,i}$ are random vectors that depend on $\boldsymbol{R}_n$ and $\boldsymbol{D}_n$, and the covariates $Z_{n,i}$ and $\widetilde{Z}_{n,i}$ are random vectors that depend only on $\boldsymbol{R}_n$.
Define
\begin{equation*}
X_{n,i}=T_{n,i} - \Lambda_{n}Z_{n,i},\quad\text{ and }\quad\widetilde{X}_{n,i}=\widetilde{T}_{n,i} - \widetilde{\Lambda}_{n}\widetilde{Z}_{n,i},
\end{equation*}
where
$$\Lambda_{n}=\left(\sum_{i=1}^n \mathbb{E}[T_{n,i}Z_{n,i}^{\prime}]\right)\left(\sum_{i=1}^n \mathbb{E}[Z_{n,i}Z_{n,i}^{\prime}]\right)^{-1},$$
and
$$\widetilde{\Lambda}_{n}=\left(\sum_{i=1}^n R_{n, i}\mathbb{E}[\widetilde{T}_{n,i}|\boldsymbol{R}_n]\widetilde{Z}_{n,i}^{\prime}\right)\left(\sum_{i=1}^n R_{n, i}\widetilde{Z}_{n,i}\widetilde{Z}_{n,i}^{\prime}\right)^{-1}.$$
That is, $X_{n,i}$ is the population residual of the regression of $T_{n,i}$ on $Z_{n,i}$, and $\widetilde{X}_{n,i}$ is the residual of the regression of $\widetilde{T}_{n,i}$ on $\widetilde{Z}_{n,i}$ using sampled units. Since we know the treatment assignment distribution with known $p_{n,i}$ and observe $\boldsymbol{R}_n$, we can calculate $\mathbb{E}[\widetilde{T}_{n,i}|\boldsymbol{R}_n]$ analytically.
(ref) summarizes the conditional expectation of widely used exposure mappings when the assignment probability is homogeneous: $D^*_{n,i}\sim \text{Bernoulli}(p_{n})$ i.i.d. The table focuses on the case where the exposure mapping is scalar. The researcher applies it element-wise for multi-dimensional cases.
For the second neighborhood, the expectation can be calculated similarly. See also (ref) below for the modification on multi-dimensional cases with the second neighborhood.
\begin{table}[h]
\caption{Conditional Expectation of Exposure Mappings Frequently Used in Applied Research}
\begin{threeparttable}
\begingroup
{5pt}
\begin{tabular}{c c c}
\hline
\hline
Exposure Mapping & $\widetilde{T}_{n,i}=g(i,\boldsymbol{D}_{n},\widetilde{\boldsymbol{A}}_{n})$ & $\mathbb{E}\left[\widetilde{T}_{n,i}\mid\boldsymbol{R}_n\right]$\\
\hline
Individual Treatment & $R_{n,i}D^*_{n,i}$ & $R_{n,i}p_{n}$\\
Treated Friends Share & $\frac{\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}\widetilde{A}_{n,i,j}}$ & $p_{n}\times\frac{\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}}{\sum_{j\neq i}\widetilde{A}_{n,i,j}}$\\
Treated Friends Number & $\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}$ & $p_{n}\times \sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}$\\
Treated Friends Existence & $\mathds{1}\left\{\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}>0\right\}$ & $1-(1-p_{n})^{\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}}$\\
\hline
\end{tabular}
\begin{tablenotes}
• {\it Note:} Assume that $R_{n,i}\sim \text{Bernoulli}(\rho_{n})$ i.i.d. and $D^*_{n,i}\sim \text{Bernoulli}(p_{n})$ i.i.d. By convention, we usually set $\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}/\sum_{j\neq i}\widetilde{A}_{n,i,j}=0$ if $\sum_{j\neq i}\widetilde{A}_{n,i,j}=0$.
\end{tablenotes}
\endgroup
\end{threeparttable}
\end{table}
To summarize relevant moments of the data, define the population matrix \(\Omega_n\) and the sample matrices \(\tilde Q_n\) and \(\tilde\Omega_n\):
$$
\Omega_n=\frac{1}{n} \sum_{i=1}^n\mathbb{E}\left[\left(\begin{array}{c}
Y_{n,i}\\
X_{n,i}\\
Z_{n,i}
\end{array}\right)\left(\begin{array}{c}
Y_{n,i}\\
X_{n,i}\\
Z_{n,i}
\end{array}\right)^{\prime}\right]\equiv\left(\begin{array}{ccc}
\Omega_n^{YY} & \Omega_n^{YX} & \Omega_n^{YZ}\\
\Omega_n^{XY} & \Omega_n^{XX} & \Omega_n^{XZ}\\
\Omega_n^{ZY} & \Omega_n^{ZX} & \Omega_n^{ZZ}
\end{array}\right),
$$
$$
\widetilde{Q}_n=\frac{1}{N} \sum_{i=1}^n R_{n, i}\left(\begin{array}{c}
Y_{n,i}\\
\widetilde{X}_{n,i}\\
\widetilde{Z}_{n,i}
\end{array}\right)\left(\begin{array}{c}
Y_{n,i}\\
\widetilde{X}_{n,i}\\
\widetilde{Z}_{n,i}
\end{array}\right)^{\prime}\equiv\left(\begin{array}{ccc}
\widetilde{Q}_n^{YY} & \widetilde{Q}_n^{YX} & \widetilde{Q}_n^{YZ}\\
\widetilde{Q}_n^{XY} & \widetilde{Q}_n^{XX} & \widetilde{Q}_n^{XZ}\\
\widetilde{Q}_n^{ZY} & \widetilde{Q}_n^{ZX} & \widetilde{Q}_n^{ZZ}
\end{array}\right),
$$
and
$$
\widetilde{\Omega}_n=\frac{1}{N} \sum_{i=1}^n R_{n, i}\mathbb{E}\left[\left(\begin{array}{c}
Y_{n,i}\\
\widetilde{X}_{n,i}\\
\widetilde{Z}_{n,i}
\end{array}\right)\left(\begin{array}{c}
Y_{n,i}\\
\widetilde{X}_{n,i}\\
\widetilde{Z}_{n,i}
\end{array}\right)^{\prime}\mid\boldsymbol{R}_n\right]\equiv\left(\begin{array}{ccc}
\widetilde{\Omega}_n^{YY} & \widetilde{\Omega}_n^{YX} & \widetilde{\Omega}_n^{YZ}\\
\widetilde{\Omega}_n^{XY} & \widetilde{\Omega}_n^{XX} & \widetilde{\Omega}_n^{XZ}\\
\widetilde{\Omega}_n^{ZY} & \widetilde{\Omega}_n^{ZX} & \widetilde{\Omega}_n^{ZZ}
\end{array}\right).
$$
Note that the expectation for $\Omega_n$ is taken over $\boldsymbol{D}_n$ and $\boldsymbol{R}_n$ while the conditional expectation for $\widetilde{\Omega}_n$ is taken over $\boldsymbol{D}_n$ conditional on $\boldsymbol{R}_n$.
Our estimands of interest are
\begin{equation}
\left(\begin{array}{c}
\theta_{n}^{\mathrm{causal}}\\
\gamma_{n}^{\mathrm{causal}}
\end{array}\right)=\left(\begin{array}{cc}
\Omega_n^{XX} & \Omega_n^{XZ}\\
\Omega_n^{ZX} & \Omega_n^{ZZ}
\end{array}\right)^{-1}\left(\begin{array}{c}
\Omega_n^{XY}\\
\Omega_n^{ZY}
\end{array}\right),
\end{equation}
and
\begin{equation}
\left(\begin{array}{c}
\theta_{n}^{\mathrm{causal,sample}}\\
\gamma_{n}^{\mathrm{causal,sample}}
\end{array}\right)=\left(\begin{array}{cc}
\widetilde{\Omega}_n^{XX} & \widetilde{\Omega}_n^{XZ}\\
\widetilde{\Omega}_n^{ZX} & \widetilde{\Omega}_n^{ZZ}
\end{array}\right)^{-1}\left(\begin{array}{c}
\widetilde{\Omega}_n^{XY}\\
\widetilde{\Omega}_n^{ZY}
\end{array}\right).
\end{equation}
These are causal estimands in the sense specified by abadie2020sampling.
$(\theta_{n}^{\mathrm{causal}},\gamma_{n}^{\mathrm{causal}})'$ concerns the population-level causal effects of intervention while $(\theta_{n}^{\mathrm{causal,sample}},\gamma_{n}^{\mathrm{causal,sample}})$ concerns the sample-level causal effects when the sampling is governed by $\boldsymbol{R}_{n}$.
$(\theta_{n}^{\mathrm{causal}},\gamma_{n}^{\mathrm{causal}})'$ is a solution for the population moment condition:
\begin{align}
\frac{1}{n} \sum_{i=1}^n \mathbb{E}\left[\left(\begin{array}{c}
X_{n,i}\\
Z_{n,i}
\end{array}\right)\left(Y_{n, i}-X_{n, i}^{\prime} \theta_{n}^{\mathrm{causal}} - Z_{n, i}^{\prime} \gamma_{n}^{\mathrm{causal}}\right)\right]=0,
\end{align}
and $(\theta_{n}^{\mathrm{causal,sample}},\gamma_{n}^{\mathrm{causal,sample}})$ is a solution for the sample moment condition:
\begin{align}
\frac{1}{N} \sum_{i=1}^n R_{n, i}\mathbb{E}\left[\left(\begin{array}{c}
\widetilde{X}_{n,i}\\
\widetilde{Z}_{n,i}
\end{array}\right)\left(Y_{n, i}-\widetilde{X}_{n, i}^{\prime} \theta_{n}^{\mathrm{causal,sample}} - \widetilde{Z}_{n, i}^{\prime} \gamma_{n}^{\mathrm{causal,sample}}\right)\mid\boldsymbol{R}_n\right]=0.
\end{align}
We study (i) whether the sample-level estimand can be estimated consistently (internal validity), and, if so, (ii) how closely it approximates the population-level estimand (external validity). We will also discuss whether each element of these estimands admits a causal interpretation, namely, whether each OLS coefficient represents a convex combination of the corresponding heterogeneous treatment effects, which is not discussed in abadie2020sampling.
For the sample-level causal estimand, we consider the ordinary least squares estimator:
\begin{equation}
\left(\begin{array}{c}
\widehat{\theta}_{n}\\
\widehat{\gamma}_{n}
\end{array}\right)=\left(\begin{array}{cc}
\widetilde{Q}_n^{XX} & \widetilde{Q}_n^{XZ}\\
\widetilde{Q}_n^{ZX} & \widetilde{Q}_n^{ZZ}
\end{array}\right)^{-1}\left(\begin{array}{c}
\widetilde{Q}_n^{XY}\\
\widetilde{Q}_n^{ZY}
\end{array}\right).
\end{equation}
Equivalently, the moment condition is
\begin{equation}
\frac{1}{n} \sum_{i=1}^n R_{n, i}\left(\begin{array}{c}
\widetilde{X}_{n,i}\\
\widetilde{Z}_{n,i}
\end{array}\right)\left(Y_{n, i}-\widetilde{X}_{n, i}^{\prime} \theta - \widetilde{Z}_{n, i}^{\prime} \gamma\right)=0.
\end{equation}
An alternative approach is to use the inverse probability weighting (IPW) estimator (e.g., leung2022causal; gao2023causal). A usual condition for the IPW estimator to work in a network experimental setting is the individual-level overlapping condition; in our notation, we need to have $\mathbb{P}[\widetilde{T}_{n,i}=t|\boldsymbol{R}_{n}]\in (\eta,1-\eta)$ almost surely for all $i\in \mathcal{N}_{n}$ and $t\in\mathcal{T}_{n}$ for some $\eta\in(0,1/2)$. This overlapping condition is difficult to maintain in the network sampling framework. For example, consider a population of two connected units. Suppose the first unit is sampled, while the second is not. The exposure mapping is defined as the number of treated neighbors. In this case, $\mathbb{P}[\widetilde{T}_{n,1}=1|\boldsymbol{R}_{n}]=0$, thereby violating the overlapping condition.
Also, it is notable that the IPW estimator typically targets a quantity that differs from our estimands, which are defined through moment conditions in (ref) and (ref).
Throughout this section, we have defined the network sampling framework, the exposure mapping, and the potential outcome model. We have also defined the population- and sample-level estimands, which are the solutions to the population and sample moment conditions, respectively. The next section provides our main theoretical results within this framework.
\section{Main Results}
In this section, we present the main results of this paper. We first discuss the population- and sample-level estimands' causal interpretation, then derive the
asymptotic properties of the OLS estimator for both.
Since we have assumed that the sequence of sampling probabilities $\rho_n$ is bounded away from $0$ ((ref)-(i)), it follows that $N>0$ a.s. for large enough $n$ ((ref) in the Supplemental Appendix). Thus, there is no additional concern for the degeneracy of the estimands and the OLS estimator in a large population, relative to the standard design-based setting with $\rho_n=1$.\footnote{This is why we have a stronger statement than abadie2020sampling who allow $\rho_n\to 0$ as $n\to\infty$ and use “with probability approaching $1$” instead of “almost surely” in their results. Note that (ref) allows $\rho_n\to 0$ as long as $\rho_n n\to \infty$.}
\subsection{Interpretability of the Causal Estimands}
We impose the following regularity conditions for the causal estimands to be well-defined.
These conditions require boundedness of the outcome, exposure mappings, and covariates, as well as full rank of the exposure mappings and covariates.
\begin{assumption}
\
\begin{enumerate}
• (Uniform Boundedness): The sequence of potential outcomes $Y_{n,i}^{*}(\cdot)$ is uniformly bounded, i.e., there exists some constant $\overline{Y}>0$ such that $|Y_{n,i}^{*}(t)|\leq \overline{Y}<\infty$ for all $n$, $i\in \mathcal{N}_{n}$, and $t\in\mathcal{T}$.
• The sequences of exposure mappings $T_{n,i}$ and $\widetilde{T}_{n,i}$ satisfy the following.
\begin{enumerate}
• (Uniform Boundedness): There exists some constant $\overline{T}$ such that $\|T_{n,i}\|,\|\widetilde{T}_{n,i}\|\leq \overline{T}<\infty$ almost surely for all $n,i\in\mathcal{N}_{n}$.
• (Variation): $(1/n)\times\sum_{i\in \mathcal{N}_{n}}\operatorname{Var}(T_{n,i})$ is invertible and $(1/N)\times\sum_{i\in \mathcal{N}_{n}}R_{n,i}\operatorname{Var}(\widetilde{T}_{n,i}\mid\boldsymbol{R}_n)$ is almost surely invertible for large enough $n$.
\end{enumerate}
• The sequences of covariates $Z_{n,i}$ and $\widetilde{Z}_{n,i}$ satisfy the following.
\begin{enumerate}
• (Uniform Boundedness): There exists some constant $\overline{Z}$ such that $\|Z_{n,i}\|,\|\widetilde{Z}_{n,i}\|\leq \overline{Z}<\infty$ almost surely for all $n,i\in\mathcal{N}_{n}$.
• (Full Rank): $(1/n)\times\sum_{i=1}^nZ_{n,i}Z_{n,i}'$ is almost surely full-rank for large enough $n$, and $(1/N)\times\sum_{i=1}^nR_{n,i}\allowbreak \widetilde{Z}_{n,i}\widetilde{Z}_{n,i}'$ is almost surely invertible for large enough $n$.
\end{enumerate}
\end{enumerate}
\end{assumption}
(ref) (iii) implies that the sequences of residualized exposure mappings $X_{n,i}$ and $\widetilde{X}_{n,i}$ satisfy the following.
\begin{enumerate}[label=(\alph*)]
• (Uniform Boundedness): There exists some constant $\overline{X}$ such that $\|X_{n,i}\|,\|\widetilde{X}_{n,i}\|\leq \overline{X}<\infty$ almost surely for all $n,i\in\mathcal{N}_{n}$.
• (Full Rank): $(1/n)\times\sum_{i\in \mathcal{N}_{n}}\mathbb{E}[X_{n,i}X_{n,i}']$ is invertible and $(1/N)\times\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{X}_{n,i}\widetilde{X}_{n,i}'|\boldsymbol{R}_{n}]$ is almost surely invertible for large enough $n$.
\end{enumerate}
The uniform boundedness of the potential outcomes in (ref) (i) is a standard assumption in the literature (e.g., leung2022causal,gao2023causal). (ref) (ii-a) rules out some network statistics in a large, dense network (e.g., a diverging degree). (ref) (ii-b) requires that the exposure mappings are not degenerate across the units. For example, in (ref), (ref) (ii-b) is violated if the network is empty, $A_{n,i,j}=0$ for all $i,j\in \mathcal{N}_{n}$, as $\mathds{1}\{\sum_{j\neq i}R_{n,j}A_{n,i,j}D^*_{n,j}>0\}=0$ for all $i\in\mathcal{N}_{n}$.
(ref) (iii-b) does not exclude the constant term in $Z_{n,i}$ and $\widetilde{Z}_{n,i}$.
(ref) (ii-b) and (iii-b) are not as restrictive as they seem since we have $N>0$ a.s. for large enough $n$.
We impose an additional condition on the exposure mapping:
\begin{assumption}
There exists a sequence of matrices $L_{n}$ such that
\begin{align*}
\mathbb{E}[T_{n,i}|\boldsymbol{R}_{n}]=L_{n}Z_{n,i}\quad\text{a.s.}
\end{align*}
for large enough $n$. Similarly, there exists a sequence of matrices $\widetilde{L}_{n}$ measurable with respect to $\sigma(\boldsymbol{R}_{n})$ such that
\begin{align*}
\mathbb{E}[\widetilde{T}_{n,i}|\boldsymbol{R}_{n}]=\widetilde{L}_{n}\widetilde{Z}_{n,i}\quad\text{a.s.}
\end{align*}
for large enough $n$.
\end{assumption}
This assumption is fairly weak, as it is automatically satisfied if $\mathbb{E}[T_{n,i}|\boldsymbol{R}_{n}]$ and $\mathbb{E}[\widetilde{T}_{n,i}|\boldsymbol{R}_{n}]$ are included in $Z_{n,i}$ and $\widetilde{Z}_{n,i}$, respectively. Typically, in a field experiment, the experimenter knows the assignment mechanism, so $\mathbb{E}[\widetilde{T}_{n,i}|\boldsymbol{R}_{n}]$ can be computed either analytically or numerically and included as covariates. As the following example shows, in some cases, it is sufficient to include some network statistics in the covariates to satisfy this assumption.
\begin{example}
Consider a variant of the exposure mapping in miguel2004worms that counts the number of treated friends:
\begin{align*}
T_{n,i}=g(i,\boldsymbol{D}_{n},\boldsymbol{A}_{n})=\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}.
\end{align*}
If there is no censoring, then $\widetilde{g}=g$.
The conditional expectations of exposure mappings are derived as $\mathbb{E}[T_{n,i}|\boldsymbol{R}_{n}]=\sum_{j\neq i}A_{n,i,j}R_{n,j}p_{n,j}$, and $\mathbb{E}[\widetilde{T}_{n,i}|\boldsymbol{R}_{n}]=\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}p_{n,j}=\sum_{j\neq i}A_{n,i,j}R_{n,j}p_{n,j}$ for $R_{n,i}=1$.
Thus, (ref) holds if the weighted degree $\sum_{j\neq i}A_{n,i,j}R_{n,j}p_{n,j}$ is included in $Z_{n,i}$ and $\widetilde{Z}_{n,i}$.
\end{example}
We obtain the following transformations of the estimands in terms of the individual causal effects $\theta_{n,i}$ in the linear potential outcome model in (ref):
\begin{theorem}
Under (ref), for large enough $n$,
\begin{align*}
\theta_{n}^{\mathrm{causal}}=\left(\sum_{i=1}^n\mathbb{E}[X_{n,i}X_{n,i}']\right)^{-1}\sum_{i=1}^n\mathbb{E}[X_{n,i}X_{n,i}']\theta_{n,i},
\end{align*}
and
\begin{align*}
\theta_{n}^{\mathrm{causal,sample}}=\left(\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{X}_{n,i}\widetilde{X}_{n,i}'|\boldsymbol{R}_{n}]\right)^{-1}\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{X}_{n,i}X_{n,i}'|\boldsymbol{R}_{n}]\theta_{n,i}\quad\text{a.s.}
\end{align*}
\end{theorem}
(ref) shows that $\theta_{n}^{\mathrm{causal}}$ is expressed as a weighted sum of causal effects $\theta_{n,i}$ induced by the exposure mapping. On the other hand, $\theta_{n}^{\mathrm{causal,sample}}$ is not necessarily a weighted sum of $\theta_{n,i}$ because of the difference in $X_{n,i}$ and $\widetilde{X}_{n,i}$ in the numerator.
Moreover, the dimension of $\theta_{n}^{\mathrm{causal,sample}}$ is $d_{\widetilde{T}}$, which can be different from $d_{T}$, the dimension of $\theta_{n,i}$.
In the absence of (ref), it is known that the formula in (ref) does not hold due to the omitted variable bias (OVB).
(ref) and (ref) suggest a takeaway for practitioners: \textit{under the linear propensity scores, the researcher can select necessary controls
easily to avoid the OVB.}
The linear propensity score assumption (ref) is a weak assumption in design-based causal inference. This assumption also appears in abadie2020sampling and borusyak2023nonrandom.
In the latter, the OVB is removed by using the recentered instruments. Theoretically, including the controls and using the recentered instruments are equivalent, but including the controls is more frequently used in practice. While borusyak2023nonrandom focuses on homogeneous treatment effects, this paper allows for heterogeneous treatment effects.
Note that in general, the $k$-th elements of $\theta_{n}^{\mathrm{causal}}$ and $\theta_{n}^{\mathrm{causal,sample}}$ do not directly correspond to the causal effect of changes in the $k$-th element of the exposure mapping on the outcomes. For example, if the exposure mapping is two-dimensional, we could have the first element of $\theta_{n}^{\mathrm{causal}}$ to be negative while the first element of $\theta_{n,i}$ is positive for all $i\in\mathcal{N}_{n}$ if the second element of it is significantly negative.
\subsection{Causal Interpretation}
To provide a causal interpretation for each element $\theta_{n,(k)}^{\mathrm{causal}}$ and $\theta_{n,(k)}^{\mathrm{causal,sample}}$, we develop an element-wise version of (ref). To this end, we let $T_{n,i,(k)}$ denote the $k$-th element of $T_{n,i}$. Similarly, we write $\widetilde{T}_{n,i,(k)},X_{n,i,(k)},\widetilde{X}_{n,i,(k)}$. For each $k$, let $U_{n,i,(k)}$ be the residual when projecting $X_{n,i,(k)}$ onto the $X_{n,i,(-k)}=(X_{n,i,(l)})_{l\neq k}$:
\begin{equation*}
U_{n,i,(k)}=X_{n,i,(k)}-\left(\sum_{i=1}^n\mathbb{E}[X_{n,i,(k)}X_{n,i,(-k)}']\right)\left(\sum_{i=1}^n\mathbb{E}[X_{n,i,(-k)}X_{n,i,(-k)}']\right)^{-1}X_{n,i,(-k)}.
\end{equation*}
Similarly, define
\begin{equation*}
\widetilde{U}_{n,i,(k)}=\widetilde{X}_{n,i,(k)}-\left(\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{X}_{n,i,(k)}\widetilde{X}_{n,i,(-k)}'|\boldsymbol{R}_{n}]\right)\left(\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{X}_{n,i,(-k)}\widetilde{X}_{n,i,(-k)}'|\boldsymbol{R}_{n}]\right)^{-1}\widetilde{X}_{n,i,(-k)}.
\end{equation*}
Then, we have the following decompositions:
\begin{corollary}
Under (ref), for large enough $n$,
\begin{align}
\theta_{n,(k)}^{\mathrm{causal}}=\frac{\sum_{i=1}^n\mathbb{E}[U_{n,i,(k)}X_{n,i,(k)}]\theta_{n,i,(k)}}{\sum_{i=1}^n\mathbb{E}[U_{n,i,(k)}^{2}]}
+\frac{\sum_{i=1}^n\mathbb{E}[U_{n,i,(k)}X_{n,i,(-k)}']\theta_{n,i,(-k)}}{\sum_{i=1}^n\mathbb{E}[U_{n,i,(k)}^{2}]}
\end{align}
for each $k=1,...,d_{T}$, and
\begin{align}
\theta_{n,(k)}^{\mathrm{causal,sample}}
=\frac{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}X_{n,i}'|\boldsymbol{R}_{n}]\theta_{n,i}}{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}^{2}|\boldsymbol{R}_{n}]}\quad\text{a.s.}
\end{align}
for each $k=1,...,d_{\widetilde{T}}$.
Under an additional assumption $d_{\widetilde{T}}=d_T$, we can simplify it into
\begin{align}
\theta_{n,(k)}^{\mathrm{causal,sample}}
=&\frac{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}X_{n,i,(k)}|\boldsymbol{R}_{n}]\theta_{n,i,(k)}}{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}^{2}|\boldsymbol{R}_{n}]}\nonumber\\
&+\frac{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}X_{n,i,(-k)}'|\boldsymbol{R}_{n}]\theta_{n,i,(-k)}}{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}^{2}|\boldsymbol{R}_{n}]}\quad\text{a.s.}
\end{align}
for each $k=1,...,d_{T}$.
\end{corollary}
(ref) shows that $\theta_{n,(k)}^{\mathrm{causal}}$ and $\theta_{n,(k)}^{\mathrm{causal,sample}}$ can be influenced by effects from other elements $\theta_{n,i,(l)}$ with $l\neq k$.
However, the residualization does not eliminate contamination bias, because the definition of $U_{n,i,(k)}$ and $\widetilde{U}_{n,i,(k)}$ only implies
$$\sum_{i=1}^n\mathbb{E}[U_{n,i,(k)}X_{n,i,(-k)}']=0 \text{ and } \sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}X_{n,i,(-k)}'|\boldsymbol{R}_{n}]=0,$$
respectively.
Moreover, $\mathbb{E}[U_{n,i,(k)}X_{n,i,(k)}]$ and $\mathbb{E}[\widetilde{U}_{n,i,(k)}X_{n,i,(k)}|\boldsymbol{R}_{n}]$ are not guaranteed to be non-negative.
\begin{example}
Suppose the true exposure mapping is the number of treated friends, $T_{n,i}= \sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}$, and that $T_{n,i}$ takes three possible values $1$, $2$, or $3$.
Suppose the researcher misspecifies the exposure mapping as dummy variables: $\widetilde{T}_{n,i}=(\widetilde{T}_{n,i,(1)},\widetilde{T}_{n,i,(2)},\widetilde{T}_{n,i,(3)})$, where $\widetilde{T}_{n,i,(k)}=\allowbreak\mathds{1}\{\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}=k\}$.
Thus, $d_{\widetilde{T}}=3>d_T=1$.
For simplicity, consider star sampling, which provides all network links in the first neighborhood. Then, $\widetilde{T}_{n,i,(k)}=\allowbreak\mathds{1}\{T_{n,i}=k\}$. Equation (ref) in (ref) implies that the coefficient for $\widetilde{T}_{n,i,(k)}$ has no contamination term because $X_{n,i}$ is a scalar in this example. However, each element of $\theta_{n}^{\mathrm{causal,sample}}$ captures a different weighted sum of $\theta_{n,i}$; thus, the interpretation is unclear.
\end{example}
\begin{remark}
\textbf{(i)} Assuming $d_{\widetilde{T}}=d_T$ requires the researcher to correctly specify the dimension of the exposure mapping ($d_{T}=d_{\widetilde{T}}$).
However, this assumption allows the researcher to misspecify the shape of $\widetilde{g}\neq g$ or mismeasure the network.
\textbf{(ii)} Our result for $\theta_n^{\mathrm{causal}}$ is a design-based analogue of Proposition 1 in goldsmith2022contamination. The main differences are that their analysis is model-based and focuses on mutually exclusive treatment indicators (e.g., K-arms).\footnote{Mutually exclusive treatments guarantee that each treatment's own effect receives a non-negative weight.}\textsuperscript{,}\footnote{
goldsmith2022contamination propose three approaches to eliminate contamination bias, but all require modeling the conditional expectation of heterogeneous treatment effects based on observed covariates. In a design-based setting with deterministic treatment effects $\theta_{n,i}$, such modeling is not appropriate. Even if the modeling assumption is justified, their methods may be unreliable for network experiments due to weak overlap in propensity scores, which is often violated for common exposure mappings.
}
In contrast, we allow more flexible treatments, including network spillovers. Our decomposition for $\theta_n^{\mathrm{causal,sample}}$ additionally accommodates both misspecification of the exposure mapping and mismeasurement of the network.
\textbf{(iii)} If the distribution of $T_{n,i}$ does not depend on $i$, a result in (ref) can be strengthened to
\begin{align*}
\theta_{n,(k)}^{\mathrm{causal}}=\frac{\sum_{i=1}^n\mathbb{E}[U_{n,i,(k)}X_{n,i,(k)}]\theta_{n,i,(k)}}{\sum_{i=1}^n\mathbb{E}[U_{n,i,(k)}^{2}]}
\end{align*}
for any $k$.
That is, we do not have a contamination bias.
However, the weight can be negative.
Moreover, the homogeneous requirement of the treatment variable $T_{n,i}$ is usually violated in design-based network experiments since the exposure mapping depends on the network information for each $i$ and the population network $\boldsymbol{A}_n$ is treated as non-random.
\textbf{(iv)} The weight for $\theta_{n}^{\mathrm{causal}}$ is clearly non-negative if the dimension of the treatment variable $T_{n,i}$ is one ($d_T=1$) because no contamination occurs when $d_{T}=1$.
This result is consistent with borusyak2024negative, but our result in (ref) is more general ($d_T>1$).
\end{remark}
\subsection{When Can We Avoid the Contamination Bias?}
The following statement provides sufficient conditions to avoid contamination bias.
Define the conditional covariance for randomvariables $W_1$ and $W_2$ given $\boldsymbol{R}_n$ as $\operatorname{Cov}(W_1,W_2|\boldsymbol{R}_n)=\mathbb{E}[(W_1-\mathbb{E}[W_1|\boldsymbol{R}_n])(W_2-\mathbb{E}[W_2|\boldsymbol{R}_n])|\boldsymbol{R}_n]$.
\begin{corollary}
Assume that (ref) and $d_{\widetilde{T}}=d_T$ hold.
Suppose that $\mathbb{E}[\operatorname{Cov}(T_{n,i,(k)},T_{n,i,(l)}|\boldsymbol{R}_n)]=0$ for all $i\in\mathcal{N}_{n}$ and for any $l\neq k$.
Then, for large enough $n$, there is no contamination bias for $\theta_{n,(k)}^{\mathrm{causal}}$, i.e.,
\begin{align*}
\theta_{n,(k)}^{\mathrm{causal}}=\frac{\sum_{i=1}^n\mathbb{E}[X_{n,i,(k)}^2]\theta_{n,i,(k)}}{\sum_{i=1}^n\mathbb{E}[X_{n,i,(k)}^{2}]}
\end{align*}
for each $k=1,...,d_{T}$. Suppose that $\operatorname{Cov}(\widetilde{T}_{n,i,(k)},T_{n,i,(l)}|\boldsymbol{R}_{n})=0$ for all $i\in\mathcal{N}_{n}$ with $R_{n,i}=1$ and for any $l\neq k$.
Then, for large enough $n$, there is no contamination bias for $\theta_{n,(k)}^{\mathrm{causal,sample}}$, i.e.,
\begin{align*}
\theta_{n,(k)}^{\mathrm{causal,sample}}
=&\frac{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{X}_{n,i,(k)}X_{n,i,(k)}|\boldsymbol{R}_{n}]\theta_{n,i,(k)}}{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{X}_{n,i,(k)}^{2}|\boldsymbol{R}_{n}]}\quad\text{a.s.}
\end{align*}
for each $k=1,...,d_{T}$.
The weights of $\theta_{n,(k)}^{\mathrm{causal}}$ for $\theta_{n,i,(k)}$ are always non-negative.
If we further assume that
$\operatorname{Cov}(\widetilde{T}_{n,i,(k)},T_{n,i,(k)}|\boldsymbol{R}_{n})\geq0$ a.s. for all $i\in\mathcal{N}_{n}$ with $R_{n,i}=1$ and for all $k=1,...,d_{T}$, then the weights of $\theta_{n,(k)}^{\mathrm{causal,sample}}$ for $\theta_{n,i,(k)}$ are non-negative, i.e.,
\begin{equation*}
\frac{R_{n,i}\mathbb{E}[\widetilde{X}_{n,i,(k)}X_{n,i,(k)}|\boldsymbol{R}_{n}]}{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{X}_{n,i,(k)}^{2}|\boldsymbol{R}_{n}]}\geq0\quad\text{a.s.}
\end{equation*}
for all $i\in\mathcal{N}_{n}$ and each $k=1,...,d_{T}$.
\end{corollary}
The zero conditional covariance assumption is satisfied if elements of $T_{n,i}$ and $\widetilde{T}_{n,i}$ are mutually independent.
The positive conditional covariance assumption is satisfied under the censored network (see (ref) below).
\begin{remark}
Under homogeneous treatment effects $\theta_{n,i}=\theta_{n}$, we have $\theta_{n}^{\mathrm{causal}}=\theta_{n}$, but
\begin{equation*}
\theta_{n}^{\mathrm{causal,sample}}=\theta_{n}-\left(\sum_{i=1}^{n}R_{n,i}\mathbb{E}[\widetilde{X}_{n,i}\widetilde{X}_{n,i}'|\boldsymbol{R}_{n}]\right)^{-1}\sum_{i=1}^{n}R_{n,i}\mathbb{E}[\widetilde{X}_{n,i}(X_{n,i}-\widetilde{X}_{n,i})'|\boldsymbol{R}_{n}]\theta_{n}.
\end{equation*}
Thus, $\theta_{n}^{\mathrm{causal}}$ does not have contamination bias for homogeneous treatment effects, but $\theta_{n}^{\mathrm{causal,sample}}$ does.
Under homogeneous treatment effects and $X_{n,i}=\widetilde{X}_{n,i}$, we have $\theta_{n}^{\mathrm{causal}}=\theta_{n}^{\mathrm{causal,sample}}=\theta_{n}$.
\end{remark}
\begin{example}
Consider the exposure mapping in (ref).
The misspecified exposure mapping is $\widetilde{T}_{n,i}=\widetilde{g}(i,\boldsymbol{D}_{n},\widetilde{\boldsymbol{A}}_{n})=\mathds{1}\left\{\sum_{j\neq i}C_{n,i,j}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}>0\right\}$.
Assume that $D^*_{n,i} \sim \text{Bernoulli}(p_{n})$ for $i = 1, \dots, n$ independently.
By adapting (ref), $\theta_{n}^{\mathrm{causal,sample}}$ is a convex combination of $\theta_{n,i}$.
Indeed, we can calculate
\begin{align*}
\theta_{n}^{\mathrm{causal,sample}}
=&\frac{\sum_{i=1}^nR_{n,i}\left(1-(1-p_n)^{\sum_{j\neq i}C_{n,i,j}\widetilde{A}_{n,i,j}R_{n,j}}\right)\theta_{n,i,(1)}}{\sum_{i=1}^nR_{n,i}\left(1-(1-p_n)^{\sum_{j\neq i}C_{n,i,j}\widetilde{A}_{n,i,j}R_{n,j}}\right)},
\end{align*}
and the weights are non-negative.
In general, if both mappings
$\widetilde{T}_{n,i,(k)}$ and $T_{n,i,(k)}$ are weakly increasing (or both weakly decreasing) in $\{D^*_{n,i}\}_{i\in\mathcal{N}_n}$, then the weights are non-negative.
Thus, censoring does not cause negative weight problems when $g$ is weakly monotone on $\{D^*_{n,i}\}_{i\in\mathcal{N}_n}$ for the first neighborhood exposure mapping.
\end{example}
\subsection{More Examples}
\begin{example}
Consider a general form of exposure mapping. For some function $q:\mathbb{R}^2\to\mathbb{R}$, let $T_{n,i}\allowbreak=(R_{n,i}D^*_{n,i},\allowbreak q(\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j},\allowbreak\sum_{j\neq i}A_{n,i,j}R_{n,j}))$ and
$\widetilde{T}_{n,i}\allowbreak=(R_{n,i}D^*_{n,i},\allowbreak q(\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j},\allowbreak\sum_{j\neq i}\widetilde{A}_{n,i,j}R_{n,j}))$. For example, the share of treated friends is covered by the following $q$:
\begin{equation*}
q\left(\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j},\sum_{j\neq i}A_{n,i,j}R_{n,j}\right)=\frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}A_{n,i,j}R_{n,j}}.
\end{equation*}
It also covers the indicator function as in (ref).
Since $D^*_{n,i}\mathrel{\perp\mspace{-10mu}\perp} D^*_{n,j}$, this satisfies the no-correlation conditions. If $q$ is non-decreasing with respect to the first argument, then $\widetilde{T}_{n,i}$ and $T_{n,i}$ are positively correlated, giving $\theta_{n}^{\mathrm{causal,sample}}$ a clear causal interpretation. This type of exposure mapping is used in cai2015social and carter2021subsidies.
As we illustrated above in the special case, the censoring $\widetilde{T}_{n,i}=(R_{n,i}D^*_{n,i},q(\sum_{j\neq i}C_{n,i,j}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j},\sum_{j\neq i}C_{n,i,j}\widetilde{A}_{n,i,j}R_{n,j}))$ does not cause negative weight problems since the exposure mapping $g$ is weakly monotone on $\{D^*_{n,i}\}_{i\in\mathcal{N}_n}$.
\end{example}
\begin{example}
Let
$T_{n,i}=(R_{n,i}D^*_{n,i}G_{n,i}, R_{n,i}D^*_{n,i}(1-G_{n,i}),(1-R_{n,i}D^*_{n,i})G_{n,i})$, where $G_{n,i}=\mathds{1}\{\sum_{j\neq i}A_{n,i,j}\allowbreak R_{n,j}D^*_{n,j}>0\}$. The exposure mapping categorizes each unit $i$ into one of three mutually exclusive exposure types, based on their own treatment status and the presence of treated friends.\footnote{The slope of the OLS estimator captures the effect associated with the group of units that are untreated and have no treated friends.} The elements are mutually exclusive but dependent, so the no-correlation conditions are violated, and we have a contamination bias. This exposure mapping is used in aronow2017estimating. For the exposure mapping with dependence among its elements, we recommend using the inverse propensity score weighting (IPW) estimators to avoid contamination bias.
\end{example}
\begin{remark}(Comparison with IPW estimators)
The causal estimand for the IPW estimators is the average treatment effect (ATE), $(1/n)\sum_{i=1}^n Y_{n,i}^*(t)$ for each $t$.
In other words, the IPW estimator and the regression estimator are for different causal estimands.
While the IPW estimator works well for cases like (ref), it is not suitable for cases like (ref) because the overlapping condition of the propensity score is easily violated.
For example, suppose that $T_{n,i}$ is the treated friends share $(\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j})/(\sum_{j\neq i}A_{n,i,j})$, and there are two units having three and two friends in the population network, respectively. The former can take $T_{n,i}=1/3$ with positive probability, but the latter never takes the value. Thus, the overlapping condition fails to hold.
Moreover, the overlapping condition can be violated in the sampled network even if it is satisfied in the population network, since the sampled network is a sub-network of the population one.
The choice between the IPW estimator and the regression should be decided by the exposure mapping formula that the researcher wants to use.
We recommend using the IPW estimators to avoid contamination bias when the overlapping condition is satisfied.
On the other hand, if there is any doubt about the overlapping condition or the exposure mapping takes (nearly) continuous values, we suggest using the regression model since it does not require the overlapping condition. We leave a more detailed comparison between the IPW estimator and the OLS estimator for future research.
\end{remark}
\begin{example}
Consider an exposure mapping
$$T_{n,i}=\left(R_{n,i}D^*_{n,i},\frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}A_{n,i,j}},\frac{\sum_{j\neq i}\sum_{k\neq i,j}A_{n,i,j}A_{n,j,k}R_{n,k}D^*_{n,k}}{\sum_{j\neq i}\sum_{k\neq i,j}A_{n,i,j}A_{n,j,k}}\right),$$
where the first element indicates whether unit $i$ is directly treated or not, the second element captures the treated friends share among $i$'s first neighbors, and the third element captures the treated friends share among $i$'s second neighbors.
There are overlaps in $\boldsymbol{D}_{n}$ in the second and third elements if there are triangles in the network, so no-correlation conditions are generally violated. (ref) shows an example of a network with triangles. The second element of $T_{n,i}$ is the average of the neighbors' treatment status including $D_{n,i_{1}}$ and $D_{n,i_{2}}$. The third element is the average of the first neighbors' treatment status, including $D_{n,i_{1}}$ and $D_{n,i_{2}}$, again. Thus, the second and third elements are correlated.
This setting is employed in cai2015social.
An easy way to avoid contamination bias is to modify the exposure mapping $g$ to eliminate the double counting. For example, we can use
\begin{align}
T_{n,i}=\left(R_{n,i}D^*_{n,i},\frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}A_{n,i,j}},\frac{\sum_{j\neq i}\sum_{k\neq i,j}A_{n,i,j}A_{n,j,k}(1-A_{n,i,k})R_{n,k}D^*_{n,k}}{\sum_{j\neq i}\sum_{k\neq i,j}A_{n,i,j}A_{n,j,k}(1-A_{n,i,k})}\right),
\end{align}
instead.
Although we miss some of the second-order links, we still manage to avoid the double counting and hence contamination bias.
\begin{figure}[h]
\caption{Networks with triangle links}
\begin{subfigure}[b]{0.38\textwidth}
\begin{tikzpicture}[
scale=.75,
transform shape,
every node/.style={
circle,
draw=blue,
fill=blue!30,
minimum size=1cm
}
]
\node (center) at (0,0) {$i$};
\node (A) at (2.7,0) {$i_1$};
\node (B) at (-2.5,-1) ;
\node (C) at (2,2) {$i_2$};
\node (D) at (-1,2) ;
\draw (center) -- (A);
\draw (center) -- (B);
\draw (center) -- (C);
\draw (center) -- (D);
\node (A1) at (5,-0.8) ;
\node (A2) at (5,0.8) ;
\draw (A) -- (A1);
\draw (A) -- (A2);
\draw (A1) -- (A2);
\draw (A) -- (C);
\draw (A1) -- ++(1,-0.2);
\draw (A1) -- ++(1,0.2);
\draw (A2) -- ++(1,0.7);
\draw (center) -- ++(0,-1.5);
\draw (A) -- ++(0.3,-1.5);
\draw (B) -- ++(-1,-0.2);
\draw (C) -- ++(0.2,1);
\draw (D) -- ++(0.2,0.9);
\draw (D) -- ++(-2,0.7);
\end{tikzpicture}
\caption{Without censored links}
\end{subfigure}
\begin{subfigure}[b]{0.38\textwidth}
\begin{tikzpicture}[
scale=.75,
transform shape,
every node/.style={
circle,
draw=blue,
fill=blue!30,
minimum size=1cm
}
]
\node (center) at (0,0) {$i$};
\node (A) at (2.7,0) {$i_1$};
\node (B) at (-2.5,-1) ;
\node (C) at (2,2) {$i_2$};
\node (D) at (-1,2) ;
\draw (center) -- (A);
\draw (center) -- (B);
\draw[dashed, gray] (center) -- (C);
\draw (center) -- (D);
\node (A1) at (5,-0.8) ;
\node (A2) at (5,0.8) ;
\draw (A) -- (A1);
\draw (A) -- (A2);
\draw (A1) -- (A2);
\draw (A) -- (C);
\draw (A1) -- ++(1,-0.2);
\draw (A1) -- ++(1,0.2);
\draw (A2) -- ++(1,0.7);
\draw (center) -- ++(0,-1.5);
\draw[dashed, gray] (A) -- ++(0.3,-1.5);
\draw (B) -- ++(-1,-0.2);
\draw (C) -- ++(0.2,1);
\draw (D) -- ++(0.2,0.9);
\draw (D) -- ++(-2,0.7);
\end{tikzpicture}
\caption{With censored (dashed) links}
\end{subfigure}
\end{figure}
\end{example}
\begin{example}
Consider the setup in (ref) but with censoring caused by naming up to four friends.
As illustrated in (ref), suppose that the sampled network link between $i_1$ and $i_2$ is not observed due to the censoring. Then, $i_2$ is misclassified as a second neighborhood friend in the observed network while $i_2$ is a first neighborhood friend in the population network.
Thus, if we consider the true exposure mapping $T_{n,i}$ as in (ref), and a misspecified exposure mapping for the sampled network
\begin{align*}
\widetilde{T}_{n,i}=&\left(R_{n,i}D^*_{n,i},
\frac{\sum_{j\neq i}C_{n,i,j}\widetilde{A}_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}C_{n,i,j}\widetilde{A}_{n,i,j}},\right.\\
&\left.\qquad
\frac{\sum_{j\neq i}\sum_{k\neq i,j}C_{n,i,j}\widetilde{A}_{n,i,j}C_{n,j,k}\widetilde{A}_{n,j,k}(1-C_{n,i,k}\widetilde{A}_{n,i,k})R_{n,k}D^*_{n,k}}{\sum_{j\neq i}\sum_{k\neq i,j}C_{n,i,j}\widetilde{A}_{n,i,j}C_{n,j,k}\widetilde{A}_{n,j,k}(1-C_{n,i,k}\widetilde{A}_{n,i,k})}\right),
\end{align*}
then, there is a correlation between $T_{n,i,(2)}$ and $\widetilde{T}_{n,i,(3)}$.
An easy way to avoid contamination bias is to modify the exposure mapping $\widetilde{g}$ so that $\widetilde{T}_{n,i,(3)}$ equals zero for individuals subject to censoring.
For example, if the censoring happens by asking up to four friends, we can eliminate the individuals with four observed links from consideration
\begin{equation*}
\widetilde{T}_{n,i,(3)}=\frac{\sum_{j\neq i}\sum_{k\neq i,j}C_{n,i,j}\widetilde{A}_{n,i,j}C_{n,j,k}\widetilde{A}_{n,j,k}(1-C_{n,i,k}\widetilde{A}_{n,i,k})R_{n,k}D^*_{n,k}}{\sum_{j\neq i}\sum_{k\neq i,j}C_{n,i,j}\widetilde{A}_{n,i,j}C_{n,j,k}\widetilde{A}_{n,j,k}(1-C_{n,i,k}\widetilde{A}_{n,i,k})}\mathds{1}\left\{\sum_{j\neq i}C_{n,i,j}\widetilde{A}_{n,i,j}<4\right\}.
\end{equation*}
Note that the censoring for $i$ does not matter for the first neighborhood element $\widetilde{T}_{n,i,(2)}$ by the same logic as (ref). Moreover, the censoring for $i_1$ does not matter for the second neighborhood element $\widetilde{T}_{n,i,(3)}$ of $i$ because it does not introduce any misclassification.
\end{example}
\subsection{Asymptotic Theory}
We mostly follow the notation of kojevnikov2021limit. Let $\mathcal{N}_{n}=\{1,...,n\}$ be the set of population units and $d_{n}(i,j)$ be the shortest distance between $i,j\in \mathcal{N}_{n}$ on $\boldsymbol{A}_{n}$ (set $d_{n}(i,i)=0$; set $d_{n}(i,j)=\infty$ if there are no paths between $i$ and $j$). Define $\mathcal{L}_{v}=\{\mathcal{L}_{v,a}:a\in\mathbb{N}\}$, where $\mathcal{L}_{v,a} = \{f:\mathbb{R}^{v\times a}\to \mathbb{R}:\|f\|_{\infty}<\infty,\operatorname{Lip}(f)<\infty\}$, $\|\cdot\|_{\infty}$ is the sup-norm, and $\operatorname{Lip}(f)$ is the Lipschitz constant of $f$. Let $\mathcal{P}_{n}(a,b;s)=\{(A,B):A,B\subset \mathcal{N}_{n},|A|=a,|B|=b,d_{n}(A,B)\geq s\}$, where $d_{n}(A,B)=\min_{i\in A}\min_{j\in B}d_{n}(i,j)$.
For each $A\subset \mathcal{N}_{n}$ and triangular array $(U_{n,i})$, let us write $U_{n,A}=(U_{n,i})_{i\in A}$.
\begin{definition}
A triangular array $\{U_{n,i}\}, n\geq 1,U_{n,i}\in \mathbb{R}^{v}$, is called {\it conditionally $\psi$-dependent given $\boldsymbol{R}_n$}, if for each $n\in\mathbb{N}$, there exists a $\sigma(\boldsymbol{R}_n)$-measurable sequence $\xi_{n}=\{\xi_{n,s}\}_{s\geq 0},\xi_{n,0}=1$, and a collection of nonrandom functions $(\psi_{a,b})_{a,b\in\mathbb{N}},\psi_{a,b}:\mathcal{L}_{v,a}\times\mathcal{L}_{v,b}\to[0,\infty)$ such that for all $(A,B)\in\mathcal{P}_{n}(a,b;s)$ with $s>0$ and all $f\in\mathcal{L}_{v,a}$ and $g\in\mathcal{L}_{v,b}$,
\begin{align*}
|\operatorname{Cov}(f(U_{n,A}),g(U_{n,B}))|\leq \psi_{a,b}(f,g)\xi_{n,s} \quad\text{a.s.}
\end{align*}
\end{definition}
Define
\begin{align*}
\mathcal{N}_{n}(i;s) =\{j\in \mathcal{N}_{n}:d_{n}(i,j)\leq s\},
\end{align*}
which is the set of $i$'s neighborhood within $s$-distance.
First, we assume that the network dependence of the exposure mappings is local.
\begin{assumption}
There exists some $K\in\mathbb{N}$ such that for any $i\in \mathcal{N}_{n},n\in\mathbb{N}$ and $\boldsymbol{d}_{n},\boldsymbol{d}_{n}^{\prime}\in\{0,1\}^{n}$ such that $\boldsymbol{d}_{n,\mathcal{N}_{n}(i,K)}=\boldsymbol{d}'_{n,\mathcal{N}_{n}(i,K)}$,
\begin{align*}
g(i,\boldsymbol{d}_{n},\boldsymbol{A}_{n})&=g(i,\boldsymbol{d}_{n}^{\prime},\boldsymbol{A}_{n}),\quad\text{ and }\quad
\widetilde{g}(i,\boldsymbol{d}_{n},\widetilde{\boldsymbol{A}}_{n})=\widetilde{g}(i,\boldsymbol{d}_{n}^{\prime},\widetilde{\boldsymbol{A}}_{n})\quad\text{a.s.}
\end{align*}
\end{assumption}
Let $\widetilde{d}_{n}(i,j)$ be the shortest distance between $i,j\in \mathcal{N}_{n}$ on $\widetilde{\boldsymbol{A}}_{n}$.
(ref) imply that $T_{n,i}\mathrel{\perp\mspace{-10mu}\perp} T_{n,j}$ if $d_{n}(i,j)>2K$. They also imply that $\widetilde{T}_{n,i}\mathrel{\perp\mspace{-10mu}\perp} \widetilde{T}_{n,j}$ if $d_{n}(i,j)>2K$ because $\widetilde{d}_{n}(i,j)\geq d_{n}(i,j)$ almost surely and because $i$ and $j$ do not share $R_{n,k}$ and $D^*_{n,k}$ for any $k\neq i,j$ in their $K$-neighborhoods.
Under the correctly specified exposure mapping, $g=\widetilde{g}$, the condition $g(i,\boldsymbol{d}_{n},\boldsymbol{A}_{n})=g(i,\boldsymbol{d}_{n}^{\prime},\boldsymbol{A}_{n})$ automatically implies $\widetilde{g}(i,\boldsymbol{d}_{n},\widetilde{\boldsymbol{A}}_{n})=\widetilde{g}(i,\boldsymbol{d}_{n}^{\prime},\widetilde{\boldsymbol{A}}_{n})$ a.s. because the distance on a sampled network is always weakly longer than that on the population network: $\widetilde{d}_{n}(i,j)\geq d_{n}(i,j)$. For the same reason, the distance on a censored network is always weakly longer than that on sampled or population networks.
Define $\mathcal{N}_{n}^{\partial}(i;s)=\{j\in \mathcal{N}_{n}:d_{n}(i,j)=s\}$, which is the set of $i$'s neighborhood with exact $s$-distance, and its $p$-th sample moment $\delta_{n}^{\partial}(s;p)=n^{-1}\sum_{i\in \mathcal{N}_{n}}|\mathcal{N}_{n}^{\partial}(i;s)|^{p}$.
The next assumption requires that the sum of these $p$-th sample moments within $2K$-distance is bounded.
\begin{assumption}
The sequence of networks $(\boldsymbol{A}_{n})$ satisfies
\begin{align*}
\sum_{1\leq s\leq 2K}\delta_{n}^{\partial}(s;1)=O(1).
\end{align*}
\end{assumption}
By a simple calculation and $\rho>0$, we can show that (ref) is equivalent to $(n\rho_n)^{-1}\sum_{i=1}^{n}\allowbreak\sum_{j\in\mathcal{N}_n(i;2K)}1=O(1)$. Also note that (ref) is weaker than the bounded network degree since this assumption only requires the boundedness on average.
Then, we show that our estimator is consistent for the sample-level causal estimand:
\begin{theorem}
Under (ref), $$\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal,sample}}\overset{p^R}{\longrightarrow}0
\quad\text{ and }\quad \widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal,sample}}\overset{p}{\to}0,$$
where $\overset{p^R}{\longrightarrow}$ denotes convergence in probability conditional on $\boldsymbol{R}_n$, that is, for any $\varepsilon>0$,
$$
\mathbb{P}\left(\|\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal,sample}}\|\leq\varepsilon\mid\boldsymbol{R}_n\right)\overset{a.s.}{\longrightarrow}1
$$
as $n\to\infty$.
\end{theorem}
(ref) establishes the internal validity of our network experiment. However, in general, $\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal}}\not\overset{p}{\to}0$ because $\theta^{\mathrm{causal}}_{n}-\theta_{n}^{\mathrm{causal,sample}}\not\overset{p}{\to}0$ due to misspecification of the exposure mapping. Moreover, as shown in (ref), $\theta_{n}^{\mathrm{causal,sample}}$ does not have a clear causal interpretation. Consequently, (ref) does not guarantee the external validity of our network experiment.
Ideally, our network experiment would satisfy $\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal}}\overset{p}{\to}0$ so that each element of $\widehat{\theta}_{n}$ can be interpreted as a causal spillover effect.
We show that this consistency is achieved when there is no misspecification and no mismeasurement ($\widetilde{T}_{n,i}=T_{n,i}$ for each $i\in\mathcal{N}_{n}$) and the observed covariates coincide with those in the population ($\widetilde{Z}_{n,i}=Z_{n,i}$ for each $i\in\mathcal{N}_{n})$.
We are essentially assuming that each $\widetilde{T}_{n,i}$ is computed by $g(i,\boldsymbol{D}_{n},\boldsymbol{A}_{n})=T_{n,i}$ where we replace $\widetilde{g}$ with $g$ and $\widetilde{\boldsymbol{A}}_{n}$ with $\boldsymbol{A}_{n}$.
Under the linear propensity scores, we can show that $X_{n,i}=\widetilde{X}_{n,i}$ a.s. ((ref)).
\begin{assumption}
\
\begin{enumerate}
• We have the following equalities almost surely for $R_{n,i}=1$: $\widetilde{T}_{n,i}=T_{n,i}$ and $\widetilde{Z}_{n,i}=Z_{n,i}$ for all $i\in\mathcal{N}_{n}$ and $n\in\mathbb{N}$.
• Each element of $T_{n,i}$ and $Z_{n,i}$ either does not depend on $R_{n,i}$, or depends on it only through a multiplicative form.
• At most one element of $T_{n,i}$ depends on $i$'s own treatment $R_{n,i}D^*_{n,i}$ and the element does not depend on $R_{n,j}$ and $D_{n,j}$ for any $j\neq i$.
\end{enumerate}
\end{assumption}
(ref) (i) holds when there is no misspecification and no mismeasurement for sampled units, i.e., $g = \widetilde{g}$ and $\widetilde{\boldsymbol{A}}_n = \boldsymbol{A}_n$ locally. For example, under star sampling with an exposure mapping restricted to the first neighborhood, all relevant links are correctly observed for sampled units. However, (ref) (i) may not hold for exposure mappings with higher-order dependence, since more global measurement of the network is then required up to the relevant order. Nonetheless, in such cases, researchers can apply our results to $\theta_{n}^{\mathrm{causal,sample}}$ instead of $\theta_{n}^{\mathrm{causal}}$ and interpret it as a convex combination of heterogeneous treatment effects.
(ref) (ii) means some components of the covariate vector are independent of the sampling indicator, while others incorporate $R_{n, i}$ in a multiplicative way---for example, $Z_{n,i,(k)}=R_{n,i}p_{n,i}$.
We can always pick covariates $Z_{n,i}$ having (ref) (ii) since $R_{n,i}$ enters only multiplicatively for $T_{n,i}$ by (ref) if we include the direct effect without any transformation. Thus, we can choose covariates $Z_{n,i}$ satisfying (ref) and (ref) simultaneously.
(ref) (iii) is satisfied if we do not include the cross term of the direct effect $R_{n,i}D^*_{n,i}$ and a spillover effect.
Excluding the cross term is also used to guarantee no contamination ((ref)).\footnote{We can allow the violation of (ref) (iii) if we modify $\widehat{\theta}_n$ in the same manner as $\widetilde{\gamma}_n$ in (ref).}
Under (ref), we can show the consistency of the OLS estimator $\widehat{\theta}_{n}$ for the population-level causal estimand $\theta_{n}^{\mathrm{causal}}$:
\begin{theorem}
Under (ref), $$\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal}}\overset{p}{\to}0.$$
\end{theorem}
It is worth noting that (ref) does not hold if $\widetilde{Z}_{n,i}\neq Z_{n,i}$, since we cannot ensure $\widetilde{X}_{n,i}\sim X_{n,i}$ asymptotically. Instead, under no misspecification, (ref) implies
\begin{align*}
\theta_{n,(k)}^{\mathrm{causal,sample}}=\frac{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}^{2}|\boldsymbol{R}_{n}]\theta_{n,i}}{\sum_{i=1}^nR_{n,i}\mathbb{E}[\widetilde{U}_{n,i,(k)}^{2}|\boldsymbol{R}_{n}]}
\end{align*}
for each $k=1,...,d_{\widetilde{T}}$. Thus, although the consistency for $\theta_{n}^{\mathrm{causal}}$ may fail in this setting, the absence of misspecification alone recovers the causal interpretability of $\theta_{n}^{\mathrm{causal,sample}}$, and by extension, that of $\widehat{\theta}_{n}$.
Next, we consider the asymptotic distribution of $\widehat{\theta}_{n}$.
Now, we introduce additional dependence measures of the network.
Define $\Delta_{n}(s,m;k)=\frac{1}{n}\sum_{i\in \mathcal{N}_{n}}\max_{j\in \mathcal{N}_{n}^{\partial}(i;s)}|\mathcal{N}_{n}(i;m)\setminus \mathcal{N}_{n}(j;s-1)|^{k}$, and $c_{n}(s,m;k)=\inf_{\alpha >1}[\Delta_{n}(s,m;k\alpha)]^{1/\alpha}\left[\delta^{\partial}_{n}\left(s;\frac{\alpha}{\alpha-1}\right)\right]^{1-1/\alpha}$.
$c_{n}(s,m;k)$ measures the density of the network and is used as a sufficient condition for the CLT.
Define
\begin{align*}
\widetilde{\varepsilon}_{n,i}&=Y_{n,i}-\widetilde{X}_{n,i}'\theta_{n}^{\mathrm{causal,sample}}-\widetilde{Z}_{n,i}'\gamma_{n}^{\mathrm{causal,sample}},\\
\varepsilon_{n,i}&=Y_{n,i}-X_{n,i}'\theta_{n}^{\mathrm{causal}}-Z_{n,i}'\gamma_{n}^{\mathrm{causal}},
\end{align*}
and
\begin{align*}
\widetilde{\Sigma}_{n}
=\operatorname{Var}\left(\sum_{i=1}^{n}R_{n,i}\widetilde{X}_{n,i}\widetilde{\varepsilon}_{n,i}\mid\boldsymbol{R}_n\right),\qquad
\Sigma_{n}
=\operatorname{Var}\left(\sum_{i=1}^{n}R_{n,i}X_{n,i}\varepsilon_{n,i}\right).
\end{align*}
We impose the following assumption, which requires a weak dependence structure in the network and rules out overly dense networks.
\begin{assumption}
There exists a positive sequence $m_{n}\to\infty$ such that for $p=1,2$,
\begin{align*}
n\widetilde{\Sigma}_{n}^{-(1+p/2)}\sum_{s=0}^{2K}c_{n}(s,m_{n};p)\overset{a.s.}{\longrightarrow} 0,\qquad
n\Sigma_{n}^{-(1+p/2)}\sum_{s=0}^{2K}c_{n}(s,m_{n};p)\to 0.
\end{align*}
\end{assumption}
Then, we show that $\widehat{\theta}_{n}$ is asymptotically normal relative to $\theta_{n}^{\mathrm{causal,sample}}$:
\begin{theorem}
Under (ref),
\begin{align*}
\widetilde{\Sigma}_{n}^{-1/2}\widetilde{Q}_n^{XX}(\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal,sample}})\overset{d^R}{\longrightarrow}\mathrm{N}(0,I_{d_{\widetilde{T}}})
\quad\text{and}\quad
\widetilde{\Sigma}_{n}^{-1/2}\widetilde{Q}_n^{XX}(\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal,sample}})\overset{d}{\to}\mathrm{N}(0,I_{d_{\widetilde{T}}}),
\end{align*}
where $\overset{d^R}{\to}$ denotes convergence in distribution conditional on $\boldsymbol{R}_n$, that is,
$$
\left|\mathbb{P}\left(\widetilde{\Sigma}_{n}^{-1/2}\widetilde{Q}_n^{XX}(\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal,sample}}) \leq t\mid\boldsymbol{R}_n\right)-F(t)\right|\overset{a.s.}{\longrightarrow}0
$$
as $n\to\infty$ for any $t\in\mathbb{R}^{d_{\widetilde{T}}}$ letting $F(t)$ be the distribution function of $\mathrm{N}(0,I_{d_{\widetilde{T}}})$.
\end{theorem}
We also show that the absence of misspecification and access to the variables in the population yield asymptotic normality of $\widehat{\theta}_{n}$ relative to $\theta_{n}^{\mathrm{causal}}$:
\begin{theorem}
Under (ref), we have
\begin{align*}
\Sigma_{n}^{-1/2}\widetilde{Q}^{XX}_{n}(\widehat{\theta}_{n}-\theta_{n}^{\mathrm{causal}})\overset{d}{\to}\mathrm{N}(0,I_{d_{T}}).
\end{align*}
\end{theorem}
\begin{remark}
When we have a homogeneous effect $\theta_{n,i}=\theta_{n}$, we have $\theta_{n}^{\mathrm{causal,sample}}=\theta_{n}^{\mathrm{causal}}\quad\text{a.s.}$
for large enough $n$ under $X_{n,i}=\widetilde{X}_{n,i}$.
Hence, we can use the same asymptotic distribution among them.
\end{remark}
\section{Variance Estimation}
In this section, we provide a conservative network heteroskedasticity- and autocorrelation-consistent (HAC) variance estimator for $\widehat{\theta}_{n}$.
Note that even when treatments and samples are randomly assigned and drawn, dependence can persist within a $2K$-neighborhood because exposure mappings $T_{n,i}$ may share elements of $\boldsymbol{D}_{n}$ and $\boldsymbol{R}_{n}$. As a result, the variance estimator must account for this local dependence structure. However, for any pair $i, j$ with $d_n(i, j) > 2K$, the exposure mappings $T_{n,i}$ and $T_{n,j}$ are independent. When the exposure mapping is correctly specified ($\widetilde{g}=g$), the researcher can directly choose a finite $K$ based on the functional form. If there is potential misspecification in $\widetilde{g}$, $K$ should be selected conservatively, reflecting the maximum range over which the exposure mapping may induce dependence.
Define
\begin{align*}
\widetilde{\mathcal{N}}_{n}(i;s) =\{j\in \mathcal{N}_{n}:\widetilde{d}_{n}(i,j)\leq s\},
\end{align*}
which is the set of $i$'s neighborhood within $s$-distance on a sampled network $\widetilde{\boldsymbol{A}}_{n}$.
Note that $\widetilde{\mathcal{N}}_{n}(i;s)$ is a random set because $\widetilde{d}_{n}(i,j)$ is a random variable depending on $\boldsymbol{R}_n$.
On the other hand, $d_{n}(i,j)$ and $\mathcal{N}_{n}(i;s)$ are non-random.
Recall that we also have $\widetilde{d}_{n}(i,j)\geq d_{n}(i,j)$ a.s., thus, $\widetilde{\mathcal{N}}_{n}(i;s)\subseteq \mathcal{N}_{n}(i;s)$ a.s.
Let
\begin{align*}
\widehat{\varepsilon}_{n,i}&=Y_{n,i}-\widetilde{X}_{n,i}^{\prime} \widehat{\theta}_n-\widetilde{Z}_{n,i}^{\prime}\widetilde{\gamma}_n,
\end{align*}
$\Psi_{n,i}=X_{n,i}\varepsilon_{n,i}$, $\widetilde{\Psi}_{n,i}=\widetilde{X}_{n,i}\widetilde{\varepsilon}_{n,i}$, and $\widehat{\Psi}_{n,i}=\widetilde{X}_{n,i}\widehat{\varepsilon}_{n,i}$, where we define $\widetilde{\gamma}_n$ later in (ref).
By orthogonality conditions, $\sum_{i=1}^{n}\mathbb{E}\left[\Psi_{n,i}\right]=0$, $\sum_{i=1}^{n}R_{n,i}\mathbb{E}\left[\widetilde{\Psi}_{n,i}\mid\boldsymbol{R}_n\right]=0$, and $\sum_{i=1}^{n}R_{n,i}\widehat{\Psi}_{n,i}=0$.
Then, the variances of interest can be written as
\begin{align*}
\frac{1}{n\rho_n}\widetilde{\Sigma}_{n}
=&\operatorname{Var}\left(\frac{1}{\sqrt{n\rho_n}}\sum_{i=1}^{n}R_{n,i}\widetilde{X}_{n,i}\widetilde{\varepsilon}_{n,i}\mid\boldsymbol{R}_n\right)\\
=&\frac{1}{n\rho_n}\sum_{i=1}^{n}\sum_{j\in\widetilde{\mathcal{N}}_{n}(i,2K)}R_{n,i}R_{n,j}\mathbb{E}\left[\left(\widetilde{\Psi}_{n,i}-\mathbb{E}\left[\widetilde{\Psi}_{n,i}\mid\boldsymbol{R}_n\right]\right)\left(\widetilde{\Psi}_{n,j}-\mathbb{E}\left[\widetilde{\Psi}_{n,j}\mid\boldsymbol{R}_n\right]\right)^{\prime}\mid\boldsymbol{R}_n\right],
\end{align*}
and
\begin{align*}
\frac{1}{n\rho_n}\Sigma_{n}
=&\operatorname{Var}\left(\frac{1}{\sqrt{n\rho_n}}\sum_{i=1}^{n}R_{n,i}X_{n,i}\varepsilon_{n,i}\right)\\
=&\frac{1}{n\rho_n}\sum_{i=1}^{n}\sum_{j\in\mathcal{N}_{n}(i,2K)}\mathbb{E}\left[\left(R_{n,i}\Psi_{n,i}-\rho_n\mathbb{E}\left[\Psi_{n,i}\right]\right)\left(R_{n,j}\Psi_{n,j}-\rho_n\mathbb{E}\left[\Psi_{n,j}\right]\right)^{\prime}\right].
\end{align*}
We consider the following feasible estimator:
\begin{align*}
\frac{1}{N}\widehat{\Sigma}_n
=&\frac{1}{N}\sum_{i=1}^{n}\sum_{j\in\widetilde{\mathcal{N}}_{n}(i,2K)}R_{n,i}R_{n,j}\widehat{\Psi}_{n,i}\widehat{\Psi}_{n,j}^{\prime}.
\end{align*}
To show the consistency of the variance estimator, we assume an additional sparsity condition.
The assumption requires a few more notations.
Let $\delta_{n}(s;p)$ be the $p$-th sample moment of the set of $i$'s neighborhood within $s$-distance: $\delta_{n}(s;p)=n^{-1}\sum_{i=1}^n|\mathcal{N}_{n}(i;s)|^{p}$.
We also define $\mathcal{J}_n(s,m)$ as the set of quadruples $(i,j,i^{\prime},j^{\prime})$ such that $i^{\prime}$ and $j^{\prime}$ are $m$-neighbors of $i$ and $j$, respectively, and the distance between $i$ and $j$ is exactly $s$:
\begin{align*}
\mathcal{J}_n(s,m)=\left\{(i,j,i^{\prime},j^{\prime})\in\mathcal{N}_n^4:i^{\prime}\in\mathcal{N}_n(i,m),j^{\prime}\in\mathcal{N}_n(j,m), d_n(i,j)=s\right\}.
\end{align*}
\begin{assumption}
(i) $\delta_{n}(2K;2)=o(n)$.
(ii) $\sum_{s=0}^{2K}|\mathcal{J}_n(s,2K)|=o(n^2)$.
\end{assumption}
(ref) is a version of Assumptions 7c and 7d of leung2022causal. This assumption is satisfied if network links are not too dense.
\begin{theorem}
Let $\widetilde{\gamma}_n=\widehat{\gamma}_n$.
Under (ref), we have
\begin{align*}
\frac{1}{N}\widehat{\Sigma}_n
&=\frac{1}{n\rho_n}\widetilde{\Sigma}_n+\widetilde{B}_{n}+o_{p^{R}}(1),
\end{align*}
where $U_n=o_{p^{R}}(1)$ means $U_n\overset{p^R}{\longrightarrow}0$, and
\begin{align*}
\widetilde{B}_{n}
=&\frac{1}{n\rho_n}\sum_{i=1}^{n}\sum_{j\in\widetilde{\mathcal{N}}_{n}(i,2K)}R_{n,i}R_{n,j}\mathbb{E}\left[\widetilde{\Psi}_{n,i}\mid\boldsymbol{R}_n\right]\mathbb{E}\left[\widetilde{\Psi}_{n,j}\mid\boldsymbol{R}_n\right]^{\prime}.
\end{align*}
Let $\widetilde{\gamma}_n=\gamma_n^{\mathrm{causal}}+o_p(1)$. If, in addition, we assume (ref) and $\widetilde{d}_n(i,j)=d_n(i,j)$ a.s. for all $(i,j)\in\mathcal{N}_n^2$ with $R_{n,i}=1$ and $R_{n,j}=1$ and for all $n\in\mathbb{N}$, then,
\begin{align*}
\frac{1}{N}\widehat{\Sigma}_n
&=\frac{1}{n\rho_n}\Sigma_n+\widehat{B}_{n}+o_p(1),
\end{align*}
where
\begin{align*}
\widehat{B}_{n}
=&\frac{1}{n}\sum_{i=1}^{n}\sum_{j\in\mathcal{N}_{n}(i,2K)}\rho_n\mathbb{E}\left[\Psi_{n,i}\right]\mathbb{E}\left[\Psi_{n,j}\right]^{\prime}.
\end{align*}
\end{theorem}
An estimator satisfying $\widetilde{\gamma}_n=\gamma_n^{\mathrm{causal}}+o_p(1)$ is given in (ref). In general, $\widehat{\gamma}_n\neq\gamma_n^{\mathrm{causal}}+o_p(1)$, and we need a modification on $\widehat{\gamma}_n$. The condition $\widetilde{d}_n(i,j)=d_n(i,j)$ a.s. for units with $R_{n,i}=1$ and $R_{n,j}=1$ is satisfied, for example, when the network is sampled using star sampling and the exposure mapping is restricted to the first neighborhood.
(ref) implies that we can only estimate the variance up to the ones with bias terms $\widetilde{B}_{n}$ and $\widehat{B}_{n}$ since there is no hope to estimate each heterogeneous expectation consistently.
This bias is inevitable in heterogeneous treatment effect settings (abadie2020sampling; leung2020treatment; gao2023causal).
Combining this convergence and the asymptotic normality, we can estimate the variance of $\widehat{\theta}_{n}$ by
\begin{equation}
{\left(\widetilde{Q}^{XX}_{n}\right)}^{-1}\left(\frac{1}{N}\widehat{\Sigma}_n\right){\left(\widetilde{Q}^{XX}_{n}\right)}^{-1}.
\end{equation}
The above variance estimator has a problem because we cannot guarantee conservativeness.
Indeed, bias matrices $\widehat{B}_{n}$ and $\widetilde{B}_{n}$ are not necessarily positive semi-definite.\footnote{Alternatively, we can implement the randomized inference as borusyak2023nonrandom. For multidimensional $\widehat{\theta}_n$, the randomized inference do not guarantee conservativeness, too.}
Conservative guarantee modification is possible.
We can write
$(1/N)\widehat{\Sigma}_n=(1/N)\widehat{R\Psi}^{\prime}_n\widetilde{K}_n\widehat{R\Psi}_n$,
where
\begin{align*}
\widehat{R\Psi}_n&=\left(R_{n,1}\widetilde{X}_{n,1}\widehat{\varepsilon}_{n,1},\cdots,R_{n,n}\widetilde{X}_{n,n}\widehat{\varepsilon}_{n,n}\right)^{\prime},\\
\widetilde{K}_{n}&=[\mathds{1}\{\widetilde{d}_n(i, j)\leq2K\}]_{i,j}.
\end{align*}
Eigendecomposition gives $\widetilde{K}_n=\mathcal{Q}_n\Xi_n\mathcal{Q}_n^{\prime}$.
By replacing $\widetilde{K}_n$ by $\widetilde{K}_n^+=\mathcal{Q}_n\max\{0,\Xi_n\}\mathcal{Q}_n^{\prime}$ ($\max$ is taken element-wise), the variance matrix estimator
\begin{align*}
\frac{1}{N}\widehat{\Sigma}_n^+
=\frac{1}{N}\widehat{R\Psi}^{\prime}_n\widetilde{K}_n^{+}\widehat{R\Psi}_n=\frac{1}{N}\sum_{i=1}^{n}\sum_{j=1}^{n}R_{n,i}R_{n,j}\widehat{\Psi}_{n,i}\widehat{\Psi}_{n,j}^{\prime}\widetilde{K}_{n,i,j}^{+}.
\end{align*}
becomes positive semi-definite.
We also have $\widetilde{K}_n^{-}=\mathcal{Q}_n|\min\{0,\Xi_n\}|\mathcal{Q}_n^{\prime}=\widetilde{K}_n^{+}-\widetilde{K}_n$.
This modification is provided by gao2023causal. The modified variance estimator is given by
\begin{equation}
{\left(\widetilde{Q}^{XX}_{n}\right)}^{-1}\left(\frac{1}{N}\widehat{\Sigma}^{+}_n\right){\left(\widetilde{Q}^{XX}_{n}\right)}^{-1}.
\end{equation}
Define $K_n=[\mathds{1}\{d_n(i, j)\leq2K\}]_{i,j}$, and define $K_n^{+}$ and $K_n^{-}$ in a similar manner to $\widetilde{K}_n^{+}$ and $\widetilde{K}_n^{-}$.
Define
\begin{align*}
\widetilde{\delta}_{n}^{-}(2K;p)=\frac{1}{n}\sum_{i=1}^{n}\left(\sum_{j=1}^{n}|\widetilde{K}_{n,i,j}^{-}|\right)^{p},
\qquad
\delta_{n}^{-}(2K;p)=\frac{1}{n}\sum_{i=1}^{n}\left(\sum_{j=1}^{n}|K_{n,i,j}^{-}|\right)^{p},
\end{align*}
and
\begin{align*}
&|\widetilde{\mathcal{J}}_n^{-}(s,2K)|=\sum_{i=1}^{n}\sum_{j=1}^{n}\mathds{1}\{d_n(i, j)=s\}\left(\sum_{i^{\prime}=1}^{n}|\widetilde{K}_{n,i,i^{\prime}}^{-}|\right)\left(\sum_{j^{\prime}=1}^{n}|\widetilde{K}_{n,j,j^{\prime}}^{-}|\right),\\
&|\mathcal{J}_n^{-}(s,2K)|=\sum_{i=1}^{n}\sum_{j=1}^{n}\mathds{1}\{d_n(i, j)=s\}\left(\sum_{i^{\prime}=1}^{n}|K_{n,i,i^{\prime}}^{-}|\right)\left(\sum_{j^{\prime}=1}^{n}|K_{n,j,j^{\prime}}^{-}|\right).
\end{align*}
\begin{assumption}
(i) $\widetilde{\delta}_{n}^{-}(2K;1)=O_{\text{a.s.}}(1)$ and $\delta_{n}^{-}(2K;1)=O(1)$.
(ii) $\widetilde{\delta}_{n}^{-}(2K;2)=O_{\text{a.s.}}(n)$ and $\delta_{n}^{-}(2K;2)=O(n)$.
(iii) $\sum_{s=0}^{2K}|\widetilde{\mathcal{J}}_n^{-}(s,2K)|=O_{\text{a.s.}}(n^2)$ and $\sum_{s=0}^{2K}|\mathcal{J}_n^{-}(s,2K)|=O(n^2)$.
\end{assumption}
(ref) is a version of Assumptions 7b-7d of gao2023causal. The assumption is a modified version of (ref) for the eigenvalue modification.
\begin{theorem}
Let $\widetilde{\gamma}_n=\widehat{\gamma}_n$.
Under (ref), we have
\begin{align*}
\frac{1}{N}\widehat{\Sigma}_n^{+}
&=\frac{1}{n\rho_n}\widetilde{\Sigma}_n+\widetilde{B}_{n}^{+}+o_p^R(1),
\end{align*}
where
\begin{align*}
\widetilde{B}_{n}^{+}
=&\frac{1}{n\rho_n}\sum_{i=1}^{n}\sum_{j=1}^{n}R_{n,i}R_{n,j}\mathbb{E}\left[\widetilde{\Psi}_{n,i}\mid\boldsymbol{R}_n\right]\mathbb{E}\left[\widetilde{\Psi}_{n,j}\mid\boldsymbol{R}_n\right]^{\prime}\widetilde{K}_{n,i,j}^{+}\\
&+\frac{1}{n\rho_n}\sum_{i=1}^{n}\sum_{j=1}^{n}R_{n,i}R_{n,j}\mathbb{E}\left[\left(\widetilde{\Psi}_{n,i}-\mathbb{E}\left[\widetilde{\Psi}_{n,i}\mid\boldsymbol{R}_n\right]\right)\left(\widetilde{\Psi}_{n,j}-\mathbb{E}\left[\widetilde{\Psi}_{n,j}\mid\boldsymbol{R}_n\right]\right)^{\prime}\mid\boldsymbol{R}_n\right]\widetilde{K}_{n,i,j}^{-}
\end{align*}
Let $\widetilde{\gamma}_n=\gamma_n^{\mathrm{causal}}+o_p(1)$.
If, in addition, we assume (ref) and $\widetilde{d}_n(i,j)=d_n(i,j)$ a.s. for all $(i,j)\in\mathcal{N}_n^2$ with $R_{n,i}=1$ and $R_{n,j}=1$ and for all $n\in\mathbb{N}$, then,
\begin{align*}
\frac{1}{N}\widehat{\Sigma}_n^{+}
&=\frac{1}{n\rho_n}\Sigma_n+\widehat{B}_{n}^{+}+o_p(1),
\end{align*}
where
\begin{align*}
\widehat{B}_{n}^{+}
=&\frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{n}\rho_n\mathbb{E}\left[\Psi_{n,i}\right]\mathbb{E}\left[\Psi_{n,j}\right]^{\prime}K_{n,i,j}^{+}\\
&+\frac{1}{n\rho_n}\sum_{i=1}^{n}\sum_{j=1}^{n}\mathbb{E}\left[\left(R_{n,i}\Psi_{n,i}-\rho_n\mathbb{E}\left[\Psi_{n,i}\right]\right)\left(R_{n,j}\Psi_{n,j}-\rho_n\mathbb{E}\left[\Psi_{n,j}\right]\right)^{\prime}\right]K_{n,i,j}^{-}.
\end{align*}
\end{theorem}
\section{Simulation}
In this section, we conduct a simulation exercise to illustrate the potential severity of contamination bias. We focus on a case where $T_{n,i} \neq \widetilde{T}_{n,i}$ and contamination bias can arise. See (ref) for results when $T_{n,i} = \widetilde{T}_{n,i}$, where no contamination bias occurs and our inference procedure is valid under correct model specification.
In the following exercise, we use network link information from banerjee2013diffusion to simulate variables based on a real-world network structure, rather than on an artificially generated population network.
That study conducted a network survey among randomly selected respondents across 75 villages in rural southern India, where respondents were asked to name 5 to 8 contacts across 12 interaction dimensions (e.g., house visits, borrowing goods). We focus on the borrowing network among individuals, specifically whether a person borrows rice or kerosene from others.\footnote{banerjee2013diffusion also collected household-level network data; we use individual-level network data, which is sparser than the household-level networks.} To illustrate the applicability of our framework to a single large network without relying on many clusters, we focus on the largest village and use its borrowing network as the population network $\boldsymbol{A}_{n}$. Basic network statistics for this village are presented in (ref):
\begin{table}[htbp]
\caption{Network Information}
\begin{tabular}{cccc}
\hline
\hline
\textbf{Nodes} & \textbf{Edges} & \textbf{Mean Degree} & \textbf{Mean 2nd Order Degree} \\\hline
1770 & 5556 & 6.28 & 11.44 \\\hline
\end{tabular}
\caption*{ \textit{Notes}: \textbf{Nodes} reports the number of individuals in the village; \textbf{Edges} reports the number of links based on borrowing relationships; \textbf{Mean Degree} reports the mean degree; \textbf{Mean 2nd Order Degree} reports the mean count of friends-of-friends not directly connected to node $i$.}
\end{table}
In this exercise, we consider a scenario in which the true and observed exposure mappings differ. The main objective is to quantify the severity of contamination bias. Specifically, we focus on a case where there is no contamination bias at the population level, but bias can arise due to the choice of $\widetilde{g}$. We specify the exposure mapping as in (ref):
\begin{align*}
T_{n,i}&=\left(R_{n,i}D^*_{n,i},\frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}A_{n,i,j}},\frac{\sum_{j\neq i}\sum_{k\neq i,j}A_{n,i,j}A_{n,j,k}(1-A_{n,i,k})R_{n,k}D^*_{n,k}}{\sum_{j\neq i}\sum_{k\neq i,j}A_{n,i,j}A_{n,j,k}(1-A_{n,i,k})}\right)\\
&=:(D_{n,i},\text{net}_{n,i},\text{weak}_{n,i}),
\end{align*}
and $\widetilde{T}_{n,i}$ is the same as $T_{n,i}$ except that its second and third elements are replaced by
\begin{align*}
\widetilde{\text{net}}_{n,i}=\frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D_{n,j}^{*}}{\sum_{j\neq i}R_{n,j}A_{n,i,j}};\quad \widetilde{\text{weak}}_{n,i}=\frac{\sum_{j\neq i}\sum_{k\neq i,j}R_{n,j}A_{n,i,j}A_{n,j,k}(1-A_{n,i,k})R_{n,k}D^*_{n,k}}{\sum_{j\neq i}\sum_{k\neq i,j}R_{n,j}A_{n,i,j}R_{n,k}A_{n,j,k}(1-A_{n,i,k})}.
\end{align*}
For comparison, we also consider $\widetilde{T}_{n,i}^{\text{overlap}}$, which is the same as $\widetilde{T}_{n,i}$ except that each $1-A_{n,i,k}$ in $\widetilde{\text{weak}}_{n,i}$ is replaced by $1$. As discussed in (ref), due to overlaps in the second and third elements, the sample-level causal estimand based on $\widetilde{T}_{n,i}^{\text{overlap}}$ will be contaminated. In contrast, the estimands based on $T_{n,i}$ and $\widetilde{T}_{n,i}$ are not, as they are free of such overlaps and correlations.
We implement the following simulation design. First, we set individual-specific parameters as
$\theta_{n,i,(1)}\sim\text{Exponential}(1/3)$ i.i.d., $\theta_{n,i,(2)}=M_{n,i}$, $\theta_{n,i,(3)}=0$, and $\nu_{n,i}\sim N(0,2)$ i.i.d., where $M_{n,i}$ is a clustering coefficient given by $M_{n,i}=(100/n)\times \sum_{k\neq i}\left(\sum_{j\neq i,k}A_{n,i,j}A_{n,j,k}\right)^{2}$.
Specifically, we draw these $\theta_{n,i}$ and $\nu_{n,i}$ once and treat them as fixed for each Monte Carlo iteration to simulate design-based and sampling-based uncertainties.
We choose $\theta_{n,i,(2)}=M_{n,i}$ to mechanically maximize contamination bias, as $M_{n,i}$ correlates with the contamination weights appearing in (ref). The average spillover effect from $\text{net}_{n,i}$ (i.e., the average of $M_{n,i}$) is about $1/2$. We also set $\theta_{n,i,(3)}=0$ for all $i$, so any deviation from $0$ can be interpreted as contamination bias.
Given the fixed population adjacency matrix $\boldsymbol{A}_{n}$ from banerjee2013diffusion, we can calculate the population-based causal estimand $\theta_{n}^{\mathrm{causal}}$.
Next, for each iteration, we draw $D^*_{n,i}\sim \text{Bernoulli}(0.5)$ i.i.d., and $R_{n,i}\sim \text{Bernoulli}(\rho_{n})$ i.i.d. for varying sampling probabilities $\rho_{n}\in\{0.1,0.5,1.0\}$ to see the impact of sampling uncertainty on inference. For each realization of $\boldsymbol{R}_{n}$, we compute $\theta_{n}^{\mathrm{causal,sample}}$. Subsequently, using each realization of $\boldsymbol{R}_{n}$ and $\boldsymbol{D}_{n}$, we estimate $\widehat{\theta}_{n}$ from the regression $Y_{n,i}\sim \widetilde{X}_{n,i} + \widetilde{Z}_{n,i}$,
where $\widetilde{Z}_{n,i}=(R_{n,i}p_{n},p_{n}\mathds{1}\{\sum_{j\neq i}R_{n,j}A_{n,i,j}>0\},p_{n}\mathds{1}\{\sum_{j\neq i}\sum_{k\neq i,j}R_{n,j}A_{n,i,j}R_{n,k}A_{n,j,k}(1-A_{n,i,k})>0\})$, restricted to units with $R_{n,i}=1$. Finally, we compute the standard errors based on (ref) with $\widetilde{\gamma}_{n}=\widehat{\gamma}_{n}$ for $\theta^{\mathrm{causal,sample}}$ and with $\widetilde{\gamma}_{n}$ from (ref) for $\theta^{\mathrm{causal}}$, as well as the conventional Eicker-Huber-White (EHW) standard errors, which are computed from the following variance estimator:
\begin{align*}
\left(\widetilde{Q}_{n}^{XX}\right)^{-1}\left(\frac{1}{N}\sum_{i=1}^{n}R_{n,i}\widetilde{X}_{n,i}\widetilde{X}_{n,i}'\widehat{\varepsilon}_{n,i}^{2}\right)\left(\widetilde{Q}_{n}^{XX}\right)^{-1}.
\end{align*}
When computing the standard errors based on (ref), we use the observed network $\widetilde{\boldsymbol{A}}_{n}=[R_{n,i}\times R_{n,j}\times A_{n,i,j}]_{i,j}$, which is the sampled network with induced subgraph links. We repeat this process 2,000 times. The overlapping case is implemented in the same manner, except that we use $\widetilde{T}_{n,i}^{\text{overlap}}$ instead of $\widetilde{T}_{n,i}$, and the third element of $\widetilde{Z}_{n,i}$ is replaced by $p_{n}\mathds{1}\{\sum_{j\neq i}\sum_{k\neq i,j}R_{n,j}A_{n,i,j}R_{n,k}A_{n,j,k}>0\}$.
Simulation results for $\rho_{n}\in\{0.1,0.5,1.0\}$ are summarized in (ref). In Panel A, we use $\widetilde{T}_{n,i}$ whose $\widetilde{\text{weak}}_{n,i}$ does not have an overlap in $D^*_{n,j}$ for any $j$ with $\widetilde{\text{net}}_{n,i}$. In Panel B, we use $\widetilde{T}^{\text{overlap}}_{n,i}$ whose $\widetilde{\text{weak}}^{\text{overlap}}_{n,i}$ does share some $D^*_{n,j}$ with $\widetilde{\text{net}}_{n,i}$. Also note that, in both panels, the true exposure mapping is fixed to $T_{n,i}$ defined above. Hence, the population-level causal estimands $\theta_{n}^{\mathrm{causal}}$ are the same regardless of which $\widetilde{T}_{n,i}$ or $\widetilde{T}_{n,i}^{\text{overlap}}$ is used.
From Panel A of (ref) (no overlap case), we can observe that the sample-level estimand and estimator largely deviate from the population-level estimand for $\text{net}_{n,i}$. This deviation is driven not by contamination, but by the difference between $\text{net}_{n,i}$ and $\widetilde{\text{net}}_{n,i}$:
\begin{align*}
\text{net}_{n,i}=\frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}A_{n,i,j}}\neq\frac{\sum_{j\neq i}A_{n,i,j}R_{n,j}D^*_{n,j}}{\sum_{j\neq i}R_{n,j}A_{n,i,j}}=\widetilde{\text{net}}_{n,i}.
\end{align*}
When $\rho_{n}$ is small, the denominator of $\widetilde{\text{weak}}_{n,i}$ tends to be smaller than that of $\text{net}_{n,i}$, which results in a downward bias.
Because of the bias, the coverage probabilities against $\theta_{n}^{\mathrm{causal}}$ are close to $0$ with both EHW standard errors and those based on (ref), especially when $\rho_{n}$ is small.
However, as $\rho_{n}$ increases, the bias and coverage probabilities tend to improve with our proposed standard errors (ref) because the difference between $T_{n,i}$ and $\widetilde{T}_{n,i}$ becomes smaller and the standard errors are designed to be conservative. In contrast, the EHW standard errors fail to capture the dependence structure and thus severely under-cover the causal estimands as $\rho_{n}$ increases.
From Panel B of (ref) (with overlap case), we can observe a similar pattern as in Panel A when $\rho_{n}$ is small. However, a crucial difference arises when $\rho_{n}=1.0$. We can observe that $\theta_{n,(3)}^{\mathrm{causal,sample}}$ and $\widehat{\theta}_{n,(3)}$ are largely biased downward compared with $\theta_{n,(3)}^{\mathrm{causal}}$, with a magnitude similar to that of $\theta_{n,(2)}^{\mathrm{causal}}$. Since the true $\theta_{n,i,(3)}=0$ for all $i$, this bias is mainly driven by contamination, as suggested by (ref). The contamination bias is also reflected in the average absolute deviation of the estimator and the coverage probabilities against $\theta_{n,(3)}^{\mathrm{causal}}$ for both EHW standard errors and those based on (ref), resulting in under-coverage.
In summary, the simulation results in (ref) show that the deviation of $\widetilde{T}_{n,i}$ from $T_{n,i}$ can lead to severe bias and under coverage for the population causal estimands. The results also highlight the potential severity of contamination bias when there is a small overlap in elements of $\widetilde{T}_{n,i}$, whose size can be comparable to the true spillover effects. This emphasizes the importance of choice of $\widetilde{g}$ in practice and calls for caution when interpreting the results based on the linear regression framework. In the next section, we discuss whether the contamination bias is present in the real data application.
\begin{table}[htbp]
\caption{Simulation Results: $T_{n,i}\neq\widetilde{T}_{n,i}$ case}
\scalebox{0.9}{
\begin{threeparttable}
\begingroup
\begin{tabular}{lccccccccc}
\hline
\hline
\multicolumn{10}{c}{\textbf{Panel A: No Overlaps}} \\
\hline
& \multicolumn{3}{c}{\(\rho = 0.1\)} & \multicolumn{3}{c}{\(\rho = 0.5\)} & \multicolumn{3}{c}{\(\rho = 1.0\)} \\
& D & net & weak & D & net & weak & D & net & weak \\
\hline
$\theta^{\mathrm{causal}}$ & 0.348 & 0.567 & 0.0 & 0.348 & 0.567 & 0.0 & 0.348 & 0.567 & 0.0 \\
$\theta^{\mathrm{causal,sample}}$ & 0.347 & 0.153 & 0.0 & 0.348 & 0.282 & 0.0 & 0.348 & 0.567 & 0.0 \\
$\widehat{\theta}$ & 0.347 & 0.139 & 0.009 & 0.346 & 0.28 & -0.004 & 0.347 & 0.565 & -0.01 \\
$\text{SE EHW}$ & 0.163 & 0.251 & 0.475 & 0.087 & 0.113 & 0.128 & 0.068 & 0.11 & 0.111 \\
$\text{SE \eqref{eq:se_mod} } \theta^{\mathrm{causal}}$ & 0.165 & 0.263 & 0.549 & 0.102 & 0.135 & 0.175 & 0.108 & 0.191 & 0.233 \\
$\text{SE \eqref{eq:se_mod} } \theta^{\mathrm{causal,sample}}$ & 0.163 & 0.248 & 0.398 & 0.1 & 0.133 & 0.153 & 0.104 & 0.198 & 0.173 \\
$|\widehat{\theta}-\theta^{\mathrm{causal}}|$ & 0.182 & 0.47 & 0.696 & 0.08 & 0.292 & 0.159 & 0.058 & 0.141 & 0.153 \\
$|\widehat{\theta}-\theta^{\mathrm{causal,sample}}|$ & 0.18 & 0.289 & 0.696 & 0.08 & 0.124 & 0.159 & 0.058 & 0.141 & 0.153 \\
$\text{Coverage EHW }\theta^{\mathrm{causal}}$ & 0.844 & 0.56 & 0.703 & 0.908 & 0.335 & 0.797 & 0.938 & 0.775 & 0.74 \\
$\text{Coverage EHW }\theta^{\mathrm{causal,sample}}$ & 0.846 & 0.819 & 0.703 & 0.909 & 0.836 & 0.797 & 0.938 & 0.775 & 0.74 \\
$\text{Coverage \eqref{eq:se_mod} }\theta^{\mathrm{causal}}$ & 0.847 & 0.577 & 0.768 & 0.942 & 0.443 & 0.922 & 0.997 & 0.968 & 0.964 \\
$\text{Coverage \eqref{eq:se_mod} }\theta^{\mathrm{causal,sample}}$ & 0.844 & 0.813 & 0.618 & 0.939 & 0.898 & 0.87 & 0.995 & 0.971 & 0.915 \\
\hline
\multicolumn{10}{c}{\textbf{Panel B: With Overlaps}} \\
\hline
& \multicolumn{3}{c}{\(\rho = 0.1\)} & \multicolumn{3}{c}{\(\rho = 0.5\)} & \multicolumn{3}{c}{\(\rho = 1.0\)} \\
& D & net & weak & D & net & weak & D & net & weak \\
\hline
$\theta^{\mathrm{causal}}$ & 0.348 & 0.567 & 0.0 & 0.348 & 0.567 & 0.0 & 0.348 & 0.567 & 0.0 \\
$\theta^{\mathrm{causal,sample}}$ & 0.347 & 0.149 & 0.032 & 0.348 & 0.279 & 0.008 & 0.348 & 0.773 & -0.356 \\
$\widehat{\theta}$ & 0.347 & 0.135 & 0.037 & 0.346 & 0.28 & 0.0 & 0.347 & 0.783 & -0.374 \\
$\text{SE EHW}$ & 0.163 & 0.269 & 0.447 & 0.087 & 0.155 & 0.176 & 0.068 & 0.249 & 0.264 \\
$\text{SE \eqref{eq:se_mod} } \theta^{\mathrm{causal}}$ & 0.165 & 0.279 & 0.492 & 0.102 & 0.186 & 0.22 & 0.105 & 0.419 & 0.452 \\
$\text{SE \eqref{eq:se_mod} } \theta^{\mathrm{causal,sample}}$ & 0.163 & 0.265 & 0.416 & 0.1 & 0.18 & 0.204 & 0.104 & 0.394 & 0.395 \\
$|\widehat{\theta}-\theta^{\mathrm{causal}}|$ & 0.181 & 0.476 & 0.59 & 0.08 & 0.297 & 0.204 & 0.058 & 0.279 & 0.439 \\
$|\widehat{\theta}-\theta^{\mathrm{causal,sample}}|$ & 0.179 & 0.295 & 0.589 & 0.08 & 0.147 & 0.204 & 0.058 & 0.211 & 0.295 \\
$\text{Coverage EHW }\theta^{\mathrm{causal}}$ & 0.845 & 0.584 & 0.756 & 0.908 & 0.528 & 0.828 & 0.936 & 0.854 & 0.65 \\
$\text{Coverage EHW }\theta^{\mathrm{causal,sample}}$ & 0.84 & 0.845 & 0.752 & 0.91 & 0.896 & 0.828 & 0.936 & 0.928 & 0.834 \\
$\text{Coverage \eqref{eq:se_mod} }\theta^{\mathrm{causal}}$ & 0.848 & 0.597 & 0.8 & 0.938 & 0.653 & 0.918 & 0.996 & 0.987 & 0.902 \\
$\text{Coverage \eqref{eq:se_mod} }\theta^{\mathrm{causal,sample}}$ & 0.84 & 0.837 & 0.714 & 0.936 & 0.933 & 0.886 & 0.995 & 0.995 & 0.96 \\
\hline
\end{tabular}
\begin{tablenotes}
• {\it Note:} Panel A reports the results when $\widetilde{T}_{n,i}$ is used while Panel B reports the results when $\widetilde{T}_{n,i}^{\text{overlap}}$ is used. The first three rows report the averages of the population and sample-level causal estimands and the OLS estimator. The fourth and fifth rows report the averages of the EHW standard errors and our proposed standard errors based on (ref). The sixth and seventh rows report the average absolute deviations of the estimator from the two causal estimands. The last four rows report the coverage probabilities of the $95\%$ confidence intervals constructed using the EHW standard errors and the standard errors based on our proposed method (ref) for the two causal estimands.
\end{tablenotes}
\endgroup
\end{threeparttable}
}
\end{table}
\section{Empirical Illustration}
In an influential study, cai2015social conducted a large-scale network experiment in which they randomly assigned information sessions on weather insurance products to rice farmers in rural villages in China. Out of 185 randomly selected villages, all rice farmers were invited to participate, and approximately 90% agreed to attend. The researchers administered both a household survey (to gather farmer characteristics) and a network survey (to collect friendship links). In the network survey, household heads were asked to list their five closest friends with whom they discussed rice production and financial matters, which provides a star sampling network. They were allowed to list friends outside of their village.\footnote{cai2015social conducted a pilot network survey in two villages without limiting the number of friends, but found that most farmers listed five or fewer friends. We take this analysis at face value and assume that there is no concern about censoring the number of friends.}
The information sessions were conducted in two rounds (first and second) and with varying intensity (simple or intensive). Farmers were randomly assigned to one of four possible sessions. The main outcome here, $Y_{n,i}$, is a test score measuring understanding of the insurance product, taking 10 values between 0 and 1 (\textbf{test}). The treatment variable, $D_{n,i}$, indicates whether a farmer was assigned to an intensive session (\textbf{intensive}). To measure the spillover/diffusion effects of the information sessions on farmers' knowledge, the researchers focused on a subsample of farmers who were not invited in the first round and defined (i) the fraction of a farmer's friends who attended an intensive session in the first round (\textbf{net}) and (ii) the fraction of those friends' friends who attended an intensive session in the first round (\textbf{weak}).
As discussed in (ref) and the simulation section, including first-order overlaps between \textbf{net} and \textbf{weak} can significantly affect inference through induced contamination bias.\footnote{We found that cai2015social included such overlaps in their version of \textbf{weak}; see the data/do/rawnet.do file in their replication folder: \url{https://www.openicpsr.org/openicpsr/project/113593/version/V1/view;jsessionid=743ABAC8AEBB3E612D4250D02BE40429}.} Here, we empirically examine whether such overlaps make a significant difference by comparing results when these overlaps are included or excluded in \textbf{net} and \textbf{weak}. Specifically, we run the following regression for the overlap and no-overlap specifications:\footnote{Note that cai2015social specified the exposure mapping as either $(\text{intensive},\text{net})$ or $(\text{weak})$, running regressions separately. Here, we consider a hypothetical scenario where both \textbf{net} and \textbf{weak} are included in the regression simultaneously, rather than replicating their original results.}
$$
\text{test} \sim \text{intensive} + \text{net} + \text{weak} + \text{controls}.
$$
For estimation, unlike in the simulation exercise above, we use all the available villages in the sample, as done in cai2015social. We control for household characteristics, village fixed effects, and network information (degree dummy) to satisfy (ref). Standard errors are calculated via our proposed method (ref), with \(K=2\).
\begin{table}[htbp]
\caption{Regression Results for cai2015social's data}
\begin{tabular}{lccc}
\hline
\hline
& With Overlaps & No Overlaps \\
\hline
intensive & 0.0752 & 0.0734 \\
& (0.0159) & (0.0164) \\
net & 0.3110 & 0.2879 \\
& (0.0527) & (0.0500) \\
weak & -0.1511 & -0.0741 \\
& (0.0453) & (0.0383) \\
\hline
\end{tabular}
\caption*{ \textit{Notes}: The number of villages is $47$, and the total sample size is $1247$. The first and second columns report estimates with and without overlaps in first-order links between \textbf{net} and \textbf{weak}. All regressions include household characteristics, village fixed effects, and network information as controls. Standard errors, computed using our proposed method (ref) with $\widetilde{\gamma}_{n}=\widehat{\gamma}_{n}$, are reported in parentheses.}
\end{table}
(ref) reports the OLS estimator $\widehat{\theta}_{n}$ and its standard errors, both with and without overlaps in the exposure mappings. When overlaps are included, the coefficient for \textbf{net} remains largely unchanged, but the estimate for \textbf{weak} becomes substantially more negative.
Specifically, the coefficient on \textbf{weak} is statistically significant at the 95% confidence level under the overlap specification, and its magnitude nearly doubles compared to the no-overlap specification\textemdash becoming comparable in size (but opposite in sign) to that of \textbf{net}.
This highlights the risk of overstating the effect of weak connections due to contamination bias, even when the true effect may be small or absent.
This pattern in the empirical results is consistent with the simulation findings in (ref), where overlaps in the exposure mapping lead to substantial contamination bias in the estimates of \textbf{weak}, while the estimates of \textbf{net} remain largely unaffected.
Overall, this exercise highlights that correlations among elements of the exposure mapping can potentially lead to misleading assessments of causal spillover effects.
\section{Conclusion}
In this paper, we study a linear regression framework for estimating causal spillover effects in network experiments. We show that, due to contamination bias, the OLS estimator for spillover effects does not bear a causal interpretation unless the exposure mapping is free of correlation among its elements. We also develop a novel asymptotic theory for inference on causal spillover effects, allowing for explicit sampling of units and networks, as well as network dependence.
Based on our theoretical analysis and simulation/empirical exercises, we recommend that researchers follow the flowchart in (ref) when estimating causal spillover effects in network experiments using linear regression. A crucial step is to ensure that the exposure mapping is free of correlations among its elements to avoid contamination bias and to ensure a causal interpretation of the OLS estimator. If the exposure mapping implied by plausible economic theories is not free of correlations but is sufficiently discrete (e.g., binary) to satisfy the overlap condition, we suggest avoiding the OLS estimator and instead using alternative methods, such as inverse probability weighting (e.g., aronow2017estimating; leung2022causal; gao2023causal), to directly estimate the causal treatment effects.
\begin{figure}[ht]
\begin{spacing}{1}
\caption{Flowchart for Valid Inference with Linear Regression}
\scalebox{0.9}{
\begin{tikzpicture}[node distance=1.2cm and 2.7cm, scale=0.95, transform shape, every node/.style={font=}]
\node (start) [startstop] {Start};
\node (dec1) [decision, below=of start] {Covariates satisfy (ref)?};
\node (dec2) [decision, below=of dec1] {Regressors satisfy the no correlation condition in (ref)?};
\node (dec3) [decision, below=of dec2] {Exposure mapping correctly specified and relevant network information observed ((ref)(i))?};
\node (dec4) [decision, below=of dec3] {Network HAC estimator correctly applied?};
\node (stop) [startstop, below=of dec4] {Valid inference: confidence interval >95% asymptotically};
\node (proc1) [process, right=of dec1] {Include\\ required covariates};
\node (proc2) [process, right=of dec2] {Modify the exposure mapping or flag results as potentially contaminated};
\node (proc3) [process, right=of dec3] {Interpret as inference for $\theta^{\mathrm{causal,sample}}$, not $\theta^{\mathrm{causal}}$};
\node (proc4) [process, right=of dec4] {Apply correct network HAC};
\coordinate (mid12) at ($(dec1.south)!0.5!(dec2.north)$);
\coordinate (mid23) at ($(dec2.south)!0.5!(dec3.north)$);
\coordinate (mid34) at ($(dec3.south)!0.5!(dec4.north)$);
\coordinate (mid45) at ($(dec4.south)!0.5!(stop.north)$);
\draw [arrow] (start) -- (dec1);
\draw [arrow] (dec1) -- node[left]{Yes} (dec2);
\draw [arrow] (dec2) -- node[left]{Yes} (dec3);
\draw [arrow] (dec3) -- node[left]{Yes} (dec4);
\draw [arrow] (dec4) -- node[left]{Yes} (stop);
\draw [arrow] (dec1.east) -- node[above]{No} (proc1.west);
\draw [arrow] (proc1.south) |- (mid12);
\draw [arrow] (dec2.east) -- node[above]{No} (proc2.west);
\draw [arrow] (proc2.south) |- (mid23);
\draw [arrow] (dec3.east) -- node[above]{No} (proc3.west);
\draw [arrow] (proc3.south) |- (mid34);
\draw [arrow] (dec4.east) -- node[above]{No} (proc4.west);
\draw [arrow] (proc4.south) |- (mid45);
\end{tikzpicture}
}
\end{spacing}
\end{figure}
While this paper establishes a comprehensive framework for network experiments on sampled networks, several avenues for future research emerge.
First, relaxing the sampling assumptions to accommodate cluster and multi-wave designs, as well as allowing more complex assignment mechanisms, would broaden applicability. The present analysis permits assignment conditional on observed covariates but excludes matched-pair and blocked randomization.
Second, a systematic comparison between regression-based estimators and inverse-probability-weighting approaches for spillover effects in network experiments is important, but lies beyond the scope of this paper.
\addcontentsline{toc}{section}{References}
\putbib[list_ref]