EconBase
← Back to paper

Spillovers of Program Benefits with Missing Network Links

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

119,118 characters · 18 sections · 48 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Spillovers of Program Benefits with Missing Network Links

{11pt} {11pt}

abstractThe issue of missing network links in partially observed networks is frequently neglected in empirical studies. This paper addresses this issue when investigating the spillovers of program benefits in the presence of network interactions. Our method is flexible enough to account for non-i.i.d. missing links. It relies on two network measures that can be easily constructed based on the incoming and outgoing links of the same observed network. The treatment and spillover effects can be point identified and consistently estimated if network degrees are bounded for all units. We also demonstrate the bias reduction property of our method if network degrees of some units are unbounded. Monte Carlo experiments and a naturalistic simulation on real-world network data are implemented to verify the finite-sample performance of our method. We also re-examine the spillover effects of home computer use on children's self-empowered learning.\\ \noindentJEL Codes: C14, C21, C25, C26, C51. \newline \noindentKeywords: Heterogeneous treatment and spillover effects; Partially observed networks; Incoming and outgoing links; Non-i.i.d. missing; Heterogeneous missing rates.

Introduction

The importance of network interactions in shaping individuals' socio-economic outcomes has led to increasing attention in empirical studies on program evaluations oster2012determinants,banerjee2013diffusion,cai2015social,paluck2016changing,carter2021subsidies. However, a first-order practical issue that is often neglected is the presence of missing links in partially observed networks. This issue is pervasive due to various reasons, such as censored peer data, incomplete survey responses, or omitted network links due to missing information. Existing studies show that even a low missing rate can lead to a sizable bias in causal effect estimates advani2018credibly. In this paper, we study the identification and estimation of the treatment and spillover effects of a randomized program intervention using a treatment response model that allows for flexible forms of heterogeneity manski2013identification,leung2020treatment. We assume that network interactions affect the outcome through two network-based random variables (hereafter referred to as NBRVs): the network degree and the indirect exposure to treated network neighbors. We demonstrate that the identifiable spillover effects that ignore the missing links are mixtures of the true effects, with possibly negative weights and an opposite sign to the true effects.

To address the missing link problem, we employ the matrix diagonalization method of hu2008identification, which requires two observed network measures that are mutually independent conditional on the true network. In practice, network data are often collected from survey responses. Therefore, unlike traditional measurement error problems where an additional measure for the true variable is rare, we can easily construct two network measures using the incoming and outgoing links of the same observed network.\footnote{Replicated measures are widely used to deal with measurement errors in the econometrics literature LI20021,mahajan2006identification,lewbel2007estimation,hu2017identification,calvi2018women,tommasi2022identifying and in the literature studying networks goldsmith2013social,comola2017missing,chang2020estimation,li2021causal.} These two observed networks are conditionally independent as long as the reporting errors made by one unit do not depend on those made by others. Our method is flexible because it accommodates non-i.i.d. missing links. It allows for arbitrary correlations among missing links of the same unit and heterogeneous missing rates among different units, depending on their true degree values.

Using these two network measures, we demonstrate that point identification of the true effects can be achieved if degrees are bounded for all units. When degrees are unbounded for some units, our method can serve as a bias reduction approach under a restriction on the extent of network sparsity. We propose a two-step semiparametric estimation method and establish its asymptotic properties. To control data correlation under network interactions, we adopt the notion of the dependency neighborhood used in chandrasekhar2021network and restrict data dependence to be local. To assess the finite-sample performance of our method, we conduct Monte Carlo experiments and also a naturalistic simulation study using school friendship data from beuermann2015one. The results of both the Monte Carlo and naturalistic simulations indicate that our method can effectively reduce the estimation bias in realistic samples compared to the naive method that ignores the missing links. Besides, we re-examine the spillover effects of winning a home laptop lottery on children's self-empowered learning outside the classroom environment, as studied by beuermann2015one. We find that failing to account for missing network links can lead to an underestimation of the spillover effects of laptop lottery winners on digital skills of others.

Recent studies have emerged to study causal effects under network interactions using missing or misclassified network data. This paper is closely related to those employing repeated network measures \citep*[e.g.,][]{li2021causal,lewbel2022estimating}. In particular, \citet*{li2021causal} study causal effects under network interactions using an exposure mapping model. They develop a bootstrap estimation method that requires at least two observed network measures, and they assume independent noises in the observed network data and a parametric degree distribution. lewbel2022estimating examine the identification and estimation of peer effects using linear-in-means models. They employ both incoming and outgoing links to address the missing link problem and assume i.i.d. missing given individuals' covariates. Different from these two studies, our method accommodates non-i.i.d. missing links, allowing for arbitrary correlation among missing links of the same unit and heterogeneous missing rates depending not only on covariates but also on the actual degree values.

This paper is also related to, but different from, other causal effect studies that deal with missing or misclassified network links using methods other than repeated network measures. Identification and estimation of peer effects through linear-in-means models are achieved, using adjusted 2SLS estimators in local-aggregate models \citep*{liu2013estimation}, assuming the existence of a consistent estimator of network distribution boucher2020estimating,herstad2023essays, and utilizing an order-invariance condition on friends' covariates griffith2022name. In addition, chandrasekhar2011econometrics consider various network-based linear regressions and propose a two-step estimation using a graphical reconstruction process. \citet*{hardy2019estimating} identify causal effects in an exposure mapping model, assuming a parametric degree distribution and random noises in the network data. A lower bound for the spillover effects is provided by he2023measuring under the restriction of nonnegative spillovers.\footnote{There is also a separate literature that studies causal effects in the presence of network interactions when the network itself is entirely unobserved \citep*[see, e.g.,][]{depaula2023identifying,lin2021uncovering,lewbel2021socialun}.}

In addition, there is a growing literature addressing the problem of missing or misclassified network links when studying network formation or network statistics butts2003network,balachandran2017propagation,comola2017missing,thirkettle2019identification,chang2020estimation,young2020bayesian,candelaria2020identification. This paper is distinct from these studies because our method does not require modeling the network formation process, and we aim to solve the missing link problem when the target parameter is the treatment and spillover effects. \footnote{Besides, this paper also relates to the literature concerned with measurement error in discrete random variables hausman1998misclassification,abrevaya1999semiparametric,li2003modeling,cameron2004modelling,molinari2008partial,chen2009nonparametric.}

The rest of this paper is organized as follows. Section (ref) introduces the model setup and the causal effects of interest. Section (ref) characterizes the bias caused by missing network links. Section (ref) presents our proposed method and main results. Section (ref) outlines the semiparametric estimation and its asymptotic properties. Section (ref) presents the results of the Monte Carlo simulation, naturalistic simulation, and empirical analysis of the `One Laptop per Child' program using real-life network data. Section (ref) concludes. All proofs are provided in the online appendix.

Model Setup

We denote by $\mathbf{A}^*=\{A^*_{ij}\}$ the true adjacency matrix corresponding to an unweighted random network over the population $\mathcal{P}$. The network links can be either directed or undirected. Let $A^*_{ij}=1$ if unit $i$ and unit $j$ are linked, and $A^*_{ij}=0$ otherwise. As a convention, self-links are ruled out, i.e., $A^*_{ii}=0$. Let $\mathcal{N}^*_i=\{j\in\mathcal{P}:A^*_{ij}=1\}$ be the set of unit $i$'s network neighbors. Consider a treatment response model for the outcome $Y_i$:

equation[equation omitted — 149 chars of source]

where $r$ is an unknown function, $D_i$ is a binary treatment variable, $Z_i$ is a vector of covariates, and $\varepsilon_i$ is a vector of unobservable error terms. We define $S_i^*=\sum_{j\in\mathcal{N}^*_i}D_j$ as the number of treated network neighbors, and $\mathcal{T}_i^*=|\mathcal{N}^*_i|$, where $|\cdot|$ denotes the cardinality of a set, as the network degree. The treatment response model in (ref) assumes that network interactions affect the outcome through two network-based random variables (hereafter referred to as NBRVs): $S_i^*$, which measures the extent of indirect exposure to the treatment, and $\mathcal{T}_i^*$, which quantifies the popularity of each unit $i$. The same model is used by leung2020treatment and viviano2024policy to capture various forms of heterogeneous treatment and spillover effects, and tests for model specifications are developed by \citet*{athey2018exact}.

We can view model (ref) as a potential outcome model, where $(D_i,S^*_i)$ acts as a multivalued treatment, and $\mathcal{T}^*_i$ and $Z_i$ are control variables. For any random variable $B_i$, let $\Omega_B$ denote its support. For any $d\in\{0,1\}$ and $(s,n,z)\in\Omega_{S^*,\mathcal{T}^*,Z}$, let us denote

align[align omitted — 110 chars of source]

By definition, $m^*(d,s,n,z)$ captures the mean value of the outcome under a counterfactual treatment $d$ and a counterfactual number of treated network neighbors $s$, given the control variables $(\mathcal{T}^*_i,Z_i)=(n,z)$. Following leung2020treatment, we refer to $m^*$ as the conditional average structural function (CASF). Given the CASF, the treatment and spillover effects can be defined as the average response to the counterfactual manipulation of a unit's own treatment status and its indirect exposure to the treatment, respectively.

definition[Treatment and Spillover Effects]For any $d\in\{0,1\}$, $(n,z)\in\Omega_{\mathcal{T}^*,Z}$, and $s,s'\in\Omega_{S^*}$, define \begin{align*} treatment effect: &\eta^*_T(s,n,z)=m^*(1,s,n,z)-m^*(0,s,n,z),\\ spillover effect: &\eta^*_S(d,s,s',n,z)=m^*(d,s,n,z)-m^*(d,s',n,z). \end{align*}

Given $(\mathcal{T}^*_i,Z_i)=(n,z)$, the treatment effect $\eta^*_T$ measures the direct treatment effect caused by the variation in a unit's own treatment status, while fixing its exposure to the treated network neighbors. The spillover effect $\eta^*_S$ measures the indirect treatment effect caused by the variation in a unit's exposure to the treated network neighbors, while fixing its own treatment status. The analysis in this paper can be easily extended to study other forms of direct and spillover effects, as long as they are defined as functions of $m^*$.

Motivation Examples

In model (ref), we assume that network interactions affect the outcome through two NBRVs, namely, $S_i^*$ and $\mathcal{T}_i^*$, which are commonly used network statistics in empirical studies of program evaluations under network interactions. We illustrate the usefulness of our model using the examples below, where the parametrization of function $r$ is used only for illustrative purposes.

example[Diffusion of a Weather Insurance Product]cai2015social study the influence of social networks on weather insurance adoption in rural China. The authors consider a model for the treatment and spillover effects at the household level of the form $$Y_{i}=\theta_0+\theta_1D_{i}+\theta_2\frac{S^*_{i}}{\mathcal{T}^*_{i}}+\theta'_3Z_{i}+\theta'_4NetSize_{i}+\varepsilon_{i },$$ where the binary outcome $Y_{i}$ indicates whether household $i$ decides to purchase the insurance, the treatment variable $D_{i}$ takes value one if a household is randomly invited to an intensive information session that introduces a new insurance product, $\frac{S^*_{i}}{\mathcal{T}^*_{i}}$ is the fraction of treated network neighbors, $NetSize_{i}$ is a set of dummies indicating network degree values, and covariates in $Z_{i}$ include household characteristics and village fixed effects.
example[Adoption of Menstrual Cups]oster2012determinants explore the role of network interactions in technology adoption using data from a randomized allocation of menstrual cups in Nepal. Their model for the binary outcome of menstrual cup adoption can be summarized as follows: $$Y_{i}=1[\theta_0+\theta_1D_i+\theta_2h(S^*_i,\mathcal{T}^*_{i})+\theta'_3Z_i+\theta'_4NetSize_{i}>\varepsilon_i],$$ where $D_i$ indicates the randomized access to menstrual cup, $h(S^*_i,\mathcal{T}^*_{i})$ represents either the number of treated friends $S^*_{i}$ or the share $\frac{S^*_{i}}{\mathcal{T}^*_{i}}$, $NetSize_{i}$ includes dummies that control for different network degree values, and $Z_i$ is a vector of other attributes and school fixed effects.
example[Subsidies and African Green Revolution]carter2021subsidies study the spillovers of a government subsidy on Green Revolution technology adoption. They estimate the following model for the technology adoption or agricultural yields: \begin{align*} y_{it}=&\theta_0+\theta_1D_{i}*Dur_t+\theta_2D_i*After_t\\ & +\theta_3SocialTreat_i*Dur_t+\theta_4SocialTreat_i*After_t+\theta_5'Z_{it}+\theta_6'NetSize_i+\varepsilon_{it}, \end{align*} where the treatment variable $D_{i}$ is one if household $i$ won the program lottery, $Dur_t$ and $After_t$ are time dummies for during and after the subsidy period, $SocialTreat_i=1[S^*_i\geq 2]$ indicates if the household has two or more lottery winner network neighbors, $NetSize_i$ is a set of dummies for different network degree values, and $Z_{it}$ consists of time and locality fixed effects.

Treatment and Spillover Effects under True Network

Let us begin by introducing the assumptions under which the treatment and spillover effects are point identified if the true network is correctly observed. For any random variables $B_i$ and $C_i$, denote $p_{B_i}(b)$ as its probability density (or mass) function and $p_{B_i|C_i=c}(b)$ as its conditional version. Let $\perp$ denote statistical independence.

assumption\begin{itemize} • (Randomized Treatment) $D_i$ and $Z_i$ are i.i.d. across $i$, and $D_i\perp(\varepsilon_j,Z_j,\mathcal{N}^*_j)$ for $\forall i,j\in\mathcal{P}$. In addition, $p_{D_i}(1)\in(\epsilon,1-\epsilon)$ for some constant $\epsilon>0$. • (Unconfounded Network) For $\forall i,j\in\mathcal{P}$, $\varepsilon_i\perp (\mathcal{N}^*_j,Z_j)\big|\mathcal{T}^*_i,Z_i$. • (Identical Error Distribution) For $\forall i,j\in\mathcal{P}$, we have $p_{\varepsilon_i|\mathcal{T}^*_i=n^*,Z_i=z}(e)=p_{\varepsilon_j|\mathcal{T}^*_j=n^*,Z_j=z}(e)$ for any $e\in\Omega_\varepsilon$, $n^*\in\Omega_{\mathcal{T}^*}$, $z\in\Omega_Z$. \end{itemize}

Assumption (ref) (a) assumes a randomized treatment allocation and i.i.d. covariates, which are relevant for a wide range of experimental contexts athey2017randomized.\footnote{We can relax the fully randomized treatment to an unconfounded treatment that is i.i.d. conditional on a subvector of individual characteristics $\tilde{Z}_i\subseteq Z_i$. This requires an additional assumption that $\tilde{Z}_i$ does not enter the network formation process of $\mathbf{A}^*$. More detailed explanations can be found in Appendix (ref). Extensions to non-i.i.d. covariates do not provide additional insights but may introduce technical complications. Hence, we omit this discussion for simplicity.} The network unconfoundedness in condition (b) permits dependence between the network and unobservable characteristics through the degree and covariates. For instance, it allows for spillover of unobservables in the form given in Example (ref) below. While the unconfounded network rules out certain types of endogenous networks, such as those with unobserved homophily johnsson2021estimation where unobserved factors may enter both the network formation and the outcome model, it is still weaker than a fully exogenous network. The identical distribution of the error term in condition (c) ensures that the expressions of the treatment and spillover effects introduced in Definition (ref) are the same for all $i\in\mathcal{P}$.

example[Spillover of Unobservables]Suppose that the error term in the outcome equation, $\varepsilon_i$, is a scalar function of $(\sum_{j\in\mathcal{N}^*_i}e_j,\mathcal{T}^*_i,e_i)$, where $e_i$ is an unobservable i.i.d. Bernoulli error that is independent of $\mathbf{A}^*$ and $\mathbf{Z}^*=\{Z_i\}_{i\in\mathcal{P}}$, and $\sum_{j\in\mathcal{N}^*_i}e_j$ measures the spillover of unobservables. One example can be $\varepsilon_i=\frac{1}{\mathcal{T}^*_i}\sum_{j\in\mathcal{N}^*_i}e_j+e_i$. In this example, $\sum_{j\in\mathcal{N}^*_i}e_j$ given $\mathcal{T}^*_i$ and $Z_i$ is a sum of a known number of i.i.d. Bernoulli variables, and it follows a binomial distribution that depends on the network only through $\mathcal{T}^*_i$.

The proposition below demonstrates that if the true network is correctly observed, then $m^*$ and the treatment and spillover effects are all point identified.

propositionUnder Assumption (ref), we have that $$m^*(d,s,n,z)=E\left[Y_i\big|D_i=d,S^*_i=s,\mathcal{T}^*_i=n,Z_i=z\right].$$

Biased Effects under Missing Network Links

Existing methods for studying spillover effects often assume no missing links in the observed network data,\footnote{See \citet*{leung2020treatment}, \citet*{vazquez2019identification}, sanchez2022network, and \citet*{viviano2024policy} for example.} but such an assumption fails to hold in many empirical applications. Suppose we randomly draw $N$ units from the population $\mathcal{P}$ and collect their network information. Denote the observable adjacency matrix as $\mathbf{A}=\{A_{ij}\}$, where self connections are dropped. For $i=1,...,N$ and $j\in\mathcal{P}$, write the observed link as

align*[align* omitted — 36 chars of source]

where $U_{ij}\in\{0,1\}$ indicates a missing link. Specifically, $U_{ij}=1$ implies that a true link $A^*_{ij}=1$ is correctly observed, whereas $U_{ij}=0$ implies that a true link $A^*_{ij}=1$ is absent. Define $\mathcal{N}_i=\{j\in\mathcal{P}:A_{ij}=1\}$ as the set of unit $i$'s observed network neighbors. Let $\mathcal{T}_i=|\mathcal{N}_i|$ be the observed network degree and $\mathcal{S}_i=\sum_{j\in\mathcal{N}_i}D_j$ be the number of observed treated network neighbors. Assume that we can observe each sampled unit $i$'s outcome, treatment status, covariates, and treatment assignments of $i$'s observable network neighbors:

align*[align* omitted — 110 chars of source]

In the presence of missing links, the observed network is a subset of the true network for all sampled units, i.e., $\mathcal{N}_i\subseteq \mathcal{N}^*_i$. Given the observed NBRVs for each unit, we define the identifiable counterpart for the CASF $m^*$ as below:

align[align omitted — 85 chars of source]

where we replace $S^*_i$ and $\mathcal{T}^*_i$ in the definition of $m^*$ with the observed $S_i$ and $\mathcal{T}_i$. If we ignore the presence of missing links and use the identifiable $m$ to compute the treatment and spillover effects of interest, we will obtain

equation[equation omitted — 173 chars of source]

We employ the following assumptions to establish the bias of $\eta_T$ and $\eta_S$ relative to $\eta^*_T$ and $\eta^*_S$.

assumption(Nondifferential Missing Links) For $\forall i,j\in\mathcal{P}$, $D_i\perp(\varepsilon_j,Z_j,\mathcal{N}^*_j,\mathcal{N}_j)$ and $\varepsilon_i\perp\left(\mathcal{N}^*_j,\mathcal{N}_j\right)\big|\mathcal{T}^*_i,Z_i$.

Assumption (ref) requires the treatment variable to be independent of missing links, which is trivially satisfied by randomized treatment assignments. In addition, it assumes that the missing links do not contain relevant information regarding the potential outcomes, given the actual degree and individual's characteristics. Following the literature bound2001measurement, missing links that satisfy these conditions can be referred to as “nondifferential” missing links.

assumption(Identical Degree Distribution) For $\forall i,j\in\mathcal{P}$, we have \begin{itemize} • $p_{\mathcal{T}^*_i|Z_i=z}(n^*)=p_{\mathcal{T}^*_j|Z_j=z}(n^*)$ for all $n^*\in\Omega_{\mathcal{T}^*}$ and $z\in\Omega_Z$; • $p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(n)=p_{\mathcal{T}_j|\mathcal{T}^*_j=n^*,Z_j=z}(n)$ for all $n\in\Omega_{\mathcal{T}}$, $n^*\in\Omega_{\mathcal{T}^*}$, and $z\in\Omega_Z$. \end{itemize}

Assumption (ref) (a) requires the distribution of the true degree to be identical for all units with the same characteristics $Z_i$. It still allows for degree heterogeneity but may rule out the possibility of strategic network interactions, under which the network formation of one unit depends on the existing links of others.\footnote{In Appendix (ref), we present Monte Carlo simulation results for our proposed method using network data generated by incorporating strategic interactions. The results demonstrate that our method can reduce the estimation bias compared to the naive estimation which ignores missing network links, even if Assumption (ref) (a) is violated.} In Example (ref), we present a network formation model that takes into account degree heterogeneity and satisfies condition (a). Condition (b) requires the observable degree to be identically distributed across all units who have the same true degree value and covariates. In Lemma (ref) of Appendix (ref), we show that at least two types of missing links are allowed under this condition: (i) missing completely at random, where $U_{ij}$ is i.i.d. across all pairs $(i,j)$ and independent of all other variables, and (ii) missing not at random, where $U_{ij}$ can exhibit arbitrary correlations across $j$ for any given unit $i$, and the missing probability can vary with the true degree and covariates. Assumption (ref) is employed to ensure that the expressions of $m$ and the identifiable effects, $\eta_T$ and $\eta_S$, are the same for all units.

example[Identical Degree Distribution] Without loss of generality, suppose there are no covariates $Z_i$. Consider a network formation model $$ A^*_{ij}= 1[\beta_1+\beta_2(\alpha_i+\alpha_j)-d(\rho_i,\rho_j)>\zeta_{ij}],$$ where $\alpha_i$ stands for the unobserved degree heterogeneity, $d(\rho_i,\rho_j)$ is the distance between two units defined using random location variables $\rho_i$ and $\rho_j$, and $\zeta_{ij}$ is a link-specific random shock. Suppose $(\alpha_i,\rho_i)$ is i.i.d. across $i$, $\zeta_{ij}$ is i.i.d. across $(i,j)$, and $\{\alpha_i,\rho_i\}_{i\in\mathcal{P}}$ and $\{\zeta_{ij}\}_{i,j\in\mathcal{P}}$ are mutually independent. Given $(\alpha_i,\rho_i)$, for any fixed $i$, $ A^*_{ij}$ becomes a function of $(\alpha_j,\rho_j,\zeta_{ij})$ and is i.i.d. across $j$. Then, $\mathcal{T}^*_i=\sum_{j\in\mathcal{P}} A^*_{ij}$, conditional on $(\alpha_i,\rho_i)$, is a sum of $(|\mathcal{P}|-1)$ i.i.d. Bernoulli variables and follows a Binomial distribution that only depends on $(\alpha_i,\rho_i)$. Because $(\alpha_i,\rho_i)$ is identically distributed across $i$, the unconditional degree distribution is also the same for all units. Detailed proofs can be found in Lemma (ref) in Appendix (ref).
theoremUnder Assumptions (ref), (ref), and (ref), $m$ is identical for all units, and it is a mixture of $m^*$, where \begin{align*} m(d,s,n,z)=\sum\limits_{(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}} m^*(d,s^*,n^*,z)\times p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*). \end{align*}

Theorem (ref) demonstrates that $m$ is a mixture of $m^*$ with the weight $p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}$ that quantifies the severity of the missing-link problem. Given the mixture expression of $m$, we can characterize the bias in $\eta_T$ and $\eta_S$.

corollary[Biased Effects under Missing Links]Under Assumptions (ref), (ref), and (ref), \begin{align*} \eta_T(s,n,z)=&\sum\limits_{(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}} \eta^*_T(s^*,n^*,z)\times p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*),\\ \eta_S(d,s,s',n,z)=&\sum\limits_{(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}}\eta_S^*(d,s^*,s',n^*,z)\nonumber\\ & \times\left[ p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*)-p_{S^*_i,\mathcal{T}^*_i|S_i=s',\mathcal{T}_i=n,Z_i=z}(s^*,n^*)\right]. \end{align*}

Corollary (ref) reveals that the identifiable treatment effect $\eta_T(s,n,z)$ is a nonnegatively-weighted average of the true treatment effects. Consequently, if the true treatment effects are all positive or all negative, $\eta_T(s,n,z)$ will maintain the same sign. However, we cannot point identify the value of $\eta^*_T$ using $\eta_T$, except in the special case where $\eta^*_T(s,n,z)$ is homogeneous in $(s,n)$. In other words, if there exist functions $m^*_1$ and $m^*_2$ such that $m^*(d,s,n,z)=m^*_1(d,z)+m^*_2(s,n,z)$, then we can point identify the true treatment effect $\eta^*_T$ by $\eta_T$, as

align*[align* omitted — 92 chars of source]

Furthermore, the identifiable spillover effect $\eta_S(d,s,s',n,z)$ is also a weighted average of the true spillover effects, albeit with possibly negative weights that sum up to zero.\footnote{Because $\sum_{(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}}[p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*)-p_{S^*_i,\mathcal{T}^*_i|S_i=s',\mathcal{T}_i=n,Z_i=z}(s^*,n^*)]=1-1=0$, the weights in $\eta_S(d,s,s',n,z)$ sum to zero, so they cannot be all positive or all negative for every $(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}$.} Therefore, $\eta_S(d,s,s',n,z)$ may have the opposite sign of $\eta^*_S(d,s,s',n,z)$, resulting in either an upward or a downward bias. In a special case where network interactions have no impact on $Y_i$, i.e., $m^*(d,s,n,z)=m^*(d,z)$, we can point identify the true spillover effect $\eta^*_S$ using $\eta_S$, since they are both zero:

align*[align* omitted — 87 chars of source]

Main Results

This section proceeds in three steps. First, we demonstrate that the weight $p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}$ in Theorem (ref), which connects $m$ to the target CASF $m^*$, is a product of two distribution functions. Second, we introduce a sparse network assumption and recover the weight by tackling the two distribution functions separately. Finally, we discuss the identification of the CASF $m^*$.

Decomposition of the Weights

We can see that the weight $p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}$ can be rewritten as a product:

align[align omitted — 199 chars of source]

Recall that $S_i=\sum_{j\in \mathcal{N}_i}D_j$ denotes the number of observed treated network neighbors among all $\mathcal{T}_i=|\mathcal{N}_i|$ observed network neighbors. Due to the randomized treatment assignment, $S_i$ given $(\mathcal{T}_i,Z_i)$ is a sum of a given number of i.i.d. binary variables. Thus, $S_i$ given $(\mathcal{T}_i,Z_i)$ follows a distribution of $Binomial(p_{D_i}(1),\mathcal{T}_i)$ and is independent of the true degree $\mathcal{T}^*_i$. Therefore, the second term on the right-hand side of (ref) reduces to $$p_{\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}=p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}. $$ Then, we can see that the weight $p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}$ is determined by two components: $p_{S^*_i|\mathcal{T}^*_i,S_i,\mathcal{T}_i,Z_i}$, which represents the dependence between the true and observed NBRVs, and $p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}$, which captures the missing probabilities in the degree. This result is formally introduced below.

theorem(Decomposition of Weight) Under Assumptions (ref), (ref), and (ref), we have \begin{align*} &p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*) =p_{S^*_i|\mathcal{T}^*_i=n^*,S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*)\times p_{\mathcal{T}^*_i|\mathcal{T}_i=n,Z_i=z}(n^*). \end{align*}

Next, we show that the first term in the weight decomposition, $p_{S^*_i|\mathcal{T}^*_i,S_i,\mathcal{T}_i,Z_i}$, is also a binomial distribution and can be point identified using observed data. Denote $$\Delta S_i:=S^*_i-S_i\text{ and }\Delta\mathcal{T}_i:=\mathcal{T}^*_i-\mathcal{T}_i.$$ The point identification of NBRV dependence relies on two key facts. First, $p_{S^*_i|S_i,\mathcal{T}_i,\mathcal{T}^*_i,Z_i}=p_{\Delta S_i|S_i,\mathcal{T}_i,\Delta\mathcal{T}_i,Z_i}$. Second, in the presence of missing links, $\Delta S_i=\sum_{j\in \mathcal{N}^*_i \bigcap \mathcal{N}^c_i}D_j$, where $\mathcal{N}^c_i$ denotes the complement of the set $\mathcal{N}_i$ and $|\mathcal{N}^*_i \bigcap \mathcal{N}^c_i|=\Delta\mathcal{T}_i$. Thus, because of the randomized treatment allocation, $\Delta S_i$ given $(\Delta\mathcal{T}_i,Z_i)$ is a sum of a given number of i.i.d. binary variables, which follows a $Binomial(p_{D_i}(1),\Delta\mathcal{T}_i)$ distribution and is independent of $(S_i,\mathcal{T}_i)$. Therefore, these two facts together imply that $$p_{S^*_i|S_i,\mathcal{T}_i,\mathcal{T}^*_i,Z_i}=p_{\Delta S_i|S_i,\mathcal{T}_i,\Delta\mathcal{T}_i,Z_i}=p_{\Delta S_i|\Delta\mathcal{T}_i ,Z_i},$$ which remains invariant to $(S^*_i,\mathcal{T}^*_i,S_i,\mathcal{T}_i)$ as long as the two differences, $\Delta S_i$ and $\Delta\mathcal{T}_i$, are fixed. This simplification dramatically reduces the dependence structure between NBRVs. Let $\binom{n}{s}$ be the number of $s$-combinations from $n$ elements.

theorem(Point Identification of NBRV Dependence) Under Assumptions (ref), (ref), and (ref), $$p_{S^*_i|S_i=s,\mathcal{T}_i=n,\mathcal{T}^*_i=n^*,Z_i}(s^*)=\begin{cases}p_{\Delta S_i|\Delta\mathcal{T}_i=\Delta n ,Z_i}(\Delta s),&\mbox{if }s^*\leq n^*,~s\leq n,\text{ and }0\leq \Delta s\leq\Delta n,\\ 0,&\mbox{otherwise},\end{cases}$$ where $\Delta s:=s^*-s$, $\Delta n:=n^*-n$, and $p_{\Delta S_i|\Delta\mathcal{T}_i=\Delta n,Z_i}(\Delta s)= \binom{\Delta n}{\Delta s}p_{D_i}(1)^{\Delta s}p_{D_i}(0)^{\Delta n-\Delta s}. $

Since the treatment probability $p_{D_i}$ is identifiable from the data, $p_{S^*_i|\mathcal{T}^*_i,S_i,\mathcal{T}_i,Z_i}$ is point identified based on Theorem (ref). Theorems (ref), (ref), and (ref) together imply that if we can identify the second term in the product expression of the weights, i.e., the missing probabilities in the degree $p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}$, we will be able to utilize the mixture model of $m$ and the weight to recover the target CASF $m^*$.

Matrix Diagonalization Method

In this section, we adopt the matrix diagonalization method hu2008identification and the matrix perturbation analysis stewart1990matrix to recover the missing probabilities in the degree, $p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}$. Our proposed method uses two network measures that can be easily constructed using the incoming and outgoing network links of the same observed network. This method is flexible as it can accommodate arbitrary correlations among missing links of the same unit and heterogeneous missing rates based on the true degree. Without loss of generality, let us denote $\mathcal{N}_i=\{j\in \mathcal{P}:~A_{ij}=1\}$ and $\tilde{\mathcal{N}}_i=\{j\in \mathcal{P}:~A_{ji}=1\}$ as the set of unit $i$'s observed outgoing and incoming network neighbors, respectively. Then, $\mathcal{T}_i=|\mathcal{N}_i|$ and $\tilde{\mathcal{T}}_i=|\tilde{\mathcal{N}}_i|$ denote the observed out-degree and in-degree.\footnote{In some empirical studies, we can only observe the incoming network neighbors among sampled units, i.e., $\tilde{\mathcal{N}}_i=\{j\in \{1,...,N\}:~A_{ji}=1\}$. If this is the case, $\tilde{\mathcal{N}}_i$ can still be defined as $\tilde{\mathcal{N}}_i=\{j\in \mathcal{P}:~A_{ji}=1\}$, as we have $A_{ji}=0$ for all $j\not\in\{1,...,N\}$.} We assume that the support of the true and observed degrees is the set of non-negative integers $\{0,1,2,...\}$.\footnote{Our method can also be applied to cases with no isolated nodes.}

Below, we introduce assumptions that are required for implementing the matrix diagonalization method. These assumptions are strengthened versions of those in hu2008identification, modified to accommodate potentially unbounded network degrees. We start with a sparse network assumption, which defines a truncated degree support. Sparse networks are common in many social science contexts, as human beings have a limited amount of time and energy to maintain their social connections.

assumption[Sparse Network]There exists a bounded integer $0<K<\infty$, such that \begin{itemize} • $\sum_{k> K}p_{\mathcal{T}^*_i|Z_i=z}(k)\leq\triangle_K$ for $\forall z\in\Omega_Z$; • there exists a $\delta^*=\delta^*(K)>0$ such that $p_{\mathcal{T}^*_i|Z_i=z}(k)>\delta^*$ and $p_{\mathcal{T}_i|Z_i=z}(k)>\delta^*$ for all $k=0,...,K$ and $\forall z\in\Omega_Z$. \end{itemize}

Assumption (ref) requires a sparse network such that the probability of having a degree larger than $K$ is bounded by $\triangle_K$. Theoretically, we require $K$ to be a known and bounded value and $\triangle_K$ to be sufficiently small. Thus, even if $\triangle_K$ may decrease to zero as $K$ increases, we do not require $K$ to go to infinity as the sample size increases. Note that it is possible to relax this assumption by allowing $K$ to depend on $Z=z$. However, we omit this dependence to ease the notation. If the degree is uniformly bounded for all units, i.e., $\max_{i\in \mathcal{P}}\{\mathcal{T}^*_i\}= K$, then this assumption holds with $\triangle_K=0$. See paula2018identifying, richards2020application, and hu2020binary2 for papers assuming bounded degrees. If only a limited number of units have unbounded degrees, this assumption also holds with $\triangle_K$ close to zero for a carefully chosen $K$. The latter is often satisfied in social network data. For example, graham2015methods documented that in a risk-sharing network among households, a small number of households have many links, whereas the vast majority have fewer than ten links. In this case, we can set $K=10$. A bounded $K$ is crucial for implementing the matrix diagonalization method. In practice, however, while the value of $K$ still needs to be bounded, it can vary with different sample sizes.

assumption[Exclusion Restriction] $\mathcal{T}_i\perp\tilde{\mathcal{T}}_i\big|\mathcal{T}^*_i,Z_i$.

Assumption (ref) states that $\tilde{\mathcal{T}}_i$ contains no extra information of $\mathcal{T}_i$ beyond what the actual degree $\mathcal{T}^*_i$ already provides. Intuitively, the exclusion restriction is satisfied by the incoming and outgoing degrees, if the missing links of one unit, $\mathbf{U}_i=\{U_{ij}\}_{j\in\mathcal{P}}$, are conditionally independent of the missing links of others, $\mathbf{U}_k=\{U_{kj}\}_{j\in\mathcal{P}}$ for all $k\neq i$, given $(\mathbf{A}^*,\mathbf{Z})$ where $\mathbf{Z}=\{Z_i\}_{i\in\mathcal{P}}$. For example, if the missing links are caused by under-reporting or by dropping links to friends with misspelled names, Assumption (ref) is satisfied if the under-reporting or spelling error made by one unit does not depend on those made by others. Nevertheless, it does not require the missing links of the same unit to be independent of each other, and it allows the probability of $\mathbf{U}_i$ to vary with the true degree $\mathcal{T}^*_i$ and $Z_i$. Therefore, this assumption can accommodate arbitrary correlation among missing links of the same unit and heterogeneous missing rates across units. See Lemma (ref) in Appendix (ref) for further illustration.\footnote{There are scenarios where Assumption (ref) may not hold. For example, when constructing networks, researchers may only include network links among sampled units or within certain geographic boundaries (e.g., schools, villages, etc.). In the former case, $U_{ik}=U_{jk}=0$ with probability one if unit $k$ is not sampled. In the latter case, $U_{ij}=U_{ji}=0$ with probability one if units $i$ and $j$ are not located within the same boundaries. In these cases, $\mathbf{U}_i$ and $\mathbf{U}_j$ are not independent, even when conditioned on $(\mathbf{A}^*,\mathbf{Z})$, which violates the exclusion restriction.}

Given the network sparsity in Assumption (ref), we can focus on the truncated degree support, $\{0, ..., K\}$. Recall that our target in this section is the missing probabilities in the degree, $p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}$. Denote by $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z}$ a $(K+1)\times (K+1)$ matrix that consists of all these probabilities in the truncated degree support:

align*[align* omitted — 370 chars of source]

Define $$\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}=\{p_{\mathcal{T}_i|\mathcal{T}^*_i=l,Z_i=z}(k)\}_{k,l=0,...,K}\text{ and }\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}=\{p_{\tilde{\mathcal{T}}_i|\mathcal{T}^*_i=l,Z_i=z}(k)\}_{k,l=0,...,K}$$ in the same way as $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}$. Let $diag(v)$ be a diagonal matrix with the elements from vector $v$ on the principal diagonal. Define the following $(K+1)\times (K+1)$ matrices:

equation*[equation* omitted — 734 chars of source]

where $\varpi:\Omega_Y\mapsto\mathbb{R}$ represents a user-specified function, to which we will impose additional restrictions in subsequent assumptions.\footnote{For example, the user-specified function can be $\varpi(y)=y$ (mean), $\varpi(y)=(y-E[Y_i])^2$ (variance), or $\varpi(y)=1[y\leq y_0]$ (quantile) for some given $y_0$. We omit the dependence of $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}$ and $\mathbf{T}_{Y|\mathcal{T}^*,Z=z}$ on $\varpi$ to ease the notation.} Define two $(K+1)\times 1$ vectors

align*[align* omitted — 215 chars of source]

From Bayes' theorem, we can write the target matrix $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}$ as

align[align omitted — 209 chars of source]

where $\mathbf{T}_{\mathcal{T}|Z=z}$ is identifiable because it is a matrix of distribution functions of observables. In what follows, we explore the identification of $\mathbf{T}_{\mathcal{T}^*|Z=z}$ and $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$.

Let $\mathbf{\Delta}_{K}$ be a matrix (or vector) with all its entries being $O(\triangle_K)$. Note that $\mathbf{\Delta}_{K}$ may stand for different matrices (or vectors) at different places. Given Assumption (ref) (network sparsity) and Assumption (ref) (exclusion restriction), applying the law of iterated expectation, we can show that

align[align omitted — 662 chars of source]

where the three identifiable matrices on the left-hand side of (ref) to (ref) differ from the first terms on the right-hand side by $\mathbf{\Delta}_{K}$. Apparently, the differences arise from the ignorance of degree values larger than $K$. Based on the matrix perturbation theory,\footnote{See Lemma (ref) for the matrix perturbation theory.} if all three matrices in the first term on the right-hand side of (ref) are invertible with bounded inverses, then $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ is also invertible, and its inverse satisfies

align[align omitted — 274 chars of source]

By post-multiplying $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}$ in (ref) by $\mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ in (ref), we can get

align[align omitted — 291 chars of source]

It then follows from (ref) and the properties of diagonalizable matrix that the normalized eigenvectors (whose entries sum to one) of $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}\times \mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ approximate the columns of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ with an approximation error of order $\triangle_{K}$. If we can further identify the ordering of these eigenvectors, then $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ can be approximated with an error of order $\triangle_{K}$. For the other unknown matrix in (ref), $\mathbf{T}_{\mathcal{T}^*|Z=z}$, pre-multiplying both sides of (ref) by $\mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ yields the following equation:

align[align omitted — 159 chars of source]

Given that $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ can be approximated using identifiable matrices and $\mathbf{F}_{\mathcal{T}|Z=z}$ is directly identifiable from the data, it follows from (ref) that $\mathbf{T}_{\mathcal{T}^*|Z=z}$ is also approximated with an error of order $\triangle_{K}$. Consequently, the missing probabilities in the degree, $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}$ in (ref), can be approximated.

Next, in Assumptions (ref) to (ref), we formalize the necessary assumptions for the matrix diagonalization method outlined above. Let $\underline{\sigma}(\mathbf{B})$ denote the smallest singular value of a matrix $\mathbf{B}$.

assumption[Invertibility]For some $\delta=\delta(K)>0$, we have $\underline{\sigma}(\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z})>\delta$, $\underline{\sigma}(\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z})>\delta$, and $\underline{\sigma}\left(\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}\right)>\delta$ for all $z\in\Omega_Z$.

Assumption (ref) requires that the choice of $K$ ensures the invertibility of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$, $\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}$, and $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$, as well as the boundedness of their inverses. Since $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ is identifiable using data on the observed degrees $\mathcal{T}$ and $\tilde{\mathcal{T}}$, this assumption can be partially verified for the chosen value of $K$ by checking the rank and the smallest singular value of $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$.

assumption[Eigen-decomposition]For some $e_Y>0$, we have $\sup\limits_{n^*\in\Omega_{\mathcal{T}^*},z\in\Omega_Z}|E[\varpi(Y_i)|\mathcal{T}^*_i=n^*,Z_i=z]|<e_Y$. In addition, for all $z\in\Omega_Z$, if $n\neq n'$, then $$E[\varpi(Y_i)|\mathcal{T}^*_i=n,Z_i=z]\neq E[\varpi(Y_i)|\mathcal{T}^*_i=n',Z_i=z].$$

Assumption (ref) rules out duplicated eigenvalues of the diagonalizable matrix $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{T}_{Y|\mathcal{T}^*,Z=z}\times \mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z=z}$, ensuring that its eigenvalues and eigenvectors are differentiable functions of the matrix itself. This smoothness condition further guarantees that the normalized eigenvectors of $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}\times \mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ approximate the columns of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ with an error of order $\triangle_K$. Note that if network interactions have no impact on $Y_i$, then $m^*(d,s,n,z)=m^*(d,z)$, and $E[\varpi(Y_i)|\mathcal{T}^*_i,Z_i]$ becomes degenerate in $\mathcal{T}^*_i$, causing Assumption (ref) to fail. Fortunately, in this case, the matrix diagonalization method is not needed, as $m^*$ can be point identified using its identifiable counterpart $m$. This is because $m^*(d,s,n,z)=m(d,s,n,z)=m^*(d,z)$ by Theorem (ref). Then, a test for whether $m(d,s,n,z)$ depends on $(s,n)$ can be used to verify Assumption (ref) and determine the necessity of the matrix diagonalization method. In Example (ref), we discuss another possible test for Assumption (ref) within commonly used network effect models.

assumption[Order of Eigenvectors] Any one of the following conditions holds for all $n^*\in\{0,...,K\}$ and $z\in\Omega_Z$. \begin{itemize} • $p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(n^*)>p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(n)$ for any $n\neq n^*$. • $p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(0)$ is strictly monotone in $n^*$ and the direction is known. • $E[\varpi(Y_i)|\mathcal{T}^*_i=n^*,Z_i=z]$ is strictly monotone in $n^*$ and the direction is known. \end{itemize}

Assumption (ref) is used to identify the order of the columns of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$. Note that any one of the conditions in Assumption (ref) is sufficient for identifying the order. Condition (a) implies that if the $l$-th entry of a column of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ is its largest entry, then this column is the $l$-th column of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$. Under condition (b) and a decreasing order of $p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(0)$ in $n^*$, if the first entry of a column of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ is the $l$-th largest among the first entries of all its columns, then this column is the $l$-th column. Condition (c) imposes an order on the eigenvalues of the diagonalizable matrix $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{T}_{Y|\mathcal{T}^*,Z=z}\times \mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z=z}$, which also implies the order of its eigenvectors.

Researchers should carefully choose the appropriate condition in Assumption (ref) based on the specific context. Conditions (a) and (b) assume that the observable degree is informative about the true degree. Lemma (ref) in Appendix (ref) provides sufficient conditions for both (a) and (b). For instance, condition (a) holds if more than half of the units have no missing links,\footnote{Similar restrictions are widely used in the measurement error literature battistin2011misclassified,battistin2014misreported,chen2011nonlinear,huSchennach2008instrumental,lewbel2007estimation,mahajan2006identification.} and it also holds under weaker conditions. Condition (b) requires that the probability of having zero observed degree strictly decreases as the true degree increases. Condition (c) imposes a shape restriction on $\varpi(Y_i)$ and is satisfied in commonly used network effect models (see Example (ref)).

exampleConsider a simplified linear-in-means model with no endogenous peer effects and no covariates: $Y_i=\theta_1D_i+\theta_2\frac{S^*}{\mathcal{T}^*_i}+\theta_3\mathcal{T}^*_i+\varepsilon_i$. We also assume no isolated units in the true network for simplicity. Let us set $\varpi(y)=y$. Lemma (ref) in Appendix (ref) shows that \begin{align*} E[Y_i|\mathcal{T}^*_i=n]=&(\theta_1+\theta_2)p_{D_i}(1)+\theta_3n,\\ E[Y_i|\mathcal{T}_i=n]=&(\theta_1+\theta_2)p_{D_i}(1)+\theta_3\sum_{n^*\in\Omega_{\mathcal{T}^*}}n^*p_{\mathcal{T}^*_i|\mathcal{T}_i=n}(n^*). \end{align*} If $\theta_3=0$, then both $E[Y_i|\mathcal{T}^*_i=n]$ and $E[Y_i|\mathcal{T}_i=n]$ remain invariant to $n$, and Assumption (ref) is violated. Thus, Assumption (ref) can be verified by a test on whether $E[Y_i|\mathcal{T}_i=n]$ depends on $n$. In addition, we can see that $E[Y_i|\mathcal{T}^*_{i}=n]$ is strictly increasing in $n$ if $\theta_3>0$ and strictly decreasing in $n$ if $\theta_3<0$, so that condition (c) in Assumption (ref) holds as long as $\theta_3\neq0$.

Denote by $\mathbf{F}^a_{\mathcal{T}|\mathcal{T}^*,Z=z}$ a matrix whose columns are the normalized eigenvectors of $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}\times \mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$, in the order implied by Assumption (ref). Based on (ref), we can show that $\mathbf{F}^a_{\mathcal{T}|\mathcal{T}^*,Z=z}$ differs from $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ by a term of order $\triangle_K$. Based on (ref), let

align*[align* omitted — 227 chars of source]

be the approximation for $\mathbf{F}_{\mathcal{T}^*|Z=z}$ and $\mathbf{T}_{\mathcal{T}^*|Z=z}$, respectively. Replacing $\mathbf{T}_{\mathcal{T}^*|Z=z}$ and $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ in (ref) with their approximations, we can get an approximation for $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}$:

align*[align* omitted — 186 chars of source]

The theorem below shows that the difference between the target matrix, the missing probabilities in the degree $\mathbf{F}_{\mathcal{T}^*|\mathcal{T}, Z=z}$, and its approximation $\mathbf{F}^a_{\mathcal{T}^*|\mathcal{T}, Z=z}$, is bounded by $\triangle_K$.

theoremSuppose Assumptions (ref) and (ref) (b) hold for both $\mathcal{N}_i$ and $\tilde{\mathcal{N}}_i$. Under Assumptions (ref)-(ref), we have $\sup\limits_{z\in\Omega_Z}\left\|\mathbf{F}^a_{\mathcal{T}^*|\mathcal{T}, Z=z}-\mathbf{F}_{\mathcal{T}^*|\mathcal{T}, Z=z}\right\|=O(\triangle_K).$

Identification of the CASF

In this section, we proceed to the identification of the CASF $m^*$ using $m$ and the weight. Let us first introduce some notations. Denote $\mathcal{G}_i=(S_i,\mathcal{T}_i)\in\Omega_{\mathcal{G}}$ and $\mathcal{G}^*_i=(S^*_i,\mathcal{T}^*_i)\in\Omega_{\mathcal{G}^*}$. Focusing on the truncated degree support, we rank the possible values of $\mathcal{G}_i$ and $\mathcal{G}^*_i$ according to the lexicographical order of integers. For $\mathfrak{g}_k=(s,n)$, define $\{\mathfrak{g}_0,\mathfrak{g}_1,...,\mathfrak{g}_{K_\mathcal{G}}\}$ as follows:

equation[equation omitted — 296 chars of source]

For $d\in\{0,1\}$ and $z\in\Omega_Z$, define two $(K_\mathcal{G}+1)\times 1$ vectors $\mathbf{M}_{d,z}$ and $\mathbf{M}_{d,z}^*$ as

equation[equation omitted — 307 chars of source]

We denote $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$ as a $(K_\mathcal{G}+1)\times (K_\mathcal{G}+1)$ matrix that consists of the weights on the truncated degree support

align*[align* omitted — 474 chars of source]

If $m^*$ is bounded, based on Theorem (ref) and the sparse network assumption, we can show that

align[align omitted — 160 chars of source]

Given the lexicographical order of $\mathfrak{g}_{k}$ and the presence of missing links, it is easy to see that $p_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g}_k,Z_i=z}(\mathfrak{g}_l)=0$ for any $k > l$, so that $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$ is a lower triangular matrix. In addition, all its diagonal elements are strictly positive and bounded away from zero under Assumption (ref). Therefore, $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$ is invertible with a bounded inverse. Below, we introduce the notation for the approximation of $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$. Recall that by Theorem (ref), we have

align*[align* omitted — 148 chars of source]

where $p_{ S^*_i|\mathcal{T}^*_i, \mathcal{T}_i,Z_i=z}$ is a binomial distribution and is point identified as shown in Theorem (ref). Let $p^a_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i=z}$ stand for the element of the matrix $\mathbf{F}^a_{\mathcal{T}^*|\mathcal{T},Z=z}$ in Theorem (ref). Define

align*[align* omitted — 152 chars of source]

Stacking all $p^a_{\mathcal{G}^*_i|\mathcal{G}_i,Z_i=z}$ into a matrix, we can get an approximation for $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$:

align[align omitted — 202 chars of source]

According to the matrix perturbation theory, for a sufficiently small $\triangle_K$, given that $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$ is invertible with a bounded inverse, its approximation $\mathbf{F}^a_{\mathcal{G}^*|\mathcal{G},Z=z}$ is also invertible with a bounded inverse. Then, based on (ref), the theorem below shows that pre-multiplying $\mathbf{M}_{d,z}$ by the inverse of $\mathbf{F}^{a'}_{\mathcal{G}^*|\mathcal{G},Z=z}$ gives us an approximation of $\mathbf{M}_{d,z}^*$.

theoremSuppose $m^*(\cdot)$ is uniformly bounded in the support of its argument. If assumptions in Theorem (ref) hold, then we have $$ \sup_{z\in\Omega_Z,d=0,1}\Big\|\left(\mathbf{F}^{a'}_{\mathcal{G}^*|\mathcal{G},Z=z}\right)^{-1}\times \mathbf{M}_{d,z}-\mathbf{M}_{d,z}^*\Big\|=O(\triangle_K). $$

As a result of Theorem (ref), if the true degree is uniformly bounded by $K$ for all units, we have $\triangle_K=0$ and the CASF is point identified.

corollary[Point Identification of CASF with Bounded Degree]Under assumptions in Theorem (ref), if there exists some integer $0<K<\infty$ such that $\triangle_K=0$ holds, then $\mathbf{M}_{d,z}^*$ is point identified for all $z\in\Omega_Z$ and $d=0,1$.

Some final remarks are in order. First, if no $K$ exists such that $\triangle_K=0$, it is still possible to approximate $m^*$ in the truncated degree support as long as $\triangle_K$ is sufficiently small. In this case, our method can yield less biased estimates of the effects of interest compared to the naive estimation method that ignores missing links. Second, researchers should carefully choose the value of $K$ to ensure that (i) $\triangle_K$ is sufficiently small, and (ii) the rank condition of $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$, $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ and $\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}$ assumed in Assumption (ref) is satisfied. In finite samples, the approximation accuracy for $m^*$ depends on both the value of $\triangle_K$ and the credibility of the rank condition. Clearly, a larger $K$ results in a smaller $\triangle_K$ but a less credible rank condition. In practice, researchers can implement our proposed method using different values of $K$.

Estimation and Inference

In this section, we propose a two-step semiparametric estimation method for the CASF $m^*$ and present its asymptotic properties. All technical details are left to Appendix (ref).

Two-Step Estimation Method

Our estimation procedure consists of two steps. First, we obtain estimators for the weights by estimating the dependence of NBRVs and the missing probabilities in the degree using a kernel estimation approach. Second, we parameterize $m^*$ and apply a least-square estimation by plugging in the first-step estimators of the weights. Imposing parametric structures on $m^*$ still allows for flexible heterogeneity in the treatment and spillover effects, which can be captured through interactions between variables and their polynomials

\noindentStep 1. Kernel Estimation for the Weights. In the first step, we present a kernel estimation method to obtain estimators of the weights. Recall that the estimation for the NBRV dependence requires estimating $p_{D_i}$, and the estimation for the missing probabilities in the degree using the matrix diagonalization method requires estimating $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}$, $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ and $\mathbf{F}_{\mathcal{T}|Z=z}$ in (ref), (ref), and (ref).\footnote{The estimated eigenvectors in the matrix diagonalization method may contain complex values. As mentioned in hu2008identification, since all the latent probabilities and densities are real and positive, we can take the real part of the estimated eigenvectors, and the probability of getting a complex value goes to zero as the sample size increases.} Let $\gamma=\gamma(z)=[\gamma_1(z),\gamma_2(z),\gamma_3(z),\gamma_4(z)]'$ be a vector that consists of all the elements required to estimate the weights, where

equation[equation omitted — 630 chars of source]

Let $W_i=(W^{c'}_i,W^{d'}_i)'\in\Omega_{W^c}\times\Omega_{W^d}$ and $W_i\subseteq(\mathcal{T}_i,\tilde{\mathcal{T}}_i,D_i,Z_i)$ denote a vector of observable variables that will be used to compute the kernel estimator for $\gamma$, where $W^c_i:=(W^c_{i,1},...,W^c_{i,Q})'$ is a $Q\times1$ vector of continuous variables in $W_i$ (if any), and $W^d_i$ contains discrete variables in $W_i$. For example, for $\gamma_2(z)$, $W_i=(\mathcal{\mathcal{T}}_i,\tilde{T}_i,Z_i)$. For $\forall w=(w^{c'},w^{d'})'\in\Omega_{W^c,W^d}$, we define the estimator for $p_{W_i}(w)$ and $E[\varpi(Y_i)|W_i=w]$ as below:

equation[equation omitted — 271 chars of source]

where $\hat{p}^{ker}_i(w):=\frac{1}{h^{Q}}\prod_{q=1}^{Q}\kappa\left(\frac{W^c_{i,q}-w^c_q}{h}\right)1\left[W^d_i=w^d\right]$, with a bandwidth $h>0$ and a univariate kernel function $\kappa(\cdot)$.\footnote{ A data-driven method for bandwidth selection is possible but it is not the focus of this paper.} Then, the estimator for $\gamma$, denoted by $\hat{\gamma}_N$, can be obtained by replacing the conditional means and probabilities in (ref) with their sample analogs in (ref). Given $\hat{\gamma}_N$, we can estimate $\mathbf{F}^a_{\mathcal{G}^*|\mathcal{G},Z=z}$ defined in (ref). Let $vec(B)$ be the vectorization of a matrix $B$. Define

align[align omitted — 267 chars of source]

where each element in $\hat{\mathbf{F}}^a_{\mathcal{G}^*|\mathcal{G},Z=z}$ is obtained by $$\hat{p}^a_{\mathcal{G}^*_i|\mathcal{G}_i,Z_i=z}=\hat{p}_{ S^*_i|\mathcal{T}^*_i, \mathcal{T}_i,Z_i=z}\times \hat{p}^a_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i=z},$$ with $\hat{p}_{ S^*_i|\mathcal{T}^*_i, \mathcal{T}_i,Z_i=z}=\hat{p}_{\Delta S_i|\Delta\mathcal{T}_i,Z_i=z}\sim Binomial(\hat{p}_{D_i}(1),\Delta\mathcal{T}_i)$ and $\hat{p}^a_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i=z}$ being estimated by applying the matrix diagonalization method using $\hat{\gamma}_N$, as outlined in Section (ref). We suppress the argument $z$ in $\gamma$, $\phi$, $\hat{\gamma}_N$, and $\hat{\phi}_N$ for notation simplicity, unless otherwise mentioned. Let $\gamma^0$ and $\phi^0$ be the true value for $\gamma$ and $\phi$.

\noindentStep 2. Semiparametric Estimation for the CASF. For any parameter $\beta$, denote $d_\beta=dim(\beta)$. In the second step, we parameterize $m^*(\cdot)=m^*(\cdot;\theta)$ to be a known function up to an unknown parameter $\theta\in\Theta\subseteq\mathbb{R}^{d_\theta}$, and we estimate $\theta$ using a plug-in estimator. Recall $\mathcal{G}_i=(S_i,\mathcal{T}_i)\in\Omega_{\mathcal{G}}$ and $\mathcal{G}^*_i=(S^*_i,\mathcal{T}^*_i)\in\Omega_{\mathcal{G}^*}$. Denote $X^*_i=(D_i,\mathcal{G}^*_i,Z_i)'$ and $X_i=(D_i,\mathcal{G}_i,Z_i)'$. Let $x_j=(d,\mathfrak{g}_j,z)$ with $j=0,...,K_{\mathcal{G}}$. Given the parameterization of $m^*$, let us rewrite $\mathbf{M}_{d,z}^*$ in (ref) as

align*[align* omitted — 123 chars of source]

Proposition (ref) implies the existence of some true value $\theta^0\in \Theta$ such that

align[align omitted — 86 chars of source]

Without loss of generality, we assume that $\theta^0$ is the unique solution to Equation (ref).\footnote{It rules out the existence of two different pairs, $(\dot{m}^{*},\dot{\theta}^{0})$ and $(\ddot{m}^{*},\ddot{\theta}^0)$, that satisfy $m^*(\cdot)=\dot{m}^{*}(\cdot;\dot{\theta}^{0})=\ddot{m}^{*}(\cdot;\ddot{\theta}^{0})$. It ensures the parametric point-identification of $\theta^0$ if the true NBRVs are observable.} However, this moment condition cannot be used to estimate $\theta^0$ because $X^*_i$ contains unobservable NBRVs. Fortunately, we have $m(X_i;\theta^0)=E\left[Y_i|X_i\right]$ by definition of $m$ in (ref), where, for $x=(d,\mathfrak{g},z)\in \Omega_X$ and the true weights $p^0_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g},Z_i=z}(\cdot)$, we have that $\theta$ enters $m$ through the mixture model:

align[align omitted — 203 chars of source]

Thus, we can obtain a moment condition based on observed NBRVs $$E\left[Y_i-m(X_i;\theta^0)\big|X_i\right]=0.$$ Let $\tau_i=1[X_i\in\mathbf{X}]$ be a fixed trimming indicator for the observed NBRVs in the truncated support, where $\mathbf{X}=\{x=(d,\mathfrak{g},z)\in\Omega_X:~\mathfrak{g}\in\{\mathfrak{g}_0,...,\mathfrak{g}_{K_{\mathcal{G}}}\}\}$. Then, we can obtain an unconditional moment condition

align*[align* omitted — 69 chars of source]

Based on the unconditional moment equation, we define the population objective function $\mathcal{L}^0_{\mathbb{P}}(\theta)$ and its minimizer $\theta^0$ as follows:

align[align omitted — 231 chars of source]

where, under the full rank condition on the Hessian matrix of $\mathcal{L}^0_{\mathbb{P}}(\theta)$ introduced later, $\theta^0$ is the unique solution to the minimization problem. Due to the possibility of unbounded degree, in the estimation for $\theta^0$, we need to further replace $m(x;\theta)$ in (ref), defined in (ref) on the whole support $\Omega_{\mathcal{G}^*}$, with its approximation $m^a(x;\theta,\phi)$ defined on the truncated support $\{\mathfrak{g}_0,...,\mathfrak{g}_{K_\mathcal{G}}\}$, where

align*[align* omitted — 226 chars of source]

Given $m^a(x;\theta,\phi)$, let us define $\theta^a\in\Theta$ to be the pseudo-true value that solves the moment equation $E\left\{\tau_i\left[Y_i-m^a\left(X_i;\theta,\phi^a\right)\right]\right\}=0$ with $\phi^a=vec\left(\mathbf{F}^a_{\mathcal{G}^*|\mathcal{G},Z=z}\right)$. Then, we have

align*[align* omitted — 226 chars of source]

where, under the full rank condition of the Hessian matrix of $\mathcal{L}^a_{\mathbb{P}}(\theta,\phi^a)$ introduced later, $\theta^a$ is the unique solution to the minimization problem. Note that the pseudo-true parameter $\theta^a$ may differ from the true value $\theta^0$ for any nonzero $\triangle_K$, and its value may depend on $K$. We omit this dependence for notation simplicity. Then, the plug-in estimator $\hat{\theta}_N$ is defined as the minimizer of the sample objective function:

align*[align* omitted — 235 chars of source]

where we replace $\phi^a$ with its estimator $\hat{\phi}_N=vec\left(\hat{\mathbf{F}}^a_{\mathcal{G}^*|\mathcal{G},Z=z}\right)$.

Asymptotic Properties

In this section, we discuss the asymptotic properties of our proposed estimator. Additional regularity assumptions are provided in Appendix (ref). For any vector $a\in\mathbb{R}^p$, let $\|a\|$ be its Euclidean norm and $\|a\|_\infty=\max_{1\leq k\leq p}|a_k|$. For a matrix $\mathbf{B}$, let $\|\mathbf{B}\|=[tr(\mathbf{B}'\mathbf{B})]^{1/2}$ be the entry-wise matrix norm. Let us partition the index set of all sampled units into $q_N$ mutually exclusive clusters. Denote these clusters as $\mathbb{S}_{1},...,\mathbb{S}_{q_N}$, where $\cup_{1\leq k\leq q_N}\mathbb{S}_k=\{1,...,N\}$. Let $\tilde{W}_i=(Y_i,X_i')'\in\Omega_{\tilde{W}}$ represent a vector of observable variables, including the outcome. For any generic measurable function $b:\Omega_{\tilde{W}}\mapsto\mathbb{R}^{d_b}$, denote the within-cluster correlation as

align[align omitted — 126 chars of source]

To control data correlation under network interactions, we introduce a modified dependency neighborhood assumption from chandrasekhar2021network in Assumption (ref) below, which restricts the data dependence to be local.\footnote{The literature on inference using network data is growing rapidly hudgens2008toward,leung2020dependence. Our assumption on data dependence is similar to those that limit data dependence to be weak or local kojevnikov2019limit,leung2019causal.} Let $\bar{r}_N=\max_{1\leq k\leq q_N}|\mathbb{S}_{k}|$ be the size of the largest cluster.

assumption$\bar{r}_N=O(1)$ is a bounded value. For any measurable function $b:\Omega_{\tilde{W}}\mapsto\mathbb{R}^{d_b}$, $$\Big\|\sum_{k=1}^{q_N}\sum_{i\in\mathbb{S}_k,j\not\in\mathbb{S}_k}Cov\left(b(\tilde{W}_i),b(\tilde{W}_j)\right)\Big\|=o(\|\Sigma^b_N\|).$$

This assumption requires that all clusters consist of a bounded number of units. Therefore, it implies that $q_N\rightarrow\infty$ as $N\rightarrow\infty$. In addition, it assumes that the correlation between units in different clusters is not necessarily zero but is weaker than the correlation between units within the same cluster. Units in different clusters can be correlated, for example, due to network interactions, spillovers of unobservables, or spatial and other forms of dependence. Note that the bounded cluster size does not require the maximal true degree to be bounded. Network connections across clusters are allowed as long as the network sparsity in Assumption (ref) and the local dependence in Assumption (ref) hold.

Next, let us introduce a dependence coefficient analogous to the strong mixing coefficient of a stochastic process. Suppose the $q_N$ clusters can be ordered in a specific manner, based on, for example, social or geographic proximity, so that units in clusters with distant indices are less likely to be correlated with each other. Without loss of generality, we assume that this order of clusters is given by $\mathbb{S}_{1},...,\mathbb{S}_{q_N}$. It is worth noting that we do not require this order to be known to researchers. Define the dependence coefficient as

align*[align* omitted — 173 chars of source]

where $\mathcal{F}_{1}^{k-2}=\sigma(\{\tilde{W}_i,~i\in\bigcup_{1\leq l\leq k-2}\mathbb{S}_l\})$ and $\mathcal{F}_{k}^{k}=\sigma(\{\tilde{W}_i,~i\in\mathbb{S}_k\})$ for $k=1,2,...,q_N$ are two $\sigma$-fields. We use $\{\alpha_{k}\}_{k=1}^{q_N}$ to control the rate at which the dependence among clusters decays, which is crucial to establish the uniform convergence of the first-step estimators. Below, we impose some restrictions on the dependence coefficient $\alpha_{k}$. With notation abuse, let $Q$ denote the number of continuous variables in $\tilde{W}_i=(Y_i,X_i')'$.

assumption[Local Dependence]For $L_N=[N/(\ln(N)h^{Q+2})]^{Q/2}$, we have the following condition holds $$\sum_{N=1}^{\infty}\Psi_N<\infty,\text{ where } \Psi_N=L_N\left(\frac{N}{\ln(N)}\right)^{1/5}\sum_{k=1}^{q_N}\alpha^{4/5}_{k}.$$

Assumption (ref) assumes that the clusters are ordered so that units in $\bigcup_{1\leq l\leq k-2}\mathbb{S}_l$ and in $\mathbb{S}_k$ tend toward being independent as the sample size increases, allowing for nonzero but decreasing local dependence across clusters. This assumption is trivially satisfied if all clusters are mutually independent, indicating that units only form networks within each cluster. In such a case, every unit has a bounded degree. This assumption also holds when a limited number of units are correlated with others from nearby clusters, so that $\alpha_{k}$ goes to zero fast enough to ensure that $\Psi_N$ is summable. In this case, the data correlation may be caused by the network interactions of a few `star' units.

Recall that $\gamma=(\gamma_1(z), \gamma_2(z), \gamma_3(z), \gamma_4(z))'$. Let $\gamma_{jl}(z)$ with $j=1,2,3,4$ be the $l$-th element in $\gamma_j(z)$. Because $\gamma$ is a function of $z$, we define $\|\gamma-\gamma^0\|_\infty=\max_{j,l}\sup_{z\in\Omega_Z}|\gamma_{jl}(z)-\gamma^0_{jl}(z)|$. The same norm is defined for $\phi$.

theorem[Uniform Convergence] Suppose assumptions in Theorem (ref), Assumptions (ref), (ref), and Assumption (ref) in Appendix (ref) hold. If $h\rightarrow0$, $Nh^Q\rightarrow\infty$, and $\ln(N)/(Nh^Q)\rightarrow0$, then \begin{itemize} • $\left\|\hat{\gamma}_N-\gamma^0\right\|_\infty=O_p(\left[\ln(N)/(Nh^Q)\right]^{1/2}+h^2)$; • for $\epsilon\rightarrow0$ as $N\rightarrow\infty$, we have \begin{align*} \sup\limits_{\|\hat{\gamma}_N-\gamma^0\|_\infty\leq\epsilon}\|\hat{\phi}_N-\phi^a\|_\infty=&O_p(\left\|\hat{\gamma}_N-\gamma^0\right\|_\infty),\\ \sup\limits_{\|\hat{\gamma}_N-\gamma^0\|_\infty\leq\epsilon}\|\hat{\phi}_N-\phi^0\|_\infty=&O_p(\left\|\hat{\gamma}_N-\gamma^0\right\|_\infty+\triangle_K). \end{align*} \end{itemize}

Theorem (ref) shows that the convergence of the estimated weights $\hat{\phi}_N$ to the true value $\phi^0$ is driven by two factors: the convergence rate of the kernel estimator $\hat{\gamma}_N$, and the approximation error of the matrix diagonalization method measured by $\triangle_K$.

theorem[Consistency]Suppose $\frac{\partial^2\mathcal{L}^0_{\mathbb{P}}(\theta)}{\partial\theta\partial\theta'}$ and $\frac{\partial^2\mathcal{L}^a_{\mathbb{P}}(\theta,\phi^a)}{\partial\theta\partial\theta'}$ are both full rank for all $\theta\in\Theta$. Under assumptions in Theorem (ref) and Assumption (ref) in Appendix (ref), we have \begin{align*}\|\hat{\theta}_N-\theta^a\|= o_p(1), \|\theta^a-\theta^0\|= O(\triangle_K), and \|\hat{\theta}_N-\theta^0\|= O_p(\triangle_K).\end{align*}

Theorem (ref) demonstrates that $\hat{\theta}_N$ is a consistent estimator for the pseudo-true parameter $\theta^a$, while its asymptotic bias with respect to the true value $\theta^0$ is governed by $$B_K:=\theta^0-\theta^a,~\text{ where }~\|B_K\|=O(\triangle_K).$$ We know that if the true degree is bounded by $K$ for all units, then $\|B_K\|=0$, and $\theta^0$ is consistently estimated. Next, we consider the asymptotic normality. Let $g(\tilde{W}_i;\theta,\phi)=\tau_i[Y_i-m^a(X_i;\theta,\phi)]\frac{\partial m^a(X_i;\theta,\phi)}{\partial\theta}$ be the score function of the sample objective function $\mathcal{L}^a_N(\theta,\phi)$. We can show that $$\frac{1}{\sqrt{N}}\sum_{i=1}^{N}g(\tilde{W}_i;\theta^a,\hat{\phi}_N)= \frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left[g(\tilde{W}_i;\theta^a,\phi^a)+\delta(\tilde{W}_i;\theta^a,\phi^a)\right]+o_p(1),$$ where $\delta(\tilde{W}_i;\theta^a,\phi^a)$ is the correction term to adjust the estimation error of the first-step kernel estimator. Denote a $d_\theta\times1$ vector $\tilde{g}_i=g(\tilde{W}_i;\theta^a,\phi^a)+\delta(\tilde{W}_i;\theta^a,\phi^a)$, where $\tilde{g}_i=(\tilde{g}_{i,1},...,\tilde{g}_{i,d_\theta})'$. Following (ref), define the within-cluster correlation for $\tilde{g}_i$ as $$\Sigma^{\tilde{g}}_N=\sum_{k=1}^{q_N}\sum_{i,j\in\mathbb{S}_k}Cov(\tilde{g}_i,\tilde{g}_j).$$

theorem[Asymptotic Normality]Suppose assumptions in Theorem (ref) and Assumptions (ref) to (ref) in Appendix (ref) hold. If further assume $\ln(N)/(N^{1/2}h^Q)\rightarrow0$ and $Nh^4\rightarrow0$ as $N\rightarrow\infty$, then $$\sqrt{N}(\hat{\theta}_N-\theta^0+B_K)\overset{d}{\rightarrow}\mathbb{N}(0,H^{-1}\Omega H^{-1}),$$ where $H=E\left[\frac{\partial g(\tilde{W}_i;\theta^a,\phi^a)}{\partial\theta'}\right]$, $\Omega=\lim\limits_{N\rightarrow\infty}\Sigma^{\tilde{g}}_N/N$, and $\mathbb{N}$ stands for a normal distribution.

Theorem (ref) implies that the bias term $B_K$ is negligible in the inference for $\theta^0$ if $\sqrt{N}\|B_K\|=O(\sqrt{N}\triangle_K)$ is sufficiently small. Theoretically, a consistent estimator of $H^{-1}\Omega H^{-1}$ can be obtained by replacing $H$ and $\Omega$ with their sample analogs. However, it is difficult to implement because the explicit formula for $\delta(\cdot;\theta,\phi)$, although exists, is complex. In practice, we suggest using the method of numerical differentiation of the influence function, as discussed in newey1994kernel, to estimate the correction term without specifying its analytic expression.\footnote{See hong2015extremum for discussions on the choice of numerical step size for the differentiation. }

Numerical and Empirical Results

Monte Carlo Simulation

In this section, we illustrate the finite-sample behavior of our method via Monte Carlo simulations. We consider two data generating processes (DGPs) for the outcome variable:

align[align omitted — 349 chars of source]

where, in both models, $D_i\overset{i.i.d.}{\sim}Bernoulli(0.3)$ is a randomized treatment and $\varepsilon_i\overset{i.i.d.}{\sim}\mathbb{N}(0,0.25)$ is an idiosyncratic error. We set $\theta=(\theta_1,\theta_2,\theta_3,\theta_4,\theta_5)'=(1,1,0.5,-0.1,1)'$. We generate data using sample size $N\in\{1000,2000,5000\}$ with replications $M=1000$.

True Network Data. We simulate the true network data using the model below:

align[align omitted — 145 chars of source]

\sloppy where $\alpha_i\overset{i.i.d.}{\sim}Bernoulli(0.5)$ stands for unobserved degree heterogeneity, $\rho_i=(\rho_{i1},\rho_{i2})\overset{i.i.d.}{\sim} Uniform([0,1]^2)$ is the random location of unit $i$, and $\zeta_{ij}=\zeta_{ji}\overset{i.i.d.\text{ across }(i,j)}{\sim}\mathbb{N}(0,1)$ is a random shock. $\alpha_i$, $\rho_i$, and $\zeta_{ij}$ are mutually independent. Let $d(\rho_i,\rho_j)$ be the distance between two units, where $d(\rho_i,\rho_j)=0$ if $\|\rho_i-\rho_j\|\leq r$ and $d(\rho_i,\rho_j)=\infty$ otherwise. We set $r=(r_{deg}/N)^{1/2}$ with $r_{deg}=3$ and $(\beta_1,\beta_2)=(-0.25,0.25)$. In this DGP design, the mean degree value is approximately 4 to 5, and the maximum degree value is about 14 to 15, as $N$ increases from $1000$ to $5000$.\footnote{We also conduct Monte Carlo simulations using an extension of model (ref) to allow for strategic network interactions, where one unit's link formation depends on the links of others. Due to space limitation, we present these simulation results in Appendix (ref). }

\noindentObserved Network Data. We generate the observed network data using $ A_{ij}=U_{ij}A^*_{ij} $. We consider three different DGPs for $U_{ij}$. The first DGP considers the case of missing completely at random: $$\text{(DGP1. random missing) }\quad U_{ij}\sim Bernoulli\left(p_U\right)\text{ are i.i.d. across all $(i,j)$}.$$ In the second DGP, the missing rate is heterogeneous and varies with the true degree value:

align*[align* omitted — 142 chars of source]

where $p_{U,i}=p_U+0.02*\log(\mathcal{T}^*_i+1)$. In the third DGP, missing indicators of the same unit $i$, $\mathbf{U}_i=(U_{i1},...,U_{iN})$, are correlated, while $\mathbf{U}_i$ and $\mathbf{U}_j$ are independent for all $i\neq j$:

align*[align* omitted — 140 chars of source]

where $\Phi(\cdot)$ denotes the standard normal CDF, $e_{ij}\overset{i.i.d.\text{ across }(i,j)}{\sim}\mathbb{N}(0,1)$, $\omega_i\overset{i.i.d.}{\sim} \mathbb{N}(0,1)$, and $e_{ij}$ and $\omega_i$ are mutually independent. In DGP3, the value of $\rho$ determines the correlation among $(U_{i1},...,U_{iN})$, and we set $\rho=0.1$. In DGP1 to DGP3, we consider $p_U\in\{0.1,0.2,0.3\}$.

\noindentEstimation Methods. We compare three estimation methods, including (1) Infeasible OLS -- the infeasible OLS regression that uses the true NBRVs; (2) Naive OLS -- the feasible OLS regression that uses the observable NBRVs and ignores the missing links; (3) SPE -- the semiparametric estimation based on the matrix diagonalization method that uses the observed incoming and outgoing degrees. For SPE method, we set $\varpi(y)=y$ and choose the value of $K$ so that the smallest singular value of $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}}$ is larger than 0.001. We order the estimated eigenvectors of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*}$ according to Assumption (ref) (c) with $E[\varpi(Y_i)|\mathcal{T}^*_i=n^*]$ strictly increasing in $n^*$.\footnote{For Model 1 in (ref), Assumption (ref) (c) holds because $\theta_5>0$ (see Example (ref)). For Model 2 in (ref), because $\mathcal{S}^*_i|\mathcal{T}^*_i=n^*$ follows a $Binomial(p_{D_i}(1),n^*)$ distribution, we can get $E[\varpi(Y_i)|\mathcal{T}^*_i=n^*]=\theta_1+\theta_2p_{D_i}(1)+\theta_3n^*p_{D_i}(1)+\theta_4[n^*p_{D_i}(1)p_{D_i}(0)+(n^*p_{D_i}(1))^2]+\theta_5n^*$. Given the value of $\theta$, Assumption (ref) (c) holds as long as $n^*$ is smaller than 62, which is a much larger value than the maximum degree in our DGP.}

\noindentEstimation Results. Our target parameter is the spillover effect $\eta^0:=\eta_S(d,s,0,n,z)= m^*(d,s,n,z)-m^*(d,0,n,z)$ at $d=0$, $s=1$, $s'=0$, and $n=4$ (no covariate $z$).\footnote{Simulation results for spillover effect defined at different values of $s$, $s'$, and $n$ display similar patterns. Therefore, we do not report them due to space limitation.} Tables (ref) and (ref) present the estimation results for Model 1 and Model 2, respectively. Denote $\hat{\eta}_j$ as the estimate of $\eta^0$ in the $j$-th simulation. We report the magnitude of the bias ($bias=\frac{1}{M}\sum_{j=1}^{M}(\hat{\eta}_j-\eta^0)$), the relative bias ($|bias/\eta^0|*100\%$), the standard deviation (sd), and the root mean squared error (rmse).

Some interesting patterns emerge. First, under DGP1 (random missing), the Infeasible OLS estimation is the least biased, with the smallest standard deviation and a relative bias of less than 0.8% in Model 1 and less than 0.3% in Model 2. Second, the Naive OLS produces the most biased estimates. When the sample size is relatively large ($N=5000$), its relative bias ranges from 11.4% to 33% in Model 1 and from 6.7% to 21% in Model 2. Third, the bias of SPE is substantially lower than that of the Naive OLS in both models. Specifically, when the sample size is relatively large ($N=5000$), the relative bias of SPE ranges from 0.5% to 16.4% in Model 1 and from 2.6% to 4.3% in Model 2. Nonetheless, the standard deviation of the SPE method exceeds that of the Naive OLS in both models, suggesting a bias-variance trade-off between these two feasible estimation methods. Lastly, similar patterns are observed in the estimation results under heterogeneous missing (DGP2) and dependent missing (DGP3). We can see that deviations from random missing, as considered in DGP2 and DGP3, lead to a slight increase in both the estimation bias and standard deviation for most cases in the Naive OLS and SPE methods.

table[table omitted — 5,507 chars of source]
table[table omitted — 5,489 chars of source]

Home Computer Use and Self-empowered Learning

In this section, we present results of a naturalistic simulation study and an empirical application using the school friendship data from beuermann2015one.\footnote{The dataset is available at https://www.aeaweb.org/articles?id=10.1257/app.20130267.} The authors conducted a randomized controlled trial in which “One Laptop per Child” (OLPC) laptops were provided to primary school students in Lima, Peru, and they examined the spillover effects of home computer use on children's self-empowered learning. In their study, fourteen treatment schools were randomly selected, and students in each class drew random lotteries to win the laptops. The baseline information, including self-reported network data, was collected in April/May 2011, before the experiment was implemented. The lottery was drawn in June/July 2011, and 1048 laptops were distributed to lottery winners. The follow-up data were collected in November 2011. We use the same sample of $N=2737$ students as beuermann2015one, which consists of students in grades 3 to 6 whose parents approved their lottery participation.

When collecting network information, students were asked to list up to 12 friends, including their closest friends, friends with whom they did homework together, and friends who visited their homes. Missing links exist for at least two reasons. First, the reported friends were top-coded to 12. Second, when constructing the network, reported friends with names that did not match those of any other students (for example, due to incomplete or misspelled names) were omitted because their treatment status could not be retrieved. Since the network data were collected before the experiment, it is reasonable to assume that the true network data and missing links are independent of the lottery results. Additionally, beuermann2015one found that once conditioning on the network degree, baseline characteristics were well balanced between students whose friends won laptops and those whose friends did not. Since observed and unobserved characteristics are often mutually dependent, this suggests that the true network and missing links are likely to be uncorrelated with unobserved characteristics.\footnote{The fact that the observed baseline characteristics are well-balanced among students with varying numbers of laptop-winner friends given the network degree implies that $S_i$ given $\mathcal{T}_i$ is uncorrelated with both observed and unobserved characteristics. This is consistent with our intermediate result proved in Lemma (ref) in Appendix (ref) under Assumption (ref) (unconfounded true network) and Assumption (ref) (nondifferential missing links).} It also indicates that the network degree is an important control variable and should be included in the regressions.

We aim to study the impacts of wining the lottery on a standardized test score for laptop digital skills (OLPC Test Score) at the follow-up stage.\footnote{The effects are studied from an intent-to-treat perspective, because 93% of students who won the lottery received laptops.} We consider the model specification:

align[align omitted — 213 chars of source]

where $Y_{ik}$ is the standardized test score for student $i$ in class $k$, $D_{ik}$ is the indicator of lottery winner, $Cov_{ik}$ includes a constant, age, sex, number of siblings, number of younger siblings, whether the father lives with the child, whether the father works at home, and whether the mother works at home, and $\mu_{k}$ represents the class fixed effect. We use the same definitions of NBRVs as in beuermann2015one, where $S_{ik}$ is the number of incoming friends who are lottery winners, $\mathcal{T}_{ik}$ is the indegree. We focus on two types of spillover effects: (i) $\theta_2$ -- the spillover of laptop winners on nonwinners (WoNW), and (ii) $\theta_2+\theta_3$ -- the spillover of laptop winners on other winners (WoW). Because the test score is standardized, the spillover effects are interpreted as the impacts on the standard deviations of OLPC test score.

Naturalistic Simulation

In the naturalistic simulation, we treat the reported network data as the true network, and the OLS coefficients for Model (ref) obtained by using the reported network data as the true coefficients. We simulate $M=1000$ experiments. In each experiment, $N$ error terms $\varepsilon_{ik}$ are randomly drawn from a normal distribution $\mathbb{N}(0,\sigma^2_\varepsilon)$, where $\sigma_\varepsilon$ is the standard deviation of the OLS residuals. Then, we generate $N$ outcome observations according to Model (ref), using the original treatment variable, NBRVs, and covariates for the $N$ students, along with the simulated error terms. Missing links are artificially introduced to the observed network $A_{ij}=U_{ij}A^*_{ij}$, where $U_{ij}$ is a binary indicator that is generated following the three DPGs considered in Section (ref), with $p_U\in\{0.1,0.2,0.3\}$.

table[table omitted — 3,700 chars of source]

We apply the two feasible estimation methods: the Naive OLS that ignores the missing links and the SPE that uses both incoming and outgoing links. Similar to the Monte Carlo simulation, for SPE method, we choose the value of $K$ so that the smallest singular value of $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}}$ is larger than 0.001, and we order the estimated eigenvectors of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*}$ according to Assumption (ref) (c). Table (ref) displays the values of the bias, relative bias (%), standard deviation (sd), and root mean squared error (rmse) of the estimated spillover effects of interest. We can see that, on average, Naive OLS underestimates the true spillover effects of WoNW and of WoW by 13% to 35% under DGP1 (random missing), 16% to 40% under DGP2 (heterogeneous missing), and 10% to 36% under DGP3 (dependent missing). The SPE estimates are less biased than those of Naive OLS, with relative bias ranging from 0.9% to 14% under DGP1, 0.5% to 19% under DGP2, and 2.5% to 12% under DGP3. As the missing probability $p_U$ increases, both estimation methods tend to underestimate the true effect in a systematic manner, and the degree of underestimation also increases.

Empirical Application

In this section, we analyze the consequences of missing network links in the study of the home computer use on self-empowered learning. We treat the reported network data as the observed network that contains missing links. We compare the estimation results of Naive OLS and SPE. For SPE, we choose different values of $K$ ($K\in\{9,10,11\}$) to define the truncated degree support, and we order the estimated eigenvectors of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*}$ according to Assumption (ref) (c). We present the estimation results of Model (ref) in Table (ref), using the sample of $N=$2737 students. The standard errors of both Naive OLS and SPE methods are clustered at the school level.\footnote{The standard errors of SPE method are calculated using the numerical method proposed in newey1994kernel. We follow hong2015extremum to choose the step size as $log(N)log(log(N))/N$ in the numerical differentiation.}

table[table omitted — 2,199 chars of source]

We find that the estimated spillover effects of WoNW and WoW using Naive OLS are positive but insignificant. Using Naive OLS, we obtain a WoNW spillover of 0.140 standard deviations with a 95% confidence interval (CI) of $[-0.052, 0.333]$, and a WoW spillover of 0.307 standard deviations with a 95% CI of $[-0.093, 0.708]$. Using SPE method, the point estimate of WoNW spillover effect ranges from 0.153 to 0.280 standard deviations, and the point estimate of WoW spillover effect ranges from 0.336 to 0.526 standard deviations, under different values of $K$. In addition, when $K=9$, the SPE estimate of the WoNW spillover effect is significant. The spillover effects for WoNW and WoW obtained by SPE are larger than those obtained by Naive OLS, indicating a possible underestimation of Naive OLS of the true spillover effects.

Conclusion

This paper investigates spillovers of program benefits in the presence of missing network links. We propose to use two network measures, that can be constructed using the incoming and outgoing links, to point identify the treatment and spillover effects in the case of bounded degree. If the degree is unbounded, our method can be used as a bias-reduction approach. We provide a two-step semiparametric estimation method and study its asymptotic properties. Monte Carlo experiments and a naturalistic simulation confirm the effectiveness of our approach in reducing estimation bias compared to the naive estimation that neglects missing links.

The literature on network effects often emphasizes the potential impacts of higher-order network connections with indirect friends (for example, friends of friends). However, incorporating these higher-order connections in the outcome model will introduce higher order missing links (for example, missing friends of friends), which complicates the dependence among observable and latent network-based random variables. Therefore, extending our method to address the missing link problem in models with higher-order network connections is nontrivial and left to future research. Furthermore, although our method focuses on missing network links, the network misclassification can be two-sided. Nonetheless, our method can still be used as bias-reduction approach when false positive links exist with a small or declining probability.

{ \setstretch{1} }

center[center omitted — 144 chars of source]