The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
119,118 characters
Spillovers of Program Benefits with Missing Network Links
\title{\LARGE{Spillovers of Program Benefits with Missing Network Links}\bigskip}
\author{Lina Zhang\thanks{Email: \texttt{[email removed]}. I would like to thank Isaiah Andrews, Sascha Becker, Denzil Fiebig, David Frazier, Jiti Gao, Frank Kleibergen, Tong Li, Francesca Molinari, Didier Nibbering, Bing Peng, Donald Poskitt, Kyungchul (Kevin) Song, Xun Tang, Denni Tommasi, Takuya Ura, Benjamin Wong, Xueyan Zhao, and participants of seminars at University of Glasgow, University of Amsterdam, University of Melbourne, Applied Young Economist Webinar, Renmin University, ESEM, ESAM, CMES, IAAE, and IPDC for helpful comments. All errors are mine.}\\University of Amsterdam and Tinbergen Institute\\}
\maketitle
\setlength{\abovedisplayskip}{11pt}
\setlength{\belowdisplayskip}{11pt}
\vspace{-0.5cm}
\begin{abstract}
\noindent The issue of missing network links in partially observed networks is frequently neglected in empirical studies. This paper addresses this issue when investigating the spillovers of program benefits in the presence of network interactions. Our method is flexible enough to account for non-i.i.d. missing links. It relies on two network measures that can be easily constructed based on the incoming and
outgoing links of the same observed network.
The treatment and spillover effects can be point identified and consistently estimated if network degrees are bounded for all units. We also demonstrate the bias reduction property of our method if network degrees of some units are unbounded. Monte Carlo experiments and a naturalistic simulation on real-world network data are implemented to verify the finite-sample performance of our method. We also re-examine the spillover effects of home computer use on children's self-empowered learning.\\
\noindent\textbf{JEL Codes}: C14, C21, C25, C26, C51.
\newline
\noindent\textbf{Keywords}: Heterogeneous treatment and spillover effects; ~Partially observed networks; ~Incoming and outgoing links; ~Non-i.i.d. missing; ~ Heterogeneous missing rates.
\end{abstract}
\newpage
\section{Introduction}\label{intro}
The importance of network interactions in shaping individuals' socio-economic outcomes has led to increasing attention in empirical studies on program evaluations
\citep[e.g.][]{oster2012determinants,banerjee2013diffusion,cai2015social,paluck2016changing,carter2021subsidies}.
However, a first-order practical issue that is often neglected is the presence of missing links in partially observed networks. This issue is pervasive due to various reasons, such as censored peer data, incomplete survey responses, or omitted network links due to missing information. Existing studies show that even a low missing rate can lead to a sizable bias in causal effect estimates \citep{advani2018credibly}. In this paper, we study the identification and estimation of the treatment and spillover effects of a randomized program intervention using a treatment response model that allows for flexible forms of heterogeneity \citep{manski2013identification,leung2020treatment}. We assume that network interactions affect the outcome through
two \emph{network-based random variables} (hereafter referred to as NBRVs): the network \emph{degree}
and the \emph{indirect exposure} to treated network neighbors.
We demonstrate that the identifiable spillover effects that ignore the missing links are mixtures of the true effects, with possibly negative weights and an opposite sign to the true effects.
To address the missing link problem, we employ the matrix diagonalization method of \citet{hu2008identification}, which requires two observed network measures that are mutually independent conditional on the true network.
In practice, network data are often collected from survey responses. Therefore, unlike traditional measurement error problems where an additional measure for the true variable is rare, we can easily construct two network measures using the incoming and outgoing links of the same observed network.\footnote{Replicated measures
are widely used to deal with measurement errors in the econometrics literature
\citep[see, e.g.][]{LI20021,mahajan2006identification,lewbel2007estimation,hu2017identification,calvi2018women,tommasi2022identifying} and
in the literature studying networks \citep[see, e.g.][]{goldsmith2013social,comola2017missing,chang2020estimation,li2021causal}.} These two observed networks are conditionally independent as long as the reporting errors made by one unit do not depend on those made by others.
Our method is flexible because it accommodates non-i.i.d. missing links. It allows for arbitrary correlations among missing links of the same unit and heterogeneous missing rates among different units, depending on their true degree values.
Using these two network measures, we demonstrate that point identification of the true effects can be achieved if degrees are bounded for all units. When degrees are unbounded for some units, our method can serve as a bias reduction approach under a restriction on the extent of network sparsity. We propose a two-step semiparametric estimation method and establish its asymptotic properties. To control data correlation under network interactions, we adopt the notion of the \emph{dependency neighborhood} used in
\citet{chandrasekhar2021network} and restrict data dependence to be local.
To assess the finite-sample performance of our method, we conduct Monte Carlo experiments and also a naturalistic simulation study using school friendship data from \citet{beuermann2015one}. The results of both the Monte Carlo and naturalistic simulations indicate that our method can effectively reduce the estimation bias in realistic samples compared to the naive method that ignores the missing links. Besides, we
re-examine the spillover effects of winning a home laptop lottery on children's self-empowered learning outside the classroom environment, as studied by \citet{beuermann2015one}.
We find that failing to account for missing network links can lead to an underestimation of the spillover effects of laptop lottery winners on digital skills of others.
Recent studies have emerged to study causal effects under network interactions using missing or misclassified network data. This paper is closely related to those employing repeated network measures \citep*[e.g.,][]{li2021causal,lewbel2022estimating}. In particular, \citet*{li2021causal} study causal effects under network interactions using an exposure mapping model. They develop a bootstrap estimation method that requires at least two observed network measures, and they assume independent noises in the observed network data and a parametric degree distribution. \cite*{lewbel2022estimating} examine the identification and estimation of peer effects using linear-in-means models. They employ both incoming and outgoing links to address the missing link problem and assume i.i.d. missing given individuals' covariates. Different from these two studies, our method accommodates non-i.i.d. missing links, allowing for arbitrary correlation among missing links of the same unit and heterogeneous missing rates depending not only on covariates but also on the actual degree values.
This paper is also related to, but different from, other causal effect studies that deal with missing or misclassified network links using methods other than repeated network measures.
Identification and estimation of peer effects through linear-in-means models are achieved, using adjusted 2SLS estimators in local-aggregate models \citep*{liu2013estimation}, assuming the existence of a consistent estimator of network distribution \citep{boucher2020estimating,herstad2023essays}, and utilizing an order-invariance condition on friends' covariates \citep{griffith2022name}. In addition, \citet{chandrasekhar2011econometrics} consider various network-based linear regressions and propose a two-step estimation using a graphical reconstruction process. \citet*{hardy2019estimating} identify causal effects in an exposure mapping model, assuming a parametric degree distribution and random noises in the network data.
A lower bound for the spillover effects is provided by \citet{he2023measuring} under the restriction of nonnegative spillovers.\footnote{There is also a separate literature that studies causal effects in the presence of network interactions when the network itself is entirely unobserved \citep*[see, e.g.,][]{depaula2023identifying,lin2021uncovering,lewbel2021socialun}.}
In addition, there is a growing literature addressing the problem of missing or misclassified network links when studying network formation or network statistics \citep[see,][for example]{butts2003network,balachandran2017propagation,comola2017missing,thirkettle2019identification,chang2020estimation,young2020bayesian,candelaria2020identification}.
This paper is distinct from these studies because our method does not require modeling the network formation process, and we aim to solve the missing link problem when the target parameter is the treatment and spillover effects.
\footnote{Besides, this paper also relates to the literature concerned with measurement error in discrete random
variables
\citep[see,][among others]{hausman1998misclassification,abrevaya1999semiparametric,li2003modeling,cameron2004modelling,molinari2008partial,chen2009nonparametric}.}
The rest of this paper is organized as follows. Section \ref{section_setup} introduces the model
setup and the causal effects of interest. Section \ref{subsection_bias} characterizes the bias caused by missing network links.
Section \ref{section_multiple_proxy} presents our proposed method and main results. Section
\ref{section_estimation_multiple} outlines the semiparametric estimation and its asymptotic properties. Section
\ref{section_numerical_empirical} presents the results of the Monte Carlo simulation, naturalistic simulation, and empirical analysis of the `One Laptop
per Child' program using real-life network data. Section \ref{section_conclusion} concludes. All proofs are provided in the online appendix.
\section{Model Setup}\label{section_setup}
We denote by $\mathbf{A}^*=\{A^*_{ij}\}$ the true adjacency matrix corresponding to an unweighted random network over the population $\mathcal{P}$. The network links can be either directed or undirected. Let $A^*_{ij}=1$ if unit $i$ and unit $j$ are linked, and $A^*_{ij}=0$ otherwise. As a convention, self-links are ruled out, i.e., $A^*_{ii}=0$. Let $\mathcal{N}^*_i=\{j\in\mathcal{P}:A^*_{ij}=1\}$ be the set of unit $i$'s network neighbors. Consider a treatment response model for the outcome $Y_i$:
\begin{equation}
\begin{aligned}\label{outcome1}
Y_i&=r(D_i,S^*_i,\mathcal{T}^*_i,Z_i,\varepsilon_i),\;\text{ for each }i\in\mathcal{P},
\end{aligned}
\end{equation}
where $r$ is an unknown function, $D_i$ is a binary treatment variable, $Z_i$ is a vector of covariates, and $\varepsilon_i$ is a vector of unobservable error terms. We define $S_i^*=\sum_{j\in\mathcal{N}^*_i}D_j$ as the number of treated network neighbors, and $\mathcal{T}_i^*=|\mathcal{N}^*_i|$, where $|\cdot|$ denotes the cardinality of a set, as the network degree. The treatment response model in \eqref{outcome1} assumes that network interactions affect the outcome through two \emph{network-based random variables} (hereafter referred to as NBRVs): $S_i^*$, which measures the extent of indirect exposure to the treatment, and $\mathcal{T}_i^*$, which quantifies the popularity of each unit $i$.
The same model is used by \citet{leung2020treatment} and \citet{viviano2024policy} to capture various forms of heterogeneous treatment and spillover effects, and tests for model specifications are developed by \citet*{athey2018exact}.
We can view model \eqref{outcome1} as a potential outcome model, where $(D_i,S^*_i)$ acts as a multivalued treatment, and $\mathcal{T}^*_i$ and $Z_i$ are control variables. For any random variable $B_i$, let $\Omega_B$ denote its support. For any $d\in\{0,1\}$ and $(s,n,z)\in\Omega_{S^*,\mathcal{T}^*,Z}$, let us denote
\begin{align}\label{def_CASF}
m^*(d,s,n,z)&=E\left[r(d,s,n,z,\varepsilon_i)\big|\mathcal{T}^*_i=n,Z_i=z\right].
\end{align}
By definition, $m^*(d,s,n,z)$ captures the mean value of the outcome under a counterfactual treatment $d$ and a counterfactual number of treated network neighbors $s$, given the control variables $(\mathcal{T}^*_i,Z_i)=(n,z)$. Following \citet{leung2020treatment}, we refer to $m^*$ as the conditional average structural function (CASF). Given the CASF, the treatment and spillover effects can be defined as the average response to the counterfactual manipulation of a unit's own treatment status and its indirect exposure to the treatment, respectively.
\begin{definition}[\textbf{Treatment and Spillover Effects}]\label{def2}For any $d\in\{0,1\}$, $(n,z)\in\Omega_{\mathcal{T}^*,Z}$, and $s,s'\in\Omega_{S^*}$, define
\begin{align*}
\text{treatment effect: }&\eta^*_T(s,n,z)=m^*(1,s,n,z)-m^*(0,s,n,z),\\
\text{spillover effect: }&\eta^*_S(d,s,s',n,z)=m^*(d,s,n,z)-m^*(d,s',n,z).
\end{align*}
\end{definition}
Given $(\mathcal{T}^*_i,Z_i)=(n,z)$, the treatment effect $\eta^*_T$ measures the direct treatment effect caused by the variation in a unit's own treatment status, while fixing its exposure to the treated network neighbors. The spillover effect $\eta^*_S$ measures the indirect treatment effect caused by the variation in a unit's exposure to the treated network neighbors, while fixing its own treatment status. The analysis in this paper can be easily extended to study other forms of direct and spillover effects, as long as they are defined as functions of $m^*$.
\subsection{Motivation Examples}
In model \eqref{outcome1}, we assume that network interactions affect the outcome through two
NBRVs, namely, $S_i^*$ and $\mathcal{T}_i^*$, which are commonly used network statistics in empirical studies of program evaluations under network interactions.
We illustrate the usefulness of our model using the examples below, where the parametrization of function $r$ is used only for illustrative purposes.
\begin{example}[\textbf{Diffusion of a Weather Insurance Product}]\label{example_motivation1}\citet{cai2015social} study the influence of social networks on weather insurance adoption in rural China. The authors consider a model for the treatment and spillover effects at the household level of the form
$$Y_{i}=\theta_0+\theta_1D_{i}+\theta_2\frac{S^*_{i}}{\mathcal{T}^*_{i}}+\theta'_3Z_{i}+\theta'_4NetSize_{i}+\varepsilon_{i
},$$
where the binary outcome $Y_{i}$ indicates whether household $i$ decides to purchase the
insurance, the treatment variable $D_{i}$ takes value one if a household is randomly invited to an intensive information session that introduces a new insurance product, $\frac{S^*_{i}}{\mathcal{T}^*_{i}}$ is the fraction of treated network neighbors, $NetSize_{i}$ is a set of dummies indicating network degree values, and covariates in $Z_{i}$ include household characteristics and village fixed effects.
\end{example}
\begin{example}[\textbf{Adoption of Menstrual Cups}]\label{example_motivation2}\citet{oster2012determinants} explore the role of network interactions in technology adoption using data from a randomized allocation of menstrual cups in Nepal. Their model for the binary outcome of menstrual cup adoption can be summarized as follows:
$$Y_{i}=1[\theta_0+\theta_1D_i+\theta_2h(S^*_i,\mathcal{T}^*_{i})+\theta'_3Z_i+\theta'_4NetSize_{i}>\varepsilon_i],$$
where $D_i$ indicates the randomized access to menstrual cup, $h(S^*_i,\mathcal{T}^*_{i})$ represents either the number of treated friends $S^*_{i}$ or the share $\frac{S^*_{i}}{\mathcal{T}^*_{i}}$, $NetSize_{i}$ includes dummies that control for different network degree values, and $Z_i$ is a vector of other attributes and school fixed effects.
\end{example}
\begin{example}[\textbf{Subsidies and African Green Revolution}]\label{example_motivation3}\citet{carter2021subsidies} study the spillovers of a government subsidy on Green Revolution technology adoption. They estimate the following model for the technology adoption or agricultural yields:
\begin{align*}
y_{it}=&\theta_0+\theta_1D_{i}*Dur_t+\theta_2D_i*After_t\\
&~~~+\theta_3SocialTreat_i*Dur_t+\theta_4SocialTreat_i*After_t+\theta_5'Z_{it}+\theta_6'NetSize_i+\varepsilon_{it},
\end{align*}
where the treatment variable $D_{i}$ is one if household $i$ won the program lottery, $Dur_t$ and $After_t$ are time dummies for during and after the subsidy period, $SocialTreat_i=1[S^*_i\geq 2]$ indicates if the household has two or more lottery winner network neighbors, $NetSize_i$ is a set of dummies for different network degree values, and $Z_{it}$ consists of time and locality fixed effects.
\end{example}
\subsection{Treatment and Spillover Effects under True Network}
Let us begin by introducing the assumptions under which the treatment and spillover effects are point identified if the true network is correctly observed. For any random variables $B_i$ and $C_i$, denote $p_{B_i}(b)$ as its probability density (or mass) function and $p_{B_i|C_i=c}(b)$ as its conditional version. Let $\perp$ denote statistical independence.
\begin{assumption}\label{unconf1}~
\begin{itemize}
\item[(a)]\emph{(\textbf{Randomized Treatment})} $D_i$ and $Z_i$ are i.i.d. across $i$, and $D_i\perp(\varepsilon_j,Z_j,\mathcal{N}^*_j)$ for $\forall i,j\in\mathcal{P}$. In addition, $p_{D_i}(1)\in(\epsilon,1-\epsilon)$ for some constant $\epsilon>0$.
\item[(b)]\emph{(\textbf{Unconfounded Network})} For $\forall i,j\in\mathcal{P}$, $\varepsilon_i\perp (\mathcal{N}^*_j,Z_j)\big|\mathcal{T}^*_i,Z_i$.
\item[(c)]\emph{(\textbf{Identical Error Distribution})} For $\forall i,j\in\mathcal{P}$, we have $p_{\varepsilon_i|\mathcal{T}^*_i=n^*,Z_i=z}(e)=p_{\varepsilon_j|\mathcal{T}^*_j=n^*,Z_j=z}(e)$ for any $e\in\Omega_\varepsilon$, $n^*\in\Omega_{\mathcal{T}^*}$, $z\in\Omega_Z$.
\end{itemize}
\end{assumption}
Assumption \ref{unconf1} (a) assumes a randomized treatment allocation and i.i.d. covariates, which are relevant for a wide range of experimental contexts \citep[see][for a review]{athey2017randomized}.\footnote{We can relax the fully randomized treatment to an unconfounded treatment that is i.i.d. conditional on a subvector of individual characteristics $\tilde{Z}_i\subseteq Z_i$. This requires an additional assumption that $\tilde{Z}_i$ does not enter the network formation process of $\mathbf{A}^*$. More detailed explanations can be found in Appendix \ref{appendix_unconfounded}. Extensions to non-i.i.d. covariates do not provide additional insights but may introduce technical complications. Hence, we omit this discussion for simplicity.}
The network unconfoundedness in condition (b) permits dependence between the network and unobservable characteristics through the degree and covariates. For instance, it allows for spillover of unobservables in the form given in Example \ref{example_sp_unob} below. While the unconfounded network rules out certain types of endogenous networks, such as those with unobserved homophily \citep[see, e.g.,][]{johnsson2021estimation} where unobserved factors may enter both the network formation and the outcome model, it is still weaker than a fully exogenous network.
The identical distribution of the error term in condition (c) ensures that the expressions of the treatment and spillover effects introduced in Definition \ref{def2} are the same for all $i\in\mathcal{P}$.
\begin{example}[\textbf{Spillover of Unobservables}]\label{example_sp_unob}Suppose that the error term in the outcome equation, $\varepsilon_i$, is a scalar function of $(\sum_{j\in\mathcal{N}^*_i}e_j,\mathcal{T}^*_i,e_i)$, where $e_i$ is an unobservable i.i.d. Bernoulli error that is independent of $\mathbf{A}^*$ and $\mathbf{Z}^*=\{Z_i\}_{i\in\mathcal{P}}$, and $\sum_{j\in\mathcal{N}^*_i}e_j$ measures the spillover of unobservables. One example can be $\varepsilon_i=\frac{1}{\mathcal{T}^*_i}\sum_{j\in\mathcal{N}^*_i}e_j+e_i$. In this example, $\sum_{j\in\mathcal{N}^*_i}e_j$ given $\mathcal{T}^*_i$ and $Z_i$ is a sum of a known number of i.i.d. Bernoulli variables, and it follows a binomial distribution that depends on the network only through $\mathcal{T}^*_i$.
\end{example}
The proposition below demonstrates that if the true network is correctly observed, then $m^*$ and the treatment and spillover effects are all point identified.
\begin{proposition}\label{lemma_unconf1}Under Assumption \ref{unconf1}, we have that
$$m^*(d,s,n,z)=E\left[Y_i\big|D_i=d,S^*_i=s,\mathcal{T}^*_i=n,Z_i=z\right].$$
\end{proposition}
\section{Biased Effects under Missing Network Links}\label{subsection_bias}
Existing methods for studying spillover effects often assume no missing links in the observed network data,\footnote{See \citet*{leung2020treatment}, \citet*{vazquez2019identification}, \citet{sanchez2022network}, and \citet*{viviano2024policy} for example.} but such an assumption fails to hold in many empirical applications. Suppose we randomly draw $N$ units from the population $\mathcal{P}$ and collect their network information. Denote the observable adjacency matrix as $\mathbf{A}=\{A_{ij}\}$, where self connections are dropped. For $i=1,...,N$ and $j\in\mathcal{P}$, write the observed link as
\begin{align*}
A_{ij}=U_{ij}A^*_{ij},
\end{align*}
where $U_{ij}\in\{0,1\}$ indicates a missing link. Specifically, $U_{ij}=1$ implies that a true link $A^*_{ij}=1$ is correctly observed, whereas $U_{ij}=0$ implies that a true link $A^*_{ij}=1$ is absent. Define $\mathcal{N}_i=\{j\in\mathcal{P}:A_{ij}=1\}$ as the set of unit $i$'s observed network neighbors. Let $\mathcal{T}_i=|\mathcal{N}_i|$ be the observed network degree and $\mathcal{S}_i=\sum_{j\in\mathcal{N}_i}D_j$ be the number of observed treated network neighbors. Assume that we can observe each sampled unit $i$'s outcome, treatment status, covariates, and treatment assignments of $i$'s observable network neighbors:
\begin{align*}
\left\{Y_i,D_i,Z_i,\mathcal{N}_i,\{D_{j}\}_{j\in \mathcal{N}_i}\right\},\text{ for }i=1,2,...,N.
\end{align*}
In the presence of missing links, the observed network is a subset of the true network for all sampled units, i.e., $\mathcal{N}_i\subseteq \mathcal{N}^*_i$. Given the observed NBRVs for each unit, we define the identifiable counterpart for the CASF $m^*$ as below:
\begin{align}\label{identfiable_m}
m(d,s,n,z)=E[Y_i|D_i=d,S_i=s,\mathcal{T}_i=n,Z_i=z],\end{align}
where we replace $S^*_i$ and $\mathcal{T}^*_i$ in the definition of $m^*$ with the observed $S_i$ and $\mathcal{T}_i$.
If we ignore the presence of missing links and use the identifiable $m$ to compute the treatment and spillover effects of interest, we will obtain
\begin{equation}\label{def_id_treatment_spillover_effects}
\begin{aligned}
\eta_T(s,n,z):=&m(1,s,n,z)-m(0,s,n,z),\\
\eta_S(d,s,s',n,z):=&m(d,s,n,z)-m(d,s',n,z).
\end{aligned} \end{equation}
We employ the following assumptions to establish the bias of $\eta_T$ and $\eta_S$ relative to $\eta^*_T$ and $\eta^*_S$.
\begin{assumption}\label{nondiff}\emph{(\textbf{Nondifferential Missing Links})} For $\forall i,j\in\mathcal{P}$, $D_i\perp(\varepsilon_j,Z_j,\mathcal{N}^*_j,\mathcal{N}_j)$ and
$\varepsilon_i\perp\left(\mathcal{N}^*_j,\mathcal{N}_j\right)\big|\mathcal{T}^*_i,Z_i$.
\end{assumption}
Assumption \ref{nondiff} requires the treatment variable to be independent of missing links, which is trivially satisfied by randomized treatment assignments. In addition, it assumes that the missing links do not contain relevant information regarding the potential outcomes, given the actual degree and individual's characteristics. Following the literature \citep[e.g.,][]{bound2001measurement}, missing links that satisfy these conditions can be referred to as ``nondifferential'' missing links.
\begin{assumption}\label{identical_degree}\emph{(\textbf{Identical Degree Distribution})} For $\forall i,j\in\mathcal{P}$, we have
\begin{itemize}
\item[(a)]$p_{\mathcal{T}^*_i|Z_i=z}(n^*)=p_{\mathcal{T}^*_j|Z_j=z}(n^*)$ for all $n^*\in\Omega_{\mathcal{T}^*}$ and $z\in\Omega_Z$;
\item[(b)]$p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(n)=p_{\mathcal{T}_j|\mathcal{T}^*_j=n^*,Z_j=z}(n)$ for all $n\in\Omega_{\mathcal{T}}$, $n^*\in\Omega_{\mathcal{T}^*}$, and $z\in\Omega_Z$.
\end{itemize}
\end{assumption}
Assumption \ref{identical_degree} (a) requires the distribution of the true degree to be identical for all units with the same characteristics $Z_i$. It still allows for degree heterogeneity but may rule out the possibility of strategic network interactions, under which the network formation of one unit depends on the existing links of others.\footnote{In Appendix \ref{appendix_simulation_add}, we present Monte Carlo simulation results for our proposed method using network data generated by incorporating strategic interactions. The results demonstrate that our method can reduce the estimation bias compared to the naive estimation which ignores missing network links, even if Assumption \ref{identical_degree} (a) is violated.} In Example \ref{example_identical_degree}, we present a network formation model that takes into account degree heterogeneity and satisfies condition (a). Condition (b) requires the observable degree to be identically distributed across all units who have the same true degree value and covariates. In Lemma \ref{lemma_degree_dis} of Appendix \ref{appendix_example}, we show that at least two types of missing links
are allowed under this condition: (i) missing completely at random, where $U_{ij}$ is i.i.d. across all pairs $(i,j)$ and independent of all other variables, and (ii) missing not at random, where $U_{ij}$ can exhibit arbitrary correlations across $j$ for any given unit $i$, and the missing probability can vary with the true degree and covariates. Assumption \ref{identical_degree} is employed to ensure that the expressions of $m$ and the identifiable effects, $\eta_T$ and $\eta_S$, are the same for all units.
\begin{example}[\textbf{Identical Degree Distribution}]\label{example_identical_degree}
Without loss of generality, suppose there are no covariates $Z_i$. Consider a network formation model $$ A^*_{ij}= 1[\beta_1+\beta_2(\alpha_i+\alpha_j)-d(\rho_i,\rho_j)>\zeta_{ij}],$$ where $\alpha_i$ stands for the unobserved degree heterogeneity, $d(\rho_i,\rho_j)$ is the distance between two units defined using random location variables $\rho_i$ and $\rho_j$, and $\zeta_{ij}$ is a link-specific random shock. Suppose $(\alpha_i,\rho_i)$ is i.i.d. across $i$, $\zeta_{ij}$ is i.i.d. across $(i,j)$, and $\{\alpha_i,\rho_i\}_{i\in\mathcal{P}}$ and $\{\zeta_{ij}\}_{i,j\in\mathcal{P}}$ are mutually independent. Given $(\alpha_i,\rho_i)$, for any fixed $i$, $ A^*_{ij}$ becomes a function of $(\alpha_j,\rho_j,\zeta_{ij})$ and is i.i.d. across $j$. Then, $\mathcal{T}^*_i=\sum_{j\in\mathcal{P}} A^*_{ij}$, conditional on $(\alpha_i,\rho_i)$, is a sum of $(|\mathcal{P}|-1)$ i.i.d. Bernoulli variables and follows a Binomial distribution that only depends on $(\alpha_i,\rho_i)$. Because $(\alpha_i,\rho_i)$ is identically distributed across $i$, the unconditional degree distribution is also the same for all units. Detailed proofs can be found in Lemma \ref{lemma_example_identical_degree_dist} in Appendix \ref{appendix_example}.
\end{example}
\begin{theorem}\label{prop_conditional_mean}Under Assumptions \ref{unconf1}, \ref{nondiff}, and \ref{identical_degree}, $m$ is identical for all units, and it is a mixture of $m^*$, where
\begin{align*}
m(d,s,n,z)=\sum\limits_{(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}} m^*(d,s^*,n^*,z)\times p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*).
\end{align*}
\end{theorem}
Theorem \ref{prop_conditional_mean} demonstrates that
$m$ is a mixture of $m^*$ with the weight $p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}$ that quantifies the severity of the missing-link problem.
Given the mixture expression of $m$, we can characterize the bias in $\eta_T$ and $\eta_S$.
\begin{corollary}[\textbf{Biased Effects under Missing Links}]\label{coro_bias_expression}Under Assumptions \ref{unconf1}, \ref{nondiff}, and \ref{identical_degree},
\begin{align*}
\eta_T(s,n,z)=&\sum\limits_{(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}} \eta^*_T(s^*,n^*,z)\times
p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*),\\
\eta_S(d,s,s',n,z)=&\sum\limits_{(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}}\eta_S^*(d,s^*,s',n^*,z)\nonumber\\
&~~~~~~~~\hspace{1cm}\times\left[
p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*)-p_{S^*_i,\mathcal{T}^*_i|S_i=s',\mathcal{T}_i=n,Z_i=z}(s^*,n^*)\right].
\end{align*}
\end{corollary}
Corollary \ref{coro_bias_expression} reveals that the identifiable treatment effect $\eta_T(s,n,z)$ is a nonnegatively-weighted average of the true treatment effects. Consequently, if the true treatment effects are all positive or all negative, $\eta_T(s,n,z)$ will maintain the same sign. However, we cannot point identify the value of $\eta^*_T$ using $\eta_T$, except in the special case where $\eta^*_T(s,n,z)$ is homogeneous in $(s,n)$. In other words, if there exist functions $m^*_1$ and $m^*_2$ such that $m^*(d,s,n,z)=m^*_1(d,z)+m^*_2(s,n,z)$, then we can point identify the true treatment effect $\eta^*_T$ by $\eta_T$, as
\begin{align*}
\eta^*_T(s,n,z)=\eta_T(s,n,z)=m^*_1(1,z)-m^*_1(0,z),\;\text{ for any }(s,n,z).
\end{align*}
Furthermore, the identifiable spillover effect $\eta_S(d,s,s',n,z)$ is also a weighted average of the true spillover effects, albeit with possibly negative weights that sum up to zero.\footnote{Because $\sum_{(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}}[p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*)-p_{S^*_i,\mathcal{T}^*_i|S_i=s',\mathcal{T}_i=n,Z_i=z}(s^*,n^*)]=1-1=0$, the weights in $\eta_S(d,s,s',n,z)$ sum to zero, so they cannot be all positive or all negative for every $(s^*,n^*)\in\Omega_{S^*,\mathcal{T}^*}$.} Therefore, $\eta_S(d,s,s',n,z)$ may have the opposite sign of $\eta^*_S(d,s,s',n,z)$, resulting in either an upward or a downward bias. In a special case where network interactions have no impact on $Y_i$, i.e., $m^*(d,s,n,z)=m^*(d,z)$, we can point identify the true spillover effect $\eta^*_S$ using $\eta_S$, since they are both zero:
\begin{align*}
\eta_S^*(d,s,s',n,z)=\eta_S(d,s,s',n,z)=0,\;\text{ for any }(d,s,s',n,z).
\end{align*}
\section{Main Results}\label{section_multiple_proxy}
This section proceeds in three steps. First, we demonstrate that the weight $p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}$ in Theorem \ref{prop_conditional_mean}, which connects $m$ to the target CASF $m^*$, is a product of two distribution functions. Second, we introduce a sparse network assumption and recover the weight by tackling the two distribution functions separately. Finally, we discuss the identification of the CASF $m^*$.
\subsection{Decomposition of the Weights}
We can see that the weight $p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}$ can be rewritten as a product:
\begin{align}\label{equation_joint_decomposition_pre}
&p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}
=p_{S^*_i|\mathcal{T}^*_i,S_i,\mathcal{T}_i,Z_i}\times p_{\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}.
\end{align}
Recall that $S_i=\sum_{j\in \mathcal{N}_i}D_j$ denotes the number of observed treated
network neighbors among all $\mathcal{T}_i=|\mathcal{N}_i|$ observed network neighbors. Due to the randomized treatment assignment, $S_i$ given $(\mathcal{T}_i,Z_i)$ is a sum of a given number of i.i.d.
binary variables. Thus, $S_i$ given $(\mathcal{T}_i,Z_i)$ follows a distribution of $Binomial(p_{D_i}(1),\mathcal{T}_i)$ and is independent of the true degree $\mathcal{T}^*_i$. Therefore, the second term on the right-hand side of \eqref{equation_joint_decomposition_pre} reduces to $$p_{\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}=p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}. $$
Then, we can see that the weight $p_{S^*_i,\mathcal{T}^*_i|S_i,\mathcal{T}_i,Z_i}$ is determined by two components: $p_{S^*_i|\mathcal{T}^*_i,S_i,\mathcal{T}_i,Z_i}$, which represents the dependence between the true and observed NBRVs, and $p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}$, which captures the missing probabilities in the degree. This result is formally introduced below.
\begin{theorem}\label{decomposition_weight}\emph{(\textbf{Decomposition of Weight})} Under Assumptions \ref{unconf1}, \ref{nondiff}, and \ref{identical_degree},
we have
\begin{align*}
&p_{S^*_i,\mathcal{T}^*_i|S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*,n^*)
=p_{S^*_i|\mathcal{T}^*_i=n^*,S_i=s,\mathcal{T}_i=n,Z_i=z}(s^*)\times p_{\mathcal{T}^*_i|\mathcal{T}_i=n,Z_i=z}(n^*).
\end{align*}
\end{theorem}
Next, we show that the first term in the weight decomposition, $p_{S^*_i|\mathcal{T}^*_i,S_i,\mathcal{T}_i,Z_i}$, is also a binomial distribution and can be point identified using observed data.
Denote $$\Delta S_i:=S^*_i-S_i\text{ and }\Delta\mathcal{T}_i:=\mathcal{T}^*_i-\mathcal{T}_i.$$
The point identification of NBRV dependence relies on two key facts. First, $p_{S^*_i|S_i,\mathcal{T}_i,\mathcal{T}^*_i,Z_i}=p_{\Delta S_i|S_i,\mathcal{T}_i,\Delta\mathcal{T}_i,Z_i}$. Second, in the presence of missing links, $\Delta S_i=\sum_{j\in \mathcal{N}^*_i \bigcap \mathcal{N}^c_i}D_j$, where $\mathcal{N}^c_i$ denotes the complement of the set $\mathcal{N}_i$ and $|\mathcal{N}^*_i \bigcap \mathcal{N}^c_i|=\Delta\mathcal{T}_i$. Thus, because of the randomized treatment allocation, $\Delta S_i$ given $(\Delta\mathcal{T}_i,Z_i)$ is a sum of a given number of i.i.d. binary
variables, which follows a $Binomial(p_{D_i}(1),\Delta\mathcal{T}_i)$ distribution and is independent of $(S_i,\mathcal{T}_i)$. Therefore, these two facts together imply that $$p_{S^*_i|S_i,\mathcal{T}_i,\mathcal{T}^*_i,Z_i}=p_{\Delta S_i|S_i,\mathcal{T}_i,\Delta\mathcal{T}_i,Z_i}=p_{\Delta S_i|\Delta\mathcal{T}_i ,Z_i},$$
which remains invariant to $(S^*_i,\mathcal{T}^*_i,S_i,\mathcal{T}_i)$ as long as the two differences, $\Delta S_i$ and $\Delta\mathcal{T}_i$, are fixed. This simplification dramatically reduces the dependence structure between NBRVs. Let $\binom{n}{s}$ be the number of $s$-combinations from $n$ elements.
\begin{theorem}\label{lemma_iden}\emph{(\textbf{Point Identification of NBRV Dependence})} Under Assumptions \ref{unconf1}, \ref{nondiff}, and \ref{identical_degree},
$$p_{S^*_i|S_i=s,\mathcal{T}_i=n,\mathcal{T}^*_i=n^*,Z_i}(s^*)=\begin{cases}p_{\Delta S_i|\Delta\mathcal{T}_i=\Delta n ,Z_i}(\Delta s),&\mbox{if }s^*\leq n^*,~s\leq n,\text{ and }0\leq \Delta s\leq\Delta n,\\
0,&\mbox{otherwise},\end{cases}$$ where $\Delta s:=s^*-s$, $\Delta n:=n^*-n$, and
$p_{\Delta S_i|\Delta\mathcal{T}_i=\Delta n,Z_i}(\Delta s)=
\binom{\Delta n}{\Delta s}p_{D_i}(1)^{\Delta s}p_{D_i}(0)^{\Delta n-\Delta s}.
$
\end{theorem}
Since the treatment probability $p_{D_i}$ is identifiable from the data, $p_{S^*_i|\mathcal{T}^*_i,S_i,\mathcal{T}_i,Z_i}$ is point identified based on Theorem \ref{lemma_iden}.
Theorems \ref{prop_conditional_mean}, \ref{decomposition_weight}, and \ref{lemma_iden} together imply that
if we can identify the second term in the product expression of the weights, i.e., the missing probabilities in the degree $p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}$, we will be able to utilize the mixture model of $m$ and the weight to recover the target CASF $m^*$.
\subsection{Matrix Diagonalization Method}\label{section_matrix_diagonalization}
In this section, we adopt the matrix diagonalization method \citep{hu2008identification} and the matrix perturbation analysis \citep{stewart1990matrix} to recover the missing probabilities in the degree, $p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}$. Our proposed method uses two network measures that can be easily constructed using the incoming and outgoing network links of the same observed network. This method is flexible as it can accommodate arbitrary correlations among missing links of the same unit and heterogeneous missing rates based on the true degree. Without loss of generality, let us denote $\mathcal{N}_i=\{j\in \mathcal{P}:~A_{ij}=1\}$ and $\tilde{\mathcal{N}}_i=\{j\in \mathcal{P}:~A_{ji}=1\}$ as the set of unit $i$'s observed outgoing and incoming network neighbors, respectively. Then, $\mathcal{T}_i=|\mathcal{N}_i|$ and $\tilde{\mathcal{T}}_i=|\tilde{\mathcal{N}}_i|$ denote the observed out-degree and in-degree.\footnote{In some empirical studies, we can only observe the incoming network neighbors among sampled units, i.e., $\tilde{\mathcal{N}}_i=\{j\in \{1,...,N\}:~A_{ji}=1\}$. If this is the case, $\tilde{\mathcal{N}}_i$ can still be defined as $\tilde{\mathcal{N}}_i=\{j\in \mathcal{P}:~A_{ji}=1\}$, as we have $A_{ji}=0$ for all $j\not\in\{1,...,N\}$.} We assume that the support of the true and observed degrees is the set of non-negative integers $\{0,1,2,...\}$.\footnote{Our method can also be applied to cases with no isolated nodes.}
Below, we introduce assumptions that are required for implementing the matrix diagonalization method. These assumptions are strengthened versions of those in \citet{hu2008identification}, modified to accommodate potentially unbounded network degrees. We start with a sparse network assumption, which defines a truncated degree support. Sparse networks are common in many social science contexts, as human beings have a limited amount of time and energy to maintain their social connections.
\begin{assumption}[\textbf{Sparse Network}]\label{ass_support}There exists a bounded integer $0<K<\infty$, such that
\begin{itemize}
\item[(i)] $\sum_{k> K}p_{\mathcal{T}^*_i|Z_i=z}(k)\leq\triangle_K$ for $\forall z\in\Omega_Z$;
\item[(ii)]there exists a $\delta^*=\delta^*(K)>0$ such that $p_{\mathcal{T}^*_i|Z_i=z}(k)>\delta^*$ and $p_{\mathcal{T}_i|Z_i=z}(k)>\delta^*$ for all $k=0,...,K$ and $\forall z\in\Omega_Z$.
\end{itemize}
\end{assumption}
Assumption \ref{ass_support} requires a sparse network such that the probability of having a degree larger than $K$ is bounded by $\triangle_K$. Theoretically, we require $K$ to be a known and bounded value and $\triangle_K$ to be sufficiently small. Thus, even if $\triangle_K$ may decrease to zero as $K$ increases, we do not require $K$ to go to infinity as the sample size increases. Note that it is possible to relax this assumption by allowing $K$ to depend on $Z=z$. However, we omit this dependence to ease the notation. If the degree is uniformly bounded for all units, i.e., $\max_{i\in \mathcal{P}}\{\mathcal{T}^*_i\}= K$, then this assumption holds with $\triangle_K=0$. See \citet{paula2018identifying}, \citet{richards2020application}, and \citet{hu2020binary2} for papers assuming bounded degrees. If only a limited number of units have unbounded degrees, this assumption also holds with $\triangle_K$ close to zero for a carefully chosen $K$. The latter is often satisfied in social network data. For example, \citet{graham2015methods} documented that in a risk-sharing network among households, a small number of households have many
links, whereas the vast majority have fewer than ten links. In this case, we can set $K=10$. A bounded $K$ is crucial for implementing the matrix diagonalization method. In practice, however, while the value of $K$ still needs to be bounded, it can vary with different sample sizes.
\begin{assumption}[\textbf{Exclusion Restriction}]\label{nondiff3} $\mathcal{T}_i\perp\tilde{\mathcal{T}}_i\big|\mathcal{T}^*_i,Z_i$.
\end{assumption}
Assumption \ref{nondiff3} states that $\tilde{\mathcal{T}}_i$ contains no extra information of $\mathcal{T}_i$ beyond what the actual degree $\mathcal{T}^*_i$ already provides.
Intuitively, the exclusion restriction is satisfied by the incoming and outgoing degrees, if the missing links of one unit, $\mathbf{U}_i=\{U_{ij}\}_{j\in\mathcal{P}}$, are conditionally independent of the missing links of others, $\mathbf{U}_k=\{U_{kj}\}_{j\in\mathcal{P}}$ for all $k\neq i$, given $(\mathbf{A}^*,\mathbf{Z})$ where $\mathbf{Z}=\{Z_i\}_{i\in\mathcal{P}}$. For example, if the missing links are caused by under-reporting or by dropping links to friends with misspelled names, Assumption \ref{nondiff3} is satisfied if the under-reporting or spelling error made by one unit does not depend on those made by others.
Nevertheless, it does not require the missing links of the same unit to be independent of each other, and it allows the probability of $\mathbf{U}_i$ to vary with the true degree $\mathcal{T}^*_i$ and $Z_i$. Therefore, this assumption can accommodate arbitrary correlation among missing links of the same unit and heterogeneous missing rates across units.
See Lemma \ref{lemma_example_noniid} in Appendix \ref{appendix_example} for further illustration.\footnote{There are scenarios where Assumption \ref{nondiff3} may not hold. For example, when constructing networks, researchers may only include network links among sampled units or within certain geographic boundaries (e.g., schools, villages, etc.). In the former case, $U_{ik}=U_{jk}=0$ with probability one if unit $k$ is not sampled. In the latter case, $U_{ij}=U_{ji}=0$ with probability one if units $i$ and $j$ are not located within the same boundaries. In these cases, $\mathbf{U}_i$ and $\mathbf{U}_j$ are not independent, even when conditioned on $(\mathbf{A}^*,\mathbf{Z})$, which violates the exclusion restriction.}
Given the network sparsity in Assumption \ref{ass_support}, we can focus on
the truncated degree support, $\{0, ..., K\}$. Recall that our target in this section is the missing probabilities in the degree, $p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i}$. Denote by $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z}$ a $(K+1)\times (K+1)$ matrix that consists of all these probabilities in the truncated degree support:
\begin{align*}
\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}&=\{p_{\mathcal{T}^*_i|\mathcal{T}_i=l,Z_i=z}(k)\}_{k,l=0,...,K}=
\begin{bmatrix}p_{\mathcal{T}^*_i|\mathcal{T}_i=0,Z_i=z}(0)&\cdots&p_{\mathcal{T}^*_i|\mathcal{T}_i=K,Z_i=z}(0)\\
\vdots&\ddots&\vdots\\
p_{\mathcal{T}^*_i|\mathcal{T}_i=0,Z_i=z}(K)&\cdots&p_{\mathcal{T}^*_i|\mathcal{T}_i=K,Z_i=z}(K)
\end{bmatrix}.
\end{align*}
Define $$\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}=\{p_{\mathcal{T}_i|\mathcal{T}^*_i=l,Z_i=z}(k)\}_{k,l=0,...,K}\text{ and }\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}=\{p_{\tilde{\mathcal{T}}_i|\mathcal{T}^*_i=l,Z_i=z}(k)\}_{k,l=0,...,K}$$ in the same way as $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}$. Let $diag(v)$ be a diagonal matrix with the elements from vector $v$ on the principal diagonal. Define the following $(K+1)\times (K+1)$ matrices:
\begin{equation*}
\begin{aligned}
&\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}=\left\{p_{\mathcal{T}_i,\tilde{\mathcal{T}}_i|Z_i=z}(k,l)\right\}_{k,l=0,...,K}\\
&\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}=\left\{E[\varpi(Y_i)|\mathcal{T}_i=k,\mathcal{\tilde{T}}_i=l,Z_i=z]p_{\mathcal{T}_i,\tilde{\mathcal{T}}_i|Z_i=z}(k,l)\right\}_{k,l=0,...,K}\\
&\mathbf{T}_{\mathcal{T}^*|Z=z}=diag\left(p_{\mathcal{T}^*_i|Z_i=z}(0),\cdots,p_{\mathcal{T}^*_i|Z_i=z}(K)\right)\\
&\mathbf{T}_{\mathcal{T}|Z=z}=diag\left(p_{\mathcal{T}_i|Z_i=z}(0),\cdots,p_{\mathcal{T}_i|Z_i=z}(K)\right)\\
&\mathbf{T}_{Y|\mathcal{T}^*,Z=z}=diag\left(E[\varpi(Y_i)|\mathcal{T}^*_i=0,Z_i=z],\cdots,E[\varpi(Y_i)|\mathcal{T}^*_i=K,Z_i=z]\right)
,\end{aligned}\end{equation*}
where $\varpi:\Omega_Y\mapsto\mathbb{R}$ represents a user-specified function, to which we will impose additional restrictions in subsequent assumptions.\footnote{For example, the user-specified function can be $\varpi(y)=y$ (mean), $\varpi(y)=(y-E[Y_i])^2$ (variance), or $\varpi(y)=1[y\leq y_0]$ (quantile) for some given $y_0$. We omit the dependence of $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}$ and $\mathbf{T}_{Y|\mathcal{T}^*,Z=z}$ on $\varpi$ to ease the notation.} Define two $(K+1)\times 1$ vectors
\begin{align*}
\mathbf{F}_{\mathcal{T}^*|Z=z}&=[p_{\mathcal{T}^*_i|Z_i=z}(0),...,p_{\mathcal{T}^*_i|Z_i=z}(K)]',\text{ and }
\mathbf{F}_{\mathcal{T}|Z=z}=[p_{\mathcal{T}_i|Z_i=z}(0),...,p_{\mathcal{T}_i| Z_i=z}(K)]'.
\end{align*}
From Bayes' theorem, we can write the target matrix $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}$ as
\begin{align}\label{degree_dist_intermediate}
\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}=\mathbf{T}_{\mathcal{T}^*|Z=z}\times \mathbf{F}'_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{T}^{-1}_{\mathcal{T}|Z=z},
\end{align}
where $\mathbf{T}_{\mathcal{T}|Z=z}$ is identifiable because it is a matrix of distribution functions of observables. In what follows, we explore the identification of $\mathbf{T}_{\mathcal{T}^*|Z=z}$ and $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$.
Let $\mathbf{\Delta}_{K}$ be a matrix (or vector) with all its entries being $O(\triangle_K)$. Note that $\mathbf{\Delta}_{K}$ may
stand for different matrices (or vectors) at different places. Given Assumption \ref{ass_support} (network sparsity) and Assumption \ref{nondiff3} (exclusion restriction), applying the law of iterated expectation, we can show that
\begin{align}\label{decomp_ynn1}
\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}=&\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{T}_{Y|\mathcal{T}^*,Z=z}\times \mathbf{T}_{\mathcal{T}^*|Z=z}\times \mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}'+\mathbf{\Delta}_{K},\\
\label{decomp_nn1}
\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}=&\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{T}_{\mathcal{T}^*|Z=z}\times \mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}'+\mathbf{\Delta}_{K},\\
\label{decomp_n1}
\mathbf{F}_{\mathcal{T}|Z=z}=&\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{F}_{\mathcal{T}^*|Z=z}+\mathbf{\Delta}_{K},
\end{align}
where the three identifiable matrices on the left-hand side of \eqref{decomp_ynn1} to \eqref{decomp_n1} differ from the first terms on the right-hand side by $\mathbf{\Delta}_{K}$. Apparently, the differences arise from the ignorance of degree values
larger than $K$.
Based on the matrix perturbation theory,\footnote{See Lemma \ref{lemma_matrix_perturbation} for the matrix perturbation theory.} if all three matrices in the first term on the right-hand side of \eqref{decomp_nn1} are invertible with bounded inverses, then $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ is also invertible, and its inverse satisfies
\begin{align}\label{decomp_nn1_inv}
\mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}=&\left(\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}'\right)^{-1}\times \mathbf{T}^{-1}_{\mathcal{T}^*|Z=z}\times \mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z}+\mathbf{\Delta}_{K}.
\end{align}
By post-multiplying $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}$ in \eqref{decomp_ynn1} by $\mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ in \eqref{decomp_nn1_inv}, we can get
\begin{align}\label{t17}
\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}\times \mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}=\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{T}_{Y|\mathcal{T}^*,Z=z}\times \mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z=z}+\mathbf{\Delta}_{K}.
\end{align}
It then follows from \eqref{t17} and the properties of diagonalizable matrix that the normalized eigenvectors (whose entries sum to one) of $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}\times \mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ approximate the columns of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ with an approximation error of order $\triangle_{K}$. If we can further identify the ordering of these eigenvectors, then $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ can be approximated with an error of order $\triangle_{K}$. For the other unknown matrix in \eqref{degree_dist_intermediate}, $\mathbf{T}_{\mathcal{T}^*|Z=z}$, pre-multiplying both sides of \eqref{decomp_n1} by $\mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ yields the following equation:
\begin{align} \label{t16}
\mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{F}_{\mathcal{T}|Z=z}=\mathbf{F}_{\mathcal{T}^*|Z=z}+\mathbf{\Delta}_{K}.
\end{align}
Given that $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ can be approximated using identifiable matrices and $\mathbf{F}_{\mathcal{T}|Z=z}$ is directly identifiable from the data, it follows from \eqref{t16} that $\mathbf{T}_{\mathcal{T}^*|Z=z}$ is also approximated with an error of order $\triangle_{K}$. Consequently, the missing probabilities in the degree, $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}$ in \eqref{degree_dist_intermediate}, can be approximated.
Next, in Assumptions \ref{ass_eigen_inv} to \ref{ass_eigen}, we formalize the necessary assumptions for the matrix diagonalization method outlined above. Let $\underline{\sigma}(\mathbf{B})$ denote the smallest singular value of a matrix $\mathbf{B}$.
\begin{assumption}[\textbf{Invertibility}]\label{ass_eigen_inv}For some $\delta=\delta(K)>0$, we have $\underline{\sigma}(\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z})>\delta$, $\underline{\sigma}(\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z})>\delta$, and $\underline{\sigma}\left(\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}\right)>\delta$ for all $z\in\Omega_Z$.
\end{assumption}
Assumption \ref{ass_eigen_inv} requires that the choice of $K$ ensures the invertibility of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$, $\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}$, and $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$, as well as the boundedness of their inverses. Since $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ is identifiable using data on the observed degrees $\mathcal{T}$ and $\tilde{\mathcal{T}}$, this assumption can be partially verified for the chosen value of $K$ by checking the rank and the smallest singular value of $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$.
\begin{assumption}[\textbf{Eigen-decomposition}]\label{ass_eigen_unique}For some $e_Y>0$, we have $\sup\limits_{n^*\in\Omega_{\mathcal{T}^*},z\in\Omega_Z}|E[\varpi(Y_i)|\mathcal{T}^*_i=n^*,Z_i=z]|<e_Y$. In addition, for all $z\in\Omega_Z$, if $n\neq n'$, then $$E[\varpi(Y_i)|\mathcal{T}^*_i=n,Z_i=z]\neq E[\varpi(Y_i)|\mathcal{T}^*_i=n',Z_i=z].$$
\end{assumption}
Assumption \ref{ass_eigen_unique} rules out duplicated eigenvalues of the diagonalizable matrix $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{T}_{Y|\mathcal{T}^*,Z=z}\times \mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z=z}$, ensuring that its eigenvalues and eigenvectors are differentiable functions of the matrix itself. This smoothness condition further guarantees that the normalized eigenvectors of $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}\times \mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ approximate the columns of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ with an error of order $\triangle_K$. Note that if network interactions have no impact on $Y_i$, then $m^*(d,s,n,z)=m^*(d,z)$, and $E[\varpi(Y_i)|\mathcal{T}^*_i,Z_i]$ becomes degenerate in $\mathcal{T}^*_i$, causing Assumption \ref{ass_eigen_unique} to fail. Fortunately, in this case, the matrix diagonalization method is not needed, as $m^*$ can be point identified using its identifiable counterpart $m$. This is because $m^*(d,s,n,z)=m(d,s,n,z)=m^*(d,z)$ by Theorem \ref{prop_conditional_mean}. Then, a test for whether $m(d,s,n,z)$ depends on $(s,n)$ can be used to verify Assumption \ref{ass_eigen_unique} and determine the necessity of the matrix diagonalization method.
In Example \ref{example_unique_eigen}, we discuss another possible test for Assumption \ref{ass_eigen_unique} within commonly used network effect models.
\begin{assumption}[\textbf{Order of Eigenvectors}]\label{ass_eigen}
Any one of the following conditions holds for all $n^*\in\{0,...,K\}$ and $z\in\Omega_Z$.
\begin{itemize}
\item[(a)]$p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(n^*)>p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(n)$ for any $n\neq n^*$.
\item[(b)]
$p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(0)$ is strictly monotone in $n^*$ and the direction is known.
\item[(c)]
$E[\varpi(Y_i)|\mathcal{T}^*_i=n^*,Z_i=z]$ is strictly monotone in $n^*$ and the direction is known.
\end{itemize}
\end{assumption}
Assumption \ref{ass_eigen} is used to identify the order of the columns of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$. Note that \emph{any one} of the conditions in Assumption \ref{ass_eigen} is sufficient for identifying the order. Condition (a) implies that if the $l$-th entry of a column of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ is its largest entry, then this column is the $l$-th column of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$. Under condition (b) and a decreasing order of $p_{\mathcal{T}_i|\mathcal{T}^*_i=n^*,Z_i=z}(0)$ in $n^*$, if the first entry of a column of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ is the $l$-th largest among the first entries of all its columns, then this column is the $l$-th column. Condition (c) imposes an order on the eigenvalues of the diagonalizable matrix $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}\times \mathbf{T}_{Y|\mathcal{T}^*,Z=z}\times \mathbf{F}^{-1}_{\mathcal{T}|\mathcal{T}^*,Z=z}$, which also implies the order of its eigenvectors.
Researchers should carefully choose the appropriate condition in Assumption \ref{ass_eigen} based on the specific context. Conditions (a) and (b) assume that the observable degree is informative about the true degree. Lemma \ref{lemma_example_eigen} in Appendix \ref{appendix_example} provides sufficient conditions for both (a) and (b). For instance, condition (a) holds if more than half of the units have no missing links,\footnote{Similar restrictions are widely used in the measurement error literature \citep[e.g.][]{battistin2011misclassified,battistin2014misreported,chen2011nonlinear,huSchennach2008instrumental,lewbel2007estimation,mahajan2006identification}.} and it also holds under weaker conditions.
Condition (b) requires that the probability of having zero observed degree strictly decreases as the true degree increases. Condition (c) imposes a shape restriction on $\varpi(Y_i)$ and is satisfied in commonly used network effect models (see Example \ref{example_unique_eigen}).
\begin{example}\label{example_unique_eigen}Consider a simplified linear-in-means model with no endogenous peer effects and no covariates: $Y_i=\theta_1D_i+\theta_2\frac{S^*}{\mathcal{T}^*_i}+\theta_3\mathcal{T}^*_i+\varepsilon_i$. We also assume no isolated units in the true network for simplicity. Let us set $\varpi(y)=y$. Lemma \ref{lemma_monotone} in Appendix \ref{appendix_example} shows that
\begin{align*}
E[Y_i|\mathcal{T}^*_i=n]=&(\theta_1+\theta_2)p_{D_i}(1)+\theta_3n,\\ E[Y_i|\mathcal{T}_i=n]=&(\theta_1+\theta_2)p_{D_i}(1)+\theta_3\sum_{n^*\in\Omega_{\mathcal{T}^*}}n^*p_{\mathcal{T}^*_i|\mathcal{T}_i=n}(n^*).
\end{align*}
If $\theta_3=0$, then both $E[Y_i|\mathcal{T}^*_i=n]$ and $E[Y_i|\mathcal{T}_i=n]$ remain invariant to $n$, and Assumption \ref{ass_eigen_unique} is violated.
Thus, Assumption \ref{ass_eigen_unique} can be verified by a test on whether $E[Y_i|\mathcal{T}_i=n]$ depends on $n$. In addition, we can see that $E[Y_i|\mathcal{T}^*_{i}=n]$ is strictly increasing in $n$ if $\theta_3>0$ and strictly decreasing in $n$ if $\theta_3<0$, so that condition (c) in Assumption \ref{ass_eigen} holds as long as $\theta_3\neq0$.
\end{example}
Denote by $\mathbf{F}^a_{\mathcal{T}|\mathcal{T}^*,Z=z}$ a matrix whose columns are the normalized eigenvectors of $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}\times \mathbf{F}^{-1}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$, in the order implied by Assumption \ref{ass_eigen}. Based on \eqref{t17}, we can show that $\mathbf{F}^a_{\mathcal{T}|\mathcal{T}^*,Z=z}$ differs from $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ by a term of order $\triangle_K$.
Based on \eqref{t16}, let
\begin{align*}
\mathbf{F}^a_{\mathcal{T}^*|Z=z}:=&(\mathbf{F}^{a}_{\mathcal{T}|\mathcal{T}^*,Z=z})^{-1}\times \mathbf{F}_{\mathcal{T}|Z=z},~~\text{ and }~~
\mathbf{T}^a_{\mathcal{T}^*|Z=z}:=diag(\mathbf{F}^a_{\mathcal{T}^*|Z=z})
\end{align*}
be the approximation for $\mathbf{F}_{\mathcal{T}^*|Z=z}$ and $\mathbf{T}_{\mathcal{T}^*|Z=z}$, respectively.
Replacing $\mathbf{T}_{\mathcal{T}^*|Z=z}$ and $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ in \eqref{degree_dist_intermediate} with their approximations, we can get an approximation for $\mathbf{F}_{\mathcal{T}^*|\mathcal{T},Z=z}$:
\begin{align*}
\mathbf{F}^a_{\mathcal{T}^*|\mathcal{T},Z=z}=\mathbf{T}^a_{\mathcal{T}^*|Z=z}\times (\mathbf{F}^a_{\mathcal{T}|\mathcal{T}^*,Z=z})'\times \mathbf{T}^{-1}_{\mathcal{T}|Z=z}.
\end{align*}
The theorem below shows that the difference between the target matrix, the missing probabilities in the degree $\mathbf{F}_{\mathcal{T}^*|\mathcal{T}, Z=z}$, and its approximation $\mathbf{F}^a_{\mathcal{T}^*|\mathcal{T}, Z=z}$, is bounded by $\triangle_K$.
\begin{theorem}\label{theorem_id_N}Suppose Assumptions \ref{nondiff} and \ref{identical_degree} (b) hold for both $\mathcal{N}_i$ and $\tilde{\mathcal{N}}_i$. Under Assumptions \ref{unconf1}-\ref{ass_eigen}, we have
$\sup\limits_{z\in\Omega_Z}\left\|\mathbf{F}^a_{\mathcal{T}^*|\mathcal{T}, Z=z}-\mathbf{F}_{\mathcal{T}^*|\mathcal{T}, Z=z}\right\|=O(\triangle_K).$
\end{theorem}
\subsection{Identification of the CASF}\label{section:id_CASF}
In this section, we proceed to the identification of the CASF $m^*$ using $m$ and the weight.
Let us first introduce some notations. Denote $\mathcal{G}_i=(S_i,\mathcal{T}_i)\in\Omega_{\mathcal{G}}$ and $\mathcal{G}^*_i=(S^*_i,\mathcal{T}^*_i)\in\Omega_{\mathcal{G}^*}$. Focusing on the truncated degree support, we rank the possible values of $\mathcal{G}_i$ and $\mathcal{G}^*_i$ according to the lexicographical order of integers. For $\mathfrak{g}_k=(s,n)$, define $\{\mathfrak{g}_0,\mathfrak{g}_1,...,\mathfrak{g}_{K_\mathcal{G}}\}$ as follows:
\begin{equation}\label{lexi_order}
\begin{aligned}
&\mathfrak{g}_0=(0,0),\\
&\mathfrak{g}_1=(0,1),~\mathfrak{g}_2=(1,1),\\
&\mathfrak{g}_3=(0,2),~\mathfrak{g}_4=(1,2),~\mathfrak{g}_5=(2,2),\\
&\vdots\\
&\mathfrak{g}_{\frac{K(K+1)}{2}}=(0,K),\cdots,\mathfrak{g}_{K_\mathcal{G}}=(K,K).
\end{aligned}
\end{equation}
For $d\in\{0,1\}$ and $z\in\Omega_Z$, define two $(K_\mathcal{G}+1)\times 1$ vectors $\mathbf{M}_{d,z}$ and $\mathbf{M}_{d,z}^*$ as
\begin{equation}\begin{aligned}
\mathbf{M}_{d,z}=&\left[m(d,\mathfrak{g}_0,z),~m(d,\mathfrak{g}_1,z),\cdots,m(d,\mathfrak{g}_{K_\mathcal{G}},z)\right]',\\
\label{t8}
\mathbf{M}_{d,z}^*=&\left[m^*(d,\mathfrak{g}_0,z),~m^*(d,\mathfrak{g}_1,z),\cdots,m^*(d,\mathfrak{g}_{K_\mathcal{G}},z)\right]'.
\end{aligned}
\end{equation}
We denote $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$ as a $(K_\mathcal{G}+1)\times (K_\mathcal{G}+1)$ matrix that consists of the weights on the truncated degree support
\begin{align*}
\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}=&\begin{bmatrix}
p_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g}_0,Z_i=z}(\mathfrak{g}_0)&\cdots&p_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g}_{K_{\mathcal{G}}},Z_i=z}(\mathfrak{g}_0)\\
\vdots&\ddots&\vdots\\
p_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g}_0,Z_i=z}(\mathfrak{g}_{K_{\mathcal{G}}})&\cdots&p_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g}_{K_{\mathcal{G}}},Z_i=z}(\mathfrak{g}_{K_{\mathcal{G}}})\end{bmatrix}.
\end{align*}
If $m^*$ is bounded, based on Theorem \ref{prop_conditional_mean} and the sparse network assumption, we can show that
\begin{align}\label{def_approx_weights0_no_inverse}
\mathbf{M}_{d,z}
=&~\mathbf{F}'_{\mathcal{G}^*|\mathcal{G},Z=z}\times \mathbf{M}_{d,z}^*+\mathbf{\Delta}_{K}.
\end{align}Given the lexicographical order of $\mathfrak{g}_{k}$ and the presence of missing links, it is easy to see that $p_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g}_k,Z_i=z}(\mathfrak{g}_l)=0$ for any $k > l$, so that $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$ is a lower triangular matrix. In addition, all its diagonal elements are strictly positive and bounded away from zero under Assumption \ref{ass_eigen_inv}. Therefore, $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$ is invertible with a bounded inverse. Below, we introduce the notation for the approximation of $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$. Recall that by Theorem \ref{decomposition_weight}, we have
\begin{align*}
p_{\mathcal{G}^*_i|\mathcal{G}_i,Z_i=z}=p_{ S^*_i|\mathcal{T}^*_i, \mathcal{T}_i,Z_i=z}\times p_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i=z},
\end{align*}
where $p_{ S^*_i|\mathcal{T}^*_i, \mathcal{T}_i,Z_i=z}$ is a binomial distribution and is point identified as shown in Theorem \ref{lemma_iden}.
Let $p^a_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i=z}$ stand for the element of the matrix $\mathbf{F}^a_{\mathcal{T}^*|\mathcal{T},Z=z}$ in Theorem \ref{theorem_id_N}. Define
\begin{align*}
p^a_{\mathcal{G}^*_i|\mathcal{G}_i,Z_i=z}=p_{ S^*_i|\mathcal{T}^*_i, \mathcal{T}_i,Z_i=z}\times p^a_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i=z}.
\end{align*}
Stacking all $p^a_{\mathcal{G}^*_i|\mathcal{G}_i,Z_i=z}$ into a matrix, we can get an approximation for $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$:
\begin{align}\label{def_approx_weights}
\mathbf{F}^a_{\mathcal{G}^*|\mathcal{G},Z=z}=\left\{p^a_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g}_{k},Z_i=z}(\mathfrak{g}_l)\right\}_{k,l=0,...,K_{\mathcal{G}}}.
\end{align}
According to the matrix perturbation theory, for a sufficiently small $\triangle_K$, given that $\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}$ is invertible with a bounded inverse, its approximation $\mathbf{F}^a_{\mathcal{G}^*|\mathcal{G},Z=z}$ is also invertible with a bounded inverse.
Then, based on \eqref{def_approx_weights0_no_inverse}, the theorem below shows that pre-multiplying $\mathbf{M}_{d,z}$ by the inverse of $\mathbf{F}^{a'}_{\mathcal{G}^*|\mathcal{G},Z=z}$ gives us
an approximation of $\mathbf{M}_{d,z}^*$.
\begin{theorem}\label{theorem_id_CASF}Suppose $m^*(\cdot)$ is uniformly bounded in the support of its argument. If assumptions in Theorem \ref{theorem_id_N} hold, then we have
$$
\sup_{z\in\Omega_Z,d=0,1}\Big\|\left(\mathbf{F}^{a'}_{\mathcal{G}^*|\mathcal{G},Z=z}\right)^{-1}\times \mathbf{M}_{d,z}-\mathbf{M}_{d,z}^*\Big\|=O(\triangle_K).
$$
\end{theorem}
As a result of Theorem \ref{theorem_id_CASF}, if the true degree is uniformly bounded by $K$ for all units, we have $\triangle_K=0$ and the CASF is point identified.
\begin{corollary}[\textbf{Point Identification of CASF with Bounded Degree}]\label{corollary_id_CASF}Under assumptions in Theorem \ref{theorem_id_CASF}, if there exists some integer $0<K<\infty$ such that $\triangle_K=0$ holds, then
$\mathbf{M}_{d,z}^*$ is point identified for all $z\in\Omega_Z$ and $d=0,1$.
\end{corollary}
Some final remarks are in order. First, if no $K$ exists such that $\triangle_K=0$, it is still possible to approximate $m^*$ in the truncated degree support as long as $\triangle_K$ is sufficiently small. In this case, our method can yield less biased estimates of the effects of interest compared to the naive estimation method that ignores missing links.
Second, researchers should carefully choose the value of $K$ to ensure that (i) $\triangle_K$ is sufficiently small, and (ii) the rank condition of $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$, $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*,Z=z}$ and $\mathbf{F}_{\tilde{\mathcal{T}}|\mathcal{T}^*,Z=z}$ assumed in Assumption \ref{ass_eigen_inv} is satisfied. In finite samples, the approximation accuracy for $m^*$ depends on both the value of $\triangle_K$ and the credibility of the rank condition. Clearly, a larger $K$ results in a smaller $\triangle_K$ but a less credible rank condition. In practice, researchers can implement our proposed method using different values of $K$.
\section{Estimation and Inference}\label{section_estimation_multiple}
In this section, we propose a two-step semiparametric estimation method for the CASF $m^*$ and present its asymptotic properties.
All technical details are left to Appendix \ref{app_section_estimation}.
\subsection{Two-Step Estimation Method}\label{section_estimation}
Our estimation procedure consists of two steps. First, we obtain estimators for the weights by estimating the dependence of NBRVs and the missing probabilities in the degree using a kernel estimation approach. Second, we parameterize $m^*$ and apply a least-square estimation by plugging in the first-step estimators of the weights. Imposing parametric structures on $m^*$ still allows for flexible heterogeneity in the treatment and spillover effects, which can be captured through interactions between variables and their polynomials
\vspace{0.2cm}
\noindent\textbf{Step 1. Kernel Estimation for the Weights}. \hspace{0.1cm} In the first step, we present a kernel estimation method to obtain estimators of the weights. Recall that the estimation for the NBRV dependence requires estimating $p_{D_i}$, and the estimation for the missing probabilities in the degree using the matrix diagonalization method requires estimating $\mathbf{E}_{\mathcal{T},\tilde{\mathcal{T}},Y|Z=z}$, $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}|Z=z}$ and $\mathbf{F}_{\mathcal{T}|Z=z}$ in \eqref{decomp_ynn1}, \eqref{decomp_nn1}, and \eqref{decomp_n1}.\footnote{The estimated eigenvectors in the matrix diagonalization method may contain complex values. As mentioned in \citet{hu2008identification}, since all the latent probabilities and densities are real and positive, we can take the real part of the estimated eigenvectors, and the probability of getting a complex value goes to zero as the sample size increases.}
Let $\gamma=\gamma(z)=[\gamma_1(z),\gamma_2(z),\gamma_3(z),\gamma_4(z)]'$ be a vector that consists of all the elements required to estimate the weights, where
\begin{equation}\label{def_kernel_estimand}
\begin{aligned}
&\gamma_1(z)=\Big[E[\varpi(Y_i)|\mathcal{T}_i=0,\mathcal{\tilde{T}}_i=0,Z_i=z],...,E[\varpi(Y_i)|\mathcal{T}_i=k,\mathcal{\tilde{T}}_i=l,Z_i=z],\\
&\hspace{8cm}...,E[\varpi(Y_i)|\mathcal{T}_i=K,\mathcal{\tilde{T}}_i=K,Z_i=z]\Big],\\
&\gamma_2(z)=\left[p_{\mathcal{\mathcal{T}}_i,\tilde{T}_i,Z_i}(0,0,z),...,p_{\mathcal{T}_i,\mathcal{\tilde{T}}_i,Z_i}(k,l,z),...,p_{\mathcal{T}_i,\mathcal{\tilde{T}}_i,Z_i}(K,K,z)\right],\\
&\gamma_3(z)=\left[p_{\mathcal{T}_i,Z_i}(0,z),...p_{\mathcal{T}_i,Z_i}(K,z)\right],\\
&\gamma_4(z)=\left[p_{Z_i}(z),p_{D_i}(1)\right].
\end{aligned}
\end{equation}
Let $W_i=(W^{c'}_i,W^{d'}_i)'\in\Omega_{W^c}\times\Omega_{W^d}$ and $W_i\subseteq(\mathcal{T}_i,\tilde{\mathcal{T}}_i,D_i,Z_i)$ denote a vector of observable variables that will be used to compute the kernel estimator for $\gamma$, where $W^c_i:=(W^c_{i,1},...,W^c_{i,Q})'$ is a $Q\times1$ vector of continuous variables in $W_i$ (if any), and $W^d_i$ contains discrete variables in $W_i$.
For example, for $\gamma_2(z)$, $W_i=(\mathcal{\mathcal{T}}_i,\tilde{T}_i,Z_i)$. For $\forall w=(w^{c'},w^{d'})'\in\Omega_{W^c,W^d}$, we define the estimator for $p_{W_i}(w)$ and $E[\varpi(Y_i)|W_i=w]$ as below:
\begin{equation}\label{t_ker}
\begin{aligned}
&\hat{p}_{W_i}(w)
:=\frac{1}{N}\sum_{i=1}^N\hat{p}^{ker}_i(w),\;\text{ and }\;
\hat{E}[\varpi(Y_i)|W_i=w]:=\frac{\frac{1}{N}\sum_{i=1}^{N}\varpi(Y_i)\hat{p}^{ker}_i(w)}{\frac{1}{N}\sum_{i=1}^N\hat{p}^{ker}_i(w)},
\end{aligned}
\end{equation}
where $\hat{p}^{ker}_i(w):=\frac{1}{h^{Q}}\prod_{q=1}^{Q}\kappa\left(\frac{W^c_{i,q}-w^c_q}{h}\right)1\left[W^d_i=w^d\right]$, with a bandwidth $h>0$ and a univariate kernel function $\kappa(\cdot)$.\footnote{
A data-driven method for bandwidth selection is possible but it is not the focus of this paper.} Then, the estimator for $\gamma$, denoted by $\hat{\gamma}_N$, can be obtained by replacing the conditional means and probabilities in \eqref{def_kernel_estimand} with their sample analogs in \eqref{t_ker}. Given $\hat{\gamma}_N$, we can estimate $\mathbf{F}^a_{\mathcal{G}^*|\mathcal{G},Z=z}$ defined in \eqref{def_approx_weights}. Let $vec(B)$ be the vectorization of a matrix $B$. Define
\begin{align}\label{expression_phi}
\phi=&vec\left(\mathbf{F}_{\mathcal{G}^*|\mathcal{G},Z=z}\right),~~
\phi^a=vec\left(\mathbf{F}^a_{\mathcal{G}^*|\mathcal{G},Z=z}\right),~~\text{ and }~~\hat{\phi}_N=vec\left(\hat{\mathbf{F}}^a_{\mathcal{G}^*|\mathcal{G},Z=z}\right),
\end{align}
where each element in $\hat{\mathbf{F}}^a_{\mathcal{G}^*|\mathcal{G},Z=z}$ is obtained by $$\hat{p}^a_{\mathcal{G}^*_i|\mathcal{G}_i,Z_i=z}=\hat{p}_{ S^*_i|\mathcal{T}^*_i, \mathcal{T}_i,Z_i=z}\times \hat{p}^a_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i=z},$$
with $\hat{p}_{ S^*_i|\mathcal{T}^*_i, \mathcal{T}_i,Z_i=z}=\hat{p}_{\Delta S_i|\Delta\mathcal{T}_i,Z_i=z}\sim Binomial(\hat{p}_{D_i}(1),\Delta\mathcal{T}_i)$ and $\hat{p}^a_{\mathcal{T}^*_i|\mathcal{T}_i,Z_i=z}$ being estimated by applying the matrix diagonalization method using $\hat{\gamma}_N$, as outlined in Section \ref{section_matrix_diagonalization}. We suppress the argument $z$ in $\gamma$, $\phi$, $\hat{\gamma}_N$, and $\hat{\phi}_N$ for notation simplicity, unless otherwise mentioned. Let $\gamma^0$ and $\phi^0$ be the true value for $\gamma$ and $\phi$.
\vspace{0.5cm}
\noindent\textbf{Step 2. Semiparametric Estimation for the CASF}. \hspace{0.1cm} For any parameter $\beta$, denote $d_\beta=dim(\beta)$. In the second step, we parameterize $m^*(\cdot)=m^*(\cdot;\theta)$ to be a known function up to an unknown parameter $\theta\in\Theta\subseteq\mathbb{R}^{d_\theta}$, and we estimate $\theta$ using a plug-in estimator. Recall $\mathcal{G}_i=(S_i,\mathcal{T}_i)\in\Omega_{\mathcal{G}}$ and $\mathcal{G}^*_i=(S^*_i,\mathcal{T}^*_i)\in\Omega_{\mathcal{G}^*}$. Denote $X^*_i=(D_i,\mathcal{G}^*_i,Z_i)'$ and $X_i=(D_i,\mathcal{G}_i,Z_i)'$. Let $x_j=(d,\mathfrak{g}_j,z)$ with $j=0,...,K_{\mathcal{G}}$. Given the parameterization of $m^*$, let us rewrite $\mathbf{M}_{d,z}^*$ in \eqref{t8} as
\begin{align*}
\mathbf{M}_{d,z}^*(\theta)=&\big[m^*(x_0;\theta),~m^*(x_1;\theta),\cdots,m^*(x_{K_\mathcal{G}};\theta)\big]'.
\end{align*}
Proposition \ref{lemma_unconf1} implies the existence of some true value $\theta^0\in \Theta$ such that
\begin{align}\label{estimation_mom0}
E\left[Y_i-m^*(X_i^*;\theta^0)\big|X^*_i\right]=0.
\end{align}
Without loss of generality, we assume that $\theta^0$ is the unique solution to Equation \eqref{estimation_mom0}.\footnote{It rules out the existence of two different pairs, $(\dot{m}^{*},\dot{\theta}^{0})$ and $(\ddot{m}^{*},\ddot{\theta}^0)$, that satisfy $m^*(\cdot)=\dot{m}^{*}(\cdot;\dot{\theta}^{0})=\ddot{m}^{*}(\cdot;\ddot{\theta}^{0})$. It ensures the parametric point-identification of $\theta^0$ if the true NBRVs are observable.} However, this moment condition cannot be used to estimate $\theta^0$ because $X^*_i$ contains unobservable NBRVs.
Fortunately, we have $m(X_i;\theta^0)=E\left[Y_i|X_i\right]$ by definition of $m$ in \eqref{identfiable_m}, where, for $x=(d,\mathfrak{g},z)\in \Omega_X$ and the true weights $p^0_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g},Z_i=z}(\cdot)$, we have that $\theta$ enters $m$ through the mixture model:
\begin{align}\label{m_on_whole_supp}
m(x;\theta)
=\sum_{\mathfrak{g}^*\in\Omega_{\mathcal{G}^*}}m^*(d,\mathfrak{g}^*,z;\theta)\times p^0_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g},Z_i=z}(\mathfrak{g}^*).
\end{align}
Thus, we can obtain a moment condition based on observed NBRVs
$$E\left[Y_i-m(X_i;\theta^0)\big|X_i\right]=0.$$
Let $\tau_i=1[X_i\in\mathbf{X}]$ be a fixed trimming indicator for the observed NBRVs in the truncated support, where $\mathbf{X}=\{x=(d,\mathfrak{g},z)\in\Omega_X:~\mathfrak{g}\in\{\mathfrak{g}_0,...,\mathfrak{g}_{K_{\mathcal{G}}}\}\}$. Then, we can obtain an unconditional moment condition
\begin{align*}
E\left[\tau_i\left(Y_i-m(X_i;\theta^0)\right)\right]=0.
\end{align*}
Based on the unconditional moment equation, we define the population objective function $\mathcal{L}^0_{\mathbb{P}}(\theta)$ and its minimizer $\theta^0$ as follows:
\begin{align}\label{estimation_uncond_obj}
\theta^0=\arg\min_{\theta\in\Theta}\mathcal{L}^0_{\mathbb{P}}(\theta),~~ \text{ and }~~ \mathcal{L}^0_{\mathbb{P}}(\theta)=E\left\{\tau_i\left[Y_i-m\left(X_i;\theta\right)\right]^2\right\},
\end{align}
where, under the full rank condition on the Hessian matrix of $\mathcal{L}^0_{\mathbb{P}}(\theta)$ introduced later, $\theta^0$ is the unique solution to the minimization problem. Due to the possibility of unbounded degree, in the estimation for $\theta^0$, we need to further replace $m(x;\theta)$ in \eqref{estimation_uncond_obj}, defined in \eqref{m_on_whole_supp} on the whole support $\Omega_{\mathcal{G}^*}$, with its approximation $m^a(x;\theta,\phi)$ defined on the truncated support $\{\mathfrak{g}_0,...,\mathfrak{g}_{K_\mathcal{G}}\}$, where
\begin{align*}
&m^a(x;\theta,\phi)
=\sum_{\mathfrak{g}^*\in \{\mathfrak{g}_0,...,\mathfrak{g}_{K_\mathcal{G}}\}}m^*(d,\mathfrak{g}^*,z;\theta)\times p_{\mathcal{G}^*_i|\mathcal{G}_i=\mathfrak{g},Z_i=z}(\mathfrak{g}^*).\nonumber
\end{align*}
Given $m^a(x;\theta,\phi)$, let us define $\theta^a\in\Theta$ to be the pseudo-true value that solves the moment equation $E\left\{\tau_i\left[Y_i-m^a\left(X_i;\theta,\phi^a\right)\right]\right\}=0$ with $\phi^a=vec\left(\mathbf{F}^a_{\mathcal{G}^*|\mathcal{G},Z=z}\right)$. Then, we have
\begin{align*}
\theta^a=&\arg\min_{\theta\in\Theta}\mathcal{L}^a_{\mathbb{P}}(\theta,\phi^a),~~ \text{ and }~~\mathcal{L}^a_{\mathbb{P}}(\theta,\phi^a)=E\left\{\tau_i\left[Y_i-m^a\left(X_i;\theta,\phi^a\right)\right]^2\right\}, \end{align*}
where, under the full rank condition of the Hessian matrix of $\mathcal{L}^a_{\mathbb{P}}(\theta,\phi^a)$ introduced later, $\theta^a$ is the unique solution to the minimization problem. Note that the pseudo-true parameter $\theta^a$ may differ
from the true value $\theta^0$ for any nonzero $\triangle_K$, and its value may depend on $K$. We omit this dependence for notation simplicity.
Then, the plug-in estimator $\hat{\theta}_N$ is defined as the minimizer of the sample objective function:
\begin{align*}
\hat{\theta}_N=&\arg\min_{\theta\in\Theta}\mathcal{L}^a_N(\theta,\hat{\phi}_N),~~ \text{ and }~~\mathcal{L}^a_N(\theta,\hat{\phi}_N)=\frac{1}{N}\sum_{i=1}^N\tau_i\left[Y_i-m^a\left(X_i;\theta,\hat{\phi}_N\right)\right]^2,
\end{align*}
where we replace $\phi^a$ with its estimator $\hat{\phi}_N=vec\left(\hat{\mathbf{F}}^a_{\mathcal{G}^*|\mathcal{G},Z=z}\right)$.
\subsection{Asymptotic Properties}\label{section_asym_prop}
In this section, we discuss the asymptotic properties of our proposed estimator.
Additional regularity assumptions are provided in Appendix \ref{app_section_estimation_details}.
For any vector $a\in\mathbb{R}^p$, let $\|a\|$ be its Euclidean norm and $\|a\|_\infty=\max_{1\leq k\leq p}|a_k|$. For a matrix $\mathbf{B}$, let $\|\mathbf{B}\|=[tr(\mathbf{B}'\mathbf{B})]^{1/2}$ be the entry-wise matrix norm.
Let us partition the index set of all sampled units into $q_N$ mutually exclusive clusters. Denote these clusters as $\mathbb{S}_{1},...,\mathbb{S}_{q_N}$, where $\cup_{1\leq k\leq q_N}\mathbb{S}_k=\{1,...,N\}$.
Let $\tilde{W}_i=(Y_i,X_i')'\in\Omega_{\tilde{W}}$ represent a vector of observable variables, including the outcome. For any generic measurable function $b:\Omega_{\tilde{W}}\mapsto\mathbb{R}^{d_b}$, denote the within-cluster correlation as
\begin{align}\label{cov_dep}
\Sigma^b_N=\sum_{k=1}^{q_N}\sum_{i,j\in\mathbb{S}_k}Cov\left(b(\tilde{W}_i),b(\tilde{W}_j)\right).
\end{align}
To control data correlation under network interactions, we introduce a modified \emph{dependency neighborhood} assumption from \citet{chandrasekhar2021network} in Assumption \ref{ass_dependency_neighbor} below, which restricts the data dependence to be local.\footnote{The literature on inference using network data is growing rapidly \citep[see, e.g.,][]{hudgens2008toward,leung2020dependence}. Our assumption on data dependence is similar to those that limit data dependence to be weak or local \citep[e.g.][]{kojevnikov2019limit,leung2019causal}.}
Let $\bar{r}_N=\max_{1\leq k\leq q_N}|\mathbb{S}_{k}|$ be the size of the largest cluster.
\begin{assumption}\label{ass_dependency_neighbor}$\bar{r}_N=O(1)$ is a bounded value. For any measurable function $b:\Omega_{\tilde{W}}\mapsto\mathbb{R}^{d_b}$,
$$\Big\|\sum_{k=1}^{q_N}\sum_{i\in\mathbb{S}_k,j\not\in\mathbb{S}_k}Cov\left(b(\tilde{W}_i),b(\tilde{W}_j)\right)\Big\|=o(\|\Sigma^b_N\|).$$
\end{assumption}
This assumption requires that all clusters consist of a bounded number of units. Therefore, it implies that $q_N\rightarrow\infty$ as $N\rightarrow\infty$. In addition, it assumes that the correlation between units in different clusters is not necessarily zero but is weaker than the correlation between units within the same cluster. Units in different clusters can be correlated, for example, due to network interactions, spillovers of unobservables, or spatial and other forms of dependence. Note that the bounded cluster size does not require the maximal true degree to be bounded. Network connections across clusters are allowed as long as the network sparsity in Assumption \ref{ass_support} and the local dependence in Assumption \ref{ass_dependency_neighbor} hold.
Next, let us introduce a dependence coefficient analogous to the strong mixing coefficient of a stochastic process. Suppose the $q_N$ clusters can be ordered in a specific manner, based on, for example, social
or geographic proximity, so that units in clusters with distant indices are less likely to be correlated with each other. Without loss of generality, we assume that this order of clusters is given by $\mathbb{S}_{1},...,\mathbb{S}_{q_N}$. It is worth noting that we do not require this order to be known to researchers. Define the dependence coefficient as
\begin{align*}
\alpha_{k}=\sup_{\mathcal{A}\in\mathcal{F}_{1}^{k-2},\mathcal{B}\in\mathcal{F}_{k}^{k}}\left|Pr(\mathcal{A},\mathcal{B})-Pr(\mathcal{A})Pr(\mathcal{B})\right|,
\end{align*}
where $\mathcal{F}_{1}^{k-2}=\sigma(\{\tilde{W}_i,~i\in\bigcup_{1\leq l\leq k-2}\mathbb{S}_l\})$ and $\mathcal{F}_{k}^{k}=\sigma(\{\tilde{W}_i,~i\in\mathbb{S}_k\})$ for $k=1,2,...,q_N$ are two $\sigma$-fields. We use $\{\alpha_{k}\}_{k=1}^{q_N}$ to control the rate at which the dependence among clusters decays, which is crucial to establish the uniform convergence of the first-step estimators.
Below, we impose some restrictions on the dependence coefficient $\alpha_{k}$. With notation abuse, let $Q$ denote the number of continuous variables in $\tilde{W}_i=(Y_i,X_i')'$.
\begin{assumption}[\textbf{Local Dependence}]\label{ass_alpha}For $L_N=[N/(\ln(N)h^{Q+2})]^{Q/2}$, we have the following condition holds
$$\sum_{N=1}^{\infty}\Psi_N<\infty,\text{ where } \Psi_N=L_N\left(\frac{N}{\ln(N)}\right)^{1/5}\sum_{k=1}^{q_N}\alpha^{4/5}_{k}.$$
\end{assumption}
Assumption \ref{ass_alpha} assumes that the clusters are ordered so that units in $\bigcup_{1\leq l\leq k-2}\mathbb{S}_l$ and in $\mathbb{S}_k$ tend toward being independent as the sample size increases, allowing for nonzero but decreasing local dependence across clusters. This assumption is trivially satisfied if all clusters are mutually independent, indicating that units only form networks within each cluster. In such a case, every unit has a bounded degree. This assumption also holds when a limited number of units are correlated with others from nearby clusters, so that $\alpha_{k}$ goes to zero fast enough to ensure that $\Psi_N$ is summable. In this case, the data correlation may be caused by the network interactions of a few `star' units.
Recall that $\gamma=(\gamma_1(z), \gamma_2(z), \gamma_3(z), \gamma_4(z))'$. Let $\gamma_{jl}(z)$ with $j=1,2,3,4$ be the $l$-th element in $\gamma_j(z)$. Because $\gamma$ is a function of $z$, we define $\|\gamma-\gamma^0\|_\infty=\max_{j,l}\sup_{z\in\Omega_Z}|\gamma_{jl}(z)-\gamma^0_{jl}(z)|$. The same norm is defined for $\phi$.
\begin{theorem}[\textbf{Uniform Convergence}]\label{theorem_ker}
Suppose assumptions in Theorem \ref{theorem_id_CASF}, Assumptions \ref{ass_dependency_neighbor}, \ref{ass_alpha}, and Assumption \ref{ass_ker} in Appendix \ref{app_section_estimation_details} hold. If $h\rightarrow0$, $Nh^Q\rightarrow\infty$, and $\ln(N)/(Nh^Q)\rightarrow0$, then
\begin{itemize}
\item[(a)] $\left\|\hat{\gamma}_N-\gamma^0\right\|_\infty=O_p(\left[\ln(N)/(Nh^Q)\right]^{1/2}+h^2)$;
\item[(b)] for $\epsilon\rightarrow0$ as $N\rightarrow\infty$, we have
\begin{align*}
\sup\limits_{\|\hat{\gamma}_N-\gamma^0\|_\infty\leq\epsilon}\|\hat{\phi}_N-\phi^a\|_\infty=&O_p(\left\|\hat{\gamma}_N-\gamma^0\right\|_\infty),\\
\sup\limits_{\|\hat{\gamma}_N-\gamma^0\|_\infty\leq\epsilon}\|\hat{\phi}_N-\phi^0\|_\infty=&O_p(\left\|\hat{\gamma}_N-\gamma^0\right\|_\infty+\triangle_K).
\end{align*}
\end{itemize}
\end{theorem}
Theorem \ref{theorem_ker} shows that the convergence of the estimated weights $\hat{\phi}_N$ to the true value $\phi^0$ is driven by two factors: the convergence rate of the kernel estimator $\hat{\gamma}_N$, and the approximation error of the matrix diagonalization method measured by $\triangle_K$.
\begin{theorem}[\textbf{Consistency}]\label{theorem_consistency}Suppose $\frac{\partial^2\mathcal{L}^0_{\mathbb{P}}(\theta)}{\partial\theta\partial\theta'}$ and $\frac{\partial^2\mathcal{L}^a_{\mathbb{P}}(\theta,\phi^a)}{\partial\theta\partial\theta'}$ are both full rank for all $\theta\in\Theta$. Under assumptions in Theorem \ref{theorem_ker} and Assumption \ref{ass_consistency_m} in Appendix \ref{app_section_estimation_details}, we have
\begin{align*}\|\hat{\theta}_N-\theta^a\|= o_p(1),~~
\|\theta^a-\theta^0\|= O(\triangle_K),\text{ and }\|\hat{\theta}_N-\theta^0\|= O_p(\triangle_K).\end{align*}
\end{theorem}
Theorem \ref{theorem_consistency} demonstrates that $\hat{\theta}_N$ is a consistent estimator for the pseudo-true parameter $\theta^a$, while its asymptotic bias with respect to the true value $\theta^0$ is governed by $$B_K:=\theta^0-\theta^a,~\text{ where }~\|B_K\|=O(\triangle_K).$$
We know that if the true degree is bounded by $K$ for all units, then $\|B_K\|=0$, and $\theta^0$ is consistently estimated. Next, we consider the asymptotic normality. Let $g(\tilde{W}_i;\theta,\phi)=\tau_i[Y_i-m^a(X_i;\theta,\phi)]\frac{\partial m^a(X_i;\theta,\phi)}{\partial\theta}$ be the score function of the sample objective function $\mathcal{L}^a_N(\theta,\phi)$. We can show that
$$\frac{1}{\sqrt{N}}\sum_{i=1}^{N}g(\tilde{W}_i;\theta^a,\hat{\phi}_N)= \frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left[g(\tilde{W}_i;\theta^a,\phi^a)+\delta(\tilde{W}_i;\theta^a,\phi^a)\right]+o_p(1),$$
where $\delta(\tilde{W}_i;\theta^a,\phi^a)$ is the correction term to adjust the estimation error of the first-step kernel estimator.
Denote a $d_\theta\times1$ vector $\tilde{g}_i=g(\tilde{W}_i;\theta^a,\phi^a)+\delta(\tilde{W}_i;\theta^a,\phi^a)$, where $\tilde{g}_i=(\tilde{g}_{i,1},...,\tilde{g}_{i,d_\theta})'$. Following \eqref{cov_dep}, define the within-cluster correlation for $\tilde{g}_i$ as
$$\Sigma^{\tilde{g}}_N=\sum_{k=1}^{q_N}\sum_{i,j\in\mathbb{S}_k}Cov(\tilde{g}_i,\tilde{g}_j).$$
\begin{theorem}[\textbf{Asymptotic Normality}]\label{theorem_normality}Suppose assumptions in Theorem \ref{theorem_consistency} and Assumptions \ref{ass_normality} to \ref{ass_stein} in Appendix \ref{app_section_estimation_details} hold. If further assume $\ln(N)/(N^{1/2}h^Q)\rightarrow0$ and $Nh^4\rightarrow0$ as $N\rightarrow\infty$, then
$$\sqrt{N}(\hat{\theta}_N-\theta^0+B_K)\overset{d}{\rightarrow}\mathbb{N}(0,H^{-1}\Omega H^{-1}),$$
where $H=E\left[\frac{\partial g(\tilde{W}_i;\theta^a,\phi^a)}{\partial\theta'}\right]$, $\Omega=\lim\limits_{N\rightarrow\infty}\Sigma^{\tilde{g}}_N/N$, and $\mathbb{N}$ stands for a normal distribution.
\end{theorem}
Theorem \ref{theorem_normality} implies that the bias term $B_K$ is negligible in the inference for $\theta^0$ if $\sqrt{N}\|B_K\|=O(\sqrt{N}\triangle_K)$ is sufficiently small.
Theoretically, a consistent estimator of $H^{-1}\Omega H^{-1}$ can be obtained by replacing $H$ and $\Omega$ with their sample analogs.
However, it is difficult to implement because the explicit formula for $\delta(\cdot;\theta,\phi)$, although exists, is complex. In practice, we suggest using the method of numerical differentiation of the
influence function, as discussed in \citet{newey1994kernel}, to estimate the correction term without specifying its analytic expression.\footnote{See \citet{hong2015extremum} for discussions on the choice of numerical step size for the differentiation.
}
\section{Numerical and Empirical Results}\label{section_numerical_empirical}
\subsection{Monte Carlo Simulation}\label{section_simulation}
In this section, we illustrate the finite-sample behavior of our method via Monte Carlo simulations. We consider two data generating processes (DGPs) for the outcome variable:
\begin{align}\label{DGP_sim}
\text{(\textbf{Model 1}) }~~Y_i= & \theta_1+\theta_2D_i+\theta_3\frac{S^*_i}{\mathcal{T}^*_i}+\theta_4D_i*\frac{S^{*}_i}{\mathcal{T}^*_i}+\theta_5\mathcal{T}^*_i+\varepsilon_i,\\
\label{DGP_sim_2}\text{(\textbf{Model 2}) }~~Y_i= & \theta_1+\theta_2D_i+\theta_3S^*_i+\theta_4S^{*2}_i+\theta_5\mathcal{T}^*_i+\varepsilon_i,
\end{align}
where, in both models, $D_i\overset{i.i.d.}{\sim}Bernoulli(0.3)$ is a randomized treatment and $\varepsilon_i\overset{i.i.d.}{\sim}\mathbb{N}(0,0.25)$ is an idiosyncratic error. We set $\theta=(\theta_1,\theta_2,\theta_3,\theta_4,\theta_5)'=(1,1,0.5,-0.1,1)'$. We generate data using sample size $N\in\{1000,2000,5000\}$ with replications $M=1000$.
\vspace{0.2cm}
\noindent
\textbf{True Network Data}. \hspace{0.1cm} We simulate the true network data using the model below:
\begin{align}\label{MC_network_DGP}
A^*_{ij}= & 1[\beta_1+\beta_2(\alpha_i+\alpha_j)-d(\rho_i,\rho_j)+\zeta_{ij}>0],\text{ for all }i,j=1,...,N,
\end{align}
\sloppy
where $\alpha_i\overset{i.i.d.}{\sim}Bernoulli(0.5)$ stands for unobserved degree heterogeneity, $\rho_i=(\rho_{i1},\rho_{i2})\overset{i.i.d.}{\sim} Uniform([0,1]^2)$ is the random location of unit $i$, and $\zeta_{ij}=\zeta_{ji}\overset{i.i.d.\text{ across }(i,j)}{\sim}\mathbb{N}(0,1)$ is a random shock. $\alpha_i$, $\rho_i$, and $\zeta_{ij}$ are mutually independent. Let $d(\rho_i,\rho_j)$ be the distance between two units, where
$d(\rho_i,\rho_j)=0$ if $\|\rho_i-\rho_j\|\leq r$ and $d(\rho_i,\rho_j)=\infty$ otherwise.
We set $r=(r_{deg}/N)^{1/2}$ with $r_{deg}=3$ and $(\beta_1,\beta_2)=(-0.25,0.25)$. In this DGP design, the mean degree value is approximately 4 to 5, and the maximum degree value is about 14 to 15, as $N$ increases from $1000$ to $5000$.\footnote{We also conduct Monte Carlo simulations using an extension of model \eqref{MC_network_DGP} to allow for strategic network interactions, where one unit's link formation depends on the links of others. Due to space limitation, we present these simulation results in Appendix \ref{appendix_simulation_add}. }
\vspace{0.2cm}
\noindent\textbf{Observed Network Data}. \hspace{0.1cm} We generate the observed network data using
$
A_{ij}=U_{ij}A^*_{ij}
$.
We consider three different DGPs for $U_{ij}$. The first DGP considers the case of missing completely at random:
$$\text{(DGP1. random missing) }\quad U_{ij}\sim Bernoulli\left(p_U\right)\text{ are i.i.d. across all $(i,j)$}.$$
In the second DGP, the missing rate is heterogeneous and varies with the true degree value:
\begin{align*}
\text{(DGP2. heterogeneous missing) }\quad U_{ij}\sim Bernoulli &\left(p_{U,i}\right)\text{ are independent across all $(i,j)$,}
\end{align*}
where $p_{U,i}=p_U+0.02*\log(\mathcal{T}^*_i+1)$. In the third DGP, missing indicators of the same unit $i$, $\mathbf{U}_i=(U_{i1},...,U_{iN})$, are correlated, while $\mathbf{U}_i$ and $\mathbf{U}_j$ are independent for all $i\neq j$:
\begin{align*}
\text{(DGP3. dependent missing) }\quad &U_{ij}=1[\Phi(U^*_{ij})>p_U]\text{ and }U^*_{ij}=\sqrt{1-\rho^2}*e_{ij}+\rho*\omega_i,
\end{align*}
where $\Phi(\cdot)$ denotes the standard normal CDF, $e_{ij}\overset{i.i.d.\text{ across }(i,j)}{\sim}\mathbb{N}(0,1)$, $\omega_i\overset{i.i.d.}{\sim} \mathbb{N}(0,1)$, and $e_{ij}$ and $\omega_i$ are mutually independent. In DGP3, the value of $\rho$ determines the correlation among $(U_{i1},...,U_{iN})$, and we set $\rho=0.1$. In DGP1 to DGP3, we consider $p_U\in\{0.1,0.2,0.3\}$.
\vspace{0.2cm}
\noindent\textbf{Estimation Methods}. \hspace{0.1cm} We compare three estimation methods, including
(1) Infeasible OLS -- the infeasible OLS regression that uses the true NBRVs;
(2) Naive OLS -- the feasible OLS regression that uses the observable NBRVs and ignores the missing links;
(3) SPE -- the semiparametric estimation based on the matrix diagonalization method that uses the observed incoming and outgoing degrees. For SPE method, we set $\varpi(y)=y$ and choose the value of $K$ so that the smallest singular value of $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}}$ is larger than 0.001. We order the estimated eigenvectors of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*}$ according to Assumption \ref{ass_eigen} (c) with $E[\varpi(Y_i)|\mathcal{T}^*_i=n^*]$ strictly increasing in $n^*$.\footnote{For Model 1 in \eqref{DGP_sim}, Assumption \ref{ass_eigen} (c) holds because $\theta_5>0$ (see Example \ref{example_unique_eigen}). For Model 2 in \eqref{DGP_sim_2}, because $\mathcal{S}^*_i|\mathcal{T}^*_i=n^*$ follows a $Binomial(p_{D_i}(1),n^*)$ distribution, we can get $E[\varpi(Y_i)|\mathcal{T}^*_i=n^*]=\theta_1+\theta_2p_{D_i}(1)+\theta_3n^*p_{D_i}(1)+\theta_4[n^*p_{D_i}(1)p_{D_i}(0)+(n^*p_{D_i}(1))^2]+\theta_5n^*$. Given the value of $\theta$, Assumption \ref{ass_eigen} (c) holds as long as $n^*$ is smaller than 62, which is a much larger value than the maximum degree in our DGP.}
\vspace{0.2cm}
\noindent\textbf{Estimation Results}. \hspace{0.1cm} Our target parameter is the spillover effect $\eta^0:=\eta_S(d,s,0,n,z)= m^*(d,s,n,z)-m^*(d,0,n,z)$ at $d=0$, $s=1$, $s'=0$, and $n=4$ (no covariate $z$).\footnote{Simulation results for spillover effect defined at different values of $s$, $s'$, and $n$ display similar patterns. Therefore, we do not report them due to space limitation.}
Tables \ref{tab:MC_model1} and \ref{tab:MC_model2} present the estimation results for Model 1 and Model 2, respectively. Denote $\hat{\eta}_j$ as the estimate of $\eta^0$ in the $j$-th simulation. We report the magnitude of the bias ($bias=\frac{1}{M}\sum_{j=1}^{M}(\hat{\eta}_j-\eta^0)$), the relative bias ($|bias/\eta^0|*100\%$), the standard deviation (sd), and the root mean squared error (rmse).
Some interesting patterns emerge. First, under DGP1 (random missing), the Infeasible OLS estimation is the least biased, with the smallest standard deviation and a relative bias of less than 0.8\% in Model 1 and less than 0.3\% in Model 2. Second, the Naive OLS produces the most biased estimates. When the sample size is relatively large ($N=5000$), its relative bias ranges from 11.4\% to 33\% in Model 1 and from 6.7\% to 21\% in Model 2. Third, the bias of SPE is substantially lower than that of the Naive OLS in both models. Specifically, when the sample size is relatively large ($N=5000$), the relative bias of SPE ranges from 0.5\% to 16.4\% in Model 1 and from 2.6\% to 4.3\% in Model 2. Nonetheless, the standard deviation of the SPE method exceeds that of the Naive OLS in both models, suggesting a bias-variance trade-off between these two feasible estimation methods. Lastly, similar patterns are observed in the estimation results under heterogeneous missing (DGP2) and dependent missing (DGP3). We can see that deviations from random missing, as considered in DGP2 and DGP3, lead to a slight increase in both the estimation bias and standard deviation for most cases in the Naive OLS and SPE methods.
\begin{table}[htbp]
\begin{center}
\caption{(\textbf{Model 1}) Estimated Spillover Effects in Monte Carlo Simulations}\label{tab:MC_model1}
\subcaption{DGP1 (random missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{cc|cccc|cccc|cccc}
\hline
\hline
\noalign{\smallskip}
\multicolumn{2}{c}{}& \multicolumn{4}{c}{{\small Infeasible OLS}} &\multicolumn{4}{c}{{\small Naïve OLS}} & \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
$p_{U}$& N&bias&\%&sd&rmse&bias&\%&sd&rmse&bias&\%&sd&rmse\\
\hline
\noalign{\smallskip}
\multirow{3}[0]{*}{0.1}
& 1k & 0.000 & 0.2\% & 0.022 & 0.022 & -0.013 & 10.3\% & 0.035 & 0.037 & 0.000 & 0.1\% & 0.062 & 0.062 \\
& 2k & -0.001 & 0.4\% & 0.015 & 0.015 & -0.015 & 11.8\% & 0.024 & 0.028 & 0.001 & 1.0\% & 0.047 & 0.047 \\
& 5k & 0.000 & 0.2\% & 0.010 & 0.010 & -0.014 & 11.4\% & 0.015 & 0.021 & 0.001 & 0.5\% & 0.023 & 0.023 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.2}
& 1k & 0.000 & 0.1\% & 0.021 & 0.021 & -0.024 & 19.3\% & 0.041 & 0.048 & -0.010 & 8.4\% & 0.101 & 0.102 \\
& 2k & 0.000 & 0.1\% & 0.015 & 0.015 & -0.026 & 21.0\% & 0.030 & 0.040 & -0.010 & 8.4\% & 0.064 & 0.065 \\
& 5k & 0.001 & 0.4\% & 0.009 & 0.009 & -0.027 & 21.6\% & 0.019 & 0.033 & -0.006 & 4.8\% & 0.047 & 0.047 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.3}
& 1k & 0.000 & 0.0\% & 0.021 & 0.021 & -0.035 & 28.2\% & 0.046 & 0.058 & -0.016 & 12.6\% & 0.156 & 0.157 \\
& 2k & 0.000 & 0.1\% & 0.016 & 0.016 & -0.039 & 31.4\% & 0.036 & 0.053 & -0.025 & 19.7\% & 0.098 & 0.101 \\
& 5k & -0.001 & 0.8\% & 0.009 & 0.009 & -0.041 & 33.0\% & 0.021 & 0.046 & -0.021 & 16.4\% & 0.065 & 0.068 \\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\smallskip
\subcaption{DGP2 (heterogeneous missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{cc|cccc|cccc|cccc}
\hline
\hline
\noalign{\smallskip}
\multicolumn{2}{c}{}& \multicolumn{4}{c}{{\small Infeasible OLS}} &\multicolumn{4}{c}{{\small Naïve OLS}} & \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
$p_{U}$& N&bias&\%&sd&rmse&bias&\%&sd&rmse&bias&\%&sd&rmse\\
\hline
\noalign{\smallskip}
\multirow{3}[0]{*}{0.1} & 1k &0.000 & 0.0\% & 0.021 & 0.021 & -0.017 & 13.7\% & 0.036 & 0.040 & -0.007 & 5.5\% & 0.087 & 0.088 \\
& 2k & 0.000 & 0.1\% & 0.015 & 0.015 & -0.018 & 14.5\% & 0.027 & 0.032 & -0.005 & 3.8\% & 0.051 & 0.052 \\
& 5k & 0.000 & 0.0\% & 0.009 & 0.009 & -0.017 & 13.8\% & 0.017 & 0.024 & 0.000 & 0.4\% & 0.031 & 0.031 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.2}& 1k &0.000 & 0.3\% & 0.021 & 0.021 & -0.030 & 23.9\% & 0.045 & 0.054 & -0.011 & 9.1\% & 0.118 & 0.119 \\
& 2k &-0.001 & 0.9\% & 0.015 & 0.015 & -0.030 & 23.6\% & 0.031 & 0.043 & -0.015 & 12.0\% & 0.079 & 0.081 \\
& 5k &0.001 & 0.5\% & 0.009 & 0.009 & -0.031 & 24.5\% & 0.020 & 0.037 & -0.011 & 8.5\% & 0.049 & 0.050 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.3}& 1k &-0.001 & 0.5\% & 0.021 & 0.021 & -0.041 & 32.6\% & 0.050 & 0.064 & -0.021 & 16.5\% & 0.159 & 0.160 \\
& 2k &0.000 & 0.2\% & 0.015 & 0.015 & -0.041 & 32.7\% & 0.037 & 0.055 & -0.021 & 16.5\% & 0.117 & 0.119 \\
& 5k &0.000 & 0.4\% & 0.009 & 0.009 & -0.043 & 34.4\% & 0.024 & 0.049 & -0.019 & 15.5\% & 0.070 & 0.073 \\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\smallskip
\subcaption{DGP3 (dependent missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{cc|cccc|cccc|cccc}
\hline
\hline
\noalign{\smallskip}
\multicolumn{2}{c}{}& \multicolumn{4}{c}{{\small Infeasible OLS}} &\multicolumn{4}{c}{{\small Naïve OLS}} & \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
$p_{U}$& N&bias&\%&sd&rmse&bias&\%&sd&rmse&bias&\%&sd&rmse\\
\hline
\noalign{\smallskip}
\multirow{3}[0]{*}{0.1}& 1k & 0.000 & 0.3\% & 0.021 & 0.021 & -0.012 & 9.5\% & 0.034 & 0.036 & 0.002 & 1.9\% & 0.065 & 0.065 \\
& 2k &0.000 & 0.1\% & 0.014 & 0.014 & -0.013 & 10.7\% & 0.025 & 0.028 & 0.001 & 0.6\% & 0.039 & 0.039 \\
& 5k &0.000 & 0.2\% & 0.010 & 0.010 & -0.015 & 11.6\% & 0.016 & 0.021 & 0.000 & 0.1\% & 0.025 & 0.025 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.2}& 1k &0.001 & 0.9\% & 0.020 & 0.021 & -0.025 & 19.8\% & 0.043 & 0.050 & -0.010 & 8.0\% & 0.119 & 0.119 \\
& 2k &0.000 & 0.4\% & 0.015 & 0.015 & -0.026 & 20.8\% & 0.031 & 0.040 & -0.010 & 7.7\% & 0.071 & 0.071 \\
& 5k &-0.001 & 0.5\% & 0.009 & 0.009 & -0.029 & 23.0\% & 0.019 & 0.035 & -0.011 & 8.6\% & 0.044 & 0.046 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.3}& 1k &0.000 & 0.3\% & 0.021 & 0.021 & -0.037 & 29.6\% & 0.048 & 0.061 & -0.026 & 21.2\% & 0.143 & 0.145 \\
& 2k &0.000 & 0.3\% & 0.015 & 0.015 & -0.040 & 32.1\% & 0.033 & 0.052 & -0.023 & 18.3\% & 0.102 & 0.105 \\
& 5k &0.000 & 0.0\% & 0.010 & 0.010 & -0.040 & 32.2\% & 0.023 & 0.046 & -0.023 & 18.7\% & 0.085 & 0.088 \\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\end{center}
\footnotesize Note: Panels (a) to (c) display the estimation results under Model 1, when the missing indicator $U_{ij}$ is generated according to DGP1 to DGP3 considered in Section \ref{section_simulation}, respectively. The target spillover effect is $\eta^0=m^*(d,s,n,z)-m^*(d,0,n,z)$ at $d=0$, $s=1$, $s'=0$, $n=4,$ and no covariate $z$. True value of $\eta^0$ is 0.125 in Model 1. The column ``\%'' lists the relative bias to $\eta^0$, and the column ``bias'' lists the magnitude of the bias with respect to $\eta^0$.
\end{table}
\begin{table}[htbp]
\begin{center}
\caption{(\textbf{Model 2}) Estimated Spillover Effects in Monte Carlo Simulations}\label{tab:MC_model2}
\subcaption{DGP1 (random missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{cc|cccc|cccc|cccc}
\hline
\hline
\noalign{\smallskip}
\multicolumn{2}{c}{}& \multicolumn{4}{c}{{\small Infeasible OLS}} &\multicolumn{4}{c}{{\small Naïve OLS}} & \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
$p_{U}$& N&bias&\%&sd&rmse&bias&\%&sd&rmse&bias&\%&sd&rmse\\
\hline
\noalign{\smallskip}
\multirow{3}[0]{*}{0.1}
& 1k & 0.000 & 0.1\% & 0.028 & 0.028 & -0.028 & 6.9\% & 0.051 & 0.058 & -0.010 & 2.4\% & 0.108 & 0.109 \\
& 2k &-0.001 & 0.3\% & 0.019 & 0.019 & -0.026 & 6.6\% & 0.035 & 0.044 & -0.009 & 2.2\% & 0.078 & 0.079 \\
& 5k &0.000 & 0.1\% & 0.012 & 0.012 & -0.027 & 6.7\% & 0.023 & 0.035 & -0.011 & 2.6\% & 0.051 & 0.052 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.2}
& 1k &0.000 & 0.1\% & 0.028 & 0.028 & -0.049 & 12.2\% & 0.067 & 0.083 & 0.026 & 6.5\% & 0.199 & 0.200 \\
& 2k &0.000 & 0.1\% & 0.019 & 0.019 & -0.054 & 13.4\% & 0.049 & 0.073 & 0.020 & 5.0\% & 0.152 & 0.153 \\
& 5k &0.000 & 0.1\% & 0.012 & 0.012 & -0.056 & 14.0\% & 0.032 & 0.064 & 0.010 & 2.5\% & 0.131 & 0.132 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.3}
& 1k &0.000 & 0.0\% & 0.027 & 0.027 & -0.077 & 19.3\% & 0.085 & 0.115 & -0.012 & 2.9\% & 0.219 & 0.219 \\
& 2k &0.001 & 0.1\% & 0.019 & 0.019 & -0.077 & 19.3\% & 0.062 & 0.099 & 0.012 & 3.0\% & 0.167 & 0.167 \\
& 5k &0.000 & 0.0\% & 0.012 & 0.012 & -0.084 & 21.0\% & 0.037 & 0.092 & 0.017 & 4.3\% & 0.158 & 0.159 \\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\smallskip
\subcaption{DGP2 (heterogeneous missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{cc|cccc|cccc|cccc}
\hline
\hline
\noalign{\smallskip}
\multicolumn{2}{c}{}& \multicolumn{4}{c}{{\small Infeasible OLS}} &\multicolumn{4}{c}{{\small Naïve OLS}} & \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
$p_{U}$& N&bias&\%&sd&rmse&bias&\%&sd&rmse&bias&\%&sd&rmse\\
\hline
\noalign{\smallskip}
\multirow{3}[0]{*}{0.1}
& 1k & 0.001 & 0.2\% & 0.028 & 0.028 & -0.033 & 8.3\% & 0.060 & 0.068 & 0.006 & 1.4\% & 0.142 & 0.142 \\
& 2k &0.000 & 0.0\% & 0.019 & 0.019 & -0.034 & 8.5\% & 0.042 & 0.054 & -0.007 & 1.7\% & 0.108 & 0.108 \\
& 5k &0.000 & 0.0\% & 0.012 & 0.012 & -0.035 & 8.7\% & 0.025 & 0.043 & -0.018 & 4.5\% & 0.073 & 0.075 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.2}
& 1k &0.001 & 0.3\% & 0.027 & 0.027 & -0.062 & 15.5\% & 0.077 & 0.099 & 0.020 & 5.0\% & 0.207 & 0.208 \\
& 2k &0.000 & 0.1\% & 0.019 & 0.019 & -0.066 & 16.4\% & 0.051 & 0.083 & 0.025 & 6.3\% & 0.152 & 0.154 \\
& 5k &0.000 & 0.1\% & 0.012 & 0.012 & -0.064 & 16.0\% & 0.033 & 0.072 & 0.020 & 4.9\% & 0.143 & 0.144 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.3}
& 1k &0.000 & 0.1\% & 0.027 & 0.027 & -0.092 & 22.9\% & 0.094 & 0.131 & -0.013 & 3.2\% & 0.220 & 0.221 \\
& 2k &0.000 & 0.1\% & 0.020 & 0.020 & -0.094 & 23.6\% & 0.066 & 0.115 & -0.013 & 3.2\% & 0.181 & 0.181 \\
& 5k &0.000 & 0.1\% & 0.012 & 0.012 & -0.097 & 24.2\% & 0.041 & 0.105 & 0.004 & 1.0\% & 0.133 & 0.133 \\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\smallskip
\subcaption{DGP3 (dependent missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{cc|cccc|cccc|cccc}
\hline
\hline
\noalign{\smallskip}
\multicolumn{2}{c}{}& \multicolumn{4}{c}{{\small Infeasible OLS}} &\multicolumn{4}{c}{{\small Naïve OLS}} & \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
$p_{U}$& N&bias&\%&sd&rmse&bias&\%&sd&rmse&bias&\%&sd&rmse\\
\hline
\noalign{\smallskip}
\multirow{3}[0]{*}{0.1}&
1k & -0.001 & 0.1\% & 0.027 & 0.027 & -0.029 & 7.2\% & 0.052 & 0.059 & -0.009 & 2.2\% & 0.110 & 0.110 \\
& 2k &0.000 & 0.0\% & 0.020 & 0.020 & -0.028 & 7.0\% & 0.037 & 0.046 & -0.011 & 2.7\% & 0.072 & 0.073 \\
& 5k &0.000 & 0.0\% & 0.012 & 0.012 & -0.027 & 6.8\% & 0.023 & 0.035 & -0.011 & 2.7\% & 0.053 & 0.054 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.1}&
1k &0.002 & 0.4\% & 0.028 & 0.028 & -0.050 & 12.6\% & 0.069 & 0.086 & 0.026 & 6.6\% & 0.193 & 0.195 \\
& 2k &-0.001 & 0.1\% & 0.020 & 0.020 & -0.053 & 13.2\% & 0.052 & 0.074 & 0.021 & 5.3\% & 0.151 & 0.153 \\
& 5k &0.000 & 0.0\% & 0.012 & 0.012 & -0.056 & 13.9\% & 0.031 & 0.064 & 0.009 & 2.2\% & 0.120 & 0.121 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{0.1}&
1k &0.000 & 0.1\% & 0.027 & 0.027 & -0.082 & 20.6\% & 0.083 & 0.117 & -0.012 & 3.1\% & 0.221 & 0.221 \\
& 2k &-0.001 & 0.3\% & 0.021 & 0.021 & -0.086 & 21.4\% & 0.060 & 0.104 & 0.006 & 1.5\% & 0.165 & 0.165 \\
& 5k &0.000 & 0.1\% & 0.012 & 0.012 & -0.083 & 20.9\% & 0.038 & 0.092 & 0.016 & 4.0\% & 0.149 & 0.150 \\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\end{center}
\footnotesize Note: Panels (a) to (c) display the estimation results under Model 2, when the missing indicator $U_{ij}$ is generated according to DGP1 to DGP3 considered in Section \ref{section_simulation}, respectively. The target spillover effect is $\eta^0=m^*(d,s,n,z)-m^*(d,0,n,z)$ at $d=0$, $s=1$, $s'=0$, $n=4,$ and no covariate $z$. True value of $\eta^0$ is 0.4 in Model 2. The column ``\%'' lists the relative bias to $\eta^0$, and the column ``bias'' lists the magnitude of the bias with respect to $\eta^0$.
\end{table}
\subsection{Home Computer Use and Self-empowered Learning}\label{section_application}
In this section, we present results of a naturalistic simulation study and an empirical application using the school friendship data from \citet{beuermann2015one}.\footnote{The dataset is available at https://www.aeaweb.org/articles?id=10.1257/app.20130267.} The authors conducted a randomized controlled trial in which ``One Laptop per Child'' (OLPC) laptops were provided to primary school students in Lima, Peru, and they examined the spillover effects of home computer use on children's self-empowered learning. In their study, fourteen treatment schools were randomly selected, and students in each class drew random lotteries to win the laptops.
The baseline information, including self-reported network
data, was collected in April/May 2011, before the experiment was implemented. The lottery was drawn in June/July
2011, and 1048 laptops were distributed to lottery winners. The follow-up data
were collected in November 2011.
We use the same sample of $N=2737$ students as \citet{beuermann2015one}, which consists of students in grades 3 to 6 whose parents approved their lottery participation.
When collecting network information, students were asked to list up to 12 friends,
including their closest friends, friends with whom they did homework together, and friends who visited
their homes. Missing links exist for at least two reasons. First, the reported friends were top-coded to 12. Second, when constructing the network, reported friends with names that did not match those of any other students (for example, due to incomplete or misspelled names) were omitted because their treatment status could not be retrieved.
Since the network data were collected before the experiment, it is reasonable to
assume that the true network data and missing links are independent of the lottery results. Additionally, \citet{beuermann2015one} found that once conditioning on the network degree, baseline characteristics were well balanced between students whose friends won laptops and those whose friends did not. Since observed and unobserved characteristics are often mutually dependent, this suggests that the true network and missing links are likely to be uncorrelated with unobserved characteristics.\footnote{The fact that the observed baseline characteristics are well-balanced among students with varying numbers of laptop-winner friends given the network degree implies that $S_i$ given $\mathcal{T}_i$ is uncorrelated with both observed and unobserved characteristics. This is consistent with our intermediate result proved in Lemma \ref{lemma_binomial} in Appendix \ref{appendix_lemma1} under Assumption \ref{unconf1} (unconfounded true network) and Assumption \ref{nondiff} (nondifferential missing links).} It also indicates that the network degree is an important control variable and should be included in the regressions.
We aim to study the impacts of wining the lottery on a standardized test score for laptop digital skills (OLPC Test Score) at the follow-up stage.\footnote{The effects are studied from an intent-to-treat perspective, because 93\% of
students who won the lottery received laptops.}
We consider the model specification:
\begin{align}\label{model_laptop}
Y_{ik}=&
\theta_1D_{ik}+\theta_2\frac{S_{ik}}{\mathcal{T}_{ik}}+\theta_3D_{ik}*\frac{S_{ik}}{\mathcal{T}_{ik}}+\theta_4\mathcal{T}_{ik}+\theta_5Cov_{ik}+\mu_{k}+\varepsilon_{ik},
\end{align}
where $Y_{ik}$ is the standardized test score for student $i$ in class $k$, $D_{ik}$ is the indicator of lottery winner, $Cov_{ik}$ includes a constant, age, sex, number of siblings, number of
younger siblings, whether the father lives with the child, whether the father works at home, and
whether the mother works at home, and $\mu_{k}$ represents the class fixed effect. We use the same definitions of NBRVs as in \citet{beuermann2015one}, where $S_{ik}$ is the number of incoming friends who are lottery winners, $\mathcal{T}_{ik}$ is the indegree.
We focus on two types of spillover effects: (i) $\theta_2$ -- the spillover of laptop winners on
nonwinners (WoNW), and (ii) $\theta_2+\theta_3$ -- the spillover of laptop winners on other winners (WoW). Because the test score is standardized, the spillover effects are interpreted as the impacts on the
standard deviations of OLPC test score.
\subsubsection{Naturalistic Simulation}\label{section_naturalistic}
In the naturalistic simulation, we treat the reported network data as the true network, and the OLS coefficients for Model \eqref{model_laptop} obtained by using the reported network data as the true coefficients. We simulate
$M=1000$ experiments. In each experiment, $N$ error terms $\varepsilon_{ik}$ are randomly drawn from a normal distribution $\mathbb{N}(0,\sigma^2_\varepsilon)$, where $\sigma_\varepsilon$ is the standard deviation of the OLS residuals. Then, we generate $N$ outcome observations according to Model \eqref{model_laptop}, using the original treatment variable, NBRVs, and covariates for the $N$ students, along with the simulated error terms. Missing links are artificially introduced to the observed network $A_{ij}=U_{ij}A^*_{ij}$, where $U_{ij}$ is a binary indicator that is generated following the three DPGs considered in Section \ref{section_simulation}, with $p_U\in\{0.1,0.2,0.3\}$.
\begin{table}[htbp]
\begin{center}
\caption{Estimated Spillover Effects using Naturalistic Simulation}\label{tab:NS_model2}
\subcaption{DGP1 (random missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{ccccccccccccccc}
\hline
\hline
\noalign{\smallskip}
&& \multicolumn{1}{c}{{\small True}} &\multicolumn{4}{c}{{\small Naive OLS}} && \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
\cline{4-7}\cline{9-12}\noalign{\smallskip}
&$p_U$&&bias&\%&sd&rmse&&bias&\%&sd&rmse\\
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{WoNW}&0.1&\multirow{3}[0]{*}{0.140}&-0.020& 14\%&0.089&0.091&&0.013&9.3\%&0.122&0.123 \\
&0.2&&-0.034&24\%&0.083&0.090 &&-0.001&0.9\%&0.120&0.120 \\
&0.3&&-0.046&33\%&0.078&0.091 &&-0.012&8.2\%&0.114&0.114 \\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{WoW}&0.1&\multirow{3}[0]{*}{0.307}&-0.041&13\%&0.156&0.161 &&0.012&4.0\%&0.224&0.225 \\
&0.2&&-0.075&24\%&0.152&0.170 &&-0.024&7.9\%&0.220&0.221 \\
&0.3&&-0.107&35\%&0.147&0.182 &&-0.042&14\%&0.215&0.220 \\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\smallskip
\subcaption{DGP2 (heterogeneous missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{ccccccccccccccc}
\hline
\hline
\noalign{\smallskip}
&& \multicolumn{1}{c}{{\small True}} &\multicolumn{4}{c}{{\small Naive OLS}} && \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
\cline{4-7}\cline{9-12}\noalign{\smallskip}
&$p_U$&&bias&\%&sd&rmse&&bias&\%&sd&rmse\\
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{WoNW}&0.1&\multirow{3}[0]{*}{0.140}&-0.022&16\% & 0.084&0.087&&0.007&5.3\%&0.116&0.117\\
&0.2&& -0.036&26\%&0.081&0.089&&-0.001&0.5\%&0.117&0.117\\
&0.3&& -0.052&37\%&0.079&0.095&&-0.018&12\%&0.111&0.112\\
\noalign{\smallskip}
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{WoW}&0.1&\multirow{3}[0]{*}{0.307}&-0.054& 18\%& 0.159&0.168&&0.002&0.7\%&0.222&0.222\\
&0.2&& -0.086&28\%&0.153&0.176&&-0.019&6.2\%&0.219&0.220\\
&0.3&& -0.124&40\%&0.143&0.189&&-0.059&19\%&0.210&0.218\\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\smallskip
\subcaption{DGP3 (dependent missing)}
\begin{adjustbox}{max width=\textwidth}
\begin{tabular}{cccccccccccccc}
\hline
\hline
\noalign{\smallskip}
&& \multicolumn{1}{c}{{\small True}} &\multicolumn{4}{c}{{\small Naive OLS}} && \multicolumn{4}{c}{{\small SPE}} \\
\noalign{\smallskip}
\cline{4-7}\cline{9-12}\noalign{\smallskip}
&$p_U$&&bias&\%&sd&rmse&&bias&\%&sd&rmse\\
\hline\noalign{\smallskip}
\multirow{3}[0]{*}{WoNW}&0.1&\multirow{3}[0]{*}{0.140}&-0.014&10\%&0.087&0.088&&0.011&7.7\%&0.121&0.122\\
&0.2&& -0.037&26\%&0.079&0.087&&0.007&5.0\%&0.122&0.122\\
&0.3&&-0.051&36\%&0.077&0.092&&-0.013&9.5\%&0.116&0.117 \\
\hline
\multirow{3}[0]{*}{WoW}&0.1&\multirow{3}[0]{*}{0.307}&-0.034&11\%&0.161&0.165&&0.008&2.5\%&0.222&0.222\\
&0.2&&-0.072&23\%& 0.155&0.171&&-0.010&3.4\%&0.222&0.222 \\
&0.3&& -0.099&32\%&0.151&0.181&&-0.038&12\%&0.218&0.221\\
\noalign{\smallskip}
\hline
\hline
\end{tabular}
\end{adjustbox}
\end{center}
\footnotesize Note: Panels (a) to (c) display the estimation results of the naturalistic simulation under Model \eqref{model_laptop}, when the missing indicator $U_{ij}$ is generated according to DGP1 to DGP3 introduced in Section \ref{section_simulation}, respectively. We consider the spillover effects of laptop winners on nonwinners (WoNW, $\theta_2$) and laptop winners on other winners (WoW, $\theta_2+\theta_3$). The true values of WoW and WoNW are given in the column ``True''. The value of the bias (``bias''), the relative bias to the true value (``\%''), the standard deviation (``sd''), and root mean squared error (``rmse'') of the estimated spillover effects are presented.
\end{table}
We apply the two feasible estimation methods: the Naive OLS that ignores the missing links and the SPE that uses both incoming and outgoing links. Similar to the Monte Carlo simulation, for SPE method, we choose the value of $K$ so that the smallest singular value of $\mathbf{F}_{\mathcal{T},\tilde{\mathcal{T}}}$ is larger than 0.001, and we order the estimated eigenvectors of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*}$ according to Assumption \ref{ass_eigen} (c). Table \ref{tab:NS_model2} displays the values of the bias, relative bias (\%), standard deviation (sd), and root mean squared error (rmse) of the estimated spillover effects of interest. We can see that, on average, Naive OLS underestimates the true spillover effects of WoNW and of WoW by 13\% to 35\% under DGP1 (random missing), 16\% to 40\% under DGP2 (heterogeneous missing), and 10\% to 36\% under DGP3 (dependent missing). The SPE estimates are less biased than those of Naive OLS, with relative bias ranging from 0.9\% to 14\% under DGP1, 0.5\% to 19\% under DGP2, and 2.5\% to 12\% under DGP3. As the missing probability $p_U$ increases, both estimation methods tend to underestimate the true effect in a systematic manner, and the degree of underestimation also increases.
\subsubsection{Empirical Application}\label{section_empirical_homelaptop}
In this section, we analyze the consequences of missing network links in the study of the home computer use on self-empowered learning. We treat the reported network data as the observed network that contains missing links. We compare the estimation results of Naive OLS and SPE.
For SPE, we choose different values of $K$ ($K\in\{9,10,11\}$) to define the truncated degree support, and we order the estimated eigenvectors of $\mathbf{F}_{\mathcal{T}|\mathcal{T}^*}$ according to Assumption \ref{ass_eigen} (c). We present the estimation results of Model \eqref{model_laptop} in Table \ref{tab:OLPC_model2}, using the sample of $N=$2737 students. The
standard errors of both Naive OLS and SPE methods are clustered at the school level.\footnote{The standard errors of SPE method are calculated using the numerical method proposed in \citet{newey1994kernel}. We follow \citet{hong2015extremum} to choose the step size as $log(N)log(log(N))/N$ in the numerical differentiation.}
\begin{table}[htbp]
\centering
\caption{Estimation Results for OLPC Test Score}\label{tab:OLPC_model2}
\begin{adjustbox}{max width=\textwidth}
\begin{threeparttable}
\begin{tabular}{lccccc}
\hline\hline
&Naive OLS& \multicolumn{3}{c}{SPE} \\
\cmidrule{3-5}
&&$K=9$&$K=10$&$K=11$\\
&(1)&(2)&(3)&(4)\\[0.05cm]
\midrule
&\multicolumn{4}{c}{Panel (a): Parameters} \\
\cmidrule{2-5}
$D_{ik}$ &0.786&0.791&0.789&0.771 \\
&(0.069)***& (0.145)***&(0.154)***&(0.226)***\\[0.1cm]
$\frac{S_{ik}}{\mathcal{T}_{ik}}$ &0.140&0.244&0.153&0.280\\
& (0.098)&(0.103)** &(0.123)&(0.258)\\[0.1cm]
$D_{ik}*\frac{S_{ik}}{\mathcal{T}_{ik}}$ &0.167&0.221&0.183&0.246\\
&(0.228)&(0.589) &(0.559)&(1.025)\\[0.15cm]
$\mathcal{T}_{ik}$ &0.051 &0.058&0.052&0.063\\
& (0.009)***&(0.011)***&(0.011)***&(0.020)***\\[0.05cm]
\midrule
&\multicolumn{4}{c}{Panel (b): Spillovers}\\
\cmidrule{2-5}
WoNW &0.140 &0.244&0.153&0.280\\
&(0.098)&(0.103)** &(0.123)&(0.258)\\
&[-0.052, 0.333]&[0.042, 0.446]&[-0.087, 0.394]&[-0.225, 0.785]\\[0.1cm]
WoW &0.307 &0.465&0.336&0.526\\
&(0.203)&(0.633)&(0.562)&(1.153)\\
&[-0.093, 0.708]&[-0.776, 1.706]&[-0.766, 1.437]&[-1.734, 2.786]\\[0.2cm]
Sample size &2737&2737&2737&2737 \\[0.05cm]
\hline\hline
\end{tabular}
\begin{tablenotes}[para,flushleft]
\smallskip
\footnotesize
\item[]Note: This table presents the estimation results of Model \eqref{model_laptop} under different values of $K$. We assume that missing indicators are independent of covariates for computational simplicity. Standard errors (s.e.) of both Naive OLS and SPE methods are clustered at the school level and reported in the parentheses. The numerical method of \citet{newey1994kernel} is applied to calculate the s.e. of SPE method, and the choice of step size in the numerical differentiation is based on \citet{hong2015extremum}. 95\% confidence intervals for spillover effects of laptop winners on nonwinners (WoNW, $\theta_2$) and laptop winners on other winners (WoW, $\theta_2+\theta_3$) are given in the brackets.
\end{tablenotes}
\end{threeparttable}
\end{adjustbox}
\end{table}
We find that the estimated spillover effects of WoNW and WoW using Naive OLS are positive but insignificant. Using Naive OLS, we obtain a WoNW spillover of 0.140 standard deviations with a 95\% confidence interval (CI) of $[-0.052, 0.333]$, and a
WoW spillover of 0.307 standard deviations with a 95\% CI of $[-0.093, 0.708]$. Using SPE method, the point estimate of WoNW spillover effect ranges from 0.153 to 0.280 standard deviations, and the point estimate of WoW spillover effect ranges from 0.336 to 0.526 standard deviations, under different values of $K$. In addition, when $K=9$, the SPE estimate of the WoNW spillover effect is significant. The spillover effects for WoNW and WoW obtained by SPE are larger than those obtained by Naive OLS, indicating a possible underestimation of Naive OLS of the true spillover effects.
\section{Conclusion}\label{section_conclusion}
This paper investigates spillovers of program benefits in the presence of missing network links. We propose to use two network measures, that can be constructed using the incoming and outgoing links, to point identify the treatment and spillover effects in the case of bounded degree. If the degree is unbounded, our method can be used as a bias-reduction approach. We provide a two-step semiparametric estimation method and study its asymptotic properties. Monte Carlo experiments and a naturalistic simulation confirm the effectiveness of our approach in reducing estimation bias compared to the naive estimation that neglects missing links.
The literature on network effects often emphasizes the potential impacts of higher-order network connections with indirect friends (for example, friends of friends). However, incorporating these higher-order connections in the outcome model will introduce higher order missing links (for example, missing friends of friends), which complicates the dependence among observable and latent network-based random variables. Therefore, extending our method to address the missing link problem in models with higher-order network connections is nontrivial and left to future research.
Furthermore, although our method focuses on missing network links, the network misclassification can be two-sided. Nonetheless, our method can still be used as bias-reduction approach when false positive links exist with a small or declining probability.
{
\setstretch{1}
\bibliographystyle{plainnat}
\bibliography{reference_SP}
}
\newpage
\begin{center}
\LARGE{Online Appendix for \\
{\Large Spillovers of Program Benefits with Missing Network Links}}
\vspace{1cm}
\large Lina Zhang
\end{center}