Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
33,653 characters · 13 sections · 21 citation commands
1Identification and Estimation of a Partially Linear Regression Model using Network Data
Most economic outcomes are not determined in isolation. Rather agents are influenced by the behaviors and characteristics of other agents. For example, a high school student's academic performance might depend on the attitudes and expectations of that student's friends and family sacerdote2011peer,bramoulle2019peer.
Incorporating this social influence into the right-hand side of an economic model may be desirable when the researcher wants to understand its impact on the outcome of interest or when it confounds the impact of another explanatory variable such as the causal effect of some nonrandomized treatment. For instance, the researcher may want to learn the causal effect of a tutoring program on academic performance in which program participation and counterfactual academic performance are both partially determined by family expectations. However, in many cases the relevant social influence is not observed by the researcher. That is, the researcher does not have access to data on the family expectations that confound the causal effect of the tutoring program and thus cannot control for this variable using conventional methods.
An increasingly popular solution to this problem is to collect network data and suppose that the unobserved social influence is revealed by linking behavior in the network. For instance, the researcher might observe pairs of students who identify as friends and believe that students with similar reported friendships have similar family expectations. It is not immediately clear, however, how one might actually use network data to account for this unobserved social influence in practice, since the number of ways in which agents can be linked in a network is typically large relative to the sample size.
This paper proposes a new way to incorporate network data into an econometric model. I specify a joint regression and network formation model, establish sufficient conditions for the parameters of the regression model to be identified, and provide consistent estimators. Large sample approximations for inference and an application to network peer effects building on work by bramoulle2009identification, de2010identification, goldsmith2013social, hsieh2014social, johnsson2015estimation,arduini2015parametric, and others is provided by auerbach2021identification.
A limitation of the framework is that the large sample approximations suppose a sequence of networks that is asymptotically dense in that the fraction of linked agent-pairs does not vanish with the sample size. The regime can fail to characterize potentially relevant features of networks in which relatively few agent pairs are linked mele2017structuralb. Potential extensions to sparse asymptotic regimes are left to future work.
Let $i$ represent an arbitrary agent from a large population. Associated with agent $i$ are an outcome $y_{i} \in \mathbb{R}$, an observed vector of explanatory variables $x_{i} \in \mathbb{R}^{k}$, and an unobserved index of social characteristics $w_{i} \in [0,1]$. The three are related by the model
where $\beta \in \mathbb{R}^{k}$ is an unknown slope parameter, $\lambda$ is an unknown measurable function, and $\varepsilon_{i}$ is an idiosyncratic error.
The researcher draws a sample of $n$ agents uniformly at random from the population. This sample is described by the sequence of independent and identically distributed random variables $\{y_{i},x_{i}, w_{i}\}_{i=1}^{n}$, although only $\{y_{i},x_{i}\}_{i=1}^{n}$ is observed as data. The researcher also observes $D$, an $n\times n$ stochastic binary adjacency matrix corresponding to an unlabeled, unweighted, and undirected random network between the $n$ agents. The existence of a link between agents $i$ and $j$ is determined by the model
in which $f$ is a symmetric measurable function satisfying the continuity condition that $\inf_{u \in [0,1]}\int \mathbbm{1}\left\{v \in [0,1]: \sup_{\tau \in [0,1]}\left|f(u,\tau)-f(v,\tau)\right| < \varepsilon\right\}dv > 0$ for every $\varepsilon > 0$ and $\{\eta_{ij}\}_{i,j=1}^{n}$ is a symmetric matrix of unobserved scalar disturbances with independent upper diagonal entries that are mutually independent of $\{x_{i},w_{i},\varepsilon_{i}\}_{i=1}^{n}$. This continuity condition is weaker than the typical assumption that $f$ is a continuous function. It is used because it allows for a variety of models where $f$ is “almost” but not quite a continuous function. For example, in the blockmodel described in Section 2.2.1 below, $f$ is a piecewise continuous function.
The regression model ((ref)) represents a pared-down version of various linear models popular in the network economics literature. For instance in the network peer effects literature, $y_{i}$ could be student $i$'s GPA, $x_{i}$ could indicate whether $i$ participates in a tutoring program, $w_{i}$ could index student $i$'s participation in various social cliques, and $\lambda(w_{i})$ could represent the influence of student $i$'s peers' expected GPA, program participation, or other characteristics on student $i$'s GPA. That is, supposing $P(D_{ij} = 1 |w_{i}) > 0$,
for some $(\gamma,\delta) \in \mathbb{R}^{k+1}$ manski1993identification. To demonstrate the proposed methodology, this paper conflates these different possible social effects into one social influence term, $\lambda(w_{i})$. This may be sufficient to identify the impact of the tutoring program holding social influence constant, predict a student's GPA under some counterfactual social influence, or test for the existence of any social influence. auerbach2021identification discusses how one can also separately identify different social effects.
The parameters of interest are $\beta$ and $\lambda(w_{i})$, the realized social influence for agent $i$. The function $\lambda: [0,1] \to \mathbb{R}$ is not a parameter of interest because it is not separately identified from $w_{i}$ (see Section 2.2.1 below). It is without loss to normalize the distribution of $w_{i}$ to be standard uniform.
Network formation ((ref)) is represented by $n \choose 2$ conditionally independent Bernoulli trials. The model is a nonparametric version of a class of dyadic regression models popular in the network formation literature. Section 6 of graham2019network or Section 3 of de2020econometric contains many examples. It is often given a discrete choice interpretation in which $f(w_{i},w_{j}) - \eta_{ij}$ represents the marginal transferable utility agents $i$ and $j$ receive from forming a link, which precludes strategic interactions between agents. The distribution of $\eta_{ij}$ is not separately identified from $f$ and so is also normalized to be standard uniform.
Under ((ref)), the observed network $D$ is almost surely dense or empty in the limit. That is, for a fixed $f$ and as $n$ tends to infinity, ${n \choose 2}^{-1}\sum_{i=1}^{n-1}\sum_{j=i+1}^{n}D_{ij}$ will either be bounded away from zero or exactly zero with probability approaching one. The framework can potentially accommodate network sparsity by allowing $f$ to vary with the sample size (see Appendix Section A.1), but a formal study of such an asymptotic regime is left to future work.
The following Assumption 1 collects key aspects of the model for reference. \newline
If $w_{i}$ were observed, ((ref)) would correspond to the partially linear regression of robinson1988root and the identification problem would be well-understood. If $w_{i}$ were unobserved but identified, one might replace $w_{i}$ with an empirical analogue as in ahn1993semiparametric. Identification strategies along these lines are considered by arduini2015parametric and johnsson2015estimation.
However, it is not generally possible to learn $w_{i}$ in the setting of this paper. The main difficulty is that many assignments of agents to social characteristics generate the same distribution of network links. Specifically, for any measure-preserving invertible $\varphi$ (that is for any measurable $A \subseteq [0,1]$, $A$ and $\varphi^{-1}(A)$ have the same measure), $\left(\{w_{i}\}_{i=1}^{n},f(\cdot,\cdot)\right)$ and $\left(\{\varphi(w_{i})\}_{i=1}^{n},f(\varphi^{-1}(\cdot),\varphi^{-1}(\cdot))\right)$ generate the same distribution of links, where $w_{i}$ and $\varphi(w_{i})$ may be very different. For example, if $\{w_{i}\}_{i=1}^{n}$ and $f(u,v) = (u+v)/2$ explain the distribution of $D$, then so too does $\{w_{i}'\}_{i=1}^{n}$ and $f'(u,v) = 1-(u+v)/2$ where $w_{i}' = 1-w_{i}$.
Furthermore, even if the researcher is willing to posit a specific $f$, the social characteristics may still not be identified. For example, in a simplified version of the blockmodel of holland1983stochastic, there exists an $l\times l$ dimensional matrix $\Theta$ such that $f(w_{i},w_{j}) = \Theta_{\lceil l w_{i}\rceil \lceil l w_{j}\rceil}.$ Intuitively, $[0,1]$ is divided into $l$ partitions (with agent $i$ assigned to partition $\lceil lw_{i}\rceil$) and the probability two agents link only depends on their partition assignments. In this case, the probability that agents link is invariant to changes in the social characteristics that do not change the agents' partition assignments, and so while the underlying partition assignments might be learned from $D$, the social characteristics that determine the partition assignments generally cannot. Notice that $f$ in this case is not a continuous function, but satisfies the continuity condition of Assumption 1.
Another example in which the social characteristics are not identified is the homophily model $f(w_{i},w_{j}) = 1-(w_{i}-w_{j})^{2}$. Intuitively, agents are more likely to form a link if their social characteristics are similar. In this case, both $\{w_{i}\}_{i=1}^{n}$ and $\{1-w_{i}\}_{i=1}^{n}$ generate the same distribution of links.
An example in which the social characteristics are identified is the nonlinear additive model $f(w_{i},w_{j}) = \Lambda(w_{i}+w_{j}),$ where $\Lambda$ is a strictly monotonic function such as the logistic function graham2019network. Intuitively, agents with larger values of $w_{i}$ are more likely to form links. In this case, $w_{i}$ is identified from $D$ because $w_{i} = P\left(\int \Lambda(w+\tau)d\tau \leq \int \Lambda(w_{i}+\tau)d\tau| w_{i}\right)$ where $\int \Lambda(w_{i}+\tau)d\tau = P\left(D_{it} = 1| w_{i}\right)$ and $w$ is an independent copy of $w_{i}$.
Since $w_{i}$ is not generally identified, I propose an alternative description about how $i$ is linked in the network that is identified. I call this alternative an agent link function and propose using link functions instead of social characteristics to identify $\beta$ and $\lambda(w_{i})$.
Agent $i$'s link function is the projection of $f$ onto $w_{i}$. That is, $f_{w_{i}}(\cdot) := f(w_{i},\cdot): [0,1] \to [0,1]$. It is the collection of probabilities that agent $i$ links to agents with each social characteristic in $[0,1]$. I consider link functions to be elements of $L^{2}([0,1])$, the usual inner product space of square integrable functions on the interval. I sometimes use $d(w_{i},w_{j}) := ||f_{w_i} - f_{w_j}||_{2}$ to refer to the pseudometric on $[0,1]$ induced by $L^2$-differences in link functions. I call this pseudometric network distance.
Conditional expectations with respect to $f_{w_{i}}$ implicitly refer to the random variable $w_{i}$. For example, $E\left[x_{i} |\hspace{1mm} f_{w_{i}}\right] := \lim_{h \to 0}E\left[x | \hspace{1mm} w \in \{u \in [0,1] : ||f_{u}-f_{w_{i}}||_{2} \leq h\},w_{i}\right]$ and \\$E\left[x_{i}x_{j}' | ||f_{w_{i}} - f_{w_{j}}||_{2} = 0\right]$ $:= \lim_{h\to0}E\left[x\tilde{x}' | (w,\tilde{w}) \in \{(u,v) \in [0,1]^{2} : ||f_{u}-f_{v}||_{2} \leq h\}\right]$ where $(x,w)$ and $(\tilde{x},\tilde{w})$ are independent copies of $(x_{i},w_{i})$. The conditional expectations on the right-hand side are well-defined for any $h > 0$ because of the continuity condition on $f$ in Assumption 1. Whenever the conditional expectation on the left-hand side is used, the relevant limit is assumed to exist.
Under ((ref)), the link function $f_{w_{i}}$ is the totality of information that $D$ contains about $w_{i}$. It describes the law of the $i$th row of $D$ and so is identified. Furthermore, $w_{i}$ is only identified when $f_{w_{i}}$ is invertible in $w_{i}$. For example, in the nonlinear additive model from Section 2.2.1, $w_{i}$ is identified because $u > v$ implies that $f_{u}(\cdot) := \Lambda(u + \cdot)$ dominates $f_{v}(\cdot) := \Lambda(v + \cdot)$. In this example, agents with different social characteristics necessarily have different probabilities of forming links to other agents in the population. In the blockmodel, $w_{i}$ is not identified because if $\lceil l u\rceil = \lceil l v\rceil$ then $f_{u}(\cdot) := \Theta_{\lceil l u\rceil \lceil l \cdot\rceil} = f_{v}(\cdot) := \Theta_{\lceil l v\rceil \lceil l \cdot\rceil}$ even if $u \neq v$. In this example, agents with different social characteristics but the same partition assignment have the same probability of forming links to other agents in the population.
The large-sample limits of many popular agent-level network statistics are determined by the agent's link function. Examples include degree $\frac{1}{n}\sum_{t=1}^{n}D_{it} \to_{p} E\left[D_{it}|w_{i}\right] = \int f_{w_{i}}(\tau)d\tau$, average peers' characteristics $\frac{\sum_{t=1}^{n}x_{t}D_{it}}{\sum_{t=1}^{n}D_{it}} \to_{p} E\left[x_{t}|D_{it} = 1,w_{i}\right] = \frac{\int E\left[x_{t}|w_{t} = \tau\right]f_{w_{i}}(\tau)d\tau}{\int f_{w_{i}}(\tau)d\tau}$, and clustering $\frac{\sum_{j=1}^{n-1}\sum_{k = j+1}^{n}D_{ij}D_{ik}D_{jk}}{\sum_{j=1}^{n-1}\sum_{k = j+1}^{n}D_{ij}D_{ik}} \to_{p} \frac{E\left[D_{ij}D_{ik}D_{jk}|w_{i}\right]}{E\left[D_{ij}D_{ik}|w_{i}\right]} = \frac{\int\int f_{w_{i}}(\tau)f_{w_{i}}(s)f(\tau,s)d\tau ds}{\left(\int f_{w_{i}}(\tau)d\tau\right)^{2}}$ (supposing $\int f_{w_{i}}(\tau)d\tau > 0$). This observation will partly inform Assumption 3 below.
If $w_{i}$ were observed or identified, the standard approach would be to first identify $\beta$ using covariation between $y_{i}$ and $x_{i}$ unrelated to $w_{i}$ and then to identify $\lambda(w_{i})$ using residual variation in $y_{i}$. This identification strategy requires variation in $x_{i}$ not explained by $w_{i}$. Let $\Xi(u) = E\left[(x_{i} - E\left[x_{i}|w_{i} \right])'(x_{i} - E\left[x_{i}|w_{i} \right])|w_{i} = u\right]$.
Assumption 2 is strong but standard. It is violated when the covariates include population analogues of network statistics such as agent degree or average peers' characteristics (or any other function of $w_{i}$). In such cases, alternative assumptions are required for identification.
Since $w_{i}$ is neither observed nor identified, the standard approach cannot be implemented. A contribution of this paper is to propose using $f_{w_{i}}$ instead of $w_{i}$ for identification. The substitution relies on an additional assumption that $\lambda(w_{i})$ is determined by $f_{w_{i}}$.
Assumption 3 is strong and new. In words, it says that agents with similar link functions have similar social influence. Since, under ((ref)), $f_{w_{i}}$ is the totality of information that $D$ contains about $w_{i}$, Assumption 3 supposes that this information is sufficient to discern $\lambda(w_{i})$. It does not restrict the function $f$.
One justification for the assumption could be that $w_{i}$ does not directly impact $y_{i}$. Instead, $w_{i}$ only influences $y_{i}$ by altering $i$'s linking behavior $f_{w_{i}}$. For example, if $w_{i}$ indexes student $i$'s participation in various social cliques, then the assumption follows if this index only directly affects which other students and teachers $i$ interacts with, and it is these interactions that ultimately determine $i$'s participation in the tutoring program and GPA.
The assumption is also satisfied when the social influence is the population analogue of one of the network statistics described in Section 2.2.2. This is the case for the network peer effects example of Section 2.1 where $\lambda(w_{i}) = E\left[x_{j}|D_{ij}=1,w_{i}\right]\gamma + \delta E\left[y_{j}|D_{ij}=1,w_{i}\right]$, because $E\left[z_{j}|D_{ij}=1,w_{i}\right] = \frac{\int E\left[z_{j} | w_{j} = \tau\right]f_{w_{i}}(\tau)d\tau}{\int f_{w_{i}}(\tau)d\tau}$ is a continuous functional of $f_{w_{i}}$.
However, the assumption may be implausible when the network is sparse (the link function is close to $0$) because every agent-pair may have network distance close to zero. As a result, under network sparsity, Assumption 3 may imply that $\lambda(w_{i})$ behaves like a constant. See Appendix Section A.1 for a discussion.
Proposition 1 states that Assumptions 1-3 are sufficient for $\beta$ and $\lambda(w_{i})$ to be identified.
I close with two examples in which $\beta$ and $\lambda(w_{i})$ are identified (Assumptions 1-3 hold) but $w_{i}$ is not. The first example is the case where links are determined by a blockmodel $f(w_{i},w_{j}) = \Theta_{\lceil lw_{i}\rceil\lceil lw_{j}\rceil}$ and social influence is determined by the agent partition assignments $y_{i} = x_{i}\beta + \alpha_{\lceil lw_{i}\rceil} + \varepsilon_{i}$. In this case, $\beta$ and $\alpha_{\lceil lw_{i}\rceil}$ are identified even though $w_{i}$ is not. The second example is the case where links are determined by a homophily model $f(w_{i},w_{j}) = 1-(w_{i}-w_{j})^{2}$ and social influence is an affine function of the agent social characteristics $y_{i} = x_{i}\beta + \rho_{1} + \rho_{2} w_{i} + \varepsilon_{i}$. In this case $\beta$ and $\rho_{1} + \rho_{2} w_{i}$ are identified even though $(\rho_{1},\rho_{2})$ and $w_{i}$ are not separately identified.
Estimation of $\beta$ and $\lambda(w_{i})$ is complicated by the fact that $f_{w_{i}}$ is unobserved and difficult to approximate directly. A contribution of this paper is to demonstrate that estimation is still possible using columns of the squared adjacency matrix. To explain the procedure, I introduce the codegree function.
Let $p$ map $(w_i,w_j)$ to the conditional probability that $i$ and $j$ have a link in common, i.e. $p(w_{i},w_{j}) := \int f_{w_{i}}(\tau)f_{w_{j}}(\tau)d\tau$. Agent $i$'s codegree function is the projection of $p$ onto $w_{i}$. That is, $p_{w_{i}}(\cdot) := p(w_{i},\cdot): [0,1] \to [0,1]$. Codegree functions are also taken to be elements of $L^{2}([0,1])$. I sometimes use $\delta$ to refer to the pseudometric on $[0,1]$ induced by $L^{2}$-differences in codegree functions,
I call this pseudometric codegree distance. Conditional expectations with respect to codegree functions are defined exactly as they are for link functions.
In contrast to link functions, the population analogues of most network statistics (including those in Section 2.2.2) cannot naturally be written as functionals of codegree functions. The use of codegree functions is instead motivated by Lemma 1 below.
I propose using codegree functions instead of link functions to construct estimators for $\beta$ and $\lambda(w_{i})$. The proposal relies on two results. The first result is that agents with similar codegree functions have similar link functions. The second result is that codegree distance can be consistently estimated using the columns of the squared adjacency matrix.
The first result is given by Lemma 1 and is related to arguments from the link prediction literature lovasz2010regularity,rohe2011spectral,zhang2015estimating.
I defer a discussion of Lemma 1 to Section 2.3.3 and emphasize here instead its implication that the parameters of interest can be expressed as functionals of the agent codegree functions. That is, under Assumptions 1-3, $\beta$ uniquely minimizes \\$E\left[\left(y_{i}-y_{j} - (x_{i}-x_{j})b\right)^{2}|||p_{w_i}-p_{w_j}||_{2} = 0\right]$ over $b \in \mathbb{R}^{k}$ and $\lambda(w_{i}) = E\left[\left(y_{i} - x_{i}\beta\right) |p_{w_{i}}\right]$.
The second result is that $\delta(w_{i},w_{j})$ can be consistently estimated by the root average squared difference in the $i$th and $j$th columns of the squared adjacency matrix,
Intuitively, the empirical codegree $\frac{1}{n}\sum_{s=1}^{n}D_{ts}D_{is}$ counts the fraction of agents that are linked to both agents $i$ and $t$, $\{\frac{1}{n}\sum_{s=1}^{n}D_{ts}D_{is}\}_{t=1}^{n}$ is the collection of empirical codegrees between agent $i$ and the other agents in the sample, and $\hat{\delta}_{ij}$ gives the root average squared difference in $i$'s and $j$'s collection of empirical codegrees. That $\hat{\delta}_{ij}$ converges uniformly to $||p_{w_{i}} - p_{w_{j}}||_{2}$ over the $n\choose 2$ distinct pairs of agents as $n\to \infty$ is shown in Appendix Section A.4 as Lemma B1.
A consequence of these two results and Assumptions 1-3 is that when the $i$th and $j$th columns of the squared adjacency matrix are similar then $(y_{i}-y_{j})$ and $(x_{i}-x_{j})\beta + (\varepsilon_{i}-\varepsilon_{j})$ are approximately equal. Under additional regularity conditions provided in Section 2.3.4, $\beta$ is consistently estimated by the pairwise difference estimator
and $\lambda(w_{i})$ is consistently estimated by the Nadaraya-Watson-type estimator
where $K$ is a kernel function and $h_{n}$ is a bandwidth parameter depending on the sample size. Since codegree functions are not finite-dimensional, the regularity conditions I provide for consistency are different than what is typically assumed. Conditions sufficient for the estimators to be asymptotically normal, consistent estimators for their variances, and more are provided by auerbach2021identification.
The proof of Lemma 1 can be found in Appendix Section A.2. The first part, that $||p_{w_{i}} - p_{w_{j}}||_{2} \leq ||f_{w_{i}} - f_{w_{j}}||_{2}$, is almost an immediate consequence of Jensen's inequality. The second part is related to Theorem 13.27 of lovasz2012large, the logic of which demonstrates that $||p_{w_{i}} - p_{w_{j}}||_{2} = 0$ implies $||f_{w_{i}} - f_{w_{j}}||_{2} = 0$ when $f$ is a continuous function. The proof is short.
Intuitively, if agents $i$ and $j$ have identical codegree functions then the difference in their link functions $(f_{w_{i}} - f_{w_{j}})$ must be uncorrelated with every other link function in the population, as indexed by $\tau$. In particular, the difference is uncorrelated with $f_{w_{i}}$ and $f_{w_{j}}$, the link functions of agents $i$ and $j$. However, this is only the case when $f_{w_{i}}$ and $f_{w_{j}}$ are perfectly correlated.
Lov\'asz's theorem demonstrates that agent-pairs with identical codegree functions have identical link functions. The estimation strategy proposed in this paper, however, requires a stronger result that agent-pairs with similar but not necessarily identical codegree functions have similar link functions. This is the statement of Lemma 1.
auerbach2021identification derives rates of convergence for the estimators under a stronger version of Lemma 1. I include the result here as it may be of independent interest.
Lemma A1 bounds the cost of using codegree distance as a substitute for network distance in the estimation of $\beta$ and $\lambda(w_{i})$. Its proof can be found in Appendix Section A.2. When $C = \alpha = 1$, the result requires an agent-pair to have a codegree distance less than $\varepsilon^{3}/8$ to guarantee that their network distance is less than $\varepsilon$. The rate of convergence of the estimators based on codegree distance may be slower than the infeasible estimators based on network distance.
The following regularity conditions are imposed. Let $r_{n}(u) := \int K\left(\frac{||p_{u}-p_{v}||^{2}_{2}}{h_{n}}\right)dv$.
The restrictions on $K$ are standard. The first two restrictions on $h_{n}$ are also standard. The third restriction on $h_{n}$, that $\inf_{u \in [0,1]} n^{\gamma/4}r_{n}(u) \to \infty$, is new. It ensures that the sums used to estimate $\hat{\beta}$ and $\widehat{\lambda(w_{i})}$ diverge with $n$. If $p_{w_i}$ was a continuously distributed $d$-dimensional random vector then, under certain conditions, $P(||p_{w_i}-p_{w_j}||^{2}_{2} \leq h_{n}| \hspace{1mm}w_{i})$ would be on the order of $h_{n}^{d/2}$. The number of agents with codegree function similar to that of agent $i$ would be on the order of $nh_{n}^{d/2}$, and $h_{n}$ could be chosen so that $nh_{n}^{d/2} \to \infty$. Such an assumption, which requires knowledge of the dimension of $p_{w_{i}}$, is standard. Since $p_{w_i}$ is an unknown function, $P(||p_{w_i}-p_{w_j}||_{2}^{2} \leq h_{n}|\hspace{1mm}w_{i})$ can not necessarily be approximated by a polynomial of $h_{n}$ of known order and so $\inf_{u \in [0,1]}n^{\gamma/4}r_{n}(u) \to \infty$ is explicitly assumed instead. One can verify it in practice (in the same sense that one can choose $h_{n}$ to satisfy the first two conditions) by computing $\min_{i=1,...,n}\frac{1}{n}\sum_{j=1}^{n}K\left(\frac{\hat{\delta}^{2}_{ij}}{h_{n}}\right)$ and choosing $h_{n}$ so that it is large relative to $n^{-\gamma/4}$.
Proposition 2 states that Assumptions 1-4 are sufficient for $\hat{\beta}$ and $\widehat{\lambda(w_{i})}$ to be consistent.
The proof of Proposition 2 is complicated by the fact that codegree functions are not finite dimensional. There is no adequate notion of a density for the distribution of codegree functions, which plays a key role in the standard theory ferraty2006nonparametric. Furthermore, even when the functions $f$ and $\lambda$ are relatively smooth, the bias of $\hat{\beta}$ may still be large relative to its variance. To make reliable inferences about $\beta$ using $\hat{\beta}$, I recommend a bias correction. See auerbach2021identification for details.
This paper proposes a new way to incorporate network data into econometric modeling. An unobserved covariate called social influence is determined by an agent's link function, which describes the collection of probabilities that the agent is linked to other agents in the population. Estimation is based on matching pairs of agents with similar columns of the squared adjacency matrix.
Understanding how to incorporate different kinds of network data into econometric modeling is an important avenue for future research. A contribution of this paper is to demonstrate that in some cases identification and estimation is possible without strong parametric assumptions about how the network is generated or exactly which features of the network determine the outcome of interest.