EconBase
← Back to paper

Identification and Estimation of a Semiparametric Logit Model using Network Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

58,234 characters · 15 sections · 36 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation of a Semiparametric Logit Model using Network Data

\onehalfspacing

\abstract

This paper studies identification and estimation in semiparametric logit models when social networks are endogenous. In many applications, unobserved individual traits shape both the outcome of interest and the formation of social ties, so standard logit specifications, including those augmented with common network controls, can be biased. I show how network data can be used to address this endogeneity without imposing a parametric structure on the link formation process. Although the outcome equation is semiparametric in this social component and the network formation process is left unspecified, the logistic distribution assumption is crucial for identification. I show that slope parameters are point identified by pairwise comparisons of agents who share identical network formation behavior. I propose feasible estimators based on matching agents using network similarity measures and establish their consistency and asymptotic normality. Monte Carlo simulations demonstrate good finite-sample performance, and an empirical application to microfinance adoption demonstrates that accounting for endogenous network formation materially affects estimated covariate effects.

Introduction

A large empirical literature documents the importance of peer and network effects in economic settings, showing that individuals’ behavior and outcomes are shaped by interactions with their peers manski1993identification, sacerdote2001peer, calvo2009peer, christakis2008collective, bramoulle2020peer,de2020consumption. These effects have been studied across contexts such as education (e.g., calvo2009peer), consumption de2020consumption, labor effort and productivity beugnot2013social, and social influence more broadly jackson2017economic, and surveys highlight the role of network structure in both shaping individual outcomes and complicating causal analysis bramoulle2020peer. In most of these applications, individuals are embedded in social networks that both shape behavior and reflect underlying unobserved heterogeneity, creating challenges for identification.

A central econometric difficulty is that social networks are rarely exogenous. Unobserved individual traits such as motivation, trust, ability, or entrepreneurial skill often influence both outcomes and the formation of social ties. When these latent characteristics enter both equations, standard estimation strategies do not generally deliver consistent estimates. In binary choice models, applied work typically relies on naive logit or probit specifications, logit models augmented with peer averages or network controls, or control function approaches based on parametric link formation models. These strategies can fail even when rich network data are available, because latent traits that govern link formation remain in the outcome error term.

This paper develops identification and estimation methods for semiparametric logit models with endogenous social networks. I study a setting in which the outcome equation includes an unknown function of an unobserved individual characteristic that also governs link formation. Such latent characteristics may represent motivation, trust, ability, or family expectations, which are difficult to measure directly but play a central role in both social interactions and individual decisions. I show that network data can instead be used to control for this source of endogeneity without imposing parametric restrictions on the network formation process.

The identification argument builds on two strands of the econometric literature. First, it is closely related to fixed effects and conditional likelihood ideas developed in CHAMBERLAIN19841247, where differencing across observationally equivalent units eliminates latent heterogeneity. Second, it draws on recent advances in the econometrics of networks, particularly auerbach2022identification, who shows that in partially linear models, agents with identical network formation types can be used to difference out unobserved heterogeneity. Auerbach’s paper makes an important contribution by demonstrating how network formation behavior itself can serve as a sufficient statistic for latent traits in linear models.

The present paper extends this logic to a nonlinear binary choice framework. This extension introduces new challenges that do not arise in linear models and require different arguments for identification and inference.. In contrast to linear models, differencing arguments in binary response models require specific distributional structure. In particular, the logistic distribution assumption is crucial for identification. The identification strategy relies on a conditional likelihood argument analogous to Chamberlain’s conditional logit, where pairwise comparisons between agents with identical network formation behavior difference out the latent social component. This argument does not extend to arbitrary single-index models, and the logit structure plays a central role in delivering point identification of slope parameters.

This framework is motivated by a wide range of empirical applications. In education, students may form friendships based on unobserved academic motivation or aspirations, while simultaneously being influenced by peers in decisions such as college major choice or persistence sacerdote2001peer, feld2022effect, pu2021peers. In social programs, households may form ties based on trust or shared norms and learn about program participation through those ties, as in microfinance diffusion banerjee2013diffusion. In professional settings, researchers with similar unobserved productivity or collaboration preferences may form coauthorship networks that also shape future research output ductor2014social. In each of these examples, unobserved traits affect both outcomes and network formation, and binary choices such as participation, adoption, or selection are central objects of interest.

I formalize these ideas using a semiparametric logit model in which social influence enters as an unknown function of a latent individual index. The network is generated by a nonparametric link formation model driven by the same latent index. Identification relies on the assumption that social influence depends on the latent characteristic only through its implied network formation behavior. Importantly, I do not require the latent index itself to be identified, nor do I impose functional form assumptions on the network formation process. The parameters of interest are identified up to equivalence classes defined by network formation behavior, which is sufficient for recovering the slope coefficients of the binary response model.

The paper contributes to the econometrics of social interactions and networks in three main ways. First, it provides identification results for semiparametric binary choice models with endogenous network formation, a setting that has received far less attention than linear outcome models. Second, it delivers feasible estimators that control for network endogeneity without specifying a parametric model of link formation. Third, it shows how network data can be used as a source of identifying variation in nonlinear models, extending matching and differencing ideas from the linear network literature to binary response settings.

Monte Carlo simulations demonstrate that the proposed estimator substantially reduces bias relative to naive logit specifications and commonly used control-based approaches, and that its performance improves with sample size across a range of network formation mechanisms. An empirical application to microfinance adoption in rural India banerjee2013diffusion illustrates the practical relevance of the approach. Using pre-treatment social network data, the estimator produces stable and precisely estimated covariate effects, both with and without village fixed effects. The results show that controlling for endogenous network formation can materially affect estimated slope coefficients, even after accounting for village-level heterogeneity.

By focusing explicitly on binary choice models, this paper complements and extends existing work on endogenous networks in linear settings. More broadly, it shows how network formation behavior can serve as a source of identification in nonlinear models without imposing parametric restrictions on link formation.

The remainder of the paper is organized as follows. In the next section, I present the literature review. Section (ref) introduces the model and presents identification and estimation results. In Section (ref), I present the results of Monte Carlo simulations. Section (ref) provides an empirical application. Section (ref) concludes.

Literature Review

This paper relates to a large literature on social interactions, peer effects, and the econometrics of networks. A central finding in this literature is that individual behavior is often shaped by peers and social environments. Early empirical work documents the importance of neighborhood and group effects in settings such as youth behavior and crime case1991company and educational outcomes sacerdote2001peer. Related evidence appears in health behaviors and social contagion, including smoking and obesity christakis2008collective. jackson2017economic provide a comprehensive survey of the economic consequences of network structure and highlight the role of social interactions across a wide range of applications.

Despite this extensive empirical evidence, identification of social interactions remains challenging. manski1993identification emphasizes the reflection problem and the difficulty of disentangling endogenous social effects from correlated effects. Subsequent work further clarifies the sources of bias arising from endogenous group formation and simultaneity moffitt2001policy, soetevent2006empirics. A recurring issue is that individuals sort into peer groups based on unobserved characteristics that also directly affect outcomes, rendering standard regression approaches difficult to interpret causally.

A growing literature addresses these concerns by explicitly modeling network formation as endogenous. In this line of work, unobserved individual characteristics affect both outcomes and the formation of social links, leading to joint models of behavior and networks. Several papers adopt parametric or semiparametric specifications for link formation in order to achieve identification. Examples include goldsmith2013social, hsieh2016social, arduini2015parametric. While these approaches can be powerful, their validity depends on the correctness of the assumed network formation model, which may be difficult to justify in many empirical settings with complex and heterogeneous linking patterns.

More recent work seeks to relax parametric assumptions on network formation and instead use network structure as a source of information about latent heterogeneity. graham2015methods surveys identification strategies in social networks and emphasizes the role of network data in controlling for unobserved individual characteristics. Most closely related to this paper, auerbach2022identification shows that parameters in a partially linear regression model with endogenous networks can be identified by pairwise differencing between agents who share identical network formation behavior. Auerbach also proposes feasible estimators based on matching agents using codegree information, drawing on results from graph limit theory lovasz2012large. johnsson2021estimation propose an alternative control function approach for peer effects in endogenous networks under different identifying assumptions, illustrating complementary ways to address network endogeneity.

The present paper contributes to this literature by focusing on binary response models. While much of the existing work emphasizes linear or partially linear outcome equations, many leading empirical questions involve discrete choices such as participation, adoption, or selection decisions. In such settings, endogeneity arising from latent traits that affect both outcomes and network formation can be particularly severe, and commonly used network controls do not generally resolve it. This paper develops identification and estimation results for a semiparametric logit model in which social influence enters the outcome equation through an unknown function of a latent individual characteristic that also governs link formation. The identification strategy builds on the idea of network type equivalence but extends it to a nonlinear binary choice framework, which introduces distinct challenges for identification and inference. I provide a conditional likelihood representation that identifies the slope coefficients through matched comparisons of observationally equivalent agents and develop feasible estimators that implement these comparisons using codegree-based matching.

Model

Consider a finite set of agents $I = \{1, 2, \cdots, n\}$, where each agent is identified by an observed vector of explanatory variables $X_i \in \mathbb{R}^k$, with a binary outcome $ y_i \in \{0,1\}$, and an unobserved index of social characteristics $\omega_i\in[0,1]$. These variables are related by the following model.

equation[equation omitted — 102 chars of source]

where $\{\varepsilon_i\}_{i=1}^{n}$ is a collection of independent and identically distributed (i.i.d.) shock with a logistic cumulative distribution function (cdf) $F$; $\beta\in\mathbb{R}^k$ is an unknown slope parameter. The explanatory variable $X_i$ does not contain an intercept. The social influence function $\lambda: [0,1]\to\mathbb{R}$ is an unknown measurable function with $\lambda(\omega_i)$ being the realized social influence for agent $i$. The social influence term $\lambda(\omega_i)$ absorbs the intercept, it is the direct effect of interacting with a particular collection of communities.

In addition to this, the researcher also observes a binary adjacency matrix $n\times n$, $D$. The element $D_{ij}=1$ if there is a direct link between agents $i$ and $j$, and $D_{ij}=0$ otherwise. By convention, any agent $i$ is not allowed to be linked to itself $D_{ii}=0$. I assume that all links are undirected, so that the adjacency matrix $D$ is symmetric, that is, $D_{ij}=D_{ji}$. Social links are generated according to a nonparametric network formation model. The existence of a link between agents $i$ and $j$ is determined by

equation[equation omitted — 114 chars of source]

where $f$ is a symmetric measurable function and $\left\{\eta_{ij}\right\}^n_{i,j=1}$ is a symmetric matrix of unobserved scalar disturbances with independent upper diagonal entries that are mutually independent of $\{x_i,\omega_i,\varepsilon_i\}_{i=1}^n$. The latent variable $\omega_i$ captures characteristics such as motivation, trust, ability, or social capital, which are typically unobserved but play a role in both outcomes and link formation.

The unobserved individual index of social characteristic $\omega_i$ can be interpreted as the social capital that increases the likelihood of forming a link (e.g.; socioeconomic status, social ability, family expectation, trust). The difference $f(\omega_i,\omega_j)-\eta_{ij}$ is interpreted as the utility agents $i$ and $j$ receive from forming a link. This implies that for different collection of indexes, $D_{ij}$ and $D_{st}$ are independent conditionally on $\omega_i, \omega_j, \omega_s$ and $\omega_t$ i.e. the utility agents receive from forming a link does not depend on the existence of other links in the sample.

The following examples illustrate applications of the model to the literature.

example[Program Participation] banerjee2013diffusion model household participation in a microfinance program in which information about the program diffuses over a social network. They measure participation using lending and trust. Their model could be generalized by adding $\omega_i$ as an index of individual trustworthiness, family expectations, and integrity in financial matters.
example[Peers Effect] The peer effects of this type would be a model for youth smoking behavior, where smokers could be more likely to form friendships with each other. Let $y_i=1$ if a student smokes and 0 if not, $X_i$ be a vector of student $i$ covariates (age, grade, gender, etc.), and $D_{ij} = 1$ if students $i$ and $j$ are friends and 0 otherwise. The number of direct neighbors of student $i$ is $s_{1i}=\sum_{j\ne i}D_{ij}$. An extension of the menzel2015strategic peers effect model is by setting $\lambda(\omega_i)=\delta\frac{1}{s_{1i}}\sum_{i\ne j}D_{ij}y_i$ which correspond to the social influence in our model.
example[Co-authorship] ductor2014social study how knowledge about the social network of a researcher, as embodied in his coauthor relations, helps them in developing a prediction of his or her future productivity. In this setting, $\omega_i$ can be interpreted as some unobserved productivity trait that induces the researcher to have more coauthors and also to be more productive at writing papers.

Identification

The parameters of interest are $\beta$ and $\lambda(\omega_i)$. The functions $\lambda$ and $f$ will be treated as nuisance functions because they cannot be identified given that $\omega_i$ is unknown. However, the realization of $\lambda$ for a particular individual, $\lambda(\omega_i)$, can be identified. Identification relies on the concept of network formation types. The parameter $\beta$ is identified if $\lambda(\omega_i)$ depends on $\omega_i$ only through the link function $f_{\omega_i}(\cdot) \equiv f (\omega_i,\cdot) : [0, 1]\to [0, 1]$, which is the conditional probability that an agent with social characteristics $\omega_i$ links with agents of every other social characteristic in $[0, 1]$. In order to compare the two agents' types, I define the integrated squared difference in the network types of agents with social characteristics $\omega_i$ and $\omega_j$ by

equation[equation omitted — 139 chars of source]

That is if the measure of the network distance between agents $i$ and $j$ equals zero ($ \rho_{ij}=0$), then there is no identifiable characteristic within the network that sets $\omega_i$ apart from $\omega_j$. Given this scenario, agents $i$ and $j$ are equally likely to form connections within any specific network configuration. Consequently, they would have identical distributions of degrees, eigenvector centralities, and average peer characteristics, as well as any other individual-specific statistic of the network representation $D$.

assumption[Sampling and disturbances] (i) $(X_i,\omega_i,\varepsilon_i)$ are i.i.d. for all $i=1, \cdots, n$; (ii) The random array $\{\eta_{ij}\}_{i,j=1}^n$ is symmetric and independent of $(X_i,\omega_i,\varepsilon_i)$ with i.i.d. entries above the diagonal; (iii) $\omega_i$ and $\eta_{ij}$ have standard uniform marginals; (iv) $\varepsilon_i$ follows a standard logistic distribution; (v) The binary outcomes $\{y_{i}\}_{i=1}^n$ and the binary adjacency matrix $D$ are given respectively by equations ((ref)) and ((ref)); (vi) $\lambda$ and $f$ are Lebesgue-measurable with $f$ being symmetric in its arguments.

Assumption (ref)(i) implies that the observables $X_i$ and the unobservable individual characteristics $(\omega_i,\varepsilon_i)$ are randomly drawn. This is a standard assumption in the network literature. Assumption (ref)(ii) assumes that the link formation error $\eta_{ij}$ is orthogonal to all other observables and unobservables in the model. This means that the dyad-specific unobservable shock $\eta_{ij}$ from the link formation process does not influence the binary outcomes $\{y_{i}\}_{i=1}^n$. The endogeneity in this model takes the form of a dependence between $X_i$ and the unobserved error $\lambda(\omega_i) + \varepsilon_i$ through $\omega_i$. From Assumption (ref)(iii), the marginal distributions of $\omega_i$ and $\eta_{ij}$ are assumed to have standard uniform marginals without loss because I cannot separately identify them from $f$. Under this assumption, $f(\omega_i,\omega_j)$ is the probability that agent $i$ and $j$ form are directly linked, $\mathbb{P}(D_{ij}=1)=f(\omega_i,\omega_j)$, which implicitly assumes that $f:[0,1]^2\to[0,1]$.

The key identifying restriction links network equivalence to social influence.

assumption[auerbach2022identification] The social influence function $\lambda$ satisfies $$\mathbb{E}\left[(\lambda(\omega_i)-\lambda(\omega_j))^2|\rho_{ij}=0\right]=0.$$

Assumption (ref) states that agents with similar network types have similar social influences. In other words, this means that if two agents interact with similar groups of people or have similar connections, they are likely to have similar opinions, behaviors, or attitudes. $\rho_{ij}=0$ implies that $f_{\omega_i}(\cdot)=f_{\omega_j}(\cdot)$ and $\lambda(\omega_i)=\lambda(\omega_j)$ under Assumption (ref), but does not imply that $\omega_i=\omega_j$.

theoremUnder Assumptions (ref) and (ref), $\beta$ and $\lambda(\omega_i)$ are uniquely identify. \begin{equation} \beta = \arg\min_{b\in\mathbb{R}^k} -\mathbb{E}\left[y_{i}\log F[(X_i-X_j)b]+y_j\log F[(X_j-X_i)b]\bigg|\rho_{ij} =0,y_i+y_j=1\right] \end{equation} and \begin{equation} \lambda(\omega_i) = \mathbb{E}\left[F^{-1}\big(\mathbb{P}(y_i=1|X_i,f_{\omega_i})\big)-X_i{\beta}\big|f_{\omega_i}\right] \end{equation}

The proof of Theorem (ref) is provided in \hyperlink{app_theo1}{Appendix}. The conditional expectation in ((ref)) is defined as follows: for any arbitrary random matrix $\Pi_i$ indexed at the agent level, $$\mathbb{E}\left[\Pi_i\left|f_{\omega_i}\right.\right]\equiv\mathbb{E}\left[\Pi_i\left|\omega_i\in\{u\in[0,1]:||f_{\omega_i}-f_u||_2=0\}\right.\right]$$

Estimation

This section provides a structured discussion regarding the estimation of the parameters of interest from Section (ref). If the individual social characteristics $\{\omega_i\}_{i=1}^n$ were observed then I can use existing tools to estimate $\beta$ and $\lambda(\omega_i)$. The link formation model in ((ref)) will not provide useful information in this case. The researcher will just have to use observations for which $\omega_i$ is close to $\omega_j$ honore1997pairwise,aradillas2007pairwise. However, since individual social characteristics $\{\omega_i\}_{i=1}^n$ are not observed and the conditional mean functions in Theorem (ref) are also not known, estimates are not feasible.

In order to make these estimates feasible, I exploit the theorem of lovasz2012large which demonstrates that pairs of individuals with identical link functions have identical codegree functions. Following auerbach2022identification, the codegree function is defined by $p(\omega_i,\omega_j)=\int f_{\omega_i}(s)f_{\omega_j}(s)ds$, which is the probability that agents $i$ and $j$ have a link in common. The agent $i$'s codegree type is defined as $p_{\omega_i}(\cdot)\equiv p(\omega_i,\cdot): [0,1]\to[0,1]$. The pseudometric codegree distance between agents $i$ and $j$ is defined by:

align*[align* omitted — 156 chars of source]

which can be consistently estimated\footnote{auerbach2022identification shows in Lemma B1 that $\hat{\delta}_{ij}$ converges uniformly to $\delta_{ij}$ over the $\binom{n}{2}$ distinct pairs of agents as $n\to\infty$.} by the root average squared difference in the $i$th and $j$th columns of the squared adjacency matrix

align*[align* omitted — 149 chars of source]

The identification argument relies on the relationship between the latent network distance $\rho_{ij}$ and the observable codegree distance $\delta_{ij}$. Under standard regularity conditions on the link function $f$, $\rho_{ij}$ and $\delta_{ij}$ are equivalent. This equivalence result follows from arguments in the graph limit and network econometrics literature, and are established formally in Lemma 1 in auerbach2022identification.

A direct consequence of this result is that if the link function $f$ is continuous and not constant almost everywhere then, $\rho_{ij}=0\iff\delta_{ij}=0$. Using this equivalence and the fact that $\delta_{ij}$ can be consistently estimated by $\hat{\delta}_{ij}$, I can therefore confidently substitute $\rho_{ij}=0$ by $\hat{\delta}_{ij}=0$ in order to consistently estimate the parameters of interest.

assumptionThe following statements about the kernel function $K$ and the bandwidth $h$ hold: \begin{enumerate}[(i)] • $K$ is non-negative, bounded, differentiable with bounded derivative $K'$, and \[\int K(u)du=1, \ \ \int \left|K(u)\right|du<\infty \ \ \mbox{ and } \ \ \int |K(u)||u|du<\infty \]$h > 0, \ \ h = o(1), \ \ h^{-1}= O(\sqrt{n})$, \ \ and \ \ $n\mathbb{E}\left[K\left(\frac{\hat{\delta}^2_{ij}}{h}\right)\right]\to\infty$. \end{enumerate}

Assumption (ref)((ref)) is the standard restriction of the kernel function $K$. The first three restrictions in Assumption (ref)((ref)) on the bandwidth sequence $h$ are also standard. The last one guarantees that as the sample size $n$ increases, the number of matches used to estimate $\beta$ also increases.

Under Assumptions (ref)-(ref), $\beta$ is consistently estimated by

equation[equation omitted — 213 chars of source]

This estimator can be interpreted as a kernel-weighted conditional logit estimator. The estimator of the social influence term, $\lambda(\omega_i)$, is then given by

equation[equation omitted — 311 chars of source]

where $$\widehat{\mathbb{P}}\left(y_i=1\left|X_i,f_{\omega_i}\right.\right)=\left[\sum_{j=1}^nK\left(\frac{\hat{\delta}^2_{ij}}{h}\right)K\left(\frac{X_j-X_i}{h}\right)\right]^{-1}\left[\sum_{j=1}^ny_jK\left(\frac{\hat{\delta}^2_{ij}}{h}\right)K\left(\frac{X_j-X_i}{h}\right)\right]$$is a consistent estimator for $\mathbb{P}\left(y_i=1\left|X_i,f_{\omega_i}\right.\right)=\mathbb{E}\left[y_i\left|X_i,f_{\omega_i}\right.\right]$; $K$ is a kernel function and $h$ is a bandwidth parameter depending on the sample size. The term $K\left(\frac{\hat{\delta}^2_{ij}}{h}\right)$ gives more weight to pairs of observations $(i, j)$ with identical codegree function. This estimator averages over agents with similar network types, thereby recovering social influence up to network-type equivalence classes.

The consistency and asymptotic normality of these estimators will be developed in the next section.

Consistency and Asymptotic Normality

This section establishes the large sample properties of the estimators introduced in Section (ref). I first show consistency of the slope coefficient estimator and the estimated social influence function, and then derive asymptotic normality of the slope estimator. Proofs are provided in the Appendix.

Let \[m(v_i,v_j,b) =-\mathbbm{1}(y_i\ne y_j)\left\{y_{i}\ln F\left[(X_i-X_j)'b\right]+y_j\ln F\left[(X_j-X_i)'b\right]\right\}\] where $v_i=(y_i,X_i)$. The sample objective function can be written as

equation[equation omitted — 146 chars of source]

The estimator $\hat{\beta}$, as defined in equation ((ref)), is derived through the minimization of the sample objective function specified in equation ((ref)). In the following two subsections, I will provide conditions on which this estimator is consistent and asymptotically normal.

Consistency

To prove that our estimator $\hat{\beta}$ is consistent, I will take advantage of the following Assumptions and theorems found in newey1994large.

Let us define the following function \[l(x,y,b)=\mathbb{E}\left[m(v_i,v_j,b)\left|v_i=x,\lambda(\omega_j)=y\right.\right].\] Note that \[l(x,\lambda(\omega_i),b)=\mathbb{E}\left[m(v_i,v_j,b)\left|v_i=x,\delta_{ij}=0\right.\right].\]

assumptionAll the following Assumptions hold \begin{enumerate} • $\mathbb{E}\left[m(v_i,v_j,b)^2\right]<\infty$; • The function $l(\cdot)$ defined above exists and is a continuous function of each of its arguments; • For all $b$, $|l(x,y, b)|\le t(x, b)$ with $\mathbb{E}\left[t(v_i, b)\right] < \infty$. \end{enumerate}
theoremIf Assumptions (ref)-(ref) hold. If the parameter space for $b$ is compact and includes the true value of $\beta$. Then, the estimator $\hat{\beta}$ of $\beta$ defined in ((ref)) and the estimator $\hat{\lambda}(\omega_i)$ of $\lambda(\omega_i)$ defined in ((ref)) are consistent.

The proof of Theorem (ref) is provided in \hyperlink{app_theo2}{Appendix}.

Asymptotic Normality

To derive asymptotic normality, note that $\hat\beta$ is an extremum estimator based on a U-statistic–type objective function. Under additional regularity conditions ensuring smoothness and nondegeneracy, the estimator admits a linear expansion around the true parameter value.

assumptionThe following holds: \begin{itemize} • $\mathbb{E}\left(||X_i||^2\right)<\infty$ • Conditional on $\delta_{ij} = 0$, $(X_i-X_j)$ has full rank; \end{itemize}
theoremUnder Assumptions (ref)-(ref), the estimator $\hat{\beta}$ of $\beta$ defined in ((ref)) is asymptotically normal with \[\sqrt{n}\left(\hat{\beta}-\beta\right)\xrightarrow{d}\mathcal{N} \left(0,4\Sigma^{-1}V\Sigma^{-1}\right)\] with $$V=Var(t(y_i,x_i,w_i))$$ where \[t(y_i,x_i,w_i)=\mathbb{E}\left[\mathbbm{1}(y_i\ne y_j)\left\{y_i-F\left((X_i-X_j)'\beta\right)\right\}(X_i-X_j)\left|y_i,x_i,\delta_{ij}=0\right.\right]\] and \begin{align*} \Sigma&=\mathbb{E}\left[\Delta_\beta(t(y_i,x_i,w_i))\right]\\ &=\mathbb{E}\left[\mathbb{E}\left[\mathbbm{1}(y_i\ne y_j)F\left((X_i-X_j)'\beta\right)F\left((X_j-X_i)'\beta\right)(X_i-X_j)(X_i-X_j)'\right]\left|y_i,x_i,\delta_{ij}=0\right.\right] \end{align*}

The asymptotic variance has the familiar sandwich form and reflects both sampling variation and uncertainty induced by matching on estimated network types. Despite the use of pairwise comparisons, the effective sample size grows at rate $n$, leading to standard $\sqrt{n}$ convergence. The complete proof of Theorem (ref) is provided in \hyperlink{app_theo3}{Appendix}.

Monte Carlo Simulation

I present in this section the results of the simulation using the estimators developed in the previous section. The simulations are designed to highlight three features of the estimator: bias reduction relative to standard logit specifications, robustness to different network formation mechanisms, and convergence toward the asymptotic approximation as sample size increases.

I generate simulated data using the model described in section (ref) in equations ((ref)) and ((ref)). The slope parameter is set to $\beta=1$, and the social influence function is specified as $\lambda(w) = (4w^3-3)/2$ which introduces nonlinear dependence on the unobserved social characteristic. To induce endogeneity, the observed regressor $x_i$ is constructed as $x_i =(3\omega_i^3+\xi_i+1)/3$, where $\xi_i\sim\mathcal{N}(0,1)$, so that $x_i$ is correlated with the unobserved social characteristic $\omega_i$.

Unobserved social characteristic $\omega_i$ and the link formation shocks, $\eta_{ij}$, are independently drawn from the uniform distribution on $[0, 1]$ for $i<j$ and $\eta_{ij}=\eta_{ji}$. The outcome shocks $\varepsilon_i$ follow a standard logistic distribution.

The adjacency matrix is generated according to equation ((ref)) using three alternative link formation functions:

enumerate• Homophily model \[f_1(x,y)=1-(x-y)^2\] • Beta model \[ f_2(x,y)=\frac{\exp(x+y)}{1+\exp(x+y)}\] • Stochastic block model \[f_3(x,y) = \left\{\begin{array}{cl} 1/3 & \mbox{ if } x\le1/3 \mbox{ and } y>1/3\\ 1/3 & \mbox{ if } 1/3<x\le2/3 \mbox{ and } y\le2/3 \\ 1/3& \mbox{ if } x>2/3 \mbox{ and } (y>2/3 \mbox{ or } y\le1/3)\\ 0& \mbox{ otherwise } \end{array}\right.\]

These designs capture qualitatively different forms of network formation, including assortative matching, degree heterogeneity, and community structure respectively.

For each link formation model, I run 500 Monte Carlo replications with sample sizes $n\in\{100, 250, 700, 1000, 2000\}$. I use the bandwidth sequence $h=n^{-1/9}/10$ and the Epanechnikov kernel $K(x)=\frac{3}{4}(1-x^2)\mathbbm{1}(x^2<1)$.

The performance of my proposed estimators is evaluated in comparison to the estimators of the following three benchmark models:

itemize• Naive Logit: $y_i = \mathbbm{1}\left\{\alpha + \beta_1x_i-\epsilon_i\ge0\right\}$, which ignores network dependence; • Logit with controls: $y_i = \mathbbm{1}\left\{\alpha +\beta_2x_i+\mu\frac{\sum_{j=1}^nD_{ij}x_j}{\sum_{j=1}^nD_{ij}} + \gamma\frac{\sum_{j=1}^nD_{ij}y_j}{\sum_{j=1}^nD_{ij}}-u_i\ge0\right\}$, which includes average neighbor covariates and outcomes. • Infeasible Logit: $y_i = \mathbbm{1}\left\{\alpha +\beta_3x_i+\theta\lambda(\omega_i)-v_i\ge0\right\}$, which includes the true social influence and therefore serves as a lower bound on achievable bias;

For each model, link function, and sample size, I report the average bias over the 500 simulations. Tables (ref)–(ref) report the average behavior over the 500 simulations results for homophily, beta, and stochastic block networks, respectively. Performance is evaluated using bias, standard deviation, root mean squared error, and coverage of nominal 95 percent confidence intervals.

table[table omitted — 1,237 chars of source]

Consitency and Convergence

A first and striking result is that naive logit estimators perform extremely poorly in all designs. The absolute bias remains large and essentially constant as the sample size increases. Under homophily networks, for example, the bias of the naive estimator remains above 0.82 even when $n = 2000$. Similar magnitudes appear under beta and stochastic block networks. The standard deviation shrinks with sample size, but the bias does not, so the RMSE remains large and coverage collapses toward zero. In large samples, coverage for the naive estimator is effectively zero across all network designs. This confirms that ignoring endogenous network formation produces asymptotically biased and severely misleading inference.

Augmenting the logit model with common network controls improves performance only marginally. While the bias is sometimes reduced relative to the naive specification, it remains substantial and does not vanish with sample size. Coverage deteriorates rapidly as $n$ increases, again approaching zero in larger samples. These findings indicate that ad hoc network controls are insufficient to eliminate bias when the covariate is correlated with latent network traits.

The infeasible logit estimator serves as a benchmark. As expected, it exhibits negligible bias across all designs and achieves coverage close to the nominal level. Its RMSE decreases and its coverage increases at the standard rate as $n$ grows. This confirms that the source of bias in the naive and control-based estimators is precisely the latent network heterogeneity entering both the outcome and link formation.

My proposed estimator substantially improves upon standard logit specifications. Although finite-sample bias is non-negligible in smaller samples—particularly under stochastic block networks—the bias decreases steadily with sample size in all designs. Under homophily networks, for example, the bias declines from 0.262 at $n = 100$ to 0.079 at $n = 2000$. Similar convergence patterns appear under beta and stochastic block models. The RMSE declines monotonically with $n$, and coverage approaches the nominal level in large samples, especially under homophily and beta networks.

Two features are worth emphasizing. First, unlike the naive and control-based estimators, my proposed estimator does not exhibit asymptotic bias. Its performance converges toward that of the infeasible benchmark as the sample size increases. Second, the relative difficulty of the design depends on the network formation model. Bias reduction is fastest under homophily networks and somewhat slower under stochastic block networks, reflecting differences in how sharply network types can be distinguished via codegree matching.

Overall, the simulations demonstrate that standard logit approaches can be severely misleading in the presence of endogenous network formation, even in large samples. In contrast, the proposed estimator restores consistency without imposing parametric structure on the link formation process and achieves reliable inference in samples of empirically relevant size.

table[table omitted — 1,246 chars of source]
table[table omitted — 1,362 chars of source]

Sensitivity to network structure

Comparing across Tables (ref)-(ref) highlights how performance depends on the structure of network formation. Under homophily networks, where similarity in latent traits translates smoothly into link probabilities, the proposed estimator exhibits the fastest decline in bias and the most rapid improvement in coverage. The continuous variation in network similarity facilitates accurate matching and reduces finite-sample distortions. Under beta networks, where degree heterogeneity drives link formation, bias reduction is slower but still monotonic. The estimator continues to converge toward the infeasible benchmark, although finite-sample bias remains more pronounced at moderate sample sizes.

Performance is weakest under stochastic block networks, where link probabilities change discontinuously across blocks. In this environment, distinguishing network types via codegree similarity is more challenging in finite samples, leading to slower convergence. Nonetheless, bias still declines as n increases, and the estimator remains substantially less distorted than naive logit specifications.

These comparisons illustrate that the estimator does not rely on a particular mode of network formation. It remains consistent under discrete block structures, degree-driven networks, and continuous homophily designs. However, the speed of convergence depends on how sharply network types can be identified from observed link patterns.

Taken together, the Monte Carlo evidence supports three conclusions. First, standard logit approaches can be severely biased in the presence of endogenous network formation, even in large samples. Second, matching on network formation behavior restores consistency without requiring parametric modeling of link formation. Third, although finite-sample bias may be non-negligible when network types are difficult to separate, the estimator converges toward the infeasible benchmark as the network size increases.

Empirical Application

This section applies the proposed estimator to household-level data on the diffusion of microfinance in rural Indian villages. The purpose of the exercise is to illustrate how the estimator can be implemented in a canonical network setting and to assess how accounting for endogenous network formation affects estimated covariate effects in a binary adoption model. The analysis is intentionally reduced-form and complements existing structural approaches to diffusion.

Data and Context: banerjee2013diffusion

The empirical application uses the village social network data collected and analyzed in banerjee2013diffusion paper. The authors studied the spread of a microfinance loan program in 43 villages of Karnataka, India. In each village, a couple of village leaders like teachers, leaders of self-help groups, and shop-keepers, were invited to a private meeting, in which they were informed of the program and encouraged to spread information about the program to their peers. The data include detailed descriptions of each household’s social relationships, including which friends or relatives visit the individual, which friends or relatives the individual visits, from whom the individual would borrow money from, to whom the individual would lend money, etc\footnote{They collected social network information across 12 dimensions, encompassing various aspects of the respondent’s social interactions. These dimensions include: 1. Individuals who visit the respondent’s home. 2. Individuals whose homes the respondent visits. 3. Family members within the same village. 4. Non-relatives with whom the respondent socializes. 5. Sources from which the respondent seeks medical advice. 6. Potential sources for borrowing money. 7. Potential individuals to whom the respondent would lend money. 8. Sources from which the respondent might borrow material goods (such as kerosene or rice). 9. Potential recipients to whom the respondent would lend material goods. 10. Advisors from whom the respondent seeks guidance. 11. Individuals to whom the respondent provides advice. 12. Individuals with whom the respondent engages in communal prayer, whether at a temple, church, or mosque.}. Information on household characteristics is also collected, e.g., whether the household has electricity or latrine, the numbers of rooms and beds, etc. The outcome of interest is a binary indicator equal to one if a household adopts microfinance during the study period.

Although social links are measured before the introduction of microfinance, unobserved household characteristics such as financial sophistication, trust in financial institutions, risk tolerance, or entrepreneurial ability plausibly affect both a household’s likelihood of adopting microfinance and its pattern of social interactions. These latent traits are not directly observed by the econometrician but may be reflected in network formation behavior.

Results

The empirical results of three regression models are reported in Table (ref). Columns (1)–(2) present estimates from a standard logit specification, columns (3)–(4) report logit estimates augmented with network controls, and columns (5)–(6) report estimates obtained using my proposed estimator. Two specifications are reported for each model. The first specification excludes village fixed effects and relies solely on network-based matching to control for unobserved heterogeneity. The second specification augments the model with village fixed effects, absorbing common village-level factors such as program intensity, local economic conditions, and institutional differences across villages. In both cases, standard errors are computed using the asymptotic variance formula derived in Section (ref).

table[table omitted — 2,874 chars of source]

Several patterns emerge from Table (ref). First, the estimates obtained using the proposed estimator (columns 5 and 6) are economically meaningful and precisely estimated across both specifications. All reported coefficients are statistically significant, indicating strong associations between observed household characteristics and microfinance adoption once latent social influence is controlled for through network formation behavior. In particular, the number of rooms and access to electricity are positively associated with adoption, while measures of household crowding such as beds per capita and rooms per capita exhibit negative effects.

Second, the inclusion of village fixed effects leads to changes in the magnitude of several coefficients but does not overturn their signs or statistical significance under the proposed estimator. This suggests that while village-level heterogeneity explains part of the variation in adoption, the network-based control for latent social traits captures an important source of within-village heterogeneity that remains even after conditioning on village fixed effects. The stability of the estimates across the two specifications indicates that identification primarily relies on within-village variation among households with similar network formation behavior, which is consistent with the identifying logic of the estimator.

Third, comparing across estimators highlights notable differences between the proposed approach and standard logit specifications. In the naive logit model, the coefficient on Has Latrine changes sign once village fixed effects are included, suggesting that the naive specification may confound household characteristics with unobserved village-level or network-related factors. Similarly, the coefficient on Beds per capita is statistically insignificant in both the naive logit and the logit with network controls, while it becomes precisely estimated under the proposed estimator. The estimated effect of No. of Beds is also less precisely estimated in the standard logit specifications relative to the proposed estimator. These differences indicate that ignoring endogenous network formation can attenuate or distort estimated relationships between household characteristics and adoption decisions.

More broadly, the magnitude of several coefficients differs across estimators. For example, the estimated effect of Has Electricity is larger under the proposed estimator than in the logit specifications, suggesting that standard logit models may understate the importance of household infrastructure once latent social influence is accounted for. Conversely, some coefficients are attenuated relative to the naive logit results, indicating that part of the variation captured by the naive specification reflects unobserved network heterogeneity rather than direct covariate effects.

Taken together, these comparisons illustrate that network-based controls and village fixed effects are complementary rather than substitutes. While village fixed effects absorb common shocks and institutional features, the network-based approach captures heterogeneity at the household level that is otherwise unobserved. This distinction is particularly relevant in settings where social interactions and individual traits vary substantially within small geographic units.

Overall, the empirical application demonstrates that the proposed estimator can be implemented in a well-known network dataset and delivers stable, interpretable results across alternative specifications. By controlling for latent social influence through network formation behavior, the estimator provides a flexible approach to estimation in binary choice models with endogenous networks. These findings complement the theoretical and simulation results and illustrate the practical relevance of the method in empirical settings with rich network data.

Conclusion

This paper studies identification and estimation in binary choice models when social networks are endogenous. In many empirical applications, unobserved individual traits influence both economic outcomes and the formation of social links, invalidating standard logit specifications and ad hoc network controls. This paper proposes an alternative approach that uses observed network data to control directly for latent social influence, without imposing parametric assumptions on network formation.

The key insight is that individuals who are observationally equivalent in their network formation behavior share the same unobserved social trait, even when that trait itself is not observed. Exploiting this equivalence, the paper establishes point identification of slope parameters in a semiparametric logit model through pairwise comparisons of agents with similar network formation behavior. A feasible estimator based on codegree information is developed, and its consistency and asymptotic normality are established under weak regularity conditions.

Monte Carlo simulations demonstrate that the proposed estimator performs well in finite samples across a range of network environments, including stochastic block, beta, and homophily designs. The estimator exhibits small bias, favorable mean squared error properties, and accurate inference in samples of empirically relevant size. Importantly, performance remains stable even when similarity in network formation behavior varies continuously rather than through discrete network types.

An empirical application to microfinance adoption banerjee2013diffusion in rural Indian villages illustrates the practical relevance of the method. Using rich pre-treatment social network data, the proposed estimator delivers stable and precisely estimated covariate effects, both with and without village fixed effects. The results indicate that network-based controls capture meaningful within-village heterogeneity in latent social influence that is not absorbed by aggregate fixed effects alone.

More broadly, the framework developed in this paper provides a flexible and transparent way to incorporate network data into nonlinear models with endogenous social interactions. While the analysis focuses on cross-sectional binary choice models, the underlying ideas extend naturally to other nonlinear settings and to models with repeated observations. Future work may consider extensions to panel data, alternative notions of network similarity, or settings with multiple outcomes and interacting networks. \hypertarget{appendix}{