Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,861 characters · 16 sections · 55 citation commands
Social Interactions Models with Latent Structures
\onehalfspacing
\onehalfspacing
Binary choice models with social interactions or peer effects\footnote{We use the terms "social interactions" and "peer effects" interchangeably.} are widely employed to study decision making processes among people connected through social networks. A typical empirical setting involves individuals distributed across many distinct groups (e.g., students assigned to classrooms with different class sizes and teaching qualities, graham2008identifying,burke2013classroom,sacerdote2011peer), where interactions occur within groups only. When estimating peer effects, a naive approach to utilizing all available data would assume homogeneous parameters (including the intercept and slopes) across groups brock2001discrete. However, the homogeneity assumption, if violated, may incur severe bias. Conversely, allowing complete heterogeneity often yields unreliable group-level estimates due to small sample sizes.
To balance the trade-off between model misspecification risk and statistical efficiency, we propose a binary choice model featuring social interactions, group fixed effects, and latent structures. In our empirical application with the National Longitudinal Study of Adolescent to Adult Health (Add Health) dataset, this model enables us to analyze heterogeneous peer effects among students across different clusters of schools. Our proposed method simultaneously identifies the latent clusters and estimates the unknown parameters, revealing that students in one of the clusters of schools exhibit statistically significant peer effects.
The binary choice model with social interactions and a latent structure is analyzed through a simultaneous game framework, where agents interact exogenously given social networks within groups. The exogenous network assumption is standard in peer effects studies. It is adopted in the linear-in-means model manski1993identification,manski2000economic and the discrete choice model brock2001discrete,lin2017estimation. The nested pseudo likelihood (NPL, aguirregabiria2007sequential,lin2017estimation) algorithm is the off-the-shelf method in estimating the structural parameters in games. The NPL algorithm sequentially estimates unknown parameters and updates the conditional choice probabilities (CCPs), thereby circumventing the computational challenges related the growing dimension of states and actions.
The data structure under consideration is similar to panel data, which is ubiquitous in econometric analysis. Tracking individuals over time, the rich information in panel data can produce more efficient estimators than those based solely on cross-sectional or time-series approaches. While parameter homogeneity is typically assumed to leverage the advantages of panel data --- including enhanced cross-sectional averaging power --- empirical studies frequently reject this assumption, see phillips2007transition, browning2007heterogeneity, su2013testing, and lu2017determining. As hsiao2022analysis notes in his Chapter 1, inferences based on invalid homogeneity assumptions may be misleading. The other side of the spectrum of specifications is complete parameter heterogeneity across all groups. However, the limited sample sizes associated with each group yield less reliable estimators.
This dilemma motivates us to examine an intermediate case through a latent structure model, where parameters are heterogeneous across clusters but homogeneous within clusters, with knowledge about neither the number of clusters nor cluster membership a priori. Many econometric papers on social interactions hinge on pre-specified cluster structures. For example, lewbel2023social analyze social effects among elementary school students in the Student/Teacher Achievement Ratio (STAR) Project in the U.S. State of Tennessee. They partition the classes (the groups) into two environments (clusters) according to class sizes and estimate the peer effects respectively for each environment (this small/large classification is also adopted by graham2008identifying to identify the peer effects). However, external variables for identifying the latent structure are not always available. To address this problem, we integrate the Classifier-Lasso (C-Lasso, \citealp*{su2016identifying}) algorithm with the NPL estimation. The C-Lasso procedure solves a penalized optimization problem where a multiplicative penalty term is appended to a primary loss function. We show that it can identify the latent structure, thereby enabling analysis of peer effects within each cluster.
While the C-Lasso procedure constitutes a non-convex problem approximable by a sequence of convex problems, its computational demands remain substantial. Moreover, the stability of the NPL algorithm may suffer when CCPs updating becomes complex. We therefore propose a multi-step estimation algorithm that decouples C-Lasso from the NPL steps: (i) apply NPL to obtain preliminary parameter and CCPs estimates for each group; (ii) plug the estimated CCPs into C-Lasso to extract cluster classifications; (iii) implement NPL to obtain post-classification estimators for each cluster. We establish large-sample results for all steps, including classification consistency for the membership and asymptotic normality of the post-classification estimators. Since group fixed effects are jointly estimated, an incidental parameter bias neyman1948consistent,hahn2004jackknife affects the estimators of the slope parameters. We are unaware that this problem has been studied in the panel data setting for the NPL estimators. We formally characterize this bias in an asymptotic framework where both the number of groups and group size grow to infinity while their ratio remains bounded. The popular split-panel jackknife estimator \citep*{dhaene2015split,li2025baggingnetwork,mei2023nickell} and other split-sample based bias-correction methods are difficult to implement in our framework because CCPs are jointly determined by all individuals in a group, preventing parameter estimation using only a subset of observations. Moreover, due to the presence of social effects, the analytical bias is too complex to estimate via simple plug-in methods. We employ higgins2024bootstrap's bootstrap inference method for panel fixed effects models to construct confidence intervals with correct coverage levels.
To summarize, our contributions are twofold. Firstly, in the classical setting, individuals over groups are homogeneous. This paper allows them to be heterogeneous, characterized by the fixed effects. We overcome the estimation bias from incidental parameters with bootstrap. Secondly, group-level heterogeneity is modeled via clusters, and we use penalized algorithms to detect such latent structures. This paper aligns with the ongoing movement in relaxing classical parsimonious econometric models to meet the challenges brought about by big data, large models, and real-world complexity.
Our paper contributes to the literature on social interactions models. manski1993identification introduced the famous linear-in-means model and demonstrated the non-identification of peer effects due to the “reflection problem”. Subsequent studies have extended and refined this framework, including lee2007identification, graham2008identifying, bramoulle2009identification, calvo2009peer, goldsmith2013social, ross2022measuring, and lin2026social, among others. We focus on binary choice models, which were pioneered by brock2001discrete,brock2001interactions. They established identification for binary choices in social interaction models where each group member's decision is influenced by a rational expectation of the average choice in the group. lee2014binary further advanced this literature by achieving identification and providing maximum likelihood estimation for a general network model with heterogeneous rational expectations. lin2017estimation,lin2021uncovering and lin2024binary investigated the estimation and inference of social interaction effects in various scenarios and proposed using NPL to reduce the computational burden of estimating CCPs. lin2025endogenous studied the endogenous treatment models with social interactions in both the treatment and outcome equations. We introduce the group fixed effects, and apply a parametric bootstrap to conduct statistical inference, which fills a gap between the literature on peer effects and panel data. Moreover, the use of latent structure modeling in this context represents a novel contribution to the strand of literature on social interactions. Heterogeneity is ubiquitous in empirical studies. In peer effects studies, people find the dependence of peer effects in a variety of background status, \citep*[e.g.][find that impact of peer weight is larger among females and adolescents with high BMI]{trogdon2008peer}. The latent structure helps us identify heterogeneous peer effects in different clusters without prior knowledge about memberships.
We also contribute to the study of fixed effects models with a latent structure. We introduce the latent structure to social interactions models, and study the classification consistency under a NPL estimation framework. The literature offers several approaches for determining an unknown cluster structure. The first approach relies on finite mixture models. For panel discrete choice models, kasahara2009nonparametric and browning2014dynamic investigated identification for a fixed number of clusters using nonparametric mixture distributions. In contrast, our model adopts a fully parametric approach and identifies the latent structure through the C-Lasso algorithm. This algorithm has been further applied in other settings by su2018identifying and su2019sieve. The second approach is based on the K-means algorithm, commonly used in statistical analysis. lin2012estimation and sarafidis2015partially applied it to linear panel data models where the slope coefficients exhibit cluster structures. bonhomme2015grouped extended this method to linear panel data models with additive fixed effects having cluster structures and studied the asymptotic properties of the estimator. This approach was further generalized in bonhomme2022discretizing to analyze general fixed effects models. wang2021identifying provided a comparison between the C-Lasso and K-means approaches in their introduction, highlighting the relative strengths and limitations of each method.
This paper proceeds as follows. In Section (ref), we introduce the heterogeneous binary choice model with latent cluster structures and social interactions and propose the estimation algorithm which combines NPL and C-Lasso. Section (ref) establishes asymptotic properties of the proposed method. Monte Carlo experiments are conducted in Section (ref) to illustrate the finite performance in classification consistency and the quality of the estimates. Using the Add Health dataset, Section (ref) presents the empirical results on the peer effects among students across different schools in their choices toward risky behaviors. The last section concludes. Proofs are rendered in Appendices (ref) and (ref).
A social interactions model with unobserved group heterogeneity naturally resembles a panel data model. In standard panel data, the textbook default is “large cross section (usually denoted as $N$) and small time dimension (usually denoted as $T$)”. In the social interactions model, a set of $n$ individuals, for example the students in a school, forms a group, and there are $G$ such groups. Therefore, in our context $G$ is the counterpart of panel data's $N$, while $n$ is the counterpart of panel data's $T$. Let $[n]$ denote $\{1,\cdots,n\}$. For simplicity, we present the mathematical formulation as a balanced panel data structure: there are $G$ independent groups, and each group has $n$ individuals interacting with each other. An extension to unbalanced panels only requires a modification of notations.
A generic individual $i$ in the $g$th group, with a vector of observable characteristic $X_{ig}$, takes a binary action $Y_{ig} \in \{ 0,1\} $. His behavior is influenced by the social network in which he lives, denoted by the adjacent friendship matrix $F^g=\{F_{ij}^g\}_{i,j \in [n]}$. If $F_{ij}^g=1$, then $i$ is influenced by $j$, and $F_{ij}^g = 0$ otherwise. By convention, we set the self links $F_{ii}^g=0$ for $i\in[n]$. Denote $F_i^g\equiv \{j: F^g_{ij}=1\}$ as the set of $i$'s influencers (or friends), and $N_i^g=\#F_i^g$ as the number of influencers. Denote $\mathbb I_g = \{X_{ig},F_i^g\}_{i=1}^n$ as the public information set. We consider discrete games with incomplete information.
Different individuals have different numbers of peers. To measure the influence of peers' actions to $i$, we collapse it into an average $$\overline{P}_{ig0}=1\{N_i^g>0\} \cdot \frac{1}{N_i^g} \sum_{j \in F_i^g}P_{jg0},$$ where $P_{jg0}\equiv P(Y_{jg}=1|\mathbb I_g,\varepsilon_{ig})=\mathbb E(Y_{jg}|\mathbb I_g,\varepsilon_{ig})$ is the conditional belief of influencer's choice. It can be simplified to $P_{jg0}= P(Y_{jg}=1|\mathbb I_g)=\mathbb E(Y_{jg}|\mathbb I_g)$ with the independence assumption on the private information we make below. The observed binary outcome $Y_{ig}$ is generated from a random utility model of social interactions:
where in the general case the slope coefficient of the peer effect $\overline\beta_{g0}$, the slope coefficient of the individual characteristics $\beta_{g0}$, and the group fixed effect $\mu_{g0}$ can all vary across groups. In the dynamic games, aguirregabiria2007sequential investigate the unobserved group/market heterogeneity and they treat them as random effects and a part of the public information. arcidiacono2012estimating study spillover effects across groups with fixed effects on continuous outcomes. Here we take $\mu$'s as fixed effects parameters in the binary choice model. In particular, the fixed effect $\mu_{g0}$ enters the single index additively. Although standard in the panel data literature, to the best of our knowledge, no paper considers fixed effects when estimating peer effects. We provide the first formal results about conducting valid statistical inference in this setting.
Model (ref) adopts an incomplete information structure for the discrete game, where the belief about influencers' choices, i.e., the conditional expectation, enters the utility function of individuals. The random error $\varepsilon_{ig}$ is independently and identically distributed (i.i.d.) across $i\in[n]$ and $g\in[G]$, with a known distribution. In this paper, we assume $\varepsilon_{ig}$ follows a standard logistic distribution; there is no difficulty in allowing other link functions, for example, the Probit. Another straightforward extension is to include the contextual effects in to the single index, as in lin2021uncovering. We focus on the current simple model to fix ideas.
To set up the building blocks for the estimation scheme, denote the data as $w_{ig}\equiv(Y_{ig},X_{ig}')'$ and the slope coefficient vector as $\theta_g\equiv(\overline\beta_g,\beta_g')'$, with dimension $1+\mathrm{dim}(X_{ig})$. The negative log-likelihood of an individual $i$ in group $g$ is
where $\Lambda(x)\equiv(1+\rm{e}^{-x})^{-1}$ is the cumulative distribution function (CDF) of the standard logistic distribution. Let $P_{g}=\{P_{ig}\}_{i\in[n]}$ be a choice probabilities profile, and then the average negative log-likelihood for the group $g$ can be written as
Furthermore, the average negative log-likelihood over the $G$ groups is
where $\bm\mu\equiv\{\mu_g\}_{g\in[G]}$, $\bm\theta=\{\theta_g\}_{g\in[G]}$, and $\bm{P}=\{P_g\}_{g\in[G]}$ concatenate the parameters over all groups.
In empirical works using fixed effects models, most applications presume homogeneous slope coefficients to capture the commonality across groups while allowing the fixed effects to be unrestricted. We start with this convention by restricting $$\theta_1=\cdots=\theta_{G}=\theta\equiv(\overline\beta,\beta')',$$ under which the negative log-likelihood function $\Psi_{n}(\bm\mu,\bm\theta;\bm{P})$ in (ref) can be written as $\Psi_n(\bm{\mu},\theta;\bm{P})$, where a common $\theta$ is shared by all groups.
Given the model (ref), the equilibrium CCP profile $\bm{P}(\bm\mu,\theta)\equiv\{P_{ig}(\mu_g,\theta)\}_{i\in[n],g\in[G]}$ is determined as fixed points of the following equations
Let $\mathcal{A} \subset \mathbb R$ be a common parameter space for all fixed effects $\mu_g$ and $\Theta \subset \mathbb{R}^{1+\mathrm{dim}(X_{ig})} $ be the common parameter space for $\theta$ --- both are assumed to be compact. Algorithm (ref) presents a sequential estimation procedure of fixed effects and common parameters jointly via the NPL algorithm. In this algorithm, parameter estimation and CCPs update are executed sequentially, which simplifies the computation of nonparametrically estimating CCPs aguirregabiria2002swapping,aguirregabiria2007sequential,lin2017estimation,lin2024binary.
We will establish the asymptotic normality of $\widehat\theta$ in Theorem (ref)(ref). However, statistical inference based on this distributional result faces two difficulties. First, similar to fixed effects estimators in panel data literature, $\widehat\theta$ suffers an incidental parameter problem neyman1948consistent. Theoretical analyses in Appendix (ref) show that the analytic form of the bias term depends not only on third derivatives of the likelihood function but also on the derivatives of the CCPs $\bm{P}(\bm\mu,\theta)$, indicating that the analytic bias-correction is too complicated to carry out. Moreover, as the CCPs are jointly determined through the social interactions, split-sample based bias-correction methods dhaene2015split do not work. Second, direct estimation of asymptotic variance performs poorly in practice. Therefore, the literature of estimating peer effects relies on (nonparametric) bootstrap to conduct statistical inference \citep*[e.g.][]{lewbel2023social}. Such a method can only provide estimates of standard errors, while the bootstrap confidence intervals are invalid due to the asymptotic bias.
While Algorithm (ref) provides point estimates of the parameters, to conduct statistical inference about the slope coefficient $\theta$, we proceed with Algorithm (ref), a version of the parametric bootstrap method for fixed effects models proposed by higgins2024bootstrap adapted to the current setting.
The previous subsection assumes that the $\theta_{g0}$ are common for all groups, leading to the standard discrete game model without group-level slope heterogeneity. However, peers' influence may vary across different schools, for example, lewbel2023social document different levels of peer effects between large and small groups. Therefore, neglecting the slope heterogeneity is questionable for valid estimation and inference. To accommodate heterogeneity among individuals, we relax the common slope assumption by allowing $\theta_{g0}$ to follow a cluster pattern of a general form
where the latent clusters $C_{k0}$, $k\in[K_0]$, consist of a partition of groups. The number of latent clusters, $K_0$, is finite. Let $G_k=\#C_{k0}$ be the number of groups in each cluster $k$, and we posit that $G_k/G\to c_k\in (0,1)$ as $G\to \infty$, i.e., all latent clusters are non-vanishing asymptotically. In practice when $K_0$ is unknown, we will provide an information criterion (IC) in Section (ref) to determine it jointly with the cluster memberships. The latent structure here is different from the mixture data assumption considered in kasahara2009nonparametric, which assumes that choice probability is a finite mixture of base distributions.
A modified Penalized Profile Likelihood estimator is adapted from the NPL with fixed effects. Let the fixed effects estimator be $$ \widehat\mu_g(\theta_g,P_g)\equiv \underset{\mu_g\in\mathcal{A}}{\arg\min}\ \Psi_{ng}(\mu_g,\theta_g;P_g). $$ For each group $g$, define the profile negative log-likelihood
Given $G$ groups, the average negative profile log-likelihood of all groups is
Let $P_{ig}(\theta_g)$ be the fixed points that solve $$ P_{ig}(\theta_g)=\Lambda(\overline{P}_{ig}\overline\beta_g+X_{ig}'\beta_g+\widehat{\mu}_g(\theta_g,P_g)),\ i\in[n],g\in[G], $$ and denote $\bm{P}(\bm\theta)=\{P_{ig}(\theta_g)\}_{i\in[n],g\in[G]}$.
To sort the individuals into groups, we employ the C-Lasso method as a classifier. Define the NPL Classifier-Lasso (NPL-C-Lasso) estimator as
where $\rho$ is a tuning parameter to determine the level of the multiplicative classification penalty. It encourages clustering among the group-specific estimates $\{\widetilde{\theta}_g\}_{g\in[G]}$ by shrinking $\theta_g$ toward the nearest cluster estimates $\{\widetilde\vartheta_k\}_{k\in[K_0]}$. The multiplicative structure $\prod_{k=1}^{K_0} \Vert\theta_g-\vartheta_k\Vert_2$ ensures that $\theta_g$ is penalized more heavily when it is far from all cluster estimates, effectively promoting exact equality $\widetilde{\theta}_g = \widetilde\vartheta_k$ for some $k$. The optimization problem yields the estimates $\{\widetilde{\theta}_g,\widetilde{\vartheta}_k\}$ for schools $g\in[G]$ and groups $k\in[K_0]$. A group $g$ is classified into the cluster $\widetilde{C}_k$ if $\widetilde\theta_g$ is closest to $\widetilde{\vartheta}_k$ in $L_2$-distance, that is $$ \widetilde{C}_k\equiv\{g\in[G]:\underset{l\in[K_0]}{\arg\min}\ \Vert \widetilde{\theta}_g-\widetilde{\vartheta}_l\Vert_2=k\} $$ for $k\in[K_0]$. Asymptotically, this rule is equivalent to the equality rule $\widetilde{C}^e_k=\{g\in[G]:\widetilde{\theta}_g=\widetilde{\vartheta}_k\}$ as in K-means.\footnote{ In principle any consistent classification method can determine the latent structure, for example K-means. In finite sample, C-Lasso makes K-means a special case if $\rho \to \infty$. The K-means algorithm is NP-hard and relies heavily on the estimation quality of $\widehat\theta_g$s, and is thus computationally heavy. } In finite sample, C-Lasso is more lenient in its restrictions.
The C-Lasso minimization problem (ref) is non-convex and it requires estimating CCPs $\bm{P}(\bm\theta)$ in each evaluation, which is computationally demanding. To make the estimation tractable and computationally efficient, we devise a new algorithm to decouple the classification step from the estimation process to enhance computational stability. It consists of three steps, as illustrated in the diagram Figure (ref):
\noindent1. Fist Step NPL: apply the NPL method (Algorithm (ref)) to each group $g$, obtaining estimators for the CCPs $P_g$, denoted as $\widetilde{P}_g$, $g\in[G]$.
\noindent2. C-Lasso: estimate the latent clusters via solving C-Lasso:
where $\widetilde{\bm{P}}=\{\widetilde{P}_g\}_{g\in[G]}$ is independent with $\bm\theta$. This allows us to classify groups into $K_0$ distinct clusters, $\widetilde{C}_1,\dots,\widetilde{C}_{K_0}$, without sequentially updating CCPs; The number of latent clusters $\widehat{K}$ is determined by an IC, see Section (ref).
\noindent3. Post-Classification Estimation: re-estimate $\theta_k$ for each cluster $k$ through NPL with FE (Algorithm (ref)) by pooling all groups' observations within the cluster. Post-classification estimators are denoted as $\widehat{\theta}_{\widetilde{C}_k}$, $k\in[K_0]$. Apply Algorithm (ref) to groups within each cluster to obtain the confidence intervals for common parameters.
We will show that C-Lasso step classifies groups into their true clusters with probability approaching one (w.p.a.1.) in Section (ref). Therefore, inference based on post-classification estimators is asymptotically valid, similar to the homogeneous model. In the empirical application concerning risky behaviors among students reveals that, ignoring the latent structure will severely under-estimate the peer effects for the cluster with significant effects and draw wrong conclusions with the intelligence variables. Therefore, we recommend the above algorithm to cope with possible slope heterogeneity in real world applications.
In our asymptotic analysis, the group size $n$ is passed to infinity. The number of groups can stay finite, or diverge to infinity. For the latter case, we adopt a triangular limit framework, viewing $G$ as a deterministic function of $n$. In asymptotic statements, we explicitly send $n\to \infty$, understanding that $G=G(n)$ and $G(n)/n\to c\in [0,\infty)$ as $n\to \infty$.
We first setup several basic assumptions. Given these assumptions, we analyze in Section (ref) the NPL estimators $\widetilde\zeta_g\equiv(\widetilde{\mu}_g,\widetilde{\theta}_g)$ and CCPs $\widetilde{P}_g$ for each group $g$, which is the building block of our adjusted algorithm. Section (ref) establishes asymptotic normality of the NPL fixed effects estimator $\widehat\theta$ and the parametric bootstrap debiased estimator $\widehat\theta^*$. We present the classification consistency and an IC for selecting the number of clusters in Sections (ref) and (ref), respectively. Let $C<\infty$ denote a finite constant.
Assumption (ref)(ref) and (ref) guarantee a finite single index, and hence $P_{ig0}\in(c,1-c)$ for some finite constant $c\in(0,1/2)$. Although technically we can relax (ref) to allow light-tail covariate variables, such an extension complicates notations but brings no theoretical insights. Part (ref) of Assumption (ref) is called the "moderate social interactions" condition in the literature glaeser2003nonmarket,horst2006interactions,xu2018social,lin2017estimation. This condition is sufficient for the existence and uniqueness of the Bayesian Nash equilibrium. Similar conditions to Parts (ref) of Assumption (ref) are imposed in lin2021uncovering and lin2024binary. This condition ensures that the NPL estimation of unknown parameters and CCPs are well-defined and consistent for all groups. Note $$ \sum_{i\in[n]}(N_i^g)^{-1}F_{ij}^g\leq \sum_{i\in[n]}F_{ij}^g=N_j^g, $$ a sufficient condition for part (ref) is that students in each group have only a finite number of friends, which usually holds in practice. This condition makes sure that the estimation errors of $\widetilde{P}_{ig}$'s do not significantly affect the criterion function through the social interactions. Therefore, we can replace $P_0$ with estimated $\widetilde{P}$ in the C-Lasso step.
Define $\zeta_g\equiv(\mu_g,\theta_g)$ and $\zeta_{g0}\equiv(\mu_{g0},\theta_{g0})$. By applying NPL to each group $g$, we obtain the estimators $\widetilde{\zeta}_{ng}$ and $\widetilde{P}_g$. Our first theoretical result, Proposition (ref), is an extension of Proposition 2 in aguirregabiria2007sequential and Theorem 1 in lin2024binary. We focus on the consistency of $\widetilde{\zeta}_{ng}$ and the convergence rate of $\widetilde{P}_g$.
Proposition (ref)(ref) shows that $\widetilde{\zeta}_g$'s are consistent uniformly for $g\in[G]$. In the steps of C-Lasso or second NPL, we can actually start from these consistent first step estimators in the optimization. By using this optimization strategy, it is also guaranteed that the post-classification estimators (and also $\widehat\theta$) are consistent. Therefore we can focus on the asymptotic normality in next subsection. Part (ref) is a non-asymptotic result, derived from Hoeffding's inequality hoeffding1963probability. Intuitively, the NPL estimator $\widetilde{\theta}_g$ is expected to be $\sqrt{n}$-consistent as each school $g$ has $n$ students. Then by Assumption (ref), $\widetilde{P}_g$ is a continuous function of $\widetilde{\theta}_g$ around $\theta_{g0}$, and hence this $\sqrt{n}$-convergence speed of $\widetilde{P}_g$ is reasonable. This result allows directly applying C-Lasso with $P_{g0}$ replaced by $\widetilde{P}_g$, as we will see that the classification consistency just requires $\sqrt{n}$-consistent estimators for the CCPs uniformly over $G$ groups. Moreover, as $P_g$ is a function of school-specific fixed effects $\mu_g$, which means the convergence speed cannot over-perform $n^{-1/2}$ in general.
We establish asymptotic normality of $\widehat\theta$ and $\widehat\theta^*$ in this subsection. The inference based on the post-classification estimators is the same in essence. The following theorem justifies statistical inference based on $\widehat\theta$ and the parametric bootstrap estimator $\widehat\theta^*$. Let $H_{0}$ be a concentrated Hessian matrix, $B_{0}$ is the asymptotic bias term, and $\Omega_{0}$ denotes the asymptotic covariance matrix, whose explicit forms are defined by Equations (ref), (ref), and (ref), respectively, in Appendix (ref).
Theorem (ref) consists of three parts. Part (ref) shows that $\widehat{\theta}$ is asymptotically normal, but it exhibits an incidental parameter bias $H_{0}^{-1}B_{0}$ that deviates its center from the true parameter $\theta_{0}$. The proof of this theorem differs from other papers discussing incidental parameter bias in that the NPL algorithm can actually affect how the estimation error from $\widehat\mu_g$ accumulates. To remove the incidental parameter bias, we provide the parametric bootstrap inference method and show that the debiased estimator $\widehat{\theta}^*$ is asymptotically unbiased. Part (ref) justifies the unbiasedness of $\widehat\theta^*$ and the correct coverage rates of the bootstrap confidence intervals, and Part (ref) reveals the underlying mechanism. Conditioning on the realized data $\boldsymbol{W}$, the “true parameter” in the world of the parametric bootstrap is $\widehat{\theta}$. The asymptotic distribution of $\widehat\theta^b-\widehat\theta$ mimics that of $\widehat\theta-\theta_0$, sharing the same asymptotic bias. As a result, $$ \sqrt{nG}\Big(\widehat{\theta}^b-\theta_{0}\Big) = \sqrt{nG}\Big( (\widehat{\theta}^b - \widehat\theta ) - (\widehat\theta -\theta_{0})\Big) $$ is properly centered at 0 by canceling out the common bias.
In this subsection, we establish the classification consistency of C-Lasso when there is a latent structure. To evaluate the classification accuracy, we define two events $$\widetilde{E}_{knG,g}=\{g\notin \widetilde{C}_k|g\in C_{k0}\}\qquad\text{and}\qquad \widetilde{F}_{knG,g}=\{g\notin C_{k0}|g\in \widetilde{C}_k\},$$ which are analogous to Type I and Type II errors in hypothesis testing, respectively. The false exclusion $\widetilde{E}_{knG,g}$ indicates a group $g$ of its true cluster identity $C_{k0}$ is assigned into a wrong cluster, and the false inclusion $\widetilde{F}_{knG,g}$ implies that a group $g$ is mistakenly assigned to $\widetilde{C}_k$. Let $$\widetilde{E}_{knG}=\cup_{g\in C_{k0}}\widetilde{E}_{knG,g} \quad \mbox{(any false exclusion from group $k$)}$$ and $$\widetilde{F}_{knG}=\cup_{g\in \widetilde{C}_k}\widetilde{F}_{knG,g} \quad \mbox{(any false inclusion into group $k$)}, $$ Theorem (ref) establishes that, when the number of groups is correctly specified, C-Lasso ensures asymptotically perfect classification.
Theorem (ref) demonstrates that all schools are classified to their true clusters w.p.a.1. if we know $K_0$ as a prior. In the next subsection, we will propose an information criteria (IC) to select the number of clusters in practice, building on the post-classification estimators.
We re-estimate the unknown parameters via NPL using within-cluster school data, denoting the resulting post-classification estimators as $$\widehat{\theta}_{\widetilde{C}_k} =(\widehat{\overline\beta}_{\widetilde{C}_k},\widehat{\beta}_{\widetilde{C}_k}')' \quad \mbox{for } k\in[K].$$ As we mentioned before, $\widehat{\theta}_{\widetilde{C}_k}$ is useful for constructing the IC and determining $\widehat{K}$. Within each cluster, we can apply the parametric bootstrap to conduct statistical inference based on Theorem (ref).
While our preceding analysis presumes prior knowledge of $K_0$, the true group number is generally unknown in empirical applications. To address this practical consideration, we posit that $K_0$ is bounded above by a finite integer $K_{\max}$ and develop a data-driven selection procedure via IC. We have the C-lasso estimates $\{\widetilde{\theta}_g(K),\widetilde{\vartheta}_k(K)\}$ of $\{\theta_{g0},\vartheta_{k0}\}$, where we make the dependence of $\widetilde{\theta}_g$ and $\widetilde{\theta}_k$ on $K$ explicitly. As above, we classify a school $g$ into group $\widetilde{C}_k(K)$ if and only if $\widetilde{\theta}_g(K)$ is the closest to the one in $\{\widetilde{\vartheta}_k(K)\}_{k \in K_{\max}}$. We propose to select $K$ by minimizing the IC $$ \mathrm{IC}(K)\equiv \frac{1}{nG}\sum_{k\in[K]}\sum_{g\in\widetilde{C}_k(K)}\sum_{i\in[n]}\psi\Big(w_{ig},\widehat{P}_g;\widehat{\theta}_{\widetilde{C}_k(K)},\widehat{\mu}_g\big(\widehat{\theta}_{\widetilde{C}_k(K)},\widehat{P}_g\big)\Big)+\lambda pK, $$ where $\lambda$ is a tuning parameter and we recommend setting $\lambda=\frac{1}{4}\log(\log n)/n$, in accordance with \citet*[Section 2.5]{su2016identifying}. Let $$\widehat{K}\equiv\arg\min_{K\in[K_{\max}]}\mathrm{IC}(K),$$ where $K_{\max}\geq K_0$ is a pre-specified integer. The following proposition justifies the use of the IC.
Proposition (ref) shows that the true number of groups can be consistently estimated, thereby we close the loop of estimation and inference. In Figure (ref), we recap the building blocks of of estimation, model selection, and inference when $K_0$ is unknown. This scheme makes the implementation feasible with real data. \afterpage{
}
We check the finite sample performance of our estimation and inference methods in Monte Carlo experiments. For each individual $i$, we first randomly select the number of friends, $N_i^g$, from a uniform distribution over ${0}\cup[N_{\max}]$ with $N_{\max}=5$. Then we randomly formulate $N_i^g$ friendship links with other individuals in the group $g$ to formulate the friend matrix $F^g$. We consider the following DGP with network interactions through $F^g$:
where $c_{g0}+\mu_{g0}$ is a composite fixed effect, with $\mu_{g0}$ being standard normal and $c_{g0}$ taking clustering specific values $c_{k0}$. The idiosyncratic error $\varepsilon_{ig}$ is standard logistic, independent across $g$ and $i$, and independent of all regressors. The CCPs, $P_{ig0}$, are solved using fixed point algorithm on
The observations in each DGP are drawn from three groups with the proportion $$G_1:G_2:G_3 = 0.3:0.3:0.4.$$ The exogenous regressor $X_{ig}=0.1\mu_{g0}+e_{ig}$ where $e_{ig}\sim i.i.d.~ N(0,1)$. We set the number of groups $G \in \{100, 200\}$, and the group size $n \in \{50,100,200\}$. We set the true coefficients $(\overline\beta_{k0},\beta_{k0},c_{k0})$ as $(1.5,-1,-0.5),(0.75,0,-0.25)$, and $(0,1,0)$, respectively, for the three groups. The number of Monte Carlo replications is 1000 for each DGP, and 500 parametric bootstrap replications are used for the bias correction.
Table (ref) summarizes simulation results of classification and reveals two key patterns as the number of students per school ($n$) increases: (i) the probability of correctly identifying the number of latent groups converges to 1; and (ii) classification accuracy improves monotonically. The simulation scenario with $(100,100)$ mirrors the structure of our empirical dataset, which comprises 119 schools with an average of 78 students each (totaling 9,262 students). In this configuration, the C-Lasso method demonstrates robust performance: it consistently selects the true number of groups, $K_0 = 3$, while achieving approximately 93% classification accuracy. These results provide favorable evidence that our approach is suitable for the empirical application to be presented in Section (ref).
Since there is no prior literature analyzing incidental parameter bias and its correction in the NPL framework, we assess the finite sample performance using the oracle grouping. The oracle grouping ensures that any finite sample bias does not arise from misclassification. We consider median bias (Bias), root mean square error (RMSE), and 95% coverage rate (95%CR; reported only for the debiased estimator). Table (ref) presents the simulation results. The debiased estimator significantly reduces bias, lowers RMSE, and achieves good coverage—especially for $\beta_k$. It provides evidence of the efficacy of the debiased procedure. \afterpage{
}
We next evaluate the performance of our post-classification estimators, restricting attention to cases where the adjusted algorithm correctly selects $K = K_0$. We also present simulation results for estimating the unknown slope coefficients using a pooled estimator based on all observations. Simulation results in Table (ref) show that the pooled estimator, which fails to capture the heterogeneity across the groups, performs poorly. On the other hand, the post-selection estimator $\widehat\theta_{\widetilde{C}_k}$ substantially enhances the finite sample estimation quality, and the bias is further reduced by $\widehat\theta_{\widetilde{C}_k}^*$. Compared with the results in Table (ref), when the true group identity is unknown, misclassification error is inevitable in finite samples, which affects the bias and statistical inference. As the group size $n$ increases, the bias decreases rapidly, and the coverage rate of the bootstrap inference improves.
The findings of this simulation exercise corroborate the asymptotic results presented in the previous section. The algorithm is capable of identifying the group identities of the individuals, and the post-selection estimators effectively approximate the oracle performance as if the group membership were known.
In terms of the computational costs of this comprehensive simulation exercise, in total we run 18 million bootstrap replicates.\footnote{There are 18 DGPs, including 2 scenarios with 6 $(n,G)$ combinations and 3 clusters. For each DGP, we replicate 1,000 times with 500 bootstrap replications.} We parallelize the process and allocate the tasks to 128 CPU cores (AMD EPYC 9754) in a high-performance computing cluster. It takes around 20 hours to finish the simulations. The following empirical application, with no need of replication, is fast. Using the same parallel scheme, it takes fewer than 20 seconds. Therefore, the computation time is affordable for real data applications.
Many empirical studies demonstrate that risky behaviors, such as smoking, binge drinking, or unsafe sexual practices, can spread across social networks christakis2007spread,centola2010spread. Peer effects can play a key role in this process, especially among students; see, for example, duncan2005peer and eisenberg2014peer. However, the existing literature usually focuses on the homogeneous peer effects. Understanding heterogeneous peer effects can help policy makers design targeted public interventions for distinct clusters. In this section, we apply our proposed procedure to the well-known Add Health dataset to study heterogeneous peer effects in risky behaviors among students.
The Add Health dataset is a longitudinal study of a nationally representative sample of over 20,000 adolescents who were in grades 7-12 during the 1994-95 school year, and have been followed for five waves to date, most recently in 2016-18.\footnote{See detailed description in \url{https://addhealth.cpc.unc.edu/}.} It is widely used in the study of peer effects. Table (ref) provides summary statistics of the students and schools. The dataset contains 119 schools and 9,262 students. The dependent variable, risky behavior, is constructed by using the self-report questionnaires in the Add Health dataset. Specifically, the survey question is
“During the past twelve months, how often did you do something dangerous because you were dared to?”
Demographic characteristics include intelligence level (low (1$\sim$2), middle (3$\sim$4), high (5$\sim$6)), race, gender, family income, age, grades, and in-school friendship networks. There are 119 schools in total. On average, each student has only one reported friend, suggesting sparse friendship networks. The share of students with a low intelligence level is much lower than the other two intelligence levels.
We first estimate the peer effects and coefficients of covariates with homogeneous groups as a benchmark. The second column in Table (ref) presents the results. The point estimate of peer effects is $0.303$ with a 95% confidence interval $[0.083,0.515]$, where the confidence interval is estimated by the parametric bootstrap. This demonstrates that, with other characteristics controlled, the peer influence will increase the possibility of students attending risky behaviors. The influence of students' intelligence level is insignificant. The white and students with higher family incomes are relatively more inclined to participate the risky activities. The female, older and students with higher GPAs are less willing to involve in risky behaviors.
We have shown that ignoring heterogeneity results in huge bias in simulations. We apply the adjusted algorithm to identify the latent structure and estimate the peer effects with heterogeneous groups. The C-Lasso procedure classifies the 119 schools into two clusters, one containing 43 schools and the other 76 schools. Table (ref) presents the summary statistics for these two clusters, respectively. We calculate the two-sample t-statistics to examine whether there are statistically significant differences in the mean values of demographic characteristics. It is found that while risky behaviors of students in different clusters are similar, some demographic characteristics vary significantly. The first cluster has a higher proportion of white and male students. The students from the first cluster are generally older, have more friends, come from families with higher income, but have a relatively lower average GPA. Notably, there is no significant difference in intelligence level between the two clusters. Moreover, the average number of students in the first cluster (93) is higher than the second cluster (69), lewbel2023social documents the influence of the group size on peer effects.
The heterogeneous panel of Table (ref) provides the post-classification estimation results of two clusters separately. The point estimates for the two clusters exhibit distinct patterns. The most significant finding is that only the point estimate of peer effects in the second cluster, 0.688, is both statistically significant and economically meaningful. The peer effects for the first cluster are not statistically different from 0. In addition to the differences in peer effects, students with higher intelligence levels in the second cluster are more inclined to make risky choices, whereas students in the first cluster are not. Female students or those with higher grades are generally less likely to engage in risky behaviors in both clusters, which aligns the economic intuition.
If ignoring the latent structures, we may draw a conclusion that the intelligence level does not affect students' risky behaviors. However, it significantly affects both the first and the second clusters with a reversed sign. Moreover, we can obtain a more precise estimate of the peer effects for groups in the second cluster, which is helpful for making policies.
In this paper, we study the social interaction models with latent structures. We model the heterogeneity across groups via the group fixed effects and cluster-specific coefficients. By combining NPL and C-Lasso, our proposed estimation method can identify the latent structures and unknown parameters simultaneously. The finite sample performance of the method is verified by Monte Carlo experiments. We analyze heterogeneous peer effects in risky behavior using the Add Health dataset. Schools are classified into two clusters, and only one of them has statistically significant peer effects. Researchers should be aware of the latent structures when studying social interaction models.