Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,965 characters · 14 sections · 47 citation commands
Count Data Models with Heterogeneous Peer Effects under Rational Expectations
{4pt} {4pt}
{3pt} {3pt}
There is a large and growing literature on peer effects in economics.\footnote{For recent reviews, see de2017econometrics and bramoulle2020peer.} Recent contributions cover, among others, models for limited dependent variables, including binary brock2001discrete, lee2014binary, lin2024binary, ordered liu2017social, xu2018social, boucher2022peer, multinomial guerra2020multinomial, and censored xu2015maximum responses. However, count variables have received little attention, despite their prevalence in survey data (e.g., the number of physician visits, the frequency of consumption of a good or service, and the frequency of participation in activities). Peer effects on these variables are often estimated using a linear-in-means (or Tobit) model or a binary model after transforming the outcome into binary data. These approaches may lead to biased estimates because the linear-in-means (or Tobit) models are developed for (censored) continuous outcomes. Moreover, transforming the outcome to a binary format does not allow the interpretation of peer influence in terms of intensive marginal effects.
This paper develops a microfounded rational-expectations model to estimate peer effects on count outcomes. The model incorporates heterogeneity, allowing for the influence of friends to vary across agents' and friends' observed groups (e.g., gender- or ethnicity-based groups). I establish conditions under which the model has a unique Bayesian Nash equilibrium (BNE). I demonstrate that the model parameters are identified under the well-known identification condition of linear-in-means models, which relies on the presence of friends' friends who are not friends in the network bramoulle2009identification. Importantly, I generalize this identification result to a large class of nonlinear models. I apply the model to student participation in extracurricular activities using the National Longitudinal Study of Adolescent to Adult Health (Add Health) data. I find that females are more responsive to their peers than males.
The proposed model is based on a static game with incomplete information, in which agents' payoffs depend on their friends' expected outcomes. I partition agents into observed groups and assume that peer effects depend on both agents' groups and peers' groups. Unlike standard games with continuous or censored outcomes, where payoffs are typically modeled as a linear-quadratic function blume2015linear, I specify a semiparametric payoff. The assumption of a linear-quadratic payoff is widely employed in the literature, as it yields linear or Tobit specifications that are easy to estimate. I show that imposing this assumption on count outcomes, which is equivalent to estimating a linear spatial autoregressive model (SAR) or SAR-Tobit model xu2015maximum, yields an ordered model with evenly spaced thresholds or cutoff points. This assumption is restrictive and may lead to misspecification issues, particularly for count outcomes with large support. By employing a more flexible payoff function, the resulting econometric model imposes fewer restrictions on the cutoff-point distribution, while remaining estimable.
Identification in peer effects models featuring rational expectations is typically established by imposing a rank condition on the matrix of explanatory variables brock2001discrete, lee2014binary, yang2017social, xu2018social. However, this condition poses challenges for empirical testing because the explanatory variable matrix includes the average of friends' expected outcomes, which is unobservable to the econometrician. I address this problem by deriving easily testable sufficient conditions. Specifically, I show that a key identification condition is that the network includes agents who have friends' friends who are not directly their friends. While this condition has previously been used to identify parameters in linear models bramoulle2009identification, I demonstrate its applicability in nonlinear models. Once identification is established, I show that the model parameters can be estimated using the nested pseudo-likelihood (NPL) method of aguirregabiria2007sequential. The count outcome is allowed to have a large support.
I empirically investigate peer effects on the number of extracurricular activities students enroll in. First, I estimate these effects without considering heterogeneity in peer effects. The results indicate that increasing the expected number of activities in which a student's friends are enrolled by 1 implies an increase of 0.08 in the expected number of activities in which the student is enrolled. In contrast, the SAR-Tobit estimate of peer effects is more than four times higher. Subsequently, I introduce heterogeneity in peer effects by gender. I find that female students are more responsive to their peers compared to male students. Using these results, I conduct a counterfactual analysis investigating the effects of school gender composition on student participation in extracurricular activities.
This paper contributes to the large literature on social interaction models manski1993identification, brock2001discrete, lee2014binary, blume2015linear, xu2015maximum, guerra2020multinomial. I investigate the case of count outcomes using a game of incomplete information. I show that estimating peer effects on these outcomes using the linear SAR or Tobit model is equivalent to estimating an ordered model with evenly spaced thresholds, potentially leading to misspecification.
The paper contributes to the literature on the identification of peer effects models bramoulle2009identification, de2010identification,blume2015linear. Necessary conditions to avoid the reflection problem manski1993identification are readily testable for linear-in-means models. However, for nonlinear models with rational expectations, identification relies on rank conditions that pose empirical testing challenges yang2017social. I demonstrate that these models can be identified if the network includes friends' friends who are not direct friends. This identification condition is similar to the restriction for linear models.
I also contribute to the literature on the heterogeneity of peer effects. Most papers allowing heterogeneity in peer effects consider a linear-in-means model with a continuous outcome peng2019heterogeneous, beugnot2019gender, comola2022heterogeneous. The proposed model in this paper allows for the influence of friends on count outcomes to vary across both agents' and friends' observed groups.
The remainder of the paper is organized as follows. Section (ref) presents the microeconomic foundation of the model. Section (ref) addresses the identification and estimation of the model parameters. Section (ref) documents the Monte Carlo experiments. Section (ref) presents the empirical results and Section (ref) concludes this paper. Proofs of propositions are provided in Appendix (ref).
This section presents the model's microfoundations. Let $\mathcal P = \{1,\dots,n\}$ be a population of $n$ agents. The outcome of agent $i \in \mathcal P$ is denoted by $y_i$, representing an integer variable known as a count variable. Let $\mathbb{N}_{R} = \{0,1,\dots,R\}$ be the support of $y_i$, where $R$ is a finite strictly positive integer.\footnote{The assumption of a finite $R$ implies a bounded outcome, which differs from outcomes in standard count data models such as the Poisson model. I make this assumption for ease of exposition and treat the case of an infinite $R$ in Online Appendix (ref). Specifically, when $R = \infty$, many summations in the paper involve infinitely many terms and care must be taken to ensure that the resulting sums converge.}
The model is based on a game of incomplete information. To introduce heterogeneity in peer effects, I partition $\mathcal P$ into $M$ observed groups (e.g., groups defined by gender, ethnicity, or the interaction of gender and ethnicity), denoted $\mathcal G_1$, \dots, $\mathcal G_M$, where $M$ is finite. Let $g_i$ be an observable variable indicating the group of individual $i$, with support $G = \{1,\dots,M\}$. For any $g \in G$, $g_i = g$ means that $i \in \mathcal G_g$. Agents act noncooperatively and interact through a directed network, in which peer influence depends on both their own and their peers’ group memberships.
Network. The network is characterized by a set of $n \times n$ adjacency matrices, $\mathcal{A} = \{\mathbf{A}^{gg^{\prime}}: g, g^{\prime} \in G\}$. For any $g, g^{\prime} \in G$, $\mathbf{A}^{gg^{\prime}} = [a_{ij}^{gg^{\prime}}]$, where $a_{ij}^{gg^{\prime}} = 1$ if $j$ is a friend of $i$, $g_i = g$, and $g_j = g^{\prime}$, and $a_{ij}^{gg^{\prime}} = 0$ otherwise. For example, if $\mathcal G_1$ and $\mathcal G_2$ denote the groups of males and females, then $a_{ij}^{12} = 1$ if $j$ is a friend of $i$, $i$ is male, and $j$ is female. Agents do not interact with themselves; i.e., $a_{ii}^{gg^{\prime}} = 0$ for all $i\in\mathcal P$. The network is directed, thus $a_{ij}^{gg^{\prime}}$ may not equal $a_{ji}^{gg^{\prime}}$.\footnote{Later, I assume that the network consists of many independent subnetworks (e.g., schools), such that two agents from different subnetworks have no connection calvo2009peer. However, for notational simplicity, I present the microeconomic model within a single subnetwork.} I introduce heterogeneity in peer effects through agents' groups: agents are influenced by their peers depending on their own and their peers’ groups.
Individual characteristics. Every agent $i$ is described by a vector of characteristics $(\phi_i,\varepsilon_i)^{\prime}$, where $\phi_i \in \mathbb R$ is a characteristic observable to all agents and $\varepsilon_i \in \mathbb R$ is a private characteristic observable only to agent $i$. Agent $i$ observes $\varepsilon_i$ and $\phi_j$ for all $j\in\mathcal P$, but does not observe $\varepsilon_j$ for $j\ne i$. Following the standard approach in the literature, I assume that the private characteristics $\varepsilon_i$'s are independently and identically distributed (i.i.d.) and that this distribution is common knowledge to all agents brock2001discrete.
Preferences. Let $\mathbf{W}^{gg^{\prime}} = [w_{ij}^{gg^{\prime}}]$ be the matrix obtained by row-normalizing $\mathbf{A}^{gg^{\prime}}$; i.e., $w_{ij}^{gg^{\prime}} = a_{ij}^{gg^{\prime}}/\sum_{j = 1}^n a_{ij}^{gg^{\prime}}$ if $a_{ij}^{gg^{\prime}} = 1$ and $w_{ij}^{gg^{\prime}} = 0$ otherwise.\footnote{An alternative way to normalize the network matrix is to define $w_{ij}^{gg^{\prime}} = \frac{a_{ij}^{gg^{\prime}}}{\sum_{g^{\prime\prime}}\sum_{j=1}^n a_{ij}^{gg^{\prime\prime}}}$ if $a_{ij}^{gg^{\prime}} = 1$ comola2022heterogeneous. However, this approach complicates the interpretation of the model because $\mathbf{W}^{gg^{\prime}} \mathbf{y}$, where $\mathbf{y} = (y_1,\dots,y_n)^{\prime}$, no longer corresponds to the average outcome among friends in group $g^{\prime}$. Consequently, the peer effect parameter ($\alpha^{gg^{\prime}}$ below) cannot be interpreted as the average influence of peers in group $g^{\prime}$. Moreover, having $\alpha^{gg^{\prime}} > \alpha^{gg^{\prime\prime}}$ does not necessarily indicate that peers in group $g^{\prime}$ are more influential, because the influence also depends on the number of friends in each group.} Since agent $i$ does not observe others' private characteristics, they also do not observe their choices. Consequently, agents' decisions depend on their beliefs about the choices of other players. However, given that the distribution of private characteristics is common knowledge, agents form rational expectations; i.e., their beliefs are consistent with the distribution of others' private characteristics brock2001discrete, blume2015linear. Specifically, preferences are described by the following additive discrete payoff function:
where $\mathbf{y}_{-i} = \left(y_1,y_2,\dots, y_{i-1},y_{i+1},\dots,y_n\right)$ and $\boldsymbol \phi = (\phi_1,\dots,\phi_n)^{\prime}$. Payoff (ref) is separable into two components. The first two terms denote the private subpayoff, where $(\phi_i + \varepsilon_i)y_i$ is the benefit and $c_{g_i}(y_i)$ is the private cost. As in blume2015linear, the benefit is linear in the agent's own observable characteristic $\phi_i$ and private characteristic $\varepsilon_i$. In the literature, the private cost is generally defined as a quadratic function of $y_i$; i.e., $c_{g_i}(y_i) = \frac{1}{2}y_i^2$ ballester2006s, calvo2009peer, blume2015linear. This is because the game results in an estimable econometric model. Leveraging the counting nature of the outcome, I show that one can relax this restriction. I model cost as a flexible nonparametric function of $y_i$ and agents' groups. I discuss in Remark (ref) below that specifying (ref) with a quadratic private cost function implies a strong restriction on the econometric model.
The last term in payoff (ref) represents an expected social cost. The expectation is with respect to $\mathbf y_{-i}$, the choices of other players, because these choices are not observed by agent $i$. As the network matrix $\mathbf{W}^{gg^{\prime}}$ is row-normalized, $\sum_{j= 1}^n w_{ij}^{g_ig^{\prime}} y_j$ is the average outcome among friends in $\mathcal G_{g^{\prime}}$. The social cost depends on the gap between agents' choices and the average choices of their friends. The effects of this gap on the preferences can be measured by the peer effect parameter $\alpha^{g_ig^{\prime}}$. There are $M^2$ peer effect parameters: $\alpha^{gg^{\prime}}$ where $g,g^{\prime}\in G$. A higher value for $\alpha^{gg^{\prime}}$ indicates that agents in $\mathcal G_g$ are more responsive to the average choice of friends in $\mathcal G_{g^{\prime}}$. A positive $\alpha^{g_ig^{\prime}}$ suggests pure conformity between agent $i$ and friends in $\mathcal G_{g^{\prime}}$, whereas a negative $\alpha^{g_ig^{\prime}}$ implies substitutability.\footnote{This specification of social cost differs from that implying complementary preferences. The choice between conformist and complementary preferences depends on the outcome being studied. However, it can be shown that both types of preferences yield the same econometric model boucher2016some.}
For any sequence $b(r)$, I denote by $\Delta b(r)$ its first difference defined by $\Delta b(r) = b(r) - b(r-1)$. I impose the inconsequential restrictions that $c_g(-1) = +\infty$ and $c_g(R + 1) = +\infty$. To show that there exists a unique choice $y_i$ that maximizes the expected payoff (ref), I introduce the following restrictions.
The strict convexity condition in Assumption (ref) implies strictly increasing differences in costs: $\Delta c_{g_i}(r+1) - \Delta c_{g_i}(r) > 0$. This nests the commonly used linear-quadratic payoff function specification. Assumption (ref) requires the overall effects of peers across all groups to be nonnegative. Although the cost function is strictly convex, large negative peer effects can lead to a non-concave payoff, making the game difficult to solve. Assumption (ref) still accommodates scenarios where friends in some groups exert negative effects, provided these effects are offset by positive effects from other groups. Assumption (ref) restricts correlation in agents' decisions to occur only through the network and observable characteristics, a standard condition in network models li2009binary, lee2014binary, yang2017social. The concavity of the payoff and the continuity of the type distribution ensure that the payoff function is maximized at a single point almost surely (a.s.).
The condition $U_i^e\left(r\right) > \max\left\{ U_i^e\left(r- 1\right), U_i^e\left(r + 1\right)\right\}$ is also equivalent to: $$ \sum_{g^{\prime}\in G}\alpha^{g_ig^{\prime}} \boldsymbol{w}_{i}^{g_ig^{\prime}} \mathbb{E}(\mathbf{y}|\mathcal{A},\boldsymbol{\phi})+ \phi_i - \gamma_{g_i}(r+ 1) < -\varepsilon_i < \sum_{g^{\prime}\in G}\alpha^{g_ig^{\prime}} \boldsymbol{w}_{i}^{g_ig^{\prime}} \mathbb{E}(\mathbf{y}|\mathcal{A},\boldsymbol{\phi})+ \phi_{i} - \gamma_{g_i}(r),$$ where $\boldsymbol{w}_{i}^{g_ig^{\prime}} = (w_{i1}^{g_ig^{\prime}},\dots,w_{in}^{g_ig^{\prime}})$ is the $i$-th row of $\mathbf{W}^{g_ig^{\prime}}$, $\mathbf{y} = (y_1,\dots,y_n)^{\prime}$, and $\textstyle \gamma_{g_i}(r) = \Delta c_{g_i}(r) + \sum_{g^{\prime}\in G}(r - 1/2)\alpha^{g_ig^{\prime}}$. Given that $\varepsilon_i$ has a symmetric distribution, Proposition (ref) implies that the probability of $\{y_i = r\}$, conditional on $\mathcal{A}$ and $\boldsymbol{\phi}$, can be expressed as:
where $F_{\varepsilon}$ is the cumulative distribution function (cdf) of $\varepsilon_i$. Equation (ref) is similar to an ordered variable model with social interactions liu2017social. The probability $\mathbb{P}(y_i = r|\mathcal{A},\boldsymbol{\phi})$ is a function of $\boldsymbol{w}_{i}^{g_ig^{\prime}} \mathbb{E}(\mathbf{y}|\mathcal{A},\boldsymbol{\phi})$ for all $g^{\prime}\in G$, which are expectations of the average outcomes among $i$'s friends in each group, conditional on $\mathcal{A}$ and $\boldsymbol{\phi}$. The distance between the thresholds (cutoff points), $\textstyle \gamma_{g_i}(r) = \Delta c_{g_i}(r) + \sum_{g^{\prime}\in G}(r - 1/2)\alpha^{g_ig^{\prime}}$, depends closely on the cost function $c_{g_i}$. A key distinction between count data models and ordered models is that, for count data, the magnitude of the outcome matters. The values taken by the outcome influence the choice probability (ref) through the term $\mathbb{E}(\mathbf{y}|\mathcal{A},\boldsymbol{\phi})+ \phi_{i} - \gamma_{g_i}(r)\big)$, whereas in ordered models, these values are irrelevant; the focus is instead on the relative ranking of the outcomes.
Under rational expectations, the belief of any agent $j \ne i$ about $\{y_i = r\}$ equals the true conditional probability $\mathbb{P}(y_i = r | \mathcal{A}, \boldsymbol{\phi})$, defined as a function of the expected outcome in Equation (ref). The conditional expected outcome of agent $j$ can be expressed as $\mathbb{E}(y_j|\mathcal{A},\boldsymbol{\phi}) = \sum_{t=1}^R t \mathbb{P}(y_j = t|\mathcal{A},\boldsymbol{\phi})$. Substituting this into (ref) expresses $\mathbb{P}(y_i = r|\mathcal{A},\boldsymbol{\phi})$ as a function of $\mathbb{P}(y_j = t|\mathcal{A},\boldsymbol{\phi})$, where $j\ne i$ and $t = 1,\dots,R$. This yields a fixed-point equation in $\mathbf{p} =\big \{\mathbb{P}(y_i = t|\mathcal{A},\boldsymbol{\phi}): ~i\in\mathcal P, t = 1,\dots,R\big\}$, which is an $nR$-dimensional vector. This fixed-point equation may have no, one, or multiple solutions.\footnote{When $R = 1$, $\mathbb{E}(y_j|\mathcal{A},\boldsymbol{\phi}) = \mathbb{P}(y_j = 1|\mathcal{A},\boldsymbol{\phi})$ and Equation (ref) implies that $ \mathbb{P}(y_i = 1|\mathcal{A},\boldsymbol{\phi}) = F_{\varepsilon}\big(\sum_{g^{\prime}\in G}\alpha^{g_ig^{\prime}} \sum_{j = 1}^n w_{ij}^{g_ig^{\prime}} \mathbb{P}(y_j = 1|\mathcal{A},\boldsymbol{\phi})+ \phi_{i} - \gamma_{g_i}(r)\big)$, which generalizes the fixed-point equation of the binary model with a single group lee2014binary.} In this section, I establish conditions ensuring the existence of a unique belief system (outcome distribution) coherent with Equation (ref).
The fixed-point equation resulting from (ref) is defined in $\mathbb R^{nR}$. Solving it may be challenging when $R$ is large. To address this issue, I focus on the conditional expected outcome $ \mathbb{E}(\mathbf{y}|\mathcal{A},\boldsymbol{\phi})$. Equation (ref) implies that the knowledge of $\mathbb{E}(\mathbf{y}|\mathcal{A},\boldsymbol{\phi})$ is sufficient to compute the underlying belief vector $\mathbf{p}$. Moreover, since $\mathbb{E}(y_i|\mathcal{A},\boldsymbol{\phi}) = \sum_{t = 1}^R t \mathbb{P}(y_i = t|\mathcal{A},\boldsymbol{\phi})$, the knowledge of $\mathbf{p}$ is sufficient to compute $\mathbb{E}(\mathbf{y}|\mathcal{A},\boldsymbol{\phi})$. I show that the rational expected outcome also verifies a fixed-point equation.
Equation (ref) expresses agents’ expected outcomes as a function of the average expected outcomes of their peers. If $\alpha^{g_ig^{\prime}} > 0$, an agent’s expected outcome is increasing in the expected outcome of peers in $\mathcal G_{g^{\prime}}$. Conversely, $\alpha^{g_ig^{\prime}} < 0$ implies that the expected outcomes of agents decrease when those of peers in $\mathcal G_{g^{\prime}}$ increase. If $\alpha^{g_ig^{\prime}} = 0$, then agents are not responsive to peers in $\mathcal G_{g^{\prime}}$.
To demonstrate that a unique rational belief system is consistent with Equation (ref), it is sufficient to show that a unique expected outcome, $\mathbb{E}(\mathbf{y}|\mathcal{A}, \boldsymbol{\phi})$, solves (ref). To achieve this result, I show that the right-hand side of Equation (ref) forms a contraction mapping by imposing the following assumption.
The problem of multiple rational expectations equilibria generally arises when peer effects are strong. Assumption (ref) is equivalent to assuming the overall marginal effect of friends' average expected outcome on agents' expected outcome is less than one. Indeed, from Equation (ref), the marginal effect of the average expected outcome among friends in $\mathcal G_{g^{\prime}}$ on the expected outcome of agent $i$ is: $$\textstyle\dfrac{\partial \mathbb{E}(y_i|\mathcal{A},\boldsymbol{\phi})}{\partial \Bar y_i^{e,g^{\prime}}} = \alpha^{g_ig^{\prime}}\sum_{t = 1}^{R}f_{\varepsilon}\big(\sum_{\hat g\in G}\alpha^{g_i\hat g} \Bar y_i^{e,\hat g}+ \phi_{i} - \gamma_{g_i}(t)\big),$$ where $ \Bar y_i^{e,g^{\prime}} = \boldsymbol{w}_{i}^{g_ig^{\prime}} \mathbb{E}(\mathbf{y}|\mathcal{A},\boldsymbol{\phi})$. The total marginal effect of the average expected outcome among friends in all groups is thus $\sum_{g^{\prime}\in G}\alpha^{g_ig^{\prime}}\sum_{t = 1}^{R}f_{\varepsilon}\big(\sum_{\hat g\in G}\alpha^{g_i\hat g} \Bar y_i^{e,\hat g}+ \phi_{i} - \gamma_{g_i}(t)\big)$, which is less than one by Assumption (ref). The interpretation of Assumption (ref) is that agents do not increase their expected choice beyond the increase in the average expected choice of their peers, ceteris paribus. This condition is standard in peer effect models and is verified in many applications lee2014binary, blume2015linear.
The uniqueness of the BNE is a direct implication of Proposition (ref). Moreover, by the contraction mapping theorem, Assumption (ref) ensures a unique expected outcome in Equation (ref), which in turn implies a unique rational belief system $\big\{\mathbb{P}(y_i = t | \mathcal{A}, \boldsymbol{\phi}): i \in \mathcal{P}, t = 1, \dots, R\big\}$ via Equation (ref).
The key condition for the uniqueness is Assumption (ref). This assumption may be violated under strong peer effects, leading to multiple equilibria. The literature offers various solutions to address this issue when the outcome is binary and the model includes a single peer effect parameter bajari2010estimating. Most solutions require enumerating equilibria and using a selection rule for estimation. The problem becomes more complex when $R$ is large or unbounded, as the solution of (ref) may diverge, as in the case of SAR models. The econometric approach presented in the next section focuses on the case where Assumption (ref) holds and the BNE is unique. I leave the issue of multiple rational expected equilibria for future research.
In this section, I present the econometric model, study parameter identification, and propose an estimation strategy. To handle dependence across agents' strategies, I assume, as is common, that the network consists of $S$ independent subnetworks (e.g., schools) that do not interact bramoulle2009identification. This allows one to define observational equivalence at the subnetwork level (see Definition (ref)), and therefore observing many subnetworks is theoretically useful rothenberg1971identification. Let $\mathcal{P}_s$ denote the set of agents in the $s$-th subnetwork and $n_s$ the number of agents in $\mathcal{P}_s$. Asymptotically, $S$ goes to infinity, while $n_s$ is bounded.\footnote{Observational equivalence can also be defined at the sample level liu2017social, yang2017social. However, additional conditions on within-network dependence are required in this case to ensure valid inference. Assuming many independent subnetworks provides a straightforward way to handle this dependence.} Throughout the paper, vectors or matrices defined at the sample level (e.g., $\mathbf y$, $\mathcal A$) carry the subscript $s$ when restricted to subnetwork $s$. However, the notation for individual variables (e.g., $y_i$, $g_i$, $\phi_i$) will not include the subscript $s$ unless necessary for clarity.
Individual characteristics are commonly modeled as $\phi_i = \beta_0 + \boldsymbol{x}_i^{\prime}\boldsymbol{\beta}_1 + \Bar{\boldsymbol{x}}_i^{\prime}\boldsymbol{\beta}_2$ for any $i\in \mathcal{P}_s$, where $\boldsymbol{x}_i \in \mathbb R^K$ is a random vector of observed characteristics (e.g., sex and age), $\Bar{\boldsymbol{x}}_i = \sum_{j \in \mathcal{P}_s} w_{ij} \boldsymbol{x}_j$ is the average of these characteristics among peers, $\beta_0\in\mathbb R$, and $\boldsymbol{\beta}_1,\boldsymbol{\beta}_2\in\mathbb R^K$ calvo2009peer. The parameter $\boldsymbol{\beta}_1$ captures the effects of own characteristics, whereas $\boldsymbol{\beta}_2$ measures the effects of average characteristics among friends (contextual effects). For simplicity, I do not introduce group-based heterogeneity in the contextual effects. Additionally, the specification of $\phi_i$ does not account for unobserved subnetwork heterogeneity because this poses non-identification problems when $S$ is large brock2007identification.\footnote{Unless $n_s$ is sufficiently large, in which case the incidental-parameter problem does not lead to biased estimates lee2014binary, guerra2020multinomial.} Let $\mathbf{X} = [\boldsymbol{x}_1 ~ \dots ~ \boldsymbol{x}_n]^{\prime}$, $\mathbf{1}_n = (1, \dots, 1)^{\prime}\in\mathbb{R}^n$, $\mathbf{Z} = \left[\mathbf{1}_n,\mathbf{X},\mathbf{W}\mathbf{X}\right]$ and $\boldsymbol{\beta} = \left(\beta_0,\boldsymbol{\beta}_1^{\prime},\boldsymbol{\beta}_2^{\prime}\right)^{\prime}$. For any $i\in\mathcal P_s$, Equation (ref) can be expressed as: \begingroup \allowdisplaybreaks
\endgroup where $\boldsymbol{w}_{s,i}$ is the row corresponding to agent $i$ in $\mathbf W_s$ and $\boldsymbol{z}_i = (1, \boldsymbol x_i^{\prime},\boldsymbol{w}_{s,i} \mathbf X_s)^{\prime}$. Similarly, the fixed-point equation (ref) can also be written as:
where $\textstyle \gamma_g(r) = \Delta c_g(r) + \sum_{g^{\prime}\in G}(r - 1/2)\alpha^{gg^{\prime}}$, for all $g$. Thus, $\gamma_g(0) = -\infty$, $\gamma_g(R + 1) = +\infty$, and $\gamma_g(r + 1) - \gamma_g(r) > \sum_{g^{\prime}\in G}\alpha^{gg^{\prime}}$.\footnote{In Online Appendix (ref), I show that the model does not impose equidispersion, as the Poisson model does, and is more flexible than the negative binomial model due to the large number of cut points.}
The free objects to be identified in the econometric model are the peer effects $\alpha^{gg^{\prime}}: ~g,g^{\prime}\in G$, the control variable parameter $\boldsymbol\beta$, the thresholds $\gamma_g(t):~g\in G, t = 1,\dots, R$, and the distribution function of agents' private characteristic, $F_{\varepsilon}$. In this section, I assume that $F_{\varepsilon}$ is known and present a general analysis that addresses the identification of $F_{\varepsilon}$ in Online Appendix (ref). As the intercept $\beta_0$ and the thresholds enter Equations (ref) only through their difference, they cannot be identified separately. I thus fix the first threshold, as in ordered models: $\gamma_g(1) = 0$ for all $g\in G$.
Let $\boldsymbol{\alpha} = (\alpha^{gg^{\prime}}: ~g,g^{\prime}\in G)^{\prime}$ be the vector of peer effects, $\boldsymbol{\gamma} = (\gamma_g(t):~g\in G, t = 2,\dots,R)^{\prime}$ be the vector of thresholds to be estimated, and $\boldsymbol\theta = (\boldsymbol{\alpha}^{\prime}, \boldsymbol{\beta}^{\prime},\boldsymbol{\gamma}^{\prime})^{\prime}$ be the vector that comprises all parameters. The parameter $\boldsymbol\theta$ is identified if any $\tilde{\boldsymbol{\theta}} \ne \boldsymbol\theta$ is not observationally equivalent to $\boldsymbol\theta$. Observational equivalence is defined as follows.
Let $\bar{\mathbf{Y}}_s = [\mathbf{W}^{gg^{\prime}}_s\mathbf y_s: ~g,g^{\prime}\in G]$ be the matrix of average peers' outcomes. For all $i\in\mathcal{P}_s$, let $\bar{\mathcal{Y}}_i$ be the row of agent $i$ in $\bar{\mathbf{Y}}_s$ and $\tilde{\boldsymbol z}_i = [\mathbb E(\bar{\mathcal{Y}}_i|\mathcal{A}_s,\mathbf Z_s),\boldsymbol{z}_i^{\prime}]^{\prime}$. The condition $p(\mathbf y_s|\mathcal{A}_s,\mathbf Z_s) = \tilde p(\mathbf y_s|\mathcal{A}_s,\mathbf Z_s)$ implies that both distributions $p(\mathbf y_s|\mathcal{A}_s,\mathbf Z_s)$ and $ \tilde p(\mathbf y_s|\mathcal{A}_s,\mathbf Z_s)$ yield the same conditional expected outcome $\mathbb E(\mathbf y_s|\mathcal{A}_s,\mathbf Z_s)$. Consequently, for $\tilde{\boldsymbol{\theta}} \ne \boldsymbol\theta$ not to be observationally equivalent to $\boldsymbol\theta$, it is necessary that the matrix of explanatory variables, $[\mathbb E(\bar{\mathbf{Y}}|\mathcal{A},\mathbf Z),\mathbf Z]$, be full rank. Specifically, I impose the following identification restrictions.
Assumption (ref) is classical in peer effect models with rational expectations brock2007identification, yang2017social,xu2018social. However, since $\mathbb E(\bar{\mathbf{Y}}_s|\mathcal{A}_s,\mathbf Z_s)$ is not observable to the econometrician, testing Assumption (ref) in practice may pose challenges. To address this issue, I establish sufficient conditions for this assumption in Proposition (ref) below. Assumption (ref) ensures that there is a positive probability that the outcome takes any value in $\mathbb {N}_R$ manski1988identification. Recall that Assumption (ref) imposes that $\varepsilon_i$'s are independent of $\{\mathcal{A}, \boldsymbol \phi, g_1, \dots,g_n\}$, which means that $\varepsilon_i$ and $\mathbf X_s$ are independent. This suggests that there is no omission of important regressors in $\mathbf{X}_s$ that can be captured by $\varepsilon_i$.
As Assumption (ref) is not easily testable, I find sufficient conditions for the cases where $M \leq 2$. These conditions do not involve the unobserved expected outcome and can then be easily tested in practice. Specifically, I demonstrate that the presence of friends' friends who are not directly friends in the network can help identify the model’s parameters. This argument has so far been used for addressing the reflection problem in linear models bramoulle2009identification. I show that it can be applied to a large class of models, where a rank condition, such as Assumption (ref), ensures identification.
Let $\boldsymbol\pi_{i} = (\pi_{i}^{1},\dots,\pi_{i}^{M})^{\prime}$ be an $M$-vector, where $\pi_{i}^{g}$ is a dummy variable that equals one if the following two conditions hold: (1) agent $i$ has friends in group $\mathcal G_{g}$; (2) agent $i$ has friends' friends (in any group) who are not directly their friends. For models without heterogeneity in the peer effects ($M = 1$), $\boldsymbol\pi_{i}$ is simply a dummy variable that equals one if $i$ has friends' friends who are not friends.
Proposition (ref) is established by assuming that the peer effect parameters $\alpha^{gg^{\prime}}$ are all nonnegative. I will later discuss cases of negative peer effects. Condition (ref) imposes that the matrix of observed regressors is of full rank. Condition (ref) implies that the network includes agents who have friends' friends who are not their friends. Moreover, when $M = 2$, certain of these agents should not have friends in all groups. For example, for gender-based groups, some males who have friends of friends who are not their own friends should have only male or only female friends. Condition (ref) ensures that the expected outcome is influenced by the contextual variable $\bar x_{i,\bar \kappa}$. The direct effect of changes in friends' exogenous characteristics on an agent’s outcome is given by $\beta_{2,\bar \kappa}$, while the effect on friends’ outcomes is $\beta_{1,\bar \kappa}$. When these parameters share the same sign, the resulting indirect effects through peer interactions reinforce the direct effect, yielding a nonzero total effect on the agent’s outcome.\footnote{A similar restriction is also set by bramoulle2009identification in linear models with a single peer effect parameter $\alpha$. In their case, the condition is $\alpha \beta_{1,\bar \kappa} + \beta_{2,\bar \kappa} \ne 0$.}
Formal proof of Proposition (ref) is presented in Appendix (ref). Below, I provide an intuition for the case of a single group. I consider a network of three agents $i_1$, $i_2$, and $i_3$, where $i_3$ is a friend's friend of $i_1$ but not a direct friend (see Figure (ref)). If $\mathbb E(\bar{\mathbf{Y}}_s|\mathcal{A}_s,\mathbf Z_s)$ and $\mathbf{Z}_s$ are not linearly independent, then for all $i$ in any subpopulation $\mathcal P_s$, I can write:
where $\check\beta_0\in\mathbb R$, $\check{\boldsymbol{\beta}}_1\in\mathbb R^K$, and $\check{\boldsymbol{\beta}}_2\in\mathbb R^K$ are constants and $\bar{\mathcal{Y}}_i = \boldsymbol w_{s,i} \mathbf{y}_s$ is the average outcome among friends, with $\boldsymbol w_{s,i}$ being the row corresponding to agent $i$ in $\mathbf W_s$. As $i_2$ is the only friend of $i_1$, $\bar{\mathcal{Y}}_{i_1} = y_{i_2}$ and $\Bar{\boldsymbol{x}}_{i_1} = \boldsymbol{x}_{i_2}$. For $i = i_1$, Equation (ref) becomes $\mathbb E(y_{i_2}|\mathcal{A}_s,\mathbf Z_s) = \check\beta_0 + \boldsymbol{x}_{i_1}^{\prime}\check{\boldsymbol{\beta}}_1 + \boldsymbol{x}_{i_2}^{\prime}\check{\boldsymbol{\beta}}_2$, which means that the expected outcome of $i_2$ is not influenced by $x_{i_3,\bar \kappa}$. This condition cannot hold because $i_3$ is a friend of $i_2$. As implied by Condition (ref), the marginal effect of $x_{i_3,\bar \kappa}$ on $\mathbb E(y_{i_2}|\mathcal{A}_s,\mathbf Z_s)$ cannot be zero. As a result, Equation (ref) cannot hold for all agents when many subnetworks include agents who have friends' friends who are not directly friends.
The positive peer effects assumed in Proposition (ref) ensure that the overall effect of $x_{i_3,\bar \kappa}$ on $\mathbb{E}(y_{i_2} | \mathcal{A}_s, \mathbf{Z}_s)$ is nonzero. Increasing $x_{i_3,\bar \kappa}$ generates direct contextual effects since $i_3$ is a friend of $i_2$. However, it may also produce indirect effects through $\mathbb{E}(y_{i_3} | \mathcal{A}_s, \mathbf{Z}_s)$ via peer interactions. If peer effects are negative, direct and indirect effects may have opposite signs, but the overall effect is unlikely to cancel out exactly. Thus, Proposition (ref) extends to negative peer effects as long as direct and indirect effects do not perfectly offset. With many subnetworks, the perfect offset of both effects is highly improbable.
This section presents an approach for estimating the model parameters. Given that the rational expected outcome depends on the cdf $F_{\varepsilon}$, estimating the parameters without specifying this cdf is challenging brock2001discrete, brock2002multinomial, lee2014binary, yang2017social, guerra2020multinomial.\footnote{An exception where semi-parametric estimation without specifying $F_{\varepsilon}$ is possible is when there is a finite number of agents (e.g., 2 or 3 firms) that are observed multiple times. In this case, the asymptotic analysis focuses on the number of times the game is repeated rather than the number of agents aradillas2010semiparametric, wan2014semiparametric.} I assume that $\varepsilon_i$ follows a standard normal distribution. The distribution's variance is set to 1 as an identification restriction, as is the case for ordered-response models.
Given $\boldsymbol{\theta}$, Equations (ref) and (ref) can be employed to compute $\mathbb E(\mathbf y_s|\mathcal{A}_s,\mathbf Z_s)$ and $\mathbb{P}(y_i = r|\mathcal{A}_s,\mathbf Z_s)$ for all $r\in\mathbb N_R$. This suggests using the maximum likelihood (ML) approach to estimate $\boldsymbol{\theta}$. However, the estimation approach may be computationally cumbersome for large samples, because one needs to solve the fixed-point equation (ref) in $\mathbb R^n$ for every guess of $\boldsymbol{\theta}$. To address this issue, I rely on the nested pseudo-likelihood (NPL) proposed by aguirregabiria2007sequential. This method eliminates the need to solve a fixed-point problem. For any $\boldsymbol u \in \mathbb R^n$, let us consider the pseudo-likelihood function defined as:
where $ p_{it}(\boldsymbol{\theta},\boldsymbol u_{s(i)}) = F_{\varepsilon}\big(\sum_{g^{\prime}\in G}\alpha^{g_ig^{\prime}} \boldsymbol{w}_{s,i}^{g_ig^{\prime}} \boldsymbol u_{s(i)}+ \boldsymbol{z}_i^{\prime}\boldsymbol{\beta} - \gamma_{g_i}(t)\big) - F_{\varepsilon}\big(\sum_{g^{\prime}\in G}\alpha^{g_ig^{\prime}} \boldsymbol{w}_{s,i}^{g_ig^{\prime}} \boldsymbol u_{s(i)}+ \boldsymbol{z}_i^{\prime}\boldsymbol{\beta} - \gamma_{g_i}(t+ 1)\big)$, $\boldsymbol{u}_{s(i)}$ denotes the vector containing the entries of $\boldsymbol{u}$ associated with the agents in the same subnetwork as $i$, and $\mathbbm{1}\{\cdot\}$ denotes the indicator function. The pseudo-likelihood function $\mathcal{L}(\boldsymbol{\theta},\boldsymbol u)$ depends on an arbitrary vector $\boldsymbol u$ and not on the true expected outcome $\mathbb E(\mathbf y_s|\mathcal{A}_s,\mathbf Z)$.
Let $\ell_i(\boldsymbol{\theta},\boldsymbol u_{s(i)}) = \sum_{t = 1}^{R}F_{\varepsilon}\big(\sum_{g^{\prime}\in G}\alpha^{g_ig^{\prime}} \boldsymbol{w}_{s,i}^{g_ig^{\prime}} \boldsymbol u_{s(i)}+ \boldsymbol{z}_i^{\prime}\boldsymbol{\beta} - \gamma_{g_i}(r)\big)$ and $\mathbf L(\boldsymbol{\theta}, \boldsymbol u) = (\ell_1(\boldsymbol{\theta}, \boldsymbol u_{s(1)}),\break \dots, \ell_n(\boldsymbol{\theta}, \boldsymbol u_{s(n)}))^{\prime}$. To describe the NPL algorithm, I also define the operators $\textstyle \tilde{\boldsymbol{\theta}}(\boldsymbol u) = \arg \max_{\boldsymbol{\theta}} \mathcal{L}(\boldsymbol{\theta},\boldsymbol u)$ and $\boldsymbol{\phi}(\boldsymbol u) = \mathbf{L}(\tilde{\boldsymbol{\theta}}(\boldsymbol u), \boldsymbol u)$. The NPL algorithm starts with a guess $\boldsymbol u^{0}$ for $\boldsymbol u$ and constructs the sequence of estimators $\mathcal{Q}_k = \{\boldsymbol{\theta}^{k}, \boldsymbol u^{k}\}$, for $k = 1,2,\dots$, where $\boldsymbol{\theta}^{k} = \tilde{\boldsymbol{\theta}}(\boldsymbol u^{k-1})$ and $\boldsymbol u^{k} = \boldsymbol{\phi}(\boldsymbol u^{k-1})$. Specifically, given the guess $\boldsymbol u^{0}$, $\boldsymbol{\theta}^{1} = \tilde{\boldsymbol{\theta}}(\boldsymbol u^{0})$ and $\boldsymbol u^{1} = \boldsymbol{\phi}(\boldsymbol u^{0})$, then $\boldsymbol{\theta}^{2} = \tilde{\boldsymbol{\theta}}(\boldsymbol u^{1})$ and $\boldsymbol u^{2} = \boldsymbol{\phi}(\boldsymbol u^{1})$, and so forth. Note that each value of $\mathcal{Q}_k$ requires evaluating the mapping $\mathbf{L}$ only once. If $\mathcal{Q}_{k}$ converges, regardless of the initial guess $\boldsymbol u^0$, its limit, which is denoted by $\{\hat{\boldsymbol{\theta}},\hat{\mathbb{E}}(\mathbf{y}|\mathcal{A},\mathbf Z)\}$, is the NPL estimator. This limit satisfies the following two properties: $\hat{\boldsymbol{\theta}}$ maximizes the pseudo-likelihood $\mathcal{L}(\boldsymbol{\theta},\hat{\mathbb{E}}(\mathbf{y}|\mathcal{A},\mathbf Z))$; and $\hat{\mathbb{E}}(\mathbf{y}|\mathcal{A},\mathbf Z) = \mathbf{L}(\hat{\boldsymbol{\theta}},\hat{\mathbb{E}}(\mathbf{y}|\mathcal{A},\mathbf Z))$.
The NPL algorithm is based on an iterative process to determine both the values of $\boldsymbol{\theta}$ and $\boldsymbol u$ that maximize (ref). The resulting estimator $\{\hat{\boldsymbol{\theta}},\hat{\mathbb{E}}(\mathbf{y}|\mathcal{A},\mathbf Z)\}$ is such that $\hat{\boldsymbol{\theta}}$ maximizes the pseudo-likelihood $\mathcal{L}(\boldsymbol{\theta},\hat{\mathbb{E}}(\mathbf{y}|\mathcal{A},\mathbf Z))$; and $\hat{\mathbb{E}}(\mathbf{y}|\mathcal{A},\mathbf Z) = \mathbf{L}(\hat{\boldsymbol{\theta}},\hat{\mathbb{E}}(\mathbf{y}|\mathcal{A},\mathbf Z))$. As demonstrated by kasahara2012sequential, a key determinant of the convergence of the NPL algorithm is the contraction property of the fixed point mapping defined by the right-hand side of Equation (ref). Specifically, this mapping is defined for all $\boldsymbol u = (u_1, \dots, u_n)^{\prime} \in \mathbb{R}^{n}$ as $\mathbf{L}^{\ast}(\boldsymbol u) = (\ell_{1}^{\ast}(\boldsymbol u_{s(1)}),\dots,\ell_{n}^{\ast}(\boldsymbol u_{s(n)}))^{\prime}$, with $\textstyle\ell_{i}^{\ast}(\boldsymbol u_{s(i)}) = \sum_{t = 1}^{R}F_{\varepsilon}\big(\sum_{g^{\prime}\in G}\alpha^{g_ig^{\prime}} \boldsymbol{w}_{s,i}^{g_ig^{\prime}} \boldsymbol u_{s(i)}+ \phi_{i} - \gamma_{g_i}(t)\big)$. As shown in Proposition (ref), this mapping is contracting under Assumption (ref), which guarantees that it has a unique fixed point.
The following proposition establishes the consistency and asymptotic distribution of $\hat{\boldsymbol{\theta}}_n$.
In this section, I conduct a Monte Carlo study to assess the finite sample performance of the model and the NPL estimator. I consider four data-generating processes (DGPs), denoted as DGPs A--D. For all DGPs, I assume that $\phi_i = 2 + 1.5 x_{1i} - 1.2 x_{2i} + 0.5 \bar{x}_{1i} - 0.9 \bar{x}_{2i}$, where $(x_{1i},x_{2i}) = \boldsymbol{x}_i'$ and $(\bar{x}_{1i},\bar{x}_{2i}) = \bar{\boldsymbol{x}}_i'$. I simulate $x_{1i}$ from a Uniform distribution over $[0,5]$ and $x_{2i}$ from a Poisson distribution with parameter 2. Following the network structure of my empirical application, I assume that each agent $i$ has $\bar n_i$ randomly selected friends, where $\bar n_i$ is chosen randomly between 0 and 10.
I set $R = 100$, although most observed $y_i$ values are relatively low. Assuming a large $R$ allows for replicating outcomes with distributions exhibiting long tails, similar to those observed in survey data. The DGPs differ in how peer effects and cutoff points (cost function) are specified. For DGPs A and B, there is a single peer effect parameter $\alpha = 0.25$ (no heterogeneity in the peer effects). DGP A considers a quadratic cost function, implying that the cutoff points are equally spaced. I set $\gamma(r+1)-\gamma(r) = 0.55$ for $r \geq 1$. For DGP B, the cost function is semiparametric. I assume that $(\gamma(2), \dots, \gamma(13)) = (2.050,1.250, 0.850,0.700,0.500,0.400,0.330,0.300,0.290, ~0.280, 0.270,0.260)$ and $\gamma(r+1)-\gamma(r) = 0.255$ for $r \geq 13$, indicating that the cost is quadratic when $r \geq 13$.
In DGPs C--D, I introduce heterogeneity in peer effects. I consider two groups: $G = \{1,2\}$, where $g_i = 1$ if $x_{1i} \leq 2.5$ and $g_i = 2$ otherwise. The peer effect parameters are defined as $\boldsymbol\alpha = (0.3,0.15,0.1,0.15)^{\prime}$ for DGP C and $\boldsymbol\alpha = (0.4,-0.1,0.2, ~0.1)$ for DGP D. In DGP D, peers in $\mathcal G_2$ have negative effects on agents in $\mathcal G_1$. In both $\mathcal G_1$ and $\mathcal G_2$, I define the cost function as in DGP B.
The parameter values are motivated by my empirical application so that the resulting outcome distributions resemble those observed in the application. Yet, the distribution is quite different for DGP A because the cost function in my application is likely non-quadratic. Figure (ref) displays examples of histograms of simulated data for $S=8$ subpopulations, where each subpopulation comprises $n_s = 250$ agents ($n=\text{2,000}$). DGP A exhibits a symmetric distribution left-truncated at zero, whereas the distributions of the other DGPs exhibit long tails.
Table (ref) summarizes the simulation results for 1,000 replications. The number of subpopulations $S \in \{2, 8\}$ and each subpopulation comprises $n_s = 250$ agents; i.e., $n \in \{\text{500, 2,000}\}$. Since the values of the parameters cannot be directly interpreted, I focus on the average direct marginal effects (DMEs).\footnote{In peer effects models, changes in exogenous characteristics directly affect the expected outcome while keeping the game equilibrium fixed, as shown in Equation (ref). This is referred to as the direct marginal effect (DME). However, changes in exogenous characteristics also influence the contextual variable (since friends’ characteristics change) and the game’s equilibrium. Consequently, the expected outcome undergoes additional changes, referred to as indirect marginal effects (IME). In this section, I focus only on DMEs, while both effects are presented for the empirical application in Section (ref).} For DGPs without heterogeneity in peer effects, I compare the estimated direct marginal effects from my model to those provided by a SAR-Tobit model. I estimate the models using a grid of values starting from $\bar R_c = 1$ and select the optimal $\bar R_c$ that minimizes the BIC (see Remark (ref)).
For DGP A, both semiparametric and quadratic cost estimators perform well. This is because the cost is quadratic in the data. Additionally, the estimated marginal effects from a SAR-Tobit model are similar to those from the count data model. This result aligns with the discussion in Remark (ref), which emphasizes that estimating peer effects on count data using a SAR-Tobit model is similar to estimating an ordered model with equally spaced thresholds.
In the case of DGP B, the quadratic-cost and SAR-Tobit approaches yield similar marginal effects. Yet, these effects are biased upward by an average of 40% for the peer effect parameters. This bias occurs because the cost is not quadratic for this DGP. In contrast, the semiparametric cost estimates are accurate. Increasing the sample size from 500 to 2,000 improves the precision of the marginal effect estimates, bringing them closer to the true marginal effects.
Furthermore, the count data model with the semiparametric cost approach performs well when peer effects are heterogeneous. The estimated marginal effects are valid for DGPs C and D. Conversely, the quadratic cost approach biases the marginal effects for both peer effects and control variables. The bias is positive for peer effects on agents in $\mathcal G_1$ and negative for peer effects on agents in $\mathcal G_2$. This occurs because agents in $\mathcal G_1$ have a lower $x_{1i}$ and, thus, a lower expected outcome. The quadratic cost approach tends to overestimate peer effects on agents with lower expected outcomes and to underestimate them on agents with higher expected outcomes.
In this section, I present an empirical illustration of the model. I estimate peer effects on the number of extracurricular activities students are enrolled in using the National Longitudinal Study of Adolescent to Adult Health (Add Health) data.
The Add Health data provide nationally representative information on \nth{7}--\nth{12} graders in the United States (US). I use the Wave I in-school data, which were collected between September 1994 and April 1995. The surveyed sample comprises 80 high schools and 52 middle schools. The data provide information on students' social and demographic characteristics, including their friendship links (best friends, up to 5 females and up to 5 males), education level, and parents' occupation, etc. The network is restricted at the school level. Two students from different schools are not friends.
After cleaning the dataset by removing observations with missing values, the sample used in this empirical study comprises $n = \text{72,291}$ students from $S = 120$ schools.\footnote{Many of the nominated friend identifiers in the dataset are missing. Numerous papers have developed methods for estimating peer effects using partial network data boucher2020estimating. To focus on the main purpose of this paper, I do not address that issue here.} The largest school has 2,156 students, and about 50% of the schools have more than 500 students. The average number of friends per student is 3.8 (1.8 male and 2.0 female friends).
The studied count variable is the number of extracurricular activities in which students are enrolled. Students were presented with a list of 33 clubs, organizations, and teams that are found in many schools. The students were asked to identify any of these activities in which they participated during the current school year or in which they planned to participate later in the school year ($R = 33$).\footnote{Since the cut points are group-dependent, estimating an ordered model here would require estimating $64$ additional parameters. Moreover, some cut points cannot be estimated because fewer than 1% of observations exceed 10. In such a case, the semiparametric cost function is a suitable alternative.} I study the influence of social interactions on the number of extracurricular activities identified by the students. Figure (ref) displays a histogram of the number of extracurricular activities, which varies from 0 to 33. The distribution has a long tail as in DGPs B--D in the simulation study, and more than 99% of observed values fall below 10. This feature suggests that the thresholds would not be evenly spaced. Therefore, peer effect estimates on these data using the SAR-Tobit model would be biased, as is the case for DGPs B, C, and D in the simulation study.
I introduce gender-based heterogeneity in the peer effects. Specifically, I estimate the influence of both male and female friends separately for male and female students. On average, students participate in 2.4 extracurricular activities, with females participating in 2.5 and males in 2.2. I control for several demographic and socioeconomic factors and contextual variables associated with these factors (see Table (ref) in Online Appendix (ref) for a data summary).
Table (ref) summarizes the estimated average marginal effects.\footnote{Comprehensive results, including all parameter estimates, are provided in Online Appendix (ref).} Models 1--6 do not account for heterogeneity in peer effects. Model 1 is based on a semiparametric cost function, Model 2 imposes a quadratic cost function, and Model 3 is a SAR-Tobit model. Models 4--6 correspond to the school fixed-effect versions of Models 1--3, respectively. These fixed effects are included as dummy variables to control for unobserved school characteristics, such as pupil-teacher ratio and school climate.\footnote{Given the large sample size ($n = 72{,}291$) and relatively few schools $(S = 120)$, including school fixed effects as dummy variables does not raise an incidental parameter problem lee2014binary, guerra2020multinomial.} Models 7–8 correspond to count data models with gender-dependent peer effects and a semiparametric cost function. Model 7 assumes a semiparametric cost function, whereas Model 8 considers a quadratic cost function.
The estimation results for Model 1 indicate that a one-unit increase in the expected number of activities in which a student’s friends are enrolled increases the expected number of activities in which the student is enrolled by 0.082. However, the quadratic cost approach (Model 2) overestimates the marginal peer effect at 0.543. As in the simulation study, this model leads to the same estimates as the SAR-Tobit approach.
Controlling for school fixed effects does not significantly change the peer effect estimate for the nonparametric cost approach (Model 4). Yet, the increase in the log-likelihood (for 119 additional regressors) indicates that unobserved school characteristics influence student participation. In contrast, controlling for school-fixed effects partially addresses the bias of the quadratic cost and SAR-Tobit approaches (Models 5 and 6). The marginal peer effects decrease by more than 30%, but remain 4.6 times higher than the estimate from the nonparametric approach.\footnote{To assess model performance in replicating observed data, I generate extracurricular activities using the estimated coefficients as parameter values (as in the Monte Carlo simulations). The quadratic cost function (Model 5) fails to match the observed distribution, whereas predictions based on semiparametric cost functions provide a closer fit (see Online Appendix (ref)).}
Furthermore, the estimation results reveal gender-based heterogeneity in peer effects, with same-sex students being less responsive to each other. In Model 7, the marginal effects of female peers on males are $0.071$, similar to the marginal effects of male peers on females ($0.069$). However, the marginal effects of female peers on females are estimated at $0.034$, while those of male peers on males are $-0.022$. The negative effects of male peers on males are perhaps surprising and suggest a decrease in males' participation when the participation of their male friends increases, holding the participation of female friends fixed. Moreover, the results show that females are more responsive and also exert a stronger influence on their friends. Indeed, the total marginal peer effects (from both males and females) on females are $0.091$, but are only half of those ($0.044$) on males. Similarly, the peer effects of females on their friends (males and females) are $0.096$, while those of males on their friends are $0.039$.
In Model 8, the quadratic cost specification overestimates the peer effects in each group. The total marginal peer effects (from both males and females) on males are $0.187$, which is on average twice as large as the effects estimated using the semiparametric cost specification. Additionally, the total marginal peer effects (from both males and females) on females are $0.116$, which is more than twice the effects estimated using the semiparametric cost specification.
In this section, I conduct a counterfactual analysis to examine the effects of the proportion of male students on participation in extracurricular activities. Genders are randomly assigned to students to achieve target proportions of males. For each proportion, I compute the expected number of extracurricular activities by solving the fixed-point equation (ref), while keeping other variables and the network constant. The parameters are fixed at their estimated values. Figure (ref) presents the average number of extracurricular activities as a function of the proportion of male students. The simulations are based on specifications that control for unobserved school heterogeneity, and the confidence intervals are obtained using the Delta method.
In Panel (A), the predictions are based on a model with a semiparametric cost function (Model 4), under the assumption of no social interactions ($a_{ij}^{gg^{\prime}} = 0$ for all $i$, $j$, $g$ and $g^{\prime}$). As the proportion of male students varies from $0\%$ to $100\%$, predicted participation declines from 2.330 to 2.056 activities. This decrease arises because male students participate in fewer activities than female students. Since there are no interactions in this scenario, there are no social multiplier effects. For example, when the proportion of males is $50\%$, which corresponds to the observed data, the average predicted outcome is 2.193 activities---lower than the observed sample mean of 2.353.
Panel (B) considers the model with a quadratic cost function but without heterogeneity in peer effects (Model 5). The results show that varying the proportion of males from $0\%$ to $100\%$ leads to a decline in participation from 2.670 to 2.364 activities. The larger decrease relative to Panel (A) is driven by social multiplier effects. Indeed, increasing the number of females raises participation through both direct effects of students' sex and indirect peer effects. However, these estimates are potentially biased because the quadratic cost specification tends to overestimate social multiplier effects. For example, when the proportion of males is $50\%$, the average predicted outcome exceeds the observed sample mean. Panel (C) is based on the model with a semiparametric cost function and without heterogeneity in peer effects (Model 4), thereby addressing the concern of biased estimates raised for Panel (B). The results show that varying the proportion of males from $0\%$ to $100\%$ leads to a decrease in participation from 2.526 to 2.205 activities.
Panel (D) incorporates heterogeneous peer effects (Model 7). Owing to this heterogeneity, the average expected outcome is not a monotone function of the proportion of males. The highest participation level is observed in coed schools with $32\%$ male representation, where students attend 2.431 extracurricular activities on average. Interestingly, although male presence generally reduces overall participation, the highest participation level is not observed in all-female schools. This arises because male students generate social multiplier effects by exerting stronger influence on female peers. However, when the proportion of males becomes too large, these effects are offset, as males tend to participate in fewer activities than females.
This paper develops a peer effect model for count data using a static game of incomplete information. I introduce heterogeneity in peer effects through agents' groups (e.g., gender-based groups). This allows for peer effects to vary across different peer and agent groups. Moreover, the counting nature of the outcome offers flexibility in specifying payoffs. Unlike the conventional linear-quadratic payoff that is imposed in linear-in-means models, I opt for a semiparametric payoff. I demonstrate that restricting to linear-quadratic payoffs leads to inconsistent estimators of peer effects on count outcomes. This result suggests that estimating peer effects on count outcomes using the linear-in-means SAR or Tobit model would be inconsistent, as these models assume linear-quadratic payoffs.
Furthermore, I provide a general identification analysis. Specifically, I demonstrate that the argument supporting the use of friends' friends who are not directly friends to identify linear-in-means models extends to a large class of nonlinear models.
The bias in peer effects estimated using the SAR-Tobit model is illustrated through a Monte Carlo study. In my empirical application, I find that the estimates are four times as high as those provided by the proposed approach. This result underscores a significant concern regarding the assumption of a linear-quadratic payoff. This assumption may be violated even for continuous outcomes, potentially leading to biased estimates of peer effects. The linear-quadratic payoff approach is adopted in the literature because it yields estimable econometric models. Introducing flexibility into the payoff structure can be challenging for continuous outcomes. The resulting econometric specification becomes semiparametric, thereby posing significant challenges for identification and estimation. Addressing these challenges is an active area of research in the current literature.