Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,787 characters · 17 sections · 73 citation commands
Heterogeneity in peer effects for binary outcomes
\spacingset{1.5}
With the growing recognition that economic behaviors are embedded in social interactions and the increasing availability of network data, the identification and estimation of peer effects have generated a rich theoretical and empirical literature zenou2025,Bramoulle2019. Theoretical models of social interactions\footnote{I use the terms “peer effects’’ and “social interactions’’ interchangeably.} typically allow for two-way heterogeneous peer effect coefficients ballester2006, meaning that the influence of individual $j$ on $i$ may differ from her influence on $k$, or from $k$’s influence on $i$. While such flexibility is conceptually appealing, econometric implementation is infeasible because estimating $n(n-1)$ parameters from $n$ behavioral equations exceeds what observed data can support. Consequently, most empirical work has relied on the more tractable linear-in-means model, in which agents react to the average behavior of their peers, and peer effects are constrained to be homogeneous. \\
This paper focuses on binary outcomes and peer effects that arise from a preference for conformity, a setting in which individuals incur a cost when deviating from the average behavior of their friends.\footnote{Equivalently, the cost may depend on a weighted sum of pairwise distances between the agent’s action and her friends’ actions, as in Boucher2016.} In such a setting, the assumption of a homogeneous conformity parameter is particularly restrictive. Consider high school students deciding whether to smoke. Under homogeneity, smoking when 70% of friends do not smoke is assumed to generate the same utility loss as not smoking when 70% of friends do smoke. Yet this symmetry is unrealistic: smokers may care little about the norm—due to addiction or a rebellious self-image—while non-smokers may face strong pressure to conform to the norm.
Imposing homogeneous peer effects can misrepresent preferences and lead to misguided policy recommendations. Suppose a social planner wishes to shift behavior toward a lower equilibrium.\footnote{Henceforth, a lower (higher) equilibrium refers to a game's outcome in which more players select the low (high) action.} She could finance campaigns that strengthen the taste for conformity when choosing the high action or that weaken the taste for conformity when choosing the low action. Examples of such policies are anti-smoking campaigns stressing the stigma of smoking (e.g., the U.S. FDA’s The Real Cost campaign) or campaigns highlighting that not drinking does not hinder social participation (e.g., “You don’t need alcohol to have fun”). For instance, an increase in the preference for conformity when smoking would decrease the utility of smoking for agents whose social norm is not to smoke. Conversely, a decrease in the preference for conformity when \textit{not} drinking alcohol would increase the utility of \textit{not} drinking for agents whose social norm is to drink. With homogeneous conformity parameters, however, the planner can act only unidimensionally, strengthening or weakening conformity pressure generically. Such policies may be counterproductive: for instance, reducing the preference for conformity may induce some non-smokers with a latent taste for smoking to start smoking.
Action-specific peer effects also imply that nudges targeting perceived norms can backfire. If agents choosing the high action do not care about conforming, a global-norm campaign such as the 2014 “Finish It’’ message—“Today, only 9% of teens smoke’’—might inadvertently lead to a higher equilibrium. Smokers, if they disregard social norms, face no additional penalty after learning that smoking is rare, while non-smokers may still incur a conformity cost if the preference for conformity when choosing the low action is strong. Under homogeneous and positive preference for conformity, by contrast, such nudges would reliably reduce smoking prevalence. \\
The main contribution of this paper is to show that action-specific peer effects can be identified and estimated for binary outcomes in an incomplete-information setting. In a conformity framework, I show that it is possible to separately identify $(i)$ the taste for conformity when the agent chooses the low action, and $(ii)$ the taste for conformity when she chooses the high action. This allows researchers to directly infer whether conformity pressures are stronger for the low or high action. Identification relies on allowing the social-distance function—measuring distance between the agent’s action and the average action of her friends—to vary depending on the action chosen by the agent. This generalizes the standard specification used in the literature akerlof1997,bernheim1984,brock_discrete_2001, which assumes that conformity pressure is homogeneous across actions. Hence, the model with action-specific preference for conformity effects nests the homogeneous specification, enabling likelihood ratio tests of whether conformity is action-specific or not.
Recent work has begun to introduce heterogeneity into empirical peer effect models. One strand conditions heterogeneity on observable characteristics such as demographics nakajima2007,hsieh2018,houndetoungan_count_2020 or network centrality lin2017. These approaches allow researchers to examine, for example, whether peer effects are stronger among individuals of the same gender or increase with the relative centrality of friends.
A second strand exploits variation across friendship groups to identify heterogeneity based on the distribution of friends’ outcomes. For continuous outcomes, boucher2023 allow norms to depend on the minimum or maximum friend outcome rather than on the mean, while houndetoungan2025 allow peer effects to vary across quantiles of friends’ outcomes. herstad2024 study a reduced-form model in which friends’ influence is heterogeneous by their rank within the distribution of friend outcomes.
A third approach, to which this paper belongs, introduces action-specific peer effects. Badev2021 develops a spillover model for binary outcomes, in which agents choose both their actions and their friends within a potential game framework under complete information. xu2018 proposes a related spillover model in an incomplete-information setting, where peer effect parameters vary across combinations of an agent’s action and her friends’ actions. However, for identification, he imposes a normalization that sets to zero all peer effect parameters involving cases where either the agent or her friends choose the low action. As a result, the model only recovers the effect of having friends choose the high action on the utility of choosing the high action. \\
The heterogeneous conformity model I develop complements xu2018 and Badev2021's papers, as both use spillover models to derive action-specific heterogeneous peer effects for binary outcomes. Spillover models assume that the utility of choosing the high action increases linearly with the share (or number) of friends choosing that same action brock_discrete_2001. Although spillover and conformity microfoundations are equivalent under homogeneous conformity brock_discrete_2001, I show that this equivalence breaks down with action-specific heterogeneity.\footnote{See Appendix (ref) for a proof.} Thus, the two approaches generate distinct best responses and I propose a specification test to evaluate whether the data is consistent with the spillover model, the conformity model, or neither, under the condition that some action-specific heterogeneity is present. This specification test relates to boucher2023, who use a similar approach for continuous outcomes. Unlike their approach, my identification strategy does not need isolated agents, but requires action-specific heterogeneity in the peer effects. Beyond inference, developing social interaction models for both spillover and conformity microfoundations is also theoretically important as they have distinct welfare implications and motivate different policy interventions brock_discrete_2001,ushchev2020. \\
Section (ref) presents the microfoundations for the conformity model as a simultaneous game on a network with incomplete information, in which players do not observe their friends' behaviors but form rational expectations based on the available information. In Section (ref), I derive a sufficient condition for the existence of a unique Bayes-Nash equilibrium (BNE) and propose its comparative statics analysis. Section (ref) derives the conditions for identification of the action-specific conformity parameters and outlines the structural estimation strategy for the heterogeneous model. Finally, I bring the model to the data in Section (ref) and provide empirical evidence of heterogeneity in conformity preferences using the Add Health data, which includes network and individual data on secondary school students from 126 U.S. schools. This analysis primarily serves as a proof of concept and revisits an important paper by lee_binary_2014, which develops a network-based estimation strategy for the homogeneous peer effects model proposed by brock_discrete_2001, focusing on students' smoking behavior. I extend the analysis to alcohol consumption, illustrating how the magnitude of the action-specific heterogeneity varies across behaviors. The empirical results show that imposing a homogeneous, rather than heterogeneous, preference for conformity leads to misleading estimates of peer effects for smoking, but not for drinking. In addition, for smoking, I show that the data is not consistent with a spillover model and cannot reject the null hypothesis that the data is consistent with a conformity model, at the 0.1% significance level.
Consider a simultaneous game of incomplete information on a network with a set of $n$ players, indexed by $i$ and denoted by $\mathcal{N}=\{1,\dots,n\}$. I focus on a binary variable, $y_i \in \mathcal{Y}_i$, which indicates the behavior or action of player $i \in \mathcal{N}$. I define $\mathcal{Y}_i:=\left\{0,1\right\}$, where a value of one is given to the high action (e.g., smoking, exerting effort, or adopting an innovation) and a value of zero to the low action. The action profile of the $n$ players, the outcome of the game, is given by $\mathbf{y}=(y_1,\dots,y_n)^{\prime}$, where $y_i \in \mathcal{Y}_i$ for each $i\in \mathcal{N}$.
Players interact through a network, which is represented as an $n\times n$ binary matrix $\mathbf{A}=[a_{ij}]$, where \(a_{ij}=1\) if players i and j are friends, and 0 otherwise. When the network is not complete, agents have heterogeneous friend groups and interactions are local. Self-influence is not allowed, i.e., $a_{ii}=0 \; \forall i \in \mathcal{N}$ and links might be directed or undirected. I denote players' number of friends by $d_i:=\sum\limits_{j\neq i}a_{ij}$ and define $\mathbf{a}_i$ as the $i^{th}$ row of $\mathbf{A}$.\footnote{Henceforth, for any matrix $\mathbf{X}$, the vector $\mathbf{x}_k$ denotes its $k^{th}$ row.} I also use a row-normalized matrix $\mathbf{G}=[g_{ij}]$, which is obtained by row-normalizing $\mathbf{A}$, i.e., $g_{ij}=\frac{a_{ij}}{d_i}$.
Following brock_discrete_2001, I specify player $i$'s utility function $U(\cdot)$ as an additive function of three components:
where $\mathds{1}\{E\}$ is the indicator function, which takes the value one if event $E$ is true and zero otherwise. The individual characteristic $\alpha_i$ is observed by all players and captures the non-strategic utility. In my empirical application, $\alpha_i$ depends on students' characteristics, the characteristics of their friends, and school fixed effects. Note that the specification of the utility function ensures that $\alpha_i$ is normalized to zero for $y_i=0$. This normalization does not affect the equilibrium's action profile but is necessary for identifying the parameters associated with $\alpha_i$.
$S(y_i,\mathbf{y}_{-i})$ is the social distance function, which, for non-isolated players, depends on their friends' actions and incorporates the strategic dimension of the game, commonly referred to as the endogenous peer effects. For instance, the quadratic distance function, $S^{quad}(y_i,\mathbf{y}_{-i})$, introduced by bernheim1984 and akerlof1997, and widely used in network games with binary or continuous action spaces brock_discrete_2001,blume_linear_2015,boucher2023,houndetoungan_count_2020,patacchini2014,ushchev2020,li2009, is given by:\footnote{A different specification is proposed in Boucher2016, where the distance function is defined as $\frac{\beta}{2}\sum\limits_{j\neq i}\left[a_{ij}(y_i-{y_{j}})^2\right]$. In this formulation, the cost of deviating from friends’ behavior scales with the number of friends a player has. Appendix (ref) shows that action-specific heterogeneity in the preference for conformity can also be identified when the social distance function computes the weighted sum of deviations from friends’ actions, rather than the deviation from the weighted sum of friends’ actions. However, identification in that case relies on the existence of isolated players, which is why I focus on the less restrictive specification.}
where $\bar{y}_i$ is the weighted sum of $i$'s friends. Specifically, if the weights are given by the row-normalized matrix $\mathbf{G}$, then $\bar{y}_i:=\sum_{j\neq i}g_{ij}y_j$, in which case $\bar{y}_i$ represents the local norm. Alternatively, if the weights are given by the adjacency matrix $\mathbf{A}$, then $\bar{y}_i:=\sum_{j\neq i}a_{ij}y_j$, so that $\bar{y}_i$ corresponds to the number of $i$'s friends who choose the high action. Because the distance function is quadratic, the marginal social distance, $\frac{\partial S^{quad}_i(y_i,\mathbf{y}_{-i})}{\partial y_i}=\beta(y_i-\bar{y_i})$, increases with the deviation from the norm when $\beta>0$. This increasing marginal social distance implies that small deviations from the norm have a relatively minor effect on a player’s utility, whereas larger deviations are increasingly costly. Although this feature of the social distance function seems realistic, I show in Appendix (ref) that a linear social distance function can also be used, although the presence of isolated players is necessary for identification in that case.
Observe that $S^{quad}(0,\mathbf{y}_{-i})=S^{quad}(1,\mathbf{1}_n-\mathbf{y}_{-i})$ and the distance function is thus symmetric. Under this homogeneous specification, the distance from the social norm is assumed to have the same effect on players’ utility whether they choose the high or the low action. If $\beta>0$, players have a preference for conformity, and the larger the social distance between a player's action and her social norm, the lower the utility she derives from selecting this particular action. However, $\beta$ is unrestricted here and may be negative, indicating that players prefer to deviate from the norm due to anti-conformity preferences, capturing, for instance, status concerns immorlica2017,fershtman1998 or competitive preferences azmat2010,lopez2021. \\
I propose an alternative heterogeneous specification of the social utility function, $S^{het}(\cdot)$, that nests the homogeneous one:
$\beta^h$ captures the taste for conformity when choosing the high action. With binary outcomes, selecting $y_i=1$ implies that $y_i\geq \bar{y}_i$ and thus that the agent's action is (weakly) above the norm. $\beta^h$ can thus be interpreted as the cost of being above the norm. Conversely, $\beta^l$ measures the taste for conformity when choosing the low action and captures the cost of being below the norm. $\beta^h$ and $\beta^l$ do not need to have the same sign, and may be negative, in which case they measure the benefit of being above or below the norm, respectively. For example, in the case of smoking among teenagers, one might expect $\beta^l>\beta^h \geq0$ if students perceive smoking as socially desirable. In this scenario, students are strongly penalized for not smoking when smoking is the norm among their friends, but they experience a weaker, or even no, penalty for smoking, even if it goes against the norm. In contrast, polluting, a socially undesirable behavior, may be characterized by $\beta^h>\beta^l\geq0$. In that case, agents are strongly penalized for polluting when the norm among their friends is to abstain, but face little or no penalty for not polluting, even if most of their friends pollute. In general, one may expect the relationship between $\beta^h$ and $\beta^l$ to depend on the social desirability of the action under scrutiny. In the limit, $\beta^h \to \infty$ for prohibited behaviors, while $\beta^l \to \infty$ for mandatory behaviors. Note that the heterogeneous social distance function is asymmetric when $\beta^h\neq\beta^l$, in the sense that $S^{het}(0,\mathbf{y}_{-i})\neq S^{het}(1,\mathbf{1}_n-\mathbf{y}_{-i})$. \\
The tuple $(\epsilon_i(y_i))_{y_i\in \mathcal{Y}_i}$ represents private preference shocks for both the high and low actions, known to the player $i$ but unknown to the other players (and the researcher). I further consider that the distributions of $(\epsilon_i(y_i))_{y_i\in \mathcal{Y}_i}$ are common knowledge and identical for all players. Players are thus facing uncertainty about the realization of their friends' types. Let $\eta_i:=\epsilon_i(0)-\epsilon_i(1)$ such that $\eta_i$ corresponds to the unobserved relative preference of player $i$ for the low action over the high one, characterizing players' types without loss of generality.
The independence of $\epsilon_i(1)$ and $\epsilon_i(0)$, in conjunction with the continuity and integrability of $f_\epsilon$, guarantee that $\eta=\epsilon_i(0)-\epsilon_i(1)$ admits a continuous density $f_{\eta}$. To ensure econometric tractability, $\epsilon(1)$ and $\epsilon(0)$ are typically assumed to follow Type-I extreme value or normal distributions, so that $\eta$ is distributed according to a logistic or a normal distribution bajari2010. The last part of Assumption (ref) implies that players' actions can only be correlated through the network and $\alpha_i$, ruling out endogeneity issues. The network is thus assumed exogenous here for simplicity, although Badev2021 or lambotte2024 propose methods to account for network endogeneity in binary games. Independence of $(\epsilon_i(y_i))_{y_i\in \mathcal{Y}_i}$, and thus of $\eta$, from $\mathbf G$ and $\boldsymbol \alpha$ also implies that knowing its own type does not provide any information on the actions of other players. \\
Because players only observe their own preference shocks $(\epsilon_i(y_i))_{y_i\in \mathcal{Y}_i}$, they are uncertain about other players' actions $\mathbf{y}_{-i}$. As such, players do not maximize the utility function of Equation ((ref)) but its expectation with respect to their beliefs about other players' actions. I assume that players form rational expectations, i.e., objective mathematical expectations about other players' actions given the information set available to them. The information set of a given player is $\{\boldsymbol \alpha,\mathbf G,\eta_i\}$. Under Assumption (ref), players' types are independent of $\mathbf G$ and $\boldsymbol \alpha$, such that, by mutual independence, the information set reduces to $\{\boldsymbol \alpha,\mathbf G\}$.
Note that the rational expectations are heterogeneous in the sense that players with different characteristics and positions in the network are expected to play different actions. I denote the rational expectation profile by $\mathbf p:=(p_1,\dots,p_n)^{\prime}$ and the expected norm faced by player $i$ as $\bar{p}_i=\sum_{j \neq i}g_{ij}p_j$.
The utility function with the heterogeneous social distance function can be written as:
The associated expected utility is then given by:
where $\bar{p}_i=\sum_{j\neq i}g_{ij}p_j$, $\boldsymbol \Sigma=\left(\mathbf p \mathbf p^{\prime}+\operatorname{diag}( \mathbf p\circ(\mathbf{1}_n-\mathbf p))\right)$, $\circ$ is the Hadamard product and the derivation of $\boldsymbol\Sigma $ follows from Assumption (ref), which implies $p_i \perp p_j \, \forall i \neq j$.\footnote{More precisely, we have
where $\operatorname{diag}(\mathbf p\circ(\mathbf{1}_n-\mathbf p))=
$. Henceforth, $\operatorname{diag}(\mathbf a)$, with $\mathbf a$ an $n\times 1$ vector, yields an $n\times n$ matrix whose diagonal entries are the elements of $\mathbf a$ and off-diagonal elements are zero, while $\operatorname{diag}(\mathbf A)$, with $\mathbf A$ an $n\times n$ matrix, extract the diagonal elements of $\mathbf A$ and return them as an $n\times 1$ vector.}
Let $\mathbb{E}[\Delta _iU_i]:=\mathbb{E}\left[U_i(y_i=1,\mathbf{y}_{-i})- U_i(y_i=0,\mathbf{y}_{-i})|\boldsymbol \alpha,\mathbf G\right]$ be player $i$'s marginal expected utility from playing $y_i=1$ instead of $y_i=0$, such that her decision rule can be written as:
where $\Delta \beta:=\beta^l-\beta^h$ captures the heterogeneity between the conformity parameters. Assuming $\beta^h>0$, if player $i$ expects that a majority of her friends play the high action, conforming to the norm and selecting $y_i=1$ would increase her marginal expected utility by $\beta^h(\bar{p}_i-\frac{1}{2})>0$, a pattern that resembles majority games jackson2015.\footnote{Note that in brock_discrete_2001 and lee_binary_2014, the majority threshold is zero instead of $\frac{1}{2}$ and does not appear explicitly, since they use an alternative coding scheme, $y_i=\{-1,1\}$.} Conversely, if player $i$ expects that a majority of her friends select the low action, then $\bar{p}_i<\frac{1}{2}$, and playing $y_i=1$ would decrease her marginal expected utility by $\beta^h(\bar{p}_i-\frac{1}{2})<0$. If $\beta^h$ is negative, players are anti-conformist and prefer to diverge from the norm. As such, selecting $y_i=1$ increases player $i$'s marginal expected utility only if she expects a minority of friends to play the high action. \\
The nonlinearity in the expected marginal utility $\mathbb{E}[\Delta _iU_i]$, captured by $\Delta \beta$, arises only when the social distance function is heterogeneous and quadratic.\footnote{If one considers a linear social distance function with heterogeneous conformity preferences, such as $S^{lin}(y_i,\mathbf{y}_{-i})=y_i\beta^h\left(\bar{y}_i-y_i\right)+ (1-y_i)\beta^l\left(y_i-\bar{y}_i\right)$, the marginal expected utility would be $\mathbb{E}[\Delta _iU_i]=\alpha_i+\mathds{1}\{d_i>0\}\left[\beta^h\left(\bar{p}_i-1\right)+\beta^l\bar{p}_i\right]-\eta_i$, which is piecewise linear in $\bar{p}_i$. However, action-specific heterogeneity can still be identified if they are isolated players in the network (see Appendix (ref)). Assuming a homogeneous quadratic distance function yields $\mathbb{E}[\Delta _iU_i]=\alpha_i+\mathds{1}\{d_i>0\}\left\{\beta\left(\bar{p}_i-\frac{1}{2}\right)\right\}-\eta_i$, which is also piecewise linear in $\bar{p}_i$.} More formally, $\operatorname{sign}(\Delta \beta)$ determines if $\mathbb{E}[\Delta _iU_i]$ is a concave or convex function of the expected norm, since $\frac{\partial^2 \mathbb{E}[\Delta _iU_i]}{\partial^2 \bar{p}_i} =\Delta \beta$. This result is intuitive: if choosing an action below the norm is more (less) costly than choosing an action above, the marginal effect of $\bar{p}_i$ on the marginal expected utility from selecting $y_i=1$ increases (decreases) as $\bar{p}_i \to 1$. $\Delta \beta$ also partially determines if $\mathbb{E}[\Delta _iU_i]$ is increasing or decreasing with the social norm, as $\frac{\partial \mathbb{E}[\Delta _iU_i]}{\partial \bar{p}_i}= \mathds{1}\{d_i>0\}\beta^h+\Delta \beta \bar{p}_i$.
Let $\Gamma_i(\mathbf p):\{0,1\}^{n} \rightarrow \{0,1\}$ be the best response function, which maps player $i$'s rational expectations about other players' actions, $\mathbf p_{-i}$, into her best-response in the strategic space, the conditional choice probabilities $\Pr(y_i=1|\boldsymbol \alpha,\mathbf G)$, such that $\Gamma_i(\mathbf p^*):=\Pr(y_i=1|\boldsymbol \alpha,\mathbf G)$. Under Assumption (ref), $\Pr(y_i=1|\boldsymbol \alpha,\mathbf G)$ can be written as:
The rational expectations profile $\mathbf{p}$ is consistent with respect to the distribution of players' types if $p_i=\Pr(y_i=1|\boldsymbol \alpha,\mathbf G) \, \forall \, j \in \mathcal{N}$, where $p_i=\mathbb{E}[y_i|\boldsymbol \alpha,\mathbf G]$ is the expectation of player $i$'s action, as defined in Assumption (ref). Recall that the best response function is given by $\Gamma_i(\mathbf p)=\Pr(y_i=1|\boldsymbol \alpha,\mathbf G)$. Thus, the existence of a consistent rational expectation profile $\mathbf{p}$ implies $p_i=\Gamma_i(\mathbf p)$.
The strategy profile $\mathbf p^*$ is a BNE if $p^*_i \in \Gamma_i(\mathbf p^*)$ for all $i \in \mathcal{N}$. Indeed, letting $\Gamma:=\{\Gamma_i\}_{i \in \mathcal{N}}$, it follows from Brouwer fixed point theorem that $\mathbf p^*$ is a fixed point of $\Gamma$ such that $\Gamma(\mathbf p^*)=\mathbf p^*$. The strategy profile $\mathbf p^*$ at a BNE can be written as:
where $\boldsymbol\Sigma ^*=\left(\mathbf p^*\mathbf p^{*\prime}+\operatorname{diag}(\mathbf p^*\circ(\mathbf 1_n-\mathbf p^*))\right)$.
To ensure the uniqueness of the equilibrium, I introduce a sufficient condition, bounding the preference for conformity and the strength of the heterogeneity.\footnote{A simpler but more restrictive sufficient condition for uniqueness can be derived from Assumption (ref): $\max\{|\beta^h|,|\beta^l|\}<\frac{1}{4\max_u f_{ \eta}(u)}$.}
If $f_\eta$ is logistic or normal, Assumption (ref) yields $|\beta^h|+\frac{3}{2} |\beta^l-\beta^h| < 4$ or $|\beta^h|+\frac{3}{2} |\beta^l-\beta^h| < \sqrt{2\pi}$, respectively. If the preference for conformity is homogeneous, i.e., $\beta^l=\beta^h=\beta$, the sufficient condition for the uniqueness of the equilibrium simplifies to $|\beta|< \frac{1}{\max_u f_{ \eta}(u)}$, which is the same condition as in lee_binary_2014. Note that a similar assumption, called Moderate Social Influence, is used to guarantee uniqueness in games with continuous action spaces horst2006,glaeser2003. With binary action spaces, incomplete information generally provides straightforward conditions for uniqueness lin2024,xu2018,liu2019. Conversely, binary games with complete information generally have multiple equilibria depaula2013,ciliberto2009. li2016 and leung2020 propose promising partial identification approaches using subnetworks and strategic neighborhoods, respectively. Other papers krauth2006,soetevent2007,bajari2010identification specify or estimate a selection rule to determine which equilibrium is played to complete the model tamer2003. Another option is to assume a sequential game with myopic players, as in nakajima2007 and Badev2021. I leave the extension of the heterogeneous conformity model to complete information settings for future research.
The proof of Proposition (ref) is provided in Appendix (ref) and is based on the contraction property of the mapping $\Gamma(\mathbf p)$. $\Gamma(\mathbf p^*)$ being a contraction mapping is a sufficient condition for the uniqueness of the fixed point $\mathbf p^*=\Gamma(\mathbf p^*)$. As the consistency of the rational expectation belief system implies that for any $\mathbf p$, $\mathbf p=\Gamma(\mathbf p)$ only if $\mathbf p$ is an equilibrium, the unique fixed point $\mathbf p^*$ is the unique equilibrium of the game. Note that Assumption (ref) is not a necessary condition for uniqueness, but is a sufficient and empirically tractable condition (see Section 5.5 in bhattacharya2024 for a related discussion).
The next corollary presents comparative statics of the equilibrium profile, emphasizing the role of action-specific heterogeneity in conformity preferences. Proofs are omitted because they follow directly from differentiating $\mathbf p^*$ with respect to $\beta^l$, $\beta^h$ and $\bar{\mathbf{p}}^*$.
Corollary (ref)$(i)$ implies that an increase in the taste for conformity when playing the low action always leads to a higher equilibrium. Indeed, agents choosing the low action are always weakly below the expected norm, and when $\beta^l$ increases they face a larger penalty from deviating. This may induce them to switch to the high action, thereby raising the overall equilibrium. Hence, if a social planner wishes to attain a higher (lower) equilibrium, increasing (decreasing) $\beta^l$ through social marketing campaigns would be an effective policy.
Corollary (ref)$(ii)$ delivers an opposite insight for $\beta^h$. Since agents who play the high action are always weakly above the expected norm, increasing (decreasing) $\beta^h$ leads to a larger penalty and thus a lower (higher) equilibrium.
Corollary (ref)$(iii)$ highlights the non-monotonic effect of the local norm on players' expected utility and best response strategies. Specifically, the impact of a shift of the expected norm on the equilibrium strategy profile can be positive for some players and negative for others, depending on the idiosyncratic value of the local norm. Note that in a homogeneous conformity model, the effect of the expected norm is linear in $\beta=\beta^h=\beta^l$ and independent of the value of the expected norm. Corollary (ref)$(iii)$ can be relaxed if the support of $\bar{p}_i$ is restricted. For instance, if all the players expect that a minority of their friends play the high action, one can show that $\left(\frac{1}{2}\mathbf{1}_n\geq \left( \bar{\mathbf{p}}^*\right)_{(1)} \geq\bar{\mathbf{p}}^*_{(0)} \right) \land \left(\beta^h \geq \beta^l\right) \land \left( \beta^h \geq 0 \right) \implies \mathbf p^*_{(1)} \geq \mathbf p^*_{(0)}$. In the limit, in societies where the network is complete or where all players face identical local norms, i.e., where $\bar{p}_i^*=p^* \, \forall i \in \mathcal{N}$, the sign of $\left(\bar{\mathbf p}^*_{(1)} - \bar{\mathbf p}^*_{(0)}\right)$ is constant across players for any value of the tastes for conformity.\footnote{To see this, assume that $\tilde{\bar{\mathbf{p}}}^*$ is a mean-preserving contraction of $\bar{\mathbf{p}}^*\implies \mathbb{E}[\bar{\mathbf{p}}^*]=\mathbb{E}[\tilde{\bar{\mathbf{p}}}^*]$ and $\max_i \tilde{\bar{\mathbf{p}}}^*_i- \min_i\tilde{\bar{\mathbf{p}}}^*_i<\max_i \bar{\mathbf{p}}^*_i- \min_i\bar{\mathbf{p}}^*_i$. If $\beta^h \geq -\frac{\beta^l\bar{p}_i}{1-\bar{p}_i}\; \forall i\in \mathcal{N}\implies \beta^h \geq\max_i \frac{\beta^l\bar{p}_i}{1-\bar{p}_i}>\max_i \frac{\beta^l\tilde{\bar{p}}_i}{1-\tilde{\bar{p}}_i}$. If $\beta^h \leq -\frac{\beta^l\bar{p}_i}{1-\bar{p}_i}\; \forall i\in \mathcal{N}\implies \beta^h \leq\min_i \frac{\beta^l\bar{p}_i}{1-\bar{p}_i}<\min_i \frac{\beta^l\tilde{\bar{p}}_i}{1-\tilde{\bar{p}}_i}$. In both cases, Corollary (ref)$(iii)$ is more likely to hold under the mean-preserving contraction.}
Corollary (ref)$(iii)$ also shows that if $\beta^h$ and $\beta^l$ are both positive (negative), shifting the expected social norm toward the high action—e.g., via a social norm nudge—enables the social planner to reach a higher (lower) equilibrium, independently of the initial value of the norm. This result is particularly useful as one may expect that the action-specific tastes for conformity are generally of the same sign. However, if the taste for conformity when choosing the high action is negative (positive) while the other parameter is positive (negative), a nudge that shifts the norm toward the high action—and thereby reduces $\frac{\mathds{1}_n - \bar{\mathbf p} }{\bar{\mathbf p}}\beta^l$-may instead lead to a lower equilibrium.
I assume that the data is collected from many networks (many networks asymptotic).\footnote{To keep the notation as simple as possible, I only use the network index $m=1,\dots,M$ when necessary.} I also use a standard empirical specification for $\alpha_i$, $\alpha_i=\mathbf{m}_i^{\prime}\boldsymbol{\gamma}_0 + x_i^{\prime}\boldsymbol\gamma_1 +\bar{\mathbf{x}}_i^{\prime}\boldsymbol\gamma_2$, where $\mathbf{m_i}\in \{0,1\}^{M}$ is a $M$-vector indicating player $i$'s membership to one of the $M$ networks and controls for correlated effects through the vector of network fixed effects $\boldsymbol \gamma_0\in \mathbb{R}^M$,\footnote{The presence of the constant $\frac{1}{2}$ in the term $\mathds{1}\{d_i>0\}\beta^h\left(\bar{p}_i-\frac{1}{2}\right)$ does not threaten the identification of the networks fixed effects, even if there are no isolated players, in which case the intercept is $\boldsymbol{\gamma_0}-\frac{\beta^h}{2}$. Indeed, $\beta^h$ is identified independently by $\bar{p}_i$ if Assumption (ref) holds.} $x_i\in \mathbb{R}^K$ is a $K$-vector of individual characteristics and $\bar{\mathbf{x}}_i=\sum_{j\neq i}g_{ij}\mathbf x_j \in\mathbb{R}^K$ are the average characteristics of $i$'s friends and capture exogenous peer effects. Without loss of generality, let $\boldsymbol\gamma := (\boldsymbol\gamma_0^{\prime}, \boldsymbol\gamma_1^{\prime}, \boldsymbol\gamma_2^{\prime})^{\prime}$ and $\mathbf{z}_i=\left(\mathbf{m}_i^{\prime}, \mathbf{x}_i^{\prime}, \bar{\mathbf{x}}_i^{\prime}\right)^{\prime}$.
In addition, let $\boldsymbol \theta=(\boldsymbol \gamma^{\prime},\beta^h, \Delta \beta)^{\prime}$ and $\mathbf k_i=\left(\mathbf z_i^{\prime}, \mathds{1}\{d_i>0\}\left(\bar{p}_i-\frac{1}{2}\right), \frac{1}{2}\mathbf g_i\boldsymbol\Sigma \mathbf g_i^{\prime}\right)^{\prime}$ gather the model's parameters and explanatory variables for a player $i$ in network $m$, respectively. Identification fails if two different vectors of parameters $\boldsymbol \theta$ and $\tilde{\boldsymbol \theta}$ yield the same equilibrium $\mathbf p_m$ in network $m$. More formally, observe that since $\boldsymbol{\alpha}_i=\mathbf{z}_i^{\prime}\boldsymbol \gamma$, $\mathbf{p}_m=\Pr\left(\mathbf{y}_m|\mathbf{Z}_m,\mathbf{G}_m;\theta\right)$, where $\mathbf{y}_m$, $\mathbf{Z}_m$ and $\mathbf{G}_m$ are observed in the data for many networks $m$. Let $\boldsymbol\theta$ and $\tilde{\boldsymbol\theta}$ be observationally equivalent at $\mathbf{Z}_m$ and $\mathbf{G}_m$ if $\Pr\left(\mathbf{y}_m|\mathbf{Z}_m,\mathbf{G}_m;\boldsymbol\theta\right)=\Pr\left(\mathbf{y}_m|\mathbf{Z}_m,\mathbf{G}_m;\tilde{\boldsymbol\theta}\right)$, or, following Equation (ref), if $F_\eta\left(\mathbf{Z}_m,\mathbf{G}_m;\boldsymbol\theta\right)=F_\eta\left(\mathbf{Z}_m,\mathbf{G}_m;\tilde{\boldsymbol\theta}\right)$. The vector of parameters $\boldsymbol \theta$ is identified if there does not exist a $\tilde{\boldsymbol \theta} \ne \boldsymbol \theta$ that is observationally equivalent to $\boldsymbol \theta$. Observational equivalence's approach to identification is standard in network games of incomplete information guerra2022,yang2017,brock2007,houndetoungan_count_2020.
Assumption (ref)$(i)$ implies that $F_\eta$ is strictly increasing, which guarantees that $F_\eta(u) = F_\eta(\tilde{u})$ only if $u = \tilde{u}$. Assumption (ref)$(ii)$ requires the model’s explanatory variables to be linearly independent, which is almost surely guaranteed by variation across networks. Especially, the second-to-last element of $\mathbf{k}_i$, $\mathds{1}\{d_i>0\}\left(\bar{p}_i-\frac{1}{2}\right)$, varies across players and independently of the other variables if there exists, for at least one network, two players $i$ and $j$ who face a different social norm, i.e., $\exists (i,j): \bar{p}_i\neq \bar{p}_j$. Similarly, $\mathbf g_i\boldsymbol\Sigma \mathbf g_i^{\prime}$ varies independently if there exists a network with two players $k$ and $\ell$ who have at least one different friend, i.e., $\exists j: g_{kj}\neq g_{\ell j}$.\footnote{Indeed, $\mathbf g_i\boldsymbol\Sigma \mathbf g_i^{\prime}$ varies independently if $\exists (k,\ell):\mathbf{g}_k \boldsymbol \Sigma \mathbf{g}_k^{\prime} - \mathbf{g}_\ell \boldsymbol \Sigma \mathbf{g}_\ell^{\prime} = (\mathbf{g}_k-\mathbf{g}_\ell)' \boldsymbol\Sigma(\mathbf{g}_k+\mathbf{g}_\ell)'= \sum_{r=1}^n \left(g_{kr}^2 - g_{\ell r}^2\right) p_r + \sum_{\substack{r,j=1 \\r \neq j}}^n \left(g_{kr} g_{kj} - g_{\ell r} g_{\ell j}\right) p_r p_j\neq 0$. The first term is different from zero if $g_{kr}\neq g_{\ell r}$, and the second term is zero only if $\sum_{r=1}^n \left(g_{kr}^2 - g_{\ell r}^2\right) p_r = \sum_{\substack{r,j=1 \\ r\neq j}}^n \left(g_{kr} g_{kj} - g_{\ell r} g_{\ell j}\right) p_r p_j$. These two conditions are unlikely to hold for all players in all networks, especially if the networks are not complete, i.e., if all players are not friends with eveveryonelse in their network. Hence $\mathbf{g}_k \boldsymbol \Sigma \mathbf{g}_k^{\prime} - \mathbf{g}_\ell \boldsymbol \Sigma \mathbf{g}_\ell^{\prime}\neq0$.} Finally, $\mathds{1}\{d_i>0\}\left(\bar{p}_i-\frac{1}{2}\right)$ and $\frac{1}{2}\mathbf g_i\boldsymbol\Sigma \mathbf g_i^{\prime}$ both depends on $\mathbf{p}_m$ and $\mathbf{G}_m$, I thus impose that there is at least one network where one player has more than one friend, i.e. $\exists i: d_i\geq2$. Indeed, if all players have at most one friend, then for any player $i$, $\exists k\neq i:g_{ik}=1$ and $\mathbf{g}_i \boldsymbol \Sigma \mathbf{g}_i^{\prime}=g_{ik}^2p_k =p_k=\bar{p}_i$. Hence, $\mathds{1}\{d_i>0\}\left(\bar{p}_i-\frac{1}{2}\right)$ and $\frac{1}{2}\mathbf g_i\boldsymbol\Sigma \mathbf g_i^{\prime}$ would be linearly dependent. All these conditions are very mild when the number of network $M$ grows large and at least one network is not complete.
The proof of Proposition (ref) is standard in games of incomplete information and is provided in Appendix (ref). \\
A limitation of Assumption (ref) is that the rank condition is imposed on variables that depend on $\mathbf p_m$, an equilibrium quantity which is not observed in the data. Assumption (ref)$(ii)$ is thus not directly testable in that setting. It also cannot be formally tested after the model has been estimated, because if the model is not identified, the implied estimate of $\mathbf{p}_m$ is not valid. houndetoungan_count_2020 proposes a solution in his conformity model for count outcomes, which uses the presence of intransitive triads, as in Bramoulle2009. I present below the intuition of this identification strategy, and the adaptation of houndetoungan_count_2020's proof to the conformity model with action-specific heterogeneity is available in Appendix (ref). Let $\pi_i\in \{0,1\}$ be a dummy variable that equals one if player $i$ has friends of friends who are not her direct friends, that is, if $i$ lies at the start of an intransitive triad: $i \rightarrow j \rightarrow k$, but $i \nrightarrow k$, where $ a\rightarrow b$ denotes the existence of a link between players $a$ and $b$.
Condition $(i)$ imposes that the observed variables in $\mathbf z_i$ are not linearly dependent and is thus less strict than Assumption (ref)$(ii)$. Condition $(ii)$ ensures that intransitive triads exist in the networks. Condition $(iii)$ guarantees that $p_i$ is correlated with the contextual variable $\bar{\mathbf{x}}_{i,\kappa}$, such that the marginal effect $\frac{d p_i}{d \mathbf{x}_{j,\kappa}}$ is not zero when $\gamma_{2,\kappa} \neq 0$ and $j$ is a friend of $i$. The condition $\gamma_{1,\kappa} \gamma_{2,\kappa} \geq 0$ is necessary to ensure that the direct marginal effects of $\mathbf{x}_{\kappa}$ through own and contextual effects do not cancel its indirect effect through peer effects. I present the intuition with a triple of players, $i$, $j$ and $k$, while the formal proof for any subnetwork configuration is given in Appendix (ref).
Assume that Assumption (ref)$(ii)$ is not verified, which implies that the matrix $\mathbf k_i$ is linearly dependent. For any player $\ell$, I can then write:
for some constants $\check{\beta}^o\in \mathbb{R}$, $\widecheck{\Delta\beta}\in \mathbb{R}$, $\check{\boldsymbol{\gamma}}_0\in \mathbb{R}^{M}$, $\check{\boldsymbol{\gamma}}_1\in \mathbb{R}^{K}$ and $\check{\boldsymbol{\gamma}}_2\in \mathbb{R}^{K}$. Here, $j$ is the only friend of $i$, such that, for $\ell=i$, Equation (ref) simplifies to:\footnote{If $j$ is the only friend of $i$, $g_{ij}=1$. Hence $\mathbf g_i\boldsymbol\Sigma \mathbf g_i^{\prime}=p_j$, $\bar{p}_i=p_j$ and $\bar{\mathbf{x}}_i=\mathbf{x}_j$.}
Thus, $p_j$, the expected action of $j$, is not influenced by $x_{k,\kappa}$ and the marginal effect $\frac{d p_j}{d \bar{\mathbf{x}}_{k,\kappa}}$ derived from Equation (ref) is zero. However, since $j$ and $k$ are peers and $\gamma_{2,\kappa} \neq 0$, $\frac{d p_j}{d \bar{\mathbf{x}}_{k,\kappa}}\neq 0$ (see Equation (ref) in Appendix (ref)) and there is a contradiction. Under the conditions given in Proposition (ref), Equation (ref) cannot hold for all players when many networks include intransitive triads and $\operatorname{plim}_{M \to \infty}\frac{1}{M}\sum_i^n \mathbf{z}_i^{\prime}\mathbf{z}_i$ has full rank.\\
This section presents the strategy for estimating the model parameters $\boldsymbol \theta$, assuming that the researcher observes $\mathbf{y}_m$, $\mathbf Z_m$ and $\mathbf G_m$ for $m=1,\dots,M$. Under Assumption (ref), rational expectations depend solely on the information set $\{\boldsymbol Z_m, \mathbf G_m\}$, i.e., the strategic decisions of player $i$'s friends do not influence player $i$'s rational expectations. Due to mutual independence, the strategic decisions of player $i$'s friends do not affect $y_i$, and each player's contribution to the global likelihood is independent. Consequently, the conditional likelihood of the observed action profile $y$ is simply the product of the likelihoods of the individual actions $y_i \, \forall i \in \mathcal{N}$, given the information set. The conditional log-likelihood function of the observed action profile $\mathbf y$, given the consistent rational expectations $\mathbf p^*$, is then expressed implicitly as $\mathcal{L}_M(\boldsymbol \theta; \mathbf p^*)$:
Let $\Theta$ denote the support set of $\boldsymbol \theta$. The MLE is $\hat{\boldsymbol \theta}_{MLE}=\operatorname{argmax}_{\theta \in \Theta} \mathcal{L}_M(\boldsymbol \theta; \mathbf p^*)$ s.t. $\mathbf p^*:=\Gamma(\boldsymbol \theta,\mathbf p^*)$, as defined in Equation ((ref)).\footnote{To estimate standard errors for $\beta^l$ rather than $\Delta \beta$, I use an alternative but equivalent form of Equation ((ref)): $\Gamma(p^*)=F_{ \eta}\left(\boldsymbol \alpha+ \mathds{1}\{\mathbf{d}>\boldsymbol 0_n\}\left\{\beta^h\left(\bar{\mathbf{p}}^*-\frac{1}{2}\mathbf{1}_n-\frac{1}{2}\operatorname{diag}\left(\mathbf{G} \boldsymbol\Sigma \mathbf{G}^{\prime}\right)\right)+\frac{1}{2}\beta^l\operatorname{diag}\left(\mathbf{G} \boldsymbol\Sigma \mathbf{G}^{\prime}\right)\right\}\right)$. For simplicity, I retain the original form of Equation ((ref)) elsewhere in the paper.} \\
However, $\mathbf p^*$ is not directly observed in the data. To address this issue, I substitute $\mathbf p^*$ with an arbitrary vector $\mathbf p$ to compute the conditional pseudo log-likelihood $\Tilde{\mathcal{L}}_M(\boldsymbol \theta; \mathbf p)$ below:
Define $\Tilde{\boldsymbol \theta}_M(\mathbf p):=\operatorname{argmax}_{\boldsymbol \theta \in \Theta} \Tilde{\mathcal{L}}_M(\boldsymbol \theta;\mathbf p)$ and $\Gamma_M(\mathbf p):=\Gamma(\Tilde{\boldsymbol \theta}_M,\mathbf p)$. Two main approaches have emerged in the literature on estimating network games with incomplete information. The first one, used by lee_binary_2014 and lin2021, adapts the Nested Fixed-Point Likelihood method of rust1987. For each candidate value of $\boldsymbol \theta$, this sequential approach solves a fixed-point problem to compute $p$ and the associated pseudo-likelihood. However, solving the fixed point at every candidate $\boldsymbol \theta$ is computationally intensive, especially for large $M$. aguirregabiria2007 propose an alternative: the Nested Pseudo-Likelihood (NPL) estimator, which updates the fixed point rather than solving it exactly. Define an NPL fixed point as a pair $(\boldsymbol \theta,\mathbf p)$ s.t. $\mathbf p=\Gamma_M(\mathbf p)$. The set of NPL fixed points is given by $\mathcal{A}_M:=\{(\boldsymbol \theta,\mathbf p) \in \mathcal{\theta}\times \mathcal{P}: \boldsymbol \theta=\Tilde{\boldsymbol \theta}_M(\mathbf p),\mathbf p=\Gamma_M(\mathbf p) \}$. The NPL estimator is then $(\hat{\boldsymbol \theta}_{NPL},\hat{\mathbf p}_{NPL})=\operatorname{argmax}_{(\boldsymbol \theta,\mathbf p) \in \mathcal{A}_M}\Tilde{\mathcal{L}}_M(\boldsymbol \theta;\mathbf p)$. The sequential estimation starts with an initial guess of $\mathbf p^{(0)}$, then estimates $\hat{\boldsymbol \theta}^{(1)}=\operatorname{argmax}_{\boldsymbol \theta \in \Theta} \Tilde{\mathcal{L}}_M(\boldsymbol \theta; \hat{\mathbf p}^{(0)})$, and updates the fixed point in one step as $\hat{\mathbf p}^{(1)}=\Gamma_M\left(\hat{\boldsymbol \theta}^{(1)},\hat{\mathbf p}^{(0)}\right)$. The estimated $\hat{\mathbf p}^{(1)}$ is then substituted into the pseudo log-likelihood function to obtain updated parameter estimates, $\hat{\boldsymbol \theta}^{(2)}=\operatorname{argmax}_{\boldsymbol \theta \in \Theta} \Tilde{\mathcal{L}}_M(\boldsymbol \theta; \hat{\mathbf p}^{(1)})$. This sequence is repeated until convergence, i.e., when $ \lVert \hat{\mathbf p}^{(k+1)}-\hat{\mathbf p}^{(k)}\rVert <c$ or $\lVert\hat{\boldsymbol \theta}^{(k+1)}-\hat{\boldsymbol \theta}^{(k)}\rVert<c$, where $c$ is a tolerance value sufficiently close to zero. $\boldsymbol \theta^{(k+1)}$ maximizes the pseudo log-likelihood and $\hat{\mathbf p}^{(k+1)}$ is, by construction, a fixed point that satisfies the consistency condition on rational expectations.
As noted by aguirregabiria2007, convergence to a fixed point in the NPL estimator is not theoretically guaranteed. However, kasahara2012 shows that convergence is guaranteed if the initial starting point lies in the neighborhood of the true parameter and the fixed point mapping is a contraction, the latter holding under Assumption (ref).\footnote{In the empirical section, I verify the convergence of the NPL estimator by also estimating the Nested Fixed-Point Likelihood estimator.} \\
Under a set of regularity assumptions, the large-sample properties of the NPL estimator, namely, $\sqrt{n}$-consistency and asymptotic normality, are established in aguirregabiria2007. Let $\nabla_{k}\Tilde{\mathcal{L}}_M$ be the Jacobian matrix $\nabla_{k} \Tilde{\mathcal{L}}_M(\boldsymbol \theta_0; \mathbf p^*)$, where $\nabla_k$ is the partial derivative with respect to $k$, and $\nabla_k\Gamma$ be the Jacobian matrix $\nabla_k\Gamma(\boldsymbol \theta_0,\mathbf p^*)$. In addition, I define the following matrices: $\Omega_{\boldsymbol\theta \boldsymbol\theta^{\prime}}=\mathbb{E}\left[ \left(\nabla_{\boldsymbol \theta}\Tilde{\mathcal{L}}_M\right) \left(\nabla_{\boldsymbol \theta}\Tilde{\mathcal{L}}_M\right)^{\prime}\right]$, and $\Omega_{\boldsymbol \theta \mathbf p^{\prime}}=\mathbb{E}\left[ \left(\nabla_{\boldsymbol \theta}\Tilde{\mathcal{L}}_M\right) \left(\nabla_{\mathbf p}\Tilde{\mathcal{L}}_M\right)^{\prime}\right]$. The asymptotic distribution of the NPL estimator is given by $\sqrt{M}(\hat{\boldsymbol\theta}_{NPL}- \boldsymbol\theta_0) \overset{d}{\to} \mathcal{N}(0,\mathbb{V}_{NPL})$ where
The empirical application leverages data from the National Longitudinal Survey of Adolescent to Adult Health (Add Health), a widely used source for studying social interactions through detailed friendship networks. The dataset covers 126 U.S. middle and high school students, capturing a broad range of behaviors and socio-demographic characteristics. The social networks are based on self-reported friendships within the same school and grade, with students allowed to nominate up to five male and five female friends. Importantly, boucheraristide show that this censoring of network links does not qualitatively affect the estimate of the peer effect coefficients. In addition, participation in the survey was mandatory unless a parent opted out in writing, ensuring high participation rates, which mitigates concerns about measurement error due to missing agents chandrasekhar2011. Finally, hsieh2020 and Badev2021 show that biases arising from network endogeneity are minimal in the Add Health study.
The main reason for using the Add Health data in the empirical application is to allow a direct comparison between the heterogeneous peer effects model and the homogeneous model of lee_binary_2014. I therefore focus on the same sample, networks, and socio-demographic variables as lee_binary_2014, who study smoking behavior. For descriptive statistics of the sample, I refer the reader to their paper.
A student is classified as a non-smoker if they report never smoking or having smoked only once or twice in the past year. I extend this analysis to also examine alcohol consumption, defining a non-drinker as a student who reports never consuming beer, wine, or liquor, or having done so only once or twice in the past year. 23.1% of the sample smokes while 55.4% drinks alcohol.
As a proof of concept, I compare the homogeneous conformity model, used, for instance, in lee_binary_2014, with the action-specific heterogeneous model developed in this paper. Table (ref) reports the estimated conformity parameters for the homogeneous and action-specific heterogeneous models, as defined in Equation ((ref)). To ensure comparability with lee_binary_2014,\footnote{lee_binary_2014 encode the binary action as $\mathcal{Y}_i=\{-1,1\}$, while I use $\mathcal{Y}_i=\{0,1\}$. The reported parameters are, by construction, four times larger in my paper.} I assume that $f_\eta$ is logistic and specify as $\boldsymbol \alpha=\mathbf M\boldsymbol \gamma_0 + \mathbf X\boldsymbol \gamma_1+\mathbf G\mathbf X\boldsymbol \gamma_2$, i.e., I control for school fixed effects, students' characteristics and exogenous peer effects.\footnote{Given the sample size and the number of students per school, including fixed effects does not result in an incidental parameters issue in this context.}
For both behaviors under scrutiny, the heterogeneous model reaches the largest likelihood, although likelihood ratio (LR) tests between the homogeneous and heterogeneous models indicate a significantly larger likelihood for smoking behavior only ($p-value<0.001$). Since the preference for conformity is homogeneous for alcohol consumption, the specification test proposed in Appendix (ref) can only be used for smoking. The null hypothesis that the pure conformity model is consistent with the data cannot be rejected at a 0.01% significance level ($p-value=0.002$), while the null hypothesis that the pure spillover is consistent with the data is strongly rejected ($p-value<10^{-16}$).\footnote{Note that since the sample size if very large (74,783), I favor a strict significance level to reduce the risk of Type I error. At a 5% or 1% significance level, the specification concludes that the pure conformity model is not consistent with the data.}
In the case of smoking, comparing the estimated coefficients in both models reveals a pronounced action-specific heterogeneity in the preference for conformity, not accounted for by the homogeneous conformity parameter $\beta$. Specifically, in the homogeneous specification for smoking behavior, $\widehat{\beta}$ is estimated at $2.392$. In contrast, the heterogeneous model yields a lower estimate for taste for conformity when choosing the high action, $\widehat{\beta^h}=1.077$, and a larger estimate for the taste for conformity when choosing the low action, $\widehat{\beta^l}=3.980$. The homogeneous conformity parameter can thus be interpreted as a weighted sum of the heterogeneous conformity parameters. The heterogeneous model implies that the preference for conformity is stronger among non-smokers than among smokers. This result is intuitive: choosing not to smoke despite having at least one friend who smokes imposes a double penalty, both from deviating from the norm and exposure to passive smoking. Given the addictive nature of smoking, the relatively low preference for conformity for smokers is expected. Note that $\widehat{\Delta \beta}$ is positive and significant, indicating a positive and convex effect of the expected social norm on the marginal expected utility from smoking. Panel (a) of Figure (ref) illustrates how the marginal effect of the expected norm on the expected utility of smoking rather than not smoking, $\frac{\partial \mathbb{E}[\Delta _iU_i]}{\partial \bar{p}_i}$, increases as the expected norm approaches one in the heterogeneous model. The more smoking becomes the norm among friends, the more costly it is to abstain, since choosing an action below the norm is more penalized than choosing an action above. In the homogeneous model, the marginal effect of the norm is constant: a 10pp shift of the expected norm from 10 to 20% has the same effect as a shift from 80 to 90%. In contrast to smoking, the estimated heterogeneous tastes for conformity for drinking alcohol are not significantly different, suggesting that individuals have homogeneous preference for conformity across actions. In this case, the specification with a homogeneous social distance function provides a satisfactory approximation (see Panel (b) of Figure (ref) for an illustration).
The marginal effects of the norm are of the standard logit form, $\frac{\partial p^*_i}{\partial (\bar{p}_i^*)}= \mathds{1}\{d_i>0\}\left(\hat{\beta^h}+\widehat{\Delta \beta}\bar{p}_i^*\right)p^*_i(1-p^*_i)$. The sample average marginal effect of an increase in the social norm is strong: for non-isolated students, expecting a friend to shift from the low to the high action increases the probability of selecting the high action in the heterogeneous model by 0.115pp and 0.154pp for smoking and alcohol consumption, respectively.\footnote{Non-isolated students have on average 3.84 friends, and thus expecting one friend to shift to the high action corresponds to an increase of 0.26pp of the norm.} In the homogeneous model, the average marginal effect of the social norm is weakly lower, reaching 0.146pp and 0.154pp for tobacco and alcohol consumption, respectively. Peer effects are thus underestimated in the homogeneous specification in the case of smoking.
This paper introduces a structural conformity model that allows for action-specific heterogeneity in the preference for conformity to the local norm. The microfoundations of this model are established within a network game of incomplete information and a binary action space. The Bayes-Nash equilibrium is unique under a mild assumption on the strength of the heterogeneous tastes for conformity. I show that the model is identified, even with heterogeneous conformity parameters. I also propose a specification test to infer if the conformity or spillover models are (separately) consistent with the data, when action-specific heterogeneity is present.
Using the NPL estimator, I estimate the heterogeneous model on smoking behavior among US high school students and illustrate how assuming a homogeneous conformity preference, as in lee_binary_2014, may lead to biased estimates of endogenous peer effects. To further characterize the richness of modeling a heterogeneous preference for conformity, I apply the model to alcohol consumption. I uncover the intuitive result that deviating from the norm is more costly for non-smokers than smokers, because of passive smoking and the addictive nature of smoking. For alcohol consumption, I find that the taste for conformity is similar when choosing to drink or not. I can therefore run the specification test only for smoking, which rejects the null hypothesis that the spillover model is consistent with the data but fails to reject the null for the conformity model at the 0.1% significance level.
Based on these empirical findings, I encourage further research efforts to incorporate action-specific heterogeneity in peer effects models for count, multinomial and continuous outcomes.