EconBase
← Back to paper

Asymmetries in Peer Effects

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

92,388 characters

Asymmetries in Peer Effects10pt10pt We are grateful to Vincent Boucher for helpful comments and discussions. We acknowledge financial support from the Social Sciences and Humanities Research Council (SSHRC) under Grant CH150174. This research uses data from Add Health, a program directed by Kathleen Mullan Harris and designed by J. Richard Udry, Peter S. Bearman, and Kathleen Mullan Harris. Special acknowledgment is given to Ronald R. Rindfuss and Barbara Entwisle for assistance in the original design. Information on how to obtain Add Health data files is available on the Add Health website. \ Email addresses: mailto:[email removed]@ecn.ulaval.ca (A. Houndetoungan), mailto:[email removed]@univ-rennes.fr (M. Lambotte) \ An R package, including all replication codes, is available at: https://github.com/MathieuLambotte/AsyPeer.


\setlength{\abovedisplayskip}{6pt}
\setlength{\belowdisplayskip}{6pt}

\begin{mytitlepage}
\maketitle


\vspace{-1cm}

\begin{abstract}
\vspace{-0.5cm}
\singlespacing
\small \noindent
Individuals are often influenced by their peers because deviating from prevailing behavior entails social costs. However, existing peer effects models typically assume that individuals respond similarly to peers who perform better or worse than they do. This paper introduces a novel structural model of asymmetric peer effects in which conformity incentives depend on whether individuals perform below or above each of their peers. We establish that the model admits a unique equilibrium and show that its parameters can be identified and estimated through simple moment conditions. Applying our method to several student outcomes, we uncover strong evidence of asymmetries in peer effects. We then demonstrate that these asymmetries are highly policy-relevant by studying targeted interventions under budget constraints. Ignoring asymmetries leads to inefficient treatment allocation and substantial welfare losses, reducing welfare to levels comparable to those in a benchmark without social interactions.

\vspace{0.5cm}

\textbf{Keywords}: Social interactions, peer effects, asymmetric conformity, welfare analysis, targeted interventions

\textbf{JEL classification}: C31, C72, D61, D85
\end{abstract}
\end{mytitlepage}

\clearpage
\section{Introduction}
Interactions with peers and friends shape many economic and social behaviors. These social influences, commonly referred to as peer effects, have long been central to economics because they help explain how behaviors spread and inform policy design \citep[see][]{Manski1993, de2017econometrics, bramoulle2020peer, zenou2025peer}. These influences often operate through social comparison, generating concerns about status, competition, social approval, or stigmatization \citep[see][]{bernheim1994theory,akerlof1997,luttmer2005neighbors,benabou2006incentives,akerlof2000economics}.

A central tenet of the social comparison literature is that the direction of comparison matters: individuals often react differently to peers who perform above them than to peers who perform below them \citep{festinger1954theory, wills1981downward, card2012inequality}. For instance, a student's study effort may depend on whether her friends study more or less than she does \citep{feld2017understanding}. Yet existing models of peer effects largely ignore this distinction by treating peers who outperform an individual and those who underperform her as exerting the same influence. The standard linear-in-means model, for example, treats deviations above and below peers symmetrically \citep{blume_linear_2015}. As a result, this model may fail to capture how individuals actually respond to their peers and may therefore provide misleading guidance for policy design. Recent contributions that introduce heterogeneity in peer effects through the distribution of peers' outcomes \citep[e.g.,][]{boucher2023, herstad2026identification} likewise do not accommodate asymmetries in peer effects.


In this paper, we bridge the gap between social comparison theory and the peer effects literature by introducing a new model in which individuals respond differently to peers who perform above them than to those who perform below them. We rationalize this asymmetry through a game of conformity in which upward and downward comparisons affect preferences differently. We establish the existence and uniqueness of equilibrium and derive an econometric framework to identify and estimate both asymmetric effects. We show that the standard linear-in-means peer effects model estimates a weighted average of the underlying asymmetric effects, with weights that may be negative. As a result, the estimated symmetric peer effect may fall outside the range defined by the two asymmetric effects. Applying the proposed model to several student outcomes, we find pervasive evidence of asymmetric peer influence: for some outcomes, peers who perform above the individual exert strong effects, while those who perform below have little or no influence; for other outcomes, the reverse occurs. We show that these asymmetries are highly policy-relevant by studying a treatment allocation problem under limited resources. We find that ignoring asymmetry in peer effects leads to an inefficient treatment allocation, resulting in substantial losses in spillovers and welfare.

In our model, individual preferences are characterized by a utility function in which the influence of each friend depends on the friend's \textit{status} as higher- or lower-performing relative to the individual's performance. We impose no restrictions on which type of peer exerts a stronger influence. Instead, we allow the direction of influence to be determined by the data. This flexibility, however, introduces several complications in the game. Unlike in many peer effects models, the outcome of individual $i$, denoted by $y_i$, cannot be written explicitly as a function of peer outcomes, because peer statuses depend not only on the distribution of peer outcomes, but also on $y_i$ itself. As a result, the best-response function is implicit. Moreover, this function is not everywhere differentiable, as peer statuses switch discretely when $y_i$ crosses the outcome of one of her peers.
To show the uniqueness of the Nash equilibrium, we exploit the game's topological structure to establish that the best-response mapping remains a contraction under reasonable conditions.



We develop an econometric framework to identify both peer effect parameters. Our identification conditions extend those of the standard model \citep{Bramoulle2009} to accommodate asymmetry in peer effects. In practice, estimating our model requires exogenous predictions of peer statuses as instruments. We rely on flexible machine learning methods to construct such instruments. To establish consistency and asymptotic normality despite the use of generated instruments, our inference builds on recent developments in double/debiased machine learning \citep{Chernozhukov2018Double}. We conduct a Monte Carlo study showing that our approach performs well in finite samples. The simulations also illustrate, through concrete examples, how the standard symmetric model may estimate peer effects that fall outside the range of the asymmetric effects.

Using data from the National Longitudinal Study of Adolescent to Adult Health (Add Health), we show that many student activities exhibit asymmetric peer effects. For three of the four outcomes we study---namely, smoking, fighting, and optimism---we find significant asymmetry, although its direction differs across outcomes. For smoking, both types of peers exert strong effects, but peers who smoke more than the individual exert a larger influence. For fighting, individuals tend to conform strongly to more aggressive peers, while less aggressive peers exert little or no influence. For optimism, individuals tend to conform to more pessimistic peers, while more optimistic peers exert little or no influence. For drinking, by contrast, we find evidence that peer effects are approximately symmetric.



We further examine the policy implications of these empirical findings by studying targeting policies in which, because of limited resources, only a subset of students is treated, with the objective of reducing smoking, fighting, or drinking, or increasing optimism \citep{ballester2006,galeotti2020}. We find that ignoring asymmetry leads to inefficient treatment allocation, generating substantial losses in spillovers. Indeed, even when the symmetric model identifies a strong peer effect, using this model to select students for treatment may nevertheless generate little or no spillover effects, similar to those obtained in a benchmark without social interactions. This occurs particularly for fighting and optimism, where peer effects exhibit large asymmetries. By ignoring these asymmetries, the symmetric model often selects students who exert little influence under the true asymmetric model. The implemented policy therefore fails to fully leverage interactions among individuals. We show that this inefficiency generates substantial welfare losses.


This paper primarily contributes to the literature on how social interactions shape individual behavior \citep[see][]{de2017econometrics, bramoulle2020peer, zenou2025peer}. Recent developments highlight the importance of heterogeneity in peer effects, either through the distribution of peer outcomes \citep[see][]{carrell2010sex, boucher2023, herstad2026identification} or through predefined groups of peers \citep{comola2025heterogeneous, houndetoungan2026count}. We contribute to this literature by developing a new structural model that identifies distinct conformity effects from higher- and lower-performing friends. Our model differs from existing approaches because heterogeneity depends not only on the distribution of peer outcomes, but also on the individual's own outcome. This allows us to empirically test behavioral mechanisms that have long been emphasized in the theory of social comparison. Some papers also document heterogeneous peer effects by interacting individuals' performance types, such as high- or low-performing status, with those of their peers \citep[see, e.g.,][]{luttmer2005neighbors,feld2017understanding,carrell2013natural}. However, these analyses are reduced-form and are not derived from an underlying behavioral model. One exception is \cite{lambotte2025}, who estimates asymmetric peer effects in a binary-outcome game. In that case, however, the binary nature of the outcome makes the distinction between higher- and lower-performing peers equivalent to estimating separate effects for each of the two actions an individual can take.

We also contribute to the empirical literature on peer effects \citep{sacerdote2014experimental} by documenting asymmetric peer influence across several outcomes. Finally, by showing that ignoring asymmetry leads to suboptimal targeting decisions and substantial welfare losses, we contribute to the literature on optimal treatment allocation \citep{kitagawa2018should, galeotti2020, viviano2025policy}.


The remainder of the paper is organized as follows. Section \ref{sec:micro} presents the microfoundation of the structural model. Section \ref{sec:metrics} presents our identification approach and estimation method, and documents the bias of the standard symmetric model. Section \ref{sec:MCsimu} reports the results of the Monte Carlo study. Section \ref{sec:empirics} provides an empirical application using the Add Health data. Section \ref{sec:conclu} concludes.

\section{Microfoundations}\label{sec:micro}
\noindent We consider a complete information game with $n$ individuals who interact through a network, where $n \geq 2$. The network is characterized by an adjacency matrix $\mathbf{G} = [g_{ij}]_{i,j=1}^n$, where $g_{ij} \geq 0$ measures the strength of the link from $i$ to $j$. We allow for flexible network specifications: the network can be weighted or unweighted, and directed or undirected. For expositional simplicity, however, we focus on row-normalized unweighted networks; that is, $g_{ij} = \frac{1}{n_i}$ if $j$ is a friend (peer) of $i$, and $g_{ij} = 0$ otherwise, where $n_i$ denotes the number of friends of $i$. Additionally,  $g_{ii} = 0$ for all $i$, implying that individuals do not interact with themselves.


Individuals choose their actions (e.g., academic effort), measured by a continuous variable $y_i \in \mathbb{R}$, taking into account their friends' actions. Preferences are represented by the following utility function:
\begin{equation}\label{eq:utility}
 U_i(y_i, \mathbf y_{-i})=\alpha_i y_i - \dfrac{y_i^2}{2} - \mathcal{S}_i(y_i, \mathbf y_{-i}),
\end{equation}
where $\mathbf{y}_{-i} = (y_1, \dots, y_{i-1}, y_{i+1}, \dots, y_n)$ is a vector of other players' actions.

The utility function \eqref{eq:utility} is additively separable into private and social components. The private component is given by $\alpha_i y_i - \frac{y_i^2}{2}$, where $\alpha_i$ denotes idiosyncratic productivity and $\frac{y_i^2}{2}$ represents the private cost of choosing $y_i$. The term $\mathcal{S}_i(y_i, \mathbf{y}_{-i})$ represents a social distance (or social cost) and captures the effect of differences between individual $i$'s effort and their friends' efforts \citep{akerlof1997}. We introduce asymmetry into the social distance function by allowing social pressure from friends to depend on whether they exert higher or lower levels of effort than $i$.

\subsection{Asymmetric Preferences for Conformity}
Most papers rely on a symmetric social distance to capture conformity peer effects \citep[see, e.g.,][]{blume_linear_2015, ushchev2020}. This social distance is defined as
\begin{equation}\label{eq:sycost}
\mathcal{S}^{sym}_i(y_i, \mathbf y_{-i}) = \dfrac{\beta}{2}\sum_{j \ne i} g_{ij}(y_i - y_j)^2.
\end{equation}

Under this specification, individuals have conformist preferences when $\beta$ is positive and anti-conformist preferences when $\beta$ is negative. This social cost implies that deviations above or below the effort exerted by a friend $j$ yield the same social penalty or benefit.
However, the social penalty may depend on whether a peer performs better or worse than the individual. To account for this scenario, we introduce an asymmetric social cost:
\begin{equation}\label{eq:asycost}
    \mathcal{S}_i(y_i, \mathbf y_{-i}) = \dfrac{\beta^l}{2}\sum_{j \ne i: y_j \leq y_i} g_{ij}(y_i - y_j)^2 + \dfrac{\beta^h}{2}\sum_{j\ne i: y_j > y_i} g_{ij}(y_i - y_j)^2.
\end{equation}

The asymmetric social distance function allows for flexible mechanisms of peer effects. Individuals may conform more strongly to friends who exert higher levels of the behavior than they do ($\beta^h > \beta^l \geq 0$), or instead to friends who exert lower levels of the behavior ($\beta^l > \beta^h \geq 0$). Which form of asymmetry arises depends on the mechanism underlying the outcome, such as role-model effects, \textit{keeping up with the Joneses} dynamics, or norm formation. The model therefore does not impose \textit{a priori} which type of peer is more influential; instead, it allows the direction of influence to be determined empirically.

Furthermore, our model can also capture situations in which $\beta^h$ and $\beta^l$ have different signs. For example, in certain competitive environments, upward comparison can be discouraging while low-performing peers may boost confidence. This scenario occurs when $\beta^h < 0$ and $\beta^l > 0$.

The asymmetric specification nests the symmetric distance function when $\beta = \beta^h = \beta^l$. We can therefore empirically assess the validity of the symmetric specification by testing whether $\beta^h = \beta^l$. Under the asymmetric social distance, the influence exerted by a friend depends on both an individual's own action and those of her friends. This contrasts with recent models that introduce heterogeneity solely through the distribution of friends' outcomes \cite[e.g.,][]{boucher2023,herstad2026identification}.


\subsection{Equilibrium}
Individuals' optimal choices (effort levels) are determined by maximizing their utility functions. To ensure that this optimization problem has a unique finite solution, we impose a lower bound on $\beta^l$ and $\beta^h$. These conditions guarantee that the utility function is strictly concave.
\begin{assumption}\label{assumption:minbeta}
The parameters $\beta^l$ and $\beta^h$ satisfy $\beta^l > -\dfrac{1}{2}$ and $\beta^h > -\dfrac{1}{2}$.
\end{assumption}
As we show below, the total marginal effect of peers' outcomes on $y_i$ lies between $\frac{\beta^l}{1+\beta^l}$ and $\frac{\beta^h}{1+\beta^h}$. Thus, Assumption \ref{assumption:minbeta} implies that the magnitude of the total marginal peer effect is less than one, which is a standard stability condition in the literature \citep[see][]{blume_linear_2015}.

Under Assumption \ref{assumption:minbeta}, we establish several results in Appendix~\ref{proof:uniqueness}. First, we show in Lemma~\ref{lemma:differentiable} that $U_i(y_i, \mathbf y_{-i})$ is continuously differentiable in $y_i$ despite discrete changes in peer status in Equation \eqref{eq:asycost}.\footnote{We refer to a peer's \textit{status} as high- or low-performing depending on whether the peer's outcome exceeds or falls below that of the individual, respectively.} We also demonstrate in Lemma \ref{lemma:concave} that $U_i(y_i, \mathbf y_{-i})$ is strictly concave in $y_i$. Finally, in Lemma~\ref{lemma:BRFcontinuous}, we show that $i$'s strategy that maximizes their utility function given the choices of other individuals (i.e., their best-response function) is given by the following implicit function:

\begin{equation}\label{eq:yi1}
    y_i = \dfrac{\alpha_i + \beta^l \bar{y}_i^l + \beta^h \bar{y}_i^h}{1 + \beta^l g_i^l + \beta^h g_i^h}.
\end{equation}
In Equation~\eqref{eq:yi1}, $\bar{y}_i^l = \sum_{j \ne i} \mathbbm{1}\{y_j \leq y_i\} g_{ij} y_j$ and $\bar{y}_i^h = \sum_{j \ne i} \mathbbm{1}\{y_j > y_i\} g_{ij} y_j$, where $\mathbbm{1}\{\cdot\}$ denotes the indicator function. Moreover, $g_i^l = \sum_{j \ne i} \mathbbm{1}\{y_j \leq y_i\} g_{ij}$ and $g_i^h = \sum_{j \ne i} \mathbbm{1}\{y_j > y_i\} g_{ij}$ are the shares of low- and high-performing friends, respectively. The outcome $y_i$ is given in an implicit form because $\bar{y}_i^l$ and $g_i^l$ depend on $y_i$ itself.

For each $i$, there are $2^{n_i}$ possible configurations of peer statuses, where $n_i$ is the number of peers of $i$. Since $\mathbbm{1}\{y_j \leq y_i\}$ is constant within a given configuration, the best-response function is linear within each configuration. Therefore, the best-response function is piecewise linear across configurations.

The marginal effect of friends' outcomes on $y_i$ is $\frac{\beta^l}{1 + \beta^l}$ when all friends are low-performing, and $\frac{\beta^h}{1 + \beta^h}$ when they are all high-performing. For intermediate configurations, the marginal effect can take up to $2^{n_i}$ distinct values between $\frac{\beta^l}{1 + \beta^l}$ and $\frac{\beta^h}{1 + \beta^h}$. This introduces substantial heterogeneity, unlike in the standard symmetric model where the marginal effect is constant and equal to $\frac{\beta}{1 + \beta}$. However, this heterogeneity comes at the cost that the best-response function is not globally differentiable.

By exploiting the identities $\bar y_i = \bar{y}_i^l + \bar{y}_i^h$ and $g_i^l + g_i^h = 1$, which hold for individuals with at least one friend (non-isolated individuals), we can substitute $\bar{y}_i^l = \bar y_i - \bar{y}_i^h$ and $g_i^l = 1 - g_i^h$ into Equation~\eqref{eq:yi1}.\footnote{We assume that all individuals are non-isolated only for expositional convenience. Isolated individuals have an equilibrium outcome $y_i = \alpha_i$ and can be accommodated within our framework.} It follows that Equation~\eqref{eq:yi1} can be rewritten as:
\begin{equation}\label{eq:yi2}
    y_i = \dfrac{\alpha_i +  \beta^l\bar y_i + (\beta^h - \beta^l)\check y_i}{1 + \beta^l},
\end{equation}
where $\check{y}_{i} = \sum_{j \ne i} g_{ij}\mathbbm{1}\{y_{j} > y_{i}\}(y_{j} - y_{i})$. As we will show later in the econometric analysis, expressing the optimal outcome as in Equation~\eqref{eq:yi2} is important for deriving a reduced-form equation that is linear in parameters.

A strategy profile $\mathbf{y} = (y_i, \dots, y_n)^{\prime}$ is a Nash equilibrium if $y_i$ verifies Equation~\eqref{eq:yi2} for all $i$. We establish the following result.
\begin{proposition}\label{propo:equilibrium}
Under Assumption \ref{assumption:minbeta}, the game described by the utility function \eqref{eq:utility} has a unique Nash equilibrium $\mathbf y^{\ast} = (y_1^{\ast}, \dots, y_n^{\ast})$, such that $y_i^{\ast}$ verifies Equations \eqref{eq:yi2}.
\end{proposition}

The proof of Proposition~\ref{propo:equilibrium} is provided in Appendix~\ref{proof:uniqueness:theo}. We establish that the best-response mapping is a contraction; that is, for any strategy profiles $\mathbf{y}$ and $\tilde{\mathbf{y}}$ and any $i$, the uniform Lipschitz condition, $\lvert b_i(\mathbf{y}_{-i}) - b_i(\tilde{\mathbf{y}}_{-i})\rvert \leq \mu \lVert \mathbf{y} - \tilde{\mathbf{y}} \rVert_{\infty}$, holds for some constant $\mu < 1$, where $b_i$ denotes the best-response function of individual $i$.\footnote{For any $\mathbf a = (a_1, \dots, a_n)^{\prime} \in \mathbb R^n$, the infinity norm of $\mathbf a$ is defined as $\displaystyle \lVert \mathbf a \rVert_{\infty} = \max_{i} \lvert a_i \rvert$.} However, since statuses may differ across $\mathbf{y}$ and $\tilde{\mathbf{y}}$, achieving this result is not straightforward. We first partition the strategy space into cells (polyhedra). Within each cell, the statuses of friends remain fixed, allowing us to establish the uniform Lipschitz condition. We then extend this argument to the case in which $\mathbf{y}$ and $\tilde{\mathbf{y}}$ do not belong to the same cell.

\subsection{Treatment Assignment and Spillover Effects} \label{microfoundation::socialmultiplier}
In the presence of social interactions, interventions that affect the productivity of treated individuals (e.g., targeted subsidy programs) can also affect nonrecipients, thereby generating social multiplier or spillover effects. In this section, we study the effects of productivity shocks in the asymmetric model. In particular, using a treatment allocation problem in which the treatment takes the form of a productivity shock, we show that ignoring asymmetries in peer effects can lead to inefficient treatment allocation. For simplicity, we assume that productivity shocks do not affect the network structure.

In the standard symmetric model, treating all individuals by uniformly increasing their productivity yields the same uniform increase in outcomes and thus generates no social multiplier effects \citep{boucherfortin2015,ushchev2020}. We first show that this result extends to the asymmetric specification.
\begin{proposition}\label{propo:socialmultiplier}
Assume a uniform shock $\bar \alpha$ to any individual's productivity $\alpha_i$, such that $\tilde{\alpha}_i = \alpha_i + \bar \alpha$ is the new productivity level. Let $y_i$ and $\tilde{y}_i$ be the equilibrium outcomes corresponding to $\alpha_i$ and $\tilde{\alpha}_i$, respectively. We have $\tilde{y}_i = y_i + \bar \alpha$ for all $i$.
\end{proposition}

\noindent The proof of Proposition~\ref{propo:socialmultiplier} is provided in Appendix~\ref{append:proof:socialmultiplier}. The result is analogous to that obtained in the symmetric case, because treating all individuals does not alter peer statuses. Consequently, the asymmetric structure of peer effects does not play a role.

However, in many settings, only a subset of individuals can be treated due to budget constraints. This is the case for targeted policies, where interventions benefit only a limited number of individuals \cite[see][]{banerjee2019using, beaman2021can}. Such policies can improve cost-effectiveness when treatment is optimally allocated.

Assume that a social planner aims to maximize aggregate outcomes $\sum_{j=1}^n y_j$ by assigning treatment to a fixed number of individuals, where a budget constraint exogenously determines this number \citep[e.g.,][]{galeotti2020}. As we show below, in the standard symmetric model, the optimal treatment-assignment rule depends only on the network structure and the peer-effect parameter, but not on the initial productivity distribution  \citep[see][]{boucherfortin2015}.

Let $\mathbf I_n$ denote the identity matrix of order $n$ and $\mathbf 1_n$ an $n$-vector of ones. Define $\hat{\boldsymbol \alpha} = (\hat\alpha_1,\dots,\hat\alpha_n)'$, where $\hat\alpha_i = \alpha_i$ if $i$ is isolated and $\hat\alpha_i = \alpha_i/(1+\beta)$ otherwise. The equilibrium outcome of the symmetric model is given by $\mathbf y = \left(\mathbf I_n - \frac{\beta}{1+\beta}\mathbf G\right)^{-1}\hat{\boldsymbol\alpha}$ \citep[see][]{blume_linear_2015}. Consequently,
$$\sum_{j=1}^n y_j = \hat{\boldsymbol\alpha}^\prime \left(\mathbf I_n - \frac{\beta}{1+\beta}\mathbf G'\right)^{-1} \mathbf 1_n.$$
Let $\mathbf e_i \in \mathbb R^n$ denote a vector of zeros except for its $i$-th entry, which is equal to $1$ if individual $i$ is isolated and $1/(1+\beta)$ otherwise.\footnote{The parameter $\beta$ is the single peer effect parameter in the symmetric model (see Equation~\eqref{eq:sycost}).} The marginal effect of increasing $\alpha_i$ by one unit on aggregate outcomes is $\frac{\partial}{\partial \alpha_i}\sum_{j=1}^n y_j = \mathbf e_i' \left(\mathbf I_n - \frac{\beta}{1+\beta}\mathbf G'\right)^{-1} \mathbf 1_n$. Stacking these marginal effects for all individuals yields $\boldsymbol{\Delta} = \mathcal E \left(\mathbf I_n - \frac{\beta}{1+\beta}\mathbf G'\right)^{-1} \mathbf 1_n$, where $\mathcal E = (\mathbf e_1,\dots,\mathbf e_n)'$. Treatment should therefore be assigned according to the ranking induced by $\boldsymbol{\Delta}$, which depends only on $\mathbf G$ and $\beta$.

In contrast, in the asymmetric model, the optimal treatment assignment depends not only on the network structure and peer effects, but also on the entire productivity distribution. For instance, if conformity to low-performing peers is stronger (high $\beta^l$ and low $\beta^h$), targeting an individual who is generally a high-performing friend may be inefficient, even though this individual is central in the network under the symmetric model. In such cases, individuals’ relative positions in the outcome distribution play a key role in determining optimal assignment.


In Figure~\ref{fig:reflection}, we illustrate different targeting schemes using a 2-star network, where $i$ is the central node. Before the intervention (panel a), the aggregated outcome is $3.66$. Since $j$ and $k$ have no friends but are both connected to $i$, increasing the productivity of $j$ or $k$ yields the same aggregate increase in the symmetric model, regardless of $\beta$. Thus, a social planner who employs the symmetric model can assign the treatment to either $j$ or $k$, expecting the same shift in the outcome distribution. However, panels b and c show that targeting $j$ rather than $k$ is not optimal. Specifically, targeting $j$ increases the aggregate outcome by 1.12, i.e., a spillover effect of 12\%, whereas targeting $k$ almost doubles the spillover effect to 22\%.\footnote{The spillover effect is measured as the additional increase in the aggregate outcome relative to the increase in the aggregate outcome in a benchmark model without social interactions, expressed as a percentage. In this benchmark model, a one-unit increase in $\alpha_i$ raises the aggregate outcome by exactly one unit.} Panel d further shows that targeting $i$ is inefficient, as there are no spillover effects on $j$~and~$k$.

\begin{figure}[!htbp]
    \centering
    \begin{tikzpicture}[scale=0.5]
            \tikzstyle{node}=[circle,draw,fill=white,thick, inner sep=0.5pt]
            \tikzstyle{nodeh}=[circle,draw,fill=blue!40, thick, inner sep=0.5pt]
            \tikzstyle{nodel}=[circle,draw,fill=red!30, thick, inner sep=0.5pt]
            \tikzstyle{link}=[->, color=black!50, thick]
            \tikzstyle{targeted}=[circle,thick]

            \node[above, align=center] at (2.3, 3.4) {\footnotesize a) Before intervention};
            \node[node, label=180:$i$] (i1) at (0,0) {\footnotesize 1.16};
            \node[nodeh, label=0:$j$] (j1) at (3,2) {\footnotesize 1.50};
            \node[nodel, label=0:$k$] (k1) at (3,-2) {\footnotesize 1.00};

            \node[above, align=center, xshift=4cm] at (2.3, 3.4) {\footnotesize b) Targeted $j$};
            \node[node, label=180:$i$, xshift=4cm] (i2) at (0,0) {\footnotesize 1.29};
            \node[nodeh, targeted, label=0:$j$, xshift=4cm] (j2) at (3,2) {\footnotesize 2.50};
            \node[nodel, label=0:$k$, xshift=4cm] (k2) at (3,-2) {\footnotesize 1.00};

            \node[above, align=center, xshift=8cm] at (2.3, 3.4) {\footnotesize c) Targeted $k$};
            \node[node, label=180:$i$, xshift=8cm] (i3) at (0,0) {\footnotesize 1.38};
            \node[nodeh, label=0:$j$, xshift=8cm] (j3) at (3,2) {\footnotesize 1.50};
            \node[nodeh, targeted, label=0:$k$, xshift=8cm] (k3) at (3,-2) {\footnotesize 2.00};

            \node[above, align=center, xshift=12cm] at (2.3, 3.4) {\footnotesize d) Targeted $i$};
            \node[node, targeted, label=180:$i$, xshift=12cm] (i4) at (0,0) {\footnotesize 1.63};
            \node[nodel, label=0:$j$, xshift=12cm] (j4) at (3,2) {\footnotesize 1.50};
            \node[nodel, label=0:$k$, xshift=12cm] (k4) at (3,-2) {\footnotesize 1.00};

            \draw[link] (i1)--(j1);
            \draw[link] (i1)--(k1);

            \draw[link] (i2)--(j2);
            \draw[link] (i2)--(k2);

            \draw[link] (i3)--(j3);
            \draw[link] (i3)--(k3);

            \draw[link] (i4)--(j4);
            \draw[link] (i4)--(k4);

        \end{tikzpicture}
    \caption{Targeted policy}
    \label{fig:reflection}
    \justifying
    \footnotesize{\noindent Note: This figure illustrates a targeted intervention in a directed network of three nodes. The values indicated inside each node represent the individuals' performance (outcomes). Red nodes denote low-performing friends, and blue nodes denote high-performing friends. $\beta^l = 1.5$, $\beta^h = 0.5$, and before the intervention (panel a), $(\alpha_i, \alpha_j, \alpha_k) = (1.2, 1.5, 1)$. The intervention consists of increasing $\alpha$ of the targeted individual by one unit.}
\end{figure}

This result highlights a fundamental difference between the symmetric and asymmetric peer effect models. In the asymmetric model, the marginal social return to treatment depends on individuals' relative positions in the outcome distribution. The symmetric model ignores these positions and may lead to inefficient allocation.

However, unlike the symmetric model, the asymmetric peer effect model does not admit a closed-form characterization of the optimal treatment-allocation ranking. Determining the optimal allocation requires an exhaustive search over all possible treatment sets, the number of which grows rapidly with both network size and the number of individuals to be treated. To make treatment assignment computationally feasible when peer effects are asymmetric, we develop algorithms in Section~\ref{application:policy} that approximate the optimal allocation. We use these algorithms in the empirical application to assign treatment in large networks.


\section{Econometric Framework} \label{sec:metrics}

In this section, we present the econometric model and study its identification and estimation. Since interactions induce dependence across individuals and may complicate identification and inference, we assume that the population is partitioned into $M$ nonoverlapping and independent networks. This assumption is commonly imposed and appropriate for many data sets, as samples typically consist of independent schools, villages, or markets \citep{Bramoulle2009, blume_linear_2015, boucher2023}. We denote by $n_m$ the number of individuals in network $m$. We assume that $n_m$ is bounded, while $M$ goes to infinity asymptotically.

Throughout the paper, the subscript $m$ refers to variables defined for network $m$; for example, $\mathbf{G}_m$ denotes the interaction matrix in network $m$. A double subscript $m,i$ (e.g., $y_{m,i}$) denotes a variable associated with individual $i$ in network $m$, while a triple subscript $m,ij$ (e.g., $g_{m,ij}$) denotes a pairwise variable.


\subsection{Econometric Model}
Individual $i$ is characterized by a vector of observable characteristics $\mathbf{x}_{m,i} \in \mathbb{R}^{d_x}$, which includes control variables such as age, sex, etc. As in \cite{blume_linear_2015}, we model productivity $\alpha_{m,i}$ as:
\begin{equation}\label{eq:alpha}
    \alpha_{m,i} = c + \mathbf{x}_{m,i}^{\prime}\boldsymbol{\gamma}_1 + \bar{\mathbf{x}}_{m,i}^{\prime}\boldsymbol{\gamma}_2 + \varepsilon_{m,i},
\end{equation}
where $\bar{\mathbf{x}}_{m,i} = \sum_{j \ne i} g_{m,ij}\mathbf{x}_{m,j}$ is a vector of contextual variables defined as the average of exogenous characteristics among friends, and $\varepsilon_{m,i}$ is an unobserved preference shock. The parameters $\boldsymbol{\gamma}_1$ and $\boldsymbol{\gamma}_2$ capture the effects of individual characteristics and contextual variables, respectively, while $c$ is an intercept term. For ease of exposition, we exclude network fixed effects from Equation~\eqref{eq:alpha}. As we show below, although the model allows for a nonlinear relationship between individual and peer outcomes, we derive from Equation~\eqref{eq:yi2} a reduced-form specification that is linear in parameters. Therefore, any network fixed effects in Equation~\eqref{eq:alpha} can be partialled out by demeaning the reduced form within each network.

From Equation~\eqref{eq:yi2}, the reduced-form equation for non-isolated individuals can be written as:
\begin{equation}\label{eq:yred}
    y_{m,i} = \tilde c + \theta_1 \bar y_{m,i} + \theta_2 \check{y}_{m,i} + \mathbf{x}_{m,i}^{\prime}\boldsymbol{\theta}_3 + \bar{\mathbf{x}}_{m,i}^{\prime}\boldsymbol{\theta}_4 + \tilde \varepsilon_{m,i},
\end{equation}
where $\tilde c = \frac{c}{1 + \beta^l}$, $\theta_1 = \frac{\beta^l}{1 + \beta^l}$, $\theta_2 = \frac{\beta^h - \beta^l}{1 + \beta^l}$, $\boldsymbol{\theta}_3 = \frac{\boldsymbol{\gamma}_1}{1 + \beta^l}$, $\boldsymbol{\theta}_4 = \frac{\boldsymbol{\gamma}_2}{1 + \beta^l}$, and $\tilde \varepsilon_{m,i} = \frac{\varepsilon_{m,i}}{1 + \beta^l}$. For isolated individuals, the reduced form simplifies to:
\begin{equation}\label{eq:yrediso}
y_{m,i} = c + \mathbf{x}_{m,i}^{\prime}\boldsymbol{\gamma}_1 + \varepsilon_{m,i}.
\end{equation}

Equations \eqref{eq:yred} and \eqref{eq:yrediso} can be combined into a single linear equation by introducing a dummy variable that indicates whether individual $i$ is isolated. However, for ease of exposition, we assume that there are no isolated individuals and focus on Equation~\eqref{eq:yred} in the remainder of this section.


Equation~\eqref{eq:yred} includes two endogenous regressors: the conventional average outcome among peers, $\bar{y}_{m,i}$, and a new endogenous variable, $\check{y}_{m,i}$. This makes the endogeneity problem more complex than in the standard model, which includes only $\bar{y}_{m,i}$. Additionally, since $\check{y}_{m,i}$ is directly linked to the dependent variable $y_{m,i}$, failing to address this endogeneity can lead to substantial bias.


\subsection{Identification}
Identifying peer effects is challenging because of the reflection problem, which arises when individuals and their peers influence one another simultaneously \citep{Manski1993}. In the standard linear model, \citet{Bramoulle2009} show that flexible restrictions on the network structure can be imposed to alleviate this problem. Unfortunately, such restrictions are generally unavailable when individual and peer behaviors are related nonlinearly. Most studies addressing this nonlinearity rely on standard rank conditions for identification, typically imposed within a generalized method of moments (GMM) framework \citep[see, e.g.,][]{brock2007identification,lee_binary_2014,boucher2023}. However, these conditions are difficult to verify empirically because they depend on unobserved variables.

In this paper, we establish identification under relatively mild conditions that differ from the standard GMM-based rank conditions. In addition to the restriction on the network structure proposed by \citet{Bramoulle2009}, our identification conditions require a novel exclusion restriction. We show that this restriction is likely to be satisfied because $\check{y}_{m,i}$ is nonlinearly related to peer outcomes.\footnote{Recall that $\check{y}_{m,i} = \sum_{j \ne i} g_{m,ij}\, \mathbbm{1}\{y_{m,j} > y_{m,i}\} (y_{m,j} - y_{m,i})$.}

For notational ease, we denote by $\mathbb E_m$ the conditional expectation given $\mathbf X_m$ and $\mathbf G_m$. We also define $$\mathbf A_m = \left[\mathbb E_m(\check{\mathbf y}_m), \, \mathbf G_m \mathbb E_m(\check{\mathbf y}_m),\, \mathbf 1_{n_m},\, \mathbf X_m,\, \mathbf G_m \mathbf X_m, \,\mathbf G_m^2 \mathbf X_m\right],$$
where $\check{\mathbf y}_m = (\check y_{m,1}, \, \dots, \, \check y_{m,n_m})^{\prime}$, $\mathbf X_m = (\mathbf x_{m,1}, \, \dots, \, \mathbf x_{m,n_m})^{\prime}$.
\begin{assumption}\label{assumption:data}~\hfill~
\begin{enumerate}[label=(\roman*), ref=\theassumption(\roman*), align=left, leftmargin=7pt, itemsep=0pt, topsep=0pt]
    \item The sequence $(n_m, \mathbf y_m, \mathbf X_m, \mathbf G_m, \boldsymbol \varepsilon_m)_{m\ge1}$ is i.i.d. and satisfies equation \eqref{eq:yred} for any $m$. \label{assumption:data:process}
    \item For all $m$ and $i$, $\mathbb{E}[\varepsilon_{m,i} \mid \mathbf X_m, \mathbf G_m] = 0$. \label{assumption:data:exogeneity}
\end{enumerate}
\end{assumption}
\begin{assumption}\label{assumption:ident}~\hfill~
\begin{enumerate}[label=(\roman*), ref=\theassumption(\roman*), align=left, leftmargin=7pt, itemsep=0pt, topsep=0pt]
    \item The reduced form parameters satisfy $\theta_1\boldsymbol{\theta}_3 + \boldsymbol{\theta}_4 \ne \mathbf 0$. \label{assumption:ident:nonzero}
    \item The matrix $\frac{1}{M}\sum_{m = 1}^M \mathbf A_m^{\prime}\mathbf A_m$ is of full rank. \label{assumption:ident:fullrank}
\end{enumerate}
\end{assumption}

Assumption~\ref{assumption:data} imposes standard conditions on the data-generating process. Assumption~\ref{assumption:data:process} requires the data sequence to be i.i.d. across $m$, while allowing for flexible dependence structures within networks. We include $n_m$ in the sequence because network size may vary across observations. Specifically, the assumption implies that $(\mathbf y_m, \mathbf X_m, \mathbf G_m, \boldsymbol \varepsilon_m)$ is i.i.d. conditional on $n_m$, and that $n_m$ itself is i.i.d. This assumption allows us to establish the equivalence between empirical averages across networks and their population counterparts using standard laws of large numbers. Assumption~\ref{assumption:data:exogeneity} requires $\mathbf X_m$ and $\mathbf G_m$ to be exogenous. We maintain this assumption to abstract from endogenous network formation.

Assumption~\ref{assumption:ident} introduces our identification restrictions. The restriction in Assumption~\ref{assumption:ident:nonzero} holds in many settings when $\mathbf{x}_{m,i}$ includes multiple variables. This restriction is also required for the standard linear model \cite[see][]{Bramoulle2009}. Assumption~\ref{assumption:ident:fullrank} is our new rank condition. Since $\mathbf{A}_m$ does not include the standard average peer outcome, but instead $\check{\mathbf{y}}_m$, which is nonlinearly related to the outcome, this restriction is easier to motivate than classical GMM rank conditions.

Specifically, Assumption~\ref{assumption:ident:fullrank} can be decomposed into two parts. First, it requires $\mathbf 1_{n_m}$, $\mathbf X_m$, $\mathbf G_m \mathbf X_m$, and $\mathbf G_m^2 \mathbf X_m$ to be linearly independent, which can be verified from the data because these variables are all observed. This requirement is related to the identification of the standard model. It holds when many individuals have friends of friends who are not direct friends \citep[see][]{Bramoulle2009}. Second, Assumption~\ref{assumption:ident:fullrank} also requires that no nonzero linear combination of $\mathbb E_m(\check{\mathbf y}_m)$ and $\mathbf G_m \mathbb E_m(\check{\mathbf y}_m)$ lies in the linear span of $\mathbf 1_{n_m}$, $\mathbf X_m$, $\mathbf G_m \mathbf X_m$, and $\mathbf G_m^2 \mathbf X_m$. While it is difficult to provide low-level conditions ensuring this restriction, the presence of the indicator function in $\check{\mathbf y}_m$ makes it unlikely that $\mathbb E_m(\check{\mathbf y}_m)$ and $\mathbf G_m \mathbb E_m(\check{\mathbf y}_m)$ lie in this linear span.\footnote{Additionally, $\mathbb E_m(\check{\mathbf y}_m)$ and $\mathbf G_m \mathbb E_m(\check{\mathbf y}_m)$ must not be collinear. This condition also generally holds when some friends of friends are not direct friends, because $\mathbf G_m \mathbb E_m(\check{\mathbf y}_m)$ involves friends of friends, whereas $\mathbb E_m(\check{\mathbf y}_m)$ depends only on direct friends.}

We establish the following result.
\begin{proposition}\label{prop:ident}
    Under Assumptions \ref{assumption:minbeta}--\ref{assumption:ident}, $\beta^l$, $\beta^h$, $c$, $\boldsymbol\gamma_1$, and $\boldsymbol\gamma_2$ are point identified.
\end{proposition}

The proof of Proposition \ref{prop:ident} is provided in Appendix \ref{append:proof:Ident}.

Let $\mathbf{V}_{m} = [\mathbf G_m \mathbf y_m, \, \check{\mathbf y}_{m}, \, \mathbf 1_{n_m}, \,\mathbf{X}_{m}, \,\mathbf G_m \mathbf{X}_{m}]$ denote the matrix of regressors in~\eqref{eq:yred}, and let $\mathbf{V}_{m}^e = \mathbb{E}_m(\mathbf{V}_{m})$ denote its conditional expectation given $\mathbf{X}_m$ and $\mathbf{G}_m$. As shown by \cite{chamberlain1987}, $\mathbf{V}_{m}^{e}$ corresponds to the optimal instrument for $\mathbf{V}_m$. However, $\mathbf{V}_{m}^{e}$ is infeasible because $\mathbb{E}_m(\mathbf y_{m})$ and $\mathbb{E}_m(\check{\mathbf y}_{m})$ are unobserved. In the next section, we discuss how to construct valid instruments for $\mathbf G_m \mathbf y_m $ and $\check{\mathbf y}_{m}$.




\subsection{Estimation} \label{sec:estim}
We propose an instrumental variable (IV) approach to estimate the model parameters. Since $\mathbb{E}_m(\bar{y}_{m,i})$ and $\mathbb{E}_m(\check{y}_{m,i})$ are unobserved and cannot be used as instruments, we instead construct their predictors that serve as instruments. Importantly, these predictors do not need to be consistent estimators. Any exogenous variables that are sufficiently informative about $\mathbb{E}_m(\bar{y}_{m,i})$ and $\mathbb{E}_m(\check{y}_{m,i})$ can serve as valid instruments. Throughout, we use the notation $\hat{\mathbb{E}}_m$ to denote a predicted counterpart of $\mathbb{E}_m$.

From Equation~\eqref{eq:yred}, for any network $m$, $\mathbb{E}_m(\mathbf{y}_m)$ can be written as:\footnote{See the derivation of $\mathbb{E}_m(\mathbf{y}_m)$ in Equation~\eqref{eq:AppendEys} in Appendix~\ref{append:proof:Ident}. Since $\lvert \theta_1 \rvert < 1$ (Assumption~\ref{assumption:minbeta}), we can express $(\mathbf{I}_{n_m} - \theta_1 \mathbf{G}_m)^{-1}$ using the Neumann series expansion $\sum_{k = 0}^{\infty} \theta_1^k \mathbf{G}_m^k$. Moreover, because there are no isolated individuals, $\mathbf{G}_m \mathbf{1}_s = \mathbf{1}_s$, which implies $\sum_{k = 0}^{\infty} \theta_1^k \mathbf{G}_m^k \mathbf{1}_s = \frac{1}{1 - \theta_1} \mathbf{1}_s$.}
\begin{equation}\label{eq:Eys}
  \mathbb E_m (\mathbf y_m) = \dfrac{\tilde c}{1 - \theta_1} \mathbf 1_{n_m}
  + \mathbf X_m\boldsymbol\theta_3
  + \sum_{k = 0}^{\infty} \theta_1^k \theta_2 \mathbf G_m^k \mathbb E_m (\check{\mathbf y}_m)
  + \sum_{k = 0}^{\infty} \theta_1^k \mathbf G_m^{k + 1} \mathbf X_m(\theta_1 \boldsymbol\theta_3 + \boldsymbol\theta_4).
\end{equation}
This representation suggests that $\check{\mathbf W}_m = [\mathbf X_m,~ \mathbf G_m \mathbf X_m,~ \mathbf G_m^2 \mathbf X_m]$ provides a natural set of predetermined predictors of $\mathbf E_m(\mathbf y_m)$.\footnote{Higher powers of $\mathbf G_m$ can be included in $\check{\mathbf W}_m$ for greater accuracy.} Let $\check{\mathbf w}_{m,i}^{\prime}$ denote the $i$-th row of $\check{\mathbf W}_m$.

We begin by predicting $\mathbb{E}_m(\check{y}_{m,i})$. Let $\Delta^y_{m,ji} = y_{m,j} - y_{m,i}$ and let $\mathbb{P}_m$ denote the probability measure conditional on $\mathbf{X}_m$ and $\mathbf{G}_m$. We rely on the following decomposition:
\begin{equation}\label{eq:Eycheck}
    \mathbb{E}_m(\check{y}_{m,i})
    = \sum_{j \ne i} g_{m,ij}\,
    \mathbb{P}_m\{\Delta^y_{m,ji} > 0\} \,
    \mathbb{E}_m(\Delta^y_{m,ji} \mid \Delta^y_{m,ji} > 0),
\end{equation}
where $\mathbb{P}_m\{\Delta^y_{m,ji} > 0\}$ can be interpreted as an extensive margin and $\mathbb{E}_m(\Delta^y_{m,ji} \mid \Delta^y_{m,ji} > 0)$ is the corresponding intensive margin. We adopt this decomposition because the two margins are easier to predict than $\mathbb{E}_m(\check{y}_{m,i})$ directly. Moreover, because both margins are defined for each pair of connected students, we can exploit a large number of dyadic observations to estimate them, which improves prediction accuracy. To construct $\hat{\mathbb{E}}_m(\check{y}_{m,i})$, we replace each margin with its predicted counterpart. We use flexible supervised machine learning methods to predict each margin. Throughout the paper, we employ random forests, although any supervised learning method could be used. For both prediction tasks, $\check{\mathbf w}_{m,j} - \check{\mathbf w}_{m,i}$ serves as the predictor vector, whereas the target variable is $\mathbbm{1}\{\Delta^y_{m,ji} > 0\}$ for the extensive margin and $\Delta^y_{m,ji}$ for the intensive margin.

To ensure that the resulting predictions are exogenous, we use a cross-fitting procedure across networks \citep{Chernozhukov2018Double}. We randomly split the networks into $L \geq 2$ folds, $\mathcal{F}_1, \dots, \mathcal{F}_L$. For any fold $\mathcal{F}_l$, the random forest model is trained using observations from $\mathcal{F}_{-l}$.\footnote{We use $\mathcal{F}_{-l}$ to denote the subsample of all observations excluding those in $\mathcal{F}_l$.} Once the model is trained, only predetermined predictors from $\mathcal{F}_l$ are used to generate predictions for $\mathcal{F}_l$. Thus, no endogenous variables from $\mathcal{F}_l$ enter either the training or prediction steps for that fold. Since the networks are independent and randomly split, this procedure ensures that the predictions for any network $m$ are independent of $\boldsymbol{\varepsilon}_m$.\footnote{Furthermore, for the intensive margin, we train the random forest on the subsample satisfying $y_{m',j} > y_{m',i}$, where $m' \in \mathcal{F}_{-l}$. This allows the prediction to be interpreted as a conditional expectation given $y_{m,j} > y_{m,i}$, consistent with the definition of the intensive margin.} For completeness, we provide a detailed description of the prediction procedure in Supplemental Appendix~\ref{append:instrument}.


For the prediction of $\mathbb{E}_m(\bar{y}_{m,i})$, we define
$$
\bar{\mathbf{W}}_m = [\mathbf{G}_m \mathbf{X}_m,\, \mathbf{G}_m^2 \mathbf{X}_m,\, \mathbf{G}_m \hat{\mathbb{E}}_m(\check{\mathbf{y}}_m),\, \mathbf{G}_m^2 \hat{\mathbb{E}}_m(\check{\mathbf{y}}_m)].
$$
We also denote by $\bar{\mathbf{w}}_{m,i}^{\prime}$ the $i$-th row of $\bar{\mathbf{W}}_m$. Premultiplying Equation~\eqref{eq:Eys} by $\mathbf{G}_m$ implies that $\bar{\mathbf{w}}_{m,i}^{\prime}$ is a natural set of predetermined predictors for $\mathbb{E}_m(\bar{y}_{m,i})$. Notably, $\bar{\mathbf{W}}_m$ includes the standard instruments $\mathbf{G}_m \mathbf{X}_m$ and $\mathbf{G}_m^2 \mathbf{X}_m$ used in the linear model \citep[see][]{Bramoulle2009}, as well as additional predictors to accommodate asymmetries in peer effects. We predict $\mathbb{E}_m(\bar{y}_{m,i})$ using random forests, where $\bar{\mathbf{w}}_{m,i}$ serves as the predictor and the target variable is $\bar{y}_{m,i}$. We also use cross-fitting across networks so that the resulting predictions are independent of $\boldsymbol{\varepsilon}_m$.


Equipped with the predictions $\hat{\mathbb{E}}_m(\check{y}_{m,i})$ and $\hat{\mathbb{E}}_m(\bar{y}_{m,i})$, we define the matrix of instruments as
$\hat{\mathbf{Z}}_m = \big[ \hat{\mathbb{E}}_m(\check{\mathbf{y}}_{m}), \, \hat{\mathbb{E}}_m(\bar{\mathbf{y}}_{m}), \, \mathbf 1_{n_m}, \, \mathbf{X}_m, \, \mathbf{G}_m \mathbf{X}_m \big]$.
The model parameters are then estimated by solving the moment condition
$$
\mathbb E \left( \hat{\mathbf{Z}}_m^{\prime} \boldsymbol{\varepsilon}_m \right)= \mathbf{0},
$$
where $\boldsymbol{\varepsilon}_m = (\varepsilon_{m,1}, \, \dots, \, \varepsilon_{m,n_m})^{\prime}$.

However, since $\hat{\mathbb{E}}_m(\check{\mathbf{y}}_{m})$ and $\hat{\mathbb{E}}_m(\bar{\mathbf{y}}_{m})$ are generated instruments, they act as nuisance parameters, and their uncertainty must be taken into account in inference. We show that the moment function is orthogonal to these nuisance parameters, and therefore their influence can be ignored. This is a particularly strong form of orthogonality, as it holds at any order \citep[see][]{mackey2018orthogonal, bonhomme2026higher}, thereby allowing the use of highly flexible machine learning methods without requiring convergence rates. As a result, our inference relies on weak regularity conditions stated in Assumption~\ref{assumption:iv} below.

Let $\check{f}_m$ and $\bar{f}_m$ be two nonstochastic functions mapping $(\mathbf X_m, \mathbf G_m)$ into $\mathbb R^{n_m}$. Let $$\mathbf Z_m = [\check f_m(\mathbf X_m, \mathbf G_m),\, \bar f_m(\mathbf X_m, \mathbf G_m),\, \mathbf 1_{n_m},\, \mathbf X_m,\, \mathbf G_m \mathbf X_m]$$ denote the asymptotic equivalent of $\hat{\mathbf Z}_m$.
\begin{assumption}\label{assumption:iv}~\hfill~
\begin{enumerate}[label=(\roman*), ref=\theassumption(\roman*), align=left, leftmargin=7pt, itemsep=0pt, topsep=0pt]
    \item $\displaystyle \max_{m} \lVert \hat{\mathbb E}_m(\check{\mathbf y}_{m}) - \check f_m(\mathbf X_m, \mathbf G_m)\rVert = o_p(1)$ and $\displaystyle \max_{m} \lVert \hat{\mathbb E}_m(\bar{\mathbf y}_{m}) - \bar f_m(\mathbf X_m, \mathbf G_m)\rVert = o_p(1)$. \label{assumption:iv:ml}
    \item The matrix $\frac{1}{M} \sum_{m = 1}^M \mathbf Z_m^{\prime}\mathbf V_m$ is asymptotically nonsingular and nonstochastic. \label{assumption:iv:fullrank}
    \item For some $\nu > 0$, $\displaystyle \max_m \lVert \mathbb E_m(\boldsymbol \varepsilon_m \boldsymbol \varepsilon_m^{\prime}) \rVert^{2 + \nu} < \infty$. \label{assumption:iv:variance}
\end{enumerate}
\end{assumption}

Assumption~\ref{assumption:iv:ml} requires that $\hat{\mathbb{E}}_m(\check{\mathbf{y}}_{m})$ and $\hat{\mathbb{E}}_m(\bar{\mathbf{y}}_{m})$ converge to nonstochastic quantities, $\check{f}_m(\mathbf{X}_m, \mathbf{G}_m)$ and $\bar{f}_m(\mathbf{X}_m, \mathbf{G}_m)$, respectively, conditional on $\mathbf{X}_m$ and $\mathbf{G}_m$. Notably, no rate of convergence is required, and $\check{f}_m(\mathbf{X}_m, \mathbf{G}_m)$ and $\bar{f}_m(\mathbf{X}_m, \mathbf{G}_m)$ need \textit{not} coincide with the estimands $\mathbb{E}_m(\bar{y}_{m,i})$ or $\mathbb{E}_m(\check{y}_{m,i})$, so the random forest predictions may be asymptotically biased. Nevertheless, the predictions must be sufficiently correlated with the endogenous variables $\bar{y}_{m,i}$ and $\check{y}_{m,i}$ to avoid weak instrument problems. This requirement is reflected in the nonsingularity condition in Assumption~\ref{assumption:iv:fullrank}. In practice, we assess instrument strength using standard weak instrument tests. We also examine whether $\frac{1}{M} \sum_{m=1}^M \hat{\mathbf{Z}}_m^{\prime} \mathbf{V}_m$ is nonsingular using rank tests \citep[e.g.,][]{kleibergen200}. Finally, Assumption~\ref{assumption:iv:variance} imposes a standard finite-moment condition required for the central limit theorem.

Let $\boldsymbol\theta = (\theta_1, \, \theta_2, \,  \tilde c, \, \boldsymbol{\theta}_3^{\prime}, \, \boldsymbol{\theta}_4^{\prime})^{\prime}$ be the vector of reduced-form parameters and $\boldsymbol{\hat \theta}$ be its estimator. Let also $\boldsymbol{\theta}_0$ denote the true value of $\boldsymbol{\theta}$. We establish the following result.
\begin{proposition}\label{prop:consistency}
     Under Assumptions~\ref{assumption:minbeta}--\ref{assumption:iv}, $\boldsymbol{\hat \theta}$ converges in probability to $\boldsymbol{\theta}_0$, and $\sqrt{M}(\boldsymbol{\hat \theta} - \boldsymbol{\theta}_0) \overset{d}{\to} N(\mathbf 0, \, \boldsymbol\Sigma)$, where $\boldsymbol\Sigma$ is given in Appendix~\ref{append:prop:consistency}.
\end{proposition}
\noindent The proof of Proposition~\ref{prop:consistency} is provided in Appendix~\ref{append:prop:consistency}. This result also extends to the structural parameters. The estimators are consistent and asymptotically normally distributed. The asymptotic variance can be obtained using the Delta method.


\subsection{Bias of the Standard Symmetric Model}\label{sec:biasSymmetric}
Given that the standard symmetric model is widely used in empirical studies, we investigate what its single peer effect parameter, $\beta$, measures when the data exhibits asymmetric peer effects. One might expect the estimate of $\beta$ to be a simple weighted average of the estimates of $\beta^l$ and $\beta^h$, which would imply that $\beta$ lies asymptotically between $\beta^l$ and $\beta^h$. However, we show that the weights in this average can be negative, implying that $\beta$ may lie outside the range defined by $\beta^l$ and $\beta^h$. This issue is common in settings where a single parameter aggregates heterogeneous effects. A well-known example is the heterogeneous treatment effects problem studied by \cite{de2020two}.

In the symmetric model \citep{blume_linear_2015}, the outcome is modeled as:
\begin{equation}\label{eq:LIM}
  \mathbf y_{m} = \delta_1 \bar{\mathbf y}_{m}
  + \mathbf{R}_{m} \boldsymbol{\delta}_2
  + \boldsymbol\eta_{m},
\end{equation}
where $\delta_1 = \frac{\beta}{1+\beta}$, $\mathbf R_m = (\mathbf 1_{n_m}, \, \mathbf X_m, \, \mathbf G_m \mathbf X_m)$ denotes the matrix of control variables, and $\boldsymbol\eta_{m}$ is an error term. The parameters $\beta$ and $\boldsymbol{\delta}_2$ can be estimated by GMM, where $\bar{\mathbf y}_{m}$ is instrumented by $\hat{\mathbb E}_m(\bar{\mathbf y}_{m})$, or by any other valid instrument, such as $\mathbf G_m^2 \mathbf X_m$ \citep[see][]{Bramoulle2009}. The resulting moment condition is
$$
\mathbb E\big(\tilde{\mathbf{Z}}_m^{\prime} (\mathbf y_{m} - \delta_1 \bar{\mathbf y}_{m} - \mathbf{R}_{m} \boldsymbol{\delta}_2)\big) = \mathbf{0},
$$
where $\tilde{\mathbf{Z}}_m = \big[\hat{\mathbb{E}}_m(\bar{\mathbf{y}}_{m}), \, \mathbf R_m \big]$ is the matrix of instruments.

For the rest of this section, we add a superscript $h$ to $\check y_{m,i}$; that is, we write $\check y_{m,i}^h = \bar y_{m,i}^h - g_{m,i}^h y_{m,i}$. We also define the analogous variable $\check y_{m,i}^l = \bar y_{m,i}^l - g_{m,i}^l y_{m,i}$, where $\bar{y}_{m,i}^l = \sum_{j \ne i} \mathbbm{1}\{y_j \leq y_i\} g_{ij} y_j$ and $g_{m,i}^l = \sum_{j \ne i} \mathbbm{1}\{y_j \leq y_i\} g_{ij}$. Let $\check{\mathbf y}_m^h = (\check y_{m,1}^h, \, \dots, \, \check y_{m,n_m}^h)^{\prime}$ and $\check{\mathbf y}_m^l = (\check y_{m,1}^l, \, \dots, \, \check y_{m,n_m}^l)^{\prime}$.

We consider the following auxiliary moment conditions:
\allowdisplaybreaks
\begin{align}\label{eq:auxiliaryIV}
\begin{split}
    \mathbb E \left( \tilde{\mathbf{Z}}_m^{\prime} \big( \check{\mathbf y}_m^h - \delta^h_1 \bar{\mathbf y}_m - \mathbf R_m\boldsymbol{\delta}^h_2\big) \right) &= \mathbf{0},\\
    \mathbb E \left( \tilde{\mathbf{Z}}_m^{\prime} \big( \check{\mathbf y}_m^l - \delta^l_1 \bar{\mathbf y}_m - \mathbf R_m\boldsymbol{\delta}^l_2 \big)\right) &= \mathbf{0},
\end{split}
\end{align}
for some unknown parameters $\delta^h_1$, $\boldsymbol{\delta}^h_2$, $\delta^l_1$, and $\boldsymbol{\delta}^l_2$. Specifically, $\delta^h_1$ is the coefficient from the IV regression of $\check{\mathbf y}_m^h$ on $\bar{\mathbf y}_m$, partialling out $\mathbf R_m$. Similarly, $\delta^l_1$ is the coefficient from the IV regression of $\check{\mathbf y}_m^l$ on $\bar{\mathbf y}_m$, partialling out $\mathbf R_m$.

\begin{proposition}\label{prop:biasSymmetric}
Under Assumptions~\ref{assumption:minbeta}--\ref{assumption:iv}, we have:
$$
\frac{\beta}{1+\beta} = \frac{\beta^l}{1+\beta^l} + \delta^h_1 \dfrac{\beta^h - \beta^l}{1+\beta^l}
\quad \text{and} \quad
\frac{\beta}{1+\beta} = \frac{\beta^h}{1+\beta^h} + \delta^l_1 \dfrac{\beta^l - \beta^h}{1+\beta^h}.
$$
\end{proposition}
The proof of Proposition~\ref{prop:biasSymmetric} is provided in Appendix~\ref{append:prop:biasSymmetric}.

A direct implication of Proposition~\ref{prop:biasSymmetric} is that the necessary and sufficient conditions for $\beta$ to lie in the range defined by $\beta^l$ and $\beta^h$ are that $\delta^l_1 \geq 0$ and $\delta^h_1 \geq 0$.\footnote{Since $f(\beta) = \frac{\beta}{1+\beta}$ is strictly increasing in $\beta$, saying that $\beta$ lies in the range defined by $\beta^l$ and $\beta^h$ is equivalent to saying that $f(\beta)$ lies in the range defined by $f(\beta^l)$ and $f(\beta^h)$ } For example, if $\beta^l < \beta^h$, then $\frac{\beta^l}{1+\beta^l} \leq \frac{\beta}{1+\beta} \leq \frac{\beta^h}{1+\beta^h}$ if and only if $\delta^l_1 \geq 0$ and $\delta^h_1 \geq 0$. A similar argument applies when $\beta^l > \beta^h$.

If $\delta^l_1$ or $\delta^h_1$ is negative, then $\beta$ falls outside the range defined by $\beta^l$ and $\beta^h$, meaning that $\beta$ is a weighted average of $\beta^l$ and $\beta^h$ with negative weights.\footnote{Proposition~\ref{prop:biasSymmetric} implies that $\dfrac{\beta}{1 + \beta} = \rho \dfrac{\beta^l}{1 + \beta^l} + (1 - \rho)\dfrac{\beta^h}{1 + \beta^h}$, where $\rho = \dfrac{\delta^l_1(1 + \beta^l)}{\delta^l_1(1 + \beta^l) + \delta^h_1(1 + \beta^h)}$. Since $f(\beta)$ is strictly increasing in $\beta$, this means that $\beta$ can also be written as an average of $\beta^l$ and $\beta^h$.} To see when this can happen, note, for example, that a negative $\delta^l_1$ implies that an increase in $\bar{y}_{m,i}$ has a larger effect on $g_{m,i}^l y_{m,i}$ than on $\bar y_{m,i}^l$. This suggests that this increase is mainly driven by $\bar y_{m,i}^h$, while $\bar y_{m,i}^l$ remains nearly unchanged. This situation arises in certain networks where higher-performing peers are predominant, and there are few or no (direct or indirect) connections from these peers to lower-performing ones. This is, for example, the case in certain segregated networks where friendships are not always reciprocated and often run from lower-status individuals to higher-status individuals \citep[e.g.,][]{ball2013friendship}. We consider such network structures in our simulation study and show that $\beta$ lies outside the range defined by $\beta^l$ and $\beta^h$.


\section{Monte Carlo Simulations}\label{sec:MCsimu}

We conduct a simulation study to assess the finite-sample performance of our estimation method. We consider a sample of $M = 50$ networks, each comprising $n_m = 50$ individuals. Each individual $i$ in network $m$ is characterized by two exogenous variables, $x_{1,m,i}$ and $x_{2,m,i}$. Productivity is defined as in Equation~\eqref{eq:alpha}, except that we introduce an unobserved network-level shock $c_m$:
\begin{equation*}
    \alpha_{m,i} = c_m + \gamma_{1,1} x_{1,m,i} + \gamma_{1,2} x_{2,m,i} + \gamma_{2,1} \bar{x}_{1,m,i} + \gamma_{2,2} \bar{x}_{2,m,i} + \varepsilon_{m,i},
\end{equation*}
where $c_m \sim \mathcal{N}(10, \,1)$ and $\varepsilon_{m,i} \sim \mathcal{N}(0, \,1)$. We simulate $x_{1,m,i}$ to be correlated with $c_m$, so that $c_m$ acts as a network fixed effect.

We consider both fully random and segregated network formation. In the case of random networks, the degree (number of friends) is randomly drawn from the empirical distribution observed in the Add Health data used in the empirical application. Individuals can have up to 10 friends, with an average of 3.47. Given the degree, links are formed uniformly at random within the network. We simulate $x_{1,m,i}$ and $x_{2,m,i}$ from $\text{Uniform}(c_m - 2, \,c_m + 2)$ and $\text{Poisson}(2)$, respectively, and set $\gamma_{1,1} = 1.4$, $\gamma_{1,2} = -0.8$, $\gamma_{2,1} = 0.7$, and $\gamma_{2,2} = -0.5$.

Under the segregated network, to easily handle peers’ relative status, we employ a model without contextual effects for both the data-generating process (DGP) and the estimation specification. We thus impose that $\gamma_{2,1} = 0$ and $\gamma_{2,2} = 0$, and keep the other parameters at their values in the random network case. To introduce segregation, we also change the way we simulate $x_{1,m,i}$. We randomly split the 50 individuals in each network $m$ into three groups, $\mathcal{G}_{m,1}$, $\mathcal{G}_{m,2}$, and $\mathcal{G}_{m,3}$, comprising 10, 10, and 30 individuals, respectively. For individuals in $\mathcal{G}_{m,1}$ and $\mathcal{G}_{m,2}$, we simulate $x_{1,m,i} \sim \text{Uniform}(c_m - 2, \, c_m + 2)$, while for those in $\mathcal{G}_{m,3}$, we simulate $x_{1,m,i} \sim \text{Uniform}(c_m + 10, \, c_m + 20)$.


Individuals in $\mathcal{G}_{m,1}$ form reciprocal links with those in $\mathcal{G}_{m,2}$. Individuals in $\mathcal{G}_{m,2}$ further form non-reciprocal links with those in $\mathcal{G}_{m,3}$, whereas individuals in $\mathcal{G}_{m,3}$ have no friends. For simplicity, the degrees from $\mathcal{G}_{m,1}$ to $\mathcal{G}_{m,2}$ and from $\mathcal{G}_{m,2}$ to $\mathcal{G}_{m,3}$ are independent and randomly drawn from the empirical Add Health distribution. Generating networks in this way ensures that higher-performing peers---who include most friends from $\mathcal{G}_{m,3}$---are predominant and have fewer connections with lower-performing peers. We therefore expect $\delta_1^l$ in Proposition~\ref{prop:biasSymmetric} to be negative for $\beta$ to fall outside the range defined by $\beta^l$ and $\beta^h$.

For each type of network, we consider four DGPs with different values of $\beta^l$ and $\beta^h$. In DGP 1, $\beta^l = \beta^h = 1$, i.e., peer effects are symmetric. In DGP 2, $\beta^l = 0.4$ and $\beta^h = 2.6$, while in DGP 3, $\beta^l = 2.6$ and $\beta^h = 0.4$. In DGP 4, we set $\beta^l = -0.4$ and $\beta^h = 2.6$. We allow for anti-conformity toward low-performing peers in DGP 4 because this pattern arises in our empirical application, although it is not statistically significant. For each DGP, we estimate both the asymmetric and the symmetric models. Table~\ref{tab:MonteCarlo} summarizes the simulation results.

\begin{table}[htbp]
\centering
\setlength{\tabcolsep}{4pt}
\begin{threeparttable}
\caption{Simulation Results}
\label{tab:MonteCarlo}
\footnotesize
\begin{tabular}{ld{4}ld{4}ld{4}ld{4}l}
\toprule
                   & \multicolumn{4}{c}{Random Network}                                       & \multicolumn{4}{c}{Segregated Network}                               \\
Model              & \multicolumn{2}{c}{Asymmetry}       & \multicolumn{2}{c}{Symmetry}       & \multicolumn{2}{c}{Asymmetry}     & \multicolumn{2}{c}{Symmetry}     \\
\midrule
\multicolumn{9}{c}{DGP 1: $\beta^l = 1.00$,   $\beta^h = 1.00$}  \\[1ex]
$\beta^l$      & 1.005         & (0.045)       &              &               & 1.009         & (0.091)       &              &               \\
$\beta^h$      & 0.998         & (0.134)       &              &               & 1.001         & (0.019)       &              &               \\
$\beta$        &               &               & 1.003        & (0.035)       &               &               & 0.999        & (0.014)       \\
$\gamma_{1,1}$ & 1.401         & (0.031)       & 1.401        & (0.029)       & 1.400         & (0.005)       & 1.400        & (0.005)       \\
$\gamma_{1,2}$ & -0.801        & (0.023)       & -0.801       & (0.021)       & -0.801        & (0.017)       & -0.800       & (0.017)       \\
$\gamma_{2,1}$ & 0.699         & (0.036)       & 0.699        & (0.035)       &               &               &              &               \\
$\gamma_{2,2}$ & -0.500        & (0.028)       & -0.500       & (0.028)       &               &               &              &               \\
\midrule
\multicolumn{9}{c}{DGP 2: $\beta^l = 0.40$,   $\beta^h = 2.60$}  \\[1ex]
$\beta^l$      & 0.403         & (0.030)       &              &               & 0.411         & (0.086)       &              &               \\
$\beta^h$      & 2.600         & (0.170)       &              &               & 2.603         & (0.045)       &              &               \\
$\beta$        &               &               & 0.757        & (0.036)       &               &               & 3.683        & (0.106)       \\
$\gamma_{1,1}$ & 1.401         & (0.031)       & 1.271        & (0.030)       & 1.400         & (0.005)       & 1.401        & (0.005)       \\
$\gamma_{1,2}$ & -0.801        & (0.021)       & -0.716       & (0.021)       & -0.801        & (0.017)       & -0.811       & (0.018)       \\
$\gamma_{2,1}$ & 0.700         & (0.036)       & 0.648        & (0.041)       &               &               &              &               \\
$\gamma_{2,2}$ & -0.500        & (0.028)       & -0.448       & (0.033)       &               &               &              &               \\
\midrule
\multicolumn{9}{c}{DGP 3: $\beta^l = 2.60$,   $\beta^h = 0.40$}  \\[1ex]
$\beta^l$      & 2.610         & (0.085)       &              &               & 2.609         & (0.123)       &              &               \\
$\beta^h$      & 0.395         & (0.119)       &              &               & 0.400         & (0.009)       &              &               \\
$\beta$        &               &               & 1.817        & (0.063)       &               &               & 0.327        & (0.008)       \\
$\gamma_{1,1}$ & 1.402         & (0.035)       & 1.414        & (0.033)       & 1.400         & (0.005)       & 1.392        & (0.005)       \\
$\gamma_{1,2}$ & -0.800        & (0.027)       & -0.819       & (0.024)       & -0.800        & (0.018)       & -0.739       & (0.016)       \\
$\gamma_{2,1}$ & 0.699         & (0.036)       & 0.816        & (0.046)       &               &               &              &               \\
$\gamma_{2,2}$ & -0.500        & (0.029)       & -0.581       & (0.040)       &               &               &              &               \\
\midrule
\multicolumn{9}{c}{DGP 4: $\beta^l = -0.40$,   $\beta^h = 2.60$} \\[1ex]
$\beta^l$      & -0.398        & (0.022)       &              &               & -0.384        & (0.071)       &              &               \\
$\beta^h$      & 2.590         & (0.167)       &              &               & 2.606         & (0.043)       &              &               \\
$\beta$        &               &               & -0.077       & (0.021)       &               &               & 4.152        & (0.178)       \\
$\gamma_{1,1}$ & 1.398         & (0.034)       & 0.956        & (0.031)       & 1.401         & (0.005)       & 1.402        & (0.005)       \\
$\gamma_{1,2}$ & -0.800        & (0.023)       & -0.532       & (0.022)       & -0.802        & (0.017)       & -0.815       & (0.018)       \\
$\gamma_{2,1}$ & 0.700         & (0.036)       & 0.791        & (0.048)       &               &               &              &               \\
$\gamma_{2,2}$ & -0.500        & (0.028)       & -0.500       & (0.037)       &               &               &              &               \\ \bottomrule
\end{tabular}
\begin{tablenotes}[para,flushleft]\footnotesize
Notes: The models are simulated and estimated 1,000 times. Values reported without parentheses correspond to mean estimates, whereas values in parentheses denote standard deviations. To generate the instruments, we use random forests with 5-fold cross-fitting (see Supplemental Appendix \ref{append:instrument}) and 1,000 trees. For each training step, 20\% of the sample is used to tune the remaining hyperparameters.
\end{tablenotes}
\end{threeparttable}
\end{table}



The results provide strong evidence that, under the asymmetric specification, our estimation method successfully recovers the structural parameters in finite samples for both network formation processes and across all DGPs. For DGP 1, the symmetric specification also performs well and is, unsurprisingly, more precise. Fortunately, once the asymmetric model is estimated, the null hypothesis $\beta^l = \beta^h$ can be easily tested using standard tests. If the null hypothesis is not rejected, the symmetric model can be used to improve efficiency.

For DGPs 2--4, the estimates of the symmetric peer-effect parameter lie between $\beta^l$ and $\beta^h$ when friendships form randomly. Yet, these estimates are overly simplistic and mask important features of the heterogeneous structure of peer effects. For instance, in DGP 4, where there is anticonformity toward lower-performing peers only, the symmetric specification implies anticonformity toward all peers.  For the segregated network, the estimates of $\beta$ always fall outside the range defined by $\beta^l$ and $\beta^h$, as expected. Across all DGPs, either $\beta < \beta^h < \beta^l$ or $\beta > \beta^h > \beta^l$, implying that $\beta$ assigns a negative weight to $\beta^l$ and a weight greater than one to $\beta^h$. This is consistent with $\delta_1^l < 0$ in Proposition~\ref{prop:biasSymmetric}.

Remarkably, the simulation results also reveal that the coefficients of the exogenous variables are biased under the symmetric specification. This bias is particularly pronounced in the randomly formed network, likely due to the inclusion of contextual effects in the specification.\footnote{Following a similar approach as in the proof of Proposition~\ref{prop:biasSymmetric}, we can derive expressions for the bias of $\boldsymbol \gamma_1$ and $\boldsymbol \gamma_2$. However, this bias is less intuitive, as it depends on how $\mathbf{x}_{m,i}$ and $\bar{\mathbf{x}}_{m,i}$ relate to $\check y_{m,i}$.} This result suggests that our asymmetric peer effect model is effective not only in estimating endogenous peer effects but also in estimating exogenous effects, such as treatment effects, in the presence of network interference when asymmetry matters.

\section{Empirical Application} \label{sec:empirics}
This section presents an empirical application using Wave I of the Add Health data. We analyze multiple outcomes to uncover different forms of asymmetry in peer effects that cannot be captured by the standard symmetric model.

\subsection{Add Health Data}
Wave I of the Add Health survey provides nationally representative and detailed information on students in grades 7--12 from 144 schools in the United States during the 1994--1995 school year. Approximately 90,000 students completed an in-school questionnaire covering demographics, family background, academic performance, health-related behaviors, and friendship networks. Each respondent could also nominate up to five male and five female best friends within the same school.

We analyze four outcomes: smoking, fighting, optimism, and drinking. Smoking and drinking are defined as the number of days per week that a student smokes tobacco or consumes alcohol, respectively. Fighting is measured as the number of times per year that a student engages in a physical fight. Optimism is constructed as the average of indicators that reflect whether students believe they will graduate from college and live to age 35. These indicators range from 0 to 8, with 8 indicating certainty about the outcome.

Although Add Health is one of the most comprehensive datasets for studying peer effects, it has some limitations. First, the observed degree may be censored because students can nominate at most 10 friends. Second, some nominated friends cannot be matched to valid student identifiers and are therefore excluded from the network, as is common in studies using this dataset.

Using the standard symmetric model, several studies have shown that degree censoring in this dataset induces only limited attenuation bias, while unmatched links may imply substantial bias \citep[see][]{griffith2022name, boucher2025estimating}. Although formally addressing the missing links problem is beyond the scope of this paper, it is important to note that the unmatched links may lead us to incorrectly classify an individual as isolated when none of their nominated friends can be matched. Because the outcome specifications differ between isolated and non-isolated individuals (see Equations~\eqref{eq:yred} and \eqref{eq:yrediso}), we conduct a robustness check in which we exclude individuals for whom none of their nominated friends can be matched (i.e., false isolates).\footnote{False isolates may still be nominated as friends by other individuals. We therefore do not remove them entirely from the network but exclude them only as focal observations. Consequently, this procedure does not generate additional missing links.} The results, reported in Table~\ref{tab:estimates_wofakeiso} in Supplemental Appendix~\ref{append:empirics}, are similar to those obtained using the full sample.

We control for a rich set of exogenous characteristics, including age, grade, sex, race, Hispanic ethnicity, and mother’s education and employment status. To mitigate the impact of missing links, we also include the number of unmatched friends and the number of matched friends as control variables. Additionally, we control for contextual variables defined as the averages of students' characteristics among friends.

After excluding observations with missing outcome values, our final sample consists of approximately 75{,}000 students from 140 schools. The average number of friends per student is 3.6, and 22\% of students have no friends. Summary statistics of the variables are reported in Supplemental Appendix~\ref{append:empirics}.

\subsection{Empirical Results}


\begin{table}[htbp]
\centering
\footnotesize
\begin{threeparttable}
\caption{Estimation Results}
\label{tab:SumEstimates}
\begin{tabular}{lcccc}
\toprule
\textbf{Outcome} & \multicolumn{2}{c}{\textbf{Smoking}} &
\multicolumn{2}{c}{\textbf{Fighting}} \\
Model & Asymmetric & Symmetric & Asymmetric & Symmetric \\
\midrule
$\beta^l$ & 1.219 & & -0.065 & \\
~ & (0.298) & & (0.087) & \\[1ex]
$\beta^h$ & 3.448 & & 1.115 & \\
~ & (0.449) & & (0.179) & \\[1ex]
$\beta$ & & 2.839 & & 0.232 \\
~ & & (0.458) & & (0.059) \\[1ex]
$\beta^h-\beta^l$ & 2.229 & & 1.180 & \\
~ & (0.281) & & (0.232) & \\[2.5ex]
Total Peer Effect & [0.549, 0.775] & 0.740 &
[-0.069, 0.527] & 0.188 \\[2.5ex]
\nth{1} stage F stat. ($\bar y$) & 88.899 & 171.597 &
42.730 & 64.382 \\
\nth{1} stage F stat. ($\check y$) & 62.752 & &
23.597 & \\[1ex]
KP LM test p-value & 0.000 & 0.000 &
0.002 & 0.000 \\
\end{tabular}
\begin{tabular}{lcccc}
\toprule
\textbf{Outcome} & \multicolumn{2}{c}{\textbf{Optimism}} &
\multicolumn{2}{c}{\textbf{Drinking}} \\
Model & Asymmetric & Symmetric & Asymmetric & Symmetric \\
\midrule
$\beta^l$ & 3.694 & & 0.842& \\
~ & (0.625) & & (0.184) & \\[1ex]
$\beta^h$ & -0.201 & & 1.206 & \\
~ & (0.162) & & (0.296) & \\[1ex]
$\beta$ & & 0.623 & & 0.897 \\
~ & & (0.093) & & (0.172) \\[1ex]
$\beta^h-\beta^l$ & -3.896 & & 0.364 & \\
~ & (0.737) & & (0.285) & \\[2.5ex]
Total Peer Effect & [-0.252, 0.787] & 0.384 &
[0.457, 0.547] & 0.473 \\[2.5ex]
\nth{1} stage F stat. ($\bar y$)  & 84.574 & 103.343 &
25.228 & 40.992
 \\
\nth{1} stage F stat. ($\check y$) & 22.324 & & 18.574 &  \\[1ex]
KP LM test p-value & 0.000 & 0.000 &
0.000 & 0.000 \\
\bottomrule
\end{tabular}
\begin{tablenotes}[para,flushleft]\footnotesize
Notes: Estimates are reported without parentheses, with standard errors (clustered at the network level) shown in parentheses. To generate the instruments, we use random forests with 5-fold cross-fitting (see Supplemental Appendix~\ref{append:instrument}) and 1,500 trees. For each training step, 20\% of the sample is used to tune the remaining hyperparameters. The row labeled ``Total Peer Effect'' corresponds to the range of marginal peer effects, defined by $\frac{\beta^l}{1 + \beta^l}$ and $\frac{\beta^h}{1 + \beta^h}$ (see Equation \eqref{eq:yi1}). The row labeled ``KP LM test p-value” reports the p-value of the LM test of \cite{kleibergen200}. The full table, including the coefficients on own and contextual control variables, is reported in the Supplemental Appendix (Table~\ref{tab:estimates}).
\end{tablenotes}
\end{threeparttable}
\end{table}

We estimate both the asymmetric and symmetric peer effects models for each outcome. For brevity, Table \ref{tab:SumEstimates} reports the estimates of the endogenous peer effects parameters, while the detailed estimation results, including the coefficients on the control variables, are reported in Supplemental Appendix \ref{append:empirics}. To assess instrument strength, we report the F-statistics from the first-stage regressions. We also report the Kleibergen--Paap (KP) rank LM test for underidentification \citep{kleibergen200}. The null hypothesis is that $\frac{1}{M} \sum_{m=1}^M \hat{\mathbf{Z}}_m^{\prime} \mathbf{V}_m$ is rank-deficient (see Assumption \ref{assumption:iv:fullrank}). Overall, these tests confirm that the instruments are relevant and rule out underidentification across all specifications and outcomes.

For three of the outcomes studied---namely, smoking, fighting, and optimism---peer effects exhibit different forms of asymmetry. For smoking, both types of peers exert strong conformity effects, although the influence of friends who smoke more than the individual is larger. Depending on peer composition, a one-unit increase in peers' smoking consumption implies an increase in the individual's smoking consumption ranging from 0.549 to 0.775 units.\footnote{The range of marginal peer effects is defined by $\frac{\beta^l}{1 + \beta^l}$ and $\frac{\beta^h}{1 + \beta^h}$ (see Equation \eqref{eq:yi1}).} The lower bound corresponds to a situation in which all friends smoke less than the individual, whereas the upper bound corresponds to a situation in which all friends smoke more than the individual. Exposure to friends who smoke more than the individual therefore has a larger effect than exposure to friends who smoke less, and this difference is statistically significant. In contrast, the standard symmetric model primarily captures the effect of exposure to friends who smoke more. Given the estimated standard errors, we cannot rule out the possibility that $\beta$ falls outside $[\beta^l,\,\beta^h]$.


Asymmetry is more pronounced for fighting and optimism, with stronger conformity toward higher- and lower-performing friends, respectively. In the case of fighting, the asymmetric model indicates that a one-unit increase in friends' fighting behavior induces peer effects ranging from $-0.069$ to $0.527$, where the lower bound is not statistically different from zero and corresponds to the effects from lower-performing peers. Students imitate friends who are more aggressive than themselves but show little or no response to less aggressive friends. This means that, depending on the status composition of the peer group, students may be either completely unresponsive or strongly responsive to peer behavior. The pattern reverses for optimism, where a one-unit increase in friends' optimism behavior index induces peer effects ranging from $-0.252$ to $0.787$. Here again, the lower bound is not statistically different from zero, but corresponds to the effects from higher-performing peers. Students show no response to friends who are more optimistic than themselves but conform strongly to more pessimistic friends.

However, for both fighting and optimism, the symmetric model masks this heterogeneity by indicating that all types of peers exert the same influence. In particular, this model suggests that students strongly conform to less aggressive friends and to more optimistic friends, although these friends actually exert little or no influence. Moreover, it underestimates the effects of the friends who actually exert strong influence, namely more aggressive friends in the case of fighting and more pessimistic friends in the case of optimism.

Lastly, for drinking, the estimated effects of friends who drink more and friends who drink less are similar in magnitude and statistically indistinguishable. In this case, the standard model provides an adequate description of peer influence, and the symmetry restriction cannot be statistically rejected.


Interestingly, all socially undesirable behaviors---namely smoking, fighting, and drinking---exhibit stronger peer effects from higher-performing peers, although the asymmetry is not statistically significant in the case of drinking. These results suggest that, for such outcomes, the most active individuals are likely the most effective targets for amplifying the effects of policy interventions aimed at mitigating these behaviors through the network. This finding provides additional insight into the results of \citet{lee2021}, who show that key players in networks involving risky behaviors, such as juvenile delinquency, are not necessarily the most active individuals. Because their analysis relies on a standard symmetric model, it does not account for the possibility that the most active individuals, although not central nodes in the network, may exert stronger influence on their peers.

Furthermore, the asymmetry in peer effects has the opposite targeting implication for optimism, a socially desirable outcome. Improving optimism may be important for fostering aspirations and self-esteem, both of which are key determinants of long-term welfare \citep{bernard2026future,genicot2020aspirations}. Our results indicate that, in this case, policy interventions should target the most pessimistic students to maximize their effectiveness.


\subsection{Policy Intervention} \label{application:policy}
We study the policy implications of the empirical results. We consider an intervention in which a social planner seeks to improve or reduce the aggregate outcome, $\sum_{i = 1}^{n_m} y_{m,i}$, in some network $m$, depending on whether the outcome is socially desirable or undesirable, respectively. This is achieved by treating students, where treatment consists of increasing productivity $\alpha_{m,i}$ by one unit when the outcome is socially desirable, or decreasing $\alpha_{m,i}$ by one unit when the outcome is socially undesirable.\footnote{Our results are not driven by the specific choice of increasing or decreasing productivity by one unit. Similar results are obtained under alternative treatment intensities.} As in \cite{galeotti2020}, we assume that resources are limited, so that only $\kappa$ students can be treated. We vary $\kappa$ from 1 to $n_m$ and show that ignoring asymmetry may lead to inefficient allocations. This misallocation can result in substantial losses in aggregate outcomes and may even render the policy counterproductive.

\subsubsection{Treatment Allocation in the Asymmetric Model}
As shown in Section \ref{microfoundation::socialmultiplier}, the optimal order in which students should be treated under the symmetric model is determined by the ranking induced by the vector $\boldsymbol{\Delta}$, which admits a closed form. Consequently, for any budget $\kappa$, the optimal set of students can be identified straightforwardly under this model. By contrast, because the asymmetric model does not admit a closed-form solution for the outcome, determining the optimal allocation is considerably more challenging. It requires an exhaustive search over all possible combinations of $\kappa$ students from a school of size $n_m$. Even for moderate values of $\kappa$ and $n_m$, the number of possible allocations quickly reaches several billions, rendering such a search computationally infeasible.\footnote{For example, treating half of the students in a school of 100 yields more than $10^{29}$ possible combinations.}

To circumvent this challenge, we propose two sequential algorithms---referred to as the ``forward'' and ``backward'' algorithms---that approximate the \textit{oracle} optimal treatment allocation. For a budget $\kappa$, we denote by $\mathcal T_{m,\kappa}$ the set of treated students selected by a given approximation algorithm.

The ``forward'' algorithm first addresses the case $\kappa = 1$, which is straightforward because it involves only $n_m$ possible allocations. It then proceeds recursively. Given $\kappa < n_m$ and $\mathcal T_{m,\kappa}$, the allocation for $\kappa + 1$ is obtained under the assumption that $\mathcal T_{m,\kappa} \subset \mathcal T_{m,\kappa+1}$. This assumption reduces the search space to $n_m - \kappa$ candidate allocations, since one only needs to identify the single student to add to $\mathcal T_{m,\kappa}$ to obtain $\mathcal T_{m,\kappa+1}$.

The ``backward'' algorithm proceeds similarly, but in reverse. It first addresses the case of $\kappa = n_m$, which involves only one possible allocation. Then, for any $\kappa > 1$, the allocation for $\kappa - 1$ is obtained under the assumption that $\mathcal T_{m,\kappa - 1} \subset \mathcal T_{m,\kappa}$, reducing the search space to $\kappa$ candidate allocations. This consists of identifying the student to remove from $\mathcal T_{m,\kappa}$ to obtain $\mathcal T_{m,\kappa - 1}$.

Implementing these algorithms for all values of $\kappa$, from 1 to $n_m$, reduces the computational complexity from $2^{n_m}$ to $n_m^2$, making the problem tractable even for large networks. However, the two approximations generally do not yield the same solution and may differ from the oracle allocation because the latter does not necessarily satisfy $\mathcal T_{m,\kappa} \subset \mathcal T_{m,\kappa+1}$. Unfortunately, it is challenging to quantify the regret of these approximations relative to the oracle allocation. Nevertheless, the degree of agreement between the two algorithms provides a useful diagnostic. When the allocations they produce are similar, the planner may have greater confidence in the robustness of the resulting recommendation. In practice, for each value of $\kappa$, the planner can select the algorithm that delivers a better-performing allocation.


\subsubsection{Intervention Results}
We consider the cases where the students to target are selected according to either the symmetric model or the forward or backward approximation of the asymmetric model. Once the selected students are treated, we evaluate the resulting aggregate outcome using the asymmetric model, which we regard as the true data-generating process because it is more general. For each outcome, we present results for two randomly selected schools (see Figure~\ref{fig:intervention}). Panel A presents the spillover generated by each targeting rule. For a socially desirable outcome, the spillover is defined as the difference between the change in the aggregate outcome and $\kappa$, expressed as a percentage of $\kappa$.\footnote{For example, for a socially desirable outcome, the spillover is 50\% when the aggregate outcome increases by 15 while 10 students are treated.} For socially undesirable outcomes, it is defined analogously with the opposite sign. Panel B reports the loss in spillover associated with the symmetric allocation relative to the asymmetric allocations.



\afterpage{\begin{landscape}
\begin{figure}
    \centering
    \includegraphics[scale = 0.58]{figures/spillovers.noFakeIso.pdf}
    \caption{Spillover losses from using the symmetric model to allocate a treatment}
    \justifying
    \footnotesize{\noindent Notes: For each outcome, two schools are randomly selected from among those with $n_m \in [\bar n - 100, \, \bar n + 100]$, where $\bar n$ is the average school size. For each school, Panel A plots the spillover effect of the intervention as a function of the intervention budget $\kappa$. Panel B displays the loss in spillover associated with the symmetric allocation relative to the forward and backward asymmetric allocations.}
    \label{fig:intervention}
\end{figure}
\end{landscape}}

In general, both targeting rules of the asymmetric model yield similar results, with the forward approximation outperforming the backward approximation for certain values of $\kappa$, and vice versa. This similarity provides reassurance regarding the robustness of the approximations.

Although asymmetry matters for smoking, the difference in spillover effects under the symmetric targeting rule is not large. This is because both higher- and lower-performing peers exert strong influence, so that ignoring asymmetry generates only limited inefficiency. Using the symmetric model for treatment assignment leads to a 5\% to 15\% loss in spillovers. Yet, even such a modest loss may translate into substantial monetary costs when the policy is implemented in a large school.

For fighting and optimism, however, asymmetry is substantially more important. The results indicate that using the symmetric model may generate no spillover effects, or may even lead to negative spillovers, which is worse than what would arise in a benchmark where students do not interact. In the case of fighting, the symmetric model assigns treatment to students who are less aggressive friends, while these students exert little or no influence. This generates no spillovers or slightly negative spillovers. In this case, the loss in the aggregate outcome when the planner uses the symmetric model exceeds 70\% and may reach 100\%. The results are similar for optimism. The symmetric model assigns treatment to more optimistic friends, even though these friends exert little or no influence. As a result, the loss in the aggregate outcome generally ranges from 40\% to 75\%.

Finally, for drinking, using the symmetric model generates virtually no loss. This is expected because peer effects are approximately symmetric for this outcome. The loss is generally below 1\% and is likely not statistically significant.



\section{Conclusion} \label{sec:conclu}

This paper develops a model of asymmetric peer effects in which individuals respond differently to peers who outperform them and to those who underperform them. We propose an econometric approach to identify and estimate both effects and show that this asymmetry is empirically pervasive and consequential for policy. Standard peer effects models overlook this distinction and instead recover an aggregate effect that is not simply the average of the two asymmetric effects. In fact, this aggregate effect can lie outside the range defined by the underlying asymmetric effects. Consequently, policy interventions based on the symmetric model may be inefficient or even harmful. Our empirical application demonstrates the relevance of these asymmetric effects across several student outcomes. In targeted interventions, we find that ignoring this asymmetry may generate substantial losses in spillovers.

This paper contributes to the recent literature on heterogeneous peer effects, where heterogeneity is often introduced through the distribution of peer outcomes \citep{boucher2023}. Although this form of heterogeneity is relevant, our results show that it does not fully capture how individual decisions are formed. Decisions also depend on how individuals evaluate themselves relative to their peers. As a result, even when facing the same distribution of peer outcomes, individuals may respond differently depending on their own performance. Since overlooking this distinction may lead to counterproductive interventions, standard peer effects models should be interpreted with caution. More broadly, our results emphasize that policy-oriented research on peer effects should rely on flexible models to better understand the mechanisms through which social interactions shape behavior.
\clearpage
\newpage