Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
37,597 characters · 7 sections · 20 citation commands
Peer effect analysis with latent processes
\thispagestyle{fancy}
Analyzing peer effects is notoriously difficult. In a seminal paper, manski1993identification discusses the reflection problem that arises when one runs a regression on the conditional expectation, which induces restrictions on the regression coefficients that lead to tautological models or identification issues.
In applications that feature small group or friendship, the channel of influence plausibly operates through actual outcomes, which leads to Spatial Autoregressive (SAR) or linear-in-means (LIM) models in the outcomes. They alleviate reflection problems insofar as they allow for identification and estimation of peer effects bramoulle2009identification, de2010identification, blume2015linear, martellosio2022non, e.g., through the use of non-overlapping peer groups that instruments away the simultaneity bias.
However, inference about the effect of peers using SAR models can be challenging. manski1993identification points out difficulties in identifying and estimating peer effects that transcend reflection regression issues. Identification can be tenuous in SAR models as it is based on functional form restrictions that are particularly sensitive to, e.g., measurement error or mechanical relationships between individual outcomes and group means gibbons2012mostly, angrist2014perils. Although identification is typically the rule when the network is known blume2015linear, identification is dependent on the mechanics of the model through its reduced form. Moreover, reliable inference about peer effects comes with caveats even in the absence of identification problems hayes2024peer, wang2025weak.
Frameworks for analyzing peer effects beyond linear models are scarce but often desirable. As summarized in sacerdote2014experimental, “researchers have shown that linear-in-means model of peer effects is often not a good description of the world, although we do not yet have an agreed-upon model to replace it.” boucher2024toward provide a recent breakthrough that extends the response class to contain mean, maximum, and minimum outcome in situations of equilibrium.
Focusing on the case of irreversible decisions, I develop a new framework by modeling the latent sequence of decisions in continuous time. In this case, the standard notion of social equilibrium\footnote{The solution for the conditional expectation or the vector of outcomes implied by the linear-in-means specification.} is inadequate as people opt in over time according to their characteristics and shocks, then cannot subsequently adjust their outcomes. This setup offers a natural framework to discuss causality, as it breaks the simultaneity by separating first-movers that potentially generate a causal reaction from the subjects of influence. I formalize the counterfactual framework using potential outcomes, and define causal peer effect parameters that lend themselves to straightforward inference via maximum likelihood.
In this context, an important object is the order in which individuals select into the absorbing state. When unknown, it is potentially an object of interest. Although the realized order cannot be identified, the probability of any order can be identified as a function of individuals' network and characteristics; I provide a framework to estimate those probabilities. Conversely, when the order is (partially) known, it carries identifying power that is useful for identifying and estimating peer effects without restricting heterogeneity, functional form, or the presence of contextual and correlated effects.
As we distinguish between sources and recipients of peer influence, exploring peer-effect heterogeneity can be more fruitful, as it provides parameters that are more natural to interpret. Heterogeneity in peer effects is an important, but under-explored topic (recent work on the subject includes mogstad2024peer). An important source of heterogeneity may be that the influence of $i$ on $j$ need not be the same as that of $j$ on $i$, whence the relevance of orderings. For instance, people may react differently upon observing a popular or highly educated peer opting in than this peer would in the converse situation. It is also likely that some characteristics make people move first with high probability.
I analyze identification under two main regimes: one large network and many networks. Large networks with observed ordering are found to be fully identified, i.e., peer effects of heterogeneous forms are identified separately from influence of covariates, contextual effects, and correlated effects under (strong) data requirements. Identification stems from information about the order of moves and is not tied to functional form restrictions; no aspect of heterogeneity needs to be restricted.
In the many, small network case, identification is more tenuous because correlated effects may create incidental parameters, which is especially problematic if those effects can interact with any other covariate or peer effect strength. In this case, they must be handled by a combination of functional form restrictions, distributional assumptions, and extra information. In the absence of correlated effects, identifying heterogeneous, nonlinear peer effects and the influence of covariates is feasible.
The paper contributes to the literature on peer effects, in particular identification and estimation issues manski1993identification, bramoulle2009identification,angrist2014perils,blume2015linear,hayes2024peer. In addition, the results on orderings make connections to the literature on targeting and diffusion banerjee2013diffusion, he2018measuring while the method may also be useful to analyze staggered treatment adoption shaikh2021randomization, athey2022design by relaxing the common assumption that treatment adoptions arise independently.
I focus on irreversible\footnote{This can be because the action cannot be undone (e.g., vaccination), is too costly to reverse, or because the focus is on first-time events.} decisions: there is an initial default state, labeled $0$, and the decision to opt-in leads to an absorbing state, labeled $1$. For instance, the outcome $y$ might represent vaccination status, technology adoption, retirement, migration decision, etc. The goal is to model and estimate peer effects, i.e. how decisions of peers alter an individual's probability of opting in. The peers are described by a network; their importance is given by a weighting matrix $W$. Although it is naturally interpreted as a row-normalized matrix giving positive weights to each individual's peers, the results apply to general matrices provided that $W$ is bounded in row and column sums.
The adoption time of an individual $i$, denoted by $T_i$, is a random variable that depends on individuals' characteristics and their expectations. Observing a peer opting in modifies the likelihood that an individual does so as well. This can be due to conformity, information transmission, or other types of social influence. This implies that there is a change in the distribution of the adoption time of individual $i$, $T_i$, upon observing the adoption of a peer.
Adapting the potential outcome notation neyman1923application, rubin1974estimating to the current setup\footnote{Potential outcomes have been used to formalize causality in peer effects in different contexts. The literature on interference (“exogeneous peer effects” or sometimes spillovers), in which one's outcome depends on neighbors' treatment, addresses (versions of) the issue and identifies direct and indirect effects under various relaxations of SUTVA toulis2013estimation, sofrygin2016semi, aronow2017estimating, arpino2017implementing, liu2019doubly, forastiere2020identification, jackson2020adjusting, sanchez2021spillovers, huber2021framework. Potential outcome that depends on peer's outcome (e.g., in egami2024identification) have been the object of less discussion. To my knowledge, the discussion of peer effects through potential outcomes in the time dimension is new.}, the adoption time is represented by $T_i(\tau)$ for $\tau \in \mathds{R}^{n-1}$: the adoption time of individual $i$ depends on the adoption times of other individuals. This dependence carries over to the outcomes $\mathds{1}_{T_i \leq t}$, which are observed for $t=S>0$: $y_i \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \mathds{1}_{T_i \leq S}$.
In general, the resulting change in the distribution of outcomes following a change in $\tau$ can be interpreted as a peer effect. The causal impact can be captured by parameters that summarize the effect of a change in $\tau$ on the distribution.
A parameter of interest is the analog of the average treatment effect: the mean change in the outcome following a change in $\tau$:
For instance, one could be interested in $\tilde{\tau}=\vec{0} \in \mathds{R}^{n-1}$ and $\tau=\infty e_i$\footnote{$e_i$ is a vector whose only nonzero entry is a $1$ in place $i$. I adopt the convention $0 \cdot \infty = 0$.} or $\tau= \vec{\infty}$ so that the parameter describes the change in adoption probability induced by the initial adoption of one or all peers compared to them never adopting.
We could also be interested in counterfactual effects such as the expected adoption time, $\mathds{E}[T_i]$, and various conditional versions, or the expected time before a fraction of the population opts in.
Individual adoption is modeled with the following continuous-time stochastic process:
The process decomposes the arrival time $T_i$ of each individual $i$ into a collection $T_i^* = (T_i^1, \ldots, T_i^n)$ of latent partial times. This process is quite general in terms of the dynamics it allows. It basically only assumes the arrow of time, ruling out feedback from the future.
Peer effects occur when the updated distributions do not coincide with the previous distribution. Latent partial times thus incorporate distributional changes due to peer effects, but are usually correlated because characteristics influence both the choice of peers and the distribution of times, inducing, e.g., homophily bias shalizi2011homophily. Controlling for these characteristics can restore independence:
where $\mathcal{F}_k$ is the relevant filtration (collecting the previous $\operatorname*{arg\,min}_i T_i^k$ and $\min_i T_i^k$, i.e., the identity of previous movers and their (partial) times. It also satisfies $\mathcal{F}_0 = \emptyset$).
The independence of latent partial times is related to a notion of unconfoundedness. Note that the potential outcomes can be expressed in terms of the latent partial times $T_i^{k}, k=1, \ldots, $ via
where $\tau_{(0)} \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} 0$, the sum ranges from $1$ to $\operatorname*{arg\,min}_k \{T_i^k < \tau_{(k)} - \tau_{(k-1)}\}$, and the latent partial times are based on $\tau$.
Then, unconfoundedness can be stated as follows.
where $T_{-i}$ is the set of times of all individuals but $i$.
Then, the following result follows immediately from the formulation of potential outcomes:
Unconfoundedness, or equivalently independence of latent partial times, ensures that some (changes in) parameters have a causal interpretation. It states that network links are not predictive of unobserved factor which influences the time when people opt in. This requires controlling for variables that affect both network formation and the probability of participation, which is a strong assumption, but similar to frequently invoked exogeneity assumptions, e.g., bramoulle2009identification.
Although beyond the scope of this paper, modeling network formation to identify or control for latent variables goldsmith2013social, graham2017econometric, auerbach2022identification, starck2025improving can help in the event of unobserved confounders. When feasible, randomization of peer groups provides a natural way to ensure unconfoundedness.
Absent information about the times when people opt in, the latent distributions are not identified. To make the model tractable and easy to interpret, I specify the distribution to be exponential and assume that the rates depend on the peers who previously opted in:
I focus on Exponential distributions for two reasons. First, exponential waiting times arise automatically under the assumption of a constant probability per unit of time, a natural point of departure. Second, exponential distributions are particularly attractive from an analytical standpoint and ensure tractability. Due to the memorylessness property, conditioning on elapsed time is irrelevant. If we assume that the peer effects depend on the identity of the previous movers (regardless of their order), the relevant filtration consists only of a set with the identity of the previous movers. In the absence of peer effects, the rates simply do not evolve: $\lambda_i^{+\mathcal{F}_k} = \lambda_i$ for all $i, k$. In what follows, I let $\lambda_i^{+\mathcal{F}_k} \equiv \lambda_i(X, W, \mathcal{F}_{k})$ with $\lambda_i \equiv \lambda_i^{+\emptyset}$.
Importantly, exponential rates can be left unrestricted as functions of covariates. Given a distribution of the outcomes, they are nonparametrically identified under few conditions. As such, the exponential specification provides a convenient framework in which peer effects are easy to interpret, while allowing for considerable heterogeneity. Under the information structure considered here (outcomes at time $S$, possibly order of adoptions), the exponential specification provides a fit to the latent time dynamics that facilitates interpretation without letting peer effect identification depend on functional form. The assumption has more bite when time dynamics is explicitly used for identification or when time analysis is of interest (e.g., diffusion analysis, generation of counterfactuals, etc.), because the exponential distribution over time then has an influence on estimates.
The exponential rates capture the heterogeneity across individuals and their changes over time reflect the influence of peers. Their changes can be directly interpreted as peer effects insofar as they capture the change in the shape of the distribution and the measure average reduction in adoption times. They also translate into practical formulae to compute parameters of interest.
The following examples consider simple homogeneous cases to illustrate the framework and build intuition for identification conditions, which are formalized later.
The likelihood induced by the stochastic process is available in closed form. This result is the object of the next theorem.
The following example shows that the process generalizes i.i.d. exponential draws. The heterogeneity in the rates induces different distributions between individuals, while the peer effects reflected in $\lambda^+ \neq \lambda$ create spatial dependence.
The sample comprises outcomes ($y_i, i=1, \ldots, n$), covariates $(x_i, i=1, \ldots, n)$, and a weighting matrix $W$ ($W_{ij} > 0$ if $i$ and $j$ are peers) that determines the peer group of each individual.
Identification relies on our ability to separate information about cross-sectional variation in rates and social influence effects. In the complete network of Example (ref), this is not possible without further information because once a single person opts in, everybody else is subject to social influence and there is no additional information about baseline rates. Most real-life networks, however, are much sparser or exhibit a block structure due to a sampling scheme that targets villages, classrooms, etc. I show how this provides the necessary information for identification under the two regimes: the network consists of components or blocks which do not interact, or are sufficiently sparse. In the latter case, I make use of the assumption that degrees are bounded, which has been invoked in the literature to consider sparser networks (e.g., de2018identifying, though it can be relaxed to a slow degree growth rate.
Although rates are nonparametrically identified, it is crucial to introduce more parsimonious specifications. This avoids severe curse-of-dimensionality issues due to the large number of rates and the dependence on possibly many continuous covariates and simplifies interpretation.
A natural model specifies the baseline rates as $\lambda_i = g(x_i, \beta)$ for a positive link function determined by a finite-dimensional vector of parameters $\beta$, and lets the updated rates be obtained by a scaling factor $\delta$. For instance, a simple specification is
Such a specification reduces the dimensionality of the problem to $(\mbox{dim}(x_i)+1)$, ensures positive rates, and lets peer effects be described by a parameter $\delta$ that reflects how rates are scaled when the members of the peer group opt in. $\delta$ controls the strength of spatial dependence.
To the first-order in $\delta$, we have
so that $\delta$ acts a scaling factor in the change in the probability of adoption induced by the peer group adopting at the onset.
It is easy to relax the parametric specification in various directions. The relevant direction is an applied choice that is likely to differ between applications. For instance, suppose that a researcher is analyzing peer effect in vaccine uptake. They might conjecture that people respond differently depending on the level of education of their peers and that the marginal peer effect is decreasing because the first few movers provide most of the information transmission or reassurance about safety. A natural way to explore heterogeneity in peer effects is then to add interaction terms such as $\sum_{j \in \mathcal{F}} W(i, j) x_j$, nonlinear terms in $\sum_{j \in \mathcal{F}} W(i, j)$ such as powers, or $p$-norms in the spirit of boucher2024toward.
For a given parameterization of the rates in terms of a parameter $\theta$ (for instance, $\theta = (\beta', \delta)'$ in ((ref))), estimation proceeds by maximum likelihood. Asymptotics can be obtained in two frameworks: multiple networks and one large network.
In the first case, I consider a sequence of networks with blocks or components of size $N_b$, $b=1, \ldots, B$. With independent blocks, the likelihood factorizes as $l \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \ln(\mathds{P}[Y=y]) = \sum_{b=1}^B l_b$, where $l_b$ is the log-likelihood of block $b$. The logarithmic likelihood of each block is given by the formula established in the theorem of the previous section. The score and Hessian are easily computed in closed form. The details are provided in the Appendix. In the second case, the network consists of a single component whose size grows and a key ingredient comes from knowledge of $\mathcal{F}$, as in the identification condition. In this case, I proceed conditionally on the covariates and the network structure.
In both cases, the estimator $\hat{\theta}$ inherits the usual properties of maximum likelihood estimators: it is consistent and asymptotically normal under regularity conditions. This implies consistency and asymptotic normality of the family of rates. Formally,
Moreover,
These theorems allow for standard inference about the underlying parameters and the rates.