EconBase
← Back to paper

Peer effect analysis with latent processes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

37,597 characters · 7 sections · 20 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Peer effect analysis with latent processes

\thispagestyle{fancy}

abstractI study peer effects that arise from irreversible decisions in the absence of a standard social equilibrium. I model a latent sequence of decisions in continuous time and obtain a closed-form expression for the likelihood, which allows to estimate proposed causal estimands. The method avoids linear-in-means regression by modeling the (possibly unobserved) realized direction of causality, whose probability is identified. I provide identification and estimation results under two settings, several networks and one large network, while allowing for various forms of peer effect heterogeneity. Under (strong) data requirements, it is possible to separate endogenous, contextual, and correlated effects while allowing for full heterogeneity and maximum likelihood methods where parameters lend themselves to standard inference. {\bf Keywords}: Peer effects, Continuous time, Heterogeneity, Causal Inference, Networks.

Introduction

Analyzing peer effects is notoriously difficult. In a seminal paper, manski1993identification discusses the reflection problem that arises when one runs a regression on the conditional expectation, which induces restrictions on the regression coefficients that lead to tautological models or identification issues.

In applications that feature small group or friendship, the channel of influence plausibly operates through actual outcomes, which leads to Spatial Autoregressive (SAR) or linear-in-means (LIM) models in the outcomes. They alleviate reflection problems insofar as they allow for identification and estimation of peer effects bramoulle2009identification, de2010identification, blume2015linear, martellosio2022non, e.g., through the use of non-overlapping peer groups that instruments away the simultaneity bias.

However, inference about the effect of peers using SAR models can be challenging. manski1993identification points out difficulties in identifying and estimating peer effects that transcend reflection regression issues. Identification can be tenuous in SAR models as it is based on functional form restrictions that are particularly sensitive to, e.g., measurement error or mechanical relationships between individual outcomes and group means gibbons2012mostly, angrist2014perils. Although identification is typically the rule when the network is known blume2015linear, identification is dependent on the mechanics of the model through its reduced form. Moreover, reliable inference about peer effects comes with caveats even in the absence of identification problems hayes2024peer, wang2025weak.

Frameworks for analyzing peer effects beyond linear models are scarce but often desirable. As summarized in sacerdote2014experimental, “researchers have shown that linear-in-means model of peer effects is often not a good description of the world, although we do not yet have an agreed-upon model to replace it.” boucher2024toward provide a recent breakthrough that extends the response class to contain mean, maximum, and minimum outcome in situations of equilibrium.

Focusing on the case of irreversible decisions, I develop a new framework by modeling the latent sequence of decisions in continuous time. In this case, the standard notion of social equilibrium\footnote{The solution for the conditional expectation or the vector of outcomes implied by the linear-in-means specification.} is inadequate as people opt in over time according to their characteristics and shocks, then cannot subsequently adjust their outcomes. This setup offers a natural framework to discuss causality, as it breaks the simultaneity by separating first-movers that potentially generate a causal reaction from the subjects of influence. I formalize the counterfactual framework using potential outcomes, and define causal peer effect parameters that lend themselves to straightforward inference via maximum likelihood.

In this context, an important object is the order in which individuals select into the absorbing state. When unknown, it is potentially an object of interest. Although the realized order cannot be identified, the probability of any order can be identified as a function of individuals' network and characteristics; I provide a framework to estimate those probabilities. Conversely, when the order is (partially) known, it carries identifying power that is useful for identifying and estimating peer effects without restricting heterogeneity, functional form, or the presence of contextual and correlated effects.

As we distinguish between sources and recipients of peer influence, exploring peer-effect heterogeneity can be more fruitful, as it provides parameters that are more natural to interpret. Heterogeneity in peer effects is an important, but under-explored topic (recent work on the subject includes mogstad2024peer). An important source of heterogeneity may be that the influence of $i$ on $j$ need not be the same as that of $j$ on $i$, whence the relevance of orderings. For instance, people may react differently upon observing a popular or highly educated peer opting in than this peer would in the converse situation. It is also likely that some characteristics make people move first with high probability.

I analyze identification under two main regimes: one large network and many networks. Large networks with observed ordering are found to be fully identified, i.e., peer effects of heterogeneous forms are identified separately from influence of covariates, contextual effects, and correlated effects under (strong) data requirements. Identification stems from information about the order of moves and is not tied to functional form restrictions; no aspect of heterogeneity needs to be restricted.

In the many, small network case, identification is more tenuous because correlated effects may create incidental parameters, which is especially problematic if those effects can interact with any other covariate or peer effect strength. In this case, they must be handled by a combination of functional form restrictions, distributional assumptions, and extra information. In the absence of correlated effects, identifying heterogeneous, nonlinear peer effects and the influence of covariates is feasible.

The paper contributes to the literature on peer effects, in particular identification and estimation issues manski1993identification, bramoulle2009identification,angrist2014perils,blume2015linear,hayes2024peer. In addition, the results on orderings make connections to the literature on targeting and diffusion banerjee2013diffusion, he2018measuring while the method may also be useful to analyze staggered treatment adoption shaikh2021randomization, athey2022design by relaxing the common assumption that treatment adoptions arise independently.

Causality, stochastic process, and likelihood

Setup and causal peer effect parameters

I focus on irreversible\footnote{This can be because the action cannot be undone (e.g., vaccination), is too costly to reverse, or because the focus is on first-time events.} decisions: there is an initial default state, labeled $0$, and the decision to opt-in leads to an absorbing state, labeled $1$. For instance, the outcome $y$ might represent vaccination status, technology adoption, retirement, migration decision, etc. The goal is to model and estimate peer effects, i.e. how decisions of peers alter an individual's probability of opting in. The peers are described by a network; their importance is given by a weighting matrix $W$. Although it is naturally interpreted as a row-normalized matrix giving positive weights to each individual's peers, the results apply to general matrices provided that $W$ is bounded in row and column sums.

The adoption time of an individual $i$, denoted by $T_i$, is a random variable that depends on individuals' characteristics and their expectations. Observing a peer opting in modifies the likelihood that an individual does so as well. This can be due to conformity, information transmission, or other types of social influence. This implies that there is a change in the distribution of the adoption time of individual $i$, $T_i$, upon observing the adoption of a peer.

Adapting the potential outcome notation neyman1923application, rubin1974estimating to the current setup\footnote{Potential outcomes have been used to formalize causality in peer effects in different contexts. The literature on interference (“exogeneous peer effects” or sometimes spillovers), in which one's outcome depends on neighbors' treatment, addresses (versions of) the issue and identifies direct and indirect effects under various relaxations of SUTVA toulis2013estimation, sofrygin2016semi, aronow2017estimating, arpino2017implementing, liu2019doubly, forastiere2020identification, jackson2020adjusting, sanchez2021spillovers, huber2021framework. Potential outcome that depends on peer's outcome (e.g., in egami2024identification) have been the object of less discussion. To my knowledge, the discussion of peer effects through potential outcomes in the time dimension is new.}, the adoption time is represented by $T_i(\tau)$ for $\tau \in \mathds{R}^{n-1}$: the adoption time of individual $i$ depends on the adoption times of other individuals. This dependence carries over to the outcomes $\mathds{1}_{T_i \leq t}$, which are observed for $t=S>0$: $y_i \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \mathds{1}_{T_i \leq S}$.

In general, the resulting change in the distribution of outcomes following a change in $\tau$ can be interpreted as a peer effect. The causal impact can be captured by parameters that summarize the effect of a change in $\tau$ on the distribution.

A parameter of interest is the analog of the average treatment effect: the mean change in the outcome following a change in $\tau$:

equation[equation omitted — 219 chars of source]

For instance, one could be interested in $\tilde{\tau}=\vec{0} \in \mathds{R}^{n-1}$ and $\tau=\infty e_i$\footnote{$e_i$ is a vector whose only nonzero entry is a $1$ in place $i$. I adopt the convention $0 \cdot \infty = 0$.} or $\tau= \vec{\infty}$ so that the parameter describes the change in adoption probability induced by the initial adoption of one or all peers compared to them never adopting.

We could also be interested in counterfactual effects such as the expected adoption time, $\mathds{E}[T_i]$, and various conditional versions, or the expected time before a fraction of the population opts in.

remark[Beyond binary outcomes] The process can be extended to account for more complex decisions. For instance, there could be a set of absorbing states to select into or there could be a gradation in behavior (say, $y_i \in \mathds{N}$ is non-decreasing over time). The analysis can be extended to such frameworks at the expense of additional notation and suitable regularity conditions.

Stochastic process and unconfoundedness

Individual adoption is modeled with the following continuous-time stochastic process:

definition[Stochastic process] Let $T_i^1$ be a set of random variables over $\mathds{R}^+$. When a first individual, say $j$, opts in (i.e., $j = \arg \min_i (T_i^1)$), she withdraws and the remaining random variables are updated to $T_i^{2}$, whose distribution may depend on the time elapsed and the event that $j$ opted in. The process then goes on with the updated distributions, and so on. The time to adoption is therefore $T_i = \sum_{k=1} T_i^k$, where the sum ranges from $1$ to the round where $i$ comes first. The outcomes are observed at time $S$: $y_i = \mathds{1}_{T_i \leq S}$.

The process decomposes the arrival time $T_i$ of each individual $i$ into a collection $T_i^* = (T_i^1, \ldots, T_i^n)$ of latent partial times. This process is quite general in terms of the dynamics it allows. It basically only assumes the arrow of time, ruling out feedback from the future.

Peer effects occur when the updated distributions do not coincide with the previous distribution. Latent partial times thus incorporate distributional changes due to peer effects, but are usually correlated because characteristics influence both the choice of peers and the distribution of times, inducing, e.g., homophily bias shalizi2011homophily. Controlling for these characteristics can restore independence:

assumption[Independence of latent partial times] The latent partial times are conditionally independent: $\forall i, T_{-i}^{k+1} \perp \!\!\! \perp T_{i}^{k+1} \vert X, W, \mathcal{F}_k$

where $\mathcal{F}_k$ is the relevant filtration (collecting the previous $\operatorname*{arg\,min}_i T_i^k$ and $\min_i T_i^k$, i.e., the identity of previous movers and their (partial) times. It also satisfies $\mathcal{F}_0 = \emptyset$).

The independence of latent partial times is related to a notion of unconfoundedness. Note that the potential outcomes can be expressed in terms of the latent partial times $T_i^{k}, k=1, \ldots, $ via

equation[equation omitted — 190 chars of source]

where $\tau_{(0)} \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} 0$, the sum ranges from $1$ to $\operatorname*{arg\,min}_k \{T_i^k < \tau_{(k)} - \tau_{(k-1)}\}$, and the latent partial times are based on $\tau$.

Then, unconfoundedness can be stated as follows.

assumption[Unconfoundedness of timings] Potential outcomes are independent of times of arrival: $\forall i, k, \tau, \ T_{-i}^{k+1} \perp \!\!\! \perp T_i(\tau) \vert X, W, \mathcal{F}_k$

where $T_{-i}$ is the set of times of all individuals but $i$.

Then, the following result follows immediately from the formulation of potential outcomes:

propositionUnconfoundedness of timings is equivalent to independence of latent partial times.

Unconfoundedness, or equivalently independence of latent partial times, ensures that some (changes in) parameters have a causal interpretation. It states that network links are not predictive of unobserved factor which influences the time when people opt in. This requires controlling for variables that affect both network formation and the probability of participation, which is a strong assumption, but similar to frequently invoked exogeneity assumptions, e.g., bramoulle2009identification.

Although beyond the scope of this paper, modeling network formation to identify or control for latent variables goldsmith2013social, graham2017econometric, auerbach2022identification, starck2025improving can help in the event of unobserved confounders. When feasible, randomization of peer groups provides a natural way to ensure unconfoundedness.

Likelihood

Absent information about the times when people opt in, the latent distributions are not identified. To make the model tractable and easy to interpret, I specify the distribution to be exponential and assume that the rates depend on the peers who previously opted in:

assumption$T_i^{k+1} \vert X, W, \mathcal{F}_{k} \sim \mbox{Exp}(\lambda_i(X, W, \mathcal{F}_{k}))$

I focus on Exponential distributions for two reasons. First, exponential waiting times arise automatically under the assumption of a constant probability per unit of time, a natural point of departure. Second, exponential distributions are particularly attractive from an analytical standpoint and ensure tractability. Due to the memorylessness property, conditioning on elapsed time is irrelevant. If we assume that the peer effects depend on the identity of the previous movers (regardless of their order), the relevant filtration consists only of a set with the identity of the previous movers. In the absence of peer effects, the rates simply do not evolve: $\lambda_i^{+\mathcal{F}_k} = \lambda_i$ for all $i, k$. In what follows, I let $\lambda_i^{+\mathcal{F}_k} \equiv \lambda_i(X, W, \mathcal{F}_{k})$ with $\lambda_i \equiv \lambda_i^{+\emptyset}$.

Importantly, exponential rates can be left unrestricted as functions of covariates. Given a distribution of the outcomes, they are nonparametrically identified under few conditions. As such, the exponential specification provides a convenient framework in which peer effects are easy to interpret, while allowing for considerable heterogeneity. Under the information structure considered here (outcomes at time $S$, possibly order of adoptions), the exponential specification provides a fit to the latent time dynamics that facilitates interpretation without letting peer effect identification depend on functional form. The assumption has more bite when time dynamics is explicitly used for identification or when time analysis is of interest (e.g., diffusion analysis, generation of counterfactuals, etc.), because the exponential distribution over time then has an influence on estimates.

The exponential rates capture the heterogeneity across individuals and their changes over time reflect the influence of peers. Their changes can be directly interpreted as peer effects insofar as they capture the change in the shape of the distribution and the measure average reduction in adoption times. They also translate into practical formulae to compute parameters of interest.

exampleConsider the change in probability of adoption: \begin{equation} \mathds{E}[Y_i(\vec{0}) - Y_i(\vec{\infty})\vert X] = e^{-\lambda_i S} (1-e^{-(\lambda_i^{+} - \lambda_i) S)}) \end{equation} where $+$ corresponds to the final level of peer effects -- all peers have previously opted in. To the first-order, this simplifies to \begin{equation} e^{-\lambda_i S} (\lambda_i^{+} - \lambda_i) S \end{equation} The first term, $e^{-\lambda_i S}$, is the baseline probability of not opting in before time $S$, while $(\lambda_i^{+} - \lambda_i)$ reflects the change in exponential rates induced by the prior adoption of friends.

The following examples consider simple homogeneous cases to illustrate the framework and build intuition for identification conditions, which are formalized later.

exampleConsider two (connected) individuals ($i=1, 2)$ and homogeneous rates $\lambda, \lambda^+$. The probabilities for the four possible outcomes are derived in the appendix. They are given by \begin{align*} \begin{split} & p_{00} \mathrel{\overset{\makebox[0pt]{\normalfont\tiny def}}{=}} \mathds{P}[y_1=y_2=0\vert W_{12} =1] = e^{-2 \lambda S} \\ & p_{10} \mathrel{\overset{\makebox[0pt]{\normalfont\tiny def}}{=}} \mathds{P}[y_1=1, y_2=0\vert W_{12} =1] = \lambda e^{-\lambda^+ S} g(2 \lambda - \lambda^+)\\ & p_{01} \mathrel{\overset{\makebox[0pt]{\normalfont\tiny def}}{=}} \mathds{P}[y_1=0, y_2=1\vert W_{12} =1] = \lambda e^{-\lambda^+ S} g(2 \lambda -\lambda^+) \\ & p_{11} \mathrel{\overset{\makebox[0pt]{\normalfont\tiny def}}{=}} \mathds{P}[y_1=y_2=1\vert W_{12} =1] = 1 - p_{00} - p_{10} - p_{01} \end{split} \end{align*} where $g(\lambda) = \frac{1-e^{-\lambda S}}{\lambda}$ if $\lambda \neq 0$ and $g(0) = S$. The probabilities are identified is one observes a sequence of independent draws of such pairs. In this case, the rates are identified since they can be recovered from the probabilities: $\lambda = -\frac{\ln(p_{00})}{2S}$ while $\lambda^+$ solves $\frac{\lambda S(p_{00}-e^{-\lambda^+ S})}{\ln(p_{00}) + \lambda^+ S} =p_{10}$, where the left-hand side is strictly decreasing. Although it is not possible to identify the identity of the first mover -- the individual who may have exerted a peer effect on the other -- when $y_1=y_2=1$, it is possible to (i) estimate the peer effect strength and (ii) determine the probability of an individual moving first, which provides a probability for each direction of causality.
exampleConsider a complete network of size $n$, i.e., the adjacency matrix corresponds to $\iota \iota' - I$ where $\iota$ is a vector of ones. People have baseline rates $\lambda$ that are updated to $\lambda^{+^k}$ when $k$ people have opted in. The outcomes $y_i \in \{0, 1\}$ only inform the fraction of people who opted in, which is insufficient to identify $\{\lambda^{+^k}, k=0, \ldots, n-1\}$ as only the $k$-th mover provides a signal about $\lambda^{+^k}$. In this regime, positive peer effects cannot be distinguished from stronger baseline rates. Because of the symmetry and perfect connectivity, knowledge about the order of moves is also uninformative. The complete network case thus requires more information, but can still be identified under additional structure. For instance, if $\lambda = e^{\beta}$ for some $\beta \in \mathds{R}$ and $\lambda^{+^k} = e^{\beta+\frac{k}{n-1} \delta}$ and the times of adoption of early adopters are known, then $\lambda$ is identified from the behavior of the density around the origin.\footnote{In this case, identification relies on the exponential specification because the identifying power of the time of adoptions needs to be used.} Given the knowledge of $\lambda$, $\delta$ is identified from the share of adopters. The source of non-identification is the lack of repeated signals about rates of a given order when everyone is connected to everyone. This suggests that one can rely on sparser networks to secure identifying information that does not rely on time data or further rate restrictions. This intuition is developed in the next Section. In particular, when the order of adoption is known, then the rates can be shown to be identified when degrees are bounded.

The likelihood induced by the stochastic process is available in closed form. This result is the object of the next theorem.

theoremIf the latent partial times are independent, then the log-likelihood for the exponential specification takes the form \begin{equation} l_N \mathrel{\overset{\makebox[0pt]{\normalfont\tiny def}}{=}} N^{-1} \ln\left(\sum_{p \in \mathcal{P}} G! \vert \mathcal{P} \vert^{-1} \left(\prod_{i=1}^G \lambda_{p_i}^{+\{p_1, \cdots, p_{i-1}\}}\right) \sum_{g=0}^{G} \frac{e^{-{\Lambda}_{p_g} S}}{\prod_{h \neq g} {\Lambda}_{p_h} - {\Lambda}_{p_g}} \right) \end{equation} where $\mathcal{P}$ contains permutations of the $G$ people who opted in, to which the remaining people are appended in any order, and $\Lambda_g \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \sum_{i=g+1}^N \lambda_{p_i}^{+\{p_1, \cdots, p_g\}}$.

The following example shows that the process generalizes i.i.d. exponential draws. The heterogeneity in the rates induces different distributions between individuals, while the peer effects reflected in $\lambda^+ \neq \lambda$ create spatial dependence.

exampleSuppose that there is no heterogeneity nor peer effects: $\lambda_i^{+\mathcal{F}} \equiv \lambda$ for all $i$ and any $\mathcal{F}$. Then $\Lambda_g = \lambda (N-g)$ and the likelihood any ordering becomes \begin{align*} \begin{split} G! \prod_{i=1}^G \lambda_i^{+\{(1), \cdots, ({i-1})\}} \sum_{g=0}^{G} \frac{e^{-\Lambda_g S}}{\prod_{h \neq g} \Lambda_h - \Lambda_g} & = G! \lambda^G \sum_{g=0}^{G} \frac{e^{- \lambda (N-g)S}}{\prod_{h \neq g} \lambda (N-h)-\lambda (N-g)} \\ & = G! e^{-\lambda (N-G) S} \sum_{g=0}^{G} \frac{e^{- \lambda (G-g)S}}{\prod_{h\neq g} g-h} \\ & = e^{-\lambda (N-G) S} \sum_{g=0}^{G} G! \frac{e^{- \lambda (G-g)S}}{g! (G-g)!} (-1)^{G-g} \\ & = e^{-\lambda (N-G) S} (1-e^{-\lambda S})^G, \end{split} \end{align*} which is the likelihood for $G$ (out of $N$) i.i.d. exponentially distributed variables falling below a cutoff $S$. Therefore, probabilities reduce to a standard exponential race with independent draws when the rates do not vary. Changes in exponential rates are a measure of dependence and social interaction; they reflect peer effects.
remark[2 people example] The probabilities in Example (ref) can be directly obtained from the theorem. For instance, noting $\Lambda_{p_1} = \lambda_1+\lambda_2$ and $\Lambda_{p_2} = \lambda_{p_2}^+$, one can verify \begin{align} \begin{split} \mathds{P}[y_1=y_2=1\vert W_{12} =1] = 1 & + \frac{\lambda_1 \lambda_2^+ e^{-(\lambda_1+\lambda_2)S}}{(\lambda_1+\lambda_2-\lambda_2^+)(\lambda_1+\lambda_2)} - \frac{\lambda_1 e^{-\lambda_2^+ S}}{\lambda_1+\lambda_2-\lambda_2^+} \\ & + \frac{\lambda_2 \lambda_1^+ e^{-(\lambda_1+\lambda_2)S}}{(\lambda_1+\lambda_2-\lambda_1^+)(\lambda_1+\lambda_2)} - \frac{\lambda_2 e^{-\lambda_1^+ S}}{\lambda_1+\lambda_2-\lambda_1^+} \\ & = 1 - p_{00} - p_{10} - p_{01} \end{split} \end{align}
remark[Computation] Although summing over all permutations can lead to an impractical computational burden, the computational cost can be lowered to practical levels. First, the likelihood factorizes based on the components of $W$, whose size can be much smaller (classrooms, villages, connected components of friends, etc.). Second, permutations are based on the number of adopters, rather than the size of each component of the network. Hence, the complexity of permutations is substantially reduced, especially when partial information about the ordering is available. Finally, approximations can further reduce the number of permutations. In particular, the average over permutations can be estimated by a random sample of permutations by the law of large numbers. The likelihood, score, and Hessian can all be approximated using this technique.

Identification

The sample comprises outcomes ($y_i, i=1, \ldots, n$), covariates $(x_i, i=1, \ldots, n)$, and a weighting matrix $W$ ($W_{ij} > 0$ if $i$ and $j$ are peers) that determines the peer group of each individual.

Identification relies on our ability to separate information about cross-sectional variation in rates and social influence effects. In the complete network of Example (ref), this is not possible without further information because once a single person opts in, everybody else is subject to social influence and there is no additional information about baseline rates. Most real-life networks, however, are much sparser or exhibit a block structure due to a sampling scheme that targets villages, classrooms, etc. I show how this provides the necessary information for identification under the two regimes: the network consists of components or blocks which do not interact, or are sufficiently sparse. In the latter case, I make use of the assumption that degrees are bounded, which has been invoked in the literature to consider sparser networks (e.g., de2018identifying, though it can be relaxed to a slow degree growth rate.

lemmaLet $x_i$ be composed of discrete and continuous variables, and let the rates be continuous in the continuous variables. Assume that latent partial times are independent and that the exponential specification holds with $\lambda_i^{+\mathcal{F}} \in A \subseteq \ ]0, \infty[$, for all $i$ and set $\mathcal{F}$. Then, \\ (i) If the network has a block structure with blocks $b=1, \cdots, B$, of size $N_b \leq \overline{N}$ for some $\overline{N} \in \mathds{N}$, then $\lambda_i^{+\mathcal{F}}$ is identified whenever $\mathds{P}[W_{b;ik}>0, \forall k \in \mathcal{F}\vert X] > 0$, $\mathds{P}[Y_b=y_b\vert W_{b;ik}>0, \forall k \in \mathcal{F} \vert X] > 0$, and the number of blocks grows large. \\ (ii) If the order of adoption is known, the degrees are bounded, i.e., $\sup_{i} d_i \leq \overline{d} < \infty$ where $d_i = \sum_{j=1}^N \mathds{1}(W(i, j) > 0)$, and $A$ is compact, then $\lambda_i^{+\mathcal{F}}$ is identified for all $i$ and any set of people $\mathcal{F}$ whose characteristics make a connection with $i$ possible with positive probability.
remark[Identification from panel information] Identification using the order of adoption can be extended to setups where the order is only partially known. For instance, monitoring the activations at regular times provides a partial order that secures identification as the frequency grows.
remark[Identification of order probabilities] Identification of the family of rates implies that the probability of any sequence of adoption is also identified. Hence, although it is not possible to recover the identity of the first movers or the realized order of adoptions, it is possible to assign a probability to such events.
remark[Contextual and correlated effects] Contextual peer effects can be straightforwardly allowed for in the large network case. Within a large network, the absence of correlated effects hinges only on the (strong) assumption of no unobserved confounders. Thus, under the maintained unconfoundedness condition, information about orderings is sufficient to identify peer effects in full heterogeneity while separating them from contextual and correlated effects. In general, note that no restriction has been placed on the way the specific network heterogeneity may interact with covariates or peer effects. It is thus unsurprising that identification generically fails: separating endogenous effects from contextual or correlated effects requires further information. In the large network case, this can take the form of the order of adoption. With several networks, which induce incidental parameters, this would require functional form restrictions (in the spirit of the SAR reduced form, for instance), distributional assumptions, or partial information about the order of adoptions combined with other restrictions.

Asymptotic Theory

Although rates are nonparametrically identified, it is crucial to introduce more parsimonious specifications. This avoids severe curse-of-dimensionality issues due to the large number of rates and the dependence on possibly many continuous covariates and simplifies interpretation.

A natural model specifies the baseline rates as $\lambda_i = g(x_i, \beta)$ for a positive link function determined by a finite-dimensional vector of parameters $\beta$, and lets the updated rates be obtained by a scaling factor $\delta$. For instance, a simple specification is

equation[equation omitted — 144 chars of source]

Such a specification reduces the dimensionality of the problem to $(\mbox{dim}(x_i)+1)$, ensures positive rates, and lets peer effects be described by a parameter $\delta$ that reflects how rates are scaled when the members of the peer group opt in. $\delta$ controls the strength of spatial dependence.

To the first-order in $\delta$, we have

equation[equation omitted — 113 chars of source]

so that $\delta$ acts a scaling factor in the change in the probability of adoption induced by the peer group adopting at the onset.

It is easy to relax the parametric specification in various directions. The relevant direction is an applied choice that is likely to differ between applications. For instance, suppose that a researcher is analyzing peer effect in vaccine uptake. They might conjecture that people respond differently depending on the level of education of their peers and that the marginal peer effect is decreasing because the first few movers provide most of the information transmission or reassurance about safety. A natural way to explore heterogeneity in peer effects is then to add interaction terms such as $\sum_{j \in \mathcal{F}} W(i, j) x_j$, nonlinear terms in $\sum_{j \in \mathcal{F}} W(i, j)$ such as powers, or $p$-norms in the spirit of boucher2024toward.

For a given parameterization of the rates in terms of a parameter $\theta$ (for instance, $\theta = (\beta', \delta)'$ in ((ref))), estimation proceeds by maximum likelihood. Asymptotics can be obtained in two frameworks: multiple networks and one large network.

In the first case, I consider a sequence of networks with blocks or components of size $N_b$, $b=1, \ldots, B$. With independent blocks, the likelihood factorizes as $l \mathrel{\overset{\makebox[0pt]{\mbox{\normalfont\tiny\sffamily def}}}{=}} \ln(\mathds{P}[Y=y]) = \sum_{b=1}^B l_b$, where $l_b$ is the log-likelihood of block $b$. The logarithmic likelihood of each block is given by the formula established in the theorem of the previous section. The score and Hessian are easily computed in closed form. The details are provided in the Appendix. In the second case, the network consists of a single component whose size grows and a key ingredient comes from knowledge of $\mathcal{F}$, as in the identification condition. In this case, I proceed conditionally on the covariates and the network structure.

In both cases, the estimator $\hat{\theta}$ inherits the usual properties of maximum likelihood estimators: it is consistent and asymptotically normal under regularity conditions. This implies consistency and asymptotic normality of the family of rates. Formally,

theorem[Consistency] Let the data-generating process be described by the stochastic process of Section (ref) and that the family of rates is identified and belongs to the interior of a compact set. Then, (i) if blocks are drawn independently with $W_b \sim \mathcal{D}_W$ such that $N_b<\overline{N}$, the maximum likelihood estimator is consistent with $\hat{\lambda}_i^{+\mathcal{F}} \overset{p}{\longrightarrow} \lambda_i^{+\mathcal{F}}$, for any $\mathcal{F} \subseteq \{1, \ldots, i-1, i+1, \ldots, N_b\}$, as $B \rightarrow \infty$. Alternatively, (ii) if the size of the network grows large and condition (ii) of Lemma (ref) is satisfied, the maximum likelihood estimator is consistent with $\hat{\lambda}_i^{+\mathcal{F}} \overset{p}{\longrightarrow} \lambda_i^{+\mathcal{F}}$, for any $\mathcal{F}$, as $N \rightarrow \infty$.

Moreover,

theorem[Asymptotic Normality] Assume that part (i) of the assumptions of the Consistency theorem hold and that the derivatives $\partial_\theta \lambda_i^{\mathcal{F}}$ are bounded. Then, the maximum likelihood estimator is asymptotically normal as $B \rightarrow \infty$ with \begin{equation} \sqrt{B} (\hat{\theta}-\theta) \overset{d}{\longrightarrow} \mathcal{N}\left(0, - H^{-1}\right) \end{equation} Alternatively, assume that part (ii) of the assumptions of the Consistency theorem hold and that the derivatives $\partial_\theta \lambda_i^{\mathcal{F}}$ are bounded. Then, the maximum likelihood estimator is asymptotically normal as $N \rightarrow \infty$ with \begin{equation} \sqrt{N} (\hat{\theta}-\theta) \vert X, W \overset{d}{\longrightarrow} \mathcal{N}\left(0, - H^{-1}\right). \end{equation} Here, the Hessian is given by \begin{equation} H \mathrel{\overset{\makebox[0pt]{\normalfont\tiny def}}{=}} - \int_0^c \frac{\partial_{\theta \theta'} r(z)}{r(z)+z^*} dz + \int_0^c \frac{\partial_{\theta} r(z) \partial_{\theta'} r(z)}{(r(z)+z^*)^2} dz - \frac{\int_0^c \frac{\partial_{\theta} r(z)}{(r(z)+z^*)^2} dz \ \int_0^c \frac{\partial_{\theta'} r(z)}{(r(z)+z^*)^2} dz}{\int_0^c \frac{1}{(r(z)+z^*)^2} dz}, \end{equation} where $r$ is a limiting function for a transformation of the rates $\Lambda_g$, $c = \lim_{N\longrightarrow \infty} \mathds{E}[G/N]$, $z^*$ satisfies a stationarity condition, as described in the Appendix. Finally, by the Delta method, the rates obey \begin{equation} \sqrt{N} (\hat{\lambda}_i^{+\mathcal{F}} - \lambda_i^{+\mathcal{F}}) \vert X, W \overset{d}{\longrightarrow} \mathcal{N}\left(0, - \partial_{\theta'} \lambda_i^{+\mathcal{F}} H^{-1} \partial_{\theta} \lambda_i^{+\mathcal{F}}\right) \end{equation}

These theorems allow for standard inference about the underlying parameters and the rates.