EconBase
← Back to paper

Estimating peer effects in noisy, low-rank networks via network smoothing

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

66,501 characters · 17 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating peer effects in noisy, low-rank networks via network smoothing

\affil[1]{ Department of Economics, Stanford University, USA} \affil[2]{ Department of Statistics, University of Wisconsin--Madison, USA}

quotePeer effect estimation requires precise network measurement, yet most empirical networks are noisy, rendering standard estimators inconsistent. To address measurement error in networks, we propose a method to estimate peer effects in networks whose expected adjacency matrix is low-rank. Our key result shows that peer effects over a true unobserved network are asymptotically equivalent to peer effects over the expected adjacency matrix. This result reduces peer effect estimation in noisy networks to low-rank matrix estimation targeting the expected adjacency matrix. We develop our theory for weighted networks observed with additive noise, but simulations suggest approach can be applied more generally when there is a low-rank estimation method suited to a particular noise structure. We demonstrate via simulations that our approach applies to egocentric samples, aggregated relational data, and networks with missing edges, each requiring a different low-rank estimation method.

Introduction

One of the fundamental challenges in estimating social effects, such as contagion, is measuring and quantifying relationships in a reliable manner. These measurement issues pose a fundamental problem, as most current techniques for estimating peer effects in social networks rely on data that precisely encodes potential social influences. For instance, in the social sciences, a popular approach is to use the “name generator” method, where each study participant is asked to name their friends or other social ties. This method is replete with complications: participants report different friends depending on the precise language of the name generator shakya2017, bidart2011, participants often forget day-to-day social contacts smieszek2012, and reports from different participants can disagree debacco2021.

When relationship measurements are noisy, traditional peer effect estimation encounters a problem. Standard network autoregressive models assume that we observe a true, noiseless network, which precisely encodes the pathways along which peer influence operates. If we observe a network via noisy measurement, these estimators may produce inconsistent or misleading results. We propose a solution to this problem: replace noisy measurements of the adjacency matrix with smoothed, de-noised estimates of network structure. Our approach is motivated by a key theoretical insight: under low-rank network models, contagion over a true, noiseless network is asymptotically equivalent to contagion over a smoothed, latent adjacency matrix.

To formalize matters, let us consider two network autoregressive models. The first is a standard peer contagion model, where influence operates over the degree-normalized adjacency matrix:

equation[equation omitted — 136 chars of source]

where $Y_i$ is an outcome for node $i$, $\W_i \in \R^p$ are observed covariates, $\Xpop_i \in \R^d$ is a latent vector encoding the behavior of node $i$, $\A_{ij} \in \R$ is the strength of the edge between nodes $i$ and $j$, and $d_i = \sum_j \A_{ij}$ is the degree of node $i$.

The second model is a latent contagion model, where influence operates over a smoothed, expected adjacency matrix:

equation[equation omitted — 162 chars of source]

Under a random dot product graph athreya2018 model, where $\E[\Xpop]{\A_{ij}} = \Xpop_i^\top \Xpop_j$, influence is inversely proportional to the cosine distance between nodes in latent space. We develop two-stage least squares estimators for both of these models, and prove that they are consistent under misspecification. That is, estimators derived under the latent contagion model consistently recover peer contagion parameters when the true data generating process follows the peer contagion model, and vice versa. In a formal and quantifiable sense, contagion in peer and latent spaces are (asymptotically) equivalent under a random dot product graph model.

This equivalence has important practical implications. The latent contagion estimators we propose are functions of the nodal covariates $\W$ and the latent positions $\Xpop$, but not the adjacency matrix $\A$ itself. This means that even when $\A$ is observed with noise, our estimators remain consistent provided we can obtain a sufficiently good estimate $\Xhat$ of $\Xpop$. Crucially, in random dot product graphs, it is often possible to obtain high-quality estimates of $\Xpop$ even when the network is observed with measurement error or when data is missing. We leverage standard tools from spectral network analysis---specifically, the adjacency spectral embedding---to estimate latent positions lyzinski2014. Under a sub-gamma model of edge-level noise, which includes binary, weighted, and count-valued networks as special cases, we show that our estimators achieve $\sqrt{n}$ convergence rates and are asymptotically normal. Since these results hold when the observed network contains sub-gamma noise, this enables inference when network measurements are corrupted by random errors.

The flexibility of our approach extends beyond the specific case of sub-gamma edge noise. While we focus on this setting in our theoretical results, the low-rank smoothing approach that we advocate opens the door to a wide range of robust estimation strategies. Any method that can reliably estimate the low-rank structure of $\E{\A}$---whether through matrix completion for missing data, debiasing for measurement error, or other spectral methods---can be combined with our latent contagion model to estimate peer effects in challenging data settings. We demonstrate this flexibility through simulations and an empirical application studying the contagiousness of smoking in an adolescent social network.

Notation

For a positive integer $n$, we write $[n] = \{1,2,\dots,n\}$. We denote the identity matrix and the zero matrix by $\mI$ and $\symbf 0$, respectively, subscripting by the intended dimension when it is not immediately clear from context. For an $n_1 \times n_2$ matrix $\H$, we denote by $\H_{\cdot j}$ the column vector formed by the $j$-th column of $\H$, and we denote by $\H_{i \cdot}$ the row vector formed by the $i$-th row of $\H$. Abusing notation slightly, we also let $\H_i \in \R^{n_2}$ denote the column vector formed by transposing the $i$-th row of $\H$, that is, $\H_i = (\H_{i \cdot})^\top$. We often consider sequences of matrices indexed by the number of nodes $n$, but suppress the $n$ in our notation to avoid notational clutter. Given any suitably specified ordering on eigenvalues of a square matrix $\H$, we let $\lambda_i(\H)$ denote the $i$-th eigenvalue of $\H$ under that ordering. Similarly, $\sigma_i(\H)$ denotes the $i$-th singular value of $\H$. We let $\norm{\H}$ denote the spectral norm of $H$ and $\norm{\H}_F$ denote the Frobenius norm. We let $\norm{\H}_{2 \rightarrow \infty}$ denote the maximum of the Euclidean norms of the rows of $\H$, so that $\norm{\H}_{2 \rightarrow \infty} = \max_i \norm*{\H_i}$. We use standard Landau notation to indicate convergence rates, as well as typical probabilistic variants $o_p$ and $\mathcal{O}_p$. Throughout, $C > 0$ will denote a constant independent of $n$ that may change from line to line.

Related work

Our work builds on several strands in the literature. Most directly related are works that model contagion in latent spaces. sweet2020 suggest a Hoff model for social influence in a latent space, while chen2023c model correlations between stock returns via the latent structure encoded by a stochastic blockmodel in a high-dimensional time series setting. These approaches are similar to our own in that they model social influence in a latent space, but they differ from the present work in that they do not consider equivalence between latent and traditional social influence, and thus cannot use the latent space models to account for measurement error in the network.

A substantial body of prior work considers network autoregression with endogeneity, where the network depends on nodal features. hayes2025b considers identification and asymptotic signal-to-noise ratios in network autoregression models, proving minimax rates that apply to the models under consideration here. mcfowland2021 establishes consistency of ordinary least squares in network autoregression models unrolled in time, which chang2024a extends to show asymptotic normality in the longitudinal setting. paul2022a presents a network autoregressive model that accounts for estimation error in latent positions, but not in the network itself, considering a quasi-maximum likelihood estimator very similar to ours. johnsson2019 proposes a non-parametric approach to account for endogeneity in cross-sectional networks using sieve estimators, and egami2021a proposes a generalized method of moments estimator using double negative controls to account for contextual effects and homophily. All of these approaches assume precise observation of a social network that encodes possible pathways for social influence, an assumption that we find unrealistic and seek to avoid in this paper.

Most relevant to our focus on measurement error is a growing literature on linear-in-means models with network mismeasurement. griffith2022 and griffith2024a consider bias in peer effect estimates when individuals are capped at reporting a maximum number of friends. To the best of our knowledge, our estimators cannot handle this kind of degree censoring and are subject to similar bias. boucher2025 considers estimation of linear-in-means models when only partial network data is available but the network distribution is known. They propose a simulated generalized method of moments approach to estimating peer effects, given a consistent estimate of edge probabilities for a network, introducing a bias-correction to account for noise due to the simulation process. In contrast, we propose a plug-in estimator for peer effects that does not require bias adjustment. However, low-rank model estimates could naturally be used with the simulated generalized method of moments. lewbel2024b shows that two-stage least squares estimators remain consistent under small amounts of network mismeasurement. The settings considered here have substantially more error than accommodated by their results. lewbel2025 proposes an adjustment for larger amounts of measurement error, based on estimating false positive and false negative rates from either an asymmetric observation of a network or multiple networks. Mechanically, both lewbel2025 and our own approach work by plugging in an estimate of the expected adjacency matrix to account for noise. Since lewbel2025 are interested in i.i.d. false positives and false negatives, they require auxiliary data to estimate false positive and false negative rates, whereas we require only a single realization of the network. li2022 also consider how multiple measurements of a network can be used to estimate and adjust for measurement error rates, in a causal exposure-mapping framework.

Related work on causal inference with network misspecification includes savje2024b, which considers causal estimation when exposure maps are misspecified, and hardy2024, who propose a mixture model over potentially misspecified treatments in linear-in-means models. These works highlight the close relationship between misspecified treatments and mismeasured networks, which can be thought of as conceptual duals. zhang2024b presents an estimator for mismeasured networks with bounded degrees in experimental settings. li2025h considers ordinary least squares estimators of linear-in-means models under parametric misspecification. lewbel2023 and yu2022a consider estimating spillover effects when the network is entirely unobserved. spohn2023, chin2019a, and leung2022 similarly consider linear models for causal interference in precisely observed networks. liu2013 considers two-stage estimators for linear-in-means models in sampled networks, and propose instruments that account for this sampling.

Finally, our work connects to the broader literature on network autoregressive models. lesage2009 describe how the typical marginal effect interpretation does not apply to linear-in-means models due to non-linearity, and present methods to compute impact scores that retain this interpretation. vazquez-bare2023 discusses when coefficients have a causal interpretation (see also leung2022,mcfowland2021). Estimation approaches are given by ord1975, kelejian1998,lee2002,lee2003,lee2004,kelejian2007,lee2010, su2012, drukker2013, lin2010a and surveyed in bivand2021. Key identification results were given by bramoulle2009, and bramoulle2020 surveys identification in network autoregressive and linear-in-means models.

Models and Estimators

In this section, we formalize two network autoregressive models for peer influence: a standard peer contagion model and our proposed latent contagion model. We then develop two-stage least squares estimators for both models, and establish our main theoretical results showing that these estimators remain consistent under model misspecification. Our results essentially show that these models and estimators are interchangeable, which is useful because the estimators derived under the latent contagion model can easily be extended to handle noisy networks.

Setup

Consider a network with $n$ nodes encoded by a symmetric, non-negative adjacency matrix $\A \in \R_+^{n \times n}$. For each node $i \in [n]$, we observe an outcome $Y_i \in \R$ and covariates $\W_i \in \R^p$. We assume each node has an unobserved latent position $\Xpop_i \in \R^d$ that governs how node $i$ forms connections to other nodes. This framing draws on a large literature on latent space models, particularly the random dot product graph athreya2018, where connection probabilities are functions of these latent positions. Let $d_i = \sum_{j \neq i} \A_{ij}$ denote the degree of node $i$. For notational convenience, we define $\D = \diag(d_1, d_2, \dots, d_n)$ and the degree-normalized adjacency matrix $\G = \D^{-1} \A \in \R^{n \times n}$. If node $i$ is isolated with $d_i = 0$, we set $\G_{i \cdot} = \symbf 0$. Similarly, in the latent space, we define $\Apop = \E[\Xpop]{\A}$, $\dtilde_i = \sum_j \Apop_{ij} = \E[\Xpop_i]{d_i}$, $\Dtilde = \diag(\dtilde_1, \dots, \dtilde_n)$, and $\Gtilde = \Dtilde^{-1} \Apop \in \R^{n \times n}$.

Two Models of Contagion

We now present two complementary specifications for how peer influence propagates through social networks.

\paragraph{Peer contagion model.} The first specification is a standard spatial autoregressive model adapted to networks via latent positions:

align[align omitted — 294 chars of source]

Here, $\betay \in (-1, 1)$ measures how peer outcomes influence focal outcomes through the observed network structure. The latent positions $\Xpop_i$ model homophily, which is crucial for both statistical identification hayes2025b and causal inference shalizi2011. We call this the peer contagion model because influence operates over the degree-normalized adjacency matrix $\G$. We assume that $\varepsilon_i$ i.i.d. and mean zero with variance $\sigma_\varepsilon^2$.

\paragraph{Latent contagion model.} Our proposed alternative replaces the observed adjacency matrix with its conditional expectation:

align[align omitted — 331 chars of source]

Under the RDPG, where $\E[\Xpop_i, \Xpop_j]{\A_{ij}} = \Xpop_i^\top \Xpop_j$, the latent contagion model takes a particularly interpretable form:

equation[equation omitted — 207 chars of source]

In this specification, peer influence is inversely proportional to the cosine distance between nodes in latent space. All pairs of nodes exert influence on one another, with the strength of influence determined by latent proximity rather than the realization of individual edges. The latent contagion model reflects a fundamentally different view of social influence: rather than contagion traveling along realized edges, it diffuses based on the propensity for connection. Nodes close in latent space exert strong influence regardless of whether a specific edge forms, while distant nodes exert negligible influence.

remarkThe parameters $\symbf \beta$ and $\symbf \theta$ have fundamentally similar interpretations, and we introduce these differing parameters in an attempt to clarify notation. We use $\symbf \beta$ to indicate parameters in the peer contagion model, and $\widehat{\symbf \beta}$ to denote estimators derived under a peer contagion working model. Correspondingly, $\symbf \theta$ indicates parameters in the latent contagion model, and $\widehat{\symbf \theta}$ to denote estimators derived under a latent contagion working model. Later on, we will make statements of the form $\sqrt{n} (\widehat{\symbf \theta} - \symbf \beta) \to \mathcal{N}(0, \Sigma)$. Here, the use of $\beta$ as the underlying parameter indicates that the true data generating process is peer contagion, but $\widehat{\symbf \theta}$ indicates that we are estimating $\symbf \beta$ using an estimator derived under a latent contagion working model.

Both the peer and latent contagion models require a stochastic model for how the network $\A$ is generated. We adopt a flexible sub-gamma framework that encompasses binary, weighted, and count-valued networks.

definition[Sub-gamma network model] Let $\A \in \R^{n \times n}$ be a random symmetric matrix. Let $\Apop = \E[\Xpop]{\A}$ be the expectation of $\A$ conditional on $\Xpop \in \R^{n \times d}$, which has independent and identically distributed rows $\Xpop_1, \dots, \Xpop_n \in \R^d$. Assume $\Apop$ has rank $d$ and is positive semi-definite with eigenvalues $\spop_1 \ge \spop_2 \ge \cdots \ge \spop_d > 0 = \spop_{d+1} = \cdots = \spop_n$. Conditional on $\Xpop$, the upper-triangular elements of $\A - \Apop$ are independent $(\nu_n, b_n)$-sub-gamma random variables.

The sub-gamma family is broad, including Bernoulli, Poisson, Exponential, Gamma, and Gaussian distributions, as well as all bounded distributions boucheron2013,vershynin2020. We provide a formal definition of sub-gamma random variables along with a handful of related technical results in Appendix (ref). The generality of this framework allows us to handle diverse edge types, including binary friendships, weighted interaction frequencies, or count-valued communication volumes. As a consequence of this generality, our results require a comparatively high network density; forthcoming work by Hayes, Chandrasekhar, McCormick and Breza shows that the density requirements are less stringent in binary networks.

Identification

Identification in spatial autoregressive models is subtle, owing to their conditional specification via $\Y_i \mid \Y_{-i}, \W, \Xpop, \A$. The model can equivalently be written in its total law form (Equations (ref) and (ref)), as is standard in the Markov random field literature besag1974, rue2005a. The following identification result is crucial for both models.

lemma[martellosio2022] Consider a network with finite numbers of nodes $n$. For the peer contagion model (ref), suppose $\bbE[\be \mid \W, \Xpop, \G] = \symbf 0$ and $\abs*{\betay} < 1$. Then $\betanaught, \betaw, \betax, \betay$ are identified if and only if \begin{equation*} \rank \begin{bmatrix} \1_n \, \, \W \, \, \Xpop \, \, \G \W \, \, \G \Xpop \end{bmatrix} > \rank \begin{bmatrix} \1_n \, \, \W \, \, \Xpop \end{bmatrix} \end{equation*} An analogous result holds for the latent contagion model (ref) with $\Gtilde$ replacing $\G$, under the additional assumption that that $d \ge 2$, which rules out collinearity between $\Xpop$ and $\Gtilde \Xpop$.
remarkThe latent positions $\Xpop$ are only identified up to an orthogonal transformation $\Q$, which implies $\betax$ is also only identified up to $\Q$ hayes2025b. In some models, $\1_n$ may lie in the column space of $\Xpop$, in which case the intercept should be dropped.
remarkLinear-in-means models can suffer from asymptotic degeneracy when $\G \Y$ (or $\Gtilde \Y$) becomes collinear with $\begin{bmatrix} \1_n \, \, \W \, \, \Xpop \end{bmatrix}$, leading to parameters that are identified but inestimable hayes2025b. We assume throughout that the network structure prevents such degeneracy. For random dot product graphs, this requires either sufficient sparsity to prevent concentration of $\G \Y$ around its expectation, or sufficient degree heterogeneity to ensure linear independence. See Example 4 in hayes2025b for details.

The following Lemma clarifies the need for a bound on the magnitude of the contagion coefficient. If there is no such bound, the spillover can cause the responses in $\Y$ to diverge.

lemmaIf $|\beta| < 1$, then $\mI-\beta \G$ is invertible with all eigenvalues in the interval $(1 - \beta, 1 + \beta)$.
proofSince $\G = \D^{-1} \A$ is row stochastic, all its eigenvalues have absolute value at most $1$. Therefore, all eigenvalues of $\beta \G$ have absolute value at most $|\beta|$, implying that all eigenvalues of $\mI - \beta \G$ lie in $1 \pm \beta$ and are bounded away from zero.

Estimators

To construct feasible estimators, we require an estimate $\Xhat$ of the latent positions $\Xpop$. We use the adjacency spectral embedding sussman2014, a spectral estimate appropriate both for precisely observed networks and networks observed with additive noise. Other types of measurement error will require distinct estimators of $\Xpop$. We consider the empirical performance of some of these estimators in the simulation study of Section (ref), but leave detailed theoretical investigation of these settings to future work.

definition[Adjacency spectral embedding] Given a network $\A$, the $d$-dimensional adjacency spectral embedding is $\Xhat = \Uhat \Shat^{1/2} \in \R^{n \times d}$, where $\Uhat \Shat \Vhat^\top$ is the rank-$d$ truncated singular value decomposition of $\A$.

The adjacency spectral embedding provides a consistent estimate of $\Xpop$ in several settings, with rate varying depending on the precise assumptions sussman2012a, lyzinski2017, levin2022a, athreya2018.

lemmaUnder suitable regularity conditions, there exists a $d \times d$ orthogonal matrix $\Q$ such that \[ \max_{i \in [n]} \, \norm*{\Q \Xhat_i - \Xpop_i} = \op{1}. \]

If we knew $\Xpop$, we could construct two-stage least squares in the typical way for linear-in-means models kelejian1998, lee2003, but since $\Xpop$ is observed only via $\mA$, we instead use plug-in estimators that replace $\Xpop$ with $\Xhat$ wherever it appears.

definition[Peer contagion estimator] Let $\Zcheck = \begin{bmatrix} \1_n \, \, \W \, \, \Xhat \, \, \G \y \end{bmatrix} \in \R^{n \times (p + d + 2)}$ and $\Hcheck = \begin{bmatrix} \W \, \, \Xhat \, \, \G \W \, \, \G \Xhat \, \, \G^2 \W \, \, \G^2 \Xhat \end{bmatrix} \in \R^{n \times (3 p + 3 d)}$, and let $\Mcheck = \Hcheck (\Hcheck^\top \Hcheck)^{-1} \Hcheck^\top$ denote the projection matrix onto the column space of $\Hcheck$. We define our peer contagion estimator according to \begin{equation} \betahattsls = (\Zcheck^\top \Mcheck \Zcheck)^{-1} \Zcheck^\top \Mcheck \y . \end{equation}

For the latent contagion estimator, we construct $\Apophat = \Xhat \Xhat^\top$, $\widehat{d}_i = \sum_j \Apophat_{ij}$, $\Dhat = \diag(\widehat{d}_1, \dots, \widehat{d}_n)$, and

equation[equation omitted — 92 chars of source]
definition[Latent contagion estimator] Let $\Zhat = \begin{bmatrix} \1_n \, \W \, \, \Xhat \, \, \Ghat \y \end{bmatrix} \in \R^{n \times (p + d + 2)}$ and $\Hhat = \begin{bmatrix} \W \, \, \Xhat \, \, \Ghat \W \, \, \Ghat \Xhat \, \, \Ghat^2 \W \, \, \Ghat^2 \Xhat \end{bmatrix} \in \R^{n \times (3 p + 3 d)}$, and let $\Mhat = \Hhat (\Hhat^\top \Hhat)^{-1} \Hhat^\top$ denote the projection matrix onto the column space of $\Hhat$. Our latent contagion estimator is given by \begin{equation} \thetahattsls = (\Zhat^\top \Mhat \Zhat)^{-1} \Zhat^\top \Mhat \y . \end{equation}

In some cases, $\Hcheck$ and $\Hhat$ will have collinear columns; any subset of columns with rank $p + d + 1$ or greater is sufficient for identification bramoulle2009.

Main Results

We begin with some preliminary theoretical results, which show that under correctly-specified models, peer and latent contagion models that adjust for latent positions can be estimated by plugging in $\Xhat$ as an estimate for $\Xpop$. These results shows that it is possible to distinguish between localized effects on a network, as parameterized by $\betax$ and $\thetax$, and diffusions, as parameterized by $\betay$ and $\thetay$.

theorem[Peer contagion estimators under peer contagion] Suppose the data are generated according to peer contagion as in Equation (ref) and $\A$ follows a sub-gamma model as in Definition (ref). Under Assumptions (ref), (ref), (ref), and (ref), detailed in the Appendix, there exists a sequence of orthogonal matrices $\Qdes \in \R^{(p+d+2) \times (p+d+2)}$ such that \begin{align*} \sqrt{n} \paren*{\Qdes \betahattsls - \symbf \beta} & \to \mathcal{N} \paren*{\symbf 0, \sigmaeps^2 \paren*{\ZpeerOracle^\top \MpeerOracle \ZpeerOracle}^{-1}} \end{align*} where $\symbf \beta = (\betanaught, \betaw^\top, \betax^\top, \betay)^\top$.
theorem[Latent contagion estimators under latent contagion] Suppose the data are generated according to latent contagion as in Equation (ref) and $\A$ follows a sub-gamma model as in Definition (ref). Under Assumptions (ref), (ref), (ref), (ref), and (ref), detailed in the Appendix, there exists a sequence of orthogonal matrices $\Qdes \in \R^{(p+d+2) \times (p+d+2)}$ such that \begin{align*} \sqrt{n} \paren*{\Qdes \thetahattsls - \symbf \theta} & \to \mathcal{N} \paren*{\symbf 0, \sigmaeps^2 \paren*{\ZlatOracle^\top \MlatOracle \ZlatOracle}^{-1}}, \end{align*} where $\symbf \theta = (\thetanaught, \thetaw^\top, \thetax^\top, \thetay)^\top$.

Our central theoretical contribution establishes that estimators derived under one model remain consistent when the true data generating process follows the other model, if the observed network is distributed as a random dot product graph. Said another way, correctly-specified two-stage least squares estimators are asymptotically normal in both the peer and latent contagion models.

theorem[Latent contagion estimators under peer contagion] Suppose the data are generated according to peer contagion as in Equation (ref) but we use latent contagion estimators from Definition (ref). Suppose $\A$ follows a sub-gamma model as in Definition (ref). Under Assumptions (ref), (ref), (ref), (ref) and (ref), detailed in the Appendix, there exists a sequence of orthogonal matrices $\Qdes \in \R^{(p+d+2) \times (p+d+2)}$ such that \begin{align*} \sqrt{n} \paren*{\Qdes \thetahattsls - \symbf \beta} & \to \mathcal{N} \paren*{\symbf 0, \sigmaeps^2 \paren*{\ZlatOracle^\top \MlatOracle \ZlatOracle}^{-1}} . \end{align*}
theorem[Peer contagion estimators under latent contagion] Suppose the data are generated according to latent contagion as in Equation (ref) but we use peer contagion estimators from Definition (ref). Suppose $\A$ follows a sub-gamma model as in Definition (ref). Under Assumptions (ref), (ref), (ref) and (ref), detailed in the Appendix, there exists a sequence of orthogonal matrices $\Qdes \in \R^{(p+d+2) \times (p+d+2)}$ such that \begin{align*} \sqrt{n} \paren*{\Qdes \betahattsls - \symbf \theta} & \to \mathcal{N} \paren*{\symbf 0, \sigmaeps^2 \paren*{\ZpeerOracle^\top \MpeerOracle \ZpeerOracle}^{-1}} . \end{align*}

Proofs of Theorems (ref), (ref), (ref) and (ref) can be found in the Appendix. The general proof strategy is to show that $\betahattsls$ and $\thetahattsls$ are close to “oracle” estimates based on using the true latent positions. Appendix (ref) discusses convergence of these oracle estimators to the true parameters. Appendices (ref) and (ref) show convergence of our estimators defined above to these oracle estimators. We include, additionally, proofs showing that the feasible ordinary least squares estimators converge to their oracle counterparts, although oracle ordinary least squares estimators are only consistent and asymptotically normal in dense networks lee2002,gupta2019.

Our misspecification results generalize to a statement about asymptotic equivalence of the corresponding functionals. Recall that the population design matrices of the peer and latent contagion models are, respectively,

equation[equation omitted — 236 chars of source]

Consider the corresponding “projection parameters”, or the population coefficients corresponding to projection of outcomes onto these design matrices, and letting $F$ be an appropriately supported cumulative distribution function (see li2025h for a discussion of these parameters in the network regression context),

align[align omitted — 435 chars of source]

Let $F_\beta$ denote the law of the peer contagion process and $F_\theta$ the law of the latent contagion process. When the expectation is taken with respect to the true generating process, the projection parameters simplify to the corresponding regression coefficients,

equation[equation omitted — 212 chars of source]

These projection parameters are not, in general, equal under the latent and peer contagion models. However, fixing either the peer or latent contagion model, the following theorem shows that they are asymptotically equivalent, in the sense that the difference between the two parameters is negligible relative to the typical estimation error.

theoremUnder Assumptions (ref), (ref), (ref) and (ref), and supposing the latent contagion model in Equation (ref) holds, then \begin{align*} \norm*{\widetilde{\tau} (F_\theta) - \tau (F_\theta)} = \norm*{\, \symbf \theta - \tau (F_\theta)} = \norm*{\bbE \brac*{\ZlatOracle^\top \ZlatOracle}^{-1} \bbE \brac*{\ZlatOracle^\top \Y} - \bbE \brac*{\ZpeerOracle^\top \ZpeerOracle}^{-1} \bbE \brac*{\ZpeerOracle^\top \Y}} = o \paren*{n^{-1/2}}. \end{align*} If the peer contagion model in Equation (ref) holds instead of the model in Equation (ref), then, with Assumption (ref) in place of Assumption (ref), \begin{align*} \norm*{\tau (F_\beta) - \widetilde{\tau} (F_\beta)} = \norm*{\, \symbf \beta - \widetilde{\tau} (F_\beta)} = \norm*{\bbE \brac*{\ZpeerOracle^\top \ZpeerOracle}^{-1} \bbE \brac*{\ZpeerOracle^\top \Y} - \bbE \brac*{ \ZlatOracle^\top \ZlatOracle}^{-1} \bbE \brac*{\ZlatOracle^\top \Y}} = o \paren*{n^{-1/2}}. \end{align*}

A proof can be found in Appendix (ref). In the above theorem, we use the “assumption lean” representation of regression coefficients, or the “projection estimands”, which are the definitions of $\tau$ and $\widetilde{\tau}$ given in (ref) and (ref), respectively. These functionals are non-parametric. That is, $\widetilde{\tau}$ is a projection parameter that one might want to estimate even if the true data generating model is the peer contagion model and the model is indexed by $\symbf \beta$, with $\symbf \theta$ undefined. The key idea of Theorem (ref) is that the true parameter and the projection parameter are within parametric estimation error of one another, so they are functionally indistinguishable in the asymptotic limit.

This result has substantial implications for estimating contagion effects in noisy networks. Suppose that outcomes are generated according to the peer contagion model in Equation (ref), but the network $\A$ is unobserved and one only has access to a noisy variant $\Anoise$. In this case, treating $\Anoise$ as the true network and plugging it into typical estimators such as $\betahattsls$ will ignore measurement error in $\Anoise$ and typically lead to inconsistent estimation, among other issues. In particular, it is challenging to target the parameter $\symbf{\beta}$, because this requires projecting onto $\ZpeerOracle$, which is a function of an unknown network $\A$. Theorem (ref), however, suggests a path forward: we can instead target the parameter $\widetilde{\tau}$, which does not equal $\symbf{\beta}$, but is functionally indistinguishable. Crucially, to estimate $\widetilde{\tau}$ we do not need to project outcomes onto $\ZpeerOracle$, but rather onto $\ZlatOracle$, which depends on the low-rank expectation $\Apop$ rather than the precise network $\A$. This is a substantially easier task: we do not need $\A$ proper, only its principal subspace. Conveniently, given a noisy observation of a network $\Anoise$, there are many off-the-shelf methods to estimate the principal subspace of $\A$.

This leads to a potentially generic recipe for estimating contagion effects in noisy networks: find a subspace estimator tailored to the noise process, and target $\widetilde{\tau}$ using an estimator like $\thetahattsls$. The variance of the corresponding estimator will depend on how fast the subspace estimator converges. For sufficiently fast estimators, such as the adjacency spectral embedding, there is no asymptotic variance penalty due to estimating the principal subspace. In fact, this approach and our results thus far are sufficient to develop an estimator for contagion effects in noisy weighted networks. Suppose that $\Anoise = \A + \mE$, where $\mE$ is mean-zero sub-gamma noise. In this case, $\Anoise$ still satisfies the assumptions of the subgamma network model (Definition (ref)), and thus the adjacency spectral embedding is still a consistent estimator of the principal subspace of $\A$, with only a slight adjustment to the sub-gamma parameters in the convergence rate. Thus, $\thetahattsls$ immediately accommodates sub-gamma noise in the adjacency matrices.

Figure (ref) offers some visual intuition for the generic nature of the network smoothing strategy, showing estimates of $\Apop$ obtained from networks with various types of corruption. Even when 10% of edges are missing, matrix completion methods produce estimates $\Apophat$ that closely approximate the true latent structure. In the subsequent section, we provide evidence via simulations that network smoothing is a viable strategy to estimate contagion in networks with missing edges, in ego-centrically sampled networks, and in networks where only aggregated relational data is available.

figure[figure omitted — 584 chars of source]

Simulation study

We now verify via simulation that $\thetahattsls$ and $\betahattsls$ are consistent and asymptotically normal estimates of contagion effects (Fig. (ref)), and that they obtain the expected $n^{-1/2}$ convergence rate predicted by our theoretical results. All networks in our simulations below are generated from a degree-corrected stochastic blockmodel (Definition (ref)) with $n$ nodes and five equally probable blocks. In particular, we use a Poisson stochastic blockmodel where edges are Poisson distributed.

figure[figure omitted — 1,021 chars of source]
definition[Poisson Degree-Corrected Stochastic Blockmodel] The Poisson degree-corrected stochastic blockmodel rohe2018, karrer2011 is an undirected model of community membership, with $d$ communities. Each node, indexed by $i \in [n]$, is assigned a block $z_i \in [d]$ with probability $\Pr({z_i = k}) = \pi_k$, and a degree-correction parameter $\xi_i$, which describes the propensity of vertex $i$ to connect with other nodes. Conditional on block memberships and degree-correction parameters, edges are generated independently between every pair of vertices in the network according to a Poisson distribution with parameter $\rho_n \, \xi_i \symbf B_{z_i, z_j} \xi_j$. That is, the expected number of edges between two vertices depends on their community memberships, their degree correction parameters, a positive semi-definite matrix $\symbf B \in [0, 1]^{d \times d}$ of inter-block edge formation probabilities, and a scaling factor $\rho_n \in [0, 1]$, which may vary with $n$.

For our simulation study, we take there to be $d = 5$ blocks, and sample diagonal elements of $\symbf B$ from a $\mathrm{Uniform}(0.75, 0.85)$ distribution and off-diagonal elements of $\symbf B$ from a $\mathrm{Uniform(0.01, 0.05)}$ distribution, such that networks are strongly assortatively, mostly forming edges within blocks. We sample degree correction parameters according to $\xi_i \sim \mathrm{Exponential(1/3)} + 1$. Shifting the distribution of $\xi_i$ away from zero limits the number of isolated nodes. The sparsity parameter $\rho_n$ is set so that the expected mean degree of the network is $n^{3/4}, n^{1/2}$ or $n^{1/4}$. Once we have generated these parameters, we compute $\Xpop = \Upop \Spop^{1/2}$ where $\Upop \Spop \Upop^\top$ is the eigendecomposition of the low-rank expectation $\E[z_1,z_2,\dots, z_n, \theta]{A}$. Errors $\be$ are sampled from a standard normal distribution, and we include three covariates $\W_1, \W_2, \W_3$, also sampled from a standard normal. We set $\betanaught = \thetanaught = 0$, $\betay = \thetay = 0.2$, $\betaw = \thetaw = (5, 5, 5)$ and $\betax = \thetax = (2, 2, 2, 2, 2)$. Then $\Y$ is generated according to Equation (ref) or Equation (ref), depending on whether we are under peer or latent contagion, respectively. For both models, we compute $\betahattsls$ and $\thetahattsls$, and measure the estimation error to the corresponding model coefficients. Since $\betax$ is only identified up to orthogonal rotation, we perform Procrustes alignment cape2019c between $\Xhat$ and $\Xpop$ to investigate recovery of $\betax$. Note that we set $\betanaught = \thetanaught = 0$ to simplify this alignment step.

In our experiments, we vary the sample size $n$ (i.e., the number of vertices) on a logarithmic scale, considering $n \in \set{100, 163, 264, 430, 698, 1135, 1845, 3000}$, and replicate our experiments $100$ times for each simulation setting. Figure (ref) shows the mean squared error of the estimated coefficients as a function of the number of nodes. Mean squared error for $\thetahattsls$ and $\betahattsls$ decreases at $n^{-1/2}$ rates in both the latent and peer contagion models, exactly as dictated by our theory.

Noisy networks

We additionally investigate how network smoothing performs when the network is mismeasured. We consider the exact same simulation setting as before, but now we observe a noisy adjacency matrix $\Anoise$ rather than $\A$. We consider only the latent contagion estimator $\thetahattsls$, since the latent space formulation allows us to smooth away noise in the network. In some cases, estimating $\Upop$ and $\Spop$ via the singular value decomposition is impossible because of the noise process (for instance, due to missing data). When this is the case, or when there is an estimator of $\Upop$ and $\Spop$ designed to handle the particular form of the noisy matrix $\A$, we use that more appropriate estimator instead of the singular value decomposition. The estimators are introduced below, alongside the various noise processes under consideration. In the first four cases, existing estimators are capable of recovering the singular values and singular vectors of $\Apop$. In the last two settings, we are unaware of principal subspace estimators, and expect difficulties with estimation.

enumerate• Baseline: The network is observed without noise, so $\Anoise = \A$ and $\Apophat$ is estimated by taking the rank $k$ singular value decomposition of $\Anoise$. $\Xhat$ is constructed via $k$-dimensional ASE applied to $\Apophat = \Uhat \Shat \Uhat$, so that $\Xhat = \Uhat \Shat^{1/2}$. $\Ghat$ is constructed by row-normalizing $\Apophat$ and the setting the diagonal elements to zero. $\Xhat$ and $\Ghat$ are then plugged in as estimates of $\Xpop$ and $\Gtilde$ in $\thetahattsls$. • Gaussian noise: We observe $\Anoise = \A + \symbf E$, where $\symbf E$ is a symmetric matrix with i.i.d. entries $\varepsilon_{ij} \sim \mathcal{N}(0, \sigma^2)$ for $i < j$ and $\varepsilon_{ii} = 0$. We compute $\Apophat$ as the rank-$k$ truncated singular value decomposition of $\Anoise$, and construct $\Ghat$, $\Xhat$ and then $\thetahattsls$ as in the baseline case. levin2022a justifies this singular value decomposition as applied to $\Anoise$ rather than $\A$. • Missing edges: We observe $\Anoise = \A$, except $30\%$ of entries of $\Anoise$ are missing at random, with missingness independent of the network. $\Upop$ and $\Spop$ are estimated via matrix completion, in particular the AdaptiveImpute algorithm of cho2019. Estimated singular values and vectors are directly substituted for the more typical $\Uhat$ and $\Shat$ from singular value decomposition in the known $\A$ case, which in turn allows computation of $\Apophat, \Ghat$ and $\Xhat$ in the same fashion as the previous two settings. • Aggregated relational data: We observe $\Anoise = \A \W$ where $\W$ are traits (also observed) that are correlated with latent $\Xpop$. $\W$ has the same dimensions as $\Xpop$, and each column of $\W$ is sampled from a multivariate normal distribution with correlation 0.8 to the corresponding column of $\Xpop$. In the aggregated relational data setting, $\Upop$ is estimated by the left singular vectors of $\Anoise$, and we denote these estimates $\bar{\Upop}$. $\Spop$ is then estimated via $\bar{\Upop}^\top \Y \W^\top \bar{\Upop} (\bar{\Upop}^\top \W \W^\top \bar{\Upop})^{-1}$, and these estimates are used as plug-in replacements for $\Uhat$ and $\Shat$. Theory and motivation for these estimators is under development in forthcoming work by Hayes, Chandrasekhar, McCormick and Breza. The aggregated relational data framework is detailed in breza2020, breza2023a. • Ego-centric data: In ego-centric sampling, half of the nodes in the network are selected at random, and only edges incident to those nodes are observed. Unobserved edges are imputed using Algorithm 1 of chan2023a, and then $\Apop$ is estimated via the full matrix recovery technique presented in the same manuscript. Let $\A_{11}$ be the ego-ego block and $\A_{12}$ be the ego-nonego block. We compute a rank-$k$ approximation $\tilde{\Apop}_{11} = \Upop_1 \symbf \Spop_1 \Vpop_1^\top$ of $\A_{11}$, and then estimate the nonego-nonego block as $\Apophat_{22} = \A_{12}^\top \tilde{\Apop}_{11}^\dagger \A_{12}$. The remaining blocks $\Apophat_{11}$ and $\Apophat_{12}$ are recovered via a rank-$k$ approximation of the observed part of the network $\begin{bmatrix} \A_{11} & \A_{12} \end{bmatrix}$. • \textbf{Degree capped}: The network is observed with censored degrees, where for each node $i$, at most $d_{\text{max}} = 20$ incident edges are available. If a node has more than $20$ edges, edges are removed uniformly at random until the constraint is met. We compute $\Apophat$ as the rank-$k$ truncated singular value decomposition of the resulting censored adjacency matrix $\Anoise$. • \textbf{Edges flipped}: $\Anoise$ is a version of $\A$ when $15\%$ of edges in the network have been flipped. To preserve the total number of edges in the network, this is implemented via random edge swapping, where $15\%$ of edges in the network are swapped with a random edge with different edge value. $\Upop$ and $\Spop$ are estimated via singular value decomposition of $\Anoise$.

The results of these simulations are in Figure (ref), which compares mean squared error for $\betay$ and $\thetay$ when using the estimates of $\Xpop$ and $\Gtilde$ defined above. We see that the network smoothing approach is able to recover both $\betay$ and $\thetay$ in settings where $\A$ is contaminated with Gaussian noise, when $\A$ has missing edges, and when only an aggregated relational data variant of $\A$ is observed. This matches our expectations that consistent estimation of peer effects is possible whenever principal subspace estimation is possible. Under degree censoring and edge flip noise mechanisms, principal subspace estimation is challenging, and the singular value decomposition is a poor estimator, such that $\betay$ and $\thetay$ cannot be recovered accurately.

figure[figure omitted — 863 chars of source]

Sparsity

We additionally investigate how $\thetahattsls$ and $\betahattsls$ perform in sparser, correctly observed networks. In Figure (ref) we report mean squared error when the average degree is $n^{1/2}$, and in Figure (ref) we report mean squared error when the average degree is $n^{1/4}$. In both cases, we observed worse performance relative to the denser $n^{3/4}$ average degree case. When the average degree is $n^{1/2}$, mean squared error is more volatile, but Figure (ref) does still suggest that $\thetahattsls$ and $\betahattsls$ are consistent under peer contagion. Under latent contagion, $\thetahattsls$ appears consistent, but $\betahattsls$ struggle to recover the intercept (coefficient $Z_1$) and the peer effect. We suspect that this is primarily due to finite sample collinearity issues, as degree heterogeneity is necessary to differentiate the intercept from the contagion term, and degree heterogeneity is less pronounced in sparser graphs. In the sparsest setting, with average degree $n^{1/4}$, the impact of sparsity is more severe. Figure (ref) suggests that all estimators experience slower rates of convergence and may not even be consistent, under both models, although $\betahattsls$ is potentially consistent under the peer contagion model. We believe that this degradation in performance is primarily attributable to noise in the estimates $\Xhat$ around $\Xpop$.

These simulations show that the sparsity can degrade the performance of the estimators that we have proposed, via two mechanisms: sparsity can reduce degree heterogeneity, leading to collinearity issues, and it may degrade the accuracy of the plug-in estimate $\Xhat$ of $\Xpop$. These simulations suggest that applied analyses using our estimators should assess the quality of estimates $\Xhat$ (namely, assess whether the estimates $\Xhat$ correspond to meaningful structure in the network) and should consider the possibility of variance inflation in regression estimates due to collinearity. In sparse networks with strong block diagonal structure and limited degree heterogenity, the intercept and $\Xpop$ may be collinear, and it may be reasonable to drop a column for collinearity reasons.

figure[figure omitted — 862 chars of source]
figure[figure omitted — 862 chars of source]

Data Application

To demonstrate our methods, we applied our estimators to network data collected during the Teenage Friends and Lifestyle Study, reported in michell1996, michell1997, michell1997a, and michell2000a. Recently, hayes2025 and dimaria2022a investigated network-mediation using this data, studying how sex influenced network position, and how network position subsequently influenced smoking behaviors in adolescents. These analyses assumed that there were no peer effects on smoking after accounting for latent network position. We re-analyzed the same data to investigate if smoking exhibits spillovers in addition to being localized within the adolescent social network.

Data

The Teenage Friends and Lifestyle Study collected three waves of survey data in a secondary school in Glasgow, beginning in January 1995. Students in the study filled out a questionnaire about their lifestyle and risk-taking behaviors, including alcohol, tobacco and drug use, and additionally were asked to list six of their friends. michell2000a found that smoking was mostly concentrated in friend groups composed of popular girls, unpopular students, and trouble-makers: “risk taking behaviour was heavily polarized within social categories so that, for instance, groups of individuals (and their peripherals) were in general either risk-taking or non-risk-taking”.

The social network was collected by asking students “who are your best friends”, and allowing adolescents to list up to six responses. We considered only data from the first wave of the survey, which included 153 adolescents. Sex and tobacco use were self-reported as nominal features with levels “Male” and “Female”; and “Never”, “Occasional,” and “Regular,” respectively. To match the analyses of hayes2025 and dimaria2022a, for the tobacco use measure, we combined “Occasional” and “Regular” into a single level, and compared smokers with non-smokers. Also to match the analysis of hayes2025, we treated age (continuous) and church attendance (nominal) as possible confounders and thus included these variables as controls.

We computed the adjacency spectral embedding of the social network $\A$. In the Glasgow data, the social network is directed: an edge $i \to j$ indicates that student $i$ listed student $j$ as friend. This directedness means that students have two distinct co-embeddings corresponding to their propensity to send out-edges and receive in-edges. Letting $\widehat{\A} \approx \Uhat \Shat \Vhat^T$ be the truncated singular value decomposition of $\A$, the left co-embedding $\widehat{\mL} = \Uhat \Shat^{1/2}$ described how students in the network send edges (i.e., claim friends), and the right co-embedding $\Xhat \equiv \Vhat \Shat^{1/2}$ described how students receive edges (i.e., are claimed as friends). Our results used the right co-embeddings $\Xhat$. We did not select any particular dimension $d$ for the latent space. Instead, we repeated our analysis for many values of $d$, to investigate the sensitivity of our results to the dimension of the latent space. Once we obtained embeddings $\Xhat$ via the singular value decomposition, we performed a multiverse analysis, estimating regression coefficients using $\thetahattsls$ and $\betahattsls$. For each estimator, we considered two variants, one including $\Xhat$ as covariates, to adjust estimates for latent positions in the network, and one without including $\Xhat$ as covariates. That is, we used two-stage least squares estimators derived under all four of the following generative models:

align[align omitted — 449 chars of source]

Recall that $\W$ consists of age and church attendance. The estimators $\thetahattsls$ required an estimate of the dimension $d$ of the latent positions $\Xpop$, as does $\betahattsls$ when $\Xpop$ are included as covariates in the model. We repeated our analysis for $2 \le d \le 25$ in order to understand how the dimension of the embedding affected estimates. For each estimate, we reported a 95% asymptotic confidence interval in Figure (ref).

figure[figure omitted — 867 chars of source]

Results

Estimates of the smoking spillovers varied moderately across the multiverse analysis. The simplest and most consistent story emerged among estimates that do not adjust for latent homophily by including $\Xpop$ in the regression specification. These estimates, visualized in blue, were consistently large and statistically differentiated from zero, suggesting that smoking does exhibit spillover effects. However, estimates that adjusted for latent homophily by including $\Xpop$ told a different story. In most cases, including $\Xpop$ in the regression reduced the point estimate of the contagion coefficient, and the associated confidence interval often contained zero. However, these confidence intervals remained wide, indicating substantially uncertainty about the presence or absence of a spillover effect after accounting for localized smoking behavior in the network via the $\Xhat$ terms.

Another set of comparisons, between the latent contagion estimators and the peer contagion estimators, was also informative. In particular, Theorems (ref), (ref), (ref), and (ref) prove that latent and peer contagion estimators are asymptotically equivalent under the random dot product graph, provided that the network is precisely observed. However, we observed some deviation between the latent contagion and peer contagion estimators, with latent contagion estimates indicating lower levels outcome spillover across most embedding dimensions. There numerous possible explanations for the difference between the latent and peer contagion estimates: (1) the peer network might have been observed with noise, (2) a random dot product model may not haven been appropriate for the network (in which case we would prefer the estimates based on the peer contagion model), or (3) the network may have been too small (recall $n = 153$ nodes) for asymptotics results to applicable.

Lastly, we observed that estimates under peer contagion model were fairly stable as a function of the embedding dimension $d$, but estimates under the latent contagion model were somewhat more volatile in the embedding dimension, with estimates of the contagion coefficient collapsing towards for the specific values $d = 10, 18, 21$. We suspect this volatility as a function of embedding dimension was related to the small sample size the limited number of smokers in the network.

Altogether, our analysis indicated that there was substantial uncertainty about the presence and scale of the contagiousness of smoking in the adolescent social network, after accounting for the localized nature of smoking in the network by control from $\Xhat$. This uncertainty mirrored previous results suggesting that peer effects and homophily are challenging to distinguish in social networks hayes2025b, shalizi2011. A possible next step to differentiate between these effects in the adolescent social network would be to consider longitudinal models of smoking behavior, such as those proposed in chang2025b.

Discussion

We have shown that under low-rank network models, contagion over a true network is asymptotically equivalent to contagion over a smoothed, latent adjacency matrix. This equivalence enables consistent peer effects estimation even when the observed network contains measurement error, provided we can reliably estimate the network's eigenspace. Any method that reliably estimates the eigenspace of $\bbE [\A]$—whether through spectral embedding for sub-gamma noise, matrix completion for missing edges, or debiasing techniques for systematic measurement error—can be combined with our latent contagion framework to estimate peer effects in challenging data settings.

While this paper focuses on parameter estimation rather than causal identification, the equivalence we establish has implications for causal inference under interference. Under appropriate conditional ignorability assumptions, network autoregressive parameters can have causal interpretations vazquez-bare2023, leung2022, mcfowland2021. Our results suggest that when peer influence operates through latent structure, causal effects may be more accurately estimated by defining exposures based on latent proximity rather than observed edges.

Acknowledgements

We thank Hyunseung Kang, Ralph Trane, Karl Rohe, Edward McFowland, Arun Chandrasekhar, Michael Newton, Tyler McCormick, Vincent Boucher, Ben Golub and the attendees of the Network Science in Economics 2026 conference, as well as attendees of the University of Wisconsin--Madison IFDS Ideas Seminar for their helpful comments and suggestions. Support for this research was provided by the University of Wisconsin--Madison, Office of the Vice Chancellor for Research and Graduate Education with funding from the Wisconsin Alumni Research Foundation, as well as NSF grants DMS 2052918 and DMS 2023239. Support for this research was also provided by American Family Insurance through a research partnership with the University of Wisconsin--Madison's American Family Insurance Data Science Institute.