EconBase
← Back to paper

Non-robustness of diffusion estimates on networks with measurement error

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

87,260 characters · 14 sections · 42 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Non-robustness of diffusion estimates on networks with measurement error

abstractNetwork diffusion models are used to study disease transmission, information spread, technology adoption, and other socio-economic processes. We show that estimates of these diffusions are highly non-robust to mismeasurement. First, even when the network is measured perfectly, small and local mismeasurement in the initial seed generates a large shift in the locations of the expected diffusion. Second, if instead the initial seed is known, even a vanishingly small share of missed links causes diffusion forecasts to be significant under-estimates. Forecast failure depends critically on the geometry of measurement error: we provide sufficient conditions for catastrophic failure when missing links bridge distant network regions (acting as shortcuts), and sufficient conditions for robustness when missing links are a uniformly, randomly thinned subset of the full network (preserving network structure). Such failures exist even when the basic reproductive number is consistently estimable. We explore difficulties implementing possible solutions and conduct simulations on synthetic and real networks.

Researchers and policymakers use network data to estimate models of diffusion---quantifying the extent of illness or technology adoption, summarizing dynamics (e.g., ${\mathcal R}_{0}$), and targeting interventions. See anderson1991infectious, jackson2009, jackson2011diffusion, and sadlere2023 for applications, including settings with strategic behavior.

This paper shows that even tiny mismeasurement of the network creates large forecast errors in diffusion. We identify a general mechanism: on networks with polynomial expansion---capturing structured interactions from geography, social groups, or institutions---even a vanishingly small number of missing links that are geometrically unaligned with the observed network change the expansion properties, generating catastrophic forecast errors. This holds even when the econometrician has perfect knowledge of either the network or the initial seed, and over policy-relevant intermediate horizons: long enough for diffusion to matter, short enough that it has not saturated the network. The numerical observation of watts1998collective---that random shortcuts on a ring lattice dramatically compress path lengths---is a special case of this principle; our framework explains why it occurs and characterizes precisely when it does and does not arise.

Crucially, not all missing links cause these problems. When missing links mirror the structure of observed links---what we term aligned error---forecasts remain asymptotically correct. Forecast failure occurs precisely when missing links are more dispersed than observed links: they act as shortcuts bridging distinct parts of the observed graph, accelerating expansion beyond what the observed network would predict. We call this unaligned error. This characterization provides actionable guidance for applied researchers: the key question is not how many links are missing, but whether data collection is more likely to miss links that resemble observed links (aligned---safe for forecasting) or links that bridge different network regions (unaligned---dangerous).

We show five key results: (i) predictions of where diffusion goes are sensitive to local uncertainty of the initial seeding (Theorem (ref)); (ii) predictions of diffusion counts will be catastrophically under-estimated with even vanishingly small unaligned measurement error of the network (Theorem (ref)); (iii) when measurement error is instead aligned, forecast failure vanishes (Theorem (ref)); (iv) while aggregated quantities such as the basic reproductive number ${\mathcal R}_0$ can be estimated consistently, they provide limited information for disaggregated forecasts; (v) because the measurement error is so small, most data augmentation approaches are ineffectual.

For intuition, consider a network where connections depend on observable proximity (geography, school, work) or latent factors hoffrh2002. A ball around the initial seed enumerates possibly affected nodes, expanding polynomially. Even with a perfectly known network, nearby seeds produce substantially diverging balls---misleading conclusions about where diffusion goes. The mismeasured seed can be within a neighborhood proportional to the forecast horizon, so the error vanishes in relative terms.

Now suppose the seed is known but some unaligned links are missed. If any missed link reaches beyond the seed's ball on the observed graph, the diffusion escapes to an unexposed region. This jump need not span a large physical distance---it must only bridge parts far apart in graph distance on the observed network. Missing links function as escape hatches that alter the effective geometry of diffusion---a mechanism that underlies the watts1998collective observation, but which operates on any polynomial-expansion network, not just ring lattices. Our general theorems allow each node to link to a vanishing fraction of the population with arbitrary structure, nesting cases of only local mismeasurement that is geometrically distinct from $L_n$.

The interaction between both issues, seed sensitivity and jump links, is multiplicative. Since different local seeds influence disjoint local regions, they encounter different error links and jump to different distant parts of the network. This causes the predictions to then diverge further as each secondary wavefront generates its own independent jumps. A local seed perturbation cascades into forecast errors spanning the entire network.

This mechanism also clarifies why aligned error is comparatively benign. When missing links are aligned, they do not bridge distant regions of the graph, so they furnish no escape hatches. Provided the number of such omissions remains small, the diffusion process stays confined to the neighborhood of the seed's ball, and forecasts remain accurate because no secondary wavefront can leap to an unexposed part of the network.

Missing links are a common concern wang2012measurement,sojourner2013identification,chandrasekharl2010,advani2018credibly,griffith2022name, but our paper highlights the impact of even the smallest errors on diffusion forecasts. Mismeasurement arises in three ways. First, aggregation into compartments (e.g., location-by-age-by-occupation) may match average interaction patterns but miss heterogeneity and cross-compartment connections acemoglu2021optimal,farboodi2021internal,fajgelbaum2021optimal,candogan2025network. Second, surveys may focus on local connections (within a school or village) while ignoring others. Third, a network snapshot may not capture links relevant to diffusion by the time the process reaches an individual -- in many cases, the network itself is forecasted from mobility or other interaction data. All three sources tend to produce unaligned error: they disproportionately miss links that bridge distant parts of the observed network while retaining local ones.

Formally, we study $n$ agents in an undirected, unweighted network $G_n$ with $n \to \infty$. A SIR diffusion proceeds for $T_n$ periods: each activated node transmits i.i.d.\ with probability $p_n$ to neighbors and is then removed. The true network $G_n = L_n \cup E_n$, where $L_n$ is the observed base network (with polynomial expansion) and $E_n$ is an unobserved error graph.\footnote{We focus on missing links, the primary concern in practice griffith2022name. When the econometrician both misses and incorrectly adds links, the problem is more complex.} Polynomial expansion is natural for latent-space models hoffrh2002; we verify our conditions for random geometric graphs in Example (ref) and Appendix (ref). We use activated to nest “infected,” “informed,” and “adopted” jackson2007diffusion.

Sections (ref) and (ref) establish the negative results: seed sensitivity (Theorem (ref) and Corollary (ref)) and catastrophic under-estimation from vanishing measurement error (Theorem (ref)). Section (ref) establishes the positive result: under aligned error, forecast failure vanishes (Theorem (ref)), contrasting sharply with Theorem (ref). Section (ref) shows that $p_n$ and ${\mathcal R}_0$ remain consistently estimable despite forecast failures, while Section (ref) shows that potential remedies are ineffectual.\footnote{alimohammadi2023epidemic make a similar point, showing that a local estimation algorithm on sampled network data asymptotically identifies the correct SIR compartment proportions.} Section (ref) examines these results through three empirical applications: a CA/NV mobility exercise showing that how links are missed matters more than how many; NYC's micro-cluster zoning, where a degree-matched mobility network identifies ${\sim}30\%$ different buffer zone neighborhoods and better predicts case growth; and the targeting methodology of beaman2021can, which is most fragile in the villages where it appears most valuable. Simulations on synthetic networks confirming finite-sample behavior appear in Appendix (ref).

Model

\paragraph*{Environment} Consider a set of $n$ observed nodes $V_n$. We model the network as a random undirected and unweighted graph $G_n \coloneqq (V_n, L_n \cup E_n)$, where $L_n$ represents the “base” links and $E_n$ represents the missing links. The base network $L_n$ is fixed and perfectly observed by the econometrician. Each link in $E_n$ is formed independently according to $\mathrm{Ber}(\beta_{ij,n})$, where the link probability $\beta_{ij,n}$ may vary across pairs. For expositional simplicity, we focus in the main text on the homogeneous case where $\beta_{ij,n} = \beta_n$ for all pairs $ij$. Appendix (ref) expands to a heterogeneous case for $\beta_{ij,n}$, which requires additional assumptions on the “expansiveness” of the support of $E_n$.\footnote{Formally, the assumption requires that the support of links in $E_n$ provides sufficiently broad coverage across the node set. Nonetheless, realized networks $E_n$ remain sparse with high probability. The i.i.d.\ case trivially satisfies this expansive support condition. See Remark (ref) for the connection to the small-worlds framework of watts1998collective.}

The links in $E_n$ are unobserved, so randomness in $G_n$ stems from the random realization of $E_n$. Our model focuses exclusively on mismeasurement due to link omission; we rule out false positives (spurious links). By a slight abuse of notation, we use $L_n$ and $E_n$ to denote both edge sets and the spanning subgraphs on $V_n$.

The diffusion process spreads over the network $G_n$ following a standard Susceptible-Infected-Removed (SIR) model with i.i.d.\ passing probability $p_n$. Each node, once activated, remains active for a single period during which it independently transmits the infection to each neighbor with probability $p_n$. After this period, the node is removed and cannot be reactivated.

To formalize the stochastic dynamics, we define $P_n(G_n)$ as a random percolation on $G_n$: a directed graph in which each undirected edge of $G_n$ generates two directed edges, each activated independently with probability $p_n$. Under this representation, the diffusion process from an initial seed is equivalent to a deterministic traversal through the realized percolation $P_n(G_n)$. Consequently, all randomness in the diffusion is fully captured by the random realization of $P_n(G_n)$. Similarly, we define $P_n(L_n)$ as the percolation restricted to the base network $L_n$, obtained by restricting $P_n(G_n)$ to edges in $L_n$.

Our analysis is asymptotic in both $n$, the number of nodes, and $T$, the forecast horizon. We consider a sequence of graphs $\{G_n\}_{n=1}^\infty = \{(L_n, E_n)\}_{n=1}^\infty$ where each $E_n$ is drawn randomly and the base network $L_n$ grows with $n$. The forecast horizon is a function of network size, $T = T(n)$, increasing in $n$. The precise growth rate of $T(n)$ is specified below; for notational simplicity, we suppress the dependence of $T$ on $n$ throughout.

The upper bound on the forecast horizon ensures diffusion has not saturated the network; the lower bound ensures a nontrivial forward-looking period. Together they isolate the policy-relevant intermediate regime.\footnote{The upper bound also ensures that diffusion wavefronts remain localized enough to analyze independently across network regions.}

ass[Forecast Horizon] For each \(n\), the forecast horizon \(T_n\) lies in a nonempty interval \([\underline{T}_n,\overline{T}_n]\) where, for the same \(q>0\) as in Assumption (ref), \[ \overline{T}_n = o\!\left(n^{\frac{1}{2q+3}}\right) \quad\text{and}\quad \underline{T}_n = \omega(1). \] Here \(a_n=o(b_n)\) means \(a_n/b_n\to 0\) as \(n\to\infty\), and \(a_n=\omega(1)\) means \(a_n\to\infty\).

We impose structural assumptions on $L_n$. First, define the activated set:

defnLet $I(i_0, T, P_n)$ denote the set of nodes activated by a diffusion process starting at seed $i_0$ and running for $T$ periods through percolation $P_n$. When clear from context, we suppress dependence on $n$ and $P_n$ for notational simplicity.

Next, we make assumptions on $L_{n}$.

ass[Base Network Structure] Fix \(q>0\). Assume: \begin{enumerate} • Polynomial volume growth. For every node \(j\) and integer \(t\ge 1\), let \(B_j(t)\) denote the ball of graph-distance radius \(t\) centered at \(j\) in \(L_n\). Then for all $t \leq C\overline T_n$ for any $0<C<\infty$: \[ a_{\min}\, t^{q+1} \le |B_j(t)| \le a_{\max}\, t^{q+1}, \] where $1<a_{\min}\leq a_{\max} < \infty$. • Supercriticality and uniform spread. The diffusion process is assumed to be supercritical on $L_n$. Furthermore, there exist constants $\alpha < 1/3, \theta \in (0,1)$ such that for any seed $j$ and all $z\in B_j((1-\alpha)t)$ and for all $t \leq \overline T_n$ \begin{align*} \mathbb{P}\bigl(|I(j, t)\cap B_z(\alpha t)| > \theta|B_z(\alpha t)|\bigr) > \varepsilon > 0. \end{align*} \end{enumerate}

Part 1 imposes polynomial expansion of order $q+1$: at $t=1$ it yields degree bounds $d_{\min} = a_{\min} - 1$ and $d_{\max} = a_{\max} -1$; for $t > 1$ it bounds reachable nodes within graph distance $t$. The condition permits substantial heterogeneity within polynomial envelopes.

Part 2 requires supercriticality ($p_n > p_c > 0$, where $p_c$ is the percolation threshold for the graph) and uniform spread: for any ball $B_z(\alpha t)$ within the $(1-\alpha)t$-neighborhood of the seed, at least fraction $\theta$ of the ball is reached with positive probability. This ensures diffusion does not systematically avoid local neighborhoods. Regular lattices and polynomial-expansion graphs satisfy this property.\footnote{We conjecture Part 2 can be derived from Part 1, supercriticality, and isoperimetric inequalities using contreras2024supercritical's renormalization techniques.} See Grimmett2002 and contreras2024supercritical for treatments of general graph classes satisfying Part 1.

remark[Role of $\alpha < 1/3$] The parameter $\alpha$ controls the resolution at which the diffusion fills space. It partitions the reachable radius $t$ into two zones: nodes within $(1-\alpha)t$ of the seed serve as candidate centers $z$, and the ball $B_z(\alpha t)$ around each must be filled with positive probability. Small $\alpha$ imposes a local filling condition---the diffusion must reach a positive fraction of moderate-sized neighborhoods throughout most of the reachable region. Large $\alpha$ instead requires filling of a single large ball from near the seed, a coarser, more global condition. The bound $\alpha < 1/3$ is used in the proof of Theorem (ref) to ensure the geometric construction admits a valid choice of parameters (see Appendix (ref)). This restriction is not binding for the graph classes of primary interest. For the random geometric graph (Example (ref)), the block-renormalization proof of Theorem (ref) establishes Part 2 for some $\alpha \in (0,1)$, but the underlying shape-theorem machinery---chemical-distance control and positive cluster density in supercritical percolation---yields uniform spread at every scale: for any $\alpha' \in (0,1)$, the conclusion holds with constants $\theta(\alpha')$ and $\varepsilon(\alpha')$ that remain positive. The same is true for integer lattices and, more broadly, any polynomial-expansion graph where supercritical percolation satisfies a shape theorem. A graph that satisfies Part 2 only for large $\alpha$ would exhibit “lumpy” diffusion---reliable spread at coarse scales but systematic gaps at finer resolution---a pathology that does not arise in spatially grounded networks.

By Assumption (ref), the radius-\(t\) ball around any node contains \(\Theta(t^{q+1})\) nodes. Consequently, at the maximal forecast horizon, \[ |B(\overline{T}_n)| = \Theta\!\left(\overline{T}_n^{q+1}\right) = o\!\left(n^{\frac{q+1}{2q+3}}\right) = o(n^{1/2}) = o(n), \] implying the diffusion explores a vanishing fraction of the network—an “intermediate” asymptotic regime. The lower bound \(\underline{T}_n=\omega(1)\) rules out trivially short horizons while permitting arbitrarily slow growth.

remark[Connection to the S-Curve] The intermediate regime corresponds to the steep, accelerating portion of the S-shaped epidemic curve, where diffusion grows as $\Theta(t^{q+1})$. Jump links supercharge expansion, creating a wedge between true and predicted S-curves that widens as $T$ grows. Once saturation dominates, forecasts reconverge. Measurement error is most damaging during the policy-critical window when interventions have the greatest potential impact.\footnote{In the simulations and empirical examples we consider in Section (ref) and Appendix (ref), the diffusion processes can be well-approximated by standard SIR curves. One can fit SIR curves to the data via moment conditions. In sample, standard SIR difference equation models fit the data well; however, forward projections using the estimated SIR model have error due to mismatch between the exponential SIR model and the underlying polynomial diffusion process.}

We now introduce notation for the expected spread of diffusion on the base network \(L_n\). Let \(\mathcal{E}_t\) denote the expected number of nodes activated by time \(t\) under percolation \(P_n\): \[ \mathcal{E}_t := \mathbb{E}_{P_n}\Bigl[\bigl|\{\, j \in V(L_n): j \text{ is activated by time } t \,\}\bigr|\Bigr]. \] We define \(\mathcal{S}_t := \mathcal{E}_t - \mathcal{E}_{t-1}\) as the expected number of new activations at integer time \(t\ge1\).

lemma[Polynomial Diffusion Growth] Suppose Assumptions (ref) and (ref) hold. Then, uniformly over \(t\in[\underline{T}_n,\overline{T}_n]\), \[ \mathcal{E}_t = \Theta\bigl(t^{q+1}\bigr). \]

All proofs appear in Appendix (ref). Supercritical percolation fills balls proportional to their volume: cumulative activations grow as $\Theta(t^{q+1})$, with the diffusion frontier expanding into the boundary of the growing ball. Uniform spread ensures diffusion does not stall in bottlenecks. This polynomial baseline is precisely what missing links disrupt: because diffusion on $L_n$ grows only polynomially, even a tiny number of jump links produces a proportionally enormous effect.

We provide two examples illustrating the scope of Assumptions (ref) and (ref), and how they can fail to hold.

exmp[Line Graph] The line graph (and ring graphs, as in the base watts1998collective case) violates both parts of Assumption (ref). Starting from seed $i_0$, volumes grow linearly, $|B_{i_0}(t)| = \Theta(t)$, which corresponds to $q=0$, which is ruled out by assumption. However, consider the set of nodes at distance $(1-2\alpha)T$ to $T$ from $i_0$. Since there exists a unique path from $i_0$ to this set, the probability the diffusion reaches it is at most $p_n^{(1-2\alpha)t}$. If $p_n < 1$, this probability vanishes as $t \to \infty$, violating the uniform spread condition. This is consistent with watts1998collective, who show that increasing “randomness” is necessary to achieve non-trivial global infection.
exmp[Complete $k$-ary Tree] Consider a complete $k$-ary tree with $k>1$ and seed $i_0$ at the root. This graph violates Assumption (ref).1. The volume of a radius-$t$ ball expands as \[ |B_{i_0}(t)| = \sum_{s=0}^t k^s = \frac{k^{t+1} - 1}{k-1} = \Theta(k^t), \] exhibiting exponential rather than polynomial growth in $t$.

Alternatively, latent-space models---where nodes link based on proximity in an underlying space hoffrh2002---satisfy our assumptions and encompass geographic contact networks, spatial interaction models, and latent-position models. We verify Assumption (ref) for the canonical case: the random geometric graph (RGG), which transparently connects latent-space geometry to expansion properties. The RGG is the most favorable case for our assumptions; if measurement error causes problems here, it will likely be worse in more complex networks.

The result extends to boundedly inhomogeneous Poisson intensities (Remark (ref) in Appendix (ref)): polynomial volume growth and uniform spread hold as long as no region is completely empty or infinitely dense. Assumption (ref) does not cover scale-free networks or networks with exponential volume growth. We consider graphs with exponential volume graph in Appendix (ref). Moreover, the data-collection procedures that produce polynomially-expanding measured networks---geographic proximity proxies, interaction thresholds, survey boundaries---are precisely those most likely to create unaligned measurement error.

exmp[Random Geometric Graph] Fix dimension $d \geq 2$ and let $G_n$ be a random geometric graph on $n$ points from a homogeneous Poisson process of intensity $\lambda$ on a $d$-dimensional torus of volume $n/\lambda$, with connection radius $r > 0$ in the supercritical regime of continuum percolation. Taking $L_n$ to be the giant component $\mathcal{C}_n$ and $q = d - 1$ so that $t^{q+1} = t^d$, both parts of Assumption (ref) hold for polylogarithmic horizons and at least a $(1-\varepsilon)$ fraction of vertices for any $\varepsilon$. The proof, a block-renormalization argument using the Liggett--Schonmann--Stacey domination theorem, appears in Appendix (ref).
remark[Typical vs.\ uniform seeds] Assumption (ref) requires polynomial volume growth and uniform spread for every node and seed. Example (ref) verifies these properties for $(1-\varepsilon)$-typical vertices in the giant component; the remaining $\varepsilon$-fraction near irregular regions (e.g., the boundary of the giant component) may violate the uniform bounds. Accordingly, the random geometric graph example should be read as validating the theory for typical seeds rather than literally for every seed. All subsequent theorems hold verbatim when Assumption (ref) is satisfied; for graph classes where the assumption holds only for typical vertices, the results apply with at most an $\varepsilon$-probability exception set that can be made arbitrarily small.

Finally, we specify the distribution of missing links $E_n$. Our baseline analysis focuses on i.i.d.\ link formation with probability $\beta_n$; Appendix (ref) extends the results to settings where each node's linking capacity is bounded. The key requirement is that the support of $E_n$ provides sufficient coverage at medium-range distances relative to $L_n$—a condition trivially satisfied under i.i.d.\ formation. Despite this expansive support, realized networks $E_n$ remain sparse with high probability.\footnote{We thank an anonymous referee whose comments on an earlier draft motivated the treatment of bounded linking capacity in Appendix (ref).}

ass[Missing Link Distribution] For all $n$ and all pairs $i,j \in V_n$, $E_{ij} \stackrel{\mathrm{iid}}{\sim} \mathrm{Ber}(\beta_n)$ with \[ \beta_n = \omega\left(\frac{1}{p_n \underline{T}_n^{q+1} n}\right),\quad \beta_n = o\left(\frac{1}{n}\right) \]

Both bounds on $\beta_n$ vanish as $n\to\infty$. Combined with Assumptions (ref) and (ref), the lower bound ensures that $p_n \underline{T}_n^{q+1} n \beta_n \to \infty$: the expected number of successful transmissions through missing links originating from the initial diffusion ball diverges, guaranteeing that the diffusion escapes its local polynomial neighborhood with probability approaching one. The upper bound $\beta_n \ll n^{-1}$ ensures sparsity: the expected degree from $E_n$ is $o(1)$, implying $E_n$ is asymptotically disconnected and contains no giant component with high probability.

To calibrate, consider a geographic contact network with $q = 2$ (balls grow as $t^3$), $n = 10^6$ nodes, and passing probability $p_n = 0.4$. The key finite-sample diagnostic is the ratio $T_n^{2q+3}/n$, which must be small for the asymptotic approximations to hold.\footnote{Asymptotic bounds involve unspecified constants, so finite-sample calibrations are illustrative. We choose parameters where $T_n^{2q+3}/n$ is well below 1.} Taking $T_n = 4$: the ratio evaluates to $4^7/10^6 = 16{,}384/10^6 \approx 0.016 \ll 1$, comfortably within the valid regime. Assumption (ref) then requires $\beta_n \gg 1/(0.4 \cdot 64 \cdot 10^6) \approx 3.9 \times 10^{-8}$ and $\beta_n \ll 10^{-6}$. Setting $\beta_n = 10^{-7}$, the expected number of missing links per node is $n\beta_n = 0.1$---roughly one missing link per ten nodes---yet our theorems show this suffices for catastrophic forecast failure. As $n$ grows the requirements become less stringent: at $n = 10^8$ the same assumptions permit $T_n$ up to $6$ (with $6^7/10^8 \approx 0.003$) and $\beta_n$ as small as $10^{-10}$.

The forecast errors we characterize arise not from dense unobserved networks or hidden giant components---settings where errors would be unsurprising---but from sparse, fragmented, idiosyncratic jump links, each connecting two otherwise distant parts of $L_n$. Substantial forecast error emerges even when mismeasurement takes its mildest form. These links are dangerous because their placement ignores the geometry of $L_n$: they connect nodes far apart in graph distance as readily as nodes nearby, functioning as escape hatches that vault diffusion past the predicted frontier.

remark[Relationship to Small-Worlds Models] The watts1998collective model is a low-dimensional analogue of our framework: $L_n$ is a ring lattice and $E_n$ is generated by uniform rewiring. The ring lattice does not satisfy Assumption (ref)---volume growth is linear ($q = 0$, excluded) and uniform spread fails (see Example 1)---so the watts1998collective model is not a literal special case of our assumptions. Nonetheless, the qualitative mechanism is the same: random shortcuts dramatically compress path lengths and accelerate diffusion. Our framework generalizes this observation to the class of polynomial-expansion networks that do satisfy Assumption (ref), and additionally characterizes when shortcuts are dangerous (geometrically unaligned with the expansion structure of $L_n$) and when they are not (when missing links are aligned with $L_n$, i.e., resemble a random thinning of the observed network)---a distinction absent from the small-worlds literature. “Long-range” here refers to graph distance on $L_n$, not necessarily physical distance---a missing link between geographically nearby nodes functions as a long-range shortcut if $L_n$ provides no short path between them.

\paragraph*{Econometrician's Goals} The econometrician has two objectives: predicting which nodes are reached by diffusion and predicting how many, both by time $T$.

Let $y_{jt}$ indicate whether node $j$ has been activated by time $t$ for a diffusion starting at seed $i_0$ under percolation $P_n(G_n)$; we suppress dependence on $G_n$, $P_n(G_n)$, and $i_0$ when clear from context. The set of activated nodes by period $T$ is \[ I_{P_n(G_n)}(i_0,T) = \{j \in V_n : y_{jT} = 1\}. \] The econometrician observes $L_n$ but not $E_n$ or $P_n$, so we study the distribution over $P_n(G_n)$ induced by both sources of randomness. We further assume the econometrician knows $T$, $q$, and $L_n$ perfectly---heroic assumptions favorable to prediction, so our impossibility results characterize a best-case scenario.

remark[Mismeasured Forecast Horizon] In practice $T$ may itself be uncertain, with effects analogous to seed mismeasurement. Under polynomial expansion, a small error $\Delta T$ yields relative forecast error of order $\Delta T / T$---potentially large when $T$ is moderate. Moreover, a longer-than-expected $T$ amplifies the impact of jump links by allowing more time for cascading secondary wavefronts. Our impossibility results, which assume $T$ is known, therefore represent a lower bound on realistic forecast errors.

Sensitive Dependence on the Seed

Small errors in identifying the initial seed generate substantial prediction errors in both the spatial extent and composition of activated nodes.

Fix a percolation realization $P := P_n(G_n)$ and consider two seeds: the true seed $i_0$ and a nearby counterfactual seed $j_0$. Let $I_P(i_0,T)$ and $I_P(j_0,T)$ denote the sets of nodes activated by period $T$ from each seed under the same transmission realization $P$.

We measure forecast sensitivity using a modified Jaccard index jaccard1901etude: \[ \Delta_n(i_0,j_0) := \frac{|I_P(i_0,T) \cap I_P(j_0,T)|}{|I_P(i_0,T) \cup I_P(j_0,T)|}. \] Values bounded away from one indicate that diffusions from nearby seeds $i_0$ and $j_0$ reach substantially different sets of nodes, despite their proximity.

theorem[Seed Sensitivity] Suppose Assumptions (ref) and (ref) hold. Fix a graph $L_n$. Then for every seed $i_0$, there exists a forecast horizon $T_n \in [\underline{T}_n, \overline{T}_n]$ and constants $C, c' \in (0,1)$ such that: \begin{enumerate} • There exists a nontrivial set of alternative seeds near $i_0$: let $U_{n,i_0} := B_{i_0}(a_n)$ for some $a_n = \Theta(T_n)$, and let $J_{n,i_0} \subseteq U_{n,i_0}$ satisfy $|J_{n,i_0}|/|U_{n,i_0}| \geq C$. • For all $j_0 \in J_{n,i_0}$, with probability bounded away from zero over percolation realizations $P_n(L_n)$: \[ \Delta_n(i_0,j_0) \leq c'. \] \end{enumerate}

All proofs appear in Appendix (ref) unless otherwise noted.

figure[figure omitted — 3,107 chars of source]

This sensitivity is intrinsic to polynomial-expansion networks---it does not require shortcuts or missing links. The intuition is spatial: on a polynomially expanding network, seeds $i_0$ and $j_0$ each command a “cone” of influence that expands outward. Because $j_0$ is displaced from $i_0$, part of $j_0$'s cone extends into regions that $i_0$'s diffusion cannot reach within $T$ steps (see Figure (ref), panel (d)). The displaced seed does not merely add noise to the activation set---it redirects a substantial portion of the diffusion into an entirely different part of the network. This divergence occurs for a constant fraction of alternative seeds within a local neighborhood---not just for adversarially chosen ones.

Network mismeasurement amplifies this sensitivity. Without missing links, the two diverging wavefronts from $i_0$ and $j_0$ remain within their respective local neighborhoods on $L_n$---they diverge, but only locally. Jump links change this picture. As each wavefront expands, it encounters unobserved connections to other parts of the network. When a jump link fires, it seeds a new wavefront in a remote region, outside the econometrician's predicted set. Because the two original wavefronts already diverged locally (Theorem (ref)), their respective jump links reach different remote regions, amplifying a local discrepancy into a global one. To formalize this notion of disjoint regions, we divide the graph into disjoint tiles (the construction of which is detailed in Lemma (ref)). Collectively, the tiles cover a constant fraction of the graph without overlap.

cor[Global Seed Sensitivity] Suppose Assumptions (ref), (ref), and (ref) hold. Consider a local perturbation from seed $i_0$ to alternative seed $j_0$ as in Theorem (ref). Then with strictly positive probability, the following events occur jointly: \begin{enumerate} • Shortcut activation. Both diffusions utilize missing links in $E_n$: there exist nodes $e_1, e_2 \in V_n$ such that the diffusion from $i_0$ reaches $e_1$ via $L_n$ and then transmits through a missing link to some node outside $B_{i_0}(T_n)$, and similarly the diffusion from $j_0$ reaches $e_2$ and transmits outside $B_{j_0}(T_n)$. • Macroscopic divergence. Let $s := \max\{d_{L_n}(i_0, e_1), d_{L_n}(j_0, e_2)\}$ denote the maximum graph distance to the shortcut nodes. Conditional on the shortcut activations in part (1), the expected cumulative number of disjoint tiles reached by the two diffusions over the remaining $T_n - s - 1$ steps is at least \[ 2C\, n\beta_n p_n (T_n - s - 1)^{q+1}, \] where $C > 0$ is the constant from Lemma (ref). \end{enumerate}

This is stated under the i.i.d.\ case (Assumption (ref)); the appendix establishes the same result under the more general bounded-capacity assumption (Corollary (ref)).

The interaction between seed sensitivity and jump links is multiplicative. Since seeds $i_0$ and $j_0$ influence disjoint local regions (Theorem (ref)), they encounter different escape hatches into $E_n$, jump to different distant parts of the network, and then diverge further as each secondary wavefront generates its own independent jumps. A local seed perturbation cascades into forecast errors spanning the entire network.

Sensitivity to Network Mismeasurement

Forecasts based on the observed network $L_n$ systematically underestimate diffusion on the true network $G_n$, even when the measurement error is vanishingly small.

Consider an econometrician who observes $i_0$ and $L_n$ perfectly but assumes $E_n \equiv \emptyset$. The forecast for expected activations is \[ \hat{Y}_T(L_n) := \mathbb{E}_{P_n(L_n)}\left[\sum_{j=1}^n y_{jT} \,\bigg\vert\, L_n, i_0\right], \] where the expectation integrates over diffusion realizations on the base network $L_n$ alone.

This estimator reflects standard practice: survey constraints truncate reported connections, mobility data impose distance or frequency thresholds, and social media studies omit offline interactions. Each effectively assumes $E_n \equiv \emptyset$, making $\hat{Y}_T(L_n)$ the natural benchmark.

Accounting for $E_n$ is infeasible even when researchers suspect its presence. Assumption (ref) requires $\beta_n = o(n^{-1})$, so the expected number of missing links is $o(n)$---as we show in Proposition (ref), there is so little error that it is hard to estimate. The misspecified forecast $\hat{Y}_T(L_n)$ is therefore not merely a stylized benchmark but the practical reality.\footnote{Note that computing the probability of activation for any given node is NP-complete shapiro2012finding.}

As a benchmark, consider an oracle econometrician who knows the distribution of $E_n$ and computes \[ \hat{Y}_T(G_n) := \mathbb{E}_{E_n, P_n(G_n)}\left[\sum_{j=1}^n y_{jT} \,\bigg\vert\, L_n, i_0\right], \] integrating over both missing link realizations and diffusion outcomes. Comparing $\hat{Y}_T(L_n)$ against this (computationally intractable) oracle isolates the effect of ignoring sparse measurement error. Despite knowing $L_n$, $i_0$, $T$, and $q$ perfectly, the misspecified forecast exhibits catastrophic error:

theorem[Asymptotic Forecast Failure] Suppose Assumptions (ref), (ref), and (ref) hold. Then as $n\to\infty$, \[ \frac{\hat{Y}_T(L_n)}{\hat{Y}_T(G_n)} \to 0. \]

On $L_n$, diffusion expands polynomially, never escaping its local neighborhood. Each missing link in $E_n$ is an escape hatch that sends diffusion to a new region, where it spreads polynomially. As the wavefront grows, it encounters more escape hatches, accelerating the cascade. Even though any given node has vanishing probability of generating an escape, the cumulative effect overwhelms the locally-predicted spread. The small-worlds phenomenon of watts1998collective---demonstrated numerically on ring lattices---is one instance of this cascade mechanism. Our framework identifies the general principle: forecast failure occurs whenever missing links are geometrically unaligned with the observed network's expansion structure.

The wavefront grows by $\Theta(t^{q+1})$ nodes at each step, each independently probing for jump links. A jump link firing at time $s$ seeds a new wavefront contributing $\Theta((T-s)^{q+1})$ additional activations. The econometrician's forecast captures only $\Theta(T^{q+1})$ from the original source, while the oracle aggregates the original plus all secondary outbreaks. The ratio diverges because cumulative satellite contributions grow faster than the single-source prediction. The proof formalizes this cascade by partitioning $L_n$ into regions and tracking how jump links seed separate wavefronts (Appendix (ref)).

When Does Mismeasurement Matter?

Not all mismeasurement is so problematic. The geometry of $E_n$---whether missing links preserve or disrupt the expansion structure of $L_n$---is critical in determining the impact of missing links. When measurement error is aligned---missing links resemble a random thinning of the observed network---forecast failure vanishes and seed sensitivity is not amplified.

ass[Aligned Edge Censoring] For each $n$, let the true network be $G_n$, where $G_n$ satisfies Assumption (ref) and the forecast horizon $T_n$ satisfies Assumption (ref). The econometrician observes a thinned network $L_n$ obtained by independent edge censoring: for each undirected edge $e\in G_n$, include $e$ in $L_n$ with probability $1-\varepsilon_n$ and delete it with probability $\varepsilon_n$, independently across edges. Define the unobserved edges as $E_n := G_n \setminus L_n$, so the true diffusion network is $G_n = L_n \cup E_n$ augmented by the missing edges $E_n$. Assume the censoring rate satisfies \[ \varepsilon_n T_n^{q+1} \;\longrightarrow\; 0 \quad\text{as } n\to\infty. \]

Assumption (ref) combines two pieces. The first component assumes correct model specification -- the researcher only misses links within the context of the correct network geometry. While there may be some long range links missed, they are long range links within the correctly specified structure. The second component ensures that there cannot be too much censoring -- this ensures we do not have dramatic changes in the geometry due to censoring.

theorem[No Asymptotic Forecast Failure under Aligned Error] Suppose Assumptions (ref), (ref), and (ref) hold. Then for any fixed seed $i_0$, \[ \mathbb{P}\bigl( I(i_0,T_n,G_n) = I(i_0,T_n, L_n)\mid P_n \bigr) \;\longrightarrow\; 1. \] In particular, the econometrician's forecast based on $L_n$ exhibits no asymptotic forecast failure: \[ \frac{\bigl| I(i_0,T_n, G_n) \triangle I(i_0,T_n, L_n) \bigr|} {\bigl| I(i_0,T_n, G_n) \bigr|} \;\xrightarrow{\mathbb P}\; 0, \quad \frac{\mathbb{E}\bigl[|I(i_0,T_n, L_n)|\bigr]} {\mathbb{E}\bigl[|I(i_0,T_n, G_n)|\bigr]} \;\longrightarrow\; 1. \] Where we consider the same percolation $P_n$ for the symmetric difference. The expectation ratio is a deterministic consequence.

The contrast with Theorem (ref) is striking: forecast failure depends on the geometry of missing links, not their quantity. Under aligned error, missing links are not escape hatches---they are redundant copies of existing paths. The true diffusion uses $O(T_n^{q+1})$ edges, each independently missing with probability $\varepsilon_n$, so the expected number of “missing but needed” edges is $O(\varepsilon_n T_n^{q+1}) \to 0$. With high probability, the two diffusions produce identical activation sets. Under Assumption (ref), by contrast, missing links bridge distant regions. The practical question is whether a data-collection process disproportionately misses links resembling observed links (aligned---safe) or links bridging different network regions (unaligned---problematic). Since activation sets coincide with high probability, any functional is preserved. In particular, aligned error cannot amplify seed sensitivity beyond what is intrinsic to the true network.

remark[Scope of the aligned/unaligned contrast] Theorems (ref) and (ref) provide sufficient conditions for forecast failure and robustness, respectively, under two distinct error-generating processes. They are not exhaustive: intermediate cases---such as missing links that are partially correlated with $L_n$, or error processes that combine aligned and unaligned components---are not covered by either theorem and remain an open question. The key qualitative insight is that the geometric relationship between missing and observed links, rather than the volume of missing links alone, governs forecast reliability.

The condition $\varepsilon_n T_n^{q+1} \to 0$ deserves interpretation. For $q = 2$ and $T_n = 10$, this requires $\varepsilon_n \ll 1/1000$. This is sufficient but not necessary: the proofs use it to ensure the diffusion path on $G_n^\star$ lies entirely in $L_n$ with high probability. In many survey settings, random nonresponse rates satisfy the condition---a $5\%$ link censoring rate works when $T_n^{q+1} < 20$. The condition becomes restrictive only when the forecast horizon is long relative to the censoring rate.

Estimation and Potential Solutions

Forecast failure occurs under unaligned error (Theorem (ref)) but not under aligned error (Theorem (ref)). Can standard econometric approaches help in the unaligned case? We consider three strategies---estimating aggregate parameters ($p_n$, ${\mathcal R}_0$), estimating the missing link probability $\beta_n$, and widespread testing---and show that each is insufficient.

Consistent Estimation of Aggregate Parameters

Despite the forecast failures above, certain aggregate parameters remain consistently estimable---though this provides no protection against the spatial and volumetric errors documented in Theorems (ref) and (ref).

Assume the econometrician observes all activations perfectly.\footnote{We do not characterize the optimal estimator, as computing maximum likelihood estimates for network diffusion is NP-complete shapiro2012finding. Instead, we demonstrate consistency using a simple moment-based approach.} Given $L_n$ and the activation history $\{y_{j,s}\}_{j \in V_n, s \leq t-1}$, a consistent estimator $\hat{p}_n$ can be constructed. Consistency of $\hat{p}_n$ immediately yields consistent estimation of ${\mathcal R}_0$, the expected number of secondary activations from a single seed in a susceptible population.\footnote{In a fully mixed population, ${\mathcal R}_0 = p_n \bar{d}$ where $\bar{d}$ is the average degree.} The natural estimator is $\hat{{\mathcal R}}_0 = \hat{p}_n d_L$, where $d_L$ is the mean degree in $L_n$. The true reproduction number ${\mathcal R}_0(G_n) = p_n(d_L + \beta_n n)$ exceeds this by the contribution of missing links.

remarkSuppose Assumptions (ref), (ref), and (ref) hold, and that ${\mathcal R}_0(L_n) = p_n d_L$ remains bounded as $n \to \infty$. We construct a consistent estimator $\hat{p}_n \to_p p_n$ below (Eq. (ref)). Given $\hat{p}_n$ and the observed mean degree $d_L$, the natural estimator $\hat{{\mathcal R}}_0 = \hat{p}_n d_L$ satisfies \begin{equation*} \frac{\hat{{\mathcal R}}_0}{{\mathcal R}_0(G_n)} \to_p 1. \end{equation*}
proofDecompose the true reproduction number as \begin{equation*} {\mathcal R}_0(G_n) = p_n d_L + p_n \beta_n n = p_n d_L \left(1 + \frac{\beta_n n}{d_L}\right). \end{equation*} By Assumption (ref), $\beta_n = o(1/n)$, which implies $\beta_n n = o(1)$ and hence $\beta_n n/d_L = o(1)$ under the maintained assumption that $d_L$ is bounded away from zero and infinity. Therefore, \begin{equation*} {\mathcal R}_0(G_n) = {\mathcal R}_0(L_n) \cdot (1 + o(1)). \end{equation*} Since $\hat{p}_n \to_p p_n$ by assumption and $d_L$ is observed, we have $\hat{{\mathcal R}}_0 \to_p {\mathcal R}_0(L_n)$ by the continuous mapping theorem. The result follows by combining these convergences.

${\mathcal R}_0$ governs whether diffusion spreads or dies out, but conditional on spread, it reveals nothing about where or how far diffusion reaches. The reason is simple: ${\mathcal R}_0$ measures the average local transmission rate, which is virtually identical on $L_n$ and $G_n$ since missing links contribute $o(1)$ to average degree. The forecast errors arise from the spatial reach of jump links, which redistribute diffusion without materially changing ${\mathcal R}_0$.

A simple consistent estimator for $p_n$ illustrates the construction. Let ${\mathcal I}(i,t) \subseteq V_n$ denote the neighbors of node $i$ in $L_n$ activated at time $t$. Consider

equation[equation omitted — 235 chars of source]

The estimator restricts attention to susceptible nodes with exactly one activated $L_n$-neighbor, which---absent hidden exposures---isolates independent Bernoulli trials with success probability $p_n$. However, the econometrician does not observe $E_n$: a node classified as having exactly one activated $L_n$-neighbor may have additional activated neighbors through unobserved missing links, creating hidden exposures that could bias the estimator. We now show this contamination is asymptotically negligible.

At any time $t \leq \overline{T}_n$, the number of activated nodes is at most ${\mathcal E}_t = O(t^{q+1})$ by Lemma (ref). For a susceptible node $i$, the probability that $i$ has any activated $E_n$-neighbor is bounded by the union bound: \[ {\mathbb P}(i \text{ has an activated } E_n\text{-neighbor at time } t) \;\leq\; {\mathcal E}_t \cdot \beta_n \;=\; O\!\left(\frac{t^{q+1}}{n}\right) \;=\; o(1), \] where we used $\beta_n = o(1/n)$ from Assumption (ref). Let $N_t^{\text{clean}}$ denote the number of susceptible nodes at time $t$ with exactly one activated $L_n$-neighbor and no activated $E_n$-neighbor (“clean” observations), and let $N_t^{\text{contam}}$ denote those with exactly one activated $L_n$-neighbor and at least one activated $E_n$-neighbor (“contaminated” observations). By the bound above and Markov's inequality, ${\mathbb E}[N_t^{\text{contam}}] / {\mathbb E}[N_t^{\text{clean}}] = o(1)$: the contaminated fraction vanishes. Each clean observation is an independent $\mathrm{Ber}(p_n)$ trial, so $\hat{p}_n \to_p p_n$ by the law of large numbers applied to the clean subsample, which dominates the denominator. This estimator exploits knowledge of $L_n$ through ${\mathcal I}(i,t)$ and is inefficient, discarding nodes with multiple activated neighbors.\footnote{The contamination bound $O(t^{q+1}/n) = o(1)$ holds uniformly over $t \leq \overline{T}_n$ since $\overline{T}_n^{q+1} = o(n^{(q+1)/(2q+3)}) = o(n^{1/2}) = o(n)$ by Assumption (ref).}

Potential Remedies and Their Limitations

A policymaker might try to (i) estimate $\beta_n$ via supplementary network surveys or (ii) detect activated regions through widespread testing. Each strategy fails.

Estimating $\beta_n$ via Network Surveys

Suppose the econometrician samples $m_n$ nodes uniformly at random and checks whether each of the $\binom{m_n}{2}$ pairs has a link in $G_n$. Since $\beta_n$ is vanishingly small, this is a best-case scenario (i.i.d.\ links are easiest to find). Two regimes emerge: with a slowly growing sample, no missing links are found; with a larger sample, some may be found but consistent estimation remains impossible.

propUnder Assumption (ref), if: \begin{enumerate} • $m_n=o(\sqrt{n})$, $ {\mathbb P}\left(\text{{\rm No links amongst }}\binom{m_n}{2} \text{ {\rm found}}\right) \rightarrow 1. $$m_n= O(1/\sqrt{\beta_n})$, there exists $\epsilon > 0$ and $c\in(0,1)$ such that ${\mathbb P}(|\hat \beta_n / \beta_n - 1| <\epsilon ) < c$. \end{enumerate}

For a population of $n = 10^6$: with $\beta_n = 1/(n(\log n)^2)$ (admissible under our assumptions with $T_n = \log n$, $q = 2$), Part (1) implies that a survey of $m_n = o(\sqrt{n}) = o(1{,}000)$ nodes yields no missing links with probability approaching one. For larger surveys up to $m_n = O(1/\sqrt{\beta_n}) \approx O(13{,}816)$ nodes, Part (2) shows that consistent estimation of $\beta_n$ remains impossible. With $\beta_n = \log(n)/(p_n n^{(2q+1)/(q+1)})$ and $q = 4$, Part (2) implies that even surveying 68{,}000 individuals---6.8% of the population---cannot produce a consistent estimator. Directly estimating $\beta_n$ through supplementary data collection is infeasible in most settings.

Widespread Testing

Can widespread testing identify which regions contain activated nodes? Using the tiling construction from Lemma (ref), we model regions as disjoint tiles and consider an idealized regime: instantaneous random tests across all $n$ nodes, each activated node detected independently with probability $\gamma_n$. This abstracts from resource constraints, compliance, and delays, providing an upper bound on feasible protocols.

theoremSuppose Assumptions (ref), (ref), and (ref) hold. Consider a time period $T$ and testing regime with detection probability $\gamma_n \to 0$ satisfying $\gamma_n = O(T^{-(q+1)})$ with implied constant satisfying $\gamma_n a_{\max} T^{q+1} < 1$. Let $K_T^\star$ denote the expected number of tiles containing at least one activated node at time $T$, and let $\hat{K}_T$ denote the expected number of tiles in which at least one activation is detected under the testing regime. If each activated individual is observed independently with probability $\gamma_n$, then as $n \to \infty$, \begin{equation*} \frac{\hat{K}_T}{K_T^\star} \leq \gamma_n \, a_{\max} \, T^{q+1} < 1. \end{equation*}

The reason is tied to the escape-hatch mechanism: jump links seed distant regions that initially contain few activated nodes. Detection requires observing at least one activation per region, but recently-seeded regions have had little time for local spread. Since jump links continue seeding new regions throughout the forecast horizon, a substantial fraction are always in this “recently seeded, hard to detect” state.

Since universal, instantaneous testing is already an upper bound on feasible protocols, and real-world testing faces resource constraints, incomplete compliance, and reporting delays, spatial detection provides limited value for medium-run forecasting. Testing is most needed precisely in the unaligned setting where shortcuts create unexpected distant activations; under aligned error, forecasts based on $L_n$ are already correct.

Empirical Applications

The following applications are designed to illustrate the mechanisms identified by our theory---sensitivity to error geometry, the relative harmlessness of aligned error, and the fragility of structured networks---in empirically relevant settings. They do not constitute formal tests of the theorems, which require asymptotic regimes ($n \to \infty$, intermediate horizons) not available in finite data. We note below where specific design choices depart from the formal assumptions.

Appendix (ref) confirms the finite-sample relevance of our asymptotic results through Monte Carlo simulations on synthetic lattice networks: with $n = 4{,}000$ and $q = 2$, the policymaker underestimates diffusion by 83% at the worst point, and nearby seeds produce nearly disjoint activation sets (${\mathcal J} = 0.055$ for $q = 4$).

Here we turn to real-world data. Three applications illustrate these results, moving from controlled simulation to policy failure: a cell-phone mobility exercise showing that how links are missed matters more than how many; New York City's micro-cluster zoning, where geography-based policy used the wrong network; and the optimal targeting methodology of beaman2021can, where measurement error is most damaging precisely where targeting is most valuable.

Network Truncation on Real Mobility Topology

First, we consider the spread of COVID-19 in the American West and Southwest. Using SafeGraph cell-phone mobility flows kang2020multiscale across California, Nevada, and part of Arizona (${\sim}11{,}000$ Census tracts), we construct $L_n$ by pruning flows below the 92nd percentile of tract to tract travel. We simulate SIR diffusion (${\mathcal R}_0 = 2.5$) on both and measure sensitive dependence and forecast underestimation (details in Appendix (ref)). The first exercise compares two error structures with the same number of missing links. The pruned $E_n = G_n \setminus L_n$ consists of geographically structured moderate-flow connections between nearby tracts. We construct $G_n$ by pruning travel flows at the 91st percentile. For the i.i.d. case, we generate $E_n$ as an Erd\H{o}s--R\'{e}nyi graph calibrated to match the edge count of the pruned $E_n$.

We begin by analyzing sensitive dependence. We compare diffusion from $i_0$ and a given $j_0$. We select $j_0$ uniformly at random from a neighborhood of radius two around $i_0$ in $L_n$, which makes up 2.5% of the graph. We then compute the expected Jaccard index of diffusion from each seed. Figure (ref) shows the sensitive dependence of the diffusion on the seed. In all cases, the average Jaccard index is well below 1. We consider the time halfway to the diameter of $L_n$ as a benchmark: for $L_n$, the average Jaccard index is 0.47 at time step ten. For $G_n$ generated via pruning, the average Jaccard index is 0.62 -- the local links cause additional overlap (corresponding to how the local linking dampens the global spread in Corollary (ref)). For $G_n$ generated via i.i.d. links, the average Jaccard index is 0.73. However, unlike in the pruned case, the Jaccard index for $L_n$ is higher than for the i.i.d. $G_n$, up to the diameter of $G_n$ -- this is consistent with Corollary (ref), where missing links increase separation.

Figure (ref) shows $\hat{Y}_T(L_n)/\hat{Y}_T(G_n)$ under both error structures (Table (ref)). Under pruning, the minimum ratio is ${\sim}0.45$ (55% underestimate). Under i.i.d.\ errors with the same edge count, it drops to ${\sim}0.24$ (76% underestimate). The i.i.d.\ errors create shortcuts between distant tracts; pruned errors add redundant connections between nearby ones. Both add the same volume of edges, but i.i.d.\ errors compress average path length to 4.03 vs.\ 5.87 for pruned (and 7.25 for $L_n$).

A natural complement is an aligned random-deletion simulation: randomly deleting edges from $G_n$ to form $L_n$. Theorem (ref) predicts negligible forecast errors, completing the comparison: i.i.d.\ (catastrophic) $\gg$ pruned (substantial) $>$ aligned (less harm). The aligned case constructs $L_n$ from $G_n$ dropping links i.i.d. at random---the missing links are a uniform thinning of the true network, preserving its geometric structure. We calibrate the probability of dropping links so that the implied $E_n$ (the gap between $L_n$ and $G_n$) has the same number of links as in the pruned and i.i.d. cases.\footnote{In a second case, we specifically set $\varepsilon_{n} = \frac{1}{\text{diam}(G_n)^3}$, to match Theorem (ref). Here, the gap between the Jaccard indices at time step 10 is 0.017, at 0.47 and 0.45, and the minimal ratio of total infections is $\sim0.85$, a much higher minimum.} Figure (ref) shows the results. Despite the same volume of links being missed, aligned errors are less damaging than the pruned case. For sensitive dependence, the gap between $L_n$ and $G_n$ is smaller at the tenth step (0.47 vs 0.33, compared to 0.47 vs 0.62)\footnote{Note that $G_n$ in the aligned case is $L_n$ in pruned and i.i.d. cases.}. Furthermore, the minimum ratio of total infections is $\sim 0.52$ ($48\%$ underestimate), a smaller underestimate.

figure[figure omitted — 931 chars of source]
figure[figure omitted — 1,044 chars of source]
figure[figure omitted — 671 chars of source]

NYC Micro-Cluster Initiative: When Policy Used the Wrong Network

In October 2020, New York State implemented a “micro-cluster” strategy to contain COVID-19 outbreaks cuomo2020microcluster, assigning each Modified ZIP Code Tabulation Area (MODZCTA) to a tier---Red, Orange, Yellow, or None---based on local case rates and those of geographically adjacent areas. Red and Orange zones faced school and business closures; Yellow zones received moderate restrictions. The program covered 177 MODZCTAs over 24 weeks (October 2020--March 2021), producing 4{,}175 MODZCTA-weeks (Figure (ref); Appendix (ref) provides the full timeline).

The implicit network was geographic adjacency---a planar-like graph with polynomial expansion (${\sim}t^2$). Our theory predicts this could be misleading: if actual transmission follows a mobility network, the absent mobility links function as jump links, connecting areas far apart in adjacency-graph distance.

We construct two alternative mobility networks from GeoDS cell-phone mobility data kang2020multiscale, which records weekly population-scaled visitor flows between Census tracts, aggregated to the MODZCTA level.\footnote{The geographic adjacency network $L$ connects every pair of MODZCTAs that share a physical boundary (${\sim}420$ edges, ${\sim}4.7$ neighbors per node, row-normalized). The mobility flow structure is stable across weeks (average pairwise Spearman $\rho = 0.93$; see Appendix (ref)).} The degree-matched mobility network $G_{\text{match}}$ retains, for each MODZCTA $i$, the top-$d_i$ mobility connections by flow weight, where $d_i$ is $i$'s geographic degree in $L$. This produces a network with the same per-node out-degree as $L$ but different edges (${\sim}70\%$ overlap), so that any difference in predictive power reflects which connections matter rather than how many. The geo-superset mobility network $G_{\text{super}}$ keeps every geographic neighbor and adds mobility connections above a per-node threshold, so $L \subset G_{\text{super}}$ by construction. This second comparison asks whether mobility information helps even when geography is fully retained.

For each network $X \in \{L, G_{\text{match}}, G_{\text{super}}\}$, we define exposure as the row-normalized weighted-average neighbor case rate:

equation[equation omitted — 109 chars of source]

where $X_{ij}$ are the row-normalized adjacency weights and $\text{CaseRate}_{j,w}$ is the weekly case rate per 100{,}000 population. Exposures are $z$-scored within week.

\paragraph*{Where the networks disagree.} Figure (ref) maps the difference between mobility-based and geographic exposure (week of October 24, 2020). Red areas face more risk through mobility than geography suggests (undertreated); blue areas face more from geography (potentially overtreated). Disagreement is greatest at the outbreak periphery---precisely where getting the network right matters most. The rank correlation between exposures averages 0.74 across weeks (Appendix Figures (ref) and (ref) provide additional detail).

We now ask which network actually predicts subsequent case growth. For each horizon $h \in \{1, 2, 3, 4\}$ weeks, we estimate:

equation[equation omitted — 227 chars of source]

where $\Delta_{h,i,w}$ is the $h$-week-ahead change in case rate, $\bm{Z}_{i,w} = (\mathbf{1}_{\text{Red}}, \mathbf{1}_{\text{Orange}}, \mathbf{1}_{\text{Yellow}})$ are zone dummies, $\alpha_i$ and $\delta_w$ are MODZCTA and week fixed effects, and standard errors are clustered at the MODZCTA level. Coefficients on exposure are standardized: cases per 100k per one-standard-deviation increase in exposure. The zone dummies absorb the treatment effect of the policy itself---so the exposure coefficients measure predictive power after controlling for the policy response.

Table (ref) reports the degree-matched horse-race. Individually, both networks predict case growth (Specs 1--2), but in the horse-race (Spec 3), $\hat\beta_G = 19.2$--$20.0$ (all $p < 0.01$) while $\hat\beta_L$ is statistically indistinguishable from zero at every horizon. Since the two networks have identical per-node degree, the gap reflects which connections matter, not how many. The geo-superset comparison tells the same story: even though $G_{\text{super}}$ includes every geographic edge, $\hat\beta_{G_s} = 20.6$--$23.8$ (all $p < 0.01$) while $\hat\beta_L \approx 0$ (Appendix Table (ref)). The predictive signal is entirely in the mobility edges.

Zone dummies confirm that Red and Orange restrictions reduced case growth at 2--4 week horizons (Appendix (ref)). MODZCTA and week fixed effects control for neighborhood characteristics and city-wide trends, though we cannot rule out all confounders and interpret the regression as predictive rather than causal.

table[table omitted — 1,257 chars of source]

We replicate the actual zoning rule---seeds designated Red, then buffer zones expanded via 1-hop (Orange) and 2-hop (Yellow) neighbors---but vary the buffer network, holding seeds fixed each week. Figure (ref) maps the results for a representative week. Replacing geography with the degree-matched mobility network ($G_{\text{match}}$) reaches only 187 buffer neighborhoods versus $L$'s 242 over the full program---a 23% reduction---with 79 zone-assignment changes. The mobility buffers reach non-contiguous areas across the city, including parts of Brooklyn and Queens that are geographically distant from the epicenter but strongly connected through commuting flows. Augmenting geography with additional mobility links ($G_{\text{super}}$) covers all 242 neighborhoods in $L$'s buffer plus 26 additional neighborhoods reached by mobility connections that do not share a border with the hotspot, for a total of 268. The Jaccard similarity between $L$ and $G_{\text{match}}$ buffers averages 0.71, indicating roughly 30% divergence in which neighborhoods receive containment resources (Appendix (ref) reports the full week-by-week comparison). The practical recommendation is augmentation: keep every geographic neighbor and add mobility connections on top.

The degree-matched comparison maps cleanly to our framework: $G_{\text{match}}$ has the same edges per node as $L$, so $G_{\text{match}} \setminus L$ consists precisely of unaligned jump links connecting MODZCTAs far apart in adjacency distance. The Yellow buffer zones are analogous to diffusion balls---regions the policymaker believes bound the outbreak. The regression evidence matches the theory: mobility exposure (capturing jump links) predicts case growth; geographic exposure does not. Two caveats: the theory targets asymptotic regimes while this is a finite setting with 177 MODZCTAs, and the ${\sim}30\%$ edge difference is not vanishingly small in the sense of $\beta_n \to 0$. Nonetheless, the theoretical mechanism is visible in the data.

figure[figure omitted — 463 chars of source]
figure[figure omitted — 512 chars of source]
figure[figure omitted — 525 chars of source]

Optimal Targeting under Measurement Error

beaman2021can design a targeting strategy for information diffusion in Malawi that identifies optimal seed nodes to maximize technology adoption. Their strategy significantly outperforming alternatives like selecting village leaders. However, their strategy depends on the network being correctly measured -- here, we examine how robust the targeting strategy is to small measurement error in the network. Our theory predicts that the answer depends on expansion properties: in well-mixed networks, seed choice matters little; in structured networks, seed choice matters enormously but is most fragile.

We take the 225 village networks from the Malawi data used by beaman2021can and calibrate an SIR percolation model to match their threshold diffusion model.\footnote{ The beaman2021can model uses a threshold rule: each agent draws a threshold $\tau$ from a truncated normal $\text{TruncN}(\lambda, 0.5; \tau > 0)$ and adopts when the number of informed neighbors exceeds $\tau$. We calibrate the per-edge infection probability in our SIR model by computing $p(\lambda) = {\mathbb P}(\tau \leq 1 \mid \tau \sim \text{TruncN}(\lambda, 0.5; \tau > 0))$, yielding $p \approx 0.46$ for $\lambda = 1$ (“simple” contagion).} For each village, we enumerate all pairs of seed nodes and run 2{,}000 SIR simulations per pair over $T = 4$ time steps to identify the optimal seed pair on the true network. We then introduce measurement error by adding false edges with probability $p_{\text{false}} = 1/n_{\text{households}}$ per non-edge (a vanishingly small rate for typical village sizes of 50--200 households), generate 1{,}000 perturbed networks per village, and re-evaluate the same seed pairs on each perturbed network.

To index the expansion properties of each village network, we compute the spectral gap---the difference between the first and second eigenvalues of the adjacency matrix. A large spectral gap indicates a well-connected, expander-like network with fast mixing; a small spectral gap indicates a more structured network with bottlenecks and slower expansion, closer to the polynomial-growth regime of our theoretical framework.

To measure sensitivity of the network to optimal seeds, we consider the change in diffusion from targeting the optimal seed (Top1 Seed) and the fifth best seed (Top5 Seed). In networks where the optimal choice matters quite a bit, the percentage difference will be large. In places where it is unimportant, this difference will be small.

\paragraph*{Results.} Figure (ref) presents five outcomes across the 225 villages, each plotted against the village's spectral gap. The results reveal a tension at the core of network-based policy. Panel (a) shows that the steepness of the objective function ($100 \times (\text{Top1 Seed} - \text{Top5 Seed})/\text{Top1 Seed}$) is 10--19% for low-spectral-gap villages and near zero for high-spectral-gap villages. Targeting matters most in bottlenecked networks---precisely where our theory predicts measurement error is most damaging.

Panels (b)--(c) confirm this. Diffusion robustness (the absolute shift in the diffusion rate, panel b) is 0.3--0.6 for low-spectral-gap villages but ${\sim}0.03$ for high-spectral-gap villages. Jaccard similarity (panel c) drops to 0.4--0.6 for low spectral gap villages---measurement error changes not just how much diffusion occurs but who is reached. The standard deviation across error draws (panel d) is large for structured villages and near zero for well-mixed ones. Panel (e) shows well-mixed villages achieve 85--90% adoption from any seed, while structured villages reach only 60%.

figure[figure omitted — 1,455 chars of source]

Low-spectral-gap villages have bottlenecked networks where the optimal seed sits at a critical bridge; a single false edge can redirect diffusion entirely. High-spectral-gap villages are near-expanders where diffusion reaches most nodes from any seed. A policymaker would invest in targeting precisely for the villages where it appears most valuable---the structured, low-spectral-gap villages. But these are exactly where targeting is most fragile: the networks with the highest potential payoff from targeting are those where measurement error most changes that payoff.

Our exercise uses i.i.d.\ false edges---the dispersed, unaligned structure our theory identifies as most damaging. More broadly, these simulations demonstrate that unaligned measurement error is particularly damaging for targeting applications, where seed choice interacts with network structure. Real survey error is plausibly unaligned and may be worse. First, non-responding households tend to differ in degree and centrality chandrasekharl2010, and may disproportionately bridge subgroups. Second, respondents are more likely to name strong, salient ties than weak ones---but weak ties bridge clusters granovetter1973. Third, surveys typically enumerate links within a defined boundary, excluding connections to neighboring villages and visiting traders.

All three channels systematically miss bridging links while retaining local ones. Our i.i.d.\ exercise may therefore understate targeting fragility, since it adds shortcuts uniformly rather than targeting bottleneck points. Some real error may be closer to aligned (random omission within well-surveyed categories), which Theorem (ref) predicts would be harmless. Researchers deploying network-based targeting should invest in understanding which error channel dominates.

Discussion

Our results establish that forecast failure in network diffusion is governed by the geometric relationship between missing and observed links. When the observed network has polynomial expansion, even vanishingly small geometrically unaligned missing links transform expansion properties, generating catastrophic forecast errors. The key mechanism is that missing links function as jump links---even those connecting geographically or socially proximate nodes---serving as escape hatches in graph distance on the observed network $L_n$, bridging regions that are macroscopically separated in the metric that governs diffusion dynamics. The small-worlds phenomenon of watts1998collective---demonstrated numerically on ring lattices---is a manifestation of this principle.

Theorem (ref) provides sufficient conditions for robustness that complement the failure conditions in Theorem (ref): forecast failure arises when missing links disrupt the expansion structure of the observed network, but not when they merely thin it. The contrast between aligned and unaligned measurement error has direct practical implications, though intermediate cases remain open (Remark (ref)). Common data-collection practices---geographic proximity proxies, interaction frequency thresholds, survey boundaries---systematically produce unaligned measurement error, because they disproportionately miss links that bridge different network regions. The most valuable investment for applied researchers may therefore be understanding the geometry of measurement error in their data, rather than estimating the rate of missing links. That said, geometry is not the only consideration: even under aligned error, Theorem (ref) requires the censoring rate to satisfy $\varepsilon_n T_n^{q+1} \to 0$, so a non-trivial volume of missing links remains problematic regardless of geometry.

Importantly, “long-range” in our framework means far apart in graph distance on $L_n$---not necessarily in geographic or social distance. A link between geographically nearby locations is a long-range shortcut if $L_n$ provides no short path between them.

Aggregates like $\mathcal{R}_0$ may be better used as descriptive than prescriptive tools. The failure of robustness would almost certainly propagate to welfare calculations that rely on the extent or location of diffusion acemoglu2021optimal,fajgelbaum2021optimal. The susceptibility to small measurement error may argue for earlier, more aggressive policy responses barnett2023epidemic. The full decision theory exercise is beyond our scope, but the statistical force we document pushes strongly in that direction.

remark[Directed Networks] Many applied settings involve directed interactions. Sensitive dependence may be stronger in directed settings, since cones of influence from nearby seeds diverge more readily when paths are asymmetric. Volume growth assumptions would apply to out-neighborhoods. We conjecture qualitatively similar results hold under analogous directed expansion conditions.

Our results complement lewbel2024estimating, who study link misclassification in peer effects regressions. Both papers show that network measurement error has more severe consequences than standard intuitions suggest, but the mechanisms differ: misclassification biases peer-effects coefficients, while in our setting vanishingly small error causes diffusion trajectories to diverge catastrophically.

Our results are specific to SIR models, but the same perturbation robustness failure may affect general models of treatment effects with spillovers aronow2017estimating, hardy2019estimating, athey2018exact---a direction we leave to future work.