Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,609 characters · 16 sections · 103 citation commands
Decomposition of Spillover Effects Under Misspecification: Pseudo-True Estimands and a Local-Global Extension
In many empirical settings, one unit’s treatment affects the outcomes of others. Vaccination programs change not only the health of vaccinated individuals, but also the infection risks of their contacts HudgensHalloran2008. Large-scale anti-poverty programs alter local markets and prices, with consequences that propagate through space egger2022ge. The challenge in these settings is that the assignment vector acts as a single, system-wide shock rather than a collection of independent unit-level treatments: the outcome for each unit can depend on many, or even all, components of the treatment vector. From the researcher’s perspective, this means that one realized assignment from a fixed design must be used to learn about this complex pattern of dependence.
Applied work typically approaches these settings through exposure mappings AronowSamii2017: low-dimensional summaries of the underlying interference structure, such as the fraction of treated neighbors in a network cai2015social, a spatial ring kernel egger2022ge, or market prices munro2021treatment. Outcomes are modeled as depending on own treatment and this exposure measure, and empirical work reports “direct effects’’ and “spillover effects’’ defined within this reduced description of the interference structure.
A central difficulty is that exposure mappings are often inevitably misspecified savje2024causal. They compress a rich pattern of interference into a simple index that omits many potentially relevant details about who is treated and how spillovers operate. This raises a basic question: when the exposure mapping is only an approximation to the true interference structure, what policy object are exposure-based estimands actually targeting, and how should we interpret their direct and spillover components relative to the underlying policy question?
This paper answers that question by starting from a primitive policy object and then working backwards to the estimands that exposure-based methods recover. We focus on the effect of marginally changing the treatment assignment rule, often called the marginal policy effect in the recent literature carneiro2010evaluating,munro2021treatment. This quantity is defined directly from the experimental or quasi-experimental design, without reference to any exposure mapping. It has a natural interpretation as a social multiplier that captures the aggregate impact of a small change in treatment intensity, and it is often more tractable than other counterfactual quantities munro2021treatment. Many recent theoretical contributions analyze marginal policy effects of this form li2022random,munro2021treatment,arkhangelsky2025evaluating,hu2022average, and they have a number of practical applications, for example as regression coefficients in linear regressions with both an own-treatment indicator and an exposure variable egger2022ge,muralidharan2023general; see also the three examples in hu2022average.
Given this policy primitive, we then study what happens when the analyst commits to using an exposure mapping chosen on the basis of domain knowledge. We show that this restriction induces a pseudo-true outcome model: among all outcome models that depend on the assignment vector only through the chosen exposure, there is a unique model that provides the best mean-squared approximation to the true outcomes. This pseudo-true model in turn induces marginal-policy, direct, and spillover effects, and these satisfy the same decomposition as in the Hu--Li--Wager identity under correct specification. Thus, our contribution is not to claim that an arbitrary misspecified exposure mapping is automatically informative about the oracle policy effect, but rather to characterize the canonical target implied by the maintained exposure restriction. We then quantify when these pseudo-true estimands are close to their oracle counterparts. In addition, under a monotonicity condition on the exposure mapping, they admit a sign-preserving representation as nonnegative linear combinations of primitive switching contrasts. Accordingly, Section (ref) provides a general framework for interpreting what exposure-based procedures target under misspecification.
So far, we have intentionally been agnostic about the detailed structure of interference. As emphasized by leung2024discussion and auerbach2024discussion, without additional structure one should not expect exposure-based analyses to recover finer channels beyond what the maintained restriction encodes. At the same time, many applied settings feature multiple distinct spillover mechanisms, which raises the question of when the general pseudo-true objects above admit a sharper interpretation. To study this, we introduce a structured model class that nests many important empirical contexts egger2022ge,AngelucciDeGiorgi2009 and theoretical work on interference li2022random,munro2021treatment. Specifically, we focus on environments in which local network spillovers and global spillovers, such as equilibrium prices, wages, or epidemic states, operate simultaneously.
This local-global model should be viewed as a structured specialization of our general misspecification framework, in which the pseudo-true interpretation sharpens into a statement about oracle channel-specific components. In this class of models, the oracle marginal policy effect admits an asymptotic three-way decomposition into a direct effect, a local spillover effect, and a global spillover effect. Specifically, a researcher who uses only a local exposure mapping can still be viewed as targeting the local component, while a researcher who uses only a global exposure mapping targets the global component, even though each omits the other first-order channel. More generally, Section (ref) continues to allow for additional omitted channels beyond these maintained local and global ones; when those residual channels are asymptotically negligible, the corresponding pseudo-true estimands remain close to the local-global oracle targets, and when they are not, they still retain their interpretation as the best exposure-based $L^2$ approximations.
An important implication is that many existing methods are more robust than previously understood once we reinterpret their targets as channel-specific components of this pseudo-true estimand. In particular, network estimators of Li--Wager type remain consistent for the local spillover component even in the presence of global spillovers. With additional sources of variation, such as augmented randomization schemes or instrumental-variable perturbations of global state variables, the global spillover component can also be separately recovered. We illustrate this idea through a semi-synthetic experiment calibrated to real data from the large-scale cash-transfer experiment studied by filmer2023cash.
\paragraph{Roadmap.} Figure (ref) summarizes the relation between the paper’s two main conceptual components. Section (ref) develops a general pseudo-true framework for misspecified exposure mappings and studies the resulting direct, indirect, and marginal policy effects. Section (ref) then discusses motivating exposure mappings and examples in which local and global spillover channels may coexist. Building on this, Section (ref) specializes the general framework to a structured local-global environment, in which the pseudo-true objects from Section (ref) admit a sharper channel-specific interpretation and can be estimated using procedures adapted to the corresponding channel. Section (ref) evaluates these ideas in simulation and semi-synthetic designs. Section (ref) concludes.
\noindentRelated literature
Misspecification in spillover estimation and policy-relevant primitives. A large literature now treats exposure mappings and randomization-based designs as the basic language for analyzing interference Aronow2012Detect,AronowSamii2017,AtheyEcklesImbens2018,ogburn2024causal. Against this backdrop, a more recent line of work takes misspecification of spillover structures seriously loomba2025policy,shuangning_additional_draft,weinstein2026causal. In particular, SavjeAronowHudgens2021 show that ADE can be estimated even under unknown interference, while savje2024causal treats exposure mappings as researcher-defined summaries and derives conditions for consistent estimation under misspecification, prompting discussion of the policy content of the resulting exposure effects auerbach2024discussion,leung2024discussion. Relatedly, leung2022causal formalizes approximate neighborhood interference, under which standard exposure-based estimators remain well behaved even when distant treatments matter, and menzel2025fixed defines conditional-on-assignment estimands that remain identified under very general interference and can be recovered by inverse-probability weighting in single-network experiments. We introduce a pseudo-true estimand perspective and formally construct it using a two-copy conditioning device. We show approximation guarantees and characterize sign-preserving conditions, building on leung2024identifying. This parallels classic pseudo-true parameter ideas in econometrics and finance, where maximum likelihood under misspecification converges to a Kullback--Leibler projection white1982maximum and Hansen--Jagannathan distance selects the stochastic discount factor that minimizes a pricing-error norm hansen1997assessing.
Spillover decompositions and a local-global extension A separate literature uses decompositions of overall policy effects into “direct” and “indirect” components to organize mechanisms. A long line of work Sobel2006,HudgensHalloran2008 has formalized direct and indirect effects under partial interference, with extensions to general exposure mappings in AronowSamii2017 and design-averaged estimands under unknown interference in SavjeAronowHudgens2021. Within this tradition, hu2022average define average direct and indirect effects under general exposure mappings and show that, in Bernoulli trials, their sum coincides with the effect of an infinitesimal increase in the treatment probability, an approach adopted in structured settings such as the market-equilibrium model of munro2021treatment. Much of this work effectively treats the indirect component as a single channel. Our analysis shows, first, that this basic direct--indirect decomposition survives misspecification once exposure effects are interpreted as pseudo-true components of a marginal policy effect. We obtain a further sharper decomposition under a structured local-global framework, where we nest the local network framework of li2022random and the equilibrium-spillover framework of munro2021treatment; see also related global-state settings in arkhangelsky2025evaluating,halloran1991direct,lin2024sir\footnote{Other complementary work includes bhattacharya2025causal, who use mean-field methods to study global treatment effects.}. We show how to interpret these path-specific settings under other forms of misspecification, and we show that contamination from the other channel is asymptotically negligible, admitting a separable decomposition into the direct, local, and global indirect effects. Recent complementary work by Ritzwoller2025Spillovers shows that regressions on proximity-weighted treatments blend multiple channels unless the proximity measure is residualized.
Consider a sample of $n$ units indexed by $\{1,\ldots,n\}$, where each unit is assigned one of two possible treatments $\{0,1\}$. The collection of all (potentially counterfactual) assignments is thus denoted as $\bw=(w_1,\ldots,w_n)\in\{0,1\}^n$. A (possibly randomized) function $y_i:\{0,1\}^n\to\RR$ gives the potential outcome for unit $i$ under a specific assignment. We impose no additional structure on $y_i(\cdot)$ until Section (ref).
Throughout, we focus on experimental designs where the actual assignment vector $\bW\in\{0,1\}^n$ is generated randomly. In particular, we consider the simplest design, a randomized controlled trial with a homogeneous treatment probability $\pi\in(0,1)$.
Our benchmark target is the oracle marginal policy effect
This estimand is defined without reference to any exposure mapping and serves as the oracle benchmark. Since $y_i(\cdot)$ is defined on the exponentially large assignment space $\{0,1\}^n$, directly estimating (ref) is generally intractable.
Sections (ref) and (ref) therefore construct pseudo-true outcome models by conditioning on a researcher-chosen exposure mapping and use them to define tractable marginal-policy, direct, and indirect effects. These estimands admit an intervention-based interpretation as components of the induced pseudo-true model and, under monotonicity, a sign-preserving representation.
Misspecification arises whenever the researcher-chosen exposure mapping \(d_i(\bW)\) is not a sufficient summary of the assignment vector for the outcome \(y_i(\bW)\). This can happen for many reasons, but two broad motivations are especially relevant for our analysis. First, the analyst may use a low-dimensional exposure mapping as a tractable approximation to a richer interference structure, for example because the true mechanism is too complex to model directly or because tuning details of the exposure are uncertain. In this case, Section (ref) shows that the resulting pseudo-true estimands remain close to the oracle targets when the chosen exposure mapping is sufficiently informative about outcomes. Second, the analyst may intentionally work with only one exposure channel in order to obtain a more interpretable decomposition of spillovers. In that case, the resulting estimands should be viewed as channel-specific pseudo-true effects rather than as attempts to recover the full oracle object; Section (ref) shows that this interpretation is especially sharp in a local-global environment.
A large empirical literature works with exposure mappings that summarize the features of the assignment vector $\bw$ that practitioners believe to be most relevant for unit $i$. Formally, the analyst specifies $d_i:\{0,1\}^n\to\cD_i$\footnote{See, among many others, HudgensHalloran2008,tchetgen2012causal,AronowSamii2017,hu2022average,li2022random,munro2021treatment for examples in epidemiology, statistics, and economics.}. Choosing $d_i$ is problem specific and requires domain expertise. We offer several examples in Section (ref). In the well-specified setting, one posits that $y_i(\bw)$ depends on $\bw$ only through $d_i(\bw)$\footnote{Technically, one typically assumes that $y_i(\bw)$ depends on $\bw$ only through $(w_i,d_i(\bw))$. Since we can redefine the exposure mapping as $\tilde d_i(\bw) := (w_i,d_i(\bw))$, this is without loss of generality; in what follows we often treat $d_i$ as already including the own-treatment component.}
However, in realistic environments with rich local and global spillovers, misspecification of exposure mappings is hard to avoid. Under such circumstances, we aim for tractable alternatives for the oracle estimands in (ref). A large body of empirical and methodological work postulates that potential outcomes depend on $\bw$ only through an exposure mapping $d_i(\bw)$, and then estimates causal effects by working with outcome models of the form $h_i\bigl(d_i(\bw)\bigr)$—for instance by pooling outcomes across units with the same or similar exposure, or by fitting flexible regressions of $Y_i$ on $d_i(\bW)$; see, e.g., AronowSamii2017,auerbach2021local,zivich2022tmle.
In the same spirit, we build exposure–based outcome models of the form $h_i\bigl(d_i(\bw)\bigr)$ and use them to define alternative estimands of interest under interference. A natural criterion is to optimize over $h_i$ so that the following square loss is minimized:
The solution is the conditional expectation:
Here we slightly abuse notation by incorporating $\pi$ as another argument of $h_i^\ast$, simply to emphasize that this optimal solution is design-induced. We then define the corresponding exposure-based outcome model
where $\bW^{(2)}$ is an independent copy of the treatment vector, introduced to average out omitted interference conditional on the exposure. Among all outcome models that depend on $\bw$ only through $d_i(\bw)$, $\tilde{y}_i$ is the unique solution that minimizes the mean-squared discrepancy from the true $y_i$ under the design. We therefore call it pseudo-true, following the misspecification literature on pseudo-true parameters in minimum-distance, likelihood, and related settings white1982maximum,hall2003large,hansen1997assessing,andrews2024misspecified.
An important practical feature of (ref)--(ref) is that, once the exposure mapping $\{d_i\}$ is fixed, the pseudo-true outcomes are functions only of the joint distribution of $(Y_i,d_i(\bW))$ under $\mathrm{RCT}(\pi)$. In particular, any flexible estimator of a conditional expectation—including classical inverse–probability-weighted and regression estimators, as well as modern machine–learning methods for nuisance functions—can be used to approximate $h_i^\ast$ and hence $\tilde y_i(\cdot;\pi)$ without modeling the full interference structure; see, for example, chernozhukov2018double,wager2018estimation.
We now use the pseudo-true outcome models (ref) to define average effects of interest. Our estimands extend the familiar ADE/AIE objects studied under correctly specified exposure mappings in hu2022average,li2022random,munro2021treatment. In the correctly specified case, the celebrated result of hu2022average shows that, under the Bernoulli design $\mathrm{RCT}(\pi)$, the marginal policy effect $\tau_{\mathrm{MPE}}(\pi)$ admits a principled decomposition into an average direct effect and an average indirect effect. This decomposition has been used both in recent theoretical analysis (e.g., munro2021treatment,arkhangelsky2025evaluating,loomba2025policy) and in applied work (e.g., behaghel2022encouraging). Here we extend it to the misspecified case by replacing oracle outcomes $y_i(\cdot)$ with the pseudo-true outcomes $\tilde y_i(\cdot;\pi)$.
\noindentMarginal policy effect. To isolate the role of $d_i$ from other unknown interference mechanisms, consider two independent assignments $\bW^{(1)}\sim\mathrm{RCT}(\pi_1)$ and $\bW^{(2)}\sim\mathrm{RCT}(\pi_2)$. Using the conditioning idea in (ref), define
Then the marginal policy effect under exposure mappings $\{d_i\}$ and $\mathrm{RCT}(\pi)$ is given by
\noindentDirect and indirect effects. The direct and indirect treatment effects under exposure mappings $\{d_i,i\in[n]\}$ and $\mathrm{RCT}(\pi)$ are defined analogously:
When exposures are correctly specified, these estimands coincide with their oracle counterparts in (ref). Our first result states that these exposure-based estimands admit exactly the same decomposition as in hu2022average. The proof is simple and is deferred to the appendix.
savje2024causal also studies misspecified exposure mappings via conditioning, but defines causal effects as contrasts between exposure labels. Motivated by the subsequent discussions in auerbach2024discussion,leung2024discussion, we take a different route. We begin from the oracle marginal policy effect in (ref) and ask what object is induced when the analyst restricts attention to outcome models that depend on the assignment vector only through a chosen exposure mapping. The resulting estimands \(\tau_{\mathrm{MPE}}(\pi)\), \(\tau_{\mathrm{ADE}}(\pi)\), and \(\tau_{\mathrm{AIE}}(\pi)\) are therefore best viewed as the marginal-policy, direct, and spillover components of the induced pseudo-true outcome model, rather than as arbitrary contrasts between exposure labels. In this sense, our contribution is not to claim that an arbitrary misspecified exposure mapping is automatically policy-informative, but to characterize the canonical target of exposure-based procedures under misspecification. Section (ref) then quantifies the distance between these induced objects and their oracle counterparts, while Section (ref) shows that in a structured local-global environment this interpretation sharpens into channel-specific approximations to the corresponding oracle components.
A separate issue is whether these pseudo-true estimands preserve the sign of the underlying single-unit treatment contrasts. Because the pseudo-true outcomes are conditional averages rather than primitive unit-level potential outcomes, this property does not hold in full generality; leung2024identifying gives explicit counterexamples. The next proposition shows that, inspired by leung2024identifying, sign preservation is nevertheless recovered under a natural monotonicity condition on the exposure mappings. Under the Bernoulli RCT considered here, the relevant design-side positive-dependence condition is automatically satisfied, so componentwise monotonicity of the exposure mappings is sufficient.
Instead of directly adopting the pseudo-true outcome models (ref), one can more generally use any collection $f=\{f_i\}_{i\in[n]}$ with $f_i:\{0,1\}^n\to\RR$ as a candidate approximation to the oracle estimands in (ref). Specifically, define the functionals $\tau_{\mathrm{MPE}}^{\mathrm{func}}(f;\pi) = \tau_{\mathrm{ADE}}^{\mathrm{func}}(f;\pi) + \tau_{\mathrm{AIE}}^{\mathrm{func}}(f;\pi)$ by
The following proposition shows that these functionals are Lipschitz in $f$ under the $L^2$ norm. The same decomposition also holds for the oracle estimand, $\tau_{\mathrm{MPE}}^{\mathrm{oracle}}=\tau_{\mathrm{ADE}}^{\mathrm{oracle}}+\tau_{\mathrm{AIE}}^{\mathrm{oracle}}$, with
The proof of this proposition is deferred to Appendix (ref). Thus, the square-loss criterion (ref) searches for the optimal $\{f_i\}_{i\in[n]}$ subject to the compositional restriction $f_i=h_i \odot d_i$ by minimizing the right-hand side of (ref). Plugging the conditional-expectation formula for $h_i^{\ast}$ in (ref) into (ref), we immediately obtain the following corollary.
This corollary makes precise in what sense our estimands approximate the oracle targets. When there is little residual variation in outcomes after conditioning on the exposure mapping, the pseudo-true direct, indirect, and total effects are necessarily close to their oracle counterparts. In other words, within the class of outcome models that depend on treatment only through the chosen exposure mapping, any method that fits individual outcomes well also delivers a good approximation to the marginal policy effect, and our pseudo-true model is the optimal such approximation in that class.
This section provides examples of exposure mappings that fit naturally within the general framework of Section (ref). Our goal is not to be exhaustive, but to highlight two broad classes of spillover mechanisms that frequently arise in practice: local network spillovers and global equilibrium spillovers. These examples motivate the structured local-global model studied in Section (ref), where both channels coexist.
\noindentLocal network spillovers. A large recent literature studies spillovers that operate through local network structure, including applications in epidemiology, peer effects, spatial externalities, and informal insurance halloran1991direct,HudgensHalloran2008,lin2024sir,ogburn2024causal,cai2015social,fafchamps2003risk. In many such settings, units are embedded in a network encoding pairwise relationships such as social ties, geographic proximity, or technological links. Let $E \in \{0,1\}^{n\times n}$ be a symmetric adjacency matrix, where $E_{ij}=1$ indicates that units $i$ and $j$ are connected. A natural exposure mapping is \( d_i(\bw) := \{ w_j : j \neq i,\; E_{ij}=1 \}, \) which records the treatment assignments of unit $i$'s neighbors. In practice, researchers often work with lower-dimensional summaries of this vector, such as the proportion of treated neighbors or an indicator for whether at least one neighbor is treated li2022random,cai2015social.
\noindentGlobal equilibrium spillovers. Other forms of interference operate through aggregate or equilibrium mechanisms that affect all units simultaneously, including herd immunity, market-clearing prices, and centralized allocation rules halloran1991direct,lin2024sir,egger2022ge,munro2025designed,arkhangelsky2025evaluating. In such settings, an individual’s outcome may depend on the full treatment assignment only through a low-dimensional global state. This motivates exposure mappings of the form \( d_i(\bw) := P_n(\bw), \, i \in [n], \) where $P_n(\bw)$ is a scalar or low-dimensional summary induced by the full assignment $\bw$, such as an epidemic threshold or an equilibrium price.
Our main interest is in environments where these two channels coexist. In such cases, individual outcomes can be written as \( y_i(\bw) = y_i\bigl(w_i,\, S_i(\bw),\, P_n(\bw)\bigr), \) where $w_i$ is an individual treatment, $S_i(\bw)$ is a local network exposure, and $P_n(\bw)$ is a global state induced by the assignment $\bw$.
1. Vaccination on networks with herd immunity. Consider a susceptible--infected--removed (SIR) model on a contact network. Vaccination of neighbors reduces unit $i$'s infection risk through local transmission channels, which can be summarized by a network exposure such as the fraction of vaccinated neighbors HudgensHalloran2008. At the same time, aggregate vaccination levels determine whether the population crosses a herd-immunity threshold, altering infection risk for all individuals through a global channel halloran1991direct,lin2024sir. Exposure mappings that focus only on local network structure therefore confound these two mechanisms whenever global epidemic conditions also matter, echoing recent work on misspecified exposure mappings and equilibrium causal estimands savje2024causal,menzel2025fixed.
2. Market equilibrium with network externalities. A similar structure arises in market environments with both equilibrium spillovers and local interactions. Individual treatments, such as subsidies or cash transfers, may affect outcomes globally through equilibrium prices determined by market clearing, but also locally through peer effects, information transmission, or technological complementarities.\footnote{munro2021treatment write: “One unit’s treatment impacts another’s outcomes only through the treatment’s impact on the equilibrium price, which rules out peer effects or other forms of network-type interference.” munro2025designed similarly note: “There are two possible sources of interference from an information treatment; the first is spillovers through the mechanism due to capacity constraints, and the second is network spillovers. The estimates in Table 4 only account for the first type of spillover.”} Local network exposures capture peer interactions holding prices fixed, while global exposures summarize equilibrium adjustments operating at the economy-wide level. Large-scale cash-transfer experiments in Kenya illustrate general-equilibrium spillovers on non-recipients via changes in local demand and prices egger2022ge, whereas AngelucciDeGiorgi2009 show the cash-transfer program in Mexico generates local network externalities through gifts, loans, and informal risk-sharing with little evidence of local price changes. Related patterns appear in other domains: job-placement programs can raise employment for treated workers but displace untreated job seekers in the same labor markets crepon2013displacement, with referrals through social networks mediating access to jobs beaman2012referral; and school-choice reforms affect aggregate sorting and housing markets hsieh2006choice while classroom peer composition generates local externalities carrell2010externalities. These examples underscore that similar interventions can trigger either or both types of spillovers depending on scale and context.
These examples illustrate that interference often arises through multiple, conceptually distinct channels. In the next section, we formalize a structured local-global model, motivated by the market-equilibrium-with-network-externalities example, in which these channels can be analyzed jointly.
We now specialize the general pseudo-true framework of Section (ref) to a structured environment in which two first-order spillover channels coexist: local network interference and global market interference. This setting is motivated by the examples in Section (ref) and nests the two well-specified benchmark environments studied separately in li2022random and munro2021treatment. Our goal is to show that, in this structured model, the pseudo-true objects from Section (ref) admit a sharper channel-specific interpretation, and that existing methods continue to estimate the corresponding local and global components.
The environment also generates implicit functionals that output the observed outcomes
In this section, we show that the oracle marginal policy effect \(\tau_{\mathrm{MPE}}^{\mathrm{oracle}}\) admits an interpretable asymptotic decomposition into direct, local, and global components. These components are defined using the general pseudo-true framework from Section (ref), specialized here to local and global exposure mappings; see Section (ref). Moreover, the existing methods of li2022random and munro2021treatment continue to provide valid estimators for the corresponding local and global components; see Section (ref). Relative to those earlier analyses, our contribution is to show that these methods remain valid in a joint local-global environment, using in particular a higher-order expansion of Z-estimators for the empirical price variable \(P_n(\bW)\).
We now define several population quantities explicitly. Let $p_\pi^{\ast}$ be the unique population-clearing price, as in munro2021treatment, which solves
In addition, define the population gradients which are evaluated at $p_{\pi}^{\ast}$,
We also impose the structural condition that the graphon model in \hyperlink{NET}{$\mathsf{(NET)}$} is of low rank. The statistical network-analysis literature has studied the spectral decay of sparse graphon models gao2015rate,chen2025minimax.
In addition to the graphon structure in \hyperlink{NET}{$\mathsf{(NET)}$}, we assume access to the same augmented randomized trial that provides instrumental variables-like variation in \hyperlink{MAR}{$\mathsf{(MAR)}$}.
We treat $(\rho_n,h_n)$ jointly as parameters of the environment.
Consistent with Section (ref), our primary object remains the oracle marginal policy effect
This notion is well-defined regardless of the interference structure. In addition, we consider three related components.
In finite samples, there is no reason to expect \(\tau_{\mathrm{MPE}}^{\mathrm{oracle}}\) to decompose exactly into these three components. Our next result shows that such a decomposition turns out to emerge asymptotically.
Within the model (ref), the local and global spillover channels are asymptotically decoupled. Intuitively, this holds because the global and local channels correspond to fluctuations of the assignment vector in very different directions. Specifically, the global spillover channel operates through a low-dimensional, “consensus” statistic of the assignment, while the local channel operates through high-dimensional ego exposures; under the Bernoulli design these directions fluctuate at order $n^{-1/2}$ and are asymptotically uncorrelated, so only the separate local and global components contribute to the welfare derivative at first order, and their interaction is second order. Appendix (ref) presents the proof in several steps.
It is also worth noting that \(\tau_{\mathrm{AIE}}^{\mathrm{L},\ast}\) coincides with the estimand in li2022random when the equilibrium price is fixed at \(p_\pi^\ast\), whereas \(\tau_{\mathrm{AIE}}^{\mathrm{G},\ast}\) coincides with the estimand in munro2021treatment when local interference is fixed at its benchmark level \(\pi\).
We also specialize Proposition (ref) to the present local-global setting.
In the market-interference setting, monotonicity of \(z_i\) with respect to the price variable \(p\) is often natural, whereas monotonicity with respect to the treatment assignment \(w\) is application specific. In settings such as cash-transfer interventions in Section (ref), the latter condition can be plausible when treatment weakly increases recipients' excess demand at any given price, so that a higher treatment intensity exerts upward pressure on the equilibrium price.
After defining several notions of treatment effects, this section presents corresponding estimators that are consistent for the asymptotic estimands. We also derive sharp convergence rates to assess the statistical efficiency of the proposed methods.
To recap the basic setup, we observe a network $\bE\in\{0,1\}^{n \times n}$, a randomized assignment $\bW\in\{0,1\}^n$ generated according to Assumption (ref), and individualized perturbations $\bU\in\RR^{n \times J}$ generated according to Assumption (ref). We then observe realized outcomes $\bY\in\RR^{n}$ and excess demands $\bZ\in\RR^{n \times J}$. Any valid estimator must be constructed only from these observables.
\noindentDirect effect. Consistent with the practice in li2022random,munro2021treatment, we employ the usual Horvitz--Thompson estimator for $\tau_{\mathrm{ADE}}^{\mathrm{oracle}}$, which is automatically unbiased under the RCT design (Assumption (ref)):
Because our model contains two distinct spillover mechanisms, the asymptotic variance of this estimator differs from the standard benchmark. We derive it explicitly in the following theorem.
\noindentLocal spillover effect. Because we adopt the same setup as li2022random for local interference, it is natural to use the same estimator. Start by forming a vector $\bnu\in\RR^n$ of raw weights:
Compute $\hat{\bPsi}\in\RR^{n\times r}$ as the normalized top-$r$ eigenvectors of the observed adjacency matrix $\bE=(E_{ij})$ with $\hat{\bPsi}^\top\hat{\bPsi}=\Ib_r$. The PC-balancing estimator is then defined by
The following remark explains the intuition behind this estimator.
Departing from the existing theory in li2022random, the next theorem deepens our understanding of the PC-balancing estimator $\hat{\tau}_{\mathrm{AIE}}^{\mathrm{L}}$. It shows that the estimator is robust to additional unspecified market interference \hyperlink{MAR}{$\mathsf{(MAR)}$}. In particular, it still targets $\tau_{\mathrm{AIE}}^{\mathrm{L},\ast}$ with the same convergence rate. The limiting variance is also similar to that in li2022random and is therefore omitted from the main text.
\noindentGlobal spillover effect. As shown in Theorem (ref), the global spillover estimand $\tau_{\mathrm{AIE}}^{\mathrm{G}}$ depends asymptotically on the price-elasticity vector $\gamma:=\xi_z^{-1}\xi_y$ and the direct effect $\tau_z:=\EE\sbr{z_i(1,p_{\pi}^{\ast})-z_i(0,p_{\pi}^{\ast})}$ on excess demand. We can therefore estimate these two objects separately and then combine them to obtain a valid estimator of $\tau_{\mathrm{AIE}}^{\mathrm{G},\ast}=-\gamma^{\top}\tau_z$.
Price elasticities have long been a central topic in econometrics houthakker1969income,chetty2009sufficient. They can be estimated using instrumental variables angrist1996identification,berry2021foundations. For conciseness, we follow the approach of munro2021treatment, in which the experimenter creates instrumental variables by augmenting the experimental design with individualized price perturbations. The construction is stated formally in Condition (ref).
Equipped with $\bU\in\{\pm h_n\}^{n \times J}$, we estimate price elasticities by \( \hat{\gamma} = \rbr{\bU^\top\bZ}^{-1}\rbr{\bU^\top\bY}. \) After constructing a Horvitz–Thompson estimator for the treatment effect of excess demands \( \hat{\tau}_z = \frac{1}{n}\sum_{i=1}^n \rbr{\frac{W_i}{\pi}-\frac{1-W_i}{1-\pi}}Z_i, \) the final estimator is $\hat{\tau}_{\mathrm{AIE}}^{\mathrm{G}}=-\hat{\gamma}^{\top} \hat{\tau}_z$. Our next theorem studies the theoretical performance of this estimator. It consistently targets the asymptotic limit $\tau_{\mathrm{AIE}}^{\mathrm{G},\ast}$, with the same convergence rate, even under unspecified local network interference \hyperlink{NET}{$\mathsf{(NET)}$}.
Taken together, Theorems (ref) and (ref) show that the PC-balancing estimator from li2022random and the augmented-trial estimator in munro2021treatment remain asymptotically valid in the full local-global environment. Each consistently recovers the corresponding local or global component of $\tau_{\mathrm{MPE}}^{\mathrm{oracle},\ast}$ singled out by our decomposition, even though the underlying exposure mapping omits the other first-order channel.
Moreover, by Theorem (ref), we can consistently estimate the oracle marginal policy effect by summing the corresponding estimators,
Since each component estimator is consistent, it follows that $\hat{\tau}_{\mathrm{MPE}}^{\mathrm{oracle}}\overset{\mathrm{p.}}{\longrightarrow}\tau_{\mathrm{MPE}}^{\mathrm{oracle},\ast}$ as $n\to\infty$. Its convergence rate, however, is governed by the slowest component estimator.
Our first simulation setup is a fixed-index model, where the outcome \(Y_i\) of each unit ultimately depends on one aggregated index \(\eta_i\). With \(\bw\in\{0,1\}^n\) being a potential assignment, the detailed model generating process is given as below.
Henceforth, this model is described by linear coefficients $(\theta_p,\theta_\ell,\theta_g,\theta_w)=(0.5,0.5,0.8,1)$, parameters \((\rho,u)\) and a link function \(g(\cdot)\). We consider five canonical choices of \(g\), including the linear link \(g(x)=x\), a quadratic link \(g(x)=x+x^2\), a cosine link \(g(x)=\cos(x)\), a logarithmic link \(g(x)=\log(1+x^2)\), and a higher-order polynomial link \(g(x)=x+x^2+x^3\). These are denoted as \(\{\mathtt{linear},\,\mathtt{quad},\,\mathtt{cos},\,\mathtt{log},\,\mathtt{poly}\}\) later. To carry out experiments, we additionally choose the treatment assigning rate \(\pi\), individualized price perturbation size \(h\), and the rank \(r\) in the PC-balancing step.
\noindentMonte Carlo experiments with fixed sample size. Throughout this part, we set \(n=1000\). When computing the estimators, we always use individualized perturbation size \(h=0.1\) and correctly specified rank \(r=1\). Figure (ref) depicts the performance of our estimators with the \(\mathtt{linear}\) link function, and varying \((u,\pi,\rho)\). In each panel we report the average direct effect (ADE) together with the local and global components of the average indirect effect (AIE): solid curves show Monte Carlo averages, and dashed curves show the corresponding oracle quantities.
Apart from the case of \(\mathtt{linear}\) link function, Figure (ref) shows the performance of our estimators in finite samples with several non-linear link functions. This time we set \(\pi=0.5\) and \(\rho=0.01\) throughout, and only vary the mixing parameter \(u\). For each choice of link in \(\{\mathtt{quad},\mathtt{cos},\mathtt{log},\mathtt{poly}\}\), the figure displays ADE, local AIE, and global AIE separately as functions of \(u\); dashed curves indicate oracle values, and solid curves indicate Monte Carlo averages.
All the experiments so far suggest that our estimators can approximate the limiting estimands well enough in finite samples.
\noindentMonte Carlo experiments of growing sample size. Now we increase the magnitude of \(n\) to numerically check the asymptotic convergence rates shown in Section (ref). In Figure (ref), we take \(n\in\{100,\ldots,10000\}\) with \(h_n=0.75\,n^{-\alpha}\) and \(\rho_n=0.75\,n^{-\kappa}\). The pair \((\kappa,\alpha)\) takes values in \(\{(0.34,0.26),\,(0.49,0.40)\}\). We set \(u=\pi=0.5\) and plot the MSE of our ADE estimator and the local and global AIE estimators against \(n\) in log-log scale. We consider the linear DGP and the cosine DGP.
We next consider a semi-synthetic design calibrated to the cash-transfer experiment in Philippine villages studied by filmer2023cash. A previous work by munro2021treatment studies the spillover effect caused by a global market-equilibrium channel operating through village-level egg prices. We extend their calibration by further adding a local network channel operated through household neighbors, so that the marginal policy effect can be decomposed into direct, local, and global components within a realistic structural environment.
In the experiments of filmer2023cash, each household (of several individuals) is randomized to receive a cash transfer with probability $\pi$ at the household level, which exactly satisfies our Assumption (ref). The target outcome is the height-for-age Z-score for children between 0 and 5 years of age.
To conduct our synthetic experiments, a concrete parametric model is specified for the outcomes, demands/supplies and network. Then we use the real dataset from filmer2023cash to estimate the involved parameters. Subsequently, fixing $\pi$ at the level of the empirical experiments, we are able to simulate multiple experiments and thus assess the ability of our estimators. munro2021treatment adopts exactly the same semi-synthetic paradigm for their calibrated evaluation.
\noindentParametric model. Besides the following brief description, Appendix (ref) provides additional details.
(a) Market equilibrium model. We use the same setup as in munro2021treatment which incorporates the market of eggs into consideration. Conditional on treatment, each individual has linear demand and supply schedules. The equilibrium price is then determined by market clearing.
(b) Block network model. To introduce local spillovers, we build a network between households whose link probabilities combine a geography-based block component (barangay and municipality codes), homophily in housing and socioeconomic characteristics (roof and wall quality, education, assets, water/sanitation, school-age presence, income, and household size), and a triadic-closure adjustment. The probability matrix is rescaled to match a target density $\rho$ before drawing edges. Related network constructions appear in fafchamps2003risk as well.
(c) Final outcome model. The outcome of each child depends linearly on eligibility of the household for cash transfer, own treatment, a scalar equilibrium price of eggs, and the share of treated neighbor households. Then the outcome of a house is the average of its children.
\noindentFitting approach. Following munro2021treatment, all the coefficients relating to the market are set equal to the structural estimates from the remote-village subsamples in filmer2023cash. To do so, $10$ moments are computed from the real dataset. Additionally, the coefficient of local exposure is calibrated from a partial-linear regression of child height-for-age on the share of treated neighbors, controlling for own treatment and village prices. A detailed description can be found in Appendix (ref).
\noindentSimulation procedure. For each Monte Carlo replication, we: (i) subsample $2{,}000$ households and draw household treatments from the Bernoulli design, (ii) resample the household network from the categorical features of the subsampled households, (iii) solve for the market-clearing price and apply a small scalar perturbation, (iv) simulate individual demand, supply, and child outcomes under the calibrated structural system, and (v) aggregate to household-level excess demand and child outcomes. We then compute the Horvitz--Thompson estimator of the average direct effect, the PC-balanced estimator of the local component, and the augmented-trial IV estimator of the global component. This process is repeated for $2{,}000$ times. Closed-form expressions for the corresponding oracle targets $\tau_{\mathrm{ADE}}^\ast,\tau_{\mathrm{AIE}}^{\mathrm{L},\ast},\tau_{\mathrm{AIE}}^{\mathrm{G},\ast}$ are derived in Appendix (ref).
\noindentResults. Table (ref) summarizes our findings. The “Truth” column reports the analytical limits from the structural model; the remaining columns report Monte Carlo means, biases, and standard deviations of the estimators. Appendix Figure (ref) provides histograms of the Monte Carlo sampling distributions for the ADE, local AIE, and global AIE estimators.
The three components are well recovered by their corresponding estimators: Monte Carlo means lie close to the oracle targets. In this calibration, the local component is sizeable and negative, reflecting that treated neighbors reduce the marginal gains from one’s own transfer, while the global component captures the additional negative effect of equilibrium price changes. The decomposition shows that looking only at the average direct effect would obscure these indirect channels, even in a setting closely matched to a real cash-transfer experiment.
\noindentEconometric insights. Ex ante, the sign of $\tau_{\mathrm{TOT}}^{\ast}$ is ambiguous. On the one hand, transfers can relax liquidity constraints and increase gifts and informal loans within the network, raising non-labor income and consumption for both treated and untreated households fafchamps2003risk. On the other hand, higher transfer intensity can generate congestion in non-priced local amenities and adverse price or equilibrium effects in thin markets, so that the indirect components $\tau_{\mathrm{AIE}}^{\mathrm{L},\ast} $ and $\tau_{\mathrm{AIE}}^{\mathrm{G},\ast} $ may be negative and potentially dominate the direct gain. Recent work on large-scale public programs documents that such general-equilibrium spillovers can be first-order and even larger than the direct effect muralidharan2023general, while market-based models of status and consumption in networks show that equilibrium adjustments can overturn the sign of aggregate welfare effects ghiglino2010keeping. Our calibration should therefore be viewed as one plausible configuration in which $\tau_{\mathrm{ADE}}^{\ast} > 0$ but $\tau_{\mathrm{TOT}}^{\ast} < 0$ because negative local and global spillovers dominate the positive direct effect.
This paper develops a framework for interpreting exposure-based spillover estimands under complex interference and misspecified exposure mappings, and for relating them to underlying policy objects. In many empirical applications, researchers summarize the influence of others’ treatments through low-dimensional exposure mappings, such as the share of treated neighbors in a network or an equilibrium price induced by aggregate treatment intensity, rather than modeling the full assignment vector AronowSamii2017,cai2015social,donaldson2016railroads,munro2021treatment,egger2022ge. We take this practice seriously and ask a basic question: when the chosen exposure mapping is only an approximation to the true interference structure, what policy object are exposure-based procedures targeting?
Our first contribution is to answer this question in a general way. For any analyst-chosen exposure mapping and treatment assignment rule, we define a pseudo-true outcome model as the best mean-squared approximation to the true potential-outcome model among functions that depend on the assignment vector only through the chosen exposure mapping. This induces corresponding pseudo-true marginal policy, direct, and spillover effects. A key result is that the familiar decomposition of the marginal policy effect into direct and spillover components continues to hold exactly for these pseudo-true objects, even when the exposure mapping is misspecified. In this sense, Section (ref) characterizes the pseudo-true targets induced by exposure-based procedures under arbitrary misspecification and the relation of these targets to the corresponding oracle policy objects.
This perspective also clarifies in what sense pseudo-true estimands approximate the oracle policy targets defined without exposure mappings. When the chosen exposure mapping is sufficiently informative, in the sense that little residual outcome variation remains after conditioning on exposure, the pseudo-true direct, spillover, and marginal policy effects are asymptotically close to their oracle counterparts. More generally, we show that approximation error in the induced policy object is controlled by approximation error in the underlying outcome model. Thus, among all outcome models restricted to depend on treatment only through the analyst’s chosen exposure mapping, the pseudo-true model delivers the closest approximation to the oracle policy estimand. Under an additional monotonicity condition on the exposure mapping, these estimands also admit a sign-preserving interpretation, which strengthens their causal interpretability.
Our second contribution specializes the general pseudo-true framework to a structured environment with two first-order spillover channels—a local network channel and a global equilibrium channel—in which outcomes depend on own treatment, a sparse-network local exposure, and an equilibrium price vector determined by aggregate excess demand. This setting nests important benchmark models from the literatures on network interference and general-equilibrium treatment effects AngelucciDeGiorgi2009,egger2022ge,li2022random,munro2021treatment, and yields an asymptotic decomposition of the oracle marginal policy effect into direct, local spillover, and global spillover components. The payoff of this structure is that it sharpens the general misspecification result into a channel-specific interpretation: exposure mappings that retain only the local channel or only the global channel can still be viewed as targeting the corresponding component of a common oracle policy object, provided any remaining omitted channels are asymptotically negligible in the sense of our approximation condition. This also provides a unified lens on methods often studied separately. Network-based procedures using local exposure mappings and spectral balancing estimators recover the local spillover component li2022random, while augmented randomized designs with exogenous price perturbations recover the global spillover component through an instrumental-variables logic munro2021treatment. Standard estimators of the direct effect remain centered on the corresponding direct component, though their variance reflects both local and global spillovers. Simulations and a semi-synthetic application calibrated to the cash-transfer experiment in filmer2023cash show that these components can be recovered in realistic experimental designs.
Taken together, our results suggest a disciplined way to think about decomposition of spillover effects under misspecification. The general pseudo-true framework characterizes the induced targets of exposure-based procedures under arbitrary misspecification and clarifies when those targets are close to the corresponding oracle objects. Additional structure can then sharpen that interpretation, as in our local-global application, where pseudo-true estimands become channel-specific components of a common oracle marginal policy effect. In this sense, exposure mappings should not be viewed only as a source of misspecification, but also as a disciplined way of isolating interpretable channels of policy transmission.
Several extensions seem particularly promising. First, ideas from doubly robust inference jiang2025new,scharfstein1999adjusting,bang2005doubly might help us extend to observational studies. Second, our analysis focuses on local marginal changes in the treatment rate. Extending the pseudo-true decomposition to other policy objects, such as global average treatment effects or larger design shifts, would clarify how far this framework can be pushed in settings where system-wide spillovers are empirically important faridani2024linear,walker2024slack. Third, the decomposition into direct, local, and global components naturally points toward new experimental designs under interference. Randomized saturation and multilevel designs, as well as network-aware assignment rules such as graph cluster randomization UganderEtAl2013, provide natural starting points for allocating statistical power across different spillover channels. Lastly, although our structured application is cast in a market-equilibrium setting, similar local-global tensions arise in epidemiology, public health, education, and social programs, where interventions generate both neighborhood-level and aggregate spillovers. Applying and adapting the pseudo-true framework in these domains may help organize a wide range of direct, spillover, and policy effects within a common language.