EconBase
← Back to paper

Identifying Treatment and Spillover Effects Using Exposure Contrasts

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

55,190 characters · 12 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identifying Treatment and Spillover Effects Using Exposure Contrasts

\onehalfspacing

abstract{\sc Abstract.} To report spillover effects, a common practice is to regress outcomes on statistics summarizing neighbors' treatments. This paper studies nonparametric analogs of these estimands, which we refer to as exposure contrasts. We demonstrate that a contrast may have the opposite sign of the unit-level effects of interest even under unconfoundedness. We then provide interpretable conditions on interference and the assignment mechanism under which exposure contrasts can be represented as convex averages of the unit-level effects and therefore avoid sign reversals. These conditions encompass cluster-randomized trials, network experiments, and observational settings with peer effects in selection into treatment. {\sc JEL Codes}: C21, C31, C57 {\sc Keywords}: causal inference, identification, interference, peer effects

Introduction

Consider a large set of $n$ units, and for each unit $i$, let $Y_i$ denote its outcome and $D_i$ a binary treatment. We study settings with interference in which outcomes may depend on the entire treatment assignment vector $\bm{D} = (D_i)_{i=1}^n \in \{0,1\}^n$. Because the causal effect of $\bm{D}$ on $Y_i$ is difficult to convey, let alone identify, a common strategy is to regress $Y_i$ on a substantially lower-dimensional vector $T_i$ that parsimoniously summarizes $\bm{D}$, for example the number of treated neighbors. We consider nonparametric analogs of these regression estimands, which take the form

equation*[equation* omitted — 150 chars of source]

If $T_i$ counts $i$'s treated neighbors, $\tau(t,t')$ compares the average outcomes of units with different numbers of treated neighbors, controlling for $\mathcal{C}_i$. The literature refers to $T_i$ as an {\em effective treatment} or {\em exposure mapping} manski2013identification,aronow2017estimating. We refer to $\tau(t,t')$ as an {\em exposure contrast} and study when and in what sense it has a causal interpretation.

lu2019place employ this empirical strategy in their study of the impact of Chinese special economic zones (SEZs) on village-level outcomes. In their setting, $D_i=1$ if village $i$ is situated in an SEZ. Because there may be spillovers from neighboring SEZs, the authors regress outcomes on an exposure mapping that includes $D_i$ and an indicator for having a treated neighbor, that is, a village in $i$'s county that lies in an SEZ. baird2018optimal, cai2015social, and miguel2004worms instead measure spillovers using the number or share of treated peers within a neighborhood or cluster.\footnote{donaldson2016railroads, kline2014local, and zheng2017birth employ similar strategies.} In all cases, the coefficient on own treatment $D_i$ is intended to capture a direct effect while the coefficient on the statistic involving neighbors' treatments is intended to capture a spillover effect. The question is whether these interpretations are warranted.

These regressions are likely intended as devices for producing summary measures of spillover effects (i.e.\ exposure contrasts) rather than as structural models of interference. Yet the causal literature predominantly treats the exposure mapping as structural in the sense that it entirely summarizes the effect of $\bm{D}$ on $Y_i$. For most exposure mappings used in the literature, this is incompatible with endogenous peer effects, a leading explanation for interference in many social and economic contexts jackson2022inequality,sacerdote2011peer, since outcomes depend on the entirety of $\bm{D}$ under the reduced form of a simultaneous-equations model. It stands in contrast to a large literature on social interactions that specifies structural models with endogenous peer effects.\footnote{E.g.\ blume2015linear, bramoulle2009identification, lazzati2015treatment, lewbel2023social, and manski1993identification.} These models have richer microfoundations, but point identification typically relies on parametric assumptions. To bridge the two approaches, this paper studies the causal interpretation of exposure contrasts in a nonparametric setting where exposures are not structural.

We demonstrate that, even when assignment is unconfounded, an exposure contrast can have the opposite sign of the unit-level effects. The sign reversal occurs if both (i) interference is more complex than what the exposure mapping dictates and (ii) treatments are correlated across units. We then provide interpretable nonparametric restrictions on either (i) or (ii) that, together with unconfoundedness, ensure that common exposure contrasts can be represented as convex averages of the unit-level effects.

Our first identification result pertains to “monotone” exposures such as counts of treated neighbors. When treatment assignments satisfy a certain positive dependence condition, we show $\tau(t,t') = n^{-1} \sum_{i=1}^n {\bf E}[Y_i(\bm{D}_{i,t}^*) - Y_i(\bm{D}_{i,t'}^*) \mid \mathcal{C}_i]$ for some monotone couplings $\bm{D}_{i,t}^* \stackrel{a.s.}\geq \bm{D}_{i,t'}^*$ (see (ref)). The result imposes no restrictions on interference. The positive dependence condition is satisfied when treatment selection is governed by a game of incomplete information or the Ising model from statistical mechanics. We provide game-theoretic microfoundations for the latter using techniques from mele2017structural.

Our second result considers estimands for cluster-randomized trials, which compare units in clusters assigned to distinct saturation levels. For example the “overall effect” compares mortality rates between clusters with different vaccination rates. These have causal interpretations under the restrictive “stratified interference” assumption that the share of treated peers entirely mediates interference, but their interpretations more generally have not been studied. We derive convex average representations without imposing any restrictions on interference within or across clusters (see (ref)).

The previous results utilize restrictions on (ii). We also provide results that leave the assignment mechanism unrestricted and instead impose assumptions on interference. These are weaker than the assumption of structural exposure contrasts and allow for endogenous peer effects (see (ref)).

A convex average representation provides a sense in which an exposure contrast can be considered causal, but whether it is “policy relevant” is a different matter auerbach2024discussion. Under the stable unit treatment value assumption (SUTVA), common policy effects are convex averages of unit-level effects with specific weights heckman2007econometric. Identifying analogous policy effects under interference would presumably require conditions that eliminate sign reversals, which this paper provides.

{\bf Related Literature.} Compared to the peer effects literature, the unit-level effects we study are reduced-form in that they do not distinguish between endogenous and exogenous peer effects manski1993identification. The upshot is that identification is possible without imposing parametric structure, in the spirit of manski2013identification. Several of our results allow for peer effects in outcomes and selection into treatment. balat2023multiple also consider strategic interactions in selection and derive bounds on the average treatment effect.

A large literature studies the causal interpretations of various regression estimands under SUTVA when treatment effects are heterogeneous blandhol2022tsls,bugni2023decomposition,de2020two,goldsmith2024contamination,small2017instrumental. Sign reversals can occur due to the use of linear regression, and reversals can be avoided by using estimators directly targeting nonparametric estimands. This is not the case in our setting. We study a nonparametric estimand, and reversals occur not due to heterogeneity but rather the combination of interference and correlated treatments across units.

savje2024causal observes that exposure mappings serve two distinct roles in the literature: “to define the effect of interest and to impose assumptions on\ldots interference.” He argues that they should only be used for the former purpose. He refers to $\tau(t,t')$ as the “expected exposure effect” and to the special case of $T_i=D_i$ as the “average distributional shift effect” (his \S S2), implying that $\tau(t,t')$ has a causal interpretation when treatment is randomized. To the contrary, we show that, even under randomized assignment, $\tau(t,t')$ can exhibit unpalatable sign reversals that are only avoided under additional restrictions.\footnote{In (ref) of the appendix, we discuss the sign preservation criterion proposed by savje2024rejoinder and how it relates to our results.}

Our identification results complement work on estimation and inference for exposure contrasts when exposure mappings are not structural. leung2022causal and leung2024graph study large-sample inference on $\tau(t,t')$ under asymptotics sending $n\rightarrow\infty$. savje2024causal states high-level conditions for consistent estimation. None formally study the causal interpretation of $\tau(t,t')$.\footnote{The first working paper version of leung2022causal provides a limited formal discussion of their causal interpretation under Bernoulli-randomized designs leung2019causal. An earlier draft of leung2024graph included some of the identification results in (ref).}

{\bf Outline.} We describe the basic setup in the next section and subsequently organize results by the classes of exposure mappings to which they pertain. Section (ref) considers monotone exposures, which are increasing in the assignment vector, and provides a motivating sign reversal example. In (ref), we turn to exposure contrasts common in the literature on cluster-randomized trials. In (ref), we study $K$-neighborhood exposure mappings, which summarize the treatment configuration within a local network neighborhood. Section (ref) concludes. Proofs of these results can be found in (ref).

Setup

Recall that $\bm{D} = (D_i)_{i=1}^n \in \{0,1\}^n$ is the observed assignment vector. We refer to its distribution as the {\em assignment mechanism}. For all $i \in \mathcal{N}_n = \{1,\ldots,n\}$, let $Y_i(\cdot)$ be a random mapping from $\{0,1\}^n$ to $\mathbb{R}$. We interpret $Y_i(\bm{d})$ as the potential outcome of unit $i$ under the counterfactual that the treatment assignment vector is $\bm{d} = (d_i)_{i=1}^n \in \{0,1\}^n$ so that the observed outcome $Y_i$ equals $Y_i(\bm{D})$. Interference arises because potential outcomes may depend not only on own assignment $D_i$ but also on the entire assignment vector.

Let $\mathcal{C}_i$ denote an array of control variables for unit $i$, the choice of which we discuss in (ref). We assume treatment assignments are unconfounded in the following sense.

myassump{UC} $Y_i(\cdot) \perp\!\!\!\perp \bm{D} \mid \mathcal{C}_i$ for all $i\in\mathcal{N}_n$.

We primarily consider exposure mappings that are deterministic functions of the assignment vector:

equation*[equation* omitted — 131 chars of source]

Most of the literature assumes the exposure mapping is {\em structural} in the following sense aronow2017estimating,forastiere2021identification,ogburn2024causal.

definitionAn exposure mapping $f$ is {\em structural} if $Y_i(\bm{d}) = Y_i(\bm{d}')$ for all $i$ and $\bm{d},\bm{d}' \in \{0,1\}^n$ such that $f(i,\bm{d}) = f(i,\bm{d}')$.

In this case, we can rewrite potential outcomes as $Y_i(f(i,\bm{d}))$, so under (ref), $\tau(t,t')$ reduces to $n^{-1} \sum_{i=1}^n {\bf E}[Y_i(t) - Y_i(t') \mid \mathcal{C}_i]$, which has a transparent causal interpretation. We will consider weaker assumptions allowing for complex forms of interference such as endogenous peer effects.

Throughout the paper we maintain the {\em overlap condition} that the conditional distribution $\bm{D} \mid T_i=s, \mathcal{C}_i=c$ exists for all $c$ in the support of $\mathcal{C}_i$, $i\in\mathcal{N}_n$, and $s \in \{t,t'\}$. For instance if $T_i$ is the number of $i$'s treated neighbors, $t=2$, and $t'=1$, overlap implies that the exposure contrast only averages over units $i$ with at least 2 neighbors.\footnote{We leave this implicit in the notation since the neighborhood structure is typically treated as fixed or conditioned upon.}

remarkThis paper is concerned with identification, and as such, we treat $\tau(t,t')$ as known to the econometrician. leung2024graph studies doubly robust estimation of $\tau(t,t')$ under network interference and asymptotics sending the number of units $n$ to infinity. leung2025cluster studies cluster-randomized trials with spatial interference under the same asymptotics. Both papers require additional conditions for weak dependence that are not necessary for most of our results. The main condition is that interference decays sufficiently quickly with distance (see (ref) in (ref)). This is substantially weaker than assuming a structural exposure mapping and allows for endogenous peer effects leung2022causal.

Monotone Exposures

This section considers the following class of “monotone” exposure mappings.

myassump{MON} For any $i \in \mathcal{N}_n$, $f(i,\cdot)$ is componentwise nondecreasing.
example[Neighborhood Counts] Call $f(i,\bm{d}) = (d_i, \sum_{j=1}^n A_{ij} d_j)$ the {\em treated neighbor count}, where $A_{ij}$ is an indicator for whether units $i$ and $j$ are neighbors. Neighbors could be social contacts, units in the same “cluster,” units within a certain geographic distance band, etc. In place of the count, the literature also uses the share of treated neighbors cai2015social and the indicator $\bm{1}\{\sum_{j=1}^n A_{ij} d_j > 0\}$ lu2019place, which are also monotone.

The exposure contrast using the treated neighbor count is intended to capture either a direct or spillover effect depending on the choices of $t$ and $t'$ and whether they vary the first or second component. The question is what justifies such an interpretation. The next subsection shows that this contrast generally does not preserve the sign of the relevant unit-level effects even under unconfoundedness. In (ref), we provide a restriction on the assignment mechanism under which $\tau(t,t')$ is a convex average of the unit-level effects and hence avoids unpalatable sign reversals. In the remaining subsections, we provide primitive sufficient conditions for the restriction, demonstrating that it allows for peer effects in selection.

Sign Reversal

Consider a setting with no control variables $\mathcal{C}_i$ and $n/4$ identical clusters, each with 4 units (“neighbors”) labeled 1--4. Within a cluster, the potential outcome of a unit $i \in \{1,\ldots,4\}$ is denoted by $Y_i(d_1,d_2,d_3,d_4)$ where $d_j$ denotes the treatment of the $j$th neighbor. We leave the cluster index implicit on account of clusters being identical.

Consider the treated neighbor count from (ref). Let $t=(1,2)$ and $t'=(1,1)$, so $\tau(t,t')$ compares average outcomes of treated units with either 2 or 1 treated neighbors. The unit-level effects of interest then include comparisons such as

equation*[equation* omitted — 47 chars of source]

This is the spillover effect for a treated unit from having one additional neighbor treated relative to a baseline of 1 treated neighbor. Contrast this with

equation*[equation* omitted — 46 chars of source]

which involves moving neighbor 4 out of treatment and the other neighbors into treatment. Our position is that the exposure contrast is intended to capture the effect of moving one additional unit into treatment, not simultaneously switching neighbors in and out of treatment. The distinguishing feature of the first comparison is that the assignment vectors are partially ordered, unlike those of the second. Thus in general, the unit-level comparisons of interest are spillover effects of the form

equation[equation omitted — 153 chars of source]

We next demonstrate that the sign of $\tau(t,t')$ can be entirely inconsistent with the signs of (ref) even under a randomized control trial. Suppose potential outcomes are given by

center[center omitted — 200 chars of source]

and $Y_i(\bm{d})=0$ for all other $\bm{d}$ and $i\neq 1$. Observe that all unit-level effects of the form (ref) are non-negative and, for unit 1, strictly positive. Consider the cluster-randomized trial that assigns treatments independently across clusters, such that the distribution of within-cluster treatments only places positive probability on the vectors $(1,1,1,0)$ and $(1,0,0,1)$. Since clusters are identical with four units each,

equation*[equation* omitted — 120 chars of source]

The sign reversal occurs because the exposure mapping is not structural, and treatment assignments are correlated across units. If the exposure mapping were structural, which is a restriction on interference, then no sign reversal would occur, per the discussion following (ref). If treatments were i.i.d., which is a restriction on the assignment mechanism, then one can calculate that $\tau(t,t') > 0$.\footnote{Let $\bm{D}_{(j)}$ be the treatment subvector of any cluster $j$. Then ${\bf P}(\bm{D}_{(j)}=\bm{d} \mid T_1=(1,2)) = 1/3$ for all $\bm{d}$ such that $f(1,\bm{d})=(1,2)$, and ${\bf P}(\bm{D}_{(j)}=\bm{d}' \mid T_1=(1,1)) = 1/3$ for all $\bm{d}'$ such that $f(1,\bm{d}')=(1,1)$, so $\tau(t,t') = (1.5/3+2.5/3+1 - 0 - 1/3 - 2/3)/4 > 0$.}

The remainder of the paper considers weaker restrictions on interference and the assignment mechanism under which $\tau(t,t')$ avoids sign reversals. The first three theorems leave interference entirely unrestricted, while the last two leave the assignment mechanism unrestricted. All results allow for endogenous peer effects, unlike the assumption of structural exposures.

Representation Result

Our first result imposes the following positive dependence condition on the assignment mechanism.

myassump{MTP} Let $p_i(\cdot \mid c)$ denote the conditional probability mass function (PMF) of $\bm{D}$ given $\mathcal{C}_i=c$. For all $i\in\mathcal{N}_n$ and $c$ in the support of $\mathcal{C}_i$, $p_i(\cdot \mid c)$ is {\em multivariate totally positive of order 2 ($\text{MTP}_2$)} in that, for all $\bm{d},\bm{d}' \in \{0,1\}^n$, \begin{equation*} p_i(\bm{d} \wedge \bm{d}' \mid c) p_i(\bm{d} \vee \bm{d}' \mid c) \geq p_i(\bm{d} \mid c) p_i(\bm{d}' \mid c).\footnote{The symbols “$\wedge$” and “$\vee$” respectively denote the componentwise minimum and maximum.} \end{equation*}

$\text{MTP}_2$ is a model of positive dependence introduced by fortuin1971correlation. By their Proposition 1, known in statistical mechanics as the “FKG theorem,” $\text{MTP}_2$ implies that $\text{Cov}(f_1(\bm{D}), f_2(\bm{D}) \mid \mathcal{C}_i=c) \geq 0$ for all componentwise nondecreasing $f_1,f_2$.

We will provide selection models that satisfy (ref). First we state the result. Let $p_{i,t}(\cdot \mid c)$ denote the conditional PMF of $\bm{D}$ given $T_i=t, \mathcal{C}_i=c$.

theoremLet $t\geq t'$. Under Assumptions (ref), (ref), and (ref), for all $i \in \mathcal{N}_n$ there exists a monotone coupling $\bm{D}_{i,t}^* \stackrel{a.s.}\geq \bm{D}_{i,t'}^*$ with $\bm{D}_{i,s}^* \sim p_{i,s}(\cdot \mid \mathcal{C}_i)$ for all $s \in \{t,t'\}$ such that \begin{equation*} \tau(t,t') = \frac{1}{n} \sum_{i=1}^n {\bf E}\big[ Y_i(\bm{D}_{i,t}^*) - Y_i(\bm{D}_{i,t'}^*) \mid \mathcal{C}_i \big]. \end{equation*}

The result states that exposure contrasts can be represented as convex averages of unit-level effects of the form $Y_i(\bm{d}) - Y_i(\bm{d}')$ for $\bm{d} \geq \bm{d}'$ such that $f(i,\bm{d}) = t$ and $f(i,\bm{d}') = t'$. These are exactly the unit-level effects in (ref). The weights in the average are determined by the conditional distributions of assignment vectors $p_{i,s}(\cdot \mid \mathcal{C}_i)$ for $s \in \{t,t'\}$.

The convex average does not include comparisons of the form $Y_i(\bm{d}) - Y_i(\bm{d}')$ for which $\bm{d},\bm{d}'$ are not partially ordered. As discussed in (ref), these are undesirable because they involve simultaneously moving units into and out of treatment. When the exposure mapping is monotone, an increase in its value pushes the conditional assignment distribution towards larger values of $\bm{D}$, as shown in the proof of (ref). The unit-level effects of interest should then correspond to a thought experiment in which we monotonically increase the assignment vector.

The representation we obtain is nontrivial precisely because we must restrict the comparisons included in the average. By definition, $\tau(t,t')$ is a difference of two convex averages, and under unconfoundedness, this can always be represented as a convex average of differences if we include all possible unit-level comparisons of the form $Y_i(\bm{d})-Y_i(\bm{d}')$ in the average (see the proof of (ref)). We require additional restrictions like (ref) to exclude the undesirable comparisons.

remarkConsider the exposure contrast in (ref). If $t=(1,0)$ and $t'=(0,0)$, it seems natural to interpret $\tau(t,t')$ as a direct effect, that is, an average of unit-level treatment effects $Y_i(1,\bm{d}_{-i}) - Y_i(0,\bm{d}_{-i})$ for $\bm{d}_{-i} \in \{0,1\}^{n-1}$. (ref) does not guarantee such an interpretation. Treatments may be positively correlated under (ref), so when $D_i=1$, more alters may be treated than under $D_i=0$. The convex average may then include comparisons of the form $Y_i(1,\bm{d}_{-i}) - Y_i(0,\bm{d}_{-i}')$ for $\bm{d}_{-i} > \bm{d}_{-i}'$, reflecting treatment {\em and} spillover effects. To interpret $\tau(t,t')$ as a direct effect, we require treatments to be conditionally independent for reasons discussed below (ref). The formal result is given in (ref) in the appendix which combines Theorems (ref) and (ref).

Conditionally Independent Assignments

If the assignment mechanism is such that $\{D_i\}_{i=1}^n$ is independently distributed conditional on $\mathcal{C}_i$ for any $i\in\mathcal{N}_n$, then (ref) is immediate. We next discuss several examples.

{\bf Experimental Data.} There is a growing literature on experimental design under network interference. Network targeting experiments may randomize treatments among the subset of nodes with a particular local network configuration, what are sometimes referred to as “seeds” or “injection points” beaman2021can,kim2015social. Treatments are then functions of the network $\bm{A}$ and possibly unit-level covariates $\bm{X}$, but remain independent conditional on $(\bm{X},\bm{A})$. Then (ref) holds with $\mathcal{C}_i = (\bm{X},\bm{A})$ for all $i$. We provide further discussion of controls of this sort below.

Many proposed designs induce correlation in assignments beyond stratification, for example the balancing designs of basse2018model, the independent-set design of karwa2018systematic, and the quasi-coloring design of jagadeesan2020designs. These papers all assume particular structural exposure mappings. (ref) provides a reason to prefer conditionally independent designs, namely to ensure that the causal interpretations of their estimands are robust to complex forms of interference. In (ref), we state results that leave the assignment mechanism unrestricted and hence allow for correlated designs.

{\bf Observational Data.} leung2024graph propose a nonparametric model of network interference that allows for strategic interactions in both the outcome stage and selection stage. Let

equation[equation omitted — 136 chars of source]

for all $i\in\mathcal{N}_n$, where $\bm{A}$ is the network (formally an $n\times n$ matrix), $\bm{X} = (X_i)_{i=1}^n$ an array of unit-level observables, $\bm{\varepsilon} = (\varepsilon_i)_{i=1}^n$ an array of outcome unobservables, $\bm{\nu} = (\nu_i)_{i=1}^n$ an array of selection unobservables, and $\{(g_n,h_n)\}_{n\in\mathbb{N}}$ a sequence of function pairs such that each $g_n(\cdot)$ has range $\mathbb{R}$ and $h_n(\cdot)$ has range $\{0,1\}$. The timing of the model is that nature draws $(\bm{A},\bm{X},\bm{\varepsilon},\bm{\nu})$; units select into treatment according to a simultaneous-equations model with reduced form $h_n(\cdot)$; and outcomes are realized according to a simultaneous-equations model with reduced form $g_n(\cdot)$.

exampleSuppose selection into treatment is determined by a game of incomplete information in which units take up treatment to maximize expected utility \begin{equation} D_i = \bm{1}\big\{ {\bf E}_i[U_i(\bm{D}_{-i}, \bm{X}, \bm{A}, \bm{\nu}) \mid \bm{X}, \bm{A}, \nu_i] > 0 \big\} \end{equation} where $(\bm{X},\bm{A},\nu_i)$ is the information set of unit $i$ and ${\bf E}_i[\cdot]$ is the expectation taken with respect to $i$'s beliefs bajari2010estimating,xu2018social. Under the usual assumption that equilibrium selection only depends on public information $(\bm{X},\bm{A})$, the selection model can be represented as \begin{equation} D_i = h_n(i,\bm{X},\bm{A},\nu_i). \end{equation}

Under model (ref), potential outcomes are given by $Y_i(\bm{d}) = g_n(i,\bm{d},\bm{X},\bm{A},\bm{\varepsilon})$. Since these and treatments depend on the entirety of $(\bm{X},\bm{A})$, to account for high-dimensional network confounding, we take

equation[equation omitted — 109 chars of source]

leung2024graph provide conditions under which doubly-robust estimation of exposure contrasts is feasible with controls (ref).

propositionSuppose the assignment mechanism is governed by model (ref), and controls are given by (ref). If $\{\nu_j\}_{j=1}^n$ is independently distributed conditional on $(\bm{X},\bm{A})$, then so is $\{D_i\}_{i=1}^n$, and (ref) holds.

The proof is straightforward and omitted. The empirical games literature studying estimation of (ref) under large-market asymptotics typically assumes private information is i.i.d.\ and independent of public information lin2021selection,lin2017estimation,xu2018social. This implies the conditional independence restriction in the theorem.

By (ref) and (ref), we can obtain a convex average representation for $\tau(t,t')$ without having to impose any restrictions on interference or the magnitude of peer effects in selection.

Ising Model

The Ising model of ferromagnetism has been applied in sociophysics to model peer effects, opinion dynamics, and other forms of collective behavior macy2024ising,mullick2025sociophysics. Under this model,

equation[equation omitted — 180 chars of source]

for some constants $\beta, h_i(c), J_{ij}(c)$. If we choose controls as in (ref), then $h_i(\mathcal{C}_i)$ may be a function of own covariates $X_i$ or a network centrality measure, while $J_{ij}(\mathcal{C}_i)$ may be a function of $A_{ij}$, the $ij$th entry of $\bm{A}$. By Proposition 3.6 of lauritzen2021total, if $J_{ij} = J_{ji}$ for all $i,j$, then (ref) is $\text{MTP}_2$ in the ferromagnetic regime $J_{ij} \geq 0$ for all $i\neq j$.

The Ising model has received relatively little attention in economics, perhaps due to a lack of apparent microfoundations. We next show that it corresponds to the stationary distribution of actions under a certain dynamic game. mele2017structural microfounds the exponential random graph model as the stationary distribution of a dynamic model of strategic network formation. The next result for the Ising model is the analog for binary games on networks.

Consider $n$ agents connected through a network $\bm{A}$, which is a symmetric, non-negative $n\times n$ matrix with zero diagonals. Each agent $i$ is endowed with covariates $X_i$ and utility function

equation*[equation* omitted — 80 chars of source]

where $d_i$ is $i$'s binary action, $\bm{d}=(d_i)_{i=1}^n$, $\phi_{ij} = \phi_{ji}$ for all $(i,j)$, and both $a_i$ and $\phi_{ij}$ may be functions of $\bm{X}=(X_i)_{i=1}^n$. The $a_i$ coefficient captures direct benefits of choosing action 1, while the $\phi_{ij}$ coefficients capture peer effects.

Actions evolve over the course of the following dynamic process. Given an initial action vector $\bm{D}^0 \in \{0,1\}^n$, at each period $t$, a single agent is randomly chosen and allowed to update their action by myopically best-responding to $\bm{D}^{t-1}$. The new action vector is denoted by $\bm{D}^t$ and only differs from $\bm{D}^{t-1}$ if the chosen agent changes their action relative to its previous state.

Formally, agent $i$ is randomly chosen in period $t$ with probability $\rho(i, \bm{D}^{t-1}_{-i}, \bm{X})$, which is strictly positive for all arguments, where $\bm{D}^{t-1}_{-i}$ is the subvector of $\bm{D}^{t-1}$ excluding component $i$ mele2017structural. The action vector is updated to $\bm{D}^t$ by replacing the $i$th component in $\bm{D}^{t-1}$ with

equation[equation omitted — 136 chars of source]

where $(d,\bm{D}_{-i}^{t-1})$ is the vector $\bm{D}^{t-1}$ with its $i$th component replaced by $d$, and the random-utility shock $\varepsilon_{it}$ has a Type I extreme value distribution and is i.i.d.\ across agents and time mele2017structural.

remarkbadev2021nash considers a similar model that additionally features endogenous link formation, meaning the randomly chosen agent reoptimizes over a subset of links. His payoff function generalizes $U_i(\bm{d})$ to include terms capturing link preferences. The result that follows, while closely related, is not a special case of his Theorem 1, which also follows the method of proof in mele2017structural. In our setting, agents take the network as given rather than optimizing over links, resulting in a different stationary distribution, and we also allow the network to be weighted. Perhaps due to this difference in setup, the connection to the Ising model was not previously noted.
propositionAs $t\rightarrow\infty$, ${\bf P}(\bm{D}_i^t=\bm{d} \mid \bm{X},\bm{A})$ converges to the unique stationary distribution given by (ref) with $h_i(c) = a_i$ and $J_{ij}(c) = A_{ij}\phi_{ij}$.

Supposing that $\bm{D}$ is a draw from the stationary distribution, the assignment mechanism is $\text{MTP}_2$ if $\phi_{ij} \geq 0$ for all $(i,j)$, which we can now interpret as a strategic complementarity condition. Like (ref), we require no restrictions on the magnitude of peer effects.

Saturation Exposures

We next turn to exposure contrasts common in the cluster-randomized trials (CRTs) literature. We suppose treatments are assigned according to a standard randomized saturation design.

myassump{CRT} For all $i\in\mathcal{N}_n$, $Y_i(\cdot) \perp\!\!\!\perp \bm{D}$. Let the clusters $\{C_j\}_{j=1}^m$ be a partition of $\mathcal{N}_n$ and the saturation levels $\{\tilde{S}_j\}_{j=1}^m$ be i.i.d.\ draws from a distribution supported on $\mathcal{P} = \{p_k\}_{k=1}^q \subseteq [0,1]$. For each $j$, $\{D_i\}_{i \in C_j} \stackrel{iid}\sim \text{Bernoulli}(p)$ conditional $\tilde{S}_j = p$.

That is, clusters are randomly assigned to saturation levels, and units within a cluster are assigned to treatment with probability equal to the saturation level. Let $S_i$ be the saturation level assigned to unit $i$'s cluster, so that $S_i = \tilde{S}_j$ if $i \in C_j$.

In this section, we depart from the setup of (ref) and define the exposure mapping as the tuple

equation*[equation* omitted — 36 chars of source]

Most of the CRT literature focuses on the following four exposure contrasts $\tau(t,t')$ hayes2017cluster.\footnote{Some papers use the number or share of treated units in $i$'s cluster in place of $S_i$ in the exposure mapping definition. (ref) in the appendix covers this case.}

enumerate• The “direct effect” sets $t=(1,p)$ and $t'=(0,p)$ for any $p \in \mathcal{P}$, meaning it conditions on the saturation level but varies own treatment assignment. Under Assumptions (ref) and (ref), this has a clear causal interpretation due to Bernoulli-randomization within cluster. This is not the case for the remaining estimands. • The “indirect effect” sets $t=(d,p)$ and $t'=(d,p')$ for any $d \in \{0,1\}$ and $p,p' \in \mathcal{P}$ with $p\geq p'$, thus varying the saturation level while conditioning on the treatment. • The “total effect” is the sum of the direct and indirect effects, which corresponds to setting $t = (1,p)$ and $t'= (0,p')$. • The “overall effect,” for lack of better notation, sets $t = (\emptyset,p)$ and $t' = (\emptyset,p')$ where we define the event $\{D_i = \emptyset\} \equiv \{D_i \in \{0,1\}\}$. In other words, the contrast varies the saturation level without conditioning on treatment assignment $D_i$.\footnote{The “group average effects” of hudgens2008toward are conceptually similar to these definitions. The main distinction is that our definition of the exposure contrast equally weights units whereas theirs equally weights clusters.}

In all cases, $\tau(t,t')$ is only a statistical comparison potentially subject to sign reversals of the sort in (ref). A common assumption in the CRT literature is {\em stratified interference}, meaning that the exposure mapping is structural basse2018analyzing,hudgens2008toward,vazquez2023identification. Then we can rewrite $Y_i(\bm{D})$ as $Y_i(T_i)$, and the four “effects” have transparent causal interpretations. However this is restrictive because it presumes units are exchangeable within cluster. In reality, units may respond differently to others depending on characteristics or if peer effects are mediated by a social network.

Perhaps for this reason, some papers do not maintain stratified interference hudgens2008toward,lee2024efficient,tchetgen2012causal, but the causal meaning of the estimands in this case has not been studied in the literature. Furthermore, virtually all references assume {\em partial interference}, that there is no interference across clusters, but cross-cluster interference is often a feature of CRTs for infectious diseases and large-scale social experiments egger2022general,leung2025cluster.

The next result shows that, for standard designs satisfying (ref), no restrictions on interference are required to ensure that the estimands can be represented as convex averages of unit-level effects. Let $C_{(i)}$ denote the cluster containing unit $i$, and for any $\bm{d} \in \{0,1\}^n$, let $\bm{d}_{(i)} = (d_j)_{j \in C_{(i)}}$ and $\bm{d}_{(-i)} = (d_j)_{j \in \mathcal{N}_n\backslash C_{(i)}}$. Finally let $p_{(i),s}(\cdot)$ denote the conditional distribution of $\bm{D}_{(i)} \mid T_i=s$ for $s \in \{t,t'\}$.

theoremUnder (ref), if $\tau(t,t')$ is the indirect, total, or overall effect with $p\geq p'$, then for all $i\in\mathcal{N}_n$ there exists a monotone coupling $\bm{D}_{(i),t}^* \stackrel{a.s.}\geq \bm{D}_{(i),t'}^*$ independent of $\bm{D}_{(-i)}$ with $\bm{D}_{(i),s}^* \sim p_{(i),s}(\cdot)$ for all $s \in \{t,t'\}$ such that \begin{equation*} \tau(t,t') = \frac{1}{n} \sum_{i=1}^n {\bf E}\big[ Y_i(\bm{D}_{(i),t}^*, \bm{D}_{(-i)}) - Y_i(\bm{D}_{(i),t'}^*, \bm{D}_{(-i)}) \big]. \end{equation*}

Because $p\geq p'$, an increase in the exposure $T_i$ from $t'$ to $t$ means an increase in the proportion treated, resulting in stochastically larger assignment vectors $\bm{D}_{(i)}$ in $i$'s cluster. The unit-level effects of interest therefore take the form $Y_i(\bm{d}_{(i)}, \bm{d}_{(-i)}) - Y_i(\bm{d}_{(i)}', \bm{d}_{(-i)})$ with $\bm{d}_{(i)} \geq \bm{d}_{(i)}'$, which are exactly those in the convex average.

$K$-Neighborhood Exposures

Suppose units are connected through a network $\bm{A}$, represented as an $n\times n$ binary matrix with $ij$th entry $A_{ij}$. Let $\mathcal{N}(i,K)$ denote unit $i$'s {\em $K$-neighborhood}, the subset of units at most path distance $K$ from $i$ in $\bm{A}$.\footnote{The path distance between two distinct units is the length of the shortest path between them if a path exists and infinite if not. The path distance between a unit and itself is zero, so $\mathcal{N}(i,0) = \{i\}$.} For any $\bm{d} \in \{0,1\}^n$, let $\bm{d}_{\mathcal{N}(i,K)} = (d_j\colon j \in \mathcal{N}(i,K))$ and $\bm{d}_{-\mathcal{N}(i,K)} = (d_j\colon j \in \mathcal{N}_n\backslash\mathcal{N}(i,K))$. It will often be convenient to partition $\bm{d}$ as $(\bm{d}_{\mathcal{N}(i,K)}, \bm{d}_{-\mathcal{N}(i,K)})$ and write $Y_i(\bm{d}_{\mathcal{N}(i,K)}, \bm{d}_{-\mathcal{N}(i,K)}) \equiv Y_i(\bm{d})$.

This section considers the setup of (ref), with the additional restriction that $f$ is a {\em $K$-neighborhood exposure mapping} in that it only depends on treatments assigned to the ego's $K$-neighborhood. Formally, $f(i,\bm{d}) = f(i,\bm{d}')$ for all $i$ and $\bm{d},\bm{d}' \in \{0,1\}^n$ such that $\bm{d}_{\mathcal{N}(i,K)} = \bm{d}_{\mathcal{N}(i,K)}'$. Abusing notation, we may abbreviate

equation*[equation* omitted — 67 chars of source]

The treated neighbor count in (ref) satisfies this restriction with $K=1$.

If $K$ is chosen large enough to encompass the entire network, this imposes no restrictions. However, the spirit of exposure mappings is to choose $K$ smaller than the typical distance between units to parsimoniously summarize $\bm{D}$. In this case, the assumption that the exposure mapping is structural typically rules out endogenous peer effects mediated by $\bm{A}$.

We consider the following class of {\em $t'$-degenerate} exposure mappings.

myassump{DEG} For any $i\in\mathcal{N}_n$ and $\bm{d}\in\{0,1\}^n$, $f(\bm{d}_{\mathcal{N}(i,K)})=t'$ implies $\bm{d}_{\mathcal{N}(i,K)} = \bm{\delta}_i$ for some $\bm{\delta}_i \in \{0,1\}^{\lvert\mathcal{N}(i,K)\rvert}$.

In other words, knowing $T_i=t'$ pins down the treatment subvector on $i$'s $K$-neighborhood.

example[Treated Neighbor Count] Consider the treated neighbor count from (ref) with $t = (d,\eta)$ and $t' = (d',\eta')$ for some $d,d' \in \{0,1\}$ and $\eta,\eta' \in \mathbb{N} \cup \{0\}$. Then the exposure is $t'$-degenerate if $\eta'=0$ since $T_i=t'$ means all neighbors are untreated. On the other hand, if $\eta' \in (0, \sum_{j=1}^n A_{ij})$, then $T_i=t'$ does not pin down which of $i$'s neighbors are treated, so $t'$-degeneracy does not hold. Thus the spirit of $t'$-degeneracy is that the exposure value $t'$ constitutes a “base case.”

While treated neighbor counts are covered by (ref), the result that follows imposes a different restriction on the assignment mechanism and allows for non-monotonic exposures, a leading case of which is the following.

example[Local Configuration] Let $\bm{A}_{\mathcal{N}(i,K)}$ be the subnetwork on $\mathcal{N}(i,K)$, that is $(A_{jk}\colon j,k \in \mathcal{N}(i,K))$. Restrict the population used in the exposure contrast to the subset of units $i$ such that $\bm{A}_{\mathcal{N}(i,K)} \cong \bm{a}$ for some network $\bm{a}$ where $\cong$ denotes graph isomorphism, and define \begin{equation*} f(\bm{d}_{\mathcal{N}(i,K)}) = \left\{ \begin{array}{ll} 1 & if (\bm{d}_{\mathcal{N}(i,K)}, \bm{A}_{\mathcal{N}(i,K)}) \cong (\bm{\delta},\bm{a}) \\ 0 & if (\bm{d}_{\mathcal{N}(i,K)}, \bm{A}_{\mathcal{N}(i,K)}) \cong (\bm{\delta}',\bm{a}) \end{array} \right. \end{equation*} which is $t'$-degenerate.\footnote{Define a {\em permutation} $\pi$ as a bijection on $\mathcal{N}_n$. Abusing notation, write $\pi(\bm{D}) = (D_{\pi(i)})_{i=1}^n$ and similarly $\pi(\bm{A}) = (A_{\pi(i)\pi(j)})_{i,j}$, which permutes the rows and columns of the matrix $\bm{A}$. If there exists a permutation $\pi$ such that $(\bm{D}_{\mathcal{N}(i,K)}, \bm{A}_{\mathcal{N}(i,K)}) = (\pi(\bm{\delta}),\pi(\bm{a}))$, then we write $(\bm{D}_{\mathcal{N}(i,K)}, \bm{A}_{\mathcal{N}(i,K)}) \cong (\bm{\delta},\bm{a})$.} Then $\tau(1,0)$ compares the subset of units with $K$-neighborhood subnetwork $\bm{a}$ and $K$-neighborhood treatment configuration $\bm{\delta}'$ vs.\ $\bm{\delta}$. This is essentially the estimand studied in \S4.2 of auerbach2023local, but whereas they assume these exposures are structural, we consider weaker restrictions on interference. Notice that if $\bm{\delta}$ and $\bm{\delta}'$ are not partially ordered, the exposure is not monotone and falls outside the scope of (ref).

Unrestricted Interference

Our first result imposes no restrictions on interference but requires the assignment mechanism to satisfy the following.

myassump{$K$-CI} For any $i\in\mathcal{N}_n$, $\bm{D}_{\mathcal{N}(i,K)} \perp\!\!\!\perp \bm{D}_{-\mathcal{N}(i,K)} \mid \mathcal{C}_i$.

This states that each unit's $K$-neighborhood treatment assignment vector is independent of the remaining assignments conditional on the controls. It holds if $\{D_i\}_{i=1}^n$ is independently distributed conditional on $\mathcal{C}_i$ for any $i$, which is the case considered in (ref). It also holds in CRTs for which treatments are independent across clusters.

Clustering corresponds to the special case in which $\bm{A}$ is block diagonal, so $\mathcal{N}(i,1)$ is the set of units in $i$'s cluster. Unlike previous theorems, we can allow for arbitrary correlation between assignments within cluster. For instance, each cluster can be a network, and within each network, one can implement the correlated designs discussed in (ref). Unlike (ref), we consider different exposure contrasts such as treated neighbor counts, which have also been used in the CRT literature miguel2004worms,vazquez2023identification.

theoremUnder Assumptions (ref) and (ref), $\tau(t,t') = \tau^*(t,t') + \mathcal{B}$ where \begin{align*} &\tau^*(t,t') = \frac{1}{n} \sum_{i=1}^n {\bf E}\big[Y_i(\bm{D}) - Y_i(\bm{\delta}_i, \bm{D}_{-\mathcal{N}(i,K)}) \mid T_i=t, \mathcal{C}_i \big] \quadand \\ &\mathcal{B} = \frac{1}{n} \sum_{i=1}^n \left( {\bf E}\big[Y_i(\bm{\delta}_i,\bm{D}_{-\mathcal{N}(i,K)}) \mid T_i=t, \mathcal{C}_i\big] - {\bf E}\big[Y_i(\bm{\delta}_i,\bm{D}_{-\mathcal{N}(i,K)}) \mid T_i=t', \mathcal{C}_i\big] \right). \end{align*} Moreover, under (ref), $\mathcal{B}=0$.

The causal estimand $\tau^*(t,t')$ is a convex average of unit-level effects of the form $Y_i(\bm{d}_{\mathcal{N}(i,K)}, \bm{d}_{-\mathcal{N}(i,K)}) - Y_i(\bm{\delta}_i, \bm{d}_{-\mathcal{N}(i,K)})$ which fix treatments outside the $K$-neighborhood while varying $K$-neighborhood treatments subject to the exposure mapping constraint $f(\bm{d}_{\mathcal{N}(i,K)})=t$.

The decomposition $\tau^*(t,t') + \mathcal{B}$ has an omitted variable bias interpretation. The first term $\tau^*(t,t')$ is the effect of variation in the “main regressor” $\bm{D}_{\mathcal{N}(i,K)}$ induced by the exposure mapping. The bias $\mathcal{B}$ is nonzero if the $K$-neighborhood exposure is correlated with the “omitted variable” $\bm{D}_{-\mathcal{N}(i,K)}$. Unconfoundedness alone provides no control over $\mathcal{B}$. Theorem 4 of sobel2006randomized and Theorem 1 of vazquez2023identification provide similar decompositions for the case of $T_i=D_i$.

Unrestricted Assignment Mechanism

The remaining results impose no restriction on the assignment mechanism other than unconfoundedness. The first result requires higher-order spillovers, meaning those induced by units beyond the ego's $K$-neighborhood, to be uniformly smaller than $K$-neighborhood spillovers. We refer to this as “neighborhood-centric interference.”

myassump{$K$-NCI} $\Delta_K > \Psi_K$ where \begin{align*} \Delta_K &= \min\big\{ \lvertY_i(\bm{d}_{\mathcal{N}(i,K)}, \bm{d}_{-\mathcal{N}(i,K)}”) - Y_i(\bm{d}_{\mathcal{N}(i,K)}', \bm{d}_{-\mathcal{N}(i,K)}”)\rvert \big\}, \\ \Psi_K &= \max\big\{ \lvertY_i(\bm{d}_{\mathcal{N}(i,K)}”, \bm{d}_{-\mathcal{N}(i,K)}) - Y_i(\bm{d}_{\mathcal{N}(i,K)}”, \bm{d}_{-\mathcal{N}(i,K)}')\rvert \big\}, \end{align*} and the max and min are taken over $i \in \mathcal{N}_n$ and $\bm{d},\bm{d}',\bm{d}'' \in \{0,1\}^n$.

The term $\Delta_K$ is the smallest $K$-neighborhood spillover effect across all units, while $\Psi_K$ is the largest higher-order spillover effect from beyond the $K$-neighborhood.

exampleIn the case of $K=0$, $\Delta_K$ is the smallest direct effect of the treatment over all units, so $K$-NCI holds if direct effects uniformly dominate spillover effects in magnitude. This is relevant for settings in which direct effects are typically larger than spillover effects, for instance online experiments viviano2023causal,yuan2021causal. In the context of vaccines, the spillover effect from reduced community transmission is often smaller than the effect of being directly vaccinated.
exampleSeveral papers assume that potential outcomes only depend on treatments within a $K$-neighborhood but without imposing a particular exposure mapping: \begin{equation} Y_i(\bm{d}) = Y_i(\bm{d}') \quadfor all\quad \bm{d},\bm{d}'\in\{0,1\}^n \quadsuch that\quad \bm{d}_{\mathcal{N}(i,K)} = \bm{d}_{\mathcal{N}(i,K)}' \end{equation} ugander2013graph,viviano2023causal. This implies $K$-NCI since $\Psi_K=0$. If $\bm{A}$ is block-diagonal, with each block representing a cluster, then partial interference corresponds to (ref) for any $K\geq 1$. $K$-NCI allows for violations of partial interference, so long as cross-cluster interference $\Psi_K$ is uniformly dominated by within-cluster interference $\Delta_K$.
theoremUnder Assumptions (ref), (ref), and (ref), $\lvert\tau^*(t,t')\rvert > \lvert\mathcal{B}\rvert$.

Recall from (ref) that $\tau^*(t,t')$ is a convex average of unit-level effects of the form $Y_i(\bm{d}_{\mathcal{N}(i,K)}, \bm{d}_{-\mathcal{N}(i,K)}) - Y_i(\bm{\delta}_i, \bm{d}_{-\mathcal{N}(i,K)})$, so if they possess the same sign for all $i$ and $\bm{d}$, so does $\tau^*(t,t')$. Since $\mathcal{B}$ is smaller in magnitude, $\tau(t,t')$ maintains the sign of $\tau^*(t,t')$, so reversals do not occur. sobel2006randomized previously noted in the context of a particular class of experiments that when $T_i=D_i$, $\tau(t,t')$ is not subject to sign reversals when direct effects dominate spillover effects. (ref) generalizes this to $t'$-degenerate exposure mappings and arbitrary designs.

The result is motivated by the omitted variable bias interpretation of (ref). By definition, $K$-neighborhood exposures directly manipulate $\bm{D}_{\mathcal{N}(i,K)}$, but since treatments are correlated, they also induce variation in $\bm{D}_{-\mathcal{N}(i,K)}$, which generates bias $\mathcal{B}$. The bias involves spillovers beyond the $K$-neighborhood, which under $K$-NCI, is smaller in magnitude than the causal estimand $\tau^*(t,t')$, so the sign of the latter dominates.\footnote{I thank a referee for comments that inspired this result.}

Our last result considers $K$-neighborhood exposure mappings with $K$ chosen relatively large, as may be the case in (ref). We require interference between units to decay with their distance. The main idea is that larger $K$ implies that the exposure mapping captures variation in the treatment subvector within a larger radius. Since interference beyond this radius is relatively small, the exposure is “approximately” structural, so $\tau(t,t')$ should approximate a quantity that has a causal interpretation. This formalizes some of the discussion in \S4 of auerbach2024discussion.

We consider the leung2024graph model (ref). For any $S \subseteq \mathcal{N}_n$, let $\bm{X}_S = (X_i)_{i\in S}$, and similarly define $\bm{\varepsilon}_S$ and $\bm{\nu}_S$. The following “approximate neighborhood interference” condition due to leung2024graph formalizes the idea of interference decaying to zero as path distance diverges.

myassump{ANI} There exists $\gamma \colon \mathbb{R}_+\rightarrow \mathbb{R}_+$ such that $\gamma(s) \stackrel{s\rightarrow\infty}\longrightarrow 0$ and \begin{multline*} \max_{i\in\mathcal{N}_n} {\bf E}\big[\lvert g_n(i, \bm{D}, \bm{X}, \bm{A}, \bm{\varepsilon}) \\ - g_{\lvert\mathcal{N}(i,s)\rvert}(i, \bm{D}_{\mathcal{N}(i,s)}, \bm{X}_{\mathcal{N}(i,s)}, \bm{A}_{\mathcal{N}(i,s)}, \bm{\varepsilon}_{\mathcal{N}(i,s)}) \rvert \mid \bm{D}, \bm{X}, \bm{A} \big] \leq \gamma(s). \end{multline*}

To understand the inequality, first recall from (ref) that $g_n(i, \bm{D}, \bm{X}, \bm{A}, \bm{\varepsilon})$ is $i$'s observed outcome $Y_i$. We interpret $g_{\lvert\mathcal{N}(i,s)\rvert}(i,\dots)$ as $i$'s outcome under a counterfactual “$s$-neighborhood model” in which the primitives and treatments are fixed at their realizations, units external to the $s$-neighborhood are excluded from the model, and the remaining units interact according to the reduced-form model $g_{\lvert\mathcal{N}(i,s)\rvert}(\cdot)$. ANI bounds the difference between $i$'s realized outcome and counterfactual $s$-neighborhood outcome by $\gamma(s)$, which is required to decay with the radius $s$. If the rate of decay is faster, then $Y_i$ is well-approximated by a model with only units in $\mathcal{N}(i,s)$ for smaller $s$, formalizing the idea that units distant from $i$ interfere less with $i$.

theoremConsider model (ref) with controls (ref). Under Assumptions (ref), (ref), and (ref), $\lvert\tau(t,t') - \tau^*(t,t')\rvert \leq \gamma(K) \rightarrow 0$ as $K\rightarrow\infty$.

The proof uses (ref) to bound the bias $\mathcal{B}$ in (ref). Under (ref), ANI holds with $\gamma(s)=0$ for all $s\geq K$, in which case $\tau(t,t')$ has an exact causal interpretation for $K$ chosen sufficiently large. leung2022causal shows that ANI can be satisfied by well-known models of social interactions with exponentially decaying $\gamma(s)$. In this case, choosing $K$ to be logarithmic in $n$ can ensure that the bias is order $n^{-c}$ for some $c>0$.

Conclusion

In settings with interference, researchers often report exposure contrasts to summarize treatment and spillover effects. A common example is to regress an outcome on own treatment assignment and the number or share of treated neighbors. Researchers typically interpret the respective coefficients as “direct” and “spillover” effects, but we show that this interpretation is not generally valid. The exposure contrast can have the opposite sign of the unit-level effects of interest even if treatment assignment is unconfounded.

Eliminating sign reversals requires restricting either interference or correlation in treatment assignments across units. The literature typically assumes exposure mappings are structural in that they entirely mediate interference. In our view, exposure mappings such as the number of treated neighbors function more as statistics of convenience and are unlikely to be structural. We propose alternative assumptions that are substantially weaker and rule out sign reversals.

Our first result considers assignments satisfying a certain positive association condition. We show that this is satisfied by stratified experiments and selection models with peer effects. Our second result concerns cluster-randomized trials, and we show that standard estimands can be written as convex averages of unit-level effects without imposing any restrictions on spillovers within or across clusters. Finally, we consider arbitrary unconfounded assignment mechanisms and show that sign reversals can be avoided under different restrictions on interference that allow for endogenous peer effects.