Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Average Direct and Indirect Causal Effects under Interference
\ifbio
\jname{Biometrika}
\jyear
\jvol{forthcoming}
\jnum
\affil{Dept. of Management Science and Engineering,
475 Via Ortega, Stanford, CA-94305, U.S.A. \email{[email removed]}}
\affil{Department of Statistics, 390 Jane Stanford Way, Stanford, CA-94305, U.S.A.
\email{[email removed]}}
\affil{Stanford Graduate School of Business, 655 Knight Way, Stanford, CA-94305, U.S.A.
\email{[email removed]}}
\else
\fi
abstractWe propose a definition for the average indirect effect of a binary treatment in the potential outcomes model for causal inference under cross-unit interference. Our definition is analogous to the standard definition of the average direct effect, and can be expressed without needing to compare outcomes across multiple randomized experiments. We show that the proposed indirect effect satisfies a decomposition theorem whereby, in a Bernoulli trial, the sum of the average direct and indirect effects always corresponds to the effect of a policy intervention that infinitesimally increases treatment probabilities. We also consider a number of parametric models for interference, and find that our non-parametric indirect effect remains a natural estimand when re-expressed in the context of these models.
\ifbio
keywordsCausal inference; Interference; Potential outcome; Randomized trial.
\fi
Introduction
The classical way of analyzing randomized trials, following neyman1923applications
and rubin1974estimating, is centered around the average treatment effect as defined using potential outcomes.
Given a sample of $i = 1, \, \ldots, \, n$ units used to study the effect of a binary treatment $W_i \in \left\{0, \, 1\right\}$,
we posit potential outcomes $Y_i(0), \, Y_i(1) \in \mathbb{R}$ corresponding to the outcome we would have measured
had we assigned the $i$-th unit to control or treatment respectively, i.e., we observe $Y_i = Y_i(W_i)$.
We then proceed by arguing that the sample average treatment effect
equation[equation omitted — 99 chars of source]
admits a simple unbiased estimator under random assignment of treatment.
One limitation of this classical approach is that it rules out interference, and instead introduces an assumption that
the observed outcome for any given unit does not depend on the treatments assigned to other units, i.e.,
$Y_i$ is not affected by $W_j$ for any $j \neq i$ halloran1995causal. However, in a wide variety of applied settings,
such interference effects not only exist, but are often of considerable scientific interest
bakshy2012role,bond201261,cai2015social,miguel2004worms,rogers2018reducing,sacerdote2001peer.
For example, in an education setting, it may be of interest to understand how a didactic innovation affects
not only certain targeted students, but also their peers.
This has led to a recent surge of interest in methods for studying randomized trials under interference
aronow2017estimating,eckles2017design,hudgens2008toward,leung2020treatment,
li2020random,manski2013identification,savje2017average,tchetgen2012causal.
A major difficulty in working under interference is that we no longer have a single obvious average
effect parameter to target as in (ref). In the general setting, each unit now has $2^n$ potential outcomes
corresponding to every possible treatment combination assigned to the $n$ units, and these can be used to formulate
effectively innumerable possible treatment effects that can arise from different assignment patterns.
As discussed further in Section (ref) below, the existing literature has mostly side-stepped this issue by
framing the estimand in terms of specific policy interventions. However, this paradigm does not provide researchers
with simple, non-parametric and agnostic average causal estimands that can be studied without spelling out a specific policy intervention of interest.
In this paper, we study a pair of averaging causal estimands, the average direct and indirect effects,
that are valid under interference yet, unlike existing targets, can be defined and estimated using a single experiment
and do not need to be defined in terms of hypothetical policy interventions. Qualitatively, the average direct effect measures the extent to which, in a given experiment and on average, the outcome $Y_i$ of a unit is affected by its own
treatment $W_i$; meanwhile, the average indirect effect measures the responsiveness of $Y_i$ to treatments $W_j$ given to other units $j \neq i$.
The average direct effect we consider is standard and has recently been discussed by a number of authors including vanderweele2011effect and \citet*{savje2017average}. Our definition
of the average indirect effect is to the best of our knowledge new, and is the main contribution of this paper.
We follow this definition with a number of results to validate it. In particular, we prove a universal decomposition theorem
whereby, in a Bernoulli trial, the sum of the average direct and indirect effects can always be interpreted as the total effect of an intuitive
policy intervention. We also interpret these estimands in the context of a number of parametric models for interference
considered by practitioners.
Treatment Effects under Interference
We study different experimental designs using the potential outcomes model.
The main difference between a setting with interference and the standard Neyman-Rubin model is that
potential outcomes for the $i$-th unit may also depend on the intervention given to the $j$-th unit with $j \neq i$
aronow2017estimating,hudgens2008toward.
For convenience, we use short-hand $Y_i(w_j = x; \, W_{-j})$ to denote the potential outcome we
would observe for the $i$-th unit if we were to assign the $j$-th unit to treatment status $x \in \left\{0, \, 1\right\}$, and
otherwise maintained all but the $j$-th unit at their realized treatments
$\smash{W_{-j} \in \left\{0, \, 1\right\}^{n-1}}$.
Expectations $E$ are over the treatment assignment
only; potential outcomes are held fixed.
assumptionFor units $i = 1, \, \ldots, \, n$, there are potential outcomes \smash{$Y_i(w) \in \mathbb{R}$},
\smash{$w \in \left\{0, \, 1\right\}^n$} such that, given a treatment vector
\smash{$W \in \left\{0, \, 1\right\}^n$}, we observe outcomes \smash{$Y_i = Y_i(W)$}.
definitionUnder Assumption (ref), the average direct effect of a binary treatment
is
\begin{equation}
\tau_{\operatorname{ADE}} = \frac{1}{n} \sum_{i=1}^n E\left\{Y_i(w_i = 1; W_{-i}) - Y_i(w_i = 0; W_{-i})\right\}.
\end{equation}
definitionUnder Assumption (ref), the average indirect effect of a binary treatment
is
\begin{equation}
\tau_{\operatorname{AIE}} = \frac{1}{n}\sum_{i = 1}^n \sum_{j \neq i} E \left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}.
\end{equation}
The definition of the direct effect $\tau_{\operatorname{ADE}}$ is standard. It follows from averaging $Y_i(w_i = 1; W_{-i}) - Y_i(w_i = 0; W_{-i})$, referred to as the direct causal effect by halloran1995causal.
savje2017average provide a recent in-depth discussion on this estimand.
This estimand measures the average effect of an intervention $W_i$
on the unit being intervened on---while marginalizing over the rest of the treatment assignments.
In a study without interference, $\tau_{\operatorname{ADE}}$ matches the usual average treatment effect (ref).
Meanwhile, our definition of the indirect effect is an immediate formal generalization of $\tau_{\operatorname{ADE}}$ to cross-unit
treatment effects. It measures the average effect of an intervention $W_i$ on all units except the one being
intervened on, again marginalizing over the rest of the process. More precisely, the term $E\left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}$ is the effect of changing unit $i$'s treatment on the outcome of unit $j$. Thus the sum, $\sum_{j \neq i} E\left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}$, would correspond to the aggregate effect of unit $i$'s treatment on all other units. Then the defined average indirect effect $\tau_{\operatorname{AIE}}$ corresponds to the average of the effects of units' treatments on other units.
The definition of $\tau_{\operatorname{AIE}}$ formally mirrors that of $\tau_{\operatorname{ADE}}$, and in the no-interference case we clearly have
$\tau_{\operatorname{AIE}} = 0$.
As a first step towards validating the definition of $\tau_{\operatorname{AIE}}$ under non-trivial interference,
we consider the average overall effect induced by adding
$\tau_{\operatorname{ADE}}$ and $\tau_{\operatorname{AIE}}$, which aggregates the marginalized effect of all treatments on all outcomes.
We then prove that, in a Bernoulli design, this matches the policy effect of infinitesimally increasing
each unit's treatment probability.
We use the term Bernoulli design to refer to an experiment where there is
a deterministic vector $\pi \in (0, \, 1)^n$ such that the treatments $W_i$
are generated as
$W_i \, \sim \, \operatorname{Bernoulli}(\pi_i)$ for all $i = 1, \, \ldots, \, n$,
independently of each other and of the potential outcomes $\left\{Y_i(w)\right\}$.
For a Bernoulli design with treatment probabilities $\pi \in [0, \, 1]^n$, we write
$E_{\pi}\left(\cdot\right)$ for expectations over the random treatment assignment, and write $\tau_{\operatorname{ADE}}(\pi)$, $\tau_{\operatorname{AIE}}(\pi)$ and
$\tau_{\operatorname{AOE}}(\pi)$ for the corresponding direct, indirect and overall effects.
definitionUnder Assumption (ref), the average overall effect of a binary treatment is
\begin{align}
\tau_{\operatorname{AOE}} = \tau_{\operatorname{ADE}} + \tau_{\operatorname{AIE}} =\frac{1}{n}\sum_{i = 1}^n \sum_{j=1}^n E\left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}.
\end{align}
definitionUnder Assumption (ref) and in a Bernoulli design, the infinitesimal policy effect is
\begin{equation}
\tau_{\operatorname{INF}}(\pi) = \mathbf{1} \cdot \nabla_{\pi} E_{\pi}\left(\frac{1}{n}\sum_{i = 1}^n Y_i\right) = \sum_{k=1}^n \frac{\partial}{\partial \pi_k} E_{\pi}\left(\frac{1}{n}\sum_{i = 1}^n Y_i\right).
\end{equation}
theoremUnder Assumption (ref) and in a Bernoulli design,
\smash{$\tau_{\operatorname{AOE}}(\pi) = \tau_{\operatorname{INF}}(\pi)$}.
By connecting our abstract notions of direct, indirect and overall effects to the effect of a concrete policy intervention,
Theorem (ref) provides an alternative lens on our definition of the indirect effect. Suppose, for example,
that a researcher knew they wanted to study nudge interventions, the total effect of which is $\tau_{\operatorname{INF}}(\pi)$,
and was also committed to the standard definition of the average direct effect given in Definition (ref).
Then, it would be natural to define an indirect effect as $\tau_{\operatorname{INF}}(\pi) - \tau_{\operatorname{ADE}}(\pi)$, i.e., to characterize
as indirect effect any effect of the nudge intervention that is not captured by the direct effect; this is, for example,
the approach implicitly taken in heckman1998general. From this perspective, Theorem (ref)
could be seen as showing that these two possible definitions of the indirect effect in fact match, i.e., that
$\tau_{\operatorname{AIE}}(\pi) = \tau_{\operatorname{INF}}(\pi) - \tau_{\operatorname{ADE}}(\pi)$. We emphasize that Theorem (ref) is a direct
consequence of Bernoulli randomization, and holds conditionally on any realization of the potential outcomes $\left\{Y_i(w)\right\}$.
We refer to $\tau_{\operatorname{INF}}(\pi)$ as a policy effect because,
in the ideal situation when one has access to observed outcomes $Y_i$ for different randomization probabilities $\pi$,
$\tau_{\operatorname{INF}}(\pi)$ is a quantity that could be measured by
averaging observed outcomes $Y_i$ for different $\pi$.
If treatment assignment probabilities are constant, i.e., there is a $\pi_0 \in (0, \, 1)$ such that $\pi_i = \pi_0$ for all $i = 1, \, \ldots, \, n$, then $\tau_{\operatorname{INF}}(\pi)$ takes on a particularly simple form,
\smash{$\tau_{\operatorname{INF}}(\pi) = {d}/{d\pi_0} \, E_{\pi_0}\left(n^{-1} \sum_{i=1}^n Y_i\right)$}.
Infinitesimal policy effects as defined above
are prevalent in social sciences
due to their ease of interpretation and desirable analytic properties; see, e.g., chetty2009sufficient, carneiro2010evaluating,
and references therein. wager2021experimenting discuss welfare implications for a social planner who
uses analogous infinitesimal policy effects to optimize a system via gradient-based methods.
remarkAccurate estimation of $\tau_{\operatorname{ADE}}$ and $\tau_{\operatorname{AIE}}$ is an interesting
question beyond the scope of this paper. In Appendix
(ref), we show that both $\tau_{\operatorname{ADE}}$ and $\tau_{\operatorname{AIE}}$
always admit unbiased estimators in Bernoulli experiments.
However, the precision of these unbiased estimators
will depend on the interference pattern. li2020random
study estimators for $\tau_{\operatorname{ADE}}$ and $\tau_{\operatorname{AIE}}$ in the context of a
random graph model for interference, including improvements in settings where simple
unbiased estimators are unstable.
Alternative Definitions and Related Work
There are also a number of other average causal effect estimands that have recently been discussed in the literature.
In the case of direct effects, the main alternative to Definition (ref) is the following proposal from
hudgens2008toward that relies on conditional expectations,
equation*[equation* omitted — 193 chars of source]
In a Bernoulli design, $\tau_{\operatorname{HH}, \operatorname{DE}} = \tau_{\operatorname{ADE}}$. However, in other designs, e.g., completely
randomized designs or stratified designs, these two estimands do not match.
As discussed in vanderweele2011effect and savje2017average, a major drawback of the definition $\tau_{\operatorname{HH}, \operatorname{DE}}$
is that it conflates the effect of setting $w_i = x$ on the $i$-th unit's outcome, and the effect of setting $w_i = x$ on
the distribution of $W_{-i}$. In particular, in completely randomized experiments, it's possible to have
$\tau_{\operatorname{HH}, \operatorname{DE}} \neq 0$ even when $Y_i(w_i = 1, \, w_{-i}) = Y_i(w_i = 0, \, w_{-i})$ for
all units and all possible treatment assignments. In contrast, $\tau_{\operatorname{ADE}}$ as defined in Definition (ref)
has a robust causal interpretation as a direct effect.
Meanwhile, as discussed in the introduction, most available notions of indirect effects rely on explicit
comparisons between two overall treatment assignment strategies.
For example, hudgens2008toward and vanderweele2011effect propose a number of
indirect effect estimands that, in the case of comparing two Bernoulli trials with randomization probabilities $\pi$ and $\pi'$, reduce to
equation*[equation* omitted — 204 chars of source]
In the case of non-Bernoulli trials, there are a number of subtleties analogous to the ones noted above; see
vanderweele2011effect for an in-depth discussion. \smash{$\tau_{\operatorname{IE}}(\pi, \, \pi')$}
is an interesting quantity to consider if we can run many independent experiments that test different overall treatment level but,
unlike $\tau_{\operatorname{AIE}}$, does not enable a researcher to describe indirect effects in a single randomized study.
Another popular approach to capturing indirect effects is via the exposure mapping approach developed in aronow2017estimating.
The main idea is to assume existence of functions $h_i : \left\{0, \, 1\right\}^n \rightarrow \left\{1, \, \ldots, \, K\right\}$ such that potential
outcomes $Y_i(w)$ only depend on $w$ via the compressed representation $h_i(w)$, i.e., $Y_i(w) = Y_i(w')$ whenever
$h_i(w) = h_i(w')$; see also karwa2018systematic, leung2020treatment
and savje2021causal for further discussions and extensions.
One can then consider estimators of
averages of potential outcome types and define treatment effects in terms of their contrasts,
equation*[equation* omitted — 185 chars of source]
Definitions of this type are again conceptually attractive and sometimes enable us to very clearly express
the answer to a natural policy question; see, e.g., basse2019randomization. However, they again require the
analyst to consider specific policy interventions to be able to even talk about indirect effects, and can also be unwieldy
to use as the number of possible exposure types $K$ gets large.
Closest to the definition of $\tau_{\operatorname{AIE}}$ is the average marginalized response of aronow2020design.
They consider a setting where treatments are assigned to points in a geographic space, and seek to estimate
the average effect of treatment at an intervention point outcomes at points that are a distance $d$ away,
equation*[equation* omitted — 277 chars of source]
where $\Delta(i, \, j)$ measures the distance between points $i$ and $j$.
This circle average bears resemblance to Definition (ref) in the sense that both of them are marginalized
over variation in $W_{-i}$ holding the treatment $W_i$ fixed. However, one key difference is the normalization factor
$|\mathcal{S}_i(d)|^{-1}$ used in $\tau_{\operatorname{AMR}}$. Adding similar normalization to $\tau_{\operatorname{AIE}}$ would invalidate Theorem (ref).
remarkOur definition of $\tau_{\operatorname{AIE}}(\pi)$ is normalized by $n$, not
by the total number of summands $n(n-1)$, and one can ask whether this scaling is always the most natural one.
We argue below that, in a number of popular models,
$\tau_{\operatorname{AIE}}(\pi)$ coincides with interesting and interpretable quantities and converges as the number of units $n$ goes to infinity.
In other models, however, our $1/n$ scaling may not be the best choice. For example,
if we have a data-generating distribution with $Y_i(w) = (\sum_{j = 1}^n w_j - n \pi) / \sqrt{n \pi (1 - \pi)}$,
the observed outcomes $Y_i$ will have a standard normal marginal distribution,
but $\tau_{\operatorname{AIE}} = \sqrt{n}$ diverges. One should also note, however, that in this example $1/n \sum_{i = 1}^n Y_i$ does not concentrate.
Models for Interference
Our discussion so far has focused on an abstract specification where direct and indirect effects are defined
via various marginalized contrasts between potential outcomes. Much of the existing applied work on
causal inference under interference, however, has focused on simpler parametric specifications that, e.g.,
connect outcomes to treatments via a linear model. The purpose of this section is to examine our abstract,
non-parametric definition of the indirect effect given in Definition (ref), and to confirm that it still
corresponds to an estimand one would want to interpret as an indirect effect once we restrict our attention
to simpler parametric models. Below, we do so in 3 examples. munro2021treatment provide an extended
study of $\tau_{\operatorname{ADE}}$ and $\tau_{\operatorname{AIE}}$ in a marketplace model where interference arises via equilibrium price formation.
The claimed expressions for $\tau_{\operatorname{ADE}}$ and $\tau_{\operatorname{AIE}}$ are derived in Appendix (ref).
exampleIn studying the spillover effects of insurance training sessions on insurance purchase, cai2015social
work in terms of a network model: There is an edge matrix $E_{ij} \in \left\{0, \,1\right\}$, such that $W_j$ can only
affect $Y_i$ if the corresponding units are connected by an edge, i.e., if $E_{ij} = 1$. They then consider a
linear-in-means model parametrized in terms of this network. For our purpose,
we focus on a simple variant of the model of cai2015social considered in leung2020treatment,
where only the effects of ego's treatment and the proportion of treated neighbors are considered as covariates.
This results in a linear model induced by the structural equation
\begin{equation}
Y_i = \beta_1+\beta_2 W_i +\beta_3 \frac{\sum_{j\ne i} E_{ij}W_j}{\sum_{j\ne i} E_{ij}} + \varepsilon_i, \ \ E_\left(\varepsilon_i \mid W\right) = 0.
\end{equation}
In words, the probability of insurance purchase is modeled as a linear function of whether the farmer attends the insurance training sessions, and the proportion of friends who attend the session. The relation (ref) should be taken as a structural model, meaning that we can generate potential outcomes $Y_i(w)$ by plugging candidate assignment vectors $w$ into (ref), i.e., \smash{$Y_i(w) = \beta_1 + \beta_2 w_i + \beta_3 \sum_{j \neq i} E_{ij} w_j / \sum_{j \neq i} E_{ij} + \varepsilon_i$}, for all $w \in \left\{0, \, 1\right\}^n$.
Under this model, it can be shown that under Assumption (ref),
$\tau_{\operatorname{ADE}} = \beta_2$, and $\tau_{\operatorname{AIE}}=\beta_3$,
i.e., the estimands from Definitions (ref) and (ref)
map exactly to the parameters in model (ref) regardless of the experimental design.
exampleThe model in Example (ref) assumes the the $i$-th unit responds in the same way to treatment assigned to any of
its neighbors. However, this restriction may be implausible in many areas; for example, in social networks, there is evidence
that some ties are stronger than others, and that peer effects are larger along strong ties bakshy2012role.
A natural generalization of Example (ref) that allows for variable strength ties is to consider a saturated structural linear model
\begin{equation}
Y_i = \alpha_i + \beta_iW_i + \sum_{j\ne i}\nu_{ij}W_j + \varepsilon_i, \ \ E_\left(\varepsilon_i \mid W\right) = 0,
\end{equation}
which allows for both unit-specific direct and indirect effects. Here the individual parameters in this model are
not identifiable; however, under (ref),
$\tau_{\operatorname{ADE}}=\frac{1}{n}\sum_{i=1}^n \beta_i$, and
$\tau_{\operatorname{AIE}}=\frac{1}{n}\sum_{i=1}^n\sum_{j\ne i} \nu_{ij}$,
i.e., our estimands can be understood as averages of the unit-level parameters, again regardless of the design.
exampleIn studying the effect of persuasion campaigns or other types of messaging, one may assume
that people respond most strongly if they get a communication directly addressed to them, but can also
respond if a member of their neighborhood or their household gets a communication. This assumption can
be formalized in terms of the following model: Each unit has 4 potential outcomes defined as
\begin{equation}
Y_i = \begin{cases}
Y_i(treated & exposed) & if $W_i = 1$ and $i$ has a treated neighbor, \\
Y_i(treated) & if $W_i = 1$ but $i$ has no treated neighbors, \\
Y_i(exposed) & if $W_i = 0$ but $i$ has a treated neighbor, \\
Y_i(\text{none}) & \text{if $W_i = 0$ and $i$ has no treated neighbors}.
\end{cases}
\end{equation}
Models of this type are considered by sinclair2012detecting for studying voter mobilization, and by
basse2018analyzing and basse2019randomization for studying anti-absenteeism interventions.
Natural treatment effect parameters to consider following (ref) include the average self-treatment
and spillover effects
\begin{alignat*}{2}
\tau_{\text{SELF},1} &= \frac{1}{n} \sum_{i = 1}^n \left\{Y_i(\text{treated & exposed}) - Y_i(\text{exposed})\right\}, \quad
\tau_{\text{SELF},0} &&= \frac{1}{n} \sum_{i = 1}^n \left\{Y_i(\text{treated}) - Y_i(\text{none})\right\},\nonumber\\
\tau_{\text{SPILL},1} &= \frac{1}{n} \sum_{i = 1}^n \left\{Y_i(\text{treated & exposed}) - Y_i(\text{treated})\right\}, \quad
\tau_{\text{SPILL},0} &&= \frac{1}{n} \sum_{i = 1}^n \left\{Y_i(\text{exposed}) - Y_i(\text{none})\right\}.
\end{alignat*}
Unlike the previous two examples, the connection between $\tau_{\operatorname{ADE}}$ and $\tau_{\operatorname{AIE}}$ to $\tau_{\text{SELF}}$ and $\tau_{\text{SPILL}}$ differs substantially across experimental designs, especially when the design introduces correlation between the units. For the purpose of illustration, we study it in a multi-stage completely randomized design considered in previous works sinclair2012detecting, basse2018analyzing, basse2019randomization. In particular, we focus on the case where there are in total $n/m$ clusters of size $m$. On the first stage, $\rho\cdot n/m$ clusters are assigned to treatment, and $(1-\rho)\cdot n/m$ clusters are assigned to control; on the second stage, a single unit in each treated cluster is randomly chosen to be treated, and all the other units are assigned to control. We can then calculate the marginal distribution of $W_{-i}$ and obtain
\begin{equation}
\begin{split}
&\tau_{\operatorname{ADE}} = \left(\rho-\frac{\rho}{m}\right)\tau_{\text{SELF},1}+\left(1-\rho+\frac{\rho}{m}\right)\tau_{\text{SELF},0},\\
&\tau_{\operatorname{AIE}} = (m-1)\left\{\frac{\rho}{m}\tau_{\text{SPILL},1}+\left(1-\rho+\frac{\rho}{m}\right)\tau_{\text{SPILL},0} \right\}.
\end{split}
\end{equation}
Therefore, our estimands can be regarded as weighted averages of the treatment effect parameters. Moreover, when the cluster size $m$ is large, $\tau_{\operatorname{ADE}}$ is approximately the average of $\tau_{\text{SELF},1}$ and $\tau_{\text{SELF},0}$, weighted by the assignment probability during the first stage, while $\tau_{\operatorname{ADE}}$ is approximately $\tau_{\text{SPILL},0}$ times the probability of being assigned to the control group during the first stage, and the factor $m-1$ simply
accounts for the fact that, any treatment will spread spillover effects to $m-1$ neighbors.
Discussion
There are many treatment effect estimation problems in which interference is present and needs to be accounted for, and indirect effects are of considerable scientific interest.
We have proposed an estimand, $\tau_{\operatorname{AIE}}$, that quantifies the typical effect of treating one unit on the outcomes of all other units,
and provides a natural counterpart to the average direct effect. Consider,
for example, an experiment seeking to measure the effect of a housing voucher on homeownership. Here,
$\tau_{\operatorname{ADE}}$ measures the marginal benefit of receiving such a voucher oneself, whereas $\tau_{\operatorname{AIE}}$
analogously measures the extent to which one person receiving a voucher crowds out other buyers.
Theorem (ref) validates this perspective, by showing that the sum of $\tau_{\operatorname{ADE}}$ and $\tau_{\operatorname{AIE}}$
matches the expected effect of giving out more housing vouchers on total homeownership.
Our definition of $\tau_{\operatorname{AIE}}$ also has the potential to help synthesize non-parametric and model-based approaches
to interference by providing a shared estimand that can be studied from both perspectives:
As discussed in Section (ref), while $\tau_{\operatorname{AIE}}$ is defined
in terms of a generic potential outcomes model, it is also a natural estimand in a number of different
structural models. munro2021treatment
pursue this agenda further and show that, in a marketplace governed by a general equilibrium
model where prices mediate interference, $\tau_{\operatorname{AIE}}$ can be expressed in terms of familiar economic quantities such
as price elasticities.
One challenge is that our estimands will in general depend on the design.
In Figure (ref), we illustrate this phenomenon in a Bernoulli experiment by plotting $\tau_{\operatorname{ADE}}(\pi)$ and $\tau_{\operatorname{AIE}}(\pi)$
as a function of $\pi$ in the following three structural models with constant treatment probabilities $\pi$,
equation[equation omitted — 499 chars of source]
where in each case $E_{}\left(\varepsilon_i \mid W\right) = 0$. Here, qualitatively, Setting 1 resembles the one considered
by cai2015social and leung2020treatment as discussed in Example 1 from Section (ref),
Setting 2 exhibits a type of herd immunity where units are more sensitive to treatment when most of their neighbors
are untreated, while Setting 3 has complicated non-linear interference effects. We then see that, in Setting 1,
$\tau_{\operatorname{ADE}}(\pi)$ and $\tau_{\operatorname{AIE}}(\pi)$ do not vary with $\pi$, but in Settings 2 and 3 they do---and may even
change signs.
\ifbio
\else
\captionsetup{singlelinecheck=off}
\fi
figure[figure omitted — 1,003 chars of source]
This potential dependence of $\tau_{\operatorname{ADE}}(\pi)$ and $\tau_{\operatorname{AIE}}(\pi)$ on $\pi$ is something that any
practitioner using these estimands needs to be aware of. However, we believe such dependence to be
largely unavoidable when seeking to define non-parametric estimands under the generality considered
here. For example, when estimating indirect effects of immunization in a population
where roughly 30% of units have been immunized, definitions of the type developed here could be
used to support non-parametric analysis of indirect effects. Now, one should recognize that any such
effects would be local to the current overall immunization rate at 30%, and would likely differ from indirect
effects we would measure at a 50% overall immunization rate. However, it seems unlikely that one could use data from
a population with 30% immunization rate to non-parametrically estimate average outcomes we might
observe at a 50%; rather, to do so, one would need to either posit a model for how infections spread
or collect different data.
\ifbio
\else
\fi
appendix\section{Unbiased Estimation}
The main focus of this paper has been on establishing definitions of average treatment effect metrics under interference, and
on verifying their interpretability under various modeling assumptions. But in order for these definitions to be useful in
applied statistical work, we of course also need these metrics to be readily identifiable and estimable under flexible
conditions. A comprehensive discussion of treatment effect estimation under interference---including distributional
results and efficiency theory---is beyond the scope of this paper. However, as one result in this direction, we note
here that unbiased estimates of the average direct and indirect effects are always available in Bernoulli-randomized
experiments via the Horvitz-Thompson construction.
As above, the case of the average direct effect is already well understood in the literature. The Horvitz-Thompson
estimator for $\tau_{\operatorname{ADE}}(\pi)$ is savje2017average
\begin{equation}
\hat{\tau}_{\operatorname{ADE}}(\pi) = \frac{1}{n} \sum_{i = 1}^n \left\{\frac{W_i Y_i}{\pi_i} - \frac{(1 - W_i) Y_i}{1 - \pi_i}\right\}.
\end{equation}
Furthermore, as is implicitly established in the technical appendix of savje2017average, this estimator is
unbiased in Bernoulli-randomized experiments, i.e., \smash{$E\left\{\hat{\tau}_{\operatorname{ADE}}(\pi)\right\} = \tau_{\operatorname{ADE}}(\pi)$}.
Next, in discussing estimators for $\tau_{\operatorname{AIE}}(\pi)$, we will work under a
network interference model whereby the analyst has access to an interference graph $E_{ij} \in \left\{0, \, 1\right\}$
and knows that the $j$-th unit's potential outcomes are unaffected by the treatment given to the $i$-th unit
whenever $E_{ij} = 0$, i.e.,
\begin{equation}
Y_j(w_i = 0; \, w_{-i}) = Y_j(w_i = 1; \, w_{-i}) whenever E_{ij} = 0,
\end{equation}
for all $i,j = 1, \, \ldots, \, n$ and $w_{-i} \in {0, \, 1}^{n-1}$. This type of assumption is not required in principle; and
in particular, the condition (ref) is vacuous if $E_{ij} = 1$ for all pairs $(i, \, j)$, i.e., if we assume that
any treatment can affect any outcome. However, assumptions of this type are ubiquitous in practice, see, e.g.,
Examples (ref) and (ref) considered in Section (ref), and when we have access to a sparse interference graph
they can considerably improve the precision with which we can estimate indirect effects.
Given this setup, the Horvitz-Thompson estimator for $\tau_{\operatorname{AIE}}(\pi)$ is
\begin{equation}
\hat{\tau}_{\operatorname{AIE}}(\pi) = \frac{1}{n} \sum_{i = 1}^n \sum_{\left\{j \neq i \, : \, E_{ij} = 1\right\}} \left\{\frac{W_i Y_j}{\pi_i} - \frac{(1 - W_i) Y_j}{1 - \pi_i}\right\}.
\end{equation}
Formally, this estimator looks like \smash{$\hat{\tau}_{\operatorname{ADE}}(\pi)$}; except now we are measuring associations between the
$i$-th unit's treatment and the $j$-th unit's outcome, for $i \neq j$. And, as in the case of the direct effect, this estimator
in unbiased in Bernoulli-randomized experiments.
\begin{theorem}
Under Assumption (ref) and in a Bernoulli trial, the Horvitz-Thompson estimator (ref) is unbiased
for the average indirect effect, \smash{$E\left\{\hat{\tau}_{\operatorname{AIE}}(\pi)\right\} = \tau_{\operatorname{AIE}}(\pi)$}.
\end{theorem}
Now, although results on unbiased estimation are helpful, they do not provide a complete
picture of what good estimators for $\tau_{\operatorname{ADE}}(\pi)$ and $\tau_{\operatorname{AIE}}(\pi)$ should look like, and what kinds of
guarantees we should expect. The direct effect estimator (ref) is further studied by savje2017average,
who provide bounds on its mean-squared error under a moderately sparse network interference model as in
(ref); roughly speaking, they assume that the graph $E_{ij}$ has degree bounded on the order of $o(\sqrt{n})$.
li2020random prove a central limit theorem for $\hat{\tau}_{\operatorname{ADE}}(\pi)$ under network interference with a random
graph generative model. Meanwhile, in the case of very sparse graphs, i.e., when vertices in the graph have bounded degrees, the indirect
effect estimator $\hat{\tau}_{\operatorname{AIE}}(\pi)$ could be studied using methods developed in aronow2017estimating and
leung2020treatment. However, in even moderately dense settings, li2020random find that unbiased estimators
of indirect effects may have vary large variance and caution against their use; they also propose alternative estimators that
are more stable---again under a random graph generative model. To the best of our knowledge, efficiency theory for treatment
effect estimation under interference remains as of now uninvestigated.
\section{Proofs}
\subsection{Proof of Theorem (ref)}
We start with a slightly more formal form of the infinitesimal policy effect,
\[\tau_{\operatorname{INF}}(\pi) = \sum_{k=1}^n \frac{\partial}{\partial \pi_k'} \left\{\frac{1}{n}\sum_{i = 1}^n E_{{\pi}'}\left( Y_i\right) \right\}_{\pi' = \pi},\]
i.e., we take derivative with respect to $\pi'$ and evaluate at $\pi$. For index $i$, we can rewrite $E_{}\left(Y_i\right)$ in terms of the potential outcomes:
\begin{align*}
E_{\pi'}\left(Y_i\right) & = \sum_{w_{-k}}\sum_{w_{k}\in\{0,1\}}Y_i(w_k ;w_{-k})pr_{\pi'_{-k}}\left(W_{-k} = w_{-k}\right) pr_{\pi_k'}\left(W_k = w_k\right)\\
& = \sum_{w_{-k}}\sum_{w_{k}\in\{0,1\}}Y_i(w_k ;w_{-k})pr_{\pi'_{-k}}\left(W_{-k} = w_{-k}\right) \left\{w_k\pi_k' + (1-w_k)(1-\pi_k')\right\}.
\end{align*}
The dependency of this term on $\pi_k'$ is clear in this form. Taking a derivative with respect to $\pi_k'$ and evaluating at $\pi_k$ yields
\begin{align*}
\frac{\partial}{\partial \pi_k'}E_{\pi'}\left(Y_i\right)\big|_{\pi' = \pi} & = \sum_{w_{-k}} \left\{Y_i(w_k=1; w_{-k}) - Y_i(w_k=0;w_{-k})\right\} pr_{\pi_{-k}}\left(W_{-k} = w_{-k}\right)\\
& = E_{\pi}\left\{Y_i(w_k=1;W_{-k})-Y_i(w_k=0;W_{-k})\right\}.
\end{align*}
If we sum over the index $i$ and $k$ and multiply the term by $1/n$, we get
\begin{align*}
\tau_{\operatorname{INF}}(\pi)
& = \frac{1}{n}\sum_{k = 1}^n \sum_{i = 1}^n E_{\pi}\left\{Y_i(w_k=1;W_{-k})-Y_i(w_k=0;W_{-k})\right\}
= \tau_{\operatorname{AOE}}(\pi).
\end{align*}
\subsection{Proof of Theorem (ref)}
For each pair of indices $i,j$, we can write $E_\pi\left\{ \frac{W_i Y_j}{\pi_i} - \frac{(1 - W_i) Y_j}{1 - \pi_i}\right\}$ as
\begin{align*}
&E_{\pi}\left( \frac{W_i Y_j }{\pi_i}\right) - E_\pi\left\{ \frac{(1 - W_i) Y_j}{1 - \pi_i}\right\}\\
& \quad \quad = E_\pi\left\{ \frac{W_i Y_j(w_i = 1; W_{-i})}{\pi_i}\right\} - E_\pi\left\{ \frac{(1 - W_i) Y_j(w_i = 0; W_{-i}) }{1 - \pi_i}\right\}\\
& \quad \quad = E_{\pi}\left(\frac{W_i}{\pi_i}\right)E_{\pi}\left\{Y_j(w_i = 1; W_{-i})\right\} - E_{\pi}\left(\frac{1-W_i}{1-\pi_i}\right) E_{\pi}\left\{ Y_j(w_i = 0; W_{-i}) \right\}\\
&\quad \quad = E_{\pi}\left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}.
\end{align*}
Thus summing over $i$ and $j$, we get
\begin{align*}
E_\pi\left\{\hat{\tau}_{\operatorname{AIE}}(\pi)\right\} &= \frac{1}{n} \sum_{i=1}^n \sum_{\{j \neq i:E_{ij}=1\}} E_\pi\left\{ \frac{W_i Y_j }{\pi_i} - \frac{(1 - W_i) Y_j}{1 - \pi_i}\right\}\\
&= \frac{1}{n}\sum_{i=1}^n \sum_{\{j \neq i:E_{ij}=1\}} E_{\pi}\left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}.
\end{align*}
When there is no edge connecting $i$ and $j$, the value of $Y_j$ does not depend on $w_i$, hence $Y_j(w_i=1;W_{-i}) = Y_j(w_i=0;W_{-i})$. Hence we can rewrite the above term:
\begin{align*}
E_\pi\left\{\hat{\tau}_{\operatorname{AIE}}(\pi)\right\}
&= \frac{1}{n}\sum_{i=1}^n \sum_{j \neq i} E_{\pi}\left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}
= \tau_{\operatorname{AIE}}(\pi).
\end{align*}
\subsection{Derivations for Example (ref)}
It's clear from (ref) that $Y_i(w_i = 1; W_{-i}) - Y_i(w_i = 0; W_{-i}) = \beta_2$ and $Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) = \beta_3 E_{ij} /{\sum_{j \neq i} E_{ij}}$. Hence
\begin{align*}
&\tau_{\operatorname{ADE}} = \frac{1}{n} \sum_{i=1}^n E_{\pi}\left\{Y_i(w_i = 1; W_{-i}) - Y_i(w_i = 0; W_{-i})\right\}
= \beta_2, \\
&\tau_{\operatorname{AIE}} = \frac{1}{n}\sum_{i = 1}^n \sum_{j \neq i} E_{\pi}\left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}
= \frac{1}{n}\sum_{i = 1}^n \sum_{j \neq i} \left(\beta_3 E_{ij}/{\sum_{j \neq i} E_{ij}}\right) = \beta_3.
\end{align*}
\subsection{Derivations for Example (ref)}
Model (ref) implies that $Y_i(w_i = 1; W_{-i}) - Y_i(w_i = 0; W_{-i}) = \beta_i$ and $Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) = \nu_{ij}$. Hence
\[
\tau_{\operatorname{ADE}} = \frac{1}{n} \sum_{i=1}^n \beta_i, \ \ \ \
\tau_{\operatorname{AIE}} = \frac{1}{n}\sum_{i = 1}^n \sum_{j \neq i} E_{\pi}\left\{Y_j(w_i=1;W_{-i})-Y_j(w_i=0;W_{-i}) \right\}
= \frac{1}{n}\sum_{i = 1}^n \sum_{j \neq i} \nu_{ij}.
\]
\subsection{Derivations for Example (ref)}
In this clustered setup, the treatment assignment of units outside of the cluster has no effect on the unit's observed outcome.
With a slight abuse of notation, we use $W_{-i}$ to denote the treatment vector of all the units in the same cluster as unit $i$, except the unit itself.
To start with, we find the distribution of $W_{-i}$ marginalized over $W_i$ under this design, where
\[
P(W_{-i}=w)=
\begin{cases}
1-\rho+\frac{\rho}{m},& \text{if } w=(0,\dots,0)\\
\frac{\rho}{m},& \text{if } w\in\{e_1,\dots,e_{m-1}\}\\
0, & \text{otherwise},
\end{cases}
\]
where $e_j\in\mathbb{R}^{m-1}$ denotes the vector with a $1$ in the $j$th coordinate and $0$'s elsewhere. Thus, writing $C_i$ for the indicator that unit $i$ in a treated cluster and recalling that, in our randomization design, if $W_i = 1$ then the $i$-th unit cannot have a treated neighbor:
\begin{align*}
\tau_{\operatorname{ADE}}(\pi) &= \frac{1}{n} \sum_{i=1}^n E\left\{Y_i(w_i = 1; W_{-i}) - Y_i(w_i = 0; W_{-i})\right\}\\
&= \frac{1}{n} \sum_{i=1}^n \left[ pr_\left(W_i = 0 \ & \ C_i = 1\right)\left\{Y_i(treated & exposed) - Y_i(exposed)\right\} \right.\\
&\qquad\qquad\left. + \left\{pr_\left(W_i = 1\right) + pr_\left(C_i = 0\right)\right\} \left\{Y_i(treated) - Y_i(none)\right\} \right ] \\
&= \left(\rho-\frac{\rho}{m}\right)\tau_{SELF,1}+\left(1-\rho+\frac{\rho}{m}\right)\tau_{\text{SELF},0}, \\ \\
\tau_{\operatorname{AIE}}(\pi) &= \frac{1}{n} \sum_{i=1}^n\sum_{j\ne i,E_{ji}=1} E\left\{Y_i(w_j = 1; W_{-j}) - Y_i(w_j = 0; W_{-j})\right\}\\
&= \frac{1}{n} \sum_{i=1}^n \left[(m - 1)pr_\left(W_i = 1\right) \left\{Y_i(\text{treated & exposed}) - Y_i(\text{treated})\right\} \right.\\
&\qquad\qquad\left. + \left\{(m - 1)pr_\left(C_i = 1\right) + pr_\left(W_i = 0 \ & \ C_i = 1\right)\right\}\left\{Y_i(\text{exposed}) - Y_i(\text{none})\right\} \right ] \\
&= (m-1)\left\{\frac{\rho}{m}\tau_{\text{SPILL},1}+\left(1-\rho+\frac{\rho}{m}\right)\tau_{\text{SPILL},0} \right\}.
\end{align*}
\ifyuchen
itemize• In total $n/m$ clusters of size $m$
• Design 1: on the first stage, $\rho\cdot n/m$ clusters are assigned to treatment, and $(1-\rho)\cdot n/m$ clusters are assigned to control; on the second stage, a single unit in each treated cluster is randomly chosen to be treated, and all the other units are assigned to control.
• Design 2: on the first stage, clusters are randomly assigned to treatment with probability $\rho$; on the second stage, units in each treated cluster are randomly chosen to be treated with probability $\pi_0$, and units in each controlled cluster are all assigned to control.
equation*[equation* omitted — 269 chars of source]
where
equation*[equation* omitted — 175 chars of source]
the probability that at least one neighbor is treated,
equation*[equation* omitted — 141 chars of source]
the probability that the unit itself is treated and every other neighbor is in control, and
equation*[equation* omitted — 158 chars of source]
the probability that none of the neighbors are treated.
\fi