Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Semiparametric Estimation of Treatment Effects in Observational Studies with Heterogeneous Partial Interference
\def\spacingset#1{
{#1}} \spacingset{1}
doublespacing\begin{abstract}
In many observational studies in social science and medicine, subjects or units are connected, and one unit's treatment and attributes may affect another's treatment and outcome, violating the stable unit treatment value assumption (SUTVA) and resulting in interference. To enable feasible estimation and inference, many previous works assume exchangeability of interfering units (neighbors). However, in many applications with distinctive units, interference is heterogeneous and needs to be modeled explicitly. In this paper, we focus on the partial interference setting, and only restrict units to be exchangeable conditional on observable characteristics. Under this framework, we propose generalized augmented inverse propensity weighted (AIPW) estimators for general causal estimands that include heterogeneous direct and spillover effects. We show that they are semiparametric efficient and robust to heterogeneous interference as well as model misspecifications. We apply our methods to the Add Health dataset to study the direct effects of alcohol consumption on academic performance and the spillover effects of parental incarceration on adolescent well-being.
{\it Keywords:} Exchangeability, SUTVA, Spillover Effects, Peer Effects, Augmented Inverse Propensity Weighting
\end{abstract}
\section{Introduction}
The classical treatment effect estimation literature typically relies on the stable unit treatment value assumption, or SUTVA rubin1974estimating,rubin1980randomization, which assumes that a unit's potential outcomes do not depend on the treatment status of other units. However, SUTVA can be inappropriate for applications where individuals or units are connected or interact with other, for example directly through social networks banerjee2013diffusion,ogburn2017vaccines and group memberships sacerdote2001peer,miguel2004worms,duflo2011peer,fadlon2019family, or indirectly through equilibrium effects heckman1998general,johari2022experimental,munro2021treatment. Motivated by these applications, there has been a growing literature on the identification and estimation of treatment and spillover effects under interference cox1958planning,hudgens2008toward, that allows a unit's potential outcomes to depend on the treatments of other units.\footnote{In observational studies, the same dependence structure that mediates interference can also lead to
correlations in treatment assignments, e.g., through the contagion of behaviors in a social network. This
implies that the classical unconfoundedness assumption (Rosenbaum and Rubin, 1983) is also likely to fail.}
Often in settings with interference, the mediating interactions among units are heterogeneous. This heterogeneity can arise from different sources, depending on the distinct nature and strength of interactions. For example, interference among close friends can matter a lot more than that among acquaintances, and units within close geographical proximity tend to have stronger interference than distant units. Another common scenario is where units have distinct types based on observable information, and interactions across different types of units are expected to be heterogeneous. For example, for a child in a household in Figure (ref), spillover effects from a treated parent may be different from the spillover effect from a treated sibling, or may depend on the gender of the treated parent. Potential heterogeneities in interference abound in empirical applications, and failure to take them into account may lead to biased estimation and invalid inference, even if one is interested in aggregate or marginal treatment effects like the ATE forastiere2020identification.
In this paper, we focus on settings with partial interference sobel2006randomized,hudgens2008toward, where units can be naturally divided into disjoint clusters (e.g., households), and interference is restricted to occur only among units within the same cluster (see Figure (ref)).
We further assume that within clusters, there is a partitioning of units into distinct types based on observable characteristics (e.g., parents vs. children or father vs. mother), and a unit's interactions with its neighbors of the same type are exchangeable. However, the interactions can be arbitrarily distinct for units of different types. We refer to this setup as the conditional exchangeability framework, which leads to an interference structure we call \emph{heterogeneous partial interference}. This structure can be expressed in the form of an \emph{exposure mapping} as established in previous works manski2013identification, aronow2017estimating,forastiere2020identification, vazquez2017identification. Nevertheless, it enjoys the benefit of intuitive interpretation of various sources of heterogeneity, including that in treatment assignment probabilities and treatment effects.
Our framework also includes the commonly used \emph{homogeneous partial interference} assumption as a special case, where all units in a cluster are exchangeable, not just those of the same type.
\begin{figure}
\begin{subfigure}{0.5\textwidth}
\caption{Household example}
\end{subfigure}
\begin{subfigure}{0.38\textwidth}
\caption{Village example}
\end{subfigure}
\caption{ Illustration of our setting under heterogeneous partial interference. In Figure (ref), each household consists of two exchangeable subsets (see Section (ref)): parents and children, so that interference from units within each subset is homogeneous (denoted by links with the same color). In Figure (ref), each village consists of leaders/influencers and villagers, and interference depends on the social status of a person in the village. Homogeneous interference among villagers is indicated by the blue-shaded ring. We allow for \emph{varying} cluster sizes in extensions.}
\end{figure}
Under the conditional exchangeability framework, we propose a class of generalized augmented inverse propensity weighted (AIPW) estimators for heterogeneous direct and spillover effects, which are shown to be doubly robust, asymptotically normal, and semiparametric efficient. To the best of our knowledge, there have not been any formal results on semiparametric efficiency and AIPW estimators in both homogeneous and heterogeneous partial interference settings for estimands defined under the corresponding exposure mapping. We further propose a data-driven framework to determine the appropriate interference structure by leveraging statistical tests using a matching-based variance estimator, which allows the practitioner to detect and account for heterogeneous partial interference in the absence of domain knowledge.
Our estimators are natural and relevant in many applications, such as the optimal design of treatment allocation rules under resource constraints, where identifying heterogeneities in direct and spillover effects is important banerjee2013diffusion.
In addition, as shown in this work, even if our primary object of interest is an aggregated marginal treatment effect like the ATE, correctly accounting for the heterogeneity in interference is crucial to obtaining a \emph{consistent} and \emph{efficient} estimate of the aggregate effect. This is particularly the case when using \emph{observational data}, which requires at least one of the outcome and treatment assignment models to be consistently estimated. When the heterogeneity in interference is not appropriately accounted for, both models may fail to be consistently estimated, resulting in biased estimates of the aggregate effect.
We demonstrate the superior finite sample properties, robustness in estimation and inference, and overall robustness of our estimators through extensive simulation studies. We also apply our methods to the Add Health data harris2009national to study two empirical questions. First, we investigate the effects of regular alcohol use on students' academic performance. Second, we examine the impact of parental incarceration on adolescent well-being. These applications showcase the practical relevance and effectiveness of our methods in real-world settings.
In summary, our results provide practitioners with a data-driven toolkit to assess and account for the impact of heterogeneous interference in a wide range of applications with observational data.
\paragraph{Related Works} This paper contributes to the growing literature on treatment effect estimation under partial interference. Partial interference in randomized experiments has been studied by halloran1995causal, sobel2006randomized, hudgens2008toward, liu2014large, vazquez2017identification, basse2018, jagadeesan2020designs, liu2023inference, often in the context of two-stage randomization. Other recent works have also studied the estimation and inference of causal quantities from \emph{observational} data tchetgen2012causal,perez2014assessing,liu2016inverse,barkley2020causal,park2020efficient,forastiere2020identification. Both these works and our work account for endogenous treatment assignments in observational data under partial interference. Our work is distinct in the following aspects. First, in many previous works, the causal quantities of interest are typically an average treatment effect over the whole population when each neighbor is \emph{hypothetically} and \emph{independently} treated with probability $\alpha$, i.e., $\alpha$-allocation strategy. In contrast, our estimands are defined under the exposure mapping derived from the conditional exchangeability framework, and also include more granular estimands defined for distinct subpopulations to explicitly account for both heterogeneous interference and treatment effects. As a result, our work could help address more general empirical questions involving a wider class of causal quantities. Second, our generalized AIPW estimators are different from those proposed in
tchetgen2012causal,perez2014assessing,liu2016inverse,barkley2020causal, which are generalized IPW estimators, and those in liu2019doubly and park2020efficient which are generalized AIPW estimators using different estimation strategies for the propensity and outcome models. Third, we provide the first semiparametric efficiency lower bound type results for estimands based on exposure mappings. Compared to the efficiency bounds based on the $\alpha$-allocation strategy in park2020efficient, our bounds apply to a different class of estimands and explicitly quantify how various sources of heterogeneity affect estimation efficiency. Lastly, we also propose a data-driven approach to detect heterogeneous interference based on hypothesis testing using a matching-based variance estimator, which was not done in prior works. Our work is closely related to forastiere2020identification, particularly with respect to the idea of defining estimands and constructing estimators through the use of a specific exposure mapping. While their work considers more general interference, our paper provides a specific interference structure under conditional exchangeability that is relevant and interpretable in many empirical settings. More importantly, our work complements this line of literature (e.g., forastiere2020identification, tortu2020modelling) by offering efficient estimators and asymptotically valid inference methods that allow the practitioner to determine the appropriate interference structure.
The rest of the paper is organized as follows. Section (ref) presents model assumptions and the conditional exchageability framework. Section (ref) defines the causal estimands of interest and proposes the estimation procedure based on generalized AIPW estimators. Section (ref) presents main asymptotic results. Sections (ref) and (ref) present simulation results and applications to the Add Health dataset. Section (ref) concludes the paper.
\section{Model Setup}
In this paper, we study the potential outcomes framework in the observational setting with partial interference, where we observe $N$ units that are divided into a large number $M=O(N)$ of clusters, and interference is restricted to among units within the same cluster. This clustering structure can arise from group memberships, geographical proximity, and sampling or randomization schemes. We assume that each cluster is drawn i.i.d. from a population distribution $\mathbb{P}$, but within each cluster, covariates and treatment assignments may have arbitrary dependence across units. In particular, we allow for phenomena such as \emph{homophily} of similar units mcpherson2001birds,bramoulle2012homophily and \emph{contagion} of treatment assignments centola2010spread,christakis2013social.
We state the standard assumptions on the data generating process in Section (ref) and introduce our framework for heterogeneous interference in Section (ref).
\subsection{General Assumptions}
Let $c=1,\dots,M$ be the index of a cluster. For exposition, assume for now that each cluster has the same size $n = N/M$. We generalize our results to varying cluster sizes in Appendix (ref). Let $\mathbf{Y}_c \in \mathbb{R}^{n}$, $\mathbf{Z}_c \in \{0,1\}^n$, and $\mathbf{X}_c \in \mathbb{R}^{n\times d_x}$, be the observed outcomes, treatment assignments, and covariates of all the $n$ units in cluster $c$, where $d_x$ is the dimenionality of each unit's covariates.
Let $Y_{c,i} \in \+R$, $Z_{c,i} \in \{0,1\}$, and $X_{c,i}\in \+R^{d_x}$ be the observed outcome, treatment assignment and covariates of unit $i$ in cluster $c$, where $Y_{c,i}$ and $Z_{c,i}$ are the $i$-th coordinate of $\mathbf{Y}$ and $\mathbf{Z}$, and $X_{c,i}$ is the $i$-th row of $\mathbf{X}_c$. We define unit $i$'s \textbf{neighbors} as all units in cluster $c$ other than unit $i$.\footnote{We can also extend to settings where there is a network structure within each cluster, so the definition of neighbors is unit-dependent.} Let $\*Y_{c,(i)} \in \+R^{n-1}$, $\*Z_{c,(i)} \in \{0,1\}^{n-1}$, and $\*X_{c,(i)} \in \+R^{(n-1) \times d_x}$ be the observed outcome, treatment assignment and covariates of $i$'s neighbors in cluster $c$.
\begin{assumption}[i.i.d. Clusters]
The $M$ clusters are i.i.d., i.e., tuples $(\c(X),\c(Y),\c(Z))$ for $c=1,\dots,M$ are drawn i.i.d. from some compactly-supported population distribution $\mathbb{P}$.
\end{assumption}
Many applications exhibit a natural clustering structure. For example, clusters may be households vazquez2017identification or dormitory rooms sacerdote2001peer. We use households as a running example throughout the paper (see Figure (ref)). The mechanism behind the clustering structure could be exogenous (e.g., assigned externally based on units' characteristics) or endogenous (e.g., through homophily). Even in general network settings where a clustering structure is not obvious, graph segmentation or sampling techniques can be used to partition the network into clusters ugander2013graph,saint2019using. To address settings where clusters are potentially correlated with each other, we discuss the generalization to weakly connected clusters in Appendix (ref).
Note that our framework allows units' covariates and treatments to be arbitrarily correlated among units in the same cluster, thus allowing for contagion of treatments in addition to homophily.
Next, we allow a unit's potential outcomes to depend on the treatment assignments of its \emph{neighbors} (in addition to its own treatment) within the same cluster, which relaxes the SUTVA assumption rubin1980randomization,rubin1986comment. However, we impose the partial interference assumption sobel2006randomized, where there is no interference between units in different clusters halloran1995causal. This assumption is reasonable for applications with disjoint or sufficiently separate clusters, e.g., in space or time.
\begin{assumption}[Partial Interference]
For any unit $i$ in cluster $c$, unit $i$'s potential outcomes can only depend on the treatment assignments of units within the same cluster $c$.
\end{assumption}
Given Assumption (ref), we can define the potential outcomes of unit $i$ in cluster $c$ as
\begin{equation*}
Y_{c,i}(\s(z),\n(\mathbf{z}))
\end{equation*}
for $c \in \{1, \cdots, M\}$, $i \in \{1, \cdots, n\}$, $\s(z) \in \{0,1\}$, and $\n(\mathbf{z}) \in \{0,1\}^{n-1}$. Here $\s(z)$ and $\n(\mathbf{z})$ denote the deterministic values that $\s(Z)$ and $\n(\mathbf{Z})$ can take.
The observed outcome of unit $i$ in cluster $c$ is
\[Y_{c,i} \equiv Y_{c,i}(Z_{c,i},\mathbf{Z}_{c,(i)}). \]
Now we consider the treatment assignment mechanism. In the classical observational setting, the unconfoundedness assumption is commonly imposed rosenbaum1983central and states that treatment assignments are free from dependence on potential outcomes \emph{conditional} on units' own covariates. In the presence of interference, this may no longer be true: as units interact with each other, one unit's treatment assignment could also depend on its neighbors' characteristics and treatment assignments. Therefore, we impose a generalized unconfoundedness assumption, which is necessary for the identification of causal effects under interference in observational studies.
\begin{assumption}[Generalized Unconfoundedness]
For any cluster $c$, unit $i$, and treatment assignment values $\big(\s(z),\n(\mathbf{z})\big)$,
\begin{align}
\s(Y)\big(\s(z), \n(\mathbf{z})\big) &\perp \big(\s(Z), \n(\mathbf{Z})\big) \mid \mathbf{X}_c.
\end{align}
\end{assumption}
Assumption (ref) extends the classical unconfoundedness assumption and implies that the treatment assignment probability satisfies
\begin{align*}
& P\big(\mathbf{Z}_c \mid \s(Y)\big(\s(z), \n(\mathbf{z})\big), \mathbf{X}_c\big) = P(\mathbf{Z}_c \mid \mathbf{X}_c) \qquad \forall i, \s(z),\n(\mathbf{z}).
\end{align*}
We define $P(\mathbf{Z}_c \mid \mathbf{X}_c)$ as the \textbf{propensity score} of cluster $c$.
Assumptions similar to Assumption (ref) have been used in the literature liu2016inverse,forastiere2020identification,park2020efficient.
Lastly, for the purpose of identification, we impose the overlap assumption on cluster-level treatment probabilities (propensity scores).
\begin{assumption}[Overlap]
There exist some $\underline{p}$ and $\overline{p}$ such that for any value of $\c(Z)$ and $\c(X)$,
\begin{eqnarray}
0 < \underline{p} \leq P(\c(Z) \mid \c(X)) \leq \overline{p} < 1.
\end{eqnarray}
\end{assumption}
Assumption (ref) implies that $P(\s(Z) \mid \c(X))$ is bounded away from $0$ and $1$, i.e., the classical overlap assumption holds for every unit $i$ in cluster $c$ if we condition on cluster covariates $\c(X)$ instead of individual covariates only.\footnote{A unit's propensity scores are sometimes stated using the unit's own covariates only in some previous works. Note that in our framework, we can also define a unit's propensity score as $P(Z_{c,i} \mid \tilde{\*X}_{c,i})$ which uses the unit's “own” covariates $\tilde{\*X}_{c,i}$, and $\tilde{\*X}_{c,i}$ is defined as $\tilde{\*X}_{c,i} \coloneqq \left(X_{c,i}, \*X_{c,(i)} \right)$ that contains $i$'s neighbors' information.}
\subsection{Conditional Exchangeability}
In many applications, units can have heterogeneous interactions and interference often depends on the particular \emph{type} of neighbors that are treated.
Building on vazquez2017identification and forastiere2020identification, we adopt a conditional exchangeability framework to formalize such heterogeneous interference. In our conditional exchangeability framework, we partition units within each cluster into $m \geq 1$ exchangeable and disjoint subsets (such as parents and children in a family), denoted by $\mathcal{I}_1, \cdots, \mathcal{I}_m$, that satisfy\footnote{Partitions can have varying sizes across clusters. $\mathcal{I}_j$ can be a singleton for any $j$. This partition implicitly imposes an ordering of subsets. In the household example with cluster size four in Figure (ref), units $i=1$ and $i=2$ are parents, and $i=3$ and $i=4$ are children for all clusters. Within each subset $\mathcal{I}_j$, the ordering of units can be arbitrary.}
\[\mathcal{I}_1 \cup \mathcal{I}_2 \cup \cdots \cup \mathcal{I}_m = \{1, 2, \cdots, n\}, \quad \text{and }\quad \mathcal{I}_j \cap \mathcal{I}_k = \emptyset \quad \text{ for $j \neq k$}. \]
We assume a unit's interference from treatments of neighboring units in the same subset are the same, but \emph{may} be different for neighboring units in distinct subsets, which is formalized in Assumption (ref) below. For ease of exposition, we assume for now that the partition is the same for \emph{all} clusters and is unit-independent, but our results can be generalized to settings where the partition is unit-dependent or cluster-dependent in a straightforward manner, as discussed in Appendix (ref).
\begin{assumption}[Conditional Exchangeability of Potential Outcomes]
For each unit $i$, its potential outcomes are exchangeable with respect to arbitrary permutations of treatment assignments of other units in the same subset, i.e.,
\begin{align}
Y_{c,i}\big(\s(z), \mathbf{z}_{c,(i),1}, \cdots, \mathbf{z}_{c,(i),m} \big) &= Y_{c,i}\big(\s(z), \pi_1(\mathbf{z}_{c,(i),1}), \cdots, \pi_m(\mathbf{z}_{c,(i),m}) \big) \, ,
\end{align}
where for $j \in \{1, \cdots, m\}$, $\mathbf{z}_{c,(i),j}$ is the treatment realization that units in $\mathcal{I}_j\backslash i$ can take, and $\pi_j(\cdot) \in \mathbb{S}^{|\mathcal{I}_j\backslash i|}$ is an arbitrary permutation. $\mathcal{I}_j\backslash i$ is the subset of units in $\mathcal{I}_j$ but not in $\{i\}$. $|\mathcal{I}_j\backslash i|$ is the cardinality of $\mathcal{I}_j\backslash i$.\footnote{If $i \not\in \mathcal{I}_j$, then $\mathcal{I}_j\backslash i = \mathcal{I}_j$. If $i \in \mathcal{I}_j$ and $\mathcal{I}_j$ is singleton, then $\mathcal{I}_j\backslash i = \emptyset$ and therefore $\mathbf{z}_{c,(i),j} = \emptyset$, $\mathbf{X}_{c,(i),j} = \emptyset$, $\pi_j( \mathbf{z}_{c,(i),j}) = \emptyset$ and $\pi_j^{-1}(\mathbf{X}_{c,(i),j}) = \emptyset$ for any permutation $\pi_j(\cdot)$.}
\end{assumption}
If a cluster only has two units or if $m = n$, then Assumption (ref) always holds. If $m = 1$, then Assumption (ref) reduces to the \textbf{fully exchangeable} assumption (hudgens2008toward calls this stratified interference). In this case, potential outcomes only depend on how many neighbors, but not which ones, are treated. This setting is commonly used in epidemiology, for example, to study the effect of vaccine coverage hudgens2008toward,tchetgen2012causal.
For general $m$, Assumption (ref) implies that potential outcomes depend on the \emph{numbers} of treated neighbors in each subset $\mathcal{I}_j$. Thus in the definition of potential outcomes, the ($n-$1)-dimensional vector of neighbors' treatment assignments $\n(\mathbf{z})$ can be summarized by the $m$-dimensional\footnote{When $m=n$, no units are exchangeable and we can identify $\mathbf{g}_{c,i}$ with the ($n-$1)-dimensional vector $\n(\mathbf{z})$ itself, instead of an $n$-dimensional vector with a trivial coordinate always equal to zero.} vector $\mathbf{g} \in \mathbb{Z}_{\geq 0}^m$ of the number of treated neighbors in each subset:
\begin{equation*}
Y_{c,i}\big(\s(z), \mathbf{z}_{c,(i),1}, \cdots, \mathbf{z}_{c,(i),m} \big) \equiv \s(Y)(\s(z), \underbrace{g_{c,1}, \cdots,g_{c,m}}_{\mathbf{g}_{c,i}} ),
\end{equation*}
for any $\mathbf{z}_{c,(i),j}$ that satisfies $\left\lVert\mathbf{z}_{c,(i),j}\right\rVert_1 = g_{c,j}$ with $g_{c,j} \in \{0, \cdots,|\mathcal{I}_j \backslash i|\}$, where $\left\lVert\mathbf{z}_{c,(i),j}\right\rVert_1$ is the $\ell_1$ norm of $\mathbf{z}_{c,(i),j}$.\footnote{In an extension in Section (ref), we consider the case where a unit's potential outcomes may depend on some, but not all, units' treatment assignments within the same cluster.}
The number of potential outcomes can be significantly reduced by using $\s(Y)(\s(z),\mathbf{g}_{c,i})$. When $m = 1$, the number of potential outcomes is reduced from $2^n$ to $2(n-1)$.
Our definition of potential outcomes $\s(Y)(\s(z),\mathbf{g}_{c,i})$ can be viewed as a form of exposure mapping aronow2017estimating from $\mathbf{z}_{c,(i)}$ to $\mathbf{g}_{c,i}$ defined through exchangeable subsets. Our conditional exchangeability framework thus complements bargagli2020heterogeneous, tortu2020modelling and forastiere2020identification which use general treatment exposure mappings to capture heterogeneous interference. In contrast, we explicitly model heterogeneous interference using the exchangeable subsets $\mathcal{I}_j$ based on observable and interpretable characteristics, which aligns with many applications where such partitions arise naturally. These subsets are also used in Section (ref) below for the definition and identification of \emph{heterogeneous} direct treatment and spillover effects for different subsets of units, whereas many previous works treat the heterogeneity of interference as a nuisance in the estimation of marginal treatment effects. A similar exchangeability condition is also discussed in vazquez2017identification in the experimental setting. Our framework is designed for the observational setting and therefore imposes exchangeability on the propensity as well.
Given the conditional exchangeability of potential outcomes imposed in Assumption (ref), it follows immediately that the unconfoundedness assumption in Assumption (ref) continues to hold when using $\s(Y)(z, \mathbf{g})$ as the definition of potential outcomes:
\begin{equation*}
\s(Y)(z, \mathbf{g}) \perp \left(\s(Z), \s(\mathbf{G})\right) \mid \mathbf{X}_c, \quad \forall c,i,z,\mathbf{g},
\end{equation*}
where $\mathbf{G}_{c,i} = \big(G_{c,i,1}, \cdots, G_{c,i,m} \big)$, and $G_{c,i,j} = \sum_{k \in \mathcal{I}_j \backslash i} Z_{c,k}$ is the number of treated neighbors of unit $i$ in subset $j$ in cluster $c$.\footnote{If $i \in \mathcal{I}_j$ and $\mathcal{I}_j$ is singleton, then $\mathcal{I}_j \backslash i = \emptyset$ and $\sum_{k \in \mathcal{I}_j \backslash i} Z_{c,k} = 0$.}
We define the \textbf{conditional outcome model} as
\begin{align*}
\mu_{i,(z,\mathbf{g})}(\c(X)) \coloneqq&\mathbb{E}[Y_{c,i}(z,\mathbf{g}) \mid \s(X),\mathbf{X}_{c,(i),1}, \cdots, \mathbf{X}_{c,(i),m}].
\end{align*}
As shown in Lemma (ref) in Appendix (ref), the conditional outcome model is well-defined and satisfies the following permutation \emph{invariance} property over covariates:
\begin{align*}
\mu_{i,(z,\mathbf{g})}(\c(X))=&\mathbb{E}[Y_{c,i}(z,\mathbf{g}) \mid \s(X) ,\pi_1(\mathbf{X}_{c,(i),1}), \cdots, \pi_m(\mathbf{X}_{c,(i),m})].
\end{align*}
In other words, the effect of neighbors' covariates on unit $i$'s potential outcomes is invariant under permutations of all the neighbors in $\mathcal{I}_j\backslash i$. It is then possible to model the effect of $\n(\mathbf{X})$ on $\mu_{i,(z,\mathbf{g})}(X_{c,i},\n(\mathbf{X}))$ using the summary statistics of covariates $\mathbf{X}_{c,(i),j}$ in each subset $j$, such as the mean or second moment.
Under our conditional exchangeability framework for observational studies, we also define the \textbf{joint propensity model} as
\begin{align*}
p_{i,(z,\mathbf{g})}(\c(X)) \coloneqq& P\big(\s(Z) = z, \mathbf{G}_{c,i}=\mathbf{g} \mid X_{c,i}, \mathbf{X}_{c,(i),1}, \cdots, \mathbf{X}_{c,(i),m}\big).
\end{align*}
Analogous to the invariance property of the conditional outcome model under conditional exchangeability, we state a similar condition for the propensity model.
\begin{assumption}[Conditional Exchangeability of Propensity Models] For arbitrary permutation $\pi_j(\cdot)$, the joint propensity model satisfies
\begin{align*}
p_{i,(z,\mathbf{g})}(\c(X)) =& P\big(\s(Z) = z, \mathbf{G}_{c,i}=\mathbf{g} \mid X_{c,i}, \pi_1(\mathbf{X}_{c,(i),1}), \cdots, \pi_m(\mathbf{X}_{c,(i),m})\big),
\end{align*}
where for each $j$, $\mathbf{X}_{c,(i),j}$ is the covariates of units in $\mathcal{I}_j\backslash i$.
\end{assumption}
Under Assumption (ref), a unit's treatment assignment probability can be impacted differently by the treatments and covariates of neighbors in \emph{distinct} subsets, but is invariant under arbitrary permutations of units within the same subset.
Lastly, the overlap assumption in Assumption (ref) can also be simplified to $ 0 < \underline{p} \leq p_{i,(z,\mathbf{g})}(\c(X)) \leq \overline{p}<1$ for any $\c(X),i,z$ and $\mathbf{g}$ under Assumptions (ref) and (ref). In practice, this overlap assumption can be harder to assess compared to the classical setting, due to the high dimensionality of treatment levels induced by interference. This problem is of independent interest and left for future works.
\section{Semiparametric Treatment Effect Estimation under Conditional Exchangeability}
In this section, we first lay out the heterogeneous causal estimands under conditional exchangeability in Section (ref), and then propose our estimators in Section (ref).
\subsection{Estimands}
We focus on two types of interference-based estimands. The first one is the \textbf{average direct effect (ADE)} for units in subset $\mathcal{I}_j$, denoted as $\beta_j(\mathbf{g})$, that measures the average direct treatment effect for units in $\mathcal{I}_j$, given that the number of treated neighbors in all subsets is $\mathbf{g}$:
\begin{equation}
\beta_j(\mathbf{g}) \coloneqq \frac{1}{|\mathcal{I}_j|} \sum_{i \in \mathcal{I}_j} \mathbb{E}[Y_{c,i}(1,\mathbf{g} )-Y_{c,i}(0,\mathbf{g})].
\end{equation}
The second one is the \textbf{average spillover effect (ASE)} for units in subset $\mathcal{I}_j$, denoted as $\tau_j(z,\mathbf{g},\mathbf{g}^\prime)$, that measures the average difference in expected outcomes for units in $\mathcal{I}_j$ when the number of treated neighbors is $\mathbf{G}_{c,i} = \mathbf{g}$ versus when $\mathbf{G}_{c,i} = \mathbf{g}^\prime$, and when unit $i$'s own treatment is $Z_{c,i} = z$:
\begin{equation}
\tau_j(z,\mathbf{g},\mathbf{g}^\prime) \coloneqq \frac{1}{|\mathcal{I}_j|} \sum_{i \in \mathcal{I}_j} \mathbb{E}[Y_{c,i}(z,\mathbf{g} )-Y_{c,i}(z, \mathbf{g}^\prime )] .
\end{equation}
ADE $\beta_j(\mathbf{g})$ relates to ASE $\tau_j(z,\mathbf{g},\mathbf{g}^\prime)$ through the equality\footnote{Both ADE and ASE are special cases of the more general class of estimands
\begin{align}
\psi_j(z,\mathbf{g}) - \psi_j(z^\prime,\mathbf{g}^\prime) = \frac{1}{|\mathcal{I}_j|}\sum_{i\in\mathcal{I}_j} \mathbb{E}[Y_{c,i}(z,\mathbf{g})-Y_{c,i}(z^\prime,\mathbf{g}^\prime)],
\end{align}
for which our proposed estimators and their asymptotic properties can be directly generalized. We provide the asymptotic distribution of these generalized estimands in Appendix (ref).}
\[\beta_j(\mathbf{g}) + \tau_j(0,\mathbf{g},\mathbf{g}^\prime) = \beta_j(\mathbf{g}^\prime)+\tau_j(1,\mathbf{g},\mathbf{g}^\prime),\]
and when $\mathbf{g}^\prime=\mathbf{0}$ and $\|\mathbf{g}\|_1=n-1$, the sum on the left can be interpreted as a “total treatment effect” for units in subset $j$ chin2018central,savje2021average.
In the household example (Figure (ref)) with $m = 2$, if $\mathcal{I}_1$ and $\mathcal{I}_2$ are the sets of parents and children, respectively, then $\beta_{2}((g_1,g_2))$ measures the ADE for children, given $g_1$ treated parents and $g_2$ treated siblings. $\tau_{2}(0, (g_1, g_2), (0,g_2))$ measures the ASE from $g_1$ treated parents to children, given the children are untreated and have $g_2$ treated siblings.
Our definitions of ADE and ASE are analogous to the direct and spillover effects defined in tchetgen2012causal,liu2016inverse,barkley2020causal,park2020efficient,forastiere2020identification,vazquez2017identification. A key distinction is that we define ADE and ASE as averages over units in a specific subset, while prior works primarily focus on marginal effects over all units.
When there is a natural and interpretable partition of clusters, our framework allows for the identification of \emph{heterogeneous} ADE and ASE across different subsets or types of units.\footnote{These subsets were defined in Section (ref) to model heterogeneities in interference. Although here they serve the dual function of
characterizing heterogeneities of treatment effects across units, in general, the subsets of units that are exchangeable in Assumption (ref) and that are used to define treatment effects do not have to depend on the same partition. For example, we can define and estimate treatment effects for a strict subset of $\mathcal{I}_j$ for any $j$.} This is important in many applications, e.g., when we want to design better treatment targeting rules based on the estimated treatment effects for different types of units from observational data.
\begin{remark}[Aggregate Estimands]
\normalfont
In some settings, we may also be interested in treatment effects that aggregate over \emph{multiple} $\beta_j(\mathbf{g})$ or $\tau_j(z,\mathbf{g},\mathbf{g}^{\prime})$. One class of such effects aggregates over multiple types of units (i.e., over $j$), such as (whenever $\mathbf{g}$ is feasible for all $j\in \mathcal{J}$),
\[\beta_{\mathcal{J}}(\mathbf{g}) := \frac{1}{| \cup_{j \in \mathcal{J}} \mathcal{I}_j|}\sum_{i \in \cup_{j \in \mathcal{J}} \mathcal{I}_j} \mathbb{E}[Y_{c,i}(1,\mathbf{g} )-Y_{c,i}(0,\mathbf{g})], \qquad \mathcal{J} \subset \{1,\cdots, m\}.\]
For example, if $\mathcal{J} = \{1,\cdots, m\}$, then $\beta_{\mathcal{J}}(\mathbf{g}) $ is the average direct effect of all units in a cluster. When $m=1$, $\beta_{\mathcal{J}}(\mathbf{g}) $ further reduces to the ADE in previous works that rely on full exchangeability assumptions.
Alternatively, we may be interested in aggregating over different $\mathbf{g}$. For example, for weights $\omega(\mathbf{g})$ satisfying $\omega(\mathbf{g}) \geq 0$ and $\sum_{\mathbf{g} \in \mathcal{G}}\omega(\mathbf{g})=1$ for some collection $\mathcal{G}$ of $\mathbf{g}$, we can define the aggregate quantity
\[ \beta_j(\mathcal{G}) \coloneqq
\sum_{\mathbf{g} \in \mathcal{G}} \omega(\mathbf{g}) \cdot \beta_j(\mathbf{g}).\]
When $\mathcal{G}$ is the set of all possible $\mathbf{g}$ for units of type $j$, this estimand can be viewed as a natural analogue of the classical ATE under interference. Finally, we may aggregate over both $j \in \mathcal{J}$ and $\mathbf{g} \in \mathcal{G}$. An important example is the direct effect defined under the $\alpha$-allocation strategy tchetgen2012causal. This is the special case with $m=n$, $\mathcal{J}=\{1,\cdots, m\}$, $\mathcal{G}=\{0,1\}^{n-1}$, and $\omega(\mathbf{g})=\alpha^{\|\mathbf{g}\|_1}\cdot (1-\alpha)^{n-1-\|\mathbf{g}\|_1}$.
In this paper, we primarily focus on the estimation of $\beta_j(\mathbf{g})$ (and $\tau_j(z,\mathbf{g},\mathbf{g}^\prime) $), which are not only of interest themselves, but also serve as important building blocks in the \emph{unbiased} and \emph{efficient} estimation of aggregate causal effects. In particular, the mis-specification of interference structures can lead to biased estimators of $\beta_{\mathcal{J}}(\mathbf{g})$. The intuition is that if the heterogeneity is overlooked, both the propensity and outcome models may be consistently estimated, resulting in biased estimates of treatment effects. On the other hand, specifying the heterogeneity in a more complicated way than needed could lead to less efficient estimates of treatment effects. We discuss this bias-efficiency trade-off in detail in Appendix (ref) and illustrate with a numerical example in Table
(ref). An important implication of our analysis is that previous estimation strategies proposed for estimands based on the $\alpha$-allocation strategy may not always be efficient.
\end{remark}
\subsection{Estimators}
Motivated by the AIPW estimators in the classical observational setting without interference robins1994estimation,robins1995analysis, we propose the generalized AIPW estimators for $\beta_j(\mathbf{g})$ and $\tau_j(z,\mathbf{g},\mathbf{g}^\prime)$:
\begin{align}
\hat{\beta}^\mathrm{aipw}_j(\mathbf{g} ) =& \hat \psi_j(1,\mathbf{g}) - \hat \psi_j(0,\mathbf{g}) \\
\hat{\tau}^\mathrm{aipw}_j(z, \mathbf{g},\mathbf{g}^\prime ) =& \hat \psi_j(z,\mathbf{g}) - \hat \psi_j(z,\mathbf{g}^\prime),
\end{align}
where
\begin{align*}
\hat \psi_j(z,\mathbf{g}) =& \frac{1}{M |\mathcal{I}_j|} \sum_{c=1}^{M} \sum_{i \in \mathcal{I}_j} \hat{\phi}_{c,i}(z, \mathbf{g})
\end{align*}
and $\hat{\phi}_{c,i}(z, \mathbf{g})$ is the estimated score of unit $i$ in cluster $c$ and is defined as
\begin{align}
\hat{\phi}_{c,i}(z, \mathbf{g}) =& \underbrace{\frac{\boldsymbol{1}\{\s(Z)=z,\mathbf{G}_{c,i}=\mathbf{g} \}\s(Y) }{\hat p_{i,(z,\mathbf{g})}(\c(X)) } }_{\mathrm{IPW} } + \underbrace{ \bigg( 1 - \frac{\boldsymbol{1}\{\s(Z)=z,\mathbf{G}_{c,i}=\mathbf{g} \} }{\hat p_{i,(z,\mathbf{g})}(\c(X)) } \bigg) \cdot \hat{\mu}_{i,(z,\mathbf{g})}(\c(X)) }_{\mathrm{augmentation} } ,
\end{align}
$\hat p_{i,(z,\mathbf{g})}(\c(X))$ and $\hat{\mu}_{i,(z,\mathbf{g})}(\c(X))$ are unit $i$'s estimated joint propensity and conditional outcome models.
Analogous to the AIPW in the classical setting robins1994estimation,robins1995analysis,
both $\hat{\beta}^\mathrm{aipw}_j(\mathbf{g})$ and $\hat{\tau}^\mathrm{aipw}_j(z, \mathbf{g},\mathbf{g}^\prime )$ can be decomposed into two parts:
the first part is the generalized IPW estimator, and the second part is an augmentation term that is a weighted average of conditional outcomes. As a result, the doubly robust property of the classical AIPW estimator carries over to $\hat{\beta}^\mathrm{aipw}_j(\mathbf{g})$ and $\hat{\tau}^\mathrm{aipw}_j(z, \mathbf{g},\mathbf{g}^\prime )$, as will be shown in Theorem (ref) in Section (ref). Double robustness means that the estimated average treatment effect is consistent if either the outcome model or the propensity model can be consistently estimated (e.g., robins1994estimation,scharfstein1999adjusting,kang2007demystifying,tsiatis2007comment).
We first discuss the estimation of the conditional outcome models $\mu_{i,(z,\mathbf{g})}(\*x)$ and joint propensity models $p_{i,(z,\mathbf{g})}(\*x)$, for generic $\*x \in \mathbb{R}^{n\times d_x}$ and $i\in \{1,\cdots, n\}$, using nonparametric series estimators (a.k.a. sieve estimators) newey1997convergence,chen2007large,hirano2003efficient,cattaneo2010efficient, on which our asymptotic results in Section (ref) are built. Sieve estimators are a sequence of estimators that progressively use more basis functions and more complex models to approximate $\mu_{i,(z,\mathbf{g})}(\*x)$ and $p_{i,(z,\mathbf{g})}(\*x)$. Let $\{r_k(\*x)\}_{k = 1}^\infty$ be such a sequence of known functions (e.g., polynomials). In the sequence of estimators for $\mu_{i,(z,\mathbf{g})}(\*x)$, let $\hat{\mu}_{i,K,(z,\mathbf{g})}(\*x)$ be the estimator that uses the first $K$ approximation functions $R_K(\*x) = \big(r_1(\*x) ~ \cdots ~ r_K(\*x) \big)^\top$ and takes the form of
\[ \hat{\mu}_{i,K,(z,\mathbf{g})}(\*x) = R_K(\*x)^\top \hat{\bm{\theta}}_{i,K,(z,\mathbf{g})}, \]
where
$\hat{\bm{\theta}}_{i,K,(z,\mathbf{g})}$ is estimated from the ordinary least squares estimator, using the outcomes of the $i$-th units across all clusters $c$ that satisfy $Z_{c,i} = z$ and $\mathbf{G}_{c,i} = \mathbf{g}$. See Internet Appendix IA.A for the formula to obtain $\hat{ \bm{\theta}}_{i,K,(z,\mathbf{g})}$.
Intuitively, $\hat{\mu}_{i,K,(z,\mathbf{g})}(\*x)$ better approximates $\mu_{i,(z,\mathbf{g})}(\*x)$ as $K$ increases.
Similarly, in the sequence of estimators for $p_{i,(z,\mathbf{g})}(\*x)$, let $\hat{p}_{i,K,(z,\mathbf{g})}(\*x)$ be the estimator that uses $K$ approximation functions, i.e., (a possibly different) $R_K(\*x)$, and satisfies
\[\ln \frac{\hat{p}_{i,K,(z,\mathbf{g})}(\*x) }{\hat{p}_{i,K,(0,\bm{0})}(\*x) } = R_K(\*x)^\top \hat{\bm{\gamma}}_{i,K,(z,\mathbf{g})},\]
where $p_{i,(0,\bm{0})}(\*x)$ is chosen as the “pivot” for identification purposes, and $\hat{\bm{\gamma}}_{i,K,(z,\mathbf{g})} $ maximizes the log-likelihood function using the treatment assignments of the $i$-th units across all clusters $c$ that satisfy $Z_{c,i} \in \{z,0\}$ and $\mathbf{G}_{c,i} \in \{\mathbf{g},\bm{0}\}$. See Internet Appendix IA.A for the objective function for $\hat{\bm{\gamma}}_{i,K,(z,\mathbf{g})}$.
\begin{remark}[Alternative Estimators]
\normalfont
We focus on nonparametric series estimators which require few functional form assumptions on the propensity and outcome models, and enjoy estimation consistency properties.
In practice, one could consider alternative parametric or nonparametric estimators, such as matching, kernel regression, and random forests, for the propensity and outcome models. It is possible to generalize our results in Section (ref) to some of these alterantive estimators, as long as the estimated conditional outcome and propensity models satisfy certain rate conditions. In Internet Appendix IA.A, we discuss some parametric simplifications of the estimation problem. In particular, when $n$ or $m$ are large so that there is a large number of pairs of $(z,\mathbf{g})$, estimating a separate model $\hat p_{i,(z,\mathbf{g})}(\c(X))$ and $\hat{\mu}_{i,(z,\mathbf{g})}(\c(X))$ for each $(z,\mathbf{g})$ can be infeasible. In this case, one may consider a universal propensity model $p(\c(X), z, \mathbf{g})$ and conditional outcome model $\mu(\c(X), z, \mathbf{g})$ for all $i,z,\mathbf{g}$.
\end{remark}
\section{Main Asymptotic Results}
In this section, we show that our generalized AIPW estimators are doubly robust, asymptotically normal, and semiparametric efficient. For exposition, we present our results for ADE $\beta_j(\mathbf{g})$. The results for ASE ${\tau}_j(z,\mathbf{g}, \mathbf{g}^\prime)$ and general causal estimands $\psi_j(z,\mathbf{g}) - \psi_j(z^\prime,\mathbf{g}^\prime)$ are conceptually the same, and are provided in Corollary (ref) and Theorem (ref) in Appendix (ref). We first show that, if either the propensity or the outcome model is estimated from the sieve estimator in Section (ref) and standard regularity conditions (Assumption (ref) in Appendix (ref)) hold, then our AIPW estimators are consistent.
\begin{theorem}[Consistency, ADE]
Suppose Assumptions (ref)-(ref) hold.
As $M \rightarrow \infty$, for any $z$ and $\mathbf{g}$,
if either the estimated joint propensity $\hat{p}_{i,(z,\mathbf{g})}(\c(X))$ or the estimated outcome $\hat \mu_{i,(z,\mathbf{g})}(\c(X))$ is uniformly consistent in $\c(X)$, then the AIPW estimators are consistent, i.e.,
\begin{align}
\hat{\beta}^\mathrm{aipw}_j(\mathbf{g}) \xrightarrow{P} \beta_j(\mathbf{g}).
\end{align}
In particular, if at least one of $\hat p_{i,(z,\mathbf{g})}(\c(X))$ and $\hat \mu_{i,(z,\mathbf{g})}(\c(X))$ is estimated from the sieve estimators in Section (ref) and the regularity conditions in Assumption (ref) in Appendix (ref) hold, then Equation (ref) holds.
\end{theorem}
The key challenge in showing Theorem (ref) is that for any two units $i$ and $i^\prime$ in cluster $c$, $(X_{c,i}, Y_{c,i}, Z_{c,i})$ and $(X_{c,i^\prime}, Y_{c,i^\prime}, Z_{c,i^\prime})$ are correlated, and consequently the estimated scores of these two units, $\hat{\phi}_{c,i}(z, \mathbf{g})$ and $\hat{\phi}_{c,i^\prime}(z, \mathbf{g})$, used in the AIPW estimators are correlated. Therefore, the independence assumption of units, which is commonly used to show the doubly robust property robins1994estimation, is violated. However, from Assumption (ref), the correlation is limited to the units within a cluster, and for units $i$ and $i^\prime$ that are in two distinct clusters $c$ and $c^\prime$, $(X_{c,i}, Y_{c,i}, Z_{c,i})$ and $(X_{c^\prime,i^\prime}, Y_{c^\prime,i^\prime}, Z_{c^\prime,i^\prime})$ are independent. Using this property, we can show that the AIPW estimators are consistent even with correlated observations within a cluster, as long as the number of clusters $M$ grows to infinity.\footnote{Alternatively, we can show that AIPW estimators constructed from matching-based estimators $\hat{\mu}_{i,(z,\mathbf{g})}(\c(X))$ of ${\mu}_{i,(z,\mathbf{g})}(\c(X))$ and kernel regression estimators $\hat p_{i,(z,\mathbf{g})}(\c(X))$ of $p_{i,(z,\mathbf{g})}(\c(X))$ (see Section (ref) for details) are consistent, even though $\hat{\mu}_{i,(z,\mathbf{g})}(\c(X))$ and $\hat p_{i,(z,\mathbf{g})}(\c(X))$ are only (uniformly) asymptotically unbiased instead of consistent.}
Next we derive the asymptotic distribution of $\hat \beta^\mathrm{aipw}_j(\mathbf{g})$. Even though the correlation among units within a cluster does not affect consistency, it affects the asymptotic variance of $\hat \beta^\mathrm{aipw}_j(\mathbf{g})$. In Theorem (ref), we provide the semiparametric efficiency bound for $\beta_j(\mathbf{g})$ in the case of correlated observations, and we show that $\hat \beta^\mathrm{aipw}_j(\mathbf{g})$ is asymptotically normal and attains this efficiency bound.
\begin{theorem}[Asymptotic Normality and Semiparametric Efficiency, ADE]
Suppose Assumptions (ref)-(ref) hold,
then $\hat \beta^\mathrm{aipw}_j(\mathbf{g})$ is asymptotically normal. If, in addition, for any $c$,
\begin{equation}
Y_{c,i}(\s(Z), \n(\mathbf{Z})) \perp Y_{c,i^\prime}(Z_{c,i^\prime}, \mathbf{Z}_{c,(i^\prime)}) \mid \c(Z), \c(X) \qquad \forall i\neq i^\prime,
\end{equation}
then as $M \rightarrow \infty$, for any subset $j$ and neighbors' treatment $\mathbf{g}$, we have
\begin{equation*}
\sqrt{M}\big(\hat{\beta}^\mathrm{aipw}_j(\mathbf{g}) - {\beta}_j(\mathbf{g}) \big) \overset{d}{\rightarrow}\mathcal{N}\big(0,V_{j,\mathbf{g}}\big),
\end{equation*}
where $V_{j,\mathbf{g}}$ is the semiparametric efficiency bound for $\beta_{j}(\mathbf{g})$, and can be decomposed into
\[V_{j,\mathbf{g}} = V_{j,\mathbf{g},\mathrm{var}} + V_{j,\mathbf{g},\mathrm{cov}}.\]
The first term $V_{j,\mathbf{g},\mathrm{var}}$ is analogous to the classical efficiency bound and is defined as
\begin{align}
V_{j,\mathbf{g},\mathrm{var}} =& \frac{1}{|\mathcal{I}_j|^2}\sum_{i \in \mathcal{I}_j}\mathbb{E}\left[\frac{\sigma_{i,(1,\mathbf{g})}^{2}(\c(X))}{p_{i,(1,\mathbf{g})}(\c(X))}+\frac{\sigma_{i,(0,\mathbf{g})}^{2}(\c(X))}{p_{i,(0,\mathbf{g})}(\c(X))}+(\beta_{i,\mathbf{g}}(\c(X))-\beta_{i,\mathbf{g}})^{2}\right], \end{align}
and the second term $V_{j,\mathbf{g},\mathrm{cov}}$ is unique to our problem that quantifies the effect of interference on estimation efficiency, and is defined as \begin{align}
V_{j,\mathbf{g},\mathrm{cov}}=&\frac{1}{|\mathcal{I}_j|^2}\sum_{i, i^\prime \in \mathcal{I}_j, i \neq i^\prime }\mathbb{E}\left[(\beta_{i,\mathbf{g}}(\c(X))-\beta_{i,\mathbf{g}})(\beta_{i',\mathbf{g}}(\c(X))-\beta_{i',\mathbf{g}})\right],
\end{align}
where
$\sigma_{i,(z,\mathbf{g})}^2(\c(X)) = \mathrm{Var}\left[\s(Y)(z,\mathbf{g})\mid \c(X)\right]$, $\mu_{i,(z ,\mathbf{g})} = \+E[\s(Y)(z,\mathbf{g})]$, $\beta_{i,\mathbf{g}}(\c(X)) = \mu_{i,(1 ,\mathbf{g})}(\c(X)) - \mu_{i,(0 ,\mathbf{g})}(\c(X)) $ and $\beta_{i,\mathbf{g}} = \mu_{i,(1,\mathbf{g})} - \mu_{i,(0,\mathbf{g})}$.
\end{theorem}
The convergence rate of $\beta^\mathrm{aipw}_j(\mathbf{g})$ is $\sqrt{M}$ because there are $M$ independent clusters. As units within a cluster are dependent, we will show that a valid influence function should be defined at the \emph{cluster} level, rather than at the individual level as in the classical setting without interference hahn1998role. Theorem (ref) states that $V_{j,\mathbf{g}}$ can be decomposed into two terms, $V_{j,\mathbf{g},\mathrm{var}}$ and $V_{j,\mathbf{g},\mathrm{cov}}$. The term $V_{j,\mathbf{g},\mathrm{var}}$ scales with $1/|\mathcal{I}_j|$ and is equal to the average of
\[ \+E\left[\left(\phi_{c,i}(1,\mathbf{g}) - \phi_{c,i}(0,\mathbf{g}) - \beta_{i, \mathbf{g}}\right)^2\right] \]
over units $i \in\mathcal{I}_j$, and is analogous to the efficiency bound derived in hahn1998role and hirano2003efficient under SUTVA. Here $\phi_{c,i}(z,\mathbf{g}) $ is the score of unit $i$ in cluster $c$ and is the population version of $\hat{\phi}_{c,i}(z,\mathbf{g}) $ defined in Equation (ref).\footnote{$\phi_{c,i}(z,\mathbf{g}) $ equals to $\hat{\phi}_{c,i}(z,\mathbf{g}) $ with $\hat{p}_{i,(z,\mathbf{g})}(\c(X))$ and $\hat{\mu}_{i,(z ,\mathbf{g})}(\c(X))$ replaced by $p_{i,(z,\mathbf{g})}(\c(X))$ and $\mu_{i,(z ,\mathbf{g})}(\c(X))$.}
In contrast, the term $V_{j,\mathbf{g},\mathrm{cov}}$ is unique to our problem and comes from the the interference between units. $V_{j,\mathbf{g},\mathrm{cov}}$ scales with the average of
\[\sigma_{i,i^\prime} \coloneqq \+E\left[\left(\phi_{c,i}(1,\mathbf{g}) - \phi_{c,i}(0,\mathbf{g}) - \beta_{i, \mathbf{g}}\right) \left(\phi_{c,i^\prime}(1,\mathbf{g}) - \phi_{c,i^\prime}(0,\mathbf{g}) - \beta_{i^\prime, \mathbf{g}}\right)\right] \]
over any two distinct units $i$ and $i^\prime$ in $\mathcal{I}_j$, where the above term has the interpretation of the “covariance” between $i$ and $i^\prime$. From the expression of $V_{j,\mathbf{g},\mathrm{cov}}$ in Theorem (ref), $\sigma_{i,i^\prime}$ is equal to the covariance between the direct effects of $i$ and $i^\prime$ conditional on $\c(X)$, i.e., $\sigma_{i,i^\prime} = \+E\left[(\beta_{i,\mathbf{g}}(\c(X))-\beta_{i,\mathbf{g}})(\beta_{i',\mathbf{g}}(\c(X))-\beta_{i',\mathbf{g}})\right].$
If $\sigma_{i,i^\prime} = 0$ for all distinct $i$ and $i^\prime$, then $V_{j,\mathbf{g},\mathrm{cov}} = 0$ and $V_{j,\mathbf{g}} = V_{j,\mathbf{g},\mathrm{var}}$. Furthermore, if units within a cluster are i.i.d.
and we set $m = 1$, then $V_{j,\mathbf{g}}$ equals to
\[V_{j,\mathbf{g}} = V_{j,\mathbf{g},\mathrm{var}} = \frac{1}{n} \cdot \mathbb{E}\left[\frac{\sigma_{i,(1,\mathbf{g})}^{2}(\c(X))}{p_{i,(1,\mathbf{g})}(\c(X))}+\frac{\sigma_{i,(0,\mathbf{g})}^{2}(\c(X))}{p_{i,(0,\mathbf{g})}(\c(X))}+(\beta_{i,\mathbf{g}}(\c(X))-\beta_{i,\mathbf{g}})^{2}\right], \]
which is identical to the efficiency bound in hahn1998role divided by $n$. The factor $n$ adjusts the rate $\sqrt{N} = \sqrt{Mn}$ in hahn1998role to the rate $\sqrt{M}$ in Theorem (ref). At the other extreme, $V_{j,\mathbf{g},\mathrm{cov}}$ is maximized when the conditional direct effects for units in a cluster are perfectly correlated, i.e., $\beta_{i,\mathbf{g}}(\c(X)) = \beta_{i^\prime,\mathbf{g}}(\c(X))$, for all distinct $i$ and $ i^\prime$, but are not constant. In this case, the effective sample size is minimized at $M$, which is the least efficient case.
\begin{remark}[Conditional Independence]
\normalfont In Theorem (ref), the condition in (ref) states that outcomes of any two units in a cluster are independent \emph{conditional} on $\c(Z)$ and $\c(X)$. This condition is similar to the assumption of independent error terms in linear models. Essentially, it requires that the available covariates capture enough information about units so that no unobserved variables can cause correlations in outcomes between different units. In some applications, we may be concerned that there are unobserved variables that invalidate Equation (ref). In this case, Theorem (ref) can still hold, but with a more complicated form of $V_{j,\mathbf{g},\mathrm{cov}}$ containing the covariance between the residuals $\s(Y) - \mu_{i,(z,\mathbf{g})}(\*X_c)$ of two units. The propensity scores will then play a role in $V_{j,\mathbf{g},\mathrm{cov}}$.
\end{remark}
\begin{remark}[Heterogeneous Interference]
\normalfont
A main motivation of our work is the heterogeneity of interference, which is captured by the conditional exchangeability framework. In practice, an important consideration is how to specify the exchangeable subsets. Not surprisingly, a trade-off arises. A more granular partition of each cluster can capture more complicated heterogeneities and reduce bias, but could result in less efficient estimators:
\begin{center}
\begin{tabular}{ll}
\multirow{5}*{\rotatebox[origin=c]{270}{{ $\xRightarrow[\text{Efficiency Loss}]{\text{Reduced Bias}}$} }} & No Interference \\
& Partial Interference with Full Exchangeability \\ & Partial Interference with Conditional Exchangeability \\ & Network Interference with Conditional Exchangeability \\ & General Interference
\end{tabular}
\end{center}
We formalize this trade-off in Section (ref) and illustrate with a numerical example in Table (ref). Our characterization of the asymptotic variance of proposed estimators paves the way for data-driven selection of the appropriate interference structure by leveraging statistical tests on the heterogeneity of interference.
We discuss these ideas in detail in Appendix (ref). As statistical tests require the use of feasible and valid variance estimators, we also propose a matching-based variance estimator in Appendix (ref) that is consistent and performs well in simulation studies. Together, these results allow practitioners to assess the impacts of heterogeneous interference in a wide array of applications, from identifying effective targets of candidate policies to constructing interference-robust treatment effect estimators.
\end{remark}
So far, we have assumed that all clusters have the same size $n$. However, in many applications, such as those with family or classroom as a cluster, clusters may have different sizes. In Appendix (ref), we extend our framework and results to the setting with varying cluster sizes under a mixture model.
\section{Simulation Studies}
In this section, we demonstrate the finite sample properties and practical relevance of hypothesis testing based on our asymptotic results, and show that our AIPW estimators are robust to model mis-specifications.\footnote{Code for implementations of our estimators is available at https://github.com/freshtaste/CausalModel as part of an actively maintained package.}
We start by introducing the data generating process for the simulated data used in this section. We generate $M = 5,000$ clusters of size $n = 4$, i.e., $N=20,000$ units in total. Each cluster has two exchangeable subsets, $\mathcal{I}_1$ and $\mathcal{I}_2$, and each subset consists of 2 units. We generate covariates from $\s(X) \stackrel{\mathrm{i.i.d.}}{\sim} \mathcal{N}(0, 0.2)$ for all $c$ and $i$. The treatment variable $Z_{c,i}$ is randomly and independently sampled from a Bernoulli distribution:
\[ P(Z_{c,i}=1|\c(X)) = \frac{1}{1 + \exp(-0.5 \s(X) - 0.5/m \cdot \sum_{j=1}^m \bar{X}_{c,j}+1) } \qquad \text{ for all $c$ and $i \in \mathcal{I}_j$,} \]
where $\bar{X}_{c,j} = \frac{1}{|\mathcal{I}_j|} \sum_{i \in \mathcal{I}_j} X_{c,i}$ is the average covariate of units in subset $\mathcal{I}_j$ in cluster $c$.
We generate the outcomes from the following model:
\begin{align}
Y_{c,i}=\omega \cdot Z_{c,i}+ \left(f(\mathbf{G}_{c,i}) + \s(X) + \bar{X}_{c,1} \right) \cdot Z_{c,i}+ \s(X) + \bar{X}_{c,1} +\varepsilon_{c,i} \qquad \text{ for all $c$ and $i$,}
\end{align}
where
$\mathbf{G}_{c,i} = \left(\sum_{i^\prime\in \mathcal{I}_1,i^\prime \neq i} Z_{c,i^\prime}, \sum_{i^\prime\in \mathcal{I}_2,i^\prime \neq i} Z_{c,i^\prime} \right)$ is the number of treated neighbors in $\mathcal{I}_1$ and $\mathcal{I}_2$, and $\varepsilon_{c,i} \stackrel{\mathrm{i.i.d.}}{\sim} \mathcal{N}(0, 1)$ for all $c$ and $i$. In addition, $\omega \in \+R$ is a parameter that governs the level of direct treatment effect, and $f(\cdot,\cdot): \+Z^{2} \rightarrow \+R$ is an interference function that specifies how a unit's outcome is affected by its treated neighbors. We will consider various functional forms of $f(\cdot,\cdot)$. In this model, the direct and spillover effects of two subsets $\mathcal{I}_1$ and $\mathcal{I}_2$ are the same. Since $X_{c,i}$ has mean zero, the average direct effects equal to $\beta_1((g_1, g_2)) = \beta_2((g_1, g_2)) = \omega + f(g_1, g_2)$, and average spillover effects equal to $\tau_1(1, (g_1, g_2)) = \tau_2(1, (g_1, g_2)) = f(g_1, g_2)$ and $\tau_1(0, (g_1, g_2)) = \tau_2(0, (g_1, g_2)) = 0$.
\subsection{Inference and Hypothesis Tests of Treatment Effects}
We examine the finite sample properties of our treatment effect estimators, and the size and power of hypothesis tests for treatment effects. We present the results for $\beta_1\left((g_1, g_2)\right)$ to conserve space. The results for other estimands, e.g., $\beta_2\left((g_1, g_2)\right)$, $\tau_1\left(z,(g_1, g_2)\right)$ and $\tau_2\left(z,(g_1, g_2)\right)$, are similar. We consider three different interference mapping functions $f(\cdot, \cdot)$ in ((ref)) for parameter $\gamma \in \+R$:
\begin{enumerate}
• (\textbf{HO}) $f(\mathbf{g}) = \gamma \cdot (g_1 + g_2)$.
• (\textbf{HE1})
$f(\mathbf{g}) = \gamma \cdot (g_1 + g_2 + g_1 g_2)$.
• (\textbf{HE2}) $f(\mathbf{g}) = \gamma \cdot (g_1 + 2 \cdot g_2)$.
\end{enumerate}
For \textbf{HO} (homogeneous interference), the direct and spillover effects do not vary with which neighbors are treated, as long as the total number of treated neighbors $g_1 + g_2$ is the same. In this case, the specification of interference structure with two exchangeable subsets $\mathcal{I}_1$ and $\mathcal{I}_2$ is in fact more granular than needed, since $m=1$ satisfies Assumption (ref) as well. For both \textbf{HE1} and \textbf{HE2} (heterogeneous interference), the direct and spillover effects vary with which neighbors are treated. Specifically, the direct and spillover effects are different for $(g_1, g_2) = (0,2)$ and $(g_1, g_2) = (1,1)$ under \textbf{HE1} and \textbf{HE2}. In addition, the direct and spillover effects are also different for $(g_1, g_2) = (0,1)$ and $(g_1, g_2) = (1,0)$ under \textbf{HE2}. Under both \textbf{HE1} and \textbf{HE2}, the specification with the partition $\mathcal{I}_1$ and $\mathcal{I}_2$ is the most parsimonious one that satisfies Assumption (ref).
We will consider tests with the following null hypotheses:
\begin{enumerate*}
• $\mathcal{H}_0: \beta_1((0,0)) = 0$.
• $\mathcal{H}_0: \beta_1((0,1)) = 0$.
• $\mathcal{H}_0: \beta_1((0,0))=\beta_1((0,1))$.
• $\mathcal{H}_0: \beta_1((0,0))=\beta_1((0,2))$.
• $\mathcal{H}_0: \beta_1((0,1))=\beta_1((1,0))$.
• $\mathcal{H}_0: \beta_1((0,2)) = \beta_1((1,1))$.
• $\mathcal{H}_0: \beta_1((0,1)) = \beta_1((1,0))$ & $\beta_1((0,2)) = \beta_1((1,1))$.
\end{enumerate*}
For each hypothesis test, we first use our generalized AIPW estimators to estimate all the $\beta_1((g_1, g_2))$ parameters involved in the hypothesis test. For example, in the third hypothesis, we need to estimate both $\beta_1((0, 0))$ and $\beta_1((0, 1))$. In the generalized AIPW estimator, we estimate the individual propensity by logistic regression, neighborhood propensity by multinomial logistic regression, and the outcome by linear regression. We use the correct propensity and outcome model specifications. Using variance estimators based on Theorems (ref) and (ref), we conduct the hypothesis test and report the rejection probability in Table (ref) for various values of treatment effect parameters $\omega$ and $\gamma$ under $K = 2,000$ Monte Carlo trials.\footnote{Details on variance estimators and hypothesis testings are provided in Section (ref) and (ref).}
In Table (ref), we find that under the null hypotheses, the rejection probability is close to the nominal level of $\alpha = 0.05$. This finding is consistent across all cases where the null hypotheses hold. Specifically, the null hypotheses are true in the following cases: (1) If $\omega = \gamma = 0$, then $\beta_1((g_1, g_2)) = 0$ for all $g_1$ and $g_2$ in \textbf{HO}, \textbf{HE1}, and \textbf{HE2}. All the seven null hypotheses are true; (2) If $\omega = 1$ and $\gamma = 0$, then $\beta_1((g_1, g_2)) = 1$ for all $g_1$ and $g_2$ in \textbf{HO}, \textbf{HE1}, and \textbf{HE2}. The third to seventh null hypotheses are true; (3) If $\omega = \gamma = 1$, then the fifth to seventh null hypotheses are true for \textbf{HO}, and the fifth null hypothesis is true for \textbf{HE1}. Additionally, we verify the asymptotic normality in Theorem (ref) with the histograms provided in Internet Appendix IA.D.
\begin{table}[t]
\captionsetup{position=top, font=normalsize, labelfont=bf, textfont=normalfont, justification=centering, margin=0mm, aboveskip=1mm, belowskip=0mm, labelsep=colon, singlelinecheck=false}\caption{Rejection probabilities of hypothesis tests}
\begin{adjustbox}{max width=\linewidth,center}
\begin{tabular}{c ccc c ccc c ccc}
\toprule
& \multicolumn{3}{c}{$\omega = \gamma = 0$} & & \multicolumn{3}{c}{$\omega=1, \gamma=0$} & & \multicolumn{3}{c}{$\omega=\gamma=1$}\\ \cmidrule{2-4} \cmidrule{6-8} \cmidrule{10-12}
$\mathcal{H}_0$ & \textbf{HO} & \textbf{HE1} & \textbf{HE2} & &\textbf{HO} & \textbf{HE1} & \textbf{HE2} & & \textbf{HO} & \textbf{HE1} & \textbf{HE2} \tabularnewline
\midrule
1 & 0.0445 & 0.0365 & 0.0360 & & 0.9740 & 0.9690 & 0.9715 & & 0.9660 & 0.9550 & 0.9655 \\
2 & 0.0470 & 0.0535 & 0.0500 & & 0.9620 & 0.9705 & 0.9705 & & 1.0000 & 1.0000 & 1.0000 \\
3 & 0.0435 & 0.0440 & 0.0420 & & 0.0525 & 0.0490 & 0.0560 & & 0.7860 & 0.7210 & 1.0000 \\
4 & 0.0455 & 0.0530 & 0.0530 & & 0.0520 & 0.0420 & 0.0555 & & 0.9995 & 0.9995 & 1.0000 \\
5 & 0.0510 & 0.0415 & 0.0440 & & 0.0445 & 0.0450 & 0.0450 & & 0.0395 & 0.0330 & 0.7845 \\
6 & 0.0530 & 0.0410 & 0.0435 & & 0.0460 & 0.0415 & 0.0425 & & 0.0520 & 0.7085 & 0.7570 \\
7 & 0.0440 & 0.0455 & 0.0500 & & 0.0575 & 0.0395 & 0.0530 & & 0.0390 & 0.7460 & 0.9670 \\
\bottomrule
\addlinespace
\end{tabular}
\end{adjustbox}
\captionsetup{position=bottom, font=footnotesize, textfont=normalfont, margin=1mm, skip=2mm, justification=justified, singlelinecheck=false}\caption*{Rejection probabilities are calculated using 5,000 clusters of size four and $K = 2,000$ Monte Carlo simulations. Significance level is set to be 0.05. }
\end{table}
\subsection{Robustness Properties of Our Estimators}
In this section, we compare the performance of our AIPW estimators with two common alternative estimators in estimating treatment effects:
\begin{enumerate}
• Ordinary least squares (OLS): Run the following linear regression
\[Y_{c,i} = \alpha + \theta_z \cdot Z_{c,i} + \bm{\theta}^\top_x \cdot \*X_c + \left( \theta_{zg1} \cdot G_{c,i,1} + \theta_{zg2} \cdot G_{c,i,2} + \bm{\theta}_{zx}^\top \cdot \*X_c \right)\cdot Z_{c,i} + \varepsilon_{c,i}, \]
where $\*X_c=\left(X_{c,i}, \bar{X}_{c,1}\right)$. Then estimate $\beta_1((g_1, g_2))$ by $\hat{\theta}_{zg1} \cdot g_1 + \hat{\theta}_{zg2} \cdot g_2$, where $\hat{\theta}_{zg1} $ and $\hat{\theta}_{zg2}$ are the estimated coefficients from the above regression.
• Orthogonal Random Forest (ORF) for CATE: Suppose SUTVA holds and let $Y_{c,i}(z)$ be the potential outcomes. Treat $\mathbf{G}_{c,i}$ as additional covariates and apply off-the-shelf forest based methods.\footnote{The CATE estimator is implemented in the python package \textsc{econml} by econml. Among all the available CATE estimators, we adopt DROrthoForest, which is the AIPW-based CATE estimator chernozhukov2018double.} Then use the fitted forest to estimate $\+E[Y_{c,i}(1) - Y_{c,i}(0)\mid G_{c,i,1} = g_1, G_{c,i,2} = g_2]$, which can be viewed as an approximation of $\beta_1((g_1, g_2))$.
\end{enumerate}
We draw simulated data $K = 1,000$ times and estimate $\beta_1((g_1, g_2))$ on each simulated data set using three different estimators. We evaluate the performance of these estimators using two metrics: mean-squared error (MSE) and coverage rate. As shown in Table (ref), our AIPW estimators consistently exhibit much smaller MSEs compared to both OLS and ORF, suggesting that our AIPW estimators can most accurately estimate $\beta_1((g_1, g_2))$ for various interference functions $f(\mathbf{g})$, even under misspecifications of the outcome model. Moreover, ORF has a much smaller MSE than OLS, indicating that in the presence of interference, tree-based methods, with their greater flexibility, tend to perform better than the more restrictive regression-based methods.
Furthermore, as shown in Table (ref), our AIPW estimators achieve the correct coverage rate (i.e., close to 95%) for various specifications of $f(\mathbf{g})$, while OLS and ORF do not. OLS tends to have a lower coverage rate when $f(\mathbf{g})$ is nonlinear (e.g., quadratic, reciprocal). This low coverage rate is likely due to its failure to obtain an accurate point estimate of the direct effects. On the other hand, ORF tends to have a higher coverage rate than the nominal rate. In fact, the coverage rate of ORF is very close to 100%, implying that the estimated confidence interval from tree-based methods can be too wide, making hypothesis tests using tree-based methods overly conservative.
Importantly, the outcome model in our AIPW estimators is misspecified when $f((g_1, g_2))$ is not linear in $g_1$ and $g_2$. However, even in such cases, our AIPW estimators can outperform OLS and ORF, thanks to the double robustness property of AIPW estimators. Therefore, we suggest that explicitly modeling the interference structure and using a doubly robust estimator can be crucial for accurate estimation and valid inference of treatment effects.
\begin{table}[t]
\captionsetup{position=top, font=normalsize, labelfont=bf, textfont=normalfont, justification=centering, margin=0mm, aboveskip=1mm, belowskip=0mm, labelsep=colon, singlelinecheck=false}\caption{Coverage rate and MSE of various estimators for $\beta_1((g_1,g_2))$}
{
\begin{tabular}{c ccc c ccc}
\toprule
& \multicolumn{3}{c}{Coverage Rate} & & \multicolumn{3}{c}{MSE} \\ \cline{2-4} \cline{6-8}
$f(\mathbf{g})$ & OLS & ORF & AIPW & & OLS & ORF & AIPW \tabularnewline
\midrule
$g_1 + 2 g_2$ & 94.52% & 98.93% & 95.60% & & 0.0003 & 0.0027 & 0.0038\tabularnewline
$\sqrt{g_1 + 2 g_2}$ & 14.37% & 99.07% & 95.08% & & 0.0482 & 0.0027 & 0.0039 \tabularnewline
$1/(g_1 + 2 g_2+1)$ & 15.35% & 98.97% & 94.88% & & 0.0191 & 0.0027 & 0.0039 \tabularnewline
$0.1 (g_1 + 2 g_2)^2 + (g_1 + 2 g_2)$& 0.47% & 99.18% & 95.53% & & 0.0627 & 0.0027 & 0.0038 \tabularnewline
\bottomrule
\addlinespace
\end{tabular}
}
\captionsetup{position=bottom, font=footnotesize, textfont=normalfont, margin=1mm, skip=2mm, justification=justified, singlelinecheck=false}\caption*{We report the average coverage rate and MSE of $\hat{\beta}((g_1,g_2))$ over $(g_1, g_2) = (0,0), (0,1), (0,2), (1,0), (1,1)$ and $(1,2)$. We run $K=1,000$ Monte Carlo simulations for OLS, ORF, and AIPW. }
\end{table}
\section{Two Applications to the Add Health Dataset}
In this section, we demonstrate our methods through two empirical applications using the National Longitudinal Study of Adolescent to Adult Health (Add Health) dataset harris2009national. The Add Health data has been frequently used in methodological and empirical studies on peer effects and interference because of its rich information on respondents' social and familial connections (e.g., bramoulle2009identification,goldsmith2013social,swisher2015paternal,forastiere2020identification).
In the first application, we investigate the effect of alcohol consumption on academic performance. We construct direct effect estimators under various specifications of the interference structure and find negative effects of regular alcohol consumption on students' academic performance. This finding is robust across different specifications of interference structures. However, the confidence intervals are wider under more complex interference structures due to finite sample efficiency loss. In the second application, we use our spillover effect estimators to study the impact of parental incarceration on adolescent well-being and find heterogeneous effects in the gender of the incarcerated parent.
\subsection{Alcohol Consumption and Academic Performance}
Prior studies have found negative associations between alcohol use and academic performance mcgrath1999academic,jeynes2002relationship,diego2003academic. Other works have also studied peer effects of alcohol use among friends clark2007wasn,eisenberg2014peer. In this paper, we adjust for the peer effects of alcohol use through the joint propensity model of alcohol use and estimate the effect of alcohol use on academic performance.
We use the Add Health data to construct a cohort of clusters, each with a small number of adolescents. Then we apply our average direct effect (ADE) estimators under various interference specifications to this cohort. We construct the cohort using adolescents' nominations of close friends. For each adolescent of our interest (which we refer to as the \textbf{centroid} of a cluster), we construct a cluster that consists of this adolescent, his/her best female friend, and his/her best male friend. Therefore each cluster has a size of three. We also perform sub-sampling to ensure minimal overlaps among clusters and obtain 7905 clusters in total. We define regular alcohol consumption as drinking at least once or twice a week. We measure academic performance by achieving a grade of B or better in mathematics, although the estimates are similar for other subjects.
We consider three interference specifications corresponding to the top three levels of interference structures in Section (ref). The first specification assumes that there is no interference, and thus an individual's alcohol use and academic performance are not affected by his/her friends' alcohol use. The second specification assumes that there is homogeneous interference, so that an adolescent's alcohol use and academic performance are allowed to be affected interchangeably by their two best friends, regardless of the friends' genders. The third specification assumes interference is heterogeneous, so that a best friend's influence on an individual's alcohol use and academic performance is potentially gender-dependent.
For each specification, we estimate the average direct effect of alcohol consumption on academic performance for centroids, using the corresponding AIPW estimator, where the propensity and outcome models are estimated under the corresponding interference structures. We adjust for covariates of both the centroids and their friends, including age, gender, frequency of skipping classes or missing school, and parents' educational backgrounds.
\begin{table}[t!]
\captionsetup{position=top, font=normalsize, labelfont=bf, textfont=normalfont, justification=centering, margin=0mm, aboveskip=1mm, belowskip=0mm, labelsep=colon, singlelinecheck=false}\caption{AIPW estimators of direct effects under different interference specifications}
{6pt}
\begin{adjustbox}{max width=\linewidth,center}
\begin{tabular}{l|c|c|c|c|c|c|c|c}
\toprule
specification& \multicolumn{8}{c}{$\hat\beta$} \tabularnewline
\midrule
no interference & \multicolumn{8}{c}{-0.049***}
\\
& \multicolumn{8}{c}{(0.010)}
\\
\midrule
& \multicolumn{4}{c|}{$\hat\beta_\mathrm{m}$} & \multicolumn{4}{c}{$\hat\beta_\mathrm{f}$} \tabularnewline
\midrule
no interference
& \multicolumn{4}{c|}{-0.054***} & \multicolumn{4}{c}{-0.044***}
\\
& \multicolumn{4}{c|}{(0.013)} & \multicolumn{4}{c}{(0.015)}
\\
\midrule \midrule
& \multicolumn{2}{c|}{$\hat\beta(0)$} & \multicolumn{4}{c|}{$\hat\beta(1)$} & \multicolumn{2}{c}{$\hat\beta(2)$} \tabularnewline
\midrule
hom. interference & \multicolumn{2}{c|}{-0.073***} & \multicolumn{4}{c|}{-0.065***} & \multicolumn{2}{c}{-0.006}
\\
& \multicolumn{2}{c|}{(0.013)} & \multicolumn{4}{c|}{(0.020)} & \multicolumn{2}{c}{(0.068)}
\\
\midrule
& $\hat\beta_\mathrm{m}(0)$ & $\hat\beta_\mathrm{f}(0)$ & \multicolumn{2}{c|}{$\hat\beta_\mathrm{m}(1)$} & \multicolumn{2}{c|}{$\hat\beta_\mathrm{f}(1)$} & $\hat\beta_\mathrm{m}(2)$ & $\hat\beta_\mathrm{f}(2)$
\tabularnewline
\midrule
hom. interference & -0.062*** & -0.084*** & \multicolumn{2}{c|}{-0.106***} & \multicolumn{2}{c|}{ -0.025} & -0.002 & -0.010
\\
& (0.018) & (0.020) & \multicolumn{2}{c|}{(0.027)} & \multicolumn{2}{c|}{(0.031)} & (0.091) & (0.102)
\\
\midrule \midrule
& $\hat\beta_\mathrm{m}(0,0)$ & $\hat\beta_\mathrm{f}(0,0)$ &
$\hat\beta_\mathrm{m}(1,0)$ &
$\hat\beta_\mathrm{m}(0,1)$ &
$\hat\beta_\mathrm{f}(1,0)$ & $\hat\beta_\mathrm{f}(0,1)$ &
$\hat\beta_\mathrm{m}(1,1)$ &
$\hat\beta_\mathrm{f}(1,1)$
\tabularnewline
\midrule
het. interference & -0.062***& -0.084*** & -0.053& -0.144***& -0.102**& 0.105 & -0.002 & -0.005
\\
&(0.018) & (0.019) & (0.033) & (0.042)& (0.037)& (0.053) & (0.098) & (0.117) \tabularnewline
\bottomrule
\end{tabular}
\end{adjustbox}
\captionsetup{position=bottom, font=footnotesize, textfont=normalfont, margin=1mm, skip=2mm, justification=justified, singlelinecheck=false}\caption*{Under the specification of no interference, $\beta$ is the ATE for all centroids, and $\beta_\mathrm{m}$ is the ATE for male centroids, and similarly for $\beta_\mathrm{f}$. Under the specification of homogeneous interference, the direct effect is defined as $\beta(g)$, where $g \in \{0,1,2\}$ denotes the number of treated friends.
Under the specification of heterogeneous interference, the direct effects are defined as $\beta_\mathrm{m}(z_{\mathrm{mf}},z_{\mathrm{ff}})$ for male centroids and $\beta_\mathrm{f}(z_{\mathrm{mf}},z_{\mathrm{ff}})$ for female centroids, where $z_{\mathrm{mf}}\in \{0,1\}$ denotes the treatment status of male friend and $ z_{\mathrm{ff}} \in \{0,1\}$ denotes the treatment status of female friend. For each specification, covariates are adjusted in the propensity and outcome models.
}
\end{table}
In Table (ref), we report the estimated direct effects under different interference structures. There are three main findings. First, all estimated treatment effects are negative, implying that regular alcohol use could have a negative effect on academic performance, which is robust to the particular specification of interference structures. Second, the magnitude of direct effects varies with the gender of centroids and their numbers of treated friends. The impact is diminished with an increase in the number of friends who drink regularly. Furthermore, direct effects are also heterogeneous in the \emph{gender} of treated friends, suggesting that interference could be heterogeneous. Third, standard errors increase with the complexity of the interference structure. There are two potential contributing factors. First, our bias-variance tradeoff analysis in Section (ref) suggests that estimators based on more complex specifications of the interference structure tend to have larger \emph{asymptotic} variances. Second, fewer samples are available when estimating each heterogeneous effect, highlighting the cost of a complex specification of interference in practice.
\subsection{Paternal Incarceration and Adolescent Well-Being}
The impact of parental incarceration on children's health, education, and economic outcomes is an important topic that has generated much attention in empirical works lee2013impact,miller2015association,wildeman2018parental,austin2022parental,jones2022mental. Following the empirical study of swisher2015paternal, we apply the spillover effect estimators from (ref) to examine the impact of paternal incarceration on children's well-being, specifically delinquency and depression (both binary outcomes). In this context, families naturally form independent clusters, with heterogeneous groups within each family comprising of the father, mother, and children. We consider three spillover effects of parental incarceration prior to measuring the outcome: only the mother incarcerated, only the father incarcerated, and both parents incarcerated ($\hat\tau(1,0)$, $\hat\tau(0,1)$, and $\hat\tau(1,1)$). The set of control variables includes age, gender, ethnicity, physical abuse, and sexual abuse.\footnote{Outcomes and covariates are obtained from Wave I while treatments are obtained from Wave IV questionnaires.} The sample size is 4692. We compare our estimators to an OLS regression similar to the original study by regressing outcomes on the three incarceration indicators and covariates. In Table (ref), we present the OLS and AIPW estimators for the spillover effects. Both estimators yield similar point estimates for delinquency across various spillover exposures. However, the standard errors associated with our AIPW estimators are consistently smaller than those of the OLS estimators, demonstrating the efficiency gain achieved with our method. In addition, our estimators reveal a larger effect on depression when the father is incarcerated compared to the OLS estimators.
\begin{table}[h]
\captionsetup{position=top, font=normalsize, labelfont=bf, textfont=normalfont, justification=centering, margin=0mm, aboveskip=1mm, belowskip=0mm, labelsep=colon, singlelinecheck=false}\caption{OLS and AIPW estimators of spillover effects}
\begin{tabular}{lccccccc}
\toprule
& \multicolumn{3}{c}{delinquency} & & \multicolumn{3}{c}{depression} \\
\cline{2-4} \cline{6-8}
& $\hat\tau(1,0)$ & $\hat\tau(0,1)$ & \multicolumn{2}{l}{$\hat\tau(1,1)$} & $\hat\tau(1,0)$ & $\hat\tau(0,1)$ & $\hat\tau(1,1)$ \\ \midrule
OLS & 0.1892*** & 0.1099*** & 0.2142*** & & -0.0013 & 0.0363 & 0.1008 \\
& (0.059) & (0.021) & (0.062) & & (0.056) & (0.021) & (0.060) \\ \\
AIPW & 0.1854*** & 0.1177*** & 0.1993*** & & 0.0077 & 0.0424*** & 0.1728*** \\
& (0.037) & (0.013) & (0.040) & & (0.037) & (0.012) & (0.039) \\
\bottomrule
\end{tabular}
\end{table}
\section{Concluding Remarks}
In this paper, we propose to explicitly model heterogeneous interference in the observational setting relevant in many empirical works through a conditional exchangeability framework. While this framework is an instance of the more general exposure mapping framework, it is applicable to many applications where heterogeneities in interference and treatment effects are determined by observable characteristics. We construct doubly robust and semiparametric efficient AIPW estimators for granular average direct and spillover effects based on the conditional exchangeability framework. Our asymptotic results provide off-the-shelf estimation and inference methods for various types of causal estimands relevant in practice, such as optimal policy targeting. We also propose a data-driven method based on hypothesis testing that allows practitioners to detect and account for heterogeneities in interference without relying heavily on domain knowledge. We demonstrate the validity and practical appeal of our estimators through extensive simulation studies and two relevant applications to the Add Health dataset. Lastly, our work is also relevant for researchers interested in estimating aggregate treatment effects, such as analogs of classical ATEs in the presence of potential interference. By aggregating our AIPW estimators for granular effects, one can construct interference-robust estimators at the cost of some efficiency loss.
\onehalfspacing
{
Acknowledgements
We are very grateful for many helpful comments and suggestions from Eric Auerbach, Jianfei Cao, Park Chan, Juan Estrada, Laura Forastiere, Hong Han, Yuchen Hu, Hyungseok Kang, Yongchan Kwon, Michael Leung, Haodong Li, Fabrizia Mealli, Markus Pelger, and Michael Pollmann, Evan Rose, Brad Ross, Martin Rotemberg, Rose Tan, Merrill Warnick, Jason Weitze, Keli Xu, Johan Ugander, and participants of the Online Causal Inference Seminar, NASMES 2022, California Econometrics Conference, the Stanford econometrics lunch seminar, and the LinkedIn Tech Talk series. We would like to thank Joshua Quan and Christopher Fraga at the Stanford Institute for Research in the Social Sciences for their assistance in accessing the Add Health dataset.
This research uses data from Add Health, a program project directed by Kathleen Mullan Harris and designed by J. Richard Udry, Peter S. Bearman, and Kathleen Mullan Harris at the University of North Carolina at Chapel Hill, and funded by grant P01-HD31921 from the Eunice Kennedy Shriver National Institute of Child Health and Human Development, with cooperative funding from 23 other federal agencies and foundations. Special acknowledgment is due Ronald R. Rindfuss and Barbara Entwisle for assistance in the original design. Information on how to obtain the Add Health data files is available on the Add Health website. No direct support was received from grant P01-HD31921 for this analysis.
}