EconBase
← Back to paper

Learning and Testing Exposure Mappings of Interference using Graph Convolutional Autoencoder

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

52,583 characters · 7 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Learning and Testing Exposure Mappings of Interference using Graph Convolutional Autoencoder

\thispagestyle{empty}

abstractInterference or spillover effects arise when an individual's outcome (e.g., health) is influenced not only by their own treatment (e.g., vaccination) but also by the treatment of others, creating challenges for evaluating treatment effects. Exposure mappings provide a framework to study such interference by explicitly modeling how the treatment statuses of contacts within an individual's network affect their outcome. Most existing research relies on a priori exposure mappings of limited complexity, which may fail to capture the full range of interference effects. In contrast, this study applies a graph convolutional autoencoder to learn exposure mappings in a data-driven way, which exploit dependencies and relations within a network to more accurately capture interference effects. As our main contribution, we introduce a machine learning-based test for the validity of exposure mappings and thus test the identification of the direct effect. In this testing approach, the learned exposure mapping is used as an instrument to test the validity of a simple, user-defined exposure mapping. The test leverages the fact that, if the user-defined exposure mapping is valid (so that all interference operates through it), then the learned exposure mapping is statistically independent of any individual's outcome, conditional on the user-defined exposure mapping. We assess the finite-sample performance of this proposed validity test through a simulation study.

Keywords: causal inference, network interference, graph convolutional autoencoder, conditional independence, instrumental variables

\setcounter{page}{1}

Introduction

Interference or spillover effects, where an individual’s outcome (e.g., health) is influenced not only by their own treatment (e.g., vaccination) but also by the treatment of others, pose substantial challenges for causal inference, particularly when arbitrary forms of interference are allowed; see Manski2013. Exposure mappings, as discussed in AronowSamii2017 and eckles2017design, provide a structured framework to study such interference by explicitly modeling the mechanisms through which the treatment statuses of social contacts within an individual's network influence their outcome. For example, an exposure mapping might specify that an interference effect occurs if at least one contact is treated, but does not depend on the number of treated contacts. By defining such mappings, researchers can separate interference effects from the direct effect of a treatment, provided that they appropriately account for differences in the probabilities of specific exposure mappings across individuals - for instance, due to variation in network structure. However, most existing research relies on exposure mappings of limited complexity that are a priori defined by the researcher in an ad hoc manner, which risks failing to capture the full range of interference effects. In this paper, we apply graph convolutional autoencoder (GCAs), which are a specific type of graph neural networks (GNNs), to learn exposure mappings in a data-driven way based on network embeddings. By exploiting network dependencies and relations, GCAs allow us to capture complex patterns of interference that may be missed by simple, a priori defined mappings. As our main contribution, we introduce a machine learning-based testing framework for the validity of exposure mappings. Our approach uses a learned, complex exposure mapping as an instrument to assess whether a given simpler, researcher-defined exposure mapping sufficiently captures all interference. Specifically, if the simpler mapping accurately captures all interference effects such that any interference operates through them, then the learned mapping should be statistically independent of the individual's outcome, conditional on the less complex mapping. By examining violations of this condition, we can assess whether less complex, researcher-defined mappings fail to capture some of the underlying interference. Although developed in the context of interference, our testing framework is related to conditional independence testing approaches in deLunaJohansson2012 and huberkueck2022, who investigate tests for joint satisfaction of conditional treatment exogeneity and instrumental variable (IV) assumptions. Analogously, we test whether the user-defined exposure mapping is both exogenous and sufficient to capture all interference effects. If the exposure mapping is correctly specified, then the treatment assignments of others affect an individual's outcome only through this mapping, implying an IV type exclusion restriction. In this case, the learned embedding should be conditionally independent of the outcome given the user-defined exposure mapping, which forms the basis of our test. The testing approach is also related to the causal framework of yao2025third, who discuss how causal representation learning (such as the learning of exposure mappings from networks) can be interpreted within a measurement model perspective, where the learned representations are viewed as proxy measurements of latent causal variables (e.g., the entire interference network).

We base the estimation of the direct effect of the treatment on the outcome, as well as the proposed testing procedure, on the double machine learning (DML) framework of Chetal2018, which builds on doubly robust score functions; see Robins+94 and RoRo95. The DML framework satisfies the Neyman1959-orthogonality condition, which implies that the resulting estimators and tests are relatively insensitive to modest approximation errors in the estimation of the exposure mapping and in the estimation of the treatment or outcome models. Consequently, the estimators and tests satisfy asymptotic normality under specific regularity conditions, in particular if the machine learning methods used to estimate these models, such as GNNs, converge at a rate of $o(n^{-1/4})$ to the respective true model. Our study contributes to the growing literature that seeks to construct exposure mappings by exploiting information contained in the data. For instance, bargaglistoffi2023heterogeneoustreatmentspillovereffects propose a tree-based method to assess heterogeneity in direct treatment and interference effects with respect to individual, neighborhood, and network characteristics. In their framework, it is assumed that interference occurs only within, but not across, predefined clusters (e.g., geographic regions) - a setting known as partial interference; see, e.g., Sobel2006, HongRaudenbush2006, and HudgensHalloran2008. However, interference may still vary depending on the network structure within clusters. The proposed method therefore aims at detecting network structures and classifying clusters that exhibit similar interference effects based on tree-based algorithms. In contrast, our method is not confined to learning exposure mappings within clusters. Instead, we learn exposure mappings as embeddings in a GCA, which constitutes a highly flexible approach capable of capturing complex network dependencies. Furthermore, we complement this estimation strategy with a novel statistical testing procedure to validate exposure mappings. pmlr-v130-ma21c also propose a GNN-based approach to learn interference from the data, focusing on the problem of learning optimal treatment policies across subgroups (see, e.g., Manski2004, HiranoPorter2008, KitagawaTetenov2018, AtheyWager2018) in the presence of interference. In contrast, the focus of our study is not on optimal policy learning but rather testing whether the exposure mapping is correctly specified which allows the identification of the direct treatment effect. NEURIPS2019_af1c25e8 consider the estimation of direct treatment effects while controlling for network-related confounders by conditioning on embeddings learned from networks using DML. To this end, they adapt the DML assumptions in Chetal2018 to account for learned embeddings and present high-level conditions under which estimation of the direct treatment effect is $\sqrt{n}$-consistent. Also 10.1145/3534678.3539299 discuss direct effect estimation when accounting for confounding induced by network interference. More closely related to our setting, leung2022graph propose using GNNs to control for network-induced confounding, with the goal of estimating both direct and interference effects and conducting statistical inference. To this end, they rely on the assumption of approximate neighborhood interference (ANI) introduced by leung2022causal, which is conceptually related to approximate sparsity as considered in lasso regression by Bellonietal2014. ANI posits that interference decays sufficiently fast with the distance between individuals in the network, thereby addressing the problem of potentially high-dimensional confounding induced by network structure. leung2022graph show that, under ANI and additional regularity conditions, the estimation of propensity score and outcome models can achieve convergence at a rate of $o(n^{-1/4})$, implying that direct and interference effect estimators based on the given exposure mappings can be $\sqrt{n}$-consistent and asymptotically normal when implemented within the DML framework. Importantly, the exposure mapping is not learned but prespecified by the researcher and the GNNs are only used to estimate the nuisance functions. These results demonstrate that some restriction on the complexity of interference is necessary to obtain well-behaved estimators. Specifically, the depth of the relevant interference network must not be too large, which allows GNNs with a limited number of layers to approximate the interference structure. baharan2025graph also propose a graph-based doubly robust estimator that uses graph neural networks to flexibly learn network confounding and to estimate both direct and interference effects using a prespecified exposure mapping. An alternative strategy for reducing complexity is proposed by belloni2022neighborhood, who permit the depth of the relevant interference network to vary across individuals but, in turn, impose additive separability of the direct and interference effects. Another relevant study in this context is wang2024graph, who - albeit not focusing specifically on estimation of interference effects via exposure mappings - establish convergence rates for GNN estimators and derive high-level conditions under which the complexity of network interference (and the approximation error when estimating it) are sufficiently small to attain $o(n^{-1/4})$ rates. The issue of estimating exposure mappings is also related to the framework of DML estimation with generated (rather than directly observed) regressors, as considered, for instance, in models with estimated control functions; see e.g. pan2024locally and escanciano2023automatic. Our study also applies the DML methodology in the context of interference and GNNs, but extends this line of research by proposing a statistical test to assess the validity of the exposure mappings.

The remainder of this study is organized as follows. Section (ref) introduces the causal framework, including the concept of exposure mappings, the causal parameters of interest - namely the direct, interference and total effects - and the identifying assumptions. Section (ref) describes how exposure mappings can be learned from the data using GCAs. Section (ref) presents the testability conditions for exposure mappings and outlines the implementation of corresponding test using DML. Section (ref) details the DML-based estimation of the causal parameters under correct specification of the exposure mapping. Section (ref) presents a simulation study investigating the finite-sample performance of the estimators of causal effects, as well as of the testing procedure. Section (ref) concludes.

Causal effects and identification

This section introduces the concept of exposure mappings, as well as the definition and identification of direct, interference and total effects under specific assumptions. Let $D_i$ denote the treatment of individual $i$ in a population of interest, and let $\mathcal{D}_{-i}$ represent the vector of treatments assigned to all other individuals (excluding individual $i$) in that population. Furthermore, let $Y_i$ denote the observed outcome. Throughout, we use uppercase letters to denote random variables and lowercase letters for specific realizations. Using the potential outcomes framework, as proposed by Neyman23 and advocated by Rubin74, the potential outcome of individual $i$ under specific treatment assignments $D_i = d$ and $\mathcal{D}_{-i} = \mathbf{d}$ is written as $Y_i(d, \mathbf{d})$. This contrasts with the standard assumption in most treatment evaluations, which impose the Stable Unit Treatment Value Assumption (SUTVA) Rubin80, Cox58. SUTVA rules out interference effects, implying that the potential outcome depends only on individual $i$'s own treatment: $Y_i(d)$. For the subsequent discussion, we assume a binary treatment, such that $d \in \{0,1\}$.

When SUTVA is violated, one pathway to identifying direct, interference, and total effects of an individual's own treatment relies on exposure mappings, see AronowSamii2017. Such mappings impose structure on how the treatment assignments of other individuals, $\mathcal{D}_{-i}$, influence the outcome of individual $i$. It is assumed that interference effects operate through an individual's social network, denoted by $\mathcal{A}$, which is observed - for example, as an adjacency matrix indicating which individuals interact with one another.\footnote{The network \(\mathcal{A}\) is typically assumed to be fixed. However, li2022random consider the network structure in a population as a random draw and propose an asymptotic framework for constructing confidence intervals for direct and interference effects under this assumption.} Exposure mappings can be viewed as sufficient statistics for capturing any interference effects within a network. More formally, the exposure for individual $i$, denoted by $Z_i$, defines strengths of interference as a function (or mapping) $\mathcal{F}$ of $i$’s network $\mathcal{A}$, the treatment assignments of other individuals, $\mathcal{D}_{-i}$, and the covariates of other individuals $\mathcal{X}_{-i}$:

eqnarray[eqnarray omitted — 98 chars of source]

The complexity of interference captured by the function $\mathcal{F}$ determines the number of possible values $Z_i$ can take. A common choice in the literature defines $\mathcal{F}$ as the number of individuals who are both treated according to $\mathcal{D}_{-i}$ and neighbors of individual $i$ in the network $\mathcal{A}$, in which case $Z_i$ takes values $z \in \{0,1,2,\ldots\}$. A simpler alternative specifies $\mathcal{F}$ as a binary indicator, where $Z_i = 1$ if at least one treated individual in $\mathcal{D}_{-i}$ is a neighbor of individual $i$ in network $\mathcal{A}$, and $Z_i = 0$ otherwise. This results in a binary exposure: $z \in \{0,1\}$. It is worth noting that we also allow the exposure mapping to depend on the covariates of other individuals.

Under a correctly specified exposure mapping in Equation (ref), the potential outcome $Y_i(d, \textbf{d})$ simplifies to $Y_i(d, z)$. This permits defining the average direct effect, interference effect, and total effect, denoted by $\gamma(z)$, $\delta(d, z, z')$, and $\Delta(z, z')$, respectively, which are functions of an individual's own treatment and the exposure:

align[align omitted — 179 chars of source]

where $z$ and $z'$ are two distinct exposures (e.g.,\ 1 and 0 in the binary case). Next, we provide the formal assumptions on the causal structure that ensure identification of the causal parameters defined in Equation (ref). We assume that we observe data $W_i=(Y_i,D_i, X_i)$ for a fixed network $\mathcal{A}$, $i=1,\dots,n$, where $(X_i,D_i)$ is $i.i.d.$, but $Y_i$ may depend on $\mathcal{D}_{-i}$ and $\mathcal{X}_{-i}$. Our first assumption concerns the exposure mapping and the treatment assignment:

assumption[Identification - Independence of $D_i$ and $Z_i$] \begin{align*} Z_i=\mathcal{F}(\mathcal{A},\mathcal{D}_{-i},\mathcal{X}_{-i}),\ and\ D_i=D_i(X_i). \end{align*}

First, Assumption (ref) states that the true exposure $Z_i=\mathcal{F}(\mathcal{A},\mathcal{D}_{-i},\mathcal{X}_{-i})$ is a function only of $\mathcal{A}$, $\mathcal{D}_{-i}$, and $\mathcal{X}_{-i}$. Second, the treatment assignment of individual $i$ may depend only on its own characteristics $X_i$, that is, $D_i=D_i(X_i)$. This implies that individual $i$'s treatment assignment $D_i$ is conditionally independent of its own exposure $Z_i$ given the covariates $\mathcal{X}=(X_i,\mathcal{X}_{-i})$ and the network structure $\mathcal{A}$. As discussed in AronowSamii2017, the propensity score under inference is defined as the joint conditional probability of the treatment and the exposure, given the network structure $\mathcal{A}$ and the observed covariates $\mathcal{X}=(X_i,\mathcal{X}_{-i})$. Under Assumption (ref), the propensity score is given by

align[align omitted — 160 chars of source]

Using a directed acyclic graph (DAG) (see, e.g. Pearl00), Figure (ref) represents a causal structure where Assumption (ref) is satisfied. In the graph, nodes represent variables, and arrows indicate causal associations between those variables. The treatment assignments of other individuals \(\mathcal{D}_{-i}\), the network structure \(\mathcal{A}\) and their covariates \(\mathcal{X}_{-i}\) jointly determine the exposure \(Z_i\), which influences the outcome \(Y_i\) in the presence of interference. The confounder $X_i$ influences individual treatment assignment \(D_i\) and the outcome \(Y_i\), whereas $\mathcal{X}_{-i}$ influences the treatment assignments of other individuals \(\mathcal{D}_{-i}\). In this structure, the direct effect refers to the effect of an individual's own treatment $D_i$ on their outcome $Y_i$, i.e., $D_i \rightarrow Y_i$. The interference effect corresponds to the impact of other individuals' treatment assignments $\mathcal{D}_{-i}$ on individual $i$'s outcome $Y_i$ through the exposure $Z_i$, i.e., $\mathcal{D}_{-i} \rightarrow Z_i \rightarrow Y_i$. We assume that both the network structure \(\mathcal{A}\) and the covariates of other individuals $\mathcal{X}_{-i}$ do not directly affect the individual $i$'s outcome $Y_i$, but only through $Z_i$.

Consistent with the causal structure in Figure (ref), identification of the direct treatment effect of $D_i$ requires that all backdoor paths from $D_i$ to $Y_i$ are blocked by the individual $i$'s observed covariates $X_i$ and the exposure $Z_i$. This motivates the following conditional exogeneity assumption:

assumption[Identification - Conditional exogeneity of $D_i$ and $Z_i$] \begin{align*} Y_i(d,z) \perp\!\!\!\perp (D_i,Z_i) | \mathcal{X},\mathcal{A} \quad \forall d \in \{0,1\},\ z\in\mathcal{Z}. \end{align*} where $\mathcal{Z}$ denotes the support of $Z_i$.

Assumption (ref) imposes that there are no unobserved variables that jointly affect $Y_i$ and the treatment assignment $D_i$, or $Y_i$ and the true exposure $Z_i$ conditional on $\mathcal{X}$ and $\mathcal{A}$. Notably, under our assumed causal structure, there is no direct effect of $\mathcal{A}$ or $\mathcal{X}_{-i}$ on the outcome $Y_i$. As a result, conditioning on $X_i$ alone would be sufficient in Assumption (ref), since all paths from $\mathcal{X}_{-i}$ and $\mathcal{A}$ to $Y_i$ are blocked by $Z_i$.

Furthermore, we require that the propensity score for any combination of treatment and exposure is strictly positive. This implies that the network structure does not deterministically determine the exposure. Thus, there exists variation in exposures, conditional on the network and the covariates, that can be leveraged to assess their effects. This leads to the following common support assumption:

assumption[Identification - Common support] \begin{align*} p_i(d, z) > 0 \quad \forall d \in \{0,1\}, z \in \mathcal{Z}. \end{align*}

Under Assumptions (ref) - (ref), the causal effects defined in Equations (ref) are identified through the propensity score. As discussed in aronow2020spillover, inverse probability weighting (IPW) Horvitz52 can be applied, reweighting observations by the inverse of the propensity score to recover the mean potential outcomes and effects:

align[align omitted — 527 chars of source]

where $I \{\cdot \}$ denotes the indicator function, which is one if its argument is satisfied and zero otherwise.

figure[figure omitted — 857 chars of source]

Learning exposure mappings

While most studies evaluating interference effects rely on researcher-specified mappings, in real-world applications the function $\mathcal{F}$, and thus the true definition of the exposure mapping, is unknown. For this reason, we aim to approximate the true exposure, $Z_i$, using a graph convolutional autoencoder (GCA) related to the graph autoencoder (GAE) of kipf2016variationalgraphautoencoders\footnote{Although inspired by the GAE of kipf2016variationalgraphautoencoders, our GCA differs in that it is trained to predict $Y_i$ in a supervised setting, rather than constructing the adjacency matrix.}. In our supervised learning approach, $\tilde{Z}_i$ is learned as the embedding from an intermediate hidden layer of a GCA trained to predict the outcome $Y_i$ from the network $\mathcal{A}$, the treatment assignments of the individuals $\mathcal{D}=(D_i,\mathcal{D}_{-i})$, and the observed covariates $\mathcal{X}=(X_i,\mathcal{X}_{-i})$. A graph autoencoder consists of an encoder that maps the input graph and node features into a low-dimensional latent representation (embedding) capturing its most relevant information, and a decoder that predicts target values from this latent representation. In our setting, the GCA combines an encoder that consists of graph convolutional network (GCN) layers kipf2017semisupervisedclassificationgraphconvolutional and a regression-based decoder. This architecture allows the model to automatically learn dependencies and relationships between nodes based on the network structure. The resulting learned exposures, denoted by $\tilde{Z}_i$, correspond to the embeddings produced by the encoder and are designed to capture how treatment assignments across the network affect individual outcomes.

We consider an undirected graph $G = (\mathcal{V},\mathcal{E})$, where the nodes are individuals $v_i \in \mathcal{V}$ with $|\mathcal{V}| = n$ and the edges are given by pairs $(v_i,v_j) \in \mathcal{E}$. The adjacency matrix $\mathcal{A} \in \mathbb{R}^{n \times n}$ represents the connections between the individuals and is defined as

equation[equation omitted — 202 chars of source]

In the adjacency matrix, $a_{ij} = 1$ indicates that node $i$ (i.e., individual $i$) is connected to node $j$, while $a_{ij} = 0$ indicates that there is no edge, i.e., no social connection, between nodes $i$ and $j$. Each node has its own features and the input feature vector consists of the treatment assignments and the covariates of all individuals. Denote this matrix by $M$ with dimensions $n \times 2$, where $d_i$ is the individual assignment and $x_i$ is the covariate of individual $i$:

equation[equation omitted — 156 chars of source]

Each layer in the encoder is indexed by $k \in \{0,\dots , K-1\}$, where $k=0$ corresponds to the raw input, $k=1$ to the first layer and $K-1$ to the final layer of the encoder. The representation of the raw input (layer $k=0$) is $H^{(0)} = M$. The graph convolutional layers $k \in \{1, \dots , K-1\}$ in the encoder follow the layer-wise propagation rule

equation[equation omitted — 136 chars of source]

where $T$ is the degree matrix. The matrix $T$ contains zeros everywhere except on the diagonal, where $t_{ii} = \sum_j a_{ij}, \forall i \in \{1, \dots, n\}.$ $\tilde{W}^{(k)}$ is a trainable and layer-specific weight matrix, and $\sigma(\cdot)$ is an activation function. $H^{(k+1)} \in \mathbb{R}^{n \times x_k}$ contains the node representations with each row corresponding to one node, where $x_k$ is the feature dimension of that layer. The final encoder layer $H^{(K)}$, with $k = K-1$, has feature dimension $x_{k}=1$, as its output serves as the learned exposure $\tilde{Z}_i$, which we model as a one-dimensional embedding, that is, $ H^{(k+1)} \in \mathbb{R}^{n \times 1}$ for $k=K-1$. This layer-wise update in the graph convolutional layers corresponds to a standard message-passing mechanism, in which each node aggregates transformed information from its neighbors according to the graph structure in $\mathcal{A}$. To obtain an outcome prediction $\hat{\mathcal{Y}} \in \mathbb{R}^{n \times 1}$, the embeddings from the encoder are passed into a regression-based decoder implemented as a linear layer, which maps the learned exposure representation into a predicted outcome:

equation[equation omitted — 87 chars of source]

where $\tilde{W}_{dec} \in \mathbb{R}^{1 \times 1}$ is the trainable weight in the decoder layer, $H^{(K)}$ is the vector of learned exposures $\tilde{Z}_i$, $i\in\{1\dots,n\}$, and $b_{dec} \in \mathbb{R}$ is the bias term. The architecture of the GCA is illustrated in Appendix (ref) Figure (ref).

The GCA is trained by minimizing the mean squared error loss function $\mathcal{L}(Y_i,\hat{Y}_i) = \frac{1}{n} \sum_i (Y_i-\hat{Y}_i)^2$, which is appropriate, because the outcome variable $Y_i$ is continuous.

It is important to note that the adjacency matrix $\mathcal{A}$ contains no self-loops, i.e., $a_{ii}=0, \forall i \in \{1, \dots, n\}$. As a result, each node aggregates information exclusively from its neighbors. Thus, the representation $H^{(k+1)}$ at any encoder layer does not use the node's own features $(D_i,X_i)$, but only those of other individuals $(\mathcal{D}_{-i},\mathcal{X}_{-i})$. This is consistent with the interpretation of the exposure $\tilde{Z}_i$, which is intended to summarize how the treatments and characteristics of others affect individual $i$'s outcome.

Testing the validity of exposure mappings

In this section, we discuss the testability of whether a researcher-defined exposure mapping is correctly specified for capturing all interference effects. Our approach uses the learned exposures $\tilde{Z}_i$, $i=1,\dots,n$, from Section (ref) as an instrument to assess whether the (simpler) exposures $\dot{Z}_i$ sufficiently captures all interference. Specifically, if the simpler mapping accurately captures all interference effects, then the exposure $\tilde{Z}_i$ is independent of the individual’s outcome, conditional on $\dot{Z}_i$. By examining violations of this condition, we can assess whether the researcher-defined mapping fails to capture some of the underlying interference. Notably, the true exposure $Z_i=\mathcal{F}(\mathcal{A},\mathcal{D}_{-i},\mathcal{X}_{-i})$, for $i=1,\dots,n$, is i.i.d. by assumption. We therefore retain the index $i$ in the learned and researcher-defined exposures $\tilde{Z}_i$ and $\dot{Z}_i$ for notational convenience. Next, we introduce the assumptions that underlie our testing approach for validating the exposure mapping. These assumptions formalize the structural restrictions on how the exposure variables may depend on one another and on the outcome. We further assume that any statistical independencies correspond to d-separation in the underlying causal model, a condition known as causal faithfulness and formally stated below.

assumption[Testing method - causal structure and faithfulness] We assume that $$ \dot{Z}_i(y)=\dot{Z}_i,\textit{ and }\tilde{Z}_i(\dot{z},y)=\tilde{Z}_i \quad \forall \dot{z} \in \mathcal{\dot{Z}} \textit{ and } y \in \mathcal{Y},$$ where $\mathcal{\dot{Z}}$ and $\mathcal{Y}$ denote the corresponding support of $\dot{Z}_i$ and $Y$ and that only variables which are d-separated in some causal model are statistically independent.

The first part of Assumption (ref) rules out reverse causal effects of outcome $Y$ on $\dot{Z}_i$ and $\tilde{Z}_i$ as well as any causal effect of $\dot{Z}_i$ on $\tilde{Z}_i$. The latter reflects the fact that $\tilde{Z}_i$ is a causal parent of $\dot{Z}_i$, which is consistent with the interpretation of $\tilde{Z}_i$ as a more complex mapping from the network structure and neighbor treatments, while $\dot{Z}_i$ represents a simpler transformation thereof. The following assumption requires that every possible combination of $\tilde{Z}_i$ and $\dot{Z}_i$ occurs with positive probability.

assumption[Testing method - common support] \begin{align*} Pr(\dot{Z}_i = \dot{z}, \tilde{Z}_i = \tilde{z}) > 0 \quad \forall \dot{z} \in \mathcal{\dot{Z}}, \tilde{z} \in \mathcal{\tilde{Z}} \end{align*} where $\mathcal{\dot{Z}}$ and $\mathcal{\tilde{Z}}$ denote the corresponding support of $\dot{Z}_i$ and $\tilde{Z}_i$.

In Assumption (ref), we impose that the simpler exposure $\dot{Z}_i$ and the learned exposure $\tilde{Z}_i$ are statistically dependent.

assumption[Testing method - dependence between $\dot{Z}_i$ and $\tilde{Z}_i$] \begin{align*} \dot{Z}_i \not\!\perp\!\!\!\perp \tilde{Z}_i. \\ \end{align*}

Together with Assumption (ref), which rules out any effect of $\dot{Z}_i$ on $\tilde{Z_i}$, Assumption (ref) ensures that $\tilde{Z}_i$ causally affects $\dot{Z}_i$. This corresponds either to a first-stage relationship in the IV literature or to the presence of (potentially unobserved) characteristics that jointly influence both $\dot{Z}_i$ and $\tilde{Z}_i$. This assumption is satisfied in Figure (ref), where $\tilde{Z}_i$ serves as a causal parent of $\dot{Z}_i$. Assumptions (ref) and (ref), allow us to apply Theorem 1 of huberkueck2022 in order to construct a test for validating the exposure mapping. It is worth noting that we still rely on the causal structure shown in Figure (ref) and Figure (ref). Most importantly, we assume that there are no unobserved confounder jointly affecting $D_i$ and $\tilde{Z_i}$ or $D_i$ and $\dot{Z_i}$.

theorem{Conditional on Assumptions (ref) and (ref), it holds that} \begin{eqnarray} &Y_i(d,\dot{z}) \perp\!\!\!\perp \dot{Z}_i ,\quad Y_i(d,\dot{z}) \perp\!\!\!\perp \tilde{Z}_i \iff & Y_i \perp\!\!\!\perp \tilde{Z}_i | \dot{Z}_i=\dot{z}, \quad\forall \dot{z} \in \mathcal{\dot{Z}}, d\in\{0,1\}. \end{eqnarray} Conditional on Assumptions (ref) and (ref), the testable implication $Y_i \perp\!\!\!\perp \tilde{Z}_i | \dot{Z}_i=\dot{z}$ is necessary and sufficient for the joint satisfaction of $Y_i(d,\dot{z}) \perp\!\!\!\perp \dot{Z}_i$ and $Y_i(d,\dot{z}) \perp\!\!\!\perp \tilde{Z}_i$ when considering potential outcomes $Y_i(d,\dot{z})$ matching $\dot{Z}_i=\dot{z}$ and $D_i=d$.

This statement follows directly from Theorem 1 of huberkueck2022, considering $\dot{Z}_i$ as the treatment variable and $\tilde{Z}_i$ as the suspected instrument. The first two conditions of Theorem (ref) state two independence conditions, which correspond to selection-on-observables for the exposure $\dot{Z}_i$ and instrument validity for the learned exposure $\tilde{Z}_i$, respectively. The first condition imposes that the potential outcome $Y_i(d,\dot{z})$ is independent of $\dot{Z}_i$ for all possible exposure levels $\dot{z} \in \mathcal{\dot{Z}}$ and $d\in\{0,1\}$, i.e., $Y_i(d,\dot{z}) \perp\!\!\!\perp \dot{Z}_i$. This rules out unobserved confounding between $\dot{Z}_i$ and $Y_i$. In Figure (ref), this corresponds to the absence of dotted arrows from the unobserved confounder $V_i$ to both $\dot{Z}_i$ and $Y_i$. The second independence condition of Theorem (ref) imposes instrument validity for the learned exposure $\tilde{Z}_i$. It requires that $Y_i(d,\dot{z})$ is independent of $\tilde{Z}_i$, i.e., $Y_i(d,\dot{z}) \perp\!\!\!\perp \tilde{Z}_i$. This means that the learned mapping does not directly affect the potential outcome and that there are no unobserved confounders jointly affecting $\tilde{Z}_i$ and $Y_i(d,\dot{z})$. In Figure (ref), this corresponds to the absence of the dotted arrows from the unobserved confounder $U_i$ to both $\tilde{Z}_i$ and $Y_i$. Therefore, the second condition of Theorem (ref) ensures that $\tilde{Z}_i$ satisfies the same type of independence requirements as a valid instrument in the IV literature. Following Theorem 1 of huberkueck2022, the testable conditional independence $Y_i \perp\!\!\!\perp \tilde{Z}_i | \dot{Z}_i=\dot{z}$ in the third condition of Equation (ref) is necessary and sufficient for the joint satisfaction of the first two independence conditions. In our setting, this provides the following interpretation of the testable implication: we assess whether the learned exposure $\tilde{Z}_i$ contains any residual association with the outcome $Y_i$ after conditioning on the predefined exposure $\dot{Z}_i$. If such an association remains, the mapping $\dot{Z}_i$ fails to capture all relevant interference, and the learned mapping offers additional, causally meaningful information. In turn, if the conditional independence holds, the simpler exposure mapping is sufficient for capturing interference effects. In Figure (ref) the dashed line (i.e., $\tilde{Z}_i \dashrightarrow Y_i$) indicates a violation of the testable condition, as $\dot{Z}_i$ does not fully capture the interference effects of $\tilde{Z}_i$ on $Y_i$ due to its definition being too simple to account for all forms of interference.

In this paper, we focus on mean conditional independence between the outcome variable and the learned exposure mapping rather than on the full outcome distribution. This is sufficient for identifying average effects, see Theorem 2 of huberkueck2022, within our framework of average direct and interference effects defined in Equation (ref). Modifying the testable implication in Equation (ref) to hold in expectation yields the following testable implication:

align[align omitted — 201 chars of source]

Equation (ref) requires that all combinations of $\dot{Z}_i$ and $\tilde{Z}_i$ occurring in the conditioning sets are observed with positive probability. This is ensured by Assumption (ref). Following huberkueck2022 and Apfel2023learning, testing whether the difference in expected outcomes in Equation (ref) equals zero requires checking this condition for all values of $\dot{Z}$ and $\tilde{Z}$ in their respective supports, which can imply infinitely many testable implications. The solution is to test Equation (ref) globally, i.e., across all values of $\dot{Z}$ and $\tilde{Z}$, using an aggregated $L_2$-type measure proposed in equation (3.4) in Apfel2023learning. Formally, we test $H_0: \theta_0 = 0$ where $\theta_0=(E[Y_i | \dot{Z}_i = \dot{z}_i, \tilde{Z}_i = \tilde{z}_i]-E[Y_i | \dot{Z}_i = \dot{z}_i])^2+(E[Y_i | \dot{Z}_i = \dot{z}_i, \tilde{Z}_i = \tilde{z}_i]-E[Y_i | \dot{Z}_i = \dot{z}_i])$ which evaluates the testable implication in Equation (ref). Relying on the DML approach, Apfel2023learning derive an estimator $\hat{\theta}_0$ that is asymptotically normal and $\sqrt{n}$-consistent under suitable regularity conditions. We use this estimator to test whether the researcher-defined exposure is correctly specified, i.e., whether Equation (ref) holds true. The technical details of the testing approach are given in Appendix (ref).

figure[figure omitted — 1,406 chars of source]

Estimation of causal effects

In this section, we outline the estimation of the causal effects defined in Section (ref). In the following the exposure mapping $Z_i$ is either the researcher-defined exposure mapping, $\dot{Z}_i$, or the learned exposure mapping, $\tilde{Z}_i$, depending on the result of the testing method introduced in Section (ref). If $H_0$ is rejected, the learned exposure mapping is used for the estimation of the causal effect, i.e., $Z_i = \tilde{Z_i}$. If $H_0$ is not rejected, the researcher-defined exposure mapping is used instead, i.e., $Z_i = \dot{Z_i}$. When $Z_i = \tilde{Z_i}$, the exposure mapping may be difficult to interpret. Hence, our focus is on the identification and estimation of the direct effect averaged over all exposure levels, $\gamma = E[\gamma(Z)]$.

To estimate these causal effects, we build on the identification assumptions introduced in Section (ref). While IPW identification is valid under correct specification of the propensity score, outcome-regression-based identification using a model for the conditional mean outcome is valid under correct specification of that model. Both methods can be combined in a doubly robust (DR) approach, which remains consistent if either the propensity score or the outcome model is correctly specified Robins+94, RoRo95. Moreover, the DR approach is first-order insensitive to small deviations of both the propensity score and the conditional mean outcome from their respective true models, a property known as Neyman-orthogonality Neyman1959. This property is key for the application of the DML framework of Chetal2018, in which propensity scores and outcome models are estimated with machine learning methods that may be prone to approximation errors. If the propensity score and outcome models are estimated with convergence rate $o(n^{-1/4})$, DML estimation of average direct, interference and total effects can attain $\sqrt{n}$-consistency under certain regularity conditions. Denoting by $\mu_i(d,z,x) = E[ Y_i | D_i = d, Z_i = z, X_i = x]$ the conditional mean outcome, the DR expression for the mean potential outcome is given by

align[align omitted — 165 chars of source]

where $\psi(W_i)=\phi(W_i)-E[Y_i(d,z)]$ is the efficient score function.

The outcome regression $\mu_i(d,z,x)$ can be estimated using any machine learning method that satisfies the required convergence rate. The propensity score $p_i(d,z)$ defined in Equation (ref) consists of two components. The first component $\Pr(D_i = d | X_i)$ can be estimated using, for example, logistic regression when the treatment is binary. The second component $\Pr(Z_i = z | \mathcal{X}_{-i}, \mathcal{A})$ is conditional on the network $\mathcal{A}$, which is why we suggest estimating it using the GNN-based propensity score estimator introduced in leung2022graph. They use a graph isomorphism network (GIN) to estimate the propensity score, that is, the conditional distribution of exposure given the covariates of the other individuals $\mathcal{X}_{-i}$ and the network structure $\mathcal{A}$. Since the exposure can take multiple values, the GIN produces a probability for each possible exposure level through a multiclass output layer. The target variable is a one-hot-encoding of the observed exposure level. The model is trained using a logistic loss, $\exp(\hat{m}_z)/\sum_{z' \in \mathcal{Z}}\exp(\hat{m}_{z'})$, to estimate the propensity score, where $\hat{m}_z$ denotes the GIN output corresponding to exposure level $z$. They show that, under approximate neighborhood interference (ANI), the propensity score can be estimated with a convergence rate of $o(n^{-1/4})$. This allows valid estimation of the average direct effect, $\gamma = E[\gamma(Z)]$, which we will also show in a simulation study in the next section.

Simulation

This section provides a simulation study to investigate the finite sample behavior of the testing approach as well as the direct effect estimation. The data are generated according to the following data generating process:

align*[align* omitted — 365 chars of source]

where $(\alpha, \delta, \gamma, \xi) = (-1, 5, 1, 1)$. The outcome $Y_i$ is a linear function of the individual treatment $D_i$, the true exposure $Z_i$, the covariate $X_i$ and an unobservable $\varepsilon_i$. The covariate $X_i$ is a binary variable, and the binary treatment $D_i$ is a function of $X_i$. Following leung2022graph, the network structure $\mathcal{A}$ is generated from a random geometric graph model. The GCA used for learning the exposure mapping consists of two graph convolutional layers in the encoder and one regression-based decoder. The learning rate is 0.01 and the number of epochs is 200. For estimating the score function, the support of the continuous learned exposure variable is partitioned based on the quartiles of its distribution, i.e., $L=4$.

We consider three settings to assess the performance of our testing approach, using $200$ Monte Carlo replications for sample sizes $n = 500, 1000, \text{ and } 2000$. In the first setting, the true exposure mapping is given by the share of treated neighbors weighted by their covariates, i.e., $Z^{S1}_i =\frac{\sum_{j \ne i}a_{ij}D_j X_j}{\sum_{j \ne i}a_{ij}}$, and the researcher-defined exposure mapping coincides with the true exposure, i.e., $\dot{Z}^{S1}_i =\frac{\sum_{j \ne i}a_{ij}D_j X_j}{\sum_{j \ne i}a_{ij}}$. Thus, the researcher-defined exposure mapping is correctly specified and the exposure learned by the GCA should not contain additional information beyond the researcher-defined mapping. In Setting 2, the true exposure mapping depends on both first-order and second-order network neighborhoods. Specifically, the true exposure is given by $Z^{S2}_i = \frac{\sum_{j \ne i}a_{ij}D_j X_j}{\sum_{j \ne i}a_{ij}} + \frac{\sum_{k \ne i}b_{ik}D_k X_k}{\sum_{k \ne i}b_{ik}}$, where $b_{ik} := I\{\sum_{i \ne j, j \ne k} a_{ij}a_{jk} > 0\} \cdot (1-a_{ik}) \cdot I\{k \ne i \}$. The first term captures again the share of treated neighbors weighted by their covariates. The second term of the true exposure captures second-degree neighbors of individual $i$, i.e., nodes that are connected to $i$ via a path of length two, excluding $i$ itself and all direct neighbors. In contrast, the researcher-defined exposure is binary and only accounts for direct neighbors, i.e., $\dot{Z}^{S2}_i =I\{ \sum_{j \ne i} a_{ij}D_j X_j > 0\}$. Thus, the researcher-defined exposure is misspecified. In the third setting, the researcher-defined exposure is equal to the one in Setting 1, i.e., $\dot{Z}^{S1}_i=\dot{Z}^{S3}_i$. However, the true exposure is a nonlinear transformation of the cumulative treated neighborhood intensity with an explicit threshold and saturation: $Z^{S3}_i=1- \exp \Big(-0.5 \; \max\!\left\{0, \;\sum_{j \ne i}a_{ij}D_j X_j - 10 \right\} \Big)$. This setting therefore also represents a case of misspecification if researchers were to apply a linear specification to model the conditional mean outcome in Equation (ref) based on $\dot{Z}^{S3}_i$.

table[table omitted — 799 chars of source]

Table (ref) reports the empirical rejection rates and mean p-values of the testing approach across the three settings. In Setting 1, where the researcher-defined exposure mapping is correctly specified, the test never rejects the null hypothesis and the p-values remain high across all sample sizes. In Settings 2 and 3, where the researcher-defined exposure mapping is misspecified, the rejection rate increase with the sample size, while the mean p-values decrease and approach the significance level $\alpha = 0.05$. Overall, these results are consistent with the theoretical implications discussed in Section (ref).

In the second part of the simulation study, we asses the estimation of the direct effect based on the learned exposure $\tilde{Z}_i$ obtained from a GCA. The direct effect is estimated using IPW, where the propensity score for the treatment assignment, $Pr(D_i =1 | X_i)$, is estimated via logistic regression, and the exposure propensity $Pr(Z_i | \mathcal{X}_{-i}, \mathcal{A})$ is approximated using an oracle estimator based on the data-generating process. The true exposure that the GCA aims to recover is defined as $Z_i =I\{ \sum_{j \ne i} a_{ij}D_j X_j > 2\}$ and we adjust the parameter for generating $a_{ij}$ to $r_n = \Big(\frac{5}{\pi n} \Big)^{(1/2)}$ in the DGP outlined in the main text.

table[table omitted — 602 chars of source]

Table (ref) reports the simulation results for the estimation of the direct effect $\gamma$ based on the learned exposure $\tilde{Z}_i$. For small sample sizes, the estimator shows noticeable variability and bias. As the sample size increases, both the bias and the standard deviation decrease. For sample sizes $n=500$ and $n=1000$, the average estimated direct effect is close to the true value $\gamma = 1$. This indicates improved precision and convergence toward the true direct effect.

Conclusion

In this paper, we develop a data-driven approach to learn exposure mappings in the presence of interference, instead of relying on a priori defined mapping. We use a graph convolutional autoencoder to learn exposure mappings that summarize how others' treatment assignments affect an individual's outcome. Since the identification of average direct effects depends crucially on the correct specification of the exposure mapping, we study whether a simple, researcher-defined exposure mapping is sufficient or a more complex, learned mapping is required. To this end, we propose a testing method based on conditional independence implications. This test evaluates whether a researcher-defined exposure mapping captures all relevant interference. Violations of the testable implication indicate that the predefined exposure mapping is misspecified and that a learned mapping should be used for estimating the direct effect. Overall, our study provides guidance on how to learn and validate exposure mappings in the presence of interference.

{ \setcounter{equation}{0}