Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
129,054 characters · 18 sections · 87 citation commands
The Network Propensity Score: Spillovers, Homophily, and Selection into Treatment
\thispagestyle{empty}
\doublespacing \selectfont
A popular strategy to identify average treatment effects in quasi-experimental settings is to compare the outcomes of individuals with similar characteristics but different treatment status. For this strategy to be valid, the researcher needs to satisfy a high-level unconfoundedness condition --also known as selection-on-observables-- which lists a set of observable covariates that delimit comparison subgroups, and a support condition, that guarantees sufficient people to compare in each subgroup. Such conditions are trivially satisfied in experiments with known assignment probabilities, but require further justification in observational settings. For example, researchers can appeal to institutional features of the program assignment rules or prior knowledge of the participants' decision-making process. While this type of strategy has a long history, researchers have traditionally ignored cases with meaningful interference/spillovers, where individual's potential outcomes depend on the treatment status of others, in addition to their own.\footnote{These situations violate the Stable Unit Treatment Value Assumption (SUTVA).} A burgeoning literature has focused on extending these notions of unconfoundedness and support to situations with network spillovers forastiere2020identification, liu2019doubly,sofrygin2017semi.\footnote{For example, job placement programs can displace non-participants from the labor market crepon2013labor, cash transfers can affect informal insurance networks meghir2020migration, and professional events can encourage the adoption of business practices fafchamps2018networks.}. In spite of these advances, much less is known about the economic content or plausibility of these assumptions outside experimental settings.
The problem is that spillovers introduce a second layer of selection through the choice of social connections. Individuals may be likely to befriend others with similar willingness to participate in the program. This implies that strategies to identify spillovers by comparing the outcomes of social groups with high and low participation rates may be misleading. Cross-sectional differences could reflect sorting patterns into high- or low-take-up groups, rather than a product of social interactions. This type of phenomenon is called homophily. Resolving this problem by comparing the outcomes of two “similar” individuals with differential friend take-up rates is a step in the right direction. However, it is difficult to define which pairs of individuals are actually comparable in a network sense, at least without further assumptions. The problem becomes more acute in real-world social networks, where individuals have non-overlapping sets of friends. Defending institutional/decision-based rationales for unconfoundedness is hard without a global, internally consistent model of collective decision-making due to the many different ways in which the network composition could affect the selection process in either layer.
This paper's main contribution is to propose primitive assumptions for unconfoundedness and support conditions, which suffice to identify causal effects in the presence of spillovers. To do so I propose an internally consistent model with selection-on-observables, bilateral network formation, and potential outcomes with spillovers (via a random coefficients specification). To keep things tractable I focus on binary networks, observed by the researchers, that do not quantity relationship intensity. Within my framework individuals select into treatment and form connections based on a combination of observed characteristics and i.i.d shocks. My non-parametric identification approach is constructive and builds on the notion of a graphon -- a function that can be used to represent a large class of exchangeable network processes. Most importantly, depending on the primitives of the model, the data can exhibit no homophily, homophily with spillovers, or homophily without spillovers. I argue that elementary building blocks can lead to rich selection patterns and produce bias of na\"{i}ve procedures, even without modeling strategic considerations explicitly.
The strength of my results depends on the generality with which the researcher decides to model spillovers. First, I focus on a class of exchangeable potential outcomes models. I allow for situations where the outcome can depend on the take-up of friends, the take-up of friends of friends, or weighted averages that depend on covariates of friends. This nests a heterogeneous version of the reduced-form linear-in-means bramoulle2009identification, interference with rooted networks auerbach2021local, and approximate network interference leung2022causal. Second, I specialize my results to models that satisfy an anonymous interactions assumption. This condition --which has been extensively analyzed in a number of recent papers aronowsamii2017estimating,leung2019treatment,spilloversnoncompliance,forastiere2020identification,liu2019doubly,sussman2017elements,tchetgen2017auto-- states that the potential outcomes spillovers only enter through direct friend connections, equally-weighted. The exchangeable and anonymous interaction models coincide when individuals are fully connected within disjoint clusters.\footnote{In that case the modeling approach is sometimes known as partial interference}. It is important to emphasize that both of these generalizations nests the Stable Unit Treatment Value (SUTVA) assumption.
I then use this framework to establish two key findings. First, the researcher can satisfy the unconfoundedness condition by choosing individual determinants of treatment take-up and relationship choices --but not friends' take-up decisions. This applies to a large class of exchangeable spillover models. Second, for the subset of models with anonymous interactions aronowsamii2017estimating there exists a three-dimensional individual statistic --that I call the network propensity score (NPS)-- which can be used as a matching variable. The first component is the individual propensity score rosenbaum1983central, the second is a measure of friend take-up rate, and the last is the number of friends. Crucially, the validity of the support condition is easy to verify and can be motivated from the patterns of association in the network. From a structural perspective, the NPS can be expressed as an integrand of the take-up process, friend preferences over traits, and the measure of traits in the population
I establish the relationship between the network propensity score and recently proposed network pseudo-metrics auerbach2022identification,zeleneev2020identification, and illustrate weak identification issues that arise from applying those approaches to study spillovers. I propose alternative assumptions to deal with unobserved heterogeneity, that encompass other strategies recently used in the literature spilloversnoncompliance,johnsson2019estimation,imbens2009identification. I propose a two-step semi-parametric estimator, that is based on inverse-weighting in a random coefficients specification graham2022semiparametrically,wooldridge1999distribution. In the asymptotics, I allow for the possibility that a subset of the control variables are unobserved but can be consistently estimated in large networks. My approach is agnostic about intra-network dependence, and hence the rate of convergence of the estimator is going to depend on the total number of groups/networks.
I apply my methodology to two empirical examples. First, I consider an intervention designed to increase political participation in Uganda eubank2019viral,ferrali2020takes. Citizens voluntarily participated in quarterly information sessions about ways to engage with local district officials. I find evidence of spillovers because individuals with a higher number of friends participating in the sessions were more likely to be politically active, after controlling for covariates. The estimates of the spillover effects under my approach are statistically significant and about twice the size of comparable ordinary least squares (OLS) regressions with additive covariates. The network propensity score matching methodology is better equipped to handle heterogeneous spillover effects that can be correlated with the endogeneous regressors.
In the second example, I analyze the effects of an intervention to increase microfinance adoption banerjee2013diffusion. This example has been analyzed extensively by the econometrics literature candelaria2020semiparametric,chandrasekhar2014tractable and has lead to many follow-up projects banerjee2017gossip,breza2019social,chandrasekhar2018social. The microfinance organization used a non-random selection rule based on occupation of household members (shopkeepers, teachers), who received in-depth information about the loans offered by the company. In practice, households with higher wealth and privileged castes were both more likely to receive treatment themselves and to be friends with others that received treatment as well. My network propensity score matching approach estimates large treatment effects but limited local spillover effects. The results suggest that while network characteristics such as centrality can affect short-term speed of information diffusion banerjee2013diffusion,akbarpour2018just, local neighbor interactions may play a smaller role on medium-term loan adoption after accounting for covariates.
Finally this paper considers applications of the network propensity score approach to stratified experiments. I analyze experiments that exogenously assign treatment probabilities across multiple networks duflo2003role,crepon2013labor,baird,vasquezbare. I find that the network propensity score has a simple form in both cases under perfect compliance. I also consider settings with non-compliance and spillovers spilloversnoncompliance,vasquezbare,imai2020causal. I discuss the applicability of the network propensity score to identify average spillover effects under non-compliance in sparse networks.
The paper is organized as follows. Section (ref) reviews prior literature. Section (ref) introduces the model. Section (ref) presents identification under exchangeability. Section (ref) introduces the network propensity score. Section (ref) proposes feasible estimator and presents the asymptotic results. Section (ref) discusses the two empirical examples. Section (ref) concludes.
There have been three recent approaches in the literature that extend propensity score methods for use with network data. The first approach uses relationship data and friend covariates to relax the selection on observables assumption. jackson2020adjusting assume that program participation is the result of a strategic game with friends (spillovers in treatment), but assume that there are no spillovers on outcomes. The second approach assumes selection on observables (without spillovers) but focuses on pairwise outcomes. arpino2015implementing, for example, compute the propensity score of adopting tariff agreements and use it to evaluate their effect on bilateral trade between countries. The third approach, which is closest to my own, incorporates spillovers by assuming anonymous interactions manski2013identification, which implies heterogeneous outcomes that depend on own treatment and the total number of treated friends. This approach is sometimes called multi-treatment matching because it assumes that individuals with different numbers of treated friends experience different intensities that satisfy unconfoundedness {forastiere2018estimating, liu2019doubly,sofrygin2017semi}. In this case a form of generalized propensity score hirano2004propensity is computed for each exposure level. Recent work in economics incorporates similar uncounfoundedness assumptions leung2019treatment,viviano2019policy,ananth2020optimal. qu2021efficient propose an efficient estimator under heterogeneous partial inference. In most of these papers, the network is typically treated as exogenous, and often times variation in the treatment is the main source of identification.
One of the key innovations is to prove unconfoundedness from a general micro-founded setting with exchangeability. manski2013identification consider a slightly broader class of social interaction models, but restrict attention to experiments. qu2021efficient propose a version of exchangeability and unconfoundedness based on partitions of the sample into disjoint groups of influence. I propose a more general version that accommodates higher-order connections auerbach2021local,leung2022causal,bramoulle2009identification or covariate-weighting to characterize peers with more influence. Covariate-weighted versions can be computed with full network data or more cost-effective approaches using aggregate relational data breza2020using,alidaee2020recovering. My results are also novel for the anonymous interactions. A wide literature, e.g. forastiere2018estimating, compute a version of generalized propensity score for network data by predicting each possible category of the endogenous variable. The rationale for unconfoundedness is often based on implicit arguments for network homophily. However, I show that if network formation/selection arguments are being used to select the covariates believed to satisfy unconfoundedness, then the dimensionality of the generalized propensity scores can be theoretically be brought down to three (the network propensity score).
To my knowledge, this is the first paper to combine an internally consistent non-parametric network model with selection-on-observables to justify unconfoundedness. Previously, goldsmith2013social studied a parametric network model and use it to identify causal effects. johnsson2019estimation extended this idea to a class of network models that satisfy a monotonicity restriction similar to gao2019nonparametric. Both papers restrict attention to cases with an exogenous treatment. However, identification of causal effects is possible for a broader class of network models. In two influential papers aldous1981representations and hoover1979relations showed that any network process whose distribution is ex-ante independent of the ordering of agents can be represented as a dyadic network with independent covariates and independent shocks. Recent work has also shown how to micro-found the dyadic model from dynamic games mele2017structural, and how to test the empirical validity of dyadic models with additive fixed effects pelican2019testing.
The key empirical challenge is whether the covariates of the Aldous-Hoover representation are actually observed or whether some of them may be latent. Imposing monotonicity as in graham2017econometric, gao2019nonparametric, johnsson2019estimation is one way to recover latent heterogeneity. Another recent literature auerbach2019identification,zeleneev2020identification focuses on a network pseudo-metric to analyze more general forms of unobserved heterogeneity. auerbach2019identification, however, argues that this form of heterogeneity cannot be separately identified from spillovers in dense networks with exogenous treatment. I show why this concern carries over the case with selection-on-observables, by establishing the relationship between the network propensity score and the network pseudo-metric. Another route to recover unobserved heterogeneity is to explore group patterns in treatment decisions. spilloversnoncompliance recover control variables in experiments with spillovers and one-sided non-compliance. I discuss conditions that allow for this type of control variables. These conditions could also nest related matching approaches that exploit exponential family forms arkhangelsky2018role.
I assume that there are $g = \{1,\ldots,G\}$ disjoint groups that contain $i = \{1,\ldots,N_g\}$ individuals each. We can interpret $g$ as the identifier for a school, village or city. Treatment status is denoted by a binary variable $D_{ig}$ that equals one if individual $\{ig\}$ is treated and zero if she is not. Each individual has a vector of socio-economic covariates $C_{ig}$, which can be stacked in a matrix $C_g$. A social network is denoted by an $N_g \times N_g$ adjacency matrix $A_g$ with binary entries. Each entry $A_{ijg}$ equals one if individuals $\{ig\}$ and $\{jg\}$ are friends and zero otherwise, using the convention that $A_{iig} = 0$. I define two additional measures: the total number of $\{ig\}'s$ friends $L_{ig} \equiv \sum_{j=1}^{N_g} A_{ijg}$ and the total number of $\{ig\}'s$ treated friends by $T_{ig} \equiv \sum_{j=1}^{N_g}A_{ijg}D_{jg}$. The variables $L_{ig}$ and $T_{ig}$ are meant to capture peer influence in $\{ig\}'s$ immediate friend circle.
I analyze a model where a scalar outcome $Y_{ig}$ is determined by\footnote{In Appendix (ref) I show how to extend non-parametric identification results to models of the form $Y_{ig}=m(X_{ig},\tau_{ig})$, where $m(\cdot)$ is an arbitrary function, possibly unknown.}
where $\tau_{ig} \equiv (\alpha_{ig},\beta_{ig},\gamma_{ig}',\delta_{ig}')' \in \mathbb{R}^{2+2k}$ is a vector of real-valued random coefficients, and $X_{ig}$ is vector of endogenous regressors defined by $$ X_{ig}' \equiv
$$ Here, $\varphi:\mathbb{Z}_{+}^2 \to \mathbb{R}^k$ is a \textit{known} function which reflects the researcher's beliefs about how the treatment of others affects unit $\{ig\}$.\footnote{The function $\varphi(\cdot)$ could accommodate non-linearities by having different basis functions.} In the most general form, it can depend on $\{ig\}$'s position within network $A_g$, a vector of treatment of indicators $D_g$ for other people in the group, and a matrix of covariates $C_g$. It is common to restrict attention to functions that depend the treatment status of immediate neighbors. For example, if $\varphi(i,A_g,C_g,N_g) = \sum_{j=1}^{N_g}A_{ijg}D_{jg}/\sum_{j=1}^{N_g} A_{ijg} = T_{ig}/L_{ig}$, then the outcome is only determined by own treatment and the proportion of treated neighbors, and the outcome equation simplifies to
I am interested in identifying the average partial effects for a target population $\mathcal{F}$, defined as
The average partial effects vector $\tau$ integrates the coefficients in (ref). The conditioning $\mathcal{F}$ is important to emphasize that the average is computed for a specific subpopulation (men or women, old or young, etc.). When the conditioning set is empty, i.e $\mathcal{F} = \emptyset$, the average is computed for the entire population.
The main barrier to identifying the average partial effect is that $\tau_{ig}$ and $X_{ig}$ might be correlated. To address this problem, I propose a control variable $V_{ig}$ that captures the main determinants of treatment and network formation. I assume that $V_{ig}$ satisfies the unconfoundedness condition $\tau_{ig} \ \indep \ X_{ig} \mid V_{ig}$ and that $\mathcal{F}$ is $V_{ig}-$ measurable. For example, $\mathcal{F}$ could include gender and $V_{ig}$ could include a finer set of variables such as gender, age and wealth. I establish primitive assumptions on the network and treatment processes that justify these conditions in the next section. Under unconfoundedness wooldridge2003further,graham2022semiparametrically,wooldridge1999distribution, we can identify average partial effects, as follows
Intuitively, Theorem (ref) states that researchers can identify average partial effects by comparing the outcomes of individuals with similar values of $V_{ig}$ but different realizations of $X_{ig}$. As shown by graham2022semiparametrically, this estimand is equivalent to computing an OLS coefficient for each subset $\{V_{ig} = v \}$ and averaging the results.
However, finding a $V_{ig}$ that meets these properties in the network setting is challenging, at least without further assumptions on the outcome, treatment, and network processes. I will propose a strategy to construct $V_{ig}$ that depends on weak notions of exchangeability. Broadly, exchangeability refers to the idea that the distribution of variables for a set of individuals is ex-ante identical, and that the distribution is invariant to relabeling of the observations in the network.
I assume that the outcome model is determined by spillovers that are exchangeable in the identities of the individuals. Formally, let $\Pi_{ij}$ be an $N_g \times N_g$ rotation matrix. For an $n \times 1$ vector $x$, the matrix is designed in such a way that $\Pi_{ij}x$ exchanges the order of the $i^{th}$ and $j^{th}$ rows. Under this definition, the matrix $\Pi_{ij}A_g\Pi_{ij}'$ is an adjacency matrix that exchanges the neighbor connections of $\{ig\}$ and $\{jg\}$. Similarly, $\Pi_{ij}C_g$ exchanges the order of covariates. I assume that social interactions are exchangeable if the following assumption holds.
\nameref{assumptxt:exchangeableinteractions} states that if $\{jg\}$ had all of the same connections as $\{ig\}$, then they would have the same value of the exposure function. Therefore, the only thing that matters is the network structure and $\{ig\}$'s relative position within the network: not any particular labeling of the dataset. This assumption is clearly satisfied if $\varphi(i,A_g,C_g,N_g)$ is a function of the total number of treated friends, giving everyone equal weight. It can also holds in more complex cases where some neighbors are be more influential than others.
I assume that the researcher has auxiliary covariates that explain $\{ig\}'s$ participation in the treatment and choice of friends. As discussed before, let $C_{ig} \in \mathbb{R}^{d_c}$ be a vector of individual characteristics that are sampled at random from a super-population and define $\Psi_g^* \in \mathbb{R}^{d_\Psi}$ to be a vector of group characteristics. I next describe assumptions on the core structure that provide guidance on the choice of $V_{ig}$.
The first part of \nameref{assumptxt:randomsampling} --stating that groups are i.i.d-- is plausible when the groups are spatially, economically or socially separated. The second part states that the covariates within a group are conditionally independent within groups, which is a common assumption in the literature on network formation johnsson2019estimation,graham2017econometric,auerbach2019identification.
The \nameref{assumptxt:confounders} assumption states that the treatment status is independent of the treatment effects, after controlling for baseline characteristics. It puts the burden on researchers to identify relevant confounding variables (such as gender, income or age) that are motivated by either theory or practice. For example, the confounders can emerge from well-defined institutional rules that constrain the assignment of slots to treatment or the stratifying variables in experiments with perfect compliance. \nameref{assumptxt:confounders} is the same assumption discussed by rosenbaum1983central, which justifies propensity score analysis.
The \nameref{assumptxt:dyadicnetwork} assumption states that friendships between pairs of individuals $\{ig\}$ and $\{jg\}$ depend on their observed characteristics $(C_{ig},C_{jg})$, a group component $\Psi_g^*$ and a pair-specific shock $U_{ijg}$. For example, let $\Vert c - c^* \Vert$ be the Euclidean distance between two sets of covariates $(c,c^*)$. In a random geometric graph, $\mathcal{L}(c,c,\Psi^*,u) = \mathbbm{1}\{ \Vert c - c^* \Vert \le u \}$, which implies that individuals are more likely to be friends if their characteristics are similar. In economics, dyadic networks have been used to analyze risk sharing agreements, political alliances and business partnerships graham2017econometric, fafchamps2007risk,attanasio2012risk,lai2000democracy, fafchamps2018networks. The function $\mathcal{L}$ can be interpreted as a decision rule that encodes preferences over friends, as a random meeting process that brings two people together mele2017structural, or a combination of both.
Dyadic networks can also be motivated as reduced form objects by appealing to exchangeability. In two influential papers, aldous1981representations and hoover1979relations showed that any network whose distribution is invariant to the ordering of the sample (exchangeability) can be represented as a dyadic network, where some of the components of $C_{ig}$ are possibly unobserved. From a practical point of view, the \nameref{assumptxt:dyadicnetwork} assumption states that the relevant determinants are indeed observed by the researcher. Therefore it can be interpreted as a network analog of the \nameref{assumptxt:confounders} assumption.
I show that $V_{ig} = (C_{ig},\Psi_g^*)$ satisfies the key unconfoundedness condition of Theorem (ref) and can be used as a matching variable to compute the average partial effect.
Theorem (ref) suggests that the researcher should include all of $\{ig\}$'s covariates that she considers relevant for treatment participation and network formation in $V_{ig}$. However, if the assumptions of (ref) hold, then there is no need to control for the covariates of others. The variables $(C_{ig},\Psi_g^*)$ control for $\{ig\}'s$ friend preferences, and hence all the residual variation in $X_{ig}$ is exogenous. In practice, observing $(C_{ig},\Psi_g^*)$ may be a strong requirement and I propose some ways to relax the assumptions in Section (ref).
Intuitively, \nameref{assumptxt:randomsampling} and \nameref{assumptxt:dyadicnetwork} imply that $(C_{ig},\Psi_g^*)$ controls for others' treatment whereas \nameref{assumptxt:confounders} ensures that it controls for own selection. That means that $\{D_{jg},C_{jg},N_g\}_{j \ne i},\{A_{jj'g}\}_{j \ne j'} \indep \tau_{ig} \mid C_{ig},\Psi_g^*$. Property (i) holds because $X_{ig}$ is a function of the left-hand side terms. Property (ii) by \nameref{assumptxt:randomsampling} and because \nameref{assumptxt:exchangeableinteractions} ensures that function $\varphi(\cdot)$ is exchangeable. Property (ii) is particularly important because it shows that if $(C_{ig},\Psi_g^*)$ is observed, then we can identify the conditional distribution of $(Y_{ig},X_{ig})$ by pooling different observations. If exchangeability did not hold, this probability would be $\{ig\}$--specific and the quantity $\mathbf{Q}_{xx}(v)$ in Theorem (ref) would not be well-defined.
Most of the recent literature has focused on a particular type of exchangeable interactions, known as anonymous interference aronowsamii2017estimating,leung2019treatment,spilloversnoncompliance,forastiere2020identification,liu2019doubly,sussman2017elements,tchetgen2017auto. This assumes that individual $\{ig\}$ is only affected by the total number (or proportion) of treated friends. In this case neighbors are given equal weight, regardless of their characteristics, and no-weight is put on higher connections. This type of assumption is reasonable in cases where the spillovers are local and there is no reason to believe that one particular friend has more influence than others.
With a slight abuse of notation, let $\varphi(t,\ell)$ be an exposure function that depends on the total number of treated friends $t$ and the total number of friends $\ell$,
By design, \nameref{assumptxt:anonymousinteractions} is a special case of \nameref{assumptxt:exchangeableinteractions}. This added structure will lead to a lower-dimensional control variable. Define the propensity score and the friend propensity score, respectively as
The scalar $p_{dig}$ is the probability of treatment given individual characteristics, whereas $p_{fig}$ is the probability that a potential friend is treated. The \nameref{assumptxt:randomsampling} assumption ensures that every friend is ex-ante identical and hence the probability does not depend on the subscript $\{jg\}$. I call the three dimensional vector $P_{ig} = (p_{dig},p_{fig},L_{ig})$ the network propensity score. Before presenting the general results I focus on a special case where $\tau$ has a closed form expression. The following result in Theorem (ref) is a special case of Theorem (ref), by setting $V_{ig} = (C_{ig},\Psi_g^*,L_{ig})$ and imposing a particular set of basis functions.
Theorem (ref) shows that the average partial effects can be identified from $(p_{dig},p_{fig})$ and $(T_{ig},L_{ig},D_{ig},Y_{ig})$ for the subsample of individuals with at least one friend. The network propensity score is not observed directly but it can be identified from the data.
Lemma (ref) shows that the distribution of $(D_{ig},T_{ig})$ given $(C_{ig},\Psi_g^*)$ can be parametrized in terms of $P_{ig}$. Part (i) is an extension of the canonical result of rosenbaum1983central, whereas par (ii) is a new result. This factorization holds regardless of the primitive function $(\mathcal{L})$ and shock distribution of network formation. The proof builds on the insight that $T_{ig}$ is a sum of conditionally independent Bernoulli variables after conditioning on the key variables of network formation. Under model (ref), $X_{ig}$ is a deterministic function of $(D_{ig},T_{ig},L_{ig})$ which means that $P_{ig}$ also parametrizes the distribution of $X_{ig} \mid C_{ig},\Psi_g,L_{ig}$.
Theorem (ref) shows that $P_{ig}$ is a suitable generalization of the propensity score to setting with spillovers and network formation by showing that inherits two key properties. First, it is a balancing score which means that two individuals with the same value of $P_{ig}$ are guaranteed to have the same distribution of covariates $(C_{ig},\Psi_g^*)$. This property is important for causal analyses because it ensures that any matching procedure based on $P_{ig}$ will compare similar individuals. Second, it shows that $P_{ig}$ satisfies the unconfoundedness property required to identify the average partial effect $\tau$ in Theorem (ref). The selection on observables ties the observed characteristics $(C_{ig},\Psi_g^*)$ to the random coefficients and is therefore crucial to prove the final step.
From an economic point of view, the network propensity score can be interpreted as a function of agents' underlying preferences. To this end, it is convenient to represent $\{ig\}'s$ treatment indicator as $D_{ig} = \mathcal{H}(C_{ig},\Psi_g^*,\eta)$ where $\mathcal{H}$ is a measurable function and $\eta_{ig} \mid C_{ig},\Psi_g^* \sim F(\eta \mid c,\Psi^*) $ is an unobserved participation shock. Since we can always define the participation shock as $\eta = D_{ig} - \mathbb{P}(D_{ig} =1 \mid C_{ig} = c,\Psi_g^* = \Psi^*)$, this form does not entail any loss of generality. The function $\mathcal{H}$ can also take the form of a threshold utility model or an institutional assignment rule based on observables. The first component of the network propensity score is the propensity score conditional on $(C_{ig},\Psi_g^*)$, which is defined as
The propensity score depends on the preference function $\mathcal{H}$ and the distribution of selection shocks. The integral averages out the individual heterogeneity $\eta$, holding the characteristics $(c,\Psi^*)$ fixed.
The friend propensity score can be written in a similar way. Let $F(\eta^*,c^*,u \mid \Psi^*)$ be the distribution of traits of a potential friend in each group $(\eta^*,c^*)$ and the friendship shock $(u)$ given $\Psi_g^*$. By Bayes' rule
The friend propensity combines $\{ig\}'s$ friendship preferences/meeting likelihood and $\{jg\}'s$ preferences for participation in the program. In the extreme case that $L = \mathbbm{1}\{c = c^*\}$, agents only befriend others with exactly the same characteristics and the friend propensity score is equal to the propensity score. At the other extreme, when $L = \mathbbm{1}\{u > 0\}$ the network is exogenous then (ref) reduces to $\int \mathcal{H}(c^*,\Psi^*,\eta) dF(\eta^*,c^* \mid \Psi^*)$, which is a group-level constant. Conversely, when the treatment is exogenous, that is when $\mathcal{H}(c^*,\Psi^*,\eta^*) = \eta$ and $\eta$ is independent of the other characteristics, then the propensity score and the friend propensity score are constant. For intermediate cases the friend propensity score will not contain the same information as the propensity score.
In the microfinance example, homophily suggests people tend to associate with people in the same caste whereas selection implies that certain castes are more likely to participate in the information sessions. Consequently, the likelihood of having a treated friend depends on a household's caste. This is where homophily interacts with selection. Pairs of friends tend to have similar traits $(c,c^*)$ on average and consequently similar partipation probabilities, via the function $h$.
One potential problem with identification via network propensity score matching is that there may be some latent confounding variables. In some cases, researchers may be able to recover some unobservables from the network structure. In this section I explain the relationship between graphon notions of network distance and the network propensity score.
auerbach2022identification defines a pseudo-metric as follows\footnote{A similar metric was used by zeleneev2020identification to identify network formation models with interactive fixed effects.}
The pseudo-metric $d_{\Psi_g^*}$ measures measures the $L_2$ difference in link functions between someone with characteristics $C_{ig}$ and $C_{jg}$. Two individuals are close in this sense, if they have the same revealed preferences for neighbors, at least in reduced-form. In practice this can be computed by counting the relative number of friends in common between $\{ig\}$ and $\{jg\}$. auerbach2022identification proposed matching estimators for a partially linear model with an exogenous treatment based on the pseudo-metric, but noted that it could not be used to study spillovers because of the failure of a particular rank condition. I show that this finding extends to a more general setting with arbitrary functional form and selection-on-observables by establishing its connection to the friend propensity score.
The difference in the friend propensity scores of two individuals is equal to
The following theorem establishes the relationship between the two metrics.
Theorem (ref) states that if two individuals are close in the pseudo-metric then they also have identical friend propensity scores. In principle, if $d_{\Psi_g^*}$ were known, then researchers could match on the pseudo-metric rather than on the friend propensity score. The converse does not necessarily hold. For example, suppose that $\mathcal{H}(C^*,\Psi_g^*) = 0.5$, then $d_f = 0$ if and only if $\int \mathcal{L}(C_{ig},C^*,\Psi_g^*) = \int \mathcal{L}(C_{jg},C^*, \Psi_g^*)$. This occurs when two individuals have the same proportion of friends. However, $d_{\Psi_g^*} > 0$ if the $\{ig\}$ and $C_{jg}$ have different preferences over specific friends.
However, the pseudo-metric is not typically known. The problem is that current methods to consistently estimate $d_{\Psi_g^*}$ require a dense network with $p_{\ell}(C_{ig},\Psi_g^*) \to \infty$ in the large-network limit. This means that the number of friends $L_{ig}$ grows proportionately with the sample size, which violates the rank conditions to identify even simple estimands as those in Theorem (ref). In essence, two individuals that are close in the pseudo=metric will have almost identical values of the regressors $X_{ig}$, and there is no residual variation to identify the average partial effects. This makes it difficult to use network structure to recover unobserved heterogeneity.
To compute the network propensity score, $(C_{ig},\Psi_g^*)$ needs to be fully observed or be consistently estimated. Unobserved heterogeneity can be addressed in a variety of ways. The researcher may still be able to identify average partial effects, even in settings with unobserved heterogeneity, under further restrictions.
Lemma (ref) provides a high-level condition stating that any residual variation in the network propensity is exogenous after conditioning on $V_{ig}$. Since the network propensity score is itself a function of $(C_{ig},\Psi_g^*)$ this means that there are exogenous shifters in individual behavior $(C_{ig})$ or group contextual factors $(\Psi_g^*)$.
There are two examples in the literature that could satisfy this requirement. johnsson2019estimation propose a restriction on the class of network models that satisfy a particular monotonicity restriction, extending prior work by goldsmith2013social. In this case, the total number of friends is a sufficient statistic for network unobservables. spilloversnoncompliance find a single-dimensional control variable in experiments with spillovers and non-compliance, which is the share of individuals that accept treatment offers in each group. I fill in some of the details in Section (ref). spilloversnoncompliance find a sufficient statistic $V_{ig}$ exploiting the binomial form of their key endogenous variable. arkhangelsky2018role show how to exploit this type of structure to identify direct effects in models with selection-on-observables and fixed effects. They rely on the idea of obtaining sufficient statistics of the unobserved heterogeneity. Similar strategies could be applied in the spillovers case by imposing particular functional forms on the network or selection processes.
Imposing the conditions of Lemma (ref) has implications for the structure of the matrix $\mathbf{Q}_{xx}(v)$ defined in Theorem (ref). Define the functions $\widetilde{\varphi}_1(p_f,l) = \mathbb{E}[\varphi(T_{ig},L_{ig}) \mid p_{fig} = p_f,L_{ig} = l]$ and $\widetilde{\varphi}_2(p_f,l) = \mathbb{E}[\varphi(T_{ig},L_{ig})\varphi(T_{ig},L_{ig})' \mid p_{fig} = p_f,L_{ig} = l]$ which are the conditional first and second moments given the friend propensity score and the total number of friends. Since Lemma (ref) shows that $(p_{fig},L_{ig})$ parametrizes the distribution of $(T_{ig},L_{ig})$ given $(C_{ig},\Psi_g^*)$, these are equivalent to conditioning on $(C_{ig},\Psi_g^*)$ directly by the decomposition axiom Constantinou2017. Lemma (ref) also implies that $\widetilde{\varphi}_1$ and $\widetilde{\varphi}_2$ are known functions that only change depending on the basis $\varphi$. In our running example, where $\varphi(t,l) = t/l$ these function take a very simple form. In this case $\widetilde{\varphi}_1(p_f,l)$ and $\widetilde{\varphi}_2(p_f,l) = \frac{p_f(1-p_f)}{l} + p_f^2$.
Lemma (ref) shows that the matrix $\mathbf{Q}_{xx}$ can be expressed as a mixture of known functions of the network propensity score.
In the special case where $V_{ig} = (C_{ig},\Psi_g^*)$ the distribution $F$ is degenerate and we can drop the integral sign. Therefore, observing the key variables for selection and network formation imposes over-identifying restrictions on the weighting matrix. The integral is non-degenerate when some of these key variables are unobserved by the researcher. This assumption is testable by comparing the entries of $\mathbf{Q}_{xx}$. For example, in a parametric model $F$ can be modeled as a latent distribution that nests the degenerate case and $(p_{dig},p_{fig})$ as link function such as probit or logit. In the empirical example, I exploit this structure to produce a feasible parametric model.
I outline a two-step procedure to estimate the causal effects for linear models as a sample analog of the estimand of $\tau$. In the first stage, I fit a parametric model for $\mathbf{Q}_{xx}$ using data from the endogenous regressors $X_{ig}$ and the control variable $V_{ig}$. In the second stage, I substitute the estimated weighting matrix $\mathbf{Q}_{xx}$ to compute $\tau$ by inverse weighting.
Notation: Let $Z_{ig}$ denote a vector of individual variables, where $Z_{ig} \equiv (X_{ig},Y_{ig},V_{ig})$ includes the endogenous regresors, the outcome and the observed control variables. I let $\sum_{ig} f(Z_{ig})$ be the sum $\sum_{g=1}^G\sum_{i=1}^{N_g} f(Z_{ig})$, where $f(\cdot)$ is an arbitrary function. I also let $\bar{n} = \frac{1}{G}\sum_{g=1}^G N_g$ denote the average group size. By construction, $\bar{n}G$ is equal to the total sample size. For convenience, let $vec(\cdot)$ denote the vectorize operator, which stacks the columns of a matrix into a single vector. I also use $\Vert x \Vert $ to denote the Euclidean norm of the vector $x$, defined as $\Vert x \Vert = \sqrt{\sum_{k=1}^{K} x_{k}^2}$.
In the first stage, I consider a parametric class of functions to model the weighting matrix, $\{\mathbf{Q}_{xx}(v,\boldsymbol{\theta}): \boldsymbol{\theta} \in \boldsymbol{\Theta} \subseteq \mathbb{R}^{d_{\theta}} \}$, that nest the true model. This means that there is a $\boldsymbol{\theta}_0 \in \boldsymbol{\Theta}$ such that $\mathbf{Q}_{xx}(v,\boldsymbol{\theta}_0) = \mathbb{E}[X_{ig}X_{ig}' \mid V_{ig} = v]$. The matrix $\mathbf{Q}_{xx}$ has to be symmetric and positive semi-definite. If \nameref{assumptxt:randomsampling}, \nameref{assumptxt:confounders} and \nameref{assumptxt:dyadicnetwork} hold, and $V_{ig} = (C_{ig},\Psi_g^*)$ the choice of parametric family can be disciplined by imposing over-identifying restrictions of the network formation model, so that $\mathbf{Q}_{xx}(v,\boldsymbol{\theta})$ can be expressed as a function of the network propensity score. Alternatively we can use the mixture model representation of Lemma (ref) to inform the choice of $\mathbf{Q}_{xx}$ for other choices of $V_{ig}$. The control variable $V_{ig}$ is valid as long as the conditions of Lemma (ref) hold.
I define the vectorized residuals, \[ r(Z_{ig},\boldsymbol{\theta}) \equiv \text{vec}(X_{ig}X_{ig}' - \mathbf{Q}_{xx}(V_{ig},\boldsymbol{\theta})). \] The residuals capture how well the control variables fit $X_{ig}$. The sample criterion function computes the average of square residuals as
The sample criterion $\widehat{\mathcal{R}}(\boldsymbol{\theta})$ is an approximation to $\mathcal{R}(\boldsymbol{\theta}) = \mathbb{E}[\Vert r(Z_{ig},\boldsymbol{\theta})\Vert^2]$. The least squares criterion is appropriate for three reasons. First, the population criterion $\mathcal{R}(\boldsymbol{\theta})$ is minimized at $\boldsymbol{\theta}_0$ because the conditional mean of $X_{ig}X_{ig}'$ given $V_{ig}$ is the optimal prediction. This provides a rationale for minimizing $\widehat{\mathcal{R}}(\boldsymbol{\theta})$. Second, joint-likelihood approaches are either impractical or infeasible without strong assumptions, particularly with more complex exposure functions. Third, quasi-likelihood approaches, such as those in tchetgen2017auto and sofrygin2017semi are valid under certain assumptions, but are more sensitive to the specification of the model. My approach is more robust than quasi-likelihood methods because it targets the conditional mean directly, which is the main object required for identification.
We can construct a feasible estimator by minimizing the sample criterion,
The estimated parameter $\widehat{\boldsymbol{\theta}}$ can be plugged-in to compute a feasible weighting matrix $\mathbf{Q}_{xx}(V_{ig},\widehat{\boldsymbol{\theta}})$. I propose the following sample analog of the inverse-weighting estimand of $\tau$.
\[ \widehat{\boldsymbol{\tau}} \equiv \frac{1}{\bar{n}G} \sum_{ig} \mathbf{Q}_{xx}(V_{ig},\widehat{\boldsymbol{\theta}})^{-1}X_{ig}Y_{ig} \] The vector $\widehat{\boldsymbol{\tau}}$ is a feasible estimator of the average partial effects defined in (ref). The estimator is subject to two sources of uncertainty. First, the sample average is an approximation to $\mathbb{E}[\mathbf{Q}_{xx}(V_{ig})^{-1}X_{ig}Y_{ig}]$. Second, the inverse weighting method is subject to first-stage uncertainty in the estimation of $\widehat{\boldsymbol{\theta}}$. Under standard regularity conditions that I list in the Appendix, $\widehat{\boldsymbol{\theta}}$ and $\widehat{\boldsymbol{\tau}}$ are consistent but the standard errors need to be adjusted. This is analogous to the first stage uncertainty in propensity score methods, that can be corrected analytically or by bootstrap procedures abadie2016matching.
To adjust the standard errors it is useful to view the first and second stages as a single system of equations. As before, let $z \equiv (x,y,v)$. I write down the first-order conditions in terms of the jacobian of the square residuals $\psi_q(z,\boldsymbol{\theta}) = \tfrac{\partial}{\partial \boldsymbol{\theta}'}\Vert r(v,\boldsymbol{\theta})\Vert^2$ and the second stage influence function $\psi_{IW}(z,\boldsymbol{\theta}) = \mathbf{Q}_{xx}(v,\boldsymbol{\theta})$. I stack the first and second stage equations in a single influence function $\psi \equiv [\psi_q,\psi_{IW}']'$. The estimated parameters solve
To this end, I define the within-group average $\overline{\psi}_g(\boldsymbol{Z}_{g},\boldsymbol{\theta}) \equiv \frac{1}{N_{gt}}\sum_{i=1}^{N_{g}} \psi(Z_{ig},\boldsymbol{\theta})$, where $\boldsymbol{Z}_{g} \equiv \{ Z_{ig}\}_{i=1}^{N_{gt}}$ is a matrix of individual covariates for each group. This allows me to decompose (ref) into group averages as $\frac{1}{\bar{n}G}\sum_{ig} \psi(Z_{ig},\widehat{\boldsymbol{\tau}},\widehat{\boldsymbol{\theta}}) = \frac{1}{G}\sum_{g=1}^G \left(\tfrac{N_{g}}{\bar{n}}\right) \overline{\psi}_g(\boldsymbol{Z}_{g},\boldsymbol{\theta})$. The fraction $(N_g/\bar{n})$ denotes the relative size of each group.
For inference, I compute heteroskedasticity-robust standard errors, clustered at the group level. Let $\widehat{\Omega}$ be an estimate of the second moments of the influence function (ref) and let $\widehat{H}$ be a sample analog of the expected jacobian, defined as
Then the covariance of the estimators is computed by the sandwich form $\widehat{\Sigma} \equiv \tfrac{1}{G}\widehat{H}^{-1}\widehat{\Omega}\widehat{H'}^{-1}$ and the standard errors can be recovered from the square root of the diagonal of $\widehat{\Sigma}$. Since the estimator $\widehat{\tau} \in \mathbb{R}^{d_{\tau}}$ only enters the second stage linearly, \[ \widehat{H} = \frac{1}{G}\sum_{g=1}^G \left(\tfrac{N_{g}}{\bar{n}}\right)
\] Here, $\overline{\psi}_{q,g}$ and $\overline{\psi}_{IW,g}$ decompose the within-group average influence functions into the first and second stages, respectively. Both $\widehat{H}$ and its inverse are lower triangular, which means that the limiting covariance matrix of $\tau$ depends on the upper-left block of $\widehat{\Omega}$ (which captures the first-stage uncertainty).
For the remainder of this section I propose inference procedures for a setting with many groups $G \to \infty$ and allow for the possibility that $N_g$ is either fixed or growing with $G$. This is intended to approximate the situation faced by empirical researchers who randomly collect data from distinct geographic units, with few individuals (classrooms) or many individuals (villages, cities), which matches the data that I use in the empirical example. Formally, I assume that there is a sequence of probability distributions that is indexed by $t$, with $G_t$ groups of unequal size $N_{gt}$, and let $N_t \equiv \mathbb{E}[N_{gt}]$ denote the expected group size. There is a triangular array of covariates for individual $\{ig\}$ for the point $t$ in the sequence, which I denote by $Z_{igt} = (X_{igt},Y_{igt},V_{igt})$. The variables $(L_{igt},T_{igt})$ are the number of treated friends and number of friends, respectively. Similarly, for each $t$, I compute estimators $(\widehat{\boldsymbol{\theta}}_t,\widehat{\boldsymbol{\tau}}_t)$. The estimator $\widehat{\boldsymbol{\tau}}_t$, in particular is compared to the population quantity $\boldsymbol{\tau}_{0t} = \mathbb{E}[\tau_{igt} \mid \mathcal{F}_t]$. Centering the estimator around the mean of the triangular array is important to derive the right rate of convergence. For simplicity, I define $\rho_{gt}$ as the relative group size. Let $0 < \underline{\rho} < \overline{\rho} < 1$ be an arbitrary constant that I use throughout the derivation.
\nameref{assumptxt:boundedrelative_n} implies that all groups are approximately the same size, within a range. It implies that the ratio of the largest to the smallest group is bounded by $\overline{\rho}/\underline{\rho}$. This assumption is automatically satisfied when $N_{gt}$ is bounded. However, if $N_{t} \to \infty$ as $t \to \infty$, then the assumption implies that the smallest group size is growing, because $\inf_{g} N_{gt} \ge \underline{\rho}N_t \to \infty$ as $N_t \to \infty$. bester2016grouped propose a weaker assumption for large unbalanced panels, where the bounds hold in the limit experiment rather than for each point along the sequence, which leads to qualitatively similar conclusions.
My asymptotic results allow for some or all of the regressors in $V_{igt}$ to be estimated. For example, johnsson2019estimation show the estimator $L_{ig}/(N_g-1)$ converges uniformly to a measure of unobserved degree heterogeneity in dense networks, at rate $\sqrt{(\log N_t) / N_t}$ in sup-norm. In related work, spilloversnoncompliance find that in randomized experiments with non-compliance, the key dimensions of heterogeneity in spillover models is unobserved but can be consistently estimated in large groups, with a $\sqrt{(\log G_t) /N_t}$ uniform rate of convergence. Finally, researchers may also want to estimate group-level averages of the covariates that are consistent in large groups. I define $V_{igt}^{0}$ as the true, but unobserved value of the regressors. My asymptotic results simply require that $\max_{g=1,\ldots,G_t}\max_{i=1,\ldots,N_{gt}} \Vert V_{igt} - V_{igt}^{0} \Vert = O_p(\lambda_t)$ and that $\sqrt{G_t}\lambda_t = o(1)$. In the two examples above, this means that the expected size of each group needs to be large relative to the number of groups. This is plausible in situations where data is collected on large villages or other geographical units. If the key confounders are observed without error then $V_{igt} = V_{igt}^0$. Otherwise the condition holds trivially and $N_t$ does not need to grow with $G$ at any particular rate.
I list additional \nameref{assumptxt:regularity} in the Appendix, where I impose conditions on the moments of $(X_{igt},Y_{igt})$ and smoothness conditions on the function $\mathbf{Q}_{xx}(\cdot,\boldsymbol{\theta})$. In particular, I provide conditions that ensure that the weighting matrix is almost surely invertible, by imposing a lower bound on the eigenvalues of the matrix. When $V_{igt} = (C_{igt},\Psi_{gt}^*)$ and \nameref{assumptxt:randomsampling}, \nameref{assumptxt:confounders} and \nameref{assumptxt:dyadicnetwork} hold, this is equivalent to saying that $L_{igt}$ is bounded, and that the remaining components of the network propensity score are bounded in a compact subset of the unit interval, i.e. $p_d(C_{igt},\Psi_{gt}^*),p_f(C_{igt},\Psi_{gt}^*) \in [\underline{\rho},\overline{\rho}] \subset (0,1)$. That avoids boundary cases, where there is not enough residual variation in the regressors after conditioning on the controls. Finally, I define two more objects, $H_{0t} \equiv \mathbb{E}[(N_{gt}/N_t)\overline{\psi}_g(\mathbf{Z}_{gt},\widehat{\boldsymbol{\tau}},\widehat{\boldsymbol{\theta}})]$ and $\Omega_{0t} \equiv \mathbb{E}[(N_{gt}/N_t)^2\overline{\psi}_g(\mathbf{Z}_{gt},\widehat{\boldsymbol{\tau}},\widehat{\boldsymbol{\theta}})\overline{\psi}_g(\mathbf{Z}_{gt},\widehat{\boldsymbol{\tau}},\widehat{\boldsymbol{\theta}})']$ that are used to compute the covariance matrix $\Sigma_t \equiv H_{0t}^{-1}\Omega_{0t}H_{0t}^{-1}$.
Theorem (ref) shows that the estimators are consistent and converge to a normal distribution. The estimator is centered around the value of $(\boldsymbol{\theta}_{0t},\boldsymbol{\beta}_{0t})$ that solves the population criterion, at each point of the sequence. This allows for estimators that are consistent, even if the networks itself does not converge to any particular structure. Theorem (ref) can be viewed as an approximation to the finite sample behavior. Researchers can construct test statistics by substituting $\Sigma_t$ with a sample analog $\widehat{\Sigma}_t$ to confidence intervals.
My results are agnostic about the dependence structure across groups, but it may be possible to improve the $\sqrt{G_t}$ to $\sqrt{G_tN_t}$ under stronger conditions. For example, kojevnikov2020limit develop a central limit theorem for network dependence and provide specific regularity conditions for a single \nameref{assumptxt:dyadicnetwork}. This requires the network to be sparse $L_{igt}$ small relative to $N_{gt}$ so that individuals far apart in the network are approximately independent. In practice, this does not change the estimation procedure but rather the way in which we construct confidence intervals. kojevnikov2020limit propose a Network-HAC estimator and kojevnikov2019bootstrap proposes a bootstrap procedure. leung2019weak proposes similar limiting theory for spillover effects when the treatment is exogenously assigned, and chandrasekhar2014tractable propose alternative limit theorems under network dependence.
The balancing property in Theorem (ref) is testable. Parametric propensity score analyses typically conduct so-called covariate balancing tests. I propose an analogous “placebo” test, where the pretreatment covariates serve as an outcome variable. Let $\widetilde{V}_{ig}\in \mathbb{R}$ be a variable in the covariate set $V_{ig} = (C_{ig},\Psi_g)$. My test relies on the simple idea that $\widetilde{V}_{ig}$ can be decomposed as \[ \widetilde{V}_{ig} = \underbrace{\widetilde{V}_{ig}}_{\widetilde{\alpha}_{ig}} + 0 \times D_{ig} + 0 \times \varphi(i,A_g,D_g,C_g) + 0 \times D_{ig} \times \varphi(i,A_g,D_g,C_g) \] Let $\widetilde{\tau}_{ig} = (\widetilde{V}_{ig},0,0,0)$ is the vector of coefficients of the placebo outcome. It is easy to verify that $X_{ig} \ \indep \ \widetilde{\tau}_{ig} \mid V_{ig}$ since $\widetilde{\tau}_{ig}$ is a measuarable function of $V_{ig}$. Therefore by Theorem (ref), $\mathbb{E}[\mathbf{Q}_{xx}(V_{ig})^{-1}X_{ig}\widetilde{V}_{ig}] = (\mathbb{E}[\widetilde{V}_{ig}],0,0,0)$. Therefore, when $\mathbf{Q}_{xx}(v)$ is properly specified the researcher can test the null hypothesis that the slope coefficients are zero. This test only uses information about the treatment, the network and the covariates but not the outcome. In practice the test could be rejected in a parametric settings if the functional form is not flexible enough. However, it could also be rejected because a violation of the over-identifying restrictions imposed by the \nameref{assumptxt:randomsampling} and \nameref{assumptxt:dyadicnetwork} assumptions. The researcher may want to check whether there are omitted variables that might influence network formation or treatment.
I evaluate the role of an intervention on political participation in Uganda eubank2019viral,ferrali2020takes. U-Bridge is a novel political communications technology that allows citizens to contact district officials via text-messages. In a pilot program, individuals in 16 villages were invited to participate in quarterly meetings, at a central location, where they received information about national service delivery standards and ways to communicate with local officials. The Governance, Accountability, Participation, and Performance (GAPP) program collected survey data on 82% of adults in the 16 villages as well as social network data. ferrali2020takes evaluated the adoption patterns of U-Bridge a couple years later. eubank2019viral study the role of social network structure on voting patterns. For my analysis, I evaluate the impact of attendance to UBridge meetings on political participation using the network propensity score matching methodology. Spillovers are likely to occur in this context because non-participants can receive information about ways to engage in politics from their friends, which can increase their own political activity.
The data collected by the researchers contains four types of social networks: Family ties, friendships, lenders and problem solvers. In my analysis, $\{ig\}$ is an identifier for an adult in the pilot villages. The indicator $A_{ijg}$ equals one if $\{ig\}$ and $\{jg\}$ have a connection along any of the four dimensions and zero otherwise. Under this definition, individuals have 10 connections on average. The indicator $D_{ig}$ equals one if $\{ig\}$ attended the Ubridge meetings, which is around 8.6% of the sample. The outcome is a continuous variable $Y_{ig}$ that denotes a political participation index constructed by ferrali2020takes. Table (ref) presents summary statistics comparing the treatment and control group. The average adult in the sample is around 40 years old. Men are more likely to attend the session than women. Individuals that a leader position and/or completed their secondary education are more likely to attend as well.
I estimate the following linear model with random coefficients.
Heterogeneity of $\beta_{ig}$ means that agents engage in varying levels of political activity after attending the meeting. In this case, we expect $\beta_{ig}$ to be close to zero because individuals that are already politically engaged are the ones opting to go to the meetings. Conversely, $\gamma_{ig}$ is the effect of peers on non-participant adults. If $\gamma_{ig} > 0$, then individuals with a larger fraction of treated friends are more politically active. The coefficient $\gamma_{ig} + \delta_{ig}$ captures the spillovers for participants. In this case we expect $\delta_{ig} <0 $ because the marginal effect of attending friends is lower because they are already receiving the information first hand.
There is a potential identification in this example because individuals select connections with similar preferences. We expect $(\gamma_{ig},\delta_{ig})$ to be correlated with $(T_{ig}/L_{ig})$. To address this problem I leverage additional covariates collected by the researcher to tease out the causal effects. The network propensity score matching methodology is the appropriate tool to identify the average partial effect $\tau$ because it allows to incorporate additional covariates while allowing for heterogeneous causal effects $\tau_{ig} = (\alpha_{ig},\beta_{ig},\gamma_{ig},\delta_{ig})$.
The propensity score in this case describes the probability of attending an Ubridge meeting given covariates $C_{ig} = (C_{ig1},\ldots,C_{igK})$. These include an indicator for holding a leadership position in the village, gender, an indicator for secondary education, a self-reported relative income measure, distance to the meeting place, number of friends and age. ferrali2020takes also incorporated a public goods question where participants were asked to donate part of their remuneration to the village that were match researchers. The donation amount is meant to capture pro-sociability attitudes.
I assume that the group-level variation $\Psi_g^*$ has an observed and an unobserved component. For the observed component, I include a vector of group-level averages of the key variables in $C_{ig}$, which I denote by $\Psi_g$. I assume that $\Psi_g^*$ has a bivariate structure with mean $(\Psi_g'\boldsymbol{\theta}_{d\Psi},\Psi_g'\boldsymbol{\theta}_{d\Psi})'$, where $(\boldsymbol{\theta}_{d\Psi},\boldsymbol{\theta}_{f\Psi})$ is a vector of parameters to be estimated. The error term of $\Psi_g^*$ follows a normally distributed random-effects structure with covariance matrix $\Sigma \equiv (\sigma_{11}^2,\sigma_{12},\sigma_{12},\sigma_{22}^2)$, that is assumed to be independent of the observed covariates and the random coefficients $\tau_{ig}$. The coefficient $\sigma_{12}$ captures the correlation between the two unobserved components of $\Psi_{g}^*$. Formally, \[ \Psi_g^* = \left(
\right) \sim \mathcal{N}(\mu_g,\Sigma), \qquad \mu_g =
, \qquad \Sigma = \left(
\right). \] I assume that the own propensity score takes the form of a logit function with an associated vector of parameters $\boldsymbol{\theta}_{d} = (\theta_{d0},\theta_{d1},\ldots,\theta_{dK})$ as follows \[ p_d(C_{ig},\Psi_g^*;\boldsymbol{\theta}_d) = \frac{\exp(\theta_{d0} + \sum_{k=1}^{K} C_{igk}\theta_{dk} + \Psi_{gd}^* ) }{1+\exp(\theta_{d0} + \sum_{k=1}^{K} C_{igk}\theta_{dk} + \Psi_{gd}^* )} \] I similarly construct the friend propensity score using a logit link function. I use the same observables variables as the friend friend propensity score with different coefficients $\boldsymbol{\theta}_{f} = (\theta_{f0},\theta_{f1},\ldots,\theta_{df})$ as follows \[ p_f(C_{ig},\Psi_g^*; \boldsymbol{\theta}_f) = \frac{\exp(\theta_{f0} + \sum_{k=1}^{K} C_{igk}\theta_{fk}+ \Psi_{gf}^* )}{1+\exp(\theta_{f0} + \sum_{k=1}^{K} C_{igk}\theta_{fk} + \Psi_{gf}^* )} \]
The full vector of parameters to be estimated is \[ \boldsymbol{\theta} \equiv (\boldsymbol{\theta}_{d\Psi},\boldsymbol{\theta}_{f\Psi},\sigma_{1}^2,\sigma_{12},\sigma_2^2,\boldsymbol{\theta}_{d},\boldsymbol{\theta}_{f}). \] Let $F(\Psi_g^* ; \boldsymbol{\theta})$ is the distribution of unobserved heterogeneity, which corresponds to that of a normal distribution with parameters $(\mu_g,\Sigma)$. I construct a weighting matrix that satisfies the mixture model representation of Lemma (ref), where $V_{ig} = (C_{ig},\Psi_g,L_{ig})$. To simplify notation I define the auxiliary matrix \[ \Lambda(C_{ig},\Psi_g^*,L_{ig};\boldsymbol{\theta}) =
\] The feasible weighting matrix is equal to
where I evaluate the integral numerically via quadrature methods and estimate the parameter $\boldsymbol{\theta}$ by minimizing the sample criterion function in (ref).
Table (ref) reports the estimated parameters. Column (2) shows the coefficients of the propensity score. None of the variables in $C_{ig}$ appears to be statistically significant. Column (3) reports the coefficients of the friend propensity score, which are far more interesting. The evidence suggests that individuals that hold a leadership position and have completed a higher education or more likely to have a treated friend. Similarly individuals in villages where individuals perceive themselves as wealthier are more likely to see engagement with the U-Bridge sessions. Figure (ref) plots the propensity score and friend propensity score, integrating out the heterogeneity $\Psi_g^*$. Each score contains complementary information about the selection patterns. Finally to test the fit of the model I run a covariate test / placebo test by replacing the outcome variable in (ref) with each of the controls used in the analysis. None of the placebo coefficients are statistically significant for 16 out of the 20 variables. There are slight imbalances on one the relative income indicators, the distance to meeting and the average sociability.
Table (ref) reports estimates of the average partial effects. Column (2) shows the coefficients under the network propensity approach. The direct effect $\beta$ is positive but not statistically significant at the 10% level. The spillover effect $\delta$ increases the participation index by 0.348 points, which is significant at the 1% level. This effect is quantitatively large relative to the standard deviation of the political participation index, which is around 0.567 points. This finding appears to suggest that the intervention had a large spillovers on non-participants, who increased their political activity. The interaction coefficient $\delta$ is negative but not statistically significant at the 1% level. The results are consistent with the idea that the intervention had limited effects direct treatment effects, but promoted spillover effects on participants' social connections. Column (3) shows benchmark coefficients from an OLS regression with additive covariates. On one hand, the OLS coefficient of $\beta$ is also not statistically significant at the 10% level. On the other hand, the OLS coefficient of $\gamma$ is statistically significant but roughly half the size of the network propensity estimate. Finally, the coefficient of $\delta$ is positive and statistically significant. The discrepancies in the results for $\gamma$ and $\delta$ can be explained by interactive spillover effects $\gamma_{ig}$ and $\delta_{ig}$ that are not captured by the additive OLS model.
In this section I re-evaluate a program that encouraged the adoption of microfinance in rural areas of Southern India, by inviting select households to participate in an information about the program banerjee2013diffusion. Participant households were more likely to take out a loan. Spillovers are likely to occur in this context due to information transmission between participants and non-participants, and peer pressure to adopt.
The outcome is a binary variable $Y_{ig}$ that is equal to one if household $\{ig\}$ took out a loan when researcher followed-up a few months later. I estimate the following linear probability model with random coefficients.
Heterogeneity of $\beta_{ig}$ in the microfinance example means that some households are more likely to take-out a loan after the information session than others. Conversely, heterogeneity of $\gamma_{ig}$ and $\delta_{ig}$ means that not every household is equally likely to get in debt after receiving information from their friends. The coefficient $\delta_{ig}$ is the difference in spillovers effects between participant and non-participant households.
Identification of the average partial effect $\tau \equiv (\alpha,\beta,\gamma,\delta)$ is particularly challenging in this setting, however, because the treatment was not randomly assigned. The microfinance organization followed a fixed targeting strategy in each village, that selected shopkeepers, teachers and related occupations. However, Table (ref) shows that treated households were wealthier; they were more likely to have stone or concrete houses as opposed to tile or thatch, have private electricity, more bedrooms, and own a latrine. For instance, the treated were 13.45% more likely to have access to some form of sanitation, with either a private or public latrine. These differences are statistically significant at the 5% level, using clustered standard errors by village. There were also significant differences by caste, a hereditary social category that still defines many social boundaries, with household of so-called “general caste” more likely to be treated as opposed to minorities.
To measure social network links, banerjee2013diffusion collected twelve different definitions of the network at baseline, including favor exchange, commensality and community activities. I choose a conservative definition of the network, such that $A_{ijg}$ is equal to one if respondents reported a link along any of the dimensions. Figure (ref) plots the resulting degree distribution, which shows that the treated had a higher number of friends. Households have around ten friends on average, which is around 5% of the average village size. Figure (ref) shows that households reported that most of their friends were in the same broad caste category. A significant portion of the households reported that all of their friends were in the same category. The histogram shows that the treated had more diversified friendships, in the sense that they had fewer friends of the same caste.
To estimate the network propensity score I use the same specification as in the example for Uganda. The second and third columns of Table (ref) show the coefficients of the own propensity score and the corresponding standard errors. The structural parameters confirm the descriptive evidence. The number of rooms in the house, as well as the access to sanitation and electricity are statistically significant at the 5% level. Individuals of general caste and more connections, are more likely to be part of the program, even after accounting for asset measures. The observed group covariates are not statistically significant at the 10% level. Conversely, the fourth and fifth columns show estimated coefficients of the friend propensity score and their standard errors. Only the sociability index and the general caste indicator are statistically significant. This suggests that caste plays a crucial role on the interplay between homophily and selection. Treated individuals of general caste are more likely to befriend other treated individuals in their same caste category. The results also show that the unobserved heterogeneity parameters are not statistically significant at the 10% level.
Table (ref) computes the treatment effects using my proposed inverse-weighting (IW) procedure and an ordinary least squares (OLS) regression that includes the covariates as additive controls. The IW results show that participants in the information session (leaders) are 8.5% more likely to take-out a microfinance loan after controlling baseline characteristics, and is significant at the 1% level. The value of the direct effect is 1% higher than the effect estimated by OLS. The OLS regression only controls for additive heterogeneity, but it does not account for the possibility of heterogeneous slopes/treatment effects. The spillover effect is not significant in either case. That means that local variation in treated friends does not affect the outcome, on average.
Many programs offered by governments and non-profit organizations are not randomly assigned. Individuals typically select into treatment based on a set of eligibility criteria. Furthermore, social networks typically exhibit homophily: individuals tend to befriend others with similar characteristics. The interaction between selection and friendship homophily is not well understood, and hence the strategies to identify spillovers from observed social networks in this context are underdeveloped. This is particularly important for policy evaluators because estimating spillovers is crucial for cost-benefit calculations and understanding potential side-effects on non-participants.
This paper proposes a novel strategy for identifying average treatment effects and average spillover effects in settings with endogenous network formation and selection on observables. In particular I show that controlling for the key determinants of friendship decisions in a dyadic network model can account for possible confounders in the estimation of spillovers. I introduce a lower dimensional statistic, the network propensity score, which summarizes the key confounders and illustrates the crucial interplay between homophily and selection. I propose a two-step semiparametric estimator of the average effects in a class of random coefficient models, which is consistent as the number and size of the network grows. I apply my estimator to an intervention to encourage political participation in Uganda where I find evidence of spillovers on non-participants, and a microfinance application in India, where I document large direct effects but no meaningful local spillovers.