Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
183,678 characters · 31 sections · 54 citation commands
Unobserved Heterogeneous Spillover Effects in Instrumental Variable Models
\onehalfspacing
{ Keywords: Unobserved heterogeneous spillover/direct effect; violation of SUTVA; causal inference; instrumental variable.
}
In econometric analyses of treatment effects, the Stable Unit Treatment Value Assumption (SUTVA) is typically imposed, requiring that each individual's potential outcomes depend only on their own treatment assignment and not on the treatment assignments of others. This assumption, however, may fail in environments involving social interactions or group structures, where one unit's treatment can influence another's outcome. In such contexts, the SUTVA is unlikely to hold. In addition, treatment assignment may be endogenous, particularly in observational studies or randomized experiments with imperfect compliance, which further complicates the identification of causal parameters.
This paper develops a new framework for identifying causal effects in environments with within-group spillovers and endogenous treatments, such as education decisions among friends or pricing choices in oligopolistic markets. The framework explicitly accounts for two sources of SUTVA violations: individual outcomes may depend on peers' treatment selection, and individual treatment decisions may be influenced by instruments assigned to other group members.
To begin, the paper introduces the generalized local average controlled spillover and direct effects (LACSEs and LACDEs), which measure treatment effects from peers and from one's own treatment, respectively, for specific subpopulations. These parameters extend the local average treatment effect (LATE) framework imbens1994identification by allowing for spillovers in both outcomes and endogenous treatment decisions. This paper is the first to formally establish conditions under which the LACSEs and LACDEs can be point identified for specific subpopulations, and to derive general point identification results that do not rely on the cardinality of the support of instrumental variables. The results characterize the precise variation in instruments required to achieve point identification of local effects in environments with within-group spillovers and endogenous treatments.
When the instrumental variables exhibit continuous variation, the analysis introduces the marginal controlled spillover and direct effects (MCSEs and MCDEs), which measure the effects of peers' and individuals' own treatments conditional on specific values of unobserved characteristics within the group. These parameters naturally extend the marginal treatment effect (MTE) framework heckman2001policy,heckman2005structural to settings with spillovers. The paper first establishes the nonparametric point identification of the MCSEs and MCDEs from continuous variation in instruments, without imposing functional form restrictions on the outcome equation or the joint distribution of unobserved characteristics across group members, thereby accommodating flexible forms of spillovers. Similar to the standard MTE, the MCSEs and MCDEs serve as building blocks for identifying a broad class of policy-relevant treatment effects (PRTEs), including the LACSEs and LACDEs, as well as other PRTEs arising from counterfactual policy changes, facilitating policy evaluation in environments with spillovers.
Section (ref) develops the framework with within-group spillovers and endogenous treatments, formally defining, identifying, and analyzing the causal parameters of interest. Section (ref) introduces an outcome model that allows each unit's potential outcome to depend flexibly on the entire vector of treatments within the group, capturing spillover effects from peers' treatment decisions on own outcomes. Treatment selection follows a single-index threshold-crossing structure, in which an individual receives treatment if her unobserved characteristic falls below a threshold function determined by her own and her peers' instrumental variables. This structure, which can be interpreted as the equilibrium behavior of a simultaneous incomplete-information game aradillas2010semiparametric, captures how peers' instruments can influence individual treatment decisions. Importantly, this framework imposes no parametric restrictions on the outcome equation or the threshold function and allows for arbitrary dependence among unobserved characteristics across group members, accommodating a broad range of environments in which spillovers may be present.
Section (ref) defines and identifies two causal parameters, the generalized local average controlled spillover and direct effects. The term “generalized local" indicates that these effects are defined for specific subpopulations of groups, while “spillover" and “direct" refer to the sources of treatment variation from peers and from the individual herself, respectively. The term “controlled" highlights that one treatment dimension, either own or peer, is held fixed when measuring the effect of the other. This section establishes general conditions under which the LACSEs and LACDEs are point identified, without requiring the instrumental variables to be either discrete or continuous, as long as they generate the required variation for identification. These identification conditions also clarify the rationale for additional restrictions, such as one-sided noncompliance, which are often imposed to achieve point identification with binary instruments (e.g., vazquez2023causal).
Section (ref) establishes that when instrumental variables exhibit continuous variation, the marginal spillover effect and marginal direct effect are nonparametrically point identified without requiring functional form assumptions on the outcome equation. In addition, the joint distribution of unobserved characteristics across group members is nonparametrically identified over the support of the observed treatment probabilities, without imposing parametric restrictions on how these unobserved factors are distributed. The MCSEs and MCDEs are defined analogously to the LACSEs and LACDEs but condition on a specific realization of unobserved characteristics within each group. By conditioning on the latent characteristics of all group members, these parameters flexibly capture heterogeneity in both direct and spillover effects that arise from variation in unobserved factors. Section (ref) further demonstrates that the marginal controlled effects form the basis for identifying a broad class of policy-relevant treatment parameters. By integrating the MCSEs and MCDEs over appropriate regions of the unobserved heterogeneity distribution, one can recover the LACSEs, LACDEs, and other treatment effect parameters associated with counterfactual policy interventions.
Section (ref) formally compares the MCSEs and MCDEs with the standard MTE and shows that, in the presence of spillovers, the conventional MTE may lose its causal interpretation, whereas in the absence of spillovers, the MCSEs and MCDEs coincide with the standard MTE. These results demonstrate that the MCSE-MCDE framework provides a natural generalization of the MTE framework to accommodate environments with spillovers. Section (ref) derives testable implications implied by the model structure and the identification assumptions.
Section (ref) develops a semiparametric estimation procedure for the MCSEs and MCDEs, extending the framework of carneiro2009estimating to accommodate within-group spillovers. The proposed approach mitigates the curse of dimensionality associated with covariates while maintaining the model's nonparametric flexibility, as the key structural components other than the covariate adjustment are left unrestricted. The section also establishes the asymptotic properties of the semiparametric estimators. Because these estimators converge at nonparametric rates, their finite sample precision may be limited in small samples or when groups include a large number of members. To address this concern, a complementary parametric framework is introduced, relying on intuitive assumptions that enable straightforward implementation and facilitate valid inference through nonparametric bootstrap methods.
Section (ref) presents both parametric simulation results and an empirical application. Section (ref) presents Monte Carlo simulation results for the parametric estimation procedure, demonstrating the strong finite-sample performance of the proposed parametric methods. Section (ref) implements the proposed framework empirically using the parametric procedure. The analysis examines how education attainment affects long-term earnings within best-friend groups, drawing on data from the National Longitudinal Study of Adolescent to Adult Health (Add Health). The results indicate positive dependence between friends' unobserved characteristics and reveal systematic heterogeneity in the marginal controlled direct and spillover effects. The estimated MCDEs of completing 16 years of education are significantly positive when the best friend has also attained this level of education across most values of the latent characteristics, but become statistically insignificant when the friend has not. Similarly, the estimated MCSEs are significantly positive for individuals who completed 16 years of education across most values of the latent characteristics, whereas for those who did not, the spillover effects are insignificant and even negative for certain ranges of unobserved heterogeneity. These results provide empirical evidence of heterogeneous spillover effects of education on long-term earnings within friendship networks, highlighting how their magnitude and direction depend on both individuals' and peers' education attainment.
The framework can be extended to accommodate additional settings. Section (ref) generalizes the analysis to cases where outcomes depend on an exposure mapping, which is a known function of group members' treatment statuses, rather than the full treatment vector. This extension is particularly relevant when group sizes vary or are large, making the full treatment representation impractical. The section formally defines the MCSEs and MCDEs under this extended setting and establishes their nonparametric point identification using continuous instrumental variables.
Recent research has devoted increasing attention to the identification and estimation of treatment effects in the presence of spillovers. This paper contributes to several key strands within this growing body of work.
A common strategy for addressing interference has been to impose parametric structures on social interactions. For instance, manski1993identification discussed the linear-in-means model, formulated as a system of linear simultaneous equations to capture endogenous, exogenous, and correlated peer effects. Building on this result, subsequent work, such as bramoulle2009identification and blume2015linear, extended the framework to more complex forms of interaction within linear models and derived conditions under which social effects can be identified. However, they fundamentally rely on correct parametric assumptions regarding the structure of social interactions. Such assumptions may lead to model misspecification, particularly in the presence of nonlinear spillovers or heterogeneity across individuals. The framework developed in this paper departs from such reliance on parametric restrictions by studying identification under a nonparametric structure in both the outcome equation and the treatment selection mechanism. This design accommodates flexible and potentially complex spillover mechanisms and provides a robust framework for causal analysis in environments with within-group interactions.
In the main setting considered, an individual's outcome depends on the full vector of treatments within the group, consistent with the treatment response framework of manski2013identification. Within the context of randomized controlled trials (RCTs), hudgens2008toward and aronow2017estimating, along with related studies, formalized design-based frameworks for analyzing interference. The framework in this paper extends this line of research by allowing for noncompliance, so that individuals may not adhere to their assigned treatments. This feature is important in observational studies, where treatment status is not fully controlled by the researcher, and in experimental settings where imperfect compliance may occur. The analysis adopts a large-sample framework rather than a design-based approach to study causal identification under endogenous treatment selection.
vazquez2023causal employed a potential outcomes framework to analyze similar settings with spillovers operating through both outcomes and treatment selection, using a binary instrumental variable for identification. His approach classifies individuals into discrete compliance types according to how their treatment choices respond to changes in instruments and focuses on identifying local average spillover and direct effects for each specific type, which are closely related to the LACSEs and LACDEs introduced in this paper. He achieved point identification by excluding certain subpopulations under one-sided noncompliance, a restriction also used in related work such as ditraglia2023identifying. This paper establishes general identification conditions for the LACSEs and LACDEs, which clarify why one-sided noncompliance is required for point identification when the instrumental variable is binary. This paper further employs a continuously distributed instrumental variable to point identify the marginal controlled direct and spillover effects, defined conditional on continuous realizations of latent characteristics within groups. This approach connects the analysis to the marginal treatment effect literature and establishes the marginal effects as fundamental components for identifying a wide class of policy-relevant treatment effects. In particular, aggregating these marginal effects recovers the local average direct and spillover effects in vazquez2023causal, as well as other causal parameters under counterfactual policy interventions.
Recent studies, including balat2023multiple and hoshino2023treatment, use instrumental variable methods to identify spillover effects in settings with direct strategic interactions among agents, where each individual's treatment choice directly depends on the treatment decisions of other group members. Frameworks with direct strategic interactions assume that an individual's treatment decision does not directly depend on the instruments assigned to other group members, thereby ruling out spillovers from peers' instruments in the treatment selection process. In contrast, the framework developed here does not model explicit strategic interactions in treatment choices but allows each individual's treatment decision to depend on instruments assigned to other group members. This structure can be interpreted as the equilibrium outcome of a simultaneous incomplete-information game, following aradillas2010semiparametric, and thus provides a complementary perspective to models that incorporate direct strategic interaction.
The spillover framework developed in this paper and the multivalued treatment framework are not nested. When the group is treated as a single decision-making unit, the group treatment vector can be reformulated as a multivalued group-level treatment. This links the setting to multivalued MTEs such as lee2018identifying. When applied to spillover contexts, identification in lee2018identifying relies on an exclusion restriction that an individual's treatment does not depend on peers' instruments, whereas the framework developed here allows and models such spillovers from peers' instruments.
I consider a sample of $G$ independent and identically distributed (i.i.d.) groups, indexed by $g = \{1, \cdots, G\}$. Each group consists of the same number of units, denoted by $n \geq 2$. For example, a group may correspond to a market with several competing firms or to a household with multiple members. Within each group, units are indexed by $i = \{0, \cdots, n-1\}$. Throughout, I assume that spillover effects operate only within groups and do not extend across groups.
Researchers are often interested in how a treatment affects an outcome. Let $Y_{ig}$ denote the outcome of interest for unit $i$ in group $g$, and let $\mathcal{Y}$ denote its support. In some settings, the outcome $Y_{ig}$ may depend not only on unit $i$'s own treatment status but also on the treatment choices of other units within the same group. For example, a firm's market share is influenced both by its own pricing decisions and by those of its competitors. Within a friendship network, an individual's labor earnings may be influenced by her best friend's education attainment, not only through direct support or access to resources, but also through information-sharing or social learning mechanisms that facilitate the transmission of knowledge about opportunities, norms, and strategies. In such contexts, the Stable Unit Treatment Value Assumption (SUTVA) may be violated, which motivates researchers to develop models that explicitly allow for spillover effects in outcomes.
The binary treatment decision of unit $i$ in group $g$ is denoted by $D_{ig} \in \{0,1\}$, where $D_{ig} = 1$ indicates that unit $i$ adopts the treatment and $D_{ig} = 0$ otherwise. In many applications, treatment decisions are not randomly assigned but instead depend on unobserved characteristics that also influence outcomes, giving rise to endogeneity concerns. For instance, a firm's pricing decision or an individual's education choice may both be endogenously determined. In group settings, units may make their decisions simultaneously, taking into account private information as well as expectations about the behavior of other group members. Each unit's decision depends on its own private information, denoted by $V_{ig}$, as well as on its expectations about the probability that other members of the group will adopt the treatment. Units form their expectations on the basis of publicly observed variables $(Z_{ig}, Z_{-ig})$, where $Z_{ig}$ denotes the random assignment received by unit $i$ in group $g$, and $Z_{-ig}$ denotes the assignments of the remaining group members. The vector $(Z_{ig}, Z_{-ig})$ thus serves as the set of instrumental variables for addressing endogeneity in treatment decisions. Since each unit's treatment choice may respond to the assignments received by other group members, these instruments can also induce spillover effects in treatment selection.
Building on the setting described above, consider the following model for unit $i$ in group $g$, where the peer of unit $i$ is denoted by $-i$. For clarity of exposition, I focus on the case in which each group $g$ consists of two units, indexed by $i = \{0,1\}$, while noting that the identification and estimation results extend straightforwardly to groups with more than two members:
The first line of Equation (ref) specifies the outcome equation. In this framework, unit $i$'s outcome $Y_{ig}$ depends on her own treatment $D_{ig}$ and on her group member's treatment, $D_{-ig}$, which explicitly models spillover effects. Importantly, I also allow $Y_{ig}$ to depend on both unit $i$'s own unobservables $U_{ig}$ and the unobservables of her group member, $U_{-ig}$. For example, in the oligopoly market, this specification captures the possibility that firm $i$'s market share $Y_{ig}$ is influenced not only by its own unobserved product characteristics $U_{ig}$ but also by the unobserved product characteristics of its rival, $U_{-ig}$.
Throughout the paper, I define the potential outcome for unit $i$ when her treatment is set to $d$ and her group member's treatment to $d'$ as $Y_{ig}(d, d') \equiv m_i(d, d', U_{ig}, U_{-ig})$. The observed outcome $Y_{ig}$ is determined according to the following equation,
This paper focuses on identifying reduced-form causal effects arising from a unit's own treatment and peers' treatments, rather than the underlying structural parameters specified in structural equations. A detailed comparison with a system of structural equations is provided in the Appendix (ref).
The framework imposes no functional form restrictions on the outcome equation $m_i$, and the subscript $i$ indicates that each group member, $i \in \{0,1\}$, may have a distinct outcome equation, meaning their functional forms are not required to be identical. It also places no restrictions on the dimension of the unobserved components $(U_{ig}, U_{-ig})$. Consequently, the influence of the peer's treatment $D_{-ig}$ on unit $i$'s outcome $Y_{ig}$ remains fully unrestricted. This generality provides a flexible structure that accommodates complex and heterogeneous spillover patterns in outcomes.
The second line of Equation (ref) characterizes the treatment selection mechanism. The treatment decision of unit $i$, $D_{ig}$, may be endogenous because it is determined by a continuous unobserved factor $V_{ig}$ that can also influence the outcome. The treatment selection $D_{ig}$ depends only on the unit's own unobservable $V_{ig}$ and not directly on her group member's unobservable $V_{-ig}$. This restriction is plausible in many applications. For example, in the oligopoly market discussed above, a firm's pricing decision is driven by its own private demand shock, while the competitor's demand shock is unobserved and therefore cannot directly affect the firm's decision rule. In the returns to education example, an individual's education decision is determined solely by her own education costs. The best friend's education costs, which are unobserved to the individual, do not directly influence her schooling decision.
Crucially, it is empirically reasonable to allow the unobserved factors $V_{ig}$ and $V_{-ig}$ to be arbitrarily dependent, since group members often share related unobserved characteristics or are exposed to common shocks. This dependence further complicates identification, and the framework accommodates it without imposing parametric restrictions on the joint distribution of unobservables. This flexibility accommodates a wide range of empirically relevant correlations. In the oligopoly setting, correlation across firms' idiosyncratic shocks arises naturally. For instance, a market-wide change in consumer tastes or a new advertising regulation may simultaneously affect how all products are perceived by consumers, thereby inducing correlation between the demand shocks $V_{ig}$ and $V_{-ig}$.
I model the treatment selection mechanism using a single threshold crossing rule: unit $i$'s chooses to take the treatment, $D_{ig} = 1$, if the unobserved factor $V_{ig} \in \mathbb{R}$ does not exceed a threshold $h_i(Z_{ig}, Z_{-ig})$, where $ h_i: \mathbb{R}^{k_i} \times \mathbb{R}^{k_{-i}} \mapsto \mathbb{R} $ is an unspecified function. I do not impose a parametric form on the threshold function $h_i$, and the subscript $i$ emphasizes that its functional form may differ across units $i \in \{0,1\}$ within the same group. In contrast to complete information games, where unit $i$'s treatment $D_{ig}$ directly depends on the treatment decision of her peer $D_{-ig}$, my framework is consistent with a simultaneous-move game with incomplete information, as studied by aradillas2010semiparametric and related papers. In this setting, $D_{ig}$ corresponds to player $i$'s action, and $(Z_{ig}, Z_{-ig})$ represent publicly observed signals that serve as instrumental variables. Each unit observes $(Z_{ig}, Z_{-ig})$ and forms beliefs about the joint treatment choices within the group, specifically the probability $\mathbb{P}(D_{ig} = 1, D_{-ig} = 1 \mid Z_{ig}, Z_{-ig})$, and then chooses her treatment based on these beliefs. Thus, treatment decisions are interdependent through expectations rather than through observing others' realized treatment choices. As shown in aradillas2010semiparametric, the optimal decision rule in such simultaneous-move incomplete information game is consistent with a single-index threshold-crossing structure of the form in Equation (ref). Appendix (ref) provides a detailed discussion.
For identification, which I discuss in detail later, the instrumental variables $(Z_{ig}, Z_{-ig})$ must be independent of the unobserved heterogeneity $(U_{ig}, U_{-ig}, V_{ig}, V_{-ig})$ in the group and must not directly affect the outcomes $(Y_{ig}, Y_{-ig})$. In studies that focus on complete information setting, such as those analyzed by balat2023multiple and hoshino2023treatment, strategic interactions between $D_{ig}$ and $D_{-ig}$ are modeled explicitly, but unit $i$'s treatment is not allowed to depend on her group member's instrument $Z_{-ig}$. A distinguishing feature of this framework is that it does not rely on any additional exclusion restrictions on instruments: I allow the treatment $D_{ig}$ to depend on both the unit's own instrument $Z_{ig}$ and the peer's instrument $Z_{-ig}$, thereby accommodating potential spillovers from instruments into treatment decisions. Moreover, the instrumental variables may take the form of common public signals observed by all group members, so that the same variable $Z$ serves as the instrument for each member, or unit-specific instruments that vary across members, $Z_{ig} \neq Z_{-ig}$. This framework accommodates both shared and individual sources of exogenous variation.
Figure (ref) presents a directed acyclic graph (DAG) that illustrates the causal relationships among the key variables within group $g$. The red arrows represent spillover channels: unit $i$'s outcome $Y_{ig}$ may depend on her peer's treatment $D_{-ig}$, and her treatment $D_{ig}$ may depend on her peer's instrument $Z_{-ig}$. Direct interaction between treatments $D_{ig}$ and $D_{-ig}$, however, is ruled out. The unobserved heterogeneity $V_{ig}$ and $V_{-ig}$ introduce endogeneity, as they may simultaneously affect both treatments and outcomes. Those are represented by the black dashed arrows. The blue dashed arrow reflects potential dependence between $V_{ig}$ and $V_{-ig}$, for which I do not impose any functional restrictions.
Assumptions (ref)-(ref) set out the maintained restrictions on the key variables that are imposed throughout the paper.
Assumption (ref) requires that instruments are randomly assigned at the group level, implying that the group-level instrument vector $(Z_{ig}, Z_{-ig})$ is independent of the unobserved heterogeneity of all units in the group. This assumption places no restrictions on the dependence structure between $Z_{ig}$ and $Z_{-ig}$ within a group. The instruments may be arbitrarily correlated across units within a group, as long as they remain jointly independent of the unobserved heterogeneity $(V_{ig}, V_{-ig}, U_{ig}, U_{-ig})$.
Assumption (ref) requires that the instruments affect the outcome only through their influence on treatment take-up, without exerting any direct effect on the outcome. This condition corresponds to the standard exclusion restriction commonly imposed in instrumental variable analyses.
Assumption (ref) requires that the unobserved heterogeneity $V_{ig}$ has a continuous distribution, which is a common condition in the literature. Under this assumption, $V_{ig}$ can be normalized to follow a uniform distribution on the interval $(0,1)$.
Another implicit restriction embedded in the treatment selection equation is a monotonicity structure. Specifically, consider the case in which instruments $Z_{ig}$ and $Z_{-ig}$ are binary, taking values in $\{z_0, z_1\}$. If the threshold function satisfies the following ordering condition:
for all group members $i$ and $-i$, then the treatment selection equation implies a corresponding monotonicity property for treatment take-up, consistent with the condition studied in the literature (e.g., vazquez2023causal):
for each $i$ and $-i$ within group $g$.
The proposed framework applies to a broad class of empirical settings where spillovers operate through both outcomes and endogenous treatment decisions. Illustrative examples include oligopoly markets, where firms' pricing decisions may influence competitors' market shares, and education contexts, where an individual's labor market outcomes depend on the best friend's schooling decision. More broadly, the framework can be extended to settings such as households, where behaviors involving risky activities generate spillover effects on the health outcomes of other members.
A central insight from the treatment effect literature is that, when treatment assignment is endogenous, causal effects can often be identified for specific subpopulations defined by the instrument, for example, the local average treatment effect (LATE) in imbens1994identification and related studies. However, in the presence of spillovers, individuals' outcomes depend not only on their own treatment but also on the treatments received by others in their group. The spillovers complicate the interpretation of conventional LATE parameters, as variation in peers' treatments introduces additional causal channels. To disentangle these channels, this section extends the LATE framework to define local average effects that separately capture the causal effect of peers' treatments on an individual's outcome and the direct effect of the individual's own treatment. These parameters retain the causal interpretability of LATE while accommodating the presence of endogenous treatment and spillovers within groups. The identification of these local average effects further motivates the development of a framework based on marginal treatment effects, which explicitly accounts for spillovers operating through both treatment selection and outcomes, as formalized in Section (ref). Since the identification analysis is conducted at the level of a super-population of groups, I suppress the group subscript $g$ throughout this section to simplify notation.
Building on the framework introduced in Section (ref), the expected potential outcome $Y_i(d, d')$ generally depends on both the individual's and her peers' unobserved characteristics, $(V_i, V_{-i})$. Following the terminology in the literature, I refer to the conditional expectations $\mathbb{E}[Y_i(d, d') \mid (V_i, V_{-i}) \in P]$, where $P$ denotes a subset of the support of $(V_i, V_{-i})$, as local average potential outcomes. These parameters capture the average potential outcomes for subpopulations defined by specific values of the group-level unobservables, which reflect the underlying unobserved heterogeneity in the population. Taking appropriate differences between local average potential outcomes under different treatment combinations $(d, d')$ yields the average spillover effects from peers' treatments and the direct effects from a unit's own treatment. Definition (ref) provides formal definitions of these parameters.
The generalized LACSEs capture counterfactual spillover effects by exogenously fixing a unit's own treatment, whereas the generalized LACDEs capture counterfactual direct effects by exogenously fixing peers' treatments. Both effects are defined conditional on a subpopulation characterized by $(V_i, V_{-i}) \in P$, in the same spirit as the LATE widely studied in the literature. These parameters possess clear causal interpretations, as they disentangle the distinct influence channels of a unit's own treatment and peers' treatments, while allowing for unobserved treatment effect heterogeneity through conditioning on group-level unobservables. I now establish the identification of the generalized LACSEs and LACDEs using instrumental variables under Assumptions (ref)-(ref). It is worth noting that this identification result accommodates both discrete instruments with multiple support points and continuously distributed instruments.
I define the propensity score function for unit $i$ as the probability of treatment conditional on the group-level instrument vector, $P_{i}(Z_i, Z_{-i}) \equiv \mathbb{P}(D_i = 1 \mid Z_i, Z_{-i}) $. I denote this function simply by $P_i$ and define the support of the propensity scores for all group members as $\mathcal{P} \equiv \text{Supp}(P_i, P_{-i})$. The propensity score function $P_i$ identifies the threshold function $h_i$ in the treatment selection equation for each group member $i$, as demonstrated in the following derivation:
Additionally, the propensity scores $(P_i, P_{-i})$ are independent of all group members' unobserved heterogeneity, since they are functions of the instruments $(Z_i, Z_{-i})$.
Building on the identification of the propensity score, Theorem (ref) establishes the identification of the generalized LACSEs and LACDEs.
The identification argument proceeds as follows, with the formal proof provided in Appendix (ref). Given a pair of observed propensity scores $(P_i, P_{-i}) = (p_0, p_1)$, and noting that the propensity score $P_i$ identifies the threshold function $h_i$, the joint treatment realizations $(D_i, D_{-i})$ partition the space of unobserved heterogeneity $(V_i, V_{-i})$ into four mutually exclusive subpopulations, separated by the thresholds $(p_0, p_1)$. The relationships are summarized as
For instance, the probability of observing $\{D_i = 1, D_{-i} = 1\}$ conditional on $(P_i, P_{-i}) = (p_0, p_1)$ identifies the share of the subpopulation with $\{V_i \leq p_0, V_{-i} \leq p_1\}$, that is,
The top left panel of Figure (ref) illustrates how these four subpopulations correspond to distinct realizations of $(D_i, D_{-i})$ given observed propensity scores $(p_0, p_1)$.
Given the data, the conditional expectation $\mathbb{E}[Y_i \mathbbm{1}\{D_i = d, D_{-i} = d'\} \mid P_i = p_0, P_{-i} = p_1]$ can be directly recovered from observables. Under the model framework and Assumptions (ref)-(ref), these observed moments identify the average potential outcomes $Y_i(d, d')$ for the subpopulations associated with the treatment realization $\{D_i = d, D_{-i} = d'\}$. These quantities form the foundation for the identification strategy of Theorem (ref). For instance, when $(D_i, D_{-i}) = (1,1)$,
which identifies the average potential outcome $Y_i(1,1)$ for the subpopulation with unobserved characteristics satisfying $\{V_i \leq p_0, V_{-i} \leq p_1\}$.
The identification of the generalized LACSEs and LACDEs exploits exogenous variation in the propensity score values. Suppose there exists another pair of observed propensity scores $(p_0, p_1')$, $p_1' > p_1$. By shifting the propensity scores from $(p_0, p_1)$ to $(p_0, p_1')$ and applying the relationships established in Equation (ref), the subpopulation with unobserved characteristics in the region $\{V_i \leq p_0, p_1 < V_{-i} \leq p_1'\}$ changes its treatment status from $(D_i, D_{-i}) = (1, 0)$ to $(1, 1)$. Likewise, the subpopulation in $\{V_i > p_0, p_1 < V_{-i} \leq p_1'\}$ changes from $(D_i, D_{-i}) = (0, 0)$ to $(0, 1)$. These two groups correspond to the blue-shaded areas in the top-right panel of Figure (ref).
In both cases, only the peer $-i$ changes her treatment status $D_{-i}$, providing variation that identifies the average spillover effect. Taking the difference between the conditional expectations, $\mathbb{E}[Y_i \mathbbm{1}\{D_i = d, D_{-i} = 1\} \mid \cdot, \cdot] - \mathbb{E}[Y_i \mathbbm{1}\{D_i = d, D_{-i} = 0\} \mid \cdot, \cdot]$, evaluated at $(p_0, p_1)$ and $(p_0, p_1')$, isolates the variation in outcomes attributable to the subpopulations that experience a change in peer treatment status. This difference identifies the local average controlled spillover effects, $\operatorname{LACSE}_i^{(1)}$ and $\operatorname{LACSE}_i^{(0)}$, for the subpopulations corresponding to the two blue-shaded regions in the top-right panel of Figure (ref). This variation yields the identification result stated in Item 1 of Theorem (ref).
Analogously, when another pair of propensity scores $(p_0', p_1)$ with $p_0' > p_0$ is observed, shifting from $(p_0, p_1)$ to $(p_0', p_1)$ and using the relationships in Equation (ref) induces changes in treatment status for unit $i$ only. Specifically, the subpopulations defined by $\{p_0 < V_i \leq p_0', V_{-i} \leq p_1\}$ and $\{p_0 < V_i \leq p_0', V_{-i} > p_1\}$ change their treatment realizations from $\{D_i = 0, D_{-i} = 1\}$ to $\{D_i = 1, D_{-i} = 1\}$ and from $\{D_i = 0, D_{-i} = 0\}$ to $\{D_i = 1, D_{-i} = 0\}$, respectively, as illustrated by the two yellow-shaded regions in the bottom-left panel of Figure (ref). Because only unit $i$ changes treatment status, the resulting variation identifies the local average controlled direct effect. This variation corresponds to the identification result presented in Item 2 of Theorem (ref).
Suppose the support of the propensity scores exhibits sufficient variation such that four distinct pairs, $(p_0, p_1)$, $(p_0, p_1')$, $(p_0', p_1)$, and $(p_0', p_1')$, are observed with $p_0' > p_0$ and $p_1' > p_1$. Applying the same logic as before, shifting the peer's propensity score from $p_1$ to $p_1'$ while fixing unit $i$'s score at $p_0'$ and, conversely, shifting unit $i$'s score from $p_0$ to $p_0'$ while fixing the peer's score at $p_1'$, identifies the corresponding LACSEs and LACDEs for subpopulations defined by these regions of $(V_i, V_{-i})$. Next, taking cross-differences of the local average effects across the four propensity-score pairs, $(p_0, p_1)$, $(p_0, p_1')$, $(p_0', p_1)$, and $(p_0', p_1')$, isolates the LACSEs and LACDEs for the subpopulation with $(V_i, V_{-i})$ lying in the rectangle $\{p_0 < V_i \leq p_0', p_1 < V_{-i} \leq p_1'\}$, illustrated by the green-shaded area in the bottom-right panel of Figure (ref).
For example, the difference between $\operatorname{LACSE}_i^{(1)}$ identified for the regions $\{V_i \leq p_0, p_1 < V_{-i} \leq p_1'\}$ and $\{V_i \leq p_0', p_1 < V_{-i} \leq p_1'\}$ yields $\operatorname{LACSE}_i^{(1)}$ for the subpopulation $\{p_0 < V_i \leq p_0', p_1 < V_{-i} \leq p_1'\}$. Similarly, analogous differences yield $\operatorname{LACSE}_i^{(0)}$ and $\operatorname{LACDE}_i^{(d)}$, $d \in \{0, 1\}$, for the same region. This result is formally presented in Item 3 of Theorem (ref), which requires the support of $(P_i, P_{-i})$ to contain four distinct points forming the “vertices” of a rectangle in the $(p_0, p_1)$ space.
Theorem (ref) demonstrates that local average effects can be identified when the instrumental variables exhibit sufficient variation to induce the necessary differences in propensity scores. As discussed in Remark (ref), when instruments take only binary values, additional restrictions are required to achieve point identification of certain local average spillover or direct effects. Together, these results highlight that adequate variation in the instruments is crucial for identifying causally interpretable parameters in settings where spillovers influence both outcomes and treatment selection.
When the instrumental variables exhibit continuous variation, they induce continuous variation in the propensity scores. In this case, one can extend the identification strategy in Theorem (ref) by taking limits as $p_0' \to p_0$ and $p_1' \to p_1$, thereby identifying the average controlled spillover and direct effects conditional on $(V_i, V_{-i})$ evaluated at a specific point $(p_0, p_1)$ within the interior of the propensity score support. The next section formalizes this idea by introducing the marginal controlled spillover and marginal controlled direct effects. These parameters serve as building blocks for identifying not only the local average controlled spillover and direct effects discussed above, but also a broader class of policy-relevant treatment effects of interest to researchers.
Definition (ref) formally defines the causal spillover and direct effects evaluated at specific values of the unobserved characteristics $(V_i, V_{-i})$.
The marginal controlled spillover effect captures the impact of changing the peer's treatment on a unit's potential outcome while holding the unit's own treatment status fixed, conditional on the unobserved characteristics $(V_i, V_{-i})$ within the group. Similarly, the marginal controlled direct effect measures the impact of changing a unit's own treatment on her potential outcome while holding the peer's treatment constant, again conditional on $(V_i, V_{-i})$. Because both effects are defined relative to the group-level unobserved heterogeneity, they capture treatment effect heterogeneity arising from latent factors within the group.
The marginal controlled spillover and direct effects are defined analogously to the standard marginal treatment effect (MTE), conditioning on continuous unobserved heterogeneity within the support of the latent variables. Unlike the conventional MTE framework, which rules out interference across units, the marginal controlled effects explicitly incorporate spillovers arising from peers' treatment decisions. As such, they extend the MTE concept to environments where spillovers exist in both outcomes and treatment selection. Section (ref) formally characterizes the connection and distinction between these effects and the standard MTE. By conditioning on the continuous unobservables, the marginal controlled effects provide the building blocks for a wide class of policy-relevant parameters. In particular, the generalized local average controlled effects introduced in Definition (ref) represent a specific class of policy-relevant parameters that can be obtained by integrating the marginal controlled effects over selected regions of the latent heterogeneity space. These connections will be discussed in detail in Section (ref). The policy-relevant effects play a central role in evaluating counterfactual policies in settings with endogenous treatment and spillovers.
As discussed in the setting, the unobserved heterogeneities $V_i$ and $V_{-i}$ within a group may be correlated. Identification of the parameters of interest requires recovering the joint density of $(V_i, V_{-i})$. Because the marginal distributions of $V_i$ and $V_{-i}$ are normalized to be uniform on the interval $(0,1)$, their joint distribution is characterized by their copula. Formally, the copula is defined as
Lemma (ref) provides identification of this copula on the support of the propensity scores $(P_i, P_{-i})$ without imposing any functional form assumptions, where $P_i$ denotes unit $i$'s propensity score as defined in Equation (ref).
Let $c_{V_i, V_{-i}}(\cdot, \cdot)$ denote the copula density of $(V_i, V_{-i})$. Since Lemma (ref) establishes identification of the copula, the copula density can be obtained provided that the conditional probability $\mathbb{P}(D_i = 1, D_{-i} = 1 \mid P_i, P_{-i})$ is twice differentiable. This differentiability condition requires that $P_i$ and $P_{-i}$ exhibit continuous variation, which in turn implies that at least some components of the instrument vector $(Z_i, Z_{-i})$ must be continuously distributed. Assumption (ref) introduces this continuity requirement.
It then follows that the copula density of $(V_i, V_{-i})$ is identified, as stated in Corollary (ref).
Following the literature, the conditional expectation of the potential outcome, given the values of the group-level unobserved characteristics $(V_i, V_{-i})$,
is defined as the marginal treatment response (MTR) function. The marginal controlled spillover effects (MCSEs) and marginal controlled direct effects (MCDEs) introduced in Definition (ref) are obtained as differences of the corresponding MTR functions. Hence, identification of the MCSEs and MCDEs requires identifying the underlying MTR functions. Theorem (ref) provides the identification of the parameters of interest, the MCSEs and MCDEs, while the detailed process for identifying MTR functions is presented in Appendix (ref).
The identification of the MCSEs and MCDEs builds on the same logic underlying the identification of the LACSEs and LACDEs in part 3 of Theorem (ref), by taking the limits $p_1' \to p_1$ and $p_0' \to p_0$. This argument is illustrated by the green-shaded region in the bottom-right panel of Figure (ref), which represents the limiting case where $p_1' \to p_1$ and $p_0' \to p_0$. The validity of this limiting argument requires the propensity scores to exhibit continuous variation in a neighborhood of $(p_0, p_1)$. Moreover, the identification framework naturally extends to settings with exogenous covariates, with the corresponding results provided in Appendix (ref).
The assumption of continuously distributed instruments in Assumption (ref) is sufficient but not necessary for identifying the MCSEs and MCDEs. When instruments have limited or discrete variation, the parametric specifications in Assumption (ref) can be used to extrapolate the identification of marginal controlled effects beyond the observed support of the propensity scores. Alternative extrapolation approaches, such as those proposed by mogstad2018using, may also be applied. However, point identification may no longer hold, and the parameters would instead be partially identified. A formal treatment of this extension is left for future research.
The MCSEs and MCDEs not only capture heterogeneous spillover and direct effects but also serve as fundamental building blocks for deriving a wide range of causal parameters commonly examined in the literature. This section illustrates several examples demonstrating how the MCSEs and MCDEs can be used to recover other treatment effect parameters of policy relevance.
Researchers are often interested in summarizing heterogeneous spillover and direct effects across individuals by aggregating them into population-level parameters (see, for example, vazquez2023identification and related studies). Within this framework, the average controlled spillover effect (ACSE) can be defined as $\operatorname{ACSE}_i(d) \equiv \mathbb{E}[Y_i(d, 1) - Y_i(d, 0)]$, which measures the expected change in unit $i$'s outcome when the peer's treatment status changes exogenously from 0 to 1, holding the unit's own treatment fixed at $D_i = d$. Similarly, the average controlled direct effect (ACDE) can be defined as $\operatorname{ACDE}_i(d) \equiv \mathbb{E}[Y_i(1, d) - Y_i(0, d)]$, which reflects the expected change in unit $i$'s outcome when her own treatment status changes exogenously from zero to one, holding her peers' treatment status fixed at $D_{-i} = d$.
Under the assumption that the propensity scores have full support, i.e., $\mathcal{P} = (0,1)^2$, the ACSEs and ACDEs are point identified using the MCSEs, the MCDEs, and the copula density of $(V_i, V_{-i})$ identified in Section (ref), by integrating the MCSEs or MCDEs weighted by the corresponding copula density:
When the propensity scores lack full support, the average controlled spillover and direct effects, as well as other policy-relevant treatment effects, remain only partially identifiable. Under the additional assumption that potential outcomes are almost surely bounded, $|Y_i(d, d')| \leq B < \infty$, the MCSEs and MCDEs are confined within the range $[-2B, 2B]$ at points outside the observed support of the propensity scores. To extend identification beyond this region, one may impose the parametric structure in Assumption (ref) or adopt an extrapolation approach similar to that proposed by mogstad2018using. A formal development of these extensions is left for future research.
Once the MCSEs and MCDEs are point identified, they can be used to recover the LACSEs and LACDEs defined in Section (ref). The following discussion illustrates how these marginal effects can be employed to obtain the local average spillover and direct effects examined in the existing literature.
Suppose there exist two values of the instrumental variable, $z_0, z_1$, such that the associated propensity scores $P_i(z, z')$, for $z, z' \in {z_0, z_1}$, can be consistently ordered across all individuals and groups. Without loss of generality, assume that $P_i(z_0, z_0) \leq P_i(z_0, z_1) \leq P_i(z_1, z_0) \leq P_i(z_1, z_1)$. Under this ordering, the treatment selection mechanism in Equation (ref) implies the following monotonicity condition:
almost surely, where $D_i(z, z')$ denotes the potential treatment received by unit $i$ when the instrument assignments are fixed exogenously at $(Z_i, Z_{-i}) = (z, z')$.
vazquez2023causal partitions the population into a finite number of compliance types based on the values of the potential treatment vector ${D_i(z, z')}_{z, z' \in \{z_0, z_1\}}$. This framework identifies the local average spillover effect $\mathbb{E}[Y_i(0,1) - Y_i(0,0) \mid T_{-i} = c]$, and the local average direct effect $\mathbb{E}[Y_i(1,0) - Y_i(0,0) \mid T_i = c]$, where $T_i = c$ denotes the complier subgroup, defined as the set of units whose unobserved heterogeneity $V_i$ lies between the two thresholds $P_i(z_0, z_1)$ and $P_i(z_1, z_0)$. Intuitively, these are individuals who would not take the treatment under $(z_0, z_1)$ but would take it under $(z_1, z_0)$. This paper also considers the setting in which a one-sided noncompliance condition holds, meaning that individuals cannot receive the treatment when assigned the instrument value $z_0$. Under one-sided noncompliance, the propensity scores satisfy $0 = P_i(z_0, z_0) = P_i(z_0, z_1) \leq P_i(z_1, z_0) \leq P_i(z_1, z_1)$, which corresponds to a special case of the sufficient identification condition stated in Part 1 of Theorem (ref).
Figure (ref) in Appendix (ref) illustrates the regions in the $(V_i, V_{-i})$ space corresponding to the subpopulations ${T_i = c}$ and ${T_{-i} = c}$. Integrating the identified MCSEs and MCDEs over these regions, using the copula density of $(V_i, V_{-i})$ as weights, recovers the local average spillover and direct effects analyzed by vazquez2023causal:
where the copula density is identified in Corollary (ref), and Theorem (ref) provides identification of $\operatorname{MCSE}_i\left(0; v_0, v_1\right)$ and $\operatorname{MCDE}_i\left(0; v_0, v_1\right)$.
In addition, the identified MCSEs and MCDEs can be used to recover the policy-relevant treatment effect (PRTE), which quantifies the expected change in outcomes induced by a policy-driven shift in the treatment selection mechanism. The PRTE aggregates the underlying marginal controlled effects across the distribution of unobserved heterogeneity, weighted by how the policy changes the propensity of treatment participation, following the interpretation of heckman2005structural and related work. This parameter provides a meaningful measure of the impact of counterfactual policy interventions in the presence of heterogeneous treatment effects, allowing for the evaluation of counterfactual interventions that modify the selection into treatment.
To formalize the link between marginal controlled effects and policy-relevant treatment effects, consider how the MCSEs and MCDEs characterize PRTEs arising from exogenous policy changes. Let $\mathcal{A}$ denote a set of feasible policies. For any policy $a \in \mathcal{A}$, I use superscript $a$ to denote the corresponding potential variables under that policy. For example, $D_i^a$ denotes the treatment status of unit $i$ that would be realized if policy $a$ were implemented. Thus, changing the policy from a to $a'$ induces changes in the distribution of instrumental variables, treatment selection, and outcomes, such as $Z_i^a \to Z_i^{a'}$, $D_i^a \to D_i^{a'}$, and $Y_i^a \to Y_i^{a'}$. These counterfactual changes form the basis for evaluating PRTEs.
Under policy $a$, the treatment decision is given by
where $P_i^a(Z_i^a, Z_{-i}^a) = \mathbb{P}(D_i^a = 1 \mid Z_i^a, Z_{-i}^a)$ denotes the policy-specific propensity score. Following the standard policy invariance assumption (as described in Assumption (ref)), the introduction of a new policy is assumed to affect only the treatment selection mechanism through changes in the propensity score, without altering the joint distribution of unobserved characteristics.
The policy-relevant treatment effect (PRTE) measures the average change in outcomes induced by a policy intervention that modifies the treatment assignment mechanism, which is defined as
where $\Delta P$ denotes the proportion of groups in which at least one member changes treatment status as a result of the policy shift from $a$ to $a'$.
Given that each pair of propensity scores $(P_i^a, P_{-i}^a)$ partitions the support of $(V_i, V_{-i})$ into four regions associated with distinct treatment realizations $(D_i, D_{-i})$, the expected outcome under policy $a$ can, under Assumption (ref), be expressed as follows:
Because $\mathbb{E}[Y_i^a]$ is represented as a weighted average of the marginal treatment response functions, and the PRTE is defined as the difference in $\mathbb{E}[Y_i^a]$ across alternative policies, the PRTE naturally admits an interpretation as a weighted average of the MCDEs and MCSEs over particular regions of the latent heterogeneity space $(V_i, V_{-i})$. Hence, the identified MCDEs and MCSEs provide the key building blocks for constructing PRTEs associated with a broad class of counterfactual policy interventions. Appendix (ref) presents explicit expressions for the PRTE under several empirically relevant types of policy changes.
This section introduces the connection between the marginal controlled effects and the standard marginal treatment effect (MTE) framework. The MCDEs and MCSEs are defined in a manner analogous to the MTE, capturing how potential outcomes vary with continuous unobserved heterogeneity. However, unlike the standard MTE that rules out spillovers, the MCDEs and MCSEs explicitly account for spillovers arising from peers' treatments as well as endogeneity in both own and peer treatment decisions. This section formally examines the relationship between the marginal controlled effects and the conventional MTE, demonstrating that the MCSEs and MCDEs naturally extend the MTE framework to settings with spillovers in outcomes and treatment selection within groups.
If spillover effects exist but are ignored and the standard MTE identification strategy is applied, the conventional estimand
fails to identify the true MTE, $\mathbb{E}[Y_i(1) - Y_i(0) \mid V_i = p_0]$.
In the presence of the spillover structure specified in Equation (ref), the conventional propensity score can be written as
where $\mathcal{Z}$ denotes the support of the peer's instrumental variable $Z_{-i}$, the second equality follows from the law of iterated expectations, and the last equality uses the independence assumption (Assumption (ref)) and the distributional normalization in Assumption (ref). This expression shows that the conventional propensity score is a weighted average of the unit's threshold function $h_i(z_0, z_1)$ over the peer's instrument $Z_{-i}$, conditional on $Z_i$.
Furthermore, Corollary (ref) demonstrates that, when spillovers are present, the conventional MTE identifier, $\partial \mathbb{E}[Y_i \mid P_i\left(Z_i\right)=p_0] / \partial p_0$ captures a weighted average of the MCDEs, augmented by a bias term arising from the dependence of the unit's treatment on the peer's instrument and from the correlation between $Z_i$ and $Z_{-i}$.
If spillover effects are absent from both the outcome and the treatment selection processes, the identification results collapse to the standard MTE framework, as shown in Corollary (ref).
To conclude, the standard MTE may lose its causal interpretation when spillovers are present, whereas the MCSEs and MCDEs coincide with the MTE under SUTVA. Hence, the framework developed in this paper provides a natural extension of the standard MTE framework, generalizing it to settings with spillovers in both outcomes and treatment selection.
The imposed spillover model structure and assumptions yield two sets of testable implications.
The first set arises from the fact that the cross-partial derivatives of the observed conditional expectations,
identify the joint distribution of potential outcomes weighted by the copula density of unobservables, where $A_1, A_2$ denote arbitrary Borel sets in the outcome support. Since both the conditional probabilities and the copula density are nonnegative, these derivatives must be weakly positive, generating a set of nesting inequalities that serve as testable restrictions.
The second set of implications follows from the index sufficiency property: the marginal treatment response functions depend only on the values of the propensity scores, not directly on the realizations of the instrumental variables. Hence, for any two instrument pairs $(z_0, z_1)$ and $(\tilde{z}_0, \tilde{z}_1)$ that yield identical propensity scores $(P_i, P_{-i})$, the conditional expectation
should remain invariant across the two sets of instruments, as they correspond to the same marginal treatment response values.
Corollary (ref) formally states the nesting inequality and index sufficiency conditions implied by the model.
Existing literature, such as carr2021testing, has developed methods that could be implemented to test the conditions in Proposition (ref). A formal application of these testing procedures to the present framework is left for future research.
Consider a random sample of $G$ groups, each consisting of n units. For each group $g = 1, \ldots, G$, the observed data
are independently and identically distributed across groups.
This section develops a semiparametric estimation procedure that extends the framework of carneiro2009estimating to settings with spillover effects in both treatment and outcome equations. The estimation section considers the covariate-augmented setting introduced in Appendix (ref). The model incorporating covariates can be expressed as
where outcomes depend on both own and peer covariates and treatments, and treatment decisions follow a threshold-crossing rule. It is assumed that covariates and instruments are randomly assigned at the group level,
where $W_{ig} \equiv (X_{ig}, Z_{ig})$.
Assumption (ref) is maintained throughout the estimation section.
To illustrate the estimation procedure, this section focuses on a simple case where each group consists of two units, i.e., $n = 2$ and $i \in \{0, 1\}$. The extension to group sizes $n > 2$ follows analogously. Building on the identification results with exogenous covariates presented in Appendix (ref), this section aims to estimate the marginal treatment response (MTR) functions,
where $\operatorname{sgn}(x)$ denotes the sign function that indicates the sign of a scalar $x$. The term ${2\operatorname{sgn}(1 - |d - d'|) - 1}$ evaluates to $1$ when $d = d'$ and to $-1$ when $d \neq d'$. The marginal controlled effects are then obtained by taking the difference between the estimated MTR functions.
The propensity score functions, $P_{0g} \equiv P_0(W_g)$ and $P_{1g} \equiv P_1(W_g)$, defined as $\mathbb{P}(D_{ig} = 1 \mid W_g)$ with $W_g \equiv (W_{0g}, W_{1g})$, are not directly observed and must be estimated from the data. The first stage of involves estimating the propensity score functions using a series regression approach. To alleviate the curse of dimensionality, a partially linear additive specification is employed:
The covariate vector $w$ includes both continuous and discrete components, denoted by $w = (w^{cts}, w^{disc})$, where $w^{cts} = (w^{cts}_1, \cdots, w^{cts}_{\ell_1})$ is an $\ell_1-$dimensional vector of continuous random variables, and $w^{disc} = (w^{disc}_1, \cdots, w^{disc}_{\ell_2})$ is an $\ell_2-$dimensional vector of discrete random variables.
To preserve flexibility, no parametric restrictions are imposed on the unknown smooth functions $\varphi_1, \ldots, \varphi_{\ell_1}$ associated with the continuous covariates, while the coefficients $\vartheta_1, \ldots, \vartheta_{\ell_2}$ on the discrete covariates remain to be estimated.
Series estimation relies on constructing a basis for smooth functions defined on $\mathbb{R}$, denoted as $\{p_k : k = 1, 2, \dots \}$, such that each continuous function $\varphi_{\ell}$, for $\ell = 1, \dots, \ell_1$, can be approximated arbitrarily well by a linear combination of these basis functions as $k \rightarrow \infty$. Commonly used basis functions include polynomial basis functions, splines, and wavelets. Given a positive integer $\kappa$, define the regressor vector
where the first $\kappa \times \ell_1$ components correspond to basis function expansions of the continuous covariates, and the remaining $\ell_2$ components include the discrete covariates in their original form.
The series estimator of the conditional probability $\mathbb{P}(D_{ig} = 1 \mid W_g)$, for $i \in \{0, 1\}$, is given by
where $\hat{\theta}_{\kappa}^i$ is obtained from the least squares optimization problem
and $\tilde{\kappa} = \kappa \ell_1 + \ell_2$ denotes the total dimension of the regressor vector $P_\kappa(W_g)$.
A finite-sample concern is that the estimated series approximation $\tilde{P}_i(W_g)$ may take values outside the admissible unit interval $[0,1]$. To address this, a standard trimming adjustment can be applied. The trimmed estimator is defined as
where $\delta > 0$ is a small constant chosen by the researcher. The resulting $\hat{P}_i(W_g)$ thus provides a feasible and bounded estimator of the propensity score $\mathbb{P}(D_{ig} = 1 \mid W_g)$, which is denoted compactly as $\hat{P}_{ig}$ in subsequent analysis.
The next step involves estimating the cross-partial derivatives appearing in both the numerator and denominator of the estimand in Equation (ref). The procedure begins with estimating the denominator of the estimand, namely the cross-partial derivative
Local polynomial regression provides a flexible and well-established approach for estimating such derivatives of conditional expectations. Following fan1996local, the polynomial order is set to $p = d + 1$, where $d$ denotes the derivative order. Since the object of interest is a second-order derivative, a local cubic regression ($p = 3$) is employed:
where $K(\cdot)$ denotes the kernel function and $h_{G1}$ is the chosen bandwidth parameter. Bandwidths for kernel-based regressions can be selected using $K$-fold cross-validation. The estimated coefficient $\hat{b}_4(p_0, p_1)$ from Equation (ref) serves as an estimator of the cross-partial derivative $\partial^2 \mathbb{E}[D_{0g} D_{1g} \mid P_{0g}=p_0, P_{1g}=p_1] / \partial p_0 \partial p_1$.
The subsequent stage focuses on estimating the cross-partial derivative that appears in the numerator of the estimand,
Since the covariate vector $X_g$ may be multidimensional, the estimation adopts the semiparametric framework to mitigate the curse of dimensionality.
Under this specification, the conditional expectation, and consequently the marginal treatment response (MTR) function, can be expressed as a semiparametric function of $(\mathbf{x}, p_0, p_1)$, separating the parametric effect of covariates from the nonparametric dependence on the propensity scores:
The conditional expectation $\mathbb{E}\left[Y_{ig} \mid D_{0g} = d, D_{1g} = d', P_{0g}, P_{1g}\right]$ can be expressed as
Therefore, conditional on the subsample with $\{D_{0g} = d, D_{1g} = d'\}$, the coefficient vector $\beta_{idd'}$ can be estimated by the least squares regression as
where $\hat{E}_h[\cdot \mid \cdot]$ represents a kernel regression estimator with selected bandwidth $h$.
The residual then follows as
which serves as an estimator of the unobserved component $U_{i d d' g}$.
In the last step, use the sample
to estimate the cross-partial derivative $\partial^2 \mathbb{E}[U_{idd'g} \mid P_{0g}=p_0, P_{1g}=p_1] / \partial p_0 \partial p_1 $ through a local polynomial regression of order three,
The resulting coefficien $\hat{c}_4(d, d'; p_0, p_1)$ consistently estimates $\partial^2 \mathbb{E}[U_{idd'g} \mid P_{0g}=p_0, P_{1g}=p_1] / \partial p_0 \partial p_1 $.
Finally, the marginal treatment response functions $m_{i g}^{(\mathbf{x}, d, d')}(p_0, p_1)$ are estimated as
where $\widehat{c}_4(d, d'; p_0, p_1)$ and $\widehat{b}_4(p_0, p_1)$ are the local polynomial estimators of the cross-partial derivatives of $\mathbb{E}[U_{idd'g} \mid P_{0g}, P_{1g}]$ and $\mathbb{E}[D_{0g} D_{1g} \mid P_{0g}, P_{1g}]$, respectively.
The next section derives the asymptotic distribution of the estimated marginal treatment response functions, abstracting from the role of covariates to focus on the sampling behavior of the nonparametric components,
Cubic spline basis functions $\{p_k: k = 1, 2, \cdots\}$ are employed to approximate the nonparametric components of the propensity score functions. The following assumptions, adapted from belloni2015some, provide the regularity conditions required to establish the uniform convergence rate of the series estimators for the propensity score functions.
As shown in Equation (ref), the proposed estimator is expressed as the ratio of two estimated cross-partial derivatives of conditional mean functions. This subsection derives the asymptotic properties these two cross-derivative estimators under the semiparametric estimation procedure described in Section (ref).
To derive the asymptotic properties of the cross-partial derivative estimators, impose the following assumption:
This assumption ensures sufficient smoothness of the underlying conditional mean functions and regularity of the kernel function, which together guarantee the validity of local polynomial approximations.
This section characterizes the asymptotic distribution of the marginal treatment response functions, abstracting from covariate effects. The estimator, defined in Equation (ref), is constructed as the ratio of two estimated cross-partial derivatives of conditional mean functions. To establish the asymptotic properties of this estimator, the following assumptions are imposed.
Finally, the asymptotic distributions of the MCSEs and MCDEs are derived. Their estimators are constructed using the estimated marginal treatment response functions:
Assuming that the differences $\hat{c}_4(d, d'; p_0, p_1) / \hat{b}_4(p_0,p_1) - c_4(d, d'; p_0, p_1) / b_4(p_0,p_1)$ are asymptotically independent across different values of $d, d' \in \{0,1\}$, the asymptotic distributions of $\widehat{\text{MCSE}}_i(\mathbf{x}, d; p_0, p_1)$ and $\widehat{\text{MCDE}}_i(\mathbf{x}, d; p_0, p_1)$ follow in Theorem (ref).
The estimators introduced in Section (ref) converge at nonparametric rates, requiring sufficiently large sample sizes to yield reliable estimates. Moreover, when the group size $n>2$, the conditional expectations in the estimand involve higher-dimensional conditioning, further slowing the rate of convergence. Consequently, in settings with limited sample sizes or large groups, it may be preferable to impose parametric assumptions to estimate parameters of interest. This section develops a parametric approach for estimation and inference procedure.
The estimation continues to rely on Assumption (ref), while introducing the following additional parametric assumptions.
The imposed parametric assumptions are standard in the marginal treatment effect (MTE) literature and provide a tractable yet flexible framework for estimation and inference. Modeling the selection rule as \(D_{ig}=\mathbbm{1}\{\widetilde V_{ig}\le h_i(W_g)\}\) with a parametric index $h_i(W_g)$ and a standard normal unobserved term corresponds to the probit-type latent index widely used in practical implementations of the MTE framework (see, e.g., carneiro2011estimating; kline2016evaluating). The polynomial specification of $h_i(\cdot)$ provides sufficient flexibility to capture nonlinear relationships between instruments and covariates.
The Gaussian copula structure for $(V_{0g},V_{1g})$ is also a common parametric choice that facilitate likelihood-based estimation and allow dependence in unobserved heterogeneity across group members. Such assumptions have been adopted in the interference literature, including hoshino2023treatment, to capture correlated unobservables within the group.
The parametric specification of the MTR function is consistent with the functional-form assumptions commonly employed in the MTE literature to achieve point identification when the available instruments provide limited variation. Under SUTVA, brinch2017beyond show that imposing a parametric structure on the MTR function allows for the identification of heterogeneous treatment effects even with discrete instruments. Analogously, in the presence of spillovers, a similar approach can be applied by specifying the MTR function $\mathbb{E}[U_{ig}(d, d') \mid V_{ig} = v_0, V_{-ig} = v_1]$ as a polynomial expansion in the unobserved heterogeneities, $\Phi^{-1}(v_0)$ and $\Phi^{-1}(v_1)$. This formulation accommodates spillover effects from peers' treatments $d'$ and captures potential dependence between group members through involving $\Phi^{-1}(v_1)$. Moreover, the parametric formulation enables extrapolation beyond the observed support of the propensity scores, thereby allowing for the identification of policy-relevant treatment effects (PRTEs) even when instrumental variables exhibit limited or discrete variation brinch2017beyond.
The objective is to estimate the marginal treatment response function, $m_{i g}^{(\mathbf{x}, d, d')}(v_0, v_1) = \mathbb{E}[Y_{ig}(\mathbf{x}, d, d') \mid V_{0g} = v_0, V_{1g} = v_1]$, for any $\mathbf{x} \in \mathcal{X}$, $d, d' \in \{0, 1\}$, and $(v_0, v_1) \in (0, 1)^2$.
As in the semiparametric case, the first step involves estimating the propensity score functions $P_{0}(W_g), P_{1}(W_g)$, where $W_{i g}=(Z_{i g}, X_{i g}), W_g=(W_{0 g}, W_{1 g}) \in \mathbb{R}^\ell$. Under the first specification in Assumption (ref), and assuming that the instruments and covariates are independent of the unobserved heterogeneity $V_{ig}$, we can express the propensity score function as
The polynomial coefficients can be estimated using standard maximum likelihood methods:
Once the polynomial coefficients $\hat{\theta}_i$ are estimated, they can be substituted into $P_{ig}$ to obtain the estimated propensity score as
The next step is to estimate the joint dependence structure of $V_{0g}$ and $V_{1g}$. Under the second specification in Assumption (ref), this dependence is modeled by a Gaussian copula with correlation parameter $\rho$. Consequently, the second step of our procedure focuses on estimating $\rho$. The identification results imply the following equations,
Therefore, $\rho$ can be estimated using the maximum likelihood, substituting the first-stage estimates $\widehat{P}_{0g}$ and $\widehat{P}_{1g}$ for the true propensity scores $P_{0g}$ and $P_{1g}$,
The final step involves estimating the marginal treatment response $\mathbb{E}[Y_{ig}(\mathbf{x}, d, d') \mid V_{0g} = v_0, V_{1g} = v_1]$. Under the third specification in Assumption (ref), this function admits the following parametric representation,
Hence, the last stage of our procedure focuses on estimating the coefficient vectors $\beta_{i d d'}$ and $c_{i d d'}$. For illustration, consider the case $d = 1$ and $d' = 1$.
By combining the identification results with the third specification in Assumption (ref), the following relationship is obtained:
In the last line, $I_{11}^1(p_0, p_1, \rho), I_{11}^2(p_0, p_1, \rho)$, and $I_{11}^3(p_0, p_1, \rho)$ denote integrals that depend on $p_0, p_1$, and the correlation parameter $\rho$ of the Gaussian copula density $c_{V_{0g}, V_{1g}}(\cdot, \cdot)$ when $d = 1$ and $d' = 1$. Given that the propensity scores and the correlation have been estimated in the previous two steps, $\widehat{P}_{0g}$, $\widehat{P}_{1g}$, and $\hat{\rho}$ are substituted for their true values. The coefficient vectors $\alpha_{i11}$ and $\beta_{i11}$ are then estimated using the following least squares regression,
where $\widetilde{X}_{g} \equiv X_g \cdot \mathbb{P}(D_{0g}D_{1g} \mid \widehat{P}_{0g}, \widehat{P}_{1g})$. A similar procedure can be applied to estimate the coefficient vectors $\beta_{i d d'}$ and $\alpha_{i d d'}$ for other treatment combinations $(d, d')$. The estimated marginal treatment response function, $\widehat{m}_{i g}^{(\mathbf{x}, d, d')}(v_0, v_1)$, is obtained by substituting $(\hat{\alpha}_{i 11}', \hat{\beta}_{i 11}')'$ for $(\alpha_{i 11}', \beta_{i 11}')'$.
Finally, the MCDEs and MCSEs are obtained by taking differences of the estimated marginal treatment response functions,
This section introduces a set of assumptions under which the parametric estimators are consistent.
Under Assumptions (ref) and (ref), the estimator $\hat{\theta}_i$ obtained in the first stage is consistent.
To establish the consistency of the second-stage estimator of the Gaussian copula correlation parameter $\rho$, the following additional assumption is imposed.
These conditions allow us to establish the consistency of the second-stage estimator of $\rho$.
Finally, the consistency of the estimated coefficient vector $(\hat{\alpha}_{i 11}', \hat{\beta}_{i 11}')'$ is established, which in turn ensures consistency of the estimated marginal controlled effects.
Inference regarding the estimator $(\hat{\alpha}_{i dd'}', \hat{\beta}_{i dd'}')'$ is performed using standard nonparametric bootstrap methodologies. Based on Assumptions (ref) and (ref), the conditions in Theorem (ref), and standard regularity conditions, the bootstrap distribution converges uniformly to the sampling distribution of the estimator\footnote{The regularity conditions include stochastic equicontinuity and a quadratic remainder condition. A formal proof is left for future work.} romano2012uniform. Therefore, standard nonparametric bootstrap methods, such as resampling the data and recomputing all stages, are expected to yield valid inference.
After establishing the consistency and the asymptotic distibutions of $(\hat{\alpha}_{i dd'}', \hat{\beta}_{i dd'}')'$, the consistency and asymptotic distributions of $\widehat{\text{MCSE}}_i(\mathbf{x}, d; p_0, p_1)$ and $\widehat{\text{MCDE}}_i(\mathbf{x}, d; p_0, p_1)$ follow directly from the continuous mapping theorem. This result obtains by the continuity of $\widehat{\text{MCSE}}_i$ and $\widehat{\text{MCDE}}_i$ as functions of $(\hat{\alpha}_{i dd'}', \hat{\beta}_{i dd'}')'$.
This section presents a Monte Carlo simulation to assess the validity of the proposed parametric estimation methods.
For each Monte Carlo replication, I generate $G$ i.i.d. groups, where each group $g$ consists of two members indexed by $i \in \{0, 1\}$. I draw the group instrument vector, $Z_{g} = (Z_{0g}, Z_{1g})$, i.i.d. from a bivariate normal distribution $N(0, \Sigma_Z)$ with $\Sigma_Z = (1, 0.1; 0.1, 1)$. The correlation of $Z_{0g}$ and $Z_{1g}$ is not zero, since I allow the instruments of group members to be correlated. I also the group-level unobserved heterogeneity vector, $(\widetilde{V}_{0g}, \widetilde{V}_{1g})$, i.i.d. from a bivariate normal distribution $N(0, \Sigma_V)$ with $\Sigma_V = (1, 0.2; 0.2, 1)$ and independent of the instrument vector $Z_g$. By construction, $\widetilde{V}_{ig}$, $i \in \{0, 1\}$, follows a standard normal distribution. Additionally, the copula linking the normalized unobserved heterogeneity $V_{0g}$ and $V_{ig}$, where $\widetilde{V}_{ig} = \Phi(\widetilde{V}_i)$, is a Gaussian copula with correlation $\rho = 0.2$. These specifications are consistent with Assumption (ref) and the first two conditions in Assumption (ref).
I construct the following model to generate individual's treatment and potential outcome.
where the group-level disturbance $U_g \in \mathbb{R}$ is generated i.i.d. from the uniform distribution $\mathcal{U}(0, 1)$ and is independent of $(Z_{0g}, Z_{1g}, \tilde{V}_{0g}, \tilde{V}_{1g})$. The observed individual outcome $Y_{ig}$ is derived from
In our data generating process, the instrument vector $Z_g$ is independent of the unobserved heterogeneities and potential outcomes, $(\widetilde{V}_{ig}, U_g)_{i, d, d' \in \{0,1\}}$, satisfying Assumption (ref). Moreover, $Z_g$ does not directly affect the outcome $Y_{ig}(d, d')$, in accordance with Assumption (ref). The threshold function $h_i(\cdot)$ in the treatment assignment equation is specified as a first-order polynomial in the instrument, satisfying the first condition in Assumption (ref). For the potential outcomes, their conditional means given $V_{0g}$ and $V_{1g}$ satisfy the third condition in Assumption (ref). Therefore, the data generating process satisfies all identification and parametric assumptions.
I apply the method in Section (ref) to estimate and construct $95\%$ confidence intervals for the marginal controlled spillover (MCSE) and direct effects (MCDE) at selected evaluated points $(p_0, p_1)$. In the final step of computation, directly evaluating the integrals $I^j_{dd'}$, $j = 0, 1, 2, 3, 4$, at each estimated $(\widehat{P}_{0g}, \widehat{P}_{1g})$ is analytically intractable. To address this, I approximate the integrals using numerical integration. Specifically, I employ the Gauss-Hermite quadrature method, which I have verified to be both accurate and computationally efficient.
I arbitrarily select the following evaluation points,
for which the true MCSEs and MCDEs can be readily computed. I conduct $500$ Monte Carlo replications for each of four sample sizes, $G = 1000, 3000, 5000, 10000$. Table (ref) reports the coverage rates for the MCSEs, MCDEs, and the correlation parameter $\rho$.
For the MCSEs and MCDEs, when the sample size is $G = 1000$, the coverage rates are already close to, but slightly above, 95% for most parameters. As the sample size increases to $G = 3000$, the coverage rates decrease slightly yet remain close to 95%, with a few parameters falling just below this threshold. For larger sample sizes, the coverage rates for all parameters stabilize around 95%. The coverage rates for $\rho$ are also close to 95% across all sample sizes. These simulation results support the validity of our identification strategy and parametric estimation methods.
In the empirical analysis, I estimate the direct and spillover effects of returns to education among the best-friend groups. I use data from the National Longitudinal Study of Adolescent to Adult Health (Add Health)\footnote{This research uses data from Add Health, funded by grant P01 HD31921 (Harris) from the Eunice Kennedy Shriver National Institute of Child Health and Human Development (NICHD), with cooperative funding from 23 other federal agencies and foundations. Add Health is currently directed by Robert A. Hummer and funded by the National Institute on Aging cooperative agreements U01 AG071448 (Hummer) and U01AG071450 (Hummer and Aiello) at the University of North Carolina at Chapel Hill. Add Health was designed by J. Richard Udry, Peter S. Bearman, and Kathleen Mullan Harris at the University of North Carolina at Chapel Hill. No direct support was received from grant P01 HD31921 for this analysis.}, a nationally representative longitudinal survey that follows a cohort of U.S. adolescents from grades 7-12 (1994-95 school year) into adulthood. The dataset contains rich information on respondents' family background and detailed friendship networks during adolescence, as well as education attainment and income in adulthood. This unique combination of longitudinal social, demographic, and economic data makes Add Health well suited for studying the long-term effects of adolescent friendships.
The Add Health dataset collects detailed friendship information during adolescence in both the in-home and in-school components of the Wave I survey. In each component, respondents are asked to list up to five male and five female friends, ranked from best to fifth best. I construct best-friend groups, each consisting of two respondents, by matching individuals who mutually nominate each other as their best friend. Following card2013peer, I first identify best-friend pairs from the Wave I in-home interviews. I then match any remaining mutually nominated best-friend pairs from the Wave I in-school interviews. Because respondents can nominate the best friend of each gender, I prioritize opposite-gender pairs: if a respondent appears in two different best-friend groups, I retain the group consisting of opposite-gender best friends.
The relationship between an individual's own education and their income has been extensively studied in the economics literature. In contrast, relatively little attention has been paid to how a best friend's education attainment influences an individual's earnings. Such an effect may operate through two competing channels.
In this empirical study, I investigate the effect of a best friend's education attainment on an individual's earnings and assess which channel, information sharing or competition, plays the dominant role within best-friend networks. Importantly, our identification framework assumes that spillover effects occur only within the same network and do not extend across different networks. In the context of returns to education, this implies that any effect of another person's education is restricted to the identified best friend, with no cross-pair spillovers. I take the total personal yearly pre-tax income from the Wave III in-home survey and apply a natural logarithm transformation to construct the outcome variable $Y$. The binary treatment variable $D$ is set to 1 if the individual has completed at least 16 years of education and 0 otherwise. I include the age, gender, race, health status, and family income as the controlled covariates $X$. I assume that, conditional on the observed covariates, the coefficients $(\alpha'_{idd'}, \beta'_{idd'})'$ in the potential outcome equations are identical for individuals $i \in \{0, 1\}$ within the same group.
For the continuous instruments $Z$, I construct measures based on the average parental education level of the individual's non-best friends, defined as all listed friends who are not ranked as the best friend. The average parental education of non-best friends may influence an individual's education attainment through channels such as shaping aspirations, fostering self-confidence, or behavioral sharing cools2019girls. Furthermore, conditional on covariates capturing demographic and socioeconomic characteristics, the family background of non-best friends is plausibly independent of the individual's unobserved heterogeneity. This is because weaker social ties, such as those with non-best friends, are less likely to exhibit the strong peer spillovers characteristic of best-friend relationships, and any residual correlation in unobservables is unlikely to persist once observed similarities are controlled for. Therefore, Assumption (ref) is likely to hold in this context.
The average parental education level of non-best friends during adolescence is unlikely to have a direct effect on an individual's yearly income in adulthood. This is because weaker social ties, such as those with non-best friends, generally lack the sustained and intensive interactions needed to shape long-term labor market outcomes. Unlike best friends, non-best friends are less likely to share close personal networks, exchange detailed career information, or provide direct referrals in the labor market. Moreover, by adulthood, many of these weaker ties from adolescence are no longer active, further limiting the scope for any direct influence on earnings. Therefore, any effect of non-best friends' parental education on the individual's income is likely to operate indirectly through its influence on the individual's own education attainment, rather than through direct channels. Hence, Assumption (ref) is plausibly satisfied in this setting.
I also require that the education decisions of best friends do not directly affect one another. This assumption is plausible because, while best friends may share aspirations or study habits, the final decision on how many years of education to pursue is typically determined by individual specific factors, such as academic ability, that are not directly changed by the best friend's decision. Therefore, any influence between best friends' education outcomes is more likely to operate indirectly through shared environments or information exchange, which aligns with the simultaneous incomplete information framework underlying our setting, rather than through direct strategic interaction in determining each other's years of schooling.
After excluding best-friend pairs in which both members have missing values for the treatment $D$ or the instrument $Z$, the sample comprises 1,019 best-friend pairs. Given the limited sample size, the parametric framework outlined in Section (ref) is applied under the parametric conditions specified in Assumption (ref). The estimated correlation between best friends' unobservables $V_{0g}$ and $V_{1g}$ is 0.36, indicating a positive dependence structure among unobservables within best-friend networks.
Figure (ref) plots the point estimates and the 90% and 95% confidence intervals of the marginal controlled spillover and direct effects by conditioning on the peer's unobservable $V_{-i} = 0.5$ and varing the value of individual's own unobservable. The covariates are fixed at their sample means. The blue solid line depicts the point estimates of the MCSEs and MCDEs, while the light and dark gray shaded areas represent the 95% and 90% confidence intervals, respectively. The red dotted line corresponds to the estimated parametric standard MTEs, which deviate substantially from the estimated MCSEs and MCDEs and lie outside their confidence intervals in most cases. This divergence provides empirical evidence of spillover effects between best friends, indicating that the standard MTE framework fails to retain a causal interpretation in the presence of such spillovers.
The results reveal substantial heterogeneity across these parameters. In particular, the estimates of the MCDEs with $d = 1$, which capture the direct effect of completing at least 16-year education given the best friend has completed at least 16 years, are positive and statistically significant at the 5% level across most values of the individual unobservable $V_i$. However, the estimates of MCDEs with $d = 0$, which measure the direct effect of completing at least 16 years of education given the best friend has not completed this level, are not statistically significant, even at the 10% level, across all values of the individual unobservable $V_i$. This discrepancy may reflect complementarities in human capital accumulation within best-friend pairs, consistent with the first channel discussed earlier: a highly educated best friend can provide valuable labor market information and opportunities that enhance the returns to one's own education. When both friends attain higher education, they may reinforce each other's labor market prospects through stronger professional networks, mutual encouragement in career development, or joint access to high-return opportunities. In contrast, when the best friend has lower education attainment, such reinforcing mechanisms may be absent, weakening the direct effect of one's own education on earnings.
Figure (ref) also presents the estimated MCSEs along with their confidence intervals. The MCSEs with $d = 1$, which capture the spillover effect of the best friend completing at least 16 years of education given the individual has completed 16 years, are significantly positive at the 5% level for some values of the individual unobservable $V_i$. In contrast, the MCSEs with $d = 0$, which measure the spillover effect of the best friend completing at least 16 years of education given the individual has not completed 16 years, are even significantly negative at the 10% level when $V_i$ is around 0.5 (approximately the value of the peer's unobservable $V_{-i}$), suggesting potential adverse spillover effects for some individuals. These patterns are consistent with the two channels through which a best friend's education attainment may affect an individual's earnings. The findings suggest that the information and opportunity channel dominates the competition channel when the individual is also highly educated, leading to positive and significant spillover effects. Conversely, when the individual has not completed 16 years of education, the competition channel appears to dominate, particularly among pairs with similar values of unobserved heterogeneity, resulting in negative estimated spillover effects. This asymmetry suggests complementarities in human capital and opportunity sharing among equally educated peers, and the potential for relative disadvantage when education attainment differs within a best-friend pair.
The baseline framework can be generalized to accommodate various settings in which spillovers occur within predefined groups. First, point identification of the marginal controlled spillover and direct effects is established when outcomes depend on an exposure mapping function, rather than the full vector of group members' treatment statuses. Subsequently, Appendix (ref) extends the analysis to environments with continuous endogenous treatments, demonstrating that the marginal controlled spillover and direct effects remain point identified in such cases.
In many applications, the predetermined groups within which spillovers occur can be large or vary in size. For example, when groups are defined at the level of schools, villages, or communities. In such cases, modeling outcomes as a function of the entire vector of group members' treatments may become infeasible. To address this issue, I instead adopt a framework in which the outcome of unit $i$ in group $g$, denoted $Y_{ig}$, depends on the unit's own treatment $D_{ig}$ and on a known function of the full vector of group treatments, denoted $H_g$. This function $H_g$ summarizes the group's effective treatment, consistent with the notion of an effective treatment in manski2013identification and the exposure mapping framework of aronow2017estimating. By reducing the dimensionality of peer treatments to an interpretable exposure measure, this approach allows for the analysis of spillovers in large or heterogeneous groups while maintaining tractable identification and interpretation.
I consider a sample of $G$ independent and identically distributed groups, indexed by $g = 1, \cdots, G$, where spillovers are restricted to occur within groups and not across them. Unlike the baseline framework, each group now consists of $n_g$ members, where the group size $n_g$ is allowed to vary across groups. To capture peer effects in this heterogeneous group size setting, I assume that the outcome of interest depends not on the full treatment vector but rather on a group-level exposure mapping, $H_g: \boldsymbol{D}_g \mapsto \mathbb{R}$, where $\boldsymbol{D}_g$ denotes the vector of individual treatment assignments within group $g$. This mapping $H_g$ is assumed to be continuous and correctly specified by the researcher. A common and tractable specification is the proportion of treated individuals in the group, given by $H_g = \sum_{i=1}^{n_g} D_{ig} / n_g$. Because treatment assignments $\boldsymbol{D}_g$ are observed, researchers can directly recover the realized values of $H_g$ for each group.
I specify the outcome for individual $i$ in group $g$ as $Y_{ig} = Y_{ig}(D_{ig}, H_g, U_{ig}, U_{-i,g})$, so that outcomes may depend on the individual's own treatment $D_{ig}$, the continuous group-level exposure $H_g$, and both the individual's unobserved characteristics $U_{ig}$ and those of her peers $U_{-i,g}$. I let $Y_{ig}(d, h)$ represents the potential outcome for unit $i$ given $D_{ig} = d$ and $H_g = h$.
I formulate the following equations to model spillovers that operate through the group-level exposure $H_g$. Throughout this section, I retain the subscript $g$ to distinguish individual level variables (indexed by $ig$) from group level variables (indexed by $g$), thereby clarifying how exposure-driven spillovers enter the model.
In Equation (ref), I formulate the outcome equation within the classical potential outcomes framework, specifying that each individual's outcome depends on her own binary treatment status, $D_{ig} \in \{0, 1\}$, as well as the group-level exposure $H_g$. A key feature of our framework is the recognition that both the individual treatment $D_{ig}$ and the group exposure $H_g$ may be endogenous. Specifically, $D_{ig}$ may correlate with unobserved individual-level characteristics that also influence the outcome. Likewise, the group exposure $H_g$, which is defined as a function of all group members' treatments, may depend on group-level unobservables that also affect the individual outcome.
To address the endogeneity of the individual treatment $D_{ig}$, I model it using a single-threshold crossing rule, analogous to the specification in the basic setting. Specifically, individual $i$ selects into treatment if the unobserved characteristic $V_{ig}$ falls below a threshold $h_i(Z_{ig}, Z_{-ig})$. The threshold function depends on the vector of instruments assigned to individual $i$, $Z_{ig}$, or additionally on the instruments assigned to other group members, $Z_{-ig}$. As before, I do not impose any functional form restrictions on the threshold function $h_i$ to preserve flexibility in how instruments affect treatment selection. In addition, the subscript $i$ allows for heterogeneity in threshold functions across individuals within the same group.
To account for the potential endogeneity of the group-level exposure $H_g$, I introduce a group instrument $Z_g$ and assume that $H_g$ follows a reduced-form relationship given by $H_g = m(Z_g, \varepsilon_g)$, where $\varepsilon_g \in \mathbb{R}$ represents an unobserved group-specific characteristic. The group instrument $Z_g$ may take various forms. For instance, it may correspond to the full vector of individual instruments $(Z_{ig})_{i \in \{1, \cdots, n_g\}}$, or to an aggregate statistic such as the average instrument level within the group. The random variable $\varepsilon_g$ captures latent group-level heterogeneity, potentially containing factors such as the group's social cohesion or the dependence structure among individual-level unobservables $(V_{ig})_{i \in \{1, \cdots, n_g\}}$. I impose no functional form restrictions on $m(\cdot)$ to maintain flexibility in the modeling of group exposure. In addition, I do not restrict the dependence structure between the individual unobservable $V_{ig}$ and the group-level unobservable $\varepsilon_g$, allowing for arbitrary correlation between individual- and group-level latent factors.
In the exposure mapping framework, I redefine the marginal controlled spillover effect (MCSE) and the marginal controlled direct effect (MCDE) relative to the basic setting by explicitly conditioning on both the individual-specific unobservable $V_{ig}$ and the group-level unobservable $\varepsilon_g$, which accommodates heterogeneity at both the individual and group levels.
Identification in this setting is achieved under Assumptions (ref)-(ref).
Assumption (ref) requires that the vector of instruments assigned to individuals and the group must be randomly assigned at the group level, such that they are independent of all potential outcomes, as well as of both individual- and group-level unobservables. Moreover, under the model structure in Equation (ref), the instruments also satisfy the exclusion restriction, in the sense that they influence outcomes only through their effect on treatment take-ups and exposure, and do not directly enter the outcome equation.
The monotonicity condition in Assumption (ref) ensures that the group-level treatment $H_g$ is a one-to-one mapping of the group-level unobservable $\varepsilon_g$, conditional on the instruments. Specifically, for any given $Z_g = z$, the reduced-form relation $H_g \mid (Z_g = z) = m(z, \varepsilon_g)$ can be inverted with respect to $\varepsilon_g$, yielding $ \varepsilon_g \mid (Z_g = z) = m_z^{-1}(H_g) $. Thus, under the random assignment assumption (ref), I obtain the control function representation $\varepsilon_g = m_{Z_g}^{-1}(H_g)$. As established in goff2024testing, if the conditional distribution $F_{H_g \mid Z_g}$ is strictly increasing and continuous, then the monotonicity condition is not an additional structural restriction but instead follows directly from the reduced-form interpretation of the exposure mapping function discussed in Remark (ref).
I define the individual-level propensity score function for unit i in group g, consistent with previous definitions, as $P_{ig}(z) \equiv \mathbb{P}(D_{ig} = 1 \mid Z_{ig} = z)$. In addition, I define the group-level propensity score function for group $g$ as $P_g(z, h) \equiv \mathbb{P}(H_g \leq h \mid Z_g = z)$. I denote by $\mathcal{P}$ the support of the joint propensity score function $(P_{ig}, P_g)$. As shown in the Appendix (ref), the individual-level propensity score function identifies the individual threshold function $h_i(\cdot)$, while the group-level propensity score function identifies the inverse of the exposure mapping, $m_{Z_g}^{-1}(H_g)$. Taken together, these results imply that both individual- and group-level propensity score functions can be used as control functions, allowing us to account for unobserved heterogeneity in the outcome equation and thereby achieve identification of the marginal controlled spillover and direct effects.
This paper develops a general framework for identifying causal effects in environments with within-group spillovers and endogenous treatment decisions. By relaxing the Stable Unit Treatment Value Assumption (SUTVA), the framework accommodates settings in which an individual's outcome depends not only on her own treatment but also on the treatment selection of group members. It further allows each individual's treatment decision to depend on instruments assigned to other group members.
The paper introduces two classes of causal parameters, he generalized local average controlled spillover and direct effects (LACSEs and LACDEs) and the marginal controlled spillover and direct effects (MCSEs and MCDEs), which extend the standard local average and marginal treatment effect frameworks to settings with spillovers. The LACSEs and LACDEs quantify peer and own treatment effects for specific subpopulations, while the MCSEs and MCDEs capture these effects conditional on continuous values of unobserved heterogeneity within groups. The paper formally establishes general conditions for the point identification of the LACSEs and LACDEs, characterizing the instrumental variation necessary for identification regardless of whether the instruments are discrete or continuous. It also shows that the MCSEs and MCDEs are nonparametrically point identified from continuous instrumental variation without imposing functional form restrictions on the outcome equation or on the joint distribution of unobserved characteristics within the group, thereby accommodating flexible forms of spillover structures.
These results extend existing approaches to causal inference with spillovers, showing that the MCSE-MCDE framework provides a natural generalization of the standard marginal treatment effect (MTE) model to spillover settings. Furthermore, the paper establishes that these marginal controlled effects serve as building blocks for policy-relevant treatment parameters (PRTEs), enabling the evaluation of both direct and spillover effects under counterfactual policy interventions.
For estimation and inference, the paper develops a semiparametric estimation strategy that builds on carneiro2009estimating, extending it to accommodate within-group spillovers while mitigating the curse of dimensionality associated with covariates. Asymptotic properties of the semiparametric estimators are derived, and a parametric estimation framework is proposed as a practical complement when sample sizes are limited or group sizes are large. Monte Carlo simulations demonstrate that the parametric estimators perform well in finite samples, providing accurate estimates and confidence interval coverage.
An empirical application using the National Longitudinal Study of Adolescent to Adult Health (Add Health) illustrates the framework's practical relevance. The analysis examines how education attainment affects long-term earnings within best-friend networks. The results indicate positive dependence between friends' unobserved characteristics and reveal heterogeneous spillover patterns, showing that both the magnitude and direction of peer influences vary with individuals' and their best friends' education attainment.
Finally, the paper extends the framework to settings with exposure mappings, where outcomes depend on a known function of group members' treatments rather than the full treatment vector. This generalization broadens the applicability of the framework to environments with varying or large group sizes, while preserving nonparametric point identification under continuous instrumental variation. The appendix further extends the analysis to settings with continuous endogenous treatments and establishes point identification results under continuous instruments.
Overall, the proposed framework offers an econometric foundation for identifying and estimating causal effects in the presence of within-group spillovers and endogenous treatments. It provides theoretical and practical tools for studying a wide range of social, education, and economic interactions. Future research could extend the framework by modeling endogenous group formation, and by developing more efficient semiparametric estimation methods to enhance finite-sample performance.
\onehalfspacing