EconBase
← Back to paper

Unobserved Heterogeneous Spillover Effects in Instrumental Variable Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

183,678 characters · 31 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Unobserved Heterogeneous Spillover Effects in Instrumental Variable Models

\onehalfspacing

abstract\begin{footnotesize} This paper develops a general framework for identifying causal effects in settings with spillovers, where both outcomes and endogenous treatment decisions are influenced by peers within a known group. It introduces the generalized local average controlled spillover and direct effects (LACSEs and LACDEs), which extend the local average treatment effect framework to settings with spillovers and establish sufficient conditions for their point identification without restricting the cardinality of the support of instrumental variables. These conditions clarify the necessity of commonly imposed restrictions to achieve point identification with binary instruments in related studies. The paper then defines the marginal controlled spillover and direct effects (MCSEs and MCDEs), which naturally extend the marginal treatment effect framework to settings with spillovers and are nonparametrically point identified from continuous variation in instruments. These marginal effects serve as building blocks for a broad class of policy-relevant treatment effects, including some causal spillover parameters in the related literature. Semiparametric and parametric estimators are developed, and an application using Add Health data reveals heterogeneity in education spillovers within best-friend networks. \end{footnotesize}

{ Keywords: Unobserved heterogeneous spillover/direct effect; violation of SUTVA; causal inference; instrumental variable.

}

Introduction

In econometric analyses of treatment effects, the Stable Unit Treatment Value Assumption (SUTVA) is typically imposed, requiring that each individual's potential outcomes depend only on their own treatment assignment and not on the treatment assignments of others. This assumption, however, may fail in environments involving social interactions or group structures, where one unit's treatment can influence another's outcome. In such contexts, the SUTVA is unlikely to hold. In addition, treatment assignment may be endogenous, particularly in observational studies or randomized experiments with imperfect compliance, which further complicates the identification of causal parameters.

This paper develops a new framework for identifying causal effects in environments with within-group spillovers and endogenous treatments, such as education decisions among friends or pricing choices in oligopolistic markets. The framework explicitly accounts for two sources of SUTVA violations: individual outcomes may depend on peers' treatment selection, and individual treatment decisions may be influenced by instruments assigned to other group members.

To begin, the paper introduces the generalized local average controlled spillover and direct effects (LACSEs and LACDEs), which measure treatment effects from peers and from one's own treatment, respectively, for specific subpopulations. These parameters extend the local average treatment effect (LATE) framework imbens1994identification by allowing for spillovers in both outcomes and endogenous treatment decisions. This paper is the first to formally establish conditions under which the LACSEs and LACDEs can be point identified for specific subpopulations, and to derive general point identification results that do not rely on the cardinality of the support of instrumental variables. The results characterize the precise variation in instruments required to achieve point identification of local effects in environments with within-group spillovers and endogenous treatments.

When the instrumental variables exhibit continuous variation, the analysis introduces the marginal controlled spillover and direct effects (MCSEs and MCDEs), which measure the effects of peers' and individuals' own treatments conditional on specific values of unobserved characteristics within the group. These parameters naturally extend the marginal treatment effect (MTE) framework heckman2001policy,heckman2005structural to settings with spillovers. The paper first establishes the nonparametric point identification of the MCSEs and MCDEs from continuous variation in instruments, without imposing functional form restrictions on the outcome equation or the joint distribution of unobserved characteristics across group members, thereby accommodating flexible forms of spillovers. Similar to the standard MTE, the MCSEs and MCDEs serve as building blocks for identifying a broad class of policy-relevant treatment effects (PRTEs), including the LACSEs and LACDEs, as well as other PRTEs arising from counterfactual policy changes, facilitating policy evaluation in environments with spillovers.

Organization of the Paper

Section (ref) develops the framework with within-group spillovers and endogenous treatments, formally defining, identifying, and analyzing the causal parameters of interest. Section (ref) introduces an outcome model that allows each unit's potential outcome to depend flexibly on the entire vector of treatments within the group, capturing spillover effects from peers' treatment decisions on own outcomes. Treatment selection follows a single-index threshold-crossing structure, in which an individual receives treatment if her unobserved characteristic falls below a threshold function determined by her own and her peers' instrumental variables. This structure, which can be interpreted as the equilibrium behavior of a simultaneous incomplete-information game aradillas2010semiparametric, captures how peers' instruments can influence individual treatment decisions. Importantly, this framework imposes no parametric restrictions on the outcome equation or the threshold function and allows for arbitrary dependence among unobserved characteristics across group members, accommodating a broad range of environments in which spillovers may be present.

Section (ref) defines and identifies two causal parameters, the generalized local average controlled spillover and direct effects. The term “generalized local" indicates that these effects are defined for specific subpopulations of groups, while “spillover" and “direct" refer to the sources of treatment variation from peers and from the individual herself, respectively. The term “controlled" highlights that one treatment dimension, either own or peer, is held fixed when measuring the effect of the other. This section establishes general conditions under which the LACSEs and LACDEs are point identified, without requiring the instrumental variables to be either discrete or continuous, as long as they generate the required variation for identification. These identification conditions also clarify the rationale for additional restrictions, such as one-sided noncompliance, which are often imposed to achieve point identification with binary instruments (e.g., vazquez2023causal).

Section (ref) establishes that when instrumental variables exhibit continuous variation, the marginal spillover effect and marginal direct effect are nonparametrically point identified without requiring functional form assumptions on the outcome equation. In addition, the joint distribution of unobserved characteristics across group members is nonparametrically identified over the support of the observed treatment probabilities, without imposing parametric restrictions on how these unobserved factors are distributed. The MCSEs and MCDEs are defined analogously to the LACSEs and LACDEs but condition on a specific realization of unobserved characteristics within each group. By conditioning on the latent characteristics of all group members, these parameters flexibly capture heterogeneity in both direct and spillover effects that arise from variation in unobserved factors. Section (ref) further demonstrates that the marginal controlled effects form the basis for identifying a broad class of policy-relevant treatment parameters. By integrating the MCSEs and MCDEs over appropriate regions of the unobserved heterogeneity distribution, one can recover the LACSEs, LACDEs, and other treatment effect parameters associated with counterfactual policy interventions.

Section (ref) formally compares the MCSEs and MCDEs with the standard MTE and shows that, in the presence of spillovers, the conventional MTE may lose its causal interpretation, whereas in the absence of spillovers, the MCSEs and MCDEs coincide with the standard MTE. These results demonstrate that the MCSE-MCDE framework provides a natural generalization of the MTE framework to accommodate environments with spillovers. Section (ref) derives testable implications implied by the model structure and the identification assumptions.

Section (ref) develops a semiparametric estimation procedure for the MCSEs and MCDEs, extending the framework of carneiro2009estimating to accommodate within-group spillovers. The proposed approach mitigates the curse of dimensionality associated with covariates while maintaining the model's nonparametric flexibility, as the key structural components other than the covariate adjustment are left unrestricted. The section also establishes the asymptotic properties of the semiparametric estimators. Because these estimators converge at nonparametric rates, their finite sample precision may be limited in small samples or when groups include a large number of members. To address this concern, a complementary parametric framework is introduced, relying on intuitive assumptions that enable straightforward implementation and facilitate valid inference through nonparametric bootstrap methods.

Section (ref) presents both parametric simulation results and an empirical application. Section (ref) presents Monte Carlo simulation results for the parametric estimation procedure, demonstrating the strong finite-sample performance of the proposed parametric methods. Section (ref) implements the proposed framework empirically using the parametric procedure. The analysis examines how education attainment affects long-term earnings within best-friend groups, drawing on data from the National Longitudinal Study of Adolescent to Adult Health (Add Health). The results indicate positive dependence between friends' unobserved characteristics and reveal systematic heterogeneity in the marginal controlled direct and spillover effects. The estimated MCDEs of completing 16 years of education are significantly positive when the best friend has also attained this level of education across most values of the latent characteristics, but become statistically insignificant when the friend has not. Similarly, the estimated MCSEs are significantly positive for individuals who completed 16 years of education across most values of the latent characteristics, whereas for those who did not, the spillover effects are insignificant and even negative for certain ranges of unobserved heterogeneity. These results provide empirical evidence of heterogeneous spillover effects of education on long-term earnings within friendship networks, highlighting how their magnitude and direction depend on both individuals' and peers' education attainment.

The framework can be extended to accommodate additional settings. Section (ref) generalizes the analysis to cases where outcomes depend on an exposure mapping, which is a known function of group members' treatment statuses, rather than the full treatment vector. This extension is particularly relevant when group sizes vary or are large, making the full treatment representation impractical. The section formally defines the MCSEs and MCDEs under this extended setting and establishes their nonparametric point identification using continuous instrumental variables.

Related literature

Recent research has devoted increasing attention to the identification and estimation of treatment effects in the presence of spillovers. This paper contributes to several key strands within this growing body of work.

A common strategy for addressing interference has been to impose parametric structures on social interactions. For instance, manski1993identification discussed the linear-in-means model, formulated as a system of linear simultaneous equations to capture endogenous, exogenous, and correlated peer effects. Building on this result, subsequent work, such as bramoulle2009identification and blume2015linear, extended the framework to more complex forms of interaction within linear models and derived conditions under which social effects can be identified. However, they fundamentally rely on correct parametric assumptions regarding the structure of social interactions. Such assumptions may lead to model misspecification, particularly in the presence of nonlinear spillovers or heterogeneity across individuals. The framework developed in this paper departs from such reliance on parametric restrictions by studying identification under a nonparametric structure in both the outcome equation and the treatment selection mechanism. This design accommodates flexible and potentially complex spillover mechanisms and provides a robust framework for causal analysis in environments with within-group interactions.

In the main setting considered, an individual's outcome depends on the full vector of treatments within the group, consistent with the treatment response framework of manski2013identification. Within the context of randomized controlled trials (RCTs), hudgens2008toward and aronow2017estimating, along with related studies, formalized design-based frameworks for analyzing interference. The framework in this paper extends this line of research by allowing for noncompliance, so that individuals may not adhere to their assigned treatments. This feature is important in observational studies, where treatment status is not fully controlled by the researcher, and in experimental settings where imperfect compliance may occur. The analysis adopts a large-sample framework rather than a design-based approach to study causal identification under endogenous treatment selection.

vazquez2023causal employed a potential outcomes framework to analyze similar settings with spillovers operating through both outcomes and treatment selection, using a binary instrumental variable for identification. His approach classifies individuals into discrete compliance types according to how their treatment choices respond to changes in instruments and focuses on identifying local average spillover and direct effects for each specific type, which are closely related to the LACSEs and LACDEs introduced in this paper. He achieved point identification by excluding certain subpopulations under one-sided noncompliance, a restriction also used in related work such as ditraglia2023identifying. This paper establishes general identification conditions for the LACSEs and LACDEs, which clarify why one-sided noncompliance is required for point identification when the instrumental variable is binary. This paper further employs a continuously distributed instrumental variable to point identify the marginal controlled direct and spillover effects, defined conditional on continuous realizations of latent characteristics within groups. This approach connects the analysis to the marginal treatment effect literature and establishes the marginal effects as fundamental components for identifying a wide class of policy-relevant treatment effects. In particular, aggregating these marginal effects recovers the local average direct and spillover effects in vazquez2023causal, as well as other causal parameters under counterfactual policy interventions.

Recent studies, including balat2023multiple and hoshino2023treatment, use instrumental variable methods to identify spillover effects in settings with direct strategic interactions among agents, where each individual's treatment choice directly depends on the treatment decisions of other group members. Frameworks with direct strategic interactions assume that an individual's treatment decision does not directly depend on the instruments assigned to other group members, thereby ruling out spillovers from peers' instruments in the treatment selection process. In contrast, the framework developed here does not model explicit strategic interactions in treatment choices but allows each individual's treatment decision to depend on instruments assigned to other group members. This structure can be interpreted as the equilibrium outcome of a simultaneous incomplete-information game, following aradillas2010semiparametric, and thus provides a complementary perspective to models that incorporate direct strategic interaction.

The spillover framework developed in this paper and the multivalued treatment framework are not nested. When the group is treated as a single decision-making unit, the group treatment vector can be reformulated as a multivalued group-level treatment. This links the setting to multivalued MTEs such as lee2018identifying. When applied to spillover contexts, identification in lee2018identifying relies on an exclusion restriction that an individual's treatment does not depend on peers' instruments, whereas the framework developed here allows and models such spillovers from peers' instruments.

Model

Setting

I consider a sample of $G$ independent and identically distributed (i.i.d.) groups, indexed by $g = \{1, \cdots, G\}$. Each group consists of the same number of units, denoted by $n \geq 2$. For example, a group may correspond to a market with several competing firms or to a household with multiple members. Within each group, units are indexed by $i = \{0, \cdots, n-1\}$. Throughout, I assume that spillover effects operate only within groups and do not extend across groups.

Researchers are often interested in how a treatment affects an outcome. Let $Y_{ig}$ denote the outcome of interest for unit $i$ in group $g$, and let $\mathcal{Y}$ denote its support. In some settings, the outcome $Y_{ig}$ may depend not only on unit $i$'s own treatment status but also on the treatment choices of other units within the same group. For example, a firm's market share is influenced both by its own pricing decisions and by those of its competitors. Within a friendship network, an individual's labor earnings may be influenced by her best friend's education attainment, not only through direct support or access to resources, but also through information-sharing or social learning mechanisms that facilitate the transmission of knowledge about opportunities, norms, and strategies. In such contexts, the Stable Unit Treatment Value Assumption (SUTVA) may be violated, which motivates researchers to develop models that explicitly allow for spillover effects in outcomes.

The binary treatment decision of unit $i$ in group $g$ is denoted by $D_{ig} \in \{0,1\}$, where $D_{ig} = 1$ indicates that unit $i$ adopts the treatment and $D_{ig} = 0$ otherwise. In many applications, treatment decisions are not randomly assigned but instead depend on unobserved characteristics that also influence outcomes, giving rise to endogeneity concerns. For instance, a firm's pricing decision or an individual's education choice may both be endogenously determined. In group settings, units may make their decisions simultaneously, taking into account private information as well as expectations about the behavior of other group members. Each unit's decision depends on its own private information, denoted by $V_{ig}$, as well as on its expectations about the probability that other members of the group will adopt the treatment. Units form their expectations on the basis of publicly observed variables $(Z_{ig}, Z_{-ig})$, where $Z_{ig}$ denotes the random assignment received by unit $i$ in group $g$, and $Z_{-ig}$ denotes the assignments of the remaining group members. The vector $(Z_{ig}, Z_{-ig})$ thus serves as the set of instrumental variables for addressing endogeneity in treatment decisions. Since each unit's treatment choice may respond to the assignments received by other group members, these instruments can also induce spillover effects in treatment selection.

Building on the setting described above, consider the following model for unit $i$ in group $g$, where the peer of unit $i$ is denoted by $-i$. For clarity of exposition, I focus on the case in which each group $g$ consists of two units, indexed by $i = \{0,1\}$, while noting that the identification and estimation results extend straightforwardly to groups with more than two members:

equation[equation omitted — 240 chars of source]

The first line of Equation (ref) specifies the outcome equation. In this framework, unit $i$'s outcome $Y_{ig}$ depends on her own treatment $D_{ig}$ and on her group member's treatment, $D_{-ig}$, which explicitly models spillover effects. Importantly, I also allow $Y_{ig}$ to depend on both unit $i$'s own unobservables $U_{ig}$ and the unobservables of her group member, $U_{-ig}$. For example, in the oligopoly market, this specification captures the possibility that firm $i$'s market share $Y_{ig}$ is influenced not only by its own unobserved product characteristics $U_{ig}$ but also by the unobserved product characteristics of its rival, $U_{-ig}$.

Throughout the paper, I define the potential outcome for unit $i$ when her treatment is set to $d$ and her group member's treatment to $d'$ as $Y_{ig}(d, d') \equiv m_i(d, d', U_{ig}, U_{-ig})$. The observed outcome $Y_{ig}$ is determined according to the following equation,

equation*[equation* omitted — 227 chars of source]

This paper focuses on identifying reduced-form causal effects arising from a unit's own treatment and peers' treatments, rather than the underlying structural parameters specified in structural equations. A detailed comparison with a system of structural equations is provided in the Appendix (ref).

The framework imposes no functional form restrictions on the outcome equation $m_i$, and the subscript $i$ indicates that each group member, $i \in \{0,1\}$, may have a distinct outcome equation, meaning their functional forms are not required to be identical. It also places no restrictions on the dimension of the unobserved components $(U_{ig}, U_{-ig})$. Consequently, the influence of the peer's treatment $D_{-ig}$ on unit $i$'s outcome $Y_{ig}$ remains fully unrestricted. This generality provides a flexible structure that accommodates complex and heterogeneous spillover patterns in outcomes.

The second line of Equation (ref) characterizes the treatment selection mechanism. The treatment decision of unit $i$, $D_{ig}$, may be endogenous because it is determined by a continuous unobserved factor $V_{ig}$ that can also influence the outcome. The treatment selection $D_{ig}$ depends only on the unit's own unobservable $V_{ig}$ and not directly on her group member's unobservable $V_{-ig}$. This restriction is plausible in many applications. For example, in the oligopoly market discussed above, a firm's pricing decision is driven by its own private demand shock, while the competitor's demand shock is unobserved and therefore cannot directly affect the firm's decision rule. In the returns to education example, an individual's education decision is determined solely by her own education costs. The best friend's education costs, which are unobserved to the individual, do not directly influence her schooling decision.

Crucially, it is empirically reasonable to allow the unobserved factors $V_{ig}$ and $V_{-ig}$ to be arbitrarily dependent, since group members often share related unobserved characteristics or are exposed to common shocks. This dependence further complicates identification, and the framework accommodates it without imposing parametric restrictions on the joint distribution of unobservables. This flexibility accommodates a wide range of empirically relevant correlations. In the oligopoly setting, correlation across firms' idiosyncratic shocks arises naturally. For instance, a market-wide change in consumer tastes or a new advertising regulation may simultaneously affect how all products are perceived by consumers, thereby inducing correlation between the demand shocks $V_{ig}$ and $V_{-ig}$.

I model the treatment selection mechanism using a single threshold crossing rule: unit $i$'s chooses to take the treatment, $D_{ig} = 1$, if the unobserved factor $V_{ig} \in \mathbb{R}$ does not exceed a threshold $h_i(Z_{ig}, Z_{-ig})$, where $ h_i: \mathbb{R}^{k_i} \times \mathbb{R}^{k_{-i}} \mapsto \mathbb{R} $ is an unspecified function. I do not impose a parametric form on the threshold function $h_i$, and the subscript $i$ emphasizes that its functional form may differ across units $i \in \{0,1\}$ within the same group. In contrast to complete information games, where unit $i$'s treatment $D_{ig}$ directly depends on the treatment decision of her peer $D_{-ig}$, my framework is consistent with a simultaneous-move game with incomplete information, as studied by aradillas2010semiparametric and related papers. In this setting, $D_{ig}$ corresponds to player $i$'s action, and $(Z_{ig}, Z_{-ig})$ represent publicly observed signals that serve as instrumental variables. Each unit observes $(Z_{ig}, Z_{-ig})$ and forms beliefs about the joint treatment choices within the group, specifically the probability $\mathbb{P}(D_{ig} = 1, D_{-ig} = 1 \mid Z_{ig}, Z_{-ig})$, and then chooses her treatment based on these beliefs. Thus, treatment decisions are interdependent through expectations rather than through observing others' realized treatment choices. As shown in aradillas2010semiparametric, the optimal decision rule in such simultaneous-move incomplete information game is consistent with a single-index threshold-crossing structure of the form in Equation (ref). Appendix (ref) provides a detailed discussion.

For identification, which I discuss in detail later, the instrumental variables $(Z_{ig}, Z_{-ig})$ must be independent of the unobserved heterogeneity $(U_{ig}, U_{-ig}, V_{ig}, V_{-ig})$ in the group and must not directly affect the outcomes $(Y_{ig}, Y_{-ig})$. In studies that focus on complete information setting, such as those analyzed by balat2023multiple and hoshino2023treatment, strategic interactions between $D_{ig}$ and $D_{-ig}$ are modeled explicitly, but unit $i$'s treatment is not allowed to depend on her group member's instrument $Z_{-ig}$. A distinguishing feature of this framework is that it does not rely on any additional exclusion restrictions on instruments: I allow the treatment $D_{ig}$ to depend on both the unit's own instrument $Z_{ig}$ and the peer's instrument $Z_{-ig}$, thereby accommodating potential spillovers from instruments into treatment decisions. Moreover, the instrumental variables may take the form of common public signals observed by all group members, so that the same variable $Z$ serves as the instrument for each member, or unit-specific instruments that vary across members, $Z_{ig} \neq Z_{-ig}$. This framework accommodates both shared and individual sources of exogenous variation.

Figure (ref) presents a directed acyclic graph (DAG) that illustrates the causal relationships among the key variables within group $g$. The red arrows represent spillover channels: unit $i$'s outcome $Y_{ig}$ may depend on her peer's treatment $D_{-ig}$, and her treatment $D_{ig}$ may depend on her peer's instrument $Z_{-ig}$. Direct interaction between treatments $D_{ig}$ and $D_{-ig}$, however, is ruled out. The unobserved heterogeneity $V_{ig}$ and $V_{-ig}$ introduce endogeneity, as they may simultaneously affect both treatments and outcomes. Those are represented by the black dashed arrows. The blue dashed arrow reflects potential dependence between $V_{ig}$ and $V_{-ig}$, for which I do not impose any functional restrictions.

figure[figure omitted — 2,523 chars of source]

Assumptions (ref)-(ref) set out the maintained restrictions on the key variables that are imposed throughout the paper.

assumption(Random assignment) The instrumental variables assigned to all members within a group are jointly independent of the group's unobserved heterogeneity: \begin{equation*} \big(Z_{ig}, Z_{-ig}\big) \perp \!\!\! \perp \big(V_{ig}, V_{-ig}, U_{ig}, U_{-ig}\big) \end{equation*} for $i, -i \in \{0, 1\}$.

Assumption (ref) requires that instruments are randomly assigned at the group level, implying that the group-level instrument vector $(Z_{ig}, Z_{-ig})$ is independent of the unobserved heterogeneity of all units in the group. This assumption places no restrictions on the dependence structure between $Z_{ig}$ and $Z_{-ig}$ within a group. The instruments may be arbitrarily correlated across units within a group, as long as they remain jointly independent of the unobserved heterogeneity $(V_{ig}, V_{-ig}, U_{ig}, U_{-ig})$.

assumption(Exclusion restriction) Given $d_0, d_1$ and $u_0, u_1$, the instrumental variables $(Z_{ig}, Z_{-ig})$ do not directly affect the outcome $Y_{ig}$: \begin{equation*} m_{i}\big(d_0, d_1, z_0, z_1, u_0, u_1\big) = m_{i}\big(d_0, d_1, z_0', z_1', u_0, u_1\big) \end{equation*} for any $z_0 \neq z'_0$ and $z_1 \neq z'_1$.

Assumption (ref) requires that the instruments affect the outcome only through their influence on treatment take-up, without exerting any direct effect on the outcome. This condition corresponds to the standard exclusion restriction commonly imposed in instrumental variable analyses.

assumption(Distribution of $V_{ig}$) The unobserved variable $V_{ig}$ is continuously distributed.

Assumption (ref) requires that the unobserved heterogeneity $V_{ig}$ has a continuous distribution, which is a common condition in the literature. Under this assumption, $V_{ig}$ can be normalized to follow a uniform distribution on the interval $(0,1)$.

Another implicit restriction embedded in the treatment selection equation is a monotonicity structure. Specifically, consider the case in which instruments $Z_{ig}$ and $Z_{-ig}$ are binary, taking values in $\{z_0, z_1\}$. If the threshold function satisfies the following ordering condition:

equation*[equation* omitted — 92 chars of source]

for all group members $i$ and $-i$, then the treatment selection equation implies a corresponding monotonicity property for treatment take-up, consistent with the condition studied in the literature (e.g., vazquez2023causal):

equation*[equation* omitted — 104 chars of source]

for each $i$ and $-i$ within group $g$.

The proposed framework applies to a broad class of empirical settings where spillovers operate through both outcomes and endogenous treatment decisions. Illustrative examples include oligopoly markets, where firms' pricing decisions may influence competitors' market shares, and education contexts, where an individual's labor market outcomes depend on the best friend's schooling decision. More broadly, the framework can be extended to settings such as households, where behaviors involving risky activities generate spillover effects on the health outcomes of other members.

example(Duopoly market: pricing decisions) To illustrate, consider an oligopoly market with two competing firms, Costco and Sam's Club. Each firm decides whether to raise the price of its membership card and is interested in how this decision affects its market share. A firm's market share depends not only on its own pricing decision but also on its competitor's pricing strategy, giving rise to spillover effects from one firm's decision to the other's outcome. Assume that pricing decisions are made simultaneously and that neither firm observes the other's choice at the decision stage. Costco's decision, denoted by $D_{ig}$, depends on a private demand shock $V_{ig}$, such as an idiosyncratic change in reputation or advertising effectiveness, that is unobserved by Sam's Club. This unobserved factor affects both Costco's incentive to raise its membership price and its resulting market share $Y_{ig}$, thereby introducing endogeneity. Although each firm does not directly observe its competitor's pricing decision, both form expectations about rival behavior based on publicly observed market signals $(Z_{ig}, Z_{-ig})$, such as industry-wide cost shocks like tariffs, which are plausibly exogenous and can serve as valid instruments. Moreover, a tariff shock affecting Sam's Club may also influence Costco's pricing decision, generating instrumental spillovers from one firm's assignment to the other's endogenous treatment.
example(Friendship Network: education decision) Consider a friendship network consisting of two best friends who decide whether to pursue higher education. Each individual's education choice may affect not only her own future earnings but also her friend's, generating spillover effects through information sharing, social learning, or mutual support mechanisms. Assume that education decisions are made simultaneously and that neither friend observes the other's choice at the decision stage. Each individual's decision depends on an idiosyncratic unobserved factor, such as intrinsic academic motivation or costs, that influences both the probability of attending college and future earnings, thereby creating endogeneity. While friends do not observe each other's choices, their education decisions may be jointly influenced by shared public characteristics, such as the average family background of classmates in their school cohort. This shared characteristic can serve as a plausible instrumental variable, as it is typically exogenous to individual-specific unobserved ability but may affect education choices through peer effects (see bifulco2011effect, bifulco2014high, cools2019girls). The individual is also exposed to public characteristics associated with her best friend, generating instrumental spillovers from the friend's assignment to the individual's endogenous education decision.

Local Average Effects and Identification Strategy

A central insight from the treatment effect literature is that, when treatment assignment is endogenous, causal effects can often be identified for specific subpopulations defined by the instrument, for example, the local average treatment effect (LATE) in imbens1994identification and related studies. However, in the presence of spillovers, individuals' outcomes depend not only on their own treatment but also on the treatments received by others in their group. The spillovers complicate the interpretation of conventional LATE parameters, as variation in peers' treatments introduces additional causal channels. To disentangle these channels, this section extends the LATE framework to define local average effects that separately capture the causal effect of peers' treatments on an individual's outcome and the direct effect of the individual's own treatment. These parameters retain the causal interpretability of LATE while accommodating the presence of endogenous treatment and spillovers within groups. The identification of these local average effects further motivates the development of a framework based on marginal treatment effects, which explicitly accounts for spillovers operating through both treatment selection and outcomes, as formalized in Section (ref). Since the identification analysis is conducted at the level of a super-population of groups, I suppress the group subscript $g$ throughout this section to simplify notation.

Building on the framework introduced in Section (ref), the expected potential outcome $Y_i(d, d')$ generally depends on both the individual's and her peers' unobserved characteristics, $(V_i, V_{-i})$. Following the terminology in the literature, I refer to the conditional expectations $\mathbb{E}[Y_i(d, d') \mid (V_i, V_{-i}) \in P]$, where $P$ denotes a subset of the support of $(V_i, V_{-i})$, as local average potential outcomes. These parameters capture the average potential outcomes for subpopulations defined by specific values of the group-level unobservables, which reflect the underlying unobserved heterogeneity in the population. Taking appropriate differences between local average potential outcomes under different treatment combinations $(d, d')$ yields the average spillover effects from peers' treatments and the direct effects from a unit's own treatment. Definition (ref) provides formal definitions of these parameters.

definition(Generalized local average controlled effects) Consider the model in Equation ((ref)). \begin{enumerate} • Fix the treatment of unit $i$ at $D_i = d$, $d \in \{0,1\}$. The generalized local average controlled spillover effect (LACSE), conditional on the group-level unobserved heterogeneity satisfying $(V_i, V_{-i}) \in P$ for some subset $P \subset (0,1)^2$, is defined as as \begin{equation*} \operatorname{LACSE}_{i}^{(d)}(P) \equiv \mathbb{E}[Y_{i}(d, 1) - Y_{i}(d, 0) \mid (V_i, V_{-i}) \in P]. \end{equation*} • For unit $i$, fix the peer's treatment at $D_{-i} = d$, where $d \in {0,1}$. The generalized local average controlled direct effect (LACDE), conditional on the group-level unobserved heterogeneity satisfying $(V_i, V_{-i}) \in P$ for some subset $P \subset (0,1)^2$, is defined as as \begin{equation*} \operatorname{LACDE}_{i}^{(d)}(P) \equiv \mathbb{E}[Y_{i}(1, d) - Y_{i}(0, d) \mid (V_i, V_{-i}) \in P]. \end{equation*} \end{enumerate}

The generalized LACSEs capture counterfactual spillover effects by exogenously fixing a unit's own treatment, whereas the generalized LACDEs capture counterfactual direct effects by exogenously fixing peers' treatments. Both effects are defined conditional on a subpopulation characterized by $(V_i, V_{-i}) \in P$, in the same spirit as the LATE widely studied in the literature. These parameters possess clear causal interpretations, as they disentangle the distinct influence channels of a unit's own treatment and peers' treatments, while allowing for unobserved treatment effect heterogeneity through conditioning on group-level unobservables. I now establish the identification of the generalized LACSEs and LACDEs using instrumental variables under Assumptions (ref)-(ref). It is worth noting that this identification result accommodates both discrete instruments with multiple support points and continuously distributed instruments.

I define the propensity score function for unit $i$ as the probability of treatment conditional on the group-level instrument vector, $P_{i}(Z_i, Z_{-i}) \equiv \mathbb{P}(D_i = 1 \mid Z_i, Z_{-i}) $. I denote this function simply by $P_i$ and define the support of the propensity scores for all group members as $\mathcal{P} \equiv \text{Supp}(P_i, P_{-i})$. The propensity score function $P_i$ identifies the threshold function $h_i$ in the treatment selection equation for each group member $i$, as demonstrated in the following derivation:

equation[equation omitted — 316 chars of source]

Additionally, the propensity scores $(P_i, P_{-i})$ are independent of all group members' unobserved heterogeneity, since they are functions of the instruments $(Z_i, Z_{-i})$.

Building on the identification of the propensity score, Theorem (ref) establishes the identification of the generalized LACSEs and LACDEs.

theorem(Identifying generalized LACSEs and LACDEs) Suppose that Assumptions (ref)-(ref) hold and let $d \in \{0,1\}$. Under the following conditions, the generalized LACSEs and LACDEs defined in Definition (ref) can be identified. \begin{enumerate} • Suppose there exist $(p_0, p_1), (p_0, p_1') \in \mathcal{P}$, $p_1' > p_1$. The local average controlled spillover effect, $\text{LACSE}_i^{(1)}(P)$ for $P =\{V_i \leq p_0, p_1 < V_{-i} \leq p_1'\}$, can be identified as \begin{equation*} LACSE_i^{(1)}(V_i \leq p_0, p_1 < V_{-i} \leq p_1') = \frac{\mu_{i,i}^{(1)}(p_0, p_1') - \mu_{i,i}^{(1)}(p_0, p_1)}{C(p_0, p_1') - C(p_0, p_1)}. \end{equation*} The local average controlled spillover effect, $\text{LACSE}_i^{(0)}(P)$ for $P =\{V_i > p_0, p_1 < V_{-i} \leq p_1'\}$, can be identified as \begin{equation*} LACSE_i^{(0)}(V_i > p_0, p_1 < V_{-i} \leq p_1') = \frac{\mu_{i,i}^{(0)}(p_0, p_1') - \mu_{i,i}^{(0)}(p_0, p_1)}{\big(p_1' - p_1\big) - \big[C(p_0, p_1') - C(p_0, p_1)\big]}, \end{equation*} where $\mu_{i,i}^{(d)}(p_0, p_1) \equiv \mathbb{E}[Y_i \mathbbm{1}\{D_i = d\} \mid P_i = p_0, P_{-i} = p_1]$, and $C(p_0, p_1) \equiv \mathbb{E}[D_i D_{-i} \mid P_i = p_0, P_{-i} = p_1]$. • Suppose there exist $(p_0, p_1), (p_0', p_1) \in \mathcal{P}$, $p_0' > p_0$. The local average controlled direct effect, $\text{LACDE}_i^{(1)}(P)$ for $P =\{p_0 < V_i \leq p_0', V_{-i} \leq p_1\}$, can be identified as \begin{equation*} LACDE_i^{(1)}(p_0 < V_i \leq p_0', V_{-i} \leq p_1) = \frac{\mu_{i,-i}^{(1)}(p_0', p_1) - \mu_{i,-i}^{(1)}(p_0, p_1)}{C(p_0', p_1) - C(p_0, p_1)}. \end{equation*} \textit{The local average controlled direct effect}, $\text{LACDE}_i^{(0)}(P)$ for $P =\{p_0 < V_i \leq p_0', V_{-i} > p_1\}$, can be identified as \begin{equation*} \text{LACDE}_i^{(0)}(p_0 < V_i \leq p_0', V_{-i} > p_1) = \frac{\mu_{i,-i}^{(0)}(p_0', p_1) - \mu_{i,-i}^{(0)}(p_0, p_1)}{\big(p_0' - p_0\big) - \big[C(p_0', p_1) - C(p_0, p_1)\big]}, \end{equation*} where $\mu_{i,-i}^{(d)}(p_0, p_1) \equiv \mathbb{E}[Y_i \mathbbm{1}\{D_{-i} = d\} \mid P_i = p_0, P_{-i} = p_1]$. • Suppose there exist $(p_0, p_1)$, $(p_0, p_1')$, $(p_0', p_1)$, $(p_0', p_1') \in \mathcal{P}$. \textit{The local average controlled spillover effect}, $\text{LACSE}_i^{(d)}(P)$ for $P =\{p_0 < V_i \leq p_0', p_1 < V_{-i} \leq p_1'\}$, can be identified as \begin{equation*} \begin{aligned} &\text{LACSE}_i^{(d)}(p_0 < V_i \leq p_0', p_1 < V_{-i} \leq p_1') \\ =& \text{sgn}(2d - 1) \cdot \frac{\big[\mu_{i, i}^{(d)}(p_0', p_1') - \mu_{i, i}^{(d)}(p_0', p_1)\big] - \big[\mu_{i, i}^{(d)}(p_0, p_1') - \mu_{i, i}^{(d)}(p_0, p_1)\big]}{\big[C(p_0', p_1') - C(p_0', p_1)\big] - \big[C(p_0, p_1') - C(p_0, p_1)\big]}. \end{aligned} \end{equation*} \textit{The local average controlled direct effect}, $\text{LACDE}_i^{(d)}(P)$ for $P =\{p_0 < V_i \leq p_0', p_1 < V_{-i} \leq p_1'\}$, can be identified as \begin{equation*} \begin{aligned} &\text{LACDE}_i^{(d)}(p_0 < V_i \leq p_0', p_1 < V_{-i} \leq p_1') \\ =& \text{sgn}(2d - 1) \cdot \frac{\big[\mu_{i, -i}^{(d)}(p_0', p_1') - \mu_{i, -i}^{(d)}(p_0', p_1)\big] - \big[\mu_{i, -i}^{(d)}(p_0, p_1') - \mu_{i, -i}^{(d)}(p_0, p_1)\big]}{\big[C(p_0', p_1') - C(p_0', p_1)\big] - \big[C(p_0, p_1') - C(p_0, p_1)\big]}. \end{aligned} \end{equation*} \end{enumerate}
proofSee Appendix (ref).

The identification argument proceeds as follows, with the formal proof provided in Appendix (ref). Given a pair of observed propensity scores $(P_i, P_{-i}) = (p_0, p_1)$, and noting that the propensity score $P_i$ identifies the threshold function $h_i$, the joint treatment realizations $(D_i, D_{-i})$ partition the space of unobserved heterogeneity $(V_i, V_{-i})$ into four mutually exclusive subpopulations, separated by the thresholds $(p_0, p_1)$. The relationships are summarized as

equation[equation omitted — 364 chars of source]

For instance, the probability of observing $\{D_i = 1, D_{-i} = 1\}$ conditional on $(P_i, P_{-i}) = (p_0, p_1)$ identifies the share of the subpopulation with $\{V_i \leq p_0, V_{-i} \leq p_1\}$, that is,

equation*[equation* omitted — 126 chars of source]

The top left panel of Figure (ref) illustrates how these four subpopulations correspond to distinct realizations of $(D_i, D_{-i})$ given observed propensity scores $(p_0, p_1)$.

figure[figure omitted — 252 chars of source]

Given the data, the conditional expectation $\mathbb{E}[Y_i \mathbbm{1}\{D_i = d, D_{-i} = d'\} \mid P_i = p_0, P_{-i} = p_1]$ can be directly recovered from observables. Under the model framework and Assumptions (ref)-(ref), these observed moments identify the average potential outcomes $Y_i(d, d')$ for the subpopulations associated with the treatment realization $\{D_i = d, D_{-i} = d'\}$. These quantities form the foundation for the identification strategy of Theorem (ref). For instance, when $(D_i, D_{-i}) = (1,1)$,

equation*[equation* omitted — 146 chars of source]

which identifies the average potential outcome $Y_i(1,1)$ for the subpopulation with unobserved characteristics satisfying $\{V_i \leq p_0, V_{-i} \leq p_1\}$.

The identification of the generalized LACSEs and LACDEs exploits exogenous variation in the propensity score values. Suppose there exists another pair of observed propensity scores $(p_0, p_1')$, $p_1' > p_1$. By shifting the propensity scores from $(p_0, p_1)$ to $(p_0, p_1')$ and applying the relationships established in Equation (ref), the subpopulation with unobserved characteristics in the region $\{V_i \leq p_0, p_1 < V_{-i} \leq p_1'\}$ changes its treatment status from $(D_i, D_{-i}) = (1, 0)$ to $(1, 1)$. Likewise, the subpopulation in $\{V_i > p_0, p_1 < V_{-i} \leq p_1'\}$ changes from $(D_i, D_{-i}) = (0, 0)$ to $(0, 1)$. These two groups correspond to the blue-shaded areas in the top-right panel of Figure (ref).

In both cases, only the peer $-i$ changes her treatment status $D_{-i}$, providing variation that identifies the average spillover effect. Taking the difference between the conditional expectations, $\mathbb{E}[Y_i \mathbbm{1}\{D_i = d, D_{-i} = 1\} \mid \cdot, \cdot] - \mathbb{E}[Y_i \mathbbm{1}\{D_i = d, D_{-i} = 0\} \mid \cdot, \cdot]$, evaluated at $(p_0, p_1)$ and $(p_0, p_1')$, isolates the variation in outcomes attributable to the subpopulations that experience a change in peer treatment status. This difference identifies the local average controlled spillover effects, $\operatorname{LACSE}_i^{(1)}$ and $\operatorname{LACSE}_i^{(0)}$, for the subpopulations corresponding to the two blue-shaded regions in the top-right panel of Figure (ref). This variation yields the identification result stated in Item 1 of Theorem (ref).

Analogously, when another pair of propensity scores $(p_0', p_1)$ with $p_0' > p_0$ is observed, shifting from $(p_0, p_1)$ to $(p_0', p_1)$ and using the relationships in Equation (ref) induces changes in treatment status for unit $i$ only. Specifically, the subpopulations defined by $\{p_0 < V_i \leq p_0', V_{-i} \leq p_1\}$ and $\{p_0 < V_i \leq p_0', V_{-i} > p_1\}$ change their treatment realizations from $\{D_i = 0, D_{-i} = 1\}$ to $\{D_i = 1, D_{-i} = 1\}$ and from $\{D_i = 0, D_{-i} = 0\}$ to $\{D_i = 1, D_{-i} = 0\}$, respectively, as illustrated by the two yellow-shaded regions in the bottom-left panel of Figure (ref). Because only unit $i$ changes treatment status, the resulting variation identifies the local average controlled direct effect. This variation corresponds to the identification result presented in Item 2 of Theorem (ref).

Suppose the support of the propensity scores exhibits sufficient variation such that four distinct pairs, $(p_0, p_1)$, $(p_0, p_1')$, $(p_0', p_1)$, and $(p_0', p_1')$, are observed with $p_0' > p_0$ and $p_1' > p_1$. Applying the same logic as before, shifting the peer's propensity score from $p_1$ to $p_1'$ while fixing unit $i$'s score at $p_0'$ and, conversely, shifting unit $i$'s score from $p_0$ to $p_0'$ while fixing the peer's score at $p_1'$, identifies the corresponding LACSEs and LACDEs for subpopulations defined by these regions of $(V_i, V_{-i})$. Next, taking cross-differences of the local average effects across the four propensity-score pairs, $(p_0, p_1)$, $(p_0, p_1')$, $(p_0', p_1)$, and $(p_0', p_1')$, isolates the LACSEs and LACDEs for the subpopulation with $(V_i, V_{-i})$ lying in the rectangle $\{p_0 < V_i \leq p_0', p_1 < V_{-i} \leq p_1'\}$, illustrated by the green-shaded area in the bottom-right panel of Figure (ref).

For example, the difference between $\operatorname{LACSE}_i^{(1)}$ identified for the regions $\{V_i \leq p_0, p_1 < V_{-i} \leq p_1'\}$ and $\{V_i \leq p_0', p_1 < V_{-i} \leq p_1'\}$ yields $\operatorname{LACSE}_i^{(1)}$ for the subpopulation $\{p_0 < V_i \leq p_0', p_1 < V_{-i} \leq p_1'\}$. Similarly, analogous differences yield $\operatorname{LACSE}_i^{(0)}$ and $\operatorname{LACDE}_i^{(d)}$, $d \in \{0, 1\}$, for the same region. This result is formally presented in Item 3 of Theorem (ref), which requires the support of $(P_i, P_{-i})$ to contain four distinct points forming the “vertices” of a rectangle in the $(p_0, p_1)$ space.

remark(Identification With Binary Instrument) The conclusions in Theorem (ref) apply to both discrete and continuous instrumental variables, provided that they generate sufficient variation in the propensity scores to satisfy the identification conditions. To illustrate, consider the case where a binary instrumental variable is assigned to both units $i$ and $-i$ within each group. In this case, four distinct pairs of propensity scores $(P_i, P_{-i})$ can be observed, corresponding to the four possible combinations of instrument assignments in each group: \begin{equation*} \begin{array}{c|cc} & Z_{-i}=0 & Z_{-i}=1 \\ \hline Z_i=0 & \left(P_i(0,0), P_{-i}(0,0)\right) & \left(P_i(0,1), P_{-i}(1,0)\right) \\ Z_i=1 & \left(P_i(1,0), P_{-i}(0,1)\right) & \left(P_i(1,1), P_{-i}(1,1)\right) \end{array} \end{equation*} According to Theorem (ref), identifying the LACSEs requires at least two pairs of propensity scores where unit $i$'s score remains fixed while the peer $-i$'s score varies, and identifying the LACDEs instead requires variation in unit $i$'s propensity score while holding the peer $-i$'s propensity score fixed. When the instruments are binary, the support of $(P_i, P_{-i})$ consists of only four points. In this case, identification relies on specific equalities among these propensity scores: identification of LACSEs requires that any two of $P_i(0,0)$, $P_i(0,1)$, $P_i(1,0)$, or $P_i(1,1)$ be equal, while identification of LACDEs requires that any two of $P_{-i}(0,0)$, $P_{-i}(0,1)$, $P_{-i}(1,0)$, or $P_{-i}(1,1)$ be equal. These conditions include the one-sided noncompliance restriction as a special example, a commonly imposed assumption in the literature to achieve identification with binary instruments kang2016peer, vazquez2023causal, ditraglia2023identifying. The one-sided noncompliance assumption requires that a unit cannot take the treatment unless it receives the instrument assignment, that is, $\mathbb{P}(D_i = 1 \mid Z_i = 0) = 0$ for all $i$. This restriction is equivalent to \begin{equation*} \begin{aligned} P_i(0, 0) = P_i(0, 1) = 0, \quad P_{-i}(0, 0) = P_{-i}(0, 1) = 0, \end{aligned} \end{equation*} which satisfies the conditions of Items 1 and 2 in Theorem (ref). Under this structure, the LACSEs and LACDEs identified by Theorem (ref) coincide with the local average spillover and direct effects studied in the existing literature vazquez2023causal. As shown in Appendix (ref), when the instrument is binary, additional restrictions such as one-sided noncompliance are therefore necessary to point identify local average effects. Interested readers are referred to Appendix (ref) for a detailed discussion.

Theorem (ref) demonstrates that local average effects can be identified when the instrumental variables exhibit sufficient variation to induce the necessary differences in propensity scores. As discussed in Remark (ref), when instruments take only binary values, additional restrictions are required to achieve point identification of certain local average spillover or direct effects. Together, these results highlight that adequate variation in the instruments is crucial for identifying causally interpretable parameters in settings where spillovers influence both outcomes and treatment selection.

When the instrumental variables exhibit continuous variation, they induce continuous variation in the propensity scores. In this case, one can extend the identification strategy in Theorem (ref) by taking limits as $p_0' \to p_0$ and $p_1' \to p_1$, thereby identifying the average controlled spillover and direct effects conditional on $(V_i, V_{-i})$ evaluated at a specific point $(p_0, p_1)$ within the interior of the propensity score support. The next section formalizes this idea by introducing the marginal controlled spillover and marginal controlled direct effects. These parameters serve as building blocks for identifying not only the local average controlled spillover and direct effects discussed above, but also a broader class of policy-relevant treatment effects of interest to researchers.

Marginal Effects and Identification Results

Definition (ref) formally defines the causal spillover and direct effects evaluated at specific values of the unobserved characteristics $(V_i, V_{-i})$.

definition(Marginal controlled spillover effects (MCSE) and marginal controlled direct effects (MCDE)) Consider the model in Equation ((ref)). \begin{enumerate} • Fix the treatment of unit $i$ at $D_i = d$, $d \in \{0,1\}$. The marginal controlled spillover effect (MCSE), given $V_i = p_0$ and $V_{-i} = p_1$, is defined as \begin{equation*} \operatorname{MCSE}_{i}^{(d)}(p_0, p_1) \equiv \mathbb{E}[Y_{i}(d, 1) - Y_{i}(d, 0) \mid V_{i} = p_0, V_{-i} = p_1], (p_0, p_1) \in (0, 1)^2. \end{equation*} • For unit $i$, fix the peer's treatment at $D_{-i} = d$, $d \in \{0,1\}$. The marginal controlled direct effect (MCDE), given $V_i = p_0$ and $V_{-i} = p_1$, is defined as \begin{equation*} \operatorname{MCDE}_{i}^{(d)}(p_0, p_1) \equiv \mathbb{E}[Y_{i}(1, d) - Y_{i}(0, d) \mid V_{i} = p_0, V_{-i} = p_1], (p_0, p_1) \in (0, 1)^2 \end{equation*} \end{enumerate}

The marginal controlled spillover effect captures the impact of changing the peer's treatment on a unit's potential outcome while holding the unit's own treatment status fixed, conditional on the unobserved characteristics $(V_i, V_{-i})$ within the group. Similarly, the marginal controlled direct effect measures the impact of changing a unit's own treatment on her potential outcome while holding the peer's treatment constant, again conditional on $(V_i, V_{-i})$. Because both effects are defined relative to the group-level unobserved heterogeneity, they capture treatment effect heterogeneity arising from latent factors within the group.

The marginal controlled spillover and direct effects are defined analogously to the standard marginal treatment effect (MTE), conditioning on continuous unobserved heterogeneity within the support of the latent variables. Unlike the conventional MTE framework, which rules out interference across units, the marginal controlled effects explicitly incorporate spillovers arising from peers' treatment decisions. As such, they extend the MTE concept to environments where spillovers exist in both outcomes and treatment selection. Section (ref) formally characterizes the connection and distinction between these effects and the standard MTE. By conditioning on the continuous unobservables, the marginal controlled effects provide the building blocks for a wide class of policy-relevant parameters. In particular, the generalized local average controlled effects introduced in Definition (ref) represent a specific class of policy-relevant parameters that can be obtained by integrating the marginal controlled effects over selected regions of the latent heterogeneity space. These connections will be discussed in detail in Section (ref). The policy-relevant effects play a central role in evaluating counterfactual policies in settings with endogenous treatment and spillovers.

As discussed in the setting, the unobserved heterogeneities $V_i$ and $V_{-i}$ within a group may be correlated. Identification of the parameters of interest requires recovering the joint density of $(V_i, V_{-i})$. Because the marginal distributions of $V_i$ and $V_{-i}$ are normalized to be uniform on the interval $(0,1)$, their joint distribution is characterized by their copula. Formally, the copula is defined as

equation*[equation* omitted — 129 chars of source]

Lemma (ref) provides identification of this copula on the support of the propensity scores $(P_i, P_{-i})$ without imposing any functional form assumptions, where $P_i$ denotes unit $i$'s propensity score as defined in Equation (ref).

lemma(Copula of $(V_{i}, V_{-i})$) Under Assumptions (ref)-(ref), the copula between $V_i$ and $V_{-i}$ is identified as \begin{equation*} C_{V_i, V_{-i}}(p_0, p_1) = \mathbb{P}\big(D_i=1, D_{-i}=1 \mid P_i=p_0, P_{-i}=p_1\big), \end{equation*} for $(p_0, p_1) \in \mathcal{P}$, where $\mathcal{P}$ denotes the support of $(P_i, P_{-i})$.
proofSee Appendix (ref).

Let $c_{V_i, V_{-i}}(\cdot, \cdot)$ denote the copula density of $(V_i, V_{-i})$. Since Lemma (ref) establishes identification of the copula, the copula density can be obtained provided that the conditional probability $\mathbb{P}(D_i = 1, D_{-i} = 1 \mid P_i, P_{-i})$ is twice differentiable. This differentiability condition requires that $P_i$ and $P_{-i}$ exhibit continuous variation, which in turn implies that at least some components of the instrument vector $(Z_i, Z_{-i})$ must be continuously distributed. Assumption (ref) introduces this continuity requirement.

assumption(Continuous instruments) At least one component of the instrumental variables $(Z_i, Z_{-i})$ is continuously distributed.

It then follows that the copula density of $(V_i, V_{-i})$ is identified, as stated in Corollary (ref).

corollary(Copula density of $(V_i, V_{-i})$) Suppose that Assumptions (ref)-(ref) hold. If $\mathbb{E}[D_i D_{-i} \mid P_i=p_0, P_{-i}=p_0]$ is twice differentiable at $(p_0, p_1) \in \mathcal{P}$, the copula density $c_{V_i, V_{-i}}(p_0, p_1)$ is identified as \begin{equation*} c_{V_i, V_{-i}}(p_0, p_1) = \frac{\partial^2 \mathbb{E}\big[D_i D_{-i} \mid P_i=p_0, P_{-i}=p_1\big]}{\partial p_0 \partial p_1}. \end{equation*}
proofSee Appendix (ref).

Following the literature, the conditional expectation of the potential outcome, given the values of the group-level unobserved characteristics $(V_i, V_{-i})$,

equation*[equation* omitted — 145 chars of source]

is defined as the marginal treatment response (MTR) function. The marginal controlled spillover effects (MCSEs) and marginal controlled direct effects (MCDEs) introduced in Definition (ref) are obtained as differences of the corresponding MTR functions. Hence, identification of the MCSEs and MCDEs requires identifying the underlying MTR functions. Theorem (ref) provides the identification of the parameters of interest, the MCSEs and MCDEs, while the detailed process for identifying MTR functions is presented in Appendix (ref).

theorem(Identifying MCSEs and MCDEs) Suppose that Assumptions (ref)-(ref) hold. For $d_0, d_1 \in \{0,1\}$ and $(p_0, p_1)$ being an interior point of $\mathcal{P}$, the following additional regularity conditions are imposed: (i) $\mathbb{E}[D_i D_{-i} \mid P_i = p_0, P_{-i} = p_1]$ and $\mathbb{E}[Y_i \mathbbm{1}\{D_i = d_0\} \mathbbm{1}\{D_{-i} = d_1\} \mid P_i = p_0, P_{-i} = p_1]$ are twice differentiable; (ii) the marginal treatment response functions $m_i^{(d_0, d_1)}(p_0, p_1)$ are continuous; and (iii) the copula density $c_{V_i, V_{-i}}(p_0, p_1)$ is bounded from above and away from zero. Then, the marginal controlled spillover effects (MCSEs), $\operatorname{MCSE}_i^{(d)}(p_0, p_1)$, are identified as \begin{equation*} \begin{aligned} sgn(2d - 1) \cdot \frac{\partial^2 \mathbb{E}\big[Y_i \mathbbm{1}\{D_i=d\} \mid P_i=p_0, P_{-i}=p_1\big]}{\partial p_0 \partial p_1} \Big/ \frac{\partial^2 \mathbb{E}\left[D_i D_{-i} \mid P_i=p_0, P_{-i}=p_1\right]}{\partial p_0 \partial p_1}, \end{aligned} \end{equation*} and the marginal controlled direct effects (MCDEs), $\operatorname{MCDE}_i^{(d)}(p_0, p_1)$, are identified as \begin{equation*} \begin{aligned} sgn(2d - 1) \cdot\frac{\partial^2 \mathbb{E}\big[Y_i \mathbbm{1}\{D_{-i}=d\} \mid P_i=p_0, P_{-i}=p_1\big]}{\partial p_0 \partial p_1} \Big/ \frac{\partial^2 \mathbb{E}\left[D_i D_{-i} \mid P_i=p_0, P_{-i}=p_1\right]}{\partial p_0 \partial p_1}, \end{aligned} \end{equation*} for $d \in \{0,1\}$ and $(p_0, p_1)$ being an interior point of $\mathcal{P}$, and the function $\text{sgn(x)}$ denotes the sign of $x$.
proofSee Appendix (ref).

The identification of the MCSEs and MCDEs builds on the same logic underlying the identification of the LACSEs and LACDEs in part 3 of Theorem (ref), by taking the limits $p_1' \to p_1$ and $p_0' \to p_0$. This argument is illustrated by the green-shaded region in the bottom-right panel of Figure (ref), which represents the limiting case where $p_1' \to p_1$ and $p_0' \to p_0$. The validity of this limiting argument requires the propensity scores to exhibit continuous variation in a neighborhood of $(p_0, p_1)$. Moreover, the identification framework naturally extends to settings with exogenous covariates, with the corresponding results provided in Appendix (ref).

The assumption of continuously distributed instruments in Assumption (ref) is sufficient but not necessary for identifying the MCSEs and MCDEs. When instruments have limited or discrete variation, the parametric specifications in Assumption (ref) can be used to extrapolate the identification of marginal controlled effects beyond the observed support of the propensity scores. Alternative extrapolation approaches, such as those proposed by mogstad2018using, may also be applied. However, point identification may no longer hold, and the parameters would instead be partially identified. A formal treatment of this extension is left for future research.

remark(Groups with multiple individuals) The identification strategy naturally extends to settings with more than two individuals per group. Consider a group of size $n < \infty$, indexed by $i \in \{1, \ldots, n\}$. In this case, the threshold function $h_i$ in Equation (ref) depends on the full vector of instrument assignments $(Z_1, \ldots, Z_n)$. The propensity score $\mathbb{P}(D_i = 1 \mid Z_1, \ldots, Z_n)$ identifies the threshold $h_i$. The joint distribution of the unobserved heterogeneities within the group is then recovered from the conditional probability $\mathbb{P}(D_1 = 1, \ldots, D_n = 1 \mid P_1, \ldots, P_n)$, where $P_i$ denotes the propensity score of individual $i$. Once this joint distribution is identified, the marginal treatment response functions can be obtained by differentiating \begin{equation*} \mathbb{E}\left[Y_{i} \mathbbm{1}\left\{D_{1} = d_1\right\} \cdots \mathbbm{1}\left\{D_{n} = d_n\right\}\mid P_{1}, \cdots, P_{n}\right] \end{equation*} with respect to $(P_1, \ldots, P_n)$, under suitable smoothness conditions. Finally, differences between the resulting MTR functions yield the MCSEs and MCDEs.
remark(Testing the spillover structure) The identification of marginal treatment response functions makes it possible to test additional structural assumptions about the nature of spillovers. For example, in addition to Assumptions (ref)-(ref), suppose that each unit's outcome depends not on the entire treatment vector $\boldsymbol{D} \equiv (D_1, \cdots, D_n)$, but instead on a lower-dimensional function $H(\cdot)$ of this vector. In the literature, $H(\cdot)$ is commonly referred to as the exposure mapping. A standard specification is the average treatment level within the group, $H(\boldsymbol{D}) = \sum_{i=1}^n D_i / n$. Under this structure, given unit $i$'s own treatment $D_i = d$, any two treatment vectors $\boldsymbol{d} = (d_1, \cdots, d_n)$ and $\boldsymbol{\tilde{d}} = (\tilde{d}_1, \cdots, \tilde{d}_n)$ that generate the same exposure level, $H(\boldsymbol{d}) = H(\boldsymbol{\tilde{d}})$, should yield identical marginal treatment response functions: \begin{equation*} \mathbb{E}\big[Y_i(d, \boldsymbol{d}) \mid V_1 = p_1, \cdots, V_n = p_n\big] = \mathbb{E}\big[Y_i(d, \boldsymbol{\tilde{d}}) \mid V_1 = p_1, \cdots, V_n = p_n\big], \end{equation*} for all $(p_1, \cdots, p_n)$ in the support of the propensity score functions. This equality provides a testable implication of the assumed spillover structure $H(\cdot)$, thereby linking identification of MTR functions to specification testing of exposure mappings.

Policy Relevant Treatment Effects

The MCSEs and MCDEs not only capture heterogeneous spillover and direct effects but also serve as fundamental building blocks for deriving a wide range of causal parameters commonly examined in the literature. This section illustrates several examples demonstrating how the MCSEs and MCDEs can be used to recover other treatment effect parameters of policy relevance.

Average Controlled Spillover and Direct Effects

Researchers are often interested in summarizing heterogeneous spillover and direct effects across individuals by aggregating them into population-level parameters (see, for example, vazquez2023identification and related studies). Within this framework, the average controlled spillover effect (ACSE) can be defined as $\operatorname{ACSE}_i(d) \equiv \mathbb{E}[Y_i(d, 1) - Y_i(d, 0)]$, which measures the expected change in unit $i$'s outcome when the peer's treatment status changes exogenously from 0 to 1, holding the unit's own treatment fixed at $D_i = d$. Similarly, the average controlled direct effect (ACDE) can be defined as $\operatorname{ACDE}_i(d) \equiv \mathbb{E}[Y_i(1, d) - Y_i(0, d)]$, which reflects the expected change in unit $i$'s outcome when her own treatment status changes exogenously from zero to one, holding her peers' treatment status fixed at $D_{-i} = d$.

Under the assumption that the propensity scores have full support, i.e., $\mathcal{P} = (0,1)^2$, the ACSEs and ACDEs are point identified using the MCSEs, the MCDEs, and the copula density of $(V_i, V_{-i})$ identified in Section (ref), by integrating the MCSEs or MCDEs weighted by the corresponding copula density:

equation*[equation* omitted — 329 chars of source]

When the propensity scores lack full support, the average controlled spillover and direct effects, as well as other policy-relevant treatment effects, remain only partially identifiable. Under the additional assumption that potential outcomes are almost surely bounded, $|Y_i(d, d')| \leq B < \infty$, the MCSEs and MCDEs are confined within the range $[-2B, 2B]$ at points outside the observed support of the propensity scores. To extend identification beyond this region, one may impose the parametric structure in Assumption (ref) or adopt an extrapolation approach similar to that proposed by mogstad2018using. A formal development of these extensions is left for future research.

Local Average Controlled Spillover and Direct Effects

Once the MCSEs and MCDEs are point identified, they can be used to recover the LACSEs and LACDEs defined in Section (ref). The following discussion illustrates how these marginal effects can be employed to obtain the local average spillover and direct effects examined in the existing literature.

Suppose there exist two values of the instrumental variable, $z_0, z_1$, such that the associated propensity scores $P_i(z, z')$, for $z, z' \in {z_0, z_1}$, can be consistently ordered across all individuals and groups. Without loss of generality, assume that $P_i(z_0, z_0) \leq P_i(z_0, z_1) \leq P_i(z_1, z_0) \leq P_i(z_1, z_1)$. Under this ordering, the treatment selection mechanism in Equation (ref) implies the following monotonicity condition:

equation*[equation* omitted — 91 chars of source]

almost surely, where $D_i(z, z')$ denotes the potential treatment received by unit $i$ when the instrument assignments are fixed exogenously at $(Z_i, Z_{-i}) = (z, z')$.

vazquez2023causal partitions the population into a finite number of compliance types based on the values of the potential treatment vector ${D_i(z, z')}_{z, z' \in \{z_0, z_1\}}$. This framework identifies the local average spillover effect $\mathbb{E}[Y_i(0,1) - Y_i(0,0) \mid T_{-i} = c]$, and the local average direct effect $\mathbb{E}[Y_i(1,0) - Y_i(0,0) \mid T_i = c]$, where $T_i = c$ denotes the complier subgroup, defined as the set of units whose unobserved heterogeneity $V_i$ lies between the two thresholds $P_i(z_0, z_1)$ and $P_i(z_1, z_0)$. Intuitively, these are individuals who would not take the treatment under $(z_0, z_1)$ but would take it under $(z_1, z_0)$. This paper also considers the setting in which a one-sided noncompliance condition holds, meaning that individuals cannot receive the treatment when assigned the instrument value $z_0$. Under one-sided noncompliance, the propensity scores satisfy $0 = P_i(z_0, z_0) = P_i(z_0, z_1) \leq P_i(z_1, z_0) \leq P_i(z_1, z_1)$, which corresponds to a special case of the sufficient identification condition stated in Part 1 of Theorem (ref).

Figure (ref) in Appendix (ref) illustrates the regions in the $(V_i, V_{-i})$ space corresponding to the subpopulations ${T_i = c}$ and ${T_{-i} = c}$. Integrating the identified MCSEs and MCDEs over these regions, using the copula density of $(V_i, V_{-i})$ as weights, recovers the local average spillover and direct effects analyzed by vazquez2023causal:

equation*[equation* omitted — 772 chars of source]

where the copula density is identified in Corollary (ref), and Theorem (ref) provides identification of $\operatorname{MCSE}_i\left(0; v_0, v_1\right)$ and $\operatorname{MCDE}_i\left(0; v_0, v_1\right)$.

Other Policy Relevant Treatment Effect

In addition, the identified MCSEs and MCDEs can be used to recover the policy-relevant treatment effect (PRTE), which quantifies the expected change in outcomes induced by a policy-driven shift in the treatment selection mechanism. The PRTE aggregates the underlying marginal controlled effects across the distribution of unobserved heterogeneity, weighted by how the policy changes the propensity of treatment participation, following the interpretation of heckman2005structural and related work. This parameter provides a meaningful measure of the impact of counterfactual policy interventions in the presence of heterogeneous treatment effects, allowing for the evaluation of counterfactual interventions that modify the selection into treatment.

To formalize the link between marginal controlled effects and policy-relevant treatment effects, consider how the MCSEs and MCDEs characterize PRTEs arising from exogenous policy changes. Let $\mathcal{A}$ denote a set of feasible policies. For any policy $a \in \mathcal{A}$, I use superscript $a$ to denote the corresponding potential variables under that policy. For example, $D_i^a$ denotes the treatment status of unit $i$ that would be realized if policy $a$ were implemented. Thus, changing the policy from a to $a'$ induces changes in the distribution of instrumental variables, treatment selection, and outcomes, such as $Z_i^a \to Z_i^{a'}$, $D_i^a \to D_i^{a'}$, and $Y_i^a \to Y_i^{a'}$. These counterfactual changes form the basis for evaluating PRTEs.

Under policy $a$, the treatment decision is given by

equation*[equation* omitted — 78 chars of source]

where $P_i^a(Z_i^a, Z_{-i}^a) = \mathbb{P}(D_i^a = 1 \mid Z_i^a, Z_{-i}^a)$ denotes the policy-specific propensity score. Following the standard policy invariance assumption (as described in Assumption (ref)), the introduction of a new policy is assumed to affect only the treatment selection mechanism through changes in the propensity score, without altering the joint distribution of unobserved characteristics.

assumption(Policy invariance) The distribution of \begin{equation*} \big(U^a_i, U^a_{-i}, V^a_i, V^a_{-i}\big) \end{equation*} is invariant to any policy $a \in \mathcal{A}$.

The policy-relevant treatment effect (PRTE) measures the average change in outcomes induced by a policy intervention that modifies the treatment assignment mechanism, which is defined as

equation*[equation* omitted — 95 chars of source]

where $\Delta P$ denotes the proportion of groups in which at least one member changes treatment status as a result of the policy shift from $a$ to $a'$.

Given that each pair of propensity scores $(P_i^a, P_{-i}^a)$ partitions the support of $(V_i, V_{-i})$ into four regions associated with distinct treatment realizations $(D_i, D_{-i})$, the expected outcome under policy $a$ can, under Assumption (ref), be expressed as follows:

equation*[equation* omitted — 509 chars of source]

Because $\mathbb{E}[Y_i^a]$ is represented as a weighted average of the marginal treatment response functions, and the PRTE is defined as the difference in $\mathbb{E}[Y_i^a]$ across alternative policies, the PRTE naturally admits an interpretation as a weighted average of the MCDEs and MCSEs over particular regions of the latent heterogeneity space $(V_i, V_{-i})$. Hence, the identified MCDEs and MCSEs provide the key building blocks for constructing PRTEs associated with a broad class of counterfactual policy interventions. Appendix (ref) presents explicit expressions for the PRTE under several empirically relevant types of policy changes.

Comparing With Marginal Treatment Effect

This section introduces the connection between the marginal controlled effects and the standard marginal treatment effect (MTE) framework. The MCDEs and MCSEs are defined in a manner analogous to the MTE, capturing how potential outcomes vary with continuous unobserved heterogeneity. However, unlike the standard MTE that rules out spillovers, the MCDEs and MCSEs explicitly account for spillovers arising from peers' treatments as well as endogeneity in both own and peer treatment decisions. This section formally examines the relationship between the marginal controlled effects and the conventional MTE, demonstrating that the MCSEs and MCDEs naturally extend the MTE framework to settings with spillovers in outcomes and treatment selection within groups.

If spillover effects exist but are ignored and the standard MTE identification strategy is applied, the conventional estimand

equation*[equation* omitted — 210 chars of source]

fails to identify the true MTE, $\mathbb{E}[Y_i(1) - Y_i(0) \mid V_i = p_0]$.

In the presence of the spillover structure specified in Equation (ref), the conventional propensity score can be written as

equation*[equation* omitted — 462 chars of source]

where $\mathcal{Z}$ denotes the support of the peer's instrumental variable $Z_{-i}$, the second equality follows from the law of iterated expectations, and the last equality uses the independence assumption (Assumption (ref)) and the distributional normalization in Assumption (ref). This expression shows that the conventional propensity score is a weighted average of the unit's threshold function $h_i(z_0, z_1)$ over the peer's instrument $Z_{-i}$, conditional on $Z_i$.

Furthermore, Corollary (ref) demonstrates that, when spillovers are present, the conventional MTE identifier, $\partial \mathbb{E}[Y_i \mid P_i\left(Z_i\right)=p_0] / \partial p_0$ captures a weighted average of the MCDEs, augmented by a bias term arising from the dependence of the unit's treatment on the peer's instrument and from the correlation between $Z_i$ and $Z_{-i}$.

corollary(Breakdown of MTE causal interpretation) Consider the model with spillovers specified in Equation (ref), and suppose that the conditions in Theorem (ref) are satisfied. Under these assumptions, the conventional MTE identifier identifies \begin{equation*} \begin{aligned} \frac{\partial \mathbb{E}[Y_i \mid P_i\left(Z_i\right)=p_0]}{\partial p_0} = &\int_{0}^{1} \int_{0}^{p_1} \operatorname{MCDE}_i(1; p_0, v_1) c_{V_i, V_{-i}}(p_0, v_1) f_{P_{-i} \mid P_i = p_0} (p_1) dv_1 dp_1 \\ +&\int_{0}^{1} \int_{p_1}^{1} \operatorname{MCDE}_i(1; p_0, v_1) c_{V_i, V_{-i}}(p_0, v_1) f_{P_{-i} \mid P_i = p_0} (p_1) dv_1 dp_1 \\ &+ \mathcal{R}. \end{aligned} \end{equation*} The bias term $\mathcal{R}$ is generally nonzero. It vanishes under two sufficient conditions: (i) the unit's treatment decision is unaffected by the peer's instrument, and (ii) the instruments $Z_i$ and $Z_{-i}$ are mutually independent within each group. The explicit form of $\mathcal{R}$ is provided in Appendix (ref). \begin{proof} See Appendix (ref). \end{proof}

If spillover effects are absent from both the outcome and the treatment selection processes, the identification results collapse to the standard MTE framework, as shown in Corollary (ref).

corollarySuppose that no spillover effects are present, and the SUTVA holds. Specifically, assume that $V_i \perp \!\!\! \perp V_{-i}$, $Y_i(D_i, d) = Y_i(D_i, d') \equiv Y_i(D_i)$, and $h_i(Z_i, z) = h_i(Z_i, z') \equiv h_i(Z_i)$. Under the conditions of Theorem (ref), the identification results for the MCSEs and MCDEs collapse to the standard MTE framework. In particular: \begin{enumerate} • The propensity score identifies the standard threshold function $h_i(Z_i)$, \begin{equation*} \mathbb{P}\big(D_i = 1 \mid Z_i = z_0, Z_{-i} = z_1) = h_i(z_0), \end{equation*} indicating that the treatment decision of unit $i$ is unaffected by the peer's instrument. • The copula density simplifies to independence, $c_{V_i, V_{-i}}(p_0, p_1) = 1$, suggesting that $V_i \perp \!\!\! \perp V_{-i}$. • The marginal controlled spillover effect is zero for all $(p_0, p_1)$ in the interior of $\mathcal{P}$, $\operatorname{MCSE}_i^{(d)}(p_0, p_1) = 0$, implying that the outcome does not depend on the peer's treatment. • The marginal controlled direct effect reduces to the MTE, $\operatorname{MCDE}_i^{(d)}(p_0, p_1) = \mathbb{E}[Y_i(1) - Y_i(0) \mid V_i = p_0]$, for all $(p_0, p_1)$ in the interior of $\mathcal{P}$. \end{enumerate}
proofSee Appendix (ref).

To conclude, the standard MTE may lose its causal interpretation when spillovers are present, whereas the MCSEs and MCDEs coincide with the MTE under SUTVA. Hence, the framework developed in this paper provides a natural extension of the standard MTE framework, generalizing it to settings with spillovers in both outcomes and treatment selection.

remark(Connection to the multivalued treatment model) If the group is treated as a single decision-making unit and the treatment vector $(D_{ig}, D_{-ig}) \in \{0,1\}^2$ can be reformulated as a multivalued group-level treatment $D_g \in \{0, 1, 2, 3\}$, the framework can be viewed as a multivalued MTE models, such as lee2018identifying. Translating their setup into the spillover context, their identification relies on an exclusion restriction requiring that unit $i$'s threshold function does not depend on peers' instruments. In contrast, the treatment selection structure developed in this paper enables point identification of the threshold functions without imposing such exclusion restrictions and accommodates settings where spillovers arise from peers' instruments. Hence, the two frameworks are not nested. A more detailed comparison with the multivalued treatment framework is provided in Appendix (ref).

Testable Implications of Identifying Assumptions

The imposed spillover model structure and assumptions yield two sets of testable implications.

The first set arises from the fact that the cross-partial derivatives of the observed conditional expectations,

equation*[equation* omitted — 208 chars of source]

identify the joint distribution of potential outcomes weighted by the copula density of unobservables, where $A_1, A_2$ denote arbitrary Borel sets in the outcome support. Since both the conditional probabilities and the copula density are nonnegative, these derivatives must be weakly positive, generating a set of nesting inequalities that serve as testable restrictions.

The second set of implications follows from the index sufficiency property: the marginal treatment response functions depend only on the values of the propensity scores, not directly on the realizations of the instrumental variables. Hence, for any two instrument pairs $(z_0, z_1)$ and $(\tilde{z}_0, \tilde{z}_1)$ that yield identical propensity scores $(P_i, P_{-i})$, the conditional expectation

equation*[equation* omitted — 199 chars of source]

should remain invariant across the two sets of instruments, as they correspond to the same marginal treatment response values.

Corollary (ref) formally states the nesting inequality and index sufficiency conditions implied by the model.

corollary(Testable implications) The following conditions constitute the testable implications of Assumptions (ref)-(ref), given the model structure specified in Equation (ref). \begin{enumerate} • (Nesting inequalities) For any Borel set $A_1, A_2 \subseteq \mathcal{Y}$, $d \in \{0,1\}$, and $(p_0, p_1) \in \mathcal{P}$, the following inequality should hold: \begin{equation*} \begin{aligned} &\frac{\partial^2}{\partial p_1 \partial p_0} \mathbb{E}\Big[\mathbbm{1} \big\{Y_i \in A_1, Y_{-i} \in A_2 \big\} \mathbbm{1}\big\{D_{i} = d, D_{-i} = d \big\} \mid P_{i} = p_0, P_{-i}= p_1 \Big] \geq 0, \\ &-\frac{\partial^2}{\partial p_1 \partial p_0} \mathbb{E}\Big[\mathbbm{1} \big\{Y_i \in A_1, Y_{-i} \in A_2 \big\} \mathbbm{1} \big\{D_{i} = d, D_{-i} = 1-d \big\} \mid P_{i} = p_0, P_{-i} = p_1\Big] \geq 0. \end{aligned} \end{equation*} • (Index sufficiency) For any Borel set $A_1, A_2 \subseteq \mathcal{Y}$, $d_0, d_1 \in \{0,1\}$, and different values of instruments, $(z_0, z_1) \neq (\tilde{z}_0, \tilde{z}_1)$, that yield the same propensity score values such that $P_i(z_0, z_1) = P_i(\tilde{z}_0, \tilde{z}_1) = p_0$ and $P_{-i}(z_0, z_1) = P_{-i}(\tilde{z}_0, \tilde{z}_1) = p_1$, the following equalities should hold: \begin{equation*} \begin{aligned} & \mathbb{E}\big[\mathbbm{1}\big\{Y_i \in A_1, Y_{-i} \in A_2 \big\} \mathbbm{1}\{D_i = d_0, D_{-i} = d_1\} \mid Z_i = z_0, Z_{-i} = z_1\big] \\ =& \mathbb{E}\big[\mathbbm{1}\big\{Y_i \in A_1, Y_{-i} \in A_2 \big\} \mathbbm{1}\{D_i = d_0, D_{-i} = d_1\} \mid Z_i = \tilde{z}_0, Z_{-i} = \tilde{z}_1\big]. \end{aligned} \end{equation*} \end{enumerate}
proofSee Appendix (ref) for the proof of the validity of these testable implications.

Existing literature, such as carr2021testing, has developed methods that could be implemented to test the conditions in Proposition (ref). A formal application of these testing procedures to the present framework is left for future research.

Estimation and Inference

Semiparametric estimation and inference

Estimation procedure

Consider a random sample of $G$ groups, each consisting of n units. For each group $g = 1, \ldots, G$, the observed data

equation*[equation* omitted — 162 chars of source]

are independently and identically distributed across groups.

This section develops a semiparametric estimation procedure that extends the framework of carneiro2009estimating to settings with spillover effects in both treatment and outcome equations. The estimation section considers the covariate-augmented setting introduced in Appendix (ref). The model incorporating covariates can be expressed as

equation*[equation* omitted — 299 chars of source]

where outcomes depend on both own and peer covariates and treatments, and treatment decisions follow a threshold-crossing rule. It is assumed that covariates and instruments are randomly assigned at the group level,

equation*[equation* omitted — 123 chars of source]

where $W_{ig} \equiv (X_{ig}, Z_{ig})$.

Assumption (ref) is maintained throughout the estimation section.

assumption(Estimation assumptions) Assume that (i) $\mathbb{E}|Y_{ig}(d, d')| < \infty, d, d' \in \{0,1\}$. (ii) Propensity scores, $P_{ig}, i \in \{0,1\}$ are nondegenerate continuous random variable. (iii) The conditional expectations $\mathbb{E}[Y_{i d d' g} \mid X_g=\mathbf{x}, P_{ig}=p_0, P_{-ig}=p_1]$, $X_g \equiv (X_{ig}, X_{-ig})$, and $\mathbb{E}[D_{ig} D_{-ig} \mid P_{ig}=p_0, P_{-ig}=p_1]$ are assumed to be twice continuously differentiable with respect to $(p_0, p_1)$. (iv) $\partial^2 \mathbb{E}[D_{ig} D_{-ig} \mid P_{ig}=p_0, P_{-ig}=p_1] / \partial p_0 \partial p_1$ is bounded from above and away from zero.

To illustrate the estimation procedure, this section focuses on a simple case where each group consists of two units, i.e., $n = 2$ and $i \in \{0, 1\}$. The extension to group sizes $n > 2$ follows analogously. Building on the identification results with exogenous covariates presented in Appendix (ref), this section aims to estimate the marginal treatment response (MTR) functions,

equation[equation omitted — 616 chars of source]

where $\operatorname{sgn}(x)$ denotes the sign function that indicates the sign of a scalar $x$. The term ${2\operatorname{sgn}(1 - |d - d'|) - 1}$ evaluates to $1$ when $d = d'$ and to $-1$ when $d \neq d'$. The marginal controlled effects are then obtained by taking the difference between the estimated MTR functions.

The propensity score functions, $P_{0g} \equiv P_0(W_g)$ and $P_{1g} \equiv P_1(W_g)$, defined as $\mathbb{P}(D_{ig} = 1 \mid W_g)$ with $W_g \equiv (W_{0g}, W_{1g})$, are not directly observed and must be estimated from the data. The first stage of involves estimating the propensity score functions using a series regression approach. To alleviate the curse of dimensionality, a partially linear additive specification is employed:

equation[equation omitted — 254 chars of source]

The covariate vector $w$ includes both continuous and discrete components, denoted by $w = (w^{cts}, w^{disc})$, where $w^{cts} = (w^{cts}_1, \cdots, w^{cts}_{\ell_1})$ is an $\ell_1-$dimensional vector of continuous random variables, and $w^{disc} = (w^{disc}_1, \cdots, w^{disc}_{\ell_2})$ is an $\ell_2-$dimensional vector of discrete random variables.

To preserve flexibility, no parametric restrictions are imposed on the unknown smooth functions $\varphi_1, \ldots, \varphi_{\ell_1}$ associated with the continuous covariates, while the coefficients $\vartheta_1, \ldots, \vartheta_{\ell_2}$ on the discrete covariates remain to be estimated.

Series estimation relies on constructing a basis for smooth functions defined on $\mathbb{R}$, denoted as $\{p_k : k = 1, 2, \dots \}$, such that each continuous function $\varphi_{\ell}$, for $\ell = 1, \dots, \ell_1$, can be approximated arbitrarily well by a linear combination of these basis functions as $k \rightarrow \infty$. Commonly used basis functions include polynomial basis functions, splines, and wavelets. Given a positive integer $\kappa$, define the regressor vector

equation*[equation* omitted — 253 chars of source]

where the first $\kappa \times \ell_1$ components correspond to basis function expansions of the continuous covariates, and the remaining $\ell_2$ components include the discrete covariates in their original form.

The series estimator of the conditional probability $\mathbb{P}(D_{ig} = 1 \mid W_g)$, for $i \in \{0, 1\}$, is given by

equation*[equation* omitted — 103 chars of source]

where $\hat{\theta}_{\kappa}^i$ is obtained from the least squares optimization problem

equation*[equation* omitted — 196 chars of source]

and $\tilde{\kappa} = \kappa \ell_1 + \ell_2$ denotes the total dimension of the regressor vector $P_\kappa(W_g)$.

remark(Series with Lasso regression) To increase flexibility in selecting relevant basis terms, one may combine nonparametric series estimation with Lasso regression, which enables automatic selection and regularization of basis terms. The estimation errors of such estimators have been studied in bickel2009simultaneous and other related works cited therein. Select a positive integer $\kappa$ such that $\kappa \gg g$. Then, the series estimator, $\tilde{P}_i\left(W_g\right)$, with $l_1$-penalization is derived as \begin{equation*} \begin{aligned} &\hat{\theta}_{\kappa}^i = \arg \min_{\theta_\kappa^i \in \mathbb{R}^{\widetilde{\kappa}}} \frac{1}{G} \sum_{g = 1}^{G} \bigg(D_{ig} - P_\kappa(W_g)' \theta_\kappa^i \bigg)^2 + 2 \lambda \frac{1}{\widetilde{\kappa}} \sum_{j=1}^{\widetilde{\kappa}} \lVert p_j \rVert _G \lvert \theta_j^i \rvert, \\ & \tilde{P}_i\left(W_g\right)=P_\kappa\left(W_g\right)^{\prime} \hat{\theta}_{\kappa}^i, \end{aligned} \end{equation*} where $\lambda > 0$ is the tuning constant, $p_j(\cdot)$ denotes the $j$-th component of the basis expansion $P_\kappa(\cdot)$, and $\lVert \cdot \rVert _G$ stands for the empirical norm, $\lVert p_j \rVert _G = \sqrt{1 / G \sum_{g=1}^G p_j^2(W_g)}$. In the estimation procedure, the tuning parameter $\lambda$ is selected using cross-validation.

A finite-sample concern is that the estimated series approximation $\tilde{P}_i(W_g)$ may take values outside the admissible unit interval $[0,1]$. To address this, a standard trimming adjustment can be applied. The trimmed estimator is defined as

equation*[equation* omitted — 344 chars of source]

where $\delta > 0$ is a small constant chosen by the researcher. The resulting $\hat{P}_i(W_g)$ thus provides a feasible and bounded estimator of the propensity score $\mathbb{P}(D_{ig} = 1 \mid W_g)$, which is denoted compactly as $\hat{P}_{ig}$ in subsequent analysis.

The next step involves estimating the cross-partial derivatives appearing in both the numerator and denominator of the estimand in Equation (ref). The procedure begins with estimating the denominator of the estimand, namely the cross-partial derivative

equation*[equation* omitted — 120 chars of source]

Local polynomial regression provides a flexible and well-established approach for estimating such derivatives of conditional expectations. Following fan1996local, the polynomial order is set to $p = d + 1$, where $d$ denotes the derivative order. Since the object of interest is a second-order derivative, a local cubic regression ($p = 3$) is employed:

equation[equation omitted — 534 chars of source]

where $K(\cdot)$ denotes the kernel function and $h_{G1}$ is the chosen bandwidth parameter. Bandwidths for kernel-based regressions can be selected using $K$-fold cross-validation. The estimated coefficient $\hat{b}_4(p_0, p_1)$ from Equation (ref) serves as an estimator of the cross-partial derivative $\partial^2 \mathbb{E}[D_{0g} D_{1g} \mid P_{0g}=p_0, P_{1g}=p_1] / \partial p_0 \partial p_1$.

The subsequent stage focuses on estimating the cross-partial derivative that appears in the numerator of the estimand,

equation*[equation* omitted — 144 chars of source]

Since the covariate vector $X_g$ may be multidimensional, the estimation adopts the semiparametric framework to mitigate the curse of dimensionality.

assumption(Partial linear outcomes) Potential outcomes satisfy a partially linear structure of the form \begin{equation*} Y_{i g}(\mathbf{x}, d, d')= \mathbf{x}'\beta_{idd'} + U_{i g}(d, d'), \end{equation*} where $\beta_{idd'}$ is a finite-dimensional parameter vector that may vary across units $i \in \{0, 1\}$ and treatment states $(d,d') \in \{0, 1\}^2$, and $U_{ig}(d,d')$ is an unrestricted nonparametric component capturing the remaining heterogeneity. The potential outcome is generated according to \begin{equation*} Y_{i g}(\mathbf{x}, d, d') \equiv m_i(\mathbf{x}, d, d', U_{ig}, U_{-ig}) \end{equation*} where the covariates are fixed at $X_g = \mathbf{x}$ and the treatment assignments are fixed at $(D_{ig}, D_{-ig}) = (d, d')$ exogenously.

Under this specification, the conditional expectation, and consequently the marginal treatment response (MTR) function, can be expressed as a semiparametric function of $(\mathbf{x}, p_0, p_1)$, separating the parametric effect of covariates from the nonparametric dependence on the propensity scores:

equation[equation omitted — 783 chars of source]

The conditional expectation $\mathbb{E}\left[Y_{ig} \mid D_{0g} = d, D_{1g} = d', P_{0g}, P_{1g}\right]$ can be expressed as

equation[equation omitted — 379 chars of source]

Therefore, conditional on the subsample with $\{D_{0g} = d, D_{1g} = d'\}$, the coefficient vector $\beta_{idd'}$ can be estimated by the least squares regression as

equation[equation omitted — 484 chars of source]

where $\hat{E}_h[\cdot \mid \cdot]$ represents a kernel regression estimator with selected bandwidth $h$.

The residual then follows as

equation*[equation* omitted — 155 chars of source]

which serves as an estimator of the unobserved component $U_{i d d' g}$.

In the last step, use the sample

equation*[equation* omitted — 115 chars of source]

to estimate the cross-partial derivative $\partial^2 \mathbb{E}[U_{idd'g} \mid P_{0g}=p_0, P_{1g}=p_1] / \partial p_0 \partial p_1 $ through a local polynomial regression of order three,

equation*[equation* omitted — 515 chars of source]

The resulting coefficien $\hat{c}_4(d, d'; p_0, p_1)$ consistently estimates $\partial^2 \mathbb{E}[U_{idd'g} \mid P_{0g}=p_0, P_{1g}=p_1] / \partial p_0 \partial p_1 $.

Finally, the marginal treatment response functions $m_{i g}^{(\mathbf{x}, d, d')}(p_0, p_1)$ are estimated as

equation*[equation* omitted — 170 chars of source]

where $\widehat{c}_4(d, d'; p_0, p_1)$ and $\widehat{b}_4(p_0, p_1)$ are the local polynomial estimators of the cross-partial derivatives of $\mathbb{E}[U_{idd'g} \mid P_{0g}, P_{1g}]$ and $\mathbb{E}[D_{0g} D_{1g} \mid P_{0g}, P_{1g}]$, respectively.

The next section derives the asymptotic distribution of the estimated marginal treatment response functions, abstracting from the role of covariates to focus on the sampling behavior of the nonparametric components,

equation[equation omitted — 102 chars of source]

Asymptotic Properties

Uniform consistent rate of propensity score

Cubic spline basis functions $\{p_k: k = 1, 2, \cdots\}$ are employed to approximate the nonparametric components of the propensity score functions. The following assumptions, adapted from belloni2015some, provide the regularity conditions required to establish the uniform convergence rate of the series estimators for the propensity score functions.

assumption(Series estimation) (i) The eigenvalues of $\mathbb{E}[P_\kappa(w_g)P_\kappa(w_g)']$ are bounded above and away from zero uniformly over $G$. (ii) Each function $\varphi_{i} \in \mathcal{G}$ in Equation (ref), where $\mathcal{G}$ is a set of functions $f$ in Hölder classes with exponent $s$, $\Sigma_s(\mathcal{W})$, such that $\|f\|_s$ is bounded from above uniformly over $\mathcal{G}$. (iii) The support of $W^{cts}$ is known and is a Cartesian product of compact connected intervals on which $W^{cts}$ has a probability density function that is bounded away from zero.
lemma(Uniform rate of propensity score) Under Assumptions (ref)-(ref), we have \begin{equation*} \max _{g=1, \ldots, G}\left|\hat{P}_i\left(W_g\right)-P_i\left(W_g\right)\right|=O_p\left[\sqrt{\frac{\kappa \log \kappa}{G}}+\kappa^{-s}\right], \end{equation*} where $\kappa \rightarrow \infty$ as $G \rightarrow \infty$, $\kappa^{m/(m-2)}\operatorname{log}\kappa / G = O(1)$ for any $m > 2$, and $\kappa^{2-2 s} / G = O(1)$.

Asymptotic properties of cross-derivative estimators

As shown in Equation (ref), the proposed estimator is expressed as the ratio of two estimated cross-partial derivatives of conditional mean functions. This subsection derives the asymptotic properties these two cross-derivative estimators under the semiparametric estimation procedure described in Section (ref).

To derive the asymptotic properties of the cross-partial derivative estimators, impose the following assumption:

assumption(Local polynomial regression) (i) $\mathbb{E}[D_{0 g} D_{1 g} \mid P_{0 g}=p_0, P_{1 g}=p_1]$ and $\mathbb{E}[U_{i d d^{\prime} g} \mid P_{0 g}=p_0, P_{1 g}=p_1]$ are $(p+1)$-time continuously differentiable, $p \geq 3$. (ii) The conditional distributions of $D_{0 g} D_{1 g} \mid P_{0 g}, P_{1 g}$ and $U_{i d d^{\prime} g} \mid P_{0 g}, P_{1 g}$ are continuous at the point $(p_0, p_1)$. (iii) The kernel $K \in L_1$ is bounded with compact support, and $\|u\|^{4 p} K(u) \in L_1$, $\|u\|^{4 p + 2} K(u) \rightarrow 0$ as $\|u\| \rightarrow \infty$.

This assumption ensures sufficient smoothness of the underlying conditional mean functions and regularity of the kernel function, which together guarantee the validity of local polynomial approximations.

lemma(Convergence rates of cross-derivative estimators) Under Assumptions (ref)-(ref), the convergence rates of estimators $\hat{b}_4(p_0, p_1)$ and $\hat{c}_4(d, d'; p_0, p_1)$, as calculated in Section (ref), can be derived as \begin{equation*} \begin{aligned} &\hat{b}_4(p_0, p_1) - \frac{\partial^2}{\partial p_0 \partial p_1} \mathbb{E}\big[D_{0g}D_{1g} \mid P_{0g} = p_0, P_{1g} = p_1\big] \\ =& O_P\Big[(G h_{G1}^6)^{-1/2} + \max _{g: 1 \leq g \leq G}|\hat{P}_{0g}-P_{0g}| + \max _{g: 1 \leq g \leq G}|\hat{P}_{1g}-P_{1g}| + h_{G1}^4\Big], \\ &\hat{c}_4(d, d'; p_0, p_1) - \frac{\partial^2}{\partial p_0 \partial p_1} \mathbb{E}\big[U_{i d d' g} \mid P_{0g} = p_0, P_{1g} = p_1\big] \\ =& O_P\Big[(G h_{G2}^6)^{-1/2} + \max _{g: 1 \leq g \leq G}|\hat{P}_{0g}-P_{0g}| + \max _{g: 1 \leq g \leq G}|\hat{P}_{1g}-P_{1g}| + h_{G2}^4\Big], \end{aligned} \end{equation*} where $h_{G1}, h_{G2}$ are the bandwidths selected to estimate $\hat{b}_4(p_0, p_1)$ and $\hat{c}_4(d, d'; p_0, p_1)$.
proofSee Appendix (ref).

Asymptotic distribution of the marginal treatment response

This section characterizes the asymptotic distribution of the marginal treatment response functions, abstracting from covariate effects. The estimator, defined in Equation (ref), is constructed as the ratio of two estimated cross-partial derivatives of conditional mean functions. To establish the asymptotic properties of this estimator, the following assumptions are imposed.

assumption(Asymptotic distribution) (i) $\max _{g=1, \ldots, G}\big|\hat{P}_i(Z_{0 g}, Z_{1 g})-P_i(Z_{0 g}, Z_{1 g})\big|=o_p\big[(G h_{G 1}^6)^{-1 / 2}\big]$. (ii) $h_{G 1}, h_{G 2} \rightarrow 0, G h_{G 1}^6, G h_{G 2}^6 \rightarrow \infty$ as $G \rightarrow \infty$, $h_{G 2}=o(h_{G 1})$, $h_{G 1}, h_{G 2}=o(G^{-1 / 10})$.
theorem(Asymptotic distributions of the ratio estimator) Under Assumptions (ref)-(ref), the asymptotic distribution of $\hat{c}_4(d, d^{\prime} ; p_0, p_1) / \hat{c}_4(d, d^{\prime} ; p_0, p_1)$, $d, d' \in \{0,1\}$, can be characterized as \begin{equation*} \begin{aligned} & \big(G h_{G2}^6\big)^{1/2}\Bigg\{\frac{\hat{c}_4(d, d'; p_0, p_1)}{\hat{b}_4(p_0,p_1)} - \frac{c_4(d, d'; p_0, p_1)}{b_4(p_0,p_1)} \Bigg\} \\ \xrightarrow{d}& N\Bigg(0, \frac{\sigma^2(p_0, p_1)}{\big(b_4(p_0, p_1)\big)^2 f(p_0, p_1)} \big(M^{-1}\Gamma M^{-1}\big)_{5,5}\Bigg), \end{aligned} \end{equation*} where $\sigma^2(d, d'; p_0, p_1) = \text{Var}(U_{i d d' g} \mid P_{0g} = p_0, P_{1g} = p_1)$, $f(p_0, p_1)$ denotes the density of $(P_{0g}, P_{1g})$ evaluated at the point $(p_0, p_1)$, and $A_{5,5}$ denotes the element located in the fifth row and fifth column of a matrix $A$. The definitions of the matrices $M$ and $\Gamma$ are presented in the Appendix (ref).
proofSee Appendix (ref).

Finally, the asymptotic distributions of the MCSEs and MCDEs are derived. Their estimators are constructed using the estimated marginal treatment response functions:

equation*[equation* omitted — 409 chars of source]

Assuming that the differences $\hat{c}_4(d, d'; p_0, p_1) / \hat{b}_4(p_0,p_1) - c_4(d, d'; p_0, p_1) / b_4(p_0,p_1)$ are asymptotically independent across different values of $d, d' \in \{0,1\}$, the asymptotic distributions of $\widehat{\text{MCSE}}_i(\mathbf{x}, d; p_0, p_1)$ and $\widehat{\text{MCDE}}_i(\mathbf{x}, d; p_0, p_1)$ follow in Theorem (ref).

corollarySuppose that $\big(\hat{c}_4(d, d'; p_0, p_1) / \hat{b}_4(p_0,p_1) - c_4(d, d'; p_0, p_1) / b_4(p_0,p_1) \big)$ are asymptotically independent across different values of $d, d' \in \{0,1\}$, and that Assumptions (ref)–(ref) are satisfied. Then, the asymptotic distributions of $\widehat{\text{MCSE}}_i(\mathbf{x}, d; p_0, p_1)$ and $\widehat{\text{MCDE}}_i(\mathbf{x}, d; p_0, p_1)$ can be characterized as \begin{equation*} \begin{aligned} &\big(G h_{G2}^6\big)^{1/2}\Bigg\{\widehat{MCSE}(\mathbf{x}, d; p_0, p_1) - MCSE(\mathbf{x}, d; p_0, p_1)\Bigg\} \\ &\xrightarrow{d} N\Bigg(0, \frac{\sigma^2(1, d; p_0, p_1) + \sigma^2(0, d; p_0, p_1)}{\big(b_4(p_0, p_1)\big)^2 f(p_0, p_1)} \big(M^{-1}\Gamma M^{-1}\big)_{5,5}\Bigg), \\ &\big(G h_{G2}^6\big)^{1/2}\Bigg\{\widehat{MCDE}(\mathbf{x}, d; p_0, p_1) - MCDE(\mathbf{x}, d; p_0, p_1)\Bigg\} \\ &\xrightarrow{d} N\Bigg(0, \frac{\sigma^2(d, 1; p_0, p_1) + \sigma^2(d, 0; p_0, p_1)}{\big(b_4(p_0, p_1)\big)^2 f(p_0, p_1)} \big(M^{-1}\Gamma M^{-1}\big)_{5,5}\Bigg) \end{aligned} \end{equation*}

Parametric procedure

Parametric estimation

The estimators introduced in Section (ref) converge at nonparametric rates, requiring sufficiently large sample sizes to yield reliable estimates. Moreover, when the group size $n>2$, the conditional expectations in the estimand involve higher-dimensional conditioning, further slowing the rate of convergence. Consequently, in settings with limited sample sizes or large groups, it may be preferable to impose parametric assumptions to estimate parameters of interest. This section develops a parametric approach for estimation and inference procedure.

The estimation continues to rely on Assumption (ref), while introducing the following additional parametric assumptions.

assumptionThe following specifications are imposed in the parametric setting. \begin{enumerate} • (Propensity score) For the treatment assignment equation $D_{ig}=\mathbbm{1}\{\widetilde{V}_{ig} \leq h_i(W_g)\}$ of individual $i$ in group $g$, $i \in \{0, 1\}$, assume that $h_i(\cdot)$ is a $K_1$-th order polynomial function of $W_g = (W_{g,1}, \cdots, W_{g,\ell})' \in \mathbb{R}^\ell$: \begin{equation*} h_i(W_g) = \sum_{\boldsymbol{k} \in \mathcal{K}_{\ell,K_1}} \theta_{i\boldsymbol{k}} \cdot \prod_{j=1}^\ell W_{g,j}^{k_j}, \end{equation*} where $\boldsymbol{k} = (k_1, \ldots, k_\ell) \in \mathbb{N}_0^\ell$ is a multi-index, $\mathcal{K}_{\ell,K_1} = \{ \boldsymbol{k} \in \mathbb{N}_0^\ell : \sum_{j=1}^\ell k_j \leq K_1 \}$, and $(\theta_{i\boldsymbol{k}})_{\boldsymbol{k} \in \mathcal{K}_{\ell,K_1}} \equiv \theta_{i}$ are polynomial coefficients. Additionally, assume that the unobserved heterogeneity $\widetilde{V}_{ig}$ follows a standard normal distribution: $\widetilde{V}_{ig} \sim N(0, 1)$. • (Copula) Assume that the joint dependence structure of the unobserved heterogeneities $V_{0g}$ and $V_{1g}$ is characterized by a Gaussian copula with correlation parameter $\rho \in [-\varepsilon, \varepsilon]$, where $\varepsilon$ is a constant such that $0< \varepsilon <1$. Specifically, let $V_{ig} = \Phi(\widetilde{V}_{ig})$ for $i \in \{0,1\}$, where $\Phi(\cdot)$ denotes the standard normal cumulative distribution function. The copula of $(V_{0g}, V_{1g})$, denoted by $C_{V_{0g},V_{1g}}(\cdot,\cdot)$, is then given by the Gaussian copula with correlation $\rho$: \begin{equation*} C_{V_{0g}, V_{1g}}(v_0, v_1) = \Phi_\rho\big(\Phi^{-1}(v_0), \Phi^{-1}(v_1)\big), \forall (v_0, v_1) \in (0, 1)^2, \end{equation*} where $\Phi_{\rho}$ is the bivariate normal CDF with zero means, unit variances, and correlation $\rho$, $\Phi^{-1}$ denotes the inverse of the standard normal CDF, and $\rho$ is an unknown parameter. • (Marginal treatment response) It is assumed that the potential outcome follows a partially linear specification in the covariates, $Y_{i g}(\mathbf{x}, d, d')=\mathbf{x}' \beta_{i d d'}+U_{i g}(d, d')$, where $U_{i g}(d, d')$ satisfies the condition stated below: \begin{equation*} \begin{aligned} \mathbb{E}[U_{i g}(d, d') \mid V_{0g} = v_0, V_{1g} = v_1] =& \alpha_{i d d', 0} + \alpha_{i d d',1} \Phi^{-1}(v_0) \\ &+ \alpha_{i d d', 2} \Phi^{-1}(v_1) + \alpha_{i d d', 3} \Phi^{-1}(v_0) \Phi^{-1}(v_1), \end{aligned} \end{equation*} for all $(v_0, v_1) \in (0, 1)^2$, and $\alpha_{i d d'} \equiv (\alpha_{i d d', 0}, \alpha_{i d d',1}, \alpha_{i d d', 2}, \alpha_{i d d', 3})'$ denotes the vector of unknown coefficients that may be heterogeneous across individuals $i$ and treatment states $(d,d')$. \end{enumerate}

The imposed parametric assumptions are standard in the marginal treatment effect (MTE) literature and provide a tractable yet flexible framework for estimation and inference. Modeling the selection rule as \(D_{ig}=\mathbbm{1}\{\widetilde V_{ig}\le h_i(W_g)\}\) with a parametric index $h_i(W_g)$ and a standard normal unobserved term corresponds to the probit-type latent index widely used in practical implementations of the MTE framework (see, e.g., carneiro2011estimating; kline2016evaluating). The polynomial specification of $h_i(\cdot)$ provides sufficient flexibility to capture nonlinear relationships between instruments and covariates.

The Gaussian copula structure for $(V_{0g},V_{1g})$ is also a common parametric choice that facilitate likelihood-based estimation and allow dependence in unobserved heterogeneity across group members. Such assumptions have been adopted in the interference literature, including hoshino2023treatment, to capture correlated unobservables within the group.

The parametric specification of the MTR function is consistent with the functional-form assumptions commonly employed in the MTE literature to achieve point identification when the available instruments provide limited variation. Under SUTVA, brinch2017beyond show that imposing a parametric structure on the MTR function allows for the identification of heterogeneous treatment effects even with discrete instruments. Analogously, in the presence of spillovers, a similar approach can be applied by specifying the MTR function $\mathbb{E}[U_{ig}(d, d') \mid V_{ig} = v_0, V_{-ig} = v_1]$ as a polynomial expansion in the unobserved heterogeneities, $\Phi^{-1}(v_0)$ and $\Phi^{-1}(v_1)$. This formulation accommodates spillover effects from peers' treatments $d'$ and captures potential dependence between group members through involving $\Phi^{-1}(v_1)$. Moreover, the parametric formulation enables extrapolation beyond the observed support of the propensity scores, thereby allowing for the identification of policy-relevant treatment effects (PRTEs) even when instrumental variables exhibit limited or discrete variation brinch2017beyond.

The objective is to estimate the marginal treatment response function, $m_{i g}^{(\mathbf{x}, d, d')}(v_0, v_1) = \mathbb{E}[Y_{ig}(\mathbf{x}, d, d') \mid V_{0g} = v_0, V_{1g} = v_1]$, for any $\mathbf{x} \in \mathcal{X}$, $d, d' \in \{0, 1\}$, and $(v_0, v_1) \in (0, 1)^2$.

As in the semiparametric case, the first step involves estimating the propensity score functions $P_{0}(W_g), P_{1}(W_g)$, where $W_{i g}=(Z_{i g}, X_{i g}), W_g=(W_{0 g}, W_{1 g}) \in \mathbb{R}^\ell$. Under the first specification in Assumption (ref), and assuming that the instruments and covariates are independent of the unobserved heterogeneity $V_{ig}$, we can express the propensity score function as

equation*[equation* omitted — 199 chars of source]

The polynomial coefficients can be estimated using standard maximum likelihood methods:

equation*[equation* omitted — 449 chars of source]

Once the polynomial coefficients $\hat{\theta}_i$ are estimated, they can be substituted into $P_{ig}$ to obtain the estimated propensity score as

equation*[equation* omitted — 179 chars of source]

The next step is to estimate the joint dependence structure of $V_{0g}$ and $V_{1g}$. Under the second specification in Assumption (ref), this dependence is modeled by a Gaussian copula with correlation parameter $\rho$. Consequently, the second step of our procedure focuses on estimating $\rho$. The identification results imply the following equations,

equation*[equation* omitted — 635 chars of source]

Therefore, $\rho$ can be estimated using the maximum likelihood, substituting the first-stage estimates $\widehat{P}_{0g}$ and $\widehat{P}_{1g}$ for the true propensity scores $P_{0g}$ and $P_{1g}$,

equation*[equation* omitted — 728 chars of source]

The final step involves estimating the marginal treatment response $\mathbb{E}[Y_{ig}(\mathbf{x}, d, d') \mid V_{0g} = v_0, V_{1g} = v_1]$. Under the third specification in Assumption (ref), this function admits the following parametric representation,

equation*[equation* omitted — 386 chars of source]

Hence, the last stage of our procedure focuses on estimating the coefficient vectors $\beta_{i d d'}$ and $c_{i d d'}$. For illustration, consider the case $d = 1$ and $d' = 1$.

By combining the identification results with the third specification in Assumption (ref), the following relationship is obtained:

equation*[equation* omitted — 1,188 chars of source]

In the last line, $I_{11}^1(p_0, p_1, \rho), I_{11}^2(p_0, p_1, \rho)$, and $I_{11}^3(p_0, p_1, \rho)$ denote integrals that depend on $p_0, p_1$, and the correlation parameter $\rho$ of the Gaussian copula density $c_{V_{0g}, V_{1g}}(\cdot, \cdot)$ when $d = 1$ and $d' = 1$. Given that the propensity scores and the correlation have been estimated in the previous two steps, $\widehat{P}_{0g}$, $\widehat{P}_{1g}$, and $\hat{\rho}$ are substituted for their true values. The coefficient vectors $\alpha_{i11}$ and $\beta_{i11}$ are then estimated using the following least squares regression,

equation*[equation* omitted — 568 chars of source]

where $\widetilde{X}_{g} \equiv X_g \cdot \mathbb{P}(D_{0g}D_{1g} \mid \widehat{P}_{0g}, \widehat{P}_{1g})$. A similar procedure can be applied to estimate the coefficient vectors $\beta_{i d d'}$ and $\alpha_{i d d'}$ for other treatment combinations $(d, d')$. The estimated marginal treatment response function, $\widehat{m}_{i g}^{(\mathbf{x}, d, d')}(v_0, v_1)$, is obtained by substituting $(\hat{\alpha}_{i 11}', \hat{\beta}_{i 11}')'$ for $(\alpha_{i 11}', \beta_{i 11}')'$.

Finally, the MCDEs and MCSEs are obtained by taking differences of the estimated marginal treatment response functions,

equation*[equation* omitted — 409 chars of source]

Parametric asymptotic results

This section introduces a set of assumptions under which the parametric estimators are consistent.

assumption(Parametric first stage) In the first stage of propensity score estimation, for each individual $i \in \{0, 1\}$, we assume that \begin{enumerate} • $\theta_i \in \Theta_i$, where the parameter space $\Theta_i$ is compact. • The true parameter $\theta_{i0}$ is unique. • Let $l(\theta_i; D_{ig}, W_g)$ denote the log-likelihood of individual $i$'s treatment in group $g$: \begin{equation*} \begin{aligned} l(\theta_i; D_{ig}, W_g) =& D_{i g} \log \Phi\bigg(\sum_{k \in \mathcal{K}_{\ell, K_1}} \theta_{i k} \cdot \prod_{j=1}^{\ell} W_{g, j}^{k_j}\bigg) \\ &+(1-D_{i g}) \log \Bigg(1-\Phi\bigg(\sum_{k \in \mathcal{K}_{\ell, K_1}} \theta_{i k} \cdot \prod_{j=1}^{\ell} W_{g, j}^{k_j}\bigg)\Bigg). \end{aligned} \end{equation*} The log-likelihood function $l(\theta_i; D_{ig}, W_g)$ satisfies the following conditions: \begin{enumerate} • $\mathbb{E}\big[\sup_{\theta_i \in \Theta_i}|l(\theta_i; D_{ig}, W_g)|\big] < \infty$. • $\mathbb{E}\big[\nabla^2_{\theta_i} l(\theta_i; D_{ig}, W_g)\big]$ exists and is invertible. • $\mathbb{E}\big[\sup_{\theta_i \in \Theta_i}|| \nabla^2_{\theta_i} l(\theta_i; D_{ig}, W_g)||\big] < \infty$. \end{enumerate} \end{enumerate}

Under Assumptions (ref) and (ref), the estimator $\hat{\theta}_i$ obtained in the first stage is consistent.

lemmaSuppose Assumptions (ref), (ref) and (ref) hold. Then, for each $i \in \{0, 1\}$, the estimator $\hat{\theta}_i \overset{a.s.}{\to} \theta_{i0}$ as $G \rightarrow \infty$.
proofSee Appendix (ref).

To establish the consistency of the second-stage estimator of the Gaussian copula correlation parameter $\rho$, the following additional assumption is imposed.

assumption(Parametric second stage) In the second stage, to estimate the correlation parameter $\rho$ of the Gaussian copula, we impose the following conditions. \begin{enumerate} • The true parameter $\rho_0$ is unique. • Define the log-likelihood of joint treatments in group $g$ as \begin{equation*} \begin{aligned} &l(\rho, \theta; D_{g}, W_{g}) = \tilde{l}(\rho;D_g, P_g)\\ \equiv& D_{0g}D_{1g} \log\bigg(\Phi_\rho\big(\Phi^{-1}(P_{0g}), \Phi^{-1}(P_{1g})\big)\bigg) +\\ & D_{0g}(1 - D_{1g}) \log\bigg(P_{0g} - \Phi_\rho\big(\Phi^{-1}(P_{0g}), \Phi^{-1}(P_{1g})\big)\bigg) + \\ & (1 - D_{0g})D_{1g} \log\bigg(P_{1g} - \Phi_\rho\big(\Phi^{-1}(P_{0g}), \Phi^{-1}(P_{1g})\big)\bigg) + \\ & (1 - D_{0g})(1 - D_{1g}) \log\bigg(1 - P_{0g} - P_{1g} + \Phi_\rho\big(\Phi^{-1}(P_{0g}), \Phi^{-1}(P_{1g})\big)\bigg), \end{aligned} \end{equation*} where $P_g \equiv (P_{0g}, P_{1g})$, $P_{ig}$, $i \in \{0, 1\}$, is the function of the first stage parameter $\theta_i$ and the variable $W_g$, $\theta \equiv (\theta_0', \theta_1')'$, and $D_g \equiv (D_{0g}, D_{1g})$. The log-likelihood needs to satisfy \begin{enumerate} • $\mathbb{E}\big[\sup_{\rho \in [-\varepsilon, \varepsilon]} |l(\rho, \theta; D_{g}, W_{g})|\big] < \infty$. • There exists a function $L(\cdot)$ with $|L(D_g)| < \infty$ almost surely such that for all $d \in \{0, 1\}^2$, $(p, p') \in (0, 1)^2$, and $\rho \in [-\varepsilon, \varepsilon]$, $|\tilde{l}(\rho; d, p) - \tilde{l}(\rho; d, p')| \leq L(d) ||p - p'||$, where $\tilde{l}(\rho; d, p)$ is defined as the second stage log-likelihood given $(P_{0g}, P_{1g}) = p$. • $\mathbb{E}\big[\partial^2 l(\rho, \theta; D_{g}, W_{g}) / \partial \rho^2\big]$ is bounded away from zero. • $\mathbb{E}\big[\sup_{\rho \in [-\varepsilon, \varepsilon]}|\partial^2 l(\rho; D_g, P_g) / \partial \rho^2|\big] < \infty$ and $\mathbb{E}\big[\sup_{\theta \in \Theta_0 \times \Theta_1}||\nabla_{\theta} \frac{\partial l(\rho, \theta; D_{g}, W_{g})}{\partial \rho} ||\big] < \infty$. \end{enumerate} \end{enumerate}

These conditions allow us to establish the consistency of the second-stage estimator of $\rho$.

lemmaSuppose Assumptions (ref), (ref), (ref), and (ref) hold. Then, $\hat{\rho} \overset{a.s.}{\to} \rho_0$ as $G \rightarrow \infty$.
proofSee Appendix (ref).

Finally, the consistency of the estimated coefficient vector $(\hat{\alpha}_{i 11}', \hat{\beta}_{i 11}')'$ is established, which in turn ensures consistency of the estimated marginal controlled effects.

theorem(Consistency of parametric MCDEs and MCSEs) Suppose that Assumptions (ref), (ref), (ref), and (ref) hold. Also, assume that \begin{enumerate} • $X'_{Pdd'}X_{Pdd'}$ and $\mathbb{E}[X'_{Pdd'}X_{Pdd'}]$ are nonsingular, where $X_{Pdd'}$ is defined as a $G \times K$ matrix with the $g$-th row as \begin{equation*} X_{Pdd'_g} \equiv \big[\Phi_{\rho}(P_{0 g}, P_{1 g}), I_{dd'}^1(P_{0 g}, P_{1 g}, \rho), I_{dd'}^2(P_{0 g}, P_{1 g}, \rho), I_{dd'}^3(P_{0 g}, P_{1 g}, \rho), \widetilde{X}_g' \big], \end{equation*} where $P_{ig}$, $i \in \{0, 1\}$, is the function of the first stage parameter $\theta_i$ and the variable $W_g$. • $\|\widehat{X}_{Pdd'} - X_{Pdd'}\|_F^2/G \overset{a.s.}{\to} 0$, where $\widehat{X}_{Pdd'}$ is obtained by replacing the true values $P_{0g}$, $P_{1g}$, and $\rho$ in $X_{Pdd'}$ with their estimates $\widehat{P}_{0g}$, $\widehat{P}_{1g}$, and $\hat{\rho}$, respectively. The notation $\|\cdot\|_F$ denotes the Frobenius norm. • Set $\varepsilon_{idd'} = Y_{ig} \mathbbm{1} \big\{D_{0g} = d \big\} \mathbbm{1} \big\{D_{1g} = d' \big\} - X_{Pdd'_g} (\alpha'_{idd'}, \beta'_{idd'})'$. Then, $\text{Var}(\varepsilon_{idd'}) = \sigma^2_{idd'} < \infty$. • Define $\psi_{idd}(\alpha_{idd'}, \beta_{idd'}, \rho, \theta;Y_{ig}, D_{0g}, D_{1g}, W_g) = (Y_{id d'g} - X_{P_g} (\alpha'_{id d'}, \beta'_{idd'})')X_{P_g}$, where $Y_{i d d' g} \equiv Y_{i g} \mathbbm{1}\{D_{0 g}=d, D_{1 g}=d'\}$. It satisfies \begin{enumerate} • $\mathbb{E}[\sup_{\rho \in [-\varepsilon, \varepsilon]} ||\nabla_\rho \psi_{idd}(\alpha_{idd'}, \beta_{idd'}, \rho, \theta;Y_{ig}, D_{0g}, D_{1g}, W_g)||] < \infty$. • $\mathbb{E}[\sup_{\theta \in \Theta_0 \times \Theta_1} ||\nabla_\theta \psi_{idd}(\alpha_{idd'}, \beta_{idd'}, \rho, \theta;Y_{ig}, D_{0g}, D_{1g}, W_g)||] < \infty$. \end{enumerate} \end{enumerate} Under the above conditions, $(\hat{\alpha}_{i dd'}', \hat{\beta}_{i dd'}')' \overset{a.s.}{\to} (\alpha_{i dd'}', \beta_{i dd'}')'$ as $G \rightarrow \infty$.
proofSee Appendix (ref).

Inference regarding the estimator $(\hat{\alpha}_{i dd'}', \hat{\beta}_{i dd'}')'$ is performed using standard nonparametric bootstrap methodologies. Based on Assumptions (ref) and (ref), the conditions in Theorem (ref), and standard regularity conditions, the bootstrap distribution converges uniformly to the sampling distribution of the estimator\footnote{The regularity conditions include stochastic equicontinuity and a quadratic remainder condition. A formal proof is left for future work.} romano2012uniform. Therefore, standard nonparametric bootstrap methods, such as resampling the data and recomputing all stages, are expected to yield valid inference.

After establishing the consistency and the asymptotic distibutions of $(\hat{\alpha}_{i dd'}', \hat{\beta}_{i dd'}')'$, the consistency and asymptotic distributions of $\widehat{\text{MCSE}}_i(\mathbf{x}, d; p_0, p_1)$ and $\widehat{\text{MCDE}}_i(\mathbf{x}, d; p_0, p_1)$ follow directly from the continuous mapping theorem. This result obtains by the continuity of $\widehat{\text{MCSE}}_i$ and $\widehat{\text{MCDE}}_i$ as functions of $(\hat{\alpha}_{i dd'}', \hat{\beta}_{i dd'}')'$.

Simulation and Application

Parametric Simulation

This section presents a Monte Carlo simulation to assess the validity of the proposed parametric estimation methods.

For each Monte Carlo replication, I generate $G$ i.i.d. groups, where each group $g$ consists of two members indexed by $i \in \{0, 1\}$. I draw the group instrument vector, $Z_{g} = (Z_{0g}, Z_{1g})$, i.i.d. from a bivariate normal distribution $N(0, \Sigma_Z)$ with $\Sigma_Z = (1, 0.1; 0.1, 1)$. The correlation of $Z_{0g}$ and $Z_{1g}$ is not zero, since I allow the instruments of group members to be correlated. I also the group-level unobserved heterogeneity vector, $(\widetilde{V}_{0g}, \widetilde{V}_{1g})$, i.i.d. from a bivariate normal distribution $N(0, \Sigma_V)$ with $\Sigma_V = (1, 0.2; 0.2, 1)$ and independent of the instrument vector $Z_g$. By construction, $\widetilde{V}_{ig}$, $i \in \{0, 1\}$, follows a standard normal distribution. Additionally, the copula linking the normalized unobserved heterogeneity $V_{0g}$ and $V_{ig}$, where $\widetilde{V}_{ig} = \Phi(\widetilde{V}_i)$, is a Gaussian copula with correlation $\rho = 0.2$. These specifications are consistent with Assumption (ref) and the first two conditions in Assumption (ref).

I construct the following model to generate individual's treatment and potential outcome.

equation*[equation* omitted — 699 chars of source]

where the group-level disturbance $U_g \in \mathbb{R}$ is generated i.i.d. from the uniform distribution $\mathcal{U}(0, 1)$ and is independent of $(Z_{0g}, Z_{1g}, \tilde{V}_{0g}, \tilde{V}_{1g})$. The observed individual outcome $Y_{ig}$ is derived from

equation*[equation* omitted — 211 chars of source]

In our data generating process, the instrument vector $Z_g$ is independent of the unobserved heterogeneities and potential outcomes, $(\widetilde{V}_{ig}, U_g)_{i, d, d' \in \{0,1\}}$, satisfying Assumption (ref). Moreover, $Z_g$ does not directly affect the outcome $Y_{ig}(d, d')$, in accordance with Assumption (ref). The threshold function $h_i(\cdot)$ in the treatment assignment equation is specified as a first-order polynomial in the instrument, satisfying the first condition in Assumption (ref). For the potential outcomes, their conditional means given $V_{0g}$ and $V_{1g}$ satisfy the third condition in Assumption (ref). Therefore, the data generating process satisfies all identification and parametric assumptions.

I apply the method in Section (ref) to estimate and construct $95\%$ confidence intervals for the marginal controlled spillover (MCSE) and direct effects (MCDE) at selected evaluated points $(p_0, p_1)$. In the final step of computation, directly evaluating the integrals $I^j_{dd'}$, $j = 0, 1, 2, 3, 4$, at each estimated $(\widehat{P}_{0g}, \widehat{P}_{1g})$ is analytically intractable. To address this, I approximate the integrals using numerical integration. Specifically, I employ the Gauss-Hermite quadrature method, which I have verified to be both accurate and computationally efficient.

I arbitrarily select the following evaluation points,

equation*[equation* omitted — 93 chars of source]

for which the true MCSEs and MCDEs can be readily computed. I conduct $500$ Monte Carlo replications for each of four sample sizes, $G = 1000, 3000, 5000, 10000$. Table (ref) reports the coverage rates for the MCSEs, MCDEs, and the correlation parameter $\rho$.

table[table omitted — 2,322 chars of source]

For the MCSEs and MCDEs, when the sample size is $G = 1000$, the coverage rates are already close to, but slightly above, 95% for most parameters. As the sample size increases to $G = 3000$, the coverage rates decrease slightly yet remain close to 95%, with a few parameters falling just below this threshold. For larger sample sizes, the coverage rates for all parameters stabilize around 95%. The coverage rates for $\rho$ are also close to 95% across all sample sizes. These simulation results support the validity of our identification strategy and parametric estimation methods.

Application: Returns to education in best-friend relationships

In the empirical analysis, I estimate the direct and spillover effects of returns to education among the best-friend groups. I use data from the National Longitudinal Study of Adolescent to Adult Health (Add Health)\footnote{This research uses data from Add Health, funded by grant P01 HD31921 (Harris) from the Eunice Kennedy Shriver National Institute of Child Health and Human Development (NICHD), with cooperative funding from 23 other federal agencies and foundations. Add Health is currently directed by Robert A. Hummer and funded by the National Institute on Aging cooperative agreements U01 AG071448 (Hummer) and U01AG071450 (Hummer and Aiello) at the University of North Carolina at Chapel Hill. Add Health was designed by J. Richard Udry, Peter S. Bearman, and Kathleen Mullan Harris at the University of North Carolina at Chapel Hill. No direct support was received from grant P01 HD31921 for this analysis.}, a nationally representative longitudinal survey that follows a cohort of U.S. adolescents from grades 7-12 (1994-95 school year) into adulthood. The dataset contains rich information on respondents' family background and detailed friendship networks during adolescence, as well as education attainment and income in adulthood. This unique combination of longitudinal social, demographic, and economic data makes Add Health well suited for studying the long-term effects of adolescent friendships.

The Add Health dataset collects detailed friendship information during adolescence in both the in-home and in-school components of the Wave I survey. In each component, respondents are asked to list up to five male and five female friends, ranked from best to fifth best. I construct best-friend groups, each consisting of two respondents, by matching individuals who mutually nominate each other as their best friend. Following card2013peer, I first identify best-friend pairs from the Wave I in-home interviews. I then match any remaining mutually nominated best-friend pairs from the Wave I in-school interviews. Because respondents can nominate the best friend of each gender, I prioritize opposite-gender pairs: if a respondent appears in two different best-friend groups, I retain the group consisting of opposite-gender best friends.

The relationship between an individual's own education and their income has been extensively studied in the economics literature. In contrast, relatively little attention has been paid to how a best friend's education attainment influences an individual's earnings. Such an effect may operate through two competing channels.

In this empirical study, I investigate the effect of a best friend's education attainment on an individual's earnings and assess which channel, information sharing or competition, plays the dominant role within best-friend networks. Importantly, our identification framework assumes that spillover effects occur only within the same network and do not extend across different networks. In the context of returns to education, this implies that any effect of another person's education is restricted to the identified best friend, with no cross-pair spillovers. I take the total personal yearly pre-tax income from the Wave III in-home survey and apply a natural logarithm transformation to construct the outcome variable $Y$. The binary treatment variable $D$ is set to 1 if the individual has completed at least 16 years of education and 0 otherwise. I include the age, gender, race, health status, and family income as the controlled covariates $X$. I assume that, conditional on the observed covariates, the coefficients $(\alpha'_{idd'}, \beta'_{idd'})'$ in the potential outcome equations are identical for individuals $i \in \{0, 1\}$ within the same group.

For the continuous instruments $Z$, I construct measures based on the average parental education level of the individual's non-best friends, defined as all listed friends who are not ranked as the best friend. The average parental education of non-best friends may influence an individual's education attainment through channels such as shaping aspirations, fostering self-confidence, or behavioral sharing cools2019girls. Furthermore, conditional on covariates capturing demographic and socioeconomic characteristics, the family background of non-best friends is plausibly independent of the individual's unobserved heterogeneity. This is because weaker social ties, such as those with non-best friends, are less likely to exhibit the strong peer spillovers characteristic of best-friend relationships, and any residual correlation in unobservables is unlikely to persist once observed similarities are controlled for. Therefore, Assumption (ref) is likely to hold in this context.

The average parental education level of non-best friends during adolescence is unlikely to have a direct effect on an individual's yearly income in adulthood. This is because weaker social ties, such as those with non-best friends, generally lack the sustained and intensive interactions needed to shape long-term labor market outcomes. Unlike best friends, non-best friends are less likely to share close personal networks, exchange detailed career information, or provide direct referrals in the labor market. Moreover, by adulthood, many of these weaker ties from adolescence are no longer active, further limiting the scope for any direct influence on earnings. Therefore, any effect of non-best friends' parental education on the individual's income is likely to operate indirectly through its influence on the individual's own education attainment, rather than through direct channels. Hence, Assumption (ref) is plausibly satisfied in this setting.

I also require that the education decisions of best friends do not directly affect one another. This assumption is plausible because, while best friends may share aspirations or study habits, the final decision on how many years of education to pursue is typically determined by individual specific factors, such as academic ability, that are not directly changed by the best friend's decision. Therefore, any influence between best friends' education outcomes is more likely to operate indirectly through shared environments or information exchange, which aligns with the simultaneous incomplete information framework underlying our setting, rather than through direct strategic interaction in determining each other's years of schooling.

After excluding best-friend pairs in which both members have missing values for the treatment $D$ or the instrument $Z$, the sample comprises 1,019 best-friend pairs. Given the limited sample size, the parametric framework outlined in Section (ref) is applied under the parametric conditions specified in Assumption (ref). The estimated correlation between best friends' unobservables $V_{0g}$ and $V_{1g}$ is 0.36, indicating a positive dependence structure among unobservables within best-friend networks.

figure[figure omitted — 554 chars of source]

Figure (ref) plots the point estimates and the 90% and 95% confidence intervals of the marginal controlled spillover and direct effects by conditioning on the peer's unobservable $V_{-i} = 0.5$ and varing the value of individual's own unobservable. The covariates are fixed at their sample means. The blue solid line depicts the point estimates of the MCSEs and MCDEs, while the light and dark gray shaded areas represent the 95% and 90% confidence intervals, respectively. The red dotted line corresponds to the estimated parametric standard MTEs, which deviate substantially from the estimated MCSEs and MCDEs and lie outside their confidence intervals in most cases. This divergence provides empirical evidence of spillover effects between best friends, indicating that the standard MTE framework fails to retain a causal interpretation in the presence of such spillovers.

The results reveal substantial heterogeneity across these parameters. In particular, the estimates of the MCDEs with $d = 1$, which capture the direct effect of completing at least 16-year education given the best friend has completed at least 16 years, are positive and statistically significant at the 5% level across most values of the individual unobservable $V_i$. However, the estimates of MCDEs with $d = 0$, which measure the direct effect of completing at least 16 years of education given the best friend has not completed this level, are not statistically significant, even at the 10% level, across all values of the individual unobservable $V_i$. This discrepancy may reflect complementarities in human capital accumulation within best-friend pairs, consistent with the first channel discussed earlier: a highly educated best friend can provide valuable labor market information and opportunities that enhance the returns to one's own education. When both friends attain higher education, they may reinforce each other's labor market prospects through stronger professional networks, mutual encouragement in career development, or joint access to high-return opportunities. In contrast, when the best friend has lower education attainment, such reinforcing mechanisms may be absent, weakening the direct effect of one's own education on earnings.

Figure (ref) also presents the estimated MCSEs along with their confidence intervals. The MCSEs with $d = 1$, which capture the spillover effect of the best friend completing at least 16 years of education given the individual has completed 16 years, are significantly positive at the 5% level for some values of the individual unobservable $V_i$. In contrast, the MCSEs with $d = 0$, which measure the spillover effect of the best friend completing at least 16 years of education given the individual has not completed 16 years, are even significantly negative at the 10% level when $V_i$ is around 0.5 (approximately the value of the peer's unobservable $V_{-i}$), suggesting potential adverse spillover effects for some individuals. These patterns are consistent with the two channels through which a best friend's education attainment may affect an individual's earnings. The findings suggest that the information and opportunity channel dominates the competition channel when the individual is also highly educated, leading to positive and significant spillover effects. Conversely, when the individual has not completed 16 years of education, the competition channel appears to dominate, particularly among pairs with similar values of unobserved heterogeneity, resulting in negative estimated spillover effects. This asymmetry suggests complementarities in human capital and opportunity sharing among equally educated peers, and the potential for relative disadvantage when education attainment differs within a best-friend pair.

Extensions

The baseline framework can be generalized to accommodate various settings in which spillovers occur within predefined groups. First, point identification of the marginal controlled spillover and direct effects is established when outcomes depend on an exposure mapping function, rather than the full vector of group members' treatment statuses. Subsequently, Appendix (ref) extends the analysis to environments with continuous endogenous treatments, demonstrating that the marginal controlled spillover and direct effects remain point identified in such cases.

Exposure to functions of peers' treatments

Setting

In many applications, the predetermined groups within which spillovers occur can be large or vary in size. For example, when groups are defined at the level of schools, villages, or communities. In such cases, modeling outcomes as a function of the entire vector of group members' treatments may become infeasible. To address this issue, I instead adopt a framework in which the outcome of unit $i$ in group $g$, denoted $Y_{ig}$, depends on the unit's own treatment $D_{ig}$ and on a known function of the full vector of group treatments, denoted $H_g$. This function $H_g$ summarizes the group's effective treatment, consistent with the notion of an effective treatment in manski2013identification and the exposure mapping framework of aronow2017estimating. By reducing the dimensionality of peer treatments to an interpretable exposure measure, this approach allows for the analysis of spillovers in large or heterogeneous groups while maintaining tractable identification and interpretation.

I consider a sample of $G$ independent and identically distributed groups, indexed by $g = 1, \cdots, G$, where spillovers are restricted to occur within groups and not across them. Unlike the baseline framework, each group now consists of $n_g$ members, where the group size $n_g$ is allowed to vary across groups. To capture peer effects in this heterogeneous group size setting, I assume that the outcome of interest depends not on the full treatment vector but rather on a group-level exposure mapping, $H_g: \boldsymbol{D}_g \mapsto \mathbb{R}$, where $\boldsymbol{D}_g$ denotes the vector of individual treatment assignments within group $g$. This mapping $H_g$ is assumed to be continuous and correctly specified by the researcher. A common and tractable specification is the proportion of treated individuals in the group, given by $H_g = \sum_{i=1}^{n_g} D_{ig} / n_g$. Because treatment assignments $\boldsymbol{D}_g$ are observed, researchers can directly recover the realized values of $H_g$ for each group.

I specify the outcome for individual $i$ in group $g$ as $Y_{ig} = Y_{ig}(D_{ig}, H_g, U_{ig}, U_{-i,g})$, so that outcomes may depend on the individual's own treatment $D_{ig}$, the continuous group-level exposure $H_g$, and both the individual's unobserved characteristics $U_{ig}$ and those of her peers $U_{-i,g}$. I let $Y_{ig}(d, h)$ represents the potential outcome for unit $i$ given $D_{ig} = d$ and $H_g = h$.

I formulate the following equations to model spillovers that operate through the group-level exposure $H_g$. Throughout this section, I retain the subscript $g$ to distinguish individual level variables (indexed by $ig$) from group level variables (indexed by $g$), thereby clarifying how exposure-driven spillovers enter the model.

equation[equation omitted — 292 chars of source]

In Equation (ref), I formulate the outcome equation within the classical potential outcomes framework, specifying that each individual's outcome depends on her own binary treatment status, $D_{ig} \in \{0, 1\}$, as well as the group-level exposure $H_g$. A key feature of our framework is the recognition that both the individual treatment $D_{ig}$ and the group exposure $H_g$ may be endogenous. Specifically, $D_{ig}$ may correlate with unobserved individual-level characteristics that also influence the outcome. Likewise, the group exposure $H_g$, which is defined as a function of all group members' treatments, may depend on group-level unobservables that also affect the individual outcome.

To address the endogeneity of the individual treatment $D_{ig}$, I model it using a single-threshold crossing rule, analogous to the specification in the basic setting. Specifically, individual $i$ selects into treatment if the unobserved characteristic $V_{ig}$ falls below a threshold $h_i(Z_{ig}, Z_{-ig})$. The threshold function depends on the vector of instruments assigned to individual $i$, $Z_{ig}$, or additionally on the instruments assigned to other group members, $Z_{-ig}$. As before, I do not impose any functional form restrictions on the threshold function $h_i$ to preserve flexibility in how instruments affect treatment selection. In addition, the subscript $i$ allows for heterogeneity in threshold functions across individuals within the same group.

To account for the potential endogeneity of the group-level exposure $H_g$, I introduce a group instrument $Z_g$ and assume that $H_g$ follows a reduced-form relationship given by $H_g = m(Z_g, \varepsilon_g)$, where $\varepsilon_g \in \mathbb{R}$ represents an unobserved group-specific characteristic. The group instrument $Z_g$ may take various forms. For instance, it may correspond to the full vector of individual instruments $(Z_{ig})_{i \in \{1, \cdots, n_g\}}$, or to an aggregate statistic such as the average instrument level within the group. The random variable $\varepsilon_g$ captures latent group-level heterogeneity, potentially containing factors such as the group's social cohesion or the dependence structure among individual-level unobservables $(V_{ig})_{i \in \{1, \cdots, n_g\}}$. I impose no functional form restrictions on $m(\cdot)$ to maintain flexibility in the modeling of group exposure. In addition, I do not restrict the dependence structure between the individual unobservable $V_{ig}$ and the group-level unobservable $\varepsilon_g$, allowing for arbitrary correlation between individual- and group-level latent factors.

remark(Reduced function of $H_g$) To explain the reduced-form function of $H_g$, consider a scenario where the exposure function $H_g$ is defined as the average treatment level within group $g$, $H_g = \sum_{i=1}^{n_g} D_{ig} / n_g$, where $n_g$ denotes the number of members in group $g$, which may vary across groups. Let $\mathcal{I}_g$ represent the set of indices for individuals in group $g$. Assume that the group comprises two types of individuals: \begin{enumerate} • Type 1: Individuals indexed by $i \in \mathcal{I}_g^1 \subseteq \mathcal{I}_g$, which have unobserved individual unobservable $V_{ig} = \varepsilon_g$, $\varepsilon_g \in (0, 1)$. • Type 2: Individuals indexed by $j \in \mathcal{I}_g^2 = \mathcal{I}_g \setminus \mathcal{I}_g^1$, with individual unobservable $V_{jg} = 1 - \varepsilon_g$. \end{enumerate} Here, $\varepsilon_g \in (0, 1)$ captures unobserved heterogeneity at the group level, influencing the individual-level unobservables for both types. Additionally, assume that Type 1 individuals constitute an $\varepsilon_g$-proportion of the group, i.e., $|\mathcal{I}_g^1| / |\mathcal{I}_g| = \varepsilon_g$. Furthermore, suppose that individual treatment decisions depend on a group-level instrument $Z_g$, $D_{ig} = \mathbbm{1}\{V_{ig} \leq h(Z_g)\}$. Under these assumptions, the exposure $H_g$ can be expressed as an explicit function of $\varepsilon_g$ and $Z_g$, \begin{equation*} \begin{aligned} H_g =& \frac{1}{n_g}\sum_{i=1}^{n_g} D_{ig} = \frac{1}{n_g} \sum_{i \in \mathcal{I}_g^1} D_{ig} + \frac{1}{n_g} \sum_{i \in \mathcal{I}_g^2} D_{ig} \\ =& \varepsilon_g \mathbbm{1}\{\varepsilon_g \leq h(Z_g)\} + (1-\varepsilon_g) \mathbbm{1}\{1-\varepsilon_g \leq h(Z_g)\}. \end{aligned} \end{equation*} In this framework, $\varepsilon_g$ not only reflects the proportion of each individual type within the group but also captures the unobserved heterogeneity among different types of group members. More generally, the exposure level $H_g$ can be represented as an unknown reduced-form function of the group-level unobservable $\varepsilon_g$ and the instrument $Z_g$, where $\varepsilon_g$ can be interpreted as a scalar latent variable that fully summarizes the group-level unobserved heterogeneity relevant for determining exposure. Formally, I write $H_g = m(Z_g, \varepsilon_g)$, with $m(\cdot)$ left unspecified. One convenient interpretation is to view $m(z,e)$ as the quantile function of the conditional distribution of exposure, $Q_{H_g \mid Z_g=z}(e)$, and to define $\varepsilon_g$ as the corresponding conditional cumulative distribution function, $\varepsilon_g = F_{H_g \mid Z_g}(H_g)$.
remark(Random saturation framework) Our exposure-mapping framework can accommodate randomized saturation designs, such as those studied in ditraglia2023identifying, where each group is randomly assigned a saturation level—defined as the proportion of individuals offered treatment within the group. In this context, the group-level instrument $Z_g$ corresponds to the randomized saturation assignment for group $g$, while the group exposure $H_g$ reflects the average treatment take-up within the group. Unlike their approach, which requires one-sided noncompliance and an individualized offer response assumption, our methodology does not impose restrictions on the compliance behavior and allow the treatment take-up to depend on other group members' treatment assignment, thereby capturing richer forms of spillovers and strategic behavior. Crucially, our identification strategy is based on a reduced-form modeling of the group exposure and outcome equations. This approach eliminates the need to specify random coefficient models or to model the distribution of compliance types explicitly. Within our framework, the group-level unobservable $\varepsilon_g$ can be viewed as a scalar proxy for the fraction of compliers in the spirit of ditraglia2023identifying, while more broadly capturing latent heterogeneity in group responsiveness without imposing restrictive parametric assumptions.

In the exposure mapping framework, I redefine the marginal controlled spillover effect (MCSE) and the marginal controlled direct effect (MCDE) relative to the basic setting by explicitly conditioning on both the individual-specific unobservable $V_{ig}$ and the group-level unobservable $\varepsilon_g$, which accommodates heterogeneity at both the individual and group levels.

definition(MCSE and MCDE: Exposure mapping setting) Consider the model specified in Equation (ref). \begin{enumerate} • Fix the treatment of unit $i$ in group $g$ to be $D_{ig} = d$, $d \in \{0, 1\}$. I define the marginal controlled spillover effect (MCSE) of changing the exposure level from $h$ to $h'$, where $h,h' \in \mathbb{R}$, conditional on the individual unobservable $V_{ig} = p_0$ and the group-level unobservable $\varepsilon_g = p_1$, as \begin{equation*} \operatorname{MCSE}_{ig}(d, h', h; p_0, p_1) \equiv \mathbb{E}\big[Y_{ig}(d, h')-Y_{ig}(d, h) \mid V_{ig}=p_0, \varepsilon_g=p_1\big]. \end{equation*} • Fix the exposure level in group $g$ as $H_g = h$, $h \in \mathbb{R}$. I define the marginal controlled direct effect (MCDE) for individual $i$, conditional on the individual unobservable $V_{ig} = p_0$ and the group-level unobservable $\varepsilon_g = p_1$, as \begin{equation*} \operatorname{MCDE}_{ig}(h; p_0, p_1) \equiv \mathbb{E}\big[Y_{ig}(1, h)-Y_{ig}(0, h) \mid V_{ig}=p_0, \varepsilon_g=p_1\big] \end{equation*} \end{enumerate}

Identification

Identification in this setting is achieved under Assumptions (ref)-(ref).

assumption(Random assignment: Exposure mapping setting) I assume that the instrument vector is randomly assigned across groups, so that for any $g \in \{1, \cdots, G\}$, \begin{equation*} (Z_{ig}, Z_g)_{i \in \{1, \cdots, n_g\}} \perp \!\!\! \perp \big(Y_{ig}(d, h), V_{ig}, \varepsilon_g\big)_{d \in \{0, 1\}, h \in \mathbb{R}, i \in \{1, \cdots, n_g\}}. \end{equation*} Additionally, the instruments $(Z_{ig}, Z_g)_{i \in \{1, \cdots, n_g\}}$ satisfy an exclusion restriction in that they do not directly affect the outcome $Y_{ig}$.

Assumption (ref) requires that the vector of instruments assigned to individuals and the group must be randomly assigned at the group level, such that they are independent of all potential outcomes, as well as of both individual- and group-level unobservables. Moreover, under the model structure in Equation (ref), the instruments also satisfy the exclusion restriction, in the sense that they influence outcomes only through their effect on treatment take-ups and exposure, and do not directly enter the outcome equation.

assumption(Monotonicity of $m$) Given the instrument values $Z_g = z \in \mathbb{R}^k$, the function $m(z, e)$ is continuous and strictly monotonic in $e$.

The monotonicity condition in Assumption (ref) ensures that the group-level treatment $H_g$ is a one-to-one mapping of the group-level unobservable $\varepsilon_g$, conditional on the instruments. Specifically, for any given $Z_g = z$, the reduced-form relation $H_g \mid (Z_g = z) = m(z, \varepsilon_g)$ can be inverted with respect to $\varepsilon_g$, yielding $ \varepsilon_g \mid (Z_g = z) = m_z^{-1}(H_g) $. Thus, under the random assignment assumption (ref), I obtain the control function representation $\varepsilon_g = m_{Z_g}^{-1}(H_g)$. As established in goff2024testing, if the conditional distribution $F_{H_g \mid Z_g}$ is strictly increasing and continuous, then the monotonicity condition is not an additional structural restriction but instead follows directly from the reduced-form interpretation of the exposure mapping function discussed in Remark (ref).

I define the individual-level propensity score function for unit i in group g, consistent with previous definitions, as $P_{ig}(z) \equiv \mathbb{P}(D_{ig} = 1 \mid Z_{ig} = z)$. In addition, I define the group-level propensity score function for group $g$ as $P_g(z, h) \equiv \mathbb{P}(H_g \leq h \mid Z_g = z)$. I denote by $\mathcal{P}$ the support of the joint propensity score function $(P_{ig}, P_g)$. As shown in the Appendix (ref), the individual-level propensity score function identifies the individual threshold function $h_i(\cdot)$, while the group-level propensity score function identifies the inverse of the exposure mapping, $m_{Z_g}^{-1}(H_g)$. Taken together, these results imply that both individual- and group-level propensity score functions can be used as control functions, allowing us to account for unobserved heterogeneity in the outcome equation and thereby achieve identification of the marginal controlled spillover and direct effects.

theorem(Identifying MCSEs and MCDEs: Exposure mapping setting) Consider the model specified in Equation (ref). Suppose that Assumptions (ref), (ref), and (ref) hold. For $d \in \{0, 1\}$, $h \in \mathbb{R}$, and $(p_0, p_1) \in \mathcal{P}$, I impose the following additional conditions: (i) $\mathbb{E}[Y_{i g} D_{i g} \mid H_g=h, P_{i g}(Z_g)=p_0, P_g(Z_g, H_g)=p_1]$ and $\mathbb{P}(D_{ig} = 1 \mid H_g = h, P_{ig}(Z_g) = p_0, P_g(Z_g, H_g) = p_1)$ are differentiable with respect to $p_0$; (ii) the marginal treatment response functions \begin{equation*} m_{ig}^{(d, h)}\left(p_0, p_1\right) \equiv \mathbb{E}\left[Y_{ig}\left(d, h\right) \mid V_{ig}=p_0, \varepsilon_g=p_1\right] \end{equation*} are continuous; and (iii) the conditional density $f_{V_{ig} \mid \epsilon_g}$ is bounded from above and away from zero. Then, $m_{ig}^{(d, h)}(p_0, p_1)$ is identified as \begin{equation*} \begin{aligned} &\frac{\partial}{\partial p_0} \mathbb{E}\left[Y_{ig} \mathbbm{1} \{D_{ig} = d \} \mid H_g = h, P_{ig}\left(Z_g\right) = p_0, P_g\left(Z_g, H_g\right) = p_1\right] \Big/ \\ &\frac{\partial}{\partial p_0} \mathbb{P}\left(D_{ig} = d \mid H_g = h, P_{ig}(Z_g) = p_0, P_g(Z_g, H_g) = p_1\right). \end{aligned} \end{equation*} By taking appropriate differences of the marginal treatment response functions, I obtain identification of the marginal controlled spillover effect (MCSE) and the marginal controlled direct effect (MCDE).
proofSee Appendix (ref).

Conclusion

This paper develops a general framework for identifying causal effects in environments with within-group spillovers and endogenous treatment decisions. By relaxing the Stable Unit Treatment Value Assumption (SUTVA), the framework accommodates settings in which an individual's outcome depends not only on her own treatment but also on the treatment selection of group members. It further allows each individual's treatment decision to depend on instruments assigned to other group members.

The paper introduces two classes of causal parameters, he generalized local average controlled spillover and direct effects (LACSEs and LACDEs) and the marginal controlled spillover and direct effects (MCSEs and MCDEs), which extend the standard local average and marginal treatment effect frameworks to settings with spillovers. The LACSEs and LACDEs quantify peer and own treatment effects for specific subpopulations, while the MCSEs and MCDEs capture these effects conditional on continuous values of unobserved heterogeneity within groups. The paper formally establishes general conditions for the point identification of the LACSEs and LACDEs, characterizing the instrumental variation necessary for identification regardless of whether the instruments are discrete or continuous. It also shows that the MCSEs and MCDEs are nonparametrically point identified from continuous instrumental variation without imposing functional form restrictions on the outcome equation or on the joint distribution of unobserved characteristics within the group, thereby accommodating flexible forms of spillover structures.

These results extend existing approaches to causal inference with spillovers, showing that the MCSE-MCDE framework provides a natural generalization of the standard marginal treatment effect (MTE) model to spillover settings. Furthermore, the paper establishes that these marginal controlled effects serve as building blocks for policy-relevant treatment parameters (PRTEs), enabling the evaluation of both direct and spillover effects under counterfactual policy interventions.

For estimation and inference, the paper develops a semiparametric estimation strategy that builds on carneiro2009estimating, extending it to accommodate within-group spillovers while mitigating the curse of dimensionality associated with covariates. Asymptotic properties of the semiparametric estimators are derived, and a parametric estimation framework is proposed as a practical complement when sample sizes are limited or group sizes are large. Monte Carlo simulations demonstrate that the parametric estimators perform well in finite samples, providing accurate estimates and confidence interval coverage.

An empirical application using the National Longitudinal Study of Adolescent to Adult Health (Add Health) illustrates the framework's practical relevance. The analysis examines how education attainment affects long-term earnings within best-friend networks. The results indicate positive dependence between friends' unobserved characteristics and reveal heterogeneous spillover patterns, showing that both the magnitude and direction of peer influences vary with individuals' and their best friends' education attainment.

Finally, the paper extends the framework to settings with exposure mappings, where outcomes depend on a known function of group members' treatments rather than the full treatment vector. This generalization broadens the applicability of the framework to environments with varying or large group sizes, while preserving nonparametric point identification under continuous instrumental variation. The appendix further extends the analysis to settings with continuous endogenous treatments and establishes point identification results under continuous instruments.

Overall, the proposed framework offers an econometric foundation for identifying and estimating causal effects in the presence of within-group spillovers and endogenous treatments. It provides theoretical and practical tools for studying a wide range of social, education, and economic interactions. Future research could extend the framework by modeling endogenous group formation, and by developing more efficient semiparametric estimation methods to enhance finite-sample performance.

\onehalfspacing