Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,314 characters · 23 sections · 70 citation commands
Regression Discontinuity Design with Spillovers
Regression discontinuity design (RDD) is a popular method for causal inference and policy evaluation, particularly in settings where experimental manipulation is not possible. However, in the continuity-based framework of hahn2001identification, identification of treatment effect at the cutoff relies on the Stable Unit Treatment Value Assumption (SUTVA). This requires each unit’s potential outcomes to be unaffected by the treatment status of other units and may be unrealistic: in many economic environments, the variables of interest are equilibrium objects shaped by strategic interactions, making SUTVA violations natural. Although RDD is known to identify treatment effects local to the cutoff under SUTVA, this need not hold once SUTVA is violated.
Following manski1993identification, the literature on social interactions typically emphasizes two types of SUTVA violations: exogenous spillovers (also known as contextual effects) and endogenous spillovers. Exogenous spillovers occur when the outcome of an agent depends directly on the treatment status of their neighbors. On the other hand, endogenous spillovers arise when the outcome of an agent depends on the outcome of their neighbors. For a concrete example, consider gonzalez2021cell, which studies the effect of cell phone coverage on voter fraud during the 2009 Afghan presidential election. Their spatial RDD compares polling stations on either side of the boundary of cell phone coverage areas and finds that cell phone coverage reduced fraud. However, the author is also explicitly concerned about potential spillovers. For example, suppose coverage at a polling station reduces fraud by allowing voters to report suspicious behavior to the electoral commission. Voters at non-covered stations could potentially walk to covered stations to report fraud, giving rise to exogenous spillovers. Alternatively, corrupt politicians might strategically allocate fraud to areas with less monitoring. In this way, higher fraud in one polling station will reduce the need to commit fraud in neighboring polling stations. This is a form of endogenous spillovers. Because spillovers appear plausible in many settings, it is important to understand their effects on the estimand of RDD.
This paper studies RDD in the presence of spillovers. To do so, we extend the framework of hahn2001identification to incorporate exogenous and endogenous spillovers that occur along the running variable. We make two important modeling assumptions. Firstly, we assume that the outcome of a given unit depends linearly on the mean treatment status and mean outcome of their neighbors. The linear-in-means specification characterizes a large literature on peer effects (see e.g. manski1993identification, bramoulle2009identification, deGiorgi2010identification,goldsmith2013social,dePaula2024identifying) and spatial autoregression models (e.g. cliff1973spatial,kelejian1998generalized, kelejian2010specification,lee2004asymptotic, lee2007gmm). Secondly, we assume that the running variable which defines the RDD also determines the social interactions in the linear-in-means model. Specifically, units with similar values of the running variables are also assumed to be neighbors. When the running variable is geographical coordinates, this model captures interaction between units that are close in space. In the classic thistlethwaite1960regression, which studies the effect of scholarships on academic achievements, the running variable is test scores. Our model then captures the idea that students with similar test scores exert peer effects on one another, for example by forming study groups.
We make two main contributions. Firstly, we show that the estimand of RDD depends on the ratio of two terms: (1) the radius over which spillovers occur and (2) the bandwidth used for local linear regression. Specifically, RDD estimates direct treatment effect at the boundary when the radius is of larger order than the bandwidth. This is the effect on a unit at the boundary when we switch their treatment status from control to treatment, keeping treatment assignment fixed for all other units. When radius is of smaller order than the bandwidth, RDD instead estimates total treatment effect, which is the effect on the unit at the boundary when we switch the treatment status of the entire population from control into treatment. In the regime where radius is of similar order as the bandwidth, the RDD estimand is a mixture of the above effects and lacks a clear interpretation. We argue, however, that this is the more realistic regime wherein the asymptotic approximation captures more features of the finite sample distribution.
As such, our second contribution is a proposal for recovering direct and spillover effects in the intermediate regime. We do so by incorporating estimated spillover terms into local linear regression, in a procedure we call the local spillover regression (LSR). We show that LSR, which is essentially the local analog of peer effects regressions, leads to consistent estimators which are asymptotically normally distributed. We also adapt the method of armstrong2020simple to construct bias-aware confidence intervals for direct and spillover effects.
Simulation results show that LSR performs well relative to local linear regression with naively chosen bandwidths. The adaptive bandwidth rules of calonico2014robust and armstrong2020simple lead to estimators that are more robust to spillovers, but our method can achieve lower MSE and better coverage particularly in settings with intermediate amounts of spillovers. We apply LSR to the setting of gonzalez2021cell and find evidence of positive endogenous spillovers in electoral fraud. The implied social multiplier may be relevant for cost-benefit analysis for fraud deterrence policies and may provide motivation for network-based interventions. Along the way, we also discuss the estimands of the popular RD donut and highlight settings under which it recovers either direct or total treatment effects.
This paper contributes to the vast literature on RDD (see cattaneo2022regression for a recent review). Empirical researchers have long been concerned with spillovers in RDD, with papers such as jardim2022boundary arguing against the use of spatial discontinuity design for policy evaluation. However, to our knowledge, prior theoretical work on this issue is limited to aronow2017rdspill, which conducts their analysis under the local randomization framework of cattaneo2015randomization, finding that RDD always recovers a weighted average of direct treatment effect. We consider RDD under the continuity-based approach of hahn2001identification and find that the estimand can exhibit more complex behavior, particularly when neighborhoods are determined by the running variable. Since the first version of this paper was posted, torrione2024regression and borusyak2024rdaggregation have also studied RDD in the presence of spillovers. torrione2024regression takes the continuity-based approach, assuming that the running variable and the variable determining neighborhood structure have a continuous joint density. They find that RDD recovers a weighted average of direct treatment effect, drawing a connection to multi-score RDD. Our paper focuses on the case where neighborhoods are completely determined by the running variable and is unique in having spillovers that appear in the asymptotic approximation, provided that radius and bandwidth are scaled appropriately. This paper also differs from aronow2017rdspill and torrione2024regression in considering endogenous spillovers, which arises in many settings of interest to economists. Finally, we note that the aforementioned papers, as well as ours, focus on spillovers between units within an RD. borusyak2024rdaggregation studies the aggregation of multiple RDDs and considers effects of spillovers across designs.
This paper joins a burgeoning body of work that considers violations of SUTVA under various research designs, such as experiments (e.g. hudgens2008toward,aronow2017estimating,savje2021average, hu2022average, leung2022causal,li2022random, gao2023causal, vazquez2023experiment, auerbach2025local), differences-in-differences (e.g. clarke2017estimating,butts2021difference, xu2023difference), synthetic control (cao2019estimation), instrumental variables (sobel2006randomized, vazquez2023iv) and other observational settings (forastiere2020causal). RDD poses unique challenges relative to these other settings because it concerns parameters that are local to the cutoff and estimation is nonparametric. Additionally, we consider endogenous spillovers, which has received relatively less attention in this literature, though exceptions include, in the experiment context, leung2022causal,munro2021treatment, li2023experimenting, munro2023efficient and faridani2024linear. In the presence of endogenous spillovers, a unit's outcome may depend on the treatment status of the entire population, even if spillovers are assumed to have a small radius. The interaction of this dependence with the discontinuity of the observed outcome function poses novel technical challenges that we address.
The rest of the paper is organized as follows. Section (ref) describes our econometric framework. Section (ref) characterizes the estimand of local linear regression under spillovers. Section (ref) in particular considers the use of donut-hole RDs. Section (ref) presents the local spillovers regression, our proposal for recovering direct treatment and spillover effects. Section (ref) applies our method to the setting of gonzalez2021cell. Section (ref) concludes. Proofs are contained in the Appendix. Simulation results and the auxiliary lemmas used in the proofs are available in the Supplemental Appendix.
The remainder of this paper uses the following notation. We write $A_n \gg B_n$ if $A_n/B_n \to \infty$, $A_n \propto B_n$ if $A_n/B_n \to c$ where $0 < c < \infty $, and $A_n \ll B_n$ if $A_n/B_n \to 0$. Let $\boldsymbol{\iota}$ and $\boldsymbol{0}$ be constant functions that take values 1 and 0 on $[-1,1]$ respectively.
In this section, we introduce the econometric framework for studying RDD with spillovers. Section (ref) presents our model for potential outcomes. Section (ref) defines the parameters of interest. Section (ref) discusses our assumed neighborhood structure. Section (ref) presents the sampling process.
Consider a continuum of agents that are indexed by their coordinates $z \in \mathcal{Z} = [-1,1]$. For convenience, we will assume that agents are uniformly distributed according to $F = \text{Uniform}(\mathcal{Z})$, so that their density with respect to the Lebesgue measure is $f(z) = \frac{1}{2}$. Our results extend easily to any density bounded away from $0$ at the cutoff, although our expressions for the bias of RDD need not apply in the more general setting.
We work in the usual potential outcomes framework with binary treatment, where the observed outcome at $z$ satisfies
Here, $d: \mathcal{Z} \to \{0,1\}$ is the treatment assignment function. $Y^+_d(z)$ and $Y^-_d(z)$ are the potential outcomes under treatment and control respectively. Because of the treatment assignment rule in RDD, defined below, we will denote quantities related to treatment with “$+$" and those related to control with “$-$".
Potential outcomes are indexed by $d$ because as a result of spillovers, they may depend on the entire treatment assignment function. Let every agent $z$ have the set of relevant neighbors $R(z) \subset \mathcal{Z}$ with measure $|R(z)|$. These are agents whose realized outcomes affect $z$. Let potential outcomes be defined as follows:
The above model extends the continuity-based framework of hahn2001identification to include two sources of spillovers, both of which are linear-in-means. The exogenous spillover term is $\gamma(z)\nu_d(z)$, where $\nu_d(z)$ is the mean treatment status of $z$'s neighbors, and $\gamma(z)$ is the effect of this term on $z$'s outcome. The endogenous spillover term is $\delta(z)\mu_d(z)$, where $\mu_d(z)$ is the mean outcome of $z$'s neighbors, and $\delta(z)$ is the effect of this term. Spillovers are assumed to enter the treated and control potential outcomes in the same way. As will become clear in Section (ref), this leads to a particularly simple formula for the parameters of interest.
The linear-in-means assumption is potentially restrictive. However, it is the workhorse model in the peer effects (e.g. manski1993identification, bramoulle2009identification, deGiorgi2010identification,goldsmith2013social,dePaula2024identifying) and spatial autoregression (e.g. cliff1973spatial,kelejian1998generalized, kelejian2010specification,lee2004asymptotic, lee2007gmm) literatures. We consider it to be a reasonable first step for analyzing endogenous spillovers, a task which is challenging in many settings. The linear-in-means assumption is in fact stronger than necessary for our characterization of the estimands. In particular, the qualitative results in Section (ref) does not require linearity in the effects of spillovers, and allows neighbors to have different weights in treated and control outcomes. The main requirement being that neighborhoods have approximately bounded support. However, as will become clear in Section (ref), the specific form of spillovers needs to be assumed in order to recover target parameters in our preferred asymptotic regime. We therefore focus on the above model. Nonetheless, it is more general than the standard peer effects model in that the effects of spillovers, $\delta$ and $\gamma$, is allowed to vary with $z$.
The remaining conditions in Assumption (ref) concerns continuity of the functions $m^+, m^-, \delta$ and $\gamma$ as well as the boundedness of $\delta$. Continuity of $m^+$ and $m^-$ around the cutoff is the main identifying assumption in RDD. Standard RDD requires only that $m^+(z)$ and $m^-(z)$ are continuous at $z = 0$. However, due to spillovers, continuity of potential outcomes also requires continuity of the two functions around the boundaries of $R_n(0)$ as well as the continuity of $\gamma$ and $\delta$. Since $R_n(0)$ is fixed under some of the regimes we consider, we assume continuity over $\mathcal{Z}$ for simplicity. We also slightly strengthen the condition to Lipschitz continuity. To ensure that $Y^+_d$ and $Y^-_d$ are well-defined, we require $\sup_{z \in \mathcal{Z}} |\delta(z)| < 1$, as well as for $\sup_{z_1,z_2 \in \mathcal{Z}} |\delta(z_1) - \delta(z_2)| \leq 1$.
In a setting with spillovers, any two treatment assignment functions $d$ and $d'$ can give rise to a treatment effect $$Y_d(z) - Y_{d'}(z)$$ that is potentially of interest. Following the literature on causal inference with spillovers (see e.g. {hudgens2008toward), we focus on the following two parameters:
The direct treatment effect on a given agent is the effect of treatment on their outcomes, keeping the treatment assignment of all other agents unchanged at some $d$. This parameter is useful for thinking about selective rollout of some policy, under which it is plausible to consider only the direct treatment effect, treating the spillovers as fixed. Here, we abuse notation in using $d$ to refer to two different treatment assign functions, although they are identical on $\mathcal{Z} \setminus {0}$. In principle, direct treatment effect depends on $d$. In RDD, it is natural to focus on:
The above treatment assignment function is the natural reference point for evaluating direct treatment effects in RDD. However, under Assumption (ref), the agent at $z = 0$ experiences the same spillover whether or not they are treated for two reasons. Firstly, $\delta(z)$ and $\gamma(z)$ are assumed to be the same regardless of treatment status. Secondly, the agent $z = 0$ is infinitesimal, so that changing their treatment status does not affect $\mu_d(z)$ or $\nu_d(z)$ for any $z \in \mathcal{Z}$. Implicitly, this is a model involving dense network asymptotics. The result of these two factors is that $\tau_\text{DIR} = m^+(0) - m^-(0)$ regardless of $d$. Additionally, $\tau_\text{DIR}$ is equal to local average treatment effect in the standard setting with no spillovers.
With infinitesimal agents, direct treatment effects are easy to define since we can change the treatment status of one individual but still keep spillovers fixed. Nonetheless, there is a coherent notion of spillovers in this model: if two treatment functions $d$ and $d'$ differ on a set of positive measure, then $\mu_d(z)$ and $\nu_d(z)$ need not be equal to $\mu_{d'}(z)$ and $\nu_{d'}(z)$. This point is also evident in the definition of our next parameter.
The total treatment effect on a given agent is the effect on their outcome when we switch the treatment status of the entire population from control to treatment. This corresponds to the effect on $z = 0$ from a large-scale rollout of treatment. In our notation, the subscripts ${\boldsymbol{\iota}}$ and ${\boldsymbol{0}}$ denote treatment assignments in which everyone and no one is treated respectively. As before, we will focus on the agent at the cutoff, $z = 0$. With total treatment effect, we are evaluating potential outcomes under different treatment assignment functions. Consequently, spillovers no longer cancel out. We remark that $\nu_{\boldsymbol{\iota}}(z) = 1$ and $\nu_{\boldsymbol{0}}(z) = 0$ so that exogenous spillovers at $z = 0$ is $\gamma(0)$.
In the presence of spillovers, the neighborhood structure is critical in determining the estimands of RDD. We focus on the case where spillovers occur along the running variable:
Here, $\lVert \cdot \rVert$ is taken to be the Euclidean distance. As such, $z$ is affected only by units whose running variable takes value within $r_n$ of $z$. All other units exert no effect. Together with an i.i.d sampling assumption introduced in the next section, the above assumption implies that the exposure map between sampled units is a random geometric graph (see e.g. penrose2003random). This is a well-studied model that is commonly used for spillovers and interference (see e.g. leung2020treatment).
When neighborhoods are defined by the running variable, nearby units that are comparable in their conditional means ($m^+, m^-$) can experience spillovers that are not comparable. Consequently, RDD may not recover any meaningful treatment effect parameters, as our analysis in Section (ref) shows. This fundamental tension is absent when neighborhoods are defined by another variable, say $w$, that is suitably continuous with respect to $z$. We will refer to such a $w$ as a background variable. Suppose for now that units are defined by their values of $(z,w)$ and that neighborhoods are defined as
Additionally, suppose $z$ and $w$ have a joint density in a neighborhood containing the line $z = 0$. Then, local to the cutoff, we have that $z$ is uncorrelated $w$, so that neighborhoods, and in turn spillovers, are uncorrelated to the RDD treatment assignment. The result is that RDD always recovers the direct treatment effect, provided that the correlation between $z$ and $w$ remain fixed in the asymptotics.\footnote{The first version of this paper presented the above intuition for spillovers on background variables (auerbach2024regression). torrione2024regression proves this result in a related setting, drawing a connection to multi-score RDD. Also see the discussion after Theorem (ref).}
In light of the above discussion, spillovers in RDD may appear pathological. However, there are many settings in which neighborhoods appear strongly correlated with the running variable. In spatial RDDs, it is often plausible that nearby units interact. For example, jardim2022boundary studies the effect of minimum wage policies on labor market outcomes at the census tract level. However, it is clear that census tracts compete in the same labor market as all other census tracts within a reasonable commute. gonzalez2021cell provides another example based on cell phone coverage and electoral fraud which we take up in greater detail in Section (ref). leung2022rate provides many more examples. The same is true with test score RDDs, such as the classic thistlethwaite1960regression which studies the effect of merit scholarships on academic achievements. Here, scholarships are assigned based on test score cutoffs and it is plausible that students who are close in test scores interact. This could happen because of homophily, because students take classes and therefore socialize with those of similar abilities, or because they compete for the same set of career opportunities. In instances like the above, our model is a useful starting point for understanding the effect of spillovers.
Having defined a sequence of models for generating potential outcomes, we close this section by considering sampling. We will assume that our data takes the following form:
In words, the data-generating process operates as follows. A continuum of agents interact and their outcomes are determined by the values of their running variables as well as spillovers. The econometrician then samples agents randomly and observes their outcomes, possibly with an error that has $0$ conditional mean. We also assume that the error has conditional variance that is continuous in $z$ -- a standard assumption. Our framework resembles dePaula2018identifying in defining interactions to occur in the population.
In this section, we show that the local linear regression estimator may not recover meaningful treatment effect parameters in the presence of spillovers, necessitating proposals such as that in Section (ref). Section (ref) defines the local linear regression estimator and Section (ref) describes its estimands in the presence of spillovers. Our results suggest that the ratio of $r_n/h_n$ is a key modeling assumption and we argue for using $r_n \propto h_n$ in Section (ref). Finally, Section (ref) considers the use of donut-hole designs as a solution to spillovers.
The local linear regression estimator for the local average treatment (equivalently, $\tau_\text{DIR}$) is defined as follows:
Second order kernels are commonly used in local linear regression. For a discussion on the properties of higher order kernels, see e.g. wand1995kernel. We additionally assume that $K$ has finite support and is bounded and non-negative.
In the absence of spillovers, it is standard practice to estimate the local average treatment effect using $\hat{\tau}_{\text{RDD}}$ (see e.g. hahn2001identification, cattaneo2022regression). In this setting, $\hat{\tau}_{\text{RDD}}$ is consistent for $\tau_\text{DIR}$, though inference is complicated by the presence of an asymptotic bias of order $h_n^2$. Various solutions are available for the problem of inference (see e.g. calonico2014robust, armstrong2018optimal). In sum, the properties of RDD are relatively well-understood in settings without spillovers.
Our first result characterizes the estimands of RDD when spillovers occur along the running variable.
When spillovers occur along the running variable, the estimand of RDD exhibits phase transition, with the phases depending on the ratio of $r_n$ to $h_n$. When $r_n \gg h_n$, RDD estimates direct treatment effect. On the other hand, when $r_n \ll h_n$, RDD estimates total treatment effect. In the intermediate regime where $r_n \propto h_n$, it estimates a parameter $\tau_*$, which is displayed in full in Equation (ref). $\tau_*$ depends on the $c$, the limiting ratio of $r_n$ and $h_n$, even though this is suppressed in the notation. Generally speaking, $\tau_*$ does not have a clear interpretation. In particular, it is not equal to either $\tau_\text{DIR}$ or $\tau_\text{TOT}$ unless either $\delta(0) = \gamma(0) = 0$ (i.e. no spillovers) or $\tau_\text{DIR} = \gamma(0) = 0$ (i.e. treatment has 0 effect), in which case $\tau_\text{DIR} = \tau_\text{TOT} = \tau_*$. It may be close to $0$ even when $\tau_\text{DIR}$ and $\tau_\text{TOT}$ are large in magnitude -- a common concern expressed by empirical papers regarding spillovers.
The above result is intuitive. The local linear regression estimator essentially treats units within the bandwidth as being comparable. If $r_n$ is large relative to $h_n$, units in the bandwidth have neighborhoods that overlap almost completely so that they experience similar spillovers. As such, units on different sides of the cutoff differ only in their treatment status and $\hat{\tau}_\text{RDD}$ estimates direct treatment effect. Conversely, when $r_n$ is small relative to $h_n$, units to the left of the cutoff essentially only has control neighbors, while units to the right of the boundary only has treated neighbors. Comparing these units therefore reveals total treatment effect. When $r_n \propto h_n$, RDD compares a mix of units, some of which have similar neighborhoods, some of which do not. The resulting estimand is therefore a complicated interpolation of direct and spillover effects.
Theorem (ref), in particular cases (b) and (c), stands in stark contrast to existing work on spillovers within an RDD. aronow2017rdspill studies spillovers under the local randomization framework of cattaneo2015randomization and find that RDD always recovers a weighted average of direct treatment effect. Under local randomization, RDD is exactly an experiment under some neighborhood of the cutoff. Consequently, treatment assignment is orthogonal to spillovers, reproducing the results seen in the literature on experiments with spillovers (see e.g. hudgens2008toward, aronow2017estimating, leung2022causal). torrione2024regression provides a similar analysis under the continuity-based approach of hahn2001identification. They assume that the running variable and the variable determining neighborhood structure have a continuous joint density, ruling out the case we consider. Given continuous joint density, they find that treatment assignment is approximately orthogonal to spillovers. As a consequence, RDD recovers a weighted average of direct treatment effect even under the continuity framework.
Our paper studies RDD under a continuity-based framework and we focus on the case where neighborhoods are completely determined by the running variable. In this setting, the extent of orthogonality between treatment assignment and spillovers depends on the ratio $r_n/h_n$. Our approach -- particularly the $r_n \propto h_n$ case -- leads to an asymptotic model in which spillovers can affect the estimands of RDD. This is not possible in the framework of aronow2017rdspill and torrione2024regression. Our model therefore provides a basis for using data to speak to the effects of spillovers, a task which we take up in Section (ref). Finally, we note that this paper is also unique relative to the above two papers in considering endogenous spillovers.
Theorem (ref) shows that how we model the relative rates between $r_n$ and $h_n$ affects the interpretation of $\hat{\tau}_\text{RDD}$. We argue that the $r_n \propto h_n$ regime is the natural choice when researchers are concerned about spillovers.
In any given application, spillovers occur over some radius that is a finite ratio $c$ of the bandwidth. The regimes $r_n \gg h_n$ or $r_n \ll h_n$ are therefore unrealistic edge cases that take $c = 0$ or $c = \infty$. Consequently, they assert that $\hat{\tau}_\text{RDD}$ must either recover direct treatment effect or total treatment effect. There is no room for the data to speak to the amount of direct effects relative to spillovers.
In contrast, the estimand of RDD depends explicitly on the ratio of $c$ under the regime $r_n \propto h_n$, which is a better approximation to how spillovers manifest in practice. It leads to an asymptotic model that preserves more features of the finite sample distribution and therefore provides greater scope for reasoning about spillovers in the data. For this reason, $r_n \propto h_n$ is our preferred framework for analyzing RDD with spillovers. We do not claim $r_n$ to be a quantity that changes with sample size, only that taking it to $0$ leads to a useful approximation. Such an approach is also taken in traditional analysis of boundary bias for kernel methods (see e.g. Section 5.5 in wand1995kernel). Here, the target point is modeled as drifting towards the boundary at rate $h_n$ to capture the notion that these two points are close.
In our preferred regime, RDD is potentially inconsistent for either treatment effects parameter. Section (ref) proposes a method for recovering both direct treatment effects and spillovers.
In the remainder of this section, we informally discuss donut-hole designs as viewed through the lens of our model. Often suggested as a solution for estimating $\tau_\text{DIR}$ in the presence of spillovers, donut-hole designs involve excluding observations close to the cutoff from estimation. The idea is that these units are contaminated by spillovers, so that removing them should lead to the recovery of $\tau_\text{DIR} = m^+(0) - m^-(0)$. This approach does not work under our model, but we state an alternative model of spillovers under which donut-hole RD can be justified, provided that there are no endogenous spillovers.
To define the donut-hole RD, let $h_n^o > 0$ be such that $h_n^o < h_n$. Then donut-hole RD estimates $\hat{\beta}^+$ by local linear regression on observations for which $Z_i \in [h_n^o, h_n]$. Similarly for $\hat{\beta}^-$ and observations for which $Z_i \in [-h_n, -h_n^o]$. Let the estimand of the donut-hole RD be denoted $\tau_o$.
Under our linear-in-means model, suppose $h_n^o = r_n$. That is, we exclude all observations that have neighbors on the other side of the cutoff. Suppose there are no endogenous spillovers ($\delta(0) = 0$). Then $\tau_o = \tau_\text{TOT}$ and not $\tau_\text{DIR}$. In our model, agents closest to the cutoff experience the most similar spillovers, so that comparing them leads to the best estimate of $\tau_\text{DIR}$. However, after excluding units in the donut hole $[-h_n^o, h_n^o]$, the remaining units only have neighbors who have the same treatment status as them so that RDD estimates $\tau_\text{TOT}$. In this sense, a donut-hole RD is exactly the wrong thing to do if researchers are interested in $\tau_\text{DIR}$.
However, suppose $\gamma(z) = \check{\gamma}(z) \mathbf{1}\{z \leq 0\}$. In other words, only control units experience spillovers based on how many of their neighbors are treated. This is reasonable, for example, in the case of information treatment, where information can diffuse from the treated to control units but not vice versa. In this case, if $h_n^o = r_n$ and $\delta(0) = 0$, the intuition in the above paragraph still applies so that $\tau_o$ estimates total treatment effect. However, in this model, total treatment effect is also equal to direct treatment effect under the reference treatment assignment function $d = 0$. Note that direct treatment effect is now dependent on $d$, since spillovers enter the treated and control potential outcomes asymmetrically. Donut-hole RDs can therefore be used to recover $\tau_\text{DIR}$ under this model.
Finally, we note that once $\delta(0) \neq 0$, donut-hole RDs do not recover $\tau_\text{DIR}$ or $\tau_\text{TOT}$ in either model. This is because under endogenous spillovers, the outcome of a given unit is affected by all other units in the domain: a given unit $z$ has an outcome that depends on the outcomes of $[z-r_n, z+r_n]$. In turn, the outcome for $z+r_n$ depends on the outcomes on $[z, z+2r_n]$, and so on. As such, it is not possible to isolate units that are “contaminated" by spillovers using a donut design.
As we argued in Section (ref), our preferred asymptotic regime features $r_n \propto h_n$, in which the RDD estimand need not have a causal interpretation. In order to disentangle direct treatment effect from spillovers, Section (ref) proposes a local analog of the peer effects regression that we term the local spillover regression (LSR). We discuss estimation consistency and inference in Section (ref). LSR requires researchers to specify a radius $r_n$. We consider choosing $r_n$ as well as $h_n$ in Section (ref).
Since spillovers are the cause of inconsistency, our proposal is to estimate $\mu_d(z)$ and $\nu_d(z)$, and then to incorporate them into the local linear regression.
Suppose for now that $r_n$ is known. For a given $Z_i$, define its observed neighbors to be:
To be consistent with our econometric model, $\lVert \cdot \rVert$ is taken to be the Euclidean distance. We then estimate the mean outcome of $Z_i$'s neighbors using a leave-$i$-out estimator:
Additionally, we define
where the only difference is that $z = 0$ is not observed. We focus on the case when the distribution of $Z$ is uniform, so that $\nu_d(Z_i)$ known. When $f$ is not uniform, it is straightforward to form the analogous leave-$i$-out estimator for $\nu_d(Z_i)$.
We propose to estimate $\tau_\text{DIR}$, $\gamma(0)$ and $\delta(0)$ using the following:
$\tilde{X}_i$ is our estimate of the unobserved $X_i$. The estimator $\tilde{\beta}$ targets the following:
The first four components of $\tilde{X}_i$ are terms coming from the standard LLR. They are obtained by rewriting the two separate local regressions in Definition (ref) as a single regression. As such, the coefficient of $D_i$ targets $\tau_\text{DIR} = {m^+(0) - m^-(0)}$. The remaining components address spillovers. $\tilde{\mu}_d(Z_i) - \tilde{\mu}_d(0)$ and ${\nu}_d(Z_i) - {\nu}_d(0)$ are our estimated mean outcome and mean treatment status in the neighborhood of $Z_i$. As such, their coefficients target $\delta(0)$ and $\gamma(0)$ respectively. $Z_i(\tilde{\mu}_d(Z_i) - \tilde{\mu}_d(0))$ and $Z_i({\nu}_d(Z_i) - {\nu}_d(0))$ are not needed for consistency but they reduce the asymptotic bias when $\delta(z)$ and $\gamma(z)$ are not constant, analogous to how introducing $Z$ reduces the bias of LLR relative to the Nadaraya-Watson estimator.
In view of the above discussion, we will also use the notation $\tilde{\tau}_\text{DIR} := \tilde{\beta}_3$, $\tilde{\delta}(0) := \tilde{\beta}_5$, $\tilde{\gamma}(0) := \tilde{\beta}_7$. When $r_n \to 0$, $\tau_\text{TOT} \to \frac{\tau_\text{DIR} + \gamma(0)}{1-\delta(0)}$. As such, we also define $\tilde{\tau}_\text{TOT} :=\frac{\tilde{\tau}_\text{DIR} + \tilde{\gamma}(0)}{1-\tilde{\delta}(0)}$.
Section (ref) presents the consistency properties of LSR. Section (ref) provides the asymptotic distribution for $\tilde{\beta}$. Section (ref) covers confidence intervals based on the bias-aware approach of armstrong2020simple.
This section presents results on the consistency of LSR. We first discuss consistency of $\tau_\text{DIR}$, which obtains under fairly unrestrictive conditions. We then turn to $\gamma(0)$ and $\delta(0)$, which are more challenging to estimate and will require more assumptions.
For direct treatment effect,
The main condition for consistency of $\tilde{\tau}_\text{DIR}$ is the Lipschitz continuity of $m^+$, $m^-$, $\delta$ and $\gamma$. This is a slight strengthening of the conditions needed for LLR, which typically assumes continuity of $m^+$ and $m^-$ at the cutoff. Including estimated spillover terms is therefore a simple and effective method to ensure that the target parameter is $\tau_\text{DIR}$.
Recovering spillover effects requires more assumptions:
Then, we have the following:
The additional assumptions here address various sources of non-identification in the model. Condition (a) is specific to $z \in \mathbb{R}$. In this case, when $c > 1$, $\nu_d(z) = z$ under the assumption of uniform $f$. Having $c < 1$ is therefore necessary for the identification of $\gamma(0)$. This collinearity does not arise when units are in e.g. $\mathbb{R}^2$. Condition (b) is needed to ensure that $\mu_d(z)$ is not collinear with $\nu_d(z)$. Our approximation result in Lemma (ref) shows that locally,
If $\delta(0) = 0$, the first and second order terms are constant, so that $\mu_d(z) - \mu_d(0) = (\nu_d(z) - \nu_d(0))$. Meanwhile, if both $ \tau_\text{DIR}$ and $\gamma(0)$, then $\mu_d(z) - \mu_d(0) \approx 0$. Intuitively, to learn about $\delta(0)$, we need treatment to create some effect ($\tau_\text{DIR}$ or $ \gamma(0) \neq 0$) that decays in $z$.
Condition (b) is relatively strong but it is reminiscent of the assumption that $\tau \gamma + \delta \neq 0$ common in the peer effects literature (see e.g. Proposition 1 in bramoulle2009identification). Finally, since the collinearities pertain to some combination of $z$, $\mu_d(z)$ and $\nu_d(z)$, $\tilde{\tau}_\text{DIR}$ is consistent for $\tau_\text{DIR}$ even when these conditions are not satisfied.
We next state a central limit theorem for $\tilde{\beta}$. For ease of exposition, we will maintain the full set of assumptions in Theorem (ref). We focus on the case when the bandwidth is assumed to have the MSE-optimal rate of $n^{-1/5}$ and introduce the following assumption:
Then,
When $h_n \propto n^{-1/5}$, there is an asymptotic bias term $\mathbf{B}_{n}$ that we address below. \(\widetilde{\mathbf{Q}}\) is the regression design matrix. The matrix \(\bm{\Omega}\) in the definition of \(\mathbf{V}\) has two components. The term \(K (Z_{i} / h) X_{i} \varepsilon_{i}\) captures the usual source of estimation error in weighted regression if the “ideal” regressors \(X_{i}\) were known. In this case, the only source of estimation error would be due to the unobserved error term \(\varepsilon_{i}\). The feasible regressors \(\widetilde{X}_{i}\), used in place of \(X_{i}\), contain an additional source of error: estimation of the unknown neighborhood average outcomes \(\mu_{d} (Z_{i})\). The term \(\eta_{X} \left( \xi_{i} \right)\) captures this additional source of estimation error, i.e. the difference \(\widetilde{X}_{i} - X_{i}\). It is straightforward to see that the sample analogue of $\bm{\Omega}$ combined with \(\widetilde{\mathbf{Q}}\) is a consistently estimates \(\mathbf{V}\). Let the plug-in estimator be denoted $\widetilde{\mathbf{V}}_{n}$.
The bias term in Theorem (ref), $\mathbf{B}_{n}$, is asymptotically non-negligible when $h_n \propto n^{-1/5}$. To form valid confidence intervals, we follow the bias-aware approach of armstrong2020simple. This entails restricting the classes of functions to which $m^+, m^-, \delta$ and $\gamma$ can potentially belong, so that an upper bound for $\mathbf{B}_{n}$ can be obtained. Confidence intervals can then be appropriately widened to ensure coverage under this worst-case bias.
To that end, define the Taylor class of order 2:
This is the class of functions whose first order Taylor approximation has error less than $Mz^2/2$. We can loosely think of the tuning parameter $M$ as the upper bound for the second derivative of $f$ at $0$. We will now assume that:
It is then straightforward to see that $|\mathbf{B}_{n,j}| \leq \overline{\mathbf{B}}_{n,j} $, where
and $w_j(Z_i) = e_j' \hat{Q}_n^{-1} \frac{1}{h_n}K\left(\frac{Z_i}{h_n}\right) \tilde{X}_i $. Since $\mu_d(Z_i)$ and $\mu_d(0)$ are unobserved, we replace them with $\tilde{\mu}_d(Z_i)$ and $\tilde{\mu}_d(0)$. Denote the resulting estimator $\widetilde{\mathbf{B}}_{n,j}$. We can then form asymptotically valid confidence intervals as follows:
The above result is an immediate consequence of Theorem (ref) and the fact that $|\mathbf{B}_{n,j}| \leq \overline{\mathbf{B}}_{n,j}$ under Assumption (ref). It yields asymptotically valid $1-\alpha$ confidence intervals for $\tau_\text{DIR}$, $\delta(0)$ and $\gamma(0)$. It is then straightforward to obtain valid confidence interval for $\tau_\text{TOT}$ by combining these confidence intervals together with a Bonferroni correction.
Results in the previous section are predicated on researcher's choices of $r_n$ and $h_n$. This section presents heuristics for choosing these parameters.
In principle, researchers should choose $r_n$ based on their beliefs about likely sources of spillovers. In practice, however, such a choice may not be obvious. Researchers may instead consider choosing $r_n$ via cross-validation as follows.
Let the fold number $L$ be specified by the user and suppose we are given some initial bandwidth $h^0_n$. Partition the data set $D$ into $L$ equal-sized subsets $D_1, ..., D_L$ and let $\tilde{\beta}^{(l)}(r)$ be the LSR estimator computed on $D \setminus D_l$ under radius $r$. Define the out-of-sample MSE as:
We can then choose
Intuitively, if spillovers exist, getting the true radius correct should lead to good out-of-sample fit. It would be difficult to do better than the true radius in MSE since this would require the misspecified model to overfit on out-of-sample $\varepsilon_i$'s, which is independent of the training sample.
Finally, to choose the initial bandwidth $h_0$, we might consider a rule of thumb such as
Here, the $ \max \, \lVert Z_i \rVert$ scaling ensures that $h^0_n$ is invariant to the units of measurement of the running variable.
The simulations in Section (ref) explores the finite sample performance of the above method. We find that cross-validation can increase the MSE of various parameters by around 4 times relative to the oracle procedure in which $r_n$ is known. This appears to be a reasonable trade-off when researchers have only weak priors about spillovers in a given setting.
The rule of thumb in Equation (ref) leads to a reasonable sequence of bandwidths, but it does not take MSE into consideration. Here, we present an intuitive adaptation of the procedure in armstrong2018optimal, which aims to reduce finite sample version of the worst-case MSE.
Given an initial bandwidth $h^0_n$ and a radius that is either pre-specified or selected via cross validation, we can perform LSR and obtain an initial estimate of the asymptotic variance based on the sample analog of $V_n$ in Theorem (ref). Call this estimate $\widetilde{\mathbf{V}}^0_n$. For any given bandwidth $h$, we can also obtain an estimate for the maximum bias for $\tilde{\beta}_j$ using the sample analog of $\overline{\mathbf{B}}_{j,n}$ in (ref). Call this term $\widetilde{\mathbf{B}}_{j,0}(h)$. We can then choose bandwidth by solving:
While any $\tilde{\beta}_j$ can potentially work with this procedure, we recommend the use of $\tilde{\beta}_3 = \tilde{\tau}_\text{DIR}$. This is because $\tilde{\tau}_\text{DIR}$ targets a parameter of immediate interest and it is consistent under weaker conditions than $\tilde{\delta}(0)$ and $\tilde{\gamma}(0)$.
In this section, we revisit gonzalez2021cell, which studies the effect of cell phone coverage on fraud in the 2009 Afghan presidential election. The analysis is conducted at the polling station level and the outcomes of interest are (1) indicator for whether fraud has likely occurred and (2) likely share of fraudulent votes. Using a spatial regression discontinuity design, the author finds that cell phone coverage reduces the probability that fraud occurs by 8 percentage points and decreases the share of fraudulent votes by 4 percentage points. gonzalez2021cell is concerned with spillovers, highlighting for example that voters may walk from a non-covered to a covered area to report fraud. To test for spillovers, the author invokes donut-hole type reasoning and compares the fraud level at non-covered polling stations between 0--2 km of the boundary to those that are 2--4 km, 4--6 km or 6--8 km away. The idea is that spillovers should lead to spikes in fraud levels at polling stations closer to the boundary. They find that the differences in outcomes are not statistically significant and conclude that there is likely no spillovers.
In the spirit of the above exercise, we apply our method in the setting of gonzalez2021cell as a robustness check. We first replicate the paper's result using either the method of armstrong2020simple (AK) or calonico2014robust (CCT). All tuning parameters are selected using the default procedures in the authors' respective R packages. Relative to the original specification, we do not include spatial fixed effects since they fall outside the scope of conventional theory. We then proceed with local spillover regression (LSR) taking the bandwidth choices of AK and CCT as given. Following gonzalez2021cell, we consider circular neighborhoods that are between 2 to 8 km in radius. In principle, we can compute $\nu_i$ given the exact boundaries of cell phone coverage. Since that is not available, we estimate $\nu_i$ using the number of treated polling station in a given station's neighborhood, although we treat it as known when conducting inference. With the running variable measured in kilometers, we chose $M_m$ to be 0.01, which implies that the local linear approximation to the conditional probability of fraud has an error up to 1 (i.e. the entire support of the variable) at the 10 km radius. Similarly, we chose $M_\delta$ and $M_\gamma$ to be 0.01. Finally, we consider the case in which we select the spillover radius by cross-validation and choose bandwidth using our rule-of-thumb (LSR-CV).
Table (ref) presents the results when using the indicator for fraud as outcome and the sample is the full set of polling stations. This is comparable to Panel A Column (1) in Table 2 of gonzalez2021cell. Across the various bandwidth and radius choices, we find $\delta(0)$ to be positive and significant at the 5% level. The estimates are also relatively stable across parameter values hovering between 0.6 to 0.75. Positive exogenous spillovers would arise, for example, if fraud is driven by local norms of corruption (fisman2007corruption; barr2010corruption and references therein). On the other hand, we do not find conclusive evidence of exogenous spillovers. These results continue to hold with data-driven choices of radius and bandwidth. Our cross-validation procedure suggests that spillovers occur across a radius of 6.1 km, which is within the plausible range of radii considered by gonzalez2021cell. Results when using fraud share as outcome are similar.
The above exercise suggests that spillovers can be a threat to the identification of treatment effect parameters in RDD. In this case, our method can be useful as a robustness check. Additionally, the parameters for endogenous and exogenous spillovers may also be of interest since they can shed light on how social interactions mediate a variable of interest.
Regression Discontinuity Design allows researchers to obtain causal estimates of treatment effects at the cutoff. It relies on the assumption of no spillovers, which may not hold in practice. When neighborhoods are determined by the running variable, we find that the estimand of RDD is sensitive to the ratio of two terms: (1) radius over which spillovers occur and (2) the bandwidth used for local linear regression. In our preferred approximation, those two terms are of similar order so that an alternative to local linear regression is needed to recover direct treatment effects and spillovers. We propose the local spillover regression -- the local analog of peer effects regression -- and show that it can be a useful tool for addressing spillovers in RDD.