Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
66,749 characters · 14 sections · 32 citation commands
Placebo Discontinuity Design
Regression discontinuity design (RDD) strategies are widely used to estimate the effect of a treatment on an outcome of interest, when the probability of treatment assignment exhibits a discontinuous jump at a known threshold thistlethwaite1960regression. Applications of RDD are broad and include studies in electoral politics hall_what_2015,spenkuch_political_2018, education cook_birthdays_2016,goodman_can_2019, housing markets kumar_restrictions_2018, and market design abdulkadiroglu_regression_2017,kong_sequential_2021, among others.
The main identifying assumption in standard RDD is the continuity of expected potential outcomes at the cutoff hahn2001identification. This assumption ensures that, in the absence of treatment, units just above and below the cutoff are comparable. As discussed in Remark (ref), this condition is closely related to the requirement that the conditional distribution of unobserved confounding given the running variable is continuous at the cutoff lee2010regression. To empirically assess the plausibility of this assumption, researchers often conduct falsification tests using placebo variables.\footnote{The literature also considers predetermined covariates, of which the placebo treatment may be viewed as a special case.} For example, imbens_regression_2008 and cattaneo_practical_2024 recommend checking for discontinuities in placebo outcomes as a diagnostic for invalid RDD. However, when these falsification tests indicate a failure of the RDD continuity assumption, current methods offer no path forward for recovering the treatment effect.
In this paper, we propose a new identification strategy that leverages a placebo treatment and a placebo outcome to recover the treatment effect even when the continuity of expected potential outcomes may fail. We introduce a local linear instrumental variable estimator that calculates the treatment effect as the difference between the right-limit of the observed outcome and an adjusted left-limit that accounts for unobserved confounding via placebo variables. The estimator decomposes into a standard RDD estimate for the target outcome, $\hat{\tau}_{\text{rdd}}^{y}$, plus an adjustment term $(\hat{\tau}_{\text{rdd}}^{w})^{\top}\hat{\gamma}_{-}$, where $\hat{\tau}_{\text{rdd}}^{w}$ is the discontinuity in the placebo outcome and $\hat{\gamma}_{-}$ is an appropriate weight.
Conceptually, we build on the insights and techniques of the placebo variable literature (also called the negative control literature) miao2018identifying,deaner2018nonparametric,tchetgen2020introduction, to answer a salient question about RDD, which is an extremely popular method in social science. To the best of our knowledge, our identification result appears to be novel and motivates a new estimation problem. For this new estimation problem, we propose and analyze what appears to be a novel variation of the widely used local linear RDD estimator fan_variable_1992, and we derive bias-corrected inference calonico_robust_2014,calonico_regression_2019. Future work may develop bias-aware inference armstrong2018optimal,imbens2019optimized,noack2024bias.
The rest of the paper proceeds as follows. Section (ref) illustrates the main idea with a real world application. Section (ref) presents our main identification assumptions and result. Section (ref) derives our estimator. Section (ref) proves consistency and normality with bias-corrected inference. Section (ref) concludes. The main text focuses on the core insights, while formal proofs are deferred to the appendix.
To motivate our approach, consider an example adapted from angrist_maimonides_2019. Suppose we are interested in estimating the effect of transitioning from a large class (39-40 pupils) to a small class (20-21 pupils), denoted by the binary treatment $A$. The outcome of interest is student achievement $Y$. A Maimonides-style rule determines class size: schools must split classes when total enrollment in a grade level $D$ exceeds a statutory threshold $d^{*} = 40$. For instance, a school with $D=42$ students in a grade level receives funding for two classes of 21 students ($A=1$), while a school with exactly $D=40$ students in a grade level keeps one class of 40 students $(A=0)$.
To apply a conventional RDD estimator in this setting, one must assume that the expected potential outcomes $\mathbb{E}\mathopen{}\mathclose\bgroup\originalleft\{Y(1, d, U, \eta_{y}) \mid D = d\aftergroup\egroup\originalright\}$ and $\mathbb{E}\mathopen{}\mathclose\bgroup\originalleft\{Y(0, d, U, \eta_{y}) \mid D = d\aftergroup\egroup\originalright\}$ are continuous at the cutoff $d^{*}$. Here, $Y(A, D, U, \eta_y)$ denotes the test score of a student in a class size of 20-21 students $(A = 1)$ or 39-40 students $(A = 0)$, given the total enrollment in the grade level $(D)$. We model unobserved heterogeneity through two components---school administrator sophistication $U$, which may influence total enrollment near the cutoff, and student-level heterogeneity $\eta_y$, which is random. The left panel of Figure (ref) depicts a valid RDD scenario under this assumption.
However, this assumption may fail if school administrators behave strategically. For instance, more sophisticated school leaders may recognize the budgetary and pedagogical advantages of exceeding the class size threshold, perhaps by enrolling just above the cutoff. These administrators may also tend to run more effective schools with higher-achieving students, thereby inducing a spurious upward jump in test scores around $d^*$ that is not due to class size per se angrist_maimonides_2019. In such a case, the unobserved aptitude of school leaders $(U)$ confounds the relationship between enrollment $(D)$ and student achievement $(Y)$ near the cutoff. As a result, schools just above $d^*$ may perform better because they are led by more capable administrators, invalidating the continuity assumption that would allow researchers to evaluate the effect of class size. This confounded scenario is depicted in the right panel of Figure (ref).
Our identification strategy recovers the treatment effect in such confounded settings by positing that continuity holds after further conditioning on the unobserved school-level aptitude $U$. As shown in Figure (ref), we require that, conditional on both $D=d$ and $U=u$, the expected potential outcomes become continuous at the cutoff. We then exploit the statistical relationship between test scores in the prior year $W$, which serve as placebo outcomes, and the availability of transfer-eligible students based on birthdates $Z$, which act as placebo treatments, to approproately adjust for the unobserved confounder $U$.
Intuitively, test scores in a previous year are unaffected by enrollment $(D)$ in the current year, making it a valid placebo outcome ($W)$. In this sense, previous test scores $(W)$ serve as a proxy for the school administrator aptitude ($U$).
Meanwhile, the availability of students able to transfer across grades should only affect student test scores $(Y)$ through its influence on total enrollment in the grade $(D)$, making it a valid placebo treatment $(Z)$. The availability of transfer students is like a local instrument, or a special type of pre-determined covariate, that can help to distill the variation due to school administrator aptitude.
In this section, we formalize the model, interpret the key assumptions, and present our main identification result: how to recover the treatment effect in an RDD setting with strategic behavior around the cutoff. Standard practice uses placebo variables to detect this issue; as our first contribution, we use placebo variables to correct this issue.
Consider the nonseparable model $$ Y=Y(A,D,U,\eta_y),$$ where $A$ is the treatment, $D$ is the running variable, $U$ is unobserved confounding, and $\eta_y$ is unobserved heterogeneity. Notice that we decompose the unobservable into the component $U$ for which we will have a model, and the component $\eta_y$ which will be exogeneous conditional on $U$.
We will study the usual causal parameter in RDD, $$\tau_0= \mathbb{E}\mathopen{}\mathclose\bgroup\originalleft\{Y(1,D,U,\eta_y)-Y(0,D,U,\eta_y)\mid D = d^*\aftergroup\egroup\originalright\},$$ which reflects the average causal effect of moving from untreated to treated at the cutoff $d^*$, conditional on the running variable.
In contrast with standard approaches, we allow for discontinuities in the conditional distribution $f(u,\eta_y \mid d)$ at $d^*$, acknowledging that strategic behavior or sorting can lead to non-smooth densities near the cutoff. However, we require the density to remain strictly positive in a neighborhood of $d^*$ to maintain identifiability.
We employ placebo variables $(Z, W)$ to handle unobserved confounding. The full structural system is written as, $$ Y = Y(A, D, Z, U, \eta_y), \quad W = W(A, D, Z, U, \eta_w), \quad A = A(D, Z, U, \eta_a). $$ We summarize the invariances of the structural system in the next section. A graphical summary of the causal structure is previewed in Figure (ref).
The following assumptions identify $\tau_0$ even when the expected potential outcomes $d\mapsto \mathbb{E}\mathopen{}\mathclose\bgroup\originalleft\{Y(0, D, U, \eta_y) \mid D = d\aftergroup\egroup\originalright\}$ and $d\mapsto \mathbb{E}\mathopen{}\mathclose\bgroup\originalleft\{Y(1, D, U, \eta_y) \mid D = d\aftergroup\egroup\originalright\}$ are discontinuous at the cutoff $d=d^*$.
Throughout, we place assumptions in some neighborhood of the threshold $d^{*}$ defined as, $$ \mathcal{D}_-(\epsilon) = (d^* - \epsilon, d^*), \quad \mathcal{D}_+(\epsilon) = (d^*, d^* + \epsilon), \quad \mathcal{D}(\epsilon) = \mathcal{D}_-(\epsilon) \cup \mathcal{D}_+(\epsilon). $$
The existence of this limit, for the expectation of the treatment assignment $A$ given the running variable $D$, is standard in the RDD identification literature hahn2001identification,lee2010regression.
In a sharp design, $A$ is fully determined by $D$, implying no heterogeneity in treatment assignment (i.e., no $\eta_a$). In this case, Assumption (ref) is trivially satisfied.
We define $\mathcal{Z}(\epsilon)$ such that $\mathbb{P}\mathopen{}\mathclose\bgroup\originalleft(Z \in \mathcal{Z}(\epsilon) \mid D \in \mathcal{D}(\epsilon)\aftergroup\egroup\originalright) = 1$.
In summary, we have the following simplification local to the cutoff: $$Y = Y(A, D, U, \eta_y), \quad W = W(U, \eta_w), \quad A = A(D, \eta_a).$$
Causal consistency rules out network effects near the cut-off.
The main condition within placebo selection on observables is the implication that $Z\perp \!\!\! \perp \eta_y, \eta_w |D,U$: conditional upon the running variable and unobserved confounder, the placebo treatment is as good as random for the target outcome and placebo outcome.\footnote{By weak union, $\eta_{w} \perp \!\!\! \perp D, Z \mid U$ implies $\eta_{w} \perp \!\!\! \perp Z \mid D,U$.} In the sharp design, $Z\perp \!\!\! \perp \eta_a |D,U$ automatically holds; in the fuzzy design, it imposes that the placebo treatment is conditionally independent of the idiosyncratic aspect of treatment assignment. As shown in Appendix (ref), $Z\perp \!\!\! \perp \eta_y, \eta_a,\eta_w |D,U$ together with an additional continuity condition suffice for our argument to work. In the main text, to simplify the exposition, we strengthen $\eta_{w} \perp \!\!\! \perp Z \mid D, U$ to $\eta_{w} \perp \!\!\! \perp D,Z \mid U$ and thereby avoid the additional continuity.
Overlap ensures there is no stratum of unobserved confounding for which the treatment, running variable, or placebo treatment are deterministic. In the language of the RDD literature, we rule out complete control but allow imprecise control lee2010regression.
Exclusion encodes the idea that the placebo treatment does not directly cause the outcome of interest; instead, it affects the outcome via the running variable. Exclusion also encodes the idea that the placebo outcome is not caused by the treatment, running variable, or placebo treatment. Finally, exclusion imposes that the treatment is pinned down by only the running variable (and possibly idiosyncratic noise in the fuzzy setting).
In the running example, suppose class size $A$ is determined by grade enrollment size $D$, and we study its effect on student achievement $Y$ in year $t$. School administrator aptitude $U$ is unobserved, and strategic behavior around the cutoff can cause a jump in $f(d \mid u ,\eta_y)$ (and hence $f( u ,\eta_y \mid d)$), since strategic administrators who are aware of the class-enrollment cutoff $d^{*}$ may manipulate enrollment so that they exceed the threshold $d^{*}$ to open an additional class. Let $Z$ be the availability of students to transfer between grade levels (affecting $D$ but not $Y$ directly), and $W$ be achievement in year $t - 1$ (unaffected by $D$ or $Z$). Overlap implies school officials cannot perfectly control enrollment levels $D$. However, there is a jump in $d\mapsto f(d \mid u,\eta_y)$ that we wish to allow for and to deal with.
Assumption (ref) is vacuously satisfied by the sharp design.
In the fuzzy design, it is analogous to hahn2001identification: it implies that conditional on $D$, the randomness in the treatment assignment mechanism is unrelated to the idiosyncratic aspect of potential outcomes and the unobserved confounding. In other words, given an enrollment level $D$, whether a school has one or two classes does not depend on the heterogeneity in students' test scores or the administrator's underlying ability. This assumption in the fuzzy design may be relaxed by imposing homogeneity of effects; see Appendix (ref).
Assumptions (ref) and (ref) can be summarized by the results below:
Assumption (ref) weakens the standard continuity assumption of the RDD literature: continuity of potential outcomes conditional upon the running variable $D=d$ near the cutoff $d^*$. Instead, it requires continuity of potential outcomes conditional upon both $D=d$ and $U=u$.
While the density of student enrollment may jump at $d^{*}$, we anticipate that potential student achievement is continuous when fixing the administrator's aptitude $U=u$.
Lemma (ref) solves for the treatment effect $\tau_{0}$ in an interpretable way: for each value of the unobserved confounded $U=u$, conduct RDD conditional upon $U=u$, then average over $U$. In our running example, this is equivalent to averaging the treatment effect of a class size $(A)$ on test scores in year-$t$ $(Y)$ for schools with enrollment levels at $d^{*}$ over different administrator-ability levels $U$.
Clearly, this calculation is infeasible, since $U$ is unobserved. We need additional assumptions to identify $\tau_{0}$ in terms of observed variables.
Section (ref) established that the causal estimand $\tau_0$ can be written in terms of conditional expectations over unobserved confounding $(U)$. The key challenge now is to express $\tau_0$ in terms of observable variables. To do this, we adapt the “confounding bridge” approach tchetgen2020introduction, which uses auxiliary variables to solve an integral equation that recovers the influence of $U$ on $Y$.
The confounding bridge assumption states that, for fixed values of the running variable $(D)$ and unobserved confounder $(U)$, the conditional expectation of the outcome $(Y)$ can be expressed as a function of a proxy variable $(W)$ that is associated with $(U)$. The functions $h_+$ and $h_-$ act as weighting kernels that summarize how $W$ captures the influence of $U$ on $Y$ on either side of the threshold.
In the running example, for schools with the same number of enrolled students $(D = d)$, there exists a function $h_+(d - d^*, \cdot)$ such that we can recover expected student achievement using a reweighting of the distribution of test scores in year $t - 1$ $(W)$. The function $h_+$ does not depend on $U$ directly, but its integral with respect to the distribution of $W$ conditional on $U$ reconstructs the confounded conditional expectation of $Y$. This formulation allows us to bypass the unobserved $U$ by using an observed variable $W$ that is a proxy for $U$.
First we show that $h_0$ is a solution to an integral equation in terms of observed variables, though it may not be unique.
Lemma (ref) translates the bridge equation from a conditional expectation over $U$ to one over observed variables $Z$ and $W$. It shows that the same $h_0$ function also solves a second integral equation where both sides of the equation are estimable from the data.
Assumption (ref) is necessary to guarantee a unique solution to the equation in Lemma (ref) (factuals). Without this assumption, there could be multiple functions $h$ that satisfy the same moment condition, making point identification of $\tau_0$ more challenging.
Future work may build on bennett2022inference to give weaker conditions under which $\tau_{0}$ can still be identified without assuming completeness.
Lemma (ref) completes the identification argument by showing that the existence of a function $h_0$ that solves the observed-data version of the bridge equation implies the existence of a confounding bridge in the structural model. This allows us to write $\tau_{0}$ in terms of functions of observed variables.
Theorem (ref) provides a complete identification strategy---it shows how the discontinuity in potential outcomes conditional on unobserved $U$ can be corrected using an observable proxy $(W)$ and an auxiliary variable $(Z)$. The functions $h_+$ and $h_-$ effectively “invert” the confounding effect of $U$ using the relationship between $U$ and $W$. This inversion is unique when completeness holds and yields a point-identified causal estimand.
In the running example, we estimate the effect of class size on test scores in year-$t$ by using test scores in year $t - 1$ ($W$) as a proxy for unobserved administrator ability ($U$), and availability of students to distort enrollment counts ($Z$) as a local instrument. If the bridge functions $h_+$ and $h_-$ can be identified from observed data, then we can reconstruct the counterfactual outcomes at the cutoff and identify $\tau_0$.
Theorem (ref) identifies the RDD treatment effect by reweighting two confounding bridges $h_{-}(d, w)$ and $h_{+}(d, w)$. Recent literature to approximate these integral equations rely on kernel methods to non-parametrically estimate these confounding bridges. However, fan_variable_1992 show that the order of bias of these kernel estimators on boundary points ($d^{*}$) is high, and advocate the use of local linear regressions to estimate points on the boundary with less bias.
As our second contribution, we propose a local linear instrumental variable estimator, extending fan_variable_1992 to accommodate instruments. For tractability, we introduce a partially linear approximation to the confounding bridge in Assumption (ref). The final estimator decomposes into two intepretable terms; it adjusts the discontinuity in the target outcome by the discontinuity in the placebo outcome.
For clarity, we focus on the sharp design, i.e. the numerator in Theorem (ref). The fuzzy design is a straightforward extension, since the denominator in Theorem (ref) is a standard RDD estimand.
For simplicity, we also focus on the exactly identified case where $\dim(z)=\dim(w)$. Our results naturally extend to the overidentified case $\dim(z)\geq \dim(w)$ by standard techniques.
In line with the identification strategy presented in Section (ref), we estimate the treatment effect $\tau_{0}$ using a two-step procedure. First, we solve for the confounding bridges $h_{+}(\cdot, \cdot)$ and $h_{-}(\cdot, \cdot)$ at the cutoff $d^{*}$ using the moment condition from Lemma (ref). Then, invoking Theorem (ref), we compute the difference in their limits, $ \tau_{0} = \lim_{\epsilon \downarrow 0} \mathopen{}\mathclose\bgroup\originalleft[ \int h_{+}(\epsilon, w) \, d\mathbb{P}(w \mid D = d^{*}) - \int h_{-}(-\epsilon, w) \, d\mathbb{P}(w \mid D = d^{*}) \aftergroup\egroup\originalright]. $
To motivate our estimator, consider the class of partially linear functions,
where $g_{+} \colon \mathcal{D}_{+}(\epsilon) \mapsto \mathbb{R}$ and $g_{-} \colon \mathcal{D}_{-}(\epsilon) \mapsto \mathbb{R}$ are any continuous functions, and $\gamma_{+}$ plays the role of the best linear projection of $h_{+}(\cdot, \cdot)$ onto $W$ at $D=d^{*}$. Proposition (ref) below gives a detailed characterization. Within this class, the moment condition in Lemma (ref) becomes,
This expression closely resembles a partially linear instrumental variable moment condition, where $W$ plays the role of the endogenous regressor and $Z$ serves as the instrument. Therefore, we propose a local instrumental variable regression approach to estimate the best approximation to $h_{+}(\cdot, \cdot)$ near $d^{*}$.
To define our estimator, we introduce some notation. Let $H_{1} = \textrm{Diag}\mathopen{}\mathclose\bgroup\originalleft(1, h_{n}\aftergroup\egroup\originalright)$ be a transformation matrix of the kernel bandwidth. Let $\alpha_{+} = (\alpha_{+, 0}, \alpha_{+, 1})^{\top} = \mathopen{}\mathclose\bgroup\originalleft(g_{+}(0), g_{+}'(0)\aftergroup\egroup\originalright)^{\top}$ concatenate the nonlinear component of $h_+$ and its derivative at the cutoff. Finally, we concatenate all of the parameters to be estimated as $\nu_{+}^{\top} = (\alpha_{+, 0}, \alpha_{+, 1}, \gamma_{+}^{\top})$.
We propose what appears to be a new local linear objective function, which we call local instrumental variable regression. Its first order condition, which yields the estimator $\hat{\nu}_+$, is
where $\omega_{i, +} = \frac{1}{h_{n}} \mathbbm{1}\mathopen{}\mathclose\bgroup\originalleft(D_{i} \geq d^{*}\aftergroup\egroup\originalright) K\mathopen{}\mathclose\bgroup\originalleft(\frac{\mathopen{}\mathclose\bgroup\originalleft\lvert D_{i} - d^{*} \aftergroup\egroup\originalright\rvert}{h_{n}}\aftergroup\egroup\originalright)$ are local kernel weights, $R_{i, 1} = \mathopen{}\mathclose\bgroup\originalleft(1, \frac{D_{i} - d^{*}}{h_{n}} \aftergroup\egroup\originalright)^{\top}$ is a transformed vector of the running variable, and $Z_{i} \in \mathbb{R}^{\dim(w)}$ is a vector of placebo treatments. This empirical moment is clearly a sample analogue of the population moment implied by equation (ref). The left limit objects are analogous.
We then use the coefficients estimated by local instrumental variable regression to construct our estimator of the treatment effect. The appropriate formula is immediate from Theorem (ref) and equation (ref). Formally, the approximate treatment effect and its estimator are
where $\hat{\beta}_{+, 0}^{w}$ estimates the right-limit $\beta_{+, 0}^{w}=\lim_{\epsilon \downarrow 0} \mathbb{E}\mathopen{}\mathclose\bgroup\originalleft\{W_{i} \mid D = d^{*} + \epsilon\aftergroup\egroup\originalright\}$ using the local linear estimator of hahn2001identification: $$H_{1} \hat{\beta}_{+}^{w_{j}} = \operatorname*{arg\,min}_{\beta} \sum_{i = 1}^{n} \omega_{i, +} (W_{i, j} - R_{i, 1}^{\top}\beta )^{2}.$$ Then, we let $\hat{\beta}_{+, 0}^{w} = \mathopen{}\mathclose\bgroup\originalleft(\hat{\beta}_{+, 0}^{w_{1}}, \dots, \hat{\beta}_{+, 0}^{w_{\dim(w)}}\aftergroup\egroup\originalright)^{\top} \in \mathbb{R}^{\dim(w)}$ concatenate the first components of these estimated coefficients. The left limit objects are analogous.
Section (ref) below shows that this estimator is consistent for $\tau_{\text{pdd}}$. After bias correction, which we defer to Section (ref), it is also asymptotically normal at the familiar rate of $n^{-2/5}$.
For any finite sample size, our proposed estimator is numerically equivalent to a highly interpretable procedure: take the standard RDD estimator based on discontinuity in the outcome ($Y$), and add an adjustment term based on the discontinuity in the placebo outcome ($W$).
The RDD literature, summarized recently by e.g. cattaneo_practical_2024, advocates for testing whether $\hat{\tau}_{\text{rdd}}^{w}$ is close to zero as a diagnostic for whether $f(d \mid u,\eta_y)$ (or, relatedly, $f(u,\eta_y \mid d)$) changes discontinuously at $d^{*}$. Our contribution is to show that this discontinuity is an essential ingredient for an adjustment term. If $\hat{\tau}_{\text{rdd}}^{w}$ is indeed negligible, our adjustment term is zero and our estimator $\hat{\tau}_{\text{pdd}}$ simplifies to the standard RDD estimate $\hat{\tau}_{\text{rdd}}^{y}$.
Within our adjustment term, the placebo outcome discontinuity ($\hat{\tau}_{\text{rdd}}^{w}$) is multiplied by weights ($\hat{\gamma}_{-}$).
These weights are increasing in the correlation between the residualized outcome ($Y^{\perp}$) and the placebo treatment ($Z$). Intuitively, if this correlation is stronger, then there is more unobserved confounding on the target outcome, and hence we increase the weights.
The weights are decreasing in the correlation between the residualized placebo outcome ($W^{\perp}$) and the placebo treatment ($Z$). Intuitively, if this correlation is weaker, then the placebo outcome is a weaker proxy for the unobserved confounding, and hence we compensate by increasing the weights on its estimated discontinuity.
Proposition (ref) clarifies the danger of a very weak proxy, which would lead to numerical instability in this final aspect of the weights. Future work may study the behavior of our estimator when using a weak proxy.
As our third contribution, we study the large sample properties of the estimator in Proposition (ref). We prove it is consistent. Then, we prove that it is asymptotically normal at the familiar rate $n^{-2/5}$ after bias correction.
While prior work used placebo outcomes and predetermined covariates to detect assumption violations, we propose a method to directly correct for them and thereby recover $\tau_{0}$. Below we state sufficient conditions for consistency.
Let $\mu_{+, y}(d - d^{*}) = \mathbb{E}\mathopen{}\mathclose\bgroup\originalleft\{Y \mid D = d\aftergroup\egroup\originalright\}$ for $d \geq d^{*}$ and $\mu_{-, y}(d - d^{*}) = \mathbb{E}\mathopen{}\mathclose\bgroup\originalleft\{Y \mid D = d\aftergroup\egroup\originalright\}$ for $d < d^{*}$. Similarly, define $\mu_{+, w_{j}}(d - d^{*})$, $\mu_{-, w_{j}}(d - d^{*})$, $\mu_{+, z_{j}}(d - d^{*})$, and $\mu_{-, z_{j}}(d - d^{*})$..
Let $\sigma^{2}_{+, y}(d - d^{*}) = \mathrm{Var}\mathopen{}\mathclose\bgroup\originalleft(Y \mid D = d \aftergroup\egroup\originalright)$ for $d \geq d^{*}$ and $\sigma^{2}_{-, y}(d - d^{*}) = \mathrm{Var}\mathopen{}\mathclose\bgroup\originalleft(Y \mid D = d \aftergroup\egroup\originalright)$ for $d < d^{*}$. Similarly, define $\sigma^{2}_{+, w_{j}}(d - d^{*})$ and $\sigma^{2}_{-, w_{j}}(d - d^{*})$.
The regularity conditions in Assumption (ref) are variations of standard assumptions in the RDD literature and are implied by those in calonico_regression_2019.
Assumption (ref) is satisfied by the window kernel $K\mathopen{}\mathclose\bgroup\originalleft(x\aftergroup\egroup\originalright) = \mathbbm{1}\mathopen{}\mathclose\bgroup\originalleft(x \leq 1\aftergroup\egroup\originalright)$ and the triangle kernel $K\mathopen{}\mathclose\bgroup\originalleft(\cdot\aftergroup\egroup\originalright) = \mathopen{}\mathclose\bgroup\originalleft(1 - x\aftergroup\egroup\originalright) \mathbbm{1}\mathopen{}\mathclose\bgroup\originalleft(x \leq 1\aftergroup\egroup\originalright)$, which are the most frequently used kernels in the RDD literature. It also allows for the Gaussian kernel $K\mathopen{}\mathclose\bgroup\originalleft(x\aftergroup\egroup\originalright) = \frac{1}{\sqrt{2\pi\lambda}}\exp{-\frac{x^{2}}{2\lambda}}$. Under these conditions we can characterize the probability limit of our local instrumental variable estimator for any vanishing bandwidth $1\gg h_n \gg n^{-1}$.
Building on Proposition (ref) (estimator equivalence), which characterizes the finite sample estimate, Proposition (ref) characterizes the asymptotic limit of the local estimator. Now, at the population level, we show that if the placebo outcome does not jump at the cutoff ($\tau^{w}_{\text{rdd}} = 0$), then the local instrumental variable estimand reduces to the traditional RDD estimand.
The population limit of the weights admits a similar interpretation as before. It also sheds light on the nature of our approximation. In particular, Proposition (ref) shows that $\gamma_{-}$ can be interpreted as a linear projection of $Y$ onto the distribution of $W$, instrumenting with $Z$, when $D$ approaches the cutoff from the left. In other words, $\gamma_{-}$ is the parameter which best approximates the confounding bridge in Lemma (ref) (factuals) as a partially linear function of the placebo outcome $W$.
If the confounding bridge $h_{-}(0, w)$ is very nonlinear in $W$ then this approximation can be poor. However, we advocate for this approach because full nonparametric estimation over $(D_{i},W_{i})$ would require us to uniformly estimate $h_{-}(0, w)$ as a function of $w$ to characterize the asymptotic bias of the local instrumental variable estimator. There would be $\dim(w)$ additional kernels and bandwidths in the estimator. As discussed by calonico_regression_2019, localizing weights for many variables would suffer from a curse of dimensionality, and hence empirical applications would become challenging. Our estimator does not suffer from this curse of dimensionality, and is therefore practical for empirical researchers. Future work may attempt to overcome these challenges, perhaps by placing additional structure on the effective dimension or sparsity of $W$ singh_kernel_2020,noack2021flexible.
Corollary (ref) shows that when the confounding bridge lies in the class of partially linear functions in Equation (ref), then the local instrumental variable estimator is consistent for the treatment effect $\tau_{0}$.
As shown in Proposition (ref), our estimator can be decomposed into a sum of two components: (i) the RDD estimate of the outcome discontinuity $\hat{\tau}_{\text{rdd}}^{y}$, and (ii) the RDD estimate of the placebo outcome discontinuity, appropriately weighted as $(\hat{\tau}_{\text{rdd}}^{w})^{\top}\hat{\gamma}_{-}$. A key insight from calonico_robust_2014 is that, when the bandwidth $h_n$ is selected to optimize mean squared error (MSE), these discontinuity estimators suffer from non-negligible bias in their sampling distributions. This bias persists in large samples and leads to invalid inference, particularly when constructing confidence intervals.
To address this issue, we extend the robust bias correction technique developed by calonico_robust_2014. In particular, we construct bias-corrected estimators that remain consistent and asymptotically normal, even when using MSE-optimal bandwidths. Specifically, we define the bias-corrected version of the local instrumental variable estimator as,
where $\hat{\tau}_{\text{rdd}}^{y\text{,bc}}$ and $\hat{\tau}_{\text{rdd}}^{w\text{,bc}}$ are bias-corrected estimators for the discontinuities of $Y$ and $W$, respectively. In particular, we explicitly estimate and subtract the first-order bias terms associated with the local linear estimates $\hat{\tau}_{\text{rdd}}^{y}$ and $\hat{\tau}_{\text{rdd}}^{w}$. The explicit forms of these bias corrections are provided in Appendix (ref).
Theorem (ref) establishes the asymptotic normality of the bias-corrected local instrumental variable estimator under the commonly used MSE-optimal bandwidth rate of $h_n\asymp n^{-1/5}$ (imbens_regression_2008). The bias correction bandwidth may be of the same order, i.e. $b_n\asymp h_n$.
More generally, Theorem (ref) tolerates any vanishing bandwidths in a wide range, i.e. any $n^{-1/7}\gg h_n\gg n^{-1}$ and any $b_n \gtrsim h_n$.
In practice, we recommend a data-driven bandwidth choice $\hat{h}_n$ along the lines of e.g. imbens2012optimal,calonico2018effect,calonico_optimal_2020. Our analysis extends accordingly. This choice is easily implemented with existing software for RDD.
Theorem (ref) justifies the use of standard Wald-type inference for $\hat{\tau}_{\text{pdd}}^{\text{bc}}$, e.g. the construction of confidence intervals and hypothesis tests.
This paper develops a new identification and estimation framework for regression discontinuity designs when the standard continuity of potential outcomes assumption fails due to unobserved confounding. We show that by leveraging a placebo treatment and a placebo outcome, it is possible to recover the causal parameter even when the running variable's distribution exhibits discontinuities at the threshold. Our identification argument relies on conditional continuity given an unobserved confounder, and it employs the concept of a confounding bridge to encode the relationship between the unobserved confounder and its observed proxy. Under a completeness condition, we show that the integrated bridge function is identified from observed data, and the treatment effect at the cutoff can be recovered using integrals of observable variables.
To operationalize this strategy, we propose a local instrumental variable estimator that approximates the confounding bridge using a partially linear specification. We demonstrate that the estimator decomposes into a standard RDD component and a correction term driven by discontinuities in the placebo outcome. This decomposition yields an interpretable adjustment that allows researchers to correct for violations of the standard RDD assumptions, rather than discarding such designs altogether. We establish conditions for consistency and bias-corrected inference, allowing optimal bandwidths. Our results extend the applicability of RDD methods to settings with strategic behavior, providing researchers with a tractable and theoretically grounded approach for handling unobserved confounding near the cutoff.