Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
53,750 characters · 11 sections · 18 citation commands
Treatment Effects in the Regression Discontinuity Model with Counterfactual Cutoff and Distorted Running Variables
Regression Discontinuity Designs (RDD) have become one of the most influential empirical strategies in applied economics, in large part due to their transparent identification of causal effects in settings where treatment assignment is determined by a cutoff rule. lee2008randomized provides a foundational interpretation of the sharp RDD by showing that, under continuity of potential outcomes at the threshold, the discontinuity in observed outcomes identifies a local average treatment effect (LATE) for units at the margin of indifference. Importantly, Lee's framework allows the running variable to be influenced by individuals' anticipation of treatment, as long as such sorting does not generate a discontinuity in potential outcomes at the cutoff. In this sense, the RDD estimand can be understood as a characteristics–weighted average treatment effect for units whose underlying attributes place them near the policy threshold.
While this framework is extremely powerful, it is limited in several important dimensions when the research or policy question extends beyond the immediate neighborhood of the cutoff. First, the classical RDD focuses exclusively on the direct effect of treatment on outcomes by conditioning on the running variable, without accounting for the possibility that the running variable itself may be endogenously distorted by the availability or anticipation of treatment. Second, identification is restricted to a local parameter at the cutoff, whereas policy makers may be interested in causal effects for individuals away from the boundary. Third, standard RDD methods are tied to the observed cutoff, leaving open the policy-relevant question of how treatment effects would differ under a counterfactual threshold.
To address these three limitations, we develop a new framework that augments the standard RDD setup with a two-period, two-group structure. We consider periods $t\in\{0,1\}$ and groups $Z_i\in\{0,1\}$. For individuals with $Z_i=0$, treatment is never available in either period. For individuals with $Z_i=1$, treatment is unavailable in period $t=0$ but becomes available in period $t=1$ depending on their running variable. We introduce a potential-outcome framework for the running variable,$ R_{it} \;=\; R_{it}^0(1-Z_i) \;+\; Z_i\Big[ R_{it}^0\mathbbm{1}(t=0) + R_{it}^1\mathbbm{1}(t=1)\Big],$ which allows the running variable to be distorted from $R_{it}^0$ to $R_{it}^1$ when treatment becomes available. By defining the total policy effect on individual $i$ in period $t=1$ as $ Y_{i1}^1(R_{i1}^1)-Y_{i1}^0(R_{i1}^0),$ our framework explicitly incorporates both the direct treatment effect and the indirect effect arising from the endogenous distortion of the running variable. This directly resolves the first limitation of the prior literature.
To address the second limitation—extrapolating treatment effects away from the cutoff—we leverage the $Z_i=0$ group as a control group. Specifically, we impose a conditional parallel trend assumption on untreated potential outcomes between the two groups, which allows us to recover the counterfactual untreated outcome $Y_{i1}^0(R_{i1}^0)$ for the $Z_i=1$ units. This strategy differs fundamentally from the two main existing extrapolation approaches: the local policy invariance method of dong2015identifying, which extrapolates using local derivatives near the cutoff, and the conditional-independence-based extrapolation of angrist2015wanna, which conditions on covariates to remove dependence between potential outcomes and the forcing variable. In contrast, our approach uses overtime variation and between-group comparisons to build the necessary counterfactual.
To address the final limitation—evaluating treatment effects under counterfactual cutoffs—we assume that the distortion of the running variable under a counterfactual policy threshold exhibits a translation-invariant structure relative to the observed distortion. Combined with the conditional parallel trend assumption, this allows us to predict both the counterfactual distribution of the distorted running variable and the implied counterfactual treatment effects at any proposed cutoff.
Building on these identification ideas, we construct a nonparametric estimator of the counterfactual total policy effect by combining kernel estimators for (i) the conditional mean of untreated potential outcomes, (ii) the conditional density of the running variable distortion, and (iii) the conditional mean of realized treated outcomes. Following the procedure outlined in Section (ref), we derive the asymptotic linear expansion of the estimator and show that, despite the multiple nonparametric components, the estimator converges at the $\sqrt{n}$ rate. This fast rate arises because the estimand is an average counterfactual policy effect that depends only on the treated subsample, and because undersmoothing renders the nonparametric bias asymptotically negligible. We evaluate the finite-sample performance through simulations and find that, due to the persistence of finite-sample bias, a two-scale bias-reduction estimator performs substantially better. Both empirical and $t$-bootstrap procedures deliver accurate confidence interval coverage in moderately sized samples.
We illustrate the usefulness of our approach by applying it to the setting of grembi2016fiscal, who study the effect of Italy's Domestic Stability Pact (DSP) rules on municipal fiscal behavior. In their setting, municipalities below a population threshold faced different fiscal rules than those above it, and population itself may respond endogenously to fiscal incentives. Section (ref) describes the institutional context and data. Our result shows that DSP decreases the municipal deficit for counterfactual cutoff set above 4,600 and can increase deficit if set below 4,600.
\paragraph{Literature} This paper contributes to the literature on RDD by extending the identification and estimation of treatment effects beyond the observed cutoff. lee2008randomized establishes the modern interpretation of the RDD as identifying a local average treatment effect under continuity of potential outcomes. dong2015identifying show how marginal changes in the policy threshold can identify the marginal threshold treatment effect under a local policy invariance assumption. angrist2015wanna develop a conditional-independence-based strategy for extrapolating RDD effects away from the cutoff using covariates.
We also contribute to the growing literature combining Difference-in-Differences ideas with RDD designs. grembi2016fiscal use a parallel trend structure to net out confounding mayoral wage policies that coincide with the same population cutoff, but their focus remains on the direct treatment effect rather than the full policy effect including the distortion of the running variable. Our framework incorporates both components and thus enriches the set of policy-relevant estimands available in RDD settings.
Consider a two-period model $t=0,1$ and a regression discontinuity design that happens in period $t=1$. Let $Y_{it}$ be the observed outcome variable for individual $i$ in period $t$, let $R_{it}$ be the running variable that will determine the treatment status in period $t=1$. Let $Z_i\in\{0,1\}$ be a binary variable that determines whether the individual is subject to RDD in period $t=1$. The treatment status $D_{it}=0$ for all individuals $i$ and $t=0$, and $D_{it}=\mathbbm{1}(R_{it}\ge c,Z_{i}=1)$ for period $t=1$. The observed outcome variable is generated by the potential outcome model:
The running variable $R_{it}$ can also be influenced by the treatment due to the firm's endogenous behavior of choosing the running variable. The econometrician cannot observe $Y_{it}^1(R_{it})$ and $Y_{it}^0(R_{it})$ simultaneously and the potential outcome can depend on $R_{it}$. The running variable can also be influenced by the potential to be treated since individuals may strategically choose the running variable when the treatment is potentially available to them. As a result, for the control group $Z_i=0$, we always observe the policy undistorted potential running variable $R_{it}^0$, while for the $Z_i=1$ group, we observe $R_{it}^0$ for $t=0$ period and the distorted potential running variable $R_{it}^1$ for the $t=1$ period. The econometrician may also observe a set of covariates $X_{it}$, but we leave the discussion to extensions. We also impose the following exogeneity condition for the potential outcome. The potential outcome framework in (ref) is general: If we impose $E[Y_{it}^d(R_{it})|X_{it}=x]\equiv E[Y_{it}^d|X_{it}=x]$ for $d=0,1,$ then we get the model in angrist2015wanna.
The literature has focused on the local average treatment effect: \[ S(r_1,r_2)=E[Y_{it}^1(R_{it})-Y_{it}^0(R_{it})|Z_i=1, c=r_1, R_{it}=r_2,t=1], \] which is the effect of treatment with treatment cutoff $c=r_1$ if an individual's running variable is fixed at $r_2>r_1$ dong2015identifying. Under suitable continuity conditions, we can identify the local treatment effect at the treatment cutoff:
The local average treatment effect is not the only object that is interesting to the policy makers: it does not consider the policy effect via the distortion of the potential running variable. To take into account the distortion of running variables, we can also define the local effects of policy \[ T(r_1,r_2)=E[Y_{i1}^1(R^1_{i1})-Y_{i1}^0(R_{i0}^0)|Z_i=1,c=r_1,R^1_{i1}=r_2]. \] The total treatment effect with cutoff $c=r_1$ focuses on a person with post-treatment running variable $R_{i1}=r_2$, and compares the treated outcome $Y_{i1}^1(R^1_{i1})$ with its untreated potential outcome evaluated at the undistorted potential running variable $Y_{i1}^0(R_{i0}^0)$. The difference $T(r_1,r_2)-S(r_1,r_2)$ is the indirect effect of policy via the distortion of the running variable.
The policy maker is interested in both evaluating policy effect under the current policy cutoff $c$, as well as the counterfactual policy effect when the policy cutoff is chosen at a different value $c^{cf}>c$. When evaluating the treatment effect at each cutoff level, the policy maker also wants to know the local direct and indirect effect on the treated unit, i.e., $S(c',r)$, $T(c',r)$ for $r\ge c'$. For a policy maker who wants to know the counterfactual overall population treatment effect on the treated, we define the quantity \[ATT(c^{cf})\equiv E[T(c^{cf},R_{it})|R_{it}\ge c^{cf}],\] which is policy relevant and can inform the policy maker about the overall benefits of policy.
To infer the treatment effect in a counterfactual cutoff location, we need to consider several steps: First, we use the over-time variation of $(Y_{it},R_{it})|Z_i=0$ to infer the trend of $Y_{it}^0(R_{it}^0)$. When we derive the trend of $Y_{it}^0(R_{it}^0)$, the potential outcome model (ref) implies the correct conditional parallel trend assumption to impose. Second, for $Z_i=1$ and $t=1$ period, we can identify the realized distributions of $(Y_{it}^0,R^1_{it})|Z_i=1,t=1,R_{it}<c$ and $(Y_{it}^1,R^1_{it})|Z_i=1,t=1,R_{it}\ge c$. The conditional distribution $R_{i1}^1|R_{i1}^0,Z_i=1$ informs us how the individuals select the distorted running variable given their past undistorted running variable. Last, to infer the treatment effect in the new cutoff place, we use the difference between the distribution $Y_{it}|Z_i=1,t=1$ and $Y_{it}^0|Z_{i}=1,t=1$ to inform us the treatment effect conditional on the running variables, and then impose additional translation invariance in the running variable distortion to inform us of the distortion of running variables under the new policy cutoff. \paragraph{Identifying the counterfactual outcome for $D_{it}=1$} We first impose the conditional parallel trend assumption in the untreated potential outcome conditional on the running variable:
We choose $R_{i0}^0$ as the conditioning variable in Assumption (ref) based on two reasons: first, we cannot observe $R_{i1}^0$ for the $Z_i=1$ group and $R_{i1}^1$ for the $Z_i=0$ group, so it is not possible to add these running variables in the parallel trend assumption; second, as we will see later, the distortion of running variable can be characterized as the conditional distribution of $R_{i1}^1$ given $R_{i0}^0$, and conditioning on $R_{i0}^0$ will deliver the right conditional unobserved potential outcome expectation so that we can take into account the distortion of the running variables.
\paragraph{Identifying the distortion of running variables} When individuals can potentially have access to the treatment, the running variable may be intentionally distorted to select into or out of treatment. The way that the running variables are distorted is characterized by the following conditional distribution: \[ F(R_{i1}^1\le r_1 |R_{i0}^0=r_0,Z_i=1). \] To characterize the counterfactual distribution of the distorted running variables, we impose the following translation invariance in the selection behavior when the counterfactual cutoff is $c^{cf}>c$:
Assumption (ref) has the following interpretation: For treated units, its distorted running variable $R_{i1}^1$ depends on its untreated potential running variable $R_{i0}^0$. When the cutoff for treatment is moved to a new location $c^{cf}=c+\Delta c$, the selection of $R_{it}^1(c^{cf})$ takes a parallel shift.
We notice that the counterfactual local average policy effect $T(c^{cf},r)$ is identified for all $r\ge c^{cf}\ge c$.
The (ref) in Theorem (ref) summarizes our identification strategy: The $E[Y_{i1}|Z_i=1,R_{i1}=r]$ corresponds to the mean of treated potential outcome with the distorted running variable, the integrand $E[Y_{i1}^0(R_{i1}^0)|Z_i=1,R_{i0}^0=r_0]$ is the untreated potential outcome under undistorted running variable, and the distribution $F(R_{i0}^0=r_0|R^{1}_{i1}(c^{cf})=r)$ characterizes the distortion of the running variable in the counterfactual policy cutoff.
Last, we can identify the treatment effect on the treated (ATT) using the counterfactual distribution of $R_{it}^1(c^{cf})$.
In the estimation procedure below, we will follow the identification formula in (ref) and (ref) to construct an estimator for the conditional mean and density components and assemble them to construct the final estimators for (ref) and (ref) .
In this section, we outline the nonparametric kernel estimation procedures for the key conditional densities and expectations required to compute the counterfactual treatment effects discussed in Section 2. Let \( K(\cdot) \) be the Epanechnikov kernel function, let $K_h(\cdot)\equiv K(\cdot/h)$ denote the bandwidth normalized kernel function, and let \( h_n, b_n \) correspondingly denote the bandwidth sequences for the $(R_{i0},R_{i1})$.
\paragraph{Estimating the Marginal Distribution $f_0(r_0):=f(R_{i0}^0=r_0|Z_i=1)$} This marginal density is later used to infer the counterfactual distribution of $R_{it}^1$ with selection behavior. Note that, for $t=0$, $R_{i0}^0=R_{i0}$, so we can estimate $f_0(r_0)$ by:
\paragraph{Estimating the Conditional Mean \( m_0(r_0):=\mathbb{E}[Y_{i0}^0(R_{i0}^0) \mid R_{i0}^0 = r_0, Z_i = 1] \)} This expectation is identified from the pre-treatment period \( t = 0 \) in the treated group by realizing that $Y_{i0}^0(R_{i0}^0)=Y_{i0}$ and $R_{i0}^0=r_0$. The corresponding estimator is:
Here \( h_n \) denotes the bandwidth for the running variable in the $Z_i=1$ group.
\paragraph{Estimating the Mean Difference \(g(r_0):=\mathbb{E}[Y_{i1}^0(R_{i1}^0) - Y_{i0}^0(R_{i0}^0) \mid R_{i0}^0 = r_0, Z_i = 0]\)} This conditional mean difference quantity is identified from the $Z_i=0$ group by realizing that $Y_{i1}^0(R_{i1}^0)=Y_{i1}$ and $Y_{i0}^0(R_{i0}^0) =Y_{i0}$ for the $Z_i=0$ individuals. Define the difference \( \Delta Y_i = Y_{i1} - Y_{i0} \), and estimate:
Here \( h_n \) denotes the bandwidths for \( R_{i0} \).
\paragraph{Estimating the Conditional Density \( f_{1|0}(r_1|r_0):=f(R_{i1}^1 \mid R_{i0}^0, Z_i = 1) \)} This density captures selection behavior in the treated group. Using observed running variable \( R_{i1} \) to replace \( R_{i1}^1 \) and $R_{i0}$ to replace $R_{i0}^0$ for the $Z_i=1$ group, we estimate:
The marginal density of counterfactual running variable $R_{i1}^1(c^{cf})$ can be calculated as \[ \widehat{f}^{cf}(r_1)=\int_{r_0} \widehat{f}_{1|0}(r_1-\Delta c|r_0-\Delta c) \hat{f}_0(r_0) dr_0 \]
\paragraph{Estimating the Conditional Mean $m_1(r):=E[Y_{i1}|R_{i1}=r,Z_i=1]$} This object corresponds to the conditional mean $E[Y^1_{i1}|R^1_{i1}=r,Z_i=1]$ if $r\ge c$. Since for the $Z_i=1$ group, we only observe $Y_{i1}^1$ for $R_{i1}\ge c$, we have a boundary issue for the nonparametric estimation. Therefore, we need to use the boundary kernel estimation. Let $\widehat{m}_1(r)$ be the nonparametric estimator such that
where the boundary kernel \(K^{\mathrm{bd}}_{r,b_n}(u)\) adjusts the moments of the base kernel \(K(u)\) on the truncated support \(\{u \ge (c-r)/b_n\}\): \[ A_0(r)=\int_{(c-r)/b_n}^{\infty} K(u)\,du,\qquad A_1(r)=\int_{(c-r)/b_n}^{\infty} u\,K(u)\,du,\qquad A_2(r)=\int_{(c-r)/b_n}^{\infty} u^2\,K(u)\,du, \] \[ a(r)=\frac{A_2(r)}{A_0(r)A_2(r)-A_1(r)^2},\qquad b(r)=-\,\frac{A_1(r)}{A_0(r)A_2(r)-A_1(r)^2}, \] \[ K^{\mathrm{bd}}_{r,b_n}(u) = \bigl[a(r)+b(r)\,u\bigr]\, K(u)\, \mathbbm{1}\!\left\{u\ge \tfrac{c-r}{b_n}\right\}. \] Here \(c\), the policy cutoff value, is the lower boundary of the support of \(R_{i1}^1\), and \(\tau>0\) (e.g.\ \(\tau=1\)) determines the width of the boundary region.
\paragraph{Estimating the Counterfactual Total Policy Effect} Given the nonparametric estimators in (ref)- (ref), we estimate the counterfactual conditional total policy effect
For the overall treatment effect on the treated, we can estimate by
where $\widehat{F}^{cf}(c^{cf})=\int_{r\ge c^{cf}} \widehat{f}^{cf}(r) dr$.
Assumption (ref)-(ref) are standard nonparametric estimation assumptions on the nonparametric estimation. The uniformly bounded density condition in Assumption (ref) is important to derive the uniform convergence of the conditional mean and conditional density to the population. Assumption (ref) allows us to derive the uniform bias from kernel estimation. Assumption (ref) is later used to derive the asymptotic normality of the estimator, and shows that the estimation variation from both $Z_i=0$ and $Z_i=1$ group matters asymptotically, which suits the common empirical context that researchers will encounter.
The asymptotic linear expansion shows several things: first, the optimal MSE for the $\widehat{T(c^{cf},r)}$ is $n^{-4/5}$ by choosing $\delta=0$. Note that even if our estimator involves conditioning on 2-dimensional variables $R_{i0},R_{i1}$, the convergence rate is the same as the standard 1-dimensional nonparametric kernel estimation, because we can undersmooth the bandwidth for $h_n$ and exploit the smoothing in the integration in the estimator; second, the linear expansion suggests the asymptotic normality of the estimator, and a fast multiplier bootstrap method for inference on the conditional total policy effect.
We achieve $\sqrt{n}$ convergence for the $\widehat{ATT}(c^{cf})$ even if we use nonparametric density estimation in the process. This is because we use the undersmoothing in the bandwidth. The ratio $ \frac{f_0(r)}{\tilde{f}_0(r)}$ reflects the covariate $R_{i0}$ density shift across the two groups.
A consistent estimator of $V$ can be found by plugging in consistent estimators to first estimate a sample $\hat{\bm{\xi}}_i$, and then find the covariance matrix of $\hat{\bm{\xi}}_i$.
If we rely on Corollary (ref) to do inference on the $ATT$, then we need to estimate $\bm{\xi}_i$, which can be complicated to implement as we need to estimate multiple nonparametric objects. We now propose the bootstrap estimator which can be easier to implement at the cost of slightly more computational power.
We simulate two groups of units, indexed by $Z_i \in \{0,1\}$, each observed over two periods $t=0,1$. For the control group ($Z_i=0$), we observe the undistorted potential running variable $R_{it}^0$ at both periods, while for the treated group ($Z_i=1$), we observe $R_{i0}^0$ at $t=0$ and the distorted running variable $R_{i1}^1$ at $t=1$ due to policy distortion of the running variable.
\paragraph{Running variable dynamics.}
where the mean shift in $R_{i1}^1$ reflects endogenous distortion when the treatment is potentially available. The simulation setting of $R_{it}$ is consistent with the translation invariant selection Assumption (ref).
\paragraph{Potential outcome functions.} For the untreated potential outcome $Y_{it}^0(R_{it}^0)$, the conditional mean functions are defined as follows.
For the control group ($Z_i=0$):
For the treated group ($Z_i=1$), the untreated potential outcomes in the first period and its counterfactual path in the second period satisfy:
The conditional trend function for the untreated potential outcome is given by $2\left(3\cos^2\!\left(\frac{\pi}{60}(r-50)\right) + 1 \right)$. The unobserved potential outcome $Y_{i1}^0$ for the $Z_i=1$ group is explicitly simulated so that we can take the difference $Y_{i1}^1-Y_{i1}^0|Z_i=1$ to calculate the total policy effect for individual $i$.
The realized post-treatment outcomes for the $Z_i=1$ group in period 1 are
\paragraph{Cutoff and counterfactual cutoff choices.} For the $Z_i=1$ group, we use $c=60$ as the actual policy cutoff so that for $R_{it}\ge 60$, we observe $Y_{i1}^1$ as $Y_{i1}$ for the $Z_i=1$ group. The counterfactual cutoff is chosen at $c^{cf}=63$.
\paragraph{Performance of the Original Estimator.} Each cell under the original columns in Tables (ref) and (ref) corresponds to the mean of ATT estimator obtained using a conventional estimator with bandwidth \[ h = \text{(row constant)} \times n^{-x}\sigma_R, \] where the exponent $x \in \{0.28, 0.30, 0.32\}$ is indicated by the column header, and the row constant is given by the leftmost row label ($1/\sqrt{2}$, $1$, or $\sqrt{2}$). For instance, in the $n=1000$ table, the value $5.980$ under the $n^{-0.28}\sigma_R$ column and the $1/\sqrt{2}$ row represents the estimate obtained using the bandwidth $h = \frac{1}{\sqrt{2}} \times 1000^{-0.28}\sigma_R$. The true total policy effect is $\text{ATT}^{true} = 6.195$. In both $n=1000$ and $n=2000$ cases, all original estimates lie below the true value, indicating negative bias. The bias of the estimators reduces with the choice of bandwidth, but is still significant when we use the smallest bandwidth $n^{-0.32}\sigma_R/\sqrt{2}$. This indicates that the asymptotic bias $\sqrt{nh^2}$ is not fully washed away in the finite sample.
\paragraph{Bias-Reduction Method.} The bias-reduced columns implement a two-scale bias-correction scheme of the form \[ \widehat{ATT}_{\text{BR}}(h) = 2\widehat{ATT}(h) - \widehat{ATT}(\sqrt{2}\,h), \] where $\widehat{ATT}(h)$ denotes the original estimator using bandwidth $h$, and we suppress the dependence on the counterfactual cutoff $c^{cf}$ in the following notation. This linear combination removes the leading-order bias term $O(\sqrt{n}h^2)$ in the estimator and leaves us with the higher-order bias $O(\sqrt{n}h^4)$ jones1995simple.
Empirically, the bias-reduced values in both tables are substantially closer to the true ATT ($6.195$) than the original estimates. For example, in the $n=1000$ table under $n^{-0.28}\sigma_R$ and $1/\sqrt{2}$, the original estimate is $5.980$ (bias $\approx -0.215$) while the bias-reduced estimate is $6.240$ (bias $\approx 0.045$). The same improvement holds across other bandwidths and constants, and the performance further improves for $n=2000$. Overall, the results confirm that the two-scale correction effectively mitigates the leading bias term, producing estimates with smaller absolute bias across different bandwidths and sample sizes.
As a result, we recommend using the $\widehat{ATT}_{\text{BR}}(h)$ in relatively small samples to avoid bias in the estimator.
\paragraph{Bootstrap Adjustment.} Since we use the bias-reduced estimator $\widehat{ATT}_{\text{BR}}(h)$, we correspondingly change the bootstrap procedure and let \[ \widehat{ATT}^{*,b}_{BR}(h)=2\widehat{ATT}^{*,b}_{BR}(h)-\widehat{ATT}^{*,b}_{BR}(\sqrt{2}h), \] where $\widehat{ATT}^{*,b}_{BR}(h)$ is estimated from the $b$-th bootstrap sample, and construct the confidence interval using the $\alpha/2$ and $1-\alpha/2$ quantile of $2\widehat{ATT}_{BR}(h)-\widehat{ATT}^{*,b}_{BR}(h)$ as the confidence interval. Alternatively, we can use the t-bootstrap to calculate the variance of the bootstrap estimator $\sigma_{ATT,BR}^*$ and construct the confidence interval as $[\widehat{ATT}_{BR}(h)\pm 1.96\sigma_{ATT,BR}^*]$.
The validity of the bootstrap is ensured by looking at the linear expansion in Theorem (ref) and Proposition (ref). Let $\xi_i(h)$ denote the leading influence term in Theorem (ref) under bandwidth of $h$, then, fixing and suppressing the counterfactual cutoff notation $c^{cf}$, we have \[
\] where $W_{i}^{*,b}$ is the resampling weights for observation $i$. The validity of the bootstrap is established as $\left[ 2\hat{\xi}_i(h)-\hat{\xi}_i(\sqrt{2}h)\right]$ mimics the variance structure of $2\xi_i(h)-\xi(\sqrt{2}h)$.
The bootstrap method here does not require us to take into account the additional estimation variation in correcting the bias term. This is crucially because we use the undersmoothing and the additional variance introduced by correcting for the bias is of order $\sqrt{n}h^2\rightarrow 0$, which is washed away in the limit and does not matter when we use the bootstrap method.
As Table (ref) shows, both empirical and t-bootstrap methods have some undercoverage problem except for the $t$-bootstrap in the $n^{-0.28}\sigma_R/\sqrt{2}$ case. Undercoverage issue is pervasive in the nonparametric inference problem calonico2014robust,armstrong2020simple. This is because, even if we use the bias-reduced estimator, we still have the stochastic error of order $O_p(h^2+b_n^2+\log n/(nh_n\sqrt{b_n}))$, which influences the coverage probability in the finite sample size. To see how this bias can influence the coverage, we present the “Oracle Bias removed Empirical Bootstrap" in Table (ref), in which we oracularly remove the bias in the original estimator. Further removing the bias makes the coverage probability even closer to the nominal coverage probability of 95%.
To apply our framework, we consider the empirical setting from grembi2016fiscal. Their study examines the effect of fiscal rules imposed by the Italian central government on local municipalities. Beginning in 1999, the Domestic Stability Pact (DSP) required all municipalities to limit the growth of their fiscal gap—defined as the difference between total expenditure net of debt service and total revenue net of transfers. In 2001, the Italian government relaxed these rules for municipalities with populations below $5{,}000$, while the rules remained binding for those above this threshold. The key policy variable thus depends on whether a municipality’s population size lies above or below this administratively determined cutoff.
In our notation, we denote the running variable by $R_{it}$, representing the population size of municipality $i$ at time $t$. The treatment indicator is defined as \[ D_{it} = \mathbbm{1}(R_{it} \le 5000)\times \mathbbm{1}(t \geq 2001), \] which equals one if the municipality is subject to the fiscal rule in year $t$ and zero otherwise. The outcome variable, $Y_{it}$, represents the fiscal balance of the municipality, measured by the per-capita deficit (total expenditure minus total revenue) as in grembi2016fiscal. The panel spans $t = 1997, \dots, 2004$, covering both pre- and post-policy periods. For our purposes, we define the pre-policy period $(t=1999,2000)$ as the control group with $Z_i=0$, which will be used to establish the conditional trend, and the early post-policy period $(t=2000,2001)$ as the policy-influenced group with $Z_i=1$. This two-period framing allows us to align the grembi2016fiscal policy variation with the structure of our baseline model in Section (ref).
grembi2016fiscal also notes that a discontinuity in mayoral wages occurs at the same $5{,}000$ population threshold, which could contaminate identification. However, in our setup, this concern is mitigated by the conditional parallel trend assumption (ref): because the mayor’s wage is a deterministic function of the running variable $R_{it}$, its effect is fully absorbed once we condition on $R_{it}$. Thus, any wage-related differences across municipalities do not bias the identification of the fiscal-rule effect.
Finally, the population variable $R_{it}$ may itself be distorted by policy incentives. While mayors cannot directly choose the population, residents may migrate toward municipalities with more relaxed fiscal constraints due to enhanced local services. In addition, fiscal rules can influence demographic changes through birth and death rates over time. Consequently, policy effects may manifest not only near the cutoff but also through broader shifts in the population distribution. Therefore, when evaluating the total policy effect, it is crucial to take into account the indirect effect via the population variation.
The results of the estimation are shown in Table (ref). Recall that the policy is imposed targeting to constrain local municipalities from fiscal expansion, so a negative number indicates that the policy is successful in reducing the deficit.
The results show that the slash-down of the deficit is decreasing as we decrease the counterfactual cutoff to the 4,000 population. After the cutoff is moved to 4,600, the total effect of policy will be positive, which shows that the deficit cutdowns are mostly from the large population towns. The result is quite different from the grembi2016fiscal as they show that towns around the policy cutoff have increased deficits due to the policy relaxation. Of course, the objects identified in grembi2016fiscal are the local direct effect of treatment on the deficit while our results also incorporate the change in the population due to the potential of treatment. The confidence intervals are wide and cover zero most of the time. This is probably because of the high variance in the response of municipal fiscal conditions to the policy and the relatively small sample size.