EconBase
← Back to paper

Improved inference for nonparametric regression and regression-discontinuity designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

83,650 characters · 16 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

improved inference for nonparametric regression -0.25cmand regression-discontinuity designs

abstractNonparametric regression and regression-discontinuity designs suffer from smoothing bias that distorts conventional confidence intervals. Solutions based on robust bias correction (RBC) are now central to the economist's toolbox. In this paper, we establish a novel connection between RBC methods and bootstrap prepivoting. Revisiting RBC through the lens of bootstrapping allows us to develop a novel bias correction procedure which delivers improved nonparametric inference. The resulting confidence intervals are 17% shorter than the usual intervals employed in curve estimation and regression discontinuity designs, without compromising asymptotic coverage. This holds regardless of evaluation point location, bandwidth choice, or regressor and error distribution. \noindentKeywords: Asymptotic bias; bootstrap; local polynomial estimation; nonparametric regression; regression-discontinuity; robust bias correction. \noindentJEL codes: C14; C21.

introduction

Economists routinely rely on nonparametric regression and regression-discontinuity designs (RDD) to estimate causal effects. Inference in this setting is challenging due to the presence of a bias component affecting the nonparametric estimators and statistics usually employed by practitioners, even asymptotically. Several solutions to this inference problem have been proposed in the econometrics and statistics literatures. These include the so-called undersmoothing and bias correction approaches; see, inter alia, Hall (1992), Imbens and Lemieux (2008), Calonico, Cattaneo, and Titiunik (2014), Calonico, Cattaneo and Farrell (2018), and the references therein. A recent leading solution is the robust bias correction (RBC) approach of Calonico et al.\ (2014, 2018), which explicitly resolves this problem by debiasing the reference nonparametric estimator, while adjusting its standard error to account for the additional uncertainty introduced by the bias correction. Bootstrap and resampling methods, though widely used by economists as a bias correction tool, are rarely employed in the present setting due to the fact that they generally fail to correct for smoothing bias (e.g.,\ Hall, 1992).

The main objective of this paper is to propose a new robust procedure for inference in the context of nonparametric regression and RDD. Our approach is based on a non-standard application of the bootstrap that implicitly corrects for the bias, thus challenging the above fact. Standard bootstrap methods fail in the present context because the bootstrap cannot correctly mimic the asymptotic bias of the nonparametric estimator, leading to confidence intervals with incorrect coverage and bootstrap p-values which are not uniformly distributed (see also Cavaliere and Georgiev, 2020). We show that prepivoting, originally proposed by Beran (1987) as a means to obtain refinements for asymptotically unbiased estimators, can provide a solution to this failure, as motivated by recent work on asymptotically biased estimators in Cavaliere et al.\ (2024). The idea is that if the asymptotic distribution of the boostrap p-value can be consistently estimated, then we can use it to restore validity by transforming (or prepivoting) a non-uniformly distributed bootstrap p-value into one that is uniformly distributed. In our context, we can use its quantiles to adjust the original bootstrap confidence intervals.

In this framework, our first contribution is to show that prepivoting performs an implicit bias correction and at the same time adjusts the standard error of the nonparametric regression estimator to account for the additional uncertainty introduced by debiasing. That is, the prepivoted bootstrap interval is asymptotically equivalent to an RBC-style interval. We also show that, for certain bootstrap schemes, bootstrap inference based on prepivoting can be computationally straightforward, as it does not require any resampling algorithm. This desirable property occurs because both the bootstrap bias correction and the standard error adjustment are known functions of the regressors (and the kernel). When the external random variable used to construct the bootstrap confidence interval is drawn from a normal distribution, the prepivoted bootstrap interval is in fact identical to an RBC-style interval.

As our second contribution, we show that the RBC interval of Calonico et al.\ (2014, 2018) is asymptotically equivalent to prepivoting a confidence interval based on the bootstrap scheme considered by Bartalotti et al.\ (2017) and He and Bartalotti (2020) for RDD. For inference based on a local linear estimator at a given evaluation point (e.g., the cutoff for RDD), this bootstrap method generates bootstrap observations using a local quadratic estimator estimated at that point. For this reason, we call this the `global polynomial' (GP) bootstrap. The GP bootstrap is asymptotically invalid when used to construct conventional bootstrap intervals. Bartalotti et al.\ (2017) and He and Bartalotti (2020) estimate the bias by a first level bootstrap, and then use a second level bootstrap to obtain standard errors. However, we show that prepivoting solves the invalidity problem directly by modifying the bootstrap quantiles to reflect the presence of bias. In particular, we show that the prepivoted GP bootstrap (hereafter, PGP) interval is asymptotically equivalent to the RBC interval of Calonico et al.\ (2014, 2018) based on local quadratic estimates of the second-order derivative that enters the bias component. More generally, by changing the estimator of the conditional mean function used in the bootstrap DGP, we obtain different bias corrections and different studentizations.

By exploiting the asymptotic equivalence between prepivoting and RBC-style intervals, our third and main contribution is to revisit the classical bootstrap in nonparametric statistics (e.g., Härdle and Bowman, 1988, Härdle and Marron, 1991, and Hall and Horowitz, 2013). This bootstrap smooths the conditional mean function at each regressor value using a local polynomial estimator, and we label it the `local polynomial' (LP) bootstrap. The inconsistency of naive, i.e., non-prepivoted, implementations of the LP bootstrap when used with `large' (including MSE-optimal) bandwidths is well documented in the statistics literature. Different solutions have been proposed, and they typically require choosing additional bandwidth or tuning parameters. We propose and analyze new prepivoted LP bootstrap schemes (hereafter, PLP), and we show that they (i) deliver novel confidence intervals of the RBC style, not explored in the existing literature, and, more importantly, (ii) lead to efficiency gains (i.e., shorter confidence intervals) with respect to the classic RBC interval. Specifically, we show that the bias correction mechanism implicitly generated by prepivoting these bootstrap methods is more efficient than the bias estimator employed in the RBC procedure, resulting in confidence intervals of shorter length and correct coverage, asymptotically. Importantly, PLP allows using the same bandwidth for point estimation and in the bootstrap data generating process (DGP). As far as we are aware, ours is the first approach in the nonparametrics literature that delivers a valid and efficient implementation of the LP bootstrap without the need for additional tuning parameters.

As is usually the case for nonparametric estimators, the asymptotic distribution of PLP statistics depends on whether the evaluation point is interior or boundary. For interior points, prepivoting can be applied directly as discussed above. However, for boundary points, the implicit bias correction mimics the leading bias term of the original statistic only up to a multiplicative factor. The consequence is that the standard prepivoting approach studied by Cavaliere et al.\ (2024) requires modification in this context. Since the multiplicative term depends on the regressors (and the kernel) only, and is therefore known, our solution involves reweighting the bootstrap statistic before applying the prepivoting approach. We show that this `modified' PLP approach (hereafter, mPLP) adapts to the nature of the evaluation point and yields asymptotically valid inference across the entire support of the regressor, including boundary points.

Importantly, the new mPLP method not only has correct asymptotic coverage, but it also delivers shorter interval length compared to the RBC interval. Intuitively, mPLP is based on a convolution of the original observations, and such convolution introduces an additional layer of smoothing. We show that this additional smoothing induces a more efficient bias-correction in the sense that the overall variance of the debiased statistic is smaller. Crucially, this efficiency result is not tied to a specific DGP, as the asymptotic relative lengths of the confidence intervals based on mPLP and RBC depend only on the choice of kernel in the nonparametric regression and on whether the evaluation point is interior or boundary. Table (ref) gives the asymptotic relative lengths of the mPLP bootstrap and \textsf{RBC} intervals for five common kernels for the special case of the local linear estimator. For the popular Epanechnikov kernel, \textsf{mPLP} bootstrap intervals are 17% shorter than \textsf{RBC} intervals for both interior and boundary evaluation points. These efficiency results are supported by our Monte Carlo simulation experiments.

table[table omitted — 637 chars of source]

The remainder of this paper is organized as follows. Section (ref) presents some general results on prepivoting in nonparametric regression and its relation to RBC-style intervals, as well as to the GP bootstrap scheme. In Section (ref), we discuss how improved inference can be based on the LP bootstrap. We show asymptotic validity of the corresponding prepivoted confidence interval and analyze its efficiency, in terms of length, compared with the RBC interval. In Section (ref) we extend our approach to boundary problems and, in particular, to (sharp) RDD. In Section (ref) we assess the performance of our methods in finite samples via Monte Carlo simulation. Section (ref) provides guidance on practical implementation for applied researchers. Finally, Section (ref) concludes. All notation, technical derivations, and proofs are included in the appendix and in the accompanying supplemental material. R packages that implement the procedures in this paper are available at \url{https://pppackages.github.io}.

robust bias correction and prepivoting

preliminaries

Given a random sample $\mathscr{D}_n := \{(y_i, x_i): i=1,\dots,n\}$, consider the problem of inference on the conditional expectation $g(x):=\operatorname{\mathbb{E}}[y_i|x_i=x]$ at a fixed point $x=\mathsf{x}$, using a local ($p$-th order, $p$ odd) polynomial estimator $\hat{g}_n(x): = \hat{g}_n(x;h,K)$, where $h=h_n>0$ is the bandwidth and $K$ the kernel function; that is, with $r_p(u)=(1,u,\ldots ,u^{p})^{\prime }$,

equation[equation omitted — 271 chars of source]

The estimator admits the linear closed-form expression $\hat{g}_n(x) = \sum_{i=1}^n w_i (x) y_i /(nh)$ with weights $w_i(x)$ depending on the data, the polynomial order $p$, and the kernel function $K$; see Appendix (ref). Inference based on $\hat{g}_n(\mathsf{x})$ is challenging due to the presence of an asymptotic bias. In particular, we typically have

equation[equation omitted — 134 chars of source]

where, letting $\mathscr{X}_n:=\{x_i : i=1,\dots,n\}$, the asymptotic bias $B:=\operatorname*{\mathrm{plim}}_{n \to \infty} B_n$, $B_n:=\operatorname{\mathbb{E}}[T_n|\mathscr{X}_n]$, is non-zero when bandwidths with MSE-optimal rate are employed. Sufficient conditions for \hyperref[{biasedCLT}]{\tagform@{\ref*{biasedCLT}}} are given by Assumptions (ref)--(ref) below; e.g., Ruppert and Wand (1994) and Fan and Gijbels (1996).

assumption$(y_i,x_i)$ are i.i.d.\ and $x_i$ has bounded support $\mathbb{S}_x$ and continuous density $f$. Let $\mathsf{N}$ be an open neighborhood of $\mathsf{x}$. Then, on $\mathsf{N} \cap \mathbb{S}_x$: (i) $f(x)>0$; (ii) $\sigma^2 (x):=\operatorname{\mathbb{V}}[y_i|x_i=x]>0$ is continuous and $\sup_x \operatorname{\mathbb{E}}[y_i^4|x_i=x]<\infty$; (iii) $g^{(p+1)}$ is H\"{o}lder continuous with exponent $\eta>0$.
assumption(i) The kernel $K$ has support $(-1,1)$ on which it is symmetric, positive, and Lipschitz continuous; (ii) the bandwidth $h=h_n$ is such that $h + (nh)^{-1} \log n \to 0$ and $nh^{2p+3}\to\kappa \in [0,+\infty)$.

Assumption (ref) covers the most widely used kernel functions, e.g., the uniform, Epanechnikov, and triangular kernels, as well as `large' ($\kappa >0$) and undersmoothing ($\kappa = 0$) bandwidths. Under Assumptions (ref)--(ref), the bias term $B_n$ satisfies

equation[equation omitted — 121 chars of source]

where $C_n =C_n (\mathsf{x})$ is a random variable satisfying $C_n \operatorname*{\to_{\mathrm{p}}} C$ for a non-zero constant $C$ defined in Appendix (ref). Note that for undersmoothing bandwidths, $B_n=\operatorname*{o_{\mathrm{p}}} (1)$ and $B=0$.

The presence of bias invalidates confidence intervals that ignore it. For instance, with $\alpha \in (0,1)$, a $100(1-\alpha)$% nominal level conventional interval that ignores the bias is

equation*[equation* omitted — 113 chars of source]

where $\hat{v}^2_{1n}$ is a consistent estimator of $v^2_1$. However, $\operatorname{\mathbb{P}} ( g(\mathsf{x}) \in \mathrm{CI}_n )$ does not converge to $1-\alpha$ as $n\to \infty$. To restore asymptotic validity of the confidence interval, the RBC approach of Calonico et al.\ (2014, 2018) recenters $\mathrm{CI}_n$ by an estimate of the leading bias term and adjusts the standard error to account for the variability of the bias estimator. With $\hat{B}_{\textsf{RBC},n}$ and $\hat{v}_{\textsf{RBC},n}$ denoting the bias estimate and the adjusted standard error, respectively, the $100(1-\alpha)$% nominal level RBC interval is

equation[equation omitted — 202 chars of source]

For local polynomial estimators, $\hat{B}_{\textsf{RBC},n}$ is based on nonparametric estimators of higher-order derivatives, which may involve the choice of additional bandwidths. It has the form $\hat{B}_{\textsf{RBC},n} = \sqrt{nh^{2p+3}}\hat{g}_n^{(p+1)}(\mathsf{x})C_n/(p+1)!$, where $C_n$ is as introduced previously, and $\hat{g}_n^{(p+1)}(\mathsf{x})$ is an estimator of $g^{(p+1)}(\mathsf{x})$. The estimator $\hat{v}^2_{\textsf{RBC},n}$ estimates the variance of the debiased statistic $\sqrt{nh}\hat{g}_n(\mathsf{x})-\hat{B}_{\textsf{RBC},n}$. As such, it is equal to $\hat{v}^2_{1n}$ (the variance estimator of $\sqrt{nh}\hat{g}_n(\mathsf{x})$) plus an additional term that estimates the variance of $\hat{B}_{\textsf{RBC},n}$ and the covariance between $\hat{B}_{\textsf{RBC},n}$ and $\sqrt{nh}\hat{g}_n(\mathsf{x})$; see Calonico et al.\ (2018, Section 3.1).

an equivalence result

As we show next, RBC intervals can be derived by prepivoting specific (invalid) bootstrap schemes. In Section (ref) we exploit the link between the two approaches to derive new RBC-type intervals with improved length properties.

Let $\hat{g}^*_n(\mathsf{x}): = \hat{g}^*_n(\mathsf{x};h,K)$ denote a bootstrap analogue of $\hat{g}_n(\mathsf{x})$ based on the same kernel function $K$ and the same bandwidth $h$. Conventional percentile bootstrap intervals are typically based on the quantiles of the bootstrap distribution of $T^*_n$, a given bootstrap analogue of $T_n$ (for instance, $T^*_n = \sqrt{nh} ( \hat{g}^*_n(\mathsf{x}) - \hat{g}_n(\mathsf{x}))$, see Section (ref)). Letting $\hat{L}_n(u):= \operatorname{\mathbb{P}}^* ( T_n^* \leq u )$, a standard $100(1-\alpha)$% nominal level equal-tailed bootstrap interval for $g(\mathsf{x})$ is

equation[equation omitted — 219 chars of source]

which, however, does not contain $g(\mathsf{x})$ with probability converging to $1-\alpha$. Just as the original nonparametric estimator $\hat{g}_n(\mathsf{x})$ is biased, its bootstrap analogue $\hat{g}^*_n(\mathsf{x})$ is also biased, implying the invalidity of $\mathrm{CI}_{\mathsf{Boot},n}$. As discussed by Cavaliere et al.\ (2024, Remark 3.2), the root of the problem lies in the fact that the standard bootstrap p-value $\hat{p}_n:=\operatorname{\mathbb{P}}^*(T^*_n\leq T_n)$ is not asymptotically distributed as $U_{[0,1]}$ (standard uniform) in the presence of asymptotic bias.

Prepivoting solves this problem by changing the level of the bootstrap quantiles. Let $\hat{H}_n$ be a (uniformly) consistent estimator of $H$, the asymptotic cumulative distribution function (cdf) of $\hat{p}_n$. A $100(1-\alpha)$% nominal level equal-tailed bootstrap interval based on prepivoting is

equation[equation omitted — 252 chars of source]

When $\hat{p}_n\operatorname*{\to_{\mathrm{d}}} U_{[0,1]}$ then $H(u)=u$ and $\hat{H}_n^{-1}(\alpha )\operatorname*{\to_{\mathrm{p}}}\alpha$, in which case the prepivoted interval $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}$ is asymptotically equivalent to $\mathrm{CI}_{\mathsf{Boot},n}$. But when asymptotic uniformity of $\hat{p}_n$ fails, $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}$ remains asymptotically valid while $\mathrm{CI}_{\mathsf{Boot},n}$ becomes invalid. The main reason why prepivoting restores validity is that if $\hat{H}_n$ is a uniformly consistent estimator of $H$ (regardless of whether the latter is the uniform cdf or not), the prepivoted p-value $\tilde{p}_n:=\hat{H}_n(\hat{p}_n)$ is asymptotically distributed as $U_{[0,1]}$ by the probability integral transform. This implies asymptotic validity of $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}$, as shown by Cavaliere et al.\ (2024, Remark 3.4); that is, for any $\alpha \in (0,1)$,

align[align omitted — 408 chars of source]

Prepivoting avoids an explicit bias correction whenever $\hat{H}_n$ does not require estimating $B$. Denote by $\hat{B}_n :=\operatorname{\mathbb{E}}^*[T_n^*]$ and $\hat{v}^2_n:=\operatorname{\mathbb{V}}^*[T_n^*]$ the bias and variance, respectively, of the bootstrap statistic $T^*_n$ under the bootstrap probability measure. As we show next, we may regard $\hat{B}_n$ as a bias estimator implicitly induced by the bootstrap. Suppose the following high level conditions hold.

\setcounter{assnLetter}{7}

assnLetter(i) $T^*_n-\hat{B}_n \operatorname*{\overset{d*}{\to_{\mathrm{p}}}} N(0,v^2)$ with $v^2=\operatorname*{\mathrm{plim}} \hat v^2_n>0$; (ii) $(T_n - B_n, \hat{B}_n - B_n)^\prime \operatorname*{\overset{d}{\to}} N(0,V)$, where $V$ is positive definite.

Assumption (ref)(i) requires the bootstrap statistic $T^*_n$ to be conditionally asymptotically Gaussian after we remove the bootstrap bias $\hat{B}_n$. Note that the bootstrap does not need to replicate the asymptotic variance of $T_n$, i.e., $v^2$ may be different from $v^2_1$. Under Assumption (ref)(i), as in Theorem 3.4 of Cavaliere et al.\ (2024), we can show that $\hat{p}_n=\Phi(\hat{v}_n^{-1}(T_n-\hat{B}_n))+\operatorname*{o_{\mathrm{p}}}(1)$. If in addition Assumption (ref)(ii) holds, then $T_n-\hat{B}_n\operatorname*{\overset{d}{\to}} N(0,v^2_\mathsf{P})$ with $v_{\mathsf{P}}^2>0$, and the asymptotic cdf of $\hat p_n$ is $H (u):= \Phi (m^{-1} \Phi^{-1} (u))$, $m:= {v}_{\mathsf{P}}/{v}$. A consistent (plug-in) estimator of $H$ can be obtained as

equation[equation omitted — 132 chars of source]

where $\hat{v}^2_{\mathsf{P},n}$ is a consistent estimator of $v^2_{\mathsf{P}}$. Different bootstrap methods induce different $\hat{B}_n$ and consequently require different estimators $\hat{v}^2_{\mathsf{P},n}$.

Although $\hat{H}_n$ of \hyperref[{Hhat_def}]{\tagform@{\ref*{Hhat_def}}}, and hence $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}$, does not involve $\hat{B}_n$ explicitly, we can show that $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}$ is asymptotically equivalent to an interval that bears a close resemblance to the RBC interval $\mathrm{CI}_{\mathsf{RBC},n}$ in that it too contains a bias correction and adjusts the standard error for the additional uncertainty appropriately. We describe this result next.

For simplicity, suppose that $T^*_n$ is conditionally Gaussian; that is, conditionally on $\mathscr{D}_n$, $T^*_n \sim N(\hat{B}_n,\hat{v}^2_n)$ and hence $\hat{L}_n(u) = \Phi ((u-\hat{B}_n )/\hat{v}_n )$. In this case, $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}$ in \hyperref[{BTprepCI}]{\tagform@{\ref*{BTprepCI}}} is exactly equal to

equation[equation omitted — 198 chars of source]

In general, $T^*_n$ is Gaussian only asymptotically, and we obtain the following result.

theoremLet $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}$, $\hat{H}_n$, and $\mathrm{CI}_{\mathsf{P} \textsf{-} \mathsf{RBC},n}$ be defined in \hyperref[{BTprepCI}]{\tagform@{\ref*{BTprepCI}}}, \hyperref[{Hhat_def}]{\tagform@{\ref*{Hhat_def}}}, and \hyperref[{CIprbc}]{\tagform@{\ref*{CIprbc}}}, respectively. \begin{itemize} • If Assumption (ref)(i) holds then $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{P} \textsf{-} \mathsf{RBC},n}+\operatorname*{o_{\mathrm{p}}}((nh)^{-1/2})$. • If Assumption (ref)(i) holds and $T_n^*$ is Gaussian conditional on $\mathscr{D}_n$, then $\mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{P} \textsf{-} \mathsf{RBC},n}$ a.s. • If Assumption (ref) holds and $\hat{v}^2_{\mathsf{P},n}\operatorname*{\to_{\mathrm{p}}} v^2_{\mathsf{P}}$, then $ \operatorname{\mathbb{P}} ( g(\mathsf{x}) \in \mathrm{CI}_{\mathsf{P}\hspace{-0.5pt},n} ) \to 1- \alpha$. \end{itemize}

An interesting implication of Theorem (ref) is that we can obtain $\mathrm{CI}_{\mathsf{P} \textsf{-} \mathsf{RBC},n}$ without any resampling as it depends only on $\hat{B}_n$ and $\hat{v}^2_{\mathsf{P},n}$ (which we derive analytically for different residual-based bootstrap methods in the next sections), and standard normal critical values $z_\alpha$. More importantly, this asymptotic equivalence result implies that we can view RBC through the lens of prepivoting (or vice versa), as we show in the next section.

standard rbc inference as bootstrap prepivoting

Assume that the bootstrap data $\mathscr{D}_n^*:=\{(y_i^*,x_i^*): i=1,\dots, n\}$ are generated by setting $x_i^* = x_i$ for $i=1,\dots, n$ (`fixed regressor' bootstrap) and

equation[equation omitted — 125 chars of source]

where $\varepsilon _i^{\ast}:=\tilde\varepsilon_i e_i^{\ast}$ with $\tilde\varepsilon_i$ being the bias-corrected HCk residuals considered in Calonico et al.\ (2018) and $e_i^{\ast}$ being (conditionally on the data)\ i.i.d.\ with $\operatorname{\mathbb{E}}^{\ast}[e_i^{\ast}]=0$ and $\operatorname{\mathbb{E}}^{\ast}[e_i^{\ast 2}]=1$. The bootstrap conditional expectation $\operatorname{\mathbb{E}}^{\ast}[y_i^{\ast}|x^*_i=x_i]$ in \hyperref[{btsDGPLQ}]{\tagform@{\ref*{btsDGPLQ}}} is

equation*[equation* omitted — 197 chars of source]

where $\hat{\beta}_{p+1,n}(\mathsf{x})$ is obtained as in \hyperref[{eq LP estimator}]{\tagform@{\ref*{eq LP estimator}}} using local polynomial regression of order $p+1$ at the fixed point $\mathsf{x}$; see Bartalotti et al.\ (2017) and He and Bartalotti (2020) for the case $p=1$. Since $\iota _{j}^{\prime }\hat{\beta}_{p+1,n}(\mathsf{x})$ is an estimator of $g^{(j)}(\mathsf{x})/j!$, the function $\tilde{g}_{\mathsf{x},n} (x)$ is an estimator of a ($p+1$)-th order Taylor expansion of the conditional expectation $g(x)=\operatorname{\mathbb{E}}[y_i|x_i=x]$ around the fixed point $\mathsf{x}$. That is, to generate the bootstrap data we use the same function $\tilde{g}_{\mathsf{x},n} (x)$ estimated locally at $\mathsf{x}$ but evaluated globally at each $x=x_i$ for $i =1,\dots, n$. Hence, we use `global polynomial' (GP) bootstrap to label the DGP \hyperref[{btsDGPLQ}]{\tagform@{\ref*{btsDGPLQ}}}. Importantly, this method generates bootstrap data using a polynomial order ($p+1$) that is larger than the order $p$ used in estimation. This mismatch is key to inducing an asymptotic bias under the bootstrap probability measure that satisfies Assumption (ref).

Let $\hat{g}_n^{\ast}(x)$ be the local ($p$-th order) polynomial estimator applied to the bootstrap sample generated as \hyperref[{btsDGPLQ}]{\tagform@{\ref*{btsDGPLQ}}}. The GP bootstrap analogue of $T_n$ is $T_{\mathsf{GP},n}^{\ast}=\sqrt{nh}(\hat{g}_n^{\ast}(\mathsf{x})-\tilde{g}_{\mathsf{x},n} (\mathsf{x}))$, and its bootstrap bias and variance are denoted $\hat{B}_{\mathsf{GP},n}:=\operatorname{\mathbb{E}}^*[T_{\mathsf{GP},n}^\ast]$ and $\hat{v}_{\mathsf{GP},n}^2:=\operatorname{\mathbb{V}}^*[T_{\mathsf{GP},n}^\ast]$, respectively. Because the GP bootstrap is a fixed-design scheme (i.e., the regressors are fixed across bootstrap samples), these moments can be computed in closed form. Specifically, we show that

equation[equation omitted — 121 chars of source]

where $\hat{g}_n^{(p+1)}(\mathsf{x}):=(p+1)!\iota_{p+1}' \hat{\beta}_{p+1,n}(\mathsf{x})$ is an estimator of $g^{(p+1)}(\mathsf{x})$, clearly implying that $\hat{B}_{\mathsf{GP},n}$ is an estimator of the dominant term of $B_n$ in \hyperref[{eq:Bn}]{\tagform@{\ref*{eq:Bn}}}. Interestingly, $\hat{B}_{\mathsf{GP},n}$ is identical to $\hat{B}_{\mathsf{RBC},n}$ as given in Calonico et al.\ (2018, Section 3); see the discussion after \hyperref[{eq RBC int}]{\tagform@{\ref*{eq RBC int}}}.

As explained in the previous section, a prepivoted bootstrap confidence interval \hyperref[{BTprepCI}]{\tagform@{\ref*{BTprepCI}}} based on $(T_n,T_{\mathsf{GP},n}^{\ast})$, say $\mathrm{CI}_{\mathsf{PGP}\hspace{-0.5pt},n}$, can be constructed by choosing $\hat{v}^2_{\mathsf{PGP},n}$ as a consistent estimator of $v^2_{\mathsf{PGP}}$, which is the asymptotic variance of $T_n-\hat{B}_{\mathsf{GP},n}$. Because $\hat{B}_{\mathsf{GP},n}=\hat{B}_{\mathsf{RBC},n}$, the natural plug-in estimator for the GP case coincides with the variance estimator in Calonico et al.\ (2018). In Theorem (ref) below we show (i) validity of such confidence interval, (ii) its asymptotic equivalence to the RBC interval $\mathrm{CI}_{\mathsf{RBC},n}$ in \hyperref[{eq RBC int}]{\tagform@{\ref*{eq RBC int}}}, and (iii) its almost sure equality to the RBC interval for Gaussian bootstraps. The proof requires verifying that Assumption (ref) holds for $(T_n,T_{\mathsf{GP},n}^{\ast})$ under Assumptions (ref) and (ref). The theorem then follows by Theorem (ref).

theoremUnder Assumptions (ref) and (ref), for any $\mathsf{x}\in \mathbb{S}_{x}$, $\operatorname{\mathbb{P}}(g(\mathsf{x})\in \mathrm{CI}_{\mathsf{PGP}\hspace{-0.5pt},n} )\to 1-\alpha$ and $\mathrm{CI}_{\mathsf{PGP}\hspace{-0.5pt},n} = \mathrm{CI}_{\mathsf{RBC},n} + \operatorname*{o_{\mathrm{p}}} ((nh)^{-1/2})$. If, in addition, $T_{\mathsf{GP},n}^{\ast}$ is Gaussian conditional on $\mathscr{D}_n$, then $\mathrm{CI}_{\mathsf{PGP}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{RBC},n}$ a.s.

improved nonparametric inference

The main feature of the GP bootstrap discussed above is that the conditional mean function under the bootstrap probability measure is based on a single local approximation of the curve $g$, using a polynomial expansion of $g$ around $\mathsf{x}$, which is then applied globally to all $x_i$'s. In contrast, in this section we consider the `local polynomial' (LP) bootstrap scheme traditionally employed in the statistics literature, where $g$ is approximated locally for each of the $x_i$'s, implying that the bootstrap conditional expectation function resembles $g$ more closely over the entire support (Härdle and Marron, 1988, Härdle and Bowman, 1991, Hall and Horowitz, 2013). This point is illustrated in Figure (ref), where we contrast the true regression curve with the GP and LP bootstrap curves under the setup considered in Section (ref).

figure[figure omitted — 703 chars of source]

In this section, we show how the prepivoted LP bootstrap leads to improved inference. We focus here on the case where $\mathsf{x}$ is an interior point. Boundary points will be considered in Section (ref).

the local polynomial bootstrap

The LP bootstrap sets the bootstrap conditional expectation to $\operatorname{\mathbb{E}}^{\ast}[y_i^{\ast}|x^*_i=x] = \hat{g}_n(x)$, where $\hat{g}_n(x)$ is as in \hyperref[{eq LP estimator}]{\tagform@{\ref*{eq LP estimator}}}. The bootstrap data $\mathscr{D}_n^*:=\{(y_i^*,x_i^*): i=1,\dots, n\}$ satisfy $x_i^* = x_i$, $i=1,\dots, n$, and

equation[equation omitted — 109 chars of source]

where $\varepsilon_i^* := \tilde\varepsilon_i e_i^*$ with $\tilde\varepsilon_i$ and $e_i^*$ as defined in Section (ref) to facilitate comparison with the existing RBC approach. Other choices are possible, e.g., $\tilde\varepsilon_i = y_i - \hat{g}_n(x_i)$, without changing our results. Bootstrap DGPs of this or similar forms have been widely applied to nonparametric regression without prepivoting; see, e.g., Härdle and Marron (1988), Härdle and Bowman (1991), and Hall and Horowitz (2013).

The LP bootstrap uses a $p$-th order local polynomial regression estimate of $g(x)$ calculated at each sample point $x=x_i$ to generate $y^*_i$, rather than extrapolating values for $g(x_i)$ from a single $(p+1)$-th order local polynomial regression estimated at $x=\mathsf{x}$, as for the GP bootstrap. In this section, we exploit this key difference and show the following. First, we demonstrate that prepivoting can also be used to restore asymptotic validity of the LP bootstrap without requiring undersmoothing or the choice of additional bandwidths, a result that stands in contrast to the existing bootstrap literature. Second, we show that the resulting prepivoted LP bootstrap (PLP) generates a novel RBC-type confidence interval which differs from the existing RBC intervals in that it does not rely on a higher-order local polynomial estimate of the derivatives entering the bias term. Third, and most importantly, we show that in large samples, the new PLP interval is shorter than the classical RBC CIs in Calonico et al.\ (2014, 2018), while maintaining correct coverage. Put differently, with respect to the usual bias estimators used in RBC procedures, the bias estimator implicitly generated by the PLP method leads to improved inference on the regression curve.

theory

Consider the prepivoted bootstrap confidence interval \hyperref[{BTprepCI}]{\tagform@{\ref*{BTprepCI}}} based on $(T_n,T_{\mathsf{LP},n}^{\ast})$, say $\mathrm{CI}_{\mathsf{PLP}\hspace{-0.5pt},n}$, where $T_{\mathsf{LP},n}^{\ast}=\sqrt{nh}(\hat{g}_n^{\ast}(\mathsf{x})-\hat{g}_n(\mathsf{x}))$ and $\hat{g}_n^{\ast}(\mathsf{x})$ is the local ($p$-th order) polynomial estimator applied to the bootstrap sample generated as \hyperref[{btsDGPLP}]{\tagform@{\ref*{btsDGPLP}}}. The bootstrap bias and variance of $T_{\mathsf{LP},n}^{\ast}$ are $\hat{B}_{\mathsf{LP},n}:=\operatorname{\mathbb{E}}^*[T_{\mathsf{LP},n}^\ast]$ and $\hat{v}_{\mathsf{LP},n}^2:=\operatorname{\mathbb{V}}^*[T_{\mathsf{LP},n}^\ast]$, respectively.

Recall from Section (ref) that in order to construct $\mathrm{CI}_{\mathsf{PLP}\hspace{-0.5pt},n}$, a consistent estimator $\hat{v}_{\mathsf{PLP},n}^2$ of the asymptotic variance $v^2_{\mathsf{PLP}}$ of $T_n - \hat{B}_{\mathsf{LP},n}$ is required. To find $\hat{v}_{\mathsf{PLP},n}^2$, we first derive the bootstrap bias $\hat{B}_{\mathsf{LP},n}$. It follows from the closed-form expression for $\hat{g}^{\ast}_n(\mathsf{x})$ that

align[align omitted — 167 chars of source]

Since $B_n = \sqrt{nh}\big((nh)^{-1}\sum_{i=1}^n w_i(\mathsf{x}) g(x_i) - g(\mathsf{x})\big)$, \hyperref[{eq:BhatLP0}]{\tagform@{\ref*{eq:BhatLP0}}} makes clear that the LP bootstrap bias $\hat{B}_{\mathsf{LP},n}$ differs from $\hat{B}_{\mathsf{RBC},n}$ by targeting the bias directly, rather than estimating the $(p+1)$-th derivative that appears in its asymptotic expansion. We can rewrite $\hat{B}_{\mathsf{LP},n}$ as

equation[equation omitted — 149 chars of source]

with weights defined by the convolution $w_{\mathsf{LP}\text{-}\mathsf{bc},i} (\mathsf{x}):= (nh)^{-1} \sum_{j=1}^n w_{j} ( \mathsf{x}) w_i ( x_j)- w_i( \mathsf{x})$. Since $T_n = \sqrt{nh}\big( (nh)^{-1} \sum_{i=1}^n w_i(\mathsf{x})y_i - g(\mathsf{x})\big)$, we obtain

equation*[equation* omitted — 276 chars of source]

Let $v^2_{\mathsf{PLP}}$ denote the asymptotic variance of $T_n - \hat{B}_{\mathsf{LP},n}$. Because the weights are measurable with respect to $\mathscr{X}_n$ and $\operatorname{\mathbb{V}} [y_i | \mathscr{X}_n] = \sigma^2 (x_i)$, by conditioning on the regressors we have $\operatorname{\mathbb{V}} [T_n - \hat{B}_{\mathsf{LP},n}| \mathscr{X}_n] = (nh)^{-1} \sum_{i=1}^n w_{\mathsf{PLP},i}(\mathsf{x})^2 \sigma^2 (x_i)$. This suggests the estimator

equation[equation omitted — 143 chars of source]

A formal proof of consistency of $\hat{v}_{\mathsf{PLP},n}^2$ under Assumptions (ref) and (ref) is provided next.

lemmaUnder Assumptions (ref) and (ref), $\hat{v}_{\mathsf{PLP},n}^2 \operatorname*{\to_{\mathrm{p}}} v_{\mathsf{PLP}}^2 >0$.

In view of Lemma (ref), an application of Theorem (ref) implies the asymptotic validity of the PLP interval $\mathrm{CI}_{\mathsf{PLP}\hspace{-0.5pt},n}$. In addition, $\mathrm{CI}_{\mathsf{PLP}\hspace{-0.5pt},n}$ is asymptotically equivalent to

equation*[equation* omitted — 188 chars of source]

which is a new $\mathsf{RBC}$-type interval that does not require resampling. The following theorem states these results.

theoremLet $\mathsf{x}$ be an interior point. Under Assumptions (ref) and (ref), $\operatorname{\mathbb{P}}(g(\mathsf{x})\in \mathrm{CI}_{\mathsf{PLP}\hspace{-0.5pt},n})\to 1-\alpha$ and $\mathrm{CI}_{\mathsf{PLP}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{PLP} \textsf{-} \mathsf{RBC},n} + \operatorname*{o_{\mathrm{p}}} ((nh)^{-1/2})$. If, in addition, $T_{\mathsf{LP},n}^{\ast}$ is Gaussian conditional on $\mathscr{D}_n$, then $\mathrm{CI}_{\mathsf{PLP}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{PLP} \textsf{-} \mathsf{RBC},n}$ a.s.

improved inference

In this section, we compare the asymptotic efficiency of the new prepivoted PLP intervals and the existing RBC-type intervals proposed by Calonico et al.\ (2014, 2018). We show that the former are asymptotically shorter than the latter. This efficiency result is based on the asymptotic equivalence between prepivoting and the RBC approach in Section (ref). In particular, Theorems (ref) and (ref) imply that the asymptotic relative length of $\mathrm{CI}_{\mathsf{PLP}\hspace{-0.5pt},n}$ compared to $\mathrm{CI}_{\mathsf{RBC},n}$ is equal to the ratio between the asymptotic standard deviations of their respective debiased statistics, i.e., $v_{\mathsf{PLP}}/v_{\mathsf{RBC}}$.

To shed some light on this ratio, let $\mathsf{w}(u), u\in \operatorname{\mathbb{R}}$, be the so-called `equivalent kernel' associated with $K$ and $p$. Specifically, for $\mathcal{X}=\mathbb{R}$, $\mathsf{w}(u) := \iota_0^{\prime}(\int_{\mathcal{X}} r_p(s) r_p^{\prime}(s)K(s) ds)^{-1} r_p(u) K(u)$, the asymptotic analogue of the sample weights $w_i(x)$; see Appendix (ref). Similarly, define the asymptotic analogue of $w_{\mathsf{PLP},i}(x)$ as $\mathsf{w}_{\mathsf{PLP}}(u) := 2 \mathsf{w}(u) - \int_{\mathcal{X}} \mathsf{w}(u-r) \mathsf{w}(r) dr$. This contrasts with that of the RBC method given by $\mathsf{w}_{\mathsf{RBC}}(u) := \mathsf{w}(u) - C\iota_{p+1}^{\prime}(\int_{\mathcal{X}} r_{p+1}(s) r_{p+1}^{\prime}(s)K(s) ds)^{-1} r_{p+1}(u) K(u)$; see Lemma (ref) and Calonico et al.\ (2022, Section 4.2).

corollaryUnder Assumptions (ref) and (ref), \begin{equation*} (v^2_{\mathsf{PLP}}, v^2_{\mathsf{RBC}})= \frac{\sigma^2 (\mathsf{x})}{f(\mathsf{x})} (\mathcal{K}_{\mathsf{PLP}}, \mathcal{K}_{\mathsf{RBC}}) , \end{equation*} where \begin{equation} \mathcal{K}_{\mathsf{PLP}} := \int_{\mathcal{X}} \mathsf{w}_{\mathsf{PLP}}(u)^2 du >0 \quadand\quad \mathcal{K}_{\mathsf{RBC}} := \int_{\mathcal{X}} \mathsf{w}_{\mathsf{RBC}}(u)^2 du >0 \end{equation} are functions only of the kernel $K$ and the polynomial order $p$.

As shown in Section (ref), $v^2_{\mathsf{PLP}}$ is the asymptotic variance of a (scaled) weighted average of $y_i$ with weights $w_{\mathsf{PLP},i}(\mathsf{x})$. Using fairly standard arguments in the nonparametric literature (e.g., Fan and Gijbels, 1996, p. 66), Corollary (ref) shows that this variance is proportional to $\sigma^2 (\mathsf{x})/f(\mathsf{x})$ where the constant of proportionality is $\mathcal{K}_{\mathsf{PLP}}$, which is a known function of the kernel and the polynomial order, $p$. A similar proportionality holds for $v^2_{\mathsf{RBC}}$ with the difference that the constant of proportionality is $\mathcal{K}_{\mathsf{RBC}}$, which is a different function of the kernel and $p$. Hence, the relative asymptotic length of the $\mathsf{PLP}$ and $\mathsf{RBC}$ intervals is the square root of $\mathcal{K}_{\mathsf{PLP}}/\mathcal{K}_{\mathsf{RBC}}$.

table[table omitted — 619 chars of source]

Table (ref) reports the values of $\mathcal{K}_{\mathsf{PLP}}$, $\mathcal{K}_{\mathsf{RBC}}$, and the square root of their ratio for five common kernels and $p=1$, which is by far the most commonly used value in practice. For all the kernels considered, the decrease in interval length is substantial, ranging from $14\%$ to $17\%$. Figure (ref) provides a comparison of the functions $\mathsf{w}_{\mathsf{PLP}}$ and $\mathsf{w}_{\mathsf{RBC}}$ for the five kernels previously considered and $p=1$. While neither function is always closer to zero than the other, $\mathsf{w}_{\mathsf{PLP}}(u)$ clearly shows less variation around zero, resulting in $\mathcal{K}_{\mathsf{PLP}}<\mathcal{K}_{\mathsf{RBC}}$.

figure[figure omitted — 442 chars of source]

inference at the boundary and rdd

In this section, we extend the results in Section (ref) to address inference on $g(\mathsf{x})$ when $\mathsf{x}$ is a boundary point. This includes standard applications of nonparametric regression to RDD, which is prominent in empirical work. As it turns out, the key challenge when $\mathsf{x}$ is on the boundary is that the LP bootstrap bias is no longer centered around the original bias $B_n$, not even asymptotically. Therefore, Assumption (ref)(ii) breaks down and Theorem (ref) no longer applies; see Section (ref). We show that a simple modification of the prepivoting method ($\mathsf{mPLP}$) restores the validity results established in Section (ref) for interior points. Importantly, confidence intervals based on the proposed $\mathsf{mPLP}$ procedure automatically account for cases where $\mathsf{x}$ is at the boundary, while remaining asymptotically equivalent to the $\mathsf{PLP}$ procedure from Section (ref) when $\mathsf{x}$ is an interior point. The proposed modification is presented in Section (ref), with its application to regression discontinuity designs discussed in Section (ref).

Challenges of standard prepivoting at the boundary

A crucial assumption underlying the results of Section (ref) is Assumption (ref)(ii), which requires that $\hat{B}_{\mathsf{LP},n} - B_n$ converges in distribution to a mean-zero Gaussian random variable. This assumption fails when $\mathsf{x}$ is a boundary point since $\hat{B}_{\mathsf{LP},n} - B_n$ is no longer (asymptotically) centered at zero. The reason is that, in contrast with the interior case, $C_n$ and its LP bootstrap analogue $C_{\mathsf{LP},n}$ do not converge to the same constant $C$ when $\mathsf{x}$ is on the boundary. Instead, Lemma (ref) shows that

equation[equation omitted — 254 chars of source]

where $\xi_{2n,\mathsf{LP}}:= (nh)^{-1/2} \sum_{i=1}^n w_{\mathsf{LP}\text{-}\mathsf{bc},i} (\mathsf{x}) \varepsilon_i$ has zero mean, $A_n:=\sqrt{nh^{2p+3}}{g^{(p+1)}(\mathsf{x})}({C}_{\mathsf{LP},n}-C_n)/(p+1)!$, $C_{\mathsf{LP},n}:=(nh)^{-1}\sum_{i=1}^{n}w_i(\mathsf{x})C_n(x_i)$ is a smoothed version of $C_n$, and $g^{(p+1)}(\mathsf{x})$ denotes the (directional) ($p+1$)-order derivative at $\mathsf{x}$. Importantly, $C \neq C_{\mathsf{LP}}$ when $\mathsf{x}$ is on the boundary, but they are uniquely determined by the kernel $K$ and the polynomial order $p$; see Lemma (ref).

Our modified method restores validity of the prepivoted bootstrap confidence interval by simply rescaling the bootstrap statistic $T^*_{\mathsf{LP},n}$ with a known function of the data, say $Q_n$. The scaling factor $Q_n$ is chosen such that the term $A_n$ is eliminated from the asymptotic distribution of the bootstrap bias. As we shall see, it does not involve any additional tuning parameters or unknown quantities.

a boundary-adaptive prepivoting method

Let $T^*_{\mathsf{mLP},n}:=Q_nT^*_{\mathsf{LP},n}$, where $Q_n:=C_n/C_{\mathsf{LP},n}$ depends only on $K$, $h$, and $\mathscr{X}_n$. By definition, $\hat{B}_{\mathsf{mLP},n} :=\operatorname{\mathbb{E}}^* [T^*_{\mathsf{mLP},n}] = Q_n \hat{B}_{\mathsf{LP},n}$ with $\hat{B}_{\mathsf{LP},n}$ as in Section (ref). In contrast with $T^*_{\mathsf{LP},n}$, Assumption (ref)(ii) holds for $T^*_{\mathsf{mLP},n}$ for any $\mathsf{x}$, including boundary points, because we have defined $Q_n$ precisely such that $Q_n C_{\mathsf{LP},n} = C_n$. Specifically,

equation[equation omitted — 263 chars of source]

Consequently, the modified LP bootstrap bias $\hat{B}_{\mathsf{mLP},n}$ is asymptotically centered at $B_n$, as required in Assumption (ref)(ii). Crucially, this result holds for any $\mathsf{x} \in \mathbb{S}_x$, including boundary points.

Taken together, these results imply that an asymptotically valid modified confidence interval, denoted $\mathrm{CI}_{\mathsf{mPLP}\hspace{-0.5pt},n}$, can be constructed provided a consistent estimator $\hat{v}_{\mathsf{mPLP},n}^2$ is available for the asymptotic variance $v^2_{\mathsf{mPLP}}$ of $T_n-\hat{B}_{\mathsf{mLP},n}$. By relying on arguments similar to those used for $\hat{v}^2_{\mathsf{PLP},n}$ in \hyperref[{vPest}]{\tagform@{\ref*{vPest}}}, we can show that a consistent estimator of $v_{\mathsf{mPLP}}^2$ is given by

align[align omitted — 264 chars of source]

where $w_{\mathsf{mLP}\text{-}\mathsf{bc},i} (\mathsf{x}):=Q_n w_{\mathsf{LP}\text{-}\mathsf{bc},i} (\mathsf{x})$ with $w_{\mathsf{LP}\text{-}\mathsf{bc},i} (\mathsf{x})$ and $\tilde\varepsilon_i$ as defined in Section (ref). The following lemma formalizes this result.

lemmaUnder Assumptions (ref) and (ref), $\hat{v}_{\mathsf{mPLP},n}^2 \operatorname*{\to_{\mathrm{p}}} v_{\mathsf{mPLP}}^2 >0$.

Because $T^*_{\mathsf{mLP},n}$ satisfies Assumption (ref), we can apply Theorem (ref) to conclude that the mPLP interval $\mathrm{CI}_{\mathsf{mPLP}\hspace{-0.5pt},n}$ is asymptotically valid and equivalent to

equation*[equation* omitted — 191 chars of source]

which is an RBC-type interval involving no resampling. This interval is based on a new bootstrap bias correction, $\hat{B}_{\mathsf{mLP},n}$, and a new studentization, $\hat{v}_{\mathsf{mPLP},n}$, and it is valid for both interior and boundary points as stated in the next theorem.

theoremUnder Assumptions (ref) and (ref), for any $\mathsf{x} \in \mathbb{S}_{x}$, $\operatorname{\mathbb{P}}(g(\mathsf{x})\in \mathrm{CI}_{\mathsf{mPLP}\hspace{-0.5pt},n})\to 1-\alpha$ and $\mathrm{CI}_{\mathsf{mPLP}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{mPLP} \textsf{-} \mathsf{RBC},n} + \operatorname*{o_{\mathrm{p}}} ((nh)^{-1/2})$. If, in addition, $T_{\mathsf{LP},n}^{\ast}$ is Gaussian conditional on $\mathscr{D}_n$, then $\mathrm{CI}_{\mathsf{mPLP}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{mPLP} \textsf{-} \mathsf{RBC},n}$ a.s.

An important feature of the mPLP approach is that it automatically adapts to the location of the point $\mathsf{x}$. That is, it coincides with the PLP approach asymptotically when $\mathsf{x}$ lies in the interior of $\mathbb{S}_x$ (where $Q_n\operatorname*{\to_{\mathrm{p}}} Q=1$), and it provides a valid generalization when $\mathsf{x}$ is on the boundary.

Finally, as in Section (ref), we can show that confidence intervals for $g(\mathsf{x})$ are asymptotically shorter when based on mPLP compared to the existing RBC intervals. This result, which holds for both interior and boundary points, follows by comparing the asymptotic variances $v^2_{\mathsf{mPLP}}$ and $v^2_{\mathsf{RBC}}$.

corollaryLet Assumptions (ref) and (ref) hold. If $\mathsf{x}$ is an interior point then $v^2_{\mathsf{mPLP}} = v^2_{\mathsf{PLP}}$, and the conclusions of Corollary (ref) hold. If $\mathsf{x}$ is a boundary point, then \begin{equation*} (v^2_{\mathsf{mPLP}}, v^2_{\mathsf{RBC}})= \frac{\sigma^2 (\mathsf{x})}{f(\mathsf{x})} (\mathcal{K}_{\mathsf{mPLP}}, \mathcal{K}_{\mathsf{RBC}}) , \end{equation*} where $\mathcal{K}_{\mathsf{mPLP}}$ and $\mathcal{K}_{\mathsf{RBC}}$ are functions only of the kernel $K$ and the polynomial order $p$.

In contrast to Corollary (ref), Corollary (ref) applies to any evaluation point $\mathsf{x}$, including both interior and boundary points. The values of the constants $\mathcal{K}$ depend on the kernel function $K$, the polynomial order $p$, as well as on whether $\mathsf{x}$ is in the interior or at the boundary of $\mathbb{S}_x$. When $\mathsf{x}$ is in the interior of $\mathbb{S}_x$, these values coincide with those of $\mathcal{K}_{\mathsf{PLP}}$ and $\mathcal{K}_{\mathsf{RBC}}$ in Corollary (ref), implying that Corollary (ref) is a special case of Corollary (ref).

table[table omitted — 646 chars of source]

When $\mathsf{x}$ is a boundary point, $\mathcal{K}_{\mathsf{RBC}}$ takes different values compared with the interior point case, and $\mathcal{K}_{\mathsf{mPLP}}$ is different from $\mathcal{K}_{\mathsf{PLP}}$. Table (ref) is the boundary analogue of Table (ref) and confirms that, for the five kernels previously considered, the modified prepivoted interval $\mathrm{CI}_{\mathsf{mPLP}\hspace{-0.5pt},n}$ is asymptotically shorter than $\mathrm{CI}_{\mathsf{RBC},n}$. As in Figure (ref), we also plot in Figure (ref) the equivalent kernels of mPLP and RBC (see Appendix (ref)). The comparison reveals that the conclusions from Section (ref) also hold for mPLP, i.e., $\mathcal{K}_{\mathsf{mPLP}}<\mathcal{K}_{\mathsf{RBC}}$.

figure[figure omitted — 460 chars of source]

application to regression-discontinuity designs

In this section, we apply the modified prepivoting approach from Section (ref) to the problem of inference in RDD. Adopting the standard potential outcome framework, let $y_i(1)$ and $y_i(0)$ denote the potential outcomes for individual $i$ ($i=1, \ldots,n$) with and without the treatment, respectively. For each individual, we observe the associated treatment indicator $d_i$, which equals 1 if unit $i$ is treated (0 otherwise), and the observed outcome, $y_i=y_i(0)+(y_i(1)-y_i(0))d_i$. We also observe the `forcing' variable $x_i$, a scalar covariate which is not affected by the treatment and determines whether $d_i=1$. Hence, the observed data are $\mathscr{D}_n := \{(y_i, x_i, d_i): i=1,\dots,n\}$.

We consider (sharp) RDD, where $d_i$ is determined by $x_i$ being above a given known cutoff $\mathsf{x}$; that is, $d_i:=\mathbb{I}_{\{x_i\geq \mathsf{x}\}}$. Interest is in estimating the average treatment effect at the cutoff, namely $\mathsf{ATE}(\mathsf{x}) := \operatorname{\mathbb{E}}[y_i(1)-y_i(0)|x_i= \mathsf{x}] = g_{+}(\mathsf{x})-g_{-}(\mathsf{x}) $, where $g_{+}(x) := \operatorname{\mathbb{E}}[y_i(1)|x_i = x]$ and $g_{-}(x) := \operatorname{\mathbb{E}}[y_i(0)|x_i = x]$ are the regression functions of the potential outcomes. We consider the following modification of Assumption (ref), where $\sigma^2_{+}(x) := \operatorname{\mathbb{V}}[y_i(1)|x_i = x]$ and $\sigma^2_{-}(x) := \operatorname{\mathbb{V}}[y_i(0)|x_i = x]$.

assumption$(y_i,x_i,d_i)$ are i.i.d.\ and $x_i$ has bounded support $\mathbb{S}_x$ and continuous density $f$. For all $x$ in an open neighborhood of $\mathsf{x}$, it holds that (i) $f(x)>0$; (ii) $\sigma_{+}^2$ and $\sigma_{-}^2$ are continuous and $\sup_{x\in\mathbb{S}_x}\operatorname{\mathbb{E}}[y_i^4|x_i=x]<\infty$; (iii) $g_{+}^{(p+1)}$ and $g_{-}^{(p+1)}$ are H\"{o}lder continuous with exponent $\eta>0$.

This assumption, which will be used for the asymptotic analysis, also identifies $\mathsf{ATE}(\mathsf{x})$ as the jump in the regression function $g$ at the cutoff $\mathsf{x}$, i.e., as $\tau(\mathsf{x}):= \lim_{x\to \mathsf{x}^{+}}g(x)-\lim_{x\to \mathsf{x}^{-}}g(x)=g_{+}(\mathsf{x})-g_{-}(\mathsf{x}) $; see Hahn et al.\ (2001) and Imbens and Kalyanaraman (2012). Hence, $\tau(\mathsf{x})$ is the parameter of interest, and the average treatment effect can be estimated as the difference of two local polynomial regressions at the cutoff $\mathsf{x}$. Formally, we can write this estimator as

equation[equation omitted — 113 chars of source]

where $\hat{g}_{+,n} (x)$ and $\hat{g}_{-,n} (x)$ are defined as in \hyperref[{eq LP estimator}]{\tagform@{\ref*{eq LP estimator}}} with $K((x_i-x)/h)$ replaced by $K_{+}((x_i-x)/h):=K((x_i-x)/h)\mathbb{I}_{\{x_i \geq \mathsf{x}\}}$ and $K_{-}((x_i-x)/h):=K((x_i-x)/h)\mathbb{I}_{\{x_i < \mathsf{x}\}}$, respectively.

To construct valid confidence intervals using the modified prepivoting method of Section (ref), let $T_n:=(nh)^{1/2}(\hat{\tau}_n(\mathsf{x})-\tau (\mathsf{x}))$ and its LP bootstrap analogue $T_n^{\ast}:=(nh)^{1/2}(\hat{\tau}_n^{\ast}(\mathsf{x})-\hat{\tau}_n(\mathsf{x}))$, where $\hat{\tau}_n^{\ast}(\mathsf{x})$ is calculated from bootstrap data, $\mathscr{D}_n^*:=\{(y_i^*,x_i^*,d_i^*): i=1,\dots, n\}$. Specifically, $(x_i^*,d_i^*) = (x_i, d_i)$, $i=1,\dots, n$, and the $y^*_i$'s are as in \hyperref[{btsDGPLP}]{\tagform@{\ref*{btsDGPLP}}} with $\hat{g}_n(x_i)$ replaced by $\hat{g}_{+,n}(x_i)\mathbb{I}_{\{x_i \geq \mathsf{x}\}}+\hat{g}_{-,n}(x_i)\mathbb{I}_{\{x_i < \mathsf{x}\}}$.

Next, we define the modified bootstrap statistic, $T_{\mathsf{rd},n}^{\ast}$. Since estimation of $\tau$ requires fitting two distinct LP regressions, it is useful to decompose $T_n=T_{+,n}-T_{-,n}$, where $T_{+,n}:=(nh)^{1/2}(\hat{g}_{+,n}(\mathsf{x})-g_{+}(\mathsf{x}))$ and $T_{-,n}:=(nh)^{1/2}(\hat{g}_{-,n}(\mathsf{x})-g_{-}(\mathsf{x}))$. In the same way, we can decompose $T_n^{\ast}=T_{+,n}^{\ast}-T_{-,n}^{\ast}$ and define the scaling factors $Q_{+,n}$ and $Q_{-,n}$ as in Section (ref) with $K$ replaced by $K_{+}$ and $K_{-}$, respectively. Thus, the modified bootstrap statistic is $T_{\mathsf{rd},n}^{\ast}:=Q_{+,n}T_{+,n}^{\ast}-Q_{-,n}T_{-,n}^{\ast}$.

The bootstrap bias is

equation[equation omitted — 142 chars of source]

where $\hat{B}_{+,n}:=\operatorname{\mathbb{E}}^{\ast}[T_{+,n}^{\ast}]$ and $\hat{B}_{-,n}:=\operatorname{\mathbb{E}}^{\ast}[T_{-,n}^{\ast}]$ are computed as in \hyperref[{eq Bhat LP}]{\tagform@{\ref*{eq Bhat LP}}} with $K$ replaced by $K_{+}$ and $K_{-}$, respectively, in the definition of the weights $w_i(x)$. Finally, we need a consistent estimator $\hat{v}_{\mathsf{rd},n}^2$ of the asymptotic variance $v_{\mathsf{rd}}^2$ of $T_n-\hat{B}_{\mathsf{rd},n}$. To this end, we note that $\operatorname{\mathbb{V}} [T_n-\hat{B}_{\mathsf{rd},n}|\mathscr{X}_n] = \operatorname{\mathbb{V}} [T_{+,n}-Q_{+,n}\hat{B}_{+,n}|\mathscr{X}_n]+\operatorname{\mathbb{V}} [T_{-,n}-Q_{-,n}\hat{B}_{-,n}|\mathscr{X}_n]$. Consistent estimators of these two variance components can be obtained as in \hyperref[{eq:hatvmPLP}]{\tagform@{\ref*{eq:hatvmPLP}}} with $K$ replaced by $K_{+}$ and $K_{-}$, respectively, thus yielding $\hat{v}_{\mathsf{rd},n}^2$.

We show in the next theorem that $T^{\ast}_{\mathsf{rd},n}$ satisfies Assumption (ref), so we can apply Theorem (ref) to conclude that a modified confidence interval for $\tau (\mathsf{x})$, denoted $\mathrm{CI}_{\mathsf{rd}\hspace{-0.5pt},n}$ and constructed using $\hat{v}_{\mathsf{rd},n}^2$ as in Sections (ref) and (ref), is asymptotically valid and equivalent to

equation*[equation* omitted — 189 chars of source]

which is an RBC-type interval involving no resampling.

theoremUnder Assumptions (ref) and (ref) it holds that (i) $\hat{v}_{\mathsf{rd},n}^2\operatorname*{\to_{\mathrm{p}}} v_{\mathsf{rd}}^2$, (ii) $\operatorname{\mathbb{P}}( \tau (\mathsf{x})\in \mathrm{CI}_{\mathsf{rd}\hspace{-0.5pt},n})\to 1-\alpha$, (iii) $\mathrm{CI}_{\mathsf{rd}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{rd} \textsf{-} \mathsf{RBC},n} + \operatorname*{o_{\mathrm{p}}} ((nh)^{-1/2})$, and (iv) if, in addition, $T_{\mathsf{rd},n}^{\ast}$ is Gaussian conditional on $\mathscr{D}_n$, then $\mathrm{CI}_{\mathsf{rd}\hspace{-0.5pt},n}=\mathrm{CI}_{\mathsf{rd} \textsf{-} \mathsf{RBC},n}$ a.s.

Finally, we note that all the results in this section can be generalized to allow for different kernel and bandwidth choices on each side of the cutoff. This changes the efficiency results in Sections (ref) and (ref) only slightly. Specifically, let $\mathcal{K}_{+,\mathsf{rd}}$ and $\mathcal{K}_{-,\mathsf{rd}}$ denote $\mathcal{K}_{\mathsf{mPLP}}$ for the kernel used to the right and to the left of the cutoff, respectively. We similarly define $\mathcal{K}_{+,\mathsf{RBC}}$ and $\mathcal{K}_{-,\mathsf{RBC}}$. By the i.i.d.\ assumption, the following result then follows from Theorem (ref) and Corollary (ref).

corollaryUnder Assumptions (ref) (applied to both sides of the cutoff) and (ref), \begin{equation*} \frac{v^2_{\mathsf{rd}}}{v^2_{\mathsf{RBC}-\mathsf{rd}}} = \frac{\sigma_{+}^2(\mathsf{x})\mathcal{K}_{+,\mathsf{rd}}+\sigma_{-}^2(\mathsf{x})\mathcal{K}_{-,\mathsf{rd}}}{\sigma_{+}^2(\mathsf{x})\mathcal{K}_{+,\mathsf{RBC}}+\sigma_{-}^2(\mathsf{x})\mathcal{K}_{-,\mathsf{RBC}}}, \end{equation*} where $v^2_{\mathsf{RBC}\textsf{-}\mathsf{rd}}$ is the asymptotic variance from the existing $\mathsf{RBC}$ interval for RDD.

As previously, the asymptotic relative length of the confidence intervals is the square root of the ratio given in Corollary (ref). Because $\mathcal{K}_{+,\mathsf{rd}}<\mathcal{K}_{+,\mathsf{RBC}}$ and $\mathcal{K}_{-,\mathsf{rd}}<\mathcal{K}_{-,\mathsf{RBC}}$ for all the kernels considered (as seen in Table (ref)), it follows from Corollary (ref) that our new confidence intervals $\mathrm{CI}_{\mathsf{rd}\hspace{-0.5pt},n}$ and $\mathrm{CI}_{\mathsf{rd} \textsf{-} \mathsf{RBC},n}$ are asymptotically shorter compared to the existing RBC intervals for RDD. Indeed, when the same kernel is applied on both sides of the cutoff, as is common in applied work, the result in Corollary (ref) simplifies to that obtained in Corollary (ref) and displayed in Table (ref). Specifically, the new intervals are shorter than the existing RBC intervals by the same amount as in Table (ref).

It is likely that these results can be extended to other types of RDDs, e.g., fuzzy or kinked RDD. Although relatively straightforward conceptually, such extensions are nontrivial. For example, the fuzzy RDD estimator is a ratio of two sharp RDD estimators. Thus, limit theory would require joint convergence of those two estimators, which in turn requires a generalization of the prepivoting theory of Cavaliere et al.\ (2024) to vector-valued statistics. With such theory in hand, the fuzzy RDD estimator can be analyzed using the arguments above combined with the delta method.

monte carlo

We now discuss the finite sample performance of the proposed CIs and compare them with the CIs based on existing RBC using Monte Carlo simulation. We also include the invalid (i.e., not prepivoted) CIs for comparison. We consider two distinct inference problems: a nonparametric regression curve evaluated at both an interior and a boundary point and a sharp RDD.

We report results for $5,000$ Monte Carlo replications and nominal level $0.95$. Estimators are based on the Epanechnikov kernel for the nonparametric regression setup and on the triangular kernel for the RDD setup, as those represent popular kernel choices. Two relevant bandwidth choices are considered: the infeasible MSE-optimal bandwidth ($h$), as a theoretical benchmark, and the feasible (plug-in based) coverage-error-optimal bandwidth ($\hat h$) of Calonico et al.\ (2018, 2020, 2022), representing the reference bandwidth if the aim is to minimize the coverage error of RBC intervals. For each bandwidth choice, we report the average bandwidth ($\bar h$) across Monte Carlo replications. We report empirical coverage probabilities and average lengths for four methods: $\mathrm{CI}_{\mathsf{GP}\hspace{-0.5pt},n}$, $\mathrm{CI}_{\mathsf{LP}\hspace{-0.5pt},n}$, $\mathrm{CI}_{\mathsf{RBC},n}$, and $\mathrm{CI}_{\mathsf{mPLP}\hspace{-0.5pt},n}$. The first two methods are based on the interval $\mathrm{CI}_{\mathsf{Boot},n}$ in \hyperref[{eq general invalid BS}]{\tagform@{\ref*{eq general invalid BS}}}, using the GP and LP bootstrap schemes, respectively. $\mathrm{CI}_{\mathsf{RBC},n}$ is the interval \hyperref[{eq RBC int}]{\tagform@{\ref*{eq RBC int}}} and $\mathrm{CI}_{\mathsf{mPLP}\hspace{-0.5pt},n}$ is the interval defined in Section (ref). For all methods, we use HC3 residuals. The bootstrap intervals are all based on a Gaussian wild bootstrap scheme, and hence implemented analytically without resampling; see Section (ref). For the same reason, results for the prepivoted GP interval $\mathrm{CI}_{\mathsf{PGP}\hspace{-0.5pt},n}$ defined in Section (ref) would be identical to $\mathrm{CI}_{\mathsf{RBC},n}$, and are thus not reported.

table[table omitted — 2,112 chars of source]

nonparametric regression. We consider i.i.d.\ data generated as $y_i =g(x_i)+\varepsilon_i$ with $x_i\sim U_{[-1,1]}$ and $\varepsilon_i \sim N(0,\sigma^2)$, where $\sigma=1$ and

equation*[equation* omitted — 76 chars of source]

This DGP was previously considered in Berry, Carroll, and Ruppert (2001), Hall and Horowitz (2013), and Calonico et al.\ (2018, 2022), among others. We consider inference both at an interior point, $\mathsf{x}=-1/3$, and a boundary point, $\mathsf{x}=-1$, denoted int and bnd, respectively, in Table (ref).

table[table omitted — 2,129 chars of source]

Results for samples of size $n \in \{250, 500, 1000, 2000\}$ are presented in Table (ref). The numerical evidence supports the anticipated decrease in interval lengths of 17%, even for small sample sizes, suggesting that the asymptotic efficiency of mPLP is rapidly achieved as \( n \) increases. We observe that RBC and mPLP perform similarly in terms of coverage under both bandwidth choices, with empirical coverage probabilities remaining close to the nominal level. A deviation from the nominal level is detected for mPLP with $\hat{h}$, when $\mathsf{x}$ is an interior point and $n$ is small, but this deviation rapidly vanishes as the sample size increases. Finally, because they are not robust to the `large' bandwidth choices considered, non-prepivoted methods exhibit much smaller average lengths, resulting in severe undercoverage even for large $n$.

rdd. In the context of RDD, we generate i.i.d.\ data from $y_i = g (x_i)+\varepsilon_i$ with $x_i\sim 2 \mathcal{B}(2,4) -1$ and $\varepsilon_i \sim N(0,\sigma^2)$, where $\mathcal B$ is the Beta distribution and $\sigma=0.1295$. We consider two choices for the regression function $g$ based on the datasets in Ludwig and Miller (2007) and Lee (2008), as in Calonico et al. (2014). Specifically, the two DGPs are

align*[align* omitted — 565 chars of source]

where, as in Calonico et al.\ (2014), the fifth-order polynomial from the Lee (2008) data in DGP2 is modified to increase the curvature of $g$ and hence increase bias.

Results for samples of size $n\in\{500,1000, 2000,4000\}$ (total sample sizes are twice those in the nonparametric regression example) are presented in Table (ref). The results confirm the conclusions drawn for the nonparametric regression example: the mPLP method produces shorter intervals than RBC for all the considered sample sizes and bandwidth choices, with coverage levels of both prepivoted methods close to the nominal value. As expected, non-prepivoted methods fail to deliver valid inference.

table[table omitted — 833 chars of source]

guidance for applied researchers

The mPLP method proposed in this paper can be used as an alternative or complement to the standard RBC interval in any nonparametric regression or RDD application. Its implementation requires no additional tuning parameters beyond those already needed for RBC: the same bandwidth $h$, kernel $K$, and polynomial order $p$ are used throughout. R packages implementing our procedures are available at \url{https://pppackages.github.io}.

Table (ref) summarizes the key choices and considerations for practitioners. We offer the following specific recommendations.

\noindentbandwidth. Our method is compatible with any bandwidth selection rule used with RBC, including the coverage-error-optimal bandwidth of Calonico et al.\ (2018) implemented in the rdrobust package, the MSE-optimal bandwidth, or cross-validation. No other bandwidth is required.

\noindentkernel. Our efficiency results hold for all standard kernels. For practitioners already using RBC with the Epanechnikov or triangular kernel, the two most common choices,mPLP delivers 17% and 16% shorter intervals, respectively (Tables (ref) and (ref)). There is no reason to switch kernels when moving from RBC to mPLP.

\noindentpolynomial order. As with RBC, local linear estimation ($p=1$) is the standard choice and the one that requires least smoothness of the conditional mean function. Higher-order polynomials are supported by the theory and can be implemented if one can assume additional smoothness.

\noindentinterior vs.\ boundary/rdd. We recommend using the modified mPLP interval for both interior evaluation points, boundary points, and RDD (where the cutoff is a boundary point) because it adapts automatically to the boundary and requires no additional input from the user.

\noindentcomputation. Because the bootstrap moments (mean and variance) entering the mPLP interval are available in closed form as functions of the kernel weights and residuals, implementation is fully analytic and requires no resampling.

conclusions

This paper proposes novel procedures for inference in nonparametric regression and regression-discontinuity designs based on non-standard implementations of the bootstrap. New confidence intervals based on the concept of prepivoting (Beran 1987, 1988; Cavaliere et al., 2024) are shown to deliver asymptotically correct coverage under general conditions that allow for the presence of bias and thus do not require undersmoothing. We show that prepivoting different choices of bootstrap DGPs yield different robust bias correction (RBC)-type confidence intervals. This connection between the prepivoting and robust bias correction approaches allows us to identify the specific bootstrap DGP underlying the RBC approach of Calonico et al.\ (2014, 2018). More importantly, it enables us to propose a new alternative approach based on a local polynomial bootstrap algorithm that delivers improved inference. Specifically, the prepivoted local polynomial method yields 14--17% shorter intervals than existing RBC intervals, depending only on the choice of kernel function.

It is worth noting that, because the conditional moments (expectation and variance) of our reference bootstrap statistics can be computed analytically, the practical implementation of our improved bias correction procedure does not require simulating any bootstrap samples. Although we could use resampling instead of the analytical formulas to make the implementation more automatic, this would be much more computationally costly.

The application of prepivoting to inference in the presence of nonnegligible bias has potential applications beyond those explored here. A natural extension is to sieve regression estimators, which have a long tradition in the nonparametrics literature (see, e.g., Andrews, 1991; Huang, 2003; Chen, 2007; Belloni et al., 2015; Chen and Christensen, 2015). Although bias is often assumed away through undersmoothing conditions, Cattaneo et al.\ (2020) have recently proposed RBC inference methods for the particular case of partitioning-based nonparametric estimators. It would be interesting to explore whether prepivoting can improve inference in this setting.

Other potential applications include two-step semiparametric estimators where the first step involves nonparametric regression, potentially affecting the asymptotic distribution of the second step estimator (e.g., Andrews, 1994; Newey, 1994; Chen, Linton, and van Keilegom, 2003). While this literature has primarily focused on how to adjust the variance of the second step estimator, Cattaneo and Jansson (2018) show that asymptotic bias emerges under `small bandwidth' asymptotics when the first step uses kernel regression. Although they show that a particular bootstrap automatically corrects for this bias, their results assume away smoothing bias. Exploring whether prepivoting can improve inference in this setting would be an interesting contribution. Finally, extensions to time series and high-dimensional frameworks (Gupta and Seo, 2023) or spatial data (Hallin et al., 2004) are also of interest.