Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
112,076 characters · 37 sections · 92 citation commands
Kernel Choice Matters for Local Polynomial Density Estimators at Boundaries
{Keywords: Asymptotic efficiency, finite sample theory, kernel selection, local polynomial fitting, manipulation tests, statistical power}
\doublespacing
The probability density function at boundary points is often of interest in empirical analyses across economics, political science, and applied statistics. For instance, certain empirical tests of economic theory rely on density at boundary points (e.g., Collin_Talbot:2023), and discontinuity detection has been used to identify social norms (Bertrand_etal:2015), study a tax-evasion behavior of taxpayers (Breunig_etal:2024), detect $p$-hacking (Elliott_etal:2022), and test the score manipulation in regression discontinuity (RD) analysis (see Cattaneo_Titiunik:2022, Cattaneo_etal:2023med).
For such purposes, where the value of the density at the boundary is of interest, one of the most widely used approaches is the local polynomial density (LPD) estimator proposed by Cattaneo_etal:2020. The LPD estimator is based on the local polynomial kernel smoothing and, as the authors note, it “enjoys all the desirable features associated with local polynomial regression estimation," including boundary adaptation. Due to this superior bias property at boundary points, the LPD estimator has been widely adopted in empirical economics (e.g., Britto_etal:ECTA, Chen_etal:JDevEcon, Connolly_Haeck:JLaborEcon, Dasgupta_etal:REStat, Gorrin_etal:2023JIE, He_etal:2020, Khanna:JPE, among others).
However, despite its popularity, our numerical simulations reveal several variance-related issues with the LPD at the boundary. First, under the commonly used triangular kernel, confidence intervals (CIs) are often wide near the boundary, and discontinuity tests can be so underpowered that they frequently fail to detect discontinuities, even in large samples. While this issue is not explicitly discussed in the literature, similar patterns of wide CIs appear in several empirical studies, e.g., Breunig_etal:2024, DeBenedetto_etal:2025, Forderer_Burtch:2025, and Keefer_Vlaicu:2025. This feature may limit our ability to draw reliable conclusions about the boundary behavior of the density function or to detect economically important empirical observations, such as bunching.
Furthermore, in addition to this large-sample variance issue, our simulation also indicates that the variance of the LPD estimator under the triangular kernel effectively diverges when the sample size is small. This finite-sample variance property further raises concerns about the credibility of boundary estimation and inference based on the LPD estimator.
Interestingly, our simulations suggest that these problematic properties of the LPD estimator are substantially mitigated by using Gaussian or Laplace kernels, although these choices are uncommon in the local polynomial smoothing literature. In particular, relative to the commonly used triangular kernel, switching to these less common kernels reduces the mean squared error (MSE), shortens confidence intervals, increases statistical power in discontinuity detection, and suppresses the finite-sample variance explosion. Our first contribution is to document this practically important observation for applied work.
Motivated by these numerical observations, this article theoretically studies and highlights a close connection between kernel choice and these properties of large- and small-sample variance. First, in sharp contrast to conventional wisdom in kernel estimation literature, we argue that kernel selection has a nontrivial effect on the LPD’s asymptotic efficiency. For example, the commonly used triangular kernel is 14% less efficient compared to our preferred Laplace kernel in terms of MSE, a sizable gap given that replacing the optimal Epanechnikov kernel with, for example, a Gaussian kernel reduces efficiency by only about 4% in the standard setting (Hansen:2022_prob). Moreover, for inference, the efficiency loss incurred by using the triangular kernel instead of the Laplace amounts to about 50%. Furthermore, in contrast to the standard local polynomial regression literature calonico2022coverage, the uniform kernel is an even worse option; its efficiency loss amounts to approximately 75%. Intuitively, these values imply that achieving the same interval length as with the Laplace kernel requires approximately 1.6-1.9 times the sample size. Building on this analysis, we further study the statistical power of the LPD-based discontinuity test against a $\sqrt{nh}$ local alternative. In contrast to the fixed alternative case studied by Cattaneo_etal:2020, the kernel choice has a first-order contribution to the asymptotic power property, and we find that the Gaussian and Laplace kernels can improve the power over the commonly employed kernel functions.
Why does the performance of the LPD estimator depend so strongly on the choice of kernel? To build intuition for these efficiency gains, we study the equivalent kernels of the LPD estimator. By deriving the infeasible optimal weighting function and comparing it with the equivalent kernels, we find that popular kernels deviate substantially from the optimum, thereby implying sizable efficiency gains from employing alternative kernel functions. Moreover, this analysis indicates that the gains are rooted in the boundary-kernel methods of Gasser_etal:1985 and Muller:1991, although the numerical importance of kernel choice has been largely overlooked in the literature. Taken together, our results provide new insights for the standard kernel-smoothing literature.
Regarding the small-sample variance inflation issue, we formally show that the LPD estimator inherits not only the “desirable features" of local polynomial smoothing techniques but also an undesirable finite-sample property: when a compactly supported kernel is used, the estimator has no finite variance Seifert_Gasser:1996. In particular, we show that the finite-sample variance is infinite when using compactly supported kernels such as the triangular and uniform kernels, whereas it is bounded when employing an unbounded support kernel such as the Laplace kernel.
This small-sample variance property is particularly consequential in score manipulation tests in RD designs, where one-sided manipulation always reduces the effective sample size on one side of the cutoff, potentially making the finite-sample variance problem dominant. In this view, the use of an unbounded support kernel is particularly recommended in manipulation testing settings---one of the most widespread applications of LPD estimators.
Taken all together, our results show that kernel choice matters both asymptotically and in finite samples. Although it is commonly believed that, for kernel-based estimators, “the choice of kernel function $K$ is not very important for the performance of the resulting estimators, both theoretically and empirically" (Fan_Gijbels:1996), the findings in this paper suggest that this (generally valid) understanding does not apply to the boundary estimation and inference using the LPD estimators. Careful selection of the kernel function can substantially enhance the performance of LPD estimation and inference. This simple yet powerful modification yields significant efficiency gains: relative to commonly used kernels, the improvements are considerable both theoretically and empirically.
\paragraph{Plan of the Article.} This paper proceeds as follows. Section (ref) presents numerical and empirical examples that motivate our analysis. Section (ref) studies the asymptotic efficiency of the LPD estimator in both estimation and inference settings, and then examines statistical power under local alternatives. We also provide intuition for the efficiency gains from kernel choice through an equivalent-kernel analysis. Section (ref) investigates the finite-sample variance properties of the LPD estimator. Section (ref) summarizes practical recommendations for empirical researchers. Section (ref) concludes. All proofs, together with additional discussion of interior points, are provided in the Online Appendix. The links to the kernel estimation literature are made explicit in Sections (ref) and (ref), and hence, to conserve space, we refrain from providing a separate literature review.
Kernel choice is often regarded as less important. To illustrate the potential importance of kernel selection in the LPD estimation and inference, we begin by presenting a piece of numerical evidence.
We begin with one of the most relevant scenarios, discontinuity detection (or manipulation testing) in RD designs. We consider the following data-generating process, which is designed to mimic a one-sided score manipulation.
That is, we consider a scenario where some observations that would otherwise fall in $(c, 0.9)$ are shifted to $[0.9, 0.9 + 0.2]$, with $0.9$ being the cutoff. We examine the following three cases: (1) $q = 0.5$ and $c = 0.8$ (2) $q = 0.5$ and $c = 0.7$, and (3) the no-manipulation case, i.e., $q=0$. A smaller value of $c$ (i.e., case (2)) implies that the manipulation occurs over a wider region. For cases (1) and (2), see Figure (ref) for example histograms.
We perform the discontinuity test by following the procedure proposed in Cattaneo_etal:2020, details of which will be introduced in a later section. The sample sizes are set to $n = 1000$, $750$, and $500$. Each experiment is repeated 2000 times. The nominal significance level is set to $0.05$.
The simulation results are reported in Table (ref). We find that the empirical power significantly differs depending on the kernel choice. In every scenario, the Gaussian and Laplace kernels show higher rejection probabilities than the triangular kernel, the most popular choice in empirical studies. Moreover, the difference in power is numerically significant. In several cases, the power under the Laplace kernel is more than doubled compared to the triangular kernel. Even with a relatively small sample size, the Laplace and Gaussian kernels continue to yield a satisfactory rejection rate, whereas the triangular kernel exhibits a marked deterioration in performance. Furthermore, in case (3) where no manipulation occurs, the rejection probability is less sensitive to the kernel selection and it is close to the nominal level in every case. These simulation results suggest that the kernel choice is consequential in discontinuity detection using the LPD estimator.
In the previous simulation, we observed that statistical power varies with the kernel choice. This result suggests that the efficiency property of the LPD estimator may crucially depend on the kernel selection. We here provide an empirical example supporting this conjecture.
Our first application adopts data from Eggers_etal:2021JME. This study employs RD designs and examines the presence of score manipulation using the LPD-based discontinuity test. Figure (ref) replicates Eggers_etal:2021JME. We can see that the CIs around the cutoff point are excessively wide, particularly on the right-hand side. This raises concerns about whether a “no-discontinuity" conclusion can be drawn with confidence, even if the test does not reject the null hypothesis of continuity. By contrast, the Gaussian and Laplace kernels (Figures (ref)-(ref)) show much tighter CIs. For example, the CIs under the Laplace kernel are $16\%$ shorter on the left and $41\%$ shorter on the right. Although the CIs are much tighter, they still exhibit overlap around the cutoff point (and the formal test does not reject the continuity). This persistent overlap, despite the narrower intervals, suggests that manipulation is unlikely and thereby strengthens our confidence in the validity of the main RD study.
As an additional empirical example, we use data from Lindo_etal:2010, which is also an RD study. Figure (ref) reports the density estimates together with CIs, where the CIs exhibit moderate overlap. Consistent with this, the discontinuity test yields a $p$-value of $0.082$, so we fail to reject continuity at the conventional nominal level. Under the Gaussian kernel, the conclusion is unchanged, while the CIs become slightly shorter. Using the Laplace kernel, the CI length decreases further (by $24\%$ on the left and $16\%$ on the right compared to the triangular kernel) with limited overlap. Consistently, the discontinuity test rejects continuity, with a $p$-value of $0.02$. This discontinuity detected under the Laplace kernel is consistent with the binomial test performed in cattaneo2023practical, which is based on the local randomization framework proposed in Cattaneo_etal:2015.
To see the variance property more directly, we next consider the standard density estimation setting at a boundary point. The data-generating process is the standard normal distribution $\mathcal{N}(0,1)$ truncated below at $-0.8$, which is adopted from Cattaneo_etal:2020. We are interested in $f(-0.8)$. We perform estimation and inference using the lpdensity package with default options.
Table (ref) reports the empirical coverage probabilities, average CI lengths, and MSEs based on 2000 iterations. Although the coverage probabilities are nearly identical across kernels and close to the nominal level, the average CI length with the Laplace kernel is substantially shorter than with the triangular kernel. In every case, the Laplace kernel yields CIs about 25% shorter than those of the triangular kernel. The Gaussian kernel also shows similar improvement, although modest compared to the Laplace. The MSEs follow a similar pattern: The Gaussian and Laplace kernels attain an MSE approximately 10-25% lower than that of the triangular kernel.
To eliminate the influence of the variation in the selected bandwidth, Table (ref) also reports the performance based on the (infeasible) asymptotically optimal bandwidths at every iteration. Once again, we observe the same pattern. The average length and MSE under the Laplace or Gaussian kernels are much shorter/smaller than under the triangular kernel.
Finally, we provide simulation evidence on the small-sample behavior of the LPD estimator. We again use the truncated normal distribution as in Section (ref), but now consider cases where the sample size is much smaller ($n\in \{400, 350, \ldots, 100, 75, 50\}$). The empirical relevance of such small-sample settings will be discussed in Section (ref); here, we simply report and interpret the simulation results.
Figure (ref) reports the variance (in the logarithmic scale) of the LPD estimator based on $1000$ iterations for each sample size. We find that the finite-sample variance exhibits a very different pattern depending on the kernel function used. When the popular triangular kernel is used, the small-sample variance increases explosively as the sample size decreases. In contrast, with the Gaussian or Laplace kernels, the variance inflation is suppressed.\footnote{A similar pattern is observed for the MSE, since the variance becomes so large that it dominates the squared bias (see the Online Appendix, Section S4.1).} Therefore, when the sample size is small---or more broadly, when the local sample size near the evaluation point is limited---the kernel selection can substantially affect the stability of the LPD estimator.
In Sections (ref) and (ref), we observed that the efficiency property depends on the kernel function used, even for large sample sizes. The statistical power was insufficient to detect the discontinuity when using the triangular kernel, the most common choice. The length of the CI was also so large under the triangular kernel that the interpretation of the statistical inference can be unclear. MSE suggested a similar conclusion. In contrast, the LPD estimator using the Gaussian or Laplace kernels performs better. The length of the CI shrinks, and the power property in discontinuity testing improves. MSE is also smaller under the Gaussian or Laplace kernels than the triangular kernel.
Similarly, Section (ref) suggests that the LPD estimator using the triangular kernel is still inferior when the sample size is small. In particular, the finite-sample variance explodes quickly, whereas this explosion is suppressed if we employ the Gaussian or Laplace kernels instead.
In the following sections, we formally investigate the root cause of why the efficiency properties of the LPD estimator depend on the choice of kernel. The analysis in Section (ref) corresponds to Sections (ref) and (ref) in that we investigate the asymptotic properties of the LPD estimator. After that, in Section (ref), we theoretically explain the observation in Section (ref) by establishing a finite-sample theory that does not rely on asymptotic approximations.
In this section, we examine how the asymptotic performance of the LPD estimator depends on the choice of kernel function. The next subsection introduces notation and several preliminary results, after which we study the asymptotic efficiency among two classes of kernel functions in Section (ref). We then investigate the statistical power property in Section (ref). After that, Section (ref) explains why this dependence arises, using the equivalent kernel framework.
Let $\mathcal{X} = (-\infty,x_R) \subset \mathbb{R}$ be the data domain. We are given an independent and identically distributed sample $X_{1:n} = (X_1,\ldots,X_n)$ defined on $\mathcal{X}$ with unknown distribution $F$ with density $f$. We are interested in the density at the boundary, $f(x_R)$. Throughout this article, we assume $f(x_R) >0$. The LPD estimator of Cattaneo_etal:2020 at the boundary point $x_R$ is given by $\hat{f}(x_R)= \bm{e}_1^\prime \hat{\bm{\beta}}(x_R)$, where
where $\bm{e}_1 = (0,1,0,\ldots,0)^\prime$; $\hat{F}(x)=n^{-1}\sum_{i=1}^{n}\mathbf{1}\left\{X_i \leq x\right\}$; $\bm{r}_p(u)=(1, u, u^2, \ldots, u^p)^\prime$; $K(\cdot)$ is a non-negative, symmetric kernel function such that $\int K(u)du=1$; $h(>0)$ is a bandwidth; and $p(\geq1)$ is the polynomial degree.
The asymptotic bias and variance of $\hat{f}(x_R)$ are obtained under standard assumptions in Cattaneo_etal:2020 as
respectively, where $\mathcal{B}_{p,K} = \bm{e}_1^\prime \bm{A}_{p,K}^{-1} \bm{c}_{p,K}$, and $\mathcal{V}_{p,K} = \bm{e}_1^\prime \bm{A}_{p,K}^{-1}\bm{B}_{p,K} \bm{A}_{p,K}^{-1} \bm{e}_1$, with
Then, the asymptotic MSE optimal bandwidth can be deduced from these as $h_p^{\mathtt{MSE}} = ( {\mathcal{V}_{p,K}}/\{C_1(p,x_R,F) \mathcal{B}_{p,K}^2\})^{1/(2p+1)} n^{-1/(2p+1)}$, where $C_1(p,x_R,F)$ is some constant that depends only on $p$, $x_R$, and $F$ Cattaneo_etal:2020. Using these results, the MSE at $h_p^{\mathtt{MSE}}$ is asymptotically given by
with some constant $C_2(p,x_R,F)$ that does not depend on $K$. We will write the constant part depending on the kernel function as $\mathcal{Q}_{p, K} = (\mathcal{B}_{p,K}^2)^{1/(2p+1)} \mathcal{V}_{p,K}^{2p/(2p+1)}$.
We begin by examining how the asymptotic variance and MSE depend on the choice of kernel function. Based on the preliminary results in the previous subsection, we can analyze this dependence through $\mathcal{V}_{p,K}$ and $\mathcal{Q}_{p,K}$, which characterize the contribution of the kernel function to the variance and MSE-efficiency properties of the LPD estimator.\footnote{The usefulness of the indicator $\mathcal{V}_{p,K}$ may be debatable, as it compares the constant term of the asymptotic variance under a fixed bandwidth, which may not reflect recent empirical practice. Even for inference, bandwidth selection rules typically account for bias in some way, so quantities such as $\Theta_K$ (defined later) are arguably more directly relevant empirically. Nevertheless, we report $\mathcal{V}_{p,K}$ here for completeness and to maintain alignment with the classical kernel smoothing literature. For example, Gasser_etal:1985 defines and derives the minimum-variance kernel that minimizes the counterpart of this quantity in standard settings.}
We compare two classes of kernel functions. Following Fan_Gijbels:1995, Fan_Gijbels:1996, our first candidate is the Beta family: $\{(1-u^2)_{+}\}^\gamma/\mathrm{Beta}(1/2, \gamma+1)$. Among this family, we consider the uniform ($\gamma=0$), the Epanechnikov ($\gamma=1$), the biweight ($\gamma=2$), and the Gaussian kernels ($\gamma\to\infty$ with appropriate scaling). In addition to these kernels, to study the case when the kernels are more centered, we also consider a generalized triangular family: $\{(1-|u|)_+\}^m (m+1)/2$. Among this family, we consider the triangular ($m=1$), what we call $2$-triangular ($m=2$), $3$-triangular ($m=3$), and the Laplace kernels ($m\to\infty$). Note that $m=0$ corresponds to the uniform kernel. Together, these two classes cover most of the popular kernel functions used in the kernel smoothing literature.
Our kernel selection is also motivated as follows. First, the uniform, triangular, and Epanechnikov kernels are implemented in the software packages provided by Cattaneo_etal:2018, Cattaneo_etal:2022, and are therefore natural benchmarks. In addition, in standard local polynomial regression, the triangular kernel is the MSE-optimal choice at a boundary point Cheng_etal:1997, while the uniform kernel is a popular choice for inference calonico2022coverage and is the MSE-optimal choice in the LPD estimation at interior points cattaneo2021local. The Epanechnikov kernel is the well-known optimal kernel in the standard kernel estimation at interior points (Epanechnikov:1969), while the biweight kernel is also optimal among kernels with a certain smoothness (Muller:1984). Furthermore, the Gaussian kernel is a standard choice in the kernel estimation literature and has several optimality properties (Cline:1988, Granovsky_Muller:1991). The Laplace kernel is the optimal choice for non-smooth densities in standard kernel estimation vanEeden:1985. Although the $2$- and $3$-triangular kernels seem less standard in the literature, we include them as they are a natural intermediate between the triangular and Laplace kernels.
The asymptotic variance ($\mathcal{V}_{p,K}$) and MSE-efficiency ($\mathcal{Q}_{p,K}$) for these kernels and $p=1,2,3$ are summarized in Table (ref).
We observe that the Laplace kernel not only exhibits better variance properties (at a fixed bandwidth) than commonly used kernels, but also dominates them in terms of asymptotic MSE (at the MSE-optimal bandwidths). Notice that, unlike usual, the kernel choice has a significant impact on the asymptotic MSE-efficiency. For the case when $p=2$, which is the most standard choice in the literature (Cattaneo_etal:2020; Fan_Gijbels:1996), the commonly used kernels are 14--26% less efficient than the Laplace kernel. Intuitively speaking, this implies that the commonly used kernels require approximately 1.2--1.3 times larger sample sizes to achieve the same performance as the Laplace kernel.
This efficiency loss is in contrast to the standard kernel smoothing at interior points, wherein the efficiency loss between the optimal Epanechnikov kernel and popular choices is quite limited. For example, the efficiency loss from using the Gaussian kernel relative to the optimal Epanechnikov kernel is only $4\%$ (e.g, Hansen:2022_prob). In this view, contrary to the common belief in kernel smoothing literature, a careful kernel choice is essential in the boundary estimation using the LPD estimator.
We then study the implications of kernel choice for inference. To perform statistical inference, we have to deal with the asymptotic smoothing bias. Cattaneo_etal:2020 propose to rely on a simple robust bias correction approach (Calonico_etal:2014; calonico2018effect). Specifically, the authors propose using local cubic estimation (i.e., $p=3$) together with the MSE optimal bandwidth for $p=2$ (i.e., $h_2^{\mathtt{MSE}}$). By doing so, the (bias-corrected) LPD estimator, $\hat{f}_{\mathrm{bc}}$, is correctly centered.
The variance of $\hat{f}_{\mathrm{bc}}$ is asymptotically approximated by
where $\Theta_{K} = (\mathcal{B}_{2,K}^2)^{1/5} \mathcal{V}_{2,K}^{-1/5} \mathcal{V}_{3,K}$. Therefore, we can examine how the asymptotic variance depends on $K$ via $\Theta_K$, a key quantity that governs the length of the CIs.
The last column of Table (ref) summarizes $\Theta_K$ for the eight kernels considered before. We observe that the Laplace kernel again performs best among the eight kernels. We also find that the triangular kernel is approximately 50% less efficient than the Laplace kernel, which is a remarkably large gap. In view of the standard local polynomial smoothing literature, one might be tempted to employ the uniform kernel for inference, as it is the CI-length-optimal kernel calonico2022coverage. However, this turns out to be an even worse choice, as it exhibits 75% less efficiency.
These magnitudes are noteworthy. The triangular and uniform kernels require sample sizes approximately 1.6 to 1.9 times larger to achieve the same performance as the Laplace kernel in terms of asymptotic variance or interval length.
Building on the discussion in Section (ref), we here investigate the statistical power property of the LPD-based discontinuity testing. Cattaneo_etal:2020 shows that, under fixed alternatives, the test is consistent (i.e., its power converges to one asymptotically) regardless of the kernel choice. While this fixed-alternative analysis is valuable for establishing consistency, it is not designed to be informative about power comparisons across kernels or about the finite-sample differences we documented in Section (ref); see vanderVaart:1998 for a general discussion. Motivated by this, and to better understand the power behavior in our setting, we study power against a local alternative.
Consider the testing problem $\mathbb{H}_0: f(0+) = f(0-)$, but the following local alternative is correct:
Let $\hat{f}_+(0)$ and $\hat{f}_-(0)$ be the LPD estimators only using the data lying over the right side region and the left side region, respectively. Following the recommendation of Cattaneo_etal:2020 and empirical practice, we specifically consider the case where $p=3$ together with the MSE-optimal bandwidths for $p=2$ (i.e., $h_{2+}^{\mathtt{MSE}}$ and $h_{2-}^{\mathtt{MSE}}$). Then $\tau_n$ can be estimated by $\hat{\tau}_n \coloneqq ({n_+}/{n}) \hat{f}_+(0) - ({n_-}/{n}) \hat{f}_-(0)$. Cattaneo_etal:2020 proposes the statistic
where $\hat{\mathscr{V}}_{+}$ and $\hat{\mathscr{V}}_{-}$ are consistent estiamtors of the asymptotic variances, ${\mathscr{V}}_{+}\coloneqq f(0+)\mathcal{V}_{3,K}$ and ${\mathscr{V}}_{-}\coloneqq f(0-)\mathcal{V}_{3,K}$. To study the power against the local alternative $\mathbb{H}_{1,\mathtt{LA}}$, we make the following assumptions. Fix $c\in\mathbb{R}\backslash\{0\}$.
The assumption (ref)(ii) is the standard regularity conditions on the data-generating process. In particular, this assumption is almost the same as Assumption 2 in the supplemental appendix for Cattaneo_etal:2020. The only notable difference is that we state it for an $n$-varying sequence $\{F_n\}$ to accommodate the local alternative. Plus, we note that we could allow $(s_l,s_r)$ to depend on $n$, while such a generalization would only lengthen the technical discussion and would not yield additional insight for our purposes.
Under these assumptions and $\mathbb{H}_{1,\mathtt{LA}}$, we have the following distributional approximation.
Lemma (ref) justifies the following approximation:
where $Z\sim \mathcal{N}(0,1)$, and $C_4(F)>0$ is a constant depending only on the underlying distribution $F$. We can rewrite as
Then, the power function can be asymptotically approximated by $\P{|T_3| > z_{1-\alpha/2}} = \pi_\alpha(\Theta_K) + o(1)$, where
which is a decreasing function of $\Theta_K$. Therefore, kernels with a smaller $\Theta_K$ are preferable from the standpoint of statistical power. Taken together with the previous section's argument that $\Theta_K$ can vary substantially across kernels, this implies that kernel choice can be of first-order importance for power, and that popular kernels can deliver lower power than, for example, the Laplace kernel.
We have observed that kernel selection has a non-trivial impact on the efficiency of the LPD estimator, affecting both the MSE, CI length, and power. Why does this happen? This subsection investigates the underlying reason for this dependence.
We first aim to develop intuition for why the MSE varies substantially with the choice of kernel function. Our discussion relies on the equivalent kernel Fan_Gijbels:1996, which is helpful in understanding how the LPD applies weights to each datum point.\footnote{The use of equivalent-kernel representations in the context of LPD estimation is not new and is rather common in the literature. For example, cattaneo2021local uses equivalent kernels to study the variance properties of their local regression distribution estimators at interior points and to show that the uniform kernel is MSE-optimal in the LPD estimation at interior points. Our focus, however, is on boundary behavior, and the discussion that follows contains several observations that are novel to both the kernel-smoothing and LPD literatures.} The LPD estimator can be rewritten as follows (see the Online Appendix Section S2 for details):
This expression suggests that the LPD estimator at boundary points is asymptotically equivalent to the standard kernel density estimator (KDE) using the following (asymmetric) kernel function,
Therefore, an investigation of $K^*_{p,K}$ will offer insights into how the LPD utilizes the information contained within the sample. The following lemma characterizes the basic properties of $K^*_{p,K}$.
Lemma (ref) shows that the equivalent kernel $K^*_{p,K}$ is the $p$-th order boundary kernel that satisfies the moment conditions of Gasser_etal:1985, reflecting that the LPD estimator is boundary adaptive. The property of $K^*_{p,K}(0)=0$ arises from our use of $\hat{F}$ as the “dependent variable." At the boundary, $\hat{F}(x_R) - \bm{r}_p(0)^\prime\hat{\bm{\beta}}$ shrinks very quickly Cattaneo_etal:2020, eliminating its contribution to the loss in (ref). As a result, the LPD asymptotically does not use information from the data at the boundary.
A well-known result from earlier studies on boundary kernel methods, such as Muller:1991, suggests that there is no optimal weighting function among kernels satisfying (ref) and $K^*_{p,K}(0)=0$. However, we can show that an optimal weighting scheme exists within an extended class of weight functions that allows $K^*_{p,K}(0) \neq 0$. The Online Appendix (Section S1.5) shows that it corresponds to the equivalent kernel of Cheng_etal:1997's (Cheng_etal:1997) local polynomial density estimator using the triangular kernel. Therefore, the closer $K_{p,k}^*$ is to this optimal weight, the more the MSE of the LPD estimator with kernel $K$ improves.
Figure (ref) displays several equivalent kernels from the generalized triangular family, scaled to have the same asymptotic variance ($\mathcal{V}_{2,K}$) for $p=2$, along with the infeasible optimal weight. We find that the triangular and uniform kernels deviate substantially from the optimal weight. This degree of deviation reflects the room for MSE-efficiency gains.
This deviation suggests that more centered kernels (i.e., those that place more weight near the evaluation point) can deliver improved MSE performance relative to commonly used kernels. Within the Beta and generalized triangular families, this can be achieved by choosing a larger $\gamma$ or $m$. Moreover, the efficiency gains do not appear to be exhausted within these families; consequently, among the kernels we consider, the Gaussian and Laplace kernels yield the best MSE performance. As a result, Figure (ref) shows that the Laplace kernel more closely matches the optimal weight and applies more weight near the evaluation point for a given variance, thereby delivering a lower MSE than standard kernels.\footnote{Figure (ref) suggests that even relative to the Laplace kernel, there remains additional scope for efficiency improvements. Indeed, for example, we can construct a kernel function based on high-order polynomials, and a numerical search can find coefficients that achieve a smaller MSE than that of the Laplace kernel. Although appealing, such a weight may not be monotonically decreasing, and it can be computationally costly when performing estimation. Hence, our preferred choice is the Laplace for its simple form and satisfactory performance in numerical studies in Section (ref). A detailed study on this issue is postponed to future research.}
The analysis above offers a new perspective on standard KDE at boundary points. Classical contributions, such as Muller:1991 and Gasser_etal:1985, develop boundary-kernel methods and find that an MSE-optimal kernel does not exist within the standard kernel class. However, they do not benchmark conventional kernels against the infeasible optimum within an enriched class of weighting functions. As a result, the fact that reasonable alternative kernels---for instance, the equivalent kernels of the Gaussian or Laplace kernels---can deliver numerically substantial MSE efficiency gains at the boundary has been largely overlooked in the kernel smoothing literature. For example, Muller:1991's (Muller:1991) preferred kernel (Muller:1991) coincides with the equivalent kernel of the LPD estimator under the uniform kernel. Nevertheless, this kernel remains far from optimal: as shown in Table (ref), the equivalent kernels induced by the other kernels attain strictly better MSE efficiency. To the best of our knowledge, the magnitude and practical relevance of this efficiency gap have not been emphasized in the existing literature.
We now turn to investigate why kernel choice plays a more prominent role in simple robust bias-corrected inference.
First, suppose the point estimation scenario with polynomial order $p=2$. Under the MSE-optimal bandwidth $h_2^{\mathtt{MSE}}$, the bias and variance are balanced so that the variance is (asymptotically) proportional to $\mathcal{Q}_{2, K}$. Now, if we increase the polynomial order to $p=3$, then the variance increases by a factor of $\mathcal{V}_{3, K} / \mathcal{V}_{2, K}$. Hence, the constant factor of the variance under simple robust bias correction is given by $\mathcal{Q}_{2, K} \times \mathcal{V}_{3, K} / \mathcal{V}_{2, K}$, which a simple calculation shows to be equal to $\Theta_K$. Therefore, in inference, the quantities $\mathcal{Q}_{2,K}$ and $\mathcal{V}_{3,K} / \mathcal{V}_{2,K}$ play central roles.
We have seen that $\mathcal{Q}_{2,K}$ is smaller for kernels with larger values of $m$ or $\gamma$. Moreover, Figures (ref) and S3 (in the Online Appendix) suggest that larger $m$ or $\gamma$ also lead to a smaller increase in variability, $\mathcal{V}_{p+1,K}/\mathcal{V}_{p,K}$, mirroring the patterns observed in standard local polynomial smoothing (Fan_Gijbels:1996).
Hence, both factors (i.e., $\mathcal{Q}_{2, K}$ and $\mathcal{V}_{3, K} / \mathcal{V}_{2, K}$) become smaller when an appropriate kernel is chosen instead of the commonly used ones. Consequently, the asymptotic variance under simple robust bias correction is much smaller with the Laplace kernel, illustrating a reason for the considerable efficiency gains reported in Table (ref).
Thus far, we have examined how the efficiency properties of the LPD estimator depend on the choice of kernel function in the asymptotic sense. However, as observed in Section (ref), there remains another important issue: the explosion of finite-sample variance. This section provides a theoretical explanation for this numerical finding. In particular, we formally show that the finite-sample variance of the LPD estimator is unbounded when compactly supported kernels are used, whereas it remains bounded when kernels with unbounded support are employed.
At first glance, the small sample property may seem less important, as the sample size is often large in recent empirical studies. However, the finite sample variance property is especially relevant in one of the most important applications: manipulation testing in RD designs. To motivate readers, we begin by explaining the unique feature of the score manipulation problem.
In the RD analysis, score manipulation is a central concern, and most of the recent RD studies report results from the LPD-based density discontinuity test to assess the presence of manipulation. In many empirical applications, manipulation is expected to be one-sided---i.e., from below the cutoff to above it (see Gerard_etal:2020). For a typical educational example, if a generous professor awards a few extra points to students whose test scores are 58 or 59, where 60 is the minimum passing threshold, then one-sided score manipulation arises.
An important implication of such one-sided manipulation is that the sample size below the cutoff necessarily decreases, as illustrated in Figure (ref). This reduction in effective sample size makes the finite-sample variance behavior observed in Section (ref) particularly relevant: The one-sided manipulation always inflates the variance of the LPD estimator on one side.
When the extent of manipulation is large, variance reduction becomes particularly important. A greater manipulation further reduces the sample size below the cutoff. Consequently, test performance does not necessarily improve with a larger effect size (or a larger jump); rather, estimation variance can become so severe that the test effectively breaks down.
This contrasts sharply with standard inference settings, in which a larger effect size typically increases detection probability, making variance reduction only of secondary importance. This is not necessarily the case in the context of manipulation testing, making it crucial to control the finite-sample variance.
We begin by providing a theoretical explanation for the variance inflation under the popular kernels, including the uniform and triangular, as observed in Figure (ref). We make the following assumptions.
Assumptions (ref) and (ref) are standard regularity conditions. Note that Assumption (ref) holds with probability one when $f$ is continuous. Under these assumptions, we have the next theorem:
Theorem (ref) states that we cannot have a bounded variance of the LPD estimator. This provides a theoretical explanation for the variance inflation under the compactly supported triangular kernel, observed in Figure (ref).
A heuristic explanation of why the variance inflates is as follows. Assume that the kernel function is compactly supported and fix $h$ and $p$. In finite samples, the local sample size within the bandwidth $h$ is not infinitely many, rather it can be only a few compared to the polynomial degree $p$.\footnote{In practice, the popular bandwidth selectors in recent empirical economics are often of the plug-in type. When such a selector is used, the bandwidth takes the form of $h=C\times n^{-1/5}$, where $n$ is the whole sample size. Consequently, the bandwidth can be small in terms of the local sample size, especially when the whole sample size $n$ is large---as is typically the case in recent empirical applications.} In such a case, the estimates can be unstable in general. Here, importantly and in contrast to the standard KDE, local polynomial fitting allows the slope term of $\hat{\bm{\beta}}$ (i.e., $\hat{f}(x_R)$) to take arbitrarily large values, especially when observations are located very closely. As an extreme case, consider $p=1$ and only two data points within the bandwidth. The LPD estimate then corresponds to the slope of the line connecting these two points (i.e., $(\hat{F}(X_{2}) -\hat{F}(X_{1}))/((X_{2} - X_{1})$ when $X_2>X_1$), which diverges as the points become closer. As a result, $\hat{f}(x_R)$ can become excessively large with a certain probability, and the finite-sample variance explodes.
Mathematically (but still heuristically), this intuition can be formulated as follows. Using the law of total variance, we have that
where $n_0$ is the local sample size within the bandwidth, that is, the number of observations in $[x_R - h,x_R)$. Here, the event $\{n_0 \leq p+2\}$ can be viewed as a mathematical formulation of the intuition that “the effective sample size is small compared to the model flexibility." Recall that the LPD estimator is a coefficient of a polynomial regression model on $[x_R - h,x_R)$. Then, we can deduce from the classic linear regression theory (e.g., Kinal:1980) that the conditional variance on the right-hand side diverges, thereby leading to $\mathbb{V}[\hat{f}(x_R)] = \infty$.
The preceding heuristics do not rely on any LPD-specific features, but rather on the general nature of local polynomial fitting. In this sense, this erratic behavior is not unique to the LPD estimator but inherent to local polynomial fitting in general. Indeed, Seifert_Gasser:1996 shows that the local linear regression estimator with a compactly supported kernel has infinite finite-sample variance under homoskedasticity. Hence, our result shows that the LPD estimator also inherits an undesirable feature associated with local polynomial smoothing techniques. (See the Online Appendix for further details.)
A simple remedy for this result is available: use a non-compactly supported kernel, as suggested by Figure (ref).\footnote{Another possible solution is to manually enlarge the bandwidth. However, doing so may introduce room for specification search and somewhat undermine one of the appealing features of the LPD approach, or its fully automatic implementation. Moreover, when conducting inference, a manually increased bandwidth can be too large, potentially rendering the procedure invalid (see related discussion in Calonico_etal:2014).} Intuitively, the infinite variance problem does not arise if we use a non-compactly supported kernel since the “local" sample size $n_0$ is equal to the whole sample size $n$. Put differently, when a kernel function is non-compactly supported, the LPD estimator can be seen as the usual weighted least squares that uses the whole sample, and we can expect it to exhibit much stable behavior. Formally, we have the following result.
This result also mirrors the finding of Seifert_Gasser:1996.
Theorems (ref) and (ref) suggest that the use of unbounded-support kernels, such as the Laplace or Gaussian kernel, is advisable. Note that, in contrast to local polynomial regression, these kernels remain a reasonable choice even asymptotically, as shown in Section (ref). This variance-stabilizing property of unbounded kernels is empirically relevant, particularly in discontinuity detection settings, as demonstrated in Section (ref).
Our message for empirical researchers is simple: kernel choice is an important design decision for estimation and inference based on the popular LPD estimator, even though this conclusion contrasts with the conventional view in the kernel-smoothing literature that kernel selection is largely irrelevant to numerical performance.
Two sets of results motivate practical guidance. First, our finite-sample theory indicates that compactly supported kernels can be vulnerable to variance inflation, whereas unbounded-support kernels such as the Gaussian and Laplace are naturally guarded against this problem. Second, our large-sample analysis shows that, at boundary points, the Laplace kernel is particularly attractive in terms of MSE, CI length, and statistical power.
That said, the choice of kernel ultimately remains a matter of researcher preference and empirical priorities. Unlike bandwidth selection, for which the optimal bandwidth is defined and a well-developed automatic rule exists Cattaneo_etal:2020, the infeasible optimal weighting scheme is not achieved by any kernel function. Accordingly, we recommend selecting the kernel to align with the main empirical goal, using the following rules of thumb.
Boundary inference. When accurate boundary estimation is the primary objective---and especially when inference at the boundary is central---the Laplace kernel is a strong default choice. In our analysis, it delivers a favorable MSE and produces shorter intervals and higher power than commonly used compactly supported kernels.
Applications requiring good performance both in the interior and at the boundary. In some empirical settings, researchers may be interested in both boundary behavior and interior fit. In such cases, the Laplace kernel may be less appealing because, at interior points, it can be substantially less MSE-efficient than compactly supported kernels such as the uniform or triangular kernels (see Table S1 in the Online Appendix). A pragmatic compromise may be the Gaussian kernel: it tends to improve boundary behavior relative to standard compact kernels while remaining reasonably competitive in the interior. Furthermore, the Gaussian kernel can avoid the finite-sample variance explosion.
Bias-sensitive applications. Finally, if bias is the primary concern, the Laplace kernel may be less attractive in some applications, as its MSE-efficiency gains can come at the cost of increased bias. In such cases, the Gaussian kernel may be a good alternative: it tends to have milder bias while still benefiting from unbounded support and is widely used in empirical work in standard settings.
Motivated by their usefulness and intuitive simplicity, recent econometrics and statistics literature extends local polynomial methods beyond the standard LPD setting to the estimation and inference of densities and related objects cattaneo2022conditional, cattaneo2021local. These extensions have also attracted increasing attention in empirical work. We conjecture that similar considerations regarding kernel choice arise in these settings as well.
\setcounter{section}{0} \setcounter{page}{1} \setcounter{equation}{0} \setcounter{lemmax}{0} \setcounter{theoremx}{0} \setcounter{propx}{0} \setcounter{corx}{0} \setcounter{table}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{remarkx}{0}
\makeatletter \makeatother
We here explain the relation between the local polynomial regression and the LPD. Write the regression function by $r(x)$, its local linear estimator by $\hat{r}(x)$, and the density of the regressor by $f(x)$. Assume for simplicity that the error satisfies $\V{\varepsilon_i}=\sigma^2$, and $[x_L,x_R)=[0,1)$. We write some bounded constants by $C_j$ below.
Utilizing the homoskedasticity assumption and the law of total variance repeatedly, Seifert_Gasser:1996 showed that
where $\mathbb{V}_{\mathrm{U},r\equiv0,f\equiv1}$ is the variance when using the uniform kernel, $r(x)\equiv0$, and $f(x)\equiv1$. Although Seifert_Gasser:1996 immediately concludes that the last term equals infinity, an important point would be in the omitted part, as it will show what happens and clarify what the connection is. Now, note that we can write
where $(Y_{(n)},\varepsilon_{(n)})$ is $(Y,\varepsilon)$ that corresponds to $X_{(n)}$. We can show that this is infinite by following similar steps as we take in our proof since ((ref)) is essentially the same as ((ref)). Now we can see that the LPD and the local linear regression share a similar structure through ((ref)) and ((ref)), which are a direct consequence of ((ref)) and ((ref)). This fact suggests that the fundamental factor for the variance property is the same: Roughly speaking, when the local sample size is small relative to the polynomial degree, the local polynomial overfits the data, and in such a case, the estimate or the slope can take an arbitrarily large value with a certain probability, since the data point can be located close enough.
In this section, we derive the infeasible optimal weight.
Throughout this section, we assume the following condition:
All the conditions are fairly standard. In particular, compactly supported $C^1$ kernel functions satisfy (ref). The Gaussian or Laplace kernels also meet the assumption.
We will see some properties of $K_{p,K}^*$ when $K\in\mathcal{K}$. We first show $\mathscr{K}_{p,K}(z) >0$ holds near $z=0$. Here, write $\bm{A}_{p,K} = \mathrm{diag}(1,-1,\ldots)\bm{D}_{p,K}\mathrm{diag}(1,-1,\ldots)$, where $\bm{D}_{p,K}\coloneqq\int_{0}^{\infty} \bm{r}_p(u)\bm{r}_p(u)^\prime K(u)\, du$. Then, Pinkus:2009 shows that $\bm{D}_{p,K}$ is strictly totally positive, and hence $\bm{A}_{p,K}^{-1}$ is also totally positive (Pinkus:2009), implying that all elements of $\bm{A}_{p,K}^{-1}$ is positive. Hence, $\bm{e}_1^\prime \bm{A}_{p,K}^{-1} \bm{r}_p$ is positively valued over $z\in[-\varepsilon,0]$ for sufficiently small $\varepsilon>0$. Combined with the assumption $K\geq0$ and $K(0)>0$, $\mathscr{K}_{p,K}(z)$ is positively valued just below the origin. Note that this implies that $K_{p, K}^*$ is also positive near the origin, since $K_{p, K}^*(u) = \int_u^0 \mathscr{K}_{p, K}(z)\,dz$ by definition.
Next, we discuss the sign change property of the equivalent kernel. The moment condition in (ref) implies $K_{p, K}^*$ changes its sign at least $p-1$ times. At the same time, $\mathscr{K}_{p, K}$ changes its sign up to $p$ times. Combined with the fact that $\int_{-\infty}^{0}\mathscr{K}_{p, K}=0$ (see (ref)), we can see that $K_{p, K}^*(u) = \int_u^0 \mathscr{K}_{p, K}(z)\,dz$ changes its sign up to $p-1$ times. To sum up, $K_{p, K}^*$ changes its sign $p-1$ times. Then, we can take zeros $u_j,j=1,\ldots,p-1$ and define $q(u) \coloneqq \prod_{j=1}^{p-1}(u - u_j)$.
With this $q$, we define $\tilde{K}(u) \coloneqq {K_{p, K}^*(u)}/(C_{p,K} \cdot q(u))$, where $C_{p,K}$ is a normalizing constant so that $\int_{-\infty}^{0}\tilde{K}=1$. Note that the sign of $K_{p, K}^*(u)$ and $q(u)$ coinsides, since $K_{p, K}^*(u)$ changes its sign exactly $p-1$ times and $K_{p, K}^*(u)>0$ near the origin. Hence, $\tilde{K}(u)\geq0$. Further, $\tilde{K}(u)$ is Lipschitz continuous. To see this, note that $u^j K(u)$ is bounded and Lipschitz for $j=0,\ldots,p$. Then, $\mathscr{K}_{p, K}$ is also bounded and Lipschitz, implying that so is $K_{p,K}^*$. Define
which is bounded and Lipschitz by noting that
and that $|k_j(u)-k_j(v)|\lesssim \int_0^1 t|u-v|\,dt = |u-v|/2$. Then, in some neighborhood $[u_j-\delta_j, u_j+\delta_j], \exists\delta_j>0$, $\tilde{K}(u) = {K_{p, K}^*(u)}/(C_{p,K} \cdot q(u)) = k_j(u)/(C_{p,K} \cdot q_{-j}(u))$, where $q(u) = (u-u_j)q_{-j}(u)$, and this is (locally) Lipschitz continuous. Based on this, the global Lipschitz continuity follows easily.
Now, rewrite $q(u) = \bm{\beta}^\top \bm{r}_{p-1}(u)$. Then, the moment condition can be rewritten as
that is $\bm{\beta}^\top =(1,0,\ldots,0)\bm{A}_{p-1, \tilde{K}}^{-1}/{C_{p,K}}$ Hence, we have that
We are now interested in the minimization of the asymptotic MSE. By the standard local polynomial regression theory, we shall consider
We have seen in the previous section that the following problem is more general than this problem:
This problem is solved by Cheng_etal:1997, which shows that the equivalent kernel of the triangular kernel, say $\mathscr{K}_{p-1, T}$, is the solution. Hence, this is the optimal weighting scheme within an enriched class of weight functions. However, it is not feasible in the LPD estimation since $\mathscr{K}_{p-1, T}(0)\neq0$.
We outline the derivation of the equivalent kernel. For notational simplicity, we write
In addition, we define
For the analysis of the boundary point $x_R$, we introduce
Note that $ \bm{S}_{(p,x_R)} = \bm{A}_{p,K}$. Finally, for the finite sample analysis below, we define
Then, the LPD estimator has the following closed form:
In the following, we assume that (i) $h\to 0$ and $nh\to\infty$ as $n\to \infty$, (ii)$K(\cdot)$ is a non-negative, symmetric kernel function such that $\int K(u)du =1$, and (iii) density function $f$ satisfies standard regularity conditions.
From Lemma 1 of the Supplemental Appendix of Cattaneo_etal:2020 and the continuous mapping theorem, it follows that
In addition, from simple algebra, we can see that
where $\hat{\mathbf{B}}_{\mathtt{LI}}$, $\hat{\mathbf{R}}$ and $\hat{\mathbf{W}}$ are the second, the third and the final terms in (ref) respectively. Lemma 2 and 4 of Supplemental Appendix of Cattaneo_etal:2020 state that $\hat{\mathbf{B}}_{\mathtt{LI}} = o_p(1)$ and $\hat{\mathbf{R}} = o_p(1)$. The convergence of $\hat{\mathbf{W}}$ is obvious from the law of large numbers. So it holds that
By (ref), (ref), and $\bm{e}_1^\prime\bm{H}^{-1} = \frac{1}{h}\bm{e}_1^\prime$, we have
That is, the equivalent kernel $K^*$ is given by
We compute the asymptotic efficiency at interior points for $p=2$. Table (ref) summarizes the result.
At interior points, the Laplace kernel performs worst in terms of MSE, while the uniform kernel delivers the best performance. This is consistent with cattaneo2021local, who show that the equivalent kernel of the LPD estimator with the uniform kernel is the Epanechnikov kernel, i.e., the MSE-optimal kernel. Among the most commonly used kernels, however, MSE is not very sensitive to kernel choice at interior points, echoing the usual observation for standard KDE.
In inference, by contrast, the Laplace kernel remains the best among the eight kernels, even at interior points. Thus, for statistical inference, the Laplace kernel is a good choice regardless of the evaluation point, although its numerical advantage is less pronounced than in the boundary case.
Figure (ref) shows how the asymptotic variance increases with the polynomial order. As in standard local polynomial regression (Fan_Gijbels:1996), the variance remains unchanged when moving from odd to even order polynomials. This observation justifies the common practice of choosing $p=2$.
Figure (ref) reports the MSE under the setup discussed in Section (ref) in the main text.
Figure (ref) is a counterpart of Figure (ref) for the Beta family kernels.