EconBase
← Back to paper

Bias correction and uniform inference for the quantile density function

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

18,757 characters · 6 sections · 14 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bias correction and uniform inference for the quantile density function

abstractFor the kernel estimator of the quantile density function (the derivative of the quantile function), I show how to perform the boundary bias correction, establish the rate of strong uniform consistency of the bias-corrected estimator, and construct the confidence bands that are asymptotically exact uniformly over the entire domain $[0,1]$. The proposed procedures rely on the pivotality of the studentized bias-corrected estimator and known anti-concentration properties of the Gaussian approximation for its supremum.

Introduction

The derivative of the quantile function, the quantile density (QD), has been long recognized as an important object in statistical inference.\footnote{This function is sometimes also called the sparsity function tukey1965part.} In particular, it arises as a factor in the asymptotically linear expansion for the quantile function bahadur1966note,kiefer1967bahadur, and hence may be used for asymptotically valid inference on quantiles csorgo1981strong,csorgo1981two,koenker_2005.

Given its importance, several estimators of the QD have been proposed in the literature. The most widely used estimator is the kernel quantile density (KQD), originally developed by siddiqui1960distribution and bloch1968simple for the case of rectangular kernel, and generalized to arbitrary kernels by falk1986estimation, welsh1988asymptotically, csorgHo1991estimating, and jones1992estimating. This estimator is simply a smoothed derivative of the empirical quantile function, where smoothing is performed via convolution with a kernel function.

Similarly to the classical case of kernel density estimation, the KQD suffers from bias close to the boundary points $\{0,1\}$ of its domain $[0,1]$, rendering the estimator inconsistent. To the best of my knowledge, no bias correction procedures have been developed for the QD.

In this paper, I show how to perform correction for the boundary bias, recovering strong uniform consistency for the resulting bias-corrected KQD (BC-KQD) estimator. The bias correction is computationally cheap and is based on the fact that the bias of the KQD is approximately equal to the integral of the localized kernel function, a quantity that only depends on the chosen kernel and bandwidth. I also develop an algorithm for construction of the uniform confidence bands around the QD on its entire domain $[0,1]$. This procedure relies on the fact that the studentized BC-KQD exhibits an influence function that is pivotal. This makes it possible to calculate the critical values by simulating from either the known influence function or the studentized BC-KQD under an alternative (pseudo) distribution of the data.

The rest of the paper is organized as follows. (ref) outlines the framework and defines the KQD estimator. (ref) introduces the BC-KQD estimator and establishes its Bahadur-Kiefer expansion. (ref) develops the uniform confidence bands based on the BC-KQD. (ref) illustrates the performance of the confidence bands in a set of Monte Carlo simulations. (ref) concludes. Proofs of theoretical results are given in the Appendix.

Setup and kernel quantile density estimator

The data consist of independent identically distributed draws $X_1,\dots,X_n$ from a distribution on $\mathbb{R}$ with a cumulative distribution function (CDF) $F$ satisfying the following assumption.

assumption[Data generating process] The distribution $F$ has compact support $[\underline x, \bar x]$ and admits a density $f=F'$ that is continuously differentiable and bounded away from zero and infinity on $[\underline x, \bar x]$.

Assumption (ref) implies that the quantile density

align[align omitted — 79 chars of source]

is continuously differentiable and bounded away from zero and infinity on the support $[\underline x, \bar x]$.

Let $X_{(1)}\le \cdots \le X_{(n)}$ be the order statistics of the sample $X_1,\dots,X_n$, and let $\hat Q$ denote the empirical quantile function,

align[align omitted — 158 chars of source]

The KQD estimator is defined as

align[align omitted — 176 chars of source]

where $K$ is a kernel function, $K_h(z) \coloneqq h^{-1}K\left(h^{-1}z\right)$, and $h>0$ is bandwidth csorgHo1991estimating. We impose the following assumptions on the kernel and bandwidth.

assumption[Kernel function] The kernel $K$ is a nonnegative function of bounded variation that is supported on $[-1/2,1/2]$, symmetric around $0$, and satisfies \begin{align} \int_\mathbb{R} K(x) \, dx = 1, \quad \int_\mathbb{R} K^2(x) \, dx < \infty. \end{align}
assumption[Bandwidth, estimation] The bandwidth $h=h_n$ is such that $h_n \to 0$ and \begin{enumerate} • $h_n^{-1} = o\left(n^{1/2} (\log n)^{-1} (\log\log n)^{-1/2}\log h^{-1} \right)$, • $h_n=o\left(n^{-1/3}(\log h^{-1})^{-1/3}\right)$. \end{enumerate}
assumption[Bandwidth, inference] The bandwidth $h=h_n$ is such that $h_n \to 0$ and \begin{enumerate} • $h_n^{-1} = o\left(n^{1/2} (\log n)^{-2} (\log\log n)^{-1/2} \right)$, • $h_n=o\left(n^{-1/3}(\log n)^{-1}\right)$. \end{enumerate}

Assumption (ref) is standard; boundedness of the total variation of $K$ ensures that the class

align[align omitted — 110 chars of source]

is a bounded VC class of measurable functions, see, e.g., nolan1987u.

Assumptions (ref) and (ref) are essentially the same, up to the log terms in the bandwidth rates, with Assumption (ref) being slightly weaker. Assumption (ref).(ref) states that the bandwidth rate is large enough (slightly larger than $n^{-1/2}$) to guarantee that the smoothed remainder of the classical Bahadur-Kiefer expansion vanishes asymptotically, see the proof of (ref) below. Assumption (ref).(ref) imposes the undersmoothing bandwidth rate (slightly smaller than $n^{-1/3}$), which ensures that the smoothing bias disappears fast enough for the confidence bands to be valid, see the proof of (ref) below.

Bias correction and Bahadur-Kiefer expansion

In this section, I introduce the bias-corrected estimator and develop its asymptotically linear expansion with an explicit a.s. uniform rate of the remainder (the Bahadur-Kiefer expansion).

To see the necessity of bias correction, note that, for $u$ close to the boundary, the kernel weights $K_h(u-i/n)$, $i=1,\dots,n-1$, do not approximately sum up to one, rendering the KQD $\hat q_h(u)$ inconsistent. Therefore, dividing the KQD by the sum of the kernel weights (or the corresponding integral of the kernel function) may eliminate the boundary bias. To this end, define

align[align omitted — 133 chars of source]

For computational purposes, note that $\psi_h$ is symmetric around $1/2$ (i.e. $\psi_h(u) = \psi_h(1-u)$ for all $u \in [0,1]$), $\psi_h\in[1/2,1]$ and $\psi_h(u) = 1$ for $u \in [h/2,1-h/2]$. The bias-corrected KQD (BC-KQD) is then defined as

align[align omitted — 206 chars of source]

The following theorem establishes that the studentized BC-KQD is approximately equal to the centered kernel density estimator with an approximation error that converges to zero a.s. at an explicit uniform rate. Since this result resembles (and relies on) the classical asymptotically linear expansion for the quantile function bahadur1966note,kiefer1967bahadur, we call it the Bahadur-Kiefer expansion for the BC-KQD. Denote $U_i=F(X_i)$, $i=1,\dots,n.$

thm[Bahadur-Kiefer expansion for the BC-KQD] Suppose Assumptions (ref) and (ref) are satisfied and $h_n \to 0$. Then the following representation holds uniformly in $u \in [0,1]$, \begin{align} Z_n^{bc}(u) = -\mathbb{G}_n(u) + O_{a.s.}\left( n^{1/2} h^{3/2} + h \log h^{-1} + h^{-1/2}n^{-1/4}(\log n)^{1/2}(\log \log n)^{1/4} \right), \end{align} where \begin{align} Z_n^{bc}(u) &\coloneqq \frac{\sqrt{nh}\left( \hat q_h^{bc}(u)-q(u) \right)}{q(u) / \psi_h(u)},\\ \mathbb{G}_n(u) &\coloneqq \frac{1}{\sqrt{nh}}\sum_{i=1}^n\left[K\left( \frac{U_i-u}{h} \right)-\mathbf{E} K\left( \frac{U_i-u}{h} \right) \right] \\ &= \sqrt{nh} \cdot \frac{1}{n}\sum_{i=1}^n \left[K_h(U_i-u) - \psi_h(u)\right]. \end{align}

This representation allows us to establish the exact rate of strong uniform consistency of the BC-KQD under a bandwidth that achieves undersmoothing (Assumption (ref).(ref)).

cor[Strong uniform consistency of BC-KQD] Suppose Assumptions (ref), (ref), and (ref) hold. Then \begin{align} \lim_{n \to \infty} \sqrt{\frac{nh_n}{2\log h_n^{-1}}} \sup_{u \in [0,1]} \left| \hat q_h^{bc}(u) - q(u) \right| = \left(\int_\mathbb{R} K^2(x) \, dx \right)^{1/2} a.s. \end{align}

One of the convenient features of the KQD (and BC-KQD) estimator is that its bandwidth has a natural scale $[0,1]$ which is independent of the data generating process. Hence, I put aside the choice of constant $c$ in the bandwidth $h=c n^{-\eta}$ and suggest setting $c=1$.

Regarding the choice of the rate $\eta$, ignoring the log terms, it is easy to establish the rate-optimal bandwidth, which is achieved whenever the rate of the smoothing bias $n^{1/2}h^{3/2}$ matches that of the remainder in the original Bahadur-Kiefer expansion $n^{-1/4}h^{-1/2}$. It follows that the nearly-optimal bandwidth is

align[align omitted — 52 chars of source]

Under this bandwidth, the exact rate of strong uniform convergence is

align[align omitted — 55 chars of source]

which is just slightly worse than the familiar “cube-root” rate kim1990cube.

Uniform confidence bands

Suppose we had access to valid approximations $c_{n,\tau}$, $c_{n,\tau}^{abs}$ to the $\tau$-quantiles of the random variables

align[align omitted — 124 chars of source]

respectively, in the sense that

align[align omitted — 165 chars of source]

Then the following confidence bands for $q(\cdot)$ would be asymptotically valid at the confidence level $1-\alpha$:

enumerate• the one-sided CB \begin{align} \left[ \frac{\hat q_h^{bc}(u)}{1 + \frac{c_{n,1-\alpha}}{\psi_h(u)\sqrt{nh}}}, \, +\infty\right), \quad u \in [0,1], \end{align} • the one-sided CB \begin{align} \left(-\infty, \, \frac{\hat q_h^{bc}(u)}{1 - \frac{c_{n,1-\alpha}}{\psi_h(u)\sqrt{nh}}} \right], \quad u \in [0,1], \end{align} • the two-sided CB \begin{align} q(u) \in \left[\frac{\hat q_h^{bc}(u)}{1 + \frac{c_{n,1-\alpha/2}^{abs}}{\psi_h(u)\sqrt{nh}}}, \,\, \frac{\hat q_h^{bc}(u)}{1 - \frac{c_{n,1-\alpha/2}^{abs}}{\psi_h(u)\sqrt{nh}}} \right], \quad u \in [0,1]. \end{align}

I propose two ways of obtaining such approximate critical values, both making use of the pivotality of the studentized bias-corrected KQD $Z_n^{bc}(u)$, see (ref). I focus on the one-sided critical value $c_{n,\tau}$ for simplicity; the proofs for the two-sided critical value are analogous.

The first approach is to let $c_{n,\tau}$ be the $\tau$-quantile of the random variable

align[align omitted — 69 chars of source]

Since $\mathbb{G}_n$ is a known process, $c_{n,\tau}$ can be obtained easily by simulation. In principle, $c_{n,\tau}$ can be tabulated for different choices of the kernel $K$ and values of the sample size $n$ and the bandwidth $h$.

The other approach is to let $c_{n,\tau}$ be the $\tau$-quantile of the random variable

align[align omitted — 76 chars of source]

where $Z_n^{bc,U[0,1]}(u)$ is equal to $Z_n^{bc}(u)$ evaluated at a pseudo-sample $\tilde X_1,\dots,\tilde X_n \sim U[0,1]$ in place of the original sample. For the uniform distribution, $q \equiv 1$, and hence

align[align omitted — 141 chars of source]

where $\tilde q_n(u)$ is the (non-bias-corrected) KQD calculated using the pseudo-sample, i.e.

align[align omitted — 145 chars of source]

The following theorem establishes that the two aforementioned approximations to the critical values are valid, implying the asymptotic validity of the confidence bands. These confidence bands are centered at an AMSE-suboptimal estimator $\hat q_h^{bc}$ and are expected to shrink at a rate slightly slower than the minimax optimal rate, as noted by chernozhukov2014anti. This is compensated for by the confidence bands exhibiting the coverage that is asymptotically exact.

thm[Exactness of confidence bands] Suppose Assumptions (ref), (ref), and (ref) hold. Then \begin{align} &\lim_{n \to \infty} \sup_{t\in\mathbb{R}}\left| \mathbb{P}\left(W_n^{bc} \le t \right) - \mathbb{P}\left(W_n^{\mathbb{G}} \le t \right) \right| = 0,\\ &\lim_{n \to \infty} \sup_{t\in\mathbb{R}}\left| \mathbb{P}\left(W_n^{bc} \le t \right) - \mathbb{P}\left(W_n^{U[0,1]} \le t \right) \right| = 0, \end{align} and hence the confidence bands (ref), (ref), and (ref) are asymptotically exact.

Monte Carlo study

In this section I study the finite-sample behavior of the proposed confidence bands in a set of Monte Carlo simulations.

I consider the following distributions of the data, all supported on the interval $[0,1]$: (i) uniform[0,1] distribution (ii) the distribution $N(1/2,1)$ truncated to $[0,1]$ (iii) the linear distribution with the PDF $f(x) = x+1/2$, $x\in[0,1]$. I set the nominal confidence level to be $1-\alpha \in \{0.8,0.9,0.95,0.99\}$ and the sample size $n \in \{100,500,1000,5000\}$. The critical values are obtained by simulating $\mathbb{G}_n(u)$ and calculating the quantiles of its supremum on the grid $u \in \{0.005,0.015,0.02,\dots,0.995\}$, with the number of simulations set to $20000$ (simulation results for the critical values based on $Z_n^{bc,U[0,1]}(u)$ are very similar, so I do not report them here). I use the kernel corresponding to the standard normal distribution truncated to $[-1/2,1/2]$ and the nearly-optimal bandwidth $h=cn^{-3/8}$, where I set $c=1$ since the scale of the bandwidth is $[0,1]$, see (ref).

In (ref), included for illustration, I plot 100 independent realizations of the $90\%$ confidence bands for the linear distribution, along with the true quantile density (in blue). (ref) contains simulated coverage values for the two-sided confidence bands. The coverage is almost invariant to the distribution of the data, but the size distortion tends to be smaller for higher nominal confidence levels.

figure[figure omitted — 319 chars of source]
table[table omitted — 1,064 chars of source]

Conclusion

To the best of my knowledge, no boundary bias correction or uniform inference procedures have been developed for the quantile density (sparsity) function. In this paper, I develop such procedures, establish their validity and show in a set of Monte Carlo simulations that they perform reasonably well in finite samples. I hope that, even when the quantile density itself is not the main inference target, these results may be employed for improving the quality of inference for other statistical objects, including the quantile function.

\part*{Appendix}