EconBase
← Back to paper

Smoothing quantile regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,150 characters · 14 sections · 11 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\citationstyle{dcu} \lineskip=.35ex \baselineskip 3.5ex \mathindent 0em {0.02in}

center[center omitted — 485 chars of source]

\vskip 2em \noindentAbstract We propose to smooth the objective function, rather than only the indicator on the check function, in a linear quantile regression context. Not only does the resulting smoothed quantile regression estimator yield a lower mean squared error and a more accurate Bahadur-Kiefer representation than the standard estimator, but it is also asymptotically differentiable. We exploit the latter to propose a quantile density estimator that does not suffer from the curse of dimensionality. This means estimating the conditional density function without worrying about the dimension of the covariate vector. It also allows for two-stage efficient quantile regression estimation. Our asymptotic theory holds uniformly with respect to the bandwidth and quantile level. Finally, we propose a rule of thumb for choosing the smoothing bandwidth that should approximate well the optimal bandwidth. Simulations confirm that our smoothed quantile regression estimator indeed performs very well in finite samples.

small\noindentJEL classification numbers C14, C21\\ Keywords asymptotic expansion, Bahadur-Kiefer representation, conditional quantile, convolution-based smoothing, data-driven bandwidth.

\lineskip=.25ex \baselineskip 2.5ex

footnotesize\noindentAcknowledgments We are grateful to Antonio Galv\ {a}o and Aureo de Paula for valuable comments, as well as to the Associate Editor and anonymous referees. Fernandes also thanks the financial support from CNPq (302278/2018-4).

\pagestyle{plain} \lineskip=.5ex \baselineskip 5ex

Introduction

Quantile regression (QR) enjoys some very appealing features. Apart from enabling some very flexible patterns of partial effects, quantile regressions are also interesting because they satisfy some equivariance and robustness principles. See \citeasnoun{koenker1978regression} and \citeasnoun{koenker2005quantile} for theoretical aspects; and \citeasnoun{koenker2000galton}, \citeasnoun{buchinsky1998recent}, \citeasnoun{koenker2001quantile}, \citeasnoun{koenker2005quantile} and references therein for applications.

There is a price to pay, though. The objective function that the standard QR estimator aims to minimize is not smooth. As established in \citeasnoun{bassett1982empirical}, it follows that the paths of this estimator have jumps, even if the underlying quantile function is very regular. Because the objective function does not have second derivatives, statistical inference is not straightforward and involves ancillary estimation of nuisance parameters (namely, the asymptotic covariance matrix depends on the population conditional density evaluated at the true quantile). See the discussions in \citeasnoun{koenker1994confidence}, \citeasnoun{buchinsky1995estimating}, \citeasnoun{koenker2005quantile}, \citeasnoun{goh2009nonstandard}, and \citeasnoun{fan2016direct}, among others. Unsurprisingly, the literature now boasts a wide array of techniques to tackle this issue, including bootstrapping horowitz1998bootstrap,machado2005bootstrap, MCMC methods chernozhukov2003mcmc, empirical likelihood otsu2008conditional,whang2006smoothed, strong approximation methods portnoy2012nearly, nonstandard inference goh2009nonstandard, and other nonparametric approaches mammen2017expansion,fan2016direct.

Further, the asymptotic normality of the standard QR estimator relies on Bahadur-Kiefer representations with poor rates of convergence. The latter is at best of order $n^{-1/4}$ for iid (homoskedastic) errors \citeaffixed{koenker1987estimation,knight2001comparing,jurevckova2012methodology}{see}. This means that the first-order linear Gaussian approximation for the distribution of the QR estimators could well fail in finite samples. \citeasnoun{portnoy2012nearly} however obtains a nearly $n^{-1/2}$ rate using a nonlinear approximation, explaining perhaps why QR inference is actually rather good in finite samples despite the poor rates of the Bahadur-Kiefer remainder.

In this paper, we employ a convolution-type smoothing of the objective function that helps produce a continuous QR estimator, which is not only less irregular and variable than the standard estimator but also more linear, in the sense that the stochastic order of the Bahadur-Kiefer remainder term is much closer to $n^{-1/2}$ for a proper bandwidth choice. As the smoothing we propose ensures the twice differentiability of the objective function with respect to the parameter vector, we can readily estimate the asymptotic covariance matrix of our smoothed QR estimator using a standard sandwich formula. In fact, the second derivative of our kernel-smoothed objective function indeed coincides with the usual kernel-based covariance matrix estimators in the standard QR setup.

Quantile estimation aims to offer a global picture of the distribution. Bearing this in mind, we contemplate an asymptotic theory that holds uniformly with respect to quantile levels, or in a functional sense. This is in line with the functional central limit theorem that \citeasnoun{koenker2002inference} establish for specification-testing purposes. Apart from uniformity in quantile level, we also focus on bandwidth uniformity in consistency and asymptotic distribution results in order to provide some robustness against bandwidth snooping. Practical applications of smoothing procedures may indeed involve several bandwidth choices, exposing practitioners to the risk of biasing inference by selecting a bandwidth that agrees with prior beliefs. As a result, our asymptotic theory considers uniformity in both quantile level and bandwidth, allowing for data-driven choices for the latter. This is in sharp contrast with related works on smoothing, such as \citeasnoun{horowitz1998bootstrap} and \citeasnoun{kaplan2017smoothed}. In particular, they restrict attention to a given quantile level and to deterministic bandwidths when studying higher-order accuracy for confidence interval coverage and type-I test errors, respectively.

Our main contributions are as follows. First, smoothing the QR objective function induces bias in the QR estimation. We derive the order of such bias, showing that it is negligible with respect to $\sqrt{n}$ for the optimal bandwidth rate, as well as for the data-driven bandwidth we use in the Monte Carlo experiments. Second, we establish that the smoothed QR estimator has a smaller Bahadur-Kiefer linearization error than the standard QR estimator as long as the bandwidth remains large enough. Third, we find that, for a proper bandwidth choice, our smoothed QR estimator yields a lower asymptotic mean squared error (AMSE) than the standard estimator. Moreover, we derive the optimal bandwidth for our smoothed QR estimator. Fourth, we provide a covariance matrix estimator that captures the variance improvement term. Fifth, we provide a functional central limit theorem analogous to \possessivecite{koenker1999goodness} that holds uniformly with respect to the bandwidth. Sixth, we show that our smoothed QR slope estimator is asymptotically differentiable. We exploit this feature to come up with conditional quantile density function (qdf) and probability density function (pdf) estimators that do not suffer from the curse of dimensionality under correct QR specification. In particular, this allows us to propose an efficient QR estimator using sample splitting weights based on the qdf estimator.

Our simulations examine confidence intervals using a random bandwidth. Although it is as usual hard to estimate it due to the derivative of the conditional pdf that appears in the bias term, we show how to circumvent that by resorting to \possessivecite{silverman1986density} rule-of-thumb bandwidth for the standard QR residuals. The simulation experiment also illustrates the estimation of the qdf and pdf in presence of a reasonably large number of explanatory variables.

We are obviously not the first to sail in the smoothing direction. In particular, \citeasnoun{nadaraya1964some} and \citeasnoun{parzen1979nonparametric} estimate unconditional quantiles respectively by inverting a smoothed estimator of the cumulative distribution function (cdf) and by smoothing the sample quantile function. \citeasnoun{xiang1994bahadur} and \citeasnoun{ralescu1997bahadur} provide Bahadur-Kiefer representations for smoothed quantile estimators. \citeasnoun{azzalini1981note} shows that the former smoothed estimator dominates sample quantiles at the second order \citeaffixed{cheung2010bootstrap}{see also}, with \citeasnoun{sheather1990kernel} establishing a similar result for the latter. For discussions on the relative deficiency of sample quantiles with respect to kernel-based quantile estimators, see also \citeasnoun{falk1984relative} and \citeasnoun{kozek2005combine}. Although the above quantile estimators are all continuous, only a few of them consider covariates in a quantile regression fashion. \citeasnoun{volgushev2019distributed} allow for covariates, ensuring quantile-level differentiability using spline smoothing.

The asymptotic covariance matrix of quantile estimators involves an unknown pdf koenker2005quantile, making inference a bit harder.\footnote{ For instance, \citeasnoun{hall1988distribution} study higher-order accuracy of confidence intervals. See also \citeasnoun{portnoy2012nearly} and references therein.} Smoothing the objective function helps in this dimension for it allows estimating the asymptotic covariance matrix using a standard sandwich formula kaplan2017smoothed. In turn, we focus more on the implications of smoothing the objective function, notably in terms of smoothness of the QR estimator, and on quantile level and bandwidth uniformity.

There is notwithstanding an intimate connection between the above papers on inference and ours. Although the smoothed objective function of \citeasnoun{horowitz1998bootstrap} differs from ours, his theoretical arguments easily extends to our framework. In addition, the first-order conditions that \citeasnoun{whang2006smoothed} employs to derive his empirical likelihood QR estimator is identical to the ones our smoothed QR estimator solves. Interestingly, the same applies to \possessivecite{kaplan2017smoothed} smoothed estimating equations (SEE) if restricting attention to QR without instrumental variables. As us, they argue that their smoothing approach can improve on \possessivecite{horowitz1998bootstrap} smoothed QR estimator. See also \citeasnoun{kaido2018decentralization} for numerical algorithms that connect standard SEE and QR estimations.

Finally, our paper also relates with the literature on qdf estimation. \citeasnoun{parzen1979nonparametric} is the first to notice the important role that the quantile density plays. See \citeasnoun{guerre2012uniform} for a recent application in the econometrics of auctions, as well as \citeasnoun{roth2018estimating} and references therein for related estimation methods.\footnote{ In particular, \citeasnoun{roth2018estimating} develop a qdf estimator based on the estimation of the second derivative of the population QR objective function.} We propose a novel efficient QR estimator based on a first-step qdf estimation, whose second-order behavior avoids the curse of dimensionality, as opposed to the estimation procedures of \citeasnoun{Newey1990efficient}, \citeasnoun{Zhao2001asymptotically}, \citeasnoun{otsu2008conditional}, and \citeasnoun{komunjer2010efficient}.

The remainder of this paper proceeds as follows. Section 2 introduces our smoothed QR estimator. Section 3 describes the main assumptions, the estimation procedure, the asymptotic covariance matrix estimator and main results. Section 4 assesses by means of a simulation study the performance of our kernel-based QR estimator relative to \possessivecite{koenker1978regression} and \possessivecite{horowitz1998bootstrap} estimators. Section 5 offers some concluding remarks. Appendix A collects the proofs of our main results, whereas we relegate the proofs of the intermediary results to an online appendix. The latter also includes an in-depth comparison with \possessivecite{horowitz1998bootstrap} smoothing approach.

Smoothed QR estimation and extensions

Let $(Y_i,X_i)$, with $i=1,\ldots,n$, denote an iid sample from $(Y,X)\in\mathbb{R}\times\mathbb{R}^d$, where the conditional quantile of the response $Y$ given the covariate $X=x$ is such that

equation[equation omitted — 100 chars of source]

with $Q(\tau\,|\, x):=\inf\big\{q:\,F(q\,|\, x)\ge\tau\big\}$ and $F(\cdot\,|\, x)$ denoting the conditional cdf of $Y$ given $X=x$, with density $f(\cdot\,|\, x)$. We henceforth assume the correct specification of the QR model in (ref).

\citeasnoun{koenker1978regression} define the (population) objective function of the quantile regression as

equation[equation omitted — 134 chars of source]

where $e(b):=Y-X'b$, $F\left(t;\,b\right)=\mathrm{Pr}\big[e(b)\le t\big]$, and $\rho_\tau(u):=u\big[\tau-\mathbb{I}(u<0)\big]$ is the usual check function with $\mathbb{I}(A)$ denoting the indicator function that takes value 1 if $A$ is true, zero otherwise. As the true parameter $\beta(\tau)$ minimizes (ref), \possessivecite{koenker1978regression} standard QR estimator $\widehat{\beta}(\tau)$ minimizes the sample analog based on the empirical distribution, namely,

equation[equation omitted — 160 chars of source]

where $e_i(b):=Y_i-X_i'b$, and $\widehat{F}(\cdot;b)$ denotes the empirical distribution function of $e_i(b)$.

The right-hand side of equation (ref) suggests that one may obtain a different QR estimator by varying the integrating measure of the check function integral, that is to say, by changing the estimator of the cdf. Instead of employing the empirical distribution as in the standard QR estimator, we shall consider a kernel-type cdf estimator as the integrating measure in a similar fashion to what \citeasnoun{nadaraya1964some} proposes for smoothing the unconditional quantile estimator.

Consider a bandwidth $h>0$ that shrinks to zero as the sample size grows and a smooth kernel function $k$ such that $\int k(v)\,\mathrm{d}v=1$. Letting $k_h(v)=\frac{1}{h}\,k(v/h)$, the kernel density and distribution estimators are given by $\widehat{f}_h(v;b):=\frac{1}{n}\sum_{i=1}^nk_h\big(v-e_i(b)\big)$ and $\widehat{F}_h(t;b):=\int_{-\infty}^t\widehat{f}_h(v;b)\,\mathrm{d}v$, respectively. We then apply kernel smoothing to the empirical objective function in (ref), yielding

equation[equation omitted — 167 chars of source]

Accordingly, the resulting smoothed QR estimator is

equation[equation omitted — 119 chars of source]

By smoothing the integrating measure of the objective function, we ensure that the mapping $b\mapsto\widehat{R}_h(b;\tau)$ is twice continuously differentiable, in opposition to the nonsmoothness of the standard objective function $b\mapsto\widehat{R}(b;\tau)$. The differentiability of the objective function is very convenient for at least two reasons. First, the smoothness of the objective function entails the regularity of the resulting QR estimator. Second, differentiability allows us to estimate the asymptotic covariance matrix of our quantile slope coefficient estimates in a canonical manner \citeaffixed{newey1994large}{see, for instance,}.

Finally, observe that

displaymath\widehat{R}_h(b;\tau)=(1-\tau)\int_{-\infty}^0\widehat{F}_h(v;b)\,\mathrm{d}v+\tau\int_0^\infty\big(1-\widehat{F}_h(v;b)\big)\,\mathrm{d}v,

so that the first- and second-order derivatives of $\widehat{R}_h(b;\tau)$ with respect to $b$ are respectively

displaymath\widehat{R}_h^{(1)}(b;\tau)=\frac{1}{n}\sum_{i=1}^n X_i\left[K\left(-\frac{e_i(b)}{h}\right)-\tau\right] and \widehat{R}_h^{(2)}(b;\tau)=\frac{1}{n}\sum_{i=1}^n X_iX_i'k_h\big(-e_i(b)\big),

where $K(t):=\int_{-\infty}^tk(v)\,\mathrm{d}v$.

Asymptotic inference

The availability of the second-order derivative $\widehat{R}_h^{(2)}\big(\widehat{\beta}_h(\tau);\tau\big)$ allows us to propose asymptotic confidence intervals for $\beta(\tau)$ in a standard fashion. We show in Section 3.4 that $\sqrt{n}\big(\widehat{\beta}_h(\tau)-\beta(\tau)\big)$ weakly converges to a Gaussian distribution with mean zero and covariance matrix $\Sigma(\tau):=D^{-1}(\tau)V(\tau)D^{-1}(\tau)$, where $V(\tau):=\tau(1-\tau)\mathbb{E}(XX')$ and

displaymathD(\tau):=R^{(2)}\big(\beta(\tau);\tau\big)=\mathbb{E}\!\left[XX'f\left(X'\beta(\tau)\,|\, X\right)\right]

is the Hessian of the objective function evaluated at the true parameter. In addition, we show that

displaymath\widehat{\Sigma}_h(\tau):=\widehat{D}_h^{-1}(\tau)\widehat{V}_h(\tau)\widehat{D}_h^{-1}(\tau),

with $\widehat{D}_h(\tau):=\widehat{R}_h^{(2)}\big(\widehat{\beta}_h(\tau);\tau\big)$ and

displaymath\widehat{V}_h(\tau):=\frac{1}{n}\sum_{i=1}^n X_iX_i'\left[K\left(-\frac{e_i\big(\widehat{\beta}_h(\tau)\big)}{h}\right)-\tau\right]^2,

is a consistent estimator of $\Sigma(\tau)$.\footnote{ Alternatively, one could simply employ the sample analog of $\tau(1-\tau)\mathbb{E}[XX']$ to estimate $V(\tau)$, which would lead to \possessivecite{powell1991} estimator of the asymptotic covariance of the standard QR estimator. See also \citeasnoun{angrist2006} and \citeasnoun{kato2012}. However, we expect that our variance estimator to entail better finite-sample properties and, accordingly, shorter confidence intervals.} This means that we may compute a $(1-\alpha)-$confidence interval for the $k$th QR coefficient $\beta_k(\tau)$ as

displaymath\mathrm{CI}_{1-\alpha}\big(\beta_k(\tau)\big):=\widehat{\beta}_{k,h}(\tau)\pm\frac{z_{\alpha/2}\,\widehat{\sigma}_{k,h}(\tau)}{\sqrt{n}},

where $\widehat{\sigma}_{k,h}(\tau)$ is the square root of the $k$th diagonal entry of $\widehat{\Sigma}_h(\tau)$, $\widehat{\beta}_{k,h}(\tau)$ is the $k$th element of the smoothed QR estimator, and $z_{\alpha}$ is the $\alpha$ quantile of a standard Gaussian distribution.

Smoothness of $\widehat{\beta}_h (\cdot)$ and extensions

Given that $\widehat{\beta}_h(\tau)$ satisfies the first-order condition $\widehat{R}_h^{(1)}\big(\widehat{\beta}_h(\tau);\tau\big)=0$, it follows from the implicit function theorem that $\widehat{\beta}_h(\tau)$ is continuously differentiable with respect to $\tau$ with

equation[equation omitted — 361 chars of source]

as long as the Hessian has an inverse, which holds with a probability tending to $1$. Because the Hessian matrix is asymptotically positive, it also follows that the QR function evaluated at the average covariate, $\bar{X}'\widehat{\beta}_h(\tau)$, is strictly increasing in $\tau$.

Figure 1 compares two different paths of our smoothed QR estimator with the ones of the standard QR estimator. The quantile paths implied by the convolution-type kernel QR estimates are not only much smoother but much less variable than the conditional quantiles based on the standard QR estimates. Theorem 2 indeed shows that our smoothed QR estimator is continuous with a probability tending to one, while the standard QR estimator is a step function. Theorem 3 also establishes that smoothing reduces the variance of the QR estimator. Even though Figure 1 evaluates the conditional quantiles at $X=0.1$ and $X=0.9$, rather than at $\bar{X}$, the quantile paths based on our smoothed QR estimator look increasing such that there is no need to apply \possessivecite{CFVG10} rearrangement procedure. This is most likely due to the smoothness of the convolution-type kernel estimator, inasmuch as the nonmonotonicity of the standard QR estimator is mostly due to small jumps.

We next consider two extensions based on the quantile density function (qdf) given by

displaymathq(\tau|x)=\frac{\partial Q(\tau|x)}{\partial\tau}=\frac{1}{f\big(Q(\tau|x)|x\big)}.

As noted by \citeasnoun{GS2017}, the curve $\tau\mapsto\big(Q(\tau|x),1/q(\tau|x)\big)=:\mathtt{f}(\tau|x)$ is the graph $y\mapsto\big(y,f(y|x)\big)$ of the conditional pdf $f(\cdot|x)$. The qdf plays a major role in first-price auctions guerre2012uniform and in semiparametric efficient QR estimation \citeaffixed{Newey1990efficient,Zhao2001asymptotically}{see, among others,}. In particular, the efficient QR slope estimator $\tilde{b}_q(\tau)=\arg\min_b\widetilde{R}_q(b;\tau)$ is infeasible because

equation[equation omitted — 125 chars of source]

depends on the conditional qdf. \citeasnoun{Newey1990efficient} and \citeasnoun{Zhao2001asymptotically} propose to estimate the conditional qdf using kernel methods. However, their estimators are bound to perform poorly if the covariate dimension is large due to the curse of dimensionality. The same drawback also applies to the efficient estimators put forth by \citeasnoun{otsu2008conditional} and \citeasnoun{komunjer2010efficient}.

\paragraph{Pdf curve estimation.} The linear QR model dictates that $q(\tau|x)=x'\beta^{(1)}(\tau)$, which we may estimate using $\widehat{\beta}_h^{(1)}(\tau)=\left[\widehat{R}_h^{(2)}\big(\widehat{\beta}_h(\tau);\tau\big)\right]^{-1}\bar{X}$ in ((ref)). The resulting conditional qdf estimator, $\widehat{q}_h(\tau|x):=x'\widehat{\beta}_h^{(1)}(\tau)$, does not suffer from the curse of dimensionality because it makes use of the linear QR structure. The pdf curve estimator

equation[equation omitted — 156 chars of source]

converges to $\mathtt{f}(\tau|x)$ at the rate of univariate kernel density estimator. If $\widehat{q}_h(\cdot|x)>0$, the pdf estimator is positive and integrates to one given that $\int_0^11/\widehat{q}_h(\tau|x)\,\mathrm{d}\widehat{Q}_h(\tau|x)=\int_0^1\mathrm{d}\tau=1$. This means that the pdf curve estimator is strict. Figure (ref) depicts an example of pdf curve estimation with three covariates using a sample of 200 observations. It also exhibits the QR and qdf estimates we use for computing the pdf estimator. The pdf curve estimator seems to capture reasonably well the underlying density despite the high dimensionality of the estimation problem relative to the sample size.

\paragraph{Semiparametric efficiency.} As in \citeasnoun{Newey1990efficient} and \citeasnoun{Zhao2001asymptotically}, we resort to sample splitting to obtain efficient QR slope estimators. To this end, we minimize a feasible counterpart of the objective function $\widetilde{R}_q(b;\tau)$ in (ref), namely,

equation[equation omitted — 143 chars of source]

where $\check{q}_h(\tau|x)=x'\check{\beta}_h^{(1)}(\tau)$ considers only the first $m<n$ observations. One could alternatively use either a smoothed version of the check function $\rho_{\tau}(\cdot)$ or a version of \possessivecite{Newey1990efficient} one-step estimator using the weights $\check{q}_h(\tau|X_i)$.

Theoretical results for pdf graph and for the efficient QR in ((ref)) readily follow from the asymptotic theory we establish for the smoothed QR estimator. We state them in Section (ref).

Alternative smoothed objective function

In the absence of covariates, the first-order condition $\widehat{R}_h^{(1)}\big(\widehat{\beta}_h(\tau);\tau\big)=0$ that results in our smoothed QR estimator coincides with the first-order condition in \citeasnoun{nadaraya1964some}, namely, $\widehat{F}_h\big(\widehat{\beta}_h(\tau)\big)=\tau$. This confirms that the smoothing we apply to the objective function boils down to estimating the conditional cumulative distribution function using a kernel approach. In turn, the smoothed objective function $\widehat{\mathfrak{R}}_h(b;\tau)$ put forth by \citeasnoun{horowitz1998bootstrap} replaces the indicator in the check function by a kernel counterpart: $\widehat{\mathfrak{R}}_h(b;\tau):=\frac{1}{n}\sum_{i=1}^ne_i(b)\Big[\tau-K\big(-e_i(b)/h\big)\Big]$. As noted by \citeasnoun{kaplan2017smoothed}, the first-order derivative of this objective function is

eqnarray[eqnarray omitted — 357 chars of source]

The resulting first-order condition differs from the one we obtain because of the additional term in (ref). We show in the Online Appendix that, as a consequence, \possessivecite{horowitz1998bootstrap} smoothed QR estimator exhibits larger bias and variance than ours.

Asymptotic theory

We start with some notation. Let $\left\Vert{\cdot}\right\Vert $ denote the Euclidean norm of a matrix, namely, $\left\Vert{A}\right\Vert=\sqrt{\mathrm{tr}(AA')}$. We denote by $f(\cdot\,|\, x)$ the conditional probability density function of $Y$ given $X=x$, with $j$th partial derivative given by $f^{(j)}(y\,|\, x):=\frac{\partial^j}{\partial y^j}\,f(y\,|\, x)$. Similarly, let $q(\tau\,|\, x):=\frac{\partial}{\partial\tau}\,Q(\tau\,|\, x)$. In what follows, we first discuss the assumptions we require to work out the asymptotic theory and then derive the asymptotic mean squared error and Bahadur-Kiefer representation for our smoothed QR estimator. We wrap up this session with some inference implications.

Assumptions

In this section, we discuss the conditions under which we derive the asymptotic theory. Apart from standard technical conditions on the covariates and kernel function, we essentially require the conditional quantile and density functions to be smooth enough. In particular, we assume the following conditions.\vskip 1em

\noindentAssumption X The components of $X$ are positive, bounded random variables, i.e. the support of $X$ is a compact subset of $\bar{\mathbb{R}}_{+*}^d$. The matrix $\mathbb{E}[XX']$ is full rank.\vskip 1ex

\noindentAssumption Q The conditional quantile and density functions $Q(\tau\,|\, x)$ and $f(y\,|\, x)$ satisfy

description• The map $\tau\mapsto\beta(\tau)$ is continuously differentiable over $(0,1)$. The conditional density $f(y\,|\, x)$ is continuous and strictly positive over $\mathbb{R}\times\mathrm{supp}(X)$. • There exists an integer $s\ge 1$ such that the derivative $f^{(s)}(\cdot\,|\,\cdot)$ is uniformly continuous in the sense that $\lim_{\epsilon\rightarrow 0}\sup_{(x,y)\in\mathbb{R}^{d+1}} \sup_{t:\,\left\vert{t}\right\vert\le\epsilon}\left\vert{f^{(s)}(y+t\,|\, x)-f^{(s)}(y\,|\, x)}\right\vert=0$, as well as such that $\sup_{(x,y)\in\mathbb{R}^{d+1}}\left\vert{f^{(j)}(y\,|\, x)}\right\vert<\infty$ and $\lim_{y\rightarrow\pm\infty}f^{(j)}(y\,|\, x)=0$ for all $j=0,\ldots,s$.

\vskip 1ex

\noindentAssumption K The kernel function $k$ and bandwidth $h$ satisfy

description• The kernel $k:\mathbb{R}\rightarrow\mathbb{R}$ is even, integrable, twice differentiable with bounded first and second derivatives, and such that $\int k(z)\,\mathrm{d}z=1$ and $0<\int_0^\infty K(z)\,[1-K(z)]\,\mathrm{d}z<\infty$. In addition, for $s$ as in Assumption Q2, $\int\left\vert{z^{s+1}k(z)}\right\vert\,\mathrm{d}z<\infty$, and $k$ is orthogonal to all nonconstant monomials of degree up to $s$, i.e., $\int z^jk(z)\,\mathrm{d}z=0$ for $j=1,\dots,s$, and $\int z^{s+1} k(z)\,\mathrm{d}z \neq 0$. • $h\in[\underaccent{\bar}{h}_n,\bar{h}_n]$ with $1/\underaccent{\bar}{h}_n=o\big((n/\ln n)^{1/3}\big)$ and $\bar{h}_n=o(1)$.

\vskip 1em

Some remarks are in order. First, observe that $R^{(2)}(b;\tau)=\mathbb{E}\!\left[XX'f\!\left(X'b\,|\, X\right)\right]$ is positive definite for all $b$ and any $\tau$ under Assumptions Q1 and X. This means that $D^{-1}(\tau)$ exists for every $\tau$. Second, Assumption Q1 also ensures that $\tau\mapsto Q(\tau\,|\, x)$ is strictly increasing over $\left(0,1\right)$, with a strictly positive derivative with respect to $\tau$ given that $q(\tau\,|\, x)=1/f\big(Q(\tau\,|\, x)\,|\, x\big)$. Third, we assume that $Y$ has support on the real line merely for notational simplicity. It is straightforward to relax it with some minor adaptations. Fourth, the reason why we choose a kernel such that $\int_0^\infty K(z)\,[1-K(z)]\,\mathrm{d}z$ is positive will become clear in Theorem 3. It makes sense because it guarantees that our smoothed QR estimator dominates the standard QR estimator in the AMSE sense. Fifth, in view that Assumption K does not preclude high-order kernels, $\widehat{f}_h\left(\cdot;b\right)$ is not necessarily a density, even if it is a consistent estimator of the density of the error term $e(b)$. It is nonetheless possible to show $\widehat{R}_h(b;\tau)=\frac{1}{n}\sum_{i=1}^n\rho_\tau*k_h\big(e_i(b)\big)$, where $*$ is the convolution operation. This means that, technically speaking, it is probably more rigorous to interpret our approach as a convolution-type smoothing. Finally, it is also important to clarify that the bandwidth $h$ implicitly depends on the sample size $n$ through its lower and upper limits in Assumption K2. This is paramount because we wish to entertain data-driven bandwidths.

Bias and Bahadur-Kiefer representation

In this section, we study the order of the asymptotic bias and obtain a Bahadur-Kiefer representation for the stochastic error of $\widehat{\beta}_h(\tau)$. As for the former, it is indeed expected from popular wisdom that smoothing the QR objective function should induce bias in finite samples. More formally, $\widehat{\beta}_h(\tau)$ actually estimates $\beta_h(\tau):=\arg\min_{b\in\mathbb{R}^d}R_h(b;\tau)$, with $R_h(b;\tau):=\mathbb{E}[\widehat{R}_h (b;\tau)]$. This yields $\beta_h(\tau)-\beta(\tau)$ as the bias term and $\widehat{\beta}_h(\tau)-\beta_h(\tau)$ as the stochastic error of our smoothed QR estimator. Our first result shows that the smoothing bias shrinks to zero as the sample size grows.\vskip 1em

\noindentTheorem 1 For $\bar{h}_n$ small enough, $\beta_h(\tau)$ is unique under Assumptions X, Q and K for every $\tau\in[\underaccent{\bar}{\tau},\bar{\tau}]$ and such that $\beta_h(\tau)=\beta(\tau)-h^{s+1}B(\tau)+o(h^{s+1})$ uniformly over $(\tau,h)\in[\underaccent{\bar}{\tau},\bar{\tau}]\times[\underaccent{\bar}{h}_n,\bar{h}_n]$, with $B(\tau)=\frac{\int z^{s+1}k(z)\,\mathrm{d}z}{(s+1)!}\, D^{-1}(\tau)\,\mathbb{E}\!\left[Xf^{(s)}\big(X'\beta(\tau)\,|\, X\big)\right]$.\vskip 1em

It is interesting to discuss the implications of Theorem 1 to the particular case of a standard linear regression model $Y_i=X_i' \beta +\varepsilon_i$ with iid errors independent of the covariates. Let $X_i=(1,\tilde{X}_i)'$, with $\tilde{X}_i\in\mathbb{R}$, so that the conditional pdf of $Y_i$ given $X_i=x$ is $f(y|x)=f_\varepsilon\big(y-x'\beta(\tau)\big)$. It then follows that $D(\tau)=f_\varepsilon(0)\mathbb{E}[XX']$ for all $\tau$ and $B(\tau)\propto\mathbb{E}^{-1}[XX']\mathbb{E}[X]=(1,0)'$, so that the first term in the bias appears only for the intercept. More generally, Theorem 1 settles the issue of possible side effects of smoothing in that $\beta_h(\tau)$ eventually becomes uniformly close to the true parameter $\beta(\tau)$.

The next result derives some convenient expansions for the stochastic error $\widehat{\beta}_h(\tau)-\beta_h(\tau)$. For this purpose, let $\widehat{S}_h(\tau):=\widehat{R}_h^{(1)}\big(\beta_h(\tau);\tau\big)$ and $D_h(\tau):=R_h^{(2)}\big(\beta_h(\tau);\tau\big)$. Note that the first-order condition $R_h^{(1)}\big(\beta_h(\tau);\tau\big)=0$ implies that the score term $\widehat{S}_h(\tau)$ has zero mean, and hence the stochastic error in the Bahadur-Kiefer representation ((ref)) is asymptotically centered.\vskip 1em

\noindentTheorem 2 Under Assumptions X, Q and K, $\widehat{\beta}_h(\cdot)$ is unique and continuous over $(\tau,h)\in[\underaccent{\bar}{\tau},\bar{\tau}]\times[\underaccent{\bar}{h}_n,\bar{h}_n]$ with probability tending to one, satisfying the following two representations:

eqnarray[eqnarray omitted — 351 chars of source]

with $\varrho_n(h)=\sqrt{\ln n/(nh)}$ and both remainder terms uniform with respect to $(\tau,h)\in[\underaccent{\bar}{\tau},\bar{\tau}]\times[\underaccent{\bar}{h}_n,\bar{h}_n]$.}\vskip 1em

Unlike the standard QR objective function, $\widehat{R}_h (\cdot;\tau)$ is not necessarily convex because higher-order kernels may take negative values. We address this issue by first showing that $\widehat{\beta}_h(\tau)$ is close to $\beta_h(\tau)$, uniformly in $\tau$ and $h$, and then proving that the smoothed objective function $\widehat{R}_h (\cdot;\tau)$ is asymptotically strictly convex in the vicinity of $\beta_h(\tau)$. This follows from the convergence of $\widehat{R}^{(2)}_h(b;\tau)$ to $R^{(2)}(b;\tau)$, uniformly for $b$ in any compact set and also in $\tau$ and $h$, as established using a powerful concentration inequality from \citeasnoun{massart2007concentration}. The first-order condition $\widehat{R}_h\big(\widehat{\beta}_h(\tau);\tau\big)=0$ implies the following integral representation

displaymath\widehat{\beta}_h(\tau)-\beta_h(\tau)=-\left[\int_0^1\widehat{R}_h^{(2)}\left(\beta_h(\tau)+u\big[\widehat{\beta}_h(\tau)-\beta_h(\tau)\big];\tau\right)\,\mathrm{d}u\right]^{-1}\widehat{S}_h^{(1)}\big(\tau\big)

for the stochastic error. Deriving the order of the score uniformly in $\tau$ and $h$ via a concentration inequality entails the representations ((ref)) and ((ref)). One could also obtain higher-order expansions for the stochastic error along the same lines by establishing uniform convergence of higher-order derivatives of the smooth objective function through concentration inequalities.

The Bahadur-Kiefer representation in (ref) shows that $\widehat{\beta}_h(\tau)$ is, in a sense, more linear than the standard QR estimator $\widehat{\beta}(\tau)$. \citeasnoun{knight2001comparing} and \citeasnoun{jurevckova2012methodology} show that, in many cases of interest, the Bahadur-Kiefer representation for the standard QR estimator is given by $\sqrt{n}\left(\widehat{\beta}(\tau)-\beta(\tau)\right)=-\sqrt{n}\,D^{-1}(\tau)\widehat{S}(\tau)+O_p(n^{-1/4})$, with $\widehat{S}(\tau)=\widehat{R}^{(1)}\big(\beta(\tau);\tau\big)$. In contrast, the remainder term in (ref) is of order nearly $O_p(n^{-1/2})$ for proper bandwidth choices if centering around $\beta_h(\tau)$. The asymptotically negligible bias given by the difference between $\beta_h(\tau)$ and $\beta(\tau)$ is the price we pay for improving the rate of the Bahadur-Kiefer representation.

The nonlinear approximation in (ref) obtains a smaller remainder term of order $n^{-1/2}$ by replacing the deterministic standardization $D_h^{-1}(\tau)$ in (ref) with $\widehat{D}_h^{-1}(\tau)$. This slightly improves the rate of the remainder term in \possessivecite{portnoy2012nearly} approximation. One could also obtain higher-order approximations involving quadratic terms under stronger bandwidth rate conditions. As in \citeasnoun{horowitz1998bootstrap}, this would however exclude bandwidths of the optimal order $n^{-1/5}$ for a kernel of order $s+1=2$, and hence we do not follow such alternative. Lemma 4 in the Appendix shows that (ref) also holds for $\widehat{R}^{(2)}\big(\beta_h(\tau);\tau\big)$ in lieu of $\widehat{D}_h(\tau)$. Using the standardization $\widehat{D}_h(\tau)$ instead of the infeasible $\widehat{R}^{(2)}\big(\beta_h(\tau);\tau\big)$ is important for practical purposes, such as computing Wald statistics.

It also follows from Theorems 1 and 2 that $\widehat{\beta}_h(\cdot)$ offers a fair global picture of $\beta(\cdot)$ in that

equation[equation omitted — 150 chars of source]

uniformly for $\tau\in[\underaccent{\bar}{\tau},\bar{\tau}]$ and $h\in[\underaccent{\bar}{h}_n,\bar{h}_n]$. Accordingly, the remainder in (ref) is of order $O_p(n^{-1/2})$ for any $h\le O\left(n^{-1/(2(s+1))}\right)$. The latter restriction is quite light, holding not only for the infeasible AMSE-optimal bandwidth in Theorem 4 but also for the corresponding rule-of-thumb bandwidth that we suggest in Section (ref).

Lastly, it is important to stress the major role that uniformity plays here. It ensures that, if a random (possibly data-driven) bandwidth process $\left\{\widehat{h}(\tau);\,\tau\in[\underaccent{\bar}{\tau},\bar{\tau}]\right\}$ has sample paths in $[\underaccent{\bar}{h}_n,\bar{h}_n]$ with a sufficiently high probability, then (ref) remains valid even if we replace $h$ with a data-driven bandwidth $\widehat{h}(\tau)$. The next result states this property in a rigorous manner.\vskip 1em

\noindentCorollary 1 If $\left\{\widehat{h}(\tau);\,\tau\in[\underaccent{\bar}{\tau},\bar{\tau}]\right\}$ satisfies $\mathrm{Pr}\!\left(\widehat{h}(\tau)\in[\underaccent{\bar}{h}_n,\bar{h}_n]~\mathrm{for~all}~\tau\right)\rightarrow 1$, then both (ref) and (ref) hold with $\widehat{h}(\tau)$ in place of $h$, uniformly in $\tau$.\vskip 1em

The asymptotic theory so far posits that our smoothed QR estimator entails a better Bahadur-Kiefer representation than the standard QR estimator and that we may employ a data-driven bandwidth that depends on the quantile level and covariates. In the next section, we complement the asymptotic theory by characterizing the AMSE of our convolution-type kernel QR estimator as well as the bandwidth choice that minimizes it.

Asymptotic mean squared error

The asymptotic covariance matrix of $\widehat{\beta}_h(\tau)$ comes from the leading term of its Bahadur-Kiefer linear representation in (ref). The next result not only characterizes this asymptotic covariance matrix but also shows that it is smaller than the asymptotic covariance matrix of $\widehat{\beta}(\tau)$. In what follows, let $\Sigma_h(\tau):=\mathbb{V}\!\left(\sqrt{n}\,D_h^{-1}(\tau)\widehat{S}_h(\tau)\right)$.\vskip 1em

\noindentTheorem 3 Assumptions X, Q and K ensure that

equation[equation omitted — 100 chars of source]

with $c_k=2\int_0^\infty K(y)[1-K(y)]\,\mathrm{d}y>0$, uniformly with respect to $(\tau,h)\in[\underaccent{\bar}{\tau},\bar{\tau}]\times[\underaccent{\bar}{h}_n,\bar{h}_n]$.}\vskip 1em

Theorem 3 shows that the asymptotic covariance matrix of the smoothed QR estimator is equal to the asymptotic covariance matrix of the standard QR estimator minus a term $c_k\,h\,D^{-1}(\tau)$ induced by smoothing. Table (ref) reports the values of $c_k$ for Gaussian-type kernels that satisfy Assumption K1. The order $\ln n/(nh)$ of the squared Bahadur-Kiefer remainder term is negligible with respect of the order $h$ of the smoothing-related term given that $1/h=o(\sqrt{n/\ln n})$ by Assumption K2.\vskip 1em

It is sometimes possible to refine the asymptotic covariance matrix expansion in Theorem 3 using strong approximation tools. For instance, assume there is no covariate and that $Y_i$ is uniform over $[0,1]$. In this case, there is no bias in that $\beta_h(\tau)=\tau$ and $D_h(\tau)=D(\tau)=1$ for $h$ small enough. The score function then reads

displaymath\widehat{S}_h(\tau)=\int\frac{1}{\sqrt{n}}\sum_{i=1}^n\Big[\mathbb{I}(Y_i\le\tau+ht)-(\tau+ht)\Big]k(t)\,\mathrm{d}t,

using $\int k(t)\,\mathrm{d}t=1$ and $\int tk(t)\,\mathrm{d}t=0$. The Koml\'os-Major-Tusn\'ady strong approximation result shows that one may reconstruct the sample jointly with a sequence of Brownian bridges $B_n(\cdot)$ such that $\sup_{y\in[0,1]}\left|\frac{1}{\sqrt{n}}\sum_{i=1}^n\Big[\mathbb{I}(Y_i\le\tau+ht)-(\tau+ht)\Big]-B_n(t)\right|=O_p\!\left(\frac{\ln n }{\sqrt{n}}\right)$; see \citeasnoun{pollard2002user}. It follows that, if $k$ has a compact support,

eqnarray[eqnarray omitted — 264 chars of source]

As the integral term that comes from smoothing vanishes for $h=0$, $B_n(\tau)$ gives the limit distribution of the standard QR estimator.

The covariance between $B_n(\tau)$ and $B_n(\tau+ht)-B_n(\tau)$ is negative of order $h$. This implies a negative covariance between $B_n(\tau)$ and the integral term in (ref). As the variance of the latter is of smaller order than this negative covariance, the asymptotic variance of $\widehat{S}_h(\tau)$ is smaller than $\mathbb{V}(B_n(\tau))=\tau(1-\tau)$ by a term of order $h$. Theorem 3 also implies that the distribution of $D_h(\tau)^{-1}\widehat{S}_h(\tau)$ is a centered normal with variance $\tau(1-\tau)-c_k h$ up to the order $\ln n /\sqrt{n}$. Because the second term in (ref) is negative, the asymptotic variance of $\widehat{\beta}_h(\tau)$ is indeed smaller than the asymptotic variance of $\widehat{\beta}(\tau)$.

We next focus on obtaining the bandwidth $h_\lambda^*$ that minimizes the asymptotic mean squared error of $\lambda'\widehat{\beta}_h(\tau)$ for a given $\lambda\in\mathbb{R}^d$: $\mathrm{AMSE}\big(\lambda'\widehat{\beta}_h(\tau)\big)=\mathbb{E}\!\left[\lambda'\big(\beta_h(\tau)-D_h^{-1}(\tau)\widehat{S}_h(\tau)-\beta(\tau)\big)\right]^2$. This approximates the mean squared error $\mathrm{MSE}\big(\lambda'\widehat{\beta}_h(\tau)\big)=\mathbb{E}\!\left[\lambda'\big(\widehat{\beta}_h(\tau)-\beta(\tau)\big)\right]^2$ by essentially ignoring the asymptotically negligible remainder term of the Bahadur-Kiefer representation. To do so, we require that the bandwidth is such that $1/h=o\!\left(\sqrt{n/\ln n}\right)$; see Theorems 2 and 3.\vskip 1em

\noindentTheorem 4 Let Assumptions X, Q and K hold. If $\lambda'B(\tau)\ne 0$, and the conditional density $f(\cdot\,|\, x)$ is $s$-times continuously differentiable for all $x$, then $\mathrm{AMSE}\big(\lambda'\widehat{\beta}_h(\tau)\big)$ is minimal for

equation[equation omitted — 162 chars of source]

and equal to $\mathrm{AMSE}\big(\lambda'\widehat{\beta}_{h_\lambda^*}(\tau)\big)=\frac{1}{n}\,\lambda'\left[\Sigma(\tau)-c_k\,h_\lambda^*\,\frac{2s+1}{2s+2}\,D^{-1}(\tau)\right] \lambda+o(h_\lambda^*/n)$.}\vskip 1em

The optimal bandwidth we derive in Theorem 4 obviously depends on $\lambda$. This is not the case of \possessivecite{kaplan2017smoothed} bandwidth choice, which minimizes $\mathbb{E}\Big[\big\Vert\widehat{\beta}(\tau)-\beta(\tau)\big\Vert^2\Big]$. Although their optimal bandwidth has the same order of $h_\lambda^*$ in (ref), it obviously does not depend on $\lambda$ given that they focus on the global estimation of $\beta(\tau)$. As such, it is suboptimal for a linear combination of the elements of $\beta(\tau)$, as we consider here. This matters for instance in the particular case of a standard linear regression in view that the optimal bandwidth is well defined only for the intercept as the leading term of the bias vanishes for the slope coefficients.

It is easy to appreciate that Theorem 4 remains valid for any bandwidth $h_\lambda=[1+o(1)]h_\lambda^*$. Although entertaining a plug-in bandwidth based on the nonparametric estimation of $D(\tau)$ and $B(\tau)$ is certainly feasible, it is perhaps not very advisable due to the presence of the pdf derivative in the bias term. We circumvent this issue by means of a rule-of-thumb approach. In particular, it follows from the objective function in ((ref)) that it is as if our convolution-kernel QR estimator were using a kernel-based estimator of the cdf of the standard QR residuals. This naturally leads to using \possessivecite{silverman1986density} rule-of-thumb bandwidth based on residual dispersion measures (e.g., sample standard deviation or interquantile range).

Inference

In this section, we derive a functional limit distribution of the smoothed QR estimator that holds for data-driven bandwidths. We also discuss how to consistently estimate the asymptotic covariance matrix of $\widehat{\beta}_h(\tau)$ in order to make inference. In particular, we establish not only that $\widehat{\Sigma}_h(\tau):=\widehat{D}_h^{-1}(\tau)\widehat{V}_h(\tau)\widehat{D}_h^{-1}(\tau)$, with $\widehat{D}_h(\tau):=\widehat{R}_h^{(2)}\big(\widehat{\beta}_h(\tau);\tau\big)$ and

displaymath\widehat{V}_h(\tau):=\frac{1}{n}\sum_{i=1}^nX_iX_i'\left[K\left(-\frac{e_i\big(\widehat{\beta}_h(\tau)\big)}{h}\right)-\tau\right]^2,

is a consistent estimator of $\Sigma(\tau)$, but also the asymptotic normality of our convolution-type kernel QR estimator.\vskip 1em

\noindentTheorem 5 Let Assumptions X, Q and K hold.

description$\Big\{\sqrt{n}\big(\widehat{\beta}_h(\tau)-\beta_h(\tau)\big):\,\tau\in[\underaccent{\bar}{\tau},\bar{\tau}]\Big\}$ converges in distribution to a centered Gaussian process $\big\{W(\tau):\,\tau\in[\underaccent{\bar}{\tau},\bar{\tau}]\big\}$ with covariance structure \begin{equation} \mathbb{V}\big(W(\tau),W(\varsigma)\big)=(\tau\wedge\varsigma-\tau\varsigma)D(\tau)^{-1}\mathbb{E}(XX')D(\varsigma)^{-1},\qquad\tau,\varsigma\in[\underaccent{\bar}{\tau},\bar{\tau}]. \end{equation} • Consider a data-driven bandwidth $\widehat{h}_n$ which satisfies $\widehat{h}_n=h_0\left\{1+o_p(1/\sqrt{h_n \ln n})\right\}$ for some deterministic $h_n$ in $[\underaccent{\bar}{h}_n,\overline{h}_n]$. $\Big\{\sqrt{n}\big(\widehat{\beta}_{\,\widehat{h}_n}(\tau)-\beta_{\,\widehat{h}_n}(\tau)\big):\,\tau\in[\underaccent{\bar}{\tau},\bar{\tau}]\Big\}$ converges in distribution to $\big\{W(\tau):\,\tau\in[\underaccent{\bar}{\tau},\bar{\tau}]\big\}$. • If $\underaccent{\bar}{h}_n=o(1/\ln n)$, the convergence result in (G) is uniform in that \begin{displaymath} \sup_{h\in[\underaccent{\bar}{h}_n,\overline{h}_n]}\left\Vert{\sqrt{n}\left(\widehat{\beta}_h(\tau)-\beta_h(\tau)-\left[\widehat{\beta}_{h_n}(\tau)-\beta_{h_n}(\tau)\right]\right)}\right\Vert=o_p(1) \end{displaymath} for any arbitrary deterministic sequence $h_n$ in $[\underaccent{\bar}{h}_n,\overline{h}_n]$. • $\widehat{\Sigma}_h(\tau)=\Sigma_h(\tau)+O_p\left(\sqrt{\ln n/(nh)}\right)+o(h^s)$ uniformly for $(\tau,h)\in[\underaccent{\bar}{\tau},\bar{\tau}]\times[\underaccent{\bar}{h}_n,\bar{h}_n]$.

\vskip 1em

Theorem 5(G) shows that the functional limit distribution of the smoothed and standard QR estimators are identical. Theorem 5(S) posits that such a functional central limit theorem (FCLT) holds for data-driven bandwidths that converge to a deterministic counterpart. Theorem 5(U) states that the FCLT holds uniformly with respect to the bandwidth. Theorem 5(V) implies not only that the confidence intervals of Section 2.1 asymptotically have the desired level, but also that $\widehat{\Sigma}_h(\tau)$ is a consistent estimator of the asymptotic covariance matrix of both standard and smoothed QR estimators. Under Assumption K2, as $s\ge 1$,

displaymath\widehat{\Sigma}_h(\tau)=\Sigma_h(\tau)+o_p(h)=\Sigma(\tau)-c_k\,h\,D^{-1}(\tau)+o_p(h)

by Theorem 3. This means that $\widehat{\Sigma}_h(\tau)$ captures the improvement in precision due to smoothing and, accordingly, the length of a confidence interval based on $\widehat{\Sigma}_h(\tau)$ for any element of $\beta(\tau)$ is asymptotically smaller than the length of a confidence interval based on the naive variance estimator $\tau(1-\tau)\widehat{D}_h^{-1}(\tau)\,\frac{1}{n}\sum_{i=1}^nX_iX_i'\widehat{D}_h^{-1}(\tau)$.

Note that both (S) and (U) make use of the implicit function theorem applied to the first-order condition $\widehat{R}_h^{(1)}\big(\widehat{\beta}_h(\tau);\tau\big)=0$ to obtain

displaymath\frac{\partial\widehat{\beta}_h(\tau)}{\partial h}=-\left[\widehat{R}_h^{(2)}\big(\widehat{\beta}_h(\tau);\tau\big)\right]^{-1}\frac{\partial\widehat{R}_h^{(1)}\big(\widehat{\beta}_h(\tau);\tau\big)}{\partial h}.

We then establish a convergence rate that holds uniformly in $h$ for $\widehat{R}_h^{(2)}\big(\widehat{\beta}_h(\tau);\tau\big)$ and $\frac{\partial}{\partial h}\,\widehat{R}_h^{(1)}\big(\widehat{\beta}_h(\tau);\tau\big)$, leading to

eqnarray*[eqnarray* omitted — 423 chars of source]

uniformly for $(h_0,h_1)\in[\underaccent{\bar}{h}_n,\overline{h}_n]^2$.

Extensions

\paragraph{Pdf curve estimation.} The next result states the uniform consistency of the pdf curve estimator. A distinctive feature of the estimation procedure we propose is that consistency holds even when the covariate value $x$ is not in the support of $X$. The key for this ability to extrapolate lies on the underlying quantile regression specification, which we assume correct.\vskip 1em

\noindentProposition 1 Under Assumptions X, Q and K, $\left\Vert{\widehat{\texttt{f}}_h(\tau|x)-\texttt{f}(\tau|x)}\right\Vert=o(h^s)+O_p\big(\sqrt{\ln n/(nh)}\big)$ uniformly for $x$ in any compact set, $\tau$ in $[\underline{\tau},\overline{\tau}]$, and $h$ in $[\underline{h}_n,\overline{h}_n]$.\vskip 1em

Proposition 1 shows that the dimension of the covariate vector $X$ does not affect the order $\sqrt{\ln n/(nh)}$ of the stochastic estimation error of $\widehat{\texttt{f}}_h(\tau|x)$, so that the curse of dimensionality does not apply. This result obviously depends on the correct specification of the quantile regression model. The bias in Proposition 1 is of order $o(h^s)$ due to the Hessian bias in Lemma 1. The latter arises instead of the usual $O(h^s)$ bias order because we calibrate the kernel order for the QR estimation, which involves a function with $s+1$ derivatives. In contrast, it is the estimation of the pdf, which has only $s$ derivatives, that drives the bias in the Hessian term.

As it turns out, the quantile density estimator $\widehat{q}_h (\tau|x)$ drives the nonparametric consistency rate. Alternatively, we could estimate the pdf by $\widehat{f}_h(y|x)=1/\widehat{q}_h\big(\widehat{F}_h(y|x)|x\big)$, where $\widehat{F}_h(y|x)$ is a conditional cdf estimator. The consistency rate of such $\widehat{f}_h(y|x)$ is the same as in Proposition 1 if $\widehat{F}_h(y|x)$ ensues from inverting $x'\widehat{\beta}_h(\tau)$ for $y$ in its range.

\paragraph{Semiparametric efficiency.} We next establish the asymptotic equivalence of the feasible QR estimator $\check{b}_q(\tau)$ in ((ref)) to its unfeasible counterpart $\tilde{b}_q(\tau)$ in ((ref)). In addition, we also document the asymptotic normality and efficiency of both estimators. Before stating the result, recall that \citeasnoun{Newey1990efficient} show that $\Sigma_q(\tau):=\tau(1-\tau)D_q^{-1}(\tau)$, with

displaymathD_q(\tau):=\mathbb{E}\left[\frac{XX'}{q^2(\tau|X)}\right]=\mathbb{E}\left\{XX'f^2[X'\beta(\tau)|X]\right\},

is the asymptotic covariance matrix of the asymptotically efficient QR estimator.\vskip 1em

\noindentProposition 2 Let $m=o(n^{1/2})$ and $h=o(1)$, with $1/h=o\big((m/\ln m)^{1/3}\big)$. Assumptions X, Q and K1 ensure that $\sqrt{n}\left(\check{b}_q(\tau)-\tilde{b}_q(\tau)\right)=O_p(\varrho_q+\varrho_s+\varrho_{bk})$, where $\varrho_q=h^s+\sqrt{\ln m/(mh)}$, $\varrho_s=\sqrt{m/n}$, and $\varrho_{bk}=(\ln n)^{3/4}\,n^{-1/4}$. In addition, both $\sqrt{n}\big(\check{b}_q(\tau)-\beta(\tau)\big)$ and $\sqrt{n}\big(\tilde{b}_q(\tau)-\beta(\tau)\big)$ converge in distribution to $\mathcal{N}\big(0,\Sigma_q(\tau)\big)$.\vskip 1em

Both $\check{b}_q(\tau)$ and $\tilde{b}_q(\tau)$ are asymptotically efficient, with asymptotic covariance matrix $\Sigma_q(\tau)$. To estimate the latter, one may employ a weighted version of the Hessian $\widehat{R}^{(2)}_h (b;\tau)$. The rate at which the difference $\sqrt{n}\left(\check{b}_q(\tau)-\tilde{b}_q(\tau)\right)$ shrinks to zero involves three components. The first is the consistency rate $\varrho_q$ of the qdf estimator $\check{q}_h(\tau|x)$ in Proposition 1. Using standard kernel estimators as in \citeasnoun{otsu2008conditional} and \citeasnoun{komunjer2010efficient} or $k-$NN smoothing as in \citeasnoun{Newey1990efficient} and \citeasnoun{Zhao2001asymptotically} would also involve a variance term of order $1/(mh^d)$ due to the curse of dimensionality. The rate $\varrho_s$ comes from the sample splitting scheme, whereas $\varrho_{bk}$ arises from the Bahadur-Kiefer linear approximation.

A natural extension of Proposition 2 is to consider smoothed versions of the objective functions corresponding to the efficient feasible and infeasible estimators, in order to obtain better AMSE performances.

Monte Carlo study

To assess how well the asymptotic theory reflects the performance of our estimator in finite samples, we run simulations for a median linear regression: $Y=X'\beta+\epsilon$, where $X=(1,\tilde{X})$, with $\tilde{X}\sim U[1,5]$, and $\beta\equiv\beta(1/2)=(1,1)$. We entertain five different specifications for the error distribution. The first three are asymmetric, with errors coming from exponential, Gumbel (type I extreme value), and $\chi_3^2$ distributions. The fourth specification displays heavy tails, with errors following a t-student with 3 degrees of freedom. The fifth distribution exhibits conditional heteroskedasticity: $\epsilon=\frac{1}{4}\,(1+\tilde{X})\,Z$, with $Z\sim\mathcal{N}(0,1)$. We recenter, if necessary, the error distributions to ensure median zero and, except to the $\chi_3^2$ distribution, also rescale them to obtain unconditional variance equal to two. Apart from the exponential distribution, the other specifications are as in \citeasnoun{horowitz1998bootstrap}, \citeasnoun{whang2006smoothed}, and \citeasnoun{kaplan2017smoothed}.

In each of the 100,000 replications, we sample $n\in\{100,250,500,1000\}$ observations and then compute \possessivecite{koenker1978regression} standard median regression estimator (MR), \possessivecite{horowitz1998bootstrap} smoothed median regression estimator (SMR), and our convolution-type kernel estimator (CKMR). We also compute the empirical coverage of their asymptotically-valid confidence intervals at the 95% and 99% levels. We gauge the latter as the proportion of replications in which the absolute value of the t-statistic is below the corresponding percentile in the standard normal distribution (namely, 1.96 and 2.58 at the 95% and 99% confidence levels, respectively).

We compute the smoothed estimators using a standard Gaussian kernel and a bandwidth grid with values ranging from $0.08$ to $0.80$, with increments of 0.02. In addition, we also evaluate them at \possessivecite{silverman1986density} rule-of-thumb bandwidth $h_{ROT}=1.06\,\widehat{s}\,n^{-1/5}$, where $\widehat{s}$ is the minimum between the sample standard deviation and the interquartile range (divided by 1.38898) of the standard MR residuals. We compute the MR standard errors by a pairs bootstrap procedure as usual in the literature,\footnote{ We omit results using standard errors as in \citeasnoun[Sections 3.4.2 and 4.10.1]{koenker2005quantile} because bootstrapping always entails better empirical coverage. They are available from the authors upon request.} whereas the SMR standard errors are as in \citeasnoun[Section 2]{horowitz1998bootstrap}. Finally, we estimate the standard errors for the CKMR estimator using the square root of the diagonal entries of $\widehat{\Sigma}_h(1/2)$.

Figures (ref) to (ref) show how each estimator performs across distributions and sample sizes. For simplicity, we focus exclusively on the slope coefficient given that we find little difference between the intercept estimates. For every figure, the first row displays the relative mean squared error (RMSE) of the smoothed estimators with respect to the standard MR estimator across different sample sizes, whereas the second row depicts the standard error of the slope estimators, within their one-standard-deviation band. The third and fourth rows portray respectively the empirical coverage of the asymptotic confidence intervals at the 95% and 99% levels.

The smoothed estimators obviously depend on the bandwidth choice, and hence we plot their results as a function of the bandwidth $h$, singling out the corresponding outcomes for the rule-of-thumb bandwidth $h_{ROT}$. As the value of the latter changes across replications, we arbitrarily place the results at the average value of $h_{ROT}$ across the 100,000 replications. The inside tick marks in the horizontal axis depict the deciles of the distribution of the rule-of-thumb bandwidth across replications. The distribution seems symmetric in every instance, with dispersion reducing drastically once we move from a sample size of 100 to 1,000 observations.

Smoothing seems to pay off as both SMR and CKMR entail lower MSE than the standard MR estimator for every distribution. Interestingly, the RMSE of the CKMR estimator is less sensitive to the sample size and bandwidth value, especially around the rule-of-thumb choice, than the RMSE of the SMR estimator. For instance, if errors are exponential as in Figure (ref), the CKMR estimator using the rule-of-thumb bandwidth yields a RMSE of about 70% for the smaller sample sizes, and between 75% and 80% for the sample sizes of 500 and 1,000 observations. Conversely, the RMSE of the SMR estimator changes from 80% to 90% as we increase the sample size from 100 to 250 observations and then cease to improve the standard MR estimator in terms of MSE for the larger sample sizes.

Regardless of the specification we consider, the key to the CKMR success lies on the much lower standard errors. It is apparent from Figures (ref) to (ref) that the standard errors of the SMR estimator that uses the rule-of-thumb bandwidth are very close to the bootstrap-based standard errors of the MR estimator. They are slightly lower for the exponential and chi-squared errors, but slightly larger for the Gumbel, t-student and heteroskedastic specifications.

The confidence intervals match very well their nominal values for every estimator at the 95% and 99% levels. This is particularly reassuring for the CKMR estimator, whose confidence intervals are significantly narrower due to the smaller standard errors. In addition, the empirical coverage of the CKMR estimator is very stable across bandwidth values, especially for samples with 250 or more observations. In contrast, the empirical coverage of the SMR estimator deteriorates for larger bandwidth values regardless of the error distribution and sample size. All in all, the CKMR estimator performs very well relative to the MR and SMR estimators, with lower MSE and tighter (asymptotic) confidence intervals that are nonetheless valid even in small samples.

Conditional qdf and pdf estimation

Next, we complement the discussion in Section (ref) by illustrating the qdf and pdf estimation of the QR model $Y=\beta_0(U)+\beta_1(U)X_1+\beta_2(U)X_2+\beta_3(U)X_3$, where $U$, $X_1$, $X_2$ and $X_3$ are independent and uniformly distributed variables on the unit interval $[0,1]$, $\beta_0(\cdot)$ and $\beta_1(\cdot)$ are respectively given by the $\mathsf{Beta}(1,16)$ and $\mathsf{Beta}(32,32)$ quantile functions, $\beta_2(\tau)=1$, and $\beta_3(\tau)=\frac{(2\pi+8)\tau-(\cos(2\pi\tau)-1)}{2\pi+8}$. The large number of regressors relative to the sample size of 1,000 observations does not bode well for standard nonparametric estimation due to the curse of dimensionality.

In each of the 1,000 replications, we estimate the pdf by means of (ref) using the rule-of-thumb bandwidth $h_{ROT}$.\footnote{ It is perhaps worth mentioning that the optimal bandwidth for the QR estimation is not necessarily optimal for the qdf estimator.} Figure (ref) displays the true quantile function in Panel (a), the true qdf in Panel (b), and the true conditional density of $Y$ given $(x_1,x_2,x_3)=(0.9,0.5,0.9)$ in Panel (c). The gray-shaded regions in Panels (a) and (b) correspond to the interquantile band from 0.05 to 0.95, i.e., the region for which at most 5% of the replications lay below its lower bound and at most 5% above its upper bound, at each fixed $\tau$.

Figure (ref)(a) suggests that the variability and bias of the QR estimator are not only quite small but also of a similar magnitude. Panel (b) reveals that the qdf estimates approximate well the poles in the tails, up to some boundary bias at $\tau=0$, with a very narrow band regardless of the quantile level $\tau$. Panel (c) displays the paths $\tau\mapsto\widehat{\texttt{f}}_h(\tau\vert x)$ for $\tau\in\{0.01,0.02,\dots,0.99\}$, corresponding to the 100 first realizations of $\widehat{\texttt{f}}_h$. We find a much higher variability for the pdf curve estimates, as typical in nonparametric estimation. As before, we evince some boundary bias at $\tau=0$ and tighter interquantile bands in the tails. Altogether, the pdf estimator seems to perform reasonably well even for three conditioning variables.

Concluding remarks

This paper proposes a convolution-type kernel QR estimator based upon smoothing the objective function. The resulting estimator improves on standard QR estimation both in terms of asymptotic mean square error (MSE) and Bahadur-Kiefer representation. Applying the implicit function theorem to the corresponding first-order condition gives way to a conditional pdf estimator that avoids the curse of dimensionality by exploiting the QR specification. We also show how to use this pdf estimator to come up with a two-step efficient QR estimator. Many of our asymptotic results are uniform with respect to the quantile level and bandwidth. The latter is convenient because it makes room for data-driven bandwidth choices, such as the rule-of-thumb bandwidth we propose. In particular, we establish a functional central limit theorem for our smoothed QR estimator that accommodates data-driven bandwidths.

Simulations show that the rule-of-thumb bandwidth works well in practice, yielding very good results in terms of confidence interval coverage. They also reveal that our smoothed QR estimator outperforms not only the standard QR estimator, but also \possessivecite{horowitz1998bootstrap} alternative smoothed estimator. Finally, our Monte Carlo study also confirms that our QR-based pdf estimator does not suffer from the curse of dimensionality.

There are many extensions to consider in future research. First, uniformity in bandwidth is useful not only for adaptive bandwidth choices as in \citeasnoun{lepski1997optimal}, but also for bandwidth-snooping-robust inference as in \citeasnoun{AK18}. Second, one could combine our QR-based pdf estimator in a data-rich environment with \possessivecite{belloni2011penalized} lasso-type methods to select the relevant covariates and with \possessivecite{volgushev2019distributed} to cope with large sample sizes. Third, one could exploit the connection between QR and quantile instrumental variable models kaido2018decentralization to obtain smoother estimators and better inferential procedures chernozhukov2005ivqte,J08,AM16,dCGK2017. Finally, one could also extend our smoothing approach to handle panel QR as in \citeasnoun{GK2016}.

\lineskip=.35ex \baselineskip 3.5ex