EconBase
← Back to paper

Uniform Convergence Results for the Local Linear Regression Estimation of the Conditional Distribution

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

26,994 characters · 5 sections · 29 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Uniform Convergence Results for the Local Linear Regression Estimation of the Conditional Distribution

\thispagestyle{empty}

abstract\doublespacing This paper examines the local linear regression (LLR) estimate of the conditional distribution function $F(y|x)$. We derive three uniform convergence results: the uniform bias expansion, the uniform convergence rate, and the uniform asymptotic linear representation. The uniformity in the above results is with respect to both $x$ and $y$ and therefore has not previously been addressed in the literature on local polynomial regression. Such uniform convergence results are especially useful when the conditional distribution estimator is the first stage of a semiparametric estimator. We demonstrate the usefulness of these uniform results with two examples: the stochastic equicontinuity condition in $y$, and the estimation of the integrated conditional distribution function. Keywords: Uniform Bias Expansion, Uniform Convergence Rate, Uniform Asymptotic Linear Representation.

Introduction

This paper studies the nonparametric estimation of the conditional distribution function. The analysis concerns a random variable $Y \in \mathbb{R}$ and a random vector of covariates $X \in \mathbb{R}^d$. The conditional distribution function of $Y$ given $X=x$ is denoted by $F(\cdot|x)$, that is,

align*[align* omitted — 71 chars of source]

When the conditional distribution function $F(\cdot|\cdot)$ is assumed to be smooth, it is natural to consider using the local linear regression (LLR) method to estimate $F$.

The main subject of this study is the uniform convergence of the LLR estimator with respect to both $y$ and $x$. In particular, we derive the uniform bias expansion, characterize the uniform convergence rate, and present the uniform asymptotic linear representation of the estimator. As explained in, for example, hansen2008uniform and kong2010uniform, these uniform results are often useful for semiparametric estimation based on nonparametrically estimated components.

The estimation of the conditional distribution is an important area of research. hansen2004nonparametric studies the asymptotic properties of both the Nadaraya-Watson (local constant) estimator and the LLR estimator, and obtains point-wise convergence results. It is well-known that the LLR estimator has the better boundary properties of the two, but unlike the Nadaraya-Watson estimator, the LLR estimator is not guaranteed to be a proper distribution function.\footnote{To solve this problem hall1999methods propose a weighted Nadaraya-Watson estimator that has the same asymptotic distribution as the LLR estimator, but these weights require extensive computation.} Recently, politis2020nonparametric propose a way to correct the LLR estimator. The conditional distribution estimation is also useful for estimating conditional quantiles. For example, yu1997thesis and yu1998local first estimate the conditional distribution function and then invert it to obtain the conditional quantile function.

The local polynomial estimators have been studied extensively, but the uniform convergence results for the estimation of $F$ are new to the literature. In the general setup of local polynomial estimators, there is only one regressand, namely, $Y$. However, in the conditional distribution estimation, there is a class of regressands, namely, $\mathbf{1}\{Y \leq y\}, y \in \mathbb{R}$. For example, masry1996multivariate establishes the uniform convergence rate for general local polynomial estimators, but the uniformity is with respect to the values of the regressors. Therefore, their results can only be applied to an estimate of $F(y|\cdot)$ for a fixed $y \in \mathbb{R}$. For the same reason, the results in kong2010uniform cannot be used to provide a uniform asymptotic linear representation for $y \in \mathbb{R}$. Our paper aims to solve these issues and prove that under suitable conditions, the desired results are uniform with respect to both $y$ and $x$. We make use of the recent discovery by fan2016multivariate on the support of the covariates, ensuring that the uniform results are valid over the entire support.

The second contribution of the paper is the presentation of a novel way of proving the uniform convergence rate via empirical process theory. This theory was developed by GINE2001consistency and GINE2002rates and supports the uniform almost sure convergence of the kernel density estimator. In this paper, we simplify their method and make it more accessible to users who are only concerned with the notion of uniform convergence in probability.

The remaining parts of the paper are organized as follows. Section (ref) introduces the statistical model and the assumptions. Section (ref) establishes the uniform bias expansion result. Section (ref) introduces empirical process theory and uses it to prove the uniform convergence rate. Section (ref) presents the uniform asymptotic linear representation and provides a simple example to illustrate the result. The proofs for the theorems in the main text are found in Appendix (ref). Appendix (ref) contains some preliminary results for empirical process theory.

Model and Assumptions

Let $\{(Y_i,X_i), 1 \leq i \leq n\}$ be a random sample of $(Y,X)$. The estimation procedure is described as follows. Let $w$ and $k$ be two kernel functions and $K(v) = \int_{-\infty}^v k(u) du$. Let $h_1 = h_{1n} = o(1)$ and $h_2 = h_{2n} = o(1)$ be two scalar sequences of bandwidths. Let $\bm{r}(u) = (1,u^\top)^\top, u \in \mathbb{R}^d$ and $\bm{e}_0 = (1, 0, \cdots, 0)$ be the first $(d+1)$-dimensional unit vector. The proposed estimator is $\hat{F} (y|x) = \bm{e}_0^\top \bm{\hat{\beta}}(y,x,h_1,h_2)$, where

align[align omitted — 412 chars of source]

Let $H_1$ be the $(d+1) \times (d+1)$ diagonal matrix with diagonal elements: $(1,h_1,\cdots,h_1)$. The first-order condition of the above minimization problem gives

align[align omitted — 134 chars of source]

where

align*[align* omitted — 378 chars of source]

In the construction of the estimator, we do not use the indicator $\mathbf{1}\{ Y_i \leq y \}$. Instead, we use the smoothed version $K\left( (y - Y_i)/h_2 \right)$, which requires the selection of another bandwidth $h_2$ and additional smoothness assumptions on the conditional distribution function. However, there are several advantages to using the smoothed version. First, the estimator constructed from the indicators is not smooth in $y$. When we believe that the true distribution function is smooth, it is customary to use the smoothed estimator. Second, from the asymptotic perspective, the indicator $\mathbf{1}\{ Y_i \leq y \}$ can be considered to be the limiting case of $K\left( (y - Y_i)/h_2 \right)$ for $h_2=0$. As hansen2004nonparametric shows, the asymptotic mean squared error is strictly decreasing for $h_2=0$; hence, there are efficiency gains from smoothing. Third, as the simulation results in yu1997thesis and yu1998local demonstrate, the estimates are not very sensitive to the value of $h_2$. Lastly, as we show in Section (ref), the smoothed estimator exhibits a stochastic equicontinuity condition in $y$. This condition is particularly useful when the conditional distribution estimation is an intermediate step in a semiparametric estimation procedure. For example, chen2003estimation provide results on using the stochastic equicontinuity condition to derive the asymptotic distribution of two-step semiparametric estimators.

The following assumptions are maintained throughout the paper.

assumptionp{X} [Distribution of $X$] The support of $X$, denoted by $\mathcal{X}$, is convex and compact. The marginal density $f_X$ is bounded away from zero on $\mathcal{X}$. The restriction of $f_X$ to $\mathcal{X}$ is twice continuously differentiable. There exist $\lambda_0, \lambda_1 \in (0,1]$ such that for any $x \in \mathcal{X}$ and all $\epsilon \in (0,\lambda_0]$, there is $x' \in \mathcal{X}$ satisfying $B(x',\lambda_1 \epsilon) \subset B(x,\epsilon) \cap \mathcal{X}$, where $B(x,\epsilon)$ denotes the ball centered at $x$ with radius $\epsilon$.
assumptionp{Y} [Conditional distribution of $Y|X$] The conditional distribution function $F(y|x)$ restricted to $\mathbb{R} \times \mathcal{X}$ is twice continuously differentiable in $y$ and $x$. Moreover, this second-order derivative of $F$ restricted to $\mathbb{R} \times \mathcal{X}$ is uniformly continuous.
assumptionp{K} [Kernel functions] \ \begin{enumerate} [label = (\roman*)] • The kernel function $w$ is a product kernel, that is, $w(u) = w_1(u_1) w_2(u_2) \cdots w_d(u_d).$ Each $w_\ell$ (1) is a symmetric density function with compact support $[-1,1]$; (2) has its second moment normalized to one, that is, $\int u_\ell^2 w_\ell(u_\ell) du_\ell = 1$; (3) is positive in the interior of the support $(-1,1)$; and (4) is of bounded variation. • The kernel function $k$ (1) is a symmetric density function with a compact support and (2) has its second moment normalized to one, that is, $\int v^2 k(v) dv = 1$. \end{enumerate}

A brief discussion of the above assumptions is in order. Assumption (ref) is introduced by fan2016multivariate as a regularity condition on the support $\mathcal{X}$. It ensures that there are sufficient observations around every estimation location, including the boundary points. Assumption (ref) imposes smoothness conditions on the conditional distribution function $F$. Under this assumption, the Hessian matrix of $F$ is uniformly continuous on the compact support $\textit{supp}(Y,X)$. Assumption (ref) contains standard conditions on the kernel functions $k$ and $w$. The bounded variation condition is imposed for the application of empirical process theory.

Uniform Bias Expansion

We denote the true value of the conditional distribution function and its derivative with respect to $x$ as

align*[align* omitted — 178 chars of source]

where $\nabla_x F (y|x) = \left( \frac{\partial}{\partial x_1} F (y|x), \cdots, \frac{\partial}{\partial x_d} F (y|x) \right)^\top$ is the gradient of $F(y|x)$ with respect to $x$. A convenient way to analyze the estimator $\hat{\bm{\beta}}(y,x,h_1,h_2)$ is to consider it as an estimator of the pseudo-true value defined by

align[align omitted — 428 chars of source]

This pseudo-true value $\bar{\bm{\beta}}$ is deterministic and converges to the true value $\bm{\beta}^*$ as $n \rightarrow \infty$.\footnote{The terminology “pseudo-true” is adopted from fan2016multivariate.} We can break the asymptotic analysis of $\hat{\bm{\beta}}(y,x,h_1,h_2) - \bm{\beta}^*(y,x)$ into two parts:

align*[align* omitted — 251 chars of source]

In this section, we study the bias term, which is the difference between the pseudo-true value and the true value. The first-order condition of ((ref)) gives an explicit expression of the pseudo-true value: $H_1 \bar{\bm{\beta}}(y,x,h_1,h_2) = \Xi(x,h_1)^{-1} \bm{\upsilon}(y,x,h_1,h_2)$, where

align*[align* omitted — 376 chars of source]

Define $\Omega(x,h_1) = \int \bm{r}(u) \bm{r}(u)^\top w(u) \mathbf{1}\{ x+ h_1 u \in \mathcal{X}\} du.$ The following lemma shows that the matrices $\Xi(x,h_1)$ and $\Omega(x,h_1)$ are always bounded and invertible.

lemmaUnder Assumptions (ref) and (ref), there exists $C > 0$ such that the eigenvalues of $\Xi(x,h_1)$ and $\Omega(x,h_1)$ are in $[1/C,C]$ for all $x \in \mathcal{X}$ and $h_1 \geq 0$ small enough.
theoremLet Assumptions (ref), (ref), and (ref) hold. Then \begin{align} & H_1 \left( \bar{\bm{\beta}}(y,x) - \bm{\beta}^*(y,x) \right) \nonumber \\ = & \frac{h_1^2}{2} \Omega(x,h_1)^{-1} \sum_{\ell,\ell' = 1}^d \frac{\partial^2 }{\partial x_\ell \partial x_{\ell'}} F(y|x) \int \bm{r}(u) u_{\ell} u_{\ell'} w(u) \mathbf{1}\{x + h_1 u \in \mathcal{X}\} du \nonumber \\ + & \frac{h_2^2}{2} \Omega(x,h_1)^{-1} \frac{\partial^2 }{\partial y^2} F(y|x) \int \bm{r}(u) w(u) \mathbf{1}\{x + h_1 u \in \mathcal{X}\} du + o(h_1^2 + h_2^2), \end{align} uniformly over $y \in \mathbb{R}$ and $x \in \mathcal{X}$. In particular, we have \begin{align} \bar{\bm{\beta}}_0(y,x) - \bm{\beta}_0^*(y,x) = \frac{h_1^2}{2} \sum_{\ell =1}^d \frac{\partial^2}{\partial x_\ell^2} F(y|x) + \frac{h_2^2}{2} \frac{\partial^2}{\partial y^2} F(y|x) + o(h_1^2 + h_2^2), \end{align} uniformly over $y \in \mathbb{R}$ and $x \in \mathring{\mathcal{X}}_{h_1}$, where $\mathring{\mathcal{X}}_{h_1} = \{ x \in \mathcal{X} : x \pm h_1 = (x_1 \pm h_1, \cdots, x_d \pm h_1) \in \mathcal{X} \}$ denotes the set of interior points with respect to the bandwidth $h_1$.

The novelty of Theorem (ref) is that it provides a uniform bias expansion for the LLR estimator over the entire region $(y,x) \in \mathbb{R} \times \mathcal{X}$. For the boundary points $x \notin \mathring{\mathcal{X}}_{h_1}$, the bias is $O(h_1^2 + h_2^2)$. For the interior points $x \in \mathring{\mathcal{X}}_{h_1}$, the bias expression ((ref)) is the same as in hansen2004nonparametric and Chapter 6 of li2007nonparametric, which contains the curvature of $F(y|x)$.

Uniform Convergence Rate

In this section, we derive the uniform convergence rate of the stochastic term $\hat{\bm{\beta}}(y,x,h_1,h_2) - \bar{\bm{\beta}}(y,x,h_1,h_2)$. We make use of empirical process theory, which is a powerful tool for studying the uniform convergence of random sequences. Some auxiliary concepts and results are introduced below.

Let $\mathcal{G}$ be a class of uniformly bounded measurable functions defined on some subset of $\mathbb{R}^d$, that is, there exists $M>0$ such that $|g| \leq M$ for all $g \in \mathcal{G}$. We say $\mathcal{G}$ is Euclidean with coefficients $(A,v)$, where $A,v>0$, if for every probability measure $P$ and every $\epsilon \in (0,1]$, $N(\mathcal{G},P,\epsilon) \leq A / \epsilon^{v}$, where $N(\mathcal{G},P,\epsilon)$ is the $\epsilon$-covering of the metric space $(\mathcal{G},L_2(P))$, that is, $N(\mathcal{G},P,\epsilon)$ is defined as the minimal number of open $\left\lVert\cdot\right\rVert_{L_2(P)}$-balls of radius $\epsilon$ and centers in $\mathcal{G}$ required to cover $\mathcal{G}$. By definition, if $\mathcal{G}$ is Euclidean with coefficients $(A,v)$, then any subset of $\mathcal{G}$ is also Euclidean with coefficients $(A,v)$.

The above definition of Euclidean classes is introduced by nolan1987uprocess. The same concept is also studied in gine1999laws, but they refer to what we call “Euclidean” as “VC.” There is a slight difference that nolan1987uprocess use the $L_1$-norm, while gine1999laws use the $L_2$-norm. We ignored the envelope in their definition because we only work with uniformly bounded $\mathcal{G}$. The following lemma is useful for deriving the uniform convergence results.

lemmaLet $\xi_1,\cdots,\xi_n$ be an iid sample of a random vector $\xi$ in $\mathbb{R}^d$. Let $\mathcal{G}_n$ be a sequence of classes of measurable real-valued functions defined on $\mathbb{R}^d$. Assume that there is a uniformly bounded Euclidean class $\mathcal{G}$ with coefficients $A$ and $v$ such that $\mathcal{G}_n \subset \mathcal{G} $ for all $n$. Let $\sigma^2_{n}$ be a positive sequence such that $\sigma^2_{n} \geq \sup_{g \in \mathcal{G}_n} \mathbb{E}[g(\xi)^2]$. Then \begin{align*} \Delta_n = \sup_{g \in \mathcal{G}_n} \left| \sum_{i=1}^n (g(\xi_i) - \mathbb{E}g(\xi_i)) \right| = O_p \left( \sqrt{n \sigma_n^2 |\log \sigma_n|} + |\log \sigma_n| \right) . \end{align*} In particular, if $n \sigma_n^2 / |\log \sigma_n| \rightarrow \infty$, then $\Delta_n = O_p \left( \sqrt{n \sigma_n^2 |\log \sigma_n|} \right) .$

The above lemma is based on the results developed by GINE2001consistency and GINE2002rates. These two papers focus on proving the almost sure convergence of kernel density estimators based on empirical process theory. We simplify their method and make it available to users who are only interested in convergence in probability. Based on Lemma (ref), deriving the uniform convergence rate of kernel-based nonparametric estimators boils down to two parts: proving the relevant function classes are Euclidean and computing a uniform bound for the variance.

Controlling the stochastic term does not require the smoothness of the conditional distribution function $F$. We only need the assumptions regarding the support $\mathcal{X}$ and kernel functions. The following theorem establishes the uniform convergence rate for the stochastic term of the LLR estimator. Then, combining this result with Theorem (ref), we obtain the uniform convergence rate of the LLR estimator as a corollary.

theoremLet Assumptions (ref) and (ref) hold. If the bandwidth satisfies $n h_1^d / |\log h_1| \rightarrow \infty$, then \begin{align} \sup_{y \in \mathbb{R}, x \in \mathcal{X}} \left| H_1 \left(\hat{\bm{\beta}}(y,x,h_1,h_2) - \bar{\bm{\beta}}(y,x,h_1,h_2) \right) \right| = O_p \left( \sqrt{\frac{|\log h_1|}{nh_1^d}} \right). \end{align}
corollaryLet Assumptions (ref), (ref), and (ref) hold. Then \begin{align*} \sup_{y \in \mathbb{R}, x \in \mathcal{X}} \left| \hat{F}(y|x) - F(y|x) \right| = O_p \left( h_1^2 + h_2^2 + \sqrt{\frac{|\log h_1|}{nh_1^d}} \right). \end{align*}

We want to compare the above uniform convergence result with the one in masry1996multivariate. In masry1996multivariate, the covariates $X$ are supported on the entire space $\mathbb{R}^d$ while the convergence result is only uniform for $x$ in a compact subset of $\mathbb{R}^d$. In our case, the support $\mathcal{X}$ is compact, and the convergence result is uniform over the entire support $(y,x) \in \mathbb{R} \times \mathcal{X}$.

Corollary (ref) shows that the uniformity over $y \in \mathbb{R}$ does not have an impact on the convergence rate. This is similar to the fact that we can uniformly estimate the unconditional distribution function under the $n^{-1/2}$-rate. The conditional distribution estimation is a nonparametric problem concerning only the covariates.

So far we have been studying the smoothed estimator. It is also of interest to study the unsmoothed version defined as $\check{F} (y|x) = \bm{e}_0^\top \bm{\check{\beta}}(y,x,h_1)$, where

align*[align* omitted — 233 chars of source]

The above minimization problem is constructed by replacing the term $K((y-Y_i)/h_2)$ by the term $\mathbf{1}\{Y_i \leq y\}$. For this estimator, we only need to chose one bandwidth $h_1$. Since $\mathbb{E}[\mathbf{1}\{Y_i \leq y\}|X=x]=F(y|x)$, the bias term of this estimator is only $O(h_1^2)$. The bias associated with the smoothing in $y$ no longer exists. The uniform convergence of $\check{F}$ can be established without the differentiability of $F$ with respect to $y$. The trade-off is that the estimator $\check{F} (y|x)$ itself is not smooth in $y$ even if $F$ is. The stochastic term can be analyzed as before. Then we obtain the following uniform convergence rate for the unsmoothed estimator. Notice that we replace Assumption (ref) by a weaker condition which only requires the smoothness of $F$ with respect to $x$.

theoremLet Assumptions (ref) and (ref) hold. Assume that $F(y|x)$ restricted to $\mathbb{R} \times \mathcal{X}$ is twice continuously differentiable in $x$, and the second-order derivative $\nabla_x^\top\nabla_x F$ is uniformly continuous on $\mathbb{R} \times \mathcal{X}$. If the bandwidth satisfies $n h_1^d / |\log h_1| \rightarrow \infty$, then \begin{align*} \sup_{y \in \mathbb{R}, x \in \mathcal{X}} \left| \check{F}(y|x) - F(y|x) \right| = O_p \left( h_1^2 + \sqrt{\frac{|\log h_1|}{nh_1^d}} \right). \end{align*}

Uniform Asymptotic Linear Representation

This section derives the uniform asymptotic linear representation of the smoothed LLR estimator. These results are particularly useful in deriving the asymptotic distribution for complicated estimators.

theoremLet Assumptions (ref), (ref), and (ref) hold. If the bandwidth satisfies that $n h_1^d / |\log h_1| \rightarrow \infty$, $nh_1^{d+4}/|\log h_1|$ bounded, and $h_2 = O(h_1)$, then \begin{align*} H_1 \left(\hat{\bm{\beta}}(y,x,h_1,h_2) - \bar{\bm{\beta}}(y,x,h_1,h_2) \right) = \Xi(x,h_1)^{-1} \frac{1}{nh_1^d} \sum_{i=1}^n \bm{s}(Y_i,X_i;y,x,h_1,h_2) + O_p \left( \frac{|\log h_1|}{nh_1^d} \right), \end{align*} uniformly over $y \in \mathbb{R}$ and $x \in \mathcal{X}$, where \begin{align*} \bm{s}(Y_i,X_i;y,x,h_1,h_2) & = \bm{r}\left( \frac{X_i - x}{h_1} \right) \left( K\left( \frac{y - Y_i}{h_2} \right) - \tilde{F}(y|X_i) \right) w \left( \frac{X_i - x}{h_1} \right), \\ \tilde{F}(y|x ) & = \mathbb{E}[ K( (y - Y)/h_2 ) \mid X=x ]. \end{align*}

The asymptotic order of the remainder term, $O_p \left( |\log h_1| / nh_1^d \right)$, is the same as in Equation (13) of kong2010uniform. Thus, once more, the uniformity over $y \in \mathbb{R}$ does not have an impact on the convergence rate. Combining the results in Theorem (ref) and (ref) and applying the central limit theorem for triangular arrays, we can show that the LLR estimator is asymptotic normal with some asymptotic bias.

We can use this asymptotic linear representation, together with the smoothness of $K$, to derive the following stochastic equicontinuity condition.

corollaryLet the assumptions of Theorem (ref) hold. Let $\delta_n = o(1)$. Then the following stochastic equicontinuity condition hold. \begin{align*} & \sup_{|y_1 - y_2| \leq \delta_n, x \in \mathcal{X}}\Big| \big(\hat{\bm{\beta}}_0(y_1,x,h_1,h_2) - \bar{\bm{\beta}}_0(y_1,x,h_1,h_2) \big) - \big(\hat{\bm{\beta}}_0(y_2,x,h_1,h_2) - \bar{\bm{\beta}}_0(y_2,x,h_1,h_2) \big) \Big| \\ =& O_p \left( \sqrt{\frac{\log h_1}{nh_1^d}} \frac{\delta_n}{h_2} + \frac{|\log h_1|}{nh_1^d} \right). \end{align*}

We study another simple example to demonstrate how the uniform asymptotic linear representation can be used. Suppose that $d=1$, $\mathcal{X} = [\underline{x},\bar{x}]$, and $Y$ is supported on $[\underline{y},\bar{y}]$. We want to estimate the integrated conditional distribution $\theta = \int_{\underline{y}}^{\bar{y}} \int_{\underline{x}}^{\bar{x}} F(y|x) dxdy$ with the estimator $\hat{\theta} = \int_{\underline{y}}^{\bar{y}} \int_{\underline{x}}^{\bar{x}} \hat{F}(y|x) dxdy$. The theorem below gives the asymptotic distribution of the estimator $\hat{\theta}$.

corollaryLet Assumptions (ref), (ref), and (ref) hold. If the bandwidth satisfies that $\sqrt{n} h_1 / |\log h_1| \rightarrow \infty$, $\sqrt{n}h_1^2 \rightarrow 0$, and $h_2 = O(h_1)$, then $\sqrt{n} (\hat{\theta} - \theta) \overset{d}{\rightarrow} N(0,V)$, where \begin{align*} V = \int \left( \int \left( \mathbf{1}\{s \leq y\} - F(y|t) \right) dy \right)^2 f(s,t) dtds, \end{align*} and $f(y,x)$ denotes the joint density of $(Y,X)$.