Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Some New Asymptotic Theory for Least Squares Series: Pointwise and Uniform Results
\address[Alexandre Belloni]{Fuqua School of Business, Duke University, United States}
\address[Victor Chernozhukov]{Department of Economics, MIT, United States}
\address[Denis Chetverikov]{Department of Economics, UCLA, United States}\email{[email removed] }
\address[Kengo Kato]{Graduate School of Economics, The University of Tokyo, Japan}
abstractIn econometric applications it is common that the exact form of a conditional expectation is unknown and having flexible functional forms can lead to improvements over a pre-specified functional form, especially if they nest some successful parametric economically-motivated forms. Series method offers exactly that by approximating the unknown function based on $k$ basis functions, where $k$ is allowed to grow with the sample size $n$ to balance the trade off between variance and bias.
In this work we consider series estimators for the conditional mean in light of four new ingredients: (i) sharp LLNs for matrices derived from the non-commutative Khinchin inequalities, (ii) bounds on the Lebesgue factor that controls the ratio between the $L^{\infty}$ and $L^{2}$-norms of approximation errors, (iii) maximal inequalities for processes whose entropy integrals diverge at some rate, and (iv) strong approximations to series-type processes.
These technical tools allow us to contribute to the series literature, specifically the seminal work of Newey1997, as follows. First, we weaken considerably the condition on the number $k$ of approximating functions used in series estimation
from the typical $k^2/n \to 0$ to $k/n \to 0$, up to
log factors, which was available only for spline series before. Second, under the same weak conditions we derive $L^{2}$ rates and pointwise central limit theorems results when the approximation error vanishes. Under an incorrectly specified model, i.e. when the approximation error does not vanish, analogous results are also shown. Third, under stronger conditions we derive uniform rates and functional central limit theorems that hold if the approximation error vanishes or not. That is, we derive the strong approximation for the entire estimate of the nonparametric function.
Finally and most importantly, from a point of view of practice, we derive uniform rates, Gaussian approximations, and uniform confidence bands for a wide collection of linear functionals of the conditional expectation function, for example, the function itself, the partial derivative function, the conditional average partial derivative function, and other similar quantities. All of these results are new.
Introduction
Series estimators have been playing a central role in various fields. In econometric applications it is common that the exact form of a conditional expectation is unknown and having a flexible functional form can lead to improvements over a pre-specified functional form, especially if it nests some successful parametric economic models. Series estimation offers exactly that by approximating the unknown function based on $k$ basis functions, where $k$ is allowed to grow with the sample size $n$ to balance the trade off between variance and bias. Moreover, the series modelling allows for convenient nesting of some theory-based models, by simply using corresponding terms as the first $k_0 \leq k$
basis functions. For instance, our series could contain linear and quadratic functions to nest the canonical Mincer equations in the context of wage equation modelling or the canonical translog demand and production functions in the context of demand and supply modelling.
Several asymptotic properties of series estimators have been investigated in the literature. The focus has been on convergence rates and asymptotic normality results vandeGeer1990,Andrews1991,EastwoodGallant1991,GallantSouza1991,Newey1997,vandeGeer2002,Huang2003b,Chen2006,CF2013.
This work revisits the topic by making use of new critical ingredients:
itemize• The sharp LLNs for matrices derived from the non-commutative Khinchin inequalities.
• The sharp bounds on the Lebesgue factor that controls the ratio between the $L^{\infty}$ and $L^{2}$-norms
of the least squares approximation of functions (which is bounded or grows like a $\log k$ in many cases).
• Sharp maximal inequalities for processes whose entropy integrals diverge at some rate.
• Strong approximations to empirical processes of series types.
To the best of our knowledge, our results are the first applications of the first ingredient to statistical estimation problems. After the use in this work, some recent working papers are also using related matrix inequalities and extending some results in different directions, e.g. Chen and Christensen (2013) allows $\beta$-mixing dependence, and Hansen (2014) handles unbounded regressors and also characterizes a trade-off between the number of finite moments and the allowable rate of expansion of the number of series terms. Regarding the second ingredient, it has already been used by Huang2003 but for splines only. All of these ingredients are critical for generating sharp results.
This approach allows us to contribute to the series literature in several directions. First, we weaken considerably the condition on the number $k$ of approximating functions used in series estimation
from the typical $k^2/n \to 0$ Newey1997 to
$$
k/n \to 0 \text{ (up to logs)}
$$
for bounded or local bases which was previously available only for spline series Huang2003, Stone1994, and recently established for local polynomial partition series CF2013. An example of a bounded basis is Fourier series; examples of local bases are spline, wavelet, and local polynomial partition series. To be more specific, for such bases we require $k\log k/n\to 0$. Note that the last condition is similar to the condition on the bandwidth value required for local polynomial (kernel) regression estimators ($h^{-d}\log (1/h)/n\to 0$ where $h=1/k^{1/d}$ is the bandwidth value).
Second, under the same weak conditions we derive $L^{2}$ rates and pointwise central limit theorems results when the approximation error vanishes. Under a misspecified model, i.e. when the approximation error does not vanish, analogous results are also shown.
Third, under stronger conditions we derive uniform rates that hold if the approximation error vanishes or not. An important contribution here is that we show that the series estimator achieves the optimal uniform rate of convergence under quite general conditions. Previously, the same result was shown only for local polynomial partition series estimator CF2013. In addition, we derive a functional central limit theorem. By the functional central limit theorem we mean here that the entire estimate of the nonparametric function is
uniformly close to a Gaussian process that can change with $n$. That is, we derive the strong approximation for the entire estimate of the nonparametric function.
Perhaps the most important contribution of the paper is a set of completely new results that
provide estimation and inference methods
for the entire linear functionals $\theta(\cdot)$ of the conditional mean function $g:\mathcal{X}\to {\Bbb{R}}$. Examples of linear functionals $\theta(\cdot)$ of interest include
itemize• the partial derivative function: \quad $x \mapsto \theta(x) = \partial_j g(x)$;
• the average partial derivative: \quad $\theta = \int \partial_j g(x) d\mu(x)$;
• the conditional average partial derivative: \ $x^s \mapsto \theta(x^s) = \int \partial_j g(x) d\mu(x|x^s)$.
where $\partial_j g(x)$ denotes the partial derivative of $g(x)$ with respect to $j$th component of $x$, $x^s$ is a subvector of $x$, and the measure $\mu$ entering the definitions above is taken as known; the result can be extended to include estimated measures. We derive uniform (in $x$) rates of convergence, large sample
distributional approximations, and inference methods for the functions above based on the Gaussian approximation. To the best of our knowledge all these results are new, especially the distributional and inferential results. For example, using these results we can now perform inference on the entire partial derivative function. The only other reference that provides analogous results but for quantile series estimator is BCF2011. Before doing uniform analysis, we also update the pointwise results of Newey1997 to weaker, more general conditions.
Notation. In what follows, all parameter values are indexed by the sample size $n$, but we omit the index whenever this does not cause
confusion. We use the notation $(a)_+ = \max\{a,0\}$, $a \vee b = \max\{ a, b\}$ and $a \wedge b = \min\{ a , b \}$. The $\ell_2$-norm of a vector $v$ is denoted by $\|v\|$, while for a matrix $Q$ the operator norm is denoted by $\|Q\|$. We also use standard notation in the empirical process literature,
$$
\mathbb{E}_n[f] = \mathbb{E}_n[f(w_i)] = \frac{1}{n}\sum_{i=1}^n f(w_i)\text{ and }\mathbb{G}_n[f]=\mathbb{G}_n[f(w_i)]= \frac{1}{\sqrt{n}}\sum_{i=1}^n (f(w_i) - E[f(w_i)])
$$
and we use the notation $a \lesssim b$ to denote $a \leqslant c b$ for some constant $c>0$ that does not depend on $n$; and
$a\lesssim_P b$ to denote $a=O_P(b)$. Moreover, for two random variables $X, Y$ we say that $X=_dY$ if they have the same probability distribution. Finally, $S^{k-1}$ denotes the space of vectors $\alpha$ in $\mathbb{R}^k$ with unit Euclidean norm: $\|\alpha\|=1$.
Set-Up
Throughout the paper, we consider a sequence of models, indexed by the sample size $n$,
equation[equation omitted — 144 chars of source]
where $y_i$ is a response variable, $x_i$ a vector of covariates (basic regressors), $\epsilon_i$ noise, and $x\mapsto g(x) = E[y_i|x_i =x]$ a regression (conditional mean) function; that is, we consider a triangular array of models with $y_i=y_{i,n}$, $x_i=x_{i,n}$, $\epsilon_i=\epsilon_{i,n}$, and $g=g_n$. We assume that $g\in\mathcal{G}$ where $\mathcal{G}$ is some class of functions. Since we consider a sequence of models indexed by $n$, we allow the function class $\mathcal{G}=\mathcal{G}_n$, where the regression function $g$ belongs to, to depend on $n$ as well. In addition, we allow $\mathcal{X}=\mathcal{X}_n$ to depend on $n$ but we assume for the sake of simplicity that the diameter of $\mathcal{X}$ is bounded from above uniformly over $n$ (dropping the uniform boundedness condition is possible at the expense of more technicalities; for example, without uniform boundedness condition, we would have an additional term $\log \text{diam}(\mathcal{X})$ in ((ref)) and ((ref)) of Lemma (ref)).
We denote $\sigma_i^2=E[\epsilon_i^2|x_i]$, $\bar{\sigma}^2:=\sup_{x\in\mathcal{X}}E[\epsilon_i^2|x_i=x]$, and $\underline{\sigma}^2:=\inf_{x\in\mathcal{X}}E[\epsilon_i^2|x_i=x]$.
For notational convenience, we omit indexing by $n$ where it does not lead to confusion.
Condition A.1 (Sample) For each $n$, random vectors $(y_i, x_i')'$, $i=1,\ldots,n,$ are i.i.d. and satisfy ((ref)).
We approximate the function $x\mapsto g(x)$ by linear forms $x \mapsto p(x)' b$, where
$$
x \mapsto p(x) := (p_{1}(x),\ldots,p_k(x))'
$$
is a vector of approximating functions that can change with $n$; in particular, $k$ may increase with $n$. We denote the regressors as
$$
p_i := p(x_i):= (p_{1}(x_i),\ldots,p_k(x_i))'.
$$
The next assumption imposes regularity conditions on the regressors.
Condition A.2 (Eigenvalues) Uniformly over all $n$, eigenvalues of $Q:= E[p_i p_i']$ are bounded above and away from zero.
Condition A.2 imposes the restriction that $p_1(x_i),\dots,p_k(x_i)$ are not too co-linear. Given this assumption, it is without loss of generality to impose the following normalization:
Normalization. To simplify notation, we normalize $Q = I$, but we shall treat $Q$ as unknown, that is
we deal with random design.
The following proposition establishes a simple sufficient condition for A.2 based on orthonormal bases with respect to some measure.
proposition[Stability of Bounds on Eigenvalues] Assume that $x_i\sim F$ where $F$ is a probability measure on $\mathcal{X}$, and that the regressors $p_1(x),\dots,p_k(x)$ are orthonormal on $(\mathcal{X}, \mu)$ for some measure $\mu$. Then
A.2 is satisfied if $dF/d\mu \text{ is bounded above and away from zero. }$
It is well known that the least squares parameter $\beta$ is defined by
$$
\beta := \arg \min_{b \in \Bbb{R}^k} E\left[(y_i - p_i'b)^2\right],
$$
which by ((ref)) also implies that $\beta=\beta_g$ where $\beta_g$ is defined by
equation[equation omitted — 108 chars of source]
We call $x \mapsto g(x)$ the target function and $x\mapsto g_k(x) = p(x)'\beta$
the surrogate function. In this setting, the surrogate function provides the best linear approximation to the target function.
For all $x\in\mathcal{X}$, let
equation[equation omitted — 78 chars of source]
denote the approximation error at the point $x$, and let
$$
r_i := r(x_i)= g(x_i) - p(x_i)'\beta_g
$$
denote the approximation error for the observation $i$.
Using this notation, we obtain a many regressors model
$$
y_i = p_i'\beta + u_i, \ \ E [u_i x_i] =0, \ \ u_i:= r_i+\epsilon_i.
$$
The least squares estimator of $\beta$ is
equation[equation omitted — 169 chars of source]
where $\widehat Q:=\mathbb{E}_n[p_i p_i']$.
The least squares estimator $\widehat{\beta}$ induces the estimator $\widehat g(x) := p(x)'\widehat \beta$ for the target function $g(x)$. Then it follows from ((ref)) that we can decompose the error in estimating the target function as
$$
\widehat g(x) - g(x) = p(x)'(\widehat \beta - \beta) - r(x),
$$
where the first term on the right-hand side is the estimation error and the second term is the approximation error.
We are also interested in various linear functionals $\theta$ of the conditional mean function. As discussed in the introduction, examples include
the partial derivative function, the average partial derivative function, and the conditional average partial derivative.
Importantly, in each example above we could be interested in estimating $\theta=\theta(w)$ simultaneously for many values $w\in\mathcal{I}$.
By the linearity of the series approximations, the above parameters can be seen as linear functions of the least squares coefficients $\beta$ up to an approximation error, that is
equation[equation omitted — 107 chars of source]
where $\ell_\theta(w)' \beta$ is the series approximation, with
$\ell_\theta(w)$ denoting the $k$-vector of loadings on the coefficients,
and $r_\theta(w)$ is the remainder term, which corresponds to the
approximation error. Indeed, the
decomposition ((ref)) arises from the application of different linear
operators $\mathcal{A}$ to the decomposition $g(\cdot) =
p(\cdot)'\beta + r(\cdot)$ and evaluating the resulting
functions at $w$:
equation[equation omitted — 156 chars of source]
Examples of the operator $\mathcal{A}$ corresponding to the cases enumerated in the introduction are given by,
respectively,
itemize• a differential operator: $(\mathcal{A}f) [x]= (\partial_j f) [x] $, so that
$$\ell_\theta(x) =\partial_j p(x), \ \ \ r_\theta(x) = \partial_j r(x);$$
• an integro-differential operator: $\mathcal{A} f=\int \partial_j f(x) d\mu(x)$, so that
$$\ell_\theta = \int \partial_j p(x) d \mu(x), \ \ \ r_\theta = \int \partial_j r(x) d \mu(x); $$
• a partial integro-differential operator: $(\mathcal{A}f) [x_{2}]=\int \partial_j f(x) d\mu(x|x^s)$, so that
$$\ell_\theta(x^s) =\int \partial_j p(x)d\mu(x|x^s), \ \ \ r_\theta(x^s) = \int \partial_j r(x) d\mu(x|x^s),$$
where $x^s$ is a subvector of $x$. For notational convenience, we use the formulation ((ref)) in the analysis, instead of the motivational formulation ((ref)).
We shall provide the inference tools
that will be valid for inference on the series approximation
$$
\ell_\theta(w)' \beta, \ \ w \in \mathcal{I}.
$$
If the approximation error $r_\theta(w), \ w \in \mathcal{I},$ is small enough as compared to the estimation error,
these tools will also be valid for inference on the functional of interest
$$
\theta(w), \ \ w \in \mathcal{I}.
$$
In this case, the series approximation $\ell_\theta(w)$ is an important intermediary
target, whereas the functional $\theta(w)$ is the ultimate target. The
inference will be based on the plug-in estimator $\widehat
\theta(w) := \ell_\theta(w)' \widehat \beta$ of the the series
approximation $\ell_\theta(w)' \beta$ and hence of the final target
$\theta(w)$.
Approximation Properties of Least Squares
Next we consider approximation properties of the least squares estimator. Not surprisingly, approximation properties must rely on the particular choice of approximating functions. At this point it is instructive to consider particular examples of relevant bases used in the literature. For each example, we state a bound on the following quantity:
$$
\xi_k:= \sup_{x \in \mathcal{X}} \| p(x)\|.
$$
This quantity will play a key role in our analysis.\footnote{{ Most results extend directly to the case that $\xi_k\geq \max_{i\leq n} \| p(x_i)\|$ holds with probability $1-o(1)$. We refer to Hansen2014 for recent results that explicit allows for unbounded regressors which required extending the concentration inequalities for matrices.}}
Excellent reviews of approximating properties of different series can also be found in huang1998 and Chen2006, where additional references are provided.
example[Polynomial series] Let $\mathcal{X}=[0,1]$ and consider a polynomial series given by
$$
\widetilde p(x) = (1, x, x^2, ..., x^{k-1})'.
$$
In order to reduce collinearity problems, it is useful to orthonormalize the polynomial series with respect to the Lebesgue measure on $[0,1]$
to get the Legendre polynomial series
$$
p(x) = (1,\sqrt{3}x, \sqrt{5/4} (3 x^2 - 1),... )'.
$$
The Legendre polynomial series satisfies
$$
\xi_k \lesssim k;
$$
see, for example, Newey1997. \qed
example[Fourier series] Let $\mathcal{X} = [0,1]$ and consider a Fourier series given by
$$
p(x) = (1, \cos(2\pi j x), \sin(2\pi j x), j=1,2,..., (k-1)/2)',
$$
for $k$ odd. Fourier series is orthonormal with respect to the Lebesgue measure on $[0,1]$ and satisfies
$$
\xi_k \lesssim \sqrt{k},
$$
which follows trivially from the fact that every element of $p(x)$ is bounded in absolute value by one.\qed
example[Spline series] Let $\mathcal{X}=[0,1]$ and consider the linear regression spline series, or regression spline series of order 1, with a finite number of
equally spaced knots $l_1, \ldots, l_{k-2}$ in $\mathcal{X}$:
$$
\widetilde p(x) = (1, x, (x-l_1)_+ ,\dots, (x-l_{k-2})_+)',
$$
or consider the cubic regression spline series, or regression spline series of order 3, with a finite number of equally spaced knots $l_1,\dots,l_{k-4}$:
$$
\widetilde p(x) = ( 1, x, x^2, x^3, (x-l_1)^3_+,
..., (x-l_{k-4})^3_+)'.
$$
Similarly, one can define the regression spline series of any order $s_0$ (here $s_0$ is a nonnegative integer).
The function $x \mapsto \widetilde p(x)'b$ constructed using regression splines of order $s_0$ is
$s_0-1$ times continuously differentiable in $x$ for any $b$. Instead of regression splines,
it is often helpful to consider B-splines $p(x)=(p_1(x),\dots,p_k(x))'$, which are
linear transformations of the regression splines with lower multicollinearity; see DeBoor01 for the introduction to the theory of splines. B-splines are local in the sense that each B-spline $p_j(x)$ is supported on the interval $[l_{j(1)},l_{j(2)}]$ for some $j(1)$ and $j(2)$ satisfying $j(2)-j(1)\lesssim 1$ and there is at most $s_0+1$ non-zero B-splines on each interval $[l_{j-1},l_j]$. From this property of B-splines, it is easy to see that B-spline series satisfies
$$
\xi_k \lesssim \sqrt{k};
$$
see, for example, Newey1997. \qed
example[Cohen-Deubechies-Vial wavelet series]
Let $\mathcal{X} = [0,1]$ and consider Cohen-Deubechies-Vial (CDV) wavelet bases; see Section 4 in CDV93, Chapter 7.5 in M09, and Chapter 7 and Appendix B in Johnstone2011 for details on CDV wavelet bases.
CDV wavelet bases is a class of orthonormal with respect to the Lebesgue measure on $[0,1]$ bases.
Each such basis is built from a Daubechies scaling function $\phi$ (defined on $\Bbb{R}$) and the wavelet $\psi$ of order $s_0$ starting from a fixed resolution level $J_{0}$ such that $2^{J_{0}} \geq 2s_0$. The functions $\phi$ and $\psi$ are supported on $[0,2s_0-1]$ and $[-s_0+1,s_0]$, respectively. Translate $\phi$ so that it has the support $[-s_0+1,s_0]$.
Let
\begin{equation*}
\phi_{l,m} (x) = 2^{l/2}\phi (2^{l}x - m), \ \psi_{l,m}(x) = 2^{l/2}\psi (2^{l} x - m), \ l,m \geq 0.
\end{equation*}
Then we can create the CDV wavelet basis from these functions as follows.
Take all the functions $\phi_{J_{0},m}, \psi_{l,m}$, $l\geq J_0$, that are supported in the interior of $[0,1]$ (these are functions $\phi_{J_{0},m}$ with $m=s_0-1,\dots,2^{J_{0}}-s_0$ and $\psi_{l,m}$ with $m=s_0-1,\dots,2^{l}-s_0, l \geq J_{0}$). Denote these functions $\widetilde{\phi}_{J_0,m}$, $\widetilde{\psi}_{l,m}$. To this set of functions, add suitable boundary corrected functions $\widetilde{\phi}_{J_0,0},\ldots,\widetilde{\phi}_{J_0,s_0-2}$, $\widetilde{\phi}_{J_0,2^{J_0}-s_0+1},\ldots,\widetilde{\phi}_{J_0,2^{J_0}-1}$,
$\widetilde{\psi}_{l,0},\ldots,\widetilde{\psi}_{l,s_0-2}$,
$\widetilde{\psi}_{l,2^{J_0}-s_0+1},\ldots,\widetilde{\psi}_{l,2^{J_0}-1}$,
$l\geq J_0$, so that
$\{ \widetilde{\phi}_{J_{0},m} \}_{0\leq m<2^{J_0}} \cup \{ \widetilde{\psi}_{l,m} \}_{0 \leq m <2^{l}, l \geq J_{0}}$ forms an orthonormal basis of $L^{2}[0,1]$. Suppose that $k = 2^{J}$ for some $J > J_{0}$. Then the CDV series takes the form:
\begin{equation*}
p(x) = (\widetilde{\phi}_{J_{0},0}(x),\ldots,\widetilde{\phi}_{J_{0},2^{J_{0}}-1}(x),\widetilde{\psi}_{J_{0},0}(x),\ldots,\widetilde{\psi}_{J-1,2^{J-1}-1}(x))'.
\end{equation*}
This series satisfies
\begin{equation*}
\xi_{k} \lesssim \sqrt{k}.
\end{equation*}
This bound can be derived by the same argument as that for B-splines Kato2013. CDV wavelet bases is a flexible tool to approximate many different function classes. See, for example, Johnstone2011, Appendix B.
\qed
example[Local polynomial partition series]
Let $\mathcal{X}=[0,1]$ and define a local polynomial partition series as follows. Let $s_0$ be a nonnegative integer. Partition $\mathcal{X}$ as $0=l_0<l_1,\dots<l_{\widetilde{k}-1}<l_{\widetilde{k}}=1$ where $\widetilde{k}:=[k/(s_0+1)]+1$ where $[ a ]$ is the largest integer that is strictly smaller than $a$. For $j=1,\dots,\widetilde{k}$, define $\delta_j:[0,1]\to\{0,1\}$ by $\delta_j(x)=1$ if $x\in(l_{j-1},l_j]$ and $0$ otherwise. For $j=1,\dots,k$, define
$$
\widetilde{p}_j(x):=\delta_{[ j/(s_0+1)]+1}(x)x^{j-1-(s_0+1)[ j/(s_0+1)]}
$$
for all $x\in\mathcal{X}$. Finally, define the local polynomial partition series $p_1(\cdot),\dots,p_k(\cdot)$ of order $s_0$ as an orthonormalization of $\widetilde{p}_1(\cdot),\dots,\widetilde{p}_k(\cdot)$ with respect to the Lebesgue (or some other) measure on $\mathcal{X}$. The local polynomial partition series estimator was analyzed in detail in CF2013. Its properties are somewhat similar to those of local polynomial estimator of Stone1982. When the partition $l_0,\dots,l_{\widetilde{k}}$ satisfies $l_j-l_{j-1}\asymp 1/\widetilde{k}$, that is there exist constants $c,C>0$ independent of $n$ and such that $c/\widetilde{k}\leq l_j-l_{j-1}\leq C/\widetilde{k}$ for all $j=1,\dots,\widetilde{k}$, and the Lebesgue measure is used, the local polynomial partition series satisfies
$$
\xi_k\lesssim \sqrt{k}.
$$
This bound can be derived by the same argument as that for B-splines.\qed
example[Tensor Products] Generalizations to multiple covariates are straightforward
using tensor products of unidimensional series. Suppose that the basic regressors
are
$$
x_i= (x_{1i}, ..., x_{di})'.
$$
Then we can create $d$
series for each basic regressor. Then we take all
interactions of functions from these $d$ series, called tensor
products, and collect them into a vector of regressors
$p_i$. If each series for a basic regressor has $J$
terms, then the final regressor has dimension $$k
= J^d,$$ which explodes exponentially in the
dimension $d$. The bounds on $\xi_k$ in terms of $k$
remain the same as in one-dimensional case.\qed
Each basis described in Examples (ref)-(ref) has different approximation properties which also depend on the particular class of functions $\mathcal{G}$. The following assumption captures the essence of this dependence into two quantities.
Condition A.3 (Approximation) For each $n$ and $k$, there are finite constants $c_k$ and $\ell_k$ such that for each $f \in \mathcal{G}$,
$$
\|r_f\|_{F,2} := \sqrt{{
array[array omitted — 36 chars of source]
r_f^2(x) d F(x)}} \leq c_k
\ \ \
{\em and} \ \ \
\|r_f\|_{F,\infty} := \sup_{x \in \mathcal{X}} |r_f(x)| \leq \ell_k c_k.
$$
Here $r_f$ is defined by ((ref)) and ((ref)) with $g$ replaced by $f$.
We call $\ell_k$ the Lebesgue factor because of its relation to the Lebesgue constant defined in Section (ref) below.
Together $c_k$ and $\ell_k$ characterize the approximation properties of the underlying class of functions under $L^2(\mathcal{X},F)$ and uniform distances. Note that constants $c_k=c_k(\mathcal{G})$ and $\ell_k=\ell_k(\mathcal{G})$ are allowed to depend $n$ but we omit indexing by $n$ for simplicity of notation. Next we discuss primitive bounds on $c_k$ and $\ell_k$.
Bounds on $c_k$
In what follows, we call the case where $c_k \to 0$ as $k \to \infty$ the correctly specified case. In particular, if the series are formed from bases that span $\mathcal{G}$, then $c_k \to 0$ as $k \to \infty$. However, if series are formed from bases that do not span $\mathcal{G}$, then $c_k \not\to 0$ as $k \to \infty$. We call any case where $c_k \not \to 0$ the incorrectly specified (misspecified) case.
To give an example of the misspecified case, suppose that $d=2$, so that $x=(x_1,x_2)'$ and $g(x)=g(x_1,x_2)$. Further, suppose that the researcher mistakenly assumes that $g(x)$ is additively separable in $x_1$ and $x_2$: $g(x_1,x_2)=g_1(x_1)+g(x_2)$. Given this assumption, the researcher forms the vector of approximating functions $p(x_1,x_2)$ such that each component of this vector depends either on $x_1$ or $x_2$ but not on both; see Newey1997 and NPV99 for the description of nonparametric series estimators of separately additive models. Then note that if the true function $g(x_1,x_2)$ is not separately additive, linear combinations $p(x_1,x_2)'b$ will not be able to accurately approximate $g(x_1,x_2)$ for any $b$, so that $c_k$ does not converge to zero as $k\to\infty$. Since analysis of misspecified models plays an important role in econometrics, we include results both for correctly and incorrectly specified models.
To provide a bound on $c_k$, note that for any $f\in\mathcal{G}$,
$$
\inf_b \| f - p'b\|_{F,2}\leq \inf_{b} \| f - p'b\|_{F, \infty},
$$
so that it suffices to set $c_k$ such that $c_k\geq\sup_{f\in\mathcal{G}}\inf_{b} \| f - p'b\|_{F, \infty}$. Next, the bounds for $\inf_{b} \| f - p'b\|_{F, \infty}$ are readily available from the Approximation Theory; see DeVoreLorentz1993. A typical example is based on the concept of $s$-smooth classes, namely H\"{o}lder classes of smoothness order $s$, $\Sigma_s(\mathcal{X})$. For $s\in(0,1]$, the H\"{o}lder class of smoothness order $s$, $\Sigma_s(\mathcal{X})$, is defined as the set of all functions $f:\mathcal{X}\to\mathbb{R}$ such that for $C>0$,
$$
|f(x)-f(\widetilde x)|\leq C\Big(\sum_{j=1}^d(x_j-\widetilde{x}_j)^2\Big)^{s/2}
$$
for all $x=(x_1,\dots,x_d)'$ and $\widetilde x=(\widetilde{x}_1,\dots,\widetilde{x}_d)'$ in $\mathcal{X}$. The smallest $C$ satisfying this inequality defines a norm of $f$ in $\Sigma_s(\mathcal{X})$, which we denote by $\|f\|_s$. For $s>1$, $\Sigma_s(\mathcal{X})$ can be defined as follows. For a $d$-tuple $\alpha=(\alpha_1,\dots,\alpha_d)$ of nonnegative integers, let
$$
D^\alpha=\partial_{x_1}^{\alpha_1}\dots\partial_{x_d}^{\alpha_d}.
$$
Let $[s]$ denote the largest integer strictly smaller than $s$. Then $\Sigma_s(\mathcal{X})$ is defined as the set of all functions $f:\mathcal{X}\to\mathbb{R}$ such that $f$ is $[s]$ times continuously differentiable and for some $C>0$,
$$
|D^{\alpha}f(x)-D^{\alpha}f(\widetilde{x})|\leq C\Big(\sum_{j=1}^d(x_j-\widetilde{x}_j)^2\Big)^{(s-[s])/2}\text{ and }|D^\beta f(x)|\leq C
$$
for all $x=(x_1,\dots,x_d)'$ and $\widetilde{x}=(\widetilde{x}_1,\dots,\widetilde{x}_d)'$ in $\mathcal{X}$ and for all $d$-tuples $\alpha=(\alpha_1,\dots,\alpha_d)$ and $\beta=(\beta_1,\dots,\beta_d)$ of nonnegative integers satisfying $\alpha_1+\dots+\alpha_d=[s]$ and $\beta_1+\dots+\beta_d\leq[s]$. Again, the smallest $C$ satisfying these inequalities defines a norm of $f$ in $\Sigma_s(\mathcal{X})$, which we denote $\|f\|_s$.
If $\mathcal{G}$ is a set of functions $f$ in $\Sigma_s(\mathcal{X})$ such that $\|f\|_s$ is bounded from above uniformly over all $f\in\mathcal{G}$ (that is, $\mathcal{G}$ is contained in a ball in $\Sigma_s(\mathcal{X})$ of finite radius), then we can take
equation[equation omitted — 83 chars of source]
for the polynomial series and
$$
c_k \lesssim k^{-(s\wedge s_0)/d}
$$
for spline, CDV wavelet, and local polynomial partition series of order $s_0$. If in addition we assume that each element of $\mathcal{G}$ can be extended to a periodic function, then ((ref)) also holds for the Fourier series. See, for example, Newey1997 and Chen2006 for references.
Bounds on $\ell_k$
We say that a least squares approximation by a particular series for the function class
$\mathcal{G}$ is co-minimal if the Lebesgue factor $\ell_k$ is small in the sense
of being a slowly varying function in $k$. A simple bound on $\ell_k$, which is independent of $\mathcal{G}$, is established in the following proposition:
propositionIf $c_k$ is chosen so that $c_k\geq \sup_{f\in\mathcal{G}}\inf_b\|f-p'b\|_{F,\infty}$, then Condition A.3 holds with
$$
\ell_k\leq 1+\xi_k.
$$
The proof of this proposition is based on the ideas of Newey1997 and is provided in the Appendix.
The advantage of the bound established in this proposition is that it is universally applicable. It is, however, not sharp in many cases because $\xi_k$ satisfies
$$
\xi_k^2\geq E[\|p(x_i)\|^2]=E[p(x_i)'p(x_i)]=k
$$
so that $\xi_k \gtrsim \sqrt{k}$ in all cases. Much sharper bounds follow from Approximation Theory for some important cases. To apply these bounds, define the Lebesgue constant:
$$
\widetilde{\ell}_k:=\sup \left ( \frac{\|p'\beta_f\|_{F,\infty}}{\|f\|_{F,\infty} }: \|f\|_{F,\infty} \neq 0, f \in \bar{\mathcal{G}} \right),
$$
where $\bar{\mathcal{G}} = \mathcal{G} + \{p'b: b \in \Bbb{R}^k\} = \{ f + p'b: f \in \mathcal{G}, b \in \Bbb{R}^k\}$. The following proposition provides a bound on $\ell_k$ in terms of $\widetilde{\ell}_k$:
propositionIf $c_k$ is chosen so that $c_k\geq\sup_{f\in\mathcal{G}}\inf_b\|f-p'b\|_{F,\infty}$, then Condition A.3 holds with
$$
\ell_k=1+\widetilde{\ell}_k.
$$
Note that in all examples above, we provided $c_k$ such that $c_k\geq\sup_{f\in\mathcal{G}}\inf_b\|f-p'b\|_{F,\infty}$, and so the results of Propositions (ref) and (ref) apply in our examples.
We now provide bounds on $\widetilde{\ell}_k$.
example[Fourier series, continued] For Fourier series on $\mathcal{X}=[0,1]$, $F = U(0,1)$, and $\mathcal{G} \subset C(\mathcal{X})$
$$
\widetilde{\ell}_k \leq C_0 \log k + C_1,
$$
where here and below $C_0$ and $C_1$ are some universal constants; see Zygmund2002.\qed
example[Spline series, continued] For continuous B-spline series on $\mathcal{X}=[0,1]$, $F=U(0,1)$, and $\mathcal{G} \subset C(\mathcal{X})$
$$
\widetilde{\ell}_k \leq C_0,
$$
under approximately uniform placement of knots; see Huang2003b. In fact, the result of Huang states that $\widetilde{\ell}_k \leq C$ whenever $F$ has the pdf on $[0,1]$ bounded from above by $\bar{a}$ and below from zero by $\underline{a}$ where $C$ is a constant that depends only on $\underline{a}$ and $\bar{a}$.\qed
example[Wavelet series, continued] For continuous CDV wavelet series on $\mathcal{X}=[0,1]$, $F=U(0,1)$, and $\mathcal{G} \subset C(\mathcal{X})$
$$
\widetilde{\ell}_k \leq C_0.
$$
The proof of this result was recently obtained by CC2013 who extended the argument of Huang2003b for B-splines to cover wavelets. In fact, the result of Chen and Christensen also shows that $\widetilde{\ell}_k \leq C$ whenever $F$ has the pdf on $[0,1]$ bounded from above by $\bar{a}$ and below from zero by $\underline{a}$ where $C$ is a constant that depends only on $\underline{a}$ and $\bar{a}$.\qed
example[Local polynomial partition series, continued]
For local polynomial partition series on $\mathcal{X}$, $F=U(0,1)$, and $\mathcal{G}\subset C(\mathcal{X})$,
$$
\widetilde{\ell_k}\leq C_0.
$$
To prove this bound, note that first order conditions imply that for any $f\in\bar{\mathcal{G}}$,
$$
\beta_f=Q^{-1}E[p(x_1)f(x_1)]=E[p(x_1)f(x_1)].
$$
Hence, for any $x\in\mathcal{X}$,
$$
|p(x)'\beta_f|=|E[p(x)'p(x_1)f(x_1)]|\lesssim \|f\|_{F,\infty}
$$
where the last inequality follows by noting that the sum $p(x)'p(x_1)=\sum_{j=1}^kp_j(x)p_j(x_1)$ contains at most $s_0+1$ nonzero terms, all nonzero terms in the sum are bounded by $\xi_k^2\lesssim k$, and $p(x)'p(x_1)=0$ outside of a set with probability bounded from above by $1/k$ up to a constant. The bound follows. Moreover, the bound $\widetilde{\ell}_k \leq C$ continues to hold whenever $F$ has the pdf on $[0,1]$ bounded from above by $\bar{a}$ and below from zero by $\underline{a}$ where $C$ is a constant that depends only on $\underline{a}$ and $\bar{a}$.
\qed
example[Polynomial series, continued] For Chebyshev polynomials with $\mathcal{X}=[0,1]$, $d F(x)/dx = 1/\sqrt{1-x^2}$, and $\mathcal{G} \subset C(\mathcal{X})$
$$
\widetilde{\ell}_k \leq C_0 \log k + C_1.
$$
This bound follows from a trigonometric representation of Chebyshev polynomials (see, for example, DeVoreLorentz1993) and Example (ref).\qed
example[Legendre Polynomials] For Legendre polynomials that form an orthonormal basis on $\mathcal{X}=[0,1]$ with respect to $F=(0,1)$, and $\mathcal{G} = C(\mathcal{X})$
$$
\widetilde{\ell}_k \geq C_0 \xi_k = C_1k,
$$
for some constants $C_0, C_1 >0$. See, for example, DeVoreLorentz1993). This means that even though some series schemes generate well-behaved uniform approximations, others -- Legendre polynomials -- do not in general. However, the following example specifies “tailored" function classes, for which Legendre and other series methods do automatically provide uniformly well-behaved approximations. \qed
example[Tailored Function Classes] For each type of series approximations, it is possible to specify function classes for which the Lebesgue factors are constant or slowly varying with $k$. Specifically, consider a collection
$$
\mathcal{G}_k = \left \{x \mapsto f(x) = p(x)'b + r(x): \int r(x) p(x) d F(x) = 0, \| r \|_{F,\infty} \leq \ell_k \| r \|_{F,2}, \| r \|_{F,2} \leq c_k \right \},
$$
where $\ell_k \leq C$ or $\ell_k \leq C \log k$. This example captures the idea, that for each type of series functions there are function classes that are well-approximated by this type. For example, Legendre polynomials may have poor Lebesgue factors in general, but there are well-defined function classes, where Legendre polynomials have well-behaved Lebesgue factors. This explains why polynomial approximations, for example, using Legendre polynomials, are frequently employed in empirical work. We provide an empirically relevant example below, where polynomial approximation works just as well as a B-spline approximation. In economic examples, both polynomial approximations and B-spline approximations are well-motivated if we consider them as more flexible forms of well-known, well-motivated functional forms in economics (for example, as more flexible versions of the linear-quadratic Mincer equations, or the more flexible versions of translog demand and production functions). \qed
The following example illustrate the performance of the series estimator using different bases for a real data set.
example[Approximations of Conditional Expected Wage Function]
Here $g(x)$ is the mean of log wage ($y$)
conditional on education $$x \in \{8,9,10,11, 12, 13, 14, 16, 17,
18, 19, 20\}.$$ The function $g(x)$ is computed using population
data -- the 1990 Census data for the U.S. men of prime age; see ACF2006 for more details. So in this example, we know the true population function $g(x)$. We would like to know how well this
function is approximated when common approximation methods
are used to form the regressors. For simplicity we assume that $x_i$
is uniformly distributed (otherwise we can weigh by the frequency).
In population, least squares estimator solves the approximation problem: $\beta=\arg\min_b E[
\{g(x_i) - p_i'b\}^2] $ for $p_i =p(x_i)$, where we form $p(x)$ as (a)
linear spline (Figure 1, left) and (b) polynomial series (Figure 1, right), such that dimension of $p(x)$ is either $k=3$ or $k=8$. It is clear from these graphs that spline and polynomial series yield similar approximations.
\begin{figure}[vt]
\begin{tabular}{cc}
&
\end{tabular}
\caption{{Conditional expectation function (cef) of $\log$ wage given education (ed) in the 1990 Census data for the U.S. men of prime age and its least squares approximation by spline (left panel) and polynomial series (right panel). Solid line - conditional expectation function; dashed line - approximation by $k=3$ series terms; dash-dot line - approximation by $k=8$ series terms}}
\end{figure}
In the table below, we also present $L^2$ and $L^\infty$ norms of approximating errors:
\begin{center}
\begin{tabular}{ c | c c c c }
&spline $k=3$ & spline $k=8$ & Poly $k=3$ & Poly $k=8$
\\
\hline
$L^2$ Error & 0.12 & 0.08 & 0.12 & 0.05 \\
$L^{\infty}$ Error & 0.29 & 0.17 & 0.30 & 0.12 \\
\hline
\end{tabular}
\end{center}
We see from the table that in this example, the Lebesgue factor, which is defined as the ratio of $L^\infty$ to $L^2$ errors, of the polynomial approximations is comparable to the Lebesgue factor of the spline approximations.
\qed
Limit Theory
$L^{2}$ Limit Theory
After we have established the set-up, we proceed to derive our results. We start with a result on the $L^{2}$ rate of convergence. Recall that $\bar{\sigma}^2=\sup_{x\in\mathcal{X}}E[\epsilon_i^2|x_i=x]$. In the theorem below, we assume that $\overline{\sigma}^2\lesssim 1$. This is a mild regularity condition.
theorem[$L^{2}$ rate of convergence] Assume that Conditions A.1-A.3 are satisfied. In addition, assume that $\xi_k^2\log k/n\to 0$ and $\overline{\sigma}^2\lesssim 1$. Then
under $c_k \to 0$,
\begin{equation}
\| \widehat g - g\|_{F,2} \lesssim_P \sqrt{k/n} + c_k,
\end{equation}
and under $c_k \not\to 0$,
\begin{equation}
\|\widehat g - p'\beta\|_{F,2}\lesssim_P \sqrt{k/n}+(\ell_kc_k\sqrt{k/n})\wedge(\xi_kc_k/\sqrt{n}),
\end{equation}
remark(i) This is our first main result in this paper. The condition $\xi_k^2 \log k /n \to 0$, which we impose, weakens (hence generalizes) the conditions imposed in Newey1997 who required $k\xi_k^2/n \to 0$. For series satisfying $\xi_k \lesssim \sqrt{k}$, the condition $\xi_k^2 \log k /n \to 0$ amounts to
\begin{equation}
k \log k/n \to 0.
\end{equation}
This condition is the same as that imposed in Stone1994, Huang2003, and recently by CF2013 but the result ((ref)) is obtained under the condition ((ref)) in Stone1994 and Huang2003 only for spline series and in CF2013 only for local polynomial partition series.
Therefore, our result improves on those in the literature by weakening the rate requirements on the growth of $k$ (with respect to $n$) and/or by allowing for a wider set of series functions.
(ii) Under the correct specification ($c_k\to 0$), the fastest $L^2$ rate of convergence is achieved by setting $k$ so that the approximation error and the sampling error are of the same order,
$$
\sqrt{k/n} \asymp c_k.
$$
One consequence of this result is that for H\"{o}lder classes of smoothness order $s$, $\Sigma_s(\mathcal{X})$, with $c_k\lesssim k^{-s/d}$, we obtain the optimal $L^2$ rate of convergence by setting $k\asymp n^{d/(d+2s)}$, which is allowed under our conditions for all $s>0$ if $\xi_k\lesssim \sqrt{k}$ (Fourier, spline, wavelet, and local polynomial partition series). On the other hand, if $\xi_k$ is growing faster than $\sqrt{k}$, then it is not possible to achieve optimal $L^2$ rate of convergence for some $s>0$. For example, for polynomial series considered above, $\xi_k\lesssim k$, and so the condition $\xi_k^2\log k/n \to 0$ becomes $k^2\log k/n \to 0$. Hence, optimal $L^2$ rate of convergence is achieved by polynomial series only if $d/(d+2s)<1/2$ or, equivalently, $s>d/2$. Even though this condition is somewhat restrictive, it weakens the condition in Newey1997 who required $k^3/n\to 0$ for polynomial series, so that optimal $L^2$ rate in his analysis could be achieved only if $d/(d+2s)\leq 1/3$ or, equivalently, $s\geq d$. Therefore, our results allow to achieve optimal $L^2$ rate of convergence in a larger set of classes of functions for particular series.
(iii) The result ((ref)) is concerned with the case when the model is misspecified ($c_k\not \to 0$). It shows that when $k/n\to 0$ and $(\ell_kc_k\sqrt{k/n})\wedge(\xi_kc_k/\sqrt{n})\to 0$, the estimator $\widehat{g}(\cdot)$ converges in $L^2$ to the surrogate function $p(\cdot)'\beta$ that provides the best linear approximation to the target function $g(\cdot)$. In this case, the estimator $\widehat{g}(\cdot)$ does not generally converge in $L^2$ to the target function $g(\cdot)$.
\qed
Pointwise Limit Theory
Next we focus on pointwise limit theory (some authors refer to pointwise limit theory as local asymptotics; see Huang2003b). That is, we study asymptotic behavior of $\sqrt{n}\alpha'(\widehat{\beta}-\beta)$ and $\sqrt{n}(\widehat{g}(x)-g(x))$ for particular $\alpha\in S^{k-1}$ and $x\in\mathcal{X}$. Here $S^{k-1}$ denotes the space of vectors $\alpha$ in $\mathbb{R}^k$ with unit Euclidean norm: $\|\alpha\|=1$. Note that both $\alpha$ and $x$ implicitly depend on $n$. As we will show, pointwise results can be achieved under weak conditions similar to those we required in Theorem (ref). The following lemma plays a key role in our asymptotic pointwise normality result.
lemma[Pointwise Linearization] Assume that Conditions A.1-A.3 are satisfied. In addition, assume that $\xi_k^2\log k/n\to 0$ and $\overline{\sigma}^2\lesssim 1$. Then for any $\alpha \in S^{k-1}$,
\begin{equation}
\sqrt{n} \alpha'( \widehat \beta - \beta) = \alpha' \mathbb{G}_n[ p_i (\epsilon_i + r_i)] + R_{1n}(\alpha),
\end{equation}
where the term $R_{1n}(\alpha)$, summarizing the impact of unknown design, obeys
\begin{equation}
R_{1n}(\alpha) \lesssim_P \sqrt{\frac{ \xi_k^2 \log k }{ n}}(1 + \sqrt{k} \ell_k c_k).
\end{equation}
Moreover,
\begin{equation}
\sqrt{n} \alpha'( \widehat \beta - \beta) = \alpha' \mathbb{G}_n[ p_i \epsilon_i] + R_{1n}(\alpha) + R_{2n}(\alpha),
\end{equation}
where the term $R_{2n}(\alpha)$, summarizing the impact of approximation error on the sampling error of the estimator, obeys
\begin{equation}
R_{2n}(\alpha) \lesssim_P \ell_k c_k.
\end{equation}
remark(i) In summary, the only condition that generally matters for linearization ((ref))-((ref)) is that $R_{1n}(\alpha) \to 0$, which holds if $\xi_k^2\log k/n\to 0$ and $k\xi_k^2\ell_k^2c_k^2\log k/n\to 0$. In particular, linearization ((ref))-((ref)) allows for misspecification ($c_k\to 0$ is not required).
In principle, linearization ((ref))-((ref)) also allows for misspecification but the bounds are only useful if the model is correctly specified, so that $\ell_k c_k \to 0$. As in the theorem on $L^2$ rate of convergence, our main condition is that $\xi_k^2\log k/n \to 0$.
(ii) We conjecture that the bound on $R_{1n}(\alpha)$ can be improved for splines to
\begin{equation}
R_{1n}(\alpha) \lesssim_P \sqrt{\frac{ \xi_k^2 \log k }{ n}}(1 + \sqrt{\log k} \cdot \ell_k c_k).
\end{equation}
since it is attained by local polynomials and splines are also similarly localized.\qed
With the help of Lemma (ref), we derive our asymptotic pointwise normality result. We will use the following additional notation:
$$\widetilde \Omega:= Q^{-1} E [ (\epsilon_i + r_i)^2 p_i p_i' ]Q^{-1}\text{ and }\Omega_0 := Q^{-1}E [\epsilon_i^2 p_i p_i' ]Q^{-1}.
$$
In the theorem below, we will impose the condition that $\sup_{x \in \mathcal{X}} E\left[\epsilon_i^2 1\{ |\epsilon_i| > M \}|x_i =x\right] \to 0$ as $M\to \infty$ uniformly over $n$. This is a mild uniform integrability condition. Specifically, it holds if for some $m>2$, $\sup_{x \in \mathcal{X}} E[|\epsilon_{i}|^{m} | x_{i} = x] \lesssim 1$. In addition, we will impose the condition that $1\lesssim \underline{\sigma}^2$. This condition is used to properly normalize the estimator.
theorem[Pointwise Normality] Assume that Conditions A.1-A.3 are satisfied. In addition, assume that (i) $\sup_{x \in \mathcal{X}} E\left[\epsilon_i^2 1\{ |\epsilon_i| > M \}|x_i =x\right] \to 0$ as $M\to \infty$ uniformly over $n$, (ii) $1\lesssim\underline{\sigma}^2$, and (iii) $(\xi_k^2\log k/n)^{1/2}(1+k^{1/2}\ell_kc_k)\to 0$. Then for any $\alpha \in S^{k-1}$,
\begin{equation}
\sqrt{n} \frac{\alpha'(\widehat \beta - \beta)}{\|\alpha' \Omega^{1/2}\|} =_d N(0,1) + o_P(1),
\end{equation}
where we set $\Omega =\widetilde \Omega$
but if $R_{2n}(\alpha) \to_P 0$, then we can set
$\Omega = \Omega_0$.
Moreover, for any $x \in \mathcal{X}$ and $s(x):= \Omega^{1/2}p(x)$,
\begin{equation}
\sqrt{n} \frac{p(x)'(\widehat \beta - \beta)}{\|s(x)\|} =_d N(0,1) + o_P(1),
\end{equation}
and if the approximation error is negligible relative to the estimation error, namely $\sqrt{n} r(x)=o(\|s(x)\|)$, then
\begin{equation}
\sqrt{n} \frac{\widehat g(x) - g(x)}{\|s(x)\|} =_d N(0,1) + o_P(1).
\end{equation}
remark(i) This is our second main result in this paper. The result delivers pointwise convergence in distribution for any sequences $\alpha=\alpha_n$ and $x=x_n$ with $\alpha\in S^{k-1}$ and $x\in\mathcal{X}$. In fact, the proof of the theorem implies that the convergence is uniform over all sequences. Note that the normalization
factor $\|s(x)\|$ is the pointwise standard error, and it is of a typical order
$ \|s(x)\| \propto \sqrt{k} $ at most points.
In this case the condition for negligibility of approximation error $\sqrt{n} r(x)/\|s(x)\|\to 0$, which can be understood as an undersmoothing condition, can be replaced by $$\sqrt{n/k}\cdot \ell_k c_k\to 0.$$
When $\ell_k c_k\lesssim k^{-s/d}\log k$, which is often the case if $\mathcal{G}$ is contained in a ball in $\Sigma_s(\mathcal{X})$ of finite radius (see our examples in the previous section), this condition substantially weakens an assumption in Newey1997 who required $\sqrt{n}k^{-s/d} \to 0$ in a similar set-up.
(ii) When applied to splines, our result is somewhat less sharp than that of Huang2003b. Specifically, Huang required that $\xi_k^2\log k/n \to 0$ and $(n/k)^{1/2}\cdot\ell_k c_k\to 0$ whereas we require $(k\xi_k^2\log k/n)^{1/2}\ell_k c_k\to 0$ in addition to Huang's conditions (see condition (iii) of the theorem). The difference can likely be explained by the fact that we use linearization bound ((ref)) whereas for splines it is likely that ((ref)) holds as well.
(iii) More generally, our asymptotic pointwise normality result, as well as other related results in this paper, applies to any problem
where the estimator of $g(x) = p(x)'\beta + r(x)$ takes the form $p(x)'\widehat \beta$, where $\widehat \beta$ admits linearization
of the form ((ref))-((ref)). \qed
Uniform Limit Theory
Finally, we turn to a uniform limit theory. Not surprising, stronger conditions are required for our results to hold when compared to the pointwise case. Let $m>2$. We will need the following assumption on the tails of the regression errors.
Condition A.4 (Disturbances) Regression errors satisfy $\sup_{x \in \mathcal{X}} E[|\epsilon_{i}|^{m} | x_{i} = x] \lesssim 1$.
It will be convenient to denote $\alpha(x):=p(x)/\|p(x)\|$ in this subsection. Moreover, denote
$$
\xi_k^L:=\sup_{x,x'\in\mathcal{X}: \, x\neq x'}\frac{\|\alpha(x)-\alpha(x')\|}{\|x - x' \|}
$$
We will also need the following assumption on the basis functions to hold with the same $m>2$ as that in Condition A.4.
Condition A.5 (Basis) Basis functions are such that (i) $\xi_{k}^{2m/(m-2)} \log k/n \lesssim 1$, (ii) $\log \xi^L_k\lesssim \log k$, and (iii) $\log \xi_k\lesssim \log k$.
The following lemma provides uniform linearization of the series estimator and plays a key role in our derivation of the uniform rate of convergence.
lemma[Uniform Linearization] Assume that Conditions A.1-A.5 are satisfied. Then
\begin{equation}
\sqrt{n} \alpha(x)'( \widehat \beta - \beta) = \alpha(x)' \mathbb{G}_n[ p_i (\epsilon_i + r_i)] + R_{1n}(\alpha(x)),
\end{equation}
where $R_{1n}(\alpha(x))$, summarizing the impact of unknown design, obeys
\begin{equation}
R_{1n}(\alpha(x)) \lesssim_P \sqrt{\frac{ \xi_k^2 \log k }{n}} (n^{1/m} \sqrt{\log k} + \sqrt{k} \cdot \ell_kc_{k})=:\bar{R}_{1n}
\end{equation}
uniformly over $x \in \mathcal{X}$.
Moreover,
\begin{equation}
\sqrt{n} \alpha(x)'( \widehat \beta - \beta) = \alpha(x)' \mathbb{G}_n[ p_i \epsilon_i ] + R_{1n}(\alpha(x)) +R_{2n}(\alpha(x)),
\end{equation}
where $R_{2n}(\alpha(x))$, summarizing the impact of approximation error on the sampling error of the estimator, obeys
\begin{equation}
R_{2n}(\alpha(x)) \lesssim_P \sqrt{\log{k}} \cdot \ell_kc_{k}=:\bar{R}_{2n}
\end{equation}
uniformly over $x\in\mathcal{X}$.
remarkAs in the case of pointwise linearization, our results on uniform linearization ((ref))-((ref)) allow for misspecification ($c_k\to 0$ is not required). In principle, linearization ((ref))-((ref)) also allows for misspecification but the bounds are most useful if the model is correctly specified so that $(\log k)^{1/2}\ell_k c_k \to 0$. We are not aware of any similar uniform linearization result in the literature. We believe that this result is useful in a variety of problems. Below we use this result to derive good uniform rate of convergence of the series estimator. Another application of this result would be in testing shape restrictions in the nonparametric model. \qed
The following theorem provides uniform rate of convergence of the series estimator.
theorem[Uniform Rate of Convergence]
Assume that Conditions A.1-A.5 are satisfied. Then
\begin{equation}
\sup_{x \in \mathcal{X}} |\alpha(x)' \mathbb{G}_n[ p_i \epsilon_i]| \lesssim_P \sqrt{\log k}.
\end{equation}
Moreover, for $\bar{R}_{1n}$ and $\bar{R}_{2n}$ given above we have
\begin{equation}
\sup_{x \in \mathcal{X}} |p(x)'( \widehat \beta - \beta)| \lesssim_P \frac{\xi_k}{\sqrt{n}}( \sqrt{\log k} + \bar{R}_{1n} + \bar{R}_{2n})
\end{equation}
and
\begin{equation}
\sup_{x \in \mathcal{X}} |\widehat g(x) - g(x)| \lesssim_P \frac{\xi_k}{\sqrt{n}}( \sqrt{\log k} + \bar{R}_{1n} + \bar{R}_{2n}) + \ell_kc_{k}.
\end{equation}
remarkThis is our third main result in this paper. Assume that $\mathcal{G}$ is a ball in $\Sigma_s(\mathcal{X})$ of finite radius, $\ell_k c_k\lesssim k^{-s/d}$, $\xi_k\lesssim \sqrt{k}$, and $\bar{R}_{1n}+\bar{R}_{2n}\lesssim (\log k)^{1/2}$. Then the bound in ((ref)) becomes
$$
\sup_{x \in \mathcal{X}} |\widehat g(x) - g(x)| \lesssim_P \sqrt{\frac{k\log k}{n}}+k^{-s/d}.
$$
Therefore, setting $k\asymp(\log n/n)^{-d/(2s+d)}$, we obtain
$$
\sup_{x \in \mathcal{X}} |\widehat g(x) - g(x)| \lesssim_P \left(\frac{\log n}{n}\right)^{s/(2s+d)},
$$
which is the optimal uniform rate of convergence in the function class $\Sigma_s(\mathcal{X})$; see Stone1982. To the best of our knowledge, our paper is the first to show that the series estimator attains the optimal uniform rate of convergence under these rather general conditions; see the next comment. We also note here that it has been known for a long time that a local polynomial (kernel) estimator achieves the same optimal uniform rate of convergence; see, for example, Tsybakov2003, and it was also shown recently by CF2013 that local polynomial partition series estimator also achieves the same rate. { Recently, in an effort to relax the independence assumption, the working paper CC2013, which appeared in ArXiv in 2013, approximately 1 year after our paper was posted to ArXiv and submitted for publication, \footnote{Our paper was submitted for publication and to ArXiv on December 3, 2012. Our result as stated here did not change since the original submission.} derived similar uniform rate of convergence result allowing for $\beta$-mixing conditions, see their Theorem 4.1 for specific conditions.}\qed
remarkPrimitive conditions leading to inequalities $\ell_k c_k\lesssim k^{-s/d}$ and $\xi_k\lesssim \sqrt{k}$ are discussed in the previous section. Also, under the assumption that $\ell_k c_k\lesssim k^{-s/d}$, inequality $\bar{R}_{2n}\lesssim (\log k)^{1/2}$ follows automatically from the definition of $\bar{R}_{2n}$.
Thus, one of the critical conditions to attain the optimal uniform rate of convergence is that we require $\bar{R}_{1n}\lesssim (\log k)^{1/2}$. Under our other assumptions, this condition holds if $k\log k/n^{1-2/m}\lesssim 1$ and $k^{2-2s/d}/n\lesssim 1$, and so we can set $k\asymp(\log n/n)^{-d/(2s+d)}$ if $d/(2s+d)<1-2/m$ and $(2d-2s)/(2s+d)<1$ or, equivalently, $m>2+d/s$ and $s/d>1/4$.
\qed
After establishing the auxiliary results on the uniform rate of convergence, we present two results on inference based on the series estimator.
The first result on inference is concerned with the strong approximation of a series process by a Gaussian process and is a (relatively) minor extension of the result obtained by CLR2008. The extension is undertaken to allow for a non-vanishing specification error to cover misspecified models. In particular, we make a distinction between $\widetilde \Omega= Q^{-1} E [ (\epsilon_i + r_i)^2 p_i p_i' ]Q^{-1},$
and $\Omega_0 = Q^{-1}E [\epsilon_i^2 p_i p_i' ]Q^{-1}$ which are potentially asymptotically different if $\bar{R}_{2n} \not\to_P 0$.
To state the result, let $a_n$ be some sequence of positive numbers satisfying $a_n \to \infty$.
theorem[Strong Approximation by a Gaussian Process]
Assume that Conditions A.1-A.5 are satisfied with $m\geq 3$. In addition, assume that (i) $\bar{R}_{1n} = o_P(a_n^{-1})$, (ii) $1\lesssim \underline{\sigma}^2$,
and (iii)
$a_n^6 k^4 \xi^2_k (1 + \ell_k^3 c_k^3)^2 \log^2 n/n \to 0.$
Then for some $\mathcal{N}_k \sim N(0, I_k)$,
\begin{equation}
\sqrt{n} \frac{\alpha(x)'(\widehat \beta - \beta)}{\|\alpha(x)' \Omega^{1/2}\|} =_d \frac{\alpha(x)' \Omega^{1/2}}{\|\alpha(x)' \Omega^{1/2}\|} \mathcal{N}_k + o_P(a_n^{-1}) in \ell^{\infty}(\mathcal{X}),
\end{equation}
so that for $s(x)= \Omega^{1/2}p(x)$,
\begin{equation}
\sqrt{n} \frac{p(x)'(\widehat \beta - \beta)}{\|s(x)\|} =_d \frac{s(x)'}{\|s(x)\|} \mathcal{N}_k + o_P(a_n^{-1}) in \ell^{\infty}(\mathcal{X}),
\end{equation}
and if $\sup_{x \in \mathcal{X}}\sqrt{n} |r(x)|/\|s(x)\| =o(a_n^{-1}) $, then
\begin{equation}
\sqrt{n} \frac{\widehat g(x) - g(x)}{\|s(x)\|} =_d \frac{s(x)'}{\|s(x)\|} \mathcal{N}_k + o_P(a_n^{-1}) in \ell^{\infty}(\mathcal{X}),
\end{equation}
where we set $\Omega =\widetilde \Omega$ but if $\bar{R}_{2n} = o_P(a_n^{-1})$, then we can set
$\Omega = \Omega_0$.
remarkOne might hope to have a result of the form
\begin{equation}
\sqrt{n} \frac{\widehat g(x) - g(x)}{\|s(x)\|}\to_d G(x) in \ell^{\infty}(\mathcal{X}),
\end{equation}
where $\{G(x):x\in\mathcal{X}\}$ is some fixed zero-mean Gaussian process. However, one can show that the process on the left-hand side of ((ref)) is not asymptotically equicontinuous, and so it does not have a limit distribution.
Instead, Theorem (ref) provides an approximation of the series process by a sequence of zero-mean Gaussian processes $\{G_k(x):x\in\mathcal{X}\}$
$$
G_k(x):= \frac{\alpha(x)' \Omega^{1/2}}{\|\alpha(x)' \Omega^{1/2}\|} \mathcal{N}_k ,
$$
with the stochastic error of size $o_P(a^{-1}_n)$. Since $a_n\to\infty$, under our conditions the theorem implies that the series process is well approximated by a Gaussian process, and so the theorem can be interpreted as saying that in large samples, the distribution of the series process depends on the distribution of the data only via covariance matrix $\Omega$; hence, it allows us to perform inference based on the whole series process. Note that the conditions of the theorem are quite strong in terms of growth requirements on $k$, but the result of the theorem is also much stronger than the pointwise normality result: it asserts that the entire series process is uniformly close to a Gaussian process of the stated form.
\qed
Our result on the strong approximation by a Gaussian process plays an important role in our second result on inference that is concerned with the weighted bootstrap. Consider a set of weights
$h_1,\ldots,h_n$ that are i.i.d. draws from the standard exponential
distribution and are independent of the data. For each draw of such weights, define the weighted
bootstrap draw of the least squares estimator as a solution to the least squares problem
weighted by $h_1,\ldots,h_n$, namely
equation[equation omitted — 139 chars of source]
For all $x\in\mathcal{X}$, denote $\widehat{g}^b(x)=p(x)'\widehat{\beta}^b$. The following theorem establishes a new result that states that the weighted bootstrap
distribution is valid for approximating the distribution of the series process.
theorem[Weighted Bootstrap Method]
(1) Assume that Conditions A.1-A.5 are satisfied. In addition, assume that $(\xi_k(\log n)^{1/2})^{2m/(m-2)}\lesssim 1$. Then the weighted bootstrap process satisfies
$$
\sqrt{n} \alpha(x)'( \widehat \beta^b - \widehat\beta) = \alpha(x)' \mathbb{G}_n[ (h_i-1)p_i (\epsilon_i + r_i)] + R_{1n}^b(\alpha(x)),
$$
where $R_{1n}^b(\alpha(x))$ obeys
\begin{equation}
R_{1n}^b(\alpha(x)) \lesssim_P \sqrt{\frac{\xi_k^2\log^3 n}{n}}(n^{1/m}\sqrt{\log n} + \sqrt{k}\cdot \ell_kc_k) =:\bar{R}_{1n}^b
\end{equation}
uniformly over $x\in\mathcal{X}$.
(2) If, in addition, Conditions A.4 and A.5 are satisfied with $m\geq 3$ and (i) $\bar{R}_{1n}^b=o_P(a_n^{-1})$, (ii) $1\lesssim \underline{\sigma}^2$, and (iii) $a_n^6 k^4 \xi^2_k (1 + \ell_k^3 c_k^3)^2 \log^2 n/n \to 0$ hold, then for $s(x)= \Omega^{1/2}p(x)$ and some $\mathcal{N}_k\sim N(0,I_k)$,
\begin{equation}
\sqrt{n} \frac{p(x)'(\widehat \beta^b - \widehat\beta)}{\|s(x)\|} =_d \frac{s(x)'}{\|s(x)\|} \mathcal{N}_k + o_P(a_n^{-1}) in \ell^{\infty}(\mathcal{X}),
\end{equation}
and so
\begin{equation}
\sqrt{n} \frac{\widehat g^b(x) - \widehat{g}(x)}{\|s(x)\|} =_d \frac{s(x)'}{\|s(x)\|} \mathcal{N}_k + o_P(a_n^{-1}) in \ell^{\infty}(\mathcal{X}).
\end{equation}
where we set $\Omega =\widetilde \Omega$, but if $\bar{R}_{2n} = o_P(a_n^{-1})$, then we can set
$\Omega = \Omega_0.$
(3) Moreover, the bounds ((ref)), ((ref)), and ((ref)) continue to hold in $P$-probability if we replace the
unconditional probability $P$ by the conditional probability computed given the data, namely if we replace $P$ by $P^*(\cdot \mid D)$ where $D=\{(x_i,y_i):i=1,\dots,n\}$.
remark(i) This is our fourth main and new result in this paper. The theorem implies that the weighted bootstrap process can be approximated by a copy of the same Gaussian process as that used to approximate original series process.
(ii) We emphasize that the theorem does not require the correct specification, that is the case $c_k\not\to 0$ is allowed.
Also, in this theorem, symbol $P$ refers to a joint probability measure with respect to the data $D=\{(x_i,y_i):i=1,\dots,n\}$ and the set of bootstrap weights $\{h_i:i=1,\dots,n\}$.\qed
We close this section by establishing sufficient conditions for consistent estimation of $\Omega$. Recall that $Q=E[p_ip_i']=I$. In addition, denote $\Sigma = E[(\epsilon_i+r_i)^2p_ip_i']$, $\widehat Q = \mathbb{E}_n[p_ip_i']$, and $\widehat \Sigma = \mathbb{E}_n[\widehat \epsilon_i^2 p_ip_i']$ where $\widehat \epsilon_i = y_i-p_i'\widehat \beta$, and let $v_n =(E[\max_{1\leq i \leq n} | \epsilon_i|^2])^{1/2}$.
theorem[Matrices Estimation]
Assume that Conditions A.1-A.5 are satisfied. In addition, assume that $\bar{R}_{1n}+\bar{R}_{2n}\lesssim(\log k)^{1/2}$. Then
$$
\|\widehat Q - Q\| \lesssim_P \sqrt{\frac{\xi_k^2\log k}{n}}=o(1) \ \ \mbox{and} \ \ \|\widehat \Sigma - \Sigma\|\lesssim_P (v_n \vee 1+\ell_kc_k) \sqrt{\frac{\xi_k^2\log k}{n}}=o(1).
$$
Moreover, for $\widehat \Omega = \widehat Q^{-1}\widehat \Sigma\widehat Q^{-1}$ and $\Omega=Q^{-1}\Sigma Q^{-1}$,
$$
\|\widehat{\Omega} - \Omega\|\lesssim_P (v_n \vee 1+\ell_kc_k) \sqrt{\frac{\xi_k^2\log k}{n}}=o(1).
$$
remarkTheorem (ref) allows for consistent estimation of the matrix $Q$ under the mild condition $\xi_k^2\log k/n \to 0$ and for consistent estimation of the matrices $\Sigma$ and $\Omega$ under somewhat more restricted conditions. Not surprisingly, the estimation of $\Sigma$ and $\Omega$ depends on the tail behavior of the error term via the value of $v_n$.
Note that under Condition A.4, we have that $v_n \lesssim n^{1/m}$. \qed
Rates and Inference on Linear Functionals
In this section, we derive rates and inference results for linear functionals $\theta(w), w\in \mathcal{I}$ of the conditional expectation function such as its derivative, average derivative, or conditional average derivative. To a large extent, with the exception of Theorem (ref), the results presented in this section can be considered as an extension of results presented in Section (ref), and so similar comments can be applied as those given in Section (ref). Theorem (ref) deals with construction of uniform confidence bands for linear functionals under weak conditions and is a new result.
By the linearity of the series approximations, the linear functionals can be seen as linear functions of the least squares coefficients $\beta$ up to an approximation error, that is
$$
\theta(w) = \ell_\theta(w)' \beta + r_\theta(w), \ \ w \in \mathcal{I},
$$
where $\ell_\theta(w)' \beta$ is the series approximation, with
$\ell_\theta(w)$ denoting the $k$-vector of loadings on the coefficients,
and $r_\theta(w)$ is the remainder term, which corresponds to the
approximation error.
Throughout this section, we assume that $\mathcal{I}$ is a subset of some Euclidean space $\mathbb{R}^l$ equipped with its usual norm $\|\cdot\|$. We allow $\mathcal{I}=\mathcal{I}_n$ to depend on $n$ but for simplicity, we assume that the diameter of $\mathcal{I}$ is bounded from above uniformly over $n$. Results allowing for the case where $\mathcal{I}$ is expanding as $n$ grows can be covered as well with slightly more technicalities.
In order to perform inference, we construct estimators of
$ \sigma_\theta^2(w) = \ell_\theta(w)'\Omega\ell_\theta(w)/n$, the
variance of the associated linear functionals, as
equation[equation omitted — 119 chars of source]
In what follows, it will be convenient to have the following result on consistency of $\widehat{\sigma}_\theta(w)$:
lemma[Variance Estimation for Linear Functionals]
Assume that Conditions A.1-A.5 are satisfied. In addition, assume that (i) $\bar{R}_{1n}+\bar{R}_{2n}\lesssim(\log k)^{1/2}$ and (ii) $1\lesssim \underline{\sigma}^2$. Then
$$
\left|\frac{\widehat{\sigma}_\theta(w)}{\sigma_\theta(w)}-1\right|\lesssim_P \|\widehat{\Omega}-\Omega\|\lesssim_P (v_n \vee 1+\ell_kc_k) \sqrt{\frac{\xi_k^2\log k}{n}}=o(1)
$$
uniformly over $w\in\mathcal{I}$.
By Lemma (ref), under our
conditions, ((ref)) is uniformly consistent for
$\sigma_\theta^2(w)$ in the sense that $\widehat
\sigma_\theta^2(w)/ \sigma_\theta^2(w) = 1 + o_P(1)$ uniformly over $w\in \mathcal{I}$.
Pointwise Limit Theory for Linear Functionals
We now present a result on pointwise rate of convergence for linear functionals. The rate we derive is $\|\ell_\theta(w)\|/\sqrt{n}$. Some examples with explicit bounds on $\|\ell_\theta(w)\|$ are given below.
theorem[Pointwise Rate of Convergence for Linear Functionals]
Assume that Conditions A.1-A.3 are satisfied. In addition, assume that (i) $\sqrt{n}|r_\theta(w)|/\|\ell_\theta(w)\| \to 0$, (ii) $\bar{\sigma}^2\lesssim 1$, (iii) $(\xi_k^2\log k/n)^{1/2}(1+k^{1/2}\ell_kc_k) \to 0$, and (iv) $\ell_kc_k\to 0$. Then
$$
| \widehat \theta(w) - \theta(w)| \lesssim_P \frac{\|\ell_{\theta}(w)\|}{\sqrt{n}}.
$$
remark(i) This theorem shows in particular that $\widehat{\theta}(w)$ is $\sqrt{n}$-consistent whenever $\|\ell_{\theta}(w)\|\lesssim 1$. A simple example of this case is $\theta=\theta(w)=E[g(x_1)]$. In this example, $\ell=\ell(w)=E[p(x_1)]$, and so $\|\ell\|=\|E[p(x_1)]\|\lesssim 1$ where the last inequality follows from the argument used in the proof of Proposition (ref). Another simple example is $\theta=\theta(w)=E[p(x_1)g(x_1)]=\beta_1$. In this example, $\ell=\ell(w)$ is a $k$-vector whose first component is 1 and all other components are 0, and so $\|\ell\|\lesssim 1$. This example trivially implies $\sqrt{n}$-consistency of the series estimator of the linear part of the partially linear model. Yet another example, which is discussed in Newey1997, is the average partial derivative.
(ii) Condition $\sqrt{n}|r_{\theta}(w)|/\|\ell_\theta(w)\|\to 0$ imposed in this theorem can be understood as undersmoothing condition. Unfortunately, to the best of our knowledge, there is no theoretically justified practical procedure in the literature that would lead to a desired level of undersmoothing. Some ad hoc suggestions include using cross validation or “plug-in” method to determine the number of series terms that would minimize the asymptotic integrated mean-square error of the series estimator Hardle1990 and then blow up the estimated number of series terms by some number that grows to infinity as the sample size increases.\qed
To perform pointwise inference, we consider the t-statistic:
$$
t(w) = \frac{ \widehat \theta(w) - \theta(w) }{ \widehat \sigma_\theta(w)}.
$$
We can carry out standard inference based on this statistic because of the following theorem.
theorem[Pointwise Inference for Linear Functionals]
Assume that the conditions of Theorem (ref)
and Lemma (ref) are satisfied. In addition, assume that $\sqrt{n}|r_\theta(w)|/\|\ell_{\theta}(w)\|\to 0$. Then
$$
t(w) \to_d N(0,1).
$$
The same comments apply here as those given in Section (ref) for pointwise results on estimating the function $g$ itself.
Uniform Limit Theory for Linear Functionals
In obtaining uniform rates of convergence and inference results for linear functionals, we will denote
$$
\xi_{k,\theta}:=\sup_{w\in\mathcal{I}}\|\ell_{\theta}(w)\|\,\,\text{ and }\,\,\xi_{k,\theta}^L:=\sup_{w,w'\in\mathcal{I}: \, w\neq w'}\frac{\|\ell_{\theta}(w)-\ell_{\theta}(w')\|}{\|w - w' \|}.
$$
The value of $\xi_{k,\theta}$ depends on the choice of the basis for
the series estimator and on the linear functional. Newey1997
and Chen2006 provide several examples. In the case of splines with $\mathcal{X}=[0,1]^d$, it has been established that
$\xi_k \lesssim \sqrt{k}$ and $\sup_{x\in\mathcal{X}}\|\partial_j^mp(x)\| \lesssim k^{1/2 + m}$; see, for example, Newey1997. With this basis we have for
itemize• the function $g$ itself: $\theta(x) = g(x)$, $\ell_{\theta}(x)=p(x)$, and $\xi_{k,\theta}\lesssim \sqrt{k}$;
• the derivatives: $\theta(x) = \partial_j g(x)$, $\ell_\theta(x)= \partial_j p(x)$, $\xi_{k,\theta} \lesssim k^{3/2}$;
• the average derivatives: $\theta = \int \partial_j g(x) d\mu(x)$, $\ell_\theta = \int \partial_j p(x) d\mu(x)$, and $\xi_{k,\theta} \lesssim 1$,
where in the last example it is assumed that ${\rm supp}(\mu) \subset {\rm int}\mathcal{X}$, $x_1$ is continuously distributed with the density bounded below from zero on ${\rm supp}(\mu)$, and $x \mapsto \partial_l \mu(x)$ is continuous
on ${\rm supp}(\mu)$ with $|\partial_l \mu(x)| \lesssim 1$ uniformly in $x \in {\rm supp}(\mu)$ for all $l=1,\dots,k$.
We will impose the following regularity condition on the loadings on the coefficients $\ell_\theta(w)$:
Condition A.6 (Loadings) Loadings on the coefficients satisfy (i) $\sup_{w\in\mathcal{I}}1/\|\ell_{\theta}(w)\|\lesssim 1$ and (ii) $\log \xi^L_{k,\theta}\lesssim \log k$.
The first part of this condition implies that the linear functional is normalized appropriately. The second part is a very mild restriction on the rate of the growth of the Lipschitz coefficient of the map $w \mapsto \theta(w)$.
Under Conditions A.1-A.6, results presented in Lemma (ref) on uniform linearization can be extended to cover general linear functionals considered here:
lemma[Uniform Linearization for Linear Functionals]
Assume that Conditions A.1-A.6 are satisfied.
Then for $\alpha_{\theta}(w)=\ell_{\theta}(w)/\|\ell_{\theta}(w)\|$,
$$
\sqrt{n}\alpha_{\theta}(w)'(\widehat{\beta}-\beta)=\alpha_{\theta}(w)'\mathbb{G}_n[p_i(\epsilon_i+r_i)]+R_{1n}(\alpha_{\theta}(w)),
$$
where $R_{1n}(\alpha_{\theta}(w))$, summarizing the impact of unknown design, obeys
$$
R_{1n}(\alpha(w)) \lesssim_P \sqrt{\frac{\xi_k^2 \log k }{n}} (n^{1/m} \sqrt{\log k} + \sqrt{k} \cdot \ell_kc_{k})=\bar{R}_{1n}
$$
uniformly over $w \in \mathcal{I}$.
Moreover,
$$
\sqrt{n} \alpha_{\theta}(w)'( \widehat \beta - \beta) = \alpha_{\theta}(w)' \mathbb{G}_n[ p_i \epsilon_i ] + R_{1n}(\alpha_{\theta}(w)) +R_{2n}(\alpha_{\theta}(w)),
$$
where $R_{2n}(\alpha_{\theta}(w))$, summarizing the impact of approximation error on the sampling error of the estimator, obeys
$$
R_{2n}(\alpha_{\theta}(w)) \lesssim_P \sqrt{\log{k}} \cdot \ell_kc_{k}=\bar{R}_{2n}
$$
uniformly over $w\in\mathcal{I}$.
From Lemma (ref), we can derive the following theorem on uniform rate of convergence for linear functionals.
theorem[Uniform Rate of Convergence for Linear Functionals]
Assume that Conditions A.1-A.6 are satisfied. Then
\begin{equation}
\sup_{w\in\mathcal{I}}\left|\alpha_{\theta}(w)'\mathbb{G}_n[p_i\epsilon_i]\right|\lesssim_P \sqrt{\log k}.
\end{equation}
If, in addition, we assume that (i) $\bar{R}_{1n}+\bar{R}_{2n}\lesssim (\log k)^{1/2}$ and (ii) $\sup_{w\in\mathcal{I}}|r_{\theta}(w)|/\|\ell_{\theta}(w)\| = o((\log k/n)^{1/2})$, then
\begin{equation}
\sup_{w\in \mathcal{I}} | \widehat \theta(w) - \theta(w)| \lesssim_P \sqrt{\frac{\xi_{k,\theta}^2\log k}{n}}.
\end{equation}
Theorem (ref) establishes uniform rates that are up to $\sqrt{\log k}$ factor agree with the pointwise rates. The requirement (ii) on the approximation error can be seen as an undersmoothing condition as discussed in Comment (ref).
Next, we consider the problem of uniform inference for linear functionals based on the series estimator.
We base our inference on the t-statistic process:
equation[equation omitted — 166 chars of source]
We present two results for inference on linear functionals. The first result is an extension of Theorem (ref) on strong approximations to cover the case of linear functionals. As we discussed in Comment (ref), in order to perform uniform in $w\in\mathcal{I}$ inference on $\theta(w)$, we would like to approximate the distribution of the whole process ((ref)). However, one can show that this process typically does not have a limit distribution in $\ell^\infty(\mathcal{I})$. Yet, we can construct a Gaussian process that would be close to the process ((ref)) for all $w\in\mathcal{I}$ simultaneously with a high probability. Specifically, we will approximate the $t$-statistic process by the following Gaussian coupling:
equation[equation omitted — 196 chars of source]
where $\mathcal{N}_k$ denotes a vector of $k$ i.i.d. $N(0,1)$ random variables.
theorem[Strong Approximation by a Gaussian Process for Linear Functionals]
Assume that the conditions of Theorem (ref) and Condition A.6 are satisfied. In addition, assume that (i) $\bar{R}_{2n}\lesssim (\log k)^{1/2}$ and (ii) $\sup_{w\in\mathcal{I}}\sqrt{n}|r_{\theta}(w)|/\|\ell_{\theta}(w)\|=o(a_n^{-1})$. Then
$$
t(w) =_d t^*(w) + o_P(a_n^{-1}) \text{ in } \ell^{\infty}(\mathcal{I}).
$$
As in the case of inference on the function $g(x)$, we could also consider the use of the weighted bootstrap method to obtain a result analogous to that in Theorem (ref). For brevity of the paper, however, we do not consider weighted bootstrap method here.
The second result on inference for linear functionals is new and concerns with the problem of constructing uniform confidence bands for the linear functional $\theta(w)$. Specifically, we are interested in the confidence bands of the form
equation[equation omitted — 231 chars of source]
where $c_n(1-\alpha)$ is chosen so that $\theta(w)\in[\dot{\iota}(w),\ddot{\iota}(w)]$ for all $w\in\mathcal{I}$ with the prescribed probability $1-\alpha$ where $\alpha\in(0,1)$ is a user-specified level. For this purpose, we would like to set $c_n(1-\alpha)$ as the $(1-\alpha)$-quantile of $\sup_{w\in\mathcal{I}}|t(w)|$.
However, this choice is infeasible because the exact distribution of $\sup_{w\in\mathcal{I}}|t(w)|$ is unknown. Instead, Theorem (ref) suggests that we can set $c_n(1-\alpha)$ as the $(1-\alpha)$-quantile of $\sup_{w\in\mathcal{I}}|t^*(w)|$ or, if $\Omega$ is unknown and has to be estimated, that we can set
equation[equation omitted — 171 chars of source]
where
$$
\widehat{t}_n^*(w):=\frac{\ell(w)^\prime\widehat{\Omega}^{1/2}\mathcal{N}_k/\sqrt{n}}{\widehat{\sigma}_\theta(w)}, \, w\in\mathcal{I}
$$
and $\mathcal{N}_k\sim N(0,I_k)$. Note that $c_n(1-\alpha)$ defined in ((ref)) can be approximated numerically by simulation. Yet, conditions of Theorem (ref) are rather strong. Fortunately, CCK2012 noticed that when we are only interested in the supremum of the process and do not need the process itself, sufficient conditions for the strong approximation can be much weaker. Specifically, we have the following theorem, which is an application of a general result obtained in CCK2012:
theorem[Strong Approximation of Suprema for Linear Functionals]
Assume that Conditions A.1-A.6 are satisfied with $m\geq 4$. In addition, assume that (i) $\bar{R}_{1n}+\bar{R}_{2n}\lesssim 1/(\log k)^{1/2}$, (ii) $\xi_k\log^2 k/n^{1/2-1/m} \to 0$, (iii) $1\lesssim \underline{\sigma}^2$, and (iv) $\sup_{w\in\mathcal{I}}\sqrt{n}|r_{\theta}(w)|/\|\ell_{\theta}(w)\|=o(1/(\log k)^{1/2})$. Then
$$
\sup_{w\in\mathcal{I}}|t(w)| =_d \sup_{t\in\mathcal{I}}|t^*(w)| + o_P\left(\frac{1}{\sqrt{\log k}}\right).
$$
Construction of uniform confidence bands also critically relies on the following anti-concentration lemma due to
CCK2012b (Corollary 2.1):
lemma[Anti-concentration for Separable Gaussian Processes]
Let $Y=(Y_t)_{t\in T}$ be a separable Gaussian process indexed by a semimetric space $T$ such that $E[Y_t]=0$ and $E[Y_t^2]=1$ for all $t\in T$. Assume that $\sup_{t\in T}Y_t<\infty$ a.s. Then $a(|Y|):=E[\sup_{t\in T}|Y_t|]<\infty$ and
$$
\sup_{x\in\mathbb{R}}P\left\{\left|\sup_{t\in T}|Y_t|-x\right|\leq \varepsilon\right\}\leq A\varepsilon a(|Y|)
$$
for all $\varepsilon\geq 0$ and some absolute constant $A$.
From Theorem (ref) and Lemma (ref), we can now derive the following result on uniform validity of confidence bands in ((ref)):
theorem[Uniform Inference for Linear Functionals]
Assume that the conditions of Theorem (ref) are satisfied. In addition, assume that $c_n(1-\alpha)$ is defined by ((ref)). Then
\begin{equation}
P\left\{\sup_{w\in\mathcal{I}}|t_n(w)|\leq c_n(1-\alpha)\right\}=1-\alpha+o(1).
\end{equation}
As a consequence, the confidence bands defined in ((ref)) satisfy
\begin{equation}
P \Big\{ \theta(w) \in [\dot{\iota}(w), \ddot{\iota}(w)], for all w \in \mathcal{I} \Big\} = 1-\alpha + o(1).
\end{equation}
The width of the confidence bands $2c_n(1-\alpha)\widehat{\sigma}_n(w)$ obeys
\begin{equation}
2c_n(1-\alpha)\widehat{\sigma}_n(w)\lesssim_P \sigma_n(w)\sqrt{\log k}\lesssim \|\ell_{\theta}(w)\|\sqrt{\frac{\log k}{n}}\lesssim \sqrt{\frac{\xi_{k,\theta}^2\log k}{n}}
\end{equation}
uniformly over $w\in\mathcal{I}$.
remark(i) This is our fifth (and last) main result in this paper. The theorem shows that the confidence bands constructed above maintain the required level asymptotically and establishes that the
uniform width of the bands is of the same order as the uniform rate
of convergence. Moreover, confidence
intervals are asymptotically similar.
(ii) The proof strategy of Theorem (ref) is similar to that proposed in CLR2008 for inference on the minimum of a function. Since the limit distribution may not exists, the insight was to use distributions provided by couplings. Because the limit distribution does not necessarily exist, it is not immediately clear
that the confidence bands are asymptotically similar or at least maintain the right asymptotic level.
Nonetheless, we show that the confidence bands are asymptotically similar with the help of anti-concentration lemma stated above.
(iii) Theorem (ref) only considers two-sided confidence bands. However, both Theorem (ref) and Lemma (ref) continue to hold if we replace suprema of absolute values of the processes by suprema of the processes itself, namely if we replace $\sup_{w\in\mathcal{I}}|t_n(w)|$ and $\sup_{w\in\mathcal{I}}|t_n^{*}(w)|$ in Theorem (ref) by $\sup_{w\in\mathcal{I}}t_n(w)$ and $\sup_{w\in\mathcal{I}}t_n^{*}(w)$, respectively, and $\sup_{t\in T}|Y_t|$ in Lemma (ref) by $\sup_{t\in T}Y_t$. Therefore, we can show that Theorem (ref) also applies for one-sided confidence bands, namely Theorem (ref) holds with $c_n(1-\alpha)$ defined as the conditional $(1-\alpha)$-quantile of $\sup_{w\in\mathcal{I}}\widehat{t}_n^{*}(w)$ given the data
and the confidence bands defined by
$[\dot{\iota}(w),\ddot{\iota}(w)]:=[\widehat{\theta}(w)-c_n(1-\alpha)\widehat{\sigma}_n(w),+\infty)$ for all $w\in\mathcal{I}$.\qed
Tools: Maximal Inequalities for Matrices and Empirical Processes
In this section we collect the main technical tools that our analysis rely upon, namely Khinchin Inequalities for Matrices and Data Dependent Maximal Inequalities.
Khinchin Inequalities for Matrices
For $p\geq 1$, consider the Schatten norm $S_p$ on symmetric $k\times k$ matrices $Q$ defined by
$$
\|Q\|_{S_p} = \left( \sum_{j=1}^k | \lambda_j(Q) |^p \right)^{1/p}
$$
where $\lambda_1(Q),\dots,\lambda_k(Q)$ is the system of eigenvalues of $Q$.
The case $p=\infty$ recovers the operator norm $\|\cdot \|$ and $p=2$ the Frobenius norm. It is obvious that for any $p\geq 1$
$$
\|Q\| \leq \|Q\|_{S_p} \leq k^{1/p} \|Q\|.
$$
Therefore, setting $p=\log k$ and observing that $k^{1/\log k}=e$ for any $k\geq 1$, we get the relation:
equation[equation omitted — 73 chars of source]
lemma[Khinchin Inequality for Matrices]
For symmetric $k \times k$-matrices $Q_i$, $i=1,\dots,n$, $2 \leq p < \infty$, and an i.i.d. sequence
of Rademacher variables $\varepsilon_1,\dots,\varepsilon_n$, we have
\begin{equation}
\left \| \left( \mathbb{E}_n [Q^2_i] \right)^{1/2} \right \|_{S_p} \leq
\left ( E_{\varepsilon} \left \| \mathbb{G}_n [\varepsilon_i Q_i] \right \|^p_{S_p} \right )^{1/p} \leq C\sqrt{p} \left \| ( \mathbb{E}_n[Q^2_i] )^{1/2} \right \|_{S_p}
\end{equation}
for some absolute constant $C$.
As a consequence, we have for $k\geq 2$
\begin{equation}
E_{\varepsilon}\left[\left\| \mathbb{G}_n[\varepsilon_i Q_i]\right\|\right] \leq C\sqrt{\log k} \left \| ( \mathbb{E}_n[ Q_i^2])^{1/2} \right\|
\end{equation}
for some (possibly different) absolute constant $C$.
This version of the Khinchin inequality is proven in Section 3 of Rudelson1999. We also provide some details of the proof in the Appendix.
The notable feature of this inequality is the $\sqrt{\log k }$ factor instead of the $\sqrt{k}$ factor expected
from the conventional maximal inequalities based on entropy. This inequality due to LustPicardPisier generalizes the Khinchin inequality for vectors. A version of this inequality was derived by GuedonRudelson using generalized entropy (majorizing measure) arguments.
This is a striking example where the use of generalized entropy yields drastic improvements over the use of entropy. Prior to this,
Talagrand1996 provided ellipsoidal examples where the difference between the two approaches was even more extreme.
LLN for Matrices
The following lemma is a variant of a fundamental result obtained by Rudelson1999.
lemma[Rudelson's LLN for Matrices] Let $Q_1,\dots,Q_n$ be a sequence of independent symmetric non-negative $k \times k$-matrix valued random variables with $k \geq 2$ such that $Q = \mathbb{E}_n[E[Q_i]]$ and $\|Q_i\| \leq M$ a.s., then for $\widehat Q = \mathbb{E}_n[Q_i]$
$$
\Delta : = E\| \widehat Q - Q \| \lesssim \frac{M\log k}{n} + \sqrt{ \frac{M \|Q\| \log k}{n}}.
$$
In particular, if $Q_i = p_i p_i'$, with $\|p_i\| \leq \xi_k$ a.s., then
$$
\Delta : = E\| \widehat Q - Q \| \lesssim \frac{\xi_k^2\log k}{n} + \sqrt{\frac{\xi_k^2 \|Q\| \log k }{n}}.
$$
For completeness, we provide the proof of this lemma in the Appendix; see also Tropp12 for a nice exposition of this result as well as many others concerning with maximal and deviation inequalities for matrices.
Maximal Inequalities
Consider a measurable space $(S, \mathcal{S})$, and a suitably measurable class of functions $\mathcal{F}$ mapping $S$ to $\Bbb{R}$, equipped with a measurable envelope function $F(z) \geq \sup_{f \in \mathcal{F}} |f(z)|$. (By “suitably measurable"
we mean the condition given in Section 2.3.1 of vdV-W; pointwise measurablity and Suslin measurability are sufficient.) The covering number $N(\mathcal{F},L^{2}(Q),\varepsilon)$ is the minimal number of $L^{2}(Q)$-balls of radius $\varepsilon$ needed to cover $\mathcal{F}$. The covering number relative to the envelope function is given by
equation[equation omitted — 99 chars of source]
The entropy is the logarithm of the covering
number.
We rely on the following result.
propositionLet $(\epsilon_{1},X_{1}), \dots,(\epsilon_{n},X_{n})$ be i.i.d. random vectors, defined on an underlying $n$-fold product probability space, in $\mathbb{R}^{d+1}$ with $E[\epsilon_{i} | X_{i}] = 0$ and $\sigma^{2} := \sup_{x \in \mathcal{X}} E[ \epsilon_{i}^{2} | X_{i} = x] < \infty$ where $\mathcal{X}$ denotes the support of $X_1$. Let $\mathcal{F}$ be a class of functions on $\mathbb{R}^{d}$ such that $E [ f(X_{1})^{2} ] = 1$ (normalization) and $\| f \|_{\infty} \leq b$ for all $f \in \mathcal{F}$.
Let $\mathcal{G} := \{ \mathbb{R}\times \mathbb{R} \ni (\epsilon,x) \mapsto \epsilon f(x) : f \in \mathcal{F} \}$. Suppose that there exist constants $A > e^{2}$ and $V \geq 2$ such that
$$\sup_{Q} N(\mathcal{G}, L^{2}(Q), \varepsilon \| G \|_{L^{2}(Q)}) \leq (A/\varepsilon)^{V}$$ for all $0 < \varepsilon \leq 1$ for the envelope $G(\epsilon,x) := | \epsilon| b$.
If for some $m > 2$ $E[ | \epsilon_{1} |^{m} ] < \infty$, then
\begin{equation*}
E \left [ \left \| \sum_{i=1}^{n} \epsilon_{i} f(X_{i}) \right \|_{\mathcal{F}} \right] \leq C \left [ (\sigma + \sqrt{E[ | \epsilon_{1} |^{m} ]}) \sqrt{ n V \log (Ab)} + V b^{m/(m-2)} \log (Ab) \right ],
\end{equation*}
where $C$ is a universal constant.
The proof is based on a truncation argument and maximal inequalities for uniformly bounded classes of functions developed in GK06. We recall its version.
theorem[GK06]
Let $\xi_{1},\dots,\xi_{n}$ be i.i.d. random variables taking values in a measurable space $(S,\mathcal{S})$ with common distribution $P$, defined on the underlying $n$-fold product probability space.
Let $\mathcal{F}$ be a suitably measurable class of functions mapping $S$ to $\Bbb{R}$ with a measurable envelope $F$. Let $\sigma^{2}$ be a constant such that $
\sup_{f \in \mathcal{F}} \operatorname{var}(f) \leq \sigma^{2} \leq \| F \|_{L^{2}(P)}^{2}$. Suppose that there exist constants $A > e^{2}$ and $V \geq 2$ such that
$\sup_{Q} N(\mathcal{F}, L^{2}(Q), \varepsilon \| F \|_{L^{2}(Q)}) \leq (A/\varepsilon)^{V}$ for all $0 < \varepsilon \leq 1$. Then,
\begin{equation*}
E \left [ \left \| \sum_{i=1}^{n} \{ f(\xi_{i}) - E[ f(\xi_{1}) ]\} \right \|_{\mathcal{F}} \right] \leq C \left [ \sqrt{n \sigma^{2} V \log \frac{A \| F \|_{L^{2}(P)}}{\sigma}} + V \| F \|_{\infty} \log \frac{A \| F \|_{L^{2}(P)}}{\sigma} \right ],
\end{equation*}
where $C$ is a universal constant.
Acknowledgements.
This paper was presented and first circulated in a series of lectures
given by Victor Chernozhukov at “Stats in the Ch\^{a}teau" Statistics Summer School on “Inverse Problems and High-Dimensional Statistics" in 2009 near Paris. Participants, especially Xiaohong Chen, and one of several referees made numerous helpful suggestions. We also thank Bruce Hansen for extremely useful comments.