Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,948 characters · 11 sections · 80 citation commands
Optimal Uniform Convergence Rates for Sieve Nonparametric Instrumental Variables Regression
\thispagestyle{empty} \setcounter{page}{0}
In economics and other social sciences one frequently encounters the relation
where $Y_{1i}$ is a response variable, $Y_{2i}$ is a predictor variable, $h_0$ is an unknown structural function of interest, and $\epsilon_i$ is an error term. However, a latent external mechanism may “determine” or “cause” $Y_{1i}$ and $Y_{2i}$ simultaneously, in which case the conditional mean restriction $E[\epsilon_i |Y_{2i}] = 0$ fails and $Y_{2i}$ is said to be endogenous.\footnote{In a canonical example of this relation, $Y_{1i}$ may be the hourly wage of person $i$ and $Y_{2i}$ may include the education level of person $i$. The latent ability of person $i$ affects both $Y_{1i}$ and $Y_{2i}$. See BlundellPowell for other examples and discussions of endogeneity in semi/nonparametric regression models.} When the regressor $Y_{2i}$ is endogenous one cannot use standard nonparametric regression techniques to consistently estimate $h_0$. In this instance one typically assumes that there exists a vector of instrumental variables $X_i$ such that $E[\epsilon_i |X_i] = 0$ and for which there is a nondegenerate relationship between $X_i$ and $Y_{2i}$. Such a setting permits estimation of $h_0$ using nonparametric instrumental variables (NPIV) techniques based on a sample $\{(X_i,Y_{1i},Y_{2i})\}_{i=1}^n$. In this paper we assume that the data is strictly stationary in that $(X_i,Y_{1i},Y_{2i})$ has the same (unknown) distribution $F_{X,Y_1,Y_2}$ as that of $(X,Y_{1},Y_{2})$ for all $i$.\footnote{The subscript $i$ denotes either the individual $i$ in a cross-sectional sample or the time period $i$ in a time-series sample. Since the sample is strictly stationary we sometimes drop the subscript $i$ without confusion.}
NPIV estimation has been the subject of much research in recent years, both because of its practical importance to applied economics and its prominent role in the literature on linear ill-posed inverse problems with unknown operators. In many economic applications the joint distribution $F_{X,Y_2}$ of $X_i$ and $Y_{2i}$ is unknown but is assumed to have a continuous density. Therefore the conditional expectation operator $Th(\cdot)=E[h(Y_{2i})|X_i = \cdot ]$ is typically unknown but compact. Model ((ref)) with $E[\epsilon_i |X_i] = 0$ can be equivalently written as
where $u_i = h_0(Y_{2i}) - Th_0(X_i) + \epsilon_i$. Model ((ref)) is called the reduced-form NPIV model if $T$ is assumed to be unknown and the nonparametric indirect regression (NPIR) model if $T$ is assumed to be known. Let $\widehat{E}[Y_{1}|X= \cdot]$ be a consistent estimator of $E[Y_{1}|X= \cdot]$. Regardless of whether the compact operator $T$ is unknown or known, nonparametric recovery of $h_0$ by inversion of the conditional expectation operator $T$ on the left-hand side of the Fredholm equation of the first kind
leads to an ill-posed inverse problem (see, e.g., Kress). Consequently, some form of regularization is required for consistent nonparametric estimation of $h_0$. In the literature there are several popular methods of NPIV estimation, including but not limited to (1) finite-dimensional sieve minimum distance estimators \citep*{NeweyPowell,AiChen2003,Blundell2007}; (2) kernel-based Tikhonov regularization estimators \citep*{HallHorowitz,Darollesetal2011,GagliardiniScaillet} and their Bayesian version FlorensSimoni; (3) orthogonal series Tikhonov regularization estimators HallHorowitz; (4) orthogonal series Galerkin-type estimators Horowitz2011; (5) general penalized sieve minimum distance estimators ChenPouzo2012 and their Bayesian version LiaoJiang. See Horowitz2011 for a recent review and additional references.
To the best of our knowledge, all the existing works on convergence rates for various NPIV estimators have only studied $L^2$-norm convergence rates. In particular, HallHorowitz are the first to establish the minimax risk lower bound in $L^2$-norm loss for a class of mildly ill-posed NPIV models, and show that their estimators attain the lower bound. ChenReiss derive the minimax risk lower bound in $L^2$-norm loss for a large class of NPIV models that could be mildly or severely ill-posed, and show that the sieve minimum distance estimator of Blundell2007 achieves the lower bound. Subsequently, some other NPIV estimators listed above have also been shown to achieve the optimal $L^2$-norm convergence rates. As yet there are no published results on sup-norm (uniform) convergence rates for any NPIV estimators, nor results on what are the minimax risk lower bounds in sup-norm loss for any class of NPIV models.
Sup-norm convergence rates for any estimators of $h_0$ are important for constructing uniform confidence bands for the unknown $h_0$ in NPIV models and for conducting inference on nonlinear functionals of $h_0$, but are currently missing. In this paper we study the uniform convergence properties of the sieve minimum distance estimator of $h_0$ for the NPIV model, which is a nonparametric series two-stage least squares regression estimator \citep*{NeweyPowell,AiChen2003,Blundell2007}. We focus on this estimator because it is easy to compute and has been used in empirical work in demand analysis \citep*{Blundell2007,ChenPouzo2009}, asset pricing Chen2009, and other applied fields in economics. Also, this class of estimators is known to achieve the optimal $L^2$-norm convergence rates for both mildly and severely ill-posed NPIV models.
We first establish a general upper bound (Theorem (ref)) on the uniform convergence rate of a sieve estimator, allowing for endogenous regressors and weakly dependent data. To provide sharp bounds on the sieve approximation error or “bias term” we extend the proof strategy of Huang2003 for sieve nonparametric least squares (LS) regression to the sieve NPIV estimator. Together, these tools yield sup-norm convergence rates for the spline and wavelet sieve NPIV estimators under i.i.d. data. Under conditions similar to those for the $L^2$-norm convergence rates for the sieve NPIV estimators, our sup-norm convergence rates coincide with the known optimal $L^2$-norm rates for severely ill-posed problems, and are power of $\log( n)$ slower than the optimal $L^2$-norm rates for mildly ill-posed problems. We then establish the minimax risk lower bound in sup-norm loss for $h_0$ in a NPIR model (i.e., ((ref)) with a known compact $T$) uniformly over H\"older balls, which in turn provides a lower bound in sup-norm loss for $h_0$ in a NPIV model uniformly over H\"older balls. The lower bound is shown to coincide with our sup-norm convergence rates for the spline and wavelet sieve NPIV estimators.
To establish the general upper bound, we first derive a new exponential inequality for sums of weakly dependent random matrices in Section (ref). This allows us to weaken conditions under which the optimal uniform convergence rates can be obtained. As an indication of the sharpness of our general upper bound result, we show that it leads to the optimal uniform convergence rates for spline and wavelet LS regression estimators with weakly dependent data and heavy-tailed error terms. Precisely, for beta-mixing dependent data and finite $(2+\delta )$-th moment error term (for $\delta \in (0,2)$), we show that the spline and wavelet nonparametric LS regression estimators attain the minimax risk lower bound in sup-norm loss of Stone1982. This result should be very useful to the literature on nonparametric estimation with financial time series.
The NPIV model falls within the class of statistical linear ill-posed inverse problems with unknown operators and additive noise. There is a vast literature on statistical linear ill-posed inverse problems with known operators and additive noise. Some recent references include but are not limited to \citet*{Cavalieretal2002}, \citet*{Cohenetal2004} and Cavalier2008, of which density deconvolution is an important and extensively-studied problem (see, e.g., CarrollHall,Zhang,Fan1991,HallMeister,LouniciNickl). There are also papers on statistical linear ill-posed inverse problems with pseudo-unknown operators (i.e., known eigenfunctions but unknown singular values) (see, e.g., CavalierHengartner, LoubesMarteau). Related papers that allow for an unknown linear operator but assume the existence of an estimator of the operator (with rate) include EfromovichKoltchinskii, HoffmannReiss and others. To the best of our knowledge, most of the published works in the statistical literature on linear ill-posed inverse problems also focus on the rate optimality in $L^2$-norm loss, except that of LouniciNickl which recently establishes the optimal sup-norm convergence rate for a wavelet density deconvolution estimator. Therefore, our minimax risk lower bounds in sup-norm loss for the NPIR and NPIV models also contribute to the large literature on statistical ill-posed inverse problems.
The rest of the paper is organized as follows. Section (ref) outlines the model and presents a general upper bound on the uniform convergence rates for a sieve estimator. Section (ref) establishes the optimal uniform convergence rates for the sieve NPIV estimators, allowing for both mildly and severely ill-posed inverse problems. Section (ref) derives the optimal uniform convergence rates for the sieve nonparametric (least squares) regression, allowing for dependent data. Section (ref) provides useful exponential inequalities for sums of random matrices, and the reinterpretation of equivalence of the theoretical and empirical $L^2$ norms as a criterion regarding convergence of a random matrix. The appendix contains a brief review of the spline and wavelet sieve spaces, proofs of all the results in the main text, and supplementary results.
\paragraph{Notation:} $ \|\cdot\|$ denotes the Euclidean norm when applied to vectors and the matrix spectral norm (largest singular value) when applied to matrices. For a random variable $Z$ let $ L^q(Z) $ denote the spaces of (equivalence classes of) measurable functions of $z$ with finite $q$-th moment if $1 \leq q < \infty$ and let $\|\cdot\|_{L^q (Z)}$ denote the $L^q (Z)$ norm. Let $L^\infty(Z)$ denote the space of measurable functions of $z$ with finite sup norm $\|\cdot\|_\infty$. If $A$ is a square matrix, $\lambda_{\min}( A)$ and $\lambda_{\max}(A)$ denote its smallest and largest eigenvalues, respectively, and $A^-$ denotes its Moore-Penrose generalized inverse. If $\{a_n:n \geq 1\}$ and $\{b_n : n \geq 1\}$ are two sequences of non-negative numbers, $a_n \lesssim b_n$ means there exists a finite positive $C$ such that $a_n \leq C b_n$ for all $n$ sufficiently large, and $a_n \asymp b_n$ means $a_n \lesssim b_n$ and $b_n \lesssim a_n$. $\#(\mathcal S)$ denotes the cardinality of a set $\mathcal S$ of finitely many elements. Let $\mbox{BSpl}(K,[0,1]^d, \gamma )$ and $\mbox{Wav}(K,[0,1]^d, \gamma )$ denote tensor-product B-spline (with smoothness $\gamma$) and wavelet (with regularity $\gamma$) sieve spaces of dimension $K$ on $[0,1]^d$ (see Appendix (ref) for details on construction of these spaces).
We begin by considering the NPIV model
where $Y_1 \in \mathbb R$ is a response variable, $Y_2$ is an endogenous regressor with support $\mathcal Y_2 \subset \mathbb R^d$ and $X$ is a vector of conditioning variables (also called instruments) with support $\mathcal{X }\subset \mathbb{R}^{d_x}$. The object of interest is the unknown structural function $h_0 : \mathcal{Y}_2 \to \mathbb{R}$ which belongs to some infinite-dimensional parameter space $\mathcal H \subset L^2(Y_2)$. It is assumed hereafter that $h_0$ is identified uniquely by the conditional moment restriction ((ref)). See NeweyPowell, \citet*{Blundell2007}, \citet*{Darollesetal2011}, Andrews2011, D'Haultfoeuille, \citet*{Chen2012b} and references therein for sufficient conditions for identification.
The sieve NPIV estimator due to NeweyPowell, AiChen2003, and Blundell2007 is a nonparametric series two-stage least squares estimator. Let the sieve spaces $\{\Psi _{J}:J\geq 1\} \subseteq L^2(Y_2)$ and $\{B_K : K \geq 1\} \subset L^2(X)$ be sequences of subspaces of dimension $J$ and $K$ spanned by sieve basis functions such that $\Psi_J$ and $B_K$ become dense in $\mathcal H \subset L^{2}(Y_{2})$ and $L^2(X)$ as $J,K\rightarrow \infty $. For given $J$ and $K$, let $\{\psi_{J1},\ldots,\psi_{JJ}\}$ and $\{b_{K1},\ldots,b_{KK}\}$ be sets of sieve basis functions whose closed linear span generates $\Psi_J$ and $B_K$ respectively. We consider sieve spaces generated by spline, wavelet or other Riesz basis functions that have nice approximation properties (see Section (ref) for details).
In the first stage, the conditional moment function $ m(x,h):\mathcal{X}\times \mathcal{H}\rightarrow \mathbb{R}$ given by
is estimated using the series (least squares) regression estimator
where
The sieve NPIV estimator $\widehat h$ is then defined as the solution to the second-stage minimization problem
which may be solved in closed form to give
where
Under mild regularity conditions (see NeweyPowell, \citet*{Blundell2007} and ChenPouzo2012), $\widehat h$ is a consistent estimator of $h_0$ (in both $\|\cdot\|_{L^2(Y_2)}$ and $\|\cdot\|_\infty$ norms) as $n,J,K \to \infty$, provided $J \leq K$ and $J$ increases appropriately slowly so as to regularize the ill-posed inverse problem.\footnote{Here we have used $K$ to denote the “smoothing parameter” (i.e. the dimension of the sieve space used to estimate the conditional moments in ((ref))) and $J$ to denote the “regularization parameter” (i.e. the dimension of the sieve space used to approximate the unknown $h_0$). Note that ChenReiss use $J$ and $m$, \citet*{Blundell2007} and ChenPouzo2012 use $J$ and $k$ to denote the smoothing and regularization parameters, respectively.} We note that the modified sieve estimator (or orthogonal series Galerkin-type estimator) of Horowitz2011 corresponds to the sieve NPIV estimator with $J=K$ and $\psi^J (\cdot) = b^K (\cdot)$ being orthonormal basis in $L^2 (Lebesgue)$.
We first present a general calculation for sup-norm convergence which will be used to obtain uniform convergence rates for both the sieve NPIV and the sieve LS estimators below.
As the sieve estimators are invariant to an invertible transformation of the sieve basis functions, we re-normalize the sieve spaces $B_K$ and $\Psi_J$ so that $\{\widetilde b_{K1},\ldots,\widetilde b_{KK}\}$ and $\{\widetilde \psi_{J1},\ldots,\widetilde \psi_{JJ}\}$ form orthonormal bases for $B_K$ and $\Psi_J$. This is achieved by setting $\widetilde b^K(x) = E[b^K(X)b^K(X)']^{-1/2}b^K(x)$ where $^{-1/2}$ denotes the inverse of the positive-definite matrix square root (which exists under Assumption (ref)(ii) below), with $\widetilde \psi^J$ similarly defined. Let
and define the $J \times K$ matrices
Let $\sigma_{JK}^2 = \lambda_{\min}(SS')$. For each $h \in \Psi_J$ define
which is the $L^2(X)$ orthogonal projection of $T h(\cdot)$ onto $B_K$. The variational characterization of singular values gives
Finally, define $P_n$ as the second-stage empirical projection operator onto the sieve space $\Psi_J$ after projecting onto the instrument space $B_K$, viz.
where $H_0 = (h_0(Y_{21}),\ldots,h_0(Y_{2n}))'$.
We first decompose the sup-norm error as
and calculate the uniform convergence rate for the “variance term” $\|\widehat h - P_n h_0\|_{\infty}$ in this section. Control of the “bias term” $\|h_0 - P_n h_0 \|_{\infty}$ is left to the subsequent sections, which will be dealt with under additional regularity conditions for the NPIV model and the LS regression model separately.
Let $Z_i = (X_i,Y_{1i},Y_{2i})$ and $\mathcal F_{i-1} = \sigma(X_i,X_{i-1},\epsilon_{i-1},X_{i-2},\epsilon_{i-2},\ldots)$.
The results stated in this section do not actually require that $\dim(X) = \dim(Y_2)$. However, most published papers on NPIV models assume $\dim(X) = \dim(Y_2) = d$ and so we follow this convention in Assumption 1(ii).
In what follows, $p>0$ indicates the smoothness of the function $h_0 (\cdot)$ (see Assumption (ref) in Section (ref)).
The preceding assumptions on the data generating process trivially nest i.i.d. sequences but also allow for quite general weakly-dependent data. In an i.i.d. setting, Assumption (ref)(ii) reduces to requiring that $E[\epsilon_i^2 |X_i = x]$ be bounded uniformly from zero and infinity which is standard (see, e.g., Newey1997,HallHorowitz). The value of $\delta$ in Assumption (ref)(iii) depends on the context. For example, $\delta \geq d/p$ will be shown to be sufficient to attain the optimal sup-norm convergence rates for series LS regression in Section (ref), whereas lower values of $\delta$ suffice to attain the optimal sup-norm convergence rates for the sieve NPIV estimator in Section (ref). Rectangular support and bounded densities of the endogenous regressor and instrument are assumed in HallHorowitz. Assumptions (ref)(i) and (ref)(i) are satisfied by many widely used sieve bases such as spline, wavelet and cosine sieves, but they rule out polynomial and power series sieves (see, e.g., Newey1997,Huang1998). The instruments sieve basis $b^K(\cdot) $ is used to approximate the conditional expectation operator $Th=E[h(Y_2)|X=\cdot)$, which is a smoothing operator. Thus Assumption (ref)(i) assumes that the sieve basis $b^K(\cdot) $ (for $Th$) is smoother than that of the sieve basis $\psi^J(\cdot) $ (for $h$).
In the next theorem, our upper bound on the “variance term” $\|\widehat h - P_n h_0\|_{\infty}$ holds under general weak dependence as captured by Condition (ii) on the convergence of the random matrices $\widetilde B'\widetilde B/n - I_K$ and $\widehat S - S$.
The restrictions on $J$, $K$ and $n$ in Conditions (i) and (ii) merit a brief explanation. The restriction $J \leq K$ merely ensures that the sieve NPIV estimator is well defined. The restriction $K \lesssim (n/\log n)^{\delta/(2+\delta)}$ is used to perform a truncation argument using the existence of $(2+\delta)$-th moment of the error terms (see Assumption (ref)). Condition (ii) ensures that $J$ increases sufficiently slowly that with probability approaching one the minimum eigenvalue of the “denominator” matrix $\Psi'B(B'B)^-B'\Psi/n$ is positive and bounded below by a multiple of $\sigma_{JK}^2$, thereby regularizing the ill-posed inverse problem. It also ensures the error in estimating the matrices $(\widetilde B'\widetilde B/n)$ and $\widehat S$ vanishes sufficiently quickly that it doesn't affect the convergence rate of the estimator.
We now exploit the specific linear structure of the sieve NPIV estimator to derive uniform convergence rates for the mildly and severely ill-posed cases. Some additional assumptions are required so as to control the “bias term” $\|h_0 - P_n h_0\|_{\infty}$ and to relate the estimator to the measure of ill-posedness.
$p$-smooth H\"older class of functions. We first impose a standard smoothness condition on the unknown structural function $h_0$ to facilitate comparison with Stone1982's minimax risk lower bound in sup-norm loss for a nonparametric regression function. Recall that $\mathcal Y_2 = [0,1]^d$. Deferring definitions to Triebel2006,Triebel2008, we let $B^{p}_{q,q}([0,1]^d)$ denote the Besov space of smoothness $p$ on the domain $[0,1]^d$ and $\|\cdot\|_{B^{p}_{q,q}}$ denote the usual Besov norm on this space. Special cases include the Sobolev class of smoothness $p$, namely $B^{p}_{2,2}([0,1]^d)$, and the H\"older-Zygmund class of smoothness $p$, namely $B^{p}_{\infty,\infty}([0,1]^d)$. Let $B(p,L)$ denote a H\"older ball of smoothness $p$ and radius $0 <L<\infty$, i.e. $B(p,L) = \{ h \in B^p_{\infty,\infty}([0,1]^d) : \|h\|_{B^p_{\infty,\infty}} \leq L\}$.
Assumptions (ref) and (ref) imply that there is $\pi_J h_0 \in \Psi_J$ such that $\|h_0 - \pi_J h_0\|_{\infty} = O(J^{-p/d})$.
Sieve measure of ill-posedness. Let $T : L^q(Y_2) \to L^q(X)$ denote the conditional expectation operator for $1\leq q \leq \infty$:
When $Y_2$ is endogenous, $T$ is compact under mild conditions on the conditional density of $Y_2$ given $X$. For $q'\geq q\geq 1$, we define a measure of ill-posedness (over a sieve space $\Psi_J$) as
The $\tau_{2,2,J}$ measure of ill-posedness is clearly related to our earlier definition of $\sigma_{JK}$. By definition
when $J \leq K$. The sieve measures of ill-posedness, $\tau_{2,2,J}$ and $\sigma_{JK}^{-1}$, are clearly non-decreasing in $J$. In \citet*{Blundell2007}, Horowitz2011 and ChenPouzo2012, the NPIV model is said to be
$\bullet$ mildly ill-posed if $\tau_{2,2,J} = O(J^{\varsigma/d})$ for some $\varsigma > 0 $;
$\bullet$ severely ill-posed if $\tau_{2,2,J} = O(\exp(\frac{1}{2} J^{\varsigma /d}))$ for some $\varsigma > 0$.
These measures of ill-posedness are not exactly the same as (but are related to) the measure of ill-posedness used in HallHorowitz and Cavalier2008. In the latter papers, it is assumed that the compact operator $T : L^2(Y_2) \to L^2(X)$ admits a singular value decomposition $\{\mu_{k};\phi_{1k},\phi_{0k}\}_{k=1}^{\infty }$, where $\{\mu_{k}\}_{k=1}^{\infty }$ are the singular numbers arranged in non-increasing order ($\mu_{k}\geq \mu_{k+1}\searrow 0$), $\{\phi_{1k}(y_{2})\}_{k=1}^{\infty }$ and $\{\phi_{0k}(x)\}_{k=1}^{\infty }$ are eigenfunction (orthonormal) bases for $ L^{2}(Y_{2})$ and $L^{2}(X)$ respectively, and ill-posedness is measured in terms of the rate of decay of the singular values towards zero. Denote $T^{\ast }$ as the adjoint operator of $T$: $\{T^{\ast }g\}(Y_{2})\equiv E[g(X)|Y_{2}]$, which maps $L^{2}(X) $ into $L^{2}(Y_{2})$. Then a compact $T$ implies that $T^{\ast }$, $T^{\ast}T$ and $TT^{\ast}$ are also compact, and that $T\phi_{1k}=\mu_{k}\phi_{0k}$ and $T^{\ast }\phi_{0k}=\mu_{k}\phi_{1k}$ for all $k$. We note that $\|Th \|_{L^2(X)}=\|(T^*T)^{1/2}h \|_{L^2(Y_2)}$ for all $h \in Dom(T)$. The following lemma provides some relations between these different measures of ill-posedness.
Lemma (ref) parts (1) and (2) is Lemma 1 of \citet*{Blundell2007}, while Lemma (ref) part (3) is proved in the Appendix. We next present a sufficient condition to bound the sieve measures of ill-posedness $\sigma_{JK}^{-1}$ and $ \tau_{2,2,J}$.
It is clear that Assumption (ref)(b) implies Assumption (ref)(a). Assumption (ref)(a) is the so-called “sieve reverse link condition” used in ChenPouzo2012, which is weaker than the “reverse link condition” imposed in ChenReiss and others in the ill-posed inverse literature: $\|Th \|_{L^2(X)}^{2}\gtrsim \sum_{j=1}^{\infty }\varphi (j^{-2/d})|E[h(Y_2)\widetilde \psi_{Jj} (Y_2)]|^{2} $ for all $h\in B(p,L)$. We immediately have the following bounds:
Given Remark (ref), in this paper we could call a NPIV model
$\bullet$ mildly ill-posed if $\sigma_{JK}^{-1} = O(J^{\varsigma/d})$ or $\varphi (t)=t^{\varsigma}$ for some $\varsigma > 0 $;
$\bullet$ severely ill-posed if $\sigma_{JK}^{-1} = O(\exp(\frac{1}{2} J^{\varsigma /d}))$ or $\varphi (t)=\exp(-t^{-\varsigma /2})$ for some $\varsigma > 0$.
Define
Assumption (ref)(ii) is a sup-norm analogue of the so-called “stability condition” imposed in the ill-posed inverse regression literature, such as Assumption 6 of \citet*{Blundell2007} and Assumption 5.2(ii) of ChenPouzo2012.
To control the “bias term” $\|P_n h_0 - h_0\|_{\infty}$, we will use spline or wavelet sieves in Assumptions (ref) and (ref) so that we can make use of sharp bounds on the approximation error due to Huang2003.\footnote{The key property of spline and wavelet sieve spaces that permits this sharp bound is their local support (see the appendix to Huang2003). Other sieve bases such as orthogonal polynomial bases do not have this property and are therefore unable to attain the optimal sup-norm convergence rates for NPIV or nonparametric series LS regression.} Control of the “bias term” $\|P_n h_0 - h_0\|_{\infty}$ is more involved in the sieve NPIV context than the sieve nonparametric LS regression context. In particular, control of this term makes use of an additional argument using exponential inequalities. To simplify presentation, the next theorem just presents the uniform convergence rate for sieve NPIV estimators under i.i.d. data.
ChenReiss show that these $L^2(Y_2) $-norm rates are optimal in the sense that they coincide with the minimax risk lower bound in $L^2(Y_2) $ loss. It is interesting to see that our sup-norm convergence rate is the same as the known optimal $L^2(Y_2) $-norm rate for the severely ill-posed case, and is only power of $\log(n)$ slower than the known optimal $L^2(Y_2) $-norm rate for the mildly ill-posed case. In the next subsection we will show that our sup-norm convergence rates are in fact optimal as well.
For severely ill-posed NPIV models, ChenReiss already showed that $(\log n)^{-p/\varsigma}$ is the minimax lower bound in $L^2(Y_2)$-norm loss uniformly over a class of functions that include the H\"older ball $B(p,L)$ as a subset. Therefore, we have for a severely ill-posed NPIV model with $\delta_n = (\log n)^{-p/\varsigma}$,
where $\inf_{\widetilde h_n}$ denotes the infimum over all estimators based on a random sample of size $n$ drawn from the NPIV model, and the finite positive constants $c, c'$ do not depend on sample size $n$. This and Remark (ref)(2) together imply that the sieve NPIV estimator attains the optimal uniform convergence rate in the severely ill-posed case.
We next show that the sup-norm rate for the sieve NPIV estimator obtained in the mildly ill-posed case is also optimal. We begin by placing a primitive smoothness condition on the conditional expectation operator $T : L^2(Y_2) \to L^2(X)$.
Assumption (ref) is a special case of the so-called “link condition” in ChenReiss for the mildly ill-posed case. It can be equivalently stated as: $\|Th \|_{L^2(X)}^{2} \lesssim \sum_{j=1}^{\infty }\varphi (j^{-2/d})|E[h(Y_2)\widetilde \psi_{Jj} (Y_2)]|^{2} $ for all $h\in B(p,L)$, with $\varphi (t)=t^{\varsigma}$ for the mildly ill-posed case. Under this assumption, $n^{-p/(2(p+\varsigma)+d)}$ is the minimax risk lower bound uniformly over the H\"older ball $B(p,L)$ in $L^2 (Y_2)$-norm loss for the mildly ill-posed NPIR and NPIV models (see ChenReiss). We next establish the corresponding minimax risk lower bound in sup-norm loss.
As in ChenReiss, Theorem (ref) is proved by (i) noting that the risk (in sup-norm loss) for the NPIV model is at least as large as the risk (in sup-norm loss) for the NPIR model, and (ii) calculating a lower bound (in sup-norm loss) for the NPIR model. We consider a Gaussian reduced-form NPIR model with known operator $T$, given by
Theorem (ref) therefore follows from a sup-norm analogue of Lemma 1 of ChenReiss and the following theorem, which establishes a lower bound on minimax risk over H\"older classes under sup-norm loss for the NPIR model.
The standard nonparametric regression model can be recovered as a special case of ((ref)) in which there is no endogeneity, i.e. $Y_2 = X$ and
in which case $h_0(x) = E[Y_{1i}|X_i = x]$.
Stone1982 (also see Tsybakov2009) establishes that $(n/\log n)^{-p/(2p +d)}$ is the minimax risk lower bound in sup-norm loss for the nonparametric LS regression model ((ref)) with $h_0 \in B(p,L)$. In this section we apply the general upper bound (Theorem (ref)) to show that spline and wavelet sieve LS estimators attain this minimax lower bound for weakly dependent data allowing for heavy-tailed error terms $\epsilon_i$.
Our proof proceeds by noticing that the sieve LS regression estimator
obtains as a special case of the NPIV estimator by setting $Y_2 = X$, $\psi^J = b^K$, $J = K$ and $\gamma = \gamma_x$. In this setting, the quantity $P_n h_0(x)$ just reduces to the orthogonal projection of $h_0$ onto the sieve space $B_K$ under the inner product induced by the empirical distribution, viz.
Moreover, in this case the $J \times K$ matrix $S$ defined in ((ref)) reduces to the $K \times K$ identity matrix $I_K$ and its smallest singular value is unity (whence $\sigma_{JK} = 1$). Therefore, the general calculation presented in Theorem (ref) can be used to control the “variance term” $\|\widehat h - P_n h_0\|_{\infty}$. The “bias term” $\|P_n h_0 - h_0\|_{\infty}$ is controlled as in Huang2003. It is worth emphasizing that no explicit weak dependence condition is placed on the regressors $\{X_i\}_{i=-\infty}^\infty$. Instead, this is implicitly captured by Condition (ii) on convergence of $\widetilde B'\widetilde B/n - I_K$.
Condition (ii) is satisfied by applying Lemma (ref) for i.i.d. data and Lemma (ref) for weakly dependent data. Theorem (ref) shows that spline and wavelet sieve LS estimators can achieve this minimax lower bound for weakly dependent data.
Corollary (ref) states that for i.i.d. data, Stone's optimal sup-norm convergence rate is achieved by spline and wavelet LS estimators whenever $\delta \geq d/p$ and $d \leq 2p$ (Assumption (ref)). If the regressors are exponentially $\beta$-mixing the optimal rate of convergence is achieved with $\delta \geq d/p$ and $d < 2p$. The restrictions $\delta \geq d/p$ and $(2+\gamma)d < 2 \gamma p$ for algebraically $\beta$-mixing (at a rate $\gamma$) reduces naturally towards the exponentially mixing conditions as the dependence becomes weaker (i.e. $\gamma$ becomes larger). In all cases, a smoother function (i.e., bigger $p$) means a lower value of $\delta$, and therefore heaver-tailed error terms $\epsilon_i$, are permitted while still obtaining the optimal sup-norm convergence rate. In particular this is achieved with $\delta =d/p \leq 2$ for i.i.d. data. Recently, BCK2013 require that the conditional $(2+\eta)$th moment (for some $\eta >0$) of $\epsilon_i$ be uniformly bounded for spline LS regression estimators to achieve the optimal sup-norm rate for i.i.d. data.\footnote{Chen would like to thank Jianhua Huang for working together on an earlier draft that does achieve the optimal sup-norm rate for a polynomial spline LS estimator with i.i.d. data, but under a stronger condition that $E[\epsilon^4_i |X_i=x]$ is uniformly bounded in $x$.} Uniform convergence rates of series LS estimators have also been studied by Newey1997, deJong2002, Song2008, LeeRobinson and others, but the sup-norm rates obtained in these papers are slower than the minimax risk lower bound in sup-norm loss of Stone1982.\footnote{See, e.g., hansen-ET, masry, CattaneoFarrell and the references therein for the optimal sup-norm convergence rates of a conditional mean function via the kernel, local linear regression and partitioning estimators of a conditional mean function.} Our result is the first such optimal sup-norm rate result for a sieve nonparametric LS estimator allowing for weakly-dependent data with heavy-tailed error terms. It should be very useful for nonparametric estimation of financial time-series models that have heavy-tailed error terms.
In this subsection a Bernstein inequality for sums of independent random matrices due to Tropp2012 is adapted to obtain convergence rates for sums of random matrices formed from $\beta$-mixing (absolutely regular) sequences, where the dimension, norm, and variance measure of the random matrices are allowed to grow with the sample size. These inequalities are particularly useful for establishing convergence rates for semi/nonparametric sieve estimators with weakly-dependent data. We first recall a result of Tropp2012.
We now provide a version of Theorem (ref) and Corollary (ref) for matrix-valued functions of $\beta$-mixing sequences. The $\beta$-mixing coefficient between two $\sigma$-algebras $ \mathcal{A}$ and $\mathcal{B}$ is defined as
with the supremum taken over all finite partitions $\{A_i\}_{i \in I}\subset \mathcal{A}$ and $\{B_J\}_{j \in J} \subset \mathcal{B}$ \citep*{DoukhanMassartRio}. The $q$th $\beta$-mixing coefficient of $ \{X_i\}_{i=-\infty}^\infty$ is defined as
The process $\{X_i\}_{i=-\infty}^\infty$ is said to be algebraically $ \beta$-mixing at rate $\gamma$ if $q^\gamma \beta(q) = o(1)$ for some $ \gamma > 1$, and geometrically $\beta$-mixing if $\beta(q) \leq c \exp(-\gamma q)$ for some $\gamma > 0$ and $c \geq 0$. The following extension of Theorem (ref) is made using a Berbee's lemma and a coupling argument (see, e.g., DoukhanMassartRio).
This subsection provides a readily verifiable condition under which, with probability approaching one (wpa1), the theoretical and empirical $L^2$ norms are equivalent over a linear sieve space. This equivalence, referred to by Huang2003 as empirical identifiability, has several applications in nonparametric sieve estimation. In the context of nonparametric series regression, empirical identifiability ensures the estimator is the orthogonal projection of $Y$ onto the sieve space under the empirical inner product and is uniquely defined Huang2003. Empirical identifiability is also used to establish the large-sample properties of sieve conditional moment estimators ChenPouzo2012. A sufficient condition for empirical identifiability is now cast in terms of convergence of a random matrix, which we verify for i.i.d. and $\beta$-mixing sequences.
A subspace $\mathcal{A }\subseteq L^2(X)$ is said to be empirically identifiable if $\frac{1}{n} \sum_{i=1}^n b(X_i)^2 = 0$ implies $b = 0$ a.e.-$[F_X]$ where $F_X$ dentoes the distribution of $X$. A sequence of spaces $\{\mathcal A_K : K \geq 1\} \subseteq L^2(X)$ is empirically identifiable wpa1 as $K = K(n) \to \infty$ with $n$ if
for any $t > 0$. Huang1998 uses a chaining argument to provide sufficient conditions for ((ref)) over the linear space $B_K$ under i.i.d. sampling. ChenPouzo2012 use this argument to establish convergence of sieve conditional moment estimators. Although easy to establish for i.i.d. sequences, it may be difficult to verify ((ref)) via chaining arguments for certain types of weakly dependent sequences. To this end, the following is a readily verifiable sufficient condition for empirical identifiability for linear sieve spaces. Let $B_K = clsp\{b_{K1},\ldots,b_{KK}\}$ denote a general linear sieve space and let $\widetilde B = (\widetilde b^K(X_1),\ldots,\widetilde b^K(X_n))'$ where $\widetilde b^K(x)$ is the orthonormalized vector of basis functions.
Condition (ref) is a sufficient condition for ((ref)) with a linear sieve space $B_K$. It should be noted that convergence is only required in the spectral norm. In the i.i.d. case this allows for $K$ to increase more quickly with $n$ than is achievable under the chaining argument of Huang1998. Let
as in Newey1997. Under regularity conditions, $\zeta_0(K) = O(\sqrt K)$ for tensor products of splines, trigonometric polynomials or wavelets and $\zeta_0(K) = O(K)$ for tensor products of power series or polynomials Newey1997,Huang1998. Under the chaining argument of Huang1998, ((ref)) is achieved under the restriction $\zeta_0(K)^2K/n = o(1)$. Huang2003 relaxes this restriction to $K (\log n)/n = o(1)$ for a polynomial spline sieve. We now generalize this result by virtue of Lemma (ref) and exponential inequalities for sums of random matrices.
The following lemma is useful to provide sufficient conditions for empirical identifiability for $\beta$-mixing sequences, which uses Theorem (ref).