EconBase
← Back to paper

Optimal Sup-norm Rates and Uniform Inference on Nonlinear Functionals of Nonparametric IV Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

121,549 characters · 17 sections · 117 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimal Sup-norm Rates and Uniform Inference on Nonlinear Functionals of Nonparametric IV Regression

\defaultbibliography{supgen-2} \defaultbibliographystyle{chicago}

bibunit\begin{abstract} \singlespacing This paper makes several important contributions to the literature about nonparametric instrumental variables (NPIV) estimation and inference on a structural function $h_0$ and its functionals. First, we derive sup-norm convergence rates for computationally simple sieve NPIV (series 2SLS) estimators of $h_0$ and its derivatives. Second, we derive a lower bound that describes the best possible (minimax) sup-norm rates of estimating $h_0$ and its derivatives, and show that the sieve NPIV estimator can attain the minimax rates when $h_0$ is approximated via a spline or wavelet sieve. Our optimal sup-norm rates surprisingly coincide with the optimal root-mean-squared rates for severely ill-posed problems, and are only a logarithmic factor slower than the optimal root-mean-squared rates for mildly ill-posed problems. Third, we use our sup-norm rates to establish the uniform Gaussian process strong approximations and the score bootstrap uniform confidence bands (UCBs) for collections of nonlinear functionals of $h_0$ under primitive conditions, allowing for mildly and severely ill-posed problems. Fourth, as applications, we obtain the first asymptotic pointwise and uniform inference results for plug-in sieve t-statistics of exact consumer surplus (CS) and deadweight loss (DL) welfare functionals under low-level conditions when demand is estimated via sieve NPIV. Empiricists could read our real data application of UCBs for exact CS and DL functionals of gasoline demand that reveals interesting patterns and is applicable to other markets. Keywords: Series 2SLS; Optimal sup-norm convergence rates; Uniform Gaussian process strong approximation; Score bootstrap uniform confidence bands; Nonlinear welfare functionals; Nonparametric demand with endogeneity. JEL codes: C13, C14, C36 \end{abstract} \pagenumbering{arabic}

Introduction

Well founded empirical evaluation of economic policy is often based upon inference on nonlinear welfare functionals of nonparametric or semiparametric structural models. This paper makes several important contributions to estimation and inference on a flexible (i.e. nonparametric) structural function $h_0$ and nonlinear functionals of $h_0$ within the framework of a nonparametric instrumental variables (NPIV) model:

equation[equation omitted — 77 chars of source]

where $h_0$ is an unknown function, $X_i$ is a vector of continuous endogenous regressors, $W_i$ is a vector of (conditional) instrumental variables, and the conditional distribution of $X_i$ given $W_i$ is unspecified.

Given a random sample $\{(Y_i, X_i, W_i)\}_{i=1}^n$ (of size $n$) from the NPIV model ((ref)), our first two main theoretical results address how well one may estimate $h_0$ and its derivatives simultaneously in sup-norm loss, i.e. we bound \[ \sup_x \Big|\wh h(x) - h_0(x)\Big| \quad \quad \mbox{and} \quad \quad \sup_x \Big|\partial^k \wh h(x) - \partial^k h_0(x)\Big| \] for estimators $\wh h$ of $h_0$, where $\partial^k h(x)$ denotes $k$-th partial derivatives of $h$ with respect to components of $x$. We first provide upper bounds on sup-norm convergence rates for the computationally simple sieve NPIV (i.e., series two stage least squares (2SLS)) estimators NeweyPowell,AiChen2003,BCK. We then derive a lower bound that describes the best possible (i.e., minimax) sup-norm convergence rates among all estimators for $h_0$ and its derivatives, and show that the sieve NPIV estimator can attain the minimax lower bound when spline or wavelet base is used to approximate $h_0$.\footnote{The optimal sup-norm rates for estimating $h_0$ were in the first version ChenChristensen-npiv. The optimal sup-norm rates for estimating derivatives of $h_0$ were in the second version ChenChristensen-adaptive-npiv.} Next, we apply our sup-norm rate results to establish the uniform Gaussian process strong approximation and the validity of score bootstrap uniform confidence bands (UCBs) for collections of possibly nonlinear functionals of $h_0$ under primitive conditions.\footnote{The uniform strong approximation and the score bootstrap UCBs results were in the second version ChenChristensen-adaptive-npiv; see Theorem B.1 and its proof in that version.} This includes valid score bootstrap UCBs for $h_0$ and its derivatives as special cases. Finally, as important applications, we establish first pointwise and uniform inference results for two leading nonlinear welfare functionals of a nonparametric demand function $h_0$ estimated via sieve NPIV, namely the exact consumer surplus (CS) and deadweight loss (DL) arising from price changes at different income levels when prices (and possibly income) are endogenous.\footnote{The pointwise inference results on exact CS and DL were in the second version ChenChristensen-adaptive-npiv.} We present two real data applications to illustrate the easy implementation and usefulness of the score bootstrap UCBs based on sieve NPIV estimators. The first is to nonparametric exact CS and DL functionals of gasoline demand and the second is to nonparametric Engel curves and their derivatives. The UCBs reveal new interesting and sensible patterns in both data applications. We note that the score bootstrap UCBs for exact CS and DL nonlinear functionals are new to the literature even when the prices might be exogenous. Empiricists could jump to Section (ref) to read the sieve score bootstrap UCBs procedure and these real data applications without the need to read the rest of more theoretical sections.

Regardless of whether the regressor $X_i$ is endogenous or not, sup-norm convergence rates provide sharper measures of how well $h_0$ and its derivatives can be estimated nonparametrically than the usual $L^2$-norm (i.e., root-mean-squared) rates. This is also why, in the existing literature on nonparametric models without endogeneity, consistent specification tests in sup-norm (i.e., Kolmogorov-Smirnov type statistics) are widely used. Further, sup-norm rates are particularly useful for controlling nonlinearity bias when conducting inference on highly nonlinear (i.e., beyond quadratic) functionals of $h_0$. In addition to being useful in constructing pointwise and uniform confidence bands for nonlinear functionals of $h_0$ via plug-in estimators, the sup-norm rates for estimating $h_0$ are also useful in semiparametric two-step procedures when $h_0$ enters the second-stage moment conditions (equalities or inequalities) nonlinearly.

Despite the usefulness of sup-norm convergence rates in nonparametric estimation and inference, as yet there are no published results on optimal sup-norm convergence rates for estimating $h_0$ or its derivatives in the NPIV model ((ref)). This is because, unlike nonparametric least squares (LS) regression (i.e. estimation of $h_0 (x) = E[Y_i |X_i =x]$ when $X_i$ is exogenous), estimation of $h_0$ in the NPIV model ((ref)) is a difficult ill-posed inverse problem with an unknown operator NeweyPowell,CFR2007. Intuitively, $h_0$ in model ((ref)) is identified by the integral equation \[ E[Y_i | W_i =w] =Th_0 (w) :=\int h_0 (x) f_{X|W}(x|w)\, \mr dx \] where $T$ must be inverted to obtain $h_0$. Since integration smoothes out features of $h_0$, a small error in estimating $E[Y_i | W_i =w]$ using the data $\{(Y_i, X_i, W_i)\}_{i=1}^n$ may lead to a large error in estimating $h_0$. In addition, the conditional density $f_{X|W}$ and hence the operator $T$ is generally unknown, so $T$ must be also estimated from the data. Due to the difficult ill-posed inverse nature, even the $L^2$-norm convergence rates for estimating $h_0$ in model ((ref)) have not been established until recently.\footnote{See e.g., HallHorowitz,BCK,ChenReiss,DFFR,Horowitz2011,ChenPouzo2012,GagliardiniScaillet,FlorensSimoni,Kato2013 and references therein.} In particular, HallHorowitz derived minimax $L^2$-norm convergence rates for mildly ill-posed NPIV models and showed that their estimators can attain the optimal $L^2$-norm rates for $h_0$. ChenReiss derived minimax $L^2$-norm convergence rates for mildly and severely ill-posed NPIV models and showed that sieve NPIV estimators can attain the optimal rates.\footnote{Appendix (ref) extends the results in ChenReiss to $L^2$-norm optimality for estimating derivatives of $h_0$.} Moreover, it is generally much harder to obtain optimal nonparametric convergence rates in sup-norm than in $L^2$-norm.\footnote{Even for the simple nonparametric LS regression of $h_0$ (without endogeneity), the optimal sup-norm rates for series LS estimators of $h_0$ were not obtained till recently in CattaneoFarrell for locally partitioning series LS, BCCK2014 for spline LS and ChenChristensen-reg for wavelet LS.}

In this paper, we derive the best possible (i.e., minimax) sup-norm convergence rates of any estimator of $h_0$ and its derivatives in mildly and severely ill-posed NPIV models. Surprisingly, the optimal sup-norm convergence rates for estimating $h_0$ and its derivatives coincide with the optimal $L^2$-norm rates for severely ill-posed problems and are only a power of $\log n$ slower than optimal $L^2$-norm rates for mildly ill-posed problems. We also obtain sup-norm convergence rates for sieve NPIV estimators of $h_0$ and its derivatives. We show that a sieve NPIV estimator using a spline or wavelet basis to approximate $h_0$ can attain the minimax sup-norm rates for estimating both $h_0$ and its derivatives. When specializing to series LS regression (without endogeneity), our results automatically imply that spline and wavelet series LS estimators will also achieve the optimal sup-norm rates of Stone1982 for estimating the derivatives of a nonparametric LS regression function, which strengthen the recent sup-norm optimality results in BCCK2014 and ChenChristensen-reg for estimating regression function $h_0$ itself. We focus on the sieve NPIV estimator because it has been used in empirical work, can be implemented as easily as 2SLS, and can reduce to simple series LS when the regressor $X_i$ is exogenous. Moreover, both $h_0$ and its derivatives may be simultaneously estimated at their respectively optimal convergence rates via a sieve NPIV estimator when the same sieve dimension is used to approximate $h_0$. This is a desirable property to practitioners. In addition, the sieve NPIV estimator for $h_0$ in model ((ref)) and our proof of its sup-norm rates could be easily extended to estimating unknown functions in other semiparametric models with nonparametric endogeneity, such as a system of shape-invariant Engel curve IV regression model BCK.

We provide two important applications of our results on sup-norm convergence rates in details; both are about inferences on nonlinear functionals of $h_0$ based on plug-in sieve NPIV estimators; see Section (ref) for discussions of additional applications. Inference on highly nonlinear (i.e., beyond quadratic) functionals of $h_0$ in a NPIV model is very difficult because of the combined effects of nonlinearity bias and the slow convergence rates (in sup-norm and $L^2$-norm) of any estimators of $h_0$. Indeed, our minimax rate results show that any estimator of $h_0$ in an ill-posed NPIV model must necessarily converge slower than their nonparametric LS counterpart. For example, the optimal sup- and $L^2$-norm rates for estimating $h_0$ in a severely ill-posed NPIV model is $(\log n )^{-\gamma}$ for some $\gamma>0$. It is well-known that a plug-in series LS estimate of a weighted quadratic functional could be root-$n$ consistent. But, a plug-in sieve NPIV estimate of a weighted quadratic functional of $h_0$ in a severely ill-posed NPIV model fails to be root-$n$ consistent ChenPouzo2014. In fact, we establish the minimax convergence rate of any estimators of a simple weighted quadratic functional of $h_0$ in a severely ill-posed NPIV model is as slow as $(\log n )^{-a}$ for some $a>0$ (see Appendix (ref)).

In the first application, we extend the seminal work of HausmanNewey1995 about pointwise inference on exact CS and DL functionals of nonparametric demand without endogeneity to allow for prices, and possibly incomes, to be endogenous. According to Hausman1981 and HausmanNewey1995,HausmanNewey2016,HausmanNewey2017, exact CS and DL functionals are the most widely used welfare and economic efficiency measures. Exact CS is a leading example of a complicated nonlinear functional of $h_0$, which is defined as the solution to a differential equation involving a demand function Hausman1981. HausmanNewey1995 were the first to establish the pointwise asymptotic normality of plug-in kernel estimators of exact CS and DL functionals of a nonparametric demand without endogeneity. Vanhems2010 was the first to estimate exact CS via the plug-in HallHorowitz kernel NPIV estimator of $h_{0}$ when price is endogenous, and derived its convergence rate in $L^2$-norm for the mildly ill-posed case, but did not establish any inference results (such as the pointwise asymptotic normality). Our paper is the first to provide low-level sufficient conditions to establish inference results for plug-in (spline and wavelet) sieve NPIV estimators of exact CS and DL functionals, allowing for both mildly and severely ill-posed NPIV models. Precisely, we use our sup-norm convergence rates for sieve NPIV estimators of $h_0$ and its derivatives to locally linearize plug-in estimators of exact CS and DL, which then leads to asymptotic normality of sieve $t$-statistics for exact CS and DL under primitive sufficient conditions. We also establish the asymptotic normality of plug-in sieve NPIV $t$-statistic for an approximate CS functional, extending Newey1997's result from nonparametric exogenous demand to endogenous demand. Recently, ChenPouzo2014 presented a set of high-level conditions for the pointwise asymptotic normality of sieve $t$-statistics of possibly nonlinear functionals of $h_0$ in a general class of nonparametric conditional moment restriction models (including the NPIV model as a special case). They verified their high-level conditions for pointwise asymptotic normality of sieve $t$-statistics for linear and quadratic functionals. But, without sup-norm convergence rate result, ChenPouzo2014 were unable to provide low-level sufficient conditions for pointwise asymptotic normality of plug-in sieve NPIV estimators for complicated nonlinear (beyond quadratic) functionals such as the exact CS functional. This was actually the original motivation for us to derive sup-norm convergence rates for sieve NPIV estimators of $h_0$ and its derivatives.

In the second important application of our sup-norm rate results, we establish the uniform Gaussian process strong approximation and the validity of score bootstrap uniform confidence bands (UCBs) for collections of possibly nonlinear functionals of $h_0$, under primitive sufficient conditions that allow for mildly and severely ill-posed NPIV models. The low-level sufficient conditions for Gaussian process strong approximation and UCBs are applied to complicated nonlinear functionals such as collections of exact CS and DL functionals of nonparametric demand with endogenous price (and possibly income). When specializing to collections of linear functionals of the NPIV function $h_0$, our Gaussian process strong approximation and sieve score bootstrap UCBs for $h_0$ and its derivatives are valid under mild sufficient conditions. In particular, for a NPIV model with a scalar endogenous regressor, our sufficient conditions are comparable to those in HorowitzLee2012 for their notion of UCBs with a growing number of grid points by interpolation for $h_0$ estimated via the modified orthogonal series NPIV estimator of Horowitz2011. When specializing to a nonparametric LS regression (with exogenous $X_i$), our results on the Gaussian strong approximation and score bootstrap UCBs for collections of nonlinear functionals of $h_0$, such as exact CS and DL functionals, are still new to the literature and complement the important results in CLR for $h_0$ and BCCK2014 for linear functionals of $h_0$ estimated via series LS.

Our sieve score bootstrap UCBs procedure is extremely easy to implement since it computes the sieve NPIV estimator only once using the data, and then perturbs the sieve score statistics by random weights that are mean zero and independent of the data. So it should be very useful to empirical researchers who conduct nonparametric estimation and inference on structural functions with endogeneity in diverse subfields of applied economics, such as consumer theory, IO, labor economics, public finance, health economics, development and trade, to name only a few. Two real data illustrations are presented in Section (ref). In the first, we construct UCBs for exact CS and DL welfare functionals for a range of gasoline taxes at different income levels. For this illustration, we use the same data set as in BlundellHorowitzParey2012,BlundellHorowitzParey2013 and estimate household gasoline demand via spline sieve NPIV (other data sets and other goods could be used). Despite the slow convergence rates of NPIV estimators, the UCBs for exact CS are particularly informative. In the second empirical illustration, we use the same data set as in BCK to estimate Engel curves for households with kids via a spline sieve NPIV and construct UCBs for Engel curves and their derivatives for various categories of household expenditure.

The rest of the paper is organized as follows. Section (ref) presents the sieve NPIV estimator, the score bootstrap UCBs procedure and two real-data applications. This section aims at empirical researchers. Section (ref) establishes the minimax optimal sup-norm rates for estimating a NPIV function $h_0$ and its derivatives. Section (ref) presents low-level sufficient conditions for the uniform Gaussian process strong approximation and sieve score bootstrap UCBs for collections of general nonlinear functionals of a NPIV function. Section (ref) deals with pointwise and uniform inferences on exact CS and DL, and approximate CS functionals in nonparametric demand estimation with endogeneity. Section (ref) concludes with discussions of additional applications of the sup-norm rates of sieve NPIV estimators. Appendix (ref) contains additional results on sup-norm convergence rates. Appendix (ref) presents optimal $L^2$-norm rates for estimating derivatives of a NPIV function under extremely weak conditions. Appendix (ref) establishes the minimax lower bounds for estimating quadratic functionals of a NPIV function. The main online supplementary appendix contains pointwise normality of sieve $t$ statistics for nonlinear functionals of NPIV under lower-level sufficient conditions than those in ChenPouzo2014 (Appendix (ref)); background material on B-spline and wavelet sieves (Appendix (ref)); and useful lemmas on random matrices (Appendix (ref)). The secondary online appendix contains additional lemmas and all of the proofs (Appendix (ref)).

Estimator and motivating applications to UCBs

This section describes the sieve NPIV estimator and a score bootstrap UCBs procedure for collections of functionals of the NPIV function. It mentions intuitively why sup-norm convergence rates of a sieve NPIV estimator are needed to formally justify the validity of the computationally simple score bootstrap UCBs procedure. It then present two real data applications of uniform inferences on functionals of a NPIV function: UCBs for exact CS and DL functionals of nonparametric demand with endogenous price, and UCBs for nonparametric Engel curves and their derivatives when the total expenditure is endogenous. This section is presented to practitioners.

Sieve NPIV estimators. Let $\{(Y_i,X_i,W_i)\}_{i=1}^n$ denote a random sample from the NPIV model ((ref)). The sieve NPIV estimator $\wh h$ of $h_0$ is simply the 2SLS estimator applied to some basis functions of $X_i$ (the endogenous regressors) and $W_i$ (the conditioning variables), namely

equation[equation omitted — 137 chars of source]

where ${Y} = (Y_1,\ldots,Y_n)'$,

align[align omitted — 219 chars of source]

and $\{\psi_{J1},\ldots,\psi_{JJ}\}$ and $\{b_{K1},\ldots,b_{KK}\}$ are collections of basis functions of dimension $J$ and $K$ for approximating $h_0$ and the instrument space, respectively BCK,ChenPouzo2012,Newey2013). The regularization parameter $J$ is the dimension of the sieve for approximating $h_0$. The smoothing parameter $K$ is the dimension of the instrument sieve. From the analogy with 2SLS, it is clear that we need $K \geq J$. BCK,ChenReiss,ChenPouzo2012 have previously shown that $\lim_{J} (K/J) = c \in [1,\infty)$ can lead to the optimal $L^2$-norm convergence rate for sieve NPIV estimator. Thus we assume that $K$ grows to infinity at the same rate as that of $J$, say $J \leq K \leq c J$ for some finite $c>1$ for simplicity.\footnote{Monte Carlo evidences in BCK,ChenPouzo2014 and others suggest that sieve NPIV estimators often perform better with $K > J$ than with $K = J$, and that the regularization parameter $J$ is important for finite sample performance while the parameter $K$ is not as important as long as it is larger than $J$. See our second version ChenChristensen-adaptive-npiv for data-driven choice of $J$.} When $K=J$ and $b^K =\psi^J$ being an orthogonal series basis, the sieve NPIV estimator becomes Horowitz2011's modified orthogonal series NPIV estimator. Note that the sieve NPIV estimator ((ref)) reduces to a series LS estimator $\wh h(x) = \psi^J(x)' [\Psi'\Psi]^- \Psi'{Y}$ when $X_i =W_i$ is exogenous, $J=K$ and $\psi^J(x)=b^K(w)$ Newey1997,Huang1998.

Uniform confidence bands for nonlinear functionals

One important motivating application is to uniform inference on a collection of nonlinear functionals $\{f_t(h_0) : t \in \mc T\}$ where $\mc T$ is an index set (e.g. an interval). Uniform inference may be performed via uniform confidence bands (UCBs) that contain the function $t \mapsto f_t(h_0)$ with prescribed coverage probability. UCBs for $h_0$ (or its derivatives) are obtained as a special case with $\mc T = \mc X$ (support of $X_i$) and $f_t(h_0) = h_0(t)$ (or $f_t(h_0) = \partial^k h_0(t)$ for $k$-th derivative). We present applications below to uniform inference on exact CS and DL functionals over a range of price changes as well as UCBs for Engel curves and their derivatives.

A $100(1-\alpha)\%$ bootstrap-based UCB for $\{f_t(h_0) : t \in \mc T\}$ is constructed as

equation[equation omitted — 197 chars of source]

In this display $f_t(\wh h)$ is the plug-in sieve NPIV estimator of $f_t(h_0)$, $\wh\sigma^2 (f_t)$ is a sieve variance estimator for $f_t(\wh h)$, and $z_{1-\alpha}^*$ is a bootstrap-based critical value to be defined below.

To compute the sieve variance estimator for $f_t(\widehat h)$ with $\wh h(x) = \psi^J(x)'\wh c$ given in ((ref)), one would first compute the 2SLS covariance matrix estimator (but applied to basis functions) for $\wh c$:

equation[equation omitted — 159 chars of source]

where $\wh S = B'\Psi/n$, $\wh G_b = B'B/n$, $\wh \Omega = n^{-1} \sum_{i=1}^n \wh u_i^2 b^K(W_i) b^K(W_i)'$ and $\wh u_i = Y_i - \wh h(X_i)$. One then compute a “delta-method” correction term, a $J \times 1$ vector $Df_t(\wh h)[\psi^J]:=\big(Df_t(\wh h)[\psi_{J1}],\ldots,Df_t(\wh h)[\psi_{JJ}]\big)'$, by calculating $Df_t(\wh h)[v] = \lim_{\delta \rightarrow 0^{+}} [\delta^{-1} f_t(\wh h +\delta v)]$, which is the (functional directional) derivative of $f_t$ at $\wh h$ in direction $v$, for $v=\psi_{J1},\ldots,\psi_{JJ}$. The sieve variance estimator for $f_t(\widehat h)$ is then

equation[equation omitted — 134 chars of source]

We use the following sieve score bootstrap procedure to calculate the critical value $z_{1-\alpha}^*$. Let $\varpi_{1},\ldots,\varpi_{n}$ be IID random variables independent of the data with mean zero, unit variance and finite 3rd moment, e.g. $N(0,1)$.\footnote{Other examples of distributions with these properties include the re-centered exponential (i.e. $\varpi_i = \mathrm{Exp}(1)-1$), Rademacher (i.e. $\pm 1$ each with probability $\frac{1}{2}$), or the two-point distribution of Mammen1993 (i.e. $(1-\sqrt 5)/2$ with probability $(\sqrt 5 + 1)/(2 \sqrt 5)$ and $(\sqrt 5+1)/\sqrt 2$ with remaining probability.} We define the bootstrap sieve $t$-statistic process $\{\mb Z_n^*(t) : t \in \mathcal T\}$ as

equation[equation omitted — 265 chars of source]

To compute $z_{1-\alpha}^*$, one would calculate $\sup_{t \in \mc T} |\mb Z_n^*(t)|$ for a large number of independent draws of $\varpi_{1},\ldots,\varpi_{n}$. The critical value $z_{1-\alpha}^*$ is the $(1-\alpha)$ quantile of $\sup_{t \in \mc T} |\mb Z_n^*(t)|$ over the draws. Note that this sieve score bootstrap procedure is different from the usual nonparametric bootstrap (based on resampling the data, then recomputing the estimator): here we only compute the estimator once, and then perturb the sieve $t$-statistic process by the innovations $\varpi_1,\ldots,\varpi_n$.

An intuitive description of why sup-norm rates are very useful to justify this procedure is as follows. Under regularity conditions, the sieve $t$-statistic for an individual functional $f_t(h_0)$ admits an expansion

equation[equation omitted — 165 chars of source]

(see equation (ref)) for the definition of $\widehat{ \mb Z}_n (t)$). The term $\widehat{ \mb Z}_n (t)$ is a CLT term, i.e. $\widehat{ \mb Z}_n (t) \to_d N(0,1)$ for each fixed $t \in \mc T$. Therefore, the sieve $t$-statistic for $f_t(h_0)$ also converges to a $N(0,1)$ random variable provided that the “nonlinear remainder term” is asymptotically negligible (i.e. $o_p(1)$) (see Assumption 3.5 in ChenPouzo2014). Our sup-norm rates are very useful for providing weak regularity conditions under which the remainder is $o_p(1)$ for fixed $t$.\footnote{ChenPouzo2014 verified their high-level Assumption 3.5 for a plug-in sieve estimator of a weighted quadratic functional example. Without sup-norm convergence rates, it is difficult to verify their Assumption 3.5 for nonlinear functionals (such as the exact CS) that are more complicated than quadratic functionals.} This justifies constructing confidence intervals for individual functionals $f_t(h_0)$ for any fixed $t \in \mc T$ by inverting the sieve $t$-statistic (on the left-hand side of display ((ref))) and using $N(0,1)$ critical values. However, for uniform inference the usual $N(0,1)$ critical values are no longer appropriate as we need to consider the sampling error in estimating the whole process $t \mapsto f_t(h_0)$. For this purpose, display ((ref)) is strengthened to be valid uniformly in $t \in \mc T$ (see Lemma (ref)). Under some regularity conditions, $\sup_{t\in \mc T} |\widehat{\mb Z}_n(t)|$ converges in distribution to the supremum of a (non-pivotal) Gaussian process. As its critical values are generally not available, we use the sieve score bootstrap procedure to estimate its critical values.

Section (ref) formally justifies the use of this procedure for constructing UCBs for $\{f_t(h_0) : t \in \mc T\}$. The sup-norm rates are useful for controlling the nonlinear remainder terms for UCBs for collections of nonlinear functionals. Theorem (ref) appears to be the first to establish the consistency of sieve score bootstrap UCBs for general nonlinear functionals of NPIV under low-level conditions, allowing for mildly and severely ill-posed problems. It includes as special cases the score bootstrap UCBs for nonlinear functionals of $h_0$ under exogeneity when $h_0$ is estimated via series LS, and the score bootstrap UCBs for the NPIV function $h_0$ and its derivatives.\footnote{One also needs to use sup-norm convergence rates of $\wh h$ to $h_0$ to build a valid UCB for $\{h_0 (t): t \in \mc X \}$.} Theorem (ref) is applied in Section (ref) to formally justify the validity of score bootstrap UCBs for exact CS and DL functionals over a range of price changes when demand is estimated nonparametrically via sieve NPIV.

Empirical application 1: UCBs for nonparametric exact CS and DL functionals

Here we apply our methodology to study the effect of gasoline price changes on household welfare. We extend the important work by HausmanNewey1995 on pointwise confidence bands for exact CS and DL of demand without endogeneity to UCBs for exact CS and DL of demand with endogeneity.

Let demand of consumer $i$ be \[ \mf Q_i = h_0(\mf P_i,\mf Y_i) + \mf u_i \] where $\mf Q_i$ is quantity, $\mf P_i$ is price, which may be endogenous, $\mf Y_i$ is income of consumer $i$, and $\mf u_i$ is an error term.\footnote{Endogeneity may also be an issue in the estimation of static models of labor supply, in which $\mf Q_i$ represents hours worked, $\mf P_i$ is the wage, and $\mf Y_i$ is other income. In this setting it is reasonable to allow for endogeneity of both $\mf P_i$ and $\mf Y_i$ (see BlundellDuncanMeghir, BlundellMaCurdyMeghir, and references therein). } Hausman1981 shows that the exact CS from a price change from $\mf p^0$ to $\mf p^1$ at income level $\mf y$, denoted $\mf S_{\mf y}(\mf p^0)$, solves

equation[equation omitted — 280 chars of source]

where $\mf p : [0,1] \to \mb R$ is a twice continuously differentiable path with $\mf p(0) = \mf p^0$ and $\mf p(1) = \mf p^1$. The corresponding DL functional $\mf D_{\mf y}(\mf p^0)$ is

equation[equation omitted — 119 chars of source]

As is evident from ((ref)) and ((ref)), exact CS and DL are (typically nonlinear) functionals of $h_0$. An exception is when demand is independent of income, in which case exact CS and DL are linear functionals of $h_0$. Let $t = (\mf p^0,\mf p^1,\mf y)$ index the initial price, final price, and income level and let $\mc T \subseteq [\underline{\mf p}^0 , \overline{\mf p}^0] \times [\underline{\mf p}^1 , \overline{\mf p}^1] \times [\underline{\mf y}, \overline{\mf y}]$ denote a range of price changes and/or incomes over which inference is to be performed. To denote dependence on $h_0$, we use the notation

eqnarray[eqnarray omitted — 194 chars of source]

so $\mf S_{\mf y}(\mf p^0) = f_{CS,t}(h_0)$ and $\mf D_{\mf y}(\mf p^0) = f_{DL,t}(h_0)$.

We estimate exact CS and DL using the plug-in estimators $f_{CS,t}(\wh h)$ and $f_{DL,t}(\wh h)$. The sieve variance estimators $\wh \sigma^2 (f_{CS,t})$ and $\wh \sigma^2 (f_{DL,t})$ are as described in ((ref)) with the delta-method correction terms

eqnarray[eqnarray omitted — 386 chars of source]

where $\mf p'(u) = \frac{\mathrm d \mf p(u)}{\mathrm d u}$, $\partial_2 h$ denotes the partial derivative of $h$ with respect to its second argument and $\widehat{\mf S}_{\mf y}(\mf p(u))$ denotes the solution to ((ref)) with $\wh h$ in place of $h_0$.

We use the 2001 National Household Travel Survey gasoline demand data from BlundellHorowitzParey2012,BlundellHorowitzParey2013.\footnote{We are grateful to Matthias Parey for sharing the dataset with us. We refer the reader to section 3 of BlundellHorowitzParey2012 for a detailed description of the data.} The main variables are annual household gasoline consumption (in gallons), average price (in dollars per gallon) in the county in which the household is located, household income, and distance from the Gulf coast to the capital of the state in which the household is located. Due to censoring, we consider the subset of households with incomes less than \$100,000 per year. To keep households somewhat homogeneous, we select household with incomes above \$25,000 per year (the 8th percentile), with at most 6 inhabitants, and 1 or 2 drivers. The resulting sample has size $n = 2753$.\footnote{We also exclude one household that reports 14,635 gallons; the next largest is 8089 gallons. Similar results are obtained using the full set of $n = 4811$ observations.} Table (ref) presents summary statistics.

table[table omitted — 413 chars of source]

We estimate the household gasoline demand function in levels via sieve NPIV using distance as instrument for price. To implement the estimator, we form $\Psi_J$ by taking a tensor product of quartic B-spline bases of dimension 5 for both price and income (so $J = 25$) and $B_K$ by taking a tensor product of quartic B-spline bases of dimension 8 for distance and 5 for income (so $K = 40$) with interior knots spaced evenly at quantiles.

We consider exact CS and DL resulting from price increases from $\mf p^0 \in [\$1.20,\$1.40]$ to $\mf p^1 = \$1.40$ at income levels of $\mf y = \$42,500$ (low) and $\mf y = \$72,500$ (high). We estimate exact CS at each initial price level by solving the ODE ((ref)) by backward differences. We construct UCBs for exact CS as described above by setting $\mc T = [\$1.20,\$1.40] \times \{\$1.40\} \times \{\$42,500\}$ for the low-income group and $\mc T = [\$1.20,\$1.40] \times \{\$1.40\} \times \{\$72,500\}$ for the high-income group, $f_t(h) = f_{CS,t}(h)$ from display ((ref)), and $D f_{t}(\wh h)[\psi^J] = D f_{CS,t}(\wh h)[\psi^J]$ from display ((ref)). The ODE ((ref)) is solved numerically by backward differences and the integrals in ((ref)) are computed numerically. UCBs for DL are formed similarly, $f_t(h) = f_{DL,t}(h)$ from display ((ref)), and $D f_{t}(\wh h)[\psi^J] = D f_{DL,t}(\wh h)[\psi^J]$ from display ((ref)). We draw the bootstrap innovations $\varpi_i$ from Mammen's two-point distribution with $1000$ bootstrap replications.

figure[figure omitted — 478 chars of source]
figure[figure omitted — 563 chars of source]

The exact CS and DL estimates are presented in Figure (ref) together with their UCBs. It is clear that exact CS is much more precisely estimated than DL. This is to be expected, since exact CS is computed by essentially integrating over one argument of the estimated demand function and is therefore smoother than the DL functional, which depends on $h_0$ estimated at the point $(\mf p^1,\mf y)$. In fact, even though the sieve NPIV $\wh h$ itself converges slowly, the UCBs for exact CS are still quite informative. At their widest point (with initial price \$1.20), the 95% UCBs for exact CS for low-income households are $[\$259,\$314]$. In terms of comparison across high- and low-income households, the exact CS estimates are higher for the high-income households whereas DL estimates are higher for the low-income households.

Figure (ref) displays estimates obtained when we treat price as exogenous and estimate demand ($h_0$) by series LS regression. This is a special case of the preceding analysis with $X_i = W_i = (\mf P_i,\mf Y_i)'$, $K = J$ and $\psi^J = b^K$. These estimates display several notable features. First, the exact CS estimates are very similar whether demand is estimated via series LS or via sieve NPIV. Second, the UCBs for exact CS estimates are of a similar width to those obtained when demand was estimated via sieve NPIV, even though NPIV is an ill-posed inverse problem whereas nonparametric LS regression is not. Third, the UCBs for DL are noticeably narrower when demand is estimated via series LS than when demand is estimated via sieve NPIV. Fourth, the DL estimates for LS and sieve NPIV are similar for high income households but quite different for low income households. This is consistent with BlundellHorowitzParey2013, who find some evidence of endogeneity in gasoline prices for low income groups.

Empirical application 2: UCBs for Engel curves and their derivatives

Engel curves describe the household budget share for expenditure categories as a function of total household expenditure. Following BCK, we use sieve NPIV to estimate Engel curves, taking log total household income as an instrument for log total household expenditure. We use data from the 1995 British Family Expenditure Survey, focusing on the subset of married or cohabitating couples with one or two children, with the head of household aged between 20 and 55 and in work. This leaves a sample of size $n = 1027$. We consider six categories of nondurables and services expenditure: food in, food out, alcohol, fuel, travel, and leisure.

We construct UCBs for Engel curves as described above by setting $\mc T = [4.75,6.25]$ (approximately the 5th to 95th percentile of log expenditure), $f_t(h) = h(t)$, and $D f_t(\wh h)[\psi^J] = \psi^J(t)$. We also construct UCBs for derivatives of the Engel curves by setting $\mc T = [4.75,6.25]$, $f_t(h)$ to be the derivative of $h$ evaluated at $t$, and $D f_t(\wh h)[\psi^J]$ to be the vector formed by taking derivatives of $\psi_{J1},\ldots,\psi_{JJ}$ evaluated at $t$. For both constructions, we use a quartic B-spline basis of dimension $J = 5$ for $\Psi_J$ and a quartic B-spline basis of dimension $K = 9$ for $B_K$, with interior knots evenly spaced at quantiles (an important feature of sieve estimators is that the same sieve dimension can be used for optimal estimation of the function and its derivatives; this is not the case for kernel-based estimators). We draw the bootstrap innovations $\varpi_i$ from Mammen's two-point distribution with $1000$ bootstrap replications.

figure[figure omitted — 396 chars of source]
figure[figure omitted — 330 chars of source]

The Engel curves presented in Figure (ref) and their derivatives presented in Figure (ref) exhibit several interesting features. The curves for food-in and fuel (necessary goods) are both downward sloping, with the curve for fuel exhibiting a pronounced downward slope at lower income levels. The derivative of the curve for fuel is negative, though the UCBs are positive at the extremities. In contrast, the curve for leisure expenditure (luxury good) is strongly upwards sloping and its derivative is positive except at low income levels. Remaining curves for food-out, alcohol and travel appear to be non-monotonic.

Optimal sup-norm convergence rates

This section presents several results on sup-norm convergence rates. Subsection (ref) presents upper bounds on sup-norm convergence rates of NPIV estimators of $h_0$ and its derivatives. Subsection (ref) presents (minimax) lower bounds. Subsection (ref) considers NPIV models with endogenous and exogenous regressors that are useful in empirical studies.

Notation: We work on a probability space $(\Omega,\mathcal F,\mb P)$. $\mathcal A^c$ denotes the complement of an event $\mathcal A \in \mathcal F$. We abbreviate “with probability approaching one” to “wpa1”, and say that a sequence of events $\{\mathcal A_n \} \subset \mathcal F$ holds wpa1 if $\mb P(\mathcal A_n^c) = o(1)$. For a random variable $X$ we define the space $L^q(X)$ as the equivalence class of all measurable functions of $X$ with finite $q$th moment if $1 \leq q < \infty$; when $q = \infty$ we denote $L^\infty(X)$ as the set of all bounded measurable functions $g : \mathcal X \to \mb R$ endowed with the sup norm $\|g\|_\infty = \sup_x |g(x)|$. Let $\langle\cdot,\cdot\rangle_X$ denote the inner product on $L^2(X)$. For matrix and vector norms, $\|\cdot\|_{\ell^q}$ denotes the vector $\ell^q$ norm when applied to vectors and the operator norm induced by the vector $\ell^q$ norm when applied to matrices. If $a$ and $b$ are scalars we let $a \vee b := \max\{a,b\}$ and $a \wedge b := \min\{a,b\}$. Minimum and maximum eigenvalues are denoted by $\lambda_{\min}$ and $\lambda_{\max}$. If $\{a_n\}$ and $\{b_n\}$ are sequences of positive numbers, we say that $a_n \lesssim b_n$ if $\limsup_{n \to \infty} a_n/b_n < \infty$ and we say that $a_n \asymp b_n$ if $a_n \lesssim b_n$ and $b_n \lesssim a_n$.

Sieve measure of ill-posedness. For a NPIV model ((ref)), an important quantity is the measure of ill-posedness which, roughly speaking, measures how much the conditional expectation $h \mapsto E[h(X_i)|W_i = w]$ smoothes out $h$. Let $T : L^2(X) \to L^2(W)$ denote the conditional expectation operator given by \[ T h(w) = E[h(X_i)|W_i = w]\,. \] Let $\Psi_J = clsp\{\psi_{J1},\ldots,\psi_{JJ}\} \subset L^2(X)$ and $B_K = clsp \{b_{K1},\ldots,b_{KK}\}\subset L^2(W)$ denote the sieve spaces for the endogenous variables and instrumental variables, respectively. Let $\Psi_{J,1} = \{h \in \Psi_J : \|h\|_{L^2(X)} = 1\}$. The sieve $L^2$ measure of ill-posedness is

equation*[equation* omitted — 155 chars of source]

Following BCK, we call a NPIV model ((ref)) with $X_i$ being a $d$-dimensional random vector: \\ (i) mildly ill-posed if $\tau_J = O(J^{\varsigma/d})$ for some $\varsigma > 0 $; and \\ (ii) severely ill-posed if $\tau_J = O(\exp(\frac{1}{2} J^{\varsigma /d}))$ for some $\varsigma > 0$.

See our second version ChenChristensen-adaptive-npiv for simple consistent estimation of the sieve measure of ill-posedness $\tau_J$.

Sup-norm convergence rates

We first introduce some basic conditions on the basic NPIV model ((ref)) and the sieve spaces.

assumption(i) $X_i$ has compact rectangular support $\mathcal X \subset \mb R^d$ with nonempty interior and the density of $X_i$ is uniformly bounded away from $0$ and $\infty$ on $\mathcal X$; (ii) $W_i$ has compact rectangular support $\mathcal W \subset \mb R^{d_w}$ and the density of $W_i$ is uniformly bounded away from $0$ and $\infty$ on $\mathcal W$; (iii) $T : L^2(X) \to L^2(W)$ is injective; and (iv) $h_0 \in \mathcal H \subset L^\infty(X)$, and ${\cup_J \Psi_J}$ is dense in $(\mathcal H,\|\cdot\|_{L^\infty (X)})$.
assumption(i) $\sup_{w\in \mathcal W} E[u_i^2|W_i = w] \leq \overline \sigma^2 < \infty$; and (ii) $E[|u_i|^{2+\delta}] < \infty$ for some $\delta > 0$.

The following assumptions concern the basis functions. Define \[

array[array omitted — 228 chars of source]

\] We assume throughout that the basis functions are not linearly dependent, i.e. $S$ has full column rank $J$ and $G_{\psi,J}$ and $G_{b,K}$ are positive definite for each $J$ and $K$, i.e. $e_J = \lambda_{\min}(G_{\psi,J}) >0$ and $e_{b,K} = \lambda_{\min}(G_{b,K}) >0$, although $e_J$ and $e_{b,K}$ could go to zero as $K\geq J$ goes to infinity. Let

align*[align* omitted — 236 chars of source]

for each $J$ and $K$ and define $\zeta = \zeta_J = \zeta_{b,K} \vee \zeta_{\psi,J}$. Note that $\zeta_{\psi,J}$ has some useful properties: $\|h\|_\infty \leq \zeta_{\psi,J} \|h\|_{L^2(X)}$ for all $h \in \Psi_J$, and $\sqrt J =(E[\|G_\psi^{-1/2} \psi^J(X)\|_{\ell^2}^2])^{1/2} \leq \zeta_{\psi,J} \leq \xi_{\psi,J} / \sqrt{e_J}$; clearly $\zeta_{b,K}$ has similar properties.

We say that the sieve basis for $\Psi_J$ is H\"older continuous if there exist finite constants $\omega \geq 0, \omega' >0$ such that $\|G_{\psi,J}^{-1/2}\{\psi^J(x) - \psi^J(x')\}\|_{\ell^2} \lesssim J^\omega \|x - x'\|_{\ell^2}^{\omega'}$ for all $x,x' \in \mathcal X$.

assumption(i) the basis spanning $\Psi_J$ is H\"older continuous; (ii) $\tau_J \zeta^2 /\sqrt n = O(1)$; and (iii) $\zeta^{(2+\delta)/\delta} \sqrt{(\log n)/n} = o(1)$.

Let $\Pi_J: L^2(X) \to \Psi_J$ denote the $L^2(X)$ orthogonal (i.e. least squares) projection onto $\Psi_J$, namely $\Pi_J h_0 = \mathrm{arg}\min_{h \in \Psi_J} \|h_0 - h\|_{L^2(X)}$ and let $\Pi_K :L^2(W) \to B_K$ denote the $L^2(W)$ orthogonal (i.e. least-squares) projection onto $B_K$. Let $Q_J h_0 = \mathrm{arg}\min_{h \in \Psi_J} \|\Pi_K T(h_0 - h)\|_{L^2(W)}$ denote the sieve 2SLS projection of $h_0$ onto $\Psi_J$. We may write $Q_J h_0 = \psi^J(\cdot)'c_{0,J}$ where \[ c_{0,J} = [S' G_b^{-1} S]^{-1} S' G_b^{-1} E[b^K(W_i)h_0 (X_i)]\,. \]

assumption(i) $\sup_{h \in \Psi_{J,1}}\|(\Pi_K T - T)h\|_{L^2(W)} = o(\tau_J^{-1})$; (ii) $\tau_J \times \|T(h_0 - \Pi_J h_0)\|_{L^2(W)} \leq \mathrm{const} \times \|h_0 - \Pi_J h_0\|_{L^2(X)}$; and (iii) $\|Q_J (h_0 - \Pi_J h_0) \|_\infty \leq O(1) \times \|h_0 - \Pi_J h_0\|_{\infty}$.

Discussion of Assumptions. Assumption (ref) is standard. Assumption (ref)(iii) is stronger than needed for convergence rates in sup-norm only. We impose it as a common sufficient condition for convergence rates in both sup-norm and $L^2$-norm (Appendix (ref)). For sup-norm convergence rate only, Assumption (ref)(iii) could be replaced by the following weaker identification condition:\\ Assumption (ref) (iii-sup) $h_0 \in \mathcal H \subset L^\infty(X)$, and $T[h-h_0]=0 \in L^2(W)$ for any $h\in \mathcal H$ implies that $\|h-h_0\|_\infty =0$. \\ This in turn is implied by the injectivity of $T:L^\infty(X)\to L^2(W)$ (or the bounded completeness), which is weaker than the injectivity of $T : L^2(X) \to L^2(W)$ (i.e., the $L^2$-completeness). Bounded completeness or $L^2$-completeness condition is often assumed in models with endogeneity (e.g. NeweyPowell,CFR2007,BCK,Andrews2011,CCLN) and is generically satisfied according to Andrews2011. The parameter space $\mathcal H$ for $h_0$ is typically taken to be a H\"older or Sobolev class of smooth functions. Assumption (ref)(i) could be relaxed to unbounded support, and the proofs need to be modified slightly using wavelet basis and weighted compact embedding results, see, e.g., BCK,ChenPouzo2012,Triebel2006 and references therein. To present the sup-norm rate results in a clean way we stick to the simplest Assumption (ref). Assumption (ref) is also imposed for sup-norm convergence rates for series LS regression under exogeneity (e.g., ChenChristensen-reg). Assumption (ref)(i) is satisfied by many commonly used sieve bases, such as splines, wavelets, and cosine bases. Assumption (ref)(ii)(iii) restrict the rate at which $J$ can grow with $n$. Upper bounds for $\zeta_{\psi,J}$ and $\zeta_{b,K}$ are known for commonly used bases. For instance, under Assumption (ref)(i)(ii), $\zeta_{b,K} = O(\sqrt K)$ and $\zeta_{\psi,J} = O(\sqrt J)$ for (tensor-product) polynomial spline, wavelet and cosine bases, and $\zeta_{b,K} = O(K)$ and $\zeta_{\psi,J} = O(J)$ for (tensor-product) orthogonal polynomial bases; see, e.g., Newey1997, Huang1998 and main online Appendix (ref). Assumption (ref)(i) is a mild condition on the approximation properties of the basis used for the instrument space and is similar to the first part of Assumption 5(iv) of Horowitz2014. In fact, $\|(\Pi_K T - T)h\|_{L^2(W)} = 0$ for all $h \in \Psi_J$ when the basis functions for $B_K$ and $\Psi_J$ form either a Riesz basis or eigenfunction basis for the conditional expectation operator. Assumption (ref)(ii) is the usual $L^2$ “stability condition” imposed in the NPIV literature (cf. Assumption 6 in BCK and Assumption 5.2(ii) in ChenPouzo2012). Assumption (ref)(iii) is a new $L^{\infty}$ “stability condition” to control the sup-norm bias. It turns out that Assumption (ref)(ii) and (ref)(iii) are also automatically satisfied by Riesz bases; see Appendix (ref) for further discussions and sufficient conditions.

To derive the sup-norm (uniform) convergence rate we split $\|\wh h - h_0\|_\infty$ into so-called “bias” and “standard deviation” terms and derive sup-norm convergence rates for the two terms. Specifically, let

equation*[equation* omitted — 137 chars of source]

where $H_0 = (h_0(X_1),\ldots,h_0(X_n))'$. We refer loosely to $\|\widetilde h - h_0\|_\infty$ as the “bias” term and $\|\wh h - \widetilde h\|_\infty$ as the “standard deviation” (or sometimes “variance”) term. Both are random quantities. We first bound the sup-norm “standard deviation” term in the following lemma.

lemmaLet Assumptions (ref)(i)(iii), (ref)(i)(ii), (ref)(ii)(iii), and (ref)(i) hold. Then:\\ (1) $\|\wh h - \widetilde h\|_\infty = O_p \big( \tau_J \xi_{\psi,J} \sqrt{ (\log J)/ (n e_J )} \big)$.\\ (2) If Assumption (ref)(i) also holds, then: $\|\wh h - \widetilde h\|_\infty = O_p \big( \tau_J \zeta_{\psi,J} \sqrt{(\log n)/n} \big)\,.$

Recall that $\sqrt{J} \leq \zeta_{\psi,J} \leq \xi_{\psi,J} / \sqrt{e_J}$. Result (2) of Lemma (ref) provides a slightly tighter upper bound on the variance term than Result (1) does, while Result (1) allows for slightly more general basis to approximate $h_0$. For splines and wavelets, we show in Appendix (ref) that $\xi_{\psi,J} / \sqrt{e_J} \lesssim \sqrt{J}$, so Results (1) and (2) produce the same tight upper bound $\|\wh h - \widetilde h\|_\infty = O_p(\tau_J \sqrt{(J\log n )/n})$ when $J \asymp n^r$ for some constant $r > 0$.

Before we present an upper bound on the “bias” term in Theorem (ref) part (1) below, we mention one more property of the sieve space $\Psi_J$ that is crucial for sharp bounds on the sup-norm bias term. Let $h_{0,J} \in \Psi_J$ denote the best approximation to $h_0$ in sup-norm, i.e. $h_{0,J}$ solves $\inf_{h \in \Psi_J}\|h_0 - h\|_\infty$. Then by Lebesgue's Lemma DeVoreLorentz: \[ \|h_0 - \Pi_J h_0 \|_\infty \leq (1 + \|\Pi_J \|_{\infty}) \times \|h_0 - h_{0,J}\|_\infty \] where $\|\Pi_J \|_{\infty}$ is the Lebesgue constant for the sieve $\Psi_J$. Recently it has been established that $\|\Pi_J\|_\infty \lesssim 1$ when $\Psi_J$ is spanned by a tensor product B-spline basis (Huang2003) or a tensor product Cohen-Daubechies-Vial (CDV) wavelet basis (ChenChristensen-reg).\footnote{See DeVoreLorentz and BCCK2014 for examples of other bases with bounded Lebesgue constant or with Lebesgue constant diverging slowly with the sieve dimension.} Boundedness of the Lebesgue constant is crucial for attaining optimal sup-norm rates.

theorem(1) Let Assumptions (ref)(iii), (ref)(ii) and (ref) hold. Then: \[ \|\widetilde h - h_0\|_\infty = O_p \left( \|h_0 - \Pi_J h_0\|_\infty \right) \,. \] (2) Let Assumptions (ref)(i)(iii)(iv), (ref)(i)(ii), (ref)(ii)(iii), and (ref) hold. Then: \[ \|\wh h - h_0\|_\infty = O_p \left( \|h_0 - \Pi_J h_0\|_\infty + \tau_J \xi_{\psi,J} \sqrt{ (\log J)/ (n e_J )} \right)\,. \] (3) Further, if the linear sieve $\Psi_J$ satisfies $\|\Pi_J\|_\infty \lesssim 1$ and $\xi_{\psi,J} / \sqrt{e_J} \lesssim \sqrt{J}$, then \[ \|\wh h - h_0\|_\infty = O_p \left( \|h_0 - h_{0,J}\|_\infty + \tau_J \sqrt{(J \log J)/n} \right)\,. \]

Theorem (ref)(2)(3) follows directly from part (1) (for bias) and Lemma (ref)(1) (for standard deviation). See Appendix (ref) for additional details about bound on sup-norm bias.

The following corollary provides concrete sup-norm convergence rates of $\wh h$ and its derivatives. To introduce the result, let $B^p_{\infty,\infty}$ denote the H\"older space of smoothness $p > 0$ and $\|\cdot\|_{B^p_{\infty,\infty}}$ denote its norm (see Section 1.11.10 of Triebel2006). Let $B_\infty(p,L) = \{h \in B^p_{\infty,\infty} : \|h\|_{B^p_{\infty,\infty}} \leq L\}$ denote a H\"older ball of smoothness $p > 0$ and radius $ L \in (0, \infty )$. Let $\alpha_1,\ldots,\alpha_d$ be non-negative integers, let $|\alpha| = \alpha_1 + \ldots + \alpha_d $, and define \[ \partial^\alpha h(x) := \frac{\partial ^{|\alpha|} h}{\partial^{\alpha_1} x_1 \cdots \partial ^{\alpha_d} x_d} h(x) \,. \] Of course, if $|\alpha| = 0$ then $\partial^\alpha h = h$.\footnote{If $|\alpha| > 0$ then we assume $h$ and its derivatives can be continuously extended to an open set containing $\mathcal X$.}

corollaryLet Assumptions (ref)(i)(ii)(iii) and (ref) hold. Let $h_0 \in B_\infty(p,L)$, $\Psi_J$ be spanned by a B-spline basis of order $\gamma > p$ or a CDV wavelet basis of regularity $\gamma > p$, $B_K$ be spanned by a cosine, spline or wavelet basis.\\ (1) If Assumption (ref)(ii) holds, then \[ \|\partial^\alpha \widetilde h - \partial^\alpha h_0 \|_\infty = O_p \Big( J^{-(p-|\alpha|)/d} \Big)~~\mbox{ for all }~~0 \leq |\alpha| < p\,. \] (2) If Assumptions (ref)(i)(ii) and (ref)(ii)(iii) hold, then \[ \|\partial^\alpha \wh h - \partial^\alpha h_0 \|_\infty = O_p \Big( J^{-(p-|\alpha|)/d} + \tau_J J^{|\alpha|/d} \sqrt{(J\log J)/n} \Big)~~\mbox{ for all }~~0 \leq |\alpha| < p\,. \] (2.a) Mildly ill-posed case: with $p \geq d/2$ and $\delta \geq d/(p + \varsigma)$, choosing $ J \asymp (n/\log n)^{d/(2(p+\varsigma)+d)}$ implies that Assumption (ref)(ii)(iii) holds and \begin{equation*} \|\partial^\alpha \wh h - \partial^\alpha h_0 \|_\infty = O_p ( (n/\log n)^{-(p-|\alpha|)/(2(p+\varsigma)+d)})\,. \end{equation*} (2.b) Severely ill-posed case: choosing $J = (c_0\log n)^{d/\varsigma}$ with $c_0 \in (0,1)$ implies that Assumption (ref)(ii)(iii) holds and \begin{equation*} \|\partial^\alpha \wh h - \partial^\alpha h_0 \|_{\infty} = O_p ( (\log n)^{-(p-|\alpha|)/\varsigma})\,. \end{equation*}

Corollary (ref) shows that, for sieve NPIV estimators, taking derivatives has the same impact on the bias and standard deviation terms in terms of the order of convergence, and that the same choice of sieve dimension $J$ can lead to optimal sup-norm convergence rates for estimating $h_0$ and its derivatives simultaneously (since they match the lower bounds in Theorem (ref) below). When specializing to series LS regression (without endogeneity, i.e., $\tau_J =1$), Corollary (ref)(2.a) with $\varsigma=0$ automatically implies that spline and wavelet series LS estimators will also achieve the optimal sup-norm rates of Stone1982 for estimating the derivatives of a nonparametric LS regression function. This strengthens the recent results in BCCK2014 and ChenChristensen-reg for sup-norm rate optimality of spline and wavelet LS estimators of the regression function $h_0$ itself. This is in contrast to kernel based LS regression estimators where different choices of bandwidth are needed for the optimal rates of estimating $h_0$ and its derivatives.

Corollary (ref) is useful for estimating functions with certain shape properties. For instance, if $h_0 : [a,b] \to \mb R$ is strictly monotone and/or strictly concave/convex, then knowing that $\partial {\wh h}(x)$ and/or $\partial^2 {\wh h}(x)$ converge uniformly to $\partial h_0 (x)$ and/or $\partial^2 h_0(x)$ implies that $\wh h$ will also be strictly monotone and/or strictly concave/convex wpa1. In this paper, we shall illustrate the usefulness of Corollary (ref) in controlling the nonlinear remainder terms for pointwise and uniform inferences on highly nonlinear (i.e., beyond quadratic) functionals of $h_0$; see Sections (ref) and (ref) for details.

Lower bounds

We now establish that the sup-norm rates obtained in Corollary (ref) are the best possible (i.e. minimax) sup-norm convergence rates for estimating $h_0$ and its derivatives.

To establish a lower bound, we require a link condition that relates smoothness of $T$ to the parameter space for $h_0$. Let $\widetilde \psi_{j,k,G}$ denote a tensor-product CDV wavelet basis for $[0,1]^d$ of regularity $\gamma > p$. Appendix (ref) provides details on the construction and properties of this basis.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Condition LB} (i) Assumption (ref)(i)--(iii) holds; (ii) $E[u_i^2 |W_i = w] \geq \underline \sigma^2 > 0$ uniformly for $w \in \mathcal W$; and (iii) there is a positive decreasing function $\nu$ s.t. $\|T h\|_{L^2(W)}^2 \lesssim \sum_{j,G,k} [\nu(2^j)]^2 \langle h, \widetilde \psi_{j,k,G} \rangle_X^2$ holds for all $h \in B_\infty(p,L)$.

Condition LB is standard in the optimal rate literature (see HallHorowitz and ChenReiss). The mildly ill-posed case corresponds to choosing $\nu(t) = t^{-\varsigma}$, and says roughly that the conditional expectation operator $T$ makes $p$-smooth functions of $X$ into $(\varsigma+p)$-smooth functions of $W$. The severely ill-posed case, which corresponds to choosing $\nu(t) = \exp(-\frac{1}{2}t^{\varsigma})$ and says roughly that $T$ maps smooth functions of $X$ into “supersmooth” functions of $W$.

theoremLet Condition LB hold for the NPIV model with a random sample $\{(X_{i},Y_{i},W_i)\}_{i=1}^n$. Then for any $0 \leq |\alpha| < p$: \begin{equation*} \liminf_{n \to \infty} \inf_{\wh g_n} \sup_{h \in B_\infty(p,L)} \mb P_h \left( \|\wh g_n - \partial^\alpha h\|_\infty \geq c r_n \right) \geq c'>0 \end{equation*} where \[ r_n = \left[ \begin{array}{ll} (n/\log n)^{-(p-|\alpha|)/(2(p+\varsigma)+d)} & \mbox{in the mildly ill-posed case } \\ (\log n)^{-(p-|\alpha|)/\varsigma} & \mbox{in the severely ill-posed case,} \end{array} \right. \] $\inf_{\wh g_n}$ denotes the infimum over all estimators of $\partial^\alpha h$ based on the sample of size $n$, $\sup_{h \in B_\infty(p,L)} \mb P_h$ denotes the sup over $h \in B_\infty(p,L)$ and distributions of $(X_i,W_i,u_i)$ that satisfy Condition LB with fixed $\nu$, and the finite positive constants $c, c'$ do not depend on $n$.

According to Theorem (ref) and Theorem (ref) (in Appendix (ref)), the minimax lower bounds in sup-norm for estimating $h_0$ and its derivatives coincide with those in $L^2$ for severely ill-posed NPIV problems, and are only a factor of $[\log( n)]^{\epsilon}$ (with $\epsilon =\frac{p-|\alpha|}{2(p+\varsigma)+d}<\frac{p}{2p+d}<\frac{1}{2}$) worse than those in $L^2$ for mildly ill-posed problems. Our proof of sup-norm lower bound for NPIV models is similar to that of ChenReiss for $L^2$-norm lower bound. Similar sup-norm lower bounds for density deconvolution were recently obtained by LouniciNickl.

Models with endogenous and exogenous regressors

In many empirical studies, some regressors might be endogenous while others are exogenous. Consider the model

equation[equation omitted — 60 chars of source]

where $X_{1i}$ is a vector of endogenous regressors and $Z_i$ is a vector of exogenous regressors. Let $X_i = (X_{1i}',Z_i')'$. Here the vector of instrumental variables $W_i$ is of the form $W_i = (W_{1i}',Z_i')'$ where $W_{1i}$ are instruments for $X_{1i}$. We refer to this as the “partially endogenous case”. The sieve NPIV estimator is implemented in exactly the same way as the “fully endogenous” setting in which $X_i$ consists only of endogenous variables, just like 2SLS with endogeneous and exogenous regressors.\footnote{ All that changes here is that $J$ may grow more quickly as the degree of ill-posedness will be smaller. In contrast, other NPIV estimators based on estimating the conditional densities of the regressors and instrumental variables must be implemented separately for each value of $z$ HallHorowitz,Horowitz2011,GagliardiniScaillet.} Our convergence rates presented in Section (ref) and Appendix (ref) apply equally to the partially endogenous model ((ref)) under the stated regularity conditions: all that differs between the two cases is the interpretation of the sieve measure of ill-posedness.

Consider first the fully endogenous case where $T : L^2(X) \to L^2(W)$ is compact under mild conditions on the conditional density of $X$ given $W$ (see, e.g., NeweyPowell,BCK,DFFR,Andrews2011). Then $T$ admits a singular value decomposition (SVD) $\{\phi_{0j},\phi_{1j},\mu_j\}_{j=1}^\infty$ where $(T^*T)^{1/2} \phi_{0j} = \mu_j \phi_{0j}$, $\mu_j \geq \mu_{j+1}$ for each $j$ and $\{\phi_{0j}\}_{j=1}^\infty$ and $\{\phi_{1j}\}_{j=1}^\infty$ are orthonormal bases for $L^2(X)$ and $L^2(W)$, respectively. Suppose that $\Psi_J$ spans $\phi_{0j},\ldots,\phi_{0J}$. Then the sieve measure of ill-posedness is $\tau_J = \mu_J^{-1}$.

Now consider the partially endogenous case. Similar to Horowitz2011, we suppose that for each value of $z$ the conditional expectation operator $T_z : L^2(X_1|Z=z) \to L^2(W_1|Z=z)$ given by $(T_z h)(w_1) = E[h(X_1)|W_{1i} = w_1,Z_i = z]$ is compact. Then each $T_z$ admits a SVD $\{\phi_{0j,z},\phi_{1j,z},\mu_{j,z}\}_{j=1}^\infty$ where $T_z \phi_{0j,z} = \mu_{j,z} \phi_{1j,z}$, $(T^*_{z}T^{\phantom *}_{z})^{1/2} \phi_{0j,z} = \mu_{j,z} \phi_{0j,z}$, $(T^{\phantom *}_{z}T^*_{z})^{1/2} \phi_{1j,z} = \mu_{j,z} \phi_{1j,z}$, $\mu_{j,z} \geq \mu_{j+1,z}$ for each $j$ and $z$, and $\{\phi_{0j,z}\}_{j=1}^\infty$ and $\{\phi_{1j,z}\}_{j=1}^\infty$ are orthonormal bases for $L^2(X_1|Z=z)$ and $L^2(W_1|Z=z)$, respectively, for each $z$. The following result adapts Lemma 1 of BCK to the partially endogenous setting.

lemmaLet $T_z$ be compact with SVD $\{\phi_{0j,z},\phi_{1j,z},\mu_{j,z}\}_{j=1}^\infty$ for each $z$. Let $\mu_j^{2}=E[\mu_{j,Z_i}^2]$ and $\phi_{0j}(\cdot,z)=\phi_{0j,z}(\cdot)$ for each $z$ and $j$. Then: (1) $\tau_J \geq \mu_J^{-1}$.\\ (2) If, in addition, $\phi_{01},\ldots,\phi_{0J} \in \Psi_J$, then: $\tau_J \leq \mu_J^{-1}$.

Consider the following partially-endogenous stylized example from HoderleinHolzmann. Let $X_{1i}$, $W_{1i}$ and $Z_i$ be scalar random variables with \[ \left(

array[array omitted — 39 chars of source]

\right) \sim N \left( \left(

array[array omitted — 27 chars of source]

\right) , \left(

array[array omitted — 104 chars of source]

\right) \right) \,. \] Then

equation[equation omitted — 372 chars of source]

where \[ \rho_{xw| z} = \frac{\rho_{xw} - \rho_{xz} \rho_{wz}}{\sqrt{(1-\rho^2_{xz})(1-\rho^2_{wz})}} \] is the partial correlation between $X_{1i}$ and $W_{1i}$ given $Z_i$. For each $j \geq 1$ let $H_j$ denote the $j$th Hermite polynomial (the Hermite polynomials form an orthonormal basis with respect to Gaussian density). Since $T_z : L^2(X_1|Z=z) \to L^2(W_1|Z=z)$ is compact for each $z$, it follows from Mehler's formula that $T_z$ has a SVD $\{\phi_{0j,z},\phi_{1j,z},\mu_{j,z}\}_{j=1}^\infty$ with \[ \phi_{0j,z}(x_1) = H_{j-1} \bigg( \frac{x_1 - \rho_{xz}z}{\sqrt{1-\rho_{xz}^2}} \bigg), \quad \phi_{1j,z}(w_1) = H_{j-1} \bigg( \frac{w_1 - \rho_{wz}z}{\sqrt{1-\rho_{wz}^2}} \bigg), \quad \mu_{j,z} = |\rho_{xw| Z}|^{j-1} \] for each $z$. Since $\mu_{J,z} = |\rho_{xw | z}|^{J-1}$ for each $z$, we have $\mu_J = |\rho_{xw | z}|^{J-1}\asymp |\rho_{xw | z}|^{J}$. If $X_{1i}$ and $W_{1i}$ are uncorrelated with $Z_i$, then $\mu_J = |\rho|^{J-1}$ where $\rho = \rho_{xw}$.

In contrast, consider the following fully-endogenous model in which $X_i$ and $W_i$ are bivariate with \[ \left(

array[array omitted — 52 chars of source]

\right) \sim N \left( \left(

array[array omitted — 32 chars of source]

\right) , \left(

array[array omitted — 107 chars of source]

\right) \right) \] where $\rho_1$ and $\rho_2$ are such that the covariance matrix is invertible. It is straightforward to verify that $T$ has singular value decomposition with \[ \phi_{0j}(x) = H_{j-1}(x_1) H_{j-1}(x_2) \, \quad \phi_{1j}(w) = H_{j-1}(w_1) H_{j-2}(w_2), \quad \mu_j = |\rho_1 \rho_2|^{j-1}\,, \] and $\mu_J = \rho^{2(J-1)}$ if $\rho_1 = \rho_2 = \rho$. Thus, the measure of ill-posedness diverges faster in the fully-endogenous case ($\mu_J = \rho^{2(J-1)}$) than that in the partially endogenous case ($\mu_J =|\rho|^{J-1}$).

Uniform inference on collections of nonlinear functionals

In this section we apply our sup-norm rate results and tight bounds on random matrices (in main online Appendix (ref)) to establish uniform Gaussian process strong approximation and the consistency of the score bootstrap UCBs defined in ((ref)) for collections of (possibly) nonlinear functionals $\{f_t(\cdot) : t \in \mathcal T\}$ of a NPIV function $h_0$. See Section (ref) for discussions of other applications.

We consider functionals $f_t : \mc H \subset L^\infty(X) \to \mb R$ for each $t \in \mathcal T$ for which $Df_t(h)[v] = \lim_{\delta \rightarrow 0^{+}} [\delta^{-1} f_t(h +\delta v)]$ exists for all $v\in \mc H - \{h_0\}$ for all $h$ in a small neighborhood of $h_0$ (where the neighborhood is independent of $t$). This is trivially true for, say, $f_t(h) = h(t)$ with $\mc T \subseteq \mc X$ for UCBs for $h_0$. Let $\Omega = E[u_i^2 b^K(W_i) b^K(W_i)']$. Then the “2SLS covariance matrix” for $\wh c$ (given in ((ref))) is $$\mho = [S' G_b^{-1} S]^{-1} S' G_b^{-1} \Omega G_b^{-1} S[ S' G_b^{-1} S]^{-1}~,$$ and the sieve variance for $f_t (\wh h)$ is

equation*[equation* omitted — 107 chars of source]

\setcounter{assumption}{1}

assumption[continued] (iii) $E[u_i^2 |W_i = w] \geq \underline \sigma^2 > 0$ uniformly for all $w \in \mathcal W$; and (iv) $\sup_w E[|u_i|^3|W_i = w] < \infty$.

Assumptions (ref)(iii)(iv) are reasonably mild conditions used to derive the uniform limit theory. Define

align*[align* omitted — 165 chars of source]

where, for each fixed $t$, $v_{n}(f_t)$ could be viewed as a “sieve 2SLS Riesz representer”. Note that $v_{n}(f_t)=\wh v_{n}(f_t)$ whenever $f_t$ is linear. Under Assumption (ref)(i)(iii) we have that

equation*[equation* omitted — 167 chars of source]

Following ChenPouzo2014, we call $f_t (\cdot)$ an irregular functional of $h_0$ (i.e., slower than $\sqrt n$-estimable) if $\sigma_n(f_t)\nearrow +\infty$ as $n \to \infty$. This includes the evaluation functionals $h_0(t)$ and $\partial^{\alpha}h_0(t)$ as well as $f_{CS,t}(h_0)$ and $f_{DL,t}(h_0)$. In this paper we shall focus on applications of sup-norm rate results to inference on irregular functionals.

\setcounter{assumption}{4}

assumptionLet $\eta_n$ and $\eta_n'$ be sequences of nonnegative numbers such that $\eta_n = o(1)$ and $\eta_n' = o(1)$. Let $\sigma_n(f_t) \nearrow +\infty$ as $n \to \infty$ for each $t \in \mathcal T$. Either (a) or (b) of the following holds: \begin{itemize} • $f_t $ is a linear functional for each $t \in \mathcal T$ and $\sup_{t \in \mathcal T}\sqrt{n} (\sigma_n(f_t))^{-1} |f_t(\widetilde{h})-f_t(h_{0})| = O_{p}(\eta_n)$; or • (i) $v\mapsto Df_t(h_{0})[v]$ is a linear functional for each $t \in \mathcal T$; (ii) \[ \sup_{t \in \mathcal T} \left| \sqrt n \frac{f_t(\wh h) - f(h_0) }{\sigma_n(f_t)} - \sqrt n \frac{Df_t(h_0)[\wh h - \widetilde h]}{\sigma_n(f_t)} \right| = O_p(\eta_n)\,; \] and (iii) $\sup_{t \in \mathcal T} \frac{\|\Pi_K T (\wh v_{n}(f_t) - v_{n}(f_t))\|_{L^2(W)}}{\sigma_n(f_t)} = O_p(\eta_n')$. \end{itemize}

Assumption (ref)(a)(b)(i)(ii) are similar to uniform-in-$t$ versions of Assumption 3.5 of ChenPouzo2014. Assumption (ref)(b)(iii) controls any additional error arising in the estimation of $\sigma_n(f_t)$ by $\wh\sigma(f_t)$ (given in equation (ref)) due to nonlinearity of $f_t(\cdot)$, and is automatically satisfied with $\eta_n'=0$ when $f_t(\cdot)$ is a linear functional.

The next remark presents a set of sufficient conditions for Assumption (ref) when $\{f_t : t \in \mc T\}$ are irregular functionals of $h_0$. Since the functionals are irregular, the quantity $\ul \sigma_n := \inf_{t \in \mathcal T} \sigma_n(f_t)$ will typically satisfy $\ul \sigma_n \nearrow +\infty$ as $n \to \infty$. Our sup-norm rates for $\wh h$ and $\wt h$, together with divergence of $\ul \sigma_n$, helps to control the nonlinearity bias terms.

remarkLet $\mathcal H_n \subseteq \mathcal H$ be a sequence of neighborhoods of $h_0$ with $\wh h,\widetilde h \in \mathcal H_n$ wpa1 and assume $\ul \sigma_n :=\inf_{t \in \mathcal T} \sigma_n(f_t) >0$ for each $n$. Then: Assumption (ref)(a) is implied by (a'), and Assumption (ref)(b) is implied by (b'), where \begin{itemize} • (i) $f_t$ is a linear functional for each $t \in \mathcal T$ and there exists $\alpha$ with $|\alpha| \geq 0$ s.t. $\sup_t |f_t(h-h_0)| \lesssim \|\partial^\alpha h- \partial^\alpha h_0 \|_\infty$ for all $h \in \mathcal H_n$; and (ii) $n^{1/2}\ul \sigma_n^{-1}\|\partial^\alpha \widetilde h - \partial^\alpha h_0\|_\infty = O_{p}(\eta_n)$; or • (i) $v\mapsto Df_t(h_{0})[v]$ is a linear functional for each $t\in \mathcal T$ and there exists $\alpha$ with $|\alpha| \geq 0$ s.t. $\sup_t|D f_t(h_{0})[h-h_0]| \lesssim \|\partial^\alpha h - \partial^\alpha h_0\|_\infty$ for all $h \in \mathcal H_n$;\\ (ii) there are $\alpha_1$, $\alpha_2$ with $|\alpha_1|,|\alpha_2| \geq 0$ s.t. \begin{align*} (ii.1)& \; \sup_t \left| f_t(\wh h) - f_t(h_0) -Df_t(h_{0})[\wh h - h_0] \right| \lesssim \|\partial^{\alpha_1} \wh h - \partial^{\alpha_1} h_0\|_\infty \|\partial ^{\alpha_2} \wh h - \partial^{\alpha_2} h_0\|_\infty\, and \\ (ii.2)& \; n^{1/2} \ul \sigma_n^{-1} \big( \|\partial^{\alpha_1}\wh h - \partial^{\alpha_1}h_0\|_\infty \|\partial ^{\alpha_2} \wh h - \partial^{\alpha_2} h_0\|_\infty + \|\partial^\alpha \widetilde h - \partial^\alpha h_0\|_\infty \big) = O_{p}(\eta_n) ; \end{align*} and (iii) $\sup_{t\in \mathcal T} \frac{(\tau_J)\sqrt{\sum_{j=1}^J \left( Df_{t}(\wh h)[(G_\psi^{-1/2} \psi^J)_j] - Df_{t}(h_0)[(G_\psi^{-1/2} \psi^J)_j] \right)^2}}{\sigma_n(f_{t})}=O_p(\eta_n')$. \end{itemize}

Condition (a')(i) is automatically satisfied by functionals of the form $f_t(h) = \partial^\alpha h(t)$ with $\mathcal T \subseteq \mathcal X$ and $\mathcal H_n = \mathcal H$. Conditions (a')(i) and (b')(i)(ii) are sufficient conditions that are formulated to take advantage of the sup-norm rate results in Section (ref). For example, condition (b')(i)(ii.1) is easily satisfied by exact CS and DL functionals (lemma A.1 of HausmanNewey1995). Condition (b')(ii.2) is simply satisfied by applying our sup-norm rate results. Condition (b')(iii) is a sufficient condition for Assumption (ref)(b)(iii), and is needed for uniform-in-$t$ consistent estimation of $\sigma_n(f_t)$ by $\wh\sigma (f_t)$ only, and is automatically satisfied with $\eta_n'=0$ when $f_t(\cdot)$ is a linear functional.

The next assumption concerns the set of normalized sieve 2SLS Riesz representers, given by \[ u_n(f_t)(x) = v_n(f_t)(x)/\sigma_n(f_t) \,. \] Let $d_n$ denote the semi-metric on $\mathcal T$ given by $d_n(t_1,t_2)^2 = E[(u_n(f_{t_1})(X_i) - u_n(f_{t_2})(X_i))^2]$ and $N(\mathcal T,d_n,\epsilon)$ be the $\epsilon$-covering number of $\mathcal T$ with respect to $d_n$. Let $\eta_n$ and $\eta_n'$ be from Assumption (ref), and $\delta_{h,n}$ be a sequence of positive constants such that $\|\wh h - h_0\|_\infty = O_p(\delta_{h,n})=o_p (1)$. Denote $\delta_{V,n} \equiv \big[ \zeta_{b,K}^{(2+\delta)/\delta} \sqrt{(\log K)/n} \big]^{\delta/(1+\delta)} + \tau_J \zeta \sqrt{(\log J)/n} + \delta_{h,n}$.

\setcounter{assumption}{5}

assumption(i) there is a sequence of finite constants $c_n \gtrsim 1$ that could grow to infinity such that \[ 1 + \int_0^\infty \sqrt{\log N(\mathcal T,d_n,\epsilon)}\,\mathrm d \epsilon = O( c_n)\,; \] and (ii) there is a sequence of constants $r_n>0$ decreasing to zero slowly such that\\ (ii.1) $r_n c_n \lesssim 1$ and $\frac{\zeta_{b,K} J^2}{r_n^3 \sqrt n} = o(1)$; and\\ (ii.2) $\tau_J \zeta \sqrt{(J \log J)/n} + \eta_n + ( \delta_{V,n} + \eta_n' )\times c_n = o(r_n)$, with $\eta_n' \equiv 0$ when $f_t(\cdot)$ is linear.

Assumption (ref)(i) is a mild regularity condition requiring that the class $\{u_n(f_t) : t \in \mathcal T\}$ not be too complex; see Remark (ref) below for sufficient conditions to bound $c_n$. Assumption (ref)(ii) strengthens conditions on the growth rate of $J$. Condition $\frac{\zeta_{b,K} J^2}{r_n^3 \sqrt n} = o(1)$ of Assumption (ref)(ii.1) is used to apply Yurinskii's coupling CLR,PollardUGMTP to derive uniform Gaussian process strong approximation to the linearized sieve process $\{\widehat{\mb Z}_n(t):t\in \mathcal T \}$ (defined in equation ((ref))). This condition could be improved if other types of strong approximation probability tools are used. Assumption (ref)(ii.2) ensures that both the nonlinear remainder terms and the error in estimating $\sigma_n(f_t)$ by $\wh \sigma(f_t)$ vanish sufficiently fast. While the consistency of $\wh \sigma(f)$ is enough for the pointwise asymptotic normality of the plug-in sieve $t$-statistic for $f(h_0)$ (see Theorem (ref) in the main online Appendix (ref)), we need the following rate of convergence for uniform inference \[ \sup_{t \in \mathcal T} \left| \frac{\sigma_n(f_t)}{\wh\sigma (f_t)} -1 \right| = O_p( \delta_{V,n} + \eta_n')~, \] which is established using our results on sup-norm convergence rates of sieve NPIV; see Lemma (ref) in the secondary online Appendix (ref).

remarkLet Assumptions (ref)(iii) and (ref)(i) hold. Let $\mathcal T$ be a compact subset in $\mb R^{d_T}$, and there exist positive sequences $\Gamma_n$ and $\gamma_n$ such that for any $t_1, t_2 \in \mathcal T$, \[ \sup_{h \in \Psi_J : \|h\|_{L^2(X)} = 1} \left| \left( Df_{t_1}(h_0)[h] - Df_{t_2}(h_0)[h] \right) \right| \leq \Gamma_n \|t_1 - t_2 \|_{\ell^2}^{\gamma_n}~. \] Then: Assumption (ref)(i) holds with $c_n = 1 + \int_0^\infty \sqrt{ \{ (d_T /\gamma_n )\log ( \Gamma_n \tau_J/(\epsilon \ul \sigma_n ))\} \vee 0 }\,\mathrm d \epsilon$.

The next lemma is about uniform Bahadur representation and uniform Gaussian process strong approximation for the sieve $t$-statistic process for (possibly) nonlinear functionals of NPIV. Define

align[align omitted — 338 chars of source]

with $\mathcal Z_n \sim N(0, G_b^{-1/2}\Omega G_b^{-1/2} )$. Note that $\mb Z_n(t)$ is a Gaussian process indexed by $t \in \mc T$.

lemmaLet Assumptions (ref)(iii), (ref), (ref)(ii)(iii), (ref)(i), (ref) and (ref) hold. Then: \begin{equation} \sup_{t \in \mc T} \left| \frac{\sqrt n (f_t(\wh h) - f_t(h_0))}{\wh\sigma (f_t)} - {\mb Z}_n(t) \right| = \sup_{t \in \mc T} \left| \frac{\sqrt n (f_t(\wh h) - f_t(h_0))}{\wh\sigma (f_t)} - \widehat{\mb Z}_n(t) \right| + o_p(r_n ) = o_p(r_n)\,. \end{equation}

Lemma (ref) is used in this paper to establish the consistency of the sieve score bootstrap for estimating the critical values of the uniform sieve $t$-statistic process, $\sup_{t \in \mathcal T} \left|\frac{\sqrt n (f_t(\wh h) - f_t(h_0))}{\wh\sigma (f_t)} \right|$, for a NPIV model. The strong approximation result, however, is also useful for various applications to testing equality and/or inequality (such as shape) constraints on $f_t(h_0)$, and is therefore of independent interest.

In what follows, $\mb P^*(\cdot)$ denotes a probability measure conditional on the data $Z^n:= \{(X_i,Y_i,W_i)\}_{i=1}^n$. Recall that $\mb Z_n^*(t)$ is defined in equation ((ref)).

theoremLet conditions of Lemma (ref) hold. Let $\eta_n' \sqrt J = o(r_n)$ for nonlinear $f_t()$. Let the bootstrap weights $\{\varpi_{i}\}_{i=1}^n$ be IID with zero mean, unit variance and finite 3rd moment, and independent of the data. Then: \begin{equation} \sup_{s \in \mb R} \left| \mb P \left( \sup_{t \in \mathcal T} \left|\frac{\sqrt n (f_t(\wh h) - f_t(h_0))}{\wh\sigma (f_t)} \right| \leq s \right) - \mb P^*\left( \sup_{t \in \mathcal T} |\mb Z_n^*(t)| \leq s\right) \right| = o_p(1)\,. \end{equation}

Theorem (ref) appears to be the first to establish consistency of a sieve score bootstrap for uniform inference on general nonlinear functionals of NPIV under low-level conditions. When specializing to collections of linear functionals, Lemma (ref), Theorem (ref) and Corollary (ref) immediately imply the following result.

corollaryConsider a collection of linear functionals $\{f_t(h_0) =\partial^\alpha h_0 (t): t\in \mathcal T \}$ of the NPIV function $h_0$, with $\mathcal T$ a compact convex subset of $\mathcal X$. Let Assumptions (ref)(i)(ii)(iii) and (ref) (with $\delta \geq 1$) hold, $h_0 \in B_\infty(p,L)$, $\Psi_J$ be formed from a B-spline basis of regularity $\gamma > (p \vee 2 + |\alpha|)$, $B_K$ be a B-spline, wavelet or cosine basis, and $\sigma_n(f_t) \asymp \tau_J J^{a}$ uniformly in $t$ with $a =\frac{1}{2} + \frac{|\alpha|}{d}$. For $\kappa \in [1/2,1]$ we set $J^5(\log n)^{6\kappa} /n = o(1)$, $\tau_J J (\log J)^{\kappa+0.5} / \sqrt{n}= o(1)$ and $J^{-p/d}=o([\log J]^{-\kappa} \tau_J \sqrt{J/n})$. Then: Results ((ref)) (with $r_n=(\log J)^{-\kappa}$) and ((ref)) hold for $f_t(h_0) =\partial^\alpha h_0 (t)$.

Recently HorowitzLee2012 developed a notion of UCBs for a NPIV function $h_0$ of a scalar endogenous regressor $X_i \in [0,1]$ based on interpolation over a growing number of uniformly generated random grid points on $[0,1]$, with $h_0$ estimated via the modified orthogonal series NPIV estimator of Horowitz2011.\footnote{Remark 4 in HorowitzLee2012 mentioned that their notion of UCB is different from the standard UCBs. They also proved the consistency of their bootstrap confidence bands over fixed finite number of grid points.} When specializing Corollary (ref) to a NPIV function of a scalar regressor (i.e., $d=1$ and $|\alpha|=0$), our sufficient conditions are comparable to theirs (see their theorem 4.1). Our score bootstrap UCBs would be computationally much simpler for a NPIV function of a multivariate endogenous regressor $X_i$, however.

When $X_i$ is exogenous, the sieve NPIV estimator $\wh h$ reduces to the series LS estimator of a nonparametric regression $h_0(x) = E[Y_i | W_i = x]$ with $X_i = W_i$, $K =J$ and $b^K = \psi^J$ with $\tau_J =1$. Lemma (ref) and Theorem (ref) immediately imply the validity of Gaussian strong approximation and sieve score bootstrap UCBs for collections of general nonlinear functionals of a nonparametric LS regression. We note that the regularity conditions in Lemma (ref) and Theorem (ref) are much weaker for models with exogenous regressors. For instance, when specializing Corollary (ref) to a nonparametric LS regression with exogenous regressor $X_i$, the conditions on $J$ simplify to $J^5(\log n)^{6\kappa} /n = o(1)$ and $J^{-p/d}=o([\log J]^{-\kappa}\sqrt{J/n} )$ for $\kappa \in [1/2,1]$, and Results ((ref)) (with $r_n = [\log J]^{-\kappa}$) and ((ref)) both hold for linear functionals $\{f_t(h_0) = \partial^\alpha h(t_0 ): t\in \mathcal T \}$ of $h_0(\cdot) =E[Y_i | X_i = \cdot]$. These conditions on $J$ are the same as those in CLR for $h_0$ (see their theorem 7) and BCCK2014 for linear functionals of $h_0$ (see their theorem 5.5 with $r_n = [\log J]^{-1/2}$) estimated via series LS.

To the best of our knowledge, there is no published work on uniform Gaussian process strong approximation and sieve score bootstrap for general nonlinear functionals of sieve NPIV or series LS regression. The results in this section are thus presented as non-trivial applications of our sup-norm rate results for sieve NPIV, and are not aimed at weakest sufficient conditions.

Monte Carlo

We now evaluate the finite sample performance of our sieve score bootstrap UCBs for $h_0$ in NPIV model ((ref)). We use the experimental design of NeweyPowell, in which IID draws are generated from \[ \left(

array[array omitted — 37 chars of source]

\right) \sim N \left( \left(

array[array omitted — 27 chars of source]

\right) , \left(

array[array omitted — 57 chars of source]

\right) \right) \] from which we then set $X_i^* = W_i^* + V_i^*$. To ensure compact support of the regressor and instrument, we rescale $X_i^*$ and $W_i^*$ by defining $X_i = \Phi(X_i^*/\sqrt 2)$ and $W_i = \Phi(W_i^*)$ where $\Phi$ is the Gaussian cdf. We use $h_0(x) = 4x-2$ for our linear design and $h_0(x) = \log(|16x-8|+1)\mathrm{sgn}(x-\frac{1}{2})$ for our nonlinear design (our nonlinear $h_0$ is a re-scaled version of the $h_0$ used in NeweyPowell). Note that $p$ for the nonlinear $h_0$ is between $1$ and $2$, so $h_0$ is not particularly smooth ($h_0'(x)$ has a kink at $x = \frac{1}{2}$).

We generate 1000 samples of length 1000 and implement our procedure using a B-spline basis for $B_K$ and $\Psi_J$. For each simulation, we calculate the 90%, 95%, and 99% uniform confidence bands for $h_0$ over the support $[0.05,0.95]$ with 1000 bootstrap replications for each simulation. We draw the bootstrap innovations $\varpi_i$ from the two-point distribution of Mammen1993. We then calculate the MC coverage probabilities of our uniform confidence bands.

table[table omitted — 941 chars of source]
figure[figure omitted — 381 chars of source]

Figure (ref) displays the estimated structural function $\wh h$ and confidence bands together with a scatterplot of the sample $(X_i,Y_i)$ data for the nonlinear design. The true function $h_0$ is seen to lie inside the UCBs. The results of this MC experiment are presented in Table (ref). Comparing the MC coverage probabilities with their nominal values, it is clear that the uniform confidence bands for the linear design are slightly too conservative. However, the uniform confidence bands for the nonlinear design using cubic B-splines to approximate $h_0$ have MC converge much closer to the nominal coverage probabilities.

Pointwise and uniform inference on nonparametric welfare functionals

We now apply our sup-norm rate results to study pointwise and uniform inference on nonlinear welfare functionals in nonparametric demand estimation with endogeneity. First, we provide mild sufficient conditions under which plug-in sieve $t$-statistics for exact CS and DL and approximate CS functionals are asymptotically $N(0,1)$, allowing for mildly and severely ill-posed NPIV models (subsections (ref) and (ref)). Second, under stronger sufficient conditions but still allowing for severely ill-posed NPIV models, the validity of uniform Gaussian process strong approximations and sieve score bootstrap UCBs for exact CS and DL over a range of taxes and/or incomes (subsection (ref)) are presented. When specialized to inference on exact CS and DL and approximate CS functionals of nonparametric demand estimation without endogeneity, our pointwise asymptotic normality results are valid under sufficient conditions weaker than those in the existing literature, while our uniform inference results appear to be new (subsection (ref)).

Previously, HausmanNewey1995 and Newey1997 provided sufficient conditions for pointwise asymptotic normality for plug-in nonparametric LS estimators of exact CS and DL functionals and of approximate CS functionals respectively, when prices and incomes are exogenous. Vanhems2010 studied consistency and convergence rates of kernel-based plug-in estimators of CS functional allowing for mildly ill-posed NPIV models. BlundellHorowitzParey2012 and HausmanNewey2016 estimated CS and DL of nonparametric gasoline demand allowing for prices to be endogenous, but did not provide theoretical justification for their inference approach under endogeneity. Therefore, although presented as applications of our sup-norm rate results, our inference results contribute nicely to the literature on nonparametric welfare analysis.

Pointwise inference on exact CS and DL with endogeneity

Here we present primitive regularity conditions for pointwise asymptotic normality of the sieve $t$-statistics for exact CS and DL. We suppress dependence of the functionals on $t = (\mf p^0, \mf p^1,\mf y)$.

Let $\mf X_i = (\mf P_i,\mf Y_i)$. We assume in what follows that the support of both $\mf P_i$ and $\mf Y_i$ is bounded away from zero. If both $\mf P_i$ and $\mf Y_i$ are endogenous, let $\mf W_i$ be a $2\times 1$ vector of instruments. Let $T : L^2(\mf X) \to L^2(\mf W)$ be compact and injective with singular value decomposition (SVD) $\{\phi_{0j},\phi_{1j},\mu_j\}_{j=1}^\infty$ where \[ T \phi_{0j} = \mu_j \phi_{1j}, \quad (T^*T)^{1/2} \phi_{0j} = \mu_j \phi_{0j}, \quad (TT^*)^{1/2} \phi_{1j} = \mu_j \phi_{1j} \] and $\{\phi_{0j}\}_{j=1}^\infty$ and $\{\phi_{0j}\}_{j=1}^\infty$ are orthonormal bases for $L^2(\mf X)$ and $L^2(\mf W)$, respectively. If $\mf P_i$ is endogenous but $\mf Y_i$ is exogenous, we take $\mf W_i = (\mf W_{1i},\mf Y_i)'$ with $\mf W_{1i}$ an instrument for $\mf P_i$. Let $T_{\mf y} : L^2(\mf P|\mf Y = \mf y) \to L^2(\mf W_1|\mf Y=\mf y)$ be compact and injective with SVD $\{\phi_{0j,\mf y},\phi_{1j,\mf y},\mu_{j,\mf y}\}_{j=1}^\infty$ for each $\mf y$ where \[ T_{\mf y} \phi_{0j,\mf y} = \mu_{j,\mf y} \phi_{1j,\mf y}, \quad (T^*_{\mf y}T^{\phantom *}_{\mf y})^{1/2} \phi_{0j,\mf y} = \mu_{j,\mf y} \phi_{0j,\mf y}, \quad (T^{\phantom *}_{\mf y}T^*_{\mf y})^{1/2} \phi_{1j,\mf y} = \mu_{j,\mf y} \phi_{1j,\mf y} \] and $\{\phi_{0j,\mf y}\}_{j=1}^\infty$ and $\{\phi_{0j,\mf y}\}_{j=1}^\infty$ are orthonormal bases for $L^2(\mf P|\mf Y = \mf y)$ and $L^2(\mf W_1|\mf Y = \mf y)$, respectively. In this case, we define $\phi_{0j}(\mf p,\mf y) = \phi_{0j,\mf y}(\mf p)$, $\phi_{1j}(\mf w_1,\mf y) = \phi_{1j,\mf y}(\mf w_1)$, and $\mu_j^2 = E[\mu_{j,\mf Y_i}^2]$ (see Section (ref) for further details).

In both cases, we follow ChenPouzo2014 and assume that $\Psi_J$ and $B_K$ are Riesz bases in that they span $\phi_{01},\ldots,\phi_{0J}$ and $\phi_{11},\ldots,\phi_{1K}$, respectively. This implies that $\tau_J \asymp \mu_J^{-1}$. For fixed $\mf p^0$, $\mf p^1$, and $\mf y$ we define

equation*[equation* omitted — 245 chars of source]

for the exact CS functional.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Assumption CS} (i) $\mf X_i$ and $\mf W_i$ both have compact rectangular support and densities bounded away from $0$ and $\infty$; (ii) $h_0 \in B_{\infty}(p,L)$ with $p > 2$ and $0 < L < \infty$; (iii) $E[\mf u_i^2|\mf W_i = w]$ is uniformly bounded away from $0$ and $\infty$, $E[|\mf u_i|^{2+\delta}]$ is finite for some $\delta>0$, and $ \sup_w E[u_i^2 \{ |u_i | > \ell(n)\}|\mf W_i = w] = o(1)$ for any positive sequence with $\ell(n) \nearrow \infty$; (iv) $\Psi_J$ is spanned by a (tensor-product) B-spline basis of order $\gamma > p$ or continuously differentiable wavelet basis of regularity $\gamma > p$ and $B_K$ is spanned by a (tensor-product) B-spline, wavelet or cosine basis; (v) $J^{(2+\delta)/(2\delta)}\sqrt{(\log n) / n} = o(1)$ and \[ \frac{\sqrt n}{ \big( \sum_{j=1}^J (a_j/\mu_j)^2 \big)^{1/2} } \times \bigg( J^{-p/2} + \mu_J^{-2} \frac{J^2 \sqrt{\log J}}{n} \bigg) = o(1)\,. \]

Assumption CS(i)--(iv) is standard even for series LS regression without endoegenity. Let $[\sigma_n(f_{CS})]^2 = \big(Df_{CS} (h_{0})[\psi^J]\big)' \mho \big(Df_{CS} (h_{0})[\psi^J]\big)$ be the sieve variance of the plug-in sieve NPIV estimator $f_{CS} (\wh h_0 )$. Then these assumptions imply that $[\sigma_n(f_{CS})]^2 \asymp \sum_{j=1}^J (a_j/\mu_j)^2 \lesssim J \mu_J^{-2}$. Assumption CS(v) is sufficient for Remark (ref)(b') for a fixed $t$.

Our first result is for exact CS functionals, established by applying Theorem (ref) in the main online Appendix (ref). Let \[ \wh \sigma^2 (f_{CS}) = Df_{CS}(\wh h)[\psi^J]'\, \wh \mho\, Df_{CS}(\wh h)[\psi^J] \] with \[ Df_{CS}(\wh h)[\psi^J] = \int_{0}^{1} \psi^J( \mf p(u),\mf y-\widehat{\mf S}_{\mf y}(\mf p(u))) e^{-\int_{0}^u \partial_2 \wh h( \mf p(v),\mf y-\widehat{\mf S}_{\mf y}(\mf p(v))) \mf p'(v)\,\mathrm dv} \mf p'(u) \,\mathrm du \,. \]

theoremLet Assumption CS hold. Then: the sieve $t$-statistic for $f_{CS}(h_0)$ is asymptotically $N(0,1)$, i.e., \[ \sqrt n\frac{ f_{CS}(\wh h) - f_{CS}(h_0)}{\wh \sigma (f_{CS})} \to_d N(0,1)\,. \]

Since $\mu_j >0$ decreases as $j$ increases, we could use the following relation

equation[equation omitted — 265 chars of source]

to provide simpler sufficient conditions for Assumption CS(v) that could be satisfied by both mildly and severely ill-posed NPIV models. Corollary (ref) provides one set of concrete sufficient conditions for Assumption CS(v).

corollaryLet Assumption CS(i)--(iv) hold and $a_j^2 \asymp j^{a}$ for $a\leq 0$. Then: $[\sigma_n(f_{CS})]^2 \asymp \sum_{j=1}^J (j^{a} \mu_j^{-2})$. (1) Mildly ill-posed case: let $\mu_j \asymp j^{-\varsigma/2}$ for $\varsigma \geq 0, a + \varsigma > -1$. Then: \[ [\sigma_n(f_{CS})]^2 \asymp J^{(a+\varsigma)+1}~; \] further, if $\delta \geq 2/(2+\varsigma- a)$, $n J^{-(p+a+\varsigma+1)} = o(1)$ and $J^{3+\varsigma - a} (\log n)/n = o(1)$, then: Assumption CS(v) is satisfied, and the sieve $t$-statistic for $f_{CS}(h_0)$ is asymptotically $N(0,1)$. (2) Severely ill-posed case: let $\mu_j \asymp \exp(-\frac{1}{2} j^{\varsigma/2})$, $\varsigma > 0$ and $J =( \log (n/(\log n)^{\varrho}))^{2/\varsigma}$ for $\varrho>0$. Then: \[ [\sigma_n(f_{CS})]^2 \gtrsim \frac{n}{(\log n)^{\varrho}} \times (\log (n/(\log n)^\varrho))^{2a/\varsigma}\,; \] further, if $\varrho>0$ is chosen such that $2 p > \varrho \varsigma - 2a$ and $\varrho \varsigma > 8-2a $, then: Assumption CS(v) is satisfied, and the sieve $t$-statistic for $f_{CS}(h_0)$ is asymptotically $N(0,1)$.

Note that in Corollary (ref), $J$ may be chosen to satisfy the stated conditions in the mildly ill-posed case whenever $p > 2-2a$, and in the severely ill-posed case whenever $p > 4 - 2a$.

Our next result is for DL functionals. Note that DL is the sum of CS and a tax receipts functional, namely $(\mf p^1 - \mf p^0)h_0(\mf p^1,\mf y)$. Note that the tax receipts functional is typically less smooth and hence converges slower than that of CS functional. Therefore, $[\sigma_n(f_{DL})]^2 = \big(Df_{DL}(h_{0})[\psi^J]\big)' \mho \big(Df_{DL}(h_{0})[\psi^J]\big)$ will typically grow at the order of $(\tau_J \sqrt J)^2$, which is the growth order of the sieve variance term for estimating the unknown NPIV function $h_0$ at a fixed point. For this reason we do not derive the joint asymptotic distribution of $f_{CS}(\wh h)$ and $f_{DL}(\wh h)$. The next result adapts Theorem (ref) to derive asymptotic normality of plug-in sieve $t$-statistics for DL functionals. Let \[ \wh \sigma^2 (f_{DL}) = Df_{DL}(\wh h)[\psi^J]'\, \wh \mho\, Df_{DL}(\wh h)[\psi^J] \] with \[ Df_{DL}(\wh h)[\psi^J] = Df_{CS}(\wh h)[\psi^J] - (\mf p^1 - \mf p^0) \psi^J(\mf p^1,\mf y)\,. \]

theoremLet Assumption CS(i)--(iv) hold. Let $\sigma_n (f_{DL}) \asymp \mu_J^{-1} \sqrt J$, $\sqrt n \mu_J J^{-(p+1)/2} = o(1)$ and $(J^{(2+\delta)/(2\delta)}\sqrt{\log n} \vee \mu_J^{-1}J^{3/2}\sqrt{\log J}) /\sqrt n= o(1)$. Then: \[ \sqrt n\frac{ f_{DL}(\wh h) - f_{DL}(h_0)}{\wh \sigma (f_{DL})} \to_d N(0,1)\,. \]

Pointwise inference on approximate CS with endogeneity

Suppose instead that demand of consumer $i$ for some good is estimated in logs, i.e.

equation[equation omitted — 87 chars of source]

As $h_0$ is the log-demand function, any linear functional of demand is a nonlinear functional of $h_0$. One such example is the weighted average demand functional of the form

equation*[equation* omitted — 87 chars of source]

where $w(\mf p)$ is a non-negative weighting function and $\mf y$ is fixed. With $w(\mf p) = 1\!\mathrm{l} \{ \underline{\mf p} \leq \mf p \leq \overline{\mf p}\}$, the functional $f(h)$ may be interpreted as the approximate CS. The functional is defined for fixed $\mf y$, so it will typically be an irregular functional of $h_0$.

The setup is similar to the previous subsection. Let $\mf X_i = (\log \mf P_i,\log \mf Y_i)$. If both $\mf P_i$ and $\mf Y_i$ are endogenous, we let $\mf W_i$ be a $2\times 1$ vector of instruments and $T : L^2(\mf X) \to L^2(\mf W)$ be compact with SVD $\{\phi_{0j},\phi_{1j},\mu_j\}_{j=1}^\infty$. If $\mf P_i$ is endogenous but $\mf Y_i$ is exogenous, we let $\mf W_i = (\mf W_{1i},\log \mf Y_i)'$ with $\mf W_{1i}$ an instrument for $\mf P_i$, and let $T_{\mf y} : L^2(\log \mf P|\log \mf Y = \log \mf y) \to L^2(\mf W_1|\log \mf Y=\log \mf y)$ be compact with SVD $\{\phi_{0j,\mf y},\phi_{1j,\mf y},\mu_{j,\mf y}\}_{j=1}^\infty$ for each $\mf y$. In this case, we define $\phi_{0j}(\log \mf p,\log \mf y) = \phi_{0j,\mf y}(\log \mf p)$, $\phi_{1j}(\mf w_1,\log \mf y) = \phi_{1j,\mf y}(\mf w_1)$, and $\mu_j^2 = E[\mu_{j,\mf Y_i}^2]$. We again assume that $\Psi_J$ and $B_K$ are Riesz bases. For each $j \geq 1$, define

align*[align* omitted — 135 chars of source]

The next result follows from Theorem (ref) (in the main online Appendix (ref)). Let \[ \wh \sigma^2 (f_A) = Df_{A}(\wh h)[\psi^J]'\, \wh \mho\, Df_{A}(\wh h)[\psi^J] \] with \[ Df_A(\wh h)[\psi^J] = \int w(\mf p) e^{\wh h(\log \mf p,\log \mf y)}\psi^J(\log \mf p,\log \mf y) \, \mathrm d \mf p \,. \]

theoremLet Assumption CS(i)--(iv) hold for the log-demand model ((ref)) with $p > 0$, and let $ J^{(2+\delta)/(2\delta)}\sqrt{(\log n)/n} = o(1)$ and \[ \frac{\sqrt n}{ \big( \sum_{j=1}^J (a_j/\mu_j)^2 \big)^{1/2} } \times \bigg( J^{-p/2} + \mu_J^{-2} \frac{J^{3/2} \sqrt{\log J}}{n} \bigg) = o(1)~. \] Then: \[ \frac{\sqrt n ( f_A(\widehat h) - f_A(h_0))}{ \wh \sigma (f_A)} \to_d N(0,1)\,. \]

Uniform inference on collections of exact CS and DL functionals with endogeneity

Here we apply Lemma (ref) and Theorem (ref) to present sufficient conditions for uniform Gaussian process strong approximations and bootstrap UCBs for exact CS and DL under endogeneity. We maintain the setup described at the beginning of Subsection (ref). We take $t = (\mf p^0,\mf p^1,\mf y) \in \mc T = [\underline{\mf p}^0 , \overline{\mf p}^0] \times [\underline{\mf p}^1 , \overline{\mf p}^1] \times [\underline{\mf y}, \overline{\mf y}]$, where the intervals $[\underline{\mf p}^0 , \overline{\mf p}^0]$ and $[\underline{\mf p}^1 , \overline{\mf p}^1]$ are in the interior of the support of $\mf P_i$ and $[\underline{\mf y}, \overline{\mf y}]$ is in the interior of the support of $\mf Y_i$. For each $t \in \mc T$ we let

equation[equation omitted — 249 chars of source]

for each $j \geq 1$ (where $\mf p(u)$ is a smooth price path from $\mf p^0 = \mf p(0)$ to $\mf p^1 = \mf p(1)$). Also define $\ul \sigma_n = \inf_{t \in \mc T}(( \sum_{j=1}^J (a_{j,t}/\mu_j)^2 )^{1/2} $.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Assumption U-CS} (i) $E[\mf u_i^2|\mf W_i = w]$ is uniformly bounded away from $0$, $E[|\mf u_i|^{2+\delta}]$ is finite with $\delta \geq 1$, and $\sup_w E[|\mf u_i|^3 |W_i = w]$ is finite; (ii) the H\"older condition in Remark (ref) holds with $\gamma_n = \gamma$ and $\Gamma_n \lesssim J^c$ for some finite positive constants $\gamma$ and $c$; (iii) $J^5(\log n)^{3} /n = o(1)$, $\frac{\sqrt {n(\log J)}}{ \ul \sigma_n } J^{-p/2} =o(1)$; (iv) let $\eta_n' = \frac{ J^{3/2} \mu_J^{-1}}{ \ul \sigma_n } \left( J^{-p/2} + \mu_J^{-1} \sqrt{J(\log J)/n} \right)$, either (iv.1) $\eta_n' (\log J)=o(1)$, or (iv.2) $\eta_n' \sqrt{J(\log J)} =o(1)$.

Assumption U-CS (i) is slightly stronger than Assumption CS(iii) (since $\delta = 1$ in Assumption U-CS(i) is enough). Assumption U-CS(ii) is made for simplicity to verify Assumption (ref)(i); other sufficient conditions could also be used. Assumption U-CS(iii)(iv.1) strengthens Assumption CS(v) to ensure uniform Gaussian process strong approximation with an error rate of $r_n =(\log J)^{-1/2}$. Again, one could use bounds on $\ul \sigma_n$ that is analogous to Relation ((ref)) to provide sufficient conditions for Assumption U-CS(iii)(iv) that could be satisfied by mildly and severely ill-posed NPIV models. See Remark (ref) below for one concrete set of such sufficient conditions.

remarkLet $\ul \sigma_n^2 \gtrsim \sum_{j=1}^J (j^{a}\mu_j^{-2})$ for $a \leq 0$.\\ (1) Mildly ill-posed case: let $\mu_j \asymp j^{-\varsigma/2}$ for $\varsigma \geq 0$ and $ a+\varsigma > -1$. Let $J^{5 \vee (4 + \varsigma-a)} (\log n)^{3} /n = o(1)$ and $nJ^{-(p+a+\varsigma+1)}{(\log J)} = o(1)$. Then Assumption U-CS(iii)(iv) holds.\\ (2) Severely ill-posed case: let $\mu_J \asymp \exp(-\frac{1}{2}j^{\varsigma/2})$, $\varsigma > 0$. Let $J = (\log(n/(\log n)^\varrho))^{2/\varsigma}$ with $\varrho > 0$ chosen such that $2p > \varrho \varsigma - 2a$ and $\varrho \varsigma > 10-2a$. Then Assumption U-CS(iii)(iv) holds.

The next results are about the uniform Gaussian process strong approximation and validity of score bootstrap UCBs for exact CS and DL functionals.

theoremLet Assumptions CS(i)(ii)(iv) and U-CS(i)(ii)(iii) hold. Then:\\ (1) If Assumption U-CS(iv.1) holds, then Result ((ref)) (with $r_n=(\log J)^{-1/2}$) holds for $f_t = f_{CS,t}$;\\ (2) If Assumption U-CS(iv.2) holds, then Result ((ref)) also holds for $f_t = f_{CS,t}$.

In the next theorem the condition $\ul \sigma_n \asymp \mu_J^{-1} \sqrt J$ is implied by the assumption that $\sigma_n (f_{DL,t}) \asymp \mu_J^{-1} \sqrt J$ uniformly for $t \in \mc T$, which is reasonable for the DL functional.

theoremLet Assumptions CS(i)(ii)(iv) and U-CS(i)(ii)(iii) hold with $\ul \sigma_n \asymp \mu_J^{-1} \sqrt J$. Then:\\ (1) If Assumption U-CS(iv.1) holds, then Result ((ref)) (with $r_n=(\log J)^{-1/2}$) holds for $f_t = f_{DL,t}$;\\ (2) If Assumption U-CS(iv.2) holds, then Result ((ref)) also holds for $f_t = f_{DL,t}$.

Inference on welfare functionals without endogeneity

This subsection specializes the pointwise and uniform inference results for welfare functionals from the preceding subsections to nonparametric demand estimation with exogenous price and income. Precisely, we let $\mf X_i = \mf W_i$, $J = K$, $b^K = \psi^J$, $\mu_J \asymp 1$, $\tau_J \asymp 1$ and so the sieve NPIV estimator reduces to the usual series LS estimator of $h_0 (x)=E[Y_i |W_i=x]$.

The next two corollaries are direct consequences of our Theorems (ref), (ref) and (ref) for pointwise asymptotic normality of sieve $t$ statistics for exact CS and DL and approximate CS functionals under exogeneity, and hence the proofs are omitted.

corollaryLet Assumption CS(i)--(iv) hold with $\mf X_i = \mf W_i$, $J = K$, $b^K = \psi^J$ and $\mu_J \asymp 1$ and let $\sum_{j=1}^J a_j^2 \gtrsim J^{a+1}$ with $ 0\geq a \geq -1$. (1) Let $n J^{-(p+a+1)} = o(1)$, $J^{3-a} (\log J)/n = o(1)$, and $\delta \geq 2/(2-a)$. Then: the sieve $t$-statistic for $f_{CS}(h_0)$ is asymptotically $N(0,1)$. (2) Let $n J^{-(p+1)} = o(1)$, $J^3 (\log J)/n = o(1)$, and $a=0$, $\delta \geq 1$. Then: the sieve $t$-statistic for $f_{DL}(h_0)$ is asymptotically $N(0,1)$.

Previously HausmanNewey1995 established the pointwise asymptotic normality of $t$-statistics for exact CS and DL based on plug-in kernel LS estimators of demand without endogeneity. They also established root-$n$ asymptotic normality of t-statistics for averaged exact CS and DL (i.e. CS/DL averaged over a range of incomes) based on plug-in power series LS estimator of demand without endogeneity, under some regularity conditions including that $\sup_{x} E[|u_i|^4|\mf X_i = x]<\infty$ (which, in our notation, implies $\delta = 2$), $p =\infty$ (i.e., $h_0$ is infinitely times differentiable) and $J^{22}/n = o(1)$. Corollary (ref) complements their work by providing conditions for the pointwise asymptotic normality of exact CS and DL functionals based on spline and wavelet LS estimators of demand.

corollaryLet Assumption CS(i)--(iv) hold for the log-demand model ((ref)) with $\mf X_i = \mf W_i$, $J = K$, $b^K = \psi^J$, $\mu_J \asymp 1$ and $p > 0$, let $\sum_{j=1}^J a_j^2 \gtrsim J^{c+1}$ with $0\geq c \geq -1$. Let $n J^{-(p+c+1)} = o(1)$, $ J^{2-c} (\log J)/n = o(1)$ and $\delta \geq 2/(1-c)$. Then: the sieve $t$ statistic for $f_{A}(h_0)$ is asymptotically $N(0,1)$.

Previously Newey1997 established the pointwise asymptotic normality of $t$-statistics for approximate CS functionals based on plug-in series LS estimators of exogenous demand under some regularity conditions including that $\sup_{x} E[|u_i|^4|\mf X_i = x]<\infty$ (which implies $\delta = 2$), $n J^{-p}=o(1)$ and either $J^6/n = o(1)$ for power series or $J^4/n=o(1)$ for splines.

The final corollary is a direct consequence of our Theorems (ref) and (ref) and Remark (ref) for uniform inferences based on sieve $t$ processes for exact CS and DL nonlinear functionals under exogeneity, and hence its proof is omitted.

corollaryLet Assumptions CS(i)(ii)(iv) and U-CS(i)(ii) hold with $\mf X_i = \mf W_i$, $J = K$, $b^K = \psi^J$ and $\mu_J \asymp 1$. Let $\ul \sigma_n^2 \gtrsim J^{a+1}$ with $ 0\geq a \geq -1$. Let $J^5(\log n)^{3} /n = o(1)$ and $nJ^{-(p+a+1)}{(\log J)} = o(1)$. Then: Results ((ref)) (with $r_n=(\log J)^{-1/2}$) and ((ref)) hold for $f_t =f_{CS,t}, f_{DL,t}$.

We note that $\ul \sigma_n^2 \asymp J$ (or $a=0$) for $f_t =f_{DL,t}$. Corollary (ref) appears to be a new addition to the existing literature. The sufficient conditions for uniform inference for collections of nonlinear exact CS and DL functionals of nonparametric demand estimation under exogeneity are mild and simple.

Conclusion

This paper makes several important contributions to inference on nonparametric models with endogeneity. We derive the minimax sup-norm convergence rates for estimating the structural NPIV function $h_0$ and its derivatives. We also provide upper bounds for sup-norm convergence rates of computationally simple sieve NPIV (series 2SLS) estimators using any sieve basis to approximate unknown $h_0$, and show that the sieve NPIV estimator using spline or wavelet basis can attain the minimax sup-norm rates. These rate results are particularly useful for establishing validity of pointwise and uniform inference procedures for nonlinear functionals of $h_0$. In particular, we use our sup-norm rates to establish the uniform Gaussian process strong approximation and the validity of score bootstrap-based UCBs for collections of nonlinear functionals of $h_0$ under primitive conditions, allowing for mildly and severely ill-posed problems. We illustrate the usefulness of our UCBs procedure with two real data applications to nonparametric demand analysis with endogeneity. We establish the pointwise and uniform limit theories for sieve $t$-statistics for exact (and approximate) CS and DL nonlinear functionals under low-level conditions when the demand function is estimated via sieve NPIV. Our theoretical and empirical results for CS and DL are new additions to the literature on nonparametric welfare analysis.

We conclude the paper by mentioning some further extensions and applications of sup-norm convergence rates of sieve NPIV estimators.

Extensions to semiparametric IV models. Although our rate results are presented for purely nonparametric IV models, the results may be adapted easily to some semiparametric models with nonparametric endogeneity, such as partially linear IV regression AiChen2003,Florens-linear, shape-invariant Engel curve IV regression BCK, and single index IV regression CCLN, to list a few. For example, consider the partially linear NPIV model \[ Y_i = X_{1i}'\beta_0 + h_0(X_{2i}) + u_i~\quad\quad E[u_i|W_{1i},W_{2i}]=0~ \] where $X_{1i}$ and $X_{2i}$ are of dimensions $d_1$ and $d_2$ and do not contain elements in common, and $W_i = (W_{1i},W_{2i})$ is the (conditional) IV. See Florens-linear,CCLN for identification of $(\beta_0, h_0)$ in this model. We can still estimate $(\beta_0, h_0)$ via sieve NPIV or series 2SLS as before, replacing $\Psi$ and $B$ in equations ((ref))--((ref)) by:

align*[align* omitted — 196 chars of source]

where $x = (x_1',x_2')'$, $w = (w_1',w_2')'$, $\psi_{J1},\ldots,\psi_{JJ}$ denotes a sieve of dimension $J$ for approximating $h_0(x_2)$ and $b_{K1},\ldots,b_{KK}$ denotes a sieve of dimension $K$ for the instrument space for $W_2$. We then partition $\wh c$ in ((ref)) into $\wh c = (\wh \beta'_{\phantom{2}},\wh c_2')'$ and set $\widehat h(x) = \psi^J_2(x)'\widehat c_2$. Note that $\wh \beta$ is root-$n$ consistent and asymptotically normal for $\beta_0$ under mild conditions (see AiChen2003,ChenPouzo2009), and hence would not affect the optimal convergence rate of $\wh h$ to $h_0$. Our rate results may be slightly altered to derive sup-norm convergence rates for $\widehat h$ and its derivatives.

Nonparametric specification testing in NPIV models. Structural models may specify a parametric form $m_{\theta_0}(x)$ where $\theta_0 \in \Theta \subseteq \mb R^{d_\theta}$ for the unknown structural function $h_0(x )$ in NPIV model ((ref)). We may be interested in testing the parametric model $\{m_\theta : \theta \in \Theta\}$ against a nonparametric alternative that only assumes some smoothness on $h_0$. Specification tests for nonparametric regression without endogeneity have typically been performed via either a quadratic-form-based statistic or a Kolmogorov-Smirnov (KS) type sup statistic.\footnote{See, e.g., Bierens82, HardleMammen, HongWhite, FanLi, LavergneVuong, StinchcombeWhite1998 and HorowitzSpokoiny to list a few.} However, specification tests for NPIV models have so far only been performed via quadratic-form-based statistics; see, e.g., Horowitz2006,Horowitz2011,Horowitz2012,BlundellHorowitz2007,Breunig2015. Equipped with our sup-norm rate and UCBs results for NPIV function and its derivatives, one could also perform specification tests in NPIV models using KS type statistics of the form \[ T_n = \sup_{x} \frac{\Big| \wh h (x) - \wh m(x,\wh \theta) \Big|}{s_n(x)} \] where $\wh \theta$ is a first-stage estimator of $\theta_0$, and $\wh m (x,\wh \theta)$ is obtained from series 2SLS regression of $m(X_1,\wh \theta),\ldots,m(X_n,\wh \theta)$ on the same basis functions as in $\wh h$, and $s_n(x)$ is a normalization factor. Alternatively, one could consider a KS statistic formed in terms of the projection of $[\wh h (x) - \wh m(x,\wh \theta)]$ onto the instrument space. Sup-norm convergence rates and uniform limit theory derived in this paper would be useful in deriving the large-sample distribution of these KS type statistics. Further, based on our rate results (in sup- and $L^2$-norm) for estimating derivatives of $h_0$ in a NPIV model, one could also perform nonparametric tests of significance by testing whether partial derivatives of NPIV function $h_0$ are identically zero, via KS or quadratic-form-based test statistics.

If one is interested in specifications or inferences on functionals directly, then one might consider KS type sup statistics for (possibly nonlinear) functionals directly. For example, if one is interested in exact CS functional of a demand and concerns about the potential endogeneity of price. Then one could estimate exact CS functional using a series LS estimated demand (under exogeneity) and series 2SLS estimated demand (under endogeneity), and then compare the two estimated exact CS functionals via KS type or quadratic-form-based test. In fact, the score bootstrap-based UCBs reported in Figure (ref) indicates that such a test based on exact CS functional directly could be quite informative.

Semiparametric 2-step procedures with NPIV first stage. Many semiparametric two-step or multi-step estimation and inference procedures involve a nonparametric first stage. There are many theoretical results when the first stage is a purely nonparametric LS regression (without endogeneity) and its sup-norm convergence rate is used to assistant subsequent analysis. For structural estimation and inference, it is natural to allow for the presence of nonparametric endogeneity in the first stage as well. For instance, if there is endogeneity present in the conditional moment inequality application of the famous intersection bound paper of CLR, one could simply use our sup-norm rate and UCBs results for sieve NPIV instead of their series LS regression in the first stage. As another example, consider semiparametric two-step GMM models $E[g(Z_i,\theta_0,h_0(X_i))] =0$, where $h_0$ is the NPIV function in model ((ref)), $g$ is a $\mb R^{d_g}$-valued vector of moment functions with $d_g \geq d_\theta$, and $\theta_0 \in \mb R^{d_\theta}$ is a finite-dimensional parameter of interest, such as the average exact CS parameter of a nonparametric demand function with endogeneity. A popular estimator $\wh \theta$ of $\theta_0$ is a solution to the semiparametric two-step GMM with a weighting matrix $\wh W$: \[ \min_{\theta } \left( \frac{1}{n} \sum_{i=1}^n g(Z_i,\theta,\wh h(X_i)) \right)' \wh W \left( \frac{1}{n} \sum_{i=1}^n g(Z_i,\theta,\wh h(X_i)) \right) \] where $\wh h$ is a sieve NPIV estimator of $h_0$. When $h_0$ enters the moment function $g(\cdot)$ nonlinearly, sup-norm convergence rates of $\wh h$ to $h_0$ are useful in deriving the asymptotic properties of $\wh \theta$.