EconBase
← Back to paper

Sieve Wald and QLR Inferences on Semi/nonparametric Conditional Moment Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

179,042 characters · 27 sections · 130 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sieve Wald and QLR Inferences on Semi/nonparametric Conditional Moment Models

frontmatter\runtitle{Sieve Wald and QLR Inference} \thankstext{r1}{ Earlier versions, some entitled \textquotedblleft On PSMD inference of functionals of nonparametric conditional moment restrictions\textquotedblright ,\ were presented in April 2009 at the Banff conference on seminonparametrics, in June 2009 at the Cemmap conference on quantile regression, in July 2009 at the SITE conference on nonparametrics, in September 2009 at the Stats in the Chateau/France, in June 2010 at the Cemmep workshop on recent developments in nonparametric instrumental variable methods, in August 2010 at the Beijing international conference on statistics and society, and econometric workshops in numerous universities. We thank a co-editor, two referees, Don Andrews, Peter Bickel, Gary Chamberlain, Tim Christensen, Michael Jansson, Jim Powell and especially Andres Santos for helpful comments. We thank Yinjia Qiu for excellent research assistant in simulations using R. Chen acknowledges financial support from National Science Foundation grant SES-0838161 and Cowles Foundation. Any errors are the responsibility of the authors.} \address{Chen: Cowles Foundation for Research in Economics, Yale University, Box 208281, New Haven, CT 06520, USA. Email: [email removed]. Pouzo: Department of Economics, UC Berkeley, 530 Evans Hall 3880, Berkeley, CA 94720, USA. Email: [email removed].} \runauthor{X. Chen and D. Pouzo} \begin{abstract} This paper considers inference on functionals of semi/nonparametric conditional moment restrictions with possibly nonsmooth generalized residuals, which include all of the (nonlinear) nonparametric instrumental variables (IV) as special cases. These models are often ill-posed and hence it is difficult to verify whether a (possibly nonlinear) functional is root-$n$ estimable or not. We provide computationally simple, unified inference procedures that are asymptotically valid regardless of whether a functional is root-$n$ estimable or not. We establish the following new useful results: (1) the asymptotic normality of a plug-in penalized sieve minimum distance (PSMD) estimator of a (possibly nonlinear) functional; (2) the consistency of simple sieve variance estimators for the plug-in PSMD estimator, and hence the asymptotic chi-square distribution of the sieve Wald statistic; (3) the asymptotic chi-square distribution of an optimally weighted sieve quasi likelihood ratio (QLR) test under the null hypothesis; (4) the asymptotic tight distribution of a non-optimally weighted sieve QLR statistic under the null; (5) the consistency of generalized residual bootstrap sieve Wald and QLR tests; (6) local power properties of sieve Wald and QLR tests and of their bootstrap versions; (7) asymptotic properties of sieve Wald and SQLR for functionals of increasing dimension. Simulation studies and an empirical illustration of a nonparametric quantile IV regression are presented. \end{abstract} \begin{keyword} Nonlinear nonparametric instrumental variables; Penalized sieve minimum distance; Irregular functional; Sieve variance estimators; Sieve Wald; Sieve quasi likelihood ratio; Generalized residual bootstrap; Local power; Wilks phenomenon. \end{keyword}

\setcounter{page}{0} \thispagestyle{empty}

\baselineskip=18pt

Introduction

This paper is about inference on functionals of the unknown true parameters $ \alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$ satisfying the semi/nonparametric conditional moment restrictions

equation[equation omitted — 80 chars of source]

where $Y$ is a vector of endogenous variables and $X$ is a vector of conditioning (or instrumental) variables. The conditional distribution of $Y$ given $X$, $F_{Y|X}$, is not specified beyond that it satisfies ((ref)). $\rho (\cdot ;\theta _{0},h_{0})$ is a $d_{\rho }\times 1-$vector of generalized residual functions whose functional forms are known up to the unknown parameters $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})\in \Theta \times \mathcal{H}$, with $\theta _{0}\equiv (\theta _{01},...,\theta _{0d_{\theta }})^{\prime }\in \Theta $ being a $d_{\theta }\times 1-$vector of finite dimensional parameters and $h_{0}\equiv (h_{01}(\cdot ),...,h_{0q}(\cdot ))\in \mathcal{H}$ being a $1\times d_{q}-$vector valued function. The arguments of each unknown function $h_{\ell }(\cdot )$ may differ across $\ell =1,...,q$, may depend on $\theta ,$ $h_{\ell ^{\prime }}(\cdot ),$ $\ell ^{\prime }\neq \ell $, $X$ and $Y$. The residual function $\rho (\cdot ;\alpha )$ could be nonlinear and pointwise non-smooth in the parameters $\alpha \equiv (\theta ^{\prime },h)\in \Theta \times \mathcal{H}$ .

The general framework ((ref)) nests many widely used nonparametric and semiparametric models in economics and finance. Well known examples include nonparametric mean instrumental variables regressions (NPIV): $ E[Y_{1}-h_{0}(Y_{2})|X]=0$ (e.g., HH_Ann05, CFR_bookchp07, BCK_Emetrica07, DFR_wp10, Horowitz_ECMA11); nonparametric quantile instrumental variables regressions (NPQIV): $ E[1\{Y_{1}\leq h_{0}(Y_{2})\}-\gamma |X]=0$ (e.g., CH_Emetrica05, CIN_JOE07, HL_Emetrica07, CP_WP07, CGS_WP08); semi/nonparametric demand models with endogeneity (e.g., BCK_Emetrica07, CP_WP07a, Souza2012); semi/nonparametric random coefficient panel data regressions (e.g., CHAMBERLAIN_ECMA92, GPowell2012); semi/nonparametric spatial models with endogeneity (e.g., Pinkse2002, MerloPaula); semi/nonparametric asset pricing models (e.g., Hansen_Richard_ECMA87, Gallant_Tauchen_ECMA89, ChenLudvigson_JAE09, ChenLudvigson2013, Sentana2013); semi/nonparametric static and dynamic game models (e.g., BHN_WP11); nonparametric optimal endogenous contract models (e.g., BMT_WP12). Additional examples of the general model ((ref)) can be found in CHAMBERLAIN_ECMA92, NP_ECMA03, AC_Emetrica03, CP_WP07, CCLN_WP10 and the references therein. In fact, model ((ref)) includes all of the (nonlinear) semi/nonparametric IV regressions when the unknown functions $ h_{0}$ depend on the endogenous variables $Y$:

equation[equation omitted — 88 chars of source]

which could lead to difficult (nonlinear) nonparametric ill-posed inverse problems with unknown operators.

Let $\left\{ Z_{i}\equiv (Y_{i}^{\prime },X_{i}^{\prime })^{\prime }\right\} _{i=1}^{n}$ be a random sample from the distribution of $Z\equiv (Y^{\prime },X^{\prime })^{\prime }$ that satisfies the conditional moment restrictions ((ref)) with a unique $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0}) $. Let $\phi :\Theta \times \mathcal{H}\rightarrow \mathbb{R} ^{d_{\phi }}$ be a (possibly nonlinear) functional with a finite $d_{\phi }\geq 1$. Typical linear functionals include an Euclidean functional $\phi (\alpha )=\theta $, a point evaluation functional $\phi (\alpha )=h( \overline{y}_{2}) $ (for $\overline{y}_{2}\in $ supp$(Y_{2})$), a weighted derivative functional $\phi (h)=\int w(y_{2})\nabla h(y_{2})dy_{2}$ and many others. Typical nonlinear functionals include a quadratic functional $\int w(y_{2})\left\vert h(y_{2})\right\vert ^{2}dy_{2}$, a quadratic derivative functional $\int w(y_{2})\left\vert \nabla h(y_{2})\right\vert ^{2}dy_{2}$, a consumer surplus or an average consumer surplus functional of an endogenous demand function $h$. We are interested in computationally simple, valid inferences on any $\phi (\alpha _{0})$ of the general model ((ref)) with i.i.d. data.\footnote{ See our Cowles Foundation Discussion Paper No. 1897 for general theory allowing for weakly dependent data.}

Although some functionals of the model ((ref)), such as the (point) evaluation functional, are known a priori to be estimated at slower than root-$n$ rates, others, such as the weighted derivative functional, are far less clear without a stare at their semiparametric efficiency bound expressions. This is because a non-singular efficiency bound is a necessary condition for $\phi (\alpha _{0})$ to be estimated at a root-$n$ rate. Unfortunately, as pointed out in CHAMBERLAIN_ECMA92 and AC_WP05 , there is generally no closed form solution for the efficiency bound of $ \phi (\alpha _{0})$ (including $\theta _{0}$) of model ((ref)), especially so when $\rho (\cdot ;\theta _{0},h_{0})$ contains several unknown functions and/or when the unknown functions $h_{0}$ of endogenous variables enter $\rho (\cdot ;\theta _{0},h_{0})$ nonlinearly. It is thus difficult to verify whether the efficiency bound for $\phi (\alpha _{0})$ is singular or not. Therefore, it is highly desirable for applied researchers to be able to conduct simple valid inferences on $\phi (\alpha _{0})$ regardless of whether it is root-$n$ estimable or not. This is the main goal of our paper.

In this paper, for the general model ((ref)) that could be nonlinearly ill-posed and for any $\phi (\alpha _{0})$ that may or may not be root-$n$ estimable, we first establish the asymptotic normality of the plug-in penalized sieve minimum distance (PSMD) estimator $\phi (\widehat{ \alpha }_{n})$ of $\phi (\alpha _{0})$. For the model ((ref)) with (pointwise) smooth residuals $\rho (Z;\alpha )$ in $\alpha _{0}$, we propose two simple consistent sieve variance estimators for possibly slower than root-$n$ estimator $\phi (\widehat{\alpha }_{n})$, which immediately leads to the asymptotic chi-square distribution of the sieve Wald statistic. However, there is no simple variance estimator for $\phi (\widehat{\alpha } _{n})$ when $\rho (Z,\alpha )$ is not pointwise smooth in $\alpha _{0}$ (without estimating an extra unknown nuisance function or using numerical derivatives). We then consider a PSMD criterion based test of the null hypothesis $\phi (\alpha _{0})=\phi _{0}$. We show that an optimally weighted sieve quasi likelihood ratio (SQLR) statistic is asymptotically chi-square distributed under the null hypothesis. This allows us to construct confidence sets for $\phi (\alpha _{0})$ by inverting the optimally weighted SQLR statistic, without the need to compute a variance estimator for $\phi (\widehat{\alpha }_{n})$. Nevertheless, in complicated real data analysis applied researchers might like to use simple but possibly non-optimally weighed PSMD procedures for estimation of and inference on $ \phi (\alpha _{0})$. We show that the non-optimally weighted SQLR statistic still has a tight limiting distribution under the null regardless of whether $\phi (\alpha _{0})$ is root-$n$ estimable or not. In addition, we establish the consistency of the generalized residual bootstrap (possibly non-optimally weighted) SQLR and sieve Wald tests under virtually the same conditions as those used to derive the limiting distributions of the original-sample statistics. The bootstrap SQLR would then lead to alternative confidence sets construction for $\phi (\alpha _{0})$ without the need to compute a variance estimator for $\phi (\widehat{\alpha }_{n})$. To ease notation burden, we present the above listed theoretical results for a scalar-valued functional in the main text. In Appendix (ref) we present the asymptotic properties of sieve Wald and SQLR for functionals of increasing dimension (i.e., $d_{\phi }=dim(\phi )$ could grow with sample size $n$). We also provide the local power properties of sieve Wald and SQLR tests as well as their bootstrap versions in Appendix (ref). Regardless of whether a possibly nonlinear functional $\phi (\alpha _{0})$ is root-$n$ estimable or not, we show that the optimally weighted SQLR is more powerful than the non-optimally weighed SQLR, and that the SQLR and the sieve Wald using the same weighting matrix have the same local power in terms of first order asymptotic theory.

To the best of our knowledge, our paper is the first to provide a unified theory about sieve Wald and SQLR inferences on (possibly nonlinear) $\phi (\alpha _{0})$ satisfying the general semi/nonparametric model ((ref) ) with possibly non-smooth residuals.\footnote{ We also provide asymptotic properties of sieve score and bootstrap sieve score statistics in the online Appendix (ref).} Our results allow applied researchers to obtain limiting distribution of the plug-in PSMD estimator $\phi (\widehat{\alpha } _{n})$ and to construct confidence sets for any $\phi (\alpha _{0})$ regardless of whether it is root-$n$ estimable or not. Our paper is also the first to provide local power properties of sieve Wald and SQLR tests and their bootstrap versions of general nonlinear hypotheses for the model ((ref)).

Roughly speaking, our results extend the classical theories on Wald and QLR tests of nonlinear hypothesis based on root-$n$ consistent parametric minimum distance estimator $\widehat{\alpha }_{n}$ to those based on slower than root-$n$ consistent nonparametric minimum distance estimator $\widehat{ \alpha }_{n}\equiv (\widehat{\theta }_{n}^{\prime },\widehat{h}_{n})$ of $ \alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$ satisfying the model ((ref)). The implementations of the sieve Wald and SQLR also resemble the classical Wald and QLR based on parametric extreme estimators and hence are computationally attractive. For example, our sieve t (Wald) test on a general nonlinear hypothesis $\phi (h_{0})=\phi _{0}$ of the NPIV model $ E[Y_{1}-h_{0}(Y_{2})|X]=0$ can be implemented as a standard t (Wald) test for a parametric linear IV model using two stage least squares (see Subsection (ref)). The proof techniques are quite different, however, because one is no longer able to rely on the root-$n$ asymptotic normality of $\widehat{\alpha }_{n}$ and then a standard \textquotedblleft delta-method\textquotedblright\ to establish the asymptotic normality of $ \sqrt{n}\left( \phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})\right) $. In our framework ((ref)), $\sqrt{n}\left( \phi (\widehat{\alpha } _{n})-\phi (\alpha _{0})\right) $ could diverge to infinity under the combined effects of (i) slower convergence rate of $\widehat{\alpha }_{n}$ to $\alpha _{0}$ due to the ill-posed inverse problem and (ii) nonlinearity of either the functional $\phi ()$ or the residual function $\rho ()$ in $h$ . Our proof strategy relies on the convergence rates of the PSMD estimator $ \widehat{\alpha }_{n}$ to $\alpha _{0}$ in both weak and strong metrics, and then the local curvatures of the functional $\phi ()$ and the criterion function under these two metrics. The weak metric is intrinsic to the variance of the linear approximation to $\phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})$, while the strong metric controls the nonlinearity (in $ \alpha $) of the functional $\phi ()$ and of the conditional mean function $ m(\cdot ,\alpha )=E[\rho (Y,X;\alpha )|X=\cdot ]$. Unfortunately the convergence rate in the strong metric could be very slow due to the illposed inverse problem. This explains why it is difficult to establish the asymptotic normality of $\phi (\widehat{\alpha }_{n})$ for a nonlinear functional $\phi ()$ even in the NPIV model. Our paper builds upon the recent results on convergence rates in CP_WP07 and others. In particular, under virtually the same conditions as those in CP_WP07, we show that our generalized residual bootstrap PSMD estimator of $\alpha _{0}$ is consistent and achieves the same convergence rates as that of the original-sample PSMD estimator $\widehat{\alpha }_{n}$. This result is then used to establish the consistency of the bootstrap sieve Wald and the bootstrap SQLR statistics under virtually the same conditions as those used to derive the limiting distributions of the original-sample statistics. \footnote{ The convergence rate of the bootstrap PSMD estimator is also very useful for the consistency of the bootstrap Wald statistic for semiparametric two-step GMM estimation of Euclidean parameters when the first-step unknown functions are estimated via a PSMD procedure. See e.g., CLvK_Emetrica03}

There are some published work about estimation of and inference on a particular linear functional, the Euclidean parameter $\phi (\alpha )=\theta $, of the general model ((ref)) when $\theta _{0}$ is assumed to be root-$n$ estimable; see AC_Emetrica03, CP_WP07a, OTSU_WP11 and others. None of the existing work allows for $\theta _{0}$ being irregular (i.e., slower than root-$n$ estimable),\footnote{ It is known that $\theta _{0}$ could have singular semiparametric efficiency bound and could not be root-$n$ estimable; see CHAMBERLAIN2010, KhanTamer2010, GPowell2012 and the references therein. Following KhanTamer2010 and GPowell2012 we call such a $\theta _{0}$ irregular. Many applied papers on complicated semi/nonparametric models simply assume that $\theta _{0}$ is root-$n$ estimable.} however. When specializing our general theory to inference on $\theta _{0}$ of the model ( (ref)), we not only recover the results of AC_Emetrica03 and CP_WP07a, but also provide local power properties of sieve Wald and SQLR as well as valid bootstrap (possibly non-optimally weighted) SQLR inference. Moreover, our results remain valid even when $\theta _{0}$ is irregular.

When specializing our theory to inference on a particular irregular linear functional, the point evaluation functional $\phi (\alpha )=h(\overline{y} _{2})$, of the semi/nonparametric IV model ((ref)), we automatically obtain the pointwise asymptotic normality of the PSMD estimator of $h_{0}( \overline{y}_{2})$ and different ways to construct its confidence set. These results are directly applicable to the NPIV example with $\rho (Y_{1};\theta _{0},h_{0}(Y_{2}))=Y_{1}-h_{0}(Y_{2})$ and to the NPQIV example with $\rho (Y_{1};\theta _{0},h_{0}(Y_{2}))=1\{Y_{1}\leq h_{0}(Y_{2})\}-\gamma $. Previously, Horowitz_07 and CGS_WP08 established the pointwise asymptotic normality of their kernel based function space Tikhonov regularization estimators of $h_{0}(\overline{y}_{2})$ for the NPIV and the NPQIV examples respectively. Immediately after our paper was first presented in April 2009 Banff/Canada conference on semiparametrics, the authors of HL_WP10 informed us that they were concurrently working on confidence bands for $h_{0}$ using a particular SMD estimator of the NPIV example. To the best of our knowledge, there is no inference results, in the existing literature, on any nonlinear functional of $h_{0}$ even for the NPIV and NPQIV examples. Our paper is the first to provide simple sieve Wald and SQLR tests for (possibly) nonlinear functionals satisfying the general semi/nonparametric IV model ((ref)).

The rest of the paper is organized as follows. Section (ref) presents the plug-in PSMD estimator $\phi (\widehat{\alpha }_{n})$ of a (possibly nonlinear) functional $\phi $ evaluated at $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$ satisfying the model ((ref)). It also provides an overview of the main asymptotic results that will be established in the subsequent sections, and illustrates the applications through a point evaluation functional $\phi (\alpha )=h(\overline{y}_{2})$, a weighted derivative functional $\phi (h)=\int w(y_{2})\nabla h(y_{2})dy_{2}$, and a quadratic functional $\phi (\alpha )=\int w(y_{2})\left\vert h(y_{2})\right\vert ^{2}dy_{2}$ of the NPIV and NPQIV examples. Section (ref) states the basic regularity conditions. Section (ref) provides the asymptotic properties of sieve t (Wald) and sieve QLR statistics. Section (ref) establishes the consistency of the bootstrap sieve t (Wald) and the bootstrap SQLR statistics. Section (ref) verifies the key regularity conditions for the asymptotic theories via the three functionals of the NPIV and NPQIV examples presented in Section (ref). Section (ref) presents simulation studies and an empirical illustration. Section (ref) briefly concludes. Appendix (ref) consists of several subsections, presenting (1) further results on sieve Riesz representation of a functional of interest; (2) the convergence rates of the bootstrap PSMD estimator $\widehat{\alpha }_{n}^{B}$ for model ((ref)); (3) the local power properties of sieve Wald and SQLR tests and of their bootstrap versions; (4) asymptotic properties of sieve Wald and SQLR for functionals of increasing dimension; (5) low level sufficient conditions with a series least squares (LS) estimated conditional mean function $m(\cdot ,\alpha )=E[\rho (Y,X;\alpha )|X=\cdot ]$; and (6) additional useful lemmas with series LS estimated $m(\cdot ,\alpha )$. Online supplemental materials consist of Appendices (ref), (ref) and (ref). Appendix (ref) contains additional theoretical results (including other consistent variance estimators and other bootstrap sieve Wald tests) and proofs of all the results stated in the main text. Appendix (ref) contains proofs of all the results stated in Appendix (ref). The online Appendix (ref) provides computationally attractive sieve score test and sieve score bootstrap.

Notation. We use \textquotedblleft $\equiv $\textquotedblright\ to implicitly define a term or introduce a notation. For any column vector $A$, we let $A^{\prime }$ denote its transpose and $||A||_{e}$ its Euclidean norm (i.e., $||A||_{e}\equiv \sqrt{A^{\prime }A}$, although sometimes we use $ |A|=||A||_{e}$ for simplicity). Let $||A||_{W}^{2}\equiv A^{\prime }WA$ for a positive definite weighting matrix $W$. Let $\lambda _{\max }(W)$ and $ \lambda _{\min }(W)$ denote the maximal and minimal eigenvalues of $W$ respectively. All random variables $Z\equiv (Y^{\prime },X^{\prime })^{\prime }$, $Z_{i}\equiv (Y_{i}^{\prime },X_{i}^{\prime })^{\prime }$ are defined on a complete probability space $(\mathcal{Z},\mathcal{B}_{Z},P_{Z})$ , where $P_{Z}$ is the joint probability distribution of $(Y^{\prime },X^{\prime })$. We define $(\mathcal{Z}^{\infty },\mathcal{B}_{Z}^{\infty },P_{Z^{\infty }})$ as the probability space of the sequences $ (Z_{1},Z_{2},...)$. For simplicity we assume that $Y$ and $X$ are continuous random variables. Let $f_{X}$ ($F_{X}$) be the marginal density (cdf) of $X$ with support $\mathcal{X}$, and $f_{Y|X}$ ($F_{Y|X}$) be the conditional density (cdf) of $Y$ given $X$. Let $E_{P}[\cdot ]$ denote the expectation with respect to a measure $P$. Sometimes we use $P$ for $P_{Z^{\infty }}$ and $E[\cdot ]$ for $E_{P_{Z^{\infty }}}[\cdot ]$. Denote $L^{p}(\Omega ,d\mu )$, $1\leq p<\infty $, as a space of measurable functions with $ ||g||_{L^{p}(\Omega ,d\mu )}\equiv \{\int_{\Omega }|g(t)|^{p}d\mu (t)\}^{1/p}<\infty $, where $\Omega $ is the support of the sigma-finite positive measure $d\mu $ (sometimes $L^{p}(d\mu )$ and $||g||_{L^{p}(d\mu )}$ are used). For any (possibly random) positive sequences $\{a_{n}\}_{n=1}^{ \infty }$ and $\{b_{n}\}_{n=1}^{\infty }$, $a_{n}=O_{P}(b_{n})$ means that $ \lim_{c\rightarrow \infty }\limsup_{n}\Pr \left( a_{n}/b_{n}>c\right) =0$; $ a_{n}=o_{P}(b_{n})$ means that for all $\varepsilon >0$, $\lim_{n\rightarrow \infty }\Pr \left( a_{n}/b_{n}>\varepsilon \right) =0$; and $a_{n}\asymp b_{n}$ means that there exist two constants $0<c_{1}\leq c_{2}<\infty $ such that $c_{1}a_{n}\leq b_{n}\leq c_{2}a_{n}$. Also, we use \textquotedblleft wpa1-$P_{Z^{\infty }}$\textquotedblright\ (or simply wpa1) for an event $ A_{n}$, to denote that $P_{Z^{\infty }}(A_{n})\rightarrow 1$ as $ n\rightarrow \infty $. We use $\mathcal{A}_{n}\equiv \mathcal{A}_{k(n)}$ and $\mathcal{H}_{n}\equiv \mathcal{H}_{k(n)}$ for various sieve spaces. We assume $\dim (\mathcal{A}_{k(n)})\asymp \dim (\mathcal{H}_{k(n)})\asymp k(n)$ for simplicity, all of which grow to infinity with the sample size $n$. We use $const.$, $c$ or $C$ to mean a positive finite constant that is independent of sample size but can take different values at different places. For sequences, $(a_{n})_{n}$, we sometimes use $a_{n}\nearrow a$ ($ a_{n}\searrow a$) to denote, that the sequence converges to $a$ and that is increasing (decreasing) sequence. For any mapping $\digamma :\mathbf{H} _{1}\rightarrow \mathbf{H}_{2}$ between two generic Banach spaces, $\frac{ d\digamma (\alpha _{0})}{d\alpha }[v]\equiv \left. \frac{\partial \digamma (\alpha _{0}+\tau v)}{\partial \tau }\right\vert _{\tau =0}$ is the pathwise (or Gateaux) derivative at $\alpha _{0}$ in the direction $v\in \mathbf{H} _{1}$. And $\frac{d\digamma (\alpha _{0})}{d\alpha }[\mathbf{v}^{\prime }]\equiv \left( \frac{d\digamma (\alpha _{0})}{d\alpha }[v_{1}],\cdot \cdot \cdot ,\frac{d\digamma (\alpha _{0})}{d\alpha }[v_{k}]\right) $ for $\mathbf{ v}^{\prime }=\left( v_{1},\cdot \cdot \cdot ,v_{k}\right) $ with $v_{j}\in \mathbf{H}_{1}$ for all $j=1,...,k$.

PSMD Estimation and Inferences: An Overview

The Penalized Sieve Minimum Distance Estimator

Let $m(X,\alpha )\equiv E\left[ \rho (Y,X;\alpha )|X\right] =\int \rho (y,X;\alpha )dF_{Y|X}(y)$ be a $d_{\rho }\times 1$ vector valued conditional mean function, $\Sigma (X)$ be a $d_{\rho }\times d_{\rho }$ positive definite ($a.s.-X$) weighting matrix, and

equation*[equation* omitted — 157 chars of source]

be the population minimum distance (MD) criterion function. Then the semi/nonparametric conditional moment model ((ref)) can be equivalently expressed as $m(X,\alpha _{0})=0$ $a.s.-X$, where $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})\in \mathcal{A}\equiv \Theta \times \mathcal{H}$, or as

equation*[equation* omitted — 74 chars of source]

Let $\Sigma _{0}(X)\equiv Var(\rho (Y,X;\alpha _{0})|X)$ be positive definite for almost all $X$. In this paper as well as in most applications $ \Sigma (X)$ is chosen to be either $I_{d_{\rho }}$ (identity) or $\Sigma _{0}(X)$ for almost all $X$. We call $Q^{0}(\alpha )\equiv E\left[ ||m(X,\alpha )||_{\Sigma _{0}^{-1}}^{2}\right] $ the population optimally weighted MD criterion function.

Let $\phi :\mathcal{A}\rightarrow \mathbb{R}^{d_{\phi }}$ be a functional with a finite $d_{\phi }\geq 1$. We are interested in inference on $\phi (\alpha _{0})$. Let

equation[equation omitted — 179 chars of source]

be a sample estimate of $Q(\alpha )$, where $\widehat{m}(X,\alpha )$ and $ \widehat{\Sigma }(X)$ are any consistent estimators of $m(X,\alpha )$ and $ \Sigma (X)$ respectively. When $\widehat{\Sigma }(X)=\widehat{\Sigma } _{0}(X) $ is a consistent estimator of the optimal weighting matrix $\Sigma _{0}(X)$, we call the corresponding $\widehat{Q}_{n}(\alpha )$ the\ sample optimally weighted MD criterion $\widehat{Q}_{n}^{0}(\alpha )$.

We estimate $\phi (\alpha _{0})$ by $\phi (\widehat{\alpha }_{n})$, where $ \widehat{\alpha }_{n}\equiv (\widehat{\theta }_{n}^{\prime },\widehat{h} _{n}) $ is an approximate penalized sieve minimum distance (PSMD) estimator of $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$, defined as

equation[equation omitted — 231 chars of source]

where $\lambda _{n}Pen(h)\geq 0$ is a penalty term such that $\lambda _{n}=o(1)$; and $\mathcal{A}_{k(n)}\equiv \Theta \times \mathcal{H}_{k(n)}$ is a finite dimensional sieve for $\mathcal{A}\equiv \Theta \times \mathcal{H }$, more precisely, $\mathcal{H}_{k(n)}$ is a finite dimensional linear sieve for $\mathcal{H}$:

equation[equation omitted — 168 chars of source]

where $\{q_{k}\}_{k=1}^{\infty }$ is a sequence of known basis functions of a Banach space $(\mathcal{H},\left\Vert \cdot \right\Vert _{\mathbf{H}})$ such as wavelets, splines, Fourier series, Hermite polynomial series, etc. And $k(n)\rightarrow \infty $ as $n\rightarrow \infty $.

For the purely nonparametric conditional moment models $E\left[ \rho (Y,X;h_{0})|X\right] =0$, CP_WP07 proposed more general approximate PSMD estimators of $h_{0}$ by allowing for possibly infinite dimensional sieves (i.e., $\dim (\mathcal{H}_{k(n)})=k(n)\leq \infty $). Nevertheless, both the theoretical properties and Monte Carlo simulations in CP_WP07 recommend the use of the PSMD procedures with slowly growing finite-dimensional linear sieves with a tiny penalty (i.e., $k(n)\rightarrow \infty ,\frac{k(n)}{n}\rightarrow 0$ as $n$ $\rightarrow \infty $ with a very small $\lambda _{n}=o(n^{-1})$, and hence the main smoothing parameter is the sieve dimension $k(n)$). This class of PSMD estimators include the original SMD estimators of NP_ECMA03 and AC_Emetrica03 as special cases, and has been used in recent empirical estimation of semiparametric structural models in microeconomics and asset pricing with endogeneity. See, e.g., BCK_Emetrica07, Horowitz_ECMA11, CP_WP07a, BHN_WP11, Souza2012, Pinkse2002, MerloPaula, BMT_WP12, ChenLudvigson_JAE09, ChenLudvigson2013, Sentana2013 and others.

In this paper we shall develop inferential theory for $\phi (\alpha _{0})$ based on the PSMD procedures with slowly growing finite-dimensional sieves $ \mathcal{A}_{k(n)}=\Theta \times \mathcal{H}_{k(n)}$. We first establish the large sample theories under a high level \textquotedblleft local quadratic approximation\textquotedblright\ (LQA) condition, which allows for any consistent nonparametric estimator $\widehat{m}(x,\alpha )$ that is linear in $\rho (Z,\alpha )$:

equation[equation omitted — 113 chars of source]

where $A_{n}(X_{i},x)$ is a known measurable function of $ \{X_{j}\}_{j=1}^{n} $ for all $x$, whose expression varies according to different nonparametric procedures such as kernel, local linear regression, series and nearest neighbors. In Appendix (ref) we provide lower level sufficient conditions for this LQA assumption when $\widehat{m} (x,\alpha )$ is the series least squares (LS) estimator ((ref)):

equation[equation omitted — 158 chars of source]

which is a linear nonparametric estimator ((ref)) with $ A_{n}(X_{i},x)=p^{J_{n}}(X_{i})^{\prime }(P^{\prime }P)^{-}p^{J_{n}}(x)$, where $\{p_{j}\}_{j=1}^{\infty }$ is a sequence of known basis functions that can approximate any square integrable functions of $X$ well, $ p^{J_{n}}(X)=(p_{1}(X),...,p_{J_{n}}(X))^{\prime }$, $ P=(p^{J_{n}}(X_{1}),...,p^{J_{n}}(X_{n}))^{\prime }$, and $(P^{\prime }P)^{-} $ is the generalized inverse of the matrix $P^{\prime }P$. Following BCK_Emetrica07 and CP_WP07a, we let $p^{J_{n}}(X)$ be a tensor-product linear sieve basis, and $J_{n}$ be the dimension of $ p^{J_{n}}(X)$ such that $J_{n}\geq d_{\theta }+k(n)\rightarrow \infty $ and $ \frac{J_{n}}{n}\rightarrow 0$ as $n$ $\rightarrow \infty $.

Preview of the Main Results for Inference

For simplicity we let $\phi :\mathbb{R}^{d_{\theta }}\times \mathcal{H} \rightarrow \mathbb{R}$ be a real-valued functional. Let $\widehat{\phi } _{n}\equiv \phi (\widehat{\alpha }_{n})$ be the plug-in PSMD estimator of $\phi (\alpha _{0})$ for $\alpha _{0}=(\theta _{0}^{\prime },h_{0})\in int(\Theta )\times \mathcal{H}$.

Sieve t (or Wald) statistic. Regardless of whether $\phi (\alpha _{0})$ is $\sqrt{n}$ estimable or not, Theorem (ref) shows that $\frac{\sqrt{n}\left\{ \phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})\right\} }{||v_{n}^{\ast }||_{sd}}$ is asymptotically standard normal, and the sieve variance $||v_{n}^{\ast }||_{sd}^{2}$ has a closed form expression resembling the \textquotedblleft delta-method\textquotedblright\ variance for a parametric MD problem:

equation[equation omitted — 258 chars of source]

where $\overline{q}^{k(n)}(\cdot )\equiv \left( \mathbf{1}_{d_{\theta }}^{\prime },q^{k(n)}(\cdot )^{\prime }\right) ^{\prime }$ is a $(d_{\theta }+k(n))\times 1$ vector with $\mathbf{1}_{d_{\theta }}$ a $d_{\theta }\times 1$ vector of $1$'s,

equation[equation omitted — 387 chars of source]

and $\gamma \equiv (\theta ^{\prime },\beta ^{\prime })^{\prime }$ are $ (d_{\theta }+k(n))\times 1$ vectors, $\frac{d\phi (\alpha _{0})}{dh} [q^{k(n)}(\cdot )^{\prime }]\equiv \frac{\partial \phi (\theta _{0},h_{0}+\beta ^{\prime }q^{k(n)}(\cdot ))}{\partial \beta } \mid_{\beta =0}$, and

equation[equation omitted — 255 chars of source]

{{

equation[equation omitted — 326 chars of source]

}} where $\frac{dm(X,\alpha _{0})}{d\alpha }[\overline{q}^{k(n)}(\cdot )^{\prime }]\equiv \frac{\partial E[\rho (Z,\theta _{0}+\theta ,h_{0}+\beta ^{\prime }q^{k(n)}(\cdot ))|X]}{\partial \gamma } \mid_{\gamma =0}$ is a $d_{\rho }\times (d_{\theta }+k(n))$ matrix. The closed form expression of $||v_{n}^{\ast }||_{sd}^{2}$ immediately leads to simple consistent plug-in sieve variance estimators; one of which is

equation[equation omitted — 339 chars of source]

where $\frac{d\phi (\widehat{\alpha }_{n})}{d\alpha }[\overline{q} ^{k(n)}(\cdot )]\equiv \frac{\partial \phi (\widehat{\theta }_{n}+\theta , \widehat{h}_{n}+\beta ^{\prime }q^{k(n)}(\cdot ))}{\partial \gamma ^{\prime } } \mid_{\gamma =0}$ and

equation[equation omitted — 355 chars of source]
equation[equation omitted — 352 chars of source]

where $\widehat{M}(X_{i}) \equiv \widehat{\Sigma }(X_{i})^{-1}\rho (Z_{i},\widehat{\alpha }_{n})\rho (Z_{i},\widehat{\alpha }_{n})^{\prime }\widehat{\Sigma } (X_{i})^{-1}$. Theorem (ref) then presents the asymptotic normality of the sieve (Student's) t statistic:\footnote{ See Theorems (ref) and (ref) for properties of bootstrap sieve t statistics.}

equation*[equation* omitted — 156 chars of source]

Sieve QLR statistic. In addition to the sieve t (or sieve Wald) statistic, we could also use sieve quasi likelihood ratio for constructing confidence set of $\phi (\alpha _{0})$ and for hypothesis testing of $ H_{0}:\phi (\alpha _{0})=\phi _{0}$ against $H_{1}:\phi (\alpha _{0})\neq \phi _{0}$. Denote

equation[equation omitted — 206 chars of source]

as the sieve quasi likelihood ratio (SQLR) statistic. It becomes an optimally weighted SQLR statistic, $\widehat{QLR}_{n}^{0}(\phi _{0}) $, when $\widehat{Q}_{n}(\alpha )$ is the optimally weighted MD criterion $\widehat{Q}_{n}^{0}(\alpha )$. Regardless of whether $\phi (\alpha _{0})$ is $\sqrt{n}$ estimable or not, Theorems (ref)(2) and (ref) show that $\widehat{QLR}_{n}^{0}(\phi _{0})$ is asymptotically chi-square distributed under the null $H_{0}$, and diverges to infinity under the fixed alternatives $H_{1}$. Theorem (ref) in Appendix (ref) states that $\widehat{QLR} _{n}^{0}(\phi _{0})$ is asymptotically noncentral chi-square distributed under local alternatives. One could compute $100(1-\tau )\%$ confidence set for $\phi (\alpha _{0})$ as

equation*[equation* omitted — 120 chars of source]

where $c_{\chi _{1}^{2}}(1-\tau )$ is the $(1-\tau )$-th quantile of the $ \chi _{1}^{2}$ distribution.

Bootstrap sieve QLR statistic. Regardless of whether $\phi (\alpha _{0})$ is $\sqrt{n}$ estimable or not, Theorems (ref)(1) and (ref) establish that the possibly non-optimally weighted SQLR statistic $\widehat{QLR}_{n}(\phi _{0})$ is stochastically bounded under the null $H_{0}$ and diverges to infinity under the fixed alternatives $H_{1}$. We then consider a bootstrap version of the SQLR statistic. Let $\widehat{QLR }_{n}^{B}$ denote a bootstrap SQLR statistic:

equation[equation omitted — 264 chars of source]

where $\widehat{\phi }_{n}\equiv \phi (\widehat{\alpha }_{n})$, and $ \widehat{Q}_{n}^{B}(\alpha )$ is a bootstrap version of $\widehat{Q} _{n}(\alpha )$:

equation[equation omitted — 193 chars of source]

where $\widehat{m}^{B}(x,\alpha )$ is a bootstrap version of $\widehat{m} (x,\alpha )$, which is computed in the same way as that of $\widehat{m} (x,\alpha )$ except that we use $\omega _{i,n}\rho (Z_{i},\alpha )$ instead of $\rho (Z_{i},\alpha )$. Here $\{\omega _{i,n}\geq 0\}_{i=1}^{n}$ is a sequence of bootstrap weights that has mean 1 and is independent of the original data $\{Z_{i}\}_{i=1}^{n}$. Typical weights include an i.i.d. weight $\{\omega _{i}\geq 0\}_{i=1}^{n}$ with $E[\omega _{i}]=1$, $E[|\omega _{i}-1|^{2}]=1$ and $E[|\omega _{i}-1|^{2+\epsilon }]<\infty $ for some $ \epsilon >0$, or a multinomial weight (i.e., $(\omega _{1,n},...,\omega _{n,n})\sim Multinomial(n;n^{-1},...,n^{-1})$). For example, if $\widehat{m} (x,\alpha )$ is a series LS estimator ((ref)) of $m(x,\alpha )$, then $ \widehat{m}^{B}(x,\alpha )$ is a bootstrap series LS estimator of $ m(x,\alpha )$, defined as:

equation[equation omitted — 184 chars of source]

We sometimes call our bootstrap procedure \textquotedblleft generalized residual bootstrap\textquotedblright\ since it is based on randomly perturbing the generalized residual function $\rho (Z,\alpha )$; see Section (ref) for details. Theorems (ref) and (ref) establish that under the null $H_{0}$, the fixed alternatives $H_{1}$ or the local alternatives,\footnote{ See Section (ref) for definition of the local alternatives and the behaviors of $\widehat{QLR}_{n}(\phi _{0})$ and $\widehat{QLR}_{n}^{B}( \widehat{\phi }_{n})$ under the local alternatives.} the conditional distribution of $\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})$ (given the data) always converges to the asymptotic null distribution of $\widehat{QLR} _{n}(\phi _{0})$. Let $\widehat{c}_{n}(a)$ be the $a-th$ quantile of the distribution of $\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})$ (conditional on the data $\{Z_{i}\}_{i=1}^{n}$). Then for any $\tau \in (0,1)$, we have $ \lim_{n\rightarrow \infty }\Pr \{\widehat{QLR}_{n}(\phi _{0})>\widehat{c} _{n}(1-\tau )\}=\tau $ under the null $H_{0}$, $\lim_{n\rightarrow \infty }\Pr \{\widehat{QLR}_{n}(\phi _{0})>\widehat{c}_{n}(1-\tau )\}=1$ under the fixed alternatives $H_{1}$, and $\lim_{n\rightarrow \infty }\Pr \{\widehat{ QLR}_{n}(\phi _{0})>\widehat{c}_{n}(1-\tau )\}>\tau $ under the local alternatives. We could also construct a $100(1-\tau )\%$ confidence set using the bootstrap critical values:

equation[equation omitted — 131 chars of source]

The bootstrap consistency holds for possibly non-optimally weighted SQLR statistic and possibly irregular functionals, without the need to compute standard errors.

Which method to use? When sieve Wald and SQLR tests are computed using the same weighting matrix $\widehat{\Sigma }$, there is no local power difference in terms of first order asymptotic theories; see Appendix (ref). As will be demonstrated in simulation Section (ref), while SQLR and bootstrap SQLR tests are useful for models ((ref)) with (pointwise) non-smooth $\rho (Z;\alpha )$, sieve Wald (or t) statistic is computationally attractive for models with smooth $ \rho (Z;\alpha )$. Empirical researchers could apply either inference method depending on whether the residual function $\rho (Z;\alpha )$ in their specific application is pointwise differentiable with respect to $\alpha $ or not.

Applications to NPIV and NPQIV models

An illustration via the NPIV model. BCK_Emetrica07 and CR_WP07 established the convergence rate of the identity weighted (i.e., $ \widehat{\Sigma }=\Sigma =1$) PSMD estimator $\widehat{h}_{n}\in \mathcal{H} _{k(n)}$ of the NPIV model:

equation[equation omitted — 73 chars of source]

By Theorem (ref)

align*[align* omitted — 109 chars of source]

with $||v_{n}^{\ast }||_{sd}^{2}=\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]^{\prime }D_{n}^{-}\mho _{n}D_{n}^{-}\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]$,

align[align omitted — 194 chars of source]

and $\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]\equiv \frac{\partial \phi (h_{0}+\beta ^{\prime }q^{k(n)}(\cdot ))}{\partial \beta ^{\prime }} \mid_{\beta =0}$. For example, for a functional $\phi (h)=h(\overline{y}_{2})$, or $=\int w(y)\nabla h(y)dy$ or $=\int w(y)\left\vert h(y)\right\vert ^{2}dy$ , we have $\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]=q^{k(n)}(\overline{y} _{2})$, or $=\int w(y)\nabla q^{k(n)}(y)dy$ or $=2\int h_{0}(y)w(y)q^{k(n)}(y)dy$.

If $0<\inf_{x}\Sigma _{0}(x)\leq \sup_{x}\Sigma _{0}(x)<\infty $ then

align*[align* omitted — 150 chars of source]

Without endogeneity (say $Y_{2}=X$) the model becomes the nonparametric LS regression

equation*[equation* omitted — 64 chars of source]

and the variance satisfies $||v_{n}^{\ast }||_{sd,ex}^{2}\asymp \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]^{\prime }D_{n,ex}^{-}\frac{d\phi (h_{0})}{dh} [q^{k(n)}(\cdot )]$, $D_{n,ex}=E[\{q^{k(n)}(Y_{2})\}\{q^{k(n)}(Y_{2})\}^{ \prime }]$. Since the conditional expectation $E[q^{k(n)}(Y_{2})|X]$ is a contraction, $D_{n}\leq D_{n,ex}$ and $||v_{n}^{\ast }||_{sd}^{2}\geq const.||v_{n}^{\ast }||_{sd,ex}^{2}$. Under mild conditions (see, e.g., NP_ECMA03, BCK_Emetrica07, DFR_wp10, Horowitz_ECMA11 ), the minimal eigenvalue of $D_{n}$, $\lambda _{\min }(D_{n})$, goes to zero while $\lambda _{\min }(D_{n,ex})$ stays strictly positive as $ k(n)\rightarrow \infty $. In fact, $D_{n,ex}=I_{k(n)}$ and $\lambda _{\min }(D_{n,ex})=1$ if $\{q_{j}\}_{j=1}^{\infty }$ is an orthonormal basis of $ L^{2}(f_{Y_{2}})$, while $\lambda _{\min }(D_{n})\asymp \exp (-k(n))$ if the conditional density of $Y_{2}$ given $X$ is normal. Therefore, while $ \lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd,ex}^{2}=\infty $ always implies $\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd}^{2}=\infty $, it is possible that $\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd,ex}^{2}<\infty $ but $\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd}^{2}=\infty $. For example, the point evaluation functional $\phi (h)=h(\overline{y}_{2})$ is known to be irregular for the nonparametric LS regression and hence for the NPIV ((ref)) as well. Under mild conditions on the weight $w()$ and the smoothness of $h_{0}$, the weighted derivative functional ($\phi (h)=\int w(y)\nabla h(y)dy$) and the quadratic functional ($\phi (h)=\int w(y)\left\vert h(y)\right\vert ^{2}dy$) of the nonparametric LS regression are typically root-$n$ estimable, but they could be irregular for the NPIV ((ref)). See Section (ref) for details.

Regardless of whether $\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd}^{2}$ is finite or infinite, Theorem (ref) shows that the sieve variance $||v_{n}^{\ast }||_{sd}^{2}$ can be consistently estimated by a plug-in sieve variance estimator $||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}$, and that $\sqrt{n}\frac{\phi (\widehat{h}_{n})-\phi (h_{0})}{||\widehat{v} _{n}^{\ast }||_{n,sd}}\Rightarrow N(0,1)$.

When the conditional mean function $m(x,h)$ is estimated by the series LS estimator ((ref)) as in NP_ECMA03, AC_Emetrica03 and BCK_Emetrica07, with $\widehat{U}_{i}=Y_{1i}-\widehat{h}_{n}(Y_{2i})$, the sieve variance estimator $||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}$ given in ( (ref)) has a more explicit expression:

equation*[equation* omitted — 295 chars of source]

$\frac{d\phi (\widehat{h}_{n})}{dh}[q^{k(n)}(\cdot )]\equiv \frac{\partial \phi (\widehat{h}_{n}+\beta ^{\prime }q^{k(n)}(\cdot ))}{\partial \beta ^{\prime }} \mid_{\beta =0}$ and

equation*[equation* omitted — 191 chars of source]
equation[equation omitted — 235 chars of source]

Interestingly, this sieve variance estimator becomes the one computed via the two stage least squares (2SLS) as if the NPIV model ((ref)) were a parametric IV regression:\footnote{This confirms a conjecture of Newey2013 for the NPIV model ((ref) ).} $Y_{1}=q^{k(n)}(Y_{2j})^{\prime }\beta _{0n}+U,$ $ E[q^{k(n)}(Y_{2})U]\neq 0,$ $E[p^{J_{n}}(X)U]=0$ and $ E[p^{J_{n}}(X)q^{k(n)}(Y_{2})^{\prime }]$ has a column rank $k(n)\leq J_{n}$ . See Subsection (ref) for simulation studies of finite sample performances of this sieve variance estimator $\widehat{V}_{1}$ for both a linear and a nonlinear functional $\phi (h)$.

An illustration via the NPQIV model. As an application of their general theory, CP_WP07 presented the consistency and the rate of convergence of the PSMD estimator $\widehat{h}_{n}\in \mathcal{H}_{k(n)}$ of the NPQIV model:

equation[equation omitted — 89 chars of source]

In this example we have $\Sigma _{0}(X)=\gamma (1-\gamma )$. So we could use $\widehat{\Sigma }(X)=\gamma (1-\gamma )$ and $\widehat{Q}_{n}(\alpha )$ given in ((ref)) becomes the optimally weighted\ MD criterion.

By Theorem (ref)

align*[align* omitted — 108 chars of source]

with $ ||v_{n}^{\ast }||_{sd}^{2}=\left( \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]\right) ^{\prime }D_{n}^{-}\left( \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]\right) $ and

equation[equation omitted — 167 chars of source]

Without endogeneity (say $Y_{2}=X$), the model becomes the nonparametric quantile regression

equation*[equation* omitted — 79 chars of source]

and the sieve variance becomes $||v_{n}^{\ast }||_{sd,ex}^{2}=\left( \frac{ d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]\right) ^{\prime }D_{n,ex}^{-}\left( \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]\right) $ with $D_{n,ex}=\frac{1}{ \gamma (1-\gamma )}E\left[ \{f_{U|Y_{2}}(0)\}^{2}\{q^{k(n)}(Y_{2})\} \{q^{k(n)}(Y_{2})\}^{\prime }\right] $. Again $D_{n}\leq D_{n,ex}$ and $ ||v_{n}^{\ast }||_{sd}^{2}\geq ||v_{n}^{\ast }||_{sd,ex}^{2}$. Under mild conditions (see, e.g., CP_WP07, CCLN_WP10), $\lambda _{\min }(D_{n})\rightarrow 0$ while $\lambda _{\min }(D_{n,ex})$ stays strictly positive as $k(n)\rightarrow \infty $. All of the above discussions for a functional $\phi (h)$ of the NPIV ((ref)) now apply to the functional of the NPQIV ((ref)). In particular, a functional $\phi (h)$ could be root-$n$ estimable for the nonparametric quantile regression ($ \lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd,ex}^{2}<\infty $) but irregular for the NPQIV ((ref)) ($\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd}^{2}=\infty $). See Section (ref) for details.

By Theorems (ref)(2) and (ref), the optimally weighted SQLR statistic $\widehat{QLR}_{n}^{0}(\phi _{0})\Rightarrow \chi _{1}^{2}$ under the null of $\phi (h_{0})=\phi _{0}$, and diverges to infinity under the alternative of $\phi (h_{0})\neq \phi _{0}$. We can compute confidence set for a functional $\phi (h)$, such as an evaluation or a weighted derivative functional, as $\left\{ r\in \mathbb{R}\colon \text{ }\widehat{QLR }_{n}^{0}(r)\leq c_{\chi _{1}^{2}}(\tau )\right\} $. See Subsection (ref) for an empirical illustration of this result to the NPQIV Engel curve regression using the British Family Survey data set that was first used in BCK_Emetrica07. Instead of using the asymptotic critical values, we could also construct a confidence set using the bootstrap critical values as in ((ref)).

Basic Regularity Conditions

Before we establish asymptotic properties of sieve t (Wald) and SQLR statistics, we need to present three sets of basic regularity conditions. The first set of assumptions allows us to establish the convergence rates of the PSMD estimator $\widehat{\alpha }_{n}$ to the true parameter value $ \alpha _{0}$ in both weak and strong metrics, which in turn allows us to concentrate on some shrinking neighborhood of $\alpha _{0}$ in the semi/nonparametric model ((ref)). The second and third regularity conditions are respectively about the local curvatures of the functional $ \phi ()$ and of the criterion function under these two metrics. The weak metric $||\cdot ||$ is closely related to the variance of the linear approximation to $\phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})$, while the strong metric $||\cdot ||_{s}$ is used to control the nonlinearity (in $ \alpha $) of the functional $\phi ()$ and of the conditional mean function $ m(x,\alpha )$. This section is mostly technical and applied researchers could skip this and directly go to the subsequent sections on the asymptotic properties of sieve Wald and SQLR statistics.

A brief discussion on the convergence rate of the PSMD estimator

For the purely nonparametric conditional moment model $E\left[ \rho (Y,X;h_{0}(\cdot ))|X\right] =0$, CP_WP07 established the consistency and the convergence rates of their various PSMD estimators of $h_{0}$. Their results can be trivially extended to establish the corresponding properties of our PSMD estimator $\widehat{\alpha }_{n}\equiv (\widehat{\theta } _{n}^{\prime },\widehat{h}_{n})$ defined in ((ref)). For the sake of easy reference and to introduce basic assumptions and notation, we present some sufficient conditions for consistency and the convergence rate here. These conditions are also needed to establish the consistency and the convergence rate of bootstrap PSMD estimators (see Lemma (ref) ). We first impose three conditions on identification, sieve spaces, penalty functions and sample criterion function. We equip the parameter space $ \mathcal{A}\equiv \Theta \times \mathcal{H}$ with a (strong) norm $ \left\Vert \alpha \right\Vert _{s}\equiv \left\Vert \theta \right\Vert _{e}+\left\Vert h\right\Vert _{\mathbf{H}}$.

assumption[Identification, sieves, criterion] (i) $E[\rho (Y,X;\alpha )|X]=0$ if and only if $\alpha \in (\mathcal{A},\left\Vert \cdot \right\Vert _{s})$ with $\left\Vert \alpha -\alpha _{0}\right\Vert _{s}=0$; (ii) For all $k\geq 1$, $\mathcal{A} _{k}\equiv \Theta \times \mathcal{H}_{k}$, $\Theta $ is a compact subset in $ \mathbb{R}^{d_{\theta }}$ with a non-empty interior, $\{\mathcal{H} _{k}:k\geq 1\}$ is a non-decreasing sequence of non-empty closed linear subsets of a Banach space $\left( \mathcal{H},\left\Vert \cdot \right\Vert _{ \mathbf{H}}\right) $ such that $\mathcal{H}=cl\left( \cup _{k}\mathcal{H} _{k}\right) $, and there is $\Pi _{n}h_{0}\in \mathcal{H}_{k(n)}$ with $ ||\Pi _{n}h_{0}-h_{0}||_{\mathbf{H}}=o(1)$; (iii) $Q:(\mathcal{A},\left\Vert \cdot \right\Vert _{s})\rightarrow \lbrack 0,\infty )$ is lower semicontinuous;\footnote{ A function $Q$ is lower semicontinuous at a point $\alpha _{o}\in \mathcal{A} $ iff $\lim_{\left\Vert \alpha -\alpha _{o}\right\Vert _{s}\rightarrow 0}Q(\alpha )\geq Q(\alpha _{o})$; is lower semicontinuous if it is lower semicontinuous at any point in $\mathcal{A}$.} (iv) $\Sigma (x)$ and $\Sigma _{0}(x)$ are positive definite, and their smallest and largest eigenvalues are finite and positive uniformly in $x\in \mathcal{X}$.
assumption[Penalty] (i) $\lambda _{n}>0$, $Q(\Pi _{n}\alpha _{0})+o(n^{-1})=O(\lambda _{n})=o(1)$; (ii) $|Pen(\Pi _{n}h_{0})-Pen(h_{0})|=O(1)$ with $Pen(h_{0})<\infty $; (iii) $Pen:(\mathcal{ H},\left\Vert \cdot \right\Vert _{\mathbf{H}})\rightarrow \lbrack 0,\infty )$ is lower semicompact.\footnote{ A function $Pen$ is lower semicompact iff for all $M$, $\{h\in \mathcal{H} \colon Pen(h)\leq M\}$ is a compact subset in $(\mathcal{H},\left\Vert \cdot \right\Vert _{\mathbf{H}})$.}

Let $\Pi _{n}\alpha \equiv (\theta ^{\prime },\Pi _{n}h)\in \mathcal{A} _{k(n)}\equiv \Theta \times \mathcal{H}_{k(n)}$. Let $\mathcal{A} _{k(n)}^{M_{0}}\equiv \Theta \times \mathcal{H}_{k(n)}^{M_{0}}\equiv \{\alpha =(\theta ^{\prime },h)\in \mathcal{A}_{k(n)}:\lambda _{n}Pen(h)\leq \lambda _{n}M_{0}\}$ for a large but finite $M_{0}$ such that $\Pi _{n}\alpha _{0}\in \mathcal{A}_{k(n)}^{M_{0}}$ and that $\widehat{\alpha } _{n}\in \mathcal{A}_{k(n)}^{M_{0}}$ with probability arbitrarily close to one for all large $n$. Let $\{\bar{\delta}_{m,n}^{2}\}_{n=1}^{\infty }$ be a sequence of positive real values that decrease to zero as $n\rightarrow \infty $.

assumption[Sample Criterion] (i) $\widehat{Q}_{n}(\Pi _{n}\alpha _{0})\leq c_{0}Q(\Pi _{n}\alpha _{0})+o_{P_{Z^{\infty }}}(n^{-1})$ for a finite constant $c_{0}>0$ ; (ii) $\widehat{Q}_{n}(\alpha )\geq cQ(\alpha )-O_{P_{Z^{\infty }}}(\bar{ \delta}_{m,n}^{2})$ uniformly over $\mathcal{A}_{k(n)}^{M_{0}}$ for some $ \bar{\delta}_{m,n}^{2}=o(1)$ and a finite constant $c>0$.

The following result is a minor modification of Theorem 3.2 of CP_WP07 .

lemmaLet $\widehat{\alpha }_{n}$ be the PSMD estimator defined in ((ref)), and Assumptions (ref), (ref) and (ref) hold. Then: $||\widehat{\alpha }_{n}-\alpha _{0}||_{s}=o_{P_{Z^{\infty }}}(1)$ and $Pen(\widehat{h}_{n})=O_{P_{Z^{\infty }}}(1)$.

Given the consistency result, the PSMD estimator belongs to any $||\cdot ||_{s}-$neighborhood around $\alpha _{0}$ wpa1. We can restrict our attention to a convex, $||\cdot ||_{s}-$neighborhood around $\alpha _{0}$, denoted as $\mathcal{A}_{os}$ such that

equation*[equation* omitted — 146 chars of source]

for a positive finite constant $M_{0}$ (the existence of a convex $\mathcal{A }_{os}$ is implied by the convexity of $\mathcal{A}$ and quasi-convexity of $ Pen(\cdot )$). For any $\alpha \in \mathcal{A}_{os}$ we define a pathwise derivative as

eqnarray*[eqnarray* omitted — 333 chars of source]

Following AC_Emetrica03 and CP_WP07a, we introduce two pseudo-metrics $||\cdot ||$ and $||\cdot ||_{0}$ on $\mathcal{A}_{os}$ as: for any $\alpha _{1},$ $\alpha _{2}\in \mathcal{A}_{os}$,

equation[equation omitted — 263 chars of source]
equation[equation omitted — 270 chars of source]

It is clear that, under Assumption (ref)(iv), these two pseudo-metrics are equivalent, i.e., $||\cdot ||\asymp ||\cdot ||_{0}$ on $ \mathcal{A}_{os}$. This is why Assumption (ref)(iv) is imposed throughout the paper.

Let $\mathcal{A}_{osn}=\mathcal{A}_{os}\cap \mathcal{A}_{k(n)}$. Let $ \{\delta _{n}\}_{n=1}^{\infty }$ be a sequence of positive real values such that $\delta _{n}=o(1)$ and $\delta _{n}\leq \bar{\delta}_{m,n}$.

assumption(i) There exists a convex $||\cdot ||_{s}-$ neighborhood of $\alpha _{0}$, $\mathcal{A}_{os}$, such that $m(\cdot ,\alpha )$ is continuously pathwise differentiable with respect to $\alpha \in \mathcal{A}_{os}$, and there is a finite constant $C>0$ such that $ ||\alpha -\alpha _{0}||\leq C||\alpha -\alpha _{0}||_{s}$ for all $\alpha \in \mathcal{A}_{os}$; (ii) $Q(\alpha )\asymp ||\alpha -\alpha _{0}||^{2}$ for all $\alpha \in \mathcal{A}_{os}$; (iii) $\widehat{Q}_{n}(\alpha )\geq cQ(\alpha )-O_{P_{Z^{\infty }}}(\delta _{n}^{2})$ uniformly over $\mathcal{A} _{osn}$, and $\max \{\delta _{n}^{2},Q(\Pi _{n}\alpha _{0}),\lambda _{n},o(n^{-1})\}=\delta _{n}^{2}$; (iv) $\lambda _{n}\times \sup_{\alpha ,\alpha ^{\prime }\in \mathcal{A}_{os}}\left\vert Pen(h)-Pen(h^{\prime })\right\vert =o(n^{-1})$ or $\lambda _{n}=o(n^{-1})$.

Assumption (ref)(ii) is about the local curvature of the population criterion $Q(\alpha )$ at $\alpha _{0}$. It can be weakened to Assumption 4.1(ii) in CP_WP07. When $\widehat{Q} _{n}(\alpha )$ is computed using the series LS estimator ((ref)), Lemma C.2 of CP_WP07 shows that $\widehat{Q}_{n}(\alpha )\asymp Q(\alpha )-O_{P_{Z^{\infty }}}(\delta _{n}^{2})$ uniformly over $\mathcal{A}_{osn}$ and hence Assumption (ref)(iii) is satisfied.

Recall the definition of the sieve measure of local ill-posedness

equation[equation omitted — 196 chars of source]

The problem of estimating $\alpha _{0}$ under $||\cdot ||_{s}$ is locally ill-posed in rate if and only if $\limsup_{n\rightarrow \infty }\tau _{n}=\infty $. We say the problem is mildly ill-posed if $\tau _{n}=O([k(n)]^{a})$, and severely ill-posed if $\tau _{n}=O(\exp \{\frac{a}{2}k(n)\})$ for some finite $a>0$. The following general rate result is a minor modification of Theorem 4.1 and Remark 4.1(i) of CP_WP07, and hence we omit its proof.

lemmaLet $\widehat{\alpha }_{n}$ be the PSMD estimator defined in ((ref)), and Assumptions (ref), (ref)(ii)(iii), (ref) and (ref)(i)(ii)(iii) hold. Then: \begin{equation*} ||\widehat{\alpha }_{n}-\alpha _{0}||=O_{P_{Z^{\infty }}}\left( \delta _{n}\right) \quad and\quad ||\widehat{\alpha }_{n}-\alpha _{0}||_{s}=O_{P_{Z^{\infty }}}\left( ||\alpha _{0}-\Pi _{n}\alpha _{0}||_{s}+\tau _{n}\delta _{n}\right) . \end{equation*}

The above convergence rate result is applicable to any nonparametric estimator $\widehat{m}(X,\alpha )$ of $m(X,\alpha )$ as soon as one could compute $\delta _{n}^{2}$, the rate at which $\widehat{Q}_{n}(\alpha )$ goes to $Q(\alpha )$. See CP_WP07 and CP_WP07a for low level sufficient conditions in terms of the series LS estimator ((ref)) of $ m(X,\alpha )$.

Let $\left\{ \delta _{s,n}:n\geq 1\right\} $ be a sequence of real positive numbers such that $\delta _{s,n}=||h_{0}-\Pi _{n}h_{0}||_{s}+\tau _{n}\delta _{n}=o(1)$. Lemma (ref) implies that $\widehat{\alpha } _{n}\in \mathcal{N}_{osn}\subseteq \mathcal{N}_{os}$ wpa1-$P_{Z^{\infty }}$, where

eqnarray*[eqnarray* omitted — 422 chars of source]

We can regard $\mathcal{N}_{os}$ as the effective parameter space and $ \mathcal{N}_{osn}$ as its sieve space in the rest of the paper. Assumption (ref)(iv) is not needed for establishing a convergence rate in Lemma (ref). but, it will be imposed in the rest of the paper so that we can ignore penalty effect in the first order local asymptotic analysis.

(Sieve) Riesz representation and (sieve) variance

We first introduce a representation of the functional of interest $\phi ()$ at $\alpha _{0}$ that is crucial for all the subsequent local asymptotic theories. Let $\phi :\mathbb{R}^{d_{\theta }}\times \mathcal{H}\rightarrow \mathbb{R}$ be continuous in $||\cdot ||_{s}$. We assume that $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]:\left( \mathbb{R}^{d_{\theta }}\times \mathcal{H},||\cdot ||_{s}\right) \rightarrow \mathbb{R}$ is a $||\cdot ||_{s}-$bounded linear functional (i.e., $\left\vert \frac{d\phi (\alpha _{0})}{d\alpha }[v]\right\vert \leq c||v||_{s}$ uniformly over $v\in \mathbb{ R}^{d_{\theta }}\times \mathcal{H}$ for a finite positive constant $c$), which could be computed as a pathwise\ (directional) derivative of the functional $\phi \left( \cdot \right) $ at $\alpha _{0}$ in the direction of $v=\alpha -\alpha _{0}\in \mathbb{R}^{d_{\theta }}\times \mathcal{H}:$

equation*[equation* omitted — 157 chars of source]

Let $\mathbf{V}$ be a linear span of $\mathcal{A}_{os}-\{\alpha _{0}\}$, which is endowed with both $||\cdot ||_{s}$ and $||\cdot ||$ (in equation ( (ref))) norms, and $||v||\leq C||v||_{s}$ for all $v\in \mathbf{V}$ (under Assumption (ref)(i)). Let $\overline{\mathbf{V}}\equiv clsp(\mathcal{A}_{os}-\{\alpha _{0}\})$, where $clsp(\cdot )$ is the closure of the linear span under $||\cdot ||$. For any $v_{1},v_{2}\in \overline{ \mathbf{V}}$, we define an inner product induced by the metric $||\cdot ||$:

equation*[equation* omitted — 211 chars of source]

and for any $v\in \overline{\mathbf{V}}$ we call $v=0$ if and only if $ ||v||=0$ (i.e., functions in $\overline{\mathbf{V}}$ are defined in an equivalent class sense according to the metric $||\cdot ||$). It is clear that $(\overline{\mathbf{V}},||\cdot ||)$ is an infinite dimensional Hilbert space (under Assumptions (ref)(i)(iii)(iv) and (ref) (i)(ii)).

If the linear functional $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is bounded on $(\mathbf{V},||\cdot ||)$, i.e.

equation*[equation* omitted — 165 chars of source]

then there is a unique extension of\ $\frac{d\phi (\alpha _{0})}{d\alpha } [\cdot ]$ from $(\mathbf{V},||\cdot ||)$ to $(\overline{\mathbf{V}},||\cdot ||)$, and a unique Riesz representer $v^{\ast }\in \overline{\mathbf{V}}$ of $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ on $(\overline{\mathbf{V}} ,||\cdot ||)$ such that\footnote{ See, e.g., page 206-207 and theorem 3.10.1 in Debnath-Hilbert.}

align[align omitted — 532 chars of source]

If $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is unbounded on $( \mathbf{V},||\cdot ||)$, i.e.

equation*[equation* omitted — 165 chars of source]

then there is no unique extension of the mapping\ $\frac{d\phi (\alpha _{0}) }{d\alpha }[\cdot ]$ from $(\mathbf{V},||\cdot ||)$ to $(\overline{\mathbf{V} },||\cdot ||)$, and nor existing any Riesz representer of\ $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ on $(\overline{\mathbf{V}},||\cdot ||)$.

Since $||v||\leq C||v||_{s}$ for all $v\in \mathbf{V}$, it is clear that a $ ||\cdot ||_{s}-$bounded linear functional $\frac{d\phi (\alpha _{0})}{ d\alpha }[\cdot ]$ could be either bounded or unbounded on $(\mathbf{V} ,||\cdot ||)$.

Sieve Riesz representation. Let $\alpha _{0,n}\in \mathbb{R} ^{d_{\theta }}\times \mathcal{H}_{k(n)}$ be such that

equation[equation omitted — 157 chars of source]

Let $\overline{\mathbf{V}}_{k(n)}\equiv clsp\left( \mathcal{A} _{osn}-\{\alpha _{0,n}\}\right) $, where $clsp\left( .\right) $ denotes the closed linear span under $\left\Vert \cdot \right\Vert $. Then $\overline{ \mathbf{V}}_{k(n)}$ is a finite dimensional Hilbert space under $\left\Vert \cdot \right\Vert $. Moreover, $\overline{\mathbf{V}}_{k(n)}$ is dense in $ \overline{\mathbf{V}}$ under $\left\Vert \cdot \right\Vert $. To simplify the presentation, we assume that $\dim (\overline{\mathbf{V}}_{k(n)})=\dim ( \mathcal{A}_{k(n)})\asymp k(n)$, all of which grow to infinity with $n$. By definition we have $\left\langle v_{n},\alpha _{0,n}-\alpha _{0}\right\rangle =0$ for all $v_{n}\in \overline{\mathbf{V}}_{k(n)}$.

Note that $\overline{\mathbf{V}}_{k(n)}$ is a finite dimensional Hilbert space. As any linear functional on a finite dimensional Hilbert space is bounded, we can invoke the Riesz representation theorem to deduce that there is a $v_{n}^{\ast }\in \overline{\mathbf{V}}_{k(n)}$ such that

equation[equation omitted — 379 chars of source]

We call $v_{n}^{\ast }$ the sieve Riesz representer of the functional\ $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ on $\overline{ \mathbf{V}}_{k(n)}$. By definition, for any non-zero linear functional $ \frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$, we have:

equation*[equation* omitted — 229 chars of source]

is non-decreasing in $k(n)$.

We emphasize that the sieve Riesz representer $v_{n}^{\ast }$ of a linear functional\ $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ on $\overline{ \mathbf{V}}_{k(n)}$ always exists regardless of whether $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is bounded on the infinite dimensional space $( \mathbf{V},||\cdot ||)$ or not. Moreover, $v_{n}^{\ast }\in \overline{ \mathbf{V}}_{k(n)}$ and its norm $\left\Vert v_{n}^{\ast }\right\Vert $ can be computed in closed form (see Subsection (ref) ). The next Lemma allows us to verify whether or not $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is bounded on $(\mathbf{V},||\cdot ||)$ by checking whether or not $\lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert <\infty $.

lemmaLet $\{\overline{\mathbf{V}} _{k}\}_{k=1}^{\infty }$ be an increasing sequence of finite dimensional Hilbert spaces that is dense in $(\overline{\mathbf{V}},\left\Vert \cdot \right\Vert )$, and $v_{n}^{\ast }\in \overline{\mathbf{V}}_{k(n)}$ be defined in ((ref)). (1) If $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is bounded on $(\mathbf{V},||\cdot ||)$, then ((ref)) holds, $ v_{n}^{\ast }=\arg \min_{v\in \overline{\mathbf{V}}_{k(n)}}\left\Vert v^{\ast }-v\right\Vert $ and $\left\Vert v^{\ast }-v_{n}^{\ast }\right\Vert \rightarrow 0$, $\lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert =\left\Vert v^{\ast }\right\Vert <\infty $; (2) Let $\frac{ d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ be bounded on $(\mathbf{V},||\cdot ||_{s})$ and $\{\overline{\mathbf{V}}_{k}\}_{k=1}^{\infty }$ be dense in $( \mathbf{V},\left\Vert \cdot \right\Vert _{s})$. If $\frac{d\phi (\alpha _{0}) }{d\alpha }[\cdot ]$ is unbounded on $(\mathbf{V},||\cdot ||)$ then $ \lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert =\infty $.

Sieve score and sieve variance. For each sieve dimension $k(n)$, we call

equation[equation omitted — 174 chars of source]

the sieve score associated with the $i$-th observation, and $ \left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}\equiv Var\left( S_{n,i}^{\ast }\right) $ as the sieve variance. Recall that $\Sigma _{0}(X)\equiv Var(\rho (Z;\alpha _{0})|X)$ a.s.-$X$. Then

align[align omitted — 331 chars of source]

(See Subsection (ref) for closed form expressions of $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}$.) Under Assumption (ref)(iv), we have $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}\asymp \left\Vert v_{n}^{\ast }\right\Vert ^{2}$, and hence $ \lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert _{sd}<\infty $ (or $=\infty $) iff $\lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert <\infty $ (or $=\infty $). Therefore, in this paper we call $\phi ()$ regular (or irregular) at $\alpha _{0}$ whenever $\lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert <\infty $ (or $=\infty $), which, by Lemma (ref), is also whenever $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is bounded (or unbounded) on $(\mathbf{V},||\cdot ||)$. It is clear that our notion of a regular $\phi \left( \cdot \right) $ at $\alpha _{0}$ is only necessary but not sufficient for the existence of root-$n$ asymptotically normal regular estimators of $\phi \left( \alpha _{0}\right) $. Moreover, if $\phi \left( \cdot \right) $ is regular at $\alpha _{0}$ then we can define

equation*[equation* omitted — 154 chars of source]

as the score associated with the $i$-th observation, and $ \left\Vert v^{\ast }\right\Vert _{sd}^{2}\equiv Var\left( S_{i}^{\ast }\right) $ as the asymptotic variance. By Lemma (ref)(1) for a regular functional we have: $\left\Vert v^{\ast }\right\Vert _{sd}^{2}\asymp \left\Vert v^{\ast }\right\Vert <\infty $ and $Var\left( S_{i}^{\ast }-S_{n,i}^{\ast }\right) \asymp \left\Vert v^{\ast }-v_{n}^{\ast }\right\Vert ^{2}\rightarrow 0$ as $k(n)\rightarrow \infty $. See Appendix (ref) for further discussions.

Two key local conditions

For all $k(n)$, let

equation[equation omitted — 111 chars of source]

be the \textquotedblleft scaled sieve Riesz representer\textquotedblright . Since $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}\asymp \left\Vert v_{n}^{\ast }\right\Vert ^{2}$ (under Assumption (ref)(iv)), we have: $\left\Vert u_{n}^{\ast }\right\Vert \asymp 1$ and $\left\Vert u_{n}^{\ast }\right\Vert _{s}\leq c\tau _{n}$ for $\tau _{n}$ defined in ( (ref)) and a finite constant $c>0$.

Let $\mathcal{T}_{n}\equiv \{t\in \mathbb{R}\colon |t|\leq 4M_{n}^{2}\delta _{n}\}$ with $M_{n}$ and $\delta _{n}$ given in the definition of $\mathcal{N }_{osn}$.

assumption[Local behavior of $\protect\phi $] (i) $v\mapsto \frac{d\phi (\alpha _{0})}{d\alpha }[v]$ is a non-zero linear functional mapping from $\mathbf{V}$ to $\mathbb{R}$; $\{ \overline{\mathbf{V}}_{k}\}_{k=1}^{\infty }$ is an increasing sequence of finite dimensional Hilbert spaces that is dense in $(\overline{\mathbf{V}} ,\left\Vert \cdot \right\Vert )$; and $\frac{\left\Vert v_{n}^{\ast }\right\Vert }{\sqrt{n}}=o(1)$; \begin{equation*} (ii)\quad \sup_{(\alpha ,t)\in \mathcal{N}_{osn}\times \mathcal{T} _{n}}\frac{\sqrt{n}\left\vert \phi \left( \alpha +tu_{n}^{\ast }\right) -\phi (\alpha _{0})-\frac{d\phi (\alpha _{0})}{d\alpha }[\alpha +tu_{n}^{\ast }-\alpha _{0}]\right\vert }{\left\Vert v_{n}^{\ast }\right\Vert }=o\left( 1\right) ; \end{equation*} (iii) $\frac{\sqrt{n}\left\vert \frac{d\phi (\alpha _{0})}{d\alpha }[\alpha _{0,n}-\alpha _{0}]\right\vert }{\left\Vert v_{n}^{\ast }\right\Vert } =o\left( 1\right) .$

Since $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}\asymp \left\Vert v_{n}^{\ast }\right\Vert ^{2}$ (under Assumption (ref)(iv)), we could rewrite Assumption (ref) using $\left\Vert v_{n}^{\ast }\right\Vert _{sd}$ instead $\left\Vert v_{n}^{\ast }\right\Vert $. As it will become clear in Theorem (ref) that $\frac{\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}}{n}$ is the variance of $\phi (\widehat{ \alpha }_{n})-\phi (\alpha _{0})$, Assumption (ref)(i) puts a restriction on how fast the sieve dimension $k(n)$ could grow with the sample size $n$.

Assumption (ref)(ii) controls the nonlinearity bias of $\phi \left( \cdot \right) $ (i.e., the linear approximation error of a possibly nonlinear functional $\phi \left( \cdot \right) $). It is automatically satisfied when $\phi \left( \cdot \right) $ is a linear functional. For a nonlinear functional $\phi \left( \cdot \right) $ (such as the quadratic functional), it can be verified using the smoothness of $\phi \left( \cdot \right) $ and the convergence rates in both $||\cdot ||$ and $||\cdot ||_{s}$ metrics (the definition of $\mathcal{N}_{osn}$). See Section (ref) for verification.

Assumption (ref)(iii) controls the linear bias part due to the finite dimensional sieve approximation of $\alpha _{0,n}$ to $\alpha _{0}$. It is a condition imposed on the growth rate of the sieve dimension $k(n)$. When $\phi \left( \cdot \right) $ is an irregular functional, we have $ \left\Vert v_{n}^{\ast }\right\Vert \nearrow \infty $. Assumption (ref)(iii) requires that the sieve bias term, $\left\vert \frac{d\phi (\alpha _{0})}{d\alpha }[\alpha _{0,n}-\alpha _{0}]\right\vert $, is of a smaller order than that of the sieve standard deviation term, $ n^{-1/2}\left\Vert v_{n}^{\ast }\right\Vert _{sd}$. This is a standard condition imposed for the asymptotic normality of any plug-in nonparametric estimator of an irregular functional (such as a point evaluation functional of a nonparametric mean regression).

remarkWhen $\phi \left( \cdot \right) $ is regular at $\alpha _{0}$ (i.e., $\left\Vert v_{n}^{\ast }\right\Vert \nearrow \left\Vert v^{\ast }\right\Vert <\infty $), since $\left\langle v_{n}^{\ast },\alpha _{0,n}-\alpha _{0}\right\rangle =0$ (by definition of $\alpha _{0,n}$) we have $\left\vert \frac{d\phi (\alpha _{0})}{d\alpha }[\alpha _{0,n}-\alpha _{0}]\right\vert \leq \left\Vert v^{\ast }-v_{n}^{\ast }\right\Vert \times \left\Vert \alpha _{0,n}-\alpha _{0}\right\Vert $. And Assumption (ref)(iii) is satisfied if \begin{equation} ||v^{\ast }-v_{n}^{\ast }||\times ||\alpha _{0,n}-\alpha _{0}||=o(n^{-1/2}). \end{equation} This is similar to assumption 4.2 in AC_Emetrica03 and assumption 3.2(iii) in CP_WP07a for the root-$n$ estimable Euclidean parameter $ \theta _{0}$ of the model ((ref)). As pointed out by CP_WP07a, Condition ((ref)) could be satisfied when $\dim (\mathcal{A} _{k(n)})\asymp k(n)$ is chosen to obtain optimal nonparametric convergence rate in $||\cdot ||_{s}$ norm. But this nice feature only applies to regular functionals.

The next assumption is about the local quadratic approximation (LQA) to the sample criterion difference along the scaled sieve Riesz representer direction $u_{n}^{\ast }=v_{n}^{\ast }/\left\Vert v_{n}^{\ast }\right\Vert _{sd}$.

For any $(\alpha ,t)\in \mathcal{N}_{osn}\times \mathcal{T}_{n}$, we let $ \widehat{\Lambda }_{n}(\alpha (t),\alpha )\equiv 0.5\{\widehat{Q}_{n}(\alpha (t))-\widehat{Q}_{n}(\alpha )\}$ with $\alpha (t)\equiv \alpha +tu_{n}^{\ast }$. Denote

equation[equation omitted — 284 chars of source]
assumption[LQA] (i) $\alpha (t)\in \mathcal{A}_{k(n)}$ for any $(\alpha ,t)\in \mathcal{N}_{osn}\times \mathcal{T}_{n}$; and with $ r_{n}(t_{n})=\left( \max \{t_{n}^{2},t_{n}n^{-1/2},o(n^{-1})\}\right) ^{-1}$ , \begin{equation*} \sup_{(\alpha ,t_{n})\in \mathcal{N}_{osn}\times \mathcal{T} _{n}}r_{n}(t_{n})\left\vert \widehat{\Lambda }_{n}(\alpha (t_{n}),\alpha )-t_{n}\left\{ \mathbb{Z}_{n}+\langle u_{n}^{\ast },\alpha -\alpha _{0}\rangle \right\} -\frac{B_{n}}{2}t_{n}^{2}\right\vert =o_{P_{Z^{\infty }}}(1), \end{equation*} where, for each $n$, $B_{n}$ is a $Z^{n}$ measurable positive random variable, and $B_{n}=O_{P_{Z^{\infty }}}(1)$; (ii) $\sqrt{n}\mathbb{Z}_{n}\Rightarrow N(0,1)$.

Assumption (ref)(ii) is a standard one, and is implied by the following Lindeberg condition: For all $\epsilon >0$,

equation[equation omitted — 287 chars of source]

which, under Lemma (ref)(1) and Assumption (ref)(iv), is satisfied when the functional $\phi (\cdot )$ is regular ($\left\Vert v_{n}^{\ast }\right\Vert _{sd}\asymp \left\Vert v_{n}^{\ast }\right\Vert \rightarrow \left\Vert v^{\ast }\right\Vert <\infty $). This is why Assumption (ref)(ii) is not imposed in AC_Emetrica03 and CP_WP07a in their root-$n$ asymptotically normal estimation of the regular functional $\phi (\alpha )=\lambda ^{\prime }\theta $.

Assumption (ref)(i) implicitly imposes restrictions on the nonparametric estimator $\widehat{m}(x,\alpha )$ of $m(x,\alpha )=E[\rho (Z,\alpha )|X=x]$ in a shrinking neighborhood of $\alpha _{0}$, so that the criterion difference could be well approximated by a quadratic form. It is trivially satisfied when $\widehat{m}(x,\alpha )$ is linear in $\alpha $, such as the series LS estimator ((ref)) when $\rho (Z,\alpha )$ is linear in $\alpha $. There are two potential difficulties in verifying this assumption for nonlinear conditional moment models with nonparametric endogeneity (such as the NPQIV\ model). First, due to the non-smooth residual function $\rho (Z,\alpha )$, the estimator $\widehat{m}(x,\alpha )$ (and hence the sample criterion $\widehat{Q}_{n}(\alpha )$) could be pointwise non-smooth with respect to $\alpha $. Second, due to the slow convergence rates in the strong norm $||\cdot ||_{s}$ present in nonlinear nonparametric ill-posed inverse problems, it could be challenging to control the remainder of a quadratic approximation. When $\widehat{m}(x,\alpha )$ is the series LS estimator ((ref)), Lemma (ref) in Section (ref) shows that Assumption (ref)(i) is satisfied by a set of relatively low level sufficient conditions (Assumptions (ref) - (ref) in Appendix (ref)). See Section (ref) for verification of these sufficient conditions for functionals of the NPQIV model.

Asymptotic Properties of Sieve Wald and SQLR Statistics

In this section, we first establish the asymptotic normality of the plug-in PSMD estimator $\phi (\widehat{\alpha }_{n})$ of $\phi (\alpha _{0})$ for the model ((ref)), regardless of whether it is root-$n$ estimable or not. We then provide a simple consistent variance estimator and hence the asymptotic standard normality of the corresponding sieve t statistic for a real-valued functional $\phi :\mathbb{R}^{d_{\theta }}\times \mathcal{H} \rightarrow \mathbb{R}$. We finally derive the asymptotic properties of SQLR tests for the hypothesis $\phi (\alpha _{0})=\phi _{0}$. See Appendix (ref) for the case of a vector-valued functional $\phi :\mathbb{R} ^{d_{\theta }}\times \mathcal{H}\rightarrow \mathbb{R}^{d_{\phi }}$ (where $ d_{\phi }$ could grow slowly with $n$).

Asymptotic normality of the plug-in PSMD estimator

The next result allows for a (possibly) nonlinear irregular functional $\phi ()$ of the general model ((ref)).

theoremLet $\widehat{\alpha }_{n}$ be the PSMD estimator ( (ref)) and Assumptions (ref) - (ref) hold. If Assumptions (ref) and (ref) hold, then: \begin{equation*} \sqrt{n}\frac{\phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})}{||v_{n}^{\ast }||_{sd}}=-\sqrt{n}\mathbb{Z}_{n}+o_{P_{Z^{\infty }}}(1)\Rightarrow N(0,1). \end{equation*}

When the functional $\phi (\cdot )$ is regular at $\alpha =\alpha _{0}$, we have $\left\Vert v_{n}^{\ast }\right\Vert _{sd}\asymp \left\Vert v_{n}^{\ast }\right\Vert =O(1)$ and $\phi (\widehat{\alpha }_{n})$ converges to $\phi (\alpha _{0})$ at the parametric rate of $1/\sqrt{n}$. When the functional $ \phi (\cdot )$ is irregular at $\alpha =\alpha _{0}$, we have $\left\Vert v_{n}^{\ast }\right\Vert _{sd}\asymp \left\Vert v_{n}^{\ast }\right\Vert \rightarrow \infty $; so the convergence rate of $\phi (\widehat{\alpha } _{n})$ becomes slower than $1/\sqrt{n}$.

For any regular functional of the semi/nonparametric model ((ref)), Theorem (ref) implies that

equation*[equation* omitted — 203 chars of source]
eqnarray*[eqnarray* omitted — 359 chars of source]

Thus, Theorem (ref) is a natural extension of the asymptotic normality results of AC_Emetrica03 and CP_WP07a for the specific regular functional $\phi (\alpha _{0})=\lambda ^{\prime }\theta _{0} $ of the model ((ref)). See Remark (ref) in Appendix (ref) for further discussions.

Closed form expressions of sieve Riesz representer and sieve variance

To apply Theorem (ref), one needs to know the sieve Riesz representer $v_{n}^{\ast }$ defined in ((ref)) and the sieve variance $ \left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}$ given in ((ref)). It turns out that both can be computed in closed form.

lemmaLet $\overline{\mathbf{V}}_{k(n)}=\mathbb{R}^{d_{\theta }}\times \{v_{h}(\cdot )=\psi ^{k(n)}(\cdot )^{\prime }\beta :\beta \in \mathbb{R}^{k(n)}\}=\{v(\cdot )=\overline{\psi }^{k(n)}(\cdot )^{\prime }\gamma :\gamma \in \mathbb{R}^{d_{\theta }+k(n)}\}$ be dense in the infinite dimensional Hilbert space $(\overline{\mathbf{V}},\left\Vert \cdot \right\Vert )$ with the norm $\left\Vert \cdot \right\Vert $ defined in ((ref)). Then: the sieve Riesz representer $v_{n}^{\ast }=(v_{\theta ,n}^{\ast \prime },v_{h,n}^{\ast }\left( \cdot \right) )^{\prime }\in \overline{\mathbf{V}}_{k(n)}$ of $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ has a closed form expression: \begin{equation} v_{n}^{\ast }=(v_{\theta ,n}^{\ast \prime },\psi ^{k(n)}(\cdot )^{\prime }\beta _{n}^{\ast })^{\prime }=\overline{\psi }^{k(n)}(\cdot )^{\prime }\gamma _{n}^{\ast }, and \gamma _{n}^{\ast }=D_{n}^{-}\digamma _{n} \end{equation} with $D_{n}=E\left[ \left( \frac{dm(X,\alpha _{0})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\Sigma (X)^{-1}\left( \frac{ dm(X,\alpha _{0})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )^{\prime }]\right) \right] $ and $\digamma _{n}=\frac{d\phi (\alpha _{0})}{d\alpha }[ \overline{\psi }^{k(n)}(\cdot )]$. Thus \begin{equation} \left\Vert v_{n}^{\ast }\right\Vert ^{2}=\gamma _{n}^{\ast \prime }D_{n}\gamma _{n}^{\ast }=\digamma _{n}^{\prime }D_{n}^{-}\digamma _{n} . \end{equation} The sieve variance ((ref)) also has a closed form expression: \begin{equation} ||v_{n}^{\ast }||_{sd}^{2}=\digamma _{n}^{\prime }D_{n}^{-}\mho _{n}D_{n}^{-}\digamma _{n}, \end{equation} {{\begin{eqnarray*} & &\mho _{n}\equiv\\ & & E\left[ \left( \frac{dm(X,\alpha _{0})}{d\alpha }[\overline{ \psi }^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\Sigma (X)^{-1}\rho (Z,\alpha _{0})\rho (Z,\alpha _{0})^{\prime }\Sigma (X)^{-1}\left( \frac{ dm(X,\alpha _{0})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )^{\prime }]\right) \right] . \end{eqnarray*}}}

Let $\mathcal{A}_{k(n)}=\Theta \times \mathcal{H}_{k(n)}$ with $\mathcal{H} _{k(n)}$ given in ((ref)). Then $\overline{\mathbf{V}} _{k(n)}=clsp\left( \mathcal{A}_{k(n)}-\{\alpha _{0,n}\}\right) $ and one could let ${\psi }^{k(n)}(\cdot )= {q}^{k(n)}(\cdot )$ in Lemma (ref), and ((ref)) becomes the sieve variance expression given in ((ref)).

Lemmas (ref) and (ref) imply that $\phi \left( \cdot \right) $ is regular (or irregular) at $ \alpha =\alpha _{0}$ iff $\lim_{k(n)\rightarrow \infty }\left( \digamma _{n}^{\prime }D_{n}^{-}\digamma _{n}\right) <\infty $ (or $ =\infty $).

According to Lemma (ref) we could use different finite dimensional linear sieve basis $\psi ^{k(n)}$ to compute sieve Riesz representer $v_{n}^{\ast }=(v_{\theta ,n}^{\ast \prime },v_{h,n}^{\ast }\left( \cdot \right) )^{\prime }\in \overline{\mathbf{V}}_{k(n)}$, $ \left\Vert v_{n}^{\ast }\right\Vert ^{2}$ and $||v_{n}^{\ast }||_{sd}^{2}$. Most typical choices include orthonormal bases and the original sieve basis $ q^{k(n)}$ (used to approximate unknown function $h_{0}$). It is typically easier to characterize the speed of $\left\Vert v_{n}^{\ast }\right\Vert ^{2}=\digamma _{n}^{\prime }D_{n}^{-}\digamma _{n}$ as a function of $k(n)$ when an orthonormal basis is used, while there is a nice interpretation in terms of sieve variance estimation when the original sieve basis $q^{k(n)}$ is used. See Sections (ref), (ref) and (ref) for related discussions.

Consistent estimator of sieve variance of $\phi (\widehat{\alpha }_{n})$

In order to apply the asymptotic normality Theorem (ref), we need an estimator of the sieve variance $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}$ defined in ((ref)). We now provide one simple consistent estimator of the sieve variance when the residual function $\rho ()$ is pointwise smooth with respect to $\alpha _{0}$. See Appendix (ref) for additional consistent variance estimators.

The theoretical sieve Riesz representer $v_{n}^{\ast }$ is unknown but can be estimated easily. Let $\left\Vert \cdot \right\Vert _{n,M}$ denote the empirical norm induced by the following empirical inner product

equation[equation omitted — 275 chars of source]

for any $v_{1},v_{2}\in \overline{\mathbf{V}}_{k(n)}$, where $M_{n,i}$ is some (almost surely) positive definite weighting matrix.

We define an empirical sieve Riesz representer $\widehat{v} _{n}^{\ast }$ of the functional $\frac{d\phi (\widehat{\alpha }_{n})}{ d\alpha }[\cdot ]$ with respect to the empirical norm $||\cdot ||_{n, \widehat{\Sigma }^{-1}}$ as

equation[equation omitted — 260 chars of source]

and

equation[equation omitted — 205 chars of source]

For $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}=E\left( S_{n,i}^{\ast }S_{n,i}^{\ast \prime }\right) $ given in ((ref)) we can define a simple plug-in sieve variance estimator:

align[align omitted — 518 chars of source]

with $\widehat{\rho }_{i}=\rho (Z_{i},\widehat{\alpha }_{n})$ and $\widehat{ \Sigma }_{i}=\widehat{\Sigma }(X_{i})$.

Under the condition stated in Lemma (ref), $\widehat{v}_{n}^{\ast }$ defined in ((ref)-(ref)) also has a closed form solution:

equation[equation omitted — 223 chars of source]

with $\widehat{D}_{n}=\frac{1}{n}\sum_{i=1}^{n}\left( \frac{d\widehat{m} (X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\widehat{\Sigma }_{i}^{-1}\left( \frac{d \widehat{m}(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi } ^{k(n)}(\cdot )^{\prime }]\right) $ and $\widehat{\digamma }_{n}=\frac{d\phi (\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )]$. Hence the sieve variance estimator given in ((ref)) now becomes

equation[equation omitted — 230 chars of source]
equation*[equation* omitted — 424 chars of source]

In particular, with $\psi ^{k(n)}=q^{k(n)}$ the sieve variance estimator $|| \widehat{v}_{n}^{\ast }||_{n,sd}^{2}$ given in ((ref)) becomes the one given in ((ref)) in Subsection (ref).

Let $\langle v_{1},v_{2}\rangle _{M}\equiv E\left[ \left( \frac{dm(X,\alpha _{0})}{d\alpha }[v_{1}]\right) ^{\prime }M\left( \frac{dm(X,\alpha _{0})}{ d\alpha }[v_{2}]\right) \right] $. Then $\langle v_{1},v_{2}\rangle _{\Sigma ^{-1}}\equiv \langle v_{1},v_{2}\rangle $ for all $v_{1},v_{2}\in \overline{ \mathbf{V}}_{k(n)}$. Denote $\overline{\mathbf{V}}_{k(n)}^{1}\equiv \{v\in \overline{\mathbf{V}}_{k(n)}\colon ||v||=1\}$.

assumption(i) $\sup_{\alpha \in \mathcal{N}_{osn}}\sup_{v\in \overline{ \mathbf{V}}_{k(n)}^{1}}\left\vert \frac{d\phi (\alpha )}{d\alpha }[v]-\frac{ d\phi (\alpha _{0})}{d\alpha }[v]\right\vert =o(1)$; (ii) for each $k(n)$ and any $\alpha \in \mathcal{N}_{osn}$, $v\in \overline{\mathbf{V}}_{k(n)}\mapsto \frac{d\widehat{m}(\cdot ,\alpha )}{ d\alpha }[v]\in L^{2}(f_{X})$ is a linear functional measurable with respect to $Z^{n}$; and\\ $\sup_{v_{1},v_{2}\in \overline{\mathbf{V}} _{k(n)}^{1}}\left\vert \langle v_{1},v_{2}\rangle _{n,\Sigma ^{-1}}-\langle v_{1},v_{2}\rangle _{\Sigma ^{-1}}\right\vert =o_{P_{Z^{\infty }}}(1)$; (iii) $\sup_{x\in \mathcal{X}}||\widehat{\Sigma }(x)-\Sigma (x)||_{e}=o_{P_{Z^{\infty }}}(1)$; (iv) $\sup_{x\in \mathcal{X}}E\left[ \sup_{\alpha \in \mathcal{N} _{osn}}||\rho (Z,\alpha )\rho (Z,\alpha )^{\prime }-\rho (Z,\alpha _{0})\rho (Z,\alpha _{0})^{\prime }||_{e}|X=x\right] =o(1)$. (v) $\sup_{v\in \overline{\mathbf{V}}_{k(n)}^{1}}\left\vert \langle v,v\rangle _{n,M}-\langle v,v\rangle _{M}\right\vert =o_{P_{Z^{\infty }}}(1)$ with $M=\Sigma ^{-1}\rho (Z,\alpha _{0})\rho (Z,\alpha _{0})^{\prime }\Sigma ^{-1}$.

Assumption (ref)(i) becomes vacuous if $\phi $ is linear; otherwise it requires smoothness of the family $\{\frac{d\phi (\alpha )}{d\alpha } [v]:\alpha \in \mathcal{N}_{osn}\}$ uniformly in $v\in \overline{\mathbf{V}} _{k(n)}^{1}$. Assumption (ref)(ii) implicitly assumes that the residual function $\rho (z,\cdot )$ is \textquotedblleft smooth\textquotedblright\ in $\alpha \in \mathcal{N}_{osn}$ (see, e.g., AC_Emetrica03) or that $\frac{d\widehat{m}(X,\widehat{\alpha }_{n})}{ d\alpha }[v]$ can be well approximated by numerical derivatives (see, e.g., HMN_WP10). Assumption (ref)(iii) assumes the existence of consistent estimators for $\Sigma $. In most applications, $\Sigma (\cdot )$ is either completely known (such as the identity matrix) or $\Sigma _{0}$; while $\Sigma _{0}(x)$ could be consistently estimated via kernel, series LS, local linear regression and other nonparametric procedures (see, e.g., AC_Emetrica03 and CP_WP07a)

theoremLet Assumptions (ref) - (ref) hold. If Assumption (ref) is satisfied, then: (1) $\left\vert \frac{||\widehat{v}_{n}^{\ast }||_{n,sd}}{||v_{n}^{\ast }||_{sd}}-1\right\vert =o_{P_{Z^{\infty }}}(1)$ for $||\widehat{v}_{n}^{\ast }||_{n,sd}$ given in ((ref)). (2) If, in addition, Assumptions (ref) and (ref) hold, then: \begin{equation*} \widehat{W}_{n}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})}{||\widehat{v}_{n}^{\ast }||_{n,sd}}=-\sqrt{n}\mathbb{Z} _{n}+o_{P_{Z^{\infty }}}(1)\Rightarrow N(0,1). \end{equation*}

Theorem (ref)(2) allows us to construct confidence sets for $\phi (\alpha _{0})$ based on a possibly non-optimally weighted plug-in PSMD estimator $\phi (\widehat{\alpha }_{n})$. A potential drawback, is that it requires a consistent estimator for $v\mapsto \frac{dm(\cdot ,\alpha _{0})}{ d\alpha }[v]$, which may be hard to compute in practice when the residual function $\rho (Z,\alpha )$ is not pointwise smooth in $\alpha \in \mathcal{N }_{osn}$ such as in the NPQIV ((ref)) example.

remarkLet $\mathcal{W}_{n}\equiv \left( \sqrt{n}\frac{\phi ( \widehat{\alpha }_{n})-\phi _{0}}{||\widehat{v}_{n}^{\ast }||_{n,sd}}\right) ^{2}=\left( \widehat{W}_{n}+\sqrt{n}\frac{\phi (\alpha _{0})-\phi _{0}}{|| \widehat{v}_{n}^{\ast }||_{n,sd}}\right) ^{2}$ be the Wald test statistic. Then Theorem (ref) (with $\frac{||v_{n}^{\ast }||_{sd}}{\sqrt{n}} \asymp \frac{||v_{n}^{\ast }||}{\sqrt{n}}=o(1)$) immediately implies the following results: Under $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$, $\mathcal{W} _{n}=\left( \widehat{W}_{n}\right) ^{2}\Rightarrow \chi _{1}^{2}$. Under $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$, $\mathcal{W} _{n}=\left( O_{P}(1)+\sqrt{n}||v_{n}^{\ast }||_{sd}^{-1}[\phi (\alpha _{0})-\phi _{0}]\left( 1+o_{P}(1)\right) \right) ^{2}\rightarrow \infty $ in probability. See Theorem (ref) in Appendix (ref) for asymptotic properties of $\mathcal{W}_{n}$ under local alternatives.

Sieve QLR statistics

We now characterize the asymptotic behaviors of the possibly non-optimally weighted SQLR statistic $\widehat{QLR}_{n}(\phi _{0})$ defined in ((ref)).

Let $\mathcal{A}_{k(n)}^{R}\equiv \{\alpha \in \mathcal{A}_{k(n)}\colon \phi (\alpha )=\phi _{0}\}$ be the restricted sieve space, and $\widehat{\alpha } _{n}^{R}\in \mathcal{A}_{k(n)}^{R}$ be a restricted approximate PSMD estimator, defined as

equation[equation omitted — 245 chars of source]

Then:

align*[align* omitted — 315 chars of source]

Recall that $u_{n}^{\ast }\equiv v_{n}^{\ast }/\left\Vert v_{n}^{\ast }\right\Vert _{sd}$, and that $\widehat{QLR}_{n}^{0}(\phi _{0})$ denotes the optimally weighted (i.e., $\Sigma =\Sigma _{0}$) SQLR statistic in Subsection (ref). We note that $||u_{n}^{\ast }||=1$ for the optimally weighted case.

theoremLet Assumptions (ref) - (ref) hold with $ \left\vert B_{n}-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{Z^{\infty }}}(1)$. If $\widehat{\alpha }_{n}^{R}\in \mathcal{N}_{osn}$ wpa1-$P_{Z^{\infty }}$, then: (1) under the null $H_{0}:\phi (\alpha _{0})=\phi _{0}$, \begin{equation*} ||u_{n}^{\ast }||^{2}\times \widehat{QLR}_{n}(\phi _{0})=\left( \sqrt{n} \mathbb{Z}_{n}\right) ^{2}+o_{P_{Z^{\infty }}}(1)\Rightarrow \chi _{1}^{2}. \end{equation*} (2) Further, let $\widehat{\alpha }_{n}$ be the optimally weighted PSMD estimator ((ref)) with $\Sigma =\Sigma _{0}$. Then: under $H_{0}:$ $ \phi (\alpha _{0})=\phi _{0}$, \begin{equation*} \widehat{QLR}_{n}^{0}(\phi _{0})=\left( \sqrt{n}\mathbb{Z}_{n}\right) ^{2}+o_{P_{Z^{\infty }}}(1)\Rightarrow \chi _{1}^{2}. \end{equation*} See Theorem (ref) in Appendix (ref) for the asymptotic behavior under local alternatives.

Compared to Theorem (ref) on the asymptotic normality of $ \phi (\widehat{\alpha }_{n})$, Theorem (ref) on the asymptotic null distribution of the SQLR statistic requires two extra conditions: $ \left\vert B_{n}-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{Z^{\infty }}}(1)$ and $\widehat{\alpha }_{n}^{R}\in \mathcal{N}_{osn}$ wpa1-$P_{Z^{\infty }}$. Both conditions are also needed even for QLR statistics in parametric extremum estimation and testing problems. Lemma (ref) in Section (ref) provides a simple sufficient condition (Assumption (ref)) for $\left\vert B_{n}-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{Z^{\infty }}}(1)$. Proposition (ref) in Appendix (ref) establishes $\widehat{\alpha }_{n}^{R}\in \mathcal{N}_{osn}$ wpa1-$P_{Z^{\infty }}$ under the null $H_{0}:\phi (\alpha _{0})=\phi _{0}$ and other conditions virtually the same as those for Lemma (ref) (i.e., $\widehat{\alpha }_{n}\in \mathcal{N}_{osn}$ wpa1-$P_{Z^{\infty }}$).

Theorem (ref)(2) recommends to construct an asymptotic $100(1-\tau )\%$ confidence set for $\phi (\alpha )$ by inverting the optimally weighted SQLR statistic: $\left\{ r\in \mathbb{R}\colon \text{ }\widehat{QLR} _{n}^{0}(r)\leq c_{\chi _{1}^{2}}(1-\tau )\right\} $. This result extends that of CP_WP07a for a regular Euclidean functional $\phi (\alpha )=\lambda ^{\prime }\theta $ to possibly irregular nonlinear functionals.

Next, we consider the asymptotic behavior of $\widehat{QLR}_{n}(\phi _{0})$ under the fixed alternatives $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$.

theoremLet Assumptions (ref), (ref) and (ref) hold. Suppose that $\sup_{h\in \mathcal{H}}Pen(h)<\infty $ and $ \phi $ is continuous in $||\cdot ||_{s}$. Then: under $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$, there is a constant $C>0$ such that \begin{equation*} \frac{\widehat{QLR}_{n}(\phi _{0})}{n}\geq C>0\quad wpa1. \end{equation*}

Inference Based on Generalized Residual Bootstrap

\setcounter{assumption}{0}

The inference procedures described in Subsections (ref) and (ref) are based on the asymptotic critical values. For many parametric models it is known that bootstrap based procedures could approximate finite sample distributions more accurately. In this section we establish the consistency of the bootstrap sieve Wald and SQLR statistics under virtually the same conditions as those imposed for the original-sample sieve Wald and SQLR statistics.

A bootstrap procedure is described by an array of \textquotedblleft weights\textquotedblright\ $\left\{ \omega _{i,n}\right\} _{i=1}^{n}$ for each $n$, where each bootstrap sample is drawn independently of the original data $\left\{ Z_{i}\right\} _{i=1}^{n}$. Different bootstrap procedures correspond to different choices of the weights $\left\{ \omega _{i,n}\right\} _{i=1}^{n}$ but all satisfy $\omega _{i,n}\geq 0$ and $ E[\omega _{i,n}]=1$. For the time being we assume that $\lim_{n\rightarrow \infty }Var(\omega _{i,n})=\sigma _{\omega }^{2}\in (0,\infty )$ for all $i$.

In this paper we focus on two types of bootstrap weights:

assumption[I.i.d Weights] Let $(\omega _{i})_{i=1}^{n}$ be a sequence such that $ \omega _{i}\in \mathbb{R}_{+}$, $\omega _{i}\sim iidP_{\omega }$, $E[\omega ]=1$, $Var(\omega )=\sigma _{\omega }^{2}$, and $\int_{0}^{\infty }\sqrt{ P(|\omega -1|\geq t)}dt<\infty $.

The condition $\int_{0}^{\infty }\sqrt{P(|\omega -1|\geq t)}dt<\infty $ is implied by $E[|\omega -1|^{2+\epsilon }]<\infty $ for some $\epsilon >0$.

assumption[Multinomial Weights] Let $(\omega _{i,n})_{i=1}^{n}$ be a triangular array of random variables such that $(\omega _{1,n},...,\omega _{n,n})\sim Multinomial(n;n^{-1},...,n^{-1})$.

We sometimes omit the $n$ subscript from the weight series. Note that under Assumption (ref), $E[\omega _{1}]=1$, $Var(\omega _{1})=(1-1/n)\rightarrow 1\equiv \sigma _{\omega }^{2}$ and $Cov(\omega _{i},\omega _{j})=-n^{-1}$ (for $i\neq j$). Finally, $n^{-1}\max_{1\leq i\leq n}(\omega _{i}-1)^{2}=o_{P_{\omega }}(1)$. We use these facts in the proofs.

Let $V_{i}\equiv (Z_{i},\omega _{i,n})$ and

equation*[equation* omitted — 82 chars of source]

be the bootstrap residual function. Let $\widehat{m}^{B}(x,\alpha )$ be a bootstrap version of $\widehat{m}(x,\alpha )$, that is, $\widehat{m} ^{B}(x,\alpha )$ is computed in the same way as that of $\widehat{m} (x,\alpha )$ except that we use $\rho ^{B}(V_{i},\alpha )$ instead of $\rho (Z_{i},\alpha )$. In particular, $\widehat{m}^{B}(x,\alpha )=\sum_{i=1}^{n}\omega _{i,n}\rho (Z_{i},\alpha )A_{n}(X_{i},x)$ for any linear estimator $\widehat{m}(x,\alpha )$ ((ref)) of $m(x,\alpha )$. For example, if $\widehat{m}(x,\alpha )$ is a series LS estimator ((ref)), then $\widehat{m}^{B}(x,\alpha )$ is the bootstrap series LS estimator ((ref)) defined in Subsection (ref).

Let $\widehat{Q}_{n}^{B}(\alpha )\equiv \frac{1}{n}\sum_{i=1}^{n}\widehat{m} ^{B}(X_{i},\alpha )^{\prime }\widehat{\Sigma }(X_{i})^{-1}\widehat{m} ^{B}(X_{i},\alpha )$ be a bootstrap version of $\widehat{Q}_{n}(\alpha )$, and $\widehat{\alpha }_{n}^{B}$ be the bootstrap PSMD estimator, i.e., $ \widehat{\alpha }_{n}^{B}$ is an approximate minimizer of $\left\{ \widehat{Q }_{n}^{B}(\alpha )+\lambda _{n}Pen(h)\right\} $ on $\mathcal{A}_{k(n)}$. Denote $\widehat{\phi }_{n}\equiv \phi (\widehat{\alpha }_{n})$. Then

equation*[equation* omitted — 222 chars of source]

is the (generalized residual) bootstrap SQLR test statistic. And $\mathcal{W} _{1,n}^{B}\equiv \left( \sqrt{n}\frac{\phi (\widehat{\alpha }_{n}^{B})- \widehat{\phi }_{n}}{\sigma _{\omega }||\widehat{v}_{n}^{\ast }||_{n,sd}} \right) ^{2}$ is one simple bootstrap Wald test statistic (see Subsection (ref) for another simple bootstrap Wald statistic).

Additional notation. To be more precise, we introduce some definitions associated with the new random variables $V_{i}\equiv (Z_{i},\omega _{i,n})$ and the enlarged probability spaces. Let $\Omega =\{\omega _{i,n}\colon i=1,...,n;~n=1,...\}$ be the space of weights, defined as a triangle array with elements in $\mathbb{R}$, the corresponding $\sigma $-algebra and probability are $(\mathcal{B}_{\Omega },P_{\Omega })$. Let $\mathcal{V}^{\infty }\equiv \mathcal{Z}^{\infty }\times \Omega $, $ \mathcal{B}^{\infty }\equiv \mathcal{B}_{Z}^{\infty }\times \mathcal{B} _{\Omega }$ be the $\sigma $-algebra, and $P_{V^{\infty }}$ be the joint probability over $\mathcal{V}^{\infty }$. Finally, for each $n$, let $ \mathcal{B}^{n}$ be the $\sigma $-algebra generated by $V^{n}\equiv Z^{n}\times (\omega _{1,n},...,\omega _{n,n})$, where each $\omega _{i,n}$ acts as a \textquotedblleft weight\textquotedblright\ of $Z_{i}$. Let $A_{n}$ be a random variable that is measurable with respect to $\mathcal{B}^{n}$, and $\mathcal{L}_{V^{\infty }|Z^{\infty }}(A_{n}|Z^{n})$ (or $P_{V^{\infty }|Z^{\infty }}\left( A_{n}\leq \cdot \mid Z^{n}\right) $) be the conditional law (or conditional distribution) of $A_{n}$ given $Z^{n}$. Let $B_{n}$ be a random variable measurable with respect to $\mathcal{B}_{Z}^{\infty }$, and $ \mathcal{L}(B_{n})$ (or $P_{Z^{\infty }}\left( B_{n}\leq \cdot \right) $) be the law (or distribution) of $B_{n}$. For two real valued random variables, $ A_{n}$ (measurable with respect to $\mathcal{B}^{n}$) and $B$ (measurable with respect to some $\sigma $-algebra $\mathcal{B}_{B}$), we say $ \left\vert \mathcal{L}_{V^{\infty }|Z^{\infty }}(A_{n}|Z^{n})-\mathcal{L} (B)\right\vert =o_{P_{Z^{\infty }}}(1)$ if for any $\delta >0$, there exists a $N(\delta )$ such that

equation*[equation* omitted — 178 chars of source]

(i.e., $\sup_{f\in BL_{1}}\left\vert E[f(A_{n})|Z^{n}]-E[f(B)]\right\vert =o_{P_{Z^{\infty }}}(1)$), where $BL_{1}$ denotes the class of uniformly bounded Lipschitz functions $f:\mathbb{R}\rightarrow \mathbb{R}$ such that $ ||f||_{L^{\infty }}\leq 1$ and $|f(z)-f(z^{\prime })|\leq |z-z^{\prime }|$. See chapter 1.12 of VdV-W_book96 (henceforth, VdV-W) for more details.

We say $\Delta _{n}$ is of order $o_{P_{V^{\infty }|Z^{\infty }}}(1)$ in $ P_{Z^{\infty }}$ probability, and denote it as $\Delta _{n}=o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$, if for any $\epsilon >0$, $ P_{Z^{\infty }}\left( P_{V^{\infty }|Z^{\infty }}\left( |\Delta _{n}|>\epsilon \mid Z^{n}\right) >\epsilon \right) \rightarrow 0$ as $ n\rightarrow \infty $.

We say $\Delta _{n}$ is of order $O_{P_{V^{\infty }|Z^{\infty }}}(1)$ in $ P_{Z^{\infty }}$ probability, and denote it as $\Delta _{n}=O_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$, if for any $\epsilon >0$ there exists a $M\in (0,\infty )$, such that $P_{Z^{\infty }}\left( P_{V^{\infty }|Z^{\infty }}\left( |\Delta _{n}|>M\mid Z^{n}\right) >\epsilon \right) \rightarrow 0$ as $n\rightarrow \infty $.

Bootstrap local quadratic approximation (LQA$^{B}$)

Lemma (ref) in Appendix (ref) shows that the bootstrap PSMD estimator $\widehat{\alpha }_{n}^{B}\in \mathcal{N}_{osn}$ wpa1 under Assumptions (ref) and (ref) - (ref). This allows us to introduce a condition that is a bootstrap version of the LQA Assumption (ref). For any $\alpha \in \mathcal{N}_{osn}$, we let $\widehat{\Lambda }_{n}^{B}(\alpha (t_{n}),\alpha )\equiv 0.5\{\widehat{Q}_{n}^{B}(\alpha (t_{n}))-\widehat{Q}_{n}^{B}(\alpha )\}$ with $\alpha (t_{n})\equiv \alpha +t_{n}u_{n}^{\ast }$ for $t_{n}\in \mathcal{T}_{n}$. For any sequence of non-negative weights $(b_{i})_{i}$, let

equation*[equation* omitted — 282 chars of source]
assumption[LQA$^{B}$] (i) $\alpha (t)\in \mathcal{A}_{k(n)}$ for any $(\alpha ,t)\in \mathcal{N}_{osn}\times \mathcal{T}_{n}$, and with $ r_{n}(t_{n})=\left( \max \{t_{n}^{2},t_{n}n^{-1/2},o(n^{-1})\}\right) ^{-1}$ , \begin{eqnarray*} & & \sup_{(\alpha ,t_{n})\in \mathcal{N}_{osn}\times \mathcal{T} _{n}}r_{n}(t_{n})\left\vert \widehat{\Lambda }_{n}^{B}(\alpha (t_{n}),\alpha )-t_{n}\left\{ \mathbb{Z}_{n}^{\omega }+\langle u_{n}^{\ast },\alpha -\alpha _{0}\rangle \right\} -\frac{B_{n}^{\omega }}{2}t_{n}^{2}\right\vert\\ & & = o_{P_{V^{\infty }|Z^{\infty }}}(1) wpa1(P_{Z^{\infty }}) \end{eqnarray*} where $B_{n}^{\omega }$ is a $V^{n}$ measurable positive random variable such that $B_{n}^{\omega }=O_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$; \begin{equation*} (ii)\quad \left\vert \mathcal{L}_{V^{\infty }|Z^{\infty }}\left( \sqrt{n}\frac{\mathbb{Z}_{n}^{\omega -1}}{\sigma _{\omega }}\mid Z^{n}\right) -\mathcal{L}\left( \mathbb{Z}\right) \right\vert =o_{P_{Z^{\infty }}}(1), \end{equation*} where $\mathbb{Z}$ is a standard normal random variable.

Assumption (ref)(i) implicitly imposes restrictions on the bootstrap estimator $\widehat{m}^{B}(x,\alpha )$ of the conditional mean function $m(x,\alpha )$. Below we provide low level sufficient conditions for Assumption (ref)(i) when $\widehat{m}^{B}(x,\alpha )$ is a bootstrap series LS estimator.

Let $g(X,u_{n}^{\ast })\equiv \{\frac{dm(X,\alpha _{0})}{d\alpha } [u_{n}^{\ast }]\}^{\prime }\Sigma (X)^{-1}$. Then $E\left[ g(X_{i},u_{n}^{\ast })\Sigma (X_{i})g(X_{i},u_{n}^{\ast })^{\prime }\right] =||u_{n}^{\ast }||^{2}$.

\setcounter{assumption}{0}

assumptionFor $\Gamma (\cdot )\in \{\Sigma (\cdot ),\Sigma _{0}(\cdot )\}$, \begin{equation*} \left\vert n^{-1}\sum_{i=1}^{n}g(X_{i},u_{n}^{\ast })\Gamma (X_{i})g(X_{i},u_{n}^{\ast })^{\prime }-E\left[ g(X_{i},u_{n}^{\ast })\Gamma (X_{i})g(X_{i},u_{n}^{\ast })^{\prime }\right] \right\vert =o_{P_{Z^{\infty }}}(1). \end{equation*}
lemmaLet Assumptions (ref) - (ref) and (ref) - (ref) hold. (1) Let $\widehat{m}$ be the series LS estimator ((ref)). Then Assumption (ref)(i) is satisfied. Further, if Assumption (ref) holds then $\left\vert B_{n}-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{Z^{\infty }}}(1)$. (2) Let $\widehat{m}^{B}(\cdot ,\alpha )$ be the bootstrap series LS estimator ((ref)), Assumption (ref), and either Assumption (ref) or (ref) hold. Then Assumption (ref)(i) holds with $B_{n}^{\omega }=B_{n}$. Further, if Assumption (ref) holds then $\left\vert B_{n}^{\omega }-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$.

Lemma (ref) indicates that the low level Assumptions (ref) - (ref) are sufficient for both the original-sample LQA Assumption (ref)(i) and the bootstrap LQA Assumption (ref)(i).

Assumption (ref)(ii) can be easily verified by applying some central limit theorems. For example, if the weights are independent (Assumption (ref)), we can use Lindeberg-Feller CLT; if the weights are multinomial (Assumption (ref)) we can apply Hayek CLT (see VdV-W_book96 p. 458 ). The next lemma provides some simple sufficient conditions for Assumption (ref)(ii).

lemmaLet either Assumption (ref) or Assumption (ref) hold. If there is a positive real sequence $(b_{n})_{n}$ such that $b_{n}=o\left( \sqrt{n}\right) $ and \begin{equation} \limsup_{n\rightarrow \infty }E\left[ \left( g(X,u_{n}^{\ast })\rho (Z,\alpha _{0})\right) ^{2}1\left\{ \frac{(g(X,u_{n}^{\ast })\rho (Z,\alpha _{0}))^{2}}{b_{n}}>1\right\} \right] =0, \end{equation} then Assumptions (ref)(ii) and (ref)(ii) hold.

Bootstrap sieve Student t statistic

\setcounter{assumption}{3}

Lemma (ref) shows that $\widehat{\alpha }_{n}^{B}\in \mathcal{N }_{osn}$ wpa1 under virtually the same conditions as those for the original-sample estimator $\widehat{\alpha }_{n}\in \mathcal{N}_{osn}$ wpa1. This would easily lead to the consistency of the simplest bootstrap sieve t statistic $\widehat{W}_{1,n}^{B}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha } _{n}^{B})-\phi (\widehat{\alpha }_{n})}{\sigma _{\omega }||\widehat{v} _{n}^{\ast }||_{n,sd}}$.

We now establish the consistency of another bootstrap sieve t statistic $ \widehat{W}_{2,n}^{B}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha } _{n}^{B})-\phi (\widehat{\alpha }_{n})}{||\widehat{v}_{n}^{\ast }||_{B,sd}}$ , where $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}$ is a bootstrap sieve variance estimator:

equation[equation omitted — 455 chars of source]

with $\varrho (V_{i},\alpha )\equiv (\omega _{i,n}-1)\rho (Z_{i},\alpha )\equiv \rho ^{B}(V_{i},\alpha )-\rho (Z_{i},\alpha )$ for any $\alpha $.

We note that $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}$ is an analog to $|| \widehat{v}_{n}^{\ast }||_{n,sd}^{2}$ defined in ((ref)) but using the bootstrapped generalized residual $\varrho (V_{i},\widehat{\alpha }_{n})$ instead of the original sample fitted residual $\rho (Z_{i},\widehat{\alpha } _{n})$. It also has a closed form expression: $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}=\widehat{\digamma }_{n}^{\prime }\widehat{D}_{n}^{-}\widehat{ \mho }_{n}^{B}\widehat{D}_{n}^{-}\widehat{\digamma }_{n}$ with

equation*[equation* omitted — 345 chars of source]

where $\widehat{M}_i \equiv \widehat{\Sigma }_{i}^{-1}\rho (Z_{i},\widehat{\alpha }_{n})\rho (Z_{i},\widehat{\alpha } _{n})^{\prime }\widehat{\Sigma }_{i}^{-1}$. That is, $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}$ is computed in the same way as $||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}=\widehat{\digamma } _{n}^{\prime }\widehat{D}_{n}^{-}\widehat{\mho }_{n}\widehat{D}_{n}^{-} \widehat{\digamma }_{n}$ given in ((ref)) except using $\widehat{ \mho }_{n}^{B}$ instead of $\widehat{\mho }_{n}$.

assumption$\sup_{v\in \overline{\mathbf{V}}_{k(n)}^{1}}|\langle v,v\rangle _{n,\widehat{M}^{B}}-\sigma _{\omega }^{2}\langle v,v\rangle _{n,\widehat{ M}}|=o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$ with $\widehat{M} _{i}^{B}=(\omega _{i,n}-1)^{2}\widehat{M}_{i}$.

This assumption can be verified given Assumptions (ref) or (ref). The following result is a bootstrap version of Theorem (ref)(1).

theoremLet Assumptions (ref) - (ref), (ref) and (ref) hold. Then: \begin{equation*} \left\vert \frac{||\widehat{v}_{n}^{\ast }||_{B,sd}}{\sigma _{\omega }||v_{n}^{\ast }||_{sd}}-1\right\vert =o_{P_{V^{\infty }|Z^{\infty }}}(1) wpa1(P_{Z^{\infty }}). \end{equation*}

Recall that $\widehat{W}_{n}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha } _{n})-\phi (\alpha _{0})}{||\widehat{v}_{n}^{\ast }||_{n,sd}}$, whose probability distribution $P_{Z^{\infty }}\left( \widehat{W}_{n}\leq \cdot \right) $ converges to the standard normal cdf $\Phi (\cdot )$. The next result is about the consistency of the bootstrap sieve t statistic $\widehat{ W}_{2,n}^{B}$.

theoremLet $\widehat{\alpha }_{n}$ be the PSMD estimator ( (ref)) and $\widehat{\alpha }_{n}^{B}$ the bootstrap PSMD estimator. Let Assumptions (ref) - (ref) and (ref) hold. Let Assumptions (ref), (ref) and (ref) hold. (1) Let Assumptions (ref) and (ref) hold. Then: \begin{equation*} \sup_{t\in \mathbb{R}}\left\vert P_{V^{\infty }|Z^{\infty }}\left( \widehat{W }_{2,n}^{B}\leq t\mid Z^{n}\right) -P_{Z^{\infty }}\left( \widehat{W} _{n}\leq t\right) \right\vert =o_{P_{V^{\infty }|Z^{\infty }}}(1) wpa1(P_{Z^{\infty }}). \end{equation*} (2) If $\phi ()$ is regular at $\alpha _{0}$, without imposing Assumptions (ref) and (ref), we have: \begin{eqnarray*} & & \sup_{t\in \mathbb{R}}\left\vert P_{V^{\infty }|Z^{\infty }}\left( \sqrt{n} \frac{\phi (\widehat{\alpha }_{n}^{B})-\phi (\widehat{\alpha }_{n})}{\sigma _{\omega }}\leq t\mid Z^{n}\right) -P_{Z^{\infty }}\left( \sqrt{n}\left( \phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})\right) \leq t\right) \right\vert \\ && = o_{P_{V^{\infty }|Z^{\infty }}}(1) wpa1(P_{Z^{\infty }}). \end{eqnarray*}

For a regular functional, Theorem (ref)(2) provides one way to construct its confidence sets without the need to compute any variance estimator. This extends the result in CP_WP07a for a regular Euclidean parameter $\lambda ^{\prime }\theta $ to a general regular functional $\phi (\alpha )$. Unfortunately for an irregular functional, we need to compute a consistent bootstrap sieve variance estimator $||\widehat{v }_{n}^{\ast }||_{B,sd}^{2}$ to apply Theorem (ref)(1). Luckily $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}$ is easy to compute when the residual function $\rho (Z_{i},\alpha )$ is pointwise smooth in $\alpha _{0}$ . Moreover, since $E\left( ||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}\mid Z^{n}\right) =\sigma _{\omega }^{2}||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}$ we suspect that the bootstrap sieve t statistic $\widehat{W}_{2,n}^{B}$ might have second order refinement property by choices of bootstrap weights $ \{\omega _{i,n}\}$. This will be a subject of future research.

The bootstrap sieve t statistic $\widehat{W}_{2,n}^{B}$ requires to compute the original sample PSMD estimator $\widehat{\alpha }_{n}$ and the bootstrap PSMD estimator $\widehat{\alpha }_{n}^{B}$. In the online Appendix (ref) we present a sieve score test and its bootstrap version, which only use the original sample restricted PSMD estimator $\widehat{ \alpha }_{n}^{R}$ and do not use $\widehat{\alpha }_{n}^{B}$, and hence are computationally simple.

remarkTheorems (ref)(2) and (ref)(1) imply that the bootstrap Wald test statistic $\mathcal{W}_{2,n}^{B}\equiv \left( \widehat{W}_{2,n}^{B}\right) ^{2}$ always has the same limiting distribution $\chi _{1}^{2}$ (conditional on the data) under the null and the alternatives. Let $\widehat{c}_{2,n}(a)$ be the $a-th$ quantile of the distribution of $\mathcal{W}_{2,n}^{B}$ (conditional on the data $ \{Z_{i}\}_{i=1}^{n}$). Let $\mathcal{W}_{n}\equiv \left( \sqrt{n}\frac{\phi ( \widehat{\alpha }_{n})-\phi _{0}}{||\widehat{v}_{n}^{\ast }||_{n,sd}}\right) ^{2}$ be the original sample Wald test statistic. Then Remark (ref) and Theorem (ref)(1) immediately imply that for any $\tau \in (0,1)$, under $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$, $\lim_{n\rightarrow \infty }\Pr \left( \mathcal{W}_{n}\geq \widehat{c}_{2,n}(1-\tau )\right) =\tau $; under $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$, $\lim_{n\rightarrow \infty }\Pr \left( \mathcal{W}_{n}\geq \widehat{c}_{2,n}(1-\tau )\right) =1.$ See Theorem (ref) in Appendix (ref) for properties under local alternatives.

See online supplemental Appendix (ref) for consistency of $\mathcal{ W}_{1,n}^{B}\equiv \left( \sqrt{n}\frac{\phi (\widehat{\alpha }_{n}^{B})- \widehat{\phi }_{n}}{\sigma _{\omega }||\widehat{v}_{n}^{\ast }||_{n,sd}} \right) ^{2}$ and other bootstrap sieve Wald (t) statistics based on different sieve variance estimators.

Bootstrap SQLR statistic

If $\Sigma \neq \Sigma _{0}$, the SQLR statistic $\widehat{QLR}_{n}(\phi _{0})=n\left( \widehat{Q}_{n}(\widehat{\alpha }_{n}^{R})-\widehat{Q}_{n}( \widehat{\alpha }_{n})\right) $ is no longer asymptotically chi-square even under the null; Theorem (ref)(1), however, implies that the SQLR statistic converges weakly to a tight limit under the null. In this subsection we show that the asymptotic null distribution of the SQLR can be consistently approximated by that of the (generalized residual) bootstrap SQLR statistic $\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})$. Recall that

equation*[equation* omitted — 230 chars of source]

where $\widehat{\phi }_{n}\equiv \phi (\widehat{\alpha }_{n})$, and $ \widehat{\alpha }_{n}^{R,B}$ is the restricted bootstrap PSMD estimator, defined as

eqnarray[eqnarray omitted — 338 chars of source]

Lemma (ref) in Appendix (ref) implies that $\widehat{ \alpha }_{n}^{R,B},$ $\widehat{\alpha }_{n}^{B}\in \mathcal{N}_{osn}$ wpa1 under both the null $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$ and the alternatives $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$. This indicates that the bootstrap SQLR statistic $\widehat{QLR}_{n}^{B}(\widehat{\phi } _{n}) $ is always properly centered and should be stochastically bounded under both the null and the alternatives, as shown in the next theorem. Let $ P_{Z^{\infty }}\left( \widehat{QLR}_{n}(\phi _{0})\leq \cdot \mid H_{0}\right) $ denote the probability distribution of $\widehat{QLR} _{n}(\phi _{0})$ under the null $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$, which would converge to the cdf of $\chi _{1}^{2}$ when $\widehat{QLR} _{n}(\phi _{0})=\widehat{QLR}_{n}^{0}(\phi _{0})$ (the optimally weighted SQLR).

theoremLet Assumptions (ref) - (ref) and (ref) hold. Let Assumptions (ref), (ref) and (ref) hold with $\left\vert B_{n}^{\omega }-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$ . Then: \begin{equation*} (1) \frac{\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})}{\sigma _{\omega }^{2}}=\left( \sqrt{n}\frac{\mathbb{Z}_{n}^{\omega -1}}{\sigma _{\omega }||u_{n}^{\ast }||}\right) ^{2}+o_{P_{V^{\infty }|Z^{\infty }}}(1)=O_{P_{V^{\infty }|Z^{\infty }}}(1) wpa1(P_{Z^{\infty }}); \end{equation*} and \begin{eqnarray*} (2) && \sup_{t\in \mathbb{R}}\left\vert P_{V^{\infty }|Z^{\infty }}\left( \frac{\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})}{\sigma _{\omega }^{2}}\leq t\mid Z^{n}\right) -P_{Z^{\infty }}\left( \widehat{QLR}_{n}(\phi _{0})\leq t\mid H_{0}\right) \right\vert \\ && =o_{P_{V^{\infty }|Z^{\infty }}}(1) wpa1(P_{Z^{\infty }}). \end{eqnarray*}

Theorem (ref) allows us to construct valid confidence sets (CS) for $\phi (\alpha _{0})$ based on inverting possibly non -optimally weighted SQLR statistic without the need to compute a variance estimator. We recommend this procedure when it is difficult to compute any consistent variance estimator for $\phi (\widehat{\alpha })$, such as in the cases when the residual function $\rho (Z;\alpha )$ is pointwise non-smooth in $\alpha _{0}$. See, e.g., AB_Emetrica00 for a thorough discussion about how to construct CS via bootstrap.

remarkLet $\widehat{c}_{n}(a)$ be the $a-th$ quantile of the distribution of $\frac{\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})}{\sigma _{\omega }^{2}}$ (conditional on the data $\{Z_{i}\}_{i=1}^{n}$). Then Theorems (ref), (ref) and (ref) immediately imply that for any $\tau \in (0,1)$, under $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$, $\lim_{n\rightarrow \infty }\Pr \left( \widehat{QLR}_{n}(\phi _{0})\geq \widehat{c}_{n}(1-\tau )\right) =\tau $; under $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$, $\lim_{n\rightarrow \infty }\Pr \left( \widehat{QLR}_{n}(\phi _{0})\geq \widehat{c}_{n}(1-\tau )\right) =1.$ See Theorem (ref) in Appendix (ref) for properties under local alternatives.

Verification of Assumptions 3.5 and 3.6

In this section, we illustrate the verification of the two key regularity conditions, Assumption (ref) and Assumption (ref)(i), via some functionals $\phi (h)$ of the (nonlinear) nonparametric IV regressions:

equation[equation omitted — 77 chars of source]

where the scalar valued residual function $\rho ()$ could be nonlinear and pointwise non-smooth in $h$. This model includes the NPIV and NPQIV as special cases. To be concrete, we consider a PSMD estimator $\widehat{h}\in \mathcal{H}_{k(n)}$ of $h_{0}$ with $\widehat{\Sigma }=\Sigma =1$, and $ \widehat{m}(\cdot ,h)$ being the series LS estimator ((ref)) of $ m(\cdot ,h)=E[\rho (Y_{1};h(Y_{2}))|X=\cdot ]$ with $J_{n}=ck(n)$ for a finite constant $c\geq 1$. We assume that $h_{0}\in \mathcal{H}=\Lambda _{c}^{\varsigma }\left( [-1,1]\right) $ with smoothness $\varsigma >1/2$ (a H \"{o}lder ball with support $[-1,1]$, see, e.g., CLvK_Emetrica03). \footnote{ This H\"{o}lder ball condition and several other conditions assumed in this subsection are for illustration only, and can be replaced by weaker sufficient conditions.} By definition, $\mathcal{H}\subset L^{2}(f_{Y_{2}})$ and we let $||\cdot ||_{s}=||\cdot ||_{L^{2}(f_{Y_{2}})}.$ We assume that $ \mathcal{H}_{k(n)}=clsp\{q_{1},...,q_{k(n)}\}$ with $\{q_{k}\}_{k=1}^{\infty }$ being a Riesz basis of $(\mathcal{H},||\cdot ||_{s})$. The convergence rates of $\widehat{h}$ to $h_{0}$ in both $||\cdot ||$ and $||\cdot ||_{s}=||\cdot ||_{L^{2}(f_{Y_{2}})}$ metrics have already been established in CP_WP07, and hence will not be repeated here.

We use $\mathcal{H}_{os}$ and $\mathcal{H}_{osn}$ for $\mathcal{A}_{os}$ and $\mathcal{A}_{osn}$ defined in Subsection (ref) (since there is no $\theta $ here). Denote $T\equiv \frac{dm(\cdot ,h_{0})}{dh}:\mathcal{H }_{os}\subset L^{2}(f_{Y_{2}})\rightarrow L^{2}(f_{X})$, i.e., for any $h\in \mathcal{H}_{os}\subset L^{2}(f_{Y_{2}})$,

equation*[equation* omitted — 124 chars of source]

Let $T^{\ast }$ be the adjoint of $T$. Then for all $h\in \mathcal{H}_{os}$, we have $||h||^{2}\equiv ||Th||_{L^{2}(f_{X})}^{2}=||(T^{\ast }T)^{1/2}h||_{L^{2}(f_{Y_{2}})}^{2}$. Under mild conditions as stated in CP_WP07, $T$ and $T^{\ast }$ are compact. Then $T$ has a singular value decomposition $\{\mu _{k};\psi _{k},\phi _{0k}\}_{k=1}^{\infty }$, where $\{\mu _{k}>0\}_{k=1}^{\infty }$ is the sequence of singular values in non-increasing order ($\mu _{k}\geq \mu _{k+1}\geq ...$) with $ \liminf_{k\rightarrow \infty }\mu _{k}=0$, $\{\psi _{k}\in L^{2}(f_{Y_{2}})\}_{k=1}^{\infty }$ and $\{\phi _{0k}\in L^{2}(f_{X})\}_{k=1}^{\infty }$ are sequences of eigenfunctions of the operators $(T^{\ast }T)^{1/2}$ and $(TT^{\ast })^{1/2}$:

equation*[equation* omitted — 197 chars of source]

Since $\{q_{k}\}_{k=1}^{\infty }$ is a Riesz basis of $(\mathcal{H},||\cdot ||_{s})$ we could also have $\mathcal{H}_{k(n)}=clsp\{\psi _{1},...,\psi _{k(n)}\}$. The sieve measure of local ill-posedness now becomes $\tau _{n}=\mu _{k(n)}^{-1}$ (see, e.g., BCK_Emetrica07 and CP_WP07 ), and hence $\left\Vert u_{n}^{\ast }\right\Vert _{s}\leq c\mu _{k(n)}^{-1}$ for a finite constant $c>0$. Also, $\Pi _{n}h_{0}\equiv \arg \min_{h\in \mathcal{H}_{k(n)}}||h-h_{0}||_{s}=\sum_{k=1}^{k(n)}\langle h_{0},\psi _{k}\rangle _{s}\psi _{k}$ is the LS projection of $h_{0}$ onto the sieve space $\mathcal{H}_{n}$ under the strong norm $||\cdot ||_{s}=||\cdot ||_{L^{2}(f_{Y_{2}})}$. Recall that $h_{0,n}\equiv \arg \min_{h\in \mathcal{H }_{k(n)}}||h-h_{0}||^{2}\equiv \arg \min_{h\in \mathcal{H} _{k(n)}}||T[h-h_{0}]||_{L^{2}(f_{X})}^{2}$. We have:

eqnarray[eqnarray omitted — 332 chars of source]

The next remark specializes Theorem (ref) to a general functional $\phi (h)$ of the model ((ref)).

remarkLet $\widehat{m}$ be the series LS estimator ((ref)) for the model ((ref)) with $\widehat{\Sigma }=\Sigma =1$ , and Assumptions (ref)(i)(ii), (ref)(ii)(iii), and (ref) hold with $\delta _{n}=O\left( \sqrt{\frac{k(n)}{n}}\right) =o(n^{-1/4})$ and $\delta _{s,n}=O\left( \{k(n)\}^{-\varsigma }+\mu _{k(n)}^{-1}\sqrt{\frac{k(n)}{n}}\right) =o(1)$. Let Assumption (ref) , equation ((ref)) and Assumptions (ref) - (ref) hold. Then: \begin{equation} \sqrt{n}\frac{\phi (\widehat{h}_{n})-\phi (h_{0})}{||v_{n}^{\ast }||_{sd}} \Rightarrow N(0,1), \end{equation} with $||v_{n}^{\ast }||_{sd}^{2}=(\frac{ d\phi (h_{0})}{dh} [q^{k(n)}(\cdot )])^{\prime }D_{n}^{-1}{\mho }_{n}D_{n}^{-1}(\frac{d\phi (h_{0})}{dh} [q^{k(n)} (\cdot )])$, and $D_{n}=E\left[ \left( T[q^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\left( T[q^{k(n)}(\cdot )^{\prime }]\right) \right] $ and $\mho _{n}=E\left[ \left( T[q^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\rho (Z,h_{0})^{2}\left( T[q^{k(n)}(\cdot )^{\prime }]\right) \right] .$

Remark (ref) includes the NPIV\ and NPQIV examples in Subsection (ref) as special cases. In particular, the sieve variance expression ((ref)) reproduces the one for the NPIV model ((ref)) with $T[q^{k(n)}(\cdot )^{\prime }]=E[q^{k(n)}(Y_{2})^{\prime }|X]$, and the one for the NPQIV model ((ref)) with $T[q^{k(n)}(\cdot )^{\prime }]=E[f_{U|Y_{2},X}(0)q^{k(n)}(Y_{2})^{\prime }|X]$.

By the result in CP_WP07, the sieve dimension $k_{n}^{\ast }$ satisfying $\{k_{n}^{\ast }\}^{-\varsigma }\asymp \mu _{k_{n}^{\ast }}^{-1}\times \sqrt{\frac{k_{n}^{\ast }}{n}}$ leads to the nonparametric optimal convergence rate of $||\widehat{h}-h_{0}||_{s}=O_{P_{Z^{\infty }}}(\delta _{s,n}^{\ast })=o(1)$ in strong norm, where $\delta _{s,n}^{\ast }\asymp \{k_{n}^{\ast }\}^{-\varsigma }$. In particular, $k_{n}^{\ast }\asymp n^{\frac{1}{2(\varsigma +a)+1}}$ and $\delta _{s,n}^{\ast }=n^{- \frac{\varsigma }{2(\varsigma +a)+1}}$ for the mildly ill-posed case $\mu _{k}\asymp k^{-a}$ for a finite $a>0$; and $\delta _{s,n}^{\ast }=\{\ln n\}^{-\varsigma }$ for the severely ill-posed case $\mu _{k}\asymp \exp \{-0.5ak\}$ for a finite $a>0$. However this paper aims at simple valid inferences on functional $\phi (h_{0})$. As will be illustrated in the next subsection, although the nonparametric optimal choice $k_{n}^{\ast }$ is compatible with the sufficient conditions for the asymptotic normality of $ \sqrt{n}(\phi (\widehat{h})-\phi (h_{0}))$ for a regular linear functional $ \phi (h_{0})$ (see Remark (ref)), it is typically ruled out by Assumption (ref)(iii) for irregular functionals.

Verification of Assumption (ref)

Let $b_{j}\equiv \frac{d\phi (h_{0})}{dh}[\psi _{j}(\cdot )]$ for all $j$. By Lemma (ref) $D_{n}=E\left[ \left( T[q^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\left( T[q^{k(n)}(\cdot )^{\prime }]\right) \right] =Diag\left\{ \mu _{1}^{2},...,\mu _{k(n)}^{2}\right\} $ and

equation[equation omitted — 240 chars of source]

By Lemma (ref), $\phi (h)$ of the model ((ref)) is regular (at $h=h_{0}$) iff $\sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}<\infty $, and is irregular (at $h=h_{0}$) iff $ \sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}=\infty $.

For the same functional $\phi (h)$ of a model ((ref)) without endogeneity:

equation[equation omitted — 87 chars of source]

we have $D_{n}\asymp I_{k(n)}$ and $||v_{n}^{\ast }||^{2}\asymp \sum_{j=1}^{k(n)}b_{j}^{2}$. Thus, $\phi (h)$ of the model ((ref)) is regular (or irregular) iff $\sum_{j=1}^{\infty }b_{j}^{2}<\infty $ (or $ =\infty $).

Since $\mu _{k(n)}\rightarrow 0$ as $k(n)\rightarrow \infty $, if a functional $\phi (h)$ is irregular for the model ((ref)) without endogeneity, then it is irregular for the model ((ref)). But, even if a functional $\phi (h)$ is regular for the model ((ref)) without endogeneity, it could still be irregular for the model ((ref)) with endogeneity.

Linear functionals of the model ((ref))

For a linear functional $\phi (h)$ of the model ((ref)), given relation ((ref)), Assumption (ref) is satisfied provided that the sieve dimension $k(n)$ satisfies ((ref)):

equation[equation omitted — 208 chars of source]

When $\phi (h)$ of the model ((ref)) is regular, Remark (ref) implies that ((ref)) is satisfied provided

equation[equation omitted — 207 chars of source]

We shall illustrate below that both these sufficient conditions allow for severely ill-posed problems.

Example 1 (evaluation functional). For $\phi (h)=h(\overline{y} _{2}) $, we have: $||v_{n}^{\ast }||^{2}=\sum_{j=1}^{k(n)}\mu _{j}^{-2}[\psi _{j}(\overline{y}_{2})]^{2}$. Let $\mathcal{H}_{k(n)}$ be the spline or the CDV wavelet sieve as described in CC_WP13, say. Then

equation*[equation* omitted — 217 chars of source]

To provide concrete sufficient condition for ((ref)), we assume $ ||v_{n}^{\ast }||^{2}\asymp E\left( \sum_{j=1}^{k(n)}\mu _{j}^{-2}[\psi _{j}(Y_{2})]^{2}\right) =\sum_{k=1}^{k(n)}\mu _{k}^{-2}$. Since $ \lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||^{2}=\infty $, the evaluation functional is irregular. Condition ((ref)) is satisfied provided that

equation[equation omitted — 280 chars of source]

Condition ((ref)) allows for both mildly and severely ill-posed cases.

(a) Mildly ill-posed: $\mu _{k}\asymp k^{-a}$ for a finite $a>0$. Then $||v_{n}^{\ast }||^{2}\asymp \{k(n)\}^{2a+1}$. Condition ((ref) ) is satisfied by a wide range of sieve dimensions, such as $k(n)\asymp n^{ \frac{1}{2(\varsigma +a)+1}}(\ln \ln n)^{\varpi }$ or $n^{\frac{1}{ 2(\varsigma +a)+1}}(\ln n)^{\varpi }$ for any finite $\varpi >0$, or $ k(n)\asymp n^{\epsilon }$ for any $\epsilon \in (\frac{1}{2(\varsigma +a)+1}, \frac{1}{2a+1})$. Note that any $k(n)$ satisfying Condition ((ref)) also ensures $\delta _{s,n}=o(1)$. However, it does require $ k(n)/k_{n}^{\ast }\rightarrow \infty $, where $k_{n}^{\ast }\asymp n^{\frac{1 }{2(\varsigma +a)+1}}$ is the choice for the nonparametric optimal convergence rate in strong norm.

(b) Severely ill-posed: $\mu _{k}\asymp \exp \{-0.5ak\}$ for a finite $a>0$. Then $||v_{n}^{\ast }||^{2}\asymp \exp \{ak(n)\}$. Condition ( (ref)) is satisfied with $k(n)\asymp a^{-1}\left[ \ln n-\varpi \ln (\ln n)\right] $ for $0<\varpi <2\varsigma $. In addition we need $\varpi >1$ (and hence $\varsigma >1/2$) to ensure $\delta _{s,n}=O\left( \{k(n)\}^{-\varsigma }+\mu _{k(n)}^{-1}\sqrt{\frac{k(n)}{n}}\right) =o(1)$.

Example 2 (weighted derivative functional). For $\phi (h)=\int w(y)\nabla h(y)dy$, where $w(y)$ is a weight satisfying the integration by part formula: $\phi (h)=\int w(y)\nabla h(y)dy=-\int h(y)\nabla w(y)dy$, we have: $||v_{n}^{\ast }||^{2}=\sum_{j=1}^{k(n)}\mu _{j}^{-2}b_{j}^{2}$ with $ b_{j}=\int \psi _{j}(y)\nabla w(y)dy$ for all $j$, and

eqnarray*[eqnarray* omitted — 251 chars of source]

provided that $E\left( \left[ \frac{\nabla w(Y_{2})}{f_{Y_{2}}(Y_{2})}\right] ^{2}\right) =\sum_{j=1}^{\infty }b_{j}^{2}=C<\infty $. That is, the weighted derivative is assumed to be regular for the model ((ref)) without endogeneity.

(i) When the weighted derivative is regular (i.e., $ \sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}<\infty $) for the model ((ref)), Condition ((ref)) is satisfied provided that $ n\times \sum_{j=k(n)+1}^{\infty }\mu _{j}^{-2}b_{j}^{2}\times \delta _{n}^{2}=o(1)$, which is the condition imposed in AC_JOE07 for their root-$n$ estimation of an average derivative of NPIV example, and is shown to allow for severely ill-posed inverse case in AC_JOE07.

(ii) When the weighted derivative is irregular (i.e., $ \sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}=\infty $) for the model ((ref)), Condition ((ref)) is satisfied provided that

equation[equation omitted — 300 chars of source]

Condition ((ref)) allows for both mildly and severely ill-posed cases. To provide concrete sufficient conditions for ((ref)) we assume $b_{j}^{2}\asymp \left( j\ln (j)\right) ^{-1}$ in the following calculations.

(a) Mildly ill-posed: $\mu _{k}\asymp k^{-a}$ for a finite $a>0$. Then $||v_{n}^{\ast }||^{2}\in \lbrack c\frac{k(n)^{2a}}{\ln (k(n))} ,c^{\prime }k(n)^{2a}]$ for some $0<c\leq c^{\prime }<\infty $. Condition ( (ref)) and $\delta _{s,n}=o(1)$ are jointly satisfied by a wide range of sieve dimensions, such as $k(n)\asymp n^{\frac{1}{2(\varsigma +a)} }(\ln n)^{\varpi }$ for any finite $\varpi >\frac{1}{2(\varsigma +a)}$, or $ k(n)\asymp n^{\epsilon }$ for any $\epsilon \in (\frac{1}{2(\varsigma +a)}, \frac{1}{2a+1})$ and $\varsigma >1/2$.

(b) Severely ill-posed: $\mu _{k}\asymp \exp \{-0.5ak\}$ for $a>0$. Then $||v_{n}^{\ast }||^{2}\in \lbrack c\frac{\exp \{ak(n)\}}{k(n)\ln (k(n))} ,c^{\prime }\frac{\exp \{ak(n)\}}{\ln (k(n))}]$ for some $0<c\leq c^{\prime }<\infty $. Condition ((ref)) and $\delta _{s,n}=o(1)$ are jointly satisfied by $k(n)\asymp a^{-1}\left[ \ln (n)-\varpi \ln (\ln (n)) \right] $ for $\varpi \in (1,2\varsigma -1)$ and $\varsigma >1$.

Nonlinear functionals

For a nonlinear functional $\phi (h)$ of the model ((ref)), Assumption (ref) is satisfied provided that the sieve dimension $ k(n) $ satisfies ((ref)) (or ((ref)) if $\phi (h)$ is regular) and Assumption (ref)(ii), which is implied by the following condition:

Assumption (ref)(ii)': there are finite non-negative constants $C\geq 0,\omega _{1},\omega _{2}\geq 0$\ such that for all $(\alpha ,t)\in \mathcal{N}_{osn}\times \mathcal{T}_{n}$,

eqnarray*[eqnarray* omitted — 301 chars of source]

and

equation*[equation* omitted — 204 chars of source]

Assumption (ref)(ii) or (ii)' controls the nonlinearity bias of $ \phi \left( \cdot \right) $ (i.e., the linear approximation error of a nonlinear functional $\phi \left( \cdot \right) $). It typically rules out nonlinear regular functionals of severely illposed inverse problems, but allows for nonlinear irregular functionals of severely illposed inverse problems.

Example 3 (weighted quadratic functional). For $\phi (h)=\frac{1}{2} \int w(y)\left\vert h(y)\right\vert ^{2}dy$, we have $||v_{n}^{\ast }||^{2}=\sum_{j=1}^{k(n)}\mu _{j}^{-2}b_{j}^{2}$ with $b_{j}=\int h_{0}(y)w(y)\psi _{j}(y)dy$ for all $j$, and

equation*[equation* omitted — 215 chars of source]

provided that $\sup_{y}\frac{w(y)}{f_{Y_{2}}(y)}<\infty $. This and $E\left( \left[ h_{0}(Y_{2})\right] ^{2}\right) <\infty $ imply that $ \sum_{j=1}^{\infty }b_{j}^{2}<\infty $. That is, the weighted quadratic functional is regular for the model ((ref)) without endogeneity. Also,

equation*[equation* omitted — 212 chars of source]

(i) When the weighted quadratic functional is regular (i.e., $ \sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}<\infty $) for the model ((ref)), Condition ((ref)) is satisfied provided that $ n\times \sum_{j=k(n)+1}^{\infty }\mu _{j}^{-2}b_{j}^{2}\times \delta _{n}^{2}=o(1)$, which allows for severely ill-posed cases. But Assumption (ref)(ii)' requires that $\sqrt{n}\times \delta _{s,n}^{2}=\sqrt{n} \times \left( \{k(n)\}^{-\varsigma }+\mu _{k(n)}^{-1}\sqrt{\frac{k(n)}{n}} \right) ^{2}=o(1)$, which clearly rules out severely ill-posed inverse case where $\mu _{k}\asymp \exp \{-0.5ak\}$ for some finite $a>0$.

(ii) When the weighted quadratic functional is irregular (i.e., $ \sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}=\infty $) for the model ((ref)), Condition ((ref)) is satisfied provided that Condition ((ref)) holds with $b_{j}=\int h_{0}(y)w(y)\psi _{j}(y)dy$ for Example 3. Assumption (ref)(ii)' is satisfied provided that

equation[equation omitted — 304 chars of source]

Any $k(n)$ satisfying Conditions ((ref)) and ((ref)) automatically satisfies $\delta _{s,n}=o(1)$. In addition, both conditions allow for mildly and severely ill-posed cases. To provide concrete sufficient conditions we assume $b_{j}^{2}\asymp \left( j\ln (j)\right) ^{-1} $ in the following calculations.

(a) Mildly ill-posed: $\mu _{k}\asymp k^{-a}$ for a finite $a>0$. Then $||v_{n}^{\ast }||^{2}\in \lbrack c\frac{k(n)^{2a}}{\ln (k(n))} ,c^{\prime }k(n)^{2a}]$ for some $0<c\leq c^{\prime }<\infty $. Conditions ( (ref)) and ((ref)) are satisfied by a wide range of sieve dimensions, such as $k(n)\asymp n^{\frac{1}{2(\varsigma +a)}}(\ln n)^{\varpi }$ for any finite $\varpi >\frac{1}{2(\varsigma +a)}$, or $ k(n)\asymp n^{\epsilon }$ for any $\epsilon \in (\frac{1}{2(\varsigma +a)}, \frac{1}{2a+2})$ and $\varsigma >1$.

(b) Severely ill-posed: $\mu _{k}\asymp \exp \{-0.5ak\}$ for $a>0$. Then $||v_{n}^{\ast }||^{2}\in \lbrack c\frac{\exp \{ak(n)\}}{k(n)\ln (k(n))} ,c^{\prime }\frac{\exp \{ak(n)\}}{\ln (k(n))}]$ for some $0<c\leq c^{\prime }<\infty $. Conditions ((ref)) and ((ref)) are satisfied with $k(n)\asymp a^{-1}\left[ \ln (n)-\varpi \ln (\ln (n))\right] $ and $\varpi \in (3,2\varsigma -1)$ for $\varsigma >2$.

Verification of Assumption (ref)(i)

By Lemma (ref)(1), to verify Assumption (ref)(i), it suffices to verify Assumptions (ref) - (ref) in Appendix (ref). Note that Assumptions (ref) and (ref) do not depend on sieve Riesz representer at all, and have already been verified in CP_WP07a, AC_JOE07 and others for (penalized) SMD estimators for the model ((ref)). Assumptions (ref) and (ref) do depend on the scaled sieve Riesz representer $u_{n}^{\ast }\equiv v_{n}^{\ast }/||v_{n}^{\ast }||_{sd}$ . Both these assumptions are also verified in AC_Emetrica03, CP_WP07a, AC_JOE07 for examples of regular functionals of the model ((ref)). Here, we present simple (albeit somewhat strong) sufficient conditions for Assumptions (ref) and (ref) for irregular functionals of the NPIV and NPQIV examples.

condition(i) $\{E[h(Y_{2})|\cdot ]:h\in \mathcal{H} \}\subseteq \Lambda _{c}^{\gamma }(\mathcal{X})$, with $\gamma >0.5$; (ii) $ \sup_{x,y_{2}}\frac{f_{Y_{2}X}(y_{2},x)}{f_{Y_{2}}(y_{2})f_{X}(x)}\leq Const.<\infty $.
propositionLet all conditions for Remark (ref) hold. Under Condition (ref), Assumptions (ref) and (ref) hold for the NPIV model ((ref)).

Proposition (ref) allows for irregular functionals of the NPIV model with severely ill-posed case.

condition(i) $\{E[F_{Y_{1}|Y_{2}X}(h(Y_{2}),Y_{2},\cdot )|\cdot ]:h\in \mathcal{H}\}\subseteq \Lambda _{c}^{\gamma }(\mathcal{X})$, with $\gamma >0.5$; (ii) $\sup_{y_{1},y_{2},x}|\frac{ df_{Y_{1}|Y_{2}X}(y_{1},y_{2},x)}{dy_{1}}|\leq C<\infty $.
condition$n(\log \log n)^{4}\delta _{s,n}^{4}=o(1)$
propositionLet all conditions for Remark (ref) hold. Under conditions (ref)(ii) and (ref)-(ref), Assumptions (ref) and (ref) hold for the NPQIV model ((ref)).

It is clear that Condition (ref) rules out severely ill-posed case, and hence Proposition (ref) only allows for irregular functionals of the NPQIV model with mildly ill-posed case.

Simulation Studies and An Empirical Illustration

This section first presents simulation studies for SQLR and sieve t tests of linear and nonlinear hypotheses for the NPQIV and NPIV models respectively. It then provides an empirical illustration of the optimally weighted SQLR inferences for a NPQIV Engel curve. In this section, we use the series LS estimator ((ref)) of $m(x,h)$ with $p^{J_{n}}(x)$ as its basis, and $ q^{k(n)}$ as the basis approximating the unknown structure function $h_{0}$. We use $p^{J}=\mathrm{P-Spline}(r,k)$ to denote $r$th degree polynomial spline with $k$ (quantile) equally spaced knots, hence $J=(r+1)+k$ is the total number of sieve terms. We use $p^{J}=\mathrm{Pol}(J)$ to denote power series up to $(J-1)$th degree. See, e.g., C_bookchp07 for definitions of these and other sieve bases.

Simulation Studies

We run Monte Carlo (MC) studies to assess the finite sample performance of SQLR and sieve t tests of linear and nonlinear hypotheses in two models: the NPQIV ((ref)) and the NPIV ((ref)).

For all cases, our design is based on the MC design of NP_ECMA03 and Santos_ECMA for a NPIV model, which we adapt to cover both NPIV and NPQIV models. Specifically, we generate i.i.d. draws of $(Y_{2},X,U^{\ast })$ from

equation*[equation* omitted — 200 chars of source]

and $Y_{2}=2(\Phi (Y_{2}^{\ast }/3)-0.5)$ and $X=2(\Phi (X^{\ast }/3)-0.5)$. The true function $h_{0}$ is given by $h_{0}(\cdot )=2\sin (\pi \cdot )$. We consider 5,000 MC repetitions and $n=750$ for each of the cases studied below. We use $Pen(h)=||h||_{L^{2}}^{2}+||\nabla h||_{L^{2}}^{2}$ in all the simulations, and have used a very small $\lambda _{n}=10^{-5}$ in most cases (except for the cases where we study the sensitivity to the choice of $\lambda _{n} $).

Summary of sensitivity checks: For NPQIV and NPIV models, for both SQLR and sieve t tests of linear and nonlinear hypotheses, as long as $ J_{n}>k(n)+1$ with not too large $k(n)$, the MC sizes of the tests are good and insensitive to the choices of basis $q^{k(n)}$ and $p^{J_{n}}$ or the very small penalty $\lambda _{n}$. This is consistent with previous MC findings in BCK_Emetrica07 and CP_WP07 for PSMD estimation of NPIV and NPQIV respectively.

NPQIV model: SQLR test for an irregular linear functional. We consider the NPQIV model $Y_{1}=h_{0}(Y_{2})+U=2\sin (\pi Y_{2})+U$ with $ U=2(\Phi (U^{\ast })-\gamma )$. This last transformation is done to ensure that $E[1\{U\leq 0\}|X]=\gamma $. To save space we only present the case with $\gamma =0.5$. The parameter of interest is $\phi (h_{0})=h_{0}(0)$, hence $\phi $ is an irregular linear functional. We study the finite sample properties of the SQLR and bootstrap-SQLR tests. The SQLR-based confidence intervals are specially well-suited for models like NPQIV where the generalized residual function is non-smooth yet the optimal weighting matrix is easy to compute.

Size. Table (ref) reports the simulated size of the SQLR test of $H_{0}\colon \phi (h_{0})=0$ as a function of the nominal size (NS), for different choices of $q^{k(n)}$ and $p^{J_{n}}$, and different values of the tuning parameters $(\lambda _{n},k(n),J_{n})$.

table[table omitted — 1,582 chars of source]

Table (ref) shows that for small value of $k(n)$, say in $(k(n),J_{n})=(4,7)$ (i.e., rows 1-3), the SQLR test performs well and is fairly insensitive to different choices of $\lambda _{n}$. For a fixed relatively small $J_{n}=7$, rows 1-6 indicate that as $k(n)$ increases, the results become a bit more sensitive to the choice of $\lambda _{n}$. For a fixed very small penalty $\lambda _{n}=10^{-5}$, rows 7-16 show that the results are fairly insensitive to different choices of $J_{n}$ and basis for $p^{J_{n}}$ and $q^{k(n)}$ as long as $J_{n}>k(n)+1$.

Local power. Figure (ref) shows the rejection probabilities at 5% (lower panel) and 1% (upper panel) level of the null hypothesis as a function of $r$ where $r\colon \phi (h_{0})=r$ for the SQLR (solid red line) and the bootstrap SQLR (dashed blue line) with multinomial weights. To save space we only report the local power results corresponding to the case of P-Spline(3,2) for $q^{k(n)}$, Pol(10) for $p^{J_{n}} $ and $\lambda _{n}= 10^{-5}$ in Table (ref). We employ 500 bootstrap evaluations per MC replication, and lower the number of MC repetitions to 1,000 to ease the computational burden. We note that since our functional $\phi (h)=h(0)$ is estimated at a slower than root-$n$ rate, the deviations considered for $r$ which are in the range of $[0,8/\sqrt{n}]$ are indeed \textquotedblleft small\textquotedblright. We can see from the figure that the bootstrap SQLR performance is similar to its non-bootstrapped counterpart. We expect that the performance will improve if we increase number of bootstrap runs. (We also run simulation studies corresponding to the case of Pol(4) for $q^{k(n)}$, Pol(7) for $p^{J_{n}} $ and $\lambda _{n}=2\times 10^{-4}$ in Table (ref), and the local power patterns are similar to the ones reported here.)

figure[figure omitted — 434 chars of source]

NPIV model: sieve variance estimators for an irregular linear functional. We now consider the NPIV model: $ Y_{1}=h_{0}(Y_{2})+0.76U=2\sin (\pi Y_{2})+0.76U$, with $U=U^{\ast }$ so the identifying condition of NPIV holds: $E[U|X]=0$. The parameter of interest is $\phi (h_{0})=h_{0}(0)$, and the null hypothesis is $H_{0}\colon \phi (h_{0})=0$. We focus on the finite sample performance of the sieve variance estimators for irregular linear functionals. We compute two sieve variance estimators:

equation*[equation* omitted — 259 chars of source]

where $\widehat{D}_{n}=n^{-1}\left( \widehat{C}_{n}(P^{\prime }P)^{-} \widehat{C}_{n}^{\prime }\right) $, $\widehat{C}_{n}\equiv \sum_{i=1}^{n}q^{k(n)}(Y_{2i})p^{J_{n}}(X_{i})^{\prime }$, $\widehat{\mho } _{n}$ is given in equation ((ref)), and $\widehat{\Omega }_{n}= \frac{1}{n}\widehat{C}_{n}(P^{\prime }P)^{-}\left( \sum_{i=1}^{n}p^{J_{n}}(X_{i})\widehat{\Sigma }_{0}(X_{i})p^{J_{n}}(X_{i})^{ \prime }\right) (P^{\prime }P)^{-}\widehat{C}_{n}^{\prime }$ with $\widehat{U }_{j}=Y_{1j}-\widehat{h}(Y_{2j})$ and $\widehat{\Sigma }_{0}(x)=\left( \sum_{j=1}^{n}\widehat{U}_{j}^{2}p^{J_{n}}(X_{j})^{\prime }\right) (P^{\prime }P)^{-}p^{J_{n}}(x)$. (See Theorem (ref) in Appendix (ref) for the definition and consistency of $\widehat{V}_{2}$ as another sieve variance estimator for any plug-in PSMD $\phi (\widehat{\alpha })$.)

table[table omitted — 2,217 chars of source]

Table (ref) reports the results for different choices of bases for $ q^{k(n)}$ and $p^{J_{n}}$, and for different values of $k(n)$ and $J_{n}$; in all cases we use a very small $\lambda _{n}=10^{-5}$. This table shows $ Med_{MC}\left[ \left\vert \frac{\widehat{V}_{j}}{||v_{n}^{\ast }||_{sd}^{2}} -1\right\vert \right] $ for $j=1,2$, where $||v_{n}^{\ast }||_{sd}$ is computed using the MC variance of $\sqrt{n}\widehat{h}_{n}(0)$ and $ Med_{MC}[\cdot ]$ is the MC median. It also shows the nominal size and MC rejection frequencies of the two sieve t tests $\widehat{t}_{j}=\sqrt{n} \frac{\widehat{h}_{n}(0)-0}{\sqrt{\widehat{V}_{j}}}$ for $j=1,2$.

We note that the two sieve variance estimators have almost identical performance and the associated sieve t tests have good rejection probabilities. These results are fairly robust to different choices of basis for $q^{k(n)}$ and $p^{J_{n}}$ and different values of $k(n)$ and $J_{n}$ as long as $J_{n}>k(n)+1$. Figure (ref) (first row) shows the QQ-Plot for the sieve t tests $\widehat{t}_{j}=\sqrt{n}\frac{\widehat{h} _{n}(0)-0}{\sqrt{\widehat{V}_{j}}}$ under the null for $j=1,2$ for the case Pol(4)-Pol(16) in the table; the right panel in the first row corresponds to $\hat{t}_{1}$ and the left panel in the first row to $\hat{t}_{2}$. Both sieve t tests are almost identical to each other and to the standard normal.

figure[figure omitted — 313 chars of source]

NPIV model: sieve variance estimators for an irregular nonlinear functional. This case is identical to the previous one for the NPIV model, except that the functional of interest is $\phi (h_{0})=\exp \{h_{0}(0)\}$, and the null hypothesis is $H_{0}\colon \phi (h_{0})=1$. This choice of $\phi $ allows us to evaluate the finite sample performance of sieve t statistics for a nonlinear functional.

table[table omitted — 2,225 chars of source]

Table (ref) shows $Med_{MC}$ and rejection probabilities for this nonlinear case. By comparing the results with those in Table (ref) we note that the results are very similar in both cases. Figure (ref) (second row) shows the QQ-Plot for the two sieve t tests for the non-linear case; the right panel in the second row corresponds to $\hat{t}_{1}$ whereas the left panel in the second row corresponds to $ \hat{t}_{2}$. These results suggest that our sieve t tests perform equally well for both functionals.

Finally we wish to point out that we have tried other bases such as Hermite polynomials and cosine series and even larger $J_{n}$ in these two NPIV MC studies, the results are all similar to the ones reported here and hence are not presented due to the lack of space.

An Empirical Application

We compute SQLR based confidence bands for nonparametric quantile IV Engel curves using the British FES data set from BCK_Emetrica07:

equation*[equation* omitted — 66 chars of source]

where $Y_{1,i}$ is the budget share of the $i-$th household on a particular non-durable goods, say food-in consumption; $Y_{2,i}$ is the log-total expenditure of the household, which is endogenous, and hence we use $X_{i}$, the gross earnings of the head of the household, to instrument it. We work with the \textquotedblleft no kids\textquotedblright\ sub-sample of the data set, which consists of $n=628$ observations. BCK_Emetrica07 estimated NPIV Engel curves using this data set. But, as explained by Koenker_2005 and others, quantile Engel curves are more informative.

We estimate $h_{0}(\cdot )$ for food-in quantile Engel curve via the optimally weighted PSMD procedure with $\widehat{\Sigma }=\Sigma _{0}=0.25$, using a polynomial spline (P-spline) sieve $\mathcal{H}_{k(n)}$ with $k(n)=4$ , $Pen(h)=||h||_{L^{2}}^{2}+||\nabla h||_{L^{2}}^{2}$ with $\lambda _{n}=0.0005$, and a Hermite polynomial LS basis $p^{J_{n}}(X)$ with $J_{n}=6$ . We also considered other bases such as P-splines as $p^{J_{n}}(X)$ and results remained essentially the same. See CP_WP07a for PSMD estimates of NPQIV Engel curves for other non-durable goods.

We use the fact that the optimally weighted SQLR of testing $\phi (h)=h(y_{2})$ (for any fixed $y_{2}$) is asymptotically $\chi _{1}^{2}$ to construct pointwise confidence bands. That is, for each $y_{2}$ in the sample we construct a grid, $(r_{i} )_{i=1}^{30}$. For each $i={1,...,30}$, we compute the value of the SQLR test statistic under $h(y_{2})=r_{i}$ for $ (r_{i})_{i=1}^{30}$. We then, take the smallest interval that included all points $r_{i}$ that yield a corresponding value of the SQLR test below the 95% percentile of $\chi _{1}^{2}$.\footnote{ The grid $(r_{i})_{i=1}^{30}$ was constructed to have $r_{15}=\widehat{h} _{n}(y_{2})$, for all $i\leq 15$, $r_{i+1}\leq r_{i}\leq r_{15}$ decreasing in steps of length $0.002$ (approx) and for all $i\geq 15$, $r_{i+1}\geq r_{i}\geq r_{15}$ increasing in steps of length $0.008$ (approx); finally, the extremes, $r_{1}$ and $r_{30}$, were chosen so the SQLR test at those points was above the 95% percentile of $\chi _{1}^{2}$. We tried different lengths and step sizes and the results remain qualitatively unchanged. For some observations, which only account for less than 4% of the sample, the confidence interval was degenerate at a point; this result was due to numerical approximation issues, and these observations were excluded from the reported results.} Figure (ref) presents the results, where the solid blue line is the point estimate and the red dashed lines are the 95% pointwise confidence bands. We can see that the confidence bands get wider towards the extremes of the sample, but are tighter in the middle.

To test whether the quantile IV Engel curve for food-in is linear or not, one can test whether $\phi (h_{0})\equiv \int \left\vert \nabla ^{2}h(y_{2})\right\vert ^{2}w(y_{2})dy_{2}=0$ using our SQLR test. Let $ w(\cdot )=(\sigma _{Y_{2}})^{-1}\exp \left( -\frac{1}{2}(\sigma _{Y_{2}}^{-1}(\cdot -\mu _{Y_{2}}))^{2}\right) 1\{t_{0.01}\leq \cdot \leq t_{0.99}\}$ where $\mu _{Y_{2}}$, $\sigma _{Y_{2}}$, $t_{0.01}$ and $ t_{0.99} $ are the sample mean, standard deviation and the 1% and 99% quantiles of $Y_{2}$. The value of the SQLR is (approx.) 38 and the p-value is smaller than 0.0001, and hence we reject the null hypothesis of linearity. \footnote{ We use the standard Riemann sum with 1000 terms to compute the integral. We also considered other choices of $w$ such that $w(\cdot )=1\{t_{0.25}\leq \cdot \leq t_{0.75}\}$ and $w(\cdot )=1\{t_{0.01}\leq \cdot \leq t_{0.99}\}$ . Although the numerical value of the SQLR test changes, all produce p-values below 0.0001.}

figure[figure omitted — 261 chars of source]

Conclusion

In this paper, we provide unified asymptotic theories for PSMD based inferences on possibly irregular parameters $\phi (\alpha _{0})$ of the general semi/nonparametric conditional moment restrictions $E[\rho (Y,X;\alpha _{0})|X]=0$. Under regularity conditions that allow for any consistent nonparametric estimator of the conditional mean function $ m(X,\alpha )\equiv E[\rho (Y,X;\alpha )|X]$, we establish the asymptotic normality of the plug-in PSMD estimator $\phi (\widehat{\alpha }_{n})$ of $ \phi (\alpha _{0})$, as well as the asymptotically tight distribution of a possibly non-optimally weighted SQLR statistic under the null hypothesis of $ \phi (\alpha _{0})=\phi _{0}$. As a simple yet useful by-product, we immediately obtain that an optimally weighted SQLR statistic is asymptotically chi-square distributed under the null hypothesis. For (pointwise) smooth residuals $\rho (Z;\alpha )$ (in $\alpha $), we propose several simple consistent sieve variance estimators for $\phi (\widehat{ \alpha }_{n})$ (in the text and in online Appendix (ref)), and establish the asymptotic chi-square distribution of sieve Wald statistics. We also establish local power properties of SQLR and sieve Wald tests in Appendix (ref). Under conditions that are virtually the same as those for the limiting distributions of the original-sample sieve Wald and SQLR statistics, we establish the consistency of the generalized residual bootstrap sieve Wald and SQLR statistics. All these results are valid regardless of whether $\phi (\alpha _{0})$ is regular or not. While SQLR and bootstrap SQLR are useful for models with (pointwise) non-smooth $\rho (Z;\alpha )$, sieve Wald statistic is computationally attractive for models with smooth $\rho (Z;\alpha )$. Monte Carlo studies and an empirical illustration of a nonparametric quantile IV regression demonstrate the good finite sample performance of our inference procedures.

This paper assumes that the semi/nonparametric conditional moment restrictions $E[\rho (Y,X;\alpha _{0})|X]=0$ uniquely identifies the unknown true parameter value $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$, and conduct inferences that are robust to whether a possibly nonlinear functional of $\alpha _{0}$ is root-$n$ estimable or not. Recently, for the NPIV model $E[Y_{1}-h_{0}(Y_{2})|X]=0$ without assuming point identification of $h_{0}$, Santos_JOE proposed a root-$n$ asymptotically normal estimation of a regular linear functional of $h_{0}$ and Santos_ECMA considered Bierens' type test of the NPIV. CPT_WP12 is extending the SQLR inference procedure to allow for partial identification of the general model $E[\rho (Y,X;\alpha _{0})|X]=0$.

\baselineskip=15pt

\setstretch{1}