The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
179,026 characters
Sieve Wald and QLR Inferences on Semi/nonparametric Conditional Moment Models
\begin{frontmatter}
\title{Sieve Wald and QLR Inferences on Semi/nonparametric Conditional
Moment Models}
\runtitle{Sieve Wald and QLR Inference}
\author{Xiaohong Chen and Demian Pouzo \thanksref{r1}}
\thankstext{r1}{
Earlier versions, some entitled \textquotedblleft On PSMD inference of
functionals of nonparametric conditional moment
restrictions\textquotedblright ,\ were presented in April 2009 at the Banff
conference on seminonparametrics, in June 2009 at the Cemmap conference on
quantile regression, in July 2009 at the SITE conference on nonparametrics,
in September 2009 at the Stats in the Chateau/France, in June 2010 at the
Cemmep workshop on recent developments in nonparametric instrumental
variable methods, in August 2010 at the Beijing international conference on
statistics and society, and econometric workshops in numerous universities.
We thank a co-editor, two referees, Don Andrews, Peter Bickel, Gary
Chamberlain, Tim Christensen, Michael Jansson, Jim Powell and especially
Andres Santos for helpful comments. We thank Yinjia Qiu for excellent
research assistant in simulations using R. Chen acknowledges financial
support from National Science Foundation grant SES-0838161 and Cowles
Foundation. Any errors are the responsibility of the authors.}
\address{Chen: Cowles Foundation for Research in Economics, Yale University, Box 208281, New Haven, CT 06520, USA. Email: [email removed]. Pouzo: Department of Economics, UC Berkeley, 530 Evans Hall 3880, Berkeley, CA 94720, USA. Email: [email removed].}
\runauthor{X. Chen and D. Pouzo}
\begin{abstract}
This paper considers inference on functionals of semi/nonparametric
conditional moment restrictions with possibly nonsmooth generalized
residuals, which include all of the (nonlinear) nonparametric instrumental
variables (IV) as special cases. These models are often ill-posed and hence
it is difficult to verify whether a (possibly nonlinear) functional is
root-$n$ estimable or not. We provide computationally simple, unified inference
procedures that are asymptotically valid regardless of whether a functional
is root-$n$ estimable or not. We establish the following new useful results:
(1) the asymptotic normality of a plug-in penalized sieve minimum distance
(PSMD) estimator of a (possibly nonlinear) functional; (2) the consistency
of simple sieve variance estimators for the plug-in PSMD estimator, and
hence the asymptotic chi-square distribution of the sieve Wald statistic;
(3) the asymptotic chi-square distribution of an optimally weighted sieve
quasi likelihood ratio (QLR) test under the null hypothesis; (4) the
asymptotic tight distribution of a non-optimally weighted sieve QLR
statistic under the null; (5) the consistency of generalized residual
bootstrap sieve Wald and QLR tests; (6) local power properties of sieve Wald
and QLR tests and of their bootstrap versions; (7) asymptotic properties of
sieve Wald and SQLR for functionals of increasing dimension. Simulation
studies and an empirical illustration of a nonparametric quantile IV
regression are presented.
\end{abstract}
\begin{keyword}
Nonlinear nonparametric instrumental variables;
Penalized sieve minimum distance; Irregular functional; Sieve variance
estimators; Sieve Wald; Sieve quasi likelihood ratio; Generalized residual
bootstrap; Local power; Wilks phenomenon.
\end{keyword}
\end{frontmatter}
\setcounter{page}{0} \thispagestyle{empty} \newpage
\baselineskip=18pt
\section{Introduction}
This paper is about inference on functionals of the unknown true parameters $
\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$ satisfying the
semi/nonparametric conditional moment restrictions
\begin{equation}
E[\rho (Y,X;\theta _{0},h_{0})|X]=0\quad a.s.-X, \label{semi00}
\end{equation}
\noindent where $Y$ is a vector of endogenous variables and $X$ is a vector
of conditioning (or instrumental) variables. The conditional distribution of
$Y$ given $X$, $F_{Y|X}$, is not specified beyond that it satisfies (\ref
{semi00}). $\rho (\cdot ;\theta _{0},h_{0})$ is a $d_{\rho }\times 1-$vector
of generalized residual functions whose functional forms are known up to the
unknown parameters $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})\in
\Theta \times \mathcal{H}$, with $\theta _{0}\equiv (\theta _{01},...,\theta
_{0d_{\theta }})^{\prime }\in \Theta $ being a $d_{\theta }\times 1-$vector
of finite dimensional parameters and $h_{0}\equiv (h_{01}(\cdot
),...,h_{0q}(\cdot ))\in \mathcal{H}$ being a $1\times d_{q}-$vector valued
function. The arguments of each unknown function $h_{\ell }(\cdot )$ may
differ across $\ell =1,...,q$, may depend on $\theta ,$ $h_{\ell ^{\prime
}}(\cdot ),$ $\ell ^{\prime }\neq \ell $, $X$ and $Y$. The residual function
$\rho (\cdot ;\alpha )$ could be nonlinear and pointwise non-smooth in the
parameters $\alpha \equiv (\theta ^{\prime },h)\in \Theta \times \mathcal{H}$
.
The general framework (\ref{semi00}) nests many widely used nonparametric
and semiparametric models in economics and finance. Well known examples
include nonparametric mean instrumental variables regressions (NPIV): $
E[Y_{1}-h_{0}(Y_{2})|X]=0$ (e.g., \cite{HH_Ann05}, \cite{CFR_bookchp07},
\cite{BCK_Emetrica07}, \cite{DFR_wp10}, \cite{Horowitz_ECMA11});
nonparametric quantile instrumental variables regressions (NPQIV): $
E[1\{Y_{1}\leq h_{0}(Y_{2})\}-\gamma |X]=0$ (e.g., \cite{CH_Emetrica05},
\cite{CIN_JOE07}, \cite{HL_Emetrica07}, \cite{CP_WP07}, \cite{CGS_WP08});
semi/nonparametric demand models with endogeneity (e.g., \cite
{BCK_Emetrica07}, \cite{CP_WP07a}, \cite{Souza2012}); semi/nonparametric
random coefficient panel data regressions (e.g., \cite{CHAMBERLAIN_ECMA92},
\cite{GPowell2012}); semi/nonparametric spatial models with endogeneity
(e.g., \cite{Pinkse2002}, \cite{MerloPaula}); semi/nonparametric asset
pricing models (e.g., \cite{Hansen_Richard_ECMA87}, \cite
{Gallant_Tauchen_ECMA89}, \cite{ChenLudvigson_JAE09}, \cite
{ChenLudvigson2013}, \cite{Sentana2013}); semi/nonparametric static and
dynamic game models (e.g., \cite{BHN_WP11}); nonparametric optimal
endogenous contract models (e.g., \cite{BMT_WP12}). Additional examples of
the general model (\ref{semi00}) can be found in \cite{CHAMBERLAIN_ECMA92},
\cite{NP_ECMA03}, \cite{AC_Emetrica03}, \cite{CP_WP07}, \cite{CCLN_WP10} and
the references therein. In fact, model (\ref{semi00}) includes all of the
(nonlinear) semi/nonparametric IV regressions when the unknown functions $
h_{0}$ depend on the endogenous variables $Y$:
\begin{equation}
E[\rho (Y_{1};\theta _{0},h_{0}(Y_{2}))|X]=0\quad a.s.-X, \label{gnpiv}
\end{equation}
which could lead to difficult (nonlinear) nonparametric ill-posed inverse
problems with unknown operators.
Let $\left\{ Z_{i}\equiv (Y_{i}^{\prime },X_{i}^{\prime })^{\prime }\right\}
_{i=1}^{n}$ be a random sample from the distribution of $Z\equiv (Y^{\prime
},X^{\prime })^{\prime }$ that satisfies the conditional moment restrictions
(\ref{semi00}) with a unique $\alpha _{0}\equiv (\theta _{0}^{\prime
},h_{0}) $. Let $\phi :\Theta \times \mathcal{H}\rightarrow \mathbb{R}
^{d_{\phi }}$ be a (possibly nonlinear) functional with a finite $d_{\phi
}\geq 1$. Typical linear functionals include an Euclidean functional $\phi
(\alpha )=\theta $, a point evaluation functional $\phi (\alpha )=h(
\overline{y}_{2}) $ (for $\overline{y}_{2}\in $ supp$(Y_{2})$), a weighted
derivative functional $\phi (h)=\int w(y_{2})\nabla h(y_{2})dy_{2}$ and many
others. Typical nonlinear functionals include a quadratic functional $\int
w(y_{2})\left\vert h(y_{2})\right\vert ^{2}dy_{2}$, a quadratic derivative
functional $\int w(y_{2})\left\vert \nabla h(y_{2})\right\vert ^{2}dy_{2}$,
a consumer surplus or an average consumer surplus functional of an
endogenous demand function $h$. We are interested in computationally simple,
valid inferences on any $\phi (\alpha _{0})$ of the general model (\ref
{semi00}) with i.i.d. data.\footnote{
See our Cowles Foundation Discussion Paper No. 1897 for general theory
allowing for weakly dependent data.}
Although some functionals of the model (\ref{semi00}), such as the (point)
evaluation functional, are known \textit{a priori} to be estimated at slower
than root-$n$ rates, others, such as the weighted derivative functional, are
far less clear without a stare at their semiparametric efficiency bound
expressions. This is because a non-singular efficiency bound is a necessary
condition for $\phi (\alpha _{0})$ to be estimated at a root-$n$ rate.
Unfortunately, as pointed out in \cite{CHAMBERLAIN_ECMA92} and \cite{AC_WP05}
, there is generally no closed form solution for the efficiency bound of $
\phi (\alpha _{0})$ (including $\theta _{0}$) of model (\ref{semi00}),
especially so when $\rho (\cdot ;\theta _{0},h_{0})$ contains several
unknown functions and/or when the unknown functions $h_{0}$ of endogenous
variables enter $\rho (\cdot ;\theta _{0},h_{0})$ nonlinearly. It is thus
difficult to verify whether the efficiency bound for $\phi (\alpha _{0})$ is
singular or not. Therefore, it is highly desirable for applied researchers
to be able to conduct simple valid inferences on $\phi (\alpha _{0})$
regardless of whether it is root-$n$ estimable or not. This is the main goal
of our paper.
In this paper, for the general model (\ref{semi00}) that could be
nonlinearly ill-posed and for any $\phi (\alpha _{0})$ that may or may not
be root-$n$ estimable, we first establish the asymptotic normality of the
plug-in penalized sieve minimum distance (PSMD) estimator $\phi (\widehat{
\alpha }_{n})$ of $\phi (\alpha _{0})$. For the model (\ref{semi00}) with
(pointwise) smooth residuals $\rho (Z;\alpha )$ in $\alpha _{0}$, we propose
two simple consistent sieve variance estimators for possibly slower than
root-$n$ estimator $\phi (\widehat{\alpha }_{n})$, which immediately leads
to the asymptotic chi-square distribution of the sieve Wald statistic.
However, there is no simple variance estimator for $\phi (\widehat{\alpha }
_{n})$ when $\rho (Z,\alpha )$ is not pointwise smooth in $\alpha _{0}$
(without estimating an extra unknown nuisance function or using numerical
derivatives). We then consider a PSMD criterion based test of the null
hypothesis $\phi (\alpha _{0})=\phi _{0}$. We show that an optimally
weighted sieve quasi likelihood ratio (SQLR) statistic is asymptotically
chi-square distributed under the null hypothesis. This allows us to
construct confidence sets for $\phi (\alpha _{0})$ by inverting the
optimally weighted SQLR statistic, without the need to compute a variance
estimator for $\phi (\widehat{\alpha }_{n})$. Nevertheless, in complicated
real data analysis applied researchers might like to use simple but possibly
non-optimally weighed PSMD procedures for estimation of and inference on $
\phi (\alpha _{0})$. We show that the non-optimally weighted SQLR statistic
still has a tight limiting distribution under the null regardless of whether
$\phi (\alpha _{0})$ is root-$n$ estimable or not. In addition, we establish
the consistency of the generalized residual bootstrap (possibly
non-optimally weighted) SQLR and sieve Wald tests under virtually the same
conditions as those used to derive the limiting distributions of the
original-sample statistics. The bootstrap SQLR would then lead to
alternative confidence sets construction for $\phi (\alpha _{0})$ without
the need to compute a variance estimator for $\phi (\widehat{\alpha }_{n})$.
To ease notation burden, we present the above listed theoretical results for
a scalar-valued functional in the main text. In Appendix \ref{app:appA} we
present the asymptotic properties of sieve Wald and SQLR for functionals of
increasing dimension (i.e., $d_{\phi }=dim(\phi )$ could grow with sample
size $n$). We also provide the local power properties of sieve Wald and SQLR
tests as well as their bootstrap versions in Appendix \ref{app:appA}.
Regardless of whether a possibly nonlinear functional $\phi (\alpha _{0})$
is root-$n$ estimable or not, we show that the optimally weighted SQLR is
more powerful than the non-optimally weighed SQLR, and that the SQLR and the
sieve Wald using the same weighting matrix have the same local power in
terms of first order asymptotic theory.
To the best of our knowledge, our paper is the first to provide a unified
theory about sieve Wald and SQLR inferences on (possibly nonlinear) $\phi
(\alpha _{0})$ satisfying the general semi/nonparametric model (\ref{semi00}
) with possibly non-smooth residuals.\footnote{
We also provide asymptotic properties of sieve score and bootstrap sieve
score statistics in the online Appendix \ref{app:appD}.} Our results allow applied researchers to obtain
limiting distribution of the plug-in PSMD estimator $\phi (\widehat{\alpha }
_{n})$ and to construct confidence sets for any $\phi (\alpha _{0})$
regardless of whether it is root-$n$ estimable or not. Our paper is also the
first to provide local power properties of sieve Wald and SQLR tests and
their bootstrap versions of general nonlinear hypotheses for the model (\ref
{semi00}).
Roughly speaking, our results extend the classical theories on Wald and QLR
tests of nonlinear hypothesis based on root-$n$ consistent parametric
minimum distance estimator $\widehat{\alpha }_{n}$ to those based on slower
than root-$n$ consistent nonparametric minimum distance estimator $\widehat{
\alpha }_{n}\equiv (\widehat{\theta }_{n}^{\prime },\widehat{h}_{n})$ of $
\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$ satisfying the model (\ref
{semi00}). The implementations of the sieve Wald and SQLR also resemble the
classical Wald and QLR based on parametric extreme estimators and hence are
computationally attractive. For example, our sieve t (Wald) test on a
general nonlinear hypothesis $\phi (h_{0})=\phi _{0}$ of the NPIV model $
E[Y_{1}-h_{0}(Y_{2})|X]=0$ can be implemented as a standard t (Wald) test
for a parametric linear IV model using two stage least squares (see
Subsection \ref{sec:NPIVex}). The proof techniques are quite different,
however, because one is no longer able to rely on the root-$n$ asymptotic
normality of $\widehat{\alpha }_{n}$ and then a standard \textquotedblleft
delta-method\textquotedblright\ to establish the asymptotic normality of $
\sqrt{n}\left( \phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})\right) $. In
our framework (\ref{gnpiv}), $\sqrt{n}\left( \phi (\widehat{\alpha }
_{n})-\phi (\alpha _{0})\right) $ could diverge to infinity under the
combined effects of (i) slower convergence rate of $\widehat{\alpha }_{n}$
to $\alpha _{0}$ due to the ill-posed inverse problem and (ii) nonlinearity
of either the functional $\phi ()$ or the residual function $\rho ()$ in $h$
. Our proof strategy relies on the convergence rates of the PSMD estimator $
\widehat{\alpha }_{n}$ to $\alpha _{0}$ in both weak and strong metrics, and
then the local curvatures of the functional $\phi ()$ and the criterion
function under these two metrics. The weak metric is intrinsic to the
variance of the linear approximation to $\phi (\widehat{\alpha }_{n})-\phi
(\alpha _{0})$, while the strong metric controls the nonlinearity (in $
\alpha $) of the functional $\phi ()$ and of the conditional mean function $
m(\cdot ,\alpha )=E[\rho (Y,X;\alpha )|X=\cdot ]$. Unfortunately the
convergence rate in the strong metric could be very slow due to the illposed
inverse problem. This explains why it is difficult to establish the
asymptotic normality of $\phi (\widehat{\alpha }_{n})$ for a nonlinear
functional $\phi ()$ even in the NPIV model. Our paper builds upon the
recent results on convergence rates in \cite{CP_WP07} and others. In
particular, under virtually the same conditions as those in \cite{CP_WP07},
we show that our generalized residual bootstrap PSMD estimator of $\alpha
_{0}$ is consistent and achieves the same convergence rates as that of the
original-sample PSMD estimator $\widehat{\alpha }_{n}$. This result is then
used to establish the consistency of the bootstrap sieve Wald and the
bootstrap SQLR statistics under virtually the same conditions as those used
to derive the limiting distributions of the original-sample statistics.
\footnote{
The convergence rate of the bootstrap PSMD estimator is also very useful for
the consistency of the bootstrap Wald statistic for semiparametric two-step
GMM estimation of Euclidean parameters when the first-step unknown functions
are estimated via a PSMD procedure. See e.g., \cite{CLvK_Emetrica03}}
There are some published work about estimation of and inference on a
particular linear functional, the Euclidean parameter $\phi (\alpha )=\theta
$, of the general model (\ref{semi00}) when $\theta _{0}$ is assumed to be
root-$n$ estimable; see \cite{AC_Emetrica03}, \cite{CP_WP07a}, \cite
{OTSU_WP11} and others. None of the existing work allows for $\theta _{0}$
being \textit{irregular} (i.e., slower than root-$n$ estimable),\footnote{
It is known that $\theta _{0}$ could have singular semiparametric efficiency
bound and could not be root-$n$ estimable; see \cite{CHAMBERLAIN2010}, \cite
{KhanTamer2010}, \cite{GPowell2012} and the references therein. Following
\cite{KhanTamer2010} and \cite{GPowell2012} we call such a $\theta _{0}$
irregular. Many applied papers on complicated semi/nonparametric models
simply assume that $\theta _{0}$ is root-$n$ estimable.} however. When
specializing our general theory to inference on $\theta _{0}$ of the model (
\ref{semi00}), we not only recover the results of \cite{AC_Emetrica03} and
\cite{CP_WP07a}, but also provide local power properties of sieve Wald and
SQLR as well as valid bootstrap (possibly non-optimally weighted) SQLR
inference. Moreover, our results remain valid even when $\theta _{0}$ is
irregular.
When specializing our theory to inference on a particular irregular linear
functional, the point evaluation functional $\phi (\alpha )=h(\overline{y}
_{2})$, of the semi/nonparametric IV model (\ref{gnpiv}), we automatically
obtain the pointwise asymptotic normality of the PSMD estimator of $h_{0}(
\overline{y}_{2})$ and different ways to construct its confidence set. These
results are directly applicable to the NPIV example with $\rho (Y_{1};\theta
_{0},h_{0}(Y_{2}))=Y_{1}-h_{0}(Y_{2})$ and to the NPQIV example with $\rho
(Y_{1};\theta _{0},h_{0}(Y_{2}))=1\{Y_{1}\leq h_{0}(Y_{2})\}-\gamma $.
Previously, \cite{Horowitz_07} and \cite{CGS_WP08} established the pointwise
asymptotic normality of their kernel based function space Tikhonov
regularization estimators of $h_{0}(\overline{y}_{2})$ for the NPIV and the
NPQIV examples respectively. Immediately after our paper was first presented
in April 2009 Banff/Canada conference on semiparametrics, the authors of
\cite{HL_WP10} informed us that they were concurrently working on confidence
bands for $h_{0}$ using a particular SMD estimator of the NPIV example. To
the best of our knowledge, there is no inference results, in the existing
literature, on any nonlinear functional of $h_{0}$ even for the NPIV and
NPQIV examples. Our paper is the first to provide simple sieve Wald and SQLR
tests for (possibly) nonlinear functionals satisfying the general
semi/nonparametric IV model (\ref{gnpiv}).
The rest of the paper is organized as follows. Section \ref{sec:model}
presents the plug-in PSMD estimator $\phi (\widehat{\alpha }_{n})$ of a
(possibly nonlinear) functional $\phi $ evaluated at $\alpha _{0}\equiv
(\theta _{0}^{\prime },h_{0})$ satisfying the model (\ref{semi00}). It also
provides an overview of the main asymptotic results that will be established
in the subsequent sections, and illustrates the applications through a point
evaluation functional $\phi (\alpha )=h(\overline{y}_{2})$, a weighted
derivative functional $\phi (h)=\int w(y_{2})\nabla h(y_{2})dy_{2}$, and a
quadratic functional $\phi (\alpha )=\int w(y_{2})\left\vert
h(y_{2})\right\vert ^{2}dy_{2}$ of the NPIV and NPQIV examples. Section \ref
{sec:conditions} states the basic regularity conditions. Section \ref
{sec:AsymDist} provides the asymptotic properties of sieve t (Wald) and
sieve QLR statistics. Section \ref{sec:bootstrap} establishes the
consistency of the bootstrap sieve t (Wald) and the bootstrap SQLR
statistics. Section \ref{sec-ex} verifies the key regularity conditions for
the asymptotic theories via the three functionals of the NPIV and NPQIV
examples presented in Section \ref{sec:model}. Section \ref
{sec:sec_simulation} presents simulation studies and an empirical
illustration. Section \ref{sec:conclusion} briefly concludes. Appendix \ref
{app:appA} consists of several subsections, presenting (1) further results
on sieve Riesz representation of a functional of interest; (2) the
convergence rates of the bootstrap PSMD estimator $\widehat{\alpha }_{n}^{B}$
for model (\ref{semi00}); (3) the local power properties of sieve Wald and
SQLR tests and of their bootstrap versions; (4) asymptotic properties of
sieve Wald and SQLR for functionals of increasing dimension; (5) low level
sufficient conditions with a series least squares (LS) estimated conditional
mean function $m(\cdot ,\alpha )=E[\rho (Y,X;\alpha )|X=\cdot ]$; and (6)
additional useful lemmas with series LS estimated $m(\cdot ,\alpha )$.
Online supplemental materials consist of Appendices \ref{app:appB}, \ref{app:appC} and \ref
{app:appD}. Appendix \ref{app:appB} contains additional theoretical results
(including other consistent variance estimators and other bootstrap sieve
Wald tests) and proofs of all the results stated in the main text. Appendix
\ref{app:appC} contains proofs of all the results stated in Appendix \ref
{app:appA}. The online Appendix \ref{app:appD} provides
computationally attractive sieve score test and sieve score bootstrap.
\textbf{Notation}. We use \textquotedblleft $\equiv $\textquotedblright\ to
implicitly define a term or introduce a notation. For any column vector $A$,
we let $A^{\prime }$ denote its transpose and $||A||_{e}$ its Euclidean norm
(i.e., $||A||_{e}\equiv \sqrt{A^{\prime }A}$, although sometimes we use $
|A|=||A||_{e}$ for simplicity). Let $||A||_{W}^{2}\equiv A^{\prime }WA$ for
a positive definite weighting matrix $W$. Let $\lambda _{\max }(W)$ and $
\lambda _{\min }(W)$ denote the maximal and minimal eigenvalues of $W$
respectively. All random variables $Z\equiv (Y^{\prime },X^{\prime
})^{\prime }$, $Z_{i}\equiv (Y_{i}^{\prime },X_{i}^{\prime })^{\prime }$ are
defined on a complete probability space $(\mathcal{Z},\mathcal{B}_{Z},P_{Z})$
, where $P_{Z}$ is the joint probability distribution of $(Y^{\prime
},X^{\prime })$. We define $(\mathcal{Z}^{\infty },\mathcal{B}_{Z}^{\infty
},P_{Z^{\infty }})$ as the probability space of the sequences $
(Z_{1},Z_{2},...)$. For simplicity we assume that $Y$ and $X$ are continuous
random variables. Let $f_{X}$ ($F_{X}$) be the marginal density (cdf) of $X$
with support $\mathcal{X}$, and $f_{Y|X}$ ($F_{Y|X}$) be the conditional
density (cdf) of $Y$ given $X$. Let $E_{P}[\cdot ]$ denote the expectation
with respect to a measure $P$. Sometimes we use $P$ for $P_{Z^{\infty }}$
and $E[\cdot ]$ for $E_{P_{Z^{\infty }}}[\cdot ]$. Denote $L^{p}(\Omega
,d\mu )$, $1\leq p<\infty $, as a space of measurable functions with $
||g||_{L^{p}(\Omega ,d\mu )}\equiv \{\int_{\Omega }|g(t)|^{p}d\mu
(t)\}^{1/p}<\infty $, where $\Omega $ is the support of the sigma-finite
positive measure $d\mu $ (sometimes $L^{p}(d\mu )$ and $||g||_{L^{p}(d\mu )}$
are used). For any (possibly random) positive sequences $\{a_{n}\}_{n=1}^{
\infty }$ and $\{b_{n}\}_{n=1}^{\infty }$, $a_{n}=O_{P}(b_{n})$ means that $
\lim_{c\rightarrow \infty }\limsup_{n}\Pr \left( a_{n}/b_{n}>c\right) =0$; $
a_{n}=o_{P}(b_{n})$ means that for all $\varepsilon >0$, $\lim_{n\rightarrow
\infty }\Pr \left( a_{n}/b_{n}>\varepsilon \right) =0$; and $a_{n}\asymp
b_{n}$ means that there exist two constants $0<c_{1}\leq c_{2}<\infty $ such
that $c_{1}a_{n}\leq b_{n}\leq c_{2}a_{n}$. Also, we use \textquotedblleft
wpa1-$P_{Z^{\infty }}$\textquotedblright\ (or simply wpa1) for an event $
A_{n}$, to denote that $P_{Z^{\infty }}(A_{n})\rightarrow 1$ as $
n\rightarrow \infty $. We use $\mathcal{A}_{n}\equiv \mathcal{A}_{k(n)}$ and
$\mathcal{H}_{n}\equiv \mathcal{H}_{k(n)}$ for various sieve spaces. We
assume $\dim (\mathcal{A}_{k(n)})\asymp \dim (\mathcal{H}_{k(n)})\asymp k(n)$
for simplicity, all of which grow to infinity with the sample size $n$. We
use $const.$, $c$ or $C$ to mean a positive finite constant that is
independent of sample size but can take different values at different
places. For sequences, $(a_{n})_{n}$, we sometimes use $a_{n}\nearrow a$ ($
a_{n}\searrow a$) to denote, that the sequence converges to $a$ and that is
increasing (decreasing) sequence. For any mapping $\digamma :\mathbf{H}
_{1}\rightarrow \mathbf{H}_{2}$ between two generic Banach spaces, $\frac{
d\digamma (\alpha _{0})}{d\alpha }[v]\equiv \left. \frac{\partial \digamma
(\alpha _{0}+\tau v)}{\partial \tau }\right\vert _{\tau =0}$ is the pathwise
(or Gateaux) derivative at $\alpha _{0}$ in the direction $v\in \mathbf{H}
_{1}$. And $\frac{d\digamma (\alpha _{0})}{d\alpha }[\mathbf{v}^{\prime
}]\equiv \left( \frac{d\digamma (\alpha _{0})}{d\alpha }[v_{1}],\cdot \cdot
\cdot ,\frac{d\digamma (\alpha _{0})}{d\alpha }[v_{k}]\right) $ for $\mathbf{
v}^{\prime }=\left( v_{1},\cdot \cdot \cdot ,v_{k}\right) $ with $v_{j}\in
\mathbf{H}_{1}$ for all $j=1,...,k$.
\section{PSMD Estimation and Inferences: An Overview}
\label{sec:model}
\subsection{The Penalized Sieve Minimum Distance Estimator}
Let $m(X,\alpha )\equiv E\left[ \rho (Y,X;\alpha )|X\right] =\int \rho
(y,X;\alpha )dF_{Y|X}(y)$ be a $d_{\rho }\times 1$ vector valued conditional
mean function, $\Sigma (X)$ be a $d_{\rho }\times d_{\rho }$ positive
definite ($a.s.-X$) weighting matrix, and
\begin{equation*}
Q(\alpha )\equiv E\left[ m(X,\alpha )^{\prime }\Sigma (X)^{-1}m(X,\alpha )
\right] \equiv E\left[ ||m(X,\alpha )||_{\Sigma ^{-1}}^{2}\right]
\end{equation*}
be the population minimum distance (MD) criterion function. Then the
semi/nonparametric conditional moment model (\ref{semi00}) can be
equivalently expressed as $m(X,\alpha _{0})=0$ $a.s.-X$, where $\alpha
_{0}\equiv (\theta _{0}^{\prime },h_{0})\in \mathcal{A}\equiv \Theta \times
\mathcal{H}$, or as
\begin{equation*}
\inf_{\alpha \in \mathcal{A}}Q(\alpha )=Q(\alpha _{0})=0.
\end{equation*}
Let $\Sigma _{0}(X)\equiv Var(\rho (Y,X;\alpha _{0})|X)$ be positive
definite for almost all $X$. In this paper as well as in most applications $
\Sigma (X)$ is chosen to be either $I_{d_{\rho }}$ (identity) or $\Sigma
_{0}(X)$ for almost all $X$. We call $Q^{0}(\alpha )\equiv E\left[
||m(X,\alpha )||_{\Sigma _{0}^{-1}}^{2}\right] $ the population optimally
weighted MD criterion function.
Let $\phi :\mathcal{A}\rightarrow \mathbb{R}^{d_{\phi }}$ be a functional
with a finite $d_{\phi }\geq 1$. We are interested in inference on $\phi
(\alpha _{0})$. Let
\begin{equation}
\widehat{Q}_{n}(\alpha )\equiv \frac{1}{n}\sum_{i=1}^{n}\widehat{m}
(X_{i},\alpha )^{\prime }\widehat{\Sigma }(X_{i})^{-1}\widehat{m}
(X_{i},\alpha ) \label{Qhat}
\end{equation}
be a sample estimate of $Q(\alpha )$, where $\widehat{m}(X,\alpha )$ and $
\widehat{\Sigma }(X)$ are any consistent estimators of $m(X,\alpha )$ and $
\Sigma (X)$ respectively. When $\widehat{\Sigma }(X)=\widehat{\Sigma }
_{0}(X) $ is a consistent estimator of the optimal weighting matrix $\Sigma
_{0}(X)$, we call the corresponding $\widehat{Q}_{n}(\alpha )$ the\ sample
optimally weighted MD criterion $\widehat{Q}_{n}^{0}(\alpha )$.
We estimate $\phi (\alpha _{0})$ by $\phi (\widehat{\alpha }_{n})$, where $
\widehat{\alpha }_{n}\equiv (\widehat{\theta }_{n}^{\prime },\widehat{h}
_{n}) $ is an approximate \textit{penalized sieve minimum distance} (PSMD)
estimator of $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$, defined as
\begin{equation}
\widehat{Q}_{n}(\widehat{\alpha }_{n})+\lambda _{n}Pen(\widehat{h}_{n})\leq
\inf_{\alpha \in \mathcal{A}_{k(n)}}\left\{ \widehat{Q}_{n}(\alpha )+\lambda
_{n}Pen(h)\right\} +o_{P_{Z^{\infty }}}(n^{-1}), \label{psmd}
\end{equation}
where $\lambda _{n}Pen(h)\geq 0$ is a penalty term such that $\lambda
_{n}=o(1)$; and $\mathcal{A}_{k(n)}\equiv \Theta \times \mathcal{H}_{k(n)}$
is a finite dimensional sieve for $\mathcal{A}\equiv \Theta \times \mathcal{H
}$, more precisely, $\mathcal{H}_{k(n)}$ is a finite dimensional \textit{
linear} sieve for $\mathcal{H}$:
\begin{equation}
\mathcal{H}_{k(n)}=\left\{ h\in \mathcal{H}:h(\cdot )=\sum_{k=1}^{k(n)}\beta
_{k}q_{k}(\cdot )=\beta ^{\prime }q^{k(n)}(\cdot )\right\} , \label{sieve}
\end{equation}
where $\{q_{k}\}_{k=1}^{\infty }$ is a sequence of known basis functions of
a Banach space $(\mathcal{H},\left\Vert \cdot \right\Vert _{\mathbf{H}})$
such as wavelets, splines, Fourier series, Hermite polynomial series, etc.
And $k(n)\rightarrow \infty $ as $n\rightarrow \infty $.
For the purely nonparametric conditional moment models $E\left[ \rho
(Y,X;h_{0})|X\right] =0$, \cite{CP_WP07} proposed more general approximate
PSMD estimators of $h_{0}$ by allowing for possibly infinite dimensional
sieves (i.e., $\dim (\mathcal{H}_{k(n)})=k(n)\leq \infty $). Nevertheless,
both the theoretical properties and Monte Carlo simulations in \cite{CP_WP07}
recommend the use of the PSMD procedures with slowly growing
finite-dimensional linear sieves with a tiny penalty (i.e., $k(n)\rightarrow
\infty ,\frac{k(n)}{n}\rightarrow 0$ as $n$ $\rightarrow \infty $ with a
very small $\lambda _{n}=o(n^{-1})$, and hence the main smoothing parameter
is the sieve dimension $k(n)$). This class of PSMD estimators include the
original SMD estimators of \cite{NP_ECMA03} and \cite{AC_Emetrica03} as
special cases, and has been used in recent empirical estimation of
semiparametric structural models in microeconomics and asset pricing with
endogeneity. See, e.g., \cite{BCK_Emetrica07}, \cite{Horowitz_ECMA11}, \cite
{CP_WP07a}, \cite{BHN_WP11}, \cite{Souza2012}, \cite{Pinkse2002}, \cite
{MerloPaula}, \cite{BMT_WP12}, \cite{ChenLudvigson_JAE09}, \cite
{ChenLudvigson2013}, \cite{Sentana2013} and others.
In this paper we shall develop inferential theory for $\phi (\alpha _{0})$
based on the PSMD procedures with slowly growing finite-dimensional sieves $
\mathcal{A}_{k(n)}=\Theta \times \mathcal{H}_{k(n)}$. We first establish the
large sample theories under a high level \textquotedblleft local quadratic
approximation\textquotedblright\ (LQA) condition, which allows for any
consistent nonparametric estimator $\widehat{m}(x,\alpha )$ that is linear
in $\rho (Z,\alpha )$:
\begin{equation}
\widehat{m}(x,\alpha )\equiv \sum_{i=1}^{n}\rho (Z_{i},\alpha )A_{n}(X_{i},x)
\label{mhat-linear}
\end{equation}
where $A_{n}(X_{i},x)$ is a known measurable function of $
\{X_{j}\}_{j=1}^{n} $ for all $x$, whose expression varies according to
different nonparametric procedures such as kernel, local linear regression,
series and nearest neighbors. In Appendix \ref{app:appA} we provide lower
level sufficient conditions for this LQA assumption when $\widehat{m}
(x,\alpha )$ is the series least squares (LS) estimator (\ref{mhat}):
\begin{equation}
\widehat{m}(x,\alpha )=\left( \sum_{i=1}^{n}\rho (Z_{i},\alpha
)p^{J_{n}}(X_{i})^{\prime }\right) (P^{\prime }P)^{-}p^{J_{n}}(x),
\label{mhat}
\end{equation}
which is a linear nonparametric estimator (\ref{mhat-linear}) with $
A_{n}(X_{i},x)=p^{J_{n}}(X_{i})^{\prime }(P^{\prime }P)^{-}p^{J_{n}}(x)$,
where $\{p_{j}\}_{j=1}^{\infty }$ is a sequence of known basis functions
that can approximate any square integrable functions of $X$ well, $
p^{J_{n}}(X)=(p_{1}(X),...,p_{J_{n}}(X))^{\prime }$, $
P=(p^{J_{n}}(X_{1}),...,p^{J_{n}}(X_{n}))^{\prime }$, and $(P^{\prime
}P)^{-} $ is the generalized inverse of the matrix $P^{\prime }P$. Following
\cite{BCK_Emetrica07} and \cite{CP_WP07a}, we let $p^{J_{n}}(X)$ be a
tensor-product linear sieve basis, and $J_{n}$ be the dimension of $
p^{J_{n}}(X)$ such that $J_{n}\geq d_{\theta }+k(n)\rightarrow \infty $ and $
\frac{J_{n}}{n}\rightarrow 0$ as $n$ $\rightarrow \infty $.
\subsection{Preview of the Main Results for Inference}
\label{sec:NPIVex}
For simplicity we let $\phi :\mathbb{R}^{d_{\theta }}\times \mathcal{H}
\rightarrow \mathbb{R}$ be a real-valued functional. Let $\widehat{\phi }
_{n}\equiv \phi (\widehat{\alpha }_{n})$ be the \textit{plug-in PSMD
estimator} of $\phi (\alpha _{0})$ for $\alpha _{0}=(\theta _{0}^{\prime
},h_{0})\in int(\Theta )\times \mathcal{H}$.
\textbf{Sieve t (or Wald) statistic}. Regardless of whether $\phi (\alpha
_{0})$ is $\sqrt{n}$ estimable or not, Theorem \ref{thm:theta_anorm} shows
that $\frac{\sqrt{n}\left\{ \phi (\widehat{\alpha }_{n})-\phi (\alpha
_{0})\right\} }{||v_{n}^{\ast }||_{sd}}$ is asymptotically standard normal,
and the sieve variance $||v_{n}^{\ast }||_{sd}^{2}$ has a \textit{closed form
} expression resembling the \textquotedblleft
delta-method\textquotedblright\ variance for a parametric MD problem:
\begin{equation}
||v_{n}^{\ast }||_{sd}^{2}=\left( \frac{d\phi (\alpha _{0})}{d\alpha }[
\overline{q}^{k(n)}(\cdot )]\right) ^{\prime }D_{n}^{-}\mho
_{n}D_{n}^{-}\left( \frac{d\phi (\alpha _{0})}{d\alpha }[\overline{q}
^{k(n)}(\cdot )]\right) , \label{P-var}
\end{equation}
where $\overline{q}^{k(n)}(\cdot )\equiv \left( \mathbf{1}_{d_{\theta
}}^{\prime },q^{k(n)}(\cdot )^{\prime }\right) ^{\prime }$ is a $(d_{\theta
}+k(n))\times 1$ vector with $\mathbf{1}_{d_{\theta }}$ a $d_{\theta }\times
1$ vector of $1$'s,
\begin{equation}
\frac{d\phi (\alpha _{0})}{d\alpha }[\overline{q}^{k(n)}(\cdot )]\equiv
\frac{\partial \phi (\theta _{0}+\theta ,h_{0}+\beta ^{\prime
}q^{k(n)}(\cdot ))}{\partial \gamma ^{\prime }} \mid_{\gamma =0}\equiv
\left( \frac{\partial \phi (\alpha _{0})}{\partial \theta ^{\prime }},\frac{
d\phi (\alpha _{0})}{dh}[q^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }
\label{P_kT_tilde}
\end{equation}
and $\gamma \equiv (\theta ^{\prime },\beta ^{\prime })^{\prime }$ are $
(d_{\theta }+k(n))\times 1$ vectors, $\frac{d\phi (\alpha _{0})}{dh}
[q^{k(n)}(\cdot )^{\prime }]\equiv \frac{\partial \phi (\theta
_{0},h_{0}+\beta ^{\prime }q^{k(n)}(\cdot ))}{\partial \beta } \mid_{\beta =0}$, and
\begin{equation}
D_{n}=E\left[ \left( \frac{dm(X,\alpha _{0})}{d\alpha }[\overline{q}
^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\Sigma (X)^{-1}\left( \frac{
dm(X,\alpha _{0})}{d\alpha }[\overline{q}^{k(n)}(\cdot )^{\prime }]\right)
\right] , \label{P-R}
\end{equation}
{\small{\begin{equation}
\mho _{n}=E\left[ \left( \frac{dm(X,\alpha _{0})}{d\alpha }[\overline{q}
^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\Sigma (X)^{-1}\rho (Z,\alpha
_{0})\rho (Z,\alpha _{0})^{\prime }\Sigma (X)^{-1}\left( \frac{dm(X,\alpha
_{0})}{d\alpha }[\overline{q}^{k(n)}(\cdot )^{\prime }]\right) \right] ,
\label{P-Omega}
\end{equation}}}
where $\frac{dm(X,\alpha _{0})}{d\alpha }[\overline{q}^{k(n)}(\cdot
)^{\prime }]\equiv \frac{\partial E[\rho (Z,\theta _{0}+\theta ,h_{0}+\beta
^{\prime }q^{k(n)}(\cdot ))|X]}{\partial \gamma } \mid_{\gamma =0}$ is
a $d_{\rho }\times (d_{\theta }+k(n))$ matrix. The closed form expression of
$||v_{n}^{\ast }||_{sd}^{2}$ immediately leads to simple consistent plug-in
sieve variance estimators; one of which is
\begin{equation}
||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}=\widehat{V}_{1}=\left( \frac{d\phi (
\widehat{\alpha }_{n})}{d\alpha }[\overline{q}^{k(n)}(\cdot )]\right)
^{\prime }\widehat{D}_{n}^{-}\widehat{\mho }_{n}\widehat{D}_{n}^{-}\left(
\frac{d\phi (\widehat{\alpha }_{n})}{d\alpha }[\overline{q}^{k(n)}(\cdot
)]\right) , \label{P-var-hat}
\end{equation}
where $\frac{d\phi (\widehat{\alpha }_{n})}{d\alpha }[\overline{q}
^{k(n)}(\cdot )]\equiv \frac{\partial \phi (\widehat{\theta }_{n}+\theta ,
\widehat{h}_{n}+\beta ^{\prime }q^{k(n)}(\cdot ))}{\partial \gamma ^{\prime }
} \mid_{\gamma =0}$ and
\begin{equation}
\widehat{D}_{n}=\frac{1}{n}\sum_{i=1}^{n}\left[ \left( \frac{d\widehat{m}
(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{q}^{k(n)}(\cdot )^{\prime
}]\right) ^{\prime }\widehat{\Sigma }(X_{i})^{-1}\left( \frac{d\widehat{m}
(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{q}^{k(n)}(\cdot )^{\prime
}]\right) \right] , \label{P-R-hat}
\end{equation}
\begin{equation}
\widehat{\mho }_{n}=\frac{1}{n}\sum_{i=1}^{n}\left[ \left( \frac{d\widehat{m}
(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{q}^{k(n)}(\cdot )^{\prime
}]\right) ^{\prime }\widehat{M}(X_{i})\left( \frac{d\widehat{m}(X_{i},\widehat{\alpha }_{n})}{d\alpha }
[\overline{q}^{k(n)}(\cdot )^{\prime }]\right) \right] . \label{P-Omega-hat}
\end{equation}
where $\widehat{M}(X_{i}) \equiv \widehat{\Sigma }(X_{i})^{-1}\rho (Z_{i},\widehat{\alpha
}_{n})\rho (Z_{i},\widehat{\alpha }_{n})^{\prime }\widehat{\Sigma }
(X_{i})^{-1}$. Theorem \ref{thm:VE} then presents the asymptotic normality of the sieve
(Student's) t statistic:\footnote{
See Theorems \ref{thm:bootstrap_2} and \ref{thm:waldB_con} for properties of
bootstrap sieve t statistics.}
\begin{equation*}
\widehat{W}_{n}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha }_{n})-\phi
(\alpha _{0})}{||\widehat{v}_{n}^{\ast }||_{n,sd}}\Rightarrow N(0,1).
\end{equation*}
\textbf{Sieve QLR statistic}. In addition to the sieve t (or sieve Wald)
statistic, we could also use sieve quasi likelihood ratio for constructing
confidence set of $\phi (\alpha _{0})$ and for hypothesis testing of $
H_{0}:\phi (\alpha _{0})=\phi _{0}$ against $H_{1}:\phi (\alpha _{0})\neq
\phi _{0}$. Denote
\begin{equation}
\widehat{QLR}_{n}(\phi _{0})\equiv n\left( \inf_{\alpha \in \mathcal{A}
_{k(n)}:\phi (\alpha )=\phi _{0}}\widehat{Q}_{n}(\alpha )-\widehat{Q}_{n}(
\widehat{\alpha }_{n})\right) \label{SQLR}
\end{equation}
as the \textit{sieve quasi likelihood ratio} (SQLR) statistic. It becomes an
\textit{optimally weighted SQLR} statistic, $\widehat{QLR}_{n}^{0}(\phi
_{0}) $, when $\widehat{Q}_{n}(\alpha )$ is the optimally weighted MD
criterion $\widehat{Q}_{n}^{0}(\alpha )$. Regardless of whether $\phi
(\alpha _{0})$ is $\sqrt{n}$ estimable or not, Theorems \ref{thm:chi2}(2)
and \ref{thm:QLR-H1} show that $\widehat{QLR}_{n}^{0}(\phi _{0})$ is
asymptotically chi-square distributed under the null $H_{0}$, and diverges
to infinity under the fixed alternatives $H_{1}$. Theorem \ref
{thm:chi2_localt} in Appendix \ref{app:appA} states that $\widehat{QLR}
_{n}^{0}(\phi _{0})$ is asymptotically noncentral chi-square distributed
under local alternatives. One could compute $100(1-\tau )\%$ confidence set
for $\phi (\alpha _{0})$ as
\begin{equation*}
\left\{ r\in \mathbb{R}\colon \text{ }\widehat{QLR}_{n}^{0}(r)\leq c_{\chi
_{1}^{2}}(1-\tau )\right\} ,
\end{equation*}
where $c_{\chi _{1}^{2}}(1-\tau )$ is the $(1-\tau )$-th quantile of the $
\chi _{1}^{2}$ distribution.
\textbf{Bootstrap sieve QLR statistic}. Regardless of whether $\phi (\alpha
_{0})$ is $\sqrt{n}$ estimable or not, Theorems \ref{thm:chi2}(1) and \ref
{thm:QLR-H1} establish that the possibly non-optimally weighted SQLR
statistic $\widehat{QLR}_{n}(\phi _{0})$ is stochastically bounded under the
null $H_{0}$ and diverges to infinity under the fixed alternatives $H_{1}$.
We then consider a bootstrap version of the SQLR statistic. Let $\widehat{QLR
}_{n}^{B}$ denote a bootstrap SQLR statistic:
\begin{equation}
\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})\equiv n\left( \inf_{\alpha \in
\mathcal{A}_{k(n)}:\phi (\alpha )=\widehat{\phi }_{n}}\widehat{Q}
_{n}^{B}(\alpha )-\inf_{\alpha \in \mathcal{A}_{k(n)}}\widehat{Q}
_{n}^{B}(\alpha )\right) , \label{SQLR-B}
\end{equation}
where $\widehat{\phi }_{n}\equiv \phi (\widehat{\alpha }_{n})$, and $
\widehat{Q}_{n}^{B}(\alpha )$ is a bootstrap version of $\widehat{Q}
_{n}(\alpha )$:
\begin{equation}
\widehat{Q}_{n}^{B}(\alpha )\equiv \frac{1}{n}\sum_{i=1}^{n}\widehat{m}
^{B}(X_{i},\alpha )^{\prime }\widehat{\Sigma }(X_{i})^{-1}\widehat{m}
^{B}(X_{i},\alpha ), \label{QhatB}
\end{equation}
where $\widehat{m}^{B}(x,\alpha )$ is a bootstrap version of $\widehat{m}
(x,\alpha )$, which is computed in the same way as that of $\widehat{m}
(x,\alpha )$ except that we use $\omega _{i,n}\rho (Z_{i},\alpha )$ instead
of $\rho (Z_{i},\alpha )$. Here $\{\omega _{i,n}\geq 0\}_{i=1}^{n}$ is a
sequence of bootstrap weights that has mean 1 and is independent of the
original data $\{Z_{i}\}_{i=1}^{n}$. Typical weights include an i.i.d.
weight $\{\omega _{i}\geq 0\}_{i=1}^{n}$ with $E[\omega _{i}]=1$, $E[|\omega
_{i}-1|^{2}]=1$ and $E[|\omega _{i}-1|^{2+\epsilon }]<\infty $ for some $
\epsilon >0$, or a multinomial weight (i.e., $(\omega _{1,n},...,\omega
_{n,n})\sim Multinomial(n;n^{-1},...,n^{-1})$). For example, if $\widehat{m}
(x,\alpha )$ is a series LS estimator (\ref{mhat}) of $m(x,\alpha )$, then $
\widehat{m}^{B}(x,\alpha )$ is a bootstrap series LS estimator of $
m(x,\alpha )$, defined as:
\begin{equation}
\widehat{m}^{B}(x,\alpha )\equiv \left( \sum_{i=1}^{n}\omega _{i,n}\rho
(Z_{i},\alpha )p^{J_{n}}(X_{i})^{\prime }\right) (P^{\prime
}P)^{-}p^{J_{n}}(x). \label{mhat_B}
\end{equation}
We sometimes call our bootstrap procedure \textquotedblleft \textit{
generalized residual bootstrap}\textquotedblright\ since it is based on
randomly perturbing the generalized residual function $\rho (Z,\alpha )$;
see Section \ref{sec:bootstrap} for details. Theorems \ref{thm:bootstrap}
and \ref{thm:BSQLR-loc-alt} establish that under the null $H_{0}$, the fixed
alternatives $H_{1}$ or the local alternatives,\footnote{
See Section \ref{subsec-A5} for definition of the local alternatives and the
behaviors of $\widehat{QLR}_{n}(\phi _{0})$ and $\widehat{QLR}_{n}^{B}(
\widehat{\phi }_{n})$ under the local alternatives.} the conditional
distribution of $\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})$ (given the
data) always converges to the asymptotic null distribution of $\widehat{QLR}
_{n}(\phi _{0})$. Let $\widehat{c}_{n}(a)$ be the $a-th$ quantile of the
distribution of $\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})$ (conditional on
the data $\{Z_{i}\}_{i=1}^{n}$). Then for any $\tau \in (0,1)$, we have $
\lim_{n\rightarrow \infty }\Pr \{\widehat{QLR}_{n}(\phi _{0})>\widehat{c}
_{n}(1-\tau )\}=\tau $ under the null $H_{0}$, $\lim_{n\rightarrow \infty
}\Pr \{\widehat{QLR}_{n}(\phi _{0})>\widehat{c}_{n}(1-\tau )\}=1$ under the
fixed alternatives $H_{1}$, and $\lim_{n\rightarrow \infty }\Pr \{\widehat{
QLR}_{n}(\phi _{0})>\widehat{c}_{n}(1-\tau )\}>\tau $ under the local
alternatives. We could also construct a $100(1-\tau )\%$ confidence set
using the bootstrap critical values:
\begin{equation}
\left\{ r\in \mathbb{R}\colon \text{ }\widehat{QLR}_{n}(r)\leq \widehat{c}
_{n}(1-\tau )\right\} . \label{boot-cs}
\end{equation}
The bootstrap consistency holds for possibly non-optimally weighted SQLR
statistic and possibly irregular functionals, without the need to compute
standard errors.
\textbf{Which method to use?} When sieve Wald and SQLR tests are computed
using the same weighting matrix $\widehat{\Sigma }$, there is no local power
difference in terms of first order asymptotic theories; see Appendix \ref
{app:appA}. As will be demonstrated in simulation Section \ref
{sec:sec_simulation}, while SQLR and bootstrap SQLR tests are useful for
models (\ref{semi00}) with (pointwise) non-smooth $\rho (Z;\alpha )$, sieve
Wald (or t) statistic is computationally attractive for models with smooth $
\rho (Z;\alpha )$. Empirical researchers could apply either inference method
depending on whether the residual function $\rho (Z;\alpha )$ in their
specific application is pointwise differentiable with respect to $\alpha $
or not.
\subsubsection{Applications to NPIV and NPQIV models\label{sec:NPIVex1}}
\textbf{An illustration via the NPIV model. }\cite{BCK_Emetrica07} and \cite
{CR_WP07} established the convergence rate of the identity weighted (i.e., $
\widehat{\Sigma }=\Sigma =1$) PSMD estimator $\widehat{h}_{n}\in \mathcal{H}
_{k(n)}$ of the NPIV model:
\begin{equation}
Y_{1}=h_{0}(Y_{2})+U,\text{\quad }E(U|X)=0. \label{npiv}
\end{equation}
By Theorem \ref{thm:theta_anorm} \begin{align*}
\sqrt{n}\frac{\phi (\widehat{h}_{n})-\phi
(h_{0})}{||v_{n}^{\ast }||_{sd}}\Rightarrow N(0,1)
\end{align*} with $||v_{n}^{\ast
}||_{sd}^{2}=\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]^{\prime
}D_{n}^{-}\mho _{n}D_{n}^{-}\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]$,
\begin{align}
D_{n}&=E\left( E[q^{k(n)}(Y_{2})|X]E[q^{k(n)}(Y_{2})|X]^{\prime }\right)
\text{, }\\
\mho _{n}&=E\left(
E[q^{k(n)}(Y_{2})|X]U^{2}E[q^{k(n)}(Y_{2})|X]^{\prime }\right)
\label{npiv-D}
\end{align}
and $\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]\equiv \frac{\partial \phi
(h_{0}+\beta ^{\prime }q^{k(n)}(\cdot ))}{\partial \beta ^{\prime }} \mid_{\beta =0}$. For example, for a functional $\phi (h)=h(\overline{y}_{2})$,
or $=\int w(y)\nabla h(y)dy$ or $=\int w(y)\left\vert h(y)\right\vert ^{2}dy$
, we have $\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]=q^{k(n)}(\overline{y}
_{2})$, or $=\int w(y)\nabla q^{k(n)}(y)dy$ or $=2\int
h_{0}(y)w(y)q^{k(n)}(y)dy$.
If $0<\inf_{x}\Sigma _{0}(x)\leq \sup_{x}\Sigma _{0}(x)<\infty $ then \begin{align*}
||v_{n}^{\ast }||_{sd}^{2}\asymp \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot
)]^{\prime }D_{n}^{-}\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]
\end{align*}
Without
endogeneity (say $Y_{2}=X$) the model becomes the nonparametric LS
regression
\begin{equation*}
Y_{1}=h_{0}(Y_{2})+U,\text{\quad }E(U|Y_{2})=0,
\end{equation*}
and the variance satisfies $||v_{n}^{\ast }||_{sd,ex}^{2}\asymp \frac{d\phi
(h_{0})}{dh}[q^{k(n)}(\cdot )]^{\prime }D_{n,ex}^{-}\frac{d\phi (h_{0})}{dh}
[q^{k(n)}(\cdot )]$, $D_{n,ex}=E[\{q^{k(n)}(Y_{2})\}\{q^{k(n)}(Y_{2})\}^{
\prime }]$. Since the conditional expectation $E[q^{k(n)}(Y_{2})|X]$ is a
contraction, $D_{n}\leq D_{n,ex}$ and $||v_{n}^{\ast }||_{sd}^{2}\geq
const.||v_{n}^{\ast }||_{sd,ex}^{2}$. Under mild conditions (see, e.g., \cite
{NP_ECMA03}, \cite{BCK_Emetrica07}, \cite{DFR_wp10}, \cite{Horowitz_ECMA11}
), the minimal eigenvalue of $D_{n}$, $\lambda _{\min }(D_{n})$, goes to
zero while $\lambda _{\min }(D_{n,ex})$ stays strictly positive as $
k(n)\rightarrow \infty $. In fact, $D_{n,ex}=I_{k(n)}$ and $\lambda _{\min
}(D_{n,ex})=1$ if $\{q_{j}\}_{j=1}^{\infty }$ is an orthonormal basis of $
L^{2}(f_{Y_{2}})$, while $\lambda _{\min }(D_{n})\asymp \exp (-k(n))$ if the
conditional density of $Y_{2}$ given $X$ is normal. Therefore, while $
\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd,ex}^{2}=\infty $ always
implies $\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd}^{2}=\infty $,
it is possible that $\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast
}||_{sd,ex}^{2}<\infty $ but $\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast
}||_{sd}^{2}=\infty $. For example, the point evaluation functional $\phi
(h)=h(\overline{y}_{2})$ is known to be irregular for the nonparametric LS
regression and hence for the NPIV (\ref{npiv}) as well. Under mild
conditions on the weight $w()$ and the smoothness of $h_{0}$, the weighted
derivative functional ($\phi (h)=\int w(y)\nabla h(y)dy$) and the quadratic
functional ($\phi (h)=\int w(y)\left\vert h(y)\right\vert ^{2}dy$) of the
nonparametric LS regression are typically root-$n$ estimable, but they could
be irregular for the NPIV (\ref{npiv}). See Section \ref{sec-ex} for details.
Regardless of whether $\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast
}||_{sd}^{2}$ is finite or infinite, Theorem \ref{thm:VE} shows that the
sieve variance $||v_{n}^{\ast }||_{sd}^{2}$ can be consistently estimated by
a plug-in sieve variance estimator $||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}$,
and that $\sqrt{n}\frac{\phi (\widehat{h}_{n})-\phi (h_{0})}{||\widehat{v}
_{n}^{\ast }||_{n,sd}}\Rightarrow N(0,1)$.
When the conditional mean function $m(x,h)$ is estimated by the series LS
estimator (\ref{mhat}) as in \cite{NP_ECMA03}, \cite{AC_Emetrica03} and \cite
{BCK_Emetrica07}, with $\widehat{U}_{i}=Y_{1i}-\widehat{h}_{n}(Y_{2i})$, the
sieve variance estimator $||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}$ given in (
\ref{P-var-hat}) has a more explicit expression:
\begin{equation*}
||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}=\widehat{V}_{1}=\left( \frac{d\phi (
\widehat{h}_{n})}{dh}[q^{k(n)}(\cdot )]\right) ^{\prime }\widehat{D}_{n}^{-}
\widehat{\mho }_{n}\widehat{D}_{n}^{-}\left( \frac{d\phi (\widehat{h}_{n})}{
dh}[q^{k(n)}(\cdot )]\right) ,\quad \text{where}
\end{equation*}
$\frac{d\phi (\widehat{h}_{n})}{dh}[q^{k(n)}(\cdot )]\equiv \frac{\partial
\phi (\widehat{h}_{n}+\beta ^{\prime }q^{k(n)}(\cdot ))}{\partial \beta
^{\prime }} \mid_{\beta =0}$ and
\begin{equation*}
\widehat{D}_{n}=\frac{1}{n}\widehat{C}_{n}(P^{\prime }P)^{-}(\widehat{C}
_{n})^{\prime },\quad \widehat{C}_{n}\equiv
\sum_{j=1}^{n}q^{k(n)}(Y_{2j})p^{J_{n}}(X_{j})^{\prime },
\end{equation*}
\begin{equation}
\widehat{\mho }_{n}=\frac{1}{n}\widehat{C}_{n}(P^{\prime }P)^{-}\left(
\sum_{i=1}^{n}p^{J_{n}}(X_{i})\widehat{U}_{i}^{2}p^{J_{n}}(X_{i})^{\prime
}\right) (P^{\prime }P)^{-}(\widehat{C}_{n})^{\prime }. \label{2sls-vhat}
\end{equation}
Interestingly, this sieve variance estimator becomes the one computed via
the two stage least squares (2SLS) as if the NPIV model (\ref{npiv}) were a
parametric IV regression:\footnote{This confirms a conjecture of \cite{Newey2013} for the NPIV model (\ref{npiv}
).} $Y_{1}=q^{k(n)}(Y_{2j})^{\prime }\beta _{0n}+U,$ $
E[q^{k(n)}(Y_{2})U]\neq 0,$ $E[p^{J_{n}}(X)U]=0$ and $
E[p^{J_{n}}(X)q^{k(n)}(Y_{2})^{\prime }]$ has a column rank $k(n)\leq J_{n}$
. See Subsection \ref{sec:sec_simulation1} for simulation studies of finite
sample performances of this sieve variance estimator $\widehat{V}_{1}$ for
both a linear and a nonlinear functional $\phi (h)$.
\medskip
\textbf{An illustration via the NPQIV model. }As an application of their
general theory, \cite{CP_WP07} presented the consistency and the rate of
convergence of the PSMD estimator $\widehat{h}_{n}\in \mathcal{H}_{k(n)}$ of
the NPQIV model:
\begin{equation}
Y_{1}=h_{0}(Y_{2})+U,\text{\quad }\Pr (U\leq 0|X)=\gamma . \label{npqiv}
\end{equation}
In this example we have $\Sigma _{0}(X)=\gamma (1-\gamma )$. So we could use
$\widehat{\Sigma }(X)=\gamma (1-\gamma )$ and $\widehat{Q}_{n}(\alpha )$
given in (\ref{Qhat}) becomes the optimally weighted\ MD criterion.
\noindent By Theorem \ref{thm:theta_anorm}
\begin{align*}
\sqrt{n}\frac{\phi (\widehat{h}
_{n})-\phi (h_{0})}{||v_{n}^{\ast }||_{sd}}\Rightarrow N(0,1)
\end{align*}
with $
||v_{n}^{\ast }||_{sd}^{2}=\left( \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot
)]\right) ^{\prime }D_{n}^{-}\left( \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot
)]\right) $ and
\begin{equation}
D_{n}=\frac{1}{\gamma (1-\gamma )}E\left(
E[f_{U|Y_{2},X}(0)q^{k(n)}(Y_{2})|X]E[f_{U|Y_{2},X}(0)q^{k(n)}(Y_{2})|X]^{
\prime }\right) . \label{npqiv-D}
\end{equation}
Without endogeneity (say $Y_{2}=X$), the model becomes the nonparametric
quantile regression
\begin{equation*}
Y_{1}=h_{0}(Y_{2})+U,\text{\quad }\Pr (U\leq 0|Y_{2})=\gamma ,
\end{equation*}
and the sieve variance becomes $||v_{n}^{\ast }||_{sd,ex}^{2}=\left( \frac{
d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]\right) ^{\prime }D_{n,ex}^{-}\left(
\frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot )]\right) $ with $D_{n,ex}=\frac{1}{
\gamma (1-\gamma )}E\left[ \{f_{U|Y_{2}}(0)\}^{2}\{q^{k(n)}(Y_{2})\}
\{q^{k(n)}(Y_{2})\}^{\prime }\right] $. Again $D_{n}\leq D_{n,ex}$ and $
||v_{n}^{\ast }||_{sd}^{2}\geq ||v_{n}^{\ast }||_{sd,ex}^{2}$. Under mild
conditions (see, e.g., \cite{CP_WP07}, \cite{CCLN_WP10}), $\lambda _{\min
}(D_{n})\rightarrow 0$ while $\lambda _{\min }(D_{n,ex})$ stays strictly
positive as $k(n)\rightarrow \infty $. All of the above discussions for a
functional $\phi (h)$ of the NPIV (\ref{npiv}) now apply to the functional
of the NPQIV (\ref{npqiv}). In particular, a functional $\phi (h)$ could be
root-$n$ estimable for the nonparametric quantile regression ($
\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||_{sd,ex}^{2}<\infty $) but
irregular for the NPQIV (\ref{npqiv}) ($\lim_{k(n)\rightarrow \infty
}||v_{n}^{\ast }||_{sd}^{2}=\infty $). See Section \ref{sec-ex} for details.
By Theorems \ref{thm:chi2}(2) and \ref{thm:QLR-H1}, the optimally weighted
SQLR statistic $\widehat{QLR}_{n}^{0}(\phi _{0})\Rightarrow \chi _{1}^{2}$
under the null of $\phi (h_{0})=\phi _{0}$, and diverges to infinity under
the alternative of $\phi (h_{0})\neq \phi _{0}$. We can compute confidence
set for a functional $\phi (h)$, such as an evaluation or a weighted
derivative functional, as $\left\{ r\in \mathbb{R}\colon \text{ }\widehat{QLR
}_{n}^{0}(r)\leq c_{\chi _{1}^{2}}(\tau )\right\} $. See Subsection \ref
{sec:application} for an empirical illustration of this result to the NPQIV
Engel curve regression using the British Family Survey data set that was
first used in \cite{BCK_Emetrica07}. Instead of using the asymptotic
critical values, we could also construct a confidence set using the
bootstrap critical values as in (\ref{boot-cs}).
\section{Basic Regularity Conditions}
\label{sec:conditions}
Before we establish asymptotic properties of sieve t (Wald) and SQLR
statistics, we need to present three sets of basic regularity conditions.
The first set of assumptions allows us to establish the convergence rates of
the PSMD estimator $\widehat{\alpha }_{n}$ to the true parameter value $
\alpha _{0}$ in both weak and strong metrics, which in turn allows us to
concentrate on some shrinking neighborhood of $\alpha _{0}$ in the
semi/nonparametric model (\ref{semi00}). The second and third regularity
conditions are respectively about the local curvatures of the functional $
\phi ()$ and of the criterion function under these two metrics. The weak
metric $||\cdot ||$ is closely related to the variance of the linear
approximation to $\phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})$, while
the strong metric $||\cdot ||_{s}$ is used to control the nonlinearity (in $
\alpha $) of the functional $\phi ()$ and of the conditional mean function $
m(x,\alpha )$. This section is mostly technical and applied researchers
could skip this and directly go to the subsequent sections on the asymptotic
properties of sieve Wald and SQLR statistics.
\subsection{A brief discussion on the convergence rate of the PSMD estimator
\label{sec:consistency}}
For the purely nonparametric conditional moment model $E\left[ \rho
(Y,X;h_{0}(\cdot ))|X\right] =0$, \cite{CP_WP07} established the consistency
and the convergence rates of their various PSMD estimators of $h_{0}$. Their
results can be trivially extended to establish the corresponding properties
of our PSMD estimator $\widehat{\alpha }_{n}\equiv (\widehat{\theta }
_{n}^{\prime },\widehat{h}_{n})$ defined in (\ref{psmd}). For the sake of
easy reference and to introduce basic assumptions and notation, we present
some sufficient conditions for consistency and the convergence rate here.
These conditions are also needed to establish the consistency and the
convergence rate of bootstrap PSMD estimators (see Lemma \ref{lem:cons_boot}
). We first impose three conditions on identification, sieve spaces, penalty
functions and sample criterion function. We equip the parameter space $
\mathcal{A}\equiv \Theta \times \mathcal{H}$ with a (strong) norm $
\left\Vert \alpha \right\Vert _{s}\equiv \left\Vert \theta \right\Vert
_{e}+\left\Vert h\right\Vert _{\mathbf{H}}$.
\begin{assumption}[Identification, sieves, criterion]
\label{ass:sieve} (i) $E[\rho (Y,X;\alpha )|X]=0$ if and only if $\alpha \in
(\mathcal{A},\left\Vert \cdot \right\Vert _{s})$ with $\left\Vert \alpha
-\alpha _{0}\right\Vert _{s}=0$; (ii) For all $k\geq 1$, $\mathcal{A}
_{k}\equiv \Theta \times \mathcal{H}_{k}$, $\Theta $ is a compact subset in $
\mathbb{R}^{d_{\theta }}$ with a non-empty interior, $\{\mathcal{H}
_{k}:k\geq 1\}$ is a non-decreasing sequence of non-empty closed linear
subsets of a Banach space $\left( \mathcal{H},\left\Vert \cdot \right\Vert _{
\mathbf{H}}\right) $ such that $\mathcal{H}=cl\left( \cup _{k}\mathcal{H}
_{k}\right) $, and there is $\Pi _{n}h_{0}\in \mathcal{H}_{k(n)}$ with $
||\Pi _{n}h_{0}-h_{0}||_{\mathbf{H}}=o(1)$; (iii) $Q:(\mathcal{A},\left\Vert
\cdot \right\Vert _{s})\rightarrow \lbrack 0,\infty )$ is lower
semicontinuous;\footnote{
A function $Q$ is lower semicontinuous at a point $\alpha _{o}\in \mathcal{A}
$ iff $\lim_{\left\Vert \alpha -\alpha _{o}\right\Vert _{s}\rightarrow
0}Q(\alpha )\geq Q(\alpha _{o})$; is lower semicontinuous if it is lower
semicontinuous at any point in $\mathcal{A}$.} (iv) $\Sigma (x)$ and $\Sigma
_{0}(x)$ are positive definite, and their smallest and largest eigenvalues
are finite and positive uniformly in $x\in \mathcal{X}$.
\end{assumption}
\begin{assumption}[Penalty]
\label{A_3.6} (i) $\lambda _{n}>0$, $Q(\Pi _{n}\alpha
_{0})+o(n^{-1})=O(\lambda _{n})=o(1)$; (ii) $|Pen(\Pi
_{n}h_{0})-Pen(h_{0})|=O(1)$ with $Pen(h_{0})<\infty $; (iii) $Pen:(\mathcal{
H},\left\Vert \cdot \right\Vert _{\mathbf{H}})\rightarrow \lbrack 0,\infty )$
is lower semicompact.\footnote{
A function $Pen$ is lower semicompact iff for all $M$, $\{h\in \mathcal{H}
\colon Pen(h)\leq M\}$ is a compact subset in $(\mathcal{H},\left\Vert \cdot
\right\Vert _{\mathbf{H}})$.}
\end{assumption}
Let $\Pi _{n}\alpha \equiv (\theta ^{\prime },\Pi _{n}h)\in \mathcal{A}
_{k(n)}\equiv \Theta \times \mathcal{H}_{k(n)}$. Let $\mathcal{A}
_{k(n)}^{M_{0}}\equiv \Theta \times \mathcal{H}_{k(n)}^{M_{0}}\equiv
\{\alpha =(\theta ^{\prime },h)\in \mathcal{A}_{k(n)}:\lambda _{n}Pen(h)\leq
\lambda _{n}M_{0}\}$ for a large but finite $M_{0}$ such that $\Pi
_{n}\alpha _{0}\in \mathcal{A}_{k(n)}^{M_{0}}$ and that $\widehat{\alpha }
_{n}\in \mathcal{A}_{k(n)}^{M_{0}}$ with probability arbitrarily close to
one for all large $n$. Let $\{\bar{\delta}_{m,n}^{2}\}_{n=1}^{\infty }$ be a
sequence of positive real values that decrease to zero as $n\rightarrow
\infty $.
\begin{assumption}[Sample Criterion]
\label{ass:rates} (i) $\widehat{Q}_{n}(\Pi _{n}\alpha _{0})\leq c_{0}Q(\Pi
_{n}\alpha _{0})+o_{P_{Z^{\infty }}}(n^{-1})$ for a finite constant $c_{0}>0$
; (ii) $\widehat{Q}_{n}(\alpha )\geq cQ(\alpha )-O_{P_{Z^{\infty }}}(\bar{
\delta}_{m,n}^{2})$ uniformly over $\mathcal{A}_{k(n)}^{M_{0}}$ for some $
\bar{\delta}_{m,n}^{2}=o(1)$ and a finite constant $c>0$.
\end{assumption}
The following result is a minor modification of Theorem 3.2 of \cite{CP_WP07}
.
\begin{lemma}
\label{thm:Thm-ill-suff_pencompact3} Let $\widehat{\alpha }_{n}$ be the PSMD
estimator defined in (\ref{psmd}), and Assumptions \ref{ass:sieve}, \ref
{A_3.6} and \ref{ass:rates} hold. Then: $||\widehat{\alpha }_{n}-\alpha
_{0}||_{s}=o_{P_{Z^{\infty }}}(1)$ and $Pen(\widehat{h}_{n})=O_{P_{Z^{\infty
}}}(1)$.
\end{lemma}
Given the consistency result, the PSMD estimator belongs to any $||\cdot
||_{s}-$neighborhood around $\alpha _{0}$ wpa1. We can restrict our
attention to a convex, $||\cdot ||_{s}-$neighborhood around $\alpha _{0}$,
denoted as $\mathcal{A}_{os}$ such that
\begin{equation*}
\mathcal{A}_{os}\subset \{\alpha \in \mathcal{A}:||\alpha -\alpha
_{0}||_{s}<M_{0},\text{ }\lambda _{n}Pen(h)<\lambda _{n}M_{0}\}
\end{equation*}
for a positive finite constant $M_{0}$ (the existence of a convex $\mathcal{A
}_{os}$ is implied by the convexity of $\mathcal{A}$ and quasi-convexity of $
Pen(\cdot )$). For any $\alpha \in \mathcal{A}_{os}$ we define a pathwise
derivative as
\begin{eqnarray*}
\frac{dm(X,\alpha _{0})}{d\alpha }[\alpha -\alpha _{0}] &\equiv &\left.
\frac{dE[\rho (Z,(1-\tau )\alpha _{0}+\tau \alpha )|X]}{d\tau }\right\vert
_{\tau =0}\quad a.s.~X \\
&=&\frac{dE[\rho (Z,\alpha _{0})|X]}{d\theta ^{\prime }}(\theta -\theta
_{0})\\
&& +\frac{dE[\rho (Z,\alpha _{0})|X]}{dh}[h-h_{0}]\quad a.s.~X.
\end{eqnarray*}
Following \cite{AC_Emetrica03} and \cite{CP_WP07a}, we introduce two
pseudo-metrics $||\cdot ||$ and $||\cdot ||_{0}$ on $\mathcal{A}_{os}$ as:
for any $\alpha _{1},$ $\alpha _{2}\in \mathcal{A}_{os}$,
\begin{equation}
||\alpha _{1}-\alpha _{2}||^{2}\equiv E\left[ \left( \frac{
dm(X,\alpha _{0})}{d\alpha }[\alpha _{1}-\alpha _{2}]\right) ^{\prime }\Sigma (X)^{-1}\left( \frac{dm(X,\alpha _{0})}{d\alpha }
[\alpha _{1}-\alpha _{2}]\right) \right] ; \label{fmetric}
\end{equation}
\begin{equation}
||\alpha _{1}-\alpha _{2}||_{0}^{2}\equiv E\left[ \left(
\frac{dm(X,\alpha _{0})}{d\alpha }[\alpha _{1}-\alpha _{2}]\right) ^{\prime }\Sigma _{0}(X)^{-1}\left( \frac{dm(X,\alpha _{0})}{
d\alpha }[\alpha _{1}-\alpha _{2}]\right) \right] .
\label{fmetric0}
\end{equation}
It is clear that, under Assumption \ref{ass:sieve}(iv), these two
pseudo-metrics are equivalent, i.e., $||\cdot ||\asymp ||\cdot ||_{0}$ on $
\mathcal{A}_{os}$. This is why Assumption \ref{ass:sieve}(iv) is imposed
throughout the paper.
Let $\mathcal{A}_{osn}=\mathcal{A}_{os}\cap \mathcal{A}_{k(n)}$. Let $
\{\delta _{n}\}_{n=1}^{\infty }$ be a sequence of positive real values such
that $\delta _{n}=o(1)$ and $\delta _{n}\leq \bar{\delta}_{m,n}$.
\begin{assumption}
\label{ass:weak_equiv} (i) There exists a convex $||\cdot ||_{s}-$
neighborhood of $\alpha _{0}$, $\mathcal{A}_{os}$, such that $m(\cdot
,\alpha )$ is continuously pathwise differentiable with respect to $\alpha
\in \mathcal{A}_{os}$, and there is a finite constant $C>0$ such that $
||\alpha -\alpha _{0}||\leq C||\alpha -\alpha _{0}||_{s}$ for all $\alpha
\in \mathcal{A}_{os}$; (ii) $Q(\alpha )\asymp ||\alpha -\alpha _{0}||^{2}$
for all $\alpha \in \mathcal{A}_{os}$; (iii) $\widehat{Q}_{n}(\alpha )\geq
cQ(\alpha )-O_{P_{Z^{\infty }}}(\delta _{n}^{2})$ uniformly over $\mathcal{A}
_{osn}$, and $\max \{\delta _{n}^{2},Q(\Pi _{n}\alpha _{0}),\lambda
_{n},o(n^{-1})\}=\delta _{n}^{2}$; (iv) $\lambda _{n}\times \sup_{\alpha
,\alpha ^{\prime }\in \mathcal{A}_{os}}\left\vert Pen(h)-Pen(h^{\prime
})\right\vert =o(n^{-1})$ or $\lambda _{n}=o(n^{-1})$.
\end{assumption}
Assumption \ref{ass:weak_equiv}(ii) is about the local curvature of the
population criterion $Q(\alpha )$ at $\alpha _{0}$. It can be weakened to Assumption 4.1(ii) in \cite{CP_WP07}. When $\widehat{Q}
_{n}(\alpha )$ is computed using the series LS estimator (\ref{mhat}), Lemma
C.2 of \cite{CP_WP07} shows that $\widehat{Q}_{n}(\alpha )\asymp Q(\alpha
)-O_{P_{Z^{\infty }}}(\delta _{n}^{2})$ uniformly over $\mathcal{A}_{osn}$
and hence Assumption \ref{ass:weak_equiv}(iii) is satisfied.
Recall the definition of the \textit{sieve measure of local ill-posedness}
\begin{equation}
\tau _{n}\equiv \sup_{\alpha \in \mathcal{A}_{osn}:||\alpha -\Pi _{n}\alpha
_{0}||\neq 0}\frac{||\alpha -\Pi _{n}\alpha _{0}||_{s}}{||\alpha -\Pi
_{n}\alpha _{0}||}. \label{tau-n}
\end{equation}
The problem of estimating $\alpha _{0}$ under $||\cdot ||_{s}$ is \textit{
locally} \textit{ill-posed in rate} if and only if $\limsup_{n\rightarrow
\infty }\tau _{n}=\infty $. We say the problem is \textit{mildly ill-posed}
if $\tau _{n}=O([k(n)]^{a})$, and \textit{severely ill-posed} if $\tau
_{n}=O(\exp \{\frac{a}{2}k(n)\})$ for some finite $a>0$. The following
general rate result is a minor modification of Theorem 4.1 and Remark 4.1(i)
of \cite{CP_WP07}, and hence we omit its proof.
\begin{lemma}
\label{thm:ThmCONVRATEGRAL} Let $\widehat{\alpha }_{n}$ be the PSMD
estimator defined in (\ref{psmd}), and Assumptions \ref{ass:sieve}, \ref
{A_3.6}(ii)(iii), \ref{ass:rates} and \ref{ass:weak_equiv}(i)(ii)(iii) hold.
Then:
\begin{equation*}
||\widehat{\alpha }_{n}-\alpha _{0}||=O_{P_{Z^{\infty }}}\left( \delta
_{n}\right) \quad \text{and}\quad ||\widehat{\alpha }_{n}-\alpha
_{0}||_{s}=O_{P_{Z^{\infty }}}\left( ||\alpha _{0}-\Pi _{n}\alpha
_{0}||_{s}+\tau _{n}\delta _{n}\right) .
\end{equation*}
\end{lemma}
The above convergence rate result is applicable to any nonparametric
estimator $\widehat{m}(X,\alpha )$ of $m(X,\alpha )$ as soon as one could
compute $\delta _{n}^{2}$, the rate at which $\widehat{Q}_{n}(\alpha )$ goes
to $Q(\alpha )$. See \cite{CP_WP07} and \cite{CP_WP07a} for low level
sufficient conditions in terms of the series LS estimator (\ref{mhat}) of $
m(X,\alpha )$.
Let $\left\{ \delta _{s,n}:n\geq 1\right\} $ be a sequence of real positive
numbers such that $\delta _{s,n}=||h_{0}-\Pi _{n}h_{0}||_{s}+\tau _{n}\delta
_{n}=o(1)$. Lemma \ref{thm:ThmCONVRATEGRAL} implies that $\widehat{\alpha }
_{n}\in \mathcal{N}_{osn}\subseteq \mathcal{N}_{os}$ wpa1-$P_{Z^{\infty }}$,
where
\begin{eqnarray*}
\mathcal{N}_{os} &\equiv &\left\{ \alpha \in \mathcal{A}\colon \text{ }
||\alpha -\alpha _{0}||\leq M_{n}\delta _{n},\text{ }||\alpha -\alpha
_{0}||_{s}\leq M_{n}\delta _{s,n},\text{ }\lambda _{n}Pen(h)\leq \lambda
_{n}M_{0}\right\} , \\
\mathcal{N}_{osn} &\equiv &\mathcal{N}_{os}\cap \mathcal{A}_{k(n)},\quad
\text{with }M_{n}\equiv \min \left\{ \log (\log (n+1)),\log ((\delta
_{s,n}^{-1}+1))\right\} .
\end{eqnarray*}
We can regard $\mathcal{N}_{os}$ as the effective parameter space and $
\mathcal{N}_{osn}$ as its sieve space in the rest of the paper. Assumption
\ref{ass:weak_equiv}(iv) is not needed for establishing a convergence rate
in Lemma \ref{thm:ThmCONVRATEGRAL}. but, it will be imposed in the rest of
the paper so that we can ignore penalty effect in the first order local
asymptotic analysis.
\subsection{(Sieve) Riesz representation and (sieve) variance}
\label{sec:phi}
We first introduce a representation of the functional of interest $\phi ()$
at $\alpha _{0}$ that is crucial for all the subsequent local asymptotic
theories. Let $\phi :\mathbb{R}^{d_{\theta }}\times \mathcal{H}\rightarrow
\mathbb{R}$ be continuous in $||\cdot ||_{s}$. We assume that $\frac{d\phi
(\alpha _{0})}{d\alpha }[\cdot ]:\left( \mathbb{R}^{d_{\theta }}\times
\mathcal{H},||\cdot ||_{s}\right) \rightarrow \mathbb{R}$ is a $||\cdot
||_{s}-$bounded linear functional (i.e., $\left\vert \frac{d\phi (\alpha
_{0})}{d\alpha }[v]\right\vert \leq c||v||_{s}$ uniformly over $v\in \mathbb{
R}^{d_{\theta }}\times \mathcal{H}$ for a finite positive constant $c$),
which could be computed as a pathwise\ (directional) derivative of the
functional $\phi \left( \cdot \right) $ at $\alpha _{0}$ in the direction of
$v=\alpha -\alpha _{0}\in \mathbb{R}^{d_{\theta }}\times \mathcal{H}:$
\begin{equation*}
\frac{d\phi (\alpha _{0})}{d\alpha }\left[ v\right] =\left. \frac{\partial
\phi (\alpha _{0}+\tau v)}{\partial \tau }\right\vert _{\tau =0}.
\end{equation*}
Let $\mathbf{V}$ be a linear span of $\mathcal{A}_{os}-\{\alpha _{0}\}$,
which is endowed with both $||\cdot ||_{s}$ and $||\cdot ||$ (in equation (
\ref{fmetric})) norms, and $||v||\leq C||v||_{s}$ for all $v\in \mathbf{V}$
(under Assumption \ref{ass:weak_equiv}(i)). Let $\overline{\mathbf{V}}\equiv
clsp(\mathcal{A}_{os}-\{\alpha _{0}\})$, where $clsp(\cdot )$ is the closure
of the linear span under $||\cdot ||$. For any $v_{1},v_{2}\in \overline{
\mathbf{V}}$, we define an inner product induced by the metric $||\cdot ||$:
\begin{equation*}
\left\langle v_{1},v_{2}\right\rangle =E\left[ \left( \frac{dm(X,\alpha _{0})
}{d\alpha }[v_{1}]\right) ^{\prime }\Sigma (X)^{-1}\left( \frac{dm(X,\alpha
_{0})}{d\alpha }[v_{2}]\right) \right] ,
\end{equation*}
and for any $v\in \overline{\mathbf{V}}$ we call $v=0$ if and only if $
||v||=0$ (i.e., functions in $\overline{\mathbf{V}}$ are defined in an
equivalent class sense according to the metric $||\cdot ||$). It is clear
that $(\overline{\mathbf{V}},||\cdot ||)$ is an infinite dimensional Hilbert
space (under Assumptions \ref{ass:sieve}(i)(iii)(iv) and \ref{ass:weak_equiv}
(i)(ii)).
If the linear functional $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is
\textit{bounded} on $(\mathbf{V},||\cdot ||)$, i.e.
\begin{equation*}
\sup_{v\in \mathbf{V},v\neq 0}\frac{\left\vert \frac{d\phi (\alpha _{0})}{
d\alpha }\left[ v\right] \right\vert }{\left\Vert v\right\Vert }<\infty ,
\end{equation*}
then there is a unique extension of\ $\frac{d\phi (\alpha _{0})}{d\alpha }
[\cdot ]$ from $(\mathbf{V},||\cdot ||)$ to $(\overline{\mathbf{V}},||\cdot
||)$, and a unique Riesz representer $v^{\ast }\in \overline{\mathbf{V}}$ of
$\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ on $(\overline{\mathbf{V}}
,||\cdot ||)$ such that\footnote{
See, e.g., page 206-207 and theorem 3.10.1 in \cite{Debnath-Hilbert}.}
\begin{align}\label{RRT-0}
&\frac{d\phi (\alpha _{0})}{d\alpha }\left[ v\right] =\left\langle v^{\ast
},v\right\rangle \text{ for all }v\in \overline{\mathbf{V}}\text{\quad
and\quad }\\ \notag
&\left\Vert v^{\ast }\right\Vert \equiv \sup_{v\in \overline{
\mathbf{V}},v\neq 0}\frac{\left\vert \frac{d\phi (\alpha _{0})}{d\alpha }
\left[ v\right] \right\vert }{\left\Vert v\right\Vert }=\sup_{v\in \mathbf{V}
,v\neq 0}\frac{\left\vert \frac{d\phi (\alpha _{0})}{d\alpha }\left[ v\right]
\right\vert }{\left\Vert v\right\Vert }<\infty .
\end{align}
If $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is \textit{unbounded} on $(
\mathbf{V},||\cdot ||)$, i.e.
\begin{equation*}
\sup_{v\in \mathbf{V},v\neq 0}\frac{\left\vert \frac{d\phi (\alpha _{0})}{
d\alpha }\left[ v\right] \right\vert }{\left\Vert v\right\Vert }=\infty ,
\end{equation*}
then there is no unique extension of the mapping\ $\frac{d\phi (\alpha _{0})
}{d\alpha }[\cdot ]$ from $(\mathbf{V},||\cdot ||)$ to $(\overline{\mathbf{V}
},||\cdot ||)$, and nor existing any Riesz representer of\ $\frac{d\phi
(\alpha _{0})}{d\alpha }[\cdot ]$ on $(\overline{\mathbf{V}},||\cdot ||)$.
Since $||v||\leq C||v||_{s}$ for all $v\in \mathbf{V}$, it is clear that a $
||\cdot ||_{s}-$bounded linear functional $\frac{d\phi (\alpha _{0})}{
d\alpha }[\cdot ]$ could be either bounded or unbounded on $(\mathbf{V}
,||\cdot ||)$.
\textbf{Sieve Riesz representation}. Let $\alpha _{0,n}\in \mathbb{R}
^{d_{\theta }}\times \mathcal{H}_{k(n)}$ be such that
\begin{equation}
||\alpha _{0,n}-\alpha _{0}||\equiv \min_{\alpha \in \mathbb{R}^{d_{\theta
}}\times \mathcal{H}_{k(n)}}||\alpha -\alpha _{0}||. \label{SP-1}
\end{equation}
Let $\overline{\mathbf{V}}_{k(n)}\equiv clsp\left( \mathcal{A}
_{osn}-\{\alpha _{0,n}\}\right) $, where $clsp\left( .\right) $ denotes the
closed linear span under $\left\Vert \cdot \right\Vert $. Then $\overline{
\mathbf{V}}_{k(n)}$ is a finite dimensional Hilbert space under $\left\Vert
\cdot \right\Vert $. Moreover, $\overline{\mathbf{V}}_{k(n)}$ is dense in $
\overline{\mathbf{V}}$ under $\left\Vert \cdot \right\Vert $. To simplify
the presentation, we assume that $\dim (\overline{\mathbf{V}}_{k(n)})=\dim (
\mathcal{A}_{k(n)})\asymp k(n)$, all of which grow to infinity with $n$. By
definition we have $\left\langle v_{n},\alpha _{0,n}-\alpha
_{0}\right\rangle =0$ for all $v_{n}\in \overline{\mathbf{V}}_{k(n)}$.
Note that $\overline{\mathbf{V}}_{k(n)}$ is a finite dimensional Hilbert
space. As any linear functional on a finite dimensional Hilbert space is
bounded, we can invoke the Riesz representation theorem to deduce that there
is a $v_{n}^{\ast }\in \overline{\mathbf{V}}_{k(n)}$ such that
\begin{equation}
\frac{d\phi (\alpha _{0})}{d\alpha }[v]=\left\langle v_{n}^{\ast
},v\right\rangle,~\forall v\in \overline{\mathbf{V}}_{k(n)},~and~\left\Vert v_{n}^{\ast }\right\Vert \equiv \sup_{v\in
\overline{\mathbf{V}}_{k(n)}:\left\Vert v\right\Vert \neq 0}\frac{\left\vert
\frac{d\phi (\alpha _{0})}{d\alpha }[v]\right\vert }{\left\Vert v\right\Vert
}<\infty . \label{RRT-1}
\end{equation}
We call $v_{n}^{\ast }$ the \textit{sieve Riesz representer} of the
functional\ $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ on $\overline{
\mathbf{V}}_{k(n)}$. By definition, for any non-zero linear functional $
\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$, we have:
\begin{equation*}
0<\left\Vert v_{n}^{\ast }\right\Vert ^{2}=E\left[ \left( \frac{dm(X,\alpha
_{0})}{d\alpha }[v_{n}^{\ast }]\right) ^{\prime }\Sigma (X)^{-1}\left( \frac{
dm(X,\alpha _{0})}{d\alpha }[v_{n}^{\ast }]\right) \right]
\end{equation*}
is non-decreasing in $k(n)$.
We emphasize that the sieve Riesz representer $v_{n}^{\ast }$ of a linear
functional\ $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ on $\overline{
\mathbf{V}}_{k(n)}$ always exists regardless of whether $\frac{d\phi (\alpha
_{0})}{d\alpha }[\cdot ]$ is bounded on the infinite dimensional space $(
\mathbf{V},||\cdot ||)$ or not. Moreover, $v_{n}^{\ast }\in \overline{
\mathbf{V}}_{k(n)}$ and its norm $\left\Vert v_{n}^{\ast }\right\Vert $ can
be computed in closed form (see Subsection \ref{closed_form_sieve_variance}
). The next Lemma allows us to verify whether or not $\frac{d\phi (\alpha
_{0})}{d\alpha }[\cdot ]$ is bounded on $(\mathbf{V},||\cdot ||)$ by
checking whether or not $\lim_{k(n)\rightarrow \infty }\left\Vert
v_{n}^{\ast }\right\Vert <\infty $.
\begin{lemma}
\label{lem:sieve-Riesz-property} Let $\{\overline{\mathbf{V}}
_{k}\}_{k=1}^{\infty }$ be an increasing sequence of finite dimensional
Hilbert spaces that is dense in $(\overline{\mathbf{V}},\left\Vert \cdot
\right\Vert )$, and $v_{n}^{\ast }\in \overline{\mathbf{V}}_{k(n)}$ be
defined in (\ref{RRT-1}). (1) If $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot
]$ is bounded on $(\mathbf{V},||\cdot ||)$, then (\ref{RRT-0}) holds, $
v_{n}^{\ast }=\arg \min_{v\in \overline{\mathbf{V}}_{k(n)}}\left\Vert
v^{\ast }-v\right\Vert $ and $\left\Vert v^{\ast }-v_{n}^{\ast }\right\Vert
\rightarrow 0$, $\lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast
}\right\Vert =\left\Vert v^{\ast }\right\Vert <\infty $; (2) Let $\frac{
d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ be bounded on $(\mathbf{V},||\cdot
||_{s})$ and $\{\overline{\mathbf{V}}_{k}\}_{k=1}^{\infty }$ be dense in $(
\mathbf{V},\left\Vert \cdot \right\Vert _{s})$. If $\frac{d\phi (\alpha _{0})
}{d\alpha }[\cdot ]$ is unbounded on $(\mathbf{V},||\cdot ||)$ then $
\lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert =\infty $.
\end{lemma}
\textbf{Sieve score and sieve variance}. For each sieve dimension $k(n)$, we
call
\begin{equation}
S_{n,i}^{\ast }\equiv \left( \frac{dm(X_{i},\alpha _{0})}{d\alpha }
[v_{n}^{\ast }]\right) ^{\prime }\Sigma (X_{i})^{-1}\rho (Z_{i},\alpha _{0})
\label{score}
\end{equation}
the \textit{sieve score} associated with the $i$-th observation, and $
\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}\equiv Var\left( S_{n,i}^{\ast
}\right) $ as the \textit{sieve variance}. Recall that $\Sigma _{0}(X)\equiv
Var(\rho (Z;\alpha _{0})|X)$ a.s.-$X$. Then
\begin{align} \label{svar}
\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2} &= E[S_{n,i}^{\ast }S_{n,i}^{\ast
\prime }]\\ \notag
& = E\left[ \left( \frac{dm(X,\alpha _{0})}{d\alpha }[v_{n}^{\ast
}]\right) ^{\prime }\Sigma (X)^{-1}\Sigma _{0}(X)\Sigma (X)^{-1}\left( \frac{
dm(X,\alpha _{0})}{d\alpha }[v_{n}^{\ast }]\right) \right] .
\end{align}
(See Subsection \ref{closed_form_sieve_variance} for closed form expressions
of $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}$.) Under Assumption \ref
{ass:sieve}(iv), we have $\left\Vert v_{n}^{\ast }\right\Vert
_{sd}^{2}\asymp \left\Vert v_{n}^{\ast }\right\Vert ^{2}$, and hence $
\lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert
_{sd}<\infty $ (or $=\infty $) iff $\lim_{k(n)\rightarrow \infty }\left\Vert
v_{n}^{\ast }\right\Vert <\infty $ (or $=\infty $). Therefore, in this paper
we call $\phi ()$ \textit{regular} (or \textit{irregular}) at $\alpha _{0}$
whenever $\lim_{k(n)\rightarrow \infty }\left\Vert v_{n}^{\ast }\right\Vert
<\infty $ (or $=\infty $), which, by Lemma \ref{lem:sieve-Riesz-property},
is also whenever $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot ]$ is bounded
(or unbounded) on $(\mathbf{V},||\cdot ||)$. It is clear that our notion of
a regular $\phi \left( \cdot \right) $ at $\alpha _{0}$ is only necessary
but not sufficient for the existence of root-$n$ asymptotically normal
regular estimators of $\phi \left( \alpha _{0}\right) $. Moreover, if $\phi
\left( \cdot \right) $ is regular at $\alpha _{0}$ then we can define
\begin{equation*}
S_{i}^{\ast }\equiv \left( \frac{dm(X_{i},\alpha _{0})}{d\alpha }[v^{\ast
}]\right) ^{\prime }\Sigma (X_{i})^{-1}\rho (Z_{i},\alpha _{0})
\end{equation*}
as the \textit{score} associated with the $i$-th observation, and $
\left\Vert v^{\ast }\right\Vert _{sd}^{2}\equiv Var\left( S_{i}^{\ast
}\right) $ as the \textit{asymptotic variance}. By Lemma \ref
{lem:sieve-Riesz-property}(1) for a regular functional we have: $\left\Vert
v^{\ast }\right\Vert _{sd}^{2}\asymp \left\Vert v^{\ast }\right\Vert <\infty
$ and $Var\left( S_{i}^{\ast }-S_{n,i}^{\ast }\right) \asymp \left\Vert
v^{\ast }-v_{n}^{\ast }\right\Vert ^{2}\rightarrow 0$ as $k(n)\rightarrow
\infty $. See Appendix \ref{app:appA} for further discussions.
\subsection{Two key local conditions}
\label{sec:LAQ}
For all $k(n)$, let
\begin{equation}
u_{n}^{\ast }\equiv \frac{v_{n}^{\ast }}{\left\Vert v_{n}^{\ast }\right\Vert
_{sd}} \label{u*}
\end{equation}
be the \textquotedblleft scaled sieve Riesz representer\textquotedblright .
Since $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}\asymp \left\Vert
v_{n}^{\ast }\right\Vert ^{2}$ (under Assumption \ref{ass:sieve}(iv)), we
have: $\left\Vert u_{n}^{\ast }\right\Vert \asymp 1$ and $\left\Vert
u_{n}^{\ast }\right\Vert _{s}\leq c\tau _{n}$ for $\tau _{n}$ defined in (
\ref{tau-n}) and a finite constant $c>0$.
Let $\mathcal{T}_{n}\equiv \{t\in \mathbb{R}\colon |t|\leq 4M_{n}^{2}\delta
_{n}\}$ with $M_{n}$ and $\delta _{n}$ given in the definition of $\mathcal{N
}_{osn}$.
\begin{assumption}[Local behavior of $\protect\phi $]
\label{ass:phi} (i) $v\mapsto \frac{d\phi (\alpha _{0})}{d\alpha }[v]$ is a
non-zero linear functional mapping from $\mathbf{V}$ to $\mathbb{R}$; $\{
\overline{\mathbf{V}}_{k}\}_{k=1}^{\infty }$ is an increasing sequence of
finite dimensional Hilbert spaces that is dense in $(\overline{\mathbf{V}}
,\left\Vert \cdot \right\Vert )$; and $\frac{\left\Vert v_{n}^{\ast
}\right\Vert }{\sqrt{n}}=o(1)$;
\begin{equation*}
\text{(ii)\quad }\sup_{(\alpha ,t)\in \mathcal{N}_{osn}\times \mathcal{T}
_{n}}\frac{\sqrt{n}\left\vert \phi \left( \alpha +tu_{n}^{\ast }\right)
-\phi (\alpha _{0})-\frac{d\phi (\alpha _{0})}{d\alpha }[\alpha
+tu_{n}^{\ast }-\alpha _{0}]\right\vert }{\left\Vert v_{n}^{\ast
}\right\Vert }=o\left( 1\right) ;
\end{equation*}
(iii) $\frac{\sqrt{n}\left\vert \frac{d\phi (\alpha _{0})}{d\alpha }[\alpha
_{0,n}-\alpha _{0}]\right\vert }{\left\Vert v_{n}^{\ast }\right\Vert }
=o\left( 1\right) .$
\end{assumption}
Since $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}\asymp \left\Vert
v_{n}^{\ast }\right\Vert ^{2}$ (under Assumption \ref{ass:sieve}(iv)), we
could rewrite Assumption \ref{ass:phi} using $\left\Vert v_{n}^{\ast
}\right\Vert _{sd}$ instead $\left\Vert v_{n}^{\ast }\right\Vert $. As it
will become clear in Theorem \ref{thm:theta_anorm} that $\frac{\left\Vert
v_{n}^{\ast }\right\Vert _{sd}^{2}}{n}$ is the variance of $\phi (\widehat{
\alpha }_{n})-\phi (\alpha _{0})$, Assumption \ref{ass:phi}(i) puts a
restriction on how fast the sieve dimension $k(n)$ could grow with the
sample size $n$.
Assumption \ref{ass:phi}(ii) controls the nonlinearity bias of $\phi \left(
\cdot \right) $ (i.e., the linear approximation error of a possibly
nonlinear functional $\phi \left( \cdot \right) $). It is automatically
satisfied when $\phi \left( \cdot \right) $ is a linear functional. For a
nonlinear functional $\phi \left( \cdot \right) $ (such as the quadratic
functional), it can be verified using the smoothness of $\phi \left( \cdot
\right) $ and the convergence rates in both $||\cdot ||$ and $||\cdot ||_{s}$
metrics (the definition of $\mathcal{N}_{osn}$). See Section \ref{sec-ex}
for verification.
Assumption \ref{ass:phi}(iii) controls the linear bias part due to the
finite dimensional sieve approximation of $\alpha _{0,n}$ to $\alpha _{0}$.
It is a condition imposed on the growth rate of the sieve dimension $k(n)$.
When $\phi \left( \cdot \right) $ is an irregular functional, we have $
\left\Vert v_{n}^{\ast }\right\Vert \nearrow \infty $. Assumption \ref
{ass:phi}(iii) requires that the sieve bias term, $\left\vert \frac{d\phi
(\alpha _{0})}{d\alpha }[\alpha _{0,n}-\alpha _{0}]\right\vert $, is of a
smaller order than that of the sieve standard deviation term, $
n^{-1/2}\left\Vert v_{n}^{\ast }\right\Vert _{sd}$. This is a standard
condition imposed for the asymptotic normality of any plug-in nonparametric
estimator of an irregular functional (such as a point evaluation functional
of a nonparametric mean regression).
\begin{remark}
\label{remark-bias} When $\phi \left( \cdot \right) $ is regular at $\alpha
_{0}$ (i.e., $\left\Vert v_{n}^{\ast }\right\Vert \nearrow \left\Vert
v^{\ast }\right\Vert <\infty $), since $\left\langle v_{n}^{\ast },\alpha
_{0,n}-\alpha _{0}\right\rangle =0$ (by definition of $\alpha _{0,n}$) we
have $\left\vert \frac{d\phi (\alpha _{0})}{d\alpha }[\alpha _{0,n}-\alpha
_{0}]\right\vert \leq \left\Vert v^{\ast }-v_{n}^{\ast }\right\Vert \times
\left\Vert \alpha _{0,n}-\alpha _{0}\right\Vert $. And Assumption \ref
{ass:phi}(iii) is satisfied if
\begin{equation}
||v^{\ast }-v_{n}^{\ast }||\times ||\alpha _{0,n}-\alpha _{0}||=o(n^{-1/2}).
\label{regular-orth}
\end{equation}
This is similar to assumption 4.2 in \cite{AC_Emetrica03} and assumption
3.2(iii) in \cite{CP_WP07a} for the root-$n$ estimable Euclidean parameter $
\theta _{0}$ of the model (\ref{semi00}). As pointed out by \cite{CP_WP07a},
Condition (\ref{regular-orth}) could be satisfied when $\dim (\mathcal{A}
_{k(n)})\asymp k(n)$ is chosen to obtain optimal nonparametric convergence
rate in $||\cdot ||_{s}$ norm. But this nice feature only applies to regular
functionals.
\end{remark}
The next assumption is about the local quadratic approximation (LQA) to the
sample criterion difference along the scaled sieve Riesz representer
direction $u_{n}^{\ast }=v_{n}^{\ast }/\left\Vert v_{n}^{\ast }\right\Vert
_{sd}$.
For any $(\alpha ,t)\in \mathcal{N}_{osn}\times \mathcal{T}_{n}$, we let $
\widehat{\Lambda }_{n}(\alpha (t),\alpha )\equiv 0.5\{\widehat{Q}_{n}(\alpha
(t))-\widehat{Q}_{n}(\alpha )\}$ with $\alpha (t)\equiv \alpha +tu_{n}^{\ast
}$. Denote
\begin{equation}
\mathbb{Z}_{n}\equiv n^{-1}\sum_{i=1}^{n}\left( \frac{dm(X_{i},\alpha _{0})}{
d\alpha }[u_{n}^{\ast }]\right) ^{\prime }\Sigma (X_{i})^{-1}\rho
(Z_{i},\alpha _{0})=n^{-1}\sum_{i=1}^{n}\frac{S_{n,i}^{\ast }}{\left\Vert
v_{n}^{\ast }\right\Vert _{sd}}. \label{effscore}
\end{equation}
\begin{assumption}[LQA]
\label{ass:LAQ} (i) $\alpha (t)\in \mathcal{A}_{k(n)}$ for any $(\alpha
,t)\in \mathcal{N}_{osn}\times \mathcal{T}_{n}$; and with $
r_{n}(t_{n})=\left( \max \{t_{n}^{2},t_{n}n^{-1/2},o(n^{-1})\}\right) ^{-1}$
,
\begin{equation*}
\sup_{(\alpha ,t_{n})\in \mathcal{N}_{osn}\times \mathcal{T}
_{n}}r_{n}(t_{n})\left\vert \widehat{\Lambda }_{n}(\alpha (t_{n}),\alpha
)-t_{n}\left\{ \mathbb{Z}_{n}+\langle u_{n}^{\ast },\alpha -\alpha
_{0}\rangle \right\} -\frac{B_{n}}{2}t_{n}^{2}\right\vert =o_{P_{Z^{\infty
}}}(1),
\end{equation*}
where, for each $n$, $B_{n}$ is a $Z^{n}$ measurable positive random
variable, and $B_{n}=O_{P_{Z^{\infty }}}(1)$;
(ii) $\sqrt{n}\mathbb{Z}_{n}\Rightarrow N(0,1)$.
\end{assumption}
Assumption \ref{ass:LAQ}(ii) is a standard one, and is implied by the
following Lindeberg condition: For all $\epsilon >0$,
\begin{equation}
\limsup_{n\rightarrow \infty }E\left[ \left( \frac{S_{n,i}^{\ast }}{
\left\Vert v_{n}^{\ast }\right\Vert _{sd}}\right) ^{2}1\left\{ \left\vert
\frac{S_{n,i}^{\ast }}{\epsilon \sqrt{n}\left\Vert v_{n}^{\ast }\right\Vert
_{sd}}\right\vert >1\right\} \right] =0, \label{LF}
\end{equation}
which, under Lemma \ref{lem:sieve-Riesz-property}(1) and Assumption \ref
{ass:sieve}(iv), is satisfied when the functional $\phi (\cdot )$ is regular
($\left\Vert v_{n}^{\ast }\right\Vert _{sd}\asymp \left\Vert v_{n}^{\ast
}\right\Vert \rightarrow \left\Vert v^{\ast }\right\Vert <\infty $). This is
why Assumption \ref{ass:LAQ}(ii) is not imposed in \cite{AC_Emetrica03} and
\cite{CP_WP07a} in their root-$n$ asymptotically normal estimation of the
regular functional $\phi (\alpha )=\lambda ^{\prime }\theta $.
Assumption \ref{ass:LAQ}(i) implicitly imposes restrictions on the
nonparametric estimator $\widehat{m}(x,\alpha )$ of $m(x,\alpha )=E[\rho
(Z,\alpha )|X=x]$ in a shrinking neighborhood of $\alpha _{0}$, so that the
criterion difference could be well approximated by a quadratic form. It is
trivially satisfied when $\widehat{m}(x,\alpha )$ is linear in $\alpha $,
such as the series LS estimator (\ref{mhat}) when $\rho (Z,\alpha )$ is
linear in $\alpha $. There are two potential difficulties in verifying this
assumption for nonlinear conditional moment models with nonparametric
endogeneity (such as the NPQIV\ model). First, due to the non-smooth
residual function $\rho (Z,\alpha )$, the estimator $\widehat{m}(x,\alpha )$
(and hence the sample criterion $\widehat{Q}_{n}(\alpha )$) could be
pointwise non-smooth with respect to $\alpha $. Second, due to the slow
convergence rates in the strong norm $||\cdot ||_{s}$ present in nonlinear
nonparametric ill-posed inverse problems, it could be challenging to control
the remainder of a quadratic approximation. When $\widehat{m}(x,\alpha )$ is
the series LS estimator (\ref{mhat}), Lemma \ref{lem:Qdiff_B} in Section \ref
{sec:bootstrap} shows that Assumption \ref{ass:LAQ}(i) is satisfied by a set
of relatively low level sufficient conditions (Assumptions \ref{ass:m_ls} -
\ref{ass:cont_diffm} in Appendix \ref{app:appA}). See Section \ref{sec-ex}
for verification of these sufficient conditions for functionals of the NPQIV
model.
\section{Asymptotic Properties of Sieve Wald and SQLR Statistics}
\label{sec:AsymDist}
In this section, we first establish the asymptotic normality of the plug-in
PSMD estimator $\phi (\widehat{\alpha }_{n})$ of $\phi (\alpha _{0})$ for
the model (\ref{semi00}), regardless of whether it is root-$n$ estimable or
not. We then provide a simple consistent variance estimator and hence the
asymptotic standard normality of the corresponding sieve t statistic for a
real-valued functional $\phi :\mathbb{R}^{d_{\theta }}\times \mathcal{H}
\rightarrow \mathbb{R}$. We finally derive the asymptotic properties of SQLR
tests for the hypothesis $\phi (\alpha _{0})=\phi _{0}$. See Appendix \ref
{app:appA} for the case of a vector-valued functional $\phi :\mathbb{R}
^{d_{\theta }}\times \mathcal{H}\rightarrow \mathbb{R}^{d_{\phi }}$ (where $
d_{\phi }$ could grow slowly with $n$).
\subsection{Asymptotic normality of the plug-in PSMD estimator}
\label{sec:anormality}
The next result allows for a (possibly) nonlinear irregular functional $\phi
()$ of the general model (\ref{semi00}).
\begin{theorem}
\label{thm:theta_anorm} Let $\widehat{\alpha }_{n}$ be the PSMD estimator (
\ref{psmd}) and Assumptions \ref{ass:sieve} - \ref{ass:weak_equiv} hold. If
Assumptions \ref{ass:phi} and \ref{ass:LAQ} hold, then:
\begin{equation*}
\sqrt{n}\frac{\phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})}{||v_{n}^{\ast
}||_{sd}}=-\sqrt{n}\mathbb{Z}_{n}+o_{P_{Z^{\infty }}}(1)\Rightarrow N(0,1).
\end{equation*}
\end{theorem}
When the functional $\phi (\cdot )$ is regular at $\alpha =\alpha _{0}$, we
have $\left\Vert v_{n}^{\ast }\right\Vert _{sd}\asymp \left\Vert v_{n}^{\ast
}\right\Vert =O(1)$ and $\phi (\widehat{\alpha }_{n})$ converges to $\phi
(\alpha _{0})$ at the parametric rate of $1/\sqrt{n}$. When the functional $
\phi (\cdot )$ is irregular at $\alpha =\alpha _{0}$, we have $\left\Vert
v_{n}^{\ast }\right\Vert _{sd}\asymp \left\Vert v_{n}^{\ast }\right\Vert
\rightarrow \infty $; so the convergence rate of $\phi (\widehat{\alpha }
_{n})$ becomes slower than $1/\sqrt{n}$.
For any regular functional of the semi/nonparametric model (\ref{semi00}),
Theorem \ref{thm:theta_anorm} implies that
\begin{equation*}
\sqrt{n}\left( \phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})\right)
=-n^{-1/2}\sum_{i=1}^{n}S_{n,i}^{\ast }+o_{P_{Z^{\infty }}}(1)\Rightarrow
N(0,\sigma _{v^{\ast }}^{2}),~~\text{with}
\end{equation*}
\begin{eqnarray*}
\sigma _{v^{\ast }}^{2} &= &\lim_{n\rightarrow \infty }\left\Vert v_{n}^{\ast
}\right\Vert _{sd}^{2}=\left\Vert v^{\ast }\right\Vert _{sd}^{2}\\
& = & E\left[
\left( \frac{dm(X,\alpha _{0})}{d\alpha }[v^{\ast }]\right) ^{\prime }\Sigma
(X)^{-1}\Sigma _{0}(X)\Sigma (X)^{-1}\left( \frac{dm(X,\alpha _{0})}{d\alpha
}[v^{\ast }]\right) \right] .
\end{eqnarray*}
Thus, Theorem \ref{thm:theta_anorm} is a natural extension of the asymptotic
normality results of \cite{AC_Emetrica03} and \cite{CP_WP07a} for the
specific regular functional $\phi (\alpha _{0})=\lambda ^{\prime }\theta
_{0} $ of the model (\ref{semi00}). See Remark \ref{remark-norm} in Appendix
\ref{app:appA} for further discussions.
\subsubsection{Closed form expressions of sieve Riesz representer and sieve
variance\label{closed_form_sieve_variance}}
To apply Theorem \ref{thm:theta_anorm}, one needs to know the sieve Riesz
representer $v_{n}^{\ast }$ defined in (\ref{RRT-1}) and the sieve variance $
\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}$ given in (\ref{svar}). It
turns out that both can be computed in closed form.
\begin{lemma}
\label{lem:P-sieve} Let $\overline{\mathbf{V}}_{k(n)}=\mathbb{R}^{d_{\theta
}}\times \{v_{h}(\cdot )=\psi ^{k(n)}(\cdot )^{\prime }\beta :\beta \in
\mathbb{R}^{k(n)}\}=\{v(\cdot )=\overline{\psi }^{k(n)}(\cdot )^{\prime
}\gamma :\gamma \in \mathbb{R}^{d_{\theta }+k(n)}\}$ be dense in the
infinite dimensional Hilbert space $(\overline{\mathbf{V}},\left\Vert \cdot
\right\Vert )$ with the norm $\left\Vert \cdot \right\Vert $ defined in (\ref
{fmetric}). Then: the sieve Riesz representer $v_{n}^{\ast }=(v_{\theta
,n}^{\ast \prime },v_{h,n}^{\ast }\left( \cdot \right) )^{\prime }\in
\overline{\mathbf{V}}_{k(n)}$ of $\frac{d\phi (\alpha _{0})}{d\alpha }[\cdot
]$ has a closed form expression:
\begin{equation}
v_{n}^{\ast }=(v_{\theta ,n}^{\ast \prime },\psi ^{k(n)}(\cdot )^{\prime
}\beta _{n}^{\ast })^{\prime }=\overline{\psi }^{k(n)}(\cdot )^{\prime
}\gamma _{n}^{\ast }\text{, and }\gamma _{n}^{\ast }=D_{n}^{-}\digamma _{n}
\label{finite-riesz-s}
\end{equation}
with $D_{n}=E\left[ \left( \frac{dm(X,\alpha _{0})}{d\alpha }[\overline{\psi
}^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\Sigma (X)^{-1}\left( \frac{
dm(X,\alpha _{0})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )^{\prime
}]\right) \right] $ and $\digamma _{n}=\frac{d\phi (\alpha _{0})}{d\alpha }[
\overline{\psi }^{k(n)}(\cdot )]$. Thus
\begin{equation}
\left\Vert v_{n}^{\ast }\right\Vert ^{2}=\gamma _{n}^{\ast \prime
}D_{n}\gamma _{n}^{\ast }=\digamma _{n}^{\prime }D_{n}^{-}\digamma _{n}\text{
.} \label{finite-riesz-s1}
\end{equation}
The sieve variance (\ref{svar}) also has a closed form expression:
\begin{equation}
||v_{n}^{\ast }||_{sd}^{2}=\digamma _{n}^{\prime }D_{n}^{-}\mho
_{n}D_{n}^{-}\digamma _{n}, \label{P-svar}
\end{equation}
{\small{\begin{eqnarray*}
& &\mho _{n}\equiv\\
& & E\left[ \left( \frac{dm(X,\alpha _{0})}{d\alpha }[\overline{
\psi }^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\Sigma (X)^{-1}\rho
(Z,\alpha _{0})\rho (Z,\alpha _{0})^{\prime }\Sigma (X)^{-1}\left( \frac{
dm(X,\alpha _{0})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )^{\prime
}]\right) \right] .
\end{eqnarray*}}}
\end{lemma}
Let $\mathcal{A}_{k(n)}=\Theta \times \mathcal{H}_{k(n)}$ with $\mathcal{H}
_{k(n)}$ given in (\ref{sieve}). Then $\overline{\mathbf{V}}
_{k(n)}=clsp\left( \mathcal{A}_{k(n)}-\{\alpha _{0,n}\}\right) $ and one
could let ${\psi }^{k(n)}(\cdot )= {q}^{k(n)}(\cdot )$ in
Lemma \ref{lem:P-sieve}, and (\ref{P-svar}) becomes the sieve variance
expression given in (\ref{P-var}).
\noindent Lemmas \ref{lem:sieve-Riesz-property} and \ref{lem:P-sieve} imply
that $\phi \left( \cdot \right) $ is \textit{regular (or irregular) at} $
\alpha =\alpha _{0}$ \textit{iff} $\lim_{k(n)\rightarrow \infty }\left(
\digamma _{n}^{\prime }D_{n}^{-}\digamma _{n}\right) <\infty $ \textit{(or }$
=\infty $\textit{)}.
According to Lemma \ref{lem:P-sieve} we could use different finite
dimensional linear sieve basis $\psi ^{k(n)}$ to compute sieve Riesz
representer $v_{n}^{\ast }=(v_{\theta ,n}^{\ast \prime },v_{h,n}^{\ast
}\left( \cdot \right) )^{\prime }\in \overline{\mathbf{V}}_{k(n)}$, $
\left\Vert v_{n}^{\ast }\right\Vert ^{2}$ and $||v_{n}^{\ast }||_{sd}^{2}$.
Most typical choices include orthonormal bases and the original sieve basis $
q^{k(n)}$ (used to approximate unknown function $h_{0}$). It is typically
easier to characterize the speed of $\left\Vert v_{n}^{\ast }\right\Vert
^{2}=\digamma _{n}^{\prime }D_{n}^{-}\digamma _{n}$ as a function of $k(n)$
when an orthonormal basis is used, while there is a nice interpretation in
terms of sieve variance estimation when the original sieve basis $q^{k(n)}$
is used. See Sections \ref{sec:NPIVex}, \ref{sec:est_avar} and \ref{sec-ex}
for related discussions.
\subsection{Consistent estimator of sieve variance of \protect$\phi (\widehat{\alpha }_{n})$}
\label{sec:est_avar}
In order to apply the asymptotic normality Theorem \ref{thm:theta_anorm}, we
need an estimator of the sieve variance $\left\Vert v_{n}^{\ast }\right\Vert
_{sd}^{2}$ defined in (\ref{svar}). We now provide one simple consistent
estimator of the sieve variance when the residual function $\rho ()$ is
pointwise smooth with respect to $\alpha _{0}$. See Appendix \ref{app:appB}
for additional consistent variance estimators.
The theoretical sieve Riesz representer $v_{n}^{\ast }$ is unknown but can
be estimated easily. Let $\left\Vert \cdot \right\Vert _{n,M}$ denote the
empirical norm induced by the following empirical inner product
\begin{equation}
\langle v_{1},v_{2}\rangle _{n,M}\equiv \frac{1}{n}\sum_{i=1}^{n}\left(
\frac{d\widehat{m}(X_{i},\widehat{\alpha }_{n})}{d\alpha }[v_{1}]\right)
^{\prime }M_{n,i}\left( \frac{d\widehat{m}(X_{i},\widehat{\alpha }_{n})}{
d\alpha }[v_{2}]\right) , \label{RRT-3}
\end{equation}
for any $v_{1},v_{2}\in \overline{\mathbf{V}}_{k(n)}$, where $M_{n,i}$ is
some (almost surely) positive definite weighting matrix.
We define an \textit{empirical sieve Riesz representer} $\widehat{v}
_{n}^{\ast }$ of the functional $\frac{d\phi (\widehat{\alpha }_{n})}{
d\alpha }[\cdot ]$ with respect to the empirical norm $||\cdot ||_{n,
\widehat{\Sigma }^{-1}}$ as
\begin{equation}
\frac{d\phi (\widehat{\alpha }_{n})}{d\alpha }[\widehat{v}_{n}^{\ast
}]=\sup_{v\in \overline{\mathbf{V}}_{k(n)},v\neq 0}\frac{|\frac{d\phi (
\widehat{\alpha }_{n})}{d\alpha }[v]|^{2}}{||v||_{n,\widehat{\Sigma }
^{-1}}^{2}}<\infty \label{RRT-4}
\end{equation}
and
\begin{equation}
\frac{d\phi (\widehat{\alpha }_{n})}{d\alpha }[v]=\langle \widehat{v}
_{n}^{\ast },v\rangle _{n,\widehat{\Sigma }^{-1}}\text{\quad for any }v\in
\overline{\mathbf{V}}_{k(n)}. \label{RRT-5}
\end{equation}
For $\left\Vert v_{n}^{\ast }\right\Vert _{sd}^{2}=E\left( S_{n,i}^{\ast
}S_{n,i}^{\ast \prime }\right) $ given in (\ref{svar}) we can define a
simple plug-in sieve variance estimator:
\begin{align}\notag
||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}= &\frac{1}{n}\sum_{i=1}^{n}\widehat{S}
_{n,i}^{\ast }\widehat{S}_{n,i}^{\ast \prime }\\
= & \frac{1}{n}
\sum_{i=1}^{n}\left( \frac{d\widehat{m}(X_{i},\widehat{\alpha }_{n})}{
d\alpha }[\widehat{v}_{n}^{\ast }]\right) ^{\prime }\widehat{\Sigma }
_{i}^{-1}\left( \widehat{\rho }_{i}\widehat{\rho }_{i}^{\prime }\right)
\widehat{\Sigma }_{i}^{-1}\left( \frac{d\widehat{m}(X_{i},\widehat{\alpha }
_{n})}{d\alpha }[\widehat{v}_{n}^{\ast }]\right) \label{svar-hat1}
\end{align}
with $\widehat{\rho }_{i}=\rho (Z_{i},\widehat{\alpha }_{n})$ and $\widehat{
\Sigma }_{i}=\widehat{\Sigma }(X_{i})$.
Under the condition stated in Lemma \ref{lem:P-sieve}, $\widehat{v}_{n}^{\ast }$
defined in (\ref{RRT-4}-\ref{RRT-5}) also has a closed form solution:
\begin{equation}
\widehat{v}_{n}^{\ast }=\overline{\psi }^{k(n)}(\cdot )^{\prime }\widehat{
\gamma }_{n}^{\ast }\text{,\quad and\quad }\widehat{\gamma }_{n}^{\ast }=
\widehat{D}_{n}^{-}\widehat{\digamma }_{n}, \label{RRT-6}
\end{equation}
with $\widehat{D}_{n}=\frac{1}{n}\sum_{i=1}^{n}\left( \frac{d\widehat{m}
(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot
)^{\prime }]\right) ^{\prime }\widehat{\Sigma }_{i}^{-1}\left( \frac{d
\widehat{m}(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }
^{k(n)}(\cdot )^{\prime }]\right) $ and $\widehat{\digamma }_{n}=\frac{d\phi
(\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )]$. Hence
the sieve variance estimator given in (\ref{svar-hat1}) now becomes
\begin{equation}
||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}=\widehat{V}_{1}\equiv \widehat{
\digamma }_{n}^{\prime }\widehat{D}_{n}^{-}\widehat{\mho }_{n}\widehat{D}
_{n}^{-}\widehat{\digamma }_{n}\text{\quad with} \label{P-svar-hat1}
\end{equation}
\begin{equation*}
\widehat{\mho }_{n}=\frac{1}{n}\sum_{i=1}^{n}\left( \frac{d\widehat{m}(X_{i},
\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )^{\prime
}]\right) ^{\prime }\widehat{\Sigma }_{i}^{-1}\left( \widehat{\rho }_{i}
\widehat{\rho }_{i}^{\prime }\right) \widehat{\Sigma }_{i}^{-1}\left( \frac{d
\widehat{m}(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }
^{k(n)}(\cdot )^{\prime }]\right) .
\end{equation*}
In particular, with $\psi ^{k(n)}=q^{k(n)}$ the sieve variance estimator $||
\widehat{v}_{n}^{\ast }||_{n,sd}^{2}$ given in (\ref{P-svar-hat1}) becomes
the one given in (\ref{P-var-hat}) in Subsection \ref{sec:NPIVex}.
Let $\langle v_{1},v_{2}\rangle _{M}\equiv E\left[ \left( \frac{dm(X,\alpha
_{0})}{d\alpha }[v_{1}]\right) ^{\prime }M\left( \frac{dm(X,\alpha _{0})}{
d\alpha }[v_{2}]\right) \right] $. Then $\langle v_{1},v_{2}\rangle _{\Sigma
^{-1}}\equiv \langle v_{1},v_{2}\rangle $ for all $v_{1},v_{2}\in \overline{
\mathbf{V}}_{k(n)}$. Denote $\overline{\mathbf{V}}_{k(n)}^{1}\equiv \{v\in
\overline{\mathbf{V}}_{k(n)}\colon ||v||=1\}$.
\begin{assumption}
\label{ass:VE} (i) $\sup_{\alpha \in \mathcal{N}_{osn}}\sup_{v\in \overline{
\mathbf{V}}_{k(n)}^{1}}\left\vert \frac{d\phi (\alpha )}{d\alpha }[v]-\frac{
d\phi (\alpha _{0})}{d\alpha }[v]\right\vert =o(1)$;
\noindent (ii) for each $k(n)$ and any $\alpha \in \mathcal{N}_{osn}$, $v\in
\overline{\mathbf{V}}_{k(n)}\mapsto \frac{d\widehat{m}(\cdot ,\alpha )}{
d\alpha }[v]\in L^{2}(f_{X})$ is a linear functional measurable with respect
to $Z^{n}$; and\\
$\sup_{v_{1},v_{2}\in \overline{\mathbf{V}}
_{k(n)}^{1}}\left\vert \langle v_{1},v_{2}\rangle _{n,\Sigma ^{-1}}-\langle
v_{1},v_{2}\rangle _{\Sigma ^{-1}}\right\vert =o_{P_{Z^{\infty }}}(1)$;
\noindent(iii) $\sup_{x\in \mathcal{X}}||\widehat{\Sigma }(x)-\Sigma
(x)||_{e}=o_{P_{Z^{\infty }}}(1)$;
\noindent (iv) $\sup_{x\in \mathcal{X}}E\left[ \sup_{\alpha \in \mathcal{N}
_{osn}}||\rho (Z,\alpha )\rho (Z,\alpha )^{\prime }-\rho (Z,\alpha _{0})\rho
(Z,\alpha _{0})^{\prime }||_{e}|X=x\right] =o(1)$.
\noindent(v) $\sup_{v\in \overline{\mathbf{V}}_{k(n)}^{1}}\left\vert \langle
v,v\rangle _{n,M}-\langle v,v\rangle _{M}\right\vert =o_{P_{Z^{\infty }}}(1)$
with $M=\Sigma ^{-1}\rho (Z,\alpha _{0})\rho (Z,\alpha _{0})^{\prime }\Sigma
^{-1}$.
\end{assumption}
Assumption \ref{ass:VE}(i) becomes vacuous if $\phi $ is linear; otherwise
it requires smoothness of the family $\{\frac{d\phi (\alpha )}{d\alpha }
[v]:\alpha \in \mathcal{N}_{osn}\}$ uniformly in $v\in \overline{\mathbf{V}}
_{k(n)}^{1}$. Assumption \ref{ass:VE}(ii) implicitly assumes that the
residual function $\rho (z,\cdot )$ is \textquotedblleft
smooth\textquotedblright\ in $\alpha \in \mathcal{N}_{osn}$ (see, e.g., \cite
{AC_Emetrica03}) or that $\frac{d\widehat{m}(X,\widehat{\alpha }_{n})}{
d\alpha }[v]$ can be well approximated by numerical derivatives (see, e.g.,
\cite{HMN_WP10}). Assumption \ref{ass:VE}(iii) assumes the existence of
consistent estimators for $\Sigma $. In most applications, $\Sigma (\cdot )$
is either completely known (such as the identity matrix) or $\Sigma _{0}$;
while $\Sigma _{0}(x)$ could be consistently estimated via kernel, series
LS, local linear regression and other nonparametric procedures (see, e.g.,
\cite{AC_Emetrica03} and \cite{CP_WP07a})
\begin{theorem}
\label{thm:VE} Let Assumptions \ref{ass:sieve} - \ref{ass:weak_equiv} hold.
If Assumption \ref{ass:VE} is satisfied, then:
(1) $\left\vert \frac{||\widehat{v}_{n}^{\ast }||_{n,sd}}{||v_{n}^{\ast
}||_{sd}}-1\right\vert =o_{P_{Z^{\infty }}}(1)$ for $||\widehat{v}_{n}^{\ast
}||_{n,sd}$ given in (\ref{svar-hat1}).
(2) If, in addition, Assumptions \ref{ass:phi} and \ref{ass:LAQ} hold, then:
\begin{equation*}
\widehat{W}_{n}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha }_{n})-\phi
(\alpha _{0})}{||\widehat{v}_{n}^{\ast }||_{n,sd}}=-\sqrt{n}\mathbb{Z}
_{n}+o_{P_{Z^{\infty }}}(1)\Rightarrow N(0,1).
\end{equation*}
\end{theorem}
Theorem \ref{thm:VE}(2) allows us to construct confidence sets for $\phi
(\alpha _{0})$ based on a possibly non-optimally weighted plug-in PSMD
estimator $\phi (\widehat{\alpha }_{n})$. A potential drawback, is that it
requires a consistent estimator for $v\mapsto \frac{dm(\cdot ,\alpha _{0})}{
d\alpha }[v]$, which may be hard to compute in practice when the residual
function $\rho (Z,\alpha )$ is not pointwise smooth in $\alpha \in \mathcal{N
}_{osn}$ such as in the NPQIV (\ref{npqiv}) example.
\begin{remark}
\label{remark-Wald} Let $\mathcal{W}_{n}\equiv \left( \sqrt{n}\frac{\phi (
\widehat{\alpha }_{n})-\phi _{0}}{||\widehat{v}_{n}^{\ast }||_{n,sd}}\right)
^{2}=\left( \widehat{W}_{n}+\sqrt{n}\frac{\phi (\alpha _{0})-\phi _{0}}{||
\widehat{v}_{n}^{\ast }||_{n,sd}}\right) ^{2}$ be the Wald test statistic.
Then Theorem \ref{thm:VE} (with $\frac{||v_{n}^{\ast }||_{sd}}{\sqrt{n}}
\asymp \frac{||v_{n}^{\ast }||}{\sqrt{n}}=o(1)$) immediately implies the
following results:
\noindent Under $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$, $\mathcal{W}
_{n}=\left( \widehat{W}_{n}\right) ^{2}\Rightarrow \chi _{1}^{2}$.
\noindent Under $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$, $\mathcal{W}
_{n}=\left( O_{P}(1)+\sqrt{n}||v_{n}^{\ast }||_{sd}^{-1}[\phi (\alpha
_{0})-\phi _{0}]\left( 1+o_{P}(1)\right) \right) ^{2}\rightarrow \infty $ in
probability.
\noindent See Theorem \ref{thm:t-localt} in Appendix \ref{app:appA} for
asymptotic properties of $\mathcal{W}_{n}$ under local alternatives.
\end{remark}
\subsection{Sieve QLR statistics\label{sec:asym_SQLR}}
We now characterize the asymptotic behaviors of the possibly \emph{
non-optimally weighted} SQLR statistic $\widehat{QLR}_{n}(\phi _{0})$
defined in (\ref{SQLR}).
Let $\mathcal{A}_{k(n)}^{R}\equiv \{\alpha \in \mathcal{A}_{k(n)}\colon \phi
(\alpha )=\phi _{0}\}$ be the restricted sieve space, and $\widehat{\alpha }
_{n}^{R}\in \mathcal{A}_{k(n)}^{R}$ be a restricted approximate PSMD
estimator, defined as
\begin{equation}
\widehat{Q}_{n}(\widehat{\alpha }_{n}^{R})+\lambda _{n}Pen(\widehat{h}
_{n}^{R})\leq \inf_{\alpha \in \mathcal{A}_{k(n)}^{R}}\left\{ \widehat{Q}
_{n}(\alpha )+\lambda _{n}Pen(h)\right\} +o_{P_{Z^{\infty }}}(n^{-1}).
\label{Rpsmd}
\end{equation}
Then:
\begin{align*}
\widehat{QLR}_{n}(\phi _{0})= &n\left( \widehat{Q}_{n}(\widehat{\alpha }
_{n}^{R})-\widehat{Q}_{n}(\widehat{\alpha }_{n})\right)\\
= & n\left(
\inf_{\alpha \in \mathcal{A}_{k(n)}^{R}}\widehat{Q}_{n}(\alpha
)-\inf_{\alpha \in \mathcal{A}_{k(n)}}\widehat{Q}_{n}(\alpha )\right) +o_{P_{Z^{\infty }}}(1).
\end{align*}
Recall that $u_{n}^{\ast }\equiv v_{n}^{\ast }/\left\Vert v_{n}^{\ast
}\right\Vert _{sd}$, and that $\widehat{QLR}_{n}^{0}(\phi _{0})$ denotes the
optimally weighted (i.e., $\Sigma =\Sigma _{0}$) SQLR statistic in
Subsection \ref{sec:NPIVex}. We note that $||u_{n}^{\ast }||=1$ for the
optimally weighted case.
\begin{theorem}
\label{thm:chi2} Let Assumptions \ref{ass:sieve} - \ref{ass:LAQ} hold with $
\left\vert B_{n}-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{Z^{\infty }}}(1)$.
If $\widehat{\alpha }_{n}^{R}\in \mathcal{N}_{osn}$ wpa1-$P_{Z^{\infty }}$,
then: (1) under the null $H_{0}:\phi (\alpha _{0})=\phi _{0}$,
\begin{equation*}
||u_{n}^{\ast }||^{2}\times \widehat{QLR}_{n}(\phi _{0})=\left( \sqrt{n}
\mathbb{Z}_{n}\right) ^{2}+o_{P_{Z^{\infty }}}(1)\Rightarrow \chi _{1}^{2}.
\end{equation*}
(2) Further, let $\widehat{\alpha }_{n}$ be the optimally weighted PSMD
estimator (\ref{psmd}) with $\Sigma =\Sigma _{0}$. Then: under $H_{0}:$ $
\phi (\alpha _{0})=\phi _{0}$,
\begin{equation*}
\widehat{QLR}_{n}^{0}(\phi _{0})=\left( \sqrt{n}\mathbb{Z}_{n}\right)
^{2}+o_{P_{Z^{\infty }}}(1)\Rightarrow \chi _{1}^{2}.
\end{equation*}
\noindent \textit{See Theorem \ref{thm:chi2_localt} in Appendix \ref
{app:appA} for the asymptotic behavior under local alternatives}.
\end{theorem}
Compared to Theorem \ref{thm:theta_anorm} on the asymptotic normality of $
\phi (\widehat{\alpha }_{n})$, Theorem \ref{thm:chi2} on the asymptotic null
distribution of the SQLR statistic requires two extra conditions: $
\left\vert B_{n}-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{Z^{\infty }}}(1)$
and $\widehat{\alpha }_{n}^{R}\in \mathcal{N}_{osn}$ wpa1-$P_{Z^{\infty }}$.
Both conditions are also needed even for QLR statistics in parametric
extremum estimation and testing problems. Lemma \ref{lem:Qdiff_B} in Section
\ref{sec:bootstrap} provides a simple sufficient condition (Assumption \ref
{ass:LLN_triangular}) for $\left\vert B_{n}-||u_{n}^{\ast }||^{2}\right\vert
=o_{P_{Z^{\infty }}}(1)$. Proposition \ref{pro:conv-rate-RPSMDE} in Appendix
\ref{app:appB} establishes $\widehat{\alpha }_{n}^{R}\in \mathcal{N}_{osn}$
wpa1-$P_{Z^{\infty }}$ under the null $H_{0}:\phi (\alpha _{0})=\phi _{0}$
and other conditions virtually the same as those for Lemma \ref
{thm:ThmCONVRATEGRAL} (i.e., $\widehat{\alpha }_{n}\in \mathcal{N}_{osn}$
wpa1-$P_{Z^{\infty }}$).
Theorem \ref{thm:chi2}(2) recommends to construct an asymptotic $100(1-\tau
)\%$ confidence set for $\phi (\alpha )$ by inverting the optimally weighted
SQLR statistic: $\left\{ r\in \mathbb{R}\colon \text{ }\widehat{QLR}
_{n}^{0}(r)\leq c_{\chi _{1}^{2}}(1-\tau )\right\} $. This result extends
that of \cite{CP_WP07a} for a regular Euclidean functional $\phi (\alpha
)=\lambda ^{\prime }\theta $ to possibly irregular nonlinear functionals.
Next, we consider the asymptotic behavior of $\widehat{QLR}_{n}(\phi _{0})$
under the fixed alternatives $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$.
\begin{theorem}
\label{thm:QLR-H1} Let Assumptions \ref{ass:sieve}, \ref{A_3.6} and \ref
{ass:rates} hold. Suppose that $\sup_{h\in \mathcal{H}}Pen(h)<\infty $ and $
\phi $ is continuous in $||\cdot ||_{s}$. Then: under $H_{1}:$ $\phi (\alpha
_{0})\neq \phi _{0}$, there is a constant $C>0$ such that
\begin{equation*}
\frac{\widehat{QLR}_{n}(\phi _{0})}{n}\geq C>0\quad \text{wpa1.}
\end{equation*}
\end{theorem}
\section{Inference Based on Generalized Residual Bootstrap}
\label{sec:bootstrap}
\setcounter{assumption}{0}
The inference procedures described in Subsections \ref{sec:est_avar} and \ref
{sec:asym_SQLR} are based on the asymptotic critical values. For many
parametric models it is known that bootstrap based procedures could
approximate finite sample distributions more accurately. In this section we
establish the consistency of the bootstrap sieve Wald and SQLR statistics
under virtually the same conditions as those imposed for the original-sample
sieve Wald and SQLR statistics.
A bootstrap procedure is described by an array of \textquotedblleft
weights\textquotedblright\ $\left\{ \omega _{i,n}\right\} _{i=1}^{n}$ for
each $n$, where each bootstrap sample is drawn independently of the original
data $\left\{ Z_{i}\right\} _{i=1}^{n}$. Different bootstrap procedures
correspond to different choices of the weights $\left\{ \omega
_{i,n}\right\} _{i=1}^{n}$ but all satisfy $\omega _{i,n}\geq 0$ and $
E[\omega _{i,n}]=1$. For the time being we assume that $\lim_{n\rightarrow
\infty }Var(\omega _{i,n})=\sigma _{\omega }^{2}\in (0,\infty )$ for all $i$.
In this paper we focus on two types of bootstrap weights:
\begin{assumption}[I.i.d Weights]
\label{ass:Wboot} Let $(\omega _{i})_{i=1}^{n}$ be a sequence such that $
\omega _{i}\in \mathbb{R}_{+}$, $\omega _{i}\sim iidP_{\omega }$, $E[\omega
]=1$, $Var(\omega )=\sigma _{\omega }^{2}$, and $\int_{0}^{\infty }\sqrt{
P(|\omega -1|\geq t)}dt<\infty $.
\end{assumption}
The condition $\int_{0}^{\infty }\sqrt{P(|\omega -1|\geq t)}dt<\infty $ is
implied by $E[|\omega -1|^{2+\epsilon }]<\infty $ for some $\epsilon >0$.
\begin{assumption}[Multinomial Weights]
\label{ass:Wboot_e} Let $(\omega _{i,n})_{i=1}^{n}$ be a triangular array of
random variables such that $(\omega _{1,n},...,\omega _{n,n})\sim
Multinomial(n;n^{-1},...,n^{-1})$.
\end{assumption}
We sometimes omit the $n$ subscript from the weight series. Note that under
Assumption \ref{ass:Wboot_e}, $E[\omega _{1}]=1$, $Var(\omega
_{1})=(1-1/n)\rightarrow 1\equiv \sigma _{\omega }^{2}$ and $Cov(\omega
_{i},\omega _{j})=-n^{-1}$ (for $i\neq j$). Finally, $n^{-1}\max_{1\leq
i\leq n}(\omega _{i}-1)^{2}=o_{P_{\omega }}(1)$. We use these facts in the
proofs.
Let $V_{i}\equiv (Z_{i},\omega _{i,n})$ and
\begin{equation*}
\rho ^{B}(V_{i},\alpha )\equiv \omega _{i,n}\rho (Z_{i},\alpha ),
\end{equation*}
be the bootstrap residual function. Let $\widehat{m}^{B}(x,\alpha )$ be a
bootstrap version of $\widehat{m}(x,\alpha )$, that is, $\widehat{m}
^{B}(x,\alpha )$ is computed in the same way as that of $\widehat{m}
(x,\alpha )$ except that we use $\rho ^{B}(V_{i},\alpha )$ instead of $\rho
(Z_{i},\alpha )$. In particular, $\widehat{m}^{B}(x,\alpha
)=\sum_{i=1}^{n}\omega _{i,n}\rho (Z_{i},\alpha )A_{n}(X_{i},x)$ for any
linear estimator $\widehat{m}(x,\alpha )$ (\ref{mhat-linear}) of $m(x,\alpha
)$. For example, if $\widehat{m}(x,\alpha )$ is a series LS estimator (\ref
{mhat}), then $\widehat{m}^{B}(x,\alpha )$ is the bootstrap series LS
estimator (\ref{mhat_B}) defined in Subsection \ref{sec:NPIVex}.
Let $\widehat{Q}_{n}^{B}(\alpha )\equiv \frac{1}{n}\sum_{i=1}^{n}\widehat{m}
^{B}(X_{i},\alpha )^{\prime }\widehat{\Sigma }(X_{i})^{-1}\widehat{m}
^{B}(X_{i},\alpha )$ be a bootstrap version of $\widehat{Q}_{n}(\alpha )$,
and $\widehat{\alpha }_{n}^{B}$ be the bootstrap PSMD estimator, i.e., $
\widehat{\alpha }_{n}^{B}$ is an approximate minimizer of $\left\{ \widehat{Q
}_{n}^{B}(\alpha )+\lambda _{n}Pen(h)\right\} $ on $\mathcal{A}_{k(n)}$.
Denote $\widehat{\phi }_{n}\equiv \phi (\widehat{\alpha }_{n})$. Then
\begin{equation*}
\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})=n\left( \inf_{\{\mathcal{A}
_{k(n)}\colon \phi (\alpha )=\widehat{\phi }_{n}\}}\widehat{Q}
_{n}^{B}(\alpha )-\widehat{Q}_{n}^{B}(\widehat{\alpha }_{n}^{B})\right)
\end{equation*}
is the (generalized residual) bootstrap SQLR test statistic. And $\mathcal{W}
_{1,n}^{B}\equiv \left( \sqrt{n}\frac{\phi (\widehat{\alpha }_{n}^{B})-
\widehat{\phi }_{n}}{\sigma _{\omega }||\widehat{v}_{n}^{\ast }||_{n,sd}}
\right) ^{2}$ is one simple bootstrap Wald test statistic (see Subsection
\ref{sub-boots-t} for another simple bootstrap Wald statistic).
\textbf{Additional notation}. To be more precise, we introduce some
definitions associated with the new random variables $V_{i}\equiv
(Z_{i},\omega _{i,n})$ and the enlarged probability spaces. Let $\Omega
=\{\omega _{i,n}\colon i=1,...,n;~n=1,...\}$ be the space of weights,
defined as a triangle array with elements in $\mathbb{R}$, the corresponding
$\sigma $-algebra and probability are $(\mathcal{B}_{\Omega },P_{\Omega })$.
Let $\mathcal{V}^{\infty }\equiv \mathcal{Z}^{\infty }\times \Omega $, $
\mathcal{B}^{\infty }\equiv \mathcal{B}_{Z}^{\infty }\times \mathcal{B}
_{\Omega }$ be the $\sigma $-algebra, and $P_{V^{\infty }}$ be the joint
probability over $\mathcal{V}^{\infty }$. Finally, for each $n$, let $
\mathcal{B}^{n}$ be the $\sigma $-algebra generated by $V^{n}\equiv
Z^{n}\times (\omega _{1,n},...,\omega _{n,n})$, where each $\omega _{i,n}$
acts as a \textquotedblleft weight\textquotedblright\ of $Z_{i}$. Let $A_{n}$
be a random variable that is measurable with respect to $\mathcal{B}^{n}$,
and $\mathcal{L}_{V^{\infty }|Z^{\infty }}(A_{n}|Z^{n})$ (or $P_{V^{\infty
}|Z^{\infty }}\left( A_{n}\leq \cdot \mid Z^{n}\right) $) be the conditional
law (or conditional distribution) of $A_{n}$ given $Z^{n}$. Let $B_{n}$ be a
random variable measurable with respect to $\mathcal{B}_{Z}^{\infty }$, and $
\mathcal{L}(B_{n})$ (or $P_{Z^{\infty }}\left( B_{n}\leq \cdot \right) $) be
the law (or distribution) of $B_{n}$. For two real valued random variables, $
A_{n}$ (measurable with respect to $\mathcal{B}^{n}$) and $B$ (measurable
with respect to some $\sigma $-algebra $\mathcal{B}_{B}$), we say $
\left\vert \mathcal{L}_{V^{\infty }|Z^{\infty }}(A_{n}|Z^{n})-\mathcal{L}
(B)\right\vert =o_{P_{Z^{\infty }}}(1)$ if for any $\delta >0$, there exists
a $N(\delta )$ such that
\begin{equation*}
P_{Z^{\infty }}\left( \sup_{f\in BL_{1}}\left\vert
E[f(A_{n})|Z^{n}]-E[f(B)]\right\vert \leq \delta \right) \geq 1-\delta \text{
\quad for all }n\geq N(\delta ),
\end{equation*}
(i.e., $\sup_{f\in BL_{1}}\left\vert E[f(A_{n})|Z^{n}]-E[f(B)]\right\vert
=o_{P_{Z^{\infty }}}(1)$), where $BL_{1}$ denotes the class of uniformly
bounded Lipschitz functions $f:\mathbb{R}\rightarrow \mathbb{R}$ such that $
||f||_{L^{\infty }}\leq 1$ and $|f(z)-f(z^{\prime })|\leq |z-z^{\prime }|$.
See chapter 1.12 of \cite{VdV-W_book96} (henceforth, VdV-W) for more details.
We say $\Delta _{n}$ is of order $o_{P_{V^{\infty }|Z^{\infty }}}(1)$ in $
P_{Z^{\infty }}$ probability, and denote it as $\Delta _{n}=o_{P_{V^{\infty
}|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$, if for any $\epsilon >0$, $
P_{Z^{\infty }}\left( P_{V^{\infty }|Z^{\infty }}\left( |\Delta
_{n}|>\epsilon \mid Z^{n}\right) >\epsilon \right) \rightarrow 0$ as $
n\rightarrow \infty $.
We say $\Delta _{n}$ is of order $O_{P_{V^{\infty }|Z^{\infty }}}(1)$ in $
P_{Z^{\infty }}$ probability, and denote it as $\Delta _{n}=O_{P_{V^{\infty
}|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$, if for any $\epsilon >0$ there
exists a $M\in (0,\infty )$, such that $P_{Z^{\infty }}\left( P_{V^{\infty
}|Z^{\infty }}\left( |\Delta _{n}|>M\mid Z^{n}\right) >\epsilon \right)
\rightarrow 0$ as $n\rightarrow \infty $.
\subsection{Bootstrap local quadratic approximation (LQA$^{B}$)}
Lemma \ref{lem:cons_boot} in Appendix \ref{app:appA} shows that the
bootstrap PSMD estimator $\widehat{\alpha }_{n}^{B}\in \mathcal{N}_{osn}$
wpa1 under Assumptions \ref{ass:rates_B} and \ref{ass:sieve} - \ref
{ass:weak_equiv}. This allows us to introduce a condition that is a
bootstrap version of the LQA Assumption \ref{ass:LAQ}. For any $\alpha \in
\mathcal{N}_{osn}$, we let $\widehat{\Lambda }_{n}^{B}(\alpha (t_{n}),\alpha
)\equiv 0.5\{\widehat{Q}_{n}^{B}(\alpha (t_{n}))-\widehat{Q}_{n}^{B}(\alpha
)\}$ with $\alpha (t_{n})\equiv \alpha +t_{n}u_{n}^{\ast }$ for $t_{n}\in
\mathcal{T}_{n}$. For any sequence of non-negative weights $(b_{i})_{i}$,
let
\begin{equation*}
\mathbb{Z}_{n}^{b}\equiv n^{-1}\sum_{i=1}^{n}b_{i}\left( \frac{
dm(X_{i},\alpha _{0})}{d\alpha }[u_{n}^{\ast }]\right) ^{\prime }\Sigma
(X_{i})^{-1}\rho (Z_{i},\alpha _{0})=n^{-1}\sum_{i=1}^{n}b_{i}\frac{
S_{n,i}^{\ast }}{\left\Vert v_{n}^{\ast }\right\Vert _{sd}}.
\end{equation*}
\begin{assumption}[LQA$^{B}$]
\label{ass:LAQ_B} (i) $\alpha (t)\in \mathcal{A}_{k(n)}$ for any $(\alpha
,t)\in \mathcal{N}_{osn}\times \mathcal{T}_{n}$, and with $
r_{n}(t_{n})=\left( \max \{t_{n}^{2},t_{n}n^{-1/2},o(n^{-1})\}\right) ^{-1}$
,
\begin{eqnarray*}
& & \sup_{(\alpha ,t_{n})\in \mathcal{N}_{osn}\times \mathcal{T}
_{n}}r_{n}(t_{n})\left\vert \widehat{\Lambda }_{n}^{B}(\alpha (t_{n}),\alpha
)-t_{n}\left\{ \mathbb{Z}_{n}^{\omega }+\langle u_{n}^{\ast },\alpha -\alpha
_{0}\rangle \right\} -\frac{B_{n}^{\omega }}{2}t_{n}^{2}\right\vert\\
& & = o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})
\end{eqnarray*}
where $B_{n}^{\omega }$ is a $V^{n}$ measurable positive random variable
such that $B_{n}^{\omega }=O_{P_{V^{\infty }|Z^{\infty
}}}(1)~wpa1(P_{Z^{\infty }})$;
\begin{equation*}
\text{(ii)\quad }\left\vert \mathcal{L}_{V^{\infty }|Z^{\infty }}\left(
\sqrt{n}\frac{\mathbb{Z}_{n}^{\omega -1}}{\sigma _{\omega }}\mid
Z^{n}\right) -\mathcal{L}\left( \mathbb{Z}\right) \right\vert
=o_{P_{Z^{\infty }}}(1),
\end{equation*}
where $\mathbb{Z}$ is a standard normal random variable.
\end{assumption}
Assumption \ref{ass:LAQ_B}(i) implicitly imposes restrictions on the
bootstrap estimator $\widehat{m}^{B}(x,\alpha )$ of the conditional mean
function $m(x,\alpha )$. Below we provide low level sufficient conditions
for Assumption \ref{ass:LAQ_B}(i) when $\widehat{m}^{B}(x,\alpha )$ is a
bootstrap series LS estimator.
Let $g(X,u_{n}^{\ast })\equiv \{\frac{dm(X,\alpha _{0})}{d\alpha }
[u_{n}^{\ast }]\}^{\prime }\Sigma (X)^{-1}$. Then $E\left[
g(X_{i},u_{n}^{\ast })\Sigma (X_{i})g(X_{i},u_{n}^{\ast })^{\prime }\right]
=||u_{n}^{\ast }||^{2}$.
\setcounter{assumption}{0}
\begin{assumption} \label{ass:LLN_triangular} For $\Gamma (\cdot )\in \{\Sigma (\cdot ),\Sigma
_{0}(\cdot )\}$,
\begin{equation*}
\left\vert n^{-1}\sum_{i=1}^{n}g(X_{i},u_{n}^{\ast })\Gamma
(X_{i})g(X_{i},u_{n}^{\ast })^{\prime }-E\left[ g(X_{i},u_{n}^{\ast })\Gamma
(X_{i})g(X_{i},u_{n}^{\ast })^{\prime }\right] \right\vert =o_{P_{Z^{\infty
}}}(1).
\end{equation*}
\end{assumption}
\begin{lemma}
\label{lem:Qdiff_B} Let Assumptions \ref{ass:sieve} - \ref{ass:weak_equiv}
and \ref{ass:m_ls} - \ref{ass:cont_diffm} hold.
(1) Let $\widehat{m}$ be the series LS estimator (\ref{mhat}). Then
Assumption \ref{ass:LAQ}(i) is satisfied. Further, if Assumption \ref
{ass:LLN_triangular} holds then $\left\vert B_{n}-||u_{n}^{\ast
}||^{2}\right\vert =o_{P_{Z^{\infty }}}(1)$.
(2) Let $\widehat{m}^{B}(\cdot ,\alpha )$ be the bootstrap series LS
estimator (\ref{mhat_B}), Assumption \ref{ass:rates_B}, and either
Assumption \ref{ass:Wboot} or \ref{ass:Wboot_e} hold. Then Assumption \ref
{ass:LAQ_B}(i) holds with $B_{n}^{\omega }=B_{n}$. Further, if Assumption
\ref{ass:LLN_triangular} holds then $\left\vert B_{n}^{\omega
}-||u_{n}^{\ast }||^{2}\right\vert =o_{P_{V^{\infty }|Z^{\infty
}}}(1)~wpa1(P_{Z^{\infty }})$.
\end{lemma}
Lemma \ref{lem:Qdiff_B} indicates that the low level Assumptions \ref
{ass:m_ls} - \ref{ass:cont_diffm} are sufficient for both the
original-sample LQA Assumption \ref{ass:LAQ}(i) and the bootstrap LQA
Assumption \ref{ass:LAQ_B}(i).
Assumption \ref{ass:LAQ_B}(ii) can be easily verified by applying some
central limit theorems. For example, if the weights are independent
(Assumption \ref{ass:Wboot}), we can use Lindeberg-Feller CLT; if the
weights are multinomial (Assumption \ref{ass:Wboot_e}) we can apply Hayek
CLT (see \cite{VdV-W_book96} p. 458 ). The next lemma provides some simple
sufficient conditions for Assumption \ref{ass:LAQ_B}(ii).
\begin{lemma}
\label{lem:clt_B} Let either Assumption \ref{ass:Wboot} or Assumption \ref
{ass:Wboot_e} hold. If there is a positive real sequence $(b_{n})_{n}$ such
that $b_{n}=o\left( \sqrt{n}\right) $ and
\begin{equation}
\limsup_{n\rightarrow \infty }E\left[ \left( g(X,u_{n}^{\ast })\rho
(Z,\alpha _{0})\right) ^{2}1\left\{ \frac{(g(X,u_{n}^{\ast })\rho (Z,\alpha
_{0}))^{2}}{b_{n}}>1\right\} \right] =0, \label{CLT_triangular}
\end{equation}
then Assumptions \ref{ass:LAQ_B}(ii) and \ref{ass:LAQ}(ii) hold.
\end{lemma}
\subsection{Bootstrap sieve Student t statistic\label{sub-boots-t}}
\setcounter{assumption}{3}
Lemma \ref{lem:cons_boot} shows that $\widehat{\alpha }_{n}^{B}\in \mathcal{N
}_{osn}$ wpa1 under virtually the same conditions as those for the
original-sample estimator $\widehat{\alpha }_{n}\in \mathcal{N}_{osn}$ wpa1.
This would easily lead to the consistency of the simplest bootstrap sieve t
statistic $\widehat{W}_{1,n}^{B}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha }
_{n}^{B})-\phi (\widehat{\alpha }_{n})}{\sigma _{\omega }||\widehat{v}
_{n}^{\ast }||_{n,sd}}$.
We now establish the consistency of another bootstrap sieve t statistic $
\widehat{W}_{2,n}^{B}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha }
_{n}^{B})-\phi (\widehat{\alpha }_{n})}{||\widehat{v}_{n}^{\ast }||_{B,sd}}$
, where $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}$ is a bootstrap sieve
variance estimator:
\begin{equation}
||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}\equiv \frac{1}{n}\sum_{i=1}^{n}\left(
\frac{d\widehat{m}(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\widehat{v}
_{n}^{\ast }]\right) ^{\prime }\widehat{\Sigma }_{i}^{-1}\varrho (V_{i},
\widehat{\alpha }_{n})\varrho (V_{i},\widehat{\alpha }_{n})^{\prime }
\widehat{\Sigma }_{i}^{-1}\left( \frac{d\widehat{m}(X_{i},\widehat{\alpha }
_{n})}{d\alpha }[\widehat{v}_{n}^{\ast }]\right) \label{svar-hat1-boot}
\end{equation}
with $\varrho (V_{i},\alpha )\equiv (\omega _{i,n}-1)\rho (Z_{i},\alpha
)\equiv \rho ^{B}(V_{i},\alpha )-\rho (Z_{i},\alpha )$ for any $\alpha $.
We note that $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}$ is an analog to $||
\widehat{v}_{n}^{\ast }||_{n,sd}^{2}$ defined in (\ref{svar-hat1}) but using
the bootstrapped generalized residual $\varrho (V_{i},\widehat{\alpha }_{n})$
instead of the original sample fitted residual $\rho (Z_{i},\widehat{\alpha }
_{n})$. It also has a closed form expression: $||\widehat{v}_{n}^{\ast
}||_{B,sd}^{2}=\widehat{\digamma }_{n}^{\prime }\widehat{D}_{n}^{-}\widehat{
\mho }_{n}^{B}\widehat{D}_{n}^{-}\widehat{\digamma }_{n}$ with
\begin{equation*}
\widehat{\mho }_{n}^{B}=\frac{1}{n}\sum_{i=1}^{n}\left( \frac{d\widehat{m}
(X_{i},\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot
)^{\prime }]\right) ^{\prime }(\omega
_{i,n}-1)^{2} \widehat{M}_i \left( \frac{d\widehat{m}(X_{i},
\widehat{\alpha }_{n})}{d\alpha }[\overline{\psi }^{k(n)}(\cdot )^{\prime
}]\right)
\end{equation*}
where $\widehat{M}_i \equiv \widehat{\Sigma }_{i}^{-1}\rho (Z_{i},\widehat{\alpha }_{n})\rho (Z_{i},\widehat{\alpha }
_{n})^{\prime }\widehat{\Sigma }_{i}^{-1}$. That is, $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}$ is computed in the same
way as $||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}=\widehat{\digamma }
_{n}^{\prime }\widehat{D}_{n}^{-}\widehat{\mho }_{n}\widehat{D}_{n}^{-}
\widehat{\digamma }_{n}$ given in (\ref{P-svar-hat1}) except using $\widehat{
\mho }_{n}^{B}$ instead of $\widehat{\mho }_{n}$.
\begin{assumption}
\label{ass:VE_boot2} $\sup_{v\in \overline{\mathbf{V}}_{k(n)}^{1}}|\langle
v,v\rangle _{n,\widehat{M}^{B}}-\sigma _{\omega }^{2}\langle v,v\rangle _{n,\widehat{
M}}|=o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$ with $\widehat{M}
_{i}^{B}=(\omega _{i,n}-1)^{2}\widehat{M}_{i}$.
\end{assumption}
This assumption can be verified given Assumptions \ref{ass:Wboot} or \ref
{ass:Wboot_e}. The following result is a bootstrap version of Theorem \ref
{thm:VE}(1).
\begin{theorem}
\label{thm:VE-boot2} Let Assumptions \ref{ass:sieve} - \ref{ass:weak_equiv},
\ref{ass:VE} and \ref{ass:VE_boot2} hold. Then:
\begin{equation*}
\left\vert \frac{||\widehat{v}_{n}^{\ast }||_{B,sd}}{\sigma _{\omega
}||v_{n}^{\ast }||_{sd}}-1\right\vert =o_{P_{V^{\infty }|Z^{\infty
}}}(1)~wpa1(P_{Z^{\infty }}).
\end{equation*}
\end{theorem}
Recall that $\widehat{W}_{n}\equiv \sqrt{n}\frac{\phi (\widehat{\alpha }
_{n})-\phi (\alpha _{0})}{||\widehat{v}_{n}^{\ast }||_{n,sd}}$, whose
probability distribution $P_{Z^{\infty }}\left( \widehat{W}_{n}\leq \cdot
\right) $ converges to the standard normal cdf $\Phi (\cdot )$. The next
result is about the consistency of the bootstrap sieve t statistic $\widehat{
W}_{2,n}^{B}$.
\begin{theorem}
\label{thm:bootstrap_2} Let $\widehat{\alpha }_{n}$ be the PSMD estimator (
\ref{psmd}) and $\widehat{\alpha }_{n}^{B}$ the bootstrap PSMD estimator.
Let Assumptions \ref{ass:sieve} - \ref{ass:weak_equiv} and \ref{ass:rates_B}
hold. Let Assumptions \ref{ass:phi}, \ref{ass:LAQ} and \ref{ass:LAQ_B} hold.
(1) Let Assumptions \ref{ass:VE} and \ref{ass:VE_boot2} hold. Then:
\begin{equation*}
\sup_{t\in \mathbb{R}}\left\vert P_{V^{\infty }|Z^{\infty }}\left( \widehat{W
}_{2,n}^{B}\leq t\mid Z^{n}\right) -P_{Z^{\infty }}\left( \widehat{W}
_{n}\leq t\right) \right\vert =o_{P_{V^{\infty }|Z^{\infty
}}}(1)~wpa1(P_{Z^{\infty }}).
\end{equation*}
(2) If $\phi ()$ is regular at $\alpha _{0}$, without imposing Assumptions
\ref{ass:VE} and \ref{ass:VE_boot2}, we have:
\begin{eqnarray*}
& & \sup_{t\in \mathbb{R}}\left\vert P_{V^{\infty }|Z^{\infty }}\left( \sqrt{n}
\frac{\phi (\widehat{\alpha }_{n}^{B})-\phi (\widehat{\alpha }_{n})}{\sigma
_{\omega }}\leq t\mid Z^{n}\right) -P_{Z^{\infty }}\left( \sqrt{n}\left(
\phi (\widehat{\alpha }_{n})-\phi (\alpha _{0})\right) \leq t\right)
\right\vert \\
&& = o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }}).
\end{eqnarray*}
\end{theorem}
For a regular functional, Theorem \ref{thm:bootstrap_2}(2) provides one way
to construct its confidence sets without the need to compute any variance
estimator. This extends the result in \cite{CP_WP07a} for a regular
Euclidean parameter $\lambda ^{\prime }\theta $ to a general regular
functional $\phi (\alpha )$. Unfortunately for an irregular functional, we
need to compute a consistent bootstrap sieve variance estimator $||\widehat{v
}_{n}^{\ast }||_{B,sd}^{2}$ to apply Theorem \ref{thm:bootstrap_2}(1).
Luckily $||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}$ is easy to compute when the
residual function $\rho (Z_{i},\alpha )$ is pointwise smooth in $\alpha _{0}$
. Moreover, since $E\left( ||\widehat{v}_{n}^{\ast }||_{B,sd}^{2}\mid
Z^{n}\right) =\sigma _{\omega }^{2}||\widehat{v}_{n}^{\ast }||_{n,sd}^{2}$
we suspect that the bootstrap sieve t statistic $\widehat{W}_{2,n}^{B}$
might have second order refinement property by choices of bootstrap weights $
\{\omega _{i,n}\}$. This will be a subject of future research.
The bootstrap sieve t statistic $\widehat{W}_{2,n}^{B}$ requires to compute
the original sample PSMD estimator $\widehat{\alpha }_{n}$ and the bootstrap
PSMD estimator $\widehat{\alpha }_{n}^{B}$. In the online Appendix \ref{app:appD} we present a sieve score test and its bootstrap version,
which only use the original sample restricted PSMD estimator $\widehat{
\alpha }_{n}^{R}$ and do not use $\widehat{\alpha }_{n}^{B}$, and hence are
computationally simple.
\begin{remark}
\label{remark-Wald-B} Theorems \ref{thm:VE}(2) and \ref{thm:bootstrap_2}(1)
imply that the bootstrap Wald test statistic $\mathcal{W}_{2,n}^{B}\equiv
\left( \widehat{W}_{2,n}^{B}\right) ^{2}$ always has the same limiting
distribution $\chi _{1}^{2}$ (conditional on the data) under the null and
the alternatives. Let $\widehat{c}_{2,n}(a)$ be the $a-th$ quantile of the
distribution of $\mathcal{W}_{2,n}^{B}$ (conditional on the data $
\{Z_{i}\}_{i=1}^{n}$). Let $\mathcal{W}_{n}\equiv \left( \sqrt{n}\frac{\phi (
\widehat{\alpha }_{n})-\phi _{0}}{||\widehat{v}_{n}^{\ast }||_{n,sd}}\right)
^{2}$ be the original sample Wald test statistic. Then Remark \ref
{remark-Wald} and Theorem \ref{thm:bootstrap_2}(1) immediately imply that
for any $\tau \in (0,1)$,
under $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$, $\lim_{n\rightarrow \infty
}\Pr \left( \mathcal{W}_{n}\geq \widehat{c}_{2,n}(1-\tau )\right) =\tau $;
under $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$, $\lim_{n\rightarrow
\infty }\Pr \left( \mathcal{W}_{n}\geq \widehat{c}_{2,n}(1-\tau )\right) =1.$
\noindent See Theorem \ref{thm:waldB_con} in Appendix \ref{app:appA} for
properties under local alternatives.
\end{remark}
See online supplemental Appendix \ref{app:appB} for consistency of $\mathcal{
W}_{1,n}^{B}\equiv \left( \sqrt{n}\frac{\phi (\widehat{\alpha }_{n}^{B})-
\widehat{\phi }_{n}}{\sigma _{\omega }||\widehat{v}_{n}^{\ast }||_{n,sd}}
\right) ^{2}$ and other bootstrap sieve Wald (t) statistics based on
different sieve variance estimators.
\subsection{Bootstrap SQLR statistic}
If $\Sigma \neq \Sigma _{0}$, the SQLR statistic $\widehat{QLR}_{n}(\phi
_{0})=n\left( \widehat{Q}_{n}(\widehat{\alpha }_{n}^{R})-\widehat{Q}_{n}(
\widehat{\alpha }_{n})\right) $ is no longer asymptotically chi-square even
under the null; Theorem \ref{thm:chi2}(1), however, implies that the SQLR
statistic converges weakly to a tight limit under the null. In this
subsection we show that the asymptotic null distribution of the SQLR can be
consistently approximated by that of the (generalized residual) bootstrap
SQLR statistic $\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})$. Recall that
\begin{equation*}
\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})=n\left( \widehat{Q}_{n}^{B}(
\widehat{\alpha }_{n}^{R,B})-\widehat{Q}_{n}^{B}(\widehat{\alpha }
_{n}^{B})\right) +o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})
\end{equation*}
where $\widehat{\phi }_{n}\equiv \phi (\widehat{\alpha }_{n})$, and $
\widehat{\alpha }_{n}^{R,B}$ is the \emph{restricted} bootstrap PSMD
estimator, defined as
\begin{eqnarray}
\widehat{Q}_{n}^{B}(\widehat{\alpha }_{n}^{R,B})+\lambda _{n}Pen(\widehat{h}
_{n}^{R,B}) &\leq& \inf_{\alpha \in \mathcal{A}_{k(n)}:\phi (\alpha )=\widehat{
\phi }_{n}}\left\{ \widehat{Q}_{n}^{B}(\alpha )+\lambda _{n}Pen(h)\right\}\\
&& +o_{P_{V^{\infty }|Z^{\infty }}}(\frac{1}{n})~wpa1(P_{Z^{\infty }}).
\label{R-B-PSMD}
\end{eqnarray}
Lemma \ref{lem:cons_boot} in Appendix \ref{app:appA} implies that $\widehat{
\alpha }_{n}^{R,B},$ $\widehat{\alpha }_{n}^{B}\in \mathcal{N}_{osn}$ wpa1
under both the null $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$ and the
alternatives $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$. This indicates
that the bootstrap SQLR statistic $\widehat{QLR}_{n}^{B}(\widehat{\phi }
_{n}) $ is always properly centered and should be stochastically bounded
under both the null and the alternatives, as shown in the next theorem. Let $
P_{Z^{\infty }}\left( \widehat{QLR}_{n}(\phi _{0})\leq \cdot \mid
H_{0}\right) $ denote the probability distribution of $\widehat{QLR}
_{n}(\phi _{0})$ under the null $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$,
which would converge to the cdf of $\chi _{1}^{2}$ when $\widehat{QLR}
_{n}(\phi _{0})=\widehat{QLR}_{n}^{0}(\phi _{0})$ (the optimally weighted
SQLR).
\begin{theorem}
\label{thm:bootstrap} Let Assumptions \ref{ass:sieve} - \ref{ass:weak_equiv}
and \ref{ass:rates_B} hold. Let Assumptions \ref{ass:phi}, \ref{ass:LAQ} and
\ref{ass:LAQ_B} hold with $\left\vert B_{n}^{\omega }-||u_{n}^{\ast
}||^{2}\right\vert =o_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }})$
. Then:
\begin{equation*}
\text{(1) }\frac{\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})}{\sigma _{\omega
}^{2}}=\left( \sqrt{n}\frac{\mathbb{Z}_{n}^{\omega -1}}{\sigma _{\omega
}||u_{n}^{\ast }||}\right) ^{2}+o_{P_{V^{\infty }|Z^{\infty
}}}(1)=O_{P_{V^{\infty }|Z^{\infty }}}(1)~wpa1(P_{Z^{\infty }});
\end{equation*}
and
\begin{eqnarray*}
\text{(2) } && \sup_{t\in \mathbb{R}}\left\vert P_{V^{\infty }|Z^{\infty
}}\left( \frac{\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})}{\sigma _{\omega
}^{2}}\leq t\mid Z^{n}\right) -P_{Z^{\infty }}\left( \widehat{QLR}_{n}(\phi
_{0})\leq t\mid H_{0}\right) \right\vert \\
&& =o_{P_{V^{\infty }|Z^{\infty
}}}(1)~wpa1(P_{Z^{\infty }}).
\end{eqnarray*}
\end{theorem}
Theorem \ref{thm:bootstrap} allows us to construct valid confidence sets
(CS) for $\phi (\alpha _{0})$ based on inverting possibly \emph{non}
-optimally weighted SQLR statistic without the need to compute a variance
estimator. We recommend this procedure when it is difficult to compute any
consistent variance estimator for $\phi (\widehat{\alpha })$, such as in the
cases when the residual function $\rho (Z;\alpha )$ is pointwise non-smooth
in $\alpha _{0}$. See, e.g., \cite{AB_Emetrica00} for a thorough discussion
about how to construct CS via bootstrap.
\begin{remark}
\label{remark-QLR-B} Let $\widehat{c}_{n}(a)$ be the $a-th$ quantile of the
distribution of $\frac{\widehat{QLR}_{n}^{B}(\widehat{\phi }_{n})}{\sigma
_{\omega }^{2}}$ (conditional on the data $\{Z_{i}\}_{i=1}^{n}$). Then
Theorems \ref{thm:chi2}, \ref{thm:QLR-H1} and \ref{thm:bootstrap}
immediately imply that for any $\tau \in (0,1)$,
under $H_{0}:$ $\phi (\alpha _{0})=\phi _{0}$, $\lim_{n\rightarrow \infty
}\Pr \left( \widehat{QLR}_{n}(\phi _{0})\geq \widehat{c}_{n}(1-\tau )\right)
=\tau $;
under $H_{1}:$ $\phi (\alpha _{0})\neq \phi _{0}$, $\lim_{n\rightarrow
\infty }\Pr \left( \widehat{QLR}_{n}(\phi _{0})\geq \widehat{c}_{n}(1-\tau
)\right) =1.$
\noindent See Theorem \ref{thm:BSQLR-loc-alt} in Appendix \ref{app:appA} for
properties under local alternatives.
\end{remark}
\section{Verification of Assumptions 3.5 and 3.6}
\label{sec-ex}
In this section, we illustrate the verification of the two key regularity
conditions, Assumption \ref{ass:phi} and Assumption \ref{ass:LAQ}(i), via
some functionals $\phi (h)$ of the (nonlinear) nonparametric IV regressions:
\begin{equation}
E[\rho (Y_{1};h_{0}(Y_{2}))|X]=0\quad a.s.-X, \label{gnpiv1}
\end{equation}
where the scalar valued residual function $\rho ()$ could be nonlinear and
pointwise non-smooth in $h$. This model includes the NPIV and NPQIV as
special cases. To be concrete, we consider a PSMD estimator $\widehat{h}\in
\mathcal{H}_{k(n)}$ of $h_{0}$ with $\widehat{\Sigma }=\Sigma =1$, and $
\widehat{m}(\cdot ,h)$ being the series LS estimator (\ref{mhat}) of $
m(\cdot ,h)=E[\rho (Y_{1};h(Y_{2}))|X=\cdot ]$ with $J_{n}=ck(n)$ for a
finite constant $c\geq 1$. We assume that $h_{0}\in \mathcal{H}=\Lambda
_{c}^{\varsigma }\left( [-1,1]\right) $ with smoothness $\varsigma >1/2$ (a H
\"{o}lder ball with support $[-1,1]$, see, e.g., \cite{CLvK_Emetrica03}).
\footnote{
This H\"{o}lder ball condition and several other conditions assumed in this
subsection are for illustration only, and can be replaced by weaker
sufficient conditions.} By definition, $\mathcal{H}\subset L^{2}(f_{Y_{2}})$
and we let $||\cdot ||_{s}=||\cdot ||_{L^{2}(f_{Y_{2}})}.$ We assume that $
\mathcal{H}_{k(n)}=clsp\{q_{1},...,q_{k(n)}\}$ with $\{q_{k}\}_{k=1}^{\infty
}$ being a Riesz basis of $(\mathcal{H},||\cdot ||_{s})$. The convergence
rates of $\widehat{h}$ to $h_{0}$ in both $||\cdot ||$ and $||\cdot
||_{s}=||\cdot ||_{L^{2}(f_{Y_{2}})}$ metrics have already been established
in \cite{CP_WP07}, and hence will not be repeated here.
We use $\mathcal{H}_{os}$ and $\mathcal{H}_{osn}$ for $\mathcal{A}_{os}$ and
$\mathcal{A}_{osn}$ defined in Subsection \ref{sec:consistency} (since there
is no $\theta $ here). Denote $T\equiv \frac{dm(\cdot ,h_{0})}{dh}:\mathcal{H
}_{os}\subset L^{2}(f_{Y_{2}})\rightarrow L^{2}(f_{X})$\textbf{, }i.e., for
any $h\in \mathcal{H}_{os}\subset L^{2}(f_{Y_{2}})$,
\begin{equation*}
Th\equiv \left. \frac{dE[\rho (Y_{1};h_{0}(Y_{2})+\tau h(Y_{2}))|X=\cdot ]}{
d\tau }\right\vert _{\tau =0}.
\end{equation*}
Let $T^{\ast }$ be the adjoint of $T$. Then for all $h\in \mathcal{H}_{os}$,
we have $||h||^{2}\equiv ||Th||_{L^{2}(f_{X})}^{2}=||(T^{\ast
}T)^{1/2}h||_{L^{2}(f_{Y_{2}})}^{2}$. Under mild conditions as stated in
\cite{CP_WP07}, $T$ and $T^{\ast }$ are compact. Then $T$ has a singular
value decomposition $\{\mu _{k};\psi _{k},\phi _{0k}\}_{k=1}^{\infty }$,
where $\{\mu _{k}>0\}_{k=1}^{\infty }$ is the sequence of singular values in
non-increasing order ($\mu _{k}\geq \mu _{k+1}\geq ...$) with $
\liminf_{k\rightarrow \infty }\mu _{k}=0$, $\{\psi _{k}\in
L^{2}(f_{Y_{2}})\}_{k=1}^{\infty }$ and $\{\phi _{0k}\in
L^{2}(f_{X})\}_{k=1}^{\infty }$ are sequences of eigenfunctions of the
operators $(T^{\ast }T)^{1/2}$ and $(TT^{\ast })^{1/2}$:
\begin{equation*}
T\psi _{k}=\mu _{k}\phi _{0k},\text{\quad }(T^{\ast }T)^{1/2}\psi _{k}=\mu
_{k}\psi _{k}\quad \text{and\quad }(TT^{\ast })^{1/2}\phi _{0k}=\mu _{k}\phi
_{0k}\quad \text{for all }k.
\end{equation*}
Since $\{q_{k}\}_{k=1}^{\infty }$ is a Riesz basis of $(\mathcal{H},||\cdot
||_{s})$ we could also have $\mathcal{H}_{k(n)}=clsp\{\psi _{1},...,\psi
_{k(n)}\}$. The sieve measure of local ill-posedness now becomes $\tau
_{n}=\mu _{k(n)}^{-1}$ (see, e.g., \cite{BCK_Emetrica07} and \cite{CP_WP07}
), and hence $\left\Vert u_{n}^{\ast }\right\Vert _{s}\leq c\mu _{k(n)}^{-1}$
for a finite constant $c>0$. Also, $\Pi _{n}h_{0}\equiv \arg \min_{h\in
\mathcal{H}_{k(n)}}||h-h_{0}||_{s}=\sum_{k=1}^{k(n)}\langle h_{0},\psi
_{k}\rangle _{s}\psi _{k}$ is the LS projection of $h_{0}$ onto the sieve
space $\mathcal{H}_{n}$ under the strong norm $||\cdot ||_{s}=||\cdot
||_{L^{2}(f_{Y_{2}})}$. Recall that $h_{0,n}\equiv \arg \min_{h\in \mathcal{H
}_{k(n)}}||h-h_{0}||^{2}\equiv \arg \min_{h\in \mathcal{H}
_{k(n)}}||T[h-h_{0}]||_{L^{2}(f_{X})}^{2}$. We have:
\begin{eqnarray} \notag
h_{0,n} &=&\arg \min_{\{a_{k}\}}\left[ \sum_{k=1}^{k(n)}\left( \langle
h_{0},\psi _{k}\rangle _{s}-a_{k}\right) ^{2}\mu
_{k}^{2}+\sum_{k=k(n)+1}^{\infty }\langle h_{0},\psi _{k}\rangle _{s}^{2}\mu
_{k}^{2}\right]\\
&= &\sum_{k=1}^{k(n)}\langle h_{0},\psi _{k}\rangle _{s}\psi
_{k}=\Pi _{n}h_{0}. \label{h0n}
\end{eqnarray}
The next remark specializes Theorem \ref{thm:theta_anorm} to a general
functional $\phi (h)$ of the model (\ref{gnpiv1}).
\begin{remark}
\label{remark-normality-gnpiv1} Let $\widehat{m}$ be the series LS estimator
(\ref{mhat}) for the model (\ref{gnpiv1}) with $\widehat{\Sigma }=\Sigma =1$
, and Assumptions \ref{ass:sieve}(i)(ii), \ref{A_3.6}(ii)(iii), and \ref
{ass:weak_equiv} hold with $\delta _{n}=O\left( \sqrt{\frac{k(n)}{n}}\right)
=o(n^{-1/4})$ and $\delta _{s,n}=O\left( \{k(n)\}^{-\varsigma }+\mu
_{k(n)}^{-1}\sqrt{\frac{k(n)}{n}}\right) =o(1)$. Let Assumption \ref{ass:phi}
, equation (\ref{LF}) and Assumptions \ref{ass:m_ls} - \ref{ass:cont_diffm}
hold. Then:
\begin{equation}
\sqrt{n}\frac{\phi (\widehat{h}_{n})-\phi (h_{0})}{||v_{n}^{\ast }||_{sd}}
\Rightarrow N(0,1), \label{P-var-gnpiv1}
\end{equation}
with $||v_{n}^{\ast }||_{sd}^{2}=(\frac{
d\phi (h_{0})}{dh} [q^{k(n)}(\cdot )])^{\prime }D_{n}^{-1}{\mho }_{n}D_{n}^{-1}(\frac{d\phi (h_{0})}{dh}
[q^{k(n)} (\cdot )])$, and $D_{n}=E\left[ \left( T[q^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\left(
T[q^{k(n)}(\cdot )^{\prime }]\right) \right] $ and $\mho _{n}=E\left[ \left(
T[q^{k(n)}(\cdot )^{\prime }]\right) ^{\prime }\rho (Z,h_{0})^{2}\left(
T[q^{k(n)}(\cdot )^{\prime }]\right) \right] .$
\end{remark}
Remark \ref{remark-normality-gnpiv1} includes the NPIV\ and NPQIV examples
in Subsection \ref{sec:NPIVex} as special cases. In particular, the sieve
variance expression (\ref{P-var-gnpiv1}) reproduces the one for the NPIV
model (\ref{npiv}) with $T[q^{k(n)}(\cdot )^{\prime
}]=E[q^{k(n)}(Y_{2})^{\prime }|X]$, and the one for the NPQIV model (\ref
{npqiv}) with $T[q^{k(n)}(\cdot )^{\prime
}]=E[f_{U|Y_{2},X}(0)q^{k(n)}(Y_{2})^{\prime }|X]$.
By the result in \cite{CP_WP07}, the sieve dimension $k_{n}^{\ast }$
satisfying $\{k_{n}^{\ast }\}^{-\varsigma }\asymp \mu _{k_{n}^{\ast
}}^{-1}\times \sqrt{\frac{k_{n}^{\ast }}{n}}$ leads to the nonparametric
optimal convergence rate of $||\widehat{h}-h_{0}||_{s}=O_{P_{Z^{\infty
}}}(\delta _{s,n}^{\ast })=o(1)$ in strong norm, where $\delta _{s,n}^{\ast
}\asymp \{k_{n}^{\ast }\}^{-\varsigma }$. In particular, $k_{n}^{\ast
}\asymp n^{\frac{1}{2(\varsigma +a)+1}}$ and $\delta _{s,n}^{\ast }=n^{-
\frac{\varsigma }{2(\varsigma +a)+1}}$ for the \textit{mildly ill-posed case}
$\mu _{k}\asymp k^{-a}$ for a finite $a>0$; and $\delta _{s,n}^{\ast }=\{\ln
n\}^{-\varsigma }$ for the \textit{severely ill-posed case} $\mu _{k}\asymp
\exp \{-0.5ak\}$ for a finite $a>0$. However this paper aims at simple valid
inferences on functional $\phi (h_{0})$. As will be illustrated in the next
subsection, although the nonparametric optimal choice $k_{n}^{\ast }$ is
compatible with the sufficient conditions for the asymptotic normality of $
\sqrt{n}(\phi (\widehat{h})-\phi (h_{0}))$ for a regular linear functional $
\phi (h_{0})$ (see Remark \ref{remark-bias}), it is typically ruled out by
Assumption \ref{ass:phi}(iii) for irregular functionals.
\subsection{Verification of Assumption \protect\ref{ass:phi}}
Let $b_{j}\equiv \frac{d\phi (h_{0})}{dh}[\psi _{j}(\cdot )]$ for all $j$.
By Lemma \ref{lem:P-sieve} $D_{n}=E\left[ \left( T[q^{k(n)}(\cdot )^{\prime
}]\right) ^{\prime }\left( T[q^{k(n)}(\cdot )^{\prime }]\right) \right]
=Diag\left\{ \mu _{1}^{2},...,\mu _{k(n)}^{2}\right\} $ and
\begin{equation}
||v_{n}^{\ast }||^{2}=\left( \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot
)]\right) ^{\prime }D_{n}^{-1}\left( \frac{d\phi (h_{0})}{dh}[q^{k(n)}(\cdot
)]\right) =\sum_{j=1}^{k(n)}\mu _{j}^{-2}b_{j}^{2}.
\label{sieve-riesz-gnpiv1}
\end{equation}
By Lemma \ref{lem:sieve-Riesz-property}, $\phi (h)$ of the model (\ref
{gnpiv1}) is regular (at $h=h_{0}$) iff $\sum_{j=1}^{\infty }\mu
_{j}^{-2}b_{j}^{2}<\infty $, and is irregular (at $h=h_{0}$) iff $
\sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}=\infty $.
For the same functional $\phi (h)$ of a model (\ref{ex-gnpiv}) without
endogeneity:
\begin{equation}
E[\rho (Y_{1};h_{0}(Y_{2}))|Y_{2}]=0\quad a.s.-Y_{2}, \label{ex-gnpiv}
\end{equation}
we have $D_{n}\asymp I_{k(n)}$ and $||v_{n}^{\ast }||^{2}\asymp
\sum_{j=1}^{k(n)}b_{j}^{2}$. Thus, $\phi (h)$ of the model (\ref{ex-gnpiv})
is regular (or irregular) iff $\sum_{j=1}^{\infty }b_{j}^{2}<\infty $ (or $
=\infty $).
Since $\mu _{k(n)}\rightarrow 0$ as $k(n)\rightarrow \infty $, if a
functional $\phi (h)$ is irregular for the model (\ref{ex-gnpiv}) without
endogeneity, then it is irregular for the model (\ref{gnpiv1}). But, even if
a functional $\phi (h)$ is regular for the model (\ref{ex-gnpiv}) without
endogeneity, it could still be irregular for the model (\ref{gnpiv1}) with
endogeneity.
\subsubsection{Linear functionals of the model (\protect\ref{gnpiv1})}
For a linear functional $\phi (h)$ of the model (\ref{gnpiv1}), given
relation (\ref{h0n}), Assumption \ref{ass:phi} is satisfied provided that
the sieve dimension $k(n)$ satisfies (\ref{a3.1-linear}):
\begin{equation}
\frac{||v_{n}^{\ast }||}{\sqrt{n}}=o(1)\text{\quad and\quad }\sqrt{n}\frac{
\left\vert \frac{d\phi (h_{0})}{dh}[\Pi _{n}h_{0}-h_{0}]\right\vert }{
||v_{n}^{\ast }||}=o(1). \label{a3.1-linear}
\end{equation}
When $\phi (h)$ of the model (\ref{gnpiv1}) is regular, Remark \ref
{remark-bias} implies that (\ref{a3.1-linear}) is satisfied provided
\begin{equation}
\sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}<\infty \text{\quad and\quad }
n\times \sum_{j=k(n)+1}^{\infty }\mu _{j}^{-2}b_{j}^{2}\times ||\Pi
_{n}h_{0}-h_{0}||^{2}=o(1). \label{a3.1-linear-r}
\end{equation}
We shall illustrate below that both these sufficient conditions allow for
severely ill-posed problems.
\smallskip
\textbf{Example 1 (evaluation functional).} For $\phi (h)=h(\overline{y}
_{2}) $, we have: $||v_{n}^{\ast }||^{2}=\sum_{j=1}^{k(n)}\mu _{j}^{-2}[\psi
_{j}(\overline{y}_{2})]^{2}$. Let $\mathcal{H}_{k(n)}$ be the spline or the CDV wavelet sieve as described in \cite{CC_WP13}, say. Then
\begin{equation*}
\left\vert \frac{d\phi (h_{0})}{dh}[\Pi _{n}h_{0}-h_{0}]\right\vert =|(\Pi
_{n}h_{0})(\overline{y}_{2})-h_{0}(\overline{y}_{2})|\leq ||\Pi
_{n}h_{0}-h_{0}||_{\infty }\leq const.\{k(n)\}^{-\varsigma }.
\end{equation*}
To provide concrete sufficient condition for (\ref{a3.1-linear}), we assume $
||v_{n}^{\ast }||^{2}\asymp E\left( \sum_{j=1}^{k(n)}\mu _{j}^{-2}[\psi
_{j}(Y_{2})]^{2}\right) =\sum_{k=1}^{k(n)}\mu _{k}^{-2}$. Since $
\lim_{k(n)\rightarrow \infty }||v_{n}^{\ast }||^{2}=\infty $, the evaluation
functional is irregular. Condition (\ref{a3.1-linear}) is satisfied provided
that
\begin{equation}
\frac{||v_{n}^{\ast }||^{2}}{n}=\frac{\sum_{k=1}^{k(n)}\mu _{k}^{-2}}{n}=o(1)
\text{\quad and\quad }\frac{\{k(n)\}^{-2\varsigma }}{\frac{1}{n}
||v_{n}^{\ast }||^{2}}=\frac{\{k(n)\}^{-2\varsigma }}{\frac{1}{n}
\sum_{k=1}^{k(n)}\mu _{k}^{-2}}=o(1). \label{Ex1-both}
\end{equation}
Condition (\ref{Ex1-both}) allows for both mildly and severely ill-posed
cases.
(a) \textit{Mildly ill-posed}: $\mu _{k}\asymp k^{-a}$ for a finite $a>0$.
Then $||v_{n}^{\ast }||^{2}\asymp \{k(n)\}^{2a+1}$. Condition (\ref{Ex1-both}
) is satisfied by a wide range of sieve dimensions, such as $k(n)\asymp n^{
\frac{1}{2(\varsigma +a)+1}}(\ln \ln n)^{\varpi }$ or $n^{\frac{1}{
2(\varsigma +a)+1}}(\ln n)^{\varpi }$ for any finite $\varpi >0$, or $
k(n)\asymp n^{\epsilon }$ for any $\epsilon \in (\frac{1}{2(\varsigma +a)+1},
\frac{1}{2a+1})$. Note that any $k(n)$ satisfying Condition (\ref{Ex1-both})
also ensures $\delta _{s,n}=o(1)$. However, it does require $
k(n)/k_{n}^{\ast }\rightarrow \infty $, where $k_{n}^{\ast }\asymp n^{\frac{1
}{2(\varsigma +a)+1}}$ is the choice for the nonparametric optimal
convergence rate in strong norm.
(b) \textit{Severely ill-posed}: $\mu _{k}\asymp \exp \{-0.5ak\}$ for a
finite $a>0$. Then $||v_{n}^{\ast }||^{2}\asymp \exp \{ak(n)\}$. Condition (
\ref{Ex1-both}) is satisfied with $k(n)\asymp a^{-1}\left[ \ln n-\varpi \ln
(\ln n)\right] $ for $0<\varpi <2\varsigma $. In addition we need $\varpi >1$
(and hence $\varsigma >1/2$) to ensure $\delta _{s,n}=O\left(
\{k(n)\}^{-\varsigma }+\mu _{k(n)}^{-1}\sqrt{\frac{k(n)}{n}}\right) =o(1)$.
\smallskip
\textbf{Example 2 (weighted derivative functional).} For $\phi (h)=\int
w(y)\nabla h(y)dy$, where $w(y)$ is a weight satisfying the integration by
part formula: $\phi (h)=\int w(y)\nabla h(y)dy=-\int h(y)\nabla w(y)dy$, we
have: $||v_{n}^{\ast }||^{2}=\sum_{j=1}^{k(n)}\mu _{j}^{-2}b_{j}^{2}$ with $
b_{j}=\int \psi _{j}(y)\nabla w(y)dy$ for all $j$, and
\begin{eqnarray*}
\left\vert \frac{d\phi (h_{0})}{dh}[\Pi _{n}h_{0}-h_{0}]\right\vert
&= &\left\vert \int [\Pi _{n}h_{0}(y)-h_{0}(y)]\nabla w(y)dy\right\vert \\
&\leq &
C\times ||\Pi _{n}h_{0}-h_{0}||_{L^{2}(f_{Y_{2}})}\leq
const.\{k(n)\}^{-\varsigma }
\end{eqnarray*}
provided that $E\left( \left[ \frac{\nabla w(Y_{2})}{f_{Y_{2}}(Y_{2})}\right]
^{2}\right) =\sum_{j=1}^{\infty }b_{j}^{2}=C<\infty $. That is, the weighted
derivative is assumed to be regular for the model (\ref{ex-gnpiv}) without
endogeneity.
\textbf{(i)} When the weighted derivative is regular (i.e., $
\sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}<\infty $) for the model (\ref
{gnpiv1}), Condition (\ref{a3.1-linear-r}) is satisfied provided that $
n\times \sum_{j=k(n)+1}^{\infty }\mu _{j}^{-2}b_{j}^{2}\times \delta
_{n}^{2}=o(1)$, which is the condition imposed in \cite{AC_JOE07} for their
root-$n$ estimation of an average derivative of NPIV example, and is shown
to allow for severely ill-posed inverse case in \cite{AC_JOE07}.
\textbf{(ii)} When the weighted derivative is irregular (i.e., $
\sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}=\infty $) for the model (\ref
{gnpiv1}), Condition (\ref{a3.1-linear}) is satisfied provided that
\begin{equation}
\frac{||v_{n}^{\ast }||^{2}}{n}=\frac{\sum_{j=1}^{k(n)}\mu _{j}^{-2}b_{j}^{2}
}{n}=o(1)\text{\quad and\quad }\frac{\{k(n)\}^{-2\varsigma }}{\frac{1}{n}
||v_{n}^{\ast }||^{2}}=\frac{\{k(n)\}^{-2\varsigma }}{\frac{1}{n}
\sum_{j=1}^{k(n)}\mu _{j}^{-2}b_{j}^{2}}=o(1). \label{Ex2-severe}
\end{equation}
Condition (\ref{Ex2-severe}) allows for both mildly and severely ill-posed
cases. To provide concrete sufficient conditions for (\ref{Ex2-severe}) we
assume $b_{j}^{2}\asymp \left( j\ln (j)\right) ^{-1}$ in the following
calculations.
(a) \textit{Mildly ill-posed}: $\mu _{k}\asymp k^{-a}$ for a finite $a>0$.
Then $||v_{n}^{\ast }||^{2}\in \lbrack c\frac{k(n)^{2a}}{\ln (k(n))}
,c^{\prime }k(n)^{2a}]$ for some $0<c\leq c^{\prime }<\infty $. Condition (
\ref{Ex2-severe}) and $\delta _{s,n}=o(1)$ are jointly satisfied by a wide
range of sieve dimensions, such as $k(n)\asymp n^{\frac{1}{2(\varsigma +a)}
}(\ln n)^{\varpi }$ for any finite $\varpi >\frac{1}{2(\varsigma +a)}$, or $
k(n)\asymp n^{\epsilon }$ for any $\epsilon \in (\frac{1}{2(\varsigma +a)},
\frac{1}{2a+1})$ and $\varsigma >1/2$.
(b) \textit{Severely ill-posed}: $\mu _{k}\asymp \exp \{-0.5ak\}$ for $a>0$.
Then $||v_{n}^{\ast }||^{2}\in \lbrack c\frac{\exp \{ak(n)\}}{k(n)\ln (k(n))}
,c^{\prime }\frac{\exp \{ak(n)\}}{\ln (k(n))}]$ for some $0<c\leq c^{\prime
}<\infty $. Condition (\ref{Ex2-severe}) and $\delta _{s,n}=o(1)$ are
jointly satisfied by $k(n)\asymp a^{-1}\left[ \ln (n)-\varpi \ln (\ln (n))
\right] $ for $\varpi \in (1,2\varsigma -1)$ and $\varsigma >1$.
\subsubsection{Nonlinear functionals}
For a nonlinear functional $\phi (h)$ of the model (\ref{gnpiv1}),
Assumption \ref{ass:phi} is satisfied provided that the sieve dimension $
k(n) $ satisfies (\ref{a3.1-linear}) (or (\ref{a3.1-linear-r}) if $\phi (h)$
is regular) and Assumption \ref{ass:phi}(ii), which is implied by the
following condition:
\noindent \textbf{Assumption \ref{ass:phi}(ii)'}: \textit{there are finite
non-negative constants }$C\geq 0,\omega _{1},\omega _{2}\geq 0$\textit{\
such that for all }$(\alpha ,t)\in \mathcal{N}_{osn}\times \mathcal{T}_{n}$,
\begin{eqnarray*}
&& \left\vert \phi (\alpha +tu_{n}^{\ast })-\phi (\alpha _{0})-\frac{d\phi
(\alpha _{0})}{d\alpha }[\alpha +tu_{n}^{\ast }-\alpha _{0}]\right\vert \\
& &\leq
C\times (||\alpha -\alpha _{0}+tu_{n}^{\ast }||^{\omega _{1}}\times ||\alpha
-\alpha _{0}+tu_{n}^{\ast }||_{s}^{\omega _{2}}),
\end{eqnarray*}
and
\begin{equation*}
C\times \frac{\sqrt{n}\times (\delta _{n}(1+M_{n}^{2}))^{\omega _{1}}\times
(\delta _{s,n}+M_{n}^{2}\delta _{n}||u_{n}^{\ast }||_{s})^{\omega _{2}}}{
||v_{n}^{\ast }||}=o\left( 1\right) .
\end{equation*}
Assumption \ref{ass:phi}(ii) or (ii)' controls the nonlinearity bias of $
\phi \left( \cdot \right) $ (i.e., the linear approximation error of a
nonlinear functional $\phi \left( \cdot \right) $). It typically rules out
nonlinear regular functionals of severely illposed inverse problems, but
allows for nonlinear irregular functionals of severely illposed inverse
problems.
\textbf{Example 3 (weighted quadratic functional).} For $\phi (h)=\frac{1}{2}
\int w(y)\left\vert h(y)\right\vert ^{2}dy$, we have $||v_{n}^{\ast
}||^{2}=\sum_{j=1}^{k(n)}\mu _{j}^{-2}b_{j}^{2}$ with $b_{j}=\int
h_{0}(y)w(y)\psi _{j}(y)dy$ for all $j$, and
\begin{equation*}
\left\vert \frac{d\phi (h_{0})}{dh}[\Pi _{n}h_{0}-h_{0}]\right\vert
=\left\vert \int w(y)h_{0}(y)[\Pi _{n}h_{0}(y)-h_{0}(y)]dy\right\vert \leq
const.\times ||\Pi _{n}h_{0}-h_{0}||_{L^{2}(f_{Y_{2}})}
\end{equation*}
provided that $\sup_{y}\frac{w(y)}{f_{Y_{2}}(y)}<\infty $. This and $E\left(
\left[ h_{0}(Y_{2})\right] ^{2}\right) <\infty $ imply that $
\sum_{j=1}^{\infty }b_{j}^{2}<\infty $. That is, the weighted quadratic
functional is regular for the model (\ref{ex-gnpiv}) without endogeneity.
Also,
\begin{equation*}
\left\vert \phi (h)-\phi (h_{0})-\frac{d\phi (h_{0})}{dh}[h-h_{0}]\right
\vert =\frac{1}{2}\int w(y)\left\vert h(y)-h_{0}(y)\right\vert ^{2}dy\leq
const.\times ||h-h_{0}||_{L^{2}(f_{Y_{2}})}^{2}.
\end{equation*}
\textbf{(i)} When the weighted quadratic functional is regular (i.e., $
\sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}<\infty $) for the model (\ref
{gnpiv1}), Condition (\ref{a3.1-linear-r}) is satisfied provided that $
n\times \sum_{j=k(n)+1}^{\infty }\mu _{j}^{-2}b_{j}^{2}\times \delta
_{n}^{2}=o(1)$, which allows for severely ill-posed cases. But Assumption
\ref{ass:phi}(ii)' requires that $\sqrt{n}\times \delta _{s,n}^{2}=\sqrt{n}
\times \left( \{k(n)\}^{-\varsigma }+\mu _{k(n)}^{-1}\sqrt{\frac{k(n)}{n}}
\right) ^{2}=o(1)$, which clearly rules out severely ill-posed inverse case
where $\mu _{k}\asymp \exp \{-0.5ak\}$ for some finite $a>0$.
\textbf{(ii)} When the weighted quadratic functional is irregular (i.e., $
\sum_{j=1}^{\infty }\mu _{j}^{-2}b_{j}^{2}=\infty $) for the model (\ref
{gnpiv1}), Condition (\ref{a3.1-linear}) is satisfied provided that
Condition (\ref{Ex2-severe}) holds with $b_{j}=\int h_{0}(y)w(y)\psi
_{j}(y)dy$ for Example 3. Assumption \ref{ass:phi}(ii)' is satisfied
provided that
\begin{equation}
\sqrt{n}\frac{\delta _{s,n}^{2}}{||v_{n}^{\ast }||}=\frac{\sqrt{n}\times
\left( \{k(n)\}^{-\varsigma }+\mu _{k(n)}^{-1}\sqrt{\frac{k(n)}{n}}\right)
^{2}}{||v_{n}^{\ast }||}\leq n^{-1/2}\frac{\mu _{k(n)}^{-2}k(n)}{\sqrt{
\sum_{j=1}^{k(n)}\mu _{j}^{-2}b_{j}^{2}}}=o(1). \label{Ex3-severe2}
\end{equation}
Any $k(n)$ satisfying Conditions (\ref{Ex2-severe}) and (\ref{Ex3-severe2})
automatically satisfies $\delta _{s,n}=o(1)$. In addition, both conditions
allow for mildly and severely ill-posed cases. To provide concrete
sufficient conditions we assume $b_{j}^{2}\asymp \left( j\ln (j)\right)
^{-1} $ in the following calculations.
(a) \textit{Mildly ill-posed}: $\mu _{k}\asymp k^{-a}$ for a finite $a>0$.
Then $||v_{n}^{\ast }||^{2}\in \lbrack c\frac{k(n)^{2a}}{\ln (k(n))}
,c^{\prime }k(n)^{2a}]$ for some $0<c\leq c^{\prime }<\infty $. Conditions (
\ref{Ex2-severe}) and (\ref{Ex3-severe2}) are satisfied by a wide range of
sieve dimensions, such as $k(n)\asymp n^{\frac{1}{2(\varsigma +a)}}(\ln
n)^{\varpi }$ for any finite $\varpi >\frac{1}{2(\varsigma +a)}$, or $
k(n)\asymp n^{\epsilon }$ for any $\epsilon \in (\frac{1}{2(\varsigma +a)},
\frac{1}{2a+2})$ and $\varsigma >1$.
(b) \textit{Severely ill-posed}: $\mu _{k}\asymp \exp \{-0.5ak\}$ for $a>0$.
Then $||v_{n}^{\ast }||^{2}\in \lbrack c\frac{\exp \{ak(n)\}}{k(n)\ln (k(n))}
,c^{\prime }\frac{\exp \{ak(n)\}}{\ln (k(n))}]$ for some $0<c\leq c^{\prime
}<\infty $. Conditions (\ref{Ex2-severe}) and (\ref{Ex3-severe2}) are
satisfied with $k(n)\asymp a^{-1}\left[ \ln (n)-\varpi \ln (\ln (n))\right] $
and $\varpi \in (3,2\varsigma -1)$ for $\varsigma >2$.
\subsection{Verification of Assumption \protect\ref{ass:LAQ}(i)}
By Lemma \ref{lem:Qdiff_B}(1), to verify Assumption \ref{ass:LAQ}(i), it
suffices to verify Assumptions \ref{ass:m_ls} - \ref{ass:cont_diffm} in
Appendix \ref{app:appA}. Note that Assumptions \ref{ass:m_ls} and \ref
{ass:rho_Donsker} do not depend on sieve Riesz representer at all, and have
already been verified in \cite{CP_WP07a}, \cite{AC_JOE07} and others for
(penalized) SMD estimators for the model (\ref{gnpiv1}). Assumptions \ref
{ass:anor-mtilde} and \ref{ass:cont_diffm} do depend on the scaled sieve
Riesz representer $u_{n}^{\ast }\equiv v_{n}^{\ast }/||v_{n}^{\ast }||_{sd}$
. Both these assumptions are also verified in \cite{AC_Emetrica03}, \cite
{CP_WP07a}, \cite{AC_JOE07} for examples of regular functionals of the model
(\ref{gnpiv1}). Here, we present simple (albeit somewhat strong) sufficient conditions for Assumptions \ref
{ass:anor-mtilde} and \ref{ass:cont_diffm} for irregular functionals of the
NPIV and NPQIV examples.
\begin{condition}
\label{eqn:NPIV-holder} (i) $\{E[h(Y_{2})|\cdot ]:h\in \mathcal{H}
\}\subseteq \Lambda _{c}^{\gamma }(\mathcal{X})$, with $\gamma >0.5$; (ii) $
\sup_{x,y_{2}}\frac{f_{Y_{2}X}(y_{2},x)}{f_{Y_{2}}(y_{2})f_{X}(x)}\leq
Const.<\infty $.
\end{condition}
\begin{proposition}
\label{pro:NPIV-suff} Let all conditions for Remark \ref
{remark-normality-gnpiv1} hold. Under Condition \ref{eqn:NPIV-holder},
Assumptions \ref{ass:anor-mtilde} and \ref{ass:cont_diffm} hold for the NPIV
model (\ref{npiv}).
\end{proposition}
Proposition \ref{pro:NPIV-suff} allows for irregular functionals of the NPIV
model with severely ill-posed case.
\begin{condition}
\label{eqn:NPQIV-holder} (i) $\{E[F_{Y_{1}|Y_{2}X}(h(Y_{2}),Y_{2},\cdot
)|\cdot ]:h\in \mathcal{H}\}\subseteq \Lambda _{c}^{\gamma }(\mathcal{X})$,
with $\gamma >0.5$; (ii) $\sup_{y_{1},y_{2},x}|\frac{
df_{Y_{1}|Y_{2}X}(y_{1},y_{2},x)}{dy_{1}}|\leq C<\infty $.
\end{condition}
\begin{condition}
\label{con:NPQIV-rate1} $n(\log \log n)^{4}\delta _{s,n}^{4}=o(1)$
\end{condition}
\begin{proposition}
\label{pro:NPQIV-suff} Let all conditions for Remark \ref
{remark-normality-gnpiv1} hold. Under conditions \ref{eqn:NPIV-holder}(ii)
and \ref{eqn:NPQIV-holder}-\ref{con:NPQIV-rate1}, Assumptions \ref
{ass:anor-mtilde} and \ref{ass:cont_diffm} hold for the NPQIV model (\ref
{npqiv}).
\end{proposition}
It is clear that Condition \ref{con:NPQIV-rate1} rules out severely
ill-posed case, and hence Proposition \ref{pro:NPQIV-suff} only allows for
irregular functionals of the NPQIV model with mildly ill-posed case.
\section{Simulation Studies and An Empirical Illustration}
\label{sec:sec_simulation}
This section first presents simulation studies for SQLR and sieve t tests of
linear and nonlinear hypotheses for the NPQIV and NPIV models respectively.
It then provides an empirical illustration of the optimally weighted SQLR
inferences for a NPQIV Engel curve. In this section, we use the series LS
estimator (\ref{mhat}) of $m(x,h)$ with $p^{J_{n}}(x)$ as its basis, and $
q^{k(n)}$ as the basis approximating the unknown structure function $h_{0}$.
We use $p^{J}=\mathrm{P-Spline}(r,k)$ to denote $r$th degree polynomial
spline with $k$ (quantile) equally spaced knots, hence $J=(r+1)+k$ is the
total number of sieve terms. We use $p^{J}=\mathrm{Pol}(J)$ to denote power
series up to $(J-1)$th degree. See, e.g., \cite{C_bookchp07} for definitions
of these and other sieve bases.
\subsection{Simulation Studies\label{sec:sec_simulation1}}
We run Monte Carlo (MC) studies to assess the finite sample performance of
SQLR and sieve t tests of linear and nonlinear hypotheses in two models: the
NPQIV (\ref{npqiv}) and the NPIV (\ref{npiv}).
For all cases, our design is based on the MC design of \cite{NP_ECMA03} and
\cite{Santos_ECMA} for a NPIV model, which we adapt to cover both NPIV and
NPQIV models. Specifically, we generate i.i.d. draws of $(Y_{2},X,U^{\ast })$
from
\begin{equation*}
\left[
\begin{array}{c}
Y_{2}^{\ast } \\
X^{\ast } \\
U^{\ast }
\end{array}
\right] \sim N\left( 0,
\begin{bmatrix}
1 & 0.8 & 0.5 \\
0.8 & 1 & 0 \\
0.5 & 0 & 1
\end{bmatrix}
\right) ,
\end{equation*}
and $Y_{2}=2(\Phi (Y_{2}^{\ast }/3)-0.5)$ and $X=2(\Phi (X^{\ast }/3)-0.5)$.
The true function $h_{0}$ is given by $h_{0}(\cdot )=2\sin (\pi \cdot )$. We
consider 5,000 MC repetitions and $n=750$ for each of the cases studied
below. We use $Pen(h)=||h||_{L^{2}}^{2}+||\nabla h||_{L^{2}}^{2}$ in all the
simulations, and have used a very small $\lambda _{n}=10^{-5}$ in most cases
(except for the cases where we study the sensitivity to the choice of $\lambda
_{n} $).
\textbf{Summary of sensitivity checks}: For NPQIV and NPIV models, for both
SQLR and sieve t tests of linear and nonlinear hypotheses, as long as $
J_{n}>k(n)+1$ with not too large $k(n)$, the MC sizes of the tests are good
and insensitive to the choices of basis $q^{k(n)}$ and $p^{J_{n}}$ or the
very small penalty $\lambda _{n}$. This is consistent with previous MC
findings in \cite{BCK_Emetrica07} and \cite{CP_WP07} for PSMD estimation of
NPIV and NPQIV respectively.
\medskip
\noindent \textbf{NPQIV model: SQLR test for an irregular linear functional.}
We consider the NPQIV model $Y_{1}=h_{0}(Y_{2})+U=2\sin (\pi Y_{2})+U$ with $
U=2(\Phi (U^{\ast })-\gamma )$. This last transformation is done to ensure
that $E[1\{U\leq 0\}|X]=\gamma $. To save space we only present the case
with $\gamma =0.5$. The parameter of interest is $\phi (h_{0})=h_{0}(0)$,
hence $\phi $ is an irregular linear functional. We study the finite sample
properties of the SQLR and bootstrap-SQLR tests. The SQLR-based confidence
intervals are specially well-suited for models like NPQIV where the
generalized residual function is non-smooth yet the optimal weighting matrix
is easy to compute.
\textbf{Size}. Table \ref{tab:NPQIV-SQLR} reports the simulated size of the
SQLR test of $H_{0}\colon \phi (h_{0})=0$ as a function of the nominal size
(NS), for different choices of $q^{k(n)}$ and $p^{J_{n}}$, and different
values of the tuning parameters $(\lambda _{n},k(n),J_{n})$.
\begin{table}[h]
\centering
\caption{Size of the SQLR test of $\protect\phi
(h_{0}) = 0$ for NPQIV model.}
\begin{tabular}{cccccc}
\hline\hline
$q^{k(n)}$ & $p^{J_{n}}$ & $\lambda_{n}$ & 10\% & 5\% & 1\% \\ \hline
\multirow{3}{*}{Pol(4)} & Pol(7) & {\small {$(1 \times 10^{-3})$ }} & 0.099 &
0.055 & 0.008 \\
& Pol(7) & {\small {$(2 \times 10^{-4})$ }} & 0.096 & 0.048 & 0.008 \\
& Pol(7) & {\small {$(4 \times 10^{-5})$ }} & 0.107 & 0.053 & 0.010 \\ \hline
\multirow{3}{*}{Pol(6)} & Pol(7) & {\small {$(1 \times 10^{-3})$ }} & 0.133 &
0.068 & 0.011 \\
& Pol(7) & {\small {$(2 \times 10^{-4})$ }} & 0.091 & 0.036 & 0.006 \\
& Pol(7) & {\small {$(4 \times 10^{-5})$ }} & 0.105 & 0.052 & 0.008 \\ \hline
\multirow{3}{*}{Pol(6)} & Pol(9) & {\small {$(1 \times 10^{-5})$ }} & 0.107 &
0.055 & 0.012 \\
& Pol(15) & {\small {$(1 \times 10^{-5})$ }} & 0.109 & 0.058 & 0.014 \\
& Pol(21) & {\small {$(1 \times 10^{-5})$ }} & 0.112 & 0.058 & 0.013 \\
\hline
\multirow{4}{*}{P-Spline(3,2)} & Pol(9) & {\small {$(1 \times 10^{-5})$ }} &
0.103 & 0.049 & 0.010 \\
& Pol(10) & {\small {$(1 \times 10^{-5})$ }} & 0.104 & 0.051 & 0.010 \\
& Pol(15) & {\small {$(1 \times 10^{-5})$ }} & 0.105 & 0.049 & 0.009 \\
& Pol(21) & {\small {$(1 \times 10^{-5})$ }} & 0.105 & 0.052 & 0.009 \\
\hline
\multirow{3}{*}{P-Spline(3,2)} & P-Spline(5,3) & {\small {$(1 \times 10^{-5})$
}} & 0.098 & 0.049 & 0.008 \\
& P-Spline(5,9) & {\small {$(1 \times 10^{-5})$ }} & 0.103 & 0.050 & 0.009
\\
& P-Spline(5,18) & {\small {$(1 \times 10^{-5})$ }} & 0.106 & 0.051 & 0.009
\\ \hline
\end{tabular}
\label{tab:NPQIV-SQLR}
\end{table}
\noindent Table \ref{tab:NPQIV-SQLR} shows that for small value of $k(n)$,
say in $(k(n),J_{n})=(4,7)$ (i.e., rows 1-3), the SQLR test performs well
and is fairly insensitive to different choices of $\lambda _{n}$. For a
fixed relatively small $J_{n}=7$, rows 1-6 indicate that as $k(n)$
increases, the results become a bit more sensitive to the choice of $\lambda
_{n}$. For a fixed very small penalty $\lambda _{n}=10^{-5}$, rows 7-16 show
that the results are fairly insensitive to different choices of $J_{n}$ and
basis for $p^{J_{n}}$ and $q^{k(n)}$ as long as $J_{n}>k(n)+1$.
\textbf{Local power}. Figure \ref{fig:SQLR-NPQIV}
shows the rejection probabilities at 5\% (lower panel) and 1\% (upper panel) level of the null hypothesis as a
function of $r$ where $r\colon \phi (h_{0})=r$ for the SQLR (solid red line) and the bootstrap SQLR (dashed blue line) with multinomial weights. To save space we only report the local power results corresponding to the case of P-Spline(3,2) for $q^{k(n)}$, Pol(10) for $p^{J_{n}} $ and $\lambda
_{n}= 10^{-5}$ in Table \ref{tab:NPQIV-SQLR}. We employ 500 bootstrap evaluations per MC replication,
and lower the number of MC repetitions to 1,000 to ease the computational
burden. We note that since our functional $\phi (h)=h(0)$ is
estimated at a slower than root-$n$ rate, the deviations considered for $r$
which are in the range of $[0,8/\sqrt{n}]$ are indeed
\textquotedblleft small\textquotedblright.
We can see from the figure that the
bootstrap SQLR performance is similar to its non-bootstrapped counterpart.
We expect that the performance will improve if we increase number of
bootstrap runs. (We also run simulation studies corresponding to the case of Pol(4) for $q^{k(n)}$, Pol(7) for $p^{J_{n}} $ and $\lambda _{n}=2\times 10^{-4}$ in Table \ref{tab:NPQIV-SQLR}, and the local power patterns are similar to the ones reported here.)
\begin{figure}[h]
\centering
\includegraphics[height=1.5in,width=4in]{power-SQLR-NPQIV-1000MC-001.pdf}
\includegraphics[height=1.5in,width=4in]{power-SQLR-NPQIV-1000MC-050.pdf}
\caption{Rejection probabilities at 1\% (upper panel) and 5\% level (lower panel) of the
null hypothesis as a function of $r = \protect\phi(h_{0})$ for the SQLR
(solid red line) and for the bootstrap SQLR (dashed blue line) for NPQIV.}
\label{fig:SQLR-NPQIV}
\end{figure}
\medskip
\noindent \textbf{NPIV model: sieve variance estimators for an irregular
linear functional.} We now consider the NPIV model: $
Y_{1}=h_{0}(Y_{2})+0.76U=2\sin (\pi Y_{2})+0.76U$, with $U=U^{\ast }$ so the
identifying condition of NPIV holds: $E[U|X]=0$. The parameter of interest
is $\phi (h_{0})=h_{0}(0)$, and the null hypothesis is $H_{0}\colon \phi
(h_{0})=0$. We focus on the finite sample performance of the sieve variance
estimators for irregular linear functionals. We compute two sieve variance
estimators:
\begin{equation*}
\widehat{V}_{1}=q^{k(n)}(0)^{\prime }\widehat{D}_{n}^{-1}\widehat{\mho }_{n}
\widehat{D}_{n}^{-1}q^{k(n)}(0)\text{\quad and\quad }\widehat{V}
_{2}=q^{k(n)}(0)^{\prime }\widehat{D}_{n}^{-1}\widehat{\Omega }_{n}\widehat{D
}_{n}^{-1}q^{k(n)}(0),
\end{equation*}
where $\widehat{D}_{n}=n^{-1}\left( \widehat{C}_{n}(P^{\prime }P)^{-}
\widehat{C}_{n}^{\prime }\right) $, $\widehat{C}_{n}\equiv
\sum_{i=1}^{n}q^{k(n)}(Y_{2i})p^{J_{n}}(X_{i})^{\prime }$, $\widehat{\mho }
_{n}$ is given in equation (\ref{2sls-vhat}), and $\widehat{\Omega }_{n}=
\frac{1}{n}\widehat{C}_{n}(P^{\prime }P)^{-}\left(
\sum_{i=1}^{n}p^{J_{n}}(X_{i})\widehat{\Sigma }_{0}(X_{i})p^{J_{n}}(X_{i})^{
\prime }\right) (P^{\prime }P)^{-}\widehat{C}_{n}^{\prime }$ with $\widehat{U
}_{j}=Y_{1j}-\widehat{h}(Y_{2j})$ and $\widehat{\Sigma }_{0}(x)=\left(
\sum_{j=1}^{n}\widehat{U}_{j}^{2}p^{J_{n}}(X_{j})^{\prime }\right)
(P^{\prime }P)^{-}p^{J_{n}}(x)$. (See Theorem \ref{thm:VE2} in Appendix \ref
{app:appB} for the definition and consistency of $\widehat{V}_{2}$ as
another sieve variance estimator for any plug-in PSMD $\phi (\widehat{\alpha
})$.)
\begin{table}[h]
\centering
\caption{Relative performance of $\hat{V}_{1}$ and $
\hat{V}_{2}$: $Med_{MC} \left[\left|\frac{\hat{V}_{j}}{||v^{
\ast}_{n}||^{2}_{sd}}-1\right| \right]$, and Nominal size and MC rejection
frequencies for t tests $\hat{t}_{j}$ for $j=1,2$ for a linear functional of
NPIV.}
\begin{tabular}{cccccccc}
\hline\hline
& & \multicolumn{2}{c}{$Med_{MC}$} & \multicolumn{2}{c}{5\%} &
\multicolumn{2}{c}{10\%} \\ \hline
$q^{k(n)}$ & $p^{J_{n}}$ & $\widehat{V}_{1}$ & $\widehat{V}_{2}$ & $\widehat{
V}_{1}$ & $\widehat{V}_{2}$ & $\widehat{V}_{1}$ & $\widehat{V}_{2}$ \\ \hline
\multirow{4}{*}{Pol(4)} & Pol(6) & 0.0946 & 0.0937 & 0.0512 & 0.0514 & 0.0980
& 0.0974 \\
& Pol(10) & 0.0922 & 0.0920 & 0.0536 & 0.0532 & 0.0992 & 0.0990 \\
& Pol(12) & 0.0918 & 0.0917 & 0.0538 & 0.0532 & 0.1002 & 0.0998 \\
& Pol(16) & 0.0911 & 0.0912 & 0.0540 & 0.0538 & 0.1000 & 0.0998 \\ \hline
\multirow{4}{*}{Pol(4)} & P-Spline(3,2) & 0.0939 & 0.0942 & 0.051 & 0.0516 &
0.0984 & 0.0986 \\
& P-Spline(3,5) & 0.0939 & 0.0920 & 0.053 & 0.0532 & 0.0990 & 0.0984 \\
& P-Spline(3,11) & 0.0923 & 0.0925 & 0.055 & 0.0548 & 0.1014 & 0.1014 \\
& P-Spline(3,17) & 0.0922 & 0.0917 & 0.0542 & 0.0538 & 0.100 & 0.1008 \\
\hline
\multirow{4}{*}{P-Spline(3,2)} & Pol(12) & 0.0938 & 0.0930 & 0.0572 & 0.0564 &
0.1082 & 0.1074 \\
& Pol(16) & 0.0936 & 0.0936 & 0.0582 & 0.0578 & 0.1082 & 0.1082 \\
& Pol(18) & 0.0936 & 0.0935 & 0.0580 & 0.0578 & 0.1088 & 0.1086 \\
& Pol(20) & 0.0936 & 0.0937 & 0.0580 & 0.0574 & 0.1086 & 0.1092 \\ \hline
\multirow{4}{*}{P-Spline(3,2)} & P-Spline(3,2) & 0.1106 & 0.1116 & 0.0606 &
0.0598 & 0.1130 & 0.1120 \\
& P-Spline(3,5) & 0.1019 & 0.1023 & 0.0584 & 0.0574 & 0.1122 & 0.1116 \\
& P-Spline(3,11) & 0.0961 & 0.0960 & 0.0572 & 0.0566 & 0.1100 & 0.1094 \\
& P-Spline(3,17) & 0.0949 & 0.0944 & 0.0570 & 0.0566 & 0.1082 & 0.1080 \\
\hline
\multirow{4}{*}{P-Spline(3,2)} & P-Spline(5,3) & 0.1007 & 0.0998 & 0.0586 &
0.0576 & 0.1102 & 0.1088 \\
& P-Spline(5,6) & 0.1011 & 0.1009 & 0.0586 & 0.0578 & 0.1100 & 0.1092 \\
& P-Spline(5,12) & 0.1007 & 0.1009 & 0.0580 & 0.0572 & 0.1110 & 0.1096 \\
& P-Spline(5,18) & 0.1009 & 0.1010 & 0.0580 & 0.0570 & 0.1106 & 0.1092 \\
\hline
\end{tabular}
\label{tab:ve}
\end{table}
Table \ref{tab:ve} reports the results for different choices of bases for $
q^{k(n)}$ and $p^{J_{n}}$, and for different values of $k(n)$ and $J_{n}$;
in all cases we use a very small $\lambda _{n}=10^{-5}$. This table shows $
Med_{MC}\left[ \left\vert \frac{\widehat{V}_{j}}{||v_{n}^{\ast }||_{sd}^{2}}
-1\right\vert \right] $ for $j=1,2$, where $||v_{n}^{\ast }||_{sd}$ is
computed using the MC variance of $\sqrt{n}\widehat{h}_{n}(0)$ and $
Med_{MC}[\cdot ]$ is the MC median. It also shows the nominal size and MC
rejection frequencies of the two sieve t tests $\widehat{t}_{j}=\sqrt{n}
\frac{\widehat{h}_{n}(0)-0}{\sqrt{\widehat{V}_{j}}}$ for $j=1,2$.
We note that the two sieve variance estimators have almost identical
performance and the associated sieve t tests have good rejection
probabilities. These results are fairly robust to different choices of basis
for $q^{k(n)}$ and $p^{J_{n}}$ and different values of $k(n)$ and $J_{n}$ as
long as $J_{n}>k(n)+1$. Figure \ref{fig:QQplot-MC3-MC4} (first row) shows
the QQ-Plot for the sieve t tests $\widehat{t}_{j}=\sqrt{n}\frac{\widehat{h}
_{n}(0)-0}{\sqrt{\widehat{V}_{j}}}$ under the null for $j=1,2$ for the case
Pol(4)-Pol(16) in the table; the right panel in the first row corresponds to
$\hat{t}_{1}$ and the left panel in the first row to $\hat{t}_{2}$. Both
sieve t tests are almost identical to each other and to the standard normal.
\begin{figure}[tbp]
\centering
\includegraphics[height=2.5in,width=4in]{qqplot-aug-16-14-crop.pdf}
\caption{QQ-Plot for t tests $\hat{t}_{j}$ for $j=1,2$ for a linear functional (first row) and a nonlinear functional (second row) of NPIV, with
$q^{k(n)}=Pol(4)$ and $p^{J_{n}}=Pol(16)$.}
\label{fig:QQplot-MC3-MC4}
\end{figure}
\medskip
\noindent \textbf{NPIV model: sieve variance estimators for an irregular
\emph{nonlinear} functional.} This case is identical to the previous one for
the NPIV model, except that the functional of interest is $\phi (h_{0})=\exp
\{h_{0}(0)\}$, and the null hypothesis is $H_{0}\colon \phi (h_{0})=1$. This
choice of $\phi $ allows us to evaluate the finite sample performance of
sieve t statistics for a nonlinear functional.
\begin{table}[h]
\centering
\caption{Relative performance of $\hat{V}_{1}$ and $
\hat{V}_{2}$: $Med_{MC} \left[\left|\frac{\hat{V}_{j}}{||v^{
\ast}_{n}||^{2}_{sd}}-1\right| \right]$, and Nominal size and MC rejection
frequencies for t tests $\hat{t}_{j}$ for $j=1,2$ for a nonlinear functional
of NPIV.}
\begin{tabular}{cccccccc}
\hline\hline
& & \multicolumn{2}{c}{$Med_{MC}$} & \multicolumn{2}{c}{5\%} &
\multicolumn{2}{c}{10\%} \\ \hline
$q^{k(n)}$ & $p^{J_{n}}$ & $\widehat{V}_{1}$ & $\widehat{V}_{2}$ & $\widehat{
V}_{1}$ & $\widehat{V}_{2}$ & $\widehat{V}_{1}$ & $\widehat{V}_{2}$ \\ \hline
\multirow{4}{*}{Pol(4)} & Pol(6) & 0.0990 & 0.0985 & 0.0528 & 0.0530 & 0.0982
& 0.0988 \\
& Pol(10) & 0.0971 & 0.0958 & 0.0524 & 0.0522 & 0.1014 & 0.1012 \\
& Pol(12) & 0.0967 & 0.0959 & 0.0526 & 0.0526 & 0.1020 & 0.1018 \\
& Pol(16) & 0.0961 & 0.0958 & 0.0524 & 0.0528 & 0.1018 & 0.1014 \\ \hline
\multirow{4}{*}{Pol(4)} & P-Spline(3,2) & 0.0996 & 0.0983 & 0.0534 & 0.0530 &
0.0978 & 0.0976 \\
& P-Spline(3,5) & 0.0982 & 0.0969 & 0.0538 & 0.0542 & 0.0990 & 0.0992 \\
& P-Spline(3,11) & 0.0985 & 0.0984 & 0.0554 & 0.0552 & 0.1014 & 0.1010 \\
& P-Spline(3,17) & 0.0982 & 0.0978 & 0.0544 & 0.0546 & 0.1010 & 0.1008 \\
\hline
\multirow{4}{*}{P-Spline(3,2)} & Pol(12) & 0.1011 & 0.1009 & 0.0580 & 0.0568 &
0.1120 & 0.1122 \\
& Pol(16) & 0.1014 & 0.1005 & 0.0588 & 0.0574 & 0.1128 & 0.1126 \\
& Pol(18) & 0.1014 & 0.1007 & 0.0582 & 0.0568 & 0.1130 & 0.1122 \\
& Pol(20) & 0.1015 & 0.1006 & 0.0580 & 0.0568 & 0.1138 & 0.1128 \\ \hline
\multirow{4}{*}{P-Spline(3,2)} & P-Spline(3,2) & 0.1191 & 0.1192 & 0.0620 &
0.0612 & 0.1132 & 0.1120 \\
& P-Spline(3,5) & 0.1090 & 0.1103 & 0.0596 & 0.0594 & 0.1140 & 0.1134 \\
& P-Spline(3,11) & 0.1028 & 0.1032 & 0.0582 & 0.0572 & 0.1130 & 0.1126 \\
& P-Spline(3,17) & 0.1029 & 0.1029 & 0.0588 & 0.0580 & 0.1124 & 0.1112 \\
\hline
\multirow{4}{*}{P-Spline(3,2)} & P-Spline(5,3) & 0.1059 & 0.1064 & 0.0594 &
0.0592 & 0.1114 & 0.1104 \\
& P-Spline(5,6) & 0.1066 & 0.1076 & 0.0598 & 0.0586 & 0.1124 & 0.1118 \\
& P-Spline(5,12) & 0.1071 & 0.1079 & 0.0594 & 0.0586 & 0.1126 & 0.1120 \\
& P-Spline(5,18) & 0.1069 & 0.1079 & 0.0594 & 0.0586 & 0.1122 & 0.1120 \\
\hline
\end{tabular}
\label{tab:ve1}
\end{table}
Table \ref{tab:ve1} shows $Med_{MC}$ and rejection probabilities for this
nonlinear case. By comparing the results with those in Table \ref{tab:ve} we
note that the results are very similar in both cases. Figure \ref
{fig:QQplot-MC3-MC4} (second row) shows the QQ-Plot for the two sieve t
tests for the non-linear case; the right panel in the second row corresponds
to $\hat{t}_{1}$ whereas the left panel in the second row corresponds to $
\hat{t}_{2}$. These results suggest that our sieve t tests perform equally
well for both functionals.
Finally we wish to point out that we have tried other bases such as Hermite
polynomials and cosine series and even larger $J_{n}$ in these two NPIV MC
studies, the results are all similar to the ones reported here and hence are
not presented due to the lack of space.
\subsection{An Empirical Application}
\label{sec:application}
We compute SQLR based confidence bands for nonparametric quantile IV Engel
curves using the British FES data set from \cite{BCK_Emetrica07}:
\begin{equation*}
E[1\{Y_{1,i}\leq h_{0}(Y_{2,i})\}\mid X_{i}]=0.5,
\end{equation*}
where $Y_{1,i}$ is the budget share of the $i-$th household on a particular
non-durable goods, say food-in consumption; $Y_{2,i}$ is the log-total
expenditure of the household, which is endogenous, and hence we use $X_{i}$,
the gross earnings of the head of the household, to instrument it. We work
with the \textquotedblleft no kids\textquotedblright\ sub-sample of the data
set, which consists of $n=628$ observations. \cite{BCK_Emetrica07} estimated
NPIV Engel curves using this data set. But, as explained by \cite
{Koenker_2005} and others, quantile Engel curves are more informative.
We estimate $h_{0}(\cdot )$ for food-in quantile Engel curve via the
optimally weighted PSMD procedure with $\widehat{\Sigma }=\Sigma _{0}=0.25$,
using a polynomial spline (P-spline) sieve $\mathcal{H}_{k(n)}$ with $k(n)=4$
, $Pen(h)=||h||_{L^{2}}^{2}+||\nabla h||_{L^{2}}^{2}$ with $\lambda
_{n}=0.0005$, and a Hermite polynomial LS basis $p^{J_{n}}(X)$ with $J_{n}=6$
. We also considered other bases such as P-splines as $p^{J_{n}}(X)$ and
results remained essentially the same. See \cite{CP_WP07a} for PSMD
estimates of NPQIV Engel curves for other non-durable goods.
We use the fact that the optimally weighted SQLR of testing $\phi
(h)=h(y_{2})$ (for any fixed $y_{2}$) is asymptotically $\chi _{1}^{2}$ to
construct pointwise confidence bands. That is, for each $y_{2}$ in the
sample we construct a grid, $(r_{i} )_{i=1}^{30}$. For each $i={1,...,30}$,
we compute the value of the SQLR test statistic under $h(y_{2})=r_{i}$ for $
(r_{i})_{i=1}^{30}$. We then, take the smallest interval that included all
points $r_{i}$ that yield a corresponding value of the SQLR test below the
95\% percentile of $\chi _{1}^{2}$.\footnote{
The grid $(r_{i})_{i=1}^{30}$ was constructed to have $r_{15}=\widehat{h}
_{n}(y_{2})$, for all $i\leq 15$, $r_{i+1}\leq r_{i}\leq r_{15}$ decreasing
in steps of length $0.002$ (approx) and for all $i\geq 15$, $r_{i+1}\geq
r_{i}\geq r_{15}$ increasing in steps of length $0.008$ (approx); finally,
the extremes, $r_{1}$ and $r_{30}$, were chosen so the SQLR test at those
points was above the 95\% percentile of $\chi _{1}^{2}$. We tried different
lengths and step sizes and the results remain qualitatively unchanged. For
some observations, which only account for less than 4\% of the sample, the
confidence interval was degenerate at a point; this result was due to
numerical approximation issues, and these observations were excluded from the reported
results.} Figure \ref{fig:EC} presents the results, where the solid blue
line is the point estimate and the red dashed lines are the 95\% pointwise
confidence bands. We can see that the confidence bands get wider towards the
extremes of the sample, but are tighter in the middle.
To test whether the quantile IV Engel curve for food-in is linear or not,
one can test whether $\phi (h_{0})\equiv \int \left\vert \nabla
^{2}h(y_{2})\right\vert ^{2}w(y_{2})dy_{2}=0$ using our SQLR test. Let $
w(\cdot )=(\sigma _{Y_{2}})^{-1}\exp \left( -\frac{1}{2}(\sigma
_{Y_{2}}^{-1}(\cdot -\mu _{Y_{2}}))^{2}\right) 1\{t_{0.01}\leq \cdot \leq
t_{0.99}\}$ where $\mu _{Y_{2}}$, $\sigma _{Y_{2}}$, $t_{0.01}$ and $
t_{0.99} $ are the sample mean, standard deviation and the 1\% and 99\%
quantiles of $Y_{2}$. The value of the SQLR is (approx.) 38 and the p-value
is smaller than 0.0001, and hence we reject the null hypothesis of linearity.
\footnote{
We use the standard Riemann sum with 1000 terms to compute the integral. We
also considered other choices of $w$ such that $w(\cdot )=1\{t_{0.25}\leq
\cdot \leq t_{0.75}\}$ and $w(\cdot )=1\{t_{0.01}\leq \cdot \leq t_{0.99}\}$
. Although the numerical value of the SQLR test changes, all produce
p-values below 0.0001.}
\begin{figure}[tbp]
\centering
\includegraphics[height=3in,width=4.5in]{feb-09-12.pdf}
\caption{{\protect\footnotesize {PSMD Estimate of the NPQIV food-in Engel
curve (blue solid line), with the 95\% pointwise confidence bands (red dash
lines).}}}
\label{fig:EC}
\end{figure}
\section{Conclusion}
\label{sec:conclusion}
In this paper, we provide unified asymptotic theories for PSMD based
inferences on possibly irregular parameters $\phi (\alpha _{0})$ of the
general semi/nonparametric conditional moment restrictions $E[\rho
(Y,X;\alpha _{0})|X]=0$. Under regularity conditions that allow for any
consistent nonparametric estimator of the conditional mean function $
m(X,\alpha )\equiv E[\rho (Y,X;\alpha )|X]$, we establish the asymptotic
normality of the plug-in PSMD estimator $\phi (\widehat{\alpha }_{n})$ of $
\phi (\alpha _{0})$, as well as the asymptotically tight distribution of a
possibly non-optimally weighted SQLR statistic under the null hypothesis of $
\phi (\alpha _{0})=\phi _{0}$. As a simple yet useful by-product, we
immediately obtain that an optimally weighted SQLR statistic is
asymptotically chi-square distributed under the null hypothesis. For
(pointwise) smooth residuals $\rho (Z;\alpha )$ (in $\alpha $), we propose
several simple consistent sieve variance estimators for $\phi (\widehat{
\alpha }_{n})$ (in the text and in online Appendix \ref{app:appB}), and
establish the asymptotic chi-square distribution of sieve Wald statistics.
We also establish local power properties of SQLR and sieve Wald tests in
Appendix \ref{app:appA}. Under conditions that are virtually the same as
those for the limiting distributions of the original-sample sieve Wald and
SQLR statistics, we establish the consistency of the generalized residual
bootstrap sieve Wald and SQLR statistics. All these results are valid
regardless of whether $\phi (\alpha _{0})$ is regular or not. While SQLR and
bootstrap SQLR are useful for models with (pointwise) non-smooth $\rho
(Z;\alpha )$, sieve Wald statistic is computationally attractive for models
with smooth $\rho (Z;\alpha )$. Monte Carlo studies and an empirical
illustration of a nonparametric quantile IV regression demonstrate the good
finite sample performance of our inference procedures.
This paper assumes that the semi/nonparametric conditional moment
restrictions $E[\rho (Y,X;\alpha _{0})|X]=0$ uniquely identifies the unknown
true parameter value $\alpha _{0}\equiv (\theta _{0}^{\prime },h_{0})$, and
conduct inferences that are robust to whether a possibly nonlinear
functional of $\alpha _{0}$ is root-$n$ estimable or not. Recently, for the
NPIV model $E[Y_{1}-h_{0}(Y_{2})|X]=0$ without assuming point identification
of $h_{0}$, \cite{Santos_JOE} proposed a root-$n$ asymptotically normal
estimation of a regular linear functional of $h_{0}$ and \cite{Santos_ECMA}
considered Bierens' type test of the NPIV. \cite{CPT_WP12} is
extending the SQLR inference procedure to allow for partial identification
of the general model $E[\rho (Y,X;\alpha _{0})|X]=0$.
\baselineskip=15pt
\bigskip
\setstretch{1}
\bibliography{mybib_FUNC}