EconBase
← Back to paper

Penalized Sieve GEL for Weighted Average Derivatives of Nonparametric Quantile IV Regressions

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

79,077 characters

Penalized Sieve GEL for Weighted Average Derivatives of Nonparametric Quantile IV Regressions



\title{Penalized Sieve GEL for Weighted Average Derivatives of Nonparametric
Quantile IV Regressions\thanks{
We appreciate discussions with Roger Koenker about quantiles and other
interesting topics. Roger's creativity, curiosity and kindness have inspired
us all these years. We thank guest editors and anonymous referees for their patience and helpful comments. Any errors are the responsibility of the authors. First version: August 2017.}}
\author{Xiaohong Chen\thanks{
Tel.: +1 203 432 5852; \textit{Email}: [email removed].} , Demian
Pouzo\thanks{
Corresponding author. Tel.: +1 510 642 6709. \textit{Email}: [email removed].},
 and James
L. Powell \thanks{
Tel.: +1 510 643 0709. \textit{Email}: [email removed].}}


\bigskip
\maketitle

\begin{abstract}
\singlespacing This paper considers estimation and inference for a
weighted average derivative (WAD) of a nonparametric quantile instrumental
variables regression (NPQIV). NPQIV is a non-separable and
nonlinear ill-posed inverse problem, which might be why there is no
published work on the asymptotic properties of any estimator of its WAD.
We first characterize the semiparametric
efficiency bound for a WAD of a NPQIV, which, unfortunately, depends on an unknown
conditional derivative operator and hence an unknown degree of ill-posedness, making
it difficult to know if the information bound is singular or not. In either case,
we propose a penalized sieve generalized empirical likelihood (GEL) estimation
and inference procedure, which is based on the unconditional WAD moment
restriction and an increasing number of unconditional moments that are implied by the conditional NPQIV restriction, where the unknown quantile function is approximated by a penalized sieve. Under some regularity conditions, we show that the self-normalized penalized sieve GEL estimator of the WAD of a NPQIV is asymptotically standard normal. We also show that the quasi likelihood ratio statistic based on the
penalized sieve GEL criterion is asymptotically chi-square distributed
regardless of whether or not the information bound is singular.

\medskip

\smallskip

\noindent \emph{JEL Classification:} C14; C22

\medskip

\noindent \emph{Keywords:} Nonparametric quantile instrumental variables;
Weighted average derivatives; Penalized sieve generalized empirical
likelihood; Semiparametric efficiency; Chi-square inference.
\end{abstract}

\newpage

\section{Introduction}

Since the seminal paper by \cite{koenker1978regression}, quantile
regressions and functionals of quantile regressions have been the subjects
of ever-expanding theoretical research and applications in economics,
statistics, biostatistics, finance, and many other science and social
science disciplines. See \cite{koenker2005quantile} and the forthcoming
Handbook of Quantile Regression (2017) for the latest theoretical advances and empirical applications.

The presence of endogenous regressors is common in many empirical
applications of structural models in economics and other social sciences.
The Nonparametric Quantile Instrumental Variable (NPQIV) regression, $
E[1\{ Y \leq h_{0}(W) \} - \tau|X]=0$, was, to our knowledge, first proposed in
\cite{chernozhukov2005iv} and \cite{CIN2007instrumental}. This model is a
leading important example of nonlinear and non-separable ill-posed inverse
problems in econometrics, which has been an active research topic following the Nonparametric (mean) Instrumental Variables (NPIV) regression, $E[ Y  - h_{0}(W)|X]=0$, studied by \cite{NP2003instrumental}, \cite{hall2005nonparametric}, \cite{BCK2007}, \cite{CFR2007linear}, \cite{darolles2011nonparametric} and others. See, for example, \cite{horowitz2007nonparametric},
\cite{CP2009efficient,CP2012estimation,CP2015sieve}, \cite{gagliardini2012nonparametric}, \cite{CH2013quantilerev}, \cite{CCLN2014local} and others for recent work on the NPQIV and its
various extensions.


In this paper, we consider
estimation and inference for a Weighted
Average Derivative (WAD) functional of a NPQIV. For models without nonparametric
endogeneity, WAD functionals of nonparametric (conditional) mean regression, $E[ Y  - h_{0}(X)|X]=0$, and of quantile regression, $E[1\{ Y \leq h_{0}(X) \} - \tau|X]=0$, have been extensively studied in both statistics and
econometrics. In particular, under some mild regularity conditions, plug-in estimators for WADs of any nonparametric mean and quantile regressions can be shown to be semiparametrically efficient
and root-$n$ asymptotically normal (where $n$ is the sample size). See, for example, \cite{newey1993efficiency},
\cite{newey1994asymptotic}, \cite{NeweyPowell1999}, \cite
{ackerberg2014asymptotic} and the references therein. Although unknown functions of endogenous
regressors occur frequently in empirical work, due to the ill-posed nature
of NPIV and NPQIV, there is not much research on WAD functionals of NPIV and
NPQIV yet. In fact, even for the simpler NPIV model that is a linear and
separable ill-posed inverse problem, it is still a difficult question whether a
linear functional of a NPIV could be estimated at the root-$n$ rate; see, e.g.,
\cite{SeveriniTripathi2012} and \cite{davezies2015existence}. Although \cite
{ai2007estimation} provide low-level sufficient conditions for a root-$n$
consistent and asymptotically normal estimator of the WAD of the NPIV model, and \cite
{ai2012semiparametric} provide a semiparametric efficient estimator of WAD
for that model, to our knowledge, there is no published work on semiparametric efficient estimation
of the WAD for the NPQIV model yet.

We first characterize the semiparametric efficiency bound for the WAD functional of a
NPQIV model. Unfortunately, the bound depends on an unknown conditional derivative
operator and hence an unknown degree of ill-posedness. Therefore, it is difficult to
know if the semiparametric information bound is singular or not. Further, even if a researcher
assumes that the information bound is non-singular and the WAD is root-$n$
consistently estimable, the results in \cite{ai2012semiparametric} and \cite
{CS2015overidentification} show that a simple plug-in estimator of a WAD
might not be semiparametrically efficient. This is in contrast to the
results of \cite{newey1993efficiency} and \cite{ackerberg2014asymptotic}
who show that plug-in estimators of a WAD of a nonparametric mean and quantile
regression are semiparametrically efficient.

We then propose penalized sieve Generalized Empirical Likelihood
(GEL) estimation of the WAD for the NPQIV model, which is based on the unconditional
WAD moment restriction and an increasing number of unconditional moments implied by the conditional moment restriction of the NPQIV model, where the unknown quantile function is approximated by a flexible penalized sieve.
Under some regularity conditions, we show that the self-normalized penalized sieve GEL estimator of the WAD of a NPQIV is asymptotically standard normal. We also show that the Quasi Likelihood Ratio (QLR) statistic based on the penalized sieve GEL criterion is asymptotically chi-squared distributed
regardless of whether the information bound is singular or not; this can be
used to construct confidence sets for the WAD of NPQIV without the need to
estimate the variance nor the need to know the precise convergence rates of
the WAD estimator.

Our estimation procedure builds upon \cite{DIN03}, who approximate a
conditional moment restriction $E[\rho (Y,\theta_0 )|X]=0$ by an increasing
sequence of unconditional moment restrictions, and then consider estimation of the Euclidean parameter $\theta_0$ (of fixed and finite dimension) and specification
tests based on GEL (and related) procedures. For the same model $E[\rho (Y,\theta_0 )|X]=0$, \cite{kitamura2004empirical}
directly estimate the conditional moment restriction via kernel and then
apply a kernel-based conditional empirical likelihood (EL) to estimate $
\theta_0$. However, the model considered in these papers does not contain any
unknown functions  (say $h()$) and the residuals $\rho (.,\theta)$ are assumed
to be twice continuously differentiable with respect to $\theta$ at $
\theta_0 $. For the semiparametric conditional moment restriction $E[\rho (Y,\theta_0,
h_0 (\cdot) )|X]=0$ when the
unknown function $h(\cdot )$ could depend on an endogenous variable, \cite{otsu2011large} and \cite{tao2013empirical} consider a
sieve conditional EL extension of \cite
{kitamura2004empirical}, and \cite{sueishi2017} provides a sieve unconditional GEL extension of \cite{DIN03}, where the unknown function $h(.)$ is approximated by a finite dimensional linear sieve (series) as in \cite{ai2003efficient}. However, like
\cite{ai2003efficient}, all these papers assume twice continuously differentiable residuals $
\rho(.,\theta, h(.))$ with respect to $(\theta_0, h_0(.))$, and hence rule out
the NPQIV model.

\cite{ParenteSmith2011gel} study GEL properties for non-smooth residuals $
g(.,.)$ in the unconditional moment models $E[g(Y,\theta _{0})]=0$, but
require the dimensions of both $g(.,.)$ and $\theta _{0}$ to be fixed and finite.
Finally, \cite{horowitz2007nonparametric}, \cite{gagliardini2012nonparametric}, \cite
{CP2009efficient,CP2012estimation,CP2015sieve}, and \cite{CNS2015constrained}  do include the NPQIV model, but none of these papers addresses the issues of
estimation and inference for the WAD of the NPQIV.

The rest of the paper is organized as follows. Section \ref{sec:model}
introduces notation and the model. Section \ref{sec:eff-bound} characterizes the
semiparametric efficiency bound for the WAD of the NPQIV model. Section \ref
{sec:PSGEL} introduces a flexible penalized sieve GEL procedure. Section \ref
{sec:conv-rate} derives the consistency and the convergence rates of the
penalized sieve GEL estimator for the NPQIV model. Section \ref{sec:ADT}
establishes the asymptotic distributions of the WAD estimator and of the QLR statistic based on penalized sieve GEL for the WAD of a NPQIV. Section \ref{sec:conclusion} concludes with a discussion
of extensions.


\section{Preliminaries and Notation}

\label{sec:model}

Let $Z\equiv (Y,W,X)$ be the observable data vector, where $Y$ is the outcome
variable, $W$ is the endogenous variable and $X$ is the instrumental variable (IV); we assume the observable data, $Z$, is distributed according to a probability distribution $\mathbf{P}$. In order to simplify the exposition, we restrict attention to real-valued continuous random variables, i.e., we assume $\mathbf{P}$ has a density $\mathbf{p}$ with support given by $\mathbb{Z}\equiv \mathbb{Y}\times \mathbb{W}\times
\mathbb{X}\subseteq \mathbb{R}^{3}$; extending our results to vector-valued endogenous and instrumental variables would be straightforward but cumbersome in terms of notation.


\textbf{Notation.} For any subset, $\mathbb{Z}$, of an Euclidean space let $\mathcal{P}(\mathbb{Z})$ be the class of Borel probability measures over $\mathbb{Z}$. For any $P\in \mathcal{P}(\mathbb{Z})$, we use $p$ to denote its probability density function (pdf) (with respect to Lebesgue (Leb) measure) and $supp(P)$ to denote its support. We also use $P_{X}$ ($p_{X}$) to denote the marginal probability (pdf) of a random variable $X$; and $P_{Y|X}$ ($p_{Y|X}$) to denote the conditional probability (pdf) of $Y$ given $X$. For expectation, we write $E_{Q}[.]$ to be explicit about the fact that $Q$ is the measure of integration; throughout we sometimes use $E[.]\equiv E_{\mathbf{P}}[.]$ when $\mathbf{P}$ is the true probability of the data. The term ``wpa1" stands for ``with probability approaching one (under $\mathbf{P}$)"; for any two real-valued sequences $(x_{n},y_{n})_{n}$ $x_{n} \precsim y_{n}$ denotes $x_{n} \leq C y_{n}$ for some $C$ finite and universal; $\succsim$ is defined analogously. For any $q \geq 1$, we use $L^{q}(Q)\equiv L^{q}(\mathbb{Z},Q)$ to denote the class of measurable functions $f : \mathbb{Z} \mapsto \mathbb{R}$ such that $||f||_{L^{q}(Q)}=\left( \int_{z\in \mathbb{Z}}
|f(z)|^{q}Q(dz)\right) ^{1/q} < \infty$; as usual $L^{\infty}(Leb)$ denotes the class of essentially bounded real-valued functions. We use $||.||_{e}$ to denote the Euclidean norm, $\mathbb{R}_{+}=[0,\infty )$ and $\mathbb{R}_{++}=(0,\infty )$.

For any subset $S$ of a vector space $(\mathbb{S},||.||_S)$, $lin \{S\}$ denotes the smallest
linear space containing $S$; for any subspace $A \subseteq \mathbb{S}$, $A^{\perp}$ denotes its orthogonal complement in $(\mathbb{S},||.||_S)$. For any linear operator, $M : (\mathbb{S}_1,||.||_1 ) \rightarrow (\mathbb{S}_2,||.||_2 )$, let $Kernel(M) \equiv \{ x\in \mathbb{S}_1 \colon M[x] = 0  \}$ and $Range(M) \equiv \{ y \in \mathbb{S}_2 \colon \exists x \in \mathbb{S}_1,~M[x] = y  \}$; it is bounded if and only if $\sup_{x\in \mathbb{S}_1 : ||x||_1=1} ||M[x]||_2 <\infty$. For any linear bounded operator $M$, $M^{+}$ denotes its generalized inverse; see, e.g., \cite{engl1996regularization}.


\subsection{The WAD of the NPQIV model}

Let $\mathbb{A}\equiv \mathbb{R}\times \mathbb{H}$, where
$\mathbb{H} = \{ h \in L^{2}(Leb) \colon h^{\prime}~exists~and~||h^{\prime}||_{L^{2}(Leb)} < \infty  \}$, i.e., $\mathbb{H}$ is a Sobolev space of order $1$, here $h^{\prime}$ should be viewed as a weak derivative of $h$ (see \cite{brezis2010functional}). We note that $\mathbb{H}$ is a Hilbert space under the norm $||h||_{\mathbb{H}}\equiv ||h||_{L^{2}(Leb)}+||h'||_{L^{2}(Leb)}$, and $\mathbb{A}$ is a Hilbert space under the norm $||(\theta, h)||_{\mathbb{A}}\equiv ||\theta||_e + ||h||_{\mathbb{H}}$. In this paper we measure convergence in $\mathbb{A}$ using another norm $||(\theta, h)||\equiv ||\theta||_e + ||h||$ for $||h||\leq ||h||_{\mathbb{H}}$ (such as $||h||=||h||_{L^{2}(Leb)}$). The parameter set is given by $\mathcal{A} \equiv \Theta \times \mathcal{H} \subseteq \mathbb{A}$, where $\Theta$ is bounded and convex  and $\mathcal{H}$ is a set that contains additional restrictions on $h \in \mathbb{H}$ which will be specified below. We assume that $\mathbf{P}$ is such that there exists a parameter $\alpha_{0} \equiv (\theta_{0},h_{0}) \in \mathcal{A}$ that satisfies
\begin{align}  \label{eqn:model1}
0 = & E_{\mathbf{P}}[1\{ Y \leq h_{0}(W) \} - \tau|X] \\
\theta_{0} = & E_{\mathbf{P}}[\mu(W)h^{\prime }_{0}(W)] \label{eqn:model2}
\end{align}
for $\tau \in (0,1)$, where $\mu$ is a nonnegative, continuously differentiable scalar function in $L^{\infty}(Leb) \cap \mathbb{H}$ and should be viewed as the weighting function of the average derivative, $\theta_{0}$, of $h_{0}$.


The following assumption ensures that the conditions above uniquely identify $\alpha_{0}$; it will be maintained throughout the paper and will not be explicitly referenced in the results below.

\begin{assumption}
\label{ass:ident} There is a unique $\alpha _{0}\in int(\mathcal{A})$ that
satisfies model (\ref{eqn:model1})-(\ref{eqn:model2}).
\end{assumption}

The interior assumption is needed only for the asymptotic distribution results in Section \ref
{sec:ADT}. In cases where $\mathcal{H}$ has an empty interior, one can use
the concept of relative interior of $\mathcal{H}$.
This assumption is clearly high level. The goal of this paper is to
characterize the asymptotic behavior of a modified GEL estimator of $\alpha $
, taking as given the identification part; for a discussion of primitive
conditions for Assumption \ref{ass:ident}, we refer the reader to \cite{CCLN2014local} and
references therein.


The following assumption
imposes additional restrictions over the primitives: $\mu $, $\mathbf{P}$ and $
\alpha _{0}$.

\begin{assumption}
\label{ass:pdf0} (i) $\mathbf{P}$ has a continuously differentiable pdf, $\mathbf{p}$, such that: the marginal density $\mathbf{p}_{W}$ of $W$ is uniformly
bounded, zero at the boundary of the support and $\mathbf{p}^{\prime}_{W} \in L^{2}(Leb)$; the marginal density $\mathbf{p}_{X}$ of $X$ is uniformly bounded away from 0 on its support; $\sup_{y,w,x \in \mathbb{Z}} \mathbf{p}_{Y|WX}(y
\mid w, x) < \infty$, $\sup_{y,w,x \in \mathbb{Z}} \frac{d\mathbf{p}_{Y|WX}(y \mid w, x)}{dy} < \infty$
;
(ii) $\mathcal{H}$ is convex and such that for all $h \in \mathcal{H}$, $\sup_{w \in \mathbb{W}} |\mu(w) h(w)| < \infty$;
(iii) $Var_{\mathbf{P}}(\mu(W)h^{\prime }(W)) > 0$ for all $h \in \mathcal{H}$ in a $||\cdot ||$-neighborhood of $h_{0}$.
\end{assumption}

Part (i) of this condition imposes differentiability and boundedness
restrictions on different elements of $\mathbf{p}$; part (ii) ensures that $\lim_{w \rightarrow \pm \infty} \mathbf{p}_{W}(w) \mu (w) h(w) = 0$ which allows for an alternative representation for $\theta_{0}$ using integration by parts (see expression \ref{eqn:int-parts}  below); part (iii) is a high level
assumption and essentially implies $Var_{\mathbf{P}}(\mu(W)h^{\prime }_{0}(W)) > 0$ as
well as continuity of $h \mapsto Var_{\mathbf{P}}(\mu(W)h^{\prime }(W))$.



\section{Efficiency Bound for $\theta_{0}$}

\label{sec:eff-bound}

By definition of $\mathbb{H}$, Assumption \ref{ass:pdf0} and integration
by parts, it follows that
\begin{align}\label{eqn:int-parts}
\theta_{0} = E[\mu(W)h_0^{\prime }(W)] = - \int \ell(w) h_0 (w) dw
\end{align}
where $$w \mapsto \ell(w) \equiv \mu^{\prime }(w) \mathbf{p}_{W}(w) + \mu(w)
\mathbf{p}_{W}^{\prime }(w).$$
For the derivations of the
efficiency bound, it is important to recall that $\ell$ depends on $p_{W}$,
so we sometimes use $\ell_{\mathbf{P}}$ to denote $\ell$. Finally, observe that under our assumptions over $\mu$ and $\mathbf{p}_{W}$, $\ell \in L^{2}(Leb)$.

The formal definition of the efficiency bound for the unknown parameter $\theta _{0}$ is given at the beginning of Appendix \ref{app:eff-bound}. Loosely speaking, the efficiency bound is a lower bound for the asymptotic variance of all locally regular and asymptotically linear estimators of $\theta_{0}$; see \cite{bickeletal1998efficient} for details and formal definitions. If it is infinite, then the parameter $\theta_{0}$ cannot be estimated at root-$n$ rate by these estimators. We now derive this bound. For this, we introduce some useful
notation. For any $(y,w,\alpha )\in \mathbb{Y}\times \mathbb{W}\times
\mathbb{A}$, let $$\rho (y,w,\alpha )\equiv \left(\rho_{1}(y,w,\alpha),\rho_{2}(y,w,h)\right)^{T} \equiv \left(\theta -\mu (w)h^{\prime}(w), 1\{y \leq h(w) \} - \tau \right)^{T}.$$
Let $\mathbf{T}:\mathbb{H}\rightarrow L^{2}(\mathbf{P}_{X})$ be given by
\begin{align*}
	\mathbf{T}[g](x)=\int	\mathbf{p}_{Y|WX}(h_{0}(w)|w,x)g(w)\mathbf{p}_{W|X}(w|x)dw
\end{align*}
for all $x\in \mathbb{X}$ and $
g\in \mathbb{H}$. The fact that $\mathbf{T}$ maps into $L^{2}(\mathbf{P}_{X})$ follows from
Jensen inequality and the fact that $\sup_{w,x} \mathbf{p}_{YW\mid X}(h_{0}(w),w\mid
x)<\infty $ (see Assumption \ref{ass:pdf0}). Its adjoint operator is denoted
as $\mathbf{T}^{\ast }:L^{2}(\mathbf{P}_{X})\rightarrow L^{2}(\mathbf{P}_{W})$. Finally, let
\[
x\mapsto \Gamma (x)\equiv E[\rho _{1}(Y,W,\alpha _{0})\rho _{2}(Y,W,h_{0})|X=x]/(\tau (1-\tau ))
\]
and $z\mapsto \epsilon (z)\equiv \rho
_{1}(y,w,\alpha _{0})-\Gamma (x)\rho _{2}(y,w,h _{0})$. Then $E[\epsilon (Z)\rho _{2}(Y,W,h_{0})|X]=0$ and $E[\epsilon (Z)]=0$.


\begin{theorem}
\label{thm:eff-bound} Suppose Assumptions \ref{ass:ident} and \ref{ass:pdf0}
hold and $\ell \in Kernel(\mathbf{T})^{\perp}$. Then

\begin{enumerate}
\item The efficiency bound of $\theta_{0}$ is finite iff $\ell \in Range
(\mathbf{T}^{\ast})$.

\item If it is finite, its efficient variance $V_0$ is given by
\begin{align*}
V_0= ||\epsilon(\cdot) ||^{2}_{L^{2}(\mathbf{P})} +
\left \Vert \mathbf{T}(\mathbf{T}^{\ast}\mathbf{T})^{+} [ \ell - \mathbf{T}^{\ast}[\Gamma] ] \right
\Vert^{2}_{L^{2}(\mathbf{P})}.
\end{align*}
\end{enumerate}
\end{theorem}

\begin{proof}
	 	See Appendix \ref{app:eff-bound}.
	 \end{proof}

The first result in Theorem \ref{thm:eff-bound} is obtained following the
approach of \cite{bickeletal1998efficient}.  The condition $\ell \in Kernel(\mathbf{T})^{\perp}$ ensures that only the ``identified part" of $h_{0}$ --- that is, the part of $h_{0}$ that is orthogonal to the kernel of $\mathbf{T}$ ---  matters for computing the weighted average derivative; we refer the reader to Appendix \ref{app:eff-bound} and the paper by  \cite{SeveriniTripathi2012} for further discussion.



\cite{SeveriniTripathi2012} provides an analogous result to Theorem \ref{thm:eff-bound}(1) for linear functionals
in a nonparametric \emph{linear} IV regression model. Our condition $\ell
\in Range(\mathbf{T}^{\ast })$, is analogous to theirs, but with a subtle yet
important difference. In \cite{SeveriniTripathi2012}, the object that plays the role of $\ell $ does not depend on $\mathbf{P}$, whereas in our case it does. This
observation changes the nature of our condition vis-a-vis theirs, because,
in our setup, $\ell \in Range(\mathbf{T}^{\ast })$ implies a restriction on $\mathbf{P}$
since both quantities, $\ell $ and $\mathbf{T}$ depend on it.\footnote{
It is worth pointing out that this restriction was not imposed as one of the
conditions that defined the model used to construct the tangent space; see
Appendix \ref{app:eff-bound} for a definition.} It is also important to note that, if $\mathbf{T}$ is compact, then the range of $
\mathbf{T}^{\ast }$ is a strict subset of $L^{2}(\mathbf{P}_{W})$ so that $\ell \in
Range(\mathbf{T}^{\ast })$ may not hold. Hence, in this case the weighted average
derivative may not be root-n estimable, and, moreover, the condition that
determines the finiteness of the efficiency bound depends on unknown
quantities. This observation highlights a difference with the no-endogeneity
case, where the efficiency bound is always finite, provided that $\ell \in
L^{2}(\mathbf{P}_{W})$ (see \cite{newey1993efficiency}).

Another discrepancy between the no-endogeneity case and ours is that in the
former case the ``plug in" is always efficient (see \cite
{newey1993efficiency}, \cite{newey1994asymptotic}) due to the fact that the
tangent space is the whole of $\{ f \in L^{2}(\mathbf{P}) \colon E[f] = 0 \}$. On
the other hand, for NPQIV \cite{CS2015overidentification} show that the
closure of the tangent space is the whole space iff the $Range(\mathbf{T})$ is dense
in $L^{2}(\mathbf{P}_{X})$, which in turn is equivalent to $Kernel(\mathbf{T}^{\ast}) = \{ 0\}$
. This last condition is comparable to a completeness condition on the
conditional distribution of the exogenous variable given the endogenous
ones, which may or may not hold for a particular $\mathbf{P}$.\footnote{
In the NPIV setting, $Kernel(\mathbf{T}^{\ast}) = \{ 0\}$ is equivalent to the pdf of
$X$ given $W$ satisfying a completeness condition.}

The second result in Theorem \ref{thm:eff-bound} follows from projecting the
influence function onto the closure of the tangent space (see \cite
{bickeletal1998efficient} and \cite{VdV2000} and references therein). So as
to shed some light on the expression for the efficiency bound, we point out
that it corresponds to the efficiency bound of the semiparametric
sequential conditional moment model via the ``orthogonalized moments" approach in \cite
{ai2012semiparametric}.
In their notation, let $\varepsilon_{2}(z,\alpha) \equiv \rho_{2}(y,w,h)$ and
$\varepsilon_{1}(z,\alpha) \equiv \rho_{1}(y,w,\alpha) - \Gamma(x)\rho_{2}(y,w,h)$.
Note that $E[\varepsilon_{1}(Z,\alpha_{0})\varepsilon_{2}(Z,\alpha_{0}) \mid X] = 0$ (and $\varepsilon_{1}(z,\alpha_0 )=\epsilon (z)$). The model (\ref{eqn:model1})-(\ref{eqn:model2}) becomes equivalent to their orthogonalized moment model:
\begin{equation}\label{eqn:model-orthog}
E[\varepsilon_{2}(Z,\alpha_0)|X] = 0~,~~~E[\varepsilon_{1}(Z,\alpha_0)]=0.
\end{equation}
The expression in our Theorem \ref{thm:eff-bound}(2) coincides with their theorem 2.3 semiparametric efficient variance bound for $\theta_0$ of the model (\ref{eqn:model-orthog}). Also see proposition 3.3 in \cite{ai2012semiparametric} for the semiparametric efficient variance bound for the WAD of a NPIV model.


\section{The Penalized-Sieve-GEL Estimator}

\label{sec:PSGEL}

In this section we introduce our estimator for $\alpha_{0}\in \mathcal{A} \equiv \Theta
\times \mathcal{H} \subseteq \mathbb{A}\equiv \Theta \times \mathbb{H}$. In order to do
this, it will be useful to define some quantities. Given the i.i.d. sample $(Z_{i})_{i=1}^{n}$, let $P_{n}$ be the corresponding empirical probability. Let $(q_{k})_{k\in \mathbb{N}}$ be a complete basis in $L^{2}(\mathbb{X},Leb)$. For any $J\in \mathbb{N}$, let $q^{J}(x)=(q_{1}(x),...,q_{J}(x))^{T}$ be $J\times 1$ vector-valued function of $x$, and for any $(z,\alpha )\in \mathbb{Z}\times \mathcal{A}$, let
\[
g_{J}(z,\alpha )\equiv  \left (\rho _{1}(y,w,\alpha ),\rho _{2}(y,w,\alpha)q^{J}(x)^{T} \right)^{T}
=\left (\theta -\mu (w)h^{\prime}(w), [ 1\{y \leq h(w) \} - \tau ]q^{J}(x)^{T} \right)^{T}~.
\]
Let $\mathcal{S}\subseteq \mathbb{R}$ be an open interval that contains $0$. For any $P \in \mathcal{P}(\mathbb{Z})$,  any $\alpha \in \mathcal{A}$ and any $J
\in \mathbb{N}$, denote $\Lambda_{J}(\alpha,P)\equiv \cap_{z \in supp (P)} \{
\lambda \in \mathbb{R}^{J+1} \colon \lambda^{T}g_{J}(z,\alpha) \in \mathcal{S}
\}$, and $\hat{\Lambda}_{J}(\alpha)\equiv \Lambda_{J}(\alpha,P_{n})$.

Let $s:\mathcal{S}\rightarrow \mathbb{R}$ be strictly
concave, twice-continuously differentiable with Lipschitz continuous second
derivative; and $s^{\prime }(0)=s^{\prime \prime }(0)=-1$; see, e.g., \cite{Smith1997} and \cite{DIN03} for examples of such $s(.)$ functions. For any $\lambda \in \Lambda_{J}(\alpha,P)$, let
\begin{align*}
S_{J}(\alpha,\lambda,P) \equiv E_{P}[s(\lambda^{T}g_{J}(Z,\alpha))] - s(0)~,~~~\hat{S}_{J}(\alpha,\lambda)\equiv S_{J}(\alpha,\lambda,P_{n}).
\end{align*}
If $\mathcal{A}$ were a finite-dimensional compact set with $\dim (\mathcal{A}) \leq J+1$, then $\alpha_0$ could be estimated by the GEL procedure: $\arg \min_{\alpha \in \mathcal{A}} \sup_{\lambda \in
\hat{\Lambda}_{J}(\alpha )}\hat{S}_{J}(\alpha ,\lambda )$ (see, e.g., \cite{DIN03}).

Due to the presence of the infinite-dimensional nuisance parameter $h_0 \in \mathcal{H}$ in the NPQIV model (\ref{eqn:model1}), the parameter space $\mathcal{A} \equiv \Theta
\times \mathcal{H}$ is  an infinite-dimensional function space that is typically non-compact subset in $(\mathbb{A}, ||.||)$ and hence the identifiable uniqueness condition needed for consistency in $||.||$-norm might fail; see, e.g., \cite{NP2003instrumental} and \cite{chen2007large}. The above GEL procedure needs to be regularized to regain consistency and/or to speed up rate of convergence in $||.||$-norm.
To this end, we introduce a \emph{regularizing structure}, which, jointly with $(q_{k})_{k \in \mathbb{N}}$, consists of a sequence of sieve spaces $(\mathcal{A}_k \equiv \Theta
\times \mathcal{H}_{k})_{k\in \mathbb{N}}$ in $(\mathbb{A}, ||.||)$, and a sequence of penalties $(\gamma_{k}\times Pen (\cdot))_{k \in \mathbb{N}}$ with tuning parameters $\gamma_{k} \downarrow 0$ and a penalty function $Pen : \mathbb{A} \rightarrow \mathbb{R}_{+}$.

The \emph{Penalized-Sieve-GEL (PSGEL)} estimator
is defined as
\begin{equation*}
\hat{\alpha}_{L,n} \in \arg \min_{\alpha \in \mathcal{A}_{K}} \left[\sup_{\lambda \in
\hat{\Lambda}_{J}(\alpha )}\hat{S}_{J}(\alpha ,\lambda )+\gamma
_{K}Pen(\alpha )\right],
\end{equation*}
for any $(L=(J,K),n)\in \mathbb{N}^{3}$. If the \textquotedblleft arg min" in the previous expression is empty, one can replace it by an approximate minimizer.

The following assumption imposes restrictions over the regularizing structure $\{(q_{k}, \mathcal{H}_{k},\gamma_{k}Pen)_{k \in \mathbb{N}}\}$. Let $(\varphi _{k})_{k\in \mathbb{N}}$ be a basis functions in $\mathbb{H}$, and $\nabla \varphi
^{K}=(\varphi _{1}^{\prime },...,\varphi _{K}^{\prime })^{T}$.

\begin{assumption}
\label{ass:reg} (i) $(q_{k})_{k\in \mathbb{N}}$ is a basis in $L^{2}(\mathbf{P}_{X})$, and $E[q^{J}(X)q^{J}(X)^{T}]=I$ for each finite $J$;
\newline (ii) For all $K $, $\mathcal{H}_{K}\subseteq lin\{\varphi _{1},...,\varphi _{K}\}$ is
closed and convex, and $\overline{\cup _{k}\mathcal{H}_{k}}\supseteq
\mathcal{H}$, i.e., for any $\alpha \in \mathcal{A} \equiv \Theta \times \mathcal{H}$ there is an $\Pi_K\alpha \in \mathcal{A}_K \equiv \Theta \times \mathcal{H}_K$ such that $||\Pi_K\alpha - \alpha ||=o(1)$; and for some finite $C\geq 1$,
$C^{-1}I\leq E\left[ \left( \varphi ^{K}(W)\right) \left(
\varphi ^{K}(W)\right) ^{T}+\left( \nabla \varphi ^{K}(W)\right) \left(
\nabla \varphi ^{K}(W)\right) ^{T}\right] \leq CI$;
\newline (iii) (a) $Pen : \mathbb{A} \rightarrow \mathbb{R}_{+}$
is lower semi-compact (in $||.||$), $|Pen (\Pi_K\alpha_0 ) - Pen (\alpha_0)|=O(1)$, $Pen (\alpha_0 )<\infty$, and $\gamma_{k} \downarrow 0$, and (b) there exists an $M<\infty
$ such that for any $m\geq M$, any $K$ and any $\alpha \in \mathcal{A}_{K}$,
if $Pen(\alpha )\leq m$ then $\sup_{w \in \mathbb{W}} |\mu(w) h^{\prime }(w)| \leq m$.
\end{assumption}

Condition (i) is mild (see \cite{DIN03} (DIN) and the discussion therein).
Condition (ii) essentially defines the sieve space. Part (a) of Condition (iii) is standard in ill-posed problems (see \cite{CP2012estimation}); Part (b) is not. If $\mathcal{H}_{K}$ is $||\cdot ||_{L^{\infty}(\mathbb{W},\mu)} $ bounded, then the condition is vacuous.
If this is not the case, then the condition requires $Pen$ to be ``stronger"
than the $||\cdot ||_{L^{\infty}(\mathbb{W},\mu)} $ norm. The need to bound $
|| h^{\prime } ||_{L^{\infty}(\mathbb{W},\mu)} $ arises from the fact that, in many
instances, in the proofs we need to control $\rho(y,w,\alpha)$ uniformly on $
(y,w)$ (e.g., see Lemma \ref{lem:Lambda-charac} in the Supplemental Material \ref{supp:conv-rate}). Additionally, in our setup, is useful to link $Pen$ to $||\cdot ||_{L^{\infty}(\mathbb{W},\mu)}$ because the structure of the problem implies a natural bound for $Pen(.)$ --- and thus, through Assumption \ref{ass:reg}(iii), a bound for $||\cdot ||_{L^{\infty}(\mathbb{W},\mu)}$ ---, as shown in the following lemma.

\begin{lemma}
\label{lem:Pen-bound} For any $L=(J,K) \in \mathbb{N}^{2}$ and any $\alpha
\in \mathcal{A}_{K}$,
\begin{align*}
\gamma_{K} Pen(\hat{\alpha}_{L,n}) \leq \sup_{\lambda \in \hat{\Lambda}
_{J}(\alpha)} \hat{S}_{J}(\alpha,\lambda) + \gamma_{K} Pen(\alpha)~~~wpa1.
\end{align*}
\end{lemma}

\begin{proof}
	See Appendix \ref{app:PSGEL}.
\end{proof}

The bound, however, may depend on $(J,K,n)$ and thus may affect the convergence
rate. Below, we will set $\alpha$ in the right-hand-side (RHS) to a particular value in $\mathcal{A}_{K}$ and use the resulting bound to construct what we call an ``effective sieve space".



\section{Consistency and Convergence Rates of the PSGEL Estimator}
\label{sec:conv-rate}

This section establishes the consistency and the rates of convergence of the PSGEL estimator $\hat{\alpha}_{L,n}$ to the true parameter $\alpha_0$ under a given norm $||.||$ over $\mathbb{A}$. In this and the next section, we note that the implicit constants inside the $O_{\mathbf{P}}$ do not depend on $(J,K,n)$.

\subsection{Effective sieve space}

Throughout the paper we use the following notation. Let $\overline{\theta} \equiv \sup_{ \theta \in \Theta} |\theta| <\infty$; and  $b_{\rho,J} \equiv (E[||q^{J}(X)||^{\rho}_{e}])^{1/\rho}$ for any $\rho > 0$. For any $L=(J,K) \in \mathbb{N}^{2}$, let
\begin{align*}
\Gamma_{L,n} \equiv \left\{ \frac{\bar{g}^{2}_{L,0}}{n} +
||E[g_{J}(Z,\Pi_K \alpha_{0})]||^{2}_{e} + \gamma_{K} Pen(\Pi_K \alpha_{0})
\right\}, ~~\bar{g}^{2}_{L,0} \equiv \overline{\theta} + ||\mu (\Pi_K h_{0})^{\prime
}||^{2}_{L^{2}(\mathbf{P})} + b_{2,J}^{2}.
\end{align*}
Let $(l_{n})_{n}$ be a slowly diverging positive sequence, e.g., $l_{n} =\log \log n$, which is introduced solely to avoid keeping track of constants. Finally we let
\begin{align*}
\bar{\mathcal{A}}_{L,n} \equiv \left\{ \alpha \in \mathcal{A}_{K} \colon
Pen(\alpha) \leq \mho_{L,n} \right\}~,~~~\text{where}~~\mho_{L,n} \equiv l_{n} \gamma^{-1}_{K} \Gamma_{L,n}~.
\end{align*}

The sequence of sets, $(\bar{\mathcal{A}}_{L,n})_{L,n}$, can be viewed as
the sequence of ``effective" sieve spaces, because, as the following lemma
shows, wpa1 the estimator (and, trivially, the sieve approximator $\Pi_K \alpha_{0}\in \mathcal{A}_{K}$) both
belong to it.

\begin{assumption}
\label{ass:rates-mild} (i) $b^{4}_{4,J}/n = o(1)$; (ii) $\delta_{n} = o(1)$, $\delta_{n} \times \mho_{L,n} = o(1)$, $ b_{\varrho,J}^{\varrho} n \delta^{\varrho}_{n}= o(1)$ for some $\varrho > 0$; (iii) $\sqrt{ \frac{\bar{g}^{2}_{L,0}}{n} + ||E[g_{J}(Z,\Pi_K \alpha_{0})]||^{2}_{e} }
= o(\delta_{n})$.
\end{assumption}

\begin{lemma}
	\label{lem:eff-sieve} Let Assumptions \ref{ass:ident}, \ref{ass:pdf0}, \ref{ass:reg} and \ref{ass:rates-mild} hold. Then, for any $L \in \mathbb{N}^{2}$, $\hat{\alpha}
	_{L,n} \in \bar{\mathcal{A}}_{L,n}$ wpa1.
\end{lemma}

\begin{proof}
	See Appendix \ref{app:conv-rate}.
\end{proof}

The proof of this Lemma follows from Lemma \ref{lem:Pen-bound} with $\alpha
= \Pi_K \alpha_{0}$ and Lemma \ref{lem:SJ-bound} with  $\alpha = \Pi_K \alpha_{0}$ and $P=P_{n}$ in
Appendix \ref{app:conv-rate}. The latter lemma provides a bound for $\sup_{\lambda \in
	\hat{\Lambda}_{J}(\Pi_K \alpha_{0})} \hat{S}_{J}(\Pi_K \alpha_{0},\lambda)$ in terms
of $||n^{-1}\sum_{i=1}^{n} g_{J}(Z_{i},\Pi_K \alpha_{0}) ||^{2}_{e}$ and $
\gamma_{K} Pen(\Pi_K \alpha_{0})$. With this in mind, the components of $
\mho_{L,n}$ are intuitive: $\bar{g}^{2}_{L,0}/n$ is related to the
``variance" of $n^{-1}\sum_{i=1}^{n} g_{J}(Z_{i},\Pi_K \alpha_{0})$, where $\bar{
	g}^{2}_{L,0}$ is a bound for $||g_{J}(.,\Pi_K \alpha_{0})||^{2}_{e}$. The term $
||E[g_{J}(Z,\Pi_K \alpha_{0})]||_{e}$ is related to the ``bias" and reflects the
fact that $\Pi_K \alpha_{0}\in \mathcal{A}_K$ is a sieve approximate to $\alpha_0$.

\begin{remark}\label{rem:bound}
	As explained above, Lemma \ref{lem:eff-sieve} and Assumption \ref
	{ass:reg}(iii) are used to ensure that $|| h^{\prime }||_{L^{\infty}(
		\mathbb{W},\mu)}$ is bounded. If the construction of $\mathcal{H}_{K}$ directly
	implies $|| h^{\prime }||_{L^{\infty}(\mathbb{W},\mu)} \leq \mho$ for some
	fixed constant $\mho < \infty$, then $\mho$ should replace $\mho_{L,n}$ in
	the definition of $\bar{\mathcal{A}}_{L,n}$. This is applicable every time $\mho_{L,n}$ appears below. $\triangle$
\end{remark}


\subsection{Relation to Penalized Sieve GMM}


As expected, the asymptotic properties of the PSGEL estimator are closely related to an approximate minimizer of a GMM criterion associated to the following expression: for any $J \in \mathbb{N}$ and any $P \in \mathcal{P}(\mathbb{Z})$, let
\begin{align*}
\alpha \mapsto Q_{J}(\alpha,P) \equiv
E_{P}[g_{J}(Z,\alpha)]^{T}H_{J}(\alpha_{0},\mathbf{P})^{-1}E_{P}[g_{J}(Z,\alpha)]
\end{align*}
where $(\alpha,P) \mapsto H_{J}(\alpha,P) \equiv E_{P}[g_{J}(Z,\alpha)
g_{J}(Z,\alpha)^{T}]$. That is, $Q_{J}(.,\mathbf{P})$ is the optimally weighted (population) GMM criterion
function associated with the vector of moments $E[g_{J}(Z,\cdot)]$.

For what follows, it will be useful to define the following intermediate
quantity which can be viewed as a (sequence) of pseudo-true parameters. For
each $L\equiv (J,K)\in \mathbb{N}^{2}$, let
\begin{equation*}
\alpha _{L,0}\equiv \arg \min_{\alpha \in \bar{\mathcal{A}}_{L,n}}Q_{J}(\alpha
,\mathbf{P}).
\end{equation*}
We note that $\alpha _{0} \in \arg \min_{\alpha \in
\mathcal{A}}Q_{J}(\alpha ,\mathbf{P})$ for any $J$; but as we
restrict to the effective sieve space $\bar{\mathcal{A}}_{L,n}$, it could be
that $\alpha _{L,0}\neq \alpha _{0}$ for any $L\in \mathbb{N}^{2}$. The following lemma guarantees that $\alpha _{L,0}$ is in fact non-empty.

\begin{lemma}
\label{lem:alpha0L-exists} Let Assumptions \ref{ass:pdf0} and \ref{ass:reg} hold. Then, for each $L=(J,K) \in \mathbb{N}^{2}$, $
\alpha_{L,0}$ is non-empty.
\end{lemma}

\begin{proof}
		See Appendix \ref{app:conv-rate}.
	\end{proof}

While this lemma shows that $\alpha _{L,0}$ is non-empty, it may not be a singleton. Nevertheless, for model (\ref{eqn:model1})-(\ref{eqn:model2}), it is easy to choose some finite-dimensional linear sieve $\mathcal{H}_{K}$ and some strict convex penalty $Pen$ such that $\alpha _{L,0}$ is in fact a singleton. Therefore the next assumption is effectively a way to suggest choices of a regularizing structure:

\begin{assumption}\label{ass:ID-pseudotrue}
	For any $L \in \mathbb{N}^{2}$, $\alpha _{L,0}$ is single-valued.
\end{assumption}

Let $m_{2}(X,\alpha) \equiv E_{\mathbf{P}}[\rho_{2}(Y,W,\alpha) \mid X]$. For each $J \in \mathbb{N}$, the $L^{2}(\mathbf{P})$ projection of $m_{2}(\cdot, \alpha)$ onto the linear span of $q^{J}(X)$ is denoted as $Proj_{J}[m_{2}(\cdot,\alpha)](X)$, where
\begin{align*}
Proj_{J}[m_{2}(\cdot,\alpha)](X) = & E_{\mathbf{P}} \left[ m_{2}(X,\alpha) q^{J}(X)^{T}  \right] (E_{\mathbf{P}}[q^{J}(X)q^{J}(X)^{T}])^{-1} q^{J}(X) \\
= &  E_{\mathbf{P}} \left[ \rho_{2}(Y,W,\alpha) q^{J}(X)^{T} \right] q^{J}(X)
\end{align*}
where $(E_{\mathbf{P}}[q^{J}(X)q^{J}(X)^{T}])^{-1} = I$ by Assumption \ref{ass:reg}.

The next lemma provides sufficient conditions that ensure convergence of $\alpha_{L,0}$ to the true parameter $\alpha_{0}$.

\begin{lemma}
\label{lem:alpha0L-consistent} Let Assumptions \ref{ass:ident}, \ref{ass:pdf0}, \ref{ass:reg} and \ref{ass:ID-pseudotrue} hold. Suppose $\lim_{n \rightarrow \infty} \sup_{ \alpha \in \bar{\mathcal{A}}_{L_{n},n}}||Proj_{J_{n}}[m_{2}(\cdot,\alpha)] - m_{2}(\cdot,\alpha)||_{L^{2}(\mathbf{P})} = 0  $. Then: $||\alpha_{L_n,0}-\alpha_0||=o(1)$.
\end{lemma}

\begin{proof}
		See Appendix \ref{app:conv-rate}.
	\end{proof}

\subsection{Convergence rates}

A crucial part of establishing the convergence rate of $\hat{\alpha}_{L,n}$ is to bound the rate of $||\hat{\alpha}_{L,n} - \alpha_{L,0}||$. For this it is important to quantify how well the population sieve GMM criterion function $Q_{J}$ separates points in $(\bar{\mathcal{A}}_{L,n},||.||)$ around $\alpha_{L,0}$. To do this, we define, for each, $(L,n)\in \mathbb{N}^{3}$, $\varpi_{L,n} : \mathbb{R}
_{+} \rightarrow \mathbb{R}_{+}$ as
\begin{align}  \label{eqn:IU}
t \mapsto \varpi_{L,n}(t) \equiv \inf_{\{ \alpha \in \bar{\mathcal{A}}_{L,n}
	\colon ||\alpha - \alpha_{L,0} || \geq t \}} Q_{J}(\alpha,\mathbf{P}) -
Q_{J}(\alpha_{L,0},\mathbf{P}).
\end{align}

The function $
\varpi_{L,n}$ is analogous to the one used in the standard identifiable
uniqueness condition (see \cite{WW1991}, \cite{newey1994large}). Within the
ill-posed inverse literature this function is akin to the notion of sieve
measure of ill-posedness used in \cite{BCK2007} and \cite{CP2012estimation,CP2015sieve}. The following lemma establishes some useful properties.

\begin{lemma}
	\label{lem:sieve-IU} Let Assumptions \ref{ass:pdf0}, \ref{ass:reg} and \ref{ass:ID-pseudotrue} hold. Then: for each $(L,n) \in \mathbb{N}^{3}$, $\varpi_{L,n}(t) = 0$
	iff $t = 0$ and $\varpi_{L,n}$ is continuous and non-decreasing in $t$.
\end{lemma}

\begin{proof}
	See Appendix \ref{app:conv-rate}.
\end{proof}

It is worth noting that even though $\varpi_{L,n}(t) > 0$ for all $t>0$,
it could happen that $\varpi_{L,n}(t) \rightarrow 0$ as $L$ diverges. This
behavior reflects the ill-posed nature of the problem.

We now present some high-level assumptions used to establish the convergence rate of the PSGEL estimator. The first of these assumptions introduces, and imposes restrictions on, a positive real-valued sequence $(\delta _{n})_{n\in \mathbb{N}}$ that is common in the GEL literature (see the Appendix in \cite{DIN03}). It ensures that the ball $\{ \lambda \in \mathbb{R}^{J+1} \colon ||\lambda||_{e}
\leq \delta_{n} \}$ belongs to $\hat{\Lambda}_{J}(\alpha)$ for any $\alpha
\in \bar{\mathcal{A}}_{L,n}$ (see Lemma \ref{lem:Lambda-charac} in the Supplemental Material \ref{supp:conv-rate}).  The assumption also restricts the rates of $(b_{\rho,J})_{\rho \in \mathbb{R},J\in \mathbb{N}},$ and the rate at which $L=(J,K) \in \mathbb{N}^{2}$ diverges relative to $n$:

\begin{assumption}
\label{ass:rates} (i) Assumption \ref{ass:rates-mild} holds; (ii)  $\delta_{n} l_{n} = o(1)$, $b^{3}_{3,J} \delta^{\varrho}_{n}= o(1)$ for some $\varrho > 0$, and $(\mho_{L,n})^{4} b^{4}_{4,J}/n = o(1)$.
\end{assumption}


Recall that the sequence $(l_{n})_{n}$ diverges
arbitrary slowly like $\log \log n$, and the bound $\mho_{L,n}$ is allowed to grow
(slowly) at the rate of $l_{n}$. Assumption \ref{ass:rates} slightly strengthens Assumption \ref{ass:rates-mild}.


The following assumption is a high-level condition that
controls the supremum of the process $f \mapsto \mathbb{G}_{n}[f] \equiv n^{-1/2} \sum_{i=1}^{n} \{f(Z_{i}) -E[f(Z_{i})]\}$ over  the classes $\bar{\mathcal{A}}_{L,n}$ and $\mathcal{G}_{L} \equiv \{ (y,w) \mapsto \rho_{2}(y,w,\alpha) \colon \alpha \in \bar{\mathcal{A}}_{L,n} \} $.

\begin{assumption}
\label{ass:Donsker} There exists a positive real-valued sequence, $
(\Delta_{L,n})_{L,n \in \mathbb{N}^{3}}$, such that, for any $L \in \mathbb{N
}^{2}$, $\sup_{(\theta,h) \in \bar{\mathcal{A}}_{L,n}} |\mathbb{G}_{n}[\mu \cdot
h^{\prime }]| =O_{\mathbf{P}}(\Delta_{L,n})$ and for all $1 \leq j \leq J$, $
\sup_{g \in \mathcal{G}_{L}} |\mathbb{G}_{n}[g \cdot q_{j}]|
=O_{\mathbf{P}}(\Delta_{L,n})$.
\end{assumption}

For instance, if $\{ \mu \cdot h^{\prime }\colon h \in
\mathcal{H} \}$ and $\{ (y,w) \mapsto 1\{ y \leq h(w) \} \colon h \in
\mathcal{H} \}$ are P-Donsker, then $(\Delta_{L,n})_{L,n \in \mathbb{N}^{3}}$
is uniformly bounded.\footnote{Restrictions on the ``complexity" of these classes are implicit restrictions on the ``complexity" of $\mathcal{H}$; see \cite{chen2003estimation} and \cite{VdV2000}.}  But if this is not the case, then $(\Delta_{L,n})_{L,n
\in \mathbb{N}^{3}}$ may diverge as $L$ (or $n$) grows.

The next theorem establishes the convergence rate of the PSGEL estimator; in particular it establishes the rate for the estimator of the infinite dimensional component $h_0 \in \mathcal{H}$.

\begin{theorem}
	\label{thm:conv-rate} Suppose Assumptions \ref{ass:ident}, \ref{ass:pdf0},
	\ref{ass:reg}, \ref{ass:ID-pseudotrue} and \ref{ass:Donsker} hold. For any $(\delta_{n},l_{n})_{n}$
	satisfying Assumption \ref{ass:rates}, there exists a finite constant $M>0$ such that
	\begin{align*}
	||\hat{\alpha}_{L,n} - \alpha_{0}|| = O_{\mathbf{P}}\left( \varpi_{L,n}^{-1}\left( M (\delta_{1,L,n} + \delta_{2,L,n}) \right) \right) + ||\alpha_{L,0} -
	\alpha_{0} ||,
	\end{align*}
	where \begin{align*}
	\delta _{1,L,n}\equiv \sqrt{\frac{J}{n}}\times \Delta _{L,n}\times
	\left( \overline{\theta}+\mho _{L,n}+b_{2,J}\right)~,~~
	\delta _{2,L,n}\equiv  \mho _{L,n}^{2}\left\{ \delta _{n}+\delta
	_{n}^{-1}\Gamma _{L,n}\right\}
	\end{align*}
\end{theorem}

\begin{proof}
	See Appendix \ref{app:Thm-conv-rate}.
\end{proof}

The rate of convergence of the PSGEL estimator is composed of two standard
terms reflecting the ``approximation error" $||\alpha_{L,0} - \alpha_{0} ||$
and the ``sampling error" $\varpi_{L,n}^{-1}\left(M(\delta_{1,L,n} +
\delta_{2,L,n})\right)$. The component $\varpi^{-1}_{L,n}(.)$, reflects the ill-posed nature of the estimation
problem. As noted previously, even though, for a fixed $L$, $\varpi_{L,n}(t) >
0$ for $t>0$, this relationship can deteriorate as $L$ diverges, which
implies that $\varpi_{L,n}^{-1}(t)$ may diverge as $L$ diverges.

Below, we present an heuristic description of the proof that sheds light on the role of the sequences $(\delta_{1,L,n},\delta_{2,L,n})_{L,n}$ and of $\varpi_{L,n}$.

\subsection{Heuristics}

By the triangle inequality it suffices to bound the rate of $||\hat{\alpha}_{L,n} - \alpha_{L,0}||$. We do this by linking the PSGEL estimator to the population sieve GMM problem defined by $Q_{J}(\cdot,\mathbf{P})$. The first step to do this is to show that the PSGEL is an \emph{approximate} minimizer of the sample sieve GMM criterion $Q_{J}(.,P_{n})$ with the rate given by $\delta _{2,L,n}$.
\begin{lemma}
	\label{lem:QJ-approx-min} Let Assumptions \ref{ass:pdf0}
	and \ref{ass:reg} hold. For any $(\delta _{n},l_{n})_{n\in \mathbb{N}}$ satisfying Assumption \ref{ass:rates}, we have:
	\begin{equation*}
Q_{J}(\hat{\alpha}_{L,n},P_{n})=O_{\mathbf{P}}(\delta _{2,L,n})~,~~~\text{with}~~\delta _{2,L,n} = \mho _{L,n}^{2}\left\{ \delta _{n}+\delta_{n}^{-1}\Gamma _{L,n}\right\}~.
	\end{equation*}
\end{lemma}

\begin{proof}
	See Appendix \ref{app:conv-rate}.
\end{proof}

The Lemma illustrates not only the role of $\delta_{2,L,n}$ but its nature. The two terms inside the curly brackets are completely analogous to those appearing in \cite{DIN03}. The
scaling by $\mho^{2}_{L,n}$ is not present in \cite{DIN03} and its
appearance here is due to the fact that the bound of $\rho(.,\alpha_{L,0})$
may depend, in principle, on $n$ and $L=(J,K)$. In \cite{DIN03}, on the
other hand, the upper bound $\mho_{L,n}$ can be taken to be a fixed constant
due to their Assumption 6.

Lemma \ref{lem:QJ-approx-min} implies that, for some finite $M$, the event $Q_{J}(\hat{\alpha}_{L,n},P_{n}) - Q_{J}(\alpha_{L,0},P_{n}) \leq M \delta_{2,L,n}$ occurs wpa1. The next step is to link the \emph{empirical} GMM criterion function, $Q_{J}(\cdot,P_{n})$, to its \emph{population} analog, $Q_{J}(\cdot,\mathbf{P})$ for which we can quantify its behavior (around $\alpha_{L,0}$) using $\varpi_{L,n}$. The next lemma provides such a link by showing that $Q_{J}(\cdot,P_{n})$ converges to its population analog.
\begin{lemma}
	\label{lem:QJ-univ-conv} Let Assumptions \ref{ass:pdf0}, \ref{ass:reg} and \ref{ass:Donsker} hold. Then: for any $L\equiv (J,K)\in
	\mathbb{N}^{2}$,
	\begin{equation*}
	\sup_{\alpha \in \bar{\mathcal{A}}_{L,n}}|Q_{J}(\alpha ,P_{n})-Q_{J}(\alpha,\mathbf{P})|=O_{\mathbf{P}}(\delta _{1,L,n})~,~~~\text{with}~~\delta _{1,L,n}=\sqrt{\frac{J}{n}}\times \Delta _{L,n}\times
	\left( \overline{\theta} +\mho _{L,n}+b_{2,J}\right).
	\end{equation*}
	\end{lemma}

\begin{proof}
	See Appendix \ref{app:conv-rate}.
\end{proof}

The rate $(\delta_{1,L,n})_{L,n}$ has several components. The component $
\sqrt{\frac{J}{n}}$ reflects the pointwise convergence rate of $||n^{-1}
\sum_{i=1}^{n} g_{J}(Z_{i},\alpha) - E[g_{J}(Z,\alpha)] ||_{e} $, while the
factor of $\Delta_{L,n}$ reflects the fact that we need \emph{uniform}
convergence of that term. Finally, the term $\left( \overline{\theta} +
\mho_{L,n} + b_{2,J} \right)$ is essentially the (uniform) bound for $\alpha
\mapsto ||n^{-1} \sum_{i=1}^{n} g_{J}(Z_{i},\alpha)||_{e}$ and $\alpha
\mapsto ||E[g_{J}(Z,\alpha)] ||_{e} $ over $\bar{\mathcal{A}}_{L,n}$.

With this result at hand and simple algebra, one can show that for some finite $M$ the set $A \equiv \{Q_{J}(\hat{\alpha}_{L,n},\mathbf{P}) - Q_{J}(\alpha_{L,0},\mathbf{P}) \leq M (\delta_{1,L,n} + \delta_{2,L,n})\}$ occurs wpa1. Therefore, by standard laws of probabilities, it follows that the probability of the set $||\hat{\alpha}_{L,n} - \alpha_{0}|| \geq M^{\prime} \varpi_{L,n}^{-1}\left( M (\delta_{1,L,n} + \delta_{2,L,n}) \right)$ (for any $M^{\prime}$) is --- up to a vanishing term --- less or equal than the probability of the intersection of the same set with $A$. Therefore, it only remains to show that the latter probability is naught for sufficiently large $M^{\prime}$. This follows because this latter probability is in turn bounded above by the probability of $\varpi_{L,n} \left( M^{\prime} \varpi_{L,n}^{-1}\left( M (\delta_{1,L,n} + \delta_{2,L,n}) \right) \right) \leq M (\delta_{1,L,n} + \delta_{2,L,n})$. By the fact that $\varpi_{L,n}$ is non-decreasing (see Lemma \ref{lem:sieve-IU}), this probability is naught by sufficiently large $M^{\prime}$, proving the result of Theorem \ref{thm:conv-rate}.

\subsection{Discussion of the elements in the Convergence Rate}
\label{sec:DiscussionRate}

We now present some observations regarding the main components of the convergence rate in Theorem \ref{thm:conv-rate}, namely, the rates $(\delta_{1,L,n},\delta_{2,L,n})_{L,n}$ and $\varpi_{L,n}$ defined in expression \ref{eqn:IU}. Regarding the latter, we first need to specify the norm $||.||$. We start by taking $(\varphi_{k})_{k \in \mathbb{N}}$ to be an orthogonal basis with respect to the Lebesgue measure over $\mathbb{H}$. Thus, for any $\alpha = (\theta,h) \in \mathbb{A}$, there exists a real-valued sequence, $(\pi_{l})_{l=0}^{\infty}$, such that   $\alpha = (\theta,h) = (\pi_{0} \bar{\varphi}_{0}, \sum_{l=1}^{\infty} \pi_{l} \bar{\varphi}_{l})$ where $\bar{\varphi}_{0} = 1$ and $\pi_{0} = \theta$, and, for any $k \geq 1$, $\bar{\varphi}_{k} = \varphi_{k}$ and $\pi_{k}$ is the ``Fourier" coefficient of $h$ with respect the basis $(\varphi_{k})_{k \in \mathbb{N}}$. This representation gives rise to the following norm over $\mathbb{A}$, $\alpha \mapsto \sqrt{ \sum_{l=0}^{\infty} \pi^{2}_{l}  } = \sqrt{\theta^{2} + ||h||^{2}_{L^{2}(Leb)}}$.  The aforementioned norm presents itself as a ``natural" norm under which we can establish convergence rate and thus we set $||.||$ as this norm; our result can be extended to norms other than this by specifying how the desired norm relates to $\alpha \mapsto \sqrt{|\theta|^{2} + ||h||^{2}_{L^{2}(Leb)}}$.


We now shed light on the behavior of $\varpi_{L,n}$ under our choice of $||.||$. In particular, we will illustrate how this function is linked to the curvature of the criterion function $\alpha \mapsto \bar{Q}_{J}(\alpha,\mathbf{P})$. To do this, it is convenient to use local approximations, so we take, for each $L=(K,L) \in \mathbb{N}^{2}$, $\mathcal{A}_{K}$ to be convex, $\alpha_{L,0}$ to be such that $ARC_{K}(\alpha_{L,0}) \equiv \{ \alpha_{L,0} + t \zeta \colon \zeta \in \mathcal{A}_{K}\setminus\{ \alpha_{L,0} \} ~and~t\in [0,1]  \} \subseteq \mathcal{A}_{K}$,  and require that $Pen$ to be convex and twice continuously differentiable. By the mean value theorem and the fact that $\alpha_{L,0}$ is a minimizer --- and thus satisfies that $\frac{d\bar{Q}_{J}(\alpha_{L,0},\mathbf{P})}{d\alpha}[\cdot]=0$ ---, it follows that for any $\alpha \in \bar{\mathcal{A}}_{L,n}$,
\begin{align*}
\bar{Q}_{J}(\alpha,\mathbf{P}) - \bar{Q}_{J}(\alpha_{L,0},\mathbf{P}) \geq \frac{1}{2}  \inf_{\eta \in ARC_{K}(\alpha_{L,0})} \frac{d^{2} \bar{Q}_{J}(\eta ,\mathbf{P})}{d\alpha^{2}}[\alpha - \alpha_{L,0},\alpha - \alpha_{L,0}].
\end{align*}
By the sieve representation discussed above, the RHS in this expression can be cast as
\begin{align*}
\bar{Q}_{J}(\alpha,\mathbf{P}) - \bar{Q}_{J}(\alpha_{L,0},\mathbf{P}) \geq &  (\pi^{K+1} - \pi^{K+1}_{L,0}) \mathcal{I}_{L} (\pi^{K+1} - \pi^{K+1}_{L,0}) ^{T},
\end{align*}
where $\pi^{K+1}$ denotes the first $K+1$ coefficients of the representation of $\alpha$; $\pi^{K+1}_{L,0}$ is the same but for $\alpha_{L,0}$, and $\mathcal{I}_{L}$ is a $(K+1) \times (K+1)$ matrix where the $(i,j)$-th component is given by
\begin{align*}
	 \mathcal{I}_{L}[i,j] \equiv \frac{1}{2} \inf_{\eta \in ARC_{K}(\alpha_{L,0})} \frac{d^{2} \bar{Q}_{J}(\eta ,\mathbf{P})}{d\alpha^{2}}[\bar{\varphi}_{i},\bar{\varphi}_{j}].
\end{align*}
This result implies that $\varpi_{L,n}(t) \geq t^{2} e_{min}(\mathcal{I}_{L} )$ ($e_{min}(A)$ is the minimal eigenvalue of the matrix $A$). If $e_{min}(\mathcal{I}_{L})>0$, then
\begin{align*}
	||\hat{\alpha}_{L,n} - \alpha_{0} || = O_{\mathbf{P}} \left( (e_{min}(\mathcal{I}_{L} ))^{-1/2} \sqrt{\delta_{1,L,n} + \delta_{2,L,n}}  +  ||\alpha_{L,0} - \alpha_{0}|| \right).
\end{align*}
The scaling factor $(e_{min}(\mathcal{I}_{L} ))^{-1/2}$ summarizes the ill-posed nature of the problem, because, even though we require $e_{min}(\mathcal{I}_{L} )>0$ for \emph{each} $L$, we do not impose this restriction \emph{uniformly} on $L$, i.e., we allow that $e_{min}(\mathcal{I}_{L} ) \rightarrow 0$ as $L \rightarrow \infty$.\footnote{The condition $e_{min}(\mathcal{I}_{L} )>0$ for each $L$ essentially ensures that the ``identifiable uniqueness" condition \ref{eqn:IU} holds for each $L$; this requirement is common in the ill-posed inverse literature (e.g., \cite{chen2007large}).} The speed at which this occurs depends on the local curvature of $\bar{Q}_{J}(\cdot, \mathbf{P})$ (at $\alpha_{L,0}$) and the growth of $\mathcal{A}_{K}$; see \cite{BCK2007} and \cite{CP2012estimation} for a more thorough discussion.

We next discuss the rate components $(\delta_{1,L,n},\delta_{2,L,n})_{L,n}$. As mentioned in Remark \ref{rem:bound} above, if $\sup_{h \in \mathcal{H}} ||h^{\prime}||_{L^{\infty}(\mathbb{W},\mu)}$ is finite, then $\mho_{L,n}$ can be replaced by a fixed constant $\mho$; this fact and some algebra implies that $\delta_{2,L,n} = O(\delta_{n} + \delta^{-1}_{n} (b^{2}_{2,J}/n + ||E[g_{J}(Z,\Pi_{K}\alpha_{0})]||^{2}_{e} + \gamma_{K} Pen(\Pi_{K}\alpha_{0}) ) )$, where $\Pi_{K} \alpha_{0}$ is the projection of $\alpha_{0}$ onto $\mathcal{A}_{K}$ (see Lemma \ref{lem:bound-hLo} in the Supplemental Material). By taking $\delta_{n}$ to balance both terms, it follows $ \delta_{2,L,n} = O\left(\sqrt{b^{2}_{2,J}/n + \{||E[g_{J}(Z,\Pi_{K}\alpha_{0})]||^{2}_{e} + \gamma_{K} Pen(\Pi_{K}\alpha_{0})\}}\right)$. Ignoring the term inside the curly brackets, this is the same rate than the one obtained by DIN in Lemma A.14 (note that in their setup $b^{2}_{2,J} \leq J$); the additional term inside the curly brackets stems from the fact that our estimation problem needs to be regularized and consequently $\alpha_{L,0}$ is not the true parameter that nullifies the moments.

 The sequence $(\delta_{1,L,n})_{L,n}$ is somewhat more standard within the semi-/non-parametric literature, e.g. \cite{chen2007large}, and its components essentially impose restrictions on the ``complexity" of $\mathcal{H}$. For instance, if $\mathcal{A} = \Theta \times \mathcal{H}$ is such that the classes $\{ \rho_{2}(\cdot,\cdot,\alpha) \colon \alpha \in \mathcal{A} \}$ and $\{ \mu h ^{\prime} \colon h \in \mathcal{H}  \}$ are P-Donsker, then $\Delta_{L,n} = O(1)$ and $\delta_{1,L,n} = O \left(  \sqrt{ \frac{J}{n}}\times b_{2,J} \right)$.

To further simplify the expression, suppose $b^{\rho}_{\rho,J} \leq J^{\rho/2}$; cf. Assumption 2 in \cite{DIN03} (see that paper for details and further references). Thus, under these conditions, the result in Theorem \ref{thm:conv-rate} simplifies to
{\small{\begin{align*}
||\hat{\alpha}_{L,n} - \alpha_{0} || = O_{\mathbf{P}} \left( (e_{min}(\mathcal{I}_{L} ))^{-1/2} \left(  \frac{J}{\sqrt{n}}  + ||E[g_{J}(Z,\Pi_{K}\alpha_{0})]||_{e} + \sqrt{\gamma_{K} Pen(\Pi_{K}\alpha_{0})}  \right)^{1/2} +  ||\alpha_{L,0} - \alpha_{0}|| \right);
\end{align*} }}
a rate governed by the degree of ill-posedness, the number $J$ of moment functions, the number $K$ of series terms and the bias arising from $\alpha_{L,0}$.



\section{Asymptotic Distribution Theory}

\label{sec:ADT} We now define the LR-type test statistic for the null
hypothesis $\theta_{0} = \nu$. For any $\nu \in \Theta$ and any $(L,n) \equiv (J,K,n)
\in \mathbb{N}^{3}$, let
\begin{align*}
\hat{\mathcal{L}}_{L,n}(\nu) \equiv 2 \left\{ \inf_{\{\alpha \in \mathcal{A}
_{K} \colon \theta = \nu\}}\left[ \sup_{\lambda \in \hat{\Lambda}_{J}(\alpha)}
\hat{S}_{J}(\alpha,\lambda) + \gamma_{K} Pen(\alpha)\right] - \inf_{\alpha \in
\mathcal{A}_{K}} \left[\sup_{\lambda \in \hat{\Lambda}_{J}(\alpha)} \hat{S}
_{J}(\alpha,\lambda) + \gamma_{K} Pen(\alpha)\right] \right\}.
\end{align*}

The goal of this section is to show that this statistic is asymptotically
chi-square distributed with one degree of freedom. The proof of this result
relies on a local quadratic approximation of the criterion function $\hat{S}
_{J}$ and a representation for the parameter of interest. To derive these
results, we define the following quantities: For any $\alpha \in \mathcal{A}
_{K}$ and any $(\theta, \zeta ) \in \mathbb{A}$, let
\begin{align*}
G(\alpha)[(\theta,\zeta)] = \frac{dE[g_{J}(Z,\alpha)]}{d\theta}\theta + \frac{dE[g_{J}(Z,\alpha)]}{dh}[\zeta] = \left[
\begin{array}{c}
\theta \\
\mathbf{0}
\end{array}
\right] + \left[
\begin{array}{c}
E[\ell(W) \zeta(W)] \\
E[\mathbf{p}_{Y|WX}(h(W)|W,X) \zeta(W) q^{J}(X) ]
\end{array}
\right]
\end{align*}
where $\mathbf{0}$ is a $J \times 1$ vector of zeros. By assumption \ref{ass:pdf0} these quantities are well-defined.


For any $L\in \mathbb{N}^{2}$ and for any $(\theta, \zeta ) \in \mathbb{A}$, we define another norm over $\mathbb{A}$ as,
\begin{equation*}
||(\theta, \zeta )||_{w}^{2}\equiv (G(\alpha _{L,0})[(\theta, \zeta )])^{T}H_{L}^{-1}(G(\alpha
_{L,0})[(\theta, \zeta ) ]),
\end{equation*}
where $H_{L}\equiv H_{J}(\alpha _{L,0},\mathbf{P})$. This norm acts as the
so-called \textquotedblleft weak norm" in \cite
{ai2003efficient,ai2007estimation}.


\subsection{Alternative Representation for the Weighted Average Derivative}


Lemma \ref{lem:weak-norm} in Appendix
\ref{app:ADT} shows that, over $lin\{\mathcal{A}_{K}\}$ for any $L=(J,K)\in
\mathbb{N}^{2}$, $||\alpha ||_{w}=0$ iff $\alpha =0$. This fact implies that
linear functionals are always bounded in the space $(lin\{\mathcal{A}
_{K}\},||.||_{w})$. Since $\theta $ can be interpreted as a linear
functional of $\alpha $, the following representation for $\theta $ holds:
For all $L=(J,K)\in \mathbb{N}^{2}$, there exists a $v_{L,n}^{\ast }\in
\mathcal{A}_{K}$ such that for any $\alpha =(\theta ,h)\in \mathcal{A}_{K}$,
\begin{equation*}
\theta =\langle v_{L,n}^{\ast },\alpha \rangle _{w},~and~||v_{L,n}^{\ast
}||_{w}=\sup_{a=(\theta ,h)\in lin\{\mathcal{A}_{K}\}, a\neq 0}\frac{|\theta|}{||a||_{w}}.
\end{equation*}

We note that, even though for each fixed $L \in \mathbb{N}^{2}$, $
||v^{\ast}_{L,n}||_{w} < \infty$, this quantity may diverge as $L$ diverges
if $\theta$ is not root-n estimable. Hence, we scale $v^{\ast}_{L,n}$ by its
norm, and define $u^{\ast}_{L,n} \equiv
v^{\ast}_{L,n}/||v^{\ast}_{L,n}||_{w} $. Then
\[
\frac{\hat{\theta}_{L,n} - \theta_{L,0}}{||v^{\ast}_{L,n}||_{w}}=\langle u^{\ast}_{L,n} , \hat{\alpha}_{L,n} - \alpha_{L,0} \rangle_{w}~.
\]


\begin{remark}[On the relationship between the Riesz representer and the Efficiency bound]\label{rem:Riesz}
	The weak norm of the Riesz representer, $||v_{L,n}^{\ast
	}||_{w}$, is the efficiency bound of $\theta_{0}$ in a model with $J+1$ unconditional moments functions, $E_{P}[g_{J}(Z,\cdot )]$, and $K+1$ parameters (which define $\alpha \in \mathcal{A}_{K}$).\footnote{\cite{ai2003efficient,ai2012semiparametric} established this claim for a richer model with conditional moments and infinite dimensional parameters.} For a suitably chosen sequence $L\equiv L(n)$ that increases as $n$ does --- since $(q_{j})_{j}$ is dense in $L^{2}(\mathbb{X},Leb)$ --- one expects the sequence of unconditional moment functions to approximate the moments (\ref{eqn:model1})-(\ref{eqn:model2}) defining the model. Thus, by the results in \cite{chamberlain1987} (see also lemma 3.3, lemma 4.1 and appendix A.1 in \cite{CP2015sieve})  one expects $(||v_{L(n),n}^{\ast}||_{w})_{n}$ to converge to the efficiency bound presented in Theorem \ref{thm:eff-bound} \emph{provided it is finite}. If the efficiency bound is infinite, the sequence $(||v_{L(n),n}^{\ast}||_{w})_{n}$ will diverge; this fact reflects the non root-n estimability of the weighted average derivative within the original model (\ref{eqn:model1})-(\ref{eqn:model2}).  $\triangle$
\end{remark}

\subsection{The Asymptotic distributions of $\hat{\theta}_{L,n}$ and LR statistic}


For any positive real-valued sequences $(\eta_{L,n},\eta_{w,L,n})_{L,n \in
\mathbb{N}^{3}}$ (they will be restricted below) and any $(L,n) \in \mathbb{N}
^{3}$, let
\begin{align*}
\mathcal{N}_{L,n} \equiv \{ \alpha \in \bar{\mathcal{A}}_{L,n} \colon
||\alpha - \alpha_{L,0} || \leq \eta_{L,n}~and~||\alpha - \alpha_{L,0}
||_{w} \leq \eta_{w,L,n} \}.
\end{align*}


In what follows, for any $(L,n) \in \mathbb{N}^{3}$, let $\hat{\alpha}
^{\nu}_{L,n}$ be the argument that minimizes the restricted criterion
function, i.e., $\hat{\alpha}^{\nu}_{L,n} \in \arg\min_{\{\alpha \in
\mathcal{A}_{K} \colon \theta = \nu\}} \sup_{\lambda \in \hat{\Lambda}
_{J}(\alpha)} \hat{S}_{J}(\alpha,\lambda) $. We impose the following
assumption that restricts the convergence rate of the unrestricted and
restricted PSGEL estimators.

\begin{assumption}
\label{ass:dev-N}
For any $L \in \mathbb{N}^{2}$ and $\alpha \in \{ \hat{\alpha}_{L,n} , \hat{
\alpha}^{\nu}_{L,n} \}$, if $\nu=\theta_{0}$: (i) $\alpha \in int(\mathcal{N}
_{L,n})$; (ii) $\gamma_{K}\sup_{t \colon |t| \leq l_{n} n^{-1/2}} |Pen(\alpha)
- Pen(\alpha + t u^{\ast}_{L,n})| = o_{\mathbf{P}}(n^{-1})$; (iii) There exists a $C< \infty$ such that for any $h \in \mathbb{H}$, $||h||_{L^{2}(Leb)}\leq C ||(0,h)||$.
\end{assumption}

Part (i) of this assumption ensure that both estimators --- the restricted
and unrestricted ones --- converge to $\alpha_{L,0}$ faster than $\eta_{L,n}$
and $\eta_{w,L,n}$ in the respective norms. One can use the results in
Section \ref{sec:conv-rate} to verify this assumption.\footnote{
The results in Section \ref{sec:conv-rate} apply to the restricted
estimator, under the null, with minimal changes.} Part (ii) ensures that the
penalty term is negligible (see also \cite{CP2015sieve}). Finally part (iii) states a relationship between the norm $h \mapsto ||(0,h)||$ --- used in Section \ref{sec:conv-rate} --- and the $L^{2}(Leb)$ norm over $\mathbb{H}$.

In the following assumption we let $\bar{\mathcal{G}}_{L,n}\equiv \{f(.,\alpha)=\rho _{2}(.,.,\alpha
)-\rho _{2}(.,.,\alpha _{L,0})\colon \alpha \in \mathcal{N}_{L,n}\}$.

\begin{assumption}
\label{ass:Donsker-LQA} There exists positive sequence, $(\Delta
_{2,L,n})_{L,n\in \mathbb{N}^{3}}$, such that, for any $L=(J,K)\in \mathbb{N}
^{2}$, $\sup_{\alpha=(\theta,h) \in \mathcal{N}_{L,n}}\mathbb{G}_{n}[\mu \cdot (h^{\prime
}-h_{L,0}^{\prime })]=O_{\mathbf{P}}(\Delta _{2,L,n})$ and for all $1\leq j\leq J$
, $\sup_{f\in \bar{\mathcal{G}}_{L,n}}\mathbb{G}_{n}[f\cdot
q_{j}]=O_{\mathbf{P}}(\Delta _{2,L,n})$.
\end{assumption}

This is a high-level assumption that controls one of the terms in the
remainder of the quadratic approximation in Lemma \ref{lem:LQA} below. As $
\mathcal{N}_{L,n}$ is shrinking, one would expect $\Delta _{2,L,n}=o(1)$; the
exact rate, however, depends on the complexity of $\bar{\mathcal{A}}_{L,n}$.

\begin{assumption}
\label{ass:HJ-sec} There exists a positive real-valued sequence, $
(\Xi_{L,n})_{L,n \in \mathbb{N}^{3}}$, such that, for any $L=(J,K) \in
\mathbb{N}^{2}$, $\sup_{\alpha \in \mathcal{N}_{L,n}} \left \Vert
H_{J}(\alpha,P_{n}) - H_{J}(\alpha_{L,0},P_{n}) - \{ H_{J}(\alpha,\mathbf{P}) -
H_{J}(\alpha_{L,0},\mathbf{P}) \} \right \Vert_{e} = O_{\mathbf{P}}(\Xi_{L,n})$.
\end{assumption}

This high-level assumption implies stochastic equi-continuity of the process
$H_{J}(\cdot ,P_{n})$, and it is used to control one of the terms in the
remainder of the quadratic approximation in Lemma \ref{lem:LQA} below.

The final two assumptions impose additional restrictions on $(\eta
_{L,n},\eta _{w,L,n})_{L,n\in \mathbb{N}^{3}}$, $(b_{\rho ,J})_{\rho \in
\mathbb{R},J\in \mathbb{N}}$, $(\delta _{n})_{n\in \mathbb{N}}$ and the rate
at which $L=(J,K)$ diverges relative to $n$.

\begin{assumption}
\label{ass:undersmooth} (i) $\frac{\sqrt{n}}{||v^{\ast}_{L,n}||_{w}}
||E[g_{J}(Z,\alpha_{L,0})]||_{e} = o(1)$; (ii) $\frac{\sqrt{n}}{
||v^{\ast}_{L,n}||_{w}} |\theta_{L,0}-\theta_{0}|= o(1)$.
\end{assumption}

This assumption implies that the ``bias" terms arising from working with $
\alpha_{L,0}$, as opposed to $\alpha_{0}$, are small relative to the rate we
are using to scale the leading term of the asymptotic expansions below $
\frac{\sqrt{n}}{||v^{\ast}_{L,n}||_{w}}$. Similar assumptions have been
imposed in the literature, e.g. \cite{CP2015sieve} and reference therein.

\begin{assumption}
\label{ass:rates-LQA} (i) $n \delta^{3}_{n} ( \mho_{L,n} + b_{3,J} )^{3} =
o(1)$, $n\delta^{2}_{n} \left( \{ \theta_{L,0}^{2} + || h^{\prime
}_{L,0}||^{2}_{L^{\infty}(\mathbb{W},\mu)} \} \sqrt{\frac{b_{4,J}}{n}} + \mho_{L,n} \eta_{L,n} + \Xi_{L,n} \right) = o(1)$ and $n\delta_{n} \left( \sqrt{\frac{J}{n}}\Delta_{2,L,n} + \eta_{L,n}^{2} b_{2,J} \right) = o(1)$; (ii) $\sqrt{
\bar{g}^{2}_{L,0}/n + ||E[g_{J}(Z,\alpha_{L,0})]||^{2}_{e} } + \eta_{w,L,n}
=o(\delta_{n})$ and $\left( \sqrt{\frac{J}{n}}\Delta_{2,L,n} +
\eta_{L,n}^{2} b_{2,J} \right) = o(\delta_{n})$; (iii) There exists a $
\varrho>0$ such that $||h^{\prime }_{L,0}||^{2+\varrho}_{L^{\infty}(\mathbb{W},\mu)}/n^{2+\varrho} = o(1)$ and $b^{2+\varrho}_{2+\varrho,J}/n^{2+
\varrho} = o(1)$; (iv) $A_{L,0} \equiv E[\mathbf{p}_{Y|WX}(h_{L,0}(W)\mid W,X)
q^{J}(X)\varphi^{K}(W)^{T} ]$ has full rank $K$ and $n^{-1/2}
e_{min}(A_{L,0}^{T}A_{L,0})^{-1} = o(\eta_{L,n})$; (v) $||h^{\prime
}_{L,0}||^{2}_{L^{\infty}(\mathbb{W}, \mu)} b^{2}_{4,J}/\sqrt{n} = o(1)$.
\end{assumption}

Part (i) ensures that the remainder term for the asymptotic quadratic
representation of $\hat{S}_{J}$ is negligible (see Lemma \ref{lem:LQA}). The sequence $(\delta _{n})_{n}$ in part (ii) was discussed after Assumption \ref
{ass:rates}. Part (iii) is used to show asymptotic normality of the leading
term in Lemma \ref{lem:QLR-A-rep} by means of a Lyapounov condition. Finally,
part (iv) ensures that the weak norm is proportional to the strong norm over
$\mathcal{A}_{K}$ (even though the constant of proportionality may vanish as
$L$ diverges) and that deviations of the form $\alpha
+l_{n}n^{-1/2}u_{L,n}^{\ast }$ stay in $\mathcal{N}_{L,n}$ (see Lemma \ref
{lem:charac-N} in Appendix \ref{app:ADT}). These deviations play a crucial
role in the proof of Lemma \ref{lem:QLR-A-rep}.

\begin{remark}[The rate restrictions of Assumption \ref{ass:rates-LQA}]
	While parts (iii)-(v) are fairly easy to check and interpret, parts (i)-(ii) are not as easy. The goal of this remark is to illustrate the restrictions imposed by these parts on the different rates $(\delta_{n},\eta_{L_{n},n},\eta_{w,L_{n},n},\Delta_{2,L_{n},n},\Xi_{L_{n},n})_{n}$ where $(L_{n})_{n}$ is a diverging sequence in $\mathbb{N}^{2}$. To do this, we take as the point of departure the setting described in Section \ref{sec:DiscussionRate}, which allows us to simplify some expressions. Under this setup, part (i) imposes $\delta_{n} = o\left( n^{-1/3} J^{-1/2}_{n}  \right)$. Given this, the restrictions in parts (i)-(ii) imply that
	 $\eta_{L_{n},n} = O(\min\{ n^{-1/2}\delta^{-1/2}_{n}J^{-1/2}_{n}, n^{-1/6} J^{-3/4}_{n}\})$
 and $\eta_{w,L_{n},n} = o(n^{-1/3} J^{-1/2}_{n})$; we note that by imposing  a polynomial rate of decay, this condition rules out the so-called severely ill-posed case wherein the rate of for $(\eta_{L_{n},n})_{n}$ decays slower than polynomial order (see \cite{CP2012estimation} and references therein). Parts (i)-(ii) also imply that $\Delta_{2,L,n} = o( (\sqrt{nJ_{n}}\delta^{2}_{n})^{-1}  )$ and $\Xi_{L,n} = o(n^{-1} \delta_{n}^{-2})$; for the ``worst case" where $\delta_{n} = \left( n^{-1/3} J^{-1/2}_{n}  \right)/l_{n}$, it follows that $\Delta_{2,L,n} = O( J_{n} n^{-1/6}  )$ and $\Xi_{L,n} = O(n^{-1/3} J_{n} )$, but the restriction can be relaxed if $(\delta_{n})_{n}$ decays faster. Finally, parts (i)-(ii) impose restrictions on the growth of $(L_{n})_{n}$: $J_{n} = O(n^{-1/6})$ and $\sqrt{J_{n}||E[g_{J_{n}}(Z,\alpha_{L_{n},0})]||^{2}_{e}} = o(n^{-1/3})$. $\triangle$
\end{remark}


The following result characterizes the asymptotic distribution of the LR test statistic under the null. This characterization holds regardless of whether the parameter $\theta_{0}$ is root-$n$ estimable or not.

\begin{theorem}
\label{thm:QLR} Let Assumptions \ref{ass:ident}-\ref{ass:ID-pseudotrue}
and \ref{ass:dev-N}-\ref{ass:rates-LQA} hold. Then, under the null $\theta_{0} = \nu$,
\begin{align*}
\hat{\mathcal{L}}_{L,n}(\theta_{0}) \Rightarrow \chi^{2}_{1}.
\end{align*}
\end{theorem}

\begin{proof}
	See Appendix \ref{app:Thm-QLR}.
\end{proof}

This result extends those in \cite
{ParenteSmith2011gel} to a non-parametric setup where the GEL is constructed
using an increasing number of moment conditions, and wherein the parameter
of interest may not be root-n estimable. Using a related estimator ---an
EL-based on conditional moments a la \cite{kitamura2004empirical} --- \cite
{tao2013empirical} derived an analogous result but her assumptions rule out non-smooth residuals, relevant for the quantile IV model considered here.

As a by-product of the derivations used to prove Theorem \ref{thm:QLR}, an asymptotic linear representation for the estimator of the WAD is obtained.

\begin{theorem}
	\label{thm:normal-thetahat} Let Assumptions \ref{ass:ident}-\ref{ass:ID-pseudotrue}, \ref{ass:dev-N} (for $\hat{\alpha}_{L,n}$), \ref{ass:Donsker-LQA}, \ref{ass:HJ-sec} and \ref{ass:rates-LQA} hold. Then
	\begin{align*}
	\frac{\hat{\theta}_{L,n} - \theta_{L,0}}{||v^{\ast}_{L,n}||_{w}}=
	n^{-1} \sum_{i=1}^{n} (G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1}
	g_{J}(Z_{i},\alpha_{L,0}) + o_{\mathbf{P}}(n^{-1/2}).
	\end{align*}
	Further, under Assumption \ref{ass:undersmooth}, we have
	\[
	\frac{\sqrt{n}(\hat{\theta}_{L,n} - \theta_{0})}{||v^{\ast}_{L,n}||_{w}}\Rightarrow N(0,1)~.
	\]
\end{theorem}


The proof is the same as that Lemma \ref{lem:ALR-thetahat} in Appendix \ref{app:ADT} so it is omitted. This result illustrates the role of $||v^{\ast}_{L,n}||_{w}$ as the appropriate scaling of our estimator. If the sequence $(||v^{\ast}_{L,n}||_{w})_{n}$ is uniformly bounded, then this theorem implies that $\hat{\theta}_{L,n}$ is $\sqrt{n}$ asymptotically Gaussian. On the other hand, if the sequence diverges, Gaussianity is still preserve but the rate is slower and given by $\sqrt{n}/||v^{\ast}_{L,n}||_{w}$.



\subsection{Heuristics}

The idea is to show that, asymptotically, $\hat{\mathcal{L}}_{L,n}$ is a quadratic form of Gaussian random variables. The first step is to provide a
quadratic approximation for the criterion function $\hat{S}_{J}(\alpha,\cdot)$ as a function of $\lambda $, as shown in the following lemma.
\begin{lemma}
\label{lem:LQA} Let Assumptions \ref{ass:ident}-\ref{ass:ID-pseudotrue}, \ref{ass:Donsker-LQA}, \ref{ass:HJ-sec} and \ref{ass:rates-LQA}(v) hold. Then uniformly over $(\alpha,\lambda) \in \mathcal{N}_{L,n} \times
\{ \lambda \in \mathbb{R}^{J+1} \colon ||\lambda||_{e} \leq \delta_{n} \}$,
for any $L=(J,K) \in \mathbb{N}^{2}$
\begin{align*}
\hat{S}_{J}(\alpha,\lambda) =& - \lambda^{T} \varDelta(\alpha) - \frac{1}{2}
\lambda^{T} H_{L} \lambda \\
& + O_{\mathbf{P}} \left( \delta^{3}_{n} ( \overline{\theta} + l_{n}\gamma_{K}^{-1}
\Gamma_{L,n} + b_{3,J} )^{3} \right) \\
& + O_{\mathbf{P}} \left( \delta^{2}_{n} \left( ( \overline{\theta} + ||h^{\prime
}_{L,0}||_{L^{\infty}(\mathbb{W},\mu)} )^{2} \sqrt{b_{4,J}/n} + \mho_{L,n}
\eta_{L,n} + \Xi_{L,n} \right) \right) \\
& + O_{\mathbf{P}} \left( \delta_{n} \left( \sqrt{\frac{J}{n}}\Delta_{2,L,n} +
\eta_{L,n}^{2} b_{2,J} \right) \right).
\end{align*}
where $\varDelta(\alpha) \equiv n^{-1} \sum_{i=1}^{n} g_{J}(Z_{i},\alpha_{L,0}) +
G(\alpha_{L,0})[\alpha - \alpha_{L,0}]$.
\end{lemma}

\begin{proof}
	See Appendix \ref{app:ADT}.
\end{proof}

The ``remainder" terms in the RHS (the $O_{\mathbf{P}}(.)$ terms) are fairly intuitive: the order $\delta _{n}^{3}$
-term requires boundedness of the third derivative of $\hat{S}_{J}(\alpha
,\cdot )$; the $\delta _{n}^{2}$-term arises because the expansion yields a
quadratic term with $H_{J}(\alpha ,P_{n})$ as opposed to $H_{L}$; and the $
\delta _{n}$-term is the error of approximating $n^{-1}
\sum_{i=1}^{n}g_{J}(Z_{i},\alpha )$ with $\varDelta(\alpha )$. This last
part handles the non-smooth nature of the residuals $\rho _{2}$ by using $
E[g_{J}(Z,\cdot )]$, which is a smooth function. Assumption \ref{ass:rates-LQA}(i) ensures that these `remainder" terms are in fact $o_{\mathbf{P}}(n^{-1})$. This fact, and the fact that $\hat{\Lambda}
_{J}(\alpha)$ contains a $\delta_{n}$-ball (see Lemma \ref{lem:Lambda-charac}
in the Supplemental Material \ref{supp:conv-rate}), imply that the expression in the Lemma provides an asymptotic characterization for
$\sup_{\lambda \in \hat{\Lambda}_{J}(\alpha)} \hat{S}_{J}(\alpha,\lambda)$
in terms of $(\varDelta(\alpha))^{T}H^{-1}_{L}(\varDelta(\alpha))$, which is
a quadratic form in $\alpha$.

With this result at hand and Assumption \ref{ass:dev-N}, one can obtain lower and upper bounds for $\hat{\mathcal{L}}_{L,n}(\theta_{0})$ of the form,
\begin{align*}
\hat{\mathcal{L}}_{L,n}(\theta_{0})  \geq (\varDelta(\hat{\alpha}_{L,n}))^{T}H^{-1}_{L}(\varDelta(\hat{\alpha}_{L,n})) - (\varDelta(\hat{\alpha}_{L,n}) + t u^{\ast}_{L,n})^{T}H^{-1}_{L}(\varDelta(\hat{\alpha}_{L,n} + t u^{\ast}_{L,n})) + o_{\mathbf{P}}(1),
\end{align*}
for appropriately chosen $t \in \mathbb{R}$, and
\begin{align*}
\hat{\mathcal{L}}_{L,n}(\theta_{0})  \leq (\varDelta(\hat{\alpha}^{\theta_{0}}_{L,n}  + t u^{\ast}_{L,n} ))^{T}H^{-1}_{L}(\varDelta(\hat{\alpha}^{\theta_{0}}_{L,n}  + t u^{\ast}_{L,n})) - (\varDelta(\hat{\alpha}^{\theta_{0}}_{L,n})^{T}H^{-1}_{L}(\varDelta(\hat{\alpha}^{\theta_{0}}_{L,n})) + o_{\mathbf{P}}(1),
\end{align*}
for appropriately chosen $t \in \mathbb{R}$. Since $\alpha \mapsto \varDelta(\alpha)$ is an affine function, the RHS in the previous expression is fairly easy to characterize. The following lemma formalizes these steps (its proof presents the explicitly choice for $t$ in the previous two displays).
\begin{lemma}
\label{lem:QLR-A-rep} Let Assumptions \ref{ass:ident}-\ref{ass:ID-pseudotrue}, \ref{ass:dev-N}-\ref{ass:rates-LQA} hold. Then, under the
null $\nu = \theta_{0}$,
\begin{align*}
&\hat{\mathcal{L}}_{L,n}(\theta_{0}) - \left( n^{-1/2} \sum_{i=1}^{n}
(G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0})
\right)^{2} \\
& \geq 2\sqrt{n}\frac{(\theta_{0} - \theta_{L,0})}{||v^{\ast}_{L,n}||_{w}}
\left( n^{-1/2} \sum_{i=1}^{n}
(G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0})
\right)+ o_{\mathbf{P}}(1).
\end{align*}
and
\begin{align*}
&\hat{\mathcal{L}}_{L,n}(\theta_{0}) - \left( n^{-1/2} \sum_{i=1}^{n}
(G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0})
\right)^{2} \\
& \leq 2\sqrt{n}\frac{(\theta_{0} - \theta_{L,0})}{||v^{\ast}_{L,n}||_{w}}
\left( n^{-1/2} \sum_{i=1}^{n}
(G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0})
\right) \\
& + \left( \sqrt{n}\frac{(\theta_{0} - \theta_{L,0})}{||v^{\ast}_{L,n}||_{w}}
\right)^{2} + o_{\mathbf{P}}(1).
\end{align*}
\end{lemma}

\begin{proof}
	See Appendix \ref{app:ADT}.
\end{proof}

This lemma shows the reason for Assumption \ref{ass:undersmooth} in our analysis,
as this assumption ensures that
\begin{align*}
\hat{\mathcal{L}}_{L,n}(\theta_{0}) = \left( n^{-1/2} \sum_{i=1}^{n}
(G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0})
\right)^{2} + o_{\mathbf{P}}(1).
\end{align*}

Under mild assumptions and Assumption \ref{ass:undersmooth}, the object
inside the parenthesis is asymptotically Normal with mean 0 and variance 1.
Here we see the importance of the \textquotedblleft optimal weight", $
H_{L}^{-1}$. If $H_{L}$ differed from $E[g_{J}(Z,\alpha
_{L,0})g_{J}(Z,\alpha _{L,0})^{T}]$, then the variance of the term inside
the parenthesis will not be equal to 1, and the test statistic will only be
proportional to a $\chi _{1}^{2}$ in the limit; see \cite{CP2015sieve} for a
more thorough discussion and results for this case.


\section{Conclusion}

\label{sec:conclusion}

Since the seminal work by Koenker and Bassett about 40 years ago (\cite{koenker1978regression}), quantile regression models have become ubiquitous in econometrics and statistics; see \cite{Koenker2018} for a recent survey. The original linear quantile regression model has been extended in several directions; in particular to the general non-parametric IV framework that allows for ``flexible functional forms" and endogeneity of the regressors. This type of model, while very general, presents technical challenges arising from the non-smooth nature of the criterion function as well as its ill-posedness. One goal of this paper is to shed some light on how the nonlinear ill-posedness of the non-parametric quantile IV (NPQIV) model affects not only the speed of convergence to the conditional quantile function but also the accuracy for estimating even simple linear functionals. For this, we derive the semiparametric efficiency bound for a particular linear functional of the NPQIV --- the weighted average derivative (WAD).

To estimate the parameters of interest --- the NPQIV function and its WAD ---  we propose a general penalized sieve GEL procedure based on the unconditional WAD moment restriction and an increasing number of
unconditional moments that are asymptotically equivalent to the conditional moment defining the NPQIV model (\ref{eqn:model1}). We show that the QLR statistic based on the penalized sieve GEL is asymptotically chi-square distributed regardless of whether or not the information bound of the WAD is singular. This result can be used to construct confidence sets for the WAD without the need to estimate the variance of the estimator of the WAD. We hope these results extend even further the scope of quantile regression models.

The penalized sieve GEL procedure is more generally applicable to any
semi/nonparametric conditional moment restrictions and unconditional moment
restrictions, say of the following form:
\begin{align}
E[\rho _{2}(Y,W;\theta _{02},h_{01}(\cdot ),...,h_{0q}(\cdot ))|X]& =0,\text{
\quad a.s.-}X\text{,}  \label{semi00} \\
E[\rho _{1}(Y,W;\theta _{01},\theta _{02},h_{01}(\cdot ),...,h_{0q}(\cdot
))]& =0\text{.}  \label{semi01}
\end{align}
Here $Y$ denotes dependent (or endogenous) variables, $X$ denotes
conditioning (or instrumental) variables and $W$ could be either endogenous or subset of $X$, $\theta =(\theta _{1}^{\prime
},\theta _{2}^{\prime })^{\prime }$ denotes a vector of finite dimensional
parameters, and $h(\cdot )=\left( h_{1}(\cdot ),...,h_{q}(\cdot )\right) $ a
$q\times 1$ vector of real-valued measurable functions of $Y$, $W$, $X$ and other
unknown parameters. The residual functions $\rho _{j}(y,w;\theta ,h(\cdot ))$
, $j=1,2$, could be nonlinear, pointwise non-smooth with respect to $(\theta
,h)$. And some of the $\theta$ could have singular information bound. This
is a valuable alternative to classical semiparametric two-step GMM
when the second step finite dimensional parameter $\theta$ might not be root-
$n$ estimable.

In an old unpublished  draft, \cite{CP2010} study the asymptotic properties of
another estimation procedure, optimally weighted penalized Sieve Minimum
Distance (SMD) based on orthogonalized residuals for model (\ref{semi00})-(
\ref{semi01}). Under a set of regularity conditions, including the assumption that the WAD of a NPQIV has a positive information bound, \cite{CP2010} establish that their optimally weighted penalized SMD
estimator of the WAD is root-$n$ asymptotically normal and semiparametrically efficient. It would
be interesting to compare this paper's estimator against theirs, and
we leave this to future work.

\bibliographystyle{plainnat}
\bibliography{QAD-biblio}