EconBase
← Back to paper

Penalized Sieve GEL for Weighted Average Derivatives of Nonparametric Quantile IV Regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,102 characters · 16 sections · 93 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Penalized Sieve GEL for Weighted Average Derivatives of Nonparametric Quantile IV Regressions

abstract\singlespacing This paper considers estimation and inference for a weighted average derivative (WAD) of a nonparametric quantile instrumental variables regression (NPQIV). NPQIV is a non-separable and nonlinear ill-posed inverse problem, which might be why there is no published work on the asymptotic properties of any estimator of its WAD. We first characterize the semiparametric efficiency bound for a WAD of a NPQIV, which, unfortunately, depends on an unknown conditional derivative operator and hence an unknown degree of ill-posedness, making it difficult to know if the information bound is singular or not. In either case, we propose a penalized sieve generalized empirical likelihood (GEL) estimation and inference procedure, which is based on the unconditional WAD moment restriction and an increasing number of unconditional moments that are implied by the conditional NPQIV restriction, where the unknown quantile function is approximated by a penalized sieve. Under some regularity conditions, we show that the self-normalized penalized sieve GEL estimator of the WAD of a NPQIV is asymptotically standard normal. We also show that the quasi likelihood ratio statistic based on the penalized sieve GEL criterion is asymptotically chi-square distributed regardless of whether or not the information bound is singular. JEL Classification: C14; C22 Keywords: Nonparametric quantile instrumental variables; Weighted average derivatives; Penalized sieve generalized empirical likelihood; Semiparametric efficiency; Chi-square inference.

Introduction

Since the seminal paper by koenker1978regression, quantile regressions and functionals of quantile regressions have been the subjects of ever-expanding theoretical research and applications in economics, statistics, biostatistics, finance, and many other science and social science disciplines. See koenker2005quantile and the forthcoming Handbook of Quantile Regression (2017) for the latest theoretical advances and empirical applications.

The presence of endogenous regressors is common in many empirical applications of structural models in economics and other social sciences. The Nonparametric Quantile Instrumental Variable (NPQIV) regression, $ E[1\{ Y \leq h_{0}(W) \} - \tau|X]=0$, was, to our knowledge, first proposed in chernozhukov2005iv and CIN2007instrumental. This model is a leading important example of nonlinear and non-separable ill-posed inverse problems in econometrics, which has been an active research topic following the Nonparametric (mean) Instrumental Variables (NPIV) regression, $E[ Y - h_{0}(W)|X]=0$, studied by NP2003instrumental, hall2005nonparametric, BCK2007, CFR2007linear, darolles2011nonparametric and others. See, for example, horowitz2007nonparametric, CP2009efficient,CP2012estimation,CP2015sieve, gagliardini2012nonparametric, CH2013quantilerev, CCLN2014local and others for recent work on the NPQIV and its various extensions.

In this paper, we consider estimation and inference for a Weighted Average Derivative (WAD) functional of a NPQIV. For models without nonparametric endogeneity, WAD functionals of nonparametric (conditional) mean regression, $E[ Y - h_{0}(X)|X]=0$, and of quantile regression, $E[1\{ Y \leq h_{0}(X) \} - \tau|X]=0$, have been extensively studied in both statistics and econometrics. In particular, under some mild regularity conditions, plug-in estimators for WADs of any nonparametric mean and quantile regressions can be shown to be semiparametrically efficient and root-$n$ asymptotically normal (where $n$ is the sample size). See, for example, newey1993efficiency, newey1994asymptotic, NeweyPowell1999, ackerberg2014asymptotic and the references therein. Although unknown functions of endogenous regressors occur frequently in empirical work, due to the ill-posed nature of NPIV and NPQIV, there is not much research on WAD functionals of NPIV and NPQIV yet. In fact, even for the simpler NPIV model that is a linear and separable ill-posed inverse problem, it is still a difficult question whether a linear functional of a NPIV could be estimated at the root-$n$ rate; see, e.g., SeveriniTripathi2012 and davezies2015existence. Although ai2007estimation provide low-level sufficient conditions for a root-$n$ consistent and asymptotically normal estimator of the WAD of the NPIV model, and ai2012semiparametric provide a semiparametric efficient estimator of WAD for that model, to our knowledge, there is no published work on semiparametric efficient estimation of the WAD for the NPQIV model yet.

We first characterize the semiparametric efficiency bound for the WAD functional of a NPQIV model. Unfortunately, the bound depends on an unknown conditional derivative operator and hence an unknown degree of ill-posedness. Therefore, it is difficult to know if the semiparametric information bound is singular or not. Further, even if a researcher assumes that the information bound is non-singular and the WAD is root-$n$ consistently estimable, the results in ai2012semiparametric and CS2015overidentification show that a simple plug-in estimator of a WAD might not be semiparametrically efficient. This is in contrast to the results of newey1993efficiency and ackerberg2014asymptotic who show that plug-in estimators of a WAD of a nonparametric mean and quantile regression are semiparametrically efficient.

We then propose penalized sieve Generalized Empirical Likelihood (GEL) estimation of the WAD for the NPQIV model, which is based on the unconditional WAD moment restriction and an increasing number of unconditional moments implied by the conditional moment restriction of the NPQIV model, where the unknown quantile function is approximated by a flexible penalized sieve. Under some regularity conditions, we show that the self-normalized penalized sieve GEL estimator of the WAD of a NPQIV is asymptotically standard normal. We also show that the Quasi Likelihood Ratio (QLR) statistic based on the penalized sieve GEL criterion is asymptotically chi-squared distributed regardless of whether the information bound is singular or not; this can be used to construct confidence sets for the WAD of NPQIV without the need to estimate the variance nor the need to know the precise convergence rates of the WAD estimator.

Our estimation procedure builds upon DIN03, who approximate a conditional moment restriction $E[\rho (Y,\theta_0 )|X]=0$ by an increasing sequence of unconditional moment restrictions, and then consider estimation of the Euclidean parameter $\theta_0$ (of fixed and finite dimension) and specification tests based on GEL (and related) procedures. For the same model $E[\rho (Y,\theta_0 )|X]=0$, kitamura2004empirical directly estimate the conditional moment restriction via kernel and then apply a kernel-based conditional empirical likelihood (EL) to estimate $ \theta_0$. However, the model considered in these papers does not contain any unknown functions (say $h()$) and the residuals $\rho (.,\theta)$ are assumed to be twice continuously differentiable with respect to $\theta$ at $ \theta_0 $. For the semiparametric conditional moment restriction $E[\rho (Y,\theta_0, h_0 (\cdot) )|X]=0$ when the unknown function $h(\cdot )$ could depend on an endogenous variable, otsu2011large and tao2013empirical consider a sieve conditional EL extension of kitamura2004empirical, and sueishi2017 provides a sieve unconditional GEL extension of DIN03, where the unknown function $h(.)$ is approximated by a finite dimensional linear sieve (series) as in ai2003efficient. However, like ai2003efficient, all these papers assume twice continuously differentiable residuals $ \rho(.,\theta, h(.))$ with respect to $(\theta_0, h_0(.))$, and hence rule out the NPQIV model.

ParenteSmith2011gel study GEL properties for non-smooth residuals $ g(.,.)$ in the unconditional moment models $E[g(Y,\theta _{0})]=0$, but require the dimensions of both $g(.,.)$ and $\theta _{0}$ to be fixed and finite. Finally, horowitz2007nonparametric, gagliardini2012nonparametric, CP2009efficient,CP2012estimation,CP2015sieve, and CNS2015constrained do include the NPQIV model, but none of these papers addresses the issues of estimation and inference for the WAD of the NPQIV.

The rest of the paper is organized as follows. Section (ref) introduces notation and the model. Section (ref) characterizes the semiparametric efficiency bound for the WAD of the NPQIV model. Section (ref) introduces a flexible penalized sieve GEL procedure. Section (ref) derives the consistency and the convergence rates of the penalized sieve GEL estimator for the NPQIV model. Section (ref) establishes the asymptotic distributions of the WAD estimator and of the QLR statistic based on penalized sieve GEL for the WAD of a NPQIV. Section (ref) concludes with a discussion of extensions.

Preliminaries and Notation

Let $Z\equiv (Y,W,X)$ be the observable data vector, where $Y$ is the outcome variable, $W$ is the endogenous variable and $X$ is the instrumental variable (IV); we assume the observable data, $Z$, is distributed according to a probability distribution $\mathbf{P}$. In order to simplify the exposition, we restrict attention to real-valued continuous random variables, i.e., we assume $\mathbf{P}$ has a density $\mathbf{p}$ with support given by $\mathbb{Z}\equiv \mathbb{Y}\times \mathbb{W}\times \mathbb{X}\subseteq \mathbb{R}^{3}$; extending our results to vector-valued endogenous and instrumental variables would be straightforward but cumbersome in terms of notation.

Notation. For any subset, $\mathbb{Z}$, of an Euclidean space let $\mathcal{P}(\mathbb{Z})$ be the class of Borel probability measures over $\mathbb{Z}$. For any $P\in \mathcal{P}(\mathbb{Z})$, we use $p$ to denote its probability density function (pdf) (with respect to Lebesgue (Leb) measure) and $supp(P)$ to denote its support. We also use $P_{X}$ ($p_{X}$) to denote the marginal probability (pdf) of a random variable $X$; and $P_{Y|X}$ ($p_{Y|X}$) to denote the conditional probability (pdf) of $Y$ given $X$. For expectation, we write $E_{Q}[.]$ to be explicit about the fact that $Q$ is the measure of integration; throughout we sometimes use $E[.]\equiv E_{\mathbf{P}}[.]$ when $\mathbf{P}$ is the true probability of the data. The term “wpa1" stands for “with probability approaching one (under $\mathbf{P}$)"; for any two real-valued sequences $(x_{n},y_{n})_{n}$ $x_{n} \precsim y_{n}$ denotes $x_{n} \leq C y_{n}$ for some $C$ finite and universal; $\succsim$ is defined analogously. For any $q \geq 1$, we use $L^{q}(Q)\equiv L^{q}(\mathbb{Z},Q)$ to denote the class of measurable functions $f : \mathbb{Z} \mapsto \mathbb{R}$ such that $||f||_{L^{q}(Q)}=\left( \int_{z\in \mathbb{Z}} |f(z)|^{q}Q(dz)\right) ^{1/q} < \infty$; as usual $L^{\infty}(Leb)$ denotes the class of essentially bounded real-valued functions. We use $||.||_{e}$ to denote the Euclidean norm, $\mathbb{R}_{+}=[0,\infty )$ and $\mathbb{R}_{++}=(0,\infty )$.

For any subset $S$ of a vector space $(\mathbb{S},||.||_S)$, $lin \{S\}$ denotes the smallest linear space containing $S$; for any subspace $A \subseteq \mathbb{S}$, $A^{\perp}$ denotes its orthogonal complement in $(\mathbb{S},||.||_S)$. For any linear operator, $M : (\mathbb{S}_1,||.||_1 ) \rightarrow (\mathbb{S}_2,||.||_2 )$, let $Kernel(M) \equiv \{ x\in \mathbb{S}_1 \colon M[x] = 0 \}$ and $Range(M) \equiv \{ y \in \mathbb{S}_2 \colon \exists x \in \mathbb{S}_1,~M[x] = y \}$; it is bounded if and only if $\sup_{x\in \mathbb{S}_1 : ||x||_1=1} ||M[x]||_2 <\infty$. For any linear bounded operator $M$, $M^{+}$ denotes its generalized inverse; see, e.g., engl1996regularization.

The WAD of the NPQIV model

Let $\mathbb{A}\equiv \mathbb{R}\times \mathbb{H}$, where $\mathbb{H} = \{ h \in L^{2}(Leb) \colon h^{\prime}~exists~and~||h^{\prime}||_{L^{2}(Leb)} < \infty \}$, i.e., $\mathbb{H}$ is a Sobolev space of order $1$, here $h^{\prime}$ should be viewed as a weak derivative of $h$ (see brezis2010functional). We note that $\mathbb{H}$ is a Hilbert space under the norm $||h||_{\mathbb{H}}\equiv ||h||_{L^{2}(Leb)}+||h'||_{L^{2}(Leb)}$, and $\mathbb{A}$ is a Hilbert space under the norm $||(\theta, h)||_{\mathbb{A}}\equiv ||\theta||_e + ||h||_{\mathbb{H}}$. In this paper we measure convergence in $\mathbb{A}$ using another norm $||(\theta, h)||\equiv ||\theta||_e + ||h||$ for $||h||\leq ||h||_{\mathbb{H}}$ (such as $||h||=||h||_{L^{2}(Leb)}$). The parameter set is given by $\mathcal{A} \equiv \Theta \times \mathcal{H} \subseteq \mathbb{A}$, where $\Theta$ is bounded and convex and $\mathcal{H}$ is a set that contains additional restrictions on $h \in \mathbb{H}$ which will be specified below. We assume that $\mathbf{P}$ is such that there exists a parameter $\alpha_{0} \equiv (\theta_{0},h_{0}) \in \mathcal{A}$ that satisfies

align[align omitted — 164 chars of source]

for $\tau \in (0,1)$, where $\mu$ is a nonnegative, continuously differentiable scalar function in $L^{\infty}(Leb) \cap \mathbb{H}$ and should be viewed as the weighting function of the average derivative, $\theta_{0}$, of $h_{0}$.

The following assumption ensures that the conditions above uniquely identify $\alpha_{0}$; it will be maintained throughout the paper and will not be explicitly referenced in the results below.

assumptionThere is a unique $\alpha _{0}\in int(\mathcal{A})$ that satisfies model ((ref))-((ref)).

The interior assumption is needed only for the asymptotic distribution results in Section (ref). In cases where $\mathcal{H}$ has an empty interior, one can use the concept of relative interior of $\mathcal{H}$. This assumption is clearly high level. The goal of this paper is to characterize the asymptotic behavior of a modified GEL estimator of $\alpha $ , taking as given the identification part; for a discussion of primitive conditions for Assumption (ref), we refer the reader to CCLN2014local and references therein.

The following assumption imposes additional restrictions over the primitives: $\mu $, $\mathbf{P}$ and $ \alpha _{0}$.

assumption(i) $\mathbf{P}$ has a continuously differentiable pdf, $\mathbf{p}$, such that: the marginal density $\mathbf{p}_{W}$ of $W$ is uniformly bounded, zero at the boundary of the support and $\mathbf{p}^{\prime}_{W} \in L^{2}(Leb)$; the marginal density $\mathbf{p}_{X}$ of $X$ is uniformly bounded away from 0 on its support; $\sup_{y,w,x \in \mathbb{Z}} \mathbf{p}_{Y|WX}(y \mid w, x) < \infty$, $\sup_{y,w,x \in \mathbb{Z}} \frac{d\mathbf{p}_{Y|WX}(y \mid w, x)}{dy} < \infty$ ; (ii) $\mathcal{H}$ is convex and such that for all $h \in \mathcal{H}$, $\sup_{w \in \mathbb{W}} |\mu(w) h(w)| < \infty$; (iii) $Var_{\mathbf{P}}(\mu(W)h^{\prime }(W)) > 0$ for all $h \in \mathcal{H}$ in a $||\cdot ||$-neighborhood of $h_{0}$.

Part (i) of this condition imposes differentiability and boundedness restrictions on different elements of $\mathbf{p}$; part (ii) ensures that $\lim_{w \rightarrow \pm \infty} \mathbf{p}_{W}(w) \mu (w) h(w) = 0$ which allows for an alternative representation for $\theta_{0}$ using integration by parts (see expression (ref) below); part (iii) is a high level assumption and essentially implies $Var_{\mathbf{P}}(\mu(W)h^{\prime }_{0}(W)) > 0$ as well as continuity of $h \mapsto Var_{\mathbf{P}}(\mu(W)h^{\prime }(W))$.

Efficiency Bound for $\theta_{0}$

By definition of $\mathbb{H}$, Assumption (ref) and integration by parts, it follows that

align[align omitted — 100 chars of source]

where $$w \mapsto \ell(w) \equiv \mu^{\prime }(w) \mathbf{p}_{W}(w) + \mu(w) \mathbf{p}_{W}^{\prime }(w).$$ For the derivations of the efficiency bound, it is important to recall that $\ell$ depends on $p_{W}$, so we sometimes use $\ell_{\mathbf{P}}$ to denote $\ell$. Finally, observe that under our assumptions over $\mu$ and $\mathbf{p}_{W}$, $\ell \in L^{2}(Leb)$.

The formal definition of the efficiency bound for the unknown parameter $\theta _{0}$ is given at the beginning of Appendix (ref). Loosely speaking, the efficiency bound is a lower bound for the asymptotic variance of all locally regular and asymptotically linear estimators of $\theta_{0}$; see bickeletal1998efficient for details and formal definitions. If it is infinite, then the parameter $\theta_{0}$ cannot be estimated at root-$n$ rate by these estimators. We now derive this bound. For this, we introduce some useful notation. For any $(y,w,\alpha )\in \mathbb{Y}\times \mathbb{W}\times \mathbb{A}$, let $$\rho (y,w,\alpha )\equiv \left(\rho_{1}(y,w,\alpha),\rho_{2}(y,w,h)\right)^{T} \equiv \left(\theta -\mu (w)h^{\prime}(w), 1\{y \leq h(w) \} - \tau \right)^{T}.$$ Let $\mathbf{T}:\mathbb{H}\rightarrow L^{2}(\mathbf{P}_{X})$ be given by

align*[align* omitted — 95 chars of source]

for all $x\in \mathbb{X}$ and $ g\in \mathbb{H}$. The fact that $\mathbf{T}$ maps into $L^{2}(\mathbf{P}_{X})$ follows from Jensen inequality and the fact that $\sup_{w,x} \mathbf{p}_{YW\mid X}(h_{0}(w),w\mid x)<\infty $ (see Assumption (ref)). Its adjoint operator is denoted as $\mathbf{T}^{\ast }:L^{2}(\mathbf{P}_{X})\rightarrow L^{2}(\mathbf{P}_{W})$. Finally, let \[ x\mapsto \Gamma (x)\equiv E[\rho _{1}(Y,W,\alpha _{0})\rho _{2}(Y,W,h_{0})|X=x]/(\tau (1-\tau )) \] and $z\mapsto \epsilon (z)\equiv \rho _{1}(y,w,\alpha _{0})-\Gamma (x)\rho _{2}(y,w,h _{0})$. Then $E[\epsilon (Z)\rho _{2}(Y,W,h_{0})|X]=0$ and $E[\epsilon (Z)]=0$.

theoremSuppose Assumptions (ref) and (ref) hold and $\ell \in Kernel(\mathbf{T})^{\perp}$. Then \begin{enumerate} • The efficiency bound of $\theta_{0}$ is finite iff $\ell \in Range (\mathbf{T}^{\ast})$. • If it is finite, its efficient variance $V_0$ is given by \begin{align*} V_0= ||\epsilon(\cdot) ||^{2}_{L^{2}(\mathbf{P})} + \left \Vert \mathbf{T}(\mathbf{T}^{\ast}\mathbf{T})^{+} [ \ell - \mathbf{T}^{\ast}[\Gamma] ] \right \Vert^{2}_{L^{2}(\mathbf{P})}. \end{align*} \end{enumerate}
proofSee Appendix (ref).

The first result in Theorem (ref) is obtained following the approach of bickeletal1998efficient. The condition $\ell \in Kernel(\mathbf{T})^{\perp}$ ensures that only the “identified part" of $h_{0}$ --- that is, the part of $h_{0}$ that is orthogonal to the kernel of $\mathbf{T}$ --- matters for computing the weighted average derivative; we refer the reader to Appendix (ref) and the paper by SeveriniTripathi2012 for further discussion.

SeveriniTripathi2012 provides an analogous result to Theorem (ref)(1) for linear functionals in a nonparametric linear IV regression model. Our condition $\ell \in Range(\mathbf{T}^{\ast })$, is analogous to theirs, but with a subtle yet important difference. In SeveriniTripathi2012, the object that plays the role of $\ell $ does not depend on $\mathbf{P}$, whereas in our case it does. This observation changes the nature of our condition vis-a-vis theirs, because, in our setup, $\ell \in Range(\mathbf{T}^{\ast })$ implies a restriction on $\mathbf{P}$ since both quantities, $\ell $ and $\mathbf{T}$ depend on it.\footnote{ It is worth pointing out that this restriction was not imposed as one of the conditions that defined the model used to construct the tangent space; see Appendix (ref) for a definition.} It is also important to note that, if $\mathbf{T}$ is compact, then the range of $ \mathbf{T}^{\ast }$ is a strict subset of $L^{2}(\mathbf{P}_{W})$ so that $\ell \in Range(\mathbf{T}^{\ast })$ may not hold. Hence, in this case the weighted average derivative may not be root-n estimable, and, moreover, the condition that determines the finiteness of the efficiency bound depends on unknown quantities. This observation highlights a difference with the no-endogeneity case, where the efficiency bound is always finite, provided that $\ell \in L^{2}(\mathbf{P}_{W})$ (see newey1993efficiency).

Another discrepancy between the no-endogeneity case and ours is that in the former case the “plug in" is always efficient (see newey1993efficiency, newey1994asymptotic) due to the fact that the tangent space is the whole of $\{ f \in L^{2}(\mathbf{P}) \colon E[f] = 0 \}$. On the other hand, for NPQIV CS2015overidentification show that the closure of the tangent space is the whole space iff the $Range(\mathbf{T})$ is dense in $L^{2}(\mathbf{P}_{X})$, which in turn is equivalent to $Kernel(\mathbf{T}^{\ast}) = \{ 0\}$ . This last condition is comparable to a completeness condition on the conditional distribution of the exogenous variable given the endogenous ones, which may or may not hold for a particular $\mathbf{P}$.\footnote{ In the NPIV setting, $Kernel(\mathbf{T}^{\ast}) = \{ 0\}$ is equivalent to the pdf of $X$ given $W$ satisfying a completeness condition.}

The second result in Theorem (ref) follows from projecting the influence function onto the closure of the tangent space (see bickeletal1998efficient and VdV2000 and references therein). So as to shed some light on the expression for the efficiency bound, we point out that it corresponds to the efficiency bound of the semiparametric sequential conditional moment model via the “orthogonalized moments" approach in ai2012semiparametric. In their notation, let $\varepsilon_{2}(z,\alpha) \equiv \rho_{2}(y,w,h)$ and $\varepsilon_{1}(z,\alpha) \equiv \rho_{1}(y,w,\alpha) - \Gamma(x)\rho_{2}(y,w,h)$. Note that $E[\varepsilon_{1}(Z,\alpha_{0})\varepsilon_{2}(Z,\alpha_{0}) \mid X] = 0$ (and $\varepsilon_{1}(z,\alpha_0 )=\epsilon (z)$). The model ((ref))-((ref)) becomes equivalent to their orthogonalized moment model:

equation[equation omitted — 114 chars of source]

The expression in our Theorem (ref)(2) coincides with their theorem 2.3 semiparametric efficient variance bound for $\theta_0$ of the model ((ref)). Also see proposition 3.3 in ai2012semiparametric for the semiparametric efficient variance bound for the WAD of a NPIV model.

The Penalized-Sieve-GEL Estimator

In this section we introduce our estimator for $\alpha_{0}\in \mathcal{A} \equiv \Theta \times \mathcal{H} \subseteq \mathbb{A}\equiv \Theta \times \mathbb{H}$. In order to do this, it will be useful to define some quantities. Given the i.i.d. sample $(Z_{i})_{i=1}^{n}$, let $P_{n}$ be the corresponding empirical probability. Let $(q_{k})_{k\in \mathbb{N}}$ be a complete basis in $L^{2}(\mathbb{X},Leb)$. For any $J\in \mathbb{N}$, let $q^{J}(x)=(q_{1}(x),...,q_{J}(x))^{T}$ be $J\times 1$ vector-valued function of $x$, and for any $(z,\alpha )\in \mathbb{Z}\times \mathcal{A}$, let \[ g_{J}(z,\alpha )\equiv \left (\rho _{1}(y,w,\alpha ),\rho _{2}(y,w,\alpha)q^{J}(x)^{T} \right)^{T} =\left (\theta -\mu (w)h^{\prime}(w), [ 1\{y \leq h(w) \} - \tau ]q^{J}(x)^{T} \right)^{T}~. \] Let $\mathcal{S}\subseteq \mathbb{R}$ be an open interval that contains $0$. For any $P \in \mathcal{P}(\mathbb{Z})$, any $\alpha \in \mathcal{A}$ and any $J \in \mathbb{N}$, denote $\Lambda_{J}(\alpha,P)\equiv \cap_{z \in supp (P)} \{ \lambda \in \mathbb{R}^{J+1} \colon \lambda^{T}g_{J}(z,\alpha) \in \mathcal{S} \}$, and $\hat{\Lambda}_{J}(\alpha)\equiv \Lambda_{J}(\alpha,P_{n})$.

Let $s:\mathcal{S}\rightarrow \mathbb{R}$ be strictly concave, twice-continuously differentiable with Lipschitz continuous second derivative; and $s^{\prime }(0)=s^{\prime \prime }(0)=-1$; see, e.g., Smith1997 and DIN03 for examples of such $s(.)$ functions. For any $\lambda \in \Lambda_{J}(\alpha,P)$, let

align*[align* omitted — 155 chars of source]

If $\mathcal{A}$ were a finite-dimensional compact set with $\dim (\mathcal{A}) \leq J+1$, then $\alpha_0$ could be estimated by the GEL procedure: $\arg \min_{\alpha \in \mathcal{A}} \sup_{\lambda \in \hat{\Lambda}_{J}(\alpha )}\hat{S}_{J}(\alpha ,\lambda )$ (see, e.g., DIN03).

Due to the presence of the infinite-dimensional nuisance parameter $h_0 \in \mathcal{H}$ in the NPQIV model ((ref)), the parameter space $\mathcal{A} \equiv \Theta \times \mathcal{H}$ is an infinite-dimensional function space that is typically non-compact subset in $(\mathbb{A}, ||.||)$ and hence the identifiable uniqueness condition needed for consistency in $||.||$-norm might fail; see, e.g., NP2003instrumental and chen2007large. The above GEL procedure needs to be regularized to regain consistency and/or to speed up rate of convergence in $||.||$-norm. To this end, we introduce a regularizing structure, which, jointly with $(q_{k})_{k \in \mathbb{N}}$, consists of a sequence of sieve spaces $(\mathcal{A}_k \equiv \Theta \times \mathcal{H}_{k})_{k\in \mathbb{N}}$ in $(\mathbb{A}, ||.||)$, and a sequence of penalties $(\gamma_{k}\times Pen (\cdot))_{k \in \mathbb{N}}$ with tuning parameters $\gamma_{k} \downarrow 0$ and a penalty function $Pen : \mathbb{A} \rightarrow \mathbb{R}_{+}$.

The Penalized-Sieve-GEL (PSGEL) estimator is defined as

equation*[equation* omitted — 191 chars of source]

for any $(L=(J,K),n)\in \mathbb{N}^{3}$. If the \textquotedblleft arg min" in the previous expression is empty, one can replace it by an approximate minimizer.

The following assumption imposes restrictions over the regularizing structure $\{(q_{k}, \mathcal{H}_{k},\gamma_{k}Pen)_{k \in \mathbb{N}}\}$. Let $(\varphi _{k})_{k\in \mathbb{N}}$ be a basis functions in $\mathbb{H}$, and $\nabla \varphi ^{K}=(\varphi _{1}^{\prime },...,\varphi _{K}^{\prime })^{T}$.

assumption(i) $(q_{k})_{k\in \mathbb{N}}$ is a basis in $L^{2}(\mathbf{P}_{X})$, and $E[q^{J}(X)q^{J}(X)^{T}]=I$ for each finite $J$; \newline (ii) For all $K $, $\mathcal{H}_{K}\subseteq lin\{\varphi _{1},...,\varphi _{K}\}$ is closed and convex, and $\overline{\cup _{k}\mathcal{H}_{k}}\supseteq \mathcal{H}$, i.e., for any $\alpha \in \mathcal{A} \equiv \Theta \times \mathcal{H}$ there is an $\Pi_K\alpha \in \mathcal{A}_K \equiv \Theta \times \mathcal{H}_K$ such that $||\Pi_K\alpha - \alpha ||=o(1)$; and for some finite $C\geq 1$, $C^{-1}I\leq E\left[ \left( \varphi ^{K}(W)\right) \left( \varphi ^{K}(W)\right) ^{T}+\left( \nabla \varphi ^{K}(W)\right) \left( \nabla \varphi ^{K}(W)\right) ^{T}\right] \leq CI$; \newline (iii) (a) $Pen : \mathbb{A} \rightarrow \mathbb{R}_{+}$ is lower semi-compact (in $||.||$), $|Pen (\Pi_K\alpha_0 ) - Pen (\alpha_0)|=O(1)$, $Pen (\alpha_0 )<\infty$, and $\gamma_{k} \downarrow 0$, and (b) there exists an $M<\infty $ such that for any $m\geq M$, any $K$ and any $\alpha \in \mathcal{A}_{K}$, if $Pen(\alpha )\leq m$ then $\sup_{w \in \mathbb{W}} |\mu(w) h^{\prime }(w)| \leq m$.

Condition (i) is mild (see DIN03 (DIN) and the discussion therein). Condition (ii) essentially defines the sieve space. Part (a) of Condition (iii) is standard in ill-posed problems (see CP2012estimation); Part (b) is not. If $\mathcal{H}_{K}$ is $||\cdot ||_{L^{\infty}(\mathbb{W},\mu)} $ bounded, then the condition is vacuous. If this is not the case, then the condition requires $Pen$ to be “stronger" than the $||\cdot ||_{L^{\infty}(\mathbb{W},\mu)} $ norm. The need to bound $ || h^{\prime } ||_{L^{\infty}(\mathbb{W},\mu)} $ arises from the fact that, in many instances, in the proofs we need to control $\rho(y,w,\alpha)$ uniformly on $ (y,w)$ (e.g., see Lemma (ref) in the Supplemental Material (ref)). Additionally, in our setup, is useful to link $Pen$ to $||\cdot ||_{L^{\infty}(\mathbb{W},\mu)}$ because the structure of the problem implies a natural bound for $Pen(.)$ --- and thus, through Assumption (ref)(iii), a bound for $||\cdot ||_{L^{\infty}(\mathbb{W},\mu)}$ ---, as shown in the following lemma.

lemmaFor any $L=(J,K) \in \mathbb{N}^{2}$ and any $\alpha \in \mathcal{A}_{K}$, \begin{align*} \gamma_{K} Pen(\hat{\alpha}_{L,n}) \leq \sup_{\lambda \in \hat{\Lambda} _{J}(\alpha)} \hat{S}_{J}(\alpha,\lambda) + \gamma_{K} Pen(\alpha) wpa1. \end{align*}
proofSee Appendix (ref).

The bound, however, may depend on $(J,K,n)$ and thus may affect the convergence rate. Below, we will set $\alpha$ in the right-hand-side (RHS) to a particular value in $\mathcal{A}_{K}$ and use the resulting bound to construct what we call an “effective sieve space".

Consistency and Convergence Rates of the PSGEL Estimator

This section establishes the consistency and the rates of convergence of the PSGEL estimator $\hat{\alpha}_{L,n}$ to the true parameter $\alpha_0$ under a given norm $||.||$ over $\mathbb{A}$. In this and the next section, we note that the implicit constants inside the $O_{\mathbf{P}}$ do not depend on $(J,K,n)$.

Effective sieve space

Throughout the paper we use the following notation. Let $\overline{\theta} \equiv \sup_{ \theta \in \Theta} |\theta| <\infty$; and $b_{\rho,J} \equiv (E[||q^{J}(X)||^{\rho}_{e}])^{1/\rho}$ for any $\rho > 0$. For any $L=(J,K) \in \mathbb{N}^{2}$, let

align*[align* omitted — 275 chars of source]

Let $(l_{n})_{n}$ be a slowly diverging positive sequence, e.g., $l_{n} =\log \log n$, which is introduced solely to avoid keeping track of constants. Finally we let

align*[align* omitted — 196 chars of source]

The sequence of sets, $(\bar{\mathcal{A}}_{L,n})_{L,n}$, can be viewed as the sequence of “effective" sieve spaces, because, as the following lemma shows, wpa1 the estimator (and, trivially, the sieve approximator $\Pi_K \alpha_{0}\in \mathcal{A}_{K}$) both belong to it.

assumption(i) $b^{4}_{4,J}/n = o(1)$; (ii) $\delta_{n} = o(1)$, $\delta_{n} \times \mho_{L,n} = o(1)$, $ b_{\varrho,J}^{\varrho} n \delta^{\varrho}_{n}= o(1)$ for some $\varrho > 0$; (iii) $\sqrt{ \frac{\bar{g}^{2}_{L,0}}{n} + ||E[g_{J}(Z,\Pi_K \alpha_{0})]||^{2}_{e} } = o(\delta_{n})$.
lemmaLet Assumptions (ref), (ref), (ref) and (ref) hold. Then, for any $L \in \mathbb{N}^{2}$, $\hat{\alpha} _{L,n} \in \bar{\mathcal{A}}_{L,n}$ wpa1.
proofSee Appendix (ref).

The proof of this Lemma follows from Lemma (ref) with $\alpha = \Pi_K \alpha_{0}$ and Lemma (ref) with $\alpha = \Pi_K \alpha_{0}$ and $P=P_{n}$ in Appendix (ref). The latter lemma provides a bound for $\sup_{\lambda \in \hat{\Lambda}_{J}(\Pi_K \alpha_{0})} \hat{S}_{J}(\Pi_K \alpha_{0},\lambda)$ in terms of $||n^{-1}\sum_{i=1}^{n} g_{J}(Z_{i},\Pi_K \alpha_{0}) ||^{2}_{e}$ and $ \gamma_{K} Pen(\Pi_K \alpha_{0})$. With this in mind, the components of $ \mho_{L,n}$ are intuitive: $\bar{g}^{2}_{L,0}/n$ is related to the “variance" of $n^{-1}\sum_{i=1}^{n} g_{J}(Z_{i},\Pi_K \alpha_{0})$, where $\bar{ g}^{2}_{L,0}$ is a bound for $||g_{J}(.,\Pi_K \alpha_{0})||^{2}_{e}$. The term $ ||E[g_{J}(Z,\Pi_K \alpha_{0})]||_{e}$ is related to the “bias" and reflects the fact that $\Pi_K \alpha_{0}\in \mathcal{A}_K$ is a sieve approximate to $\alpha_0$.

remarkAs explained above, Lemma (ref) and Assumption (ref)(iii) are used to ensure that $|| h^{\prime }||_{L^{\infty}( \mathbb{W},\mu)}$ is bounded. If the construction of $\mathcal{H}_{K}$ directly implies $|| h^{\prime }||_{L^{\infty}(\mathbb{W},\mu)} \leq \mho$ for some fixed constant $\mho < \infty$, then $\mho$ should replace $\mho_{L,n}$ in the definition of $\bar{\mathcal{A}}_{L,n}$. This is applicable every time $\mho_{L,n}$ appears below. $\triangle$

Relation to Penalized Sieve GMM

As expected, the asymptotic properties of the PSGEL estimator are closely related to an approximate minimizer of a GMM criterion associated to the following expression: for any $J \in \mathbb{N}$ and any $P \in \mathcal{P}(\mathbb{Z})$, let

align*[align* omitted — 133 chars of source]

where $(\alpha,P) \mapsto H_{J}(\alpha,P) \equiv E_{P}[g_{J}(Z,\alpha) g_{J}(Z,\alpha)^{T}]$. That is, $Q_{J}(.,\mathbf{P})$ is the optimally weighted (population) GMM criterion function associated with the vector of moments $E[g_{J}(Z,\cdot)]$.

For what follows, it will be useful to define the following intermediate quantity which can be viewed as a (sequence) of pseudo-true parameters. For each $L\equiv (J,K)\in \mathbb{N}^{2}$, let

equation*[equation* omitted — 109 chars of source]

We note that $\alpha _{0} \in \arg \min_{\alpha \in \mathcal{A}}Q_{J}(\alpha ,\mathbf{P})$ for any $J$; but as we restrict to the effective sieve space $\bar{\mathcal{A}}_{L,n}$, it could be that $\alpha _{L,0}\neq \alpha _{0}$ for any $L\in \mathbb{N}^{2}$. The following lemma guarantees that $\alpha _{L,0}$ is in fact non-empty.

lemmaLet Assumptions (ref) and (ref) hold. Then, for each $L=(J,K) \in \mathbb{N}^{2}$, $ \alpha_{L,0}$ is non-empty.
proofSee Appendix (ref).

While this lemma shows that $\alpha _{L,0}$ is non-empty, it may not be a singleton. Nevertheless, for model ((ref))-((ref)), it is easy to choose some finite-dimensional linear sieve $\mathcal{H}_{K}$ and some strict convex penalty $Pen$ such that $\alpha _{L,0}$ is in fact a singleton. Therefore the next assumption is effectively a way to suggest choices of a regularizing structure:

assumptionFor any $L \in \mathbb{N}^{2}$, $\alpha _{L,0}$ is single-valued.

Let $m_{2}(X,\alpha) \equiv E_{\mathbf{P}}[\rho_{2}(Y,W,\alpha) \mid X]$. For each $J \in \mathbb{N}$, the $L^{2}(\mathbf{P})$ projection of $m_{2}(\cdot, \alpha)$ onto the linear span of $q^{J}(X)$ is denoted as $Proj_{J}[m_{2}(\cdot,\alpha)](X)$, where

align*[align* omitted — 244 chars of source]

where $(E_{\mathbf{P}}[q^{J}(X)q^{J}(X)^{T}])^{-1} = I$ by Assumption (ref).

The next lemma provides sufficient conditions that ensure convergence of $\alpha_{L,0}$ to the true parameter $\alpha_{0}$.

lemmaLet Assumptions (ref), (ref), (ref) and (ref) hold. Suppose $\lim_{n \rightarrow \infty} \sup_{ \alpha \in \bar{\mathcal{A}}_{L_{n},n}}||Proj_{J_{n}}[m_{2}(\cdot,\alpha)] - m_{2}(\cdot,\alpha)||_{L^{2}(\mathbf{P})} = 0 $. Then: $||\alpha_{L_n,0}-\alpha_0||=o(1)$.
proofSee Appendix (ref).

Convergence rates

A crucial part of establishing the convergence rate of $\hat{\alpha}_{L,n}$ is to bound the rate of $||\hat{\alpha}_{L,n} - \alpha_{L,0}||$. For this it is important to quantify how well the population sieve GMM criterion function $Q_{J}$ separates points in $(\bar{\mathcal{A}}_{L,n},||.||)$ around $\alpha_{L,0}$. To do this, we define, for each, $(L,n)\in \mathbb{N}^{3}$, $\varpi_{L,n} : \mathbb{R} _{+} \rightarrow \mathbb{R}_{+}$ as

align[align omitted — 210 chars of source]

The function $ \varpi_{L,n}$ is analogous to the one used in the standard identifiable uniqueness condition (see WW1991, newey1994large). Within the ill-posed inverse literature this function is akin to the notion of sieve measure of ill-posedness used in BCK2007 and CP2012estimation,CP2015sieve. The following lemma establishes some useful properties.

lemmaLet Assumptions (ref), (ref) and (ref) hold. Then: for each $(L,n) \in \mathbb{N}^{3}$, $\varpi_{L,n}(t) = 0$ iff $t = 0$ and $\varpi_{L,n}$ is continuous and non-decreasing in $t$.
proofSee Appendix (ref).

It is worth noting that even though $\varpi_{L,n}(t) > 0$ for all $t>0$, it could happen that $\varpi_{L,n}(t) \rightarrow 0$ as $L$ diverges. This behavior reflects the ill-posed nature of the problem.

We now present some high-level assumptions used to establish the convergence rate of the PSGEL estimator. The first of these assumptions introduces, and imposes restrictions on, a positive real-valued sequence $(\delta _{n})_{n\in \mathbb{N}}$ that is common in the GEL literature (see the Appendix in DIN03). It ensures that the ball $\{ \lambda \in \mathbb{R}^{J+1} \colon ||\lambda||_{e} \leq \delta_{n} \}$ belongs to $\hat{\Lambda}_{J}(\alpha)$ for any $\alpha \in \bar{\mathcal{A}}_{L,n}$ (see Lemma (ref) in the Supplemental Material (ref)). The assumption also restricts the rates of $(b_{\rho,J})_{\rho \in \mathbb{R},J\in \mathbb{N}},$ and the rate at which $L=(J,K) \in \mathbb{N}^{2}$ diverges relative to $n$:

assumption(i) Assumption (ref) holds; (ii) $\delta_{n} l_{n} = o(1)$, $b^{3}_{3,J} \delta^{\varrho}_{n}= o(1)$ for some $\varrho > 0$, and $(\mho_{L,n})^{4} b^{4}_{4,J}/n = o(1)$.

Recall that the sequence $(l_{n})_{n}$ diverges arbitrary slowly like $\log \log n$, and the bound $\mho_{L,n}$ is allowed to grow (slowly) at the rate of $l_{n}$. Assumption (ref) slightly strengthens Assumption (ref).

The following assumption is a high-level condition that controls the supremum of the process $f \mapsto \mathbb{G}_{n}[f] \equiv n^{-1/2} \sum_{i=1}^{n} \{f(Z_{i}) -E[f(Z_{i})]\}$ over the classes $\bar{\mathcal{A}}_{L,n}$ and $\mathcal{G}_{L} \equiv \{ (y,w) \mapsto \rho_{2}(y,w,\alpha) \colon \alpha \in \bar{\mathcal{A}}_{L,n} \} $.

assumptionThere exists a positive real-valued sequence, $ (\Delta_{L,n})_{L,n \in \mathbb{N}^{3}}$, such that, for any $L \in \mathbb{N }^{2}$, $\sup_{(\theta,h) \in \bar{\mathcal{A}}_{L,n}} |\mathbb{G}_{n}[\mu \cdot h^{\prime }]| =O_{\mathbf{P}}(\Delta_{L,n})$ and for all $1 \leq j \leq J$, $ \sup_{g \in \mathcal{G}_{L}} |\mathbb{G}_{n}[g \cdot q_{j}]| =O_{\mathbf{P}}(\Delta_{L,n})$.

For instance, if $\{ \mu \cdot h^{\prime }\colon h \in \mathcal{H} \}$ and $\{ (y,w) \mapsto 1\{ y \leq h(w) \} \colon h \in \mathcal{H} \}$ are P-Donsker, then $(\Delta_{L,n})_{L,n \in \mathbb{N}^{3}}$ is uniformly bounded.\footnote{Restrictions on the “complexity" of these classes are implicit restrictions on the “complexity" of $\mathcal{H}$; see chen2003estimation and VdV2000.} But if this is not the case, then $(\Delta_{L,n})_{L,n \in \mathbb{N}^{3}}$ may diverge as $L$ (or $n$) grows.

The next theorem establishes the convergence rate of the PSGEL estimator; in particular it establishes the rate for the estimator of the infinite dimensional component $h_0 \in \mathcal{H}$.

theoremSuppose Assumptions (ref), (ref), (ref), (ref) and (ref) hold. For any $(\delta_{n},l_{n})_{n}$ satisfying Assumption (ref), there exists a finite constant $M>0$ such that \begin{align*} ||\hat{\alpha}_{L,n} - \alpha_{0}|| = O_{\mathbf{P}}\left( \varpi_{L,n}^{-1}\left( M (\delta_{1,L,n} + \delta_{2,L,n}) \right) \right) + ||\alpha_{L,0} - \alpha_{0} ||, \end{align*} where \begin{align*} \delta _{1,L,n}\equiv \sqrt{\frac{J}{n}}\times \Delta _{L,n}\times \left( \overline{\theta}+\mho _{L,n}+b_{2,J}\right) , \delta _{2,L,n}\equiv \mho _{L,n}^{2}\left\{ \delta _{n}+\delta _{n}^{-1}\Gamma _{L,n}\right\} \end{align*}
proofSee Appendix (ref).

The rate of convergence of the PSGEL estimator is composed of two standard terms reflecting the “approximation error" $||\alpha_{L,0} - \alpha_{0} ||$ and the “sampling error" $\varpi_{L,n}^{-1}\left(M(\delta_{1,L,n} + \delta_{2,L,n})\right)$. The component $\varpi^{-1}_{L,n}(.)$, reflects the ill-posed nature of the estimation problem. As noted previously, even though, for a fixed $L$, $\varpi_{L,n}(t) > 0$ for $t>0$, this relationship can deteriorate as $L$ diverges, which implies that $\varpi_{L,n}^{-1}(t)$ may diverge as $L$ diverges.

Below, we present an heuristic description of the proof that sheds light on the role of the sequences $(\delta_{1,L,n},\delta_{2,L,n})_{L,n}$ and of $\varpi_{L,n}$.

Heuristics

By the triangle inequality it suffices to bound the rate of $||\hat{\alpha}_{L,n} - \alpha_{L,0}||$. We do this by linking the PSGEL estimator to the population sieve GMM problem defined by $Q_{J}(\cdot,\mathbf{P})$. The first step to do this is to show that the PSGEL is an approximate minimizer of the sample sieve GMM criterion $Q_{J}(.,P_{n})$ with the rate given by $\delta _{2,L,n}$.

lemmaLet Assumptions (ref) and (ref) hold. For any $(\delta _{n},l_{n})_{n\in \mathbb{N}}$ satisfying Assumption (ref), we have: \begin{equation*} Q_{J}(\hat{\alpha}_{L,n},P_{n})=O_{\mathbf{P}}(\delta _{2,L,n}) , with \delta _{2,L,n} = \mho _{L,n}^{2}\left\{ \delta _{n}+\delta_{n}^{-1}\Gamma _{L,n}\right\} . \end{equation*}
proofSee Appendix (ref).

The Lemma illustrates not only the role of $\delta_{2,L,n}$ but its nature. The two terms inside the curly brackets are completely analogous to those appearing in DIN03. The scaling by $\mho^{2}_{L,n}$ is not present in DIN03 and its appearance here is due to the fact that the bound of $\rho(.,\alpha_{L,0})$ may depend, in principle, on $n$ and $L=(J,K)$. In DIN03, on the other hand, the upper bound $\mho_{L,n}$ can be taken to be a fixed constant due to their Assumption 6.

Lemma (ref) implies that, for some finite $M$, the event $Q_{J}(\hat{\alpha}_{L,n},P_{n}) - Q_{J}(\alpha_{L,0},P_{n}) \leq M \delta_{2,L,n}$ occurs wpa1. The next step is to link the empirical GMM criterion function, $Q_{J}(\cdot,P_{n})$, to its population analog, $Q_{J}(\cdot,\mathbf{P})$ for which we can quantify its behavior (around $\alpha_{L,0}$) using $\varpi_{L,n}$. The next lemma provides such a link by showing that $Q_{J}(\cdot,P_{n})$ converges to its population analog.

lemmaLet Assumptions (ref), (ref) and (ref) hold. Then: for any $L\equiv (J,K)\in \mathbb{N}^{2}$, \begin{equation*} \sup_{\alpha \in \bar{\mathcal{A}}_{L,n}}|Q_{J}(\alpha ,P_{n})-Q_{J}(\alpha,\mathbf{P})|=O_{\mathbf{P}}(\delta _{1,L,n}) , with \delta _{1,L,n}=\sqrt{\frac{J}{n}}\times \Delta _{L,n}\times \left( \overline{\theta} +\mho _{L,n}+b_{2,J}\right). \end{equation*}
proofSee Appendix (ref).

The rate $(\delta_{1,L,n})_{L,n}$ has several components. The component $ \sqrt{\frac{J}{n}}$ reflects the pointwise convergence rate of $||n^{-1} \sum_{i=1}^{n} g_{J}(Z_{i},\alpha) - E[g_{J}(Z,\alpha)] ||_{e} $, while the factor of $\Delta_{L,n}$ reflects the fact that we need uniform convergence of that term. Finally, the term $\left( \overline{\theta} + \mho_{L,n} + b_{2,J} \right)$ is essentially the (uniform) bound for $\alpha \mapsto ||n^{-1} \sum_{i=1}^{n} g_{J}(Z_{i},\alpha)||_{e}$ and $\alpha \mapsto ||E[g_{J}(Z,\alpha)] ||_{e} $ over $\bar{\mathcal{A}}_{L,n}$.

With this result at hand and simple algebra, one can show that for some finite $M$ the set $A \equiv \{Q_{J}(\hat{\alpha}_{L,n},\mathbf{P}) - Q_{J}(\alpha_{L,0},\mathbf{P}) \leq M (\delta_{1,L,n} + \delta_{2,L,n})\}$ occurs wpa1. Therefore, by standard laws of probabilities, it follows that the probability of the set $||\hat{\alpha}_{L,n} - \alpha_{0}|| \geq M^{\prime} \varpi_{L,n}^{-1}\left( M (\delta_{1,L,n} + \delta_{2,L,n}) \right)$ (for any $M^{\prime}$) is --- up to a vanishing term --- less or equal than the probability of the intersection of the same set with $A$. Therefore, it only remains to show that the latter probability is naught for sufficiently large $M^{\prime}$. This follows because this latter probability is in turn bounded above by the probability of $\varpi_{L,n} \left( M^{\prime} \varpi_{L,n}^{-1}\left( M (\delta_{1,L,n} + \delta_{2,L,n}) \right) \right) \leq M (\delta_{1,L,n} + \delta_{2,L,n})$. By the fact that $\varpi_{L,n}$ is non-decreasing (see Lemma (ref)), this probability is naught by sufficiently large $M^{\prime}$, proving the result of Theorem (ref).

Discussion of the elements in the Convergence Rate

We now present some observations regarding the main components of the convergence rate in Theorem (ref), namely, the rates $(\delta_{1,L,n},\delta_{2,L,n})_{L,n}$ and $\varpi_{L,n}$ defined in expression (ref). Regarding the latter, we first need to specify the norm $||.||$. We start by taking $(\varphi_{k})_{k \in \mathbb{N}}$ to be an orthogonal basis with respect to the Lebesgue measure over $\mathbb{H}$. Thus, for any $\alpha = (\theta,h) \in \mathbb{A}$, there exists a real-valued sequence, $(\pi_{l})_{l=0}^{\infty}$, such that $\alpha = (\theta,h) = (\pi_{0} \bar{\varphi}_{0}, \sum_{l=1}^{\infty} \pi_{l} \bar{\varphi}_{l})$ where $\bar{\varphi}_{0} = 1$ and $\pi_{0} = \theta$, and, for any $k \geq 1$, $\bar{\varphi}_{k} = \varphi_{k}$ and $\pi_{k}$ is the “Fourier" coefficient of $h$ with respect the basis $(\varphi_{k})_{k \in \mathbb{N}}$. This representation gives rise to the following norm over $\mathbb{A}$, $\alpha \mapsto \sqrt{ \sum_{l=0}^{\infty} \pi^{2}_{l} } = \sqrt{\theta^{2} + ||h||^{2}_{L^{2}(Leb)}}$. The aforementioned norm presents itself as a “natural" norm under which we can establish convergence rate and thus we set $||.||$ as this norm; our result can be extended to norms other than this by specifying how the desired norm relates to $\alpha \mapsto \sqrt{|\theta|^{2} + ||h||^{2}_{L^{2}(Leb)}}$.

We now shed light on the behavior of $\varpi_{L,n}$ under our choice of $||.||$. In particular, we will illustrate how this function is linked to the curvature of the criterion function $\alpha \mapsto \bar{Q}_{J}(\alpha,\mathbf{P})$. To do this, it is convenient to use local approximations, so we take, for each $L=(K,L) \in \mathbb{N}^{2}$, $\mathcal{A}_{K}$ to be convex, $\alpha_{L,0}$ to be such that $ARC_{K}(\alpha_{L,0}) \equiv \{ \alpha_{L,0} + t \zeta \colon \zeta \in \mathcal{A}_{K}\setminus\{ \alpha_{L,0} \} ~and~t\in [0,1] \} \subseteq \mathcal{A}_{K}$, and require that $Pen$ to be convex and twice continuously differentiable. By the mean value theorem and the fact that $\alpha_{L,0}$ is a minimizer --- and thus satisfies that $\frac{d\bar{Q}_{J}(\alpha_{L,0},\mathbf{P})}{d\alpha}[\cdot]=0$ ---, it follows that for any $\alpha \in \bar{\mathcal{A}}_{L,n}$,

align*[align* omitted — 241 chars of source]

By the sieve representation discussed above, the RHS in this expression can be cast as

align*[align* omitted — 173 chars of source]

where $\pi^{K+1}$ denotes the first $K+1$ coefficients of the representation of $\alpha$; $\pi^{K+1}_{L,0}$ is the same but for $\alpha_{L,0}$, and $\mathcal{I}_{L}$ is a $(K+1) \times (K+1)$ matrix where the $(i,j)$-th component is given by

align*[align* omitted — 187 chars of source]

This result implies that $\varpi_{L,n}(t) \geq t^{2} e_{min}(\mathcal{I}_{L} )$ ($e_{min}(A)$ is the minimal eigenvalue of the matrix $A$). If $e_{min}(\mathcal{I}_{L})>0$, then

align*[align* omitted — 192 chars of source]

The scaling factor $(e_{min}(\mathcal{I}_{L} ))^{-1/2}$ summarizes the ill-posed nature of the problem, because, even though we require $e_{min}(\mathcal{I}_{L} )>0$ for each $L$, we do not impose this restriction uniformly on $L$, i.e., we allow that $e_{min}(\mathcal{I}_{L} ) \rightarrow 0$ as $L \rightarrow \infty$.\footnote{The condition $e_{min}(\mathcal{I}_{L} )>0$ for each $L$ essentially ensures that the “identifiable uniqueness" condition (ref) holds for each $L$; this requirement is common in the ill-posed inverse literature (e.g., chen2007large).} The speed at which this occurs depends on the local curvature of $\bar{Q}_{J}(\cdot, \mathbf{P})$ (at $\alpha_{L,0}$) and the growth of $\mathcal{A}_{K}$; see BCK2007 and CP2012estimation for a more thorough discussion.

We next discuss the rate components $(\delta_{1,L,n},\delta_{2,L,n})_{L,n}$. As mentioned in Remark (ref) above, if $\sup_{h \in \mathcal{H}} ||h^{\prime}||_{L^{\infty}(\mathbb{W},\mu)}$ is finite, then $\mho_{L,n}$ can be replaced by a fixed constant $\mho$; this fact and some algebra implies that $\delta_{2,L,n} = O(\delta_{n} + \delta^{-1}_{n} (b^{2}_{2,J}/n + ||E[g_{J}(Z,\Pi_{K}\alpha_{0})]||^{2}_{e} + \gamma_{K} Pen(\Pi_{K}\alpha_{0}) ) )$, where $\Pi_{K} \alpha_{0}$ is the projection of $\alpha_{0}$ onto $\mathcal{A}_{K}$ (see Lemma (ref) in the Supplemental Material). By taking $\delta_{n}$ to balance both terms, it follows $ \delta_{2,L,n} = O\left(\sqrt{b^{2}_{2,J}/n + \{||E[g_{J}(Z,\Pi_{K}\alpha_{0})]||^{2}_{e} + \gamma_{K} Pen(\Pi_{K}\alpha_{0})\}}\right)$. Ignoring the term inside the curly brackets, this is the same rate than the one obtained by DIN in Lemma A.14 (note that in their setup $b^{2}_{2,J} \leq J$); the additional term inside the curly brackets stems from the fact that our estimation problem needs to be regularized and consequently $\alpha_{L,0}$ is not the true parameter that nullifies the moments.

The sequence $(\delta_{1,L,n})_{L,n}$ is somewhat more standard within the semi-/non-parametric literature, e.g. chen2007large, and its components essentially impose restrictions on the “complexity" of $\mathcal{H}$. For instance, if $\mathcal{A} = \Theta \times \mathcal{H}$ is such that the classes $\{ \rho_{2}(\cdot,\cdot,\alpha) \colon \alpha \in \mathcal{A} \}$ and $\{ \mu h ^{\prime} \colon h \in \mathcal{H} \}$ are P-Donsker, then $\Delta_{L,n} = O(1)$ and $\delta_{1,L,n} = O \left( \sqrt{ \frac{J}{n}}\times b_{2,J} \right)$.

To further simplify the expression, suppose $b^{\rho}_{\rho,J} \leq J^{\rho/2}$; cf. Assumption 2 in DIN03 (see that paper for details and further references). Thus, under these conditions, the result in Theorem (ref) simplifies to {{

align*[align* omitted — 277 chars of source]

}} a rate governed by the degree of ill-posedness, the number $J$ of moment functions, the number $K$ of series terms and the bias arising from $\alpha_{L,0}$.

Asymptotic Distribution Theory

We now define the LR-type test statistic for the null hypothesis $\theta_{0} = \nu$. For any $\nu \in \Theta$ and any $(L,n) \equiv (J,K,n) \in \mathbb{N}^{3}$, let

align*[align* omitted — 387 chars of source]

The goal of this section is to show that this statistic is asymptotically chi-square distributed with one degree of freedom. The proof of this result relies on a local quadratic approximation of the criterion function $\hat{S} _{J}$ and a representation for the parameter of interest. To derive these results, we define the following quantities: For any $\alpha \in \mathcal{A} _{K}$ and any $(\theta, \zeta ) \in \mathbb{A}$, let

align*[align* omitted — 309 chars of source]

where $\mathbf{0}$ is a $J \times 1$ vector of zeros. By assumption (ref) these quantities are well-defined.

For any $L\in \mathbb{N}^{2}$ and for any $(\theta, \zeta ) \in \mathbb{A}$, we define another norm over $\mathbb{A}$ as,

equation*[equation* omitted — 140 chars of source]

where $H_{L}\equiv H_{J}(\alpha _{L,0},\mathbf{P})$. This norm acts as the so-called \textquotedblleft weak norm" in ai2003efficient,ai2007estimation.

Alternative Representation for the Weighted Average Derivative

Lemma (ref) in Appendix (ref) shows that, over $lin\{\mathcal{A}_{K}\}$ for any $L=(J,K)\in \mathbb{N}^{2}$, $||\alpha ||_{w}=0$ iff $\alpha =0$. This fact implies that linear functionals are always bounded in the space $(lin\{\mathcal{A} _{K}\},||.||_{w})$. Since $\theta $ can be interpreted as a linear functional of $\alpha $, the following representation for $\theta $ holds: For all $L=(J,K)\in \mathbb{N}^{2}$, there exists a $v_{L,n}^{\ast }\in \mathcal{A}_{K}$ such that for any $\alpha =(\theta ,h)\in \mathcal{A}_{K}$,

equation*[equation* omitted — 180 chars of source]

We note that, even though for each fixed $L \in \mathbb{N}^{2}$, $ ||v^{\ast}_{L,n}||_{w} < \infty$, this quantity may diverge as $L$ diverges if $\theta$ is not root-n estimable. Hence, we scale $v^{\ast}_{L,n}$ by its norm, and define $u^{\ast}_{L,n} \equiv v^{\ast}_{L,n}/||v^{\ast}_{L,n}||_{w} $. Then \[ \frac{\hat{\theta}_{L,n} - \theta_{L,0}}{||v^{\ast}_{L,n}||_{w}}=\langle u^{\ast}_{L,n} , \hat{\alpha}_{L,n} - \alpha_{L,0} \rangle_{w}~. \]

remark[On the relationship between the Riesz representer and the Efficiency bound] The weak norm of the Riesz representer, $||v_{L,n}^{\ast }||_{w}$, is the efficiency bound of $\theta_{0}$ in a model with $J+1$ unconditional moments functions, $E_{P}[g_{J}(Z,\cdot )]$, and $K+1$ parameters (which define $\alpha \in \mathcal{A}_{K}$).\footnote{ai2003efficient,ai2012semiparametric established this claim for a richer model with conditional moments and infinite dimensional parameters.} For a suitably chosen sequence $L\equiv L(n)$ that increases as $n$ does --- since $(q_{j})_{j}$ is dense in $L^{2}(\mathbb{X},Leb)$ --- one expects the sequence of unconditional moment functions to approximate the moments ((ref))-((ref)) defining the model. Thus, by the results in chamberlain1987 (see also lemma 3.3, lemma 4.1 and appendix A.1 in CP2015sieve) one expects $(||v_{L(n),n}^{\ast}||_{w})_{n}$ to converge to the efficiency bound presented in Theorem (ref) provided it is finite. If the efficiency bound is infinite, the sequence $(||v_{L(n),n}^{\ast}||_{w})_{n}$ will diverge; this fact reflects the non root-n estimability of the weighted average derivative within the original model ((ref))-((ref)). $\triangle$

The Asymptotic distributions of $\hat{\theta}_{L,n}$ and LR statistic

For any positive real-valued sequences $(\eta_{L,n},\eta_{w,L,n})_{L,n \in \mathbb{N}^{3}}$ (they will be restricted below) and any $(L,n) \in \mathbb{N} ^{3}$, let

align*[align* omitted — 183 chars of source]

In what follows, for any $(L,n) \in \mathbb{N}^{3}$, let $\hat{\alpha} ^{\nu}_{L,n}$ be the argument that minimizes the restricted criterion function, i.e., $\hat{\alpha}^{\nu}_{L,n} \in \arg\min_{\{\alpha \in \mathcal{A}_{K} \colon \theta = \nu\}} \sup_{\lambda \in \hat{\Lambda} _{J}(\alpha)} \hat{S}_{J}(\alpha,\lambda) $. We impose the following assumption that restricts the convergence rate of the unrestricted and restricted PSGEL estimators.

assumptionFor any $L \in \mathbb{N}^{2}$ and $\alpha \in \{ \hat{\alpha}_{L,n} , \hat{ \alpha}^{\nu}_{L,n} \}$, if $\nu=\theta_{0}$: (i) $\alpha \in int(\mathcal{N} _{L,n})$; (ii) $\gamma_{K}\sup_{t \colon |t| \leq l_{n} n^{-1/2}} |Pen(\alpha) - Pen(\alpha + t u^{\ast}_{L,n})| = o_{\mathbf{P}}(n^{-1})$; (iii) There exists a $C< \infty$ such that for any $h \in \mathbb{H}$, $||h||_{L^{2}(Leb)}\leq C ||(0,h)||$.

Part (i) of this assumption ensure that both estimators --- the restricted and unrestricted ones --- converge to $\alpha_{L,0}$ faster than $\eta_{L,n}$ and $\eta_{w,L,n}$ in the respective norms. One can use the results in Section (ref) to verify this assumption.\footnote{ The results in Section (ref) apply to the restricted estimator, under the null, with minimal changes.} Part (ii) ensures that the penalty term is negligible (see also CP2015sieve). Finally part (iii) states a relationship between the norm $h \mapsto ||(0,h)||$ --- used in Section (ref) --- and the $L^{2}(Leb)$ norm over $\mathbb{H}$.

In the following assumption we let $\bar{\mathcal{G}}_{L,n}\equiv \{f(.,\alpha)=\rho _{2}(.,.,\alpha )-\rho _{2}(.,.,\alpha _{L,0})\colon \alpha \in \mathcal{N}_{L,n}\}$.

assumptionThere exists positive sequence, $(\Delta _{2,L,n})_{L,n\in \mathbb{N}^{3}}$, such that, for any $L=(J,K)\in \mathbb{N} ^{2}$, $\sup_{\alpha=(\theta,h) \in \mathcal{N}_{L,n}}\mathbb{G}_{n}[\mu \cdot (h^{\prime }-h_{L,0}^{\prime })]=O_{\mathbf{P}}(\Delta _{2,L,n})$ and for all $1\leq j\leq J$ , $\sup_{f\in \bar{\mathcal{G}}_{L,n}}\mathbb{G}_{n}[f\cdot q_{j}]=O_{\mathbf{P}}(\Delta _{2,L,n})$.

This is a high-level assumption that controls one of the terms in the remainder of the quadratic approximation in Lemma (ref) below. As $ \mathcal{N}_{L,n}$ is shrinking, one would expect $\Delta _{2,L,n}=o(1)$; the exact rate, however, depends on the complexity of $\bar{\mathcal{A}}_{L,n}$.

assumptionThere exists a positive real-valued sequence, $ (\Xi_{L,n})_{L,n \in \mathbb{N}^{3}}$, such that, for any $L=(J,K) \in \mathbb{N}^{2}$, $\sup_{\alpha \in \mathcal{N}_{L,n}} \left \Vert H_{J}(\alpha,P_{n}) - H_{J}(\alpha_{L,0},P_{n}) - \{ H_{J}(\alpha,\mathbf{P}) - H_{J}(\alpha_{L,0},\mathbf{P}) \} \right \Vert_{e} = O_{\mathbf{P}}(\Xi_{L,n})$.

This high-level assumption implies stochastic equi-continuity of the process $H_{J}(\cdot ,P_{n})$, and it is used to control one of the terms in the remainder of the quadratic approximation in Lemma (ref) below.

The final two assumptions impose additional restrictions on $(\eta _{L,n},\eta _{w,L,n})_{L,n\in \mathbb{N}^{3}}$, $(b_{\rho ,J})_{\rho \in \mathbb{R},J\in \mathbb{N}}$, $(\delta _{n})_{n\in \mathbb{N}}$ and the rate at which $L=(J,K)$ diverges relative to $n$.

assumption(i) $\frac{\sqrt{n}}{||v^{\ast}_{L,n}||_{w}} ||E[g_{J}(Z,\alpha_{L,0})]||_{e} = o(1)$; (ii) $\frac{\sqrt{n}}{ ||v^{\ast}_{L,n}||_{w}} |\theta_{L,0}-\theta_{0}|= o(1)$.

This assumption implies that the “bias" terms arising from working with $ \alpha_{L,0}$, as opposed to $\alpha_{0}$, are small relative to the rate we are using to scale the leading term of the asymptotic expansions below $ \frac{\sqrt{n}}{||v^{\ast}_{L,n}||_{w}}$. Similar assumptions have been imposed in the literature, e.g. CP2015sieve and reference therein.

assumption(i) $n \delta^{3}_{n} ( \mho_{L,n} + b_{3,J} )^{3} = o(1)$, $n\delta^{2}_{n} \left( \{ \theta_{L,0}^{2} + || h^{\prime }_{L,0}||^{2}_{L^{\infty}(\mathbb{W},\mu)} \} \sqrt{\frac{b_{4,J}}{n}} + \mho_{L,n} \eta_{L,n} + \Xi_{L,n} \right) = o(1)$ and $n\delta_{n} \left( \sqrt{\frac{J}{n}}\Delta_{2,L,n} + \eta_{L,n}^{2} b_{2,J} \right) = o(1)$; (ii) $\sqrt{ \bar{g}^{2}_{L,0}/n + ||E[g_{J}(Z,\alpha_{L,0})]||^{2}_{e} } + \eta_{w,L,n} =o(\delta_{n})$ and $\left( \sqrt{\frac{J}{n}}\Delta_{2,L,n} + \eta_{L,n}^{2} b_{2,J} \right) = o(\delta_{n})$; (iii) There exists a $ \varrho>0$ such that $||h^{\prime }_{L,0}||^{2+\varrho}_{L^{\infty}(\mathbb{W},\mu)}/n^{2+\varrho} = o(1)$ and $b^{2+\varrho}_{2+\varrho,J}/n^{2+ \varrho} = o(1)$; (iv) $A_{L,0} \equiv E[\mathbf{p}_{Y|WX}(h_{L,0}(W)\mid W,X) q^{J}(X)\varphi^{K}(W)^{T} ]$ has full rank $K$ and $n^{-1/2} e_{min}(A_{L,0}^{T}A_{L,0})^{-1} = o(\eta_{L,n})$; (v) $||h^{\prime }_{L,0}||^{2}_{L^{\infty}(\mathbb{W}, \mu)} b^{2}_{4,J}/\sqrt{n} = o(1)$.

Part (i) ensures that the remainder term for the asymptotic quadratic representation of $\hat{S}_{J}$ is negligible (see Lemma (ref)). The sequence $(\delta _{n})_{n}$ in part (ii) was discussed after Assumption (ref). Part (iii) is used to show asymptotic normality of the leading term in Lemma (ref) by means of a Lyapounov condition. Finally, part (iv) ensures that the weak norm is proportional to the strong norm over $\mathcal{A}_{K}$ (even though the constant of proportionality may vanish as $L$ diverges) and that deviations of the form $\alpha +l_{n}n^{-1/2}u_{L,n}^{\ast }$ stay in $\mathcal{N}_{L,n}$ (see Lemma (ref) in Appendix (ref)). These deviations play a crucial role in the proof of Lemma (ref).

remark[The rate restrictions of Assumption (ref)] While parts (iii)-(v) are fairly easy to check and interpret, parts (i)-(ii) are not as easy. The goal of this remark is to illustrate the restrictions imposed by these parts on the different rates $(\delta_{n},\eta_{L_{n},n},\eta_{w,L_{n},n},\Delta_{2,L_{n},n},\Xi_{L_{n},n})_{n}$ where $(L_{n})_{n}$ is a diverging sequence in $\mathbb{N}^{2}$. To do this, we take as the point of departure the setting described in Section (ref), which allows us to simplify some expressions. Under this setup, part (i) imposes $\delta_{n} = o\left( n^{-1/3} J^{-1/2}_{n} \right)$. Given this, the restrictions in parts (i)-(ii) imply that $\eta_{L_{n},n} = O(\min\{ n^{-1/2}\delta^{-1/2}_{n}J^{-1/2}_{n}, n^{-1/6} J^{-3/4}_{n}\})$ and $\eta_{w,L_{n},n} = o(n^{-1/3} J^{-1/2}_{n})$; we note that by imposing a polynomial rate of decay, this condition rules out the so-called severely ill-posed case wherein the rate of for $(\eta_{L_{n},n})_{n}$ decays slower than polynomial order (see CP2012estimation and references therein). Parts (i)-(ii) also imply that $\Delta_{2,L,n} = o( (\sqrt{nJ_{n}}\delta^{2}_{n})^{-1} )$ and $\Xi_{L,n} = o(n^{-1} \delta_{n}^{-2})$; for the “worst case" where $\delta_{n} = \left( n^{-1/3} J^{-1/2}_{n} \right)/l_{n}$, it follows that $\Delta_{2,L,n} = O( J_{n} n^{-1/6} )$ and $\Xi_{L,n} = O(n^{-1/3} J_{n} )$, but the restriction can be relaxed if $(\delta_{n})_{n}$ decays faster. Finally, parts (i)-(ii) impose restrictions on the growth of $(L_{n})_{n}$: $J_{n} = O(n^{-1/6})$ and $\sqrt{J_{n}||E[g_{J_{n}}(Z,\alpha_{L_{n},0})]||^{2}_{e}} = o(n^{-1/3})$. $\triangle$

The following result characterizes the asymptotic distribution of the LR test statistic under the null. This characterization holds regardless of whether the parameter $\theta_{0}$ is root-$n$ estimable or not.

theoremLet Assumptions (ref)-(ref) and (ref)-(ref) hold. Then, under the null $\theta_{0} = \nu$, \begin{align*} \hat{\mathcal{L}}_{L,n}(\theta_{0}) \Rightarrow \chi^{2}_{1}. \end{align*}
proofSee Appendix (ref).

This result extends those in ParenteSmith2011gel to a non-parametric setup where the GEL is constructed using an increasing number of moment conditions, and wherein the parameter of interest may not be root-n estimable. Using a related estimator ---an EL-based on conditional moments a la kitamura2004empirical --- tao2013empirical derived an analogous result but her assumptions rule out non-smooth residuals, relevant for the quantile IV model considered here.

As a by-product of the derivations used to prove Theorem (ref), an asymptotic linear representation for the estimator of the WAD is obtained.

theoremLet Assumptions (ref)-(ref), (ref) (for $\hat{\alpha}_{L,n}$), (ref), (ref) and (ref) hold. Then \begin{align*} \frac{\hat{\theta}_{L,n} - \theta_{L,0}}{||v^{\ast}_{L,n}||_{w}}= n^{-1} \sum_{i=1}^{n} (G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0}) + o_{\mathbf{P}}(n^{-1/2}). \end{align*} Further, under Assumption (ref), we have \[ \frac{\sqrt{n}(\hat{\theta}_{L,n} - \theta_{0})}{||v^{\ast}_{L,n}||_{w}}\Rightarrow N(0,1)~. \]

The proof is the same as that Lemma (ref) in Appendix (ref) so it is omitted. This result illustrates the role of $||v^{\ast}_{L,n}||_{w}$ as the appropriate scaling of our estimator. If the sequence $(||v^{\ast}_{L,n}||_{w})_{n}$ is uniformly bounded, then this theorem implies that $\hat{\theta}_{L,n}$ is $\sqrt{n}$ asymptotically Gaussian. On the other hand, if the sequence diverges, Gaussianity is still preserve but the rate is slower and given by $\sqrt{n}/||v^{\ast}_{L,n}||_{w}$.

Heuristics

The idea is to show that, asymptotically, $\hat{\mathcal{L}}_{L,n}$ is a quadratic form of Gaussian random variables. The first step is to provide a quadratic approximation for the criterion function $\hat{S}_{J}(\alpha,\cdot)$ as a function of $\lambda $, as shown in the following lemma.

lemmaLet Assumptions (ref)-(ref), (ref), (ref) and (ref)(v) hold. Then uniformly over $(\alpha,\lambda) \in \mathcal{N}_{L,n} \times \{ \lambda \in \mathbb{R}^{J+1} \colon ||\lambda||_{e} \leq \delta_{n} \}$, for any $L=(J,K) \in \mathbb{N}^{2}$ \begin{align*} \hat{S}_{J}(\alpha,\lambda) =& - \lambda^{T} \varDelta(\alpha) - \frac{1}{2} \lambda^{T} H_{L} \lambda \\ & + O_{\mathbf{P}} \left( \delta^{3}_{n} ( \overline{\theta} + l_{n}\gamma_{K}^{-1} \Gamma_{L,n} + b_{3,J} )^{3} \right) \\ & + O_{\mathbf{P}} \left( \delta^{2}_{n} \left( ( \overline{\theta} + ||h^{\prime }_{L,0}||_{L^{\infty}(\mathbb{W},\mu)} )^{2} \sqrt{b_{4,J}/n} + \mho_{L,n} \eta_{L,n} + \Xi_{L,n} \right) \right) \\ & + O_{\mathbf{P}} \left( \delta_{n} \left( \sqrt{\frac{J}{n}}\Delta_{2,L,n} + \eta_{L,n}^{2} b_{2,J} \right) \right). \end{align*} where $\varDelta(\alpha) \equiv n^{-1} \sum_{i=1}^{n} g_{J}(Z_{i},\alpha_{L,0}) + G(\alpha_{L,0})[\alpha - \alpha_{L,0}]$.
proofSee Appendix (ref).

The “remainder" terms in the RHS (the $O_{\mathbf{P}}(.)$ terms) are fairly intuitive: the order $\delta _{n}^{3}$ -term requires boundedness of the third derivative of $\hat{S}_{J}(\alpha ,\cdot )$; the $\delta _{n}^{2}$-term arises because the expansion yields a quadratic term with $H_{J}(\alpha ,P_{n})$ as opposed to $H_{L}$; and the $ \delta _{n}$-term is the error of approximating $n^{-1} \sum_{i=1}^{n}g_{J}(Z_{i},\alpha )$ with $\varDelta(\alpha )$. This last part handles the non-smooth nature of the residuals $\rho _{2}$ by using $ E[g_{J}(Z,\cdot )]$, which is a smooth function. Assumption (ref)(i) ensures that these `remainder" terms are in fact $o_{\mathbf{P}}(n^{-1})$. This fact, and the fact that $\hat{\Lambda} _{J}(\alpha)$ contains a $\delta_{n}$-ball (see Lemma (ref) in the Supplemental Material (ref)), imply that the expression in the Lemma provides an asymptotic characterization for $\sup_{\lambda \in \hat{\Lambda}_{J}(\alpha)} \hat{S}_{J}(\alpha,\lambda)$ in terms of $(\varDelta(\alpha))^{T}H^{-1}_{L}(\varDelta(\alpha))$, which is a quadratic form in $\alpha$.

With this result at hand and Assumption (ref), one can obtain lower and upper bounds for $\hat{\mathcal{L}}_{L,n}(\theta_{0})$ of the form,

align*[align* omitted — 270 chars of source]

for appropriately chosen $t \in \mathbb{R}$, and

align*[align* omitted — 324 chars of source]

for appropriately chosen $t \in \mathbb{R}$. Since $\alpha \mapsto \varDelta(\alpha)$ is an affine function, the RHS in the previous expression is fairly easy to characterize. The following lemma formalizes these steps (its proof presents the explicitly choice for $t$ in the previous two displays).

lemmaLet Assumptions (ref)-(ref), (ref)-(ref) hold. Then, under the null $\nu = \theta_{0}$, \begin{align*} &\hat{\mathcal{L}}_{L,n}(\theta_{0}) - \left( n^{-1/2} \sum_{i=1}^{n} (G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0}) \right)^{2} \\ & \geq 2\sqrt{n}\frac{(\theta_{0} - \theta_{L,0})}{||v^{\ast}_{L,n}||_{w}} \left( n^{-1/2} \sum_{i=1}^{n} (G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0}) \right)+ o_{\mathbf{P}}(1). \end{align*} and \begin{align*} &\hat{\mathcal{L}}_{L,n}(\theta_{0}) - \left( n^{-1/2} \sum_{i=1}^{n} (G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0}) \right)^{2} \\ & \leq 2\sqrt{n}\frac{(\theta_{0} - \theta_{L,0})}{||v^{\ast}_{L,n}||_{w}} \left( n^{-1/2} \sum_{i=1}^{n} (G(\alpha_{L,0})[u^{\ast}_{L,n}])^{T}H_{L}^{-1} g_{J}(Z_{i},\alpha_{L,0}) \right) \\ & + \left( \sqrt{n}\frac{(\theta_{0} - \theta_{L,0})}{||v^{\ast}_{L,n}||_{w}} \right)^{2} + o_{\mathbf{P}}(1). \end{align*}
proofSee Appendix (ref).

This lemma shows the reason for Assumption (ref) in our analysis, as this assumption ensures that

align*[align* omitted — 189 chars of source]

Under mild assumptions and Assumption (ref), the object inside the parenthesis is asymptotically Normal with mean 0 and variance 1. Here we see the importance of the \textquotedblleft optimal weight", $ H_{L}^{-1}$. If $H_{L}$ differed from $E[g_{J}(Z,\alpha _{L,0})g_{J}(Z,\alpha _{L,0})^{T}]$, then the variance of the term inside the parenthesis will not be equal to 1, and the test statistic will only be proportional to a $\chi _{1}^{2}$ in the limit; see CP2015sieve for a more thorough discussion and results for this case.

Conclusion

Since the seminal work by Koenker and Bassett about 40 years ago (koenker1978regression), quantile regression models have become ubiquitous in econometrics and statistics; see Koenker2018 for a recent survey. The original linear quantile regression model has been extended in several directions; in particular to the general non-parametric IV framework that allows for “flexible functional forms" and endogeneity of the regressors. This type of model, while very general, presents technical challenges arising from the non-smooth nature of the criterion function as well as its ill-posedness. One goal of this paper is to shed some light on how the nonlinear ill-posedness of the non-parametric quantile IV (NPQIV) model affects not only the speed of convergence to the conditional quantile function but also the accuracy for estimating even simple linear functionals. For this, we derive the semiparametric efficiency bound for a particular linear functional of the NPQIV --- the weighted average derivative (WAD).

To estimate the parameters of interest --- the NPQIV function and its WAD --- we propose a general penalized sieve GEL procedure based on the unconditional WAD moment restriction and an increasing number of unconditional moments that are asymptotically equivalent to the conditional moment defining the NPQIV model ((ref)). We show that the QLR statistic based on the penalized sieve GEL is asymptotically chi-square distributed regardless of whether or not the information bound of the WAD is singular. This result can be used to construct confidence sets for the WAD without the need to estimate the variance of the estimator of the WAD. We hope these results extend even further the scope of quantile regression models.

The penalized sieve GEL procedure is more generally applicable to any semi/nonparametric conditional moment restrictions and unconditional moment restrictions, say of the following form:

align[align omitted — 237 chars of source]

Here $Y$ denotes dependent (or endogenous) variables, $X$ denotes conditioning (or instrumental) variables and $W$ could be either endogenous or subset of $X$, $\theta =(\theta _{1}^{\prime },\theta _{2}^{\prime })^{\prime }$ denotes a vector of finite dimensional parameters, and $h(\cdot )=\left( h_{1}(\cdot ),...,h_{q}(\cdot )\right) $ a $q\times 1$ vector of real-valued measurable functions of $Y$, $W$, $X$ and other unknown parameters. The residual functions $\rho _{j}(y,w;\theta ,h(\cdot ))$ , $j=1,2$, could be nonlinear, pointwise non-smooth with respect to $(\theta ,h)$. And some of the $\theta$ could have singular information bound. This is a valuable alternative to classical semiparametric two-step GMM when the second step finite dimensional parameter $\theta$ might not be root- $n$ estimable.

In an old unpublished draft, CP2010 study the asymptotic properties of another estimation procedure, optimally weighted penalized Sieve Minimum Distance (SMD) based on orthogonalized residuals for model ((ref))-( (ref)). Under a set of regularity conditions, including the assumption that the WAD of a NPQIV has a positive information bound, CP2010 establish that their optimally weighted penalized SMD estimator of the WAD is root-$n$ asymptotically normal and semiparametrically efficient. It would be interesting to compare this paper's estimator against theirs, and we leave this to future work.