EconBase
← Back to paper

Towards a General Large Sample Theory for Regularized Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,633 characters · 12 sections · 83 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Towards a General Large Sample Theory for Regularized Estimators

abstractWe present a general framework for studying regularized estimators; such estimators are pervasive in estimation problems wherein “plug-in" type estimators are either ill-defined or ill-behaved. Within this framework, we derive, under primitive conditions, consistency and a generalization of the asymptotic linearity property. We also provide data-driven methods for choosing tuning parameters that, under some conditions, achieve the aforementioned properties. We illustrate the scope of our approach by presenting a wide range of applications.

\startcontents[section1] \printcontents[section1]{1}{

Table of Contents

}

{2.5pt} {2.5pt}

Introduction

It was noted as early as Stein Stein1956 that in many complex models, the parameter mapping, $\psi$, that links the probability distribution generating the data, $P$, to some parameter space may be ill-behaved or even ill-defined when evaluated at the empirical distribution. The widespread solution in these cases is to regularize the problem. Regularization procedures are ubiquitous in statistics and elsewhere, examples of these include kernel-based estimators; series-based estimators and penalization-based estimators among many others.\footnote{Examples of regularizations are so ubiquitous that providing a thorough review is outside the scope of the paper; see e.g. BickelLi06, buhlmann2011statistics, HardleLinton1994 and Chen2007 for excellent reviews of several regularization methods.} Even though there has been an enormous amount of work in statistics and other sciences studying the properties of these procedures, they are viewed, by and large, as separate and unrelated. In particular, results like consistency or large sample distribution theory, when they exists, they have only been derived in a case-by-case basis; to our knowledge, there is no general theory or systematic approach. The goal of this paper is to fill this gap by providing the basis for an unifying large sample theory for regularized estimators that will allow us to make systematic progress in studying their large sample properties.

Our point of departure is the general conceptual framework put forward by Bickel and Li (BickelLi06), wherein the authors propose a general definition of regularization. According to their framework, a regularization can be viewed sequence of parameter mappings, $(\psi_{k})_{k=1}^{\infty}$ that replaces the original parameter mapping, $\psi$, each element is well-behaved, and its limit coincides with the original mapping. The index of this sequence (denoted by $k$) represents what is often referred as the tuning (or regularization) parameter; e.g. it is the (inverse of the) bandwidth for kernels, the number of terms in a series expansion, or the (inverse of the) scale parameter in penalizations. While Bickel and Li's framework encompasses many examples and applications, it is unclear what type of asymptotic properties can be obtained in such a general framework. We provide two set of general theorems under intuitive conditions that establish large sample properties for regularized estimators. One set of results establishes consistency and rate of convergence, and a data-driven method for choosing the tuning parameter that achieves these rates. Another set of results provide foundations for large sample distribution theory by deriving a generalization of the classical asymptotic linearity property.

Our approach to obtain consistency and convergence rate results is akin to the one used in the standard large sample theory for “plug-in" estimators, in the sense that it relies on continuity of the mapping used for estimation (see Wolfowitz1957, DonohoLiu1991). The key difference is that in plug-in estimation this mapping is $\psi$, but for regularized estimators the natural mapping is the (sequence of) regularized parameter mappings, $(\psi_{k})_{k=1}^{\infty}$; this difference --- in particular, the fact that we have a sequence of mappings --- introduces nuances that are not present in the standard “plug-in" estimation case. We show that the key component of the convergence rate is the modulus of continuity of the regularized mapping, which, typically, will deteriorate as one moves further into the sequence of regularized mappings, thus yielding a generalized version of the well-known “noise-bias" trade-off. While this result, by itself, does not constitute a big leap from Bickel and Li's framework, we use the underlying insights to propose a data-driven method to choose the tuning parameter that under some conditions yields convergence rates proportional to the “oracle" ones, i.e., those implied by the choice that balances the “noise-bias" trade-off. This method is an extension of the Lepski method as presented in PereverzevSchock2006 for ill-posed inverse problems.\footnote{Similar versions has been used in several particular applications. Closest to our examples are the work: by Pouzo2017 for regularized M-estimators; by ChenChristensen2015 in non-parametric IV regressions; by GineNickl2008 for estimation of the integrated square density; by gaillac2019adaptive in a random coefficient model; by lepski1997 for estimation of a function at a point.}

Our second set of results are concerned with obtaining a type of asymptotic linear representation for regularized estimators. The property of asymptotic linearity is well-known in the literature and is the cornerstone of large sample distribution theory. This property states that the estimator, once centered at the true parameter, is equal to a sample average of a mean zero function --- referred as the influence function --- plus an asymptotically negligible term.

In parametric models, asymptotic linearity is typically satisfied by commonly used estimators like the “plug-in" estimator. In more complex settings such semi-/non-parametric models, however, this is not longer true. In such cases, there are no estimators satisfying this property, because, for instance, the efficiency bound of the parameter of interest is infinite, or more generally, the parameter is not root-n estimable. For these situations, asymptotic representations analogous to asymptotic linearity have been obtained in specific examples for specific regularizations, but, to our knowledge, there is no general approach. This is specially problematic as there is no systematic method for properly standardizing the estimator in situations where the parameter is not root-n estimable.\footnote{For density and regression estimation problems there is a large literature, especially for particular functionals like evaluation at a point; e.g. see EggermontLariccia2001 Vol I and II for references and results. In more general contexts such as M-estimation and GMM-based models, to our knowledge, the literature is much more sparse with only a few papers allowing for slower than root-n parameters in particular settings. Closest to ours are the papers by ChenLiao2014, chen2014sieve in the context of M-estimation models with series/sieve-based estimators; Newey1994 in a two-stage moment model using kernel-based estimators; ChenPouzo2015 in conditional moment models with sieve-based estimators; cattaneo2013optimal in partitioning estimators of the conditional expectation function and its derivatives.} Our goal is to propose a systematic approach by considering a generalization of asymptotic linearity that relaxes certain features of the standard property but still provides a useful asymptotic characterization of the estimator. This property, which is already present in many examples and we refer to as Generalized Asymptotic Linearity (GAL for short), relaxes the standard one in two dimensions: It allows for the location term to be different from the true parameter, and it allows for this centering and the influence function to vary with the sample size. Each of these relaxations attempts to capture different nuances that already exists in the many scattered examples in the literature. Our results, which we now describe, will shed more light on the role and necessity of each.

We provide sufficient conditions for regularized estimators to satisfy GAL. Analogously to the theory of asymptotic linearity for plug-in estimators, our results rely on a notion of differentiability, but contrary to plug-in estimators, it relies on differentiability of each element in the sequence of regularized mappings, $(\psi_{k})_{k=1}^{\infty}$ not on differentiability of the original mapping $\psi$.

As a consequence of this approach, GAL for regularized estimators exhibits two simplified features. First, the location term is given by $\psi_{k}(P)$ which can be interpreted as a psuedo-true parameter. The second simplified feature concerns the influence function and its dependence on the sample size. As in the location term, the dependence on the sample size of the influence function arises only through the dependence of the tuning parameter, $k$, on the sample size. Thus, the relevant object is a sequence of influence functions, each related to the derivative of the elements in $(\psi_{k})_{k=1}^{\infty}$. We view this quantity as the natural departure from the traditional influence function as it is the sequence of regularized mappings, $(\psi_{k})_{k=1}^{\infty}$, an not the original mapping, $\psi$, the one used for constructing the estimator. This last feature allows us to propose a natural and systematic way of standardizing the estimator regardless of whether root-n consistency holds. To explain this, we first note that in situations where asymptotic linearity holds, the proper standardization is given by square root of the sample size divided by the standard error of the value of the influence function. Under GAL the standarization of the regularized estimators turns out to be analogous except that in this case the influence function is indexed by the tuning parameter which at the same time depends on the sample size. Whether the standardization is root-n or slower depends on the behavior of the standard error of the value influence function as we move further into the sequence of regularized mappings (i.e., as $k$ diverges).

Throughout the paper we present examples not to break new ground but to illustrate our assumptions and results, and their scope.

Notation. The term “wpa1-$P$" is short for with probability approaching 1 under $P$, so for a generic sequence of IID random variables $(Z_{n})_{n}$ with $Z_{n} \sim P$, the phrase “$Z_{n} \in A$ wpa1-$P$" formally means $P(Z_{n} \notin A) = o(1)$. For any random variables $(X,Y)$ we use $p_{X}$ and $p_{XY}$ to denote the pdf (w.r.t. Lebesgue) corresponding to $X$ and $X,Y$ resp. For any linear normed spaces $(A,||.||_{A})$ and $(B,||.||_{B})$, let $A^{\ast}$ be the dual of $A$, and for any continuous, homogeneous of degree 1 function $f : (A,||.||_{A}) \mapsto (B,||.||_{B})$, $||f||_{\ast} = \sup_{a \in A \colon ||a||_{A} \ne 1} ||f(a)||_{B}$. For a Euclidean set $S$, we use $L^{p}(S)$ to denotes the set of $L^{p}$ functions with respect to Lebesgue. For any other measure $\mu$, we use $L^{p}(S,\mu)$ or $L^{p}(\mu)$. The norm $||.||$ denotes the Euclidean norm and when applied to matrices it corresponds to the operator norm. For any matrix $A$, let $e_{min}(A)$ denote the minimal eigenvalue. The symbol $\precsim$ denotes less or equal up to universal constants; $\succsim$ is defined analogously.

Setup

Let $\mathbb{Z} \subseteq \mathbb{R}^{d}$ and let $\boldsymbol{z} \equiv (z_{1},z_{2},...) \in \mathbb{Z}^{\infty}$ denote a sequence of IID data drawn from some $P \in \mathcal{P}(\mathbb{Z}) \subset ca(\mathbb{Z})$, where $\mathcal{P}(\mathbb{Z})$ is the set of Borel probability measures over $\mathbb{Z}$ and $ca(\mathbb{Z})$ is the space of signed Borel measures of finite variation. For each $P \in \mathcal{P}(\mathbb{Z})$, let $\mathbf{P}$ be the induced probability over $\mathbb{Z}^{\infty}$. A model is defined as a subset of $\mathcal{P}(\mathbb{Z})$; and it will typically be denoted as $\mathcal{M}$.

remarkSince we only consider IID random variables, it is enough to define a model as a family of probabilities over marginal probabilities. For richer data structures, one would have to define the model as a family of probabilities over $(Z_{1},Z_{2},...)$. See Appendix (ref) for a discussion about how to extend our results to general stationary models. $\triangle$

A parameter on model $\mathcal{M}$ is a mapping $\psi : \mathcal{M} \rightarrow \Theta$ with $(\Theta,||.||_{\Theta})$ being a normed space.\footnote{If the mapping does not point-identified an element of $\Theta$, i.e., $\psi$ is one-to-many, our results go through with minimal changes that account for the fact that $\psi(P)$ is a set in $\Theta$.}

For the results in this paper, we need to endow $\mathcal{M}$ with some topology. For the results in Section (ref) it suffices to work with a distance, $d$, under which the empirical distribution (defined below) converges to $P$. For the results in Section (ref) and beyond, however, it is convenient to have more structure on the distance function, and thus, we work with a distance of the form

align*[align* omitted — 122 chars of source]

where $\mathcal{S}$ is some class of Borel measurable and uniformly bounded functions (bounded by one). For instance, the total variation norm can be viewed as taking $\mathcal{S}$ as the class of indicator functions over Borel sets, and its denoted directly as $||.||_{TV}$; the weak topology over $\mathcal{P}(\mathbb{Z})$ is metricized by taking $\mathcal{S}=LB$ --- the space of bounded Lipschitz functions --- and its norm is denoted directly as $||.||_{LB}$; see VdV-W1996 for a more thorough discussion.

Regularization

Let $\mathcal{D} \subseteq \mathcal{P}(\mathbb{Z})$ be the set of all discretely supported probability distributions. Let $P_{n} \in \mathcal{D}$ be the empirical distribution, where $P_{n}(A) = n^{-1} \sum_{i=1}^{n} 1\{ Z_{i} \in A \}$ for any $A \subseteq \mathbb{Z}$. As illustrated by our examples, in many situations --- especially in non-/semi-parametric models --- the parameter mapping might be either ill-defined (e.g., if $P_{n} \notin \mathcal{M}$) or ill-behaved when evaluated at the empirical distribution $P_{n}$, so it has to be regularized.

The following definition of regularization is based on the first part of the definition in BickelLi06 p. 7. To state it, we define a tuning set as any subset of $\mathbb{R}_{+}$ that is unbounded from above, and the approximation error function as $k \mapsto B_{k}(P) \equiv ||\psi_{k}(P) - \psi(P)||_{\Theta}$.

definitionGiven a model $\mathcal{M}$, a regularization of the parameter mapping $\psi$ is a sequence $\boldsymbol{\psi} \equiv (\psi_{k})_{k \in \mathbb{K}}$ such that $\mathbb{K}$ is a tuning set and \begin{enumerate} • For any $k \in \mathbb{K}$, $\psi_{k} : \mathbb{D}_{\psi} \subseteq ca(\mathbb{Z}) \rightarrow \Theta$ where $\mathbb{D}_{\psi} \supseteq \mathcal{M} \cup \mathcal{D}$. • For any $P \in \mathcal{M}$, $\lim_{k \rightarrow \infty} B_{k}(P) = 0 $. \end{enumerate}

Condition 1 ensures that $\psi_{k}(P_{n})$ and $\psi_{k}(P)$ are well-defined and that they are singletons for all $k \in \mathbb{K}$. Condition 2 ensures that, in the limit, the regularization approximates the original parameter mapping; the limit is warranted as the tuning set $\mathbb{K}$ is unbounded from above. In many applications the tuning set is given by $\mathbb{N}$ but there are applications such as kernel-based estimators, where it is more natural to use a (uncountable) subset of $\mathbb{R}_{+}$.

For each $k \in \mathbb{K}$, the implied estimator is given by $\psi_{k}(P_{n})$ which --- like the “plug-in" estimator --- is permutation invariant. While, this restriction still encompasses a wide array of commonly used methods, it does rule out some estimation methods, notably those that rely on non-trivial sample-splitting procedures. We briefly discuss how to extend our framework to these cases in Appendix (ref).

Conditions 1 and 2 are not enough to obtain “nice" asymptotic properties of the regularized estimator such as consistency and asymptotic normality. In analogy to the standard asymptotic theory for “plug-in" estimators, these properties will be obtained by essentially imposing different degrees of smoothness on the regularization.

Examples

The following examples complement those in BickelLi06 to illustrate that the Definition (ref) encompasses a wide array of commonly used methods.

example[Non-Parametric IV Regression (NPIV)] This example studies a popular regression model used in economics called the Non-parametric Instrumental Variable (IV) model, that belongs to the class of ill-posed inverse problems; see DarollesFanFlorensRenault2011,HallHorowitz2005,AiChen2003,ai2007estimation,NeweyPowell2003,Florens2003,BCK2007 among others. The model is given by \begin{align} E[Y - h(W) \mid X] = 0, \end{align} where $h$ is such that $E[|h(W)|^{2}]<\infty$, $Y$ is the outcome variable, $W$ is the endogenous regressor and $X$ is the IV. We show how our method encompasses commonly used regularizations schemes such as sieves-based and penalized-based ones. For a given subspace of $L^{2}([0,1],p_{W})$, $\Theta$, the model $\mathcal{M}$ is defined as the class of probabilities over $Z = (Y,W,X) \in \mathbb{R} \times [0,1]^{2}$ with pdf with respect to Lebesgue, $p$, such that:\footnote{This restriction is mild and can be changed to accommodate discrete variables simply by requiring pdf's with respect to the counting measure.} (1) $p_{X} = p_{W} = U(0,1)$, $E[|Y|^{2}]<\infty$ and $||p_{XW}||_{L^{\infty}}<\infty$; and (2) there exists a unique $h \in \Theta$ that satisfies (ref). The restriction (1) can be relaxed and is made for simplicity so we can focus on the objects of interest that are $h$ and $P$; it implies that $L^{2}([0,1],p_{X}) = L^{2}([0,1],p_{W}) = L^{2}([0,1])$ which simplifies the derivations.\footnote{To restrict the support to $[0,1]$ is common in the literature (e.g. HallHorowitz2005). At this level of generality, one can always re-define $h$ as $h \circ F_{W}^{-1}$ so that $p_{W} = U(0,1)$; of course this will affect the smoothness properties of $h$. The restriction $p_{X} = U(0,1)$ is really about $p_{X}$ being known, since in that case, one can always take $F_{X}(X)$ as the instrument.} The restriction (2) is what defines an IV non-parametric model. It implies that for any $P \in \mathcal{M}$, $r_{P}(\cdot) \equiv \int y P_{YX}(dy,\cdot)$ is well-defined and belongs to the range of the operator $T_{P} : \Theta \subseteq L^{2}([0,1]) \rightarrow L^{2}([0,1])$ given by $T_{P}[h](\cdot) = \int h(w) p_{WX}(w,\cdot) dw $ for any $h \in L^{2}([0,1])$.\footnote{Alternatively, we can define $T_{P}[h](X)= \int h(w) p(w|X)dw$ and $r_{P}(X)= \int y p(y|X)dy$. Depending on the type of the regularization one has at hand, it is more convenient to use one or the other.} Thus, for any $P \in \mathcal{M}$, $\psi(P)$ is the (unique) solution of $r_{P} = T_{P}[h]$. To illustrate our method, we consider the estimation of a linear functional of $\psi(P)$ of the form $\gamma(P) \equiv \int \pi(w) \psi(P)(w) dw$ for some $\pi \in L^{2}([0,1])$, which by the Riesz representation theorem covers any linear bounded functional on $L^{2}([0,1])$. It is well-known that the estimation problem needs to be regularized. First, we need to regularize the “first stage parameters" --- the operator $T_{P}$ and $r_{P}$; second, given the regularization of $T_{P}$ and $r_{P}$, the inverse problem for finding $\psi(P)$ typically needs to be regularized; e.g. when $T_{P}$ is compact or when $\psi(P)$ is not a singleton. By setting $\mathbb{K}=\mathbb{N}$, the regularization of the “first stage" is given by a sequence of mappings $(T_{k,P},r_{k,P})_{k \in \mathbb{N}}$ such that, for any $k \in \mathbb{N}$, $T_{k,P} : \Theta \rightarrow L^{2}([0,1])$ and $r_{k,P} \in L^{2}([0,1])$. The “second stage" regularization is summarized by an operator $\mathcal{R}_{k,P} : L^{2}([0,1]) \rightarrow L^{2}([0,1])$ for which \begin{align} \psi_{k}(P) = \mathcal{R}_{k,P} [T^{\ast}_{k,P} [r_{k,P}]], \forall P \in lin(\mathcal{M} \cup \mathcal{D}). \end{align} We assume that the regularization structure $(T_{k,P},r_{k,P},\mathcal{R}_{k,P})_{k \in \mathbb{N}}$ is such that: (1) $\lim_{k \rightarrow \infty}||\mathcal{R}_{k,P}[T^{\ast}_{k,P}[g]] - (T^{\ast}_{P} T_{P})^{-1}T^{\ast}_{P}[g]||_{L^{2}([0,1])} = 0$ pointwise over $g \in L^{2}([0,1])$; (2) $\lim_{k \rightarrow \infty}||\mathcal{R}_{k,P}[T^{\ast}_{k,P}[r_{k,P} - r_{P}]]||_{L^{2}([0,1])} =0$. We relegate a more thorough discussion and particular examples of the regularization to Appendix (ref). For now, it suffices to note that the first stage regularization encompasses commonly used regularizations such as the Kernel-based (e.g., DarollesFanFlorensRenault2011, HallHorowitz2005) and the Series-Based (e.g., AiChen2003 and NeweyPowell2003) regularizations, and the second stage regularization encompasses commonly used regularizations such as Tikhonov-/Penalization-based regularization (e.g., DarollesFanFlorensRenault2011 and HallHorowitz2005) and Series-based regularization (e.g., AiChen2003 and NeweyPowell2003). For these combinations, conditions (1)-(2) haven been verified, under primitive conditions, in the literature; e.g. see EnglEtAl1996 Ch. 3-4. It is easy to see that under conditions (1)-(2), the expression in (ref) is in fact a regularization for $\psi(P)$ with $\mathbb{D}_{\psi} \supseteq \mathcal{M} \cup \mathcal{D}$ being a linear subspace specified in expression (ref) in Appendix (ref). From this result, it also follows that $\{\gamma_{k}(P) \equiv \int \pi(w) \psi_{k}(P)(w) dw\}_{k \in \mathbb{N}}$ is a regularization for $\gamma(P)$ (in this case, $\Theta = \mathbb{R}$). $\triangle$

The next is not an example but rather a commonly used estimation technique that also fits in our framework.

example[Regularized M-Estimators] Given some model $\mathcal{M}$, the parameter mapping is defined as \begin{align*} \psi(P) = \arg\min_{\theta \in \Theta} E_{P}[\phi(Z,\theta)], \forall P \in \mathcal{M}, \end{align*} where $\Theta$ and $\phi : \mathbb{Z} \times \Theta \rightarrow \mathbb{R}_{+}$ are primitives of the problem and are such that the argmin is non-empty for any $P \in \mathcal{M}$. We impose the following assumptions over $(\mathcal{M},\Theta,\phi)$: $\Theta$ is a subspace of $L^{q}$ where $L^{q} \equiv L^{q}(\mathbb{Z},\mu)$ for any $q \in [1,\infty)$ and some finite measure $\mu$, and for $q=\infty$, $L^{\infty} = \mathbb{C}(\mathbb{Z},\mathbb{R})$;\footnote{The class $\mathbb{C}(\mathbb{Z},\mathbb{R})$ is the class of continuous and uniformly bounded real-valued functions on $\mathbb{Z}$.} and $\theta \mapsto E_{P}[|\phi(Z,\theta)|]$ bounded and continuous, for all $P \in \mathcal{M}$. The regularization is lifted from Pouzo2017 and is defined using $\mathbb{K} = \mathbb{N}$ by: a sequence of nested linear subspaces of $L^{q}$, $(\Theta_{k})_{k \in \mathbb{N}}$, such that $dim(\Theta_{k}) = k$ and the union is dense in $\Theta$; a vanishing real-valued sequence $(\lambda_{k})_{k \in \mathbb{N}}$ with $\lambda_{k} \in (0,1]$ and a lower-semi compact function $Pen : L^{q} \rightarrow \mathbb{R}_{+}$ such that, for each $k \in \mathbb{N}$\footnote{A lower-semi compact function is one with compact lower contour sets.} \begin{align*} \psi_{k}(P) \equiv \arg\min_{\theta \in \Theta_{k}} E_{P}[\phi(Z,\theta)] + \lambda_{k} Pen(\theta) \end{align*} is a singleton for any $P \in \mathcal{M} \cup \mathcal{D}$. It is clear that condition 1 in Definition (ref) holds; we now show by contradiction that Condition 2 also holds. Suppose that there exists a $\epsilon>0$ such that $||\psi_{k}(P)-\psi(P)||_{\Theta} \geq \epsilon$ for all $k$ large. Let $\Pi_{k} \psi(P)$ be the projection of $\psi(P)$ onto $\Theta_{k}$; for sufficiently large $k$, $||\Pi_{k} \psi(P)-\psi(P)||_{\Theta} \leq \epsilon$. Then, by optimality of $\psi_{k}(P)$ and some algebra, for large $k$, $\inf_{\theta \in \Theta \colon ||\theta-\psi(P)||_{\Theta} \geq \epsilon} E_{P}[\phi(Z,\theta)] \leq E_{P}[\phi(Z,\psi(P))] - \{ E_{P}[\phi(Z,\psi(P))-\phi(Z,\Pi_{k}\psi(P))] + \lambda_{k} Pen(\Pi_{k}\psi(P))\}$. By continuity of $E_{P}[\phi(Z,\cdot)]$, $\lambda_{k} \downarrow 0$ and convergence of $\Pi_{k}\psi(P)$ to $\psi(P)$ the term in the curly bracket vanishes as $k$ diverges, leading to the contradiction $\inf_{\theta \in \Theta \colon ||\theta-\psi(P)||_{\Theta} \geq \epsilon} E_{P}[\phi(Z,\theta)] \leq E_{P}[\phi(Z,\psi(P))]$. $\triangle$

As these examples, those below and those in BickelLi06 illustrate, the formulation of regularization is very flexible. Indeed, Definition (ref) is sufficiently mild that one may wonder whether any results can be obtained at the present level of generality. In what follows, we establish some useful asymptotic properties by imposing additional smoothness restrictions on the regularization.

Consistency for Regularized Estimators

We say a regularization is consistent if for each $k \in \mathbb{K}$, $\psi_{k}(P_{n}) = \psi_{k}(P) + o_{P}(1)$. Using Definition (ref) it is straightforward to show that for consistent regularizations there exists a tuning parameter sequence $(k_{n})_{n}$ such that $\psi_{k_{n}}(P_{n}) = \psi(P) + o_{P}(1)$. This claim, however, is silent about how to choose the tuning parameter sequence and what is the corresponding rate. This feature does not seem to be a shortcoming of the claim, but rather a manifestation of the fact that the consistency requirement is very mild. In other words, it would appear that some strengthening of this requirement is needed if we want obtain stronger conclusions. In this section and the next one, our goal is to give general but “easy-to-interpret" conditions that will enable us to obtain convergence rates and a data-driven choice of tuning parameter sequence that achieves such rates, thus providing a general guideline for establishing asymptotic concentration results for regularized estimators.

In order to do this, we take as given that the empirical distribution converges to the truth at rate, $(r_{n})_{n}$ --- a diverging positive real-valued sequence --- under a distance $d$, i.e., $d(P_{n},P) = O_{P}(r^{-1}_{n})$. Such results can typically be obtained from empirical process theory and are not the focus of this paper. Under this condition, from the classical results of Wolfowitz1957, continuity of the regularization (with respect to $d$) presents itself as a natural condition to establish consistency of the regularization. To define the notion of continuity formally we say a function $f : \mathbb{R}_{+} \rightarrow \mathbb{R}_{+}$ is a modulus of continuity if $f$ is continuous, non-decreasing and such that $f(t)=0$ iff $t=0$.

definition[Continuous Regularization] A regularization $\boldsymbol{\psi}$ of $\psi$ is continuous at $P \in \mathbb{D}_{\psi}$ with respect to $d$, if there exists a family of modulus of continuity $(\delta_{k})_{k \in \mathbb{R}}$ such that for any $k \in \mathbb{K}$ \begin{align} || \psi_{k}(P') - \psi_{k}(P) ||_{\Theta} \leq \delta_{k}(d(P',P)) \end{align} for any $P' \in \mathbb{D}_{\psi}$.

The definition is equivalent to the standard “$\delta/\epsilon$"-definition of continuity because the modulus of continuity of $\psi_{k}$, $\delta_{k}$, can converge to $0$ arbitrary slowly. Admittedly, this condition is not necessary to obtain our results as it requires continuity to hold for any deviation of $P$, but the notion of consistent regularization requires continuity only along those deviations taken by $P_{n}$ for all $n$ large enough. However, we still view this condition as a reasonable “initial step" to obtain consistency results. Moreover, the definition does not impose any uniform bounds on $\delta_{k}$ across different $k \in \mathbb{K}$. While such restriction would simplify the proofs considerably, it is too strong for many applications. Recall that the regularization is introduced precisely due to the poor behavior of $\psi$ at $P$.

As the lemma (ref) in Appendix (ref) shows, when the regularization is continuous, the “sampling error" term, $||\psi_{k}(P_{n}) - \psi_{k}(P)||_{\Theta}$ is of order $\delta_{k}(r^{-1}_{n})$ in probability. From this result and the fact that $\limsup_{n \rightarrow \infty} \delta_{k}(r^{-1}_{n}) = 0$ for every $k \in \mathbb{K}$, a simple diagonalization argument establishes existence of a tuning parameter sequence, $(k_{n})_{n}$ for which $(\psi_{k_{n}}(P_{n}))_{n \in \mathbb{N}}$ is consistent. The next theorem formalizes this claim.

theorem[Consistency of Regularized Estimators] Suppose a regularization, $\boldsymbol{\psi}$, is continuous (at $P$) with respect to $d$ such that $d(P_{n},P) = O_{P}(r^{-1}_{n})$. Then there exists a $(k_{n})_{n \in \mathbb{N}}$ in $\mathbb{K}$ such that \begin{align*} B_{k_{n}}(P) = o(1) and \delta_{k_{n}}(d(P_{n},P)) = o_{P}(1), \end{align*} and \begin{align*} ||\psi_{k_{n}}(P_{n}) - \psi(P) ||_{\Theta} = o_{P}(1). \end{align*}
proofSee Appendix (ref).

Although constructive, the consistency result outlined above suffers from the potential problem that the associated tuning parameter sequence makes no attempt to ensure that the magnitude of the estimation error is in some sense minimal. Indeed, because the tuning parameter sequence has been designed to satisfy only the minimal requirement that $B_{k_{n}}(P) + \delta_{k_{n}}(d(P_{n},P)) = o_{P}(1)$; the only property that can be claimed on the part of the regularized estimator is consistency. The goal of the next section is to improve on this as well as to construct a data-driven choice of tuning parameter.

Data-driven Choice of Tuning Parameter

Let $\mathbb{K}_{n} $ be the (user-specified) set over which the tuning parameter is chosen; it is assumed to be a finite subset of $\mathbb{K}$.\footnote{In Appendix (ref) we extend the main theorem of this section to the case where $\mathbb{K}_{n}$ is any closed set of $\mathbb{K}$, not necessarily finite.} Also, in this section, we strengthen the consistency of $P_{n}$ to $d(P_{n},P) = o_{P}(r^{-1}_{n})$ (remark (ref) below discusses the reason behind this choice).

By Theorem (ref) and the triangle inequality it is easy to see that the distance between the regularized estimator, $\psi_{k_{n}}(P_{n})$, and the true parameter is bounded by the sum of two terms: the “sampling error", $\delta_{k_{n}}(r^{-1}_{n})$ and the “approximation error", $B_{k_{n}}(P)$, which generalizes the well-known “noise-bias" trade-off present in many applications. Thus, this observation suggests the following criterion to construct tuning sequence $(k_{n})_{n}$ that yields a consistent estimators:

align*[align* omitted — 86 chars of source]

which minimizes the trade-off between the approximation and the sampling errors. This choice represents commonly used heuristics and it is a good prescription to obtain approximately optimal rate of convergences (BirgeMassart1998). However, often times it is unfeasible, since it relies on knowledge of the approximation error, which is typically unknown because it depends on features of the unknown $P$.

It is thus desirable to construct a choice of tuning parameter that sidestep this issue while still providing similar rates of convergence. We propose an adaptation of the Lepski method that provides a data-driven choice that satisfies these properties, without requiring knowledge of $B_{k}(P)$. Due to the nature of the Lepski method, in order to establish the desired results we need monotonicity of the sampling and approximation errors as functions of the tuning parameter. Since these functions may not be monotonic, we replace them by monotonic majorants. Formally, let $k \mapsto \bar{B}_{k}(P)$ be a non-increasing function from $\mathbb{R}_{+}$ to itself such that $\bar{B}_{k}(P) \geq ||\psi_{k}(P) - \psi(P)||_{\Theta}$ for all $k \geq 0$, $\lim_{k \rightarrow \infty} \bar{B}_{k}(P) = 0$, and, for each $n \in \mathbb{N}$, let $ k \mapsto \bar{\delta}_{k}(r^{-1}_{n})$ be a non-decreasing function from $\mathbb{R}_{+}$ to itself such that $\bar{\delta}_{k}(r^{-1}_{n}) \geq \delta_{k}(r^{-1}_{n})$.

For each $n \in \mathbb{N}$, given $\mathbb{K}_{n}$ and a non-decreasing function $\upsilon_{n} : \mathbb{K}_{n} \rightarrow \mathbb{R}_{+}$ such that $k \mapsto \upsilon_{n}(k) = 4\bar{\delta}_{k}(r^{-1}_{n})$, the Lepski choice is given by

align*[align* omitted — 79 chars of source]

where

align[align omitted — 220 chars of source]

The following theorem is the main result of this section.

theoremSuppose a regularization, $\boldsymbol{\psi}$, is continuous (at $P$) with respect to $d$ and there exists a real-valued positive diverging sequence $(r_{n})_{n \in \mathbb{N}}$ such that $d(P_{n},P) = o_{P}(r^{-1}_{n})$. Then \begin{align*} || \psi_{\tilde{k}_{n}}(P_{n}) - \psi(P) ||_{\Theta} = O_{P} \left( \inf_{k \in \mathbb{K}_{n}} \{ \bar{\delta}_{k}(r^{-1}_{n}) + \bar{B}_{k}(P) \} \right). \end{align*}
proofSee Appendix (ref).
remarkThe rate $(r_{n})_{n}$ is defined as $d(P_{n},P) = o_{P}(r^{-1}_{n})$, as opposed to $d(P_{n},P) = O_{P}(r^{-1}_{n})$ as in Theorem (ref). That is, $r_{n}$ diverges (arbitrary) slower than the usual rates for $P_{n}$ --- which is typically given by $\sqrt{n}$ in our context. This type of lost is common when studying choice of tuning parameters (cf. GineNickl2008 and references therein). In our setup, it stems from the following fact: Take a rate $(s_{n})_{n}$ such that $d(P_{n},P) = O_{P}(s^{-1}_{n})$. For this rate, there are unknown constants (e.g. $M$ in Lemma (ref)) which will render our data-driven choice infeasible. So to avoid them it suffices to replace $s^{-1}_{n}$ by a (arbitrary) slower rate, e.g. $r^{-1}_{n} = \log (1+n) s^{-1}_{n}$ or $r^{-1}_{n} = \log (\log (1+n)) s^{-1}_{n}$. $\triangle$
remarkThe rate of convergence does not depend on the “complexity" of the set $\mathbb{K}_{n}$. This result stems from a certain “separability" property of the estimator: The probability statements stem from the behavior of $d(P_{n},P)$ which does not depend on $k$ nor on $\mathbb{K}_{n}$, the tuning parameter $k$ only appear through the topological properties of the regularization. $\triangle$
remark[Heuristics of the proof of Theorem (ref)] Heuristically, for any $k \in \mathbb{K}_{n}$ that is larger or equal than $\tilde{k}_{n}$ it follows that $||\psi_{\tilde{k}_{n}}(P_{n}) - \psi(P) ||_{\Theta}$ is bounded above (up to constants) by $\bar{\delta}_{k}(r^{-1}_{n}) + \bar{B}_{k}(P)$ with probability approaching one. Lemma (ref) in Appendix (ref) formalizes this observation and shows that in order to establish the claim of the theorem it suffices to show existence of a tuning parameter in $\mathbb{K}_{n}$ that is larger or equal than $\tilde{k}_{n}$ (with probability approaching one) and minimizes (up to constants) $k \mapsto \{ \bar{\delta}_{k}(r^{-1}_{n}) + \bar{B}_{k}(P) \}$ over $\mathbb{K}_{n}$. Moreover, since $\tilde{k}_{n}$ is chosen as the minimal value in $\mathcal{L}_{n}$, to obtain the former condition it suffices to show that the tuning parameter belongs to $\mathcal{L}_{n}$ (with high probability). By studying “projections" onto $\mathbb{K}_{n}$ of the tuning parameter that balances $ k \mapsto \bar{\delta}_{k}(r^{-1}_{n})$ and $k \mapsto \bar{B}_{k}(P)$ we are able to explicitly construct a sequence of tuning parameters that satisfies these conditions; it is in this part that the monotonicity properties of these mappings are used. See Lemmas (ref) and (ref) in Appendix (ref). $\triangle$

The following corollary is a direct consequence of Theorem (ref) and its proof is omitted.

corollarySuppose $k \mapsto B_{k}(P)$ and $k \mapsto \delta_{k}(r^{-1}_{n})$ are continuous, and non-increasing and non-decreasing resp.. Then under the conditions of Theorem (ref), it follows \begin{align*} || \psi_{\tilde{k}_{n}(r_{n})}(P_{n}) - \psi(P) ||_{\Theta} = O_{P} \left( \inf_{k \in \mathbb{K}_{n}} \{ \delta_{k}(r^{-1}_{n}) + B_{k}(P) \} \right). \end{align*}

Theorem (ref) and its corollary show that our data-driven choice of tuning parameter achieves the same rate as the one corresponding to the “infeasible" choice, provided the monotonicity conditions hold. Hence, this result, Theorem (ref) and Theorem (ref) offer a general road-map for establishing consistency and convergence rates of regularized estimators based on continuity of the regularization.\footnote{Without further restriction on $\mathbb{K}_{n}$ there is no guarantee that $\inf_{k \in \mathbb{K}_{n}} \{ \delta_{k}(r^{-1}_{n}) + B_{k}(P) \}$ is of the same magnitude as $\inf_{k \in \mathbb{K}} \{ \delta_{k}(r^{-1}_{n}) + B_{k}(P) \}$. In Appendix (ref), Proposition (ref) gives conditions on $\mathbb{K}_{n}$ that guarantee this result.}

Examples

The following examples illustrate how to apply our results to existing applications and in the process establish a new result. Example (ref) considers the case of bootstrapping the mean of a distribution when it is known to be non-negative, and it is based on andrews2000inconsistency. In this paper, the author showed inconsistency of the bootstrap and proposed several consistent alternatives; we take one --- the “k-out-of-n" bootstrap (BickelFreedman1981) --- and illustrate how our methods can be used to derive the rate of convergence of this procedure and to choose the tuning parameter $k$ that achieves this rate. To our knowledge this last result is novel.\footnote{BickelLi06 and bickel2008choice perform a similar exercise but for a different case: estimation of largest order statistic.} Example (ref) provides primitive conditions for establishing continuity in M-estimation problems.

example[Bootstrap when the parameter is on the boundary] Let $\mathcal{M}$ be the class of Borel probability measures over $\mathbb{R}$ with non-negative mean, unit variance and finite third moments; the non-negativity of the mean is a formalization that captures the issue of a parameter at the boundary. The object of interest is the law of an estimator of the mean, $\boldsymbol{z} \mapsto T_{n}(\boldsymbol{z},P) = \sqrt{n}( \max\{ n^{-1} \sum_{i=1}^{n} z_{i} , 0 \} - \max \{ E_{P}[Z] , 0 \} )$. Thus, let, for each $k \in \mathbb{K} = \mathbb{N}$, $\psi_{k} : \mathcal{P}(\mathbb{R}) \rightarrow \mathcal{P}(\mathbb{R})$ be defined as \begin{align*} \psi_{k}(P)(A) \equiv \mathbf{P} \left( \{ \boldsymbol{z} \colon T_{k}(\boldsymbol{z},P) \in A \} \right) , \forall A Borel. \end{align*} In particular, for $P=P_{n}$, it follows that \begin{align*} \psi_{k}(P_{n})(A) = \mathbf{P}_{n} \left( \sqrt{k} \left( \max\{ k^{-1}\sum_{i=1}^{k} Z^{\ast}_{i} , 0 \} - \max\{ n^{-1}\sum_{i=1}^{n} Z_{i} , 0 \} \right) \in A \right) , \forall A Borel \end{align*} where $(Z^{\ast}_{i})_{i=1}^{n}$ is an IID sample drawn from $P_{n}$ and $\mathbf{P}_{n}$ is the probability over $\mathbb{Z}^{\infty}$ induced by $P_{n}$. It is easy to see that $\psi_{n}(P_{n})$ is the standard bootstrap estimator while $\psi_{k}(P_{n})$ for $k < n$ is the $k$-out-of-$n$ bootstrap estimator. andrews2000inconsistency showed that the “plug-in estimator", $\psi_{n}(P_{n})$, while well-defined, fails to approximate the law of $T_{n}$, $\psi_{n}(P)$, even in the limit; but he showed that for certain sequences, $(k_{n})_{n}$, $\psi_{k_{n}}(P_{n})-\psi_{n}(P)$ converge to zero as $n$ diverges. We now recast this result using the tools developed in this paper; by doing so we are able to provide a data-driven choice of the tuning parameter $k_{n}$. To do this, we first show that the $(\psi_{k})_{k \in \mathbb{N}}$ is continuous in the sense of Definition (ref). Let $\Theta = \mathcal{P}(\mathbb{R})$ and let $|| \cdot ||_{\Theta} \equiv || \cdot ||_{LB}$, where recall $LB$ is the class of real-valued Lipschitz with constant one function. This norm is one of the notions of distance typically used to establish validity of the Bootstrap. Also, let $\mathcal{W}(\cdot ,\cdot ) $ denote the Wassertein distance over $\mathcal{P}(\mathbb{Z})$, that is $\mathcal{W}(P,Q) \equiv \inf_{\zeta \in H(P,Q)} \int |z - z'| \zeta(dz,dz') $, where $H(P,Q)$ is the set of Borel probabilities over $\mathbb{Z}^{2}$ with marginals $P$ and $Q$. The following proposition suggests the form of the modulus of continuity $\delta_{k}$. \begin{proposition} For any $ k \in \mathbb{N}$, $||\psi_{k}(P) - \psi_{k}(Q)||_{\Theta} \leq 2 \sqrt{k} \mathcal{W}(P,Q)$ for any $P$ and $Q$ in $\mathcal{M} \cup \mathcal{D}$. \end{proposition} \begin{proof} See Appendix (ref). \end{proof} The previous results suggests $\mathcal{W}$ as the natural distance over $\mathcal{P}(\mathbb{Z})$. In addition, the result also indicates that $\delta_{k}(t) = 2 \sqrt{k} t$ for all $t \in \mathbb{R}_{+}$, which is increasing and continuous as a function of $(t,k)$. We now apply the results in Theorem (ref) to choose the number of draws for the k-out-n bootstrap. Theorem 1 in FournierHal2015 (their results are applied with $d=1$, $p=1$ and $q=2$) shows that $\mathcal{W}(P_{n},P) = O_{P}(n^{-1/2})$. Therefore, we take $r^{-1}_{n} = l_{n}n^{-1/2}$ where $(l_{n})_{n}$ diverges arbitrary slowly. We also take $\mathbb{K}_{n} = \{1,...,n\}$; it is clear that $k \mapsto \bar{B}_{k}(P) = B_{k}(P)$ and $k \mapsto \bar{\delta}_{k}(r^{-1}_{n}) = \delta_{k}(r^{-1}_{n})$. Given these choices, for each $n \in \mathbb{N}$, let $\tilde{k}_{n}$ be the choice of tuning parameter proposed above. Theorem (ref) imply the following result. \begin{proposition} $ || \psi_{\tilde{k}_{n}}(P_{n}) - \psi_{n}(P) ||_{LB} = O_{P} \left( \inf_{k \in \{1,...,n\}} \{ l_{n} \sqrt{k} n^{-1/2} + k^{-1/2} E_{P}[|Z|^{3}] \} \right)$. \end{proposition} \begin{proof} See Appendix (ref). \end{proof} The RHS of the expression implies that the rate of convergence is given by $\sqrt{l_{n}} n^{-1/4}$. To our knowledge there is no data-driven method to choose the tuning parameter in this example. bickel2008choice propose a similar method to ours in a different example: Inference on the extrema of an IID sample. The authors obtain polynomial rates of convergence that are slower than ours but for a stronger norm. $\triangle$
example[Regularized M-Estimators (cont.)] The following proposition shows that the regularization is continuous and more importantly it provides a “natural" choice of distance and illustrates the role of the regularization structure $\langle (\lambda_{k},\Theta_{k})_{k},Pen \rangle$ and primitives $(\Theta,\phi)$ for determining the rate of convergence of the regularized estimator. Henceforth, let $(\theta,P,k) \mapsto Q_{k}(P,\theta) \equiv E_{P}[\phi(Z,\theta)] + \lambda_{k} Pen(\theta)$. \begin{proposition} For each $k \in \mathbb{N}$ and $P \in \mathcal{M} \cup \mathcal{D}$, \begin{align*} ||\psi_{k}(P) - \psi_{k}(P') ||_{L^{q}} \leq \Gamma^{-1}_{k}(d(P,P')), \forall P' \in \mathcal{M} \cup \mathcal{D}, \end{align*} where for all $t > 0$ \begin{align*} \Gamma_{k}(t) = \inf_{s \geq t} \left\{ \min_{\theta \in \Theta_{k} \colon ||\theta - \psi_{k}(P)||_{L^{q}} \geq s} \frac{ Q_{k}(P,\theta) - Q_{k}(P,\psi_{k}(P)) }{s} \right\} \end{align*} and $d(P,P') \equiv \max_{k \in \mathbb{N}} ||P-P'||_{\mathcal{S}_{k}}$, where $\mathcal{S}_{k} \equiv \left\{ \frac{\phi(.,\theta) - \phi(.,\psi_{k}(P))}{||\theta - \psi_{k}(P)||_{\Theta}} \colon \theta \in \Theta_{k} \right\} $.\footnote{We define $\Gamma_{k}(0)=0$. The “$\inf_{s \geq t}$" ensures that $\Gamma_{k}$ is non-decreasing; it can be omitted if such property is not needed. The “$max_{k \in \mathbb{N}}$" comes from the fact that $d$ cannot depend on $k$ in the definition of continuity.} \end{proposition} \begin{proof} See Appendix (ref) \end{proof} Following ShenWong1995, the proof applies the standard arguments due to Wald --- for establishing consistency of estimators --- to “strips" of the sieve set $\Theta_{k}$; by doing so, one improves the rates obtained from the standard Wald approach. The proposition suggests the natural notion of distance over the space of probabilities, that is defined by the class of “test functions" given by $\left( \frac{\phi(z,\theta) - \phi(z,\psi_{k}(P))}{||\theta - \psi_{k}(P)||_{\Theta}} \right)_{\theta \in \Theta_{k}}$. By imposing additional conditions on $\phi$ and $\Theta_{k}$ one can embed the class $\mathcal{S}_{k}$ into well-known classes of functions for which one has a bound for the supremum of the empirical process $f \mapsto n^{-1}\sum_{i=1}^{n} f(Z_{i}) - E_{P}[f(Z)] $, and thus bounds for $d(P_{n},P)$. For instance, if $\theta \mapsto \frac{d\phi(z,\theta)}{dz}$ is Lipschitz uniformly in $z$, then by using the mean value theorem and some algebra it follows that $\mathcal{S}_{k} \subseteq LB$ for every $k$, and thus $d(P_{n},P) = O_{P}(n^{-1/2})$ (see VdV-W1996). The modulus of continuity, $\Gamma^{-1}_{k}$ is non-decreasing and is continuous over $t >0$ (see the proof), and by definition $\Gamma_{k}(0) = 0$. Its behavior is determined by how well the criterion separates points in $\Theta_{k}$ relative to the norm $||.||_{L^{q}}$; the flatter $Q_{k}(P,\cdot)$ is around its minimizer, the larger $\Gamma^{-1}_{k}$. Importantly, even though $\Gamma_{k}(t) > 0$ for each $k$ (recall that $\psi_{k}(P)$ is assumed to be unique), as $k$ diverges, $\Gamma_{k}(t)$ may approach zero. This phenomena relates to the potential ill-posedness of the original problem, and will affect the rate of convergence of the estimator. To shed some more light on the behavior of $\Gamma_{k}$ and on the potential ill-posedness, consider the case where, $q=2$, $Q(P,\cdot)$ is strictly concave and smooth, and $Pen(.) = ||.||^{2}_{L^{2}}$. Since $\psi_{k}(P)$ is a minimizer, $Q_{k}(P,\cdot)$ behaves locally as a quadratic function, in particular $\Gamma_{k}(t) \geq 0.5 (C_{k} + \lambda_{k}) t$ for some non-negative constant $C_{k}$ related to the Hessian of $Q(P,\cdot)$, and thus $\Gamma_{k}^{-1}(t) \precsim (C_{k} + \lambda_{k})^{-1}t$. If $C_{k} \geq c>0$ then $ \Gamma_{k}^{-1}(t) \precsim t$; we deem this case to be well-posed as $||\psi_{k}(P') - \psi_{k}(P)||_{L^{q}} \precsim d(P',P)$.\footnote{This case relates to the so-called identifiable uniqueness condition (see WW1991).} On the other hand, if $\lim\inf_{k \rightarrow \infty} C_{k} = 0$ then, while the previous bound for the modulus of continuity is not possible, the following bound $\Gamma_{k}^{-1}(t) \precsim \lambda^{-1}_{k} t$ is. This case is deemed to be ill-posed and $||\psi_{k}(P') - \psi_{k}(P)||_{L^{q}} \precsim \lambda_{k}^{-1} d(P',P)$. Finally, under the conditions discussed in the previous paragraph, in the ill-posed case, $k \mapsto \bar{\delta}_{k}(.) = \delta_{k}(.)$ if $k \mapsto \lambda_{k}$ is chosen to be non-increasing and continuous.\footnote{For the well-posed case the condition holds trivially.} Thus Theorem (ref) delivers a choice of tuning parameter that achieves consistency and a rate of $\min_{k \in \mathbb{N}} \{ \lambda^{-1}_{k} \times r^{-1}_{n} + \inf_{l \geq k}||\psi_{l}(P) - \psi(P)||_{L^{q}} \}$, where $(r_{n})_{n}$ is such that $\max_{k \in \mathbb{N}} ||P_{n} - P||_{\mathcal{S}_{k}} = o_{P}(r^{-1}_{n})$. $\triangle$

Asymptotic Representations for Regularized Estimators

The goal of this section is to provide “easy-to-interpret" sufficient conditions to obtain a generalized asymptotic linear (GAL) representation for regularized estimators, assuming that $P_{n}$ converges to $P$ in some sense. Throughout this section we assume $\Theta \subseteq \mathbb{R}$ to simplify the exposition; the results can be easily be extended to vector-valued parameters.\footnote{In other cases where the parameter of interest is infinite-dimensional GAL is too weak and a stronger notion is needed; we refer the reader to a previous version of this paper jansson2017general for this case.}

We say a regularization, $\boldsymbol{\psi}$, is asymptotically linear if there exists $\boldsymbol{\nu} \equiv (\nu_{k})_{k\in \mathbb{K}}$ such that, for all $k \in \mathbb{K}$, $\nu_{k} \in L^{2}_{0}(P) \equiv \{ f\in L^{2}(P)\setminus\{0 \} \colon E_{P}[f(Z)] = 0 \}$ and

align*[align* omitted — 102 chars of source]

In this case, we say $\boldsymbol{\nu}$ is the influence of the regularization. It is straightforward to show that for such regularizations there exists a sequence $(k_{n})_{n}$ for which the following property is satisfied:

definition[Generalized Asymptotic Linearity: GAL($\boldsymbol{k}$)] A regularization $\boldsymbol{\psi}$ satisfies generalized asymptotic linearity for $\boldsymbol{k} \colon \mathbb{N} \rightarrow \mathbb{K}$ at $P \in \mathbb{D}_{\psi}$ with influence $\boldsymbol{\nu}$, if \begin{align} \left| \psi_{\boldsymbol{k}(n)}(P_{n})-\psi_{\boldsymbol{k}(n)}(P) - n^{-1} \sum_{i=1}^{n} \nu_{\boldsymbol{k}(n)}(Z_{i}) \right| = o_{P}(n^{-1/2} ||\nu_{\boldsymbol{k}(n)}||_{L^{2}(P)} ). \end{align}

If a regularization satisfies GAL($\boldsymbol{k}$) then, in order to study its asymptotic behavior, it suffices to study the behavior of $ n^{-1/2} \sum_{i=1}^{n} \frac{\nu_{\boldsymbol{k}(n)}(Z_{i})}{||\nu_{\boldsymbol{k}(n)}||_{L^{2}(P)}}$. Moreover, this property suggests a systematic way for how to scale the estimator: Using by $\sqrt{n}/||\nu_{\boldsymbol{k}(n)}||_{L^{2}(P)}$ as opposed to just $\sqrt{n}$. This insight is particularly useful in situations where root-n estimation is not possible. Thus, this property can be viewed as extending the standard asymptotic linearity one for root-n estimable parameters to a larger class of problems.

Representations akin to GAL are already present in many examples; our contribution, as we view it, is to put these insights on a common framework so they can be applied more generally and to provide primitive properties on the structure of the regularization that guarantee asymptotic linearity --- and consequently, guarantee GAL. Regarding this last point, it is well known that in cases where the “plug-in" method is used, differentiabilty is the natural property, the definition of asymptotic linear regularization suggests differentiability of $\psi_{k}$ for each $k \in \mathbb{K}$ as reasonable starting point. Before presenting the definition of differentiable regularization we present a classical example that illustrates GAL, the influence and the scaling when the parameter is not root-n estimable

example[Density Evaluation] The parameter of interest is the density function evaluated at a point, which can be formally viewed as a mapping from the space of probability distributions to $\mathbb{R}$, given by $P \mapsto \psi(P) = p(0)$, where $p$ denotes the pdf of $P$. It is well known that this problem needs to be regularized. The standard estimator is given by $n^{-1} \sum_{i=1}^{n} \kappa_{k}(Z_{i})$ where $\kappa_{k}( \cdot) = k \kappa(k \cdot )$, $\kappa$ is a kernel (i.e., a smooth function over $\mathbb{R} \setminus \{0\}$ that ingrates to one); $1/k$ acts as the bandwidth of the kernel estimator. This estimator can be cast as $\psi_{k}(P_{n})$ where \begin{align*} P \mapsto \psi_{k}(P) = (\kappa_{k} \star P)(0) \equiv \int_{\mathbb{R}} \kappa_{k} \left( z \right) P(dz), \forall k \in \mathbb{N}. \end{align*} It is well-known that the parameter, $p(0)=\psi(P)$, is not root-n estimable and thus the proposed estimator is not asymptotically linear. The following representation, however, does hold: \begin{align*} \psi_{k}(P_{n}) - \psi_{k}(P) = & n^{-1} \sum_{i=1}^{n} \{ \kappa_{k}(Z_{i}) - E_{P}[\kappa_{k}(Z)] \} \end{align*} which can be viewed as a generalization of asymptotic linearity in which the estimator $\psi_{k}(P_{n})$ is centered at $\psi_{k}(P)=\int \kappa_{k}(z)p(z) dz $ instead of $\psi(P)=p(0)$. Moreover, by drawing an analogy with the standard approach for root-n estimable parameters (see HampelEtAl2011, bickeletal1998efficient, newey1990semiparametric), for each $k$, the term in the curly brackets can be thought as an influence function. This term plays a crucial role on determining the asymptotic distribution of the estimator and on determining the proper way of standardizing it. For general regularized estimators, exact representations of this form are not always possible; however, in this section we identify a class of regularizations --- satisfying a certain differentiability notion (see Definition (ref)) --- that admit, asymptotically, an analogous representation, with the influence function being a function of the derivative of the regularization. It is well-known that the scaling is given by $\sqrt{n / k_{n}}$ which is slower (for some $k_{n}$ that diverges with $n$) than the “standard" $\sqrt{n}$. The $\sqrt{k_{n}}$ correction arises because it is the correct order of the influence function, i.e., $\sqrt{Var_{P} \left( n^{-1/2} \sum_{i=1}^{n} \{ \kappa_{k_{n}}(Z_{i}) - E_{P}[\kappa_{k_{n}}(Z)] \} \right)} = \sqrt{Var_{P}\left( \kappa_{k_{n}}(Z) \right)} \asymp \sqrt{k_{n}}$. Our results extend this simple observation to a large class of regularizations, thereby providing a systematic way for “standardizing" the estimator: By using $\sqrt{n}$ divided the standard deviation of the influence function, which it depends on $n$ through the tuning parameter. $\triangle$

We now present the definition of differentiable regularization. For this, let $\mathcal{T}_{P} \equiv \{ a \mu \colon a\geq 0~and~ \mu \in \mathcal{D} - \{P\}\}$ and let $\tau$ be any locally convex topology over $ca(\mathbb{Z})$ dominated by $||.||_{TV}$. \footnote{Since we are working with measures, and not probabilities, it is convenient to allow for (non-metrizable) topologies. Locally convex topology means that it is constructed in terms of a family of semi-norms; dominated by $||.||_{TV}$ means that for any semi-norm, $\rho$, $\rho(Q) =O( ||Q||_{TV})$ for all $Q \in ca(\mathbb{Z})$.}

definition[Differentiable Regularization: DIFF($P,\mathcal{C}$)] A regularization $\boldsymbol{\psi}$ is differentiable at $P \in \mathbb{D}_{\psi}$ tangential to $\mathcal{T}_{P}$ under the class $\mathcal{C} \subseteq 2^{\mathcal{T}_{P}}$, if for any $k \in \mathbb{K}$, there exists a $D \psi_{k}(P) : \mathcal{T}_{P} \rightarrow \Theta$ $\tau$-continuous and linear such that for any $U \in \mathcal{C}$, \begin{align} \lim_{t \downarrow 0} \sup_{Q\in U} |\eta_{k}(tQ)|/t =0, where Q \mapsto \eta_{k}(Q) \equiv \psi_{k}(P+Q) - \psi_{k}(P) - D \psi_{k}(P)[Q]. \end{align}
remarkThe functional $D \psi_{k}(P)$ acts as the gradient of $\psi_{k}$ at $P$. The set $\mathcal{T}_{P}$ is the tangent set, i.e., the set that contains all the directions of the curves at $P$ that we are considering. It turns out that to obtain an asymptotic linear representation for the regularization, it is enough to consider curves of the form $t \mapsto P + t \sqrt{n} (P_{n} - P)$. So, the choice of tangent set seems to be the most natural one. Of course, larger tangent sets will also deliver the desired results but establishing differentiability under them can be harder. The definition does not impose any linear structure on $\mathcal{T}_{P}$ and $t$ is restricted to be non-negative. This feature of the definition is analogous to the idea of directional derivative in Shapiro1990 which has been shown to be sufficient for showing the validity of the Delta Method (see Shapiro1990), and turns out to be enough to also carry out our analysis. See also FangSantos2014 and cho_white_2017 for further references, examples and discussion. $\triangle$
remarkThe class $\mathcal{C}$ determines the degree of uniformity of the limit and thus defines different notions of differentiability. It is known that common notions of differentiability can be obtained from different choices of $\mathcal{C}$; see dudley2010frechet for a discussion. We now enumerate a few: \begin{enumerate} • $\tau$-Gateaux: $\mathcal{C}$ is the class of finite subsets of $\mathcal{T}_{P}$; denoted by $\mathcal{J}_{\tau}$. • $\tau$-Hadamard: $\mathcal{C}$ is the class of $\tau$-compact subsets of $\mathcal{T}_{P}$; denoted by $\mathcal{H}_{\tau}$. • $\tau$-Frechet: $\mathcal{C}$ is the class of $\tau$-bounded subsets of $\mathcal{T}_{P}$; denoted by $\mathcal{E}_{\tau}$. \end{enumerate} $\triangle$

The following is the main result of this section.

theoremSuppose there exists a class $\mathcal{C} \subseteq 2^{\mathcal{T}}$ such that $\boldsymbol{\psi}$ is $DIFF(P,\mathcal{C})$ and \footnote{Implicit in the differentiability condition lies the assumption that for any $Q \in \mathcal{T}_{P}$, $t \mapsto P + t Q \in \mathbb{D}_{\psi}$. For this to hold, it is sufficient that $P$ belongs to the algebraic interior of $\mathcal{M}$ relative to $\mathcal{T}_{P}$. However, by inspection of the proof of the Theorem, it can be seen that this assumption is not really needed since we only consider curves of the form $t \mapsto P + t_{n} a_{n}(P_{n}-P)$ where $(t_{n},a_{n})$ are such that the curve equals $P_{n}$ which is in $\mathbb{D}_{\psi}$.} \begin{enumerate} • For any $\epsilon>0$, there exists a $U \in \mathcal{C}$ and a $N$ such that $\mathbf{P} \left( \sqrt{n} (P_{n} - P) \in U \right) \geq 1 - \epsilon$ for all $n \geq N$. \end{enumerate} Then, there exists a $\boldsymbol{k} \colon \mathbb{N} \rightarrow \mathbb{K}$ for which $\boldsymbol{\psi}$ satisfies $GAL(\boldsymbol{k})$ and $\lim_{n \rightarrow \infty} \mathbf{k}(n) = \infty$.
proofSee Appendix (ref).

It is easy to check that the influence of the regularization implied by the theorem is given by the sequence of $L^{2}_{0}(P)$ mappings, $(\varphi_{k}(P))_{k \in \mathbb{N}}$ where

align*[align* omitted — 77 chars of source]

While the theorem shows existence of a sequence of tuning parameters for which generalized asymptotic linearity holds, it is silent about how to construct such sequence; we discuss this in Section (ref).

remark[Heuristics of the Proof] The proof is straightforward and is comprised of two steps. First, it is shown that $\boldsymbol{\psi}$ satisfies $GAL(\mathbf{k})$ for any fixed $k$, i.e., $\mathbf{k}(n) = k$. Showing this result is analogous to showing that the regularization is asymptotic linear in the sense defined above, thus, it suffices to show that the reminder of the linear approximation is asymptotically negligible for each fixed $k$, i.e., \begin{align} \eta_{k}(P_{n}-P) = o_{P}(n^{-1/2}). \end{align} This is a standard condition for “plug-in" estimators (e.g. VdV2000), and the restriction over the class $\mathcal{C}$ and the definition of differentiability imply it. In some cases, however, it might be straightforward to verify condition ((ref)) directly or by other means. Second, a diagonalization argument is used to show existence of a diverging sequence. $\triangle$
remarkA common way of using Theorem (ref) is by finding a class $\mathcal{S}$ that is P-Donsker, which implies that $(\sqrt{n}(P_{n} - P))_{n \in \mathbb{N}}$ is $||.||_{\mathcal{S}}$-compact (see Lemma (ref) in Appendix (ref)), and ensuring $||.||_{\mathcal{S}}$-Hadamard differentiability; e.g. VdV2000 Ch. 20. dudley2010frechet proposes an alternative way of using this result by showing that $\sqrt{n}(P_{n} - P)$ belongs, with high probability, to bounded p-variation sets, so the relevant notion of differentiability is Frechet differentiability (under the p-variation norm). $\triangle$

Examples

We now present a series of examples. The goal here is not to break new ground but to use classical examples to illustrate the conditions and components of Theorem (ref).

The following example is the celebrated integrated square density, for which BickelRitov1990 showed that even though the efficiency bound is finite, no estimator converges at root-n rate; thereby illustrating that in some circumstances studying the local shape of $\psi$ can be quite misleading. Our approach does not suffer from this criticism since it directly captures the (local) behavior of the estimator at hand. It is also general enough to encompass many of the proposed estimators in the literature, including “leave-one-out" types.

example[Integrated Square Density] The parameter of interest in this case is given by \begin{align*} P \mapsto \psi(P) = \int p(x)^{2} dx. \end{align*} The model is defined as the class of probability measures, $P$, with Lebesgue density, $p$, such that $p \in L^{\infty}(\mathbb{R})$ and \begin{align} |p(x+t) - p(x) | \leq C(x) |t|^{\varrho}, \forall t,x \in \mathbb{R}, \end{align} with $C \in L^{2}(\mathbb{R})$ and $\varrho \in (0,0.5)$. This restriction is rather mild and is similar to those used in the literature, e.g. BickelRitov88, HallMarron87 and powell1989semiparametric. The mapping $\psi$ is well-defined over the model, but not when evaluated at the empirical probability distribution, $P_{n}$, since $P_{n}$ does not have a density; it thus needs to be regularized. We consider a class of regularizations given by \begin{align} P \mapsto \psi_{k}(P) = \int (\kappa_{k} \star P)(x) P(dx), \forall k \in \mathbb{K}, \end{align} where $\kappa$ is a kernel such that $\int |\kappa(u)| |u|^{\varrho} du < \infty$, and $t \mapsto \kappa_{k}(t) \equiv k \kappa(k t)$. Thus, $1/k$ acts as the bandwidth for each $k \in \mathbb{K}$ which is a tuning set in $\mathbb{R}_{++}$. Depending on the form of $\kappa$ this regularization encompasses many estimators proposed in the literature. For instance, when $\kappa = \rho + \lambda (\rho - \rho \star \rho)$ with some $\lambda \in \mathbb{R}$ and some kernel $\rho$, and, for any $k>0$, $z \mapsto \hat{p}_{k}(z) = \frac{1}{n} \sum_{i=1}^{n} \rho_{k}(Z_{i} - z)$, it follows that\footnote{Details of the claims 1-3 and 1'-3' below are shown in the Appendix (ref).} \begin{enumerate} • For $\lambda = 0$, the implied estimator is $n^{-1}\sum_{i=1}^{n} \hat{p}_{k}(Z_{i}) = n^{-2}\sum_{i,j} \rho_{k}(Z_{i} - Z_{j})$. • For $\lambda = -1$, the implied estimator is $\int (\hat{p}_{k}(z))^{2} dz = n^{-2}\sum_{i,j} (\rho \star \rho)_{k}(Z_{i} - Z_{j})$. • For $\lambda = 1$, the implied estimator is $2n^{-1}\sum_{i=1}^{n} \hat{p}_{k}(Z_{i}) - \int (\hat{p}_{k}(z))^{2} dz = n^{-2}\sum_{i,j} (2\rho_{k} - (\rho \star \rho)_{k})(Z_{i} - Z_{j})$. \end{enumerate} The first two estimators are standard; the third estimator is inspired by the one considered in NeweyHsiehRobbins2004, wherein $\kappa$ is a twicing kernel. Moreover, the formalization in display ((ref)) captures commonly used “leave-one-out" estimators by simply imposing $\kappa(0)=0$. For instance, the “leave-one-out" versions of the estimators 1-3 are given by \begin{enumerate} • For $\lambda = 0$, the implied estimator is $n^{-2}\sum_{i \ne j} \rho_{k}(Z_{i} - Z_{j})$. • For $\lambda = -1$, the implied estimator is $n^{-2}\sum_{i \ne j} (\rho \star \rho)_{k} (Z_{i} - Z_{j})$. • For $\lambda = 1$, the implied estimator is $n^{-2}\sum_{i \ne j} (2\rho_{k} - (\rho \star \rho)_{k})(Z_{i} - Z_{j})$ \end{enumerate} These estimators are essentially the ones considered by GineNickl2008 and HallMarron87 (see also POWELL1996 and references therein); the estimator 3' is also a somewhat simplified version of the one considered in BickelRitov88. We now show that Definition (ref) is satisfied by our class of regularizations and also establish a rate for the remainder term, $\eta_{k}(P_{n} - P)$ which is used to verify for which sequence of tuning parameter condition (ref) holds. \begin{proposition} For any $P \in \mathcal{M}$, the regularization defined in expression ((ref)) is $DIFF(P,\mathcal{E}_{||.||_{LB}})$. For each $k \in \mathbb{K}$, \begin{align*} Q \mapsto D\psi_{k}(P)[Q] = 2\int (\kappa_{k} \star P)(z) Q(dz), \end{align*} and $Q \mapsto \eta_{k}(Q) = \int (\kappa_{k} \star Q)(z) Q(dz)$ is such that there exists a $L_{k} < \infty $ such that $|\eta_{k}(Q)| \leq L_{k} ||Q||^{2}_{LB}$ for all $Q \in ca(\mathbb{Z})$. \end{proposition} \begin{proof} See Appendix (ref). \end{proof} This proposition implies that for each $k \in \mathbb{K}$, $\psi_{k}$ is $||.||_{LB}$-Frechet differentiable, and since $LB$ is P-Donsker, the conditions in Theorem (ref) are met. The influence is given by $z \mapsto \varphi_{k}(P)(z) \equiv 2\{ ( \kappa_{k} \star P)(z) - E_{P}[( \kappa_{k} \star P)(Z)]\}$, and since $\sup_{k} ||\varphi_{k}(P)||_{L^{2}(P)} \leq 2 ||p||_{L^{\infty}(\mathbb{R})} ||\kappa ||_{L^{1}(\mathbb{R})}$ (see Lemma (ref) in Appendix (ref)), the natural scaling for GAL is $\sqrt{n}$. $\triangle$

Next, we consider the NPIV example. It is not hard to see that the influence of $\boldsymbol{\gamma}$ will be given by $z \mapsto \int D \psi_{k}(P)^{\ast}[\pi](z) - E_{P}[D \psi_{k}(P)^{\ast}[\pi](Z)]$ provided $D \psi_{k}(P) : \mathcal{T}_{P}^{\ast} \rightarrow L^{2}([0,1])$ and its adjoint $D \psi_{k}(P)^{\ast} : L^{2}([0,1]) \rightarrow \mathcal{T}_{P}^{\ast}$ exists ($\mathcal{T}_{P}^{\ast}$ is the dual of $\mathcal{T}_{P}$). For sieve-based and penalization-based regularization schemes, we characterize $D \psi_{k}(P)^{\ast}$ and show how its standard deviation can be used to appropriately scale the estimator to obtain a generalized asymptotic linear representation regardless of whether the parameter is root-n estimable or not. This last result, illustrates how our method can be used to generalize the approach proposed in ChenPouzo2015 to general regularizations. As a by-product, we extend the results in AckerbergChenHahnLiao2014 and link the influence function of the sieve-based regularization to simpler, fully parametric, misspecified GMM models.

example[NPIV (cont.): The sieve-based Case] We study the sieve-based regularization approach, which is constructed using two basis for $L^{2}([0,1])$, $(u_{k},v_{k})_{k \in \mathbb{N}}$, and two indices $k \mapsto (J(k),L(k))$ such that \begin{align*} & (g,x) \mapsto T_{k,P}[g](x) = (u^{J(k)}(x))^{T} Q_{uu}^{-1} E_{P} \left[ u^{J(k)}(X) g(W) \right], \\ & x \mapsto r_{k,P}(x) = (u^{J(k)}(x))^{T} Q_{uu}^{-1} E_{P}[u^{J(k)}(X) Y],\\ & \mathcal{R}_{k,P} = (\Pi_{k}^{\ast}T^{\ast}_{k,P} T_{k,P} \Pi_{k} )^{-1} \end{align*} where $u^{k}(x) \equiv (u_{1}(x),...,u_{k}(x))$, $v^{k}(w) \equiv (v_{1}(w),...,v_{k}(w))$, $\Pi_{k} : L^{2}([0,1]) \rightarrow lin\{ v^{L(k)} \} \subseteq L^{2}([0,1])$ is the projection operator, $g \mapsto \Pi_{k}[g] = (v^{L(k)})^{T}Q^{-1}_{vv} \int v^{L(k)}(w) g(w) dw$, and $Q_{uu} \equiv E_{Leb}[u^{k}(X)(u^{k}(X))^{T}]$, $Q_{uv} \equiv E_{P}[u^{k}(X)(v^{k}(W))^{T}]$ and $Q_{vv} \equiv E_{Leb}[v^{k}(W)(v^{k}(W))^{T}]$. The next proposition proves differentiable of the regularization $\boldsymbol{\gamma}$ and provides the expression for the derivative. \begin{proposition} For any $P \in \mathcal{M}$, the sieve-based regularization $\boldsymbol{\gamma}$ is DIFF$(P,\mathcal{E}_{||.||_{LB}})$. For each $k \in \mathbb{N}$, \begin{align*} Q \mapsto D \gamma_{k}(P)[Q] = \int D \psi_{k}(P)^{\ast}[\pi](z) Q(dz) \end{align*}where \begin{align}\notag D \psi_{k}(P)^{\ast}[\pi](y,w,x) =& (y - \psi_{k}(P)(w)) (u^{J(k)}(x))^{T} Q^{-1}_{uu} Q_{uv} (Q^{T}_{uv}Q^{-1}_{uu}Q_{uv})^{-1} E_{Leb}[v^{L(k)}(W) \pi(W)] \\ \notag &+ \left\{ E_{P}[(\psi(P)(W) - \psi_{k}(P)(W))(u^{J(k)}(X))^{T}] Q^{-1}_{uu} u^{J(k)}(x) \right. \\ & \left. \times (v^{L(k)}(w))^{T} (Q^{T}_{uv}Q^{-1}_{uu}Q_{uv})^{-1} E_{Leb}[v^{L(k)}(W) \pi(W)] \right\}. \end{align} And, for each $k \in \mathbb{N}$, the reminder of $\gamma_{k}$, $\eta_{k}$, is such that $ |\eta_{k}(\zeta)| = o(||\zeta||_{LB})$ for any $\zeta \in \mathbb{D}_{\psi}$.\footnote{The “o" function may depend on $k$.} \end{proposition} \begin{proof} See Appendix (ref). \end{proof} Even though expression for $D \psi_{k}(P)^{\ast}[\pi]$ may look cumbersome, it has an intuitive interpretation: It is identical to the influence function of the parameter $\int \theta^{T} v^{L(k)}(w)\pi(w) dw$ where $\theta$ is the estimand of a misspecified linear GMM model where the “endogenous variables" are $v^{L(k)}(W)$ and the “instrumental variables" are $u^{J(k)}(X)$; cf. HallInoue2003. The first term in the RHS of expression (ref) also has an intuitive interpretation: It is the influence function of the parameter $\int \theta^{T} v^{L(k)}(w)\pi(w) dw$ but in well-specified linear GMM model. The proposition implies that for the “fix-$k$" case, expression (ref) is the proper influence function to be considered. However, one can ask whether as $k$ diverges, the second term (the one in curly brackets) in RHS of expression (ref) can be ignored. To shed light on this matter, it is convenient to use operator notation for expression (ref): \begin{align}\notag D\psi^{\ast}_{k}(P)[\pi](y,w,x) = & T_{k,P}\mathcal{R}_{k,P}\Pi_{k}[\pi](x) \times (y - \psi_{k}(P)(w))\\ & + \mathcal{R}_{k,P }\Pi_{k}[\pi](w) \times T_{k,P}[\psi(P) - \psi_{k}(P)](x) \end{align} (we derive this equality in expression (ref) in Appendix (ref)). The term $T_{k,P}[\psi(P) - \psi_{k}(P)]$ is multiplied by $\mathcal{R}_{k,P}\Pi_{k} [\pi]$, which is different to $T_{k,P}\mathcal{R}_{k,P}\Pi_{k}[\pi]$ --- the factor multiplying $(y- \psi_{k}(P)(w))$. If $\pi \in Range (T_{P})$ both multiplying factors converge to bounded quantities as $k$ diverges. Thus, since $T_{k,P}[\psi(P) - \psi_{k}(P)]$ vanishes, the first summand in the RHS of expression (ref) “asymptotically dominates" the second one. This is framework considered in AckerbergChenHahnLiao2014. However, if $\pi \notin Range (T_{P})$ --- and thus $\gamma(P)$ is not root-estimable (see SeveriniTripathi2012) --- the situation is more subtle and without additional assumptions it is not clear which term in expression (ref) dominates. The reason is that the aforementioned multiplying factors will no longer converge to a bounded quantity, and moreover, the rate of growth of $T_{k,P}\mathcal{R}_{k,P}\Pi_{k}[\pi]$ can can be dominated by the rate of $\mathcal{R}_{k,P}\Pi_{k} [\pi]$. For this last case of $\pi \notin Range (T_{P})$, the results closest to ours are those in ChenPouzo2015 wherein the influence function for slower than root-n sieve estimators is derived. Their expression for the influence function is simpler than ours, but this arises from a different set of assumptions and, more importantly, a different approach that directly focus on expressions for “diverging $k$". $\triangle$
example[NPIV (cont.): The Penalization-based Case] We study the penalization-based regularization case given by \begin{align*} & (x,g) \mapsto T_{k,P}[g](x) \equiv \int \kappa_{k}(x'-x) \int g(w) P(dw,dx') \\ & x \mapsto r_{k,P}(x) \equiv \int \kappa_{k}(x'-x) \int y P(dy,dx')\\ & \mathcal{R}_{k,P} = (T_{k,P}^{\ast} T_{k,P} + \lambda_{k} I )^{-1} \end{align*} where $\kappa_{k}(\cdot) = k \kappa(k \cdot )$ and $\kappa$ is a smooth, symmetric around 0 pdf. As opposed to the previous case, there is no obvious link to a “simpler" problem like GMM and thus it is not obvious a-priori what the influence function would be and what the proper scaling should be when $\gamma(P)$ is not root-n estimable. Theorem (ref) suggests $D \psi_{k}^{\ast}(P)[\pi]$ and $\sqrt{n/Var_{P}( D \psi_{k}^{\ast}(P)[\pi])}$ as the influence function and scaling factor resp.; the next proposition characterizes it. \begin{proposition} For any $P \in \mathcal{M}$, the Penalization-based regularization $\boldsymbol{\gamma}$ is DIFF$(P,\mathcal{E}_{||.||_{LB}})$. For each $k \in \mathbb{N}$, $D \gamma_{k}(P)[\zeta] = \int D \psi_{k}(P)^{\ast}[\pi](z) \zeta(dz)$, where \begin{align} D \psi_{k}^{\ast}(P)[\pi](y,w,x) = & \mathcal{K}_{k}^{2}T_{P}(T_{P}^{\ast}\mathcal{K}_{k}^{2}T_{P} + \lambda_{k} I )^{-1}[\pi](x) \times (y - \psi_{k}(P)(w)) \\ \notag & + \lambda_{k} (T_{P}^{\ast}\mathcal{K}_{k}^{2}T_{P} + \lambda_{k} I )^{-1}[\pi](w) \times \mathcal{K}^{2}_{k} T_{P} (T_{P}^{\ast}\mathcal{K}_{k}^{2}T_{P} + \lambda_{k} I )^{-1} [\psi_{id}(P)] (x). \end{align} where $\mathcal{K}_{k}$ is the convolution operator $g \mapsto \mathcal{K}_{k}[g] = \kappa_{k} \star g$. And, for each $k \in \mathbb{N}$, the reminder of $\gamma_{k}$, $\eta_{k}$, is such that $ |\eta_{k}(\zeta)| = o( ||\zeta||_{LB})$ for any $\zeta \in \mathbb{D}_{\psi}$.\footnote{The “o" function may depend on $k$.} \end{proposition} \begin{proof} See Appendix (ref). \end{proof} If $\pi \in Range(T_{P})$, then the variance term converges to $|| T_{P}[v^{\ast}](X) (Y - \psi(P)(W)) ||^{2}_{L^{2}(P)} = E_{P}[(T_{P}(T_{P}^{\ast}T_{P})^{-1}[\pi](X))^{2} E_{P}[(Y - \psi(P)(W))^{2}|X] ]$ as $k$ diverges, where $v^{\ast} \equiv (T_{P}^{\ast}T_{P})^{-1}[\pi]$. The function $(y,w,x) \mapsto T_{P}[v^{\ast}](x) (y - \psi(P)(w)) $ is the influence function one would obtained by employing the methods in AiChen2007 (with identity weighting) and $v^{\ast}$ is the Riesz representer of the functional $w \mapsto \int \pi(w) g(w) dw$ using their weak norm $||T_{P}[\cdot]||_{L^{2}(P)}$. If $\pi \notin Range(T_{P})$, the variance diverges, and, as in the sieve case, without additional assumptions it is not clear which term dominates the variance term $Var_{P}( D \psi_{k}^{\ast}(P)[\pi])$, as $k$ diverges. This case illustrates how our results can be used to extend the results in ChenPouzo2015 for irregular sieve-based estimators to more general regularization schemes. $\triangle$

Data-driven Choice of Tuning Parameter and Undersmoothing

Theorem (ref) implies existence of a $n \mapsto k(n)$ such that\footnote{The display hold provided $\liminf_{n \rightarrow \infty} ||\varphi_{k(n)}(P)||_{L^{2}(P)} >0 $ For the applications we have in mind, this restriction is natural and non-binding. Our results are not designed for cases where $\lim_{k \rightarrow \infty} ||\varphi_{k}(P)||_{L^{2}(P)} = 0$; this case can be handled separately --- and rather easily --- since both the approximation error and the rate of $k \mapsto \eta_{k}(P_{n}-P)$ decrease as $k$ increases.}

align[align omitted — 321 chars of source]

I.e., the asymptotic behavior of the regularized estimator --- once scaled and centered --- is characterized by a term due to the approximation error and a stochastic term. For obtaining asymptotic distributions, it is common practice to try to find sequences $(k(n))_{n}$ satisfying Theorem (ref) for which the approximation term in expression ((ref)) vanishes. Unfortunately it is known that such sequences do not always exists at this level of generality; e.g. BickelRitov88 and HallMarron87. In view of this remark it is natural to seek choices of tuning parameter that make the terms in the RHS of expression ((ref)) as “small as possible". Such choices will guarantee that GAL and the asymptotic negligibility of the approximation error both hold when possible, and otherwise, will at least yield good rates of convergence for $\psi_{k}(P_{n}) - \psi(P)$.

The result in this section shows that the data-driven way of choosing tuning parameters described in Section (ref) satisfies this property. For each $n \in \mathbb{N}$, the data-driven choice of tuning parameter is of the form $ \tilde{k}_{n} = \arg\min\{ k \colon k \in \mathcal{L}_{n}(\Lambda) \}$ for a suitable chosen function $ k \mapsto \Lambda(k)$. In section (ref), the relevant function was $k \mapsto 4 \bar{\delta}_{k}(r^{-1}_{n})$; in this case, however, the structure of the problem is different. In particular, in addition to the reminder term $(\eta_{k})_{k}$ implied by differentiability and the scaled approximation error, there is the additional term given by $n^{-1/2} \sum_{i=1}^{n} \frac{ \varphi_{k(n)}(P)(Z_{i}) }{|| \varphi_{k(n)}(P) ||_{L^{2}(P)} }$. The following assumption introduces the quantities to construct $(\Lambda_{k})_{k}$. For each $n$, let $\mathbb{K}_{n}$ be the grid defined as in Section (ref).

assumptionThere exists a $(n,k) \mapsto \bar{\delta}_{j,k}(n)$ for $j \in \{1,2\}$ non-decreasing and a $N \in \mathbb{N}$ such that \begin{enumerate} • $\sup_{k \in \mathbb{K}_{n}} \frac{ |\eta_{k}(P_{n} - P) |}{\bar{\delta}_{1,k}(n)} \leq 1$ wpa1-$P$. • $\frac{|\mathbb{K}_{n}|}{\sqrt{n}} \sup_{k' \geq k~in~\mathbb{K}_{n}} \frac{||\varphi_{k'}(P) - \varphi_{k}(P)||_{L^{2}(P)}}{\bar{\delta}_{2,k'}(n)} \leq 1$ for all $n \geq N$. \end{enumerate}

As the proof of Lemma (ref) in Appendix (ref) suggests, the sequence that defines our tuning parameter, for each $n \in \mathbb{N}$, is given by $k \mapsto \Lambda(k) \equiv 4 (\bar{\delta}_{1,k}(n) + \bar{\delta}_{2,k}(n))$.

remark[Discussion of Assumption (ref)] Part (ii) implies that $(\bar{\delta}_{2,k}(n))_{n,k}$ acts as a growth rate for an object that, on the hand, involves the complexity of $\mathbb{K}_{n}$ --- given by $|\mathbb{K}_{n}|$ --- and on the other hand, involves the “length" of $\mathbb{K}_{n}$ --- measured by $k \mapsto ||\varphi_{k}(P)||_{L^{2}(P)}$. In cases where $(||\varphi_{k}(P)||_{L^{2}(P)})_{k}$ is uniformly bounded, the “length" of $\mathbb{K}_{n}$ (measured by $k \mapsto ||\varphi_{k}(P)||_{L^{2}(P)}$) is also uniformly bounded and part (ii) boils down to $\frac{|\mathbb{K}_{n}|}{\sqrt{n}} \leq \inf_{k \in \mathbb{K}_{n}} \delta_{2,k}(n))$. Part (i) implies that $(\bar{\delta}_{1,k}(n))_{n,k}$ also acts as the growth rate, but of a very different quantity: The reminder term of GAL, uniformly on $k \in \mathbb{K}_{n}$. To shed more light on Assumption (ref), suppose there exists a norm $||.||_{\mathcal{S}}$ such that \begin{enumerate} • There exists, for each $k \in \mathbb{K}$ a modulus of continuity $\bar{\eta}_{k} : \mathbb{R}_{+} \rightarrow \mathbb{R}_{+}$ such that $\eta_{k}(Q) = \bar{\eta}_{k}(||Q||_{\mathcal{S}})$. • There exists a real-valued positive diverging sequence $(r_{n})_{n}$ such that $||P_{n} -P||_{\mathcal{S}} = o_{P}(r^{-1}_{n})$. \end{enumerate} Condition C1 states that $\eta_{k}$ is continuous with respect to some norm $||.||_{\mathcal{S}}$ and C2 ensure convergence of $P_{n}$ to $P$ under this norm. These conditions are analogous to the assumptions used to show Theorem (ref). Under these conditions it is easy to see that Part (i) follows by choosing $\bar{\delta}_{1,k}(n) = \bar{\eta}_{k}(r^{-1}_{n})$, which acts as $\delta_{k}(r^{-1}_{n})$ in Theorem (ref), and, in particular, it does not depend on the grid. Part (ii), however, is not necessarily implied by this choice. If the growth rate of the reminder, $ \bar{\eta}_{k}(r^{-1}_{n})$, is small compared to $\frac{|\mathbb{K}_{n}|}{\sqrt{n}} \sup_{k' \geq k~in~\mathbb{K}_{n}} ||\varphi_{k'}(P) - \varphi_{k}(P)||_{L^{2}(P)}$ then part (ii) requires that $\bar{\delta}_{2,k}(n)$ to be larger than the latter, i.e., $\bar{\delta}_{2,k}(n) \geq \frac{|\mathbb{K}_{n}|}{\sqrt{n}} \sup_{k' \geq k~in~\mathbb{K}_{n}} ||\varphi_{k'}(P) - \varphi_{k}(P)||_{L^{2}(P)} $. Below we illustrate how to verify these assumptions in the context of Example (ref).$\triangle$
propositionSuppose all the conditions of Theorem (ref) hold, and Assumption (ref) holds. Then \begin{align*} \left| \frac{\sqrt{n} ( \psi_{\tilde{k}_{n}}(P_{n}) - \psi(P) ) }{|| \varphi_{\tilde{k}_{n}}(P)||_{L^{2}(P)}} - \frac{n^{-1/2} \sum_{i=1}^{n} \varphi_{\tilde{k}_{n}}(P)(Z_{i})}{|| \varphi_{\tilde{k}_{n}}(P)||_{L^{2}(P)}} \right| = O_{P} \left( C^{2}_{n} \sqrt{n} \inf_{k \in \mathbb{K}_{n}} \left\{ \frac{\bar{\delta}_{1,k}(n) + \bar{\delta}_{2,k}(n) + \bar{B}_{k}(P) }{|| \varphi_{k}(P)||_{L^{2}(P)}} \right\} \right), \end{align*} where $C_{n} \equiv \sup_{k',k~in~\mathbb{K}_{n}} \frac{|| \varphi_{k}(P)||_{L^{2}(P)} }{|| \varphi_{k'}(P)||_{L^{2}(P)} } $.
proofSee Appendix (ref).

The rate in the proposition is --- up to $C^{2}_{n}$ factor --- the minimum value of the sum of two terms: $\sqrt{n} \frac{\bar{\delta}_{1,k}(n) + \bar{\delta}_{2,k}(n)}{||\varphi_{k}(P)||_{L^{2}(P)}}$, that controls the reminder term of GAL and another one, $\sqrt{n} \frac{B_{k}(P)}{||\varphi_{k}(P)||_{L^{2}(P)}} $, that controls the approximation error term. Therefore, if there exists a choice of tuning parameter for which both these terms are asymptotically negligible, our result implies that

align*[align* omitted — 237 chars of source]

That is, the asymptotic distribution of $ \sqrt{n} \frac{ \psi_{\tilde{k}_{n}}(P_{n}) - \psi(P)}{|| \varphi_{\tilde{k}_{n}}(P)||_{L^{2}(P)} }$ is given that of $n^{-1/2} \sum_{i=1}^{n} \frac{ \varphi_{\tilde{k}_{n}}(P)(Z_{i})}{|| \varphi_{\tilde{k}_{n}}(P)||_{L^{2}(P)} }$. On the other hand, if no such sequence exists, the proposition readily implies a rate of convergence of the form $\left| \frac{ \psi_{\tilde{k}_{n}}(P_{n}) - \psi(P) }{|| \varphi_{\tilde{k}_{n}}(P)||_{L^{2}(P)}} \right| = O_{P} \left( n^{-1/2} + C^{2}_{n} \inf_{k \in \mathbb{K}_{n} } \left\{ \frac{\bar{\delta}_{1,k}(n) + \bar{\delta}_{2,k}(n) + \bar{B}_{k}(P) }{|| \varphi_{k}(P)||_{L^{2}(P)}} \right\} \right)$.

The sequence $(C_{n})_{n}$ quantifies the discrepancy of $k \mapsto || \varphi_{k}(P)||_{L^{2}(P)}$ within the grid. In cases where Assumption (ref)(ii) holds for all $k'$ and $k$ in $\mathbb{K}_{n}$, it readily follows that $C_{n} = 1 + |\mathbb{K}_{n}|^{-1} \sup_{k \in \mathbb{K}_{n}} \bar{\delta}_{2,k}(n)/||\varphi_{k}(P)||_{L^{2}(P)}$.

In order to shed more light on these expressions and the assumptions, we applied our results to the estimation of the integrated square PDF (example (ref)). In this setting, GineNickl2008 already provide a data-driven method to choose the bandwidth which is akin to ours. In fact, our method can be viewed as generalization of theirs to general regularizations, and the example illustrates that, at least in their setup, we do not have to pay an extra price for the added generality.

example[Integrated Square Density (cont.)] In this example the relevant tuning parameter is the bandwidth of the kernel, so we let $k \mapsto k^{-1}$, and as the grid, $\mathbb{K}_{n}$, we use the one proposed by Gine and Nickl (GineNickl2008), i.e., $\mathbb{K}_{n} = \{ k \colon k^{-1} \in \mathcal{H}_{n} \}$ where \begin{align*} \mathcal{H}_{n} = \left\{ h \in \left[ \frac{(\log n )^{4}}{n^{2}} , \frac{1}{n^{1-\delta}} \right] \colon h_{0} = \frac{1}{n^{1-\delta}}, h_{1} = \frac{\log n}{n}, h_{2} = \frac{l^{-1}_{n}}{n}, h_{k+1} = h_{k}/a, \forall k=2,3,... \right\} \end{align*} where $a > 1$, $(l_{n})_{n}$ diverges to infinity slower than $\log n$ and $l^{-1}_{n} < \log n$ and $\delta>0$ is arbitrary close to 1; in particular $\delta>0$ is such that $2\varrho < \frac{1+\delta}{2(1-\delta)}$. Of importance to our analysis are the fact that $|\mathbb{K}_{n}| = O(\log n)$ and that for sufficiently large $n$, any two consecutive elements in $\mathcal{H}_{n}$ are such that $h_{k+1}/h_{k} \leq 1/a$. The following lemma suggests an expression for the functions $(n,k) \mapsto \bar{\delta}_{i,k}(n)$ for $i \in \{1,2\}$. \begin{lemma} For any $M>0$, there exists a $N$ such that for all $n \geq N$, \begin{align*} \sup_{h' \leq h in \mathcal{H}_{n}} ||\varphi_{1/h}(P) - \varphi_{1/h'}(P)||_{L^{2}(P)} \leq 4 ||C||_{L^{2}(P)} h^{\varrho} E_{|\kappa|}[ |U|^{m+\varrho}] . \end{align*} where the function $C$ is the one in expression (ref) in Example (ref), and \begin{align*} \mathbf{P} \left( \sup_{h \in \mathcal{H}_{n}} \sqrt{n} |\eta_{1/h}(P_{n} - P) |\geq M \left( \frac{\kappa(0)}{\sqrt{n} h} + \frac{1}{ \sqrt{n h} } \right) \right) \leq |\mathcal{H}_{n}| M^{-1}. \end{align*} \end{lemma} \begin{proof} See Appendix (ref). \end{proof} Therefore, $\{(n,k) \mapsto \bar{\delta}_{i,k}(n)\}_{i=1,2}$ can be chosen as \begin{align*} (n,k) \mapsto \bar{\delta}_{1,k} \equiv (\log n)^{3} \frac{ k\kappa(0) + \sqrt{k} }{ n}, and (n,k) \mapsto \bar{\delta}_{2,k} \equiv \frac{(\log n)^{3} k^{-(m+\varrho)}}{\sqrt{n}}. \end{align*} The lemma and this display illustrate the different nature of Assumptions (ref)(i)(ii). Part (i) bounds the reminder of the linear approximation and it increases with $k$ and decreases with $n$; this is reflected in the term $\frac{ \left( k\kappa(0)+ \sqrt{k} \right) }{ n }$ in the display. Part (ii) on the other hand essentially requires that the bandwidths in the grid $\mathcal{H}_{n}$ are not “too far apart". In particular, it depends on the size of the bandwidths in $\mathcal{H}_{n}$; this is reflected in the term $k^{-\varrho} $ in the display. It follows that $\sup_{n \in \mathbb{N}} C_{n} < \infty$ because there exists a constant $C>1$ such that $k \mapsto ||\varphi_{k}(P)||_{L^{2}(P)} \in [C^{-1},C]$ and is continuous for all $k \geq 1$. We verified that all assumptions of Proposition (ref) hold. Moreover, Proposition (ref) in Appendix (ref) implies that $h \mapsto \bar{B}_{1/h}(P) = O(h^{2 \varrho})$. Thus, the rate of Proposition (ref) is given by\\ $\inf_{h \in \mathcal{H}_{n}} \{ (\log n)^{3} \left( \frac{ \kappa(0)/h + 1/\sqrt{h} }{\sqrt{n}} + h^{\varrho} \right) + \sqrt{n} h^{2 \varrho} \} $. In fact, given our choice of $\mathcal{H}_{n}$ and $\delta$, some straightforward algebra shows that, at least for large $n$, the infimum over $\mathbb{K}_{n}$ and be replaced by the infimum over $\mathbb{R}_{+}$. Therefore, \begin{align*} \left| \frac{\sqrt{n} ( \psi_{\tilde{k}_{n}}(P_{n}) - \psi(P) ) }{|| \varphi_{\tilde{k}_{n}}(P)||_{L^{2}(P)}} - \frac{n^{-1/2} \sum_{i=1}^{n} \varphi_{\tilde{k}_{n}}(P)(Z_{i})}{|| \varphi_{\tilde{k}_{n}}(P)||_{L^{2}(P)}} \right| = \left\{ \begin{array}{cc} O_{P} \left( \left(\frac{\log n }{n} \right)^{\frac{4\varrho}{1 + 4\varrho} - 0.5} \right) & if \kappa(0) = 0 \\ O_{P} \left( \left(\frac{\log n }{n} \right)^{\frac{2\varrho}{1 + 2\varrho} - 0.5} \right) & if \kappa(0) > 0 \end{array} \right. \end{align*} For the case $\kappa(0) = 0$, we replicate the results by GineNickl2008: if $\varrho > 0.25$, the reminder is negligible and root-n consistency follows, otherwise the optimal convergence rate is achieved. $\triangle$

Conclusion

We propose an unifying framework to study the large sample properties of regularized estimators that extends the scope of the existing large sample theory for “plug-in" estimators to a large class containing regularized estimators. Our results suggest that the large sample theory for regularized estimators does not constitute a large departure from the existing large sample theory for “plug-in" estimators, in the sense that both are based on local properties of the mappings used for constructing the estimator. This last observation indicates that other large sample results developed for “plug-in" estimators can also be extended to the more general setting of regularized estimators; e.g., estimation of the asymptotic variance of the estimator and, more generally, inference procedure like the bootstrap. We view this as a potentially worthwhile avenue for future research.