Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Improved Density and Distribution Function Estimation
tabular[tabular omitted — 544 chars of source]
}
abstractGiven additional distributional information in the form of moment restrictions, kernel density and distribution function estimators with implied generalised empirical likelihood probabilities as weights achieve a reduction in variance due to the systematic use of this extra information. The particular interest here is the estimation of densities or distributions of (generalised) residuals in semi-parametric models defined by a finite number of moment restrictions. Such estimates are of great practical interest, being potentially of use for diagnostic purposes, including tests of parametric assumptions on an error distribution, goodness-of-fit tests or tests of overidentifying moment restrictions. The paper gives conditions for the consistency and describes the asymptotic mean squared error properties of the kernel density and distribution estimators proposed in the paper. A simulation study evaluates the small sample performance of these estimators. Supplements provide analytic examples to illustrate situations where kernel weighting provides a reduction in variance together with proofs of the
results in the paper.
{\bf Keywords}: Moment conditions, residuals, mean squared error, bandwidth.
{\bf MSC 2010 subject classifications}: Primary 62G07, secondary 62G05, 62G20.
\numberwithin{remark}{section}
\numberwithin{equation}{section}
\numberwithin{assumption}{section}
\numberwithin{theorem}{section}
Introduction
In many statistical and economic applications, additional distributional information about the data observation $d_{z}$-vector $z$ may be available in the form of moment restrictions on its distribution. These constraints may arise from a particular economic or physical law, e.g., chen1997, be implied by estimating equations, qin1994, or correspond to known population moments of another observable random vector correlated with $z$, e.g., in survey samples with auxiliary population information available from census data, e.g., chen1993 and qin1994.
The primary purpose of the paper is to explore the advantages of this additional information for the estimation of the density and distribution function of a scalar residual-like function of $z$ which may depend on unknown parameters.
To this end, let $g(z,\beta)$ denote a $d_{g}$-vector of known functions of the data observation $d_{z}$-vector $z\in\mathcal{Z}$ and the $d_{\beta}$-vector $\beta\in\mathcal{B}$ of parameters where the sample space $\mathcal{Z}\subseteq\mathbbm{R}^{d_{z}}$ and parameter space $\mathcal{B}\subset\mathbbm{R}^{d_{\beta}}$ with $d_{\beta}\leq d_{g}$. The moment indicator vector $g(z,\beta)$ will form the basis for inference in the following discussion and analysis. In particular, it is assumed that the true value $\beta_{0}$ taken by $\beta$ uniquely satisfies the population unconditional moment equality condition
equation[equation omitted — 77 chars of source]
where $\operatorname{E}[\cdot]$ denotes expectation taken with respect to the true population probability law of $z$. The true parameter value $\beta_{0}$ is generally unknown, but can also be fully or partially known in particular applications.
Models specified in the form of unconditional moment restrictions (ref) convey partial information about the distribution $F^{z}$ of $z$ and are ubiquitous in economics; see, e.g., the monographs hall2005 and matyas1999. Many other commonly used models lead to estimators that can be reformulated as solutions to a set of moment restrictions. Clearly, models given by conditional moment restrictions imply (ref). Traditionally, such models are estimated by the generalised method of moments (GMM). However, the performance of GMM estimators and associated test statistics is
often poor in finite samples, which has lead to the development of a number of (information-theoretic) alternatives to GMM.
This paper focuses on the class of generalised (G) empirical likelihood (EL) estimators, which has attractive large sample properties; see, e.g., newey2004, smith1997,smith2011, and parente2014 for a recent review. Special cases of GEL include EL, owen1988,owen1990, qin1994, exponentially tilting (ET), corcoran1998, kitamura1997, imbens1998, and continuous-updating (GMM) estimators (CUE), hansen1996; see also Euclidean EL, antoine2007. Of these estimators, EL has the attractive property of being Bartlett-correctable; see chen2007.
When the parameter vector $\beta_{0}$ is overidentified by the moment restriction (ref), i.e., $d_{\beta}<d_{g}$, these constraints generally carry useful additional information about $F^{z}$. Given a random sample $z_{i}$, $i=1,\ldots,n$, of observations on $z$, such information is captured by the associated (G)EL implied probabilities $\pi_{i}$, $i=1,\ldots,n$, which enable a nonparametric description of $F^{z}$ satisfying the moment condition (ref) given by the estimator $F_{\pi}^{z}(z)=\sum_{i=1}^{n}\pi_{i}\mathbbm{1}\{z_{i}\leq z\}$, where $\mathbbm{1}\{\cdot\}$ denotes the indicator function, back1993, qin1994. In the absence of the moment information (ref) or when $\beta_{0}$ is just identified, $d_{\beta}=d_{g}$, $F_{\pi}^{z}(z)$ reduces to the empirical distribution function (EDF) $F_{n}^{z}(z)=n^{-1}\sum_{i=1}^{n}\mathbbm{1}\{z_{i}\leq z\}$. In general, if $d_{\beta}<d_{g}$, $F_{\pi}^{z}(z)$ is a more efficient estimator of $F^{z}$ than the EDF $F_{n}^{z}(z)$ reflecting the value of the overidentifying information in (ref). This observation suggests therefore that estimation of the functionals of $F^{z}$, $T(F^{z})$, by $T(F_{\pi}^{z})$ rather than $T(F_{n}^{z})$ will be similarly more efficient. Indeed this is the case when estimating expectations of certain known functions of $z$, see brown1998. A similar advantage is apparent for EL estimation of quantile functions with known $\beta_{0}$, e.g., chen1993 and zhang1995, general EL-based quantile estimation, yuan2014, and EL-based kernel estimation of a univariate density function, e.g., chen1997 and zhang1998.
The concern of this paper is with efficient kernel estimation of the probability density (p.d.f.) and distribution (c.d.f.) functions of a scalar-valued function $u(z,\beta_{0})$ of the data observation $z$ with either known or unknown parameter vector $\beta_{0}$. The former case, when $\beta_{0}$ is known, is the classical situation briefly mentioned above. The central case of interest, when $\beta_{0}$ is unknown, is estimation of the p.d.f and c.d.f. of an error term based on the estimated residuals. Such estimates are routinely computed by practitioners and are used for both visual diagnostics, e.g., potentially revealing omitted structure such as multimodality or other features of interest, and formal diagnostic tests, e.g., goodness-of-fit and tests of parametric assumptions on the error distribution. The importance of obtaining residual density estimates with
good (higher order) properties can hardly be understated. Yet, as discussed below, simply applying standard kernel estimators with default bandwidths to estimated residuals may result in an inconsistent p.d.f. or c.d.f. estimators as further conditions on the kernel function and bandwidth are generally required. Similar conclusions have been reached elsewhere in related literature on residual density estimation in nonparametric regression and other settings; see, e.g.,
ahmad1992, cheng2004, kiwitt2008, gyorfi2012 and the discussion and references in bott2013.
When $\beta_{0}$ is known, kernel density and distribution function estimators exploiting the (G)EL implied probabilities instead of the uniform EDF $n^{-1}$ weights achieve a reduction of higher order variance due to the systematic use of the extra moment information in (ref). The efficiency gains are first order asymptotically in the c.d.f. case and
second order for p.d.f. estimation. In contradistinction, for residual p.d.f. and c.d.f. estimation, such gains will not always be realised. One can, however, expect efficiency gains from the knowledge that the mean of residuals is zero.
The outline of the paper is as follows. Section (ref) briefly describes (G)EL estimation and the associated (G)EL implied probabilities. The main results concerning p.d.f. and c.d.f. estimators are given in Sections (ref) and (ref) for both known and unknown $\beta_{0}$ cases. The finite sample performance of the proposed estimators is evaluated via a simulation study reported in Section (ref). Section (ref) concludes.
Supplement Supplement (ref): Proofs and (ref): Examples in the Supplementary Information respectively
details some additional assumptions for and the proofs of the results in the main text and
analyses a number of examples to illustrate the the properties of the estimators developed in the paper.
Generalised Empirical Likelihood
The GEL class of estimators for $\beta_{0}$ is defined in terms of a real valued scalar carrier function $\rho:\mathcal{V}\mapsto\mathbbm{R}$ that is concave on an open interval $\mathcal{V}$ containing zero with derivatives $\rho^{(j)}(v)=\mathrm{d}^{j}\rho(v)/\mathrm{d}v^{j}$ and $\rho_{j}=\rho^{(j)}(0)$, $j=1,2,\ldots$, normalized without loss of generality such that $\rho_{1}=\rho_{2}=-1$. The special cases $\rho(v)=\ln(1-v)$ for $\mathcal{V}=(-\infty, 1)$, $\rho(v)=-\exp(v)$ and $\rho(v)=-v^{2}/2-v$ correspond to EL, ET and CUE respectively and are all members of the cressie1984 family where $\rho(v)=-(1+\gamma v)^{(\gamma +1)/\gamma}/(\gamma+1)$.
Given a random sample $z_{i}$, $i=1,\ldots,n$, of size $n$ of observations on the $d_{z}$-dimensional vector $z$, let $g_{i}(\beta)=g(z_{i},\beta)$, $g_{i}=g_{i}(\beta_{0})$, and $G_{i}(\beta)=\partial g(z_{i},\beta)/\partial\beta^{\top}$, $G_{i}=G_{i}(\beta_{0})$, $i=1,\ldots,n$. Also let $\Lambda _{n}(\beta)=\{\lambda : \lambda^{\top}g_{i}(\beta)\in\mathcal{V}, i=1,\ldots,n\}$. The GEL criterion $P_{n}(\beta,\lambda)$ is defined by $P_{n}(\beta,\lambda)=n^{-1}\sum_{i=1}^{n}\rho(\lambda^{\top}g_{i}(\beta))-\rho(0)$, with $\lambda$ a $d_{g}$-vector of auxiliary parameters, each element of which corresponding to an
element of the moment function vector $g(z,\beta)$; for members of the cressie1984 family of power divergence criteria $\lambda$ is the Lagrange multiplier vector associated with imposition of the moment restriction (ref). The GEL estimator $\hat{\beta}$ is the solution to the saddle point problem
equation[equation omitted — 163 chars of source]
If Supplement (ref): Assumptions (ref) and (ref) are satisfied, in particular, the
population Jacobian $G=\operatorname{E}[\partial g(z,\beta_{0})/\partial\beta^{\top}]$ and variance $\Omega =\operatorname{E}[g(z,\beta_{0})g(z,\beta_{0})^{\top}]$ matrices are full column rank and positive definite respectively, then all GEL estimators share the same first order large sample properties, see, e.g., newey2004, i.e., $n^{1/2}(\hat{\beta}
-\beta _{0})\xrightarrow{d}N(0,\Sigma)$, achieving the semiparametric efficiency lower bound $\Sigma =(G^{\top}\Omega^{-1}G)^{-1}$, chamberlain1987. Furthermore, if the additional Supplement (ref): Assumption (ref) is imposed, defining $H=\Sigma G^{\top}\Omega^{-1}$ and $P=\Omega^{-1}-\Omega^{-1}G(G^{\top}\Omega^{-1}G)^{-1}G^{\top}\Omega^{-1}$, the second order bias of $\hat{\beta}$ is
$\operatorname{E}[\hat{\beta}]-\beta _{0}=n^{-1}H\zeta_{\lambda}+\mathit{O}_{}(n^{-2})$,
where
equation[equation omitted — 172 chars of source]
with $c_{\rho}=1+\rho_{3}/2$ and $a$ a $d_{g}$-vector with elements $a^{j}=\mathop{\mathrm{tr}}(\Sigma\operatorname{E}[\partial^{2}g^{j}(z,\beta_{0})/\partial\beta\partial\beta^{\top}])$, $j=1,\ldots,d_{g}$; see newey2004.
remarkThe validity of the higher order bias and variance calculations, and hence the validity of the results reported below can be
formally justified by that of an Edgeworth expansion of order $\mathit{o}_{}(n^{-1})$ for the distribution of GEL parameter estimators. If $z$ is continuously distributed, appropriate conditions may be found in bhattacharya1978 for general smooth functions of sample moments and kundhi2012 for Edgeworth expansions for (G)EL estimators. If some of
the elements of $z$ are discretely distributed, jensen1989 provides appropriate conditions.
For given $\beta$, the auxiliary parameter estimator is defined by $\lambda(\beta)=\operatorname*{argmax}_{\lambda\in\Lambda_{n}(\beta)}P_{n}(\beta,\lambda)$. Whenever the constraint in $\lambda\in\Lambda_{n}(\beta)$ is not binding, $\lambda(\beta)$ solves the first-order conditions \\ $n^{-1}\sum_{i=1}^{n}\rho^{(1)}(\lambda(\beta)^{\top}g_{i}(\beta))g_{i}(\beta)=0$.
The GEL implied probabilities are then
equation*[equation* omitted — 204 chars of source]
The sample moment constraint $\sum_{i=1}^{n}\pi_{i}(\beta)g_{i}(\beta)=0$ holds whenever the first order conditions for $\lambda
(\beta)$ hold. In what follows, $\hat{\pi}_{i}=\pi_{i}(\hat{\beta})$, $i=1,\ldots,n$, corresponds to the solution $\hat{\lambda}=\lambda(\hat{\beta})$, and, if $\beta_{0}$ is known, $\tilde{\pi}_{i}=\pi_{i}(\beta_{0})$, $i=1,\ldots,n$, with auxiliary parameter estimator $\tilde{\lambda}=\lambda(\beta_{0})$. The generic notation $\pi_{i}$, $i=1,\ldots,n$, is used whenever the distinction is unnecessary.
remarkProperties of the GEL implied probabilities relevant to the subsequent developments are summarized in Supplement (ref): Lemmas (ref) and (ref). Although $\pi_{i}(\beta)$, $i=1,\ldots,n$, sum to unity and are positive if $\pi_{i}(\beta)g_{i}(\beta)$ is small uniformly in $i$, they are not guaranteed to be non-negative. The shrinkage estimator $\pi_{i}^{\varepsilon}=(\pi_{i}+\varepsilon_{n})/\sum_{j=1}^{n}(\pi_{j}+\varepsilon_{n})$, $i=1,\ldots,n$, where $\varepsilon_{n}=-\min[\min_{1\leq i\leq n}\pi_{i},0]$, see antoine2007, smith2011,
ensures non-negativity $\pi_{i}^{\varepsilon}\geq0$, $i=1,\ldots,n$, and $\sum_{i=1}^{n}\pi_{i}^{\varepsilon}=1$. Alternative solutions relevant to probability density and distribution function estimation respectively are discussed in Sections (ref) and (ref).
remarkThe implied probabilities were given for EL by owen1988, for ET by kitamura1997, for quadratic $\rho(\cdot)$ by back1993, and for the general case in the 1992 working paper version of brown2002;
see also smith1997. For any function $a(z,\beta)$ and GEL estimator $\hat{\beta}$ the implied probabilities can be used to form a semiparametrically efficient estimator $\sum_{i=1}^{n}\hat{\pi}_{i}a(z_{i},\hat{\beta})$ of $\operatorname{E}[a(z,\beta_{0})]$ as in brown1998.
GEL-Based Density Estimation
Suppose the p.d.f. $f(\cdot)$ of the scalar random variable $u=u(z,\beta _{0})$ is of interest, where the scalar function $u:\mathcal{Z}\times\mathcal{B}\mapsto\mathcal{U}\subseteq\mathbbm{R}$ is known up to the parameter vector $\beta_{0}$.
Let $\mathcal{N}$ denote an open neighbourhood of $\beta_{0}$.
assumptionFor all $\beta\in\mathcal{N}$ there exists a function $v:\mathcal{Z}\times\mathcal{B}\mapsto\mathcal{V}\subseteq\mathbbm{R}^{d_{z}-1}$ such that the vector of functions $(u(z,\beta),\: v(z,\beta)^{\top})^{\top}$ is a bijection between $\mathcal{Z}$ and $\mathcal{U}\times \mathcal{V}$.
remarkEquivalently Assumption (ref) may be restated as requiring that for every $\beta\in \mathcal{N}$ there exists a bijection between $z$ and some $d_{z}$-vector $w=w(z,\beta)$ such that, given $\{w^{j}(z,\beta)\}_{j=2}^{d_{z}}$, $u(z,\beta)$ and $w^{1}(z,\beta)$ are bijective. That is to say, $z$ may be solved for uniquely given values for $u$, $v$ and $\beta$.
remarkA function $u(z,\beta)$ satisfying Assumption (ref) may be thought of as defining a generalised residual in the sense of cox1968 and loynes1969, with $\hat{u}_{i}=u(z_{i},\hat{\beta})$, $i=1,\ldots,n$, the estimated residuals. Of course, other possibilities of interest are included, e.g., estimating the density of an element of $z$ subject to the extra information available in the moment condition (ref).
Known \texorpdfstring{$\beta_{0}$}{beta}
Suppose that $u_{i}=u(z_{i},\beta _{0})$, $i=1,\ldots,n$, are observed. Then the classical kernel density estimator for the p.d.f. $f$ of $u=u(z,\beta_{0})$ can be employed; viz.
equation[equation omitted — 139 chars of source]
where $k_{b}(x)=k(x/b)/b$, $k(\cdot)$ is a kernel function and $b=b_{n}>0$ is a bandwidth sequence; see rosenblatt1956 and parzen1962. The estimator $\tilde{f}$ (ref) will serve as a benchmark for later comparisons.
The properties of $\tilde{f}$ are well known and can be formally established under different combinations of smoothness and integrability conditions on the kernel $k$ and density $f$; see, e.g., rao1983. A standard set of such conditions is given in Assumption (ref) below. If $k$ is square integrable, but not absolutely integrable, as is the case for the sinc kernel, conditions such as those in tsybakov2009 can be imposed.
Let $R(k)=\int_{-\infty}^{\infty}k(x)^{2}\mathrm{d}x$ for any square integrable function $k$; the limits of integration are omitted whenever there is little scope for confusion. Also let $f^{(j)}(u)=\mathrm{d}^{j}f(u)/\mathrm{d}u^{j}$ for any $j$th order differentiable function $f$.
assumption(a)(i) $\sup_{-\infty<x<\infty}\lvertk(x)\rvert<\infty$, $\int\lvertk(x)\rvert\mathrm{d}x<\infty$,
$\int k(x)\mathrm{d}x=1$, and $\lim_{\lvertx\rvert\to\infty}\lvertxk(x)\rvert =0$;
(ii) $k$ is a $(2r)$th order kernel, i.e., an even function such that, for some $r\geq1$, $\mu_{0}(k)=1$, $\mu_{j}(k)=0$, $j=1,\ldots,2r-1$, and $\mu_{2r}(k)<\infty$, where $\mu_{j}(k)=\int x^{j}k(x)\mathrm{d}x$; (iii) $R(k)<\infty$;
(b) $f(\cdot)$ is $s$ times continuously differentiable and $R(f^{(j)})<\infty$, $j=0,1,\ldots,s$.
(c) as $n\to\infty$, $b\to0$ and $nb\to\infty$.
remarkIf Assumption (ref)(a)(i) holds, then by Supplement (ref): Lemma (ref), $\operatorname{E}[\tilde{f}(u)]\to f(u)$ as $b\to0$ at all points $u$ of continuity of $f$ and if, in addition, Assumption (ref)(c) holds, then the mean squared error (MSE), $\operatorname{MSE}[\tilde{f}(u)]=\operatorname{E}[(\tilde{f}(u)-f(u))^{2}]\to 0$ as $n\to\infty$; see, e.g., parzen1962.
remarkHigher order approximations to $\operatorname{MSE}[\tilde{f}(u)]$ can be obtained if $f$ is sufficiently smooth. See, e.g. rao1983, wand1995 or pagan1999. The idea of using higher order kernels as a bias reduction technique originates at least as far back as bartlett1963.
Let $1\leq r<\infty$. Suppose that Assumptions (ref)(a)(ii), (ref)(b) with $s=2r+2$, (ref)(c) together with $\mu_{2r+2}(k)<\infty$ and $\int x^{2}k(x)^{2}\mathrm{d}x<\infty$ hold. Then
align*[align* omitted — 211 chars of source]
Hence,
equation[equation omitted — 210 chars of source]
remarkIf $k$ is a $(2r)$th order kernel and Assumption (ref)(b) holds with $s=2r$, the remainder term in
$\operatorname{E}[\hat{f}(u)]$ is $\mathit{o}_{}(b^{2r})$. The $\sim n^{-1}$ term is kept explicit with $\mathit{O}_{}$ remainder for reasons that will become apparent below.
The mean integrated squared error (MISE), $\operatorname{MISE}[\tilde{f}]=\operatorname{E}[\int(\tilde{f}(u)-f(u))^{2}\mathrm{d}u]$, is a commonly used global measure of performance. The optimal bandwidth is then defined as that value of $b>0$ minimising MISE, or an approximation thereof. In particular, the asymptotically optimal bandwidth is defined as the value $b^{\ast}$ minimising the two leading terms in the expansion
equation[equation omitted — 206 chars of source]
i.e., $b^{\ast}=cn^{-1/(4r+1)}$ where $c=[(2r)!^{2}R(k)/(4r\mu_{2r}(k)^{2}R(f^{(2r)})]^{1/(4r+1)}$. The asymptotically optimal MISE is thereby
equation*[equation* omitted — 159 chars of source]
remarkIf $k$ is of order greater than two, it necessarily takes negative values. Hence $\tilde{f}$ (ref) itself need not be a density function. Note, however, that the positive part estimator, $\tilde{f}^{+}(u)=\max[\tilde{f}(u),0]$ has MSE at most equal to $\operatorname{MSE}[\tilde{f}(u)]$. Further modifications that ensure integration to unity can be applied as described in glad2003.
The GEL-based kernel density estimator incorporates the information embedded in the moment restriction (ref) replacing the sample EDF weights $n^{-1}$ in the construction of $\tilde{f}(u)$ (ref) by the
implied probabilities $\tilde{\pi}_{i}$, $i=1,\ldots,n$; viz.
equation[equation omitted — 156 chars of source]
remarkThe GEL-based kernel density estimator $\tilde{f}_{\rho}(u)$ (ref) is the estimator of $f(u)$ obtained from the revised GEL criterion $\sum_{i=1}^{n}[\rho(\eta (f(u)-k_{b}(u-u_{i}))+\lambda^{\top}g_{i}(\beta))-\rho(0)]/n$ with the implicit moment condition $\operatorname{E}[k_{b}(u-u_{i})]=f(u)$ and associated auxiliary parameter $\eta$; see smith2011.
remarkIf the validity of the moment restriction (ref) is in doubt, a pre-test can be conducted using the GEL-based criterion (ref) paralleling the classical likelihood ratio test; see, e.g.,
kitamura1997, imbens1998 and smith1997,smith2011. For example, under the null hypothesis that (ref) holds for some unique $\beta_{0}\in\mathcal{B}$, the normalised GEL criterion (ref) evaluated at the estimated parameters, $2nP_{n}(\hat{\beta},\hat{\lambda})$, is asymptotically chi-square distributed with $d_{g}-d_{\beta}$ degrees of freedom. The parametric null hypothesis of known $\beta_{0}=\beta^{0}$ can be tested at the $\alpha$ level using the critical region $\{2nP_{n}(\beta^{0},\tilde{\lambda})\geq \chi_{d_{\beta}}^{2}(\alpha)\}$.
To describe the properties of GEL-based kernel density estimator $\tilde{f}_{\rho }(u)$ (ref), the shorthand notation, e.g., $\operatorname{E}[g_{i}|u]=\operatorname{E}[g(z,\beta_{0})|\{z: u(z,\beta_{0})=u\}]$, for conditional expectations given $u$ is adopted.
theoremIf Supplement (ref): Assumptions (ref)--(ref) and (ref)(a)(i) and (c) are satisfied, then $\tilde{f}_{\rho}(u)=\tilde{f}(u)+\mathit{o}_{p}(1)$ for all $u$ such that $f(u)<\infty$.
If, in addition, Assumption (ref) is satisfied, then
\begin{align}
\mkern-10mu\operatorname{E}[\tilde{f}_{\rho}(u)] & = \operatorname{E}[\tilde{f}(u)]
+ n^{-1}c_{\rho}\left(-\operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}|u]
+ \operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}g_{i}^{\top}]\Omega^{-1}\operatorname{E}[g_{i}|u]
+d_{g}\right)f(u) + \mathit{o}_(n^{-1}), \\
\mkern-10mu\operatorname{Var}[\tilde{f}_{\rho}(u)] &= \operatorname{Var}[\tilde{f}(u)] - n^{-1}\operatorname{E}[g_{i}|u]^{\top}\Omega^{-1}\operatorname{E}[g_{i}|u]f(u)^{2}
+ \mathit{o}_(n^{-1}).
\end{align}
Thus, the estimators $\tilde{f}$ and $\tilde{f}_{\rho}$ are asymptotically first-order equivalent, and the asymptotically optimal bandwidth for $\tilde{f}_{\rho}$ is identical to that of $\tilde{f}$, i.e., $b^{\ast}$.
Whenever $c_{\rho}=0$, as is the case for (G)EL with $\rho_{3}=-2$ , e.g., EL, the $n^{-1}$ bias term in (ref) vanishes. In general, provided the bandwidth does not go to zero faster than $n^{-1/(2r)}$, and certainly when $b=b^{\ast}\sim n^{-1/(4r+1)}$, this bias term is at most third order. Its contribution to MISE is via the integrated squared bias (ISB)
align*[align* omitted — 389 chars of source]
with the $\mathit{O}_{}(n^{-1}b^{2r})$ term generally non-zero and either positive or negative. With the asymptotically optimal bandwidth, $n^{-1}(b^{\ast})^{2r}\sim n^{-3/2+1/(8r+2)}$, which approaches $n^{-3/2}$ arbitrarily closely as $r$ increases, whereas the leading terms in $\operatorname{MISE}[\tilde{f}(\cdot;b^{\ast})]$ becomes arbitrarily close to $n^{-1}$.
As long as $\operatorname{E}[g_{i}|u]\neq 0$, the GEL-based estimator $\tilde{f}_{\rho}$ enjoys a second-order reduction in variance due to the $n^{-1}$ term in (ref), which does not depend on the choice of GEL carrier function $\rho(\cdot)$. Hence
equation*[equation* omitted — 243 chars of source]
While this reduction is negligible asymptotically, the leading term in $\operatorname{MISE}[\tilde{f}]$ approaches zero only a little more slowly than $n^{-1}$. Hence the effect could be substantial in small samples.
Unknown \texorpdfstring{$\beta_{0}$}{beta 0}
Suppose now that $\beta_{0}$ is unknown. Then, after substitution of the estimators $\hat{u}_{i}=u(z_{i},\hat{\beta})$ for $u_{i}$, $i=1,\ldots,n$, in $\tilde{f}$ and $\tilde{f}_{\rho}$ in (ref) and (ref), the analogous estimators of $f(u)$ are
align[align omitted — 297 chars of source]
respectively. Because $u_{i}$, $i=1,\ldots,n$, are not directly observable, the behaviour of the estimation error $\hat{u}_{i}-u_{i}$, $i=1,\ldots,n$, needs to be constrained with additional restrictions imposed on $k$ and $b$. Assumption (ref) gives a set of mild sufficient conditions, see, e.g., vanryzin1969 and ahmad1992; similar conditions have also been considered
in, e.g., cheng2005 and kiwitt2008.
assumption(a) $k$ is H\"{o}lder continuous with exponent $0<\tau\leq1$;
(b) there exists $d(z)\geq0$ with $\operatorname{E}[d(z)^{\tau}]<\infty$ such that, for some $0<\alpha\leq1$, $\lvertu(z,\beta)-u(z,\beta_{0})\rvert\leq d(z)\lVert\beta-\beta_{0}\rVert^{\alpha}$ for all $z$ and for all $\beta\in\mathcal{N}$;
(c) $b\to0$ and $n^{\alpha\tau/2}b^{1+\tau}\to\infty$ as $n\to\infty$.
The uniform $\alpha$-H\"{o}lder condition Assumption (ref)(b) on $u(z,\beta)$, also known as a Lipschitz condition of order $\alpha$, is an appropriate way to quantify the `degree of continuity' of $u(z,\beta)$; see zygmund2003. Many kernels used in practice are Lipschitz continuous, and hence satisfy Assumption (ref)(a) with $\tau=1$. For example, a kernel that satisfies Assumption (ref)(a) for any $0<\tau\leq\gamma$ but not for $\gamma<\tau\leq1$ is $k(x) = (1+\gamma)(1-\lvertx\rvert)^{\gamma}/2$ if $\lvertx\rvert\leq 1$ and $0$ otherwise, yielding the Bartlett (triangular) kernel if $\gamma=1$. Assumption (ref)(c) is important as it prevents the bandwidth from being too small. Intuitively, if $b$ is very small, the kernel $k_{b}(u-\hat{u}_{i})$ is very narrowly centered around the incorrect value $\hat{u}_{i}$ potentially excluding the true value $u_{i}$; see, e.g., silverman1986 for a generic illustration. Assumption (ref)(c) requires $nb^{4}\to\infty$ regardless of the values of $\tau$ and $\alpha$ and $b=n^{-1/4}$ is the fastest rate achievable when $\alpha=\tau=1$. Note that the optimal bandwidth $b^{\ast}$ is excluded if $[\alpha(4r+1)-2]\tau <2$.
Under these conditions, Theorem (ref) establishes that the differences between the kernel density estimators $\hat{f}$ (ref) and $\hat{f}_{\rho }$ (ref) and their counterparts $\tilde{f}$ (ref) and $\tilde{f}_{\rho}$ (ref) based on observable $u_{i}$, $i=1,\ldots,n$, are negligible asymptotically.
theoremIf Supplement (ref): Assumptions (ref)--(ref) and (ref) are satisfied, then
$\hat{f}(u)=\tilde{f}(u)+\mathit{o}_{p}(1)$ and
$\hat{f}_{\rho}(u) = \sum_{i=1}^{n}\hat{\pi}_{i}k_{b}(u-u_{i})+\mathit{o}_{p}(1)$
for all $u$.
If, in addition, Assumption (ref)(a)(i) holds, $\hat{f}_{\rho}(u) = \tilde{f}(u)+\mathit{o}_{p}(1)$ a.e.
To obtain higher order expansions for the mean and variance of $\hat{f}(u)$ (ref) and $\hat{f}_{\rho}(u)$ (ref) requires a further strengthening of the assumptions. Let $\nabla u(z,\beta)$ and $\nabla^{2}u(z,\beta)$ denote respectively the $d_{\beta}$-vector and $d_{\beta}\times d_{\beta}$ matrix of the first and second derivatives of $u(z,\beta)$ with respect to $\beta$. Also let $\nabla u_{i}=\nabla u(z_{i},\beta _{0})$ and $\nabla^{2}u_{i}=\nabla^{2}u(z_{i},\beta _{0})$.
assumption(a) $k$ is twice differentiable and $k^{(2)}$ is H\"{o}lder continuous with exponent $0<\tau\leq1$, $k$, $k^{(1)}$, and $k^{(2)}$ are absolutely integrable; $\lim_{\lvertx\rvert\to\infty}\lvertx^{s}k^{(s-1)}(x)\rvert=0$, $s=1,2,3$, and
$\int k(x)\mathrm{d}x=1$;
(b) $u(z,\beta)$ is twice differentiable for all $\beta\in\mathcal{N}$,
$\operatorname{E}[\lVert\nabla u_{i}\rVert^{4}]<\infty$, $\operatorname{E}[\lVert\nabla^{2}u_{i}\rVert^{4}]<\infty$, and there exists $d(z)\geq0$ with $\operatorname{E}[d(z)^{4}]<\infty$ such that, for some $0<\alpha\leq1$, $\lVert\nabla^{2}u(z,\beta)-\nabla^{2}u(z,\beta_{0})\rVert\leq d(z)\lVert\beta-\beta_{0}\rVert^{\alpha}$ for all $z$ and for all $\beta\in\mathcal{N}$;
(c) $b\to0$ as $n\to\infty$, $n^{\tau/2}b^{3+\tau}\to\infty$, and $n^{\alpha/2}b^{5/4}\to\infty$;
(d)(i) $f$ is twice differentiable;
(ii) $\operatorname{E}[\nabla u_{i}|u]$, $\operatorname{E}[\nabla^{\top}u_{i}Hg_{i}|u]$, and $\operatorname{E}[\nabla^{2}u_{i}|u]$ are differentiable in $u$ and
$\operatorname{E}[\nabla u_{i}\nabla^{\top}u_{i}|u]$ is twice differentiable in $u$;
(iii) $\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u] f(u)\}/\mathrm{d}u$, $\mathrm{d}\{\operatorname{E}[\nabla^{\top}u_{i}Hg_{i}|u]f(u)\}/\mathrm{d}u$, $\mathrm{d}\{\operatorname{E}[\nabla^{2}u_{i}|u]f(u)\}/\mathrm{d}u$, and $\mathrm{d}^{2}\{\operatorname{E}[\nabla u_{i}\nabla^{\top}u_{i}|u]f(u)\}/\mathrm{d}u^{2}$ are absolutely integrable functions of $u$.
Assumption (ref)(a)(b) implies Assumption (ref)(a)(b) holds with $\alpha=\tau=1$ with the requirement in Assumption (ref)(c) rendered as $n^{1/2}b^{2}\to\infty$. Note that Assumption (ref)(a) also implies Assumption (ref)(a)(i). Assumption (ref)(d) imposes additional smoothness and integrability conditions on $f$ and $u(z,\beta)$. Assumption (ref)(c) is much stronger than Assumption (ref)(c) requiring $nb^{8}\to\infty$ regardless of the values of $\tau$ and $\alpha$ thereby prohibiting the asymptotically optimal bandwidth $b^{\ast}$ when $k$ is a second order kernel. For $r\geq2$, $b^{\ast}$ is permissible as long as $\tau>6/(4r-1)$ and $\alpha>5/(8r+2)$. Note that, if $\alpha>5/16$, $n^{\tau/2}b^{3+\tau}\to\infty$ implies $n^{\alpha/2}b^{5/4}\to\infty$.
theoremIf Supplement (ref): Assumptions (ref)--(ref), (ref), and (ref) are satisfied, then
$\operatorname{E}[\hat{f}(u)]=\operatorname{E}[\tilde{f}(u)]+n^{-1}\delta(u)+\mathit{o}_{}(n^{-1})$
and $\operatorname{E}[\hat{f}_{\rho}(u)]= \operatorname{E}[\tilde{f}(u)]+n^{-1}\delta(u)+n^{-1}\delta_{\rho}(u)+\mathit{o}_{}(n^{-1})$,
where
\begin{align}
\notag
\delta(u) & = \mathrm{d}\{\operatorname{E}[\nabla^{\top}u_{i}Hg_{i}|u]f(u)\}/\mathrm{d}u - \zeta_{\lambda}^{\top}H^{\top}[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u] f(u)\}/\mathrm{d}u]\\
&\quad +\tfrac{1}{2}\mathop{\mathrm{tr}}(\Sigma [\mathrm{d}^{2}\{\operatorname{E}[\nabla u_{i}\nabla^{\top}u_{i}|u]f(u)\}/\mathrm{d}u^{2}
-\mathrm{d}\{\operatorname{E}[\nabla^{2}u_{i}|u]f(u)\}/\mathrm{d}u])\\
\intertext{and}
\delta_{\rho}(u) & = (-c_{\rho}\operatorname{E}[g_{i}^{\top}Pg_{i}|u] + c_{\rho}(d_{g}-d_{\beta}) +\zeta_{\lambda}^{\top}P\operatorname{E}[g_{i}|u] )f(u).
\end{align}
Also
\begin{align}
\notag
\operatorname{Var}[\hat{f}(u)] &= \operatorname{Var}[\tilde{f}(u)]
+ n^{-1}[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u] f(u)\}/\mathrm{d}u]^{\top}\Sigma[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u] f(u)\}/\mathrm{d}u]\\
&\quad + n^{-1}2[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u] f(u)\}/\mathrm{d}u]^{\top}H\operatorname{E}[g_{i}|u]f(u) + \mathit{o}_(n^{-1}), \\
\operatorname{Var}[\hat{f}_{\rho}(u)] &= \operatorname{Var}[\hat{f}(u)] -n^{-1}\operatorname{E}[g_{i}|u]^{\top}P\operatorname{E}[g_{i}|u]f(u)^{2} + \mathit{o}_(n^{-1}).
\end{align}
remarkThe general conclusion of Theorem (ref) for both bias and variance is identical to that of Theorem (ref), i.e., the estimation effects of substituting $\hat{u}_{i}$ for $u_{i}$, $i=1,\ldots,n$, and the GEL implied probabilities $\hat{\pi}_{i}$ for $\tilde{\pi}_{i}$, $i=1,\ldots,n$, are both of order $n^{-1}$. The bias term in $\hat{f}$ induced by estimation is similar to that for $\tilde{f}$ in Theorem (ref) except that $P$ in (ref)
replaces $\Omega^{-1}$ in (ref)
and two extra terms enter via $\zeta_{\lambda}$, viz. $-a$ and $\operatorname{E}[G_{i}Hg_{i}]$ in (ref). These latter terms appear in the higher order asymptotic bias $n^{-1}H(-a+\operatorname{E}[G_{i}Hg_{i}])$ for the infeasible GEL estimator based on the optimal moment indicator vector $G^{\top}\Omega^{-1}g(z,\beta)$, see newey2004, and are inherited by all GEL estimators. Unlike Theorem (ref) for the known $\beta_{0}$ case, this term no longer vanishes for a particular choice of a carrier function $\rho$. The replacement of $\Omega^{-1}$ by $P$ represents the loss of information occasioned by the estimation of $\beta_{0}$. In a number of cases, the term $\operatorname{E}[g_{i}|u]^{\top}P\operatorname{E}[g_{i}|u]$ may vanish, see, e.g., Supplement (ref): Example (ref). This of course always occurs for an exactly identified model $d_{g}=d_{\beta}$ since $\hat{\pi}_{i}=n^{-1}$ and $\hat{f}_{\rho}$ (ref) and $\hat{f}$ (ref) are identical. However, see Supplement (ref): Example (ref), in general $\hat{f}_{\rho}$ may still enjoy a second-order reduction in variance due to the systematic use of overidentifying moment information (ref).
The extra bias term $\delta (u)$ (ref) for $\hat{f}_{\rho}$ and those terms appearing in $\operatorname{Var}[\hat{f}(u)]$ (ref) primarily arise due to the substitution of $\hat{u}_{i}$ for $u_{i}$, $i=1,\ldots,n$. Supplement (ref): Examples (ref) and (ref) examine these terms in more detail for regression on a constant and (G)EL with a constant and zero mean condition respectively.
Here, although $\int [\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u] f(u)\}/\mathrm{d}u]^{\top}\Sigma[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u] f(u)\}/\mathrm{d}u]\mathrm{d}u$ is non-negative, the term $\int [\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u] f(u)\}/\mathrm{d}u]^{\top}H\operatorname{E}[g_{i}|u]f(u) \mathrm{d}u$ can be negative, as can be the ISB term due to the additional $\delta(u)$ (ref).
Bias Correction
While the contribution from the $n^{-1}$ bias terms to MISE is of a lower order than the contribution from the variance terms, the effect of bias can be substantial in small and moderate samples, potentially offsetting any reduction in variance. The direction of the bias cannot of course be known a priori. Hence it may be advisable to bias-correct the density
estimates by estimating and subtracting the $n^{-1}$ bias term.
To be more specific, the bias-corrected estimates are defined as
align*[align* omitted — 188 chars of source]
where $\hat{\delta}(u)$ and $\hat{\delta}_{\rho}(u)$ are suitable (asymptotically) unbiased estimators of $\delta (u)$ (ref) and $\delta_{\rho }(u)$ (ref). The implied probabilities $\hat{\pi}_{i}$, $i=1,\ldots,n$, can be used to obtain efficient estimators of the component quantities entering $\delta(u)$ and $\delta_{\rho }(u)$ with the modifications described in glad2003 applied to ensure that the bias-corrected estimate is a density.
remarkWhen $\beta_{0}$ is known, bias-correction requires the estimation of the $n^{-1}$ term in (ref) unless $c_{\rho }=0$, i.e., $\rho_{3}=-2$.
GEL-Based Distribution Function Estimation
The results for distribution function estimation parallel those given in Section (ref) for density estimation but can be shown to hold under much weaker conditions, and so are given here separately.
Known \texorpdfstring{$\beta_{0}$}{beta}
When $u_{i}$, $i=1,\ldots,n$, are observed, the c.d.f. $F$ of $u(z,\beta_{0})$ can be estimated by
equation[equation omitted — 143 chars of source]
with $K(u)=\int_{-\infty}^{u}k(x)\mathrm{d}x$; see nadaraya1964b and watson1964. The kernel distribution function estimator (ref) can be obtained by integrating (ref) or motivated as a smoothed version of the EDF.
Assumption (ref)(a)(i) is sufficient for $\widetilde{F}$ to be an asymptotically unbiased and consistent estimator of $F$ at all continuity points of $F$ if $b\to0$ as $n\to\infty$. In addition, if $F$ is continuous then $\widetilde{F}$ converges to $F$ uniformly with probability $1$ (w.p.$1$.); see yamato1973. If $k$ satisfies Assumption (ref)(a)(ii) with $\mu_{2r+2}(k)<\infty$ for some $r\geq1$, $f$ satisfies Assumption (ref)(b) with $s=2r+1$, and $b\to0$
as $n\to\infty$ (Assumption (ref)(c) is not required here), then
align*[align* omitted — 256 chars of source]
where $\psi(k) = 2\int xK(x)k(x)\mathrm{d}x$. Hence
equation[equation omitted — 239 chars of source]
where $V_{F}={\begingroup\textstyle\int\endgroup} F(u)(1- F(u))\mathrm{d}u$.
Provided $\psi(k)>0$, the asymptotically optimal bandwidth minimising the leading terms in (ref) is $b^{\ast}=\varsigma n^{-1/(4r-1)}$, where
$\varsigma=[(2r)!^{2}\psi(k)/(4r\mu_{2r}(k)^{2}R(f^{(2r-1)}))]^{1/(4r-1)}$, and the asymptotically optimal MISE is
equation*[equation* omitted — 182 chars of source]
remarkThe leading term $n^{-1}V_{F}$ in (ref) is the integrated variance and, hence, the MISE of EDF. Thus, whenever $\psi(k)>0$ and $b$ approaches zero at least as fast as $n^{-1/(4r-1)}$, kernel smoothing provides a second order asymptotic improvement in MISE relative to the EDF. Smoothness of the kernel estimates and the reduction in MISE are
the two main reasons to prefer the kernel distribution function estimator (ref) over the EDF. The condition $\psi(k)>0$ is satisfied if $k$ is a symmetric second order kernel, since in this case $\psi(k) = \int K(x)(1-K(x))\mathrm{d}x>0$. Although $\psi(k)$ need not be positive in general, this property holds for certain classes of kernels, including Gaussian kernels of arbitrary order; see oryshchenko2017.
remarkIf $k$ is of order greater than two, $K$ is not monotone, and the resultant estimates may not themselves be distribution functions. However, if necessary, the estimates can be corrected by rearrangement; see chernozhukov2009. The MISE of the rearranged estimator can be at most equal to, and is often strictly smaller, than the MISE of the original estimator.
The modified GEL kernel distribution function estimator corresponding to $\tilde{f}_{\rho}$ (ref) which incorporates the information embedded in the moment restrictions (ref) is
equation[equation omitted — 164 chars of source]
theoremIf Supplement (ref): Assumptions (ref)--(ref) and (ref)(a)(i) are satisfied and $b\to0$ as $n\to\infty$, then $\widetilde{F}_{\rho}(u)=\widetilde{F}(u)+\mathit{o}_{p}(1)$ at all points of continuity of $F$. If, in addition, Assumption (ref) is satisfied, then
\begin{equation}
\operatorname{E}[\widetilde{F}_{\rho}(u)] = \operatorname{E}[\widetilde{F}(u)]
+\tfrac{c_{\rho}}{n}{\begingroup\textstyle\int\endgroup}_{-\infty}^{u}\left( -\operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}|t] +\operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}g_{i}^{\top}]\Omega^{-1}\operatorname{E}[g_{i}|t]+d_{g}\right)\mathrm{d}F(t)
+ \mathit{o}_(n^{-1}).
\end{equation}
If also $\lim_{\lvertx\rvert\to\infty}\lvertx^{2}k(x)\rvert=0$, then
\begin{equation}
\operatorname{Var}[\widetilde{F}_{\rho}(u)] = \operatorname{Var}[\widetilde{F}(u)]
- n^{-1}[{\begingroup\textstyle\int\endgroup}_{-\infty}^{u}\operatorname{E}[g_{i}|t]\mathrm{d}F(t)]^{\top}\Omega^{-1}[{\begingroup\textstyle\int\endgroup}_{-\infty}^{u}\operatorname{E}[g_{i}|t]\mathrm{d}F(t)]
+ \mathit{o}_(n^{-1}b).
\end{equation}
These results are qualitatively similar to Theorem (ref), the important difference being that the reduction in variance is now first-order asymptotically, whereas the contribution from the $n^{-1}$ bias term in (ref) to MISE is of order $n^{-1}b^{2r}$. Ceteris paribus, the asymptotically optimal c.d.f. bandwidth converges to zero at a faster rate than that for density estimation. Hence the additional bias effect can be
expected to be of less importance.
Unknown \texorpdfstring{$\beta_{0}$}{beta}
When $\beta_{0}$ is unknown, the analogues of $\widetilde{F}$ and $\widetilde{F}_{\rho}$ are
align[align omitted — 313 chars of source]
respectively.
theoremIf Supplement (ref): Assumptions (ref)--(ref) and (ref)(a)(i) are satisfied, Assumption (ref)(b) holds with $\tau=1$ for some $0<\alpha\leq1$, and $b\to0$ and $n^{\alpha/2}b\to\infty$ as $n\to\infty$, then
$\widehat{F}(u)=\widetilde{F}(u)+\mathit{o}_{p}(1)$,
$\widehat{F}_{\rho}(u)=\widetilde{F}(u)+\mathit{o}_{p}(1)$, and
$\widehat{F}_{\rho}(u) = \sum_{i=1}^{n}\hat{\pi}_{i}K((u-u_{i})/b)+\mathit{o}_{p}(1)$
for all $u$.
Similar to Theorem (ref), Theorem (ref) establishes that the differences between $\widehat{F}$ (ref) and $\widehat{F}_{\rho}$ (ref) and their counterparts based on observable $u_{i}$, $i=1,\ldots,n$, are negligible asymptotically. No additional requirements are placed on $k$ beyond the standard conditions in (ref)(a)(i) and the restriction on the bandwidth is thus weaker than Assumption (ref)(c).
Higher order expansions similar to those in Theorem (ref) may be obtained under the following conditions.
assumptionSuppose Assumption (ref)(b) holds.
(a) $k$ is differentiable and $k^{(1)}$ is H\"{o}lder continuous with exponent $0<\tau\leq1$, $k$ and $k^{(1)}$ are absolutely integrable, $\lim_{\lvertx\rvert\to\infty}\lvertx^{2}k(x)\rvert=0$, $\lim_{\lvertx\rvert\to\infty}\lvertx^{2}k^{(1)}(x)\rvert=0$,
and $\int k(x)\mathrm{d}x=1$;
(b) $b\to0$ as $n\to\infty$, $n^{\tau/2}b^{2+\tau}\to\infty$, and $n^{\alpha/2}b^{1/4}\to\infty$;
(c) (i) $f(u)$ and $\operatorname{E}[\nabla u_{i}\nabla^{\top}u_{i}|u]$ are differentiable in $u$;
(ii) $\mathrm{d}\{\operatorname{E}[\nabla u_{i}\nabla^{\top}u_{i}|u]f(u)\}/\mathrm{d}u$ is an absolutely integrable function of $u$.
theoremIf Supplement (ref): Assumptions (ref)--(ref), (ref), and (ref) are satisfied, then as $n\to\infty$,
$\operatorname{E}[\widehat{F}(u)] = \operatorname{E}[\widetilde{F}(u)]+n^{-1}\Delta(u)+\mathit{o}_{}(n^{-1})$ and
$\operatorname{E}[\widehat{F}_{\rho}(u)] = \operatorname{E}[\widetilde{F}(u)]+n^{-1}\Delta(u)+n^{-1}\Delta_{\rho}(u)+\mathit{o}_{}(n^{-1})$, where
\begin{align}
\notag
\Delta(u) &= \operatorname{E}[\nabla^{\top}u_{i}Hg_{i}|u]f(u)
-\zeta_{\lambda}^{\top}H^{\top}\operatorname{E}[\nabla u_{i}|u]^{\top}f(u) \\
&\quad +\tfrac{1}{2}\mathop{\mathrm{tr}}\left(\Sigma[\mathrm{d}\{\operatorname{E}[\nabla u_{i}\nabla^{\top}u_{i}|u]f(u)\}/\mathrm{d}u
-\operatorname{E}[\nabla^{2}u_{i}|u]f(u)]\right)\\
\intertext{and}
\Delta_{\rho}(u) & = {\begingroup\textstyle\int\endgroup}_{-\infty}^{u}(-c_{\rho}\operatorname{E}[g_{i}^{\top}Pg_{i}|t]+ c_{\rho}(d_{g}-d_{\beta}) +\zeta_{\lambda}^{\top}P\operatorname{E}[g_{i}|t] )\mathrm{d}F(t) = {\begingroup\textstyle\int\endgroup}_{-\infty}^{u}\delta_{\rho}(t)\mathrm{d}t.
\end{align}
Also
\begin{align}
\notag
\operatorname{Var}[\widehat{F}(u)] & = \operatorname{Var}[\widetilde{F}(u)] + n^{-1}\operatorname{E}[\nabla u_{i}|u]^{\top}\Sigma \operatorname{E}[\nabla u_{i}|u]f(u)^{2} \\
&\quad + 2n^{-1}\operatorname{E}[\nabla u_{i}|u]^{\top}H[{\begingroup\textstyle\int\endgroup}_{-\infty}^{u}\operatorname{E}[g_{i}|t]\mathrm{d}F(t)]f(u) +\mathit{o}_(n^{-1}),\\
\operatorname{Var}[\widehat{F}_{\rho}(u)] & = \operatorname{Var}[\widehat{F}(u)] -n^{-1}[{\begingroup\textstyle\int\endgroup}_{-\infty}^{u}\operatorname{E}[g_{i}|t]\mathrm{d}F(t)]^{\top}P[{\begingroup\textstyle\int\endgroup}_{-\infty}^{u}\operatorname{E}[g_{i}|t]\mathrm{d}F(t)] +\mathit{o}_(n^{-1}b).
\end{align}
If, in addition, $\mathrm{d}\{ \operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u$ is absolutely integrable, the remainder term of $\operatorname{Var}[\widehat{F}(u)]$ is $\mathit{o}_{}(n^{-1}b)$.
remarkIf $\delta (u)$ in Theorem (ref) is defined, then $\Delta(u)=\int_{-\infty}^{u}\delta(t)\mathrm{d}t$, but there is no requirement that $\Delta(u)$ is absolutely continuous in Theorem (ref). Otherwise, the interpretation is exactly the same as in Theorem (ref). In particular, the main qualitative conclusions in Supplement (ref): Examples (ref) and (ref) still hold.
Simulation Evidence
Preliminaries
Consider the inverse hyperbolic sine (IHS) transformation model
equation[equation omitted — 174 chars of source]
here $\beta = (\delta,\gamma,\theta)^{\top}$ and $z=(y,x)^{\top}$. The IHS transformation has been proposed in johnson1949 as an alternative to the Box-Cox power transform, $(y^{\lambda}-1)/\lambda$, $y\geq 0$, and developed in burbidge1988 and mackinnon1990; see also, e.g., ramirez1994, brown2015 and the references therein for recent applications in statistics and econometrics, and tsai2017 for comparisons with other transformations. When $\theta =0$, the IHS transform is defined as the limiting value, $\lim_{\theta\to0}\operatorname{arsinh}(\theta y)/\theta=y$, which corresponds to the Box-Cox transform with $\lambda=1$; when $\theta\neq0$, the shapes of the IHS transforms are similar to those of the Box-Cox with $\lambda<1$. The advantage of the IHS transform is that it is a smooth function of $y\in\mathbbm{R}$ and $\theta\in\mathbbm{R}$ with values at $\theta=0$ defined as the corresponding limits.
The infeasible optimal instruments in the IHS transformation model (ref) are
equation*[equation* omitted — 244 chars of source]
see robinson1991. The last element of $S(x;\beta_{0})$, $s_{3}(x;\beta_{0})$, depends on the conditional distribution of $u$ given $x$, and, in general, there is little reason to argue for a particular scalar function of $x$ as a good approximant. For example, if $u|x\sim N(0,\sigma ^{2})$, based on $\tanh (x)\simeq 2\Phi ((\pi /2)^{1/2}x)-1$ twice, $s_{3}(x;\beta_{0})$
is approximately $\tanh\left(\theta_{0}(\delta_{0}+\gamma_{0}x)/(\pi\theta_{0}^{2}\sigma^{2}/2+1)^{1/2}\right)/\theta_{0}^{2} -(\delta_{0}+\gamma_{0}x)/\theta_{0}$ which suggests the use of odd degree polynomials in $x$ as instruments; other and better approximations are of course available.
In all cases the true parameters are $\delta_{0}=1$, $\gamma_{0}=2$ and $\theta_{0}=0.08$ which yield a signal-to-noise ratio of $\gamma_{0}^{2}/(1+\gamma_{0}^{2})=4/5=0.8$ somewhat more stringent than that of $16/17=0.941$ in robinson1991.
Design
Given the uncertainty concerning the conditional distribution $u|x$ the approach adopted here is to simply compare estimators based on moment conditions $\operatorname{E}[g(z,\beta_{0})]=0$ (ref) where
equation*[equation* omitted — 79 chars of source]
for $d_{g}=3$ (exactly identified), $4$ and $5$ (over-identified).
Three data generating processes for $(x,u)$ are considered.
\noindentScenario 1. $x$ and $u$ are distributed as independent standard normal $N(0,1)$; cf. robinson1991.
remarkScenario 1 satisfies the conditions of Supplement (ref): Example (ref). Hence
$\operatorname{IVar}[\hat{f}_{\rho}]=\operatorname{IVar}[\hat{f}]+\mathit{o}_{}(n^{-1})$ and the relative integrated variance (IVar)
\begin{equation}
\operatorname{IVar}[\hat{f}]/\operatorname{IVar}[\tilde{f}] = 1
-\tfrac{b}{4\pi^{1/2}R(k)}
+ \tfrac{b}{\tau^{\top}D\tau R(k)}{\begingroup\textstyle\int\endgroup}\left(\mathrm{d}\{(\tau_{0|u}(u)-\tau_{0})f(u)\}/\mathrm{d}u\right)^{2}\mathrm{d}u
+\mathit{o}_(b),
\end{equation}
where $\tau_{0|u}(u) = \operatorname{E}[\tanh(\theta_{0}(u +\delta_{0}+\gamma_{0} x))|u]/\theta_{0}^{2}-(\delta_{0}+u)/\theta_{0}$,
$\tau_{j}=\operatorname{E}[x^{j}s_{3}(x,\beta_{0})]$, $j=0,1,2,\ldots$, $\tau=(\tau_{0},\tau_{1},\ldots,\tau_{d_{g}-1})^{\top}$, and
$D = M^{-1}-\mathop{\mathrm{diag}}(I_{2},0)$
with $M=\{M_{ij}\}_{i,j=1}^{d_{g}}$, $M_{ij}=\operatorname{E}[x^{i+j-2}]$, $i,j=1,\ldots,d_{g}$.
The term $-b/(4\pi^{1/2}R(k))$ does not depend on the number of moment conditions $d_{g}$ and is the asymptotic reduction in integrated variance due to the constraint that the mean of $u$ is zero; see also Supplement (ref): Example (ref). The second term in $b$ is non-negative and represents the increase in integrated variance due to estimation of $\gamma _{0}$ and $\theta _{0}$; it decreases as the number of moment condition increases; e.g. for $d_{g}=4$, $5$, $10$, $20$, $\tau^{\top}D\tau =9.8092$, $9.8514$, $9.9857$ and $9.9859$, respectively.
\noindentScenarios 2 and 3. $x$ and $u$ have joint density $f_{ux}(u,x)=xf_{NM}(ux)f_{x}(x)$ where $x$ is a generalised gamma random variable, stacy1962, with parameters $p=2$, $d=\nu$ and $a=(2/\nu)^{1/2}$ for some $\nu>4$ and $f_{NM}$ is the normal mixture density with $m$ components, viz. $f_{NM}(w) = \sum_{j=1}^{m}\omega_{j}\phi_{\sigma_{j}}(w-\mu_{j})$, $-\infty<\mu _{j}<\infty$, $\sigma_{j}>0$, $j=1,\ldots,m$, $\sum_{j=1}^{m}\omega_{j}=1$, and $\sum_{j=1}^{m}\omega_{j}\mu_{j}=0$, i.e., $\operatorname{E}[w]=0$.
Here $\phi(x)$ denotes the standard normal p.d.f. and $\phi_{\sigma}(x)=\phi(x/\sigma)/\sigma$.
The joint density $f_{ux}$ is the density of $u=w/x$ and $x$ where $w$ and $x$ are independent. The conditional density of $u$ given $x$ is
$f_{u|x}(u|x) = xf_{NM}(ux) = \sum_{j=1}^{m}\omega_{j}\phi_{\sigma_{j}/x}(u-\mu_{j}/x)$.
Hence, $\operatorname{E}[u|x]=0$ and
$\operatorname{E}[u_{i}^{2}|x] = \sum_{j=1}^{m}\omega_{j}(\sigma_{j}^{2}+\mu_{j}^{2})/x^{2}$.
The marginal density of $u$ is a mixture of noncentral $t$ densities
$f_{u}(u) = \sum_{j=1}^{m}\omega_{j}t_{\nu}(u/\sigma_{j};\: \mu_{i}/\sigma_{i})/\sigma_{j}$
where $t_{\nu}(\cdot;\eta)$ is the density of a noncentral $t$-distributed random variable with $\nu$ degrees of freedom and noncentrality parameter $\eta$ allowing a wide variety of shapes for $f_{u}$ by varying the mixture $f_{NM}$. The skewed unimodal and bimodal densities shown in Figure (ref) describe the NM densities for Scenarios 2 and 3 respectively, i.e., the mixture densities marron1992 centered to have zero mean.
figure[figure omitted — 726 chars of source]
Kernel Functions and Bandwidths
Fourth order Gaussian-based kernels, $k(x)=(3-x^{2})\phi(x)/2$ and $K(x)=\Phi(x)+x\phi (x)/2$, $\Phi(x)=\int_{-\infty}^{x}\phi(u)\mathrm{d}u$, are employed; see wand1990 and oryshchenko2017 respectively. Thus the choices of the asymptotically optimal bandwidths $(27/4\sqrt{\pi})^{1/9}R(f^{(4)})^{-1/9}n^{-1/9}$ and $(7/2\sqrt{\pi })^{1/7}R(f^{(3)})^{-1/7}n^{-1/7}$ for p.d.f. and c.d.f. estimation respectively are permitted, thereby satisfying Assumptions (ref)(c) and (ref)(b). The practical issue of estimating the derivatives of $f$ required for the computation of $R(f^{(j)})$, $j=3,4$, is ignored and the respective true values used. For the standard normal distribution these are $R(\phi^{(3)})=15/(16\sqrt{\pi})$ and $R(\phi^{(4)})=105/(32\sqrt{\pi})$; for the mixture distributions, approximate values are shown in Figure (ref).
Results
The study compares the performance of GEL-based kernel density p.d.f. and c.d.f. estimators. The GEL parameter estimators are CUE, EL and ET, the most notable special cases of the GEL family. For each estimator the mean and variance were computed on a grid $1000$ of points between $-5$ and $5$ and are reported as the integrated squared bias and integrated variance relative
to those of the corresponding infeasible estimator based on the true $u$, i.e., $\tilde{f}$ and $\widetilde{F}$.
Tables (ref), (ref) and (ref) report results for Scenarios 1, 2 and 3 respectively. The ISB, IVar and MISE (all $\times10^5$) for the infeasible $\tilde{f}$ and $\widetilde{F}$ are presented. Rows ISB, IVar, and MISE are the ISB, IVar, and MISE of $\hat{f}$, $\hat{f}_{\rho}$ ($\widehat{F}$, $\widehat{F}_{\rho}$) relative to the infeasible $\tilde{f}$ ($\widetilde{F}$), respectively; row `vs $d_{g}=3$' is the MISE of $\hat{f}$, $\hat{f}_{\rho}$ ($\hat{F}$, $\widehat{F}_{\rho}$) relative to the corresponding value for $d_{g}=3$; row `w. vs unw.' is the MISE of $\hat{f}_{\rho}$ ($\widehat{F}_{\rho}$) relative to $\hat{f}$ ($\widehat{F}$). Rows MISE, `vs $d_{g}=3$', and `w. vs unw.' examine the significance of the paired $t$-statistics in a two-sided test for equality of the respective ISE means, e.g., $\int (\hat{f}(u)-f(u))^{2}\mathrm{d}u$; the symbol \dag\ indicates that the $p$-value is between $0.01$ and $0.05$ whereas \ddag\ that it is less than $0.01$ and in all other cases the $p$-value is greater than $0.05$. Values of relative MISE less than $1$ are emphasised in bold.
Sample sizes $n=100$, $500$, $1000$, and $2000$ are examined.
All computations were carried out in MATLAB; the relevant code and additional results, including the properties of GEL estimators, are available from the first named author upon request. All results are based on $10,000$ random draws.
{5pt}
table[table omitted — 12,808 chars of source]
table[table omitted — 12,643 chars of source]
table[table omitted — 11,995 chars of source]
{6pt}
Scenario 1
The first $\sim b$ term in eq. (ref) is approximately $-0.321n^{-1/9}$, which for $n=100$, $500$, $1000$, and $2000$ is approximately $-0.192$, $-0.161$, $-0.149$, and $-0.138$ respectively. The second $\sim b$ term is
approximately $0.04728n^{-1/9}$ for $d_{g}=4$ and $0.04708n^{-1/9}$ for $d_{g}=5$, which offsets the reduction in variance slightly. The predicted relative IVar of $\hat{f}$ and $\hat{f}_{\rho}$ up to order $\mathit{o}_{}(b)$ is thus $0.836$, $0.863$, $0.873$ and $0.882$ for $n=100$, $500$, $1000$, and $2000$ respectively and is identical within three digit precision for $d_{g}=4$ and $5$.
The results reported in Table (ref) confirm these predictions. In fact, the reduction in variance is even larger than expected in small and medium samples due to the $\mathit{o}_{}(b)$ effects. Furthermore, estimators $\hat{f}$ and $\hat{f}_{\rho}$ have smaller ISB relative to $\tilde{f}$. A comparison of $\hat{f}$ and $\hat{f}_{\rho}$ between $d_{g}=3$ (just-identified) and $d_{g}=4,5$ (over-identified) for moderate and larger sample sizes emphasises further the contribution of additional moment information. Hence $\hat{f}$ and $\hat{f}_{\rho}$ enjoy a reduction in MISE of as much as $21\%$ for $n=100$ and $10\%$ for $n=2000$ relative to $\tilde{f}$. The benefits are even more pronounced for c.d.f. estimation, where the reduction in MISE can be as much as $56\%$ for $n=100$ and around $53\%$ in moderate samples. There are also small but statistically significant benefits to re-weighting which are mostly due to the smaller biases of $\hat{f}_{\rho}$ and $\widehat{F}
_{\rho}$ relative to $\hat{f}$ and $\widehat{F}$ at moderate and larger sample sizes. There is some deterioration in ISB, IVar and, thus, MISE with increases in $d_{g}$ which can be contributed to the increased importance of outliers.
Finally, while in moderate and large samples the performances of CUE, EL, and ET are virtually identical, in small samples ET can be unstable with larger $d_{g}$.
Scenarios 2 and 3
Scenarios 2 and 3 with densities of $(x,u)$ which are heavy-tailed and also, e.g., skewed and bimodal, illustrate the many difficulties for both GEL estimation and kernel p.d.f. and c.d.f. estimation which are absent in the relatively benign Scenario 1.
The performance of CUE in small samples is generally worse than that of EL and ET. It ranks last by MSE in both scenarios with $n=100$ and $500$, except Scenario 3 with $n=100$ where ET underperforms. In a number of cases increasing with $d_{g}$ the optimisation routine for ET failed. Somewhat surprisingly, although it is known to be sensitive to outliers, EL appears to deliver good results in the simulation experiments. It ranks first by MSE in Scenario 3 with $d_{g}=5$ and alternates with ET otherwise. These differences become very small with $n=1,000$ and greater.
The conclusion about the inferior performance of CUE in small samples holds true for CUE-based kernel density p.d.f. and c.d.f. estimators as well; see Tables (ref) and (ref), in particular, the ISBs of $\hat{f}$ and $\hat{f}_{\rho}$ with $d_{g}=4,5$ in Table (ref).
However, the ranking of EL and ET-based kernel density p.d.f. and c.d.f. estimators by MISE does not always correspond to the ranking of the underlying EL and ET estimators of $\beta_{0}$ by MSE.
In particular, the sensitivity of EL to outliers adversely affects the estimators $\hat{f}_{\rho}$ and $\widehat{F}_{\rho}$ via the implied probabilities in Scenario 3 with $n=500$ and greater; see Table (ref). ET and CUE perform better in those cases.
Unlike Scenario 1, in Scenario 3 none of the feasible kernel density estimators have smaller MISE than their infeasible counterparts for the sample sizes considered. In Scenario 2, with less complicated distributional features, these estimators do achieve a reduction in MISE with $d_{g}=4,5$. The same is true for the feasible kernel c.d.f. estimators in Scenario 2 with $d_{g}=3,4,5$, and more often than not in Scenario 3 as well, with the few exceptions mentioned above.
Importantly, it is generally beneficial to increase the number of moment conditions beyond those necessary to identify the parameters except when stability of GEL estimators of $\beta_{0}$ is likely to deteriorate.
Finally, the benefits of re-weighting are present, but not universal, and as expected, are quite small; cf. Supplement (ref): Example (ref).
Summary and Conclusions
Large sample results and simulation evidence reported in this paper suggest that it is generally sensible to apply either the standard or re-weighted kernel estimators to estimate the p.d.f. or c.d.f. of a scalar residual $u(z,\beta_{0})$ in a variety of situations, provided error associated with the estimation of $\beta_{0}$ satisfies some mild regularity conditions and care is taken to ensure the bandwidth is not too small. If the assumptions on $u(z,\beta)$ prove difficult to verify in practice, using fourth or higher order kernels and the corresponding asymptotically optimal bandwidths will generally assist with ensuring the appropriate regularity conditions hold.
Incorporating information from overidentifying moment conditions by re-weighting the estimators using GEL implied probabilities offers efficiency gains which are realised in regular situations. However, if the model is highly nonlinear and the distribution of the data is heavy-tailed or contaminated with outliers, the methods proposed in this paper, including GEL, should be applied with some caution in very small samples. Robustified hybrid estimators such as the exponentially tilted empirical likelihood,
see, e.g., schennach2007, may prove useful in these circumstances.
While the results in this paper were presented only for the scalar-valued $u(z,\beta)$, generalisations to the vector case are relatively straightforward provided an analogue of the bijection Assumption (ref) holds.
An issue for future research to usefully address is the construction of tests for overidentifying moment conditions or parametric restrictions based on the differences between the kernel p.d.f. estimators $\hat{f}_{\rho}$ and $\hat{f}$ or $\tilde{f}_{\rho}$ and $\tilde{f}$ for known $\beta_{0}$. Test statistics of the Bickel-Rosenblatt type based on the integrated squared difference $\int(\hat{f}_{\rho}(u)-\hat{f}(u))^{2}\mathrm{d}u$, bickel1973, fan1994,fan1998, or the integrated absolute difference, cao2005, would be of interest. Alternatively, Kolmogorov-Smirnov or Cram\'{e}r-von Mises-type tests could be constructed based on the differences between kernel c.d.f. estimators.
\phantomsection
\addcontentsline{toc}{section}{References}
{0pt plus 0.3ex}
thebibliography{xx}
\harvarditem{Ahmad}{1992}{ahmad1992}
Ahmad, I. A. \harvardyearleft 1992\harvardyearright , `Residuals density
estimation in nonparametric regression', {\em Statistics & Probability
Letters} {\bf 14}(2), 133--139.
doi:
\href{http://dx.doi.org/10.1016/0167-7152(92)90077-I}{10.1016/0167-7152(92)90077-I}
\harvarditem[Antoine et al.]{Antoine, Bonnal \harvardand\
Renault}{2007}{antoine2007}
Antoine, B., Bonnal, H. \harvardand\ Renault, E. \harvardyearleft
2007\harvardyearright , `On the efficient use of the informational content of
estimating equations: Implied probabilities and {E}uclidean empirical
likelihood', {\em Journal of Econometrics} {\bf 138}(2), 461--487.
doi:
\href{http://dx.doi.org/10.1016/j.jeconom.2006.05.005}{10.1016/j.jeconom.2006.05.005}
\harvarditem{Back \harvardand\ Brown}{1993}{back1993}
Back, K. \harvardand\ Brown, D. P. \harvardyearleft 1993\harvardyearright ,
`Implied probabilities in {GMM} estimators', {\em Econometrica} {\bf
61}(4), 971--975.
doi: \href{http://dx.doi.org/10.2307/2951771}{10.2307/2951771}
\harvarditem{Bartlett}{1963}{bartlett1963}
Bartlett, M. S. \harvardyearleft 1963\harvardyearright , `Statistical
estimation of density functions', {\em Sankhy\={a}: The Indian Journal of
Statistics, Series A} {\bf 25}(3), 245--254.
URL: \url{www.jstor.org/stable/25049271}
\harvarditem{Bhattacharya \harvardand\ Ghosh}{1978}{bhattacharya1978}
Bhattacharya, R. N. \harvardand\ Ghosh, J. K. \harvardyearleft
1978\harvardyearright , `On the validity of the formal {E}dgeworth
expansion', {\em The Annals of Statistics} {\bf 6}(2), 434--451.
doi: \href{http://dx.doi.org/10.1214/aos/1176344134}{10.1214/aos/1176344134}
\harvarditem{Bickel \harvardand\ Rosenblatt}{1973}{bickel1973}
Bickel, P. J. \harvardand\ Rosenblatt, M. \harvardyearleft
1973\harvardyearright , `On some global measures of the deviations of density
function estimates', {\em The Annals of Statistics} {\bf 1}(6), 1071--1095.
doi: \href{http://dx.doi.org/10.1214/aos/1176342558}{10.1214/aos/1176342558}
\harvarditem{Bochner}{1955}{bochner1955}
Bochner, S. \harvardyearleft 1955\harvardyearright , {\em Harmonic analysis
and the theory of probability}, University of California Press.
\harvarditem[Bott et al.]{Bott, Devroye \harvardand\ Kohler}{2013}{bott2013}
Bott, A.-K., Devroye, L. \harvardand\ Kohler, M. \harvardyearleft
2013\harvardyearright , `Estimation of a distribution from data with small
measurement errors', {\em Electronic Journal of Statistics} {\bf
7}, 2457--2476.
doi: \href{http://dx.doi.org/10.1214/13-EJS850}{10.1214/13-EJS850}
\harvarditem{Brown \harvardand\ Newey}{1998}{brown1998}
Brown, B. W. \harvardand\ Newey, W. K. \harvardyearleft 1998\harvardyearright
, `Efficient semiparametric estimation of expectations', {\em Econometrica}
{\bf 66}(2), 453--464.
doi: \href{http://dx.doi.org/10.2307/2998566}{10.2307/2998566}
\harvarditem{Brown \harvardand\ Newey}{2002}{brown2002}
Brown, B. W. \harvardand\ Newey, W. K. \harvardyearleft 2002\harvardyearright
, `Generalized method of moments, efficient bootstrapping, and improved
inference', {\em Journal of Business & Economic Statistics} {\bf
20}(4), 507--517.
doi:
\href{http://dx.doi.org/10.1198/073500102288618649}{10.1198/073500102288618649}
\harvarditem[Brown et al.]{Brown, Greene, Harris \harvardand\
Taylor}{2015}{brown2015}
Brown, S., Greene, W. H., Harris, M. N. \harvardand\ Taylor, K.
\harvardyearleft 2015\harvardyearright , `An inverse hyperbolic sine
heteroskedastic latent class panel tobit model: An application to modelling
charitable donations', {\em Economic Modelling} {\bf 50}, 228--236.
doi:
\href{http://dx.doi.org/10.1016/j.econmod.2015.06.018}{10.1016/j.econmod.2015.06.018}
\harvarditem[Burbidge et al.]{Burbidge, Magee \harvardand\
Robb}{1988}{burbidge1988}
Burbidge, J. B., Magee, L. \harvardand\ Robb, A. L. \harvardyearleft
1988\harvardyearright , `Alternative transformations to handle extreme values
of the dependent variable', {\em Journal of the American Statistical
Association} {\bf 83}(401), 123--127.
doi:
\href{http://dx.doi.org/10.1080/01621459.1988.10478575}{10.1080/01621459.1988.10478575}
\harvarditem{Cao \harvardand\ Lugosi}{2005}{cao2005}
Cao, R. \harvardand\ Lugosi, G. \harvardyearleft 2005\harvardyearright ,
`Goodness-of-fit tests based on the kernel density estimator', {\em
Scandinavian Journal of Statistics} {\bf 32}(4), 599--616.
doi:
\href{http://dx.doi.org/10.1111/j.1467-9469.2005.00471.x}{10.1111/j.1467-9469.2005.00471.x}
\harvarditem{Chamberlain}{1987}{chamberlain1987}
Chamberlain, G. \harvardyearleft 1987\harvardyearright , `Asymptotic
efficiency in estimation with conditional moment restrictions', {\em Journal
of Econometrics} {\bf 34}(3), 305--334.
doi:
\href{http://dx.doi.org/10.1016/0304-4076(87)90015-7}{10.1016/0304-4076(87)90015-7}
\harvarditem{Chen \harvardand\ Qin}{1993}{chen1993}
Chen, J. \harvardand\ Qin, J. \harvardyearleft 1993\harvardyearright ,
`Empirical likelihood estimation for finite populations and the effective
usage of auxiliary information', {\em Biometrika} {\bf 80}(1), 107--116.
doi: \href{http://dx.doi.org/10.1093/biomet/80.1.107}{10.1093/biomet/80.1.107}
\harvarditem{Chen}{1997}{chen1997}
Chen, S. X. \harvardyearleft 1997\harvardyearright , `Empirical
likelihood-based kernel density estimation', {\em Australian and New Zealand
Journal of Statistics} {\bf 39}(1), 47--56.
doi:
\href{http://dx.doi.org/10.1111/j.1467-842X.1997.tb00522.x}{10.1111/j.1467-842X.1997.tb00522.x}
\harvarditem{Chen \harvardand\ Cui}{2007}{chen2007}
Chen, S. X. \harvardand\ Cui, H. \harvardyearleft 2007\harvardyearright , `On
the second-order properties of empirical likelihood with moment
restrictions', {\em Journal of Econometrics} {\bf 141}(2), 492--516.
doi:
\href{http://dx.doi.org/10.1016/j.jeconom.2006.10.006}{10.1016/j.jeconom.2006.10.006}
\harvarditem{Cheng}{2004}{cheng2004}
Cheng, F. \harvardyearleft 2004\harvardyearright , `Weak and strong uniform
consistency of a kernel error density estimator in nonparametric regression',
{\em Journal of Statistical Planning and Inference} {\bf 119}(1), 95--107.
doi:
\href{http://dx.doi.org/10.1016/S0378-3758(02)00417-2}{10.1016/S0378-3758(02)00417-2}
\harvarditem{Cheng}{2005}{cheng2005}
Cheng, F. \harvardyearleft 2005\harvardyearright , `Asymptotic distributions
of error density estimators in first-order autoregressive models', {\em
Sankhy\={a}: The Indian Journal of Statistics} {\bf 67}(3), 553--567.
URL: \url{http://www.jstor.org/stable/25053449}
\harvarditem[Chernozhukov et al.]{Chernozhukov, Fern\'{a}ndez-Val \harvardand\
Galichon}{2009}{chernozhukov2009}
Chernozhukov, V., Fern\'{a}ndez-Val, I. \harvardand\ Galichon, A.
\harvardyearleft 2009\harvardyearright , `Improving point and interval
estimators of monotone functions by rearrangement', {\em Biometrika} {\bf
96}(3), 559--575.
doi: \href{http://dx.doi.org/10.1093/biomet/asp030}{10.1093/biomet/asp030}
\harvarditem{Corcoran}{1998}{corcoran1998}
Corcoran, S. A. \harvardyearleft 1998\harvardyearright , `Bartlett adjustment
of empirical discrepancy statistics', {\em Biometrika} {\bf 85}(4), 967--972.
doi: \href{http://dx.doi.org/10.1093/biomet/85.4.967}{10.1093/biomet/85.4.967}
\harvarditem{Cox \harvardand\ Snell}{1968}{cox1968}
Cox, D. R. \harvardand\ Snell, E. J. \harvardyearleft 1968\harvardyearright ,
`A general definition of residuals', {\em Journal of the Royal Statistical
Society. Series B} {\bf 30}(2), 248--275.
URL: \url{www.jstor.org/stable/2984505}
\harvarditem{Cressie \harvardand\ Read}{1984}{cressie1984}
Cressie, N. \harvardand\ Read, T. R. C. \harvardyearleft 1984\harvardyearright
, `Multinomial goodness-of-fit tests', {\em Journal of the Royal Statistical
Society. Series B} {\bf 46}(3), 440--464.
URL: \url{http://www.jstor.org/stable/2345686}
\harvarditem{Fan}{1994}{fan1994}
Fan, Y. \harvardyearleft 1994\harvardyearright , `Testing the goodness of fit
of a parametric density function by kernel method', {\em Econometric Theory}
{\bf 10}(2), 316--356.
doi:
\href{http://dx.doi.org/10.1017/S0266466600008434}{10.1017/S0266466600008434}
\harvarditem{Fan}{1998}{fan1998}
Fan, Y. \harvardyearleft 1998\harvardyearright , `Goodness-of-fit tests based
on kernel density estimators with fixed smoothing parameters', {\em
Econometric Theory} {\bf 14}(5), 604--621.
doi:
\href{http://dx.doi.org/10.1017/s0266466698145036}{10.1017/s0266466698145036}
\harvarditem[Glad et al.]{Glad, Hjort \harvardand\ Ushakov}{2003}{glad2003}
Glad, I. K., Hjort, N. L. \harvardand\ Ushakov, N. G. \harvardyearleft
2003\harvardyearright , `Correction of density estimators that are not
densities', {\em Scandinavian Journal of Statistics} {\bf 30}(2), 415--427.
doi: \href{http://dx.doi.org/10.1111/1467-9469.00339}{10.1111/1467-9469.00339}
\harvarditem{Gy\"{o}rfi \harvardand\ Walk}{2012}{gyorfi2012}
Gy\"{o}rfi, L. \harvardand\ Walk, H. \harvardyearleft 2012\harvardyearright ,
`Strongly consistent density estimation of the regression residual', {\em
Statistics & Probability Letters} {\bf 82}(11), 1923--1929.
doi:
\href{http://dx.doi.org/10.1016/j.spl.2012.06.021}{10.1016/j.spl.2012.06.021}
\harvarditem{Hall}{2005}{hall2005}
Hall, A. R. \harvardyearleft 2005\harvardyearright , {\em Generalized method
of moments}, Oxford University Press.
\harvarditem[Hansen et al.]{Hansen, Heaton \harvardand\
Yaron}{1996}{hansen1996}
Hansen, L. P., Heaton, J. \harvardand\ Yaron, A. \harvardyearleft
1996\harvardyearright , `Finite-sample properties of some alternative {GMM}
estimators', {\em Journal of Business & Economic Statistics} {\bf
14}(3), 262--280.
doi: \href{http://dx.doi.org/10.2307/1392442}{10.2307/1392442}
\harvarditem[Imbens et al.]{Imbens, Spady \harvardand\
Johnson}{1998}{imbens1998}
Imbens, G. W., Spady, R. H. \harvardand\ Johnson, P. \harvardyearleft
1998\harvardyearright , `Information theoretic approaches to inference in
moment condition models', {\em Econometrica} {\bf 66}(2), 333--357.
doi: \href{http://dx.doi.org/10.2307/2998561}{10.2307/2998561}
\harvarditem{Jensen}{1989}{jensen1989}
Jensen, J. L. \harvardyearleft 1989\harvardyearright , `Validity of the formal
{E}dgeworth expansion when the underlying distribution is partly discrete',
{\em Probability Theory and Related Fields} {\bf 81}(4), 507--519.
doi: \href{http://dx.doi.org/10.1007/BF00367300}{10.1007/BF00367300}
\harvarditem{Johnson}{1949}{johnson1949}
Johnson, N. L. \harvardyearleft 1949\harvardyearright , `Systems of frequency
curves generated by methods of translation', {\em Biometrika} {\bf
36}(1--2), 149--176.
doi:
\href{http://dx.doi.org/10.1093/biomet/36.1-2.149}{10.1093/biomet/36.1-2.149}
\harvarditem{Kitamura \harvardand\ Stutzer}{1997}{kitamura1997}
Kitamura, Y. \harvardand\ Stutzer, M. \harvardyearleft 1997\harvardyearright ,
`An information-theoretic alternative to generalized method of moments
estimation', {\em Econometrica} {\bf 65}(4), 861--874.
doi: \href{http://dx.doi.org/10.2307/2171942}{10.2307/2171942}
\harvarditem[Kiwitt et al.]{Kiwitt, Nagel \harvardand\
Neumeyer}{2008}{kiwitt2008}
Kiwitt, S., Nagel, E. \harvardand\ Neumeyer, N. \harvardyearleft
2008\harvardyearright , `Empirical likelihood estimators for the error
distribution in nonparametric regression models', {\em Mathematical Methods
of Statistics} {\bf 17}(3), 241--260.
doi:
\href{http://dx.doi.org/10.3103/S1066530708030058}{10.3103/S1066530708030058}
\harvarditem{Kundhi \harvardand\ Rilstone}{2012}{kundhi2012}
Kundhi, G. \harvardand\ Rilstone, P. \harvardyearleft 2012\harvardyearright ,
`Edgeworth expansions for {GEL} estimators', {\em Journal of Multivariate
Analysis} {\bf 106}, 118--146.
doi:
\href{http://dx.doi.org/10.1016/j.jmva.2011.11.005}{10.1016/j.jmva.2011.11.005}
\harvarditem{Loynes}{1969}{loynes1969}
Loynes, R. M. \harvardyearleft 1969\harvardyearright , `On {C}ox and {S}nell's
general definition of residuals', {\em Journal of the Royal Statistical
Society. Series B} {\bf 31}(1), 103--106.
URL: \url{www.jstor.org/stable/2984331}
\harvarditem{MacKinnon \harvardand\ Magee}{1990}{mackinnon1990}
MacKinnon, J. G. \harvardand\ Magee, L. \harvardyearleft 1990\harvardyearright
, `Transforming the dependent variable in regression models', {\em
International Economic Review} {\bf 31}(2), 315--339.
doi: \href{http://dx.doi.org/10.2307/2526842}{10.2307/2526842}
\harvarditem{Marron \harvardand\ Wand}{1992}{marron1992}
Marron, J. S. \harvardand\ Wand, M. P. \harvardyearleft 1992\harvardyearright
, `Exact mean integrated squared error', {\em The Annals of Statistics} {\bf
20}(2), 712--736.
doi: \href{http://dx.doi.org/10.1214/aos/1176348653}{10.1214/aos/1176348653}
\harvarditem{M\'{a}ty\'{a}s}{1999}{matyas1999}
M\'{a}ty\'{a}s, L., ed. \harvardyearleft 1999\harvardyearright , {\em
Generalized method of moments estimation}, Cambridge University Press.
\harvarditem{Muhsal \harvardand\ Neumeyer}{2010}{mushal2010}
Muhsal, B. \harvardand\ Neumeyer, N. \harvardyearleft 2010\harvardyearright ,
`A note on residual-based empirical likelihood kernel density estimation',
{\em Electronic Journal of Statistics} {\bf 4}, 1386--1401.
doi: \href{http://dx.doi.org/10.1214/10-EJS586}{10.1214/10-EJS586}
\harvarditem{Nadaraya}{1964}{nadaraya1964b}
Nadaraya, E. A. \harvardyearleft 1964\harvardyearright , `Some new estimates
for distribution functions', {\em Theory of Probability and its Applications}
{\bf 9}(3), 497--500.
doi: \href{http://dx.doi.org/10.1137/1109069}{10.1137/1109069}
\harvarditem{Newey \harvardand\ Smith}{2004}{newey2004}
Newey, W. K. \harvardand\ Smith, R. J. \harvardyearleft 2004\harvardyearright
, `Higher order properties of {GMM} and generalized empirical likelihood
estimators', {\em Econometrica} {\bf 72}(1), 219--255.
doi:
\href{http://dx.doi.org/10.1111/j.1468-0262.2004.00482.x}{10.1111/j.1468-0262.2004.00482.x}
\harvarditem{Oryshchenko}{2017}{oryshchenko2017}
Oryshchenko, V. \harvardyearleft 2017\harvardyearright , Exact mean integrated
squared error and bandwidth selection for kernel distribution function
estimators, Working paper, University of Manchester.
URL: \url{https://arxiv.org/abs/1606.06993}
\harvarditem{Owen}{1988}{owen1988}
Owen, A. \harvardyearleft 1988\harvardyearright , `Empirical likelihood ratio
confidence intervals for a single functional', {\em Biometrika} {\bf
75}(2), 237--249.
doi: \href{http://dx.doi.org/10.1093/biomet/75.2.237}{10.1093/biomet/75.2.237}
\harvarditem{Owen}{1990}{owen1990}
Owen, A. \harvardyearleft 1990\harvardyearright , `Empirical likelihood ratio
confidence regions', {\em The Annals of Statistics} {\bf 18}(1), 90--120.
doi: \href{http://dx.doi.org/10.1214/aos/1176347494}{10.1214/aos/1176347494}
\harvarditem{Pagan \harvardand\ Ullah}{1999}{pagan1999}
Pagan, A. \harvardand\ Ullah, A. \harvardyearleft 1999\harvardyearright , {\em
Nonparametric econometrics}, Cambridge University Press.
\harvarditem{Parente \harvardand\ Smith}{2014}{parente2014}
Parente, P. M. \harvardand\ Smith, R. J. \harvardyearleft
2014\harvardyearright , `Recent developments in empirical likelihood and
related methods', {\em Annual Review of Economics} {\bf 6}, 77--102.
doi:
\href{http://dx.doi.org/10.1146/annurev-economics-080511-110925}{10.1146/annurev-economics-080511-110925}
\harvarditem{Parzen}{1962}{parzen1962}
Parzen, E. \harvardyearleft 1962\harvardyearright , `On estimation of a
probability density function and mode', {\em The Annals of Mathematical
Statistics} {\bf 33}(3), 1065--1076.
doi: \href{http://dx.doi.org/10.1214/aoms/1177704472}{10.1214/aoms/1177704472}
\harvarditem{Qin \harvardand\ Lawless}{1994}{qin1994}
Qin, J. \harvardand\ Lawless, J. \harvardyearleft 1994\harvardyearright ,
`Empirical likelihood and general estimating equations', {\em The Annals of
Statistics} {\bf 22}(1), 300--325.
doi: \href{http://dx.doi.org/10.1214/aos/1176325370}{10.1214/aos/1176325370}
\harvarditem[Ramirez et al.]{Ramirez, Moss \harvardand\
Boggess}{1994}{ramirez1994}
Ramirez, O. A., Moss, C. B. \harvardand\ Boggess, W. G. \harvardyearleft
1994\harvardyearright , `Estimation and use of the inverse hyperbolic sine
transformation to model non-normal correlated random variables', {\em Journal
of Applied Statistics} {\bf 21}(4), 289--304.
doi: \href{http://dx.doi.org/10.1080/757583872}{10.1080/757583872}
\harvarditem{Rao}{1983}{rao1983}
Rao, B. L. S. P. \harvardyearleft 1983\harvardyearright , {\em Nonparametric
functional estimation}, Academic Press.
\harvarditem{Robinson}{1991}{robinson1991}
Robinson, P. M. \harvardyearleft 1991\harvardyearright , `Best nonlinear
three-stage least squares estimation of certain econometric models', {\em
Econometrica} {\bf 59}(3), 755--786.
doi: \href{http://dx.doi.org/10.2307/2938227}{10.2307/2938227}
\harvarditem{Rosenblatt}{1956}{rosenblatt1956}
Rosenblatt, M. \harvardyearleft 1956\harvardyearright , `Remarks on some
nonparametric estimates of a density function', {\em The Annals of
Mathematical Statistics} {\bf 27}(3), 832--837.
doi: \href{http://dx.doi.org/10.1214/aoms/1177728190}{10.1214/aoms/1177728190}
\harvarditem{Schennach}{2007}{schennach2007}
Schennach, S. M. \harvardyearleft 2007\harvardyearright , `Point estimation
with exponentially tilted empirical likelihood', {\em The Annals of
Statistics} {\bf 35}(2), 634--672.
doi:
\href{http://dx.doi.org/10.1214/009053606000001208}{10.1214/009053606000001208}
\harvarditem{Silverman}{1986}{silverman1986}
Silverman, B. W. \harvardyearleft 1986\harvardyearright , {\em Density
Estimation for Statistics and Data Analysis}, Chapman & Hall.
\harvarditem{Smith}{1997}{smith1997}
Smith, R. J. \harvardyearleft 1997\harvardyearright , `Alternative
semi-parametric likelihood approaches to generalised method of moments
estimation', {\em The Economic Journal} {\bf 107}(441), 503--519.
doi:
\href{http://dx.doi.org/10.1111/j.0013-0133.1997.174.x}{10.1111/j.0013-0133.1997.174.x}
\harvarditem{Smith}{2011}{smith2011}
Smith, R. J. \harvardyearleft 2011\harvardyearright , `{GEL} criteria for
moment condition models', {\em Econometric Theory} {\bf 27}(6), 1192--1235.
doi:
\href{http://dx.doi.org/10.1017/S026646661100003X}{10.1017/S026646661100003X}
\harvarditem{Stacy}{1962}{stacy1962}
Stacy, E. W. \harvardyearleft 1962\harvardyearright , `A generalization of the
gamma distribution', {\em The Annals of Mathematical Statistics} {\bf
33}(3), 1187--1192.
doi: \href{http://dx.doi.org/10.1214/aoms/1177704481}{10.1214/aoms/1177704481}
\harvarditem[Tsai et al.]{Tsai, Liou, Simak \harvardand\ Cheng}{2017}{tsai2017}
Tsai, A. C., Liou, M., Simak, M. \harvardand\ Cheng, P. E. \harvardyearleft
2017\harvardyearright , `On hyperbolic transformations to normality', {\em
Computational Statistics & Data Analysis} {\bf 115}, 250--266.
doi:
\href{http://dx.doi.org/10.1016/j.csda.2017.06.001}{10.1016/j.csda.2017.06.001}
\harvarditem{Tsybakov}{2009}{tsybakov2009}
Tsybakov, A. B. \harvardyearleft 2009\harvardyearright , {\em Introduction to
Nonparametric Estimation}, (Springer Series in Statistics), Springer.
doi: \href{http://dx.doi.org/10.1007/b13794}{10.1007/b13794}
\harvarditem{{Van Ryzin}}{1969}{vanryzin1969}
{Van Ryzin}, J. \harvardyearleft 1969\harvardyearright , `On strong
consistency of density estimates', {\em The Annals of Mathematical
Statistics} {\bf 40}(5), 1765--1772.
doi: \href{http://dx.doi.org/10.1214/aoms/1177697388}{10.1214/aoms/1177697388}
\harvarditem{Wand \harvardand\ Jones}{1995}{wand1995}
Wand, M. P. \harvardand\ Jones, M. C. \harvardyearleft 1995\harvardyearright ,
{\em Kernel Smoothing}, Chapman & Hall.
\harvarditem{Wand \harvardand\ Schucany}{1990}{wand1990}
Wand, M. P. \harvardand\ Schucany, W. R. \harvardyearleft
1990\harvardyearright , `Gaussian-based kernels', {\em Canadian Journal of
Statistics} {\bf 18}(3), 197--204.
doi: \href{http://dx.doi.org/10.2307/3315450}{10.2307/3315450}
\harvarditem{Watson \harvardand\ Leadbetter}{1964}{watson1964}
Watson, G. S. \harvardand\ Leadbetter, M. R. \harvardyearleft
1964\harvardyearright , `Hazard analysis {II}', {\em Sankhy\={a}: The Indian
Journal of Statistics, Series A} {\bf 26}(1), 101--116.
URL: \url{http://www.jstor.org/stable/25049316}
\harvarditem{Yamato}{1973}{yamato1973}
Yamato, H. \harvardyearleft 1973\harvardyearright , `Uniform convergence of an
estimator of a distribution function', {\em Bulletin of Mathematical
Statistics} {\bf 15}(3-4), 69--78.
URL: \url{http://ci.nii.ac.jp/naid/120001036895/}
\harvarditem[Yuan et al.]{Yuan, Xu \harvardand\ Zheng}{2014}{yuan2014}
Yuan, A., Xu, J. \harvardand\ Zheng, G. \harvardyearleft 2014\harvardyearright
, `On empirical likelihood statistical functions', {\em Journal of
Econometrics} {\bf 178}(3), 613--623.
doi:
\href{http://dx.doi.org/10.1016/j.jeconom.2013.08.037}{10.1016/j.jeconom.2013.08.037}
\harvarditem{Zhang}{1995}{zhang1995}
Zhang, B. \harvardyearleft 1995\harvardyearright , `M-estimation and quantile
estimation in the presence of auxiliary information', {\em Journal of
Statistical Planning and Inference} {\bf 44}(1), 77--94.
doi:
\href{http://dx.doi.org/10.1016/0378-3758(94)00040-3}{10.1016/0378-3758(94)00040-3}
\harvarditem{Zhang}{1998}{zhang1998}
Zhang, B. \harvardyearleft 1998\harvardyearright , `A note on kernel density
estimation with auxiliary information', {\em Communications in
Statistics---Theory and Methods} {\bf 27}(1), 1--11.
doi:
\href{http://dx.doi.org/10.1080/03610929808832647}{10.1080/03610929808832647}
\harvarditem{Zygmund}{2003}{zygmund2003}
Zygmund, A. \harvardyearleft 2003\harvardyearright , {\em Trigonometric
series}, 3rd edn, Cambridge University Press.
\setcounter{page}{1}
\setcounter{section}{0}
tabular[tabular omitted — 345 chars of source]
}
\phantomsection
\addcontentsline{toc}{section}{Supplement A: Proofs}
\textcolor{white}{
\protected@write \@auxout {\string \newlabel {Supp:Proofs}{{A}{[B.\arabic{page}]}{A}{Supp:Proofs}} }
\hypertarget{Supp:Proofs}{A}
}
Throughout the Appendix, $0<C<\infty$ and $0\leq \omega\leq1$ will denote generic constants that may be different in different uses. CS, T, and H refer to the Cauchy-Schwarz, triangle, and H\"{o}lder inequalities, respectively with LIE and WLLN the law of iterated expectations and Khintchine's i.i.d. weak law of large numbers. MVT is the mean value theorem.
In addition, $\mathop{\mathrm{int}}(\cdot)$ denotes the interior of $\cdot$, w.p.(a.)1 with probability (approaching) 1, and $\mathcal{N}$ is an open neighbourhood of $\beta_{0}$.
GEL Stochastic Expansions
The following identification and regularity conditions are imposed.
assumption(a) $\beta_{0}\in\mathcal{B}$ is the unique solution to $\operatorname{E}[g(z,\beta)]=0$;
(b) $\mathcal{B}$ is compact;
(c) $g(z,\beta)$ is continuous at each $\beta\in\mathcal{B}$ w.p.1;
(d) $\operatorname{E}[\sup_{\beta\in\mathcal{B}}\lVertg(z,\beta)\rVert^{2}]<\infty$;
(e) $\Omega$ is nonsingular;
(f) $\rho(v)$ is twice continuously differentiable in a neighbourhood of zero.
Assumption (ref) is newey2004 and is sufficient for the consistency of $\hat{\beta}$. Moreover, $\hat{\lambda}=\operatorname*{argmax}_{\lambda\in\Lambda_{n}(\hat{\beta})}P_{n}(\hat{\beta},\lambda)$ exists w.p.a.1 and $\hat{\lambda}=\mathit{O}_{p}(n^{-1/2})$; see newey2004.
assumption(a) $\beta_{0}\in\mathop{\mathrm{int}}(\mathcal{B})$;
(b) $g(z,\beta)$ is continuously differentiable for $\beta\in\mathcal{N}$ and \\
$\operatorname{E}[\sup_{\beta\in\mathcal{N}}\lVert\nabla g(z,\beta)\rVert]<\infty$;
(c) $\mathop{\mathrm{rank}}(G)=d_{\beta}$.
Assumption (ref) is newey2004. If Assumptions (ref) and (ref) hold then $n^{1/2}((\hat{\beta}-\beta_{0})^{\top},\:\hat{\lambda}^{\top})^{\top}\xrightarrow{d}N\left(0,\mathop{\mathrm{diag}}(\Sigma,P)\right)$; see newey2004.
Let $\nabla^{2}g(z,\beta)$ denote a vector of all distinct second order partial derivatives with respect to $\beta$.
assumption(a) $\operatorname{E}[\lVertg(z,\beta_{0})\rVert^{6}]<\infty$;
(b) $g(z,\beta)$ is twice differentiable for $\beta\in\mathcal{N}$,
$\operatorname{E}[\lVert\nabla g(z,\beta_{0})\rVert^{4}]$ $<\infty$,
$\operatorname{E}[\lVert\nabla^{2}g(z,\beta_{0})\rVert^{2}]<\infty$;
(c) there exists $d(z)\geq0$ with $\operatorname{E}[d(z)^{2}]<\infty$ such that $\lVert\nabla^{2}g(z,\beta)-\nabla^{2}g(z,\beta_{0})\rVert\leq d(z)\lVert\beta-\beta_{0}\rVert$ for all $z$ and $\beta\in\mathcal{N}$;
(d) $\rho(v)$ is four times differentiable with Lipschitz fourth derivative in a neighbourhood of zero.
Cf. newey2004.
Write $\tilde{g}=n^{-1}\sum_{i=1}^{n}g_{i}$, $\widetilde{G}=n^{-1}\sum_{i=1}^{n}G_{i}-G$, and $\widetilde{\Omega}=n^{-1}\sum_{i=1}^{n}g_{i}g_{i}^{\top}-\Omega$. Also let $g_{i}^{j}=\partial g(z_{i},\beta_{0})/\partial\beta_{j}$ and $G_{i}^{j}=\partial^{2}g(z_{i},\beta_{0})/\partial\beta_{j}\partial\beta^{\top}$, $j=1,\ldots,d_{\beta}$.
From the proof of Theorem 3.4 in newey2004, GEL estimators satisfy the following stochastic expansion
equation[equation omitted — 256 chars of source]
where
align*[align* omitted — 635 chars of source]
remarkWrite $\tilde{\zeta}=(\tilde{\zeta}_{\beta}^{\top},\tilde{\zeta}_{\lambda}^{\top})^{\top}$ partitioned conformably with $\beta$ and $\lambda$. Then $\operatorname{E}[\tilde{\zeta}_{\beta}]=0$ and $\operatorname{E}[\tilde{\zeta}_{\lambda}]=\zeta_{\lambda}$ given in eq. (ref).
If $\beta_{0}$ is known, the stochastic expansion for $\tilde{\lambda}$ is identical to that in eq. (ref) except $H$ is set to zero and $\Omega^{-1}$ replaces $P$, i.e.,
$\tilde{\lambda}=-\Omega^{-1}\tilde{g} + \Omega^{-1}\tilde{\zeta}_{\lambda}+\mathit{O}_{p}(n^{-3/2})$, where
$\tilde{\zeta}_{\lambda} = \widetilde{\Omega}\Omega^{-1}\tilde{g}
+\rho_{3}\sum_{j=1}^{d_{\beta}}[\Omega^{-1}\tilde{g}]_{j}\operatorname{E}[g_{ij}g_{i}g_{i}^{\top}]\Omega^{-1}\tilde{g}/2$. Thus, in expectation, the first two terms in eq. (ref) are eliminated and
$\operatorname{E}[\tilde{\zeta}_{\lambda}] = n^{-1}c_{\rho}\operatorname{E}[g_{i}g_{i}^{\top}\Omega^{-1}g_{i}]$.
remarkWhen $\beta_{0}$ is known, Assumptions (ref)(b,c) can be relaxed to $g(z,\beta)$ is continuously differentiable for $\beta\in\mathcal{N}$, $\operatorname{E}[\sup_{\beta\in\mathcal{N}}\lVert\nabla g(z,\beta_{0})\rVert]<\infty$, and there exists $d(z)\geq0$ with $\operatorname{E}[d(z)]<\infty$ such that $\lVert\nabla g(z,\beta)-\nabla g(z,\beta_{0})\rVert\leq d(z)\lVert\beta-\beta_{0}\rVert$ for all $z$ and $\beta\in\mathcal{N}$. The Lipschitz condition in Assumptions (ref)(b,c,d) can also be relaxed to $\alpha$-H\"{o}lder for some $0<\alpha\leq1$ and changing the remainder terms from $\mathit{O}_{}(n^{-3/2})$ to $\mathit{O}_{}(n^{-1-\alpha/2})$.
remarkThe two-step GMM estimator is defined as $\hat{\beta}_{GMM}=\operatorname*{argmin}_{\beta\in\mathcal{B}}\hat{g}(\beta)^{\top}\widehat{\Omega}(\tilde{\beta})^{-1}\hat{g}(\beta)$ where $\tilde{\beta}$ is a $\sqrt{n}$-consistent preliminary estimator of $\beta_{0}$. If the preliminary estimator $\tilde{\beta}$ is first order efficient, i.e., $\tilde{\beta}-\beta_{0}=-H\tilde{g}+\mathit{O}_{p}(n^{-1})$, then, if Assumptions (ref)--(ref) hold, all GMM estimators $\hat{\beta}_{GMM}$ admit the same expansion to order $\mathit{O}_{p}(n^{-3/2})$; see newey2004. Moreover, defining $\hat{\lambda}_{GMM}=-\widehat{\Omega}(\tilde{\beta})^{-1}\hat{g}(\hat{\beta}_{GMM})$, the expansion is
\begin{equation*}
\begin{bmatrix}\hat{\beta}_{GMM}-\beta_{0}\\\hat{\lambda}_{GMM}\end{bmatrix}
= -\begin{bmatrix}H\\P\end{bmatrix}\tilde{g} + \begin{bmatrix}-\Sigma & H\\ H^{\top} & P\end{bmatrix}\tilde{\zeta}^{GMM}+\mathit{O}_{p}(n^{-3/2}),
\end{equation*}
where
\begin{align*}
\tilde{\zeta}^{GMM} & = \left\{\begin{bmatrix}0& \widetilde{G}^{\top} \\\widetilde{G} &\widetilde{\Omega}-\sum_{j=1}^{d_{\beta}}\operatorname{E}[g_{i}^{j}g_{i}^{\top}+g_{i}g_{i}^{j\top}]e_{j}^{\top}H\tilde{g}\end{bmatrix}
-\frac{1}{2}\sum_{j=1}^{d_{\beta}}[H\tilde{g}]_{j}\begin{bmatrix}
0 & E[G_{i}^{j}]^{\top}
\\ E[G_{i}^{j}] & 0 \end{bmatrix}
\right. \\ & \qquad \left.
-\frac{1}{2}\sum_{j=1}^{d_{g}}[P\tilde{g}]_{j}\begin{bmatrix}
E[\partial^{2}g_{ij}/\partial\beta\partial\beta^{\top}] &
0 \\ 0 & 0
\end{bmatrix}
\right\}\begin{bmatrix}H\\P\end{bmatrix}\tilde{g}.
\end{align*}
Writing $\tilde{\zeta}^{GMM}=(\tilde{\zeta}_{\beta}^{GMM\top},\tilde{\zeta}_{\lambda}^{GMM\top})^{\top}$ partitioned conformably with $\beta$ and $\lambda$, $\zeta_{\beta}^{GMM}=\operatorname{E}[\tilde{\zeta}_{\beta}^{GMM}]=\operatorname{E}[G_{i}^{\top}Pg_{i}]$ and $\zeta_{\lambda}^{GMM}=\operatorname{E}[\tilde{\zeta}_{\lambda}]=-a+\operatorname{E}[G_{i}Hg_{i}]+\operatorname{E}[g_{i}g_{i}^{\top}Pg_{i}]$. Hence, the second order bias of $\hat{\beta}_{GMM}$, newey2004, is given by
\begin{equation*}
\operatorname{E}[\hat{\beta}_{GMM}]-\beta _{0}=-n^{-1}\Sigma\zeta_{\beta}^{GMM}+n^{-1}H\zeta_{\lambda}^{GMM}+\mathit{O}_(n^{-3/2}),
\end{equation*}
the notable difference with GEL being the additional term $-n^{-1}\Sigma\zeta _{\beta}^{GMM}$ with the term $n^{-1}H\zeta_{\lambda}^{GMM}$ identical to CUE.
Preliminary Lemmas
lemmaIf Assumptions (ref)--(ref) are satisfied, then
\begin{equation}
n\hat{\pi}_{i}=1 -g_{i}^{\top}P\tilde{g} -\tfrac{\rho_{3}}{2}(g_{i}^{\top}P\tilde{g})^{2}
+ g_{i}^{\top}\begin{bmatrix}H^{\top}&P\end{bmatrix}\tilde{\zeta}
+\tilde{g}^{\top}PG_{i}H\tilde{g}
+c_{\rho}\tilde{g}^{\top}P\tilde{g}
+ \mathit{o}_{p}(n^{-1})
\end{equation}
uniformly $i=1,\ldots,n$.
\noindentProof. Let $\hat{v}_{i}=\hat{\lambda}^{\top}g(z_{i},\hat{\beta})$.
A third order Taylor expansion of $\rho^{(1)}(\hat{v}_{i})$ around $0$ yields
equation*[equation* omitted — 148 chars of source]
noting $\lvert\hat{v}_{i}\rvert\xrightarrow{p}0$ uniformly $i=1,\ldots,n$ by newey2004.
A Taylor expansion from eq. (ref) of $g(z_{i},\hat{\beta})$ about $\beta_{0}$ yields
$g(z_{i},\hat{\beta})=g_{i}-G_{i}H\tilde{g}+\mathit{o}_{p}(n^{-1/2})$ uniformly $i=1,\ldots,n$ by owen1990. Hence, substituting, using eq. (ref),
equation*[equation* omitted — 259 chars of source]
From a similar expansion, using $n^{-1}\sum_{i=1}^{n}g(z_{i},\hat{\beta})=\Omega P\tilde{g}+\mathit{O}_{p}(n^{-1})$,
eq. (ref), and $P\Omega P=P$,
equation*[equation* omitted — 286 chars of source]
Hence,
$[{\begingroup\textstyle\sum\endgroup}_{i=1}^{n}\rho^{(1)}(\hat{v}_{i})]^{-1}
= -n^{-1}[1+c_{\rho}\tilde{g}^{\top}P\tilde{g}+\mathit{O}_{p}(n^{-3/2})]$ and
equation*[equation* omitted — 284 chars of source]
uniformly $i=1,\ldots,n$. $\blacksquare$
corollary[Known $\beta_{0}$] If Assumptions (ref)--(ref) are satisfied, then
\begin{equation}
n\tilde{\pi}_{i}=1 -g_{i}^{\top}\Omega^{-1}\tilde{g} -\tfrac{\rho_{3}}{2}(g_{i}^{\top}\Omega^{-1}\tilde{g})^{2}
+ g_{i}^{\top}\Omega^{-1}\tilde{\zeta}_{\lambda}
+c_{\rho}\tilde{g}^{\top}\Omega^{-1}\tilde{g}
+ \mathit{o}_{p}(n^{-1})
\end{equation}
uniformly $i=1,\ldots,n$.
Let $a(z)$ denote a real scalar function of $z$ such that $\operatorname{E}[a(z)^{2}]<\infty$. Write $a_{i}=a(z_{i})$, $i=1,\ldots,n$.
lemmaIf Assumptions (ref)--(ref) are satisfied, then
\begin{equation}
\operatorname{E}[(n\hat{\pi}_{i}-1)a_{i}] = n^{-1}\left(
-c_{\rho}\operatorname{E}[a_{i}g_{i}^{\top}Pg_{i}]
+ \operatorname{E}[a_{i}g_{i}^{\top}]P\zeta_{\lambda}
+c_{\rho}(d_{g}-d_{\beta})\operatorname{E}[a_{i}]\right)
+ \mathit{o}_(n^{-1})
\end{equation}
uniformly $i=1,\ldots,n$. For $i\neq j$,
\begin{equation}
\operatorname{E}[(n\hat{\pi}_{i}-1)a_{i}a_{j}] = \operatorname{E}[(n\hat{\pi}_{i}-1)a_{i}]\operatorname{E}[a_{j}]-n^{-1}\operatorname{E}[a_{i}g_{i}^{\top}]P\operatorname{E}[g_{j}a_{j}] +\mathit{O}_(n^{-2}),
\end{equation}
\begin{equation}
\operatorname{E}[(n\hat{\pi}_{i}-1)(n\hat{\pi}_{j}-1)a_{i}a_{j}] = n^{-1}\operatorname{E}[a_{i}g_{i}^{\top}]P\operatorname{E}[g_{j}a_{j}] +\mathit{O}_(n^{-2}).
\end{equation}
Let $\bar{a}=n^{-1}\sum_{i=1}^{n}a_{i}$ and $\hat{a}=\sum_{i=1}^{n}\hat{\pi}_{i}a_{i}$. Then,
\begin{equation}
\operatorname{Var}[\hat{a}] = \operatorname{Var}[\bar{a}] - n^{-1}\operatorname{E}[a_{i}g_{i}^{\top}]P\operatorname{E}[g_{j}a_{j}] + \mathit{O}_(n^{-2}).
\end{equation}
\noindentProof. The first result follows from the expansion for $\hat{\pi}_{i}$ in Lemma (ref). In particular, noting $\operatorname{E}[g_{i}]=0$ and $\operatorname{E}[a_{i}\mathit{o}_{p}(n^{-1})]=\mathit{o}_{}(n^{-1})$ by uniformity of $\mathit{o}_{p}(n^{-1})$, then, by independence,
align*[align* omitted — 752 chars of source]
uniformly $i=1,\ldots,n$, using $\operatorname{E}[\tilde{\zeta}]=(0^{\top},n^{-1}\zeta_{\lambda}^{\top})^{\top}$, $P\Omega P=P$, $H\Omega P=0$, and $\mathop{\mathrm{tr}}(\Omega P)=d_{g}-d_{\beta}$.
Eqs. (ref) and (ref) follow by a similar argument.
Finally note that $\hat{a}-\bar{a}=n^{-1}\sum_{i=1}^{n}(n\hat{\pi}_{i}-1)a_{i}$. Hence, $\operatorname{Var}[\hat{a}] = \operatorname{Var}[\bar{a}]+\operatorname{Var}[\hat{a}-\bar{a}]+2\operatorname{Cov}[\hat{a}-\bar{a},\bar{a}]$. Now, from above, $\operatorname{E}[\hat{a}-\bar{a}]=\mathit{O}_{}(n^{-1})$. Hence,
equation*[equation* omitted — 250 chars of source]
Also,
align*[align* omitted — 354 chars of source]
corollary[Known $\beta_{0}$] If Assumptions (ref)--(ref) are satisfied, then
\begin{equation}
\operatorname{E}[(n\tilde{\pi}_{i}-1)a_{i}] = n^{-1}c_{\rho}\left(
-\operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i} a_{i}]
+ \operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}g_{i}^{\top}]\Omega^{-1}\operatorname{E}[g_{i}a_{i}]
+q\operatorname{E}[a_{i}]\right)
+ \mathit{o}_(n^{-1})
\end{equation}
uniformly $i=1,\ldots,n$. Lemma (ref) remains valid with $\Omega^{-1}$ replacing $P$.
Repeated use is made of the following lemma; see bochner1955 and parzen1962. See also pagan1999.
lemmaSuppose that $f:\mathbbm{R}\mapsto\mathbbm{R}$ and $k:\mathbbm{R}\mapsto\mathbbm{R}$ are Borel functions satisfying
(a) $\int_{-\infty}^{\infty}\lvertf(x)\rvert\mathrm{d}x < \infty$;
(b) $\sup_{-\infty<x<\infty}\lvertk(x)\rvert<\infty$, $\int_{-\infty}^{\infty}\lvertk(x)\rvert\mathrm{d}x<\infty$, and $\lim_{\lvertx\rvert\to\infty}\lvertxk(x)\rvert=0$.
Then ${\begingroup\textstyle\int\endgroup}_{-\infty}^{\infty} b^{-1}\lvertk((y-x)/b)\rvert\lvertf(x)\rvert\mathrm{d}x<\infty$ a.e. and
\begin{equation}
\lim_{b\downarrow0}\lvert{\begingroup\textstyle\int\endgroup}_{-\infty}^{\infty} b^{-1}k((y-x)/b)f(x)\mathrm{d}x - f(y){\begingroup\textstyle\int\endgroup}_{-\infty}^{\infty} k(t)\mathrm{d}t\rvert=0
\end{equation}
at every continuity point $y$ of $f$; if $f$ is uniformly continuous, then convergence is uniform.
Under the same conditions
$\lim_{b\downarrow0}\lvert\int_{-\infty}^{\infty} b^{-1}k((y-x)/b)^{r}f(x)\mathrm{d}x - f(y)\int_{-\infty}^{\infty} k(t)^{r}\mathrm{d}t\rvert=0$ at every continuity point $y$ of $f$ for any $r\geq1$.
If $\sup_{-\infty<x<\infty}\lvertf(x)\rvert<\infty$, $\int_{-\infty}^{\infty}\lvertk(x)\rvert\mathrm{d}x<\infty$ is sufficient for (ref) to hold.
$\square$
remarkIf $k$ is H\"{o}lder continuous with exponent $0<\tau\leq1$ and, thus, uniformly continuous, and absolutely integrable, then it is bounded.
Proofs of Theorems
\noindentProof of Theorem (ref). Write
$\tilde{f}_{\rho}(u) = \tilde{f}(u) + n^{-1}\sum_{i=1}^{n}(n\tilde{\pi}_{i}-1)k_{b}(u-u_{i})$. By Corollary (ref) and owen1990, $\max_{1\leq i\leq n}\lvertn\tilde{\pi}_{i}-1\rvert=\mathit{o}_{p}(1)$. By Lemma (ref), $\operatorname{E}[\lvertk_{b}(u-u_{i})\rvert]<\infty$ whenever $\lvertf(u)\rvert<\infty$ which holds a.e. Thus, $k_{b}(u-u_{i})$, $i=1,\ldots,n$, satisfies the conditions for WLLN. Hence, the first conclusion follows.
From Assumption (ref)(a)(i), $b\operatorname{E}[k_{b}(u-u_{i})^{2}]<\infty$ a.e.
By CS, invoking Assumptions (ref)(e) and (ref)(a), $\operatorname{E}[\lvertg_{i}b^{1/2}k_{b}(u-u_{i})\rvert]<\infty$ and $\operatorname{E}[\lvertg_{i}^{\top}\Omega^{-1}g_{i}b^{1/2}k_{b}(u-u_{i})\rvert]<\infty$. Hence, by
Corollary (ref),
setting $a_{i}=b^{1/2}k_{b}(u-u_{i})$,
align*[align* omitted — 329 chars of source]
Under Assumption (ref)(a)(i), $\operatorname{E}[k_{b}(u-u_{i})]=f(u)+\mathit{o}_{}(1)$.
Invoking Assumption (ref) and the change of variables $z\mapsto(u,v^{\top})^{\top}$,
then, by LIE and Lemma (ref),
$\operatorname{E}[g_{i}k_{b}(u-u_{i})]
= \int \operatorname{E}[g_{i}|t]f(t)k_{b}(u-t)\mathrm{d}t
= \operatorname{E}[g_{i}|u]f(u) + \mathit{o}_{}(1)$.
Similarly,
$\operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}k_{b}(u-u_{i})] = \operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}|u]f(u)+\mathit{o}_{}(1)$. The final result is a direct consequence of Lemma (ref) and the same argument. $\blacksquare$
Set
align[align omitted — 483 chars of source]
Note $\hat{f}(u) = \tilde{f}(u) + \hat{\delta}_{1}(u)$ and
$\hat{f}_{\rho}(u) = \hat{f}(u) +\hat{\delta}_{2}(u)+\hat{\delta}_{3}(u)$.
\noindentProof of Theorem (ref).
Under Assumptions (ref) and (ref), $\hat{\beta}\in\mathcal{N}$ w.p.a.1 and $n^{1/2}(\hat{\beta}-\beta_{0})=\mathit{O}_{p}(1)$. First, by Assumption (ref)(a,b), from eq. (ref),
align*[align* omitted — 465 chars of source]
since $n^{-1}\sum_{i=1}^{n}d(z_{i})^{\tau}=\mathit{O}_{p}(1)$ by WLLN and $n^{\alpha\tau/2}b^{1+\tau}\to\infty$ from
Assumption (ref)(c).
Next, $\max_{1\leq i\leq n}\lvertn\hat{\pi}_{i}-1\rvert=\mathit{o}_{p}(1)$ by Lemma (ref) and owen1990, from eq. (ref),
align*[align* omitted — 436 chars of source]
Hence, the first conclusion follows.
The final result follows from eq. (ref), by noting that also
equation*[equation* omitted — 316 chars of source]
by WLLN since $\operatorname{E}[\lvertk_{b}(u-u_{i})\rvert]<\infty$ a.e. by Lemma (ref).
$\blacksquare$
\noindentProof of Theorem (ref). Preliminaries.
From a second order Taylor expansion around $\beta_{0}$,
align*[align* omitted — 314 chars of source]
where $k_{b}^{(j)}(x)=k^{(j)}(x/b)/b^{j+1}$, $j=1,2$, and $\bar{u}_{i}=u(z_{i},\bar{\beta})$, $i=1,\ldots,n$, with $\bar{\beta}$ on the line segment joining $\hat{\beta}$ and $\beta_{0}$; and $\nabla\bar{u}_{i}$ and $\nabla^{2}\bar{u}_{i}$, $i=1,\ldots,n$, are defined analogously. Note that
$\lVert\bar{\beta}-\beta_{0}\rVert\leq \lVert\hat{\beta}-\beta_{0}\rVert=\mathit{O}_{p}(n^{-1/2})$.
Assumption (ref)(b) and twice differentiability of $u(z,\beta)$ for $\beta\in\mathcal{N}$ implies there exist $d_{0}(z)\geq0$ with $\operatorname{E}[d_{0}(z)^{4}]<\infty$ and $d_{1}(z)\geq0$ with $\operatorname{E}[d_{1}(z)^{4}]<\infty$
such that $\lvertu(z,\beta)-u(z,\beta_{0})\rvert\leq d_{0}(z)\lVert\beta-\beta_{0}\rVert$ and
$\lVert\nabla u(z,\beta)-\nabla u(z,\beta_{0})\rVert\leq d_{1}(z)\lVert\beta-\beta_{0}\rVert$ for all $z$ and $\beta\in\mathcal{N}$.
Thus, by T,
$\lVert\nabla \bar{u}_{i}\nabla^{\top} \bar{u}_{i}-\nabla u_{i}\nabla^{\top}u_{i}\rVert
\leq 2d_{1}(z_{i})\lVert\nabla u_{i}\rVert\lVert\hat{\beta}-\beta_{0}\rVert +d_{1}(z_{i})^{2}\lVert\hat{\beta}-\beta_{0}\rVert^{2}$. By owen1990, $\max_{1\leq i\leq n}d_{1}(z_{i})^{2}=\mathit{o}_{p}(n^{1/2})$ and
$\max_{1\leq i\leq n}d_{1}(z_{i})\lVert\nabla u_{i}\rVert=\mathit{o}_{p}(n^{1/2})$.
Hence, $\lVert\nabla \bar{u}_{i}\nabla^{\top} \bar{u}_{i}-\nabla u_{i}\nabla^{\top}u_{i}\rVert
\leq n^{-1/2}[d_{1}(z_{i})\lVert\nabla u_{i}\rVert+\mathit{o}_{p}(1)]\mathit{O}_{p}(1)$.
By CS, from Assumption (ref)(b), for $0<\tau\leq1$,
$\operatorname{E}[d_{0}(z_{i})^{\tau}\lVert\nabla u_{i}\rVert^{2}]<\infty$ and $\operatorname{E}[d_{1}(z_{i})^{2}\lVert\nabla u_{i}\rVert^{2}]<\infty$.
Thus, $\operatorname{E}[\lvertb^{-1/2}k^{(2)}((u-u_{i})/b)\rvertd_{1}(z_{i})\lVert\nabla u_{i}\rVert]<\infty$ since
$\operatorname{E}[b^{-1}k^{(2)}((u-u_{i})/b)^{2}]<\infty$ also by CS and using Lemma (ref). Hence, by T, and noting
$n^{\tau/2}b^{3+\tau}\to\infty$, $0<\tau\leq1$, from Assumption (ref)(c),
align*[align* omitted — 722 chars of source]
Assumption (ref)(a) implies $k^{(1)}$ is Lipschitz, and hence, invoking
Assumption (ref)(b), for all mean values $\bar{\beta}$ between $\hat{\beta}$ and $\beta_{0}$,
$\lvertk_{b}^{(1)}(u-\bar{u}_{i}) -k_{b}^{(1)}(u-u_{i})\rvert\leq b^{-3}Cd_{0}(z_{i})\lVert\hat{\beta}-\beta_{0}\rVert$ w.p.a.1. By Assumption (ref)(a) and Lemma (ref), $\operatorname{E}[b^{-1}\vert k^{(1)}((u-u_{i})/b)\vert^{4/3}]<\infty$ a.e., and as $\operatorname{E}[d(z_{i})^{4}]<\infty$, $\operatorname{E}[\lvertb^{-3/4}k^{(1)}((u-u_{i})/b)\rvertd(z_{i})]<\infty$ using the H\"{o}lder inequality with exponents $4/3$ and $4$.
Therefore, by the same argument as above,
align*[align* omitted — 649 chars of source]
Using expansion eq. (ref) and Lemma (ref) eq. (ref), from eq. (ref),
align[align omitted — 570 chars of source]
from eq. (ref),
equation[equation omitted — 229 chars of source]
and, from eq. (ref),
align[align omitted — 409 chars of source]
\noindentExpectation. Since $H\Omega H^{\top}=\Sigma$, from eq. (ref),
align*[align* omitted — 407 chars of source]
Assumption (ref)(a) states $\lim_{\lvertx\rvert\to\infty}\lvertx^{2}k^{(1)}(x)\rvert=0$ and implies that $\int k^{(1)}(x)\mathrm{d}x=0$, $\int xk^{(1)}(x)\mathrm{d}x=-1$, and $xk^{(1)}(x)$ satisfies the hypotheses of Lemma (ref), i.e., it is bounded and absolutely integrable.
Thus, invoking Assumption (ref)(d), by MVT and Lemma (ref),
align[align omitted — 592 chars of source]
Similarly, $\operatorname{E}[k_{b}^{(1)}(u-u_{i})\nabla^{\top}u_{i}Hg_{i}] = \mathrm{d}\{\operatorname{E}[\nabla^{\top}u_{i}Hg_{i}|u]f(u)\}/\mathrm{d}u+ \mathit{o}_{}(1)$ and
$\operatorname{E}[k_{b}^{(1)}(u-u_{i})\nabla^{2}u_{i}] = \mathrm{d}\{\operatorname{E}[\nabla^{2}u_{i}|u]f(u)\}/\mathrm{d}u+ \mathit{o}_{}(1)$.
Furthermore, Assumption (ref)(a) also implies that $\int k^{(2)}(x)\mathrm{d}x=0$, $\int xk^{(2)}(x)\mathrm{d}x=0$, $\int x^{2}k^{(2)}(x)\mathrm{d}x=2$, and $x^{2}k^{(2)}(x)$ satisfies the hypotheses of Lemma (ref). Thus, by a second order Taylor expansion and a similar argument to eq. (ref),
align*[align* omitted — 824 chars of source]
Since $H\Omega P=0$, from eq. (ref), $\operatorname{E}[\hat{\delta}_{2}(u)]=\mathit{o}_{}(n^{-1})$. By Lemma (ref) eq. (ref) and the same argument used in the proof of Theorem (ref),
$\operatorname{E}[\hat{\delta}_{3}(u)] = n^{-1}\{
-c_{\rho}\operatorname{E}[g_{i}^{\top}Pg_{i}|u] +\operatorname{E}[g_{i}|u]^{\top}P\zeta_{\lambda} + c_{\rho}(d_{g}-d_{\beta})\}f(u)
+\mathit{o}_{}(n^{-1})$.
\noindentVariance. Since $\operatorname{E}[\hat{\delta}_{1}(u)]=\mathit{O}_{}(n^{-1})$, from eq. (ref),
align*[align* omitted — 481 chars of source]
Similarly, noting $\operatorname{E}[\hat{\delta}_{2}(u)]=\mathit{o}_{}(n^{-1})$, from Lemma (ref), it is straightforward to verify that $\operatorname{Var}[\hat{\delta}_{2}(u)]=\mathit{o}_{}(n^{-1})$. Furthermore, also using Lemma (ref), as $\operatorname{E}[\hat{\delta}_{3}(u)]=\mathit{O}_{}(n^{-1})$ and $\operatorname{E}[k_{b}(u-u_{i})g_{i}]=\operatorname{E}[g_{i}|u]f(u)$,
$\operatorname{Var}(\hat{\delta}_{3}(u)) = n^{-1}\operatorname{E}[g_{i}|u]^{\top}P\operatorname{E}[g_{i}|u]f(u)^{2}+\mathit{o}_{}(n^{-1})$.
It is straightforward to verify that
$\operatorname{Cov}[\hat{\delta}_{1},\hat{\delta}_{2}]=\mathit{o}_{}(n^{-1})$, recalling $H\Omega P=0$,
equation*[equation* omitted — 327 chars of source]
align*[align* omitted — 335 chars of source]
$\operatorname{Cov}[\hat{\delta}_{2}(u),\hat{\delta}_{3}(u)]=\mathit{o}_{}(n^{-1})$, $\operatorname{Cov}[\hat{\delta}_{2}(u),\tilde{f}(u)]=\mathit{o}_{}(n^{-1})$, noting again $H\Omega P=0$, and finally,
equation*[equation* omitted — 168 chars of source]
Combining these results gives eqs.
(ref)--(ref).
$\blacksquare$
\noindentProof of Theorem (ref).
Since $\lim_{x\to-\infty}K(x)=0$ and $\lim_{x\to\infty}K(x)=1$, $2\int K(x)k(x)\mathrm{d}x=1$, and $\int\lvertK(x)k(x)\rvert\mathrm{d}x<\infty$,
$\operatorname{E}[K((u-u_{i})/b)] = F(u)+\int k(t)[F(u-bt)-F(u)]\mathrm{d}t$
and
$\operatorname{E}[K((u-u_{i})/b)^{2}] = F(u)+2\int K(t)k(t)[F(u-bt)-F(u)]\mathrm{d}t$.
$F$ as a c.d.f. is bounded and hence $\operatorname{E}[K((u-u_{i})/b)^{2}]<\infty$, and $\operatorname{E}[K((u-u_{i})/b)] = F(u)+\mathit{o}_{}(1)$ and $\operatorname{E}[K((u-u_{i})/b)^{2}] = F(u)+\mathit{o}_{}(1)$ as $b\to0$ and at all points of continuity of $F$.
Therefore, cf. the proof of Theorem (ref),
$\lvert\widetilde{F}_{\rho}(u)-\widetilde{F}(u)\rvert=\mathit{o}_{p}(1)$.
Equation (ref) follows by Corollary (ref) with $a_{i}=K((u-u_{i})/b)$, $i=1,\ldots,n$.
Assumptions (ref)(a)(i) and $\lim_{\lvertx\rvert\to\infty}\lvertx^{2}k(x)\rvert=0$ imply that $xk(x)$ satisfies conditions of Lemma (ref). Since $\int xk(x)=0$ and $\operatorname{E}[\lvert\operatorname{E}[g_{i}|u]\rvert]<\infty$, integration by parts and an application of MVT give
align[align omitted — 682 chars of source]
Similarly,
$\operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}K((u-u_{i})/b)] =\int_{-\infty}^{u}\operatorname{E}[g_{i}^{\top}\Omega^{-1}g_{i}|t]\mathrm{d}F(t)$. Eq. (ref) follows by Corollary (ref) and eq. (ref). $\blacksquare$
Set
align[align omitted — 469 chars of source]
Note
$\widetilde{F}(u) = \widehat{F}(u)+ \widehat{\Delta}_{1}(u)$ and
$\widehat{F}_{\rho}(u) = \widehat{F}(u) +\widehat{\Delta}_{2}(u)+\widehat{\Delta}_{3}(u)$.
\noindentProof of Theorem (ref).
Since $k$ is bounded, $K$ is Lipschitz continuous and, by the proof of Theorem (ref), $\operatorname{E}[\lvertK((u-u_i)/b)\rvert]<\infty$ for all $u$. Then, as in the proof of Theorem (ref), invoking Assumptions (ref)--(ref) and (ref)(b), from eqs.
(ref)--(ref),
align*[align* omitted — 693 chars of source]
\noindentProof of Theorem (ref). Preliminaries.
From a second order Taylor expansion around $\beta_{0}$,
align*[align* omitted — 301 chars of source]
where $\bar{u}_{i}=u(z_{i},\bar{\beta})$, $i=1,\ldots,n$, with $\bar{\beta}$ on the line segment joining $\hat{\beta}$ and $\beta_{0}$; $\nabla\bar{u}_{i}$ and $\nabla^{2}\bar{u}_{i}$, $i=1,\ldots,n$, are defined analogously.
By the same argument as in the proof of Theorem (ref), noting that Assumption (ref)(a) implies $k$ is Lipschitz and $nb^{6}\to\infty$ as $n^{\tau/2}b^{2+\tau}\to\infty$, invoking Assumption (ref)(b),
align*[align* omitted — 717 chars of source]
and
align*[align* omitted — 636 chars of source]
Therefore, using expansion eq. (ref) and Lemma (ref),
align[align omitted — 1,157 chars of source]
\noindentExpectation. Similarly to the proof of Theorem (ref), from eq. (ref),
align*[align* omitted — 393 chars of source]
Assumption (ref)(a) implies $k(x)$ satisfies the hypotheses of of Lemma (ref). Hence
$\operatorname{E}[k_{b}(u-u_{i})\nabla u_{i}]
= \operatorname{E}[\nabla u_{i}|u]f(u)+\mathit{o}_{}(1)$,
$\operatorname{E}[k_{b}(u-u_{i})\nabla^{\top} u_{i}Hg_{i}] = \operatorname{E}[\nabla^{\top}u_{i}Hg_{i}|u]f(u)+\mathit{o}_{}(1)$, and
$\operatorname{E}[k_{b}(u-u_{i})\nabla^{2}u_{i}] = \operatorname{E}[\nabla^{2}u_{i}|u]f(u)+\mathit{o}_{}(1)$.
Assumption (ref)(a) also implies $xk^{(1)}(x)$ satisfies the hypotheses of Lemma (ref). Hence, by MVT as in eq. (ref),
$\operatorname{E}[k_{b}^{(1)}(u-u_{i})\nabla u_{i}\nabla^{\top}u_{i}]
= \mathrm{d}\{\operatorname{E}[\nabla u_{i}\nabla^{\top}u_{i}|u]f(u)\}/\mathrm{d}u+\mathit{o}_{}(1)$.
Therefore, $\operatorname{E}[\widehat{\Delta}_{1}(u)]=n^{-1}\Delta(u)+\mathit{o}_{}(n^{-1})$ as required.
Likewise, as in the proof of Theorem (ref), from eq. (ref), $\operatorname{E}[\widehat{\Delta}_{2}(u)]=\mathit{o}_{}(n^{-1})$. Finally, by Lemma (ref) and proof of Theorem (ref), $\operatorname{E}[\widehat{\Delta}_{3}(u)] = n^{-1}\Delta_{\rho}(u) +\mathit{o}_{}(n^{-1})$.
\noindentVariance. Using expansions eqs. (ref)--(ref) for $\widehat{\Delta}_{j}(u)$, $j=1,2,3$,
$\operatorname{Cov}[\widehat{\Delta}_{1}(u),\widehat{\Delta}_{2}(u)]$, \\
$\operatorname{Cov}[\widehat{\Delta}_{1}(u),\widehat{\Delta}_{3}(u)]$,
$\operatorname{Cov}[\widehat{\Delta}_{2}(u),\widehat{\Delta}_{3}(u)]$,
$\operatorname{Cov}[\widetilde{F}(u),\widehat{\Delta}_{2}(u)]$, and
$\operatorname{Var}[\widehat{\Delta}_{2}(u)]$ are all $\mathit{O}_{}(n^{-2})$.
Also,
align*[align* omitted — 730 chars of source]
Eqs. (ref) and (ref) then follow immediately using eq. (ref) and $\operatorname{E}[k_{b}(u-u_{i})\nabla u_{i}] = \operatorname{E}[\nabla u_{i}|u]f(u)+\mathit{o}_{}(1)$.
If $\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u$ is absolutely integrable, then, using Lemma (ref),
$\operatorname{E}[k_{b}(u-u_{i})\nabla u_{i}] = \operatorname{E}[\nabla u_{i}|u]f(u) - b\int(\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u-\omega bt]f(u-\omega bt)\}/\mathrm{d}u)tk(t)\mathrm{d}t =\operatorname{E}[\nabla u_{i}|u]f(u)+\mathit{o}_{}(b)$.
$\blacksquare$
\setcounter{page}{1}
\setcounter{section}{0}
tabular[tabular omitted — 345 chars of source]
}
\phantomsection
\addcontentsline{toc}{section}{Supplement B: Examples}
\textcolor{white}{
\protected@write \@auxout {\string \newlabel {Supp:Examples}{{B}{[B.\arabic{page}]}{B}{Supp:Examples}} }
\hypertarget{Supp:Examples}{B}
}
\phantomsection
\addcontentsline{toc}{section}{Example (ref): \texorpdfstring{$u$}{u} is not a function of \texorpdfstring{$\beta$}{beta}}
example[$u$ is not a function of $\beta$]
When $u=u(z)$ is a function of $z$ but not of $\beta$, $u_{i}$, $i=1,\ldots,n$, is of course observable. Hence the estimators $\tilde{f}$ eq. (ref) and $\hat{f}$ eq. (ref) are identical and the terms $\hat{\delta}_{1}$ and $\hat{\delta}_{2}$ in the proof of Theorem (ref) are zero.
The density estimators $\tilde{f}_{\rho}$ eq. (ref) and $\hat{f}_{\rho}$ eq. (ref) use different implied probabilities, $\tilde{\pi}_{i}$ versus $\hat{\pi}_{i}$, $i=1,\ldots,n$. Thus, Theorem (ref) with known $\beta_{0}$ is unchanged whereas, in Theorem (ref) with estimated $\beta_{0}$,
$\operatorname{E}[\hat{f}_{\rho}(u)]= \operatorname{E}[\tilde{f}(u)]+n^{-1}\delta_{\rho}(u)+\mathit{o}_{}(n^{-1})$ with $\delta_{\rho}(u)$ defined in eq. (ref). Eq. (ref) also holds with $\tilde{f}$ replacing $\hat{f}$.
Classical examples wherein, e.g., a mean, variance, or a third moment of $u$ are either fully or partially known, are included here. For instance, symmetry can be imposed by the moment condition that the third moment around an unknown mean is known to be zero.
This set-up also allows for situation in which the interest is in the density of $u(z_{1})$, say, but the remaining $d_{z}-1$ variates $z_{2}$ satisfy moment conditions $\operatorname{E}[g(z_{2},\beta_{0})]=0$. Provided $u(z_{1})$ and $g(z_{2},\beta_{0})$ are not independent, (G)EL-based estimators for $f$ will generally enjoy a reduction in variance due to the extra information from the moment condition $\operatorname{E}[g(z_{2},\beta_{0})]=0$.
\phantomsection
\addcontentsline{toc}{section}{Example (ref): Regression on a constant}
example[Regression On A Constant]
To explain the method behind the proof of Theorem (ref) and to provide the background for Example (ref) below, the estimation of the density of the residual $u$ from a regression on a constant is examined, viz.,
$y=\beta_{0}+u$, with $\beta_{0}$ estimated by the sample average $\hat{\beta}= \bar{y} = n^{-1}\sum_{i=1}^{n}y_{i} = \beta_{0}+\bar{u}$.
The estimated residuals are
$\hat{u}_{i} = y_{i}-\hat{\beta} = u_{i} - \bar{u}$, $i=1,\ldots,n$. If Assumption (ref)(a) holds, $\hat{f}(u) = \tilde{f}(u) + \hat{\delta}_{1}(u)$, where,
for some $0\leq\omega\leq1$,
\begin{equation*}
\hat{\delta}_{1}(u)=
n^{-1}{\begingroup\textstyle\sum\endgroup}_{i=1}^{n}k_{b}^{(1)}(u-u_{i})\bar{u}
+\tfrac{1}{2}n^{-1}{\begingroup\textstyle\sum\endgroup}_{i=1}^{n}k_{b}^{(2)}(u-u_{i})\bar{u}^{2}
+\tfrac{1}{2}n^{-1}{\begingroup\textstyle\sum\endgroup}_{i=1}^{n}[k_{b}^{(2)}(u-u_{i}+\omega\bar{u})-k_{b}^{(2)}(u-u_{i})]\bar{u}^{2}.
\end{equation*}
By H\"{o}lder continuity of $k^{(2)}$, for some $0<C<\infty$,
$\lvertk_{b}^{(2)}(u-u_{i}+\omega\bar{u})-k_{b}^{(2)}(u-u_{i})\rvert\leq
C\lvertn^{1/2}\bar{u}\rvert^{\tau}/n^{\tau/2}b^{3+\tau}\\\to0$ in probability if $n^{\tau/2}b^{3+\tau}\to\infty$, and in mean square if $\operatorname{E}[u^{4}]<\infty$.
Furthermore, for some $\epsilon>0$, $n^{(1-\epsilon)/2}\bar{u}^{2}$ is essentially bounded w.p.1 as $n\to\infty$.
To see this, suppose $\operatorname{E}[X_{n}^{2}]<\infty$. Then, for any $\epsilon>0$ and $0<B<\infty$, by Chebyshev inequality, $\sum_{n=1}^{\infty}P(\lvertX_{n}\rvert\geq n^{(1+\epsilon)/2}B)\leq \operatorname{E}[X_{n}^{2}]B^{-2}\sum_{n=1}^{\infty}n^{-(1+\epsilon)}<\infty$. Thus, by the first Borel-Cantelli Lemma, $P(n^{-(1+\epsilon)/2}\lvertX_{n}\rvert\geq B\quad \text{i.o.})=0$, i.e., $n^{-(1+\epsilon)/2}\lvertX_{n}\rvert$ is essentially bounded w.p.1 as $n\to\infty$.
Since $\operatorname{E}[u^{4}]<\infty$ by assumption, for some $\epsilon>0$, however small,
$n^{(1-\epsilon)/2}\bar{u}^{2}
=n^{-(1+\epsilon)/2}(n^{1/2}\bar{u})^{2}$
is essentially bounded w.p.1 as $n\to\infty$. Next,
\begin{align*}
\operatorname{E}[(n^{-1}{\begingroup\textstyle\sum\endgroup}_{i=1}^{n}[k_{b}^{(2)}(u-u_{i}+\omega\bar{u})-k_{b}^{(2)}(u-u_{i})])^{2}\bar{u}^{4}]
&\leq \operatorname{E}[(\max_{1\leq i\leq n}\lvertk_{b}^{(2)}(u-u_{i}+\omega\bar{u})-k_{b}^{(2)}(u-u_{i})\rvert)^{2}\bar{u}^{4}] \\
& \leq C^{2}(n^{\tau/2}b^{3+\tau})^{-2}n^{\tau(1+\epsilon)/2}\operatorname{E}[(n^{(1-\epsilon)/2}\lvert\bar{u}\rvert^{2})^{\tau}\bar{u}^{4}]\\ & = \mathit{o}_(n^{-2+\tau(1+\epsilon)/2}) = \mathit{o}_(n^{-1}b^{3})\quad w.p.1.
\end{align*}
The first inequality follows from $n^{-1}\sum_{i}a_{i}^{2}\leq \max_{1\leq i\leq n}a_{i}^{2}$, the second by H\"{o}lder continuity of $k^{(2)}$ as above and writing
$\lvertn^{1/2}\bar{u}\rvert^{2\tau}
=n^{\tau(1+\epsilon)/2}(\lvertn^{(1-\epsilon)/2}\bar{u}\rvert^{2})^{\tau}$, the third as, by Assumption (ref)(c), $n^{\tau/2}b^{3+\tau}\to\infty$ and,
by the extremal H\"{o}lder inequality with exponents $\infty$ and $1$,
$\operatorname{E}[(n^{(1-\epsilon)/2}\lvert\bar{u}\rvert^{2})^{\tau}\bar{u}^{4}] \leq
\mathit{O}_{}(n^{-2})$ noting that $n^{(1-\epsilon)/2}\lvert\bar{u}\rvert^{2}$ is essentially bounded w.p.1 as $n\to\infty$ and $\operatorname{E}[\bar{u}^{4}]=\mathit{O}_{}(n^{-2})$ and, finally, as
$\mathit{o}_{}(n^{-2+\tau(1+\epsilon)/2}) =
\mathit{o}_{}(n^{-1}b^{3})n^{(\tau-1)/2+9(\tau-1)/[8(3+\tau)]+(4\tau\epsilon-1)/8}$
because $n^{-3\tau/[2(3+\tau)]}b^{-3}\to0$ by Assumption (ref)(c), choosing $\epsilon\leq1/4\tau$ gives the result.
If $f$ is twice differentiable and $f^{(2)}(u)$ and $uf^{(1)}(u)$ are absolutely integrable, applying Lemma (ref),
\begin{align*}
\operatorname{E}[\hat{\delta}_{1}(u)]
& = n^{-1}\operatorname{E}[u_{i}k_{b}^{(1)}(u-u_{i})]
+\tfrac{1}{2}\sigma^{2}n^{-1}\operatorname{E}[k_{b}^{(2)}(u-u_{i})]
+\mathit{o}_(n^{-1})\\
& = n^{-1}\left(f(u)+uf^{(1)}(u) +\tfrac{1}{2}\sigma^{2}f^{(2)}(u)\right) +\mathit{o}_(n^{-1}),
\end{align*}
where $\sigma^{2}=\operatorname{E}[u^{2}]$.
Since $\operatorname{Var}[\tilde{f}(u)] \sim (nb)^{-1}$, the covariance between $\tilde{f}(u)$ and the remainder term in $\hat{\delta}_{1}(u)$ is of order $\mathit{o}_{}(n^{-1}b)$, and, hence,
\begin{align}
\mkern-20mu\operatorname{Cov}[\tilde{f}(u),\hat{\delta}_{1}(u)]
& = n^{-1}\operatorname{E}[k_{b}^{(1)}(u-u_{i})]\operatorname{E}[k_{b}(u-u_{j})u_{j}] +\mathit{o}_(n^{-1}b)
= n^{-1}uf^{(1)}(u)f(u) + \mathit{o}_(n^{-1}b),\\
\operatorname{Var}[\hat{\delta}_{1}(u)] & = n^{-1}\operatorname{E}[k_{b}^{(1)}(u-u_{i})]^{2}\operatorname{E}[u_{j}^{2}] + \mathit{o}_(n^{-1}b^{3})
= n^{-1}\sigma^{2}f^{(1)}(u)^{2}+ \mathit{o}_(n^{-1}b).
\end{align}
Note that $\zeta_{\lambda}=0$, $\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u=-f^{(1)}(u)$, $\mathrm{d}\{\operatorname{E}[\nabla^{\top}u_{i}Hg_{i}|u]f(u)\}/\mathrm{d}u=f(u)+uf^{(1)}(u)$,
$\mathrm{d}\{\operatorname{E}[\nabla^{2}u_{i}|u]f(u)\}/\mathrm{d} u=0$, and
$\mathrm{d}^{2}\{\operatorname{E}[\nabla u_{i}\nabla^{\top}u_{i}|u]f(u)\}/\mathrm{d}u^{2}=f^{(2)}(u)$
from the unbiasedness of $\hat{\beta}$ and linearity of $u(z,\beta)$; cf. Theorem (ref).
Assuming $f^{(1)}(u)$ is square integrable, and if $\lim_{\lvertu\rvert\to\infty}uf(u)^{2}=0$, $\int uf^{(1)}(u)f(u)\mathrm{d}u=-\frac{1}{2}R(f)$ and, thus,
\begin{equation*}
\operatorname{IVar}[\hat{f}] = \operatorname{IVar}[\tilde{f}] - n^{-1}(R(f)-\sigma^{2}R(f^{(1)})) + \mathit{o}_(n^{-1}).
\end{equation*}
Hence, whenever $R(f)>\sigma^{2}R(f^{(1)})$, $\hat{f}$ achieves a second order reduction in variance relative to $\tilde{f}$. While this may appear as a `free' reduction in variance, it is not so. Construction of $\hat{f}$ explicitly assumes that $\operatorname{E}[u]$ exists, and the validity of the above result requires the first four moments of $u$ to exist whereas that of $\tilde{f}$ makes no such assumptions.
When the mean $\operatorname{E}[u]$ is known, the (G)EL-reweighted estimator $\tilde{f}_{\rho}$ eq. (ref) imposing the constraint $\operatorname{E}[u]=0$ will achieve a second order reduction in variance of $n^{-1}\sigma^{-2}u^{2}f(u)^{2}$, i.e., $\operatorname{IVar}[\tilde{f}_{\rho}] = \operatorname{IVar}[\tilde{f}] - n^{-1}\sigma^{-2}\int u^{2}f(u)^{2}\mathrm{d}u+ \mathit{o}_{}(n^{-1})$; see, e.g., chen1997.
In particular, for normally distributed $u$,
$R(\phi_{\sigma})-\sigma^{2}R(\phi_{\sigma}^{(1)})=1/4\sqrt{\pi}\sigma$, which equals
$\sigma^{-2}\int u^{2}\phi_{\sigma}(u)^{2}\mathrm{d}u$ exactly.
For the Student $t$ distribution with $\nu>2$ degrees of freedom,
$R(t_{\nu}) -\sigma^{2}R(t_{\nu}^{(1)})= R(t_{\nu})(2\nu^{2}-3\nu -17)/4(\nu^{2}-4)$,
which is positive for $\nu>4$, the condition for the first four moments of $u$ to exist, whereas
$\sigma^{-2}\int u^{2}t_{\nu}(u)^{2}\mathrm{d}u = R(t_{\nu})(\nu-2)/(2\nu-1)$ which is always larger than $R(t_{\nu})-\sigma^{2}R(t_{\nu}^{(1)})$. This difference may be interpreted as the cost of having to estimate the mean of $u$.
The same or similar terms appear in the expansions for the variance of $\hat{f}$ in other contexts (the $\mathit{O}_{}(n^{-1})$ bias terms tend to be ignored as their contribution to MISE is $\mathit{o}_{}(n^{-1})$);
cf. mushal2010. As the next example demonstrates, these same effects appear in a large class of parametric moment condition models.
\phantomsection
\addcontentsline{toc}{section}{Example (ref): GEL with a constant and zero mean restriction}
example[GEL With A Constant And Zero Mean Restriction]
Consider GEL estimation based on moment indicator functions of the form $g(z,\beta) = u(z,\beta)\alpha(w)$ where $u(z,\beta)$ is scalar, $\beta$ a $d_{\beta}$-vector of parameters, and $\alpha(w)$ a $d_{g}$-vector of functions of $w$. Suppose that $u(z,\beta_{0})$ is independent of $w$,
Assumption (ref) holds,
and the moment condition $\operatorname{E}[g(z,\beta_{0})]=0$ includes the restriction $\operatorname{E}[u(z,\beta_{0})]=0$.
Furthermore, it is assumed that $u(z,\beta)$ contains a constant; the inclusion of an explicit constant is not essential as the results here continue to hold if $\operatorname{E}[\partial u(z,\beta_{0})/\partial\beta^{\top}|w]\gamma=c$ for some non-zero vector $\gamma$ and scalar $c$, in which case $\operatorname{E}[\alpha(w)]=G\gamma/c$.
Without loss of generality let $\alpha_{1}(w)=1$ and $\partial u(z,\beta_{0})/\partial\beta_{1} =-1$.
Since $u$ and $w$ are independent, $\operatorname{E}[g_{i}|u]=u\operatorname{E}[\alpha(w)]$, $\Omega=\sigma^{2}\operatorname{E}[\alpha(w)\alpha(w)^{\top}]$, where $\sigma^{2}=\operatorname{E}[u^{2}|w]=\operatorname{E}[u^{2}]$. Then, because the first column of $G$ is $-\operatorname{E}[\alpha(w)]$, as $PG=0$, $\operatorname{E}[g_{i}|u]^{\top}P\operatorname{E}[g_{i}|u]=0$. That is, there is no second order reduction in variance due to re-weighting.
Since the first column (and row) of $\Omega$ is $\sigma^{2}\operatorname{E}[\alpha(w)]$,
$\Omega^{-1}\operatorname{E}[g_{i}|u] = u\sigma^{-2}e_{1}$, where $e_{j}$ is the $j$th unit $d_{g}$-vector, $j=1,\ldots,d_{g}$.
For an $n\times m$ matrix $A$, let $A_{(s:t)}$, $1\leq s\leq t\leq m$, denote the $n\times (t-s+1)$ submatrix comprised of columns $j=s,\ldots,t$ of $A$. Noting that $\Omega^{-1}\operatorname{E}[\alpha(w)]=e_{1}/\sigma^{2}$ and $\operatorname{E}[\alpha(w)]^{\top}\Omega^{-1}\operatorname{E}[\alpha(w)]=1/\sigma^{2}$, partition $\Sigma$ and $H$ as
\begin{equation*}
\Sigma
=\sigma^{2}\left[\begin{smallmatrix}1+e_{1}^{\top}G_{(2:p)}Q^{-1}G_{(2:p)}^{\top}e_{1}&e_{1}^{\top}G_{(2:p)}Q^{-1}\\Q^{-1}G_{(2:p)}^{\top}e_{1}&Q^{-1}\end{smallmatrix}\right], \qquad
H =-\left[\begin{smallmatrix}
e_{1}^{\top}
-e_{1}^{\top}G_{(2:p)}Q^{-1}G_{(2:p)}^{\top}(\sigma^{2}\Omega^{-1}-e_{1}e_{1}^{\top}) \\
-Q^{-1}G_{(2:p)}^{\top}(\sigma^{2}\Omega^{-1}-e_{1}e_{1}^{\top})\end{smallmatrix}\right],
\end{equation*}
where $Q = G_{(2:p)}^{\top}(\sigma^{2}\Omega^{-1}-e_{1}e_{1}^{\top})G_{(2:p)}$ and
$e_{1}^{\top}G_{(2:p)}= (\operatorname{E}[\partial u(z,\beta_{0})/\partial\beta_{2}],\ldots,\operatorname{E}[\partial u(z,\beta_{0})/\partial\beta_{p}])$ is the first row of $G_{(2:p)}$.
Thus, $H\operatorname{E}[g_{i}|u] =-ue_{1}$.
As the first element of $\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u$ is $-f^{(1)}(u)$, $[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]^{\top}H\operatorname{E}[g_{i}|u]f(u) = uf^{(1)}(u)f(u)$, same as eq. (ref).
Partition $\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u$ as $\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u=\left(-f^{(1)}(u),\right. \\ \left.[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]_{(2:p)}\right)^{\top}$. Hence,
\begin{align}\notag
& \mkern-10mu [\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]^{\top}\Sigma [\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]
=\sigma^{2}f^{(1)}(u)^{2}
+\sigma^{2}f^{(1)}(u)^{2}e_{1}^{\top}G_{(2:p)}Q^{-1}G_{(2:p)}^{\top}e_{1}
\\ \notag & \mkern230mu
-2\sigma^{2}f^{(1)}(u)e_{1}^{\top}G_{(2:p)}Q^{-1}[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]_{(2:p)}
\\ & \mkern230mu
+ \sigma^{2}[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]_{(2:p)}^{\top}Q^{-1}[\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]_{(2:p)}
\end{align}
The first term in (ref) is the same as the main term in (ref). The remaining terms represent the additional increase in the variance of $\hat{f}(u)$ due to the estimation error in $\beta_{2},\ldots,\beta_{p}$.
The independence of $u$ and $w$ is crucial to the above argument implying $\operatorname{E}[g_{i}|u]=u\operatorname{E}[\alpha(w)]$, and $P$ annihilates $\operatorname{E}[\alpha(w)]$. The next example illustrates that these relationships need not hold in the dependent case.
\phantomsection
\addcontentsline{toc}{section}{Example (ref): Linear regression model with
\texorpdfstring{$\operatorname{E}[u|x]=0$}{E[u|x]=0} but dependent \texorpdfstring{$u$}{u} and \texorpdfstring{$x$}{x}}
example[Linear Regression Model With
\texorpdfstring{$\operatorname{E}[u|x]=0$}{E[u|x]=0} But Dependent \texorpdfstring{$u$}{u} And \texorpdfstring{$x$}{x}]
For simplicity, consider the linear regression model
\begin{equation}
y = \delta_{0} + \gamma_{0} x + u,
\end{equation}
where $\operatorname{E}[u|x]=0$. Here $\beta = (\delta,\gamma)^{\top}$ and $z=(y,x)^{\top}$.
Estimation of $\beta_{0}$ may be based on the unconditional moment restriction $\operatorname{E}[g(z,\beta_{0})]=0$ where
\begin{equation}
g(z,\beta) = u(z,\beta)(1,\:x,\:x^{2},\:\ldots,\:x^{q-1})^{\top}, \quad q\geq2
\end{equation}
Suppose now that $u$ and $x$ are distributed with joint density
\begin{equation}
f_{U,X}(u,x) = \frac{2(\nu/2)^{\nu/2}}{\sqrt{2\pi}\omega\Gamma(\nu/2)} x^{\nu}\mathrm{e}^{-x^{2}(\nu+u^{2}/\omega^{2})/2},
\quad x\geq0, \: -\infty<u<\infty, \: \nu>0, \:\omega>0.
\end{equation}
The marginal distributions of $u$ and $x$ are the non-standardized Student $t$ distribution with $\nu$ degrees of freedom and scale parameter $\omega$, and the generalized gamma stacy1962 with parameters $p=2$, $d=\nu$, and $a=(2/\nu)^{1/2}$. \footnote{$x$ is distributed as $(w/\nu)^{1/2}$ where $w\sim\chi_{\nu}^{2}$ and, if $\nu=1$, as a standard half-normal random variable. The joint density eq. (ref) is that of $u=z/x$ and $x$, where $z\sim N(0,\omega^{2})$, independent of $x$.}
The moments of $x$ are $m_{k}=\operatorname{E}[x^{k}] = (2/\nu)^{k/2}\Gamma((\nu+k)/2)/\Gamma(\nu/2)$, $k>-\nu$, and satisfy the recursion $m_{k+2}= (1+k/\nu)m_{k}$.
The odd moments of $u$ of order $k<\nu$ are zero, while the even moments are $\operatorname{E}[u^{2k}] = \omega^{2k}\pi^{-1/2}\nu^{k}\Gamma(\nu/2-k)\Gamma(k+1/2)/\Gamma(\nu/2)$, $k<\nu/2$.
The conditional density of $u$ given $x$ is $f_{U|X}(u,x)=\phi_{\omega/x}(u)$ and, hence, $\operatorname{E}[u|x] = 0$, but $u$ and $x$ are not independent. If $\nu>2$, $\operatorname{E}[u^{2}|x]=\omega^{2}/x^{2}$.
The conditional moments of $x$ given $u$ are
$m_{k|u}(u) =\operatorname{E}[x^{k}|u] = m_{k+1}/m_{1}(1+(u/\omega)^{2}/\nu)^{k/2}$, $k>-\nu-1$. The transformation in Assumption (ref) has $v(z,\beta)=x$ and, hence, $\operatorname{E}[g_{i}|u] = u(1,\: m_{1|u}(u),\: m_{2|u}(u),\: \ldots,\: m_{q-1|u}(u))^{\top}$.
To describe the quantities involved, let $\prescript{q}{s}{\mathrm{M}}^{}_{}=\{m_{i+j-2-s}\}_{i,j=1}^{q}$ be a $q\times q$ matrix composed of the $(i+j-2-s)$th moments of $x$. Note that, if $q>2$, then $\prescript{q}{0}{\mathrm{M}}^{\top}_{(s)} \prescript{q}{2}{\mathrm{M}}^{-1}_{}\prescript{q}{0}{\mathrm{M}}^{}_{(t)} =m_{s+t}$ for $s,t=1,2,\ldots$, $(s\wedge t)\leq q-2$, and $\prescript{q}{2}{\mathrm{M}}^{-1}_{}\prescript{q}{0}{\mathrm{M}}^{}_{(t)} = e_{t+2}$ for $1\leq t\leq q-2$.
The relevant (G)EL matrices are $\Omega = \omega^{2}\prescript{q}{2}{\mathrm{M}}^{}_{}$, $G = -\prescript{q}{0}{\mathrm{M}}^{}_{(1:2)}$ and, if $q\geq4$,
\begin{equation*}
\Sigma = \tfrac{\omega^{2}\nu}{\nu(\nu+2)-(\nu+1)^{2}m_{1}^{2}}
\left[\begin{smallmatrix} \nu+2& -(\nu+1)m_{1} \\ -(\nu+1)m_{1} & \nu \end{smallmatrix}\right], \qquad
H = -\tfrac{1}{\omega^{2}}\Sigma \left[\begin{smallmatrix}e_{q,3}^{\top} \\ e_{q,4}^{\top} \end{smallmatrix} \right],
\end{equation*}
and
\begin{equation*}
P = \tfrac{1}{\omega^{2}}\prescript{q}{2}{\mathrm{M}}^{-1}_-\tfrac{1}{\omega^{2}}\tfrac{\nu}{\nu(\nu+2)-(\nu+1)^{2}m_{1}^{2}}
\left[ (\nu+2)e_{3}e_{3}^{\top}
-(\nu+1)m_{1}\left( e_{3}e_{4}^{\top} +e_{4}e_{3}^{\top} \right)
+ \nu e_{4}e_{4}^{\top} \right].
\end{equation*}
\begin{remark}
For the exactly identified case, $q=2$, $G$ is square and invertible. Hence, $\Sigma = G^{-1}\Omega G^{\top-1}$, $H = G^{-1}$, and $P=0$. Closed form expressions for $\Sigma$, $H$, and $P$ when $q=3$ can be obtained in a straightforward fashion. That $\Sigma$ remains unaltered as $q$ increases above $4$ is of course due to the special form of the conditional variance of $u$.
Figure (ref) displays the relative efficiency of $\hat{\beta}$ based on the first $q$ compared with the first $q^{\prime}$ moment conditions, $[\det(\Sigma_{q})/\det(\Sigma_{q^{\prime}})]^{1/p}$, for various values of $\nu$.
\end{remark}
\begin{figure}[htbp]
\caption{Relative efficiency of $\hat{\beta}$ based on
$q\geq 4$ vs. $q=2$ (solid line),
$q=3$ vs. $q=2$ (dashed line),
and $q\geq 4$ vs. $q=3$ (dash-dotted line) moment conditions as a function of $\nu$.}
\end{figure}
If $q\geq4$, only the moment indicators $x^{j-1}u(z,\beta)$, $j=3,4$, are used to estimate $\beta_{0}$. Information in the remaining moment conditions, however, can be usefully exploited to improve the efficiency of the density estimators $\hat{f}$ and $\hat{f}_{\rho}$. The quantities entering the integrated variance eqs. (ref) and (ref) can be computed as
$\mathop{\mathrm{tr}}\left(\Sigma {\begingroup\textstyle\int\endgroup} [\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u][\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]^{\top}\mathrm{d}u\right)$,
$\mathop{\mathrm{tr}}\left(H \int\operatorname{E}[g_{i}|u]\right.\\\left.\times [\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]^{\top}f(u)\mathrm{d}u \right)$, and
$\mathop{\mathrm{tr}}\left(P \int\operatorname{E}[g_{i}|u]\operatorname{E}[g_{i}|u]^{\top}f(u)^{2}\mathrm{d}u \right)$, where
\begin{equation*}
\int \tfrac{\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}}{\mathrm{d}u}\left(\tfrac{\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}}{\mathrm{d}u}\right)^{\top} \mathrm{d}u
= \tfrac{\Gamma((\nu+3)/2)}{\omega^{3}\pi^{1/2}\nu^{3/2}\Gamma(\nu/2)}\left[\begin{smallmatrix}
\tfrac{\nu\Gamma(\nu+3/2)\Gamma((\nu+3)/2)}{\Gamma(\nu/2+1)\Gamma(\nu+3)} &
\tfrac{\nu^{1/2}\Gamma(\nu+3)}{2^{1/2}\Gamma(\nu+7/2)}\\
\tfrac{\nu^{1/2}\Gamma(\nu+3)}{2^{1/2}\Gamma(\nu+7/2)}&
\tfrac{\Gamma(\nu+5/2)\Gamma(\nu/2+2)}{2\Gamma(\nu+2)\Gamma((\nu+5)/2)} \end{smallmatrix}\right],
\end{equation*}
the $q\times 2$ matrix ${\begingroup\textstyle\int\endgroup}\operatorname{E}[g_{i}|u][\mathrm{d}\{\operatorname{E}[\nabla u_{i}|u]f(u)\}/\mathrm{d}u]^{\top}f(u)\mathrm{d}u$ has rows
\begin{equation*}
\tfrac{1}{(2\pi)^{1/2}\omega}\left[
\tfrac{(2/\nu)^{i/2}\Gamma((\nu+3)/2)\Gamma(\nu+i/2)\Gamma((\nu+i)/2))}{\Gamma(\nu/2)^{2}\Gamma(\nu+(i+3)/2)}
, \quad
\tfrac{(2/\nu)^{(i-1)/2}(\nu+2)\Gamma((\nu+1)/2)\Gamma(\nu+(i+1)/2)}{2\Gamma(\nu/2)\Gamma(\nu+i/2+2)}
\right], \quad i=1,\ldots,q,
\end{equation*}
and the $q\times q$ matrix $\int\operatorname{E}[g_{i}|u]\operatorname{E}[g_{i}|u]^{\top}f(u)^{2}\mathrm{d}u$ with $(i,j)$th element
\begin{equation*}
\omega m_{i}m_{j}\tfrac{\nu^{3/2} \Gamma(\nu+(i+j-3)/2)}{4\pi^{1/2}\Gamma(\nu+(i+j)/2)}, \quad i,j=1,\ldots,q.
\end{equation*}
\begin{remark}
Figure (ref) shows the values of the above quantities and the overall effect on the integrated variance for selected values of $q$ and $\nu>2$; note that the validity of asymptotic expansions requires $\nu>4$, but variance is defined for $\nu>2$.
While the main reduction in variance is still due to the zero mean restriction as in Example (ref) (Panels A and B), there are small additional gains due to re-weighting (Panel C). The latter do increase as more moment conditions are added.
\end{remark}
figure[figure omitted — 1,288 chars of source]