EconBase
← Back to paper

Smoothed estimating equations for instrumental variables quantile regression

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

136,506 characters

Smoothed Estimating Equations for Instrumental Variables Quantile Regression


\title[Smoothed Estimating Equations for {IV-QR}]{Smoothed Estimating Equations \\ for Instrumental Variables Quantile Regression}




\thanks{
Thanks to Victor Chernozhukov (co-editor) and an anonymous referee for insightful comments and references, and thanks to Peter C.\ B.\ Phillips (editor) for additional editorial help.  Thanks to Xiaohong Chen, Brendan
Beare, Andres Santos, and active seminar and conference participants for insightful
questions and comments.}

\maketitle

\begin{center}
{\sc David M.\ Kaplan}

\textit{Department of Economics, University of Missouri}

\textit{118 Professional Bldg, 909 University Ave, Columbia, MO 65211-6040}

E-mail: \texttt{[email removed]}

~\\

{\sc Yixiao Sun}

\textit{Department of Economics, University of California, San Diego}

E-mail: \texttt{[email removed]}
\end{center}






\paragraph{\sc Abstract}

The moment conditions or estimating equations for instrumental variables
quantile regression involve the discontinuous indicator function. We instead
use smoothed estimating equations (SEE), with bandwidth $h$. We show that
the mean squared error (MSE) of the vector of the SEE is minimized for some $h>0$,
leading
to smaller asymptotic
MSE
of the estimating
equations and associated parameter estimators. The same MSE-optimal $h$ also
minimizes the higher-order type I error of a SEE-based $\chi ^{2}$ test
and increases
size-adjusted power in large samples. Computation of the SEE
estimator also becomes simpler and more reliable, especially with (more)
endogenous regressors. Monte Carlo simulations demonstrate all of these
superior properties in finite samples, and we apply our estimator to JTPA data.
Smoothing the estimating equations is
not just a technical operation for establishing Edgeworth expansions and
bootstrap refinements; it also brings the real benefits of having more
precise estimators and more powerful tests.






\allowdisplaybreaks[4]


\section{Introduction}

Many econometric models are specified by moment conditions or estimating
equations. An advantage of this approach is that the full distribution of
the data does not have to be parameterized. In this paper, we consider
estimating equations that are not smooth in the parameter of interest. We
focus on instrumental variables quantile regression (IV-QR), which
includes the usual quantile regression as a special case. Instead of using
the estimating equations that involve the nonsmooth indicator function, we
propose to smooth the indicator function, leading to our smoothed estimating
equations (SEE) and SEE estimator.

Our SEE estimator has several advantages. First, from a computational point
of view, the SEE estimator can be computed using any standard iterative
algorithm that requires smoothness. This is especially attractive in IV-QR
where simplex methods for the usual QR are not applicable. In fact,
the SEE approach has been used in
\citet{ChenPouzo2009,ChenPouzo2012} for computing their nonparametric sieve
estimators in the presence of nonsmooth moments or generalized residuals.
However, a rigorous investigation is currently lacking. Our paper can be
regarded as a first step towards justifying the SEE approach in
nonparametric settings.
Relatedly, \citet[][\S7.1]{FanLiao2014} have employed
the same strategy of smoothing the indicator function to
reduce the computational burden of their focused GMM approach.
Second, from a technical point of view, smoothing
the estimating equations enables us to establish high-order properties of
the estimator. This motivates \citet{Horowitz1998}, for instance, to examine
a smoothed objective function for median regression, to show high-order
bootstrap refinement. Instead of smoothing the objective function, we show
that there is an advantage to smoothing the estimating equations. This point
has not been recognized and emphasized in the literature. For QR estimation
and inference via empirical likelihood, \citet{Otsu2008} and
\citet{Whang2006} also examine smoothed estimators. To the best of our
knowledge, no paper has examined smoothing the estimating equations for the
usual QR estimator, let alone IV-QR. Third, from a statistical point of
view, the SEE estimator is a flexible class of estimators that includes the
IV/OLS mean regression estimators and median and quantile regression
estimators as special cases. Depending on the smoothing parameter, the SEE
estimator can have different degrees of robustness in the sense of
\citet{Huber1964}. By selecting the smoothing parameter appropriately, we
can harness the advantages of both the mean regression estimator and the
median/quantile regression estimator. Fourth and most importantly, from an
econometric point of view, smoothing can reduce the mean squared error (MSE)
of the SEE, which in turn leads to a smaller asymptotic MSE of the parameter
estimator and to more powerful tests. We seem to be the first to establish
these advantages.

In addition to investigating the asymptotic properties of the SEE estimator,
we provide a smoothing parameter choice that minimizes different criteria:
the MSE of the SEE, the type I error of a chi-square test subject to exact
asymptotic size control, and the approximate MSE of the parameter estimator. We show
that the first two criteria produce the same optimal smoothing parameter,
which is also optimal under a variant of the third criterion. With the
data-driven smoothing parameter choice, we show that the statistical and
econometric advantages of the SEE estimator are reflected clearly in our
simulation results.

There is a growing literature on IV-QR.  For a recent review, see \citet{ChernozhukovHansen2013}.  Our paper is built upon \citet{ChernozhukovHansen2005}, which establishes a structural framework for IV-QR and provides primitive identification conditions.  Within this framework, \citet{ChernozhukovHansen2006} and \citet{ChernozhukovEtAl2009} develop estimation and inference procedures under strong identification.  For inference procedures that are robust to weak identification, see \citet{ChernozhukovHansen2008} and \citet{Jun2008}, for example. IV-QR can also reduce bias for dynamic panel fixed effects estimation as in \citet{Galvao2011}.
None of these papers considers smoothing the IV-QR estimating equations; that idea (along with minimal first-order theory) seemingly first appeared in an unpublished draft by \citet{MaCurdyHong1999}, although the idea of smoothing the indicator function in general appears even earlier, as in \citet{Horowitz1992} for the smoothed maximum score estimator.
An alternative approach to overcome the computational obstacles in the presence of a nonsmooth objective function is to explore the asymptotic equivalence of the Bayesian and classical methods for regular models and use the MCMC approach to obtain the classical extremum estimator; see \citet{ChernozhukovHong2003}, whose Example 3 is IV-QR.  As a complement, our approach deals with the computation problem in the classical framework directly.

The rest of the paper is organized as follows. Section \ref{sec:see}
describes our setup and discusses some illuminating connections with other
estimators. Sections \ref{sec:mse}, \ref{sec:eI}, and \ref{sec:est}
calculate the MSE of the SEE, the type I and type II errors of a chi-square
test, and the approximate MSE of the parameter estimator, respectively.
Section \ref{sec:emp} applies our estimator to JTPA data, and Section \ref{sec:sim} presents simulation results before we conclude. Longer
proofs and calculations are gathered in the appendix.

\section{Smoothed Estimating Equations\label{sec:see}}

\subsection{Setup}

We are interested in estimating the instrumental variables quantile
regression (IV-QR) model
\begin{equation*}
Y_{j}=X_{j}^{\prime }\beta _{0}+U_{j}
\end{equation*}
where $\mathbb E\mathopen{}\mathclose \bgroup \originalleft[Z_{j}\mathopen{}\mathclose \bgroup \originalleft( 1\{U_{j}<0\}-q\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] =0$ for instrument vector $
Z_{j}\in \mathbb{R}^{d}$ and $1\{ \cdot \}$ is the indicator function.
Instruments are taken as given; this does not preclude first determining the
efficient set of instruments as in \citet{Newey2004} or
\citet{NeweyPowell1990}, for example.
We restrict attention to the ``just identified'' case $X_{j}\in \mathbb{R}
^{d}$ and iid data for simpler exposition; for the overidentified case, see
\eqref{eqn:EE-overID} below.

A special case of this model is exogenous QR with $Z_{j}=X_{j}$, which is
typically estimated by minimizing a criterion function:
\begin{equation*}
\hat{\beta}_{Q}\equiv \mathop{\rm arg\,min}_{\beta }\frac{1}{n}
\sum_{j=1}^{n}\rho _{q}(Y_{j}-X_{j}^{\prime }\beta ),
\end{equation*}
where $\rho _{q}(u)\equiv \mathopen{}\mathclose \bgroup \originalleft( q-1\{u<0\}\aftergroup \egroup \originalright) u$ is the check function.
Since the objective function is not smooth, it is not easy to obtain a
high-order approximation to the sampling distribution of $\hat{\beta}_{Q}$.
To avoid this technical difficulty, \citet{Horowitz1998} proposes to smooth
the objective function to obtain
\begin{equation*}
\hat{\beta}_{H}=\mathop{\rm arg\,min}_{\beta }\frac{1}{n}\sum_{j=1}^{n}\rho
_{q}^{H}(Y_{j}-X_{j}^{\prime }\beta ),\quad \rho _{q}^{H}(u)\equiv \mathopen{}\mathclose \bgroup \originalleft[
q-G\mathopen{}\mathclose \bgroup \originalleft( -u/h\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] u,
\end{equation*}
where $G(\cdot )$ is a smooth function and $h$ is the smoothing parameter or
bandwidth. Instead of smoothing the objective function, we smooth the
underlying moment condition and define $\hat{\beta}$ to be the solution of
the vector of smoothed estimating equations (SEE) $m_{n}(\hat{\beta})=0$,
where\footnote{
It suffices to have $m_{n}(\hat{\beta})=o_{p}(1)$, which allows for a small
error when $\hat{\beta}$ is not the exact solution to $m_{n}(\hat{\beta})=0$.
}
\begin{equation*}
m_{n}(\beta )\equiv \frac{1}{\sqrt{n}}\sum_{j=1}^{n}W_{j}(\beta )\text{ and }
W_{j}(\beta )\equiv Z_{j}\mathopen{}\mathclose \bgroup \originalleft[ G\mathopen{}\mathclose \bgroup \originalleft( \frac{X_{j}^{\prime }\beta -Y_{j}}{h}
\aftergroup \egroup \originalright) -q\aftergroup \egroup \originalright] .
\end{equation*}

Our approach is related to kernel-based nonparametric conditional quantile
estimators. The moment condition there is $\mathbb E\mathopen{}\mathclose \bgroup \originalleft[ 1\{X=x\} \mathopen{}\mathclose \bgroup \originalleft(
1\{Y<\beta \}-q\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] =0$. Usually the $1\{X=x\}$ indicator
function is ``smoothed'' with a kernel, while the latter term is not. This
yields the nonparametric conditional quantile estimator $\hat{\beta}_{q}(x)=
\mathop{\rm arg\,min}_{b}\sum_{i=1}^{n}\rho _{q}(Y_{i}-b)K[(x-X_{i})/h]$ for
the conditional $q$-quantile at $X=x$, estimated with kernel $K(\cdot)$ and
bandwidth $h$. Our approach is different in that we smooth the indicator $
1\{Y<\beta \}$ rather than $1\{X=x\}$. Smoothing both terms may help but is
beyond the scope of this paper.

Estimating $\hat{\beta}$ from the SEE is computationally easy: $d$ equations
for $d$ parameters, and a known, analytic Jacobian. Computationally, solving
our problem is faster and more reliable than the IV-QR method in
\citet{ChernozhukovHansen2006}, which requires specification of a grid of
endogenous coefficient values to search over, computing a conventional QR
estimator for each grid point. This advantage is important particularly when
there are more endogenous variables.

If the model is overidentified with $\dim (Z_{j})>\dim (X_{j})$, we can use
a $\dim (X_{j})\times \dim (Z_{j})$ matrix $\mathbb{W}$ to transform the
original moment conditions $\mathbb E\mathopen{}\mathclose \bgroup \originalleft[Z_{j}\mathopen{}\mathclose \bgroup \originalleft(q-1\mathopen{}\mathclose \bgroup \originalleft\{Y_{j}<X_{j}^{\prime }\beta \aftergroup \egroup \originalright\}\aftergroup \egroup \originalright)\aftergroup \egroup \originalright]=0$
into
\begin{equation}  \label{eqn:EE-overID}
\mathbb E\mathopen{}\mathclose \bgroup \originalleft[ \tilde{Z}_{j}\mathopen{}\mathclose \bgroup \originalleft( q-1\mathopen{}\mathclose \bgroup \originalleft\{Y_{j}<X_{j}^{\prime }\beta \aftergroup \egroup \originalright\} \aftergroup \egroup \originalright)
\aftergroup \egroup \originalright] =0,\text{ for }\tilde{Z}_{j}=\mathbb{W}Z_{j}\in \mathbb{R}^{\dim
(X_{j})}.
\end{equation}
Then we have an exactly identified model with transformed instrument vector $
\tilde{Z}_{j}$, and our asymptotic analysis can be applied to
\eqref{eqn:EE-overID}.

By the theory of optimal estimating equations or efficient two-step GMM, the
optimal $\mathbb{W}$ takes the following form:
\begin{align*}
\mathbb{W} &= \mathopen{}\mathclose \bgroup \originalleft. \frac{\partial }{\partial \beta }\mathbb E\mathopen{}\mathclose \bgroup \originalleft[ Z^{\prime
}\mathopen{}\mathclose \bgroup \originalleft( q-1\mathopen{}\mathclose \bgroup \originalleft\{Y<X^{\prime }\beta \aftergroup \egroup \originalright\} \aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright \vert _{\beta
=\beta _{0}}\mathrm{Var}\mathopen{}\mathclose \bgroup \originalleft[Z\mathopen{}\mathclose \bgroup \originalleft(q-1\{Y<X^{\prime }\beta _{0}\} \aftergroup \egroup \originalright) \aftergroup \egroup \originalright]^{-1}
\\
&= \mathbb E \mathopen{}\mathclose \bgroup \originalleft[XZ^{\prime }f_{U|Z,X}(0)\aftergroup \egroup \originalright] \mathopen{}\mathclose \bgroup \originalleft\{ \mathbb E\mathopen{}\mathclose \bgroup \originalleft[ZZ^{\prime }\sigma
^{2}\mathopen{}\mathclose \bgroup \originalleft( Z\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]\aftergroup \egroup \originalright\} ^{-1} ,
\end{align*}
where $f_{U|Z,X}(0)$ is the conditional PDF of $U$ evaluated at $U=0$ given $
\mathopen{}\mathclose \bgroup \originalleft( Z,X\aftergroup \egroup \originalright) $ and $\sigma ^{2}\mathopen{}\mathclose \bgroup \originalleft( Z\aftergroup \egroup \originalright) =\mathrm{Var}
\mathopen{}\mathclose \bgroup \originalleft(1\mathopen{}\mathclose \bgroup \originalleft \{ U<0\aftergroup \egroup \originalright \} \mid Z\aftergroup \egroup \originalright)$. The standard two-step approach
requires an initial estimator of $\beta _{0}$ and nonparametric estimators
of $f_{U|Z,X}(0)$ and $\sigma ^{2}\mathopen{}\mathclose \bgroup \originalleft( Z\aftergroup \egroup \originalright)$. The underlying
nonparametric estimation error may outweigh the benefit of having an optimal
weighting matrix. This is especially a concern when the dimensions of $X$
and $Z$ are large.  The problem is similar to what \citet{HwangSun2015} consider in a time series GMM framework where the optimal weighting matrix
is estimated using a nonparametric HAC approach.  Under the alternative and
more accurate asymptotics that captures the estimation error of the
weighting matrix, they show that the conventionally optimal two-step
approach does not necessarily outperform a first-step approach that does not
employ a nonparametric weighting matrix estimator. While we expect a similar
qualitative message here, we leave a rigorous analysis to future research.

In practice, a simple procedure is to ignore $
f_{U|Z,X}(0) $ and $\sigma ^{2}\mathopen{}\mathclose \bgroup \originalleft( Z\aftergroup \egroup \originalright) $ (or assume that they are
constants) and employ the following empirical weighting matrix,
\begin{equation*}
\mathbb{W}_{n} = \mathopen{}\mathclose \bgroup \originalleft( \frac{1}{n}\sum_{j=1}^{n}X_{j}Z_{j}^{\prime }\aftergroup \egroup \originalright)
\mathopen{}\mathclose \bgroup \originalleft( \frac{1}{n}\sum_{j=1}^{n}Z_{j}Z_{j}^{\prime }\aftergroup \egroup \originalright) ^{-1}.
\end{equation*}
This choice of $\mathbb{W}_{n}$ is in the spirit of the influential work of
\citet{LiangZeger1986} who advocate the use of a working correlation matrix
in constructing the weighting matrix. Given the above choice of $\mathbb{W}
_{n}$, $\tilde{Z}_{j}$ is the least squares projection of $X_{j}$ on $Z_{j}$
. It is easy to show that with some notational changes our asymptotic
results remain valid in this case.

An example of an overidentified model is the conditional moment
model
\begin{equation*}
\mathbb E\mathopen{}\mathclose \bgroup \originalleft[ \mathopen{}\mathclose \bgroup \originalleft( 1\{U_{j}<0\}-q\aftergroup \egroup \originalright) \mid Z_{j}\aftergroup \egroup \originalright] =0.
\end{equation*}
In this case, any measurable function of $Z_{j}$ can be
used as an instrument. As a result, the model could be overidentified.
According to \citet{Chamberlain1987} and \citet{Newey1990}, the optimal set of instruments in our setting is given by
\begin{equation*}
\mathopen{}\mathclose \bgroup \originalleft. \mathopen{}\mathclose \bgroup \originalleft[ \frac{\partial }{\partial \beta }\mathbb E\mathopen{}\mathclose \bgroup \originalleft( 1\{Y_{j}-X_{j}^{
\prime }\beta <0\}\mid Z_{j}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright \vert _{\beta =\beta _{0}}.
\end{equation*}
Let $F_{U|Z,X}\mathopen{}\mathclose \bgroup \originalleft( u\mid z,x\aftergroup \egroup \originalright) $ and $
f_{U|Z,X}\mathopen{}\mathclose \bgroup \originalleft( u\mid z,x\aftergroup \egroup \originalright) $ be the conditional
distribution function and density function of $U$ given $\mathopen{}\mathclose \bgroup \originalleft(
Z,X\aftergroup \egroup \originalright) =(z,x)$.  Then under some regularity conditions,
\begin{align*}
\mathopen{}\mathclose \bgroup \originalleft. \mathopen{}\mathclose \bgroup \originalleft[ \frac{\partial }{\partial \beta }\mathbb E\mathopen{}\mathclose \bgroup \originalleft( 1\{Y_{j}-X_{j}^{
\prime }\beta <0\}\mid Z_{j}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright \vert _{\beta =\beta _{0}}
  &= \mathopen{}\mathclose \bgroup \originalleft. \mathopen{}\mathclose \bgroup \originalleft\{ \frac{\partial }{\partial \beta }\mathbb E\mathopen{}\mathclose \bgroup \originalleft[ \mathbb E\mathopen{}\mathclose \bgroup \originalleft(
1\{Y_{j}-X_{j}^{\prime }\beta <0\} \mid Z_{j},X_{j}\aftergroup \egroup \originalright) \mathrel{\big|} Z_{j}\aftergroup \egroup \originalright]
\aftergroup \egroup \originalright\} \aftergroup \egroup \originalright \vert _{\beta =\beta _{0}} \\
  &= \mathbb E\mathopen{}\mathclose \bgroup \originalleft\{ \mathopen{}\mathclose \bgroup \originalleft. \mathopen{}\mathclose \bgroup \originalleft[\frac{\partial }{\partial \beta }
F_{U_{j}|Z_{j},X_{j}}\mathopen{}\mathclose \bgroup \originalleft( X_{j}\mathopen{}\mathclose \bgroup \originalleft( \beta -\beta _{0}\aftergroup \egroup \originalright) \mid
Z_{j},X_{j}\aftergroup \egroup \originalright)\aftergroup \egroup \originalright]\aftergroup \egroup \originalright\vert _{\beta =\beta_{0}} \mathrel{\bigg|} Z_{j}\aftergroup \egroup \originalright\}  \\
  &= \mathbb E\mathopen{}\mathclose \bgroup \originalleft[  f_{U_{j}|Z_{j},X_{j}}(0\mid Z_{j},X_{j})X_{j} \mathrel{\big|} Z_{j}\aftergroup \egroup \originalright] .
\end{align*}

The optimal instruments involve the conditional density
$f_{U|Z,X}\mathopen{}\mathclose \bgroup \originalleft( u\mid z,x\aftergroup \egroup \originalright) $ and a conditional expectation. In
principle, these objects can be estimated nonparametrically. However, the
nonparametric estimation uncertainty can be very high,
adversely affecting the reliability of inference. A simple and practical
strategy\footnote{We are not alone in recommending this simple strategy for empirical work.
\citet{ChernozhukovHansen2006} make the same recommendation in their Remark 5 and use this strategy in
their empirical application. See also \citet{Kwak2010}.}
is to construct the optimal instruments as the OLS projection of each $
X_{j}$ onto some sieve basis functions $\Phi ^{K}\mathopen{}\mathclose \bgroup \originalleft(
Z_{j}\aftergroup \egroup \originalright) \equiv \mathopen{}\mathclose \bgroup \originalleft[ \Phi
_{1}(Z_{j}),\ldots,\Phi _{K}(Z_{j})\aftergroup \egroup \originalright] ^{\prime }$, leading to
\begin{equation*}
\tilde{Z}_{j}=\mathopen{}\mathclose \bgroup \originalleft[ \frac{1}{n}\sum_{j=1}^{n}X_{j}\Phi ^{K}\mathopen{}\mathclose \bgroup \originalleft(
Z_{j}\aftergroup \egroup \originalright) ^{\prime }\aftergroup \egroup \originalright] \mathopen{}\mathclose \bgroup \originalleft[ \frac{1}{n}\sum_{j=1}^{n}\Phi
^{K}\mathopen{}\mathclose \bgroup \originalleft( Z_{j}\aftergroup \egroup \originalright) \Phi ^{K}\mathopen{}\mathclose \bgroup \originalleft( Z_{j}\aftergroup \egroup \originalright) ^{\prime }\aftergroup \egroup \originalright]
^{-1}\Phi ^{K}\mathopen{}\mathclose \bgroup \originalleft( Z_{j}\aftergroup \egroup \originalright) \in \mathbb{R}^{\dim (X_{j})}
\end{equation*}
as the instruments. Here $\mathopen{}\mathclose \bgroup \originalleft \{ \Phi _{i}\mathopen{}\mathclose \bgroup \originalleft( \cdot \aftergroup \egroup \originalright)
\aftergroup \egroup \originalright \} $ are the basis functions such as power functions. Since
the dimension of $\tilde{Z}_{j}$ is the same as the dimension of $
X_{j}$, our asymptotic analysis can be applied for any fixed value
of $K$.\footnote{A theoretically efficient estimator can be obtained using the sieve minimum
distance approach. It entails first estimating the conditional expectation $\mathbb E
\mathopen{}\mathclose \bgroup \originalleft[ \mathopen{}\mathclose \bgroup \originalleft( 1\{Y_{j}<X_{j}\beta \}-q\aftergroup \egroup \originalright) \mid Z_{j}\aftergroup \egroup \originalright] $ using $\Phi
^{K}\mathopen{}\mathclose \bgroup \originalleft( Z_{j}\aftergroup \egroup \originalright) $ as the basis functions and then choosing $\beta $
to minimize a weighted sum of squared conditional expectations. See, for
example, \citet{ChenPouzo2009,ChenPouzo2012}. To achieve the semiparametric
efficiency bound, $K$ has to grow with the sample size at an appropriate
rate. In work in progress, we consider nonparametric quantile regression
with endogeneity and allow $K$ to diverge, which is necessary for both
identification and efficiency. Here we are content with a fixed $K$ for
empirical convenience at the cost of possible efficiency loss.}

\subsection{Comparison with other estimators\label{sec:SEE-comp}}


\subsubsection*{Smoothed criterion function}

For the special case $Z_{j}=X_{j}$, we compare the SEE with the estimating equations derived
from smoothing the criterion function as in \citet{Horowitz1998}. The first
order condition of the smoothed criterion function, evaluated at the true $
\beta _{0}$, is
\begin{align}
\notag
0& =\mathopen{}\mathclose \bgroup \originalleft. \frac{\partial }{\partial \beta }\aftergroup \egroup \originalright \vert _{\beta =\beta
_{0}}n^{-1}\sum_{i=1}^{n}\mathopen{}\mathclose \bgroup \originalleft[ q-G\mathopen{}\mathclose \bgroup \originalleft( \frac{X_{i}^{\prime }\beta -Y_{i}}{
h}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] (Y_{i}-X_{i}^{\prime }\beta ) \\
\notag
& =n^{-1}\sum_{i=1}^{n}\Big[-qX_{i}-G^{\prime
}(-U_{i}/h)(X_{i}/h)Y_{i}+G^{\prime }(-U_{i}/h)(X_{i}/h)X_{i}^{\prime }\beta
_{0}+G(-U_{i}/h)X_{i}\Big] \\
\notag
& =n^{-1}\sum_{i=1}^{n}X_{i}\mathopen{}\mathclose \bgroup \originalleft[ G(-U_{i}/h)-q\aftergroup \egroup \originalright] +n^{-1}
\sum_{i=1}^{n}G^{\prime }(-U_{i}/h)\mathopen{}\mathclose \bgroup \originalleft[ (X_{i}/h)X_{i}^{\prime }\beta
_{0}-(X_{i}/h)Y_{i}\aftergroup \egroup \originalright] \\
\label{eqn:SCF-EE}
& =n^{-1}\sum_{i=1}^{n}X_{i}\mathopen{}\mathclose \bgroup \originalleft[ G(-U_{i}/h)-q\aftergroup \egroup \originalright] +n^{-1}
\sum_{i=1}^{n}(1/h)G^{\prime }(-U_{i}/h)[-X_{i}U_{i}].
\end{align}
The first term agrees with our proposed SEE. Technically, it should be
easier to establish high-order results for our SEE estimator since it has
one fewer term. Later we show that the absolute bias of our SEE estimator is
smaller, too. Another subtle point is that our SEE requires only the
estimating equation $\mathbb E\mathopen{}\mathclose \bgroup \originalleft[X_{j}\mathopen{}\mathclose \bgroup \originalleft( 1\{U_{j}<0\}-q\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] =0$, whereas
\citet{Horowitz1998} has to impose an additional condition to ensure that
the second term in the FOC is approximately mean zero.


\subsubsection*{IV mean regression\label{sec:see-comp-IV}}

When $h\rightarrow \infty $, $G(\cdot )$ only takes arguments near zero and
thus can be approximated well linearly. For example, with the $G(\cdot )$
from \citet{Whang2006} and \citet{Horowitz1998}, $G(v)=0.5+(105/64)v+O(v^{3})
$ as $v\rightarrow 0$. Ignoring the $O(v^{3})$, the corresponding estimator $
\hat{\beta}_{\infty }$ is defined by
\begin{align*}
0& =\sum_{i=1}^{n}Z_{i}\mathopen{}\mathclose \bgroup \originalleft[ G\mathopen{}\mathclose \bgroup \originalleft( \frac{X_{i}^{\prime }\hat{\beta}
_{\infty }-Y_{i}}{h}\aftergroup \egroup \originalright) -q\aftergroup \egroup \originalright]  \\
& \doteq \sum_{i=1}^{n}Z_{i}\mathopen{}\mathclose \bgroup \originalleft[ \mathopen{}\mathclose \bgroup \originalleft( 0.5+(105/64)\frac{X_{i}^{\prime }
\hat{\beta}_{\infty }-Y_{i}}{h}\aftergroup \egroup \originalright) -q\aftergroup \egroup \originalright]  \\
& =(105/64h)Z^{\prime }X\hat{\beta}_{\infty }-(105/64h)Z^{\prime
}Y+(0.5-q)Z^{\prime }\mathbf{1}_{n,1} \\
& =(105/64h)Z^{\prime }X\hat{\beta}_{\infty }-(105/64h)Z^{\prime
}Y+(0.5-q)Z^{\prime }(Xe_{1}) ,
\end{align*}
where $e_{1}=(1,0,\ldots ,0)^{\prime }$ is $d\times 1$, $\mathbf{1}
_{n,1}=(1,1,\ldots ,1)^{\prime }$ is $n\times 1$, $X$ and $Z$ are $n\times d$
with respective rows $X_{i}^{\prime }$ and $Z_{i}^{\prime }$, and using the
fact that the first column of $X$ is $\mathbf{1}_{n,1}$ so that $Xe_{1}=
\mathbf{1}_{n,1}$. It then follows that
\begin{equation*}
\hat{\beta}_{\infty }=\hat{\beta}_{IV}+\mathopen{}\mathclose \bgroup \originalleft( (64h/105)(q-0.5),0,\ldots
,0\aftergroup \egroup \originalright) ^{\prime }.
\end{equation*}
As $h$ grows large, the smoothed QR estimator approaches the IV estimator
plus an adjustment to the intercept term that depends on $q$, the bandwidth,
and the slope of $G(\cdot )$ at zero. In the special case $Z_{j}=X_{j}$, the
IV estimator is the OLS estimator.\footnote{
This is different from \citet{ZhouEtAl2011}, who add the $d$ OLS moment
conditions to the $d$ median regression moment conditions before estimation;
our connection to IV/OLS emerges naturally from smoothing the (IV)QR
estimating equations.}

The intercept is often not of interest, and when $q=0.5$, the adjustment is
zero anyway. The class of SEE estimators is a continuum (indexed by $h$)
with two well-known special cases at the extremes: unsmoothed IV-QR and mean
IV. For $q=0.5$ and $Z_j=X_j$, this is median regression and mean regression
(OLS). Well known are the relative efficiency advantages of
the median and the mean for different error distributions. Our estimator
with a data-driven bandwidth can harness the advantages of both, without
requiring the practitioner to make guesses about the unknown error
distribution.


\subsubsection*{Robust estimation}

With $Z_j=X_j$, the result that our SEE can yield OLS when $h\to \infty$ or
median regression when $h=0$ calls to mind robust estimators like the
trimmed or Winsorized mean (and corresponding regression estimators).
Setting the trimming/Winsorization parameter to zero generates the mean
while the other extreme generates the median. However, our SEE mechanism is
different and more general/flexible; trimming/Winsorization is not directly
applicable to $q\ne0.5$; our method to select the smoothing parameter is
novel; and the motivations for QR extend beyond (though include) robustness.

With $X_{i}=1$ and $q=0.5$ (population median estimation), our SEE becomes
\begin{equation*}
0=n^{-1}\sum_{i=1}^{n}\mathopen{}\mathclose \bgroup \originalleft[ 2G\mathopen{}\mathclose \bgroup \originalleft(\frac{\beta -Y_{i}}{h}\aftergroup \egroup \originalright)-1\aftergroup \egroup \originalright] .
\end{equation*}
If $G'(u)=1\{-1\le u\le 1\}/2$ (the uniform kernel), then $H(u)\equiv 2G(u)-1=u$ for $u\in[-1,1]$, $H(u)=1$ for $u>1$, and $H(u)=-1$ for $u<-1$.  The SEE is then $0=\sum_{i=1}^{n}\psi\mathopen{}\mathclose \bgroup \originalleft(Y_i;\beta\aftergroup \egroup \originalright)$ with $\psi\mathopen{}\mathclose \bgroup \originalleft(Y_i;\beta\aftergroup \egroup \originalright)=H\mathopen{}\mathclose \bgroup \originalleft(\mathopen{}\mathclose \bgroup \originalleft(\beta-Y_i\aftergroup \egroup \originalright)/h\aftergroup \egroup \originalright)$.  This produces the Winsorized mean estimator of the type in \citet[example (iii), p.\ 79]{Huber1964}.\footnote{
For a strict mapping, multiply by $h$ to get $\psi (Y_{i};\beta )=hH[(\beta
-Y_{i})/h]$. The solution is equivalent since $\sum h\psi (Y_{i};\beta )=0$
is the same as $\sum \psi (Y_{i};\beta )=0$ for any nonzero constant $h$.}



Further theoretical comparison of our SEE-QR with trimmed/Winsorized mean
regression (and the IV versions) would be interesting but is beyond the
scope of this paper. For more on robust location and regression estimators,
see for example \citet{Huber1964}, \citet{KoenkerBassett1978}, and
\citet{RuppertCarroll1980}.

\section{MSE of the SEE\label{sec:mse}}

Since statistical inference can be made based on the estimating equations
(EEs), we examine the mean squared error (MSE) of the SEE.
An advantage of using EEs directly is that inference can be made robust to the
strength of identification. Our focus on the EEs is also in the same
spirit of the large literature on optimal estimating equations. For the
historical developments of EEs and their applications in econometrics, see
\citet{BeraEtAl2006}. The MSE of the SEE is also related to the estimator MSE
and inference properties both intuitively and (as we will show) theoretically. Such results may provide
helpful guidance in contexts where the SEE MSE is easier to compute than the
estimator MSE, and it provides insight into how smoothing works in the QR
model as well as results that will be used in subsequent sections.

We maintain different subsets of the following assumptions for different
results.
We write $f_{U|Z}(\cdot \mid z)$ and $F_{U|Z}(\cdot \mid z)$ as the
conditional PDF and CDF of $U$ given $Z=z$. We define $f_{U|Z,X}(\cdot \mid
z,x)$ and $F_{U|Z,X}(\cdot \mid z,x) $ similarly.

\begin{assumption}
\label{a:sampling} $(X_{j}^{\prime },Z_{j}^{\prime },Y_{j})$ is iid across $
j=1,2,\ldots ,n$, where $Y_{j}=X_{j}^{\prime }\beta _{0}+U_{j}$, $X_{j}$ is
an observed $d\times 1$ vector of stochastic regressors that can include a
constant, $\beta _{0}$ is an unknown $d\times 1$ constant vector, $U_{j}$ is
an unobserved random scalar, and $Z_{j}$ is an observed $d\times 1$ vector
of instruments such that $\mathbb E\mathopen{}\mathclose \bgroup \originalleft[Z_{j}\mathopen{}\mathclose \bgroup \originalleft( 1\{U_{j}<0\}-q\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] =0$.
\end{assumption}

\begin{assumption}
\label{a:rank} (i) $Z_{j}$ has bounded support. (ii) $\mathbb E\mathopen{}\mathclose \bgroup \originalleft(
Z_{j}Z_{j}^{\prime }\aftergroup \egroup \originalright) $ is nonsingular.
\end{assumption}

\begin{assumption}
\label{a:fUZ}(i) $P(U_{j}<0\mid Z_{j}=z)=q$ for almost all $z\in \mathcal{Z}$
, the support of $Z$. (ii) For all $u$ in a neighborhood of zero and almost
all $z\in \mathcal{Z}$, $f_{U|Z}(u\mid z)$ exists, is bounded away from
zero, and is $r$ times continuously differentiable with $r\geq 2$. (iii)
There exists a function $C(z)$ such that $\mathopen{}\mathclose \bgroup \originalleft \vert f_{U|Z}^{(s)}(u\mid
z)\aftergroup \egroup \originalright
\vert \leq C(z)$ for $s=0,2,\ldots ,r$, almost all $z\in \mathcal{Z
}$ and $u$ in a neighborhood of zero, and $\mathbb E\mathopen{}\mathclose \bgroup \originalleft[ C(Z)\mathopen{}\mathclose \bgroup \originalleft
\Vert
Z\aftergroup \egroup \originalright
\Vert ^{2}\aftergroup \egroup \originalright] <\infty $.
\end{assumption}

\begin{assumption}
\label{a:G} (i) $G(v)$ is a bounded function satisfying $G(v)=0$ for $
v\leq-1 $, $G(v)=1$ for $v\geq 1$, and $1-\int_{-1}^{1}G^{2}(u)du>0$. (ii) $
G^{\prime }(\cdot )$ is a symmetric and bounded $r$th order kernel with $
r\geq 2$ so that $\int_{-1}^{1}G^{\prime }(v)dv=1$, $\int_{-1}^{1}v^{k}G^{
\prime }(v)dv=0$ for $k=1,2,\ldots ,r-1$, $\int_{-1}^{1}\mathopen{}\mathclose \bgroup \originalleft \vert
v^{r}G^{\prime }(v)\aftergroup \egroup \originalright \vert dv<\infty $, and $\int_{-1}^{1}v^{r}G^{
\prime }(v)dv\neq 0$. (iii) Let $\tilde{G}(u)=\mathopen{}\mathclose \bgroup \originalleft( G(u),[G(u)]^{2},\ldots
,[G(u)]^{L+1}\aftergroup \egroup \originalright) ^{\prime }$ for some $L\geq 1$. For any $\theta \in
\mathbb{R}^{L+1}$ satisfying $\mathopen{}\mathclose \bgroup \originalleft \Vert \theta \aftergroup \egroup \originalright \Vert =1$, there is
a partition of $[-1,1]$ given by $-1=a_{0}<a_{1}<\cdots <a_{\tilde L}=1$ for some finite $\tilde L$ such
that $\theta ^{\prime }\tilde{G}(u)$ is either strictly positive or strictly
negative on the intervals $(a_{i-1},a_{i})$ for $i=1,2,\ldots ,\tilde L$.
\end{assumption}

\begin{assumption}
\label{a:h} $h\propto n^{-\kappa }$ for $1/\mathopen{}\mathclose \bgroup \originalleft( 2r\aftergroup \egroup \originalright) <\kappa <1$.
\end{assumption}

\begin{assumption}
\label{a:beta} $\beta=\beta_0$ uniquely solves $\mathbb E\mathopen{}\mathclose \bgroup \originalleft[ Z_{j}\mathopen{}\mathclose \bgroup \originalleft(
q-1\{Y_{j}<X_{j}^{\prime }\beta \} \aftergroup \egroup \originalright) \aftergroup \egroup \originalright] =0$ over $\beta \in
\mathcal{B}$.
\end{assumption}

\begin{assumption}
\label{a:power_mse}(i) $f_{U|Z,X}(u\mid z,x)$ is $r$ times continuously
differentiable in $u$ in a neighborhood of zero for almost all $x\in
\mathcal{X}$ and $z\in \mathcal{Z}$ for $r>2$. (ii) $\Sigma _{ZX}\equiv \mathbb E
\mathopen{}\mathclose \bgroup \originalleft[ Z_{j}X_{j}^{\prime }f_{U|Z,X}(0\mid Z_{j},X_{j})\aftergroup \egroup \originalright] $ is
nonsingular.
\end{assumption}


Assumption \ref{a:sampling} describes the sampling process. Assumption \ref
{a:rank} is analogous to Assumption 3 in both \citet{Horowitz1998} and
\citet{Whang2006}. As discussed in these two papers, the boundedness
assumption for $Z_{j}$, which is a technical condition, is made only for
convenience and can be dropped at the cost of more complicated proofs.

Assumption \ref{a:fUZ}(i) allows us to use the law of iterated expectations
to simplify the asymptotic variance. Our qualitative conclusions do not rely
on this assumption. Assumption \ref{a:fUZ}(ii) is critical. If we are not
willing to make such an assumption, then smoothing will be of no benefit.
Inversely, with some small degree of smoothness of the conditional error
density, smoothing can leverage this into the advantages described here.
Also note that \citet{Horowitz1998} assumes $r\geq 4$, which is sufficient
for the estimator MSE result in Section \ref{sec:est}.

Assumptions \ref{a:G}(i--ii) are analogous to the standard high-order kernel
conditions in the kernel smoothing literature. The integral condition in (i)
ensures that smoothing reduces (rather than increases) variance.  Note that
\begin{align*}
1-\int_{-1}^{1}G^{2}(u)du& =2\int_{-1}^{1}uG(u)G^{\prime }(u)du \\
& =2\int_{0}^{1}uG(u)G^{\prime }(u)du+2\int_{-1}^{0}uG(u)G^{\prime }(u)du \\
& =2\int_{0}^{1}uG(u)G^{\prime }(u)du-2\int_{0}^{1}vG(-v)G^{\prime }(-v)dv \\
& =2\int_{0}^{1}uG^{\prime }(u)\mathopen{}\mathclose \bgroup \originalleft[ G(u)-G(-u)\aftergroup \egroup \originalright] du,
\end{align*}
using the evenness of $G'(u)$.  When $r=2$, we can use any $G(u)$ such that $G'(u)$ is a symmetric PDF on $[-1,1]$.
In this case, $1-\int_{-1}^1 G^2(u)du>0$ holds automatically.
When $r>2$, $G'(u)<0$ for some $u$, and $G(u)$ is not monotonic.  It is not easy to sign $1-\int_{-1}^1 G^2(u)du$ generally, but it is simple to calculate this quantity for any chosen $G(\cdot)$.  For example, consider $r=4$ and the $G(\cdot)$ function in \citet{Horowitz1998} and \citet{Whang2006} shown in Figure \ref{fig:G}:
\begin{equation}
G(u)=\mathopen{}\mathclose \bgroup \originalleft \{
\begin{array}{ll}
0, & u\leq -1 \\
0.5+\frac{105}{64}\mathopen{}\mathclose \bgroup \originalleft( u-\frac{5}{3}u^{3}+\frac{7}{5}u^{5}-\frac{3}{7}
u^{7}\aftergroup \egroup \originalright) , & u\in \lbrack -1,1] \\
1 & u\geq 1
\end{array}
\aftergroup \egroup \originalright.   \label{eqn:G_fun}
\end{equation}
The range of the function is outside $[0,1]$.  Simple calculations show that $1-\int_{-1}^1 G^2(u)du>0$.

\begin{figure}[htbp]
\begin{center}
\includegraphics[width=0.7\textwidth,clip=true,trim=35 35 20 70]{Graph_G_from_HorowitzWhang.pdf}
\end{center}
\caption{Graph of $G(u)=0.5+\frac{105}{64}\mathopen{}\mathclose \bgroup \originalleft( u-\frac{5}{3}u^{3}+\frac{7}{
5}u^{5}-\frac{3}{7}u^{7}\aftergroup \egroup \originalright) $ (solid line) and its derivative (broken).}
\label{fig:G}
\end{figure}

Assumption \ref{a:G}(iii) is needed for the Edgeworth expansion. As \citet{Horowitz1998} and
\citet{Whang2006} discuss, Assumption \ref{a:G}(iii) is a technical
assumption that (along with Assumption \ref{a:h}) leads to a form of Cram\'{e
}r's condition, which is needed to justify the Edgeworth expansion used in
Section \ref{sec:eI}. Any $G(u)$ constructed by integrating polynomial kernels in \citet{Muller1984} satisfies Assumption \ref{a:G}(iii).  In fact, $G(u)$ in \eqref{eqn:G_fun} is obtained by integrating a fourth-order kernel given in Table 1 of \citet{Muller1984}.
Assumption \ref{a:h} ensures that the bias of the SEE is of
smaller order than its variance. It is needed for the asymptotic normality
of the SEE as well as the Edgeworth expansion.

Assumption \ref{a:beta} is an identification assumption. See Theorem 2 of
\citet{ChernozhukovHansen2006} for more primitive conditions. It ensures the
consistency of the SEE estimator. Assumption \ref{a:power_mse} is necessary
for the $\sqrt{n}$-consistency and asymptotic normality of the SEE estimator.

Define
\begin{equation*}
W_{j}\equiv W_{j}(\beta _{0})=Z_{j}\mathopen{}\mathclose \bgroup \originalleft[ G(-U_{j}/h)-q\aftergroup \egroup \originalright]
\end{equation*}
and abbreviate $m_{n}\equiv m_{n}(\beta _{0})=n^{-1/2}\sum_{j=1}^{n}W_{j}$.
The theorem below gives the first two moments of $W_{j}$ and the first-order
asymptotic distribution of $m_{n}$.

\begin{theorem}
\label{thm:Wj} Let Assumptions \ref{a:rank}(i), \ref{a:fUZ}, and \ref{a:G}
(i--ii) hold. Then
\begin{align}
\mathbb E(W_{j})& =\frac{(-h)^{r}}{r!}\mathopen{}\mathclose \bgroup \originalleft[ \int_{-1}^{1}G^{\prime
}(v)v^{r}dv\aftergroup \egroup \originalright] \mathbb E\mathopen{}\mathclose \bgroup \originalleft[ f_{U|Z}^{(r-1)}(0\mid Z_{j})Z_{j}\aftergroup \egroup \originalright] +o\mathopen{}\mathclose \bgroup \originalleft(
h^{r}\aftergroup \egroup \originalright) ,  \label{eqn:see-bias} \\
\mathbb E(W_{j}^{\prime }W_{j})
  &= q(1-q)\mathbb E\mathopen{}\mathclose \bgroup \originalleft(Z_{j}^{\prime }Z_{j}\aftergroup \egroup \originalright)
    -h\mathopen{}\mathclose \bgroup \originalleft[1-\int_{-1}^{1}G^{2}(u)du\aftergroup \egroup \originalright] \mathbb E\mathopen{}\mathclose \bgroup \originalleft[f_{U|Z}(0\mid Z_{j})Z_{j}^{\prime
}Z_{j}\aftergroup \egroup \originalright]+O(h^{2}),  \label{eqn:W-var-trace} \\
\mathbb E(W_{j}W_{j}^{\prime })& =q(1-q)\mathbb E\mathopen{}\mathclose \bgroup \originalleft(Z_{j}Z_{j}^{\prime }\aftergroup \egroup \originalright)-h\mathopen{}\mathclose \bgroup \originalleft[
1-\int_{-1}^{1}G^{2}(u)du\aftergroup \egroup \originalright] \mathbb E\mathopen{}\mathclose \bgroup \originalleft[f_{U|Z}(0\mid Z_{j})Z_{j}Z_{j}^{\prime
}\aftergroup \egroup \originalright]+O(h^{2}).  \notag
\end{align}
If additionally Assumptions \ref{a:sampling} and \ref{a:h} hold, then
\begin{equation*}
m_{n}\overset{d}{\rightarrow }N(0,V),\quad V\equiv \lim_{n\rightarrow \infty
}\mathbb E\mathopen{}\mathclose \bgroup \originalleft\{ \mathopen{}\mathclose \bgroup \originalleft[W_{j}-\mathbb E(W_{j})\aftergroup \egroup \originalright]\mathopen{}\mathclose \bgroup \originalleft[W_{j}-\mathbb E(W_{j})\aftergroup \egroup \originalright]^{\prime }\aftergroup \egroup \originalright\}
=q(1-q)\mathbb E\mathopen{}\mathclose \bgroup \originalleft(Z_{j}Z_{j}^{\prime }\aftergroup \egroup \originalright).
\end{equation*}
\end{theorem}

Compared with the EE derived from smoothing the criterion function as in \citet{Horowitz1998}, our SEE has smaller bias and variance, and these differences affect the bias and variance of the parameter estimator.
The former approach only applies to exogenous QR with $Z_j=X_j$.
The EE derived from smoothing the criterion function in \eqref{eqn:SCF-EE} for $Z_j=X_j$ can be written
\begin{align}\label{eqn:SCF-EE-Wj}
0 &= n^{-1}\sum_{j=1}^n W_j, \quad
W_j \equiv X_j\mathopen{}\mathclose \bgroup \originalleft[G(-U_j/h)-q\aftergroup \egroup \originalright] + (1/h)G'(-U_j/h)(-X_jU_j) .
\end{align}
Consequently, as calculated in the appendix,
\begin{align}\label{eqn:SCF-EWj}
\mathbb E(W_j)
  & =(r+1)\frac{(-h)^r}{r!} \mathopen{}\mathclose \bgroup \originalleft[
\int G^{\prime }(v)v^{r}dv\aftergroup \egroup \originalright] \mathbb E\mathopen{}\mathclose \bgroup \originalleft[ f_{U|Z}^{(r-1)}(0\mid Z_{j})Z_{j}
\aftergroup \egroup \originalright] +o\mathopen{}\mathclose \bgroup \originalleft( h^{r}\aftergroup \egroup \originalright) , \\
\mathbb E(W_jW_j') \label{eqn:SCF-EWjWj}
  &= q(1-q)\mathbb E(X_jX_j')
     +h \int_{-1}^1[G'(v)v]^2dv \, \mathbb E\mathopen{}\mathclose \bgroup \originalleft[f_{U|X}(0\mid X_j)X_jX_j'\aftergroup \egroup \originalright]
     +O(h^2) , \\
\mathbb E & \mathopen{}\mathclose \bgroup \originalleft[\frac{\partial }{\partial \beta'} n^{-1/2}m_n(\beta_0)\aftergroup \egroup \originalright] \label{eqn:SCF-EdBmn}
   = \mathbb E\mathopen{}\mathclose \bgroup \originalleft[f_{U|X}(0\mid X_j)X_jX_j'\aftergroup \egroup \originalright]
    -h \mathbb E\mathopen{}\mathclose \bgroup \originalleft[f_{U|X}'(0\mid X_j)X_jX_j'\aftergroup \egroup \originalright]
    +O(h^2) .
\end{align}
The dominating term of the bias of our SEE in \eqref{eqn:see-bias} is $r+1$ times smaller in absolute value than that of the EE derived from a smoothed criterion function in \eqref{eqn:SCF-EWj}.
A larger bias can lead to less accurate confidence regions if the same
variance estimator is used.
Additionally, the smoothed criterion function analog of $\mathbb E(W_jW_j')$ in \eqref{eqn:SCF-EWjWj} has a positive $O(h)$ term instead of the negative $O(h)$ term for SEE.
The connection between these terms and the estimator's asymptotic mean squared error (AMSE) is shown in Section \ref{sec:est} to rely on the inverse of the matrix in equation \eqref{eqn:SCF-EdBmn}.  Here, though, the sign of the $O(h)$ term is indeterminant since it depends on a PDF derivative.  (A negative $O(h)$ term implies higher AMSE since this matrix is inverted in the AMSE expression, and positive implies lower.)  If $U=0$ is a mode of the conditional (on $X$) distribution, then the $O(h)$ term is zero and the AMSE comparison is driven by $\mathbb E(W_j)$ and $\mathbb E(W_jW_j')$.  Since SEE yields smaller $\mathbb E(W_jW_j')$ and smaller absolute $\mathbb E(W_j)$, it will have smaller estimator AMSE in such cases.  Simulation results in Section \ref{sec:sim} add evidence that the SEE estimator usually has smaller MSE in practice.

The first-order asymptotic variance $V$ is the same as the asymptotic
variance of
\begin{equation*}
n^{-1/2}\sum_{j=1}^n Z_j \mathopen{}\mathclose \bgroup \originalleft(1\{U_j<0\}-q\aftergroup \egroup \originalright) ,
\end{equation*}
the scaled EE of the unsmoothed IV-QR. The effect of smoothing to reduce
variance is captured by the term of order $h$, where $1-
\int_{-1}^{1}G^2(u)du>0$ by Assumption \ref{a:G}(i). This reduction in
variance is not surprising. Replacing the discontinuous indicator function $
1\{U<0\}$ by a smooth function $G(-U/h)$ pushes the dichotomous values of
zero and one into some values in between, leading to a smaller variance. The
idea is similar to \citeauthor{Breiman1994}'s (\citeyear{Breiman1994})
bagging (bootstrap aggregating), among others.

Define the MSE of the SEE to be $\mathbb E\mathopen{}\mathclose \bgroup \originalleft(m_{n}^{\prime }V^{-1}m_{n}\aftergroup \egroup \originalright)$. Building
upon \eqref{eqn:see-bias} and \eqref{eqn:W-var-trace}, and using $W_{i}
\mathpalette{\protect \independenT}{\perp}W_{j}$ for $i\neq j$, we have:
\begin{align}
& \mathbb E\mathopen{}\mathclose \bgroup \originalleft(m_{n}^{\prime }V^{-1}m_{n}\aftergroup \egroup \originalright)  \notag \\
& =\frac{1}{n}\sum_{j=1}^{n}\mathbb E\mathopen{}\mathclose \bgroup \originalleft(W_{j}^{\prime }V^{-1}W_{j}\aftergroup \egroup \originalright)+\frac{1}{n}
\sum_{j=1}^{n}\sum_{i\neq j}\mathbb E\mathopen{}\mathclose \bgroup \originalleft( W_{i}^{\prime }V^{-1}W_{j}\aftergroup \egroup \originalright)
\notag \\
& =\frac{1}{n}\sum_{j=1}^{n}\mathbb E\mathopen{}\mathclose \bgroup \originalleft(W_{j}^{\prime }V^{-1}W_{j}\aftergroup \egroup \originalright)+\frac{1}{n}
n(n-1)\mathbb E(W_{j}^{\prime })V^{-1}\mathbb E(W_{j})  \notag \\
& =q(1-q)\mathbb E\mathopen{}\mathclose \bgroup \originalleft(Z_{j}^{\prime }V^{-1}Z_{j}\aftergroup \egroup \originalright)+nh^{2r} \mathbb E(B)'\mathbb E(B) -h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[ \mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA^{\prime }\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]
+o\mathopen{}\mathclose \bgroup \originalleft(h+nh^{2r}\aftergroup \egroup \originalright),  \notag \\
& =d+nh^{2r} \mathbb E(B)'\mathbb E(B) -h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[
\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA^{\prime }\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] +o\mathopen{}\mathclose \bgroup \originalleft(h+nh^{2r}\aftergroup \egroup \originalright),  \label{eqn:mse}
\end{align}
where
\begin{align*}
A& \equiv \mathopen{}\mathclose \bgroup \originalleft[ 1-\int_{-1}^{1}G^{2}(u)du\aftergroup \egroup \originalright] ^{1/2}\mathopen{}\mathclose \bgroup \originalleft[ f_{U|Z}(0\mid
Z)\aftergroup \egroup \originalright] ^{1/2}V^{-1/2}Z, \\
B& \equiv \mathopen{}\mathclose \bgroup \originalleft[ \frac{1}{r!}\int_{-1}^{1}G^{\prime }(v)v^{r}dv\aftergroup \egroup \originalright]
f_{U|Z}^{(r-1)}(0\mid Z)V^{-1/2}Z.
\end{align*}

Ignoring the $o(\cdot )$ term, we obtain the asymptotic MSE of the SEE. We
select the smoothing parameter to minimize the asymptotic MSE:
\begin{equation}
h_{\text{SEE}}^{\ast }
   \equiv \mathop{\rm arg\,min}_{h}
              nh^{2r}\mathbb E(B)'\mathbb E(B)
            - h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[ \mathbb E\mathopen{}\mathclose \bgroup \originalleft(AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] .  \label{eqn:def_h_SEE}
\end{equation}
The proposition below gives the optimal smoothing parameter $h_{\text{SEE}
}^{\ast }$.

\begin{proposition}
\label{prop:hSEE} Let Assumptions \ref{a:sampling}, \ref{a:rank}, \ref{a:fUZ}
, and \ref{a:G}(i--ii) hold. The bandwidth that minimizes the asymptotic MSE
of the SEE is
\begin{equation*}
h_{\text{SEE}}^{\ast }
   = \mathopen{}\mathclose \bgroup \originalleft( \frac{\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[ \mathbb E\mathopen{}\mathclose \bgroup \originalleft(AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] }{\mathbb E(B)'\mathbb E(B) }\frac{1}{2nr}\aftergroup \egroup \originalright) ^{
\frac{1}{2r-1}}.
\end{equation*}
Under the stronger assumption $U\mathpalette{\protect \independenT}{\perp} Z$
,
\begin{equation*}
h_{\text{SEE}}^{\ast }=\mathopen{}\mathclose \bgroup \originalleft( \frac{\mathopen{}\mathclose \bgroup \originalleft( r!\aftergroup \egroup \originalright) ^{2}\mathopen{}\mathclose \bgroup \originalleft[
1-\int_{-1}^{1}G^{2}(u)du\aftergroup \egroup \originalright] f_{U}(0)}{2r\mathopen{}\mathclose \bgroup \originalleft[ \int_{-1}^{1}G^{\prime
}(v)v^{r}dv\aftergroup \egroup \originalright] ^{2}\mathopen{}\mathclose \bgroup \originalleft[ f_{U}^{\mathopen{}\mathclose \bgroup \originalleft( r-1\aftergroup \egroup \originalright) }(0)\aftergroup \egroup \originalright] ^{2}}
\frac{d}{n}\aftergroup \egroup \originalright) ^{\frac{1}{2r-1}}.
\end{equation*}
\end{proposition}

When $r=2$, the MSE-optimal $h_{\text{SEE}}^{\ast }\asymp
n^{-1/(2r-1)}=n^{-1/3}$. This is smaller than $n^{-1/5}$, the rate that
minimizes the MSE of estimated standard errors of the usual regression
quantiles. Since nonparametric estimators of $f_{U}^{(r-1)}(0)$ converge
slowly, we propose a parametric plug-in described in Section \ref{sec:sim}.

We point out in passing that the optimal smoothing parameter $h_{\text{SEE}
}^{\ast }$ is invariant to rotation and translation of the (non-constant)
regressors. This may not be obvious but can be proved easily.

For the unsmoothed IV-QR, let
\begin{equation*}
\tilde{m}_{n} = \frac{1}{\sqrt{n}}\sum_{j=1}^{n}Z_{j}\mathopen{}\mathclose \bgroup \originalleft( 1\mathopen{}\mathclose \bgroup \originalleft \{
Y_{j}\leq X_{j}'\beta \aftergroup \egroup \originalright \} -q\aftergroup \egroup \originalright) ,
\end{equation*}
then the MSE of the estimating equations is $\mathbb E\mathopen{}\mathclose \bgroup \originalleft(\tilde{m}_{n}^{\prime }V^{-1}
\tilde{m}_{n}\aftergroup \egroup \originalright)=d$. Comparing this to the MSE of the SEE given in
\eqref{eqn:mse}, we find that the SEE has a smaller MSE when $h=h_{\text{SEE}
}^{\ast }$ because
\begin{equation*}
n(h_{\text{SEE}}^{\ast })^{2r} \mathbb E(B)'\mathbb E(B)
-h_{\text{SEE}}^{\ast }\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[ \mathbb E\mathopen{}\mathclose \bgroup \originalleft(AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]
=-h_{\text{SEE}}^{\ast }\mathopen{}\mathclose \bgroup \originalleft( 1-\frac{1}{2r}\aftergroup \egroup \originalright) \mathrm{tr}
\mathopen{}\mathclose \bgroup \originalleft[ \mathbb E\mathopen{}\mathclose \bgroup \originalleft(AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] <0.
\end{equation*}
In terms of MSE, it is advantageous to smooth the estimating equations. To
the best of our knowledge, this point has never been discussed before in the
literature.

\section{Type {I} and Type {II} Errors of a Chi-square Test\label{sec:eI}}

In this section, we explore the effect of smoothing on a chi-square test.
Other alternatives for inference exist, such as the Bernoulli-based
MCMC-computed method from \citet{ChernozhukovEtAl2009}, empirical likelihood
as in \citet{Whang2006}, and bootstrap as in \citet{Horowitz1998}, where the
latter two also use smoothing. Intuitively, when we minimize the MSE, we may
expect lower type I error: the $\chi ^{2}$ critical value is from the
unsmoothed distribution, and smoothing to minimize MSE makes large values
(that cause the test to reject) less likely.
The reduced MSE also makes it easier to distinguish the null hypothesis from
some given alternative. This combination leads to improved size-adjusted
power. As seen in our simulations, this is true especially for the IV case.

Using the results in Section \ref{sec:mse} and under Assumption \ref{a:h}, we
have
\begin{equation*}
m_{n}^{\prime }V^{-1}m_{n}\overset{d}{\rightarrow }\chi _{d}^{2},
\end{equation*}
where we continue to use the notation $m_{n}\equiv m_{n}(\beta _{0})$. From
this asymptotic result, we can construct a hypothesis test that rejects the
null hypothesis $H_{0}:\beta =\beta _{0}$ when
\begin{equation*}
S_{n}\equiv m_{n}^{\prime }\hat{V}^{-1}m_{n}>c_{\alpha },
\end{equation*}
where
\begin{equation*}
\hat{V}=q(1-q)\frac{1}{n}\sum_{j=1}^{n}Z_{j}Z_{j}^{\prime }
\end{equation*}
is a consistent estimator of $V$ and $c_{\alpha }\equiv \chi _{d,1-\alpha
}^{2}$ is the $1-\alpha $ quantile of the chi-square distribution
with $d$ degrees of freedom. As desired, the asymptotic size is
\begin{equation*}
\lim_{n\rightarrow \infty }P\mathopen{}\mathclose \bgroup \originalleft( S_{n}>c_{\alpha }\aftergroup \egroup \originalright) =\alpha .
\end{equation*}
Here $P\equiv P_{\beta _{0}}$ is the probability measure under the true
model parameter $\beta _{0}$. We suppress the subscript $\beta _{0}$ when
there is no confusion.

It is important to point out that the above result does not rely on the
strong identification of $\beta _{0}$. It still holds if $\beta _{0}$ is
weakly identified or even unidentified. This is an advantage of focusing on
the estimating equations instead of the parameter estimator. When a direct
inference method based on the asymptotic normality of $\hat{\beta}$ is used,
we have to impose Assumptions \ref{a:beta} and \ref{a:power_mse}.

\subsection{Type {I} error and the associated optimal bandwidth}

To more precisely measure the type I error $P\mathopen{}\mathclose \bgroup \originalleft( S_{n}>c_{\alpha }\aftergroup \egroup \originalright)
$, we first develop a high-order stochastic expansion of $S_{n}$. Let $
V_{n}\equiv \mathrm{Var}\mathopen{}\mathclose \bgroup \originalleft( m_{n}\aftergroup \egroup \originalright) $. Following the same
calculation as in \eqref{eqn:mse}, we have
\begin{align*}
V_{n}& =V-h\mathopen{}\mathclose \bgroup \originalleft[ 1-\int_{-1}^{1}G^{2}(u)du\aftergroup \egroup \originalright] \mathbb E\mathopen{}\mathclose \bgroup \originalleft[f_{U|Z}(0\mid
Z_{j})Z_{j}Z_{j}^{\prime }\aftergroup \egroup \originalright]+O(h^{2}) \\
& =V^{1/2}\mathopen{}\mathclose \bgroup \originalleft[ I_{d}-h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) +O\mathopen{}\mathclose \bgroup \originalleft(h^{2}\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] \mathopen{}\mathclose \bgroup \originalleft(
V^{1/2}\aftergroup \egroup \originalright) ^{\prime },
\end{align*}
where $V^{1/2}$ is the matrix square root of $V$ such that $V^{1/2}\mathopen{}\mathclose \bgroup \originalleft(
V^{1/2}\aftergroup \egroup \originalright) ^{\prime }=V$. We can choose $V^{1/2}$ to be symmetric but do
not have to.

Details of the following are in the appendix; here we outline our strategy
and highlight key results. Letting
\begin{equation}
\Lambda _{n}=V^{1/2}\mathopen{}\mathclose \bgroup \originalleft[ I_{d}-h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) +O\mathopen{}\mathclose \bgroup \originalleft(h^{2}\aftergroup \egroup \originalright)
\aftergroup \egroup \originalright] ^{1/2}  \label{Lambda_n}
\end{equation}
such that $\Lambda _{n}\Lambda _{n}^{\prime }=V_{n}$, and defining
\begin{equation}
\bar{W}_{n}^{\ast }\equiv \frac{1}{n}\sum_{j=1}^{n}W_{j}^{\ast }\text{ and }
W_{j}^{\ast }=\Lambda _{n}^{-1}Z_{j}\mathopen{}\mathclose \bgroup \originalleft[ G(-U_{j}/h)-q\aftergroup \egroup \originalright] ,
\label{define_W_star}
\end{equation}
we can approximate the test statistic as
$S_{n}=S_{n}^{L}+e_{n}$,
where
\begin{equation*}
S_{n}^{L}=\mathopen{}\mathclose \bgroup \originalleft( \sqrt{n}\bar{W}_{n}^{\ast }\aftergroup \egroup \originalright) ^{\prime }\mathopen{}\mathclose \bgroup \originalleft( \sqrt{n}
\bar{W}_{n}^{\ast }\aftergroup \egroup \originalright) -h\mathopen{}\mathclose \bgroup \originalleft( \sqrt{n}\bar{W}_{n}^{\ast }\aftergroup \egroup \originalright)
^{\prime }\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \mathopen{}\mathclose \bgroup \originalleft( \sqrt{n}\bar{W}_{n}^{\ast
}\aftergroup \egroup \originalright)
\end{equation*}
and $e_{n}$ is the remainder term satisfying $P\mathopen{}\mathclose \bgroup \originalleft( \mathopen{}\mathclose \bgroup \originalleft \vert
e_{n}\aftergroup \egroup \originalright \vert >O\mathopen{}\mathclose \bgroup \originalleft( h^{2}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright) =O\mathopen{}\mathclose \bgroup \originalleft( h^{2}\aftergroup \egroup \originalright) $.

The stochastic expansion above allows us to approximate the characteristic
function of $S_{n}$ with that of $S_{n}^{L}$. Taking the Fourier--Stieltjes
inverse of the characteristic function yields an approximation of the
distribution function, from which we can calculate the type I error by
plugging in the critical value $c_{\alpha}$.

\begin{theorem}
\label{thm:inf} Under Assumptions \ref{a:sampling}--\ref{a:h}, we have
\begin{align*}
P\mathopen{}\mathclose \bgroup \originalleft(S_{n}^{L}<x\aftergroup \egroup \originalright)
  &= \mathcal{G}_{d}(x) -\mathcal{G}_{d+2}^{\prime}(x) \mathopen{}\mathclose \bgroup \originalleft \{
nh^{2r}\mathbb E(B)'\mathbb E(B) -h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft
[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright \} +R_n, \\
P\mathopen{}\mathclose \bgroup \originalleft( S_{n}>c_{\alpha }\aftergroup \egroup \originalright) &= \alpha +\mathcal{G}_{d+2}^{\prime
}(c_{\alpha }) \mathopen{}\mathclose \bgroup \originalleft \{ nh^{2r}\mathbb E(B)'\mathbb E(B) -h\mathrm{tr}
\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright \} +R_{n},
\end{align*}
where $R_{n}=O\mathopen{}\mathclose \bgroup \originalleft(h^{2}+nh^{2r+1}\aftergroup \egroup \originalright)$ and $\mathcal{G}_{d}(x)$ is the CDF of the $
\chi _{d}^{2}$ distribution.
\end{theorem}

From Theorem \ref{thm:inf}, an approximate measure of the type I error of
the SEE-based chi-square test is
\begin{equation*}
\alpha +\mathcal{G}_{d+2}^{\prime }(c_{\alpha })\mathopen{}\mathclose \bgroup \originalleft\{ nh^{2r}\mathbb E(B)'\mathbb E(B) -h\mathrm{tr}
\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright\} ,
\end{equation*}
and an approximate measure of the coverage probability error (CPE) is
\footnote{
The CPE is defined to be the nominal coverage minus the true coverage
probability, which may be different from the usual definition. Under this
definition, smaller CPE corresponds to higher coverage probability (and
smaller type I error).}
\begin{equation*}
\mathrm{CPE}=\mathcal{G}_{d+2}^{\prime }(c_{\alpha })\mathopen{}\mathclose \bgroup \originalleft\{ nh^{2r}\mathbb E(B)'\mathbb E(B) -h\mathrm{tr}
\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright\} ,
\end{equation*}
which is also the error in rejection probability under the null.

Up to smaller-order terms, the term $nh^{2r}\mathbb E(B)'\mathbb E(B)$
characterizes the bias effect from smoothing. The bias increases type I
error and reduces coverage probability. The term $h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]$ characterizes the variance
effect from smoothing. The variance reduction decreases type I error and
increases coverage probability.
The type I error is $\alpha $ up to order $O\mathopen{}\mathclose \bgroup \originalleft(h+nh^{2r}\aftergroup \egroup \originalright)$.
There exists some $h>0$ that makes bias and variance effects cancel, leaving
type I error equal to $\alpha$ up to smaller-order terms in $R_{n}$.

Note that $nh^{2r}\mathbb E(B)'\mathbb E(B) -h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]$ is identical to the
high-order term in the asymptotic MSE of the SEE in \eqref{eqn:mse}. The $h_{
\text{CPE}}^{\ast }$ that minimizes type I error is the same as $h_{\text{SEE
}}^{\ast }$.

\begin{proposition}
\label{prop:hCPE} Let Assumptions \ref{a:sampling}--\ref{a:h} hold. The
bandwidth that minimizes the approximate type I error of the chi-square test
based on the test statistic $S_{n}$ is
\begin{equation*}
h_{\text{CPE}}^{\ast }=h_{\text{SEE}}^{\ast }=\mathopen{}\mathclose \bgroup \originalleft( \frac{\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]}{\mathbb E(B)'\mathbb E(B)}
\frac{1}{2nr}\aftergroup \egroup \originalright) ^{\frac{1}{2r-1}}.
\end{equation*}
\end{proposition}

The result that $h_{\text{CPE}}^{\ast }=h_{\text{SEE}}^{\ast }$ is
intuitive. Since $h_{\text{SEE}}^{\ast }$ minimizes $\mathbb E\mathopen{}\mathclose \bgroup \originalleft(m_{n}^{\prime
}V^{-1}m_{n}\aftergroup \egroup \originalright)$, for a test with $c_{\alpha }$ and $\hat{V}$ both invariant
to $h$, the null rejection probability $P\mathopen{}\mathclose \bgroup \originalleft(m_{n}^{\prime }\hat{V}
^{-1}m_{n}>c_{\alpha }\aftergroup \egroup \originalright)$ should be smaller when the SEE's MSE is smaller.

When $h=h_{\text{CPE}}^{\ast }$,
\begin{equation*}
P\mathopen{}\mathclose \bgroup \originalleft( S_{n}>c_{\alpha }\aftergroup \egroup \originalright) =\alpha -C^{+}\mathcal{G}_{d+2}^{\prime
}(c_{\alpha })h_{\text{CPE}}^{\ast }\mathopen{}\mathclose \bgroup \originalleft[1+o(1)\aftergroup \egroup \originalright]
\end{equation*}
where $C^{+}=\mathopen{}\mathclose \bgroup \originalleft( 1-\frac{1}{2r}\aftergroup \egroup \originalright) \mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] >0$. If instead we construct the test statistic
based on the unsmoothed estimating equations, $\tilde{S}_{n}=\tilde{m}
_{n}^{\prime }\hat{V}^{-1}\tilde{m}_{n}$, then it can be shown that
\begin{equation*}
P\mathopen{}\mathclose \bgroup \originalleft( \tilde{S}_{n}>c_{\alpha }\aftergroup \egroup \originalright) =\alpha +Cn^{-1/2}\mathopen{}\mathclose \bgroup \originalleft[1+o(1)\aftergroup \egroup \originalright]
\end{equation*}
for some constant $C$, which is in general not equal to zero. Given that $
n^{-1/2}=o(h_{\text{CPE}}^{\ast })$ and $C^{+}>0$, we can expect the
SEE-based chi-square test to have a smaller type I error in large samples.




\subsection{Type II error and local asymptotic power}

To obtain the local asymptotic power of the $S_{n}$ test, we let the true
parameter value be $\beta _{n}=\beta _{0}-\delta /\sqrt{n}$, where $\beta _{0}
$ is the parameter value that satisfies the null hypothesis $H_{0}$. In this
case,
\begin{equation*}
m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0}\aftergroup \egroup \originalright) =\frac{1}{\sqrt{n}}\sum_{j=1}^{n}Z_{j}\mathopen{}\mathclose \bgroup \originalleft[
G\mathopen{}\mathclose \bgroup \originalleft( \frac{X_{j}^{\prime }\delta /\sqrt{n}-U_{j}}{h}\aftergroup \egroup \originalright) -q\aftergroup \egroup \originalright] .
\end{equation*}
In the proof of Theorem \ref{thm:power}, we show that
\begin{align*}
\mathbb E\mathopen{}\mathclose \bgroup \originalleft[m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0}\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] & =\Sigma _{ZX}\delta +\sqrt{n}
(-h)^{r}V^{1/2}\mathbb E(B)+O\mathopen{}\mathclose \bgroup \originalleft( n^{-1/2}+\sqrt{n}h^{r+1}\aftergroup \egroup \originalright) , \\
V_{n}& =\mathrm{Var}\mathopen{}\mathclose \bgroup \originalleft[ m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]
=V-hV^{1/2}\mathopen{}\mathclose \bgroup \originalleft[ \mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] (V^{1/2})^{\prime }+O\mathopen{}\mathclose \bgroup \originalleft( n^{-1/2}+h^{2}\aftergroup \egroup \originalright) .
\end{align*}

\begin{theorem}
\label{thm:power} Let Assumptions \ref{a:sampling}--\ref{a:h} and \ref{a:power_mse}(i) hold. Define $\Delta \equiv \mathbb E\mathopen{}\mathclose \bgroup \originalleft[V_{n}^{-1/2}m_{n}(\beta _{0})\aftergroup \egroup \originalright]$ and $\tilde{\delta}\equiv V^{-1/2}\Sigma_{ZX}\delta $. We have
\begin{align*}
P_{\beta _{n}}\mathopen{}\mathclose \bgroup \originalleft( S_{n}<x\aftergroup \egroup \originalright)
  &= \mathcal{G}_{d}\mathopen{}\mathclose \bgroup \originalleft( x;\mathopen{}\mathclose \bgroup \originalleft \Vert\Delta \aftergroup \egroup \originalright \Vert ^{2}\aftergroup \egroup \originalright)
    +\mathcal{G}_{d+2}^{\prime }\mathopen{}\mathclose \bgroup \originalleft(x;\Vert \Delta\Vert ^{2}\aftergroup \egroup \originalright)
       h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]  \\
& \quad +\mathcal{G}_{d+4}^{\prime }\mathopen{}\mathclose \bgroup \originalleft(x;\mathopen{}\mathclose \bgroup \originalleft \Vert \Delta \aftergroup \egroup \originalright \Vert ^{2}\aftergroup \egroup \originalright)h
\mathopen{}\mathclose \bgroup \originalleft[ \Delta ^{\prime }\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \Delta \aftergroup \egroup \originalright] +O\mathopen{}\mathclose \bgroup \originalleft(
h^{2}+n^{-1/2}\aftergroup \egroup \originalright)  \\
  &= \mathcal{G}_{d}\mathopen{}\mathclose \bgroup \originalleft( x;\Vert \tilde{\delta}\Vert ^{2}\aftergroup \egroup \originalright)
    -\mathcal{G}_{d+2}^{\prime }\mathopen{}\mathclose \bgroup \originalleft(x;\Vert \tilde{\delta}\Vert ^{2}\aftergroup \egroup \originalright)
     \mathopen{}\mathclose \bgroup \originalleft\{ nh^{2r}\mathbb E(B)'\mathbb E(B)-h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)\aftergroup \egroup \originalright]\aftergroup \egroup \originalright\}\\
& \quad +\mathopen{}\mathclose \bgroup \originalleft[ \mathcal{G}_{d+4}^{\prime }\mathopen{}\mathclose \bgroup \originalleft(x;\Vert \tilde{\delta}\Vert
^{2}\aftergroup \egroup \originalright)-\mathcal{G}_{d+2}^{\prime }\mathopen{}\mathclose \bgroup \originalleft(x;\Vert \tilde{\delta}\Vert ^{2}\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] h
\mathopen{}\mathclose \bgroup \originalleft[ \tilde{\delta}^{\prime }\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \tilde{\delta}
\aftergroup \egroup \originalright]  \\
& \quad -\mathcal{G}_{d+2}^{\prime }\mathopen{}\mathclose \bgroup \originalleft(x;\Vert \tilde{\delta}\Vert ^{2}\aftergroup \egroup \originalright)2
\tilde{\delta}^{\prime }\sqrt{n}(-h)^{r}\mathbb E(B)+O\mathopen{}\mathclose \bgroup \originalleft( h^{2}+n^{-1/2}\aftergroup \egroup \originalright) ,
\end{align*}
where $\mathcal{G}_{d}(x;\lambda )$ is the CDF of the noncentral chi-square
distribution with degrees of freedom $d$ and noncentrality parameter $
\lambda $. If we further assume that $\tilde{\delta}$ is uniformly
distributed on the sphere $\mathcal{S}_{d}(\tau )=\{ \tilde{\delta}\in
\mathbb{R}^{d}:\Vert \tilde{\delta}\Vert =\tau \}$, then
\begin{align*}
\mathbb E_{\tilde{\delta}}& \mathopen{}\mathclose \bgroup \originalleft[P_{\beta _{n}}\mathopen{}\mathclose \bgroup \originalleft( S_{n}>c_{\alpha }\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \\
& =1-\mathcal{G}_{d}\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha };\tau ^{2}\aftergroup \egroup \originalright) +\mathcal{G}
_{d+2}^{\prime }(c_{\alpha };\tau ^{2})\mathopen{}\mathclose \bgroup \originalleft\{ nh^{2r}\mathbb E(B)'\mathbb E(B)
-h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)\aftergroup \egroup \originalright]\aftergroup \egroup \originalright\}  \\
& \quad -\mathopen{}\mathclose \bgroup \originalleft[ \mathcal{G}_{d+4}^{\prime }(c_{\alpha };\tau ^{2})-\mathcal{G
}_{d+2}^{\prime }(c_{\alpha };\tau ^{2})\aftergroup \egroup \originalright] \frac{\tau ^{2}}{d}h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)\aftergroup \egroup \originalright]
+O\mathopen{}\mathclose \bgroup \originalleft(h^{2}+n^{-1/2}\aftergroup \egroup \originalright)
\end{align*}
where $\mathbb E_{\tilde{\delta}}$ takes the average uniformly over the sphere $
\mathcal{S}_{d}(\tau )$.
\end{theorem}

When $\delta=0$, which implies $\tau=0$,
the expansion in Theorem \ref{thm:power} reduces to that in Theorem \ref
{thm:inf}.

When $h=h_{\text{SEE}}^{\ast }$, it follows from Theorem \ref{thm:inf} that
\begin{align*}
P_{\beta _{0}}\mathopen{}\mathclose \bgroup \originalleft( S_{n}>c_{\alpha }\aftergroup \egroup \originalright) &= 1-\mathcal{G}_{d}\mathopen{}\mathclose \bgroup \originalleft(
c_{\alpha }\aftergroup \egroup \originalright) -C^{+}\mathcal{G}_{d+2}^{\prime }(c_{\alpha })h_{\text{SEE
}}^{\ast }+o(h_{\text{SEE}}^{\ast }) \\
&= \alpha -C^{+}\mathcal{G}_{d+2}^{\prime }(c_{\alpha })h_{\text{SEE}}^{\ast
}+o(h_{\text{SEE}}^{\ast }).
\end{align*}
To remove the error in rejection probability of order $h_{\text{SEE}}^{\ast
} $, we make a correction to the critical value $c_{\alpha }$. Let $
c_{\alpha }^{\ast }$ be a high-order corrected critical value such that $
P_{\beta_{0}}\mathopen{}\mathclose \bgroup \originalleft( S_{n}>c_{\alpha }^{\ast }\aftergroup \egroup \originalright) =\alpha +o(h_{\text{SEE}
}^{\ast})$. Simple calculation shows that
\begin{equation*}
c_{\alpha }^{\ast }=c_{\alpha }-\frac{\mathcal{G}_{d+2}^{\prime }(c_{\alpha
})}{\mathcal{G}_{d}^{\prime }\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha }\aftergroup \egroup \originalright) }C^{+}h_{\text{SEE}
}^{\ast }
\end{equation*}
meets the requirement.

To approximate the size-adjusted power of the $S_{n}$ test, we use $
c_{\alpha }^{\ast }$ rather than $c_{\alpha }$ because $c_{\alpha }^{\ast }$
leads to a more accurate test in large samples. Using Theorem \ref{thm:power}
, we can prove the following corollary.

\begin{corollary}
\label{cor:power} Let the assumptions in Theorem \ref{thm:power} hold.
Then for $h=h_{\text{SEE}}^{\ast }$,
\begin{equation}
\begin{split}
\mathbb E_{\tilde{\delta}}& \mathopen{}\mathclose \bgroup \originalleft[P_{\beta _{n}}\mathopen{}\mathclose \bgroup \originalleft( S_{n}>c_{\alpha }^{\ast }\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] \\
& =1-\mathcal{G}_{d}\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha };\tau ^{2}\aftergroup \egroup \originalright) +Q_{d}\mathopen{}\mathclose \bgroup \originalleft(
c_{\alpha },\tau ^{2},r\aftergroup \egroup \originalright) \mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)\aftergroup \egroup \originalright]
h_{\text{SEE}}^{\ast }+O\mathopen{}\mathclose \bgroup \originalleft( h_{\text{SEE}}^{\ast 2}+n^{-1/2}\aftergroup \egroup \originalright) ,
\end{split}
\label{eqn:asym-local-power}
\end{equation}
where
\begin{align*}
Q_{d}\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha },\tau ^{2},r\aftergroup \egroup \originalright) & =\mathopen{}\mathclose \bgroup \originalleft( 1-\frac{1}{2r}\aftergroup \egroup \originalright)
\mathopen{}\mathclose \bgroup \originalleft[ \mathcal{G}_{d}^{\prime }\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha };\tau ^{2}\aftergroup \egroup \originalright) \frac{
\mathcal{G}_{d+2}^{\prime }(c_{\alpha })}{\mathcal{G}_{d}^{\prime }\mathopen{}\mathclose \bgroup \originalleft(
c_{\alpha }\aftergroup \egroup \originalright) }-\mathcal{G}_{d+2}^{\prime }(c_{\alpha };\tau ^{2})
\aftergroup \egroup \originalright] \\
& \quad -\frac{1}{d}\mathopen{}\mathclose \bgroup \originalleft[ \mathcal{G}_{d+4}^{\prime }(c_{\alpha };\tau
^{2})-\mathcal{G}_{d+2}^{\prime }(c_{\alpha };\tau ^{2})\aftergroup \egroup \originalright] \tau ^{2}.
\end{align*}
\end{corollary}

In the asymptotic expansion of the local power function in
\eqref{eqn:asym-local-power}, $1-\mathcal{G}_{d}\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha };\tau
^{2}\aftergroup \egroup \originalright) $ is the usual first-order power of a standard chi-square test.
The next term of order $O(h_{\text{SEE}}^{\ast })$ captures the effect of
smoothing the estimating equations. To sign this effect, we plot the
function $Q_{d}\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha },\tau ^{2},r\aftergroup \egroup \originalright) $ against $\tau ^{2}$
for $r=2$, $\alpha =10\%$, and different values of $d$ in Figure \ref
{fig:power}. Figures for other values of $r$ and $\alpha $ are qualitatively
similar. The range of $\tau ^{2}$ considered in Figure \ref{fig:power} is
relevant as the first-order local asymptotic power, i.e.,\ $1-\mathcal{G}
_{d}\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha };\tau ^{2}\aftergroup \egroup \originalright) $, increases from $10\%$ to about $
94\%$, $96\%$, $97\%$, and $99\%$, respectively for $d=1,2,3,4$. It is clear
from this figure that $Q_{d}\mathopen{}\mathclose \bgroup \originalleft( c_{\alpha },\tau ^{2},r\aftergroup \egroup \originalright) >0$ for
any $\tau ^{2}>0$. This indicates that smoothing leads to a test with
improved power. The power improvement increases with $r$. The smoother the
conditional PDF of $U$ in a neighborhood of the origin is, the larger the
power improvement is.

\begin{figure}[tbp]
\centering
\includegraphics[width=0.65\textwidth,clip=true,trim=40 190 80 200]{local_asym_power_10_better.pdf}
\caption{Plots of $Q_{d}\mathopen{}\mathclose \bgroup \originalleft( c_{\protect \alpha },\protect \tau
^{2},2\aftergroup \egroup \originalright) $ against $\protect \tau ^{2}$ for different values of $d$ with
$\protect \alpha =10\%$.}
\label{fig:power}
\end{figure}

\section{MSE of the Parameter Estimator\label{sec:est}}

In this section, we examine the approximate MSE of the parameter estimator.
The approximate MSE, being a Nagar-type approximation \citep{Nagar1959}, can
be motivated from the theory of optimal estimating equations, as presented
in \citet{Heyde1997}, for example.

The SEE estimator $\hat{\beta}$ satisfies $m_{n}(\hat{\beta})=0$. In Lemma
\ref{lem:stochastic_expansion_beta} in the appendix, we show that
\begin{equation}
\sqrt{n}\mathopen{}\mathclose \bgroup \originalleft( \hat{\beta}-\beta _{0}\aftergroup \egroup \originalright)
  = -\mathopen{}\mathclose \bgroup \originalleft \{ \mathbb E\mathopen{}\mathclose \bgroup \originalleft[\frac{\partial }{
\partial \beta ^{\prime }}\frac{1}{\sqrt{n}}m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0}\aftergroup \egroup \originalright)
\aftergroup \egroup \originalright]\aftergroup \egroup \originalright\} ^{-1}m_{n}+O_{p}\mathopen{}\mathclose \bgroup \originalleft( \frac{1}{\sqrt{nh}}\aftergroup \egroup \originalright)
\label{eqn:expansion1}
\end{equation}
and
\begin{equation}
\mathbb E\mathopen{}\mathclose \bgroup \originalleft[\frac{\partial }{\partial \beta ^{\prime }}\frac{1}{\sqrt{n}}m_{n}\mathopen{}\mathclose \bgroup \originalleft(
\beta _{0}\aftergroup \egroup \originalright)\aftergroup \egroup \originalright]
  = \mathbb E\mathopen{}\mathclose \bgroup \originalleft[ Z_{j}X_{j}^{\prime }f_{U|Z,X}(0\mid Z_{j},X_{j})
\aftergroup \egroup \originalright] +O\mathopen{}\mathclose \bgroup \originalleft(h^{r}\aftergroup \egroup \originalright).  \label{eqn:expansion2}
\end{equation}
Consequently, the approximate MSE (AMSE) of $\sqrt{n}\mathopen{}\mathclose \bgroup \originalleft( \hat{\beta}
-\beta _{0}\aftergroup \egroup \originalright) $ is\footnote{
Here we follow a common practice in the estimation of nonparametric and
nonlinear models and define the AMSE to be the MSE of $\sqrt{n}\mathopen{}\mathclose \bgroup \originalleft( \hat{
\beta}-\beta _{0}\aftergroup \egroup \originalright) $ after dropping some smaller-order terms. So the
asymptotic MSE we define here is a Nagar-type approximate MSE. See
\citet{Nagar1959}.}
\begin{align*}
\mathrm{AMSE}_{\beta }
  &= \mathopen{}\mathclose \bgroup \originalleft \{ \mathbb E\mathopen{}\mathclose \bgroup \originalleft[\frac{\partial }{\partial \beta ^{\prime
}}\frac{1}{\sqrt{n}}m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright\} ^{-1}
\mathbb E\mathopen{}\mathclose \bgroup \originalleft(m_n m_n'\aftergroup \egroup \originalright) \mathopen{}\mathclose \bgroup \originalleft\{ \mathbb E\mathopen{}\mathclose \bgroup \originalleft[\frac{\partial }{\partial \beta
^{\prime }}\frac{1}{\sqrt{n}}m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]\aftergroup \egroup \originalright\}
^{-1\prime } \\
&= \Sigma _{ZX}^{-1}V\Sigma _{XZ}^{-1}+\Sigma _{ZX}^{-1}V^{1/2}\mathopen{}\mathclose \bgroup \originalleft[
nh^{2r} \mathbb E(B)\mathbb E(B') -h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \mathopen{}\mathclose \bgroup \originalleft( V^{1/2}\aftergroup \egroup \originalright) ^{\prime }\Sigma _{XZ}^{-1} \\
&\quad+O\mathopen{}\mathclose \bgroup \originalleft(h^r\aftergroup \egroup \originalright)+o\mathopen{}\mathclose \bgroup \originalleft(h+nh^{2r}\aftergroup \egroup \originalright),
\end{align*}
where
\begin{equation*}
\Sigma _{ZX} = \mathbb E\mathopen{}\mathclose \bgroup \originalleft[ Z_{j}X_{j}^{\prime }f_{U|Z,X}(0\mid Z_{j},X_{j})\aftergroup \egroup \originalright]
\text{ and }\Sigma _{XZ}=\Sigma _{ZX}^{\prime }.
\end{equation*}

The first term of $\mathrm{AMSE}_{\beta }$ is the asymptotic variance of the
unsmoothed QR estimator. The second term captures the higher-order effect of
smoothing on the AMSE of $\sqrt{n}(\hat{\beta}-\beta _{0})$. When $
nh^{r}\rightarrow \infty $ and $n^{3}h^{4r+1}\rightarrow \infty$, we have $
h^{r}=o\mathopen{}\mathclose \bgroup \originalleft( nh^{2r}\aftergroup \egroup \originalright) $ and $1/\sqrt{nh}=o\mathopen{}\mathclose \bgroup \originalleft( nh^{2r}\aftergroup \egroup \originalright) $, so
the terms of order $O_{p}(1/\sqrt{nh})$ in \eqref{eqn:expansion1} and of
order $O\mathopen{}\mathclose \bgroup \originalleft( h^{r}\aftergroup \egroup \originalright) $ in \eqref{eqn:expansion2} are of smaller order
than the $O(nh^{2r})$ and $O(h)$ terms in the AMSE. If $h\asymp
n^{-1/(2r-1)} $ as before, these rate conditions are satisfied when $r>2$.

\begin{theorem}
\label{thm:est-MSE} Let Assumptions \ref{a:sampling}--\ref{a:G}(i--ii), \ref
{a:beta}, and \ref{a:power_mse} hold. If $nh^{r}\rightarrow \infty $ and $
n^{3}h^{4r+1}\rightarrow \infty $, then the AMSE of $\sqrt{n}(\hat{\beta}
-\beta _{0})$ is
\begin{equation*}
\Sigma _{ZX}^{-1} V^{1/2}
\mathopen{}\mathclose \bgroup \originalleft[ I_{d}+nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]
\mathopen{}\mathclose \bgroup \originalleft(V^{1/2}\aftergroup \egroup \originalright) ^{\prime }\mathopen{}\mathclose \bgroup \originalleft( \Sigma _{ZX}^{\prime }\aftergroup \egroup \originalright)^{-1}
+O\mathopen{}\mathclose \bgroup \originalleft(h^r\aftergroup \egroup \originalright)+o\mathopen{}\mathclose \bgroup \originalleft(h+nh^{2r}\aftergroup \egroup \originalright).
\end{equation*}
\end{theorem}

The optimal $h^{\ast }$ that minimizes the high-order AMSE satisfies
\begin{align*}
&\Sigma _{ZX}^{-1} V^{1/2} \mathopen{}\mathclose \bgroup \originalleft[ n\mathopen{}\mathclose \bgroup \originalleft( h^{\ast }\aftergroup \egroup \originalright) ^{2r}\mathbb E(B)\mathbb E(B') - h^{\ast }\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]
\mathopen{}\mathclose \bgroup \originalleft( V^{1/2} \aftergroup \egroup \originalright)^\prime \mathopen{}\mathclose \bgroup \originalleft( \Sigma _{ZX}^{\prime }\aftergroup \egroup \originalright) ^{-1} \\
&\quad \leq \Sigma _{ZX}^{-1} V^{1/2} \mathopen{}\mathclose \bgroup \originalleft[ nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]
\mathopen{}\mathclose \bgroup \originalleft( V^{1/2} \aftergroup \egroup \originalright)^\prime \mathopen{}\mathclose \bgroup \originalleft( \Sigma _{ZX}^{\prime
}\aftergroup \egroup \originalright) ^{-1}
\end{align*}
in the sense that the difference between the two sides is nonpositive
definite for all $h$. This is equivalent to
\begin{equation*}
n\mathopen{}\mathclose \bgroup \originalleft( h^{\ast }\aftergroup \egroup \originalright) ^{2r}\mathbb E(B)\mathbb E(B') - h^{\ast }\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)
\leq nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) .
\end{equation*}

This choice of $h$ can also be motivated from the theory of optimal
estimating equations. Given the estimating equations $m_{n}=0$, we follow
\citet{Heyde1997} and define the standardized version of $m_{n}$ by
\begin{equation*}
m_{n}^{s}(\beta _{0},h)
  = -\mathbb E\mathopen{}\mathclose \bgroup \originalleft\{\frac{\partial }{\partial \beta ^{\prime }}
m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0}\aftergroup \egroup \originalright) \mathopen{}\mathclose \bgroup \originalleft[ \mathbb E(m_{n}m_{n}^{\prime })\aftergroup \egroup \originalright]
^{-1}m_{n} \aftergroup \egroup \originalright\}.
\end{equation*}
We include $h$ as an argument of $m_{n}^{s}$ to emphasize the dependence of $
m_{n}^{s}$ on $h$. The standardization can be motivated from the following
considerations. On one hand, the estimating equations need to be close to
zero when evaluated at the true parameter value. Thus we want $
\mathbb E(m_{n}m_{n}^{\prime })$ to be as small as possible. On the other hand, we
want $m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta +\delta \beta \aftergroup \egroup \originalright) $ to differ as much as
possible from $m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta \aftergroup \egroup \originalright) $ when $\beta $ is the true value.
That is, we want $\mathbb E\frac{\partial }{\partial \beta ^{\prime }}m_{n}\mathopen{}\mathclose \bgroup \originalleft(
\beta _{0}\aftergroup \egroup \originalright) $ to be as large as possible. To meet these requirements,
we choose $h$ to maximize
\begin{equation*}
\mathbb E\mathopen{}\mathclose \bgroup \originalleft\{m_{n}^{s}(\beta _{0},h)\mathopen{}\mathclose \bgroup \originalleft[ m_{n}^{s}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0},h\aftergroup \egroup \originalright) \aftergroup \egroup \originalright]
^{\prime }\aftergroup \egroup \originalright\}
  = \mathopen{}\mathclose \bgroup \originalleft[ \mathbb E\frac{\partial }{\partial \beta ^{\prime }}m_{n}\mathopen{}\mathclose \bgroup \originalleft(
\beta _{0}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \mathopen{}\mathclose \bgroup \originalleft[ \mathbb E(m_{n}m_{n}^{\prime })\aftergroup \egroup \originalright] ^{-1}\mathopen{}\mathclose \bgroup \originalleft[ \mathbb E
\frac{\partial }{\partial \beta ^{\prime }}m_{n}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0}\aftergroup \egroup \originalright)
\aftergroup \egroup \originalright] ^{\prime }.
\end{equation*}
More specifically, $h^{\ast }$ is optimal if
\begin{equation*}
\mathbb E\mathopen{}\mathclose \bgroup \originalleft\{m_{n}^{s}(\beta _{0},h^{\ast })\mathopen{}\mathclose \bgroup \originalleft[ m_{n}^{s}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0},h^{\ast
}\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] ^{\prime }\aftergroup \egroup \originalright\}
  - \mathbb E\mathopen{}\mathclose \bgroup \originalleft\{m_{n}^{s}(\beta _{0},h)\mathopen{}\mathclose \bgroup \originalleft[
m_{n}^{s}\mathopen{}\mathclose \bgroup \originalleft( \beta _{0},h\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] ^{\prime }\aftergroup \egroup \originalright\}
\end{equation*}
is nonnegative definite for all $h\in \mathbb{R}^{+}$. But $
\mathbb E\mathopen{}\mathclose \bgroup \originalleft[m_{n}^{s}\mathopen{}\mathclose \bgroup \originalleft( m_{n}^{s}\aftergroup \egroup \originalright) ^{\prime }\aftergroup \egroup \originalright]=\mathopen{}\mathclose \bgroup \originalleft( \mathrm{AMSE}_{\beta
}\aftergroup \egroup \originalright) ^{-1}$, so maximizing $\mathbb E\mathopen{}\mathclose \bgroup \originalleft[m_{n}^{s}\mathopen{}\mathclose \bgroup \originalleft( m_{n}^{s}\aftergroup \egroup \originalright) ^{\prime
}\aftergroup \egroup \originalright]$ is equivalent to minimizing $\mathrm{AMSE}_{\beta }$.

The question is whether such an optimal $h$ exists. If it does, then the
optimal $h^{\ast }$ satisfies
\begin{equation}  \label{AMSE_obj}
h^{\ast } = \mathop{\rm arg\,min}_{h} u^{\prime }\mathopen{}\mathclose \bgroup \originalleft[ nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] u
\end{equation}
for all $u\in \mathbb{R}^{d}$, by the definition of nonpositive definite
plus the fact that the above yields a unique minimizer for any $u$. Using
unit vectors $e_1=(1,0,\ldots,0)$, $e_2=(0,1,0,\ldots,0)$, etc.,\ for $u$,
and noting that $\mathrm{tr}(A)=e_1^{\prime
}Ae_1+\cdots+e_d^{\prime }Ae_d$ for $d\times d$ matrix $A$, this implies
that
\begin{align*}
h^{\ast } &= \mathop{\rm arg\,min}_{h}\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[ nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \\
&= \mathop{\rm arg\,min}_{h}\mathopen{}\mathclose \bgroup \originalleft\{ nh^{2r}\mathbb E(B)'\mathbb E(B) - h\mathrm{tr}\mathopen{}\mathclose \bgroup \originalleft[\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] \aftergroup \egroup \originalright\} .
\end{align*}
In view of \eqref{eqn:def_h_SEE}, $h_\text{SEE}^*=h^*$ if $h^*$ exists.
Unfortunately, it is easy to show that no single $h$ can minimize the
objective function in \eqref{AMSE_obj} for all $u\in \mathbb{R}^{d}$. Thus,
we have to redefine the optimality with respect to the direction of $u$. The
direction depends on which linear combination of $\beta $ is the focus of
interest, as $u^{\prime }\mathopen{}\mathclose \bgroup \originalleft[ nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] u$ is the high-order AMSE of
$c^{\prime }\sqrt{n}(\hat{\beta}-\beta _{0})$ for $c=\Sigma_{XZ} \mathopen{}\mathclose \bgroup \originalleft(
V^{-1/2}\aftergroup \egroup \originalright)^{\prime} u$.

Suppose we are interested in only one linear combination. Let $h_{c}^{\ast }$
be the optimal $h$ that minimizes the high-order AMSE of $c^{\prime }\sqrt{n}
(\hat{\beta}-\beta _{0})$. Then
\begin{equation*}
h_{c}^{\ast }=\mathopen{}\mathclose \bgroup \originalleft( \frac{u^{\prime }\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA^{\prime }\aftergroup \egroup \originalright) u}{
u^{\prime }\mathbb E(B)\mathbb E(B')u}\frac{1}{2nr}
\aftergroup \egroup \originalright) ^{\frac{1}{2r-1}}
\end{equation*}
for $u=\mathopen{}\mathclose \bgroup \originalleft( V^{1/2}\aftergroup \egroup \originalright) ^{\prime } \Sigma_{XZ}^{-1} c$. Some algebra
shows that
\begin{equation*}
h_{c}^{\ast }\geq \mathopen{}\mathclose \bgroup \originalleft( \frac{1}{\mathbb E(B)'\mathopen{}\mathclose \bgroup \originalleft[
\mathbb E\mathopen{}\mathclose \bgroup \originalleft(AA'\aftergroup \egroup \originalright)\aftergroup \egroup \originalright]^{-1}\mathbb E(B)}\frac{1}{2nr}\aftergroup \egroup \originalright) ^{\frac{1}{2r-1}}>0.
\end{equation*}
So although $h_{c}^{\ast }$ depends on $c$ via $u$, it is nevertheless
greater than zero.

Now suppose without loss of generality we are interested in $d$ directions $
\mathopen{}\mathclose \bgroup \originalleft( c_{1},\ldots ,c_{d}\aftergroup \egroup \originalright) $ jointly where $c_{i}\in \mathbb{R}^{d}$.
In this case, it is reasonable to choose $h_{c_{1},\ldots ,c_{d}}^{\ast }$
to minimize the sum of direction-wise AMSEs, i.e.,
\begin{equation*}
h_{c_{1},\ldots ,c_{d}}^{\ast }=\mathop{\rm arg\,min}_{h}
\sum_{i=1}^{d}u_{i}^{\prime }\mathopen{}\mathclose \bgroup \originalleft[ nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) \aftergroup \egroup \originalright] u_{i},
\end{equation*}
where $u_{i}=\mathopen{}\mathclose \bgroup \originalleft( V^{1/2}\aftergroup \egroup \originalright) ^{\prime }\Sigma _{XZ}^{-1}c_{i}$. It is
easy to show that
\begin{equation*}
h_{c_{1},\ldots ,c_{d}}^{\ast }
  = \mathopen{}\mathclose \bgroup \originalleft[ \frac{\sum_{i=1}^{d}u_{i}^{\prime
}\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) u_{i}}{\sum_{i=1}^{d}u_{i}^{\prime }\mathbb E(B)\mathbb E(B')u_{i}}\frac{1}{2nr}\aftergroup \egroup \originalright]^{\frac{1}{
2r-1}}.
\end{equation*}

As an example, consider $u_{i}=e_{i}=\mathopen{}\mathclose \bgroup \originalleft( 0,\ldots ,1,\ldots ,0\aftergroup \egroup \originalright) $,
the $i$th unit vector in $\mathbb{R}^{d}$. Correspondingly
\begin{equation*}
\mathopen{}\mathclose \bgroup \originalleft( \tilde{c}_{1},\ldots,\tilde{c}_{d}\aftergroup \egroup \originalright) = \Sigma_{XZ} \mathopen{}\mathclose \bgroup \originalleft(
V^{-1/2}\aftergroup \egroup \originalright)^{\prime } \mathopen{}\mathclose \bgroup \originalleft( e_{1},\ldots,e_{d}\aftergroup \egroup \originalright) .
\end{equation*}
It is clear that
\begin{equation*}
h_{\tilde{c}_{1},\ldots ,\tilde{c}_{d}}^{\ast }=h_{\text{SEE}}^{\ast }=h_{
\text{CPE}}^{\ast } ,
\end{equation*}
so all three selections coincide with each other. A special case of interest
is when $Z=X$, non-constant regressors are pairwise independent and
normalized to mean zero and variance one, and $U
\mathpalette{\protect
\independenT}{\perp} X$. Then $u_{i}=c_{i}=e_{i}$ and the $d$ linear
combinations reduce to the individual elements of $\beta $.

The above example illustrates the relationship between $h_{c_{1},\ldots
,c_{d}}^{\ast }$ and $h_{\text{SEE}}^{\ast }$. While $h_{c_{1},\ldots
,c_{d}}^{\ast }$ is tailored toward the flexible linear combinations $
\mathopen{}\mathclose \bgroup \originalleft(c_{1},\ldots ,c_{d}\aftergroup \egroup \originalright)$ of the parameter vector, $h_{\text{SEE}
}^{\ast }$ is tailored toward the fixed $\mathopen{}\mathclose \bgroup \originalleft( \tilde{c}_{1},\ldots ,\tilde{
c}_{d}\aftergroup \egroup \originalright) $. While $h_{c_{1},\ldots ,c_{d}}^{\ast }$ and $h_{\text{SEE}
}^{\ast }$ are of the same order of magnitude, in general there is no
analytic relationship between $h_{c_{1},\ldots ,c_{d}}^{\ast }$ and $h_{
\text{SEE}}^{\ast }$.

To shed further light on the relationship between $h_{c_{1},\ldots
,c_{d}}^{\ast }$ and $h_{\text{SEE}}^{\ast }$, let $\mathopen{}\mathclose \bgroup \originalleft \{ \lambda
_{k},k=1,\ldots ,d\aftergroup \egroup \originalright \} $ be the eigenvalues of $nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)$ with the
corresponding orthonormal eigenvectors $\mathopen{}\mathclose \bgroup \originalleft \{ \ell _{k},k=1,\ldots
,d\aftergroup \egroup \originalright \} $. Then we have $nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright) =\sum_{k=1}^{d}\lambda _{k}\ell
_{k}\ell _{k}^{\prime }$ and $u_{i}=\sum_{j=1}^{d}u_{ij}\ell _{j}$ for $
u_{ij}=u_{i}^{\prime }\ell _{j}$. Using these representations, the objective
function underlying $h_{c_{1},\ldots ,c_{d}}^{\ast }$ becomes
\begin{align*}
& \sum_{i=1}^{d}u_{i}^{\prime }\mathopen{}\mathclose \bgroup \originalleft[ nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)\aftergroup \egroup \originalright] u_{i} \\
& =\sum_{i=1}^{d}\mathopen{}\mathclose \bgroup \originalleft( \sum_{j=1}^{d}u_{ij}\ell _{j}^{\prime }\aftergroup \egroup \originalright)
\mathopen{}\mathclose \bgroup \originalleft( \sum_{k=1}^{d}\lambda _{k}\ell _{k}\ell _{k}^{\prime }\aftergroup \egroup \originalright) \mathopen{}\mathclose \bgroup \originalleft(
\sum_{\tilde{j}=1}^{d}u_{i\tilde{j}}\ell _{\tilde{j}}\aftergroup \egroup \originalright) \\
& =\sum_{j=1}^{d}\mathopen{}\mathclose \bgroup \originalleft( \sum_{i=1}^{d}u_{ij}^{2}\aftergroup \egroup \originalright) \lambda _{j}.
\end{align*}
That is, $h_{c_{1},\ldots ,c_{d}}^{\ast }$ minimizes a weighted sum of the
eigenvalues of $nh^{2r}\mathbb E(B)\mathbb E(B') - h\mathbb E\mathopen{}\mathclose \bgroup \originalleft( AA'\aftergroup \egroup \originalright)$ with weights depending on $c_{1},\ldots ,c_{d}$. By
definition, $h_{\text{SEE}}^{\ast }$ minimizes the simple unweighted sum of
the eigenvalues, viz.\ $\sum_{j=1}^{d}\lambda _{j}$. While $h_{\text{SEE}
}^{\ast }$ may not be ideal if we know the linear combination(s) of
interest, it is a reasonable choice otherwise.

In empirical applications, we can estimate $h_{c_{1},\ldots ,c_{d}}^{\ast }$
using a parametric plug-in approach similar to our plug-in implementation of
$h_{\text{SEE}}^{\ast }$. If we want to be agnostic about the directional
vectors $c_{1},\ldots ,c_{d}$, we can simply use $h_{\text{SEE}}^{\ast }$.













\section{Empirical example: JTPA\label{sec:emp}}

We revisit the IV-QR analysis of Job Training Partnership Act (JTPA) data in
\citet{AbadieEtAl2002}, specifically their Table III.\footnote{
Their data and Matlab code for replication are helpfully provided online in
the Angrist Data Archive,
\url{http://economics.mit.edu/faculty/angrist/data1/data/abangim02}.} They
use 30-month earnings as the outcome, randomized offer of JTPA services as
the instrument, and actual enrollment for services as the endogenous
treatment variable. Of those offered services, only around 60 percent
accepted, so self-selection into treatment is likely. Section 4 of
\citet{AbadieEtAl2002} provides much more background and descriptive
statistics.

We compare estimates from a variety of methods.\footnote{Code and data for replication is available on the first author's website.}
 ``AAI'' is the original paper's
estimator. AAI restricts $X$ to have finite support (see condition (iii) in
their Theorem 3.1), which is why all the regressors in their example are
binary. Our fully automated plug-in estimator is ``SEE ($\hat{h}$).''
``CH'' is \citet{ChernozhukovHansen2006}. Method ``tiny $h$'' uses $h=400$ (compared with our plug-in values on the order of $10\,000$), while ``huge $h$'' uses $h=5\times 10^{6}$. 2SLS is the usual (mean) two-stage least squares estimator, put in the $q=0.5$ column only for convenience of comparison.

\begin{table}[htbp]
\centering
\caption{\label{tab:emp}IV-QR estimates of coefficients for certain regressors as in Table III of \citet{AbadieEtAl2002} for adult men.}
\begin{tabular}[c]{ccS[table-format=9.0]S[table-format=8.0]S[table-format=6.0]S[table-format=7.0]S[table-format=8.0]}
\hline\hline
 &  &  \multicolumn{5}{c}{Quantile} \\
 \cline{3-7}
Regressor    &     Method     &   \multicolumn{1}{c}{0.15}  &   \multicolumn{1}{c}{0.25}  &   \multicolumn{1}{c}{0.50}  &   \multicolumn{1}{c}{0.75}  &   \multicolumn{1}{c}{0.85}  \\
\hline
Training     & AAI            &      121 &      702 &  1544 &  3131 &  3378 \\
Training     & SEE ($\hat h$) &       57 &      381 &  1080 &  2630 &  2744 \\
Training     & CH             &     -125 &      341 &      385 &  2557 &  3137 \\
Training     & tiny $h$       &     -129 &      500 &      381 &  2760 &  3114 \\
Training     & huge $h$       &  1579 &  1584 &  1593 &  1602 &  1607 \\
Training     & 2SLS           &          &          &  1593 &          &          \\
HS or GED    & AAI            &      714 &  1752 &  4024 &  5392 &  5954 \\
HS or GED    & SEE ($\hat h$) &      812 &  1498 &  3598 &  6183 &  6753 \\
HS or GED    & CH             &      482 &  1396 &  3761 &  6127 &  6078 \\
HS or GED    & tiny $h$       &      463 &  1393 &  3767 &  6144 &  6085 \\
HS or GED    & huge $h$       &  4054 &  4062 &  4075 &  4088 &  4096 \\
HS or GED    & 2SLS           &          &          &  4075 &          &          \\
Black        & AAI            &     -171 &     -377 & -2656 & -4182 & -3523 \\
Black        & SEE ($\hat h$) &     -202 &     -546 & -1954 & -3273 & -3653 \\
Black        & CH             &      -38 &     -109 & -2083 & -3233 & -2934 \\
Black        & tiny $h$       &      -18 &     -139 & -2121 & -3337 & -2884 \\
Black        & huge $h$       & -2336 & -2341 & -2349 & -2357 & -2362 \\
Black        & 2SLS           &          &          & -2349 &          &          \\
Married      & AAI            &  1564 &  3190 &  7683 &  9509 & 10185 \\
Married      & SEE ($\hat h$) &  1132 &  2357 &  7163 & 10174 & 10431 \\
Married      & CH             &      504 &  2396 &  7722 & 10463 & 10484 \\
Married      & tiny $h$       &      504 &  2358 &  7696 & 10465 & 10439 \\
Married      & huge $h$       &  6611 &  6624 &  6647 &  6670 &  6683 \\
Married      & 2SLS           &          &          &  6647 &          &          \\
Constant     & AAI            &     -134 &  1049 &  7689 & 14901 & 22412 \\
Constant     & SEE ($\hat h$) &      -88 &  1268 &  7092 & 15480 & 22708 \\
Constant     & CH             &      242 &  1033 &  7516 & 14352 & 22518 \\
Constant     & tiny $h$       &      294 &  1000 &  7493 & 14434 & 22559 \\
Constant     & huge $h$       & -1157554    & -784046 &  10641 & 805329 & 1178836 \\
Constant     & 2SLS           &          &          & 10641 &          &          \\
\hline
\end{tabular}
\end{table}

Table \ref{tab:emp} shows results from the sample of $5102$ adult men, for a
subset of the regressors used in the model. Not shown in the table are
coefficient estimates for dummies for Hispanic, working less than 13 weeks
in the past year, five age groups, originally recommended service strategy,
and whether earnings were from the second follow-up survey. CH is very close
to ``tiny $h$''; that is,
simply using the smallest possible $h$ with SEE provides a good
approximation of the unsmoothed estimator in this case. Demonstrating our theoretical results in Section \ref{sec:see-comp-IV}, ``huge $h$'' is very close to
2SLS for everything except the constant term for $q\neq 0.5$. The IVQR-SEE estimator using
our plug-in bandwidth has some economically significant differences with the
unsmoothed estimator. Focusing on the treatment variable (``Training''), the unsmoothed median effect estimate is below
$400$ (dollars), whereas SEE$(\hat{h})$ yields $1080$, both of
which are smaller than AAI's $1544$ (AAI is the most positive at all
quantiles). For the $0.15$-quantile effect, the unsmoothed estimates are
actually slightly negative, while SEE$(\hat{h})$ and AAI are
slightly positive.
For $q=0.85$, though, the SEE$(\hat{h})$ estimate is smaller than
the unsmoothed one, and the two are quite similar for $q=0.25$ and $q=0.75$;
there is no systematic ordering.

Computationally, our code takes only one second total to calculate the plug-in bandwidths and coefficient estimates at all five quantiles.  Using the fixed $h=400$ or $h=5\times10^6$, computation is immediate.

\begin{table}[htbp]
\centering
\caption{\label{tab:emp2}IV-QR estimates similar to Table \ref{tab:emp}, but replacing age dummies with a quartic polynomial in age and adding baseline measures of weekly hours worked and wage.}
\begin{tabular}[c]{ccS[table-format=4.0]S[table-format=4.0]S[table-format=4.0]S[table-format=4.0]S[table-format=4.0]}
\hline\hline
 &  &  \multicolumn{5}{c}{Quantile} \\
 \cline{3-7}
Regressor    &     Method     &   \multicolumn{1}{c}{0.15}  &   \multicolumn{1}{c}{0.25}  &   \multicolumn{1}{c}{0.50}  &   \multicolumn{1}{c}{0.75}  &   \multicolumn{1}{c}{0.85}  \\
\hline\rule{0pt}{12pt}
Training     & SEE ($\hat h$) &       74 &      398 &  1045 &  2748 &  2974 \\
Training     & CH             &      -20 &      451 &      911 &  2577 &  3415 \\
Training     & tiny $h$       &      -50 &      416 &      721 &  2706 &  3555 \\
Training     & huge $h$       &  1568 &  1573 &  1582 &  1590 &  1595 \\
Training     & 2SLS           &          &          &  1582 &          &          \\
\hline
\end{tabular}
\end{table}

Table \ref{tab:emp2} shows estimates of the endogenous coefficient when various ``continuous'' control variables are added, specifically a quartic polynomial in age (replacing the age range dummies), baseline weekly hours worked, and baseline hourly wage.\footnote{Additional JTPA data downloaded from the W.E.\ Upjohn Institute at \url{http://upjohn.org/services/resources/employment-research-data-center/national-jtpa-study}; variables are named \texttt{age}, \texttt{bfhrswrk}, and \texttt{bfwage} in file \texttt{expbif.dta}.}  The estimates do not change much; the biggest difference is for the unsmoothed estimate at the median.  Our code again computes the plug-in bandwidth and SEE coefficient estimates at all five quantiles in one second.  Using the small $h=400$ bandwidth now takes nine seconds total (more iterations of \texttt{fsolve} are needed); $h=5\times10^6$ still computes almost immediately.










\section{Simulations\label{sec:sim}}

For our simulation study,\footnote{
Code to replicate our simulations is available on the first author's
website.} we use $G(u)$ given in \eqref{eqn:G_fun} as in
\citet{Horowitz1998} and \citet{Whang2006}. This satisfies Assumption \ref
{a:G} with $r=4$. Using (the integral of) an Epanechnikov
kernel with $r=2$ also worked well in the cases we consider here, though
never better than $r=4$. Our error distributions always have at
least four derivatives, so $r=4$ working somewhat better is expected.
Selection of optimal $r$ and $G(\cdot)$, and the quantitative impact
thereof, remain open questions.

We implement a plug-in version ($\hat h$) of the infeasible $h^{\ast }\equiv h_{\text{
SEE}}^{\ast }$. We make the plug-in assumption $U
\mathpalette{\protect
\independenT}{\perp} Z$ and parameterize the distribution of $U$. Our
current method, which has proven quite accurate and stable, fits the
residuals from an initial $h=(2nr)^{-1/(2r-1)}$ IV-QR to Gaussian, $t$,
gamma, and generalized extreme value distributions via maximum likelihood.
With the distribution parameter estimates, $f_{U}(0)$ and $f_{U}^{(r-1)}(0)$
can be computed and plugged in to calculate $\hat{h}$.  With larger $n$, a nonparametric kernel estimator may perform better, but nonparametric estimation of $f_U^{(r-1)}(0)$ will have high variance in smaller samples.
Viewing the unsmoothed estimator as a reference point,
potential regret (of using $\hat h$ instead of $h=0$) is largest when $\hat h$ is too large,
so we separately calculate $\hat{h}$
for each of the four distributions and take the smallest. Note that this
particular plug-in approach works well even under heteroskedasticity and/or
misspecification of the error distribution: DGPs 3.1--3.6 in Section \ref{sec:sim-additional} have error
distributions other than these four, and DGPs 1.3, 2.2, 3.3--3.6 are
heteroskedastic, as are the JTPA-based simulations. For the infeasible $h^{\ast }$, if the PDF derivative in
the denominator is zero, it is replaced by $0.01$ to avoid $h^{\ast }=\infty
$.

For the unsmoothed IV-QR estimator, we use code based on
\citet{ChernozhukovHansen2006} from the latter author's website.
We use the option to let their code determine the
grid of possible endogenous coefficient values from the data.
This code in turn uses the interior point method in \texttt{rq.m} (developed
by Roger Koenker, Daniel Morillo, and Paul Eilers) to solve exogenous QR
linear programs.


\subsection{JTPA-based simulations\label{sec:sim-JTPA}}

We use two DGPs based on the JTPA data examined in Section \ref
{sec:emp}. The first DGP corresponds to the variables used in the original
analysis in \citet{AbadieEtAl2002}. For individual $i$, let $Y_{i}$ be the
scalar outcome (30-month earnings), $X_{i}$ be the vector of exogenous
regressors, $D_{i}$ be the scalar endogenous training dummy, $Z_{i}$ be the
scalar instrument of randomized training offer, and $U_{i}\sim \text{Unif}
(0,1)$ be a scalar unobservable term. We draw $X_{i}$ from the joint
distribution estimated from the JTPA data. We randomize $Z_{i}=1$ with
probability $0.67$ and zero otherwise. If $Z_{i}=0$, then we set the
endogenous training dummy $D_{i}=0$ (ignoring that in reality, a few percent
still got services). If $Z_{i}=1$, we set $D_{i}=1$ with a probability
increasing in $U_{i}$. Specifically, $P(D_{i}=1\mid Z_{i}=1,U_{i}=u)=\min
\{1,u/0.75\}$, which roughly matches the $P(D_{i}=1\mid Z_{i}=1)=0.62$ in
the data. This corresponds to a high degree of self-selection into treatment
(and thus endogeneity). Then, $Y_{i}=X_{i}\beta _{X}+D_{i}\beta
_{D}(U_{i})+G^{-1}(U_{i})$, where $\beta _{X}$ is the IVQR-SEE $\hat{\beta}
_{X}$ from the JTPA data (rounded to the nearest $500$), the function $\beta
_{D}(U_{i})=2000U_{i}$ matches $\hat{\beta}_{D}(0.5)$ and the increasing
pattern of other $\hat{\beta}_{D}(q)$, and $G^{-1}(\cdot )$ is a recentered
gamma distribution quantile function with parameters estimated to match the
distribution of residuals from the IVQR-SEE estimate with JTPA data.
In each of $1000$ simulation replications, we generate $n=5102$ iid observations.

For the second DGP, we add a second endogenous regressor (and instrument)
and four exogenous regressors, all with normal distributions. Including the
intercept and two endogenous regressors, there are $20$ regressors. The
second instrument is $Z_{2i}\overset{iid}{\sim}N(0,1)$, and the second
endogenous regressor is $D_{2i}=0.8Z_{2i}+0.2\Phi^{-1}(U_i)$. The
coefficient on $D_{2i}$ is $1000$ at all quantiles. The new exogenous
regressors are all standard normal and have coefficients of $500$ at all
quantiles. To make the asymptotic bias of 2SLS relatively more important,
the sample size is increased to $n=50\,000$.

\begin{table}[htbp]
\centering
\caption{\label{tab:sim-JTPA1}Simulation results for endogenous coefficient estimators with first JTPA-based DGP.  ``Robust MSE'' is squared median-bias plus the square of the interquartile range divided by $1.349$, $\textrm{Bias}_{\textrm{median}}^{2}+(\textrm{IQR}/1.349)^2$; it is shown in units of $10^5$, so $7.8$ means $7.8\times10^5$, for example.  ``Unsmoothed'' is the estimator from \citet{ChernozhukovHansen2006}.}
\begin{tabular}[c]{cS[table-format=2.1]S[table-format=2.1]S[table-format=2.1]cS[table-format=5.1]S[table-format=4.1]S[table-format=5.1]}
\hline\hline\rule{0pt}{12pt}
 & \multicolumn{3}{c}{Robust MSE / $10^5$} & & \multicolumn{3}{c}{Median Bias} \\
  \cline{2-4}\cline{6-8} \rule{0pt}{14pt}
$q$ & \multicolumn{1}{c}{Unsmoothed} & \multicolumn{1}{c}{SEE ($\hat h$)} & \multicolumn{1}{c}{2SLS} & &
      \multicolumn{1}{c}{Unsmoothed} & \multicolumn{1}{c}{SEE ($\hat h$)} & \multicolumn{1}{c}{2SLS} \\
\hline
$0.15$ &  78.2 &  43.4 &  18.2  &&  -237.6 &   8.7 & 1040.6 \\%262.6 &  54.2 &  17.4  &  -1199.8 & -272.5 & 1001.4
$0.25$ &  30.5 &  18.9 &  14.4  &&  -122.2 &  16.3 & 840.6 \\% 27.2 &  18.8 &  13.8  &  -125.5 &   8.7 & 801.4
$0.50$ &   9.7 &   7.8 &   8.5  &&   24.1 &  -8.5 & 340.6 \\%  9.2 &   7.9 &   8.3  &  -18.0 & -11.7 & 301.4
$0.75$ &   7.5 &   7.7 &   7.6  &&   -5.8 & -48.1 & -159.4 \\%  8.5 &   7.3 &   7.8  &  -19.1 & -28.4 & -198.6
$0.85$ &  11.7 &   9.4 &   8.6  &&   49.9 & -17.7 & -359.4 \\% 12.2 &   9.7 &   9.0  &   10.5 & -11.5 & -398.6
\hline
\end{tabular}
\end{table}

Table \ref{tab:sim-JTPA1} shows results for the first JTPA-based DGP, for three estimators of the endogenous coefficient: \citet{ChernozhukovHansen2006}; SEE with our data-dependent $\hat h$; and 2SLS.  The first and third can be viewed as limits of IVQR-SEE estimators as $h\to0$ and $h\to\infty$, respectively.  We show median bias and ``robust MSE,'' which is squared median bias plus the square of the interquartile range divided by $1.349$, $\textrm{Bias}_{\textrm{median}}^{2}+(\textrm{IQR}/1.349)^2$.  We report these ``robust'' versions of bias and MSE since the (mean) IV estimator does not even possess a first moment in finite samples \citep{Kinal1980}.  We are unaware of an analogous result for IV-QR but remain wary of presenting bias and MSE results for IV-QR, too, especially since the IV estimator is the limit of the SEE IV-QR estimator as $h\to\infty$.  At all quantiles, for all methods, the robust MSE is dominated by the IQR rather than bias.  Consequently, even though the 2SLS median bias is quite large for $q=0.15$, it has less than half the robust MSE of SEE($\hat h$), which in turn has half the robust MSE of the unsmoothed estimator.  With only a couple exceptions, this is the ordering among the three methods' robust MSE at all quantiles.  Although the much larger bias of 2SLS than that of SEE($\hat h$) or the unsmoothed estimator is expected, the smaller median bias of SEE($\hat h$) than that of the unsmoothed estimator is surprising.  However, the differences are not big, and they may be partly due to the much larger variance of the unsmoothed estimator inflating the simulation error in the simulated median bias, especially for $q=0.15$.  The bigger difference is the reduction in variance from smoothing.

\begin{table}[htbp]
\centering
\caption{\label{tab:sim-JTPA2}Simulation results for endogenous coefficient estimators with second JTPA-based DGP.}
\begin{tabular}[c]{cS[table-format=6.0]S[table-format=6.0]S[table-format=7.0]cS[table-format=4.1]S[table-format=4.1]S[table-format=5.1]}
\hline\hline\rule{0pt}{12pt}
 & \multicolumn{3}{c}{Robust MSE} & & \multicolumn{3}{c}{Median Bias} \\
  \cline{2-4}\cline{6-8} \rule{0pt}{14pt}
$q$ & \multicolumn{1}{c}{$h=400$} & \multicolumn{1}{c}{SEE ($\hat h$)} & \multicolumn{1}{c}{2SLS} & &
      \multicolumn{1}{c}{$h=400$} & \multicolumn{1}{c}{SEE ($\hat h$)} & \multicolumn{1}{c}{2SLS} \\
\hline
\multicolumn{8}{c}{\textit{Estimators of binary endogenous regressor's coefficient}} \\
$0.15$ &   780624 &   539542 &  1071377  &&  -35.7 &  10.6 & 993.6 \\%  723856 &   524639 &  1060738  &  -69.5 & -24.9 & 989.3
$0.25$ &   302562 &   227508 &   713952  &&  -18.5 &  17.9 & 793.6 \\%  273749 &   219483 &   705021  &  -15.6 &  17.4 & 789.3
$0.50$ &   101433 &    96350 &   170390  &&  -14.9 & -22.0 & 293.6 \\%   98549 &    89963 &   165726  &  -16.2 & -12.3 & 289.3
$0.75$ &    85845 &    90785 &   126828  &&   -9.8 & -22.5 & -206.4 \\%   87771 &    83331 &   126432  &  -16.0 & -22.4 & -210.7
$0.85$ &   147525 &   119810 &   249404  &&  -15.7 & -17.4 & -406.4 \\%  138270 &   115398 &   250714  &  -15.8 & -17.1 & -410.7
\multicolumn{8}{c}{\textit{Estimators of continuous endogenous regressor's coefficient}} \\
$0.15$ &     9360 &     7593 &    11434  &&   -3.3 &  -5.0 &  -7.0 \\
$0.25$ &    10641 &     9469 &    11434  &&   -3.4 &  -3.5 &  -7.0 \\
$0.50$ &    13991 &    12426 &    11434  &&   -5.7 &  -9.8 &  -7.0 \\
$0.75$ &    28114 &    25489 &    11434  &&  -12.3 & -17.9 &  -7.0 \\
$0.85$ &    43890 &    37507 &    11434  &&  -17.2 & -17.8 &  -7.0 \\
\hline
\end{tabular}
\end{table}

Table \ref{tab:sim-JTPA2} shows results from the second JTPA-based DGP.  The first estimator is now a nearly-unsmoothed SEE estimator instead of the unsmoothed \citet{ChernozhukovHansen2006} estimator.  Although in principle \citet{ChernozhukovHansen2006} can be used with multiple endogenous coefficients, the provided code allows only one, and Tables \ref{tab:emp} and \ref{tab:emp2} show that SEE with $h=400$ produces very similar results in the JTPA data.  For the binary endogenous regressor's coefficient, the 2SLS estimator now has the largest robust MSE since the larger sample size reduces the variance of all three estimators but does not reduce the 2SLS median bias (since it has first-order asymptotic bias).  The plug-in bandwidth yields smaller robust MSE than the nearly-unsmoothed $h=400$ at four of five quantiles.  At the median, for example, compared with $h=400$, $\hat h$ slightly increases the median bias but greatly reduces the dispersion, so the net effect is to reduce robust MSE.  This is consistent with the theoretical results.  For the continuous endogenous regressor's coefficient, the same pattern holds for the nearly-unsmoothed and $\hat h$-smoothed estimators.  Since this coefficient is constant across quantiles, the 2SLS estimator is consistent and very similar to the SEE estimators with $q=0.5$.



\subsection{Comparison of SEE and smoothed criterion function\label{sec:sim-SCF}}

For exogenous QR, smoothing the criterion function (SCF) is a different
approach, as discussed. The following simulations compare the MSE of our SEE
estimator with that of the SCF estimator. All DGPs have $n=50$, $X_{i}
\overset{iid}{\sim }\text{Unif}(1,5)$, $U_{i}\overset{iid}{\sim }N(0,1)$, $
X_{i}\mathpalette{\protect \independenT}{\perp}U_{i}$, and $
Y_{i}=1+X_{i}+\sigma (X_{i})\mathopen{}\mathclose \bgroup \originalleft( U_{i}-\Phi ^{-1}(q)\aftergroup \egroup \originalright) $. DGP 1 has $
q=0.5$ and $\sigma (X_{i})=5$. DGP 2 has $q=0.25$ and $\sigma
(X_{i})=1+X_{i} $. DGP 3 has $q=0.75$ and $\sigma (X_{i})=1+X_{i}$. In
addition to using our plug-in $\hat{h}$, we also compute the estimators for
a much smaller bandwidth in each DGP: $h=1$, $h=0.8$, and $h=0.8$,
respectively. Each simulation ran $1000$ replications. We
compare only the slope coefficient estimators.

\begin{table}[htbp]
\center
\caption{\label{tab:sim-SCF}Simulation results comparing SEE and SCF exogenous QR estimators.}
\begin{tabular}[c]{cS[table-format=1.3]S[table-format=1.3]cS[table-format=1.3]S[table-format=1.3]cS[table-format=2.3]S[table-format=2.3]cS[table-format=2.3]S[table-format=2.3]}
\hline\hline
 & \multicolumn{5}{c}{MSE} & & \multicolumn{5}{c}{Bias} \\
  \cline{2-6}\cline{8-12}
\rule{0pt}{14pt}
 & \multicolumn{2}{c}{Plug-in $\hat h$} && \multicolumn{2}{c}{small $h$} &&
   \multicolumn{2}{c}{Plug-in $\hat h$} && \multicolumn{2}{c}{small $h$} \\
  \cline{2-3}\cline{5-6} \cline{8-9}\cline{11-12}
\rule{0pt}{12pt}
DGP & \multicolumn{1}{c}{SEE} & \multicolumn{1}{c}{SCF} &
    & \multicolumn{1}{c}{SEE} & \multicolumn{1}{c}{SCF} &
    & \multicolumn{1}{c}{SEE} & \multicolumn{1}{c}{SCF} &
    & \multicolumn{1}{c}{SEE} & \multicolumn{1}{c}{SCF} \\
\hline
1 & 0.423 & 0.533 && 0.554 & 0.560 &&  -0.011 & -0.013 && -0.011 & -0.009 \\% &&  0.423 & 0.533 && 0.554 & 0.560
2 & 0.342 & 0.433 && 0.424 & 0.430 &&  0.092 & -0.025 && 0.012 & 0.012 \\% &&  0.334 & 0.432 && 0.424 & 0.430
3 & 0.146 & 0.124 && 0.127 & 0.121 &&  -0.292 & -0.245 && -0.250 & -0.232 \\% &&  0.061 & 0.065 && 0.064 & 0.067
\hline
\end{tabular}
\end{table}

Table \ref{tab:sim-SCF} shows MSE and bias for the SEE and SCF estimators, for our plug-in $\hat h$ as well as the small, fixed $h$ mentioned above.  The SCF estimator can have slightly lower MSE, as in the third DGP ($q=0.75$ with heteroskedasticity), but the SEE estimator has more substantially lower MSE in more DGPs, including the homoskedastic conditional median DGP.  The differences are quite small with the small $h$, as expected.  Deriving and implementing an MSE-optimal bandwidth for the SCF estimator could shrink the differences, but based on these simulations and the theoretical comparison in Section \ref{sec:see}, such an effort seems unlikely to yield improvement over the SEE estimator.



\subsection{Additional simulations\label{sec:sim-additional}}

We tried additional data generating processes (DGPs).  The first three DGPs are for exogenous QR, taken directly from \citet{Horowitz1998}.  In each case, $q=0.5$, $Y_i=X_i'\beta_0+U_i$, $\beta_0=(1,1)'$, $X_i=(1,x_i)'$ with $x_i\stackrel{iid}{\sim}\textrm{Uniform}(1,5)$, and $n=50$.  In DGP 1.1, the $U_i$ are sampled iid from a $t_3$ distribution scaled to have variance two.  In DGP 1.2, the $U_i$ are iid from a type I extreme value distribution again scaled and centered to have median zero and variance two.  In DGP 1.3, $U_i=(1+x_i)V/4$ where $V_i\stackrel{iid}{\sim}N(0,1)$.

DGPs 2.1, 2.2, and 3.1--3.6 are shown in the working paper version; they include variants of the \citet{Horowitz1998} DGPs with $q\ne0.5$, different error distributions, and another regressor.

DGPs 4.1--4.3 have endogeneity.  DGP 4.1 has $q=0.5$, $n=20$, and $\beta _{0}=(0,1)^{\prime }$.  It uses the reduced form
equations in \citet[equation 2]{CattaneoEtAl2012} with $\gamma _{1}=\gamma
_{2}=1$, $x_{i}=1$, $z_{i}\sim N(0,1)$, and $\pi =0.5$. Similar to their
simulations, we set $\rho =0.5$, $(\tilde{v}_{1i},\tilde{v}_{2i})$ iid $
N(0,1)$, and $(v_{1i},v_{2i})^{\prime }=\mathopen{}\mathclose \bgroup \originalleft(\tilde{v}_{1i},\sqrt{1-\rho ^{2}}
\tilde{v}_{2i}+\rho \tilde{v}_{1i}\aftergroup \egroup \originalright)'$.  DGP 4.2 is similar to DGP 4.1 but with $(\tilde v_{1i},\tilde v_{2i})'$ iid Cauchy, $n=250$, and $\beta_0=\mathopen{}\mathclose \bgroup \originalleft(0,\mathopen{}\mathclose \bgroup \originalleft[\rho-\sqrt{1-\rho^2}\aftergroup \egroup \originalright]^{-1}\aftergroup \egroup \originalright)'$.
DGP 4.3 is the same as DGP 4.1 but with $q=0.35$ (and consequent re-centering of the error term) and $n=30$.

We compare MSE for our SEE estimator using the plug-in $\hat h$ and estimators using different (fixed) values of $h$.  We include $h=0$ by using unsmoothed QR or the method in \citet{ChernozhukovHansen2006} for the endogenous DGPs.  We also include $h=\infty$ (although not in graphs) by using the usual IV estimator.  For the endogenous DGPs, we consider both MSE and the ``robust MSE'' defined in Section \ref{sec:sim-JTPA} as $\textrm{Bias}_{\textrm{median}}^{2}+(\textrm{IQR}/1.349)^2$.

For ``size-adjusted'' power (SAP) of a test with nominal size $\alpha$, the
critical value is picked as the $(1-\alpha)$-quantile of the simulated test
statistic distribution. This is for demonstration, not practice. The size
adjustment fixes the left endpoint of the size-adjusted power curve to the
null rejection probability $\alpha $. The resulting size-adjusted power
curve is one way to try to visualize a combination of type I and type II
errors, in the absence of an explicit loss function. One shortcoming is that
it does not reflect the variability/uniformity of size and power over the
space of parameter values and DGPs.

Regarding notation in the size-adjusted power figures, the vertical axis in the size-adjusted
power figures shows the simulated rejection probability. The horizontal axis
shows the magnitude of deviation from the null hypothesis, where a
randomized alternative is generated in each simulation iteration as that
magnitude times a random point on the unit sphere in $\mathbb{R}^{d}$, where
$\beta \in \mathbb{R}^{d}$. As the legend shows, the dashed line corresponds
to the unsmoothed estimator ($h=0$), the dotted line to the infeasible $h_{
\text{SEE}}^{\ast }$, and the solid line to the plug-in $\hat{h}$.

For the MSE graphs, the flat horizontal solid and dashed lines are the MSE
of the intercept and slope estimators (respectively) using feasible plug-in $
\hat{h}$ (recomputed each replication). The other solid and dashed lines
(that vary with $h$) are the MSE when using the value of $h$ from the
horizontal axis. The left vertical axis shows the MSE values for the
intercept parameter; the right vertical axis shows the MSE for slope
parameter(s); and the horizontal axis shows a log transformation of the
bandwidth, $\log _{10}(1+h)$.

Our plug-in bandwidth is quite stable.  The range of $\hat{h}$ values over the
simulation replications is usually less than a factor of $10$, and the range
from $0.05$ to $0.95$ empirical quantiles is around a factor of two. This corresponds to
a very small impact on MSE; note the log transformation in the x-axis in the MSE
graphs.


\begin{figure}[tbph]
\begin{center}
\includegraphics[clip=true,trim=25 175 145 200,width=.49\textwidth]
	{mse_psi1.1_noh_better.pdf}
	\hfill
\includegraphics[clip=true,trim=25 175 145 200,width=.49\textwidth]
	{mse_psi1.3_noh_better.pdf}
\end{center}
\caption{MSE for DGPs 1.1 (left) and 1.3 (right).}
\label{fig:mse-1a}
\end{figure}
\begin{figure}[tbph]
\begin{center}
\includegraphics[clip=true,trim=30 175 145 200,width=.48\textwidth]
	{adjpwr_psi1.1_better.pdf}
	\hfill
\includegraphics[clip=true,trim=30 175 145 200,width=.48\textwidth]
	{adjpwr_psi1.3_better.pdf}
\end{center}
\caption{Size-adjusted power for DGPs 1.1 (left) and 1.3 (right).}
\label{fig:adjpwr-1a}
\end{figure}

In DGPs 1.1--1.3, SEE($\hat h$) has smaller MSE than either the unsmoothed estimator or OLS, for both the intercept and slope coefficients.  Figure \ref{fig:mse-1a} shows MSE for DGPs 1.1 and 1.3.  It shows that the MSE of SEE($\hat h$) is very close to that of the best estimator with a fixed $h$.  In principle, a data-dependent $\hat h$ can attain MSE even lower than any fixed $h$.  SAP for SEE($\hat h$) is similar to that with $h=0$; see Figure \ref{fig:adjpwr-1a} for DGPs 1.1 and 1.3.

\begin{figure}[tbph]
\begin{center}
\includegraphics[clip=true,trim=0 175 125 200,width=.49\textwidth]
	{mse_psi5.2_noh_better.pdf}
	\hfill
\includegraphics[clip=true,trim=0 175 125 200,width=.49\textwidth]
	{medmse_psi5.2_noh_better.pdf}
\end{center}
\caption{For DGP 4.2, MSE (left) and ``robust
MSE'' (right): squared median-bias plus the square of the
interquartile range divided by $1.349$, $\text{Bias}_{\mathrm{median}}^{2}+(
\text{IQR}/1.349)^{2}$.}
\label{fig:mse-4.2}
\end{figure}

\begin{figure}[tbph]
\centering
\includegraphics[clip=true,trim=25 175 125 200,width=.49\textwidth]
	{mse_psi5.3_noh_better.pdf}
	\hfill
\includegraphics[clip=true,trim=25 175 125 200,width=.49\textwidth]
	{medmse_psi5.3_noh_better.pdf}
\caption{Similar to Figure \ref{fig:mse-4.2}, MSE (left) and ``robust MSE'' (right) for DGP 4.3.}
\label{fig:mse-4.3}
\end{figure}

Figures \ref{fig:mse-4.2} and \ref{fig:mse-4.3} show MSE and ``robust MSE'' for two DGPs with endogeneity.  Graphs for the other endogenous DGP (4.1) are similar to those for the slope estimator in DGP 4.3 but with larger MSE; they may be found in the working paper.  The MSE graph for DGP 4.2 is not as informative since it is sensitive to very large outliers that occur in only a few replications.  However, as shown, the MSE for SEE($\hat h$) is still better than that for the unsmoothed IV-QR estimator, and it is nearly the same as the MSE for the mean IV estimator (not shown: $1.1\times10^6$ for $\beta_1$, $2.1\times10^5$ for $\beta_2$).
For robust MSE, SEE($\hat h$) is again always better than the unsmoothed estimator.  For DGP 4.3 with normal errors and $q=0.35$, it is similar to the IV estimator, slightly worse for the slope coefficient and slightly better for the intercept, as expected.  Also as expected, for DGP 4.2 with Cauchy errors, SEE($\hat h$) is orders of magnitude better than the mean IV estimator.
Overall, using $\hat{h}$ appears
to consistently reduce the MSE of all estimator components compared with $h=0$ and with IV ($h=\infty $). Almost always, the exception is cases where
MSE is monotonically decreasing with $h$ (mean regression is more
efficient), in which $\hat{h}$ is much better than $h=0$ but not quite large
enough to match $h=\infty $.

\begin{figure}[tbph]
\begin{center}
\includegraphics[clip=true,trim=30 175 145 200,width=.48\textwidth]
	{adjpwr_psi5.1_better.pdf}
	\hfill
\includegraphics[clip=true,trim=30 175 145 200,width=.48\textwidth]
	{adjpwr_psi5.3_better.pdf}
\end{center}
\caption{Size-adjusted power for DGPs 4.1 (left) and 4.3 (right).}
\label{fig:adjpwr-4a}
\end{figure}

Figure \ref{fig:adjpwr-4a} shows SAP for DGPs 4.1 and 4.3.  The gain from smoothing is more substantial than in the exogenous DGPs, close to $10$ percentage points for a range of deviations.  Here, the randomness in $\hat h$ is not helpful.  In DGP 4.2 (not shown), the SAP for $\hat h$ is actually a few percentage points \emph{below} that for $h=0$ (which in turn is below the infeasible $h^*$), and in DGP 4.1, the SAP improvement from using the infeasible $h^*$ instead of $\hat h$ is similar in magnitude to the improvement from using $\hat h$ instead of $h=0$. Depending on one's loss function of type I and type II errors, the SEE-based test may be preferred or not.









\section{Conclusion\label{sec:conclusion}}

We have presented a new estimator for quantile regression with or without
instrumental variables. Smoothing the estimating equations (moment
conditions) has multiple advantages beyond the known advantage of allowing
higher-order expansions. It can reduce the MSE of both the estimating
equations and the parameter estimator, minimize type I error and improve
size-adjusted power of a chi-square test, and allow more reliable
computation of the instrumental variables quantile regression estimator
especially when the number of endogenous regressors is larger. We have given
the theoretical bandwidth that optimizes these properties, and simulations
show our plug-in bandwidth to reproduce all these advantages over the
unsmoothed estimator. Links to mean instrumental variables regression and
robust estimation are insightful and of practical use.

The strategy of smoothing the estimating equations can be applied to any
model with nonsmooth estimating equations; there is nothing peculiar to the
quantile regression model that we have exploited. For example, this strategy
could be applied to censored quantile regression, or to select the optimal
smoothing parameter in \citeauthor{Horowitz2002}'s (\citeyear{Horowitz2002})
smoothed maximum score estimator. The present paper has focused on
parametric and linear IV quantile regression; extensions to nonlinear IV
quantile regression and nonparametric IV quantile regression along the lines
of \citet{ChenPouzo2009,ChenPouzo2012} are currently under development.


\theendnotes

\bibliographystyle{chicago}
\begin{thebibliography}{}

\bibitem[\protect\citeauthoryear{Abadie, Angrist, and Imbens}{Abadie
  et~al.}{2002}]{AbadieEtAl2002}
Abadie, A., J.~Angrist, \& G.~Imbens (2002) Instrumental variables estimates
  of the effect of subsidized training on the quantiles of trainee earnings.
\newblock {\em Econometrica\/}~70, 91--117.

\bibitem[\protect\citeauthoryear{Bera, Bilias, and Simlai}{Bera
  et~al.}{2006}]{BeraEtAl2006}
Bera, A.~K., Y.~Bilias, \& P.~Simlai (2006)
\newblock Estimating functions and equations: An essay on historical
  developments with applications to econometrics.
\newblock In T.~C. Mills and K.~Patterson (Eds.), {\em Palgrave Handbook of
  Econometrics: Volume 1 Econometric Theory}, pp.\  427--476. Palgrave
  MacMillan.

\bibitem[\protect\citeauthoryear{Breiman}{Breiman}{1994}]{Breiman1994}
Breiman, L. (1994)
\newblock Bagging predictors.
\newblock Technical Report 421, Department of Statistics, University of
  California, Berkeley.

\bibitem[\protect\citeauthoryear{Cattaneo, Crump, and Jansson}{Cattaneo
  et~al.}{2012}]{CattaneoEtAl2012}
Cattaneo, M.~D., R.~K. Crump, \& M.~Jansson (2012) Optimal inference for
  instrumental variables regression with non-{G}aussian errors.
\newblock {\em Journal of Econometrics\/}~167, 1--15.

\bibitem[\protect\citeauthoryear{Chamberlain}{Chamberlain}{1987}]{Chamberlain1987}
Chamberlain, G. (1987) Asymptotic efficiency in estimation with conditional
  moment restrictions.
\newblock {\em Journal of Econometrics\/}~34, 305--334.

\bibitem[\protect\citeauthoryear{Chen and Pouzo}{Chen and
  Pouzo}{2009}]{ChenPouzo2009}
Chen, X. \& D.~Pouzo (2009) Efficient estimation of semiparametric
  conditional moment models with possibly nonsmooth residuals.
\newblock {\em Journal of Econometrics\/}~152, 46--60.

\bibitem[\protect\citeauthoryear{Chen and Pouzo}{Chen and
  Pouzo}{2012}]{ChenPouzo2012}
Chen, X. \& D.~Pouzo (2012) Estimation of nonparametric conditional moment
  models with possibly nonsmooth moments.
\newblock {\em Econometrica\/}~80, 277--322.

\bibitem[\protect\citeauthoryear{Chernozhukov, Hansen, and
  Jansson}{Chernozhukov et~al.}{2009}]{ChernozhukovEtAl2009}
Chernozhukov, V., C.~Hansen, \& M.~Jansson (2009) Finite sample inference for
  quantile regression models.
\newblock {\em Journal of Econometrics\/}~152, 93--103.

\bibitem[\protect\citeauthoryear{Chernozhukov and Hansen}{Chernozhukov and
  Hansen}{2005}]{ChernozhukovHansen2005}
Chernozhukov, V. \& C.~B. Hansen (2005) An {IV} model of quantile treatment
  effects.
\newblock {\em Econometrica\/}~73, 245--261.

\bibitem[\protect\citeauthoryear{Chernozhukov and Hansen}{Chernozhukov and
  Hansen}{2006}]{ChernozhukovHansen2006}
Chernozhukov, V. \& C.~B. Hansen (2006) Instrumental quantile regression
  inference for structural and treatment effect models.
\newblock {\em Journal of Econometrics\/}~132, 491--525.

\bibitem[\protect\citeauthoryear{Chernozhukov and Hansen}{Chernozhukov and
  Hansen}{2008}]{ChernozhukovHansen2008}
Chernozhukov, V. \& C.~B. Hansen (2008) Instrumental variable quantile
  regression: A robust inference approach.
\newblock {\em Journal of Econometrics\/}~142, 379--398.

\bibitem[\protect\citeauthoryear{Chernozhukov and Hansen}{Chernozhukov and
  Hansen}{2013}]{ChernozhukovHansen2013}
Chernozhukov, V. \& C.~B. Hansen (2013) Quantile models with endogeneity.
\newblock {\em Annual Review of Economics\/}~5, 57--81.

\bibitem[\protect\citeauthoryear{Chernozhukov and Hong}{Chernozhukov and
  Hong}{2003}]{ChernozhukovHong2003}
Chernozhukov, V. \& H.~Hong (2003) An {MCMC} approach to classical
  estimation.
\newblock {\em Journal of Econometrics\/}~115, 293--346.

\bibitem[\protect\citeauthoryear{Fan and Liao}{Fan and
  Liao}{2014}]{FanLiao2014}
Fan, J. \& Y.~Liao (2014) Endogeneity in high dimensions.
\newblock {\em Annals of Statistics\/}~42, 872--917.

\bibitem[\protect\citeauthoryear{Galvao}{Galvao}{2011}]{Galvao2011}
Galvao, A.~F. (2011) Quantile regression for dynamic panel data with fixed
  effects.
\newblock {\em Journal of Econometrics\/}~164, 142--157.

\bibitem[\protect\citeauthoryear{Hall}{Hall}{1992}]{Hall1992}
Hall, P. (1992)
\newblock {\em Bootstrap and {Edgeworth} Expansion}.
\newblock Springer Series in Statistics. New York: Springer-{V}erlag.

\bibitem[\protect\citeauthoryear{Heyde}{Heyde}{1997}]{Heyde1997}
Heyde, C.~C. (1997)
\newblock {\em Quasi-Likelihood and Its Application: A General Approach to
  Optimal Parameter Estimation}.
\newblock Springer Series in Statistics. New York: Springer.

\bibitem[\protect\citeauthoryear{Horowitz}{Horowitz}{1992}]{Horowitz1992}
Horowitz, J.~L. (1992) A smoothed maximum score estimator for the binary
  response model.
\newblock {\em Econometrica\/}~60, 505--531.

\bibitem[\protect\citeauthoryear{Horowitz}{Horowitz}{1998}]{Horowitz1998}
Horowitz, J.~L. (1998) Bootstrap methods for median regression models.
\newblock {\em Econometrica\/}~66, 1327--1351.

\bibitem[\protect\citeauthoryear{Horowitz}{Horowitz}{2002}]{Horowitz2002}
Horowitz, J.~L. (2002) Bootstrap critical values for tests based on the
  smoothed maximum score estimator.
\newblock {\em Journal of Econometrics\/}~111, 141--167.

\bibitem[\protect\citeauthoryear{Huber}{Huber}{1964}]{Huber1964}
Huber, P.~J. (1964) Robust estimation of a location parameter.
\newblock {\em The Annals of Mathematical Statistics\/}~35, 73--101.

\bibitem[\protect\citeauthoryear{Hwang and Sun}{Hwang and
  Sun}{2015}]{HwangSun2015}
Hwang, J. \& Y.~Sun (2015)
\newblock Should we go one step further? {A}n accurate comparison of one-step
  and two-step procedures in a generalized method of moments framework.
\newblock Working paper, Department of Economics, UC San Diego.

\bibitem[\protect\citeauthoryear{Jun}{Jun}{2008}]{Jun2008}
Jun, S.~J. (2008) Weak identification robust tests in an instrumental quantile
  model.
\newblock {\em Journal of Econometrics\/}~144, 118--138.

\bibitem[\protect\citeauthoryear{Kinal}{Kinal}{1980}]{Kinal1980}
Kinal, T.~W. (1980) The existence of moments of k-class estimators.
\newblock {\em Econometrica\/}~48, 241--249.

\bibitem[\protect\citeauthoryear{Koenker and Bassett}{Koenker and
  Bassett}{1978}]{KoenkerBassett1978}
Koenker, R. \& G.~Bassett, Jr. (1978) Regression quantiles.
\newblock {\em Econometrica\/}~46, 33--50.

\bibitem[\protect\citeauthoryear{Kwak}{Kwak}{2010}]{Kwak2010}
Kwak, D.~W. (2010)
\newblock Implementation of instrumental variable quantile regression ({IVQR})
  methods.
\newblock Working paper, Michigan State University.

\bibitem[\protect\citeauthoryear{Liang and Zeger}{Liang and
  Zeger}{1986}]{LiangZeger1986}
Liang, K.-Y. \& S.~Zeger (1986) Longitudinal data analysis using generalized
  linear models.
\newblock {\em Biometrika\/}~73, 13--22.

\bibitem[\protect\citeauthoryear{MaCurdy and Hong}{MaCurdy and
  Hong}{1999}]{MaCurdyHong1999}
MaCurdy, T. \& H.~Hong (1999)
\newblock Smoothed quantile regression in generalized method of moments.
\newblock Working paper, Stanford University.

\bibitem[\protect\citeauthoryear{M{\"u}ller}{M{\"u}ller}{1984}]{Muller1984}
M{\"u}ller, H.-G. (1984) Smooth optimum kernel estimators of densities,
  regression curves and modes.
\newblock {\em The Annals of Statistics\/}~12, 766--774.

\bibitem[\protect\citeauthoryear{Nagar}{Nagar}{1959}]{Nagar1959}
Nagar, A.~L. (1959) The bias and moment matrix of the general k-class
  estimators of the parameters in simultaneous equations.
\newblock {\em Econometrica\/}~27, 573--595.

\bibitem[\protect\citeauthoryear{Newey}{Newey}{1990}]{Newey1990}
Newey, W.~K. (1990) Efficient instrumental variables estimation of nonlinear
  models.
\newblock {\em Econometrica\/}~58, 809--837.

\bibitem[\protect\citeauthoryear{Newey}{Newey}{2004}]{Newey2004}
Newey, W.~K. (2004) Efficient semiparametric estimation via moment
  restrictions.
\newblock {\em Econometrica\/}~72, 1877--1897.

\bibitem[\protect\citeauthoryear{Newey and Powell}{Newey and
  Powell}{1990}]{NeweyPowell1990}
Newey, W.~K. \& J.~L. Powell (1990) Efficient estimation of linear and type
  {I} censored regression models under conditional quantile restrictions.
\newblock {\em Econometric Theory\/}~6, 295--317.

\bibitem[\protect\citeauthoryear{Otsu}{Otsu}{2008}]{Otsu2008}
Otsu, T. (2008) Conditional empirical likelihood estimation and inference for
  quantile regression models.
\newblock {\em Journal of Econometrics\/}~142, 508--538.

\bibitem[\protect\citeauthoryear{Phillips}{Phillips}{1982}]{Phillips1982}
Phillips, P. C.~B. (1982)
\newblock Small sample distribution theory in econometric models of
  simultaneous equations.
\newblock Cowles Foundation Discussion Paper 617, Yale University.

\bibitem[\protect\citeauthoryear{Ruppert and Carroll}{Ruppert and
  Carroll}{1980}]{RuppertCarroll1980}
Ruppert, D. \& R.~J. Carroll (1980) Trimmed least squares estimation in the
  linear model.
\newblock {\em Journal of the American Statistical Association\/}~75, 828--838.

\bibitem[\protect\citeauthoryear{van~der Vaart}{van~der
  Vaart}{1998}]{vanderVaart1998}
van~der Vaart, A.~W. (1998)
\newblock {\em Asymptotic Statistics}.
\newblock Cambridge: Cambridge University Press.

\bibitem[\protect\citeauthoryear{Whang}{Whang}{2006}]{Whang2006}
Whang, Y.-J. (2006) Smoothed empirical likelihood methods for quantile
  regression models.
\newblock {\em Econometric Theory\/}~22, 173--205.

\bibitem[\protect\citeauthoryear{Zhou, Wan, and Yuan}{Zhou
  et~al.}{2011}]{ZhouEtAl2011}
Zhou, Y., A.~T.~K. Wan, \& Y.~Yuan (2011) Combining least-squares and
  quantile regressions.
\newblock {\em Journal of Statistical Planning and Inference\/}~141,
  3814--3828.

\end{thebibliography}