EconBase
← Back to paper

On semiparametric estimation of the intercept of the sample selection model: a kernel approach

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

89,604 characters

On semiparametric estimation of the intercept of the sample selection model: a kernel approach



\title{\textbf{On semiparametric estimation of the intercept of the sample
selection model: \\
a kernel approach}}
\author{Zhewen Pan \\
School of Economics, Zhejiang University of Finance \& Economics \\
{\normalsize 18 Xueyuan Street, Xiasha, Hangzhou 310018, China. Email:
[email removed]}}
\maketitle

\begin{abstract}
This paper presents a new perspective on the identification at infinity for
the intercept of the sample selection model as identification at the
boundary via a transformation of the selection index. This perspective
suggests generalizations of estimation at infinity to kernel regression
estimation at the boundary and further to local linear estimation at the
boundary. The proposed kernel-type estimators with an estimated
transformation are proven to be nonparametric-rate consistent and
asymptotically normal under mild regularity conditions. A fully data-driven
method of selecting the optimal bandwidths for the estimators is developed.
The Monte Carlo simulation shows the desirable finite sample properties of
the proposed estimators and bandwidth selection procedures.
\end{abstract}

\noindent \emph{Keywords}{\small : identification at infinity, kernel
regression estimation, boundary effect, local linear estimation, bandwidth
selection}

\noindent \emph{JEL codes}{\small : C01, C13, C14, C34}

\newpage

\section{Introduction}

\label{sec:intro}Since it was first introduced by the seminal papers of \cite
{heckman1974shadow}, \cite{gronau1974wage}$,$ and \cite{lewis1974comments},
the sample selection model has been increasingly widely applied in empirical
studies to address potentially nonrandom samples that may arise from a
variety of causes, such as improper sampling design, self-selectivity,
nonresponse on survey questions, and attrition from social programs. Several
important applications of the sample selection model, such as the estimation
of average treatment effects and the decomposition of wage differentials,
rely on an estimate of the intercept.

The intercept of the sample selection model is conventionally estimated
along with the slope coefficients by means of the parametric maximum
likelihood approach or the likelihood-based two-step procedure
\citep{heckman1979sample}. When the distribution of the model disturbance is
misspecified, however, the parametric likelihood-based estimators for
limited dependent variable models may possess evident bias and are likely
inconsistent for the true value of interest
\citep[e.g.,][]{arabmazar1982investigation}. This finding spawned
considerable and influential literature on semiparametric identification and
estimation approaches that do not rest on parametric specification of the
disturbance distribution, among which \cite{chamberlain1986asymptotic}
invented the notion of \textquotedblleft identification at
infinity\textquotedblright , building identification upon the unbounded
support of the regressor distribution. According to the identification at
infinity, \cite{heckman1990varieties} proposed a semiparametric estimator
for the intercept of the sample selection model. However,
\citeauthor{heckman1990varieties}'s estimator contains discontinuous
indicator functions, which greatly complicate the analysis of the
statistical properties. \cite{andrews1998semiparametric} suggested replacing
the indicator function with a smoothed version and established the
consistency and asymptotic normality of their modified estimator. Since
then, the identification-at-infinity method has been standard for
semiparametrically estimating the intercept of the sample selection model
\citep[see][among others]{schafgans1998ethnic, schafgans2000gender,
hussinger2008r, mulligan2008selection, liu2009maternal, shen2013determinants}
.

Nevertheless, identification-at-infinity estimation must choose a smoothing
parameter or bandwidth controlling for the proportion of observations used,
and no bandwidth selection algorithm that is both theoretically valid and
practically tractable in the context of the intercept estimation has yet
been developed. This is mainly because the asymptotic biases and variances
of the estimators of \cite{heckman1990varieties} and \cite
{andrews1998semiparametric} are all implicit functions of the bandwidth. To
choose a theoretically valid bandwidth, one must impose high-level
assumptions on the tail behaviors of the regressor and disturbance
distributions in the selection equation and then carefully estimate some
sort of tail indices, as in \cite{klein2015estimation}. However, a precise
estimate of the tail index of the selection disturbance distribution is
difficult to obtain for two reasons. First, only the information in the
tails, which is limited, is useful for estimating tail indices. Second,
there is additional information loss when estimating the tail index of the
selection disturbance distribution because one can observe only the binary
selection outcome instead of the continuous selection disturbance. The
imprecise estimation of the tail index makes it difficult to choose an
eligible bandwidth in practice \citep{tan2018root}. In empirical studies,
researchers typically report several estimates of the intercept according to
different values of bandwidth selected by simple rules, such as by sample
quantiles of the selection linear index. However, these simple rules lack
theoretical justification, and an inappropriately selected bandwidth may
lead to sizable estimation bias and misleading inference results
\citep{schafgans2004finite}. Additionally, confusion will arise if different
bandwidths lead to contradictory conclusions.

In this paper, I extend the identification-at-infinity estimators to a
kernel regression estimator at the boundary and provide a simple bandwidth
selection algorithm that is theoretically optimal in the sense of minimizing
the asymptotic mean squared estimation error. By transforming the
identification at infinity into identification at the boundary, I propose a
kernel regression estimator for the intercept of the sample selection model
that includes the estimators of \cite{heckman1990varieties} and \cite
{andrews1998semiparametric} as special cases by taking particular forms of
transformation. I suggest using the cumulative distribution function (CDF)
of the selection linear index as the transformation, which makes the
proposed kernel estimator substantially different from the existing
identification-at-infinity estimators. Under this specific form of
transformation, the asymptotic bias and variance of the kernel estimator are
explicit functions of the bandwidth, motivating a simple plug-in procedure
of optimal bandwidth selection. In practice, however, the CDF of the linear
index is unknown. I adopt the empirical CDF estimation and show that the
induced estimation error is asymptotically negligible under regularity
conditions, implying that the kernel estimator with the empirical CDF
follows the same asymptotic distribution as if the true CDF were known. As a
result, the proposed bandwidth selection algorithm is theoretically valid.

Careful consideration of the asymptotic distribution of the kernel
regression estimator for the intercept indicates a boundary effect in that
the asymptotic bias has a larger order than the nonparametric kernel
regression estimator in the interior because the estimator is, in nature, a
Nadaraya-Watson or local constant estimator at the boundary. Hence, I
further consider the local linear estimation of the intercept for bias
reduction and show that it achieves a univariate nonparametric rate. The
consistency and asymptotic normality of the estimator are established, and
an optimal bandwidth selection algorithm tailored to the local linear case
is presented.

The main contributions of this paper are as follows. First, I provide a
novel interpretation of identification at infinity. Although infinity is
intractable or at least irregular in most econometric models, it is merely
an ordinary boundary point of the extended real line. Under the monotonic
transformation that serves to construct a metric, the identification at
infinity is converted into identification at the boundary, and kernel-type
estimators are naturally derived. This approach may be helpful in other
irregular identification problems involving infinity
\citep[see,
e.g.,][]{khan2010irregular}. Second, I provide a theoretically justified
procedure for bandwidth selection that is easy to implement. By implementing
a special form of transformation, the asymptotic biases and variances of the
kernel-type estimators become explicit functions of the bandwidth, based on
which a fully data-driven algorithm is developed to choose the optimal
bandwidth. The regularization method in \cite{imbens2012optimal} is employed
to keep the random denominator of the estimated bandwidth away from zero.
Third, in the course of developing the statistical properties of the
estimators, I provide a\ rigorous treatment of the estimation error induced
by the empirical CDF. By means of a decomposition into an empirical process
component and a discontinuous error component, the estimation error of the
empirical CDF can be delicately controlled by the uniform rate of
convergence of the empirical process over shrinking intervals
\citep[e.g.,][]{stute1982oscillation} and by the convergence results
developed by \cite{schafgans2002intercept}.

The rest of this paper is organized as follows. Section \ref{sec:model}
presents the sample selection model and reviews the semiparametric
estimators for the intercept in the literature. Section \ref{sec:KRE}
motivates the kernel regression estimator, provides the regularity
conditions under which the estimator is consistent and asymptotically
normal, discusses the choice of the transformation, and presents the optimal
bandwidth selection method. Section \ref{sec:LL} proposes local linear
estimation to eliminate the boundary effect. Section \ref{sec:simulation}
reports the results from a small simulation and Section \ref{sec:conclution}
concludes. All the technical proofs are left to the Appendix.

\section{The model and existing estimators}

\label{sec:model}

\subsection{The model}

The sample selection model has the following three-equation form:
\begin{eqnarray}
D_{i} &=&1\left\{ X_{i}^{\prime }\beta _{0}>\varepsilon _{i}\right\} ,
\notag \\
Y_{i}^{\ast } &=&\mu _{0}+Z_{i}^{\prime }\theta _{0}+U_{i},  \label{model} \\
Y_{i} &=&Y_{i}^{\ast }D_{i},  \notag
\end{eqnarray}
where $D_{i}$ is a binary selection variable, $Y_{i}^{\ast }$ is the latent
outcome, and $Y_{i}$ is the observed outcome. $X_{i}$ and $Z_{i}$ are random
vectors of regressors (excluding the constant term) of the selection and
outcome equations, respectively. $\varepsilon _{i}$ and $U_{i}$ are scalar
disturbances that are possibly correlated with each other. $\beta _{0}$ and $
\theta _{0}$ are vectors of the slope coefficients, and $\mu _{0}$ is a
scalar intercept coefficient, both of which are the estimands of the model. $
U_{i}$ is assumed to have zero mean to ensure identifiability of the
intercept.

This paper is primarily concerned with the estimation of $\mu _{0}$ in the
sample selection model (\ref{model}) without imposing a normal (or any other
parametric) distribution restriction on the disturbances. The estimation of $
\mu _{0}$, however, relies on preliminary $\sqrt{n}$-consistent estimators
for $\beta _{0}$ and $\theta _{0}$, which are denoted by $\hat{\beta}$ and $
\hat{\theta}$. Several such estimators are available in the literature. For
instance, $\hat{\beta}$ can be one of the semiparametric estimators for the
binary response model
\citep[e.g.,][]{klein1993efficient,lewbel2000semiparametric} or for the
single-index model
\citep[e.g.,][]{powell1989semiparametric,ichimura1993semiparametric}, and $
\hat{\theta}$ can be one of the semiparametric slope estimators for the
sample selection model
\citep[e.g.,][to name a few]{gallant1987semi, chen1998efficient,
powell2001semiparametric, lewbel2007endogenous, newey2009twostep, chen2010semiparametric}
. For notational convenience, denote $W_{i}=X_{i}^{\prime }\beta _{0}$ and $
\hat{W}_{i}=X_{i}^{\prime }\hat{\beta}$, hence $D_{i}=1\left\{
W_{i}>\varepsilon _{i}\right\} $.

\subsection{The identification-at-infinity estimators}

A natural strategy for identifying the intercept $\mu _{0}$ starts from the
expectation of $Y_{i}-Z_{i}^{\prime }\theta _{0}$ conditional on the
observation being selected:
\begin{equation*}
E\left[ Y_{i}-Z_{i}^{\prime }\theta _{0}\left\vert D_{i}=1\right. \right] =E
\left[ Y_{i}^{\ast }-Z_{i}^{\prime }\theta _{0}\left\vert D_{i}=1\right.
\right] =\mu _{0}+E\left[ U_{i}\left\vert D_{i}=1\right. \right] .
\end{equation*}
In the context of sample selection, where $U_{i}$ is correlated with $
\varepsilon _{i}$, the selectivity bias term $E\left[ U_{i}\left\vert
D_{i}=1\right. \right] $ is generally nonzero and contaminates the
identification of $\mu _{0}$. Inspired by the fact that $E\left[
U_{i}\left\vert D_{i}=1\right. \right] $ will become arbitrarily close to
zero for those $W_{i}$ such that $\Pr \left( D_{i}=1\left\vert W_{i}\right.
\right) $ is arbitrarily close to unity, \cite{heckman1990varieties}
suggested identifying $\mu _{0}$ at infinity, i.e., when $W_{i}$ takes
arbitrarily large values. This idea can be heuristically formulated as
\begin{equation*}
E\left[ Y_{i}-Z_{i}^{\prime }\theta _{0}\left\vert D_{i}=1,W_{i}=+\infty
\right. \right] =\mu _{0}+E\left[ U_{i}\left\vert D_{i}=1\right.
,W_{i}=+\infty \right] =\mu _{0}+E\left[ U_{i}\left\vert W_{i}=+\infty
\right. \right] =\mu _{0},
\end{equation*}
provided the conditional mean independence is assumed. An intuitive
estimator based on the above identification strategy is
\begin{equation}
\hat{\mu}^{H}=\frac{\sum_{i=1}^{n}\left( Y_{i}-Z_{i}^{\prime }\hat{\theta}
\right) D_{i}\cdot 1\left\{ \hat{W}_{i}>\gamma _{n}\right\} }{
\sum_{i=1}^{n}D_{i}\cdot 1\left\{ \hat{W}_{i}>\gamma _{n}\right\} },
\label{H1990Est}
\end{equation}
where $\left\{ \gamma _{n}\right\} $ is a sequence of positive smoothing
parameters such that $\gamma _{n}\rightarrow \infty $ as $n\rightarrow
\infty $. The estimator $\hat{\mu}^{H}$ is essentially a sample average of $
\mu _{0}+U_{i}$ over a decreasingly small fraction of all observations. The
effective sample size depends on the proportion of data censoring, the
degree of tail heaviness of $\hat{W}_{i}$'s distribution, and the choice of $
\gamma _{n}$.

To facilitate the development of distribution theory of the
identification-at-infinity estimation, \cite{andrews1998semiparametric}
replaced the indicator function in \citeauthor{heckman1990varieties}'s
estimator (\ref{H1990Est}) with a smoothed function and proposed a weighted
sample average estimator:
\begin{equation}
\hat{\mu}^{AS}=\frac{\sum_{i=1}^{n}\left( Y_{i}-Z_{i}^{\prime }\hat{\theta}
\right) D_{i}\cdot s\left( \hat{W}_{i}-\gamma _{n}\right) }{
\sum_{i=1}^{n}D_{i}\cdot s\left( \hat{W}_{i}-\gamma _{n}\right) },
\label{AS1998Est}
\end{equation}
where $s\left( \cdot \right) $ is a nondecreasing $\left[ 0,1\right] $
-valued function that has a third derivative bounded over $R$ and satisfies $
s\left( w\right) =0$ for $w\leq 0$ and $s\left( w\right) =1$ for $w\geq b$,
for some $b>0$. By taking advantage of the smoothness of $s\left( \cdot
\right) $, \cite{andrews1998semiparametric} showed the asymptotic
negligibility of the estimation error induced by the preliminary estimators $
\hat{\beta}$ and $\hat{\theta}$ and established the consistency and
asymptotic normality of $\hat{\mu}^{AS}$ under mild regularity conditions.
The rate of convergence of $\hat{\mu}^{AS}$ depends both upon the tail
heaviness of $W_{i}$'s distribution and upon the rate of divergence of $
\gamma _{n}$. However, theoretically valid procedures for choosing $\gamma
_{n}$ have not yet been developed.

\subsection{Other existing estimators}

The semiparametric literature has neglected, to a certain extent, the
intercept estimation of the sample selection model. This is mainly because
the intercept is absorbed into the selectivity bias correction term in the
course of estimating the slope coefficients. The only exceptions, besides $
\hat{\mu}^{H}$ and $\hat{\mu}^{AS}$, are the estimators of \cite
{gallant1987semi}, \cite{lewbel2007endogenous}, and \cite
{chen2010semiparametric}.

By employing Hermite series to approximate the unknown bivariate density
function of the disturbances, \cite{gallant1987semi} considered
semiparametric maximum likelihood estimation of the intercept along with the
slopes. The consistency of their estimator requires complicated continuity
conditions on the distributions of the disturbances and regressors that are
difficult to verify. Moreover, the asymptotic distribution of their
estimator has not been established. \cite{lewbel2007endogenous} achieved
identification of the intercept by the presence of a special regressor that
has a large support. When the support is infinite, his identification
strategy would be essentially equivalent to the identification at infinity.
\cite{chen2010semiparametric} proposed a kernel-weighted
pairwise-difference-type estimator for the intercept and showed its
consistency and asymptotic normality. However, their estimation method
requires the disturbances to be jointly symmetrically distributed; this
joint symmetry condition may not hold in practice.

\section{Kernel regression estimation}

\label{sec:KRE}The kernel approach to estimating the intercept $\mu _{0}$ in
the sample selection model (\ref{model}) is grounded in the identification
at infinity:
\begin{equation}
\mu _{0}=E\left[ Y_{i}-Z_{i}^{\prime }\theta _{0}\left\vert
D_{i}=1,W_{i}=+\infty \right. \right] .  \label{IdenAtInf}
\end{equation}
Let $F\left( \cdot \right) $ be an absolutely continuous CDF that is
strictly increasing over $R$; then, the infinity condition $W_{i}=+\infty $
is equivalent to a boundary condition $F\left( W_{i}\right) =F\left( +\infty
\right) =1$. Therefore, the identification at infinity (\ref{IdenAtInf}) can
be written as identification at the boundary:
\begin{equation}
\mu _{0}=E\left[ Y_{i}-Z_{i}^{\prime }\theta _{0}\left\vert D_{i}=1,F\left(
W_{i}\right) =1\right. \right] .  \label{IdenAtBoundary}
\end{equation}
In other words, $\mu _{0}$ is identified by the conditional expectation of $
Y_{i}-Z_{i}^{\prime }\theta _{0}$ given that the observation is selected and
that the transformation of $W_{i}$, $F\left( W_{i}\right) $, takes a
boundary value. The identification-at-boundary equation (\ref{IdenAtBoundary}
) motivates the kernel regression estimation of $\mu _{0}$:
\begin{equation}
\tilde{\mu}=\frac{\displaystyle\sum_{i=1}^{n}\left( Y_{i}-Z_{i}^{\prime }
\hat{\theta}\right) D_{i}k\left( \frac{1-F\left( \hat{W}_{i}\right) }{h_{n}}
\right) }{\displaystyle\sum_{i=1}^{n}D_{i}k\left( \frac{1-F\left( \hat{W}
_{i}\right) }{h_{n}}\right) },  \label{LC}
\end{equation}
where $\hat{W}_{i}=X_{i}^{\prime }\hat{\beta}$, $\hat{\beta}$ and $\hat{
\theta}$ are preliminary estimators for $\beta _{0}$ and $\theta _{0}$, $
k\left( \cdot \right) $ is a kernel function defined on $\left[ 0,\infty
\right) $, and $\left\{ h_{n}\right\} $ is a sequence of positive bandwidth
parameters such that $h_{n}\rightarrow 0$ as $n\rightarrow \infty $. The
values of the kernel over the negative reals are irrelevant since the
argument of it is always positive. Some examples of $k\left( \cdot \right) $
are given in Table \ref{table:kernel}. $\tilde{\mu}$ is a kernel regression
estimator at the boundary. Alternatively, it can be viewed as a kernel
regression estimator at infinity over the extended real line $\bar{R}=\left[
-\infty ,+\infty \right] =R\cup \left\{ -\infty ,+\infty \right\} $. A
distance function with which $\bar{R}$ would become a (compact) metric space
can be defined as $d\left( w_{1},w_{2}\right) =\left\vert F\left(
w_{1}\right) -F\left( w_{2}\right) \right\vert $\ for any $w_{1},w_{2}\in
\bar{R}$. Accordingly, the numerator within the kernel is the distance
between $\hat{W}_{i}$ and infinity since $d\left( \hat{W}_{i},+\infty
\right) =\left\vert F\left( \hat{W}_{i}\right) -1\right\vert =1-F\left( \hat{
W}_{i}\right) $. From this perspective, our estimation method is an
extension of kernel regression to the extended real line.

Similar to the identification-at-infinity estimators $\hat{\mu}^{H}$ and $
\hat{\mu}^{AS}$ defined in (\ref{H1990Est}) and (\ref{AS1998Est}), the
kernel estimator $\tilde{\mu}$ is also a (weighted) sample average over a
vanishingly small subset of the entire data set. However, the local average
in $\tilde{\mu}$ is over observations for $F\left( \hat{W}_{i}\right) $ in a
left neighborhood of the right boundary point, rather than for $\hat{W}_{i}$
near positive infinity. Notably, $\tilde{\mu}$ generalizes the
identification-at-infinity estimators in the sense that it nests $\hat{\mu}
^{H}$ and $\hat{\mu}^{AS}$ as special cases by taking specific forms of $
F\left( \cdot \right) $, $k\left( \cdot \right) $ and $h_{n}$.

\begin{prop}
(i) For any strictly increasing $F\left( \cdot \right) $, if $k\left(
u\right) =1\left\{ 0\leq u\leq 1\right\} $ and $h_{n}=1-F\left( \gamma
_{n}\right) $, then $\tilde{\mu}=\hat{\mu}^{H}$. (ii) If $F\left( \cdot
\right) $ is the standard Laplacian CDF such that $F\left( w\right)
=1-\left( 1\left/ 2\right. \right) \exp \left( -w\right) $ for $w\geq 0$, if
$k\left( u\right) =s\left( -\log u\right) $, and if $h_{n}=1-F\left( \gamma
_{n}\right) =\left( 1\left/ 2\right. \right) \exp \left( -\gamma _{n}\right)
$, then $\tilde{\mu}=\hat{\mu}^{AS}$.
\end{prop}

\subsection{Asymptotic property}

This subsection investigates the consistency and asymptotic normality of the
kernel estimator $\tilde{\mu}$ under general $F\left( \cdot \right) $.
Several regularity assumptions are made first. Denote $k_{ni}=k\left( \left.
\left( 1-F\left( W_{i}\right) \right) \right/ h_{n}\right) $, where $
W_{i}=X_{i}^{\prime }\beta _{0}$.

\bigskip

\noindent\textbf{Assumption 1.} (i) $\left\{ \left(
Y_{i},D_{i},X_{i},Z_{i}\right) \right\} _{i=1}^{n}$ is a random sample of $n$
observations from Model (\ref{model}). (ii) The model disturbance $\left(
\varepsilon _{i},U_{i}\right) $ is independent of the regressors $X_{i}$ and
$Z_{i}$. (iii) $EU_{i}=0$. (iv) There exists a positive constant $c_{1}$
such that $E\left| U_{i}\right| ^{2+c_{1}}<\infty $, $E\left\| X_{i}\right\|
^{4+c_{1}}<\infty $, and $E\left\| Z_{i}\right\| ^{2+c_{1}}<\infty $.

\noindent\textbf{Assumption 2.} The kernel function $k\left( \cdot \right) $
is defined on $\left[ 0,\infty \right) $ and satisfies that (i) it is
nonnegative and supported on $\left[ 0,1\right] $, (ii) it is bounded above
by $\bar{k}>0$, and (iii) it is twice continuously differentiable over $
\left[ 0,\infty \right) $ and its derivatives $k^{\prime }\left( \cdot
\right) $ and $k^{\prime \prime }\left( \cdot \right) $ are bounded above by
$\bar{k^{\prime }}$ and $\bar{k^{\prime \prime }}$, respectively.

\noindent\textbf{Assumption 3.} (i) The transformation function $F\left(
\cdot \right) $ is a CDF for a continuously distributed random variable
whose support is right-unbounded. (ii) The probability density function
(PDF) $f\left( \cdot \right) $ corresponding to $F\left( \cdot \right) $ is
absolutely continuous with respect to the Lebesgue measure, with its
derivative $f^{\prime }\left( \cdot \right) $ bounded almost everywhere by $
\bar{f^{\prime }}$. (iii) For a large positive constant $C$, the tail hazard
rate $H_{C}\left( \cdot \right) =1\left\{ \cdot >C\right\} f\left( \cdot
\right) \left/ \left[ 1-F\left( \cdot \right) \right] \right. $ and the tail
score (of location parameter) $S_{C}\left( \cdot \right) =1\left\{ \cdot
>C\right\} \left. f^{\prime }\left( \cdot \right) \right/ f\left( \cdot
\right) $ corresponding to $F\left( \cdot \right) $ satisfy $E\left[
\sup_{\beta \in \mathcal{N}\left( \beta _{0}\right) }H_{C}^{4}\left(
X_{i}^{\prime }\beta \right) \right] <\infty $ and $E\left[ \sup_{\beta \in
\mathcal{N}\left( \beta _{0}\right) }S_{C}^{4}\left( X_{i}^{\prime }\beta
\right) \right] <\infty $ for a neighborhood $\mathcal{N}\left( \beta
_{0}\right) $ of $\beta _{0}$.

\noindent\textbf{Assumption 4.} The preliminary estimators $\hat{\beta}$ and
$\hat{\theta}$ are $\sqrt{n}$-consistent, that is, $\sqrt{n}\left\| \hat{
\beta}-\beta _{0}\right\| =O_{p}\left( 1\right) $ and $\sqrt{n}\left\| \hat{
\theta}-\theta _{0}\right\| =O_{p}\left( 1\right) $.

\noindent \textbf{Assumption 5.} The bandwidth satisfies that $0<h_{n}\leq
1/2$ and that as $n\rightarrow \infty $, (i) $h_{n}\rightarrow 0$, (ii) $
nEk_{ni}^{2}\rightarrow \infty $, and (iii) $\left. \left[ \Pr \left(
F\left( W_{i}\right) >1-h_{n}\right) \right] ^{1+c}\right/
Ek_{ni}^{2}\rightarrow 0$ for any $c>0$.

\bigskip

Assumption 1 describes the model and data. Assumptions 2 and 3 impose
smoothness and boundedness conditions on the kernel function $k\left( \cdot
\right) $ and the transformation function $F\left( \cdot \right) $. The
compactness of the kernel's support is assumed to simplify the technical
proofs, and the theoretical results in this paper are still supposed to hold
when using kernels that decay sufficiently fast in the tails. Assumption
3.(iii) is a joint condition on the tail behaviors of $F\left( \cdot \right)
$ and $X_{i}$'s distribution. When $F\left( \cdot \right) $ has a power-type
upper tail, e.g., Pareto($\lambda $) tail such that\footnote{
Here and below, $g\left( t\right) \sim h\left( t\right) $ means that the
ratios $g\left( t\right) \left/ h\left( t\right) \right. $ and $h\left(
t\right) \left/ g\left( t\right) \right. $ are $O\left( 1\right) $ as $
t\rightarrow +\infty $.} $1-F\left( t\right) \sim t^{-\lambda }$ for some $
\lambda >0$, we have $H_{C}\left( t\right) \sim t^{-1}$ and $S_{C}\left(
t\right) \sim t^{-1}$, and Assumption 3.(iii) holds for any distribution of $
X_{i}$. When $F\left( \cdot \right) $ has an exponential-type upper tail,
e.g., Weibull($\lambda $) tail such that $1-F\left( t\right) \sim \exp
\left( -c_{0}t^{\lambda }\right) $ for some $\lambda >0$ and $c_{0}>0$, we
have $H_{C}\left( t\right) \sim t^{\lambda -1}$ and $S_{C}\left( t\right)
\sim t^{\lambda -1}$. If $0<\lambda \leq 1$, Assumption 3.(iii) still holds
for any distribution of $X_{i}$. For example, in the special case of \cite
{andrews1998semiparametric}'s estimator that corresponds to a Laplacian $
F\left( \cdot \right) $ (i.e., $\lambda =1$), Assumption 3.(iii) is
automatically satisfied. If $\lambda >1$, then Assumption 3.(iii) will
require $E\left\Vert X_{i}\right\Vert ^{4\left( \lambda -1\right) }<\infty $
. For example, if $F\left( \cdot \right) $ is the normal CDF, we need $
E\left\Vert X_{i}\right\Vert ^{4}<\infty $, which is already guaranteed by
Assumption 1.(iv). When the tail of $F\left( \cdot \right) $ decays more
rapidly as $1-F\left( t\right) \sim \exp \left( -\exp \left( c_{0}t^{\lambda
}\right) \right) $, a sufficient condition of Assumption 3.(iii) is that the
distribution of $\left\Vert X_{i}\right\Vert $ has a Weibull($\theta $) tail
with $\theta >\lambda $. Assumption 4 imposes the $\sqrt{n}$-consistency of
the preliminary slope estimators. Examples of such estimators are listed in
Subsection \ref{sec:model}.1.

According to Assumption 5, the bandwidth $h_{n}$ is required to go to zero
as the sample size goes to infinity, but the speed of the decline is not
allowed to be excessively fast. To see this, note that $Ek_{ni}^{2}\leq \bar{
k}^{2}\Pr \left( F\left( W_{i}\right) >1-h_{n}\right) $ by Assumption 2. As $
h_{n}$ goes to zero, $\Pr \left( F\left( W_{i}\right) >1-h_{n}\right) $ will
also go to zero. If the speed at which $h_{n}$ declines is so fast that $\Pr
\left( F\left( W_{i}\right) >1-h_{n}\right) $ goes to zero at a rate faster
than $n^{-1}$, we will have $nEk_{ni}^{2}\rightarrow 0$, which violates
Assumption 5.(ii). Another implicit requirement of Assumption 5.(ii) is the
upper unboundedness of $W_{i}$'s support, for if $W_{i}$ has bounded support
from above, $\Pr \left( F\left( W_{i}\right) >1-h_{n}\right) $ will be
exactly zero for sufficiently small $h_{n}$. Assumption 5.(iii) requires the
distribution of $W_{i}$ to be not too thin upper tailed relative to $F\left(
\cdot \right) $, as illustrated by the following example.

\begin{examp}
Suppose $F\left( \cdot \right) $ is the standard Laplacian CDF, as is the
case with $\hat{\mu}^{AS}$. Then, Assumption 5.(iii) holds if

(i) the distribution of $W_{i}$ has a power-type upper tail;

(ii) the distribution of $W_{i}$ has an exponential-type upper tail;

(iii) the distribution of $W_{i}$ has a ``double''-exponential-type but not
excessively thin upper tail, namely, $\Pr \left( W_{i}>t\right) \sim \exp
\left( -\exp \left( c_{0}t^{\lambda }\right) \right) $ with $\lambda <1$.
\end{examp}

\begin{theorem}
\label{Thm:lite}Under Assumptions 1-5, the kernel regression estimator $
\tilde{\mu}$ defined in (\ref{LC}) is consistent and asymptotically normal:
\begin{equation*}
\frac{\sqrt{n}Ek_{ni}}{\sqrt{Ek_{ni}^{2}}}\left( \tilde{\mu}-\mu _{0}-\frac{
EU_{i}D_{i}k_{ni}}{ED_{i}k_{ni}}\right) \rightarrow N\left( 0,\sigma
_{U}^{2}\right) ,
\end{equation*}
where $\sigma _{U}^{2}=EU_{i}^{2}$.
\end{theorem}

Theorem \ref{Thm:lite} generalizes Theorems 1-2 of \cite
{andrews1998semiparametric} by allowing general forms of the transformation
function $F\left( \cdot \right) $. At first sight, introducing a general $
F\left( \cdot \right) $ appears to help little in the choice of $h_{n}$
because the asymptotic bias and variance of $\tilde{\mu}$ are still unknown
functions of $h_{n}$. However, as will be seen below, a particular
data-dependent choice of $F\left( \cdot \right) $ can largely facilitate the
subsequent choice of $h_{n}$.

\subsection{Choice of the transformation $F\left( \cdot \right) $}

A natural choice of $F\left( \cdot \right) $ would be the CDF of $
W_{i}=X_{i}^{\prime }\beta _{0}$, under which $F\left( W_{i}\right) $
follows a uniform distribution and several assumptions of Theorem \ref
{Thm:lite} have a simpler form. For example, Assumption 3.(iii) will be
automatically fulfilled in this case as long as $W_{i}$ is continuously
distributed, and Assumption 3 is implied by

\bigskip

\noindent\textbf{Assumption 3'.} $W_{i}$ is continuously distributed with
right-unbounded support. In addition, $W_{i}$'s PDF $f_{W}\left( \cdot
\right) $ is absolutely continuous with a derivative that is bounded almost
everywhere.

\bigskip

Assumption 3' presumes that the underlying data is continuous, as in
standard kernel methods. When encountering discrete regressors, the kernel
estimator can be adapted using the frequency-based method or the smoothing
method \citep{racine2004nonparametric} at the cost of more complicated
notations. Assumption 5.(iii) will also be automatically fulfilled for $
F\left( \cdot \right) =F_{W}\left( \cdot \right) $ because in this case $\Pr
\left( F\left( W_{i}\right) >1-h_{n}\right) =h_{n}$ and
\begin{equation}
Ek_{ni}^{r}=\int_{0}^{1}k^{r}\left( \frac{1-t}{h_{n}}\right)
dt=h_{n}\int_{0}^{1\left/ h_{n}\right. }k^{r}\left( s\right)
ds=h_{n}\int_{0}^{1}k^{r}\left( s\right) ds  \label{knir}
\end{equation}
for any $r>0$. Therefore, Assumption 5 is simplified to $h_{n}\rightarrow 0$
and $nh_{n}\rightarrow \infty $ as $n\rightarrow \infty $.

Now, consider the asymptotic bias of $\tilde{\mu}$, $EU_{i}D_{i}k_{ni}\left/
ED_{i}k_{ni}\right. $, under $F\left( \cdot \right) =F_{W}\left( \cdot
\right) $. Define
\begin{equation}
G\left( t\right) =E\left[ U_{i}\left| D_{i}=1,F_{W}\left( W_{i}\right)
=t\right. \right] =E\left[ U_{i}\left| \varepsilon _{i}<F_{W}^{-1}\left(
t\right) \right. \right]  \label{G(t)}
\end{equation}
for $t\in \left( 0,1\right) $, and define $G\left( 1\right)
=\lim_{t\rightarrow 1}G\left( t\right) =0$.

\begin{lemm}
Let $F\left( \cdot \right) =F_{W}\left( \cdot \right) =\Pr \left( W_{i}\leq
\cdot \right) $. Suppose $G\left( t\right) $ is continuously differentiable
over $t\in \left( 1-\delta ,1\right] $ for a small $\delta >0$, with its
derivative function $g\left( t\right) $ being bounded above over $t\in
\left( 1-\delta ,1\right] $. Then, we have
\begin{equation*}
\frac{EU_{i}D_{i}k_{ni}}{ED_{i}k_{ni}}=-\frac{\kappa _{1}g\left( 1\right) }{
\kappa _{0}}h_{n}+o\left( h_{n}\right) ,
\end{equation*}
where $\kappa _{r}=\int_{0}^{1}t^{r}k\left( t\right) dt$.
\end{lemm}

Since by a calculation
\begin{equation*}
g\left( t\right) =\frac{dG\left( t\right) }{dt}=\frac{f_{\varepsilon }\left(
F_{W}^{-1}\left( t\right) \right) }{f_{W}\left( F_{W}^{-1}\left( t\right)
\right) }\cdot \frac{E\left[ U_{i}\left\vert \varepsilon
_{i}=F_{W}^{-1}\left( t\right) \right. \right] -G\left( t\right) }{
F_{\varepsilon }\left( F_{W}^{-1}\left( t\right) \right) },
\end{equation*}
it can be seen that the finiteness of $g\left( 1\right) =\lim_{t\rightarrow
1}g\left( t\right) $ assumed by the Lemma essentially requires the selection
index $W_{i}$ to have a heavier upper tail than the selection disturbance $
\varepsilon _{i}$. The relatively heavy tail of $W_{i}$ is also required by $
\hat{\mu}^{AS}$, as illustrated by the examples of \cite
{andrews1998semiparametric}.

\begin{cor}
Let $F\left( \cdot \right) =F_{W}\left( \cdot \right) =\Pr \left( W_{i}\leq
\cdot \right) $. Suppose (i) Assumptions 1-5 hold, (ii) the assumption of
the Lemma holds, and (iii) $nh_{n}^{3}=O\left( 1\right) $. Then, the kernel
regression estimator $\tilde{\mu}$ defined in (\ref{LC})\ is consistent and
asymptotically normal:
\begin{equation*}
\sqrt{nh_{n}}\left( \tilde{\mu}-\mu _{0}+\frac{\kappa _{1}g\left( 1\right) }{
\kappa _{0}}h_{n}\right) \rightarrow N\left( 0,\frac{\chi _{0}\sigma _{U}^{2}
}{\kappa _{0}^{2}}\right) ,
\end{equation*}
where $\chi _{0}=\int_{0}^{1}k^{2}\left( t\right) dt$.
\end{cor}

The corollary shows that, by setting $F\left( \cdot \right) =F_{W}\left(
\cdot \right) $, the asymptotic bias and variance of the kernel estimator
become explicit functions of $h_{n}$, based on which we can select a
theoretically optimal bandwidth. In practice, however, $F_{W}\left( \cdot
\right) $ is unknown and must be estimated beforehand. A simple estimator
for $F_{W}\left( \cdot \right) $ is the empirical CDF of $\hat{W}
_{i}=X_{i}^{\prime }\hat{\beta}$. With this specific choice, the kernel
estimator becomes
\begin{equation}
\hat{\mu}=\frac{\displaystyle\sum_{i=1}^{n}\left( Y_{i}-Z_{i}^{\prime }\hat{
\theta}\right) D_{i}k\left( \frac{1-\hat{F}_{n}\left( \hat{W}_{i}\right) }{
h_{n}}\right) }{\displaystyle\sum_{i=1}^{n}D_{i}k\left( \frac{1-\hat{F}
_{n}\left( \hat{W}_{i}\right) }{h_{n}}\right) },  \label{FLC}
\end{equation}
where $\hat{F}_{n}\left( \hat{W}_{i}\right) =\frac{1}{n-1}\sum_{j\neq
i}1\left\{ \hat{W}_{j}\leq \hat{W}_{i}\right\} $. To control the estimation
error induced by the empirical CDF, the existing assumptions should be
strengthened, and one new assumption should be imposed. Partition $
X_{i}=\left( X_{i1},X_{i,(-1)}^{\prime }\right) ^{\prime }$, $\beta
_{0}=\left( \beta _{01},\beta _{0,(-1)}^{\prime }\right) ^{\prime }$, and $
\hat{\beta}=\left( \hat{\beta}_{1},\hat{\beta}_{(-1)}^{\prime }\right)
^{\prime }$, with $X_{i1}$, $\beta _{01}$, and $\hat{\beta}_{1}$\ being the
respective first components.

\bigskip

\noindent\textbf{Assumption 1'.} Assumption 1 holds and $E\left\|
X_{i}\right\| ^{6}<\infty $; $\beta _{01}\neq 0$. Without loss of
generality, set $\beta _{01}=1$ to achieve scale normalization of $\beta
_{0} $.

\noindent \textbf{Assumption 2'.} Assumption 2 holds and $k\left( \cdot
\right) $ is six times continuously differentiable over $\left[ 0,\infty
\right) $ with all of its derivatives bounded above.

\noindent\textbf{Assumption 4'.} Assumption 4 holds and $\hat{\beta}_{1}=1$.

\noindent \textbf{Assumption 5'.} The bandwidth $h_{n}\in \left( 0,1/2\right]
$ satisfies (i) $h_{n}\rightarrow 0$, (ii) $nh_{n}^{3}=O\left( 1\right) $,
and (iii) $nh_{n}^{13/5}\rightarrow \infty $, as $n\rightarrow \infty $.

\noindent \textbf{Assumption 6.} The PDF of $W_{i}$, $f_{W}\left( \cdot
\right) $, and the conditional PDF of $W_{i}$ given $X_{i,(-1)}$, $
f_{W\left| X_{(-1)}\right. }\left( \cdot \left| X_{i,(-1)}\right. \right) $,
satisfies that (i) $\left. f_{W}\left( F_{W}^{-1}\left( 1-h_{n}\right)
-n^{-1/10}h_{n}^{-1/18}\right) \right/ h_{n}^{2/3}\rightarrow 0$, (ii) $
f_{W\left| X_{(-1)}\right. }\left( \cdot \left| X_{i,(-1)}\right. \right) $
is uniformly bounded by $\overline{f_{W\left| X_{(-1)}\right. }}$, and (iii)
there exists a large constant $C$ such that $f_{W\left| X_{(-1)}\right.
}\left( w\left| X_{i,(-1)}\right. \right) $ declines monotonically for $w>C$
almost surely.

\bigskip

Since $\beta _{0}$ is identified only up to scale, the normalization $\beta
_{01}=\hat{\beta}_{1}=1$ is postulated by most semiparametric estimators for
the binary response model. Assumption 6, which is comparable to Assumption A
of \cite{schafgans2002intercept}, is imposed to address the
non-differentiable indicator function contained in the empirical CDF.
Because the expectation of the generalized derivative of the indicator
function equals the value of the PDF, Assumption 6 involves restrictions on
the PDF and conditional PDF of $W_{i}$. Assumption 6.(i) relates to the
upper tail behavior of the distribution of $W_{i}$. Under the assumption of $
nh_{n}^{13/5}\rightarrow \infty $ that implies $n^{-1/10}h_{n}^{-1/18}
\rightarrow 0$, it can be shown that Assumption 6.(i) is satisfied if $W_{i}$
has a power- or exponential-type upper tail.

\begin{theorem}
\label{Thm:FLC}Under Assumptions 1'-5' and 6, the kernel regression
estimator $\hat{\mu}$ defined in (\ref{FLC}) is consistent and
asymptotically normal:
\begin{equation*}
\sqrt{nh_{n}}\left( \hat{\mu}-\mu _{0}+\frac{\kappa _{1}g\left( 1\right) }{
\kappa _{0}}h_{n}\right) \rightarrow N\left( 0,\frac{\chi _{0}\sigma _{U}^{2}
}{\kappa _{0}^{2}}\right) ,
\end{equation*}
where $\kappa _{r}=\int_{0}^{1}t^{r}k\left( t\right) dt$ and $\chi
_{r}=\int_{0}^{1}t^{r}k^{2}\left( t\right) dt$.
\end{theorem}

Theorem \ref{Thm:FLC} reveals that, under slightly stronger conditions, the
estimation error induced by the empirical CDF is asymptotically negligible;
thus, the asymptotic distribution of $\hat{\mu}$ is the same as if the true
CDF $F_{W}\left( \cdot \right) $ were known.

\subsection{Bandwidth selection}

An important implication of Theorem \ref{Thm:FLC} is that, under the
specific choice of $F\left( \cdot \right) =F_{W}\left( \cdot \right) $, the
kernel estimator for the intercept follows a standard asymptotic
distribution as an ordinary kernel regression estimator at the boundary
point. As a result, we can borrow approaches of bandwidth selection from the
nonparametric regression literature \citep[see, e.g.,][]{li2007nonparametric}
. However, the widely used cross-validation method based on the integrated
mean squared error (MSE) criteria takes into account the global performance
of the regression function estimation and is thus not suitable for the
problem at hand. A closely related problem is the choice of bandwidth for
the nonparametric regression discontinuity estimator, which is the
difference between two regression estimators evaluated at boundary points. I
follow the plug-in method proposed by \cite{imbens2012optimal} and suggest a
data-dependent bandwidth selection procedure that is tailored to $\hat{\mu}$.

By Theorem \ref{Thm:FLC}, we know that
\begin{equation*}
\text{Abias}\left( \hat{\mu}\right) =\left( -\frac{\kappa _{1}g\left(
1\right) }{\kappa _{0}}\right) h_{n},\text{ \ Avar}\left( \hat{\mu}\right)
=\left( \frac{\chi _{0}\sigma _{U}^{2}}{\kappa _{0}^{2}}\right) \frac{1}{
nh_{n}},
\end{equation*}
where $g\left( 1\right) =\left. \left. dG\left( t\right) \right/
dt\right\vert _{t=1}$ with $G\left( t\right) $ defined in (\ref{G(t)}), and $
\sigma _{U}^{2}=Var\left( U_{i}\right) $. Provided $g\left( 1\right) \neq 0$
, the optimal bandwidth for $\hat{\mu}$ is defined as a minimizer of its
asymptotic MSE:
\begin{eqnarray*}
h_{opt} &=&\arg \min_{h}\left\{ \text{AMSE}_{h}\left( \hat{\mu}\right)
=\left( \frac{\kappa _{1}g\left( 1\right) }{\kappa _{0}}\right)
^{2}h^{2}+\left( \frac{\chi _{0}\sigma _{U}^{2}}{\kappa _{0}^{2}}\right)
\frac{1}{nh}\right\} \\
&=&c_{k}\left( \sigma _{U}^{2}\left/ g^{2}\left( 1\right) \right. \right)
^{1/3}n^{-1/3},
\end{eqnarray*}
where $c_{k}=\left( \chi _{0}\left/ \left( 2\kappa _{1}^{2}\right) \right.
\right) ^{1/3}$ is a functional of the kernel $k\left( \cdot \right) $. If $
g\left( 1\right) =0$, the bias converges to zero faster, allowing for
estimation of the intercept at a faster rate of convergence. However, it is
difficult to exploit the improved convergence rate resulting from this
condition in practice; hence, I focus on the optimal bandwidth given $
g\left( 1\right) \neq 0$.

A natural choice of the estimator for the optimal bandwidth $h_{opt}$ is to
replace $\sigma _{U}^{2}$ and $g\left( 1\right) $ with their consistent
estimators $\hat{\sigma}_{U}^{2}$ and $\hat{g}\left( 1\right) $,
respectively. One potential problem with this choice is that $\hat{g}\left(
1\right) $ may occasionally be very close to zero due to the stochastic
estimation error, even if $g\left( 1\right) \neq 0$. In such cases, the
estimated bandwidth will be imprecisely and unstably large, which may in
turn lead to large finite sample bias of $\hat{\mu}$ because observations
that are far from the boundary will be included in the kernel estimation. To
alleviate this problem, I employ the regularization method
\citep[Subsection
4.1.1]{imbens2012optimal} and propose the following bandwidth estimator:
\begin{equation}
\hat{h}_{opt}=c_{k}\left( \frac{\hat{\sigma}_{U}^{2}}{\hat{g}^{2}\left(
1\right) +3\widehat{Var}\left( \hat{g}\left( 1\right) \right) }\right)
^{1/3}n^{-1/3}.  \label{hat_h}
\end{equation}
By means of regularization, $\hat{h}_{opt}$ will not become infinite even in
the case of $\hat{g}\left( 1\right) =0$. Moreover, the leading term of $E
\left[ 1\left/ \left( \hat{g}^{2}\left( 1\right) +3Var\left( \hat{g}\left(
1\right) \right) \right) \right. -1\left/ \hat{g}^{2}\left( 1\right) \right.
\right] $ cancels out the leading term of $E\left[ 1\left/ \hat{g}^{2}\left(
1\right) \right. -1\left/ g^{2}\left( 1\right) \right. \right] $. Therefore,
the bias of $1\left/ \left( \hat{g}^{2}\left( 1\right) +3Var\left( \hat{g}
\left( 1\right) \right) \right) \right. $ for the reciprocal of $g^{2}\left(
1\right) $ is of lower order than the bias of the naive $1\left/ \hat{g}
^{2}\left( 1\right) \right. $. It remains to construct consistent estimators
for the components of the plug-in bandwidth, namely, $\hat{\sigma}_{U}^{2}$,
$\hat{g}^{2}\left( 1\right) $, and $\widehat{Var}\left( \hat{g}\left(
1\right) \right) $, which is deferred to the next section for expositional
convenience.

\section{Local linear estimation}

\label{sec:LL}Theorem \ref{Thm:FLC} finds that the kernel regression
estimator $\hat{\mu}$ for the intercept suffers from a boundary effect in
the sense that its order of bias, $O\left( h_{n}\right) $, is larger than
that of the nonparametric regression estimator in the interior, which is
typically $O\left( h_{n}^{2}\right) $. This is simply because $\hat{\mu}$ is
in nature a Nadaraya-Watson or local constant estimator at the boundary. In
this section, I resort to the local polynomial regression method
\citep{fan1996local} for bias reduction. In particular, I focus on the local
linear estimation because of its asymptotic minimax efficiency properties
\citep{cheng1997automatic} and attractive practical performance
\citep{gelman2019why}.

Denote $\hat{F}_{n}\left( \hat{W}_{i}\right) =\frac{1}{n-1}\sum_{j\neq
i}1\left\{ \hat{W}_{j}\leq \hat{W}_{i}\right\} $ as the empirical estimate
of $F_{W}\left( W_{i}\right) $ as before. The local linear estimator $\hat{
\mu}^{L}$ for the intercept is defined via locally weighted least squares
regression:
\begin{equation}
\left( \hat{\mu}^{L},\hat{b}^{L}\right) =\arg \min_{\mu ,b}\sum_{i=1}^{n}
\left[ Y_{i}-Z_{i}^{\prime }\hat{\theta}-\mu -\left( \hat{F}_{n}\left( \hat{W
}_{i}\right) -1\right) b\right] ^{2}D_{i}k\left( \frac{1-\hat{F}_{n}\left(
\hat{W}_{i}\right) }{h_{n}}\right) .  \label{FLL}
\end{equation}
To establish the asymptotic properties of $\hat{\mu}^{L}$, Assumption 5'
must be modified to accommodate the local linear case.

\bigskip

\noindent\textbf{Assumption 5''.} The bandwidth $h_{n}\in \left( 0,1/2\right]
$ satisfies (i) $h_{n}\rightarrow 0$, (ii) $nh_{n}^{5}=O\left( 1\right) $,
and (iii) $nh_{n}^{3}\rightarrow \infty $, as $n\rightarrow \infty $.

\begin{theorem}
\label{Thm:FLL}Suppose Assumptions 1'-4', 5\textquotedblright , and 6 hold.
In addition, suppose $G\left( t\right) $ given in (\ref{G(t)}) is twice
continuously differentiable over $t\in \left( 1-\delta ,1\right] $ for a
small $\delta >0$, with its first and second derivative functions $g\left(
t\right) $ and $g^{\prime }\left( t\right) $ being bounded above over $t\in
\left( 1-\delta ,1\right] $. Then, the local linear estimator defined in (
\ref{FLL}) is consistent for $\left( \mu _{0},g\left( 1\right) \right) $ and
asymptotically normal:
\begin{equation*}
\left(
\begin{array}{cc}
\sqrt{nh_{n}} &  \\
& \sqrt{nh_{n}^{3}}
\end{array}
\right) \left[ \left(
\begin{array}{c}
\hat{\mu}^{L} \\
\hat{b}^{L}
\end{array}
\right) -\left(
\begin{array}{c}
\mu _{0} \\
g\left( 1\right)
\end{array}
\right) +\left(
\begin{array}{c}
\frac{\left( \kappa _{1}\kappa _{3}-\kappa _{2}^{2}\right) g^{\prime }\left(
1\right) }{2\left( \kappa _{0}\kappa _{2}-\kappa _{1}^{2}\right) }h_{n}^{2}
\\
\frac{\left( \kappa _{0}\kappa _{3}-\kappa _{1}\kappa _{2}\right) g^{\prime
}\left( 1\right) }{2\left( \kappa _{0}\kappa _{2}-\kappa _{1}^{2}\right) }
h_{n}
\end{array}
\right) \right] \rightarrow N\left( 0,\sigma _{U}^{2}\Omega ^{L}\right) ,
\end{equation*}
where $\sigma _{U}^{2}=EU_{i}^{2}$ and
\begin{equation*}
\Omega ^{L}=\frac{1}{\left( \kappa _{0}\kappa _{2}-\kappa _{1}^{2}\right)
^{2}}\left(
\begin{array}{cc}
\kappa _{2}^{2}\chi _{0}+\kappa _{1}^{2}\chi _{2}-2\kappa _{1}\kappa
_{2}\chi _{1} & \kappa _{1}\kappa _{2}\chi _{0}+\kappa _{0}\kappa _{1}\chi
_{2}-\left( \kappa _{0}\kappa _{2}+\kappa _{1}^{2}\right) \chi _{1} \\
\kappa _{1}\kappa _{2}\chi _{0}+\kappa _{0}\kappa _{1}\chi _{2}-\left(
\kappa _{0}\kappa _{2}+\kappa _{1}^{2}\right) \chi _{1} & \kappa
_{1}^{2}\chi _{0}+\kappa _{0}^{2}\chi _{2}-2\kappa _{0}\kappa _{1}\chi _{1}
\end{array}
\right) .
\end{equation*}
\end{theorem}

\medskip

Theorem \ref{Thm:FLL} shows that the local linear estimation of the
intercept automatically eliminates the boundary effect, and its asymptotic
bias is of the same order as that in the interior. More interestingly, the
local linear procedure generates a consistent estimate of $g\left( 1\right) $
as a byproduct, which is a key ingredient of the optimal bandwidth for the
kernel estimator $\hat{\mu}$. There would be no technical difficulty in
extending the results of Theorem \ref{Thm:FLL} to local polynomial
estimation with higher order, except for more complicated notation and more
involved mathematical derivations.

\subsection{Bandwidth selection}

Provided $g^{\prime }\left( 1\right) \neq 0$, the optimal bandwidth for $
\hat{\mu}^{L}$ is analogously defined by minimizing the asymptotic MSE of $
\hat{\mu}^{L}$:
\begin{eqnarray*}
h_{opt}^{L} &=&\arg \min_{h}\left\{ \text{AMSE}_{h}\left( \hat{\mu}
^{L}\right) =\left[ \frac{\left( \kappa _{1}\kappa _{3}-\kappa
_{2}^{2}\right) g^{\prime }\left( 1\right) }{2\left( \kappa _{0}\kappa
_{2}-\kappa _{1}^{2}\right) }\right] ^{2}h^{4}+\frac{\left( \kappa
_{2}^{2}\chi _{0}+\kappa _{1}^{2}\chi _{2}-2\kappa _{1}\kappa _{2}\chi
_{1}\right) \sigma _{U}^{2}}{\left( \kappa _{0}\kappa _{2}-\kappa
_{1}^{2}\right) ^{2}}\frac{1}{nh}\right\} \\
&=&c_{k}^{L}\left( \sigma _{U}^{2}\left/ \left[ g^{\prime }\left( 1\right)
\right] ^{2}\right. \right) ^{1/5}n^{-1/5},
\end{eqnarray*}
where
\begin{equation}
c_{k}^{L}=\left[ \frac{\kappa _{2}^{2}\chi _{0}+\kappa _{1}^{2}\chi
_{2}-2\kappa _{1}\kappa _{2}\chi _{1}}{\left( \kappa _{1}\kappa _{3}-\kappa
_{2}^{2}\right) ^{2}}\right] ^{1/5}.  \label{cL_k}
\end{equation}
As in Subsection \ref{sec:KRE}.3, I adopt the regularization method and
propose the following estimator for $h_{opt}^{L}$:
\begin{equation}
\hat{h}_{opt}^{L}=c_{k}^{L}\left( \frac{\hat{\sigma}_{U}^{2}}{\left[ \hat{g}
^{\prime }\left( 1\right) \right] ^{2}+3\widehat{Var}\left( \hat{g}^{\prime
}\left( 1\right) \right) }\right) ^{1/5}n^{-1/5}.  \label{hat_h_L}
\end{equation}

To implement the plug-in bandwidth selection procedures (\ref{hat_h}) and (
\ref{hat_h_L}), one must estimate the limits of the derivative functions, $
g\left( 1\right) $ and $g^{\prime }\left( 1\right) $, the regularization
terms, $Var\left( \hat{g}\left( 1\right) \right) $ and $Var\left( \hat{g}
^{\prime }\left( 1\right) \right) $, and the variance of the disturbance, $
\sigma _{U}^{2}$. I first construct $\hat{g}\left( 1\right) $ and $\hat{g}
^{\prime }\left( 1\right) $ by fitting a quadratic function to the
observations near the boundary:
\begin{equation}
\left( \hat{\mu}^{Q},\hat{g}\left( 1\right) ,\hat{g}^{\prime }\left(
1\right) \right) =\arg \min_{b_{0},b_{1},b_{2}}\sum_{i=1}^{n}\left[
Y_{i}-Z_{i}^{\prime }\hat{\theta}-\sum_{r=0}^{2}\frac{\left( \hat{F}
_{n}\left( \hat{W}_{i}\right) -1\right) ^{r}}{r!}b_{r}\right]
^{2}D_{i}k\left( \frac{1-\hat{F}_{n}\left( \hat{W}_{i}\right) }{h_{1n}}
\right) ,  \label{LQE}
\end{equation}
where $h_{1n}$ is a pilot bandwidth. Similar to Theorem \ref{Thm:FLL},\ one
can show that under $nh_{1n}^{7}=O\left( 1\right) $, $nh_{1n}^{5}\rightarrow
\infty $, and the regularity conditions, the local quadratic estimator is
consistent and asymptotically normal:
\begin{equation*}
\left(
\begin{array}{ccc}
\sqrt{nh_{1n}} &  &  \\
& \sqrt{nh_{1n}^{3}} &  \\
&  & \sqrt{nh_{1n}^{5}}
\end{array}
\right) \left[ \left(
\begin{array}{c}
\hat{\mu}^{Q} \\
\hat{g}\left( 1\right) \\
\hat{g}^{\prime }\left( 1\right)
\end{array}
\right) -\left(
\begin{array}{c}
\mu _{0} \\
g\left( 1\right) \\
g^{\prime }\left( 1\right)
\end{array}
\right) +\left(
\begin{array}{c}
B_{1}^{Q}g^{\prime \prime }\left( 1\right) h_{1n}^{3} \\
B_{2}^{Q}g^{\prime \prime }\left( 1\right) h_{1n}^{2} \\
B_{3}^{Q}g^{\prime \prime }\left( 1\right) h_{1n}
\end{array}
\right) \right] \rightarrow N\left( 0,\sigma _{U}^{2}\Omega ^{Q}\right) .
\end{equation*}
As a result, the regularization terms can be estimated by
\begin{equation*}
\widehat{Var}\left( \hat{g}\left( 1\right) \right) =\frac{1}{nh_{1n}^{3}}
\hat{\sigma}_{U}^{2}\Omega _{22}^{Q},\text{ \ }\widehat{Var}\left( \hat{g}
^{\prime }\left( 1\right) \right) =\frac{1}{nh_{1n}^{5}}\hat{\sigma}
_{U}^{2}\Omega _{33}^{Q},
\end{equation*}
where
\begin{eqnarray}
\Omega _{22}^{Q} &=&\frac{\left\{
\begin{array}{c}
\left( \kappa _{1}\kappa _{4}-\kappa _{2}\kappa _{3}\right) ^{2}\chi
_{0}-2\left( \kappa _{1}\kappa _{4}-\kappa _{2}\kappa _{3}\right) \left(
\kappa _{0}\kappa _{4}-\kappa _{2}^{2}\right) \chi _{1} \\
+\left[ \left( \kappa _{0}\kappa _{4}-\kappa _{2}^{2}\right) ^{2}+2\left(
\kappa _{1}\kappa _{4}-\kappa _{2}\kappa _{3}\right) \left( \kappa
_{0}\kappa _{3}-\kappa _{1}\kappa _{2}\right) \right] \chi _{2} \\
-\left( \kappa _{0}\kappa _{4}-\kappa _{2}^{2}\right) \left( \kappa
_{0}\kappa _{3}-\kappa _{1}\kappa _{2}\right) \chi _{3}+\left( \kappa
_{0}\kappa _{3}-\kappa _{1}\kappa _{2}\right) ^{2}\chi _{4}
\end{array}
\right\} }{\left( \kappa _{0}\kappa _{2}\kappa _{4}-\kappa _{0}\kappa
_{3}^{2}-\kappa _{1}^{2}\kappa _{4}+2\kappa _{1}\kappa _{2}\kappa
_{3}-\kappa _{2}^{3}\right) ^{2}},  \label{omigaQ_22} \\
\Omega _{33}^{Q} &=&\frac{\left\{
\begin{array}{c}
\left( \kappa _{1}\kappa _{3}-\kappa _{2}^{2}\right) ^{2}\chi _{0}-2\left(
\kappa _{1}\kappa _{3}-\kappa _{2}^{2}\right) \left( \kappa _{0}\kappa
_{3}-\kappa _{1}\kappa _{2}\right) \chi _{1} \\
+\left[ \left( \kappa _{0}\kappa _{3}-\kappa _{1}\kappa _{2}\right)
^{2}+2\left( \kappa _{1}\kappa _{3}-\kappa _{2}^{2}\right) \left( \kappa
_{0}\kappa _{2}-\kappa _{1}^{2}\right) \right] \chi _{2} \\
-\left( \kappa _{0}\kappa _{3}-\kappa _{1}\kappa _{2}\right) \left( \kappa
_{0}\kappa _{2}-\kappa _{1}^{2}\right) \chi _{3}+\left( \kappa _{0}\kappa
_{2}-\kappa _{1}^{2}\right) ^{2}\chi _{4}
\end{array}
\right\} }{\left( \kappa _{0}\kappa _{2}\kappa _{4}-\kappa _{0}\kappa
_{3}^{2}-\kappa _{1}^{2}\kappa _{4}+2\kappa _{1}\kappa _{2}\kappa
_{3}-\kappa _{2}^{3}\right) ^{2}}.  \label{omigaQ_33}
\end{eqnarray}
Last, I estimate the variance of the disturbance by
\begin{equation}
\hat{\sigma}_{U}^{2}=\frac{\displaystyle\sum_{i=1}^{n}\left( Y_{i}-\hat{\mu}
^{Q}-Z_{i}^{\prime }\hat{\theta}\right) ^{2}D_{i}k\left( \frac{1-\hat{F}
_{n}\left( \hat{W}_{i}\right) }{h_{2n}}\right) }{\displaystyle
\sum_{i=1}^{n}D_{i}k\left( \frac{1-\hat{F}_{n}\left( \hat{W}_{i}\right) }{
h_{2n}}\right) },  \label{sigma2_hat}
\end{equation}
where $\hat{\mu}^{Q}$ is the initial estimate of $\mu _{0}$ given in (\ref
{LQE}) and $h_{2n}$ is another pilot bandwidth. Following Theorem \ref
{Thm:FLC}, it can be shown that
\begin{equation*}
\hat{\sigma}_{U}^{2}\overset{p}{\rightarrow }E\left[ \left. \left( Y_{i}-\mu
_{0}-Z_{i}^{\prime }\theta _{0}\right) ^{2}\right\vert D_{i}=1,F_{W}\left(
W_{i}\right) =1\right] =EU_{i}^{2}=\sigma _{U}^{2}.
\end{equation*}
For the pilot bandwidths, simply setting
\begin{equation*}
h_{1n}=n^{-1/7},\text{ }h_{2n}=n^{-1/3}
\end{equation*}
is sufficient to ensure the consistency of $\hat{h}_{opt}$ for $h_{opt}$ and
of $\hat{h}_{opt}^{L}$ for $h_{opt}^{L}$. In practice, the suggested
bandwidth selection algorithm is fairly robust to the choice of pilot
bandwidth, which is not surprising given the presence of the power $1/3$ or $
1/5$ in the expressions for the optimal bandwidths.

\section{Simulation}

\label{sec:simulation}This section examines the finite sample properties of
the kernel regression estimator $\hat{\mu}$ defined in (\ref{FLC}) and the
local linear estimator $\hat{\mu}^{L}$ defined in (\ref{FLL}), in comparison
with the parametric two-step estimator \citep{heckman1979sample} and the
semiparametric identification-at-infinity estimators
\citep{heckman1990varieties, andrews1998semiparametric}. The parametric
two-step procedure implements probit estimation for the selection equation
in the first step and least squares estimation for the outcome equation with
a correction term using the uncensored observations in the second step. This
approach is commonly applied in empirical studies due to its computational
ease. However, it is likely to be inconsistent when the true distribution of
the model disturbance is nonnormal. In contrast, the consistency of \cite
{heckman1990varieties}'s estimator (henceforth the Heckman estimator) and
\cite{andrews1998semiparametric}'s estimator (henceforth the AS estimator)
does not rely on parametric specification of the disturbance distribution,
but it is difficult to choose an appropriate smoothing parameter for these
estimators. To investigate the robustness of their practical performance, a
wide range of smoothing parameters is considered. Following the literature,
the choices considered are based on the percentage of uncensored
observations used in the estimation. Specifically, the smoothing parameter
takes the values of various quantiles of the selection linear index in the
uncensored subsample.

The simulation setting mainly follows \cite{schafgans2004finite}, and the
data generating process is
\begin{eqnarray*}
D_{i} &=&1\left\{ c_{0}+X_{1i}+X_{2i}>\varepsilon _{i}\right\} , \\
Y_{i}^{\ast } &=&\mu _{0}+U_{i}, \\
Y_{i} &=&Y_{i}^{\ast }D_{i},\text{ }i=1,2,\cdots ,n,
\end{eqnarray*}
where $\mu _{0}=0$, $X_{1i}$ follows the standard normal distribution, $
X_{2i}$ follows the standardized (zero-mean, unit-variance) Student's $t$
distribution with three degrees of freedom, $\varepsilon _{i}$ and $U_{i}$
are zero-mean random variables described below, and only $\left(
Y_{i},D_{i},X_{1i},X_{2i}\right) $ is observed. In the simulation, the
outcome equation does not contain any nonconstant regressors, implying that
the intercept $\mu _{0}$ of primary concern represents the population mean
of the latent outcome $Y_{i}^{\ast }$. Different designs are constructed by
varying the disturbance distribution and the value of $c_{0}$. $\varepsilon
_{i}$ follows three different distributions, namely, the standard normal
distribution, the standardized Student's $t$ distribution with three degrees
of freedom, and the standardized chi-square distribution with three degrees
of freedom. $U_{i}$ is generated by $U_{i}=\varepsilon _{i}+e_{i}$, where $
e_{i}$ is a standard normal random variable independent of $\varepsilon _{i}$
. The constant $c_{0}$ controls for the amount of censoring. In the
benchmark design, $c_{0}=0$, producing approximately 50\% censoring.
Different values of $c_{0}$ are chosen so that $\Pr \left(
c_{0}+X_{1i}+X_{2i}>\varepsilon _{i}\right) $ is equal to 0.8 and 0.2,
corresponding to 20\% and 80\% proportions of zero observations,
respectively. The sample size $n$ is set to 250, 1000, 4000, and the
simulation is replicated 1000 times for each design.

Before calculating the semiparametric estimators for $\mu _{0}$, a
distribution-free estimate for the selection equation is necessary. Since
the simulation results of \cite{schafgans2004finite} show little sensitivity
to the particular choice of this estimate, I employ the average derivative
estimation \citep{powell1989semiparametric} for computational convenience.
When implementing the Heckman and AS estimators, the smoothing parameter is
equal to the 0.99, 0.95, 0.9, 0.8, 0.7, and 0.5 quantiles of $\hat{W}_{i}=
\hat{\beta}_{1}X_{1i}+\hat{\beta}_{2}X_{2i}$ in the uncensored subsample,
corresponding to 1\%, 5\%, 10\%, 20\%, 30\%, and 50\% uncensored
observations used in the estimation. The value of the smoothing parameter
declines with the proportion of uncensored observations. For consistency,
the smoothing parameter is required to approach infinity as $n$ goes to
infinity such that the estimation is based on only the observations for
which $\Pr \left( \left. D_{i}=1\right\vert X_{i}\right) $ is close to one
and in the limit is equal to one. Following the suggestion of \cite
{andrews1998semiparametric}, the smoothed function in the AS estimator is
\begin{equation*}
s\left( w\right) =\left\{
\begin{array}{c}
0 \\
1-\exp \left( -\frac{w}{b-w}\right) \\
1
\end{array}
\begin{array}{l}
\text{for }w\leq 0, \\
\text{for }0<w\leq b, \\
\text{for }w>b,
\end{array}
\right.
\end{equation*}
where $b$ is set equal to 1. Note that the AS estimator with $b=0$ is
equivalent to the Heckman estimator. For the proposed kernel-type
estimators, the bandwidth is chosen by the plug-in algorithm given in
Subsection \ref{sec:LL}.1. The commonly used Gaussian and Epanechnikov
kernel functions are applied. However, the Gaussian kernel is not compactly
supported, and the Epanechnikov kernel is not smooth. Therefore, I also
consider the polynomial and polyweight kernels, both of degree seven, which
possess sufficient smoothness required by Assumption 2'. The definition of
these kernel functions and their relevant functionals are presented in Table
\ref{table:kernel}.

\begin{sidewaystable}[ptb]
\caption{Several kernel functions and their relevant functionals}
\bigskip
\label{table:kernel}
\resizebox{\linewidth}{!}{
\begin{tabular}{C{2.2cm}ccc}
\hline\hline
Kernel & Definition & $\kappa _{r}=\int_{0}^{\infty }t^{r}k\left( t\right) dt
$, $r\in \mathbb{N}$ & $\chi _{r}=\int_{0}^{\infty }t^{r}k^{2}\left(
t\right) dt$, $r\in \mathbb{N}$ \\ \hline
Gaussian & $\displaystyle k\left( t\right) =\frac{1}{\sqrt{2\pi }}\exp
\left( -\frac{t^{2}}{2}\right)1\left\{t\geq 0\right\} $ & $\displaystyle\sqrt{\frac{2^{r-2}}{\pi }}
\gamma \left( \frac{r+1}{2}\right) $ & $\displaystyle\frac{1}{4\pi }\gamma
\left( \frac{r+1}{2}\right) $ \\
Epanechnikov & $\displaystyle k\left( t\right) =\left\{ \QATOP{1-t^{2}\text{
if }0\leq t\leq 1}{0\text{ \ if }t>1}\right. $ & $\displaystyle\frac{2}{
\left( r+1\right) \left( r+3\right) }$ & $\displaystyle\frac{8}{\left(
r+1\right) \left( r+3\right) \left( r+5\right) }$ \\
Polynomial of degree 7 & $\displaystyle k\left( t\right) =\left\{ \QATOP{
\left( 1-t\right) ^{7}\text{ if }0\leq t\leq 1}{\text{ \ }0\text{\ \ \ if }
t>1}\right. $ & $\displaystyle B\left( r+1,8\right) =\frac{7!r!}{\left(
r+8\right) !}$ & $\displaystyle B\left( r+1,15\right) =\frac{14!r!}{\left(
r+15\right) !}$ \\
Polyweight of degree 7 & $\displaystyle k\left( t\right) =\left\{ \QATOP{
\left( 1-t^{2}\right) ^{7}\text{ if }0\leq t\leq 1}{\text{ \ }0\text{\ \ \ \
if }t>1}\right. $ & $\displaystyle\frac{1}{2}B\left( \frac{r+1}{2},8\right) =
\frac{2^{7}7!\left( r-1\right) !!}{\left( r+15\right) !!}$ & $\displaystyle
\frac{1}{2}B\left( \frac{r+1}{2},15\right) =\frac{2^{14}14!\left( r-1\right)
!!}{\left( r+29\right) !!}$ \\ \hline
Kernel & $c_{k}^{L}$ in (\ref{cL_k}) & $\Omega _{22}^{Q}$ in (\ref{omigaQ_22}
) & $\Omega _{33}^{Q}$ in (\ref{omigaQ_33}) \\ \hline
Gaussian & $\left[ \frac{\left( \pi +1-2\sqrt{2}\right) \sqrt{\pi }}{\left(
4-\pi \right) ^{2}}\right] ^{1/5}\approx \allowbreak 1.\allowbreak 259$ & $
\frac{\left( 4\pi +11-12\sqrt{2}\right) \sqrt{\pi }}{8\left( \pi -3\right)
^{2}}\approx \allowbreak 72.\allowbreak 89$ & $\frac{3\pi ^{2}-\left( 16-4
\sqrt{2}\right) \pi +\left( 44-24\sqrt{2}\right) }{16\sqrt{\pi }\left( \pi
-3\right) ^{2}}\approx \allowbreak 12.\allowbreak 62$ \\
Epanechnikov & $3.200$ & $4913.0$ & $6327.8$ \\
Polynomial of degree 7 & $8.175$ & $8477.5$ & $48857.0$ \\
Polyweight of degree 7 & $5.396$ & $8645.0$ & $29139.5$ \\ \hline\hline
\end{tabular}
}
{\small {Note: $r!!$, $\gamma \left( r\right) $, and $B\left(
r_{1},r_{2}\right) $ denote the double factorial, gamma function, and beta
function, respectively.} }
\end{sidewaystable}

The summary statistics for the simulation are the estimators' bias, standard
deviation (SD), root mean squared error (RMSE) ratio, and rejection rate of
the t test, over 1000 replications. The RMSE can be calculated from the bias
and SD and is thus omitted from the tables. Instead, the RMSE ratio defined
by the RMSE of the semiparametric estimators over that of the parametric
two-step estimator is reported. Although the RMSE (ratio) is the primary
criterion used to compare consistent estimators, it is useless if there are
both consistent and inconsistent estimators. An inconsistent estimator
having a small RMSE implies that the estimator is highly concentrated in a
narrow interval centered at a biased value, consequently leading to
incorrect inference. As a complement to the RMSE (ratio), I consider the
simulated probability of rejecting the null hypothesis $H_{0}:\mu _{0}=0$
against $H_{1}:\mu _{0}\neq 0$ at a 5\% level of significance using t tests
(or, more precisely, z tests), based on the asymptotic variances given in
\cite{heckman1979sample}, \cite{schafgans2002intercept}, \cite
{andrews1998semiparametric}, and Theorems \ref{Thm:FLC}-\ref{Thm:FLL}.

Tables \ref{table:normal}-\ref{table:chi2} report the simulation results
when the model disturbance follows a normal distribution, a $t\left(
3\right) $ distribution that is symmetric but fat-tailed, and a $\chi
^{2}\left( 3\right) $ distribution that is skewed, respectively, under
approximately 50\% censoring. Table \ref{table:normal} shows that, under
normal disturbance, the parametric two-step estimator is asymptotically
unbiased, as expected, and converges at a $\sqrt{n}$ rate (its SD halves
when the sample size quadruples). In contrast, all the considered
semiparametric estimators converge at slower than $\sqrt{n}$ rates, as their
RMSE ratios increase with the sample size. For the Heckman and AS
estimators, the bias increases and the SD decreases when the proportion of
uncensored observations used in the estimation increases or, equivalently,
the smoothing parameter decreases. The optimal smoothing parameter in terms
of RMSE depends upon the estimator, the sample size, and, by comparing
across Tables \ref{table:normal}-\ref{table:chi2}, the disturbance
distribution. If the smoothing parameter is improperly chosen, the RMSE may
be several times larger than the smallest RMSE, and the rejection rate may
be far larger than the specified level of significance. Moreover, the
optimal smoothing parameter in terms of RMSE usually disagrees with that in
terms of the rejection rate. For instance, in the case of $n=1000$ for the
Heckman estimator and the case of $n=4000$ for the AS estimator, the optimal
smoothing parameter in terms of RMSE would use 30\% uncensored observations
in the estimation, but the corresponding simulated rejection rates are more
than twice the real level. In these two cases, a more sensible choice would
be to use 20\% uncensored observations, sacrificing a little RMSE but
leading to a rejection rate very close to 0.05. In practice, however, the
RMSE and rejection rate are not known; therefore, we are never aware of
whether we have made a good choice.

\begin{sidewaystable}[ptb]
\caption{Simulation results when $\varepsilon _{i}\sim N\left(0,1\right) $}
\medskip
\label{table:normal}
\resizebox{\linewidth}{!}{
\begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}}
\hline\hline
$\Pr \left( Y_{i}=0\right) =0.5$ &  & \multicolumn{4}{c}{$n=250$} &  &
\multicolumn{4}{c}{$n=1000$} &  & \multicolumn{4}{c}{$n=4000$} \\
\cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16}
Estimator &  & Bias & SD & RMSE ratio & Rejection rate &  & Bias & SD & RMSE
ratio & Rejection rate &  & Bias & SD & RMSE ratio & Rejection rate \\ \hline
\multicolumn{16}{l}{\cite{heckman1979sample}'s parametric two-step estimator}
\\
&  & -0.004 & 0.186 & 1 & 0.052 &  & -0.002 & 0.093 & 1 & 0.048 &  & 0.000 &
0.044 & 1 & 0.043 \\
\multicolumn{16}{l}{\cite{heckman1990varieties}'s semiparametric estimator
with various smoothing parameters} \\
$1\%$ observations &  & -0.007 & 1.025 & 5.519 & 0.390 &  & 0.028 & 0.599 &
6.444 & 0.128 &  & 0.006 & 0.301 & 6.812 & 0.062 \\
$5\%$ observations &  & -0.003 & 0.549 & 2.955 & 0.125 &  & 0.001 & 0.286 &
3.074 & 0.076 &  & -0.006 & 0.140 & 3.173 & 0.056 \\
$10\%$ observations &  & -0.016 & 0.389 & 2.093 & 0.088 &  & -0.021 & 0.195
& 2.108 & 0.064 &  & -0.016 & 0.096 & 2.209 & 0.062 \\
$20\%$ observations &  & -0.046 & 0.275 & 1.503 & 0.070 &  & -0.049 & 0.136
& 1.558 & 0.059 &  & -0.044 & 0.067 & 1.801 & 0.088 \\
$30\%$ observations &  & -0.077 & 0.228 & 1.294 & 0.082 &  & -0.085 & 0.111
& 1.499 & 0.114 &  & -0.079 & 0.055 & 2.185 & 0.278 \\
$50\%$ observations &  & -0.168 & 0.175 & 1.306 & 0.178 &  & -0.164 & 0.084
& 1.982 & 0.518 &  & -0.164 & 0.041 & 3.821 & 0.983 \\
\multicolumn{16}{l}{\cite{andrews1998semiparametric}'s semiparametric
estimator with various smoothing parameters} \\
$1\%$ observations &  & -0.055 & 1.302 & 7.013 & 0.616 &  & 0.037 & 0.697 &
7.498 & 0.144 &  & 0.010 & 0.341 & 7.710 & 0.065 \\
$5\%$ observations &  & -0.007 & 0.689 & 3.711 & 0.165 &  & 0.007 & 0.338 &
3.632 & 0.077 &  & -0.005 & 0.165 & 3.737 & 0.048 \\
$10\%$ observations &  & -0.012 & 0.480 & 2.585 & 0.106 &  & -0.007 & 0.241
& 2.589 & 0.074 &  & -0.008 & 0.116 & 2.628 & 0.050 \\
$20\%$ observations &  & -0.027 & 0.331 & 1.787 & 0.076 &  & -0.029 & 0.163
& 1.775 & 0.064 &  & -0.024 & 0.079 & 1.857 & 0.051 \\
$30\%$ observations &  & -0.045 & 0.263 & 1.438 & 0.076 &  & -0.052 & 0.130
& 1.508 & 0.070 &  & -0.046 & 0.063 & 1.769 & 0.101 \\
$50\%$ observations &  & -0.106 & 0.198 & 1.211 & 0.111 &  & -0.110 & 0.096
& 1.569 & 0.214 &  & -0.106 & 0.047 & 2.629 & 0.596 \\
\multicolumn{16}{l}{Kernel regression (local constant) estimator with
various kernel functions} \\
Gaussian &  & -0.031 & 0.302 & 1.635 & 0.058 &  & -0.027 & 0.165 & 1.797 &
0.072 &  & -0.016 & 0.090 & 2.073 & 0.058 \\
Epanechnikov &  & -0.009 & 0.550 & 2.959 & 0.038 &  & 0.004 & 0.312 & 3.357
& 0.052 &  & -0.005 & 0.171 & 3.874 & 0.042 \\
7th polynomial &  & -0.017 & 0.653 & 3.514 & 0.064 &  & 0.012 & 0.357 & 3.843
& 0.052 &  & -0.001 & 0.192 & 4.333 & 0.044 \\
7th polyweight &  & -0.012 & 0.605 & 3.259 & 0.046 &  & 0.009 & 0.344 & 3.696
& 0.049 &  & -0.004 & 0.187 & 4.222 & 0.039 \\
\multicolumn{16}{l}{Local linear estimator with various kernel functions} \\
Gaussian &  & 0.032 & 0.284 & 1.539 & 0.092 &  & 0.029 & 0.148 & 1.616 &
0.091 &  & 0.032 & 0.077 & 1.894 & 0.102 \\
Epanechnikov &  & 0.019 & 0.424 & 2.284 & 0.057 &  & 0.014 & 0.235 & 2.535 &
0.053 &  & 0.014 & 0.125 & 2.843 & 0.045 \\
7th polynomial &  & -0.001 & 0.553 & 2.979 & 0.089 &  & 0.015 & 0.312 & 3.353
& 0.076 &  & 0.006 & 0.167 & 3.791 & 0.053 \\
7th polyweight &  & 0.005 & 0.491 & 2.643 & 0.065 &  & 0.013 & 0.278 & 2.989
& 0.059 &  & 0.008 & 0.151 & 3.424 & 0.063 \\ \hline\hline
\end{tabular}
}
\end{sidewaystable}

\begin{sidewaystable}[ptb]
\caption{Simulation results when $\varepsilon _{i}\sim t\left(3\right) $}
\medskip
\label{table:t}
\resizebox{\linewidth}{!}{
\begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}}
\hline\hline
$\Pr \left( Y_{i}=0\right) =0.5$ &  & \multicolumn{4}{c}{$n=250$} &  &
\multicolumn{4}{c}{$n=1000$} &  & \multicolumn{4}{c}{$n=4000$} \\
\cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16}
Estimator &  & Bias & SD & RMSE ratio & Rejection rate &  & Bias & SD & RMSE
ratio & Rejection rate &  & Bias & SD & RMSE ratio & Rejection rate \\ \hline
\multicolumn{16}{l}{\cite{heckman1979sample}'s parametric two-step estimator}
\\
&  & 0.014 & 0.181 & 1 & 0.059 &  & 0.027 & 0.089 & 1 & 0.058 &  & 0.025 &
0.047 & 1 & 0.114 \\
\multicolumn{16}{l}{\cite{heckman1990varieties}'s semiparametric estimator
with various smoothing parameters} \\
$1\%$ observations &  & -0.052 & 0.981 & 5.403 & 0.386 &  & 0.003 & 0.592 &
6.363 & 0.155 &  & -0.001 & 0.296 & 5.594 & 0.069 \\
$5\%$ observations &  & -0.050 & 0.533 & 2.945 & 0.121 &  & -0.030 & 0.271 &
2.934 & 0.070 &  & -0.029 & 0.135 & 2.603 & 0.056 \\
$10\%$ observations &  & -0.049 & 0.367 & 2.035 & 0.073 &  & -0.033 & 0.186
& 2.029 & 0.056 &  & -0.039 & 0.097 & 1.975 & 0.081 \\
$20\%$ observations &  & -0.061 & 0.260 & 1.468 & 0.063 &  & -0.052 & 0.133
& 1.532 & 0.065 &  & -0.056 & 0.067 & 1.654 & 0.137 \\
$30\%$ observations &  & -0.076 & 0.210 & 1.226 & 0.073 &  & -0.068 & 0.105
& 1.351 & 0.092 &  & -0.072 & 0.054 & 1.701 & 0.274 \\
$50\%$ observations &  & -0.120 & 0.163 & 1.115 & 0.113 &  & -0.109 & 0.080
& 1.452 & 0.257 &  & -0.111 & 0.042 & 2.238 & 0.770 \\
\multicolumn{16}{l}{\cite{andrews1998semiparametric}'s semiparametric
estimator with various smoothing parameters} \\
$1\%$ observations &  & -0.089 & 1.285 & 7.083 & 0.627 &  & 0.006 & 0.711 &
7.646 & 0.197 &  & 0.005 & 0.346 & 6.541 & 0.072 \\
$5\%$ observations &  & -0.059 & 0.659 & 3.641 & 0.137 &  & -0.018 & 0.325 &
3.503 & 0.069 &  & -0.022 & 0.157 & 2.999 & 0.060 \\
$10\%$ observations &  & -0.054 & 0.476 & 2.633 & 0.089 &  & -0.028 & 0.227
& 2.460 & 0.060 &  & -0.031 & 0.113 & 2.209 & 0.068 \\
$20\%$ observations &  & -0.051 & 0.311 & 1.733 & 0.069 &  & -0.039 & 0.156
& 1.727 & 0.058 &  & -0.045 & 0.080 & 1.727 & 0.087 \\
$30\%$ observations &  & -0.061 & 0.246 & 1.395 & 0.063 &  & -0.052 & 0.123
& 1.440 & 0.062 &  & -0.057 & 0.064 & 1.616 & 0.161 \\
$50\%$ observations &  & -0.089 & 0.184 & 1.123 & 0.075 &  & -0.080 & 0.093
& 1.319 & 0.142 &  & -0.083 & 0.048 & 1.813 & 0.438 \\
\multicolumn{16}{l}{Kernel regression (local constant) estimator with
various kernel functions} \\
Gaussian &  & -0.061 & 0.278 & 1.564 & 0.051 &  & -0.041 & 0.153 & 1.705 &
0.066 &  & -0.042 & 0.086 & 1.798 & 0.090 \\
Epanechnikov &  & -0.053 & 0.547 & 3.021 & 0.041 &  & -0.021 & 0.303 & 3.264
& 0.048 &  & -0.020 & 0.163 & 3.099 & 0.058 \\
7th polynomial &  & -0.058 & 0.642 & 3.545 & 0.061 &  & -0.013 & 0.355 &
3.824 & 0.059 &  & -0.014 & 0.190 & 3.596 & 0.055 \\
7th polyweight &  & -0.055 & 0.605 & 3.341 & 0.037 &  & -0.016 & 0.335 &
3.610 & 0.053 &  & -0.017 & 0.181 & 3.427 & 0.054 \\
\multicolumn{16}{l}{Local linear estimator with various kernel functions} \\
Gaussian &  & -0.020 & 0.270 & 1.488 & 0.072 &  & -0.009 & 0.147 & 1.583 &
0.069 &  & -0.018 & 0.082 & 1.587 & 0.067 \\
Epanechnikov &  & -0.034 & 0.399 & 2.200 & 0.046 &  & -0.016 & 0.222 & 2.397
& 0.046 &  & -0.022 & 0.125 & 2.403 & 0.060 \\
7th polynomial &  & -0.044 & 0.541 & 2.987 & 0.076 &  & -0.013 & 0.301 &
3.238 & 0.071 &  & -0.015 & 0.164 & 3.112 & 0.070 \\
7th polyweight &  & -0.039 & 0.474 & 2.616 & 0.050 &  & -0.014 & 0.265 &
2.850 & 0.056 &  & -0.019 & 0.147 & 2.803 & 0.068 \\ \hline\hline
\end{tabular}
}
\end{sidewaystable}

\begin{sidewaystable}[ptb]
\caption{Simulation results when $\varepsilon _{i}\sim \chi^2\left(3\right) $}
\medskip
\label{table:chi2}
\resizebox{\linewidth}{!}{
\begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}}
\hline\hline
$\Pr \left( Y_{i}=0\right) =0.5$ &  & \multicolumn{4}{c}{$n=250$} &  &
\multicolumn{4}{c}{$n=1000$} &  & \multicolumn{4}{c}{$n=4000$} \\
\cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16}
Estimator &  & Bias & SD & RMSE ratio & Rejection rate &  & Bias & SD & RMSE
ratio & Rejection rate &  & Bias & SD & RMSE ratio & Rejection rate \\ \hline
\multicolumn{16}{l}{\cite{heckman1979sample}'s parametric two-step estimator}
\\
&  & -0.146 & 0.167 & 1 & 0.180 &  & -0.145 & 0.082 & 1 & 0.459 &  & -0.146
& 0.041 & 1 & 0.913 \\
\multicolumn{16}{l}{\cite{heckman1990varieties}'s semiparametric estimator
with various smoothing parameters} \\
$1\%$ observations &  & -0.008 & 0.977 & 4.405 & 0.413 &  & 0.004 & 0.578 &
3.465 & 0.134 &  & -0.024 & 0.299 & 1.978 & 0.076 \\
$5\%$ observations &  & -0.016 & 0.514 & 2.321 & 0.125 &  & -0.047 & 0.258 &
1.569 & 0.067 &  & -0.052 & 0.128 & 0.913 & 0.073 \\
$10\%$ observations &  & -0.057 & 0.351 & 1.603 & 0.080 &  & -0.084 & 0.183
& 1.206 & 0.087 &  & -0.083 & 0.088 & 0.797 & 0.140 \\
$20\%$ observations &  & -0.117 & 0.243 & 1.215 & 0.099 &  & -0.129 & 0.124
& 1.072 & 0.189 &  & -0.126 & 0.059 & 0.917 & 0.530 \\
$30\%$ observations &  & -0.162 & 0.200 & 1.160 & 0.152 &  & -0.165 & 0.098
& 1.148 & 0.395 &  & -0.163 & 0.048 & 1.124 & 0.911 \\
$50\%$ observations &  & -0.229 & 0.149 & 1.232 & 0.354 &  & -0.230 & 0.073
& 1.445 & 0.869 &  & -0.229 & 0.037 & 1.535 & 1.000 \\
\multicolumn{16}{l}{\cite{andrews1998semiparametric}'s semiparametric
estimator with various smoothing parameters} \\
$1\%$ observations &  & -0.009 & 1.219 & 5.497 & 0.605 &  & 0.021 & 0.694 &
4.162 & 0.180 &  & -0.009 & 0.350 & 2.315 & 0.078 \\
$5\%$ observations &  & -0.009 & 0.624 & 2.816 & 0.153 &  & -0.026 & 0.315 &
1.893 & 0.067 &  & -0.040 & 0.156 & 1.061 & 0.068 \\
$10\%$ observations &  & -0.037 & 0.440 & 1.990 & 0.102 &  & -0.057 & 0.217
& 1.343 & 0.066 &  & -0.063 & 0.108 & 0.822 & 0.097 \\
$20\%$ observations &  & -0.084 & 0.291 & 1.365 & 0.081 &  & -0.099 & 0.148
& 1.068 & 0.113 &  & -0.097 & 0.072 & 0.799 & 0.251 \\
$30\%$ observations &  & -0.119 & 0.231 & 1.173 & 0.106 &  & -0.131 & 0.116
& 1.048 & 0.210 &  & -0.127 & 0.056 & 0.919 & 0.609 \\
$50\%$ observations &  & -0.182 & 0.171 & 1.124 & 0.206 &  & -0.185 & 0.084
& 1.217 & 0.582 &  & -0.185 & 0.042 & 1.250 & 0.991 \\
\multicolumn{16}{l}{Kernel regression (local constant) estimator with
various kernel functions} \\
Gaussian &  & -0.087 & 0.278 & 1.316 & 0.093 &  & -0.092 & 0.157 & 1.091 &
0.131 &  & -0.079 & 0.088 & 0.783 & 0.199 \\
Epanechnikov &  & -0.019 & 0.526 & 2.372 & 0.051 &  & -0.028 & 0.302 & 1.817
& 0.067 &  & -0.038 & 0.167 & 1.128 & 0.068 \\
7th polynomial &  & -0.015 & 0.624 & 2.816 & 0.071 &  & -0.017 & 0.355 &
2.131 & 0.072 &  & -0.030 & 0.196 & 1.308 & 0.062 \\
7th polyweight &  & -0.016 & 0.580 & 2.615 & 0.045 &  & -0.020 & 0.337 &
2.024 & 0.069 &  & -0.033 & 0.186 & 1.245 & 0.066 \\
\multicolumn{16}{l}{Local linear estimator with various kernel functions} \\
Gaussian &  & -0.066 & 0.258 & 1.202 & 0.095 &  & -0.070 & 0.134 & 0.907 &
0.122 &  & -0.065 & 0.069 & 0.627 & 0.183 \\
Epanechnikov &  & -0.027 & 0.387 & 1.751 & 0.056 &  & -0.044 & 0.222 & 1.355
& 0.066 &  & -0.039 & 0.122 & 0.846 & 0.071 \\
7th polynomial &  & -0.007 & 0.530 & 2.392 & 0.090 &  & -0.019 & 0.302 &
1.813 & 0.072 &  & -0.026 & 0.167 & 1.116 & 0.071 \\
7th polyweight &  & -0.012 & 0.466 & 2.102 & 0.063 &  & -0.029 & 0.266 &
1.605 & 0.059 &  & -0.031 & 0.148 & 0.997 & 0.072 \\ \hline\hline
\end{tabular}
}
\end{sidewaystable}

By contrast, the kernel regression and local linear estimators have no
difficulty in choosing the bandwidth parameter because fully data-driven
procedures of selecting the optimal bandwidths have been explicitly
proposed. Table \ref{table:normal} shows that the proposed bandwidths
perform satisfactorily in that the rejection rates for the kernel regression
and local linear estimators are all close to 0.05 for different sample sizes
and different kernel functions. In terms of RMSE, the local linear estimator
is superior to the kernel regression estimator, as expected. The RMSE of the
local linear estimator with Gaussian kernel is comparable to the optimal
RMSE of the Heckman and AS estimators. However, the Gaussian kernel may not
be the best choice because over-rejection for the t test appears to be a
problem. As an alternative, the local linear estimator with Epanechnikov
kernel achieves a satisfactory balance between RMSE and rejection rate. The
polynomial and polyweight kernels also do not lead to size distortion of the
t test but are clearly outperformed by the simple Epanechnikov kernel in
terms of RMSE.

Table \ref{table:t} investigates the finite sample behavior of the
estimators under $t\left( 3\right) $ distributed disturbance. The parametric
two-step estimator performs well for small sample sizes but deteriorates as
the sample size increases because its nonvanishing bias due to nonnormality
becomes large relative to its declining SD. By contrast, the performance of
the semiparametric estimators is robust to nonnormality, and the main
findings are almost the same as in the normal case. First, if the smoothing
parameter of the Heckman and AS estimators is improperly chosen, their RMSEs
may be increased by several factors and the t tests may yield misleading
inferences. Second, the smoothing parameter giving rise to the smallest RMSE
may lead to evident over-rejection for the t test. These facts indicate the
importance of selecting a proper smoothing parameter for the
identification-at-infinity estimators, but this task is difficult since the
RMSE and rejection rate are unobserved in practice. Third, the bandwidth
selection algorithms proposed for the kernel-type estimators perform well in
terms of rejection rate. Fourth, the local linear estimator dominates the
kernel regression estimator. Fifth, the RMSE of the local linear estimator
with Gaussian kernel is comparable to (when $n=4000$, even smaller than) the
smallest RMSE of the identification-at-infinity estimators. Sixth, using the
Epanechnikov kernel results in more accurate rejection rate.

Table \ref{table:chi2} considers the case of $\chi ^{2}\left( 3\right) $
distributed disturbance. In this case, the parametric two-step estimator has
notable bias, and the probability of making a type I error rapidly
approaches one. Comparison with Table \ref{table:t} indicates that the
skewness of the nonnormal disturbance exerts a worse influence on the
parametric estimator than does the fat tail. The Heckman and AS estimators
become less robust to the smoothing parameter under this design, in the
sense that the rejection rate is below 0.1 for only a narrow range of
smoothing parameters. On the other hand, the local linear estimator with
Epanechnikov kernel still has desirable finite sample properties in terms of
both RMSE and rejection rate.

Tables \ref{table:small}-\ref{table:large} investigate the effect of the
proportion of censoring. Table \ref{table:small} shows that, under mild
censoring, all the considered estimators behave better. The parametric
two-step estimator becomes less biased in nonnormal designs. The Heckman and
AS estimators are more robust to the smoothing parameter, and the rejection
rates for the kernel regression and local linear estimators are closer to
the true level, even when the Gaussian kernel is used. Table \ref
{table:large} reveals the reverse side of the coin, with an unchanged
conclusion being the superiority of the local linear estimator with
Epanechnikov kernel.

\begin{sidewaystable}[ptb]
\caption{Simulation results when $n=1000$ and the amount of zero observations is small}
\medskip
\label{table:small}
\resizebox{\linewidth}{!}{
\begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}}
\hline\hline
$\Pr \left( Y_{i}=0\right) =0.2$ &  & \multicolumn{4}{c}{$\varepsilon
_{i}\sim N\left( 0,1\right) $} &  & \multicolumn{4}{c}{$\varepsilon _{i}\sim
t\left( 3\right) $} &  & \multicolumn{4}{c}{$\varepsilon _{i}\sim \chi
^{2}\left( 3\right) $} \\ \cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16}
Estimator &  & Bias & SD & RMSE ratio & Rejection rate &  & Bias & SD & RMSE
ratio & Rejection rate &  & Bias & SD & RMSE ratio & Rejection rate \\ \hline
\multicolumn{16}{l}{\cite{heckman1979sample}'s parametric two-step estimator}
\\
&  & -0.001 & 0.065 & 1 & 0.055 &  & -0.015 & 0.068 & 1 & 0.070 &  & -0.071
& 0.066 & 1 & 0.239 \\
\multicolumn{16}{l}{\cite{heckman1990varieties}'s semiparametric estimator
with various smoothing parameters} \\
$1\%$ observations &  & -0.006 & 0.483 & 7.447 & 0.103 &  & -0.014 & 0.466 &
6.719 & 0.099 &  & -0.004 & 0.477 & 4.920 & 0.102 \\
$5\%$ observations &  & 0.006 & 0.224 & 3.449 & 0.065 &  & -0.017 & 0.218 &
3.145 & 0.061 &  & -0.021 & 0.206 & 2.134 & 0.057 \\
$10\%$ observations &  & 0.002 & 0.163 & 2.510 & 0.063 &  & -0.027 & 0.153 &
2.231 & 0.049 &  & -0.036 & 0.145 & 1.545 & 0.054 \\
$20\%$ observations &  & -0.001 & 0.113 & 1.743 & 0.056 &  & -0.035 & 0.132
& 1.970 & 0.058 &  & -0.056 & 0.106 & 1.242 & 0.081 \\
$30\%$ observations &  & -0.012 & 0.095 & 1.474 & 0.057 &  & -0.042 & 0.102
& 1.589 & 0.082 &  & -0.075 & 0.087 & 1.183 & 0.153 \\
$50\%$ observations &  & -0.036 & 0.073 & 1.252 & 0.093 &  & -0.054 & 0.074
& 1.317 & 0.130 &  & -0.111 & 0.067 & 1.336 & 0.434 \\
\multicolumn{16}{l}{\cite{andrews1998semiparametric}'s semiparametric
estimator with various smoothing parameters} \\
$1\%$ observations &  & -0.006 & 0.573 & 8.841 & 0.126 &  & 0.002 & 0.536 &
7.723 & 0.109 &  & -0.016 & 0.573 & 5.917 & 0.136 \\
$5\%$ observations &  & 0.002 & 0.264 & 4.077 & 0.061 &  & -0.012 & 0.258 &
3.714 & 0.056 &  & -0.010 & 0.258 & 2.661 & 0.070 \\
$10\%$ observations &  & 0.006 & 0.190 & 2.924 & 0.061 &  & -0.021 & 0.182 &
2.642 & 0.061 &  & -0.026 & 0.172 & 1.797 & 0.051 \\
$20\%$ observations &  & 0.000 & 0.132 & 2.037 & 0.060 &  & -0.030 & 0.130 &
1.918 & 0.060 &  & -0.042 & 0.121 & 1.325 & 0.064 \\
$30\%$ observations &  & -0.005 & 0.107 & 1.651 & 0.054 &  & -0.036 & 0.116
& 1.755 & 0.074 &  & -0.057 & 0.100 & 1.185 & 0.095 \\
$50\%$ observations &  & -0.020 & 0.081 & 1.285 & 0.066 &  & -0.046 & 0.088
& 1.423 & 0.091 &  & -0.088 & 0.074 & 1.189 & 0.243 \\
\multicolumn{16}{l}{Kernel regression (local constant) estimator with
various kernel functions} \\
Gaussian &  & 0.002 & 0.160 & 2.466 & 0.051 &  & -0.025 & 0.155 & 2.259 &
0.063 &  & -0.031 & 0.147 & 1.553 & 0.049 \\
Epanechnikov &  & -0.002 & 0.302 & 4.663 & 0.043 &  & -0.011 & 0.300 & 4.322
& 0.054 &  & -0.005 & 0.305 & 3.150 & 0.059 \\
7th polynomial &  & -0.004 & 0.345 & 5.321 & 0.056 &  & -0.007 & 0.342 &
4.920 & 0.057 &  & -0.010 & 0.359 & 3.711 & 0.074 \\
7th polyweight &  & -0.004 & 0.330 & 5.094 & 0.043 &  & -0.012 & 0.329 &
4.748 & 0.056 &  & -0.005 & 0.337 & 3.482 & 0.063 \\
\multicolumn{16}{l}{Local linear estimator with various kernel functions} \\
Gaussian &  & 0.017 & 0.143 & 2.222 & 0.070 &  & -0.024 & 0.147 & 2.151 &
0.067 &  & -0.020 & 0.128 & 1.334 & 0.057 \\
Epanechnikov &  & 0.005 & 0.234 & 3.608 & 0.057 &  & -0.016 & 0.227 & 3.280
& 0.053 &  & -0.011 & 0.217 & 2.247 & 0.044 \\
7th polynomial &  & 0.003 & 0.301 & 4.635 & 0.056 &  & -0.008 & 0.299 & 4.310
& 0.063 &  & -0.005 & 0.304 & 3.135 & 0.077 \\
7th polyweight &  & 0.008 & 0.274 & 4.227 & 0.054 &  & -0.010 & 0.273 & 3.933
& 0.056 &  & -0.008 & 0.265 & 2.734 & 0.061 \\ \hline\hline
\end{tabular}
}
\end{sidewaystable}

\begin{sidewaystable}[ptb]
\caption{Simulation results when $n=1000$ and the amount of zero observations is large}
\medskip
\label{table:large}
\resizebox{\linewidth}{!}{
\begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}}
\hline\hline
$\Pr \left( Y_{i}=0\right) =0.8$ &  & \multicolumn{4}{c}{$\varepsilon
_{i}\sim N\left( 0,1\right) $} &  & \multicolumn{4}{c}{$\varepsilon _{i}\sim
t\left( 3\right) $} &  & \multicolumn{4}{c}{$\varepsilon _{i}\sim \chi
^{2}\left( 3\right) $} \\ \cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16}
Estimator &  & Bias & SD & RMSE ratio & Rejection rate &  & Bias & SD & RMSE
ratio & Rejection rate &  & Bias & SD & RMSE ratio & Rejection rate \\ \hline
\multicolumn{16}{l}{\cite{heckman1979sample}'s parametric two-step estimator}
\\
&  & 0.001 & 0.168 & 1 & 0.045 &  & 0.195 & 0.231 & 1 & 0.193 &  & -0.234 &
0.132 & 1 & 0.428 \\
\multicolumn{16}{l}{\cite{heckman1990varieties}'s semiparametric estimator
with various smoothing parameters} \\
$1\%$ observations &  & -0.040 & 0.889 & 5.285 & 0.316 &  & 0.030 & 0.879 &
2.907 & 0.343 &  & -0.067 & 0.855 & 3.186 & 0.325 \\
$5\%$ observations &  & -0.035 & 0.403 & 2.404 & 0.086 &  & -0.044 & 0.415 &
1.380 & 0.097 &  & -0.081 & 0.391 & 1.485 & 0.098 \\
$10\%$ observations &  & -0.066 & 0.289 & 1.762 & 0.080 &  & -0.056 & 0.292
& 0.984 & 0.064 &  & -0.133 & 0.277 & 1.141 & 0.114 \\
$20\%$ observations &  & -0.154 & 0.208 & 1.539 & 0.149 &  & -0.099 & 0.212
& 0.772 & 0.089 &  & -0.212 & 0.188 & 1.053 & 0.219 \\
$30\%$ observations &  & -0.237 & 0.166 & 1.716 & 0.324 &  & -0.143 & 0.171
& 0.737 & 0.145 &  & -0.269 & 0.148 & 1.140 & 0.445 \\
$50\%$ observations &  & -0.388 & 0.124 & 2.419 & 0.876 &  & -0.228 & 0.134
& 0.875 & 0.452 &  & -0.362 & 0.113 & 1.409 & 0.891 \\
\multicolumn{16}{l}{\cite{andrews1998semiparametric}'s semiparametric
estimator with various smoothing parameters} \\
$1\%$ observations &  & -0.021 & 1.085 & 6.449 & 0.428 &  & 0.024 & 1.011 &
3.345 & 0.442 &  & -0.062 & 1.039 & 3.869 & 0.436 \\
$5\%$ observations &  & -0.022 & 0.497 & 2.955 & 0.105 &  & -0.032 & 0.487 &
1.614 & 0.104 &  & -0.056 & 0.485 & 1.814 & 0.108 \\
$10\%$ observations &  & -0.039 & 0.347 & 2.072 & 0.075 &  & -0.048 & 0.353
& 1.178 & 0.077 &  & -0.094 & 0.335 & 1.294 & 0.094 \\
$20\%$ observations &  & -0.093 & 0.243 & 1.545 & 0.080 &  & -0.070 & 0.251
& 0.860 & 0.071 &  & -0.159 & 0.230 & 1.040 & 0.136 \\
$30\%$ observations &  & -0.155 & 0.195 & 1.480 & 0.150 &  & -0.098 & 0.203
& 0.744 & 0.088 &  & -0.209 & 0.181 & 1.027 & 0.233 \\
$50\%$ observations &  & -0.282 & 0.144 & 1.878 & 0.526 &  & -0.162 & 0.152
& 0.736 & 0.199 &  & -0.291 & 0.131 & 1.187 & 0.610 \\
\multicolumn{16}{l}{Kernel regression (local constant) estimator with
various kernel functions} \\
Gaussian &  & -0.142 & 0.216 & 1.535 & 0.207 &  & -0.110 & 0.193 & 0.734 &
0.156 &  & -0.201 & 0.198 & 1.050 & 0.349 \\
Epanechnikov &  & -0.052 & 0.306 & 1.843 & 0.062 &  & -0.055 & 0.296 & 0.994
& 0.066 &  & -0.115 & 0.294 & 1.174 & 0.098 \\
7th polynomial &  & -0.041 & 0.355 & 2.122 & 0.069 &  & -0.043 & 0.343 &
1.142 & 0.062 &  & -0.094 & 0.359 & 1.380 & 0.094 \\
7th polyweight &  & -0.042 & 0.335 & 2.005 & 0.062 &  & -0.050 & 0.325 &
1.087 & 0.060 &  & -0.098 & 0.329 & 1.276 & 0.089 \\
\multicolumn{16}{l}{Local linear estimator with various kernel functions} \\
Gaussian &  & -0.101 & 0.188 & 1.264 & 0.224 &  & -0.012 & 0.182 & 0.602 &
0.139 &  & -0.169 & 0.181 & 0.920 & 0.328 \\
Epanechnikov &  & -0.035 & 0.244 & 1.467 & 0.087 &  & -0.023 & 0.230 & 0.764
& 0.068 &  & -0.134 & 0.217 & 0.947 & 0.136 \\
7th polynomial &  & -0.005 & 0.317 & 1.885 & 0.082 &  & -0.022 & 0.306 &
1.015 & 0.078 &  & -0.087 & 0.306 & 1.182 & 0.104 \\
7th polyweight &  & -0.011 & 0.284 & 1.686 & 0.081 &  & -0.023 & 0.272 &
0.904 & 0.067 &  & -0.106 & 0.263 & 1.054 & 0.112 \\ \hline\hline
\end{tabular}
}
\end{sidewaystable}

\section{Conclusion}

\label{sec:conclution}This paper rephrases the identification at infinity
into an identification at the boundary via a CDF transformation and
accordingly proposes a kernel approach to semiparametrically estimate the
intercept of the sample selection model. The proposed kernel regression
estimator with generic transformation is a generalization of the
identification-at-infinity estimators and thus inherits the disadvantage
that the asymptotic bias and variance are implicit functions of the
bandwidth parameter. To select a bandwidth that minimizes the asymptotic
mean squared error, I use a specific transformation, namely, the empirical
CDF of the selection index, under which the asymptotic bias and variance
become explicit with respect to the bandwidth. A plug-in bandwidth selection
algorithm with regularization is therefore suggested. For the purpose of
bias reduction, I further propose a local linear estimator and an associated
analogous bandwidth selection algorithm. A simulation study illustrates the
effectiveness of the selected bandwidths. Comparison of the finite sample
performance of the estimators indicates that the local linear estimator with
Epanechnikov kernel is superior to the parametric two-step estimator under
nonnormal disturbance and to the identification-at-infinity estimators in
most cases.

 \setcounter{theorem}{0}