EconBase
← Back to paper

On semiparametric estimation of the intercept of the sample selection model: a kernel approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

89,608 characters · 13 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On semiparametric estimation of the intercept of the sample selection model: a kernel approach

abstractThis paper presents a new perspective on the identification at infinity for the intercept of the sample selection model as identification at the boundary via a transformation of the selection index. This perspective suggests generalizations of estimation at infinity to kernel regression estimation at the boundary and further to local linear estimation at the boundary. The proposed kernel-type estimators with an estimated transformation are proven to be nonparametric-rate consistent and asymptotically normal under mild regularity conditions. A fully data-driven method of selecting the optimal bandwidths for the estimators is developed. The Monte Carlo simulation shows the desirable finite sample properties of the proposed estimators and bandwidth selection procedures.

Keywords{ : identification at infinity, kernel regression estimation, boundary effect, local linear estimation, bandwidth selection}

JEL codes{ : C01, C13, C14, C34}

Introduction

Since it was first introduced by the seminal papers of heckman1974shadow, gronau1974wage$,$ and lewis1974comments, the sample selection model has been increasingly widely applied in empirical studies to address potentially nonrandom samples that may arise from a variety of causes, such as improper sampling design, self-selectivity, nonresponse on survey questions, and attrition from social programs. Several important applications of the sample selection model, such as the estimation of average treatment effects and the decomposition of wage differentials, rely on an estimate of the intercept.

The intercept of the sample selection model is conventionally estimated along with the slope coefficients by means of the parametric maximum likelihood approach or the likelihood-based two-step procedure heckman1979sample. When the distribution of the model disturbance is misspecified, however, the parametric likelihood-based estimators for limited dependent variable models may possess evident bias and are likely inconsistent for the true value of interest arabmazar1982investigation. This finding spawned considerable and influential literature on semiparametric identification and estimation approaches that do not rest on parametric specification of the disturbance distribution, among which chamberlain1986asymptotic invented the notion of \textquotedblleft identification at infinity\textquotedblright , building identification upon the unbounded support of the regressor distribution. According to the identification at infinity, heckman1990varieties proposed a semiparametric estimator for the intercept of the sample selection model. However, heckman1990varieties's estimator contains discontinuous indicator functions, which greatly complicate the analysis of the statistical properties. andrews1998semiparametric suggested replacing the indicator function with a smoothed version and established the consistency and asymptotic normality of their modified estimator. Since then, the identification-at-infinity method has been standard for semiparametrically estimating the intercept of the sample selection model schafgans1998ethnic, schafgans2000gender, hussinger2008r, mulligan2008selection, liu2009maternal, shen2013determinants .

Nevertheless, identification-at-infinity estimation must choose a smoothing parameter or bandwidth controlling for the proportion of observations used, and no bandwidth selection algorithm that is both theoretically valid and practically tractable in the context of the intercept estimation has yet been developed. This is mainly because the asymptotic biases and variances of the estimators of heckman1990varieties and andrews1998semiparametric are all implicit functions of the bandwidth. To choose a theoretically valid bandwidth, one must impose high-level assumptions on the tail behaviors of the regressor and disturbance distributions in the selection equation and then carefully estimate some sort of tail indices, as in klein2015estimation. However, a precise estimate of the tail index of the selection disturbance distribution is difficult to obtain for two reasons. First, only the information in the tails, which is limited, is useful for estimating tail indices. Second, there is additional information loss when estimating the tail index of the selection disturbance distribution because one can observe only the binary selection outcome instead of the continuous selection disturbance. The imprecise estimation of the tail index makes it difficult to choose an eligible bandwidth in practice tan2018root. In empirical studies, researchers typically report several estimates of the intercept according to different values of bandwidth selected by simple rules, such as by sample quantiles of the selection linear index. However, these simple rules lack theoretical justification, and an inappropriately selected bandwidth may lead to sizable estimation bias and misleading inference results schafgans2004finite. Additionally, confusion will arise if different bandwidths lead to contradictory conclusions.

In this paper, I extend the identification-at-infinity estimators to a kernel regression estimator at the boundary and provide a simple bandwidth selection algorithm that is theoretically optimal in the sense of minimizing the asymptotic mean squared estimation error. By transforming the identification at infinity into identification at the boundary, I propose a kernel regression estimator for the intercept of the sample selection model that includes the estimators of heckman1990varieties and andrews1998semiparametric as special cases by taking particular forms of transformation. I suggest using the cumulative distribution function (CDF) of the selection linear index as the transformation, which makes the proposed kernel estimator substantially different from the existing identification-at-infinity estimators. Under this specific form of transformation, the asymptotic bias and variance of the kernel estimator are explicit functions of the bandwidth, motivating a simple plug-in procedure of optimal bandwidth selection. In practice, however, the CDF of the linear index is unknown. I adopt the empirical CDF estimation and show that the induced estimation error is asymptotically negligible under regularity conditions, implying that the kernel estimator with the empirical CDF follows the same asymptotic distribution as if the true CDF were known. As a result, the proposed bandwidth selection algorithm is theoretically valid.

Careful consideration of the asymptotic distribution of the kernel regression estimator for the intercept indicates a boundary effect in that the asymptotic bias has a larger order than the nonparametric kernel regression estimator in the interior because the estimator is, in nature, a Nadaraya-Watson or local constant estimator at the boundary. Hence, I further consider the local linear estimation of the intercept for bias reduction and show that it achieves a univariate nonparametric rate. The consistency and asymptotic normality of the estimator are established, and an optimal bandwidth selection algorithm tailored to the local linear case is presented.

The main contributions of this paper are as follows. First, I provide a novel interpretation of identification at infinity. Although infinity is intractable or at least irregular in most econometric models, it is merely an ordinary boundary point of the extended real line. Under the monotonic transformation that serves to construct a metric, the identification at infinity is converted into identification at the boundary, and kernel-type estimators are naturally derived. This approach may be helpful in other irregular identification problems involving infinity khan2010irregular. Second, I provide a theoretically justified procedure for bandwidth selection that is easy to implement. By implementing a special form of transformation, the asymptotic biases and variances of the kernel-type estimators become explicit functions of the bandwidth, based on which a fully data-driven algorithm is developed to choose the optimal bandwidth. The regularization method in imbens2012optimal is employed to keep the random denominator of the estimated bandwidth away from zero. Third, in the course of developing the statistical properties of the estimators, I provide a\ rigorous treatment of the estimation error induced by the empirical CDF. By means of a decomposition into an empirical process component and a discontinuous error component, the estimation error of the empirical CDF can be delicately controlled by the uniform rate of convergence of the empirical process over shrinking intervals stute1982oscillation and by the convergence results developed by schafgans2002intercept.

The rest of this paper is organized as follows. Section (ref) presents the sample selection model and reviews the semiparametric estimators for the intercept in the literature. Section (ref) motivates the kernel regression estimator, provides the regularity conditions under which the estimator is consistent and asymptotically normal, discusses the choice of the transformation, and presents the optimal bandwidth selection method. Section (ref) proposes local linear estimation to eliminate the boundary effect. Section (ref) reports the results from a small simulation and Section (ref) concludes. All the technical proofs are left to the Appendix.

The model and existing estimators

The model

The sample selection model has the following three-equation form:

eqnarray[eqnarray omitted — 211 chars of source]

where $D_{i}$ is a binary selection variable, $Y_{i}^{\ast }$ is the latent outcome, and $Y_{i}$ is the observed outcome. $X_{i}$ and $Z_{i}$ are random vectors of regressors (excluding the constant term) of the selection and outcome equations, respectively. $\varepsilon _{i}$ and $U_{i}$ are scalar disturbances that are possibly correlated with each other. $\beta _{0}$ and $ \theta _{0}$ are vectors of the slope coefficients, and $\mu _{0}$ is a scalar intercept coefficient, both of which are the estimands of the model. $ U_{i}$ is assumed to have zero mean to ensure identifiability of the intercept.

This paper is primarily concerned with the estimation of $\mu _{0}$ in the sample selection model ((ref)) without imposing a normal (or any other parametric) distribution restriction on the disturbances. The estimation of $ \mu _{0}$, however, relies on preliminary $\sqrt{n}$-consistent estimators for $\beta _{0}$ and $\theta _{0}$, which are denoted by $\hat{\beta}$ and $ \hat{\theta}$. Several such estimators are available in the literature. For instance, $\hat{\beta}$ can be one of the semiparametric estimators for the binary response model klein1993efficient,lewbel2000semiparametric or for the single-index model powell1989semiparametric,ichimura1993semiparametric, and $ \hat{\theta}$ can be one of the semiparametric slope estimators for the sample selection model gallant1987semi, chen1998efficient, powell2001semiparametric, lewbel2007endogenous, newey2009twostep, chen2010semiparametric . For notational convenience, denote $W_{i}=X_{i}^{\prime }\beta _{0}$ and $ \hat{W}_{i}=X_{i}^{\prime }\hat{\beta}$, hence $D_{i}=1\left\{ W_{i}>\varepsilon _{i}\right\} $.

The identification-at-infinity estimators

A natural strategy for identifying the intercept $\mu _{0}$ starts from the expectation of $Y_{i}-Z_{i}^{\prime }\theta _{0}$ conditional on the observation being selected:

equation*[equation* omitted — 233 chars of source]

In the context of sample selection, where $U_{i}$ is correlated with $ \varepsilon _{i}$, the selectivity bias term $E\left[ U_{i}\left\vert D_{i}=1\right. \right] $ is generally nonzero and contaminates the identification of $\mu _{0}$. Inspired by the fact that $E\left[ U_{i}\left\vert D_{i}=1\right. \right] $ will become arbitrarily close to zero for those $W_{i}$ such that $\Pr \left( D_{i}=1\left\vert W_{i}\right. \right) $ is arbitrarily close to unity, heckman1990varieties suggested identifying $\mu _{0}$ at infinity, i.e., when $W_{i}$ takes arbitrarily large values. This idea can be heuristically formulated as

equation*[equation* omitted — 252 chars of source]

provided the conditional mean independence is assumed. An intuitive estimator based on the above identification strategy is

equation[equation omitted — 239 chars of source]

where $\left\{ \gamma _{n}\right\} $ is a sequence of positive smoothing parameters such that $\gamma _{n}\rightarrow \infty $ as $n\rightarrow \infty $. The estimator $\hat{\mu}^{H}$ is essentially a sample average of $ \mu _{0}+U_{i}$ over a decreasingly small fraction of all observations. The effective sample size depends on the proportion of data censoring, the degree of tail heaviness of $\hat{W}_{i}$'s distribution, and the choice of $ \gamma _{n}$.

To facilitate the development of distribution theory of the identification-at-infinity estimation, andrews1998semiparametric replaced the indicator function in heckman1990varieties's estimator ((ref)) with a smoothed function and proposed a weighted sample average estimator:

equation[equation omitted — 237 chars of source]

where $s\left( \cdot \right) $ is a nondecreasing $\left[ 0,1\right] $ -valued function that has a third derivative bounded over $R$ and satisfies $ s\left( w\right) =0$ for $w\leq 0$ and $s\left( w\right) =1$ for $w\geq b$, for some $b>0$. By taking advantage of the smoothness of $s\left( \cdot \right) $, andrews1998semiparametric showed the asymptotic negligibility of the estimation error induced by the preliminary estimators $ \hat{\beta}$ and $\hat{\theta}$ and established the consistency and asymptotic normality of $\hat{\mu}^{AS}$ under mild regularity conditions. The rate of convergence of $\hat{\mu}^{AS}$ depends both upon the tail heaviness of $W_{i}$'s distribution and upon the rate of divergence of $ \gamma _{n}$. However, theoretically valid procedures for choosing $\gamma _{n}$ have not yet been developed.

Other existing estimators

The semiparametric literature has neglected, to a certain extent, the intercept estimation of the sample selection model. This is mainly because the intercept is absorbed into the selectivity bias correction term in the course of estimating the slope coefficients. The only exceptions, besides $ \hat{\mu}^{H}$ and $\hat{\mu}^{AS}$, are the estimators of gallant1987semi, lewbel2007endogenous, and chen2010semiparametric.

By employing Hermite series to approximate the unknown bivariate density function of the disturbances, gallant1987semi considered semiparametric maximum likelihood estimation of the intercept along with the slopes. The consistency of their estimator requires complicated continuity conditions on the distributions of the disturbances and regressors that are difficult to verify. Moreover, the asymptotic distribution of their estimator has not been established. lewbel2007endogenous achieved identification of the intercept by the presence of a special regressor that has a large support. When the support is infinite, his identification strategy would be essentially equivalent to the identification at infinity. chen2010semiparametric proposed a kernel-weighted pairwise-difference-type estimator for the intercept and showed its consistency and asymptotic normality. However, their estimation method requires the disturbances to be jointly symmetrically distributed; this joint symmetry condition may not hold in practice.

Kernel regression estimation

The kernel approach to estimating the intercept $\mu _{0}$ in the sample selection model ((ref)) is grounded in the identification at infinity:

equation[equation omitted — 134 chars of source]

Let $F\left( \cdot \right) $ be an absolutely continuous CDF that is strictly increasing over $R$; then, the infinity condition $W_{i}=+\infty $ is equivalent to a boundary condition $F\left( W_{i}\right) =F\left( +\infty \right) =1$. Therefore, the identification at infinity ((ref)) can be written as identification at the boundary:

equation[equation omitted — 148 chars of source]

In other words, $\mu _{0}$ is identified by the conditional expectation of $ Y_{i}-Z_{i}^{\prime }\theta _{0}$ given that the observation is selected and that the transformation of $W_{i}$, $F\left( W_{i}\right) $, takes a boundary value. The identification-at-boundary equation ((ref) ) motivates the kernel regression estimation of $\mu _{0}$:

equation[equation omitted — 283 chars of source]

where $\hat{W}_{i}=X_{i}^{\prime }\hat{\beta}$, $\hat{\beta}$ and $\hat{ \theta}$ are preliminary estimators for $\beta _{0}$ and $\theta _{0}$, $ k\left( \cdot \right) $ is a kernel function defined on $\left[ 0,\infty \right) $, and $\left\{ h_{n}\right\} $ is a sequence of positive bandwidth parameters such that $h_{n}\rightarrow 0$ as $n\rightarrow \infty $. The values of the kernel over the negative reals are irrelevant since the argument of it is always positive. Some examples of $k\left( \cdot \right) $ are given in Table (ref). $\tilde{\mu}$ is a kernel regression estimator at the boundary. Alternatively, it can be viewed as a kernel regression estimator at infinity over the extended real line $\bar{R}=\left[ -\infty ,+\infty \right] =R\cup \left\{ -\infty ,+\infty \right\} $. A distance function with which $\bar{R}$ would become a (compact) metric space can be defined as $d\left( w_{1},w_{2}\right) =\left\vert F\left( w_{1}\right) -F\left( w_{2}\right) \right\vert $\ for any $w_{1},w_{2}\in \bar{R}$. Accordingly, the numerator within the kernel is the distance between $\hat{W}_{i}$ and infinity since $d\left( \hat{W}_{i},+\infty \right) =\left\vert F\left( \hat{W}_{i}\right) -1\right\vert =1-F\left( \hat{ W}_{i}\right) $. From this perspective, our estimation method is an extension of kernel regression to the extended real line.

Similar to the identification-at-infinity estimators $\hat{\mu}^{H}$ and $ \hat{\mu}^{AS}$ defined in ((ref)) and ((ref)), the kernel estimator $\tilde{\mu}$ is also a (weighted) sample average over a vanishingly small subset of the entire data set. However, the local average in $\tilde{\mu}$ is over observations for $F\left( \hat{W}_{i}\right) $ in a left neighborhood of the right boundary point, rather than for $\hat{W}_{i}$ near positive infinity. Notably, $\tilde{\mu}$ generalizes the identification-at-infinity estimators in the sense that it nests $\hat{\mu} ^{H}$ and $\hat{\mu}^{AS}$ as special cases by taking specific forms of $ F\left( \cdot \right) $, $k\left( \cdot \right) $ and $h_{n}$.

prop(i) For any strictly increasing $F\left( \cdot \right) $, if $k\left( u\right) =1\left\{ 0\leq u\leq 1\right\} $ and $h_{n}=1-F\left( \gamma _{n}\right) $, then $\tilde{\mu}=\hat{\mu}^{H}$. (ii) If $F\left( \cdot \right) $ is the standard Laplacian CDF such that $F\left( w\right) =1-\left( 1\left/ 2\right. \right) \exp \left( -w\right) $ for $w\geq 0$, if $k\left( u\right) =s\left( -\log u\right) $, and if $h_{n}=1-F\left( \gamma _{n}\right) =\left( 1\left/ 2\right. \right) \exp \left( -\gamma _{n}\right) $, then $\tilde{\mu}=\hat{\mu}^{AS}$.

Asymptotic property

This subsection investigates the consistency and asymptotic normality of the kernel estimator $\tilde{\mu}$ under general $F\left( \cdot \right) $. Several regularity assumptions are made first. Denote $k_{ni}=k\left( \left. \left( 1-F\left( W_{i}\right) \right) \right/ h_{n}\right) $, where $ W_{i}=X_{i}^{\prime }\beta _{0}$.

\noindentAssumption 1. (i) $\left\{ \left( Y_{i},D_{i},X_{i},Z_{i}\right) \right\} _{i=1}^{n}$ is a random sample of $n$ observations from Model ((ref)). (ii) The model disturbance $\left( \varepsilon _{i},U_{i}\right) $ is independent of the regressors $X_{i}$ and $Z_{i}$. (iii) $EU_{i}=0$. (iv) There exists a positive constant $c_{1}$ such that $E\left| U_{i}\right| ^{2+c_{1}}<\infty $, $E\left\| X_{i}\right\| ^{4+c_{1}}<\infty $, and $E\left\| Z_{i}\right\| ^{2+c_{1}}<\infty $.

\noindentAssumption 2. The kernel function $k\left( \cdot \right) $ is defined on $\left[ 0,\infty \right) $ and satisfies that (i) it is nonnegative and supported on $\left[ 0,1\right] $, (ii) it is bounded above by $\bar{k}>0$, and (iii) it is twice continuously differentiable over $ \left[ 0,\infty \right) $ and its derivatives $k^{\prime }\left( \cdot \right) $ and $k^{\prime \prime }\left( \cdot \right) $ are bounded above by $\bar{k^{\prime }}$ and $\bar{k^{\prime \prime }}$, respectively.

\noindentAssumption 3. (i) The transformation function $F\left( \cdot \right) $ is a CDF for a continuously distributed random variable whose support is right-unbounded. (ii) The probability density function (PDF) $f\left( \cdot \right) $ corresponding to $F\left( \cdot \right) $ is absolutely continuous with respect to the Lebesgue measure, with its derivative $f^{\prime }\left( \cdot \right) $ bounded almost everywhere by $ \bar{f^{\prime }}$. (iii) For a large positive constant $C$, the tail hazard rate $H_{C}\left( \cdot \right) =1\left\{ \cdot >C\right\} f\left( \cdot \right) \left/ \left[ 1-F\left( \cdot \right) \right] \right. $ and the tail score (of location parameter) $S_{C}\left( \cdot \right) =1\left\{ \cdot >C\right\} \left. f^{\prime }\left( \cdot \right) \right/ f\left( \cdot \right) $ corresponding to $F\left( \cdot \right) $ satisfy $E\left[ \sup_{\beta \in \mathcal{N}\left( \beta _{0}\right) }H_{C}^{4}\left( X_{i}^{\prime }\beta \right) \right] <\infty $ and $E\left[ \sup_{\beta \in \mathcal{N}\left( \beta _{0}\right) }S_{C}^{4}\left( X_{i}^{\prime }\beta \right) \right] <\infty $ for a neighborhood $\mathcal{N}\left( \beta _{0}\right) $ of $\beta _{0}$.

\noindentAssumption 4. The preliminary estimators $\hat{\beta}$ and $\hat{\theta}$ are $\sqrt{n}$-consistent, that is, $\sqrt{n}\left\| \hat{ \beta}-\beta _{0}\right\| =O_{p}\left( 1\right) $ and $\sqrt{n}\left\| \hat{ \theta}-\theta _{0}\right\| =O_{p}\left( 1\right) $.

Assumption 5. The bandwidth satisfies that $0<h_{n}\leq 1/2$ and that as $n\rightarrow \infty $, (i) $h_{n}\rightarrow 0$, (ii) $ nEk_{ni}^{2}\rightarrow \infty $, and (iii) $\left. \left[ \Pr \left( F\left( W_{i}\right) >1-h_{n}\right) \right] ^{1+c}\right/ Ek_{ni}^{2}\rightarrow 0$ for any $c>0$.

Assumption 1 describes the model and data. Assumptions 2 and 3 impose smoothness and boundedness conditions on the kernel function $k\left( \cdot \right) $ and the transformation function $F\left( \cdot \right) $. The compactness of the kernel's support is assumed to simplify the technical proofs, and the theoretical results in this paper are still supposed to hold when using kernels that decay sufficiently fast in the tails. Assumption 3.(iii) is a joint condition on the tail behaviors of $F\left( \cdot \right) $ and $X_{i}$'s distribution. When $F\left( \cdot \right) $ has a power-type upper tail, e.g., Pareto($\lambda $) tail such that\footnote{ Here and below, $g\left( t\right) \sim h\left( t\right) $ means that the ratios $g\left( t\right) \left/ h\left( t\right) \right. $ and $h\left( t\right) \left/ g\left( t\right) \right. $ are $O\left( 1\right) $ as $ t\rightarrow +\infty $.} $1-F\left( t\right) \sim t^{-\lambda }$ for some $ \lambda >0$, we have $H_{C}\left( t\right) \sim t^{-1}$ and $S_{C}\left( t\right) \sim t^{-1}$, and Assumption 3.(iii) holds for any distribution of $ X_{i}$. When $F\left( \cdot \right) $ has an exponential-type upper tail, e.g., Weibull($\lambda $) tail such that $1-F\left( t\right) \sim \exp \left( -c_{0}t^{\lambda }\right) $ for some $\lambda >0$ and $c_{0}>0$, we have $H_{C}\left( t\right) \sim t^{\lambda -1}$ and $S_{C}\left( t\right) \sim t^{\lambda -1}$. If $0<\lambda \leq 1$, Assumption 3.(iii) still holds for any distribution of $X_{i}$. For example, in the special case of andrews1998semiparametric's estimator that corresponds to a Laplacian $ F\left( \cdot \right) $ (i.e., $\lambda =1$), Assumption 3.(iii) is automatically satisfied. If $\lambda >1$, then Assumption 3.(iii) will require $E\left\Vert X_{i}\right\Vert ^{4\left( \lambda -1\right) }<\infty $ . For example, if $F\left( \cdot \right) $ is the normal CDF, we need $ E\left\Vert X_{i}\right\Vert ^{4}<\infty $, which is already guaranteed by Assumption 1.(iv). When the tail of $F\left( \cdot \right) $ decays more rapidly as $1-F\left( t\right) \sim \exp \left( -\exp \left( c_{0}t^{\lambda }\right) \right) $, a sufficient condition of Assumption 3.(iii) is that the distribution of $\left\Vert X_{i}\right\Vert $ has a Weibull($\theta $) tail with $\theta >\lambda $. Assumption 4 imposes the $\sqrt{n}$-consistency of the preliminary slope estimators. Examples of such estimators are listed in Subsection (ref).1.

According to Assumption 5, the bandwidth $h_{n}$ is required to go to zero as the sample size goes to infinity, but the speed of the decline is not allowed to be excessively fast. To see this, note that $Ek_{ni}^{2}\leq \bar{ k}^{2}\Pr \left( F\left( W_{i}\right) >1-h_{n}\right) $ by Assumption 2. As $ h_{n}$ goes to zero, $\Pr \left( F\left( W_{i}\right) >1-h_{n}\right) $ will also go to zero. If the speed at which $h_{n}$ declines is so fast that $\Pr \left( F\left( W_{i}\right) >1-h_{n}\right) $ goes to zero at a rate faster than $n^{-1}$, we will have $nEk_{ni}^{2}\rightarrow 0$, which violates Assumption 5.(ii). Another implicit requirement of Assumption 5.(ii) is the upper unboundedness of $W_{i}$'s support, for if $W_{i}$ has bounded support from above, $\Pr \left( F\left( W_{i}\right) >1-h_{n}\right) $ will be exactly zero for sufficiently small $h_{n}$. Assumption 5.(iii) requires the distribution of $W_{i}$ to be not too thin upper tailed relative to $F\left( \cdot \right) $, as illustrated by the following example.

exampSuppose $F\left( \cdot \right) $ is the standard Laplacian CDF, as is the case with $\hat{\mu}^{AS}$. Then, Assumption 5.(iii) holds if (i) the distribution of $W_{i}$ has a power-type upper tail; (ii) the distribution of $W_{i}$ has an exponential-type upper tail; (iii) the distribution of $W_{i}$ has a “double”-exponential-type but not excessively thin upper tail, namely, $\Pr \left( W_{i}>t\right) \sim \exp \left( -\exp \left( c_{0}t^{\lambda }\right) \right) $ with $\lambda <1$.
theoremUnder Assumptions 1-5, the kernel regression estimator $ \tilde{\mu}$ defined in ((ref)) is consistent and asymptotically normal: \begin{equation*} \frac{\sqrt{n}Ek_{ni}}{\sqrt{Ek_{ni}^{2}}}\left( \tilde{\mu}-\mu _{0}-\frac{ EU_{i}D_{i}k_{ni}}{ED_{i}k_{ni}}\right) \rightarrow N\left( 0,\sigma _{U}^{2}\right) , \end{equation*} where $\sigma _{U}^{2}=EU_{i}^{2}$.

Theorem (ref) generalizes Theorems 1-2 of andrews1998semiparametric by allowing general forms of the transformation function $F\left( \cdot \right) $. At first sight, introducing a general $ F\left( \cdot \right) $ appears to help little in the choice of $h_{n}$ because the asymptotic bias and variance of $\tilde{\mu}$ are still unknown functions of $h_{n}$. However, as will be seen below, a particular data-dependent choice of $F\left( \cdot \right) $ can largely facilitate the subsequent choice of $h_{n}$.

Choice of the transformation $F\left( \cdot \right) $

A natural choice of $F\left( \cdot \right) $ would be the CDF of $ W_{i}=X_{i}^{\prime }\beta _{0}$, under which $F\left( W_{i}\right) $ follows a uniform distribution and several assumptions of Theorem (ref) have a simpler form. For example, Assumption 3.(iii) will be automatically fulfilled in this case as long as $W_{i}$ is continuously distributed, and Assumption 3 is implied by

\noindentAssumption 3'. $W_{i}$ is continuously distributed with right-unbounded support. In addition, $W_{i}$'s PDF $f_{W}\left( \cdot \right) $ is absolutely continuous with a derivative that is bounded almost everywhere.

Assumption 3' presumes that the underlying data is continuous, as in standard kernel methods. When encountering discrete regressors, the kernel estimator can be adapted using the frequency-based method or the smoothing method racine2004nonparametric at the cost of more complicated notations. Assumption 5.(iii) will also be automatically fulfilled for $ F\left( \cdot \right) =F_{W}\left( \cdot \right) $ because in this case $\Pr \left( F\left( W_{i}\right) >1-h_{n}\right) =h_{n}$ and

equation[equation omitted — 195 chars of source]

for any $r>0$. Therefore, Assumption 5 is simplified to $h_{n}\rightarrow 0$ and $nh_{n}\rightarrow \infty $ as $n\rightarrow \infty $.

Now, consider the asymptotic bias of $\tilde{\mu}$, $EU_{i}D_{i}k_{ni}\left/ ED_{i}k_{ni}\right. $, under $F\left( \cdot \right) =F_{W}\left( \cdot \right) $. Define

equation[equation omitted — 198 chars of source]

for $t\in \left( 0,1\right) $, and define $G\left( 1\right) =\lim_{t\rightarrow 1}G\left( t\right) =0$.

lemmLet $F\left( \cdot \right) =F_{W}\left( \cdot \right) =\Pr \left( W_{i}\leq \cdot \right) $. Suppose $G\left( t\right) $ is continuously differentiable over $t\in \left( 1-\delta ,1\right] $ for a small $\delta >0$, with its derivative function $g\left( t\right) $ being bounded above over $t\in \left( 1-\delta ,1\right] $. Then, we have \begin{equation*} \frac{EU_{i}D_{i}k_{ni}}{ED_{i}k_{ni}}=-\frac{\kappa _{1}g\left( 1\right) }{ \kappa _{0}}h_{n}+o\left( h_{n}\right) , \end{equation*} where $\kappa _{r}=\int_{0}^{1}t^{r}k\left( t\right) dt$.

Since by a calculation

equation*[equation* omitted — 352 chars of source]

it can be seen that the finiteness of $g\left( 1\right) =\lim_{t\rightarrow 1}g\left( t\right) $ assumed by the Lemma essentially requires the selection index $W_{i}$ to have a heavier upper tail than the selection disturbance $ \varepsilon _{i}$. The relatively heavy tail of $W_{i}$ is also required by $ \hat{\mu}^{AS}$, as illustrated by the examples of andrews1998semiparametric.

corLet $F\left( \cdot \right) =F_{W}\left( \cdot \right) =\Pr \left( W_{i}\leq \cdot \right) $. Suppose (i) Assumptions 1-5 hold, (ii) the assumption of the Lemma holds, and (iii) $nh_{n}^{3}=O\left( 1\right) $. Then, the kernel regression estimator $\tilde{\mu}$ defined in ((ref))\ is consistent and asymptotically normal: \begin{equation*} \sqrt{nh_{n}}\left( \tilde{\mu}-\mu _{0}+\frac{\kappa _{1}g\left( 1\right) }{ \kappa _{0}}h_{n}\right) \rightarrow N\left( 0,\frac{\chi _{0}\sigma _{U}^{2} }{\kappa _{0}^{2}}\right) , \end{equation*} where $\chi _{0}=\int_{0}^{1}k^{2}\left( t\right) dt$.

The corollary shows that, by setting $F\left( \cdot \right) =F_{W}\left( \cdot \right) $, the asymptotic bias and variance of the kernel estimator become explicit functions of $h_{n}$, based on which we can select a theoretically optimal bandwidth. In practice, however, $F_{W}\left( \cdot \right) $ is unknown and must be estimated beforehand. A simple estimator for $F_{W}\left( \cdot \right) $ is the empirical CDF of $\hat{W} _{i}=X_{i}^{\prime }\hat{\beta}$. With this specific choice, the kernel estimator becomes

equation[equation omitted — 302 chars of source]

where $\hat{F}_{n}\left( \hat{W}_{i}\right) =\frac{1}{n-1}\sum_{j\neq i}1\left\{ \hat{W}_{j}\leq \hat{W}_{i}\right\} $. To control the estimation error induced by the empirical CDF, the existing assumptions should be strengthened, and one new assumption should be imposed. Partition $ X_{i}=\left( X_{i1},X_{i,(-1)}^{\prime }\right) ^{\prime }$, $\beta _{0}=\left( \beta _{01},\beta _{0,(-1)}^{\prime }\right) ^{\prime }$, and $ \hat{\beta}=\left( \hat{\beta}_{1},\hat{\beta}_{(-1)}^{\prime }\right) ^{\prime }$, with $X_{i1}$, $\beta _{01}$, and $\hat{\beta}_{1}$\ being the respective first components.

\noindentAssumption 1'. Assumption 1 holds and $E\left\| X_{i}\right\| ^{6}<\infty $; $\beta _{01}\neq 0$. Without loss of generality, set $\beta _{01}=1$ to achieve scale normalization of $\beta _{0} $.

Assumption 2'. Assumption 2 holds and $k\left( \cdot \right) $ is six times continuously differentiable over $\left[ 0,\infty \right) $ with all of its derivatives bounded above.

\noindentAssumption 4'. Assumption 4 holds and $\hat{\beta}_{1}=1$.

Assumption 5'. The bandwidth $h_{n}\in \left( 0,1/2\right] $ satisfies (i) $h_{n}\rightarrow 0$, (ii) $nh_{n}^{3}=O\left( 1\right) $, and (iii) $nh_{n}^{13/5}\rightarrow \infty $, as $n\rightarrow \infty $.

Assumption 6. The PDF of $W_{i}$, $f_{W}\left( \cdot \right) $, and the conditional PDF of $W_{i}$ given $X_{i,(-1)}$, $ f_{W\left| X_{(-1)}\right. }\left( \cdot \left| X_{i,(-1)}\right. \right) $, satisfies that (i) $\left. f_{W}\left( F_{W}^{-1}\left( 1-h_{n}\right) -n^{-1/10}h_{n}^{-1/18}\right) \right/ h_{n}^{2/3}\rightarrow 0$, (ii) $ f_{W\left| X_{(-1)}\right. }\left( \cdot \left| X_{i,(-1)}\right. \right) $ is uniformly bounded by $\overline{f_{W\left| X_{(-1)}\right. }}$, and (iii) there exists a large constant $C$ such that $f_{W\left| X_{(-1)}\right. }\left( w\left| X_{i,(-1)}\right. \right) $ declines monotonically for $w>C$ almost surely.

Since $\beta _{0}$ is identified only up to scale, the normalization $\beta _{01}=\hat{\beta}_{1}=1$ is postulated by most semiparametric estimators for the binary response model. Assumption 6, which is comparable to Assumption A of schafgans2002intercept, is imposed to address the non-differentiable indicator function contained in the empirical CDF. Because the expectation of the generalized derivative of the indicator function equals the value of the PDF, Assumption 6 involves restrictions on the PDF and conditional PDF of $W_{i}$. Assumption 6.(i) relates to the upper tail behavior of the distribution of $W_{i}$. Under the assumption of $ nh_{n}^{13/5}\rightarrow \infty $ that implies $n^{-1/10}h_{n}^{-1/18} \rightarrow 0$, it can be shown that Assumption 6.(i) is satisfied if $W_{i}$ has a power- or exponential-type upper tail.

theoremUnder Assumptions 1'-5' and 6, the kernel regression estimator $\hat{\mu}$ defined in ((ref)) is consistent and asymptotically normal: \begin{equation*} \sqrt{nh_{n}}\left( \hat{\mu}-\mu _{0}+\frac{\kappa _{1}g\left( 1\right) }{ \kappa _{0}}h_{n}\right) \rightarrow N\left( 0,\frac{\chi _{0}\sigma _{U}^{2} }{\kappa _{0}^{2}}\right) , \end{equation*} where $\kappa _{r}=\int_{0}^{1}t^{r}k\left( t\right) dt$ and $\chi _{r}=\int_{0}^{1}t^{r}k^{2}\left( t\right) dt$.

Theorem (ref) reveals that, under slightly stronger conditions, the estimation error induced by the empirical CDF is asymptotically negligible; thus, the asymptotic distribution of $\hat{\mu}$ is the same as if the true CDF $F_{W}\left( \cdot \right) $ were known.

Bandwidth selection

An important implication of Theorem (ref) is that, under the specific choice of $F\left( \cdot \right) =F_{W}\left( \cdot \right) $, the kernel estimator for the intercept follows a standard asymptotic distribution as an ordinary kernel regression estimator at the boundary point. As a result, we can borrow approaches of bandwidth selection from the nonparametric regression literature li2007nonparametric . However, the widely used cross-validation method based on the integrated mean squared error (MSE) criteria takes into account the global performance of the regression function estimation and is thus not suitable for the problem at hand. A closely related problem is the choice of bandwidth for the nonparametric regression discontinuity estimator, which is the difference between two regression estimators evaluated at boundary points. I follow the plug-in method proposed by imbens2012optimal and suggest a data-dependent bandwidth selection procedure that is tailored to $\hat{\mu}$.

By Theorem (ref), we know that

equation*[equation* omitted — 244 chars of source]

where $g\left( 1\right) =\left. \left. dG\left( t\right) \right/ dt\right\vert _{t=1}$ with $G\left( t\right) $ defined in ((ref)), and $ \sigma _{U}^{2}=Var\left( U_{i}\right) $. Provided $g\left( 1\right) \neq 0$ , the optimal bandwidth for $\hat{\mu}$ is defined as a minimizer of its asymptotic MSE:

eqnarray*[eqnarray* omitted — 338 chars of source]

where $c_{k}=\left( \chi _{0}\left/ \left( 2\kappa _{1}^{2}\right) \right. \right) ^{1/3}$ is a functional of the kernel $k\left( \cdot \right) $. If $ g\left( 1\right) =0$, the bias converges to zero faster, allowing for estimation of the intercept at a faster rate of convergence. However, it is difficult to exploit the improved convergence rate resulting from this condition in practice; hence, I focus on the optimal bandwidth given $ g\left( 1\right) \neq 0$.

A natural choice of the estimator for the optimal bandwidth $h_{opt}$ is to replace $\sigma _{U}^{2}$ and $g\left( 1\right) $ with their consistent estimators $\hat{\sigma}_{U}^{2}$ and $\hat{g}\left( 1\right) $, respectively. One potential problem with this choice is that $\hat{g}\left( 1\right) $ may occasionally be very close to zero due to the stochastic estimation error, even if $g\left( 1\right) \neq 0$. In such cases, the estimated bandwidth will be imprecisely and unstably large, which may in turn lead to large finite sample bias of $\hat{\mu}$ because observations that are far from the boundary will be included in the kernel estimation. To alleviate this problem, I employ the regularization method imbens2012optimal and propose the following bandwidth estimator:

equation[equation omitted — 189 chars of source]

By means of regularization, $\hat{h}_{opt}$ will not become infinite even in the case of $\hat{g}\left( 1\right) =0$. Moreover, the leading term of $E \left[ 1\left/ \left( \hat{g}^{2}\left( 1\right) +3Var\left( \hat{g}\left( 1\right) \right) \right) \right. -1\left/ \hat{g}^{2}\left( 1\right) \right. \right] $ cancels out the leading term of $E\left[ 1\left/ \hat{g}^{2}\left( 1\right) \right. -1\left/ g^{2}\left( 1\right) \right. \right] $. Therefore, the bias of $1\left/ \left( \hat{g}^{2}\left( 1\right) +3Var\left( \hat{g} \left( 1\right) \right) \right) \right. $ for the reciprocal of $g^{2}\left( 1\right) $ is of lower order than the bias of the naive $1\left/ \hat{g} ^{2}\left( 1\right) \right. $. It remains to construct consistent estimators for the components of the plug-in bandwidth, namely, $\hat{\sigma}_{U}^{2}$, $\hat{g}^{2}\left( 1\right) $, and $\widehat{Var}\left( \hat{g}\left( 1\right) \right) $, which is deferred to the next section for expositional convenience.

Local linear estimation

Theorem (ref) finds that the kernel regression estimator $\hat{\mu}$ for the intercept suffers from a boundary effect in the sense that its order of bias, $O\left( h_{n}\right) $, is larger than that of the nonparametric regression estimator in the interior, which is typically $O\left( h_{n}^{2}\right) $. This is simply because $\hat{\mu}$ is in nature a Nadaraya-Watson or local constant estimator at the boundary. In this section, I resort to the local polynomial regression method fan1996local for bias reduction. In particular, I focus on the local linear estimation because of its asymptotic minimax efficiency properties cheng1997automatic and attractive practical performance gelman2019why.

Denote $\hat{F}_{n}\left( \hat{W}_{i}\right) =\frac{1}{n-1}\sum_{j\neq i}1\left\{ \hat{W}_{j}\leq \hat{W}_{i}\right\} $ as the empirical estimate of $F_{W}\left( W_{i}\right) $ as before. The local linear estimator $\hat{ \mu}^{L}$ for the intercept is defined via locally weighted least squares regression:

equation[equation omitted — 292 chars of source]

To establish the asymptotic properties of $\hat{\mu}^{L}$, Assumption 5' must be modified to accommodate the local linear case.

\noindentAssumption 5”. The bandwidth $h_{n}\in \left( 0,1/2\right] $ satisfies (i) $h_{n}\rightarrow 0$, (ii) $nh_{n}^{5}=O\left( 1\right) $, and (iii) $nh_{n}^{3}\rightarrow \infty $, as $n\rightarrow \infty $.

theoremSuppose Assumptions 1'-4', 5\textquotedblright , and 6 hold. In addition, suppose $G\left( t\right) $ given in ((ref)) is twice continuously differentiable over $t\in \left( 1-\delta ,1\right] $ for a small $\delta >0$, with its first and second derivative functions $g\left( t\right) $ and $g^{\prime }\left( t\right) $ being bounded above over $t\in \left( 1-\delta ,1\right] $. Then, the local linear estimator defined in ( (ref)) is consistent for $\left( \mu _{0},g\left( 1\right) \right) $ and asymptotically normal: \begin{equation*} \left( \begin{array}{cc} \sqrt{nh_{n}} & \\ & \sqrt{nh_{n}^{3}} \end{array} \right) \left[ \left( \begin{array}{c} \hat{\mu}^{L} \\ \hat{b}^{L} \end{array} \right) -\left( \begin{array}{c} \mu _{0} \\ g\left( 1\right) \end{array} \right) +\left( \begin{array}{c} \frac{\left( \kappa _{1}\kappa _{3}-\kappa _{2}^{2}\right) g^{\prime }\left( 1\right) }{2\left( \kappa _{0}\kappa _{2}-\kappa _{1}^{2}\right) }h_{n}^{2} \\ \frac{\left( \kappa _{0}\kappa _{3}-\kappa _{1}\kappa _{2}\right) g^{\prime }\left( 1\right) }{2\left( \kappa _{0}\kappa _{2}-\kappa _{1}^{2}\right) } h_{n} \end{array} \right) \right] \rightarrow N\left( 0,\sigma _{U}^{2}\Omega ^{L}\right) , \end{equation*} where $\sigma _{U}^{2}=EU_{i}^{2}$ and \begin{equation*} \Omega ^{L}=\frac{1}{\left( \kappa _{0}\kappa _{2}-\kappa _{1}^{2}\right) ^{2}}\left( \begin{array}{cc} \kappa _{2}^{2}\chi _{0}+\kappa _{1}^{2}\chi _{2}-2\kappa _{1}\kappa _{2}\chi _{1} & \kappa _{1}\kappa _{2}\chi _{0}+\kappa _{0}\kappa _{1}\chi _{2}-\left( \kappa _{0}\kappa _{2}+\kappa _{1}^{2}\right) \chi _{1} \\ \kappa _{1}\kappa _{2}\chi _{0}+\kappa _{0}\kappa _{1}\chi _{2}-\left( \kappa _{0}\kappa _{2}+\kappa _{1}^{2}\right) \chi _{1} & \kappa _{1}^{2}\chi _{0}+\kappa _{0}^{2}\chi _{2}-2\kappa _{0}\kappa _{1}\chi _{1} \end{array} \right) . \end{equation*}

Theorem (ref) shows that the local linear estimation of the intercept automatically eliminates the boundary effect, and its asymptotic bias is of the same order as that in the interior. More interestingly, the local linear procedure generates a consistent estimate of $g\left( 1\right) $ as a byproduct, which is a key ingredient of the optimal bandwidth for the kernel estimator $\hat{\mu}$. There would be no technical difficulty in extending the results of Theorem (ref) to local polynomial estimation with higher order, except for more complicated notation and more involved mathematical derivations.

Bandwidth selection

Provided $g^{\prime }\left( 1\right) \neq 0$, the optimal bandwidth for $ \hat{\mu}^{L}$ is analogously defined by minimizing the asymptotic MSE of $ \hat{\mu}^{L}$:

eqnarray*[eqnarray* omitted — 586 chars of source]

where

equation[equation omitted — 210 chars of source]

As in Subsection (ref).3, I adopt the regularization method and propose the following estimator for $h_{opt}^{L}$:

equation[equation omitted — 235 chars of source]

To implement the plug-in bandwidth selection procedures ((ref)) and ( (ref)), one must estimate the limits of the derivative functions, $ g\left( 1\right) $ and $g^{\prime }\left( 1\right) $, the regularization terms, $Var\left( \hat{g}\left( 1\right) \right) $ and $Var\left( \hat{g} ^{\prime }\left( 1\right) \right) $, and the variance of the disturbance, $ \sigma _{U}^{2}$. I first construct $\hat{g}\left( 1\right) $ and $\hat{g} ^{\prime }\left( 1\right) $ by fitting a quadratic function to the observations near the boundary:

equation[equation omitted — 378 chars of source]

where $h_{1n}$ is a pilot bandwidth. Similar to Theorem (ref),\ one can show that under $nh_{1n}^{7}=O\left( 1\right) $, $nh_{1n}^{5}\rightarrow \infty $, and the regularity conditions, the local quadratic estimator is consistent and asymptotically normal:

equation*[equation* omitted — 643 chars of source]

As a result, the regularization terms can be estimated by

equation*[equation* omitted — 260 chars of source]

where

eqnarray[eqnarray omitted — 1,717 chars of source]

Last, I estimate the variance of the disturbance by

equation[equation omitted — 341 chars of source]

where $\hat{\mu}^{Q}$ is the initial estimate of $\mu _{0}$ given in ((ref)) and $h_{2n}$ is another pilot bandwidth. Following Theorem (ref), it can be shown that

equation*[equation* omitted — 220 chars of source]

For the pilot bandwidths, simply setting

equation*[equation* omitted — 56 chars of source]

is sufficient to ensure the consistency of $\hat{h}_{opt}$ for $h_{opt}$ and of $\hat{h}_{opt}^{L}$ for $h_{opt}^{L}$. In practice, the suggested bandwidth selection algorithm is fairly robust to the choice of pilot bandwidth, which is not surprising given the presence of the power $1/3$ or $ 1/5$ in the expressions for the optimal bandwidths.

Simulation

This section examines the finite sample properties of the kernel regression estimator $\hat{\mu}$ defined in ((ref)) and the local linear estimator $\hat{\mu}^{L}$ defined in ((ref)), in comparison with the parametric two-step estimator heckman1979sample and the semiparametric identification-at-infinity estimators heckman1990varieties, andrews1998semiparametric. The parametric two-step procedure implements probit estimation for the selection equation in the first step and least squares estimation for the outcome equation with a correction term using the uncensored observations in the second step. This approach is commonly applied in empirical studies due to its computational ease. However, it is likely to be inconsistent when the true distribution of the model disturbance is nonnormal. In contrast, the consistency of heckman1990varieties's estimator (henceforth the Heckman estimator) and andrews1998semiparametric's estimator (henceforth the AS estimator) does not rely on parametric specification of the disturbance distribution, but it is difficult to choose an appropriate smoothing parameter for these estimators. To investigate the robustness of their practical performance, a wide range of smoothing parameters is considered. Following the literature, the choices considered are based on the percentage of uncensored observations used in the estimation. Specifically, the smoothing parameter takes the values of various quantiles of the selection linear index in the uncensored subsample.

The simulation setting mainly follows schafgans2004finite, and the data generating process is

eqnarray*[eqnarray* omitted — 173 chars of source]

where $\mu _{0}=0$, $X_{1i}$ follows the standard normal distribution, $ X_{2i}$ follows the standardized (zero-mean, unit-variance) Student's $t$ distribution with three degrees of freedom, $\varepsilon _{i}$ and $U_{i}$ are zero-mean random variables described below, and only $\left( Y_{i},D_{i},X_{1i},X_{2i}\right) $ is observed. In the simulation, the outcome equation does not contain any nonconstant regressors, implying that the intercept $\mu _{0}$ of primary concern represents the population mean of the latent outcome $Y_{i}^{\ast }$. Different designs are constructed by varying the disturbance distribution and the value of $c_{0}$. $\varepsilon _{i}$ follows three different distributions, namely, the standard normal distribution, the standardized Student's $t$ distribution with three degrees of freedom, and the standardized chi-square distribution with three degrees of freedom. $U_{i}$ is generated by $U_{i}=\varepsilon _{i}+e_{i}$, where $ e_{i}$ is a standard normal random variable independent of $\varepsilon _{i}$ . The constant $c_{0}$ controls for the amount of censoring. In the benchmark design, $c_{0}=0$, producing approximately 50% censoring. Different values of $c_{0}$ are chosen so that $\Pr \left( c_{0}+X_{1i}+X_{2i}>\varepsilon _{i}\right) $ is equal to 0.8 and 0.2, corresponding to 20% and 80% proportions of zero observations, respectively. The sample size $n$ is set to 250, 1000, 4000, and the simulation is replicated 1000 times for each design.

Before calculating the semiparametric estimators for $\mu _{0}$, a distribution-free estimate for the selection equation is necessary. Since the simulation results of schafgans2004finite show little sensitivity to the particular choice of this estimate, I employ the average derivative estimation powell1989semiparametric for computational convenience. When implementing the Heckman and AS estimators, the smoothing parameter is equal to the 0.99, 0.95, 0.9, 0.8, 0.7, and 0.5 quantiles of $\hat{W}_{i}= \hat{\beta}_{1}X_{1i}+\hat{\beta}_{2}X_{2i}$ in the uncensored subsample, corresponding to 1%, 5%, 10%, 20%, 30%, and 50% uncensored observations used in the estimation. The value of the smoothing parameter declines with the proportion of uncensored observations. For consistency, the smoothing parameter is required to approach infinity as $n$ goes to infinity such that the estimation is based on only the observations for which $\Pr \left( \left. D_{i}=1\right\vert X_{i}\right) $ is close to one and in the limit is equal to one. Following the suggestion of andrews1998semiparametric, the smoothed function in the AS estimator is

equation*[equation* omitted — 218 chars of source]

where $b$ is set equal to 1. Note that the AS estimator with $b=0$ is equivalent to the Heckman estimator. For the proposed kernel-type estimators, the bandwidth is chosen by the plug-in algorithm given in Subsection (ref).1. The commonly used Gaussian and Epanechnikov kernel functions are applied. However, the Gaussian kernel is not compactly supported, and the Epanechnikov kernel is not smooth. Therefore, I also consider the polynomial and polyweight kernels, both of degree seven, which possess sufficient smoothness required by Assumption 2'. The definition of these kernel functions and their relevant functionals are presented in Table (ref).

sidewaystable[ptb] \caption{Several kernel functions and their relevant functionals} \resizebox{\linewidth}{!}{ \begin{tabular}{C{2.2cm}ccc} \hline\hline Kernel & Definition & $\kappa _{r}=\int_{0}^{\infty }t^{r}k\left( t\right) dt $, $r\in \mathbb{N}$ & $\chi _{r}=\int_{0}^{\infty }t^{r}k^{2}\left( t\right) dt$, $r\in \mathbb{N}$ \\ \hline Gaussian & $\displaystyle k\left( t\right) =\frac{1}{\sqrt{2\pi }}\exp \left( -\frac{t^{2}}{2}\right)1\left\{t\geq 0\right\} $ & $\displaystyle\sqrt{\frac{2^{r-2}}{\pi }} \gamma \left( \frac{r+1}{2}\right) $ & $\displaystyle\frac{1}{4\pi }\gamma \left( \frac{r+1}{2}\right) $ \\ Epanechnikov & $\displaystyle k\left( t\right) =\left\{ \QATOP{1-t^{2}\text{ if }0\leq t\leq 1}{0\text{ \ if }t>1}\right. $ & $\displaystyle\frac{2}{ \left( r+1\right) \left( r+3\right) }$ & $\displaystyle\frac{8}{\left( r+1\right) \left( r+3\right) \left( r+5\right) }$ \\ Polynomial of degree 7 & $\displaystyle k\left( t\right) =\left\{ \QATOP{ \left( 1-t\right) ^{7}\text{ if }0\leq t\leq 1}{\text{ \ }0\text{\ \ \ if } t>1}\right. $ & $\displaystyle B\left( r+1,8\right) =\frac{7!r!}{\left( r+8\right) !}$ & $\displaystyle B\left( r+1,15\right) =\frac{14!r!}{\left( r+15\right) !}$ \\ Polyweight of degree 7 & $\displaystyle k\left( t\right) =\left\{ \QATOP{ \left( 1-t^{2}\right) ^{7}\text{ if }0\leq t\leq 1}{\text{ \ }0\text{\ \ \ \ if }t>1}\right. $ & $\displaystyle\frac{1}{2}B\left( \frac{r+1}{2},8\right) = \frac{2^{7}7!\left( r-1\right) !!}{\left( r+15\right) !!}$ & $\displaystyle \frac{1}{2}B\left( \frac{r+1}{2},15\right) =\frac{2^{14}14!\left( r-1\right) !!}{\left( r+29\right) !!}$ \\ \hline Kernel & $c_{k}^{L}$ in ((ref)) & $\Omega _{22}^{Q}$ in ((ref) ) & $\Omega _{33}^{Q}$ in ((ref)) \\ \hline Gaussian & $\left[ \frac{\left( \pi +1-2\sqrt{2}\right) \sqrt{\pi }}{\left( 4-\pi \right) ^{2}}\right] ^{1/5}\approx \allowbreak 1.\allowbreak 259$ & $ \frac{\left( 4\pi +11-12\sqrt{2}\right) \sqrt{\pi }}{8\left( \pi -3\right) ^{2}}\approx \allowbreak 72.\allowbreak 89$ & $\frac{3\pi ^{2}-\left( 16-4 \sqrt{2}\right) \pi +\left( 44-24\sqrt{2}\right) }{16\sqrt{\pi }\left( \pi -3\right) ^{2}}\approx \allowbreak 12.\allowbreak 62$ \\ Epanechnikov & $3.200$ & $4913.0$ & $6327.8$ \\ Polynomial of degree 7 & $8.175$ & $8477.5$ & $48857.0$ \\ Polyweight of degree 7 & $5.396$ & $8645.0$ & $29139.5$ \\ \hline\hline \end{tabular} } { {Note: $r!!$, $\gamma \left( r\right) $, and $B\left( r_{1},r_{2}\right) $ denote the double factorial, gamma function, and beta function, respectively.} }

The summary statistics for the simulation are the estimators' bias, standard deviation (SD), root mean squared error (RMSE) ratio, and rejection rate of the t test, over 1000 replications. The RMSE can be calculated from the bias and SD and is thus omitted from the tables. Instead, the RMSE ratio defined by the RMSE of the semiparametric estimators over that of the parametric two-step estimator is reported. Although the RMSE (ratio) is the primary criterion used to compare consistent estimators, it is useless if there are both consistent and inconsistent estimators. An inconsistent estimator having a small RMSE implies that the estimator is highly concentrated in a narrow interval centered at a biased value, consequently leading to incorrect inference. As a complement to the RMSE (ratio), I consider the simulated probability of rejecting the null hypothesis $H_{0}:\mu _{0}=0$ against $H_{1}:\mu _{0}\neq 0$ at a 5% level of significance using t tests (or, more precisely, z tests), based on the asymptotic variances given in heckman1979sample, schafgans2002intercept, andrews1998semiparametric, and Theorems (ref)-(ref).

Tables (ref)-(ref) report the simulation results when the model disturbance follows a normal distribution, a $t\left( 3\right) $ distribution that is symmetric but fat-tailed, and a $\chi ^{2}\left( 3\right) $ distribution that is skewed, respectively, under approximately 50% censoring. Table (ref) shows that, under normal disturbance, the parametric two-step estimator is asymptotically unbiased, as expected, and converges at a $\sqrt{n}$ rate (its SD halves when the sample size quadruples). In contrast, all the considered semiparametric estimators converge at slower than $\sqrt{n}$ rates, as their RMSE ratios increase with the sample size. For the Heckman and AS estimators, the bias increases and the SD decreases when the proportion of uncensored observations used in the estimation increases or, equivalently, the smoothing parameter decreases. The optimal smoothing parameter in terms of RMSE depends upon the estimator, the sample size, and, by comparing across Tables (ref)-(ref), the disturbance distribution. If the smoothing parameter is improperly chosen, the RMSE may be several times larger than the smallest RMSE, and the rejection rate may be far larger than the specified level of significance. Moreover, the optimal smoothing parameter in terms of RMSE usually disagrees with that in terms of the rejection rate. For instance, in the case of $n=1000$ for the Heckman estimator and the case of $n=4000$ for the AS estimator, the optimal smoothing parameter in terms of RMSE would use 30% uncensored observations in the estimation, but the corresponding simulated rejection rates are more than twice the real level. In these two cases, a more sensible choice would be to use 20% uncensored observations, sacrificing a little RMSE but leading to a rejection rate very close to 0.05. In practice, however, the RMSE and rejection rate are not known; therefore, we are never aware of whether we have made a good choice.

sidewaystable[ptb] \caption{Simulation results when $\varepsilon _{i}\sim N\left(0,1\right) $} \resizebox{\linewidth}{!}{ \begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}} \hline\hline $\Pr \left( Y_{i}=0\right) =0.5$ & & \multicolumn{4}{c}{$n=250$} & & \multicolumn{4}{c}{$n=1000$} & & \multicolumn{4}{c}{$n=4000$} \\ \cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16} Estimator & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate \\ \hline \multicolumn{16}{l}{heckman1979sample's parametric two-step estimator} \\ & & -0.004 & 0.186 & 1 & 0.052 & & -0.002 & 0.093 & 1 & 0.048 & & 0.000 & 0.044 & 1 & 0.043 \\ \multicolumn{16}{l}{heckman1990varieties's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.007 & 1.025 & 5.519 & 0.390 & & 0.028 & 0.599 & 6.444 & 0.128 & & 0.006 & 0.301 & 6.812 & 0.062 \\ $5\%$ observations & & -0.003 & 0.549 & 2.955 & 0.125 & & 0.001 & 0.286 & 3.074 & 0.076 & & -0.006 & 0.140 & 3.173 & 0.056 \\ $10\%$ observations & & -0.016 & 0.389 & 2.093 & 0.088 & & -0.021 & 0.195 & 2.108 & 0.064 & & -0.016 & 0.096 & 2.209 & 0.062 \\ $20\%$ observations & & -0.046 & 0.275 & 1.503 & 0.070 & & -0.049 & 0.136 & 1.558 & 0.059 & & -0.044 & 0.067 & 1.801 & 0.088 \\ $30\%$ observations & & -0.077 & 0.228 & 1.294 & 0.082 & & -0.085 & 0.111 & 1.499 & 0.114 & & -0.079 & 0.055 & 2.185 & 0.278 \\ $50\%$ observations & & -0.168 & 0.175 & 1.306 & 0.178 & & -0.164 & 0.084 & 1.982 & 0.518 & & -0.164 & 0.041 & 3.821 & 0.983 \\ \multicolumn{16}{l}{andrews1998semiparametric's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.055 & 1.302 & 7.013 & 0.616 & & 0.037 & 0.697 & 7.498 & 0.144 & & 0.010 & 0.341 & 7.710 & 0.065 \\ $5\%$ observations & & -0.007 & 0.689 & 3.711 & 0.165 & & 0.007 & 0.338 & 3.632 & 0.077 & & -0.005 & 0.165 & 3.737 & 0.048 \\ $10\%$ observations & & -0.012 & 0.480 & 2.585 & 0.106 & & -0.007 & 0.241 & 2.589 & 0.074 & & -0.008 & 0.116 & 2.628 & 0.050 \\ $20\%$ observations & & -0.027 & 0.331 & 1.787 & 0.076 & & -0.029 & 0.163 & 1.775 & 0.064 & & -0.024 & 0.079 & 1.857 & 0.051 \\ $30\%$ observations & & -0.045 & 0.263 & 1.438 & 0.076 & & -0.052 & 0.130 & 1.508 & 0.070 & & -0.046 & 0.063 & 1.769 & 0.101 \\ $50\%$ observations & & -0.106 & 0.198 & 1.211 & 0.111 & & -0.110 & 0.096 & 1.569 & 0.214 & & -0.106 & 0.047 & 2.629 & 0.596 \\ \multicolumn{16}{l}{Kernel regression (local constant) estimator with various kernel functions} \\ Gaussian & & -0.031 & 0.302 & 1.635 & 0.058 & & -0.027 & 0.165 & 1.797 & 0.072 & & -0.016 & 0.090 & 2.073 & 0.058 \\ Epanechnikov & & -0.009 & 0.550 & 2.959 & 0.038 & & 0.004 & 0.312 & 3.357 & 0.052 & & -0.005 & 0.171 & 3.874 & 0.042 \\ 7th polynomial & & -0.017 & 0.653 & 3.514 & 0.064 & & 0.012 & 0.357 & 3.843 & 0.052 & & -0.001 & 0.192 & 4.333 & 0.044 \\ 7th polyweight & & -0.012 & 0.605 & 3.259 & 0.046 & & 0.009 & 0.344 & 3.696 & 0.049 & & -0.004 & 0.187 & 4.222 & 0.039 \\ \multicolumn{16}{l}{Local linear estimator with various kernel functions} \\ Gaussian & & 0.032 & 0.284 & 1.539 & 0.092 & & 0.029 & 0.148 & 1.616 & 0.091 & & 0.032 & 0.077 & 1.894 & 0.102 \\ Epanechnikov & & 0.019 & 0.424 & 2.284 & 0.057 & & 0.014 & 0.235 & 2.535 & 0.053 & & 0.014 & 0.125 & 2.843 & 0.045 \\ 7th polynomial & & -0.001 & 0.553 & 2.979 & 0.089 & & 0.015 & 0.312 & 3.353 & 0.076 & & 0.006 & 0.167 & 3.791 & 0.053 \\ 7th polyweight & & 0.005 & 0.491 & 2.643 & 0.065 & & 0.013 & 0.278 & 2.989 & 0.059 & & 0.008 & 0.151 & 3.424 & 0.063 \\ \hline\hline \end{tabular} }
sidewaystable[ptb] \caption{Simulation results when $\varepsilon _{i}\sim t\left(3\right) $} \resizebox{\linewidth}{!}{ \begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}} \hline\hline $\Pr \left( Y_{i}=0\right) =0.5$ & & \multicolumn{4}{c}{$n=250$} & & \multicolumn{4}{c}{$n=1000$} & & \multicolumn{4}{c}{$n=4000$} \\ \cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16} Estimator & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate \\ \hline \multicolumn{16}{l}{heckman1979sample's parametric two-step estimator} \\ & & 0.014 & 0.181 & 1 & 0.059 & & 0.027 & 0.089 & 1 & 0.058 & & 0.025 & 0.047 & 1 & 0.114 \\ \multicolumn{16}{l}{heckman1990varieties's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.052 & 0.981 & 5.403 & 0.386 & & 0.003 & 0.592 & 6.363 & 0.155 & & -0.001 & 0.296 & 5.594 & 0.069 \\ $5\%$ observations & & -0.050 & 0.533 & 2.945 & 0.121 & & -0.030 & 0.271 & 2.934 & 0.070 & & -0.029 & 0.135 & 2.603 & 0.056 \\ $10\%$ observations & & -0.049 & 0.367 & 2.035 & 0.073 & & -0.033 & 0.186 & 2.029 & 0.056 & & -0.039 & 0.097 & 1.975 & 0.081 \\ $20\%$ observations & & -0.061 & 0.260 & 1.468 & 0.063 & & -0.052 & 0.133 & 1.532 & 0.065 & & -0.056 & 0.067 & 1.654 & 0.137 \\ $30\%$ observations & & -0.076 & 0.210 & 1.226 & 0.073 & & -0.068 & 0.105 & 1.351 & 0.092 & & -0.072 & 0.054 & 1.701 & 0.274 \\ $50\%$ observations & & -0.120 & 0.163 & 1.115 & 0.113 & & -0.109 & 0.080 & 1.452 & 0.257 & & -0.111 & 0.042 & 2.238 & 0.770 \\ \multicolumn{16}{l}{andrews1998semiparametric's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.089 & 1.285 & 7.083 & 0.627 & & 0.006 & 0.711 & 7.646 & 0.197 & & 0.005 & 0.346 & 6.541 & 0.072 \\ $5\%$ observations & & -0.059 & 0.659 & 3.641 & 0.137 & & -0.018 & 0.325 & 3.503 & 0.069 & & -0.022 & 0.157 & 2.999 & 0.060 \\ $10\%$ observations & & -0.054 & 0.476 & 2.633 & 0.089 & & -0.028 & 0.227 & 2.460 & 0.060 & & -0.031 & 0.113 & 2.209 & 0.068 \\ $20\%$ observations & & -0.051 & 0.311 & 1.733 & 0.069 & & -0.039 & 0.156 & 1.727 & 0.058 & & -0.045 & 0.080 & 1.727 & 0.087 \\ $30\%$ observations & & -0.061 & 0.246 & 1.395 & 0.063 & & -0.052 & 0.123 & 1.440 & 0.062 & & -0.057 & 0.064 & 1.616 & 0.161 \\ $50\%$ observations & & -0.089 & 0.184 & 1.123 & 0.075 & & -0.080 & 0.093 & 1.319 & 0.142 & & -0.083 & 0.048 & 1.813 & 0.438 \\ \multicolumn{16}{l}{Kernel regression (local constant) estimator with various kernel functions} \\ Gaussian & & -0.061 & 0.278 & 1.564 & 0.051 & & -0.041 & 0.153 & 1.705 & 0.066 & & -0.042 & 0.086 & 1.798 & 0.090 \\ Epanechnikov & & -0.053 & 0.547 & 3.021 & 0.041 & & -0.021 & 0.303 & 3.264 & 0.048 & & -0.020 & 0.163 & 3.099 & 0.058 \\ 7th polynomial & & -0.058 & 0.642 & 3.545 & 0.061 & & -0.013 & 0.355 & 3.824 & 0.059 & & -0.014 & 0.190 & 3.596 & 0.055 \\ 7th polyweight & & -0.055 & 0.605 & 3.341 & 0.037 & & -0.016 & 0.335 & 3.610 & 0.053 & & -0.017 & 0.181 & 3.427 & 0.054 \\ \multicolumn{16}{l}{Local linear estimator with various kernel functions} \\ Gaussian & & -0.020 & 0.270 & 1.488 & 0.072 & & -0.009 & 0.147 & 1.583 & 0.069 & & -0.018 & 0.082 & 1.587 & 0.067 \\ Epanechnikov & & -0.034 & 0.399 & 2.200 & 0.046 & & -0.016 & 0.222 & 2.397 & 0.046 & & -0.022 & 0.125 & 2.403 & 0.060 \\ 7th polynomial & & -0.044 & 0.541 & 2.987 & 0.076 & & -0.013 & 0.301 & 3.238 & 0.071 & & -0.015 & 0.164 & 3.112 & 0.070 \\ 7th polyweight & & -0.039 & 0.474 & 2.616 & 0.050 & & -0.014 & 0.265 & 2.850 & 0.056 & & -0.019 & 0.147 & 2.803 & 0.068 \\ \hline\hline \end{tabular} }
sidewaystable[ptb] \caption{Simulation results when $\varepsilon _{i}\sim \chi^2\left(3\right) $} \resizebox{\linewidth}{!}{ \begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}} \hline\hline $\Pr \left( Y_{i}=0\right) =0.5$ & & \multicolumn{4}{c}{$n=250$} & & \multicolumn{4}{c}{$n=1000$} & & \multicolumn{4}{c}{$n=4000$} \\ \cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16} Estimator & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate \\ \hline \multicolumn{16}{l}{heckman1979sample's parametric two-step estimator} \\ & & -0.146 & 0.167 & 1 & 0.180 & & -0.145 & 0.082 & 1 & 0.459 & & -0.146 & 0.041 & 1 & 0.913 \\ \multicolumn{16}{l}{heckman1990varieties's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.008 & 0.977 & 4.405 & 0.413 & & 0.004 & 0.578 & 3.465 & 0.134 & & -0.024 & 0.299 & 1.978 & 0.076 \\ $5\%$ observations & & -0.016 & 0.514 & 2.321 & 0.125 & & -0.047 & 0.258 & 1.569 & 0.067 & & -0.052 & 0.128 & 0.913 & 0.073 \\ $10\%$ observations & & -0.057 & 0.351 & 1.603 & 0.080 & & -0.084 & 0.183 & 1.206 & 0.087 & & -0.083 & 0.088 & 0.797 & 0.140 \\ $20\%$ observations & & -0.117 & 0.243 & 1.215 & 0.099 & & -0.129 & 0.124 & 1.072 & 0.189 & & -0.126 & 0.059 & 0.917 & 0.530 \\ $30\%$ observations & & -0.162 & 0.200 & 1.160 & 0.152 & & -0.165 & 0.098 & 1.148 & 0.395 & & -0.163 & 0.048 & 1.124 & 0.911 \\ $50\%$ observations & & -0.229 & 0.149 & 1.232 & 0.354 & & -0.230 & 0.073 & 1.445 & 0.869 & & -0.229 & 0.037 & 1.535 & 1.000 \\ \multicolumn{16}{l}{andrews1998semiparametric's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.009 & 1.219 & 5.497 & 0.605 & & 0.021 & 0.694 & 4.162 & 0.180 & & -0.009 & 0.350 & 2.315 & 0.078 \\ $5\%$ observations & & -0.009 & 0.624 & 2.816 & 0.153 & & -0.026 & 0.315 & 1.893 & 0.067 & & -0.040 & 0.156 & 1.061 & 0.068 \\ $10\%$ observations & & -0.037 & 0.440 & 1.990 & 0.102 & & -0.057 & 0.217 & 1.343 & 0.066 & & -0.063 & 0.108 & 0.822 & 0.097 \\ $20\%$ observations & & -0.084 & 0.291 & 1.365 & 0.081 & & -0.099 & 0.148 & 1.068 & 0.113 & & -0.097 & 0.072 & 0.799 & 0.251 \\ $30\%$ observations & & -0.119 & 0.231 & 1.173 & 0.106 & & -0.131 & 0.116 & 1.048 & 0.210 & & -0.127 & 0.056 & 0.919 & 0.609 \\ $50\%$ observations & & -0.182 & 0.171 & 1.124 & 0.206 & & -0.185 & 0.084 & 1.217 & 0.582 & & -0.185 & 0.042 & 1.250 & 0.991 \\ \multicolumn{16}{l}{Kernel regression (local constant) estimator with various kernel functions} \\ Gaussian & & -0.087 & 0.278 & 1.316 & 0.093 & & -0.092 & 0.157 & 1.091 & 0.131 & & -0.079 & 0.088 & 0.783 & 0.199 \\ Epanechnikov & & -0.019 & 0.526 & 2.372 & 0.051 & & -0.028 & 0.302 & 1.817 & 0.067 & & -0.038 & 0.167 & 1.128 & 0.068 \\ 7th polynomial & & -0.015 & 0.624 & 2.816 & 0.071 & & -0.017 & 0.355 & 2.131 & 0.072 & & -0.030 & 0.196 & 1.308 & 0.062 \\ 7th polyweight & & -0.016 & 0.580 & 2.615 & 0.045 & & -0.020 & 0.337 & 2.024 & 0.069 & & -0.033 & 0.186 & 1.245 & 0.066 \\ \multicolumn{16}{l}{Local linear estimator with various kernel functions} \\ Gaussian & & -0.066 & 0.258 & 1.202 & 0.095 & & -0.070 & 0.134 & 0.907 & 0.122 & & -0.065 & 0.069 & 0.627 & 0.183 \\ Epanechnikov & & -0.027 & 0.387 & 1.751 & 0.056 & & -0.044 & 0.222 & 1.355 & 0.066 & & -0.039 & 0.122 & 0.846 & 0.071 \\ 7th polynomial & & -0.007 & 0.530 & 2.392 & 0.090 & & -0.019 & 0.302 & 1.813 & 0.072 & & -0.026 & 0.167 & 1.116 & 0.071 \\ 7th polyweight & & -0.012 & 0.466 & 2.102 & 0.063 & & -0.029 & 0.266 & 1.605 & 0.059 & & -0.031 & 0.148 & 0.997 & 0.072 \\ \hline\hline \end{tabular} }

By contrast, the kernel regression and local linear estimators have no difficulty in choosing the bandwidth parameter because fully data-driven procedures of selecting the optimal bandwidths have been explicitly proposed. Table (ref) shows that the proposed bandwidths perform satisfactorily in that the rejection rates for the kernel regression and local linear estimators are all close to 0.05 for different sample sizes and different kernel functions. In terms of RMSE, the local linear estimator is superior to the kernel regression estimator, as expected. The RMSE of the local linear estimator with Gaussian kernel is comparable to the optimal RMSE of the Heckman and AS estimators. However, the Gaussian kernel may not be the best choice because over-rejection for the t test appears to be a problem. As an alternative, the local linear estimator with Epanechnikov kernel achieves a satisfactory balance between RMSE and rejection rate. The polynomial and polyweight kernels also do not lead to size distortion of the t test but are clearly outperformed by the simple Epanechnikov kernel in terms of RMSE.

Table (ref) investigates the finite sample behavior of the estimators under $t\left( 3\right) $ distributed disturbance. The parametric two-step estimator performs well for small sample sizes but deteriorates as the sample size increases because its nonvanishing bias due to nonnormality becomes large relative to its declining SD. By contrast, the performance of the semiparametric estimators is robust to nonnormality, and the main findings are almost the same as in the normal case. First, if the smoothing parameter of the Heckman and AS estimators is improperly chosen, their RMSEs may be increased by several factors and the t tests may yield misleading inferences. Second, the smoothing parameter giving rise to the smallest RMSE may lead to evident over-rejection for the t test. These facts indicate the importance of selecting a proper smoothing parameter for the identification-at-infinity estimators, but this task is difficult since the RMSE and rejection rate are unobserved in practice. Third, the bandwidth selection algorithms proposed for the kernel-type estimators perform well in terms of rejection rate. Fourth, the local linear estimator dominates the kernel regression estimator. Fifth, the RMSE of the local linear estimator with Gaussian kernel is comparable to (when $n=4000$, even smaller than) the smallest RMSE of the identification-at-infinity estimators. Sixth, using the Epanechnikov kernel results in more accurate rejection rate.

Table (ref) considers the case of $\chi ^{2}\left( 3\right) $ distributed disturbance. In this case, the parametric two-step estimator has notable bias, and the probability of making a type I error rapidly approaches one. Comparison with Table (ref) indicates that the skewness of the nonnormal disturbance exerts a worse influence on the parametric estimator than does the fat tail. The Heckman and AS estimators become less robust to the smoothing parameter under this design, in the sense that the rejection rate is below 0.1 for only a narrow range of smoothing parameters. On the other hand, the local linear estimator with Epanechnikov kernel still has desirable finite sample properties in terms of both RMSE and rejection rate.

Tables (ref)-(ref) investigate the effect of the proportion of censoring. Table (ref) shows that, under mild censoring, all the considered estimators behave better. The parametric two-step estimator becomes less biased in nonnormal designs. The Heckman and AS estimators are more robust to the smoothing parameter, and the rejection rates for the kernel regression and local linear estimators are closer to the true level, even when the Gaussian kernel is used. Table (ref) reveals the reverse side of the coin, with an unchanged conclusion being the superiority of the local linear estimator with Epanechnikov kernel.

sidewaystable[ptb] \caption{Simulation results when $n=1000$ and the amount of zero observations is small} \resizebox{\linewidth}{!}{ \begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}} \hline\hline $\Pr \left( Y_{i}=0\right) =0.2$ & & \multicolumn{4}{c}{$\varepsilon _{i}\sim N\left( 0,1\right) $} & & \multicolumn{4}{c}{$\varepsilon _{i}\sim t\left( 3\right) $} & & \multicolumn{4}{c}{$\varepsilon _{i}\sim \chi ^{2}\left( 3\right) $} \\ \cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16} Estimator & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate \\ \hline \multicolumn{16}{l}{heckman1979sample's parametric two-step estimator} \\ & & -0.001 & 0.065 & 1 & 0.055 & & -0.015 & 0.068 & 1 & 0.070 & & -0.071 & 0.066 & 1 & 0.239 \\ \multicolumn{16}{l}{heckman1990varieties's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.006 & 0.483 & 7.447 & 0.103 & & -0.014 & 0.466 & 6.719 & 0.099 & & -0.004 & 0.477 & 4.920 & 0.102 \\ $5\%$ observations & & 0.006 & 0.224 & 3.449 & 0.065 & & -0.017 & 0.218 & 3.145 & 0.061 & & -0.021 & 0.206 & 2.134 & 0.057 \\ $10\%$ observations & & 0.002 & 0.163 & 2.510 & 0.063 & & -0.027 & 0.153 & 2.231 & 0.049 & & -0.036 & 0.145 & 1.545 & 0.054 \\ $20\%$ observations & & -0.001 & 0.113 & 1.743 & 0.056 & & -0.035 & 0.132 & 1.970 & 0.058 & & -0.056 & 0.106 & 1.242 & 0.081 \\ $30\%$ observations & & -0.012 & 0.095 & 1.474 & 0.057 & & -0.042 & 0.102 & 1.589 & 0.082 & & -0.075 & 0.087 & 1.183 & 0.153 \\ $50\%$ observations & & -0.036 & 0.073 & 1.252 & 0.093 & & -0.054 & 0.074 & 1.317 & 0.130 & & -0.111 & 0.067 & 1.336 & 0.434 \\ \multicolumn{16}{l}{andrews1998semiparametric's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.006 & 0.573 & 8.841 & 0.126 & & 0.002 & 0.536 & 7.723 & 0.109 & & -0.016 & 0.573 & 5.917 & 0.136 \\ $5\%$ observations & & 0.002 & 0.264 & 4.077 & 0.061 & & -0.012 & 0.258 & 3.714 & 0.056 & & -0.010 & 0.258 & 2.661 & 0.070 \\ $10\%$ observations & & 0.006 & 0.190 & 2.924 & 0.061 & & -0.021 & 0.182 & 2.642 & 0.061 & & -0.026 & 0.172 & 1.797 & 0.051 \\ $20\%$ observations & & 0.000 & 0.132 & 2.037 & 0.060 & & -0.030 & 0.130 & 1.918 & 0.060 & & -0.042 & 0.121 & 1.325 & 0.064 \\ $30\%$ observations & & -0.005 & 0.107 & 1.651 & 0.054 & & -0.036 & 0.116 & 1.755 & 0.074 & & -0.057 & 0.100 & 1.185 & 0.095 \\ $50\%$ observations & & -0.020 & 0.081 & 1.285 & 0.066 & & -0.046 & 0.088 & 1.423 & 0.091 & & -0.088 & 0.074 & 1.189 & 0.243 \\ \multicolumn{16}{l}{Kernel regression (local constant) estimator with various kernel functions} \\ Gaussian & & 0.002 & 0.160 & 2.466 & 0.051 & & -0.025 & 0.155 & 2.259 & 0.063 & & -0.031 & 0.147 & 1.553 & 0.049 \\ Epanechnikov & & -0.002 & 0.302 & 4.663 & 0.043 & & -0.011 & 0.300 & 4.322 & 0.054 & & -0.005 & 0.305 & 3.150 & 0.059 \\ 7th polynomial & & -0.004 & 0.345 & 5.321 & 0.056 & & -0.007 & 0.342 & 4.920 & 0.057 & & -0.010 & 0.359 & 3.711 & 0.074 \\ 7th polyweight & & -0.004 & 0.330 & 5.094 & 0.043 & & -0.012 & 0.329 & 4.748 & 0.056 & & -0.005 & 0.337 & 3.482 & 0.063 \\ \multicolumn{16}{l}{Local linear estimator with various kernel functions} \\ Gaussian & & 0.017 & 0.143 & 2.222 & 0.070 & & -0.024 & 0.147 & 2.151 & 0.067 & & -0.020 & 0.128 & 1.334 & 0.057 \\ Epanechnikov & & 0.005 & 0.234 & 3.608 & 0.057 & & -0.016 & 0.227 & 3.280 & 0.053 & & -0.011 & 0.217 & 2.247 & 0.044 \\ 7th polynomial & & 0.003 & 0.301 & 4.635 & 0.056 & & -0.008 & 0.299 & 4.310 & 0.063 & & -0.005 & 0.304 & 3.135 & 0.077 \\ 7th polyweight & & 0.008 & 0.274 & 4.227 & 0.054 & & -0.010 & 0.273 & 3.933 & 0.056 & & -0.008 & 0.265 & 2.734 & 0.061 \\ \hline\hline \end{tabular} }
sidewaystable[ptb] \caption{Simulation results when $n=1000$ and the amount of zero observations is large} \resizebox{\linewidth}{!}{ \begin{tabular}{ccccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}cccC{1.2cm}C{1.5cm}} \hline\hline $\Pr \left( Y_{i}=0\right) =0.8$ & & \multicolumn{4}{c}{$\varepsilon _{i}\sim N\left( 0,1\right) $} & & \multicolumn{4}{c}{$\varepsilon _{i}\sim t\left( 3\right) $} & & \multicolumn{4}{c}{$\varepsilon _{i}\sim \chi ^{2}\left( 3\right) $} \\ \cline{1-1}\cline{3-6}\cline{8-11}\cline{13-16} Estimator & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate & & Bias & SD & RMSE ratio & Rejection rate \\ \hline \multicolumn{16}{l}{heckman1979sample's parametric two-step estimator} \\ & & 0.001 & 0.168 & 1 & 0.045 & & 0.195 & 0.231 & 1 & 0.193 & & -0.234 & 0.132 & 1 & 0.428 \\ \multicolumn{16}{l}{heckman1990varieties's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.040 & 0.889 & 5.285 & 0.316 & & 0.030 & 0.879 & 2.907 & 0.343 & & -0.067 & 0.855 & 3.186 & 0.325 \\ $5\%$ observations & & -0.035 & 0.403 & 2.404 & 0.086 & & -0.044 & 0.415 & 1.380 & 0.097 & & -0.081 & 0.391 & 1.485 & 0.098 \\ $10\%$ observations & & -0.066 & 0.289 & 1.762 & 0.080 & & -0.056 & 0.292 & 0.984 & 0.064 & & -0.133 & 0.277 & 1.141 & 0.114 \\ $20\%$ observations & & -0.154 & 0.208 & 1.539 & 0.149 & & -0.099 & 0.212 & 0.772 & 0.089 & & -0.212 & 0.188 & 1.053 & 0.219 \\ $30\%$ observations & & -0.237 & 0.166 & 1.716 & 0.324 & & -0.143 & 0.171 & 0.737 & 0.145 & & -0.269 & 0.148 & 1.140 & 0.445 \\ $50\%$ observations & & -0.388 & 0.124 & 2.419 & 0.876 & & -0.228 & 0.134 & 0.875 & 0.452 & & -0.362 & 0.113 & 1.409 & 0.891 \\ \multicolumn{16}{l}{andrews1998semiparametric's semiparametric estimator with various smoothing parameters} \\ $1\%$ observations & & -0.021 & 1.085 & 6.449 & 0.428 & & 0.024 & 1.011 & 3.345 & 0.442 & & -0.062 & 1.039 & 3.869 & 0.436 \\ $5\%$ observations & & -0.022 & 0.497 & 2.955 & 0.105 & & -0.032 & 0.487 & 1.614 & 0.104 & & -0.056 & 0.485 & 1.814 & 0.108 \\ $10\%$ observations & & -0.039 & 0.347 & 2.072 & 0.075 & & -0.048 & 0.353 & 1.178 & 0.077 & & -0.094 & 0.335 & 1.294 & 0.094 \\ $20\%$ observations & & -0.093 & 0.243 & 1.545 & 0.080 & & -0.070 & 0.251 & 0.860 & 0.071 & & -0.159 & 0.230 & 1.040 & 0.136 \\ $30\%$ observations & & -0.155 & 0.195 & 1.480 & 0.150 & & -0.098 & 0.203 & 0.744 & 0.088 & & -0.209 & 0.181 & 1.027 & 0.233 \\ $50\%$ observations & & -0.282 & 0.144 & 1.878 & 0.526 & & -0.162 & 0.152 & 0.736 & 0.199 & & -0.291 & 0.131 & 1.187 & 0.610 \\ \multicolumn{16}{l}{Kernel regression (local constant) estimator with various kernel functions} \\ Gaussian & & -0.142 & 0.216 & 1.535 & 0.207 & & -0.110 & 0.193 & 0.734 & 0.156 & & -0.201 & 0.198 & 1.050 & 0.349 \\ Epanechnikov & & -0.052 & 0.306 & 1.843 & 0.062 & & -0.055 & 0.296 & 0.994 & 0.066 & & -0.115 & 0.294 & 1.174 & 0.098 \\ 7th polynomial & & -0.041 & 0.355 & 2.122 & 0.069 & & -0.043 & 0.343 & 1.142 & 0.062 & & -0.094 & 0.359 & 1.380 & 0.094 \\ 7th polyweight & & -0.042 & 0.335 & 2.005 & 0.062 & & -0.050 & 0.325 & 1.087 & 0.060 & & -0.098 & 0.329 & 1.276 & 0.089 \\ \multicolumn{16}{l}{Local linear estimator with various kernel functions} \\ Gaussian & & -0.101 & 0.188 & 1.264 & 0.224 & & -0.012 & 0.182 & 0.602 & 0.139 & & -0.169 & 0.181 & 0.920 & 0.328 \\ Epanechnikov & & -0.035 & 0.244 & 1.467 & 0.087 & & -0.023 & 0.230 & 0.764 & 0.068 & & -0.134 & 0.217 & 0.947 & 0.136 \\ 7th polynomial & & -0.005 & 0.317 & 1.885 & 0.082 & & -0.022 & 0.306 & 1.015 & 0.078 & & -0.087 & 0.306 & 1.182 & 0.104 \\ 7th polyweight & & -0.011 & 0.284 & 1.686 & 0.081 & & -0.023 & 0.272 & 0.904 & 0.067 & & -0.106 & 0.263 & 1.054 & 0.112 \\ \hline\hline \end{tabular} }

Conclusion

This paper rephrases the identification at infinity into an identification at the boundary via a CDF transformation and accordingly proposes a kernel approach to semiparametrically estimate the intercept of the sample selection model. The proposed kernel regression estimator with generic transformation is a generalization of the identification-at-infinity estimators and thus inherits the disadvantage that the asymptotic bias and variance are implicit functions of the bandwidth parameter. To select a bandwidth that minimizes the asymptotic mean squared error, I use a specific transformation, namely, the empirical CDF of the selection index, under which the asymptotic bias and variance become explicit with respect to the bandwidth. A plug-in bandwidth selection algorithm with regularization is therefore suggested. For the purpose of bias reduction, I further propose a local linear estimator and an associated analogous bandwidth selection algorithm. A simulation study illustrates the effectiveness of the selected bandwidths. Comparison of the finite sample performance of the estimators indicates that the local linear estimator with Epanechnikov kernel is superior to the parametric two-step estimator under nonnormal disturbance and to the identification-at-infinity estimators in most cases.

\setcounter{theorem}{0}