EconBase
← Back to paper

Nonseparable Sample Selection Models with Censored Selection Rules

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

112,655 characters · 16 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonseparable Sample Selection Models with Censored Selection Rules

abstractWe consider identification and estimation of nonseparable sample selection models with censored selection rules. We employ a control function approach and discuss different objects of interest based on (1) local effects conditional on the control function, and (2) global effects obtained from integration over ranges of values of the control function. We derive the conditions for the identification of these different objects and suggest strategies for estimation. Moreover, we provide the associated asymptotic theory. These strategies are illustrated in an empirical investigation of the determinants of female wages in the United Kingdom. \thispagestyle{empty}

\jelcodes{C14, C21, C24}

\setcounter{page}{1}

Introduction

This paper considers a nonseparable sample selection model with a censored selection rule. The most common example is a selection rule with censoring at zero, also referred to in the parametric setting as tobit type 3, although other forms of censored selection rules are permissible. A leading empirical example is estimating the determinants of wages when workers report working hours rather than the binary work/not work decision. We account for selection via an appropriately constructed control function and propose a three-step estimation procedure which first employs distribution regression to compute the appropriate control function. The second step is estimated by least squares, distribution or quantile regression employing the estimated control function as an additional conditioning variable. The estimands of interest are obtained in the third step as functionals of the second step and control function estimates.

Our paper contributes to the growing literatures on nonseparable models with endogeneity (see, for example, Chesher 2003, Ma and Koenker 2006, Florens et al. 2008, Imbens and Newey 2009, Jun 2009 and Masten and Torgovitsky, 2016) and nonseparable sample selection models (for example, Newey 2007). An important focus of our paper is on the identification and estimation of local effects. This local approach to identification is popular in many contexts (see, for example, Chesher 2003, and Heckman and Vytlacil 2005). We show that for any population observation with positive probability of being selected, selection is irrelevant for the distribution of the outcome variable conditional on the control function. Hence, our local objects of interest are identified for the whole population and not just the selected one. We can also estimate global objects by integrating over the distribution of the control function in the selected population. However, global objects are only nonparametrically identified under support assumptions for the explanatory variables which may be difficult to satisfy in empirical applications. Accordingly, we also consider global effects \textquotedblleft for the treated\textquotedblright\ that are identified under weaker assumptions.

We provide estimators of the global and local effects and their associated asymptotic theory. These are based on semiparametric specifications of the first and second steps of the estimation procedure. The semiparametric structure overcomes the curse of dimensionality and the reliance on support assumptions required to nonparametrically identify the global effects (Newey and Stouli, 2018, Chernozhukov et al., 2020). It can also be used as the basis for sieve methods by increasing the dimension of the specification with the sample size (Chen, 2007). We recommend that the estimates of the global effects are supplemented with bounds that only require the semi-parametric structure be correctly specified within the support, whereas the point estimators rely on extrapolation outside the support. We provide estimates for these bounds.

Our paper is closest to Newey (2007) who uses a control function approach to correct for sample selection in a nonseparable model. Nevertheless, there are several differences. First, Newey (2007) examines a binary selection rule whereas we consider a censored selection rule. Second, in addition to considering global parameters for the selected population, as is done in Newey (2007), we also consider local parameters for the entire population and identification of the effects \textquotedblleft for the treated\textquotedblright . Third, we follow the suggestion by Newey (2007) to construct bounds for the global parameters for the entire population. Finally, we provide estimation and inference results in addition to addressing identification issues.

Our paper is also related to Caetano et al. (2016), which studied a nonseparable model with an endogenous regressor featuring a mass point at some known value. This includes censoring as a special case. They build on Caetano (2015) by developing a test for the validity of the control function approach. These papers, and ours, are motivated by the Imbens and Newey (2009) approach although each paper provides a treatment of important but different issues. They focus on specification testing, whereas we analyze the estimation and identification of local effects in the presence of sample selection. We also consider the use of bounds for the global effects. Thus, we view our paper as complementary to Caetano et al. (2016).

This paper also contributes to the literature on quantile selection models. Arellano and Bonhomme (2017) addressed selection by modeling the copula of the error terms in the outcome and selection equations. The most important distinction to this paper is that they consider the conventional binary, rather than a censored, selection equation. Thus, we require more information about the selection process. However, this has the advantage that one can consider local effects conditional on the control function which are identified under weaker conditions.

We acknowledge that the binary selection model is more frequently employed in empirical work than the censored selection model. However, this partially reflects the remarkable popularity of the Heckman (1979) two-step procedure. In fact, in many empirical investigations, researchers dichotomize censored selection variables in order to employ the Heckman procedure. Examples of such commonly encountered censoring selection variables include unemployment duration, training program length and the magnitude or the length of the welfare benefit receipt. In the parametric setting, there are no substantial benefits in using the censored, rather than the dichotomized form of the selection variable other than that the selection variable can appear in the outcome equation as a regressor and the variation in this variable provides an additional form of identification. However, in the nonseparable setting, a censored selection rule enables inference for both the selected and non-selected populations and the investigation of local effects. Finally, we illustrate that the additional information implicit in the censored selection setting can be exploited to reduce the bounds for the global effects. Thus, it could be argued that the censored selection approach should be employed when possible.

The following section outlines the model and some related literature. Section (ref) defines the control function and provides identification results regarding the objects of interest in the model. Section (ref) provides estimators of these objects and discusses inference. Section (ref) illustrates some of our estimands focusing on the determinants of female wages in the United Kingdom.

Model

The model has the following structure:

eqnarray[eqnarray omitted — 123 chars of source]

where $Y$ and $C$ are observable random variables, and $X$, $Z_{1}$ and $ Z:=(Z_{1},Z_{2})$ are vectors of observable explanatory variables. The variables included in $X$ are a subset of those included in $Z_{2}$, i.e. $X \subseteq Z_2$. We do not need to impose an exclusion restriction on $Z_{2}$ with respect to the elements of $X$, although our nonparametric identification assumptions will be more plausible with such a restriction. We separate $X$ from $Z_{1}$ in (ref) to distinguish between the variables of interest or treatments, $X$, from the rest of the explanatory variables, $Z_{1}$, which play the role of controls. The functions $g$ and $h $ are unknown and $\varepsilon $ and $\eta $ are a vector and a scalar of potentially mutually dependent unobservables, respectively. We shall impose restrictions on the stochastic properties of these unobservables. The primary objective is to estimate functionals related to $g$ noting that $Y$ is only observed when $C$ is above some known threshold normalized to be zero. The non observability of $Y $ for specific values of $C$ induces the possibility of selection bias. We refer to (ref) as the outcome equation and (ref) as the selection equation.

The model is a nonparametric and nonseparable representation of the tobit type-3 model and is a variant of the Heckman (1979) selection model. It was initially examined in a fully parametric setting, imposing additivity and normality, and estimated by maximum likelihood (see Amemiya, 1978, 1979). Vella (1993) provided a two-step estimator based on estimating the generalized residual from the selection equation and including it as a control function in the outcome equation. Honor\'{e} et al. (1997), Chen (1997) and Lee and Vella (2006) provide semi-parametric estimators for this model.

The model can be extended in several directions. For example, the selection variable $C$ could be censored in a number of ways provided that there are some region(s) for which it is continuously observed. This allows for top, middle and/or bottom censoring. In addition, although we do not explicitly consider it here, our approach is applicable when the outcome variable $Y$ is also censored. For example:

equation*[equation* omitted — 66 chars of source]

It is also applicable with random censoring in the selection equation. That is, where (ref) is replaced with:

equation*[equation* omitted — 60 chars of source]

where $\varsigma $ is a random variable which is independent of ($ \varepsilon ,\eta $) conditional on $Z.\footnote{ The appropriate control function can be obtained via the Kaplan Meier approach conditional on $Z$. The proof is available from the authors upon request. We are grateful to the Associate Editor for this suggestion.}$ The model can also be extended to include $C$ in the outcome equation as an explanatory variable provided that there is an exclusion restriction in $ Z_{2}$ with respect to $X$. This extension, which corresponds to the triangular system of Imbens and Newey (2009) with censoring in the first-stage equation and selection in the second, is not considered here as it is not relevant for our empirical application. However, note that the Imbens and Newey (2009) estimator does not apply to this case as they do not consider the selection model.

We highlighted above that we follow a local approach to identification such as proposed by Heckman and Vytlacil (2005) who consider a binary treatment/selection rule and a separable selection equation. While our focus is also, in part, on local effects, our model differs with respect to the selection rule and the possible presence of nonseparability.

Identification of objects of interest

We account for selection bias via the use of an appropriately constructed control function. We establish the existence of such a function for this model and then define some objects of interest incorporated in ((ref))-( (ref)).

Let $\perp\!\!\!\perp$ denote stochastic independence. We begin with the following assumption:

assumption[Control Function] $\left( \varepsilon ,\eta \right) \perp\!\!\!\perp Z $ , $\eta$ is a continuously distributed random variable with strictly increasing CDF on the support of $\eta$, and $t \mapsto h(Z,t)$ is strictly increasing a.s.

This assumption allows for endogeneity between $(X,Z_1)$ and $\varepsilon $ in the selected population with $C>0$, since, in general, $\varepsilon $ and $\eta $ are dependent, i.e. $\varepsilon \not \perp\!\!\!\perp(X,Z_1)\mid C>0$. It allows for a non-monotonic relationship between $\varepsilon $ and $C$ because $\varepsilon $ and $\eta $ are allowed to be non-monotonically dependent. Under Assumption (ref), we can normalize the distribution of $\eta $ to be uniform on $[0,1]$ without loss of generality (Matzkin, 2003).\footnote{ Indeed if $t\mapsto h(z,t)$ is strictly increasing, and $\eta $ is continuously distributed with $\eta \sim F_{\eta }$, then $\tilde{h}(z, \tilde{\eta})=h(z,F_{\eta }(\tilde{\eta}))$ is such that $t\mapsto \tilde{h} (z,t)$ is strictly increasing and $\tilde{\eta}\sim U(0,1)$.}

The following lemma shows the existence of a control function for the selected population in this setting. That is, there is a function of the observable data such that the unobservable component is independent of the explanatory variables in the outcome equation for the selected population conditional on this function. Let $V:=F_{C|Z}(C\mid Z)$ where $F_{C|Z}(\cdot \mid z)$ denotes the CDF of $C$ conditional on $Z=z$.

lemma[Existence of Control Function] Under the model in ((ref))-((ref)) and Assumption (ref): \begin{equation*} \varepsilon \perp \!\!\!\perp Z\mid V,C>0. \end{equation*}

All proofs are provided in the Appendix. The intuition behind Lemma (ref) is based on three observations. First, $V=\eta$ when $C>0,$ so that conditioning on $V$ is identical to conditioning on $\eta$ in the selected population. Second, conditioning on $Z$ and $\eta $ makes the selection, i.e. $ C>0 $, deterministic. Therefore, the distribution of $\varepsilon $, conditional on $Z$ and $\eta $, does not depend on the condition that $C>0$. The final observation, namely our assumption that $(\varepsilon,\eta) \perp \!\!\!\perp Z,$ is sufficient to prove the Lemma.

We consider two classes of objects of interest. These are: (1) local effects conditional on the value of the control function, and (2) global effects based on integration over the control function.

Local effects

We consider local effects on $Y$ for given values of $X$ conditional on the control function $V$. Let $\mathcal{Z}$, $\mathcal{X}$, $\mathcal{Z}_1$, and $\mathcal{V}$ denote the marginal supports of $Z$, $X$, $Z_1$ and $V$ in the selected population, respectively. We define $\mathcal{XZ}_1\mathcal{V}$ as the joint support of $X$, $Z_1$ and $V$ in the selected population:

definition[Identification set] Define \begin{equation*} \mathcal{XZ}_1\mathcal{V}:=\left\{ (x,z_1,v)\in \mathcal{X}\times \mathcal{Z} _1 \times \mathcal{V}:h(z,v)>0,z\in \mathcal{Z}(x,z_1)\right\}, \end{equation*} where $\mathcal{Z}(x,z_1)=\{z\in \mathcal{Z}:(x,z_1)\subseteq z\},$ i.e. the set of values of $Z$ with the components $X=x$ and $Z_1=z_1$.

Depending on the values of $(X,Z_1,\eta)$, we can classify the units of observation into 3 groups: (1) always selected units when $h(z,t)>0$ for all $z\in \mathcal{Z}(x,z_1)$, (2) switchers when $h(z,t)>0$ for some $z\in \mathcal{Z}(x,z_1)$ and $h(z,t)\leq 0$ for some $z\in \mathcal{Z}(x,z_1)$, and (3) never selected units when $h(z,t)\leq 0$ for all $z\in \mathcal{Z} (x,z_1)$. The set $\mathcal{XZ}_1\mathcal{V}$ only includes always selected units and switchers, i.e. units with $(X,Z_1,V)$ such that they are observed for some values of $Z$. When $X=Z_2$, there are no switchers because the set $\mathcal{Z}(x)$ is a singleton. Otherwise the size of the set $\mathcal{XZ}_1\mathcal{V}$ increases with the support of the excluded variables and their strength in the selection equation.

We now define the local average structural function.

definition[LASF] The local average structural function (LASF) at $(x,z_1,v)$ is: \begin{equation*} \mu (x,z_1,v)=\mathbb{E}(g(x,z_1,\varepsilon ) \mid V=v). \end{equation*}

The LASF gives the expected value of the potential outcome $ g(x,z_1,\varepsilon ) $ obtained by fixing $(X,Z_1)$ at $(x,z_1)$ conditional on $V=v$ for the entire population. It is useful for measuring the effect of $X$ on the mean of $Y$. For example, the average treatment effect of changing $X$ from $x_{0}$ to $x_{1}$ at $Z_1 = z_1$ conditional on $V=v$ is:

equation*[equation* omitted — 53 chars of source]

The following result shows that $\mu (x,z_1,v)$ is identified for all $ (x,z_1,v)\in \mathcal{XZ}_1\mathcal{V}$.

theorem[Identification of LASF] Under the model ((ref))-((ref)), Assumption (ref) and $\mathbb{E} |Y| < \infty$, for $(x,z_1,v)\in \mathcal{XZ} _1\mathcal{V}$, \begin{equation} \mu (x,z_1,v)= \mathbb{E}(Y \mid X=x,Z_1 = z_1,V=v,C>0). \end{equation}

According to Theorem (ref), the LASF is identical to the expected value of $Y$ conditional on $(X,Z_1,V)=(x,z_1,v)$ in the selected population. The proof is based on Assumption (ref) that allows for the LASF to be conditional on $(Z,V)=(z,v)$. Since $(x,z_1,v)\in \mathcal{XZ}_1\mathcal{V}$, there is a $z\in \mathcal{Z}(x,z_1)$ such that $ h(z,v)>0$, the expected value of $g(X,Z_1,\varepsilon )$ conditional on $V=v$ for the entire population, i.e. the LASF, is the same as that expected value of $Y$ for the selected population conditional on $(X,Z_1,V) = (x,z_1,v)$.

We can consider the average derivative of $g(x,z_1,\varepsilon )$ with respect to $x$ conditional on the control function when $X$ is continuous and $x\mapsto g(x,z_1,\varepsilon)$ is differentiable almost surely.

definition[LADF] The local average derivative function (LADF) at $(x,z_1,v)$ is: \begin{equation} \delta (x,z_1,v)=\mathbb{E}[\partial _{x}g(x,z_1,\varepsilon )\mid V=v],\ \ \ \partial _{x}:=\partial /\partial x. \end{equation}

The LADF is the first-order derivative of the LASF with respect to $x$, provided that we can interchange differentiation and integration in ((ref)). This is made formal in the next corollary which shows that the LADF is identified for all $(x,z_1,v)\in \mathcal{XZ}_1 \mathcal{V}$.

corollary[Identification of LADF] Assume that for all $(x,z_1) \in \mathcal{X}\times\mathcal{Z}_1$, $x \mapsto g(x,z_1,\varepsilon )$ is continuously differentiable a.s., $\mathbb{E}[|g(x,z_1,\varepsilon )|]<\infty $, and $\mathbb{E}[|\partial _{x}g(x,$ $z_1,\varepsilon )|]<\infty $. Under the conditions of Theorem (ref), for $ (x,z_1,v)\in \mathcal{XZ}_1\mathcal{V}$, \begin{equation*} \delta (x,z_1,v)= \partial_x \mu (x,z_1,v)=\partial _{x}\mathbb{E}(Y \mid X=x,Z_1=z_1,V=v,C>0). \end{equation*}

The local effects extend in a straightforward manner to distributions and quantiles.

definition[LDSF and LQSF] The local distribution structural function (LDSF) at $(y,x,z_1,v)$ is: \begin{equation*} G(y,x,z_1,v)=\mathbb{E}[1\left\{ g(x,z_1,\varepsilon )\leq y\right\} \mid V=v]. \end{equation*} The local quantile structural function (LQSF) at $(\tau,x,z_1,v)$ is: \begin{equation*} q(\tau ,x,z_1,v):=\inf \{y\in \mathbb{R}:G(y,x,z_1,v)\geq \tau \}. \end{equation*}

The LDSF is the distribution function of the potential outcome $ g(x,z_1,\varepsilon )$ conditional on the value of the control function for the entire population. The LQSF is the left-inverse function of $y\mapsto G(y,x,z_1,v)$ and corresponds to the quantiles of $g(x,z_1,\varepsilon )$. The differences of the LQSF across levels of $x$ correspond to quantile treatment effects conditional on $V$ for the entire population. For example, the $\tau $-quantile treatment effect of changing $X$ from $x_{0}$ to $x_{1}$ is:

equation*[equation* omitted — 59 chars of source]

The identification of the LDSF follows by the same argument as the identification of the LASF, replacing $g(x,z_1,\varepsilon )$ (as in Definition (ref)) by $1\left\{ g(x,z_1,\varepsilon )\leq y\right\} $ and $Y$ (as in equation ((ref))) by $1\{Y\leq y\}$. Thus, under Assumption (ref), for $(x,z_1,v)\in \mathcal{XZ}_1\mathcal{V}$,

equation*[equation* omitted — 116 chars of source]

The LQSF is then identified by the left-inverse function of $y\mapsto F_{Y|X,Z_1,V,C>0}(y\mid x,z_1,v)$, the conditional quantile function $\tau \mapsto \mathbb{Q}_{Y}[\tau \mid X=x,Z_1=z_1,V=v,C>0]$, i.e., for $ (x,z_1,v)\in \mathcal{XZ}_1\mathcal{V}$,

equation*[equation* omitted — 80 chars of source]

We also consider the derivative of $q(\tau ,x,z_1,v)$ with respect to $x$ and call it the local quantile derivative function (LQDF). This object corresponds to the average derivative of $g(x,z_1,\varepsilon )$ with respect to $x$ at the quantile $q(\tau ,x,z_1,v)$ conditional on $V=v$ under suitable regularity conditions; see Hoderlein and Mammen (2011). Thus, for $ (\tau ,x,z_1,v)\in \lbrack 0,1]\times \mathcal{XZ}_1\mathcal{V},$

equation*[equation* omitted — 167 chars of source]

By an analogous argument to Corollary (ref) , the LQDF is identified at $(\tau ,x,z_1,v)\in \lbrack 0,1]\times \mathcal{ XZ}_1\mathcal{V}$ by:

equation*[equation* omitted — 101 chars of source]

provided that $x\mapsto \mathbb{Q}_{Y}[\tau \mid X=x,Z_1=z_1,V=v,C>0]$ is differentiable and other regularity conditions hold.

remark[Exclusion restrictions] The identification of local effects does not explicitly require exclusion restrictions in $Z_2$ with respect to $X$ although the size of the identification set $\mathcal{XZ}_1\mathcal{V}$ depends on such restrictions. For example, dropping $Z_1$, if $h(z,\eta )=z+\Phi ^{-1}(\eta )$ where $\Phi $ is the standard normal distribution and $X=Z$, then $\mathcal{XV} =\{(x,v)\in \mathcal{X}\times \mathcal{V}]:x> -\Phi ^{-1}(v)\}\subset \mathcal{X}\times \mathcal{V}$; whereas if $h(z,\eta )=x+\tilde z+\Phi ^{-1}(\eta )$ for $Z=(X,\tilde Z)$, then $\mathcal{XV}=\{(x,v)\in \mathcal{X} \times \mathcal{V}:x> -\Phi ^{-1}(v)-\tilde z,\tilde z \in \mathcal{Z}(x)\}$ , such that $\mathcal{XV}=\mathcal{X}\times \mathcal{V}$ if $\tilde Z$ is independent of $X$ and supported in $\mathbb{R}$.

Global effects

We expand our set of estimands by examining the global counterparts of the local effects obtained by integration over the control function and the observed controls in the selected population. A typical global effect at $ x\in \mathcal{X}$ is:

equation[equation omitted — 89 chars of source]

where $\theta (x,z_1,v)$ can be any of the local objects defined above and $ F_{Z_1,V \mid C>0} $ is the joint distribution of $(Z_1,V)$ in the selected population. The identification of $\theta_S (x)$ requires identification of $ \theta (x,z_1,v)$ over $\mathcal{Z}_1\mathcal{V}$, the support of $(Z_1,V)$ in the selected population.

For example, the average structural function (ASF):

equation*[equation* omitted — 72 chars of source]

gives the average of the potential outcome $g(x,Z_1,\varepsilon )$ in the selected population. By the law of iterated expectations, this is a special case of the global effect (ref) with $\theta (x,z_1,v)=\mu (x,z_1,v)$, the LASF. The average treatment effect of changing $X$ from $x_{0}$ to $ x_{1} $ in the selected population is:

equation[equation omitted — 70 chars of source]

Similarly, one can consider the distribution structural function (DSF) in the selected population as in Newey (2007), i.e:

equation*[equation* omitted — 82 chars of source]

which gives the distribution of the potential outcome $g(x,Z_1,\varepsilon )$ at $y$ in the selected population. This is also a special case of the global effect (ref) with $\theta (x,z_1,v)=G(y,x,z_1,v)$. We can then construct the quantile structural function (QSF) in the selected population as the left-inverse of $y\mapsto G_{S}(y,x)$, that is:

equation*[equation* omitted — 79 chars of source]

The QSF gives the quantiles of $g(x,Z_1,\varepsilon )$. Unlike $G_{S}(y,x)$, $q_{S}(\tau ,x)$ cannot be obtained by integration of the corresponding local effect, $q(\tau ,x,z_1,v)$, because we cannot interchange quantiles and expectations. The $\tau $-quantile treatment effect of changing $X$ from $x_{0}$ to $x_{1}$ in the selected population is:

equation*[equation* omitted — 55 chars of source]

The global counterparts of the LADF and LQSF are obtained by taking the derivatives of $\mu _{S}(x)$ and $q_{S}(\tau ,x)$ with respect to $x$.

As in Newey (2007), the identification of the global effects in the selected population requires a condition on the support of the control function. Let $ \mathcal{Z}_1\mathcal{V}(x)$ denote the support of $(Z_1,V)$ conditional on $ X=x$, i.e. $\mathcal{Z}_1\mathcal{V}(x) := \{(z_1,v) \in \mathcal{Z} _1 \times \mathcal{V}: (x,z_1,v) \in \mathcal{XZ}_1\mathcal{V} \}$.

assumption[Common Support] $\mathcal{Z}_1\mathcal{V}(x) = \mathcal{Z}_1\mathcal{V} $.

The main implication of common support is the identification of $\theta(x)$ from the identification of $\theta(x,z_1,v)$ in $(z_1,v) \in \mathcal{Z}_1 \mathcal{V}(x) = \mathcal{Z}_1\mathcal{V}$. Assumption (ref) is only plausible under exclusion restrictions on $Z_2$ with respect to $X$; see the example in Remark (ref) . We now establish the identification of the typical global effect (ref) .

theorem[Identification of Global Effects] If $\theta(x,z_1,v)$ is identified for all $(x,z_1,v) \in \mathcal{XZ}_1 \mathcal{V}$, then $\theta_S(x)$ is identified for all $x \in \mathcal{X}$ that satisfy Assumption (ref).

We can now apply this result to show the identification of global effects in the selected population, because under Assumption (ref), the local effects are identified over $\mathcal{XZ}_1\mathcal{V}$, which is the support of $(X,Z_1,V)$ in the selected population.

remark[Global Effects in the Entire Population] The effects in the selected population generally differ from the effects in the entire population, except under the additional support condition: \begin{equation} \mathcal{Z}_1\mathcal{V}=\mathcal{Z}_1\mathcal{V}^0, \end{equation} where $\mathcal{Z}_1\mathcal{V}^0$ is the support of $(Z_1,V)$ in the entire population. This condition requires an excluded variable in $Z_2$ with sufficient variation to make $h(Z,\eta )>0$ for any $(z_1,\eta)$.

Global effects for the treated and average derivatives

Assumption (ref) might be too restrictive for empirical applications where an excluded variable with large support is not available. Without this assumption the global effects are not point identified. The alternative generic global effect:

equation[equation omitted — 119 chars of source]

is point identified under weaker support conditions than ((ref)). Examples of (ref) include the ASF conditional on $X=x_{0}$ in the selected population:

equation*[equation* omitted — 91 chars of source]

which is a special case of (ref) with $\theta (x,z_{1},v)=\mu (x,z_{1},v)$. This ASF measures the mean of the potential outcome $ g(x,Z_{1},\varepsilon )$ for the selected individuals with $X=x_{0}$, and is useful to construct the average treatment effect on the treated of changing $ X$ from $x_{0}$ to $x_{1}$:

equation[equation omitted — 91 chars of source]

The object in (ref) is identified in the selected population under the following support condition:

assumption[Weak Common Support] $\mathcal{Z}_1\mathcal{V}(x) \supseteq \mathcal{Z}_1 \mathcal{V}(x_0)$.

Assumption (ref) is weaker than Assumption (ref) because $\mathcal{Z}_1\mathcal{V}(x_{0}) \subseteq \mathcal{Z}_1\mathcal{V}$ . In particular, if the selection equation (ref) is increasing in $X$ and $\mathcal{X}$ is bounded from below, then Assumption (ref) is satisfied by setting $x_0$ lower than $x$.

We define the $\tau$-quantile treatment on the treated as:

equation*[equation* omitted — 70 chars of source]

where $q_{S}(\tau, x \mid x_0)$ is the left-inverse of the DSF conditional on $X=x_0$ in the selected population,

equation*[equation* omitted — 98 chars of source]

which is a special case of the effect (ref) with $\theta(x,z_1,v) = G(y,x,z_1,v)$.

We now establish the identification of the typical global effect (ref).

theorem[Identification of Global effects for the Treated] If $\theta (x,z_1,v)$ is identified for all $(x,z_1,v)\in \mathcal{XZ}_1 \mathcal{V}$, then $\theta _{S}(x\mid x_{0})$ is identified for all $x\in \mathcal{X}$ that satisfy Assumption (ref).

We can define global objects in the selected population that are identified without a common support assumption when $X$ is continuous and $x\mapsto g(x,\cdot )$ is differentiable. One example is the average derivative conditional on $X=x$ in the selected population:

equation*[equation* omitted — 73 chars of source]

which is a special case of the effect (ref) with $\theta (x,z_1,v)=\delta (x,v)$ and $x_{0}=x$. This object is point identified in the selected population under Assumption (ref) because the integral is over $\mathcal{Z}_1\mathcal{V}(x)$, the support of $(Z_1,V)$ conditional on $X=x$ in the selected population. Another example is the average derivative in the selected population:

equation*[equation* omitted — 66 chars of source]

which is point identified under Assumption (ref) because the integral is over $\mathcal{XZ}_1\mathcal{V}$, the support of $(X,Z_1,V)$ in the selected population. This is a special case of the generic global effect:

equation[equation omitted — 93 chars of source]

Bounds for the global effects

Another approach to relaxing the support condition required by Assumption (ref) is to construct bounds on the structural functions. As bounds are easily obtained for the LDSF, which is restricted between 0 and 1, we only discuss the bounds for the DSF noting that those for the ASF are similar.

We start by separating the identified part of the DSF from the unidentified part:

equation*[equation* omitted — 199 chars of source]

where $\overline{\mathcal{Z}_{1}\mathcal{V}}(x)=\mathcal{Z}_{1}\mathcal{V} \setminus \mathcal{Z}_{1}\mathcal{V}(x).$ Then, following Imbens and Newey (2009), we form the bounds as:

multline*[multline* omitted — 276 chars of source]

since $0\leq G(y,x,z_{1},v)\leq 1$. This implies that the width of the bounds equals $P[(Z_{1},V)\in \overline{\mathcal{Z}_{1}\mathcal{V}}(x)\mid C>0]$. An alternative approach, suggested by Newey (2007) for binary selection, employs the propensity score $P=P(C>0\mid Z)$ as a control function. In Appendix (ref), we show that our bounds are tighter than the bounds obtained from the propensity score. This highlights another benefit of using a censored, rather than a binary selection equation, when possible.

It is also possible to calculate bounds for the DSF for the entire population since:

multline*[multline* omitted — 262 chars of source]

and hence:

multline*[multline* omitted — 265 chars of source]

The ability to construct bounds for the entire population reflects that the local effects are not conditional on the selected population.

Estimation and inference

The effects of interest are all identified by functionals of the distribution of the observed variables and the control function in the selected population. The control function is the distribution of the censoring variable $C$ conditional on all explanatory variables $Z$. We propose a multistep semiparametric method based on least squares, distribution and quantile regressions to estimate the effects. The reduced form specifications used in each step can be motivated by parametric restrictions on the model (ref)--(ref). We refer to Chernozhukov et al. (2020) for examples of such restrictions. The semiparametric structure produces precise estimators and dispenses with the support assumptions for the identification of the global effects (Newey and Stouli, 2018). In practice, we recommend to supplement the estimators of the global effects with estimators of their bounds that do not rely on the semiparametric structure to avoid the support assumptions. We provide an example of this robustness check in Section (ref).

Throughout this section, we assume that we have a random sample of size $n$, $\{(Y_{i}\times 1(C_{i}>0),C_{i},Z_{i})\}_{i=1}^{n}$, of the random variables $(Y\times 1(C>0),C,Z)$, where $Y\times 1(C>0)$ indicates that $Y$ is observed only when $C>0$.

Step 1: Estimation of the control function

We estimate the control function using logistic distribution regression (Foresi and Peracchi, 1995, and Chernozhukov et al., 2013). For every observation in the selected sample, we set:

equation*[equation* omitted — 129 chars of source]

where, for $c\in \mathcal{C}_{n},$ the empirical support of $C$,

equation*[equation* omitted — 213 chars of source]

$\Lambda $ is the logistic distribution and $r(z)$ is a $d_{r}$-dimensional vector of transformations of $z$ with good approximating properties such as polynomials, B-splines and interactions.

Step 2: Estimation of local objects

We can estimate the local average, distribution and quantile structural functions using flexibly parametrized least squares, distribution and quantile regressions, where we replace the control function by its estimator from the previous step. For technical reasons explained in Section (ref), our estimation method is based on a trimmed sample with respect to the censoring variable $C $. Therefore, we introduce the following trimming indicator among the selected sample:

equation*[equation* omitted — 49 chars of source]

where $\overline{\mathcal{C}}=(0,\overline{c}]$ for some $0<\overline{c} <\infty $, such that $P(T=1)>0$.

The estimator of the LASF is $\widehat{\mu }(x,z_{1},v)=w(x,z_{1},v)^{ \mathrm{T}}\widehat{\beta },$ where $w(x,z_{1},v)$ is a $d_{w}$-dimensional vector of transformations of $(x,z_{1},v)$ with good approximating properties, and $\widehat{\beta }$ is the ordinary least squares estimator: \footnote{ An alternative approach is to follow Jun (2009) and Masten and Torgovitsky (2016). These papers acknowledge that with an index restriction the parameters of interest can be estimated in the presence of a control function by estimation over subsamples for which the control function has a similar value. While each of these papers considers a random coefficients model with endogeneity their approach is applicable here.}

equation*[equation* omitted — 210 chars of source]

provided that $\sum_{i=1}^{n}\widehat{W}_{i}\widehat{W}_{i}^{\mathrm{T} }T_{i} $ is invertible. The estimator of the LDSF is $\widehat{G} (y,x,z_{1},v)=\Lambda (w(x,z_{1},v)^{\mathrm{T}}\widehat{\beta }(y))$, where $\widehat{\beta }(y)$ is the logistic distribution regression estimator:

equation*[equation* omitted — 232 chars of source]

Similarly, the estimator of the LQSF is $\widehat{q}(\tau ,x,z_{1},v)=w(x,z_{1},v)^{\mathrm{T}}\widehat{\beta }(\tau )$, where $ \widehat{\beta }(\tau )$ is the Koenker and Bassett (1978) quantile regression estimator:

equation*[equation* omitted — 150 chars of source]

Estimators of the local derivatives are obtained by taking derivatives of the estimators of the local structural functions. Thus, the estimator of the LADF is:

equation*[equation* omitted — 102 chars of source]

and the estimator of the LQDF is:

equation*[equation* omitted — 117 chars of source]

Step 3: Estimation of global effects

We obtain estimators of the generic global effects by approximating the integrals over the control function by averages of the estimated local effects evaluated at the estimated control function. The estimator of the effect (ref) is:

equation*[equation* omitted — 126 chars of source]

This yields the estimators of the ASF for $\widehat{\theta }(x,z_1,v)= \widehat{\mu }(x,z_1,v)$ and DSF at $y$ for $\widehat{\theta }(x,z_1,v)= \widehat{G}(y,x,z_1,v)$. The estimator of the QSF is then obtained by inversion of the estimator of the DSF.\footnote{ We use the generalized inverse

equation*[equation* omitted — 147 chars of source]

which does not require that the estimator of the DSF $y\mapsto \widehat{G} _{S}(y,x)$ be monotone.} We form an estimator of the effect (ref) as:

equation*[equation* omitted — 160 chars of source]

for $K_{i}(x_{0})=1(X_{i}=x_{0})$ when $X$ is discrete or $ K_{i}(x_{0})=k_{h}(X_{i}-x_{0})$ when $X$ is continuous, where $ k_{h}(u)=k(u/h)/h$, $k$ is a kernel, and $h$ is a bandwidth such as $ h\rightarrow 0$ as $n\rightarrow \infty$. Finally, the estimator of the effect (ref) is:

equation*[equation* omitted — 127 chars of source]

Estimation of bounds on global effects

We provide estimators of the bounds of $G(y,x)$. Estimators of the bounds of $G_S(y,x)$ can be constructed similarly. We start by rewriting the width of the bounds as

equation*[equation* omitted — 212 chars of source]

where $v_u(x,z_1) = \sup\{v : (x,z_1,v)\in \mathcal{XZ}_1\mathcal{V}\}$ and $ v_{\ell}(x,z_1) = \inf\{v : (x,z_1,v)\in \mathcal{XZ}_1\mathcal{V}\}$. By the monotonicity in Assumption (ref), $v_u(x,z_1) = 1$ and $ v_{\ell}(x,z_1) = \inf\{ F_{Z}(0 \mid Z=z) : z \in \mathcal{Z} _{2}(x,z_{1})\} $, where $\mathcal{Z}_{2}(x,z_{1})$ is the support of $Z_{2} \mid X=x,Z_{1}=z_{1}$. We can then estimate $v_{\ell}(x,z_1)$ by

equation*[equation* omitted — 155 chars of source]

Finally, we can form estimators of the lower and upper bounds of $G(y,x)$ as:

equation*[equation* omitted — 129 chars of source]

and

equation*[equation* omitted — 211 chars of source]

Inference

We use weighted bootstrap to make inference on all objects of interest (Praestgaard and Wellner, 1993; Hahn, 1995). This method obtains the bootstrap version of the estimator of interest by repeating all estimation steps including random draws from a distribution as sampling weights. The weights should be positive and come from a distribution with unit mean and variance such as the standard exponential. The weighted bootstrap has some theoretical and practical advantages over the empirical bootstrap. Thus, it is appealing that the consistency can be proven following the strategy set forth by Ma and Kosorok (2005) and the smoothness induced by the weights helps dealing with discrete covariates with small cell sizes. The implementation of the bootstrap for the local and global effects is summarized in the following algorithm:

algorithm[algorithm omitted — 2,630 chars of source]

Asymptotic theory

We derive large sample theory for the local and global effects. We focus on the average effects for the sake of brevity. The theory for distribution and quantile effects can be derived using similar arguments; see, for example, Chernozhukov et al. (2015) and Chernozhukov et al. (2020).\footnote{ Chernozhukov et al. (2015) and Chernozhukov et al. (2020) considered models with endogeneity and censoring, and models with endogeneity, respectively. They derived theory for two-step estimators of these models where the first step is estimated by distribution regression and the second step by quantile or distribution regression. We derive theory for two-step estimators of sample selection models where the first step is estimated by distribution regression and the second step by least squares.} Throughout the analysis, we assume that the flexible specifications used in all the steps are correct and treat their dimensions as fixed, so that all model parameters are estimable at a $\sqrt{n} $-rate and the estimation of the global effects does not require any support assumptions.

In what follows we use the following notation. We let the random vector $ A=(Y\times 1(C>0),C,Z,V)$ live on some probability space $(\Omega _{0}, \mathcal{F}_{0},P)$. Thus, the probability measure $P$ determines the law of $A$ or any of its elements. We also let $A_{1},...,A_{n}$ be i.i.d. copies of $A$ and live on the complete probability space $(\Omega ,\mathcal{F},\mathbb{P} ) $, which contains the infinite product of $(\Omega _{0},\mathcal{F}_{0},P)$ . Moreover, this probability space can be suitably enriched to also carry the random weights that appear in the weighted bootstrap. The distinction between the two laws $P$ and $\mathbb{P} $ is helpful to simplify the notation in the proofs and in the analysis. Calligraphic letters such as $\mathcal{Y}$ and $\mathcal{X}$ denote the supports of $Y\times 1(C>0)$ and $X$; and $ \mathcal{YX}$ denotes the joint support of $(Y,X)$. Unless explicitly mentioned, all functions appearing in the statements are assumed to be measurable.

We now formally state the assumptions. The first assumption is about sampling and the bootstrap weights.

condition[Sampling and Bootstrap Weights] (a) Sampling: the data $\{Y_{i}\times 1(C_{i}>0),C_{i},Z_{i}\}_{i=1}^{n}$ are a sample of size $n$ of independent and identically distributed observations from the random vector $(Y\times 1(C>0),C,Z).$ (b) Bootstrap weights: $(\omega_{1},...,\omega_{n})$ are i.i.d. draws from a random variable $\omega \geq 0$, with ${\mathbb{E}} _{P}[\omega]=1$, $\mathrm{Var}_{P}[\omega]=1,$ and ${\mathbb{E}} _{P}|\omega|^{2+\delta }<\infty $ for some $\delta >0 $; live on the probability space $(\Omega ,\mathcal{F},\mathbb{P} )$; and are independent of the data $\{Y_{i}\times 1(C_{i}>0),C_{i},Z_{i}\}_{i=1}^{n}$ for all $n$.

The second assumption concerns the first stage in which we estimate the control function:

equation*[equation* omitted — 53 chars of source]

We assume a logistic distribution regression model for the conditional distribution of $C$ in the trimmed support, $\overline{\mathcal{C}}$, that excludes censored and extreme values of $C$. The purpose of the upper trimming is to avoid the upper tail in the modeling and estimation of the control function, and to make the eigenvalue assumption in Condition (ref)(b) more plausible. We consider a fixed trimming rule, which greatly simplifies the derivation of the asymptotic properties.\footnote{ We view the fixed trimming as a convenient theoretical device that is not restrictive in practical applications. Indeed, our estimators perform well without trimming in numerical simulations. Alternative random, data driven rules are possible at the cost of more complicated proofs and technical conditions on the upper tail of $F_{C}(c\mid z)$.} Throughout this section, we use bars to denote trimmed supports with respect to $C$, e.g., $\overline{ \mathcal{C}\mathcal{Z}}=\{(c,z)\in \mathcal{C}\mathcal{Z}:c\in \overline{ \mathcal{C}}\}$, and $\overline{\mathcal{V}}=\{\vartheta _{0}(c,z):(c,z)\in \overline{\mathcal{C}\mathcal{Z}}\}$.

condition[First Stage] (a) Trimming: we consider the trimming rule as defined by the indicator $T=\mathbf{1}(C \in \overline{ \mathcal{C}})$. (b) Model: the distribution of $C$ conditional on $Z$ follows the distribution regression model in the trimmed support $\overline{\mathcal{C}}$ , i.e., \begin{equation*} F_C(c \mid Z) = F_C(c \mid R) = \Lambda(R^{\mathrm{T}}\pi_0(c)), \ \ R = r(Z), \end{equation*} for all $c \in \overline{\mathcal{C}}$, where $\Lambda$ is the logit link function; the coefficients $c \mapsto \pi_0(c)$ are three times continuously differentiable with uniformly bounded derivatives; $\overline{\mathcal{R}}$ is compact; and the minimum eigenvalue of ${\mathbb{E}}_P \left[\Lambda(R^{ \mathrm{T}}\pi_0(c))[1- \Lambda(R^{\mathrm{T}}\pi_0(c))] RR^{\mathrm{T}} \right] $ is bounded away from zero uniformly over $c \in \overline{\mathcal{ C}}$.

For $c\in \overline{\mathcal{C}}$, let:

equation*[equation* omitted — 214 chars of source]

where either $\omega _{i}=1$ for the unweighted sample, to obtain the estimator, or $\omega _{i}$ are the bootstrap weights for obtaining bootstrap draws of the estimator. Then set:

equation*[equation* omitted — 151 chars of source]

if $(c,r)\in \overline{\mathcal{C}\mathcal{R}},$ and $\vartheta _{0}(c,r)= \widehat{\vartheta }^{b}(c,r)=0$ otherwise.

Theorem 4 of Chernozhukov et al. (2015) established the asymptotic properties of the DR estimator of the control function. We repeat the result here as a lemma for completeness and to introduce notation that will be used in the results below. Let $\|f\|_{T,\infty} := \sup_{a \in \mathcal{A}} |T(c) f(a)|$ for any function $f : \mathcal{A} \mapsto \mathbb{R}$, $ \ell^{\infty}(\mathcal{A})$ be the set of bounded functions on $\mathcal{A}$ equipped with the norm $\| \cdot \|_{T,\infty}$, and $\lambda = \Lambda(1-\Lambda)$, the density of the logistic distribution.

lemma[First Stage] Suppose that Conditions (ref) and (ref) hold. Then, (1) \begin{eqnarray*} \sqrt{n}(\widehat{\vartheta }^{b}(c,r)-\vartheta _{0}(c,r)) &=&\frac{1}{ \sqrt{n}}\sum_{i=1}^{n}e_{i}\ell (A_{i},c,r)+o_{\mathbb{P} }(1)\rightsquigarrow \Delta ^{b}(c,r) in \ell ^{\infty }(\overline{\mathcal{C}\mathcal{R}} ), \\ \ell (A,c,r):= &&\lambda (r^{\mathrm{T}}\pi _{0}(c))[1\{C\leq c\}-\Lambda (R^{\mathrm{T}}\pi _{0}(c))]\times \\ &&\times r^{\mathrm{T}}{\mathbb{E}}_{P}\left\{ \Lambda (R^{\mathrm{T}}\pi _{0}(c))[1-\Lambda (R^{\mathrm{T}}\pi _{0}(c))]RR^{\mathrm{T}}\right\} ^{-1}R, \\ {\mathbb{E}}_{P}[\ell (A,c,r)] &=&0,{\mathbb{E}}_{P}[T\ell (A,C,R)^{2}]<\infty , \end{eqnarray*} where $(c,r)\mapsto \Delta ^{b}(c,r)$ is a Gaussian process with uniformly continuous sample paths and covariance function given by ${\mathbb{E}} _{P}[\ell (A,c,r)\ell (A,\tilde{c},\tilde{r})^{\mathrm{T}}]$. (2) There exists $\widetilde{\vartheta }^{b}:\overline{\mathcal{CR}}\mapsto \lbrack 0,1]$ that obeys the same first-order representation uniformly over $ \overline{\mathcal{CR}}$, is close to $\widehat{\vartheta }^{b}$ in the sense that $\Vert \widetilde{\vartheta }^{b}-\widehat{\vartheta }^{b}\Vert _{T,\infty }=o_{\mathbb{P} }(1/\sqrt{n})$ and, with probability approaching one, belongs to a bounded function class $\Upsilon $ such that the covering entropy satisfies:\footnote{ See Appendix (ref) for a definition of the covering entropy.} \begin{equation*} \log N(\epsilon ,\Upsilon ,\Vert \cdot \Vert _{T,\infty })\lesssim \epsilon ^{-1/2},\ \ 0<\epsilon <1. \end{equation*}

The next assumptions are about the second stage. We assume a flexible linear model for the conditional distribution of $Y$ given $(X,V)$ in the trimmed support $C \in \overline{\mathcal{C}}$, impose compactness conditions, and provide sufficient conditions for the identification of the parameters. Compactness is imposed over the trimmed support and can be relaxed at the cost of more complicated and cumbersome proofs.

condition[Second Stage] (a) Model: the expectation of $Y$ conditional on $(X,Z_1,$ $V)$ in the trimmed support $C\in \overline{\mathcal{C}}$ is: \begin{equation*} \mathbb{E}(Y\mid X,Z_1,V,C\in \overline{\mathcal{C}})=W^{\mathrm{T}}\beta _{0},\ \ V=F_{C|Z}(C\mid Z),\ \ W=w(X,Z_1,V). \end{equation*} (b) Compactness and moments: the set $\overline{\mathcal{W}}$ is compact; the derivative vector $\partial _{v}w(x,z_1,v)$ exists and its components are uniformly continuous in $v\in \overline{\mathcal{V}}$, uniformly in $ (x,z_1)\in \overline{\mathcal{XZ}}_1$, and are bounded in absolute value by a constant, uniformly in $(x,z_1,v)\in \overline{\mathcal{XZ}_1\mathcal{V}}$ ; $\mathbb{E}(Y^{2}\mid C\in \overline{\mathcal{C}})<\infty $; and $\beta _{0}\in \mathcal{B}$, where $\mathcal{B}$ is a compact subset of $\mathbb{R} ^{d_{w}}$. (c) Identification and nondegeneracy: the matrix $J:={\mathbb{E}} _{P}[WW^{\mathrm{T}}\ T]$ is of full rank; and the matrix $\Omega :=\mathrm{ Var}_{P}[f_{1}(A)+f_{2}(A)]$ is finite and is of full rank, where: \begin{equation*} f_{1}(A):=\{W^{\mathrm{T}}\beta _{0}-Y\}WT, \end{equation*} and, for $\dot{W}=\partial _{v}w(X,Z_1,v)|_{v=V}$, \begin{equation*} f_{2}(A):={\mathbb{E}}_{P}[\{[W^{\mathrm{T}}\beta _{0}-Y]\dot{W}+W^{\mathrm{T }}\beta _{0}W\}T\ell (a,C,R)]\big|_{a=A}. \end{equation*}

Let

equation*[equation* omitted — 251 chars of source]

where $\widehat \vartheta$ is the estimator of the control function in the unweighted sample; and

equation*[equation* omitted — 277 chars of source]

where $\widehat \vartheta^b$ is the estimator of the control function in the weighted sample. The following lemma establishes a central limit theorem and a central limit theorem for the bootstrap for the estimator of the coefficients in the second stage. Let $\rightsquigarrow_{\mathbb{P}}$ denote the bootstrap consistency, i.e. weak convergence conditional on the data in probability as defined in Appendix (ref).

lemma[CLT and Bootstrap FCLT for $\protect\widehat{\protect\beta }$] Under Conditions (ref)--(ref), in $ \mathbb{R}^{d_w}$, \begin{equation*} \sqrt{n}(\widehat\beta - \beta_0) \rightsquigarrow J^{-1} G, \ \ and \ \ \sqrt{n}(\widehat\beta^{b} - \widehat \beta) \rightsquigarrow_{\mathbb{P}} J^{-1} G, \end{equation*} where $G \sim N(0,\Omega)$ and $J$ and $\Omega$ are defined in Assumption (ref)(c).

The properties of the estimator of the LASF, $\widehat \mu(x,z_1,v) = w(x,z_1,v)^{\mathrm{T}} \widehat \beta$, and its bootstrap version, $ \widehat \mu^b(x,z_1,v) = w(x,z_1,v)^{\mathrm{T}} \widehat \beta^b$, constitute a corollary of Lemma (ref).

corollary[FCLT and Bootstrap FCLT for LASF] Under Assumptions (ref)--(ref), in $\ell (\overline{\mathcal{XZ}_1\mathcal{V}})$, \begin{equation*} \sqrt{n}(\widehat{\mu }(x,z_1,v)-\mu (x,z_1,v))\rightsquigarrow Z(x,z_1,v) and \sqrt{n}(\widehat{\mu }^{b}(x,z_1,v)-\widehat{\mu } (x,z_1,v))\rightsquigarrow _{\mathbb{P} }Z(x,z_1,v), \end{equation*} where $(x,z_1,v)\mapsto Z(x,z_1,v):=w(x,z_1,v)^{\mathrm{T}}J^{-1}G$ is a zero-mean Gaussian process with covariance function: \begin{equation*} \mathrm{Cov} _{P}[Z(x_{0},z_{10},v_{0}),Z(x_{1},z_{11},v_{1})]=w(x_{0},z_{10},v_{0})^{ \mathrm{T}}J^{-1}\Omega J^{-1}w(x_{1},z_{11},v_{1}). \end{equation*}

To obtain the properties of the estimator of the ASFs, we define $ W_{x}:=w(x,Z_1,V)$, $\widehat{W}_{x}:=w(x, Z_1,\widehat{V})$, and $\widehat{W }_{x}^{b}:=w(x,Z_1,\widehat{V}^{b})$. The estimator and its bootstrap draw of the ASF in the trimmed support, $\mu_S(x)={\ \mathbb{E}}_{P}\{\beta_{0}^{ \mathrm{T}} W_{x} \mid T = 1\}$, are $\widehat{\mu}_S(x)=\sum_{i=1}^{n} \widehat{\beta}^{\mathrm{T}} \widehat{W}_{xi} T_{i}/n_T$, and $\widehat{\mu} _S^{b}(x)=\sum_{i=1}^{n}e_{i} \widehat{\beta}^{b T }\widehat{W}_{xi}^{b} T_{i}/n_T^b$, where $n_T = \sum_{i=1}^n T_i$ and $n^b_T = \sum_{i=1}^n e_i T_i$. The estimator and its bootstrap draw of the ASF on the treated in the trimmed support, $\mu_S(x \mid x_0)={\ \mathbb{E}}_{P}\{\beta_{0}^{\mathrm{T} } W_{x} \mid T = 1, X = x_0\}$, are $\widehat{\mu}_S(x \mid x_0)=\sum_{i=1}^{n}\widehat{\beta}^{\mathrm{T}} \widehat{W}_{xi} K_i(x_0) T_{i}/n_T(x_0)$, and $\widehat{\mu}_S^{b}(x) = \sum_{i=1}^{n}e_{i} \widehat{ \beta}^{b T}\widehat{W}_{xi}^{b}$ $K_i(x_0)T_{i} / n_T^b(x_0)$, where $ n_T(x_0) = \sum_{i=1}^n K_i(x_0) T_i $ and $n^b_T(x_0) = \sum_{i=1}^n e_i K_i(x_0) T_i$. Let $p_T := P(T=1)$ and $p_T(x) := P(T=1, X=x)$. The next result gives the large sample theory for these estimators. The theory for the ASF on the treated is derived for $X$ discrete, which is the relevant case in our empirical application.

theorem[FCLT and Bootstrap FCLT for ASF] Under Assumptions (ref)--(ref), in $ \ell (\overline{\mathcal{X}})$, \begin{equation*} \sqrt{np_{T}}(\widehat{\mu }_{S}(x)-\mu _{S}(x))\rightsquigarrow Z(x) and \sqrt{np_{T}}(\widehat{\mu }_{S}^{b}(x)-\widehat{\mu } _{S}(x))\rightsquigarrow _{\mathbb{P} }Z(x), \end{equation*} where $x\mapsto Z(x)$ is a zero-mean Gaussian process with covariance function: \begin{equation*} \mathrm{Cov}_{P}[Z(x_{0}),Z(x_{1})]=\mathrm{Cov}_{P}[W_{x_{0}}^{\mathrm{T} }\beta _{0}+\sigma _{x_{0}}(A),W_{x_{1}}^{\mathrm{T}}\beta _{0}(v)+\sigma _{x_{1}}(A)\mid T=1], \end{equation*} with: \begin{equation*} \sigma _{x}(A)={\mathbb{E}}_{P}\{W_{x}^{\mathrm{T}}T \}J^{-1}[f_{1}(A)+f_{2}(A)]+{\mathbb{E}}_{P}\{\dot{W}_{x}^{\mathrm{T}}\beta _{0}T\ell (a,C,R)\}\big|_{a=A}. \end{equation*} Also, if $p_{T}(x_{0})>0$, in $\ell (\overline{\mathcal{X}})$, \begin{equation*} \begin{split} \sqrt{np_{T}(x_{0})}(\widehat{\mu }_{S}(x& \mid x_{0})-\mu _{S}(x\mid x_{0}))\rightsquigarrow Z(x\mid x_{0}) and \\ \sqrt{np_{T}(x_{0})}(\widehat{\mu }_{S}^{b}(x& \mid x_{0})-\widehat{\mu } _{S}(x\mid x_{0}))\rightsquigarrow _{\mathbb{P} }Z(x\mid x_{0}), \end{split} \end{equation*} where $x\mapsto Z(x\mid x_{0})$ is a zero-mean Gaussian process with covariance function: \begin{equation*} \mathrm{Cov}_{P}[Z(x_1\mid x_{0}),Z(x_2\mid x_{0})]=\mathrm{Cov} _{P}[W_{x_1}^{\mathrm{T}}\beta _{0}+\sigma_{x_1}(A),W_{x_2}^{\mathrm{T} }\beta _{0}+\sigma _{x_2}(A)\mid T=1,X=x_{0}], \end{equation*} \begin{comment} with \begin{equation*} \psi_{x}(A)={\mathbb{E}}_{P}\{ W^{\mathrm{T}}_{x}T\}J^{ -1}[f_1(A)+f_2(A)]+ { \mathbb{E}}_{P}\{\dot{W}_{x}^{\mathrm{T}}\beta_{0}T\ell(a,X,R)\}\big|_{a=A}. \end{equation*} \end{comment}

Theorem (ref) can be used to construct confidence bands for the ASFs, $x \mapsto \mu_S(x)$ and $x \mapsto \mu_S(x \mid x_0)$, over regions of values of $x$ via Kolmogorov-Smirnov type statistics and weighted bootstrap, and to construct confidence intervals for average treatment effects, $\mu(x_1) - \mu(x_0)$ and $\mu(x_1 \mid x_0) - \mu(x_0 \mid x_0)$, via t-statistics and weighted bootstrap.

Application: United Kingdom wage regressions

We investigate the impact of human capital on the wage level of female workers in the United Kingdom (UK). We use data from the Family Expenditure Survey (FES) for the years 1978 to 1999. Blundell et al. (2003) study male wage growth and Blundell et al. (2007) examine wage inequality for both males and females using the same data source. We employ the same data selection rules and refer the reader to these earlier papers for details. The FES is a repeated cross section of households and contains detailed information on the number of weekly hours worked and the hourly wage of the individual. We restrict the data to those who report an education level and only include working females who report working weekly hours of 70 or less and an hourly wage of at least 0.01 pounds. This reduces the total number of observations from 96,402 to 94,985, corresponding to over 4,100 individuals per year with an average of 2,600 working.

The outcome variable, $Y$, is the log-hourly wage defined as the nominal weekly earnings divided by the number of hours worked and deflated by the quarterly UK retail price index. The number of hours worked, $C$, is defined as the usual number of hours worked per week{.}\footnote{Some workers who work very few hours per year may report zero hours as their usual number of hours per week worked. We do not account for this source of measurement error.} Following Blundell et al. (2003) we use the simulated out-of-work benefits income as an exclusion restriction in the hours equation. We refer to their paper for details and note that the UK benefits system makes this restriction appropriate as, in contrast to other European countries, unemployment benefits are not related to income prior to the period out of work. Blundell et al. (2007) argue in an application studying the behavior of both males and females that the system of housing benefits may have a positive relationship with the in-work potential. However, we do not address this possibility here, although we note that Blundell et al. (2007) suggest a potential solution using a monotonicity restriction rather than an exclusion restriction in the hours equation. \footnote{ Given this concern regarding the validity of the exclusion restriction we exploit the ability to identify the effects of interest via the variability in hours in the absence of an exclusion restriction. Accordingly we re-estimated the model with the simulated out-of-work benefits included and excluded from both steps. Although we do not report them here, the differences in results across these two identification schemes and those reported in the paper are minor. } Equations ((ref)) and ((ref)) characterize the model of Blundell et al. (2003) when $g$ and $h$ are linear and separable and $\varepsilon $ and $ \eta $ are normally distributed.

Figure (ref)A reports the female participation rate over the sample period. Figure (ref)B reports the average number of hours for all females and those reporting positive hours, respectively. Recall that the control function exploits the variation in both the extensive and the intensive margins of the hours decisions. Figure (ref)A illustrates that participation was around 65 percent in the years before the recession at the beginning of the 1980s. Participation drops to a sample period low of 58 percent in 1982, but subsequently increases and almost reaches 70 percent at the end of the sample period. The figures for the average hours show similar trends but most notably there is significant variation in average hours over time for the sample of workers.

We employ the following variables for our empirical analysis. Three different education levels; (1) a dummy variable indicating that the individual left school at the age of 16 years or younger, (2) a dummy variable for left school at the age of 17 or 18 years, and (3) a dummy variable for left school at the age of 19 years or older. We include age and age squared and interact these with the level of education. In addition, we use a dummy variable indicating that the individual lives together with a partner and 12 dummy variables indicating the region in the UK in which the individual lives. We pool data for four consecutive years, i.e. 1978-81, 1982-1985, 1986-1989, 1990-1993, 1994-1997 and 1998-2000, noting that the last period is only 3 years and we assume that the model's assumptions are satisfied for each of the pooled samples.

\pgfplotstableread{hours1.txt}{\hours} \pgfplotstableread{hours2.txt}{ \hoursa} \pgfplotstableread{result_part.txt}{\participation}

figure[figure omitted — 858 chars of source]
table[table omitted — 1,142 chars of source]

We first examine the returns to schooling.\footnote{ We treat the level of education as exogenous in this empirical example. Our approach could be extended to incorporate the endogeneity of conditioning variables but this is beyond the scope of this paper.} Table (ref) reports the impact of education on wages unadjusted for selection. The results in the first column are the average treatment effects of the difference between the lowest level of education and each of the two higher levels. Similarly, columns 2 to 4 report the absolute values of the quantile treatment effects.\footnote{ We calculate the average treatment effects for the middle education level as the difference between $\frac{1}{n}\sum_{i=1}w(Z_{1i},X_{i}=``middle^{\prime \prime })^{\mathrm{T}}\widehat{\beta }$ and $\frac{1}{n} \sum_{i=1}w(Z_{1i},X_{i}=``low^{\prime \prime })^{\mathrm{T}}\widehat{\beta } $, where $w$ is a polynomial and $\widehat{\beta }$ the linear-least squares estimator. We use distribution regression for the quantile treatment effect to estimate the distribution and calculate the quantile of that distribution. } The results in Table (ref) indicate that there is generally a larger coefficient at higher quantiles. There is also evidence of an increase in the return to education over time at some quantiles, although there is no evidence of a general increase in the impact of education over the sample period.

\captionsetup[subfigure]{labelformat=empty}

figure[figure omitted — 1,508 chars of source]

Before discussing the results produced by our approach we examine the estimates of the control function. Figure (ref) presents the joint densities of the control function and age for the different educational categories. Age and education are the two economically most interesting variables included in the wage equations. The nonparametric identification of the global effects discussed in Section (ref) (Assumption (ref)) requires that the supports of the control function are identical across the education levels considered. This appears to be satisfied and it is especially true for the beginning of the sample period. After 1990 the support for the lowest education level begins between 0.2 and 0.3 while that for the other education levels begins between 0.1 and 0.2.

table[table omitted — 2,069 chars of source]

Table (ref) reports the results of the local average and quantile treatment effects. We report the effects relative to the lowest level of education. These local effects report the increase in terms of the mean or quantile when we compare the wages that all individuals would receive with the lowest level of education to those they would receive if they had a higher level of education. These local effects correspond to the estimates described in Section (ref). For example, the impact at the mean for $V=0.5$ and \textquotedblleft Leaving school at the age of 17-18 years" is the average increase in wages of all individuals with $V=0.5$ if their education level were to increase from the lowest to the middle education level. This impact is hypothetical since not everyone with $V=0.5$ has either a low or middle education level. With respect to the discussion in Section (ref), it equals $\int \left\{ \mu (x_{1},z_{1},v)-\mu (x_{0},z_{1},v)\right\} dF_{Z_{1}|C>0}(z_{1})$, with $x_{1}$ the middle level of education and $x_{0}$ the lowest level of education. The quantile treatment effects are their corresponding quantiles. We account for sample selection by including $V$ and $V^{2}$ and the interaction of $V$ with all the other regressors discussed above. We report results for values of $V$ at the median and higher since it appears that our nonparametric identification requirements are not satisfied at lower quantiles.

Table (ref) reveals that the impact of education varies by quantile and the value of $V$ at which it is evaluated. The results for the mean reveal some variation in the returns to education across different values of $V$. However, the evidence is not strong statistically. This, in addition to the similarity of the results to the unadjusted results, suggests that there are no clear effects from selection bias.

To explicitly investigate for the presence of a selection bias, Table (ref) provides the average derivative of wages with respect to $V$ for each of the pooled samples. This represents a test for selection bias although a zero estimate for the average derivative does not exclude the presence of selection since the derivative may change signs depending on the value where it is evaluated.\footnote{ Deriving the optimal test for selection in this setting is beyond the scope of this paper.} The table shows that for each of the pooled samples $V$ has a positive and statistically significant effect on wages. This indicates that the unobservable characteristics which increase hours of work are positively correlated with those which increase wages. Thus, there is statistical evidence supporting the presence of selection bias and the impact of $V$ increases over time. As the support of $V$ is generally from 0.2 to 1, a coefficient of around 0.27 indicates that the impact of $V$ on wages is economically important.

table[table omitted — 362 chars of source]

Table (ref) reports the global treatment effects as presented in Section (ref). For example, the mean results for \textquotedblleft leaving school at the age of 17-18 years\textquotedblright\ is the impact on the average wage for the selected population when their education level changes from the low to the middle education level. From Section (ref), it equals ((ref)) for $x_{1}$ the middle level of education and $x_{0}$ the low education level. We do not find large differences between the results presented in Table (ref) and those reported in Table (ref). This also suggests that selection does not have an important effect on the returns to education.

The global treatment effects are nonparametrically identified under the strong support condition presented in Assumption (ref) and as discussed above this may not be satisfied for the later years of our analysis. We accordingly estimate bounds for the global effects. These are also reported in Table (ref) and are tight in this particular example. Since it is not apparent whether this reflects features of the data or the additional information inherent in the censored selection rule, we also constructed the corresponding bounds employing the propensity scores. This did not produce informative bounds.

table[table omitted — 1,718 chars of source]

\pgfplotstableread{global_5yrs_new/src/results_global1_1_5.txt}{\resultsa} \pgfplotstableread{global_5yrs_new/src/results_global1_2_5.txt}{\resultsb} \pgfplotstableread{global_5yrs_new/src/results_global1_3_5.txt}{\resultsc} \pgfplotstableread{global_5yrs_new/src/results_global2_1_5.txt}{\resultsaa} \pgfplotstableread{global_5yrs_new/src/results_global2_2_5.txt}{\resultsba} \pgfplotstableread{global_5yrs_new/src/results_global2_3_5.txt}{\resultsca} \pgfplotstableread{global_5yrs_new/src/results_global3_1_5.txt}{\resultsab} \pgfplotstableread{global_5yrs_new/src/results_global3_2_5.txt}{\resultsbb} \pgfplotstableread{global_5yrs_new/src/results_global3_3_5.txt}{\resultscb}

figure[figure omitted — 1,399 chars of source]

We also explore the influence of education by deriving the average and quantile impacts of obtaining a higher level of education for specific groups. These are the treatment effects for the treated discussed in Section (ref) and are shown in Figures (ref) and (ref). The labels \textquotedblleft Low\textquotedblright , \textquotedblleft Middle\textquotedblright\ and \textquotedblleft High\textquotedblright\ capture the three education groups. These effects are nonparametrically identified under the weak common support condition in Assumption (ref). This requirement seems satisfied as the support of the control function for the “Middle” and “High” education levels are at least as large as that of the education level “Low”, while the support of the control function for the education level “High” is at least as large as that of the education level “Medium” (see Figure (ref)). The estimates are based on pooling the data as described above. The figure \textquotedblleft Low versus middle\textquotedblright\ in Figure (ref) displays the average increase in wages when individuals of the low education group have an education level equal to the middle education level. These are the estimates described in Section (ref) and correspond to the marginal increase in the education attainment. In particular, we estimate (ref). These contrast to those in Table (ref) which report the impact on the mean or quantiles if everyone in the selected sample went from a low to a higher education level, irrespective of the observed education levels of these individuals. The figure displayed for $\tau =0.25$ in Figure (ref) presents this marginal increase at the first quartile of the distribution. The magnitude of the average impact of education for the various educational comparisons is consistent with the estimates discussed above and the plot over time appears to reveal some cyclical behavior.

\pgfplotstableread{global1_5yrs_new/src/results_global1_1_25.txt}{\resultsa} \pgfplotstableread{global1_5yrs_new/src/results_global1_2_25.txt}{\resultsb} \pgfplotstableread{global1_5yrs_new/src/results_global1_3_25.txt}{\resultsc} \pgfplotstableread{global1_5yrs_new/src/results_global2_1_25.txt}{\resultsaa} \pgfplotstableread{global1_5yrs_new/src/results_global2_2_25.txt}{\resultsba} \pgfplotstableread{global1_5yrs_new/src/results_global2_3_25.txt}{\resultsca} \pgfplotstableread{global1_5yrs_new/src/results_global3_1_25.txt}{\resultsab} \pgfplotstableread{global1_5yrs_new/src/results_global3_2_25.txt}{\resultsbb} \pgfplotstableread{global1_5yrs_new/src/results_global3_3_25.txt}{\resultscb}

\pgfplotstableread{global1_5yrs_new/src/results_global1_1_5.txt}{\resultsac} \pgfplotstableread{global1_5yrs_new/src/results_global1_2_5.txt}{\resultsbc} \pgfplotstableread{global1_5yrs_new/src/results_global1_3_5.txt}{\resultscc} \pgfplotstableread{global1_5yrs_new/src/results_global2_1_5.txt}{\resultsad} \pgfplotstableread{global1_5yrs_new/src/results_global2_2_5.txt}{\resultsbd} \pgfplotstableread{global1_5yrs_new/src/results_global2_3_5.txt}{\resultscd} \pgfplotstableread{global1_5yrs_new/src/results_global3_1_5.txt}{\resultsae} \pgfplotstableread{global1_5yrs_new/src/results_global3_2_5.txt}{\resultsbe} \pgfplotstableread{global1_5yrs_new/src/results_global3_3_5.txt}{\resultsce}

\pgfplotstableread{global1_5yrs_new/src/results_global1_1_75.txt}{\resultsaf} \pgfplotstableread{global1_5yrs_new/src/results_global1_2_75.txt}{\resultsbf} \pgfplotstableread{global1_5yrs_new/src/results_global1_3_75.txt}{\resultscf} \pgfplotstableread{global1_5yrs_new/src/results_global2_1_75.txt}{\resultsag} \pgfplotstableread{global1_5yrs_new/src/results_global2_2_75.txt}{\resultsbg} \pgfplotstableread{global1_5yrs_new/src/results_global2_3_75.txt}{\resultscg} \pgfplotstableread{global1_5yrs_new/src/results_global3_1_75.txt}{\resultsah} \pgfplotstableread{global1_5yrs_new/src/results_global3_2_75.txt}{\resultsbh} \pgfplotstableread{global1_5yrs_new/src/results_global3_3_75.txt}{\resultsch}

figure[figure omitted — 3,146 chars of source]

The different estimates of the impact of education may partially capture that the returns to education vary by birth cohort. For example, the increasing numbers of college educated females may reflect that selection into such education has changed (i.e. has become less or more demanding) and may result in lower or higher returns to education over time. We investigate this by estimating birth cohort specific age profiles. We find that older cohorts have slightly higher wages but that the differences are small.

We also explore how the return to experience has varied by education group over the sample period by estimating the average derivative with respect to age. Figure (ref) presents the derivative for different education levels. The figures show that there is a drastic increase to the return to experience during the 1990s. They also reveal that there is a drastic difference in the rate of wage growth across education groups. Figure (ref) reports these derivatives evaluated at ages 25, 40, and 55 years and these represent the local average responses. There is a strong positive relationship between wage growth and age at 25 years and the effect is particularly strong for the highest educated. Moreover, the effect increases notably over the sample period with large increases in the 1990s. The effect is notably lower although still positive at the age of 40 years. The differences by education groups are less dramatic. At the age of 55 years wages do not appear to be generally increasing with age. In fact, there appears to be evidence that the real wage is decreasing for the highest education group.

\pgfplotstableread{average_derivative/src/results_global1_1_5.txt}{\resultsa} \pgfplotstableread{average_derivative/src/results_global1_2_5.txt}{\resultsb} \pgfplotstableread{average_derivative/src/results_global1_3_5.txt}{\resultsc} \pgfplotstableread{average_derivative/src/results_global2_1_5.txt}{ \resultsaa} \pgfplotstableread{average_derivative/src/results_global2_2_5.txt}{ \resultsba} \pgfplotstableread{average_derivative/src/results_global2_3_5.txt}{ \resultsca} \pgfplotstableread{average_derivative/src/results_global3_1_5.txt}{ \resultsab} \pgfplotstableread{average_derivative/src/results_global3_2_5.txt}{ \resultsbb} \pgfplotstableread{average_derivative/src/results_global3_3_5.txt}{ \resultscb} \pgfplotstableread{average_derivative/src/results_global4_1_5.txt}{ \resultsac} \pgfplotstableread{average_derivative/src/results_global4_2_5.txt}{ \resultsbc} \pgfplotstableread{average_derivative/src/results_global4_3_5.txt}{ \resultscc}

figure[figure omitted — 1,718 chars of source]

\pgfplotstableread{local_average_response/src/results_global1_1_5_25.txt}{ \resultsa} \pgfplotstableread{local_average_response/src/results_global1_2_5_25.txt}{ \resultsb} \pgfplotstableread{local_average_response/src/results_global1_3_5_25.txt}{ \resultsc} \pgfplotstableread{local_average_response/src/results_global2_1_5_25.txt}{ \resultsaa} \pgfplotstableread{local_average_response/src/results_global2_2_5_25.txt}{ \resultsba} \pgfplotstableread{local_average_response/src/results_global2_3_5_25.txt}{ \resultsca} \pgfplotstableread{local_average_response/src/results_global3_1_5_25.txt}{ \resultsab} \pgfplotstableread{local_average_response/src/results_global3_2_5_25.txt}{ \resultsbb} \pgfplotstableread{local_average_response/src/results_global3_3_5_25.txt}{ \resultscb} \pgfplotstableread{local_average_response/src/results_global4_1_5_25.txt}{ \resultsac} \pgfplotstableread{local_average_response/src/results_global4_2_5_25.txt}{ \resultsbc} \pgfplotstableread{local_average_response/src/results_global4_3_5_25.txt}{ \resultscc}

\pgfplotstableread{local_average_response/src/results_global1_1_5_40.txt}{ \resultsamiddle} \pgfplotstableread{local_average_response/src/results_global1_2_5_40.txt}{ \resultsbmiddle} \pgfplotstableread{local_average_response/src/results_global1_3_5_40.txt}{ \resultscmiddle} \pgfplotstableread{local_average_response/src/results_global2_1_5_40.txt}{ \resultsaamiddle} \pgfplotstableread{local_average_response/src/results_global2_2_5_40.txt}{ \resultsbamiddle} \pgfplotstableread{local_average_response/src/results_global2_3_5_40.txt}{ \resultscamiddle} \pgfplotstableread{local_average_response/src/results_global3_1_5_40.txt}{ \resultsabmiddle} \pgfplotstableread{local_average_response/src/results_global3_2_5_40.txt}{ \resultsbbmiddle} \pgfplotstableread{local_average_response/src/results_global3_3_5_40.txt}{ \resultscbmiddle} \pgfplotstableread{local_average_response/src/results_global4_1_5_40.txt}{ \resultsacmiddle} \pgfplotstableread{local_average_response/src/results_global4_2_5_40.txt}{ \resultsbcmiddle} \pgfplotstableread{local_average_response/src/results_global4_3_5_40.txt}{ \resultsccmiddle}

\pgfplotstableread{local_average_response/src/results_global1_1_5_55.txt}{ \resultsahighest} \pgfplotstableread{local_average_response/src/results_global1_2_5_55.txt}{ \resultsbhighest} \pgfplotstableread{local_average_response/src/results_global1_3_5_55.txt}{ \resultschighest} \pgfplotstableread{local_average_response/src/results_global2_1_5_55.txt}{ \resultsaahighest} \pgfplotstableread{local_average_response/src/results_global2_2_5_55.txt}{ \resultsbahighest} \pgfplotstableread{local_average_response/src/results_global2_3_5_55.txt}{ \resultscahighest} \pgfplotstableread{local_average_response/src/results_global3_1_5_55.txt}{ \resultsabhighest} \pgfplotstableread{local_average_response/src/results_global3_2_5_55.txt}{ \resultsbbhighest} \pgfplotstableread{local_average_response/src/results_global3_3_5_55.txt}{ \resultscbhighest} \pgfplotstableread{local_average_response/src/results_global4_1_5_55.txt}{ \resultsachighest} \pgfplotstableread{local_average_response/src/results_global4_2_5_55.txt}{ \resultsbchighest} \pgfplotstableread{local_average_response/src/results_global4_3_5_55.txt}{ \resultscchighest}

figure[figure omitted — 3,281 chars of source]

This empirical investigation uncovers a number of interesting features of the data and highlights several benefits of our approach. First, despite the large change in the participation rates of females over this period there are no corresponding changes in the return to education. However, there is evidence of an increase in the return to experience in the 1990s. Our approach uncovers the presence of a selection bias and although this does not appear to influence the estimates of the return to education, there is an increase in the return to unobservables which influence hours over the sample period. Finally, through this empirical example we have illustrated the substantial reduction in bounds which results from the use of the control function associated with a censored selection rule in contrast to those from the use of the propensity score associated with a binary selection rule.

Conclusion

This paper examines a nonseparable sample selection model with a selection equation which is based on a censored outcome. We account for selection by conditioning on an appropriately constructed control function. We show that we are able to identify several economically interesting objects. We categorize these as local effects which represent estimands conditional on a specific outcome of the control function and global effects which represent estimands evaluated over a range of values of the control function. For both effects we provide identification results and estimation methods and the related asymptotic theory. We illustrate the utility of our approach in an empirical application focusing on the determinants of wages for a sample of females in the UK over a period of increasing labor force participation at both the intensive and extensive margins.

We provide $\sqrt{n}$-consistent estimators of the effects of interest under the correct semiparametric specification of the control function. This specification could also be used as the basis for sieve methods that are more robust to functional form assumptions (see, for example, Chen 2007). Analyzing the asymptotic properties of the resulting estimators whose dimensions grow with the sample size would require a different proof strategy. We rely on weak dependence of the estimator of the control function to apply the delta method, which is not available for sieve estimators. It might be possible to derive theory for sieve estimators by using different techniques such as strong approximations. This is a useful extension that we leave to future work.

thebibliography{99} \bibitem Amemiya, T. (1978), “The estimation of a simultaneous equation generalized probit model”, Econometrica 46, 1193--1205. \bibitem Amemiya, T. (1979), \textquotedblleft The estimation of a simultaneous tobit model\textquotedblright , International Economic Review 20, 169--81. \bibitem \textsc{Arellano, M. and S. Bonhomme} (2017), \textquotedblleft Quantile selection models with an application to understanding changes in wage inequality\textquotedblright , \emph{Econometrica} \textbf{85}, 1--28. \bibitem \textsc{Blundell, R., A. Gosling, H. Ichimura and C. Meghir} (2007), “Changes in the distribution of male and female wages accounting for employment composition using bounds”, \emph{Econometrica} \textbf{75}, 323--63. \bibitem \textsc{Blundell, R., H. Reed and T.M. Stoker} (2003), “Interpreting aggregate wage growth: the role of labor market participation”, \emph{American Economic Review} \textbf{93}, 1114--31. \bibitem \textsc{Buchinsky, M.} (1998), “The dynamics of changes in the female wage distribution in the USA: a quantile regression approach”, \emph{ Journal of Applied Econometrics} \textbf{13}, 1--30. \bibitem \textsc{Caetano, C.} (2015), \textquotedblleft A test of exogeneity without instrumental variables in models with bunching\textquotedblright , \textit{Econometrica} \textbf{83}, 1581--1600. \bibitem \textsc{Caetano, C., C. Rothe and N. Yildiz }(2016), “A discontinuity test for identification in triangular nonseparable models", \emph{Journal of Econometrics} \textbf{193}, 113--22. \bibitem \textsc{Chen, S.} (1997), “Semiparametric estimation of the Type-3 Tobit model”, \emph{Journal of Econometrics} \textbf{80}, 1--34. \bibitem \textsc{Chen, X.} (2007), “Large sample sieve estimation of semi-nonparametric models”, chapter 76 in Heckman, J.J. and Leamer, E.E. eds., \emph{Handbook of Econometrics}, vol. 6B, Elsevier, 5550--5632. \bibitem \textsc{Chernozhukov, V., I. Fern\'andez-Val and A. Kowalski} (2015), “Quantile regression with censoring and endogeneity”, \emph{ Journal of Econometrics} \textbf{186}, 201--21. \bibitem \textsc{Chernozhukov, V., I. Fern\'{a}ndez-Val and B. Melly} (2013), \textquotedblleft Inference on counterfactual distributions\textquotedblright , \emph{Econometrica} \textbf{81}, 2205--68. \bibitem \textsc{Chernozhukov, V., I. Fern\'{a}ndez-Val, W.K. Newey, S. Stouli and F. Vella} (2020), \textquotedblleft Semiparametric estimation of structural functions in nonseparable triangular models\textquotedblright , \textit{Quantitative Economics} \textbf{11}, 503--33. \bibitem \textsc{Chesher, A.} (2003), \textquotedblleft Identification in nonseparable models\textquotedblright , \textit{Econometrica} \textbf{71}, 1401--44. \bibitem \textsc{Dinardo, J., N.M. Fortin and T. Lemieux} (1996), \textquotedblleft Labor market institutions and the distribution of wages, 1973-1992: a semiparametric approach\textquotedblright , \emph{Econometrica} \textbf{64}, 1001--44. \bibitem \textsc{Firpo, S., N.M. Fortin and T. Lemieux} (2011), \textquotedblleft Decomposition methods in economics\textquotedblright\, chapter 1 in D. Card and O. Ashenfelter, eds., Handbook of Labor Economics, 4th Edition, Elsevier North Holland, 1--102. \bibitem \textsc{Florens, J., J.J. Heckman, C. Meghir and E. Vytlacil} (2008), \textquotedblleft Identification of treatment effects using control functions in models with continuous, endogenous treatment and heterogeneous effects\textquotedblright, \textit{Econometrica} \textbf{76}, 1191--1206. \bibitem \textsc{Foresi, S. and F. Peracchi} (1995), “The conditional distribution of excess returns: an empirical analysis”, \emph{Journal of the American Statistical Association} \textbf{90}, 451--66. \bibitem \textsc{Hahn, J.} (1995), \textquotedblleft Bootstrapping quantile regression estimators\textquotedblright, \emph{Econometric Theory} \textbf{11}, 105--21. \bibitem \textsc{Heckman, J.J.} (1979), \textquotedblleft Sample selection bias as a specification error\textquotedblright , \textit{Econometrica} \textbf{47}, 153--61. \bibitem \textsc{Heckman, J.J. and E. Vytlacil} (2005), \textquotedblleft Structural equations, treatment effects, and econometric policy evaluation\textquotedblright , \emph{Econometrica} \textbf{73}, 669--738. \bibitem \textsc{Hoderlein, S. and E. Mammen} (2007), “Identification of marginal effects in nonseparable models without monotonicity", \textit{\ Econometrica} \textbf{75}, 1513--18. \bibitem \textsc{Honor\'e, B., E. Kyriazidou and C. Udry} (1997), “Estimation of type 3 tobit models using symmetric trimming and pairwise comparisons”, \emph{Journal of Econometrics} \textbf{76}, 107--28. \bibitem \textsc{Imbens, G.W. and W.K. Newey} (2009), “Identification and estimation of triangular equations models without additivity”, \emph{ Econometrica} \textbf{77}, 1481--1512. \bibitem \textsc{Jun, S.J.} (2009), \textquotedblleft Local structural quantile effects in a model with a nonseparable control variable\textquotedblright , \emph{Journal of Econometrics} \textbf{151}, 82--97. \bibitem \textsc{Koenker, R. and G. Bassett} (1978), \textquotedblleft Regression quantiles\textquotedblright , \textit{Econometrica} \textbf{46}, 33--51. \bibitem \textsc{Lee, M.J. and F. Vella} (2006), “A semi-parametric estimator for censored selection models with endogeneity”, \emph{Journal of Econometrics} \textbf{130}, 235--52. \bibitem \textsc{Ma, L. and R. Koenker} (2006), \textquotedblleft Quantile regression methods for recursive structural equation models\textquotedblright, \emph{Journal of Econometrics} \textbf{134}, 471--506. \bibitem \textsc{Ma, S. and M. Kosorok} (2005), “Robust semiparametric M-estimation and the weighted bootstrap”, \emph{Journal of Multivariate Analysis} \textbf{96}, 190--217. \bibitem \textsc{Masten, M. and A. Torgovitsky} (2016), \textquotedblleft Instrumental variables estimation of a generalized correlated random coefficients model\textquotedblright , \emph{Review of Economics and Statistics} \textbf{98}, 1001--1005. \bibitem \textsc{Matzkin, R.} (2003), \textquotedblleft Nonparametric estimation of nonadditive random functions\textquotedblright, \textit{ Econometrica} \textbf{71}, 1339--75. \bibitem \textsc{Newey, W.K.} (2007), \textquotedblleft Nonparametric continuous/discrete choice models\textquotedblright , \emph{International Economic Review} \textbf{48}, 1429--39. \bibitem \textsc{Newey, W.K. and S. Stouli} (2018), “Control variables, discrete instruments, and identification of structural functions”, arXiv preprint arXiv:1809.05706. \bibitem \textsc{Praestgaard, J. and J. Wellner} (1993), “Exchangeably weighted bootstraps of the general empirical process”, \emph{Annals of Probability} \textbf{21}, 2053--86. \bibitem \textsc{Van der Vaart, A.W.} (1998), \textquotedblleft \emph{ Asymptotic statistics}\textquotedblright, Cambridge University Press. \bibitem \textsc{Van der Vaart, A.W. and J.A. Wellner} (1996), \textquotedblleft \emph{Weak convergence and empirical processes} \textquotedblright, Springer, New York. \bibitem \textsc{Vella, F.} (1993), \textquotedblleft A simple estimator for simultaneous models with censored endogenous regressors\textquotedblright , \emph{International Economic Review} \textbf{ 34}, 441--57.