Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,528 characters · 7 sections · 63 citation commands
Testing for Quantile Sample Selection
Abstract This paper provides nonparametric tests for detecting sample selection in conditional quantile functions. The first test is an omitted predictor test with the propensity score as the omitted variable, which holds uniformly across a compact set of quantile ranks. As with any omnibus test, in the case of rejection we cannot distinguish between rejection due to genuine selection or to misspecification. Thus, we suggest a second test to provide supporting evidence whether the cause for rejection at the first stage was solely due to selection or not. Using only individuals with propensity score close to one, this second test relies on an `identification at infinity' argument, but accommodates cases of irregular identification. Importantly, neither of the two tests requires parametric assumptions on the selection equation nor a continuous exclusion restriction. Data-driven bandwidth procedures for both tests are proposed, and simulation evidence suggests a good finite sample performance in particular of the first test. We apply our procedure to test for selection in log hourly wages using UK Family Expenditure Survey data as AB2017. Our findings indicate that some of the evidence for sample selection among males may instead reflect misspecification of the quantile functions. Finally, we also derive an extension of the first test to conditional mean functions.\newline \newline
Key-Words: Nonparametric Estimation, Conditional Quantile Function, Irregular Identification, Specification Test, Selection into Employment. \newline
\noindentJEL Classification: C12, C14, C21.\newline
\doublespacing
Empirical studies using non-experimental data are often plagued by the presence of non-random sample selection: individuals typically self select themselves into employment, training programs etc. on the basis of characteristics which are believed to be non-randomly distributed and unobservable to the researcher(s) G1974,H1974. In fact, it is well known that ignoring selection in conditional mean models induces a bias in the estimation, which can be additive H1979,DNV2003 or multiplicative J15 depending on the functional form of the model. In both cases, one can deal with the selection bias by adopting a control function approach. On the other hand, until recently little was known about identification and estimation of conditional quantile models in the presence of sample selection, see the recent survey by ABsurvey2017. A notable exception is the case of sample selection in `location shift' models, where only the intercept is allowed to vary across quantile ranks, and the control function approach still applies HM2015. In all other cases, however, including linear quantile regression models, the presence of sample selection is more difficult to deal with and control function methods can no longer be used. In fact, AB2017 proposed an alternative three-step estimator for parametric linear conditional quantile models, which has recently gained popularity in the literature. However, identification and estimation in AB2017 require either parametric restrictions on the joint distribution of the unobservable error terms or functional form conditions on the latter, which are hard to verify and still rely on a parametric choice in practice.
In this paper, we therefore propose two different nonparametric tests for detecting sample selection in conditional quantile functions. These tests only impose a minimal set of functional form assumptions on both the outcome and the selection equation(s). In fact, the only additional assumption is that selection (if present) affects the outcome through the propensity score, the probability to be in the selected sample, a standard assumption in the sample selection literature DNV2003. Importantly, the tests do also not require a continuous exclusion restriction such as a continuous instrumental variable in the selection equation. This is of particular relevance for empirical work as various examples from the literature have used the three-step estimator of AB2017 without relying on such a variable. For example, recent studies examining the effects of sample selection in conditional wage distributions from the Current Population Survey (CPS) in the context of nonresponse BHHZ2019 and the gender wage gap MW2019 rely on discrete instrumental variables such as `month-observed-in-sample' and `number and presence of children', respectively.
Formally, the first test we propose is a test for omitted predictors, where the omitted predictor is the propensity score. The test statistic is similar to that of VBDN2013, but holds holds uniformly over a compact subset $\mathcal{T}\subset (0,1)$ of quantile ranks. The statistic is constructed as a weighted average of conditional quantile errors. Since the latter are not observed, we need to estimate them in the first place, which in turn requires estimation of the nonparametric conditional quantile function under the null hypothesis. This is done using the local polynomial estimator of GS2012. We derive an asymptotic representation for the statistic, which features estimation error from the conditional quantile function, but not from the propensity score EJCL2014. Importantly, we can test the null of no selection over different subsets of quantile ranks simultaneously. As the limiting distribution is non-pivotal and depends on features of the Data Generating Process, we establish the first order validity of wild bootstrap critical values. Moreover, to select the bandwidth parameter(s) for this test, we suggest a data-driven way that relies on a procedure suggested by LR2008 and that allows for discrete covariates. Finally, in Appendix B we also provide an extension of this test to nonparametric conditional mean functions. The latter are commonly used in practice and tests for sample selection have so far relied on the correct (semi-)parametric specification of these mean function under the null BJLST2015.
As with other omnibus tests, if we reject the null of the first test, we cannot distinguish between a rejection due to genuine sample selection or to omission of other relevant predictor variable(s), not independent of the propensity score. This distinction is crucial when estimation of nonparametric conditional quantile functions is the ultimate goal. In fact, sample selection generally leads to a loss of point identification AB2017, but consistent estimation and inference may still be carried out on a subset of observations with propensity score close to one. By contrast, omitting relevant predictors impedes consistent estimation and inference altogether. To understand the heuristics of our second test, note that dependence of the conditional quantile error on the propensity score when the latter is (close to) one hints at the presence of omitted relevant predictors. This is so because individuals with a propensity score equal to one are selected into the sample almost surely and thus sample selection bias is not present in that case. By contrast, our conditional quantile function is likely to miss out on relevant predictor(s) correlated with the propensity score if it depends on the latter, regardless of whether it takes on values in the interior of the unit interval or close to one. The null hypothesis of the second test formalizes this heuristic argument.
The test statistic of the second test is constructed as a weighted average of `quantile errors' from individuals with (estimated) propensity score close to one, so that selection bias does no longer `bite'.\footnote{ Thus, to obtain power against misspecification in this test we require that the omitted predictor(s) are correlated with the propensity score, even when the latter is at or close to one.} However, while the second test relies on a so called `identification at infinity' argument and thus requires observations with estimated propensity score `close to one' (in a nonparametric sense), our test does not need continuous instrument(s) and allows for a thin set of observations close to the boundary, thus accommodating cases of so called irregular identification KT2010. In fact, the rate of convergence of the second test depends on the degree of irregularity of the marginal density of the propensity score. We therefore suggest a studentized version of the test statistic, which is rate adaptive and converges weakly even if numerator and denominator of the statistic diverge individually at the same rate.
We conduct a Monte Carlo study to asses the finite sample properties of our two tests: while the first test performs well in terms of size and power throughout all designs even when the exclusion restriction consists of a discrete instrument and the tuning parameters are chosen in a completely automated manner, the results of the second test appear to be somewhat more sensitive to the choice of the tuning parameters as well as to the continuity (or discreteness) of the instrument(s). This reflects the irregularity of the underlying problem.
We apply our testing procedure to test for selection in log hourly wages of females and males in the UK using data from the UK Family Expenditure Survey from 1995 to 2000. The same data was recently also used by AB2017 to analyze gender wage inequality in the UK. We run our testing procedure on two different sub-periods of different economic performance, namely 1995-1997 and 1998-2000. As a preview of the results, we cannot find evidence for selection among females for the 1995-1997 period, but only for the 1998-2000 period. By contrast, while we reject the null of the first test for males with data from 1995 to 1997, our second test strongly suggests that this rejection may actually be due to misspecification of the quantile function, a feature that might have remained undetected without our testing procedure.
The rest of the paper is organized as follows. Section (ref) outlines the set-up. Section (ref) then establishes the limiting behavior of the first test for omitted variables, and the first order validity of inference based on wild bootstrap critical values. Section (ref) on the other hand derives the same results for the second test, and Section (ref) reports the findings of our Monte Carlo Study. Finally, Section (ref) provides an empirical illustration in which we apply our procedure to log hourly wages of females and males in the UK and Section (ref) concludes. The definition of various nonparametric estimators used in the construction of the test statistics, and proofs of the main Theorems 1 and 2 are provided in Appendix A. On the other hand, a formal outline of the first test for the conditional mean, together with a set of Monte Carlo simulations can be found in Appendix B.
The proofs of several technical Lemmas as well as of results on the first order asymptotic validity of bootstrapped critical values have been relegated to the supplementary material. In addition, the supplementary material contains: (i) a formal argument for a data-driven bandwidth choice in the second test; (ii) a two-step testing procedure with a formal decision rule; (iii) a formal proof that the classification errors of this procedure are asymptotically controlled at pre-specified levels; (iv) proofs of the conditional mean test.
We begin by outlining the data generating process. As it is customary in the sample selection literature, we postulate that the continuous outcome variable of interest, $y_{i}$, is observed if and only if $s_{i}=1$, where $ s_{i}$ denotes a binary selection indicator. A standard application example of this set-up is for instance to the study of wages and employment: (log) wages $y_{i}$ are only observed for individuals who participate in the labor market and who are employed ($s_{i}=1$), and different sub-groups (e.g., males and females) may differ in terms of their unobservable labor market attachment. Thus, conventional measures of wage gaps or wage inequality may be biased H1974,H1979. For every individual $i$, we observe $x_{i}$ and $z_{i}$. The variable(s) in $x_{i}$ affect observed outcome $y_{i}$, while the variable(s) in $z_{i}$ predict the selection variable $s_{i}$, but not directly observed outcome $y_{i}$ once we condition on $x_{i}$.\footnote{The assumption that $z_{i}$ is statistically independent $y_{i}$ once we condition on $x_{i}$ and selection $s_{i}=1$ is testable using for example K2010. } Note that the variables in $x_{i}$ and $z_{i}$ need not be disjoint as in most applications, although our testing procedure requires at least one of the variable(s) in $z_{i}$ to be excluded from $x_{i}$ (cf. Assumption A.1 below). This is typically referred to as instrumental exclusion restriction in the literature and we will therefore refer to the element(s) of $z_{i}$ that are excluded from $x_{i}$ as instrument(s) or instrumental variable(s) in what follows. Importantly, however, none of the instrument(s) in $z_{i}$ needs to be continuous, an aspect that we will explore in greater detail in the simulations in Section (ref).
Throughout the paper, the maintained assumptions are that (i) the variable(s) in $z_{i}$ and (ii) non-random selection (if present) enter the conditional quantile function only through the propensity score $ \Pr \left( s_{i}=1|z_{i}\right) =p(z_{i})\equiv p_{i}$, the probability to be in the selected sample for a given $z_{i}$. Formally, this can be expressed as follows:
A.Q For all $\tau $,
holds almost surely, where $q_{\tau }\left( x_{i}\right) $ denotes the conditional $\tau $-quantile of $y_{i}$ given $x_{i}$ and selection $s_{i}=1$.
Assumption A.Q is a high-level assumption and indeed the quantile equivalent of Assumption 2.1(i) in DNV2003. It is for instance implied by the model of AB2017 with the threshold crossing selection equation $s_{i}=1\{p(z_{i})>v_{i}\}$ and $x_{i}\subset z_{i}$.\footnote{Indeed, AB2017 require that at least one element in $z_{i}$ exists, which is not in $x_{i}$.} In their paper:
with $y_{i}^{\ast}=q(x_{i},u_{i})$ where the unobservables $(u_{i},v_{i})$ are jointly statistically independent of $z_{i}$ given $x_{i}$ and are absolutely continuous (with standard uniform marginals). In this set-up, provided $p_{i}>0$ with probability one, we obtain that:
Note that in fact $q_{\tau }(x_{i})$, the `observed' $\tau$ quantile of $y_{i}$ given $x_{i}$ (and $s_{i}=1$), coincides with the $\tau $ quantile of $y_{i}^{\ast }$ given $x_{i}$ and selection when selection is random.
Given the existence of a valid instrument, we test the hypothesis that the propensity score is not an omitted predictor, against its negation. In what follows, let $\mathcal{T}=[\underline{\tau },\overline{\tau }]$ denote the compact set of quantile ranks to be examined, where $0<\underline{ \tau }\leq \overline{\tau }<1$. Also, we use $\mathcal{X}$ to denote a compact set in the interior of the union of the supports of covariates $ R_{x} $, and $\mathcal{P}$ to denote a compact subset of the support of $ p(z_{i})$. In the first step we test $H_{0,q}^{(1)}$ versus $H_{A,q}^{(1)}$ using the subset of selected individuals for which $s_{i}=1$, i.e.
versus
The logic behind $H_{0,q}^{(1)}$ vs. $H_{A,q}^{(1)}$ is that, given ((ref)),
if and only if $\Pr \left( y_{i}\leq q_{\tau }(x_{i})|x_{i},p_{i},s_{i}=1\right) =\Pr \left( y_{i}\leq q_{\tau }(x_{i})|x_{i},s_{i}=1\right)$. Note that the null hypothesis of no omitted predictor in ((ref)) could have been also stated in terms of Conditional Distribution Functions (CDFs), i.e.
for all $x\in \mathcal{X}$, $p\in \mathcal{P}$, and $y\in \mathcal{Y}=\{q_{\tau}(x):x\in\mathcal{X},\tau\in\mathcal{T}\}$, subset of the support of $y_{i}$. A test for the null in ((ref)) could for instance be based on the difference between CDFs estimated using a larger and a smaller information set. However, while this may circumvent the issue of extreme quantile estimation, in the context of sample selection, interest often lies in specific quantiles or a subset of quantile ranks. For instance, we might only be interested in testing for sample selection in the (log) wage distribution of males and females from lower conditional quantiles such as from the 10% to the 25% quantiles, or of individuals that earn below the (conditional) median wage etc.. To carry out this type of analysis in the conditional distribution function context would require finding corresponding values say $y_{1}$ and $y_{2}$ to examine all $y$ such that $y_{1}\leq y\leq y_{2}$. These values are typically unknown and require estimating the conditional quantiles in the first place.
We now introduce a statistic for testing $H_{0,q}^{(1)}$ vs. $H_{A,q}^{(1)}$ , as defined in ((ref)) and ((ref)). Moreover, for notational simplicity, from here onwards we assume that all components of $x_{i}$ are continuous, while we make the possibility of discrete elements in $z_{i}$ explicit by partitioning it into $z_{i}=(z_{i}^{c},z_{i}^{d})$, where `$c$' (`$d$') denotes the sub-vector of continuous (discrete) elements. The extension to discrete elements in $x_{i}$ is immediate at the cost of more complicated notation and more lengthy arguments in the proofs. In fact, from HRL2004 and LR2008 we know that discrete covariates do not contribute to the rate at which the variance approaches zero, and hence they do not add to the curse of dimensionality. We discuss the extension to the case of discrete elements in $x_{i}$ in more detail in the empirical application section. In addition, note that partitioning the vector $z_{i}$ into continuous and discrete elements allows for the possibility that $ x_{i}=z_{i}^{c}$, and that only $z_{i}^{d}$ satisfies the exclusion restriction.
To implement our test, we rely on a statistic very close to that of VBDN2013. This statistic has the advantage of requiring an estimate of the conditional quantile function only under the null hypothesis, i.e. where the conditional quantile is a function of $x_{i}$ only. To estimate the conditional quantile function(s) at some point $x_{i}=x$, we use an $r$ -th order local polynomial estimator, which we denote by $\widehat{q}_{\tau }(x)$, while its corresponding probability limit is denoted by $q_{\tau }^{\dag }(x)$, which are formally defined in the Appendix Equations ((ref)) and ((ref)). Moreover, define { $\widehat{u} _{\tau }(x_{i})\equiv y_{i}-\widehat{q}_{\tau }(x_{i})$, $u_{\tau }(x_{i})\equiv y_{i}-q_{\tau }^{\dag }(x_{i})$, and let }$\underline{x}=( \underline{x}^{1},...,\underline{x}^{d_{x}}),$ and $\overline{x}=\left( \overline{x}^{1},...,\overline{x}^{d_{x}}\right) $, $\underline{x},\overline{ x}\in \mathcal{X}$, where $d_{x}$ denotes the dimension of $x_{i}$. Finally, for notational simplicity, we assume that $\mathcal{P}$ is a connected interval with boundary points $(\underline{p},\overline{p})$.\footnote{ The extension to a disconnected interval is immediate at the cost of more complex notation.} The test statistic is given by:
where
The statistc $Z_{1,n}\left( \tau ,\underline{x},\overline{x},\underline{p}, \overline{p}\right) $ differs from VBDN2013 in two aspects. First and most importantly, our test statistic is constructed taking the supremum also w.r.t. $\tau $ (over $\mathcal{T}$). We therefore test for selection across all quantile ranks in a compact set $\mathcal{T}$ simultaneously, while the test of VBDN2013 would have only allowed for a pointwise search. Second, the omitted regressor $p_{i}$ is not observable and thus replaced by a nonparametric estimator, $\widehat{p}_{i}$. Heuristically, the uniformity of our test is achieved via the use of a local polynomial quantile estimator for which GS2012 established a Bahadur representation uniform over compact sets $\mathcal{X}$ and $ \mathcal{T}$. As for the nonparametric estimator of the propensity score, under regularity and bandwidth conditions outlined below, we show that the estimation error arising from $\widehat{p}_{i}$ is asymptotically negligible. This is a well known result for estimates affecting the statistic only through a weight function EJCL2014.
Finally, note that one could construct an alternative statistic in which the quantile estimator is replaced by a conditional CDF estimator, say:
with $\widehat{F}_{y|x,s_{i=1}}(y|x_{i})$ denoting a nonparametric estimator of the CDF $F _{y|x,s_{i=1}}(y|x_{i})$. However, as discussed in the Section (ref), such a CDF based formulation will not allow to test the null of no selection for specific subsets of the (conditional) distribution.
In the sequel, we make the following assumptions:
A.1 $(y_{i},x_{i}^{\prime },z_{i}^{c\prime },z_{i}^{d\prime },s_{i})\subset R_{y}\times R_{x}\times R_{z}^{c}\times R_{z}^{d}\times \{0,1\}$ are identically and independently distributed. Let $ \mathcal{X}\equiv \mathcal{X}_{1}\times \ldots \times \mathcal{X}_{d_{x}}$ denote a compact subset of the interior of $R_{x}$. $z_{i}$ contains at least one variable which is not contained in $x_{i}$ and which is not $x_i$ -measurable. The variables $x_{i}$ and $z_{i}^{c}$ (conditional on all values of $z_{i}^{d}$) have probability density functions with respect to Lebesgue measure which are strictly positive and continuously differentiable (with bounded derivatives) over the interior of their respective support. Also, assume that the joint density function of $y_{i}$, $x_{i}$ and $p_{i}$ is uniformly bounded everywhere, and that $\Pr (s_{i}=1|x,p)=\Pr (s_{i}=1|p)>0$ for all $x\in\mathcal{X}$ and $p\in\mathcal{P}$.
A.2 The distribution function $ F_{y|x,s=1}(\cdot |\cdot ,\cdot )$ of $y_{i}$ given $x_{i}$ and selection $ s_{i}=1$ has a continuous probability density function $f_{y|x,s=1}(y|x,s=1)$ w.r.t. Lebesgue measure which is strictly positive and bounded for all $y\in R_{y}$, $x\in \mathcal{X}$. The partial derivative(s) $\nabla _{x}F_{y|x,s=1}(y|x,s=1) $ are continuous on $R_{y}\times \mathcal{X}$. Moreover, there exists a positive constant $C_{1}$ such that:
for all $(y,x),(y^{\prime },x^{\prime })\in R_{y}\times \mathcal{X}$. Also assume that $q_{\tau }(x)$ is $r+1-$th times continuously differentiable on $ \mathcal{X}$ for all $\tau \in \mathcal{T}$ with $r>\frac{1}{2}d_{x}$.
A.3 There exists an estimator $\widehat{p} (z_{i}^{c},z_{i}^{d})$ such that for any value in the support of $ z_{i}^{d}=z^{d}$, it holds that $\sup_{z^{c}\in\mathcal{Z}}|\widehat{p} (z^{c},z^{d})-p(z^{c},z^{d})|=o_{p}(n^{-\frac{1}{4}})$ with $\mathcal{Z}$ a compact subset of $R_{z}^{c}$, and that:
A.4 For some positive constant $C_{2}$, it holds that:
for all $\tau \in \mathcal{T}$, $(p,p^{\prime })\in \mathcal{P}$, and $ (x,x^{\prime })\in \mathcal{X}$, where $F_{p|x,u_{\tau },s=1}(p|x,0,s=1)$ denotes the conditional distribution function of $p_{i}$ given $x_{i}=x$, $ u_{\tau }(x)=0$, and $s_{i}=1$.
A.5 The non-negative kernel function $K(\cdot )$ is a bounded, continuously differentiable function with uniformly bounded derivative and compact support on $[-1,1]$. It satisfies $\int K(v)dv=1$ as well as $\int vK(v)dv=0$.
Assumption A.1 imposes the existence of at least one, continuous or discrete, element excluded from $x_{i}$. It guarantees the existence of selected observations for all values in $\mathcal{X}$ and $ \mathcal{P}$. Assumptions A.2 and A.4 on the other hand are rather standard smoothness assumptions, while A.3 is a high-level condition, which ensures that $p(z_{i}^{c},z_{i}^{d})$ can be estimated at a specific rate uniformly over $\mathcal{Z}$ so that estimation error in $\widehat{p}(z_{i}^{c},z_{i}^{d})$ is asymptotically negligible, while the second part of A.3 ensures that values of $z_{i}^{c}$ outside that set are asymptotically negligible for the estimation of the propensity score. In fact, in the case where $\widehat{p}(z_{i}^{c},z_{i}^{d})$ is a local constant kernel estimator, the use of a second order kernel imposes restrictions on the dimensionality of the number of continuous regressors, $ d_{z}$, namely $d_{z}<4$.\footnote{ Let $h_{z}=(h_{zc},h_{zd})$ and $d_{z}=(d_{zc},d_{zd})$. If we estimate $ \widehat{p}_{i}$ using a local constant estimator and set $h_{zc}=O(n^{- \frac{1}{4+d_{zc}}})$ and $h_{zd}=O(n^{-\frac{2}{4+d_{zc}}}),$ then for $ d_{zc}<4$ Assumption A.3 is satisfied. This is because for $d_{zc}<4 $, the bias of the continuous component is of order $n^{-\frac{2}{4+d_{zc}} }=o\left( n^{-1/4}\right) $, and the standard deviation for the continuous component is of order $(\sqrt{nh_{d_{zc}}})^{-1}=o\left( n^{-1/4}\right) $. As for the discrete component, the bias is of order $n^{-\frac{1}{4+d_{zc}} }=o\left( n^{-1/4}\right) $ and it does not contribute to the variance, see e.g. Theorem 2.1 in LR2008.} Note also that a sufficient condition for the second part of \textbf{A.3} is the existence of sufficient moments. Letting `$\Rightarrow $' denote weak convergence, we establish the asymptotic behavior of $Z_{1,n}^{q}$.
In the sequel, let $h_{x}$ be the bandwidth used in the estimation of the conditional quantile.
Theorem 1: Let Assumptions A.1-A.5 and A.Q hold. Moreover, if as $n\rightarrow \infty ,$ $ (nh_{x}^{2d_{x}})/\log n\rightarrow \infty $, $nh_{x}^{2r}\rightarrow 0,$ then
(i) under $H_{0,q}^{(1)}$,
where $Z_{1}^{q}$ is the supremum of the absolute value of a zero mean Gaussian process whose covariance kernel is defined in the proof of Theorem 1.
(ii) under $H^{(1)}_{A,q},$ there exists $\varepsilon >0,$ such that
The results of Theorem 1 rely on an appropriate choice of $h_{x}$. As common in the nonparametric testing literature, our rate conditions require undersmoothing, and thus cross-validation is not directly applicable in our setting.\footnote{ In fact, the order of the bandwidth selected by cross-validation is too large for $nh_{x}^{2r}\rightarrow 0$.} However, to still pick $h_{x}$ in a data-driven manner ensuring minimal bias at the same time, one possibility to select $h_{x}$ in practice could be to choose $h_{x}$ on the basis of cross-validation for a local polynomial estimator of order smaller than the one assumed for the test. For instance, if the assumed polynomial order for the test was $r=3$ as an example, $h_{x}$ could be chosen by cross-validation for a local linear estimator, i.e. $h_{x}=O\left( n^{-\frac{ 1}{4+d_{x}}}\right) $. This in turn implies that $nh_{x}^{2r}\rightarrow 0$ as well as $nh_{x}^{2d_{x}}/\log (n)\rightarrow \infty $ whenever $d_{x}<4$. In the Monte Carlo simulations of Section (ref), we demonstrate that this procedure in fact seems to perform well for the chosen designs.
Finally, note that in the proof of Theorem 1 we show that, under $ H_{0,q}^{(1)}$, $Z_{1,n}^{q}\left( \tau ,\underline{x},\overline{x}, \underline{p},\overline{p}\right) $ has the following asymptotic representation:
which holds uniformly over $\mathcal{T}$, $\mathcal{X}$ and $\mathcal{P}$. The term involving the difference of conditional distribution functions $ (F_{p|x,u_{\tau },s=1}(\overline{p}|x_{i},0,s_{i}=1)-F_{p|x,u_{\tau },s=1}( \underline{p}|x_{i},0,s_{i}=1))$ stems from the contribution of the quantile estimation error to the asymptotic representation. It therefore becomes evident that estimation error from the estimated propensity score, $\widehat{ p}_{i}$, on the other hand does not play a role in this representation, a finding that was also corroborated by EJCL2014.
Since the limiting distribution $Z_{1}^{q}$ depends on features of the data generating process, we derive a bootstrap approximation for it. In particular, we follow HZ2003, and use the bootstrap statistic:
where $B_{i,\tau }=1\left\{ U_{i}\leq \tau \right\} $ with $U_{i}\overset{ i.i.d.}{\sim }U(0,1)$, and independent of the sample, and $\widehat{F} _{p|x,u_{\tau },s=1}\left( p|x_{i,}0,s_{i}=1\right) $ denotes a nonparamemtric kernel estimator with corresponding bandwidth sequence $h_{F}$ satisfaying $h_{F}\rightarrow 0$ as $n\rightarrow \infty $ (see Equation ( (ref)) in the Appendix for a formal definition). The bootstrap test statistic is then given by:
Let $c_{(1-\alpha ),n,R}^{\ast (1)}$ be the $(1-\alpha )$ percentile of the empirical distribution of $Z_{1,n}^{\ast q,1},...,Z_{1,n}^{\ast q,R},$where $ R$ is the number of bootstrap replications. The following Theorem establishes the first order validity of inference based on the bootstrap critical values, $c_{(1-\alpha ),n,R}^{\ast }$.
Theorem 1$^{\ast }$: Let Assumption A.1- A.5 and A.Q hold. If as $n\rightarrow \infty$, { $ (nh_{x}^{2d_{x}})/\log n\rightarrow \infty $, $nh_{x}^{2r}\rightarrow 0,$ }$ h_{F}\rightarrow 0,$ $nh_{F}^{d_{x}+1}\rightarrow \infty ,$ { and }$R\rightarrow \infty ,$ then
(i) under $H_{0,q}^{(1)}$
(ii) under $H_{A,q}^{(1)}$
Under ((ref)) and the assumptions outlined in the previous section, failure to reject $H_{0,q}^{(1)}$ rules out sample selection asymptotically, with probability approaching one. By contrast, rejection in this first test could in principle occur either due to genuine sample selection or due to an omitted variable in the outcome equation, which happens to be correlated with the propensity score. This is so since the omitted predictor test, as any omnibus test, does not possess directed power against specific alternatives. Since this distinction is crucial for the estimation of nonparametric conditional quantile functions as outlined in the Introduction, we design a second test which has directed power against detecting misspecification. To this end, let $\widetilde{q}_{\tau }(x_{i},\pi _{i})$ and $q_{\tau }(x_{i})$ denote probability limits of two local polynomial quantile estimators for the selected subsample of $y_{i}$ on $x_{i}$ and $\pi _{i}$ as well as on $x_{i}$ only, respectively. We say that $\pi _{i}$ is a relevant predictor if for some $\tau\in\mathcal{T}$ and some value $\pi$ in the support of $\pi_{i}$, $\widetilde{q}_{\tau }(x,\pi )\neq q_{\tau }(x)$ for at least all $x$ in a subset of $\mathcal{X}$ with non-zero Lebesgue measure.
We want to disentangle selection from relevant omitted predictors correlated with the propensity score, which may (or may not) be present simultaneously with sample selection. In order to impose no-selection as maintained hypothesis, we require that at least one value $z$ in the support of $z_{i}$ s.t. $p(z)=1$ exists, which is indeed one of the identification assumptions in AB2017. Then, in the absence of relevant omitted predictors $ \pi _{i}$ whenever $p\rightarrow 1$, the selection bias approaches zero, and $\lim_{p\rightarrow 1}\Pr \left( y_{i}\leq q_{\tau }(x_{i})|p_{i}=p\right) =\tau $. By contrast, when $\pi _{i}$ is also a relevant omitted predictor, correlated with $p_{i},$ when the latter is close to one, then $ \lim_{p\rightarrow 1}\Pr \left( y_{i}\leq q_{\tau }(x_{i})|p_{i}=p\right) \neq \tau $ with positive probability.
The requirement that $p(z)=1$ for at least one value $z$ is typically labelled `identification at infinity' in the nonparametric identification literature C1986. In fact, note that `identification at infinity' may in principle hold when all elements of $x_{i}$ and $z_{i}$ are the same, see MR2008 for an application of `identification at infinity' to quantile models without an exclusion restriction. However, as simulation evidence suggested a rather poor test performance in this case we continue to maintain the assumption of a discrete or continuous exclusion restriction.
Since on the event $p_{i}=1$, the individual is selected into the sample with certainty, and so selection is not present, in a second step, we test the null hypothesis that the propensity score is an omitted predictor when $ p_{i}=1$.\footnote{ Note that the test has power when the omitted predictor is relevant also when $p$ is close to one.} That is, we test that:
versus
Thus, if selection is the sole cause for rejection of $H_{0,q}^{(1)}$, we do not expect to reject $H_{0,q}^{(2)}$ (at least asymptotically). By contrast, if we reject $H_{0,q}^{(2)}$, we take this as an indication that misspecification was likely to be the or at least one driver of the rejection at the first stage. Of course, as we discuss in the supplement in the context of the Decision Rule, we cannot rule out selection if both misspecification and selection are present and lead to a rejection simultaneously.
A common concern in the context of `identification at infinity' is so called irregular identification KT2010, where, although conditional quantiles are point identified, they cannot be estimated at a regular convergence rate as the marginal density of $p_{i}$ may not be bounded away from zero at the evaluation point $p(z)=1$. That is, heuristically, even if `identification at infinity' holds, and for some value $z\in R_{z}$, $p(z)$ can reach one, it is still possible that observations in the neighborhood of one are very sparse in practice (`thin density set'), and so convergence occurs at an irregular rate KT2010. To address this issue, we only use observations from parts of the support where the density of $p_{i}$ is bounded away from zero. Formally, this is implemented by introducing a trimming sequence, converging to zero at a sufficiently slow rate so that irregular identification is no longer a concern. Thus, let $\delta =1-H$ with $H\rightarrow 0$ and $H/h_{p}\rightarrow \infty $ as $n\rightarrow \infty $, where $H$ governs the speed of the trimming sequence $\delta $, while $h_{p}$ defines the window width around $\delta $. Then, reaclling the definition of $\widehat{u}_{\tau}(x_{i})=y_{i}-\widehat{q}_{\tau}(x_{i})$, the second test is based on the statistic
This statistic only uses observations with (estimated) propensity score $ \widehat{p}_{i}\in (1-h_{p}-H,1+h_{p}-H)$, and thus overcomes the issue of possible irregular identification as long as a sufficient number of observations are assumed to exist in this set (see below). Note here that the convergence speed of $H$ is inherently pegged to the tail behavior of the density of $p_{i}$ in the neighborhood of $p=1$, which is of course unknown in practice. That is, the thinner the density tail of $p_{i}$, the slower $H$ has to go to zero. We discuss below this issue and a potential data-driven way to select $H$ and $h_{p}$. Finally, as pointed out above, in order for the test to possess directed power against misspecification we require that relevant omitted predictor(s) $\pi _{i}$, if present, are correlated with the event $\{p_{i}=1\}$, or, more specifically $\{p_{i}\in (1-h_{p}-H,1+h_{p}-H)\}$.
In what follows, let $G_{u_{\tau}}(\tau ,1-H)\equiv \Pr (u_{\tau }(x_{i})\leq 0|p_{i}=1-H)$, and note that under $H_{0,q}^{(2)}$, it holds that $\lim_{H\rightarrow 0}G_{u_{\tau}}(\tau ,1-H)=\tau $ for every $\tau \in \mathcal{T}$, and that $\lim_{H\rightarrow 0}\Pr (s_{i}=1|p_{i}=1-H)=1$. We make the following additional assumptions:
{ A.6} There exists at least one $ z\in R_{z}$ such that $p(z)=1$. Moreover, there exists a strictly positive, continuous, and integrable function $g_{u_{\tau },p}(u_{\tau },1)$ and $ g_{p}(1)$ such that for all $\tau \in \mathcal{T}$:
as $n\rightarrow \infty $ for some $0\leq \eta <\overline{\eta }<1$, where $f_{u_{\tau},p}(\cdot,\cdot)$ and $f_{p}(\cdot)$ are the joint and marginal densities of $u_{\tau}$ and $p$, respectively.
{ A.7 The distribution function $ F_{u_{\tau}|p,s=1}(\cdot |\cdot ,\cdot )$ of $u_{\tau}(x_{i})$ given $p_{i}$ , and selection $s_{i}=1$ has a continuous probability density function $ f_{u_{\tau}|p,s=1}(u_{\tau}|p,s=1)$ w.r.t. Lebesgue measure for all $\tau\in \mathcal{T}$. The functions $f_{u_{\tau}|p,s=1}(u_{\tau}(x_{i})|p,s=1)$, $ \Pr(s_{i}=1|p) $, and $f_{p}(p)$ are continuously differentiable w.r.t. to $p $ for all $\tau\in\mathcal{T}$ with bounded partial derivatives. Moreover, assume that these functions are left-continuous at $p=1$. }
A.8 Assume that for all $\tau \in \mathcal{T}$, there exist positive constants $C$ and $C^{\prime}$ such that:
as well as
A.9 Assume that the support of $x_{i}$, $R_{x}$, is equal to $\mathcal{X}$, and that: \[ \sup_{(y,x)\in R_{y}\times\mathcal{X}}\left\lvert \frac{\partial f_{y|x,s=1}(y|x,s=1)}{\partial y}\right\rvert<\infty\quad\text{and}\quad \sup_{(y,x)\in R_{y}\times\mathcal{X}}\left\lvert \frac{\partial f_{y|x,s=1}(y|x,s=1)}{\partial x}\right\rvert<\infty. \]
Assumption A.6 requires identification at infinity, for at least one value $z$ of $z_{i}$, and this can be achieved provided some common covariate and/or the exclusion restriction are continuous.\footnote{Note that assumption A.6 could be relaxed to unbounded supports at the cost of more complex notation and proofs.} Importantly, our set-up does deal with the pratically relevant case of irregular identification, in the sense that $f_{p}(1)$ may not necessarily be bounded away from zero at $p=1$. More precisely, when $\eta =0 $, $\lim_{H\rightarrow 0}f_{p}(1-H)$ is bounded away from zero, while $\eta >0$ corresponds to the case of irregular support (with a larger value of $ \eta $ representing thinner tails). That is, if $\eta >0$, we allow for a thin set of observations with a propensity score close to one. Similarly, when $\eta =0$, the first part A.8 becomes a standard Lipschitz condition, while as $\eta $ gets closer to one and the tails of the densities in A.6 become thinner, we allow $G_{u_{\tau }}(\tau ,1-H)$ and $\Pr (s_{i}=1|1-H)$ to approach $G_{u_{\tau }}(\tau ,1)$ and 1, respectively, at a slower rate.\footnote{ Note that under $H_{0,q}^{(2)}$, we have that $G_{u_{\tau }}(\tau ,1)=\tau $ almost surely.} Finally, Assumption A.9 is a technical condition that ensures that the uniform local Bahadur representation continues to hold also at the boundary of the support of $x_{i}$. More specifically, together with \textbf{A.1}, \textbf{A.2}, and \textbf{A.5}, it allows the application of Proposition 4.4 in FG2016 and avoids the use of trimming, which would in turn require further steps in the proofs and more complex assumptions.
The rate of convergence of the numerator in ((ref)) depends on $H^{\eta }$ , the tail behavior of the density $f_{p}(p)$ around $p=1$, which is of course unknown in practice. In fact, we are generally ignorant about the rate of convergence given by $\sqrt{nh_{p}H^{\eta }}$, which may in principle be as fast as $\sqrt{nh_{p}}$. To address this problem, we use a studentized statistic, which allows the convergence rate to vary depending on the sparsity of observations around $p$ close to one. That is, as we cannot infer the appropriate scaling factor, it is crucial that, regardless the degree of thinness of the set of observations with propensity score close to one, $Z_{2,n,\delta}\left( \tau \right) $ and $\sqrt{\widehat{var} (Z_{2,n,\delta }\left( \tau \right) )}$ diverge at the same rate, so that the ratio still converges in distribution.
In the sequel, we study the asymptotic behavior of $\frac{Z_{2,n,\delta }^{q}\left( \tau \right) }{\sqrt{\widehat{var}(Z_{2,n,\delta }^{q}\left( \tau \right) )}}$ as an empirical process over $\tau \in \mathcal{T}$.
Theorem 2: Let Assumption A.1, A.3, A.5, A.6, A.7, \textbf{A.8}, \textbf{A.9}, and \textbf{A.Q} hold. If as $n\rightarrow \infty ,$ $(nh_{x}^{2d_{x}})/\log n\rightarrow \infty $, $nh_{x}^{2r} \rightarrow 0,$ $H\rightarrow 0,$ $H/h_{p}\rightarrow \infty $, $nh_{p}H^{2-\eta }\rightarrow 0$, and $nh_{p}H^{\eta }\rightarrow \infty ,$ then
(i) under $H_{0,q}^{(2)}$ ,
where $Z_{2}^{q}$ is the supremum of the absolute value of a zero mean Gaussian process with covariance kernel defined in the proof of Theorem 2.
(ii) under $H_{A,q}^{(2)}$ , there exists $\varepsilon >0,$ such that
Theorem 2 establishes the limiting distribution of the studentized statistic. As the theoretical results crucially hinge on the tuning parameters $H$ and $h_{p}$, whose rates depend in turn on the unknown $\eta $ , a discussion of their choice in practice is warranted. In fact, a possible data-driven choice of these parameters, without claiming optimality of a specific kind, could be as follows: as shown in the supplementary material, one may re-write $h_{p}$ and $H$, which is a function of $h_{p}$ itself, as functions of $\eta $ only, i.e. $h_{p}\left( \eta \right) =Cn^{-\frac{1+\varepsilon}{ 1 +\varepsilon+ \eta }}\log (n)$ and $H(\eta )=h_{p}(\eta )^{1/(1+\varepsilon )}$ for some arbitrary $\varepsilon >0$ and $ \eta <\overline{\eta }$, and some scaling constant $C$. Here, $\overline{ \eta }$ represents the threshold value with the slowest possible convergence rate still satisfying the rate conditions of Theorem 2. Thus, in order to `choose' the smallest possible $\eta $ in practice, which in turn corresponds to the fastest possible convergence rate, one could for instance plot $\frac{1}{\left( nh_{p}\left( \eta \right) \right) ^{1-\eta }} \sum_{i=1}^{n}K\left( \frac{\widehat{p}_{i}-(1-H(\eta ))}{h_{p}\left( \eta \right) }\right) $ for a given $\varepsilon >0$ (e.g., $\varepsilon =.1$) on a grid of different $\eta $ values with $\eta \in \lbrack 0,1)$, and select $ \hat{\eta}$ as the smallest value for which the estimated density is bounded away from zero, e.g. above a minimum threshold value such as $0.1$. In fact, with this procedure, if the set of $\widehat{p}_{i}$ close to $1$ is not `thin', we would expect to select $\hat{\eta}=0$ in large enough samples. We investigate this procedure further in Section (ref).
As outlined in the proof of Theorem 2, quantile estimation error vanishes. This is because under appropriate rate conditions it approaches zero at a rate which is faster than the convergence rate of the statistic$.$ Hence, when constructing the wild bootstrap statistic we do not have to `subtract' an estimator of the conditional distribution of $p_{i}.$ On the other hand, as the rate of convergence depends on the `degree' of irregular identification at $p$ close to 1, we also need an appropriately studentized bootstrap statistic, i.e.
where
with $B_{i,\tau }=1\left\{ U_{i}\leq \tau \right\} $ with $U_{i}\overset{ i.i.d.}{\sim }U(0,1)$ and independent of the sample, and
By noting that $\frac{1}{n}\sum_{i=1}^{n}(B_{i,\tau }-\tau )^{2}=\tau (1-\tau )+o_{p}^{\ast }(1),$ given ((ref)), we see that whenever `identification at infinity' holds and the number of observations with propensity score in the interval $(1-h_{p}-H,1+h_{p}-H)$ grows at rate $ nh_{p},$ then both numerator and denominator in ((ref)) are bounded in probability, otherwise they diverge at the same rate.
Let $c_{(1-\gamma ),n,R}^{\ast (2)}$ be the $(1-\gamma )$ percentile of the empirical distribution of $Z_{2,n}^{\ast q,1},...,Z_{2,n}^{\ast q,R},$where $ R$ is the number of bootstrap replications. The following Theorem establishes the first order validity of inference based on the bootstrap critical values, $c_{(1-\gamma ),n,R}^{\ast (2)}$.
Theorem 2$^{\ast }$: Let Assumption A.1, A.3, A.5, A.6, A.7, \textbf{A.8}, \textbf{A.9}, and \textbf{A.Q} hold. If as $n\rightarrow \infty ,$ $(nh_{1}^{2d_{x}})/\log n\rightarrow \infty $, $nh_{x}^{2r}\rightarrow 0,$ $H\rightarrow 0,$ $ H/h_{p}\rightarrow \infty $, $nh_{p}H^{2-\eta }\rightarrow 0$, $ nh_{p}H^{\eta }\rightarrow \infty$, and $R\rightarrow \infty$, then
(i) under $H_{0,q}^{(2)}$
(ii) under $H_{A,q}^{(2)}$
Theorem 2* establishes the first order validity of inference based on wild bootstrap critical values. Under $H_{0,q}^{(2)}$, the studentized statistic and its bootstrap counterpart have the same limiting distribution. Under, $H_{A,q}^{(2)},$ the statistic diverges, as the numerator is of larger probability order than the denominator, while the bootstrap statistic remains bounded in probability.
As a final remark, note that when the ultimate goal is to estimate the conditional quantile function(s) nonparametrically, both tests may be used in a testing procedure (see the supplement for a formal outline). That is, if we fail to reject the null hypothesis of the first test, one may interpret this as evidence against sample selection and decide to rely on nonparametric estimators of the conditional quantiles using all selected individuals in the data. On the other hand, if we reject the first test, but fail to reject the second one, estimation of the conditional quantile function(s) may still be carried out using only individuals with propensity score close to one, e.g. as in ((ref)) in the Appendix. By contrast, if the null hypotheses of both tests are rejected, there is evidence for relevant omitted predictor(s) (and possibly endogenous selection), and neither the estimator using all selected individuals nor the one using only those with propensity score close to one will deliver estimates consistent for the conditional quantile function(s) of interest. Of course, as pointed out before, the distinction between sample selection and relevant omitted predictors is only possible when the omitted predictors are correlated with $p_{i}$ when $p_{i}$ takes on values close to one. However, even when omitted predictors are present, but uncorrelated with the event $p_{i}$ close to one, one may make the `right' decision deciding in favor of selection, and by estimating the quantiles using only observations with propensity score close to one. This is so as for observations with propensity score close to one, omitted predictor bias is not present.
In this section we examine the finite sample properties of our tests via a Monte Carlo study. Results for the conditional mean can be found in the supplement. The outcome equation of our simulation design (given selection $ s_{i}=1$) is given by:
where $x_{i}\sim U(0,1)$, the distribution of $\widetilde{z}_{i}$ varies according to the design (see below), and the marginal distribution of $\varepsilon_{i}$ is standard normal. In the above equation, the parameter $\gamma_{1}$ determines the level of misspecification. Thus, when $\gamma_{1}$ is non-zero, $\widetilde{z}_{i}$ becomes an omitted relevant predictor as outlined in the previous section. On the other hand, when $\gamma_{1}$ is zero, the conditional quantile function $q_{\tau}(x_{i})$ is given by the third order polynomial function:
where $q_{\tau}(\varepsilon)$ denotes the unconditional $\tau$ quantile of $\varepsilon_{i}$. Selection enters into this set-up via:
with:
Thus, $\rho$ controls the degree of selection and we consider three scenarios, namely the case of `no selection' ($\rho=0$), `moderate selection' ($\rho=0.25$), and `strong selection' ($\rho=0.5$), while we set the scaling factor $\sigma$ of $v_{i}$ to one and $\gamma_{1}=0$ for most cases of the first test. The instrument $\widetilde{z}_{i}$ is simulated according to one of the following four designs:
Design (i) is the benchmark scenario and will illustrate the performance of the first test when the instrument exhibits continuous variation as for instance in the empirical illustration of Section (ref). Design (ii) on the other hand represents the opposite (extreme) case where the excluded instrumental variable $\widetilde{z}_{i}$ only takes on two values. Designs (iii) and (iv) are intermediate scenarios where $\widetilde{z}_{i}$ follows a (discrete) Poisson distribution with an average of around 7 support points, and a discrete uniform distribution with 7 support points (and equally distributed point mass).\footnote{ The relationship between the `signal' variance of $0.75(x_{i}-0.5)+0.75 \widetilde{z}_{i} $ and the `noise' variance of $v_{i}$ therefore ranges from around 0.2 to 0.6 across cases.} Importantly, the last two distributions are meant to be stylized examples of set-ups with discrete instruments such as `number of kids' or `month-observed-in-the-sample'.
We consider three quantiles $\mathcal{T}=\{0.3,0.5,07\}$ and two sample sizes $n=\{1000 , 2000\}$, which, given a selection probability of approximately $0.5$, imply an effective sample size for the outcome equation of around 500 to 1000 observations, respectively. Throughout, we estimate $ q_{\tau}(x_{i})$ using the selected sample and a third order local polynomial estimator in line with ((ref)) and the conditions of Theorem 1.\footnote{ In order to restrict ourselves to a compact subset $\mathcal{X}$, we trim the outer 2.5% observations of the selected sample.} In total, we consider six different cases: since our theoretical results suggest that estimation error from the propensity score does not feature in the asymptotic representation of ((ref)) under $H_{0,q}^{(1)}$, in the first four cases, Cases I-IV, we consider designs (i) to (iv) using the oracle propensity score. For these cases, we choose the bandwidths $h_{x}$ and $h_{F}$ ad-hoc in line with Theorem 1 and 1* as $h_{x}=c\cdot \text{sd}(x_{i})\cdot n^{-\frac{1}{3}}$, $ c=\{3.5,4,4.5\}$, and $h_{F}=2.2\cdot \text{sd}(x_{i},\widehat{u} _{\tau}(x_{i}))\cdot n^{-\frac{1}{6}}$, respectively, where $\text{sd}(\cdot)$ denotes the standard deviation. In Case V, we still consider the oracle propensity score using design (i), but instead choose $ h_{F}$ and $h_{x}$ via cross-validation ($\widehat{h}_{x}$). More specifically, for $\widehat{h}_{x}$, we employ the method described after Theorem 1 and pick the bandwidth running cross-validation for a lower local polynomial order (i.e., $\widetilde{r}=1$) than the one assumed. In Case VI, we replace the oracle propensity score from Case V by an estimate using a standard local constant estimator with second order Epanechnikov kernel and a cross-validated bandwidth as permitted by our theory when the number of continuous variables in the selection equation is less than four (cf. footnote 6).\footnote{ To construct this estimator as well as the estimators for the conditional mean in the supplement, we use routines from the np package of HR2008. This package allows to construct the bandwidth according the cross-validation procedure outlined in Section 2 of LR2008. For the local polynomial quantile estimator, we use a routine from the quantreg package K2017.}
Finally, in Case VII and VIII we examine the power of the first test under no selection, but misspecification. More specifically, we set $\rho=0$ and vary $ \gamma_{1} $, the misspecification parameter, for the design of Case I in Case VII and of Case II in Case VIII with cross-validation, respectively. Throughout, we use the `warp speed' procedure of GPW2013 with 999 Monte Carlo replications and fix the nominal level of the test to $ \alpha_{1}=0.05$ and $\alpha_{2}=0.1$. The results of the first test can be found in Table (ref) below.
Turning to the results, observe that under $H_{0,q}^{(1)}$ (i.e., $\rho=0$), we have overall a good size control and rejection rates converge to the nominal levels as the sample size increases. Interestingly, the performance of test in Cases III and IV with a discrete multi-valued instrument is slightly worse in terms of size than for the binary instrument of Case II. Moreover, we can see that the Cases V and VI, where the bandwidths are chosen in an automated manner, are in line with the rest, suggesting that the data-driven procedures to pick the tuning parameters may be a good alternative for the application of the test in practice. Turning to power ($ \rho=0.25$ and $\rho=0.5$), note that power picks up rather quickly with the sample size when sample selection is `moderate' ($\rho=0.25$) and is close to one when $\rho=0.5$ throughout. As expected, we observe that power is generally lower when discrete instruments are used, although the gap between Case II (binary instrument) and Case I (continuous instrument) is never above 10 percentage points when $n=2,000$. Finally, note that when we turn to the misspecification in Cases VII and VIII, we observe that even at $ \gamma_{1}=0.25$ the rejection rate of the first test is immediately very high and very close to one suggesting that misspecification may be a concern when potentially relevant predictors have been omitted.
For the second test, we consider the case of misspecification by manipulating $\gamma_{1}$ while operating under the alternative of the first test. More specifically, we set $\rho=0.25$ throughout and examine three different values of $\gamma_{1}$ when $\widetilde{z}_{i}$ is normally distributed (CASE I $^{\ast}$-III$^{\ast}$) with $\widetilde{z}_{i}\sim N(0,0.5)$, and three alternative cases with $\widetilde{z}_{i}\sim Binom(0.5)-0.5$ (Cases IV$^{\ast}$-VI$^{\ast}$), respectively. For each distribution of $\widetilde{z}_{i}$, we start with the case under the null of the second test ($\gamma_{1}=0$), and then consider cases under the alternative with `moderate' misspecification ($\gamma_{1}=0.25$) and `strong' misspecification ($\gamma_{1}=0.5$), respectively. In the case where $\widetilde{z}_{i}\sim N(0,0.5)$ we choose $h_{p}\in \{0.075,0.05,0.03,0.02\}$ and $\delta\in\{0.95,0.975,0.98\}$. On the other hand, for the binomial variable, we set $h_{p}\in\{0.075,0.05\}$ and $ \delta\in\{0.95,0.975\}$ reflecting the fact that the discrete nature of $ \widetilde{z}_{i}$ in the Cases IV$^{\ast}$-VI$^{\ast}$ leads to less observations with a propensity score value close to one.\footnote{ We also set the scaling factor $\sigma$ to $0.5$ in this case.} Finally, to analyze the performance under the data-driven procedure outlined after Theorem 2, we added to each case the performance when $\widehat{h}_{x}$ and $\widehat{h}_{p}$ are chosen in a data-driven manner. More specifically, $\widehat{h}_{x}$ is chosen via cross-validation as outlined for the first test, while for $\widehat{h}_{p}$ we set $h_{p}(\eta)=\log(n)n^{-\frac{1+\varepsilon}{1+\varepsilon+\eta}}$ and $H(\eta)=h_{p}^{1/(1+\varepsilon)}$ for $\varepsilon=0.1$ and select the smallest possible $\eta$ from a grid $\{0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9\}$ such that $\frac{1}{\left( nh_{p}\left( \eta \right) \right) ^{1-\eta }} \sum_{i=1}^{n}K\left( \frac{p_{i}-(1-H(\eta ))}{h_{p}\left( \eta \right) }\right) >0.1$. All results can be found in Table (ref).
Starting with the size results, observe that in Case I$^{\ast}$ with $ z_{i}\sim N(0,0.5)$ the empirical rejection rate of the test appears to be somewhat sensitive to the choices of $h_{p}$ and $\delta$, respectively. More specifically, note that when $h_{p}=0.075$ and $\delta=.95$ the test over-sizes substantially at $n=2,000$. In fact, as expected by the theoretical results, this problem is alleviated when $\delta$ is chosen closer to one and $h_{p}$ smaller as the sample size increases, see for instance $\delta=0.98$ and $h_{p}=0.02$. By contrast, when $\widetilde{z}_{i}\sim Binom(0.5)-0.5$, the size results become generally very conservative and are close to zero. Turning to power, we can observe that power is generally very good when $\widetilde{z}_{i}$ is normally distributed, but rather poor when $ \gamma_{1}=0.25$ and $\widetilde{z}_{i}$ follows the binomial distribution. This is of course to be expected and suggests that when only elements in $x_{i}$ are continuous, larger sample sizes might be required for good power results. Finally, observe that the data-driven choice of $h_{p}$ and $h_{x}$ delivers results that are largely of a similar order of magnitude in terms of size and power as the results with a fixed bandwidth, suggesting that this choice might also be an option in practice.
Summarizing this section, we obtain a somewhat mixed picture of our tests in finite samples. While the first test performs well throughout the designs even when the instrument is discrete and the tuning parameters are chosen in a completely automated manner, the second test appears, as expected, to be somewhat more sensitive to the choice of the tuning parameters $\delta$ and $h_{p}$ as well as to the continuity (or discreteness) of $\widetilde{z}_{i}$, which seems to reflect the irregularity of the underlying problem.
Our illustration is based on a subsample of the UK wage data from the Family Expenditure Survey used by AB2017.\footnote{ For the exact construction of the sample see their paper and references therein.} As pointed out by these authors, due to changes in employment rates over time, simply examining wage inequality for females and males at work over time may provide a distorted picture of market-level wage inequality. We will therefore run our selection testing procedure on two different subsets of the data, namely 1995 to 1997, a period of increasing gross domestic product (GDP) growth rates, and 1998 to 2000, a period of high, but stable GDP growth rates. Unlike AB2017, however, our testing procedure for selection will not rely on a parametric specification of the conditional log-wage quantile functions, but remain completely nonparametric.
The covariates we include in $x_{i}$ are dummies for marital status, education (end of schooling at 17 or 18, and end of schooling after 18), location (eleven regional dummies), number of kids (split by six age categories), time (year dummies), as well as age in years. This set of covariates is identical to the one used by AB2017, but for the fact that the latter used cohort dummies instead of age in years. The continuous instrumental variable is given by the measure of potential out-of-work (welfare) income, interacted with marital status. This variable, which was also used by AB2017, builds on BRS2003 and is constructed for each individual in the sample (employed and non-employed) using the Institute of Fiscal Studies (IFS) tax and welfare-benefit simulation model.
The final sample for the years 1995-1997 comprises 21,263 individuals, 11,647 of which are females and 9,616 of which are males, respectively. The number of working females (males) with a positive log hourly wage in that sample is 7,761 (7,623). By contrast, for the 1998-2000 period we obtain 16,350 observations, 8,904 females and 7,446 males. The number of working females (males) in that sample are 5,931 (6,058).
All estimates are constructed using routines from the np package of HR2008. That is, we estimate the propensity score $ \Pr(s_{i}=1|z_{i})$ fully nonparametrically using a standard local constant estimator with an Epanechnikov second order kernel function for the continuous instrument and discrete kernel functions as in Equations (1) and (2) of LR2008 for the remaining discrete controls. The bandwidth is determined in a data-driven manner using cross-validation according to the procedure in Section 2 of LR2008 as permitted by A.3 since $d_{zc}<4$ (cf. footnote 6). More specifically, to reduce computation complexity, we follow the method outlined in R1993 conducting cross-validation on random subsets of the data (size $n=450$), and select the median values over 50 replications.
The conditional quantile function $q_{\tau}(x_{i})$ on the other hand is estimated as in Equation (19) of LR2008. Note that in this case, all covariates in $x_{i}$ are discrete, and the rate conditions of Theorem 1 and 1$^{\ast}$ do therefore not apply directly as discrete predictors do not contribute to the asymptotic variance. As a result, we choose to select the bandwidth parameters for the discrete kernel functions according to the same method outlined in R1993 using cross-validation for discrete covariates as automated by the np package. The same holds true for the conditional distribution function $F_{p|x,u_{\tau},s=1}(\cdot|\cdot,\cdot,\cdot)$, which is constructed as in Equation (4) of the same paper. Finally, the quantile grid is chosen to be $\mathcal{T}=\{.1,.2,.3,.4,.5,.6,.7,.8,.9\}$.
To provide the reader with a better illustration of the potential magnitudes of selection into work, we replicate the predictions from the (estimated) parametric conditional quantile functions from Figure 1 of AB2017 for the sub-periods 1995-1997 and 1998-2000, see Figures (ref) and (ref), respectively. In these pictures, solid lines represent estimated uncorrected (for selection) conditional log-wage quantile functions, while dashed lines are the ones corrected for sample selection.\footnote{ For the exact specification used, see AB2017.} Throughout, female quantile lines lie below the male quantile lines. The figures, which display selection corrections based on a linear quantile regression model and parametric selection correction, show little difference between the original and the corrected lines for both males and females for the 1995-1997 period (except for males at lower percentile levels), but more pronounced differences for the subsequent 1998-2000 period. In terms of magnitude, the effect of correction appears to be generally bigger for males than for females in both subperiods.
Turning to the test results in Table (ref), we see that while we cannot find any evidence of selection during the 1995-1997 period for females at conventional significance levels, there is some evidence for males at the 10% significance level. In fact, taking a closer look at the results in Table (ref), we observe that rejection for males occurs on the basis of the 10th percentile, which is in line with the graphical evidence in Figure (ref). Switching over to the right panel of Table (ref), however, we obtain a different picture: for females, $ H_{0,q}^{(1)}$ is rejected at any conventional level, and rejection is most pronounced at the 20th and the 30th percentile. On the other hand, we cannot reject $H_{0,q}^{(1)}$ for males. This failure to reject $H_{0,q}^{(1)}$ for males is in contrast to the graphical evidence in Figure (ref), and highlights the importance of formal testing under a more flexible specification.
To determine whether misspecification may be the cause of rejection, we perform the second test for males in the 1995-1997 period, and for females in the 1998-2000 period (see Table (ref)). Turning to the results, we strongly reject the null of no misspecification of the conditional quantile function $q_{\tau}(x_{i})$ for males, but fail to reject that null for females at any conventional levels. Both results appear to be robust to different choices of $\delta$ and $h_{p}$. Thus, under the assumption that out-of-work income is a valid instrument and indeed selection enters outcome as postulated in Equation ((ref)), our test results suggest that there is evidence for selection among females for the 1998-2000, but not for males. In fact, what appears to be selection among males during the 1995-1997 period could actually be attributed to misspecification of the conditional quantile function.\footnote{ In a related paper, K2010 tested for the validity of the same instrumental variable (but for the interaction with marital status) in a similar data set on the basis of the UK Family Expenditure Survey used by BGIM2007. Although his test results are not directly informative here as his test is run on a much coarser set of covariates $x_{i}$ not including e.g. regional, martial, or family information, his evidence suggested that the conditional independence of the instrument and outcome (given $x_{i}$ and selection $s_{i}=1$) may indeed be violated for some sub-groups (in particular, younger males with moderate levels of education). Thus, rejection in the second test for males could indeed be related to this feature.}
This paper introduces two tests to detect sample selection in conditional quantile functions, without imposing parametric assumptions on either the outcome or the selection equation. The first test is an omitted predictor test, with the estimated propensity score as omitted predictor. This test may be of particular interest to practitioners who rely on the three step estimator of AB2017, but can check for the presence of sample selection without relying on functional form assumptions. Another feature of this test is that it may be carried out selecting the bandwidth parameters in a data-driven manner.
As with any omnibus test, rejection in the first test can be due to either selection or to the omission of a predictor which is correlated with the estimated propensity score. Since selection and misspecification have very different implications for the estimation of nonparametric (conditional) quantile functions, we aim at disentangling the two if we reject in the first step. That is, after rejection in the first test we proceed to the second test using only observations with (estimated) propensity score close to one. A rejection in this case indicates the presence of misspecification, possibly in conjunction with selection.
Importantly, neither of our tests rely on a continuous exclusion restriction and may be executed over a (compact) subset of quantile ranks. Indeed, Monte Carlo evidence suggests that in particular the first test has good finite sample properties when the exclusion restriction is discrete, which renders it attractive for applied work. In Appendix B, we also develop the first test for nonparametric conditional mean functions. Simulations confirm the favorable performance in this case, too.
In our empirical illustration, we test for sample selection in log hourly wages of females and males in the UK using data from the UK Family Expenditure Survey. We find evidence for selection among females for the 1998-2000, but not for males. In fact, what appears to be selection among males during the 1995-1997 period may actually be attributed to misspecification of the conditional quantile function.
\singlespacing