Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
68,702 characters · 16 sections · 57 citation commands
Nonparametric Regression with Selectively Missing Covariates
{
}
Sample selection is a central challenge for empirical evaluation studies. Nonrandom selection can affect the empirical analysis in many ways, for example through nonrandom selection into treatment programs, selective measurement error or through selective nonresponse or missingness of data. In this paper, we propose an instrumental variable approach to address the problem of nonrandom selection. We provide constructive nonparametric identification results and build on those to establish a novel fractional probability weighting estimator.
While the methodology is general and applicable to many situations in which selection might be problematic the leading example in this paper will be selective nonresponse and selective missing data. We are interested in the identification and estimation of the nonparametric regression function $g(x)={\mathbb E}[Y|X^*=x]$ where $Y$ is always observed but $X^*$ is only selectively observed. In this case, parts of the information of the covariates are missing not at random for some sampling units. Without accounting for selectivity of responses, statements about individual behavior based on such incomplete data might be severely biased.
In this paper, we establish identification of the nonparametric regression function $g(x)={\mathbb E}[Y|X^*=x]$ based on instrumental variables that explain variation in the latent covariates but have no direct effect on selection. Such an instrumental variable approach is well suited when selection is driven by the latent variables $X^*$. We show that the regression function $g$ can be written as a weighted version of its observed counterpart. The weighting function is determined by a fraction of selection probabilities, i.e., fractional probability weights (FPW), that depends on latent variables. We propose a novel identification restriction, the so called partial completeness assumption, which implies identification of the FPW and thus of the nonparametric regression function $g$. In contrast to usual completeness assumptions required for identification of nonparametric instrumental variable models, we are able to provide primitive, functional form conditions for the partial completeness assumption to hold. We emphasize that these functional form conditions do not imply (semi-)parametric restrictions but only impose a nonparametric structure in different forms of separability of $Y$ and $X^*$. Specifically, we show that a nonparametric generalized additive structure of the selection probability is sufficient to obtain identification of the function $g$.
Based on the constructive nonparametric identification result we propose a novel nonparametric FPW series estimator that is convenient for implementation. We show that our estimator has a rate of convergence that coincides with usual nonparametric regression estimators, i.e., the asymptotic performance of the estimator is the same as of an estimator with full information of the underlying selection mechanism. We establish asymptotic normality of the estimator and show that the asymptotic variance is not necessarily enlarged by FPW estimation. We also propose a bootstrap procedure to construct uniform confidence bands. A Monte Carlo simulation study demonstrates the improvements of our approach over missing at random (MAR) estimators. In particular, we highlight our contribution also in a finite sample analysis of linear regression with alternative inverse probability weighting estimators (IPW).
Finally, we use the method in two different empirical applications. Both applications are important for the discussion about income inequality and highly relevant for public policy. First, we use the developed methodology to analyze the association between income and the risk of bad health. The empirical analysis is based on linked data from the Survey of Health, Aging and Retirement in Europe (SHARE). We exploit a specific feature of the data which allows us to link a sub-sample of the survey data to administrative data of the German pension insurance.\footnote{BinMar_17 compare self reported information about income and education from SHARE with matched information from Danish administrative data in order to study the implications of measurement error.} We find that income in the SHARE data is selectively missing and we demonstrate that standard methods that do not account for the nonrandom selection process are biased, specifically for individuals with low incomes. Assuming linearity of the regression function $g$ the point estimate of income is significantly negative when imposing MAR however it is not significantly different from zero when accounting for nonrandom nonresponse. In the second example we analyze how housing varies with income. We quantify the relationship between labor earnings and the probability to own a house. For this empirical analysis we use data from the German Socio Economic Panel (SOEP) and exploit information about the regional purchasing power collected on the residential block level (Sub-Zip code level) as an instrument. Again we find that earnings are selectively missing and we demonstrate that the estimates derived in standard methods are biased for individuals with earnings in the lowest decile which is a specifically relevant group for public policy. For individuals with higher earnings the estimates from the different methods do not differ significantly.
Our paper is linked to several strands of the literature. The most common way to deal with missing data is to assume missing at random pioneered by little2002. In the context of selectively missing covariates, a sieve semiparametric maximum likelihood estimator was proposed by chen2007. In contrast, an instrumental variable strategy, as proposed in the paper, was used so far only to deal with endogenous missingness of dependent variables, see, for instance, tang2003, ramalho2013, and 2010Hault. Also breunig2015 consider the problem of nonparametric regression with selective nonresponse of the dependent variables; zhao2015 focus on a semiparametric approach. There only has been minor attention to selectively observed covariates. One example is zhaostatistica who consider a semiparametric approach to deal with selectively missing covariates that is crucially different from ours. While zhaostatistica require a parametric specification of the distribution of outcome given potential covariates, we leave these conditional distribution unrestricted. We establish nonparametric identification of the regression function and hence ensure that the identification is not due to specific functional form restrictions that might be violated in practice.
The paper adds as well to the literature on income inequality, more specifically to studies on the income gradient on health outcomes and mortality e.g. Pre_1975, DeatPax_98, CutDeaLLe_06, or CutLleVog_11, and on the effect of income on housing and on housing demand, see e.g. QuiRap2004, Albouyetal2016 or DusFitZim2018. In general these studies are based on survey data in which wealth, income, health and housing information and further demographic variables are self reported. As shown in breunig2017 information on income or earnings in surveys is likely to suffer from nonrandom selection which might result in biased estimates of the association between income and health or income and home ownership. In this respect this study extends the previous literature as we account for nonrandom nonresponse of the income information.
The paper is organized as follows. In Section (ref) we establish identification of our nonparametric model. In Section (ref) we derive the FPW series estimator, establish its rate of convergence and its asymptotic normality, and derive uniform bootstrap confidence bands. Section (ref) provides Monte Carlo simulations and discusses implications of FPW estimation to unconditional moments. In Section (ref) we apply our methodology to the two different empirical applications. All proofs can be found in Appendix (ref). Appendix (ref) provides some technical results. Finally, Appendix (ref) provides an extension when also the dependent variable is selectively missing.
This section consists of two subsections. In Subsection (ref), we provide assumptions required for identification. In particular, we introduce a novel restriction, i.e., the partial completeness assumption, and provide primitive conditions for it. Subsection (ref) establishes identification of the nonparametric regression function.
Given an observable outcome variable $Y$ and latent covariates $X^*$ our interest lies in the regression function $g(x)={\mathbb E}[Y|X^*=x]$. Identification relies on instrumental variables $W$ that explain variations of the latent variable $X^*$ but are not directly related to the selection mechanism $D$. This is formalized in the following. Throughout the paper, we assume that a sample $(D_1,Y_1,W_1),\dots,(D_n,Y_n,W_n)$ of $(D,Y,W)$ is observed for each individual. A $d_x$-- dimensional vector of covariates $X^*$ is only fully observed depending on a binary indicator variable $D$, i.e., $X^*$ is observed when $D=1$ and missing when $D=0$. We write $X=D X^*$.\footnote{The situation can be easily extended to a multivariate version where $D$ denotes a $d_x$--dimensional vector of missing data indicators. In order to keep the notation simple we do not treat this case explicitly.} Under the assumptions presented below we see that the selection probability conditional on $(Y,X^*)$, i.e., $\mathbb{P}(D=1|Y,X^*)$, is only partially identified but still point identification of the regression function $g$ is established.
Assumption (ref) states an exclusion restriction of the random vector $W$. It excludes any relation between $W$ and the selection mechanism $D$ that is not channeled through $(Y,X^*)$. The setting corresponds to the measurement error set up, where instrumental variables are required to drive the latent, true variable but not the variable that is observed with error. However, identification with nonclassical measurement error requires an additional exclusion restriction which restricts $W$ to have no information on $Y$ that is not captured in $X^*$, see Assumption 2 $(i)$ in hu2008. Interestingly, nonrandom selection as extreme form of nonclassical measurement error simplifies the exclusion restriction imposed on the instruments.
We also emphasize that Assumption (ref) allows for dependence of $D$ and $Y$. Thus, our approach captures selection on unobservables that do not only stem from latent characteristics in $X^*$ but also from unobservables that are unexplained by the regression function $g(X^*)$.\footnote{In the model $Y=g(X^*)+U$, not only $X^*$ but also unobservables $U$ are allowed to directly affect the selection mechanism $D$.} This is an important feature of our framework, as in many economic environments, selection variables can be driven by unobserved individual characteristics. Related literature on nonrandom nonresponse of covariates does not allow for such a general selection mechanism, see zhao2015. We introduce the function class $\mathcal B =\{\phi:\, {\mathbb E}|\phi(Y,X^*)|<\infty \text{ and }\inf_{y,x}\phi(y,x)>0\}$.
Assumption (ref) is less restrictive than the usual completeness assumption which assumes that ${\mathbb E}[\phi(Y,X^*)|Y,W]=0$ implies $\phi(Y,X^*)=0$. This assumption is commonly imposed to ensure identification in nonparametric instrumental variable models, see for instance NP03econometrica. In the context of endogenous selection such completeness assumptions were considered by 2010Hault and breunig2015. On the other hand, the partial completeness assumption holds under mild functional form assumptions as shown below.
Assumption (ref) is automatically satisfied if $\phi$ does not depend on $Y$. Indeed, if the selection probability $\mathbb{P}(D=1|Y,X^*)$ does not depend on $Y$ the regression function $g$ is identified as we see in the next subsection and thus, the partial completeness assumption is well suited for our particular selection problem. Moreover, the next result provides functional form restriction under which partial completeness holds. Throughout the paper, $f_V$ denotes the probability density function of a random variable $V$.
Proposition (ref) requires that $Y$ does not provide information on $X^*$ that is not contained in the vector $W$. Given this mild restriction we see from Proposition (ref) that functional form restrictions imply the partial completeness assumption to hold. We emphasize that these functional form restrictions do not imply (semi-)parametric specifications but only impose a nonparametric structure in different forms of separability of $Y$ and $X^*$. Note that we subtract by one as the exclusion restriction in Assumption (ref) implies the conditional mean restriction ${\mathbb E}[D/\mathbb{P}(D=1|Y,X^*)-1|Y,W]=0$. The selection probability $\mathbb{P}(D=1|Y,X^*)$ is not necessarily point identified through the former conditional mean restriction given Assumption (ref). Equation (ref) provides a generalized additive type restriction on the functions of interest. The restriction ${\mathbb E}[\psi^{(j)}\left(\phi_1(y)+\phi_2(X^*)\right)|W]\neq 0$, for some integer $j\geq 1$ and all $y$, is satisfied by a broad class of link functions $\psi$. This condition holds, in particular, when the first derivative $\psi'\left(\phi_1(y)+\phi_2(X^*)\right)$ is strictly positive for all $y$ in the support of $Y$. Consequently, we conclude from Proposition (ref) that partial completeness holds under a generalized additive restriction when the link function $\psi$ coincides with any analytic, cumulative distribution function supported on the whole real line. Another case where ${\mathbb E}[\psi^{(j)}\left(\phi_1(y)+\phi_2(X^*)\right)|W]\neq 0$, for some $j\geq 1$, holds is when $\psi$ is a non-constant polynomial function.
While Proposition (ref) provides a broad class of functions which satisfy the partial completeness assumption, partial completeness can fail for functions which vary in $y$ with vanishing means and when instruments are independent of $X^*$. This is demonstrated in the following illustrative example.
Assumption (ref) can rule out a selection when it is a deterministic function of $Y$ and $X^*$, such as certain indicator functions. We also emphasize that Assumption (ref) can be relaxed if we are only interested in point evaluation at some $x_0$ in the support of $X$. That is, identification of ${\mathbb E}[Y|X^*=x_0]$ requires only $\mathbb{P}(D=1|Y=y,X^*=x_0)>0$ uniformly over all $y$ in the support of $Y$.
In this subsection, we establish identification of the nonparametric regression function $g(x)={\mathbb E}[Y|X^*=x]$. We show that the function $g$ can be identified via a fractional probability weight (FPW). In addition, we show that the FPW is identified by making use of instrumental variables $W$ which satisfy the previous assumptions. In the next result, we document that the regression function $g$ can be written as
where the fractional probability weight (FPW) function $\omega$ is given by
Further, we establish identification of the nonparametric regression function $g$.
The previous result shows that the FPW function given in (ref) is point identified although the selection probabilities conditional on latent variables are only partially identified. This is an implication of partial completeness imposed in Assumption (ref). Corollary (ref) presents a useful property of the FPW function $\omega$. This result is an immediate consequence of the proof of Theorem (ref) and hence we omit its proof.
In empirical applications also the dependent variable might be selectively missing. We discuss this case in Appendix (ref). The remark below highlights the difference between selectively missing covariates which requires FPW and selectively missing dependent variables which requires inverse probability weighting.
While Remark (ref) highlights the difference between fractional and inverse probability weighting, we note that for unconditional moment estimation both approaches are feasible (see Section (ref)). Still our FPW approach is preferable in this case due to a less restrictive identification requirement and finite sample improvements as shown in Section (ref).
This section consists of four subsections. In Subsection (ref), we derive the FPW series estimator which stems from our constructive identification result. Subsection (ref) provides the rate of convergence of the estimator. We establish pointwise asymptotic normality of the FPW series estimator in Subsection (ref) and provide asymptotic validity of uniform bootstrap confidence bands in Subsection (ref).
We define the conditional selection probability by
In particular, the selection probability conditional on the latent regressors $X^*$ is determined by
see the proof of Theorem (ref). We thus obtain the following expression for the FPW function $\omega$:
We estimate the regression function $g(x)={\mathbb E}[Y|X^*=x]={\mathbb E}[Y\omega(Y,x)|D=1, X^*=x]$ using a plug-in series least squares estimator. To do so, we introduce a vector of basis functions $p^K(\cdot)=\big(p_1(\cdot)\dots,p_K(\cdot)\big)'$; $K=K(n)$ is an integer which increases with the sample size $n$. We further introduce the $n\times K$--matrix $\mathbf X=\big(D_1 p^K(X_1),\dots,D_n p^K(X_n)\big)$. We estimate the FPW function $\omega$ via
It is common in the context of inverse probability weighting, to normalize the weights to sum up to one. In our context, we normalize $\omega$ as follows. Employing Corollary (ref), i.e., ${\mathbb E}[\omega(Y, x)|D=1,X^*=x]=1$, we obtain
Replacing ${\mathbb E}[D\, p^K(X)\, \omega(Y,X)\, p^K(X)']$ by the empirical matrix $\mathbf X'\, \boldsymbol \omega(\widehat \varphi)\,\mathbf X$ we obtain the FPW series estimator of the regression function $g$ given by
where $\boldsymbol \omega(\phi)=\text{diag}\big(\widehat \omega(Y_1,X_1;\phi),\dots,\widehat \omega(Y_n,X_n;\phi)\big)$. Here, $\widehat\varphi$ is a restricted sieve minimum distance estimator of the selection probability $\varphi$ given as follows. We have the conditional moment restriction induced by the exclusion restriction imposed in Assumption (ref), that is,
(Here, we use that $\varphi(Y,X^*)=\varphi(Y,X)$ whenever $D=1$.) Consider a vector of tensor product basis functions $q^L(\cdot,\cdot)=\big(q_1(\cdot,\cdot)\dots,q_L(\cdot,\cdot)\big)'$ used to approximate the conditional mean in (ref); $L=L(n)$ is an integer which increases with the sample size $n$. We estimate the conditional mean $m(y,w; \phi)={\mathbb E}[D/\phi(Y,X)-1|Y=y,W=w]$ by the series least squares estimator
then we the constrained sieve minimum distance estimator
where $\mathcal B_L=\{\phi=1/(\beta'q^L):\, \min_{y,x}\beta'q^L(y,x)\geq 1\}$ following chenpouzo2012. We may assume that $\mathcal B_L$ becomes dense in $\mathcal B$ as $L$ tends to infinity. Note that in the case of nonparametric estimation, any estimator of $\varphi$ has a slow rate of convergence since the conditional mean restriction yields in general to a so called ill-posed inverse problem, see NP03econometrica and BCK07econometrica. In our case, $\varphi$ is not identified through the conditional mean equation (ref) but we can always ensure uniqueness of the estimator, for instance, by considering the minimal norm estimator of equation (ref). Finally, note that FPW series estimation is convenient since control variables which enter the model linearly can be simply included in the empirical matrix $\mathbf X$. This allows to treat partially linear models as considered in our empirical applications in Section (ref).
We now introduce some assumptions. The support of $X$ is denoted by $\mathcal X$. We also introduce the $L^2_X$--norm $\|\phi\|_X=\sqrt{{\mathbb E}\phi^2(X)}$ and $\|\cdot\|$ denotes the Euclidean norm. We make use of the notation $U=Y-g(X^*)$ and $h(x,\phi)={\mathbb E} \left[1/\phi(V)|D=1, X^*=x\right]$. Recall that $\omega(V,\varphi)=(\varphi(V)h(X,\varphi))^{-1}$ is identified due to Theorem (ref). We introduce function class
Note that $h(\cdot,\varphi)\in{\cal H}$ due to Assumption (ref). In order to achieve the rate of convergence, we have to restrict the complexity of the function classes $\mathcal B$ and $\mathcal H$, by imposing a finite entropy integral. We first provide primitive conditions by imposing smoothness restrictions on the function classes under consideration.
For any vector $a=(a_1,\dots, a_d)$ of $d$ integers, where $d=\dim(X,Y)$, define $D^a = \partial^{|a|}/\partial a_{a_1}\dots \partial x_{a_d}$, where $|a|=\sum_{l=1}^d a_l$. Let $\mathcal R_d$ be a bounded, convex subset of $\mathbb R^d$ with nonempty interior. For a function $\phi: \mathcal R_d\to\mathbb R$ and some $\alpha>0$, let $\underline\alpha$ be the largest integer smaller than $\alpha$, and
We further restrict the function class $\mathcal B$ in the next assumption to contain only functions with bounded $\|\cdot\|_{\infty,\alpha}$ norm and bounded supremum norm $\|\cdot\|_\infty$. To do so, we introduce
for some $\alpha>0$ and constant $C>0$. Similarly for $\mathcal H$ we introduce a smooth function class
for some $\alpha>0$ and constant $C>0$, which is a subset of $\mathcal H$ for any constant $C>0$.
Assumption (ref) $(ii)-(iii)$ restricts the magnitude of the approximating functions $\{p_j\}_{j\geq1}$ and imposes nonsingularity of their second moment matrix. Assumption (ref) $(ii)$ holds for instance for polynomial splines, Fourier series and wavelet bases. Assumption (ref) $(iii)$ is satisfied if $p^K$ is a vector of orthonormal basis functions and the probability density function of $X^*$ given $D=1$ is uniformly bounded away from zero on its support. Assumption (ref) $(iv)$ determines the sieve approximation error which in turn characterizes the bias of the estimated regression function $g$; see also chen2007 for further discussions on sieve bases. Assumption (ref) $(v)$ imposes smoothness conditions on the function classes $\mathcal B$ and ${\cal H}$ in order to reduce the complexity of these classes.
The next result establishes the rate of convergence for the FPW series estimator $\widehat g$.
From Theorem (ref) we observe that the estimator $\widehat g$ attains the usual bias and variance term in integrated mean square error for nonparametric series regression. If $K$ is chosen to level variance and bias, i.e., $K\sim n^{d_x/(2\alpha+d_x)}$, then the convergence rate given in Theorem (ref) coincides with $n^{-2\alpha/(2\alpha+d_x)}$. Consequently, we obtain the optimal nonparametric rate of convergence as in the situation where the covariates $X$ are completely observed, that is, we do not obtain a slower rate of convergence due to the estimation of the FPW function. This result is obtained by the regularity conditions imposed in Assumption (ref) $ (v)$, which restrict the complexity of the underlying fractional probability function $\omega$.
This subsection discusses the inference of the estimator of regression function $g$ evaluated at some point of the support of $X$. In applications, such asymptotic distribution results can be useful to construct approximate confidence intervals. Before stating the result we make the following additional assumptions, in particular, with respect to the error term $U=Y-g(X^*)$.
The bounds imposed in Assumption (ref) $(i)$ are not stronger than the one imposed in Newey1997. We also note that it is possible to relax these conditions as noted by belloni2015 or chen2015optimal. Assumption (ref) $(ii)$ is a condition on the basis functions and satisfied for B-splines or wavelets, see Appendix E of chen2018optimal.
To obtain asymptotic normality of our estimator we require a normalization factor. Therefore, we introduce the sieve variance given by
In contrast to the usual series regression in Newey1997, we see that the sieve variance also contains the FPW function $\omega$. As $\omega$ can take values smaller than one, the sieve variance for our FPW series estimator can be even smaller than the one associated to the usual series estimator. This is in contrast to estimators based on weighting via inverse selection probabilities that always lead to larger sieve variances, see breunig2015 for selective outcomes or das2003 for propensity score weighting. We replace the sieve variance $ \textsl{v}_K(x)$ by the estimator
where $\widehat U_i=Y_i-\widehat g(X_i)$. We now establish the asymptotic distribution of the estimator $\widehat g$ evaluated at some point $x$ in the support of $X$. Similarly, asymptotic distribution results for linear functionals of $g$ can be obtained. We introduce the supremum norm $\|\phi\|_\infty=\sup_{x\in\mathcal X}|\phi(x)|$.
Condition (ref) requires the estimator of $g$ to be undersmoothed. This ensures that the sieve approximation bias in the second step estimation procedure becomes asymptotically negligible. Theorem (ref) can be also used to construct pointwise confidence intervals for $g(x)$ but can also be extended to construct uniform bootstrap confidence bands, as the following remark illustrates.
This section provides a bootstrap procedure to construct uniform confidence bands for $g$ and establishes asymptotic validity of it. Let $(\varepsilon_1,\ldots,\varepsilon_n)$ be a bootstrap sequence of i.i.d. random variables drawn independently of the data $\{(D_i,Y_i,X_i,W_i)\}_{1\leq i\leq n}$, with ${\mathbb E}[\varepsilon_{i}] = 0$, ${\mathbb E}[\varepsilon_{i}^2] = 1$, with bounded moments. Common choices of distributions for $\varepsilon_i$ include the standard Normal, Rademacher, and the two-point distribution of mammen1993. Further, $\mathbb{P}^*$ denotes the probability distribution of the bootstrap innovations $(\varepsilon_1,\ldots,\varepsilon_n)$ conditional on the data. We introduce the bootstrap process
Under regularity conditions it can be shown that the bootstrap process provides a uniform approximation of the influence function of the estimator $\widehat g$ and thus, can be used to construct uniform confidence bands (see also chen2018optimal).
For the implementation of a $100(1-\alpha)\%$ uniform confidence bands, we compute the critical value $z^B_{1-\alpha}$ as the $(1-\alpha)$ quantile of $\sup_{x}| \mathbb Z^B(x)|$ for a large number of independent bootstrap draws. The resulting $100(1-\alpha)\%$ uniform confidence band is then given by
In the following, we assume that $\mathcal C$ is a closed subset of $\mathbb R^{d_x}$. Let $\delta(\cdot,\cdot)$ be the standard deviation semimetric on $\mathcal{C}$ of the Gaussian Process $\mathbb Z(x) = p^K(x)'\mathcal Z/\sqrt{\textsl{v}_K(x)}$ with $\mathcal Z\sim \mathcal N(0,\Sigma)$ and $\Sigma={\mathbb E}[p^K(X^*) {\mathbb V}ar(U D\,\omega(Y,X^*)|X^*)\,p^K(X^*)']$ defined as $\delta(x_1,x_2) = ({\mathbb E}[(\mathbb Z(x_1) - \mathbb Z(x_2))^2])^{1/2}$, see e.g. Vaart2000. Further, we introduce the notation $N(\varepsilon,\mathcal{F},\|\cdot\|_{\mathcal F})$ for the $\epsilon$-entropy of $\mathcal{F}$ with respect to a norm $\|\cdot\|_{\mathcal F}$.
Assumption (ref) is similar to chen2018optimal who establish asymptotic validity of uniform confidence bands in nonparametric instrumental variable estimation. Assumption (ref) $(ii)$ is a mild regularity assumption, see also chen2018optimal for sufficient conditions.
The next theorem establishes asymptotic validity of the bootstrap for constructing uniform confidence bands for the regression function $g$. Below, ${\mathbb{P}}^*$ denotes a probability measure conditional on the data $\{(D_i, Y_i,X_i,W_i)\}_{i=1}^n$.
In this section, we study the finite-sample performance of our estimator by presenting the results of a Monte Carlo simulation. We first focus on the estimation of nonlinear conditional moments, then we turn to linear regression. We perform $1000$ Monte Carlo replications in each experiment and the sample size is $n=1000$.
We consider the estimation of the regression function $g$ under the following simulation design. The data are generated by $W=\Phi(\xi)$ and $X^*=\Phi(\chi)$ where $\chi=\rho\,\xi+\sqrt{1-\rho^2}\,\nu$ and $(\xi,\nu)'\sim{\cal N}(0,I_2)$. Here, $\rho$ characterizes the strength of the instruments and is varied in the experiments below. Further, we draw $Y$ from the model
where $g(x)=\Phi\big(c_g(x-0.5)\big)$ with standard normal distribution function $\Phi$, a constant $c_g\in\{5,20\}$, and $U\sim {\cal N}(0,\sigma_U^2)$, where $\sigma_U^2\in\{0.5,1\}$.
We generate realizations of the selection variable $D$ from the Bernoulli distribution
Consequently, the selection probability is a function of the latent covariates $X^*$ and also unobservables $U$. The selection probability $\varphi$ is estimated using the sieve minimum distance procedure in (ref) with tensor product of quadratic B-splines and zero knots (hence $L=9$). We estimate the function $g$ by using the FPW series estimator $\widehat g$ given in (ref). As basis functions we use quadratic B-splines with 2 knots (hence $K=5$) when $c_g=5$ (see Figure (ref)) and quadratic B-splines with 7 knots (hence $K=10$) when $c_g=20$ (see Figure (ref)). For the B-splines of $X$ note that we choose the interior knots located at the quantiles of the observed subset of $\{X_i\}_{D_i=1,1\leq i\leq n}$.
Figures (ref) and (ref) depict the median of the FPW series estimator $\widehat g$ together with its 95% pointwise confidence bands and a series estimator under the missing at random (MAR) assumption based on listwise deletion under different simulation designs. We vary the parameters $\rho$ and $\sigma_U$; the first row shows results with $\sigma_U^2=0.5$, the second row with $\sigma_U^2=1$, the first column with $\rho=0.2$, and the second column with $\rho=0.6$. In all cases the median of the FPW series estimator is close to the true regression function $g$ and the MAR series estimator is severely biased. From Figures (ref) and (ref) we see that the strength of instruments only has a moderate influence on the performance of the FPW series estimator. This is in line with our theoretical results that the asymptotic performance of the estimator is not driven by the correlation of the instruments to the latent covariates. On the other hand, we see that the variance of estimation becomes much larger as $\sigma_U^2$ increases from $0.5$ to $1$. For larger values of $\sigma_U$ the problem of selection on unobservables becomes more severe. In this case, the confidence intervals of the FPW series estimator become larger but also the bias of the MAR series estimator increases.
For estimators based on unconditional moments, an alternative approach to FPW is given by IPW. Yet this subsection demonstrates that the FPW approach leads to more accurate estimation results even in linear regression models.
We generate the data as described in the previous subsection. As we are interested in the unconditional mean ${\mathbb E}[YX^*]$ we could also make use of IPW. Indeed, making use of the notation $\widetilde \omega(y,x)=1/\mathbb{P}(D=1|Y=y,X^*=x)$ we obtain by the law of iterated expectations that
Alternatively, we can apply the FPW function $\omega(y, x)=\mathbb{P}(D=1|X^*=x)/\mathbb{P}(D=1|Y=y, X^*=x)$ and obtain by our nonparametric identification results that
Estimating the unconditional mean by FPW has the advantage over IPW that identification of the inverse selection probability is more restrictive than identification of the FPW function $\omega$. That is, for identification of the IPW the usual completeness assumption is required. In addition, we demonstrate in the following the finite sample properties of both approaches in a finite sample analysis. To do so, we consider the linear model
where $\beta_0=1$ and $\beta_1=3$. The data is generated as described in the previous subsection with $\rho=0.2$ and $\sigma_U=1$. Below, we analyze the absolute median bias and the coverage at the 95% nominal coverage rate for the FPW estimator, the IPW estimator, the MAR estimator based on listwise deletion and the estimator when there is no missing data, that is, $D\equiv 1$. We estimate the weights for FPW and IPW nonparametrically as described in the previous subsection. The FPW and IPW estimators coincide then with weighted ordinary least squares (OLS) estimators.
In Table (ref) we compare the IPW and FPW estimators with the OLS estimator under the MAR hypothesis and the OLS estimator when there is no missing data. The second and third row show the absolute median bias of the estimators of the intercept and the slope parameter. We see that the FPW estimator has smaller median bias than the IPW estimator for both parameters. Not surprisingly, the bias dramatically increases when we ignore selection and consider the MAR estimator. The last two rows depict the coverage of the confidence interval for the intercept and the slope parameter. We see that the FPW estimator has more accurate coverage than the IPW estimator. Yet there is undercoverage of the FPW estimator which is due to the severity of the selectivity of the nonresponse mechanism. Note that the 95% confidence interval of the MAR estimator contains the true intercept only in 4 out of 1000 Monte Carlo Iterations. If we relax the severity of the selectivity then the coverage of the FPW estimator is more accurate. For instance, if in (ref) the variable $U/2$ is replaced by $U/3$ then the empirical coverage of the FPW estimator for $\beta_1$ increases from 0.798 to 0.887.
In the final section of the paper we apply the developed methodology to study to relevant economic questions. First, we focus on the association between income and health. In the second application we analyze how income affects the demand for housing.
As mentioned in the introduction, a large body of literature has documented a positive correlation between income and health, see e.g. DeatPax_98.\footnote{In general, it is difficult to identify the causal effect of income on health, therefore most studies focus on the association on income and health. We follow these studies. Notable exceptions are studies that focus on the effect of income or wealth shocks on health, see e.g. Schwand_18. } However, in general these studies are based on survey data in which income, health information and further demographic variables are self reported. As shown e.g. in breunig2017 information on income in surveys is likely to suffer from nonrandom selection which might result in biased estimates of the association between income and health.
The empirical analysis is based on linked data from the German sample of the Survey of Health, Aging and Retirement in Europe (SHARE, Wave 5, collected in 2013) and the German pension insurance. SHARE is a multinational survey of the elderly population aged 50 and above in Europe, for more information see abs_2013. The survey includes standard demographic characteristics and self reported information about different income measures and various subjective and objective health outcomes. The key variables for our analysis are individual income and health outcomes. We use a broad definition of income. For non-retired individuals the income includes labor earnings, income from self employment and transfers for unemployed. For retired individuals the income is composed of own pensions, and if applicable widowers pension and additional labor earnings. The health status is described by an objective measurement of the hand grip strength. Previous studies have documented that hand grip strength is a good measure of physical functioning and a predictor of morbidity, disability and mortality, see e.g. rantanen1999midlife, bohannon2015muscle, or dodds2014grip. From the grip strength we construct a binary variable which indicates bad health status if the grip strength is below the 25th percentile.
For our analysis we exploit a specific feature of the data which allows us to link a subsample\footnote{The linkage of the data requires the consent of the individuals, about 2/3 of individuals agreed to the linkage.} of the survey data to administrative data of the German pension insurance. Thus, in addition to the self-reported income information which might suffer from nonrandom nonresponse the data includes official information about pension entitlements. For pensioners we observe the full pension entitlements, i.e. number of pension points, they have earned during their working life; for non retired individuals we observe the entitlements they have collected so far. Pension entitlements are a deterministic function of the full individual earnings history. The earnings history is a good predictor of current income, however, it contains no direct information about the response behavior for current income. Therefore, this information allows us to construct a suitable instrument to account for potential nonrandom nonresponse of the current self reported income. In fact this instrument is superior to instruments based on self reported lagged employment outcomes which are often used, see e.g. breunig2017. First, the instrument is not affected by transitory shocks since it combines information about the full working life instead of using information of only one period. Second, self reported past information might as well suffer from nonresponse. This is not the case for the information about pension entitlements in administrative data.
In the empirical analysis we concentrate on $3340$ individuals which are younger than 80 years and who have agreed to the linkage of the survey data and the information of the pension insurance. Out of this sample, $12.34\%$ do not respond to the income information question.\footnote{In our sample only 4.8% do not provide information about grip strength - we assume that this information is missing at random. Using the test of missing (completely) at random by breunig2017 we obtain the value of the test statistic 0.053 with $0.05$-- level critical value of 0.10 and hence, we fail to reject the MAR hypothesis.} Table (ref) provides summary statistics of the relevant variables for the analysis.
To quantify the association between income and health we use the following semiparametric model
where the function $g$ and the parameters $\alpha_0$ and $\beta_0$ are unknown. We assume that $U_i$ is conditional mean independent of the explanatory variables, $\log(Income_i^*)$, $Age_i$, and $Gender_i$. We apply the FPW estimator as described in the previous section, i.e., we estimate the nonparametric selection probability using quadratic B-spline basis functions for least square approximations. Specifically, the selection probability $\varphi$ is estimated using the sieve minimum distance procedure described in (ref) with tensor product of quadratic B-splines and zero knots, where we additionally control for age and gender. We estimate the function $g$ using the FPW series estimator $\widehat g$ given in (ref) using quadratic B-splines with 1 knot (placed at the median of observed income) and again controlling for age and gender. The MAR estimator uses the same choice of B-spline basis functions with the same knot placement.
Figure (ref) depicts the FPW series estimator with the 95% uniform confidence bands together with the MAR series estimator evaluated at the median age. The uniform confidence bands are computed using the bootstrap procedure as described in Section (ref) with 1000 bootstrap iterations. For the MAR series estimator, we consider listwise deletion of missing values. The FPW series estimator shows a negative association between income and bad health measured by the grip strength which is moderate. For example we find, that the risk of bad health for individuals at the 25th percentile of income (about 8.9 log income) is about 40% while for individuals at the 75th percentile (log income of 9.9) it is slightly lower (about 37.5%). However, according to the confidence interval this difference is not significant. For higher incomes the risk only changes moderately and changes are again not statistically different. Importantly, our analysis shows that the MAR assumption leads to biased results and potentially erroneous conclusions about the association between income and health. The negative association obtained in the MAR estimator is far more pronounced than in the estimator which accounts for the nonrandom nonresponse. Specifically, with missing at random we predict a risk of bad health at the 25th percentile of about 44% which is only close to 40% at the 75th percentile of the income distribution. Note, the confidence intervals show that the results of the two different estimators are significantly different for incomes below the median. For higher incomes the differences are not significantly different.
In addition to the non-linear analysis, we assume a linear $g$ function and present results from a linear model. This linear specification has been used in the literature to test the absolute income hypothesis derived in Pre_1975, see e.g. adeline2017.
In Table (ref) we depict the results using ordinary least squares estimators with and without probability weighting to account for nonrandom nonresponse. For the FPW estimator we leave the functional form of the selection probability completely unrestricted. Overall, this application underlines the importance to account for nonrandom nonresponse in income information when studying the link between income and health. Importantly, while the MAR finds a negative relation between bad health and income which is significant at the 5% level, the estimators which account for nonrandom nonresponse reject a significant relation between bad health and income in a linear model. Finally, we note that overall the FPW estimator does not lead to larger standard errors relative to MAR, even if the selection probability is estimated via nonparametric instrumental variable method.
In the second application we revisit the question how income affects the demand for housing, for previous studies see e.g. QuiRap2004, Albouyetal2016 or DusFitZim2018. For example, DusFitZim2018 show for Germany that about 70% of households in the lowest income quintile are renters whereas in the highest quintile the share is with 30% markedly lower. As in the application of health and income, studies on housing demand are in general based on survey data with self reported income which suffer from nonrandom selection. In the following we estimate the relationship between income and the probability to own a house and quantify the bias when not accounting for nonrandom nonresponse.
The empirical analysis is based on the data of the SOEP. The SOEP is a longitudinal household survey of the German population, for more information see WagFriSch_07. The survey includes self reported standard socio-demographic characteristics including housing and information about different income measures. In contrast to the previous application we focus on a narrow definition of income, labor earnings, which is the most important income component for most individuals. Since a larger fraction of women does not have positive labor earnings, we restrict the analyses to men.
The individual SOEP data can be linked to regional data with information about the average socio-economic situation at the ZIP-code level or even at the residential block. The regional data is provided by a private marketing company which uses administrative information from tax records in combination with credit card information and information about local infrastructure, for more details see Goebel_etal_2014. From the regional data, we use the information about the average purchasing power of households living in a specific residential block to construct an instrument for potentially non-random missings of the self-reported earnings information. This information is well suited to construct an instrument. First there exists a strong positive correlation (0.3025) between the individual labor earning and the average purchasing power of households living in a specific residential block. Second, the regional information is available for all individuals such that the instrument does by definition not suffer from nonresponse.
In the empirical analysis we concentrate on $11735$ employed men which are younger than 65 years. Table (ref) provides summary statistics of the relevant variables for the analysis.
To quantify the association between earnings and home ownership we use again a semiparametric model with a binary indicator of home ownership and age as depend variable and log monthly earnings as explanatory variable.
where the function $g$ and parameter $\alpha_0$ are unknown. $U_i$ is conditional mean independent of the explanatory variables, $\log(Earning_i^*)$ and $Age_i$.
Figure (ref) depicts the FPW series estimator with the 95% uniform confidence bands (again described in Section (ref) using 1000 bootstrap iterations) together with the MAR series estimator evaluated at the median age. The selection probability $\varphi$ is again estimated using the sieve minimum distance procedure in (ref) with tensor product of cubic B-splines and one knot. We estimate the function $g$ using the FPW series estimator $\widehat g$ given in (ref) using cubic B-splines with two knots placed at the quantiles of the observed values of earnings. The MAR estimator is also based on cubic B-splines with the same number and placement of knots. The FPW series estimator shows a strong positive but non-linear relation between earnings and home ownership. For example we find, that the probability to own a house amounts to about 30% at monthly earnings below the median earnings in the sample (about 2400 Euros per months). The probability markedly increases to 60% at monthly earnings at the 75% percentile (8.25 log points or about 3800 Euro per months) and further to close to 70% at earnings above 6500 Euros (8.8 log points). Importantly, our analysis shows that the MAR assumption might lead to biased results about the association between earnings and the home ownership rate for individuals with very low earnings, i.e. individuals in the lowest decile, which is a key group for public policy. Specifically, we find significantly lower probabilities of home ownerships when not accounting for nonrandom nonresponse. The difference is with about 10 percentage points economically important. Interestingly, for the rest of the earnings distribution the two estimators do not significantly differ and point at a positive relationship between earnings and home ownership
Finally, we assume again a linear $g$ function and consider a linear model (see Table (ref)). Since the missing at random assumption only affected a small part of the earnings distribution in the nonlinear application, it is not surprising that the coefficients in the linear model do not significantly differ. As expected we find a slightly larger point estimator (0.127) when assuming that nonresponse is random.
In this paper we derive a nonparametric estimators that addresses the problem of nonrandom selection that can be related to nonrandom selection into treatment programs, selective measurement error or through selective nonresponse or missingness of data. Identification of the regression function relies on instrumental variables that are independent of selection conditional on potential covariates. We obtain identification of our nonparametric regression function without restricting the selection probability to belong to a parametric class of functions via a novel partial completeness assumption and provide primitive conditions for it. We achieve optimal rates of nonparametric rates of convergence of our estimator. Moreover, the variance of our estimator is not larger than in the case where the variables are fully observed.
We demonstrate the usefulness and relevance of our method in survey data with nonrandom missingness in two different applications with different instruments. First, we analyze the association between bad health and income. We show that standard methods that do not account for the nonrandom selection process are strongly upward biased for individuals with below-median income. Moreover, in a linear model the standard estimator finds a negative relation which is significant at the 5% level, however the estimators which account for nonrandom non-response reject a significant relation between bad health and income. In the second application we focus on the relation between housing and earnings. While the different estimators with and without the assumption of random missingness lead to the similar estimates in a linear model and for a large share of the earnings distribution, we document significant and important differences for individuals in the lowest earnings decile, which is a central group for public policy.