Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,136 characters · 13 sections · 64 citation commands
Varying Random Coefficient Models
Heterogeneity in individual behavior is a common source of variation in microeconometric applications. Thus, in recent years it became increasingly popular to explicitly model unobserved heterogeneity, for instance, by introducing random coefficients. Yet identification in random coefficient models requires functional form restrictions. This paper shows that one can be much more flexible with respect to observed characteristics which can essentially influence the shape of the density of random coefficients. Specifically we extend the ordinary random coefficient model to allow for nonlinearities in observed heterogeneity, captured by varying coefficients.
The {\it varying random coefficient} (VRC) model is given by
where $Y$ is a scalar dependent variable and $X$ is a vector of covariates of dimension $d-1$. The VRC vector $B=(B_0,B_1')'$ satisfies
for some covariates $W$ and unobservables $A=(A_0,A_{1,1}, \dots, A_{1,d-1})'$. The varying coefficient functions $g_0(\cdot),\dots,g_{d-1}(\cdot)$ are unknown and capture nonlinearities in observed heterogeneity. The vectors $X$ and $W$ may have elements in common but, to ensure identification of the varying coefficient functions in general, we rule out that $X$ is a subvector of $W$. In this model, the {\it varying random slope} (VRS) given by $B_1$ represents observed and unobserved heterogeneity in the dependence of $Y$ on $X$. The VRC is thus more general than the ordinary random coefficient (RC) model where all varying coefficient functions vanish and hence, $B=A$. Identification of the VRC model is based on full independence of $A$ and $X$ but only conditional mean independence of $A$ and $(X,W)$. While identification of the model requires $X$ to have enough variation, our setup permits $W$ to be discrete.
This paper is concerned with inference on the density of the VRC vector $B$ holding observed characteristics $W$ fixed. This density contains all the information of the underlying heterogeneity in the model and many functionals of it are of interest, e.g., the distribution of potential outcomes. A density estimator based on weighted sieve minimum distance is proposed that builds on a conditional characteristic function equation of the model. The estimation criterion is minimized over a finite dimensional sieve space which is also convenient to impose shape restrictions on the estimator. The estimator is of closed form if no constraints are imposed on the sieve space and then coincides with a double series least squares estimator. An initial weighting step is used in the estimation criterion to stabilize the procedure. This is important, as it is well known that estimation of the joint RC density in ordinary RC models leads to an ill-posed inverse problem. One insight of this paper is that our procedure allows us to separate estimation of the VRS density from estimation of the varying random intercept $B_0$. In particular, we show that estimation of the VRS density does not suffer from the ill-posed estimation problem once we impose finite dimensional restrictions on the density of $B_0$.
For the sieve minimum distance estimator, inference results are established that go beyond what has been obtained in ordinary RC models. The rate of convergence of the estimator is derived, which coincides with the usual ill-posed rate of convergence when estimating the joint VRC density and corresponds to the usual well-posed, nonparametric rate when only the density of VRS is of interest and semiparametric restrictions on the random intercept are imposed. Many important objects of interest, such as the distribution of the potential outcome, are functionals of the density of $B$. For a plug-in estimator of such linear functionals we establish pointwise asymptotic normality. This paper also provides a bootstrap procedure to construct uniform confidence bands of the estimator. The inference results in this paper make explicit how the marginal distribution of $X$ affects the asymptotic behavior of the estimator.
Identification of the model requires a large support condition of $X$ in general. Yet under mild assumptions, identification can be also achieved under bounded support via extrapolation. This motivates the use of global basis functions for the sieve minimum distance estimator, such as the Hermite functions. When the sieve space is spanned by Hermite functions, we see that the estimator considerably simplifies as these basis functions form eigenfunctions of the Fourier transform. The estimator performs well in finite samples even when covariates $X$ are far from being heavy tailed. This is demonstrated in Monte Carlo simulations that also clarify how the variance of $X$ affects the mean integrated squared error of the estimator in finite samples.
The estimation procedure is also applied to analyze heterogeneity in income elasticity of housing demand using German survey data. In our specification, estimated observed housing characteristics exhibit a nonlinear shape. The estimated density of heterogeneous income elasticity is unimodal with mode close to zero and positively skewed. Uniform confidence bands allow to make significant statements about the shape of the estimated density. The empirical application demonstrates that our proposed methodology can be useful to analyze complex heterogeneity using cross sectional data.
Nonparametric identification and estimation of ordinary RC models is considered by beran1992, beran1996, and hoderlein2010. For testing of qualitative features of the ordinary RC models see dunker2017. lewbel2017unobserved generalize these models to allow for nonlinear index functions and in the next section we will provide a more detailed comparison of it to the VRC model. The ordinary RC models can be extended to {\it conditional random coefficient models} that assume model equation (ref) together with the condition that $X$ is independent of $B$ conditional on $W$. This model is more flexible than $(\ref{mod:gen}$--$\ref{mod:gen:rc})$, but requires that the conditional density of $X$ given $W$ satisfies a large support condition for all realizations of $W$. Beyond this more restrictive support restriction, estimation in conditional RC models also suffers from the curse of dimensionality of $W$. This is why functional form restrictions are typically employed rather than considering the more general conditional RC models, see for instance, lewbel2017unobserved. Recently, {\it correlated random coefficient models} were studied in the literature which allow for full dependence of random coefficients and covariates $X$. In this setting, instruments are available that drive the covariates $X$ but not the random coefficients in (ref). These types of models are analyzed by masten2014 and hoderlein2014. While such a model is clearly more general than the VRC model, identification of the CRC models can be challenging with more restrictive exclusion assumptions and large support conditions on the instruments.
The methodology of sieve estimation became increasingly popular in recent years. For sieve estimation of conditional moment restrictions models see NP03econometrica and AC03econometrica. The VRC model does not fall into this category. For sieve estimation of ordinary RC models with discrete outcome see fox2011simple and fox2016simple. In binary choice models, gautier2011 proposed an estimator based on needlet thresholding. In the context of specification testing, a sieve approach was used by breunig2016. In the literature on varying coefficients, series estimators were analyzed by xia1999, fan2003, xue2012, or ma2015.
The remainder of the paper is organized as follows. Section (ref) provides the setup, motivating examples, and sufficient conditions for nonparametric identification. In Section (ref), the estimation procedure based on sieve minimum distance is introduced and its asymptotic properties are established. Section (ref) is concerned with the finite sample properties of the estimator analyzed via Monte Carlo simulation and an empirical illustration. Section (ref) concludes. All proofs can be found in the appendix.
This section consists of two subsections. Subsection (ref) recalls the varying random coefficients (VRC) model, outlines its key properties, and provides motivating examples for it. Subsection (ref) provides an identification result of the joint density of the VRC vector $B$.
Consider again equations $(\ref{mod:gen}$--$\ref{mod:gen:rc})$, the VRC model is given by
As stated above $X$ and $W$ may have elements in common. Yet without further functional form restrictions on the varying coefficient functions $g_0,\dots,g_{d-1}$ we need to rule out that $X$ has only elements that are contained in $W$ (see also Assumption (ref) and the discussion thereafter). The covariates $X$ are assumed to be independent of $A$ and the vector of covariates $(X',W')'$ is restricted to be mean independent of $A$, i.e., $\operatorname{\mathbb{E}}[A|X, W]=0$ (see Assumption (ref) below). All the results in this paper will hold if there were no varying coefficients, i.e., model $(\ref{mod:gen}$--$\ref{mod:gen:rc})$ is the ordinary RC model. While identification of the model requires $X$ to have enough variation, our setup permits $W$ to be discrete.
Under conditional mean independence of $A$ and $(X,W)$, model $(\ref{mod:gen}$--$\ref{mod:gen:rc})$ implies the varying coefficient model
using the notation $S=(X',W')'$. In the conditional mean restriction (ref), the varying coefficient functions $g_l$ are identified through a rank condition, see also below. For an overview article of varying coefficient models see park2015varying. We emphasize that the varying coefficients specification in (ref) also derives from the additive separability of the random coefficients $A$ in equation (ref). Without imposing such an additively separable structure, our model is in general not identified, see masten2014. On the other hand, under further assumptions, lewbel2017unobserved establish identification when equation (ref) is replaced by $Y=B_0+\sum_{l=1}^{d-1}G_l(B_{1,l}X_l)$ for unknown functions $G_l$ and $A$ is independent of $(X,W)$ (it is thus non-nested to our VRC model). While the sieve minimum distance approach in this paper allows for estimation of additional nonlinear index functions, we do not consider this extension.
Identification of the varying coefficient functions $g_0,\dots,g_{d-1}$ in the VRC model $(\ref{mod:gen}$--$\ref{mod:gen:rc})$ can be obtained in the following two scenarios: First, under a full rank condition on $\operatorname{\mathbb{E}}[XX'|W=w]$ or, second, if $g_1\equiv\dots\equiv g_{d-1}\equiv0$. In the first case, the VRC model has the interpretation similar to a conditional random coefficient model but, instead of leaving the distribution of $X$ and $W$ unrestricted, it imposes more structure on it. Also similarly to conditional random coefficient models, the role of $W$ is to ensure that independence between $A$ and $X$ is more plausible. The variables $W$ can also serve as control function residuals, which then allows $X$ to be endogenous, i.e., $A$ is unconditionally correlated with $X$.\footnote{Following imbens2009 assume that endogenous regressors $X$ are related to observed instruments $Z$ via $X=\psi(Z,W)$ where $\psi$ is strictly monotonic in scalar $W$ and $Z\perp(B,W)$. The model implies independence of $X$ and $A$ conditional on $W=F_{X|Z}(X|Z)$.} In the second case, only $g_0$ differs from the zero function and thus the VRC model has the interpretation of a nonlinear model in $X$, when $X=W$, but with random coefficients only on the linear term.
This paper is concerned with estimation of the VRC density holding the observed characteristics fixed at some potential realization $w$, i.e., the density $f_B(\cdot,w)$ of
where $\mathbf g(\cdot)=(g_0(\cdot),\dots,g_{d-1}(\cdot))'$ denotes the vector of varying coefficient functions. In the case where $X$ and $W$ have no common elements, $B^w$ captures heterogeneous marginal effects. If $X$ and $W$ have joint elements then, to obtain marginal effects, replace $\mathbf g(w)$ with the vector of partial derivatives of $\mathbf g(w)$ w.r.t. $x$. The results in this paper are still valid in this case but it is not made explicit in order to keep the notation simple. By holding observed characteristics $W$ fixed, the density $f_B(\cdot,w)$ contains all information on unobserved heterogeneity. Many objects of interest, such as the density of potential outcomes of $Y$, are linear functionals of $f_B(\cdot,w)$.
Economic theory and empirical findings suggest nonlinearities in many applications of interest. For instance, through a nonparametric analysis of Engel curves to analyze nonlinearities in total expenditure, banks1997 suggest quadratic terms in the logarithm of total expenditure. While a random coefficient version of their nonlinear model is not identified one might still account for unobserved heterogeneity by allowing only the coefficient for the linear term to vary among individuals. The following two examples provide a relation of VRC to measurement error models and show that ignoring nonlinearities in the varying coefficients may have severe consequences in a standard Monte Carlo exercise setting.
This section provides assumptions under which the density $f_B(\cdot,w)$ of $B^w$ is identified for any value $w$ in the support of $W$. We now impose restrictions on the observed and unobserved variables of our model.
An independence assumption similar to Assumption (ref) is common in the literature on RC models with cross-sectional data (see, for instance, beran1993, beran1996, hoderlein2010). It should be also emphasized that the independence assumption can often be justified in our model if the information in $W$ about heterogeneity is rich enough and in this sense is milder than in the ordinary RC model where $\mathbf g\equiv 0$. Clearly, if $W$ contains only elements in $X$ and $\operatorname{\mathbb{E}}[A]=0$ then Assumption (ref) (ii) is implied by (i).
Assumption (ref) (iii) is automatically satisfied if $g_l\equiv 0$ for all $1\leq l\leq d-1$. Otherwise, Assumption (ref) (iii) is satisfied if for all $w$ in the support of $W$, the smallest eigenvalue of $\operatorname{\mathbb{E}}[XX'|W=w]$ is bounded away from zero. This rank condition is commonly imposed in the varying coefficient literature. In particular, it rules out that the vector of regressors $X$ has only values that are also contained in the vector $W$. It is also possible to relax the rank condition by imposing functional form restrictions on $g$, see fan2003. Throughout the paper, the conditional characteristic function of $Y-g(S)$ given $X$ is denoted by $h(x,t;g)=\operatorname{\mathbb{E}}\big[\exp\big(it(Y-g(S))\big)\big|X=x\big]$.
While large support conditions are often required in econometrics to ensure identification, Assumption (ref) (i) can be relaxed. If the distribution of $A$ has finite absolute moments and is uniquely determined by its moments, then identification with bounded support of $X$ can be achieved by extrapolation, see masten2014 and hoderlein2014. Assumption (ref) (ii) imposes a mild regularity assumption on the conditional characteristic function $h$.
The next result establishes identification of the VRC density. We make use of identification of the varying coefficients through the conditional mean restriction (ref). Consequently, by employing the relation $f_{ B}(b,w)= f_A(b-\mathbf g(w))$ the following result is due to Fourier inversion.
Lemma (ref) shows that the density $f_B(\cdot,w)$ can be written as a transform of varying coefficient functions. Besides the shift of the conditional characteristic function $\operatorname{\mathbb{E}}[\exp(itY)|X=x]$ to $h(x,t;g)$ there is also a location shift by $\mathbf g(w)$. This corresponds to shifts in frequency and time domain for the Fourier transform.
This section presents an estimator of the VRC density based on sieve minimum distance and establishes its asymptotic properties. Subsection (ref) introduces the weighted sieve minimum distance estimator and motivates the use of Hermite functions as sieve basis. In Subsection (ref), the rate of convergence of the estimator of $f_B$ is derived. Subsection (ref) establishes the pointwise asymptotic distribution of linear functionals of $f_B$, where the density of potential outcome is one particular example of interest. Subsection (ref) presents a Bootstrap procedure to construct uniform confidence bands.
Estimation builds on the conditional characteristic function equation induced by the model $(\ref{mod:gen}$--$\ref{mod:gen:rc})$. We denote the Fourier transform by $(\mathcal{F}\phi )(u):= \int_{\mathbb R^d} \exp (iu'z)\phi (z)dz$ for any absolutely integrable function $\phi$. Recall the notation of the conditional characteristic function $h(x,t;g)=\operatorname{\mathbb{E}}[\exp(it(Y-g(S)))|X=x]$. Independence of $X$ and $A$, as imposed in Assumption (ref), immediately implies
and hence the relation
for all $t\in\mathbb R$ and $x\in\mathbb{R}^{d-1}$. Moreover, relation (ref) leads to
for some measure $\nu$. Using this $L^2$ criterion we construct a sieve minimum distance estimator of $f_A$ and use a plug-in approach to estimate the VRC density $f_B$. Below, we show that the choice of the log-normal distribution as weighting measure $\nu$ is well suited for our estimation problem.
The proposed sieve minimum distance estimator of the VRC density $f_{B}(\cdot,w)$ is based on the relation $f_{B}(b,w)= f_A(b-\mathbf g(w))$ for any $w$ in the support of $W$: Consider the plug-in estimator
where $\widehat f_A$ is a sieve minimum distance estimator of $f_A$ given by
and $\mathcal{A} _K$ is a sieve space of dimension $K=K(n)<\infty $ with basis functions $\{q_k\}_{k\geq 1}$. The sieve dimension $K=K(n)$ grows slowly with sample size $n$. Sieve estimation is also convenient to impose additional constraints on the unknown functions. These constraints, such as positivity, can be directly imposed on the sieve space $\mathcal A_K$. When the constraints are not binding, we may consider $\mathcal A_K$ without constraints, which then coincides with the linear sieve space $\mathcal{A} _K=\big\{\phi(\cdot)=\beta' q^K(\cdot):\,\beta\in\mathbb{R}^K\big\}$ where $q^K(\cdot)=\big(q_1(\cdot),\dots,q_K(\cdot)\big)'$.
The unknown conditional characteristic function $h$ is replaced by the plug-in series least squares estimator
where $p^K(\cdot)=\big(p_1(\cdot),\dots,p_K(\cdot)\big)'$ is a vector of basis functions and we use the notation $\widehat P= n^{-1}\sum_{j=1}^np^K(X_j)p^K(X_j)'$. Thus, the same sieve dimension $K$ (as for $\mathcal A_K$) is used to approximate the conditional characteristic function $h$. The estimator $\widehat g$ of the regression function $g$ is based on the conditional mean restriction (ref). We do not impose an explicit form of this estimator but rather impose a rate condition on the estimator $\widehat g$ to obtain our asymptotic results. In particular, $\widehat g$ can account for generated regressors when $W$ are control function residuals. Below, Example (ref) provides an illustration of estimating $g$ via series least squares. Although estimation of the density $f_B$ involves two preliminary steps (estimation of $h$ and $g$) it should be emphasized that the estimation procedure is equivalent to a one-step sieve minimum distance estimator, which, additionally involves the conditional characteristic equation $h(X,t;g)=\operatorname{\mathbb{E}}\big[\exp\big(it(Y-g(S))\big)\big|X\big]$ and the conditional moment equation (ref).
When no constraints are imposed, the sieve minimum distance estimator (ref) is of closed form. In this case, $\widehat f_{ B}$ coincides with the {\it double series least squares estimator}
where
is assumed to be nonsingular (at least for $K$ sufficiently large). Nonsingularity of $Q$ is satisfied by Hermite functions (under mild conditions) which, as the following example illustrates, are a convenient choice of bases for the sieve space $\mathcal A_K$.
We now consider the case of tensor product basis functions. We make use of the notation $\text{I}_{K_1}$ for the $K_1$-dimensional identity matrix and $\otimes$ for the Kronecker product.
Lemma (ref) shows that given tensor product Hermite functions in a linear sieve space only the dimension parameter $K_0$ used to approximate the random intercept induces potentially small eigenvalues of $Q$ and hence, slow accuracy of estimators. Thus, if we are willing to restrict the complexity of sieve estimation in terms of $K_0$, for instance, by imposing a fixed dimension $K_0$, then the rate of convergence does not suffer from an ill-posed inverse problem. This semiparametric specification provides a numerically stable estimation procedure as we see in our Monte Carlo simulation section.
\paragraph{Notation.} Introduce the norms $\|\phi \| _{\nu }=\big(\int_{\mathbb{R}} \int_{\mathbb{R}^{d-1}}|\phi (x,t)|^{2}dx d\nu(t)\big)^{1/2}$ and $\| \phi\|_{\mathbb{R}^d}=\big(\int_{\mathbb{R}^d}|\phi (u)|^{2}du\big)^{1/2}$ with associated Hilbert spaces $L^2_\nu={\left\lbrace \phi:\|\phi\|_\nu<\infty\right\rbrace }$ and $L^2(\mathbb{R}^d)={\left\lbrace \phi:\|\phi\|_{\mathbb{R}^d}<\infty\right\rbrace }$. Moreover, $\|\cdot\|$ and $\|\cdot\|_\infty$ denote the $\ell_2$--norm and the supremum norm. For a matrix $A$, the operator norm is denoted by $\|A\|$. For a random vector $V$ the corresponding calligraphic capital letter $\mathcal V$ denotes its support. Let $\Pi_K:L^2(\mathbb{R}^d)\to\mathcal A_K$ denote the least squares projection onto $\mathcal A_K$, i.e., $\Pi_K f=\operatorname*{arg\,min}_{\phi\in\mathcal A_K}\|\phi-f\|_{\mathbb{R}^d}$ for all $f\in L^2(\mathbb{R}^d)$. Further, define the function $\gamma:\mathbb{R}\to\mathbb{R}^K$ given by $\gamma(t)=P^{-1}\operatorname{\mathbb{E}}[\exp\big(it(Y-g(S))\big)p^K(X)]$ where we recall $P=\operatorname{\mathbb{E}}[ p^K(X)p^K(X)']$. We use the notation $a_n\sim b_n$ for $c b_n\leq a_n\leq C b_n$ given two constant $c,C>0$ and all $n\geq 1$.
This subsection provides rates of convergence of the proposed estimators. In particular, we see that only the rate of the VRC density hinges on the minimal eigenvalues of $Q$. On the other hand, it is shown below that the rate of convergence of the estimator of the varying random slope $B_1$ is not affected by the choice of $\nu$.
We introduce an assumption required for our asymptotic results. Below, $\lambda_{\max}(M)$ denotes the maximal eigenvalue of a matrix $M$. Throughout the paper, $(\tau _k)_{k\geq 1}$ and $(\lambda _{k})_{k\geq 1}$ denote nonincreasing sequences.
Assumption (ref) (ii) imposes a mild restriction on the weighting measure $\nu$ and is satisfied, for instance, if $\nu$ is the log-normal distribution. The maximal eigenvalue restriction imposed in Assumption (ref) (iii) is automatically satisfied if $\{p_k\}_{k\geq 1}$ and $\{q_k\}_{k\geq 1}$ form orthonormal bases in $L^2(\mathbb{R}^{d-1})$ and $L^2(\mathbb{R}^d)$, respectively. Assumption (ref) (iv) allows the minimal eigenvalues of $P$ to tend to zero, see also chen2015, who consider a similar restriction on the growth of basis functions. This condition holds for Hermite functions, which are uniformly bounded, but also for polynomial splines, Fourier series and wavelet bases, see also belloni2012 for further discussion and extensions of newey1997. It is well known, that the smallest eigenvalue of $P$ is uniformly bounded away from zero if ${\left\lbrace p_l\right\rbrace }_{l\geq 1}$ forms an orthonormal basis and $f_X$ is uniformly bounded away from zero on the support of $X$. Thus, Assumption (ref) (iv) is particularly suited in our case where a large support assumption of $X$ is required for identification. Here, $\lambda_K$ has the interpretation of a truncation parameter in estimation problems to ensure that the denominator is bounded away from zero, see breunig2016. Assumption (ref) (v) imposes a rate restriction on the minimal eigenvalue of the matrix $Q$ relative to $(\tau _k)_{k\geq 1}$. Assumption (ref) (v) further specifies a link of the sieve approximation error on the RC density $f_A$ between the “strong” norm $\| \cdot\|_{\mathbb{R}^d}$ and the “weak” norm $\|\cdot\|_\nu$. Note that Assumption (ref) (v) is only required for estimating the VRC density but not for estimating the VRS density. Such stability conditions are commonly imposed in the literature on nonparametric instrumental variable estimation, see, for instance, BCK07econometrica or Chen08, but rely on mapping properties of an unknown conditional expectation operator. In addition, in VRC models it is possible to provide primitive conditions, as we see below.
In the case of Hermite functions, condition (ref) requires $\nu$ to be sufficiently heavy tailed while condition (ref) requires $\int_\mathbb R\big|(\mathcal F q_l)(t) \big|^2\,\widetilde \nu(t)d\mu(t)$ to be sufficiently small for all $l>K$. Both conditions impose restrictions on the weighting measure $\nu$ and the basis functions $\{q_k\}_{k\geq 1}$ but do not on the unknown density $f_A$ itself. The following example discusses the log-normal distribution as a weighting measure and provides primitive conditions for inequality (ref).
We impose regularity conditions on the unknown functions of interest. To do so, we introduce the space of square integrable functions $L_S^2=\big\{\phi:\|\phi\|_S:=\sqrt{\operatorname{\mathbb{E}}\phi^2(S)}<\infty\big\}$. Further, we define the class of functions $\mathcal G:=\mathcal G(n):=\{\phi=\beta'p_d^K:\|\phi-g\|_\infty\leq \sqrt{K/n}\}$. The covering number $N(\mathcal G , \|\cdot\|_S, \epsilon)$ is the minimal number of $L_S^2$--balls of radius $\epsilon$ needed to cover $\mathcal G$.
Assumption (ref) (i) captures regularity conditions on the density $f_A$ and the conditional characteristic function $h$ via sieve approximation errors. Example (ref) characterizes the sieve approximation condition when using Hermite basis functions relative to the smoothness of the unknown function of interest. For further discussion and other examples of sieve bases, see Chen07. Assumption (ref) (ii) imposes a mild smoothness condition on the density $f_A$. Assumption (ref) (iii) imposes rate restrictions on the varying coefficients estimator. Assumption (ref) (iv) is a regularity condition on the function class $\mathcal G$ and was also imposed in the literature, for instance, in Chen08.
Consider first estimation of the joint density of the VRC vector $B$ holding observed characteristics fixed. The following theorem provides the rate of convergence in $L^2(\mathbb{R}^d)$--norm of the estimator $\widehat f_{B}(\cdot,w)$ given in (ref) for some fixed $w\in\mathcal W$.
We see from the previous result that the parameters $\tau_K$ and $\lambda_K$, which capture the minimal eigenvalues of $Q$ and $P$, respectively, slow down the rate of convergence. As the weighting measure $\nu$ influences the eigenvalues of $Q$ it does also affect the rate of convergence of the joint density of the varying random coefficient $B$. Note that the optimal choice of the dimension parameter $K$ is lower than for usual nonparametric estimation problems and relative to the value of $\tau_K$ and $\lambda_K$.
We now provide the rate of convergence in $L^2(\mathbb{R}^d)$--norm of $\widehat f_B$ when the dimension parameter $K$ is chosen such that the variance part $(n\tau_K\lambda_K)^{-1}K$ and the squared bias part $K^{-2\zeta/d}$ are of the same order. For ease of exposition we assume that $\lambda_K$ is uniformly bounded away from zero, which requires a sufficiently large support of regressors $X$. We further consider the case that $n^{-1}K$ dominates the squared bias part of estimating $h$ which is $K^{-2\rho/(d-1)}$. The next result immediately follows from Theorem (ref) and hence, its proof is omitted.
As we see from the previous result, the usual rate of convergence for ill-posed inverse problems is obtained if $K$ levels variance and squared bias, see also hoderlein2010 in the case of kernel estimation in ordinary RC models. Yet in contrast to hoderlein2010 (see their Section 4.3), no heavy-tailedness of covariates $X$ is required for our convergence results. In particular, hohmann2016 showed under Gaussian white noise that the rate of convergence can be much slower, even severely ill-posed, under lighter tails of $X$. Also note that the condition (ref) ensures that $n^{-1}K$ dominates the squared bias of estimating $h$ and is only a mild restriction. Indeed, if, for instance, $d=2$ then the usual rate of convergence is obtained if $2\rho\geq \zeta+\alpha$.
Below, we establish a rate of convergence for estimating the density of the varying random slope $B_1$ holding observed characteristics $W$ fixed. The next results show that for estimating the VRS vector $B_1$ when we restrict the sieve dimension, to approximate the random intercept, to be bounded. In this case the rate of convergence does not depend on the eigenvalue behavior of the matrix $Q$.
Note that the condition $K_0=O(1)$ together with Assumption (ref) (i), that is, $\| \Pi_K f_A-f_A\|_{\mathbb R^d}=O(K^{-\zeta/d})$, imposes a semiparametric restriction on the RC density $f_A$. We emphasize that the parametric restriction concerns only the random intercept $A_0$ but not the random slope $A_1$. Also the rate of convergence derived in the previous result depends on the dimension parameter $K_1$ only. This is due to the definition of our estimator where the dimension of basis functions used to estimate the conditional characteristic function $h$ coincides with the dimension parameter $K_1$, see also Remark (ref). Similarly to Theorem (ref), convergence rates of subvector VRS densities can be derived.
The following result provides the rate of convergence when the dimension parameter $K_1$ such that the variance part $(n\lambda_{K_1})^{-1}K$ and the squared bias $K_1^{-2\zeta/d}$ are of the same order and $n^{-1}K_1$ dominates the squared bias of estimating $h$. Again we consider the case where $\lambda_{K_1}$ is uniformly bounded away from zero. The next result immediately follows from Theorem (ref) and hence, its proof is omitted.
As we see from the previous result, the rate of convergence coincides with the usual optimal rate in nonparametric density of dimension $d-1$. We see that the rate for estimating the density of the varying random slope $B_1$ is not affected by the ill-posedness under the condition $K_0=O(1)$.
This subsection is about inference of linear functionals $\ell(\cdot):L^2(\mathbb R^d)\to\mathbb R$ of the density $f_B(\cdot,w)$ for some realization $w$ of $W$. Examples of linear functionals are, but are not limited to, the point-evaluation functional or (weighted) averages of $f_B(\cdot,w)$. The linear functional $\ell(f_B(\cdot,w))$ is estimated below using the plug-in sieve estimate estimator $\ell(\widehat f_B(\cdot,w))$. This subsection provides a limit theory for this plug-in estimator.
As mentioned in the introduction, an important example of a linear functional of $f_B$ is the density of the potential outcome $Y^{s}=\mathbf g(w)'x+A'x$, where $s=(x',w')'$. Indeed, the density of $Y^s$ can be written as\footnote{This relation follows immediately by the well known property $f_{A_0+A_1'x}(c)=\int_{\mathbb R^{d-1}} f_A\big(c-u'x,u\big)du$.}
Also the weighted averages of potential outcomes $\int_\mathbb R \omega(y) f_Y\big(y,s\big)dy$ for some function $\omega$ is a linear functional of $f_B(\cdot,w)$. In particular, we obtain the distribution of the potential outcomes which is relevant also in the context of quantile treatment effects.
The asymptotic distribution result below requires the following additional notation and assumptions. We introduce the sieve variance
where
using the notation $\rho(t)=\exp(it(Y-g(S)))-h(X,t;g)$ and the matrix valued function $R(t)=Q^{-1/2}\int_{\mathbb{R}^{d-1}} (\mathcal F q^K)(t, tx)p^K(x)'dx$. We replace the sieve variance $\textsl{v}_K$ by the estimator
where the matrix $\Sigma$ is replaced by
with $\widehat \rho_j(t)=\exp(it(Y_j-\widehat g(S_j)))-\widehat h(X_j,t;\widehat g)$. In order to derive the asymptotic distribution of our estimator we require addition assumptions. Below, we denote the inverse Fourier transform by $(\mathcal F^{-1}\phi)(u)=(2\pi)^{-d}\int_{\mathbb{R}^d}\exp(-itu)\phi(t)dt$.
Assumption (ref) (i) implies a lower bound on the sieve variance which we require to achieve asymptotic distribution results of the estimator. This condition implies that the sieve variance attains the lower bound $\textsl{v}_K(w)\geq C_\Sigma\, \|\ell\big(q^K(\cdot -\mathbf g(w))\big)' Q^{-1/2}\|^2/\lambda_K$ for some constant $C_\Sigma>0$. For instance, when $\ell(\cdot)$ is the point evaluation functional and $Q$ is a diagonal matrix with polynomial decay of order $-2\alpha/d$ we obtain $\textsl{v}_K(w)\geq C_\Sigma\, K^{(2\alpha+d)/d}$ provided that $\lambda_K$ is bounded from below and $\|q^K(a -\mathbf g(w))\|^2\geq K$ which holds at most points $a\in\mathcal A$ (see belloni2012). Consequently the lower bound of the sieve variance corresponds to the mildly ill-posed case in chen2013. Assumption (ref) (ii) imposes a stronger rate requirement on $K$ than the one required for consistency. Assumption (ref) (iii) specifies a pointwise sieve approximation error and has the interpretation of an undersmoothing condition, which is required to ensure that the approximation bias is asymptotically negligible. Assumption (ref) (iv) restricts the sieve space to be linear and, in particular, that constraints on density estimation, such as positivity, are not binding. If such constraints are binding asymptotically then such shape restriction can lead to non-normal distributions. Assumption (ref) (v) is automatically satisfied if the basis under consideration is given by Hermite functions.
Assumption (ref) strengthens the rate conditions imposed in Assumption (ref) and is required for consistent estimation of the sieve variance $\textsl{v}_K(w)$. The next result establishes the asymptotic distribution of the estimator $\ell(\widehat f_B(\cdot,w)) $.
A direct implication of the previous results concerns inference on functionals of the VRS density $f_{B_1}(\cdot,w)$. The next result is based on plug-in series estimator (ref) using tensor product Hermite functions. In this case, the sieve variance simplifies to
where the matrix $\Sigma_1$ is coincides with $\Sigma$ except that $K$ replaced by $K_1$ and $R(t)$ replaced by $\int_{\mathbb{R}^{d-1}} b_{K_0}(t)\widetilde q^{K_1}(tx)p^{K_1}(x)'dx$, where the function $b_{K_0}$ is given in Remark (ref). Now if $K_0=O(1)$ a lower bound for the sieve variance is $\textsl{v}_{K_1}(w)\geq C_{\Sigma_1} \,\|\ell\big(q^{K_1}(\cdot -\mathbf g(w))\big)\|^2/\lambda_{K_1}$ for some constant $C_{\Sigma_1}>0$. When $ \lambda_{K_1}$ is uniformly bounded away from zero this corresponds to the usual lower bound in well posed estimation problems, see newey1997 or belloni2012. An estimator $\widehat{\textsl{v}}_{K_1}(w)$ of the sieve variance $\textsl{v}_{K_1}(w)$ is obtained by replacing the covariance matrix $\Sigma_1$ by $\widehat \Sigma_1$ which is analog to the definition of $\widehat \Sigma$. The next result is an immediate consequence of Theorem (ref) and hence its proof is omitted.
This subsection provides a bootstrap procedure to construct uniform confidence bands for $f_B(\cdot,w)$ and establishes asymptotic validity of it. The multiplier bootstrap procedure is as follows. Let $(\varepsilon_1,\ldots,\varepsilon_n)$ be a bootstrap sequence of i.i.d. random variables drawn independently of the data $\{(Y_1,X_1,W_1), \ldots, (Y_n,X_n,W_n)\}$, with $\operatorname{\mathbb{E}}[\varepsilon_j] = 0$, $\operatorname{\mathbb{E}}[\varepsilon_j^2] = 1$, $\operatorname{\mathbb{E}}[|\varepsilon_j|^{3}] < \infty$ for all $1\leq j\leq n$. Common choices of distributions for $\varepsilon_j$ include the standard Normal, Rademacher, and the two-point distribution of mammen1993. Further, $\mathbb{P}^*$ denotes the probability distribution of the bootstrap innovations $(\varepsilon_1,\ldots,\varepsilon_n)$ conditional on the data. For any $w\in\mathcal W$, we introduce the bootstrap process
Here, we use the notation $\textsl{v}_K(b,w)$ and $\widehat{\textsl{v}}_K(b,w)$ for the sieve variance and its estimator given in (ref) and (ref) in case of point-evaluation functionals. Let $\mathcal C$ be a closed subset of $\mathbb R^d$. Let $\Delta$ be the standard deviation semimetric on $\mathcal{C}$ of the Gaussian Process $\mathbb Z(b,w) = q^K(b-\mathbf g(w))'Q^{-1/2}\mathcal Z/\sqrt{\textsl{v}_K(b,w)}$ with $\mathcal Z\sim \mathcal N(0,\Sigma)$ defined as $\Delta_w(b_1,b_2) = (\operatorname{\mathbb{E}}[(\mathbb Z(b_1,w) - \mathbb Z(b_2,w))^2])^{1/2}$, see e.g. Vaart2000. To do so, we introduce the notation $N(\mathcal{C},\Delta,\epsilon)$ for the $\epsilon$-entropy of $\mathcal{C}$ with respect to norm $\Delta$.
Assumption (ref) is similar to ChenChristensen2015 who establish asymptotic validity of uniform confidence bands in nonparametric instrumental variable estimation. Assumption (ref) (ii) is a mild regularity assumption, see also ChenChristensen2015 for sufficient conditions. Assumption (ref) (iii) strengthens the rate conditions imposed on the dimension parameter $K$ and imposes a uniform sieve approximation error.
The next theorem establishes the validity of the bootstrap for constructing uniform confidence bands for the VRC density $f_B(\cdot,w)$. The proof of this result is based on strong approximation of a series process by a Gaussian process, and uses an anti-concentration inequality for the supremum of the approximating Gaussian process obtained in chernozhukov2014. For sieve minimum distance estimation in nonparametric instrumental variable estimation this is also exploited by ChenChristensen2015.
Theorem (ref) establishes consistency of the sieve score bootstrap procedure for estimating the critical values of the uniform sieve t-statistic process for the VRC model. The result also contributes to the literature on ordinary RC models, as up to now, only asymptotic validity of pointwise confidence intervals is established and complements the results of dunker2017 who proposed tests for qualitative features of the ordinary RC's density.
This section provides a finite sample analysis of the proposed estimator. Subsection (ref) presents the finite sample performance of the proposed estimator in Monte Carlo simulations. Subsection (ref) applies the procedure to analyze heterogeneity in income elasticity of the willingness to pay for rent in an empirical illustration.
This subsection presents the finite-sample performance of the estimator of the varying random slope in a Monte Carlo simulation study. The experiments use a sample size of $1000$ and $1000$ Monte Carlo replications in each setting. In each experiment, i.i.d. draws of regressors $(X,W)$ are generated from
where the variance $\sigma^2$ is varied in the experiments. Realizations of $A$ are generated independently of $(X,W)$ as follows: The random slope parameter $A_1$ is drawn either by a mixture of normal distributions, i.e., $\mathcal{N}\left( -1.5,1\right) $ and $\mathcal{N}\left( 1.5,0.5\right) $ with weights $0.5$, or by the Gamma distribution $\Gamma(3,1)$. In each case, the random intercept is generated independently of the random slope by $A_0\sim\mathcal{N}( 0,1)$. Realizations of the dependent variable $Y$ are obtained by
with varying random coefficients
where in the experiments either $g_1(w)=\sin(w)$ or $g_1(w)=\exp(|w|)-1$.
In this simulation study, a linear sieve space with Hermite functions is used and no additional constraints are imposed. Consequently, the resulting estimator of the density of $B_1$ is of closed form as given in (ref). Numerical integration is used to compute the integrals in equation (ref) based on the Adaptive Gauss-Kronrod quadrature (using the pracma package in R). We keep the dimension for the random intercept fixed with $K_0=1$. The conditional characteristic function $h$ is estimated via series least squares as in equation (ref) where the basis functions coincide with the Hermite functions of dimension $K_1$. For the estimation of the varying coefficient function $g_1$ we use quadratic B-spline bases with interior knots placed evenly. More precisely, we use two interior knots when using the function $g_1(w)=\sin(w)$ and four knots in the case of $g_1(w)=\exp(|w|)-1$.
As weighting measure $\nu$, we choose the log-normal distribution as motivated in Example (ref), more precisely, $\nu$ is given by $\textsl{Lognormal}(0,\sigma_\nu^2)$ where $\sigma_\nu=1/4$.
For the implementation of the estimator it is required to choose the dimension parameter $K_1$, given that we have fixed the dimension $K_0=1$. In the Monte Carlo experiments we choose these parameters to minimize the mean integrated squared error. Table (ref) reports the mean integrated squared error of the estimator $\widehat f_{B_1}(\cdot,w)$ of $f_{B_1}(\cdot,w)$ when $w=0$ which is given by
where the expectation is over the mean of all simulations.
Table (ref) depicts values of $\text{MISE}(\widehat f_{B_1})$ for different values of $\sigma^2$ and different functions $g_1$. In each case, the $\text{MISE}(\widehat f_{B_1})$ is provided for different sieve dimensions $K_1$, where the minimized value (w.r.t. the sieve dimension $K_1$) is shown in bold letters. From Table (ref) we see, not surprisingly, that the $\text{MISE}(\widehat f_{B_1})$ decreases as the variance of $X$ becomes larger. In particular, the optimal choice of the dimension parameter $K_1$ increases as the variance $\sigma^2$ becomes larger. This is in line with the rate of convergence as derived in Theorem (ref), see also the discussion thereafter. The MISE when $g_1(w)=\exp(|w|)-1$ is larger as in the other case which is due to irregularity of the function $g_1$ at zero.
Figure (ref) shows the estimation results when the true distribution of $B_1$ is given by an equally weighted mixture of $\mathcal{N}\left( -1.5,2\right) $ and $\mathcal{N}\left( 1.5,1\right) $. The solid line depicts the true density $f_{B_1}(\cdot,w)$ with $w=0$, the thick dashed line depicts the median of the estimators, and the thin dashed lines show the $95\%$ percent pointwise confidence intervals. The number of Hermite basis functions $K_1$ is chosen to minimize the $\text{MISE}(\widehat f_{B_1})$ as shown in Table (ref), i.e., $K_1=5$ when $\sigma^2=1$ and $K_1=6$ when $\sigma^2=2$. We see that as the variance $\sigma^2$ increases (lower panel of the figure) the median of the estimators is closer to the true density. Although there is no positivity constraint imposed, the closed form median of the estimator together with their $95\%$ confidence intervals are non-negative between $-5$ and $5$. As the variance of $X$ increases the pointwise confidence intervals become more narrow, which is in line with the pointwise asymptotic theory. This indicates the difficulty of estimating random coefficients in the case where $X$ has light tails. Nevertheless, we see that the estimator performs well even if $X$ has a small variance and is far from heavy tailed. Figure (ref) also shows that the procedure is robust even against irregularities of the varying coefficient function, i.e., when $g_1$ coincides with $g_1(w)=\exp(|w|)-1$.
Figure (ref) depicts the estimator of $f_{B_1}$ in the case when $A_1$ is generated by the Gamma distribution $\Gamma(3,1)$. For the implementation of the estimator we use the same choice of tuning parameters as described above in the normal mixture case. Again we find the $95\%$ confidence interval and the median are more accurate when the variance of $X$ is increased from $1$ to $2$ and $g_1$ coincides with the sine function. In all cases, we see that the true density function lies outside of the confidence intervals for $b_1\in[7,9]$. This bias is due to a larger variance of $A_1$, i.e., $\mathop{{\mathbb V}ar}\nolimits(A_1)=3$, which implies that higher order Hermite functions are required to fully accurately capture the finite sample support of $A_1$.
We finally present estimation results when not only the random slope (again we consider $A_1\sim\Gamma(3,1)$) but also the random intercept is not normally distributed but $A_0\sim \mathcal U(0,1)$. We use the same implementation as above but normalize the VRS density estimator to integrate to one. Since we only use the Hermite basis function of order zero to account for the random slope the estimator is misspecified in this direction. The estimation results are shown in Figure (ref). From this figure we see that misspecification of the density of $A_0$ has only a minor effect on the accuracy of the VRS density estimator after normalization. This is in contrast to misspecification of the functional form of varying coefficient functions $g_l$ which can lead to severe nonlinear biases, see Example (ref).
In this subsection, the methodology is applied to analyze heterogeneity in income elasticity of demand for housing. Heterogeneity plays an important role in classical consumer demand and might be driven by unobserved heterogenous preferences. In the empirical illustration we use data from the German Socio-Economic Panel (SOEP). While the SOEP is a longitudinal survey we restrict ourselves to the year 2013. We only consider individuals who do not have missing observations in rent, income, and size of the apartment, which results in a sample of size $n=7230$.
We are interested in assessing the heterogenous effect of household income on households' willingness to pay for rent. Formally, let us consider the empirical VRC model
and
where $Y$ denotes the log monthly rent, $X$ denotes the logarithm of the household net income per month, and $W$ is the logarithm of the size of the housing unit in square meters.\footnote{As stated in harrison1978, rental prices reflect the market's current valuation of housing attributes, while housing values reflect expectations about future as well as present housing conditions. Hence, conceptually it is more appropriate to use rental prices when estimating hedonic functions for housing demand.} The empirical VRC model thus imposes functional forms rather than letting the conditional distribution of unobserved heterogeneity given housing characteristics unrestricted. The following table provides summary statistics of the relevant variables.
The interpretation of $B_1$ is that of a heterogeneous elasticity. Independence of the heterogeneous income elasticity of demand and income itself might be difficult to justify if no additional covariates are included to explain $B_1$. We compute the variance of $B_1$ from the empirical analog of $\operatorname{\mathbb{E}}[(Y-g_1(W))^2 X]-\operatorname{\mathbb{E}}[(Y-g_1(W))X]^2$ which yields the value of $0.0431$, where $g_1$ is estimated using B-splines as described below. The log size of housing $W$ explains much of the variation in $B_1$, i.e., if $g_1\equiv 0$ then variance of $B_1$ is given by $0.5899$. The small variance does not prevent estimating the density of $B_1$ using global basis functions, such as Hermite functions, because we can transform the model such that $B_1$ has , e.g., variance one and then back--transform the density function of $B_1$. Here, $g_1$ is replaced by a B-spline estimator as explained below.
The estimator $\widehat f_{B_1}$ is implemented as described in the previous section. The number of Hermite functions used is $K_0=1$ and $K_1=7$. The weighting measure $\nu$ is again given by $\textsl{Lognormal}(0,\sigma_\nu^2)$ with $\sigma_\nu=1/4$, as in the Monte Carlo section. For estimation of the functions $g_0$ and $g_1$ we use again quadratic B-spline bases functions with three interior knots and follow Example (ref). Figure (ref) depicts the B-spline estimators for the varying coefficient functions $g_0$ and $g_1$. We see that both estimators are nonlinear on the support of $W$.
For the bootstrap uniform confidence bands, we consider one representative sample and generate the bootstrap innovations $\varepsilon$ according to the two-point distribution suggested by mammen1993, i.e., $\varepsilon$ equals $(1-\sqrt 5)/2$ with probability $(1+\sqrt 5)/(2\sqrt 5)$ and $(1+\sqrt 5)/2$ with probability $1-(1+\sqrt 5)/(2\sqrt 5)$. Based on the estimator we generate the bootstrap process $\mathbb Z^*(\cdot|w)$ as described in Subsection (ref). The results are based on $1000$ bootstrap iterations.
Figure (ref) depicts the estimator for the density of $B_1$ evaluated at the mean of $W$ which is $w=4.253$. Note that $B_1$ can be directly interpreted as heterogenous marginal effect. From this figure we see that the estimated density has support between $-0.2$ and $0.8$. The uniform confidence bands show that the support is significantly positive (at $0.05$ nominal level) only at $-0.05$ and $0.6$. The estimated density is clearly not symmetric. We also see that the density is positively skewed and is more heavy tailed on the right hand side. This is reasonable as one would expect the response of a marginal increase of income to be skewed. It is also interesting to see that the $95\%$ uniform confidence sets are bounded away from zero.
This paper analyzes heterogeneity in VRC models. This model generalizes ordinary RC models by including nonlinearities in observed characteristics, which might stem, for instance, from measurement errors or control function residuals. A novel estimator of the VRC density based on weighted sieve minimum distance is proposed. Under semiparametric restrictions on the random intercept, our estimator of the VRS density is not affected by the ill-posedness that is associated with the nonparametric estimation of the joint VRC density. We establish novel inference results, such as uniform confidence bands, to adress uncertainty in VRC density estimation which goes beyond what has been shown in ordinary RC models. We find that finite sample estimation results are surprisingly stable when the sieve space is spanned by Hermite functions. This also advocates the use of the proposed methodology in the context of ordinary RC models. Finally, the methodology is applied to estimate the density of heterogeneous income elasticity of demand for housing, which is shown to be highly skewed. The proposed estimator can also be extended to include nonlinear index functions as in lewbel2017unobserved. Yet the analysis of its asymptotic properties is left to future research.