Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,000 characters · 11 sections · 87 citation commands
Estimation of Conditional Random Coefficient Models using Machine Learning Techniques
In recent years microeconometric models aimed at capturing complex heterogeneity in individual behavior. In particular, it has become relevant to include observed and/or unobserved heterogeneity when modeling (average) partial effects in regression models.
An important model that accounts very flexibly for such heterogeneity is the nonparametric random coefficient model
where $(B_{0}, B_{1})$ is a vector of $p+1$ random variables and $W \in \mathbb{R}^p$ is a vector of regressors. Let $X$ be a set of additional control variables that may be related to $(W,B_{0}, B_{1})$. $B_{1}$ is the individual effect of a change in $W$ on $Y$ and the distribution of $B_1$ reflects the heterogeneity of individual effects in the population.
The model is nonparametric in that it does not impose distributional assumptions on partial effects $B_1$ and the nuisance $B_0$. The primary goal for this class of models is to identify and estimate the entire distribution of the vector of random coefficients $(B_0, B_1)$ such as the joint density function. The marginal density of the random slope parameter $B_1$ is of special interest in economic applications as $B_1$ can be interpreted as an average partial effect of a change in $W$ on the outcome. Its distribution reflects the heterogeneity of this effect in the underlying population. If $W$ is an exogenous treatment then the expected value of $B_1$ corresponds to an average treatment effect, its conditional expectation to an conditional average treatment effect and so on.
A crucial identifying assumption in the literature is full independence of covariates $W$ and random coefficients $(B_0,B_1)$. This is satisfied if $W$ is randomly assigned, such as in experimental data settings, and the marginal slope density captures the entire heterogeneity of a partial effect in the population. This knowledge does not allow to link the heterogeneity to any observable characteristics. For instance, the shape of the RC-density may vary across observable individual characteristics. Learning the RC-density conditional on a set of control variables $X$ provides additional insight on how the shape of the heterogeneity varies across subpopulations with differing observable characteristics. This allows to disentangle heterogeneity in observable and unobservable heterogeneity.
Furthermore, when dealing with observational data there is always room for a potential dependence between $W$ and control variables $X$. The random intercept $B_0$ subsumes the effects of $X$ on $Y$ which violates full independence between $W$ and random coefficients. In this work the identifying restriction can be weakened to allow for conditional independence, i.e. $B \Perp W |X$, which corresponds to a selection-on-observables assumption. In most economic applications the set of control variables $X$ is of considerable size or even high-dimensional and there is generally no prior knowledge about which variables in $X$ drive heterogeneous partial effects. Modern Machine Learning (ML) methods allow to deal with large dimensional set of controls and perform some form of variable selection to identify those elements of $X$ inducing different shapes of heterogeneity. Recently ML-techniques have proven useful in relating features of the distribution of partial effects $B_1$ to additional observable characteristics, such as in the estimation of conditional average treatment effects, also referred to as heterogeneous treatment effect. In this work, I link the entire distribution of random coefficients to observable characteristics by studying a conditional random coefficient model. This allows to uncover (i) which variables generally drive heterogeneity in partial effects and (ii) how the distribution of partial effects varies across subpopulations with different observable characteristics. \\
First I begin by providing an identification statement for the conditional RC-density. Deriving from this identification statement I can formulate a sieve approximation to the RC-density conditional on a fixed value of $X$. This sieve approximation has a closed form expression and each sieve coefficient can be expressed as a conditional expectation function of some nonlinear transformation of $Y$ and $W$ varying with controls $X$.
Generic ML-methods can be used to estimate this set of conditional expectation functions. Various practical considerations to be outlined later require to orthogonalize both the outcome $Y$ and treatment $W$ using ML technqiues before estimating the sieve coefficients itself. This requires the use of iterated sample splitting to deal with nested ML-steps.
Following the outline of the estimation strategy I derive the $L_2$-convergence rate of the final conditional RC-density estimator. There convergence rate is crucially determined by the asymptotic properties of the ML-methods employed. Given that the slowest ML-estimator employed converges at a polynomial rate the $L_2$-convergence rate of the RC-density estimator is by a factor slower than that of the slowest ML-estimator. This factor hinges on the degree of ill-posedness of the underlying random coefficient problem and the overall smoothness of the density.
In addition to the asymptotic properties of the estimator, I introduce a cross-validation strategy to inform the choice of tuning parameters.
Finally I apply the estimator to study behavioral heterogeneity in an economic experiment on portfolio choice. Survey respondents of the german socio-economic panel (SOEP) are asked to invest an hypothetical monetary amount into a riskfree asset or into an risky asset with payoff depending on the return of a stock market index. I study the effect of stock market beliefs on the investment decision with a random coefficient model. I find a bi-modal RC-density that reflects the presence of two types in the population. One type complies with economic theory and the stock market beliefs have a positive impact on the amount invested in the risky asset. For a second type this is not the case and the effect is centered around zero. This type-division prevails when varying the values of controls. This suggests that the assignment to types is largely based on unobservable characteristics not in the data. An exemption is age, as for a subpopulation of elder respondents the share of non-compliers is substantially larger.
\paragraph{Related literature} Identification and estimation of the nonparametric random coefficient model is studied in beran1992, beran1994, beran1996 and hodermam_RC_2010. See also masten17_RC for a refined identification result. All of these works operate under the assumption that random coefficients and regressors are fully independent. Breunig_VRC considers a so called varying random coefficient model where each random coefficient is made up of a nonparametric function of control variables and an additively separable random component which is fully independent of controls. Further the number of control variables affecting random coefficients is effectively smaller than the number of random coefficients itself whereas in the present setting the number of controls is allowed to be larger. Sieve estimation for random coefficient models is used for the testing procedure in BreunigHoder18 and in Breunig_VRC.
Machine Learning estimation and econometric models have been paired frequently in recent years. ChernPostSelec, ChernDebiased17 and chernozhukov2020debiased study estimation and inference on parameters and linear functionals of parameters in high-dimensional linear models. Thereby highlighting the importance of iterated ML estimation and sample splitting for achieving consistency and asymptotic normality of parameter estimates.
Of particular importance for this work is the estimation of heterogeneous treatment effects using machine learning methods as considered in AtheyWager18 and grf, see also the references therein. This is due to the fact that a heterogeneous treatment effect can be viewed as the conditional expectation function of a random slope coefficient. In contrast, the conditional RC-density studied here is informative about the entire distribution of a treatment effect in a given (sub-)population. This also includes learning the form of unobservable heterogeneity which remains otherwise unknown when only the mean of a random coefficient is studied. ChernGenericML considers identification and estimation of particular subfeatures of the conditional expectation function of a random slope.
The theory of the conditional RC density estimate developed here holds for generic machine learning techniques. However I make use of the causal forest algorithms of grf and the popular ML-tool of random forests introduced by Breiman01 in the implementation of the estimator and in the asymptotic theory. The asymptotic theory of random forest estimators has been studied in scornet2015, WagerWalther, AtheyWager18 and grf. \\
The remainder of this paper is organized as follows. Section (ref) introduces the main model and discusses identification of conditional RC-densities. Section (ref) outlines the estimation strategy. Section (ref) presents the asymptotic properties of the estimator and Section (ref) asymptotic inference on the conditional RC density estimates. Section (ref) addresses auxiliary topics, i.e. marginal density estimation, variable importance measures and the choice of tuning parameters. Section (ref) contains a Monte Carlo simulation study. Section (ref) contains an empirical application of the estimation procedure using survey data. Section (ref) concludes.
This paper considers the following random coefficient model
where $B=(B_0, B_1)$ consists of two scalar random variables and $W$ is a scalar regressor of interest to the researcher. Here $B_1$ is the average partial effect of a change in $W$ on the outcome $Y$. Within the model framework there exists a vector of additional covariates $X \subseteq \mathbb{R}^d$ that may affect both the random coefficients as well as the regressor $W$. The goal of this section is to give conditions under which we can identify the conditional random coefficient density $f_{B|X=x}$. That is the random coefficient density for a given subpopulation with characteristics $X=x$. Note that the results of this paper can be readily extended to cover the case where $W$ is multivariate, though for conciseness I focus on the scalar $W$ case which is of the most practical relevance.
The model ((ref)) can be interpreted as a reduced form of a more general multivariate random coefficient model with the random intercept absorbing all the (heterogenous) effects of other covariates on the outcome. Without loss of generality the random coefficients satisfy
which illustrates that ((ref)) nests many popular mean regression models. For instance it can be viewed as an extension of Robinson88 with a random coefficient instead of a deterministic one. \\ For obtaining identification the following assumptions are imposed.
Assumption (ref) (i) requires that the regressor of interest $W$ is independent of random coefficients conditional on controls $X$. This restriction allows for some dependence between $B$ and $W$ and is thereby weaker then the full independence assumption typically encountered in random coefficient models, see beran1992, hodermam_RC_2010 and masten17_RC. It can also be interpreted as an exogeneity condition on the regressor $W$. If $W$ is a (quasi-)experimental intervention then (i), or more precisely the part $B_0 \Perp W \;|\;X $, corresponds to a selection on observables assumption. It is also one of the main assumptions for identifying heterogenous treatment effects as in the RC model in Section 6 of grf.
Assumption (ref) (ii) strengthens the common large support restriction from the random coefficients literature (see again e.g. hodermam_RC_2010) to also hold conditional on $X=x$. The assumption rules out the case where $W$ is a deterministic function of $X$ and may be problematic if otherwise certain realizations $x$ provide strong information about $W$. A workaround for this assumptions is provided by the varying RC model of Breunig_VRC. There however functional form restrictions on the random coefficients need to be made which outlines a tradeoff between the above full support assumption and further restrictions on the random coefficient model. If this assumption is violated for some realizations of $X$ then nevertheless identification for different $x$-values which satisfy (ii) can be established. masten17_RC discusses identification in the case of bounded support of the regressor $W$. Taking this results in to account the following identification result can be formulated.
The proof of Lemma (ref) immediately follows from extending classical identification results for random coefficient models as synthesized in masten17_RC to the conditional case. This section concludes with a remark on identification for the case where $W$ is a discrete, binary variable.
This section introduces the estimation strategy for conditional RC-densities. Subsection (ref) outlines the principal idea and introduces the main notation whereas Subsection (ref) is of main practical relevance. There a demeaned random coefficient model is discussed and further details on machine learning estimators and sample splitting are provided. Before moving forward the following general notation needs to be introduced. Let
denote the characteristic function of $Y$ conditional on $X=x$. The Fourier transformation $\mathcal{F}$ and the inverse Fourier transformation $\mathcal{F}^{-1}$ are defined as
for some functions $f,g: \mathbb{R}^d \to \mathbb{R}$. The operators $\mathcal{F}: \mathbb{R}^d \to \mathbb{C}^d$ and $\mathcal{F}^{-1}: \mathbb{C}^d \to \mathbb{R}^d$ relate the characteristic function of a random variable to its probability density function, given the latter exists. For some random variable $A$ with density $f_A$ it holds that $\phi_{A}(t) = (\mathcal{F}f_{A})(t)$ and vice versa $(\mathcal{F}^{-1}\phi_{A})(a)= f_{A}(a)$.
The essential implication of Assumption (ref) that will be leveraged for estimation is the identity
which holds for every $x$ in the support of $X$ and $t,w \in \mathbb{R}$. The identity in turn implies the following $L_2$-condition
where $v,\mu$ are arbitrary probability measures on $\mathbb{R}$ that are discussed later in more detail. Following Breunig_VRC the $L_2$-criterion in ((ref)) can be used to construct a sieve estimator of the density $f_{B|X=x}$.
To this end let $q^K = (q_1,\dots, q_K)$ denote a $K=K(n)$-dimensional vector of known basis functions that span the linear sieve space $\mathcal{B}_K=\{\phi(\cdot)=q^K(\cdot)'\pi \}$. As $f_{B|X=x}$ is bivariate, $q^K$ typically is a tensor product of univariate basis functions, i.e. $q^K(b_0, b_1)=q^{K_1}(b_0) \otimes q^{K_2}(b_1)$ with $K=K_1 \cdot K_2$.
If the characteristic function $\phi_{Y|X, W}(\cdot|x,w)$ were known then a sieve estimator of the conditional random coefficient density is
which has the following closed form expression
and where
Note that the estimator in ((ref)) is not feasible as $\phi_{Y|X, W}(t|x,w)$ is not known. Breunig_VRC proceeds by replacing the unknown characteristic function with a nonparametric plug-in estimate. A general problem is the presence of the possibly large-dimensional set of controls $X$ that cannot be reduced a priori in most practical applications. Breunig_VRC studies a varying random coefficient model which puts additional structure on the relationship of random coefficients and controls and where $X$ is low-dimensional.
Modern Machine Learning (ML-)estimators are well suited for the estimation of conditional expectation functions like $ \phi_{Y|X, W}(t|x,w)=\mathop{{\mathbf E}\null}\nolimits[\exp(itY) |X=x, W=w]$ in the presence of possibly high-dimensional controls $X$. However an additional issue arises in that we would require to perform different Machine Learning steps over a continuum of values for $t$.
In order to enable the use of Machine Learning techniques a different strategy is needed. Further rearranging of ((ref)) yields
where the operator $T(w, y) := \int_{\mathbb{R}} (\mathcal{F}q^K )(-t, -tw)\exp(ity)d\nu(t)$ is short-hand for the nonlinear mapping $ T: \mathcal{W}\times \mathcal{Y} \to \mathbb{R}^K $. The operator $T$ can be computed via numeric integration and thus I consider it deterministic for the remainder of this paper. Also for clarification let $T(W,Y)=(T_1(W, Y), \dots, T_K(W, Y))$ with functions $T_k:\mathcal{W}\times \mathcal{Y} \to \mathbb{R}$, for $k=1,\dots, K$. Further define the functions
for $k=1,\dots, K$ with $\pi(x,w)=(\pi_1(x,w), \dots, \pi_K(x,w)$. A distinguishing feature of the sieve approximation in ((ref)) is that sieve coefficients are the relevant quantity that varies in $X$.
Finally I construct a feasible estimator by replacing $Q$ with a sample mean and $\pi(x,w)$ with a vector of Machine Learning estimates. The resulting RC-density estimator is
where $\widehat \pi$ is a generic ML-estimate of the unknown function $\pi$ and
where the $(M_i, N_i)$'s are $n$ Monte Carlo draws from the probability distributions $\mu$ and $\nu$ that are specified by the researcher. In principal $Q$ can be calculated directly via numerical methods for most measures $\mu, \nu$, however this representation is introduced here, as we will later consider the case where $\mu$ is the distribution of $W$ and use sample realizations $W_i$ instead of the generated $M_i$.
Notice that ((ref)) is a direct estimate of the closed-form sieve projection in ((ref)). The sieve coefficients can be expressed in terms of different conditional expectation functions which can in turn be conveniently estimated by generic machine learning routines even if the set of controls $X$ is high-dimensional.
\paragraph{Choice of weighting measures}
The choice of the measure $\mu$ leaves room for further simplification of the estimator in ((ref)). If we choose $f_{W|X=x}$ as the density of the weighting measure $\mu$ then it holds that $\int_{\mathbb{R}} \pi(x,w)d\mu(w) = \mathop{{\mathbf E}\null}\nolimits[T(W,Y)|X=x]$ and we only need to consider ML-estimation of a conditional expectation function and there is no need for additional weighting of this ML-estimator. This does slightly ease computation and simplifies the asymptotic analysis as asymptotic properties of ML-estimators of conditional expectation functions are readily available. The properties of the transformed ML-estimator $\int_{\mathbb{R}} \widehat \pi(x,w)d\mu(w)$ used to calculate the estimator in ((ref)) have not been studied explicitly.
A caveat of choosing $f_{W|X=x}$ is that the matrix $Q$ will vary in $x$ which is problematic from both the computational as well as the theoretical stance\footnote{Each of the $K^2$ elements of $Q$ would need to be estimated by a ML-step. Further the asymptotic properties of a machine-learned matrix whose dimensions increase with the sample size remain unclear.}. A possible workaround is to orthogonalize $W$ and is detailed in the next section. However, this workaround will slow down the convergence of the RC-density estimator, as will be shown in Section (ref). Hence choosing $d\mu/dw=f_{W|X=x}$ as weighting measure is problematic.
In any case choosing a $\mu$ that is related to the distribution of $W$ is appropriate. ML-estimators like $\widehat{\Pi}(x,w)$ perform best for points from the center of the distribution and will be less accurate in the tails. Moving forward I will focus on the case where $d\mu(w)/dw=f_W(w)$ which automatically weighs down areas where the ML-estimates may be less accurate. This respects the finite sample behavior of each machine learned sieve coefficient and reduces the problem of estimating $Q$ to a simple sample mean. The additional weighting of the ML-estimate $\widehat{\pi}(x,w)$ is a minor issue compared to the difficulties arising from alternative choices for $\mu$. \\
Following Breunig_VRC, I choose $\nu$ to follow a $\text{lognormal}(0, \sigma_t)$ distribution where $\sigma_t > 0$ is then the second tuning parameter to be chosen by the researcher, along with $K$. This particular choice of weighting measure works well in the settings of Breunig_VRC and also in the simulations and applications in this work. Theoretical justifications are given in BreunigHoder18, Breunig_VRC and in Section (ref) but these do not preclude other choices of weighting distributions. \paragraph{Choice of sieve basis functions} Throughout this paper I again follow Breunig_VRC and choose Hermite functions as sieve basis $q^K$.
Hermite functions are a $L_2$-basis and have appealing theoretical properties. They are eigenfunctions of the Fourier transformation and satisfy $\mathcal{F}q_k(a,b)=\sqrt{2\pi}i^{k-1}q_k(a,b)$. This property simplifies the computation of the estimator considerably. A drawback of Hermite functions is that most of the support concentrates around zero even if $K$ is moderately large. Thus any moderately sized sieve approximation will fail to be a good approximation of a density function that is centered away from zero and/or has a particularly large support. This is a major motivation for considering a demeaned random coefficient model in the next subsection.\\
This section discusses estimation of a demeaned version of the random coefficient model in ((ref)). The reason is that if marginal densities of $B_0$ and $B_1$ are centered away from zero, estimation of the bivariate density function with a Hermite function sieve will require a possibly large choice of $K$ and thus prohibitively many ML-steps. The computational cost associated with each ML-step and the general ill-posedness of the RC-density estimation problem lead to a strong preference for a coarse choice of $K$.
Second, as the random slope density is of particular interest in economics, specific ML-routines have already been developed which provide high quality estimates for the conditional expectation $\mathop{{\mathbf E}\null}\nolimits[B_1 |X]$, see AtheyWager18 and grf. By demeaning we can separate estimation of the conditional expectation from the remaining conditional shape of the RC-density. This enables the use of ML-tools that are tailored to the specific predictive tasks such as the estimators in grf for conditional expectation function of $B_1$. Further this direct estimators of the conditional expectation will perform better than those one can infer from an indirect estimate via integrating the entire conditional density function. \\
We can reformulate the original RC-model ((ref)),
and estimate the joint density of the demeaned random coefficients $f_{A|X=x}$ with the procedure outlined in the previous section. To this end let $\beta(x)=\mathop{{\mathbf E}\null}\nolimits[B|X=x]$ with $\beta(x)=(\beta_0(x), \beta_1(x))$ denoting the conditional expectation of the intercept and slope respectively. Further let $m(x,w)=\mathop{{\mathbf E}\null}\nolimits[Y|X=x, W=w]$. Then the closed form of the sieve approximation analogous to the previous section is
with
Further define
with $\widehat m$ denoting a generic (ML)-estimator for the unknown function $m$. By choosing $d\mu(w)/dw=f_W(w)$ an estimator is
where
and $\widehat{\Pi}_{dm}(x)$ is an ML-estimate of $\Pi_{dm}(x)$. The conditional expectation $\mathop{{\mathbf E}\null}\nolimits[T(W, Y- \widehat m(X,W )) | X=x, W=w]$ can be conveniently estimated with ML methods, but it remains to construct an estimate for the quantity $\int_{\mathbb{R}}\mathop{{\mathbf E}\null}\nolimits[T(W, Y- \widehat m(X,W )) | X=x, W=w]d\mu(w)$. I suggest to use
where $\widehat \mathop{{\mathbf E}\null}\nolimits[T(W, Y- \widehat m(X,W )) | X=x, W=w]$ is an ML-estimator of the respective conditional expectation function and thus $\widehat{\Pi}_{dm}(x)$ is a sample average of different predictions from the ML-estimators over a hold-out sample of size $R$ which has not been used in the estimation before. Another possibility, that does not rely on a hold-out sample is to use a leave-one-out ML-estimator. In the applications and simulation studies I simply calculate ((ref)) on the entire sample observations for $W$. This is theoretically not valid, yet in simulations there is practically no difference between using ((ref)) on the entire sample for $W$ or an equally-sized hold out sample of $W$.
Therefore, I assume for the remainder of the paper that
so when studying the asymptotics of ((ref)) in the next Section, it is implicitly assumed that the rather slow ML-estimators dominate the asymptotic behavior of $\widehat{\Pi}_{dm}$, i.e. convergence of the sample mean to the integral is negligible compared to the convergence of ML-estimators.
The estimator ((ref)) nests several ML-estimates and it is therefore apparent that sample splitting is required to achieve consistency.
A particular requirement is that $\widehat m$ is calculated on a different sample than $\widehat \Pi_{dm}$ which takes $\widehat m$ as input. \\ The use of sample splitting is somewhat mandatory for nested ML-estimators, see ChernPostSelec.
The subsequent paragraph outlines the precise use of sample splitting along with a concise summary of the estimation procedure. \paragraph{Estimation Procedure} \\ \\ The sample is $(X_i, W_i, Y_i)$ with $i=1,\dots,n$. Set tuning parameters $K$ and $\nu$.
Then perform cross-fitting, i.e. iterate steps 2-4 a number of $M$ times to obtain $M$ different estimates $\widehat \Pi_{dm, m}(x)$ for $m=1,\dots, M$. Then aggregate these to a final conditional RC-density estimate
The cross-fitting procedure is optional, yet highly recommended as it stabilizes the estimates considerably. \\
In many works linking causal inference with machine learning methods, orthogonalization of treatments $W$ is often mandatory to achieve consistent estimation of causal effects, see e.g. ChernPostSelec or at least desirable for the performance of machine learning methods, see Section 6.1.1. in grf. Therefore this section ends with a brief discussion on how to handle the case of a demeaned $W$ in the estimation. In the next section it is shown that orthogonalization of the treatment $W$ is not innocuous in the RC model, as it slows down the convergence rate of the RC-density estimator.
The following section develops the asymptotic theory of the estimator in ((ref)) with its composite parts in ((ref)) and ((ref)). Estimation in the orthogonalized $W$ case outlined in Remark (ref) is also considered. The following notation is required. Let $\lambda_{min}(\Omega), \lambda_{\max}(\Omega)$ denote the smallest and largest eigenvalues of a matrix $\Omega$. Further define the $L_2$-norm $\mynorm{g}=\int |g(a)|^2 da$ and the weighted $L_2$-norm $\mynorm{g}_{\nu, \mu}=\int |g(t,x)|^2 d\nu(t)d\mu(x)$ for a generic, possibly complex-valued function $g$. To avoid confusion at some points, $\mynorm{\cdot}_E$ denotes the euclidean norm of a (complex) vector. $P_K g$ denotes the $L_2$-projection on a linear sieve space $\mathcal{B}_K$, i.e. $P_K g=\arg \min_{f\in \mathcal{B}_K} \mynorm{g - f}$. The relation $a \lesssim b$ is shorthand for $a \leq C\cdot b$ for some constant $C>0$.
The following set of assumptions is necessary.
Assumption ((ref)) (i) is satisfied for the most commonly employed sieve bases such as splines, fourier series or wavelets, see e.g.belloni2012 as well as for Hermite functions. Sufficient conditions for Assumptions (ii) and (iii) are given in Breunig_VRC in the absence of the measure $\mu$ in ((ref)). With the additional measure the eigenvalue decay $\tau_K$ will generally depend on $\mu$, i.e. in my preferred specification on the distribution of $W$. Simulations show that the decay is faster the more light-tailed the distribution of $W$ and the smaller its support. Condition (iii) is a typical assumption on the approximating properties of the basis functions and the parameter $\alpha$ is solely related to the smoothness of the density functions $f_{A|X=x}$ as the dimension of the random coefficient vector is not of interest in this analysis, see e.g. Chen07 for a review of approximation properties of various sieve bases across different smoothness classes. The remaining parts (iv) and (v) are standard regularity conditions on the density $f_{A|X=x}$, in particular (iv) imposes that for any $x$ the $L_2$-projection of $f_{A|X=x}$ has bounded coefficients.
Assumption $\ref{A::MLRates}$ (i) states an abstract upper bound for the pointwise convergence rates of various ML-estimates. Thereby one can abstract from considering different convergence rates for each ML-estimator by simply focusing on the slowest rate among those ML-estimators employed. Typically for any ML-method $\varphi< 1/2$. Part (ii) is a common rate restriction in the series estimation literature to achieve that $\mynorm{\widehat{Q}^{-1}-Q^{-1}} \to 0 $, see BellChern_Series. The last part (iii) is an additional rate restriction that is trivially satisfied if $\tau_K^{-1}= K^\gamma$ with $\gamma >1$, i.e. if the eigenvalue decay $\tau_K$ is sufficiently fast. It holds more generally if $\varphi$ is sufficiently small.
In general the particular convergence rates depend on factors such as the effective dimension and the smoothness of the conditional expectation function that is to be estimated. An additional aspect is the proper choice of tuning parameters for any ML-method that is applied and the precise notion of high-dimensionality, i.e. the rates at which $\dim(X)$ may go to infinity relative to the sample size. In order to derive these convergence rates additional restrictions will be required.
Thus Assumption $\ref{A::MLRates}$ abstracts from many theoretical and practical details of the ML-techniques employed. However, the complexity of any given ML-method makes the joint parameter choice of our model tuning parameters $K$ and $\sigma_t$ along with other parameters of the ML-routines impractical. Therefore I suppose that any ML-estimator used has been properly tuned by e.g. data-driven methods to the prediction task at hand. Thus the rate in Assumption (ref) can be viewed as the best rate achievable given a set of ML-estimators that are properly tuned to their respective estimation problem. This abstraction is in line with other works using generic machine learning techniques such as ChernPostSelec or ChernGenericML.
Below Assumption (ref) is discussed in the context of regression forests which will be used in the applied segments of this paper.
The Remark above gives convergence rates for random forest estimators of conditional expectation functions. Further Assumption (ref) imposes a convergence rate on $\widehat{\Pi}_{dm,k}(x)$ which is by definition ((ref)) itself an average of different random forest estimates by averaging over $W$. Thus its convergence rate can be expected to be faster then the rate of the random forest estimate $\widehat \mathop{{\mathbf E}\null}\nolimits[T(W, Y- \widehat m(X,W )) | X=x, W=w]$ for any fixed $w$. \\
Finally the following convergence rate result holds.
Here the decay of eigenvalues $\tau_K$ of the matrix $Q$ serves as a measure of ill-posedness. The estimation of the RC density is known to be an ill-posed problem which implies a slower convergence rate of estimators, see hodermam_RC_2010 or Breunig_VRC. If $\tau_K$ decays polynomially, e.g. $\tau_K \sim K^{- \gamma/2}$ then choosing $K\sim n^{\frac{2\varphi}{1+\gamma+\alpha}}$ balances bias and variance and results in the convergence rate
This shows that the convergence rate $\varphi$ of the generic machine learning estimators is slowed down by a factor $\alpha/(1+\alpha+\gamma) < 1$. This loss of speed is increasing in the eigenvalue decay parameter $\gamma$ and decreasing in the smoothness parameter $\alpha$. Note that the density considered here is bivariate and thus there is no explicit parameter for the number of random coefficients. If $W$ is multidimensional, its dimension enters the factor and further slows down convergence.
This rate result is not sharp in that the convergence rate can be improved for a given ML-technique. As $\varphi$ depends on tuning parameters specific to the chosen ML- technique, a joint choice of $K$ and the tuning parameters subsumed in $\varphi$ may improve the rate of convergence. However calculating these exact rates may prove difficult and in practice, joint tuning of $K$ along with the parameters of the specific ML-techniques leads to excessive computational costs. Further note that using the result from WagerWalther in Remark (ref) a rate for $\mathop{{\mathbf E}\null}\nolimits[\int \left[\widehat{f}_{B|X}(b|X) - {f}_{B|X}(b|X)\right]^2db]$ can be derived analogously. \\
In Remark (ref), estimation with orthogonalized treatment $W$ is considered. This is important for some ML estimators applied in causal inference. For the estimation of $\beta_1(x)=\mathop{{\mathbf E}\null}\nolimits[B_1|X=x]$, grf suggest orthogonalization of both the outcome $Y$ and treatment $W$ to improve the performance of the random forest routines involved, see Section 6.1.1 in grf. The subsequent Corollary shows that in the context of random coefficient models orthogonalization of $W$ is not innocuous, as it slows down the convergence rate of the RC-density estimator. This is in contrast to Theorem (ref) where sole orthogonalization of the outcome $Y$ does not result in a slower rate.
This slower convergence rate is due to the fact, that the "generated regressor" $W-\widehat{g}(X)$ appears within Hermite functions $q^K$. As is shown in the proof, the derivatives of Hermite functions $q^K$ diverge in $K$ and thus an additional $K$-term appears in the derivation of the convergence rate. This is a general issue, that does not seem to have been noticed so far. The convergence rate of any Hermite function sieve estimator of a RC density is slower when $W$ is a generated regressor.
This section discusses inference for the conditional RC-density estimator, in particular pointwise inference of the conditional RC density function estimate. Asymptotic normality of the RC-density estimator follows from asymptotic normality of ML-estimators linked with theory from the series estimation literature. For random forests such asymptotic normality results have been recently provided by grf and AtheyWager18. The main issue is how to establish the asymptotic covariance of the $K$ different ML-estimators, i.e. to obtain estimates for $\mathop{{\mathbf E}\null}\nolimits[\widehat{\Pi}_{dm,j}(x)\cdot\widehat{\Pi}_{dm,l}(x)]$ for any pair $1\leq j,l \leq K$. This issue can be overcome by introducing an additional layer of sample splitting. If we split the sample $\mathcal{R}$ on which sieve coefficients are learned into $K$ equally sized subsamples and use a different subsample for estimating each sieve coefficient then naturally these estimators are stochastically independent. Then the asymptotic variance of each sieve coefficient can be obtained by resigning to established variance estimators for the respective ML-method. This approach requires the block size $|\mathcal{R}|/K $ to be of meaningful size in practice, yet again cross-fitting, i.e. iterated sample splitting and averaging of estimates stabilizes the results and reduces any losses in efficiency. \\
More precisely, moving forward I amend Step 4. of the estimation procedure at the end of Section (ref).
Introduce the notation $r_p(n)=n^\varphi$. The resulting estimator does not coincide with the one in the previous section. The convergence rate is similar with the term $r_p(n/K)$ appearing in the convergence rate rather then $r_p(n)$ which is due to the fact that each sieve coefficients is learned on a sample of size $n/K$. Nevertheless in implementations the finite sample performance is comparable to the estimator in the previous section.
Assumption (ref) (i) establishes asymptotic normality of various ML-estimators. Such asymptotic normality results are standard for many ML-methods. For (honest) random forests such results have been established by AtheyWager18. For asymptotic normality results for different tree-based algorithms see the references therein and also grf. $\sigma_{dm,k}$ is the individual standard error of the $k$-th ML step and subsumes both the convergence rate and the residual standard deviation. Part (ii) of the above assumption imposes that the residual standard deviation in any of the $K$ ML-regressions is bounded away from zero and infinity. This assumption can be weakened to allow for standard deviations to diverge as $K$ grows at the cost of introducing an additional rate parameter that further strengthens rate restrictions. In Assumption (ref) (iii) $v_n(b,x)$ serves as the standard error of the conditional RC-density estimate. It holds that
which is bounded away from zero and approaching zero asymptotically under the rate restriction in (iv) which is required for consistency of the RC-density estimate, see Theorem (ref).
Part (i) and (iii) are the most intricate conditions in Assumption (ref) and typically involve additional regularity conditions and restrictions on the growth of $K$. The following Lemma gives conditions such that Assumption (ref) (i) and (iii) are satisfied for honest regression forests.
Assumptions (i) to (iii) in the Lemma above are as in Theorem 3.1. of AtheyWager18 that establishes asymptotic normality of a single (honest) random forest estimator of a mean regression function. Assumption (iv) is an additional rate restriction that is needed to achieve asymptotic normality of the sieve coefficient estimates. Additional rate restrictions are common in the series estimation literature, see e.g. Theorem 4.2. (iii) in BellChern_Series and also appear in Assumption 5 (ii) of Breunig_VRC for an RC-density estimate. Here the rate restriction is milder then the one in (ref) (ii) if for instance $\delta>2$ and the convergence rate $b$ is sufficiently fast such that $-\beta^{\ast}/\beta_{\min} > 1$. In this case the rate restriction in Assumption (ref) (iv) that guarantees consistency of the RC-density estimate is already sufficient for Assumption (ref) (iii). If however $\delta$ is rather small and the convergence rate $b$ close to the worst-case $\beta_{\min}$, there can be cases where the rate restriction of Lemma (ref) is stronger compared to the one in (ref) (ii), especially if the decay of $\tau_K$ is slow.
The following additional assumption is required.
Assumption (ref) is an undersmoothing condition that is standard for pointwise inference of a series estimator, see BellChern_Series (4.18). Note that similar rate restrictions are not needed for $\widehat{\beta}(x), \widehat{g}, \widehat{m} $, as these are calculated on a sample proportional to $n$ and thus converge at rate $r_p(n)$ which is always faster then the standard error rate $v_n(b,w)$. An estimator for $v_n(b,x)$ is
where $\widehat{\Sigma}_n(x) = diag\left(\widehat\sigma_{dm,1}(x), \dots, \widehat\sigma_{dm,K}(x) \right)$ and the individual standard error estimates $\widehat\sigma_{dm,k}(x)$ are specific to the employed ML-method. For random forests these can be obtained from applying the infinitisimal jackknife procedure of Efron2014, see also the discussion in AtheyWager18. Additional rate restrictions required for consistent estimation of the standard error $v_n(b,x)$ are not required. In the proof of the subsequent Theorem (ref) it is shown that $\widehat{v}_n(b,x)$ is consistent for $v_n(b,x)$ under the assumptions given so far.
The following pointwise asymptotic normality result holds
This determines the asymptotic normality of the estimator conditional on a given sample split, i.e. the case $M=1$. To handle the cross-fitting case $M>1$ and additional uncertainty due to sample splitting we can follow the variational inference approach of ChernGenericML. The idea summarizes as follows.
Suppose there are $M$ different estimates $\widehat{f}_{B|X}^l(b|x)$ for $l=1,\dots, M$. For each estimate it is possible to construct a (1-$\alpha$)-confidence interval $[L_{1-\alpha, l}, U_{1-\alpha, l}]$ from Theorem (ref) with $L_l=\widehat{f}_{B|X}^l(b|x)-c_{1-\alpha} \cdot \widehat{v}_n(b,w)$ and $U_l=\widehat{f}_{B|X}^l(b|x)+c_{1-\alpha} \cdot \widehat{v}_n(b,w)$ and $c_{1-\alpha}$ denoting the respective $1-\alpha$ quantile of the standard normal distribution.
To construct an asymptotically valid $1-\alpha$- confidence intervals for $f_{B|X}(b|x)$ ChernGenericML propose $[ \underline{Med}(\{L_{1-\alpha/2, l}\}_{l=1}^{M}), \overline{Med}(\{U_{1-\alpha/2, l}\}_{l=1}^{M})]$ with $\underline{Med}$ denoting the lower median and $\overline{Med}$ the upper median. The confidence level of each single interval needs to be discounted to $1-\alpha/2$. ChernGenericML provide a similar reasoning for constructing adjusted p-values.
This section touches on additional important aspects for the practical application of the estimation procedure. First, I discuss how to construct estimates of the marginal random coefficient density $f_{B}$. Second, I present a measure of variable importance that assigns an importance score to every variable in $X$. This is an important descriptive tool for uncovering which variables in $X$ drive the heterogeneity in conditional RC densities. Lastly, I discuss a cross-validation procedure for a data-driven choice of tuning parameters. \paragraph{Estimating marginal RC densities} There are various direct estimators for marginal RC densities in the literature such as the Radon transform estimator of hodermam_RC_2010 or an adaptation of the sieve estimation strategy from BreunigHoder18 and Breunig_VRC. The common identifying restriction is however full independence between $B$ and $W$ which is difficult to maintain in non-experimental data settings.
Maintaining the weaker conditional independence condition in Assumption (ref) (i) estimates of the marginal density can be readily constructed by averaging over leave-one-out estimates of conditional density estimates $\widehat{f}_{-i, B|X}$. Here the estimate is calculated without using the $i$-th datapoint $(Y_i, W_i. X_i)$. A consistent estimator for the marginal density is
The estimator $\widehat{f}_B$ will inherit its asymptotic properties from the conditional estimate $\widehat{f}_{B|X}$ which is discussed in the previous section. Thus the convergence rate is slower compared to direct marginal RC density estimators making use of full independence between random coefficients and covariates. To the best of my knowledge there are however currently no alternative estimators for the marginal random coefficient density that operate under Assumption (ref) (i). \paragraph{Variable Importance Measures} The estimation procedure outlined so far yields consistent estimates of $f_{B|X=x}$ for any given point $x$. An important question in applications is to identify those variables in the set of controls $X$ that drive heterogeneity in conditional RC densities, i.e. a criterion to guide the choice of interesting points $x$ on which to evaluate the estimate $\widehat f_{B | X}(b |x)$.
We focus here on our running example that makes use of regression forests. Note that for regression forests generally no post-selection inference problems arise as variable selection is done within in the various ML-steps of the estimation procedure. The goal is to find points $x$ that reveal interesting heterogeneities to the researcher. This is analogous to the role of variable importance measures for the causal forests of grf.
For each of the ML-estimators used we can calculate a measure of variable importance that assigns an importance score to each covariate that is normalized to sum to one. This score is informative on how often a specific variable has been used for placing splits in the growing of the forest.
First, I focus on the ML-estimates $\widehat{\Pi}_{dm}$ which constitute the sieve coefficients and thus determine the shape of the RC density. For each $k=1,\dots, K$ let $VI_k(X_l)$ denote an importance score assigned to covariate $X_l \in X$ by the regression forest estimator $\widehat{\Pi}_{dm}$. Here any normalized measure such as, cite can be used.
To obtain a global measure of variable importance for the shape of the function $f_{A|X=x}$ we can simply average over $K$. Thus define the variable importance of $X_l$ for the shape of the density as
A measure of variable importance for the conditional expectation of random coefficients $\beta(x)$ is directly available by considering the variable importance measure for causal forests as implemented in the grf-package. In contrast to $VI_{shape}$, this measure of variable importance for the center of the density will be henceforth referred to as $VI_{mean}$.
\paragraph{Parameter Tuning} In this paragraph I propose a cross-validation procedure for the choice of tuning parameters. Analogous to classical density estimation tuning parameters are chosen by minimization of the integrated squared error
which is equivalent to minimizing the criterion
The first part is simply the integrated squared RC density estimate. The second term is typically estimated via cross-validation. However it is not possible to observe realizations of random coefficients. The following Lemma links the second part to an expression that can be estimated via leave-one-out cross validation.
Here again a weighting for $w$ should be considered for practical reasons. So defining
it is equivalent to consider the integral $\int_{\mathbb{R}} \mathop{{\mathbf E}\null}\nolimits[V(Y,W)'\widehat Q^{-1}\widehat{\Pi}_{dm}(X) |X=x, W=w]f_W(w)dw$. Using a plug-in estimate for the unknown density $f_W$ the function $V$ can be computed and cross-validation used to estimate the conditional expectation with a machine learning estimator. Here either a subsample of observations that has not been used for calculating $\widehat{f}_{B|X}$ can be used for the prediction task or a leave-one-out estimator for $\widehat{f}_{B|X}$. Standard practices of cross-validation apply.
This section evaluates the finite sample performance of the RC-density estimator outlined in the earlier sections. The following data generating process is studied first,
where $A_0, V$ are standard normal random variables and $A_1$ is a mixture of a $N(-1.5,1)$ and a $N(1.5, \sqrt{1/2})$ random variable with weights $1/2$. In this setting the density of the random slope $B_1$ is bi-modal and any testpoint $X=x$ solely determines the center of the density function. The controls $X$ are a $p$-dimensional vector of iid standard normal variables. Here I set $p=10$ but as we see from the setup above, only variables $X_1, X_2, X_3$ are of importance in this toy model. This reflects the common practical problem that there is a large set of control variables but only some of them drive the heterogeneity in $B_1$ or may otherwise affect the outcome $Y$. Further I introduce a form of heteroskedasticity in the equation for $W$, such that we do not only consider the clean case where orthogonalization removes all dependence between $W$ and $X$. Further there is some form of dependence between the regressor $W$ and the random slope $B_1$ as both depend on the regressor $X_3$.
The goal is to estimate the density of the random slope $B_1$ conditional on some testpoint $X=x$. Here I implement the estimator in ((ref)) with the algorithm outlined at the end of Section (ref). In this setting $W$ does not need to be orthogonalized.
The other parameters of the estimation problem are chosen as follows. I set $K_1=K_2=3$ and thus there are a total number of $K=9$ basis functions. Hermite polynomials are used as sieve basis $q^K$ and the weighting measure follows a log-normal law, i.e. $\mu \sim lognormal(0, \sigma_t)$ with $\sigma_t=1$. In practice, when only the slope parameter is of interest, $K_1$ should be fixed and cross-validation performed to guide the choice of $K_2$ and $\sigma_t$. Simulations show that $K_2$ is the more relevant parameter for estimates compared to $\sigma_t$, so sole cross-validation of $K_2$ may be sufficient if computation time is a concern. To reduce computational effort parameters in this simulation study are not chosen via cross-validation and there is no cross-fitting as well. So $M=1$ and the sample is split only once in equally sized parts $\mathcal{R}$,$\mathcal{D}$ of size $n/2$ and RC-density estimates are computed only once per Monte Carlo iteration. The testpoint is chosen as $x=(0,0.3,0,\dots,0)$, so the correct density is centered around $0.3$.
All ML estimates are obtained from using honest regression forests, respectively causal forests for the quantity $\beta_1(x)$, see AtheyWager18 and grf, with the implementation taken from the $\textbf{grf}$-package in $R$. Each random forest is tuned using implemented data-driven routines, the number of trees in each forest is set to 2000, which is the packages default setting. In general, I find that tuning of internal forest parameters does improve the quality of estimates but is only of secondary importance for the overall shape of the density estimate.\\
The sample size is $n=1000$ and 100 Monte Carlo draws of the model in ((ref)) are performed. The simulation results for $\widehat f_{B_1|X=x}$ are presented in Figure (ref).
Figure (ref) shows a favorable performance of the estimator even for a moderate sample size and for a coarse choice of $K_2$. \\
The second data generating process is,
where all random variables are chosen as before and $B_1$ is a mixture distribution like $A_1$ in the first setting, but now with weights $\Phi(X_2), 1-\Phi(X_2)$. So in this setting $X$ determines the entire shape of the density function as opposed to the first setting where $X$ only determines the center of the density. For the testpoint $x=(0,0.3,0,\dots,0)$, the conditional density is again bi-modal but now the mode on the negative part of the domain is more pronounced. All parameters and hyperparameters of the ML-procedures are as before but now $K_2=7$ to illustrate the performance of the estimator for a more complex model. As the density of $B_1$ is more dispersed compared to the first setting this increase in complexity can be rationalized. Note that the support of each of the Hermite basis functions increases with $K$. This leads to the suggestion to increase $K$ for highly dispersed densities or to otherwise scale down $Y$ and $W$ accordingly to control the maximal dispersion of the density. The simulation results are presented in Figure (ref).
Through the larger number of basis functions the bias is comparably lower then in Figure (ref) at the expense of increased confidence intervals. As there are no shape constraints, we see that density estimates can in principal have negative parts. Yet, the method can detect the conditional density reliably even when the entire shape of the conditional density varies with $X$.
In this section the estimation strategy is applied to study heterogeneous effects of stock market expectations on portfolio choice. I make use of the innovation sample of the german socio-economic panel (SOEP-IS). Therein survey respondents were supplied with a hypothetical amount of $50,000$ Euros and asked to split their investment among one risk-free and one risky asset with returns paid out one year later. The risk-free asset is a state claim with a fixed annual interest rate of $4\%$ whereas the risky asset's return hinges on the return of the german stock market index (DAX) within the next year.
This experiment has been previously analyzed in HuckWeiz15. Economic theory suggests that stock market expectations and risk preferences are the main determinants of the portfolio choice task at hand.
The goal is to study the effect of stock market expectations on the investment in the risky asset. Formulated as a random coefficient model I study the following econometric model,
where $Y_i$ denotes the individual investment in the risky asset, $W_i$ is the individual belief on the development of the DAX for the next year and $B_{1,i}$ is the individual effect of interest. The random intercept $B_{0,i}$ subsumes the effects of other controls $X$ and further unobservable characteristics on the outcome $Y_i$.
The set of controls is quite rich and contains 75 variables including information on socio-demographics such as gender, age or tertiary degrees as well as self-assessed measures of risk aversion, personality traits or skills in mathematical calculations. Summary statistics for the main variables are provided in Table (ref).
In order to apply the estimation method we need to assume that the conditional independence restriction of Assumption (ref) (i) holds. Applied to the present setting this implies that stock market beliefs are exogenous conditional on the set of controls $X$. The data is observational and beliefs are self-reported, so we cannot rule out relations between beliefs and other controls which rules out considering the standard, unconditional RC model that relies on full independence of random coefficients and controls.
Further beliefs $W$ must vary sufficiently in the population to plausibly fulfill the support restrictions in Assumption (ref) (ii), which is the case in this setting. \\
The analysis begins by choosing the tuning parameters $K$ and $\sigma_t$. As the random slope is of main interest, I fix $K_1=3$ and $\sigma_t=1$ and vary the choice of $K_2$. Figure (ref) presents various estimates for different choices of $K_2$ and a suitable choice of $K_2$ can be eyeballed. For most choices of $K_2$ the two modes of the density are centered as for the case $K_2=5$ which thus appears to be a reasonable and coarse choice for the remainder of this analysis.
Next, I set a testpoint $x$ that corresponds to the medians of the variables in $X$. For those individuals with "median" characteristics $X=x$ an estimate of the random slope density $f_{B_1|X=x}$ is presented in (ref) below.
Most notable is the bi-modal shape of the density with one mode centered around zero and another around 2. Note that variables $Y$ and $W$ have been rescaled such that a value of 2 can be interpreted in the following way. A $1\%$-point increase in beliefs is associated with investing 1000 Euro (that is $2\%$ of available funds) more into the risky asset.
Such a bi-modal density corresponds to the existence of two types in the population. For one part of the population stock market expectations are actually linked to investment in the stock index as predicted by economic theory. Higher expectations also lead to a larger investment in the stock index. This does not seem to be true for a second group in the population where the marginal effect centers around zero. This part of the population may follow different, e.g. heuristic decision rules in their portfolio choice and their stated beliefs do not appear to be a relevant constituent of the investment decision.
This appearance of types is in line with other results from the portfolio choice literature such as GaudeckerME. They establish a link between the precision of subjective beliefs and the predictive power of economic models. Whenever beliefs are rather crude and imprecise they are likely not determinants of a rational portfolio choice. For individuals with such beliefs, economic theory has a rather low power in predicting their stock market participation. This is in line with the finding here that for some part of the population their stated, subjective beliefs do not seem to influence their investment decision. \\
So far the finding indicates the existence of two, equally-large groups in the population. One group for which beliefs seem to have an impact on investment choice and one where it does not. Next I study heterogeneity of the random slope densities, i.e. consider estimates of $f_{B_1|X=x}$ evaluated at different testpoints $x$. This is interesting because the type distribution may vary across subpopulations with different observable characteristics $X$.
To get an idea of which variables may drive the heterogeneity I report a variable importance measure for the density's shape and center, as outlined in Section (ref). The largest variable importance scores among the 75 control variables are reported in Table (ref).
There does not appear to be much variation in densities across different controls. The most importance is given to the age variable followed by "daxnetto1" and "daxnetto2" which are randomly selected information on past annual DAX returns that were presented to the survey respondents before the investment game. The other two are a measure of risk-aversion and a measure of self-assessed patience.
Taking these results allows to investigate heterogeneity with respect to age and the historic information.
Figure (ref) shows the heterogeneity in random slope densities for different age groups. Therefore the conditional density estimate is evaluated at three different testpoints. The variable age is varied but all other points are set to the respective sample median value of the variables. So $x$ is as in Figure (ref) except that age is varied.
The most prominent descriptive fact here is that the type composure in the population seems to vary with age. For the young and medium aged subpopulations both types are equal in size. For the elder subpopulation fewer people behave according to economic theory.
Next I also consider heterogeneity with respect to the historic information that has been displayed to the respondents. Again there are three testpoints. There is one testpoint where both historic informations have been very positive (return of 35%), one where both informations are negative (return of -5%) and one mixed with a positive first information and a negative second. The results are displayed in Figure (ref).
Here there is no apparent or robust heterogeneity with respect to the historic information. We therefore stop the analysis and do not vary according to variables with a lower variable importance score than that of $daxnetto2$.
The random coefficient analysis suggests the presence of two, roughly equally-sized types in the population. One group of individuals complies with economic theory in that their stock market expectation also explains their investment in a risky asset. A second group seems to follow different decision rules, their beliefs have no impact on their investment decision. Due to a possible correlation of beliefs and random coefficients this finding cannot be inferred from estimating standard, marginal random coefficient models.
Regarding heterogeneity I find that the mixture of types in the populuation may depend on age but fail to uncover more interesting heterogeneity with the given data.
It appears that the main determinants of type membership are unobservables that are not captured in the given data set.
The present paper discusses estimation of conditional random coefficient densities when the set of conditioning variables is large. The very general conditional RC model has rarely been studied in both theory and application. This paper provides a general sieve estimation strategy for estimating conditional RC densities. The approach enables the use of generic machine learning methods to estimate sieve coefficients in the presence of a large dimensional set of control variables. Therefore the estimator is applicable in many economic settings in which a continuous treatment variable is available. Theoretical results of the paper include convergence rate and inference results for the conditional sieve RC density estimator which combine asymptotic theories of sieve estimators and machine learning methods, in particular applying results on (honest) random forests.
The finite sample properties of the estimator are illustrated in a Monte Carlo simulation study and an empirical application. The application reveals behavioral heterogeneity in an experimental portfolio choice task which is in line with recent empirical findings in the literature.