EconBase
← Back to paper

Identification and Estimation of Categorical Random Coefficient Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

150,592 characters · 13 sections · 37 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation of Categorical Random Coefficient Models

abstractThis paper proposes a linear categorical random coefficient model, in which the random coefficients follow parametric categorical distributions. The distributional parameters are identified based on a linear recurrence structure of moments of the random coefficients. A Generalized Method of Moments estimation procedure is proposed also employed by Peter Schmidt and his coauthors to address heterogeneity in time effects in panel data models. Using Monte Carlo simulations, we find that moments of the random coefficients can be estimated reasonably accurately, but large samples are required for estimation of the parameters of the underlying categorical distribution. The utility of the proposed estimator is illustrated by estimating the distribution of returns to education in the U.S. by gender and educational levels. We find that rising heterogeneity between educational groups is mainly due to the increasing returns to education for those with postsecondary education, whereas within group heterogeneity has been rising mostly in the case of individuals with high school or less education.

Keywords: Random coefficient models, categorical distribution, return to education

JEL Code: C01, C21, C13, C46, J30

\thispagestyle{empty}

\setcounter{page}{1}

Introduction

Random coefficient models have been used extensively in time series, cross-section and panel regressions. nicholls198516 consider the estimation of first and second moments of the random coefficient ${\Greekmath 010C} _{i}$ and the error term $u_{i}$, in a linear regression model. In the seminal work, beran1992estimating establish the conditions of identifying and estimating the distribution of ${\Greekmath 010C} _{i}$ and $u_{i}$ non-parametrically. The baseline linear univariate regression in beran1992estimating has been extended in non-parametric framework by \citet*{beran1993semiparametric, beran1994minimum, beran1996nonparametric, hoderlein2010analyzing, hoderlein2017triangular} and breunig2018specification, to just name a few. hsiao2008random survey random coefficient models in linear panel data models.

In some econometric applications, hausman1981exact,hausman1995nonparametric,foster_hahn2000consistent for examples, the main interest is to estimate the consumer surplus distribution based on a linear demand system where the coefficient associated with the price is random. In such settings, the distribution of the random coefficients is needed when computing the consumer surplus function, and the non-parametric estimation is more general, flexible and suitable for the purpose. On the other hand, parametric models may be favored in applications in which the implied economic meaning of the distribution of the random coefficients is of interests. Examples include estimation of the return to education lemieux2006postnber,lemieux2006postsecondary and the labor supply equation \citep*{bick2020hours}.

In this paper, we consider a linear regression model with a random coefficient ${\Greekmath 010C} _{i}$ that is assumed to follow a categorical distribution, i.e. ${\Greekmath 010C} _{i}$ has a discrete support $\left\{ b_{1},b_{2},\cdots ,b_{K}\right\} $, and ${\Greekmath 010C} _{i}=b_{k}$ with probability ${\Greekmath 0119} _{k}$. The discretization of the support of the random coefficient $ {\Greekmath 010C} _{i}$ naturally corresponds to the interpretation that each individual belongs to a certain category, or group, $k$ with probability ${\Greekmath 0119} _{k}$. Compared to a non-parametric distribution with continuous support, assuming a categorical distribution allows us not only to model the heterogeneous responses across individuals but also to interpret the results with sharper economic meaning. As we will illustrate in the empirical application in Section (ref), it is hard to clearly interpret the distribution of returns to education without imposing some form of parametric restrictions.

In addition, with the categorical distribution imposed, the identification and estimation of the distribution of ${\Greekmath 010C} _{i}$ do not rely on identically distributed error terms $u_{i}$ and regressors $\mathbf{w}_i$, as shown in Section (ref) and (ref). Heterogeneously generated errors can be allowed, which is important in many empirical applications. To the best of our knowledge, this is the first identification result in linear random coefficient model without a strict IID setting.

The identification of the distribution of ${\Greekmath 010C} _{i}$ is established in this paper based on the identification of the moments of ${\Greekmath 010C} _{i}$, which coincides with the identification condition in beran1992estimating that the distribution of ${\Greekmath 010C} _{i}$ is uniquely determined by its moments, which is assumed to exist up to an arbitrary order. Since under our setup the distribution of ${\Greekmath 010C} _{i}$ is parametrically specified, the moments of ${\Greekmath 010C} _{i}$ exist and can be derived explicitly. The parameters of the assumed categorical distribution can then be uniquely determined by a system of equations in terms of the moments, as in Theorem (ref). The parameters of the categorical distribution are then estimated consistently by the generalized method of moments (GMM). The estimation procedure based on moment conditions shares similar spirits as in \citet*{ahn2001gmm,ahn2013panel} in which Peter Schmidt and coauthors study panel data models with interactive effects where they allow for the time effects to vary across individual units. Comparing to alternative non-parametric random coefficient models, the standard GMM estimation is easy to implement, and the identified categorical structure has a clear economic interpretation.

Using Monte Carlo (MC) simulations, we find that moments of the random coefficients can be estimated reasonably accurately, but large samples are required for estimation of the parameters of the underlying categorical distributions. Our theoretical and MC results also suggest that our method is suitable when the number heterogeneous coefficients and the number of categories are small (2 or 3). With the number of categories rising the burden on identification from the moments to the parameters of the categorical distribution also rises rapidly. The quality of identification also deteriorates as we need to rely on higher and higher moments to identify a larger number of categories, since the information content of the moments tend to decline with their order.

The proposed method is also illustrated by providing estimates of the distribution of returns to education in the U.S. by gender and educational levels, using the May and Outgoing Rotation Group (ORG) supplements of the Current Population Survey (CPS) data. Comparing the estimates obtained over the sub-periods 1973-75 and 2001-03, we find that rising between group heterogeneity is largely due to rising returns to education in the case of individuals with postsecondary education, whilst within group heterogeneity has been rising in the case of individuals with high school or less education.

Related Literature: This paper draws mainly upon the literature of random coefficient models. As already mentioned, the main body of the recent literature is focused on non-parametric identification and estimation. Following beran1992estimating, beran1993semiparametric and beran1994minimum extend the model to a linear semi-parametric model with a multivariate setup and propose a minimum distance estimator for the unknown distribution. foster_hahn2000consistent extend the identification results in beran1992estimating and apply the minimum distance estimator to a gasoline consumption data to estimate the consumer surplus function. \citet*{beran1996nonparametric} and \citet*{ hoderlein2010analyzing} propose kernel density estimators based on the Radon inverse transformation in linear models.

In addition to linear models, ichimura1998maximum and gautier2013nonparametric incorporate the random coefficients in binary choice models. gautier2011triangular and \citet*{ hoderlein2017triangular} consider triangular models with random coefficients allowing for causal inference. matzkin2012identification and masten2018random discuss the identification of random coefficients in simultaneous equation models. breunig2018specification propose a general specification test in a variety of random coefficient models. Random coefficients are also widely studied in panel data models, for example hsiao2008random and arellano2012identifying

The rest of the paper is organized as follows: Section (ref) establishes the main identification results. The GMM estimation procedure is proposed and discussed in Section (ref). An extension to a multivariate setting is considered in Section (ref). Small sample properties of the proposed estimator are investigated in Section (ref), using Monte Carlo techniques under different regressor and error distributions. Section (ref) presents and discusses our empirical application to the return to education. Section (ref) provides some concluding remarks and suggestions for future work. Technical proofs are given in Appendix (ref).

Notations: Largest and smallest eigenvalues of the $ p\times p$ matrix $\mathbf{A}=\left( a_{ij}\right) $ are denoted by ${\Greekmath 0115} _{\max }\left( \mathbf{A}\right) $ and ${\Greekmath 0115} _{\min }\left( \mathbf{A} \right) ,$ respectively, its spectral norm by $\left\Vert \mathbf{A} \right\Vert ={\Greekmath 0115} _{\max }^{1/2}\left( \mathbf{A}^{\prime }\mathbf{A} \right) $, $\mathbf{A}\succ 0$ means that $\mathbf{A}$ is positive definite, $\mathrm{vech\left( \mathbf{A}\right) }$ denotes the vectorization of distinct elements of $\mathbf{A}$, $\mathbf{0}$ denotes zero matrix (or vector). For $\mathbf{a}\in \mathbb{R}^{p}$, $\mathrm{diag}\left( \mathbf{a} \right) $ represents the diagonal matrix with elements of $ a_{1},a_{2},\cdots ,a_{p}$. For random variables (or vectors) $u$ and $v$, $ u\perp v$ represents $u$ is independent of $v$. We use $c$ ($C$) to denote some small (large) positive constants. For a differentiable real-valued function $f\left( \mathbf{{\Greekmath 0112} }\right) $, ${\Greekmath 0272} _{\mathbf{{\Greekmath 0112} } }f\left( \mathbf{{\Greekmath 0112} }\right) $ denotes the gradient vector. Operator $ \rightarrow _{p}$ denotes convergence in probability, and $\rightarrow _{d}$ convergence in distribution. The symbols $O(1)$, and $O_{p}(1)$ denote asymptotically bounded deterministic and random sequences, respectively.

Categorical random coefficient model

We suppose the single cross-section observations, $\left\{ y_{i},x_{i}, \mathbf{z}_{i}\right\} _{i=1}^{n}$, follow the categorical random coefficient model

equation[equation omitted — 131 chars of source]

where $y_{i},x_{i}\in \mathbb{R}$, $\mathbf{z}_{i}\in \mathbb{R}^{p_z},$ and ${\Greekmath 010C} _{i}\in \left\{ b_{1},b_{2},\cdots ,b_{K}\right\} $ admits the following $K$-categorical distribution,

equation[equation omitted — 430 chars of source]

w.p. denotes “with probability”, ${\Greekmath 0119}_{k}\in \left( 0,1\right) $, $ \sum_{k=1}^{K}{\Greekmath 0119} _{k}=1$, $b_{1}<b_{2}<\cdots <b_{K}$, $\mathbf{{\Greekmath 010D} }\in \mathbb{R}^{p_z}$ is homogeneous and $\mathbf{z}_{i}$ could include an intercept term as its first element. It is assumed that ${\Greekmath 010C} _{i}\perp \mathbf{w}_i = \left( x_{i},\mathbf{z}_{i}^{\prime }\right) ^{\prime }$, and the idiosyncratic errors $u_{i}$ are independently distributed with mean $0$.

remarkThe model can be extended to allow $\mathbf{x}_{i},\mathbf{{\Greekmath 010C} }_{i}\in \mathbb{R}^{p}$, with $\mathbf{{\Greekmath 010C} }_{i}$ following a multivariate categorical distribution, though with more complicated notations. We will consider possible extensions in Section (ref).
remarkSince we consider a pure cross-sectional setting, the key assumption that $ {\Greekmath 010C} _{i}$ and $x_{i}$ are independently distributed cannot be relaxed. Allowing ${\Greekmath 010C} _{i}$ to vary with $w_{i}$, without any further restrictions, is tantamount to assuming $y_{i}$ is a general function of $ w_{i}$, in effect rendering a nonparametric specification.
remarkThe number of categories $K$ is assumed to be fixed and known. Conditions $ \sum_{k=1}^{K}{\Greekmath 0119} _{k}=1$, $b_{1}<b_{2}<\cdots <b_{K},$ and ${\Greekmath 0119} _{k}\in \left( 0,1\right) $ together are sufficient for the existence of $K$ categories. For example, if $b_{k}=b_{k^{\prime }}$, then we can merge categories $k$ and $k^{\prime }$, and the number of categories reduces to $ K-1$. Similarly, if ${\Greekmath 0119} _{k}=0$ for some $k$, then category $k$ can be deleted, and the number of categories is again reduced to $K-1$. Information criteria can be used to determine $K$, but this will not be pursued in this paper. Model specification tests could also be considered. See, for examples, andrews2001testing and breunig2018specification.

In the rest of this section, we focus on the model ((ref)) and establish the conditions under which the distribution of ${\Greekmath 010C} _{i}$ is identified.

Identifying the moments of $\protect{\Greekmath 010C} _{i}$

assumption\begin{enumerate} • (i) $u_{i}$ is distributed independently of $\mathbf{w} _{i}=\left( x_{i},\mathbf{z}_{i}^{\prime }\right) ^{\prime }$ and ${\Greekmath 010C} _{i} $. (ii) $\sup_{i}\mathrm{E}\left( \left\vert u_{i}^{r}\right\vert \right) <C$, $r=1,2,\cdots ,2K-1$. (iii) $n^{-1}\sum_{i=1}^{n} u_i^4 = O_p(1) $. • (i) Let $\mathbf{Q}_{n,ww}=n^{-1}\sum_{i=1}^{n}\mathbf{w}_{i} \mathbf{w}_{i}^{\prime }$, and $\mathbf{q}_{n,wy}=n^{-1}\sum_{i=1}^{n} \mathbf{w}_{i}y_{i}$. Then $\left\Vert \mathrm{E}\left( \mathbf{Q} _{n,ww}\right) \right\Vert <C<\infty $, and $\left\Vert \mathrm{E}\left( \mathbf{q}_{n,wy}\right) \right\Vert <C<\infty $, and there exists $n_{0}\in \mathbb{N}$ such that for all $n\geq n_{0}$, \begin{equation*} 0<c<{\Greekmath 0115} _{\min }\left( \mathbf{Q}_{n,ww}\right) <{\Greekmath 0115} _{\max }\left( \mathbf{Q}_{n,ww}\right) <C<\infty . \end{equation*} (ii) $\sup_{i} \mathrm{E}\left( \left\Vert \mathbf{w}_{i}\right\Vert ^{r}\right) <C<\infty $, $r=1,2,\cdots ,4K-2$.\newline (iii) $n^{-1} \sum_{i=1}^{n} \left\Vert \mathbf{w}_{i} \right\Vert^{4} = O_p(1)$. • $\left\Vert \mathbf{Q}_{n,ww}-\mathrm{E} \left( \mathbf{Q} _{n,ww}\right) \right\Vert =O_p\left( n^{-1/2}\right)$, $\left\Vert \mathbf{q }_{n,wy}- \mathrm{E} \left( \mathbf{q}_{n,wy}\right) \right\Vert =O_p\left( n^{-1/2}\right)$, and \begin{equation*} \mathrm{E} \left( \mathbf{Q}_{n,ww}\right) =n^{-1}\sum_{i=1}^{n}\mathrm{E} \left( \mathbf{w}_{i}\mathbf{w}_{i}^{\prime }\right) \succ 0. \end{equation*} • $\left\Vert \mathrm{E} \left( \mathbf{Q}_{n,ww}\right) -\mathbf{Q} _{ww}\right\Vert =O\left( n^{-1/2}\right) $, $\left\Vert \mathrm{E} \left( \mathbf{q}_{n,wy}\right) -\mathbf{q}_{wy}\right\Vert =O\left( n^{-1/2}\right) $, where $\mathbf{q}_{wy} = \lim\limits_{n \to \infty} \mathrm{E} \left( \mathbf{q}_{n, wy} \right)$, $\mathbf{Q}_{ww} = \lim\limits_{n \to \infty} \mathrm{E} \left( \mathbf{Q}_{n, ww} \right) $ and $\mathbf{Q}_{ww}\succ 0$. \end{enumerate}
remarkPart (a) of Assumption (ref) relaxes the assumption that $u_{i}$ is identically distributed, and allows for heterogeneously generated errors. For identification of the distribution of ${\Greekmath 010C} _{i}$, we require $u_{i}$ to be distributed independently of $ \mathbf{w}_{i}$ and ${\Greekmath 010C} _{i}$, which rules out conditional heteroskedasticity. However, estimation and inference involving $\mathrm{E} \left( {\Greekmath 010C} _{i}\right) $ and $\mathbf{{\Greekmath 010D} }$ can be carried out in presence of conditionally error heteroskedastic, as shown in Theorem (ref). Parts (c) and (d) of Assumption (ref) relax the condition that $\mathbf{ w}_{i}$ is identically distributed across $i$. As we proceed, only ${\Greekmath 010C} _{i}$, whose distribution is of interest, is assumed to be IID across $i$, and it is not required for $\mathbf{w}_{i}$ and $u_{i}$ to be identically distributed over $i$.
remarkThe high level conditions in Assumption (ref), concerning the convergence in probability of averages such as $Q_{n,ww}=n^{-1}\sum_{i=1}^{n} \mathbf{w}_{i}\mathbf{w}_{i}^{\prime }$, can be verified under weak cross-sectional dependence. Let $f_{i}=f\left( \mathbf{w}_{i},{\Greekmath 010C} _{i},u_{i}\right) $ be a generic function of $\mathbf{w}_{i}$, ${\Greekmath 010C} _{i}$ and $u_{i}$.\footnote{$f_{i}$ is assumed to be a scalar, and we can apply the analysis element-by-element to a matrix, for example $\mathbf{w}_{i} \mathbf{w}_{i}^{\prime }$.} Assume that $\sup_{i}\mathrm{E}\left( f_{i}^{2}\right) <C$, and $\sup_{j}\sum_{i=1}^{n}\left\vert \mathrm{cov} \left( f_{i},f_{j}\right) \right\vert <C$, for some fixed $C<\infty $. Then, \begin{equation*} \mathrm{var}\left( \frac{1}{n}\sum_{i=1}^{n}f_{i}\right) \leq \frac{1}{n^{2}} \sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert \mathrm{cov}\left( f_{i},f_{j}\right) \right\vert \leq \frac{1}{n}\sup_{j}\sum_{i=1}^{n}\left\vert \mathrm{cov} \left( f_{i},f_{j}\right) \right\vert \leq \frac{C}{n}. \end{equation*} By Chebyshev's inequality, for any ${\Greekmath 0122} >0$, we have $M_{{\Greekmath 0122} }>\sqrt{C/{\Greekmath 0122} }$ such that \begin{equation*} \Pr \left( \sqrt{n}\left\vert \frac{1}{n}\sum_{i=1}^{n}\left[ f_{i}-\mathrm{E }\left( f_{i}\right) \right] \right\vert >M_{{\Greekmath 0122} }\right) \leq \frac{ n\mathrm{var}\left( n^{-1}\sum_{i=1}^{n}f_{i}\right) }{C}{\Greekmath 0122} \leq {\Greekmath 0122} , \end{equation*} i.e. $n^{-1}\sum_{i=1}^{n}\left[ f_{i}-\mathrm{E}\left( f_{i}\right) \right] =O_{p}\left( n^{-1/2}\right) $.

Denote $\mathbf{{\Greekmath 011E} }_{i}=\left( {\Greekmath 010C} _{i},\mathbf{{\Greekmath 010D} }^{\prime }\right) ^{\prime }$ and $\mathbf{{\Greekmath 011E} }=\mathrm{E}\left( \mathbf{{\Greekmath 011E} } _{i}\right) =\left( \mathrm{E}\left( {\Greekmath 010C} _{i}\right) ,\mathbf{{\Greekmath 010D} } ^{\prime }\right) ^{\prime }$. Consider the moment condition,

equation[equation omitted — 171 chars of source]

and sum (ref) over $i$

equation[equation omitted — 237 chars of source]

Let $n\to \infty$, then $\mathbf{{\Greekmath 011E} }$ is identified by

equation[equation omitted — 97 chars of source]

under Assumption (ref).

assumptionLet $\tilde{y}_{i}=y_{i}-\mathbf{z}_{i}^{\prime } \mathbf{{\Greekmath 010D} }$. \begin{enumerate} • {$\left\vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( \tilde{y} _{i}^{r}x_{i}^{s}\right) -{\Greekmath 011A} _{r,s}\right\vert =O\left( n^{-1/2}\right) ,$ and $\left\vert {\Greekmath 011A} _{r,s}\right\vert <\infty ,$ for $r,s=0,1,\cdots ,2K-1$ . } • {$\left\vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( u_{i}^{r}\right) -{\Greekmath 011B} _{r}\right\vert =O\left( n^{-1/2}\right) ,$ and $ \left\vert {\Greekmath 011B} _{r}\right\vert <\infty $, for $r=2,3,\cdots ,2K-1$. } • $n^{-1}\sum_{i=1}^{n}\left[ \mathrm{var}(x_{i}^{r})-\left( {\Greekmath 011A} _{0,2r}-{\Greekmath 011A} _{0,r}^{2}\right) \right] =O\left( n^{-1/2}\right) $ where $ {\Greekmath 011A} _{0,2r}-{\Greekmath 011A} _{0,r}^{2}>0,$ for $r=2,3,\cdots ,2K-1$. \end{enumerate}
remarkThe above assumption allows for a limited degree of heterogeneity of the moments. As an example, let $\mathrm{E}\left( u_{i}^{r}\right) ={\Greekmath 011B} _{ir}$ and denote the heterogeneity of the $r^{th}$ moment of $u_{i}$ by $e_{ir}={\Greekmath 011B} _{ir}-{\Greekmath 011B} _{r}$. Then \begin{equation*} \left\vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( u_{i}^{r}\right) -{\Greekmath 011B} _{r}\right\vert \leq n^{-1}\sum_{i=1}^{n}\left\vert e_{ir}\right\vert , \end{equation*} and condition (b)\ of Assumption (ref) is met if $\ \sum_{i=1}^{n}\left\vert e_{ir}\right\vert =O(n^{{\Greekmath 010B} _{r}})$ with ${\Greekmath 010B} _{r}<1/2$. ${\Greekmath 010B} _{r}$ measures the degree of heterogeneity with ${\Greekmath 010B} _{r}=1$ representing the highest degree of heterogeneity. A similar idea is used by pesaran2018pool in their analysis of poolability in panel data models.
theoremUnder Assumptions (ref) and (ref), $ \mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) $ and ${\Greekmath 011B} _{r}$, $r=2,3,\cdots ,2K-1$ are identified.
proofFor $r=2,\cdots ,2K-1$, \begin{align} \mathrm{E}\left( \tilde{y}_{i}^{r}\right) & =\mathrm{E}\left( x_{i}^{r}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) +\mathrm{E}\left( u_{i}^{r}\right) +\sum_{q=2}^{r-1}\binom{r}{q}\mathrm{E}\left( x_{i}^{r-q}\right) \mathrm{E}\left( u_{i}^{q}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{r-q}\right) , \\ \mathrm{E}\left( \tilde{y}_{i}^{r}x_{i}^{r}\right) & =\mathrm{E}\left( x_{i}^{2r}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) +\mathrm{E}\left( x_{i}^{r}\right) \mathrm{E}\left( u_{i}^{r}\right) +\sum_{q=2}^{r-1}\binom{r }{q}\mathrm{E}\left( x_{i}^{2r-q}\right) \mathrm{E}\left( u_{i}^{q}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{r-q}\right) . \end{align} where $\binom{r}{q}=\frac{r!}{q!(r-q)!}$ are binomial coefficients, for non-negative integers $q\leq r$. Sum over $i$, then by parts (a) and (b) of Assumption (ref) , \begin{align} {\Greekmath 011A} _{0,r}\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) +{\Greekmath 011B} _{r}& ={\Greekmath 011A} _{r,0}-\sum_{q=2}^{r-1}\binom{r}{q}{\Greekmath 011A} _{0,r-q}{\Greekmath 011B} _{q}\mathrm{E}\left( {\Greekmath 010C} _{i}^{r-q}\right) , \\ {\Greekmath 011A} _{0,2r}\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) +{\Greekmath 011A} _{0,r}{\Greekmath 011B} _{r}& ={\Greekmath 011A} _{r,r}-\sum_{q=2}^{r-1}\binom{r}{q}{\Greekmath 011A} _{0,2r-q}{\Greekmath 011B} _{q}\mathrm{E} \left( {\Greekmath 010C} _{i}^{r-q}\right) . \end{align} Derivation details are relegated to Appendix (ref). By part (c) of Assumption (ref), the matrix $ \begin{pmatrix} {\Greekmath 011A} _{0,r} & 1 \\ {\Greekmath 011A} _{0,2r} & {\Greekmath 011A} _{0,r} \end{pmatrix} $ is invertible for $r=2,3,\cdots ,2K-1$. As a result, we can sequentially solve (ref) and (ref) for $\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) $ and ${\Greekmath 011B} _{r}$, for $r=2,3,\cdots ,2K-1$.

Identifying the distribution of $\protect{\Greekmath 010C} _{i}$

beran1992estimating prove the identification of the distribution of the random coefficient, ${\Greekmath 010C} _{i}$, in a canonical model without covariates, $z_{i}$, under the condition that the distribution of ${\Greekmath 010C} _{i}$ is uniquely determined by its moments. We show the identification of moments of ${\Greekmath 010C}_i$ holds more generally when $x_i$ and $ u_i$ are not identically distributed and the distribution of ${\Greekmath 010C}_i$ is identified if it follows a categorical distribution. Note that under ((ref)),

equation[equation omitted — 152 chars of source]

with $\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) $ identified under Assumption (ref). To identify $\mathbf{{\Greekmath 0119} } =\left( {\Greekmath 0119} _{1},{\Greekmath 0119} _{2},...,{\Greekmath 0119} _{K}\right) ^{\prime }$ and $\mathbf{b} =\left( b_{1},b_{2},...,b_{K}\right) ^{\prime }$, we need to verify that the system of $2K$ equations in (ref) has a unique solution if $ b_{1}<b_{2}<\cdots <b_{K}$, and ${\Greekmath 0119} _{k}\in \left( 0,1\right) $. In the proof, we construct a linear recurrence relation and make use of the corresponding characteristic polynomial.

theoremConsider the random coefficient regression model (ref), suppose that Assumptions (ref) and (ref) hold. Then $\mathbf{{\Greekmath 0112} }=\left( \mathbf{{\Greekmath 0119} }^{\prime },\mathbf{b}^{\prime }\right) ^{\prime }$ is identified subject to $b_{1}<b_{2}<\cdots <b_{K}$ and ${\Greekmath 0119} _{k}\in \left( 0,1\right) $, for all $k=1,2,\cdots ,K$.
proofWe motivate the key idea of the proof in the special case where $K=2,$ and relegate the proof of the general case to the Appendix (ref). Let $b_{1}={\Greekmath 010C} _{L}$, $b_{2}={\Greekmath 010C} _{H}$, ${\Greekmath 0119} _{1}={\Greekmath 0119} $ and ${\Greekmath 0119} _{2}=1-{\Greekmath 0119} $ . Note that \begin{align} \mathrm{E}\left( {\Greekmath 010C} _{i}\right) & ={\Greekmath 0119} {\Greekmath 010C} _{L}+\left( 1-{\Greekmath 0119} \right) {\Greekmath 010C} _{H}, \\ \mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) & ={\Greekmath 0119} {\Greekmath 010C} _{L}^{2}+\left( 1-{\Greekmath 0119} \right) {\Greekmath 010C} _{H}^{2}, \\ \mathrm{E}\left( {\Greekmath 010C} _{i}^{3}\right) & ={\Greekmath 0119} {\Greekmath 010C} _{L}^{3}+\left( 1-{\Greekmath 0119} \right) {\Greekmath 010C} _{H}^{3}, \end{align} and $\mathrm{E}\left( {\Greekmath 010C} _{i}^{k}\right) $, $k=1,2,3$ are identified. $ \left( {\Greekmath 0119} ,{\Greekmath 010C} _{L},{\Greekmath 010C} _{H}\right) $ can be identified if the system of equations (ref) to (ref), has a unique solution. By (ref), \begin{equation} {\Greekmath 0119} =\frac{{\Greekmath 010C} _{H}-\mathrm{E}\left( {\Greekmath 010C} _{i}\right) }{{\Greekmath 010C} _{H}-{\Greekmath 010C} _{L}},\; \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{and}\; 1-{\Greekmath 0119} =\frac{\mathrm{E}\left( {\Greekmath 010C} _{i}\right) -{\Greekmath 010C} _{L}}{{\Greekmath 010C} _{H}-{\Greekmath 010C} _{L}}. \end{equation} Plug (ref) into (ref) and (ref), \begin{align} \mathrm{E}\left( {\Greekmath 010C} _{i}\right) \left( {\Greekmath 010C} _{L}+{\Greekmath 010C} _{H}\right) -{\Greekmath 010C} _{L}{\Greekmath 010C} _{H}& =\mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) , \\ \mathrm{E}\left( {\Greekmath 010C} _{i}^2\right) \left( {\Greekmath 010C}_L + {\Greekmath 010C}_H\right) - \mathrm{E}\left({\Greekmath 010C}_i\right) {\Greekmath 010C} _{L}{\Greekmath 010C} _{H} & =\mathrm{E}\left( {\Greekmath 010C} _{i}^{3}\right) . \end{align} Denote ${\Greekmath 010C}_{L+H} = {\Greekmath 010C}_{L} + {\Greekmath 010C}_{H}$ and ${\Greekmath 010C}_{LH} = {\Greekmath 010C}_{L}{\Greekmath 010C}_{H}$, and write (ref) and (ref) in matrix form, \begin{equation} \mathbf{M} \mathbf{D} \mathbf{b}^{\ast} = \mathbf{m}, \end{equation} where \begin{equation*} \mathbf{M} = \begin{pmatrix} 1 & \mathrm{E}\left({\Greekmath 010C}_i\right) \\ \mathrm{E}\left({\Greekmath 010C}_i\right) & \mathrm{E}\left({\Greekmath 010C}_i^2 \right) \end{pmatrix} , \, \mathbf{D} = \begin{pmatrix} -1 & 0 \\ 0 & 1 \end{pmatrix} ,\, \mathbf{b}^{\ast} = \begin{pmatrix} {\Greekmath 010C}_{LH} \\ {\Greekmath 010C}_{L+H} \end{pmatrix} , \;\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{and}\; \mathbf{m} = \begin{pmatrix} \mathrm{E}\left({\Greekmath 010C}_i^2\right) \\ \mathrm{E}\left({\Greekmath 010C}_i^3\right) \end{pmatrix} . \end{equation*} Under the conditions $0 < {\Greekmath 0119} < 1$ and ${\Greekmath 010C}_H > {\Greekmath 010C}_L$, \begin{equation*} \det\left( \mathbf{M} \right) = \mathrm{var}\left( {\Greekmath 010C} _{i}\right) = \mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) -\mathrm{E}\left( {\Greekmath 010C} _{i}\right) ^{2} = {\Greekmath 0119}\left( 1-{\Greekmath 0119} \right)\left( {\Greekmath 010C}_H - {\Greekmath 010C}_L \right)^2 > 0. \end{equation*} As a result, we can solve (ref) for ${\Greekmath 010C}_{L+H}$ and $ {\Greekmath 010C}_{LH}$ as \begin{align} {\Greekmath 010C} _{L+H}& =\frac{\mathrm{E}\left( {\Greekmath 010C} _{i}^{3}\right) -\mathrm{E} \left( {\Greekmath 010C} _{i}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) }{\mathrm{var }\left( {\Greekmath 010C}_i \right)}, \\ {\Greekmath 010C} _{LH}& =\frac{\mathrm{E}\left( {\Greekmath 010C} _{i}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{3}\right) -\mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) ^{2}}{\mathrm{ var}\left( {\Greekmath 010C}_i \right)}. \end{align} ${\Greekmath 010C} _{L}$ and ${\Greekmath 010C} _{H}$ are solutions to the quadratic equation, \begin{equation} {\Greekmath 010C} ^{2}-{\Greekmath 010C} _{L+H}{\Greekmath 010C} +{\Greekmath 010C} _{LH}=0. \end{equation} We can verify that $\Delta ={\Greekmath 010C} _{L+H}^{2}-4{\Greekmath 010C} _{LH}>0$ by direct calculation using (ref) and (ref). Simplifying $\Delta $ in terms of $\mathrm{E}\left( {\Greekmath 010C} _{i}^{k}\right) $ and then plugging in (ref), (ref) and (ref), \begin{align*} \Delta =& \frac{ \left[ \mathrm{E}\left( {\Greekmath 010C} _{i}^{3}\right) - \mathrm{E} \left( {\Greekmath 010C} _{i}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) \right] ^{2} -4 \mathrm{var}\left( {\Greekmath 010C}_i \right) \left[ \mathrm{E}\left( {\Greekmath 010C} _{i}\right) \mathrm{E}\left({\Greekmath 010C} _{i}^{3}\right) - \mathrm{E}\left( {\Greekmath 010C}_{i}^{2}\right) ^{2} \right] }{ \left[ \mathrm{var}\left( {\Greekmath 010C}_i \right) \right] ^{2} } \\ & = \left( {\Greekmath 010C}_H - {\Greekmath 010C}_L \right)^2 > 0. \end{align*} Then, we obtain the unique solutions, \begin{align} {\Greekmath 010C} _{L}& =\frac{1}{2}\left( {\Greekmath 010C} _{L+H}-\sqrt{{\Greekmath 010C} _{L+H}^{2}-4{\Greekmath 010C} _{LH}}\right) , \\ {\Greekmath 010C} _{H}& =\frac{1}{2}\left( {\Greekmath 010C} _{L+H}+\sqrt{{\Greekmath 010C} _{L+H}^{2}-4{\Greekmath 010C} _{LH}}\right) , \end{align} and ${\Greekmath 0119} $ can be determined by ((ref)) correspondingly.
remarkThe key identifying assumption in ((ref)) is the assumed existence of the strict ordinal relation $b_{1}<b_{2}<\cdots <b_{K}$ so that $b_{k}$ and $b_{k^{\prime }}$ are not symmetric for $k\neq k^{\prime }$, and $0<{\Greekmath 0119} _{k}<1$ so that the distribution of ${\Greekmath 010C} _{i}$ does not degenerate. When $K=2$, the conditions $b_{1}<b_{2}<\cdots <b_{K}$, and ${\Greekmath 0119} _{k}\in \left( 0,1\right) $, are equivalent to $\mathrm{var}\left( {\Greekmath 010C} _{i}\right) ={\Greekmath 0119} _{1}\left( 1-{\Greekmath 0119} _{1}\right) \left( b_{2}-b_{1}\right) ^{2}>0$. In other words, not surprisingly, the categorical distribution of ${\Greekmath 010C} _{i}$ are identified only if $\mathrm{var} \left( {\Greekmath 010C} _{i}\right) >0$. In practice, a test for $\mathbb{H}_{0}:\mathrm{var}\left( {\Greekmath 010C} _{i}\right) =0$ is possible, by noting that $\mathrm{var}\left( {\Greekmath 010C} _{i}\right) =0$ is equivalent to \begin{equation*} {\Greekmath 0114} ^{2}=\frac{\mathrm{E}\left( {\Greekmath 010C} _{i}\right) ^{2}}{\mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) }=1, \end{equation*} where ${\Greekmath 0114} ^{2}$ is well-defined as long as ${\Greekmath 010C} _{i}\not\equiv 0$. One important advantage of basing the test of slope homogeneity on ${\Greekmath 0114} ^{2}$ rather than on $var({\Greekmath 010C} _{i})=0$, is that ${\Greekmath 0114} ^{2}\,$is scale-invariant. $\mathrm{E}\left( {\Greekmath 010C} _{i}\right) $ and $\mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) $ are identified as in Section (ref), whose consistent estimation does not require $\mathrm{var}\left( {\Greekmath 010C} _{i}\right) >0$. Consequently, in principle it is possible to test slope homogeneity by testing $\mathbb{H}_{0}:{\Greekmath 0114} ^{2}=1$ . However, the problem becomes much more complicated when there are more than two categories and/or there are more than one regressor under consideration. A full treatment of testing slope homogeneity in such general settings is beyond the scope of the present paper.
remarkNote that in the special case of the proof of Theorem (ref) where $K=2$, ${\Greekmath 010C}_{L+H}={\Greekmath 010C}_{L}+ {\Greekmath 010C}_{H}$ and ${\Greekmath 010C}_{LH}={\Greekmath 010C}_{L}{\Greekmath 010C}_{H}$ corresponds to the $ b_{1}^{\ast}$ and $b_{2}^{\ast}$ and (ref) is the same as (ref) when $K = 2$. The special case illustrates the procedure of identification: identify $\left(b_{k}^{\ast} \right)_{k=1}^{K}$ by the moments of ${\Greekmath 010C}_{i}$, then solve for $ \left(b_{k}\right)_{k=1}^{K}$ and finally identify $\left({\Greekmath 0119}_{k} \right)_{k=1}^{K}$.

Estimation

In this section, we propose a generalized method of moments estimator for the distributional parameters of ${\Greekmath 010C}_i$. To reduce the complexity of the moment equations, we first obtain a $\sqrt{n}$-consistent estimator of $ {\Greekmath 010D}$ and consider the estimation of the distribution of ${\Greekmath 010C}_i$ by replacing ${\Greekmath 010D}$ by $\hat{{\Greekmath 010D}}$.

Estimation of $\protect{\Greekmath 010D}$

Let $\mathbf{{\Greekmath 011E} }=\left( \mathrm{E}\left( {\Greekmath 010C} _{i}\right) ,\mathbf{ {\Greekmath 010D} }^{\prime }\right) ^{\prime }$, $v_{i}={\Greekmath 010C} _{i}-\mathrm{E}\left( {\Greekmath 010C} _{i}\right) $ and using the notation in Assumption (ref), (ref) can be written as

equation[equation omitted — 121 chars of source]

where ${\Greekmath 0118} _{i}=u_{i}+x_{i}v_{i}$. Then $\mathbf{{\Greekmath 011E} }$ can be estimated consistently by $\hat{\mathbf{{\Greekmath 011E} }}=\mathbf{Q}_{n,ww}^{-1}\mathbf{q} _{n,wy} $ where $\mathbf{Q}_{n,ww}$ and $\mathbf{q}_{n,wy}$ are defined in Assumption (ref).

assumption$\left\Vert n^{-1}\sum_{i=1}^{n}\mathrm{E} \left( \mathbf{w}_i\mathbf{w}_i^\prime {\Greekmath 0118}_i^2\right) - \mathbf{V}_{w{\Greekmath 0118}} \right\Vert = O\left( n^{-1/2} \right) $, $\mathbf{V}_{w{\Greekmath 0118}}\succ 0, $ and \begin{equation} \left\Vert \frac{1}{n}\sum_{i=1}^{n} \mathbf{w}_i\mathbf{w}_i^\prime {\Greekmath 0118}_i^2 - \frac{1}{n}\sum_{i=1}^{n}\mathrm{E}\left( \mathbf{w}_i\mathbf{w}_i^\prime {\Greekmath 0118}_i^2\right) \right\Vert = O_p\left( n^{-1/2} \right). \end{equation}
remarkAs in the case of Assumption (ref), the high level condition (ref) can be shown to hold under weak cross-sectional dependence, assuming that elements of $\mathbf{w}_{i} \mathbf{w}_{i}^{\prime }{\Greekmath 0118} _{i}^{2}$ are cross-sectionally weakly correlated over $i$. See Remark (ref) .
theoremUnder Assumption (ref), $\hat{\mathbf{{\Greekmath 011E}}}$ is a consistent estimator for $\mathbf{{\Greekmath 011E}}$. In addition, under Assumptions (ref) and (ref) , as $n\to \infty$, \begin{equation} \sqrt{n}\left( \hat{\mathbf{{\Greekmath 011E}}} - \mathbf{{\Greekmath 011E}} \right) \rightarrow_{d} N\left( \mathbf{0}, \mathbf{V}_{\Greekmath 011E} \right), \end{equation} where $\mathbf{V}_{{\Greekmath 011E}} = \mathbf{Q}_{w w}^{-1} \mathbf{V}_{w{\Greekmath 0118}} \mathbf{Q} _{w w}^{-1}. $ $\mathbf{V}_{{\Greekmath 011E}}$ is consistently estimated by \begin{equation*} \hat{\mathbf{V}}_{{\Greekmath 011E}} = \mathbf{Q}_{n,ww}^{-1}\hat{\mathbf{V}}_{w{\Greekmath 0118}} \mathbf{Q}_{n,ww}^{-1}\to_p \mathbf{V}_{{\Greekmath 011E}}, \end{equation*} as $n\to\infty$, where $\hat{\mathbf{V}}_{w{\Greekmath 0118}} = n^{-1}\sum_{i=1}^{n} \mathbf{w}_i\mathbf{w}_i^\prime \hat{{\Greekmath 0118}}_{i}^2$, and $\hat{{\Greekmath 0118}}_i = y_i - \mathbf{w}_i^\prime \hat{\mathbf{{\Greekmath 011E}}}$.

The proof of Theorem (ref) is provided in Section (ref) in the online supplement.

Estimation of the distribution of $\protect{\Greekmath 010C}_i$

Denote the moments of ${\Greekmath 010C} _{i}$ on the right-hand side of ((ref)) by

equation*[equation* omitted — 416 chars of source]

and note that

equation[equation omitted — 498 chars of source]

so in general we can write $\mathbf{m}_{{\Greekmath 010C} }\triangleq h\left( \mathbf{ {\Greekmath 0112} }\right) ,$ where $\mathbf{{\Greekmath 0112} }=\left( \mathbf{{\Greekmath 0119} }^{\prime }, \mathbf{b}^{\prime }\right) ^{\prime }\in \Theta $, and $\mathbf{{\Greekmath 0112} }$ can be uniquely determined in terms of $\mathbf{m}_{{\Greekmath 010C} }$ by Theorem (ref). To estimate $\mathbf{{\Greekmath 0112} }$, we consider moment conditions following a similar procedure as in Section (ref), and propose a generalized method of moments (GMM) estimator.

We consider the following moment conditions

equation*[equation* omitted — 233 chars of source]

and

equation[equation omitted — 203 chars of source]

where $\mathrm{E}\left( u_{i}\right) =0$, $\tilde{y}_{i}=y_{i}-\mathbf{z} _{i}^{\prime }\mathbf{{\Greekmath 010D} }$, $r=1,2,...,2K-1$, and $s_{r}=0,1,\cdots ,S-r $, where $S$ is a user-specific tuning parameter, chosen such that the highest order moments of $x_{i}$ included is at most $S$, where $S>2K-1$. \footnote{ For identification, we require the moments of $x_{i}$ to exist up to order $ 4K-2$. $S$ can take values between $2K$ to $4K-2$. In practice, the choice of $S$ affects the trade-off between bias and efficiency.}

Let ${\Greekmath 011B} _{0}=1$ and ${\Greekmath 011B} _{1}=0$ such that ${\Greekmath 011B} _{r}$ is well-defined for $r=0,1,\cdots ,2K-1$.\ Sum (ref) over $i$ and rearrange terms,

align[align omitted — 537 chars of source]

where

equation*[equation* omitted — 270 chars of source]

as shown in the proof of Theorem (ref).

Taking $n\to \infty$ in (ref),

equation[equation omitted — 158 chars of source]

by Assumption (ref). We stack the left-hand side of (ref) over $r=1,2,...,2K-1$, and $s_{r}=0,1,\cdots ,S-r$ and transform $\mathbf{m}_{\Greekmath 010C} = h\left( \mathbf{{\Greekmath 0112}} \right)$ to get $ \mathbf{g}_0\left(\mathbf{{\Greekmath 0112} }, \mathbf{{\Greekmath 011B}}, \mathbf{{\Greekmath 010D} } \right) $.

To implement the GMM estimation we replace $\tilde{y}_{i}$, by $\hat{\tilde{y }}_{i}=y_{i}-\mathbf{z}_{i}^{\prime }\hat{\mathbf{{\Greekmath 010D} }}$, and ${\Greekmath 011A} _{r,s_{r}}$ by $n^{-1}\sum_{i=1}^{n}\hat{\tilde{y}}_{i}^{r}x_{i}^{s_{r}}$. Noting that $\mathbf{m}_{{\Greekmath 010C} }=h\left( \mathbf{{\Greekmath 0112} }\right) $, denote the sample version of the left-hand side of (ref) by

equation[equation omitted — 312 chars of source]

where

equation*[equation* omitted — 328 chars of source]

and $\mathbf{{\Greekmath 011B} }=\left( {\Greekmath 011B} _{2},{\Greekmath 011B} _{3},\cdots ,{\Greekmath 011B} _{2K-1}\right) ^{\prime }$. Stack the equations in ((ref)), over $ r=0,1,...,2K-1$ and $s_{r}=0,1,\cdots ,S-r$ ($S>2K-1$), in vector notations we have

equation[equation omitted — 304 chars of source]

Given $\hat{\mathbf{{\Greekmath 010D} }}$, the GMM estimator of $\left( \mathbf{{\Greekmath 0112} } ^{\prime },\mathbf{{\Greekmath 011B} }^{\prime }\right) ^{\prime }$ is now computed as

equation*[equation* omitted — 344 chars of source]

where $\hat{\Phi}_{n}=\mathbf{\hat{g}}_{n}\left( \mathbf{{\Greekmath 0112} },\mathbf{ {\Greekmath 011B} },\hat{\mathbf{{\Greekmath 010D} }}\right) ^{\prime }\mathbf{A}_{n}\mathbf{\hat{g }}_{n}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\hat{\mathbf{{\Greekmath 010D} }}\right) $, and $\mathbf{A}_{n}$ is a positive definite matrix. We follow the GMM literature using the following choice of $\mathbf{A}_{n}$,

equation[equation omitted — 460 chars of source]

where $\mathbf{\bar{g}}_{n}=\frac{1}{n}\sum_{i=1}^{n}\mathbf{\hat{g}} _{i}\left( \tilde{\mathbf{{\Greekmath 0112} }},\tilde{\mathbf{{\Greekmath 011B} }},\hat{\mathbf{ {\Greekmath 010D} }}\right) $, and $\tilde{\mathbf{{\Greekmath 0112} }}$ and $\tilde{\mathbf{ {\Greekmath 011B} }}$ are preliminary estimators.

assumptionDenote the true values of $\mathbf{{\Greekmath 0112} }$, $\mathbf{{\Greekmath 011B}}$ and $\mathbf{{\Greekmath 010D}}$ by $\mathbf{{\Greekmath 0112} }_{0}$, $ \mathbf{{\Greekmath 011B}}_0$ and $\mathbf{{\Greekmath 010D}}_0 $. \begin{enumerate} • $\Theta $ and $\mathcal{S}$ are compact. $\mathbf{{\Greekmath 0112} }_{0}\in \mathrm{int}\left( \Theta \right)$ and $\mathbf{{\Greekmath 011B}}_0 \in \mathrm{int} \left( \mathcal{S} \right)$. • $\mathbf{A}_{n}\rightarrow _{p}\mathbf{A}$ as $n\rightarrow \infty $, where $\mathbf{A}$ is some positive definite matrix. • $\,$ \begin{equation*} \frac{1}{n}\sum_{i=1}^{n}\left[ \hat{\tilde{y}}_{i}^{r}x_{i}^{s_{r}}-\mathrm{ E}\left( \tilde{y}_{i}^{r}x_{i}^{s_{r}}\right) \right] =O_{p}\left( n^{-1/2}\right) , \end{equation*} for $r=0,1,2,\cdots ,2K-1$, $s_{r}=0,1,\cdots ,S-r,$ and $S>2K-1.$ \end{enumerate}
remarkParts (a) and (b) of Assumption (ref) are standard regularity conditions in the GMM literature. Part (c) together with Assumption (ref) are high-level regularity conditions which allow us to generalize the usual IID assumption and nest the IID data generation process as a special case. The sample analogue terms in (c) include $\hat{\tilde{y}}_{i}=y_{i}-\mathbf{z}_{i}^{\prime }\hat{\mathbf{ {\Greekmath 010D} }}$, instead of the infeasible $\tilde{y}_{i}=y_{i}-\mathbf{z} _{i}^{\prime }\mathbf{{\Greekmath 010D} }$. The $\sqrt{n}$-consistency of $\hat{\mathbf{ {\Greekmath 010D} }}$ shown in Theorem (ref) ensures that replacing $\tilde{y}_{i}$ by $\hat{\tilde{y}}_{i}$ does not alter the convergence rate.
theoremLet $\mathbf{{\Greekmath 0111} }=\left( \mathbf{{\Greekmath 0112} }^{\prime }, \mathbf{{\Greekmath 011B} }^{\prime }\right) ^{\prime }$ and $\mathbf{{\Greekmath 0111} }_{0}=\left( \mathbf{{\Greekmath 0112} }_{0}^{\prime },\mathbf{{\Greekmath 011B} }_{0}^{\prime }\right) ^{\prime }$. Under Assumptions (ref) , (ref), and (ref), $\hat{ \mathbf{{\Greekmath 0111} }}\rightarrow _{p}\mathbf{{\Greekmath 0111} }_{0}$ as $n\rightarrow \infty $.

The proof of Theorem (ref) is provided in Appendix (ref).

assumptionFollow the notations as in Assumption (ref) and in addition denote $\mathbf{G}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\mathbf{{\Greekmath 010D} }\right) ={\Greekmath 0272} _{\left( \mathbf{{\Greekmath 0112} }^{\prime }, \mathbf{{\Greekmath 011B} }^{\prime }\right) ^{\prime }} \mathbf{g}_{0}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\mathbf{{\Greekmath 010D} } \right) $, $\mathbf{G}_{0}=\mathbf{G}\left( \mathbf{{\Greekmath 0112} }_{0},\mathbf{ {\Greekmath 011B} }_{0},\mathbf{{\Greekmath 010D} }_{0}\right) $, $\mathbf{G}_{{\Greekmath 010D} }\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\mathbf{{\Greekmath 010D} }\right) ={\Greekmath 0272} _{\mathbf{ {\Greekmath 010D} }}\mathbf{g}_{0}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\mathbf{ {\Greekmath 010D} }\right) $, $\mathbf{G}_{0, {\Greekmath 010D}}=\mathbf{G}_{{\Greekmath 010D}} \left( \mathbf{{\Greekmath 0112} }_{0},\mathbf{{\Greekmath 011B} }_{0},\mathbf{{\Greekmath 010D} }_{0}\right) $. \begin{enumerate} • $\sqrt{n}\mathbf{\hat{g}}_{n}\left( \mathbf{{\Greekmath 0112} }_{0},\mathbf{ {\Greekmath 011B} }_{0},\mathbf{{\Greekmath 010D}}_0\right) \rightarrow _{d}\mathbf{{\Greekmath 0110} }\sim N\left( 0,\mathbf{V}\right) $ as $n\to\infty$. • $\mathbf{G}_{0}^{\prime }\mathbf{AG}_{0}\succ 0$. \end{enumerate}
remarkIn Assumption (ref), parts (a) is the high level condition required to ensure the asymptotic normality of $\mathbf{\hat{g}}_{n}\left( \mathbf{{\Greekmath 0112} }_{0},\mathbf{{\Greekmath 011B} }_{0},\mathbf{{\Greekmath 010D}}_0\right) $, which can be verified by Lindeberg central limit theorem under low-level regularity conditions. Part (c) of Assumption (ref) represents the full-rank condition on $\mathbf{G}_{0}$, required for identification of $\mathbf{{\Greekmath 0112} }_{0}$ and $\mathbf{{\Greekmath 011B} }_{0}$.

By Theorem (ref), we have $\sqrt{n}\left( \hat{ \mathbf{{\Greekmath 010D}}} - \mathbf{{\Greekmath 010D}} \right) \rightarrow_d {\Greekmath 0110}_{\Greekmath 010D} \sim N(0, V_{\Greekmath 010D})$. The following theorem shows the asymptotic normality of the GMM estimator $\hat{\mathbf{{\Greekmath 0111}}}$.

theoremUnder Assumptions (ref), (ref), (ref) and (ref), \begin{equation*} \sqrt{n}\left( \hat{\mathbf{{\Greekmath 0111} }}-\mathbf{{\Greekmath 0111} }_{0}\right) \rightarrow _{d}\left( \mathbf{G}_{0}^{\prime }\mathbf{A}\mathbf{G}_{0}\right) ^{-1} \mathbf{G}_{0}^{\prime }\mathbf{A}\left( \mathbf{{\Greekmath 0110} }+\mathbf{G}_{0, {\Greekmath 010D}}\mathbf{{\Greekmath 0110} }_{{\Greekmath 010D} }\right), \end{equation*} as $n\rightarrow \infty $.

The proof of Theorem (ref) is provided in Appendix (ref).

remarkIn practice, we estimate the variance of the asymptotic distribution of $ \hat{{\Greekmath 0111}}$ by \begin{equation} \hat{\mathbf{V}}_{{\Greekmath 0111}} = \left( \hat{\mathbf{G}}^\prime \hat{\mathbf{A}}_n \hat{\mathbf{G}} \right)^{-1} \hat{\mathbf{G}}^\prime \hat{\mathbf{A}}_n \hat{\mathbf{V}}_{{\Greekmath 0110}} \hat{\mathbf{A}}_n^\prime \hat{\mathbf{G}} \left( \hat{\mathbf{G}}^\prime \hat{\mathbf{A}}_n \hat{\mathbf{G}} \right)^{-1}, \end{equation} where $\hat{\mathbf{G}} = {\Greekmath 0272}_{\left( \mathbf{{\Greekmath 011B} }^{\prime },\mathbf{ {\Greekmath 0112} }^{\prime }\right) ^{\prime }} \hat{\mathbf{g}}_{n}\left( \hat{ \mathbf{{\Greekmath 0112}}}, \hat{\mathbf{{\Greekmath 011B}}}, \hat{{\Greekmath 010D}} \right)$, $\hat{ \mathbf{A}}_n$ is given by (ref), and \begin{equation*} \hat{\mathbf{V}}_{\Greekmath 0110} = \frac{1}{n} \sum_{i=1}^{n} \mathbf{{\Greekmath 0120}}_{n,i} \mathbf{{\Greekmath 0120}}_{n,i}^\prime, \end{equation*} where \begin{equation*} \mathbf{{\Greekmath 0120}}_{n,i} = \hat{\mathbf{g}}_i\left( \hat{\mathbf{{\Greekmath 0112}}}, \hat{ \mathbf{{\Greekmath 011B}}}, \hat{{\Greekmath 010D}} \right) + {\Greekmath 0272} _{\mathbf{{\Greekmath 010D} }}\hat{ \mathbf{g}}_{n}\left( \hat{\mathbf{{\Greekmath 0112}}}, \hat{\mathbf{{\Greekmath 011B}}}, \hat{ {\Greekmath 010D}}\right) \mathbf{L} \mathbf{Q}_{n,ww}^{-1}\left( \mathbf{w}_i\hat{{\Greekmath 0118}}_i \right), \end{equation*} and $\mathbf{L} = \begin{pmatrix} \mathbf{0}_{p_z\times 1} & \mathbf{I}_{p_z} \end{pmatrix} $ is the loading matrix that selects $\mathbf{{\Greekmath 010D}}$ out of $\mathbf{{\Greekmath 011E}}$ .

Multiple regressors with random coefficients

One important extension of the regression model (ref) is to allow for multiple regressors with random coefficients having categorical distribution. With this in mind consider

equation[equation omitted — 166 chars of source]

where the $p\times 1$ vector of random coefficients, $\mathbf{{\Greekmath 010C} }_{i}\in \mathbb{R}^{p}$ follows the multivariate distribution\footnote{ We assume the number of categories $K$ is homogeneous across $j=1,2,\cdots ,p $. This is for notational simplicity, and can be readily generalized to allow for $K_{j}\neq K_{j^{\prime }}$ without affecting the main results.}

equation[equation omitted — 227 chars of source]

with $k_{j}\in \left\{ 1,2,\cdots ,K\right\} $, $b_{j1}<b_{j2}<\cdots <b_{jK} $, and

equation*[equation* omitted — 132 chars of source]

As in Section (ref), $\mathbf{{\Greekmath 010D} }\in \mathbb{R} ^{p_{z}}$, $\mathbf{w}_{i}=\left( \mathbf{x}_{i}^{\prime },\mathbf{z} _{i}^{\prime }\right) ^{\prime }$, $\mathbf{{\Greekmath 010C} }_{i}\perp \mathbf{w}_{i}$ , $u_{i}\perp \mathbf{w}_{i}$, and $u_{i}$ are independently distributed over $i$ with mean $0.$

Example 1 Consider the simple case with $p = 2$ and $K = 2$. For $j = 1, 2$, denote two categories as $\left\{ L, H \right\}$ . The probabilities of four possible combinations of realized $\mathbf{{\Greekmath 010C}} _i$ is summarized in Table (ref), where ${\Greekmath 0119}_{LL} + {\Greekmath 0119}_{LH} + {\Greekmath 0119}_{HL} + {\Greekmath 0119}_{HH} = 1$.

table[table omitted — 734 chars of source]

We first identify the moments of $\mathbf{{\Greekmath 010C} }_{i}$. As in Section (ref), $\mathbf{{\Greekmath 011E} }=\left( \mathrm{E}\left( \mathbf{{\Greekmath 010C} } _{i}\right) ^{\prime },\mathbf{{\Greekmath 010D} }^{\prime }\right) ^{\prime }$ is identified by

equation[equation omitted — 113 chars of source]

under Assumption (ref). We now consider the identification of the higher order moments of $\mathbf{{\Greekmath 010C} } _{i}$ up to the finite order $2K-1$.

Since $\mathbf{{\Greekmath 010D} }$ is identified as in (ref), we treat it as known and let $\tilde{y}_{i}^{r}=y_{i}-\mathbf{z}_{i}^{\prime } \mathbf{{\Greekmath 010D} }$. For $r=2,3,\cdots ,2K-1$, consider the moment conditions

align[align omitted — 506 chars of source]

Note that $\mathbf{x}_{i}^{\prime }\mathbf{{\Greekmath 010C} }_{i}=\sum_{j=1}^{p}{\Greekmath 010C} _{ij}x_{ij}$, and

equation*[equation* omitted — 280 chars of source]

where $\binom{r}{\mathbf{q}}=\frac{r!}{q_{1}!q_{2}!\cdots q_{p}!}$, for non-negative integers $r$, $q_{1}$, $\cdots $, $q_{p}$ with $ r=\sum_{j=1}^{p}q_{j}$, denotes the multinomial coefficients. We stack $ \prod_{j=1}^{p}x_{ij}^{q_{j}}$ with $\mathbf{q}\in \left\{ \mathbf{q}\in \left\{ 0,1,\cdots r\right\} ^{p}:\sum_{j=1}^{p}q_{j}=r\right\} $ in a vector form by denoting \footnote{ For $\mathbf{x}\in \mathbb{R}^{p}$, note that $\mathbf{{\Greekmath 011C} }_{0}\left( \mathbf{x}\right) =1$, $\mathbf{{\Greekmath 011C} }_{1}\left( \mathbf{x}\right) =\mathbf{x }$ and $\mathbf{{\Greekmath 011C} }_{2}\left( \mathbf{x}\right) =\mathrm{vech}\left( \mathbf{x}\mathbf{x}^{\prime }\right) $.}

equation*[equation* omitted — 322 chars of source]

where ${\Greekmath 0127} \left( \mathbf{x}_{i},\mathbf{q}\right) =\prod_{j=1}^{p}x_{ij}^{q_{j}}$ and ${\Greekmath 0117} _{r}=\binom{r+p-1}{p-1}$ is the number of distinct monomials of degree $r$ on the variables $ x_{i1},x_{i2},\cdots ,x_{ip}$. Similarly,

equation*[equation* omitted — 391 chars of source]

where ${\Greekmath 0127} \left( \mathbf{{\Greekmath 010C} }_{i},\mathbf{q}\right) =\prod_{j=1}^{p}{\Greekmath 010C} _{ij}^{q_{j}}$.

Example 2 Consider $p = 2$ and $r = 2$, we have

align*[align* omitted — 328 chars of source]

and

align*[align* omitted — 970 chars of source]

where $\mathbf{\Lambda}_2 = \mathrm{diag}\left[ \left( 1, 2, 1 \right)^\prime \right]$.}

Then the moment condition (ref) can be written as

align[align omitted — 665 chars of source]

where $\mathbf{\Lambda }_{r}=\mathrm{diag}\left[ \left[ \binom{r}{\mathbf{q}} \right] _{\sum_{j=1}^{p}q_{j}=r}\right] $ is the ${\Greekmath 0117} _{r}\times {\Greekmath 0117} _{r}$ diagonal matrix of multinomial coefficients. We further consider the moment conditions

align[align omitted — 936 chars of source]

$r=2,3,\cdots ,2K-1$. (ref) and (ref) reduce to (ref) and (ref) when $p=1$.

assumption$\,$ \begin{enumerate} • $\left\Vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( \tilde{y}_{i}^{r} \mathbf{{\Greekmath 011C} }_{s}\left( \mathbf{x}_{i}\right) \right) -\mathbf{{\Greekmath 011A} } _{r,s}\right\Vert =O\left( n^{-1/2}\right) ,$ and $\left\Vert \mathbf{{\Greekmath 011A} } _{r,s}\right\Vert <\infty $, $r,s=0,1,\cdots ,2K-1$. • $\left\Vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} } _{r}\left( \mathbf{x}_{i}\right) \mathbf{{\Greekmath 011C} }_{s}\left( \mathbf{x} _{i}\right) ^{\prime }\right] -\mathbf{\Xi }_{r,s}\right\Vert =O\left( n^{-1/2}\right) ,$ and $\left\Vert \mathbf{\Xi }_{r,s}\right\Vert <\infty $, $r,s=0,1,\cdots ,2K-1$. • $\left\vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( u_{i}^{r}\right) -{\Greekmath 011B} _{r}\right\vert =O\left( n^{-1/2}\right) ,$ and $\left\vert {\Greekmath 011B} _{r}\right\vert <\infty $ for $r=2,3,\cdots ,2K-1$. • $\left\Vert n^{-1}\sum_{i=1}^{n} \left[ \mathrm{var} \left( \mathbf{{\Greekmath 011C}}_r \left(\mathbf{x_i} \right) \right) - \left( \mathbf{\Xi} _{r,r} - \mathbf{{\Greekmath 011A}}_{0, r}\mathbf{{\Greekmath 011A}}_{0, r}^\prime \right) \right] \right\Vert = O(n^{-1/2})$, where $\mathbf{\Xi}_{r,r} - \mathbf{{\Greekmath 011A}}_{0, r} \mathbf{{\Greekmath 011A}}_{0, r}^\prime \succ 0$ for $r=2,3\cdots, 2K-1$. \end{enumerate}
theoremFor any $\mathbf{q} \in \left\{ \mathbf{q}\in \left\{ 0, 1, \cdots r \right\} ^{p}: \sum_{j=1}^p q_j = r\right\}$ and $r = 2, 3,\cdots, 2K-1$, $\mathrm{E} \left(\prod_{j=1}^{p} {\Greekmath 010C}_{ij}^{q_{j}}\right)$ and ${\Greekmath 011B}_r$ are identified under Assumptions (ref) and (ref).
proofFor $r=2,3,\cdots ,2K-1$, sum (ref) and (ref) over $i,$ go through the same steps as in the proof of Theorem (ref), then by Assumptions (ref)(a) to (c), we have (for $ n\rightarrow \infty $) \begin{align} \mathbf{{\Greekmath 011A} }_{r,0}^{\prime }\mathbf{\Lambda }_{r}\mathrm{E}\left[ \mathbf{ {\Greekmath 011C} }_{r}\left( \mathbf{{\Greekmath 010C} }_{i}\right) \right] +{\Greekmath 011B} _{r}& =\mathbf{ {\Greekmath 011A} }_{r,0}-\sum_{s=2}^{r-1}\binom{r}{s}\mathbf{{\Greekmath 011A} }_{0,r-s}\mathbf{ \Lambda }_{r-s}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r-s}\left( \mathbf{{\Greekmath 010C} } _{i}\right) \right] {\Greekmath 011B} _{s}, \\ \mathbf{\Xi }_{r,r}\mathbf{\Lambda }_{r}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} } _{r}\left( \mathbf{{\Greekmath 010C} }_{i}\right) \right] +\mathbf{{\Greekmath 011A} }_{0,r}{\Greekmath 011B} _{r}& =\mathbf{{\Greekmath 011A} }_{r,r}-\sum_{s=2}^{r-1}\binom{r}{s}\mathbf{\Xi }_{r,r-s} \mathbf{\Lambda }_{r-s}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r-s}\left( \mathbf{ {\Greekmath 010C} }_{i}\right) \right] {\Greekmath 011B} _{s}. \end{align} Note that \begin{equation*} \mathbf{M}_{r}= \begin{pmatrix} \mathbf{\Xi }_{r,r} & \mathbf{{\Greekmath 011A} }_{0,r} \\ \mathbf{{\Greekmath 011A} }_{0,r}^{\prime } & 1 \end{pmatrix} \begin{pmatrix} \mathbf{\Lambda }_{r} & \mathbf{0} \\ \mathbf{0} & 1 \end{pmatrix} , \end{equation*} is invertible since $\det \left( \mathbf{M}_{r}\right) =\det \left( \mathbf{ \Xi }_{r,r}-\mathbf{{\Greekmath 011A} }_{0,r}\mathbf{{\Greekmath 011A} }_{0,r}^{\prime }\right) \det \left( \mathbf{\Lambda }_{r}\right) >0,$ for $r=2,3,\cdots ,R$, by Assumption (ref)(d). As a result, we can sequentially solve (ref) and (ref) for $\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{{\Greekmath 010C} }_{i}\right) \right] $ and ${\Greekmath 011B} _{r}$, for $r=2,3,\cdots ,2K-1$.

We now move from the moments of $\mathbf{{\Greekmath 010C} }_{i}$ to the distribution of $\mathbf{{\Greekmath 010C} }_{i}$. We first focus on the identification of the marginal probabilities obtained from ((ref)) by averaging out the effects of the other coefficients except for ${\Greekmath 010C} _{ij}$, namely we initially focus on identification of ${\Greekmath 0115} _{jk}=\Pr \left( {\Greekmath 010C} _{ij}=b_{jk}\right) $, for $k=1,2,\cdots ,K,$ and $j=1,2,\cdots ,p$.

remarkFocusing on the marginal distribution of ${\Greekmath 010C} _{i}$ is similar to focusing on estimation of partial derivatives in the context of non-parametric estimation, where the curse of dimensionality applies. Consider the estimation of regressing $y_{i}$ on $\mathbf{x}_{i}=\left( x_{i1},x_{i2},\cdots ,x_{ip}\right) ^{\prime }$, \begin{equation*} y_{i}=F\left( x_{i1},x_{i2},\cdots .x_{ip}\right) +u_{i}. \end{equation*} Then if $F\left( x_{1},x_{i2},\cdots ,x_{ip}\right) $ is a homogeneous function (of degree $1/{\Greekmath 0116} $), then \begin{equation*} y_{i}=\sum_{j=1}^{p}\left( {\Greekmath 0116} \frac{\partial F\left( \cdot \right) }{ \partial x_{ij}}\right) x_{ij}+u_{i}, \end{equation*} and under certain conditions we can treat ${\Greekmath 0116} \frac{\partial F\left( \cdot \right) }{\partial x_{ij}}\equiv {\Greekmath 010C} _{ij}$.

By Theorem (ref), $\mathrm{E}\left( {\Greekmath 010C} _{ij}^{r}\right) $ is identified for $r=1,2,\cdots ,2K-1$ under Assumptions (ref) and (ref). By (ref), we have equations

equation[equation omitted — 139 chars of source]

$r=0,1,\cdots ,2K-1$, which is of the same form as (ref) and (ref). To identify $\mathbf{{\Greekmath 0115} }_{j}=\left( {\Greekmath 0115} _{j1},{\Greekmath 0115} _{j2},\cdots ,{\Greekmath 0115} _{jK}\right) ^{\prime }$ and $\mathbf{b} _{j}=\left( b_{j1},b_{j2},\cdots ,b_{jK}\right) ^{\prime }$, we can verify the system of $2K$ equations in (ref) has a unique solution if $b_{j1}<b_{j2}<\cdots <b_{jK}$ and ${\Greekmath 0115} _{jk}\in \left( 0,1\right) $. The following corollary is a direct application of Theorem (ref).

corollaryConsider the model (ref) and suppose that Assumptions (ref) and (ref) hold. Then the parameters $\mathbf{{\Greekmath 0112}}_j = \left( \mathbf{{\Greekmath 0115}}_{j}^\prime, \mathbf{b}_{j}^\prime \right)^\prime$ of the marginal distribution of ${\Greekmath 010C}_i$ with respect to ${\Greekmath 010C}_{ij}$ is identified subject to $b_{j1} < b_{j2} < \cdots < b_{jK}$ and ${\Greekmath 0115}_{jk} \in \left( 0, 1 \right)$ for $j = 1, 2, \cdots, p$.

The problem of identification and estimation of the joint distribution of $ \mathbf{{\Greekmath 010C} }_{i}$ is subject to the curse of dimensionality. We have $ K^{p}-1$ probability weights, ${\Greekmath 0119} _{k_{1},k_{2},\cdots ,k_{p}}$, to be identified in addition to the $pK$ categorical coefficients $b_{ij}$ that are identified by Corollary (ref). The number of parameters increases rapidly with $p$. Even in the simplest case with $K=2$, the total number of unknown parameters is $2p+2^{p}-1$, which grows exponentially.

Note that the marginal probabilities ${\Greekmath 0115} _{jk}$ are related to the joint distribution by

equation[equation omitted — 228 chars of source]

$k=1,2,\cdots ,K$ and $j=1,2,\cdots ,p$. The number of linearly independent equations in (ref) is $pK-(p-1)$.

\noindentExample 3 \ Consider the same setup as in Example 1 with $p = 2$ and $K = 2$. The marginal probabilities are obtained by

align[align omitted — 633 chars of source]

Note that any equation in (ref) can be expressed as a linear combination of other three equations, for example ${\Greekmath 0115}_{2H} = {\Greekmath 0115}_{1L} + {\Greekmath 0115}_{1H} - {\Greekmath 0115}_{2L}$. }

The equations corresponding to the cross-moments, $\mathrm{E}\left( \prod_{j=1}^{p}{\Greekmath 010C} _{ij}^{q_{j}}\right) $, are

equation[equation omitted — 274 chars of source]

for $\mathbf{q}\in \left\{ \mathbf{q}\in \left\{ 0,1,\cdots r-1\right\} ^{p}:\sum_{j=1}^{p}q_{j}=r\right\} $, $r=2,\cdots ,2K-1$. The linear system (ref) has

equation*[equation* omitted — 60 chars of source]

equations. Then the total number of equations in (ref) and (ref) that can be utilized to identify joint probabilities is $C_{r}=\sum_{r=1}^{2K-1}\binom{ r+p-1}{p-1}-pK$, which is smaller than the number of joint probabilities $ K^{p}-1$ for large $p$. When $K=2$, $C_{r}<K^{p}-1$ for $p\geq 7$.

Identification and estimation of the joint distribution of $\mathbf{{\Greekmath 010C} } _{i}$ in the general setting will not be pursued in this paper due to the curse of dimensionality. Instead, we consider special cases, that are empirically relevant, in which identification of the joint distribution of $ \mathbf{{\Greekmath 010C} }_{i}$ can be readily established. We first consider small $p$ and $K$, in particular $p=2$ and $K=2$ as in Example 1.

Example 4 Consider the same setup as in Example 1 with $p=2$ and $K=2$. In addition to (ref), consider the cross-moment,

equation[equation omitted — 262 chars of source]

Writing (ref) and (ref) in matrix form, we have

equation*[equation* omitted — 83 chars of source]

where

equation*[equation* omitted — 550 chars of source]

Note that $\mathrm{E}\left( {\Greekmath 010C} _{i1}{\Greekmath 010C} _{i2}\right) $ is identified by Theorem (ref), and $b_{jk_{j}}$ and $ {\Greekmath 0115} _{jk_{j}}$ are identified by Corollary (ref), and matrix $\mathbf{B}$ is invertible given that $b_{1L}<b_{1H}$ and $ b_{2L}<b_{2H}$. (See Appendix (ref)). As a result, the joint probabilities, $\mathbf{{\Greekmath 0119} },$ are identified.

remarkThe argument in Example 4 is applicable for identification of the joint distribution of $\left( {\Greekmath 010C}_{ij}, {\Greekmath 010C}_{i,j^\prime} \right)^\prime$ for $ j\neq j^\prime$ when $p > 2$ and $K = 2$.

Finite sample properties using Monte Carlo experiments

We examine the finite sample performance of the categorical coefficient estimator proposed in Section (ref) by Monte Carlo experiments.

Data generating processes

We generate $y_{i}$ as

equation[equation omitted — 237 chars of source]

with ${\Greekmath 010C} _{i}$ distributed as in ((ref)) with $K=2,$ and the parameters ${\Greekmath 0119} ,{\Greekmath 010C} _{L}$ and ${\Greekmath 010C} _{H}$.\footnote{ A Monte Carlo experiment with $K=3$ is relegated to Section (ref) in the online supplement.}

We draw ${\Greekmath 010C} _{i}$ for each individual $i$ independently by setting ${\Greekmath 010C} _{i}={\Greekmath 010C} _{L}$ with probability ${\Greekmath 0119} $ and ${\Greekmath 010C} _{i}={\Greekmath 010C} _{H}$ with probability $1-{\Greekmath 0119} $, through a sequence of independent Bernoulli draws. \ We consider two sets of parameters in all DGPs, denoted as high variance and low variance parametrization, respectively,

equation[equation omitted — 382 chars of source]

${\Greekmath 010C} _{H}/{\Greekmath 010C} _{L}=2$ for the high variance parametrization, and ${\Greekmath 010C} _{H}/{\Greekmath 010C} _{L} = 2.69$, for the low variance parametrization, which is motivated by the estimates in our empirical illustration in Section (ref).\footnote{ The estimates for ${\Greekmath 010C}_H / {\Greekmath 010C}_L$ in our empirical analysis range from 1.50 to 2.79.} The values of E$({\Greekmath 010C} _{i})$ and $\mathrm{var}\left( {\Greekmath 010C} _{i}\right) $ are obtained noting that E$({\Greekmath 010C} _{i})={\Greekmath 0119} {\Greekmath 010C} _{L}+(1-{\Greekmath 0119} ){\Greekmath 010C} _{H}$, and $\mathrm{var}\left( {\Greekmath 010C} _{i}\right) ={\Greekmath 0119} (1-{\Greekmath 0119} )({\Greekmath 010C} _{H}-{\Greekmath 010C} _{L})^{2}$. The remaining parameters are set as ${\Greekmath 010B} =0.25$, and $\mathbf{{\Greekmath 010D} }=\left( 1,1\right) ^{\prime },$ across DGPs.

We generate the regressors and the error terms as follows.

DGP 1 (Baseline) We first generate $\tilde{x}_{i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}(2)$, and then set $x_{i}=(\tilde{x}_{i}-2)/2$ so that $ x_{i}$ has $0$ mean and unit variance. The additional regressors, $z_{ij}$, for $j=1,2$ with homogeneous slopes are generated as

equation*[equation* omitted — 129 chars of source]

with $v_{ij}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID }N\left( 0,1\right) $, for $j=1,2$. This ensures that the regressors are sufficiently correlated. The error term, $u_{i}$, is generated as $u_{i}={\Greekmath 011B} _{i}{\Greekmath 0122} _{i}$, where ${\Greekmath 011B} _{i}^{2}$ are generated as $0.5(1+\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}(1))$, and ${\Greekmath 0122} _{i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}N(0,1)$. Note that ${\Greekmath 0122} _{i}$ and ${\Greekmath 011B} _{i}^{2}$ are generated independently, and $E(u_{i}^{2})=1$.

DGP 2 (Categorical $x$) This setup deviates from the baseline DGP, and allows the distribution of $x_{i}$ to differ across $i$. Accordingly, we generate $x_{i}=\left( \tilde{x}_{1i}-2\right) /2$ where $ \tilde{x}_{1i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}\left( 2\right) $ for $i=1,2,\cdots ,\lfloor n/2\rfloor $, and $x_{i}=\left( \tilde{x}_{2i}-2\right) /4$ where $ \tilde{x}_{2i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}\left( 4\right) $, for $i=\lfloor n/2\rfloor +1,\cdots ,n$. The additional regressors, $z_{ij}$, for $j=1,2$ with homogeneous slopes are generated as

equation*[equation* omitted — 129 chars of source]

with $v_{ij}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID }N\left( 0,1\right) $, for $j=1,2$. The error term $u_{i}$ is generated the same as in DGP 1.

DGP 3 (Categorical $u$) We generate $x_{i}$ and $\mathbf{z} _{i}$ the same as in DGP 1, but allow the error term $u_{i}$ to have a heterogeneous distribution over $i$. For $i=1,2,\cdots ,\lfloor n/2\rfloor $ , we set $u_{i}={\Greekmath 011B} _{i}{\Greekmath 0122} _{i},$ where ${\Greekmath 011B} _{i}^{2}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}\left( 2\right) $ and ${\Greekmath 0122} _{i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID} N(0,1)$, and for $i=\lfloor n/2\rfloor +1,\cdots ,n$, we set $u_{i}=\left( \tilde{u}_{i}-2\right) /2$, where $\tilde{u}_{i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}\left( 2\right) $.

We investigate the finite sample performance of the estimator proposed in Section (ref) across DGP 1 to 3 with low variance and high variance scenarios.\footnote{ We can consider a DGP with conditional heteroskedasticity, in which we follow the baseline DGP and generate the error term as $u_{i}=x_{i} {\Greekmath 0122} _{i}$, where ${\Greekmath 0122} _{i}\sim N(0,1)$. The least square estimator for $\mathbf{{\Greekmath 011E} }$ is valid in this setup in terms of estimation and inference, whereas the GMM estimator for the distributional parameters $ \mathbf{{\Greekmath 0112} }$ breaks down, which is to be expected since we can only identify the first moment of ${\Greekmath 010C} _{i}$ under conditional heteroskedasticity. The results are available on request.} Details of the computational algorithm used to carry out the Monte Carlo experiments (and the empirical results that follow) are given in Section (ref) of the online supplement. An accompanying R package is available at \url{https://github.com/zhan-gao/ccrm}.

Summary of the MC results

table[table omitted — 8,263 chars of source]
figure[figure omitted — 2,099 chars of source]
figure[figure omitted — 2,093 chars of source]
table[table omitted — 8,205 chars of source]
figure[figure omitted — 2,186 chars of source]
figure[figure omitted — 2,184 chars of source]

For each sample size $n = 100$, $1,000$, $2,000$, $5,000$, $10,000$ and $ 100,000$ we run $5,000$ replications of experiments for DGP 1 (baseline), DGP 2 (categorical $x$) and DGP 3 (categorical $u$) with high variance and low variance parametrization, as set out in (ref).

We first investigate the finite sample performance of $\hat{\mathbf{{\Greekmath 011E} }}$ , as an estimator of $\mathbf{{\Greekmath 011E} }=\left( \mathrm{E}\left( {\Greekmath 010C} _{i}\right) ,\mathbf{{\Greekmath 010D} }^{\prime }\right) ^{\prime }$. Bias, root mean squared errors (RMSE) for estimation of $\mathrm{E}\left( {\Greekmath 010C} _{i}\right) $ , ${\Greekmath 010D} _{1}$ and ${\Greekmath 010D} _{2}$, as well as size of testing of the null values at the 5 per cent nominal value are reported in Table (ref). In addition, we plot the associated empirical power functions in Figure (ref) and (ref), for cases of high and low $var({\Greekmath 010C} _{i})$. The results show that $\hat{\mathbf{{\Greekmath 011E} }}$ has very good small sample properties with small bias and RMSEs, with size very close to the nominal value of 5 per cent across all DGPs and parametrization, even when sample size is relatively small. The power of the test increases steadily as the sample size increases.

Then, we turn to the GMM estimator for the distributional parameters of $ {\Greekmath 010C} _{i}$ proposed in Section (ref). The bias, RMSE, and the test size based on the asymptotic distribution given in Theorem (ref), for ${\Greekmath 0119} $, ${\Greekmath 010C} _{L}$ and ${\Greekmath 010C} _{H}$, are reported in Table (ref). The empirical power functions are reported in Figure (ref) and (ref). The reported results are based on $S=4$, where $ S$ $(>2K-1=3)$ denotes the highest order of moments of $x_{i}$ included in estimation.\footnote{ We also tried estimation based on a larger number of moments (using $S=5$ and $S=6$). In the case of current Monte Carlo results, adding more moments does not seem to add much to the precision of the estimates and could be counter-productive when $n$ is not sufficiently large. The results are available in Section (ref) in the online supplement.}

The upper panel of this table reports the results of the high variance and the lower panel for the low variance parametrization, as set out in ((ref)). For all parameters and under all DGPs, the bias and RMSE decline steadily with the sample size as predicted by Theorem (ref), and confirm the robustness of the GMM estimates to the heterogeneity in the regressor and the error processes. But for a given sample size, the relative precision of the estimates depends on the variability of ${\Greekmath 010C} _{i}$, as characterized by the true value of $\mathrm{ var} ({\Greekmath 010C} _{i})$. The precision of the estimates with high variance parametrization is relatively higher than that with low variance parametrization. This is to be expected since, unlike $\mathrm{E} ({\Greekmath 010C} _{i}),$ the distributional parameters are only identified if $\mathrm{var} ({\Greekmath 010C} _{i})>0$. As shown in (ref) and (ref) for the current case of $K=2$, $\mathrm{var}({\Greekmath 010C} _{i})$ is in the denominator when we recover the distributional parameters from the moments of ${\Greekmath 010C} _{i}$. When $\mathrm{var}({\Greekmath 010C} _{i})$ is small, estimation errors in the moments of $ {\Greekmath 010C} _{i}$ can be amplified in the estimation of ${\Greekmath 0119} $, ${\Greekmath 010C} _{L}$ and $ {\Greekmath 010C} _{H}$. On the other hand, the larger the variance the more precisely $ {\Greekmath 0119} $, ${\Greekmath 010C} _{H}$ and ${\Greekmath 010C} _{L}$ can be estimated for a given $n$. \footnote{ Section (ref) in the online supplement presents parametrization with $\mathrm{var}\left( {\Greekmath 010C}_i \right) = 6.35$ and $18.95$ , which further confirms the pattern that the larger the variance the more precisely ${\Greekmath 0119} $, ${\Greekmath 010C} _{H}$ and ${\Greekmath 010C} _{L}$ can be estimated for a given $n$.} The size and power also depends on the parametrization. With both high variance and low variance parametrization, we can achieve correct size and reasonable power when $n$ is quite large ($ n=100,000 $). We plot the empirical power functions for $n\geq 5,000$ for $ {\Greekmath 0119} $, ${\Greekmath 010C} _{H}$ and ${\Greekmath 010C} _{L}$ since the size is far above 5 per cent for smaller values of $n$, and power comparisons are not meaningful in such cases.

remarkNote that GMM estimators of moments of ${\Greekmath 010C} _{i}$, namely $\mathbf{m}_{\mathbf{{\Greekmath 010C} }}$, can be obtained using the moment conditions in (ref),and the transformations $\mathbf{m}_{ \mathbf{{\Greekmath 010C} }}=h\left( \mathbf{{\Greekmath 0112} }\right) $ in (ref) are required only to derive the estimators of $\mathbf{{\Greekmath 0112} }$, the parameters of the underlying categorical distribution. The Monte Carlo results in Section (ref) in the online supplement show that $\mathbf{m}_{\mathbf{{\Greekmath 010C} }}$ can be accurately estimated with relatively small sample sizes. In the estimation of both $\mathbf{m}_{ \mathbf{{\Greekmath 010C} }}$ and $\mathbf{{\Greekmath 0112} }$, the same set of moment conditions are included, so the estimation of distributional parameters $\mathbf{{\Greekmath 0112} }$ essentially relies on the relation $\mathbf{{\Greekmath 0112} }=h^{-1}\left( \mathbf{ m}_{\mathbf{{\Greekmath 010C} }}\right) $. Sampling uncertainties in the estimation of $ \mathbf{m}_{\mathbf{{\Greekmath 010C} }}$, particularly in higher order moments, are potentially amplified through the inverse transformation $h^{-1}$ that involves matrix inversion, which causes the difficulties in estimation and inference of $\mathbf{{\Greekmath 0112} }$ when sample sizes are small. This is analogous to the problem of precision matrix estimation from an estimated covariance matrix. In practice, estimation of the categorical parameters is recommended for applications where the sample size is relatively large, otherwise it is advisable to focus on estimates of the lower order moments of ${\Greekmath 010C} _{i}$.

Heterogeneous return to education: An empirical application

Since the pioneering work by becker1962investment,becker1964book on the effects of investments in human capital, estimating returns to education has been one of the focal points of labor economics research. In his pioneering contribution mincer1974book models the logarithm of earnings as a function of years of education and years of potential labor market experience (age minus years of education minus six), which can be written in a generic form:

equation[equation omitted — 315 chars of source]

as in \citet*[Equation (1)]{heckman2018returns}, where $\mathbf{z}_{i}$ includes the labor market experience and other relevant control variables. The above wage equation, also known as the \textquotedblleft Mincer equation\textquotedblright, has become of the workhorse of the empirical works on estimating the return to education. In the most widely used specification of the Mincer equation (ref),

equation*[equation* omitted — 336 chars of source]

where $\tilde{\mathbf{z}}_{i}$ is the vector of control variables other than potential labor market experience.

Along with the advancement of empirical research on this topic, there has been a growing awareness of the importance of heterogeneity in individual cognitive and non-cognitive abilities heckman2001jpenobellecture and their significance for explaining the observed heterogeneity in return to education. Accordingly, it is important to allow the parameters of the wage equation to differ across individuals. In equation (ref) we allow ${\Greekmath 010B} _{i}$ and ${\Greekmath 010C} _{i}$ to differ across individuals, but assume that ${\Greekmath 011E} \left( \mathbf{z}_{i}\right) $ can be approximated as non-linear functions of experience and other control variables with homogeneous coefficients.

Specifically, following lemieux2006postnber,lemieux2006postsecondary we also allow for time variations in the parameters of the wage equation and consider the following categorical coefficient model over a given cross-section sample indexed by $t$:\footnote{ Some investigators have suggested including higher powers of the experience variable in the wage equation. lemieux2006mincerbookchapter, for example, proposes using a quartic rather than a quadratic function. As a robustness check we also provide estimation results with quartic experience specification in Section (ref) in the online supplement.}

equation[equation omitted — 555 chars of source]

where the return to education follows the categorical distribution,

equation*[equation* omitted — 282 chars of source]

and $\tilde{\mathbf{z}}_{it}$ includes gender, martial status and race. $ {\Greekmath 010B} _{it}={\Greekmath 010B} _{t}+{\Greekmath 010E} _{it}$ where ${\Greekmath 010E} _{it}$ is mean $0$ random variable assumed to be distributed independently of $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{edu}_{it}$ and $\mathbf{z}_{it}=\left( \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{exper}_{it},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{exper}_{it}^{2}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, } \tilde{\mathbf{z}}_{t}^{\prime }\right) ^{\prime }$. Let $u_{it}={\Greekmath 0122} _{it}+{\Greekmath 010E} _{it}$, and write (ref) as

equation[equation omitted — 530 chars of source]

The correlation between ${\Greekmath 010B} _{it}$ and $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{edu}_{it}$ in (ref) is the source of \textquotedblleft ability bias\textquotedblright\ griliches1977estimating. Given the pure cross-sectional nature of our analysis, we do not allow for the endogeneity from \textquotedblleft ability bias\textquotedblright\ or dynamics. To allow for non-zero correlations between ${\Greekmath 010B} _{it}$, edu$_{it}$ and $\mathbf{z} _{it}$, a panel data approach is required, which has its own challenges, as education and experience variables tend to very slow moving (if at all) for many individuals in the panel. Time delays between changes in education and experience, and the wage outcomes also further complicate the interpretation of the mean estimates of ${\Greekmath 010C} _{it}$ which we shall be reporting. To partially address the possible dynamic spillover effects, we provide estimates of the distribution of ${\Greekmath 010C} _{it}$ using cross-sectional data from two different sample periods, and investigate the extent to which the distribution of return to education has changed over time, by gender and the level of educational achievements.\footnote{ Time variations in return to education has also been investigated in the literature as a possible explanation of increasing wage inequality in the U.S. See, for example, the papers by lemieux2006postnber,lemieux2006postsecondary.}

We estimate the categorical distribution of the return to education in (ref) using the May and Outgoing Rotation Group (ORG) supplements of the Current Population Survey (CPS) data, as in lemieux2006postnber,lemieux2006postsecondary.\footnote{ The data is retrieved from \url{https://www.openicpsr.org/openicpsr/project/116216/version/V1/view}.} We pool observations from 1973 to 1975 for the first sample period, $ t=\left\{ 1973-1975\right\} $ and observations from 2001 to 2003 for the second sample period, $t=\left\{ 2001-2003\right\} $. Following lemieux2006postnber, we consider sub-samples of those with less than 12 years of education, \textquotedblleft high school or less", and those with more than 12 years of education, \textquotedblleft postsecondary education", as well as the combined sample. We also present results by gender. The summary statistics are reported in Table (ref). As to be expected, the mean log wages are higher for those with postsecondary education (for male and female), with the number of years of schooling and experience rising by about one year across the two sub-period samples. There are also important differences across male and female, and the two educational groupings, which we hope to capture in our estimation.

table[table omitted — 6,716 chars of source]
table[table omitted — 5,712 chars of source]
table[table omitted — 5,270 chars of source]

We treat the cross-section observations in the two sample periods, $ t=\left\{ 1973-1975\right\} $ and $\left\{ 2001-2003\right\} $, as repeated cross-sections, rather than a panel data since the data in these two periods do not cover the same individuals, and represent random samples from the population of wage earners in two periods. It should also be noted that sample sizes $(n_{t})$, although quite large, are much larger during $ \left\{ 2001-2003\right\} $, which could be a factor when we come to compare estimates from the two sample periods. For example, for both male and female $n_{73-75}=111,632$ as compared to $n_{01-03}=511,819$, a difference which becomes more pronounced when we consider the number observations in postsecondary/female category - which rises from $12,882$ for the first period to $100,007$ in the second period.

We report estimates of ${\Greekmath 0119}_{t}$, ${\Greekmath 010C} _{L,t}$ and ${\Greekmath 010C} _{H,t}$, as well as corresponding mean and standard deviations (denoted by s.d.($\hat{{\Greekmath 010C}} _{it}$)) of the return to education (${\Greekmath 010C}_{it}$) for $t=\left\{ 1973-1975\right\} $ and $\left\{ 2001-2003\right\} $. For a given ${\Greekmath 0119} _{t}$ , the ratio ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ provides a measure of within group heterogeneity and allows us to augment information on changes in mean with changes in the distribution of return of education. The estimates for the distribution of the return to education (${\Greekmath 010C} _{it}$) are summarized in Table (ref), with the estimation results for control variables (such as experience, experienced squared, and other individual specific characteristic) reported in Table (ref).

As can be seen from Table (ref), estimates of $ \mathrm{s.d.}\left( {\Greekmath 010C} _{it}\right) $ are strictly positive for all sub-groups, except for the \textquotedblleft high school or less\textquotedblright\ group during the first sample period. For this group during the first period the estimate of $\mathrm{s.d.}\left( {\Greekmath 010C} _{it}\right) $ for the male sub-sample is zero, ${\Greekmath 0119} $ is not identified, and we have identical estimates for ${\Greekmath 010C} _{L}$ and ${\Greekmath 010C} _{H}$. For this sub-sample, the associated estimates and their standard errors are shown as unavailable ($n/a$). In case of the female sub-sample as well as both male and female sub-samples where the estimates of s.d.($\hat{{\Greekmath 010C}}_{it}$) are close to zero and ${\Greekmath 0119} $ is poorly estimated, only the mean of the return to education is informative. In the case of the samples where the estimates of $ \mathrm{s.d.}\left( {\Greekmath 010C} _{it}\right) $ are strictly positive, the estimate of the ratio ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ provides a good measure of within group heterogeneity of return to education. The estimates of ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$, lie between $1.50$ to $2.79$, with the high estimate obtained for the females with high school or less education during $\left\{ { 2001-03}\right\} $, and the low estimate is obtained for females with postsecondary education during the same period.

As our theory suggests the mean estimates of return to education, $E\left( {\Greekmath 010C} _{it}\right) $, are very precisely estimated and inferences involving them tend to be robust to conditional error heteroskedasticity. The results in Table (ref) show that estimates of $E\left( {\Greekmath 010C} _{it}\right) $ have increased over the two sample periods $t=\left\{ 1973-75\right\} $ to $t=\left\{ 2001-03\right\} $, regardless of gender or educational grouping. The postsecondary educational group show larger increases in the estimates of $E\left( {\Greekmath 010C} _{it}\right) $ as compared to those with high school or less. Estimates of $\mathrm{E}\left( {\Greekmath 010C} _{it}\right) $ increases by $36$ per cent for the postsecondary group while the estimates of mean return to education rises only by around $5$ per cent in the case of those with high school or less. This result holds for both genders. Comparing the mean returns across the two educational groups, we find that mean return to education of individuals with postsecondary education is $45$ per cent higher than those with high school or less in the $\{1973 - 75\}$ period, but this gap increases to $87$ per cent in the second period, $\left\{ 2001-03\right\} $. Similar patterns are observed in the sub-samples by gender. The estimates suggest rising between group heterogeneity, which is mainly due to the increasing returns to education for the postsecondary group.

Turning to within group heterogeneity, we focus on the estimates of ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ and first note that over the two periods, within group heterogeneity has been rising mainly in the case of those with high school or less, for both male and female. For the combined male and female samples and the male sub-sample, there is little evidence of within group heterogeneity for the first period $\left\{ 1973-75\right\} $. However, for the second\ period $\left\{ 2001-03\right\} $ we find a sizeable degree of within group heterogeneity where ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ is estimated to be around $2.41$, with $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{s.d.}\left( {\Greekmath 010C} _{it}\right) \approx 0.03$. For the female sub-sample with high school or less, little evidence of heterogeneity was found for the first period, estimates of ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ increases to $2.79$ for the second sample period, that corresponds to a commensurate rise in $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{s.d.}\left( {\Greekmath 010C} _{i}\right) $ to $0.032$. The pattern of within group heterogeneity is very different for those with postsecondary educational. For this group we in fact observe a slight decline in the estimates of ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ by gender and over two sample periods.

Overall, our between and within estimates of mean return to education are in line with the evidence of rising wage inequality documented in the literature corak2013income.

Conclusion

In this paper we consider random coefficient models for repeated cross-sections in which the random coefficients follow categorical distributions. Identification is established using moments of the random coefficients in terms of the moments of the underlying observations. We propose two-step generalized method of moments to estimate the parameters of the categorical distributions. The consistency and asymptotic normality of the GMM estimators are established without the IID assumption typically assumed in the literature. Small sample properties of the proposed estimator are investigated by means of Monte Carlo experiments and shown to be robust to heterogeneously generated regressors and errors, although relatively large samples are required to estimate the parameters of the underling categorical distributions. This is largely due to the highly non-linear mapping between the parameters of the categorical distribution and the higher order moments of the coefficients. This problem is likely to become more pronounced with a larger number of categories and coefficients.

In the empirical application, we apply the model to study the evolution of returns to education over two sub-periods, also considered in the literature by lemieux2006postnber. Our estimates show that mean (ex post) returns to education have risen over the periods from 1973 - 75 to 2001 - 2003 mainly in the case of individuals with postsecondary education, and this result is robust by gender. We find evidence of within group heterogeneity in the case of high school or less educational group as compared to those with postsecondary education.

In our model specification, the number of categories, $K$, is treated as a tuning parameter and assumed to be known. An information criterion, as in bonhomme_manersa2015fixedeffect and \citet*{ssp2016classo}, to determine $K$ could be considered. Further investigation of models with multiple regressors subject to parameter heterogeneity is also required. These and other related issues are topics for future research.