EconBase
← Back to paper

Identification and Estimation of Categorical Random Coefficient Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

150,592 characters

Identification and Estimation of Categorical Random Coefficient Models



\title{Identification and Estimation of Categorical Random Coefficient Models\thanks{We would like to thank Timothy Armstrong, Hidehiko
Ichimura, Esfandiar Maasoumi, Geert Ridder, Ron Smith and Hayun Song for helpful
comments, and the editor and two anonymous referees for constructive comments and
suggestions.}
}
\author{ Zhan Gao\thanks{
Ph.D. student, Department of Economics, University of Southern California,
3620 South Vermont Avenue, Los Angeles, CA 90089, USA. Email: \texttt{
[email removed]}.} \and M. Hashem Pesaran\thanks{
1. Department of Economics, University of Southern California, 3620 South
Vermont Avenue, Los Angeles, CA 90089, USA. Email: \texttt{[email removed]}.
2. Trinity College, Cambridge, United Kingdom} }
\maketitle

\begin{abstract}
This paper proposes a linear categorical random coefficient model, in which
the random coefficients follow parametric categorical distributions. The
distributional parameters are identified based on a linear recurrence
structure of moments of the random coefficients. A Generalized Method of
Moments estimation procedure is proposed also employed by Peter Schmidt and
his coauthors to address heterogeneity in time effects in panel data models.
Using Monte Carlo simulations, we find that moments of the random
coefficients can be estimated reasonably accurately, but large samples are
required for estimation of the parameters of the underlying categorical
distribution. The utility of the proposed estimator is illustrated by
estimating the distribution of returns to education in the U.S. by gender
and educational levels. We find that rising heterogeneity between
educational groups is mainly due to the increasing returns to education for
those with postsecondary education, whereas within group heterogeneity has
been rising mostly in the case of individuals with high school or less
education.
\end{abstract}

\bigskip{}

\textbf{Keywords:} Random coefficient models, categorical distribution,
return to education

\textbf{JEL Code:} C01, C21, C13, C46, J30

\thispagestyle{empty}

\newpage \setcounter{page}{1}

\section{Introduction}

Random coefficient models have been used extensively in time series,
cross-section and panel regressions. \citet{nicholls198516} consider the
estimation of first and second moments of the random coefficient ${\Greekmath 010C} _{i}$
and the error term $u_{i}$, in a linear regression model. In the seminal
work, \citet{beran1992estimating} establish the conditions of identifying
and estimating the distribution of ${\Greekmath 010C} _{i}$ and $u_{i}$
non-parametrically. The baseline linear univariate regression in
\citet{beran1992estimating} has been extended in non-parametric framework by
\citet*{beran1993semiparametric, beran1994minimum, beran1996nonparametric,
hoderlein2010analyzing, hoderlein2017triangular} and
\citet{breunig2018specification}, to just name a few. \citet{hsiao2008random}
survey random coefficient models in linear panel data models.

In some econometric applications,
\citet{hausman1981exact,hausman1995nonparametric,foster_hahn2000consistent}
for examples, the main interest is to estimate the consumer surplus
distribution based on a linear demand system where the coefficient
associated with the price is random. In such settings, the distribution of
the random coefficients is needed when computing the consumer surplus
function, and the non-parametric estimation is more general, flexible and
suitable for the purpose. On the other hand, parametric models may be
favored in applications in which the implied economic meaning of the
distribution of the random coefficients is of interests. Examples include
estimation of the return to education
\citep{lemieux2006postnber,lemieux2006postsecondary} and the labor supply
equation \citep*{bick2020hours}.

In this paper, we consider a linear regression model with a random
coefficient ${\Greekmath 010C} _{i}$ that is assumed to follow a categorical
distribution, i.e. ${\Greekmath 010C} _{i}$ has a discrete support $\left\{
b_{1},b_{2},\cdots ,b_{K}\right\} $, and ${\Greekmath 010C} _{i}=b_{k}$ with probability
${\Greekmath 0119} _{k}$. The discretization of the support of the random coefficient $
{\Greekmath 010C} _{i}$ naturally corresponds to the interpretation that each individual
belongs to a certain category, or group, $k$ with probability ${\Greekmath 0119} _{k}$.
Compared to a non-parametric distribution with continuous support, assuming
a categorical distribution allows us not only to model the heterogeneous
responses across individuals but also to interpret the results with sharper
economic meaning. As we will illustrate in the empirical application in
Section \ref{sec:Empirical-Application}, it is hard to clearly interpret the
distribution of returns to education without imposing some form of
parametric restrictions.

In addition, with the categorical distribution imposed, the identification
and estimation of the distribution of ${\Greekmath 010C} _{i}$ do not rely on
identically distributed error terms $u_{i}$ and regressors $\mathbf{w}_i$,
as shown in Section \ref{sec:Identification} and \ref{sec:Estimation}.
Heterogeneously generated errors can be allowed, which is important in many
empirical applications. To the best of our knowledge, this is the first
identification result in linear random coefficient model without a strict
IID setting.

The identification of the distribution of ${\Greekmath 010C} _{i}$ is established in
this paper based on the identification of the moments of ${\Greekmath 010C} _{i}$, which
coincides with the identification condition in \citet{beran1992estimating}
that the distribution of ${\Greekmath 010C} _{i}$ is uniquely determined by its moments,
which is assumed to exist up to an arbitrary order. Since under our setup
the distribution of ${\Greekmath 010C} _{i}$ is parametrically specified, the moments of
${\Greekmath 010C} _{i}$ exist and can be derived explicitly. The parameters of the
assumed categorical distribution can then be uniquely determined by a system
of equations in terms of the moments, as in Theorem \ref
{prop:identification_of_beta_L_H}. The parameters of the categorical
distribution are then estimated consistently by the generalized method of
moments (GMM). The estimation procedure based on moment conditions shares
similar spirits as in \citet*{ahn2001gmm,ahn2013panel} in which Peter
Schmidt and coauthors study panel data models with interactive effects where
they allow for the time effects to vary across individual units. Comparing
to alternative non-parametric random coefficient models, the standard GMM
estimation is easy to implement, and the identified categorical structure
has a clear economic interpretation.

Using Monte Carlo (MC) simulations, we find that moments of the random
coefficients can be estimated reasonably accurately, but large samples are
required for estimation of the parameters of the underlying categorical
distributions. Our theoretical and MC results also suggest that our method
is suitable when the number heterogeneous coefficients and the number of
categories are small (2 or 3). With the number of categories rising the
burden on identification from the moments to the parameters of the
categorical distribution also rises rapidly. The quality of identification
also deteriorates as we need to rely on higher and higher moments to
identify a larger number of categories, since the information content of the
moments tend to decline with their order.

The proposed method is also illustrated by providing estimates of the
distribution of returns to education in the U.S. by gender and educational
levels, using the May and Outgoing Rotation Group (ORG) supplements of the
Current Population Survey (CPS) data. Comparing the estimates obtained over
the sub-periods 1973-75 and 2001-03, we find that rising between group
heterogeneity is largely due to rising returns to education in the case of
individuals with postsecondary education, whilst within group heterogeneity
has been rising in the case of individuals with high school or less
education.

\textbf{Related Literature:} This paper draws mainly upon the literature of
random coefficient models. As already mentioned, the main body of the recent
literature is focused on non-parametric identification and estimation.
Following \citet{beran1992estimating}, \citet{beran1993semiparametric} and
\citet{beran1994minimum} extend the model to a linear semi-parametric model
with a multivariate setup and propose a minimum distance estimator for the
unknown distribution. \citet{foster_hahn2000consistent} extend the
identification results in \citet{beran1992estimating} and apply the minimum
distance estimator to a gasoline consumption data to estimate the consumer
surplus function. \citet*{beran1996nonparametric} and \citet*{
hoderlein2010analyzing} propose kernel density estimators based on the Radon
inverse transformation in linear models.

In addition to linear models, \citet{ichimura1998maximum} and
\citet{gautier2013nonparametric} incorporate the random coefficients in
binary choice models. \citet{gautier2011triangular} and \citet*{
hoderlein2017triangular} consider triangular models with random coefficients
allowing for causal inference. \citet{matzkin2012identification} and
\citet{masten2018random} discuss the identification of random coefficients
in simultaneous equation models. \citet{breunig2018specification} propose a
general specification test in a variety of random coefficient models. Random
coefficients are also widely studied in panel data models, for example
\citet{hsiao2008random} and \citet{arellano2012identifying}

\medskip

The rest of the paper is organized as follows: Section \ref
{sec:Identification} establishes the main identification results. The GMM
estimation procedure is proposed and discussed in Section \ref
{sec:Estimation}. An extension to a multivariate setting is considered in
Section \ref{sec:Extensions}. Small sample properties of the proposed
estimator are investigated in Section \ref{sec:Monte-Carlo-Simulation},
using Monte Carlo techniques under different regressor and error
distributions. Section \ref{sec:Empirical-Application} presents and
discusses our empirical application to the return to education. Section \ref
{sec:Conclusion} provides some concluding remarks and suggestions for future
work. Technical proofs are given in Appendix \ref{sec:Proofs}.

\medskip

\noindent \textbf{Notations:} Largest and smallest eigenvalues of the $
p\times p$ matrix $\mathbf{A}=\left( a_{ij}\right) $ are denoted by ${\Greekmath 0115}
_{\max }\left( \mathbf{A}\right) $ and ${\Greekmath 0115} _{\min }\left( \mathbf{A}
\right) ,$ respectively, its spectral norm by $\left\Vert \mathbf{A}
\right\Vert ={\Greekmath 0115} _{\max }^{1/2}\left( \mathbf{A}^{\prime }\mathbf{A}
\right) $, $\mathbf{A}\succ 0$ means that $\mathbf{A}$ is positive definite,
$\mathrm{vech\left( \mathbf{A}\right) }$ denotes the vectorization of
distinct elements of $\mathbf{A}$, $\mathbf{0}$ denotes zero matrix (or
vector). For $\mathbf{a}\in \mathbb{R}^{p}$, $\mathrm{diag}\left( \mathbf{a}
\right) $ represents the diagonal matrix with elements of $
a_{1},a_{2},\cdots ,a_{p}$. For random variables (or vectors) $u$ and $v$, $
u\perp v$ represents $u$ is independent of $v$. We use $c$ ($C$) to denote
some small (large) positive constants. For a differentiable real-valued
function $f\left( \mathbf{{\Greekmath 0112} }\right) $, ${\Greekmath 0272} _{\mathbf{{\Greekmath 0112} }
}f\left( \mathbf{{\Greekmath 0112} }\right) $ denotes the gradient vector. Operator $
\rightarrow _{p}$ denotes convergence in probability, and $\rightarrow _{d}$
convergence in distribution. The symbols $O(1)$, and $O_{p}(1)$ denote
asymptotically bounded deterministic and random sequences, respectively.

\section{Categorical random coefficient model\label{sec:Identification}}

We suppose the single cross-section observations, $\left\{ y_{i},x_{i},
\mathbf{z}_{i}\right\} _{i=1}^{n}$, follow the categorical random
coefficient model
\begin{equation}
y_{i}=x_{i}{\Greekmath 010C} _{i}+\mathbf{z}_{i}^{\prime }\mathbf{{\Greekmath 010D} }+u_{i},
\label{eq:basic_model}
\end{equation}
where $y_{i},x_{i}\in \mathbb{R}$, $\mathbf{z}_{i}\in \mathbb{R}^{p_z},$ and
${\Greekmath 010C} _{i}\in \left\{ b_{1},b_{2},\cdots ,b_{K}\right\} $ admits the
following $K$-categorical distribution,
\begin{equation}
{\Greekmath 010C} _{i}=
\begin{cases}
b_{1}, & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{w.p. }{\Greekmath 0119} _{1}, \\
b_{2}, & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{w.p. }{\Greekmath 0119} _{2}, \\
\vdots & \vdots \\
b_{K}, & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{w.p. }{\Greekmath 0119} _{K},
\end{cases}
\label{eq:category_dist}
\end{equation}
w.p. denotes ``with probability'', ${\Greekmath 0119}_{k}\in \left( 0,1\right) $, $
\sum_{k=1}^{K}{\Greekmath 0119} _{k}=1$, $b_{1}<b_{2}<\cdots <b_{K}$, $\mathbf{{\Greekmath 010D} }\in
\mathbb{R}^{p_z}$ is homogeneous and $\mathbf{z}_{i}$ could include an
intercept term as its first element. It is assumed that ${\Greekmath 010C} _{i}\perp
\mathbf{w}_i = \left( x_{i},\mathbf{z}_{i}^{\prime }\right) ^{\prime }$, and
the idiosyncratic errors $u_{i}$ are independently distributed with mean $0$.

\begin{remark}
The model can be extended to allow $\mathbf{x}_{i},\mathbf{{\Greekmath 010C} }_{i}\in
\mathbb{R}^{p}$, with $\mathbf{{\Greekmath 010C} }_{i}$ following a multivariate
categorical distribution, though with more complicated notations. We will
consider possible extensions in Section \ref{sec:Extensions}.
\end{remark}

\begin{remark}
Since we consider a pure cross-sectional setting, the key assumption that $
{\Greekmath 010C} _{i}$ and $x_{i}$ are independently distributed cannot be relaxed.
Allowing ${\Greekmath 010C} _{i}$ to vary with $w_{i}$, without any further
restrictions, is tantamount to assuming $y_{i}$ is a general function of $
w_{i}$, in effect rendering a nonparametric specification.
\end{remark}

\begin{remark}
The number of categories $K$ is assumed to be fixed and known. Conditions $
\sum_{k=1}^{K}{\Greekmath 0119} _{k}=1$, $b_{1}<b_{2}<\cdots <b_{K},$ and ${\Greekmath 0119} _{k}\in
\left( 0,1\right) $ together are sufficient for the existence of $K$
categories. For example, if $b_{k}=b_{k^{\prime }}$, then we can merge
categories $k$ and $k^{\prime }$, and the number of categories reduces to $
K-1$. Similarly, if ${\Greekmath 0119} _{k}=0$ for some $k$, then category $k$ can be
deleted, and the number of categories is again reduced to $K-1$. Information
criteria can be used to determine $K$, but this will not be pursued in this
paper. Model specification tests could also be considered. See, for
examples, \citet{andrews2001testing} and \citet{breunig2018specification}.
\end{remark}

In the rest of this section, we focus on the model (\ref{eq:basic_model})
and establish the conditions under which the distribution of ${\Greekmath 010C} _{i}$ is
identified.

\subsection{Identifying the moments of $\protect{\Greekmath 010C} _{i}$\label
{subsec:identify_moments_beta}}

\begin{assumption}
\label{assu:identification_regularity_condition}

\begin{enumerate}
\item[(a)] (i) $u_{i}$ is distributed independently of $\mathbf{w}
_{i}=\left( x_{i},\mathbf{z}_{i}^{\prime }\right) ^{\prime }$ and ${\Greekmath 010C}
_{i} $. (ii) $\sup_{i}\mathrm{E}\left( \left\vert u_{i}^{r}\right\vert
\right) <C$, $r=1,2,\cdots ,2K-1$. (iii) $n^{-1}\sum_{i=1}^{n} u_i^4 =
O_p(1) $.

\item[(b)] (i) Let $\mathbf{Q}_{n,ww}=n^{-1}\sum_{i=1}^{n}\mathbf{w}_{i}
\mathbf{w}_{i}^{\prime }$, and $\mathbf{q}_{n,wy}=n^{-1}\sum_{i=1}^{n}
\mathbf{w}_{i}y_{i}$. Then $\left\Vert \mathrm{E}\left( \mathbf{Q}
_{n,ww}\right) \right\Vert <C<\infty $, and $\left\Vert \mathrm{E}\left(
\mathbf{q}_{n,wy}\right) \right\Vert <C<\infty $, and there exists $n_{0}\in
\mathbb{N}$ such that for all $n\geq n_{0}$,
\begin{equation*}
0<c<{\Greekmath 0115} _{\min }\left( \mathbf{Q}_{n,ww}\right) <{\Greekmath 0115} _{\max }\left(
\mathbf{Q}_{n,ww}\right) <C<\infty .
\end{equation*}
(ii) $\sup_{i} \mathrm{E}\left( \left\Vert \mathbf{w}_{i}\right\Vert
^{r}\right) <C<\infty $, $r=1,2,\cdots ,4K-2$.\newline
(iii) $n^{-1} \sum_{i=1}^{n} \left\Vert \mathbf{w}_{i} \right\Vert^{4} =
O_p(1)$.

\item[(c)] $\left\Vert \mathbf{Q}_{n,ww}-\mathrm{E} \left( \mathbf{Q}
_{n,ww}\right) \right\Vert =O_p\left( n^{-1/2}\right)$, $\left\Vert \mathbf{q
}_{n,wy}- \mathrm{E} \left( \mathbf{q}_{n,wy}\right) \right\Vert =O_p\left(
n^{-1/2}\right)$, and
\begin{equation*}
\mathrm{E} \left( \mathbf{Q}_{n,ww}\right) =n^{-1}\sum_{i=1}^{n}\mathrm{E}
\left( \mathbf{w}_{i}\mathbf{w}_{i}^{\prime }\right) \succ 0.
\end{equation*}

\item[(d)] $\left\Vert \mathrm{E} \left( \mathbf{Q}_{n,ww}\right) -\mathbf{Q}
_{ww}\right\Vert =O\left( n^{-1/2}\right) $, $\left\Vert \mathrm{E} \left(
\mathbf{q}_{n,wy}\right) -\mathbf{q}_{wy}\right\Vert =O\left(
n^{-1/2}\right) $, where $\mathbf{q}_{wy} = \lim\limits_{n \to \infty}
\mathrm{E} \left( \mathbf{q}_{n, wy} \right)$, $\mathbf{Q}_{ww} =
\lim\limits_{n \to \infty} \mathrm{E} \left( \mathbf{Q}_{n, ww} \right) $
and $\mathbf{Q}_{ww}\succ 0$.
\end{enumerate}
\end{assumption}

\begin{remark}
Part (a) of Assumption \ref{assu:identification_regularity_condition}
relaxes the assumption that $u_{i}$ is identically distributed, and allows
for heterogeneously generated errors. For identification of the distribution
of ${\Greekmath 010C} _{i}$, we require $u_{i}$ to be distributed independently of $
\mathbf{w}_{i}$ and ${\Greekmath 010C} _{i}$, which rules out conditional
heteroskedasticity. However, estimation and inference involving $\mathrm{E}
\left( {\Greekmath 010C} _{i}\right) $ and $\mathbf{{\Greekmath 010D} }$ can be carried out in
presence of conditionally error heteroskedastic, as shown in Theorem \ref
{lem:gamma_est_consistency}. Parts (c) and (d) of Assumption \ref
{assu:identification_regularity_condition} relax the condition that $\mathbf{
w}_{i}$ is identically distributed across $i$. As we proceed, only ${\Greekmath 010C}
_{i}$, whose distribution is of interest, is assumed to be IID across $i$,
and it is not required for $\mathbf{w}_{i}$ and $u_{i}$ to be identically
distributed over $i$.
\end{remark}

\begin{remark}
\label{rem:weak_crosssectional_depen_condition} The high level conditions in
Assumption \ref{assu:identification_regularity_condition}, concerning the
convergence in probability of averages such as $Q_{n,ww}=n^{-1}\sum_{i=1}^{n}
\mathbf{w}_{i}\mathbf{w}_{i}^{\prime }$, can be verified under weak
cross-sectional dependence. Let $f_{i}=f\left( \mathbf{w}_{i},{\Greekmath 010C}
_{i},u_{i}\right) $ be a generic function of $\mathbf{w}_{i}$, ${\Greekmath 010C} _{i}$
and $u_{i}$.\footnote{$f_{i}$ is assumed to be a scalar, and we can apply
the analysis element-by-element to a matrix, for example $\mathbf{w}_{i}
\mathbf{w}_{i}^{\prime }$.} Assume that $\sup_{i}\mathrm{E}\left(
f_{i}^{2}\right) <C$, and $\sup_{j}\sum_{i=1}^{n}\left\vert \mathrm{cov}
\left( f_{i},f_{j}\right) \right\vert <C$, for some fixed $C<\infty $. Then,
\begin{equation*}
\mathrm{var}\left( \frac{1}{n}\sum_{i=1}^{n}f_{i}\right) \leq \frac{1}{n^{2}}
\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert \mathrm{cov}\left( f_{i},f_{j}\right)
\right\vert \leq \frac{1}{n}\sup_{j}\sum_{i=1}^{n}\left\vert \mathrm{cov}
\left( f_{i},f_{j}\right) \right\vert \leq \frac{C}{n}.
\end{equation*}
By Chebyshev's inequality, for any ${\Greekmath 0122} >0$, we have $M_{{\Greekmath 0122}
}>\sqrt{C/{\Greekmath 0122} }$ such that
\begin{equation*}
\Pr \left( \sqrt{n}\left\vert \frac{1}{n}\sum_{i=1}^{n}\left[ f_{i}-\mathrm{E
}\left( f_{i}\right) \right] \right\vert >M_{{\Greekmath 0122} }\right) \leq \frac{
n\mathrm{var}\left( n^{-1}\sum_{i=1}^{n}f_{i}\right) }{C}{\Greekmath 0122} \leq
{\Greekmath 0122} ,
\end{equation*}
i.e. $n^{-1}\sum_{i=1}^{n}\left[ f_{i}-\mathrm{E}\left( f_{i}\right) \right]
=O_{p}\left( n^{-1/2}\right) $.
\end{remark}

Denote $\mathbf{{\Greekmath 011E} }_{i}=\left( {\Greekmath 010C} _{i},\mathbf{{\Greekmath 010D} }^{\prime
}\right) ^{\prime }$ and $\mathbf{{\Greekmath 011E} }=\mathrm{E}\left( \mathbf{{\Greekmath 011E} }
_{i}\right) =\left( \mathrm{E}\left( {\Greekmath 010C} _{i}\right) ,\mathbf{{\Greekmath 010D} }
^{\prime }\right) ^{\prime }$. Consider the moment condition,
\begin{equation}
\mathrm{E}\left( \mathbf{w}_{i}y_{i}\right) =\mathrm{E}\left( \mathbf{w}_{i}
\mathbf{w}_{i}^{\prime }\right) \mathbf{{\Greekmath 011E} },  \label{mc_y1_w1}
\end{equation}
and sum \eqref{mc_y1_w1} over $i$
\begin{equation}
\frac{1}{n}\sum_{i=1}^{n}\mathrm{E}\left( \mathbf{w}_{i}y_{i}\right) =\left[
\frac{1}{n}\sum_{i=1}^{n}\mathrm{E}\left( \mathbf{w}_{i}\mathbf{w}
_{i}^{\prime }\right) \right] \mathbf{{\Greekmath 011E} }.  \label{mc_y1w1_n}
\end{equation}
Let $n\to \infty$, then $\mathbf{{\Greekmath 011E} }$ is identified by
\begin{equation}
\mathbf{{\Greekmath 011E} } = \mathbf{Q}_{ww} ^{-1} \mathbf{q}_{wy},  \label{phi}
\end{equation}
under Assumption \ref{assu:identification_regularity_condition}.

\begin{assumption}
\label{assu:conv_moment} Let $\tilde{y}_{i}=y_{i}-\mathbf{z}_{i}^{\prime }
\mathbf{{\Greekmath 010D} }$.

\begin{enumerate}
\item[(a)] {$\left\vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( \tilde{y}
_{i}^{r}x_{i}^{s}\right) -{\Greekmath 011A} _{r,s}\right\vert =O\left( n^{-1/2}\right) ,$
and $\left\vert {\Greekmath 011A} _{r,s}\right\vert <\infty ,$ for $r,s=0,1,\cdots ,2K-1$
. }

\item[(b)] {$\left\vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left(
u_{i}^{r}\right) -{\Greekmath 011B} _{r}\right\vert =O\left( n^{-1/2}\right) ,$ and $
\left\vert {\Greekmath 011B} _{r}\right\vert <\infty $, for $r=2,3,\cdots ,2K-1$. }

\item[(c)] $n^{-1}\sum_{i=1}^{n}\left[ \mathrm{var}(x_{i}^{r})-\left( {\Greekmath 011A}
_{0,2r}-{\Greekmath 011A} _{0,r}^{2}\right) \right] =O\left( n^{-1/2}\right) $ where $
{\Greekmath 011A} _{0,2r}-{\Greekmath 011A} _{0,r}^{2}>0,$ for $r=2,3,\cdots ,2K-1$.
\end{enumerate}
\end{assumption}

\begin{remark}
\label{heterogeniety of moments} The above assumption allows for a limited
degree of heterogeneity of the moments. As an example, let $\mathrm{E}\left(
u_{i}^{r}\right) ={\Greekmath 011B} _{ir}$ and denote the heterogeneity of the $r^{th}$
moment of $u_{i}$ by $e_{ir}={\Greekmath 011B} _{ir}-{\Greekmath 011B} _{r}$. Then
\begin{equation*}
\left\vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( u_{i}^{r}\right) -{\Greekmath 011B}
_{r}\right\vert \leq n^{-1}\sum_{i=1}^{n}\left\vert e_{ir}\right\vert ,
\end{equation*}
and condition (b)\ of Assumption \ref{assu:conv_moment} is met if $\
\sum_{i=1}^{n}\left\vert e_{ir}\right\vert =O(n^{{\Greekmath 010B} _{r}})$ with ${\Greekmath 010B}
_{r}<1/2$. ${\Greekmath 010B} _{r}$ measures the degree of heterogeneity with ${\Greekmath 010B}
_{r}=1$ representing the highest degree of heterogeneity. A similar idea is
used by \citet{pesaran2018pool} in their analysis of poolability in panel
data models.
\end{remark}

\begin{theorem}
\label{lem: identification_moments} Under Assumptions \ref
{assu:identification_regularity_condition} and \ref{assu:conv_moment}, $
\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) $ and ${\Greekmath 011B} _{r}$, $r=2,3,\cdots
,2K-1$ are identified.
\end{theorem}

\begin{proof}
For $r=2,\cdots ,2K-1$,
\begin{align}
\mathrm{E}\left( \tilde{y}_{i}^{r}\right) & =\mathrm{E}\left(
x_{i}^{r}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) +\mathrm{E}\left(
u_{i}^{r}\right) +\sum_{q=2}^{r-1}\binom{r}{q}\mathrm{E}\left(
x_{i}^{r-q}\right) \mathrm{E}\left( u_{i}^{q}\right) \mathrm{E}\left( {\Greekmath 010C}
_{i}^{r-q}\right) ,  \label{eq:mc_n_r} \\
\mathrm{E}\left( \tilde{y}_{i}^{r}x_{i}^{r}\right) & =\mathrm{E}\left(
x_{i}^{2r}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) +\mathrm{E}\left(
x_{i}^{r}\right) \mathrm{E}\left( u_{i}^{r}\right) +\sum_{q=2}^{r-1}\binom{r
}{q}\mathrm{E}\left( x_{i}^{2r-q}\right) \mathrm{E}\left( u_{i}^{q}\right)
\mathrm{E}\left( {\Greekmath 010C} _{i}^{r-q}\right) .  \label{eq:mc_n_2r}
\end{align}
where $\binom{r}{q}=\frac{r!}{q!(r-q)!}$ are binomial coefficients, for
non-negative integers $q\leq r$.

Sum over $i$, then by parts (a) and (b) of Assumption \ref{assu:conv_moment}
,
\begin{align}
{\Greekmath 011A} _{0,r}\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) +{\Greekmath 011B} _{r}& ={\Greekmath 011A}
_{r,0}-\sum_{q=2}^{r-1}\binom{r}{q}{\Greekmath 011A} _{0,r-q}{\Greekmath 011B} _{q}\mathrm{E}\left(
{\Greekmath 010C} _{i}^{r-q}\right) ,  \label{eq:mc_limit_r} \\
{\Greekmath 011A} _{0,2r}\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) +{\Greekmath 011A} _{0,r}{\Greekmath 011B} _{r}&
={\Greekmath 011A} _{r,r}-\sum_{q=2}^{r-1}\binom{r}{q}{\Greekmath 011A} _{0,2r-q}{\Greekmath 011B} _{q}\mathrm{E}
\left( {\Greekmath 010C} _{i}^{r-q}\right) .  \label{eq:mc_limit_2r}
\end{align}
Derivation details are relegated to Appendix \ref{sec:Proofs}. By part (c)
of Assumption \ref{assu:conv_moment}, the matrix $
\begin{pmatrix}
{\Greekmath 011A} _{0,r} & 1 \\
{\Greekmath 011A} _{0,2r} & {\Greekmath 011A} _{0,r}
\end{pmatrix}
$ is invertible for $r=2,3,\cdots ,2K-1$. As a result, we can sequentially
solve \eqref{eq:mc_limit_r} and \eqref{eq:mc_limit_2r} for $\mathrm{E}\left(
{\Greekmath 010C} _{i}^{r}\right) $ and ${\Greekmath 011B} _{r}$, for $r=2,3,\cdots ,2K-1$.
\end{proof}

\subsection{Identifying the distribution of $\protect{\Greekmath 010C} _{i}$}

\citet[Theorem 2.1, pp. 1972]{beran1992estimating} prove the identification
of the distribution of the random coefficient, ${\Greekmath 010C} _{i}$, in a canonical
model without covariates, $z_{i}$, under the condition that the distribution
of ${\Greekmath 010C} _{i}$ is uniquely determined by its moments. We show the
identification of moments of ${\Greekmath 010C}_i$ holds more generally when $x_i$ and $
u_i$ are not identically distributed and the distribution of ${\Greekmath 010C}_i$ is
identified if it follows a categorical distribution. Note that under (\ref
{eq:category_dist}),
\begin{equation}
\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) =\sum_{k=1}^{K}{\Greekmath 0119}
_{k}b_{k}^{r},\;r=0,1,2,\cdots ,2K-1,  \label{mbeta}
\end{equation}
with $\mathrm{E}\left( {\Greekmath 010C} _{i}^{r}\right) $ identified under Assumption
\ref{assu:identification_regularity_condition}. To identify $\mathbf{{\Greekmath 0119} }
=\left( {\Greekmath 0119} _{1},{\Greekmath 0119} _{2},...,{\Greekmath 0119} _{K}\right) ^{\prime }$ and $\mathbf{b}
=\left( b_{1},b_{2},...,b_{K}\right) ^{\prime }$, we need to verify that the
system of $2K$ equations in \eqref{mbeta} has a unique solution if $
b_{1}<b_{2}<\cdots <b_{K}$, and ${\Greekmath 0119} _{k}\in \left( 0,1\right) $. In the
proof, we construct a linear recurrence relation and make use of the
corresponding characteristic polynomial.

\begin{theorem}
\label{prop:identification_of_beta_L_H} Consider the random coefficient
regression model \eqref{eq:basic_model}, suppose that Assumptions \ref
{assu:identification_regularity_condition} and \ref{assu:conv_moment} hold.
Then $\mathbf{{\Greekmath 0112} }=\left( \mathbf{{\Greekmath 0119} }^{\prime },\mathbf{b}^{\prime
}\right) ^{\prime }$ is identified subject to $b_{1}<b_{2}<\cdots <b_{K}$
and ${\Greekmath 0119} _{k}\in \left( 0,1\right) $, for all $k=1,2,\cdots ,K$.
\end{theorem}

\begin{proof}
We motivate the key idea of the proof in the special case where $K=2,$ and
relegate the proof of the general case to the Appendix \ref{sec:Proofs}. Let
$b_{1}={\Greekmath 010C} _{L}$, $b_{2}={\Greekmath 010C} _{H}$, ${\Greekmath 0119} _{1}={\Greekmath 0119} $ and ${\Greekmath 0119} _{2}=1-{\Greekmath 0119} $
. Note that
\begin{align}
\mathrm{E}\left( {\Greekmath 010C} _{i}\right) & ={\Greekmath 0119} {\Greekmath 010C} _{L}+\left( 1-{\Greekmath 0119} \right)
{\Greekmath 010C} _{H},  \label{id1} \\
\mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) & ={\Greekmath 0119} {\Greekmath 010C} _{L}^{2}+\left( 1-{\Greekmath 0119}
\right) {\Greekmath 010C} _{H}^{2},  \label{id2} \\
\mathrm{E}\left( {\Greekmath 010C} _{i}^{3}\right) & ={\Greekmath 0119} {\Greekmath 010C} _{L}^{3}+\left( 1-{\Greekmath 0119}
\right) {\Greekmath 010C} _{H}^{3},  \label{id3}
\end{align}
and $\mathrm{E}\left( {\Greekmath 010C} _{i}^{k}\right) $, $k=1,2,3$ are identified. $
\left( {\Greekmath 0119} ,{\Greekmath 010C} _{L},{\Greekmath 010C} _{H}\right) $ can be identified if the system
of equations \eqref{id1} to \eqref{id3}, has a unique solution. By
\eqref{id1},
\begin{equation}
{\Greekmath 0119} =\frac{{\Greekmath 010C} _{H}-\mathrm{E}\left( {\Greekmath 010C} _{i}\right) }{{\Greekmath 010C} _{H}-{\Greekmath 010C}
_{L}},\; \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{and}\; 1-{\Greekmath 0119} =\frac{\mathrm{E}\left( {\Greekmath 010C} _{i}\right) -{\Greekmath 010C}
_{L}}{{\Greekmath 010C} _{H}-{\Greekmath 010C} _{L}}.  \label{pi_expression}
\end{equation}
Plug \eqref{pi_expression} into \eqref{id2} and \eqref{id3},
\begin{align}
\mathrm{E}\left( {\Greekmath 010C} _{i}\right) \left( {\Greekmath 010C} _{L}+{\Greekmath 010C} _{H}\right)
-{\Greekmath 010C} _{L}{\Greekmath 010C} _{H}& =\mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) ,
\label{sys_1} \\
\mathrm{E}\left( {\Greekmath 010C} _{i}^2\right) \left( {\Greekmath 010C}_L + {\Greekmath 010C}_H\right) -
\mathrm{E}\left({\Greekmath 010C}_i\right) {\Greekmath 010C} _{L}{\Greekmath 010C} _{H} & =\mathrm{E}\left(
{\Greekmath 010C} _{i}^{3}\right) .  \label{sys_2}
\end{align}
Denote ${\Greekmath 010C}_{L+H} = {\Greekmath 010C}_{L} + {\Greekmath 010C}_{H}$ and ${\Greekmath 010C}_{LH} =
{\Greekmath 010C}_{L}{\Greekmath 010C}_{H}$, and write \eqref{sys_1} and \eqref{sys_2} in matrix
form,
\begin{equation}
\mathbf{M} \mathbf{D} \mathbf{b}^{\ast} = \mathbf{m},
\label{eq:sys_mat_form}
\end{equation}
where
\begin{equation*}
\mathbf{M} =
\begin{pmatrix}
1 & \mathrm{E}\left({\Greekmath 010C}_i\right) \\
\mathrm{E}\left({\Greekmath 010C}_i\right) & \mathrm{E}\left({\Greekmath 010C}_i^2 \right)
\end{pmatrix}
, \, \mathbf{D} =
\begin{pmatrix}
-1 & 0 \\
0 & 1
\end{pmatrix}
,\, \mathbf{b}^{\ast} =
\begin{pmatrix}
{\Greekmath 010C}_{LH} \\
{\Greekmath 010C}_{L+H}
\end{pmatrix}
, \;\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{and}\; \mathbf{m} =
\begin{pmatrix}
\mathrm{E}\left({\Greekmath 010C}_i^2\right) \\
\mathrm{E}\left({\Greekmath 010C}_i^3\right)
\end{pmatrix}
.
\end{equation*}
Under the conditions $0 < {\Greekmath 0119} < 1$ and ${\Greekmath 010C}_H > {\Greekmath 010C}_L$,
\begin{equation*}
\det\left( \mathbf{M} \right) = \mathrm{var}\left( {\Greekmath 010C} _{i}\right) =
\mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) -\mathrm{E}\left( {\Greekmath 010C} _{i}\right)
^{2} = {\Greekmath 0119}\left( 1-{\Greekmath 0119} \right)\left( {\Greekmath 010C}_H - {\Greekmath 010C}_L \right)^2 > 0.
\end{equation*}
As a result, we can solve \eqref{eq:sys_mat_form} for ${\Greekmath 010C}_{L+H}$ and $
{\Greekmath 010C}_{LH}$ as
\begin{align}
{\Greekmath 010C} _{L+H}& =\frac{\mathrm{E}\left( {\Greekmath 010C} _{i}^{3}\right) -\mathrm{E}
\left( {\Greekmath 010C} _{i}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) }{\mathrm{var
}\left( {\Greekmath 010C}_i \right)},  \label{sys_3} \\
{\Greekmath 010C} _{LH}& =\frac{\mathrm{E}\left( {\Greekmath 010C} _{i}\right) \mathrm{E}\left(
{\Greekmath 010C} _{i}^{3}\right) -\mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) ^{2}}{\mathrm{
var}\left( {\Greekmath 010C}_i \right)}.  \label{sys_4}
\end{align}
${\Greekmath 010C} _{L}$ and ${\Greekmath 010C} _{H}$ are solutions to the quadratic equation,
\begin{equation}
{\Greekmath 010C} ^{2}-{\Greekmath 010C} _{L+H}{\Greekmath 010C} +{\Greekmath 010C} _{LH}=0.
\end{equation}
We can verify that $\Delta ={\Greekmath 010C} _{L+H}^{2}-4{\Greekmath 010C} _{LH}>0$ by direct
calculation using \eqref{sys_3} and \eqref{sys_4}. Simplifying $\Delta $ in
terms of $\mathrm{E}\left( {\Greekmath 010C} _{i}^{k}\right) $ and then plugging in
\eqref{id1}, \eqref{id2} and \eqref{id3},
\begin{align*}
\Delta =& \frac{ \left[ \mathrm{E}\left( {\Greekmath 010C} _{i}^{3}\right) - \mathrm{E}
\left( {\Greekmath 010C} _{i}\right) \mathrm{E}\left( {\Greekmath 010C} _{i}^{2}\right) \right] ^{2}
-4 \mathrm{var}\left( {\Greekmath 010C}_i \right) \left[ \mathrm{E}\left( {\Greekmath 010C}
_{i}\right) \mathrm{E}\left({\Greekmath 010C} _{i}^{3}\right) - \mathrm{E}\left(
{\Greekmath 010C}_{i}^{2}\right) ^{2} \right] }{ \left[ \mathrm{var}\left( {\Greekmath 010C}_i
\right) \right] ^{2} } \\
& = \left( {\Greekmath 010C}_H - {\Greekmath 010C}_L \right)^2 > 0.
\end{align*}
Then, we obtain the unique solutions,
\begin{align}
{\Greekmath 010C} _{L}& =\frac{1}{2}\left( {\Greekmath 010C} _{L+H}-\sqrt{{\Greekmath 010C} _{L+H}^{2}-4{\Greekmath 010C}
_{LH}}\right) ,  \label{solution_beta_L} \\
{\Greekmath 010C} _{H}& =\frac{1}{2}\left( {\Greekmath 010C} _{L+H}+\sqrt{{\Greekmath 010C} _{L+H}^{2}-4{\Greekmath 010C}
_{LH}}\right) ,  \label{solution_beta_H}
\end{align}
and ${\Greekmath 0119} $ can be determined by (\ref{pi_expression}) correspondingly.
\end{proof}

\begin{remark}
The key identifying assumption in (\ref{prop:identification_of_beta_L_H}) is
the assumed existence of the strict ordinal relation $b_{1}<b_{2}<\cdots
<b_{K}$ so that $b_{k}$ and $b_{k^{\prime }}$ are not symmetric for $k\neq
k^{\prime }$, and $0<{\Greekmath 0119} _{k}<1$ so that the distribution of ${\Greekmath 010C} _{i}$
does not degenerate. When $K=2$, the conditions $b_{1}<b_{2}<\cdots <b_{K}$,
and ${\Greekmath 0119} _{k}\in \left( 0,1\right) $, are equivalent to $\mathrm{var}\left(
{\Greekmath 010C} _{i}\right) ={\Greekmath 0119} _{1}\left( 1-{\Greekmath 0119} _{1}\right) \left(
b_{2}-b_{1}\right) ^{2}>0$. In other words, not surprisingly, the
categorical distribution of ${\Greekmath 010C} _{i}$ are identified only if $\mathrm{var}
\left( {\Greekmath 010C} _{i}\right) >0$.

In practice, a test for $\mathbb{H}_{0}:\mathrm{var}\left( {\Greekmath 010C} _{i}\right)
=0$ is possible, by noting that $\mathrm{var}\left( {\Greekmath 010C} _{i}\right) =0$ is
equivalent to
\begin{equation*}
{\Greekmath 0114} ^{2}=\frac{\mathrm{E}\left( {\Greekmath 010C} _{i}\right) ^{2}}{\mathrm{E}\left(
{\Greekmath 010C} _{i}^{2}\right) }=1,
\end{equation*}
where ${\Greekmath 0114} ^{2}$ is well-defined as long as ${\Greekmath 010C} _{i}\not\equiv 0$. One
important advantage of basing the test of slope homogeneity on ${\Greekmath 0114} ^{2}$
rather than on $var({\Greekmath 010C} _{i})=0$, is that ${\Greekmath 0114} ^{2}\,$is
scale-invariant. $\mathrm{E}\left( {\Greekmath 010C} _{i}\right) $ and $\mathrm{E}\left(
{\Greekmath 010C} _{i}^{2}\right) $ are identified as in Section \ref
{subsec:identify_moments_beta}, whose consistent estimation does not require
$\mathrm{var}\left( {\Greekmath 010C} _{i}\right) >0$. Consequently, in principle it is
possible to test slope homogeneity by testing $\mathbb{H}_{0}:{\Greekmath 0114} ^{2}=1$
. However, the problem becomes much more complicated when there are more
than two categories and/or there are more than one regressor under
consideration. A full treatment of testing slope homogeneity in such general
settings is beyond the scope of the present paper.
\end{remark}


\begin{remark}
Note that in the special case of the proof of Theorem \ref
{prop:identification_of_beta_L_H} where $K=2$, ${\Greekmath 010C}_{L+H}={\Greekmath 010C}_{L}+
{\Greekmath 010C}_{H}$ and ${\Greekmath 010C}_{LH}={\Greekmath 010C}_{L}{\Greekmath 010C}_{H}$ corresponds to the $
b_{1}^{\ast}$ and $b_{2}^{\ast}$ and \eqref{eq:sys_mat_form} is the same as
\eqref{eq: linear_recurrence_matrix} when $K = 2$. The special case
illustrates the procedure of identification: identify $\left(b_{k}^{\ast}
\right)_{k=1}^{K}$ by the moments of ${\Greekmath 010C}_{i}$, then solve for $
\left(b_{k}\right)_{k=1}^{K}$ and finally identify $\left({\Greekmath 0119}_{k}
\right)_{k=1}^{K}$.
\end{remark}


\section{Estimation\label{sec:Estimation}}

In this section, we propose a generalized method of moments estimator for
the distributional parameters of ${\Greekmath 010C}_i$. To reduce the complexity of the
moment equations, we first obtain a $\sqrt{n}$-consistent estimator of $
{\Greekmath 010D}$ and consider the estimation of the distribution of ${\Greekmath 010C}_i$ by
replacing ${\Greekmath 010D}$ by $\hat{{\Greekmath 010D}}$.

\subsection{Estimation of $\protect{\Greekmath 010D}$\label{subsec:Estimation-of-gamma}}

Let $\mathbf{{\Greekmath 011E} }=\left( \mathrm{E}\left( {\Greekmath 010C} _{i}\right) ,\mathbf{
{\Greekmath 010D} }^{\prime }\right) ^{\prime }$, $v_{i}={\Greekmath 010C} _{i}-\mathrm{E}\left(
{\Greekmath 010C} _{i}\right) $ and using the notation in Assumption \ref
{assu:identification_regularity_condition}, \eqref{eq:basic_model} can be
written as
\begin{equation}
y_{i}=\mathbf{w}_{i}^{\prime }\mathbf{{\Greekmath 011E} }+{\Greekmath 0118} _{i},
\label{eq: model_w_phi}
\end{equation}
where ${\Greekmath 0118} _{i}=u_{i}+x_{i}v_{i}$. Then $\mathbf{{\Greekmath 011E} }$ can be estimated
consistently by $\hat{\mathbf{{\Greekmath 011E} }}=\mathbf{Q}_{n,ww}^{-1}\mathbf{q}
_{n,wy} $ where $\mathbf{Q}_{n,ww}$ and $\mathbf{q}_{n,wy}$ are defined in
Assumption \ref{assu:identification_regularity_condition}.

\begin{assumption}
\label{assu:gamma_est_normality} $\left\Vert n^{-1}\sum_{i=1}^{n}\mathrm{E}
\left( \mathbf{w}_i\mathbf{w}_i^\prime {\Greekmath 0118}_i^2\right) - \mathbf{V}_{w{\Greekmath 0118}}
\right\Vert = O\left( n^{-1/2} \right) $, $\mathbf{V}_{w{\Greekmath 0118}}\succ 0, $ and
\begin{equation}
\left\Vert \frac{1}{n}\sum_{i=1}^{n} \mathbf{w}_i\mathbf{w}_i^\prime {\Greekmath 0118}_i^2
- \frac{1}{n}\sum_{i=1}^{n}\mathrm{E}\left( \mathbf{w}_i\mathbf{w}_i^\prime
{\Greekmath 0118}_i^2\right) \right\Vert = O_p\left( n^{-1/2} \right).
\label{eq:conv_of_wwxi}
\end{equation}
\end{assumption}

\begin{remark}
As in the case of Assumption \ref{assu:identification_regularity_condition},
the high level condition \eqref{eq:conv_of_wwxi} can be shown to hold under
weak cross-sectional dependence, assuming that elements of $\mathbf{w}_{i}
\mathbf{w}_{i}^{\prime }{\Greekmath 0118} _{i}^{2}$ are cross-sectionally weakly
correlated over $i$. See Remark \ref{rem:weak_crosssectional_depen_condition}
.
\end{remark}

\begin{theorem}
\label{lem:gamma_est_consistency} Under Assumption \ref
{assu:identification_regularity_condition}, $\hat{\mathbf{{\Greekmath 011E}}}$ is a
consistent estimator for $\mathbf{{\Greekmath 011E}}$. In addition, under Assumptions \ref
{assu:identification_regularity_condition} and \ref{assu:gamma_est_normality}
, as $n\to \infty$,
\begin{equation}
\sqrt{n}\left( \hat{\mathbf{{\Greekmath 011E}}} - \mathbf{{\Greekmath 011E}} \right) \rightarrow_{d}
N\left( \mathbf{0}, \mathbf{V}_{\Greekmath 011E} \right),
\end{equation}
where $\mathbf{V}_{{\Greekmath 011E}} = \mathbf{Q}_{w w}^{-1} \mathbf{V}_{w{\Greekmath 0118}} \mathbf{Q}
_{w w}^{-1}. $ $\mathbf{V}_{{\Greekmath 011E}}$ is consistently estimated by
\begin{equation*}
\hat{\mathbf{V}}_{{\Greekmath 011E}} = \mathbf{Q}_{n,ww}^{-1}\hat{\mathbf{V}}_{w{\Greekmath 0118}}
\mathbf{Q}_{n,ww}^{-1}\to_p \mathbf{V}_{{\Greekmath 011E}},
\end{equation*}
as $n\to\infty$, where $\hat{\mathbf{V}}_{w{\Greekmath 0118}} = n^{-1}\sum_{i=1}^{n}
\mathbf{w}_i\mathbf{w}_i^\prime \hat{{\Greekmath 0118}}_{i}^2$, and $\hat{{\Greekmath 0118}}_i = y_i -
\mathbf{w}_i^\prime \hat{\mathbf{{\Greekmath 011E}}}$.
\end{theorem}

The proof of Theorem \ref{lem:gamma_est_consistency} is provided in Section
\ref{suppsec:Proofs} in the online supplement.

\subsection{Estimation of the distribution of $\protect{\Greekmath 010C}_i$\label
{subsec:estimation_beta}}

Denote the moments of ${\Greekmath 010C} _{i}$ on the right-hand side of (\ref{mbeta})
by
\begin{equation*}
\mathbf{m}_{{\Greekmath 010C} }=(m_{1},m_{2},...,m_{2K-1})^{\prime }=\left[ \mathrm{E}
\left( {\Greekmath 010C} _{i}^{r}\right) \right] _{r=1}^{2K-1}\in \Theta _{m}\subset
\left\{ \mathbf{m}_{{\Greekmath 010C} }\in \mathbb{R}^{2K-1}:m_{r}\geq 0,\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }r\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
is even}\right\} ,
\end{equation*}
and note that
\begin{equation}
\mathbf{m}_{{\Greekmath 010C} }=\left(
\begin{array}{c}
m_{1} \\
m_{2} \\
\vdots \\
m_{2K-1}
\end{array}
\right) =\left(
\begin{array}{cccc}
b_{1} & b_{2} & \cdots & b_{K} \\
b_{1}^{2} & b_{2}^{2} & \cdots & b_{K}^{2} \\
\vdots & \vdots & \vdots & \vdots \\
b_{1}^{2K-1} & b_{2}^{2K-1} & \cdots & b_{K}^{2K-1}
\end{array}
\right) \left(
\begin{array}{c}
{\Greekmath 0119} _{1} \\
{\Greekmath 0119} _{2} \\
\vdots \\
{\Greekmath 0119} _{K}
\end{array}
\right) ,  \label{eq:mbeta_mat}
\end{equation}
so in general we can write $\mathbf{m}_{{\Greekmath 010C} }\triangleq h\left( \mathbf{
{\Greekmath 0112} }\right) ,$ where $\mathbf{{\Greekmath 0112} }=\left( \mathbf{{\Greekmath 0119} }^{\prime },
\mathbf{b}^{\prime }\right) ^{\prime }\in \Theta $, and $\mathbf{{\Greekmath 0112} }$
can be uniquely determined in terms of $\mathbf{m}_{{\Greekmath 010C} }$ by Theorem \ref
{prop:identification_of_beta_L_H}. To estimate $\mathbf{{\Greekmath 0112} }$, we
consider moment conditions following a similar procedure as in Section \ref
{sec:Identification}, and propose a generalized method of moments (GMM)
estimator.

We consider the following moment conditions
\begin{equation*}
\mathrm{E}\left( \tilde{y}_{i}^{r}\right) =\sum_{q=0}^{r}\binom{r}{q}\mathrm{
E}\left( x_{i}^{r-q}\right) \mathrm{E}\left( u_{i}^{q}\right) m_{r-q},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
}
\end{equation*}
and
\begin{equation}
\mathrm{E}\left( \tilde{y}_{i}^{r}x_{i}^{s_{r}}\right) =\sum_{q=0}^{r}\binom{
r}{q}\mathrm{E}\left( x_{i}^{r-q+s_{r}}\right) \mathrm{E}\left(
u_{i}^{q}\right) m_{r-q},  \label{eq:mc_yrxs}
\end{equation}
where $\mathrm{E}\left( u_{i}\right) =0$, $\tilde{y}_{i}=y_{i}-\mathbf{z}
_{i}^{\prime }\mathbf{{\Greekmath 010D} }$, $r=1,2,...,2K-1$, and $s_{r}=0,1,\cdots
,S-r $, where $S$ is a user-specific tuning parameter, chosen such that the
highest order moments of $x_{i}$ included is at most $S$, where $S>2K-1$.
\footnote{
For identification, we require the moments of $x_{i}$ to exist up to order $
4K-2$. $S$ can take values between $2K$ to $4K-2$. In practice, the choice
of $S$ affects the trade-off between bias and efficiency.}

Let ${\Greekmath 011B} _{0}=1$ and ${\Greekmath 011B} _{1}=0$ such that ${\Greekmath 011B} _{r}$ is
well-defined for $r=0,1,\cdots ,2K-1$.\textbf{\ }Sum \eqref{eq:mc_yrxs} over
$i$ and rearrange terms,
\begin{align}
0& =\sum_{q=0}^{r}\binom{r}{q}\left[ \frac{1}{n}\sum_{i=1}^{n}\mathrm{E}
\left( x_{i}^{r-q+s_{r}}\right) \mathrm{E}\left( u_{i}^{q}\right) \right]
m_{r-q}-\frac{1}{n}\sum_{i=1}^{n}\mathrm{E}\left( \tilde{y}
_{i}^{r}x_{i}^{s_{r}}\right)  \notag \\
& =\sum_{q=0}^{r}\binom{r}{q} \left[ \frac{1}{n}\sum_{i=1}^{n}\mathrm{E}
\left( x_{i}^{r-q+s_{r}}\right) \right] {\Greekmath 011B}_{q}m_{r-q} - \frac{1}{n}
\sum_{i=1}^{n}\mathrm{E}\left( \tilde{y}_{i}^{r}x_{i}^{s_{r}}\right)+{\Greekmath 010E}
_{n}^{(r,s_{r})},  \label{eq:mc_n}
\end{align}
where
\begin{equation*}
{\Greekmath 010E} _{n}^{(r,s_{r})} =\sum_{q=0}^{r}\binom{r}{q}\left[ \frac{1}{n}
\sum_{i=1}^{n}\mathrm{E}\left( x_{i}^{r-q+s_{r}}\right) \left[ \mathrm{E}
\left( u_{i}^{q}\right) -{\Greekmath 011B} _{q}\right] \right] m_{r-q} = O\left(
n^{-1/2} \right),
\end{equation*}
as shown in the proof of Theorem \ref{lem: identification_moments}.

Taking $n\to \infty$ in \eqref{eq:mc_n},
\begin{equation}
\sum_{q=0}^{r}\binom{r}{q}{\Greekmath 011A} _{0,r-q+s_{r}}{\Greekmath 011B} _{q}m_{r-q}-{\Greekmath 011A}
_{r,s_{r}} = 0,  \label{eq:mc_limit}
\end{equation}
by Assumption \ref{assu:conv_moment}. We stack the left-hand side of
\eqref{eq:mc_limit} over $r=1,2,...,2K-1$, and $s_{r}=0,1,\cdots ,S-r$ and
transform $\mathbf{m}_{\Greekmath 010C} = h\left( \mathbf{{\Greekmath 0112}} \right)$ to get $
\mathbf{g}_0\left(\mathbf{{\Greekmath 0112} }, \mathbf{{\Greekmath 011B}}, \mathbf{{\Greekmath 010D} }
\right) $.

To implement the GMM estimation we replace $\tilde{y}_{i}$, by $\hat{\tilde{y
}}_{i}=y_{i}-\mathbf{z}_{i}^{\prime }\hat{\mathbf{{\Greekmath 010D} }}$, and ${\Greekmath 011A}
_{r,s_{r}}$ by $n^{-1}\sum_{i=1}^{n}\hat{\tilde{y}}_{i}^{r}x_{i}^{s_{r}}$.
Noting that $\mathbf{m}_{{\Greekmath 010C} }=h\left( \mathbf{{\Greekmath 0112} }\right) $, denote
the sample version of the left-hand side of \eqref{eq:mc_limit} by
\begin{equation}
\hat{g}_{n}^{(r,s_{r})}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\hat{\mathbf{
{\Greekmath 010D} }}\right) =\frac{1}{n}\sum_{i=1}^{n}\hat{g}_{i}^{(r,s_{r})}\left(
\mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\hat{\mathbf{{\Greekmath 010D} }}\right) ,
\label{gbarr}
\end{equation}
where
\begin{equation*}
\hat{g}_{i}^{\left( r,s_{r}\right) }\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },
\hat{\mathbf{{\Greekmath 010D} }}\right) =\sum_{q=0}^{r}\binom{r}{q}x_{i}^{r-q+s_{r}}
{\Greekmath 011B} _{q}\left[ h\left( \mathbf{{\Greekmath 0112} }\right) \right] _{r-q}-\hat{\tilde{
y}}_{i}^{r}x_{i}^{s_{r}},
\end{equation*}
and $\mathbf{{\Greekmath 011B} }=\left( {\Greekmath 011B} _{2},{\Greekmath 011B} _{3},\cdots ,{\Greekmath 011B}
_{2K-1}\right) ^{\prime }$. Stack the equations in (\ref{gbarr}), over $
r=0,1,...,2K-1$ and $s_{r}=0,1,\cdots ,S-r$ ($S>2K-1$), in vector notations
we have
\begin{equation}
\mathbf{\hat{g}}_{n}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\hat{\mathbf{
{\Greekmath 010D} }}\right) =\frac{1}{n}\sum_{i=1}^{n}\mathbf{\hat{g}}_{i}\left(
\mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\hat{\mathbf{{\Greekmath 010D} }}\right) .  \label{gn}
\end{equation}
Given $\hat{\mathbf{{\Greekmath 010D} }}$, the GMM estimator of $\left( \mathbf{{\Greekmath 0112} }
^{\prime },\mathbf{{\Greekmath 011B} }^{\prime }\right) ^{\prime }$ is now computed as
\begin{equation*}
\left( \hat{\mathbf{{\Greekmath 0112} }}^{\prime },\hat{\mathbf{{\Greekmath 011B} }}^{\prime
}\right) ^{\prime }=\arg \min_{\mathbf{{\Greekmath 0112} }\in \Theta ,\mathbf{{\Greekmath 011B} }
\in \mathcal{S}}\hat{\Phi}_{n}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\hat{
\mathbf{{\Greekmath 010D} }}\right) ,
\end{equation*}
where $\hat{\Phi}_{n}=\mathbf{\hat{g}}_{n}\left( \mathbf{{\Greekmath 0112} },\mathbf{
{\Greekmath 011B} },\hat{\mathbf{{\Greekmath 010D} }}\right) ^{\prime }\mathbf{A}_{n}\mathbf{\hat{g
}}_{n}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\hat{\mathbf{{\Greekmath 010D} }}\right)
$, and $\mathbf{A}_{n}$ is a positive definite matrix. We follow the GMM
literature using the following choice of $\mathbf{A}_{n}$,
\begin{equation}
\mathbf{\hat{A}}_{n}=\left[ \frac{1}{n}\sum_{i=1}^{n}\mathbf{\hat{g}}
_{i}\left( \tilde{\mathbf{{\Greekmath 0112} }},\tilde{\mathbf{{\Greekmath 011B} }},\hat{\mathbf{
{\Greekmath 010D} }}\right) \mathbf{\hat{g}}_{i}\left( \tilde{\mathbf{{\Greekmath 0112} }},\tilde{
\mathbf{{\Greekmath 011B} }},\hat{\mathbf{{\Greekmath 010D} }}\right) ^{\prime }-\mathbf{\bar{g}}
_{n}\mathbf{\bar{g}}_{n}^{\prime }\right] ^{-1},
\label{eq:efficient_weight_mat}
\end{equation}
where $\mathbf{\bar{g}}_{n}=\frac{1}{n}\sum_{i=1}^{n}\mathbf{\hat{g}}
_{i}\left( \tilde{\mathbf{{\Greekmath 0112} }},\tilde{\mathbf{{\Greekmath 011B} }},\hat{\mathbf{
{\Greekmath 010D} }}\right) $, and $\tilde{\mathbf{{\Greekmath 0112} }}$ and $\tilde{\mathbf{
{\Greekmath 011B} }}$ are preliminary estimators.

\begin{assumption}
\label{assu:Consistency Assumption} Denote the true values of $\mathbf{{\Greekmath 0112}
}$, $\mathbf{{\Greekmath 011B}}$ and $\mathbf{{\Greekmath 010D}}$ by $\mathbf{{\Greekmath 0112} }_{0}$, $
\mathbf{{\Greekmath 011B}}_0$ and $\mathbf{{\Greekmath 010D}}_0 $.

\begin{enumerate}
\item[(a)] $\Theta $ and $\mathcal{S}$ are compact. $\mathbf{{\Greekmath 0112} }_{0}\in
\mathrm{int}\left( \Theta \right)$ and $\mathbf{{\Greekmath 011B}}_0 \in \mathrm{int}
\left( \mathcal{S} \right)$.

\item[(b)] $\mathbf{A}_{n}\rightarrow _{p}\mathbf{A}$ as $n\rightarrow
\infty $, where $\mathbf{A}$ is some positive definite matrix.

\item[(c)] $\,$
\begin{equation*}
\frac{1}{n}\sum_{i=1}^{n}\left[ \hat{\tilde{y}}_{i}^{r}x_{i}^{s_{r}}-\mathrm{
E}\left( \tilde{y}_{i}^{r}x_{i}^{s_{r}}\right) \right] =O_{p}\left(
n^{-1/2}\right) ,
\end{equation*}
for $r=0,1,2,\cdots ,2K-1$, $s_{r}=0,1,\cdots ,S-r,$ and $S>2K-1.$
\end{enumerate}
\end{assumption}

\begin{remark}
Parts (a) and (b) of Assumption \ref{assu:Consistency Assumption} are
standard regularity conditions in the GMM literature. Part (c) together with
Assumption \ref{assu:conv_moment} are high-level regularity conditions which
allow us to generalize the usual IID assumption and nest the IID data
generation process as a special case. The sample analogue terms in (c)
include $\hat{\tilde{y}}_{i}=y_{i}-\mathbf{z}_{i}^{\prime }\hat{\mathbf{
{\Greekmath 010D} }}$, instead of the infeasible $\tilde{y}_{i}=y_{i}-\mathbf{z}
_{i}^{\prime }\mathbf{{\Greekmath 010D} }$. The $\sqrt{n}$-consistency of $\hat{\mathbf{
{\Greekmath 010D} }}$ shown in Theorem \ref{lem:gamma_est_consistency} ensures that
replacing $\tilde{y}_{i}$ by $\hat{\tilde{y}}_{i}$ does not alter the
convergence rate.
\end{remark}

\begin{theorem}
\label{thm:consisteny} Let $\mathbf{{\Greekmath 0111} }=\left( \mathbf{{\Greekmath 0112} }^{\prime },
\mathbf{{\Greekmath 011B} }^{\prime }\right) ^{\prime }$ and $\mathbf{{\Greekmath 0111} }_{0}=\left(
\mathbf{{\Greekmath 0112} }_{0}^{\prime },\mathbf{{\Greekmath 011B} }_{0}^{\prime }\right)
^{\prime }$. Under Assumptions \ref{assu:identification_regularity_condition}
, \ref{assu:conv_moment}, and \ref{assu:Consistency Assumption}, $\hat{
\mathbf{{\Greekmath 0111} }}\rightarrow _{p}\mathbf{{\Greekmath 0111} }_{0}$ as $n\rightarrow \infty $.
\end{theorem}

The proof of Theorem \ref{thm:consisteny} is provided in Appendix \ref
{sec:Proofs}.

\begin{assumption}
\label{assu:normality} Follow the notations as in Assumption \ref
{assu:Consistency Assumption} and in addition denote $\mathbf{G}\left(
\mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\mathbf{{\Greekmath 010D} }\right) ={\Greekmath 0272} _{\left(
\mathbf{{\Greekmath 0112} }^{\prime }, \mathbf{{\Greekmath 011B} }^{\prime }\right) ^{\prime }}
\mathbf{g}_{0}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\mathbf{{\Greekmath 010D} }
\right) $, $\mathbf{G}_{0}=\mathbf{G}\left( \mathbf{{\Greekmath 0112} }_{0},\mathbf{
{\Greekmath 011B} }_{0},\mathbf{{\Greekmath 010D} }_{0}\right) $, $\mathbf{G}_{{\Greekmath 010D} }\left(
\mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\mathbf{{\Greekmath 010D} }\right) ={\Greekmath 0272} _{\mathbf{
{\Greekmath 010D} }}\mathbf{g}_{0}\left( \mathbf{{\Greekmath 0112} },\mathbf{{\Greekmath 011B} },\mathbf{
{\Greekmath 010D} }\right) $, $\mathbf{G}_{0, {\Greekmath 010D}}=\mathbf{G}_{{\Greekmath 010D}} \left(
\mathbf{{\Greekmath 0112} }_{0},\mathbf{{\Greekmath 011B} }_{0},\mathbf{{\Greekmath 010D} }_{0}\right) $.

\begin{enumerate}
\item[(a)] $\sqrt{n}\mathbf{\hat{g}}_{n}\left( \mathbf{{\Greekmath 0112} }_{0},\mathbf{
{\Greekmath 011B} }_{0},\mathbf{{\Greekmath 010D}}_0\right) \rightarrow _{d}\mathbf{{\Greekmath 0110} }\sim
N\left( 0,\mathbf{V}\right) $ as $n\to\infty$.

\item[(b)] $\mathbf{G}_{0}^{\prime }\mathbf{AG}_{0}\succ 0$.
\end{enumerate}
\end{assumption}

\begin{remark}
In Assumption \ref{assu:normality}, parts (a) is the high level condition
required to ensure the asymptotic normality of $\mathbf{\hat{g}}_{n}\left(
\mathbf{{\Greekmath 0112} }_{0},\mathbf{{\Greekmath 011B} }_{0},\mathbf{{\Greekmath 010D}}_0\right) $, which
can be verified by Lindeberg central limit theorem under low-level
regularity conditions. Part (c) of Assumption \ref{assu:normality}
represents the full-rank condition on $\mathbf{G}_{0}$, required for
identification of $\mathbf{{\Greekmath 0112} }_{0}$ and $\mathbf{{\Greekmath 011B} }_{0}$.
\end{remark}

By Theorem \ref{lem:gamma_est_consistency}, we have $\sqrt{n}\left( \hat{
\mathbf{{\Greekmath 010D}}} - \mathbf{{\Greekmath 010D}} \right) \rightarrow_d {\Greekmath 0110}_{\Greekmath 010D} \sim
N(0, V_{\Greekmath 010D})$. The following theorem shows the asymptotic normality of the
GMM estimator $\hat{\mathbf{{\Greekmath 0111}}}$.

\begin{theorem}
\label{thm:normality} Under Assumptions \ref
{assu:identification_regularity_condition}, \ref{assu:gamma_est_normality},
\ref{assu:Consistency Assumption} and \ref{assu:normality},
\begin{equation*}
\sqrt{n}\left( \hat{\mathbf{{\Greekmath 0111} }}-\mathbf{{\Greekmath 0111} }_{0}\right) \rightarrow
_{d}\left( \mathbf{G}_{0}^{\prime }\mathbf{A}\mathbf{G}_{0}\right) ^{-1}
\mathbf{G}_{0}^{\prime }\mathbf{A}\left( \mathbf{{\Greekmath 0110} }+\mathbf{G}_{0,
{\Greekmath 010D}}\mathbf{{\Greekmath 0110} }_{{\Greekmath 010D} }\right),
\end{equation*}
as $n\rightarrow \infty $.
\end{theorem}

The proof of Theorem \ref{thm:normality} is provided in Appendix \ref
{sec:Proofs}.

\begin{remark}
In practice, we estimate the variance of the asymptotic distribution of $
\hat{{\Greekmath 0111}}$ by
\begin{equation}  \label{eq:est_variance_matrix_gmm}
\hat{\mathbf{V}}_{{\Greekmath 0111}} = \left( \hat{\mathbf{G}}^\prime \hat{\mathbf{A}}_n
\hat{\mathbf{G}} \right)^{-1} \hat{\mathbf{G}}^\prime \hat{\mathbf{A}}_n
\hat{\mathbf{V}}_{{\Greekmath 0110}} \hat{\mathbf{A}}_n^\prime \hat{\mathbf{G}} \left(
\hat{\mathbf{G}}^\prime \hat{\mathbf{A}}_n \hat{\mathbf{G}} \right)^{-1},
\end{equation}
where $\hat{\mathbf{G}} = {\Greekmath 0272}_{\left( \mathbf{{\Greekmath 011B} }^{\prime },\mathbf{
{\Greekmath 0112} }^{\prime }\right) ^{\prime }} \hat{\mathbf{g}}_{n}\left( \hat{
\mathbf{{\Greekmath 0112}}}, \hat{\mathbf{{\Greekmath 011B}}}, \hat{{\Greekmath 010D}} \right)$, $\hat{
\mathbf{A}}_n$ is given by \eqref{eq:efficient_weight_mat}, and
\begin{equation*}
\hat{\mathbf{V}}_{\Greekmath 0110} = \frac{1}{n} \sum_{i=1}^{n} \mathbf{{\Greekmath 0120}}_{n,i}
\mathbf{{\Greekmath 0120}}_{n,i}^\prime,
\end{equation*}
where
\begin{equation*}
\mathbf{{\Greekmath 0120}}_{n,i} = \hat{\mathbf{g}}_i\left( \hat{\mathbf{{\Greekmath 0112}}}, \hat{
\mathbf{{\Greekmath 011B}}}, \hat{{\Greekmath 010D}} \right) + {\Greekmath 0272} _{\mathbf{{\Greekmath 010D} }}\hat{
\mathbf{g}}_{n}\left( \hat{\mathbf{{\Greekmath 0112}}}, \hat{\mathbf{{\Greekmath 011B}}}, \hat{
{\Greekmath 010D}}\right) \mathbf{L} \mathbf{Q}_{n,ww}^{-1}\left( \mathbf{w}_i\hat{{\Greekmath 0118}}_i
\right),
\end{equation*}
and $\mathbf{L} =
\begin{pmatrix}
\mathbf{0}_{p_z\times 1} & \mathbf{I}_{p_z}
\end{pmatrix}
$ is the loading matrix that selects $\mathbf{{\Greekmath 010D}}$ out of $\mathbf{{\Greekmath 011E}}$
.
\end{remark}


\section{Multiple regressors with random coefficients\label{sec:Extensions}}

One important extension of the regression model \eqref{eq:basic_model} is to
allow for multiple regressors with random coefficients having categorical
distribution. With this in mind consider
\begin{equation}
y_{i}=\mathbf{x}_{i}^{\prime }\mathbf{{\Greekmath 010C} }_{i}+\mathbf{z}_{i}^{\prime }
\mathbf{{\Greekmath 010D} }+u_{i},  \label{eq:multiple_x_model}
\end{equation}
where the $p\times 1$ vector of random coefficients, $\mathbf{{\Greekmath 010C} }_{i}\in
\mathbb{R}^{p}$ follows the multivariate distribution\footnote{
We assume the number of categories $K$ is homogeneous across $j=1,2,\cdots
,p $. This is for notational simplicity, and can be readily generalized to
allow for $K_{j}\neq K_{j^{\prime }}$ without affecting the main results.}
\begin{equation}
\mathrm{Pr}\left( {\Greekmath 010C} _{i1}=b_{1k_{1}},{\Greekmath 010C} _{i2}=b_{2k_{2}},\cdots
,{\Greekmath 010C} _{ip}=b_{pk_{p}}\right) ={\Greekmath 0119} _{k_{1},k_{2},\cdots ,k_{p}},
\label{eq:prob_dist_ext}
\end{equation}
with $k_{j}\in \left\{ 1,2,\cdots ,K\right\} $, $b_{j1}<b_{j2}<\cdots
<b_{jK} $, and
\begin{equation*}
\sum_{k_{1},k_{2},\cdots ,k_{p}\in \left\{ 1,2,\cdots ,K\right\} }{\Greekmath 0119}
_{k_{1},k_{2},\cdots ,k_{p}}=1.
\end{equation*}
As in Section \ref{sec:Identification}, $\mathbf{{\Greekmath 010D} }\in \mathbb{R}
^{p_{z}}$, $\mathbf{w}_{i}=\left( \mathbf{x}_{i}^{\prime },\mathbf{z}
_{i}^{\prime }\right) ^{\prime }$, $\mathbf{{\Greekmath 010C} }_{i}\perp \mathbf{w}_{i}$
, $u_{i}\perp \mathbf{w}_{i}$, and $u_{i}$ are independently distributed
over $i$ with mean $0.$

\medskip
\noindent \textbf{Example 1} \textit{Consider the simple case with $p = 2$
and $K = 2$. For $j = 1, 2$, denote two categories as $\left\{ L, H \right\}$
. The probabilities of four possible combinations of realized $\mathbf{{\Greekmath 010C}}
_i$ is summarized in Table \ref{tab:dist_beta_p2_k2}, where ${\Greekmath 0119}_{LL} +
{\Greekmath 0119}_{LH} + {\Greekmath 0119}_{HL} + {\Greekmath 0119}_{HH} = 1$. }
\begin{table}[htbp]
\caption{Distribution of $\mathbf{\protect{\Greekmath 010C}}_i$ with $p = 2$ and $K = 2$}
\label{tab:dist_beta_p2_k2}
\begin{center}
\begin{tabular}{c|c|c}
\hline
& $k_2 = L$ & $k_2 = H$ \\ \hline
$k_1 = L$ & ${\Greekmath 0119}_{LL} = \Pr \left( {\Greekmath 010C}_{i1} = b_{1L}, {\Greekmath 010C}_{i2} = b_{2L}
\right)$ & ${\Greekmath 0119}_{LH} = \Pr \left( {\Greekmath 010C}_{i1} = b_{1L}, {\Greekmath 010C}_{i2} = b_{2H}
\right)$ \\ \hline
$k_1 = H$ & ${\Greekmath 0119}_{HL} = \Pr \left( {\Greekmath 010C}_{i1} = b_{1H}, {\Greekmath 010C}_{i2} = b_{2L}
\right)$ & ${\Greekmath 0119}_{HH} = \Pr \left( {\Greekmath 010C}_{i1} = b_{1H}, {\Greekmath 010C}_{i2} = b_{2H}
\right)$ \\ \hline
\end{tabular}
\end{center}
\end{table}
\medskip

We first identify the moments of $\mathbf{{\Greekmath 010C} }_{i}$. As in Section \ref
{sec:Identification}, $\mathbf{{\Greekmath 011E} }=\left( \mathrm{E}\left( \mathbf{{\Greekmath 010C} }
_{i}\right) ^{\prime },\mathbf{{\Greekmath 010D} }^{\prime }\right) ^{\prime }$ is
identified by
\begin{equation}
\mathbf{{\Greekmath 011E} } = \mathbf{Q}_{ww} ^{-1} \mathbf{q}_{wy} ,
\label{eq:identify_phi_ext}
\end{equation}
under Assumption \ref{assu:identification_regularity_condition}. We now
consider the identification of the higher order moments of $\mathbf{{\Greekmath 010C} }
_{i}$ up to the finite order $2K-1$.

Since $\mathbf{{\Greekmath 010D} }$ is identified as in \eqref{eq:identify_phi_ext}, we
treat it as known and let $\tilde{y}_{i}^{r}=y_{i}-\mathbf{z}_{i}^{\prime }
\mathbf{{\Greekmath 010D} }$. For $r=2,3,\cdots ,2K-1$, consider the moment conditions
\begin{align}
\mathrm{E}\left( \tilde{y}_{i}^{r}\right) & =\mathrm{E}\left[ \left( \mathbf{
x}_{i}^{\prime }\mathbf{{\Greekmath 010C} }_{i}+u_{i}\right) ^{r}\right]  \notag \\
& =\mathrm{E}\left[ \left( \mathbf{x}_{i}^{\prime }\mathbf{{\Greekmath 010C} }
_{i}\right) ^{r}\right] +\mathrm{E}\left( u_{i}^{r}\right) +\sum_{s=2}^{r-1}
\binom{r}{s}\mathrm{E}\left[ \left( \mathbf{x}_{i}^{\prime }\mathbf{{\Greekmath 010C} }
_{i}\right) ^{r-s}\right] \mathrm{E}\left( u_{i}^{s}\right) .
\label{eq:moment_y_r_ext}
\end{align}
Note that $\mathbf{x}_{i}^{\prime }\mathbf{{\Greekmath 010C} }_{i}=\sum_{j=1}^{p}{\Greekmath 010C}
_{ij}x_{ij}$, and
\begin{equation*}
\mathrm{E}\left[ \left( \sum_{j=1}^{p}{\Greekmath 010C} _{ij}x_{ij}\right) ^{r}\right]
=\sum_{\sum_{j=1}^{p}q_{j}=r}\binom{r}{\mathbf{q}}\mathrm{E}\left(
\prod_{j=1}^{p}x_{ij}^{q_{j}}\right) \mathrm{E}\left( \prod_{j=1}^{p}{\Greekmath 010C}
_{ij}^{q_{j}}\right) ,
\end{equation*}
where $\binom{r}{\mathbf{q}}=\frac{r!}{q_{1}!q_{2}!\cdots q_{p}!}$, for
non-negative integers $r$, $q_{1}$, $\cdots $, $q_{p}$ with $
r=\sum_{j=1}^{p}q_{j}$, denotes the multinomial coefficients. We stack $
\prod_{j=1}^{p}x_{ij}^{q_{j}}$ with $\mathbf{q}\in \left\{ \mathbf{q}\in
\left\{ 0,1,\cdots r\right\} ^{p}:\sum_{j=1}^{p}q_{j}=r\right\} $ in a
vector form by denoting \footnote{
For $\mathbf{x}\in \mathbb{R}^{p}$, note that $\mathbf{{\Greekmath 011C} }_{0}\left(
\mathbf{x}\right) =1$, $\mathbf{{\Greekmath 011C} }_{1}\left( \mathbf{x}\right) =\mathbf{x
}$ and $\mathbf{{\Greekmath 011C} }_{2}\left( \mathbf{x}\right) =\mathrm{vech}\left(
\mathbf{x}\mathbf{x}^{\prime }\right) $.}
\begin{equation*}
\mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{x}_{i}\right) =\left[ {\Greekmath 0127} \left(
\mathbf{x}_{i},\mathbf{q}_{1}\right) ,{\Greekmath 0127} \left( \mathbf{x}_{i},\mathbf{q
}_{2}\right) ,\cdots ,{\Greekmath 0127} \left( \mathbf{x}_{i},\mathbf{q}_{{\Greekmath 0117}
_{r}}\right) \right] ^{\prime },
\end{equation*}
where ${\Greekmath 0127} \left( \mathbf{x}_{i},\mathbf{q}\right)
=\prod_{j=1}^{p}x_{ij}^{q_{j}}$ and ${\Greekmath 0117} _{r}=\binom{r+p-1}{p-1}$ is the
number of distinct monomials of degree $r$ on the variables $
x_{i1},x_{i2},\cdots ,x_{ip}$. Similarly,
\begin{equation*}
\mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{{\Greekmath 010C} }_{i}\right) =\left[ {\Greekmath 0127} \left(
\mathbf{{\Greekmath 010C} }_{i},\mathbf{q}_{1}\right) ,{\Greekmath 0127} \left( \mathbf{{\Greekmath 010C} }
_{i},\mathbf{q}_{2}\right) ,\cdots ,{\Greekmath 0127} \left( \mathbf{{\Greekmath 010C} }_{i},
\mathbf{q}_{{\Greekmath 0117} _{r}}\right) \right] ^{\prime },
\end{equation*}
where ${\Greekmath 0127} \left( \mathbf{{\Greekmath 010C} }_{i},\mathbf{q}\right)
=\prod_{j=1}^{p}{\Greekmath 010C} _{ij}^{q_{j}}$.

\medskip
\noindent \textbf{Example 2} \textit{Consider $p = 2$ and $r = 2$, we have
\begin{align*}
\mathbf{{\Greekmath 011C}}_2\left( \mathbf{x}_i \right) & = \left( x_{i1}^2,
x_{i1}x_{i2}, x_{i2}^2 \right)^\prime, \\
\mathbf{{\Greekmath 011C}}_2\left( \mathbf{{\Greekmath 010C}}_i \right) & = \left( {\Greekmath 010C}_{i1}^2,
{\Greekmath 010C}_{i1}{\Greekmath 010C}_{i2}, {\Greekmath 010C}_{i2}^2 \right)^\prime,
\end{align*}
and
\begin{align*}
\mathrm{E}\left[ \left( x_{i1}{\Greekmath 010C}_{i1} + x_{i2}{\Greekmath 010C}_{i2} \right)^2 \right]
& = \mathrm{E}\left(x_{i1}^2\right)\mathrm{E}\left({\Greekmath 010C}_{i1}^2\right) + 2
\mathrm{E}\left(x_{i1}x_{i2}\right)\mathrm{E}\left({\Greekmath 010C}_{i1}{\Greekmath 010C}_{i2}
\right) + \mathrm{E}\left(x_{i2}^2 \right)\mathrm{E}\left({\Greekmath 010C}_{i2}^2\right)
\\
& = \left[ \mathrm{E}\left(x_{i1}^2\right), \mathrm{E}\left(x_{i1}x_{i2}
\right), \mathrm{E}\left(x_{i2}^2 \right) \right] \mathrm{diag}\left[ \left(
1, 2, 1 \right)^\prime \right] \left[ \mathrm{E}\left({\Greekmath 010C}_{i1}^2\right),
\mathrm{E}\left({\Greekmath 010C}_{i1}{\Greekmath 010C}_{i2}\right), \mathrm{E}\left({\Greekmath 010C}_{i2}^2
\right) \right]^\prime \\
& = \mathrm{E}\left[\mathbf{{\Greekmath 011C}}_2\left( \mathbf{x}_i \right)\right]^\prime
\mathbf{\Lambda}_2 \mathrm{E}\left[\mathbf{{\Greekmath 011C}}_2\left( \mathbf{{\Greekmath 010C}}_i
\right)\right] ,
\end{align*}
where $\mathbf{\Lambda}_2 = \mathrm{diag}\left[ \left( 1, 2, 1
\right)^\prime \right]$.}
\medskip

Then the moment condition \eqref{eq:moment_y_r_ext} can be written as
\begin{align}
\mathrm{E}\left( \tilde{y}_{i}^{r}\right) & =\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }
_{r}\left( \mathbf{x}_{i}\right) \right] ^{\prime }\mathbf{\Lambda }_{r}
\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{{\Greekmath 010C} }_{i}\right) \right]
+\mathrm{E}\left( u_{i}^{r}\right)  \notag \\
& \quad \quad \quad +\sum_{s=2}^{r-1}\binom{r}{s}\mathrm{E}\left[ \mathbf{
{\Greekmath 011C} }_{r-s}\left( \mathbf{x}_{i}\right) \right] ^{\prime }\mathbf{\Lambda }
_{r-s}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r-s}\left( \mathbf{{\Greekmath 010C} }_{i}\right)
\right] \mathrm{E}\left( u_{i}^{s}\right) ,  \label{eq:moment_y_r_ext_vec}
\end{align}
where $\mathbf{\Lambda }_{r}=\mathrm{diag}\left[ \left[ \binom{r}{\mathbf{q}}
\right] _{\sum_{j=1}^{p}q_{j}=r}\right] $ is the ${\Greekmath 0117} _{r}\times {\Greekmath 0117} _{r}$
diagonal matrix of multinomial coefficients. We further consider the moment
conditions
\begin{align}
\mathrm{E}\left( \tilde{y}_{i}^{r}\mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{x}
_{i}\right) \right) & =\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{x}
_{i}\right) \mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{x}_{i}\right) ^{\prime }\right]
\mathbf{\Lambda }_{r}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{{\Greekmath 010C}
}_{i}\right) \right] +\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r}\left( \mathbf{x}
_{i}\right) \right] \mathrm{E}\left( u_{i}^{r}\right)  \notag \\
& \quad \quad \quad +\sum_{s=2}^{r-1}\binom{r}{s}\mathrm{E}\left[ \mathbf{
{\Greekmath 011C} }_{r}\left( \mathbf{x}_{i}\right) \mathbf{{\Greekmath 011C} }_{r-s}\left( \mathbf{x}
_{i}\right) ^{\prime }\right] \mathbf{\Lambda }_{r-s}\mathrm{E}\left[
\mathbf{{\Greekmath 011C} }_{r-s}\left( \mathbf{{\Greekmath 010C} }_{i}\right) \right] \mathrm{E}
\left( u_{i}^{s}\right) ,  \label{eq:moment_y_r_x_r_ext_vec}
\end{align}
$r=2,3,\cdots ,2K-1$. \eqref{eq:moment_y_r_ext_vec} and
\eqref{eq:moment_y_r_x_r_ext_vec} reduce to \eqref{eq:mc_n_r} and
\eqref{eq:mc_n_2r} when $p=1$.

\begin{assumption}
\label{assu:idenfication_regularity_condition_ext} $\,$

\begin{enumerate}
\item[(a)] $\left\Vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( \tilde{y}_{i}^{r}
\mathbf{{\Greekmath 011C} }_{s}\left( \mathbf{x}_{i}\right) \right) -\mathbf{{\Greekmath 011A} }
_{r,s}\right\Vert =O\left( n^{-1/2}\right) ,$ and $\left\Vert \mathbf{{\Greekmath 011A} }
_{r,s}\right\Vert <\infty $, $r,s=0,1,\cdots ,2K-1$.

\item[(b)] $\left\Vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }
_{r}\left( \mathbf{x}_{i}\right) \mathbf{{\Greekmath 011C} }_{s}\left( \mathbf{x}
_{i}\right) ^{\prime }\right] -\mathbf{\Xi }_{r,s}\right\Vert =O\left(
n^{-1/2}\right) ,$ and $\left\Vert \mathbf{\Xi }_{r,s}\right\Vert <\infty $,
$r,s=0,1,\cdots ,2K-1$.

\item[(c)] $\left\vert n^{-1}\sum_{i=1}^{n}\mathrm{E}\left( u_{i}^{r}\right)
-{\Greekmath 011B} _{r}\right\vert =O\left( n^{-1/2}\right) ,$ and $\left\vert {\Greekmath 011B}
_{r}\right\vert <\infty $ for $r=2,3,\cdots ,2K-1$.

\item[(d)] $\left\Vert n^{-1}\sum_{i=1}^{n} \left[ \mathrm{var} \left(
\mathbf{{\Greekmath 011C}}_r \left(\mathbf{x_i} \right) \right) - \left( \mathbf{\Xi}
_{r,r} - \mathbf{{\Greekmath 011A}}_{0, r}\mathbf{{\Greekmath 011A}}_{0, r}^\prime \right) \right]
\right\Vert = O(n^{-1/2})$, where $\mathbf{\Xi}_{r,r} - \mathbf{{\Greekmath 011A}}_{0, r}
\mathbf{{\Greekmath 011A}}_{0, r}^\prime \succ 0$ for $r=2,3\cdots, 2K-1$.
\end{enumerate}
\end{assumption}

\begin{theorem}
\label{thm: idenfitication_moment_beta_ext} For any $\mathbf{q} \in \left\{
\mathbf{q}\in \left\{ 0, 1, \cdots r \right\} ^{p}: \sum_{j=1}^p q_j =
r\right\}$ and $r = 2, 3,\cdots, 2K-1$, $\mathrm{E} \left(\prod_{j=1}^{p}
{\Greekmath 010C}_{ij}^{q_{j}}\right)$ and ${\Greekmath 011B}_r$ are identified under Assumptions
\ref{assu:identification_regularity_condition} and \ref
{assu:idenfication_regularity_condition_ext}.
\end{theorem}

\begin{proof}
For $r=2,3,\cdots ,2K-1$, sum \eqref{eq:moment_y_r_ext_vec} and
\eqref{eq:moment_y_r_x_r_ext_vec} over $i,$ go through the same steps as in
the proof of Theorem \ref{lem: identification_moments}, then by Assumptions
\ref{assu:idenfication_regularity_condition_ext}(a) to (c), we have (for $
n\rightarrow \infty $)
\begin{align}
\mathbf{{\Greekmath 011A} }_{r,0}^{\prime }\mathbf{\Lambda }_{r}\mathrm{E}\left[ \mathbf{
{\Greekmath 011C} }_{r}\left( \mathbf{{\Greekmath 010C} }_{i}\right) \right] +{\Greekmath 011B} _{r}& =\mathbf{
{\Greekmath 011A} }_{r,0}-\sum_{s=2}^{r-1}\binom{r}{s}\mathbf{{\Greekmath 011A} }_{0,r-s}\mathbf{
\Lambda }_{r-s}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r-s}\left( \mathbf{{\Greekmath 010C} }
_{i}\right) \right] {\Greekmath 011B} _{s},  \label{eq:mc_limit_r_ext} \\
\mathbf{\Xi }_{r,r}\mathbf{\Lambda }_{r}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }
_{r}\left( \mathbf{{\Greekmath 010C} }_{i}\right) \right] +\mathbf{{\Greekmath 011A} }_{0,r}{\Greekmath 011B}
_{r}& =\mathbf{{\Greekmath 011A} }_{r,r}-\sum_{s=2}^{r-1}\binom{r}{s}\mathbf{\Xi }_{r,r-s}
\mathbf{\Lambda }_{r-s}\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r-s}\left( \mathbf{
{\Greekmath 010C} }_{i}\right) \right] {\Greekmath 011B} _{s}.  \label{eq:mc_limit_2r_ext}
\end{align}
Note that
\begin{equation*}
\mathbf{M}_{r}=
\begin{pmatrix}
\mathbf{\Xi }_{r,r} & \mathbf{{\Greekmath 011A} }_{0,r} \\
\mathbf{{\Greekmath 011A} }_{0,r}^{\prime } & 1
\end{pmatrix}
\begin{pmatrix}
\mathbf{\Lambda }_{r} & \mathbf{0} \\
\mathbf{0} & 1
\end{pmatrix}
,
\end{equation*}
is invertible since $\det \left( \mathbf{M}_{r}\right) =\det \left( \mathbf{
\Xi }_{r,r}-\mathbf{{\Greekmath 011A} }_{0,r}\mathbf{{\Greekmath 011A} }_{0,r}^{\prime }\right) \det
\left( \mathbf{\Lambda }_{r}\right) >0,$ for $r=2,3,\cdots ,R$, by
Assumption \ref{assu:idenfication_regularity_condition_ext}(d). As a result,
we can sequentially solve \eqref{eq:mc_limit_r_ext} and
\eqref{eq:mc_limit_2r_ext} for $\mathrm{E}\left[ \mathbf{{\Greekmath 011C} }_{r}\left(
\mathbf{{\Greekmath 010C} }_{i}\right) \right] $ and ${\Greekmath 011B} _{r}$, for $r=2,3,\cdots
,2K-1$.
\end{proof}


We now move from the moments of $\mathbf{{\Greekmath 010C} }_{i}$ to the distribution of
$\mathbf{{\Greekmath 010C} }_{i}$. We first focus on the identification of the marginal
probabilities obtained from (\ref{eq:prob_dist_ext}) by averaging out the
effects of the other coefficients except for ${\Greekmath 010C} _{ij}$, namely we
initially focus on identification of ${\Greekmath 0115} _{jk}=\Pr \left( {\Greekmath 010C}
_{ij}=b_{jk}\right) $, for $k=1,2,\cdots ,K,$ and $j=1,2,\cdots ,p$.

\begin{remark}
Focusing on the marginal distribution of ${\Greekmath 010C} _{i}$ is similar to focusing
on estimation of partial derivatives in the context of non-parametric
estimation, where the curse of dimensionality applies. Consider the
estimation of regressing $y_{i}$ on $\mathbf{x}_{i}=\left(
x_{i1},x_{i2},\cdots ,x_{ip}\right) ^{\prime }$,
\begin{equation*}
y_{i}=F\left( x_{i1},x_{i2},\cdots .x_{ip}\right) +u_{i}.
\end{equation*}
Then if $F\left( x_{1},x_{i2},\cdots ,x_{ip}\right) $ is a homogeneous
function (of degree $1/{\Greekmath 0116} $), then
\begin{equation*}
y_{i}=\sum_{j=1}^{p}\left( {\Greekmath 0116} \frac{\partial F\left( \cdot \right) }{
\partial x_{ij}}\right) x_{ij}+u_{i},
\end{equation*}
and under certain conditions we can treat ${\Greekmath 0116} \frac{\partial F\left( \cdot
\right) }{\partial x_{ij}}\equiv {\Greekmath 010C} _{ij}$.
\end{remark}

By Theorem \ref{thm: idenfitication_moment_beta_ext}, $\mathrm{E}\left(
{\Greekmath 010C} _{ij}^{r}\right) $ is identified for $r=1,2,\cdots ,2K-1$ under
Assumptions \ref{assu:identification_regularity_condition} and \ref
{assu:idenfication_regularity_condition_ext}. By \eqref{eq:prob_dist_ext},
we have equations
\begin{equation}
\mathrm{E}\left( {\Greekmath 010C} _{ij}^{r}\right) =\sum_{k=1}^{K}{\Greekmath 0115}
_{jk}b_{jk}^{r},  \label{eq:mbeta_ext}
\end{equation}
$r=0,1,\cdots ,2K-1$, which is of the same form as \eqref{mbeta} and
\eqref{eq:mbeta_mat}. To identify $\mathbf{{\Greekmath 0115} }_{j}=\left( {\Greekmath 0115}
_{j1},{\Greekmath 0115} _{j2},\cdots ,{\Greekmath 0115} _{jK}\right) ^{\prime }$ and $\mathbf{b}
_{j}=\left( b_{j1},b_{j2},\cdots ,b_{jK}\right) ^{\prime }$, we can verify
the system of $2K$ equations in \eqref{eq:mbeta_ext} has a unique solution
if $b_{j1}<b_{j2}<\cdots <b_{jK}$ and ${\Greekmath 0115} _{jk}\in \left( 0,1\right) $.
The following corollary is a direct application of Theorem \ref
{prop:identification_of_beta_L_H}.
\begin{corollary}
    \label{core:marginal_dist}
    Consider the model \eqref{eq:multiple_x_model} and suppose that Assumptions \ref{assu:identification_regularity_condition} and \ref{assu:idenfication_regularity_condition_ext} hold. Then the parameters $\mathbf{{\Greekmath 0112}}_j = \left( \mathbf{{\Greekmath 0115}}_{j}^\prime, \mathbf{b}_{j}^\prime \right)^\prime$ of the marginal distribution of ${\Greekmath 010C}_i$ with respect to ${\Greekmath 010C}_{ij}$ is identified subject to $b_{j1} < b_{j2} < \cdots < b_{jK}$ and ${\Greekmath 0115}_{jk} \in \left( 0, 1 \right)$ for $j = 1, 2, \cdots, p$.
\end{corollary}

The problem of identification and estimation of the joint distribution of $
\mathbf{{\Greekmath 010C} }_{i}$ is subject to the curse of dimensionality. We have $
K^{p}-1$ probability weights, ${\Greekmath 0119} _{k_{1},k_{2},\cdots ,k_{p}}$, to be
identified in addition to the $pK$ categorical coefficients $b_{ij}$ that
are identified by Corollary \ref{core:marginal_dist}. The number of
parameters increases rapidly with $p$. Even in the simplest case with $K=2$,
the total number of unknown parameters is $2p+2^{p}-1$, which grows
exponentially.

Note that the marginal probabilities ${\Greekmath 0115} _{jk}$ are related to the
joint distribution by
\begin{equation}
{\Greekmath 0115} _{jk}=\sum_{k_{1},\cdots ,k_{j-1},k_{j+1},\cdots ,k_{p}\in \left\{
1,2,\cdots ,K\right\} }{\Greekmath 0119} _{k_{1},k_{2},\cdots ,k_{j-1},k,k_{j+1},\cdots
,k_{p}},  \label{eq:marginal_dist_prob}
\end{equation}
$k=1,2,\cdots ,K$ and $j=1,2,\cdots ,p$. The number of linearly independent
equations in \eqref{eq:marginal_dist_prob} is $pK-(p-1)$.

\medskip
\noindent\textbf{Example 3} \textit{\ Consider the same setup as in Example
1 with $p = 2$ and $K = 2$. The marginal probabilities are obtained by
\begin{align}
{\Greekmath 0115}_{1L} = \Pr\left( {\Greekmath 010C}_{i1} = b_{1L} \right) = {\Greekmath 0119}_{LL} + {\Greekmath 0119}_{LH},
& \quad {\Greekmath 0115}_{1H} = \Pr\left( {\Greekmath 010C}_{i1} = b_{1H} \right) = 1 -
{\Greekmath 0115}_{1L} = {\Greekmath 0119}_{HL} + {\Greekmath 0119}_{HH},  \notag \\
{\Greekmath 0115}_{2L} = \Pr\left( {\Greekmath 010C}_{i2} = b_{2L} \right) = {\Greekmath 0119}_{LL} + {\Greekmath 0119}_{HL},
& \quad {\Greekmath 0115}_{2H} = \Pr\left( {\Greekmath 010C}_{i2} = b_{2H} \right) = 1 -
{\Greekmath 0115}_{2L} = {\Greekmath 0119}_{LH} + {\Greekmath 0119}_{HH} .  \label{eq:marginal_prob_p2k2}
\end{align}
Note that any equation in \eqref{eq:marginal_prob_p2k2} can be expressed as
a linear combination of other three equations, for example ${\Greekmath 0115}_{2H} =
{\Greekmath 0115}_{1L} + {\Greekmath 0115}_{1H} - {\Greekmath 0115}_{2L}$. }
\medskip

The equations corresponding to the cross-moments, $\mathrm{E}\left(
\prod_{j=1}^{p}{\Greekmath 010C} _{ij}^{q_{j}}\right) $, are
\begin{equation}
\mathrm{E}\left( \prod_{j=1}^{p}{\Greekmath 010C} _{ij}^{q_{j}}\right)
=\sum_{k_{1},k_{2},\cdots ,k_{p}\in \left\{ 1,2,\cdots ,K\right\} }\left(
\prod_{j=1}^{p}b_{jk_{j}}^{q_{j}}\right) {\Greekmath 0119} _{k_{1},k_{2},\cdots ,k_{p}},
\label{eq:cross_moments}
\end{equation}
for $\mathbf{q}\in \left\{ \mathbf{q}\in \left\{ 0,1,\cdots r-1\right\}
^{p}:\sum_{j=1}^{p}q_{j}=r\right\} $, $r=2,\cdots ,2K-1$. The linear system
\eqref{eq:cross_moments} has
\begin{equation*}
\sum_{r=1}^{2K-1}\binom{r+p-1}{p-1}-p(2K-1)
\end{equation*}
equations. Then the total number of equations in
\eqref{eq:marginal_dist_prob} and \eqref{eq:cross_moments} that can be
utilized to identify joint probabilities is $C_{r}=\sum_{r=1}^{2K-1}\binom{
r+p-1}{p-1}-pK$, which is smaller than the number of joint probabilities $
K^{p}-1$ for large $p$. When $K=2$, $C_{r}<K^{p}-1$ for $p\geq 7$.

Identification and estimation of the joint distribution of $\mathbf{{\Greekmath 010C} }
_{i}$ in the general setting will not be pursued in this paper due to the
curse of dimensionality. Instead, we consider special cases, that are
empirically relevant, in which identification of the joint distribution of $
\mathbf{{\Greekmath 010C} }_{i}$ can be readily established. We first consider small $p$
and $K$, in particular $p=2$ and $K=2$ as in Example 1.

\medskip \noindent \textbf{Example 4}
Consider the same setup as in Example 1 with $p=2$ and $K=2$. In addition to
\eqref{eq:marginal_prob_p2k2}, consider the cross-moment,
\begin{equation}
\mathrm{E}\left( {\Greekmath 010C} _{i1}{\Greekmath 010C} _{i2}\right) =b_{1L}b_{2L}{\Greekmath 0119}
_{LL}+b_{1L}b_{2H}{\Greekmath 0119} _{LH}+b_{1H}b_{2L}{\Greekmath 0119} _{HL}+b_{1H}b_{2H}{\Greekmath 0119} _{HH}.
\label{eq:cross_moments_p2k2}
\end{equation}
Writing \eqref{eq:marginal_prob_p2k2} and \eqref{eq:cross_moments_p2k2} in
matrix form, we have
\begin{equation*}
\mathbf{B}\mathbf{{\Greekmath 0119} }=\mathbf{{\Greekmath 0115} },
\end{equation*}
where
\begin{equation*}
\mathbf{B}=
\begin{pmatrix}
1 & 1 & 0 & 0 \\
0 & 0 & 1 & 1 \\
1 & 0 & 1 & 0 \\
b_{1L}b_{2L} & b_{1L}b_{2H} & b_{1H}b_{2L} & b_{1H}b_{2H}
\end{pmatrix}
,\,\mathbf{{\Greekmath 0119} }=
\begin{pmatrix}
{\Greekmath 0119} _{LL} \\
{\Greekmath 0119} _{LH} \\
{\Greekmath 0119} _{HL} \\
{\Greekmath 0119} _{HH}
\end{pmatrix}
,\,\mathbf{{\Greekmath 0115} }=
\begin{pmatrix}
{\Greekmath 0115} _{1L} \\
{\Greekmath 0115} _{1H} \\
{\Greekmath 0115} _{2L} \\
\mathrm{E}\left( {\Greekmath 010C} _{i1}{\Greekmath 010C} _{i2}\right)
\end{pmatrix}
.
\end{equation*}
Note that $\mathrm{E}\left( {\Greekmath 010C} _{i1}{\Greekmath 010C} _{i2}\right) $ is identified by
Theorem \ref{thm: idenfitication_moment_beta_ext}, and $b_{jk_{j}}$ and $
{\Greekmath 0115} _{jk_{j}}$ are identified by Corollary \ref{core:marginal_dist}, and
matrix $\mathbf{B}$ is invertible given that $b_{1L}<b_{1H}$ and $
b_{2L}<b_{2H}$. (See Appendix \ref{sec:Proofs}). As a result, the joint
probabilities, $\mathbf{{\Greekmath 0119} },$ are identified.

\medskip

\begin{remark}
The argument in Example 4 is applicable for identification of the joint
distribution of $\left( {\Greekmath 010C}_{ij}, {\Greekmath 010C}_{i,j^\prime} \right)^\prime$ for $
j\neq j^\prime$ when $p > 2$ and $K = 2$.
\end{remark}


\section{Finite sample properties using Monte Carlo experiments\label
{sec:Monte-Carlo-Simulation}}

We examine the finite sample performance of the categorical coefficient
estimator proposed in Section \ref{sec:Estimation} by Monte Carlo
experiments.

\subsection{Data generating processes\label{subsec:dgp}}

We generate $y_{i}$ as
\begin{equation}
y_{i}={\Greekmath 010B} +x_{i}{\Greekmath 010C} _{i}+z_{i1}{\Greekmath 010D} _{1}+z_{i2}{\Greekmath 010D} _{2}+u_{i},
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for }i=1,2,...,n,  \label{eq:mc_dgp}
\end{equation}
with ${\Greekmath 010C} _{i}$ distributed as in (\ref{eq:category_dist}) with $K=2,$ and
the parameters ${\Greekmath 0119} ,{\Greekmath 010C} _{L}$ and ${\Greekmath 010C} _{H}$.\footnote{
A Monte Carlo experiment with $K=3$ is relegated to Section \ref
{subsec:mc_k3} in the online supplement.}

We draw ${\Greekmath 010C} _{i}$ for each individual $i$ independently by setting ${\Greekmath 010C}
_{i}={\Greekmath 010C} _{L}$ with probability ${\Greekmath 0119} $ and ${\Greekmath 010C} _{i}={\Greekmath 010C} _{H}$ with
probability $1-{\Greekmath 0119} $, through a sequence of independent Bernoulli draws. \
We consider two sets of parameters in all DGPs, denoted as \textit{high
variance} and \textit{low variance} parametrization, respectively,
\begin{equation}
\left( {\Greekmath 0119} ,{\Greekmath 010C} _{L},{\Greekmath 010C} _{H},\mathrm{E}\left( {\Greekmath 010C} _{i}\right) ,
\mathrm{var}\left( {\Greekmath 010C} _{i}\right) \right) =
\begin{cases}
\left( 0.5,1,2,1.5,0.25\right) & \left( high\,variance\right) \\
\left( 0.3,0.5,1.345,1.0915,0.15\right) & \left( low\,variance\right)
\end{cases}
.  \label{eq:mc_dgp_para}
\end{equation}
${\Greekmath 010C} _{H}/{\Greekmath 010C} _{L}=2$ for the \textit{high variance} parametrization,
and ${\Greekmath 010C} _{H}/{\Greekmath 010C} _{L} = 2.69$, for the \textit{low variance}
parametrization, which is motivated by the estimates in our empirical
illustration in Section \ref{sec:Empirical-Application}.\footnote{
The estimates for ${\Greekmath 010C}_H / {\Greekmath 010C}_L$ in our empirical analysis range from
1.50 to 2.79.} The values of E$({\Greekmath 010C} _{i})$ and $\mathrm{var}\left( {\Greekmath 010C}
_{i}\right) $ are obtained noting that E$({\Greekmath 010C} _{i})={\Greekmath 0119} {\Greekmath 010C} _{L}+(1-{\Greekmath 0119}
){\Greekmath 010C} _{H}$, and $\mathrm{var}\left( {\Greekmath 010C} _{i}\right) ={\Greekmath 0119} (1-{\Greekmath 0119} )({\Greekmath 010C}
_{H}-{\Greekmath 010C} _{L})^{2}$. The remaining parameters are set as ${\Greekmath 010B} =0.25$,
and $\mathbf{{\Greekmath 010D} }=\left( 1,1\right) ^{\prime },$ across DGPs.

We generate the regressors and the error terms as follows.

\medskip \textbf{DGP 1 (Baseline)} We first generate $\tilde{x}_{i}\sim
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}(2)$, and then set $x_{i}=(\tilde{x}_{i}-2)/2$ so that $
x_{i}$ has $0$ mean and unit variance. The additional regressors, $z_{ij}$,
for $j=1,2$ with homogeneous slopes are generated as
\begin{equation*}
z_{i1}=x_{i}+v_{i1}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }z_{i2}=z_{i1}+v_{i2},
\end{equation*}
with $v_{ij}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID }N\left( 0,1\right) $, for $j=1,2$. This ensures
that the regressors are sufficiently correlated. The error term, $u_{i}$, is
generated as $u_{i}={\Greekmath 011B} _{i}{\Greekmath 0122} _{i}$, where ${\Greekmath 011B} _{i}^{2}$
are generated as $0.5(1+\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}(1))$, and ${\Greekmath 0122} _{i}\sim
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}N(0,1)$. Note that ${\Greekmath 0122} _{i}$ and ${\Greekmath 011B} _{i}^{2}$ are
generated independently, and $E(u_{i}^{2})=1$.

\medskip \textbf{DGP 2 (Categorical $x$)} This setup deviates from the
baseline DGP, and allows the distribution of $x_{i}$ to differ across $i$.
Accordingly, we generate $x_{i}=\left( \tilde{x}_{1i}-2\right) /2$ where $
\tilde{x}_{1i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}\left( 2\right) $ for $i=1,2,\cdots
,\lfloor n/2\rfloor $, and $x_{i}=\left( \tilde{x}_{2i}-2\right) /4$ where $
\tilde{x}_{2i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}\left( 4\right) $, for $i=\lfloor
n/2\rfloor +1,\cdots ,n$. The additional regressors, $z_{ij}$, for $j=1,2$
with homogeneous slopes are generated as
\begin{equation*}
z_{i1}=x_{i}+v_{i1}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }z_{i2}=z_{i1}+v_{i2},
\end{equation*}
with $v_{ij}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID }N\left( 0,1\right) $, for $j=1,2$. The error
term $u_{i}$ is generated the same as in DGP 1.

\medskip \textbf{DGP 3 (Categorical $u$)} We generate $x_{i}$ and $\mathbf{z}
_{i}$ the same as in DGP 1, but allow the error term $u_{i}$ to have a
heterogeneous distribution over $i$. For $i=1,2,\cdots ,\lfloor n/2\rfloor $
, we set $u_{i}={\Greekmath 011B} _{i}{\Greekmath 0122} _{i},$ where ${\Greekmath 011B} _{i}^{2}\sim
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F} ^{2}\left( 2\right) $ and ${\Greekmath 0122} _{i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}
N(0,1)$, and for $i=\lfloor n/2\rfloor +1,\cdots ,n$, we set $u_{i}=\left(
\tilde{u}_{i}-2\right) /2$, where $\tilde{u}_{i}\sim \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{IID}{\Greekmath 011F}
^{2}\left( 2\right) $.

\medskip We investigate the finite sample performance of the estimator
proposed in Section \ref{sec:Estimation} across DGP 1 to 3 with \textit{low
variance} and \textit{high variance} scenarios.\footnote{
We can consider a DGP with conditional heteroskedasticity, in which we
follow the baseline DGP and generate the error term as $u_{i}=x_{i}
{\Greekmath 0122} _{i}$, where ${\Greekmath 0122} _{i}\sim N(0,1)$. The least square
estimator for $\mathbf{{\Greekmath 011E} }$ is valid in this setup in terms of estimation
and inference, whereas the GMM estimator for the distributional parameters $
\mathbf{{\Greekmath 0112} }$ breaks down, which is to be expected since we can only
identify the first moment of ${\Greekmath 010C} _{i}$ under conditional
heteroskedasticity. The results are available on request.} Details of the
computational algorithm used to carry out the Monte Carlo experiments (and
the empirical results that follow) are given in Section \ref{sec:computation}
of the online supplement. An accompanying R package is available at
\url{https://github.com/zhan-gao/ccrm}.

\subsection{Summary of the MC results}

\begin{table}[tbp]
\caption{Bias, RMSE and size of the least square estimator $\hat{\mathbf{
\protect{\Greekmath 011E}}}$}
\label{tab:mc_bias_rmse_gamma}
\begin{center}
{\small
\begin{tabular}{rrrrrrrrrrr}
\hline
\multicolumn{2}{r|}{DGP} & \multicolumn{3}{c|}{Baseline} &
\multicolumn{3}{c|}{Categorical $x$} & \multicolumn{3}{c}{Categorical $u$}
\\ \hline
\multicolumn{2}{r|}{Sample size $n$} & \multicolumn{1}{c}{Bias} &
\multicolumn{1}{c}{RMSE} & \multicolumn{1}{c|}{Size} & \multicolumn{1}{c}{
Bias} & \multicolumn{1}{c}{RMSE} & \multicolumn{1}{c|}{Size} &
\multicolumn{1}{c}{Bias} & \multicolumn{1}{c}{RMSE} & \multicolumn{1}{c}{Size
} \\ \hline
\multicolumn{11}{c}{\textit{high variance}: $\mathrm{var}\left( {\Greekmath 010C}_i
\right) = 0.25$} \\ \hline
\multirow{7}{*}{\begin{turn}{90} $\mathrm{E}\left({\Greekmath 010C}_i\right) = 1.5$
\end{turn}} & \multicolumn{1}{r|}{100} & -0.0024 & 0.2035 &
\multicolumn{1}{r|}{0.0966} & -0.0037 & 0.2035 & \multicolumn{1}{r|}{0.0858}
& -0.0042 & 0.2268 & 0.0920 \\
& \multicolumn{1}{r|}{1,000} & -0.0017 & 0.0669 & \multicolumn{1}{r|}{0.0568}
& -0.0002 & 0.0657 & \multicolumn{1}{r|}{0.0540} & -0.0019 & 0.0738 & 0.0540
\\
& \multicolumn{1}{r|}{2,000} & -0.0008 & 0.0463 & \multicolumn{1}{r|}{0.0512}
& -0.0015 & 0.0475 & \multicolumn{1}{r|}{0.0534} & -0.0010 & 0.0523 & 0.0522
\\
& \multicolumn{1}{r|}{5,000} & -0.0004 & 0.0301 & \multicolumn{1}{r|}{0.0540}
& -0.0008 & 0.0300 & \multicolumn{1}{r|}{0.0546} & -0.0007 & 0.0335 & 0.0560
\\
& \multicolumn{1}{r|}{10,000} & 0.0002 & 0.0214 & \multicolumn{1}{r|}{0.0508}
& 0.0000 & 0.0212 & \multicolumn{1}{r|}{0.0510} & 0.0000 & 0.0229 & 0.0456
\\
& \multicolumn{1}{r|}{100,000} & -0.0001 & 0.0066 & \multicolumn{1}{r|}{
0.0472} & 0.0000 & 0.0066 & \multicolumn{1}{r|}{0.0460} & 0.0000 & 0.0075 &
0.0506 \\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 010D}_1 = 1$ \end{turn}} &
\multicolumn{1}{r|}{100} & -0.0022 & 0.1571 & \multicolumn{1}{r|}{0.0604} &
-0.0006 & 0.1598 & \multicolumn{1}{r|}{0.0666} & 0.0018 & 0.1912 & 0.0656 \\
& \multicolumn{1}{r|}{1,000} & 0.0004 & 0.0501 & \multicolumn{1}{r|}{0.0496}
& -0.0005 & 0.0496 & \multicolumn{1}{r|}{0.0508} & 0.0000 & 0.0600 & 0.0530
\\
& \multicolumn{1}{r|}{2,000} & 0.0003 & 0.0352 & \multicolumn{1}{r|}{0.0530}
& -0.0004 & 0.0350 & \multicolumn{1}{r|}{0.0544} & 0.0002 & 0.0432 & 0.0602
\\
& \multicolumn{1}{r|}{5,000} & -0.0001 & 0.0222 & \multicolumn{1}{r|}{0.0470}
& 0.0005 & 0.0225 & \multicolumn{1}{r|}{0.0548} & 0.0007 & 0.0267 & 0.0522
\\
& \multicolumn{1}{r|}{10,000} & -0.0004 & 0.0157 & \multicolumn{1}{r|}{0.0470
} & 0.0002 & 0.0157 & \multicolumn{1}{r|}{0.0512} & 0.0000 & 0.0188 & 0.0504
\\
& \multicolumn{1}{r|}{100,000} & -0.0001 & 0.0049 & \multicolumn{1}{r|}{
0.0494} & 0.0000 & 0.0049 & \multicolumn{1}{r|}{0.0468} & 0.0000 & 0.0059 &
0.0500 \\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 010D}_2 = 1$ \end{turn}} &
\multicolumn{1}{r|}{100} & 0.0011 & 0.1115 & \multicolumn{1}{r|}{0.0616} &
0.0016 & 0.1121 & \multicolumn{1}{r|}{0.0654} & -0.0002 & 0.1364 & 0.0700 \\
& \multicolumn{1}{r|}{1,000} & -0.0003 & 0.0358 & \multicolumn{1}{r|}{0.0558}
& 0.0001 & 0.0354 & \multicolumn{1}{r|}{0.0550} & 0.0006 & 0.0421 & 0.0508
\\
& \multicolumn{1}{r|}{2,000} & -0.0001 & 0.0253 & \multicolumn{1}{r|}{0.0522}
& 0.0006 & 0.0246 & \multicolumn{1}{r|}{0.0502} & -0.0003 & 0.0302 & 0.0560
\\
& \multicolumn{1}{r|}{5,000} & 0.0000 & 0.0158 & \multicolumn{1}{r|}{0.0480}
& 0.0000 & 0.0159 & \multicolumn{1}{r|}{0.0570} & -0.0003 & 0.0185 & 0.0470
\\
& \multicolumn{1}{r|}{10,000} & 0.0002 & 0.0111 & \multicolumn{1}{r|}{0.0494}
& -0.0002 & 0.0111 & \multicolumn{1}{r|}{0.0530} & -0.0001 & 0.0134 & 0.0522
\\
& \multicolumn{1}{r|}{100,000} & 0.0001 & 0.0035 & \multicolumn{1}{r|}{0.0488
} & 0.0000 & 0.0034 & \multicolumn{1}{r|}{0.0446} & 0.0000 & 0.0042 & 0.0496
\\ \hline
\multicolumn{11}{c}{\textit{low variance}: $\mathrm{var}\left( {\Greekmath 010C}_i
\right) = 0.15$} \\ \hline
\multirow{7}{*}{\begin{turn}{90} $\mathrm{E}\left({\Greekmath 010C}_i\right) = 1.0915$
\end{turn}} & \multicolumn{1}{r|}{100} & -0.0006 & 0.1829 &
\multicolumn{1}{r|}{0.0810} & -0.0023 & 0.1855 & \multicolumn{1}{r|}{0.0766}
& -0.0025 & 0.2094 & 0.0828 \\
& \multicolumn{1}{r|}{1,000} & -0.0005 & 0.0597 & \multicolumn{1}{r|}{0.0610}
& 0.0005 & 0.0590 & \multicolumn{1}{r|}{0.0478} & -0.0006 & 0.0670 & 0.0542
\\
& \multicolumn{1}{r|}{2,000} & -0.0002 & 0.0408 & \multicolumn{1}{r|}{0.0516}
& -0.0007 & 0.0427 & \multicolumn{1}{r|}{0.0606} & -0.0004 & 0.0475 & 0.0544
\\
& \multicolumn{1}{r|}{5,000} & -0.0002 & 0.0264 & \multicolumn{1}{r|}{0.0530}
& -0.0006 & 0.0266 & \multicolumn{1}{r|}{0.0480} & -0.0005 & 0.0302 & 0.0538
\\
& \multicolumn{1}{r|}{10,000} & 0.0000 & 0.0189 & \multicolumn{1}{r|}{0.0546}
& -0.0002 & 0.0188 & \multicolumn{1}{r|}{0.0486} & -0.0002 & 0.0208 & 0.0482
\\
& \multicolumn{1}{r|}{100,000} & -0.0001 & 0.0059 & \multicolumn{1}{r|}{
0.0474} & 0.0000 & 0.0059 & \multicolumn{1}{r|}{0.0494} & 0.0000 & 0.0068 &
0.0508 \\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 010D}_1 = 1$ \end{turn}} &
\multicolumn{1}{r|}{100} & -0.0027 & 0.1521 & \multicolumn{1}{r|}{0.0614} &
-0.0001 & 0.1538 & \multicolumn{1}{r|}{0.0622} & 0.0014 & 0.1847 & 0.0624 \\
& \multicolumn{1}{r|}{1,000} & 0.0001 & 0.0480 & \multicolumn{1}{r|}{0.0520}
& -0.0007 & 0.0481 & \multicolumn{1}{r|}{0.0542} & -0.0003 & 0.0584 & 0.0570
\\
& \multicolumn{1}{r|}{2,000} & 0.0002 & 0.0338 & \multicolumn{1}{r|}{0.0514}
& -0.0006 & 0.0334 & \multicolumn{1}{r|}{0.0512} & 0.0001 & 0.0417 & 0.0572
\\
& \multicolumn{1}{r|}{5,000} & -0.0002 & 0.0213 & \multicolumn{1}{r|}{0.0474}
& 0.0003 & 0.0216 & \multicolumn{1}{r|}{0.0532} & 0.0007 & 0.0257 & 0.0498
\\
& \multicolumn{1}{r|}{10,000} & -0.0003 & 0.0150 & \multicolumn{1}{r|}{0.0466
} & 0.0002 & 0.0152 & \multicolumn{1}{r|}{0.0542} & 0.0001 & 0.0183 & 0.0518
\\
& \multicolumn{1}{r|}{100,000} & -0.0001 & 0.0047 & \multicolumn{1}{r|}{
0.0482} & 0.0000 & 0.0047 & \multicolumn{1}{r|}{0.0474} & 0.0000 & 0.0057 &
0.0500 \\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 010D}_2 = 1$ \end{turn}} &
\multicolumn{1}{r|}{100} & 0.0011 & 0.1081 & \multicolumn{1}{r|}{0.0592} &
0.0013 & 0.1079 & \multicolumn{1}{r|}{0.0622} & -0.0002 & 0.1323 & 0.0674 \\
& \multicolumn{1}{r|}{1,000} & -0.0003 & 0.0345 & \multicolumn{1}{r|}{0.0594}
& 0.0003 & 0.0342 & \multicolumn{1}{r|}{0.0556} & 0.0006 & 0.0409 & 0.0500
\\
& \multicolumn{1}{r|}{2,000} & 0.0000 & 0.0243 & \multicolumn{1}{r|}{0.0534}
& 0.0006 & 0.0235 & \multicolumn{1}{r|}{0.0450} & -0.0001 & 0.0292 & 0.0576
\\
& \multicolumn{1}{r|}{5,000} & 0.0001 & 0.0152 & \multicolumn{1}{r|}{0.0490}
& 0.0001 & 0.0152 & \multicolumn{1}{r|}{0.0552} & -0.0002 & 0.0179 & 0.0470
\\
& \multicolumn{1}{r|}{10,000} & 0.0002 & 0.0106 & \multicolumn{1}{r|}{0.0454}
& -0.0002 & 0.0107 & \multicolumn{1}{r|}{0.0528} & -0.0002 & 0.0131 & 0.0526
\\
& \multicolumn{1}{r|}{100,000} & 0.0001 & 0.0033 & \multicolumn{1}{r|}{0.0442
} & 0.0000 & 0.0033 & \multicolumn{1}{r|}{0.0448} & 0.0000 & 0.0040 & 0.0486
\\ \hline
\end{tabular}
}
\end{center}
\par
{\footnotesize \textit{Notes:} The data generating process is
\eqref{eq:mc_dgp}. \textit{high variance} and \textit{low variance}
parametrization are described in \eqref{eq:mc_dgp_para}. ``Baseline'',
``Categorical $x$'' and ``Categorical $u$'' refer to DGP 1 to 3 as in
Section \ref{subsec:dgp}. Generically, bias, RMSE and size are calculated by
$R^{-1}\sum_{r=1}^R \left( \hat{ {\Greekmath 0112}}^{(r)} - {\Greekmath 0112}_0 \right)$, $\sqrt{
R^{-1}\sum_{r= 1}^R \left( \hat{{\Greekmath 0112}}^{(r)} -{\Greekmath 0112}_0 \right)^2}$, and $
R^{-1}\sum_{r=1}^R \mathbf{1}\left[ \left\vert \hat{{\Greekmath 0112}}^{(r)} - {\Greekmath 0112}_0
\right\vert / \hat{{\Greekmath 011B}}_{\hat{{\Greekmath 0112}}}^{(r)} > \mathrm{cv}_{0.05} \right]
$, respectively, for true parameter ${\Greekmath 0112}_0$, its estimate $\hat{{\Greekmath 0112}}
^{(r)}$, the estimated standard error of $\hat{{\Greekmath 0112}}^{(r)}$, $\hat{{\Greekmath 011B}}
_{\hat{{\Greekmath 0112}}}^{(r)}$, and the critical value $\mathrm{cv}_{0.05} =
\Phi^{-1}\left( 0.975 \right)$ across $R = 5,000$ replications, where $
\Phi\left( \cdot \right)$ is the cumulative distribution function of
standard normal distribution.}
\end{table}

\begin{figure}[tbp]
\caption{Empirical power functions for the least square estimator $\hat{
\mathbf{\protect{\Greekmath 011E}}}$ with the \textit{high variance} parametrization ($
\mathrm{var}\left( \protect{\Greekmath 010C}_i \right) = 0.25$)}
\label{fig:power_function_gamma_high}
\begin{center}
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Baseline}
            \includegraphics[width=\textwidth, height=1.9in]{fig1.png}
            \label{fig:power_function_gamma_baseline_high}
        \end{subfigure}\\[0pt]
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Categorical $x$}
            \includegraphics[width=\textwidth, height=1.9in]{fig2.png}
            \label{fig:power_function_gamma_cat_x_high}
        \end{subfigure}\\[0pt]
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Categorical $u$}
            \includegraphics[width=\textwidth, height=1.9in]{fig3.png}
            \label{fig:power_function_gamma_cat_u_high}
        \end{subfigure}
\end{center}
\par
{\footnotesize \textit{Notes:} The data generating process is
\eqref{eq:mc_dgp} with \textit{high variance} parametrization that is
described in \eqref{eq:mc_dgp_para}. ``Baseline'', ``Categorical $x$'' and
``Categorical $u$'' refer to DGP 1 to 3 as in Section \ref{subsec:dgp}.
Generically, power is calculated by $R^{-1}\sum_{r=1}^R \mathbf{1}\left[
\left\vert \hat{{\Greekmath 0112}}^{(r)} - {\Greekmath 0112}_{\Greekmath 010E} \right\vert / \hat{{\Greekmath 011B}}_{
\hat{{\Greekmath 0112}}}^{(r)} > \mathrm{cv}_{0.05} \right] $, for ${\Greekmath 0112}_{\Greekmath 010E}$ in a
symmetric neighborhood of the true parameter ${\Greekmath 0112}_0$, the estimate $\hat{
{\Greekmath 0112}}^{(r)}$, the estimated standard error of $\hat{{\Greekmath 0112}}^{(r)}$, $\hat{
{\Greekmath 011B}}_{\hat{{\Greekmath 0112}}}^{(r)}$, and the critical value $\mathrm{cv}_{0.05} =
\Phi^{-1}\left( 0.975 \right)$ across $R = 5,000$ replications, where $
\Phi\left( \cdot \right)$ is the cumulative distribution function of
standard normal distribution. }
\end{figure}

\begin{figure}[tbp]
\caption{Empirical power functions for the least square estimator $\hat{
\mathbf{\protect{\Greekmath 011E}}}$ with the \textit{low variance} parametrization ($
\mathrm{var}\left( \protect{\Greekmath 010C}_i \right) = 0.15$)}
\label{fig:power_function_gamma_low}
\begin{center}
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Baseline}
            \includegraphics[width=\textwidth, height=1.9in]{fig4.png}
            \label{fig:power_function_gamma_baseline_low}
        \end{subfigure}\\[0pt]
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Categorical $x$}
            \includegraphics[width=\textwidth, height=1.9in]{fig5.png}
            \label{fig:power_function_gamma_cat_x_low}
        \end{subfigure}\\[0pt]
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Categorical $u$}
            \includegraphics[width=\textwidth, height=1.9in]{fig6.png}
            \label{fig:power_function_gamma_cat_u_low}
        \end{subfigure}
\end{center}
\par
{\footnotesize \textit{Notes:} The data generating process is
\eqref{eq:mc_dgp} with \textit{low variance} parametrization that is
described in \eqref{eq:mc_dgp_para}. ``Baseline'', ``Categorical $x$'' and
``Categorical $u$'' refer to DGP 1 to 3 as in Section \ref{subsec:dgp}.
Generically, power is calculated by $R^{-1}\sum_{r=1}^R \mathbf{1}\left[
\left\vert \hat{{\Greekmath 0112}}^{(r)} - {\Greekmath 0112}_{\Greekmath 010E} \right\vert / \hat{{\Greekmath 011B}}_{
\hat{{\Greekmath 0112}}}^{(r)} > \mathrm{cv}_{0.05} \right] $, for ${\Greekmath 0112}_{\Greekmath 010E}$ in a
symmetric neighborhood of the true parameter ${\Greekmath 0112}_0$, the estimate $\hat{
{\Greekmath 0112}}^{(r)}$, the estimated standard error of $\hat{{\Greekmath 0112}}^{(r)}$, $\hat{
{\Greekmath 011B}}_{\hat{{\Greekmath 0112}}}^{(r)}$, and the critical value $\mathrm{cv}_{0.05} =
\Phi^{-1}\left( 0.975 \right)$ across $R = 5,000$ replications, where $
\Phi\left( \cdot \right)$ is the cumulative distribution function of
standard normal distribution. }
\end{figure}

\begin{table}[tbp]
\caption{Bias, RMSE and size of the GMM estimator for distributional
parameters of $\protect{\Greekmath 010C}$ }
\label{tab:mc_S4}
\begin{center}
{\small \
\begin{tabular}{rrrrrrrrrrr}
\hline
\multicolumn{2}{r|}{DGP} & \multicolumn{3}{c|}{Baseline} &
\multicolumn{3}{c|}{Categorical $x$} & \multicolumn{3}{c}{Categorical $u$}
\\ \hline
\multicolumn{2}{r|}{Sample size $n$} & \multicolumn{1}{c}{Bias} &
\multicolumn{1}{c}{RMSE} & \multicolumn{1}{c|}{Size} & \multicolumn{1}{c}{
Bias} & \multicolumn{1}{c}{RMSE} & \multicolumn{1}{c|}{Size} &
\multicolumn{1}{c}{Bias} & \multicolumn{1}{c}{RMSE} & \multicolumn{1}{c}{Size
} \\ \hline
\multicolumn{11}{c}{\textit{high variance}: $\mathrm{var}\left( {\Greekmath 010C}_i
\right) = 0.25$} \\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 0119}=0.5$ \end{turn}} & \multicolumn{1}{r|}{
100} & 0.0457 & 0.2291 & \multicolumn{1}{r|}{0.1737} & 0.0363 & 0.2410 &
\multicolumn{1}{r|}{0.2130} & 0.0235 & 0.2361 & 0.2231 \\
& \multicolumn{1}{r|}{1,000} & 0.0018 & 0.1019 & \multicolumn{1}{r|}{0.1308}
& 0.0033 & 0.1178 & \multicolumn{1}{r|}{0.1437} & -0.0270 & 0.1741 & 0.2033
\\
& \multicolumn{1}{r|}{2,000} & 0.0017 & 0.0688 & \multicolumn{1}{r|}{0.1084}
& 0.0015 & 0.0826 & \multicolumn{1}{r|}{0.1199} & -0.0174 & 0.1273 & 0.1545
\\
& \multicolumn{1}{r|}{5,000} & -0.0003 & 0.0416 & \multicolumn{1}{r|}{0.0936}
& -0.0015 & 0.0495 & \multicolumn{1}{r|}{0.0908} & -0.0089 & 0.0810 & 0.1048
\\
& \multicolumn{1}{r|}{10,000} & 0.0002 & 0.0301 & \multicolumn{1}{r|}{0.0774}
& -0.0006 & 0.0351 & \multicolumn{1}{r|}{0.0780} & -0.0052 & 0.0582 & 0.0864
\\
& \multicolumn{1}{r|}{100,000} & -0.0001 & 0.0096 & \multicolumn{1}{r|}{
0.0550} & 0.0002 & 0.0114 & \multicolumn{1}{r|}{0.0576} & -0.0009 & 0.0194 &
0.0582 \\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 010C}_L = 1$ \end{turn}} &
\multicolumn{1}{r|}{100} & 0.1415 & 0.4749 & \multicolumn{1}{r|}{0.2472} &
0.1099 & 0.5110 & \multicolumn{1}{r|}{0.2138} & 0.1151 & 0.5961 & 0.1820 \\
& \multicolumn{1}{r|}{1,000} & 0.0207 & 0.1242 & \multicolumn{1}{r|}{0.1501}
& 0.0200 & 0.1454 & \multicolumn{1}{r|}{0.1433} & -0.0256 & 0.2373 & 0.1225
\\
& \multicolumn{1}{r|}{2,000} & 0.0129 & 0.0819 & \multicolumn{1}{r|}{0.1344}
& 0.0116 & 0.1007 & \multicolumn{1}{r|}{0.1355} & -0.0094 & 0.1486 & 0.1094
\\
& \multicolumn{1}{r|}{5,000} & 0.0048 & 0.0512 & \multicolumn{1}{r|}{0.1052}
& 0.0027 & 0.0607 & \multicolumn{1}{r|}{0.1000} & -0.0053 & 0.0897 & 0.0850
\\
& \multicolumn{1}{r|}{10,000} & 0.0031 & 0.0365 & \multicolumn{1}{r|}{0.0854}
& 0.0021 & 0.0428 & \multicolumn{1}{r|}{0.0900} & -0.0020 & 0.0633 & 0.0714
\\
& \multicolumn{1}{r|}{100,000} & 0.0002 & 0.0112 & \multicolumn{1}{r|}{0.0534
} & 0.0007 & 0.0135 & \multicolumn{1}{r|}{0.0584} & -0.0002 & 0.0207 & 0.0574
\\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 010C}_H = 2$ \end{turn}} &
\multicolumn{1}{r|}{100} & -0.0996 & 0.5609 & \multicolumn{1}{r|}{0.2014} &
-0.0873 & 0.6154 & \multicolumn{1}{r|}{0.1963} & -0.1071 & 0.6996 & 0.1866
\\
& \multicolumn{1}{r|}{1,000} & -0.0193 & 0.1407 & \multicolumn{1}{r|}{0.1864}
& -0.0128 & 0.1581 & \multicolumn{1}{r|}{0.1661} & -0.0319 & 0.2400 & 0.2093
\\
& \multicolumn{1}{r|}{2,000} & -0.0099 & 0.0893 & \multicolumn{1}{r|}{0.1486}
& -0.0099 & 0.1094 & \multicolumn{1}{r|}{0.1467} & -0.0239 & 0.1663 & 0.1673
\\
& \multicolumn{1}{r|}{5,000} & -0.0053 & 0.0519 & \multicolumn{1}{r|}{0.1092}
& -0.0072 & 0.0622 & \multicolumn{1}{r|}{0.1082} & -0.0127 & 0.1019 & 0.1156
\\
& \multicolumn{1}{r|}{10,000} & -0.0020 & 0.0362 & \multicolumn{1}{r|}{0.0878
} & -0.0033 & 0.0430 & \multicolumn{1}{r|}{0.0880} & -0.0080 & 0.0718 &
0.0986 \\
& \multicolumn{1}{r|}{100,000} & -0.0005 & 0.0114 & \multicolumn{1}{r|}{
0.0530} & -0.0003 & 0.0134 & \multicolumn{1}{r|}{0.0548} & -0.0017 & 0.0236
& 0.0646 \\ \hline
\multicolumn{11}{c}{\textit{low variance}: $\mathrm{var}\left( {\Greekmath 010C}_i
\right) = 0.15$} \\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 0119}=0.3$ \end{turn}} & \multicolumn{1}{r|}{
100} & 0.2175 & 0.3084 & \multicolumn{1}{r|}{0.2183} & 0.2227 & 0.3187 &
\multicolumn{1}{r|}{0.2464} & 0.2294 & 0.3157 & 0.2500 \\
& \multicolumn{1}{r|}{1,000} & 0.0170 & 0.1536 & \multicolumn{1}{r|}{0.1873}
& 0.0307 & 0.1837 & \multicolumn{1}{r|}{0.2063} & 0.0511 & 0.2295 & 0.2493
\\
& \multicolumn{1}{r|}{2,000} & 0.0014 & 0.1010 & \multicolumn{1}{r|}{0.1426}
& 0.0105 & 0.1290 & \multicolumn{1}{r|}{0.1601} & 0.0181 & 0.1815 & 0.2102
\\
& \multicolumn{1}{r|}{5,000} & -0.0002 & 0.0590 & \multicolumn{1}{r|}{0.1084}
& 0.0010 & 0.0737 & \multicolumn{1}{r|}{0.1158} & 0.0085 & 0.1232 & 0.1468
\\
& \multicolumn{1}{r|}{10,000} & -0.0001 & 0.0415 & \multicolumn{1}{r|}{0.0894
} & 0.0005 & 0.0515 & \multicolumn{1}{r|}{0.0928} & 0.0067 & 0.0906 & 0.1046
\\
& \multicolumn{1}{r|}{100,000} & -0.0001 & 0.0129 & \multicolumn{1}{r|}{
0.0594} & 0.0003 & 0.0158 & \multicolumn{1}{r|}{0.0536} & 0.0108 & 0.0349 &
0.0776 \\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 010C}_L = 0.5$ \end{turn}} &
\multicolumn{1}{r|}{100} & 0.3365 & 0.5905 & \multicolumn{1}{r|}{0.2426} &
0.3153 & 0.6042 & \multicolumn{1}{r|}{0.2432} & 0.3384 & 0.6746 & 0.2005 \\
& \multicolumn{1}{r|}{1,000} & 0.0352 & 0.2334 & \multicolumn{1}{r|}{0.1560}
& 0.0290 & 0.2813 & \multicolumn{1}{r|}{0.1544} & 0.0131 & 0.4141 & 0.1233
\\
& \multicolumn{1}{r|}{2,000} & 0.0175 & 0.1414 & \multicolumn{1}{r|}{0.1310}
& 0.0131 & 0.1835 & \multicolumn{1}{r|}{0.1382} & -0.0157 & 0.2988 & 0.1037
\\
& \multicolumn{1}{r|}{5,000} & 0.0085 & 0.0830 & \multicolumn{1}{r|}{0.1082}
& 0.0041 & 0.1052 & \multicolumn{1}{r|}{0.1118} & -0.0057 & 0.1798 & 0.0928
\\
& \multicolumn{1}{r|}{10,000} & 0.0055 & 0.0577 & \multicolumn{1}{r|}{0.0966}
& 0.0031 & 0.0730 & \multicolumn{1}{r|}{0.0934} & 0.0019 & 0.1231 & 0.0760
\\
& \multicolumn{1}{r|}{100,000} & 0.0005 & 0.0180 & \multicolumn{1}{r|}{0.0596
} & 0.0011 & 0.0222 & \multicolumn{1}{r|}{0.0582} & 0.0130 & 0.0443 & 0.0962
\\ \hline
\multirow{7}{*}{\begin{turn}{90} ${\Greekmath 010C}_H = 1.345$ \end{turn}} &
\multicolumn{1}{r|}{100} & 0.0023 & 0.4727 & \multicolumn{1}{r|}{0.1377} &
0.0238 & 0.5290 & \multicolumn{1}{r|}{0.1453} & 0.0185 & 0.6500 & 0.1461 \\
& \multicolumn{1}{r|}{1,000} & -0.0081 & 0.1265 & \multicolumn{1}{r|}{0.1737}
& 0.0042 & 0.1621 & \multicolumn{1}{r|}{0.1655} & 0.0120 & 0.2353 & 0.1738
\\
& \multicolumn{1}{r|}{2,000} & -0.0092 & 0.0828 & \multicolumn{1}{r|}{0.1428}
& -0.0026 & 0.1045 & \multicolumn{1}{r|}{0.1475} & 0.0029 & 0.1607 & 0.1710
\\
& \multicolumn{1}{r|}{5,000} & -0.0048 & 0.0489 & \multicolumn{1}{r|}{0.1028}
& -0.0041 & 0.0586 & \multicolumn{1}{r|}{0.1034} & 0.0006 & 0.0970 & 0.1172
\\
& \multicolumn{1}{r|}{10,000} & -0.0025 & 0.0340 & \multicolumn{1}{r|}{0.0808
} & -0.0024 & 0.0412 & \multicolumn{1}{r|}{0.0942} & 0.0019 & 0.0706 & 0.0958
\\
& \multicolumn{1}{r|}{100,000} & -0.0004 & 0.0105 & \multicolumn{1}{r|}{
0.0486} & -0.0002 & 0.0125 & \multicolumn{1}{r|}{0.0548} & 0.0073 & 0.0262 &
0.0696 \\ \hline
\end{tabular}
}
\end{center}
\par
{\footnotesize \textit{Notes:} The data generating process is
\eqref{eq:mc_dgp}. \textit{high variance} and \textit{low variance}
parametrization are described in \eqref{eq:mc_dgp_para}. ``Baseline'',
``Categorical $x$'' and ``Categorical $u$'' refer to DGP 1 to 3 as in
Section \ref{subsec:dgp}. Generically, bias, RMSE and size are calculated by
$R^{-1}\sum_{r=1}^R \left( \hat{ {\Greekmath 0112}}^{(r)} - {\Greekmath 0112}_0 \right)$, $\sqrt{
R^{-1}\sum_{r= 1}^R \left( \hat{{\Greekmath 0112}}^{(r)} -{\Greekmath 0112}_0 \right)^2}$, and $
R^{-1}\sum_{r=1}^R \mathbf{1}\left[ \left\vert \hat{{\Greekmath 0112}}^{(r)} - {\Greekmath 0112}_0
\right\vert / \hat{{\Greekmath 011B}}_{\hat{{\Greekmath 0112}}}^{(r)} > \mathrm{cv}_{0.05} \right]
$, respectively, for true parameter ${\Greekmath 0112}_0$, its estimate $\hat{{\Greekmath 0112}}
^{(r)}$, the estimated standard error of $\hat{{\Greekmath 0112}}^{(r)}$, $\hat{{\Greekmath 011B}}
_{\hat{{\Greekmath 0112}}}^{(r)}$, and the critical value $\mathrm{cv}_{0.05} =
\Phi^{-1}\left( 0.975 \right)$ across $R = 5,000$ replications, where $
\Phi\left( \cdot \right)$ is the cumulative distribution function of
standard normal distribution. }
\end{table}

\begin{figure}[tbp]
\caption{Empirical power functions for the GMM estimator of distributional
parameters of $\protect{\Greekmath 010C}$ with the \textit{high variance}
parametrization($\mathrm{var}\left( \protect{\Greekmath 010C}_i \right) = 0.25$)}
\label{fig:power_function_S4_high}
\begin{center}
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Baseline}
            \includegraphics[width=\textwidth, height=1.9in]{fig7.png}
            \label{fig:power_function_S4_baseline_high}
        \end{subfigure}\\[0pt]
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Categorical $x$}
            \includegraphics[width=\textwidth, height=1.9in]{fig8.png}
            \label{fig:power_function_S4_cat_x_high}
        \end{subfigure}\\[0pt]
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Categorical $u$}
            \includegraphics[width=\textwidth, height=1.9in]{fig9.png}
            \label{fig:power_function_S4_cat_u_high}
        \end{subfigure}
\end{center}
\par
{\footnotesize \textit{Notes:} The data generating process is
\eqref{eq:mc_dgp} with \textit{high variance} parametrization that is
described in \eqref{eq:mc_dgp_para}. ``Baseline'', ``Categorical $x$'' and
``Categorical $u$'' refer to DGP 1 to 3 as in Section \ref{subsec:dgp}. The
model is estimated with $S = 4$, the highest order of moments of $x_i$ used
in estimation. Generically, power is calculated by $R^{-1}\sum_{r=1}^R
\mathbf{1}\left[ \left\vert \hat{{\Greekmath 0112}}^{(r)} - {\Greekmath 0112}_{\Greekmath 010E} \right\vert /
\hat{{\Greekmath 011B}}_{\hat{{\Greekmath 0112}}}^{(r)} > \mathrm{cv}_{0.05} \right] $, for $
{\Greekmath 0112}_{\Greekmath 010E}$ in a symmetric neighborhood of the true parameter ${\Greekmath 0112}_0$,
the estimate $\hat{{\Greekmath 0112}}^{(r)}$, the estimated standard error of $\hat{
{\Greekmath 0112}}^{(r)}$, $\hat{{\Greekmath 011B}}_{\hat{{\Greekmath 0112}}}^{(r)}$, and the critical value $
\mathrm{cv}_{0.05} = \Phi^{-1}\left( 0.975 \right)$ across $R = 5,000$
replications, where $\Phi\left( \cdot \right)$ is the cumulative
distribution function of standard normal distribution. }
\end{figure}

\begin{figure}[tbp]
\caption{Empirical power functions for the GMM estimator of distributional
parameters of $\protect{\Greekmath 010C}$ with the \textit{low variance} parametrization
($\mathrm{var}\left( \protect{\Greekmath 010C}_i \right) = 0.15$)}
\label{fig:power_function_S4_low}
\begin{center}
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Baseline}
            \includegraphics[width=\textwidth, height=1.9in]{fig10.png}
            \label{fig:power_function_S4_baseline_low}
        \end{subfigure}\\[0pt]
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Categorical $x$}
            \includegraphics[width=\textwidth, height=1.9in]{fig11.png}
            \label{fig:power_function_S4_cat_x_low}
        \end{subfigure}\\[0pt]
\begin{subfigure}[b]{\textwidth}
            \centering
            \caption{Categorical $u$}
            \includegraphics[width=\textwidth, height=1.9in]{fig12.png}
            \label{fig:power_function_S4_cat_u_low}
        \end{subfigure}
\end{center}
\par
{\footnotesize \textit{Notes:} The data generating process is
\eqref{eq:mc_dgp} with \textit{low variance} parametrization that is
described in \eqref{eq:mc_dgp_para}. ``Baseline'', ``Categorical $x$'' and
``Categorical $u$'' refer to DGP 1 to 3 as in Section \ref{subsec:dgp}. The
model is estimated with $S = 4$, the highest order of moments of $x_i$ used
in estimation. Generically, power is calculated by $R^{-1}\sum_{r=1}^R
\mathbf{1}\left[ \left\vert \hat{{\Greekmath 0112}}^{(r)} - {\Greekmath 0112}_{\Greekmath 010E} \right\vert /
\hat{{\Greekmath 011B}}_{\hat{{\Greekmath 0112}}}^{(r)} > \mathrm{cv}_{0.05} \right] $, for $
{\Greekmath 0112}_{\Greekmath 010E}$ in a symmetric neighborhood of the true parameter ${\Greekmath 0112}_0$,
the estimate $\hat{{\Greekmath 0112}}^{(r)}$, the estimated standard error of $\hat{
{\Greekmath 0112}}^{(r)}$, $\hat{{\Greekmath 011B}}_{\hat{{\Greekmath 0112}}}^{(r)}$, and the critical value $
\mathrm{cv}_{0.05} = \Phi^{-1}\left( 0.975 \right)$ across $R = 5,000$
replications, where $\Phi\left( \cdot \right)$ is the cumulative
distribution function of standard normal distribution. }
\end{figure}

For each sample size $n = 100$, $1,000$, $2,000$, $5,000$, $10,000$ and $
100,000$ we run $5,000$ replications of experiments for DGP 1 (baseline),
DGP 2 (categorical $x$) and DGP 3 (categorical $u$) with \textit{high
variance} and \textit{low variance} parametrization, as set out in
\eqref{eq:mc_dgp_para}.

We first investigate the finite sample performance of $\hat{\mathbf{{\Greekmath 011E} }}$
, as an estimator of $\mathbf{{\Greekmath 011E} }=\left( \mathrm{E}\left( {\Greekmath 010C}
_{i}\right) ,\mathbf{{\Greekmath 010D} }^{\prime }\right) ^{\prime }$. Bias, root mean
squared errors (RMSE) for estimation of $\mathrm{E}\left( {\Greekmath 010C} _{i}\right) $
, ${\Greekmath 010D} _{1}$ and ${\Greekmath 010D} _{2}$, as well as size of testing of the null
values at the 5 per cent nominal value are reported in Table \ref
{tab:mc_bias_rmse_gamma}. In addition, we plot the associated empirical
power functions in Figure \ref{fig:power_function_gamma_high} and \ref
{fig:power_function_gamma_low}, for cases of high and low $var({\Greekmath 010C} _{i})$.
The results show that $\hat{\mathbf{{\Greekmath 011E} }}$ has very good small sample
properties with small bias and RMSEs, with size very close to the nominal
value of 5 per cent across all DGPs and parametrization, even when sample
size is relatively small. The power of the test increases steadily as the
sample size increases.

Then, we turn to the GMM estimator for the distributional parameters of $
{\Greekmath 010C} _{i}$ proposed in Section \ref{subsec:estimation_beta}. The bias,
RMSE, and the test size based on the asymptotic distribution given in
Theorem \ref{thm:normality}, for ${\Greekmath 0119} $, ${\Greekmath 010C} _{L}$ and ${\Greekmath 010C} _{H}$, are
reported in Table \ref{tab:mc_S4}. The empirical power functions are
reported in Figure \ref{fig:power_function_S4_high} and \ref
{fig:power_function_S4_low}. The reported results are based on $S=4$, where $
S$ $(>2K-1=3)$ denotes the highest order of moments of $x_{i}$ included in
estimation.\footnote{
We also tried estimation based on a larger number of moments (using $S=5$
and $S=6$). In the case of current Monte Carlo results, adding more moments
does not seem to add much to the precision of the estimates and could be
counter-productive when $n$ is not sufficiently large. The results are
available in Section \ref{subsec:mc_s5_s6} in the online supplement.}

The upper panel of this table reports the results of the high variance and
the lower panel for the low variance parametrization, as set out in (\ref
{eq:mc_dgp_para}). For all parameters and under all DGPs, the bias and RMSE
decline steadily with the sample size as predicted by Theorem \ref
{thm:consisteny}, and confirm the robustness of the GMM estimates to the
heterogeneity in the regressor and the error processes. But for a given
sample size, the relative precision of the estimates depends on the
variability of ${\Greekmath 010C} _{i}$, as characterized by the true value of $\mathrm{
var} ({\Greekmath 010C} _{i})$. The precision of the estimates with \textit{high variance
} parametrization is relatively higher than that with \textit{low variance}
parametrization. This is to be expected since, unlike $\mathrm{E} ({\Greekmath 010C}
_{i}),$ the distributional parameters are only identified if $\mathrm{var}
({\Greekmath 010C} _{i})>0$. As shown in \eqref{sys_3} and \eqref{sys_4} for the current
case of $K=2$, $\mathrm{var}({\Greekmath 010C} _{i})$ is in the denominator when we
recover the distributional parameters from the moments of ${\Greekmath 010C} _{i}$. When
$\mathrm{var}({\Greekmath 010C} _{i})$ is small, estimation errors in the moments of $
{\Greekmath 010C} _{i}$ can be amplified in the estimation of ${\Greekmath 0119} $, ${\Greekmath 010C} _{L}$ and $
{\Greekmath 010C} _{H}$. On the other hand, the larger the variance the more precisely $
{\Greekmath 0119} $, ${\Greekmath 010C} _{H}$ and ${\Greekmath 010C} _{L}$ can be estimated for a given $n$.
\footnote{
Section \ref{subsec:mc_highvar} in the online supplement presents
parametrization with $\mathrm{var}\left( {\Greekmath 010C}_i \right) = 6.35$ and $18.95$
, which further confirms the pattern that the larger the variance the more
precisely ${\Greekmath 0119} $, ${\Greekmath 010C} _{H}$ and ${\Greekmath 010C} _{L}$ can be estimated for a given
$n$.} The size and power also depends on the parametrization. With both
\textit{high variance} and \textit{low variance} parametrization, we can
achieve correct size and reasonable power when $n$ is quite large ($
n=100,000 $). We plot the empirical power functions for $n\geq 5,000$ for $
{\Greekmath 0119} $, ${\Greekmath 010C} _{H}$ and ${\Greekmath 010C} _{L}$ since the size is far above 5 per cent
for smaller values of $n$, and power comparisons are not meaningful in such
cases.

\begin{remark}
\label{rem:practical} Note that GMM estimators of moments of ${\Greekmath 010C} _{i}$,
namely $\mathbf{m}_{\mathbf{{\Greekmath 010C} }}$, can be obtained using the moment
conditions in \eqref{eq:mc_limit},and the transformations $\mathbf{m}_{
\mathbf{{\Greekmath 010C} }}=h\left( \mathbf{{\Greekmath 0112} }\right) $ in \eqref{eq:mbeta_mat}
are required only to derive the estimators of $\mathbf{{\Greekmath 0112} }$, the
parameters of the underlying categorical distribution. The Monte Carlo
results in Section \ref{subsec:mc_gmm_moments} in the online supplement show
that $\mathbf{m}_{\mathbf{{\Greekmath 010C} }}$ can be accurately estimated with
relatively small sample sizes. In the estimation of both $\mathbf{m}_{
\mathbf{{\Greekmath 010C} }}$ and $\mathbf{{\Greekmath 0112} }$, the same set of moment conditions
are included, so the estimation of distributional parameters $\mathbf{{\Greekmath 0112}
}$ essentially relies on the relation $\mathbf{{\Greekmath 0112} }=h^{-1}\left( \mathbf{
m}_{\mathbf{{\Greekmath 010C} }}\right) $. Sampling uncertainties in the estimation of $
\mathbf{m}_{\mathbf{{\Greekmath 010C} }}$, particularly in higher order moments, are
potentially amplified through the inverse transformation $h^{-1}$ that
involves matrix inversion, which causes the difficulties in estimation and
inference of $\mathbf{{\Greekmath 0112} }$ when sample sizes are small. This is
analogous to the problem of precision matrix estimation from an estimated
covariance matrix. In practice, estimation of the categorical parameters is
recommended for applications where the sample size is relatively large,
otherwise it is advisable to focus on estimates of the lower order moments
of ${\Greekmath 010C} _{i}$.
\end{remark}


\section{Heterogeneous return to education: An empirical application \label
{sec:Empirical-Application}}

Since the pioneering work by \citet{becker1962investment,becker1964book} on
the effects of investments in human capital, estimating returns to education
has been one of the focal points of labor economics research. In his
pioneering contribution \citet{mincer1974book} models the logarithm of
earnings as a function of years of education and years of potential labor
market experience (age minus years of education minus six), which can be
written in a generic form:
\begin{equation}
\log \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{wage}_{i}={\Greekmath 010B} _{i}+{\Greekmath 010C} _{i}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{edu}_{i}+{\Greekmath 011E} \left(
\mathbf{z}_{i}\right) +{\Greekmath 0122} _{i},  \label{eq:heckman2018eq1}
\end{equation}
as in \citet*[Equation (1)]{heckman2018returns}, where $\mathbf{z}_{i}$
includes the labor market experience and other relevant control variables.
The above wage equation, also known as the \textquotedblleft Mincer
equation\textquotedblright, has become of the workhorse of the empirical
works on estimating the return to education. In the most widely used
specification of the Mincer equation \eqref{eq:heckman2018eq1},
\begin{equation*}
{\Greekmath 011E} \left( \mathbf{z}_{i}\right) ={\Greekmath 011A} _{1}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{exper}_{i}+{\Greekmath 011A} _{2}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
exper}_{i}^{2}+\tilde{\mathbf{z}}_{i}^{\prime }\tilde{\mathbf{{\Greekmath 010D} }},
\end{equation*}
where $\tilde{\mathbf{z}}_{i}$ is the vector of control variables other than
potential labor market experience.

Along with the advancement of empirical research on this topic, there has
been a growing awareness of the importance of heterogeneity in individual
cognitive and non-cognitive abilities \citep{heckman2001jpenobellecture} and
their significance for explaining the observed heterogeneity in return to
education. Accordingly, it is important to allow the parameters of the wage
equation to differ across individuals. In equation \eqref{eq:heckman2018eq1}
we allow ${\Greekmath 010B} _{i}$ and ${\Greekmath 010C} _{i}$ to differ across individuals, but
assume that ${\Greekmath 011E} \left( \mathbf{z}_{i}\right) $ can be approximated as
non-linear functions of experience and other control variables with
homogeneous coefficients.

Specifically, following \citet{lemieux2006postnber,lemieux2006postsecondary}
we also allow for time variations in the parameters of the wage equation and
consider the following categorical coefficient model over a given
cross-section sample indexed by $t$:\footnote{
Some investigators have suggested including higher powers of the experience
variable in the wage equation. \citet{lemieux2006mincerbookchapter}, for
example, proposes using a quartic rather than a quadratic function. As a
robustness check we also provide estimation results with quartic experience
specification in Section \ref{sec:empirical_supp} in the online supplement.}
\begin{equation}
\log \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{wage}_{it}={\Greekmath 010B} _{it}+{\Greekmath 010C} _{it}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{edu}_{it}+{\Greekmath 011A} _{1t}
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{exper}_{it}+{\Greekmath 011A} _{2t}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{exper}_{it}^{2}+\tilde{\mathbf{z}}
_{it}^{\prime }\tilde{\mathbf{{\Greekmath 010D} }}_{t}+{\Greekmath 0122} _{it},
\label{eq:ccrm_spec_alpha_i}
\end{equation}
where the return to education follows the categorical distribution,
\begin{equation*}
{\Greekmath 010C} _{it}=
\begin{cases}
b_{tL} & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{w.p. }{\Greekmath 0119} _{t}, \\
b_{tH} & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{w.p. }1-{\Greekmath 0119} _{t},
\end{cases}
\end{equation*}
and $\tilde{\mathbf{z}}_{it}$ includes gender, martial status and race. $
{\Greekmath 010B} _{it}={\Greekmath 010B} _{t}+{\Greekmath 010E} _{it}$ where ${\Greekmath 010E} _{it}$ is mean $0$
random variable assumed to be distributed independently of $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{edu}_{it}$
and $\mathbf{z}_{it}=\left( \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{exper}_{it},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{exper}_{it}^{2}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, }
\tilde{\mathbf{z}}_{t}^{\prime }\right) ^{\prime }$. Let $u_{it}={\Greekmath 0122}
_{it}+{\Greekmath 010E} _{it}$, and write \eqref{eq:ccrm_spec_alpha_i} as
\begin{equation}
\log \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{wage}_{it}={\Greekmath 010B} _{t}+{\Greekmath 010C} _{it}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{edu}_{it}+{\Greekmath 011A} _{1t}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
exper}_{it}+{\Greekmath 011A} _{2t}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{exper}_{it}^{2}+\tilde{\mathbf{z}}_{it}^{\prime }
\tilde{\mathbf{{\Greekmath 010D} }}_{t}+u_{it}.  \label{eq:ccrm_spec}
\end{equation}
The correlation between ${\Greekmath 010B} _{it}$ and $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{edu}_{it}$ in
\eqref{eq:heckman2018eq1} is the source of \textquotedblleft ability
bias\textquotedblright\ \citep{griliches1977estimating}. Given the pure
cross-sectional nature of our analysis, we do not allow for the endogeneity
from \textquotedblleft ability bias\textquotedblright\ or dynamics. To allow
for non-zero correlations between ${\Greekmath 010B} _{it}$, edu$_{it}$ and $\mathbf{z}
_{it}$, a panel data approach is required, which has its own challenges, as
education and experience variables tend to very slow moving (if at all) for
many individuals in the panel. Time delays between changes in education and
experience, and the wage outcomes also further complicate the interpretation
of the mean estimates of ${\Greekmath 010C} _{it}$ which we shall be reporting. To
partially address the possible dynamic spillover effects, we provide
estimates of the distribution of ${\Greekmath 010C} _{it}$ using cross-sectional data
from two different sample periods, and investigate the extent to which the
distribution of return to education has changed over time, by gender and the
level of educational achievements.\footnote{
Time variations in return to education has also been investigated in the
literature as a possible explanation of increasing wage inequality in the
U.S. See, for example, the papers by
\citet{lemieux2006postnber,lemieux2006postsecondary}.}

We estimate the categorical distribution of the return to education in
\eqref{eq:ccrm_spec} using the May and Outgoing Rotation Group (ORG)
supplements of the Current Population Survey (CPS) data, as in
\citet{lemieux2006postnber,lemieux2006postsecondary}.\footnote{
The data is retrieved from
\url{https://www.openicpsr.org/openicpsr/project/116216/version/V1/view}.}
We pool observations from 1973 to 1975 for the first sample period, $
t=\left\{ 1973-1975\right\} $ and observations from 2001 to 2003 for the
second sample period, $t=\left\{ 2001-2003\right\} $. Following
\citet{lemieux2006postnber}, we consider sub-samples of those with less than
12 years of education, \textquotedblleft high school or less", and those
with more than 12 years of education, \textquotedblleft postsecondary
education", as well as the combined sample. We also present results by
gender. The summary statistics are reported in Table \ref
{tab:Descriptive_Statistics}. As to be expected, the mean log wages are
higher for those with postsecondary education (for male and female), with
the number of years of schooling and experience rising by about one year
across the two sub-period samples. There are also important differences
across male and female, and the two educational groupings, which we hope to
capture in our estimation.

\begin{table}[htbp]
\caption{Summary Statistics of the May and Outgoing Rotation Group (ORG)
supplements of the Current Population Survey (CPS) data across two periods,
1973 - 75 and 2001 - 03, by years of education and gender}
\label{tab:Descriptive_Statistics}\vspace{-10pt}
\par
\begin{center}
{\small
\begin{tabular}{rccccccc}
\hline
& \multicolumn{3}{c}{1973 - 75} &  & \multicolumn{3}{c}{2001 - 03} \\
\cline{2-4}\cline{6-8}
& \multicolumn{1}{c}{High School} & Postsecondary & \multirow{2}{*}{All} &
& High School & Postsecondary & \multirow{2}{*}{All} \\
& \multicolumn{1}{c}{or Less} & Education &  &  & or Less & Education &  \\
\hline
& \multicolumn{7}{c}{\textit{Both male and female}} \\ \hline
\texttt{log wage} & 1.59 & 1.94 & 1.69 &  & 1.47 & 1.88 & 1.71 \\
& {\footnotesize \textit{(0.50)}} & {\footnotesize \textit{(0.53)}} &
{\footnotesize \textit{(0.53)}} &  & {\footnotesize \textit{(0.47)}} &
{\footnotesize \textit{(0.57)}} & {\footnotesize \textit{(0.57)}} \\
\texttt{edu.} & 10.64 & 15.21 & 12.02 &  & 11.29 & 14.96 & 13.41 \\
& {\footnotesize \textit{(2.11)}} & {\footnotesize \textit{(1.65)}} &
{\footnotesize \textit{(2.89)}} &  & {\footnotesize \textit{(1.68)}} &
{\footnotesize \textit{(1.82)}} & {\footnotesize \textit{(2.53)}} \\
\texttt{age} & 36.74 & 34.90 & 36.18 &  & 37.96 & 39.87 & 39.06 \\
& {\footnotesize \textit{(13.85)}} & {\footnotesize \textit{(11.58)}} & (
{\footnotesize \textit{{13.23})}} &  & {\footnotesize \textit{(12.93)}} &
{\footnotesize \textit{(11.33)}} & {\footnotesize \textit{(12.07)}} \\
\texttt{expr.} & 20.10 & 13.69 & 18.17 &  & 20.67 & 18.91 & 19.65 \\
& {\footnotesize \textit{(14.44)}} & {\footnotesize \textit{(11.41)}} &
{\footnotesize \textit{(13.91)}} &  & {\footnotesize \textit{(12.95)}} &
{\footnotesize \textit{(11.17)}} & {\footnotesize \textit{(11.98)}} \\
\texttt{marriage} & 0.67 & 0.70 & 0.68 &  & 0.52 & 0.60 & 0.57 \\
& {\footnotesize \textit{(0.47)}} & {\footnotesize \textit{(0.46)}} &
{\footnotesize \textit{(0.47)}} &  & {\footnotesize \textit{(0.50)}} &
{\footnotesize \textit{(0.49)}} & {\footnotesize \textit{(0.50)}} \\
\texttt{nonwhite} & 0.11 & 0.08 & 0.10 &  & 0.15 & 0.14 & 0.15 \\
& {\footnotesize \textit{(0.32)}} & {\footnotesize \textit{(0.27)}} &
{\footnotesize \textit{(0.30)}} &  & {\footnotesize \textit{(0.36)}} &
{\footnotesize \textit{(0.35)}} & {\footnotesize \textit{(0.35)}} \\
$n$ & 77,899 & 33,733 & 111,632 &  & 216,136 & 295,683 & 511,819 \\ \hline
& \multicolumn{7}{c}{\textit{Male}} \\ \hline
\texttt{log wage} & 1.76 & 2.07 & 1.86 &  & 1.57 & 2.00 & 1.81 \\
& {\footnotesize \textit{(0.48)}} & {\footnotesize \textit{(0.53)}} &
{\footnotesize \textit{(0.52)}} &  & {\footnotesize \textit{(0.48)}} &
{\footnotesize \textit{(0.58)}} & {\footnotesize \textit{(0.58)}} \\
\texttt{edu.} & 10.44 & 15.29 & 12.00 &  & 11.19 & 15.02 & 13.31 \\
& {\footnotesize \textit{(2.26)}} & {\footnotesize \textit{(1.69)}} &
{\footnotesize \textit{(3.08)}} &  & {\footnotesize \textit{(1.82)}} &
{\footnotesize \textit{(1.84)}} & {\footnotesize \textit{(2.64)}} \\
\texttt{age} & 36.79 & 35.29 & 36.31 &  & 37.21 & 40.24 & 38.89 \\
& {\footnotesize \textit{(13.82)}} & {\footnotesize \textit{(11.24)}} &
{\footnotesize \textit{(13.07)}} &  & {\footnotesize \textit{(12.70)}} &
{\footnotesize \textit{(11.30)}} & {\footnotesize \textit{(12.04)}} \\
\texttt{expr.} & 20.35 & 14.00 & 18.32 &  & 20.02 & 19.22 & 19.58 \\
& {\footnotesize \textit{(14.49)}} & {\footnotesize \textit{(11.06)}} &
{\footnotesize \textit{(13.81)}} &  & {\footnotesize \textit{(12.75)}} &
{\footnotesize \textit{(11.08)}} & {\footnotesize \textit{(11.86)}} \\
\texttt{marriage} & 0.73 & 0.76 & 0.74 &  & 0.53 & 0.64 & 0.59 \\
& {\footnotesize \textit{(0.44)}} & {\footnotesize \textit{(0.43)}} &
{\footnotesize \textit{(0.44)}} &  & {\footnotesize \textit{(0.50)}} &
{\footnotesize \textit{(0.48)}} & {\footnotesize \textit{(0.49)}} \\
\texttt{nonwhite} & 0.10 & 0.06 & 0.09 &  & 0.14 & 0.13 & 0.13 \\
& {\footnotesize \textit{(0.30)}} & {\footnotesize \textit{(0.24)}} &
{\footnotesize \textit{(0.29)}} &  & {\footnotesize \textit{(0.34)}} &
{\footnotesize \textit{(0.33)}} & {\footnotesize \textit{(0.34)}} \\
$n$ & 44,299 & 20,851 & 65,150 &  & 116,129 & 144,138 & 260,267 \\ \hline
& \multicolumn{7}{c}{\textit{Female}} \\ \hline
\texttt{log wage} & 1.35 & 1.71 & 1.45 &  & 1.77 & 1.36 & 1.61 \\
& {\footnotesize \textit{(0.41)}} & {\footnotesize \textit{(0.47)}} &
{\footnotesize \textit{(0.46)}} &  & {\footnotesize \textit{(0.54)}} &
{\footnotesize \textit{(0.43)}} & {\footnotesize \textit{(0.54)}} \\
\texttt{edu.} & 10.89 & 15.08 & 12.05 &  & 14.90 & 11.42 & 13.52 \\
& {\footnotesize \textit{(1.87)}} & {\footnotesize \textit{(1.59)}} &
{\footnotesize \textit{(2.60)}} &  & {\footnotesize \textit{(1.79)}} &
{\footnotesize \textit{(1.49)}} & {\footnotesize \textit{(2.40)}} \\
\texttt{age} & 36.67 & 34.27 & 36.01 &  & 38.83 & 39.52 & 39.24 \\
& {\footnotesize \textit{(13.88)}} & {\footnotesize \textit{(12.09)}} &
{\footnotesize \textit{(13.45)}} &  & {\footnotesize \textit{(13.14)}} &
{\footnotesize \textit{(11.35)}} & {\footnotesize \textit{(12.10)}} \\
\texttt{expr.} & 19.78 & 13.19 & 17.96 &  & 18.61 & 21.41 & 19.73 \\
& {\footnotesize \textit{(14.36)}} & {\footnotesize \textit{(11.94)}} &
{\footnotesize \textit{(14.04)}} &  & {\footnotesize \textit{(11.24)}} &
{\footnotesize \textit{(13.13)}} & {\footnotesize \textit{(12.11)}} \\
\texttt{marriage} & 0.60 & 0.60 & 0.60 &  & 0.56 & 0.51 & 0.54 \\
& {\footnotesize \textit{(0.49)}} & {\footnotesize \textit{(0.49)}} &
{\footnotesize \textit{(0.49)}} &  & {\footnotesize \textit{(0.50)}} &
{\footnotesize \textit{(0.50)}} & {\footnotesize \textit{(0.50)}} \\
\texttt{nonwhite} & 0.13 & 0.10 & 0.12 &  & 0.15 & 0.17 & 0.16 \\
& {\footnotesize \textit{(0.33)}} & {\footnotesize \textit{(0.30)}} &
{\footnotesize \textit{(0.33)}} &  & {\footnotesize \textit{(0.36)}} &
{\footnotesize \textit{(0.38)}} & {\footnotesize \textit{(0.37)}} \\
$n$ & 33,600 & 12,882 & 46,482 &  & 151,545 & 100,007 & 251,552 \\ \hline
\end{tabular}
}
\end{center}
\par
{\footnotesize \textit{Notes:} ``Postsecondary Education'' stands for the
sub-sample with years of education higher than 12 and ``High School or
Less'' stands for sub-sample with years of education less than or equal to
12). \texttt{edu.} and \texttt{exper.} are in years. \texttt{marriage} and
\texttt{nonwhite} are dummy variables. $n$ is the sample size. We report
mean and standard deviation (in parentheses) of each variable. The data is
from the May and Outgoing Rotation Group (ORG) supplements of the Current
Population Survey (CPS) data retrived from
\url{https://www.openicpsr.org/openicpsr/project/116216/version/V1/view}.}
\end{table}

\begin{table}[htbp]
\caption{Estimates of the distribution of the return to education across two
periods, 1973 - 75 and 2001 - 03, by years of education and gender}
\label{tab:Return-to-Education}
\begin{center}
\begin{tabular}{rrrrrrrrr}
\hline
& \multicolumn{2}{c}{High School or Less} &  & \multicolumn{2}{c}{
Postsecondary Edu.} &  & \multicolumn{2}{c}{All} \\
\cline{2-3}\cline{5-6}\cline{8-9}
& \multicolumn{1}{c}{1973 - 75} & \multicolumn{1}{c}{2001 - 03} &  &
\multicolumn{1}{c}{1973 - 75} & \multicolumn{1}{c}{2001 - 03} &  &
\multicolumn{1}{c}{1973 - 75} & \multicolumn{1}{c}{2001 - 03} \\ \hline
& \multicolumn{8}{c}{Both Male and Female} \\ \hline
${\Greekmath 0119}$ & 0.4843 & 0.5069 &  & 0.4398 & 0.3537 &  & 0.4719 & 0.3463 \\
& {\small \textit{(4188.8)}} & {\small \textit{(0.0269)}} &  & {\small
\textit{(0.0502)}} & {\small \textit{(0.0091)}} &  & {\small \textit{(0.0485)
}} & \textit{(0.0047)} \\
${\Greekmath 010C}_L$ & 0.0608 & 0.0382 &  & 0.0624 & 0.0866 &  & 0.0558 & 0.0645 \\
& {\small \textit{(5.0939)}} & {\small \textit{(0.0014)}} &  & {\small
\textit{(0.0035)}} & {\small \textit{(0.0009)}} &  & {\small \textit{(0.0020)
}} & {\small \textit{(0.0004)}} \\
${\Greekmath 010C}_H$ & 0.0619 & 0.0920 &  & 0.1103 & 0.1401 &  & 0.0941 & 0.1263 \\
& {\small \textit{(4.8132)}} & {\small \textit{(0.0019)}} &  & {\small
\textit{(0.0032)}} & {\small \textit{(0.0007)}} &  & {\small \textit{(0.0022)
}} & {\small \textit{(0.0004)}} \\
${\Greekmath 010C}_H / {\Greekmath 010C}_L$ & 1.0194 & 2.4102 &  & 1.7680 & 1.6178 &  & 1.6879 &
1.9567 \\
& {\small \textit{(6.2938)}} & {\small \textit{(0.0428)}} &  & {\small
\textit{(0.0618)}} & {\small \textit{(0.0111)}} &  & {\small \textit{(0.0295)
}} & {\small \textit{(0.0080)}} \\
$\mathrm{E}\left({\Greekmath 010C}_i\right)$ & 0.0614 & 0.0647 &  & 0.0893 & 0.1212 &  &
0.0760 & 0.1049 \\
$\mathrm{s.d.}\left({\Greekmath 010C}_i\right)$ & 0.0006 & 0.0269 &  & 0.0238 & 0.0256 &
& 0.0191 & 0.0294 \\
$n$ & 77,899 & 216,136 &  & 33,733 & 295,683 &  & 111,632 & 511,819 \\ \hline
& \multicolumn{8}{c}{Male} \\ \hline
${\Greekmath 0119}$ & n/a & 0.4939 &  & 0.4706 & 0.3201 &  & 0.4802 & 0.3290 \\
& n/a & {\small \textit{(0.0399)}} &  & {\small \textit{(0.0707)}} & {\small
\textit{(0.0104)}} &  & {\small \textit{(0.0815)}} & {\small \textit{(0.0053)
}} \\
${\Greekmath 010C}_L$ & 0.0637 & 0.0404 &  & 0.0534 & 0.0743 &  & 0.0536 & 0.0548 \\
& n/a & {\small \textit{(0.0019)}} &  & {\small \textit{(0.0046)}} & {\small
\textit{(0.0012)}} &  & {\small \textit{(0.0030)}} & {\small \textit{(0.0005)
}} \\
${\Greekmath 010C}_H$ & 0.0637 & 0.0911 &  & 0.0995 & 0.1308 &  & 0.0875 & 0.1192 \\
& n/a & {\small \textit{(0.0026)}} &  & {\small \textit{(0.0042)}} & {\small
\textit{(0.0009)}} &  & {\small \textit{(0.0031)}} & {\small \textit{(0.0005)
}} \\
${\Greekmath 010C}_H / {\Greekmath 010C}_L$ & 1.0000 & 2.2526 &  & 1.8641 & 1.7603 &  & 1.6312 &
2.1772 \\
& n/a & {\small \textit{(0.0534)}} &  & {\small \textit{(0.1038)}} & {\small
\textit{(0.0209)}} &  & {\small \textit{(0.0459)}} & {\small \textit{(0.0144)
}} \\
$\mathrm{E}\left({\Greekmath 010C}_i\right)$ & 0.0637 & 0.0661 &  & 0.0778 & 0.1128 &  &
0.0712 & 0.0980 \\
$\mathrm{s.d.}\left({\Greekmath 010C}_i\right)$ & 0.0000 & 0.0253 &  & 0.0230 & 0.0264 &
& 0.0169 & 0.0303 \\
$n$ & 44,299 & 116,129 &  & 20,851 & 144,138 &  & 65,150 & 260,267 \\ \hline
& \multicolumn{8}{c}{Female} \\ \hline
${\Greekmath 0119}$ & 0.4999 & 0.5166 &  & 0.4526 & 0.3906 &  & 0.4566 & 0.3608 \\
& {\small \textit{(0.5047)}} & {\small \textit{(0.0283)}} &  & {\small
\textit{(0.0829)}} & {\small \textit{(0.0167)}} &  & {\small \textit{(0.0810)
}} & {\small \textit{(0.0086)}} \\
${\Greekmath 010C}_L$ & 0.0441 & 0.0348 &  & 0.0823 & 0.0979 &  & 0.0628 & 0.0751 \\
& {\small \textit{(0.0133)}} & {\small \textit{(0.0016)}} &  & {\small
\textit{(0.0053)}} & {\small \textit{(0.0013)}} &  & {\small \textit{(0.0033)
}} & {\small \textit{(0.0007)}} \\
${\Greekmath 010C}_H$ & 0.0723 & 0.0972 &  & 0.1310 & 0.1473 &  & 0.1028 & 0.1333 \\
& {\small \textit{(0.0159)}} & {\small \textit{(0.0025)}} &  & {\small
\textit{(0.0055)}} & {\small \textit{(0.0011)}} &  & {\small \textit{(0.0038)
}} & {\small \textit{(0.0007)}} \\
${\Greekmath 010C}_H / {\Greekmath 010C}_L$ & 1.6392 & 2.7934 &  & 1.5913 & 1.5048 &  & 1.6357 &
1.7756 \\
& {\small \textit{(0.1565)}} & {\small \textit{(0.0700)}} &  & {\small
\textit{(0.0539)}} & {\small \textit{(0.0121)}} &  & {\small \textit{(0.0353)
}} & \textit{(0.0090)} \\
$\mathrm{E}\left({\Greekmath 010C}_i\right)$ & 0.0582 & 0.0650 &  & 0.1090 & 0.1280 &  &
0.0845 & 0.1123 \\
$\mathrm{s.d.}\left({\Greekmath 010C}_i\right)$ & 0.0141 & 0.0312 &  & 0.0242 & 0.0241 &
& 0.0199 & 0.0280 \\
$n$ & 33,600 & 100,007 &  & 12,882 & 151,545 &  & 46,482 & 251,552 \\ \hline
\end{tabular}
\end{center}
\par
{\footnotesize \textit{Notes:} This table reports the estimates of the
distribution of ${\Greekmath 010C} _{i}$ with the quadratic in experience specification
\eqref{eq:ccrm_spec_alpha_i}, using $S = 4$ order moments of $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{edu}_i$.
\textquotedblleft Postsecondary Edu.\textquotedblright\ stands for the
sub-sample with years of education higher than 12 and \textquotedblleft High
School or Less\textquotedblright\ stands for those with years of education
less than or equal to 12. $\mathrm{s.d.}\left({\Greekmath 010C} _{i}\right)$ corresponds
to the square root of estimated $\mathrm{var}\left( {\Greekmath 010C} _{i}\right) $. $n$
is the sample size. \textquotedblleft n/a" is inserted when the estimates
show homogeneity of ${\Greekmath 010C} _{i}$ and ${\Greekmath 0119} $ is not identified and cannot be
estimated.}
\end{table}

\begin{table}[p!]
\caption{Estimates of $\mathbf{\protect{\Greekmath 010D}}$ associated with control
variables $\mathbf{z}_i$ with specification \eqref{eq:ccrm_spec_alpha_i}
across two periods, 1973 - 75 and 2001 - 03, by years of education and
gender, which complements Table \protect\ref{tab:Return-to-Education}}
\label{tab:gamma_est_quadratic}
\begin{center}
\begin{tabular}{rrrrrrlrr}
\hline
& \multicolumn{2}{c}{High School or Less} & \multicolumn{1}{l}{} &
\multicolumn{2}{c}{Postsecondary Edu.} &  & \multicolumn{2}{c}{All} \\
\cline{2-3}\cline{5-6}\cline{8-9}
& \multicolumn{1}{c}{1973 - 75} & \multicolumn{1}{c}{2001 - 03} &
\multicolumn{1}{l}{} & \multicolumn{1}{c}{1973 - 75} & \multicolumn{1}{c}{
2001 - 03} &  & \multicolumn{1}{c}{1973 - 75} & \multicolumn{1}{c}{2001 - 03}
\\ \hline
& \multicolumn{8}{c}{\textit{Both male and female}} \\ \hline
\texttt{exper.} & 0.0305 & 0.0319 &  & 0.0415 & 0.0354 &  & 0.0310 & 0.0321
\\
& {\small \textit{(0.0004)}} & {\small \textit{(0.0002)}} &  & {\small
\textit{(0.0008)}} & {\small \textit{(0.0003)}} &  & {\small \textit{
(0.0003) }} & {\small \textit{(0.0002)}} \\
$\mathtt{exper.}^2$ ($\times 10^2$) & -0.0490 & -0.0505 &  & -0.0826 &
-0.0652 &  & -0.0499 & -0.0537 \\
& {\small \textit{(0.0009)}} & {\small \textit{(0.0005)}} &  & {\small
\textit{(0.0022)}} & {\small \textit{(0.0007)}} &  & {\small \textit{
(0.0008) }} & {\small \textit{(0.0005)}} \\
\texttt{marriage} & 0.1120 & 0.0751 &  & 0.0886 & 0.0770 &  & 0.1085 & 0.0818
\\
& {\small \textit{(0.0036)}} & {\small \textit{(0.0020)}} &  & {\small
\textit{(0.0059)}} & {\small \textit{(0.0020)}} &  & {\small \textit{
(0.0031) }} & {\small \textit{(0.0014)}} \\
\texttt{nonwhite} & -0.0922 & -0.0775 &  & -0.0424 & -0.0571 &  & -0.0715 &
-0.0667 \\
& {\small \textit{(0.0047)}} & {\small \textit{(0.0024)}} &  & {\small
\textit{(0.0088)}} & {\small \textit{(0.0025)}} &  & {\small \textit{
(0.0042) }} & {\small \textit{(0.0018)}} \\
\texttt{gender} & 0.4157 & 0.2298 &  & 0.2962 & 0.2023 &  & 0.3892 & 0.2167
\\
& {\small \textit{(0.0029)}} & {\small \textit{(0.0017)}} &  & {\small
\textit{(0.0050)}} & {\small \textit{(0.0018)}} &  & {\small \textit{
(0.0025) }} & {\small \textit{(0.0013)}} \\
$n$ & 77,899 & 216,136 &  & 33,733 & 295,683 &  & 111,632 & 511,819 \\ \hline
& \multicolumn{8}{c}{\textit{Male}} \\ \hline
\texttt{exper.} & 0.0369 & 0.0366 &  & 0.0516 & 0.0405 &  & 0.0389 & 0.0371
\\
& {\small \textit{(0.0005)}} & {\small \textit{(0.0003)}} &  & {\small
\textit{(0.0011)}} & {\small \textit{(0.0005)}} &  & {\small \textit{
(0.0005) }} & {\small \textit{(0.0003)}} \\
$\mathtt{exper.}^2$ ($\times 10^2$) & -0.0589 & -0.0589 &  & -0.1016 &
-0.0752 &  & -0.0635 & -0.0629 \\
& {\small \textit{(0.0012)}} & {\small \textit{(0.0008)}} &  & {\small
\textit{(0.0029)}} & {\small \textit{(0.0011)}} &  & {\small \textit{
(0.0010) }} & {\small \textit{(0.0007)}} \\
\texttt{marriage} & 0.1940 & 0.1123 &  & 0.1497 & 0.1344 &  & 0.1828 & 0.1316
\\
& {\small \textit{(0.0053)}} & {\small \textit{(0.0028)}} &  & {\small
\textit{(0.0085)}} & {\small \textit{(0.0031)}} &  & {\small \textit{
(0.0045) }} & {\small \textit{(0.0021)}} \\
\texttt{nonwhite} & -0.1241 & -0.1165 &  & -0.1172 & -0.1010 &  & -0.1178 &
-0.1093 \\
& {\small \textit{(0.0065)}} & {\small \textit{(0.0035)}} &  & {\small
\textit{(0.0127)}} & {\small \textit{(0.0039)}} &  & {\small \textit{
(0.0058) }} & {\small \textit{(0.0027)}} \\
$n$ & 44,299 & 116,129 &  & 20,851 & 144,138 &  & 65,150 & 260,267 \\ \hline
& \multicolumn{8}{c}{\textit{Female}} \\ \hline
\texttt{exper.} & 0.0223 & 0.0265 &  & 0.0271 & 0.0313 &  & 0.0208 & 0.0272
\\
& {\small \textit{(0.0006)}} & {\small \textit{(0.0003)}} &  & {\small
\textit{(0.0011)}} & {\small \textit{(0.0004)}} &  & {\small \textit{
(0.0005) }} & {\small \textit{(0.0003)}} \\
$\mathtt{exper.}^2$ ($\times 10^2$) & -0.0376 & -0.0411 &  & -0.0564 &
-0.0576 &  & -0.0338 & -0.0450 \\
& {\small \textit{(0.0013)}} & {\small \textit{(0.0008)}} &  & {\small
\textit{(0.0030)}} & {\small \textit{(0.0010)}} &  & {\small \textit{
(0.0012) }} & {\small \textit{(0.0006)}} \\
\texttt{marriage} & 0.0115 & 0.0317 &  & -0.0005 & 0.0262 &  & 0.0118 &
0.0322 \\
& {\small \textit{(0.0048)}} & {\small \textit{(0.0028)}} &  & {\small
\textit{(0.0079)}} & {\small \textit{(0.0026)}} &  & {\small \textit{
(0.0041) }} & {\small \textit{(0.0019)}} \\
\texttt{nonwhite} & -0.0581 & -0.0441 &  & 0.0395 & -0.0236 &  & -0.0202 &
-0.0315 \\
& {\small \textit{(0.0065)}} & {\small \textit{(0.0033)}} &  & {\small
\textit{(0.0117)}} & {\small \textit{(0.0033)}} &  & {\small \textit{
(0.0058) }} & {\small \textit{(0.0024)}} \\
$n$ & 33,600 & 100,007 &  & 12,882 & 151,545 &  & 46,482 & 251,552 \\ \hline
\end{tabular}
\end{center}
\par
{\footnotesize \textit{Notes:} This table reports the estimates of $\mathbf{
\ {\Greekmath 010D}}$ in \eqref{eq:ccrm_spec_alpha_i}. ``Postsecondary Edu.'' stands
for the sub-sample with years of education higher than 12 and ``High School
or Less'' stands for those with years of education less than or equal to 12.
The standard error of estimates of coefficients associated with control
variables are estimated based on Theorem \ref{lem:gamma_est_consistency} and
reported in parentheses. $n$ is the sample size.}
\end{table}

We treat the cross-section observations in the two sample periods, $
t=\left\{ 1973-1975\right\} $ and $\left\{ 2001-2003\right\} $, as \textit{
repeated} cross-sections, rather than a panel data since the data in these
two periods do not cover the same individuals, and represent random samples
from the population of wage earners in two periods. It should also be noted
that sample sizes $(n_{t})$, although quite large, are much larger during $
\left\{ 2001-2003\right\} $, which could be a factor when we come to compare
estimates from the two sample periods. For example, for both male and female
$n_{73-75}=111,632$ as compared to $n_{01-03}=511,819$, a difference which
becomes more pronounced when we consider the number observations in
postsecondary/female category - which rises from $12,882$ for the first
period to $100,007$ in the second period.

We report estimates of ${\Greekmath 0119}_{t}$, ${\Greekmath 010C} _{L,t}$ and ${\Greekmath 010C} _{H,t}$, as well
as corresponding mean and standard deviations (denoted by s.d.($\hat{{\Greekmath 010C}}
_{it}$)) of the return to education (${\Greekmath 010C}_{it}$) for $t=\left\{
1973-1975\right\} $ and $\left\{ 2001-2003\right\} $. For a given ${\Greekmath 0119} _{t}$
, the ratio ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ provides a measure of within group
heterogeneity and allows us to augment information on changes in mean with
changes in the distribution of return of education. The estimates for the
distribution of the return to education (${\Greekmath 010C} _{it}$) are summarized in
Table \ref{tab:Return-to-Education}, with the estimation results for control
variables (such as experience, experienced squared, and other individual
specific characteristic) reported in Table \ref{tab:gamma_est_quadratic}.

As can be seen from Table \ref{tab:Return-to-Education}, estimates of $
\mathrm{s.d.}\left( {\Greekmath 010C} _{it}\right) $ are strictly positive for all
sub-groups, except for the \textquotedblleft high school or
less\textquotedblright\ group during the first sample period. For this group
during the first period the estimate of $\mathrm{s.d.}\left( {\Greekmath 010C}
_{it}\right) $ for the male sub-sample is zero, ${\Greekmath 0119} $ is not identified,
and we have identical estimates for ${\Greekmath 010C} _{L}$ and ${\Greekmath 010C} _{H}$. For this
sub-sample, the associated estimates and their standard errors are shown as
unavailable ($n/a$). In case of the female sub-sample as well as both male
and female sub-samples where the estimates of s.d.($\hat{{\Greekmath 010C}}_{it}$) are
close to zero and ${\Greekmath 0119} $ is poorly estimated, only the mean of the return to
education is informative. In the case of the samples where the estimates of $
\mathrm{s.d.}\left( {\Greekmath 010C} _{it}\right) $ are strictly positive, the estimate
of the ratio ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ provides a good measure of within
group heterogeneity of return to education. The estimates of ${\Greekmath 010C}
_{H,t}/{\Greekmath 010C} _{L,t}$, lie between $1.50$ to $2.79$, with the high estimate
obtained for the females with high school or less education during $\left\{ {
2001-03}\right\} $, and the low estimate is obtained for females with
postsecondary education during the same period.

As our theory suggests the mean estimates of return to education, $E\left(
{\Greekmath 010C} _{it}\right) $, are very precisely estimated and inferences involving
them tend to be robust to conditional error heteroskedasticity. The results
in Table \ref{tab:Return-to-Education} show that estimates of $E\left( {\Greekmath 010C}
_{it}\right) $ have increased over the two sample periods $t=\left\{
1973-75\right\} $ to $t=\left\{ 2001-03\right\} $, regardless of gender or
educational grouping. The postsecondary educational group show larger
increases in the estimates of $E\left( {\Greekmath 010C} _{it}\right) $ as compared to
those with high school or less. Estimates of $\mathrm{E}\left( {\Greekmath 010C}
_{it}\right) $ increases by $36$ per cent for the postsecondary group while
the estimates of mean return to education rises only by around $5$ per cent
in the case of those with high school or less. This result holds for both
genders. Comparing the mean returns across the two educational groups, we
find that mean return to education of individuals with postsecondary
education is $45$ per cent higher than those with high school or less in the
$\{1973 - 75\}$ period, but this gap increases to $87$ per cent in the
second period, $\left\{ 2001-03\right\} $. Similar patterns are observed in
the sub-samples by gender. The estimates suggest rising between group
heterogeneity, which is mainly due to the increasing returns to education
for the postsecondary group.

Turning to within group heterogeneity, we focus on the estimates of ${\Greekmath 010C}
_{H,t}/{\Greekmath 010C} _{L,t}$ and first note that over the two periods, within group
heterogeneity has been rising mainly in the case of those with high school
or less, for both male and female. For the combined male and female samples
and the male sub-sample, there is little evidence of within group
heterogeneity for the first period $\left\{ 1973-75\right\} $. However, for
the second\ period $\left\{ 2001-03\right\} $ we find a sizeable degree of
within group heterogeneity where ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ is estimated to
be around $2.41$, with $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{s.d.}\left( {\Greekmath 010C} _{it}\right) \approx 0.03$.
For the female sub-sample with high school or less, little evidence of
heterogeneity was found for the first period, estimates of ${\Greekmath 010C}
_{H,t}/{\Greekmath 010C} _{L,t}$ increases to $2.79$ for the second sample period, that
corresponds to a commensurate rise in $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{s.d.}\left( {\Greekmath 010C} _{i}\right) $
to $0.032$. The pattern of within group heterogeneity is very different for
those with postsecondary educational. For this group we in fact observe a
slight decline in the estimates of ${\Greekmath 010C} _{H,t}/{\Greekmath 010C} _{L,t}$ by gender and
over two sample periods.

Overall, our between and within estimates of mean return to education are in
line with the evidence of rising wage inequality documented in the
literature \citep{corak2013income}.


\section{Conclusion\label{sec:Conclusion}}

In this paper we consider random coefficient models for repeated
cross-sections in which the random coefficients follow categorical
distributions. Identification is established using moments of the random
coefficients in terms of the moments of the underlying observations. We
propose two-step generalized method of moments to estimate the parameters of
the categorical distributions. The consistency and asymptotic normality of
the GMM estimators are established without the IID assumption typically
assumed in the literature. Small sample properties of the proposed estimator
are investigated by means of Monte Carlo experiments and shown to be robust
to heterogeneously generated regressors and errors, although relatively
large samples are required to estimate the parameters of the underling
categorical distributions. This is largely due to the highly non-linear
mapping between the parameters of the categorical distribution and the
higher order moments of the coefficients. This problem is likely to become
more pronounced with a larger number of categories and coefficients.

In the empirical application, we apply the model to study the evolution of
returns to education over two sub-periods, also considered in the literature
by \citet{lemieux2006postnber}. Our estimates show that mean (ex post)
returns to education have risen over the periods from 1973 - 75 to 2001 -
2003 mainly in the case of individuals with postsecondary education, and
this result is robust by gender. We find evidence of within group
heterogeneity in the case of high school or less educational group as
compared to those with postsecondary education.

In our model specification, the number of categories, $K$, is treated as a
tuning parameter and assumed to be known. An information criterion, as in
\citet{bonhomme_manersa2015fixedeffect} and \citet*{ssp2016classo}, to
determine $K$ could be considered. Further investigation of models with
multiple regressors subject to parameter heterogeneity is also required.
These and other related issues are topics for future research.

\newpage