EconBase
← Back to paper

Identification and estimation of multinomial choice models with latent special covariates

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

53,692 characters

Identification and estimation of multinomial choice models with latent special covariates



\title{Identification and estimation of multinomial choice models with latent special covariates\thanks{A previous version of this paper was circulated under the title ``Identification and estimation of discrete outcome models with latent special covariates.'' I thank the editor and two anonymous referees for comments and suggestions that have greatly improved the manuscript. I also thank Roy Allen, Victor H. Aguiar, Tim Conley, Nirav Mehta, Salvador Navarro, Joris Pinkse, David Rivers, and Bruno Salcedo for their useful comments and discussions.}}
\author{
	Nail Kashaev
	\thanks{Department of Economics, University of Western Ontario.}\\[email removed]
	}
\date{First version: November 13, 2018\\
 This version: March 21, 2022}

\maketitle

\begin{abstract}
Identification of multinomial choice models is often established by using special covariates that have full support. This paper shows how these identification results can be extended to a large class of multinomial choice models when all covariates are bounded. I also provide a new $\sqrt{n}$-consistent asymptotically normal estimator of the finite-dimensional parameters of the model.  \bigskip

\noindent JEL classification numbers: C50, C57
\bigskip

\noindent Keywords: Multinomial choice, random coefficients, special covariate, identification at infinity, bundles
\end{abstract}

\section{Introduction}
This paper studies identification and estimation of random coefficients multinomial choice models with covariates that have bounded support. Often \emph{some} latent variables in these models have full support (i.e. supported on the whole Euclidean space). Under common restrictions on the distribution of these unobservables, I constructively identify it and show how these latent variables can be used to construct special covariates (i.e., artificial observables with full support) to nonparametrically identify the distribution of \emph{all} the other unobservables. Identification of all parts of the structural model is crucial for welfare analysis (e.g., aggregate welfare changes between two choice situations).
My identification technique is constructive and leads to an asymptotically normal estimator of the finite-dimensional parameters of the model.
\par
The results of this paper rest on two commonly used assumptions. First, I assume existence of excluded covariates that affect the distribution over choices via a random coefficient. Using variation in these excluded covariates I can identify the distribution of the random coefficient. Second, I assume that the distribution of the random coefficient is sufficiently ``rich''. ``Richness'' of the random coefficient distribution is formalized by a notion of bounded completeness.\footnote{Completeness of a family of distributions is a well-known concept in the Statistics and Econometrics literature. See, for example, \citet{mattner1993some, newey03, chernozhukov2005iv, blundell07, chernozhukov2007instrumental, hu2008instrumental, andrews11, darolles11}, and \citet{d2011completeness}.}
As a result, I show how to identify the distribution over outcomes conditional on the realization of the observed covariates and the latent random coefficient nonparametrically. Since the latent random coefficient often has full support, I can treat it as an observed covariate with full support and apply \emph{any} identification technique that requires existence of such covariates to identify the rest of the model parameters (e.g., the distribution of other latent variables).
\par
I provide two nonnested identification results. The first result does not make any parametric assumptions about the distribution of latent variables. It, however, imposes some restrictions on the support of observables. In particular, I require the support of some covariates to contain zero. It also requires some smoothness of the distribution of the latent variables. To the best of my knowledge, this is the first result in the literature that nonparametrically identifies the distribution of all latent variables in multinomial choice settings with bounded covariates. The second result uses one of the most popular parameterizations in applied work - a Gaussian distribution of the latent random coefficient. But, in contrast to the first result, it does not require zero in the support of covariates and leaves the distribution of other latent variables completely unrestricted. The second result also leads to an easy to implement asymptotically normal estimator of the finite-dimensional parameters of the model. Similar to \citet{powell1989semiparametric}, this estimator is $\sqrt{n}$-consistent since it is based on average derivatives of an estimable object.
\par
I contribute to the discrete outcome literature in several respects. I show how existing results that use full-support-excluded covariates with monotonicity restrictions\footnote{See, for example, \citet{manski1985semiparametric, manski1988identification, heckman1990varieties, matzkin1992nonparametric, ichimura1998maximum, lewbel1998semiparametric, lewbel2000semiparametric, tamer03, matzkin2007heterogeneous, berry2009nonparametric, BHR, gautier2013nonparametric, gautier2015triangular, fox2016nonparametric, dunker2017nonparametric, fox2017note, fox2018unobserved, fox2020note}, and \citet{kashaev2020discerning}.} can be directly used in environments with bounded covariates. Formally, I demonstrate that my setting inherits all identifying properties of the setting with a special covariate. I also contribute to the literature on semiparametric models by showing that common parametric restrictions can be used instead of covariates that have full support (e.g., \citealp{fox2012random}). This paper is also related to the literature on identification of finite-dimensional parameters in discrete outcome models with bounded covariates.\footnote{E.g., \citet{magnac2007identification, chen2016informational, kline16}, and \citet{lewbel2021semiparametric}.} The main difference from that literature is that in my framework the distribution of latent variables (e.g., the random intercept) can be nonparametrically identified even if these latent variables have full support, but covariates are bounded.
\par
My approach is complementary to existing methods. Since as an input my framework requires the average structural demand function (i.e., the choice probability function) for one good, my results may be combined with the ones in \citet{berry2020nonparametric} to nonparametricaly identify the distribution of unobserved individual level heterogeneity. Moreover, in situations where the researcher is not sure whether covariates have full support and is willing to impose mild restrictions because of tractability or data limitations, my approach can provide an additional reassurance of identification. Also, the results in this paper provide a more solid econometric foundation to the models with at least one normally distributed random coefficient (e.g. \citealp{nevo2000practitioner}).
\par
The paper is organized as follows.
In Section~\ref{sec: model}, I describe the setting. Sections~\ref{sec:nonparam} and~\ref{sec:normal error} provide two identification results.
I show how my identification results can be extended to bundles model in Section~\ref{sub: bundles}. In Sections~\ref{sec: estimation} and~\ref{sec: simulations}, I propose a new estimator of the finite-dimensional parameters and evaluate its performance in simulations. Section~\ref{sec: empirical application} provides an empirical illustration. Section~\ref{sec:conclusion} concludes. All proofs can be found in Appendix~\ref{app: proofs}. Appendix~\ref{app: simulations} provides additional simulation evidence.


\section{Multinomial Choice}\label{sec: model}
Consider the following random coefficients model. The agent maximizes (indirect) utility by choosing between $J$ inside goods (e.g., different brands of cereals) and an outside option of no purchase. The choice set is denoted by $Y=\{0,1,\dots, J\}$. I normalize the utility from alternative $y=0$ to $0$. The random utility from choosing an alternative $y\neq 0$ is\footnote{Deterministic vectors are denoted by lower-case regular font Latin letters (e.g., $x$) and random objects by bold letters (e.g., $\mathbf{x}$). Capital letters are usually used to denote supports of random variables (e.g., $\mathbf{x}\in X$). I denote the support of a conditional distribution of $\mathbf{x}$ conditional on $\mathbf{z}=z$ by $X_z$.
The cumulative distribution function (c.d.f.) and the probability density function (p.d.f.) of $\mathbf{x}$ are denoted by $F_{\mathbf{x}}$ and $f_{\mathbf{x}}$. $F_{\mathbf{x}|\mathbf{z}}$ ($f_{\mathbf{x}|\mathbf{z}}$) denotes the c.d.f. (p.d.f.) of $\mathbf{x}$ conditional on $\mathbf{z}=z$.
}
\begin{align*}\label{eq:random coeff parametrization}
    &\mathbf{z}_{y}\big[\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}\big]+\boldsymbol{\varepsilon}_y,
\end{align*}
where $\mathbf{z}_y\in Z_y\subseteq{\mathds{R}}$ is a product-specific observed covariate that can be different for different consumers (e.g., fiber content or price); $\mathbf{d}\in D\subseteq{\mathds{R}}$ is observed (demographic) individual-specific taste shifter (e.g., age or income); $\mathbf{w}\in W\subseteq{\mathds{R}}^{d_w}$ is a vector of all other observable covariates, which may include the rest of product/agent characteristics; $\mathbf{e}\in E\subseteq{\mathds{R}}$ is a latent taste shock. The latent random vector $\boldsymbol{\varepsilon}=(\boldsymbol{\varepsilon}_{y})_{y\in Y\setminus\{0\}}$ captures all other sources of unobserved heterogeneity (e.g., $\boldsymbol{\varepsilon}_y=\boldsymbol{\theta}^{{\sf T}}\mathbf{w}_y+\boldsymbol{\epsilon}_y\:\ensuremath{\mathrm{a.s.}}$, where $\boldsymbol{\theta}$ and $\boldsymbol{\epsilon}_y$ are random coefficients). The observed covariates are $\mathbf{x}=(\mathbf{d}, \mathbf{z},\mathbf{w})$, where $\mathbf{z}=(\mathbf{z}_{y})_{y\in Y\setminus\{0\}}$.
\par
The random coefficient $\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}$ represents individual specific heterogeneous tastes associated with the product characteristic $\mathbf{z}_{y}$ (i.e., the marginal utility from the product characteristic $\mathbf{z}_{y}$). This specification of random coefficients is common in applied work (see, for instance, \citealp{berry1995automobile, nevo2000practitioner,nevo2001measuring, berry2004differentiated}). The functions $\beta_0,\beta_1:W\to{\mathds{R}}$ are unknown to the researcher and $\beta_1(w)\neq 0$ for all $w\in W$. I assume that $\mathbf{d}$ (and $\beta_1(\mathbf{w})$) is scalar without loss of generality since if $\mathbf{d}$ is a vector, then all components of it but one can be absorbed by $w$. In this case, one would need to use variation in those absorbed components to identify the coefficients in front of them. Similarly to the existing treatment of random coefficients model, I assume that the random coefficients in front of $\mathbf{z}_{y}$ are the same for each alternative $y$. However, I do not impose sign restrictions on $\big[\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}\big]$.\footnote{ Since
$\Pr(\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}> 0|\mathbf{x}=x)=1-F_{\mathbf{e}|\mathbf{x}}(-\beta_0(w)-\beta_{1}(w)d|x)$ and there are no restrictions on $\beta_0(\cdot)$, the random coefficient $\big[\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}\big]$ can be positive (negative) with probability that is arbitrarily close to $1$ if the support of $\mathbf{e}$ conditional on $\mathbf{x}=x$ is unbounded.}
\par
I start by stating two assumptions that will be used throughout the paper. The first one is a data requirement, the second one is a shape constraint on the distribution of latent variables.

\begin{assumption}[Data]\label{ass:observables}
The analyst can identify $p_0(x)=\Pr(\mathbf{y}=0|\mathbf{x}=x)$
for all $x\in X$
\end{assumption}

Assumption~\ref{ass:observables} implies that I only need to observe whether a consumer bought a product or not without knowing the identity of the product (see also, for instance, \citealp{thompson1989identification,lewbel2000semiparametric,fox2012random}).\footnote{The outcome $y=0$ can be replaced by any outcome. In this case, one will just need to renormalize the utility from that outcome to zero.} If the information on the identity of the purchases is also available, then this information (i) may improve the efficiency of an estimator; (ii) can help to satisfy the assumptions needed for identification (e.g., in my empirical illustration, I use one product to identify the sign of $\beta_1$ and I use another one to estimate it); and (iii) can be used to weaken the assumption that the random slope coefficient $\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}$ is the same across inside goods.

\begin{assumption}[Exclusion Restrictions]\label{ass:exclusion restriction}
For all $w\in W$
\begin{enumerate}
    \item $\boldsymbol{\varepsilon}$ is conditionally independent of $(\mathbf{e},\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$;
    \item $\mathbf{e}$ is conditionally independent of $(\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$.
\end{enumerate}
\end{assumption}
Assumption~\ref{ass:exclusion restriction} is an exclusion restriction that requires latent shocks $\mathbf{e}$ and $\boldsymbol{\varepsilon}$ to be independent of each other (condition (i)) and independent of excluded covariates $(\mathbf{d},\mathbf{z})$ (condition (ii)) after conditioning on $\mathbf{w}$. Assumption~\ref{ass:exclusion restriction} allows any form of dependence between $(\boldsymbol{\varepsilon},\mathbf{e})$ and nonexcluded covariates $\mathbf{w}$. That is, $\boldsymbol{\varepsilon}$ may contain latent product characteristics (e.g., unobserved quality) that can be correlated with nonexcluded covariates (e.g., market-product identifier).\footnote{Since, for identification and estimation, I require the average structural function $p_0$, some forms of endogeneity (i.e., correlation between $\mathbf{x}$ and $\boldsymbol{\varepsilon}$) can be addressed using suitable instruments and control function residuals as in \citet{blundell2004endogeneity} (see also \citealp{berry1994estimating, berry1995automobile, berry2014identification} for identification of structural demand function using aggregate data and instruments).  I leave the detailed analysis of this case for future research.}  In general, since I only require the identification of the structural demand function $p_0$, one can use the results in \citet{berry2020nonparametric} to identify $p_0$ and treat market-product level unobservables as a part of $\mathbf{w}$.
\par
Next, I provide two nonnested sets of conditions that allow for identification of $\beta_0$, $\beta_1$, and the distribution of $\mathbf{e}$ and $\boldsymbol{\varepsilon}$. In Section~\ref{sec:nonparam}, I impose no parametric assumptions on latent $\mathbf{e}$ and $\boldsymbol{\varepsilon}$ but assume some smoothness on the c.d.f. of $\boldsymbol{\varepsilon}$ and restrict the support of covariates. In Section~\ref{sec:normal error}, I identify the model when $\mathbf{e}$ is normally distributed, without any additional restrictions on the distribution of $\boldsymbol{\varepsilon}$ and with minimal support restrictions on covariates.

\section{Identification}\label{sec: identification}
\subsection{Nonparametric Identification}\label{sec:nonparam}
\begin{assumption}\label{ass:random coeff nonparametric}
For all $w\in W$
\begin{enumerate}
\item Conditional on $\mathbf{w}=w$, $\mathbf{e}$ has mean zero and variance one;
\item $F_{\boldsymbol{\varepsilon}|\mathbf{w}}(\cdot|w)$ has bounded partial derivatives up to order $\kappa$ for some $\bar{y}$ and $\partial^l_{\varepsilon^l_{\bar{y}}}F_{\boldsymbol{\varepsilon}|\mathbf{w}}(\cdot|w)|_{\varepsilon=0}\neq 0$ for all $l\leq\kappa$;
\item There exists $d^*$ such that the support of $(\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$ contains $(d^*,0)$ with an open neighborhood.
\end{enumerate}
\end{assumption}
Assumption~\ref{ass:random coeff nonparametric}(i) is a scale and location normalization. It restricts $\mathbf{e}$ conditional on $\mathbf{w}=w$ to have a finite expectation and a nonzero variance for all $w$. Assumption~\ref{ass:random coeff nonparametric}(ii) requires the conditional  distribution $F_{\boldsymbol{\varepsilon}|\mathbf{w}}$ to be sufficiently smooth in one component of $\varepsilon$ in the neighborhood of zero and have different from zero higher order partial derivatives. Since $\mathds{E}\left[\,\boldsymbol{\varepsilon}\,\right]$ is not assumed to be zero, if, for instance, $\boldsymbol{\varepsilon}$ is multivariate normal with component-wise nonzero mean, then Assumption~\ref{ass:random coeff nonparametric}(ii) is automatically satisfied. It is also generically satisfied when at least one component of $\boldsymbol{\varepsilon}$ is independent of the others and has a type I extreme value distribution \citep{fox2012random}. However, Assumption~\ref{ass:random coeff nonparametric}(ii) rules out cases when $\boldsymbol{\varepsilon}$ is a constant. Another example of violation of Assumption~\ref{ass:random coeff nonparametric}(ii) is when $\kappa$ is infinite and $F_{\boldsymbol{\varepsilon}|\mathbf{w}}$ is a polynomial function of any finite degree. (In Section~\ref{sec:normal error}, I provide an alternative result that does not restrict $F_{\boldsymbol{\varepsilon}|\mathbf{w}}$.) Assumption~\ref{ass:random coeff nonparametric}(iii) requires the support of $\mathbf{z}$ to contain zero with some open neighborhood. Assumptions similar to Assumptions~\ref{ass:random coeff nonparametric}(ii)-(iii) are common in the literature on identification of random coefficients models (e.g., Assumptions 8 and 10 in \citealp{fox2012random} and Assumption 4 in \citealp{allen2020identification}).

\begin{proposition}\label{prop:nonparameric identification1}
If Assumptions~\ref{ass:observables}-~\ref{ass:random coeff nonparametric} hold, then
$\beta_0(w)$, $\beta_1(w)$, and $\mathds{E}\left[\,\mathbf{e}^l|\mathbf{w}=w\,\right]$, $0\leq l\leq\kappa$, are identified for all $w\in W$.
\end{proposition}
Identification of $\kappa\leq\infty$ moments of the conditional distribution of $\mathbf{e}$ conditional on $\mathbf{w}$ is often sufficient for nonparametric identification of it. For example, Assumption 7 in \citet{fox2012random} uses the Carleman condition.\footnote{For more detailed discussion of the problem of identification of the distribution from its moments see, for instance, \citet{kleiber2013multivariate} and references therein.} Thus, under minimal restrictions, I can nonparametically identify the conditional c.d.f. $F_{\mathbf{v}|\mathbf{x}}$, where $\mathbf{v}=\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}$.
\par
To establish the next identification result I need the following definition.
\begin{definition}[Bounded completeness]\label{def:bounded completeness general}
The family of distributions $\left\{F_{\mathbf{v}|\mathbf{x}}(\cdot|x),x\in X'\right\}$ is boundedly complete if
\[
\forall x\in X',\:\int_{V}g(t)dF_{\mathbf{v}|\mathbf{x}}(t|x)=0\implies g(\mathbf{v})=0 \:\ensuremath{\mathrm{a.s.}},
\]
for any bounded function $g$.
\end{definition}
Completeness assumptions have been widely used in econometric analysis. Completeness is typically imposed on the distribution of observables (e.g., \citealp{newey03}). However, many commonly used parametric restrictions on the distribution of unobservables imply bounded completeness. For instance, it is satisfied for normal distributions and the Gumbel distribution.\footnote{For testability of the completeness assumptions see \citet{canay13}.}
\par
Combining bounded completeness with the identified distribution of the index $\mathbf{v}$, I have the following result.
\begin{proposition}\label{prop:nonparameric identification2}
If $F_{\mathbf{v}|\mathbf{x}}$ is identified and
    $
    \left\{F_{\mathbf{v}|\mathbf{x}}(\cdot|(d,z,w)),d\in D_{(z,w)}\right\}
    $
    is boundedly complete for all $(z,w)$ in the support, then the above model inherits all identifying properties of the random coefficients model with utilities
    $
        \mathds{1}\left(\,y\neq0\,\right)(\mathbf{r}_{y}+\boldsymbol{\varepsilon}_{y}).
    $
    The vector $\mathbf{r}=(\mathbf{r}_{y})_{y\in Y\setminus\{0\}}$ is an observed covariate conditionally independent of $\boldsymbol{\varepsilon}=(\boldsymbol{\varepsilon}_y)_{y\in Y\setminus\{0\}}$ conditional on $\mathbf{w}=w$ with the conditional support
    $
    R_w=\left\{r\in{\mathds{R}}^J\::\:r=v z,\: z\in Z_w, v\in V_w\right\},
    $
    where $V_w$ is the support of $\mathbf{v}$ conditional on $\mathbf{w}=w$.
    In particular, $F_{\boldsymbol{\varepsilon}|\mathbf{w}}$ is identified over $R_w$.
\end{proposition}
The proof of Proposition~\ref{prop:nonparameric identification2} is similar to the proof of Theorem 11 in \citet{fox2012random}. The main difference is that, instead of parametric restrictions, Proposition~\ref{prop:nonparameric identification2} uses the interaction between $\mathbf{d}$ and $\mathbf{z}$.
\par
Proposition~\ref{prop:nonparameric identification2} implies that the original random coefficient model can be represented in the ``special-covariate-with-full-support'' framework without assuming existence of such covariates. Moreover, if the set of directions that $z/\left\lVertz\right\rVert$ can cover is sufficiently rich and the support of $\mathbf{e}$ conditional on $\mathbf{w}=w$ is ${\mathds{R}}$, then $R_{w}={\mathds{R}}^{J}$ and all the identification results that require existence of special covariates with full support (e.g., \citealp{lewbel2000semiparametric, berry2009nonparametric, gautier2015triangular, fox2016nonparametric}, and \citealp{fox2020note}) can be applied.
\par
Combining the results in Propositions~\ref{prop:nonparameric identification1} and~\ref{prop:nonparameric identification2} with Theorem~1 in \citet{fox2020note}, I can establish the following result.
\begin{corollary}\label{corr: foxnote}
For all $y\neq 0$, let $\boldsymbol{\varepsilon}_y=\boldsymbol{\theta}^{{\sf T}}\mathbf{w}_y+\boldsymbol{\zeta}_y$, where $\boldsymbol{\theta}$ and $\boldsymbol{\zeta}=(\boldsymbol{\zeta}_y)_{y\inY\setminus\{0\}}$ are random coefficients, and $\mathbf{w}_y$ is the vector of product-$y$-specific covariates. Suppose
\begin{enumerate}
\item The assumptions of Propositions~\ref{prop:nonparameric identification1} and~\ref{prop:nonparameric identification2} hold;
\item $R_{w}={\mathds{R}}^{J}$ for all $w\in W$;
\item $(\boldsymbol{\theta},\boldsymbol{\zeta})$ and $\mathbf{w}=(\mathbf{w}_y)_{y\inY\setminus\{0\}}$ are independent;
\item The support of $\mathbf{w}$ contains an open ball of dimensionality of $\mathbf{w}$;
\item $(\boldsymbol{\theta},\boldsymbol{\zeta})$ has finite absolute moments and its distribution is uniquely determined by its moments;
\end{enumerate}
then $\beta_0$, $\beta_1$, and the distribution of $(\mathbf{e}_1,\boldsymbol{\theta},\boldsymbol{\epsilon})$ are identified.
\end{corollary}

To the best of my knowledge, Corollary~\ref{corr: foxnote} is the first result that establishes nonparametric identification of the whole distribution of the random coefficients in the multinomial choice environments without assuming the existence of special covariates. \citet{fox2012random,allen2020identification}, and \citet{lewbel2021semiparametric} also allow for bounded covariates. However, they either do not fully identify the distribution of the random intercept $\boldsymbol{\varepsilon}$ \citep{allen2020identification,lewbel2021semiparametric} or impose parametric restrictions on it \citep{fox2012random}.

\subsection{Normal Taste Shock}\label{sec:normal error}
\begin{assumption}\label{ass:normal error random coef}
For all $w\in W$
\begin{enumerate}
\item Conditional on $\mathbf{w}=w$, $\mathbf{e}$ is a standard normal random variable;
\item there exists $(d^*,z^{*\sf T})$ in the interior of the support of $(\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$ such that $z^*_{y}>0$ for all $y\in Y$;
\item there exists $(d^{**},z^{**\sf T})$ in the interior of the support of $(\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$ such that $p_0((d,z^{**},w))$ is neither an exponential nor an affine function of $d$ on some open set.
\end{enumerate}
\end{assumption}
Assumption~\ref{ass:normal error random coef}(i) requires $\mathbf{e}$ to be normally distributed with nonzero variance. With nonzero variance, the assumption that $\mathds{E}\left[\,\mathbf{e}^2\,\right]=1$ is just a scale normalization. The assumption is common in applied work (e.g., \citealp{nevo2000practitioner,nevo2001measuring}) and allows me to relax Assumptions~\ref{ass:random coeff nonparametric}(ii)-(iii). Assumption~\ref{ass:normal error random coef}(ii) is only needed for identification of the sign of $\beta_1(w)$. Assumption~\ref{ass:normal error random coef}(iii) means that if I fix all covariates but the one that shifts the random coefficient, then the probability of the default conditional on covariates is neither an affine nor an exponential function of this nonfixed covariate. Assumption~\ref{ass:normal error random coef}(iii) is not very restrictive since it rules out only some exponential and linear probability models. Moreover, it is testable.

\begin{proposition}\label{prop: normal mult choice}
If Assumptions~\ref{ass:observables},~\ref{ass:exclusion restriction}, and~\ref{ass:normal error random coef} hold, then
\begin{enumerate}
    \item $\beta_0(w)$ and $\beta_1(w)$ are identified for all $w\in W$;
    \item The conditions of Proposition~\ref{prop:nonparameric identification2} are satisfied.
\end{enumerate}
\end{proposition}
The proof of the identification of $\beta_{0}$ and $\beta_{1}$ uses the multiplicative structure of $d$ and $z$, and properties of the standard normal p.d.f. Informally, note that
\[
\beta_{0}(\mathbf{w})\mathbf{z}+\beta_{1}(\mathbf{w})\mathbf{d}\mathbf{z}+\mathbf{e}\mathbf{z}.
\]
Since $d$ and $z$ can be moved independently, I can use variation in $d$ while keeping $dz$ by varying $z$ to identify $\beta_{0}(w)$. Then, by varying $z$, I can identify $\beta_{1}(w)$. Proposition~\ref{prop: normal mult choice}(ii) follows from $\beta_0(w)$ and $\beta_1(w)$ being identified and $\mathbf{e}$ being standard normal (i.e, $\beta_0(\mathbf{w})+\beta_1(\mathbf{w})\mathbf{d}+\mathbf{e}$ conditional on $\mathbf{x}=x$ generates a boundedly complete family of distributions).
\par
Note that the only restriction on $\boldsymbol{\varepsilon}$ needed for Proposition~\ref{prop: normal mult choice} is the conditional independence assumption (Assumption~\ref{ass:exclusion restriction}). The random intercept $\boldsymbol{\varepsilon}$ is allowed be continuously or discretely distributed (e.g., it may be a constant). Hence, I can extend Theorem 2 in \citet{fox2016nonparametric} to environments with bounded covariates.
\begin{corollary}
For all $y\neq 0$ let $\boldsymbol{\varepsilon}_y=\boldsymbol{\theta}_y(\mathbf{w})$, where $\boldsymbol{\theta}_{y}$ is a random function such that its realization $\theta_y$ is a map from $W$ to ${\mathds{R}}$. Suppose
\begin{enumerate}
\item Assumptions of Proposition~\ref{prop: normal mult choice} hold;
\item $R_{w}={\mathds{R}}^{J}$ for all $w\in W$;
\item $\boldsymbol{\theta}=(\boldsymbol{\theta}_y)_{y\neq 0}$ and $\mathbf{w}$ are independent;
\item The support of $\boldsymbol{\theta}$, $\Theta$, satisfies Assumption 4 in \citet{fox2016nonparametric};
\end{enumerate}
then $\beta_0$, $\beta_1$, and the distribution of $\boldsymbol{\theta}$ are identified.
\end{corollary}

\subsection{Bundles}\label{sub: bundles}
Note that since I do not assume independence among $\boldsymbol{\varepsilon}_{y}$ across $y$, the multinomial choice model I study covers some bundles models \citep{gentzkow2007valuing, dunker2017nonparametric,fox2017note}. In particular, assume that there are $\tilde{J}$ goods and the agent can purchase any bundle consisting of these goods. The vector $\tilde{y}$ describes the purchasing decision of the agent. That is, $\tilde{y}\in\tilde{Y}=\{0,1\}^{\tilde{J}}$. For instance, $\tilde{y}=(0,1,0,1,0,\dots,0)$ corresponds to the case when the agent purchased a bundle of goods $2$ and $4$. The random utility from choosing bundle $\tilde{y}\neq 0$ is of the form
\begin{align*}
    (\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}) \sum_{j=1}^{\tilde{J}} \tilde{y}_j \mathbf{\tilde{z}}_{j}+\boldsymbol{\varepsilon}_{\tilde{y}},
\end{align*}
and the utility from buying nothing is zero. I can rewrite the above utilities from bundles as as the utilities form the multinomial choice problem since there are finitely ($2^{\tilde{J}}$) possible bundles. Indeed, I can enumerate them all with $y=0$ corresponding to $\tilde{y}=0\in{\mathds{R}}^{\tilde{J}}$ (i.e., $Y=\{0,1,2,\dots,2^{\tilde{J}}\}$) and define $z_y=\sum_{j=1}^{\tilde{J}} \tilde{y}_j \tilde{z}_{j}$. As a result, I can extend the conclusions of Theorem 1 in \citet{fox2017note} to environments with bounded covariates
\begin{corollary}
Let $J=2$ and
\begin{align*}
    \boldsymbol{\varepsilon}_{(1,0)}&=\theta_1(\mathbf{w})+\boldsymbol{\epsilon}_1,\quad\quad
    \boldsymbol{\varepsilon}_{(0,1)}=\theta_2(\mathbf{w})+\boldsymbol{\epsilon}_2,\\
    \boldsymbol{\varepsilon}_{(1,1)}&=\boldsymbol{\varepsilon}_{(1,0)}+\boldsymbol{\varepsilon}_{(0,1)}+\boldsymbol{\xi}\theta_3(\mathbf{w}),
\end{align*}
where $\theta_i(\cdot)$, $i=1,2,3$, are some unknown functions, and $(\boldsymbol{\epsilon}_1,\boldsymbol{\epsilon}_2,\boldsymbol{\xi})\in{\mathds{R}}^2\times{\mathds{R}}_{+}$. Suppose
\begin{enumerate}
    \item Assumptions of Propositions~\ref{prop:nonparameric identification1} and~\ref{prop:nonparameric identification2} or Proposition~\ref{prop: normal mult choice} hold;
    \item $R_w={\mathds{R}}$ for all $w$;
    \item $(\boldsymbol{\epsilon}_1,\boldsymbol{\epsilon}_2)|\mathbf{w}=w$ has an everywhere positive Lebesgue density on its support for all $w\in W$;
    \item $\mathds{E}\left[\,\boldsymbol{\epsilon}_i|\mathbf{w}=w\,\right]=0$ and $\mathds{E}\left[\,\boldsymbol{\xi}|\mathbf{w}=w\,\right]=1$ for all $w\in W$ and $i=1,2$,
\end{enumerate}
then $\theta_i(\cdot)$, $i=1,2,3$, and the c.d.fs $F_{\boldsymbol{\epsilon}_i|\mathbf{w}}$, $i=1,2$, and $F_{\boldsymbol{\xi}|\mathbf{w}}$ are identified.
\end{corollary}

\section{Estimation of \texorpdfstring{$\beta$}{the Finite-Dimensional parameter}}\label{sec: estimation}
Proposition~\ref{prop: normal mult choice} constructively identifies $\beta_{0}$ and $\beta_{1}$. In this section, I use it to estimate these parameters. That is, I focus on the multinomial choice model with random coefficients with normally distributed $\mathbf{e}$.\footnote{Proposition~\ref{prop:nonparameric identification1} also provides a constructive identification for $\beta_0$ and $\beta_1$. However, Assumption~\ref{ass:random coeff nonparametric}(iii) fails to hold in my illustrative application presented in Section~\ref{sec: empirical application}. Additionally, Proposition~\ref{prop:nonparameric identification1} uses limits of derivatives of identifiable functions at a single point, thus, most likely, leading to a consistent estimator with nonparametric rate of convergence.} Moreover, to simplify the exposition, I assume that there are no nonexcluded covariates $\mathbf{w}$ (i.e., $\beta_0(\cdot)$ and $\beta_1(\cdot)$ are constant functions). Note that, even though $\beta_0$ and $\beta_1$ are finite-dimensional parameters and the distribution of $\mathbf{e}$ is assumed to be known, the model is still semiparametric since the distribution of $\boldsymbol{\varepsilon}$ is not parametric.
\par
The first ingredient of the estimator is a nonparametric estimator of $p_0(\cdot)=\Pr(\mathbf{y}=0|\mathbf{x}=\cdot)$, $\hat{p}_0(\cdot)$. Any consistent and smooth enough estimator $\hat{p}_0$ will deliver a consistent estimator of $\beta=(\beta_1,\beta_0)$.\footnote{The normality of $\mathbf{e}$ implies that $p_0$ has continuous derivatives of any order. See Appendix~\ref{app: proof of est} for details.} For concreteness, I work with the series estimator based on products of powers of components of $x=(d,z)$ (polynomial regressions). That is, given a sample of independent identically distributed (i.i.d.) observations on covariates and a binary random variable that indicates whether the product was purchased or not $\left\{\mathds{1}\left(\,\mathbf{y}^{(i)}=0\,\right),\mathbf{x}^{(i)}\right\}_{i=1}^n$, define
\begin{align*}
\hat{p}_0(x)&=\psi^K(x)^{{\sf T}} \left(\Psi^{{\sf T}}\Psi\right)^{-}\sum_{i=1}^n\psi^K\left(\mathbf{x}^{(i)}\right)\mathds{1}\left(\,\mathbf{y}^{(i)}=0\,\right),
\end{align*}
where $\psi^K(\cdot)$ is a vector of orthonormal basis functions based on products of powers of components of $x$,
$\Psi=\left(\psi^K\left(\mathbf{x}^{(1)}\right),\psi^K\left(\mathbf{x}^{(2)}\right),\dots,\psi^K\left(\mathbf{x}^{(n)}\right)\right)^{{\sf T}}$, and $\left(\Psi^{{\sf T}}\Psi\right)^{-}$ is the Moore-Penrose generalized inverse. I assume that the sum of powers of components of $x$ in $\psi^K$ is monotonically increasing in $K$.
\par
The sign of $\beta_1$ can be trivially estimated from $\hat{p}_0$ since
\[
\mathrm{sign}(\beta_1)=\mathrm{sign}\left(p_0((d',z))-p_0((d,z))\right)\mathrm{sign}(z_{y^*})\mathrm{sign}(d'-d)
\]
if $z\geq 0$ or $z\leq 0$ with $z_{y^*}\neq 0$. Hence, for simplicity I assume that $\beta_1>0$.
\par
The identification result in Proposition~\ref{prop: normal mult choice} is constructive and provides a closed form expression for $\beta$ as a functional of $p_0$ (see Appendix~\ref{app: proof of est}). Given the nonparametric power series estimator $\hat{p}_0$, the plug-in estimator of $\beta$ is
\begin{align*}\label{eq:beta^2}
    \hat{\beta}_1&=\sqrt{\frac{\displaystyle\sum\limits_{i=1}^n\hat{p}_{111}\left(\mathbf{x}^{(i)}\right)\hat{p}_1\left(\mathbf{x}^{(i)}\right)-\hat{p}_{11}\left(\mathbf{x}^{(i)}\right)^2}{\displaystyle\sum\limits_{i=1}^n\hat{p}_{12}\left(\mathbf{x}^{(i)}\right)\hat{p}_1\left(\mathbf{x}^{(i)}\right)-\hat{p}_2\left(\mathbf{x}^{(i)}\right)\hat{p}_{11}\left(\mathbf{x}^{(i)}\right)-\hat{p}_1\left(\mathbf{x}^{(i)}\right)^2}},\\
    \hat{\beta}_0&=\hat{\beta}_1\dfrac{\displaystyle\sum\limits_{i=1}^n\hat{p}_{2}\left(\mathbf{x}^{(i)}\right)-\mathbf{d}^{(i)}\hat{p}_{1}\left(\mathbf{x}^{(i)}\right)}{\displaystyle\sum\limits_{i=1}^n\hat{p}_{1}\left(\mathbf{x}^{(i)}\right)}-\frac{1}{\hat{\beta}_1}\dfrac{\displaystyle\sum\limits_{i=1}^n\hat{p}_{11}\left(\mathbf{x}^{(i)}\right)}{\displaystyle\sum\limits_{i=1}^n\hat{p}_{1}\left(\mathbf{x}^{(i)}\right)},
\end{align*}
where
\begin{align*}
    \hat{p}_{1}(x)&=\partial_{d}\hat{p}_0(x),\quad
    \hat{p}_{11}(x)=\partial^2_{d^2}\hat{p}_0(x),\quad
    \hat{p}_{111}(x)=\partial^3_{d^3}\hat{p}_0(x),\\
    \hat{p}_2(x)&=\sum_{y=1}^{J}z_{y}\partial_{z_{y}}\hat{p}_0(x),\quad
    \hat{p}_{12}(x)=\partial_{d}\hat{p}_2(x).
\end{align*}
Note that $\hat\beta$ is essentially a nonlinear function of sample averages of different derivatives of estimated $\hat{p}_0$. Following \citet{newey1994asymptotic, newey1997convergence}, to achieve $\sqrt{n}$-consistency and asymptotic normality of the proposed estimator, I will have to establish existence of the Reisz representer of a particular directional derivative. Let
\begin{align*}
    \bar{v}_1(x)&=-\Big[4p_{1111}(x)f_{\mathbf{x}}(x)+8p_{111}(x)\partial_{d}f_{\mathbf{x}}(x)+5p_{11}(x)\partial^2_{d^2}f_{\mathbf{x}}(x)+p_{1}(x)\partial^3_{d^3}f_{\mathbf{x}}(x)\Big]/f_{\mathbf{x}}(x),\\
    \bar{v}_2(x)&=\Big[\beta_1\{(1-J)f_{\mathbf{x}}(x)+d\partial_{d}f_{\mathbf{x}}(x)-\sum_{y}z_{y}\partial_{z_{y}}f_{\mathbf{x}}(x)\}-\partial^2_{d^2}f_{\mathbf{x}}(x)\Big]/f_{\mathbf{x}}(x),\\
    \bar{v}(x)&=(\bar{v}_1(x),\bar{v}_2(x)),
\end{align*}
where $f_{\mathbf{x}}$ is the p.d.f. of $\mathbf{x}$, and $p_1$, $p_{11}$, $p_{111}$, and $p_{1111}$ are first, second, third, and forth derivatives of $p_0$ with respect to $d$, respectively.
\begin{assumption}\label{ass: regularity}
\begin{enumerate}
    \item The support of $\mathbf{x}$, $X$, is a Cartesian product of compact connected nonsingleton intervals in ${\mathds{R}}$.
    \item $f_{\mathbf{x}}$ is bounded away from zero on the interior of $X$;
    \item $f_{\mathbf{x}}$, $\partial_{d}f_{\mathbf{x}}$, $\partial_{z_{y}}f_{\mathbf{x}}$, and $\partial^2_{d^2}f_{\mathbf{x}}$ equal to zero at the boundary of $X$ for all $y$;
    \item $\mathds{E}\left[\,\bar{v}(\mathbf{x})\bar{v}(\mathbf{x})^{{\sf T}}\,\right]$ is finite and nonsingular.
\end{enumerate}
\end{assumption}
Assumptions~\ref{ass: regularity}(i)-(ii) are standard in the literature on nonparametric estimation of conditional expectations.
Similarly to the average derivative estimator of \citet{powell1989semiparametric}, to achieve $\sqrt{n}$-consistency the estimator I need to impose restrictions on the behavior of $f_{\mathbf{x}}$ on the boundary of its support. Since \citet{powell1989semiparametric} work with the first derivative they only require $f_{\mathbf{x}}$ to vanish on the boundary. My estimator involves derivatives up to order 3, thus, leading to Assumption~\ref{ass: regularity}(iii). Assumption~\ref{ass: regularity}(iv) is the mean-square continuity condition that requires the variance of the score function of $\mathbf{x}$ (i.e $\log f_{\mathbf{x}}$) and derivatives of it to be finite.
\par
The following proposition establishes asymptotic normality of my estimator and is based on Theorem 6 in \citet{newey1997convergence}. Denote
\[
G=\left(\begin{array}{cc}
    2\beta_1 &  0\\
    0 & 1
\end{array}\right)\left(\begin{array}{cc}
    {\mathds{E}\left[\,p_{12}(\mathbf{x})p_{1}(\mathbf{x})-p_{2}(\mathbf{x})p_{11}(\mathbf{x})-p_{1}(\mathbf{x})^2\,\right]} &  0\\
    0 & {\beta_1\mathds{E}\left[\,p_{1}(\mathbf{x})\,\right]}
\end{array}\right)^{-1},
\]

\begin{proposition}\label{prop: estimator}
If (i) $\left\{\mathds{1}\left(\,\mathbf{y}^{(i)}=0\,\right),\mathbf{x}^{(i)}\right\}_{i=1}^n$ are i.i.d.; (ii) Assumptions~\ref{ass:exclusion restriction},~\ref{ass:normal error random coef} and~\ref{ass: regularity} are satisfied, and Assumption~\ref{ass:normal error random coef}(iii) is satisfied for all $x^{**}=(d^{**},z^{**})\in X$; (iii) $K^6/n\to_{n\to\infty}0$, then
\[
\sqrt{n}(\hat\beta-\beta)\to_{d} \mathrm{N}(0,V),
\]
where $V=G\mathds{E}\left[\,\bar{v}(\mathbf{x})\bar{v}(\mathbf{x})^{{\sf T}} p_0(\mathbf{x})(1-p_0(\mathbf{x}))\,\right]G^{{\sf T}}$.
\end{proposition}
In the proof of Proposition~\ref{prop: estimator}, I also provide a consistent estimator of the asymptotic variance matrix $V$ that is based on the estimator proposed in \citet{newey1997convergence}.
\par
I conclude this section by noting that after $\beta$ is estimated, one can construct a sieve maximum-likelihood estimator of $F_{\boldsymbol{\varepsilon}}$ since
\[
\Pr(\mathbf{y}=0|\mathbf{x}=x)=\int_{\mathds{R}} F_{\boldsymbol{\varepsilon}}(tz_1,tz_2,\dots,tz_{J})\phi \left(t+\beta_0+\beta_1d\right)dt
\]
where $\phi(\cdot)$ is the standard normal p.d.f. Thus, one can find the maximizer of
\begin{align*}
\max_{F\in\mathcal{F}_n}\sum_{i=1}^n &\mathds{1}\left(\,\mathbf{y}^{(i)}=0\,\right)\log\left(\int_{\mathds{R}} F(tz_1,tz_2,\dots,tz_{J})\phi \left(t+\hat{\beta}_0+\hat{\beta}_1d\right)dt\right)+\\
&\mathds{1}\left(\,\mathbf{y}^{(i)}\neq0\,\right)\log\left(1-\int_{\mathds{R}} F(tz_1,tz_2,\dots,tz_{J})\phi \left(t+\hat{\beta}_0+\hat{\beta}_1d\right)dt\right),
\end{align*}
where $\{\mathcal{F}_n\}_{n=1}^{\infty}$ is a sequence of sieve spaces for $F_{\boldsymbol{\varepsilon}}$. Inference on known functionals of $\beta$ and $F_{\boldsymbol{\varepsilon}}$ (e.g., counterfactuals) can be done using likelihood-ratio type statistic (see, for instance, \citealp{shen2005sieve,chen2014sieve}).\footnote{Both $\beta$ and $F_{\boldsymbol{\varepsilon}}$ can be estimated in one step by the sieve maximum-likelihood estimator. In this case, however, the estimator of $\beta$ may not be $\sqrt{n}$ consistent.}
\section{Monte-Carlo Simulations}\label{sec: simulations}
In this section, I assess the performance of my estimator in finite samples. I consider the binary choice model:
\[
\mathbf{y}=\mathds{1}\left(\,(\beta_0+\beta_1\mathbf{d}+\mathbf{e})\mathbf{z}+\beta_3+\boldsymbol{\varepsilon}\geq 0\,\right),
\]
where $\beta_0=-0.5$, $\beta_1=1$, and $\mathbf{e}$ is a standard normal random variable. The random intercept $\beta_3+\boldsymbol{\varepsilon}$ is independent from $\mathbf{x}$ and $\mathbf{e}$ with mean $\beta_3=0.5$. The observed covariates $\mathbf{x}=(\mathbf{d},\mathbf{z})$ are distributed according to a monotone transformation of a bivariate normal distribution: $\mathbf{x}=5(\mathrm{arctan}(\mathbf{\tilde{x}})/\pi+0.5)$, where $\mathbf{\tilde{x}}$ is a mean-zero normal random vector such that each component of it has variance 1 and the correlation between components is $0.1$. Note that $\mathbf{x}$ has bounded support.
\par
I consider several data generating processes (DGPs). The first one (DGP-$0$) is when $\boldsymbol{\varepsilon}$ is a standard normal random variable. The next five DGPs correspond to $\boldsymbol{\varepsilon}$ being an equally weighted  mixture of three unit-variance normal distributions with mean $-t$, $0$, and $t$ for $t\in\{1,2,3,4,5\}$ (DGP-$t$). For every $t$ the distribution of $\boldsymbol{\varepsilon}$ is symmetric. However, the variance is growing with $t$ and the distribution changes from a unimodal distribution to a distribution with three modes. Finally, DGP-L corresponds to the case with logistically distributed $\boldsymbol{\varepsilon}$.
\par
Each experiment is conducted $1000$ times for every DGP for 3 sample sizes $n\in\{10^3,5\cdot10^3,10^4\}$. I use a tensor product of cubic polynomials in estimation of the conditional probability $p_0$.\footnote{The results are qualitatively the same for higher order polynomials.} The results for the mean deviation (bias) of the estimator of $\beta_1$ are presented in Table~\ref{table: bias my}. As expected, the bias decreases with the sample size.\footnote{The mean absolute deviation of the estimator also decreases with the sample size. See, Appendix~\ref{app: simulations} for further details.} However, there is not much variation across DGPs.\footnote{For comparison of my estimator with two alternative potentially misspecified parametric estimators, see Appendix~\ref{app: simulations}.}
\begin{table}[h]
\centering
\begin{threeparttable}
\centering
\caption{Bias}\label{table: bias my}
\begin{tabular}{c|ccccccc}
\hline
\hline
Sample Size & DGP-$0$ & DGP-$1$ & DGP-$2$ & DGP-$3$ & DGP-$4$ & DGP-$5$ & DGP-L\\
\hline
1000            & 1.08    & 1.10    & 1.18    & 1.46    & 1.51    & 1.50    & 1.12 \\
5000            & 0.36    & 0.54    & 0.89    & 1.05    & 1.25    & 1.23    & 0.62 \\
10000           & 0.17    & 0.26    & 0.57    & 0.84    & 1.09    & 1.17    & 0.38 \\
\hline
\end{tabular}
\vspace{1ex}
\end{threeparttable}
\end{table}

\section{Illustrative Empirical Application}\label{sec: empirical application}
To illustrate the empirical importance of the relaxation of the parametric assumptions about the distribution of $F_{\boldsymbol{\varepsilon}}$ and the proposed estimation procedure, I analyze margarine purchasing decisions of households from Springfield, MO, USA, using the multinomial choice model with normally distributed $\mathbf{e}$. I find substantial differences between estimates obtained by employing my semiparametric estimator and a fully parametric multinomial-logit-type estimator.

\subsection*{Data}
The original dataset, constructed by \citet{allenby1991quality}, is a panel of 9196 purchases of 10 brands of stick and tube margarine by 517 households from Springfield, MO, USA, extracted from an ERIM (A.C. Nielsen) scanner dataset. The dataset contains information on the shelf prices of each brand that is constructed using the actual price paid and the value of any redeemed coupon. The household demographics contain information on the household income.\footnote{See \citet{allenby1991quality} for specific details of the dataset construction.} \citet{benoit2016outlier} focused on 5 brands instead of 10 and transformed this dataset to a cross-section with 242 households. In particular, every observation contains only information on the household annual income, which I use as the agent-specific covariate $d$, agent choices ($y$), and product-specific prices $p_y$.\footnote{Income and prices are measured in thousands of US dollars and US dollars, respectively.} There are 5 brands: Generic ($y=0$), Blue Bonnet ($y=1$), House Brand ($y=2$), Shed Spread ($y=3$), and Fleischmann's ($y=4$).
\par
Income varies from 2.5k to 130k, with the median and average income being 26.75k and 22.5k, respectively. Table~\ref{table: sum1} summarizes  the share and price information for different products. There is a variation in prices across brands with Generic being on average the cheapest and Fleischmann's being the most expensive. At the same time, Fleischmann's is the least demanded product.
\begin{table}[h]
\centering
\begin{threeparttable}
\centering
\caption{Summary Statistics for Products}\label{table: sum1}
\begin{tabular}{c|ccccc}
\hline
\hline
Brand         & Share & Average Price & Median Price & Min Price & Max Price\\
\hline
Generic       & 0.17  & 0.37          & 0.36         & 0.33      & 0.53\\
Blue Bonnet   & 0.30  & 0.58          & 0.61         & 0.19      & 0.76\\
House Brand   & 0.19  & 0.51          & 0.57         & 0.19      & 0.58\\
Shed Spread   & 0.21  & 0.83          & 0.85         & 0.50      & 0.98\\
Fleischmann's & 0.12  & 1.04          & 1.08         & 0.99      & 1.13\\
\hline
\end{tabular}
\vspace{1ex}
\end{threeparttable}
\end{table}
\subsection*{Utility}
I follow \citet{nevo2000practitioner,nevo2001measuring} and model the utility from purchasing brand $y\in\{0,1,2,3,4\}$ as
\[
\boldsymbol{\delta}\mathbf{d}+(\beta_0+\beta_1 \mathbf{d}+\mathbf{e}) \mathbf{p}_{y}+\tilde{\boldsymbol{\varepsilon}}_y.
\]
The random coefficient $\boldsymbol{\delta}$ captures the direct marginal effect of income on utility from consumption of margarine (i.e., it is the same for all brands). The coefficient $\beta_0+\beta_1 \mathbf{d}$ can be thought of as the average marginal utility with respect to price. It captures the sensitivity of agents with respect to prices and is expected to be negative. Agents with different incomes may react differently to price changes. Note that no assumptions are made about $\tilde{\boldsymbol{\varepsilon}}_y$ (e.g., it is not assumed that it is has zero mean).\footnote{Estimation using $\log(\mathbf{p}_y)$ instead of $\mathbf{p}_y$ gives qualitatively similar results.} This utility specification correspond to the ``preference shifter'' specification in \citet{griffith2018income}.
There is no information about those who did not purchase any margarine products, thus, I analyze the choices of those who already decided to purchase a margarine product.
If I treat the utility from consuming Generic brand as the baseline utility and subtract it from all utilities, the normalized utility from purchasing different brands for $y=1,2,3,4$ is
\[
(\beta_0+\beta_1 \mathbf{d}+\mathbf{e}) [\mathbf{p}_{y}-\mathbf{p}_{0}]+\tilde{\boldsymbol{\varepsilon}}_y-\tilde{\boldsymbol{\varepsilon}}_0,
\]
and the utility from purchasing Generic brand is $0$. Hence, I can define $\mathbf{z}_{y}=\mathbf{p}_{y}-\mathbf{p}_{0}$ and $\boldsymbol{\varepsilon}_y=\tilde{\boldsymbol{\varepsilon}}_y-\tilde{\boldsymbol{\varepsilon}}_0$, $y=1,2,3,4$, where $\mathbf{p}_0$ is the price of Generic margarine.
\par
Given that I am considering margarine products, it is not surprising that the support for $\mathbf{z}_y$ is far from being full. In particular, $\max_{y}\max_{i}z^{(i)}_y=0.78$ and  $\min_{y}\min_{i}z^{(i)}_y=-0.15$. At the same time, there is still variation in relative prices $\mathbf{z}_y$ and income $\mathbf{d}$. This variation allows me to recover $\beta$ without specifying the distribution of $\boldsymbol{\varepsilon}$.
\par
In the current application, I use a minimal amount of information: there are only two covariates. If one has more demographic and product data, it can be easily incorporated into the current framework via $\mathbf{w}$. For instance,  $\mathbf{w}$ may contain nonprice marketing variables, packet size dummies, saturated fat content, household size, age of the household head, household location (e.g. zip-code).

\subsection*{Parametric Estimation}
First, I assume the most common parametric specification for the random intercept -- multinomial logit. Formally, I estimate the following specification for normalized utility:
\[
\mathds{1}\left(\,y\neq0\,\right)[\gamma_y+(\beta_0+\beta_1 \mathbf{d}+\mathbf{e}) \mathbf{z}_{y}]+\alpha\boldsymbol{\varepsilon}_y,
\]
where $\{\boldsymbol{\varepsilon}_y\}_{y=0}^4$ are i.i.d. Gumbel across $y$ that are also independent from $\mathbf{x}=(\mathbf{d},\mathbf{z})$; $\mathbf{e}$ is a standard normal random variable. (Parameter $\alpha$ captures the scale of $\boldsymbol{\varepsilon}_y$ since the variance of $\mathbf{e}$ is set to 1.) Although, price $\mathbf{p}_y$ is probably correlated with unobserved part of the utility $\tilde{\boldsymbol{\varepsilon}}_y$ (e.g., unobserved quality), the price difference $\mathbf{z}_y=\mathbf{p}_y-\mathbf{p}_0$ may be independent from $\tilde{\boldsymbol{\varepsilon}}_y-\tilde{\boldsymbol{\varepsilon}}_0$.
\par
The estimates of $\beta_0$ and $\beta_1$ are $\bar{\beta}_0=-6331.94$ (standard error$=17.19$) and $\bar{\beta}_1=-19.69$ (standard error$=514.48$), respectively. As expected, the sign of $\bar{\beta}_0$ is negative. The coefficient in front of the income variable, $\bar{\beta}_1$, is negative and not significant at the $5$ percent significance level. Although income does not matter much, the overall sensitivity to prices (mostly captured by $\bar\beta_0$ in this case) is substantial. The effect of income on marginal disutility from the price increase is not surprising given that margarine constitutes a small share of household expenditures on groceries.\footnote{E.g., in UK households spend about one percent of their grocery expenditures on margarine and butter \citep{griffith2018income}.}


\subsection*{Semiparametric Estimation}
Next, I apply the estimator proposed in Section~\ref{sec: estimation}. Formally, I estimate the following specification for normalized utility:
\[
\mathds{1}\left(\,y\neq0\,\right)[(\beta_0+\beta_1 \mathbf{d}+\mathbf{e}) \mathbf{z}_{y}+\boldsymbol{\varepsilon}_y],
\]
where $\mathbf{e}$ is a standard normal random variable. The random intercept $\boldsymbol{\varepsilon}=(\boldsymbol{\varepsilon}_y)_{y=1,2,3,4}$ is assumed to be independent from $\mathbf{x}$. There are \emph{no} other restrictions on the joint distribution of $\boldsymbol{\varepsilon}$. This specification nests the logit specification estimated in the previous section. Hence, if the assumptions of multinomial logit are correct, then the results of parametric and semiparametric estimators should not differ much.
\par
The estimates of $\beta_0$ and $\beta_1$ are $\hat\beta_0=-39.1$ (standard error$=43.8$) and $\hat\beta_1=-16.7\times 10^{-3}$ (standard error$=3.97\times 10^{-6}$).
\footnote{I use the tensor product of the $4$-th degree Chebyshev polynomials for $d$ and the $1$-st degree Chebyshev polynomials for every $z_{y}$.}
Similar to the multinomial logit estimator, the sign of $\hat\beta_0$ is negative. The coefficient in front of the income variable is negative and significant at the $5$ percent significance level. However, the maximal value that $\hat{\beta}_1\mathbf{d}$ can take in the sample is substantially smaller than $\hat\beta_0$
($\max_i\left(\mathbf{d}^{(i)}\hat{\beta}_1 /\hat{\beta}_0\right)=0.055$, standard error$=0.062$). The latter indicates that, similarly to the fully parametric specification, income does not affect marginal disutility from price increase much. However, the estimate of $\beta_0$ is substantially lower than the one in the fully parametric case. This indicates that consumers may be less sensitive to price changes than one would think after estimating the logit-type model.
\par
Interestingly, the difference between the estimates obtained using the fully parametric logit estimator $\bar{\beta}$ and my semiparametric estimator $\hat{\beta}$ is substantial (e.g., $\bar\beta_1/\hat\beta_1>10^3$). That is, the parametric estimator overestimates the magnitude of the agents sensitivity to relative price changes of margarine.  This suggests that the multinomial logit structure most likely fails to hold, emphasizing the importance of semiparametric estimation.\footnote{This empirical finding is in line with the simulation results, presented in Appendix~\ref{app: simulations}.}

\section{Conclusion}\label{sec:conclusion}
This paper shows that commonly used exclusion restrictions and richness assumptions about the distribution of some unobservables may lead to full nonparametric identification in discrete outcome models even when covariates are bounded. The proposed identification framework extends the results from a large literature that uses special covariates with full support to environments where such full-support covariates are not available. It also leads to an asymptotically normal estimator of the finite-dimensional parameters of the model.

\bibliography{references}