EconBase
← Back to paper

Identification and estimation of multinomial choice models with latent special covariates

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

53,692 characters · 14 sections · 47 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and estimation of multinomial choice models with latent special covariates

abstractIdentification of multinomial choice models is often established by using special covariates that have full support. This paper shows how these identification results can be extended to a large class of multinomial choice models when all covariates are bounded. I also provide a new $\sqrt{n}$-consistent asymptotically normal estimator of the finite-dimensional parameters of the model. JEL classification numbers: C50, C57 Keywords: Multinomial choice, random coefficients, special covariate, identification at infinity, bundles

Introduction

This paper studies identification and estimation of random coefficients multinomial choice models with covariates that have bounded support. Often some latent variables in these models have full support (i.e. supported on the whole Euclidean space). Under common restrictions on the distribution of these unobservables, I constructively identify it and show how these latent variables can be used to construct special covariates (i.e., artificial observables with full support) to nonparametrically identify the distribution of all the other unobservables. Identification of all parts of the structural model is crucial for welfare analysis (e.g., aggregate welfare changes between two choice situations). My identification technique is constructive and leads to an asymptotically normal estimator of the finite-dimensional parameters of the model. The results of this paper rest on two commonly used assumptions. First, I assume existence of excluded covariates that affect the distribution over choices via a random coefficient. Using variation in these excluded covariates I can identify the distribution of the random coefficient. Second, I assume that the distribution of the random coefficient is sufficiently “rich”. “Richness” of the random coefficient distribution is formalized by a notion of bounded completeness.\footnote{Completeness of a family of distributions is a well-known concept in the Statistics and Econometrics literature. See, for example, mattner1993some, newey03, chernozhukov2005iv, blundell07, chernozhukov2007instrumental, hu2008instrumental, andrews11, darolles11, and d2011completeness.} As a result, I show how to identify the distribution over outcomes conditional on the realization of the observed covariates and the latent random coefficient nonparametrically. Since the latent random coefficient often has full support, I can treat it as an observed covariate with full support and apply any identification technique that requires existence of such covariates to identify the rest of the model parameters (e.g., the distribution of other latent variables). I provide two nonnested identification results. The first result does not make any parametric assumptions about the distribution of latent variables. It, however, imposes some restrictions on the support of observables. In particular, I require the support of some covariates to contain zero. It also requires some smoothness of the distribution of the latent variables. To the best of my knowledge, this is the first result in the literature that nonparametrically identifies the distribution of all latent variables in multinomial choice settings with bounded covariates. The second result uses one of the most popular parameterizations in applied work - a Gaussian distribution of the latent random coefficient. But, in contrast to the first result, it does not require zero in the support of covariates and leaves the distribution of other latent variables completely unrestricted. The second result also leads to an easy to implement asymptotically normal estimator of the finite-dimensional parameters of the model. Similar to powell1989semiparametric, this estimator is $\sqrt{n}$-consistent since it is based on average derivatives of an estimable object. I contribute to the discrete outcome literature in several respects. I show how existing results that use full-support-excluded covariates with monotonicity restrictions\footnote{See, for example, manski1985semiparametric, manski1988identification, heckman1990varieties, matzkin1992nonparametric, ichimura1998maximum, lewbel1998semiparametric, lewbel2000semiparametric, tamer03, matzkin2007heterogeneous, berry2009nonparametric, BHR, gautier2013nonparametric, gautier2015triangular, fox2016nonparametric, dunker2017nonparametric, fox2017note, fox2018unobserved, fox2020note, and kashaev2020discerning.} can be directly used in environments with bounded covariates. Formally, I demonstrate that my setting inherits all identifying properties of the setting with a special covariate. I also contribute to the literature on semiparametric models by showing that common parametric restrictions can be used instead of covariates that have full support (e.g., fox2012random). This paper is also related to the literature on identification of finite-dimensional parameters in discrete outcome models with bounded covariates.\footnote{E.g., magnac2007identification, chen2016informational, kline16, and lewbel2021semiparametric.} The main difference from that literature is that in my framework the distribution of latent variables (e.g., the random intercept) can be nonparametrically identified even if these latent variables have full support, but covariates are bounded. My approach is complementary to existing methods. Since as an input my framework requires the average structural demand function (i.e., the choice probability function) for one good, my results may be combined with the ones in berry2020nonparametric to nonparametricaly identify the distribution of unobserved individual level heterogeneity. Moreover, in situations where the researcher is not sure whether covariates have full support and is willing to impose mild restrictions because of tractability or data limitations, my approach can provide an additional reassurance of identification. Also, the results in this paper provide a more solid econometric foundation to the models with at least one normally distributed random coefficient (e.g. nevo2000practitioner). The paper is organized as follows. In Section (ref), I describe the setting. Sections (ref) and (ref) provide two identification results. I show how my identification results can be extended to bundles model in Section (ref). In Sections (ref) and (ref), I propose a new estimator of the finite-dimensional parameters and evaluate its performance in simulations. Section (ref) provides an empirical illustration. Section (ref) concludes. All proofs can be found in Appendix (ref). Appendix (ref) provides additional simulation evidence.

Multinomial Choice

Consider the following random coefficients model. The agent maximizes (indirect) utility by choosing between $J$ inside goods (e.g., different brands of cereals) and an outside option of no purchase. The choice set is denoted by $Y=\{0,1,\dots, J\}$. I normalize the utility from alternative $y=0$ to $0$. The random utility from choosing an alternative $y\neq 0$ is\footnote{Deterministic vectors are denoted by lower-case regular font Latin letters (e.g., $x$) and random objects by bold letters (e.g., $\mathbf{x}$). Capital letters are usually used to denote supports of random variables (e.g., $\mathbf{x}\in X$). I denote the support of a conditional distribution of $\mathbf{x}$ conditional on $\mathbf{z}=z$ by $X_z$. The cumulative distribution function (c.d.f.) and the probability density function (p.d.f.) of $\mathbf{x}$ are denoted by $F_{\mathbf{x}}$ and $f_{\mathbf{x}}$. $F_{\mathbf{x}|\mathbf{z}}$ ($f_{\mathbf{x}|\mathbf{z}}$) denotes the c.d.f. (p.d.f.) of $\mathbf{x}$ conditional on $\mathbf{z}=z$. }

align*[align* omitted — 172 chars of source]

where $\mathbf{z}_y\in Z_y\subseteq{\mathds{R}}$ is a product-specific observed covariate that can be different for different consumers (e.g., fiber content or price); $\mathbf{d}\in D\subseteq{\mathds{R}}$ is observed (demographic) individual-specific taste shifter (e.g., age or income); $\mathbf{w}\in W\subseteq{\mathds{R}}^{d_w}$ is a vector of all other observable covariates, which may include the rest of product/agent characteristics; $\mathbf{e}\in E\subseteq{\mathds{R}}$ is a latent taste shock. The latent random vector $\boldsymbol{\varepsilon}=(\boldsymbol{\varepsilon}_{y})_{y\in Y\setminus\{0\}}$ captures all other sources of unobserved heterogeneity (e.g., $\boldsymbol{\varepsilon}_y=\boldsymbol{\theta}^{{\sf T}}\mathbf{w}_y+\boldsymbol{\epsilon}_y\:\ensuremath{\mathrm{a.s.}}$, where $\boldsymbol{\theta}$ and $\boldsymbol{\epsilon}_y$ are random coefficients). The observed covariates are $\mathbf{x}=(\mathbf{d}, \mathbf{z},\mathbf{w})$, where $\mathbf{z}=(\mathbf{z}_{y})_{y\in Y\setminus\{0\}}$. The random coefficient $\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}$ represents individual specific heterogeneous tastes associated with the product characteristic $\mathbf{z}_{y}$ (i.e., the marginal utility from the product characteristic $\mathbf{z}_{y}$). This specification of random coefficients is common in applied work (see, for instance, berry1995automobile, nevo2000practitioner,nevo2001measuring, berry2004differentiated). The functions $\beta_0,\beta_1:W\to{\mathds{R}}$ are unknown to the researcher and $\beta_1(w)\neq 0$ for all $w\in W$. I assume that $\mathbf{d}$ (and $\beta_1(\mathbf{w})$) is scalar without loss of generality since if $\mathbf{d}$ is a vector, then all components of it but one can be absorbed by $w$. In this case, one would need to use variation in those absorbed components to identify the coefficients in front of them. Similarly to the existing treatment of random coefficients model, I assume that the random coefficients in front of $\mathbf{z}_{y}$ are the same for each alternative $y$. However, I do not impose sign restrictions on $\big[\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}\big]$.\footnote{ Since $\Pr(\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}> 0|\mathbf{x}=x)=1-F_{\mathbf{e}|\mathbf{x}}(-\beta_0(w)-\beta_{1}(w)d|x)$ and there are no restrictions on $\beta_0(\cdot)$, the random coefficient $\big[\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}\big]$ can be positive (negative) with probability that is arbitrarily close to $1$ if the support of $\mathbf{e}$ conditional on $\mathbf{x}=x$ is unbounded.} I start by stating two assumptions that will be used throughout the paper. The first one is a data requirement, the second one is a shape constraint on the distribution of latent variables.

assumption[Data] The analyst can identify $p_0(x)=\Pr(\mathbf{y}=0|\mathbf{x}=x)$ for all $x\in X$

Assumption (ref) implies that I only need to observe whether a consumer bought a product or not without knowing the identity of the product (see also, for instance, thompson1989identification,lewbel2000semiparametric,fox2012random).\footnote{The outcome $y=0$ can be replaced by any outcome. In this case, one will just need to renormalize the utility from that outcome to zero.} If the information on the identity of the purchases is also available, then this information (i) may improve the efficiency of an estimator; (ii) can help to satisfy the assumptions needed for identification (e.g., in my empirical illustration, I use one product to identify the sign of $\beta_1$ and I use another one to estimate it); and (iii) can be used to weaken the assumption that the random slope coefficient $\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}$ is the same across inside goods.

assumption[Exclusion Restrictions] For all $w\in W$ \begin{enumerate} • $\boldsymbol{\varepsilon}$ is conditionally independent of $(\mathbf{e},\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$; • $\mathbf{e}$ is conditionally independent of $(\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$. \end{enumerate}

Assumption (ref) is an exclusion restriction that requires latent shocks $\mathbf{e}$ and $\boldsymbol{\varepsilon}$ to be independent of each other (condition (i)) and independent of excluded covariates $(\mathbf{d},\mathbf{z})$ (condition (ii)) after conditioning on $\mathbf{w}$. Assumption (ref) allows any form of dependence between $(\boldsymbol{\varepsilon},\mathbf{e})$ and nonexcluded covariates $\mathbf{w}$. That is, $\boldsymbol{\varepsilon}$ may contain latent product characteristics (e.g., unobserved quality) that can be correlated with nonexcluded covariates (e.g., market-product identifier).\footnote{Since, for identification and estimation, I require the average structural function $p_0$, some forms of endogeneity (i.e., correlation between $\mathbf{x}$ and $\boldsymbol{\varepsilon}$) can be addressed using suitable instruments and control function residuals as in blundell2004endogeneity (see also berry1994estimating, berry1995automobile, berry2014identification for identification of structural demand function using aggregate data and instruments). I leave the detailed analysis of this case for future research.} In general, since I only require the identification of the structural demand function $p_0$, one can use the results in berry2020nonparametric to identify $p_0$ and treat market-product level unobservables as a part of $\mathbf{w}$. Next, I provide two nonnested sets of conditions that allow for identification of $\beta_0$, $\beta_1$, and the distribution of $\mathbf{e}$ and $\boldsymbol{\varepsilon}$. In Section (ref), I impose no parametric assumptions on latent $\mathbf{e}$ and $\boldsymbol{\varepsilon}$ but assume some smoothness on the c.d.f. of $\boldsymbol{\varepsilon}$ and restrict the support of covariates. In Section (ref), I identify the model when $\mathbf{e}$ is normally distributed, without any additional restrictions on the distribution of $\boldsymbol{\varepsilon}$ and with minimal support restrictions on covariates.

Identification

Nonparametric Identification

assumptionFor all $w\in W$ \begin{enumerate} • Conditional on $\mathbf{w}=w$, $\mathbf{e}$ has mean zero and variance one; • $F_{\boldsymbol{\varepsilon}|\mathbf{w}}(\cdot|w)$ has bounded partial derivatives up to order $\kappa$ for some $\bar{y}$ and $\partial^l_{\varepsilon^l_{\bar{y}}}F_{\boldsymbol{\varepsilon}|\mathbf{w}}(\cdot|w)|_{\varepsilon=0}\neq 0$ for all $l\leq\kappa$; • There exists $d^*$ such that the support of $(\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$ contains $(d^*,0)$ with an open neighborhood. \end{enumerate}

Assumption (ref)(i) is a scale and location normalization. It restricts $\mathbf{e}$ conditional on $\mathbf{w}=w$ to have a finite expectation and a nonzero variance for all $w$. Assumption (ref)(ii) requires the conditional distribution $F_{\boldsymbol{\varepsilon}|\mathbf{w}}$ to be sufficiently smooth in one component of $\varepsilon$ in the neighborhood of zero and have different from zero higher order partial derivatives. Since $\mathds{E}\left[\,\boldsymbol{\varepsilon}\,\right]$ is not assumed to be zero, if, for instance, $\boldsymbol{\varepsilon}$ is multivariate normal with component-wise nonzero mean, then Assumption (ref)(ii) is automatically satisfied. It is also generically satisfied when at least one component of $\boldsymbol{\varepsilon}$ is independent of the others and has a type I extreme value distribution fox2012random. However, Assumption (ref)(ii) rules out cases when $\boldsymbol{\varepsilon}$ is a constant. Another example of violation of Assumption (ref)(ii) is when $\kappa$ is infinite and $F_{\boldsymbol{\varepsilon}|\mathbf{w}}$ is a polynomial function of any finite degree. (In Section (ref), I provide an alternative result that does not restrict $F_{\boldsymbol{\varepsilon}|\mathbf{w}}$.) Assumption (ref)(iii) requires the support of $\mathbf{z}$ to contain zero with some open neighborhood. Assumptions similar to Assumptions (ref)(ii)-(iii) are common in the literature on identification of random coefficients models (e.g., Assumptions 8 and 10 in fox2012random and Assumption 4 in allen2020identification).

propositionIf Assumptions (ref)- (ref) hold, then $\beta_0(w)$, $\beta_1(w)$, and $\mathds{E}\left[\,\mathbf{e}^l|\mathbf{w}=w\,\right]$, $0\leq l\leq\kappa$, are identified for all $w\in W$.

Identification of $\kappa\leq\infty$ moments of the conditional distribution of $\mathbf{e}$ conditional on $\mathbf{w}$ is often sufficient for nonparametric identification of it. For example, Assumption 7 in fox2012random uses the Carleman condition.\footnote{For more detailed discussion of the problem of identification of the distribution from its moments see, for instance, kleiber2013multivariate and references therein.} Thus, under minimal restrictions, I can nonparametically identify the conditional c.d.f. $F_{\mathbf{v}|\mathbf{x}}$, where $\mathbf{v}=\beta_0(\mathbf{w})+\beta_{1}(\mathbf{w})\mathbf{d}+\mathbf{e}$. To establish the next identification result I need the following definition.

definition[Bounded completeness] The family of distributions $\left\{F_{\mathbf{v}|\mathbf{x}}(\cdot|x),x\in X'\right\}$ is boundedly complete if \[ \forall x\in X',\:\int_{V}g(t)dF_{\mathbf{v}|\mathbf{x}}(t|x)=0\implies g(\mathbf{v})=0 \:\ensuremath{\mathrm{a.s.}}, \] for any bounded function $g$.

Completeness assumptions have been widely used in econometric analysis. Completeness is typically imposed on the distribution of observables (e.g., newey03). However, many commonly used parametric restrictions on the distribution of unobservables imply bounded completeness. For instance, it is satisfied for normal distributions and the Gumbel distribution.\footnote{For testability of the completeness assumptions see canay13.} Combining bounded completeness with the identified distribution of the index $\mathbf{v}$, I have the following result.

propositionIf $F_{\mathbf{v}|\mathbf{x}}$ is identified and $ \left\{F_{\mathbf{v}|\mathbf{x}}(\cdot|(d,z,w)),d\in D_{(z,w)}\right\} $ is boundedly complete for all $(z,w)$ in the support, then the above model inherits all identifying properties of the random coefficients model with utilities $ \mathds{1}\left(\,y\neq0\,\right)(\mathbf{r}_{y}+\boldsymbol{\varepsilon}_{y}). $ The vector $\mathbf{r}=(\mathbf{r}_{y})_{y\in Y\setminus\{0\}}$ is an observed covariate conditionally independent of $\boldsymbol{\varepsilon}=(\boldsymbol{\varepsilon}_y)_{y\in Y\setminus\{0\}}$ conditional on $\mathbf{w}=w$ with the conditional support $ R_w=\left\{r\in{\mathds{R}}^J\::\:r=v z,\: z\in Z_w, v\in V_w\right\}, $ where $V_w$ is the support of $\mathbf{v}$ conditional on $\mathbf{w}=w$. In particular, $F_{\boldsymbol{\varepsilon}|\mathbf{w}}$ is identified over $R_w$.

The proof of Proposition (ref) is similar to the proof of Theorem 11 in fox2012random. The main difference is that, instead of parametric restrictions, Proposition (ref) uses the interaction between $\mathbf{d}$ and $\mathbf{z}$. Proposition (ref) implies that the original random coefficient model can be represented in the “special-covariate-with-full-support” framework without assuming existence of such covariates. Moreover, if the set of directions that $z/\left\lVertz\right\rVert$ can cover is sufficiently rich and the support of $\mathbf{e}$ conditional on $\mathbf{w}=w$ is ${\mathds{R}}$, then $R_{w}={\mathds{R}}^{J}$ and all the identification results that require existence of special covariates with full support (e.g., lewbel2000semiparametric, berry2009nonparametric, gautier2015triangular, fox2016nonparametric, and fox2020note) can be applied. Combining the results in Propositions (ref) and (ref) with Theorem 1 in fox2020note, I can establish the following result.

corollaryFor all $y\neq 0$, let $\boldsymbol{\varepsilon}_y=\boldsymbol{\theta}^{{\sf T}}\mathbf{w}_y+\boldsymbol{\zeta}_y$, where $\boldsymbol{\theta}$ and $\boldsymbol{\zeta}=(\boldsymbol{\zeta}_y)_{y\inY\setminus\{0\}}$ are random coefficients, and $\mathbf{w}_y$ is the vector of product-$y$-specific covariates. Suppose \begin{enumerate} • The assumptions of Propositions (ref) and (ref) hold; • $R_{w}={\mathds{R}}^{J}$ for all $w\in W$; • $(\boldsymbol{\theta},\boldsymbol{\zeta})$ and $\mathbf{w}=(\mathbf{w}_y)_{y\inY\setminus\{0\}}$ are independent; • The support of $\mathbf{w}$ contains an open ball of dimensionality of $\mathbf{w}$; • $(\boldsymbol{\theta},\boldsymbol{\zeta})$ has finite absolute moments and its distribution is uniquely determined by its moments; \end{enumerate} then $\beta_0$, $\beta_1$, and the distribution of $(\mathbf{e}_1,\boldsymbol{\theta},\boldsymbol{\epsilon})$ are identified.

To the best of my knowledge, Corollary (ref) is the first result that establishes nonparametric identification of the whole distribution of the random coefficients in the multinomial choice environments without assuming the existence of special covariates. fox2012random,allen2020identification, and lewbel2021semiparametric also allow for bounded covariates. However, they either do not fully identify the distribution of the random intercept $\boldsymbol{\varepsilon}$ allen2020identification,lewbel2021semiparametric or impose parametric restrictions on it fox2012random.

Normal Taste Shock

assumptionFor all $w\in W$ \begin{enumerate} • Conditional on $\mathbf{w}=w$, $\mathbf{e}$ is a standard normal random variable; • there exists $(d^*,z^{*\sf T})$ in the interior of the support of $(\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$ such that $z^*_{y}>0$ for all $y\in Y$; • there exists $(d^{**},z^{**\sf T})$ in the interior of the support of $(\mathbf{d},\mathbf{z})$ conditional on $\mathbf{w}=w$ such that $p_0((d,z^{**},w))$ is neither an exponential nor an affine function of $d$ on some open set. \end{enumerate}

Assumption (ref)(i) requires $\mathbf{e}$ to be normally distributed with nonzero variance. With nonzero variance, the assumption that $\mathds{E}\left[\,\mathbf{e}^2\,\right]=1$ is just a scale normalization. The assumption is common in applied work (e.g., nevo2000practitioner,nevo2001measuring) and allows me to relax Assumptions (ref)(ii)-(iii). Assumption (ref)(ii) is only needed for identification of the sign of $\beta_1(w)$. Assumption (ref)(iii) means that if I fix all covariates but the one that shifts the random coefficient, then the probability of the default conditional on covariates is neither an affine nor an exponential function of this nonfixed covariate. Assumption (ref)(iii) is not very restrictive since it rules out only some exponential and linear probability models. Moreover, it is testable.

propositionIf Assumptions (ref), (ref), and (ref) hold, then \begin{enumerate} • $\beta_0(w)$ and $\beta_1(w)$ are identified for all $w\in W$; • The conditions of Proposition (ref) are satisfied. \end{enumerate}

The proof of the identification of $\beta_{0}$ and $\beta_{1}$ uses the multiplicative structure of $d$ and $z$, and properties of the standard normal p.d.f. Informally, note that \[ \beta_{0}(\mathbf{w})\mathbf{z}+\beta_{1}(\mathbf{w})\mathbf{d}\mathbf{z}+\mathbf{e}\mathbf{z}. \] Since $d$ and $z$ can be moved independently, I can use variation in $d$ while keeping $dz$ by varying $z$ to identify $\beta_{0}(w)$. Then, by varying $z$, I can identify $\beta_{1}(w)$. Proposition (ref)(ii) follows from $\beta_0(w)$ and $\beta_1(w)$ being identified and $\mathbf{e}$ being standard normal (i.e, $\beta_0(\mathbf{w})+\beta_1(\mathbf{w})\mathbf{d}+\mathbf{e}$ conditional on $\mathbf{x}=x$ generates a boundedly complete family of distributions). Note that the only restriction on $\boldsymbol{\varepsilon}$ needed for Proposition (ref) is the conditional independence assumption (Assumption (ref)). The random intercept $\boldsymbol{\varepsilon}$ is allowed be continuously or discretely distributed (e.g., it may be a constant). Hence, I can extend Theorem 2 in fox2016nonparametric to environments with bounded covariates.

corollaryFor all $y\neq 0$ let $\boldsymbol{\varepsilon}_y=\boldsymbol{\theta}_y(\mathbf{w})$, where $\boldsymbol{\theta}_{y}$ is a random function such that its realization $\theta_y$ is a map from $W$ to ${\mathds{R}}$. Suppose \begin{enumerate} • Assumptions of Proposition (ref) hold; • $R_{w}={\mathds{R}}^{J}$ for all $w\in W$; • $\boldsymbol{\theta}=(\boldsymbol{\theta}_y)_{y\neq 0}$ and $\mathbf{w}$ are independent; • The support of $\boldsymbol{\theta}$, $\Theta$, satisfies Assumption 4 in fox2016nonparametric; \end{enumerate} then $\beta_0$, $\beta_1$, and the distribution of $\boldsymbol{\theta}$ are identified.

Bundles

Note that since I do not assume independence among $\boldsymbol{\varepsilon}_{y}$ across $y$, the multinomial choice model I study covers some bundles models gentzkow2007valuing, dunker2017nonparametric,fox2017note. In particular, assume that there are $\tilde{J}$ goods and the agent can purchase any bundle consisting of these goods. The vector $\tilde{y}$ describes the purchasing decision of the agent. That is, $\tilde{y}\in\tilde{Y}=\{0,1\}^{\tilde{J}}$. For instance, $\tilde{y}=(0,1,0,1,0,\dots,0)$ corresponds to the case when the agent purchased a bundle of goods $2$ and $4$. The random utility from choosing bundle $\tilde{y}\neq 0$ is of the form

align*[align* omitted — 178 chars of source]

and the utility from buying nothing is zero. I can rewrite the above utilities from bundles as as the utilities form the multinomial choice problem since there are finitely ($2^{\tilde{J}}$) possible bundles. Indeed, I can enumerate them all with $y=0$ corresponding to $\tilde{y}=0\in{\mathds{R}}^{\tilde{J}}$ (i.e., $Y=\{0,1,2,\dots,2^{\tilde{J}}\}$) and define $z_y=\sum_{j=1}^{\tilde{J}} \tilde{y}_j \tilde{z}_{j}$. As a result, I can extend the conclusions of Theorem 1 in fox2017note to environments with bounded covariates

corollaryLet $J=2$ and \begin{align*} \boldsymbol{\varepsilon}_{(1,0)}&=\theta_1(\mathbf{w})+\boldsymbol{\epsilon}_1,\quad\quad \boldsymbol{\varepsilon}_{(0,1)}=\theta_2(\mathbf{w})+\boldsymbol{\epsilon}_2,\\ \boldsymbol{\varepsilon}_{(1,1)}&=\boldsymbol{\varepsilon}_{(1,0)}+\boldsymbol{\varepsilon}_{(0,1)}+\boldsymbol{\xi}\theta_3(\mathbf{w}), \end{align*} where $\theta_i(\cdot)$, $i=1,2,3$, are some unknown functions, and $(\boldsymbol{\epsilon}_1,\boldsymbol{\epsilon}_2,\boldsymbol{\xi})\in{\mathds{R}}^2\times{\mathds{R}}_{+}$. Suppose \begin{enumerate} • Assumptions of Propositions (ref) and (ref) or Proposition (ref) hold; • $R_w={\mathds{R}}$ for all $w$; • $(\boldsymbol{\epsilon}_1,\boldsymbol{\epsilon}_2)|\mathbf{w}=w$ has an everywhere positive Lebesgue density on its support for all $w\in W$; • $\mathds{E}\left[\,\boldsymbol{\epsilon}_i|\mathbf{w}=w\,\right]=0$ and $\mathds{E}\left[\,\boldsymbol{\xi}|\mathbf{w}=w\,\right]=1$ for all $w\in W$ and $i=1,2$, \end{enumerate} then $\theta_i(\cdot)$, $i=1,2,3$, and the c.d.fs $F_{\boldsymbol{\epsilon}_i|\mathbf{w}}$, $i=1,2$, and $F_{\boldsymbol{\xi}|\mathbf{w}}$ are identified.

Estimation of \texorpdfstring{$\beta$}{the Finite-Dimensional parameter}

Proposition (ref) constructively identifies $\beta_{0}$ and $\beta_{1}$. In this section, I use it to estimate these parameters. That is, I focus on the multinomial choice model with random coefficients with normally distributed $\mathbf{e}$.\footnote{Proposition (ref) also provides a constructive identification for $\beta_0$ and $\beta_1$. However, Assumption (ref)(iii) fails to hold in my illustrative application presented in Section (ref). Additionally, Proposition (ref) uses limits of derivatives of identifiable functions at a single point, thus, most likely, leading to a consistent estimator with nonparametric rate of convergence.} Moreover, to simplify the exposition, I assume that there are no nonexcluded covariates $\mathbf{w}$ (i.e., $\beta_0(\cdot)$ and $\beta_1(\cdot)$ are constant functions). Note that, even though $\beta_0$ and $\beta_1$ are finite-dimensional parameters and the distribution of $\mathbf{e}$ is assumed to be known, the model is still semiparametric since the distribution of $\boldsymbol{\varepsilon}$ is not parametric. The first ingredient of the estimator is a nonparametric estimator of $p_0(\cdot)=\Pr(\mathbf{y}=0|\mathbf{x}=\cdot)$, $\hat{p}_0(\cdot)$. Any consistent and smooth enough estimator $\hat{p}_0$ will deliver a consistent estimator of $\beta=(\beta_1,\beta_0)$.\footnote{The normality of $\mathbf{e}$ implies that $p_0$ has continuous derivatives of any order. See Appendix (ref) for details.} For concreteness, I work with the series estimator based on products of powers of components of $x=(d,z)$ (polynomial regressions). That is, given a sample of independent identically distributed (i.i.d.) observations on covariates and a binary random variable that indicates whether the product was purchased or not $\left\{\mathds{1}\left(\,\mathbf{y}^{(i)}=0\,\right),\mathbf{x}^{(i)}\right\}_{i=1}^n$, define

align*[align* omitted — 176 chars of source]

where $\psi^K(\cdot)$ is a vector of orthonormal basis functions based on products of powers of components of $x$, $\Psi=\left(\psi^K\left(\mathbf{x}^{(1)}\right),\psi^K\left(\mathbf{x}^{(2)}\right),\dots,\psi^K\left(\mathbf{x}^{(n)}\right)\right)^{{\sf T}}$, and $\left(\Psi^{{\sf T}}\Psi\right)^{-}$ is the Moore-Penrose generalized inverse. I assume that the sum of powers of components of $x$ in $\psi^K$ is monotonically increasing in $K$. The sign of $\beta_1$ can be trivially estimated from $\hat{p}_0$ since \[ \mathrm{sign}(\beta_1)=\mathrm{sign}\left(p_0((d',z))-p_0((d,z))\right)\mathrm{sign}(z_{y^*})\mathrm{sign}(d'-d) \] if $z\geq 0$ or $z\leq 0$ with $z_{y^*}\neq 0$. Hence, for simplicity I assume that $\beta_1>0$. The identification result in Proposition (ref) is constructive and provides a closed form expression for $\beta$ as a functional of $p_0$ (see Appendix (ref)). Given the nonparametric power series estimator $\hat{p}_0$, the plug-in estimator of $\beta$ is

align*[align* omitted — 881 chars of source]

where

align*[align* omitted — 290 chars of source]

Note that $\hat\beta$ is essentially a nonlinear function of sample averages of different derivatives of estimated $\hat{p}_0$. Following newey1994asymptotic, newey1997convergence, to achieve $\sqrt{n}$-consistency and asymptotic normality of the proposed estimator, I will have to establish existence of the Reisz representer of a particular directional derivative. Let

align*[align* omitted — 460 chars of source]

where $f_{\mathbf{x}}$ is the p.d.f. of $\mathbf{x}$, and $p_1$, $p_{11}$, $p_{111}$, and $p_{1111}$ are first, second, third, and forth derivatives of $p_0$ with respect to $d$, respectively.

assumption\begin{enumerate} • The support of $\mathbf{x}$, $X$, is a Cartesian product of compact connected nonsingleton intervals in ${\mathds{R}}$. • $f_{\mathbf{x}}$ is bounded away from zero on the interior of $X$; • $f_{\mathbf{x}}$, $\partial_{d}f_{\mathbf{x}}$, $\partial_{z_{y}}f_{\mathbf{x}}$, and $\partial^2_{d^2}f_{\mathbf{x}}$ equal to zero at the boundary of $X$ for all $y$; • $\mathds{E}\left[\,\bar{v}(\mathbf{x})\bar{v}(\mathbf{x})^{{\sf T}}\,\right]$ is finite and nonsingular. \end{enumerate}

Assumptions (ref)(i)-(ii) are standard in the literature on nonparametric estimation of conditional expectations. Similarly to the average derivative estimator of powell1989semiparametric, to achieve $\sqrt{n}$-consistency the estimator I need to impose restrictions on the behavior of $f_{\mathbf{x}}$ on the boundary of its support. Since powell1989semiparametric work with the first derivative they only require $f_{\mathbf{x}}$ to vanish on the boundary. My estimator involves derivatives up to order 3, thus, leading to Assumption (ref)(iii). Assumption (ref)(iv) is the mean-square continuity condition that requires the variance of the score function of $\mathbf{x}$ (i.e $\log f_{\mathbf{x}}$) and derivatives of it to be finite. The following proposition establishes asymptotic normality of my estimator and is based on Theorem 6 in newey1997convergence. Denote \[ G=\left(

array[array omitted — 46 chars of source]

\right)\left(

array[array omitted — 210 chars of source]

\right)^{-1}, \]

propositionIf (i) $\left\{\mathds{1}\left(\,\mathbf{y}^{(i)}=0\,\right),\mathbf{x}^{(i)}\right\}_{i=1}^n$ are i.i.d.; (ii) Assumptions (ref), (ref) and (ref) are satisfied, and Assumption (ref)(iii) is satisfied for all $x^{**}=(d^{**},z^{**})\in X$; (iii) $K^6/n\to_{n\to\infty}0$, then \[ \sqrt{n}(\hat\beta-\beta)\to_{d} \mathrm{N}(0,V), \] where $V=G\mathds{E}\left[\,\bar{v}(\mathbf{x})\bar{v}(\mathbf{x})^{{\sf T}} p_0(\mathbf{x})(1-p_0(\mathbf{x}))\,\right]G^{{\sf T}}$.

In the proof of Proposition (ref), I also provide a consistent estimator of the asymptotic variance matrix $V$ that is based on the estimator proposed in newey1997convergence. I conclude this section by noting that after $\beta$ is estimated, one can construct a sieve maximum-likelihood estimator of $F_{\boldsymbol{\varepsilon}}$ since \[ \Pr(\mathbf{y}=0|\mathbf{x}=x)=\int_{\mathds{R}} F_{\boldsymbol{\varepsilon}}(tz_1,tz_2,\dots,tz_{J})\phi \left(t+\beta_0+\beta_1d\right)dt \] where $\phi(\cdot)$ is the standard normal p.d.f. Thus, one can find the maximizer of

align*[align* omitted — 373 chars of source]

where $\{\mathcal{F}_n\}_{n=1}^{\infty}$ is a sequence of sieve spaces for $F_{\boldsymbol{\varepsilon}}$. Inference on known functionals of $\beta$ and $F_{\boldsymbol{\varepsilon}}$ (e.g., counterfactuals) can be done using likelihood-ratio type statistic (see, for instance, shen2005sieve,chen2014sieve).\footnote{Both $\beta$ and $F_{\boldsymbol{\varepsilon}}$ can be estimated in one step by the sieve maximum-likelihood estimator. In this case, however, the estimator of $\beta$ may not be $\sqrt{n}$ consistent.}

Monte-Carlo Simulations

In this section, I assess the performance of my estimator in finite samples. I consider the binary choice model: \[ \mathbf{y}=\mathds{1}\left(\,(\beta_0+\beta_1\mathbf{d}+\mathbf{e})\mathbf{z}+\beta_3+\boldsymbol{\varepsilon}\geq 0\,\right), \] where $\beta_0=-0.5$, $\beta_1=1$, and $\mathbf{e}$ is a standard normal random variable. The random intercept $\beta_3+\boldsymbol{\varepsilon}$ is independent from $\mathbf{x}$ and $\mathbf{e}$ with mean $\beta_3=0.5$. The observed covariates $\mathbf{x}=(\mathbf{d},\mathbf{z})$ are distributed according to a monotone transformation of a bivariate normal distribution: $\mathbf{x}=5(\mathrm{arctan}(\mathbf{\tilde{x}})/\pi+0.5)$, where $\mathbf{\tilde{x}}$ is a mean-zero normal random vector such that each component of it has variance 1 and the correlation between components is $0.1$. Note that $\mathbf{x}$ has bounded support. I consider several data generating processes (DGPs). The first one (DGP-$0$) is when $\boldsymbol{\varepsilon}$ is a standard normal random variable. The next five DGPs correspond to $\boldsymbol{\varepsilon}$ being an equally weighted mixture of three unit-variance normal distributions with mean $-t$, $0$, and $t$ for $t\in\{1,2,3,4,5\}$ (DGP-$t$). For every $t$ the distribution of $\boldsymbol{\varepsilon}$ is symmetric. However, the variance is growing with $t$ and the distribution changes from a unimodal distribution to a distribution with three modes. Finally, DGP-L corresponds to the case with logistically distributed $\boldsymbol{\varepsilon}$. Each experiment is conducted $1000$ times for every DGP for 3 sample sizes $n\in\{10^3,5\cdot10^3,10^4\}$. I use a tensor product of cubic polynomials in estimation of the conditional probability $p_0$.\footnote{The results are qualitatively the same for higher order polynomials.} The results for the mean deviation (bias) of the estimator of $\beta_1$ are presented in Table (ref). As expected, the bias decreases with the sample size.\footnote{The mean absolute deviation of the estimator also decreases with the sample size. See, Appendix (ref) for further details.} However, there is not much variation across DGPs.\footnote{For comparison of my estimator with two alternative potentially misspecified parametric estimators, see Appendix (ref).}

table[table omitted — 540 chars of source]

Illustrative Empirical Application

To illustrate the empirical importance of the relaxation of the parametric assumptions about the distribution of $F_{\boldsymbol{\varepsilon}}$ and the proposed estimation procedure, I analyze margarine purchasing decisions of households from Springfield, MO, USA, using the multinomial choice model with normally distributed $\mathbf{e}$. I find substantial differences between estimates obtained by employing my semiparametric estimator and a fully parametric multinomial-logit-type estimator.

Data

The original dataset, constructed by allenby1991quality, is a panel of 9196 purchases of 10 brands of stick and tube margarine by 517 households from Springfield, MO, USA, extracted from an ERIM (A.C. Nielsen) scanner dataset. The dataset contains information on the shelf prices of each brand that is constructed using the actual price paid and the value of any redeemed coupon. The household demographics contain information on the household income.\footnote{See allenby1991quality for specific details of the dataset construction.} benoit2016outlier focused on 5 brands instead of 10 and transformed this dataset to a cross-section with 242 households. In particular, every observation contains only information on the household annual income, which I use as the agent-specific covariate $d$, agent choices ($y$), and product-specific prices $p_y$.\footnote{Income and prices are measured in thousands of US dollars and US dollars, respectively.} There are 5 brands: Generic ($y=0$), Blue Bonnet ($y=1$), House Brand ($y=2$), Shed Spread ($y=3$), and Fleischmann's ($y=4$). Income varies from 2.5k to 130k, with the median and average income being 26.75k and 22.5k, respectively. Table (ref) summarizes the share and price information for different products. There is a variation in prices across brands with Generic being on average the cheapest and Fleischmann's being the most expensive. At the same time, Fleischmann's is the least demanded product.

table[table omitted — 671 chars of source]

Utility

I follow nevo2000practitioner,nevo2001measuring and model the utility from purchasing brand $y\in\{0,1,2,3,4\}$ as \[ \boldsymbol{\delta}\mathbf{d}+(\beta_0+\beta_1 \mathbf{d}+\mathbf{e}) \mathbf{p}_{y}+\tilde{\boldsymbol{\varepsilon}}_y. \] The random coefficient $\boldsymbol{\delta}$ captures the direct marginal effect of income on utility from consumption of margarine (i.e., it is the same for all brands). The coefficient $\beta_0+\beta_1 \mathbf{d}$ can be thought of as the average marginal utility with respect to price. It captures the sensitivity of agents with respect to prices and is expected to be negative. Agents with different incomes may react differently to price changes. Note that no assumptions are made about $\tilde{\boldsymbol{\varepsilon}}_y$ (e.g., it is not assumed that it is has zero mean).\footnote{Estimation using $\log(\mathbf{p}_y)$ instead of $\mathbf{p}_y$ gives qualitatively similar results.} This utility specification correspond to the “preference shifter” specification in griffith2018income. There is no information about those who did not purchase any margarine products, thus, I analyze the choices of those who already decided to purchase a margarine product. If I treat the utility from consuming Generic brand as the baseline utility and subtract it from all utilities, the normalized utility from purchasing different brands for $y=1,2,3,4$ is \[ (\beta_0+\beta_1 \mathbf{d}+\mathbf{e}) [\mathbf{p}_{y}-\mathbf{p}_{0}]+\tilde{\boldsymbol{\varepsilon}}_y-\tilde{\boldsymbol{\varepsilon}}_0, \] and the utility from purchasing Generic brand is $0$. Hence, I can define $\mathbf{z}_{y}=\mathbf{p}_{y}-\mathbf{p}_{0}$ and $\boldsymbol{\varepsilon}_y=\tilde{\boldsymbol{\varepsilon}}_y-\tilde{\boldsymbol{\varepsilon}}_0$, $y=1,2,3,4$, where $\mathbf{p}_0$ is the price of Generic margarine. Given that I am considering margarine products, it is not surprising that the support for $\mathbf{z}_y$ is far from being full. In particular, $\max_{y}\max_{i}z^{(i)}_y=0.78$ and $\min_{y}\min_{i}z^{(i)}_y=-0.15$. At the same time, there is still variation in relative prices $\mathbf{z}_y$ and income $\mathbf{d}$. This variation allows me to recover $\beta$ without specifying the distribution of $\boldsymbol{\varepsilon}$. In the current application, I use a minimal amount of information: there are only two covariates. If one has more demographic and product data, it can be easily incorporated into the current framework via $\mathbf{w}$. For instance, $\mathbf{w}$ may contain nonprice marketing variables, packet size dummies, saturated fat content, household size, age of the household head, household location (e.g. zip-code).

Parametric Estimation

First, I assume the most common parametric specification for the random intercept -- multinomial logit. Formally, I estimate the following specification for normalized utility: \[ \mathds{1}\left(\,y\neq0\,\right)[\gamma_y+(\beta_0+\beta_1 \mathbf{d}+\mathbf{e}) \mathbf{z}_{y}]+\alpha\boldsymbol{\varepsilon}_y, \] where $\{\boldsymbol{\varepsilon}_y\}_{y=0}^4$ are i.i.d. Gumbel across $y$ that are also independent from $\mathbf{x}=(\mathbf{d},\mathbf{z})$; $\mathbf{e}$ is a standard normal random variable. (Parameter $\alpha$ captures the scale of $\boldsymbol{\varepsilon}_y$ since the variance of $\mathbf{e}$ is set to 1.) Although, price $\mathbf{p}_y$ is probably correlated with unobserved part of the utility $\tilde{\boldsymbol{\varepsilon}}_y$ (e.g., unobserved quality), the price difference $\mathbf{z}_y=\mathbf{p}_y-\mathbf{p}_0$ may be independent from $\tilde{\boldsymbol{\varepsilon}}_y-\tilde{\boldsymbol{\varepsilon}}_0$. The estimates of $\beta_0$ and $\beta_1$ are $\bar{\beta}_0=-6331.94$ (standard error$=17.19$) and $\bar{\beta}_1=-19.69$ (standard error$=514.48$), respectively. As expected, the sign of $\bar{\beta}_0$ is negative. The coefficient in front of the income variable, $\bar{\beta}_1$, is negative and not significant at the $5$ percent significance level. Although income does not matter much, the overall sensitivity to prices (mostly captured by $\bar\beta_0$ in this case) is substantial. The effect of income on marginal disutility from the price increase is not surprising given that margarine constitutes a small share of household expenditures on groceries.\footnote{E.g., in UK households spend about one percent of their grocery expenditures on margarine and butter griffith2018income.}

Semiparametric Estimation

Next, I apply the estimator proposed in Section (ref). Formally, I estimate the following specification for normalized utility: \[ \mathds{1}\left(\,y\neq0\,\right)[(\beta_0+\beta_1 \mathbf{d}+\mathbf{e}) \mathbf{z}_{y}+\boldsymbol{\varepsilon}_y], \] where $\mathbf{e}$ is a standard normal random variable. The random intercept $\boldsymbol{\varepsilon}=(\boldsymbol{\varepsilon}_y)_{y=1,2,3,4}$ is assumed to be independent from $\mathbf{x}$. There are no other restrictions on the joint distribution of $\boldsymbol{\varepsilon}$. This specification nests the logit specification estimated in the previous section. Hence, if the assumptions of multinomial logit are correct, then the results of parametric and semiparametric estimators should not differ much. The estimates of $\beta_0$ and $\beta_1$ are $\hat\beta_0=-39.1$ (standard error$=43.8$) and $\hat\beta_1=-16.7\times 10^{-3}$ (standard error$=3.97\times 10^{-6}$). \footnote{I use the tensor product of the $4$-th degree Chebyshev polynomials for $d$ and the $1$-st degree Chebyshev polynomials for every $z_{y}$.} Similar to the multinomial logit estimator, the sign of $\hat\beta_0$ is negative. The coefficient in front of the income variable is negative and significant at the $5$ percent significance level. However, the maximal value that $\hat{\beta}_1\mathbf{d}$ can take in the sample is substantially smaller than $\hat\beta_0$ ($\max_i\left(\mathbf{d}^{(i)}\hat{\beta}_1 /\hat{\beta}_0\right)=0.055$, standard error$=0.062$). The latter indicates that, similarly to the fully parametric specification, income does not affect marginal disutility from price increase much. However, the estimate of $\beta_0$ is substantially lower than the one in the fully parametric case. This indicates that consumers may be less sensitive to price changes than one would think after estimating the logit-type model. Interestingly, the difference between the estimates obtained using the fully parametric logit estimator $\bar{\beta}$ and my semiparametric estimator $\hat{\beta}$ is substantial (e.g., $\bar\beta_1/\hat\beta_1>10^3$). That is, the parametric estimator overestimates the magnitude of the agents sensitivity to relative price changes of margarine. This suggests that the multinomial logit structure most likely fails to hold, emphasizing the importance of semiparametric estimation.\footnote{This empirical finding is in line with the simulation results, presented in Appendix (ref).}

Conclusion

This paper shows that commonly used exclusion restrictions and richness assumptions about the distribution of some unobservables may lead to full nonparametric identification in discrete outcome models even when covariates are bounded. The proposed identification framework extends the results from a large literature that uses special covariates with full support to environments where such full-support covariates are not available. It also leads to an asymptotically normal estimator of the finite-dimensional parameters of the model.