EconBase
← Back to paper

Consumer Theory with Non-Parametric Taste Uncertainty and Individual Heterogeneity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

111,255 characters · 28 sections · 132 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Consumer Theory with Non-Parametric Taste Uncertainty and Individual Heterogeneity

center[center omitted — 149 chars of source]
abstractWe introduce two models of non-parametric random utility for demand systems: the stochastic absolute risk aversion (SARA) model, and the stochastic safety-first (SSF) model. In each model, individual-level heterogeneity is characterized by a distribution $\pi\in\Pi$ of taste parameters, and heterogeneity across consumers is introduced using a distribution $F$ over the distributions in $\Pi$. Demand is non-separable and heterogeneity is infinite-dimensional. Both models admit corner solutions. We consider two frameworks for estimation: a Bayesian framework in which $F$ is known, and a hyperparametric (or empirical Bayesian) framework in which $F$ is a member of a known parametric family. Our methods are illustrated by an application to a large U.S. panel of scanner data on alcohol consumption.

Keywords: Consumer Theory, Scanner Data, Stochastic Demand, Taste Heteroge- neity, Non-Parametric Model, Bayesian Approach.

Introduction

The recent availability of databases containing all dated purchases made by a large nu- mber of consumers (28,036 in our application) presents a modern challenge for the eco- nometrics of demand systems, requiring new models and estimation approaches (see, for example, burda-2008, burda-2008, burda-2012, for discrete choice, and guha-ng, guha-ng, cher-newey, cher-newey, and dobronyi, dobronyi, for the first analyses of such data in the demand literature). This type of data is commonly called scanner data because its collection involves retailers or households scanning each purchased good on the date of purchase. This paper introduces two models of random utility for scanner data: the stochastic absolute risk aversion (SARA) model, and the stochastic safety-first (SSF) model. These models have the following advantages in comparison with the existing literature:

enumerate[(i)] • Both models are consistent with consumer theory: Every consumer maximizes a strictly increasing and strictly quasi-concave utility function. The latter prop- erty is not accommodated by existing approximations of the utility function like the quadratic approximation of the utility function (theil-neudecker, theil-neudecker; barten, barten), the translog utility model (johansen-relationship, johansen-relationship; transcendental, transcendental), or the Almost Ideal Demand System aids and its extensions (bbl, bbl; moschini, moschini). • Both models are non-parametric. In each model, the utility function is indexed by a functional parameter characterizing the individual heterogeneity, allowing for infinite-dimensional heterogeneity. In this respect, our paper differs from the existing literature when finite-dimensional heterogeneity is considered (see beckert-blundell, beckert-blundell, blom-kumar-more, blom-kumar-more, \hyper@link{cite}{cite.blundell\@extra@b@citeb}{Blundell, Horowitz, and Parey}, \hyper@link{cite}{cite.blundell\@extra@b@citeb}{2017}, and \hyper@link{cite}{cite.blundell-wp\@extra@b@citeb}{Blundell, Kristensen, and Matzkin}, \hyper@link{cite}{cite.blundell-wp\@extra@b@citeb}{2017}, for some examples of finite-dimensional restrictions). Our approach is in line with dette who write, “in general the multivariate demand function is a non-monotonic function of an infinite-dimensional unobservable---the individual's preference ordering.” • Both models yield demand functions with non-separable heterogeneity (see the discussions in brown, brown, beckert-blundell, beckert-blundell, and dette, dette). They are also endowed with precise structural interpretations, as heterogeneity is introduced by means of a distribution $\pi$ of taste parameters, so that we can imagine consumers facing \emph{taste uncertainty}, which they eliminate using expected utility. • Both models are \emph{identified} under weak restrictions. Identification follows from the use of panel data. Without such data, we lose identification (hn-individual-het, hn-individual-het). Of course, the structure of scanner data is extremely important.

Each model is characterized by a basis of functions. This basis is used to generate a family of utility functions. A distribution is, then, placed over this family. To be precise, we start with a basis of increasing and concave functions. Let $U(x;a)$ denote an element of this basis, where $x$ is a bundle and $a\in \mathscr{A}$ is a finite-dimensional vector of taste parameters. A family of utility functions is generated by taking the convex hull of the basis. Let $U(x;\pi)=\mathbb{E}_{\pi}\big[U(x;a)\big]$ denote an element of this family, where $\pi\in \Pi$ is a distribution on $\mathscr{A}$. This family is indexed by a functional parameter $\pi$, which can be structurally interpreted as taste uncertainty (resolved after the consumer makes her decisions). The heterogeneity across consumers is introduced using a distribution $F$ on the set $\Pi$ of probability distributions $\pi$ on $\mathscr{A}$. Therefore, each model combines uncertainty and heterogeneity: the uncertainty in taste for a given consumer is represented by $\pi$, and the heterogeneity across consumers is captured by $F$.

The paper considers a two-good framework with continuous support for $x$. It is organized as follows: Section (ref) introduces the stochastic absolute risk aversion (SARA) model and Section (ref) introduces the stochastic safety-first (SSF) model. For each mo- del, we derive conditions on $\Pi$ under which there exists a unique demand system, for each $\pi\in \Pi$. In Section (ref), the distribution of heterogeneity $F$ is introduced. When $F$ is known, we obtain a Bayesian framework in which the functional parameter $\pi\in \Pi$ has to be estimated. When $F$ is a member of a known parametric family, indexed by $\theta$, we obtain an empirical Bayesian framework with a hyperparameter $\theta$ that has to be estimated, and a stochastic functional parameter $\pi$ that has to be filtered. In Section (ref), we consider the identification of the taste distribution $\pi$ within each model. Next, we examine if it is possible to distinguish between stochastic risk aversion and stochastic safety-first. In Section (ref), we use the Nielsen Homescan Consumer Panel to illustrate our methodology in an application to the consumption of alcohol. Section (ref) concludes. The details of the Dirichlet process are in Appendix (ref); integrability is discussed in Appendix (ref); an optimization procedure for filtering the taste distributions $\pi$ after estimating $F$ is in Appendix (ref); details of the data are placed in Appendix (ref).

A Model with Stochastic Risk Aversion

This section introduces the first utility specification that we consider. It first describes the set of utility functions, then derives conditions under which there exists a unique demand system. The taste uncertainty is introduced using risk aversion parameters.

The Set of Utility Functions

There are two goods, denoted $1$ and $2$. Let $\bar{R}=\mathbb{R}_{+}^2$ denote the non-negative orthant with interior $R$. A consumer has preferences over the bundles in $\bar{R}$. Her preferences are summarized by a utility function of the form:

equation[equation omitted — 80 chars of source]

for every $x$ such that $x_1,x_2\geq 0$, where $A=(A_1,A_2)$ is a positive stochastic parameter characterizing the consumer's degrees of absolute risk aversion with respect to goods 1 and 2, and $\pi$ is a joint distribution for this pair of stochastic taste parameters. Her preferences are, as a result, contained in a broad family of utility functions, indexed by a functional parameter $\pi$. There are two interpretations of specification (ref): (i) the preferences are summarized by a deterministic utility function in the convex hull gen- erated by a parametric family, or (ii) the consumer faces “taste uncertainty” and she resolves this uncertainty by using expected utility. We call these preferences stochastic absolute risk aversion (SARA) preferences.\footnote{These preferences differ from those used to describe consumer behaviour when facing ambiguity or uncertainty, as in, say, halevy-feltkamp.}

If $\pi$ is a point mass at $a=(a_1,a_2)$ such that $a_1,a_2>0$, the stochastic parameters are constant, and $U(x;\pi)$ reduces to $U(x;a)=-\exp(-a'x)$. This function is strictly increasing because we have:\footnote{Here, $>0$ means each component is strictly larger than $0$.}

equation[equation omitted — 125 chars of source]

at each $x$ such that $x_1,x_2>0$, and concave (although not necessarily strictly concave) because the Hessian associated with the utility function:

equation[equation omitted — 167 chars of source]

is negative semi-definite, at each $x$ such that $x_1,x_2>0$. This matrix is related to a bivariate measure of absolute risk aversion\footnote{Such a measure can be defined as: \[ -\left(\text{diag}\frac{\partial U(x;\pi)}{\partial x}\right)^{-1/2}\frac{\partial^2 U(x;a)}{\partial x\partial x'}\left(\text{diag}\frac{\partial U(x;\pi)}{\partial x}\right)^{-1/2}, \] where $\text{diag}\frac{\partial U(x;\pi)}{\partial x}$ is the diagonal matrix whose diagonal elements are the first derivatives of $U(x;\pi)$.} (richard-1975, richard-1975; karni-1979, karni-1979, karni-1983; grant, grant). These properties translate into properties of the more general function: $U(x;\pi)$.

propositionIf preferences are SARA and the consumer's taste distribution $\pi$ is not the mixture of point masses $a,a'\in R$ where $a$ is proportional to $a'$, then the utility \makebox[\textwidth][s]{function $U(x;\pi)$ is strictly increasing with a negative definite Hessian everywhere on $R$.}
proofThe utility function $U(x;\pi)$ is strictly increasing on $R$ because: \begin{equation} \frac{\partial}{\partial x}\mathbb{E}_{\pi}\left[U(x;A)\right]=\mathbb{E}_{\pi}\left[\frac{\partial U(x;A)}{\partial x}\right]>0, \end{equation} at every $x$ such that $x_1,x_2>0$. Its Hessian is negative definite on $R$ because the sum of two 2-by-2 matrices of rank 1, whose columns are not proportional, has full rank.

Proposition (ref) implies that we have effectively constructed a family of well-behaved utility functions $\{U(x;\pi):\pi\in\Pi\}$ indexed by a functional parameter $\pi$, describing the taste uncertainty, instead of the standard finite-dimensional parameter usually con- sidered in the literature.

Let $g_{\pi}(\cdot)$ denote the function defined by the implicit equation:

equation[equation omitted — 45 chars of source]

for every $x_1\geq 0$, and each (attainable) level of utility $u<0$. This implicit equation has a unique solution because $U(x;\pi)$ is strictly increasing on $\bar{R}$. The function $g_{\pi}(\cdot,u)$ is the indifference curve associated with the functional parameter $\pi$ and a utility level of $u$---$g_{\pi}(\cdot,u)$ maps every value of $x_1$ to a value of $x_2$ for which $(x_1,x_2)$ attains a utility level of $u$ given $\pi$. The implicit function theorem implies that $g_{\pi}(\cdot)$ is twice-continuo- usly-differentiable with respect to $x_1$ and:

equation[equation omitted — 116 chars of source]

on $R$ where $\text{MRS}(x;\pi)\equiv \frac{\partial U(x;\pi)/\partial x_1}{\partial U(x;\pi)/\partial x_2}$ denotes the marginal rate of substitution at $x$---the rate at which the consumer is willing to exchange good 1 for good 2 given $x$ and $\pi$. The indifference curve $g_{\pi}(\cdot,u)$ is strictly convex such that:

equation[equation omitted — 68 chars of source]

at every $x_1>0$, since the Hessian of $U(x;\pi)$ is negative definite everywhere on $R$ (see Lemma 1 in dobronyi, dobronyi). This property is stronger than the sta- ndard assumption of strict quasi-concavity, which allows this derivative to be zero on a nowhere dense set (katzner, katzner). This distinction is important for what follows.

Note that, after integrating out the taste uncertainty, the absolute risk aversions will depend on the consumption level. For instance, when $A_1$ and $A_2$ are independent with distributions $\pi_1$ and $\pi_2$, the risk aversion for good 1 becomes:

equation[equation omitted — 178 chars of source]

where $U_1(x_1;\pi_1)$ denotes $\mathbb{E}_{\pi_1}[\exp(-A_1x_1)]$, the portion of the utility function $U(x;\pi)$ corresponding to good 1. Clearly, $A_1(x_1)$ depends on $x_1$. Indeed, it is the average of $A_1$ given the following modified density:

equation[equation omitted — 79 chars of source]

with respect to $\pi_1$.

The Demand Function

Let $z\in R$ denote a pair $z=(y,p)$ in which $y$ denotes expenditure and $p$ denotes the price of good 1, both normalized by the price of good 2. The consumer can purchase a bundle $x\in \bar{R}$ if, and only if, $px_1+x_2\leq y$. She chooses a bundle $x\in\bar{R}$ that solves:

equation[equation omitted — 107 chars of source]

Let $X^*(z;\pi)$ denote the solution to:

equation[equation omitted — 149 chars of source]

While (ref) is restricted to bundles in the non-negative orthant, (ref) allows for neg- ative values. The solution to (ref) is characterized by the following system of first- order conditions:

equation[equation omitted — 178 chars of source]

The first equality says that the marginal rate of substitution equals the relative price $p$. The second equality says that the budget constraint holds with equality. Equivalently, we can solve the equality:

equation[equation omitted — 99 chars of source]

for the first component $X_1^*(z;\pi)$, and then use the budget constraint in (ref) to solve for $X_2^*(z;\pi)$. As long as $A_1-pA_2$ is not almost surely equal to zero, the first-order partial derivative of the left side of this equality with respect to $x_1$ is strictly negative:

equation[equation omitted — 90 chars of source]

The function on the left side of (ref) is, therefore, strictly decreasing in $x_1$, implying that there exists a unique solution $X_1^*(z;\pi)$ to (ref), and a unique solution $X^*(z;\pi)$ to (ref). If $X^*(z;\pi)$ is in $\bar{R}$, then $X^*(z;\pi)$ coincides with the solution to (ref). Else, the solution to (ref) is on the boundary of $\bar{R}$. Let $X(z;\pi)$ denote the solution to (ref) given both $z$ and $\pi$. There are three regimes of demand in the design space:

equation[equation omitted — 222 chars of source]

Because the utility function $U(x;\pi)$ has strictly convex indifference curves everywhere on $R$, the demand function $X(z;\pi)$ is invertible in the second regime (see Proposition 2 in dobronyi, dobronyi).

propositionIf preferences are SARA and the consumer's taste distribution $\pi$ is not the mixture of point masses $a,a'\in R$ where $a$ is proportional to $a'$, then there exists a unique solution $X(z;\pi)$ to the maximization problem in (ref) given $z$ and $\pi$, for every $z\in R$, almost surely, for every $\pi$. There are three regimes of demand defined by (ref). The resulting demand function $X(z;\pi)$ is invertible in the second regime.

As a final remark, let us consider a risk-neutral consumer. In particular, let us ass- ume that $A_1$ and $A_2$ tend stochastically to zero, with means that tend to zero so that $\mathbb{E}_{\pi}[A_1]/\mathbb{E}_{\pi}[A_2]$ converges to a non-degenerate $a_0$. By considering the Taylor expansion of utility, it can be shown that these preferences are represented by:

equation[equation omitted — 37 chars of source]

This representation is unique up to an increasing transformation. For this risk-neutral consumer, goods are considered to be perfect substitutes. It is known that such a consumer will consume only good 1 whenever $p<a_0$, and only good 2 whenever $p>a_0$.

Gamma Taste Uncertainty

As an illustration, let us assume that $A_1$ and $A_2$ are independent and that $A_j$ has a Gamma distribution $\gamma(\nu_j,\alpha_j)$ with degree of freedom $\nu_j>0$ and scale factor $\alpha_j>0$, for $j=1,2$. Under this specification, $\pi=\gamma(\nu_1,\alpha_1)\otimes\gamma(\nu_2,\alpha_2)$, where $\otimes$ denotes the tensor product of distributions. By the Laplace transform of the Gamma distribution:

equation[equation omitted — 128 chars of source]

Under this specification, the absolute risk aversion for good 1 in (ref) becomes:

equation[equation omitted — 53 chars of source]

which is hyperbolic in $x_1$. The indifference curve $g_{\pi}(\cdot)$ associated with utility level $u$ is:

equation[equation omitted — 163 chars of source]

for every $x_1\geq 0$ and $u\in(-1,0)$ such that:

equation[equation omitted — 114 chars of source]

It is easily shown that the second derivative of the indifference curve $g_{\pi}(\cdot,u)$ equals:

equation[equation omitted — 120 chars of source]

for some $c>0$. This inequality confirms that the indifference curve $g_{\pi}(\cdot,u)$ is strictly convex. Furthermore, the MRS is equal to:

equation[equation omitted — 88 chars of source]

The unconstrained solution $X_1^*(z;\pi)$ to the first-order condition in (ref) is equal to:

equation[equation omitted — 157 chars of source]

The second component $X_2^*(z;\pi)$ is deduced from the budget constraint in (ref). By equation (ref), the demand function $X(z;\pi)$ coincides with $X^*(z;\pi)$ over the set $\mathcal{Z}$ of pairs $z$ such that:

equation[equation omitted — 110 chars of source]

The three regimes of demand are illustrated in Figure (ref) in the design space. The strict convexity of the indifference curve $g_{\pi}(\cdot,u)$ on $\mathcal{Z}$ implies that the demand function $X(\cdot;\pi)$ associated with this utility function is invertible on $\mathcal{Z}$.

figure[figure omitted — 1,483 chars of source]

A Model with Stochastic Safety-First

We now consider a model with taste parameters that have a safety-first interpretation.

The Set of Utility Functions

In Section (ref), we constructed a family of well-behaved utility functions by taking the convex hull generated by a particular basis. In this section, we consider another basis, consisting of functions with the form:

equation[equation omitted — 107 chars of source]

for every $x_1,x_2\geq 0$, where $x^+=\max\{0,x\}$ and $a_1,a_2>0$. This function corresponds to the “safety-first” criterion, introduced into the literature on portfolio management by roy-safety. In order to illustrate, let us consider the consumption of alcohol, as in dobronyi. Suppose that there are two groups of goods: group 1 consisting of drinks with low alcohol by volume such as beers and ciders, and group 2 consisting of drinks with high alcohol by volume such as wines and liquors. Assume that the quantities are measured in identical units such as volume of alcohol---that is, the total volume of the drink in litres multiplied by the alcohol by volume of the drink.\footnote{Quantities could be, alternatively, measured in calories.} We can, then, add these volumes to aggregate two drinks with different sizes and/or percentages of alcohol. Here, $a_1$ is the consumer's relative preference between the two groups of drinks, and $a_2$ is a “control” parameter, specifying her attempt to limit her intake of alcohol.

Now, let us introduce a distribution $\pi$ such that $\mathbb{E}_{\pi}[A_j]<\infty$, $j=1,2$, and define:

equation[equation omitted — 84 chars of source]

By the law of iterated expectations, we obtain:

equation[equation omitted — 132 chars of source]

We call these preferences stochastic safety-first (SSF) preferences.

Under mild regularity conditions:

align[align omitted — 441 chars of source]

for every $x$ such that $x_1,x_2>0$. These partial derivatives are strictly positive when $\pi$ has full support: $\pi(a_1,a_2)>0$, for $a_1,a_2>0$. By taking the second-order derivatives:

equation[equation omitted — 235 chars of source]

for every $x$ such that $x_1,x_2>0$, where $\pi_0\equiv\pi(x_1+A_1x_2|A_1)$ in which $\pi(\cdot|A_1)$ denotes the conditional density of $A_2$ given $A_1$, assuming that such a density exists. This matrix is both symmetric and negative definite when $\pi(\cdot|A_1)$ is continuous and $A_1$ is not constant. This result follows from the positivity of $\mathbb{E}_{\pi}[\pi_0]$ and the following equality:

equation[equation omitted — 123 chars of source]

which holds for every $x$ such that $x_1,x_2>0$, in which $\tilde{\pi}$ denotes the modified density:

equation[equation omitted — 90 chars of source]
propositionIf preferences are SSF and the consumer's taste distribution $\pi$ is con- tinuous with full support given $A_1$, then the utility function $U(x;\pi)$ is strictly increasing with a negative definite Hessian everywhere on $R$.

Consequently, we have constructed another family of well-behaved utility functions $\{U(x;\pi):\pi\in\Pi\}$ indexed by a functional parameter $\pi$, describing taste uncertainty.

The Demand Function

Let us revisit the utility maximization problem in (ref). Under the safety-first spec- ification, the analogue of the unconstrained first-order condition in (ref) is given by:

equation[equation omitted — 99 chars of source]

We obtain this equality by equating the marginal rate of substitution with the relative price $p$, and then using the budget constraint to replace $x_2$ with $y-px_1$. Under the regularity conditions from above, the left-hand side is strictly monotone in $x_1$ given $\pi$, so that there exists a unique solution to the first-order condition. As in Section (ref), we let $X_1^*(z;\pi)$ denote this solution, and let $X_2^*(z;\pi)$ denote the quantity $y-pX_1^*(z;\pi)$.

propositionIf preferences are SSF and the consumer's taste distribution $\pi$ is con- tinuous with full support given $A_1$, then there exists a unique solution $X(z;\pi)$ to the maximization problem in (ref) given $z$ and $\pi$, for every $z\in R$, almost surely, for every $\pi$. There are three regimes of demand defined by (ref). The resulting demand function $X(z;\pi)$ is invertible in the second regime.

When the consumer's preferences are SSF, the MRS has the form:

equation[equation omitted — 201 chars of source]

Thus, the rate at which the consumer is willing to exchange good 1 for good 2 given $x$ and $\pi$ is equal to the inverse of the expectation of her relative preference between goods $A_1$, conditional on not surpassing her control parameter $A_2$.

Some functionals of the distribution $\pi$ can be especially interesting. For instance, in an application to the consumption of alcohol, we might expect the conditional distribution of $A_2$ given $A_1=a_1$ to be concentrated around a single mode, characterizing an implicit alcohol limit for this consumer. Then, we can ask the following questions:

enumerate[(i)] • Is this limit positively correlated with $A_1$? In other words, is there a positive relationship between this limit and a preference for strong alcoholic beverages? • Does a change in the maximum blood alcohol level for driving affect this limit?

These are questions that cannot be answered using classical demand systems like the Almost Ideal Demand System aids. In fact, tests based on the Almost Ideal Demand System have rejected rationality in applications to alcohol consumption all-ferg-stew. Clearly, it is possible that the Almost Ideal Demand System is misspecified.

Exponential Threshold Taste Uncertainty

In general, the first-order condition in (ref) has no closed-form solution. However, its expression can be simplified for some taste distributions $\pi$. As an illustration, let us assume that:

enumerate[(i)] • $A_1$ and $A_2$ are independent. • $A_2$ follows an exponential distribution $\gamma(1,\lambda)$ with survival function: \begin{equation} P(A_2>a_2) = \exp(-\lambda a_2). \end{equation} • $A_1$ follows a distribution with Laplace transform: $\Psi(v)=\mathbb{E}[\exp(-vA_1)]$, $v\geq 0$.

Under this specification, we can first integrate with respect to $A_2$ within the expecta- tion in (ref) in order to obtain the following condition:

equation[equation omitted — 86 chars of source]

Equivalently, we obtain:

equation[equation omitted — 133 chars of source]

This equation can be written in terms of the Laplace transform $\Psi$ for $A_1$. This yields:

equation[equation omitted — 92 chars of source]

which can also be written as:

equation[equation omitted — 77 chars of source]

Finally, by inverting this expression and rearranging the terms, we get:

equation[equation omitted — 138 chars of source]

The second component $X_2^*(z;\pi)$ of the unconstrained solution in (ref) is deduced from the budget constraint. It follows from equation (ref) that the demand function $X(z;\pi)$ coincides with $X^*(z;\pi)$ if, and only if:

equation[equation omitted — 111 chars of source]

For instance, if $A_1$ follows a gamma distribution $\gamma(\nu,\alpha)$, then $\log \Psi(v)=-\nu\log(1+v/\alpha)$, and we obtain:

equation[equation omitted — 64 chars of source]

for $v\geq 0$. Moreover, by inverting this function, we get:

equation[equation omitted — 99 chars of source]

Therefore, the solution $X_1^*(z;\pi)$ has the form:

equation[equation omitted — 88 chars of source]

and demand $X(z;\pi)$ coincides with $X^*(z;\pi)$ if, and only if:

equation[equation omitted — 58 chars of source]

The regimes of demand are illustrated in Figure (ref) in the design space. Note, we can also verify that the Slutsky coefficient is strictly negative\footnote{This property holds for any Laplace transform $\Psi$ of $A_1$ (see Appendix (ref)).} such that:

equation[equation omitted — 150 chars of source]

ensuring that the demand function $X(\cdot;\pi)$ is invertible over the set $\mathcal{Z}$ of pairs $z$ on which demand is strictly positive (see Section 2 in dobronyi, dobronyi).

figure[figure omitted — 1,569 chars of source]

Individual Heterogeneity

Sections (ref) and (ref) introduced two utility specifications, both indexed by the functional parameter $\pi$. Of course, different consumers can have different functional parameters. This individual heterogeneity is introduced in a second layer, by specifying a distribution $F$ over the set $\Pi$ of distributions on $R$, such as the Dirichlet process (see, for example, navarro, navarro, for an application of the Dirichlet process in modelling individual differences). More precisely, we make the following theoretical assumption:

assumptiona[Latent Stochastic Model] \phantom{1} \begin{enumerate}[(i)] • There are $n\geq 1$ consumers. • Consumers are segmented into $M$ homogeneous groups. • Consumers in group $m$ have the utility function $U(x;\pi_m)$, for all $m=1,\dots,M$. • The taste parameters $(\pi_m)$ are independently drawn from a Dirichlet process $F$. \end{enumerate}

Assumption (ref) introduces a distribution $F$ over the functional taste parameter $\pi$. This distribution $F$ characterizes the heterogeneity across homogeneous groups. It can encompass, for example, regional or demographic differences in preferences. This infinite-dimensional heterogeneity is non-separable in the stochastic demand equation.

The Dirichlet process can be constructed in three steps:

steps• Consider the set of (Bernoulli) distributions on $\{0,1\}$. This set is characterized by $q\in \bar{R}$ such that $q_1+q_2=1$. A distribution defined on this set of distributions is a distribution defined on this parameter set. We can, for instance, introduce a beta distribution, denoted $B(\alpha_1,\alpha_2)$. The distribution $B(\alpha_1,\alpha_2)$ has a continuous density: \begin{equation} f(q)=\frac{\Gamma(\alpha_1+\alpha_2)q_1^{\alpha_1}q_2^{\alpha_2}}{\Gamma(\alpha_1)\Gamma(\alpha_2)}, \end{equation} with respect to the Lebesgue measure over the simplex $\{(q_1,q_2)\geq 0:q_1+q_2=1\}$, where $\Gamma$ denotes the gamma function,\footnote{The gamma function $\Gamma$ is defined by $\Gamma(\alpha)=\int_0^{\infty}\exp(-x)x^{\alpha-1}dx$, for each $\alpha>0$.} and $\alpha_1,\alpha_2>0$ are positive scalar parameters. • The beta distribution can be extended to define a distribution on the set of discrete distributions with weights $q_j\geq 0$, $j=1,\dots,J$, such that $\sum_{j=1}^Jq_j=1$. This procedure leads to the Dirichlet distribution, denoted $D(\alpha)$. The res- ulting distribution $D(\alpha)$ has continuous density: \begin{equation} f(q)=\frac{\Gamma\big(\sum_{j=1}^J\alpha_j\big)\prod_{j=1}^J q_j^{\alpha_j}}{\prod_{j=1}^J\Gamma(\alpha_j)}, \end{equation} with respect to the Lebesgue measure over the simplex: \begin{equation} \left\{q\in\mathbb{R}_{+}^J:$\sum_{j=1}^Jq_j=1$ and $q_j\geq 0$, $\forall j$\right\}, \end{equation} (see, for example, kotz, kotz, page 485, and lin-dirichlet, lin-dirichlet, for details). • Then, the Dirichlet distribution can be extended to define a distribution on a large set of distributions\footnote{The realizations of a Dirichlet process are, almost surely, discrete distributions. Although we assumed continuity to prove the existence of a unique demand system in Section (ref), these realizations can approximate any continuous distribution. This discrepancy has no practical implications.} defined on $\bar{R}$ (see Appendix (ref)). This procedure leads to the Dirichlet process. The Dirichlet process is characterized by a distribution $\mu$ on $\bar{R}$ and a scaling parameter $c>0$. The distribution $\mu$ can be thought of as the mean of the Dirichlet process, while the parameter $c$ manages its degree of discretization (see Appendix (ref)). This extension of the Dirichlet distribution is much more complicated than the Dirichlet distribution, especially because the notion of the Lebesgue measure on the set of distributions, and the notion of a density, no longer exist (see ferguson, ferguson, rolin, rolin, and sethuraman, sethuraman).

Let us now discuss implications of Assumption (ref): If the functional and scaling parameters of the Dirichlet process are known, then we are in a Bayesian framework (see, for example, geweke, geweke, for a Bayesian analysis of revealed preference) in which the taste distribution $\pi\in \Pi$ has to be estimated. Otherwise, we can assume that the mean $\mu$ of our process $F$ is characterized by a finite-dimensional hyperparameter $\theta$. Naturally, the hyperparametric model has two types of parameters: the hyperparameter $\theta$ to be estimated, and the functional parameters $(\pi_m)$ to be filtered.

Non-Parametric Identification

In this section, we consider the identification of the functional parameter $\pi$ within each model from the observation of a demand function. Then, we examine if we can distin- guish between the SARA and SSF models.

Intuitively, a consumer's demand function is identified if we observe her making a lot of consumption decisions at a variety of designs $z$. Clearly, we can identify her demand function if (i) her preferences are constant over time and we observe a large panel or experiment,\footnote{In this case, when the number of dates $T$ is large, we can have a segment $m$ for each consumer $i$.} or (ii) she belongs to a large homogeneous segment of consum- ers with identical preferences. This explains the form of Assumption (ref) (as it allows for either interpretation). Later, we apply the segmented approach to scanner data in the application to the consumption of alcohol in Section (ref).

figure[figure omitted — 2,385 chars of source]

With panel data, one no longer requires the assumption that demand is monotonic with respect to unobserved heterogeneity in order to achieve identification (see Figure (ref), and the role of this assumption in brown-matzkin, brown-matzkin, matzkin-nonadd, matzkin-nonadd, and hn-individual-het, hn-individual-het).

Within Model Identification

In the models introduced in Sections (ref) and (ref), and for any $\pi$ such that demand is inv- ertible, we can derive the inverse demand function, whose second component coincides with the MRS which can be integrated to obtain a unique preference ordering. Indeed, by construction, the integrability conditions (needed to recover a unique well-behaved preference ordering) are satisfied, implying that preferences are recoverable (see samuelson1948, samuelson1948, for a seminal discussion of integrability in the case of two goods, and samuelson, samuelson, hurwicz-uzawa, hurwicz-uzawa, and hosoya2016, hosoya2016, for general approaches). However, the possibility to recover preferences from a consumer's demand function does not imply that the distribution of taste uncertainty $\pi$ is identified. Indeed, two distinct taste distributions could produce an identical MRS.

For identification, we only consider the information contained in the demand function $X(\cdot;\pi)$ on the set $\mathcal{Z}$ of designs $z$ for which the components of the demand function are strictly positive. This restriction disregards some information that may be available in the first or third regimes of (ref). In most datasets, when a component of the demand function equals zero, the price $p$ is not observed.

Stochastic Absolute Risk Aversion

In the stochastic absolute risk aversion (SARA) model, the identification condition is:

equation[equation omitted — 230 chars of source]

In the degenerate case in which $A$ is deterministic and equal to $(a_1,a_2)$, the MRS red- uces to $a_1/a_2$. Thus, in this special case, the two-dimensional parameter $a=(a_1,a_2)$ is identified up to a positive factor. This reasoning leads us to a question: Does this lack of identification also exist in an extended setting?

Let us first remark that the utility function $U(x;\pi)$ is equal to the moment generating function for $\pi$ with a negative sign: $\Phi(x;\pi)=-U(x;\pi)$. Because this moment generating function characterizes $\pi$ when the stochastic parameter $A$ is non-negative (see Theorem 1a in Chapter 13 on Tauberian Theorems in feller-1968, feller-1968), it is equivalent to consider the identification of either $\pi$, or $\Phi(x;\pi)$.\footnote{Note, the existence of the moment generating function does not imply the existence of all power moments and, even if all power moments exist, they do not necessarily characterize the distribution. A known example is the log-normal distribution used in the application heyde.} As mentioned, we can always integrate the MRS to recover a unique preference ordering. That is, we can recover $U(x;\pi)$ up to a monotonic transformation. We still need to discern the conditions on $\pi$ under which we can recover $\Phi(x;\pi)$. Indeed, moment generating functions have properties that are not necessarily preserved under monotonic transformations.

We obtain the following result:

propositionIf preferences are SARA, then $\Phi(x;\pi)$ and $\Phi(x;\pi)^{\nu}$ lead to the same preference ordering, for all positive scalars $\nu>0$.
proofLet $U(x;\pi)=-\Phi(x;\pi)$ and $\tilde{U}(x;\pi)=-\Phi(x;\pi)^{\nu}$ denote the utility functions associated with $\Phi(x;\pi)$ and $\Phi(x;\pi)^{\nu}$, respectively. Then, by definition, we must have: \begin{equation} \tilde{U}(x;\pi)=-\Phi(x;\pi)^{\nu}=-(-U(x;\pi))^{\nu}=\phi_{\nu}(U(x;\pi)), \end{equation} where $\phi_{\nu}(u)=-(-u)^{\nu}$ is strictly increasing for $u<0$. Since $\tilde{U}(x;\pi)$ is a monotonic transformation of $U(x;\pi)$, these utility functions yield the same preference ordering.

This means that we can, at most, identify the class of moment generating functions $\mathscr{C}(\Phi)=\{\Phi^{\nu}:\nu>0\}$. Note that, for any moment generating function $\Phi$, the transf- ormed function $\Phi^{\nu}$ is also a moment generating function.

Let us now consider identification when $A_1$ and $A_2$ are independent:

propositionLet $\Phi_j$ denote the marginal moment generating function for $A_j$, for $j=1,2$. If preferences are SARA, and $A_1$ and $A_2$ are independent, then $(\Phi_1,\Phi_2)$ and $(\Phi_1^*,\Phi_2^*)$ lead to the same preference ordering if, and only if, for some $\nu >0$, we have: \[ \Phi_1^*=\Phi_1^{\nu} \; \; \text{and} \; \; \Phi_2^*=\Phi_2^{\nu}. \]
proofThe identification criterion becomes: \[ \left(\frac{\partial\Phi_1(x_1)}{\partial x_1}\Phi_2(x_2)\right)\left(\Phi_1(x_1)\frac{\partial\Phi_2(x_2)}{\partial x_2}\right)^{-1}=\left(\frac{\partial\Phi_1^*(x_1)}{\partial x_1}\Phi_2^*(x_2)\right)\left(\Phi_1^*(x_1)\frac{\partial\Phi_2^*(x_2)}{\partial x_2}\right)^{-1}, \] for all $x\in R$. This criterion can, then, be written as: \[ \frac{\partial\log \Phi_1(x_1)}{\partial x_1}\left(\frac{\partial\log \Phi_1^*(x_1)}{\partial x_1}\right)^{-1}=\frac{\partial\log \Phi_2(x_2)}{\partial x_2}\left(\frac{\partial\log \Phi_2^*(x_2)}{\partial x_2}\right)^{-1}, \] for all $x\in R$. Thus, we deduce that, if these distributions yield the same MRS, then: \[ \frac{\partial\log \Phi_j^*(x_j)}{\partial x_j}=\nu\frac{\partial\log \Phi_j(x_j)}{\partial x_j}, \] for some $\nu >0$, at every $x_j\geq 0$, for both $j=1,2$. Because the log-transform of the moment generating function at zero equals zero, by integrating this equation, we get: \begin{equation} \log \Phi_j^*(x_j)=\nu\log\Phi_j(x_j), \end{equation} at every $x_j\geq 0$, for both $j=1,2$. Equivalently, $\Phi_1^*=\Phi_1^{\nu}$ and $\Phi_2^*=\Phi_2^{\nu}$.

Proposition (ref) implies that $\mathscr{C}(\Phi)$ is identified under the independence of $A_1$ and $A_2$. Indeed, we can recover the consumer's preference ordering using traditional methods, and use the fact that all admissible preference orderings map to a unique class $\mathscr{C}(\Phi)$.

Of course, independence is a strong restriction. In the SARA model, it is equivalent to the additive separability of the utility function.\footnote{In the case of two goods, additive separability is stronger than separability.} To see this result, notice that, under independence, we obtain:

equation[equation omitted — 103 chars of source]

Since utility functions are unique up to strictly increasing transformations, this utility function is equivalent to:

equation[equation omitted — 139 chars of source]

which is an additively separable utility function. In Appendix (ref), we prove a generalization of Proposition (ref) where stochastic taste parameters have a common component.

Stochastic Safety-First

In the SSF model, the identification condition is:

equation[equation omitted — 228 chars of source]

Let us now consider the validity of this condition under an independence assumption. Note, in the SSF model, independence is no longer equivalent to additive separability.

propositionIf preferences are SSF, $A_1$ and $A_2$ are independent, and the marginal distribution of $A_2$ is continuous, then $\mathbb{E}[A_1]$ is identified, and the marginal distribution of $A_2$ is identified up to some positive power transformation of its survival function.
proofIn the SSF model, the MRS is identified, and it satisfies: \begin{equation} \mathbb{E}_{\pi}[A_1S(x_1+A_1x_2)]=MRS(x;\pi)\mathbb{E}_{\pi}[S(x_1+A_1x_2)], \end{equation} where $S(\cdot)$ denotes the survival function of $A_2$. \begin{enumerate}[(i)] • The expectation $\mathbb{E}_{\pi}[A_1]$ is identified because $\text{MRS}(x_1,0;\pi)=\mathbb{E}_{\pi}[A_1]$. • By differentiating (ref) with respect to $x_2$, we get: \[ \begin{gathered} \mathbb{E}_{\pi}[A_1^2S'(x_1+A_1x_2)]=\text{MRS}(x;\pi)\mathbb{E}_{\pi}[A_1S'(x_1+A_1x_2)] \\ +\frac{\partial \text{MRS}}{\partial x_2}(x;\pi)\mathbb{E}_{\pi}[S(x_1+A_1x_2)]. \end{gathered} \] When $x_2=0$, this equation becomes: \[ S'(x_1)\mathbb{E}_{\pi}[A_1^2]=\text{MRS}(x_1,0;\pi)S'(x_1)\mathbb{E}_{\pi}[A_1]+\frac{\partial \text{MRS}}{\partial x_2}(x_1,0;\pi)S(x_1). \] By rearranging, we get: \[ \frac{\partial \text{MRS}}{\partial x_2}(x_1,0;\pi)=\frac{S'(x_1)}{S(x_1)}\left(\mathbb{E}_{\pi}[A_1^2]-\text{MRS}(x_1,0;\pi)\mathbb{E}_{\pi}[A_1]\right)=\frac{S'(x_1)}{S(x_1)}V(A_1). \] Because the partial derivative of the MRS with respect to $x_2$ is identified, the hazard function $\lambda(x_1)=-S'(x_1)/S(x_1)$ of the distribution of $A_2$ is identified up to a positive factor. Since $S(x_1)=\exp\{-\Lambda(x_1)\}$, where $\Lambda(x_1)=\int_0^{x_1}\lambda(t)dt$ is the cumulative hazard function of the distribution of $A_2$, we can identify $S(\cdot)$ up to a positive power transformation. \end{enumerate}

Proposition (ref) provides no information on the identifiability of the distribution of $A_1$ beyond its first moment. It seems difficult to obtain a general identification result, but insights into our identification problem can be obtained by considering the two primary families of distributions that are invariant to positive power transformations, that are, the exponential family and the Pareto family.

enumerate[(i)] • Exponential family: Suppose that the marginal distribution of $A_2$ belongs to the exponential family, and that we have identified its survival function up to a positive power transformation such that $S(x)=\exp\{-cx\}$, for some unknown $c>0$. The MRS in (ref) becomes: \begin{equation} MRS(x;\pi)=\frac{\mathbb{E}_{\pi}[A_1\exp\{-cx_2A_1\}]}{\mathbb{E}_{\pi}[\exp\{-cx_2A_1\}]}\equiv G_0(x_2). \end{equation} This expression does not depend on $x_1$. Now, let $\Psi(u)=\mathbb{E}_{\pi}[\exp\{-uA_1\}]$ denote the Laplace transform of $A_1$. Under this notation, the equality in (ref) implies: \[ G_0(x_2)=\frac{d\log \Psi}{du}(cx_2). \] Or, equivalently, $G_0(u/c)=d\log \Psi(u)/du$. By integrating, we obtain: \[ \log\Psi(u)=c\big[H(u/c)-H(0)\big], \] where $H(\cdot)$ is a primitive of the MRS. Therefore: \begin{corollary} Under the conditions of Proposition (ref), if the marginal distribution of $A_2$ belongs to the exponential family, the following results hold: \begin{enumerate}[(a)] • The power transform $c$ is not identified. • The distribution of $A_1$ is identified under an identification restriction on $c$. \end{enumerate} \end{corollary} We conclude that, under the conditions of Corollary (ref), the distributions of $A_1$ and $A_2$ are non-parametrically identified up to a single scalar parameter $c>0$. • Pareto family: Let us now examine whether a similar result can be obtained for the Pareto family, in which $S(x)=x^{-\alpha}$, for some $\alpha>0$. The parameter $\alpha$ characterizes the fat tails of the distribution of $A_2$ and the power transformation on the MRS. This survival function produces: \[ \text{MRS}(x;\pi)=\frac{\mathbb{E}_{\pi}[A_1(x_1+A_1x_2)^{-\alpha}]}{\mathbb{E}_{\pi}[(x_1+A_1x_2)^{-\alpha}]}=\frac{\mathbb{E}_{\pi}[A_1(x_0+A_1)^{-\alpha}]}{\mathbb{E}_{\pi}[(x_0+A_1)^{-\alpha}]}, \] where $x_0\equiv x_1/x_2$ denotes a ratio of quantities. Equivalently, we get: \begin{equation} MRS(x;\pi)=\frac{\mathbb{E}_{\pi}[(x_0+A_1)^{-\alpha+1}]}{\mathbb{E}_{\pi}[(x_0+A_1)^{-\alpha}]}-x_0\equiv G_0(x_0), \end{equation} which only depends on the ratio $x_0$. Therefore, we have constructed homothetic preferences. By equation (ref): \[ e(x)\equiv\frac{d}{dx}\log \mathbb{E}_{\pi}[(x+A_1)^{-\alpha+1}], \] is identified up to a multiplicative constant. Therefore, by integration, $\mathbb{E}_{\pi}[(x+A_1)^{-\alpha+1}]$ is identified up to $\alpha$ and a multiplicative constant $\kappa$. However, as $x$ tends to infinity, this expression is equivalent to $\kappa x^{-\alpha+1}\exp E(x)$, where $E(\cdot)$ is a primitive of $e(\cdot)$. This tail behaviour provides both the identification of $\alpha$ and $\kappa$. This analysis is summarized by the following result: \begin{corollary} Under the conditions of Proposition (ref), if the marginal distribution of $A_2$ belongs to the Pareto family, the distributions of $A_1$ and $A_2$ are both non-parametrically identified. \end{corollary}

Between Model Identification

Once the identification of the consumer's taste distribution $\pi$ within each model is solved, we still need to consider the identification between the models. This analysis is needed to test whether preferences are consistent with SARA, or SSF, or both. It is important to know whether these two classes of preferences are nested or non-nested. If they are non-nested, we need to characterize their intersection and define a general class encompassing both types of preferences.

To illustrate, suppose that the consumer has SSF preferences:

equation[equation omitted — 104 chars of source]

where (i) $A_1$ and $A_2$ are independent, (ii) $A_1$ has distribution $\pi_2$, and (iii) $A_2$ follows an exponential distribution (with unit intensity). Under this specification, we obtain:

equation[equation omitted — 90 chars of source]

To clarify this result, observe that, by conditioning on $A_1$, we are left with the expec- tation of the minimum of a set containing a constant and a random variable with an exponential distribution. This utility function is a strictly increasing transformation of a SARA utility function:

equation[equation omitted — 74 chars of source]

where (i) $B_1$ follows a point mass at $1$, and (ii) $B_2$ has distribution $\pi_2$. Consequently, these utility functions, one SARA, and the other SSF, induce the same preference ordering over the consumption set.

Discussion

The possible lack of identification of each consumer's taste distribution $\pi_m$ has to be taken into account in the economic interpretation of the results. However, it has to be noted that it does not create difficulties for structural inference, where the (scalar or functional) parameters of interest are the parameters characterizing the MRS, rather than the parameters characterizing the utility function.

The lack of identification is due to the special structure of the cone of increasing and concave functions defined on $R$, and of the extremal elements of this cone. For finite increasing concave functions defined on $\mathbb{R}_+$, it is well-known that the extremal functions are of the type:

equation[equation omitted — 81 chars of source]

in which $(\alpha_j,\beta_j)\in \bar{R}$, for $j=1,2$ (see blaschke-pick, blaschke-pick), and that any finite positive increasing concave function can be written as:

equation[equation omitted — 81 chars of source]

where $b$ is a positive scalar and $\pi$ is the distribution of $A$. Such functions are charac- terized by $b$ and $\pi$. The set of extremal functions in (ref) is a minimal set of extremal points generating the cone.

Such a property no longer holds for finite positive increasing concave functions defined on $\bar{R}$. johansen has described a large set of extremal points of the type:

equation[equation omitted — 85 chars of source]

for which $h_1(\cdot)$ induces a covering with vertices of order 3 (see page 62 in johansen, johansen), and has shown that this set is dense in the cone of finite continuous convex functions defined on a convex set in $R$ (see Theorem 2 in johansen, johansen). A minimal set of extremal points generating this cone does not exist. This argument explains why Sections (ref) and (ref) consider specific convex subsets generated by parametric functions.

While we restrict our attention to SARA and SSF preferences (because the stochastic taste parameters have clear interpretations in these models), other convex hulls co- uld have been considered. For example:

enumerate[(i)] • The convex hull generated by the union of the SARA and SSF models---that is, the smallest structural model containing both of the models in Sections (ref) and (ref). • The convex hull generated by a basis of the form: \begin{equation} U(x;a,\nu)=\frac{a_1}{\nu_1}x_1^{\nu_1}+\frac{a_2}{\nu_2}x_2^{\nu_2}, \end{equation} for every $x\in\bar{R}$ in which $a\in R$ and $\nu\in(0,1)^2$. This basis corresponds to a first-order expansion of a utility function johansen-relationship, and contains a Stone-Geary utility function as a limiting case. Indeed, as $\nu$ approaches zero, we obtain: $U(x;a)=a_1\log x_1+a_2\log x_2$. However, the convex hull generated by this basis is not flexible enough because it only contains weighted combinations of $x_1^{\nu_1}$ and $x_2^{\nu_2}$. Similarly, the convex hull generated by a Stone-Geary basis only contains Stone-Geary utility functions, where the weights are the means of the taste parameters: \begin{equation} U(x;\pi)=\mathbb{E}_{\pi}\big[A_1\log x_1+A_2\log x_2\big] = \mathbb{E}_{\pi}\big[A_1]\log x_1+\mathbb{E}_{\pi}\big[A_2]\log x_2. \end{equation}

The Stone-Geary basis $U(x;a)$ above can be adjusted to define another parametric basis. In particular, let us apply the transformation $\varphi(x)=-\exp(-x)$ to the Stone-Geary utility function. This transformation yields:

equation[equation omitted — 84 chars of source]

This utility function forms a well-behaved basis because it is strictly increasing with a negative semi-definite Hessian. While $U(x;a)$ and $\tilde{U}(x;a)$ represent the same preference ordering, they will generate different families due to the strict concavity of $\varphi(\cdot)$. To illustrate, suppose that the stochastic parameters, $A_1$ and $A_2$, are independently distributed with respect to uniform distributions on $[0,1]$. This specification produces:

equation[equation omitted — 190 chars of source]

While $U(x;\pi)$ is a Stone-Geary utility function, $\tilde{U}(x;\pi)$ is a complicated non-linear function of $x_1$ and $x_2$. Consequently, an uninteresting basis has been transformed into an interesting one. This procedure can be completed for any increasing, concave, and twice-differentiable transformation $\varphi(\cdot)$.

An Illustration

This section shows how to use the SARA and SSF models in a non-parametric framework. First, we specify the statistical model by introducing an assumption on the obs- ervations, and then we discuss statistical inference. The methodology is illustrated in an application to alcohol consumption using scanner data concerning individual purchase histories.

Assumptions on Observations

The behavioural models introduced in the previous sections can be completed with an assumption on the available observations. We consider panel data, indexed by the consumer $i$ and date $t$. After a preliminary treatment of the purchase histories, we have a large number $n$ of consumers and a fixed number $T$ of observed dates. In the preliminary treatment, the goods are aggregated into two groups using a common quantity unit and the dated purchases are aggregated by month (see Section (ref)). Recall that, under Assumption (ref), we have $M$ segments of homogeneous consumers.

We introduce the following assumption on the observations:

assumptiona[Observations] \phantom{1} \begin{enumerate}[(i)] • We jointly observe $(x_{it},z_{it})$, for all $i=1,\dots,n$ and $t=1,\dots,T$, when $x_{it}>0$. • The individual histories $(x_{it},z_{it})_{t=1}^T$ are independent given all $\pi_m$, $m=1,\dots,M$. • Designs $(z_{it})$ are exogenous (independent of taste distributions $\pi_m$). \end{enumerate}

Assumption (ref) describes the structure of the observations. It implies that we can imagine taste parameters $(\pi_m)$ being independently drawn from a Dirichlet process $F$, designs $(z_{it})$ being independently drawn from some distribution, and consumption $x_{it}$ satisfying $x_{it}=X(z_{it};\pi_{m_i})$, where $m_i$ is the group of consumer $i$. Many papers assume that consumption $x_{it}$ is positive (see Section IV.A in \hyper@link{cite}{cite.blundell\@extra@b@citeb}{Blundell, Horowitz, and Parey}, \hyper@link{cite}{cite.blundell\@extra@b@citeb}{2017}, for this assumption in an application to gasoline demand, as well as Assumption A5 in dobronyi, dobronyi, for this assumption in an application to the consumption of alcohol); the SARA and SSF models allow for corner solutions. However, in many datasets (including the dataset used in the application in Section (ref)), there is a problem of partial observability. Let $\tilde{y}$ denote the expenditure (prior to normalization), and let $\tilde{p}_j$ denote the price of good $j$ (prior to normalization). Usually, we only observe the price $\tilde{p}_j$ of a good $j$ when the consumer buys a positive quantity of good $j$. Then, we only observe (normalized) expenditure $y$ when the consumer buys a positive quantity of good 2, and we only observe the (normalized) price $p$ when the consumer buys a positive quantity of both goods (see craw-pol, craw-pol, for an approach to revealed preference that deals with this partial observability problem).\, This problem explains the specific form of Assumption \hyperref[ass:2]{2(i)}.

For deriving the asymptotic properties of estimators, it is also necessary to specify the type of asymptotics to be considered:

assumptionaLet $n_m$ denote the size of the $m^{th}$ homogeneous group. \begin{enumerate}[(i)] • $n_mT\rightarrow\infty$, as $n\rightarrow\infty$, for all $m=1,\dots,M$. • $n_mT\sim \lambda_m n$, for some $\lambda_m\in(\lambda_{\ell},\lambda_h)$, where $0<\lambda_{\ell}<\lambda_h<1$, for $m=1,\dots,M$. • $M\rightarrow\infty$, as $n\rightarrow\infty$. \end{enumerate}

Assumptions \hyperref[ass:3]{A3(i)} and \hyperref[ass:3]{A3(ii)} ensure that there are enough observations to non-parametrically estimate the demand function associated with the functional parameter $\pi_m$ on a sufficiently large subset $\mathcal{Z}_m$ of designs $z$. Assumption \hyperref[ass:3]{A3(iii)} guarantees enough filtered parameters $\hat{\pi}_m$ to estimate the underlying Dirichlet process $F$. In some special circumstances, $T$ is large, and Assumption (ref) can be used with $m=i$ and $M=n$---that is, a single consumer per group. Otherwise, grouping of homogeneous consumers is needed to identify the demand functions on sufficiently large subsets $\mathcal{Z}_m$.

Estimation Method

The Dirichlet process is common in Bayesian estimation (see, for instance, ferguson, ferguson, and li-2019, li-2019). This process is useful because it is flexible and, if observations are independently and identically drawn from an unknown distribution, the posterior distribution of this distribution has a closed-form expression. However, our framework is much more complicated for two reasons:

enumerate[(i)] • The observed consumption choices $(X_{ijt})$ are not identically distributed because consumers make decisions at different expenditures and prices. • It is difficult to derive a closed-form expression for the demand, as a function of the expenditure, the price, and the functional parameter characterizing taste uncertainty $\pi$. It is, therefore, difficult to derive a closed-form expression for the distribution of $X_{ijt}$ conditional on $Z_{it}$.

These features of our model explain why estimation requires specific numerical algor- ithms. These specific algorithms have to be able to deal with the non-linear and high- dimensional features of the models. In the Bayesian framework, the Dirichlet process is fixed. In the hyperparametric framework, it is parameterized by a vector $\theta$. These parameters have to be estimated and the functional parameters $(\pi_m)$ have to be filter- ed. These estimation approaches are described below.

Bayesian Framework

In a pure Bayesian framework, a Dirichlet process is fixed by selecting a mean distribution $\mu$ and a scaling parameter $c$ (see Appendix (ref)). This distribution defines the common prior for the functional taste parameters $(\pi_m)$. After, the data are used to compute the posterior distribution for the functional parameters $(\pi_m)$. Under Assu- mption (ref), the posterior distribution can be computed separately for each homogeneous group of consumers: \[ \ell(\pi_m|x_{it},z_{it},x_{it}>0,i\in\Lambda_m, t=1,\dots,T), \] where $\Lambda_m$ denotes the group of consumers with preferences characterized by the taste parameter $\pi_m$. This approach does not have to account for the potential identification problem discussed in Section (ref). If a specific characteristic of $\pi_m$ is weakly identified, its posterior distribution will be close to the prior distribution.

In our framework, the observations $(x_{it},z_{it})$, conditional on $x_{it}>0$, must satisfy the deterministic first-order conditions implied by the model. These conditions have the following form:

equation[equation omitted — 71 chars of source]

for any observed pair $(x_{it},z_{it})$. Equivalently:

equation[equation omitted — 193 chars of source]

for any observed pair $(x_{it},z_{it})$. These conditions are moment restrictions, called MRS restrictions. In our big data framework, the number of MRS restrictions is very large, typically several hundred to a thousand. The posterior of $\pi_m$ is simply the distribution of $\pi_m$ given these deterministic restrictions on $\pi_m$. If the taste parameters, $A_1$ and $A_2$, are independent with marginal distributions, $\pi_1$ and $\pi_2$, respectively, then the MRS restrictions are bilinear in $\pi_1$ and $\pi_2$---specifically, these restrictions are linear in $\pi_1$ given $\pi_2$, and linear in $\pi_2$ given $\pi_1$. Later, this property is used to construct a numerically efficient optimization algorithm for filtering all the $\pi_m$ (see Appendix (ref)).

Hyperparametric (or Empirical Bayesian) Framework

The hyperparametric (or empirical Bayesian) framework is a complicated non-linear state-space model with two layers of latent state variables. Such a framework can be characterized as follows:

enumerate[(i)] • Deep layer: Functional parameters $(\pi_m)$, drawn from $F$ (parameterized by $\theta$); • Surface layer: Demand functions $X(\cdot;\pi_m)$ deduced from $\pi_m$; • Measurement equations: Observed pairs $(x_{it},z_{it})$, given $x_{it}>0$.

We have partial observability of the demand function because the value of demand $X(z;\pi_m)$ is observed at finitely many designs $z$. Furthermore, unlike most state-space models, the state variables are infinite-dimensional.

Estimating the Hyperparameter

While it is difficult to derive analytically the distribution of $X_{it}$ given $Z_{it}$, it is easy to simulate its distribution for a given value of $\theta$ (see Appendix (ref) for simulations from the Dirichlet distribution). Therefore, $\theta$ can be estimated by the method of simulated moments (MSM), or indirect inference (see mcfadden-msm, mcfadden-msm, pakes-optimization, pakes-optimization, and gou-mon, gou-mon). That is, $\theta$ is estimated by matching some sample and simulated moments of the pair $(X_{it},Z_{it})$.

To illustrate, consider a pure panel such that $M=n$.\footnote{When $M<n$, we simulate $n_mT$ observations for the $m^{th}$ draw from the Dirichlet process.} The steps are the following:

steps• Simulate $s=1,\dots,n$ draws from a Dirichlet process given the parameter $\theta$. Each draw $\pi^s(\theta)$ is associated with an individual consumer $i$ such that $s=i$. • Compute simulated consumption $x_{it}^s(\theta)$ by solving the first-order condition in (ref) with respect to $x_1$ and applying the transformation in (ref) given $z_{it}=(y_{it},p_{it})$ and $\pi^i(\theta)$. • Construct a collection of $K$ moments from the observed and simulated data: \[ m\equiv\left[\frac{1}{nT}\sum_{i=1}^n\sum_{t=1}^T m_k(x_{it},z_{it})\right]_k \; \; \text{and} \; \; m(\theta)\equiv\left[\frac{1}{nT}\sum_{i=1}^n\sum_{t=1}^T m_k(x_{it}^s(\theta),z_{it})\right]_k. \] Then, numerically solve the following problem: \begin{equation} \operatorname*{argmin}_{\theta} \; \big|\big|m-m(\theta)\big|\big|, \end{equation} in which $||\cdot||$ is a Euclidean norm with the form $||m||^2=m'\Omega m$, for some positive-definite $K\times K$ matrix $\Omega$.

Given the estimated hyperparameter $\hat{\theta}$, the taste distributions $(\pi_m)$ must be filtered. This step is equivalent to applying the Bayesian approach with the estimated Dirichlet distribution as the prior distribution (see Appendix (ref)).

Under Assumptions (ref) to (ref), the estimator for $\theta$ is consistent and asymptotically normal, and it converges at a speed of $1/\sqrt{nT}$. The derivation of the asymptotic pro- perties of the filtered functional parameter $\hat{\pi}_m$ is out of the scope of this paper and left for future research.

Filtering the Taste Distributions

Once the hyperparameter $\theta$ is estimated, we can filter $\pi_m$ by using the following steps:

steps• Draw a taste distribution $\tilde{\pi}_m$ from the Dirichlet process given $\hat{\theta}$. Then, by construction, the taste distribution $\tilde{\pi}_m$ is a draw from the prior distribution. • Discretize $\tilde{\pi}_m$ on a grid of values for the taste parameters, $A_1$ and $A_2$. Let $\bar{\pi}_m$ denote the result. The aim of this step is to put $\tilde{\pi}_m$ on a grid for optimization. • Solve the minimization problem: \[ \min_{\pi} \; ||\pi-\bar{\pi}_m|| \; \; \text{s.t. MRS restictions \eqref{mrsrestrictions} and unit mass restrictions.} \] Let $\hat{\pi}_m^*$ denote the solution. This solution approximates a drawing from the posterior. • Replicate these steps to obtain a sequence of solutions: $\hat{\pi}_{m,s}^*$, $s=1,\dots,S$, where $S$ is the number of replications. The filtered $\hat{\pi}_m$ is obtained by averaging over all simulations such that: \[ \hat{\pi}_m=\frac{1}{S}\sum_{s=1}^S\hat{\pi}_{m,s}^* \]

This procedure involves a high-dimensional argument $\pi_m$, and a very large number of MRS restrictions. Indeed, we need several hundred grid points for $\pi_m$, and, in the application, we have about one-thousand MRS restrictions, for each $m=1,\dots,M$. If the taste parameters, $A_1$ and $A_2$, are independent, this procedure can be numerically simplified by using the fact that these restrictions are bilinear (see Section (ref) and Appendix (ref)).

The Data

We use the Nielsen Homescan Consumer Panel (NHCP). Nielsen provides a sample of households with barcode scanners. Households are asked to scan all purchased goods on the date of each purchase. The prices are entered by the households or linked to retailer data by The Nielsen Company. The households that agree to participate are compensated through benefits and lotteries.

We focus on the consumption of alcoholic drinks (see manning-et-al, manning-et-al, for an application to alcohol consumption in economics). We classify drinks by type. Good 1 contains beers and ciders.\footnote{The NHCP classifies ciders as wine, by default. We reclassify these beverages using product desc- riptions because most ciders have a low alcohol by volume (ABV).} Good 2 contains wines and liquors. We disregard all non-alcoholic beers, ciders, and wines. We are left with 30,635 beers and ciders, and 108,439 wines and liquors, for a total of 139,074 drinks. We convert all measurement units to litres of alcohol by first converting all units to litres and then multiplying by the standard alcohol by volume (ABV) in each subgroup---specifically, 4.5% for beer and cider, 11.6% for wine, and 37% for liquor. For example, if a household buys two packs of six bottles of beer and each bottle contains 355 millilitres of beer, then the household buys 4.26 litres of this beer, or $4.26\times 0.045=0.231$ litres of alcohol. We use the standard ABV in each subgroup as a result of data limitations. Our sample only contains purchases made at stores, not purchases made at bars, or restaurants.

Measuring quantities in litres of alcohol has at least three advantages: (i) it can account for a quality effect, (ii) it is appropriate for analyzing most relevant structural objects (e.g. the effect of a change in taxation on alcohol consumption), and (iii) it yi- elds continuous quantities, permitting the application of standard tools in consumer theory (which could not be used if quantities were measured in, for example, bottles), and avoiding some common identification issues in the literature.

We restrict our sample to purchases made from August to November in 2016. This relatively short window is used to diminish the impact of changing tastes and product availability, and to avoid most federal holidays in the United States that are often associated with alcohol consumption such as Independence Day, Christmas Day, and New Year's Eve. Our sample contains 28,036 households. Some additional details of this restricted sample are placed in Appendix (ref).

table[table omitted — 1,191 chars of source]

The dated purchases are aggregated by month. For each household and month, the prices are constructed by dividing the total expenditure for each aggregate good (after accounting for the value of coupons) by the amount of alcohol of that aggregate good purchased by the household, when this amount is strictly positive. Then, we norm- alize by the price of good 2. This procedure yields four monthly observations per hou- sehold for a total of 112,144. A total of 63,936 observations have positive consumption such that $x_{it}> 0$. Table (ref) gives summary statistics conditional on $x_{it}\neq 0$. The prices $(p_{it})$ are conditional on being well-defined (see the discussion of partial observability on pages 23 and 24). For the interpretation of the results, recall that $\tilde{y}$ is the expenditure (prior to normalization), and that $\tilde{p}_j$ is the price of good $j$ (prior to normal- ization).

table[table omitted — 353 chars of source]

There are four regimes of observations: (i) zero expenditure on all goods, (ii) zero expenditure on good 1 and strictly positive expenditure on good 2, (iii) strictly pos- itive expenditure on good 1 and zero expenditure on good 2, and (iv) strictly positive expenditure on all goods. Table (ref) provides the proportion of observations in each regi- me, and shows a large proportion of observations with zero expenditure. Recall, under Assumption (ref), designs $z_{it}$ are drawn from a distribution. Therefore, we can interpret this result as a mass at zero in the marginal distribution of expenditure.

Figure (ref) displays the sample distribution of expenditure $\tilde{y}_{it}$ by regime: the distribution of expenditure $\tilde{y}_{it}$ conditional on $x_{it}>0$ is on the left; the sample distributions of expenditure $\tilde{y}_{it}$ for the two other regimes with positive expenditure are on the right. The shape of the sample distribution of expenditure $\tilde{y}_{it}$ does not appear to vary all that much with the regime. That being said, the sample distribution conditional on $x_{it}>0$ has more probability attributed to higher expenditures.

Figure (ref) compares the sample distributions of prices $\tilde{p}_j$ by regime: the sample dis- tributions of $\tilde{p}_1$ are on the left; the sample distributions of $\tilde{p}_2$ are on the right. Although the sample distribution of $\tilde{p}_1$ differs from the sample distribution of $\tilde{p}_2$, these distributions do not seem to be affected by the regime.

Figure (ref) displays the sample distributions of (normalized) designs $z_{it}=(y_{it},p_{it})$ and the components of consumption $x_{it}$ given $x_{it}>0$. As expected, the components of consumption $x_{it}$ are increasing in expenditure $y_{it}$. Furthermore, the first component of consumption $x_{it}$ is more affected by changes in the price $p_{it}$ than the second component.

figure[figure omitted — 707 chars of source]
figure[figure omitted — 768 chars of source]
figure[figure omitted — 574 chars of source]

Since we consider a rather short window of time, we follow the segmented population approach. We segment the population by state. Large states (e.g. California) are segmented again by county. Specifically, a county is given its own segment if it has more than 70 observations with positive consumption and it is in a state with more than 1,000 observations with positive consumption. We are left with a total of 65 seg- ments, each corresponding to a state or county. The smallest segment is Wyoming, containing 15 observations with positive consumption; the largest state is Florida (after removing Broward, Hillsborough, Palm Beach, Pinellas, and Miami-Dade counties), containing 880 observations with positive expenditure; the mean number of observations with positive consumption per segment is approximately 226.

Figure (ref) displays the sample distributions of (normalized) designs $z_{it}=(y_{it},p_{it})$ and the first component of consumption $x_{it}$ given $x_{it}>0$ in two of the larger segments: California (after removing Almeda, Los Angeles, Orange, Riverside, Sacramento, San Bernardino, and San Diego counties), and Florida (after removing Broward, Hillsborough, Palm Beach, Pinellas, and Miami-Dade counties).

figure[figure omitted — 857 chars of source]

Figure (ref) displays the Nadaraya-Watson (kernel) estimates of the demand function for beer conditional on $x_{it}>0$ in California and Florida over a subset of the domain of designs. Demand for beer in California is lower and less responsive to price changes than in Florida. Figure (ref) displays Engel curves for good 1 in California and Florida given $p\equiv\tilde{p}_1/\tilde{p}_2=4$.\footnote{This price is chosen to be in a sufficiently dense region of the sample distribution (see Figure (ref)).} These Engel curves cross.

figure[figure omitted — 493 chars of source]
figure[figure omitted — 261 chars of source]

Estimation Results

As an illustration, we consider the SARA model in the hyperparametric framework. We assume that the taste parameters, $A_1$ and $A_2$, are independent. Under this assum- ption, the taste uncertainty is characterized by the marginal distributions, $\pi_1$ and $\pi_2$. The marginal distribution $\pi_j$ of $A_j$ is independently drawn from a Dirichlet process $F_j$, $j=1,2$. The mean of $F_j$ is a log-normal distribution with parameters $\mu_j$ and $\sigma_j$, and the scale parameter of $F_j$ is $c_j$. The utility function corresponding to this log-normal mean distribution, say $\bar{\pi}_j$, has a quasi closed-form expression. Indeed, under this distribution, we obtain: \[ \log(A_j)=\mu_j+\sigma_j\varepsilon_j, \; \; \forall j=1,2, \] where $\varepsilon_j$ is distributed with respect to a standard normal distribution. Then: \[

gathered\mathbb{E}_{\bar{\pi}_j}[\exp(-A_jx_j)]=\mathbb{E}_{\bar{\pi}_j}[\exp(-\exp(\mu_j+\sigma_j\varepsilon_j)x_j)], \\ =\frac{1}{\sqrt{1+w(x_j\exp(\mu_j)\sigma_j^2)}}\exp\left\{-\frac{1}{2\sigma_j^2}w(x_j\exp(\mu_j)\sigma_j^2)^2-\frac{1}{\sigma^2}w(x_j\exp(\mu_j)\sigma_j^2)\right\},

\] where $w(x)$ is the Lambert function, defined by the implicit equation: \[ w(x)\exp(w(x))=x, \] [see equation (1.3) in laplace-lognormal, laplace-lognormal]. By drawing from the Dirichlet process, we will draw a stochastic utility function around the closed-form expression above. The hyperparameter $\theta$ has six components such that: \[ \theta=(\mu_1,\sigma_1,c_1,\mu_2,\sigma_2,c_2). \]

The Hyperparameter

As described in Section (ref), the first step involves estimating the hyperparameter $\theta$ using the Method of Simulated Moments (MSM). The hyperparameter $\theta$ is calibrated by using the following (sample and simulated) moments computed for all of the 63,936 observations with positive consumption:

enumerate[(i)] • marginal moments of $(X_{it})$; • cross-moments of $(\log X_{it},\log P_{it})$ and $(\log X_{it},\log Y_{it})$; • cross-moments of $(X_{it}, P_{it})$, $(X_{it},Y_{it})$, $(X_{it},\log P_{it})$, and $(X_{it},\log Y_{it})$.

The moments in (ii) are the moments used in the Almost Ideal Demand System aids; the moments in (iii) are introduced in order to capture risk effects by comparison with the moments in (ii). The optimum is found using a random search algorithm over a sufficiently big support.\footnote{Random search is more efficient than grid search in hyperparameter optimization bergstra.}

To apply MSM, it is necessary to compute simulated consumption $x_{it}^s(\theta)$ for every observation, at each step of the optimization algorithm. This procedure is computati- onally costly. Note that, the number of simulated observations with positive consumption is stochastic, and not necessarily equal to the number of observations with positive consumption in the sample. This aspect has no impact on the consistency of the MSM estimator.

The estimated hyperparameter is:

equation[equation omitted — 80 chars of source]

Therefore, the median level of risk aversion for the mean of the Dirichlet process\footnote{This is not the absolute risk aversion of the utility function for the log-normal mean distribution which depends on the consumption level and has to be computed with a modified density.} for $A_1$ is $\exp(0.5495)\simeq 2.2226$, and the median level of risk aversion for the mean of the Dirichlet process for $A_2$ is $\exp(0.8738)\simeq 1.1276$. The fact that $\mu_1$ is smaller than $\mu_2$ is expected: Since quantities are measured in terms of volume of alcohol, this result is consistent with the faster overall intake of alcohol when consuming drinks with a higher ABV. Moreover, the distribution $\pi_1$ of $A_1$ is much more concentrated around its mean than the distribution $\pi_2$ of $A_2$, as the scaling parameter $c_1=45.0951$ for $\pi_1$ is much larger than the scaling parameter $c_2=3.5544$ for $\pi_2$.

We do not report any standard errors because they are automatically small from the large number of observations. Indeed, the standard significance test procedures (such as comparing a t-statistic to the critical value of a standard normal at the 5% significance level) are not relevant in this big data framework. The highest degree of uncertainty concerns the filtered functional parameters $(\hat{\pi}_m)$ since $\pi_m$ is a high-dim- ensional parameter and the number of observations in each segment is much smaller.

The means of these Dirichlet processes are displayed in the left panel in Figure (ref). The right panel displays the indifference curves associated with utility levels $-0.1000$, $-0.0800$, and $-0.0680$ for a draw from the Dirichlet process given $\hat{\theta}$.

figure[figure omitted — 680 chars of source]

Figure (ref) displays the Q-Q plots for two draws $(\pi_1^s,\pi_2^s)$, $s=1,2$, from the Dirichlet process given $\hat{\theta}$. In particular, we plot the quantiles of the realization of the distribution $\pi_j$ of $\log(A_j)$ against the quantiles of the normal distribution given the estimated hyperparameters $(\hat{\mu}_j,\hat{\sigma}_j)$, for $j=1,2$. If these quantiles coincide exactly, they will lie on the 45-degree line. As expected, these Q-Q plots lie approximately around the 45-degree line. The draws $(\pi_1^s)$, $s=1,2$, for $\pi_1$ are closer the 45-degree line and “more continuous” than the draws $(\pi_2^s)$, $s=1,2$, for $\pi_2$ since $c_1>c_2$.

figure[figure omitted — 747 chars of source]

Figure (ref) illustrates how one might use the (estimated) hyperparameter for interpretation. Specifically, it is used to deduce the mean of the Dirichlet process, which is used as a benchmark for comparison with a drawn or filtered functional parameter $\pi_m$.

Taste Distributions

This section uses the filtering approach described in Section (ref) to recover $\pi_m$. In the SARA model, the MRS restriction in (ref) is: \[ \mathbb{E}_{\pi}[A_1\exp(-A'x_{it})]=p_{it}\, \mathbb{E}_{\pi}[A_2\exp(-A'x_{it})]. \] When $A_1$ and $A_2$ are independent, this expression becomes:

equation[equation omitted — 218 chars of source]

To filter $\pi_m$, these restrictions have to be imposed for every observation with positive consumption $x_{it}$ associated with segment $\Lambda_m$. In California, there are 688 MRS rest- rictions, and, in Florida, there are 880. Appendix (ref) shows how to numerically solve the resulting optimization problem given the bilinearity of the MRS restrictions under independence.

The marginal taste distributions were filtered using a grid with 500 points between $\exp(-10)$ and $\exp(10)$, equally spaced on the log-scale. All draws from the estimated prior were simulated by the stick-breaking method given $J=1000$ breaks (see Appen- dix (ref)). Exactly $S=100$ draws from the posterior were used to filter each distribution.

figure[figure omitted — 759 chars of source]

Figure (ref) displays the Q-Q plots for the filtered taste parameters $\hat{\pi}_m$ for California and Florida. As in Figure (ref), the (estimated) hyperparameter is used to construct a benchmark for comparison. As expected, the filtered taste parameters are rather diff- erent from this benchmark. Here, the role of the estimated prior distribution diminishes with the number of observations. In both states, the slope on the left is steeper than the 45-degree line, suggesting that the posterior mean distribution for $A_1$ is more “dispersed” than its estimated prior mean distribution. The convexity of these curves also suggests fatter tails.

For the structural interpretation of these plots, assume that (i) the preferences are SARA, (ii) the taste parameters are independent, (iii) the marginal distribution of $A_1$ is the same in both states, and (iv) the marginal distribution of $A_2$ “shifts” such that $\pi_2^*(A_2) = \pi_2(cA_2)$, where $\pi_2$ and $\pi_2^*$ denote the marginal distributions of $A_2$ in these states. Under these assumptions:

equation[equation omitted — 50 chars of source]

and solving the utility maximization problem in (ref) yields:

equation[equation omitted — 129 chars of source]

Similarly, if there is a “shift” in the marginal distribution of $A_1$ and the marginal dist- ribution of $A_2$ is the same in both states, we obtain:

equation[equation omitted — 173 chars of source]

The relationships given in (ref) and (ref) suggest that there exists a complicated non-linear relationship between such demand functions. Therefore, we cannot immediately deduce from Figure (ref) which state has a higher demand for beer. For a more formal analysis, the utility functions associated with each posterior mean taste distribution must be used to derive a posterior MRS, or a posterior demand function.

figure[figure omitted — 509 chars of source]

This analysis has to be completed with a discussion of accuracy. In this non-param- etric framework, the posterior distributions of $\pi_1$ and $\pi_2$ are infinite-indimesional and cannot be represented. However, posterior distributions of any scalar transformation of $\pi_1$ and $\pi_2$ can be derived using simulation. In this respect, it is important to know which scalar objects are of interest. Typically, we are interested in the MRS evalu- ated at a specific bundle, say $x_0$, or counterfactual demand, corresponding to a particular design, say $z_0=(y_0,p_0)$. Figure (ref) displays the posterior distributions of the MRS, evaluated at two bundles, $(1,1)$ and $(1,2)$, for California and Florida. In both states, these distributions are approximately log-normal (with is consistent with dobronyi, dobronyi), and the posterior for $\text{MRS}(1,2;\pi)$ has a much longer tail than the posterior for $\text{MRS}(1,1;\pi)$, implying that, the quantity of good 2 that must be given to the consumer in order to compensate her for one unit of good 2 (and keep her just as happy) is larger, on average, when she has more of good 2. This tail is longer in California.

The filtered taste distributions in Figure (ref) are obtained by applying the algorithm in Appendix (ref) and forcing the density $\pi_m$ to be non-negative at each iteration. The existence of negative “probabilities” can be a result of numerical uncertainty, the choice of grid, or misspecification. Specifically, it can arise if the consumer in segment $\Lambda_m$ does not maximize her SARA/SSF utility function (or any utility function) subject to the linear budget constraint. By analyzing these negative probabilities, we can construct a measure of the deviation from rationality. To illustrate, let $\pi_k^+=\max\{0,\pi_k\}$ and $\pi_k^-=\max\{0,-\pi_k\}$, respectively, denote the positive and negative components of the elementary probability $\pi_k$ associated with the $k^{th}$ grid point. The following ratio:

equation[equation omitted — 73 chars of source]

is a measure of bounded rationality. This ratio ranges between 0 and 1. The closer this ratio is to 1, the less compatible the data are with the hundreds of MRS restrictions imposed by the chosen model. This ratio is related to a subset of the literature conce- rned with such measures. Existing measures include Afriat's Efficiency Index (afriat, afriat; varian-1990, varian-1990), and the Money Pump Index (echenique, echenique). In general, these indices are used to measure a single consumer's deviation from rationality by evaluating how “close” her choices are to satisfying the Generalized Axiom of Revealed Preference (GARP), a necessary and sufficient condition for a finite number of choices to be consistent with the maximization of any locally non-satiated utility function. In our framework, the BR ratio can be used to measure the violation of the homogeneous segment assumption. Table (ref) displays the BR ratios for California and Florida. The BR ratio for $\pi_1$ is smaller than the ratio for $\pi_2$ in each state; these ratios are roughly the same across states.

table[table omitted — 331 chars of source]

Concluding Remarks

This paper is one among pioneering papers attempting to tackle the challenges of performing structural demand analysis with scanner data (see also burda-2008, burda-2008, burda-2012, craw-pol, craw-pol, guha-ng, guha-ng, cher-newey, cher-newey, and \hyper@link{cite}{cite.dobronyi\@extra@b@citeb}{Dobronyi and Gouri\'{e}roux}, dobronyi). The recent availability of scanner data permits new developments in the analysis of consumer behaviour. Here, we have shown that, by introducing homogeneous segments of consumers, we can consider a model of consumption with non-parametric preferences and infinite-dimensional heterogeneity, not only from a theoretical point-of-view, but also from a practical one. The distribution of individual heterogeneity in the population can be estimated, and the underlying non-parametric preferences can be filtered by using appropriate algorithms.

We developed an analysis for two goods for exposition. This feature of our analysis leaves the question: Can the methods developed in this paper be extended to a framework with, say, 100 goods? A completely unconstrained non-parametric analysis wou- ld encounter the curse of dimensionality. Specifically, we would need to estimate the distribution of the utility function (a non-parametric function with, in this scenario, 100 arguments). This task would be infeasible, even in our big data framework. But, the SARA model with independent taste parameters is a constrained non-parametric model. The structure of the SARA model reduces the non-parametric dimension of the problem, making it feasible. Indeed, when taste parameters are independent, we only need to estimate 100 one-dimensional distributions. A similar remark applies to the algorithm used to filter the taste distributions: The two steps based on the bilinear form of the MRS restrictions in a two good setting can be replaced with 100 successive steps based on the multilinear form of MRS restrictions in a 100 good setting, without increasing the numerical complexity.

Many of the results in this paper require taste parameters to be independent, but this requirement can be relaxed. For example, we can always consider a SARA model with the following form:

equation[equation omitted — 62 chars of source]

where $A_c$ is a common component, and $A_j$ is a good-specific taste parameter, for each $j=1,2$. In such a framework, independence between $A_c$, $A_1$, and $A_2$ does not imply independence between the parameters:

equation[equation omitted — 68 chars of source]

but it does reduce the dimensionality of the problem: Instead of introducing a joint distribution $\pi$ on a space of dimension 2, the model only depends on three distributions on a space of dimension 1. This specification avoids the curse of dimensionality. (See Appendix (ref) for a discussion of identification in this case with taste dependence.)

In this paper, consumers are assumed to be rational, and divided into homogeneous segments. Since, in each segment, the demand function can be non-parametrically estimated over a subset of its domain, the analysis can be continued to develop a test of the homogeneity of each segment, or, more generally, a non-parametric method for constructing homogeneous segments.

The approach developed in this paper uses standard ideas from consumer theory to make inference. This approach is appropriate when both quantities and prices have continuous supports. This feature makes this approach valid for some levels of good, consumer, and date aggregation. Hence, this approach can be used for, say, evaluating the effect of alcohol tax on alcohol consumption, but unreasonable for analyzing how a particular consumer will choose between hundreds of different brands of whiskey. To our knowledge, the tools needed to solve such a problem have not been developed yet.

\nocite{rock,manning-et-al,brown,karni-1983,grant,blundell-wp,cher-newey,ng,guha-ng,bilinear,bilinear-vanrosen,ledoit,heyde,laplace-lognormal,blundell}