Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,255 characters · 28 sections · 132 citation commands
Consumer Theory with Non-Parametric Taste Uncertainty and Individual Heterogeneity
Keywords: Consumer Theory, Scanner Data, Stochastic Demand, Taste Heteroge- neity, Non-Parametric Model, Bayesian Approach.
The recent availability of databases containing all dated purchases made by a large nu- mber of consumers (28,036 in our application) presents a modern challenge for the eco- nometrics of demand systems, requiring new models and estimation approaches (see, for example, burda-2008, burda-2008, burda-2012, for discrete choice, and guha-ng, guha-ng, cher-newey, cher-newey, and dobronyi, dobronyi, for the first analyses of such data in the demand literature). This type of data is commonly called scanner data because its collection involves retailers or households scanning each purchased good on the date of purchase. This paper introduces two models of random utility for scanner data: the stochastic absolute risk aversion (SARA) model, and the stochastic safety-first (SSF) model. These models have the following advantages in comparison with the existing literature:
Each model is characterized by a basis of functions. This basis is used to generate a family of utility functions. A distribution is, then, placed over this family. To be precise, we start with a basis of increasing and concave functions. Let $U(x;a)$ denote an element of this basis, where $x$ is a bundle and $a\in \mathscr{A}$ is a finite-dimensional vector of taste parameters. A family of utility functions is generated by taking the convex hull of the basis. Let $U(x;\pi)=\mathbb{E}_{\pi}\big[U(x;a)\big]$ denote an element of this family, where $\pi\in \Pi$ is a distribution on $\mathscr{A}$. This family is indexed by a functional parameter $\pi$, which can be structurally interpreted as taste uncertainty (resolved after the consumer makes her decisions). The heterogeneity across consumers is introduced using a distribution $F$ on the set $\Pi$ of probability distributions $\pi$ on $\mathscr{A}$. Therefore, each model combines uncertainty and heterogeneity: the uncertainty in taste for a given consumer is represented by $\pi$, and the heterogeneity across consumers is captured by $F$.
The paper considers a two-good framework with continuous support for $x$. It is organized as follows: Section (ref) introduces the stochastic absolute risk aversion (SARA) model and Section (ref) introduces the stochastic safety-first (SSF) model. For each mo- del, we derive conditions on $\Pi$ under which there exists a unique demand system, for each $\pi\in \Pi$. In Section (ref), the distribution of heterogeneity $F$ is introduced. When $F$ is known, we obtain a Bayesian framework in which the functional parameter $\pi\in \Pi$ has to be estimated. When $F$ is a member of a known parametric family, indexed by $\theta$, we obtain an empirical Bayesian framework with a hyperparameter $\theta$ that has to be estimated, and a stochastic functional parameter $\pi$ that has to be filtered. In Section (ref), we consider the identification of the taste distribution $\pi$ within each model. Next, we examine if it is possible to distinguish between stochastic risk aversion and stochastic safety-first. In Section (ref), we use the Nielsen Homescan Consumer Panel to illustrate our methodology in an application to the consumption of alcohol. Section (ref) concludes. The details of the Dirichlet process are in Appendix (ref); integrability is discussed in Appendix (ref); an optimization procedure for filtering the taste distributions $\pi$ after estimating $F$ is in Appendix (ref); details of the data are placed in Appendix (ref).
This section introduces the first utility specification that we consider. It first describes the set of utility functions, then derives conditions under which there exists a unique demand system. The taste uncertainty is introduced using risk aversion parameters.
There are two goods, denoted $1$ and $2$. Let $\bar{R}=\mathbb{R}_{+}^2$ denote the non-negative orthant with interior $R$. A consumer has preferences over the bundles in $\bar{R}$. Her preferences are summarized by a utility function of the form:
for every $x$ such that $x_1,x_2\geq 0$, where $A=(A_1,A_2)$ is a positive stochastic parameter characterizing the consumer's degrees of absolute risk aversion with respect to goods 1 and 2, and $\pi$ is a joint distribution for this pair of stochastic taste parameters. Her preferences are, as a result, contained in a broad family of utility functions, indexed by a functional parameter $\pi$. There are two interpretations of specification (ref): (i) the preferences are summarized by a deterministic utility function in the convex hull gen- erated by a parametric family, or (ii) the consumer faces “taste uncertainty” and she resolves this uncertainty by using expected utility. We call these preferences stochastic absolute risk aversion (SARA) preferences.\footnote{These preferences differ from those used to describe consumer behaviour when facing ambiguity or uncertainty, as in, say, halevy-feltkamp.}
If $\pi$ is a point mass at $a=(a_1,a_2)$ such that $a_1,a_2>0$, the stochastic parameters are constant, and $U(x;\pi)$ reduces to $U(x;a)=-\exp(-a'x)$. This function is strictly increasing because we have:\footnote{Here, $>0$ means each component is strictly larger than $0$.}
at each $x$ such that $x_1,x_2>0$, and concave (although not necessarily strictly concave) because the Hessian associated with the utility function:
is negative semi-definite, at each $x$ such that $x_1,x_2>0$. This matrix is related to a bivariate measure of absolute risk aversion\footnote{Such a measure can be defined as: \[ -\left(\text{diag}\frac{\partial U(x;\pi)}{\partial x}\right)^{-1/2}\frac{\partial^2 U(x;a)}{\partial x\partial x'}\left(\text{diag}\frac{\partial U(x;\pi)}{\partial x}\right)^{-1/2}, \] where $\text{diag}\frac{\partial U(x;\pi)}{\partial x}$ is the diagonal matrix whose diagonal elements are the first derivatives of $U(x;\pi)$.} (richard-1975, richard-1975; karni-1979, karni-1979, karni-1983; grant, grant). These properties translate into properties of the more general function: $U(x;\pi)$.
Proposition (ref) implies that we have effectively constructed a family of well-behaved utility functions $\{U(x;\pi):\pi\in\Pi\}$ indexed by a functional parameter $\pi$, describing the taste uncertainty, instead of the standard finite-dimensional parameter usually con- sidered in the literature.
Let $g_{\pi}(\cdot)$ denote the function defined by the implicit equation:
for every $x_1\geq 0$, and each (attainable) level of utility $u<0$. This implicit equation has a unique solution because $U(x;\pi)$ is strictly increasing on $\bar{R}$. The function $g_{\pi}(\cdot,u)$ is the indifference curve associated with the functional parameter $\pi$ and a utility level of $u$---$g_{\pi}(\cdot,u)$ maps every value of $x_1$ to a value of $x_2$ for which $(x_1,x_2)$ attains a utility level of $u$ given $\pi$. The implicit function theorem implies that $g_{\pi}(\cdot)$ is twice-continuo- usly-differentiable with respect to $x_1$ and:
on $R$ where $\text{MRS}(x;\pi)\equiv \frac{\partial U(x;\pi)/\partial x_1}{\partial U(x;\pi)/\partial x_2}$ denotes the marginal rate of substitution at $x$---the rate at which the consumer is willing to exchange good 1 for good 2 given $x$ and $\pi$. The indifference curve $g_{\pi}(\cdot,u)$ is strictly convex such that:
at every $x_1>0$, since the Hessian of $U(x;\pi)$ is negative definite everywhere on $R$ (see Lemma 1 in dobronyi, dobronyi). This property is stronger than the sta- ndard assumption of strict quasi-concavity, which allows this derivative to be zero on a nowhere dense set (katzner, katzner). This distinction is important for what follows.
Note that, after integrating out the taste uncertainty, the absolute risk aversions will depend on the consumption level. For instance, when $A_1$ and $A_2$ are independent with distributions $\pi_1$ and $\pi_2$, the risk aversion for good 1 becomes:
where $U_1(x_1;\pi_1)$ denotes $\mathbb{E}_{\pi_1}[\exp(-A_1x_1)]$, the portion of the utility function $U(x;\pi)$ corresponding to good 1. Clearly, $A_1(x_1)$ depends on $x_1$. Indeed, it is the average of $A_1$ given the following modified density:
with respect to $\pi_1$.
Let $z\in R$ denote a pair $z=(y,p)$ in which $y$ denotes expenditure and $p$ denotes the price of good 1, both normalized by the price of good 2. The consumer can purchase a bundle $x\in \bar{R}$ if, and only if, $px_1+x_2\leq y$. She chooses a bundle $x\in\bar{R}$ that solves:
Let $X^*(z;\pi)$ denote the solution to:
While (ref) is restricted to bundles in the non-negative orthant, (ref) allows for neg- ative values. The solution to (ref) is characterized by the following system of first- order conditions:
The first equality says that the marginal rate of substitution equals the relative price $p$. The second equality says that the budget constraint holds with equality. Equivalently, we can solve the equality:
for the first component $X_1^*(z;\pi)$, and then use the budget constraint in (ref) to solve for $X_2^*(z;\pi)$. As long as $A_1-pA_2$ is not almost surely equal to zero, the first-order partial derivative of the left side of this equality with respect to $x_1$ is strictly negative:
The function on the left side of (ref) is, therefore, strictly decreasing in $x_1$, implying that there exists a unique solution $X_1^*(z;\pi)$ to (ref), and a unique solution $X^*(z;\pi)$ to (ref). If $X^*(z;\pi)$ is in $\bar{R}$, then $X^*(z;\pi)$ coincides with the solution to (ref). Else, the solution to (ref) is on the boundary of $\bar{R}$. Let $X(z;\pi)$ denote the solution to (ref) given both $z$ and $\pi$. There are three regimes of demand in the design space:
Because the utility function $U(x;\pi)$ has strictly convex indifference curves everywhere on $R$, the demand function $X(z;\pi)$ is invertible in the second regime (see Proposition 2 in dobronyi, dobronyi).
As a final remark, let us consider a risk-neutral consumer. In particular, let us ass- ume that $A_1$ and $A_2$ tend stochastically to zero, with means that tend to zero so that $\mathbb{E}_{\pi}[A_1]/\mathbb{E}_{\pi}[A_2]$ converges to a non-degenerate $a_0$. By considering the Taylor expansion of utility, it can be shown that these preferences are represented by:
This representation is unique up to an increasing transformation. For this risk-neutral consumer, goods are considered to be perfect substitutes. It is known that such a consumer will consume only good 1 whenever $p<a_0$, and only good 2 whenever $p>a_0$.
As an illustration, let us assume that $A_1$ and $A_2$ are independent and that $A_j$ has a Gamma distribution $\gamma(\nu_j,\alpha_j)$ with degree of freedom $\nu_j>0$ and scale factor $\alpha_j>0$, for $j=1,2$. Under this specification, $\pi=\gamma(\nu_1,\alpha_1)\otimes\gamma(\nu_2,\alpha_2)$, where $\otimes$ denotes the tensor product of distributions. By the Laplace transform of the Gamma distribution:
Under this specification, the absolute risk aversion for good 1 in (ref) becomes:
which is hyperbolic in $x_1$. The indifference curve $g_{\pi}(\cdot)$ associated with utility level $u$ is:
for every $x_1\geq 0$ and $u\in(-1,0)$ such that:
It is easily shown that the second derivative of the indifference curve $g_{\pi}(\cdot,u)$ equals:
for some $c>0$. This inequality confirms that the indifference curve $g_{\pi}(\cdot,u)$ is strictly convex. Furthermore, the MRS is equal to:
The unconstrained solution $X_1^*(z;\pi)$ to the first-order condition in (ref) is equal to:
The second component $X_2^*(z;\pi)$ is deduced from the budget constraint in (ref). By equation (ref), the demand function $X(z;\pi)$ coincides with $X^*(z;\pi)$ over the set $\mathcal{Z}$ of pairs $z$ such that:
The three regimes of demand are illustrated in Figure (ref) in the design space. The strict convexity of the indifference curve $g_{\pi}(\cdot,u)$ on $\mathcal{Z}$ implies that the demand function $X(\cdot;\pi)$ associated with this utility function is invertible on $\mathcal{Z}$.
We now consider a model with taste parameters that have a safety-first interpretation.
In Section (ref), we constructed a family of well-behaved utility functions by taking the convex hull generated by a particular basis. In this section, we consider another basis, consisting of functions with the form:
for every $x_1,x_2\geq 0$, where $x^+=\max\{0,x\}$ and $a_1,a_2>0$. This function corresponds to the “safety-first” criterion, introduced into the literature on portfolio management by roy-safety. In order to illustrate, let us consider the consumption of alcohol, as in dobronyi. Suppose that there are two groups of goods: group 1 consisting of drinks with low alcohol by volume such as beers and ciders, and group 2 consisting of drinks with high alcohol by volume such as wines and liquors. Assume that the quantities are measured in identical units such as volume of alcohol---that is, the total volume of the drink in litres multiplied by the alcohol by volume of the drink.\footnote{Quantities could be, alternatively, measured in calories.} We can, then, add these volumes to aggregate two drinks with different sizes and/or percentages of alcohol. Here, $a_1$ is the consumer's relative preference between the two groups of drinks, and $a_2$ is a “control” parameter, specifying her attempt to limit her intake of alcohol.
Now, let us introduce a distribution $\pi$ such that $\mathbb{E}_{\pi}[A_j]<\infty$, $j=1,2$, and define:
By the law of iterated expectations, we obtain:
We call these preferences stochastic safety-first (SSF) preferences.
Under mild regularity conditions:
for every $x$ such that $x_1,x_2>0$. These partial derivatives are strictly positive when $\pi$ has full support: $\pi(a_1,a_2)>0$, for $a_1,a_2>0$. By taking the second-order derivatives:
for every $x$ such that $x_1,x_2>0$, where $\pi_0\equiv\pi(x_1+A_1x_2|A_1)$ in which $\pi(\cdot|A_1)$ denotes the conditional density of $A_2$ given $A_1$, assuming that such a density exists. This matrix is both symmetric and negative definite when $\pi(\cdot|A_1)$ is continuous and $A_1$ is not constant. This result follows from the positivity of $\mathbb{E}_{\pi}[\pi_0]$ and the following equality:
which holds for every $x$ such that $x_1,x_2>0$, in which $\tilde{\pi}$ denotes the modified density:
Consequently, we have constructed another family of well-behaved utility functions $\{U(x;\pi):\pi\in\Pi\}$ indexed by a functional parameter $\pi$, describing taste uncertainty.
Let us revisit the utility maximization problem in (ref). Under the safety-first spec- ification, the analogue of the unconstrained first-order condition in (ref) is given by:
We obtain this equality by equating the marginal rate of substitution with the relative price $p$, and then using the budget constraint to replace $x_2$ with $y-px_1$. Under the regularity conditions from above, the left-hand side is strictly monotone in $x_1$ given $\pi$, so that there exists a unique solution to the first-order condition. As in Section (ref), we let $X_1^*(z;\pi)$ denote this solution, and let $X_2^*(z;\pi)$ denote the quantity $y-pX_1^*(z;\pi)$.
When the consumer's preferences are SSF, the MRS has the form:
Thus, the rate at which the consumer is willing to exchange good 1 for good 2 given $x$ and $\pi$ is equal to the inverse of the expectation of her relative preference between goods $A_1$, conditional on not surpassing her control parameter $A_2$.
Some functionals of the distribution $\pi$ can be especially interesting. For instance, in an application to the consumption of alcohol, we might expect the conditional distribution of $A_2$ given $A_1=a_1$ to be concentrated around a single mode, characterizing an implicit alcohol limit for this consumer. Then, we can ask the following questions:
These are questions that cannot be answered using classical demand systems like the Almost Ideal Demand System aids. In fact, tests based on the Almost Ideal Demand System have rejected rationality in applications to alcohol consumption all-ferg-stew. Clearly, it is possible that the Almost Ideal Demand System is misspecified.
In general, the first-order condition in (ref) has no closed-form solution. However, its expression can be simplified for some taste distributions $\pi$. As an illustration, let us assume that:
Under this specification, we can first integrate with respect to $A_2$ within the expecta- tion in (ref) in order to obtain the following condition:
Equivalently, we obtain:
This equation can be written in terms of the Laplace transform $\Psi$ for $A_1$. This yields:
which can also be written as:
Finally, by inverting this expression and rearranging the terms, we get:
The second component $X_2^*(z;\pi)$ of the unconstrained solution in (ref) is deduced from the budget constraint. It follows from equation (ref) that the demand function $X(z;\pi)$ coincides with $X^*(z;\pi)$ if, and only if:
For instance, if $A_1$ follows a gamma distribution $\gamma(\nu,\alpha)$, then $\log \Psi(v)=-\nu\log(1+v/\alpha)$, and we obtain:
for $v\geq 0$. Moreover, by inverting this function, we get:
Therefore, the solution $X_1^*(z;\pi)$ has the form:
and demand $X(z;\pi)$ coincides with $X^*(z;\pi)$ if, and only if:
The regimes of demand are illustrated in Figure (ref) in the design space. Note, we can also verify that the Slutsky coefficient is strictly negative\footnote{This property holds for any Laplace transform $\Psi$ of $A_1$ (see Appendix (ref)).} such that:
ensuring that the demand function $X(\cdot;\pi)$ is invertible over the set $\mathcal{Z}$ of pairs $z$ on which demand is strictly positive (see Section 2 in dobronyi, dobronyi).
Sections (ref) and (ref) introduced two utility specifications, both indexed by the functional parameter $\pi$. Of course, different consumers can have different functional parameters. This individual heterogeneity is introduced in a second layer, by specifying a distribution $F$ over the set $\Pi$ of distributions on $R$, such as the Dirichlet process (see, for example, navarro, navarro, for an application of the Dirichlet process in modelling individual differences). More precisely, we make the following theoretical assumption:
Assumption (ref) introduces a distribution $F$ over the functional taste parameter $\pi$. This distribution $F$ characterizes the heterogeneity across homogeneous groups. It can encompass, for example, regional or demographic differences in preferences. This infinite-dimensional heterogeneity is non-separable in the stochastic demand equation.
The Dirichlet process can be constructed in three steps:
Let us now discuss implications of Assumption (ref): If the functional and scaling parameters of the Dirichlet process are known, then we are in a Bayesian framework (see, for example, geweke, geweke, for a Bayesian analysis of revealed preference) in which the taste distribution $\pi\in \Pi$ has to be estimated. Otherwise, we can assume that the mean $\mu$ of our process $F$ is characterized by a finite-dimensional hyperparameter $\theta$. Naturally, the hyperparametric model has two types of parameters: the hyperparameter $\theta$ to be estimated, and the functional parameters $(\pi_m)$ to be filtered.
In this section, we consider the identification of the functional parameter $\pi$ within each model from the observation of a demand function. Then, we examine if we can distin- guish between the SARA and SSF models.
Intuitively, a consumer's demand function is identified if we observe her making a lot of consumption decisions at a variety of designs $z$. Clearly, we can identify her demand function if (i) her preferences are constant over time and we observe a large panel or experiment,\footnote{In this case, when the number of dates $T$ is large, we can have a segment $m$ for each consumer $i$.} or (ii) she belongs to a large homogeneous segment of consum- ers with identical preferences. This explains the form of Assumption (ref) (as it allows for either interpretation). Later, we apply the segmented approach to scanner data in the application to the consumption of alcohol in Section (ref).
With panel data, one no longer requires the assumption that demand is monotonic with respect to unobserved heterogeneity in order to achieve identification (see Figure (ref), and the role of this assumption in brown-matzkin, brown-matzkin, matzkin-nonadd, matzkin-nonadd, and hn-individual-het, hn-individual-het).
In the models introduced in Sections (ref) and (ref), and for any $\pi$ such that demand is inv- ertible, we can derive the inverse demand function, whose second component coincides with the MRS which can be integrated to obtain a unique preference ordering. Indeed, by construction, the integrability conditions (needed to recover a unique well-behaved preference ordering) are satisfied, implying that preferences are recoverable (see samuelson1948, samuelson1948, for a seminal discussion of integrability in the case of two goods, and samuelson, samuelson, hurwicz-uzawa, hurwicz-uzawa, and hosoya2016, hosoya2016, for general approaches). However, the possibility to recover preferences from a consumer's demand function does not imply that the distribution of taste uncertainty $\pi$ is identified. Indeed, two distinct taste distributions could produce an identical MRS.
For identification, we only consider the information contained in the demand function $X(\cdot;\pi)$ on the set $\mathcal{Z}$ of designs $z$ for which the components of the demand function are strictly positive. This restriction disregards some information that may be available in the first or third regimes of (ref). In most datasets, when a component of the demand function equals zero, the price $p$ is not observed.
In the stochastic absolute risk aversion (SARA) model, the identification condition is:
In the degenerate case in which $A$ is deterministic and equal to $(a_1,a_2)$, the MRS red- uces to $a_1/a_2$. Thus, in this special case, the two-dimensional parameter $a=(a_1,a_2)$ is identified up to a positive factor. This reasoning leads us to a question: Does this lack of identification also exist in an extended setting?
Let us first remark that the utility function $U(x;\pi)$ is equal to the moment generating function for $\pi$ with a negative sign: $\Phi(x;\pi)=-U(x;\pi)$. Because this moment generating function characterizes $\pi$ when the stochastic parameter $A$ is non-negative (see Theorem 1a in Chapter 13 on Tauberian Theorems in feller-1968, feller-1968), it is equivalent to consider the identification of either $\pi$, or $\Phi(x;\pi)$.\footnote{Note, the existence of the moment generating function does not imply the existence of all power moments and, even if all power moments exist, they do not necessarily characterize the distribution. A known example is the log-normal distribution used in the application heyde.} As mentioned, we can always integrate the MRS to recover a unique preference ordering. That is, we can recover $U(x;\pi)$ up to a monotonic transformation. We still need to discern the conditions on $\pi$ under which we can recover $\Phi(x;\pi)$. Indeed, moment generating functions have properties that are not necessarily preserved under monotonic transformations.
We obtain the following result:
This means that we can, at most, identify the class of moment generating functions $\mathscr{C}(\Phi)=\{\Phi^{\nu}:\nu>0\}$. Note that, for any moment generating function $\Phi$, the transf- ormed function $\Phi^{\nu}$ is also a moment generating function.
Let us now consider identification when $A_1$ and $A_2$ are independent:
Proposition (ref) implies that $\mathscr{C}(\Phi)$ is identified under the independence of $A_1$ and $A_2$. Indeed, we can recover the consumer's preference ordering using traditional methods, and use the fact that all admissible preference orderings map to a unique class $\mathscr{C}(\Phi)$.
Of course, independence is a strong restriction. In the SARA model, it is equivalent to the additive separability of the utility function.\footnote{In the case of two goods, additive separability is stronger than separability.} To see this result, notice that, under independence, we obtain:
Since utility functions are unique up to strictly increasing transformations, this utility function is equivalent to:
which is an additively separable utility function. In Appendix (ref), we prove a generalization of Proposition (ref) where stochastic taste parameters have a common component.
In the SSF model, the identification condition is:
Let us now consider the validity of this condition under an independence assumption. Note, in the SSF model, independence is no longer equivalent to additive separability.
Proposition (ref) provides no information on the identifiability of the distribution of $A_1$ beyond its first moment. It seems difficult to obtain a general identification result, but insights into our identification problem can be obtained by considering the two primary families of distributions that are invariant to positive power transformations, that are, the exponential family and the Pareto family.
Once the identification of the consumer's taste distribution $\pi$ within each model is solved, we still need to consider the identification between the models. This analysis is needed to test whether preferences are consistent with SARA, or SSF, or both. It is important to know whether these two classes of preferences are nested or non-nested. If they are non-nested, we need to characterize their intersection and define a general class encompassing both types of preferences.
To illustrate, suppose that the consumer has SSF preferences:
where (i) $A_1$ and $A_2$ are independent, (ii) $A_1$ has distribution $\pi_2$, and (iii) $A_2$ follows an exponential distribution (with unit intensity). Under this specification, we obtain:
To clarify this result, observe that, by conditioning on $A_1$, we are left with the expec- tation of the minimum of a set containing a constant and a random variable with an exponential distribution. This utility function is a strictly increasing transformation of a SARA utility function:
where (i) $B_1$ follows a point mass at $1$, and (ii) $B_2$ has distribution $\pi_2$. Consequently, these utility functions, one SARA, and the other SSF, induce the same preference ordering over the consumption set.
The possible lack of identification of each consumer's taste distribution $\pi_m$ has to be taken into account in the economic interpretation of the results. However, it has to be noted that it does not create difficulties for structural inference, where the (scalar or functional) parameters of interest are the parameters characterizing the MRS, rather than the parameters characterizing the utility function.
The lack of identification is due to the special structure of the cone of increasing and concave functions defined on $R$, and of the extremal elements of this cone. For finite increasing concave functions defined on $\mathbb{R}_+$, it is well-known that the extremal functions are of the type:
in which $(\alpha_j,\beta_j)\in \bar{R}$, for $j=1,2$ (see blaschke-pick, blaschke-pick), and that any finite positive increasing concave function can be written as:
where $b$ is a positive scalar and $\pi$ is the distribution of $A$. Such functions are charac- terized by $b$ and $\pi$. The set of extremal functions in (ref) is a minimal set of extremal points generating the cone.
Such a property no longer holds for finite positive increasing concave functions defined on $\bar{R}$. johansen has described a large set of extremal points of the type:
for which $h_1(\cdot)$ induces a covering with vertices of order 3 (see page 62 in johansen, johansen), and has shown that this set is dense in the cone of finite continuous convex functions defined on a convex set in $R$ (see Theorem 2 in johansen, johansen). A minimal set of extremal points generating this cone does not exist. This argument explains why Sections (ref) and (ref) consider specific convex subsets generated by parametric functions.
While we restrict our attention to SARA and SSF preferences (because the stochastic taste parameters have clear interpretations in these models), other convex hulls co- uld have been considered. For example:
The Stone-Geary basis $U(x;a)$ above can be adjusted to define another parametric basis. In particular, let us apply the transformation $\varphi(x)=-\exp(-x)$ to the Stone-Geary utility function. This transformation yields:
This utility function forms a well-behaved basis because it is strictly increasing with a negative semi-definite Hessian. While $U(x;a)$ and $\tilde{U}(x;a)$ represent the same preference ordering, they will generate different families due to the strict concavity of $\varphi(\cdot)$. To illustrate, suppose that the stochastic parameters, $A_1$ and $A_2$, are independently distributed with respect to uniform distributions on $[0,1]$. This specification produces:
While $U(x;\pi)$ is a Stone-Geary utility function, $\tilde{U}(x;\pi)$ is a complicated non-linear function of $x_1$ and $x_2$. Consequently, an uninteresting basis has been transformed into an interesting one. This procedure can be completed for any increasing, concave, and twice-differentiable transformation $\varphi(\cdot)$.
This section shows how to use the SARA and SSF models in a non-parametric framework. First, we specify the statistical model by introducing an assumption on the obs- ervations, and then we discuss statistical inference. The methodology is illustrated in an application to alcohol consumption using scanner data concerning individual purchase histories.
The behavioural models introduced in the previous sections can be completed with an assumption on the available observations. We consider panel data, indexed by the consumer $i$ and date $t$. After a preliminary treatment of the purchase histories, we have a large number $n$ of consumers and a fixed number $T$ of observed dates. In the preliminary treatment, the goods are aggregated into two groups using a common quantity unit and the dated purchases are aggregated by month (see Section (ref)). Recall that, under Assumption (ref), we have $M$ segments of homogeneous consumers.
We introduce the following assumption on the observations:
Assumption (ref) describes the structure of the observations. It implies that we can imagine taste parameters $(\pi_m)$ being independently drawn from a Dirichlet process $F$, designs $(z_{it})$ being independently drawn from some distribution, and consumption $x_{it}$ satisfying $x_{it}=X(z_{it};\pi_{m_i})$, where $m_i$ is the group of consumer $i$. Many papers assume that consumption $x_{it}$ is positive (see Section IV.A in \hyper@link{cite}{cite.blundell\@extra@b@citeb}{Blundell, Horowitz, and Parey}, \hyper@link{cite}{cite.blundell\@extra@b@citeb}{2017}, for this assumption in an application to gasoline demand, as well as Assumption A5 in dobronyi, dobronyi, for this assumption in an application to the consumption of alcohol); the SARA and SSF models allow for corner solutions. However, in many datasets (including the dataset used in the application in Section (ref)), there is a problem of partial observability. Let $\tilde{y}$ denote the expenditure (prior to normalization), and let $\tilde{p}_j$ denote the price of good $j$ (prior to normalization). Usually, we only observe the price $\tilde{p}_j$ of a good $j$ when the consumer buys a positive quantity of good $j$. Then, we only observe (normalized) expenditure $y$ when the consumer buys a positive quantity of good 2, and we only observe the (normalized) price $p$ when the consumer buys a positive quantity of both goods (see craw-pol, craw-pol, for an approach to revealed preference that deals with this partial observability problem).\, This problem explains the specific form of Assumption \hyperref[ass:2]{2(i)}.
For deriving the asymptotic properties of estimators, it is also necessary to specify the type of asymptotics to be considered:
Assumptions \hyperref[ass:3]{A3(i)} and \hyperref[ass:3]{A3(ii)} ensure that there are enough observations to non-parametrically estimate the demand function associated with the functional parameter $\pi_m$ on a sufficiently large subset $\mathcal{Z}_m$ of designs $z$. Assumption \hyperref[ass:3]{A3(iii)} guarantees enough filtered parameters $\hat{\pi}_m$ to estimate the underlying Dirichlet process $F$. In some special circumstances, $T$ is large, and Assumption (ref) can be used with $m=i$ and $M=n$---that is, a single consumer per group. Otherwise, grouping of homogeneous consumers is needed to identify the demand functions on sufficiently large subsets $\mathcal{Z}_m$.
The Dirichlet process is common in Bayesian estimation (see, for instance, ferguson, ferguson, and li-2019, li-2019). This process is useful because it is flexible and, if observations are independently and identically drawn from an unknown distribution, the posterior distribution of this distribution has a closed-form expression. However, our framework is much more complicated for two reasons:
These features of our model explain why estimation requires specific numerical algor- ithms. These specific algorithms have to be able to deal with the non-linear and high- dimensional features of the models. In the Bayesian framework, the Dirichlet process is fixed. In the hyperparametric framework, it is parameterized by a vector $\theta$. These parameters have to be estimated and the functional parameters $(\pi_m)$ have to be filter- ed. These estimation approaches are described below.
In a pure Bayesian framework, a Dirichlet process is fixed by selecting a mean distribution $\mu$ and a scaling parameter $c$ (see Appendix (ref)). This distribution defines the common prior for the functional taste parameters $(\pi_m)$. After, the data are used to compute the posterior distribution for the functional parameters $(\pi_m)$. Under Assu- mption (ref), the posterior distribution can be computed separately for each homogeneous group of consumers: \[ \ell(\pi_m|x_{it},z_{it},x_{it}>0,i\in\Lambda_m, t=1,\dots,T), \] where $\Lambda_m$ denotes the group of consumers with preferences characterized by the taste parameter $\pi_m$. This approach does not have to account for the potential identification problem discussed in Section (ref). If a specific characteristic of $\pi_m$ is weakly identified, its posterior distribution will be close to the prior distribution.
In our framework, the observations $(x_{it},z_{it})$, conditional on $x_{it}>0$, must satisfy the deterministic first-order conditions implied by the model. These conditions have the following form:
for any observed pair $(x_{it},z_{it})$. Equivalently:
for any observed pair $(x_{it},z_{it})$. These conditions are moment restrictions, called MRS restrictions. In our big data framework, the number of MRS restrictions is very large, typically several hundred to a thousand. The posterior of $\pi_m$ is simply the distribution of $\pi_m$ given these deterministic restrictions on $\pi_m$. If the taste parameters, $A_1$ and $A_2$, are independent with marginal distributions, $\pi_1$ and $\pi_2$, respectively, then the MRS restrictions are bilinear in $\pi_1$ and $\pi_2$---specifically, these restrictions are linear in $\pi_1$ given $\pi_2$, and linear in $\pi_2$ given $\pi_1$. Later, this property is used to construct a numerically efficient optimization algorithm for filtering all the $\pi_m$ (see Appendix (ref)).
The hyperparametric (or empirical Bayesian) framework is a complicated non-linear state-space model with two layers of latent state variables. Such a framework can be characterized as follows:
We have partial observability of the demand function because the value of demand $X(z;\pi_m)$ is observed at finitely many designs $z$. Furthermore, unlike most state-space models, the state variables are infinite-dimensional.
While it is difficult to derive analytically the distribution of $X_{it}$ given $Z_{it}$, it is easy to simulate its distribution for a given value of $\theta$ (see Appendix (ref) for simulations from the Dirichlet distribution). Therefore, $\theta$ can be estimated by the method of simulated moments (MSM), or indirect inference (see mcfadden-msm, mcfadden-msm, pakes-optimization, pakes-optimization, and gou-mon, gou-mon). That is, $\theta$ is estimated by matching some sample and simulated moments of the pair $(X_{it},Z_{it})$.
To illustrate, consider a pure panel such that $M=n$.\footnote{When $M<n$, we simulate $n_mT$ observations for the $m^{th}$ draw from the Dirichlet process.} The steps are the following:
Given the estimated hyperparameter $\hat{\theta}$, the taste distributions $(\pi_m)$ must be filtered. This step is equivalent to applying the Bayesian approach with the estimated Dirichlet distribution as the prior distribution (see Appendix (ref)).
Under Assumptions (ref) to (ref), the estimator for $\theta$ is consistent and asymptotically normal, and it converges at a speed of $1/\sqrt{nT}$. The derivation of the asymptotic pro- perties of the filtered functional parameter $\hat{\pi}_m$ is out of the scope of this paper and left for future research.
Once the hyperparameter $\theta$ is estimated, we can filter $\pi_m$ by using the following steps:
This procedure involves a high-dimensional argument $\pi_m$, and a very large number of MRS restrictions. Indeed, we need several hundred grid points for $\pi_m$, and, in the application, we have about one-thousand MRS restrictions, for each $m=1,\dots,M$. If the taste parameters, $A_1$ and $A_2$, are independent, this procedure can be numerically simplified by using the fact that these restrictions are bilinear (see Section (ref) and Appendix (ref)).
We use the Nielsen Homescan Consumer Panel (NHCP). Nielsen provides a sample of households with barcode scanners. Households are asked to scan all purchased goods on the date of each purchase. The prices are entered by the households or linked to retailer data by The Nielsen Company. The households that agree to participate are compensated through benefits and lotteries.
We focus on the consumption of alcoholic drinks (see manning-et-al, manning-et-al, for an application to alcohol consumption in economics). We classify drinks by type. Good 1 contains beers and ciders.\footnote{The NHCP classifies ciders as wine, by default. We reclassify these beverages using product desc- riptions because most ciders have a low alcohol by volume (ABV).} Good 2 contains wines and liquors. We disregard all non-alcoholic beers, ciders, and wines. We are left with 30,635 beers and ciders, and 108,439 wines and liquors, for a total of 139,074 drinks. We convert all measurement units to litres of alcohol by first converting all units to litres and then multiplying by the standard alcohol by volume (ABV) in each subgroup---specifically, 4.5% for beer and cider, 11.6% for wine, and 37% for liquor. For example, if a household buys two packs of six bottles of beer and each bottle contains 355 millilitres of beer, then the household buys 4.26 litres of this beer, or $4.26\times 0.045=0.231$ litres of alcohol. We use the standard ABV in each subgroup as a result of data limitations. Our sample only contains purchases made at stores, not purchases made at bars, or restaurants.
Measuring quantities in litres of alcohol has at least three advantages: (i) it can account for a quality effect, (ii) it is appropriate for analyzing most relevant structural objects (e.g. the effect of a change in taxation on alcohol consumption), and (iii) it yi- elds continuous quantities, permitting the application of standard tools in consumer theory (which could not be used if quantities were measured in, for example, bottles), and avoiding some common identification issues in the literature.
We restrict our sample to purchases made from August to November in 2016. This relatively short window is used to diminish the impact of changing tastes and product availability, and to avoid most federal holidays in the United States that are often associated with alcohol consumption such as Independence Day, Christmas Day, and New Year's Eve. Our sample contains 28,036 households. Some additional details of this restricted sample are placed in Appendix (ref).
The dated purchases are aggregated by month. For each household and month, the prices are constructed by dividing the total expenditure for each aggregate good (after accounting for the value of coupons) by the amount of alcohol of that aggregate good purchased by the household, when this amount is strictly positive. Then, we norm- alize by the price of good 2. This procedure yields four monthly observations per hou- sehold for a total of 112,144. A total of 63,936 observations have positive consumption such that $x_{it}> 0$. Table (ref) gives summary statistics conditional on $x_{it}\neq 0$. The prices $(p_{it})$ are conditional on being well-defined (see the discussion of partial observability on pages 23 and 24). For the interpretation of the results, recall that $\tilde{y}$ is the expenditure (prior to normalization), and that $\tilde{p}_j$ is the price of good $j$ (prior to normal- ization).
There are four regimes of observations: (i) zero expenditure on all goods, (ii) zero expenditure on good 1 and strictly positive expenditure on good 2, (iii) strictly pos- itive expenditure on good 1 and zero expenditure on good 2, and (iv) strictly positive expenditure on all goods. Table (ref) provides the proportion of observations in each regi- me, and shows a large proportion of observations with zero expenditure. Recall, under Assumption (ref), designs $z_{it}$ are drawn from a distribution. Therefore, we can interpret this result as a mass at zero in the marginal distribution of expenditure.
Figure (ref) displays the sample distribution of expenditure $\tilde{y}_{it}$ by regime: the distribution of expenditure $\tilde{y}_{it}$ conditional on $x_{it}>0$ is on the left; the sample distributions of expenditure $\tilde{y}_{it}$ for the two other regimes with positive expenditure are on the right. The shape of the sample distribution of expenditure $\tilde{y}_{it}$ does not appear to vary all that much with the regime. That being said, the sample distribution conditional on $x_{it}>0$ has more probability attributed to higher expenditures.
Figure (ref) compares the sample distributions of prices $\tilde{p}_j$ by regime: the sample dis- tributions of $\tilde{p}_1$ are on the left; the sample distributions of $\tilde{p}_2$ are on the right. Although the sample distribution of $\tilde{p}_1$ differs from the sample distribution of $\tilde{p}_2$, these distributions do not seem to be affected by the regime.
Figure (ref) displays the sample distributions of (normalized) designs $z_{it}=(y_{it},p_{it})$ and the components of consumption $x_{it}$ given $x_{it}>0$. As expected, the components of consumption $x_{it}$ are increasing in expenditure $y_{it}$. Furthermore, the first component of consumption $x_{it}$ is more affected by changes in the price $p_{it}$ than the second component.
Since we consider a rather short window of time, we follow the segmented population approach. We segment the population by state. Large states (e.g. California) are segmented again by county. Specifically, a county is given its own segment if it has more than 70 observations with positive consumption and it is in a state with more than 1,000 observations with positive consumption. We are left with a total of 65 seg- ments, each corresponding to a state or county. The smallest segment is Wyoming, containing 15 observations with positive consumption; the largest state is Florida (after removing Broward, Hillsborough, Palm Beach, Pinellas, and Miami-Dade counties), containing 880 observations with positive expenditure; the mean number of observations with positive consumption per segment is approximately 226.
Figure (ref) displays the sample distributions of (normalized) designs $z_{it}=(y_{it},p_{it})$ and the first component of consumption $x_{it}$ given $x_{it}>0$ in two of the larger segments: California (after removing Almeda, Los Angeles, Orange, Riverside, Sacramento, San Bernardino, and San Diego counties), and Florida (after removing Broward, Hillsborough, Palm Beach, Pinellas, and Miami-Dade counties).
Figure (ref) displays the Nadaraya-Watson (kernel) estimates of the demand function for beer conditional on $x_{it}>0$ in California and Florida over a subset of the domain of designs. Demand for beer in California is lower and less responsive to price changes than in Florida. Figure (ref) displays Engel curves for good 1 in California and Florida given $p\equiv\tilde{p}_1/\tilde{p}_2=4$.\footnote{This price is chosen to be in a sufficiently dense region of the sample distribution (see Figure (ref)).} These Engel curves cross.
As an illustration, we consider the SARA model in the hyperparametric framework. We assume that the taste parameters, $A_1$ and $A_2$, are independent. Under this assum- ption, the taste uncertainty is characterized by the marginal distributions, $\pi_1$ and $\pi_2$. The marginal distribution $\pi_j$ of $A_j$ is independently drawn from a Dirichlet process $F_j$, $j=1,2$. The mean of $F_j$ is a log-normal distribution with parameters $\mu_j$ and $\sigma_j$, and the scale parameter of $F_j$ is $c_j$. The utility function corresponding to this log-normal mean distribution, say $\bar{\pi}_j$, has a quasi closed-form expression. Indeed, under this distribution, we obtain: \[ \log(A_j)=\mu_j+\sigma_j\varepsilon_j, \; \; \forall j=1,2, \] where $\varepsilon_j$ is distributed with respect to a standard normal distribution. Then: \[
\] where $w(x)$ is the Lambert function, defined by the implicit equation: \[ w(x)\exp(w(x))=x, \] [see equation (1.3) in laplace-lognormal, laplace-lognormal]. By drawing from the Dirichlet process, we will draw a stochastic utility function around the closed-form expression above. The hyperparameter $\theta$ has six components such that: \[ \theta=(\mu_1,\sigma_1,c_1,\mu_2,\sigma_2,c_2). \]
As described in Section (ref), the first step involves estimating the hyperparameter $\theta$ using the Method of Simulated Moments (MSM). The hyperparameter $\theta$ is calibrated by using the following (sample and simulated) moments computed for all of the 63,936 observations with positive consumption:
The moments in (ii) are the moments used in the Almost Ideal Demand System aids; the moments in (iii) are introduced in order to capture risk effects by comparison with the moments in (ii). The optimum is found using a random search algorithm over a sufficiently big support.\footnote{Random search is more efficient than grid search in hyperparameter optimization bergstra.}
To apply MSM, it is necessary to compute simulated consumption $x_{it}^s(\theta)$ for every observation, at each step of the optimization algorithm. This procedure is computati- onally costly. Note that, the number of simulated observations with positive consumption is stochastic, and not necessarily equal to the number of observations with positive consumption in the sample. This aspect has no impact on the consistency of the MSM estimator.
The estimated hyperparameter is:
Therefore, the median level of risk aversion for the mean of the Dirichlet process\footnote{This is not the absolute risk aversion of the utility function for the log-normal mean distribution which depends on the consumption level and has to be computed with a modified density.} for $A_1$ is $\exp(0.5495)\simeq 2.2226$, and the median level of risk aversion for the mean of the Dirichlet process for $A_2$ is $\exp(0.8738)\simeq 1.1276$. The fact that $\mu_1$ is smaller than $\mu_2$ is expected: Since quantities are measured in terms of volume of alcohol, this result is consistent with the faster overall intake of alcohol when consuming drinks with a higher ABV. Moreover, the distribution $\pi_1$ of $A_1$ is much more concentrated around its mean than the distribution $\pi_2$ of $A_2$, as the scaling parameter $c_1=45.0951$ for $\pi_1$ is much larger than the scaling parameter $c_2=3.5544$ for $\pi_2$.
We do not report any standard errors because they are automatically small from the large number of observations. Indeed, the standard significance test procedures (such as comparing a t-statistic to the critical value of a standard normal at the 5% significance level) are not relevant in this big data framework. The highest degree of uncertainty concerns the filtered functional parameters $(\hat{\pi}_m)$ since $\pi_m$ is a high-dim- ensional parameter and the number of observations in each segment is much smaller.
The means of these Dirichlet processes are displayed in the left panel in Figure (ref). The right panel displays the indifference curves associated with utility levels $-0.1000$, $-0.0800$, and $-0.0680$ for a draw from the Dirichlet process given $\hat{\theta}$.
Figure (ref) displays the Q-Q plots for two draws $(\pi_1^s,\pi_2^s)$, $s=1,2$, from the Dirichlet process given $\hat{\theta}$. In particular, we plot the quantiles of the realization of the distribution $\pi_j$ of $\log(A_j)$ against the quantiles of the normal distribution given the estimated hyperparameters $(\hat{\mu}_j,\hat{\sigma}_j)$, for $j=1,2$. If these quantiles coincide exactly, they will lie on the 45-degree line. As expected, these Q-Q plots lie approximately around the 45-degree line. The draws $(\pi_1^s)$, $s=1,2$, for $\pi_1$ are closer the 45-degree line and “more continuous” than the draws $(\pi_2^s)$, $s=1,2$, for $\pi_2$ since $c_1>c_2$.
Figure (ref) illustrates how one might use the (estimated) hyperparameter for interpretation. Specifically, it is used to deduce the mean of the Dirichlet process, which is used as a benchmark for comparison with a drawn or filtered functional parameter $\pi_m$.
This section uses the filtering approach described in Section (ref) to recover $\pi_m$. In the SARA model, the MRS restriction in (ref) is: \[ \mathbb{E}_{\pi}[A_1\exp(-A'x_{it})]=p_{it}\, \mathbb{E}_{\pi}[A_2\exp(-A'x_{it})]. \] When $A_1$ and $A_2$ are independent, this expression becomes:
To filter $\pi_m$, these restrictions have to be imposed for every observation with positive consumption $x_{it}$ associated with segment $\Lambda_m$. In California, there are 688 MRS rest- rictions, and, in Florida, there are 880. Appendix (ref) shows how to numerically solve the resulting optimization problem given the bilinearity of the MRS restrictions under independence.
The marginal taste distributions were filtered using a grid with 500 points between $\exp(-10)$ and $\exp(10)$, equally spaced on the log-scale. All draws from the estimated prior were simulated by the stick-breaking method given $J=1000$ breaks (see Appen- dix (ref)). Exactly $S=100$ draws from the posterior were used to filter each distribution.
Figure (ref) displays the Q-Q plots for the filtered taste parameters $\hat{\pi}_m$ for California and Florida. As in Figure (ref), the (estimated) hyperparameter is used to construct a benchmark for comparison. As expected, the filtered taste parameters are rather diff- erent from this benchmark. Here, the role of the estimated prior distribution diminishes with the number of observations. In both states, the slope on the left is steeper than the 45-degree line, suggesting that the posterior mean distribution for $A_1$ is more “dispersed” than its estimated prior mean distribution. The convexity of these curves also suggests fatter tails.
For the structural interpretation of these plots, assume that (i) the preferences are SARA, (ii) the taste parameters are independent, (iii) the marginal distribution of $A_1$ is the same in both states, and (iv) the marginal distribution of $A_2$ “shifts” such that $\pi_2^*(A_2) = \pi_2(cA_2)$, where $\pi_2$ and $\pi_2^*$ denote the marginal distributions of $A_2$ in these states. Under these assumptions:
and solving the utility maximization problem in (ref) yields:
Similarly, if there is a “shift” in the marginal distribution of $A_1$ and the marginal dist- ribution of $A_2$ is the same in both states, we obtain:
The relationships given in (ref) and (ref) suggest that there exists a complicated non-linear relationship between such demand functions. Therefore, we cannot immediately deduce from Figure (ref) which state has a higher demand for beer. For a more formal analysis, the utility functions associated with each posterior mean taste distribution must be used to derive a posterior MRS, or a posterior demand function.
This analysis has to be completed with a discussion of accuracy. In this non-param- etric framework, the posterior distributions of $\pi_1$ and $\pi_2$ are infinite-indimesional and cannot be represented. However, posterior distributions of any scalar transformation of $\pi_1$ and $\pi_2$ can be derived using simulation. In this respect, it is important to know which scalar objects are of interest. Typically, we are interested in the MRS evalu- ated at a specific bundle, say $x_0$, or counterfactual demand, corresponding to a particular design, say $z_0=(y_0,p_0)$. Figure (ref) displays the posterior distributions of the MRS, evaluated at two bundles, $(1,1)$ and $(1,2)$, for California and Florida. In both states, these distributions are approximately log-normal (with is consistent with dobronyi, dobronyi), and the posterior for $\text{MRS}(1,2;\pi)$ has a much longer tail than the posterior for $\text{MRS}(1,1;\pi)$, implying that, the quantity of good 2 that must be given to the consumer in order to compensate her for one unit of good 2 (and keep her just as happy) is larger, on average, when she has more of good 2. This tail is longer in California.
The filtered taste distributions in Figure (ref) are obtained by applying the algorithm in Appendix (ref) and forcing the density $\pi_m$ to be non-negative at each iteration. The existence of negative “probabilities” can be a result of numerical uncertainty, the choice of grid, or misspecification. Specifically, it can arise if the consumer in segment $\Lambda_m$ does not maximize her SARA/SSF utility function (or any utility function) subject to the linear budget constraint. By analyzing these negative probabilities, we can construct a measure of the deviation from rationality. To illustrate, let $\pi_k^+=\max\{0,\pi_k\}$ and $\pi_k^-=\max\{0,-\pi_k\}$, respectively, denote the positive and negative components of the elementary probability $\pi_k$ associated with the $k^{th}$ grid point. The following ratio:
is a measure of bounded rationality. This ratio ranges between 0 and 1. The closer this ratio is to 1, the less compatible the data are with the hundreds of MRS restrictions imposed by the chosen model. This ratio is related to a subset of the literature conce- rned with such measures. Existing measures include Afriat's Efficiency Index (afriat, afriat; varian-1990, varian-1990), and the Money Pump Index (echenique, echenique). In general, these indices are used to measure a single consumer's deviation from rationality by evaluating how “close” her choices are to satisfying the Generalized Axiom of Revealed Preference (GARP), a necessary and sufficient condition for a finite number of choices to be consistent with the maximization of any locally non-satiated utility function. In our framework, the BR ratio can be used to measure the violation of the homogeneous segment assumption. Table (ref) displays the BR ratios for California and Florida. The BR ratio for $\pi_1$ is smaller than the ratio for $\pi_2$ in each state; these ratios are roughly the same across states.
This paper is one among pioneering papers attempting to tackle the challenges of performing structural demand analysis with scanner data (see also burda-2008, burda-2008, burda-2012, craw-pol, craw-pol, guha-ng, guha-ng, cher-newey, cher-newey, and \hyper@link{cite}{cite.dobronyi\@extra@b@citeb}{Dobronyi and Gouri\'{e}roux}, dobronyi). The recent availability of scanner data permits new developments in the analysis of consumer behaviour. Here, we have shown that, by introducing homogeneous segments of consumers, we can consider a model of consumption with non-parametric preferences and infinite-dimensional heterogeneity, not only from a theoretical point-of-view, but also from a practical one. The distribution of individual heterogeneity in the population can be estimated, and the underlying non-parametric preferences can be filtered by using appropriate algorithms.
We developed an analysis for two goods for exposition. This feature of our analysis leaves the question: Can the methods developed in this paper be extended to a framework with, say, 100 goods? A completely unconstrained non-parametric analysis wou- ld encounter the curse of dimensionality. Specifically, we would need to estimate the distribution of the utility function (a non-parametric function with, in this scenario, 100 arguments). This task would be infeasible, even in our big data framework. But, the SARA model with independent taste parameters is a constrained non-parametric model. The structure of the SARA model reduces the non-parametric dimension of the problem, making it feasible. Indeed, when taste parameters are independent, we only need to estimate 100 one-dimensional distributions. A similar remark applies to the algorithm used to filter the taste distributions: The two steps based on the bilinear form of the MRS restrictions in a two good setting can be replaced with 100 successive steps based on the multilinear form of MRS restrictions in a 100 good setting, without increasing the numerical complexity.
Many of the results in this paper require taste parameters to be independent, but this requirement can be relaxed. For example, we can always consider a SARA model with the following form:
where $A_c$ is a common component, and $A_j$ is a good-specific taste parameter, for each $j=1,2$. In such a framework, independence between $A_c$, $A_1$, and $A_2$ does not imply independence between the parameters:
but it does reduce the dimensionality of the problem: Instead of introducing a joint distribution $\pi$ on a space of dimension 2, the model only depends on three distributions on a space of dimension 1. This specification avoids the curse of dimensionality. (See Appendix (ref) for a discussion of identification in this case with taste dependence.)
In this paper, consumers are assumed to be rational, and divided into homogeneous segments. Since, in each segment, the demand function can be non-parametrically estimated over a subset of its domain, the analysis can be continued to develop a test of the homogeneity of each segment, or, more generally, a non-parametric method for constructing homogeneous segments.
The approach developed in this paper uses standard ideas from consumer theory to make inference. This approach is appropriate when both quantities and prices have continuous supports. This feature makes this approach valid for some levels of good, consumer, and date aggregation. Hence, this approach can be used for, say, evaluating the effect of alcohol tax on alcohol consumption, but unreasonable for analyzing how a particular consumer will choose between hundreds of different brands of whiskey. To our knowledge, the tools needed to solve such a problem have not been developed yet.
\nocite{rock,manning-et-al,brown,karni-1983,grant,blundell-wp,cher-newey,ng,guha-ng,bilinear,bilinear-vanrosen,ledoit,heyde,laplace-lognormal,blundell}