Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,140 characters · 11 sections · 110 citation commands
Identification of Random Coefficient Latent Utility Models
Latent utility models with linear random coefficients have been extensively used. They have a long history in discrete choice,\footnote{heckman2001micro attributes the first use to domenich1975urban in economics.} and have become increasingly popular due to computational advances (see e.g. train2009discrete). They now form the core demand system of most applied work involving demand for differentiated products following berry1995automobile. Progress has been made on identification of these models in discrete choice, but gaps remain, even in semiparametric settings. For example, nonparametric identification of the distribution of random coefficients in the random coefficients nested logit model has not been established without unbounded regressors.\footnote{See nevo2000practitioner, p. 524-526. il2014identification has established identification in the special case where intercept location coefficients are $0$.} More broadly, there is growing interest in models that allow complementarity, but even less is known about identification of random coefficients in these models.\footnote{Recent work includes gentzkow2007, mcfadden2012theory, fosgerau2019inverse, allen2019identification, ershov2018mergers, monardo2019flexible, iaria2019, and wang2020. This work is an outgrowth of the discrete choice additive random utility model mcfadden1981 and differs from classic continuous demand systems (e.g. deaton1980almost) by focusing on characteristic variation rather than variation in a budget constraint.}
The main contribution of this paper establishes nonparametric identification for the moments of random coefficients in a general class of latent utility models. The framework applies to discrete and continuous choices. As a special case, we establish identification for a bundles model with limited consideration of either alternatives or characteristics (Example (ref)). Identification only depends on the average structural function blundell2003endogeneity. Thus the results can be applied when one observes the average demands of individuals without observing whether goods are chosen together. Leveraging the main result, we can identify the distribution of random coefficients when it is characterized by its moments (e.g. normal distributions). Specialized to discrete choice, the main contribution is new since it does not require any regressor to be unbounded.
Two key ingredients let us get traction for identification. First, we assume independence between random slopes and random intercepts. This is a standard assumption in the widely-used random coefficients logit model. We use this assumption to integrate out the random intercepts. This smooths out demand when conditioning on non-intercept components of the utility function (regressors and random coefficients). This also allows us to treat discrete and continuous choice models in a common framework. Second, we exploit theoretical restrictions since choices arise from optimization. Without using this structure, the model would resemble general random coefficients index models studied in fox2012random and lewbel2017unobserved. While they assume a function governing the mapping from indices to choices is known, we do not.\footnote{In discrete choice, assuming this function is known translates to the distribution of random intercepts being known (e.g. logit). lewbel2017unobserved show that one can drop the assumption that this mapping is known in some settings if one imposes an additional additive separability assumption.}
The fundamental shape restriction we exploit is that, after integrating out random intercepts, integrated mean choices are the derivative of a convex function. This follows from an application of the envelope theorem. Similar tools from convex analysis have also been used for identification of hedonic models ekeland2002identifying,ekeland2004identification,heckman2010nonparametric,chernozhukov2019single, matching galichon2015cupid, dynamic discrete choice chiong2016duality, discrete choice panel models shi2018estimating, and perturbed models with additively separable heterogeneity allen2019identification, among others.\footnote{matzkin1994restrictions reviews other identification results using shape restrictions motivated by economic theory. See also work on optimal transport, as in galichon2018optimal.}
By exploiting the envelope theorem, we can treat several models in a common framework. There is little work on identification of optimizing models with linear random coefficients outside of discrete choice. Exceptions include dunker2018nonparametric for discrete games and dunker2017nonparametric and iaria2019 for a random coefficients version of the gentzkow2007 discrete bundles model. We differ by requiring identification of only the average structural function (“mean demands”), without needing to observe the frequency with which goods are chosen together.\footnote{wang2020 also works with the average structural function but does not identify the distribution of random coefficients.} Identification with linear random coefficients has also been established in settings without assuming an optimizing model. See for example the simultaneous equations analysis in masten2017random and references therein.
Identification of linear random coefficients has been extensively studied in discrete choice. Despite this, nonparametric identification has only been established either requiring a regressor with large support or assuming the distribution of random intercepts is either known or parametric. One reason we do not require a large support assumption is that we focus on identification of the distribution of random coefficients without identifying random intercepts. Papers that make use of large support regressors with heterogeneity that is not additively separable include ichimura1998maximum, berry2009nonparametric, briesch2010nonparametric, gautier2013nonparametric, fox2016nonparametric, dunker2017nonparametric, and fox2017note.\footnote{Exceptions include kashaev2018identification and matzkin2019constructive, but neither paper studies nonparametric identification for the distribution of linear random coefficients.} Several of these papers additionally assume large support regressors also have the same coefficient across goods. In contrast, fox2012random and il2014identification do not assume large support or a homogeneous regressor, but assume the distribution of the random intercept is known (e.g. logit). chernozhukov2019nonseparable discusses identification of ratios of certain moments of the distribution of random coefficients without requiring large support, but do not provide conditions under which the full distribution is identified.
The remainder of the paper proceeds as follows. Section (ref) provides details on the class of latent utility models we study and examples of behavior that this covers. Section (ref) provides the main result, which identifies arbitrary order moments of random coefficients and shows that a single independence and scale assumption can be used to identify all other moments. Section (ref) discusses how to recover different welfare objects and perform counterfactuals. Finally, Section (ref) discusses relations to some existing papers, shows how the results can be taken to settings with non-linear random coefficients, and discusses some testable properties of the framework.
This paper studies the random coefficients perturbed utility model, in which optimizing choices satisfy
mcfadden2012theory and allen2019identification have studied related frameworks without random coefficients. We interpret $Y(\cdot)$ as the quantity vector for $K$ different goods. The vector $X_k = (X_{k,1}, \ldots, X_{k,d_k})'$ denotes observable shifters of the desirability of good $k$, and $\beta_k = (\beta_{k,1}, \ldots, \beta_{k,d_k})'$ denotes random coefficients on these shifters, which may be good-specific. The index $\beta'_k X_k$ shifts the marginal utility of good $k$. We collect $X = (X'_1, \ldots, X'_K)'$ and $\beta = (\beta'_1, \ldots, \beta'_K)'$. The term $D(y,\varepsilon)$ is a disturbance that depends on unobservables $\varepsilon$ of unrestricted dimension. When $D(y,\varepsilon)=\sum_{k=1}^K \varepsilon_k y_k$, $\varepsilon_k$ can be interpreted as a random intercept for the desirability of the $k$-th good. In general, we refer to $D(y,\varepsilon)$ as the random intercept. The set $B \subseteq \mathbb{R}^K$ is a feasibility set. This is introduced purely for exposition, since $D(y,\varepsilon)$ can be $-\infty$ which allows random feasibility sets.
The focus of this paper is on identification of moments of the distribution of $\beta$. Our results do not require specification of the budget $B$, the disturbance $D$, or the distribution of over $\varepsilon$. For concreteness, we provide some examples.
This paper establishes identification of moments of $\beta$ using the average structural function blundell2003endogeneity \[ \overline{Y}(x) = \int Y(x, \beta, \varepsilon) d \tau(\beta,\varepsilon) \] for some probability measure $\tau$ that does not depend on covariates $x$. We assume that the measure $\tau$ satisfies a key independence condition.
While independence between $\beta$ and $\varepsilon$ is restrictive, it is a standard assumption in applications of the random coefficients logit model in discrete choice. It has been exploited for identification in fox2012random and chernozhukov2019nonseparable.\footnote{However, independence is not imposed in some papers studying identification. For example, ichimura1998maximum or gautier2013nonparametric do not impose independence of the slope and intercept.}
With this assumption, we can write \[ \overline{Y}(x) = \int \int Y(x, \beta, \varepsilon) d\mu(\varepsilon) d\nu (\beta) \] for some probability measures $\mu$ and $\nu$. Technically, full independence is not needed as long as we can factor the average structural function in this way.
For an example of an average structural function, suppose $(Y,X,\beta,\varepsilon)$ are random variables that satisfy $Y = Y(X, \beta, \varepsilon)$ almost surely. Moreover, assume $X$, $\beta$, and $\varepsilon$ are all independent. In addition to independence, suppose a continuous version of the conditional mean of $Y$ given $X$ exists. Then \[ \mathbb{E}[ Y \mid X = x] = \overline{Y}(x) = \int \int Y(x, \beta, \varepsilon) d\mu(\varepsilon) d\nu (\beta) \] for $x$ in the support of $X$,\footnote{Recall that the support of $X$ is the smallest closed set $S$ such that $P(X \in S) = 1$.} where $\mu$ is the marginal distribution of $\varepsilon$ and $\nu$ is the marginal distribution of $\beta$.
The results in this paper apply to general average structural functions $\overline{Y}(x)$, not only the conditional mean. Thus, while slope-intercept independence is important for our results, independence between $X$ and $(\beta,\varepsilon)$ is not. Therefore, the results in this paper are relevant for settings with endogeneity.
The goal of this paper is not to provide a new method to identify the average structural function, but rather to use the function to identify other features of a utility maximizing model. There is a large literature on identifying structural functions. blundell2003endogeneity describe how to use control functions to identify the average structural function $\overline{Y}(x)$. altonji2005cross identify derivatives of the average structural function using certain conditional independence or symmetry conditions. berry1994estimating, berry1995automobile, newey2003instrumental, berry2014identification, and dunker2017nonparametric among others use instrumental variables to identify an average structural function from aggregate data.\footnote{A key step to apply these methods is injectivity in a market-level observable to a vector of unobservable endogenous vectors, usually denoted $\xi$. See allen2019injectivity or Lemma 3 in allen2019identification for injectivity results that cover the present model when the utility index for good $k$ is $\beta'_k x_k + \xi_k$. Related injectivity results have appeared in galichon2015cupid and chiong2017counterfactual.}
An important feature of the analysis is that only the average structural function is required to be identified over an appropriate region. Thus, the full distribution of $Y(x,\cdot,\cdot)$ induced by the product measure $\mu \times \nu$ over $(\beta, \varepsilon)$ is not necessary for identification. For common discrete choice models the average structural function and the full distribution of $Y(x, \cdot,\cdot)$ contain the same information, but this is not true in general. This is particularly important when combining this analysis with work allowing endogeneity between $X$ and $(\beta, \varepsilon)$. In particular, there are well-understood methods to identify the average structural function in the presence of endogeneity as mentioned earlier. In contrast, less is known about identification of the entire distribution of $Y(x,\cdot,\cdot)$ in the presence of endogeneity.\footnote{imbens2009identification identify average and quantile structural functions with multidimensional heterogeneity in the outcome equation. torgovitsky2015identification and d2015identification identify the entire structural function with one-dimensional unobservable heterogeneity in the outcome equation. A multidimensional counterpart has been studied in fguns2019. These papers all identify features of structural functions in the presence of endogeneity.}
In addition, requiring only the average structural function implies that the analysis can be applied to settings outside of discrete choice without observing whether goods are chosen together. Of course if the full distribution of $Y(x,\cdot,\cdot)$ is identified, then these results apply as well. We recall that this paper does not study identification of the distribution of $\varepsilon$ in the original latent utility model ((ref)). However, it is possible to identify the distribution of $\varepsilon$ in some cases. For example, dunker2017nonparametric show how to identify the distribution of random intercepts in a full-consideration random coefficients bundles model, provided the analyst has aggregate data on the frequency with which goods are chosen together.
We make use of an aggregation result that first integrates out the distribution of $\varepsilon$.
Note that assuming $Y(\cdot)$ satisfies ((ref)) requires the argmax set to be nonempty. This is a behavioral restriction that imposes sufficient structure for the theorem to go through, and imposes minimal restrictions on $D$. In particular, $D$ can be $-\infty$ for certain combinations of $(y,\varepsilon)$ and need not be continuous. This allows us to treat limited consideration models as in Example (ref).
We leverage the aggregation result from Lemma (ref) to use calculus-based techniques for identification. To illustrate how aggregation can lead to smoothness, recall that $Y(x,\beta,\varepsilon)$ in discrete choice is a vector of indicators denoting which good is chosen (assuming no ties). Derivatives with respect to $x$ either do not exist at certain points, or are zero and contain little information.
We smooth choices by working with \[ \overline{Y}(x, \beta) := \int Y(x, \beta, \varepsilon) d \mu (\varepsilon) \] with $\mu$ as in Assumption (ref). In discrete choice, when $\varepsilon$ is integrated out $\overline{Y}(x,\beta)$ can be interpreted as the vector of probabilities conditional on only the utility indices. However, this general framework allows us to use the same tools to address discrete and continuous choice. For example, choices could involve a single discrete choice, discrete bundle choice, a prospective matching, continuous quantities of several goods, or time use among other settings.
We places some additional high-level sufficient conditions relative to the conclusions of Lemma (ref).
allen2019identification provide lower-level conditions that, when combined with Lemma (ref), imply this assumption. Part (i) strengthens the conclusion of Lemma (ref) to obtain a unique maximizer. Concavity in part (iii) is milder than it first appears, and delivers no additional restrictions on $\overline{Y}(x,\beta)$ when the other assumptions are maintained. See the discussion in allen2019identification.
To further present the foundation of the identification results, we present a version of the envelope theorem.
Here, $\overline{Y}_k(x,\beta)$ is the $k$-th component of $\overline{Y}(x,\beta)$ and $\partial_k V(\beta'_1 x_1, \ldots, \beta'_K x_K)$ is the derivative with respect to the $k$-th dimension of $V$ evaluated at the point $(\beta'_1 x_1, \ldots, \beta'_K x_K)'$. We use similar notation for the rest of the paper. Differentiability of $V$ is implied by the fact that $\overline{Y}$ is the unique maximizer. This is the primary implication of Assumption (ref) that we use for this paper.
fox2012random use a structure similar to ((ref)), showing that when $V$ is known, it is possible to identify moments of the distribution of $\beta$. We differ because we do not require an analyst to specify $V$. Instead, we require certain moments to be nonzero as a relevance condition. Appendix (ref) provides further details and a comparison with their approach. A related structure is considered in lewbel2017unobserved, who identify the distribution of random coefficients when an analogue of $\partial_k V$ is known in advance or additively separable in arguments. We do not impose this structure.
We also leverage a symmetry property of mixed partial derivatives that results from the optimizing behavior in Assumption (ref). For a vector of indices $\gamma = (\gamma_1, \ldots, \gamma_M) \in \{1, \ldots, K \}^M$ and a sufficiently differentiable function $f : \mathbb{R}^K \rightarrow \mathbb{R}$, let \[ \partial_{\gamma} f := \partial_{\gamma_1} \cdots \partial_{\gamma_M} f. \]
This result states that the order in which we take partial derivatives does not matter. For example, when $M = 2$ we have the usual symmetry property of mixed partial derivatives with respect to dimensions $j,k \in \{1,\ldots,K\}$ that \[ \partial_{j,k} V(\vec{u}) = \partial_{k,j} V(\vec{u}). \] The lemma follows by repeated application of the $M = 2$ case.
With the foundations in place, we now turn to the task of identifying moments of random coefficients. We focus on conditions where certain $M$-th order moments of the distribution of $\beta$ are identified. In particular, if the assumptions hold for all $M$, then all moments of the distribution of random coefficients are identified.
We assume regressors are continuous and satisfy an exclusion restriction.
We now provide some intuition for the main result (Theorem (ref)). We consider identifying second moments of $\beta$ ($M = 2$) when there are two goods ($K = 2$) and each good has a single covariate ($d_k = 1$). We focus on second moments since this example captures the power of the results in the simplest non-trivial setting. We write the partial derivative of a function, $f$, with respect to the covariates of the $j$-th good, $x_j$, as $\partial_{x_j}f$. Differentiating the envelope theorem (Lemma (ref)) and evaluating at $x = 0$ we obtain \[ \partial_{x_j} \overline{Y}_k(0, \beta) = \partial_{j,k} V(0) \beta_j.\footnote{Here we abuse notation and for the function $f$, we let $\partial_{s} f(0) = \left. \partial_{s} f(z) \right|_{z=0}$.} \] This uses the fact that $x_j$ is continuous and excluded from the utility index of other goods. This can be repeated with other mixed partial derivatives. Importantly, by evaluating derivatives at the point $x = 0$, the terms $\partial_{j,k}V(0)$ do not depend on $\beta$. Thus, when integrating over the values of the random coefficients, the term involving $V$ passes outside of the integral. In particular, integrating over $\beta$ yields the following system of equations
where we have implicitly assumed that differentiation and integration can be interchanged.
Assume that the derivatives of $\overline{Y}$ are identified. At first glance, this is a system of four equations with seven unknowns (clearly the $\beta_1 \beta_2$ and $\beta_2 \beta_1$ moments are equal). However, when $V$ is sufficiently differentiable, partial derivatives of $V$ do not depend on the order of differentiation (Lemma (ref)), which eliminates two unknowns. Using a scale assumption that $\int \beta^2_1 d\nu (\beta)$ is known a priori will eliminate an unknown and gives a system with $4$ equations and $4$ unknowns. We show that this is enough to identify all second moments of $\beta$.
To constructively see how the moments are identified, note that using symmetry of derivatives, the first and third equations identify $\int \beta_1 \beta_2 d\nu ( \beta)$. Using this, we identify $\partial_{1,2,2} V(0)$ using the second equation. Again using symmetry of derivatives and combining this with the last equation identifies $\int \beta^2_2 d\nu (\beta)$. Once all moments are identified, the remaining third order derivatives of $V$ can be identified at $0$.\footnote{This part also requires the equations
to identify $\partial_{1,1,1} V(0)$ and $\partial_{2,2,2} V(0)$.}
We now provide formal conditions that justify the intuitive argument for any number of goods, covariates, and order of moment $M$.
Assumption (ref) holds if we set $\beta_{1,1} = 1$, for example, but is considerably more general. It allows heterogeneity in the sign of $\beta_{1,1}$, for example. In general, if one wants to identify all moments of $\beta$ using the main result, then for every $M$ Assumption (ref) must hold. This assumption holds when the distribution of $\beta_{1,1}$ is known a priori and the distribution has nonzero moments of all orders. If Assumption (ref) is dropped, the results in this paper establish identification of the ratio of any nonzero $M$-th order moments. Thus, Assumption (ref) can be appropriately modified by instead holding fixed the value of some other nonzero $M$-th order moment of the form $\int \beta_{k_1, \ell_1} \cdots \beta_{k_M, \ell_M} d \nu (\beta)$. We show in Section (ref) that if $\beta_{1,1}$ is independent of all other components of $\beta$, then identification is possible using a single scale assumption on the first moment.
Recall that with minor abuse of notation we set \[ \overline{Y}(x) = \int \overline{Y}(x, \beta) d\nu (\beta). \] We require the following regularity conditions.
These regularity conditions parallel assumptions in fox2012random. To interpret part (i), note that $\nu$ can be a discrete probability measure over $\beta$ with finite support. For discrete measures, (i) holds whenever $\overline{Y}_k(x,\beta)$ is $M$-times differentiable in $x$ for every $\beta$ in its support. Part (ii) formalizes that the moments we wish to identify exist and are finite.
Parts (iii) and (iv) can be linked to derivatives of the function $\overline{Y}(x)$ via the envelope theorem (Lemma (ref)). Indeed, differentiating the envelope theorem for the $k$-th good, evaluating the derivative of $\bar{Y}_k$ with respect to $x_{j,\ell}$ at $x = 0$, and taking expectations yields
Thus, when $V$ is $(M+1)$-times continuously differentiable, it follows that $\overline{Y}$ is $M$-times continuously differentiable. Moreover, if one sees empirically that $\frac{\partial {\overline{Y}_k(0)}}{\partial x_{j,\ell}} \neq 0$, then it follows that that $\partial_{j,k} V(0) \neq 0$ (whenever this derivative exists).
fox2012random show that condition (iv) holds for random coefficients logit, for “most” values of nonrandom intercepts. Specifically, the set of intercepts that violate (iv) for some $\gamma$ has Lebesgue measure $0$. In general, whether (iv) holds depends on features of the distribution of $\varepsilon$ and choice of $D$, which are example specific. For example, part (iv) rules out pure characteristic discrete choice models as in berry2007pure and dunker2017nonparametric. These models do not include a random intercept, and so the value function for the pure characteristics model, $V^{PC}$, can be written as \[ V^{PC}(\beta'_1 x_1, \ldots, \beta'_K x_K) = \sup_{y \in \overline{B}} \sum_{k = 1}^K y_k (\beta_k'x_k) \] without the additive disturbance $\overline{D}$, where $\overline{B}$ is the probability simplex. $V^{PC}$ does not have a non-zero derivative at $x=0$ for any $M \ge 1$ with this constraint set. This choice of $V^{PC}$ also does not always induce a unique maximizer.
More generally, condition (iv) requires that goods in the demand system are related. For example, if the original $K$-good demand system can be written as $K$ separate $1$-good demand systems, then derivatives of the form $V_{j,k}(0)$ will be zero for $j \neq k$. This is because under this separability assumption, the utility index of good $j$ does not alter the demand for the $k$-th good. In general, (iv) cannot be relaxed for the main result to hold without additional assumptions. For example, if we impose that $\beta_j = \beta_k$ (a.s.) for all $j,k \in \{1,\ldots,K\}$, then one can identify ratios of moments under the weaker assumption that $\partial^{M+1}_j V(0) \neq 0$ for some $j$. See Appendix (ref).
Condition (v) states that $\overline{Y}(x)$ is identified over a small region near $x = 0$. The constructive identification results in fox2012random and chernozhukov2019nonseparable have also made use of variation around zero. In contrast, most of the literature instead requires identification of $\overline{Y}$ either for all $x$ or for a set over which $x$ is unbounded along some dimensions.
To interpret condition (v), suppose that $X$, $\beta$, and $\varepsilon$ are all independent, and we identify $\overline{Y}$ from a continuous version of the conditional mean of $Y$ given $X$. For this case, condition (v) is implied when the support of $X$ contains an open ball around $x = 0$. The second more general part of (v) highlights that the results also apply when the average structural function is identified over a weakly positive region. Thus, our results do not rule out prices. We can handle this case because we only need to identify certain derivatives of $\overline{Y}$ at $0$. These derivatives of $\overline{Y}$ at $0$ are identified in this case by calculating derivatives from “one-sided” limits involving non-negative numbers.
The final assumption used for identification is that a sufficiently rich set of $M$-th order moments of $\beta$ are nonzero.
This is a relevance condition. It is not necessary to know which indices $(\ell_1, \ldots, \ell_M)$ satisfy this condition in advance.\footnote{See the discussion after Lemma (ref) in Appendix (ref).} A sufficient condition for this is that for every $k$-th good there is a regressor $\ell_k \in \{1,\ldots,d_k\}$ such that either $\beta_{k,\ell_k} \geq 0$ almost surely or $\beta_{k,\ell_k} \leq 0$ almost surely, with positive probability that the inequality is strict. A stronger condition that implies this is that $\beta_{k,1} = 1$ (a.s.) for every $k$-th good by setting $\ell_1 = \cdots = \ell_M = 1$. This is a common assumption in the literature berry2009nonparametric,briesch2010nonparametric,dunker2017nonparametric. However, ichimura1998maximum and gautier2013nonparametric establish identification of random coefficients models for binary discrete choice using a more general halfspace condition.
With these assumptions, we can now state the main result of the paper.
This result establishes nonparametric identification of certain moments of $\beta$. It can be directly used to establish semiparametric identification of the distribution of $\beta$ for certain parametric families without specifying other objects (e.g. $V$). For example, if $\beta$ is normally distributed then Theorem (ref) identifies the distribution when the assumptions hold for $M \in \{1, 2\}$ because normal distributions are characterized by means and covariances. Recall that while we identify non-centered moments, we can use this information to identify centered moments. More generally, for any distribution of $\beta$ that is defined by its moments up to order $M$ this result estabilishes identification of the distribution.
fox2012random and il2014identification describe a sufficient condition for a distribution to be determined by its moments. Distributions with compact finite support are determined by their moments. Lognormal distributions are an example of distributions that are not determined by integer moments heyde1963property. That is, there are other nonparametric distributions that can match the same moments. However, the parameters may still be identified within the lognormal class.
Identifying the distribution of $\beta$ using Corollary (ref) requires Assumption (ref), which specifies all moments of the form $\int \beta^M_{1,1} d\nu (\beta)$. When the distribution of $\beta$ is identified from its moments, one must specify the marginal distribution of $\beta_{1,1}$ in advance to apply Corollary (ref). While the common assumption $\beta_{1,1} = 1$ (a.s.) implies Assumption (ref), one may not want to impose either this assumption or the weaker assumption that $\beta_{1,1}$ has a known distribution. This section describes an alternative assumption that ensures identification of moments of $\beta$. In particular, we assume $\beta_{1,1}$ is independent of other components of $\beta$.
With this independence assumption, we show that a single scale assumption on the first moment of $\beta_{1,1}$ allows us to identify a rich collection of moments. This contrasts with Theorem (ref), which uses an assumption on the $M$-th order moment of $\beta_{1,1}$ to identify only $M$-th order moments of $\beta$.
Alternatively, one could set the absolute value of some other order moment of $\beta_{1,1}$, but we focus on the first moment since it facilitates interpretation. Independence between $\beta_{1,1}$ and other components is considerably weaker than assuming $\beta_{1,1} = 1$ almost surely. For example, this allows $\beta_{1,1}$ to be sometimes negative and sometimes positive. Thus, different individuals can be repelled or attracted to higher values of $x_{1,1}$.
Replacing Assumption (ref) with Assumption (ref), we obtain the following counterpart of Theorem (ref).
Relative to Theorem (ref), independence of $\beta_{1,1}$ from the other components allows us to relate the $M$-th and $M-1$ order moments. To see this, consider some $M$-th order moment in which $\beta_{1,1}$ appears exactly once. Using independence, we obtain \[ \int \beta_{1,1} \beta_{k_2, \ell_2} \cdots \beta_{k_M, \ell_M} d\nu (\beta) = \int \beta_{1,1} d\nu (\beta) \int \beta_{k_2, \ell_2} \cdots \beta_{k_M, \ell_M} d\nu (\beta). \] In the proof, we show that $\int \beta_{1,1} d\nu (\beta)$ can be identified when $| \int \beta_{1,1} d\nu (\beta)|$ is finite, known a priori, and non-zero. With this knowledge, we can identify the ratio of all $M$-th order moments to all $(M-1)$-th order moments and apply induction to identify all $M \le \bar{M}$ order moments.
We now turn to identification of certain welfare and counterfactual objects. Identification is established given identification of certain features of $V$, which is the indirect utility function obtained when random intercepts are integrated out. We first provide three results that identify differences in $V$. Using these results, we discuss welfare analysis and counterfactuals.
The reason we identify $V$ is that we can use the envelope theorem to determine certain average choices \[ \overline{Y}(x,\beta) = \nabla V(\beta'_1 x_1, \ldots, \beta'_K x_K). \] We require identification of the right hand side at values other than $0$ to consider counterfactuals at new values of covariates.
We first provide conditions under which identification of partial derivatives of $V$ at $0$ allows us to directly extrapolate the function. Specifically, we assume $V$ is a real analytic function. That is, $V$ has derivatives of all orders and agrees with its Taylor series in a neighborhood of every point. Real analytic functions have the important property that local information can be used to reconstruct the function globally by extrapolating. This is similar to common parametric classes of functions. However, the set of real analytic functions is infinite dimensional.
One way to drop the assumption that $V$ is a real analytic function is to instead assume $\beta_{k,1} = 1$ almost surely for each $k$. With this assumption, let $\tilde{x}$ be a value that is zero for every characteristic except the first characteristic of each good. Then the envelope theorem (Lemma (ref)) specializes to \[ \overline{Y}_k(\tilde{x},\beta) = \partial_k V(\tilde{x}_{1,1}, \ldots, \tilde{x}_{K,1}). \] This does not depend on $\beta$, and so by taking expectations, the average structural function identifies the derivative of $V$ at the point $(\tilde{x}_{1,1}, \ldots, \tilde{x}_{K,1})$. By integrating the derivatives we can identify differences in $V$, as we now formalize.
The results on welfare and counterfactual analysis require that derivatives of $V$ be identified at certain values $(\beta'_1 x_1, \cdots, \beta'_K x_K)$. If the support of $\beta$ is compact, then it is not necessary to identify $V$ everywhere, and so it is not necessary to have $\underline{x}_{k,1} = -\infty$ and $\overline{x}_{k,1} = \infty$ to apply Proposition (ref).
Finally, we mention a third way to identify differences in $V$. A key distinction is that it requires identification of the distribution of $\overline{Y}(x,\beta)$ for fixed $x$, rather than identification of the average structural function as in the rest of the paper. We adapt the following lemma.
This result can be applied to our setting by adapting the envelope theorem, \[ \overline{Y} (x,\beta) = \nabla V(\beta_1 ' x_1, \ldots, \beta'_K x_K). \] Interpret $W = \overline{Y}(x,\beta)$, $f = \nabla V$, and $\eta = (\beta_1' x_1, \ldots, \beta_K' x_K)$. When $x$ is fixed and the distribution of $\beta$ is identified (from previous arguments), the distribution of $\eta$ is known. The function $V$ is convex, and so the lemma provides conditions under which $\nabla V$ is identified. Importantly, the lemma can be applied at a single $x$, so it is not necessary to have full support of covariates to apply the result. Such an $x$ cannot be arbitrary. For example, when $x = 0$ the distribution corresponding to $\eta$ is not absolutely continuous. The lemma can still be applied for $x$ near but not equal to $0$. Moreover, to apply the lemma, the distribution of $\beta$ cannot be degenerate, i.e. there must be truly “random” coefficients. If $\beta$ is almost surely equal to a constant, then the distribution corresponding to $\eta$ is not absolutely continuous and the lemma does not apply.
Importantly, to apply Lemma (ref) in our setting, the distribution of $\overline{Y}(x,\beta)$ must be identified at some fixed $x$. One example in which this lemma can be applied is when $\varepsilon$ in the original latent utility model is not present, so that $\overline{Y}(x,\beta)$ corresponds to the observable choices given $x$ and $\beta$. Such structure could be appropriate in a continuous choice model in which all unobservable heterogeneity is controlled by the random slopes $\beta$, and in which the choices (rather than e.g. average choices for a group of individuals) are observed.
We now describe how identification of $V$ leads to identification of certain welfare objects. First, recall that $V$ may be interpreted as the indirect utility conditional on the utility index (Lemma (ref)) where the random intercept $\varepsilon$ under the measure $\mu$ is integrated out. To interpret $V(\cdot)$ as a welfare object, suppose $\beta$ is an individual-specific term that is random across the population but constant across decisions for the same individual. Interpret the random intercept $\varepsilon$ as an idiosyncratic taste shock across decision problems. Then $V(\beta'_1 x_1, \ldots, \beta'_K x_K)$ is an individual-specific (integrated) indirect utility. The conditions of Corollary (ref) identify the distribution of $V(\beta'_1 x_1, \ldots, \beta'_K x_K)$ under the measure $\nu$ up to an additive constant once the values of covariates are fixed.
Thus, we can identify the distribution of individual-specific indirect utilities. By further integrating out the distribution of $\beta$, we also identify differences in the average indirect utility via Lemma (ref):
This result holds regardless of whether the distribution of $\varepsilon$ is identified since $V$ is a welfare-relevant summary measure of the distribution of $\varepsilon$. Indeed, we do not establish identification of the distribution of \[ \sup_{y \in B} \sum_{k = 1}^K y_k (\beta'_k x_k) + D(y, \varepsilon) \] according to the product measure $\mu \times \nu$ over $(\beta,\varepsilon)$. In particular, since this paper does not study identification of $\varepsilon$, we do not identify the distribution of indirect utilities including random intercepts.
Note that the units of ((ref)) are relative to the scale assumption used to identify the distribution of $\beta$. If we impose the scale assumptions in Theorem (ref) to apply Corollary (ref), then the distribution of $\beta_{1,1}$ is fixed. Thus, the units of Equation (ref) are set by the distribution of the conversion rate between $x_{1,1}$ and utils. In contrast, if we impose the conditions of Proposition (ref), then the scale is determined by $\left| \int \beta_{1,1} \nu (\beta) \right|$. For this case, the units of Equation (ref) are relative to the average conversion rate between $x_{1,1}$ and utils.
An alternative measure of average indirect utility is
This sets the conversion rate of $x_{1,1}$ and utils to $\pm 1$. Importantly, this preserves whether the first characteristic is desirable or undesirable. It also forces the intensity of preference to be constant across individuals. This welfare measure is most interpretable when the regressor has a homogeneous sign. For example, if $x_{1,1}$ is the (negative) price of good $1$, then $\beta_{1,1} < 0$ is a natural assumption, and the units of Equation (ref) are in dollars.
Once $V$ is identified, we can also answer certain counterfactual questions involving quantities at new values of covariates. To this end, recall from Lemma (ref) that
Here, $\overline{Y}_k(x, \beta)$ is the demand for good $k$ fixing covariates and the random intercept, but integrating out the distribution of $\varepsilon$. We interpret $\beta$ as an individual-specific parameter that is constant across decision problems, while $\varepsilon$ is an idiosyncratic shock that can vary across decision problems. Thus, $\overline{Y}_k(x, \beta)$ is the individual-specific average quantity of the $k$-th good. Once $V$ and the distribution of $\beta$ are identified, we can identify the distribution of $\overline{Y}_k(x,\cdot)$ from Equation (ref).
Conceptually, this shows it is possible to start with identification of the average structural function (“mean choices”) around $x = 0$ to identify the integrated choices $\overline{Y}_k(x,\beta)$ at all values of the covariates. This also implies that any value $x$ at which $\overline{Y}$ can be identified directly from data (as opposed to the theoretical analysis just described) provides overidentifying information.
We now provide additional discussion of the main results in the paper.\\