EconBase
← Back to paper

Nonparametric Identification of Random Coefficients in Endogenous and Heterogeneous Aggregate Demand Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

89,116 characters

Nonparametric Identification of Random Coefficients in Endogenous and Heterogeneous Aggregate Demand Models



\title{Nonparametric Identification of Random Coefficients in Endogenous and
Heterogeneous Aggregate Demand Models}
 \author{
 \begin{tabular}{ccccc}
 Fabian Dunker\thanks{
 School of Mathematics and Statistics, University of Canterbury, Private
 Bag 4800, Christchurch 8140, New Zealand, Email: [email removed]} &  & Stefan Hoderlein\thanks{
 Department of Economics, Emory University, 1602 Fishburn Dr, Atlanta, GA 30322, USA, Email: [email removed]. } &  & Hiroaki Kaido
 \thanks{
 Department of Economics, Boston University, 270 Bay State Road, Boston, MA
 02215, USA, Email: [email removed]. We thank Steve Berry, Jeremy Fox, Amit
 Gandhi, Phil Haile, Kei Hirano, Arthur Lewbel, Marc Rysman, and Elie Tamer
 for their helpful comments. We also thank seminar and conference
 participants at Harvard, BU-BC joint mini-conference, the Demand Estimation
 and Modeling Conference 2013 and the Cowles Summer Conference 2014. Kaido gratefully acknowledges financial
 support from NSF Grant SES-1357643.} \\
 {\tmpsmall\sc University of Canterbury } &  & {\tmpsmall\sc Emory University } &  &
 {\tmpsmall\sc Boston University} \\
 &  &  &  &  \\
 &  &  &  &
 \end{tabular}
 }
 \date{\today }
\maketitle

\begin{abstract}
This paper studies nonparametric identification in market level demand
models for differentiated products with heterogeneous consumers. We consider
a general class of models that allows for the individual specific
coefficients to vary continuously across the population and give conditions
under which the density of these coefficients, and hence also functionals
such as welfare measures, is identified.
A key finding is that two leading models, the BLP-model
(Berry, Levinsohn, and Pakes, 1995) and the pure characteristics model
(Berry and Pakes, 2007), require considerably different conditions on the
support of the product characteristics.


\bf{Keywords:} Random Coefficients, Aggregate Demand, Nonparametric Identification
\end{abstract}


\section{Introduction}

Modeling consumer demand for products that are bought in single or discrete
units has a long and colorful history in applied economics, dating back to
at least the foundational work of McFadden (1974, 1981). While allowing for
heterogeneity, much of the earlier work on this topic, however, was not able
to deal with the fact that in particular the own price is endogenous. In a
seminal paper that provides the foundation for much of contemporaneous work
on discrete choice consumer demand, Berry, Levinsohn and Pakes (1994, BLP)
have proposed a solution to the endogeneity problem. Indeed, this work is so
appealing that it is not just applied in discrete choice demand and
empirical IO, but also increasingly in many adjacent fields, such as health,
urban or education economics, and many others. From a methodological
perspective, this line of work is quite different from traditional
multivariate choice, as it uses data on the aggregate level and integrates
out individual characteristics\footnote{
There are extensions of the BLP framework that allow for the use of
Microdata, see Berry, Levinsohn and Pakes (2004, MicroBLP). In this paper,
we focus on the aggregate demand version of BLP, and leave an analogous work
to MicroBLP\ for future research.} to obtain a system of nonseparable
equations. This system is then inverted for unobservables for which in turn
a moment condition is then supposed to hold.

Descending in parts from the parametric work of McFadden (1974, 1981),
market-level demand models share many of its features, in particular
(parametric) distributional assumptions, but also a linear random
coefficients (RCs) structure for the latent utility. Not surprisingly, there
is increasing interest in the properties of the model, in particular which
features of the model are nonparametrically point identified, and how the
structural assumptions affect identification of the parameters of interest.
Why is the answer to these questions important? Because an empiricist
working with this model wants to understand whether the results she obtained
are a consequence of the specific parametric assumptions she invoked, or
whether they are at least qualitatively robust. In addition, nonparametric
identification provides some guidance on essential model structure and on
data requirements, in particular about instruments. Finally, understanding
the basic structure of the model makes it easier to understand how the model
can be extended. Extensions of the BLP framework that are desirable are in
particular to allow for consumption of bundles and multiple units of a product  without modeling every choice as a new separate alternative.

We are not the first to ask the nonparametric identification question for
market demand models. In a series of elegant papers, Berry and Haile (2014, BH henceforth), Berry and Haile (2020) provide important answers to many of the identification
questions. In particular, they establish conditions under which the
\textquotedblleft Berry inversion\textquotedblright , a core building block
of the BLP\ model named after Berry (1994), which allows to solve for
unobserved product characteristics, as well as the distribution of a
heterogeneous utility index are nonparametrically identified.

Our work complements this line of work in that we follow more closely the
original BLP specification and assume in addition that the utility index has
a linear random coefficients (RCs) structure. More specifically, we show how
to nonparametrically identify the distribution of random coefficients in
this framework. This result does not just close the remaining gap in the
proof of nonparametric identification of the original BLP model, but is also
important for applications because the distribution of random coefficients
allows to characterize the distribution of the changes in welfare due to a
change in observable characteristics, in particular the own price (to borrow an analogy from
the treatment effect literature, if we think of a price as a treatment, BH
recover the treatment effect on the distribution, while we recover the
distribution of treatment effects). For example, consider a change in the
characteristics of a good. The change may be due to a new regulation, an
improvement of the quality of a product, or an introduction of a new
product. Knowledge of the random coefficient density allows the researcher
to calculate the distribution of the welfare effects. This allows one to
answer various questions. For example, one may investigate whether the
change gives rise to a Pareto improvement. This is possible because, with
the distribution of the random coefficients being identified, one can track
each individual's welfare before and after the change. If a change in one of
the product characteristics is not Pareto improving, one can also calculate
the proportion of individuals who would benefit from the change and
therefore prefers the product with new characteristics.\footnote{
Note that simultaneous changes in product characteristics and price are
allowed. Hence, one can investigate how much price change is required
to compensate for a change (e.g. downgrading of a feature) in one of the
product characteristics to let a certain fraction of individuals receive a
non-negative utility change, i.e. $P(\Delta U_{ijt}\geq 0)\geq \tau $ for
some prespecified $\tau \in \lbrack 0,1]$, where $\Delta U_{ijt}$ denotes
the utility change.} Identification of the random coefficient distribution
allows one to conduct various types of welfare analysis that are not
possible by only identifying the demand function. Our focus therefore will
be on the set of conditions under which one can uniquely identify the random
coefficient distribution from the observed demand.

Naturally, identification will depend crucially on the specific model at
hand. As it turns out, there are important differences between the classical
BLP and the pure characteristics model (see Berry and Pakes (2007), PCM\
henceforth) that stem from the presence of an alternative, individual and
market specific error, typically assumed to be logistically distributed and
hence called \textquotedblleft logit error\textquotedblright\ in the
following. A lucid discussion about the pros and cons of both approaches can
be found in Berry and Pakes (2007). One advantage of the PCM we would like
to emphasize at this point is that it is well-suited for the analysis of
welfare changes when a new product with a particular characteristic is
introduced to the market. Moreover, the pure characteristics model also
predicts a reasonable substitution pattern when the number of products is
large, while the BLP-type model may give counter-intuitive predictions. In
addition to these important economic differences, the identification
strategies including the required assumptions also differ significantly
across the two models. In particular, in the BLP model, one needs to rely on
an identification at infinity argument to isolate the unobservable for each
product. In remarkable contrast, in the PCM\ one does not require such an
argument (and therefore does not have to employ some restrictive
assumptions). Instead, in the PCM, we demonstrate that one may combine
demand on products across different markets to construct a function that
depends on the random coefficients through a single index so that we can
recover the distribution of unobserved heterogeneity without relying on
identification at infinity. We call this construction \emph{marginalization
(or aggregation) of demand}. This is possible due to the unique structure of
the PCM in which  only the product characteristics (but not the tastes for
products) determine the demand. To our knowledge, this identification
strategy is novel.

The arguments in establishing nonparametric identification of these changes
are constructive and permit the construction of sample counterparts
estimators, using the theory in Hoderlein, Klemel\"{a} and Mammen (2010) or
Dunker Mendoza and Reale (2021). This theory reveals that the random
coefficients density is only weakly identified, suggesting that numerical
instabilities and problems frequently reported and discussed in the BLP
literature, e.g., Dube, Fox and Su (2013), are caused or aggravated by this
feature of the model.

Another contribution in this paper is that we use the insights obtained from
the identification results to extend the market demand\ framework to cover
bundle choice (i.e., consume complementary goods together). Note that bundles can
in principle be accommodated within the BLP framework by treating them as
separate alternatives. However, this is not parsimonious as the number of
alternatives increases rapidly and with it the number of unobserved product
characteristics, making the system quickly intractable. To fix ideas,
suppose there were two goods, say good A and B. First, we allow for the
joint consumption of goods A and B, and second, we allow for the consumption
of several units of either A and/or B, without labeling it a separate
alternative. We model the utility of each bundle as a combination of the
utilities for each good and an extra utility from consuming the bundle. This
structure in turn implies that the dimension of the unobservable product
characteristic equals the number of goods $J$ instead of the number of
bundles. There are three conclusions we draw from this contribution:\ first,
depending on the type of model, the data requirements vary. In particular,
to identify all structural parts of the model, in, say, the model on bundle
choice, market shares are not the correct dependent variable any more.
Second, depending on the object of interest, the data requirements and
assumptions may vary depending on whether we want to just recover demand
elasticities, or the entire distribution of random coefficients. Third, the
parsimonious features of the structural model result in significant
overidentification of the model, which opens up the way for specification
testing, and efficient estimation. As in the classical BLP setup, in all
setups we may use the identification argument to propose a nonparametric
sample counterpart estimators.\footnote{In DHK (2018), we also use the insights obtained to
propose a parametric estimator for models where there had not been an
estimator before.}

\textbf{Related literature: }as discussed above, this paper is closely
related to both the original BLP line of work (Berry, Levinsohn and Pakes
(1994, 2004)), as well as to the recent identification analysis of Berry and
Haile (2014, 2020). Because of its generality, our approach also provides
identification analysis for the \textquotedblleft pure
characteristics\textquotedblright\ model of Berry and Pakes (2007), see also
Ackerberg, Benkard, Berry and Pakes (2007) for an overview. Other important
work in this literature that is completely or partially covered by the
identification results in this paper include Petrin (2002) and Nevo (2001).
Moreover, from a methodological perspective, we note that BLP continues a
line of work that emanates from a broader literature which in turn was
pioneered by McFadden (1974, 1981); some of our identification results
extend therefore beyond the specific market demand model at hand. Other
important recent contributions in discrete choice demand include
Gowrisankaran and Rysman (2012), Armstrong (2016) and Moon, Shum, and
Weidner (2018). Less closely related is the literature on hedonic models,
see Heckman, Matzkin and Nesheim (2010), and references therein.

In addition to this line of work, we also share some commonalities with the
work on bundle choice in IO, most notably Gentzkow (2007), and Fox and
Lazzati (2017). For some of the examples discussed in this paper, we use
Gale-Nikaido inversion results, which are related to arguments in Berry,
Gandhi and Haile (2013). Because of the endogeneity, our approach also
relates to nonparametric IV, in particular to Newey and Powell (2003),
Andrews (2017), and Dunker, Florens, Hohage, Johannes, and Mammen (2014).
Finally, our arguments are related to the literature on random coefficients
in discrete choice model, see Ichimura and Thompson (1995), Gautier and
Kitamura (2013), Fox and Gandhi (2016), Dunker, Hoderlein, Kaido, and Sherman (2018),
and Matzkin (2012). Since we use the Radon transform introduced by
Hoderlein, Klemel\"{a} and Mammen (2010, HKM) into Econometrics,
this work is particularly close to the literature that uses the Radon
transform, in particular HKM\ and Gautier and Hoderlein (2015). Finally, the
class of models we consider is related but differs from the mixed logit
model (without endogeneity) analyzed by Fox, Kim, Ryan, and Bajari (2012)
who established the identification of the distribution of the random
coefficients from micro-level data, while maintaining the logit assumption
on the tastes for products. Our focus here is on market-level models with
endogeneity with the main goal being the identification of the distribution
of all random coefficients without any parametric assumption. As such, our
identification strategy differs significantly from theirs. Finally, after the original version of this paper, there have been recent developments on the nonparametric identification of aggregate demand models. Allen and Rehbeck (2020) study partial identification of latent complementarity in an aggregate demand model of bundles. Their focus is on what can be learned about latent complementarity when the variation of demand shifters is limited. Lu, Shi, and Tao (2019) study identification and semiparametric estimation of random coefficient logit demand models in a related but different environment, in which consumers face a growing number of products.

\textbf{Structure of the paper}: The second section lays out preliminaries
we require for our main result: We first introduce the class of models and
detail the structure of our two main setups. Still in the same section, for
completeness we quickly recapitulate the results of Berry and Haile (2014)
concerning the identification of structural demands, adapted to our setup.
The third section contains the key novel result in this paper, the
nonparametric (point-)identification of the distribution of random
coefficients in the class of discrete choice demand model with
endogeneity, which includes the BLP and PCM models. In the fourth section,
we discuss the identification in the bundles
case, including how the structural demand identification results of Berry
and Haile (2014) have to be adapted, but again focusing on the random
coefficients density. We then end with an outlook.

\section{Preliminaries}

\subsection{Model}

We begin with a setting where a consumer faces $J\in\mathbb{N}$ products and
an outside good which is labeled good 0. Throughout, we index individuals by
$i$, products by $j$ and markets by $t$. We use upper-case letters, e.g. $
X_{jt}$, for random variables (or vectors) that vary across markets and
lower-case letters, e.g. $x_{j}$, for particular values the random variables
(vectors) can take. In addition, we use letters without a subscript for
products e.g. $X_{t}$ to represent vectors e.g. $(X_{1t},\cdots ,X_{Jt})$.
For individual $i$ in market $t$, the (indirect) utility from consuming good
$j$ depends on its (log) price $P_{jt}$, a vector of observable
characteristics $X_{jt}\in \mathbb{R}^{d_X}$, and an unobservable scalar
characteristic $\Xi _{jt}\in \mathbb{R}$. We model the utility from
consuming good $j$ using the linear random coefficient specification:
\begin{equation}
U_{ijt}^{\ast }\equiv X_{jt}^{\prime }\beta _{it}+\alpha _{it}P_{jt}+\Xi
_{jt}+\sigma_\epsilon\epsilon_{ijt},~j=1,\cdots ,J~,  \label{eq:randomu}
\end{equation}
where $(\alpha _{it},\beta _{it})^{\prime }\in \mathbb{R}^{d_X+1}$ is a
vector of random coefficients representing the tastes for the product
characteristics. For each $j$, $\epsilon_{ijt}$ represents the ``taste for
the product'' itself. Following Berry and Pakes (2007), we consider a class
of general market-level demand models that nests models with tastes for products
($\sigma_\epsilon=1$) and without tastes for products ($\sigma_\epsilon=0$).
The models with tastes for products include the random coefficient logit model
used in BLP, in which case $\sigma_\epsilon=1$ and $\epsilon_{ijt},j=1,\cdots,J$
are i.i.d. Type-I
extreme value random variables. When $\sigma_\epsilon=0$, the model is
called the \textit{pure characteristic model} (PCM). The two models are
known to have different theoretical properties. For example, the BLP model
predicts that even with a large number of products, the mark-up remains
positive implying there is always an incentive to develop a new product. As
the number of new products grows, each individual's utility tends to
infinity.  On the other hand, in PCM, the model approaches competitive
equilibrium and the incentive to develop a new product diminishes as the
number of products increases.\footnote{
See Berry and Pakes (2007) for more details.} As we will show below, the two
models also differ in terms of empirical contents.

Throughout, we assume that $X_{jt}$ is exogenous, while $P_{jt}$ can be
correlated with the unobserved product characteristic $\Xi _{jt}$ in an
arbitrary way. Without loss of generality, we normalize the utility from the
outside good to 0. This mirrors the setup considered in BH (2014).

We think of a large sample of individuals as $iid$ copies of this population
model. The random coefficients $\theta_{it}\equiv(\alpha_{it}
,\beta_{it},\epsilon_{i1t},\cdots,\epsilon_{iJt})^{\prime }$ vary across
individuals in any given market (or, alternatively, have a distribution in
any given market in the population), while the product characteristics vary
solely across markets. These coefficients are assumed to follow a
distribution with a density function $f_{\theta }$ with respect to Lebesgue
measure, i.e., be continuously distributed.\footnote{
This assumption is not crucial but made for the ease of exposition. Our main
identification results (Theorems \ref{thm:fbeta} and \ref{thm:fbeta_pure})
hold for any Borel measure.} This density is assumed to
be common across markets, and is therefore not indexed by $t.$ As we will
show, an important aspect of our identification argument is that, once the
demand function is identified, one may recover $\Xi_{t}$ from the market
shares and other product characteristics $(X_{t},P_{t})$. Then, by creating
exogenous variations in the product characteristics and exploiting the
linear random coefficients structure, one may trace out the distribution $
f_\theta$ of the preference that is common across markets. We note that we
can allow for the coefficients $(\alpha_{it} ,\beta_{it})$ to be alternative
$j$ specific, and will do so in the online supplement. However, parts of the analysis
will subsequently change.

Having specified the model on the individual level, the outcomes of individual
decisions are then aggregated in every market. The econometrician observes
exactly these market level outcomes $S_{l,t}$, where $l$ belongs to some
index set denoted by $\mathbb{L}$. Below, we give two examples. The first
example is the setting of the BLP and pure characteristics models, where
individuals choose a single good out of multiple products, while the second
is about the demand for bundles.

\begin{example}[Multinomial choice]
\textrm{\textrm{\textrm{Each individual chooses the product that maximizes
her utility out of $J\in\mathbb{N}$ products. Hence, product $j$ is chosen
if
\begin{align}
U^*_{jt}>U^*_{kt}~,~~\forall k\ne j~.
\end{align}
The demand for good $j$ in market $t$ is obtained by aggregating the
individual demand with respect to the distribution of individual
preferences.
\begin{align} \label{eq:blp1}
\begin{split}
\varphi_j(X_t,P_t,\Xi_t)&=\int 1\{X_{jt}^{\prime }b+aP_{jt}+\sigma_\epsilon
e_{j}>-\Xi_{jt}\}\\
&\times 1\{(X_{jt}-X_{1t})^{\prime
}b+a(P_{jt}-P_{1t})+\sigma_\epsilon(e_{j}-e_{1})>-(\Xi_{jt}-\Xi_{1t})\} \\
\cdots 1\{(X_{jt}-&X_{Jt})^{\prime
}b+a(P_{jt}-P_{Jt})+\sigma_\epsilon(e_j-e_J)>-(\Xi_{jt}-\Xi_{Jt})\}f_
\theta(b,a,e)d\theta~,
\end{split}
\end{align}
for $j=1,\cdots,J$, while the aggregate demand for good 0 is given by
\begin{multline*}
\varphi_0(X_t,P_t,\Xi_t)=\int 1\{X_{1t}^{\prime }b+aP_{1t}+\sigma_\epsilon
e_1<-\Xi_{1t}\}\\[-6pt]
\cdots 1\{X_{Jt}^{\prime }b+aP_{Jt}+\sigma_\epsilon
e_J<-\Xi_{Jt}\}f_\theta(b,a,e)d\theta~,
\end{multline*}
where $(b,a,e_1,\cdots,e_J)$ are placeholders for the coefficients $
\theta_{it}=(\beta_{it},\alpha_{it},\epsilon_{i1t},\cdots,\epsilon_{iJt})$.
The researcher then observes the market shares of products $
S_{lt}=\varphi_l(X_t,P_t,\Xi_t), l\in \mathbb{L}$, where $\mathbb{L}
=\{0,1,\cdots, J\}.$ } } }
\end{example}

The second of example considers discrete choice, but allows for the
choice of bundles.

\begin{example}[Bundles]\label{ex:bundles}\rm
Each individual faces $J=2$
products and decides whether or not to consume a single unit of each of the
products. There are therefore four possible combinations $(Y_1,Y_2)$ of
consumption units, which we call \emph{bundles}. In addition to the utility
from consuming each good as in \eqref{eq:randomu}, the individuals gain
additional utility (or disutility) $\Delta_{it}$ if the two goods are
consumed simultaneously. Here, $\Delta_{it}$ is also allowed to vary across
individuals. The utility $U^*_{i,(Y_1,Y_2),t}$ from each bundle is therefore
specified as:
\begin{align}
&U^*_{i,(0,0),t}=0,  \notag \\
&U^*_{i,(1,0),t}=X_{1t}^{\prime }\beta_{it}+\alpha_{it}
P_{1t}+\Xi_{1t}+\sigma_\epsilon\epsilon_{i1t},~~
U^*_{i,(0,1),t}=X_{2t}^{\prime }\beta_{it}+\alpha_{it}
P_{2t}+\Xi_{2t}+\sigma_\epsilon\epsilon_{i2t},  \notag \\
&U^*_{i,(1,1),t}=X_{1t}^{\prime }\beta_{it}+X_{2t}^{\prime
}\beta_{it}+\alpha_{it} P_{1t}+\alpha_{it} P_{2t}
+\Xi_{1t}+\Xi_{2t}+\sigma_\epsilon\epsilon_{i1t}+\sigma_\epsilon
\epsilon_{i2t}+\Delta_{it}~,  \label{eq:bundleutil}
\end{align}
Each individual chooses a bundle that maximizes her utility. Hence, bundle $
(y_1,y_2)$ is chosen when $U^*_{i,(y_1,y_2),t}>U^*_{i,(y^{\prime
}_1,y^{\prime }_2),t}$ for all $(y^{\prime }_1,y^{\prime }_2)\ne (y_1,y_2)$.
For example, bundle $(1,0)$ is chosen if
\begin{align}\label{eq:10demand}
\begin{split}
&X_{1t}^{\prime }\beta_{it}+\alpha_{it}
P_{1t}+\Xi_{1t}+\sigma_\epsilon\epsilon_{i1t}> 0,\\
&X_{1t}^{\prime }\beta_{it}+\alpha_{it}
P_{1t}+\Xi_{1t}+\sigma_\epsilon\epsilon_{i1t}>X_{2t}^{\prime
}\beta_{it}+\alpha_{it} P_{2t}+\Xi_{2t}+\sigma_\epsilon\epsilon_{i2t},\\
&X_{1t}^{\prime }\beta_{it}+\alpha_{it}
P_{1t}+\Xi_{1t}+\sigma_\epsilon\epsilon_{i1t}\\
&\hspace{0.4in}>X_{1t}^{\prime
}\beta_{it}+\alpha_{it}
P_{1t}+\Xi_{1t}+\sigma_\epsilon\epsilon_{i1t}+X_{2t}^{\prime
}\beta_{it}+\alpha_{it}
P_{2t}+\Xi_{2t}+\sigma_\epsilon\epsilon_{i2t}+\Delta_{it}~.
\end{split}
\end{align}
Suppose the random coefficients $\theta_{it}=(\beta_{it}^{\prime
},\alpha_{it},\Delta_{it},\epsilon_{i1t},\epsilon_{i2t})$ have a joint
density $f_\theta$. The aggregate structural demand for $(1,0)$ can then be
obtained by integrating over the set of individuals satisfying
\eqref{eq:10demand} with respect to the distribution of the random
coefficients:
\begin{align}\label{eq:10agdemand}
\begin{split}
\varphi_{(1,0)}(X_t,P_t,\Xi_t)=\int 1\{&X_{1t}^{\prime }b+a
P_{1t}+\sigma_\epsilon e_1>-\Xi_{1t}\} \\
\times 1\{(X_{1t}&-X_{2t})^{\prime
}b+a(P_{1t}-P_{2t})+\sigma_\epsilon(e_1-e_2)>\Xi_{2t}-\Xi_{1t}\} \\
&\times 1\{X_{2t}^{\prime }b+a P_{2t}+\sigma_\epsilon
e_2+\Delta<-\Xi_{2t}\}f_\theta(b,a,\Delta,e)d\theta~.
\end{split}
\end{align}
The aggregate demand on other bundles can be obtained similarly. The
econometrician then observes a vector of aggregate demand on the bundles: $
S_{l,t}=\varphi_{l}(X_t,P_t,\Xi_t),l\in\mathbb{L}$ where $\mathbb{L}\equiv
\{(0,0),(1,0),(0,1),(1,1)\}$.
\end{example}

In Examples 2, we assume that the econometrician observes the aggregate
demand for all the respective bundles. We emphasize this point as it changes
the data requirement, and an interesting open question arises about what
happens if these requirements are not met. Examples of data sets that would
satisfy these requirements are when 1. individual observations are collected
through direct survey or scanner data on individual consumption (in every
market), 2. aggregate variables (market shares) are collected, but augmented
with a survey that asks individuals whether they consume each good
separately or as a bundle. 3. Finally, another possible data source are
producer's direct record of sales of bundles, provided each bundles are
recorded separately (e.g., when they are sold through promotional
activities). When discussing Example 2, we henceforth tacitly assume to have access to such data in
principle.

\subsection{Structural Demand}

\label{sec:inversion}

The first step toward identification of $f_\theta$ is to use a set of moment
conditions generated by instrumental variables to identify the aggregate
demand function $\varphi$. This section summarizes the identification result
obtained by BH (2014). Following BH (2014), we partition the covariates
as $X_{jt}=(X_{jt}^{(1)},X_{jt}^{(2)})\in \mathbb{R}\times\mathbb{R}
^{d_X-1}, $ and make the following assumption.

\begin{ass}
\label{as:index} The coefficient $\beta^{(1)}_{ij}$
is non-random for all $j$ and is normalized to 1.
\end{ass}

Assumption \ref{as:index} requires that at least one coefficient on the
covariates is non-random. Since we may freely choose the scale of utility,
we normalize the utility by setting $\beta^{(1)}_{ij}=1$ for all $j$. Under
Assumption \ref{as:index}, the utility for product $j$ can be written as $
U^*_{jt}=X_{jt}^{(2)^{\prime }}{\beta_{ij}}^{(2)}+\alpha_{ij}
P_{jt}+\sigma_\epsilon \epsilon_{ijt}+D_{jt}$, where $D_{jt}\equiv
X_{jt}^{(1)}+\Xi_{jt}$ is the part of the utility that is common across
individuals. Assumption \ref{as:index} (i) is arguably strong but will
provide a way to obtain valid instruments required to identify the
structural demand (see BH, 2014, Section 7 for details). Under this
assumption, $U^*_{ijt}$ is strictly increasing in $D_{jt}$ but unaffected by
$D_{kt}$ for all $k\ne j.$ In Example 1, together with a mild regularity
condition, this is sufficient for inverting the demand system to obtain $
\Xi_t$ as a function of the market shares $S_t$, price $P_t$, and exogenous
covariates $X_t$ (Berry, Gandhi, and Haile, 2013). In what follows, we
redefine the aggregate demand as a function of $(X_t^{(2)},P_t,D_t)$ instead
of $(X_t,P_t,\Xi_t)$ by
\begin{equation*}
    \phi(X_t^{(2)},P_t,D_t)\equiv
\varphi(X_t,P_t,\Xi_t),
\end{equation*} where $X_t=(X_t^{(1)},X_t^{(2)})$ and $
D_t=\Xi_t+X^{(1)}_t$ and make the following assumption

\begin{ass}
\label{as:inv1} For some subset $\tilde{\mathbb{L}}$ of $\mathbb{L}$ whose
cardinality is $J$, there exists a unique function $\psi:\mathbb{R}^{J\times
(d_X-1)}\times\mathbb{R}^J\times\mathbb{R}^{J}\to\mathbb{R}^{J}$ such that $
D_{jt}=\psi_j(X_t^{(2)},P_t,\tilde S_t)$ for $j=1,\cdots, J$, where $\tilde
S_t$ is a subvector of $S_t$, which stacks the components of $S_t$ whose
indices belong to $\tilde{\mathbb{L}}$.
\end{ass}

Under Assumption \ref{as:inv1}, we may write
\begin{equation}
\Xi _{jt}=\psi _{j}(X_{t}^{(2)},P_{t},\tilde{S}_{t})-X_{jt}^{(1)}.
\label{eq:invxi}
\end{equation}
This can be used to generate moment conditions in order to identify the
aggregate demand function.

\setcounter{example}{0}

\begin{example}[BLP, continued]
\textrm{\textrm{Let $\tilde{\mathbb{L}}=\{1,\cdots ,J\}$. In this setting,
the inversion discussed above is the standard Berry inversion. A key
condition for the inversion is that the products are \emph{connected
substitutes} (Berry, Gandhi, and Haile (2013)). The linear random
coefficient specification as in \eqref{eq:randomu} is known to satisfy this
condition. Then, Assumption \ref{as:inv1} follows. } }
\end{example}

In Example \ref{ex:bundles}, one may employ an alternative inversion
strategy to obtain $\psi$ in \eqref{eq:invxi} using only subsystems of
demand such as $\tilde{\mathbb{L}}=\{(1,0),(1,1)\}$ or $\tilde{\mathbb{L}}
=\{(0,0),(0,1)\}$. We defer details on this case to Section \ref{sec:bundles}
.

The inverted system in \eqref{eq:invxi}, together with the following
assumption, yields a set of moment conditions the researcher can use to
identify the structural demand.

\begin{ass}
\label{as:npiv1} There is a vector of instrumental variables $Z_{t}\in
\mathbb{R}^{d_Z}$ such that (i) $0=E[\Xi _{jt}|Z_{t},X_{t}]$, a.s.; (ii) for
any $B:\mathbb{R}^{Jk_{2}}\times \mathbb{R}^{J}\times \mathbb{R}
^{J}\rightarrow \mathbb{R}$ with $E[|B(X_{t}^{(2)},P_{t},\tilde{S}
_{t})|]<\infty $, it holds that
\begin{equation*}
E[B(X_{t}^{(2)},P_{t},\tilde{S}_{t})|Z_{t},X_{t}]=0\Longrightarrow
B(X_{t}^{(2)},P_{t},\tilde{S}_{t})=0,~a.s.
\end{equation*}
\end{ass}

Assumption \ref{as:npiv1} (i) is a mean independence assumption on $\Xi
_{jt} $ given a set of instruments $Z_{t}$, which also normalizes the
location of $\Xi _{jt}$. Assumption \ref{as:npiv1} (ii) is a completeness
condition, which is common in the nonparametric IV literature, see BH\
(2014) for a detailed discussion. However, the role it plays here is
slightly different, as the moment condition leads to an integral equation
which is different from nonparametric IV (Newey and Powell 2003), and more
resembles GMM.
In the online supplement Section \ref{sec:fullind}, we discuss an approach based on a
strengthening of the mean independence condition to full independence. In
case such a strengthening is economically palatable, we still retain the sum
$X_{jt}^{(1)}+\Xi _{jt}$. Where $X_{jt}^{(1)}$ is similar to a
dependent variable in nonparametric IV.

Given Assumption \ref{as:npiv1} and \eqref{eq:invxi}, the unknown function $
\psi $ can be identified through the following conditional moment
restrictions:
\begin{equation}
E[\psi
_{j}(X_{t}^{(2)},P_{t},S_{t})-X_{jt}^{(1)}|Z_{t},X_{t}]=0,~~j=1,\cdots ,J.
\end{equation}
We here state this result as a theorem. It is essentially Theorem 1 of BH (2014).

\begin{theorem}
\label{thm:npiv1} Suppose Assumptions \ref{as:index}-\ref{as:npiv1} hold.
Then, $\psi$ is identified.
\end{theorem}

Once $\psi$ is identified, the structural demand $\phi$ can be identified
nonparametrically in Examples 1 and 2.

\setcounter{example}{0}

\begin{example}[Multinomial choice, continued]
\textrm{\textrm{\textrm{Recall that $\psi$ is a unique function s.t.
\begin{align}
S_{jt}=\phi_j(X^{(2)}_t,P_t,D_t),~j=1,\cdots, J~\Leftrightarrow
~\Xi_{jt}=\psi_j(X^{(2)}_t,P_t,\tilde S_t)-X^{(1)}_{jt},~j=1,\cdots, J,
\label{eq:blpeq}
\end{align}
where $\tilde S_t=(S_{1t},\cdots, S_{Jt}).$ Hence, the structural demand $
(\phi_1,\cdots,\phi_J)$ is identified by Theorem \ref{thm:npiv1} and the
equivalence relation above. In addition, $\phi_0$ is identified through the
identity: $\phi_0=1-\sum_{j=1}^J\phi_j$. } } }
\end{example}

\begin{example}[Bundles, continued]
\textrm{\textrm{\textrm{Let $\tilde{\mathbb{L}}=\{(1,0),(1,1)\}$. Then $\psi $ is a unique function such that.
\begin{equation*}
S_{lt}=\phi _{l}(X_{t}^{(2)},P_{t},D_{t}),~l\in \tilde{\mathbb{L}}
~~~\Leftrightarrow ~~~\Xi _{jt}=\psi _{j}(X_{t}^{(2)},P_{t},\tilde{S}
_{t})-X_{t}^{(1)},~~j=1,2,
\end{equation*}
where $\tilde{S}_{t}=(S_{(1,0),t},S_{(1,1),t}).$ Theorem \ref{thm:npiv1} and
the equivalence relation above then identify the demand for bundles $(1,0)$
and $(1,1)$. This, therefore, only identifies subcomponents of $\phi $.
Although these subcomponents are sufficient for recovering the random
coefficient density, one may also identify the rest of the subcomponents by
taking $\tilde{\mathbb{L}}=\{(0,0),(0,1)\}$ and applying Theorem \ref
{thm:npiv1} again. } } }
\end{example}

\section{Identification of the Random Coefficient Density}

\label{sec:rcid}

This section contains the main innovation in this paper: We establish that
the density of random coefficients in the market-level demand models is
nonparametrically identified. Our strategy for identification of the random
coefficient density is to construct a function from the structural demand,
which is related to the density through an integral transform known as the
\emph{Radon transform}. More precisely, we construct a function $\Phi (w,u)$
such that
\begin{equation}
\frac{\partial \Phi (w,u)}{\partial u}=\mathcal{R}[f](w,u)~,
\label{eq:radonrep}
\end{equation}
where $f$ is the density of interest, $w$ is a vector in $\mathbb{R}^{q}$
(with $q$ the dimension of the random coefficients), normalized to have unit
length, and $u\in \mathbb{R}$ is a scalar. In what follows, we let $\mathbb{S
}^{q}\equiv \{v\in \mathbb{R}^{q}:\Vert v\Vert =1\}$ denote the unit sphere
in $\mathbb{R}^{q}$. $\mathcal{R}$ is the Radon transform defined pointwise
by
\begin{equation}
\mathcal{R}[f](w,u)=\int_{P_{w,u}}f(v)d\mu _{w,u}(v).  \label{eq:radon}
\end{equation}
where $P_{w,u}$ denotes the hyperplane $\{v\in \mathbb{R}^{q}:v^{\prime
}w=u\},$ and $\mu _{{w,u}}$ is the Lebesgue measure on $P_{w,u}$. See for
example Helgason (1999) for details on the properties of the Radon transform
including its injectivity. Our identification strategy is constructive and
will therefore suggest a natural nonparametric estimator. Applications of
the Radon transform to random coefficients models have been studied in
Beran, Feuerverger, and Hall (1996), Hoderlein, Klemel\"{a}, and Mammen
(2010), and Gautier and Hoderlein (2015).

Throughout, we maintain the following assumption.

\begin{ass}
\label{as:covariates} (i) For all $j\in\{1,\cdots, J\}$, $( X^{(2)}_{jt},
P_{jt}, D_{jt})$ are absolutely continuous with respect to Lebesgue measure
on $\mathbb{R}^{d_X-1}\times \mathbb{R}\times\mathbb{R}$; (ii) the random
coefficients $\theta$ are independent of $(X_{t},P_t,D_t)$.
\end{ass}

Assumption \ref{as:covariates} (i) requires that $
(X_{jt}^{(2)},P_{jt},D_{jt})$ are continuously distributed for all $j$. By
Assumption \ref{as:covariates} (ii), we assume that the covariates $
(X_{t},P_{t},D_{t})$ are exogenous to the individual heterogeneity. These
conditions are used to invert the Radon transform.

Before proceeding further, we overview our identification strategy in
relation to the key differences between the BLP and pure characteristics
models. Heuristically, for a given $(w,u)\in\mathbb{S}^q\times\mathbb{R}$,
the Radon transform aggregates individuals whose coefficients are on the
hyperplane $P_{w,u}$. For each $(w,u)$, we relate this aggregate value to a
feature of the demand with a specific product characteristics. By varying $
(w,u)$ and inverting the map $\mathcal{R}$ in \eqref{eq:radonrep}, we may
then recover the distribution of the random coefficients. A key step in this
identification argument is the construction of a function $\Phi$ satisfying
\eqref{eq:radonrep}. The two demand models suggest different strategies to
construct $\Phi.$ In the BLP model, we construct $\Phi$ for each product $j$
and recover the joint distribution of the coefficients $(\beta^{(2)}_{it},
\alpha_{it},\epsilon_{ijt})$. We take this approach because the presence of
the tastes for products requires us to isolate the demand for each product
from the rest. On the other hand, the pure characteristics model does not
require such an approach. Furthermore, both models allow the researcher to
combine demand across different markets to construct $\Phi$.

\subsection{BLP model}

Throughout this section, we let $\sigma_\epsilon=1$.\footnote{
Here, the scale of the taste for product $\epsilon_{ijt}$ is normalized
relative to the scale of $X^{(1)}_{jt}$ as we set the coefficient on $
X^{(1)}_{jt}$ to 1 in Assumption \ref{as:index}.} Recall that the demand for
good $j$ with the product characteristics $(X_{t},P_{t},\Xi _{t})$ is as
given in \eqref{eq:blp1}. Since $D_{t}=X_{t}^{(1)}+\Xi _{t}$, the demand in
market $t$ with $(X_{t}^{(2)},P_{t},D_{t})=(x^{(2)},p,\delta )$ is given by:
\begin{align}\label{eq:multnom}
\begin{split}
\phi _{j}(x^{(2)},p,\delta )=&\int 1\{x_{j}^{(2)}{}^{\prime }{b}
^{(2)}+ap_{j}+e_j>-\delta _{j}\}\\
&\times 1\{(x_{j}^{(2)}-x_{1}^{(2)}){}^{\prime }{b}
^{(2)}+a(p_{j}-p_{1})+(e_j-e_1)>-(\delta _{j}-\delta _{1})\} \\
\cdots 1\{(x_{j}^{(2)}-&x_{J}^{(2)})^{\prime }{b}
^{(2)}+a(p_{j}-p_{J})+(e_j-e_J)>-(\delta _{j}-\delta _{J})\}f_{\theta
}(b^{(2)},a,e)d\theta ~.
\end{split}
\end{align}
Suppose the vertical characteristics $\{D_{kt},k\ne j\}$ (for products other
than $j$) have a large enough support so that  $
(X_{jt}^{(2)}-X_{kt}^{(2)})^{\prime }{\beta}^{
(2)}_{it}+\alpha_{it}(P_{jt}-P_{kt})+(\epsilon_{ijt}-
\epsilon_{ikt})-D_{jt}>D_{kt}$ for all $k\ne j$ for some values of $
D_{kt},k\ne j$. The demand for good $j$ for such values of $D_{kt},k\ne j$
is then
\begin{align}  \label{eq:blp_phitilde}
\begin{split}
\tilde\Phi_j(x_{j}^{(2)},p_{j},\delta_{j})&=\lim_{\delta_1,\ldots,\delta_{j-1},
\delta_{j+1},\ldots, \delta_J \rightarrow -\infty} \phi_j(x^{(2)},p,\delta)\\
&=\int 1\{x_{j}^{(2)}{}^{\prime }{b}^{
(2)}+ap_{j}+e_j<-\delta_{j}\}f_{\vartheta_j }(b^{(2)},a,e_j)d\vartheta_j,
\end{split}
\end{align}
where $f_{\vartheta_j }$ is the joint density of the subvector $
\vartheta_{ijt}\equiv(\beta^{(2)}_{it},\alpha_{it},\epsilon_{ijt} )$ of the
random coefficients. Let $w\equiv (x_{j}^{(2)},p_{j},1)/\Vert
(x_{j}^{(2)},p_{j},1)\Vert $ and $u\equiv \delta_j/\Vert
(x_{j}^{(2)},p_{j},1)\Vert $. Define
\begin{align}\label{eq:Phi}
\begin{split}
\Phi(w,u)&\equiv\tilde\Phi_j\Bigg(\frac{x_j^{(2)}}{\Vert
(x_{j}^{(2)},p_{j},1)\Vert},\frac{p_j}{\Vert (x_{j}^{(2)},p_{j},1)\Vert},
\frac{\delta_j}{\Vert (x_{j}^{(2)},p_{j},1)\Vert}\Bigg) \\
&=\tilde\Phi_j(x_{j}^{(2)},p_{j},\delta _{j}),~~(x_{j}^{(2)},p_{j},\delta
_{j})\in \mathrm{supp}\,(X_{jt}^{(2)},P_{jt},D_{jt}),
\end{split}
\end{align}
where the second equality holds because normalizing the scale of $
(x_{j}^{(2)},p_{j},\delta _{j})$ does not change the value of $\tilde\Phi_j$
. $\Phi$ then satisfies
\begin{align}\label{eq:sum1}
\begin{split}
\Phi(w,&u )=-\int 1\{w^{\prime }\vartheta_j<-u \}f_{\vartheta_j
}(b^{(2)},a,e_j)d\vartheta_j \\
&=-\int_{-\infty }^{-u}\int_{P_{w,r}}f_{\vartheta_j }(b^{(2)},a,e_j)d\mu _{{
w,r}}(b^{(2)},a,e_j)dr=-\int_{-\infty }^{-u}\mathcal{R}[f_{\vartheta_j
}](w,r)dr~,
\end{split}
\end{align}
Hence, by taking a derivative with respect to $u$, we may relate $\Phi$ to
the random coefficient density through the Radon transform:
\begin{equation}
\frac{\partial \Phi(w,u)}{\partial u}=\mathcal{R}[f_{\vartheta_j }](w,u).
\label{eq:radon1}
\end{equation}
Note that since the structural demand $\phi $ is identified by Theorem \ref
{thm:npiv1}, $\Phi$ is nonparametrically identified as well. Hence, Eq.
\eqref{eq:radon1} gives an operator that maps the random coefficient density
to an object identified by the moment condition studied in the previous
section. To construct $\Phi$ described above and to invert the Radon
transform, we formally make the following assumptions. Below, for each $
1\le,j,k\le J$, we let $V_{jk} =(X_{jt}^{(2)}-X_{kt}^{(2)})^{\prime }{\beta}
^{
(2)}_{it}+\alpha_{it}(P_{jt}-P_{kt})+(\epsilon_{ijt}-\epsilon_{ikt})-D_{jt}$
and make the following assumptions on the support of the product
characteristics.\footnote{$V_{jk}$ is a random variable that varies across
individuals and markets and hence should be denoted as $V_{ijkt}$ in
principle. For conciseness, we drop subscripts $i$ and $t$ below.}

\begin{ass}
\label{as:rc_blp} Let $\mathcal{J}$ be a nonempty subset of $\{1,\cdots,J\}.$
For each $j\in\mathcal{J}$, let $\mathrm{supp}\,(V_{jk},k\ne j)\subset \mathrm{
supp}\,(D_{kt},k\ne j)$.
\end{ass}

Assumption \ref{as:rc_blp} requires that one may vary the vertical
characteristics of the alternative products $\{D_{kt},k\neq j\}$ on a large
enough support so that the demand for product $j$ is determined through its
choice between product $j$ and the outside good. This identification
argument therefore uses a \textquotedblleft thin\textquotedblright\
(lower-dimensional) subset of the support of the covariates, which is due to
the presence of the tastes for products. This is in remarkable contrast with
the identification of the random coefficients density in the PCM (analyzed in the next section) which does
not rely on thin sets.

\begin{ass}
\label{as:analytic} One of the following conditions hold

\begin{enumerate}
\item[(i)] $\bigcup_{j\in \mathcal{J}}\mathrm{supp}\,(X_{jt}^{(2)},P_{jt},D_{jt})$
has full support in $\mathbb{R}^{d_X-1}\times\mathbb{R}\times\mathbb{R}$.

\item[(ii)] $\bigcup_{j\in \mathcal{J}}\mathrm{supp}\,(X_{jt}^{(2)},P_{jt})$
contains an open ball $B_\mathcal{J} \subset \mathbb{R}^{d_X-1}\times\mathbb{
R}$. For every $(x,p) \in B_\mathcal{J}$ and every $(b^{(2)},a,e_j) \in
\mathrm{supp}\,(\vartheta_j)$ it holds that
\begin{align}
\big(x,p,-x^{\prime} b^{(2)} - ap - e_j\big) \in \bigcup_{j\in \mathcal{J}}
\mathrm{supp}\,\left(X_{jt}^{(2)},P_{jt},D_{jt}\right).  \label{eq:supp_blp}
\end{align}
Furthermore, all the absolute moments of each component of $\theta_{it}$ are
finite, and for any fixed $z\in \mathbb{R}_+$, $\lim_{l\to\infty}\frac{z^l
}{l!} E[(|\theta_{it}^{(1)}|+\cdots+|\theta^{(d_\theta)}_{it}|)^l]=0$.
\end{enumerate}
\end{ass}

Assumption \ref{as:analytic} (i) is our benchmark assumption. Under this
assumption, no restrictions on $\theta_{it}$ are necessary for
identification. In fact, the identification strategy would be valid for
arbitrary Borel measures and may also be applied to settings where $
\theta_{it}$ does not have a density.\footnote{
More precisely, the Radon transform $\mathcal{R}[f_{\vartheta _{j}}](w,u)$
gives $f_{\vartheta _{j}}$'s integral along each hyperplane $P_{w,u}=\{v\in
\mathbb{R}^{d_{\theta }}:v^{\prime }w=u\}$ defined by the \emph{angle} $
w=(x_{j}^{(2)},p_{j},1)/\Vert (x_{j}^{(2)},p_{j},1)\Vert $ and \emph{offset}
$u=\delta _{j}/\Vert (x_{j}^{(2)},p_{j},1)\Vert $. For recovering $
f_{\vartheta _{j}}$ from its Radon transform, one needs exogenous variations
in both. Our proof uses the fact that varying $w$ over the hemisphere $
\mathbb{H}_{+}\equiv \{w=(w_{1},w_{2},\cdots ,w_{d_{\vartheta _{j}}})\in
\mathbb{S}^{d_{\vartheta _{j}}-1}:w_{d_{\vartheta _{j}}}\geq 0\}$ and $u$
over $\mathbb{R}$ suffices to recover $f_{\vartheta _{j}}$.} However, this
large support assumption is stringent and may be violated by various product
characteristics and prices used in practice. Hence, it should be viewed as a
benchmark to understand what the model requires to identify the distribution
$f_\theta$ of $\theta_{it}$ if one does not impose any restriction on it.

Assumption \ref{as:analytic} (ii) is an alternative condition, which relaxes
the support requirement significantly. Instead of a large support, it is
enough for the product characteristics to have a properly combined support
that contains a (possibly small) open ball $B_{\mathcal{J}}$ in it. This
includes as a special case where a single product's characteristics $
(X_{jt}^{(2)},P_{jt})$
contains an open ball, which can be met in various applications. Even if
such a product does not exist, identification of the random coefficient
density is possible as long as the required support condition is met by
combining the supports of multiple products belonging to $\mathcal{J}$. This
means that our identification strategy may use variations of $
(X_{jt}^{(2)},P_{jt})$ across products. To illustrate, consider three
products $J=3$. If $(D_{2t},D_{3t})$ have a large support in the sense of
Assumption \ref{as:rc_blp} ($\mathcal{J}=\{1\}$ in this case),
identification of the random coefficient density is possible as long as the
characteristics of good 1 contains an open ball. If all $\{D_{jt}\}_{j=1}^3$
jointly have a large support (this implies $\mathcal{J}=\{1,2,3\}$), our
requirement on $(X_{jt}^{(2)},P_{jt})$ becomes even milder as we only need
to construct an open ball by combining the characteristics of all three products.

The condition in \eqref{eq:supp_blp} allows for a bounded support of $
(X_{jt}^{(2)},P_{jt})$. Further, if $\vartheta _{j}$ has a bounded
support, Assumption \ref{as:analytic} (ii) will allow for a bounded support of $D_{j}$. The price to pay for this relaxation of the
support requirement is a regularity assumption on the moments of $\theta
_{it}$. This rules out heavy tailed distributions that are not determined by
their moments. A sufficient, yet stronger than necessary, condition for this
assumption is a compact support of $f_{\theta }$. Under Assumption \ref
{as:analytic} (ii), the characteristic function $w\mapsto \varphi
_{\vartheta _{j}}(tw)$ of $\vartheta _{ijt}$ (a key element of the Radon
inversion) is analytic and thereby uniquely determined by its restriction to
a non-empty full dimensional subset of its domain.\footnote{
This type of moment condition on $\theta _{it}$ is common in the recent
literature. See, for example, Hoderlein, Holzmann, and Meister (2014) and
Masten (2014).}
Hence, $f_{\vartheta _{j}}$ can be identified if one varies $
(X_{jt}^{(2)},P_{jt})$ on a full dimensional subset.


Under the conditions given in the theorem below, the Radon inversion
identifies $f_{\vartheta_j}$. If one is interested in the joint density of
the coefficients on the product characteristics $(\beta^{(2)}_{it},
\alpha_{it})$, one may stop here as marginalizing $f_{\vartheta_j}$ gives
the desired density. The joint distribution of the coefficients including
the tastes for products can be identified under an additional independence
assumption. We state this result in the following theorem.

\begin{theorem}
\label{thm:fbeta} Suppose Assumptions \ref{as:index}-\ref{as:analytic} hold.
Suppose the conditional distribution of $\epsilon_{ijt}$ given $
(\beta^{(2)}_{it},\alpha_{it})$ is identical for all $j\in\mathcal{J}$.
Then, (i) for each $j\in\mathcal{J}$, the density $f_{\vartheta_j}$ is
identified, where $\vartheta_{ijt}=(\beta^{(2)}_{it},\alpha_{it},
\epsilon_{ijt})$; (ii) If, in addition, $\{\epsilon_{ijt},j\in\mathcal{J}\}$
are independently distributed (across $j$) conditional on $
(\beta^{(2)}_{it},\alpha_{it})$, the joint density $f_{\theta_{\mathcal{J}}}$
of $\theta_{\mathcal{J}}=(\beta^{(2)}_{it},\alpha_{it},\{\epsilon_{ijt}\}_{j
\in\mathcal{J}})$ is identified.
\end{theorem}

An immediate corollary is the following.

\begin{corollary}
\label{cor:fbeta}  Suppose Assumptions \ref{as:index}-\ref{as:analytic}
hold. Let $\{\epsilon_{ijt}\}_{j=1}^J$ be i.i.d. (across $j$)
conditional on $(\beta^{(2)}_{it},\alpha_{it})$. Then, the joint density $
f_\theta$ of all random coefficients $\theta_{it}=(\beta^{(2)}_{it},
\alpha_{it},\{\epsilon_{ijt}\}_{j=1}^J)$ is identified.
\end{corollary}


\begin{remark}\rm
\textrm{Theorems \ref{thm:npiv1} and \ref{thm:fbeta} shed light on the roles
played by the key features of the BLP-type demand model: the invertibility
of the demand system, instrumental variables, and the linear random
coefficients specification. In Theorem \ref{thm:npiv1}, the invertibility
and instrumental variables play key roles in identifying the demand. Once
the demand is identified, one may ``observe'' the vector $
(X^{(2)}_t,P_t,D_t) $ of product characteristics. This is possible because
the invertibility of demand allows one to recover the unobserved product
characteristics $\Xi_t$ from the market shares $S_t$ (together with other
covariates). One may then vary $(X^{(2)}_t,P_t,D_t)$ across markets in a
manner that is exogenous to the individual heterogeneity $\theta_{it}$.
Theorem \ref{thm:fbeta} and Corollary \ref{cor:fbeta} show that this
exogenous variation combined with the linear random coefficients
specification allows to trace out the distribution of $\theta_{it}$. }
\end{remark}

\begin{remark}\rm
\textrm{\label{rem:iid} The identical distribution assumption on the tastes
for products  in Theorem \ref{thm:fbeta} is compatible with commonly used
utility specifications and can also be relaxed at the cost of a stronger
support condition on the product characteristics. In applications, it is
often assumed that the utility of product $j$ is
\begin{align}
U_{ijt}^{\ast }=\beta^0_{it}+\tilde X_{jt}^{\prime }\beta _{it}+\alpha
_{it}P_{jt}+\Xi _{jt}+\tilde \epsilon_{ijt},
\end{align}
where $\tilde X_{jt}$ is a vector of non-constant product characteristics, $
\beta^0_{it}$ is an individual specific intercept, which measures the
utility difference between inside goods and the outside good, and $
\tilde\epsilon_{ijt}$ is a mean zero error that follows the Type-I extreme
value distribution. The requirement that $\epsilon_{ijt}=\beta^0_{it}+\tilde
\epsilon_{ijt}$ are i.i.d. across $j$ (conditional on $(\beta_{it},
\alpha_{it})$) can be met if $\tilde\epsilon_{ijt}$ are i.i.d. across $j$. }

\textrm{If for each $j$, $(X_{jt}^{(2)},P_{jt},D_{jt})$ fulfills the support
condition in Assumption \ref{as:analytic} (i) or Assumption \ref{as:analytic}
(ii), one can drop the identical distribution assumption. This is because
one can identify $f_{\vartheta _{j}}$ for all $j$ by inverting the Radon
transform in \eqref{eq:radon1} repeatedly. This in turn implies that the
distribution of $\epsilon _{ijt}$ conditional on $(\beta _{it}^{(2)},\alpha
_{it})$ is identified for each $j$. If the tastes for products $
\{\epsilon_{ijt}\}_{j=1}^J$ are mutually independent (conditional on $(\beta
_{it}^{(2)},\alpha _{it})$), as is commonly assumed in BLP, the joint density $f_{\theta }$ is identified. }

\textrm{Finally, we comment on what an additional parametric assumption may
add to our result. If one assumes that the tastes for products are i.i.d.
and follows a parametric distribution, Eq \eqref{eq:multnom} reduces to $
\phi _{j}(x^{(2)},p,\delta )=\int L(x^{(2)}{}^{\prime
(2)}+ap+\delta)f_{(\beta,\alpha)}(b,a)dbda,$ for some function $L$, e.g. $L$
is the logit function when $\{\epsilon_{ijt}\}$ follows a Type-I extreme
value distribution. This type of integral equation is considered in Fox,
Kim, Ryan, and Bajari (2012) in the context of individual-level demand model
without endogeneity. Given that $\phi$ is identified, we believe that it is
possible to extend their framework to the market-level demand model with
endogeneity and identify $f_{(\beta,\alpha)}$ semiparametrically.  This
approach may allow us to relax some of the support conditions. To keep a
tight focus on nonparametric identification, we leave this extension
for future work.}
\end{remark}

\begin{remark}
\label{rem:support} \rm Our identification result reveals the nature of
the BLP-type demand model. A positive aspect of our result is that the
preference is nonparametrically identified if one observes full dimensional
variations in the consumers' choice sets (represented by $
(X^{(2)}_{jt},P_{jt},D_{jt})$) across markets. The identifying power is
quite strong, if the product characteristics jointly span a full support,
i.e. $\bigcup_{j\in\mathcal{J}}(X^{(2)}_{jt},P_{jt},D_{jt})=\mathbb{R}
^{d_X-1}\times\mathbb{R}\times\mathbb{R}.$ On the other hand, if the product
characteristics have limited variations, the identifying power of the model
on the distribution of preferences may be limited. In particular,
identification is not achieved only with discrete covariates. Hence, for
such settings, one needs to augment the model structure with a parametric
specification. Another interesting direction would be to conduct partial
identification analysis on functionals of $f_\theta$, while imposing weak
support restrictions. We leave this possibility for future research.
\end{remark}

\subsection{Pure Characteristics Demand Models}\label{sec:PCM}

Throughout this section, we consider the following utility
specification where each product's utility is fully determined by the tastes
for the product characteristics:
\begin{equation}
U_{ijt}^{\ast }\equiv X_{jt}^{\prime }\beta _{it}+\alpha _{it}P_{jt}+\Xi
_{jt},~j=1,\cdots ,J~.  \label{eq:randomu_pure}
\end{equation}
In other words, we set $\sigma _{\epsilon }=0$ in \eqref{eq:randomu}. For
this model, we employ a different, and arguably less restrictive, strategy
from the one adopted in the previous section to construct $\Phi $ in
\eqref{eq:radonrep}. Below, we maintain Assumptions \ref{as:index}-\ref
{as:npiv1}, which ensure the identification of demand by Theorem \ref
{thm:npiv1}. The demand for good $j$ with the product characteristics $(X_{t},P_{t},\Xi
_{t})$ is as given in \eqref{eq:blp1} but with $\sigma_\epsilon=0$. Since $
D_{t}=X_{t}^{(1)}+\Xi _{t}$, the demand in market $t$ with $
(X_{t}^{(2)},P_{t},D_{t})=(x^{(2)},p,\delta )$ is given by:
\begin{multline}
\phi _{j}(x^{(2)},p,\delta )=\int 1\{x_{j}^{(2)}{}^{\prime }{b}
^{(2)}+ap_{j}>-\delta _{j}\}1\{(x_{j}^{(2)}-x_{1}^{(2)}){}^{\prime }{b}
^{(2)}+a(p_{j}-p_{1})>-(\delta _{j}-\delta _{1})\} \\
\cdots 1\{(x_{j}^{(2)}-x_{J}^{(2)}){}^{\prime }{b}^{(2)}+a(p_{j}-p_{J})>-(
\delta _{j}-\delta _{J})\}f_{\theta }(b^{(2)},a)d\theta ~.
\end{multline}

For any subset $\mathcal{J}$ of $\{1,\cdots ,J\}\setminus \{j\}$, let $
\mathcal{M}_{\mathcal{J}}$ denote the map $(x^{(2)},p,\delta )\mapsto (
\acute{x}^{(2)},\acute{p},\acute{\delta})$ that is uniquely defined by the
following properties:
\begin{align*}
(\acute{x}_{j}^{(2)}-\acute{x}_{i}^{(2)},\acute{p}_{j}-\acute{p}_{i},\acute{
\delta}_{j}-\acute{\delta}_{i})&
=-(x_{j}^{(2)}-x_{i}^{(2)},p_{j}-p_{i},\delta _{j}-\delta _{i}),~\forall
i\in \mathcal{J}~, \\
(\acute{x}_{i}^{(2)},\acute{p}_{i},\acute{\delta}_{i})&
=(x_{i}^{(2)},p_{i},\delta _{i}),~\forall i\notin \mathcal{J}~.
\end{align*}
In words, for a given product $j$ and product characteristics $(x^{(2)},p,\delta )$, this map finds another value $(
\acute{x}^{(2)},\acute{p},\acute{\delta})$ of  product characteristics such that, for products $i$ belonging to $\mathcal J$, the difference in the product characteristics (e.g. $\acute{x}_{j}^{(2)}-\acute{x}_{i}^{(2)}$) coincides with the original value (e.g., $x_{j}^{(2)}-x_{i}^{(2)}$) in terms of magnitude but has an opposite sign. For products $i$ not belonging to $\mathcal J$, the map sets their product characteristics the original value $(x^{(2)}_i,p_i,\delta_i)$.

Consider the
composition $\phi _{j}\circ \mathcal{M}_{\mathcal{J}}(x^{(2)},p,\delta )$.
If $( \acute{x}^{(2)},\acute{p},\acute{\delta})$ is in the support, this
corresponds to the demand of product $j$ in some market (say $t^{\prime }$)
with $(X_{t^{\prime }}^{(2)},P_{t^{\prime }},D_{t^{\prime }})=( \acute{x}
^{(2)},\acute{p},\acute{\delta})$. We then define
\begin{equation}
\tilde{\Phi}_{j}(x_{j}^{(2)},p_{j},\delta _{j})\equiv -\sum_{\mathcal{J}
\subseteq \{1,\cdots J\}\setminus \{j\}}\phi _{j}\circ \mathcal{M}_{\mathcal{
J}}(x^{(2)},p,\delta )~.  \label{eq:sum_pure}
\end{equation}
Eq \eqref{eq:sum_pure} aggregates the structural demand function for good $j$
in different markets to define a function, which can be related to the
random coefficient density in a simple way. This operation can be easily
understood when $J=2$, where for example demand for product 1 is given by
\begin{multline*}
\phi _{1}(x^{(2)},p,\delta )=\int 1\{x_{1}^{(2)}{}^{\prime }{b}
^{(2)}+ap_{1}<-\delta _{1}\} \\
\times 1\{(x_{1}^{(2)}{}-x_{2}^{(2)}){}{}^{\prime }{b}
^{(2)}+a(p_{1}-p_{2})<-(\delta _{1}-\delta _{2})\}f_{\theta
}(b^{(2)},a)d\theta ~.
\end{multline*}
Then, $\tilde{\Phi}_{1}$ is given by
\begin{align}\label{eq:sum1_pure}
\begin{split}
&\tilde{\Phi}_{1}(x_{1}^{(2)},p_{1},\delta _{1}) =-\phi _{1}\circ \mathcal{M}
_{\emptyset }(x^{(2)},p,\delta)-\phi _{1}\circ \mathcal{M}
_{\{2\}}(x^{(2)},p,\delta) \\
& =-\int 1\{x_{1}^{(2)}{}^{\prime }b^{(2)}+ap_{1}<-\delta _{1}\}\Big(
1\{(x_{1}^{(2)}-x_{2}^{(2)})^{\prime }b^{(2)}+a(p_{1}-p_{2})<-(\delta
_{1}-\delta _{2})\}\\
& \hspace{76pt} +1\{(x_{1}^{(2)}{}-x_{2}^{(2)})^{\prime
}b^{(2)}+a(p_{1}-p_{2})>-(\delta _{1}-\delta _{2})\}\Big)f_{\theta
}(b^{(2)},a)d\theta \\
& =-\int 1\{x_{1}^{(2)}{}^{\prime }b^{(2)}+ap_{1}<-\delta _{1}\}f_{\theta
}(b^{(2)},a)d\theta
\end{split}
\end{align}
This shows that aggregating the demand in the two markets with $
(X_{t}^{(2)},P_{t},D_{t})=(x^{(2)},p,\delta )$ and $(X_{t^{\prime
}}^{(2)},P_{t^{\prime }},D_{t^{\prime }})=(\acute{x}^{(2)},\acute{p},\acute{
\delta})$ yields a function $\tilde{\Phi}_{1}$ that depends only on product
1's characteristic $(x_{1}^{(2)},p_{1},\delta _{1})$ through a single index
in \eqref{eq:sum1_pure}. This then allows us to trace out the random
coefficients density by varying product 1's characteristic as done in the
BLP model. Since the operation above yields a function that depends only on
the characteristic of a single product, we call it \emph{marginalization of
demand}.\footnote{
Note however that this marginalized demand still depends on the joint
distribution of the entire random coefficient vector.}

Eq. \eqref{eq:sum_pure} generalizes this argument to settings with $J\geq 2$
. For the marginalization of demand to work, the product characteristic $(
\acute{x}^{(2)},\acute{p},\acute{\delta})=\mathcal{M}_{\mathcal{J}}(x^{(2)},
p,\delta)$ needs to be an observable value, meaning it must be in the
support. Formally, a value of the product characteristic $(x^{(2)},
p,\delta)\in \text{supp}( X^{(2)}_t, P_{t}, D_{t})$ is said to permit \emph{
marginalization of demand with respect to product $j$} if
\begin{align}
\mathcal{M}_{\mathcal{J}}( x^{(2)}, p,\delta)\in \text{supp}( X^{(2)}_t,
P_{t}, D_{t}),~\forall \mathcal{J}\subseteq\{1,\cdots, J\}\setminus\{j\}.
\end{align}

As done in the BLP setting, we will only require that a rich enough set to
recover $f_\theta$ can be constructed by combining the supports of multiple
products' characteristics. Toward this end, for each $j\in \{1,\cdots,J\}$,
let $\pi_j$ be the projection map such that $(x^{(2)}_{j},p_j,\delta_j)=
\pi_j(x^{(2)}, p,\delta)$, and define the following sets:
\begin{align*}
\mathcal{H}_j&\equiv \{(x^{(2)}, p,\delta)\in \text{supp}( X^{(2)}_t,
P_{t},D_{t}): \mathcal{M}_{\mathcal{J}}( x^{(2)}, p,\delta)\in \text{supp}
(X^{(2)}_t, P_{t}, D_{t}), \\
&\hspace{0.3in}\text{ for all }\mathcal{J}\subseteq\{1,\cdots,J\}\setminus
\{j\}\}, \\
\mathcal{S}_j&\equiv\{(x^{(2)}_{j},p_j,\delta_j)\in \text{supp}(
X^{(2)}_{jt}, P_{jt},D_{jt}):(x^{(2)}_{j},p_j,\delta_j)=\pi_j(x^{(2)},
p,\delta),\\
&\hspace{0.3in}~\text{ for some }(x^{(2)}, p,\delta)\in \mathcal{H}_j\}.
\end{align*}
In words, $\mathcal{H}_j$ is the set of the entire product characteristic
vectors for which marginalization with respect to product $j$ is permitted. $
\mathcal{S}_j$ is the coordinate projection of $\mathcal{H}_j$ onto the
space of product $j$'s characteristics. We then make the following
assumption.

\begin{ass}
\label{as:support_pure} One of the following conditions hold:
\begin{enumerate}
\item[(i)] $\bigcup_{j=1}^J\mathcal{S}_j=\mathbb{R}^{d_X-1}\times\mathbb{R}\times
\mathbb{R}$;

\item[(ii)] $\bigcup_{j=1}^J\mathcal{S}_j=\mathbb{E}\times \mathbb{D}$, where $
\mathbb{E}$ contains an open ball $B \subset \mathbb{R}^{d_X-1}\times\mathbb{
R}$, and $\mathbb{D}\subseteq \mathbb{R}$. For every $(x,p) \in B$ and every
$(b^{(2)},a) \in \mathrm{supp}\,(\theta_{it})$, it holds that $
(x,p,-x^{\prime} b^{(2)} - ap) \in \bigcup_{j=1}^J\mathcal{S}_j$.

Furthermore, all the absolute moments of each component of $\theta_{it}$ are
finite, and for any fixed $z\in \mathbb{R}_+$, $0=\lim_{l\to\infty}\frac{z^l
}{l!}(E[|\theta_{it}^{(1)}|^l]+\cdots+E[|\theta^{(d_\theta)}_{it}|^l])$.
\end{enumerate}
\end{ass}

The idea behind Assumption \ref{as:support_pure} is as follows. For the
moment, suppose we don't impose any moment condition on the random
coefficient density. Also, fix a benchmark product $j$. For any $
(x^{(2)}_{j},p_j,\delta_j)\in\mathcal{S}_j$, one may find a vector $
(x^{(2)}, p,\delta)$ of all product characteristics for which
marginalization of demand is allowed. Then, one would wish to vary $
(x^{(2)}_{j},p_j,\delta_j)$ to trace out the random coefficient density.
This is possible, of course, if $\mathcal{S}_j=\mathbb{R}^{d_X-1}\times
\mathbb{R}\times\mathbb{R}$, meaning that marginalization is possible
everywhere with respect to product $j.$ However, this assumption may be too
strong in empirical applications. One may not be able to find any single
product, for which this condition is satisfied.  Assumption \ref
{as:support_pure} (i) relaxes this requirement substantially using the
structure of the model. Observe that the identification argument is
symmetric across products because only the characteristics matter. Hence,
the argument is valid as long as, for each $(\mathbf{x}^{(2)},\mathbf{p},
\mathbf{d})\in \mathbb{R}^{d_X-1}\times\mathbb{R}\times \mathbb{R}$, one can
find \emph{some} product for which marginalization is permitted. This is the
reason why it is enough to ``patch'' $\mathcal{S}_j$s together to $
\mathbb{R}^{d_X-1}\times\mathbb{R}\times \mathbb{R}$ in Assumption \ref
{as:support_pure} (i). This condition can be made even weaker with the help
of an additional moment condition. In Assumption \ref{as:support_pure} (ii),
we only require that $\mathcal{S}_j$s combined together contain an open ball (in terms of $(\mathbf{x}^{(2)},\mathbf{p})$).
This support requirement is quite mild, and hence it can be satisfied even
if each product's characteristic has limited variation across markets. Note
also that, if $\mathrm{supp}\,(\theta_{it})$ is compact, the support of $
D_{jt}$ can be compact as well.

It is important to note that we construct $\tilde\Phi_j$ without relying on
any ``thin'' (lower-dimensional) subset of the support of the product
characteristics as done in the BLP model. Instead, we construct $\tilde\Phi_j
$ in \eqref{eq:sum_pure} by combining the demand in different markets.
This is desirable as estimators that rely on thin or irregular
identification may have a slow rate of convergence (Khan and Tamer, 2010).
In the pure characteristics model, the individuals have varying tastes
(random coefficients) over the product characteristics but not over the
products themselves. This is the key feature of the model that allows us to
identify the random coefficients through the variation of the product
characteristics $( X^{(2)}_t, P_{t}, D_{t})$. In contrast, in the BLP model,
there was an additional taste for the product itself, which was the main
reason for using the thin set to isolate the demand for each product.

Given Assumption \ref{as:support_pure}, we now construct $\Phi$ in Eq.
\eqref{eq:radonrep}. For each $(\mathbf{x}^{(2)},\mathbf{p},\mathbf{d})\in
\bigcup_{j=1}^J\mathcal{S}_j$, let $w\equiv (\mathbf{x}^{(2)},\mathbf{p}
)/\Vert (\mathbf{x}^{(2)},\mathbf{p})\Vert $ and $u\equiv \mathbf{d}/\Vert (
\mathbf{x}^{(2)},\mathbf{p})\Vert $. Define
\begin{align}
\Phi(w,u)\equiv\tilde\Phi_j\Big(\frac{\mathbf{x}^{(2)}}{\Vert (\mathbf{x}
^{(2)},\mathbf{p})\Vert},\frac{\mathbf{p}}{\Vert (\mathbf{x}^{(2)},\mathbf{p}
)\Vert},\frac{\mathbf{d}}{\Vert (\mathbf{x}^{(2)},\mathbf{p})\Vert}\Big),~
\text{where}~(\mathbf{x}^{(2)},\mathbf{p},\mathbf{d})\in \mathcal{S}_j.
\label{eq:Phi_pure}
\end{align}
Here, for each $(\mathbf{x}^{(2)},\mathbf{p},\mathbf{d})$, any $j$ can be
used to construct $\tilde\Phi_j$ through marginalization as long as $
\mathcal{S}_j$ contains $(\mathbf{x}^{(2)},\mathbf{p},\mathbf{d})$. Then $
\Phi$ is defined on a set that is rich enough to invert the Radon (or
limited angle Radon) transform. The rest of the analysis parallels our
analysis of the BLP model.\footnote{
Note that the additional independence (or i.i.d.) assumptions on $
(\epsilon_{i1t},\cdots,\epsilon_{iJt})$ is not needed in the pure
characteristics model.} We therefore obtain the following point
identification result.

\begin{theorem}
\label{thm:fbeta_pure} Suppose Assumptions \ref{as:index}-\ref{as:covariates}
, and \ref{as:support_pure} hold. Then, $f_\theta$ is identified in the pure
characteristics demand model, where $\theta_{it}=(\beta^{(2)}_{it},
\alpha_{it}).$
\end{theorem}



\subsection{Bundle choice (Example \protect\ref{ex:bundles})}\label{sec:bundles}

In this section, we consider $\sigma_\epsilon=1$, but it is also possible to analyze the case without the tastes for products.\footnote{For the setting without the tastes for products, we refer to Dunker, Hoderlein, and Kaido (2013), an earlier version of the paper.} We consider an alternative procedure for inverting the demand in Example \ref{ex:bundles}.  This is because this example (and also the example in the next section)
has a specific structure. We note that the inversion of Berry, Gandhi, and
Haile (2013) can still be applied to bundles if one treats each bundle as a
separate good and recast the bundle choice problem into a standard
multinomial choice problem. However, as can be seen from
\eqref{eq:bundleutil}, Example \ref{ex:bundles} has the additional structure
that the utility of a bundle is the combination of the utilities for each
good and extra utilities, and hence the model does not involve any bundle
specific unobserved characteristic. This structure in turn implies that the
dimension of the unobservable product characteristic $\Xi_t$ equals the
number of goods $J$, while the econometrician in general observes $\dim(S)=
2^J$ aggregate choice probabilities over bundles, which leads to a system of equations
 whose number of restrictions exceeds the number of unknown quantities. This suggests that (i) using only a part
of the demand system is sufficient for obtaining an inversion, which can be
used to identify $f_\theta$ and (ii) using additional subcomponents of $S$,
one may potentially overidentify the parameter of interest. We therefore
consider an inversion that exploits a monotonicity property of the demand
system that follows from this structure.\footnote{
The additional structure can potentially be tested. In Example 2, one may
identify the demand for bundles (1,0) and (1,1) using the inversion
described below under the hypothesis that eq. \eqref{eq:bundleutil} holds.
Further, treating (1,0), (0,1), and (1,1) as three separate goods (and (0,0)
as an outside good) and applying the inversion of Berry, Gandhi, and Haile
(2013), one may identify the demand for bundles (1,0) and (1,1) without
imposing \eqref{eq:bundleutil}. The specification can then be tested by
comparing the demand functions obtained from these distinct inversions. We
are indebted to Phil Haile for this point.} For this, we assume that the
following condition is met.


\begin{condition}
\label{cond:bundles} The random coefficient density $f_{\theta}$ is
continuously differentiable. In addition, $(\epsilon_{i1t},\epsilon_{i2t})$ and $
(D_{1t},D_{2t})$ have full supports in $\mathbb{R}^2$ respectively.
\end{condition}

Let $\tilde{\mathbb{L}}=\{(1,0),(1,1)\}.$ From \eqref{eq:10agdemand}, it is
straightforward to show that $\varphi_{(1,0)}$ is strictly increasing in $
D_{1t}$ but is strictly decreasing in $D_{2t}$, while $\varphi_{(1,1)}$ is
strictly increasing both in $D_{1t}$ and $D_{2t}$. Hence, the Jacobian
matrix is non-degenerate. Together with a mild support condition on $
(D_{1t},D_{2t})$, this allows to invert the demand (sub)system and write $
\Xi_{jt}=\psi_j(X^{(2)}_t,P_t,\tilde S_t)-X^{(1)}_{jt},$ where $\tilde
S_t=(S_{(1,0),t},S_{(1,1),t})$. This ensures Assumption \ref{as:inv1} in
this example (see Lemma \ref{lem:inv1} given in the appendix). By Theorem
\ref{thm:npiv1}, one can then nonparametrically identify subcomponents $
(\varphi_{(1,0)},\varphi_{(1,1)})$ of the demand function $\varphi$.

One may alternatively choose $\tilde{\mathbb{L}}=\{(0,0),(0,1)\}$, and the
argument is similar, which then identifies $(\varphi_{(0,0)},
\varphi_{(0,1)}) $, and hence all components of the demand function $\varphi$
are identified. This inversion is valid even if the two goods are
complements. This is because the inversion uses the monotonicity property of
the aggregate choice probabilities on bundles (e.g. $\phi_{(1,0)}$ and $
\phi_{(1,1)}$) with respect to $(D_{1t}, D_{2t})$. Hence, even if the
aggregate share of each good (e.g. aggregate share on good 1: $
\sigma_1=\phi_{(1,0)}+\phi_{(1,1)}$) is not invertible in the price $P_t$
due to the presence of complementary goods, one can still obtain a useful
inversion provided that aggregate choice probabilities on bundles are
observed.

Given the demand for bundles, we now analyze identification of the random
coefficient density. By \eqref{eq:bundleutil}, the demand for bundle (0,0)
is given by
\begin{align}  \label{eq:dem00}
&\phi_{(0,0)}(x^{(2)},p,\delta) =\int 1\{x^{(2)}_1{}^{\prime
}b^{(2)}+ap_1+e_1<-\delta_1\}1\{x^{(2)}_2{}^{\prime
}b^{(2)}+ap_2+e_2<-\delta_2\}   \\
&\qquad\times1\{(x_1^{(2)}+x_2^{(2)})^{\prime
}b^{(2)}+a(p_1+p_2)+(e_1+e_2)+\Delta<-\delta_1-\delta_2\}f_{
\theta}(b^{(2)},a,e,\Delta)d\theta.\notag
\end{align}
Given product $j\in\{1,2\}$, let $-j$ denote the other product. We then
define $\tilde\Phi_l$ with $l=(0,0)$ as in the BLP example by letting $
D_{-jt}$ take a large negative value. For each $(x^{(2)},p,\delta)$, let
\begin{align}
\tilde\Phi_{(0,0)}(x_{j}^{(2)},p_{j},\delta_{j})\equiv
-\lim_{\delta_{-j}\to-\infty}\phi_{(0,0)}(x^{(2)},p,\delta),~j=1,2.
\label{eq:sumbund}
\end{align}
We then define $\Phi_{(0,0)}$ as in \eqref{eq:Phi}.\footnote{
In the BLP example, we invert a Radon transform only once. Hence $\Phi$ in
\eqref{eq:Phi} does not have any subscript. In Examples 2 and 3, we invert
Radon transforms multiple times, and to make this point clear we add
subscripts to $\Phi$ (e.g. $\Phi_{(0,0)}$ and $\Phi_{(1,1)}$).} Consider for
the moment $j=1$ in \eqref{eq:sumbund}. Then, $\Phi_{(0,0)}$ is related to
the joint density $f_{\vartheta_1}$ of $\vartheta_{i1t}\equiv(
\beta^{(2)}_{it},\alpha_{it},\epsilon_{i1t})$ through a Radon transform.
\footnote{
Since the bundle effect $\Delta_{it}$ does not appear in \eqref{eq:dem00},
one may only identify the joint density of the subvector $
(\beta^{(2)}_{it},\alpha_{it},\epsilon_{i1t})$ from the demand for bundle
(0,0).} Arguing as in \eqref{eq:sum1}, it is straightforward to show that $
\partial\Phi_{(0,0)}(w,u)/\partial u=\mathcal{R}[f_{\vartheta_1}](w,u)~$
with $w\equiv (x^{(2)}_1,p_1,1)/\|(x^{(2)}_1,p_1,1)\|$ and $u\equiv
\delta_1/\|(x^{(2)}_1,p_1,1)\|$. Hence, one may identify $f_{\vartheta_1}$
by inverting the Radon transform under Assumptions \ref{as:covariates} and
\ref{as:rc_blp} with $J=2$.

If the researcher is only interested in the distribution of $
(\beta^{(2)}_{it},\alpha_{it},\epsilon_{ijt})$ but not in the bundle effect,
the demand for $(0,0)$ is enough for recovering their density. However, $
\Delta_{it}$ is often of primary interest. The demand on (1,1) can be used
to recover its distribution by the following argument.

The demand for bundle (1,1) is given by
\begin{align}\label{eq:bundle11}
&\phi_{(1,1)}(x^{(2)},p,\delta) \notag \\
&=\int 1\{x^{(2)}_1{}^{\prime
}b^{(2)}+ap_1+e_1+\Delta>-\delta_1\}1\{x^{(2)}_2{}^{\prime
}b^{(2)}+ap_2+e_2+\Delta>-\delta_2\} \\
&\qquad \times1\{(x_1^{(2)}+x_2^{(2)})^{\prime
}b^{(2)}+a(p_1+p_2)+(e_1+e_2)+\Delta>-\delta_1-\delta_2\}f_{
\theta}(b^{(2)},a,e,\Delta)d\theta. \notag
\end{align}
Note that $\Delta_{it}$ can be viewed as an additional random coefficient on
the constant whose sign is fixed. Hence, the set of covariates includes a
constant. Again, conditioning on an event where $D_{-jt}$ takes a large
negative value and normalizing the arguments by the norm of $(
x^{(2)}_j,p_j,1)$ yield a function $\Phi_{(1,1)}$ that is related to the
density of $\eta_{ijt}\equiv(\beta^{(2)}_{it},\alpha_{it},\Delta_{it}+
\epsilon_{ijt})$ through the Radon transform in \eqref{eq:radon}. Note that
the last component of $\eta_j$ and $\vartheta_j$ differ only in the bundle
effect $\Delta_{it}$. Hence, if $\epsilon_{ijt}$ is independent of $
\Delta_{it}$ conditional on $(\beta^{(2)}_{it},\alpha_{it})$, the
distribution of $\Delta_{it}$ can be identified via deconvolution. For this,
let $\Psi_{\epsilon_{j}|(\beta^{(2)},\alpha)}$ denote the characteristic
function of $\epsilon_{ijt}$ conditional on $(\beta^{(2)}_{it},\alpha_{it})$
. We summarize these results in the following theorems.

\begin{theorem}
\label{thm:idsub} Suppose Assumptions \ref{as:index}-\ref{as:rc_blp}, \ref
{as:support_pure} and Condition \ref{cond:bundles} hold with $J=2$ and $
\theta_{it}=(\beta^{(2)}_{it},\alpha_{it},\Delta_{it},\epsilon_{i1t},
\epsilon_{i2t})$. Suppose the conditional distribution of $\epsilon_{ijt}$
given $(\beta^{(2)}_{it},\alpha_{it})$ is identical for $j=1,2$.

Then, (a) $f_{\vartheta_j},f_{\eta_j}$ are nonparametrically identified in
Example 2; (b) If, in addition, $\Delta_{it}\perp
\epsilon_{ijt}|(\beta^{(2)}_{it},\alpha_{it})$ and $\Psi_{\epsilon_{j}|(
\beta^{(2)},\alpha)}(t)\ne 0$ for almost all $t\in\mathbb{R}$ and for some $
j $, and $\epsilon_{ijt},j=1,2$ are independently distributed (across $j$)
conditional on $(\beta^{(2)}_{it},\alpha_{it})$, then $f_\theta$ is
nonparametrically identified in Example 2.
\end{theorem}

The identification of the distribution of the bundle effect requires the
characteristic function of $\epsilon_{ijt}$ to have isolated zeros (see e.g.
Devroye, 1989, Carrasco and Florens, 2010). This condition can be satisfied
by various distributions including the Type-I extreme value distribution and
normal distribution.

\begin{remark}\rm
\textrm{Note that the conditions of Theorem \ref{thm:idsub} do not impose
any sign restriction on $\Delta_{it}$. Hence, the two goods can be
substitutes $(\Delta_{it}<0)$ for some individuals and complements ($
\Delta_{it}>0$) for others. This feature, therefore, can be useful for
analyzing bundles of goods whose substitution pattern can significantly
differ across individuals (e.g. E-books and print books). }
\end{remark}

\begin{remark}\rm
\textrm{We note that the utility specification adopted in the pure
characteristics model can also be combined with the bundle choice and
multiple units of consumption studied in Section \ref{sec:mult} in the
online supplement. The identification of the random
coefficients can be achieved using arguments similar to the ones in Section
\ref{sec:PCM}
.}
\end{remark}


\section{Outlook}

This paper is concerned with the nonparametric identification of models of
market demand. It provides a general framework that nests several important
models, including the workhorse BLP model, and provides conditions under
which these models are point identified. Important conclusions include that
the assumption necessary to recover various objects differ; in particular,
it is easier to identify demand elasticities and more difficult to identify
the individual specific random coefficient densities. Moreover, the data
requirements are also shown to vary with the model considered. The
identification analysis is constructive, extends the classical nonparametric
BLP identification as analyzed in BH to other models, and opens up the way
for future research on sample counterpart estimation. A particularly
intriguing part hereby is the estimation of the demand elasticities, as the
moment condition is different from the one used in nonparametric IV.
Understanding the properties of these estimators, and evaluating their
usefulness in an application, is an open research question that we hope this
paper stimulates.

\ifx\undefined\BySame
\fi
\ifx\undefined \let\tmpsmall{\tmpsmall\sc  \fi


\begin{thebibliography}{99}
{\tmpsmall\sc
\harvarditem[Ackerberg, Benkard, Berry, and Pakes]{Ackerberg, Benkard, Berry,
  and Pakes}{2007}{Ackerberg:2007qe} \textsc{Ackerberg, D., C.~L. Benkard,
S.~Berry, and A.~Pakes}  (2007): ``Chapter 63 Econometric Tools for
Analyzing Market Outcomes,'' vol.  6, Part A of \emph{Handbook of
Econometrics\/}, pp. 4171 -- 4276. Elsevier. }



{\tmpsmall\sc \harvarditem[Allen and Rehbeck]{Allen and Rehbeck}{2020}{Allen:2020a} \textsc{Allen, R. and J. Rehbeck} (2020): ``Latent Complementarity in Bundles Models,''  \emph{SSRN preprint\/}, 3257028.}



{\tmpsmall\sc \harvarditem[Andrews]{Andrews}{2017}{Andrews:2017} \textsc{
Andrews, D.} (2017): ``Examples of L2-Complete and Boundedly-complete
Distributions,'' \emph{Journal of Econometrics}, 199(2) 213--220. }

{\tmpsmall\sc \harvarditem[Armstrong]{Armstrong}{2013}{Armstrong:2013} \textsc{
Armstrong, T.} (2016): ``Large Market Asymptotics for Differentiated Product Demand Estimators with Economic Models of Supply,'' \emph{Econometrica}, 84(5), 1961--1980.}


{\tmpsmall\sc \harvarditem[Beran, Feuerverger, and Hall]{Beran, Feuerverger, and Hall}{1996}{BeranFeuervergerHall1996}
\textsc{Beran, R., A.~Feuerverger, and P.~Hall} (1996): ``On nonparametric estimation of intercept and slope distributions in random coefficient regression,'' \emph{The Annals of
Statistics\/}, 24(6), 2569--2592. }

{\tmpsmall\sc \harvarditem[Beran and Hall]{Beran and Hall}{1992}{BeranHall1992AS}
\textsc{Beran, R., and P.~Hall} (1992): ``Estimating coefficient
distributions in random coefficient regressions,'' \emph{The Annals of
Statistics\/}, 20(4), 1970--1984. }

{\tmpsmall\sc
\harvarditem[Berry, Gandhi, and Haile]{Berry, Gandhi, and
  Haile}{2013}{Berry:2013jw} \textsc{Berry, S.~T., A.~Gandhi, and P.~A. Haile
} (2013):  ``Connected Substitutes and Invertibility of Demand,'' \emph{
Econometrica\/},  81(5), 2087--2111. }


{\tmpsmall\sc \harvarditem[Berry and Haile]{Berry and Haile}{2020}{Berry:2020}
\textsc{Berry, S.~T., and P.~A. Haile} (2020): ``Nonparametric Identification of Differentiated Products Demand Using Micro Data,'' NBER Working Paper 27704.}

{\tmpsmall\sc \harvarditem[Berry and Haile]{Berry and Haile}{2014}{BerryHaile2013}
\textsc{\leavevmode\rule[.5ex]{3em}{.5pt}\ {}} (2014): ``Identification in
Differentiated Products Markets  Using Market Level Data,''
\emph{Econometrica}, 82(5), 1749--1797. }

{\tmpsmall\sc
\harvarditem[Berry, Levinsohn, and Pakes]{Berry, Levinsohn, and
  Pakes}{1995}{Berry:1995rt} \textsc{Berry, S.~T., J.~A. Levinsohn, and
A.~Pakes} (1995):  ``Automobile Prices in Market Equilibrium,'' \emph{
Econometrica: Journal of the Econometric Society\/}, pp. 841--890. }

{\tmpsmall\sc
\harvarditem[Berry, Levinsohn, and Pakes]{Berry, Levinsohn, and
  Pakes}{2004}{Berry:2004wb} \textsc{\leavevmode\rule[.5ex]{3em}{.5pt}\ {}}
(2004): ``Differentiated Products Demand Systems from a  Combination of
Micro and Macro Data: The New Car Market,'' \emph{Journal of Political
Economy\/}, 112(1), 68--105. }

{\tmpsmall\sc \harvarditem[Berry and Pakes]{Berry and Pakes}{2007}{Berry:2007ys}
\textsc{Berry, S.~T., and A.~Pakes} (2007): ``The Pure  Characteristics
Demand Model,'' \emph{International Economic Review\/}, 48(4),  1193--1225. }


{\tmpsmall\sc \harvarditem[Carrasco and Florens]{Carrasco and
Florens}{2010}{Carrasco:2010uq} \textsc{Carrasco, M., and J.~Florens}
(2010): ``A spectral method  for deconvolving a density,'' \emph{Econometric
Theory\/}, 27(3), 546--581. }

{\tmpsmall\sc \harvarditem[Cramer and Wold]{Cram\'er and Wold}{1936}{CW:36} \textsc{
Cram\'er, H., and H.~Wold} (1936): ``Some Theorems on Distribution Functions'' \emph{Journal of the London Mathematical Society\/}, s1-11(4), 290--294.
}

{\tmpsmall\sc \harvarditem[Dub{\'e}, Fox, and Su]{Dub{\'e}, Fox, and
Su}{2012}{Dube:2012yj} \textsc{Dub{\'e}, J.-P., J.~T. Fox, and C.-L. Su}
(2012):  ``Improving the Numerical Performance of Static and Dynamic
Aggregate  Discrete Choice Random Coefficients Demand Estimation,'' \emph{
Econometrica\/},  80(5), 2231--2267. }

{\tmpsmall\sc
\harvarditem[Dunker]{Dunker}{2021}{Dunker:2021} \textsc{Dunker, F.} (2021): ``Adaptive estimation for some nonparametric instrumental variable models with full independence,'' \emph{Electronic Journal of Statistics}, forthcoming.}

{\tmpsmall\sc
\harvarditem[Dunker, Florens, Hohage, Johannes, and Mammen]{Dunker, Florens,
  Hohage, Johannes, and Mammen}{2014}{Dunker:2014km} \textsc{Dunker, F.,
J.-P. Florens, T.~Hohage, J.~Johannes, and E.~Mammen} (2014): ``Iterative
Estimation of Solutions to Noisy Nonlinear  Operator Equations in
Nonparametric Instrumental Regression,'' \emph{Journal of Econometrics\/},
178(3), 444 -- 455. }

{\tmpsmall\sc
\harvarditem[Dunker, Hoderlein, and Kaido]{Dunker, Hoderlein, and
  Kaido}{2013}{Dunker:2018xr} \textsc{Dunker, F., S.~Hoderlein, H.~Kaido, and R.~Sherman}
(2018): ``Nonparametric Identification of the Distribution of Random Coefficients in Binary Response Static Games of Complete Information,''
\emph{Journal of Econometrics}, 206(1), 83--102. }

{\tmpsmall\sc
\harvarditem[Dunker, Hoderlein, and Kaido2]{Dunker, Hoderlein, and
  Kaido}{2014}{Dunker:2014xr} \textsc{Dunker, F., S.~Hoderlein, and H.~Kaido}
(2014): ``Nonparametric Identification of Endogenous and Heterogeneous Aggregate demand models: Complements, Bundles and the Market Level,''
CEMMAP Working Paper No. CWP51/15. }

{\tmpsmall\sc
\harvarditem[Dunker, Mendoza, Reale]{Dunker, Mendoza, and Reale}{2021}{DMR:2021} \textsc{Dunker, F., E.~Mendoza, and M.~Reale} (2021): ``Regularized Maximum Likelihood Estimation for the Random Coefficients Model ,'' \emph{ArXiv preprints}, arXiv:2104.08402.}




{\tmpsmall\sc \harvarditem[Fox et al]{Fox et al}{2012}{Foxetal:2012}  \textsc{
Fox, J., K. ~Kim, S.~Ryan, and P.~Bajari} (2012): ``The random coefficients logit model is identified,'' \emph{Journal of Econometrics},
106, 204--212.
}


{\tmpsmall\sc \harvarditem[Fox and Lazzati]{Fox and Lazzati}{2017}{Fox:2017}
\textsc{Fox, J., and N.~Lazzati} (2017): ``A Note on Identification of  Discrete
Choice Models for Bundles and Binary Games,'' \emph{Quantitative Economics}, 8(3), 1021--1036. }

{\tmpsmall\sc \harvarditem[Fox and Gandhi]{Fox and Gandhi}{2016}{Fox:2016}
\textsc{Fox, J.~T., and A.~Gandhi} (2016): ``Nonparametric Identification and Estimation of Random Coefficients in Multinomial Choice Models,'' \emph{The RAND Journal of Economics}, 47, 118--139.}

{\tmpsmall\sc \harvarditem[Gale and Nikaido]{Gale and Nikaido}{1965}{GN:1965}
\textsc{Gale, D., and J.~Nikaido} (1965): ``The Jacobian matrix and global
univalence of mappings,'' \emph{Math. Ann.\/\/}, 159(2), 81--93. }

{\tmpsmall\sc
\harvarditem[Gautier and Hoderlein]{Gautier and
  Hoderlein}{2015}{Gautier:2015} \textsc{Gautier, E., and S.~Hoderlein}
(2015): ``A Triangular Treatment Effect Model with Random Coefficients in the Selection Equation,'' \emph{Arxiv
preprint} arXiv:1109.0362. }

{\tmpsmall\sc \harvarditem[Gautier and Kitamura]{Gautier and
Kitamura}{2013}{Gautier:2013ys} \textsc{Gautier, E., and Y.~Kitamura}
(2013): ``Nonparametric  Estimation in Random Coefficients Binary Choice
Models,''  \emph{Econometrica\/}, 81(2), 581--607. }

{\tmpsmall\sc \harvarditem[Gentzkow]{Gentzkow}{2007}{Gentzkow:2007sp} \textsc{
Gentzkow, M.~A.} (2007): ``Valuing New Goods in a Model with
Complementarity: Online Newspapers,'' \emph{American Economic Review\/},
97(3),  713--744. }

{\tmpsmall\sc \harvarditem[GR]{Gowrisankaran and Rysman}{2012}{}
\textsc{Gowrisankaran, G and M.~Rysman} (2012):
``Dynamics of Consumer Demand for New Durable Goods,'' \emph{Journal of Political Economy\/}, 120, 1173-1219.}

{\tmpsmall\sc \harvarditem[HMN]{Heckman, Matzkin, and Nesheim}{2010}{} \textsc{Heckman, J.~J., R. L. Matzkin and L.~Nesheim} (2010): ``Nonparametric Identification and Estimation of Nonadditive Hedonic Models,'' \emph{Econometrica\/}, 78(5), 1569--1591.}

{\tmpsmall\sc \harvarditem[Helgason]{Helgason}{1999}{Helgason:1999jw} \textsc{
Helgason, S.} (1999): \emph{The Radon Transform\/}. Birkhauser,
Boston-Basel-Berlin, 2nd edition edn. }

{\tmpsmall\sc
\harvarditem[Hoderlein, Holzmann, and Meister]{Hoderlein, Holzmann, and
  Meister}{2017}{Hoderlein:2017} \textsc{Hoderlein, S., H.~Holzmann, and
A.~Meister} (2017): `` The Triangular Model with Random
Coefficients,'' \emph{Journal of Econometrics}, 201(1), 144--169.}

{\tmpsmall\sc
\harvarditem[Hoderlein, Klemel{\"a}, and Mammen]{Hoderlein, Klemel{\"a}, and
  Mammen}{2010}{HoderleinKlemelaMammen2010ET} \textsc{Hoderlein, S.,
J.~Klemel{\"a}, and E.~Mammen} (2010):  ``Analyzing the Random Coefficient
Model Nonparametrically,''  \emph{Econometric Theory\/}, 26(03), 804--837. }

{\tmpsmall\sc
\harvarditem[Ichimura and Thompson]{Ichimura and
  Thompson}{1998}{IchimuraThompson1998JE} \textsc{Ichimura, H., and
T.~Thompson} (1998): ``Maximum Likelihood  Estimation of a Binary Choice
Model with Random Coefficients of Unknown  Distribution,'' \emph{Journal of
Econometrics\/}, 86(2), 269--295. }


{\tmpsmall\sc \harvarditem[Keane]{Keane}{1992}{Keane:1992ek} \textsc{Keane, M.~P.}
(1992): ``A note on identification in the multinomial  probit model,'' \emph{
Journal of Business \& Economic Statistics\/}, 10(2),  193--200. }

{\tmpsmall\sc \harvarditem[KT]{Khan and Tamer}{2010}{KhanTamer2010ECMA} \textsc{Khan, S., and
E.~Tamer} (2010): ``Irregular Identification, Support Conditions, and Inverse Weight Estimation,''
\emph{Econometrica\/}, 78(6), 2021--2042.}

{\tmpsmall\sc \harvarditem[Krantz and Parks]{Krantz and Parks}{2002}{Krantz:2002nq}
\textsc{Krantz, S.~G., and H.~R. Parks} (2002): \emph{The implicit function
theorem: history, theory, and applications\/}. Birkh{\"a}user, Boston,  USA.
}

{\tmpsmall\sc \harvarditem[LuShiTao]{Lu, Shi, and Tao}{2019}{LuShiTao:2019}
\textsc{Lu, Z., X. Shi, and J. Tao} (2019): ``Semi-Nonparametric Estimation of Random Coefficient Logit Model for Aggregate Demand,'' \emph{SSRN preprint\/}, 3503560.
}

{\tmpsmall\sc \harvarditem[Matzkin]{Matzkin}{2012}{Matzkin:2012bv} \textsc{
Matzkin, R.~L.} (2012): ``Identification in Nonparametric Limited  Dependent
Variable Models with Simultaneity and Unobserved Heterogeneity,''  \emph{
Journal of Econometrics\/}, 166(1), 106--115. }

{\tmpsmall\sc \harvarditem[McFadden]{McFadden}{1974}{McFADDEN:1974hb} \textsc{
McFadden, D.~L.} (1974): \textquotedblleft Conditional Logit Analysis of
Qualitative Choice Behavior,\textquotedblright\  In Zarembka, P., editor, \emph{Frontiers in
Econometrics\/}, Academic Press, pp. 105--142. }

{\tmpsmall\sc \harvarditem[McFadden]{McFadden}{1981}{McFADDEN:1981hb} \textsc{
McFadden, D.~L.} (1981): \textquotedblleft Structural Discrete Probability
Models Derived from Theories of Choice,\textquotedblright\ In Charles Manski
and Daniel McFadden, editors, \emph{Structural Analysis of Discrete Data and
Econometric Applications }, \ The MIT\ Press, Cambridge.}

{\tmpsmall\sc \harvarditem[Moon, Shum, Weidner]{Moon}{2018}{Moon:2018} \textsc{
Moon, H.~R., M. Shum, and M. Weidner} (2018):
``Estimation of Random Coefficients Logit Demand Models with Interactive Fixed Effects,'' \emph{Journal of Econometrics} 206(2), 613--644.}


{\tmpsmall\sc \harvarditem[Nevo]{Nevo}{2001}{Nevo:2001ng} \textsc{Nevo, A.}
(2001): ``Measuring Market Power in the Ready-to-Eat Cereal  Industry,''
\emph{Econometrica\/}, 69, 307--342. }

{\tmpsmall\sc \harvarditem[Newey and Powell]{Newey and Powell}{2003}{Newey:2003ya}
\textsc{Newey, W.~K., and J.~L. Powell} (2003): ``Instrumental  Variable
Estimation of Nonparametric Models,'' \emph{Econometrica\/}, 71(5),
1565--1578. }

{\tmpsmall\sc \harvarditem[Petrin]{Petrin}{2002}{Petrin:2002ys} \textsc{Petrin, A.}
(2002): ``Quantifying the Benefits of New Products: The  Case of the
Minivan,'' \emph{Journal of Political Economy\/}, 110(4), 705--729. }


\end{thebibliography}