EconBase
← Back to paper

Estimating Discrete Choice Demand Models with Sparse Market-Product Shocks

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

111,300 characters · 21 sections · 77 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating Discrete Choice Demand Models with Sparse Market-Product Shocks

abstractWe propose a new approach to estimating the random coefficient logit demand model for differentiated products when the vector of market-product level shocks is sparse. Assuming sparsity, we establish nonparametric identification of the distribution of random coefficients and demand shocks under mild conditions. Then we develop a Bayesian estimation procedure, which exploits the sparsity structure using shrinkage priors, to conduct inference about the model parameters and counterfactual quantities. Comparing to the standard BLP (berry1995automobile) method, our approach does not require demand inversion or instrumental variables (IVs), and thus provides a compelling alternative when IVs are not available or their validity is questionable. Monte Carlo simulations validate our theoretical findings and demonstrate the effectiveness of our approach, while empirical applications reveal evidence of sparse demand shocks in well-known datasets. \noindentKeywords: Demand Estimation, Sparsity, Bayesian Inference, Shrinkage Prior \noindentJEL Codes: C1, C3, L00, D1

Introduction

Since Daniel L. McFadden's seminal work \citep*{mcfadden2001economic}, discrete choice models have become a basic tool for understanding consumer demand for differentiated products in many empirical contexts. As an important advancement of the literature, the BLP framework (\citealp*{berry1994estimating}, \citealp*{berry1995automobile}) allows researchers to estimate a flexible demand model that incorporates consumer preference heterogeneity and addresses price endogeneity, using market-product level aggregate data on price, quantity, and other variables.

A prominent feature of the BLP framework is the inclusion of market-product level demand shocks (a.k.a.\ unobserved market-product characteristics), which are observed by economic agents but not by econometricians. The dependence of price (or other endogenous variables) on these demand shocks provides a natural way to model the price endogeneity problem. However, this modeling strategy introduces a challenging estimation problem for two main reasons: (1) demand shocks enter the demand system nonlinearly, and (2) a dimensionality problem arises, as the number of parameters to estimate (including the demand shocks) exceeds the number of equations in the demand system. To address these challenges, BLP proposes inverting the demand system to recover the demand shocks, which are then interacted with a set of instrumental variables (IVs) to construct moment conditions for GMM estimation. While effective in many cases, this IV-based approach raises practical concerns regarding the availability and validity of suitable instruments.

In this paper, we propose an alternative approach to estimating the BLP model that eliminates the need for IVs by leveraging a sparsity assumption on market-product-level demand shocks. Specifically, we assume that in some markets, the demand shocks for certain products take the same market-specific values. The sparsity assumption effectively reduces the number of unknown parameters, enabling identification directly from the constraints in the model without relying on IV-based restrictions. Under this assumption and mild regularity conditions, we demonstrate that the demand shocks can be identified as model parameters, alongside other parameters characterizing consumer preferences in the random coefficients logit demand model.

Our identification strategy, based on the sparsity assumption, naturally leads to a likelihood-based inference framework, in contrast to the traditional BLP method that relies on demand inversion and IV-based moment conditions. We develop a Bayesian shrinkage approach that incorporates the sparsity assumption to estimate model parameters, including the sparsity structure of demand shocks, as well as counterfactual quantities such as price elasticities.

To handle the potentially very high-dimensional space of sparsity patterns for the market-product demand shocks, we employ a type of shrinkage priors, a variable selection technique in Bayesian statistics. Shrinkage priors have their roots in high-dimensional statistics and machine learning literature and are connected to penalized likelihood estimators such as LASSO (Tibshirani1996lasso) in the sense that the posterior modes can be considered equivalent to these estimators (KyungGillGhoshCasella2010). Shrinkage priors have been used successfully in linear econometric models such as vector autoregression (VAR) (e.g.,\ giannone2015prior, GiannoneLenzaPrimiceri2021); see, e.g.,\ KorobilisShimizu2022 for a review of shrinkage priors and their applications to linear models in economics.

However, despite this growing body of work, their application in non-nonlinear structural models -- such as the random coefficients logit model in the BLP framework -- remains limited. This may stem from the lack of theoretical results on how introducing sparsity can facilitate identification in such non-linear settings. Our paper addresses this gap by first establishing how sparsity contributes to identification in discrete choice demand models, which in turn motivates the use of shrinkage priors as a practical device. The insight and method developed here may also be valuable for estimating other structural models.

The proposed approach is both conceptually straightforward -- returning to the likelihood framework for classic multinomial choice models -- and computationally efficient, offering a one-stop estimator for both preference parameters and the sparsity structure of demand shocks. It also avoids the computationally intensive demand inversion procedure required by the BLP estimator. As such, it provides a compelling alternative to the BLP approach, particularly when valid IVs are difficult to find or when researchers wish to evaluate the robustness of specific IV choices.

Simulation results show that when the sparsity assumption holds in the data generating process (DGP), our approach performs similarly to the BLP estimator with strong IVs and outperforms the BLP estimator with potentially weak IVs. This supports our theoretical results on identification. Additionally, we examine cases where the demand shocks are not strictly sparse in the DGP, and we find that our estimator still performs reasonably well in estimating the preference parameters, demonstrating robustness to mild misspecifications.

Depending on the empirical context, the sparsity assumption can have natural interpretations. For example, in supermarket scanner data applications where markets are defined by “store-week” pairs, the demand shocks may reflect unobserved promotion efforts at the store-week level for different products, usually identified by UPCs (Universal Product Code), after controlling for more aggregate fixed effects, such as brand, city, or quarter. In such cases, a store can only promote a selected subset of products in a given week due to limited shelf space (e.g., end-of-aisle displays), so the demand shocks for other products in the store-week will share the same value (either zero or the store-week level value). Similarly, in berry1995automobile's automotive market application, the demand shocks largely capture unobservable advertising efforts, which can vary across brands and models in different markets. Some brands or models may engage in active advertising campaigns in a specific market, while others may choose to maintain a more “standard level” of marketing effort. In addition, in some contexts, when viewing the demand shocks as unobserved product level preferences, the sparsity assumption also lends itself to intuitive interpretations. For example, in the automotive applications, the “go-to" or “standard” models would share similar values of the additive unobserved preferences compared to other models, after controlling for the observables.

We explore these interpretations by applying our approach to two empirical applications. First, we analyze a supermarket scanner dataset, focusing on the yogurt category, where we interpret the demand shocks as unobserved store-week-level promotion efforts. Our results demonstrate that the approach effectively captures the sparsity in promotion patterns across products, revealing interesting insights into consumer demand and store-level marketing strategies. Second, we revisit the automotive market data from berry1995automobile to assess the performance of our method in a well-documented market setting. Across both applications, we find empirical evidence of sparsity in demand shocks, underscoring the practical relevance of our identification strategy. Moreover, our estimation results are largely consistent with those from the standard BLP method, but our approach has the advantage of not requiring IVs. These findings highlight the value of our approach as a robust alternative to a traditional IV-based method.

Related Literature

As emphasized by BerryHaile2014ECMA, the identification and estimation of the BLP model heavily depend on the availability of valid instrumental variables (IVs). In practice, finding suitable IVs for a specific empirical application is often challenging and requires consideration of data structure and availability, economic theory, institutional knowledge, and other contextual factors. Consequently, the literature has proposed and employed a wide range of IVs, including cost shifters (berry1999voluntary, goldberg2001evolution), BLP IVs (berry1995automobile), Hausman IVs (hausman1994valuation, nevo2001measuring), optimal IVs (berry1999voluntary, reynaert2014improving), differential IVs (gandhi2019measuring), and time-series or panel data-based IVs (sweeting2013dynamic, jin2021flagship), among others. However, even with this variety of alternatives, practitioners often face difficulties in selecting appropriate IVs, especially when estimation results are highly sensitive to the choice of IVs. This sensitivity may arise from the weak IV problem, a common concern in empirical research that can also emerge theoretically under specific model assumptions \citep*{armstrong2016large}.

Our identification result is related to several findings in the literature on the identification of random coefficients in BLP and/or classic multinomial choice models, including, among others, fox2012random; fox2016nonparametric; lu2023semi; and dunker2023nonparametric. While these results are developed under different assumptions and with distinct arguments, they are not directly applicable to our setting, where the demand shocks are sparse and the number of both products and markets is growing.

MoonShumWeidner2018blp propose an approach to estimating the BLP model by modeling demand shocks as interactive fixed effects. While our paper shares a similar spirit of imposing structure on demand shocks, our identification and estimation strategies are fundamentally different. Their identification relies on additional exogenous variables and moment conditions, whereas ours depends solely on the sparsity condition described earlier. Moreover, their estimation strategy builds on least squares and minimum distance methods, while ours employs a Bayesian shrinkage approach. A related work by GillenMonteroMoonShum2019blp introduces a LASSO-type estimator to select from a large number of control variables in the BLP model. In contrast, our focus is on addressing the challenge of high-dimensional demand shocks. Moreover, our estimation strategy differs: their approach involves multiple steps of variable selection and requires post-selection inference, whereas ours is a Bayesian approach that delivers all results in a single pass. ByrneImaiJainSarafidis2022ivfree uses cost data to establish identification of the BLP model. The cost variables do not have to be instruments, meaning that they do not need to be orthogonal to the unobserved demand shocks. While we also aim to achieve identification and estimation without instruments, our assumptions and data requirement are different. They impose assumptions on the cost function while we assume sparsity structure on demand shocks. Importantly, we do not require cost-side data.

Previous papers have proposed Bayesian estimation procedures of demand models for aggregate data. The approach introduced by Jiang2009bayesian can be seen as a “Bayesian BLP,” where the likelihood is constructed via demand inversion. Our approach differs in that, while they treat the market-product shocks as econometric residuals as in BLP, we treat them as parameters. As a result, our method does not require demand inversion and would be more scalable with respect to the number of products and markets than their approach, which requires demand inversion at each Markov chain Monte Carlo (MCMC) iteration.

In contrast to this, other Bayesian approaches, such as those by YangChenAllenby2003bayesian and MusalemBradlowRaju2009BayesBLP, construct the likelihood by assigning an artificial set of consumer choices proportionally to the market shares. While this facilitates the estimation of the multinomial logit model for individual demand and the simulation of random coefficients, the sampling noise introduced at the data creation stage can potentially affect inference. Our approach avoids this issue by not requiring artificially assigned choices. Furthermore, as pointed out by berry2003comment, methods that assign artificial choices often lack a thorough discussion of identification and its relationship to prior restrictions. In contrast, our approach establishes a tight connection between identification under sparsity and estimation using shrinkage priors as a practical tool, effectively putting the identification argument into action.

Lastly, because our approach attempts to explore a large dimensional space of sparsity pattern of the market-product shocks, broadly speaking, this paper also contributes to the expanding literature on high-dimensional demand estimation (e.g. ChiongShum2019MS: random projection for aggregate demand model with many products; SmithAllenby2019JASA: random partitions of products; LoaizaNibbering2022ScalableProbit_JBES: high-dimensional probit models; Jiang2024high_MS: graphical lasso for flexible substitution patterns; WangIaria2024_JEEA: model of demand for bundles; Ershov2024RAND: estimation of complementarity with many products; and ChibShimizu2025scalable: scalable estimation of consideration set models).

The rest of the paper is organized as follows. Section (ref) introduces a sparsity condition and establishes the identification of the model. Section (ref) proposes a shrinkage-prior-based estimation method. Section (ref) investigates the performance of the proposed approach through Monte Carlo simulations. Section (ref) applies the proposed approach to two well-known real datasets and finds empirical evidence of sparsity in both. Section (ref) concludes.

Model and Identification

Model

We consider a stylized random coefficient logit demand model for aggregate data, in the spirit of berry1995automobile. There are \(T\) markets, indexed by \(t = 1, \dots, T\), each consisting of \(J_t + 1\) products, indexed by \(j = 0, 1, \dots, J_t\), and \(N_t\) consumers, indexed by \(i = 1, \dots, N_t\). The products indexed by \(j > 0\) are “inside goods,” and product \(0\) is the “outside option.”

Each consumer \(i\)'s utility from product \(j\) in market \(t\) is given by

equation[equation omitted — 98 chars of source]

where \(X_{jt} \in \mathbb{R}^{d_X}\) is a vector of observed market-product characteristics, \(\beta_i\) represents consumer-specific taste parameters (i.e., random coefficients), which are i.i.d. across consumers and follow the distribution \(f \in \mathcal{F}\), \(\xi_{jt}\) is the market-product level demand shock (a.k.a. unobserved characteristic), and \(\varepsilon_{ijt}\) is an i.i.d. idiosyncratic preference shock across \(i\), \(j\), and \(t\), following the standard Gumbel distribution. To normalize the level of the random utility, the product characteristics and demand shock of the outside option, \(X_{0t}, \xi_{0t}\), are set to zero.

As is typical in aggregate demand modeling, we allow certain variables in \(X_{jt}\), such as price, to be endogenous. This endogeneity arises because these variables may depend on \(\xi_{jt}\), which the firm can observe when setting prices or making other decisions, but econometricians cannot. We do not impose a specific supply-side model to characterize this dependence but will return to it later when discussing the identification assumptions.

In each market \(t\), each consumer chooses the product that maximizes their utility, and aggregating consumer choices gives the market share of each product \(j\) as follows:

equation[equation omitted — 228 chars of source]

where we denote the \(J_t\)-dimensional vector by \(\xi_t = (\xi_{1t}, \dots, \xi_{J_t t})^{\top}\). Our setup encompasses the typical specification in most empirical applications, where some coefficients in \(X_{jt}\) are fixed (i.e., these random coefficients follow degenerate distributions). We will consider these special cases in the Monte Carlo simulations and empirical applications.

The observed market share of product \(j\) in market \(t\) is \(s_{jt}=(1/N_t)\sum_{i=1}^{N_{t}}y_{ijt}\), where \(y_{ijt}\) is an indicator that equals to 1 if \(i\) chooses product \(j\) in market \(t\) and 0 otherwise. By definition, the market share vector \(\left(s_{0t},...,s_{J_{t}t}\right)\in\Delta^{J_{t}}\), where \(\Delta^{J_{t}}\) is the standard \(J_t\)-simplex. The likelihood function of the observed choices is

equation[equation omitted — 278 chars of source]

where \(q_{jt}=\sum_{i=1}^{N_{t}}y_{ijt}\) is the total quantity of product \(j\) in market \(t\). The aggregation across consumers (second equality) is due to the fact that the aggregate data do not contain individual-level attributes, e.g., demographics. When individual-level data are available, our approach can be modified easily to incorporate this information (such an extension is available upon request).

Without additional restrictions, estimating \(\left(f,\xi_{1},...,\xi_{T}\right)\) based on the likelihood function ((ref)) is infeasible, as there are only \(\sum_{t=1}^{T}J_{t}\) linearly independent first-order conditions while the number of unknown parameters is \(\sum_{t=1}^{T}J_{t}+\dim\left(f\right)\), where \(\dim\left(f\right)\) denotes the dimensionality of \(f\). In particular, these conditions effectively constitute the demand system

equation[equation omitted — 97 chars of source]

and it is evident that the system is underidentified.

In the following, we first review the standard BLP approach to addressing this dimensionality problem and then introduce our new strategy to resolve it.

The BLP Approach

The standard BLP approach to addressing the identification problem consists of two main components. First, for a given \(f\), the demand system in ((ref)) is inverted (invertibility is established in \citet*{berry1994estimating} and \citet*{Berry2013invertibility}) to obtain

equation[equation omitted — 73 chars of source]

Next, the \(\xi_{jt}\)'s are treated as econometric residuals, satisfying the following conditional moment restrictions:

equation[equation omitted — 148 chars of source]

where \(Z_{jt}\) is a vector of instrumental variables (IVs). If the IVs provide sufficient variation, then the distribution \(f\) is identified and can be estimated via GMM.

Because there are endogenous product characteristics, such as price, and market shares in the moment conditions, the IVs \(Z_{jt}\) must include exogenous variables that are excluded from \(X_{jt}\) (see \citet*{BerryHaile2014ECMA}). However, finding and constructing valid IVs remains a persistent challenge in empirical applications, often requiring careful consideration of data structure and economic context. While the literature offers various strategies to address this, there is still considerable debate over their effectiveness and applicability in different settings.

Sparsity Assumption on Demand Shocks

We propose an alternative approach to the identification problem based on the demand system ((ref)). Instead of imposing conditional moment restrictions (or other distributional assumptions) on \(\xi_{jt}\)'s, like ((ref)), we treat \(\xi_{t}\)'s as parameters to be estimated and assume they exhibit a sparsity structure: for some market \(t\), a sub-vector of \(\xi_{t}\) shares the same value, as formally stated in Assumption (ref). In the background, we consider a population of many markets, with each market containing a population of many products. By many, we mean a countable infinity.

assumption[Sparsity] There exists an infinite subset of markets \( \mathcal{S} \) such that, for each \( t \in \mathcal{S} \), there exists an infinite subset of products \( \mathcal{K}_t \) satisfying \(\xi_{jt}=\xi_{kt}\) for any \(j,k\in \mathcal{K}_{t}\).
commentIn the background, we consider a framework with both a large number of products per market \((J_t \to \infty)\) and a large number of markets \((T \to \infty)\). \begin{assumption} [Sparsity] There exist a set of markets \(\mathcal{S}\subset\left\{ 1,...,T\right\}\) and a set of products \( \mathcal{K}_{t}\subset\left\{ 1,...,J_{t}\right\} \) for each \(t\in \mathcal{S}\), such that \(\xi_{jt}=\xi_{kt}\) for any \(j,k\in \mathcal{K}_{t}\). Furthermore, \(\left|\mathcal{S}\right|\rightarrow\infty\) and \(\left|\mathcal{K}_{t}\right|\rightarrow\infty\) for each \(t\in \mathcal{S}\). \end{assumption}

Assumption (ref) basically says that for any market \(t\) in the subset \(\mathcal{S}\), the \(\xi_{jt}\)'s for products \(j\) in the sparse set \(\mathcal{K}_{t}\) have the same value, which is denoted as \(\nu_t\in\mathbf{R}\), while those not in \(\mathcal{K}_{t}\) are unrestricted. So the number of unknowns in \(\xi_t\) is reduced by \(\left|\mathcal{K}_{t}\right|-1\) (from \(J_t\)), which in turn implies that the number of unknowns in the demand system ((ref)) decreases by \(\sum_{t\in\mathcal{S}}\left(\left|\mathcal{K}_{t}\right|-1\right)\). Intuitively, the reduction in the number of unknown parameters can circumvent the dimensionality problem and restore the identification of \(\left(f,\xi_{1},...,\xi_{T}\right)\), where the sparsity condition is imposed on the vector \(\xi_t\) for those markets \(t\) in \(\mathcal{S}\).

To ensure a sufficient reduction in the number of parameters, Assumption (ref) requires that sparsity occurs in an infinite number of places in the population: (1) for each market \( t \in \mathcal{S} \), the number of products sharing the same value of \( \xi_{jt} \) is infinite; and (2) the number of such sparse markets is infinite. These conditions are sufficient for the nonparametric identification of \( (f, \xi_1, \dots, \xi_T) \), and they can be relaxed if \( f \) is parametrically specified (e.g., Gaussian).

Importantly, Assumption (ref) does not require sparsity to be widespread across the population of markets. The number of sparse markets \( |\mathcal{S}| \) increases with the total number of markets, but may grow arbitrarily slowly. That is, \( \mathcal{S} \) can be small relative to the full set of markets. Likewise, for each \( t \in \mathcal{S} \), the number of products with identical unobservables, \( |\mathcal{K}_t| \), may grow slowly relative to \( J_t \). The complement of these sparse sets - the non-sparse markets and products - can remain large or even constitute the majority of the data.

Assumption (ref) has important implications for the canonical price (and/or other product characteristics) endogeneity problem. To see this, let us consider a concrete example.

exampleSuppose one element of \( X_{jt} \) is price \( P_{jt} \), which is determined by the following linear model: \[ P_{jt} = Z_{jt}^{\top} \rho + \omega_{jt}, \] where \( Z_{jt} \) is a vector of IVs, \( \rho \) is a vector of parameters, and \( \omega_{jt} \) is the error term in the pricing equation. To capture potential price endogeneity, assume that \( \omega_{jt} \) is related to the unobserved demand shock \( \xi_{jt} \) via \[ \omega_{jt} = \phi \xi_{jt} + \epsilon_{jt}, \] where \( \phi > 0 \) and \( \epsilon_{jt} \) is an idiosyncratic pricing shock that is independent of \( \xi_{jt} \). This structure implies that price is endogenous in the demand equation, as it is correlated with \( \xi_{jt} \) through the pricing residual \( \omega_{jt} \). Now suppose the unobserved demand shock \( \xi_{jt} \) follows a sparse structure: \[ \xi_{jt} = \begin{cases} \nu_t, & \text{with probability } \varphi, \\ \xi_{jt}^*, & \text{with probability } 1 - \varphi, \end{cases} \] where \( \nu_t \) and \( \xi_{jt}^* \) are mean-zero random variables. The parameter \( \varphi \in [0,1] \) governs the degree of sparsity: with probability \( \varphi \), all products in market \( t \) share the same unobservable \( \nu_t \); with probability \( 1 - \varphi \), product-level unobservables \( \xi_{jt}^* \) apply. The conditional covariance between price and the unobserved demand shock (given \( Z_t \)) is, suppressing \(Z_t\) for notational simplicity, \[ \text{Cov}(P_{jt}, \xi_{jt}) = \text{Cov}(\omega_{jt}, \xi_{jt}) = \phi \cdot \text{Var}(\xi_{jt}), \] where by the zero-mean assumptions \[ \text{Var}(\xi_{jt}) = \varphi \, \text{Var}(\nu_t) + (1 - \varphi) \, \text{Var}(\xi_{jt}^*). \] For the moment, suppose \( \text{Var}(\nu_t) \) is large relative to \( \text{Var}(\xi_{jt}^*) \). Then, increasing the degree of sparsity \(\varphi\) leads to a higher correlation between price and unobserved demand shocks, exacerbating the severity of endogeneity. Conversely, if \( \text{Var}(\xi_{jt}^*) \) is more variable, a higher degree of sparsity can decrease endogeneity. Thus, the sparsity structure in \( \xi_{jt} \) has implications for the nature and extent of price endogeneity. However, increased sparsity does not necessarily reduce endogeneity - it depends on the relative magnitudes of \( \text{Var}(\nu_t) \) and \( \text{Var}(\xi_{jt}^*) \).
comment\begin{example} Suppose one element of \(X_{jt}\) is price \(P_{jt}\) and it is determined by a linear function \[ P_{jt}=Z_{jt}^{\top}\rho+\xi_{jt}, \] where \(Z_{jt}\) is the vector of IVs and \(\rho\) is a vector of parameters. Also, for any market \(t\), \(\xi_{jt}\) is generated as \[ \xi_{jt}=\begin{cases} \nu_{t}, & \text{with prob }\varphi\\ \xi_{jt}^{*}, & \text{with prob }1-\varphi, \end{cases} \] where \(\nu_{t}\) and \(\xi_{jt}^{*}\) are two mean-zero random variables that are independent of each other. So the probability \(\varphi\) captures the degree of sparsity in \(\xi\)'s. In this case, the price endogeneity problem can be measured by the conditional (given \(Z_t\)) covariance of \(P_{jt}\) and \(\xi_{jt}\). Suppressing the conditioning variables \(Z_t\) for notational simplicity, we have \begin{equation} Cov\left(P_{jt},\xi_{jt}\right)=Var\left(\xi_{jt}\right)=E\left(\xi_{jt}^{2}\right)=\varphiVar\left(\nu_{t}\right)+\left(1-\varphi\right)Var\left(\xi_{jt}^{*}\right). \end{equation} When \(\varphi\) is small, i.e., there is limited sparsity in \(\xi\)'s, the second term in ((ref)) dominates so price endogeneity is captured by the variance of \(\xi_{jt}^{*}\). As \(\varphi\) gets larger, i.e., more sparsity in \(\xi\)'s, the variance of \(\nu_{t}\) becomes more important in determining the severity of the endogeneity problem. Thus, if the across market variation of \(\nu_{t}\) is small comparing to the across market-product variation of \(\xi_{jt}^{*}\), then more sparsity implies less endogeneity. Of course, the converse is also true. So the sparsity assumption has certain implications on the price endogeneity problem, but does not necessarily reduce its severity. \end{example}

The sparsity assumption and the conditional mean restriction in Equation (ref) represent two distinct approaches to addressing endogeneity, and neither is strictly stronger nor weaker than the other. The standard moment condition (ref) places no structural restriction on the form of endogeneity, but it requires valid instruments that are excluded from the unobservable \( \xi_{jt} \). In contrast, the sparsity assumption restricts the dependence structure of \( \xi_{jt} \), which can, under suitable conditions, enable identification without instruments. However, this comes at the cost of imposing a specific sparsity structure on the unobservables, which may or may not hold in certain applications.

One particularly interesting feature of the sparsity assumption is that it does not impose any restrictions on the \(\xi_{jt}\)'s in the non-sparse set. In particular, these \(\xi_{jt}\)'s can either be realizations of any continuous or discrete distributions. Also, they can arbitrarily depend on \(X_{jt}\) and \(Z_{jt}\) (these IVs are not valid in this case). On the contrary, typical statistical assumptions on \(\xi_{jt}\)'s, such as ((ref)) or other distributional assumptions, imply much stronger restrictions on the non-sparse set.

Next, we shall explore how Assumption (ref) can help identify the model. The next subsection will establish the main identification result that shows that both \(\xi_t\)'s and \(f\) are identified by the demand system ((ref)) under the sparsity assumption and some other conditions. Before diving into the formal result, it is instructive to illustrate its key idea via a nested-logit example.

example[Nested-Logit with Sparse \(\xi\)] Consider a nested-logit model (as in berry1994estimating) of consumer demand for 4 inside goods and an outside option. For convenience, we focus on a single market and omit the subscript \(t\). The products are grouped into mutually exclusive nests, with the outside option \(0\) being the only member in its own nest. The utility function of consumer \(i\) can be written as \[u_{ij}=\begin{cases} \beta X_{j}+\xi_{j}+\zeta_{ig(j)}+\left(1-\lambda\right)\epsilon_{ij}, & j=1,2,3,4\\ \zeta_{i0}+\left(1-\lambda\right)\epsilon_{i0}, & j=0, \end{cases}\] where \(X_j\in \textbf{R}\) is an observed product characteristic, \(\zeta_{ig(j)}\) is a random coefficient following a specific distribution \citep*{cardell1997variance}, \(g(j)\) labels the nest of product \(j\), \(\lambda\) is the “nesting parameter,” and \(\epsilon_{ij}\) is the “logit error” following the standard Gumbel distribution. Suppose we know the sparsity pattern is \(\xi_2=\xi_3=\xi_4=\nu \). Then the parameters to be identified are \(\beta,\lambda, \xi_1\), and \(\nu\). Given the close-form inversion of nested-logit model, we obtain the following linear system: \begin{align*} \xi_{1}+\beta X_{1}+\lambda\log\left(\bar{s}_{1|g(1)}\right) & =\log\left(\frac{s_{1}}{s_{0}}\right)\\ \nu+\beta X_{2}+\lambda\log\left(\bar{s}_{2|g(2)}\right) & =\log\left(\frac{s_{2}}{s_{0}}\right)\\ \nu+\beta X_{3}+\lambda\log\left(\bar{s}_{3|g(3)}\right) & =\log\left(\frac{s_{3}}{s_{0}}\right)\\ \nu+\beta X_{4}+\lambda\log\left(\bar{s}_{4|g(4)}\right) & =\log\left(\frac{s_{4}}{s_{0}}\right), \end{align*} where \(s_j\) is the market share of product \(j\) and \(\bar{s}_{j|g(j)}\) denotes the within-group share of product \(j\) in its nest \(g(j)\). There are 4 equations and 4 unknowns, so the parameters are determined by \begin{equation} \begin{pmatrix}\xi_{1}\\ \nu\\ \beta\\ \lambda \end{pmatrix}=\begin{bmatrix}1 & 0 & X_{1} & \log\left(\bar{s}_{1|g(1)}\right)\\ 0 & 1 & X_{2} & \log\left(\bar{s}_{2|g(2)}\right)\\ 0 & 1 & X_{3} & \log\left(\bar{s}_{3|g(3)}\right)\\ 0 & 1 & X_{4} & \log\left(\bar{s}_{4|g(4)}\right) \end{bmatrix}^{-1}\begin{bmatrix}\log\left(\frac{s_{1}}{s_{0}}\right)\\ \log\left(\frac{s_{2}}{s_{0}}\right)\\ \log\left(\frac{s_{3}}{s_{0}}\right)\\ \log\left(\frac{s_{4}}{s_{0}}\right) \end{bmatrix}. \end{equation} We can see that the identification condition in this special case boils down to the invertibility of the matrix in ((ref)). The invertibility requires that the vector \(X\) cannot be collinear with the indicator variables for the sparse set (the first two columns in the matrix), which automatically holds when \(X\) is continuous.

This example highlights the key role of the sparsity assumption in identification: it reduces the number of unknown parameters from 6 (\(\beta,\lambda\) and all the \(\xi\)'s) down to 4 so we have sufficient number of equations. Based on this insight, we can expect that, to identify a more complicated model with more parameters, we will need data on more products and/or markets, as well as a sufficient degree of sparsity. Also, the nested-logit model, which is a special case of the general random coefficient logit model ((ref)), has a close-form inversion in \(\xi\)'s, so we can derive an explicit solution for the parameters. However, for the general model that does not have close-form inversion, establishing identification requires additional technical conditions and arguments.

Non-parametric Identification with Sparse Demand Shocks

In this subsection, we establish the formal non-parametric identification result, Theorem (ref), under the sparsity assumption. Given Assumption (ref), without loss of generality, suppose \(\mathcal{K}_{t} = \left\{K_{t}+1,...,J_t\right\} \), i.e., the \(\xi_{jt}\)'s for the last \(J_t - K_t\) products takes the same value \(\nu_t\). Note that \(\mathcal{K}_{t} = \varnothing\) for \(t\notin\mathcal{S}\). Then for each market \(t\), let \[ \Xi_{t}=\left\{ \xi_{t}\in\mathbf{R}^{J_{t}}:\xi_{jt}=\nu_{t}\text{ for any }j\in\mathcal{K}_{t}\right\} \] denote the space of \(\xi_{t}\)'s restricted by Assumption (ref).

For a given \(f\in\mathcal{F}\) and any market \(t\), there are \(K_t+1\) unknowns, \(\left(\xi_{1t},...,\xi_{K_{t}t}\right)\) and \(\nu_t\), and \(J_t\) equations in the demand system ((ref)). So, intuitively we may only need the first \(K_t + 1\) equations, i.e.,

equation[equation omitted — 102 chars of source]

to solve for the vector \(\left(\xi_{1t},...,\xi_{K_{t}t}, \nu_t \right)\). Lemma (ref) confirms that this is the case so the demand system ((ref)), and hence ((ref)), is invertible in \(\xi_{t}\in\Xi_{t}\) for any \(f\in\mathcal{F}\), which is a slight modification of the invertibility result in berry1994estimating.

lemmaSuppose \(s_{jt}>0\) for any \(j,t\). Then for any \(f\in\mathcal{F}\) and any \(t\), there is a unique \(\xi_{t}\in\Xi_{t}\) that satisfies the demand system \(s_{jt}=\sigma_{jt}\left(\xi_{t},f\right), \: j=1,...,J_t\). Moreover, for any \(t\), \(\xi_{jt}=\tilde{\sigma}_{jt}^{-1}\left(\tilde{s}_{t},f\right)\) for any \(j=1,...,K_{t}+1\), where \(\tilde{s}_{t}=\left(s_{1t},...,s_{K_{t}+1,t}\right)\) and \(\left[\tilde{\sigma}_{1t}^{-1}\left(\tilde{s}_{t},f\right),...,\tilde{\sigma}_{K_{t}+1,t}^{-1}\left(\tilde{s}_{t},f\right)\right]\) denotes the solution of the system ((ref)).
proofSee Appendix (ref).

Lemma (ref) gives us the inversion of the subsystem ((ref)), i.e., \(\xi_{jt}=\tilde{\sigma}_{jt}^{-1}\left(\tilde{s}_{t},f\right)\), \(j=1,...,K_{t}+1\). We can substitute these inverse demand functions into the last \(J_t - K_t - 1\) equations of ((ref)) to obtain

equation[equation omitted — 163 chars of source]

where \(\tilde{\sigma}_{t}^{-1}\left(\tilde{s}_{t},f\right)=\left[\sigma_{1t}^{-1}\left(\tilde{s}_{t},f\right),...,\sigma_{K_{t}+1,t}^{-1}\left(\tilde{s}_{t},f\right)\right]^{\top}\). Note that the only unknown object in ((ref)) is \(f\), so the identification problem becomes whether \(f\) is uniquely determined by ((ref)). To establish identification, we need to introduce additional assumptions.

assumptionThe random vector \(X_{jt}\) is independent across \(j,t\), and has continuous and full support in \(\mathbf{R}^{d_{X}}\).

Assumption (ref) rules out the cases where \(X_{jt}\) has discrete or bounded support. This is not surprising if we want to identify a continuous \(f\) nonparametrically. Similar continuous support assumptions are also imposed in related studies: see, among others, the Assumption 2 of fox2012random and the Assumption 7 of lu2023semi. The continuous support assumption can be relaxed when a sub-vector of the coefficients on \(X_{jt}\) are fixed and/or \(f\) is parameterized by a finite number of parameters, as commonly estimated models in practice.

Also, Assumption (ref) requires that \(X_{jt}\) has full support, so that we can establish identification based on the “identification at infinity” argument, as in lewbel2000semiparametric and khan2009inference, among others.

assumptionFor any \(f\in\mathcal{F}\), \(\sup_{x\in\mathcal{B}}\int\exp\left(x^{\top}\beta\right)f\left(\beta\right)d\beta<\infty\) for any bounded open \(\mathbf{R}^{d_{X}}\)-ball \(\mathcal{B}\).

Assumption (ref) is a regularity condition restricting the tail of \(f\) to be exponential or subexponential, which is satisfied by Gaussian distributions, for example. This assumption is necessary for our identification argument based on the uniqueness of Laplace transform; it can be viewed as an alternative restriction on the shape of \(f\) to the bounded support assumption imposed in fox2012random.

As we shall show in the proof of Theorem (ref), Assumption (ref), (ref), and (ref) imply that \(f\) can be nonparametrically identified by the subsystem (ref). Combining this result with Lemma (ref), we can conclude that both \(\xi_{jt}\)'s and \(f\) are identified by the demand system ((ref)), as stated in Theorem (ref).

theoremIf the conditions of Lemma (ref), Assumptions (ref), (ref) and (ref) hold, then \(\xi_{t}\in\Xi_{t}\) for all \(t=1,...,T\) and \(f\in\mathcal{F}\) are identified by the demand system ((ref)).
proofSee Appendix (ref).
remarkComparing with the canonical identification result for BLP in BerryHaile2014ECMA and dunker2023nonparametric, which relies on demand inversion and IVs, Theorem (ref) exploits the sparsity structure in \(\xi\)'s but does not require IVs. Note that demand inversion is still used in the proof of Theorem (ref): we employ the inversion of the subsystem ((ref)) to establish the identification of \(\xi\)'s for a given \(f\). However, in our context, the inversion is only used as a theoretical device in the proof; as we shall see later, our estimation procedure does not require explicitly computing the demand inversion. This stands in contrast with the BLP estimation strategy, where demand inversion is computed repeatably in the estimation procedure.

Bayesian Shrinkage Approach to Estimation

The identification result established by Theorem (ref) is conditional on the sparsity structure defined by \(\Xi_{t}\)'s. If the sparsity structure were known, we could simply implement the maximum likelihood estimation (MLE) using ((ref)) with the restrictions on \(\xi_t\)'s defined by \(\Xi_{t}\)'s. However, in practice, we typically do not know the sparsity structure ex-ante, just as the situation with the classical high-dimensional regression with many predictors. In this section, we shall borrow insights from the high-dimensional Bayesian statistics literature to design an inference procedure that can both uncover the latent sparsity structure and deliver estimates of the parameters of interests.

The latent sparsity structure is a high-dimensional object comprised of the following components:

enumerate$\mathcal{S} \subset \{1,\ldots,T\}$, the set of “sparse” markets exhibiting sparsity in \(\xi\)'s; • $\mathcal{K}_t \subset \{1,\ldots,J_t\}$ for each $t \in \mathcal{S}$, the set of “sparse” products such that $\xi_{jt}=\nu_t$ for any $j \in \mathcal{K}_t$.

For each market $t$, there are essentially $2^{J_t}$ possible configurations of $\mathcal{K}_t$\footnote{For convenience, if $\mathcal{K}_t= \varnothing $, i.e. all the $\xi_{jt}$'s are unrestricted, we think of $t$ as not belonging to $\mathcal{S}$. Hence, the estimation of $\mathcal{K}_t$ subsumes that of both (A) and (B).}. For example, in the automobile market application of \citet*{berry1995automobile}, there are on average 110 products in each of the 20 markets, indicating a very large dimensionality of the parameter space to be explored.

The problem of finding such sparsity structure is akin to the variable selection problem in high-dimensional linear regression with many predictors for which frequentist penalized likelihood methods such as LASSO (Tibshirani1996lasso) and Bayesian shrinkage prior methods are widely used. In this paper, we propose a stochastic search method based on shrinkage priors for the ease of uncertainty quantification (KyungGillGhoshCasella2010; Womack2014inference; PorwalReftery2022comparing). For recent applications of shrinkage priors in econometrics, primarily on linear models, see, e.g.,\ GiannoneLenzaPrimiceri2021; KoopKorobilis2023IER; and SmithGriffin2023shrinkage.

Conceptually, we ex-ante consider all the $2^{J_t}$ possible configurations of $\mathcal{K}_t$, including those “dense” ones where the majority of the $\xi_{jt}$'s are unrestricted. The data then informs us which of the $\xi_{jt}$'s “significantly” deviate from $\nu_t$, just as how the penalized likelihood estimator selects non-zero slopes in linear regression models with many predictors.

In particular, we employ a type of spike-and-slab priors (MitchellBeauchampll1988; GeorgeMcCulloch1993; GeorgeMcCulloch1997; IshwaranRao2005AoS; NarisettyHe2014AoS; RovckovaGeorge2018JASA). With the spike-and-slab priors, one can easily obtain a probabilistic statement about sparsity (i.e.,\ $\xi_{jt} = \nu_t$) in contrast to the penalized likelihood approaches, which are based on constrained optimization problems, and other Bayesian alternatives.

The likelihood

We introduce our estimation procedure with a commonly used parametric specification of the distribution of random coefficients (\(f\)), for the ease of exposition. A non-parametric extension using sieve approximation, as in lu2023semi and WANG2023325, is possible, given the general non-parametric identification result of Theorem (ref); however, it is beyond the scope of the current paper, so we leave it for future research.

Specifically, let the random coefficients follow a mutually independent joint normal distribution: \[ \beta \sim N_{d_X}(\bar{\beta},\Sigma), \] where $\Sigma=\text{diag}(\sigma^2_{1},\ldots,\sigma^2_{d_X})$. Again we assume independence for the ease of exposition: it is straightforward to include non-zero off diagonal elements in $\Sigma$\footnote{berry1995automobile also imposes independence.}. Note that this formulation nests the common case where only some of the $d_X$ covariates are assigned random coefficients. For example, with $d_X=3$ and if only the first variable has random coefficients, we have $\Sigma=\text{diag}(\sigma^2_1,0,0)$.

For convenience, we reparametrize the non-negative elements in $\Sigma$ as in Jiang2009bayesian. First, we decompose $\Sigma=RR'$, where $R=\text{diag}(\sigma_{1},\ldots,\sigma_{d_X})$. Then let $r=(r_1,\ldots,r_{d_X})'$ be the log standard deviations of the random coefficients, i.e., $r_{k}=\log(\sigma_k)$. Consequently, $R=\text{diag}(e^{r_1},\ldots,e^{r_{d_X}})$. Next, for the unobserved product characteristics, let

equation[equation omitted — 65 chars of source]

where $\bar{\xi}_t$ is the market-level shock in $t$ and $\eta_{jt}$ is the market-product $(j,t)$- specific deviation from $\bar{\xi}_t$. Note that a sparse vector \(\xi_t\) (as in Theorem (ref)) is a special case of this formulation with \(\bar{\xi}_t \equiv \nu_t\). If product $j$ belongs to the sparse set $\mathcal{K}_t$, then $\eta_{jt}=0$ and hence $\xi_{jt}=\bar{\xi}_t$; otherwise, $\eta_{jt}$ is unrestricted so that $\xi_{jt}$ can freely deviate from $\bar{\xi}_t$.

Given the above parameterization, the utility function (ref) can be rewritten as \[ u_{ijt}=\delta_{jt} + \mu_{ijt} +\varepsilon_{ijt}, \] where $\delta_{jt}=X_{jt}^{\top}\bar{\beta}+\xi_{jt}$ and $\mu_{ijt}=X_{jt}^{\top}R v_i$. The $d_X$-dimensional vector $v_i$ is i.i.d. and follows the product of $d_X$ independent standard normal distributions. Consequently, the predicted market share is

equation[equation omitted — 246 chars of source]

where $\phi\left(\cdot\vert 0, I \right)$ is the $d_X$-dimensional standard normal density. We approximate the integral based on $R_0$ i.i.d.\ draws of $v_i$ from the normal distribution. The likelihood is defined as

equation[equation omitted — 241 chars of source]

where $\eta_t=(\eta_{1t},\ldots,\eta_{J_tt})^{\top}$, $q=\{q_1,\ldots,q_T \}$ and $q_t =( q_{1t},\ldots,q_{J_tt})^{\top}$.

Note that we could define $\bar{\xi}_t$'s as part of $\bar{\beta}$ by modifying the covariates appropriately, but especially when $T$ is large, it is computationally more efficient to separately update the market-specific intercepts from the slopes. We therefore treat them separately in what follows.

Prior

As previously discussed, the sparsity structure of \(\xi\)'s is latent and needs to be learned from data. To this end, we employ a spike-and-slab prior on the deviation term $\eta_{jt}$'s in (ref). The (continuous) spike-and-slab prior (GeorgeMcCulloch1993; GeorgeMcCulloch1997) is a popular method for stochastic search variable selection. Other types of priors can be also used in our framework, but the unique feature of this class of priors is the ease of interpretation of the estimated sparsity structure, as we will see shortly.

Specifically, we define the prior on the deviations of the unobserved market-product shocks as

equation[equation omitted — 117 chars of source]

independently over $j$ and $t$, where $0<\tau^2_0\ll\tau^2_1$ are the prior variances in the two component mixture of normals with $\tau^2_0$ taking a very small value and $\tau^2_1$ a large value\footnote{Note that the original spike-and-slab prior has the dirac-delta function in place of $N(0,\tau^2_0)$ (MitchellBeauchampll1988). While the formulation (ref) is an approximation to the original spike-and-slab prior, the mixture of two normals formulation has become popular due to its computational simplicity. Note that, by choosing $\tau_0^2$ small enough, the spike component can made arbitrarily close to the dirac-delta function.}. The binary indicator $\gamma_{jt}$ equals to 0 if the $\eta_{jt}$ belongs to the “spike” component and $1$ if it is in the “slab” component. Intuitively, if $\gamma_{jt}=0$, $\eta_{jt}$ is shrunk toward zero (i.e.,\ $\xi_{jt}$ is shrunk toward $\bar{\xi}_t$ and hence is sparse). On the other hand, when $\gamma_{jt}=1$, it is unrestricted (i.e.,\ $\xi_{jt}$ freely deviates from $\bar{\xi}_t$). The posterior mean of $\gamma_{jt}$ will inform us of the ex-post uncertainty of whether $\eta_{jt}$ is zero or not, i.e.,\ $(jt)\in\mathcal{K}_t$ or not. For each market $t$, the $J_t$-dimensional vector $\gamma_{t}=(\gamma_{1t},\ldots,\gamma_{J_tt})'$ summarizes the $2^{J_t}$ possible sparsity patterns. Ex-ante, all of the $2^{J_t}$ configurations, including the “dense” ones where the majority of the $\xi_{jt}$'s are unrestricted, are considered. The data then informs us the posterior belief over them.

A priori, the binary variable $\gamma_{jt}$'s are i.i.d. and follow a Bernoulli distribution with the prior inclusion probability $\phi_t$ that is specific to market $t$ to allow for different degrees of sparsity across markets $ t=1,\ldots,T$: \[ \gamma_{jt} \overset{iid}{\sim}\text{Bernoulli}(\phi_t), \ j=1,\ldots,J_t. \] We follow the convention and specify a beta prior on $\phi_t$: \[ \phi_t \overset{iid}{\sim} \text{Beta}(\underline{a}_\phi,\underline{b}_\phi), \ t=1,\ldots,T. \] Markets $(t)$ in the sparse set $\mathcal{S}$ are associated with small values of $\phi_t$, and market-product pairs $(jt)$ with $\gamma_{jt}=0$ are in the sparse product set $\mathcal{K}_t$. We set $(\underline{a}_\phi,\underline{b}_\phi)=(1,1)$ so that ex-ante, all the market-specific prior inclusion probability $\phi_t$ has mean of 0.5. In other words, the prior probability that each market-product shock $\xi_{jt}$ is sparse is 50%.

For the remaining parameters, we employ standard priors independently:

align*[align* omitted — 282 chars of source]

a $d_X$-dimensional normal prior for the slope vector $\bar{\beta}$, a normal prior for the market-specific product shock $\bar{\xi}_t$ independently over markets, and a normal prior for the log standard deviation of the random coefficients independently over covariates. We let $(\underline{\mu}_\beta,\underline{V}_\beta)=(0, 10 \cdot I_{d_X})$ and $(\underline{\mu}_{\xi_t},\underline{V}_{\xi_t})=(0, 10)$ for all markets $t$ to give sufficiently uninformative priors on the fixed slope parameter $\bar{\beta}$ and the market-specific shocks $\bar{\xi}_t$. For the log standard deviations, we let $\underline{V}_{r,k}=0.5$ for all $k$ to give sufficiently uninformative prior on $\Sigma$.

The hyperparameters $(\tau^2_0,\tau^2_1)$ in the spike-and-slab prior (ref) are chosen by the researcher. In the literature of shrinkage priors, it is known that computational problems can arise when the ratio $\tau^2_1 / \tau^2_0$ is too large. They can be avoided when $\tau^2_1 / \tau^2_0 \leq 10,000$ (GeorgeMcCulloch1993; GeorgeMcCulloch1997). We recommend to fix them as $(\tau^2_0,\tau^2_1)=(10^{-3},1)$ and use these values in our simulation studies and empirical applications below. The prior variance in the slab component $\tau^2_1=1$ is a reasonably large value. For example, in a canned-tuna category data, Jiang2009bayesian found the posterior mean of the (uniform) variance of the market-product shocks to be around 0.33; and in a facial tissue application, MusalemBradlowRaju2009BayesBLP found the corresponding value to be around 0.72. A semi-automated approach would be to fit the model under $\eta_{jt} \sim N(0,\tau^2)$ i.i.d. for all $(jt)$ with an uninformative prior on $\tau^2$, and let e.g.,\ $\tau^2_0=10^{-2}\hat{\tau}^2$ and $\tau^2_1=10\hat{\tau}^2$ where $\hat{\tau}^2$ is the posterior mean. In our experience, the default option works better.

Posterior inference

We have the slope parameters $\bar{\beta}$, the log standard deviations for the random coefficients $r=(r_1,\ldots,r_{d_X})'$, the market-specific intercepts $ \bar{\xi}=\{\bar{\xi}_1,\ldots,\bar{\xi}_T\}$, the market-product specific deviations $\eta=\{ \eta_1,\ldots,\eta_T \}$, where $\eta_t=(\eta_{1t},\ldots,\eta_{J_t t})'$, the binary indicator variables $\Gamma =\{ \gamma_1,\ldots,\gamma_T \}$, where $\gamma_t=(\gamma_{1t},\ldots,\gamma_{J_t t})'$, and the inclusion probabilities $\phi=(\phi_1,\ldots,\phi_T)'$. The data contains the quantity demanded $q=\{q_1,\ldots,q_T \}$, where $q_t = \{ q_{1t},\ldots,q_{J_tt}\}$ and the market-level covariates $X=\{ X_1,\ldots, X_T\}$. Then, from Bayes' theorem, suppressing the dependency on the covariates, the posterior density of interest is defined as

equation[equation omitted — 276 chars of source]

The first term on the right-hand side is the likelihood function (ref), and the second term in (ref) gives the prior on the parameters and factors as $ p\left( \bar{\beta}, r, \bar{\xi}, \eta, \Gamma, \phi \right) = p( \bar{\beta}) p(r) p(\bar{\xi}) p(\eta, \Gamma, \phi ) $. The first three terms are the prior on $\bar{\beta}$, $\bar{\xi}$, and $r$, respectively. The last term defines the prior on $\eta_{jt}$'s and

equation[equation omitted — 381 chars of source]

where $\pi(\phi_t)$ is the prior on $\phi_t$.

The model is estimated via Markov chain Monte Carlo (MCMC). We obtain a posterior sample $\{ \bar{\beta}^{(g)}, r^{(g)}, \bar{\xi}^{(g)}, \eta^{(g)}, \Gamma^{(g)}, \phi^{(g)} \}_{g=1}^G$, where $G$ is the total number of MCMC draws (after discarding an appropriate burn-in draws). Using the posterior sample, one can easily conduct inference on any functions of the model parameters, such as elasticity.

Roughly speaking, our MCMC algorithm for sampling from the joint posterior distribution iterates between two sets of conditional distributions. The first set of conditionals is used for updating $(\bar{\beta}, r, \bar{\xi},\eta)$ the utility parameters common across markets and products ($\bar{\beta}, r$) as well as the market-specific intercepts and the market-product specific shocks ($\bar{\xi},\eta$). The second set of conditionals is for the parameters related to the latent sparsity structure $(\Gamma,\phi)$. The two sets of conditionals are

align*[align* omitted — 402 chars of source]

The first step can be implemented using the tailored Metropolis-Hasting algorithm ChibGreenberg1995understandingMH, whose efficient sampling is made possible by exploiting the existence of gradients and Hessian matrices of the log-likelihood function with respect to the relevant parameters. The second set of conditionals can be implemented based on the conjugacy known in the high-dimensional Bayesian statistics literature. The computational details of the algorithm can be found in the Appendix.

remarkOur proposed inference procedure has several appealing features. First of all, it is conceptually simple as it is based on the standard likelihood ((ref)) (i.e., McFadden's classic framework) coupled with shrinkage priors on \(\xi\)'s; in particular, it does not rely on the BLP machinery of demand inversion or IVs as additional identification restrictions. Moreover, because no demand inversion is needed, our method has two practical advantages over alternative approaches that are based on the inversion. First, our method can accommodate zeros in market shares data, which is an important empirical problem in many applications, offering an alternative to existing approaches such as gandhi2023estimating. Second, our method is computationally more scalable than alternative Bayesian procedures like Jiang2009bayesian and Hortaccsu2023Bayes, where the inversion needs to be computed in each MCMC iteration; this advantage becomes more prominent as $T$ and/or $J_t$'s become large. Finally, our method can conveniently deliver inference results (e.g., credible intervals) using posterior draws, for model parameters, including the sparsity structure of \(\xi\)'s, and counterfactual quantities such as price elasticities. The computation of price elasticities is described in the Appendix and demonstrated in an empirical application in Section (ref).

Monte Carlo Simulations

In this section, we examine the performance of our proposed approach via a series of Monte Carlo experiments and compare with the standard BLP estimator with alternative IV choices.

Simulation Design

We generate data from the following random coefficient logit model, where the utility of consumer $i$ for product $j$ in market $t$ is specified as \[ u_{ijt} = \beta_{pi} p_{jt} + \beta_w^\ast w_{jt} + \xi_{jt}^\ast + \varepsilon_{ijt}, \] where $\beta_{pi} \sim N(\beta_p^\ast,\sigma^{\ast 2})$ is the random coefficient on the endogenous variable price \(p_{jt}\), \(\beta_w^\ast\) is a fixed coefficient on the exogenous product characteristic $w_{jt}$, $\varepsilon_{ijt}$ is i.i.d.\ across $i,j,t$ following the standard Gumbel distribution.

The exogenous product characteristic $w_{jt}$ is i.i.d. across $j,t$ and generated from $U(1,2)$, the uniform distribution with support $(1,2)$. The endogenous variable price is generated as \[ p_{jt}= \alpha_{jt}^\ast+ 0.3 w_{jt} + u_{jt}, \] where $u_{jt}$ can be interpreted as a “cost shock” that is i.i.d. across $j,t$ and drawn from a $N(0,.7^2)$. The unobserved market-product characteristics are generated as \[ \xi_{jt}^\ast=\bar{\xi}_t^\ast+\eta_{jt}^\ast, \] where $\bar{\xi}_t^\ast$ is fixed at $-1$ for all $t$.

The key parameters of interest are $\beta_p^\ast=-1$, $\beta_w^\ast=0.5$, and $\sigma^\ast=1.5$. The specification of $\alpha_{jt}^\ast$ and $\eta_{jt}^\ast$ varies by the following four DGP designs: sparse \(\xi\) with exogenous \(p\) (DGP1), sparse \(\xi\) with endogenous \(p\) (DGP2), non-sparse \(\xi\) with exogenous \(p\) (DGP3), non-sparse \(\xi\) with endogenous \(p\) (DGP4).

In DGP1, for each $t$, the first 40% of the elements in the vector $\eta_t^\ast=(\eta_{1t}^\ast,\ldots,\eta_{Jt}^\ast)^{\top}$ are non-zero: the odd components are set to $1$ while the even ones are set to $-1$. The remaining 60% of the components are set to zero. The $\alpha_{jt}^\ast$ in the price equation is set to 0 for each $(j,t)$, so the vector $\alpha_{t}^\ast = (\alpha_{1t}^\ast,\ldots,\alpha_{Jt}^\ast)^{\top} $ is independent of $\eta_t^\ast$.

In DGP2, we introduce price endogeneity by letting \(\alpha_{jt}^\ast\) depend on $\xi_{jt}^\ast$. In particular, we set $\eta_{jt}^\ast$'s the same as in DGP1 and let $\alpha_{jt}^\ast=0.3$ if $\eta_{jt}^\ast=1$, $\alpha_{jt}^\ast=-0.3$ if $\eta_{jt}^\ast=-1$, and $\alpha_{jt}^\ast=0$ otherwise. This implies a positive correlation between price \(p_{jt}\) and the unobserved characteristics $\xi_{jt}^\ast$.

In DGP3 and DGP4, we consider a non-sparse structure of $\eta_{jt}^\ast$. In particular, they are i.i.d.\ draws from the normal distribution with zero mean and standard deviation $1/3$, i.e., $\eta^*_{jt} \sim N(0,(1/3)^2)$. The distribution has a large mass around zero, which can be regarded as approximately sparse. The purpose of this design is to examine how our approach works when the sparsity assumption is mildly violated. The $\alpha_{jt}^\ast$ in DGP3 is the same as DGP 1. For DGP4, price is endogenous and positively correlated with \(\xi\): we let $\alpha_{jt}^\ast=0.3$ if $\eta_{jt}^\ast\geq 1/3$, $\alpha_{jt}^\ast=-0.3$ if $\eta_{jt}^\ast\leq -1/3$, and $\alpha_{jt}^\ast=0$ otherwise.

Given the specification of the utility function, the market shares and quantities are simulated based on ((ref)), using \(N_t=1000\) consumer draws from the distribution of the random coefficient on price. The number of products is the same across the markets, i.e., $J_t=J$ $\forall t$. We consider different numbers of markets and products: $T\in \{25,100\}$ and $J\in\{5,15\}$. We simulate 50 data sets $\{ (q^{(r)},X^{(r)}) \}_{r=1}^{50}$ for each case, and implement the following three estimation strategies:

itemize• The BLP estimator that uses $(1,w_{jt}, w_{jt}^2, u_{jt}, u_{jt}^2)$ as instruments, labeled as “BLP (with cost IV)”, where $u_{jt}$ is the exogenous cost shock in the price equation which is typically unobservable in real data. This estimator uses a set of valid IVs and provides a benchmark for comparing the other two approaches. • The BLP estimator that uses $(1,w_{jt}, w_{jt}^2, w_{jt}^3, w_{jt}^4)$ as IVs, labeled as “BLP (without cost IV)". This set of IVs is a natural choice in practice when the only observed exogenous variable is \(w_{jt}\) and the cost shock is unavailable to the researcher. We also tried other IVs, including the BLP type of IVs, e.g., the sum of other products' \(w\)'s, and they perform similarly or worse than our current choice. Note that this choice undermines the IV rank condition because the moments of \(w_{jt}\) tend to be highly correlated with each other. We use this case to illustrate the identification problem caused by poor choices of IVs, e.g., weak IVs, which may happen in practice. • Our proposed Bayesian shrinkage approach with a spike-and-slab prior (“Shrinkage”). We use the priors described earlier and $R_0=200$ i.i.d.\ draws from the $d_X$-dimensional independent normal distribution for approximating the choice probabilities.

We report the estimation results of $\bar{\beta}$, $\sigma$, and $\xi$ from the repeated study. For the Bayesian shrinkage approach, we use the posterior mean as the point estimator to make it comparable with the BLP estimator. For the BLP estimator, we estimate $\xi$ by solving for the mean utility $\hat{\delta}_{jt}$'s at the estimated $\hat{\sigma}$ and define $\hat{\xi}_{jt}=\hat{\delta}_{jt}- x_{jt}^{\top}\hat{\beta}$.

Results

Table (ref) reports bias and standard deviation of the estimators under DGP1 and DGP2. As expected, in general, BLP (with cost IV) and Shrinkage outperform BLP (without cost IV). In particular, BLP (without cost IV) has large biases and standard deviations in many cases, highlighting the potentially severe identification and estimation issues caused by weak or invalid IVs.

In both exogenous (panel (a)) and endogenous (panel (b)) cases in Table (ref), the shrinkage approach clearly outperforms the BLP (without cost IV) in terms of bias and standard deviation; in many cases, it achieves similar or even better performance to the benchmark estimator BLP (with cost IV), especially in terms of estimating \(\sigma\) and \(\xi\)'s. This result supports our identification strategy that exploits the sparsity of \(\xi\) instead of relying on IVs; also, it shows that the Bayesian shrinkage inference procedure works well in the current setting.

To further confirm that our inference procedure works as expected, we examine the estimated sparsity pattern of \(\xi\). The nice feature of the spike-and-slab prior is that it allows us to compute the posterior probability that $\eta_{jt}$ is nonzero (i.e.,\ $\xi_{jt}$ deviates from the market-specific common shock $\bar{\xi}_t$), which is equivalent to the event $\gamma_{jt}=1$. The last column of Table (ref) reports the probability that $\gamma_{jt}=1$ when the true value $\eta_{jt}^\ast$ is indeed nonzero (first row) and the probability of the same event when $\eta_{jt}^\ast$ is zero (second row). Overall, our procedure can uncover the sparsity structure in $\xi$ reasonably well, giving a higher probability for the market-product pair $(j,t)$ when the true value of $\eta_{jt}^\ast$ is nonzero and a lower probability otherwise.

In DGP3 and DGP4, the market-product shocks are not sparse but approximately so. We consider such cases to examine the robustness of our approach when the sparsity assumption is mildly violated. Table (ref) shows the results in the same format as Table (ref). We can see that even with non-sparse \(\xi\)'s in the data-generating process, the proposed method outperforms BLP (without cost IV) and is comparable to BLP (with cost IV) in most cases.

In summary, the simulation studies indicate that the proposed approach effectively uncovers the latent sparsity structure in the unobserved market-product shocks $\xi$'s when they are sparse. It also provides reliable estimates for other structural parameters under both sparse and non-sparse $\xi$'s. Furthermore, our approach often matches the performance of the BLP estimator with strong but impractical IVs and outperforms the BLP estimator when poor IVs are used, making it a compelling alternative when good IVs are unavailable or their validity is uncertain.

\FloatBarrier

table[table omitted — 4,413 chars of source]

\setcounter{table}{1}

table[table omitted — 4,191 chars of source]

\FloatBarrier

Empirical Applications

In this section, we begin by applying the proposed method to analyze consumer demand and store promotion strategies in the yogurt market using the IRI dataset.\footnote{See bronnenberg2008iri for a description of the IRI marketing dataset.} This application emphasizes the ability of our method to uncover the sparse patterns in store promotion activities, where only a few products receive special promotions in a given store during a specific week due to space constraints. In the second application, we revisit the U.S. auto market dataset from berry1995automobile to assess the performance of our method in a well-documented market setting. In both applications, we find evidence of sparsity in the market-product level demand shocks, highlighting the relevance of our approach for capturing such latent structure and exploiting it for identification. Furthermore, in both data sets, the estimated structural parameters are roughly comparable to those from the standard BLP method, but our approach does not require IVs, which underscores the value of our approach as a robust alternative to the traditional IV-based method.

Consumer Demand and Store Promotion in Yogurt Market

Data

This analysis focuses on the yogurt category, using data from 95 stores located in the New York market (defined by the IRI dataset) for a single week, the week of June 25 to July 1, 2012. This sample selection allows us to make the sample size manageable while retaining sufficient variation in product characteristics, prices, and promotional activities. Specifically, the data provides detailed UPC level information, including weekly price, quantity, product characteristics, and marketing mix variables, for each store in the sample.

We aggregate the UPCs into “products,” which are defined by a combination of brand, size category (size 1 to size 4 defined using three thresholds: 0.9, 1.3, and 1.9 pints), and product characteristics -- such as flavored or not, low-fat, Greek, organic, etc. -- as shown in Table (ref). Product prices are calculated as quantity-weighted averages, while quantities are obtained through simple summation across UPCs within each product. The marketing mix indicator variables, Display and Feature, are marked as active (equal to 1) if any UPC within the product is active. This aggregation decreases the number of observations while preserving the essential variation in product attributes.

A market is defined by a store, with the consumers' choice set comprising all products available in that store. The market share of a product is calculated as the quantity sold divided by the population size within the local area surrounding the store, as provided by the IRI dataset. In total, there are 5,927 unique market-product pairs in the data.

Table (ref) presents summary statistics for several randomly selected products. The “No. of Markets” column indicates the markets where each product is available, highlighting substantial variations in consumers' choice sets. The “Market Share (%)” and “Price” columns, as well as the marketing mix columns, report the averages across different markets for each product. One pattern that stands out is the greater variation in market shares across markets compared to prices, as reflected in the mean-to-standard-deviation ratio.

table[table omitted — 2,334 chars of source]

Model Specification

With the above data, we consider the following discrete choice demand model, where the utility function is specified as

equation[equation omitted — 90 chars of source]

where \(X_{jt}\) is a 24-dimensional (i.e., \(d_X = 24\)) vector of product characteristics, including price, dummy variables for 16 brands, 4 product sizes, and indicators for whether the product is flavored, non-fat, low-fat, Greek, or organic. We introduce random coefficients on price and the organic indicator. Specifically, \(\beta_i \sim N_{24}(\bar{\beta}, \Sigma)\), where \(\Sigma = \text{diag}(\sigma^2_1, 0, \ldots, 0, \sigma^2_{d_X})\).

Recall that the market-product demand shocks, which capture promotion efforts, are modeled as

equation[equation omitted — 61 chars of source]

where \(\bar{\xi}_t\) represents a store-level demand shock, potentially reflecting overall store-level promotions, and \(\eta_{jt}\) denotes the product-specific deviation, driven by promotional efforts that could originate from the manufacturer or the store itself.

The product-specific promotion is naturally sparse due to the space constraints of stores, as a store can only promote a limited number of products in a given week. In the data, we observe variables such as display and feature, which partially capture these promotional efforts (and, notably, these variables already exhibit a sparse pattern). However, they are noisy measures of the actual promotion effort, meaning some promotional activities are not recorded. Using our approach, we aim to directly estimate the underlying promotion efforts and, ex-post, evaluate how well the observed display and feature variables explain the estimated promotion.

Recall that there are 5,927 market-product pairs in the data, implying 5,927 independent first-order conditions if one were to use the MLE to estimate the model defined by ((ref)). However, the number of parameters to estimate is $5927 (\eta_{jt}) + 95 (\bar{\xi}_{t}) + 24 (\bar{\beta}) +2 (\Sigma)$, making the model underidentified. As our theoretical result shows, introducing sparsity on $\xi_{jt}$'s can restore identification, and we implement our Bayesian shrinkage approach to estimate the model.

We use the priors defined in Section (ref) and $R_0=200$ i.i.d.\ draws from the standard normal distribution for approximating the integrals in the choice probabilities. The MCMC procedure consists of 10,000 draws, with the first 3,000 discarded for burn-in, leaving 7,000 draws for estimation and inference.

For comparison, we implement the standard BLP GMM estimator using the same model specification. The instrumental variables (IVs) are constructed by interacting lagged prices (along with other product characteristics) with market dummies.\footnote{We construct the BLP GMM estimator based on \(E\left[\eta_{jt}\left|Z_{jt}\right.\right]=0\), where \(Z_{jt}\) is the vector of chosen IVs, and estimate \(\bar{\xi}_{t}\)'s as market fixed effects.} Additionally, we include results from simple logit specifications estimated using both OLS and IV methods for comparison.

Estimation Results

table[table omitted — 4,711 chars of source]

\FloatBarrier

The estimation results for the preference parameters (\(\bar{\beta}\) and \(\Sigma\)) are presented in Table (ref). The estimated slope coefficients for price and the organic indicator in the proposed approach have reasonable signs and magnitudes, aligning closely with the results from the BLP approach. Both approaches also provide significant evidence of dispersion in the random coefficients for these variables. Additionally, well-known brands, such as Chobani, Fage Total, and Stonyfield Organic Oikos, exhibit relatively larger brand fixed effects in the consumer utility function. Finally, the random coefficient model (either Bayesian or BLP estimates) implies a more elastic demand compared to the model without random coefficients, as indicated by the last two rows of the table. This difference is primarily driven by the dispersion in the random coefficients on price, which captures heterogeneity in consumer sensitivity to price changes.

Overall, our Bayesian approach produces similar results to the BLP approach in this case. We emphasize that our Bayesian shrinkage approach does not rely on IVs, and the agreement between the two approaches here validates the BLP results that rely on IVs. However, such agreement is not guaranteed in general; in cases where the two approaches diverge, it becomes essential to assess which underlying assumption -- sparsity or the validity of IVs -- is more plausible in the specific context.

We now turn to discuss the latent sparsity structure of the market-product shocks \(\xi_{jt}\) uncovered by our procedure, as summarized in Figure (ref). The solid lines in Figure (ref) represent the posterior means of \(\phi_t\)'s, which indicate the degree of sparsity in each of the 95 markets. Many of them fall within the low range of 5% to 20%, implying that these markets are likely to belong to the sparse set \(\mathcal{S}\). Importantly, this is a substantial ex-post evidence of sparsity as the prior mean of $\phi_t$'s is set to 50% (see Section (ref)).

figure[figure omitted — 1,126 chars of source]

\FloatBarrier

We can learn about the estimated sparsity structure at product-level by looking at the posterior means of the binary indicator $\gamma_{jt}$, represented by the colored dots in Figure (ref). Different colors correspond to different markets. Recall that the product is sparse, i.e.\ $j\in\mathcal{K}_t$, if $\gamma_{jt}=0$. For many of the market-product pairs $(jt)$, the posterior means of $\gamma_{jt}$ are below 0.1, implying that, ex-post, they are likely to belong to the sparse set $\mathcal{K}_t$. This does not imply that all pairs are sparse; in fact, 658 pairs exceed the prior mean of 0.5, indicating a substantial subset of dense pairs.

Figure (ref) illustrates the estimated values of the demand shocks, $\xi_{jt}=\bar{\xi}_t+\eta_{jt}$. The solid lines correspond to the posterior means of the market-specific values \(\bar{\xi}_t\)'s and the colored dots are those of the deviations $\eta_{jt}$'s. The posterior mean of \(\xi_{jt}\) is obtained by vertically summing those of \(\bar{\xi}_t\) and \(\eta_{jt}\). As expected, many \(\eta_{jt}\)'s are close to zero, confirming a high degree of sparsity in the data. Again, we emphasize that not all of the $\xi_{jt}$'s are shrunk to the market-level $\bar{\xi}_t$, as characterized by significant deviations of $\eta_{jt}$'s away from zero for some pairs $(jt)$. Lastly, an interesting pattern emerges: the distribution of \(\eta_{jt}\) is right-skewed, with more \((jt)\) pairs exhibiting positive \(\eta_{jt}\)'s than negative ones. These positive \(\eta_{jt}\)'s likely capture store-product level promotion efforts, as discussed earlier.

The estimated market-product demand shocks reflect promotional efforts. To evaluate how much of these efforts are explained by the observed marketing mix variables, namely display and feature indicators, we regress the estimated \(\eta_{jt}\)'s on these variables. Table (ref) presents the regression results for both the shrinkage and BLP approaches. The slopes on the marketing mix variables are positive and significant in both cases, suggesting that \(\eta_{jt}\) effectively captures store-product-level promotion activities. However, based on the \(R^2\), the marketing mix indicator variables explain only 2.5% of the estimated promotion efforts, highlighting the presence of potentially substantial unobserved store-level marketing activities. This finding underscores the noisiness of the marketing mix variables in capturing store promotions and emphasizes the importance of incorporating the market-product demand shocks \(\xi_{jt}\)'s into the model.

table[table omitted — 786 chars of source]

Finally, we examine product-level price elasticities, which are a key output of demand estimation. As mentioned earlier, once MCMC draws are obtained, computing point and interval estimates of elasticity is straightforward, representing a key advantage of the proposed approach. Table (ref) presents posterior means and 95% credible intervals of own-price elasticity for selected products in a market, sorted by price. For comparison, the table also includes elasticity estimates based on the BLP approach and the simple logit model (IV and OLS).

One notable observation is that when random coefficients (particularly on price) are incorporated, as in the shrinkage and BLP approaches, a U-shaped relationship between price and elasticity emerges, consistent with findings in the literature, such as berry1995automobile. In contrast, the simple logit model shows a monotonic increase in elasticity with price. Furthermore, the magnitudes of the elasticities estimated by our procedure are reasonable, and all products exhibit elastic demand (i.e., greater than 1).

table[table omitted — 1,624 chars of source]

Revisit the BLP Auto Data

We revisit the classic BLP application to the U.S. automobile market. The BLP auto dataset contains product-level prices, quantities, and characteristics for major car models in the U.S. market for each year from 1971 to 1990. Following berry1995automobile, we define each year as a market, resulting in 20 markets (\(T=20\)) and an average of approximately 110 products per market. A detailed description of the dataset is provided in berry1995automobile.

This additional application is valuable for several reasons. First, the industry context differs sharply: the automobile market involves durable goods and large, infrequent purchases, whereas the yogurt market represents low-cost grocery items and frequent consumption. Second, the scope of the data is distinct: the auto dataset captures national-level demand for automobiles, while the yogurt application focuses on highly disaggregated store-level activities. By applying our method to these two contrasting settings, we demonstrate its flexibility in uncovering sparse demand shocks in distinct scenarios with diverse market structures and datasets.

We consider a similar model structure as in the yogurt application. The product characteristics (\(X_{jt}\)) include price, horsepower/weight (log), weight (log), size (log), dollar/mile (log), and indicators for air conditioning, power steering, automatic transmission, and forward drive (\(d_X = 9\)). Random coefficients on \(X_{jt}\) follow \(\beta_i \sim N_{d_X}(\bar{\beta}, \Sigma)\), where \(\Sigma = \text{diag}(\sigma_1^2, \ldots, \sigma_{d_X}^2)\). The market-product shocks (\(\xi_{jt}\)) are decomposed into market fixed effects (\(\bar{\xi}_t\)) and product-specific deviations (\(\eta_{jt}\)).

This specification differs from berry1995automobile in two key ways: (1) it excludes a supply-side model, and (2) it includes market fixed effects to account for market-level heterogeneity. Our goal is not to replicate their results but to demonstrate how our approach can uncover sparsity in this classic dataset.

Compared to the yogurt application, the auto dataset differs in several aspects. While the yogurt data feature more markets and fewer products per market, the auto dataset includes fewer markets (20 years) but an average of 110 products per market, totaling 2,217 market-product pairs. As in the yogurt case, the number of parameters exceeds the number of independent first-order conditions, making sparsity assumptions on \(\eta_{jt}\) essential to restore identification. We use our shrinkage approach for estimation, following the same MCMC procedure. Again, for comparison, we include the standard BLP GMM estimator, with IVs constructed by interacting the BLP IVs with market dummies, as well as simple logit specifications (OLS and IV).

table[table omitted — 3,059 chars of source]

The estimation results for the preference parameters (\(\bar{\beta}\) and \(\Sigma\)) are presented in Table (ref). For the mean random coefficients, our Bayesian shrinkage approach yields estimates with reasonable signs and magnitudes, closely aligning with those of the standard BLP estimates. Regarding the standard deviations (SDs) of random coefficients, the Bayesian shrinkage approach indicates considerable dispersion for all random coefficients, suggesting rich heterogeneity in consumers' tastes across all product characteristics. In contrast, several SDs from the BLP estimates, including those for weight, size, power steering, and automatic transmission, are virtually zero. These near-zero estimates may be attributed to the weak IV problem, as highlighted by reynaert2014improving. Furthermore, while the BLP estimator is sensitive to the choice of IVs -- based on our experiments with the data, though specific results are not reported here -- our Bayesian shrinkage approach is immune to this issue, making it a particularly advantageous tool in practice.

\FloatBarrier

figure[figure omitted — 1,130 chars of source]

\FloatBarrier

Now we turn to the latent sparsity structure of the market-product shocks \(\eta_{jt}\) identified by our procedure, as summarized in Figure (ref). This figure serves as the counterpart to Figure (ref) from the yogurt application. Overall, we find stronger evidence of sparsity in this dataset compared to the yogurt data. The solid lines in Figure (ref) show the posterior means of \(\phi_t\)'s, which indicate sparse markets. The average posterior mean is notably small, at 9.8%, with particularly low values (less than 0.05) observed in the 1982, 1983, and 1990 markets. Among the 2,217 market-product pairs in this dataset, only 133 have a posterior mean of \(\gamma_{jt}\) greater than 0.5, represented by the colored dots. Additionally, Figure (ref) reveals that the distribution of \(\eta_{jt}\) is right-skewed, a pattern consistent with the yogurt application.

Conclusion

In this paper, we have proposed a new approach to estimating the random coefficient logit demand model with sparse market-product level demand shocks. Our approach eliminates the need for instrumental variables (IVs), which are required in the standard BLP GMM method. We show that, under certain regularity conditions, the demand shocks and their sparsity structure can be identified along with other model parameters. We also propose a Bayesian shrinkage estimation procedure that offers a scalable and flexible alternative to existing methods.

We demonstrate the applicability of our approach through two empirical applications. First, in the context of supermarket scanner data, we interpret the demand shocks as unobserved promotion efforts at the store-week level, capturing the sparsity in promotional activities across products. Second, we revisit the automotive market to assess the performance of our estimator in a well-documented dataset. In both cases, we find strong evidence of sparsity in demand shocks, supporting the relevance of the sparsity assumption in real-world data.

spacing{1.15}