EconBase
← Back to paper

Choice Models and Permutation Invariance: Demand Estimation in Differentiated Products Markets

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

118,913 characters · 19 sections · 92 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Choice Models and Permutation Invariance: Demand Estimation in Differentiated Products Markets

\affil[1]{University of Washington}

abstractChoice modeling is at the core of understanding how changes to the competitive landscape affect consumer choices and reshape market equilibria. In this paper, we propose a fundamental characterization of choice functions that encompasses a wide variety of extant choice models. We demonstrate how non-parametric estimators like neural nets can easily approximate such functionals and overcome the curse of dimensionality that is inherent in the non-parametric estimation of choice functions. We demonstrate through extensive simulations that our proposed functionals can flexibly capture underlying consumer behavior in a completely data-driven fashion and outperform traditional parametric models. As demand settings often exhibit endogenous features, we extend our framework to incorporate estimation under endogenous features. Further, we also describe a formal inference procedure to construct valid confidence intervals on objects of interest like price elasticity. Finally, to assess the practical applicability of our estimator, we utilize a real-world dataset from berry1995automobile. Our empirical analysis confirms that the estimator generates realistic and comparable own- and cross-price elasticities that are consistent with the observations reported in the existing literature. \noindentKeywords: Choice Models, Demand Estimation, Permutation Invariance, Set Functions, Neural networks

\setcounter{page}{1} \thispagestyle{empty}

Introduction

Demand estimation is a critical component in fields of marketing, operations, and economics, enabling practitioners to model consumer choice behavior and understand how consumers react to changes in a market. This understanding helps policymakers and businesses to make informed decisions, whether it be about launching new products, adjusting pricing strategies, or analyzing the consequences of mergers nevo2000mergers, petrin2002quantifying, nevo2003new, wollmann2018trucks. Over the years, various approaches, both parametric and non-parametric, have been developed to address the complexities inherent in demand estimation.

Parametric methods, based on logit or probit assumptions, often require strong assumptions about the underlying choice process, limiting their ability to capture the true complexity of consumer preferences. As such, the substantive and policy (counterfactual) implications from such models could be largely biased and/or misleading. Nevertheless, they have remained popular because of a few reasons -- (1) simplicity, interpretability, and scalability to a large product set, (2) the ability to model counterfactual demand in situations involving new product introductions or mergers, and (3) the ability to handle endogenous product features.

Non-parametric methods, on the other hand, offer a more flexible approach to demand estimation, allowing for more nuanced representations of consumer preferences without restrictive assumptions about the underlying distributions of observables and unobservables. However, despite their potential advantages, non-parametric approaches face significant challenges that have limited their widespread adoption in practice. Firstly, a common limitation across all non-parametric demand estimation methods is the “curse of dimensionality," where the computational complexity of estimating choice functions increases exponentially with the number of products. Additionally, a large subset of these models, specifically those that are `black box' in nature and lack additional economic structure, are unable to perform counterfactual predictions, which are crucial for important tasks such as policy simulations. Furthermore, most non-parametric methods cannot properly handle endogenous product features. This inability to account for endogeneity and scalability constraints, along with the drawback of limited counterfactual analysis, significantly undermines their utility in practice, where such capabilities are often the primary objective of demand estimation.

In this paper, we make significant strides in bridging the gap between the flexibility of non-parametric methods and the tractability of parametric models. We achieve this by introducing a fundamental characterization of choice models. Our work begins by considering a broad set of choice functions, encompassing most existing choice modeling approaches in the empirical literature. We demonstrate that most choice functions exhibit specific symmetry properties. We then leverage recent advances in computer science and mathematics literature (for instance, han2019universal, zaheer2017deep, wagstaff2019limitations) to characterize these functions. Our characterization allows us to build on the strengths of the non-parametric approaches (e.g., not assuming a model of consumer behavior) and overcome the challenges associated with them. First, it addresses the curse of dimensionality in choice systems and enables the flexible estimation of choice functions via non-parametric estimators. By leveraging the inherent permutation-invariant structure of choice models, our characterization can close the estimator's generalization gap (for instance, see sannai2019improved in the context of neural networks) by a factor of $\sqrt{J!}$ (where $J$ is the number of products). Second, we recognize that real-world demand systems often contain unobserved demand shocks correlated with observable product features, such as prices. These shocks can lead to endogeneity issues, which can bias the estimated choice functions if not properly addressed. To tackle this challenge, we extend our framework to accommodate endogeneity. Third, our proposed approach successfully estimates counterfactual demand for scenarios that involve changes in the product set (e.g., new product introduction, mergers, product exit). Finally, we build upon recent advances in automatic debiased machine learning and provide an inference procedure for constructing valid confidence intervals on objects of interest, such as the average effect of price.

We demonstrate the effectiveness of our proposed approach using a series of numerical simulations. We consider a variety of data-generating processes in our simulations -- multinomial logit model with linear utility, random coefficients logit with linear utility, random coefficient logit with non-linear utilities, and a setting where some consumers have inattention (i.e., ignore certain products in the market). Across all these scenarios, we show that our approach can predict market shares, own-, and cross-price elasticities with relatively high accuracy, even though we do not make any assumptions about consumer behavior. Further, even when the underlying DGP is complex, our model can recover market shares and elasticities similar to oracle estimators (that are assumed to know the true DGP). This allows researchers and managers to use our approach in general-purpose situations without making ad-hoc assumptions about consumer behavior.

Next, we consider counterfactual analyses and empirically show that our model can generate accurate counterfactuals in the case of new product introductions. Finally, we showcase the performance of the automatic debiasing procedure and show that we can provide consistent inference and confidence intervals over the average effect of price on demand across all the products. In sum, our extensive numerical simulations cover a wide range of DGPs and show that our approach -- (1) is able to accurately predict market shares and price elasticities for a variety of scenarios, (2) can generate realistic counterfactual predictions in cases where the product set changes, and (3) can provide inference over economic objects of interest.

Finally, to showcase the effectiveness and applicability of our approach to real-life datasets, we use the berry1995automobile automobile dataset and estimate the price elasticities using our non-parametric estimator (while correcting for price endogeneity). The results of our analysis align with existing literature and demonstrate the practical utility of our approach, underscoring its potential for adoption in real-world demand estimation settings. In summary, given the theoretical, practical, and computational advantages of our approach, we expect it to be easily applicable to a wide variety of demand estimation problems in both research and practice.

Literature Review

Parametric Discrete Choice Models

Discrete choice models play a crucial role in various fields, including economics, marketing, and operations management, as they describe decision-making processes when individuals face multiple alternatives. These models typically rely on a random utility maximization assumption, which can be traced back to thurstone1927law and marschak1959binary.

Over time, various choice models have emerged under different specifications for the density of unobserved utility, following the general framework in marschak1959binary. The (Multinomial) logit model, first proposed by luce1959possible, is widely used for capturing systematic taste variance based on observed characteristics across alternatives. However, the logit model assumes that error terms are independent of each other and of the characteristics of the alternative. Additionally, the independent extreme value distribution results in the irrelevant alternatives (IIA) property, implying proportional substitution across alternatives.

To enable more flexible substitution patterns, Generalized Extreme Value (GEV) models were introduced, allowing for correlated unobserved utility across alternatives. The most commonly used GEV model is the nested logit model train1987demand, forinash1993application. In nested logit models, the unobserved error of alternatives within a nest is specified as correlated, while the marginal distribution of each unobserved error remains as a univariate extreme value.

The mixed logit model mcfadden2000mixed, also known as random coefficient logit (RCL), is an even more versatile model that allows for randomness in both unobserved factors and coefficients of observed characteristics. The model was first applied by boyd1980effect and cardell1980measuring, with random coefficients typically specified as normal or lognormal ben1993estimation, mehndiratta1996time, revelt1998mixed but also other distributions such as uniform and triangular greene2003latent, train2001comparison. When random coefficients follow a mixture distribution, the model becomes the well-known mixed logit latent class model.

Aside from logit family models (MNL, GEV, mixed logit) that have an extreme value distributed random component, the probit model hausman1978conditional assumes a jointly normal distribution for the error term. This allows for any pattern of substitution and can handle random taste variation. blanchet2016markov proposed a model where substitution between alternatives is a state transition in a Markov chain, which can approximate MNL, probit, and mixed logit models.

As the decision space grows, assuming that all alternatives are considered becomes unrealistic, leading to research in consideration sets and consumer search honka2019empirical,jiang2021consumer. These models typically impose strong parametric assumptions on the decision rule, including whether decisions are simultaneous or sequential, the stopping rule for searches, the size of the consideration set, and the functional form of the match value.

Further, real-world demand systems often contain unobserved demand shocks correlated with observable product features, such as prices. These shocks can lead to endogeneity issues, which can bias the estimated parameters if not properly addressed. Thus, various approaches have been proposed to address these issues. For instance, berry1995automobile proposed a generalized method of moments-based estimator to estimate a random-coefficients logit model of demand using instrumental variables. Similarly, petrin2010control demonstrated the application of control functions to resolve endogeneity in the random coefficient logit model of demand.

Two primary concerns related to discrete choice models are the need for correct model specification and the distribution of unobserved factors. To address these issues, nonparametric demand estimation models have been developed.

Nonparametric Demand Estimation

While parametric discrete choice models make assumptions about the distribution of unobserved variables and specify functional forms for utilities, recent research has developed more flexible semi-parametric and nonparametric approaches for demand estimation. These newer methods ease some of the restrictive assumptions and improve the computational efficiency of choice models while retaining some structure.

Early semi-parametric work manski1987semiparametric, lewbel2000semiparametric, honore2000panel, abrevaya2000rank focused on relaxing the distribution of random shocks in individual-level binary choice models. More recent research khan2021inference, shi2018estimating, pakes2022moment has extended this approach to relax the parametric assumptions on random shocks in individual-level multinomial choice models, allowing for increased flexibility, such as individual-level fixed effects. fox2016nonparametric, briesch2010nonparametric, allen2019identification, fosgerau2021identification,chitla2022nonparametric,lu2023semi,wang2023sieve concentrate on nonparametric identification and estimation of distributions of heterogeneous unobservables, like random coefficients, in various demand models, relaxing assumptions about heterogeneity distribution. More recent work has looked at keeping the indirect utility or choice probability functions largely unspecified while retaining parametric error terms. For instance bentz2000neural,wang2020deep,han2022neural,sifringer2020enhancing,wong2021reslogit,aouad2022representing use neural networks to characterize the indirect utility while retaining the logit error structure. Our work, in comparison, provides a completely flexible mapping from observed product and consumer characteristics to observed demand without relying on any assumptions regarding the choice making process.

Extant research in nonparametric methods aims to completely avoid any parametric assumptions to remove any potential source of misspecification. berry2014identification, berry2020nonparametric presented the identification results for nonparametric estimation of demand from market-level and individual-level data, respectively. Studies by hausman2016individual, blundell2017nonparametric, chen2018optimal focus on individual-level data. In particular, hausman2016individual employs a nonparametric approach to estimate consumer surplus bounds, while blundell2017nonparametric introduces a method for consistently estimating demand functions with nonseparable unobserved taste heterogeneity, subject to the shape restriction imposed by the Slutsky inequality. chen2018optimal concentrates on nonparametric instrumental variables and inference in individual-level data. This paper, alongside compiani2022market and tebaldi2023nonparametric, examines market-level data. compiani2022market develops a nonparametric method based on Bernstein polynomials, drawing on the identification result of berry2014identification. Meanwhile, tebaldi2023nonparametric proposes a technique for deriving nonparametric bounds on demand counterfactuals and applies it to California's health insurance market.

The common challenge across all extant nonparametric demand estimation work has been that the computational complexity of estimating nonparametric demand models increases exponentially with the number of products. For instance, compiani2022market could estimate their nonparametric model for just two products. Similarly, cai2022deep needed data on more than 100,000 markets to reasonably estimate demand with 50 products. To put this in perspective, economic datasets are much smaller; for instance, the dataset in berry1995automobile had 150 products across only 20 markets. This curse of dimensionality has been the primary obstacle preventing the widespread adoption of nonparametric methods in demand estimation. To overcome this barrier, our work demonstrates how to exploit the inherent permutation invariant structure of choice functions to break this curse of dimensionality and flexibly estimate demand in markets with large a number of products. In addition, our work is also related to recent work in marketing wei2022estimating exploring the use of neural networks to estimate parameters of structural models.

Finally, our work leverages recent work in mathematics and computer science (for instance, han2019universal, zaheer2017deep, wagstaff2019limitations), which investigate the universal approximation of symmetric and antisymmetric functions, offering fundamental characterizations for functions defined on sets. We build on this literature to characterize a general class of choice functions in demand systems.

Theory

Choice Models

In this section, we provide a general characterization of consumer choice functions. In particular, we focus on a scenario where researchers have access only to aggregate market-level demand data, while individual-level choices and characteristics remain unobserved. Aggregate demand models have been extensively studied in marketing and economics berry1995automobile, besanko1998logit, sudhir2001competitive, chintagunta2001endogeneity, albuquerque2009estimating, compiani2022market, and are particularly useful when individual-level demand data is not available.

Suppose consumers in a market $t$ face an offer set $\mathcal{S}_t$ that can comprise any subset of $J_t$ distinct products ($\{ 1, 2, \dots, J_t\} $). We use $u_{ijt}$ to represent the index tuple $\{X_{jt}, p_{jt}, I_{it}, \varepsilon_{ijt}\}$, where $X_{jt} \in \mathbb{C}^{d}$ denotes $d$ non-price features belonging to some countable universe $\mathbb{C}^{d}$; $p_{jt} \in \mathbb{C}$ denotes the price of the product; $I_{it} \in \mathbb{C}^{l}$ denotes demographics of consumer $i$ in market $t$, we assume there are $l$ features and belong to some countable universe $\mathbb{C}^{l}$, and $\varepsilon_{ijt}$ denotes random idiosyncratic components pertinent to consumer $i$ for product $j$ in market $t$ that are not unobservable to the researcher but observable to consumers.

definition[Choice Function] Given the offer set $\mathcal{S}_t \subset \{1, 2, 3, \dots, J_t\}$, we define a function $\pi: \{ u_{ijt}: j \in \mathcal{S}_t\} \rightarrow \mathbb{R}^{|\mathcal{S}_t|} $ that maps a set of index tuples $\{u_{ijt}\}_{j \in \mathcal{S}_t}$ to a $|\mathcal{S}_t|$-dimensional probability vector. Each element in the $\pi (\cdot)$ vector represents the probability of consumer $i$ choosing product $j$ in market $t$.

Here we present a very general characterization of choice functions that maps the observable and unobservable components of product and individual characteristics to observed choices through some choice function $\pi$. Note that, traditionally, $u_{ijt}$ is a scalar that represents utility in choice models. However, in our framework, $u_{ijt}$ does not necessarily represent utility. Further, we have not yet imposed any assumption on $\pi$, i.e., how consumers make choices.

We now specify a set of assumptions on the model and data-generating process below.

assumption[Exogeneity] The unobserved error term $\varepsilon_{ijt}$ is independent and identically distributed (i.i.d.) across all products. This can be expressed as follows:

$$\mathbb{P}(\varepsilon_{ijt} \mid X_{\cdot t}, p_{\cdot t}) = \mathbb{P}(\varepsilon_{ijt})$$ This assumption implies that the error term $\varepsilon_{ijt}$ is not correlated with any of the observed variables $X_{\cdot t}$ and $p_{\cdot t}$. As such, it precludes the possibility of endogenous prices and/or marketing-mix variables, as is common in observational data. We start with the basic case with exogenous covariates in this section and later in $\S$(ref), we relax this assumption and allow for endogenous covariates.

assumption[Identity Independence] For any product $j\in \mathcal{S}_t$ and any market $t$, we assume the choice function $\pi$ does not depend on the identity of the product ($jt$). That is:

$$\pi_{ijt}(\{u_{ikt}\}_{k \in \mathcal{S}_t}) = \pi_{ijt}(u_{ijt}, \{u_{ikt}\}_{k \in \mathcal{S}_t, k \neq j}) = \pi(u_{ijt}, \{u_{ikt}\}_{k \in \mathcal{S}_t, k \neq j})$$ This assumption implies two things: first, the functional form of the choice probability for different products and markets is the same; second, for any market-level heterogeneity (e.g., in the distribution $F_t(I_{it}, \varepsilon_{ijt})$), we can include them in $X_{jt}$ as features. Intuitively, this assumption suggests that conditional on product and consumer features and the unobserved error term, the choice probabilities are not functions of the identities of the products themselves.

assumption[Permutation Invariance] The choice function $\pi$ is invariant under any permutation function $\sigma_j()$ that rearranges the indices of the competitors of product $j$, such that:

$$\pi_{ijt} = \pi(u_{ijt}, \{u_{i\sigma_{j}(k)t}\}_{k \in \mathcal{S}_t, k \neq j})$$ In this assumption, we state that the choice function for product $j$ is invariant to all permutations of its competitors. This implies that the individual's choice for product $j$ is not affected by the order or identity of the other products in the market, and it only depends on the set of competitors' characteristics.

Since researchers only observe aggregate data, we next define the aggregate demand function. In aggregate demand settings, individual-level choices are not observable and only aggregate demand is observable. It is often the case that the market-specific individual features are not observable and are assumed to be exogenously drawn from some distribution $\mathcal{F}(m_t)$, where $m_t$ represents the market-level characteristics. For the sake of notional simplicity, we let $m_t$ to be the same across all markets. One can easily incorporate market-specific user demographics in the choice function. Thus the demand of product $j$ in market $t$ denoted by $\pi_{jt}$ can be expressed as follows:

equation[equation omitted — 167 chars of source]

where $\mathcal{G} ( \varepsilon_{ijt})$ denotes the CDF of unobserved errors $\varepsilon_{ijt}$. Since $u_{ijt}$ is determined by $\{X_{jt}, p_{jt}, I_{it}, \varepsilon_{ijt}\}$ and $I_{it}, \varepsilon_{ijt}$ are integrated out in a market. Hence, we can express $\pi_{jt}$ as a function of only the observable product characteristics --

equation[equation omitted — 89 chars of source]
lemmaFor any choice function that satisfies Assumption (ref) and (ref), the aggregate demand function is permutation invariant.

This permutation invariance of the aggregate demand function exists because, under the exogeneity assumption, the aggregate demand function is simply the sum (or integral) of individual choice functions that are themselves invariant to permutation. Hence, changes to the order of competitors have no impact on the aggregated result. When the assumption of exogeneity is not satisfied, the aggregate demand function does not retain the permutation invariance, even though the individual-level choice function exhibits permutation invariance. We will return to this issue in $\S$(ref).

Our Assumptions (ref) (identity independence) and (ref) (permutation invariance) are fairly standard in the choice modeling literature, although they might not always be explicitly stated as such. Table (ref) summarizes models that satisfy these assumptions. Please see Web Appendix (ref) for detailed derivations of how these models satisfy permutation invariance.

table[table omitted — 874 chars of source]
restatable{thm}{mainthm} For any offer set $\mathcal{S}_t \subset \{1, 2, 3, \dots, J_t\}$, if a choice function $\pi: \{ u_{ijt}: j \in \mathcal{S}_t\} \rightarrow \mathbb{R}^{|\mathcal{S}_t|}$ where $u_{ijt}$ represents the index tuple $\{X_{jt}, p_{jt}, I_{it}, \varepsilon_{ijt}\}$ satisfies Assumption (ref), (ref) and (ref), then there exists suitable $\rho$, $\phi_1$ and $\phi_2$ such that $$\pi_{jt} = \rho (\phi_1 (X_{jt}, p_{jt}) + \sum_{k\neq j, k \in \mathcal{S}_t} \phi_2 (X_{kt}, p_{kt})),$$

Proof: See Web Appendix (ref).

This result is the generalization of the results shown in zaheer2017deep and can be shown following similar arguments. The above result is very powerful and has two important takeaways: (i) The input space of the choice function does not grow with the number of products in the assortment. Rather, the input space of the choice function (i.e., $\phi_1$ and $\phi_2$) grows only as a function of the number of features of the products in consideration, and (ii) the same transformations ($\rho$, $\phi_1$, and $\phi_2$) remain valid for all offer sets, denoted by $\mathcal{S}$, irrespective of their size. This property allows us to easily simulate the demand and entry of new products or changes in market structure, as one does with traditional parametric models. As an example, assuming $v_{jt}$ is the utility of product $j$ at market t, for the multinomial logit model one possible set of transformations could be $\phi_1(v_{jt}) =

bmatrix[bmatrix omitted — 31 chars of source]

$ and $\phi_2(v_{kt}) =

bmatrix[bmatrix omitted — 31 chars of source]

$ that generate two-dimensional vectors, and the function $\rho\left(

bmatrix[bmatrix omitted — 82 chars of source]

\right) = \frac{\phi_1(v_{jt})}{\phi_1(v_{jt}) + \sum_{k \neq j, k \in \mathcal{S}_t} \phi_2(v_{kt})}$ operates on these vectors. \footnote{$\rho\left(\phi_1(v_{jt}) + \sum_{k \neq j} \phi_2(v_{kt})\right) = \rho\left(

bmatrix[bmatrix omitted — 30 chars of source]

+

bmatrix[bmatrix omitted — 47 chars of source]

\right) = \rho\left(

bmatrix[bmatrix omitted — 56 chars of source]

\right) = \frac{exp(v_{jt})}{exp(v_{jt}) + \sum_{k \neq j} exp(v_{kt})}$}

Endogenous Covariates

In this section, we relax the exogeneity assumption and handle the potential endogeneity issue that is commonplace in demand settings. Note that, when the price (or other product characteristics or market-mix variables, such as promotions, correlate with unobserved variables ($\varepsilon_{ijt}$), Assumption (ref) (exogeneity) is compromised. As a result, it becomes infeasible to integrate out $\varepsilon_{ijt}$ in the aggregate demand function, as we did in Equation (ref). This means that the aggregate demand function loses its property of permutation invariance with respect to the observable characteristics of competitors. To address this, we build on the approach developed in petrin2010control to allow for endogenous observable features. Without loss of generality, we assume that price $p_{jt} \in \mathbb{C}$ is the endogenous variable and all other characteristics of the product $X_{jt} \in \mathbb{C}^{d}$ are exogenous variables. i.e., $$\mathbb{E}[p_{jt}\cdot \varepsilon_{ijt}] \ne 0 \quad\text{and}\quad \mathbb{E}[X_{jt}\cdot \varepsilon_{ijt}] = 0.$$

Given valid instruments $IV_{jt}$, we can express $p_{jt}$ as

equation[equation omitted — 73 chars of source]

At this point, no specific assumptions are made regarding the function $\gamma$. However, in the subsequent inference section, we will discuss that the estimator of $\gamma$ must be estimable at $n^{-1/2}$ in order to construct valid confidence intervals. Next, to address the issue of price endogeneity, we impose a mild restriction on the space of choice functions we consider.

comment\begin{assumption} [Linear Separability] We assume that the index tuple $u_{ijt}$ affects the individual-level choice function $\pi_{ijt}$ only through some scalar projection $\tilde{u}_{ijt} = f_i(X_{jt}, p_{jt}) + \varepsilon_{ijt} $ such that $$\pi(u_{ijt}, \{u_{ikt}\}_{k \ne j}) = \tilde{\pi} (\tilde{u}_{ijt}, \{\tilde{u}_{ikt}\}_{k \ne j}),$$ \end{assumption} The above assumption implies the demand function $\pi$ can be alternatively expressed as $\tilde{\pi}$ which is only the function of scalar projections ($\tilde{u}$) of index tuples.
assumption[Linear Separability] The unobserved product characteristics can be expressed as the sum of an endogenous ($\text{CF}$) and exogenous component \begin{equation} \varepsilon_{ijt}=CF\left(\mu_{jt} ; \lambda\right)+\tilde{\varepsilon}_{ijt}, \end{equation} where $\mathbb{E}[p_{jt}\cdot \tilde{\varepsilon}_{ijt} ] = 0$.

This assumption implies that, after controlling for $\mu_{jt}$ using the control function $CF$, the endogenous variable $p_{jt}$ is uncorrelated with the error term $\varepsilon_{ijt}$ in the model, thus it becomes exogenous. Then, we can re-write the index tuple $u_{ijt}$ as

equation[equation omitted — 107 chars of source]

such that $\mathbb{E}\Big[\tilde{\varepsilon}_{ijt}|(X_{jt}, p_{jt}, \mu_{jt}) \Big]=0$

restatable{thm}{endothm} For any offer set $\mathcal{S}_t \subset \{1, 2, 3, \dots, J_t\}$, if a choice function $\pi: \{ u_{ijt}: j \in \mathcal{S}_t\} \rightarrow \mathbb{R}^{|\mathcal{S}_t|}$ where $u_{ijt}$ represents the index tuple $\{X_{jt}, p_{jt}, I_{it}, \varepsilon_{ijt}\}$ satisfies Assumptions (ref) to (ref). Then under the condition of knowing the true function ($\gamma_0$) of $\gamma$, there exists suitable $\rho$, $\phi_1$ and $\phi_2$ such that $$\pi_{jt} = \rho (\phi_1 (X_{jt}, p_{jt}, \mu_{jt}(\gamma_0)) + \sum_{k\neq j, k \in \mathcal{S}} \phi_2 (X_{kt}, p_{jt}, \mu_{kt}(\gamma_0))),$$

The result follows straightforwardly from the observation that after controlling for $CF(\mu_{jt}; \lambda)$ the unobservable component $\tilde{\varepsilon}$ is exogenous. This implies the aggregate demand function is invariant under any permutation applied to the competitors of product $j$. The result demonstrates that endogeneity can be addressed by using the residuals from Equation (ref) as an additional set of features along with observable product characteristics.

Inference

This paper aims to estimate choice functions flexibly using non-parametric estimators. However, often in social science contexts, researchers and managers are also interested in conducting inference over some economic objects. Note that because non-parametric regression functions are estimated at a slower rate compared to parametric regressions, it is often infeasible to construct confidence intervals directly on the estimated $\hat{\pi}$. However, it is generally possible to perform inference and construct valid confidence intervals for specific economic objects that are functions of $\pi$. In this section, we will provide an example of one such important economic object and demonstrate how to construct valid confidence intervals for it. This will be done by leveraging the recent advances in automatic debiased machine learning as shown in the works of ichimura2022influence, chernozhukov2022automatic, chernozhukov2022locally, chernozhukov2021automatic, and others. However, unlike existing automatic debiased machine learning setups we also have to account for an additional first-stage estimator $\hat{\gamma}$.

In demand estimation, researchers are often interested in estimating the average effect of a price change on the demand for a product, as it can significantly influence market dynamics, pricing strategies, and regulatory decisions. To proceed with our analysis, let $w_{jt} = (y_{jt}, p_{jt}, X_{jt}, \{p_{kt},X_{kt}\}_{k \ne j})$ and $z_{jt} = (p_{jt}, X_{jt}, \{p_{kt},X_{kt}\}_{k \ne j})$ represent the variables associated with product $j$ in market $t$. Here, $p_{jt} \in \mathbb{C}$ denotes the observed prices, $X_{jt} \in \mathbb{C}^{d-1}$ represents other product characteristics, and $y_{jt} \in \mathbb{R}$ refers to the observed demand for product $j$ in market $t$, such as market shares or log shares. Note that either the observed price ($p_{jt}$) or other characteristics ($X_{jt}$) could be endogenous. For simplicity and without loss of generality, we focus on $p_{jt}$ as the endogenous variable in the following analysis.

The average effect of a price change\footnote{The expression for the average effect of a price change can be adapted to represent average price elasticity by placing the known and fixed value of $\Delta p_{jt}$ in the denominator.} can be expressed as the difference between the demand function $\pi_{jt}(\cdot;\gamma)$ evaluated at the original price $p_{jt}$ and at the price incremented by $\Delta p_{jt}$, given by the following expression:

align*[align* omitted — 176 chars of source]

The parameter of interest, $\theta_0$, is the expected value of this price change effect over the true population distribution\footnote{We assume the data reflects the true population.} of $w_{jt}$, which can be calculated as:

align*[align* omitted — 195 chars of source]

In summary, the average effect of a price change on demand, denoted by $\theta_0$, is calculated by evaluating the difference between the demand function at the original price and at the price incremented by $\Delta p_{jt}$, and then computing the expected value of this difference.

In practice, we estimate $\theta_0$ by computing its empirical analog using the estimated demand function $\hat{\pi}$ and first-stage estimator $\hat{\gamma}$, i.e.,

equation[equation omitted — 101 chars of source]

where $n$ is the number of observations. When parametric methods are employed to estimate $\hat{\pi}$ and $\hat{\gamma}$, the estimator for $\hat{\theta}$ is generally $\sqrt{n}$-consistent, assuming that the model is correctly specified. However, $\sqrt{n}$-consistency may not hold when non-parametric estimators are used, particularly if the first-order bias does not vanish at a rate of $\sqrt{n}$. Irrespective of the method used to estimate $\pi$, this is often the case, as flexible estimation of $\pi$ always requires some form of regularization and/or model selection. Debiasing techniques are required to mitigate the effects of regularization and/or model selection when learning flexible demand models. These approaches can help improve the performance of the estimator and facilitate valid inference with $\hat{\theta}$. We therefore adapt recent debiasing techniques developed in recent automatic debiased machine learning literature (see chernozhukov2022automatic). Specifically, we will focus on problems where there exists a square-integrable random variable $\alpha_0(z)$ such that $\forall$ $||\gamma - \gamma_0||$ small enough --

align*[align* omitted — 197 chars of source]

By the Riesz representation theorem, the existence of such $\alpha_0(z_{jt})$ is equivalent to $\mathbb{E}[m(w_{jt},\pi(z_{jt};\gamma))] $ being a mean square continuous functional of $\pi(z_{jt};\gamma)$. Henceforth, we refer to $\alpha_0(z)$ as Riesz representer (or RR). newey1994asymptotic shows that the mean square continuity of $\mathbb{E}[m(w_{jt},\pi_{jt}(z_{jt};\gamma))]$ is equivalent to the semiparametric efficiency bound of $\theta_0$ being finite. Thus, our approach focuses on regular functionals. Similar uses of the Riesz representation theorem can be found in ai2007estimation, ackerberg2014asymptotic, hirshberg2020debiased, and chernozhukov2022automatic among others. The debiasing term in this case takes the form $\alpha(z_{jt})(y_{jt}-\pi(z_{jt};\gamma))$. To see that, consider the score $m(w_{jt},\pi(z_{jt};\gamma)) +\alpha(z_{jt})(y_{jt}-\pi(z_{jt};\gamma))-\theta_0$. It satisfies the following mixed bias property: $$

aligned\mathbb{E}[m(w_{jt},\pi(z_{jt};\gamma)) & +\alpha(z_{jt})(y_{jt}-\pi(z_{jt};\gamma))-\theta_0] \\ & =-\mathbb{E}\left[\left(\alpha(z_{jt})-\alpha_0(z_{jt})\right)\left(\pi(z_{jt})-y_{jt}\right)\right] .

$$ This property implies double robustness \citep{robins1994estimation, funk2011doubly} of the score. That is, if either $ \alpha(z_{jt}) $ is correctly estimated, which would mean $ \alpha(z_{jt}) - \alpha_0(z_{jt}) = 0 $, or $ \pi(z_{jt}) $ is correctly estimated, implying $ \pi(z_{jt}) - y_{jt} = 0 $, then the term $ (\alpha(z_{jt}) - \alpha_0(z_{jt})) (\pi(z_{jt}) - y_{jt}) $ will be zero. This results in the score going to zero, thereby making the estimator consistent for $ \theta_0 $. A debiased machine learning estimator of $\theta_0$ can be constructed from this score and first-stage learners $\widehat{\pi}$ and $\widehat{\alpha}$. Let $\mathbb{E}_n[\cdot]$ denote the empirical expectation over a sample of size $n$, i.e., $\mathbb{E}_n[x_i]=\frac{1}{n} \sum_{i=1}^n x_i$. We consider: $$ \widehat{\theta}=\mathbb{E}_n[m(w_{jt}; \widehat{\pi})+\widehat{\alpha}(z_{jt})(y_{jt}-\widehat{\pi}(z_{jt}))] . $$ The mixed bias property implies that the bias of this estimator will vanish at a rate equal to the product of the mean-square convergence rates of $\widehat{\alpha}$ and $\widehat{\pi}$. Therefore, in cases where the demand function $\pi$ can be estimated very well, the rate requirements on $\widehat{\alpha}$ will be less strict, and vice versa. More notably, whenever the product of the mean-square convergence rates of $\widehat{\alpha}$ and $\widehat{f}$ is larger than $\sqrt{n}$, we have that $\sqrt{n}\left(\widehat{\theta}-\theta_0\right)$ converges in distribution to centered normal law $N\left(0, \mathbb{E}\left[\psi_0(w_{jt})^2\right]\right)$, where $$ \psi_0(w_{jt}):=m\left(w_{jt} ; \pi_0\right)+\alpha_0(z_{jt})\left(y_{jt}-\pi_0(z_{jt})\right)-\theta_0, $$ as proven formally in Theorem 3 of \cite{chernozhukov2022automatic}. Results in \cite{newey1994asymptotic} imply that $\mathbb{E}\left[\psi_0(w_i)^2\right]$ is a semiparametric efficient variance bound for $\theta_0$, and therefore the estimator achieves this bound.

restatable{thm}{mainthm3} [chernozhukov2021automatic] One can view the Riesz representer as the minimizer of the loss function: $$ \begin{aligned} \alpha_0 & =\underset{\alpha}{\arg \min } \mathbb{E}\left[\left(\alpha(z_{jt})-\alpha_0(z_{jt})\right)^2\right] \\ & =\underset{\alpha}{\arg \min } \mathbb{E}\left[\alpha(z_{jt})^2-2 \alpha_0(z_{jt}) \alpha(z_{jt})+\alpha_0(z_{jt})^2\right] \\ & =\underset{\alpha}{\arg \min } \mathbb{E}\left[\alpha(z_{jt})^2-2 m(w_{jt} ; \alpha)\right], \end{aligned} $$

In our earlier discussions, we employed the moment function of $\pi$, whereas in Theorem (ref), we focus on the moment function of $\alpha$. This shift is justified by the Riesz Representation Theorem, which implies $\mathbb{E}[m(w_{jt}; \pi)] = \mathbb{E} [\alpha_0(z_{jt}) \pi(z_{jt})]$. Given that $\pi$ can represent any function, substituting $\alpha$ for $\pi$ is permissible, thereby validating the transition from the second to the third line in Theorem (ref). We use the above theorem to flexibly estimate the RR. The advantage of this approach is that it eliminates the need to derive an analytical form for the RR estimator, allowing it to be addressed as a simple computational optimization problem.

restatable{thm}{mainthm4} [chernozhukov2021automatic] Let $\delta_n$ be an upper bound on the critical radius (wainwright2019high) of the function spaces:

$$

gathered\left\{z \mapsto \zeta\left(\alpha(z)-\alpha_0(z)\right): \alpha \in \mathcal{A}_n, \zeta \in[0,1]\right\} and \\ \left\{w \mapsto \zeta\left(m(w ; \alpha)-m\left(w ; \alpha_0\right)\right): \alpha \in \mathcal{A}_n, \zeta \in[0,1]\right\}

$$ and suppose that for all $f$ in the spaces above: $\|f\|_{\infty} \leq 1$. Suppose, furthermore, that $m$ satisfies the mean-squared continuity property: $$ \mathbb{E}\left[\left(m(w ; \alpha)-m\left(w ; \alpha^{\prime}\right)\right)^2\right] \leq M\left\|\alpha-\alpha^{\prime}\right\|_2^2 $$ for all $\alpha, \alpha^{\prime} \in \mathcal{A}_n$ and some $M \geq 1$. Then for some universal constant $C$, we have that w.p. $1-\zeta$ : $$

aligned\left\|\widehat{\alpha}-\alpha_0\right\|_2^2 \leq C( & \delta_n^2 M+\frac{M \log (1 / \zeta)}{n} \\ & \left.+\inf _{\alpha_* \in \mathcal{A}_n}\left\|\alpha_*-\alpha_0\right\|_2^2\right)

$$ The critical radius has been widely studied in various function spaces, such as high-dimensional linear functions, neural networks, and superficial regression trees, often showing $\delta_n = O\left(d_n n^{-1 / 2}\right)$, where $d_n$ represents the effective dimensions of the hypothesis spaces (chernozhukov2021automatic). In our research, we focus on applying Theorem 3 from an application standpoint to neural networks.

To that end, we make the following assumptions.

assumption(i) $\alpha_0(z)$ is bounded, (ii) $\forall$ $||\gamma - \gamma_0||$ small enough, $\mathbb{E}[(y-\pi_0(z_{jt}; \gamma))^2|z_{jt}]$ is bounded, and (iii) $\mathbb{E}[m(w_{jt}, \pi_0(z_{jt};\gamma_0))^2] < \infty$.

These assumptions are standard regularity conditions used in the automatic machine learning literature.

assumptioni) $\forall$ $||\gamma - \gamma_0||$ small enough $||\hat{\pi}(;\gamma) - \pi_0(;\gamma)|| \xrightarrow[]{p} 0$ and $ ||\hat{\alpha} - \alpha_0|| \xrightarrow[]{p} 0$; ii) $\sqrt{n}||\hat{\alpha} - \alpha||||(\hat{\pi}(;\gamma) - \pi_0(;\gamma)|| \xrightarrow[]{p} 0$; iii) $\hat{\alpha}$ is bounded; (iv) $\sqrt{n}||\hat{\gamma}-\gamma_0||\xrightarrow[]{p} 0$

Intuitively these assumptions mean that (i) the estimator of both $\pi$ and $\alpha$ should be consistent for values of $\gamma$ in a close enough neighborhood of $\gamma_0$. Further, it requires that the product of mean square error of $\hat{\alpha}$ and mean square error of $\pi$ should vanish at $\sqrt{n}-$ rate. This can be achieved if both these terms converge at least at $n^{-1/4}$ rate. Finally, we also assume that the first stage estimator $\hat{\gamma}$ is estimable at $n^{-1/2}$ rate. This limits the class of functions one can use to estimate $\gamma$.

assumption$m(w, \pi)$ is linear in $\pi$ and there is $C>0$ such that $$ \left|E\left[m(w, \pi)-\theta_{0}+\alpha_{0}(z) (y-\pi(z;\gamma))\right]\right| \leq C\left\|\pi-\pi_{0}\right\|^{2} $$
comment\begin{corollary} The Riesz representer $\alpha(p_{jt}, X_{jt}, \{p_{kt},X_{kt}\}_{k \ne j})$ is permutation invariant under any permutation function $\sigma_j$ that rearranges the indices of the competitors of product $j$, such that $$\alpha(p_{jt}, X_{jt}, \{p_{kt},X_{kt}\}_{k \ne j}) = \alpha(p_{jt}, X_{jt}, \{p_{\sigma_{j}(k)t},X_{\sigma_{j}(k)t}\}_{k \ne j})$$ \end{corollary} This follows directly from the Riesz representation theorem and assumption (ref).
restatable{prop}{mainprop} If Assumptions 5-7 are satisfied then for $V = E[\{m(w, \pi_0(z;\gamma_0)) - \theta_0 $\\ $+ \alpha_0(z)(y-\pi_0(z;\gamma_0))\}^2]$, \begin{equation*} \sqrt{n}(\hat{\theta} - \theta_0) \xrightarrow[]{D} N(0, V ), \hat{V} \xrightarrow[]{p} V. \end{equation*}

We show the proof in Web Appendix (ref). This theorem shows that if $\hat{\gamma}$ is estimable at a fast enough rate one can still construct valid confidence intervals for $\hat{\theta}$. This result can be shown following similar arguments as in chernozhukov2022locally. Finally, we note that while the above arguments focus on the estimation of the average effect of a price change on demand, we can follow the same arguments to derive inference results for other economic quantities of interest, e.g., the effect of changing some product features on demand.

comment\subsection{Summary of Theory}

Estimation Procedure

Based on theoretical results presented above, we now outline an estimation procedure for both the choice function ($\pi$) and the average effects of price changes ($\theta$).

Consider a dataset, where $\{y_t, z_t, IV_t \}_{t=1}^{n}$ are independently and identically distributed. Here, $y_t$ is the vector of market shares in market $t$, $z_t$ is $J \times (d+1)$ matrix of product features and $IV_t$ is the $J $-dimension vector of instrumental variables.

itemize• Stage 0 (Data partition): We randomly split the observed markets into L folds such that the data $D_l := \{y_t, z_t \}_{t \in l}$, where $l$ denotes the $l^{th}$ partition. Note that all the observations for one market are always in one fold. • Stage 1 (Estimate $\hat{\gamma}$): For each fold $l$, we estimate $\gamma_l$ by regressing the endogenous variable on the exogenous instruments on the left out data $D_{l}^c := \{y_t, z_t \}_{t \notin l}$. We then use the cross-fitting technique, same as chernozhukov2021automatic, to calculate the residual $\hat{\mu}_l$ of fold $l$ with estimated $\hat{\gamma}_l$ on $D_{l}^c$. • Stage 2 (Estimate $\hat{\pi}$ and $\hat{\theta}$): \begin{itemize} • Stage 2a (Estimation): In the second stage, for each fold $l$, we estimate both the choice function ($\hat{\pi}$) and the Riesz estimator ($\hat{\alpha}$) on the left out data $D_{l}^c := \{y_t, z_t \}_{t \notin l}$ \begin{equation} \hat{\pi}_l=\operatorname*{arg\,min}_{\pi \in \mathcal{F}} \frac{1}{\sum_{t \in D_{l}^c} J_t}\sum_{t \in D_{l}^c}\sum_{j \in J_t}[(y_{jt}-\pi(z_{jt};\hat{\gamma})) ^2] \end{equation} \begin{equation} \hat{\alpha}_l = \operatorname*{arg\,min}_{\alpha \in \mathcal{A}} \frac{1}{\sum_{t \in D_{l}^c} J_t}\sum_{t \in D_{l}^c}\sum_{j \in J_t} \left[\alpha(z_{jt})^2-2 m(w_{jt} ; \alpha)\right]. \end{equation} Based on Theorem (ref), instead of directly estimating the function $\pi$, we decompose the estimation into three sub-components: $\rho, \phi_1$, and $\phi_2$. Specifically, for each component of our model ($\phi_1$, $\phi_2$, and $\rho$), we use a standard 3-layer neural network\footnote{For $\phi_1$ and $\phi_2$, each of the three layers consists of 64 neurons, with the output vector also featuring 64 neurons. For $\rho$, the layers are configured with 300, 100, and 64 neurons for the first, second, and third layers, respectively.}, and this is implemented without further hyperparameter tuning. We implement ReLU activation function at each layer as it is standard in feedforward designs due to simplicity and computation efficiency in gradients. Figure (ref) illustrates how we pass the data to the neural networks. Specifically, we pass the focal product's characteristics (price $p_{jt}$ and other product features $X_{jt}$) and the residuals ($\hat{\mu}_{jt}$) estimated from the first stage regression to the $\phi_1$. In parallel, we pass all the products' characteristics of the other products in the same market ($p_{jt}$ and $X_{kt}$) and the corresponding residuals ($\hat{\mu}_{kt}$) to the same $\phi_2$, and then sum the output up. The output of $\phi_1$ and $\phi_2$ have the same data structure (e.g., a 64-dimension vector). Next, we pass the summation of the output of $\phi_1$ and $\phi_2$ to a third neural network $\rho$. The output of $\rho$ is a scalar which represents the market share of the focal product $jt$. We use the same neural network structure (with the three sub-components same as $\phi_1$, $\phi_2$, and $\rho$) to estimate $\alpha$. The only difference is that the loss function of $\alpha$ is not based on the difference between the observed and the predicted market share as in Equation (ref). Instead, the loss function is based on the squared difference between $\alpha$ and the moment function of $\alpha$ as stated in Equation (ref) and Theorem (ref). \begin{figure} \caption{Illustration of Neural Network Architecture} \end{figure} • Stage 2b (Cross-fitting): Now we again use the cross-fitting technique to reduce the bias when estimating $\hat{\theta}$. Specifically, we use the estimators ($\hat{\pi}$ and $\hat{\alpha}$) estimated on $D_{l}^{c}$ to estimate the $\hat{\theta}_l$ of $l$. By applying cross-fitting, we ensure that the nuisance functions and the parameters are estimated on separate, non-overlapping datasets. This approach diminishes the risk of overfitting and enhances the robustness of our estimation. And finally, to estimate $\hat{\theta}$, we randomly select one observation $t^*$ in each market $t$ and average it out across all folds. Thus the estimator for $\theta_0$ and its variance can be given as follows: \begin{comment} \begin{eqnarray} \hat{\theta} &=& \frac{1}{n} \sum_{\ell=1}^{L} \sum_{t \in D_{\ell}^c}\left\{m\left(w_{t}, \hat{\pi}_{\ell}\right)+\hat{\alpha}_{\ell}\left(z_{t}\right) \left(y_t-\hat{\pi}_{\ell}(z_{t};\hat{\gamma})\right)\right\} \\ \hat{V} &=& \frac{1}{n} \sum_{\ell=1}^{L} \sum_{t \in D_{\ell}^c} \hat{\psi}_{t \ell}^{2}, \quad \hat{\psi}_{t \ell}=m\left(w_{t}, \hat{\pi}_{\ell}\right)-\hat{\theta}+\hat{\alpha}_{\ell}\left(z_{t}\right) \left(y_t-\hat{\pi}_{\ell}(z_{t};\hat{\gamma})\right) \end{eqnarray} \end{comment} \begin{eqnarray} \hat{\theta} &=& \frac{1}{n} \sum_{\ell=1}^{L} \sum_{t \in D_{\ell}^c}\left\{m\left(w_{t^*}, \hat{\pi}_{\ell}\right)+\hat{\alpha}_{\ell}\left(z_{t^*}\right) \left(y_{t^*}-\hat{\pi}_{\ell}(z_{t^*};\hat{\gamma})\right)\right\} \\ \hat{V} &=& \frac{1}{n} \sum_{\ell=1}^{L} \sum_{t \in D_{\ell}^c} \hat{\psi}_{t^* \ell}^{2}, \quad \hat{\psi}_{t^* \ell}=m\left(w_{t^*}, \hat{\pi}_{\ell}\right)-\hat{\theta}+\hat{\alpha}_{\ell}\left(z_{t^*}\right) \left(y_{t^*}-\hat{\pi}_{\ell}(z_{t^*};\hat{\gamma})\right) \end{eqnarray} \end{itemize}

Numerical Experiments

We now present a series of simulation studies that establish the numerical performance of our approach. First, in $\S$(ref), we examine the predictive performance of our model on a series of models, including stylized discrete choice models with linear utilities as well as more general models that allow non-linear utilities and realistic consumer behaviors such as inattention. Next, in $\S$(ref), we present numerical experiments that demonstrate our approach's ability to simulate counterfactuals. Finally, in $\S$(ref), we demonstrate the applicability of the inference procedure proposed in $\S$(ref).

Predictive Performance

In this section, we show that our approach can recover demand and elasticity estimates for a wide variety of settings without the knowledge of the underlying choice model and/or making parametric assumptions on consumers' behaviors. Before describing the simuations, we first describe the metrics used for comparing the performance of different estimators and the benchmark estimation approaches used.

First, to assess the predictive performance of our approach, we focus on three quantities of interest:

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • Market share ($\hat{\pi}_{jt}$) • Own price elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$) • Cross-price elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{k\neq j t}/ p_{k\neq j t} }$)

In each simulation, we compare the performance of our model with the predictive performance of four baseline models:

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • Multinomial Logit model (MNL) • Random Coefficient Logit model (RCL) • A standard neural network-based non-parametric method (NP). Here we simply use all the features of all the products in the market as input and predict a market share vector without using the permutation invariance property or correcting for endogeneity. This is similar to the approaches used by gabel2022product and cai2022deep, which simply use a large neural network for demand prediction. For this baseline, we tune the hyperparameters of the neural network, including the number of layers, number of nodes in each layer, learning rate, and the number of epochs using 5-fold cross-validation for each data generation. We detail the space of hyperparameters in Web Appendix (ref). We also apply the ReLU activation for each layer. • A Mean Predictor (MP). Here we predict a uniform market share for all observed products within a market, excluding the outside option. This predicted share is set to the average of all observed market shares in that market\footnote{For instance, consider a market with three observed products having market shares of 0.2, 0.3, and 0.4, respectively. This implies the outside option holds a market share of 0.1. In such a scenario, the MP would predict the market share for each of the three products to be 0.3, which is the average ((0.2 + 0.3 + 0.4) / 3).}. This serves as a baseline estimator as it does not account for individual product characteristics, providing a benchmark for the simplest prediction scenario. Additionally, comparing the performance of other models with MP also enables us to quantify the natural decrease in the MAE and RMSE when the number of products increases because the magnitude of the market shares decreases as the number of products increases\footnote{For instance, in a market with 100 products, the MAE and RMSE for any estimator are expected to be quite low. Comparing with the MP in such situations helps establish a baseline for MAE or RMSE.}.

Multinomial Logit and Random Coefficient Logit with Linear Utility

We first consider the two standard Data Generating Processes (DGPs) used in the demand estimation literature that use linear utility-based choice models -- (1) Multinomial Logit model, and (2) the Random Coefficient Logit model. For both cases, we consider a setting with 10 products ($J = 10$), 100 markets ($T = 100$), one price feature, and 10 non-price features ($d = 10$). We define the utility $u_{ijt}$ that consumer $i$ in market $t$ derives from product $j$ as the following linear function:

equation[equation omitted — 104 chars of source]

where $\varepsilon_{ijt}$ represents an independently and identically distributed (iid) Type-I extreme value across products and consumers. $X_{jt} \in \mathbb{C}^{d}$ denotes the non-price features of the product. $\alpha_i, \beta_{i}$ are the model coefficients, which are kept constant for all consumers in the MNL, while in the RCL, they are normally distributed across consumers. The probability distribution of features and coefficients used are shown in Web Appendix (ref). Also, the mean utility from the outside option is normalized to 0.

We denote the market share of product $j$ in market $t$ generated from MNL by $\pi_{jt}^{MNL}$ and the market share generated from RCL by $ \pi_{jt}^{RCL}$. For each market, we generate the market shares of each product by simulating $N = 10,000$ individual choices and aggregating by each market as shown below.

equation[equation omitted — 106 chars of source]
equation[equation omitted — 209 chars of source]

Note that for MNL, instead of simulating each individual's choice probability, we simulate each individual's choice based on the utility maximization principle. This approach ensures that when we use MNL (true model) for estimation, it does not reproduce the data perfectly.

table[table omitted — 6,634 chars of source]

For each DGP, we split the generated data into training data (80%) and test data (20%). We use the training data for estimation (both our model and the benchmark models described above).\footnote{In our simulations, we train both our model and NP with the log market shares ($log(\pi_{jt})$). The performance metrics reported for predicted market shares are computed based on the exponential values of the predicted log market shares, bringing these metrics back to market shares ($\pi_{jt}$). The performance metrics reported for elasticities are computed based on the relative change of the predicted log market shares ($\partial log(\pi_{jt})$) divided by the percentage change of the price ($\partial p_{jt}/p_{jt}$ for own-elasticity and $\partial p_{kt}/p_{kt}$ for cross-elasticity), which is equivalent to the elasticity calculated directly using the market share ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$ for own-elasticity and $\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{k\neq j t}/ p_{k\neq j t} }$ for cross-elasticity).} For the predicted market share, we present all the model results and comparisons on the test data. For the predicted own- and cross-elasticities, we present all the model results and comparisons on the training data.\footnote{The reason we only use test data to report predicted market share accuracy is to demonstrate the model's predictive performance on unseen data. In contrast, we use training data to report accuracy in elasticities to mimic the real empirical setting where we use full data to estimate elasticity.}

Tables (ref) shows the Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) in the predicted market share ($\hat{\pi}_{jt}$) for our approach as well as the baseline models.We see that when the true model is MNL, our model cannot beat RCL or MNL, which is as expected; but the error of our model is quite close to the true model. When the true model is RCL, our model can beat MNL consistently and the performance of our model is also close to the true model. Importantly, we find that our model consistently outperforms the benchmark Non-Parametric (NP) method in all data generation processes. This is despite extensive hyperparameter tuning for the NP method.

There are two key reasons why our approach outperforms the standard neural network-based non-parametric method, especially as the number of products increases. First, unlike the standard neural network, our method can circumvent the curse of dimensionality that arises with the increase in the number of products. The standard neural network uses the stacked product features (d-dimension non-price features $X_{jt}$ and price $p_{jt}$) as input and has $(J \times (d +1) ) \times h_1$ parameters in the input layer, where $h_1$ denotes the size of the first hidden layer.\footnote{In this section, we consider only cases where there are no correlated unobservables. When there are potential endogeneity concerns, then we can include a residual $\mu_{jt}$ estimated from a first-stage regression, and then size of the input layer becomes $J \times (d+2)$. We consider settings with endogeneity in Web Appendix $\S$(ref) and in the experiments on inference in $\S$(ref).} In contrast, our model only uses product features as the input (with dimension $d+1$); see Figure (ref). Therefore, the parameters for our model do not scale with the number of products. Thus, as the number of products increases, our method is able to exploit this information to improve its performance, whereas the standard NP is unable to do so. Second, our model can leverage product-level market-share data more effectively. Note that one observation in the standard NN consists of one market, whereas one observation for our method consists of one product in a market. Hence, given data on $T$ markets, the number of samples available for the NP method is $T$, whereas the sample available for our model is $T \times J$. Together, these strengths of our approach lead to significantly better performance compared to a naive neural network.

We observe both sources of performance improvement in the numerical simulation results in Table (ref). As we vary the number of products (5, 10, and 20), the MAE of the predicted market shares from our model decreases monotonically (see simulation numbers 0, 2, and 4 for MNL, and 5, 7, and 9 for RCL). In contrast, the performance of the benchmark non-parametric estimator deteriorates as the number of products increases and becomes even worse than the Mean Prediction for the 20 product cases (simulation numbers 4 and 9). These findings demonstrate the overfitting problems and the curse of dimensionality issues discussed above. Indeed, this limitation has also been theoretically established by sannai2019improved, who showed that for neural networks that do not take into account the inherent invariance structure, the generalization gap increases in proportion to the number of possible permutations, which is $\sqrt{J!}$ in our case. Further, we examine the performance of our model and the other benchmarks by varying the number of markets (20, 100, and 200) while keeping the number of products constant at 10; see simulations 1, 2, and 3 for MNL and 6, 7, and 8 for RCL. Although both our model and the NP method show improved performance with more markets, the non-parametric estimator is more adversely affected by a decrease in market numbers due to a more significant reduction in its sample size. This is particularly problematic in scenarios with one market, as the NP method becomes infeasible for estimation for only one sample.

Finally, in Tables (ref) and (ref), we show the predictive performance of own-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$) and cross-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{k\neq j t}/ p_{k\neq j t} }$), respectively. Again, we find that our model consistently outperforms the NP method in all scenarios for both own- and cross-elasticity predictions. When the true model is MNL, our model underperforms compared to RCL, given that RCL inherently captures the substitution pattern among products in MNL\footnote{In an extreme case, when the variance of random coefficients is zero, RCL is equivalent to MNL.}. When the true model is RCL, our model is the closest one to the true model. Unlike the market share predictions, the accuracy of own-elasticity predictions does not exhibit a clear monotonical improvement as the number of products increases. We observe a similar pattern even for the true model--the accuracy of own-elasticity decreases as the number of products increases. This suggests that it is not a deficiency of our model but due to the inherent complexities in estimating own-elasticities in markets with many products. In the prediction of cross-elasticity, our model shows a monotonical improvement in accuracy with an increase in the number of products. Additionally, for both own- and cross-elasticities, as the number of markets increases, the performance of our model is better due to the increase in the sample size.

So far, in the above simulations, we did not consider any endogenous explanatory variables. As discussed in $\S$(ref), our approach can easily account for endogeneity following the estimation steps in $\S$(ref). For interested readers, we present a set of numerical experiments with endogeneous explanatory variables in Web Appendix $\S$(ref). The key takeaway from this analysis is that ignoring endogeneity can lead to significant biases in the estimates of own- and cross-price elasticities. Therefore, in the application to real data, we take care of endogeneity carefully; please see $\S$(ref) for further details.

Random Coefficient Logit with Non-Linear Utility

In the previous section, we focused on linear utility specifications and standard choice behaviors. Recent literature has highlighted that the oversight of non-linear relationships between features and utilities can introduce biases in the estimates allenby2004choice. Conversely, non-parametric estimators are adept at capturing these non-linear patterns directly from the data. As a result, there has been a growing trend towards the adoption of non-linear utility functions. Therefore, we now focus on data generated from a random coefficient logit model with non-linear transformations applied to observable features. We consider a case with two product features -- price and a non-price feature $x$. We apply a non-linear transformation $g(x)$. Following bakhitov2022causal,we consider two functions for $g(x)$:

enumerate• log(): $g(x) = log(|16x - 8| + 1)\text{sign}(x - 0.5)$ • sin(): $g(x) = sin(x)$

The log transformation is common in empirical studies, to capture a diminishing sensitivity of a feature on the market share. The sine transformation captures periodic or cyclical effects. For example, when the feature represents the time of year, normalized from 0 to 1, then using sine transformation can effectively capture the seasonal variations in consumer preferences.

The utility that consumer $i$ in market $t$ derives from product $j$ then has the following non-linear form:

equation[equation omitted — 84 chars of source]

The marketshares then follow a similar structure to that from Equation (ref).

As before, we generate data using this model and estimate the marketshares and elasticities using both our approach and the baseline models. When estimating the baseline MNL and RCL models, we assume that the researcher does not have knowledge of the non-linearities in the utility function and hence uses the simple linear utility(s) in their estimation (as shown in Equation (ref)). The results for the predicted market shares ($\hat{\pi}_{jt}$), own-price elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$) and cross-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{k\neq j t}/ p_{k\neq j t} }$) are presented in Tables (ref), (ref), and (ref), respectively.

Regarding the MAE of predicted market shares, our model surpasses the RCL model by a factor of 8X and 4X across transformations (a) and (b), respectively. Similarly, considering the MAE of predicted own-elasticity in transformations (a) and (b), our model outperforms RCL by factors of 20X and 2.5X, respectively. For the MAE of predicted cross-elasticity, our model is 2X and 1.5X superior to RCL across transformations (a) and (b), respectively. It's worth noting that while our model consistently outperforms the NP method across metrics, the NP method still shows better performance than both RCL and MNL in terms of MAE and RMSE for estimated own-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$), underscoring the strengths of neural network-driven approaches in navigating non-linearities.

table[table omitted — 2,590 chars of source]

Models with Consumer Inattention and Consideration Set Formation

Finally, we consider a scenario where consumers do not pay attention to all the products and/or are not fully informed of all the alternatives in the choice set. Recent literature has shown that this is often the case in many empirical settings goeree2008limited, gabaix2019behavioral, honka2019empirical, abaluck2020method, compiani2022market. However, such cases violate a standard assumption of the choice model: that consumers are informed and consider all options when they make purchase decisions. In some parametric models van2010retrieving, this issue is managed by constructing a consumer-level consideration set. However, consideration sets are usually unobserved in data; so these approaches often require assumptions on how consideration sets are formed, which might not always be appropriate or reflective of actual consumer behavior. Another way to manage this issue is to model search costs, i.e., allow consumers to ignore certain products because search is costly weitzman1978optimal, mehta2003price, hortaccsu2004product, kim2010online. Similarly, it also requires researchers to specify how search cost enters the utility function and decision process. In contrast to these models, our approach refrains from making any parametric assumptions, which allows for a potentially more flexible representation of consumer behavior in the case of consumer attention. To demonstrate how our model can capture the inattentive behavior, we look at a scenario where consumers are inattentive and deviate from the traditional random coefficient logit model.

figure[figure omitted — 1,175 chars of source]

We consider a simple model of inattention where a portion of consumers ($1 - \frac{1}{1+p_{ht}}$) ignore the product with the highest price, following compiani2022market. In other words, the consideration set of $1 - \frac{1}{1+p_{ht}}$ consumers in market $t$ excludes the highest-price product $h$. So when the price increases, the portion of inattentive consumers increases. Suppose there is only one feature, price, then, the choice probability of the highest price product $h$ in the market $t$ is: \[ \frac{1}{1+p_{h t}}\frac{exp(\alpha_i p_{ht} )}{ \sum_k exp(\alpha_i p_{kt} )}. \] The choice probability of other products $j \neq h$ are given by: \[ \frac{1}{1+p_{h t}}\frac{exp(\alpha_i p_{j t} )}{ \sum_k exp(\alpha_i p_{kt}) } + (1- \frac{1}{1+p_{h t}}) \frac{exp(\alpha_i p_{jt})}{ \sum_{k \neq h t } exp(\alpha_i p_{kt} )}. \]

table[table omitted — 3,044 chars of source]

We first present detailed results on the own- and cross-elasticity for the two-product case in Figure (ref). We consider the number of markets to be 1,000 so that we can observe enough variance in our data. Figure (ref) shows how the estimated own-elasticity for the highest-priced product (which is ignored by a subset of consumers), for a range of prices. Note that in our model, when the price is higher, the portion of inattentive consumers is higher. Thus when we change the price, the change in market share is smaller than the case without inattention. While our model and the fully NP model are able to capture this pattern, both the parametric models (MNL and RCL) are unable to do so. Figure (ref) shows how the estimated cross-elasticity of the other products vary with the price of the highest-priced product. Similarly, due to the ignorance of inattention, both MNL and RCL overestimate the magnitude of the elasticity of the other product. In contrast, our model and the fully NP are able to capture the true cross-price elasticity and are close to the true model.

Next, we show a more comprehensive set of results for all three metrics (market-shares, own-, and cross-price elasticities) when there are more products (2, 5, and 10) and fewer markets (100) in Table (ref), (ref), and (ref). We find that our approach consistently outperforms RCL and MNL, as expected. Further, the performance of our model improves as the number of products increases while that of the NP model monotonically decreases with the number of products (for the reasons discussed in $\S$(ref)).

In summary, we find that our approach adapts well even as the underlying model of consumer behavior changes without the need to impose any specific assumptions on consumer decision-making.

Counterfactual Analysis

In general, a key advantage of parametric models like MNL and RCL, or a structural approach to consumer behavior is their ability to predict outcomes in counterfactual scenarios outside the distribution of the data used for estimation. In contrast, fully non-parametric approaches like the naive neural network model that simply uses all product features as inputs cannot be used for counterfactual predictions\footnote{Both our model and fully NP can do counterfactuals when shocks only result in the change in features $\{X_{jt}, P_{jt}\}$. For example, consider a choice model where ranking is a feature. In such a scenario, researchers would like to see how the demand would change if the ranking policy is changed. To estimate this counterfactual scenario, one would update the value assigned to the ranking feature within a model. Therefore, both our model and fully NP are capable of doing such analysis, offering insights into how changes in specific features could impact demand.}. By leveraging the choice invariance property, our method adds some structure to the neural network architecture; and as a result, accommodates certain types of counterfactual models. In particular, we focus on a specific type of counterfactual that is of interest to firms and policy-makers -- demand estimation when the choice set changes; e.g., through the introduction of a new product, the exit of an existing product, or the merger of two firms nevo2000mergers, petrin2002quantifying, nevo2003new, wollmann2018trucks. Our model can easily handle such counterfactuals since the choice function is specified as a function of the focal product features and the features of a set of competing products (see Theorem (ref) and Figure (ref)). In contrast, estimating counterfactual demand when the choice set changes is infeasible with a standard neural network estimator due to its structural constraints on the input space. A change in the choice set would result in the change of size of the input vector, making such estimation infeasible.

To showcase the capability of our model to estimate counterfactuals, we consider a setting where a new product is introduced to the market. For comparison, we will only consider an MNL and RCL estimator since it is infeasible for the NP method to estimate the counterfactual. We use two data generation processes -- Multinomial Logit (MNL) and Random Coefficient Logit (RCL), the same as $\S$(ref) with 100 markets, 10 products (and one outside option)m and 10 features. For each of these two cases, we consider a counterfactual where an 11th product is introduced to each market. The observable characteristics of the new product are simulated from the same distribution as other products.

In Table (ref), we present the error in the estimated market share of all products when a new product is introduced to the market. We find that our model does quite well. When the true model is MNL, both MNL and RCL outperform our method, similar to what we find in $\S$(ref). Our model also outperforms the MNL when the underlying data generation process is RCL, and produces results comparable to the true model. Overall, these results suggest that our approach can be used for counterfactual estimation, which is often a key focus of many substantive studies in marketing and economics.

table[table omitted — 930 chars of source]

Inference and Coverage Analysis

We now demonstrate the performance of the debiasing and inference procedure discussed in $\S$(ref). The objective is to demonstrate the validity of the estimated confidence intervals. To this end, we estimate the average effect of a 1% change in own price on demand over all products ($\hat{\theta}$) and compute the corresponding confidence intervals of this effect. The difference between this section and $\S$(ref) lies in both the estimators and the methods. In terms of estimators, in $\S$(ref), we predict the market share ($\hat{\pi}_{jt}$), own-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$) and cross-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{k\neq j t}/ p_{k\neq j t} }$) for individual products. In contrast, the object of interest in this section is the average effect of price on demand across all products ($\hat{\theta}$). As a result, in $\S$(ref), we did not use the debiasing techniques that we apply here. It is important to emphasize that, in our approach, constructing a confidence interval is viable only for aggregate measures, not for individual observations.

To simulate the data, we consider a random coefficient logit model of demand with 3 products across 100 markets. We set the true model parameters to be $\beta_{ik} \sim \mathcal{N}(1, 0.5), \alpha_{i} \sim \mathcal{N}(-1,0.5)$. The effect of a 1% increase in a product's price is given by $$\theta_0 = \mathbb{E}[m(w_{jt},\pi)] = \mathbb{E}[\pi(p_{jt}*(1.01), X_{jt}, \{p_{kt}, X_{kt}\}_{k \ne j}) - \pi(p_{jt}, X_{jt}, \{p_{kt}, X_{kt}\}_{k \ne j})],$$ As discussed earlier, one way to estimate this effect is to compute the sample analog of this using the estimated $\hat{\pi}$, such that $\hat{\theta} = \frac{1}{n}\sum_{i = 1}^{n}m(w_{jt},\hat{\pi})$. However, as we pointed out earlier, the distribution of $\hat{\theta}$ might not be asymptotically normal. To demonstrate this, in Figure (ref), we display the histogram of the estimated effect across 50 random samples by using the plug-in method. We standardize the estimates by subtracting the mean and then dividing by the standard deviation and plot them against the standard normal distribution. As one can observe the distribution appears multi-modal and deviates from a standard normal distribution. Next, we use our proposed debiased estimator and plot the standardized estimates across 50 samples of draws in Figure (ref). The resultant distribution with the debiased estimator is much closer to a standard normal. Finally, we calculate the 95% confidence intervals using our debiased estimator across 50 random draws. In Table (ref), we report the bias (mean absolute error from all draws) and the coverage i.e., the percentage of times the true parameter is covered in the estimated confidence intervals. We find that bias across both data-generating processes (RCL and MNL) is notably low, reflecting only a -0.0001 difference from the true effect. The coverage rate of the 95% confidence interval in our corrected model is 90%, indicating good coverage. This shows that our debiased estimator can be used to conduct valid inference in finite samples.

figure[figure omitted — 1,178 chars of source]
table[table omitted — 954 chars of source]

Emprical Data Analysis: US Automobile Data (1971 - 1990)

In this section, we apply our model to a real-world dataset. We use the “US Automobile Data (1971 -- 1990)" from berry1995automobile. The dataset features cars in the US market from 1971 to 1990, with each year regarded as a market. The number of cars varies from 86 to 150 each year. For each car, the dataset provides information such as the car's name, the manufacturing company, factory region, market share, price, and four exogenous car characteristics: horsepower, space, mileage per dollar, and the presence of an air conditioning device.

Even though the dataset is relatively small, it presents three key challenges that make it difficult to use naive non-parametric estimators: (i) the dataset features markets with more than 100 products and only 20 markets in total, (ii) the product assortment in each year or market varies; (iii) the feature “price" has the endogeneity issue, which was not considered in numerical experiments in $\S$(ref). In this section, we demonstrate the use of our estimator, which is capable of effectively addressing such challenges posed in real-world datasets.

For each component of our model ($\phi_1$, $\phi_2$, and $\rho$), we use a standard 3-layer neural network, and this is implemented without further hyperparameter tuning. We implement ReLU activation function at each layer. This architecture is the same as the one we used in our numerical experiments. For comparison, we replicate the random coefficient logit model (with only the demand side) used by berry1995automobile using the Python package pyblp PyBLP. In our replication, we allow for heterogeneity in random coefficients across all variables. Our findings show that the estimates obtained from our model are comparable to the random coefficient logit estimation presented in berry1995automobile.

We estimate our model both without and with consideration of endogeneity. To address endogeneity, we utilize three sets of IVs -- (i) the sum of characteristics of all car models, excluding the product in focus, produced by the same firm in the same year; (ii) the sum of characteristics of all car models, excluding the product in focus, produced by rival firms in the same year; and (iii) cost shifters, which encompass the wage and exchange rate prevalent in the year and region where the factory is located. The utilization of traditional BLP-style instruments, as discussed by gandhi2019measuring, can be problematic due to their relative weakness, often resulting in considerable bias in the estimation of parameters. These issues are significantly exacerbated in non-parametric models. Thus, to counter potential concerns related to weak instruments, we employ a machine-learning-based IV methodology (MLIV) as proposed by singh2020machine. We detail the estimation procedure and results using BLP style IVs in Web Appendix (ref).

In Figure (ref), we present the estimated own-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$) of our model without IV and with IV. The x-axis represents the price of the focal product, while the y-axis shows the product's own-elasticity. Each point corresponds to a product in a market, resulting in 2,217 observations. We report the estimated elasticity based on the same price variation used in the BLP paper (a 1,000-dollar change). In Figure (ref), we observe that the majority of low-priced products (priced below 6,000 dollars) exhibit positive estimated own-elasticity, demonstrating the existence of the endogeneity. We also notice that this issue is attenuated when we correct for endogeneity (see Figure (ref)), i.e., which suggests that our approach is able to handle situations with endogenous features.

figure[figure omitted — 1,444 chars of source]

We report the own-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$) and cross-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{k\neq j t}/ p_{k\neq j t} }$) estimated in our model and random coefficient logit model with a sample of 13 cars in the 1990 market in Tables (ref) and (ref). The sample of 13 cars is the same as the one reported in berry1995automobile. Overall, our results are very similar and comparable to berry1995automobile. We also plot the distributions of the estimated own-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{jt}/p_{jt}}$) and cross-elasticity ($\frac{\partial \hat{\pi}_{jt} / \pi_{jt} }{\partial p_{k\neq j t}/ p_{k\neq j t} }$) obtained from our model and the BLP model in Figure (ref). The filled areas in the violin plots represent the complete range of the elasticities, while the text labels next to the line indicate the mean values. The estimated mean own- and cross-elasticities appear to be similar between our model and the BLP model, though our model exhibits a larger RMSE in the estimated elasticity values compared to the BLP model.

sidewaystable*\scalebox{0.7}{ \begin{tabular}{lccccccccccccc} \hline & Acura & BMW & Buick & Cadillac & Chevy & Ford & Ford & Honda & Lexus & Lincoln & Mazda & Nissan & Nissan \\ & Legend & 735i & Century & Seville & Cavalier & Escort & Taurus & Accord & LS400 & Town Car & 323 & Maxima & Sentra \\ \hline Acura Legend & -5.6060 & 0.1993 & 0.2198 & 0.2221 & 0.0632 & 0.2317 & 0.2199 & 0.2337 & 0.2144 & 0.2187 & 0.2595 & 0.2354 & 0.2595 \\ BMW 735i & 0.4095 & -6.1528 & 0.3653 & 0.4093 & 0.0352 & 0.3525 & 0.3655 & 0.3807 & 0.4120 & 0.4084 & 0.3547 & 0.4161 & 0.3547 \\ Buick Century & 0.1234 & 0.1020 & -6.1023 & 0.1213 & 0.0376 & 0.1294 & 0.1215 & 0.1205 & 0.1175 & 0.1191 & 0.1400 & 0.1299 & 0.1400 \\ Cadillac Seville & 0.2895 & 0.2179 & 0.2631 & -7.2896 & 0.0647 & 0.2692 & 0.2631 & 0.2566 & 0.2810 & 0.2818 & 0.2892 & 0.2849 & 0.2892 \\ Chevy Cavalier & 0.0142 & -0.0013 & 0.0202 & 0.0167 & -1.3447 & 0.0171 & 0.0202 & 0.0353 & 0.0126 & 0.0167 & 0.0354 & 0.0291 & 0.0355 \\ Ford Escort & 0.0410 & 0.0245 & 0.0353 & 0.0413 & -0.0148 & -1.8519 & 0.0353 & 0.0520 & 0.0392 & 0.0413 & 0.0494 & 0.0513 & 0.0494 \\ Ford Taurus & 0.1166 & 0.0914 & 0.1188 & 0.1160 & 0.0183 & 0.1199 & -6.1473 & 0.1258 & 0.1066 & 0.1157 & 0.1451 & 0.1290 & 0.1451 \\ Honda Accord & 0.0975 & 0.0647 & 0.0968 & 0.1006 & -0.0003 & 0.0945 & 0.0969 & -5.7438 & 0.0914 & 0.1006 & 0.1166 & 0.1140 & 0.1165 \\ Lexus LS400 & 0.3357 & 0.2606 & 0.3136 & 0.3325 & 0.0822 & 0.3235 & 0.3137 & 0.3126 & -6.8791 & 0.3271 & 0.3495 & 0.3348 & 0.3494 \\ Lincoln Town Car & 0.2681 & 0.2310 & 0.2548 & 0.2656 & 0.0713 & 0.2663 & 0.2548 & 0.2648 & 0.2586 & -5.3996 & 0.3009 & 0.2719 & 0.3009 \\ Mazda 323 & 0.0361 & 0.0212 & 0.0272 & 0.0326 & -0.0127 & 0.0249 & 0.0272 & 0.0363 & 0.0323 & 0.0326 & -2.6589 & 0.0404 & 0.0357 \\ Nissan Maxima & 0.1579 & 0.1367 & 0.1589 & 0.1555 & 0.0425 & 0.1670 & 0.1589 & 0.1689 & 0.1484 & 0.1534 & 0.1884 & -7.2216 & 0.1884 \\ Nissan Sentra & 0.0386 & 0.0239 & 0.0304 & 0.0384 & -0.0202 & 0.0294 & 0.0304 & 0.0496 & 0.0375 & 0.0383 & 0.0439 & 0.0470 & -1.8754 \\ \hline \end{tabular} } \caption{Estimated own- and cross-elasticities of a sample of automobile data using our method} Note: This table presents the estimated own- and cross-elasticity of a sample of 13 cars in the 1990 market using our model. The selected cars are the same as berry1995automobile reports. Each entry with row index $i$ and column index $j$ gives the percentage change in demand divided by the percentage change in price (based on \$1,000 change in the price of $i$).
sidewaystable*\scalebox{0.7}{ \begin{tabular}{lccccccccccccc} \hline & Acura & BMW & Buick & Cadillac & Chevy & Ford & Ford & Honda & Lexus & Lincoln & Mazda & Nissan & Nissan \\ & Legend & 735i & Century & Seville & Cavalier & Escort & Taurus & Accord & LS400 & Town Car & 323 & Maxima & Sentra \\ \hline Acura Legend & -5.4677 & 0.0489 & 0.0205 & 0.1029 & 0.0221 & 0.0220 & 0.0143 & 0.1477 & 0.1503 & 0.0273 & 0.0013 & 0.1359 & 0.0039 \\ BMW 735i & 0.1267 & -9.8502 & 0.0156 & 0.1058 & 0.0122 & 0.0121 & 0.0057 & 0.1313 & 0.1546 & 0.0278 & 0.0006 & 0.1375 & 0.0022 \\ Buick Century & 0.0184 & 0.0054 & -5.1978 & 0.0124 & 0.1153 & 0.1043 & 0.1475 & 0.1982 & 0.0165 & 0.0800 & 0.0076 & 0.0379 & 0.0174 \\ Cadillac Seville & 0.1293 & 0.0513 & 0.0175 & -6.6819 & 0.0151 & 0.0150 & 0.0099 & 0.1393 & 0.1576 & 0.0271 & 0.0008 & 0.1406 & 0.0027 \\ Chevy Cavalier & 0.0131 & 0.0028 & 0.0766 & 0.0071 & -3.1163 & 0.1421 & 0.0849 & 0.2608 & 0.0086 & 0.0404 & 0.0100 & 0.0395 & 0.0241 \\ Ford Escort & 0.0137 & 0.0029 & 0.0726 & 0.0074 & 0.1487 & -3.0590 & 0.0603 & 0.2781 & 0.0090 & 0.0263 & 0.0106 & 0.0419 & 0.0258 \\ Ford Taurus & 0.0048 & 0.0007 & 0.0554 & 0.0027 & 0.0479 & 0.0326 & -4.0258 & 0.0727 & 0.0017 & 0.1779 & 0.0026 & 0.0122 & 0.0057 \\ Honda Accord & 0.0387 & 0.0133 & 0.0582 & 0.0291 & 0.1151 & 0.1173 & 0.0568 & -4.3399 & 0.0409 & 0.0297 & 0.0081 & 0.0618 & 0.0200 \\ Lexus LS400 & 0.1350 & 0.0536 & 0.0166 & 0.1126 & 0.0130 & 0.0130 & 0.0045 & 0.1400 & -7.4316 & 0.0243 & 0.0006 & 0.1464 & 0.0024 \\ Lincoln Town Car & 0.0087 & 0.0034 & 0.0286 & 0.0069 & 0.0217 & 0.0135 & 0.1693 & 0.0362 & 0.0086 & -5.6139 & 0.0011 & 0.0123 & 0.0024 \\ Mazda 323 & 0.0114 & 0.0020 & 0.0743 & 0.0056 & 0.1476 & 0.1494 & 0.0679 & 0.2723 & 0.0063 & 0.0304 & -2.8631 & 0.0390 & 0.0254 \\ Nissan Maxima & 0.1008 & 0.0393 & 0.0314 & 0.0831 & 0.0493 & 0.0500 & 0.0271 & 0.1749 & 0.1209 & 0.0286 & 0.0033 & -4.7872 & 0.0086 \\ Nissan Sentra & 0.0140 & 0.0031 & 0.0707 & 0.0078 & 0.1471 & 0.1504 & 0.0617 & 0.2763 & 0.0095 & 0.0271 & 0.0105 & 0.0422 & -3.1799 \\ \hline \end{tabular} } \caption{Estimated own and cross-elasticities of a sample of automobile data using the BLP model} Note: This table presents the estimated own- and cross-elasticity of a sample of 13 cars in the 1990 market using the BLP model. The selected cars are the same as berry1995automobile reports. Each entry with row index $i$ and column index $j$ gives the percentage change in demand divided by the percentage change in price (based on \$1,000 change in the price of $i$).
figure[figure omitted — 833 chars of source]

We further estimate the average own-elasticity ($\hat{\theta}$) for high-priced, medium-priced, and low-priced cars and construct a confidence interval for each category using our inference procedure. We present our result in Table (ref). Both our model and the BLP model indicate that the average own-elasticity is highest (in terms of the absolute value of own-elasticity) for high-priced cars and lowest for low-priced cars. Moreover, the 95% confidence intervals for all three categories do not include zero. This also demonstrates the efficiency of our model even when there is a limited sample of only 20 observations.

table[table omitted — 1,143 chars of source]

The empirical analysis demonstrates the applicability and effectiveness of our model in a real-world setting, addressing challenges such as limited sample size, variability in product assortments, and endogeneity. The comparable results with established econometric models, such as BLP model, help in validating the robustness and reliability of our approach.

Conclusion

Choice models are fundamental in understanding consumer behavior and informing business decisions. Over the years, various methods, both parametric and non-parametric, have been developed to represent consumer behavior. While parametric methods, such as logit or probit-based models, are favored for their simplicity and interpretability, their restrictive assumptions can limit their ability to fully capture consumer preferences' intricacies. On the other hand, non-parametric methods offer a more flexible approach, but they often suffer from the “curse of dimensionality", where the complexity of estimating choice functions escalates exponentially with an increase in the number of products.

In this paper, we propose a fundamental characterization of choice models that combines the tractability of traditional choice models and the flexibility of non-parametric estimators. This characterization specifically tackles the challenge of high dimensionality in choice systems and facilitates flexible estimation of choice functions. Through extensive simulations, we validate the efficacy of our model, demonstrating its superior ability to capture a range of consumer behaviors that traditional choice models fail to capture. We also show how to address the endogeneity issue and estimate counterfactuals in our characterization. Furthermore, leveraging the recent strides in the automatic debiased machine learning literature, we offer an inference procedure that constructs confidence intervals on relevant objects, such as price elasticities. Finally, we apply our method to the automobile dataset from berry1995automobile. Our empirical analysis affirms that our model produces results that align well with the extant literature.

Our paper opens many avenues for future research. We focus on using neural network-based estimators. However, estimators, such as Gaussian processes and Gradient boosting-based estimators can be adopted to estimate the proposed functionals. Also, we believe more experimentation needs to be done on the neural network design side.

Competing Interests Declaration

Author(s) have no competing interests to declare.

{0.3em}