arXiv 25 Mar 2026 · Statistics — Methodology · 1 citations (OpenAlex)
arXiv:2603.24705 · PDF · DOI · OpenAlex · Extracted main text
Discrete choice models are fundamental tools in management science, economics, and marketing for understanding and predicting decision-making. Logit-based models are dominant in applied work, largely due to their convenient closed-form expressions for choice probabilities. However, these models entail restrictive assumptions on the stochastic utility component, constraining our ability to capture realistic and theoretically grounded choice behavior$-$most notably, substitution patterns. In this work, we propose an amortized inference approach using a neural network emulator to approximate choice probabilities for general error distributions, including those with correlated errors. Our proposal includes a specialized neural network architecture and accompanying training procedures designed to respect the invariance properties of discrete choice models. We provide group-theoretic foundations for the architecture, including a proof of universal approximation given a minimal set of invariant features. Once trained, the emulator enables rapid likelihood evaluation and gradient computation. We use Sobolev training, augmenting the likelihood loss with a gradient-matching penalty so that the emulator learns both choice probabilities and their derivatives. We show that emulator-based maximum likelihood estimators are consistent and asymptotically normal under mild approximation conditions, and we provide sandwich standard errors that remain valid even with imperfect likelihood approximation. Simulations show significant gains over the GHK simulator in accuracy and speed.
appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Zaheer, Manzil and Kottur, Satwik and Ravanbakhsh, Siamak and Poczos… (2017) Deep Sets | 1.000 | 6 | 4 | 100% |
| 2 | Blum-Smith, Ben and Huang, Ningyuan (Teresa) and Cuturi, Marco and V… (2025) Functions on Symmetric Matrices and Point Clouds via Lightweight Invariant Features from Galois Theory | 0.961 | 9 | 5 | 89% |
| 3 | Czarnecki, Wojciech Marian and Osindero, Simon and Jaderberg, Max an… (2017) Sobolev training for neural networks | 0.941 | 6 | 4 | 83% |
| 4 | Norets, Andriy (2012) Estimation of Dynamic Discrete Choice Models Using Artificial Neural Network Approximations | 0.644 | 2 | 2 | 100% |
| 5 | Singh, Amandeep and Liu, Ye and Yoganarasimhan, Hema (2023) Choice models and permutation invariance: Demand estimation in differentiated products markets | 0.644 | 2 | 2 | 100% |
| 6 | Aouad, Ali and Antoine Desir (2025) Representing Random Utility Choice Models with Nueral Networks | 0.585 | 3 | 1 | 100% |
| 7 | McFadden, Daniel and Train, Kenneth (2000) Mixed MNL models for discrete response | 0.585 | 3 | 1 | 100% |
| 8 | Glen E. Bredon (1972) Introduction to Compact Transformation Groups | 0.511 | 3 | 2 | 33% |
| 9 | Cybenko, G (1989) Approximation by Superpositions of a Sigmoidal Function | 0.511 | 3 | 2 | 33% |
| 10 | Hornik, Kurt and Stinchcombe, Maxwell and White, Halbert (1989) Multilayer Feedforward Networks are Universal Approximators | 0.511 | 3 | 2 | 33% |
Showing the top 10 of 59 scored citations.