Daniel F. Villarraga, Ricardo A. Daziano
arXiv 23 May 2025 · Statistics — Machine Learning · 1 citations (OpenAlex)
arXiv:2505.18077 · PDF · DOI · OpenAlex · Extracted main text
Discrete choice models (DCMs) are used to analyze individual decision-making in contexts such as transportation choices, political elections, and consumer preferences. DCMs play a central role in applied econometrics by enabling inference on key economic variables, such as marginal rates of substitution, rather than focusing solely on predicting choices on new unlabeled data. However, while traditional DCMs offer high interpretability and support for point and interval estimation of economic quantities, these models often underperform in predictive tasks compared to deep learning (DL) models. Despite their predictive advantages, DL models remain largely underutilized in discrete choice due to concerns about their lack of interpretability, unstable parameter estimates, and the absence of established methods for uncertainty quantification. Here, we introduce a deep learning model architecture specifically designed to integrate with approximate Bayesian inference methods, such as Stochastic Gradient Langevin Dynamics (SGLD). Our proposed model collapses to behaviorally informed hypotheses when data is limited, mitigating overfitting and instability in underspecified settings while retaining the flexibility to capture complex nonlinear relationships when sufficient data is available. We demonstrate our approach using SGLD through a Monte Carlo simulation study, evaluating both predictive metrics--such as out-of-sample balanced accuracy--and inferential metrics--such as empirical coverage for marginal rates of substitution interval estimates. Additionally, we present results from two empirical case studies: one using revealed mode choice data in NYC, and the other based on the widely used Swiss train choice stated preference data.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Villarraga, Daniel F and Daziano, Ricardo A (2025) Designing Graph Convolutional Neural Networks for Discrete Choice with Network Effects self | 1.000 | 9 | 5 | 100% |
| 2 | Wang, Shenhao and Wang, Qingyi and Zhao, Jinhua (2020) Deep neural networks for choice analysis: Extracting complete economic information for interpretation | 1.000 | 6 | 4 | 100% |
| 3 | Welling, Max and Teh, Yee W (2011) Bayesian learning via stochastic gradient Langevin dynamics | 0.811 | 4 | 2 | 100% |
| 4 | Arkoudi, Ioanna and Krueger, Rico and Azevedo, Carlos Lima and Perei… (2023) Combining discrete choice models and neural networks through embeddings: Formulation, interpretability and performance | 0.737 | 3 | 2 | 100% |
| 5 | Papamarkou, Theodore and Skoularidou, Maria and Palla, Konstantina a… (2024) Position: Bayesian deep learning is needed in the age of large-scale AI | 0.737 | 3 | 2 | 100% |
| 6 | Wilson, Andrew G and Izmailov, Pavel (2020) Bayesian deep learning and a probabilistic perspective of generalization | 0.737 | 3 | 2 | 100% |
| 7 | Wilson, Andrew Gordon and Izmailov, Pavel (2020) Bayesian deep learning and a probabilistic perspective of generalization | 0.693 | 5 | 1 | 100% |
| 8 | Izmailov, Pavel and Podoprikhin, Dmitrii and Garipov, Timur and Vetr… (2018) Averaging weights leads to wider optima and better generalization | 0.693 | 5 | 1 | 100% |
| 9 | Villarraga, Daniel F. and Daziano, Ricardo A (2025) Hierarchical Nearest Neighbor Gaussian Process models for discrete choice: Mode choice in New York City | 0.644 | 2 | 2 | 100% |
| 10 | Sifringer, Brian and Lurkin, Virginie and Alahi, Alexandre (2018) Enhancing discrete choice models with neural networks | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 33 scored citations.