Prateek Bansal, Rico Krueger, Michel Bierlaire, Ricardo A. Daziano, Taha H. Rashidi
arXiv 7 Apr 2019 · Statistics — Machine Learning · publishedTransportation Research Part B Methodological (2019) · 37 citations (OpenAlex)
arXiv:1904.03647 · PDF · DOI · OpenAlex · Extracted main text
Variational Bayes (VB) methods have emerged as a fast and computationally-efficient alternative to Markov chain Monte Carlo (MCMC) methods for scalable Bayesian estimation of mixed multinomial logit (MMNL) models. It has been established that VB is substantially faster than MCMC at practically no compromises in predictive accuracy. In this paper, we address two critical gaps concerning the usage and understanding of VB for MMNL. First, extant VB methods are limited to utility specifications involving only individual-specific taste parameters. Second, the finite-sample properties of VB estimators and the relative performance of VB, MCMC and maximum simulated likelihood estimation (MSLE) are not known. To address the former, this study extends several VB methods for MMNL to admit utility specifications including both fixed and random utility parameters. To address the latter, we conduct an extensive simulation-based evaluation to benchmark the extended VB methods against MCMC and MSLE in terms of estimation times, parameter recovery and predictive accuracy. The results suggest that all VB variants with the exception of the ones relying on an alternative variational lower bound constructed with the help of the modified Jensen's inequality perform as well as MCMC and MSLE at prediction and parameter recovery. In particular, VB with nonconjugate variational message passing and the delta-method (VB-NCVMP-Delta) is up to 16 times faster than MCMC and MSLE. Thus, VB-NCVMP-Delta can be an attractive alternative to MCMC and MSLE for fast, scalable and accurate estimation of MMNL models.
appendix boundary found by appendix_command · 95% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Tan, L. S. L (2017) Stochastic variational inference for large-scale discrete choice models using adaptive batch sizes | 1.000 | 11 | 5 | 100% |
| 2 | Depraetere, N. and Vandebroek, M (2017) A comparison of variational approximations for fast inference in mixed logit models | 1.000 | 11 | 3 | 100% |
| 3 | Train, K. E (2009) Discrete Choice Methods with Simulation | 1.000 | 9 | 4 | 100% |
| 4 | Braun, M. and McAuliffe, J (2010) Variational Inference for Large-Scale Models of Discrete Choice | 1.000 | 8 | 3 | 100% |
| 5 | Nocedal, J. and Wright, S (2006) Numerical optimization | 1.000 | 6 | 3 | 100% |
| 6 | Rossi, P. E., Allenby, G. M., and McCulloch, R (2012) Bayesian statistics and marketing | 1.000 | 6 | 3 | 100% |
| 7 | Knowles, D. A. and Minka, T (2011) Non-conjugate variational message passing for multinomial and binary regression | 0.874 | 7 | 2 | 100% |
| 8 | Blei, D. M., Kucukelbir, A., and McAuliffe, J. D (2017) Variational Inference: A Review for Statisticians | 0.874 | 6 | 2 | 100% |
| 9 | Akinc, D. and Vandebroek, M (2018) Bayesian estimation of mixed logit models: Selecting an appropriate prior for the covariance matrix | 0.811 | 4 | 2 | 100% |
| 10 | Bhat, C. R. and Lavieri, P. S (2018) A new mixed mnp model accommodating a variety of dependent non-normal coefficient distributions | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 49 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | 1905.00419 | 1.000 | 8 | 4 |
| 2 | ResLogit: A residual neural network logit model for data-driven choice modelling | 0.644 | 2 | 2 |
| 3 | Discrete Choice Analysis with Machine Learning Capabilities | 0.405 | 1 | 1 |