Shenhao Wang, Baichuan Mo, Yunhan Zheng, Stephane Hess, Jinhua Zhao
arXiv 1 Feb 2021 · Machine Learning · 21 citations (OpenAlex)
arXiv:2102.01130 · PDF · DOI · OpenAlex · Extracted main text
Numerous studies have compared machine learning (ML) and discrete choice models (DCMs) in predicting travel demand. However, these studies often lack generalizability as they compare models deterministically without considering contextual variations. To address this limitation, our study develops an empirical benchmark by designing a tournament model, thus efficiently summarizing a large number of experiments, quantifying the randomness in model comparisons, and using formal statistical tests to differentiate between the model and contextual effects. This benchmark study compares two large-scale data sources: a database compiled from literature review summarizing 136 experiments from 35 studies, and our own experiment data, encompassing a total of 6,970 experiments from 105 models and 12 model families. This benchmark study yields two key findings. Firstly, many ML models, particularly the ensemble methods and deep learning, statistically outperform the DCM family (i.e., multinomial, nested, and mixed logit models). However, this study also highlights the crucial role of the contextual factors (i.e., data sources, inputs and choice categories), which can explain models' predictive performance more effectively than the differences in model types alone. Model performance varies significantly with data sources, improving with larger sample sizes and lower dimensional alternative sets. After controlling all the model and contextual factors, significant randomness still remains, implying inherent uncertainty in such model comparisons. Overall, we suggest that future researchers shift more focus from context-specific model comparisons towards examining model transferability across contexts and characterizing the inherent uncertainty in ML, thus creating more robust and generalizable next-generation travel demand models.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Wang S, Mo B, Zhao J (2020) Deep neural networks for choice analysis: Architecture design with alternative-specific utility functions [Journal Article] | 0.928 | 4 | 3 | 100% |
| 2 | Wang S, Wang Q, Zhao J (2020) Deep neural networks for choice analysis: Extracting complete economic information for interpretation [Journal Article] | 0.928 | 4 | 3 | 100% |
| 3 | Cantarella GE, de Luca S (2005) Multilayer feedforward networks for transportation mode choice analysis: An analysis and a comparison with random utility models | 0.644 | 2 | 2 | 100% |
| 4 | Cheng L, Chen X, De Vos J, Lai X, Witlox F (2019) Applying a random forest method approach to model travel mode choice behavior | 0.644 | 2 | 2 | 100% |
| 5 | Hagenauer J, Helbich M (2017) A comparative study of machine learning classifiers for modeling travel mode choice | 0.644 | 2 | 2 | 100% |
| 6 | Wang F, Ross CL (2018) Machine learning travel mode choices: Comparing the performance of an extreme gradient boosting model with a multinomial logit m… | 0.644 | 2 | 2 | 100% |
| 7 | Wang S, Wang Q, Zhao J (2020) Multitask learning deep neural networks to combine revealed and stated preference data [Journal Article] | 0.644 | 2 | 2 | 100% |
| 8 | Shafique MA, Hato E (2015) Use of acceleration data for transportation mode prediction [Journal Article] | 0.585 | 3 | 1 | 100% |
| 9 | Aha DW, Kibler D, Albert MK (1991) Instance-based learning algorithms | 0.511 | 2 | 1 | 100% |
| 10 | Zhang Y, Xie Y (2076) Travel mode choice modeling with support vector machines | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 83 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Bayesian Deep Learning for Discrete Choice | 0.511 | 2 | 1 |
| 2 | The Mixed Aggregate Preference Logit Model: A Machine Learning Approach to Modeling Unobserved Heterogeneity in Discrete Choice Analysis | 0.405 | 1 | 1 |