Shenhao Wang, Baichuan Mo, Jinhua Zhao
arXiv 22 Oct 2020 · Machine Learning · publishedTransportation Research Part B Methodological (2021) · 55 citations (OpenAlex)
arXiv:2010.11644 · PDF · DOI · OpenAlex · Extracted main text
Researchers often treat data-driven and theory-driven models as two disparate or even conflicting methods in travel behavior analysis. However, the two methods are highly complementary because data-driven methods are more predictive but less interpretable and robust, while theory-driven methods are more interpretable and robust but less predictive. Using their complementary nature, this study designs a theory-based residual neural network (TB-ResNet) framework, which synergizes discrete choice models (DCMs) and deep neural networks (DNNs) based on their shared utility interpretation. The TB-ResNet framework is simple, as it uses a ($\delta$, 1-$\delta$) weighting to take advantage of DCMs' simplicity and DNNs' richness, and to prevent underfitting from the DCMs and overfitting from the DNNs. This framework is also flexible: three instances of TB-ResNets are designed based on multinomial logit model (MNL-ResNets), prospect theory (PT-ResNets), and hyperbolic discounting (HD-ResNets), which are tested on three data sets. Compared to pure DCMs, the TB-ResNets provide greater prediction accuracy and reveal a richer set of behavioral mechanisms owing to the utility function augmented by the DNN component in the TB-ResNets. Compared to pure DNNs, the TB-ResNets can modestly improve prediction and significantly improve interpretation and robustness, because the DCM component in the TB-ResNets stabilizes the utility functions and input gradients. Overall, this study demonstrates that it is both feasible and desirable to synergize DCMs and DNNs by combining their utility specifications under a TB-ResNet framework. Although some limitations remain, this TB-ResNet framework is an important first step to create mutual benefits between DCMs and DNNs for travel behavior modeling, with joint improvement in prediction, interpretation, and robustness.
appendix boundary found by appendix_titled_section at “Appendix I: Proof of Propositions 1 and 2” · 80% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Wang, Shenhao, Wang, Qingyi, Zhao, Jinhua (2020) Deep neural networks for choice analysis: Extracting complete economic information for interpretation self | 1.000 | 8 | 3 | 100% |
| 2 | Tanaka, Tomomi, Camerer, Colin F, Nguyen, Quang (2010) Risk and time preferences: linking experimental and household survey data from Vietnam | 1.000 | 6 | 4 | 100% |
| 3 | Wang, Shenhao, Mo, Baichuan, Zhao, Jinhua (2020) Deep neural networks for choice analysis: Architecture design with alternative-specific utility functions self | 1.000 | 5 | 4 | 100% |
| 4 | Lipton, Zachary C (2016) The mythos of model interpretability | 1.000 | 5 | 3 | 100% |
| 5 | Wang, Shenhao, Wang, Qingyi, Bailey, Nate, Zhao, Jinhua (2018) Deep Neural Networks for Choice Analysis: A Statistical Learning Theory Perspective self | 1.000 | 5 | 3 | 100% |
| 6 | McFadden, Daniel (1974) Conditional logit analysis of qualitative choice behavior | 0.941 | 6 | 4 | 83% |
| 7 | Ribeiro, Marco Tulio, Singh, Sameer, Guestrin, Carlos (2016) Why should i trust you?: Explaining the predictions of any classifier | 0.843 | 3 | 3 | 100% |
| 8 | Wang, Shenhao, Wang, Qingyi, Zhao, Jinhua (2020) Multitask learning deep neural networks to combine revealed and stated preference data self | 0.843 | 3 | 3 | 100% |
| 9 | Ben-Akiva, Moshe E, Lerman, Steven R (1985) Discrete choice analysis: theory and application to travel demand | 0.737 | 3 | 3 | 67% |
| 10 | Train, Kenneth E (2009) Discrete choice methods with simulation | 0.737 | 3 | 3 | 67% |
Showing the top 10 of 84 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Bayesian Deep Learning for Discrete Choice | 0.405 | 1 | 1 |
| 2 | Amortized Inference for Correlated Discrete Choice Models via Equivariant Neural Networks | 0.405 | 1 | 1 |