EconBase
← All papers

Combine and conquer: model averaging for out-of-distribution forecasting

Stephane Hess, Sander van Cranenburgh

arXiv 4 Jun 2025 · Econometrics

arXiv:2506.03693 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Travel behaviour modellers have an increasingly diverse set of models at their disposal, ranging from traditional econometric structures to models from mathematical psychology and data-driven approaches from machine learning. A key question arises as to how well these different models perform in prediction, especially when considering trips of different characteristics from those used in estimation, i.e. out-of-distribution prediction, and whether better predictions can be obtained by combining insights from the different models. Across two case studies, we show that while data-driven approaches excel in predicting mode choice for trips within the distance bands used in estimation, beyond that range, the picture is fuzzy. To leverage the relative advantages of the different model families and capitalise on the notion that multiple `weak' models can result in more robust models, we put forward the use of a model averaging approach that allocates weights to different model families as a function of the distance between the characteristics of the trip for which predictions are made, and those used in model estimation. Overall, we see that the model averaging approach gives larger weight to models with stronger behavioural or econometric underpinnings the more we move outside the interval of trip distances covered in estimation. Across both case studies, we show that our model averaging approach obtains improved performance both on the estimation and validation data, and crucially also when predicting mode choices for trips of distances outside the range used in estimation.

Citation extraction

32
references
35
in-text mentions
32
distinct cited
11
self-citations
9,352
main-text words

appendix boundary found by appendix_command · 94% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Hancock, Thomas O. and Hess, Stephane and Marley, A.A.J. and Choudhu… An accumulation of preference: Two alternative dynamic models for understanding transport choices self0.64422100%
2Hastie, Trevor and Tibshirani, Robert and Friedman, Jerome (2009) The Elements of Statistical Learning: Data Mining, Inference, and Prediction0.5112250%
3Fox, James and Daly, Andrew and Hess, Stephane and Miller, Eric (2014) Temporal transferability of models of mode-destination choice for the Greater Toronto and Hamilton Area self0.51121100%
4A. Daly and S. Zachary (1978) Improved multiple choice models0.40511100%
5Chen, Tianqi and Guestrin, Carlos (2016) XGBoost: A Scalable Tree Boosting System0.40511100%
6Julian Hagenauer and Marco Helbich (2017) A comparative study of machine learning classifiers for modeling travel mode choice0.40511100%
7Yafei Han and Francisco Camara Pereira and Moshe Ben-Akiva and Chris… (2022) A neural-embedded discrete choice model: Learning taste representation with strengthened interpretability0.40511100%
8Thomas O. Hancock and Jan Broekaert and Stephane Hess and Charisma F… (2020) Quantum probability: A new method for modelling travel behaviour self0.40511100%
9Hancock, Thomas O and Hess, Stephane and Daly, Andrew J. and Fox, Ja… (2020) Using a sequential latent class approach for model averaging: Benefits in forecasting and behavioural insights self0.40511100%
10Paszke, Adam and Gross, Sam and Massa, Francisco and Lerer, Adam and… (2019) PyTorch: An Imperative Style, High-Performance Deep Learning Library0.40511100%

Showing the top 10 of 32 scored citations.