EconBase
← All papers

Using machine learning metrics to provide deeper insights into the performance of choice models

Lorenzo Muñoz, Stephane Hess, Thomas O. Hancock, Georges Sfeir

arXiv 17 Sep 2026 · Econometrics

arXiv:2609.20655 · PDF · Extracted main text

Abstract

Machine learning (ML) techniques are increasingly drawing interest in the choice modelling (CM) field. The focus has primarily been on comparing the performance of these contrasting approaches or on improving behavioural insights for ML techniques, rather than translating ideas from one field into the other. In the present paper, we specifically focus on knowledge transfer from ML into CM in the context of model performance evaluation. In CM, model performance is typically evaluated using log-likelihood and related indicators, which are aggregate fit metrics that focus on overall fit. Conversely, in ML, the focus is on alternative-level misclassifications and correct classifications, which provide a more nuanced view of the results. To bridge these approaches, we explore the use of a probabilistic version of the confusion matrix, which reports the average probability of the model predicting each alternative, conditional on which alternative was observed to be chosen, across all choice tasks. This enables the computation of probabilistic ML metrics for both classic choice models and ML algorithms. We analyse model performance jointly in terms of overall fit and alternative-level predictions. Our findings demonstrate that models with similar log-likelihood can exhibit substantially different confusion matrices, revealing different probability patterns that aggregate metrics cannot capture. This framework identifies where models systematically `confuse' alternatives, highlighting trade-offs between alternatives, and potentially guiding model specification. Furthermore, evaluating these matrices and metrics out-of-sample reveals alternative-level prediction shifts that significantly impact forecasting performance.

Citation extraction

65
references
86
in-text mentions
65
distinct cited
9
self-citations
12,822
main-text words

appendix boundary found by appendix_command · 88% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Hillel, Tim and Bierlaire, Michel and Elshafie, Mohammed ZEB and Jin… (2021) A systematic review of machine learning classification methodologies for modelling passenger mode choice0.64441100%
2Yacouby, Reda and Axman, Dustin (2020) Probabilistic Extension of Precision, Recall, and F1 Score for More Thorough Evaluation of Classification Models0.64441100%
3Sander van Cranenburgh and Shenhao Wang and Akshay Vij and Francisco… (2022) Choice modelling in the age of machine learning - Discussion paper0.58531100%
4Zhang, Jiawei and Yang, Yuhong and Ding, Jie (2023) Information criteria for model selection0.58531100%
5Akaike, H (1974) A new look at the statistical model identification0.51121100%
6Azam Ali and Arash Kalatian and Charisma F. Choudhury (2023) Comparing and contrasting choice model and machine learning techniques in the context of vehicle ownership decisions0.51121100%
7Gideon Schwarz (1978) Estimating the Dimension of a Model0.51121100%
8Hess, Stephane and Lancsar, Emily and Mariel, Petr and Meyerhoff, Jü… (2022) The path towards herd immunity: Predicting COVID-19 vaccination uptake through results from a stated choice study across six con… self0.51121100%
9Ortelli, Nicola and Hillel, Tim and Pereira, Francisco C and de Lapp… (2021) Assisted specification of discrete choice models0.51121100%
10Terven, Juan and Cordova-Esparza, Diana M and Ramirez-Pedraza, Alfon… (2023) Loss functions and metrics in deep learning0.51121100%

Showing the top 10 of 65 scored citations.