EconBase
← All papers

Tabular Foundation Models for Discrete Choice Estimation

Liu Liu, Dan Zhang

arXiv 14 Jul 2026 · Machine Learning

arXiv:2607.13314 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Tabular foundation models (TFMs) generate predictions on structured data via in-context learning, without task-specific estimation. We ask whether TFMs can be effectively applied to discrete choice, a central demand estimation framework in marketing and operations, and find that directly applying TFMs yields limited performance. The gap is structural: TFMs assume row-independent observations, whereas discrete choice is inherently set-valued and subject to persistent consumer preference heterogeneity. We propose a reformulation that encodes both choice-set dependence and individual heterogeneity within a row-based learning framework. Evaluated on a yogurt scanner panel, individual-level heterogeneity encoding is the dominant driver of predictive accuracy. The best reformulation outperforms hierarchical Bayesian estimation on both holdout log-likelihood and hit rate, running 16 times faster, a practical advantage for large-scale demand estimation. The advantage is largest in the medium-data regime (10--40 purchase occasions per consumer), where parametric Bayesian shrinkage most distorts estimates for atypical consumers. Fine-tuning on population choice data provides additional gains for consumers with shallow purchase histories, where in-context learning has limited individual-specific signal to condition on. These results establish a principled approach for applying foundation models to consumer choice problems more broadly.

Citation extraction

60
references
97
in-text mentions
60
distinct cited
1
self-citations
12,701
main-text words

appendix boundary found by appendix_command · 95% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Hollmann, Noah and Müller, Samuel and Purucker, Lennart and Krishnak… (2025) Accurate predictions on small data with a tabular foundation model1.00064100%
2Rossi, Peter E and Allenby, Greg M and McCulloch, Robert (2005) Bayesian Statistics and Marketing0.9507686%
3Allenby, Greg M and Rossi, Peter E (1998) Marketing Models of Consumer Heterogeneity0.9285480%
4Train, Kenneth E (2009) Discrete choice methods with simulation0.8434475%
5Hollmann, Noah and Müller, Samuel and Eggensperger, Katharina and Hu… (2023) TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second0.84333100%
6Evgeniou, Theodoros and Pontil, Massimiliano and Toubia, Olivier (2007) A Convex Optimization Approach to Modeling Consumer Heterogeneity in Conjoint Estimation0.73732100%
7Müller, Samuel and Hollmann, Noah and Arango, Sebastian Pineda and G… (2022) Transformers Can Do Bayesian Inference0.73732100%
8Nagler, Thomas (2023) Statistical Foundations of Prior-Data Fitted Networks0.73732100%
9Brown, Tom B and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie… (2020) Language Models are Few-Shot Learners0.64422100%
10Eremeev, Dmitry and Bazhenov, Gleb and Platonov, Oleg and Babenko, A… (2025) Turning Tabular Foundation Models into Graph Foundation Models0.64422100%

Showing the top 10 of 60 scored citations.