EconBase
← All papers

Foundation Priors

Sanjog Misra

arXiv 30 Nov 2025 · Artificial Intelligence

arXiv:2512.01107 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Foundation models, and in particular large language models, can generate highly informative responses, prompting growing interest in using these ”synthetic” outputs as data in empirical research and decision-making. This paper introduces the idea of a foundation prior, which shows that model-generated outputs are not as real observations, but draws from the foundation prior induced prior predictive distribution. As such synthetic data reflects both the model's learned patterns and the user's subjective priors, expectations, and biases. We model the subjectivity of the generative process by making explicit the dependence of synthetic outputs on the user's anticipated data distribution, the prompt-engineering process, and the trust placed in the foundation model. We derive the foundation prior as an exponential-tilted, generalized Bayesian update of the user's primitive prior, where a trust parameter governs the weight assigned to synthetic data. We then show how synthetic data and the associated foundation prior can be incorporated into standard statistical and econometric workflows, and discuss their use in applications such as refining complex models, informing latent constructs, guiding experimental design, and augmenting random-coefficient and partially linear specifications. By treating generative outputs as structured, explicitly subjective priors rather than as empirical observations, the framework offers a principled way to harness foundation models in empirical work while avoiding the conflation of synthetic ”facts” with real data.

Citation extraction

32
references
38
in-text mentions
32
distinct cited
0
self-citations
9,754
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Ludwig, Jens and Mullainathan, Sendhil and Rambachan, Ashesh (2025) Large Language Models: An Applied Econometric Framework0.84333100%
2Arai, Yuta and others (2025) Generative AI Priors Increase Effective Sample Size in Clinical Trials0.64422100%
3Bajari, Patrick and Fox, Jeremy T. and Ryan, Stephen P (2007) Linear Regression Estimation of Discrete Choice Models with Nonparametric Distributions of Random Coefficients0.51121100%
4Csiszár, Imre (1975) I-divergence geometry of probability distributions and minimization problems0.51121100%
5Kullback, Solomon (1959) Information Theory and Statistics0.51121100%
6Angelopoulos, Anastasios N. and Bates, Stephen and Fannjiang, Clara… (2023) Prediction-Powered Inference0.40511100%
7Argyle, Lisa P. and Busby, Ethan C. and Fulda, Nancy and Gubler, Jos… (2023) Out of One, Many: Using Language Models to Simulate Human Samples0.40511100%
8Berger, James O. and Pericchi, Luis R (1996) Intrinsic Bayes factors for model selection and prediction0.40511100%
9Berry, Steven and Levinsohn, James and Pakes, Ariel (2004) Differentiated Products Demand Systems from a Combination of Micro and Macro Data: The New Car Market0.40511100%
10Bissiri, Peter G. and Holmes, Christopher C. and Walker, Stephen G (2016) A General Framework for Updating Belief Distributions0.40511100%

Showing the top 10 of 32 scored citations.