arXiv 30 Nov 2025 · Artificial Intelligence
arXiv:2512.01107 · PDF · DOI · OpenAlex · Extracted main text
Foundation models, and in particular large language models, can generate highly informative responses, prompting growing interest in using these ”synthetic” outputs as data in empirical research and decision-making. This paper introduces the idea of a foundation prior, which shows that model-generated outputs are not as real observations, but draws from the foundation prior induced prior predictive distribution. As such synthetic data reflects both the model's learned patterns and the user's subjective priors, expectations, and biases. We model the subjectivity of the generative process by making explicit the dependence of synthetic outputs on the user's anticipated data distribution, the prompt-engineering process, and the trust placed in the foundation model. We derive the foundation prior as an exponential-tilted, generalized Bayesian update of the user's primitive prior, where a trust parameter governs the weight assigned to synthetic data. We then show how synthetic data and the associated foundation prior can be incorporated into standard statistical and econometric workflows, and discuss their use in applications such as refining complex models, informing latent constructs, guiding experimental design, and augmenting random-coefficient and partially linear specifications. By treating generative outputs as structured, explicitly subjective priors rather than as empirical observations, the framework offers a principled way to harness foundation models in empirical work while avoiding the conflation of synthetic ”facts” with real data.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ludwig, Jens and Mullainathan, Sendhil and Rambachan, Ashesh (2025) Large Language Models: An Applied Econometric Framework | 0.843 | 3 | 3 | 100% |
| 2 | Arai, Yuta and others (2025) Generative AI Priors Increase Effective Sample Size in Clinical Trials | 0.644 | 2 | 2 | 100% |
| 3 | Bajari, Patrick and Fox, Jeremy T. and Ryan, Stephen P (2007) Linear Regression Estimation of Discrete Choice Models with Nonparametric Distributions of Random Coefficients | 0.511 | 2 | 1 | 100% |
| 4 | Csiszár, Imre (1975) I-divergence geometry of probability distributions and minimization problems | 0.511 | 2 | 1 | 100% |
| 5 | Kullback, Solomon (1959) Information Theory and Statistics | 0.511 | 2 | 1 | 100% |
| 6 | Angelopoulos, Anastasios N. and Bates, Stephen and Fannjiang, Clara… (2023) Prediction-Powered Inference | 0.405 | 1 | 1 | 100% |
| 7 | Argyle, Lisa P. and Busby, Ethan C. and Fulda, Nancy and Gubler, Jos… (2023) Out of One, Many: Using Language Models to Simulate Human Samples | 0.405 | 1 | 1 | 100% |
| 8 | Berger, James O. and Pericchi, Luis R (1996) Intrinsic Bayes factors for model selection and prediction | 0.405 | 1 | 1 | 100% |
| 9 | Berry, Steven and Levinsohn, James and Pakes, Ariel (2004) Differentiated Products Demand Systems from a Combination of Micro and Macro Data: The New Car Market | 0.405 | 1 | 1 | 100% |
| 10 | Bissiri, Peter G. and Holmes, Christopher C. and Walker, Stephen G (2016) A General Framework for Updating Belief Distributions | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 32 scored citations.