EconBase
← All papers

Large Language Models: An Applied Econometric Framework

Jens Ludwig, Sendhil Mullainathan, Ashesh Rambachan

arXiv 9 Dec 2024 · Econometrics · publishedAnnual Review of Economics (2026) · 19 citations (OpenAlex)

arXiv:2412.07031 · PDF · DOI · OpenAlex · Extracted main text

Abstract

How can we use the novel capacities of large language models (LLMs) in empirical research? And how can we do so while accounting for their limitations, which are themselves only poorly understood? We develop an econometric framework to answer this question that distinguishes between two types of empirical tasks. Using LLMs for prediction problems (including hypothesis generation) is valid under one condition: no “leakage” between the LLM's training dataset and the researcher's sample. No leakage can be ensured by using open-source LLMs with documented training data and published weights. Using LLM outputs for estimation problems to automate the measurement of some economic concept (expressed either by some text or from human subjects) requires the researcher to collect at least some validation data: without such data, the errors of the LLM's automation cannot be assessed and accounted for. As long as these steps are taken, LLM outputs can be used in empirical research with the familiar econometric guarantees we desire. Using two illustrative applications to finance and political economy, we find that these requirements are stringent; when they are violated, the limitations of LLMs now result in unreliable empirical estimates. Our results suggest the excitement around the empirical uses of LLMs is warranted -- they allow researchers to effectively use even small amounts of language data for both prediction and estimation -- but only with these safeguards in place.

Citation extraction

110
references
139
in-text mentions
110
distinct cited
12
self-citations
16,126
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Ludwig, Jens and Mullainathan, Sendhil (2024) Machine learning as a tool for hypothesis generation self0.84333100%
2Paul Glasserman and Caden Lin (2023) Assessing Look-Ahead Bias in Stock Return Predictions Generated By GPT Sentiment Analysis0.81142100%
3Anastasios N. Angelopoulos and Stephen Bates and Clara Fannjiang and… (2023) Prediction-powered inference0.73732100%
4Sarkar, Suproteem K and Vafa, Keyon (2024) Lookahead bias in pretrained language models0.73732100%
5Egami, Naoki and Hinck, Musashi and Stewart, Brandon M. and Wei, Han… (2024) Using imperfect surrogates for downstream inference: design-based supervised learning for social science applications of large l…0.64422100%
6Lung-fei Lee and Jungsywan H. Sepanski (1995) Estimation of Linear and Nonlinear Errors-in-Variables Models Using Validation Data0.64422100%
7Liu, Pengfei and Yuan, Weizhe and Fu, Jinlan and Jiang, Zhengbao and… (2023) Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing0.64422100%
8Suproteem Sarkar (2024) StoriesLM: A Family of Language Models With Time-Indexed Training Data0.64422100%
9Siruo Wang and Tyler H. McCormick and Jeffrey T. Leek (2020) Methods for correcting inference based on outcomes predicted by machine learning0.64422100%
10Jacob Carlson and Melissa Dell (2025) A Unifying Framework for Robust and Efficient Inference with Unstructured Data0.64422100%

Showing the top 10 of 110 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Causal Inference on Outcomes Learned from Text0.92843
2Foundation Priors0.84333
3A Unifying Framework for Robust and Efficient Inference with Unstructured Data0.81142
4Who Saw It Coming? Historical Experience and the 2021 Inflation Forecast Failure0.81142
5Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach0.64422
6Inflation Attitudes of Large Language Models0.64422
7Fake Date Tests: Can We Trust In-sample Accuracy of LLMs in Macroeconomic Forecasting?0.64422
8How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge0.64422
9The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective0.51121
10Inference for Regression with Variables Generated by AI or Machine Learning0.40511