Alexander Eliseev, Sergei Seleznev
arXiv 12 Jan 2026 · Econometrics
arXiv:2601.07992 · PDF · DOI · OpenAlex · Extracted main text
Large language models (LLMs) are a type of machine learning tool that economists have started to apply in their empirical research. One such application is macroeconomic forecasting with backtesting of LLMs, even though they are trained on the same data that is used to estimate their forecasting performance. Can these in-sample accuracy results be extrapolated to the model's out-of-sample performance? To answer this question, we developed a family of prompt sensitivity tests and two members of this family, which we call the fake date tests. These tests aim to detect two types of biases in LLMs' in-sample forecasts: lookahead bias and context bias. According to the empirical results, none of the modern LLMs tested in this study passed our first test, signaling the presence of lookahead bias in their in-sample forecasts.
appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Hansen, Anne Lundgaard and Horton, John J. and Kazinnik, Sophia and… (2025) Simulating the Survey of Professional Forecasters | 0.644 | 4 | 1 | 100% |
| 2 | Ludwig, Jens and Mullainathan, Sendhil and Rambachan, Ashesh (2025) Large Language Models: An Applied Econometric Framework | 0.644 | 2 | 2 | 100% |
| 3 | Faria-e-Castro, Miguel and Leibovici, Fernando (2024) Artificial Intelligence and Inflation Forecasts | 0.511 | 2 | 1 | 100% |
| 4 | Lin, Jianhao and Sun, Lexuan and Yan, Yixin (2025) Simulating Macroeconomic Expectations using LLM Agents | 0.511 | 2 | 1 | 100% |
| 5 | Paleka, Daniel and Goel, Shashwat and Geiping, Jonas and Tramèr, Flo… (2025) Pitfalls in Evaluating Language Model Forecasters | 0.511 | 2 | 1 | 100% |
| 6 | Ritzwoller, David M. and Romano, Joseph P. and Shaikh, Azeem M (2025) Randomization Inference: Theory and Applications | 0.511 | 2 | 1 | 100% |
| 7 | Sarkar, Suproteem K. and Vafa, Keyon (2024) Lookahead Bias in Pretrained Language Models | 0.511 | 2 | 1 | 100% |
| 8 | Tomáš, Adam and Aleš, Michl and Josef, Svéda (2025) First use of AI in inflation forecasting at the CNB | 0.511 | 2 | 1 | 100% |
| 9 | Zarifhonarvar, Ali (2026) Generating inflation expectations with large language models | 0.511 | 2 | 1 | 100% |
| 10 | Acemoglu, Daron (2025) The simple macroeconomics of AI | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 40 scored citations.