Giuseppe Arbia, Luca Morandini, Vincenzo Nardelli
arXiv 4 Jun 2025 · Computers and Society
arXiv:2506.06377 · PDF · DOI · OpenAlex · Extracted main text
This paper investigates Large Language Models (LLMs) ability to assess the economic soundness and theoretical consistency of empirical findings in spatial econometrics. We created original and deliberately altered "counterfactual" summaries from 28 published papers (2005-2024), which were evaluated by a diverse set of LLMs. The LLMs provided qualitative assessments and structured binary classifications on variable choice, coefficient plausibility, and publication suitability. The results indicate that while LLMs can expertly assess the coherence of variable choices (with top models like GPT-4o achieving an overall F1 score of 0.87), their performance varies significantly when evaluating deeper aspects such as coefficient plausibility and overall publication suitability. The results further revealed that the choice of LLM, the specific characteristics of the paper and the interaction between these two factors significantly influence the accuracy of the assessment, particularly for nuanced judgments. These findings highlight LLMs' current strengths in assisting with initial, more surface-level checks and their limitations in performing comprehensive, deep economic reasoning, suggesting a potential assistive role in peer review that still necessitates robust human oversight.
appendix boundary found by appendix_command · 97% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | DeepSeek AI (2025) Deepseek-r1 | 0.693 | 6 | 1 | 100% |
| 2 | Arbia, G (2014) A Primer for Spatial Econometrics, Volume 230 of Palgrave Texts in Econometrics self | 0.644 | 2 | 2 | 100% |
| 3 | Anthropic (2025) Claude 3.7 sonnet and claude code | 0.585 | 3 | 1 | 100% |
| 4 | Anthropic (2025) Claude 3.7 sonnet system card | 0.585 | 3 | 1 | 100% |
| 5 | DeepSeek AI (2025) Deepseek-v3 | 0.585 | 3 | 1 | 100% |
| 6 | Llama Team (2025) Llama 3.3 model card and prompt format | 0.585 | 3 | 1 | 100% |
| 7 | OpenAI (2025a, March 25) (2025) Introducing 4o image generation | 0.585 | 3 | 1 | 100% |
| 8 | Birhane, A., A. Kasirzadeh, D. Leslie, and S. Wachter (2023) Science in the age of large language models | 0.511 | 2 | 1 | 100% |
| 9 | DeepSeek API Docs (2025, March 25) (2025) Deepseek-v3-0324 release | 0.511 | 2 | 1 | 100% |
| 10 | Hosseini, M. and S. P. Horbach (2023) Fighting reviewer fatigue or amplifying bias? considerations and recommendations for use of ChatGPT and other large language mod… | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 27 scored citations.