Masoud Soleimani
arXiv 26 Nov 2025 · Finance — Risk Management
arXiv:2512.07867 · PDF · Extracted main text
We develop a transparent and fully auditable LLM-based pipeline for macro-financial stress testing, combining structured prompting with optional retrieval of country fundamentals and news. The system generates machine-readable macroeconomic scenarios for the G7, which cover GDP growth, inflation, and policy rates, and are translated into portfolio losses through a factor-based mapping that enables Value-at-Risk and Expected Shortfall assessment relative to classical econometric baselines. Across models, countries, and retrieval settings, the LLMs produce coherent and country-specific stress narratives, yielding stable tail-risk amplification with limited sensitivity to retrieval choices. Comprehensive plausibility checks, scenario diagnostics, and ANOVA-based variance decomposition show that risk variation is driven primarily by portfolio composition and prompt design rather than by the retrieval mechanism. The pipeline incorporates snapshotting, deterministic modes, and hash-verified artifacts to ensure reproducibility and auditability. Overall, the results demonstrate that LLM-generated macro scenarios, when paired with transparent structure and rigorous validation, can provide a scalable and interpretable complement to traditional stress-testing frameworks.
appendix boundary found by appendix_command · 91% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Aikman, David and Angotti, Riccardo and Budnik, Katarzyna (2024) Stress Testing with Multiple Scenarios: A Tale on Tails and Reverse Stress Scenarios | 0.644 | 2 | 2 | 100% |
| 2 | Staudinger, Moritz and Kusa, Wojciech and Piroi, Florina and Lipani,… (2024) A reproducibility and generalizability study of large language models for query generation | 0.644 | 2 | 2 | 100% |
| 3 | Jorion, Philippe (1997) Value at risk: the new benchmark for managing financial risk | 0.511 | 2 | 1 | 100% |
| 4 | Alfaro, Rodrigo A. and Drehmann, Mathias (2009) Macro Stress Tests and Crises: What Can We Learn? | 0.405 | 1 | 1 | 100% |
| 5 | Araci, Dogu (2019) FinBERT: Financial Sentiment Analysis with Pre-Trained Language Models | 0.405 | 1 | 1 | 100% |
| 6 | Asai, Akari and Wu, Zeqiu and Wang, Yizhong and Sil, Avirup and Haji… (2024) Self-RAG: Learning to Retrieve, Generate, and Critique Through Self-Reflection | 0.405 | 1 | 1 | 100% |
| 7 | Baer, Michael and Gasparini, Marta and Lancaster, Robin and Ranger,… (2023) “All Scenarios Are Wrong, but Some Are Useful”–-Toward a Framework for Assessing and Using Current Climate Risk Scenarios Within… | 0.405 | 1 | 1 | 100% |
| 8 | Best, Philip (2000) Implementing Value at Risk | 0.405 | 1 | 1 | 100% |
| 9 | Bollerslev, Tim (1986) Generalized Autoregressive Conditional Heteroskedasticity | 0.405 | 1 | 1 | 100% |
| 10 | Borio, Claudio and Drehmann, Mathias and Tsatsaronis, Kostas (2014) Stress-Testing Macro Stress Testing: Does It Live Up to Expectations? | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 78 scored citations.