arXiv 10 May 2026 · Econometrics
arXiv:2605.09712 · PDF · DOI · OpenAlex · Extracted main text
Average forecast accuracy is not the same as forecast reliability. I treat forecast loss differentials relative to a benchmark as a return series. I then evaluate these returns using risk-adjusted performance measures from finance, including the Sharpe ratio, Sortino ratio, Omega ratio, and drawdown-based metrics. I also introduce the Edge Ratio capturing a model's propensity to deliver uniquely informative predictions relative to the forecasting frontier. I apply this framework to U.S. macroeconomic forecasting, comparing econometric benchmarks, machine learning models, a foundation model (TabPFN), and the Survey of Professional Forecasters. While it is often feasible to beat professional forecasters in terms of average accuracy, it is much harder to beat them on a risk-adjusted basis. They rarely exhibit catastrophic failures and often achieve high Edge Ratios, plausibly reflecting the value of contextual judgment. Nonetheless, selected machine learning methods deliver attractive risk profiles for specific targets. The framework naturally extends to meta-analyses across targets, horizons, and samples, illustrated with a density forecast evaluation and the M4 competition.
appendix boundary found by appendix_command · 66% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Goulet Coulombe, P., Frenette, M., and Klieber, K (2026) From reactive to proactive volatility modeling with hemisphere neural networks | 0.894 | 7 | 3 | 71% |
| 2 | Hollmann, N., Müller, S., Eggensperger, K., and Hutter, F (2022) TabPFN: A transformer that solves small tabular classification problems in a second | 0.737 | 3 | 2 | 100% |
| 3 | Goulet Coulombe, P (2026) LGB+: A macroeconomic forecasting road test | 0.737 | 3 | 2 | 100% |
| 4 | Alam, M. J., Boyle, S., Li, H., and Sekhposyan, T (2025) ChatMacro: Evaluating inflation forecasts of generative AI | 0.644 | 2 | 2 | 100% |
| 5 | Chekhlov, A., Uryasev, S., and Zabarankin, M (2005) Drawdown measure in portfolio optimization | 0.644 | 2 | 2 | 100% |
| 6 | Goulet Coulombe, P., Göbel, M., and Klieber, K (2025) Dual interpretation of machine learning forecasts | 0.644 | 2 | 2 | 100% |
| 7 | Gneiting, T (2011) Making and evaluating point forecasts | 0.644 | 2 | 2 | 100% |
| 8 | Keating, C. and Shadwick, W. F (2002) A universal performance measure | 0.644 | 2 | 2 | 100% |
| 9 | Magdon-Ismail, M. and Atiya, A. F (2004) Maximum drawdown | 0.644 | 2 | 2 | 100% |
| 10 | Makridakis, S., Spiliotis, E., and Assimakopoulos, V (2018) The M4 competition: Results, findings, conclusion and way forward | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 56 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | 0.5cm dpd LGB+: A Macroeconomic Forecasting Road Test . 0.25cm | 0.405 | 1 | 1 |