EconBase
← All papers

Uncertainty Quantification in Forecast Comparisons

Marc-Oliver Pohle, Tanja Zahn, Sebastian Lerch

arXiv 5 May 2026 · Statistics — Methodology

arXiv:2605.03997 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Skill scores, which measure the relative improvement of a forecasting method over a benchmark via consistent scoring functions and proper scoring rules, are a standard tool in forecast evaluation, yet their sampling uncertainty is rarely rigorously quantified. With modern forecasting applications being increasingly multivariate and involving evaluations across multiple horizons, variables, spatial locations, and forecasting methods, standard tools like the pairwise Diebold-Mariano forecast accuracy test or pointwise confidence intervals fail to account for the multiple comparison problem, leading to inflated Type I error rates and invalid joint inference. To address the lack of a coherent, statistically rigorous framework for quantifying uncertainty across these multi-dimensional evaluation problems, we introduce simultaneous confidence bands for expected scores and skill scores. Our framework provides a versatile tool for joint inference that is applicable to any forecast type from mean and quantile to full distributional forecasts. We develop a bootstrap implementation and show that our bands are valid under multivariate extensions of the classical Diebold-Mariano assumptions. We demonstrate the practical utility of the approach in two case studies by quantifying the benefits of time-varying parameter models for macroeconomic forecasting, and by comparing data-driven and physics-based models in probabilistic weather forecasting.

Citation extraction

47
references
96
in-text mentions
47
distinct cited
7
self-citations
13,662
main-text words

appendix boundary found by appendix_command · 75% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Fitzenberger, Bernd (1997) The moving blocks bootstrap and robust inference for linear least squares and quantile regressions0.92844100%
2Knüppel, Malte and Krüger, Fabian and Pohle, Marc-Oliver (2022) Score-based calibration testing for multivariate forecast distributions self0.84333100%
3Lahiri, Soumendra Nath (2003) Resampling methods for dependent data0.8115280%
4Gneiting, Tilmann and Raftery, Adrian E (2007) Strictly proper scoring rules, prediction, and estimation0.81142100%
5Gneiting, Tilmann (2011) Making and evaluating point forecasts0.81142100%
6Primiceri, Giorgio E (2005) Time Varying Structural Vector Autoregressions and Monetary Policy0.81142100%
7Bi, Kaifeng and Xie, Lingxi and Zhang, Hengheng and Chen, Xin and Gu… (2023) Accurate medium-range global weather forecasting with 3D neural networks0.73732100%
8Bülte, Christopher and Horat, Nina and Quinting, Julian and Lerch, S… (2026) Uncertainty quantification for data-driven weather models self0.73732100%
9Diebold, Francis X and Mariano, Roberto S (1995) Comparing Predictive Accuracy0.73732100%
10Diebold, Francis X (2015) Comparing predictive accuracy, twenty years later: A personal perspective on the use and abuse of Diebold–Mariano tests0.73732100%

Showing the top 10 of 47 scored citations.