EconBase
← All papers

Statistical Inference for Score Decompositions

Timo Dimitriadis, Marius Puke

arXiv 4 Mar 2026 · Econometrics

arXiv:2603.04275 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We introduce inference methods for score decompositions, which partition scoring functions for predictive assessment into three interpretable components: miscalibration, discrimination, and uncertainty. Our estimation and inference relies on a linear recalibration of the forecasts, which is applicable to general multi-step ahead point forecasts such as means and quantiles due to its validity for both smooth and non-smooth scoring functions. This approach ensures desirable finite-sample properties, enables asymptotic inference, and establishes a direct connection to the classical Mincer-Zarnowitz regression. The resulting inference framework facilitates tests for equal forecast calibration or discrimination, which yield three key advantages. They enhance the information content of predictive ability tests by decomposing scores, deliver higher statistical power in certain scenarios, and formally connect scoring-function-based evaluation to traditional calibration tests, such as financial backtests. Applications demonstrate the method's utility. We find that for survey inflation forecasts, discrimination abilities can differ significantly even when overall predictive ability does not. In an application to financial risk models, our tests provide deeper insights into the calibration and information content of volatility and Value-at-Risk forecasts. By disentangling forecast accuracy from backtest performance, the method exposes critical shortcomings in current banking regulation.

Citation extraction

106
references
261
in-text mentions
106
distinct cited
10
self-citations
17,560
main-text words

appendix boundary found by appendix_command · 47% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Wagner Piazza Gaglianone and Luiz Renato Lima and Oliver Linton and… (2011) Evaluating Value-at-Risk Models via Quantile Regression1.00084100%
2Dimitriadis, Timo and Gneiting, Tilmann and Jordan, Alexander I (2021) Stable reliability diagrams for probabilistic classifiers self1.00083100%
3Tilmann Gneiting and Johannes Resin (2023) Regression diagnostics meets forecast evaluation: Conditional calibration, reliability diagrams, and coefficient of determination0.96911391%
4Marc-Oliver Pohle (2020) The Murphy Decomposition and the Calibration-Resolution Principle: A New Perspective on Forecast Evaluation0.9285480%
5Bayer, Sebastian and Dimitriadis, Timo (2022) Regression-based Expected Shortfall backtesting self0.92844100%
6Olivier Coibion and Yuriy Gorodnichenko (2015) Information Rigidity and the Expectations Formation Process: A Simple Framework and New Facts0.92844100%
7Basel Committee (2019) Minimum capital requirements for market risk0.92843100%
8Ehm, Werner and Gneiting, Tilmann and Jordan, Alexander and Krüger,… (2016) Of Quantiles and Expectiles: Consistent Scoring Functions, Choquet Representations and Forecast Rankings0.87452100%
9Tilmann Gneiting (2011) Making and Evaluating Point Forecasts0.87452100%
10Andrew J. Patton (2020) Comparing Possibly Misspecified Forecasts0.87452100%

Showing the top 10 of 106 scored citations.