Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
127,583 characters · 23 sections · 137 citation commands
Statistical Inference for Score Decompositions
Forecasting---whether of macroeconomic conditions, financial market movements, or other uncertain outcomes---has become an indispensable element of modern economic analysis. From monetary policy design to investment planning and risk management, economic decisions increasingly rely on the ability to anticipate what lies ahead. The value of these decisions, however, depends critically on the accuracy and credibility of the underlying forecasts. Therefore, establishing rigorous standards for evaluating predictions is essential, ensuring that economic forecasts are not only informative but also empirically grounded and verifiable.
Two main concepts have emerged for evaluating a forecast $X_t$ of an uncertain future real-valued outcome $Y_t$. First, the statistical theory of forecast comparison ranks competing forecasts according to their empirical score (or loss) S_1971,G_2011a. The appropriate scoring function depends on the forecast target---e.g., whether the aim is to predict the full distribution of $Y_t$, its mean, or a particular quantile. For meaningful comparisons, these scoring functions must be strictly consistent: the ideal forecast should minimize the expected score, ensuring that lower score values reliably indicate superior predictive performance.
Second, forecasts are evaluated in terms of their alignment with the realized outcomes---a property referred to as calibration in statistics Gneiting2007probabilistic, GR_2023, forecast optimality, rationality or efficiency in economics D_1998, Elliott2016, and backtesting in finance B_2023_vola,B_2019. For example, when forecasting the mean, the average forecast should coincide with the average realized outcome—both unconditionally and conditional on the forecast itself, the latter being referred to as auto-calibration. In economics, this property is most commonly assessed using the test of MZ_1969, which examines whether the forecast $X_t$ equals the conditional expectation of $Y_t$ given $X_t$ through a linear regression framework. Extensions to other targets, such as quantiles, follow naturally by replacing the mean regression with the corresponding quantile regression Gaglianone2011, Guler2017, BD_2020.
These two evaluation paradigms are known to occasionally produce counterintuitive results, a point illustrated by our following motivational example. In this paper, we develop asymptotic inference for the components of so-called score decompositions, which provide a transparent link between the two forecast evaluation regimes described above. Beyond delivering a sharper theoretical understanding of the respective strengths and limitations of these approaches, the motivational example underscores their practical relevance by demonstrating how, in applied settings, the two methods can lead to seemingly contradictory conclusions.
Consider the daily logarithmic returns of the S&P E-mini future as the world’s most actively traded equity index futures contract. We evaluate one-day ahead forecasts for its variance as well as its $1\%$-quantile, known as the Value-at-Risk (VaR), between January 1, 2008 and January 1, 2022. Both forecast types are important tools in banking regulation; see B_2023_vola and B_2019 for volatility and VaR backtesting, respectively. Section (ref) provides additional details on this application.
We use the following state of the art forecasting models: First, the Heterogeneous Auto-Regressive (HAR) model of C_2009 uses intraday trading information through the Realized Variance (RV) estimator to forecast the variance through a linear model. Second, we employ a baseline GARCH model of B_1986 and for both the HAR and GARCH model, we obtain quantile forecasts through a Gaussianity assumption on the model residuals. For the case of VaR forecasts, we further employ the basic Historical Simulation (HS) approach that uses the unconditional quantile of the past 250 trading days as its quantile forecast.
Table (ref) shows that, for both forecast types, the HAR model achieves the lowest average scores. These improvements are statistically significant at the $1\%$ level according to the DM_1995 test, with the sole exception of the HAR–GARCH comparison for VaR forecasts. Thus, the HAR forecasts deliver the strongest predictive performance. At the same time, the MZ test decisively rejects auto-calibration of the HAR forecasts. In contrast, the GARCH variance forecasts yield significantly higher scores, yet auto-calibration cannot be rejected. A similar pattern emerges for the HS VaR forecasts: they record the largest scores but constitute the only forecasting sequence for which calibration is not rejected.
This example illustrates that good predictive performance and proper auto-calibration do not necessarily coincide and may diverge sharply. Since both evaluation approaches are widely applied in practice, developing tools to reconcile and interpret these seemingly contradictory conclusions is of central importance.
We shed light on these results by examining the score decomposition reported in Table (ref). Score decompositions have a long tradition in meteorology M_1973,Dawid1986,MW_1987,B_2009,BF_2014 and have recently attracted renewed interest in statistics DGJ_2021,GR_2023,G_etal_2023, yet they have received comparatively little attention in economics. In essence, the average score $\mathsf{S}_i$ of forecaster $i$ can be decomposed additively into three non-negative components capturing miscalibration ($\mathsf{MCB}_i$), discrimination ($\mathsf{DSC}_i$), and a forecast-independent uncertainty term ($\mathsf{UNC}$):
Intuitively, the $\mathsf{MCB}$ component measures the extent of miscalibration relative to the scoring function by quantifying how much the score would improve if the forecasts were perfectly calibrated---thus capturing a concept closely related to what the MZ test evaluates. Likewise, the $\mathsf{DSC}$ component reflects the forecaster’s ability to distinguish (or “discriminate”) between higher and lower outcome realizations, conditional on the forecasts being calibrated. The remaining $\mathsf{UNC}$ term does not depend on the forecast itself and therefore serves as a normalization.
The last columns of Table (ref) help reconcile the discrepancies between the score-based and calibration-based assessments. For both forecast types, the HAR model attains the highest (and thus best) $\mathsf{DSC}$ values, which is intuitive given that it incorporates the richest information set through high-frequency returns. However, it also yields the largest (worst) $\mathsf{MCB}$ values, accounting for its rejection in the MZ test. By contrast, the GARCH model---and even more so the HS model---exhibits weaker discrimination but achieves substantially better calibration. Yet, despite the clear interpretive value of these decomposition terms, formal inference methods for them are not available, limiting their practical usefulness as reflected in ET_2016, who point out that “am important limitation of these decompositions is that there are no objective measures for how large or small the terms should be”.
This gap is particularly problematic in settings such as banking regulation under the Basel framework, where calibration-based backtesting plays a central role. Such procedures risk creating misguided incentives: the best-performing HAR forecasts may be labeled “invalid,” whereas HS forecasts---whose poor discriminatory ability makes them slow to respond to emerging crises---may be deemed acceptable. These tensions highlight the need for both a deeper understanding of the distinct properties captured by the decomposition and the development of rigorous statistical inference for its components---needs that this paper directly addresses.
In this paper, we propose augmenting scoring-function-based forecast comparisons with the corresponding decomposition terms, as illustrated in Table (ref), and we are the first to develop asymptotic inference for these components. To estimate the decomposition components in (ref), we advocate the use of a simple linear regression---either mean or quantile, depending on the forecast target. This specification is motivated by several considerations: its stability and resistance to overfitting; its good empirical fit in our applications; its close conceptual connection to the MZ_1969 framework; its attractive non-negativity guarantees for the resulting decomposition terms; and, crucially, its tractability for developing valid asymptotic inference procedures. Recent contributions have suggested nonparametric (isotonic) regression approaches instead DGJ_2021, GR_2023, arnold2024decompositions. However, their asymptotic properties remain unclear---particularly in time-series settings---and they are prone to overfitting, which can introduce biases that further complicate or even undermine asymptotic inference.
We derive the joint asymptotic distribution of the estimated $\mathsf{MCB}$ and $\mathsf{DSC}$ components for competing forecasts allowing for possibly non-smooth scoring functions, thereby enabling formal tests of equal miscalibration or equal discrimination for, e.g., mean or quantile forecasts. These tests extend the classical framework of DM_1995 by incorporating two layers of uncertainty: the sampling variation arising from averaging over time and the additional variation introduced by the linear “recalibration regression” used in constructing the $\mathsf{MCB}$ and $\mathsf{DSC}$ components. A key complication is that the limiting distribution depends on whether the population values of $\mathsf{MCB}$ or $\mathsf{DSC}$ are zero or strictly positive. Our testing procedures account for these cases through a $p$-value adjustment inspired by a hybrid intersection–union/union–intersection (IU/UI) testing principles Berger1997likelihood.
Our simulations demonstrate that the proposed tests are valid—though potentially conservative due to the underlying IU/UI principles---under their respective null hypotheses for both mean and quantile forecasting settings. We also show that the tests exhibit strong power properties, and in particular achieve a substantial power increase relative to the classical DM test in fair comparisons where differences in average scores stem exclusively from a single component, either $\mathsf{MCB}$ or $\mathsf{DSC}$. This perhaps surprising power gain arises because our tests isolate a single aspect of forecast performance---either discrimination or miscalibration---whereas score-based tests must aggregate both components into a single, and consequently noisier, metric.
Our first application compares the predictive performance of two survey forecasts for U.S. inflation—the Survey of Professional Forecasters (SPF) and the Michigan Survey of Consumers. While professional forecasters unsurprisingly attain better average scores than consumers, these differences are typically insignificant EGJK_2016, P_2020. In contrast, the score decomposition shows that both groups exhibit similar calibration, whereas our new tests detect a significant difference in discrimination. This finding is consistent with the power gains demonstrated in our simulations. Given that professional forecasters draw on a richer information set, their superior discrimination is a natural consequence.
The second application extends the motivating example of Section (ref) and clarifies the relationship between scoring-function-based forecast evaluation and backtests for both, variance and VaR forecasts. Moreover, we disentangle several widely used VaR backtests by characterizing them through the conditioning sets they implicitly impose within the framework of conditional quantile calibration, thereby revealing the precise properties they actually test.
Early approaches to score decompositions resembling (ref) date back to sanders1963subjective, T_1966, and M_1973, and primarily focus on the squared error in the context of probability forecasts for binary outcomes, which can be interpreted as a mean. degroot1981assessing and B_2009 demonstrated that such decompositions can be derived for any strictly proper scoring rule, including those for full distributional forecasts. More recently, GR_2023 developed a general framework for score decompositions for scoring functions, extending these concepts to point forecasts. This line of research has spurred a growing body of both theoretical and applied work; see, for example, BF_2014, S_2017, Pohle2020, FLM_2022, Wetal_2025, Betal_2024. For a more detailed historical overview of score decompositions, see mitchell2019score.
In practice, the score decompositions rely on estimates of recalibrated forecasts, and the choice of estimation method remains an active area of research. Recently, isotonic regression—which estimates, for example, a conditional mean nonparametrically under a monotonicity constraint—has been widely used for both point and distributional forecasts DGJ_2021, G_etal_2023, DGJV_2024, arnold2024decompositions, allen2025assessing. A key reason for its popularity is that it satisfies attractive finite-sample non-negativity conditions, which facilitate the interpretation of the decomposition terms; these conditions are generally not satisfied by alternative methods such as binning or kernel regressions DGJ_2021. We show that, under mild conditions, these non-negativity properties are also preserved when recalibration is performed using linear regressions.
As we develop asymptotic inference methods for the estimated decomposition terms on the right-hand side of (ref), our approach is closely related to statistical tests for predictive performance, which focus on the sampling variability of the left-hand side of (ref). The seminal contribution of DM_1995 laid the foundation for a large econometric literature on testing for equal forecast performance. Numerous extensions have since been proposed, including methods for model-based forecasts West1996asymptotic, clark2001tests, multiple forecast comparisons Hansen2005test, HLN_2011, conditional predictive ability Giacomini2006tests, Li2022conditional, and anytime-valid procedures Henzi2022valid, Choe2024comparing. For comprehensive surveys of this literature, see, for example, west2006forecast, ClarkMcCracken2013, and Diebold2015comparing.
In contrast, the assessment of sampling uncertainty for the right-hand side of (ref) has received relatively little attention. Existing contributions, such as S_2014 and GR_2023, propose ad-hoc (resampling) methods to quantify the uncertainty for a single forecaster, but they do not provide tools for comparative inference across multiple forecasters. Our work addresses this gap by developing a unified framework for inference on the decomposition components.
VaR backtests are widely used to evaluate the reliability, i.e., the calibration, of risk models via their quantile forecasts. Classic procedures include the unconditional coverage test of K_1995, which underlies the traffic-light system of B_1996, B_2019, as well as the conditional coverage test of C_1998, which additionally assesses the independence of exceedances. A large body of subsequent work tests slightly differing notions of conditional calibration BCP_2011, Gaglianone2011, BCP_2011, HD_2023, FGP_2023, wang2025backtesting; see HKP_2022 for a recent overview. Our score decompositions---and in particular the application in Section (ref)---show that well-calibrated forecasts do not necessarily exhibit strong predictive performance. This underscores the importance of evaluating VaR forecasts through both calibration and discrimination, rather than relying solely on traditional backtesting metrics; also see Section (ref).
A growing strand of the macroeconomic literature---including, among many others, CG_2015, BGMS_2020, kohlhas2021asymmetric, AHS_2021, broer2024forecaster, and FNS_2024---examines how forecast errors respond to observable information. This line of work effectively tests conditional calibration, asking whether forecast errors are predictable given the information available at the time forecasts are made. Our methodology suggests a natural refinement of these approaches by additionally assessing forecast discrimination, since calibration captures only one dimension of predictive performance. This refinement is particularly relevant because forecasts that discriminate well between outcomes are often found to be miscalibrated, and thus appear to deviate from rational expectations more frequently.
Section (ref) introduces key concepts of forecast evaluation. The theory in Section (ref) develops asymptotic inference under high-level conditions, which are subsequently verified for mean and quantile forecast. Section (ref) contains simulations and we present two applications in Section (ref). Section (ref) concludes. Appendix (ref) contains all proofs and we discuss the relation to further decompositions in Appendix (ref). Appendix (ref) presents technical details for the simulations and Appendix (ref) contains additional figures.
We provide replication material at \OldHref{https://github.com/marius-cp/SDI_replications}{https://github.com/marius-cp/SDI_replications} for the simulations and applications, which draws on the SDI package for the statistical software R R2026, available at \OldHref{https://github.com/marius-cp/SDI}{https://github.com/marius-cp/SDI}.
After fixing notation in Section (ref), we discuss relative and absolute forecast evaluation in Sections (ref) and (ref).
We follow G_2011a and HogaDimi2023 and consider a complete probability space $(\Omega, \mathcal{F}, \mathbb{P})$ with a series of stationary random variables $Y_t: \Omega \to \mathsf{O} \subseteq \mathbb{R}$, $t \in \mathbb N$, that take values in the interval-valued observation domain $\mathsf{O}$. The setting allows for time series (where $t$ represents time points) as well as cross-sectional (where $t$ represents cross-sectional units) applications. For some information set (in the form of a sigma-algebra) $\mathcal{F}_t \subset \mathcal{F}$, where $Y_t$ is in general not $\mathcal{F}_t$-measurable, we denote the conditional distribution of $Y_t$ given $\mathcal{F}_t$ by $F_t(\bullet, \omega) = \mathbb P(Y_t \le \bullet \mid \mathcal{F}_t)(\omega)$, where we assume that $F_t(\bullet, \omega) \in \mathcal{P}$ for all $\omega \in \Omega$, for some convex class of distributions $\mathcal{P}$ taking values in $\mathsf{O}$, which contains the Dirac measures $\delta_y$ for all $y \in \mathsf{O}$. We write $F_t(\bullet)$ for the random variable defined by $\omega \mapsto F_t(\omega, \bullet)$. The information set $\mathcal{F}_t$ represents the full information that could be used for forecast $Y_t$, but competing forecasters might have less information at hand.
For an interval-valued action domain $\mathsf{A} \subseteq \mathbb{R}$, we consider a stationary sequence of forecasts for some target functional $\Gamma: \mathcal{P} \to \mathsf{A}$ of the conditional distribution $F_t$. Prominent functionals are the conditional expectation $\mathbb{E} [ Y_t \mid \mathcal{F}_{t}]$, or a conditional $\alpha$-quantile $Q_\alpha(Y_t \mid \mathcal{F}_{t})$ for some $\alpha \in (0,1)$, which is assumed to be unique for simplicity. We sometimes abuse notation and write $\Gamma( Y_t \mid \mathcal{F}_{t})$ instead of $\Gamma({F}_{t})$ for a functional of the conditional distribution.
The point forecasts are denoted by $X_{it}$ where the index $i=1,2$ captures a series of competing forecasts. These forecasts are assumed to be $\mathcal{F}_{t}$-measurable throughout the article but they are in general based on different (less) information, $\mathcal{I}_{it} \subseteq \mathcal{F}_{t}$, or might be misspecified through other aspects such as model mispecification or estimation noise.
Throughout, (in)equalities of random variables are meant to hold almost surely (a.s.) unless stated otherwise, and for a vector $\mathbf{z} \in \mathbb{R}^k$, $k \in \mathbb{N}$, $\lVert \mathbf{z} \rVert$ denotes the Euclidean norm and $\mathbf{z}^\top$ its transpose. We write $\nabla_{\mathbf{z}} f$ for the (column-vector) gradient of the mapping $\mathbf{z} \mapsto f(\mathbf{z}) \in \mathbb{R}$ and $\nabla_{\mathbf{z} \mathbf{z}} f$ for its second derivative, the Hessian matrix. In the univariate case, $f'(z)$ denotes the derivative of $f(z)$ with respect to $z \in \mathbb{R}$. For functions with multiple real-valued arguments, a prime indicates partial differentiation with respect to the first argument, and multiple primes denote higher-order derivatives.
Following G_2011a, point forecasts are compared by strictly consistent scoring (or loss) functions $\mathsf{S}: \mathsf{A} \times \mathsf{O} \to \mathbb{R}$, where lower values indicate higher predictive ability. A scoring function $\mathsf{S}$ is said to be consistent for $\Gamma(F)$ with respect to a (usually large) class of distributions $\mathcal{P}$ if
for all $F \in \mathcal{P}$ and all (deterministic) $x \in \mathsf{A}$, where $\mathbb{E}_{Y \sim F}[\,\cdot\,]$ denotes the expectation with respect to $Y \sim F \in \mathcal{P}$. A scoring function is said to be strictly consistent if equality in (ref) implies $x = \Gamma(F)$ such that the optimal forecast $\Gamma(F)$ uniquely obtains the lowest expected score. A functional is said to be elicitable if a strictly consistent scoring function exists.
HE_2014 allows to define consistency of a scoring function relative to a generic information set $\mathcal{G}_t \subseteq \mathcal{F}_t$ as
for all $Y_t \mid \mathcal G_t \in \mathcal{P}$, and for all $\mathcal{G}_t$-measurable random variables $\widetilde{X}_t$. The scoring function is strictly consistent if almost sure equality in (ref) implies that $\widetilde{X}_t = \Gamma(Y_t \mid \mathcal{G}_t)$ a.s..
If the target functional $\Gamma$ is the conditional mean, under mild regularity conditions, all strictly consistent scoring functions are characterized by the so-called Bregman class,
parametrized by a strictly convex function $\phi$ with subgradient $\phi'$ G_2011a. The omnipresent squared error
arises for $\phi_{\textsf{SE}}(x) = x^2$. For $\phi_{\textsf{QL}}(x) = -\log(x)$ we obtain the QLIKE scoring function,
which is regularly used for the evaluation of variance forecasts in financial econometrics with positive action and observation domain, $\mathsf{A} = \mathsf{O} = (0, \infty)$; see e.g., P_2011. Under the standard assumption of a zero mean, the variance coincides with the second moment, thereby rendering the Bregman class applicable.
If the target functional is the conditional $\alpha$-quantile, $\alpha \in (0,1)$, all strictly consistent scoring functions are given by the class of generalized piecewise linear (GPL) loss functions
where the function $g$ is non-decreasing on $\mathsf{A} = \mathsf{O} \subseteq \mathbb{R}$ G_2011a. All GPL loss functions are non-differentiable at $x=y$ due to the indicator function. The check loss $\mathsf{S}_\textsf{CL}(x,y) = (\mathds{1}\{y \le x\} - \alpha) (x - y)$ as the most popular candidate arises for $g_{\textsf{CL}}(x) = x$.
In practice, scoring functions are employed to compare forecast sequences $X_{it}$, $i=1,2$, $t=1,\dots,T$, for the target variable $Y_t$ through their average scores $\widehat{\mathsf{S}}_{iT} = \frac{1}{T} \sum_{t=1}^{T} \mathsf{S}(X_{it}, Y_t)$. The sample average serves as an empirical approximation to the expectation appearing in (ref)--(ref). Following DM_1995, a rich body of asymptotic inference methods is available for constructing tests based on such average score differences.
A known practical limitation of purely score-based forecast evaluations lies in their inability to identify specific deficiencies of the forecasts, such as systematic biases, which we formalize through the theory of forecast calibration. A forecast $X_{it}$ for $\Gamma(Y_t)$ is said to be conditionally calibrated with respect to the information set $\mathcal{G}_{it} \subseteq \mathcal{F}_t$ if
This property asserts a certain compatibility of the forecast $X_{it}$ with its forecasting target $\Gamma(Y_t \mid \mathcal{G}_{it})$. A conditionally calibrated forecast w.r.t.\ the full information set $\mathcal{F}_t$ minimizes the $\mathcal{F}_t$-conditional expectation of the scoring function by (ref), such that the econometric literature also calls it optimal or efficient ET_2016. As the forecast might not have used---or even have access to---all information in $\mathcal{F}_t$, it is, however, often more appropriate to consider calibration based on some reduced information set $\mathcal{G}_{it} \subset \mathcal{F}_t$ in (ref).
An important special case is that of an auto-calibrated forecast, obtained for $\mathcal{G}_{it} = \sigma\{X_{it}\}$, i.e.,
Conditioning on the forecast itself guarantees that the information encoded in the forecast was indeed used, and it remains feasible even when the evaluator does not have access to the full information set of the forecaster $\mathcal{I}_{it}$.
A common tool to assess auto-calibration is to plot $x \in \mathsf{A}$ against an estimate of $\Gamma(Y_t \mid X_{it} = x)$, where deviations from the diagonal indicate miscalibration. In the meteorological literature, this is known as a reliability diagram, typically estimated nonparametrically MurphyWinkler1992, Pohle2020, DGJ_2021. In economic mean forecast evaluation, this idea goes back to MZ_1969, who regress $Y_t$ on $X_{it}$ and test whether the intercept is zero and the slope is one---essentially a linear reliability diagram. Gaglianone2011, Guler2017, BD_2020 cover extensions to other functionals.
A further important tool for calibration assessment are identification functions. Under smoothness conditions, they correspond to derivatives of the scoring function NZ_2017. Formally, an identification function $\mathsf{V}: \mathsf{A} \times \mathsf{O} \to \mathbb{R}$ is increasing and left-continuous in its first argument, and it is called strict if $\mathbb{E}_{Y \sim F} \big[ \mathsf{V}(x, Y) \big] = 0$ if and only if $x = \Gamma(Y)$. For example, $\mathsf{V}(x, y) = 2(x - y)$ is the canonical identification function that arises as the derivative of the squared error loss and for the $\alpha$-quantile, $\mathsf{V}(x, y) = \mathds{1}\{y \le x\} - \alpha$ is the almost sure derivative of the check loss. Identification functions allow for a (subject to regularity conditions) equivalent definition of conditional calibration from (ref) as
which specializes to auto-calibration when $\mathcal{G}_{it} = \sigma\{X_{it}\}$. Taking expectations on both sides of (ref) allows to define unconditional calibration of $X_{it}$ as $\mathbb{E} \big[ \mathsf{V}(X_{it}, Y_t) \big] = 0$.
We introduce linear score decompositions in Section (ref), develop asymptotic inference in Section (ref), discuss testable hypotheses in Section (ref), and derive low-level inference conditions for mean and quantile forecasts in Section (ref). Section (ref) discusses the connection to VaR backtests. Related decompositions are discussed in Section (ref).
For an elicitable functional $\Gamma$ and corresponding strictly consistent scoring function $\mathsf{S}$, we denote the expected score of the forecasts $X_{it}$ by $\overline{\mathsf{S}}_{i} := \mathbb{E}\big[ \mathsf{S}(X_{it},Y_t) \big]$. Throughout this paper, an overline denotes population quantities.
To illustrate the mechanics of the score decomposition, we introduce an “idealized” version of the forecast: the population reCalibrated forecast $X_{it}^\mathsf{C} = \Gamma(Y_t \mid \mathcal{W}_{it})$, where $\mathcal{W}_{it}$ denotes the recalibration information set satisfying $\sigma\{X_{it}\} \subseteq \mathcal{W}_{it} \subseteq \mathcal{F}_t$. By construction, $X_{it}^\mathsf{C}$ uses at least as much information as the original forecast, but is $\mathcal{W}_{it}$-conditional calibrated by definition; see (ref). The canonical auto-calibration case---in line with MZ_1969---arises for $\mathcal{W}_{it} = \sigma\{X_{it}\}$. Richer information sets allow connections to stronger calibration concepts, VaR backtests, and related macroeconomic applications CG_2015. The corresponding reCalibrated score is denoted by $\overline{\mathsf{S}}_{i}^\mathsf{C} = \mathbb{E} \big[ \mathsf{S}(X_{it}^\mathsf{C}, Y_t) \big]$. We further define the Reference forecast $\bar{r} = \Gamma(Y_t)$, which is unconditionally calibrated but constant, and the corresponding Reference score $\overline{\mathsf{S}}^{\mathsf{R}} = \mathbb{E} \big[ \mathsf{S}( \bar{r}, Y_t) \big]$.
With these ingredients, we define the population score decomposition, consisting of non-negative components GR_2023:
In the canonical case $\mathcal{W}_{it} = \sigma\{X_{it}\}$, the component $\overline{\mathsf{MCB}}_{i} = \overline{\mathsf{S}}_{i} - \overline{\mathsf{S}}_{i}^{\mathsf{C}}$ quantifies the increase in the score's predictive ability that can be achieved by recalibration, hence providing a measure of miscalibration such that low $\overline{\mathsf{MCB}}_{i}$ values are desirable. As the recalibrated forecast is a measurable transformation of $X_{it}$, it incorporates no additional information; thus, any score gains stem purely from improved calibration. If $\overline{\mathsf{MCB}}_{i} = 0$, recalibration yields no score improvement, indicating that the original forecast was already calibrated.
The component $\overline{\mathsf{DSC}}_{i} = \overline{\mathsf{S}}^\mathsf{R} - \overline{\mathsf{S}}_{i}^\mathsf{C}$ compares the scores of two recalibrated forecasts: $\overline{\mathsf{S}}_{i}^\mathsf{C}$ based on information in $X_{it}$, and $\overline{\mathsf{S}}^\mathsf{R}$ unconditionally without additional information. As this “controls for calibration”, the remaining score improvement in $\overline{\mathsf{DSC}}_{i}$ reflects the forecast's information content, that is, its discrimination ability. Larger values are therefore desirable. The final component $\overline{\mathsf{UNC}} = \overline{\mathsf{S}}^{\mathsf{R}}$ is forecast-independent and serves as a normalizing constant.
In practice, the components in (ref) require estimated versions of both the recalibrated and reference forecasts. While estimating the reference forecast as the unconditional functional is straightforward, the estimation of the recalibrated forecast remains debated, with proposals including binning methods MW_1987, kernel regression Pohle2020, and more recently isotonic regression DGJ_2021, GR_2023. We instead propose the use of linear regression, which mitigates biases from nonparametric overfitting that can hinder valid inference, particularly in settings of close to auto-calibration, aligns with MZ_1969, and easily accommodates additional covariates beyond the forecasts themselves.
For this, we consider the recalibration variables $\bm W_{it} = (1, X_{it}, \dots)^\top \in \mathbb{R}^k$ that generate the recalibration information $\mathcal{W}_{it} = \sigma\{\bm W_{it}\}$ and impose the following:
Besides stationarity, Assumption (ref) imposes linearity of the recalibration function $X_{it}^\mathsf{C} = \Gamma(Y_t \mid \bm W_{it}) = \bm W_{it}^\top \bar{\boldsymbol{\theta}}_i$, which is required for correct specification of the linear “recalibration regression”. For the finite-sample statements of Theorem (ref) below, it is crucial that we use M-estimation for the regression parameters by using the same scoring function that is used for the decomposition\footnote{While DFZ_CharMest show that strictly consistent loss functions (and only these) produce consistent parameter estimates, the resulting estimates generally differ in value across loss functions Elliott2016, Muhlemann2021, motivating the need to employ the same scoring function in (ref) and (ref).}
where we assume for simplicity that the arg\,min in (ref) is unique. Moreover, the empirical reference forecast $\widehat{r}_T$ is the unconditional functional $\Gamma$ of the sample $Y_1, \dots, Y_T$. Then, the empirical counterpart of the score decomposition in (ref) is given by
where
The following theorem ensures (strict) positivity of the sampling versions of the decomposition in (ref). While similar results for the population decomposition (ref) are established in GR_2023, it is noteworthy that these results extend to the sampling versions under linear recalibration, even without imposing correct specification in Assumption (ref).
The non-negativity guarantees of Theorem (ref) are essential, as zero miscalibration and discrimination components are routinely interpreted as perfect calibration and no discrimination, respectively DGJ_2021, GR_2023, allen2025assessing. The proof of Theorem (ref) applies to any recalibration method estimated by global score minimization and which is flexible enough to nest both a constant and the identity line, such as isotonic regression DGJ_2021, GR_2023. In contrast, kernel regressions Pohle2020, logistic regressions and Beta-CDF models for binary events S_2017, ranjan2010combining do not offer such guarantees. Although kernels can represent constants and the identity line, they are not fitted via global score minimization; Beta-CDFs, being strictly increasing, cannot produce constants; and logistic regressions generally cannot reproduce the identity line.
We now provide new distributional convergence results for the considered linear score decomposition, which are derived under the following high-level regularity conditions, for which we derive more explicit conditions for mean and quantile forecasts in Section (ref).
Assumption (ref) provides high-level conditions on the scoring function and the parametric estimators. These conditions can accommodate a wide range of settings, including various estimation methods, target functionals, and forms of temporal dependence. They are verified for the mean and quantile cases in Section (ref) below, and their generality allows for straightforward extensions to other functionals. In a nutshell, conditions (ref) and (ref) can be verified based on classical dependence and moment conditions on the forecast-realization pairs. Conditions (ref) and (ref) ensure applicability to non-smooth scoring functions, such as the check loss, by imposing smoothness on the expected score and controlling empirical deviations via stochastic equicontinuity Andrews1994, wellner2013weak.
The joint asymptotic approximation of Theorem (ref) is required for the construction of “comparative” confidence intervals and hypothesis tests---for instance, tests for equal miscalibration or discriminatory ability across the two forecast sequences. Notably, the resulting asymptotic distribution---with covariance matrix $\boldsymbol{\Omega}_T$ from (ref)---is unaffected by the estimation uncertainty of $\widehat{\boldsymbol{\theta}}_{iT}$ and $\widehat{r}_T$, due to an orthogonality between these estimators and the sampling error from replacing expected scores with their empirical averages.
Theorem (ref) requires strict positivity of the population values $\overline{\mathsf{MCB}}_{i}, \overline{\mathsf{DSC}}_{i} > 0$, hence excluding the boundary cases $\overline{\mathsf{MCB}}_{i} = 0$ and $\overline{\mathsf{DSC}}_{i} = 0$, which necessitate a separate treatment for which we impose the following high-level assumptions.
As before, the high-level conditions in Assumption (ref) allow for case-by-case verifications for specific functionals, scoring functions, and estimators as exemplified in Section (ref). Condition (ref) provides a standard asymptotic linear representation of the estimators, slightly strengthening Assumption (ref) (ref). Notably, the $\mathsf{MCB}$ result below could be derived using only the representation for $\widehat{\boldsymbol{\theta}}_{iT}$. Condition (ref) applies to smooth loss functions and imposes a uniform LLN type requirement, although this condition is substantially weakened by the $T^{-1/2}$ scaling factor.
Theorem (ref) shows for smooth scoring functions that when the population quantities $\overline{\mathsf{MCB}}_{i}$ or $\overline{\mathsf{DSC}}_{i}$ attain their boundary value of zero, the non-negative test statistics $T\,\widehat{\mathsf{MCB}}_{iT}$ and $T\,\widehat{\mathsf{DSC}}_{iT}$ (see Theorem (ref)) converge to generalized $\chi^2$-distributions. Extending Theorem (ref) to non-smooth scoring functions via the proof strategy of Theorem (ref) would necessitate establishing stochastic equicontinuity of $\sqrt{T} \, \nu_{iT} (\boldsymbol{\theta}) = \sum_{t=1}^T \mathsf{S}( \bm W_{it}^\top \boldsymbol{\theta}, Y_t) - \mathbb{E} \big[ \mathsf{S} (\bm W_{it}^\top \boldsymbol{\theta} , Y_t ) \big]$. To the best of our knowledge, established techniques for such a derivation are currently unavailable in this context.
Recall that the asymptotic variance in the Gaussian case of Theorem (ref) is unaffected by the estimation error in $\widehat{r}_T$ and $\widehat{\boldsymbol{\theta}}_{iT}$; it depends solely on the sampling uncertainty from approximating expected scores by empirical averages, as captured by $\boldsymbol{\Omega}_T$ in (ref). In contrast, the limiting distributions in Theorem (ref) depend exclusively on the estimation effects, as summarized by $\bm\Upsilon_{i}$, $\bm H_{i}$, and $\bm\Pi_{iT}$, which correspond to the population counterparts of the asymptotic expansion in (ref). Technically, this distinction arises because the asymptotic distribution in Theorem (ref) is driven by first-order Taylor expansion terms, whereas the distributions in Theorem (ref) are governed by second-order terms; see the proofs for details.
Feasible inference requires consistent estimators of the population matrices $\boldsymbol{\Omega}_T$, $\bm{\Pi}_{iT}$, $\bm\Upsilon_{i}^{-1}$, and $\bm{H}_{i}^{-1}$ so that the asymptotic results in Theorems (ref) and (ref) remain valid when replacing population quantities by their estimates. While $\bm\Upsilon_{i}^{-1}$ and $\bm H^{-1}_{i}$ admit consistent estimation via sample averages, we use HAC estimators for the long-run covariance matrices $\boldsymbol{\Omega}_T$ and $\bm{\Pi}_{iT}$ to accommodate serial dependence NW_1987,A_1991. Consistency follows by standard arguments, see e.g., W_2001 for smooth and GY_2024 for non-smooth scoring functions.
For “one-sample” hypotheses of the form $\overline{\mathsf{MCB}}_i = c$ or $\overline{\mathsf{DSC}}_i = c$ for some $c \ge 0$, the appropriate asymptotic distribution---Gaussian or generalized $\chi^2$---can be selected directly based on the value of $c$. Under “two-sample” hypotheses such as $\overline{\mathsf{MCB}}_1 = \overline{\mathsf{MCB}}_2$ or $\overline{\mathsf{DSC}}_1 = \overline{\mathsf{DSC}}_2$, however, the population values $\overline{\mathsf{MCB}}_{i}$ and $\overline{\mathsf{DSC}}_{i}$ may be either zero or strictly positive, so it is a priori unclear whether the Gaussian or the generalized $\chi^2$ approximation applies. A valid testing procedure must therefore remain robust to both possibilities. We achieve this by employing a hybrid of intersection--union and union--intersection testing principles Berger1997likelihood.
For example, consider the hypothesis $\mathcal H_\mathsf{MCB} : \overline{\mathsf{MCB}}_{1} = \overline{\mathsf{MCB}}_{2}$, which we can rewrite as
For the hypothesis $\mathcal H^+_{\mathsf{MCB}}$, we employ the selection vector $\boldsymbol{\omega} = (1,0,-1,0)^\top$ such that Theorem (ref) yields
which immediately gives an asymptotically valid $p$-value called $p^{+}_{\mathsf{MCB}}$.
For smooth scoring functions, we test the hypothesis $\mathcal H^0_{{\mathsf{MCB}}_i}$ for $i=1,2$ based on the test statistic $T \, \widehat{\mathsf{MCB}}_{iT}$ and the asymptotic approximation from Theorem (ref). For a realized value of $T \, \widehat{\mathsf{MCB}}_{iT}$, we compute the associated $p$-values $p^0_{{\mathsf{MCB}_i}} = \mathbb P \bigl( 0.5 \bm N_{iT}^\top \bm\Upsilon_{i}^{-1} \bm N_{iT} > T \, \widehat{\mathsf{MCB}}_{iT} \bigr)$ by using the method of I_1961, which uses an inversion formula to express the probability as a definite integral of the characteristic function that is solved via numerical integration by using the R package CompQuadForm CompQuadFormPackage. For non-smooth scoring functions as the GPL score from (ref), we test the hypothesis $\mathcal H^0_{{\mathsf{MCB}}_i}$ by using the VQR test of Gaglianone2011 and denote its $p$-value by $p^0_{{\mathsf{MCB}_i}}$.
Using (ref), an asymptotically valid $p$-value for the combined null hypothesis $\mathcal H_{\mathsf{MCB}}$ is
Similarly, we obtain an asymptotically valid $p$-value, $p_{\mathsf{DSC}}$, for testing equal discrimination,
by adapting (ref) and (ref). The Gaussian $p$-value follows from (ref) by setting $\boldsymbol{\omega} = (0,1,0,-1)^\top$, while for smooth scoring functions, the corresponding generalized $\chi^2$-probability is computed analogously using the method of I_1961. For non-smooth scoring functions, we test $\overline{\mathsf{DSC}}_i = 0$ through the MZ-type regression of the VQR test, but by simply testing the hypothesis of a zero slope.
A classical DM test for equal expected scores, $\mathcal H_\mathsf{S} = \left\{ \overline{\mathsf{S}}_{1} = \overline{\mathsf{S}}_{2} \right\}$, arises as a special case of Theorem (ref) for $\boldsymbol{\omega} = (1,-1,-1,1)^\top$. Confidence intervals for the individual decomposition terms as well as their differences can be obtained by inverting the above tests.
We begin with mean forecasts, where $\Gamma$ corresponds to the expectation. In this case, all strictly consistent scoring functions belong to the Bregman class given in (ref). For this class, we verify the high-level Assumptions (ref) and (ref)---and hence the validity of Theorems (ref) and (ref)---in a general time-series setting.
Assumption (ref) resembles standard time series conditions for mean regression. Item (ref) is a smoothness and convexity condition on the scoring function. For the SE scoring function, we often use $\mathsf{A}=\mathsf{O} =\mathbb{R}$, whereas the QLIKE score requires positive values such as $\mathsf{A}=\mathsf{O} = (0, \infty)$. Item (ref) is a classical mixing condition that ensures that suitable LLNs and CLTs apply and items (ref) and (ref) are standard in M-estimation NM_1994. Item (ref) collects moment conditions, which simplify considerably for the common choices $\phi_{\textsf{SE}}(z) = z^2$ and $\phi_{\textsf{QL}}(z) = -\log(z)$ of the SE and QLIKE scores in (ref) and (ref).
Positive definiteness of $\boldsymbol{\Omega}_T$ is stated in the separate Assumption (ref) as it is required for the asymptotic normality result in Theorem (ref), but is violated in the cases $\overline{\mathsf{MCB}}_i = 0$ and $\overline{\mathsf{DSC}}_i =0$, and is as such not required for the verification of the conditions of Theorem (ref). The positive definiteness of $\boldsymbol{\Omega}_T$ essentially imposes non-collinearity conditions for the (sum over time of the) original, recalibrated, and reference scores. Violations arise, for instance, if the recalibrated score is a constant multiple of the original score for all $t \in \mathbb{N}$, in which case $\overline{\mathsf{MCB}}_i=0$ follows.
We continue with the case where $\Gamma$ is the $\alpha$-quantile for some $\alpha \in (0,1)$, where all strictly consistent loss functions are given by the GPL class in (ref). For a linear GPL score decomposition, we employ (generalized) quantile regression by combining the estimator (ref) with the loss in (ref) and impose the following conditions.
The conditions of Assumption (ref) are standard regularity conditions for quantile regression, but also cover the case of estimation through the more general GPL loss functions. The required moments of order $s > k$ are required for the derivation of stochastic equicontinuity for strong mixing processes through hansen1996stochastic. As we consider low dimensionalities of $\boldsymbol{\Theta} \subset \mathbb{R}^k$ such as $k=2$ in our applications, this is not overly restrictive.
Compared to Assumption (ref), the conditions in Assumption (ref) require the conditional distribution to be absolutely continuous, as specified in item (ref); this is a standard requirement for quantile regression frameworks. Furthermore, item (ref) is included here because the subsequent proposition specifically verifies the conditions for the asymptotic normality result established in Theorem (ref), but does not treat the case of Theorem (ref).
Proposition (ref) verifies the high-level conditions of Theorem (ref) through standard conditions from (time series) quantile regression. Notice that Theorem (ref) only applies to smooth loss functions, and hence, the cases $\overline{\mathsf{MCB}}_{i} = 0$ and $\overline{\mathsf{DSC}}_{i} = 0$ require a different treatment for the case of quantile forecasts. For the case $\overline{\mathsf{MCB}}_{i} = 0$, we employ the quantile-specific MZ-test of Gaglianone2011, which we adapt for the case of $\overline{\mathsf{DSC}}_{i} = 0$ by simply testing for a zero slope in the MZ-regression. The resulting $p$-values are accordingly used in the combination method in (ref).
VaR backtesting is a fundamental pillar of financial risk management, used to assess whether a model accurately predicts the frequency and magnitude of potential losses. Traditional backtests such as by K_1995 and the traffic light system of B_1996, B_2019 typically focus on unconditional counts for the “violations” or “hits”, where realized losses exceed the VaR predictions as captured through the quantile identification functions, $\mathsf{V}_{it} = \mathsf{V}(X_{it}, Y_t) = \mathds{1}\{Y_t \le X_{it}\} - \alpha$, from (ref). These unconditional tests, however, remain insensitive to the clustering of violations---a phenomenon particularly critical during financial crises, when consecutive days of unpredicted losses can lead to rapid capital depletion.
Following the seminal work of C_1998, a plethora of conditional backtests has been proposed over the past decades. See e.g., campbell2007review and nieto2016frontiers for comprehensive reviews on VaR backtesting. Most of these methods test hypotheses closely related to the conditional calibration property in (ref) by utilizing varying information sets $\mathcal{G}_{it} = \sigma(\bm{G}_{it})$. Table (ref) provides a non-exhaustive list of prominent VaR backtests, illustrating the intrinsic connection between their choice of covariates (or instruments) $\bm{G}_{it}$ and our recalibration variables $\bm W_{it}$. Notably, the equivalence between (ref) and (ref) permits specifications based on either quantile regressions for $Y_t$, or mean regressions for the respective identification function.
As a canonical case of assessing auto-calibration, the quantile-specific adaptation of the MZ backtest by Gaglianone2011 utilizes the same recalibration technique as our linear score decomposition when $\bm W_{it}^\top = (1,X_{it})$. By matching the recalibration variables $\bm W_{it}$ with the covariates $\bm{G}_{it}$, both the score decompositions and the underlying inference can be tailored to the specific notion of conditional calibration underlying the other backtests presented in Table (ref).
We now assess the finite-sample performance of the proposed tests from Section (ref). Section (ref) describes the simulation setup and Section (ref) reports the results.
To empirically validate our tests for equal miscalibration, with null hypothesis $\mathcal H_{\mathsf{MCB}}: \overline{\mathsf{MCB}}_{1} = \overline{\mathsf{MCB}}_{2}$, or for equal discrimination, with $\mathcal H_{\mathsf{DSC}}: \overline{\mathsf{DSC}}_{1} = \overline{\mathsf{DSC}}_{2}$, between competing mean or quantile forecasts, we use the following data-generating process (DGP) for the realizations,
with $\varepsilon_{Y,t} \stackrel{\text{iid}}{\sim} \mathcal{N}(0,1)$. The predictor variables in $\mathbf{V}_t$ follow independent autoregressive AR(1) processes $K_t = \beta K_{t-1} + \varepsilon_{K,t},\ L_t = \beta L_{t-1} + \varepsilon_{L,t},\ M_t = \beta M_{t-1} + \varepsilon_{M,t},\ N_t = \beta N_{t-1} + \varepsilon_{N,t}$, with fixed $\beta = 0.25$ and $ \varepsilon_{K,t}, \varepsilon_{L,t}, \varepsilon_{M,t}, \varepsilon_{N,t} \overset{\text{iid}}{\sim} \mathcal{N}(0,1)$, which are all mutually independent.
Forecaster $X_{1t}$ has exclusive access to $K_t$ and forecaster $X_{2t}$ to $L_t$, while both forecasters share information on $M_t$ and $N_t$ as formalized by $\mathcal{I}_{1t} = \sigma \left\{K_{t}, M_{t}, N_{t} \right\}$ and $\mathcal{I}_{2t} = \sigma \left\{ L_{t}, M_{t}, N_{t} \right\}$. Notice that our setup requires both “joint” predictors $M_{t}$ and $N_{t}$ in order to construct non-collinear forecasts with equal $\overline{\mathsf{MCB}}$ and/or $\overline{\mathsf{DSC}}$ values.
Based on the distinct information sets $\mathcal{I}_{1t}$ and $\mathcal{I}_{2t}$, the forecasters issue mean forecasts
and $\delta_0, \xi_0 \in \mathbb{R}$. The imposed zeros in $\bm{\delta}$ and $\bm{\xi}$ reflect the absence of the corresponding predictor variables in the information sets of the respective forecaster. The coefficients $(\delta_0, \bm{\delta}^\top)$ and $(\xi_0, \bm{\xi}^\top)$ and their relation to $\bm{\gamma}$ allows to generate forecasts with different calibration and discrimination properties, as formalized in the following proposition for the squared error score from (ref).
The expressions in Proposition (ref) illustrate that for non-constant forecasts, $\overline{\mathsf{MCB}}_1=0$ holds if and only if $\bm{\delta}^\top \bm{\delta} = \bm{\delta}^\top \bm{\gamma}$ and $\delta_0=0$. Hence, perfect calibration, $\overline{\mathsf{MCB}}_1=0$, can be achieved even if the forecaster omits relevant variables. In fact, $\overline{\mathsf{MCB}}_1=0$ implies that if a predictor (say $M_t$) is used, it must be used with exactly the right parameter value ($\delta_M = \gamma_M$). This also implies that the use (e.g., $\delta_M \not= 0$) of irrelevant predictor variables (e.g., $\gamma_M = 0$) causes miscalibrated forecasts.
The $\overline{\mathsf{DSC}}_1$ component is unaffected by the intercept $\delta_0$ that does not carry any information. $\overline{\mathsf{DSC}}_1$ is positive whenever $\bm{\delta}^\top \bm{\gamma} \neq 0$, that is, whenever $X_{1t}$ reacts to the signal in $Y_t$, even incorrectly such as in the wrong direction. Moreover, by Cauchy-Schwarz, $\frac{(\bm{\delta}^\top \bm{\gamma})^2}{\bm{\delta}^\top \bm{\delta}} \leq \bm{\gamma}^\top\bm{\gamma}$, such that the maximal value of $\overline{\mathsf{DSC}}_1$ is $\varsigma\,\bm{\gamma}^\top \bm{\gamma}$, which is attained if and only if the parameters $\bm\delta$ are chosen proportionally to $\bm\gamma$, i.e., $\bm{\delta} = c \bm{\gamma}$ for some $c \in \mathbb{R}$ as then, $\frac{(\bm{\delta}^\top \bm{\gamma})^2}{\bm{\delta}^\top \bm{\delta}} = \frac{c^2 (\bm{\gamma}^\top \bm{\gamma})^2}{c^2 \bm{\gamma}^\top \bm{\gamma}} = \bm{\gamma}^\top \bm{\gamma}$. Hence, maximal discrimination can arise even for miscalibrated forecasts. If only some relevant variables are observed, the best achievable discrimination is obtained when the available predictors are used correctly. Analogous conclusions hold for the second forecaster $X_{2t}$ by replacing $\bm{\delta}$ with $\bm{\xi}$.
The closed-form expressions in Proposition (ref) enable diverse specifications of $\overline{\mathsf{MCB}}_i$ and $\overline{\mathsf{DSC}}_i$, allowing us to study the size and power of our proposed tests across different settings by appropriately choosing the parameters $\bm{\gamma}$, $(\delta_0, \bm{\delta}^\top)$, and $(\xi_0, \bm{\xi}^\top)$. In detail, Table (ref) specifies twelve different parameterizations that result in illustrative settings for $\overline{\mathsf{MCB}}_i$ and $\overline{\mathsf{DSC}}_i$ as given in the facet labels of Figure (ref). To analyze test power, we continuously misspecify the respective null hypotheses through a parameter $k>0$ that affects $\bm{\gamma}$, $(\delta_0, \bm{\delta}^\top)$, and $(\xi_0, \bm{\xi}^\top)$. The specifications of Table (ref) are chosen to guarantee that either $\overline{\mathsf{MCB}}_2(k)$ or $\overline{\mathsf{DSC}}_2(k)$ depend on $k$ and equal $\frac{8k^2}{15}$ for all $k \ge 0$. Figure (ref) in Appendix (ref) plots the population values $\overline{\mathsf{MCB}}_i$ and $\overline{\mathsf{DSC}}_i$, $i=1,2$ as functions of $k$ for all twelve setups.
For example, we generate the scenario $\overline{\mathsf{MCB}}_{1} = 0$ and $\overline{\mathsf{MCB}}_{2}(k) = \tfrac{8k^2}{15}$ while maintaining $0 < \overline{\mathsf{DSC}}_1 < \overline{\mathsf{DSC}}_2$ independent of $k$ by setting $\bm\gamma^\top = (0, \tfrac{1}{4}, \tfrac{1}{4}, 0)$ as well as $(\delta_0, \bm\delta^\top) = (0, 0, 0, \tfrac{1}{4}, 0)$ and $(\xi_0, \bm\xi^\top) = (0,0, \tfrac{1}{4} + \tfrac{k}{2}, \tfrac{1}{4} + \tfrac{k}{2}, 0)$. Here, the realizations are only driven by $L_t$ and $M_t$ with equal magnitudes, $\gamma_L = \gamma_M$. The first forecaster $X_{1t}$ does not have access to $L_t$, but she correctly uses $M_t$ through $\delta_M = \gamma_M$ such that her forecasts are perfectly calibrated, $\overline{\mathsf{MCB}}_1 = 0$. The sole access to $M_t$ however limits her discrimination to $\overline{\mathsf{DSC}}_1 = \tfrac{1}{15}$. By contrast, the second forecaster $X_{2t}$ uses both relevant variables $L_t$ and $M_t$, and hence achieves a higher discrimination $\overline{\mathsf{DSC}}_2 = \tfrac{2}{15} > \overline{\mathsf{DSC}}_1$ for any $k \ge 0$. However, for values $k > 0$, she uses the information incorrectly with $\bm{\xi} \not= \bm{\gamma}$ such that $\overline{\mathsf{MCB}}_2(k) = \tfrac{8k^2}{15}$ increases with $k$.
For the evaluation of our tests for quantile forecasts, we use the same DGP in (ref). As the information in $\mathbf{V}_t$ only affects the conditional mean of $Y_t$ in (ref), we generate the quantile forecasts at level $\alpha \in (0,1)$ as
using the inverse of the standard normal CDF, denoted by $\Phi^{-1}(\cdot)$.
We deliberately use the pure location process (ref) for quantile forecasts, allowing us to explicitly control the population miscalibration and discrimination values for the check loss in Proposition (ref), albeit these closed-form expressions become much more complicated as for the squared error. Due to this additional complication, we only consider four informative parameterizations as given in Table (ref). For these, Appendix (ref) provides the closed-form expressions for $\overline{\mathsf{MCB}}_i$ and $\overline{\mathsf{DSC}}_i$ as explicit functions of $k, \alpha$ and the parameters of the DGP. Notably, for the last parametrization, (approximate) equality of the miscalibration values is achieved by a numerical choice for the values of $\xi_0 = \xi_0(k,\alpha)$; see Figure (ref) and the text of Appendix (ref).
All results are based on 5000 Monte Carlo replications, a nominal significance level of $10\%$ and $T=500$. We estimate all long-run covariance matrices by the HAC estimator of A_1991, adapted to quantile regressions by GY_2024. For the implementation, we follow the default vcovHAC procedure from the R package sandwich of Z_2004.
For mean forecasts and the associated squared error score, Figure (ref) shows the rejection rates of the test for equal miscalibration $\mathcal H_\mathsf{MCB}: \overline{\mathsf{MCB}}_{1} = \overline{\mathsf{MCB}}_{2}$ in the upper panel in blue color, and of the test for equal discrimination $\mathcal H_\mathsf{DSC}: \overline{\mathsf{DSC}}_{1} = \overline{\mathsf{DSC}}_{2}$ in orange color in the lower panel. For comparison, we also show rejection rates of the DM test of equal predictive power $\mathcal H_{\mathsf{S}}: \overline{\mathsf{S}}_{1} = \overline{\mathsf{S}}_{2}$ with gray and dotted lines. We only show these rejection rates when the underlying population score difference coincides with the respective difference in $\overline{\mathsf{MCB}}$ or $\overline{\mathsf{DSC}}$ values. E.g., in the lower row of panel a), we ensure $\overline{\mathsf{DSC}}_{1} = \overline{\mathsf{DSC}}_{2}$ such that $\overline{\mathsf{S}}_{1} - \overline{\mathsf{S}}_{2} = \overline{\mathsf{MCB}}_1 - \overline{\mathsf{MCB}}_2$ and the tests of the hypotheses $\mathcal H_{\mathsf{S}}$ and $\mathcal H_\mathsf{MCB}$ are comparable.
Throughout the four plots in the left column, as well as for $k=0$ in all twelve plots, the data is generated under the respective null hypotheses of equal miscalibration in panel a), and equal discrimination in panel b). Here, we find correct or conservative rejection rates below the nominal level of $10\%$ for our tests, which is to be expected given the Bonferroni-type corrections from Section (ref). In the left panel, our tests approach the correct size for larger values of $k$ as e.g., for the $\mathsf{MCB}$-test, miscalibration of both forecasts increases with $k$ such that $2 \min \{ p^0_{\mathsf{MCB}_1}, p^0_{\mathsf{MCB}_2} \}$ becomes much smaller than $p^{+}_{\mathsf{MCB}}$, which operates under its null hypothesis and hence generates asymptotically exact size control.
The middle and right plot columns show that for increasing $k$, power increases naturally in all subplots. This illustrates the robustness of our testing procedure to zero or positive values of the respective population $\overline{\mathsf{MCB}}_i$ and $\overline{\mathsf{DSC}}_i$ values. Importantly, in the cases where score differences coincide with $\overline{\mathsf{MCB}}$ ($\overline{\mathsf{DSC}}$) differences in the upper (lower) panel, the component tests' power increases substantially with respect to the DM test. This can be explained as in these cases, the DM test's (estimated) variance is affected by fluctuations of both, (estimated) $\mathsf{MCB}$ and $\mathsf{DSC}$ values whereas the $\mathsf{MCB}$ test is unaffected by noise in the $\mathsf{DSC}$ component, and vice versa. Hence, we can expect power gains by testing on the respective components, as will become evident in our applications in Section (ref).
Figure (ref) displays rejection rates for quantile forecasts at levels $\alpha \in \{0.01, 0.05, 0.1, 0.25, 0.5\}$ and within the four most informative parameterizations that allow for comparison with the DM test. These four scenarios correspond to the lower left and middle plots in panels a) and b) from Figure (ref). The results for median forecasts with $\alpha = 0.5$ are comparable to the mean forecasting results from Figure (ref): The tests are conservative under the null, develop power with $k > 0$ and despite their conservative size\footnote{In unreported simulations, we find that the elevated type I error of the DSC test in the third column for $\alpha = 0.01$ does not increase with $k$ and declines as the sample size grows. This pattern suggests that the distortion is attributable to finite-sample effects.}, our tests dominate the DM-test in terms of power. For more extreme quantile levels, power naturally declines. Importantly, and especially for the DSC test, the dominance of the DM test is even more pronounced for extreme quantile levels.
Here, we apply the proposed tests for equal miscalibration and discrimination in two central forecasting domains: inflation rates in Section (ref) and financial risks in Section (ref). Replication code and details on the data are available at \OldHref{https://github.com/marius-cp/SDI_replications}{https://github.com/marius-cp/SDI_replications}.
Inflation is renowned for being difficult to forecast, and survey forecasts often outperform sophisticated econometric models FW_2013. We compare two popular one-year-ahead survey forecasts for quarterly U.S.\ CPI inflation: the Survey of Professional Forecasters (SPF) of the Federal Reserve Bank of Philadelphia, and the Michigan Survey of Consumers.\footnote{Data sources: SPF forecasts (Individual_CPI.xlsx) from \OldHref{https://www.philadelphiafed.org/surveys-and-data/cpi-spf}{Philadelphia Fed}; Michigan forecasts (Table 32) from their \OldHref{https://data.sca.isr.umich.edu/tables.php}{data site}; and CPI data (cpiQvMd.xlsx) from \OldHref{https://www.philadelphiafed.org/surveys-and-data/real-time-data-research/cpi}{Philadelphia Fed}.} The SPF is considered an expert forecast, as its panelists are professionals with extensive macroeconomic forecasting experience. In contrast, the Michigan survey randomly samples U.S.\ cell phone numbers, without targeting professionals. One might therefore expect the experts to outperform consumers, yet overall forecast comparisons are often inconclusive as e.g., in the empirical applications of EGJK_2016 and P_2020.
The aim of this application is to demonstrate that, even when inference on the overall scores is inconclusive, the improved test power of the individual $\mathsf{MCB}$ and $\mathsf{DSC}$ decomposition terms illustrated in Section (ref) can produce significant results with clear and interpretable insights while being applicable in general multi-step ahead forecast environments.
We replicate the analyses of EGJK_2016 and P_2020, comparing the predictive performance of the SPF and Michigan survey forecasts, and extend their approach by applying our novel tests for equal forecast calibration and discrimination. Our evaluation sample covers $T = 154$ quarters, from 1982:Q3 to 2020:Q2, thereby excluding the high-inflation period following the COVID-19 pandemic. Figure (ref) displays the forecasts alongside the corresponding realizations.
Michigan survey respondents report their expectations of the percentage change in prices over the next 12 months. To ensure comparability, we follow EGJK_2016 and P_2020 and construct a corresponding SPF measure by averaging the one- to four-quarter-ahead consensus forecasts (where the consensus is defined as the median across respondents). Averaging is appropriate because SPF forecasts are stated in terms of annualized quarter-over-quarter inflation rates. As the evaluation target $Y_t$, we use the four-quarter logarithmic growth rate of the consumer price index (CPI), based on the 2023:Q3 vintage, over the year following the forecast date.
The left panel of Figure (ref) reports the average squared error score together with its decomposition into miscalibration and discrimination components for the two competing forecasts. Calibration appears broadly similar across the two surveys. In contrast, the SPF forecasts display a markedly higher degree of discrimination---a difference that is not visible in the purely scoring-function-based analyses of EGJK_2016 and P_2020. The right panel corroborates their main finding: the overall difference in predictive performance is statistically insignificant, with a $p$-value of 0.284.
Consistent with the visual impression from the left panel, the difference in miscalibration is not statistically significant. However, the SPF forecasts exhibit significantly stronger discrimination, with a $p$-value of 0.037. For the Michigan survey forecasts, we cannot even reject the null hypothesis of no discrimination. These findings suggest that professional forecasters’ predictions contain more information about future inflation outcomes, which is plausible given their field-specific expertise and experience.
Figure (ref) in Appendix (ref) reports diagnostic checks for the linear recalibration regression and reveals no evidence of misspecification, supporting the use of linear recalibration for inference on score decompositions.
The macroeconomic literature following CG_2015, as discussed in the introduction, examines the sensitivity of forecast errors to additional covariates. These covariates can be integrated into our framework through the recalibration vector $\bm W_{it}$. Consequently, their tests for rational expectations focus on a specific dimension of forecast calibration. However, such evaluations fail to account for the underlying forecast discrimination---a component that our application to survey forecasts reveals to be fundamentally distinct from calibration performance.
From a methodological perspective, this application underscores the utility of our proposed inference framework. Beyond partitioning forecast performance into interpretable economic components, our tests can exhibit superior statistical power compared to the DM test in these settings. Furthermore, the framework is applicable in general time-series environments, including multi-step-ahead forecasting horizons.
Here, we extend the motivational example from Section (ref) and consider one-day-ahead forecasts of the variance and the $\alpha$-quantiles ($\alpha$-VaR), with $\alpha \in \{1\%, 5\%\}$, of financial log returns $R_t$. This application demonstrates the value of inference based on score decompositions, as it helps reconcile the partly contradictory findings that often arise between traditional backtests reviewed in Section (ref) and scoring-function-based forecast evaluations in applied work such as in Section (ref).
Given the standard assumption $\mathbb{E}[R_t] = 0$ such that $\operatorname{Var}(R_t) = \mathbb{E}[R_t^2]$, the variance forecasts can be treated as forecasts for the expectation of $R_t^2$, or as introduced below, the expectation of an unbiased RV estimator. Hence, we follow P_2011 and evaluate the forecasts with the squared error and QLIKE loss functions that are special cases of the Bregman class (ref), considered in Proposition (ref). For the VaR forecasts, we use the check loss from (ref), as covered by Proposition (ref).
For $R_t$, we use logarithmic returns of the S&P 500 E-mini futures, a highly liquid instrument central to practical portfolio management, which is widely used in the financial variance forecasting literature C_2009, LPS_2015, BHHP_2018. E-mini contracts trade almost continuously from Sunday evening to Friday afternoon, with brief daily interruptions from 3:15–3:30 p.m.\ and 4:00–5:00 p.m.\ CT. A trading day is defined as the interval spanning from 5:00 p.m.\ CT of the preceding calendar day to 4:59 p.m.\ CT. We use intraday price observations from the data provider Refinitiv (ticker: ESc1), spanning the period from January 1, 2000, to January 1, 2022. From this, we compute the daily open-to-close return $R_t$ and the 5-minute realized variance, $\operatorname{RV}_t := \sum_{m=1}^M r_{tm}^2$, where \(r_{tm}\) is the \(m\)-th intraday return on day \(t\) and \(M\) denotes the number of equally spaced observations per trading day.
We employ several standard models to generate one-step-ahead variance forecasts. Specifically, we consider the Gaussian GARCH model of B_1986 and its asymmetric extension, the GJR-GARCH model of GJR_1993, estimated with Student’s $t$ residuals, given by
We further use two models that use high-frequency information through the RV estimator. The third model is the Heterogeneous Autoregressive (HAR) model proposed by C_2009,
where $\overline{\operatorname{RV}}_{w,t}^{1/2} := \frac{1}{5}\sum_{j=0}^4 {\operatorname{RV}_{t-j}^{1/2}}$ and $\overline{\operatorname{RV}}_{m,t}^{1/2} := \frac{1}{22}\sum_{j=0}^{21} {\operatorname{RV}}_{t-j}^{1/2}$ are weekly and monthly averages, respectively. Fourth, we employ the MIDAS model of GSV_2006, GKZ_2016, where ${\operatorname{RV}_t^{1/2}}$ is predicted based on a weighted average of its lagged values,
where the MIDAS weights $w_l$ are parameterized through the exponential Almon polynomial.
We construct VaR forecasts for these four models by using the $\alpha$-quantile of the imposed residual distribution multiplied by the square root of the respective variance forecast. For quantile forecasts, we additionally use the classical Historical Simulation (HS) method, that uses the empirical $\alpha$-quantile of the preceding 250 trading days as a quantile forecast.
For the first four models, we estimate the parameters by maximum likelihood using the imposed residual distributions, based on a fixed estimation window from January 1, 2000, to December 31, 2007, corresponding to approximately the first third of the sample. The subsequent period from January 1, 2008 to January 1, 2022 serves as evaluation period, yielding $T=3617$ one-step ahead forecasts and corresponding realizations. Figure (ref) in Appendix (ref) shows the variance and VaR forecasts together with their evaluation targets.
Figure (ref) presents the evaluation of variance forecasts under the QLIKE and SE scoring functions. The left panels, following DGJV_2024, plot the estimated miscalibration and discrimination of the four forecasting methods along the axes. Gray iso-lines represent average score levels, with observations closer to the upper-left corner indicating superior performance. We find that the RV-based forecasting methods exhibit markedly stronger discrimination, which translates into better overall predictive accuracy as reflected in the average scores.
The right panel of Figure (ref) reports the corresponding test results: Gray-shaded boxes display $p$-values for tests of zero miscalibration and zero discrimination. While all models show significantly positive discrimination, the null of correct calibration is rejected more frequently for the RV-based models than for the two GARCH-type specifications.
The remaining off-diagonal panels present pairwise differences in average scores ($\Delta \widehat{\mathsf{S}}$), miscalibration ($\Delta \widehat{\mathsf{MCB}}$), and discrimination ($\Delta \widehat{\mathsf{DSC}}$), with significance indicated by 0 to 3 stars corresponding to the 10%, 5%, and 1% levels. Focusing on the lower column, which compares the GARCH model to the three alternatives, most differences are highly significant. In particular, the inferior performance of the GARCH model is primarily driven by its significantly weaker discrimination, consistent with the informational advantage of RV-based models that exploit high-frequency data.
Turning to the VaR evaluation results, Table (ref) reports the $p$-values of seven standard VaR backtests---described in Section (ref) and Table (ref)---along with the unconditional hit frequency in the final column. Strikingly, particularly at the 1% level, the simple HS model appears to be the best calibrated, whereas all competing models are consistently rejected across the backtests. Although evidence at the 5% level and for conditional calibration backtests is weak or inconclusive, the 1% level remains the most relevant for risk management B_2019.
Score decompositions with their associated inference help explain the seemingly superior performance of the HS model. Figure (ref) presents these decompositions, analogous to Figure (ref), using the check loss function for 1% and 5% VaR forecasts.
At the 1% level, the $\mathsf{MCB}$–$\mathsf{DSC}$ plot in Figure (ref) shows that, although the HS model exhibits the best calibration, it has by far the weakest discrimination. This results in clearly inferior predictive performance according to the check loss, as further confirmed by the strongly significant results in the right panel. These findings resolve the seemingly contradictory results highlighted in the motivational example from Section (ref) and underscore the risk of evaluating forecasts solely on calibration. As such, they point to a potentially serious shortcoming in current banking regulation.
At the 5% level, the HS model shows markedly poorer conditional calibration, as reflected both in the backtest results in Table (ref) and in the score decomposition in Figure (ref). The relative ordering of the other four models, however, is similar to that observed at the 1% level, with generally stronger significance, which can be attributed to lower noise at less extreme probability levels.
A general pattern evident in the $\mathsf{MCB}$–$\mathsf{DSC}$ plot at the 1% level---and, to a lesser extent, at the 5% level and for the variance forecasts in Figure (ref)---is that models incorporating more information can convert it into higher discrimination. However, this often comes at the cost of increased miscalibration. Notably, achieving auto-calibration conditional on $\bm W_{it} = (1, X_{it})$ is relatively easy when forecasts $X_{it}$ are nearly constant, but much more difficult for highly discriminating forecasts that vary substantially over time. This observation calls into question the common practice of evaluating risk measure forecasts primarily through MZ-type backtests Gaglianone2011, Guler2017, BD_2020, which focus on auto-calibration.
Figures (ref)--(ref) in Appendix (ref) present diagnostic checks for the linear recalibration regressions underlying the score decompositions of the variance and VaR forecasts. The pointwise confidence bands largely cover the null specification, providing no evidence against correct model specification. This supports the use of linear recalibration for score decompositions in these applications, combining robustness and adequate fit with the ability to conduct inference, which is essential in the settings considered.
This paper has established a comprehensive inferential framework for score decompositions, enabling a more granular assessment of predictive performance through the lenses of miscalibration, discrimination, and uncertainty. By utilizing a linear recalibration technique, we have shown that it is possible to derive robust inference for a wide range of point forecasts---including those governed by non-smooth loss functions---while maintaining attractive finite sample non-negativity conditions and a formal bridge to the classical MZ_1969 regression. Our results demonstrate that this decomposition approach not only enriches the informational content of traditional predictive ability tests but also offers tangible gains in statistical power. Through our empirical applications, we have highlighted how these methods can reveal significant latent differences in forecast discrimination and, crucially, identify systemic deficiencies in financial risk backtesting that current banking regulations fail to capture.
The assumption of linear recalibration is guaranteed to hold under the null hypotheses of zero discrimination or perfect calibration, though it may be violated in more general settings. Nevertheless, this linear specification provides essential stability and allows for the inclusion of further recalibration variables, allowing for direct generalizations to e.g., cross-calibration SZ_2017, and enabling direct links to VaR backtests and the literature on macroeconomic survey forecasts following CG_2015. The linearity assumption further guards against the overfitting that often plagues nonparametric techniques in this context. Such overfitting typically manifests as a positive bias in the non-negative $\mathsf{MCB}$ and $\mathsf{DSC}$ components---a risk that is particularly acute in the small-to-medium sample sizes characteristic of our applications. Notably, we find no empirical evidence against the linearity assumption in our data, which supports the robustness of our approach.
Nonetheless, particularly in data-rich environments, semiparametric extensions that estimate the recalibration curve nonparametrically---using kernel, spline, or isotonic regression---represent promising avenues for future research. Such developments could build on the seminal work of newey1994asymptotic regarding two-step semiparametric inference and more recent advances in isotonic plug-in estimation, such as Xu_IsoPlugIn_2022.
We treat the possibly model-based forecasts as the finalized output of a specific sequence of modeling choices, including parameter estimation. Consequently, we formulate our hypotheses based on the realized quality of these forecasts rather than on the properties of the underlying models themselves. In contrast, the framework of West1996asymptotic assesses whether two forecasting models based on their pseudo-true parameters would exhibit equal predictive accuracy, hence requiring that the test statistic's standardization explicitly accounts for the asymptotic noise introduced by parameter estimation. This extension can be integrated into our framework by redefining the population $\overline{\mathsf{MCB}}$ and $\overline{\mathsf{DSC}}$ components in a model-based environment and replacing the central limit theorem in Assumption (ref) (ref) with the asymptotic approximations of West1996asymptotic, or with the numerous refinements and generalizations proposed in subsequent work.
Generalizations of our framework to other elicitable functionals, such as expectiles, are straightforward. However, functionals that are merely jointly or multi-objective elicitable following FZ_2016 and FisslerHoga2024---such as the variance (when the mean is unknown), the Expected Shortfall, or the systemic risk measures Co-Value-at-Risk and Marginal Expected Shortfall---require additional care. In these cases, the underlying recalibration regressions necessitate the estimation of conditional models for the associated nuisance components, such as the mean or the VaR patton2019dynamic, BD_2020, dimitriadis2025regressions. The propagation of estimation risk from these nuisance quantities into the final score decomposition represents a non-trivial extension of the current inferential framework.
In conclusion, the contributions of this paper extend beyond their immediate applications in econometrics and finance. The principles of score decomposition and linear recalibration constitute a versatile toolkit for any domain where predictive accuracy is paramount, from meteorology to machine learning. The framework proposed herein represents a fundamental shift in perspective: from merely asking which forecast is better to diagnosing why it is better. As the complexity of our models and the richness of our data continue to grow, this diagnostic approach will be essential for robust and reliable decision-making.
Timo Dimitriadis acknowledges funding by the German Research Foundation (DFG) through the projects 502572912 and 568876076. We gratefully acknowledge the Hohenheim Datalab (DALAHO) for providing access to Refinitiv TickHistory. We thank all seminar and conference participants who provided feedback on earlier versions of this paper. In particular, we are grateful to Sam Allen, Tillman Gneiting, Alexander Jordan, Robert Jung, Fabian Krüger, Andrew Patton, Winfried Pohlmeier, Johannes Resin, Karsten Schweikert, and Johanna Ziegel for their valuable comments and suggestions.