EconBase
← Back to paper

Sequentially valid inference for probabilistic inflation forecasts

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

90,638 characters

Sequentially valid inference for probabilistic inflation forecasts



\def\spacingset#1{\renewcommand{\baselinestretch}
	{#1}\small\normalsize} \spacingset{1}
\spacingset{1.84}

\bibliographystyle{chicago}


\if00
{
	\title{\textbf{Sequentially valid inference for probabilistic inflation forecasts}}
\author{
\textsc{Amadeo Grob} \footnotemark[1]
\and
\textsc{Maurizio Daniele} \footnotemark[2]
\and
\textsc{Johanna Ziegel} \footnotemark[3]
}
\date{\today}
\footnotetext[1]{Corresponding Author. University of St. Gallen, Mathematics and Statistics Division, Rosenbergstrasse 22, 9000 St. Gallen, Switzerland. Email: [email removed]}
\footnotetext[2]{ETH Zürich, KOF Swiss Economic Institute, 8092 Zurich, Switzerland. Email: [email removed]}
\footnotetext[3]{ETH Zürich, Seminar for Statistics, Rämistrasse 101, 8092 Zürich, Switzerland. Email: [email removed]}
\maketitle
} \fi

\if10
{
\bigskip
\bigskip
\bigskip
$ $
\vspace{1cm}
$ $
\begin{center}
	{\LARGE\bf Sequentially valid inference for probabilistic inflation forecasts} \\
	\bigskip
	\today
\end{center}
\medskip
} \fi

\bigskip
\begin{abstract}\footnotesize
Traditional statistical tests are poorly suited for the sequential evaluation of probabilistic forecast calibration. We address this limitation in macroeconomic forecasting by applying a new sequential testing method based on e-values. The e-value-based methodology enables anytime-valid inference. It allows practitioners to test against calibration continuously without invalidating statistical guarantees. To illustrate the framework's practical value, we apply it to probabilistic inflation forecasts for the United States, the Euro Area, and Switzerland. Our analysis shows that the sequential approach gives detailed insights into the timing and nature of forecast misspecification. We find these diagnostics are particularly insightful during major structural breaks. During these events, we find evidence against calibration that static, full-sample tests often miss. Therefore, this work shows that e-value-based tests are a practical method for the evaluation of forecast calibration in empirical macroeconomics.\\

\noindent{\em Keywords: Sequential testing, e-values, probabilistic forecast calibration, macroeconomic forecasting}
\end{abstract}

\begin{comment}
\section*{Discussion}

\begin{enumerate}
    \item (JZ) We evaluate calibration sequentially, but forecast performance with respect to loss functions/proper scoring rules only at one time point (end of the sample period). Why is this sensible?\newline
    (AG) Focus on calibration from the start. Comparison as first summary statistic. Maybe consider average performance statistics for interesting time periods that we found out through calibration methodology (high level discussion!).
    \item (JZ) Which questions would we like to answer in the paper?\newline
    (AG) Goal to demonstrate methodology to a new audience. Method gives new insights during regime changes/shocks in real time.
    \item (JZ) What is our target outlet? IJF.
    \item (JZ) Naive last is always a bit worse than ARIMA(1,1,0), which is why we consider further and don't look much more at naive last. BVAR models are only considered further since they are prominent in applications/literature but we have to emphasize that their performance does is not competitive.
    \item (JZ) Possibly: Consider score decompositions for US/CH data. EU 292 data points, US 784, CH 508. Orthogonal perspective on BVAR models?
    \item Main focus US, one graphic for EUR. Rest into appendix.
\end{enumerate}

\begin{figure}[t]
  \centering
  \includegraphics[width=\textwidth]{Pictures/plots/US/us_sdi_mcb_dsc_by_horizon.pdf}
  \vspace{-2.25em}
  \caption{TBW}
  \label{fig:dscmcbus}
\end{figure}

\begin{table}[h]
  \centering
  \setlength{\tabcolsep}{3pt} \begin{tabular}{l c c c c c c c c c c c c}
\toprule
 & \multicolumn{3}{c}{\textbf{h=1}} & \multicolumn{3}{c}{\textbf{h=3}} & \multicolumn{3}{c}{\textbf{h=6}} & \multicolumn{3}{c}{\textbf{h=12}} \\
\cmidrule(lr){2-4} \cmidrule(lr){5-7} \cmidrule(lr){8-10} \cmidrule(lr){11-13}
\textbf{Model} & \textbf{MCB} & \textbf{DSC} & \makecell{\textbf{DSC} \\ \textbf{($\mathbf{p}$)}} & \textbf{MCB} & \textbf{DSC} & \makecell{\textbf{DSC} \\ \textbf{($\mathbf{p}$)}} & \textbf{MCB} & \textbf{DSC} & \makecell{\textbf{DSC} \\ \textbf{($\mathbf{p}$)}} & \textbf{MCB} & \textbf{DSC} & \makecell{\textbf{DSC} \\ \textbf{($\mathbf{p}$)}} \\
\midrule
\multicolumn{13}{l}{\textit{Baseline}} \\
ARIMA(1,1,0) & 0.003 & 2.431 & -- & 0.057 & 1.979 & -- & 0.256 & 1.287 & -- & 1.220 & 0.313 & -- \\
\midrule
\multicolumn{13}{l}{\textit{BVAR}} \\
BVAR (Minnesota) & 0.015 & 2.340 & 0.002 & 0.100 & 1.876 & 0.005 & 0.369 & 1.214 & 0.239 & 1.389 & 0.288 & 0.594 \\
\midrule
\multicolumn{13}{l}{\textit{Advanced}} \\
DFM & 0.005 & 2.433 & 0.247 & 0.073 & 1.984 & 0.336 & 0.296 & 1.307 & 0.033 & 1.366 & 0.318 & 0.429 \\
DRF & 0.007 & 2.229 & 0.025 & 0.019 & 1.758 & 0.009 & 0.052 & 1.268 & 0.833 & 0.021 & 0.826 & 0.013 \\
\bottomrule
\end{tabular}
  \caption{Miscalibration (MCB) and discrimination (DSC) decomposition of the mean squared error for the US forecasts, relative to the ARIMA baseline, by horizon. $p$-values test equal discrimination against ARIMA (SDI package, Dimitriadis \& Puke).}
  \label{tab:us_sdi_mcb_dsc}
\end{table}
\end{comment}



\section{Introduction}\label{sec:intro}
Central banks rely on accurate inflation forecasts to guide monetary policy and communication. While point forecasts remain the most visible form of forecast communication, probabilistic forecasts provide a richer representation of uncertainty by describing the full distribution of possible future outcomes. In practice, however, major central banks differ in how they communicate this uncertainty. The Bank of England, for example, pioneered density forecasts by publishing inflation fan charts since 1997 \citep{bernanke_forecasting_2024, monetary_policy_committee_boe_monetary_2025}. In contrast, other central banks such as the US Federal Reserve (Fed), the European Central Bank (ECB), and the Swiss National Bank (SNB) provide less explicit distributional information. This variation in communication practices also reflects a broader gap in the evaluation of probabilistic inflation forecasts. Existing research has focused on the calibration of the Bank of England's fan charts, with some studies finding evidence of systematic miscalibration \citep{clements_evaluating_2004, wallis_assessment_2004}, while others provide a more favorable assessment \citep{elder_assessing_2005}. In contrast, relatively little empirical evidence exists on the calibration of density forecasts for other major central banks, including the Fed, ECB, and SNB \citep{fomc_minutes_2007, reifschneider_gauging_2007, chahad_update_2024, european_central_bank_eurosystem_2025}. This paper contributes to closing this gap by developing, monitoring, and evaluating probabilistic inflation forecasts for the United States, the Euro Area, and Switzerland. In particular, we focus on forecast calibration, a fundamental property of probabilistic forecasts that determines whether reported uncertainty accurately reflects the likelihood of future outcomes.

A probabilistic forecast is calibrated if observed outcomes are consistent with the forecast distributions. In other words, the forecast probabilities should coincide with the empirical frequencies of the realized events. A widely used diagnostic tool to assess calibration is the probability integral transform (PIT) \citep{diebold_comparing_1995, gneiting_probabilistic_2007}. The PIT is defined as the value of the forecast cumulative distribution function evaluated at the observed outcome. Under probabilistic calibration, the resulting PIT values are expected to be distributed according to a standard uniform distribution. In practice, calibration is commonly assessed using graphical diagnostics such as PIT histograms. A histogram that is approximately uniform provides evidence of well-calibrated forecasts, whereas systematic deviations from uniformity indicate calibration deficiencies. The shape of the histogram can also provide insight into the nature of the miscalibration. For example, a skewed histogram suggests systematic biases, while an inverse U-shaped histogram is indicative of overdispersion in the predictive distributions, meaning that the predictive distributions are excessively diffuse and therefore overestimate the true uncertainty.

Despite their widespread use, traditional calibration tools have a fundamental limitation: they are designed for fixed-sample settings and therefore ignore the sequential nature of forecast evaluation. In practice, particularly in macroeconomic forecasting, the model performance is monitored continuously rather than assessed only after a predetermined evaluation period has ended \citep{arnold_sequentially_2023}. Classical goodness-of-fit tests for PIT uniformity, such as the Kolmogorov–Smirnov test, are valid only when the evaluation sample size is specified \textit{ex ante}. This assumption is rarely satisfied in practice. Instead, analysts often evaluate the forecast performance whenever concerns about the model adequacy arise. Such a data-dependent monitoring, known as optional stopping, invalidates the nominal p-values of conventional hypothesis tests \citep{shafer_game-theoretic_2019, arnold_sequentially_2023}. Consequently, repeated testing over time inflates the probability of false rejections and undermines the statistical validity of classical calibration tests. A further limitation is that forecast calibration is rarely constant over time. Structural changes in the data-generating process may cause a forecasting model to alternate between periods of good calibration and periods of systematic miscalibration. Moreover, biases in opposite directions can even cancel out over an extended sample. For example, a period of under-prediction followed by over-prediction could yield a PIT distribution that looks deceptively uniform \citep{arnold_sequentially_2023}. A static, full-sample test may fail to detect important temporal changes in the forecast performance. These issues highlight the need for a new testing framework. This framework should remain remain valid under sequential testing, adapt to structural changes, and allow for real-time calibration assessment without violating statistical error rates.

To address these limitations, we adopt the sequential inference framework for forecast calibration proposed by \citet{arnold_sequentially_2023}. Rather than relying on classical p-values, this framework quantifies statistical evidence using e-values. An e-value is a non-negative random variable with expectation at most one under the null hypothesis \citep{shafer_testing_2021, arnold_sequentially_2023}. Consequently, large e-values provide evidence against the null. A key advantage of e-values is that they can be combined multiplicatively over time to form a test supermartingale under the null hypothesis, which guarantees anytime-valid inference \citep{grunwald_safe_2024}. Therefore, practitioners can use the test supermartingale to assess calibration at any point in time. This allows us to monitor forecast calibration continuously with valid Type~I error control, even with optional stopping. When the forecasts are well-calibrated, the test supermartingale is unlikely to attain large values. In contrast, persistent miscalibration causes the test supermartingale to grow, thereby accumulating evidence against the null hypothesis. This approach is therefore particularly well-suited for the real-time, adaptive environments, such as those encountered by central banks and other institutions that routinely update and evaluate predictive models.

We construct a suite of probabilistic inflation forecasting models and evaluate the calibration of their predictive distributions using the sequential e-value framework. Our empirical analysis focuses on monthly headline inflation forecasts at horizons of 1, 3, 6, and 12 months for the United States, the Euro Area, and Switzerland. For each region, we generate predictive distributions using a diverse set of forecasting models, ranging from simple benchmark methods, such as naive forecasts and autoregressive (AR) models, to more sophisticated approaches, including Bayesian vector autoregressions (BVARs), a dynamic factor model (DFM), and the distributional random forest, a modern nonparametric forecasting method \citep{cevid_distributional_2022}. Forecast calibration is assessed using sequential e-value tests. These tests use PIT and rank values to check for continuous and discrete uniformity. To complement the calibration analysis, we also evaluate point and distributional forecast accuracy using conventional performance measures, including the root mean squared error (RMSE), mean absolute error (MAE), and the continuous ranked probability score (CRPS). Together, these metrics provide a comprehensive assessment of both forecast accuracy and probabilistic calibration. The sequential e-value framework enables continuous monitoring of the forecast performance over time. Our empirical results show that calibration deficiencies are concentrated in distinct episodes rather than persisting uniformly throughout the evaluation sample. Simple benchmark models, such as the Gaussian no-change forecast and univariate AR models, generally exhibit satisfactory calibration. In contrast, the BVAR models tend to be systematically overconfident, displaying pronounced calibration failures during periods of structural change, including the global financial crisis in the United States and the post-pandemic inflation surge in the Euro Area. Although the DFM performs more favorably than the BVAR specifications, the sequential diagnostics nevertheless reveal horizon-dependent calibration deficiencies that are not consistently detected by conventional full-sample tests. These findings illustrate the practical value of the e-value framework for real-time forecast evaluation. Beyond identifying whether a forecasting model is miscalibrated, the sequential approach pinpoints when calibration deteriorates and whether the deterioration is transient or persistent. This ability to detect and monitor time-varying changes in forecast calibration constitutes a key advantage over traditional fixed-sample evaluation methods.

The rest of the paper is organized as follows. Section~\ref{sec:evaluation_theory} introduces the theoretical framework for evaluating probabilistic forecasts, defines calibration in the sequential setting, and presents the e-value tests used in our analysis. Section~\ref{sec:forecast_design} describes the macroeconomic datasets, forecast design, and the considered forecasting models. Section~\ref{sec:results} reports our main empirical findings on forecast calibration across all regions, models, and horizons. Finally, Section~\ref{sec:conclusion} summarizes the results and discusses their implications.


\section{Evaluating probabilistic forecasts}
\label{sec:evaluation_theory}

First, we review classical methods for forecast evaluation, and then introduce e-values as a modern tool to test the calibration of probabilistic forecasts. Unlike classical tests, e-values remain valid if a forecaster monitors calibration continuously or stops testing at any point. The methods and definitions we present are based mainly on the work of \citet{arnold_sequentially_2023}.

\subsection{Notions of calibration}
\label{sec:notionscalibration}

We model the year-over-year inflation rate as a stochastic process $(Y_t)_{t \in \mathbb{N}}$ on a probability space $(\Omega, \mathcal{F}, \mathcal{P})$. The available information over time is represented by a filtration $(\mathcal{F}_t)_{t \in \mathbb{N}}$, a sequence of nested $\sigma$-algebras where $\mathcal{F}_t \subset \mathcal{F}_{t+1} \subset \mathcal{F}$. We assume the process $Y_t$ is adapted to the filtration, that is, $Y_t$ is $\mathcal{F}_t$-measurable for each $t$, or in other words, at time point $t$, we know the value of $Y_t$. A probabilistic forecast for $Y_{t+h}$ provides a full predictive probability distribution based on the available information $\mathcal{F}_t$ at the time of forecasting \citep{gneiting_probabilistic_2014}. This approach allows to explicitly quantify uncertainty of the future outcome.

Let $h \in \mathbb{N}$ be the forecast horizon. The true conditional cumulative distribution function (CDF) of a future value $Y_{t+h}$, given the information $\mathcal{F}_t$, is
\[
F_{t,h}(y) = \mathbb{P}(Y_{t+h} \le y \mid \mathcal{F}_t), \quad y \in \mathbb{R}.
\]
This distribution is the theoretical, unobservable ground truth. A probabilistic forecast, which we denote by the predictive CDF $\hat{F}_{t,h}$, is an estimate of this true conditional CDF.
This predictive CDF, $\hat{F}_{t,h}$, is our central object of interest. In our application, we denote the realization of the random variable $Y_{t+h}$ by the standard econometric convention, $\pi_{t+h}$. We can derive any desired statistical functional from the predictive CDF. For example, a point forecast for the median is the 50\%-quantile, and a 90\% central prediction interval is the range $[\hat{F}_{t,h}^{-1}(0.05), \hat{F}_{t,h}^{-1}(0.95)]$, where $\hat{F}^{-1}_{t,h}$ denotes the generalized inverse, or quantile function.


\begin{comment}
XXX: [This can possibly go in the appendix but probably it is not needed at all.]
Various measures exist for point forecast performance, but no single measure suits all purposes \citep{hyndman_another_2006, kolassa_why_2020}. A good point forecast is close to the actual observation. A smaller forecast error, therefore, means a more accurate forecast. From the predictive CDF, $\hat{F}_{t,h}$, we can extract two common point forecasts: the mean and the median.

The forecast mean, denoted $\bar{y}_{t+h|t}$, is the expected value of the predictive distribution:
\[
\bar{y}_{t+h|t} = \int_{-\infty}^{\infty} y \, d\hat{F}_{t,h}(y).
\]
The forecast median, denoted $\tilde{y}_{t+h|t}$, is the 50\%-quantile of the predictive distribution:
\[
\tilde{y}_{t+h|t} = \hat{F}_{t,h}^{-1}(0.5).
\]
These point forecasts should be evaluated with an appropriate scoring function. A scoring function $S$ is consistent for a statistical functional $T(\cdot)$ if, for any probability distribution $F$, the functional is the unique minimizer of the expected score \citep{gneiting_making_2011}:
\[
T(F) = \operatorname*{arg\,min}_{x \in \mathbb{R}} \mathbb{E}_{Y \sim F} [S(x,Y)].
\]
The squared error is consistent for the mean and is evaluated as:
\[
S(\bar{y}_{t+h|t}, \pi_{t+h}) = (\bar{y}_{t+h|t}-\pi_{t+h})^2.
\] The absolute error is consistent for the median and is evaluated as:
\[
S(\tilde{y}_{t+h|t}, \pi_{t+h}) = |\tilde{y}_{t+h|t}-\pi_{t+h}|.
\]
To evaluate performance over a sample, we use the root mean squared error (RMSE) and the mean absolute error (MAE) \citep{gneiting_making_2011}. These measures aggregate the performance by averaging the consistent scoring functions. For an evaluation sample of size $N$, the RMSE averages the squared error to evaluate the mean forecast:
\[
\text{RMSE}_h = \sqrt{\frac{1}{N}\sum_{i=1}^{N} (\pi_{t_i+h} - \bar{y}_{t_i+h|t_i})^2}.
\]
The MAE averages the absolute error to evaluate the median forecast:
\[
\text{MAE}_h = \frac{1}{N}\sum_{i=1}^{N} |\pi_{t_i+h} - \tilde{y}_{t_i+h|t_i}|.
\]
While the RMSE and MAE provide metrics for point forecast accuracy, they are insufficient for evaluating a full predictive distribution.
\end{comment}

The quality of a probabilistic forecast is determined by two key properties: calibration and sharpness \citep{gneiting_probabilistic_2007}. Calibration refers to the statistical consistency between the predictive distributions and the observed outcomes. It is a joint property of the forecasts and the realizations. A calibrated forecast is ``statistically honest'' because its stated probabilities match the long-run frequencies of events. Sharpness, in contrast, is a property of the forecast alone. It refers to the concentration of the predictive distribution; narrower forecasts are sharper. Sharpness is essential for a forecast to be informative. The forecaster's goal is to achieve the greatest possible sharpness subject to maintaining calibration.


\begin{comment}
To assess the overall quality of a probabilistic forecast, we need a metric that rewards both calibration and sharpness simultaneously. Proper scoring rules are the appropriate tools for this task, as they provide a single, comprehensive summary of performance \citep{gneiting_probabilistic_2007, gneiting_probabilistic_2014}. A widely used proper scoring rule is the continuous ranked probability score (CRPS) \citep{waghmare_proper_2025}. For a predictive CDF $\hat{F}_{t,h}$ and a realization $\pi_{t+h}$, the CRPS is the integrated squared difference between the forecast CDF and the empirical CDF of the observation:
\[
\text{CRPS}(\hat{F}_{t,h}, \pi_{t+h}) = \int_{-\infty}^{\infty} (\hat{F}_{t,h}(y) - \mathbf{1}_{\{y\ge\pi_{t+h}\}})^2 \, dy.
\]
The CRPS penalizes errors in both the location and the scale of the predictive distribution. A convenient property of the CRPS is that it generalizes the MAE. If the forecast $\hat{F}_{t,h}$ is a deterministic point forecast, the CRPS reduces to the absolute error \citep{gneiting_probabilistic_2007}. This allows the CRPS to be used to compare the performance of both probabilistic and point forecasts. The CRPS is also expressed in the same unit as the observation (e.g., percentage points for inflation), which aids interpretation. A lower CRPS value indicates a better forecast.
\end{comment}


In this article, we focus on forecast calibration, more specifically on probabilistic calibration, which is arguably the most common notion of calibration considered in practice \citep{Dawid1984, diebold_evaluating_1998, gneiting_probabilistic_2007}. For a given predictive CDF $\hat{F}$, and a real-valued outcome $Y$, the probability integral transform (PIT) is defined as
\[
Z_{\hat{F}}(Y) = \hat{F}(Y-) + V(\hat{F}(Y) - \hat{F}(Y-)),
\]
where $\hat{F}(y-) = \lim_{z \downarrow y} \hat{F}(z)$ is the left-limit of the CDF, and $V$ is an auxiliary random variable with $V \sim \text{UNIF}(0,1)$, independent of $(\hat{F}, Y)$. The forecast $\hat{F}$ is \emph{probabilistically calibrated} if its PIT is uniformly distributed on the unit interval, i.e., $Z_{\hat{F}}(Y) \sim \text{UNIF}(0,1)$. The randomization of the PIT through $V$ is necessary to handle potential discontinuities in the predictive CDF, ensuring that the resulting PIT is continuous.


An ensemble forecast $\mathcal{X} = (X_1, \dots, X_m)$ is a collection of $m$ possible values of the future outcome $Y$, which are typically interpreted as random draws of the conditional distribution of $Y$. The calibration of ensemble forecasts is usually assessed through rank histograms, which are closely related to PIT histograms. For simplicity, we assume that the conditional distribution of $Y$ is continuous, which is a natural assumption in our application.
Then, the rank of the outcome $Y$ with respect to the ensemble forecast $\mathcal{X}$ is
\[
\text{rank}_{\mathcal{X}}(Y) = 1 + \sum_{i=1}^m \mathbf{1}_{\{X_i < Y\}}.
\]
An ensemble forecast $\mathcal{X}$ is \emph{rank calibrated} if its rank is uniformly distributed $\{1, \dots, m+1\}$.

In our application, we observe a time series of forecasts and outcomes, $(\hat{F}_{t,h}, \pi_{t+h})_{t \in \mathbb{N}}$. Here, $\hat{F}_{t,h}$ is a forecast made at time $t$ for the realization $\pi_{t+h}$, where $h \geq 1$ is a fixed forecast horizon.

Following \citet[Definition 3.5]{arnold_sequentially_2023}, we say that
the sequence of probabilistic forecasts $(\hat{F}_{t,h})_{t \in \mathbb{N}}$ is \emph{probabilistically calibrated at horizon $h$} if the resulting sequence of PITs, $Z_t = Z_{\hat{F}_{t,h}}(\pi_{t+h})$, satisfies
    \[
    \mathcal{L}(Z_t \mid Z_j, 0 \le j \le t-h) = \text{UNIF}(0,1), \quad \text{for all  $t \in \mathbb{N}$},
    \]
    where the left-hand side denotes the conditional distribution of $Z_t$ given $Z_j$, $0 \le j \le t-h$. A sequence of ensemble forecasts $(\mathcal{X}_{t})_{t \in \mathbb{N}}$ of size $m$ is \emph{rank calibrated at horizon $h$} if the resulting sequence of ranks, $R_t = \text{rank}_{\mathcal{X}_{t}}(Y_{t+h})$, satisfies
    \[
    \mathcal{L}(R_t \mid R_j, 0 \le j \le t-h) = \text{UNIF}(\{1, \dots, m+1\}), \quad \text{for all $t \in \mathbb{N}$}.
    \]
When $t \le h$, the conditioning set is empty, so the statements are unconditional.

For the one-step-ahead case where $h=1$, the conditional distribution of each PIT, $Z_t$, is the standard uniform distribution, irrespective of the values of the previous PITs. This is precisely the condition for the sequences of PITs $(Z_t)$ and ranks $(R_t)$ to be independent and identically distributed (i.i.d.). This connects the sequential framework directly to the classical notion of calibration as formulated by \citet{diebold_evaluating_1998}, which requires an i.i.d.~uniform sequence of PITs. For forecast horizons $h > 1$, the definition is more flexible. The conditioning is on PITs that are at least $h$ time steps in the past, which allows for potential dependence between test statistics that are less than $h$ time steps apart. This reflects the overlapping information content of multi-step-ahead forecasts.

Probabilistic calibration generally does not imply that the predictive distribution $\hat{F}_{t,h}$ matches the true conditional distribution of the outcome, i.e., $\mathcal{L}(Y_{t+h}|\mathcal{F}_t) = \hat{F}_{t,h}$ \citep{gneiting_combining_2013}. If the latter was the case, we would call the forecast \emph{ideal}. We focus on testing for probabilistic calibration, as it is more practical and a necessary condition for ideal forecasts. If a sequence of forecasts fails this test, it cannot be ideal. As noted by \citet{arnold_sequentially_2023}, probabilistically calibrated forecasts are generally not ideal when forecasts use side information in addition to the target variable's own history. The forecasting methods outlined in Section~\ref{sec:forecast_design} use a wide range of macroeconomic predictors and thus fall into this category.


\begin{comment}
The e-value is a new tool for sequential testing. This section introduces the concepts of e-values and e-processes, their sequential extensions, and related concepts. Our presentation mainly follows \citet{vovk_e-values_2021} and \citet{arnold_sequentially_2023}. In all subsequent definitions, let $(\Omega, \mathcal{F})$ be a measurable space and $\mathcal{P}$ be a set of probability distributions on it. We first define a general e-value.

\begin{definition}\citep[Def.~4.1]{arnold_sequentially_2023}\label{def:evalue}
Let $\mathcal{H}_0,\mathcal{H}_1 \subset \mathcal{P}$, where $\mathcal{H}_0$ is the (possibly composite) null hypothesis. An \emph{e-value} for $\mathcal{H}_0$ is a non-negative random variable $E$ such that
\[
\mathbb{E}_P[E] \le 1 \text{ for all } P\in\mathcal{H}_0.
\]
It \emph{tests} $\mathcal{H}_0$ against $\mathcal{H}_1$ if $\mathbb{E}_Q[E]>1$ for all $Q\in\mathcal{H}_1$.
\end{definition}

Under any distribution $P$ in the null hypothesis, the e-value's expectation is at most one. A realized e-value much larger than one therefore provides evidence against $\mathcal{H}_0$. Markov's inequality allows us to convert an e-value into a test with a guaranteed Type~I error rate. For any test size $\alpha \in (0,1)$, the inequality guarantees that $\mathbb{P}_P(E \ge 1/\alpha) \le \alpha$ for all $P \in \mathcal{H}_0$. Thus, rejecting $\mathcal{H}_0$ whenever $E \ge 1/\alpha$ is a valid level-$\alpha$ test. The reciprocal $1/E$ can be used to construct a p-value. However, since $1/E$ may be larger than 1, a valid p-value is typically defined as $p' = \min(1, 1/E)$. An advantage of e-values is that they are simple to combine \citep{vovk_e-values_2021, ramdas_game-theoretic_2023}. The product of independent e-values is a valid e-value. Furthermore, any convex combination of e-values is also a valid e-value, even without independence.


To extend the e-value to a sequential setting, we use test supermartingales. Let $(\Omega, \mathcal{F})$ be a measurable space equipped with a filtration $(\mathcal{F}_t)_{t \in \mathbb{N}_0}$.

\begin{definition}\citep[Sec.~2.2]{ramdas_game-theoretic_2023}\label{def:martingale}
A stochastic process $\mathcal{M} = (\mathcal{M}_t)_{t \in \mathbb{N}_0}$, adapted to the filtration $(\mathcal{F}_t)_{t \in \mathbb{N}_0}$, is a \emph{martingale} for $\mathcal{H}_0$ if each $\mathcal{M}_t$ is integrable and
\[
\mathbb{E}_P[\mathcal{M}_t | \mathcal{F}_{t-1}] = \mathcal{M}_{t-1} \quad \text{for all } P \in \mathcal{H}_0 \text{ and for all } t \ge 1. \tag{1}\label{tag:martingale}
\]
The process $\mathcal{M}$ is a \emph{supermartingale} for $\mathcal{H}_0$ if \eqref{tag:martingale} holds with ``$\leq$'' instead of ``$=$''. If $\mathcal{M}$ is non-negative and $\mathcal{M}_0 = 1$, it is called a test supermartingale.
\end{definition}

The conditions of non-negativity and an initial value of one establish $\mathcal{M}$ as a test process for $\mathcal{H}_0$. Intuitively, $\mathcal{M}_t$ represents the capital of a gambler who repeatedly bets against the null hypothesis $\mathcal{H}_0$. Under $\mathcal{H}_0$, the supermartingale property ensures the gambler's expected capital does not grow. A large increase in $\mathcal{M}_t$ is therefore unexpected under the null and signals evidence against it.

Ville's inequality is the fundamental tool that links supermartingales to sequential testing \citep{ville_etude_1939}. We can interpret the theorem as a time-uniform extension of Markov's inequality \citep[Eq.~7abc]{ramdas_admissible_2020}.

\begin{theorem}\label{thm:ville}
If $\mathcal{M}$ is a test supermartingale for $\mathcal{H}_0$, then for every threshold $\gamma>0$,
\[
\sup_{P \in \mathcal{H}_0} \mathbb{P}_P \left( \sup_{t \in \mathbb{N}} \mathcal{M}_t \ge \gamma \right) \le \frac{1}{\gamma}.
\]
\end{theorem}

By setting $\gamma = 1/\alpha$ for some $\alpha \in (0,1)$, this implies that the probability of the process $\mathcal{M}$ ever crossing the threshold $1/\alpha$ is at most $\alpha$:
\[
\mathbb{P}_P \left( \exists\, t \in \mathbb{N}: \mathcal{M}_t \ge \frac{1}{\alpha} \right) \le \alpha \quad \text{for all } P \in \mathcal{H}_0.
\]
This inequality justifies anytime-valid testing. We can monitor $\mathcal{M}_t$ continuously for $t \ge 1$ and stop the first time it crosses $1/\alpha$. The Type~I error of this procedure is guaranteed to be no more than $\alpha$. Hence, every test supermartingale defines a sequential hypothesis test.

A related result, Doob's Optional Stopping Theorem, states that for any stopping time $\tau \in \mathbb{N}_0$, $\mathbb{E}_P[\mathcal{M}_\tau] \le \mathcal{M}_0 = 1$. By Markov's inequality, this implies $\mathbb{P}_P(\mathcal{M}_\tau \ge 1/\alpha) \le \alpha$. This confirms that we can stop at any data-dependent time and the resulting value $\mathcal{M}_\tau$ still provides valid evidence. Both results show that crossing the threshold $1/\alpha$ at any time is as significant as a single e-value exceeding $1/\alpha$. This yields an anytime-valid test, as we do not need to fix the sample size in advance.

An example from \citet{arnold_sequentially_2023} illustrates this principle. Consider a sequence of e-values $(E_t)_{t \in \mathbb{N}}$ adapted to a filtration $(\mathcal{F}_t)_{t \in \mathbb{N}}$.\footnote{To avoid any ambiguity, we stress that this ``sequence of e-values'' is not the same object as ``sequential e-values'' as in Definition~\ref{def:seq_evalue}. We use ``sequential e-values'' to mean ``sequential e-values at lag $h$''.} We can construct a process by taking the cumulative product:
\[
e_t = \prod^t_{i=1} E_i \text{ for } t \in \mathbb{N}. \tag{2}\label{tag:illustration}
\]
By the law of iterated expectations, the process $(e_t)_{t \in \mathbb{N}}$ is a non-negative supermartingale, or test supermartingale, for the null hypothesis $\mathcal{H}_0$. This holds because each $E_i$ in the sequence satisfies the condition $\mathbb{E}_{P}[E_i \mid \mathcal{F}_{i-1}] \le 1$ for all $P \in \mathcal{H}_0$. At each time $t$, the value $e_t$ is therefore a valid e-value for the joint experiment. This construction can also be viewed in reverse. Any test supermartingale $\mathcal{M}_t$ can be decomposed into a product of sequential increments. By defining $E_t := \mathcal{M}_t/\mathcal{M}_{t-1}$ (for $\mathcal{M}_{t-1}>0$), we obtain a sequence where each term satisfies $\mathbb{E}[E_t \mid \mathcal{F}_{t-1}] \le 1$ under $\mathcal{H}_0$. A test supermartingale is therefore equivalent to the cumulative product of such increments. Its existence ensures we can test sequentially without inflating the Type~I error.

An e-process is the time-indexed analogue of an e-value.

\begin{definition}\citep[Prop.~2]{grunwald_safe_2024}\label{def:eprocess}
Let $(\mathcal{F}_t)_{t\in\mathbb{N}}$ be a filtration. An adapted, non-negative process $\mathcal{E}=(E_t)_{t\in\mathbb{N}}$ with $E_0=1$ is an \emph{e-process} for $\mathcal{H}_0$ if its value at any almost surely finite stopping time $\tau$ is an e-value. That is, for every stopping time $\tau$ with respect to $\mathcal{F}_t$,
\[
\mathbb{E}_P[E_\tau] \le 1 \quad \text{for all } P\in\mathcal{H}_0.
\]
\end{definition}

The e-process property implies that we can use a data-dependent stopping rule and still trust the resulting value $E_{\tau}$ as a valid e-value. This works without inflating the Type~I error \citep{ramdas_testing_2022}. Like static e-values, e-processes yield a level-$\alpha$ test that rejects $\mathcal{H}_0$ when $E_{\tau} \ge 1/\alpha$, but this test is valid at any stopping time.

An e-process that is not a supermartingale still satisfies a Ville-type inequality. The reason is that any e-process can be dominated by a suitable test supermartingale, or constructed as an infimum over a family of them. This property leads to the composite Ville inequality, which states that for any e-process $\mathcal{E}$ for $\mathcal{H}_0$ \citep{ramdas_hypothesis_2025},
\[
\sup_{P\in \mathcal{H}_0} \mathbb{P}_{P} \left( \exists t \in \mathbb{N}: E_t \ge \frac{1}{\alpha} \right) \le \alpha.
\]

Every non-negative test supermartingale $\mathcal{M}$ is an e-process. This is because, for any stopping time $\tau$, the stopped value $\mathcal{M}_{\tau}$ inherits the property $\mathbb{E}_P[\mathcal{M}_{\tau}]\le1$ from the definition of a supermartingale. Test supermartingales are therefore a subset of e-processes. They satisfy the e-process property in addition to the stepwise supermartingale property. E-processes, however, are more general because they need not satisfy this stepwise property.

E-processes are particularly useful for testing composite null hypotheses. For example, one can construct an e-process from a family of test supermartingales, with one supermartingale, $\mathcal{M}^{(P)}$, for each distribution $P$ in the null hypothesis $\mathcal{H}_0$. The construction takes the pointwise infimum:
\[
E_t = \inf_{P\in \mathcal{H}_0} \mathcal{M}_t^{(P)}.
\]
The value $E_t$ represents the worst-case betting capital over all possible null distributions. By construction, the resulting process $(E_t)_{t \in \mathbb{N}}$ is an e-process for the composite null $\mathcal{H}_0$ \citep{ramdas_admissible_2020}. This property holds even if the infimum process is not itself a supermartingale, which highlights the generality of the e-process framework.

We now introduce sequential e-values, which are our primary tool for testing the calibration of probabilistic forecasts.

\begin{definition}\citep[Def.~4.2]{arnold_sequentially_2023}\label{def:seq_evalue}
Let $(\mathcal{F}_t)_{t\in\mathbb{N}}$ be a filtration, $h$ be a positive integer, and $\mathcal{H}_0, \mathcal{H}_1 \subset \mathcal{P}$. A sequence of adapted, non-negative random variables $(E_t)_{t \in \mathbb{N}}$ is called a sequence of \emph{sequential e-values} for $\mathcal{H}_0$ at lag $h$ if
\[
\mathbb{E}_P(E_t | \mathcal{F}_{t-h}) \le 1 \quad \text{for all } P \in \mathcal{H}_0 \text{ and for all } t \in \mathbb{N}.
\]
The sequence tests $\mathcal{H}_0$ against $\mathcal{H}_1$ if $\mathbb{E}_Q(E_t | \mathcal{F}_{t-h}) > 1$ for all $Q \in \mathcal{H}_1$ and for all $t \in \mathbb{N}$. For $t \le h$, the conditional expectation is understood unconditionally.
\end{definition}

This definition is designed for evaluating forecasts with a prediction horizon of $h$, as in Definition~\ref{def:seq_pit}. The lag-$h$ structure directly mirrors the practice of forecast evaluation. A fair evaluation of a forecast for time $t$ uses only the information available when the forecast was made, at time $t-h$. The definition, therefore, requires the corresponding e-value, $E_t$, to have an expectation of at most one, conditional on the filtration $\mathcal{F}_{t-h}$.

Combining sequential e-values depends on the lag $h$. For a lag of $h=1$, the condition becomes $\mathbb{E}_P(E_t|\mathcal{F}_{t-1}) \le 1$. This case allows us to combine the terms by their cumulative product to form a test supermartingale, as shown in (\ref{tag:illustration}). For longer horizons ($h>1$), however, this simple product rule no longer applies. Aggregating the e-values then requires a more advanced method, such as a U-statistics approach \citep{vovk_e-values_2021}.

\end{comment}





\subsection{Anytime-valid inference on calibration}
\label{sec:evalsforcalibration}

\subsubsection{Preliminaries on e-values}
E-values allow to draw inference continuously over time without compromising type~I error guarantees. They have attracted great interest in the statistical literature in recent years for several reasons including their favorable behavior under optional stopping or continuation of experiments, and since they are easy to combine in contrast to p-values \citep{ramdas_game-theoretic_2023}. An e-value is a non-negative random variable with expectation less or equal to one under the null hypothesis. By Markov's inequality, the reciprocal of an e-value is a (typically conservative) p-value. Therefore, e-values reject the null hypothesis when they are large, so they can be interpreted as the amount of evidence against the null hypothesis. Indeed, \citet{shafer_testing_2021} nicely illustrate how an e-value can be seen as a bet with unit capital against the null hypothesis. If the null hypothesis is true, one can never make money on average.


Anytime-valid inference on calibration is based on sequential e-values. That is, we consider a sequence of adapted, non-negative random variables $(E_t)_{t \in \mathbb{N}}$ that satisfy
\[
\mathbb{E}(E_t | \mathcal{F}_{t-h}) \le 1 \quad  \text{ and for all  $t \in \mathbb{N}$ under calibration}.
\]
We call $E_t$ a \emph{sequential e-value} for calibration at lag $h$. For $t \le h$, the conditional expectation is understood unconditionally \citep[Definition 4.2]{arnold_sequentially_2023}.

For lag $h=1$, we construct a process by taking the cumulative product:
\begin{equation}
e_t = \prod^t_{i=1} E_i \text{ for } t \in \mathbb{N}. \label{tag:illustration}
\end{equation}
By the law of iterated expectations, the process $(e_t)_{t \in \mathbb{N}}$ is a non-negative supermartingale with expectation less or equal to one under the null hypothesis of calibration. Such a process is called a test supermartingale for the null hypothesis of calibration. Ville's inequality \citep{ville_etude_1939}, which holds for any test supermartingale, states that, under the null hypothesis, the probability of the process $(e_t)_{t \in \mathbb{N}}$ ever crossing the threshold $1/\alpha$ is at most $\alpha$:
\[
\mathbb{P} \left( \exists\, t \in \mathbb{N}: e_t \ge \frac{1}{\alpha} \right) \le \alpha.
\]
This inequality justifies anytime-valid inference. We can monitor $e_t$ continuously for $t \ge 1$ and stop the first time it crosses $1/\alpha$. The type-I error of this procedure is guaranteed to be no more than $\alpha$. Hence, every test supermartingale defines a sequential hypothesis test $\phi_t = \mathbf{1}_{\{e_t \ge 1/\alpha\}}$.

For longer forecast horizons ($h > 1$), the simple product as at \eqref{tag:illustration} does not allow for anytime-valid inference.

\begin{prop}\citep[Prop. 4.3]{arnold_sequentially_2023}\label{prop:combine_seq_evalue}
Let $(E_t)_{t \in \mathbb{N}}$ be a sequence of sequential e-values for $\mathcal{H}_0 \subset \mathcal{P}$ at lag $h$, adapted to the filtration $(\mathcal{F}_t)_{t \in \mathbb{N}}$. For any time $t \ge h+1$, define the index sets $I_k(t) = \{k+hs : s = 0, \dots, \lfloor(t-k)/h\rfloor\}$. The aggregated value
\begin{equation}
e_t = \frac{1}{h} \sum_{k=1}^h \prod_{l \in I_k(t)} E_l \label{tag:combine_evals}
\end{equation}
is an $\mathcal{F}_t$-measurable e-value for $\mathcal{H}_0$. The process $(e_t)_{t \in \mathbb{N}}$ satisfies
\[
\mathbb{E}_P(e_{\tau+h-1}) \le 1, \quad \text{for any stopping time } \tau \text{ and any } P \in \mathcal{H}_0.
\]
\end{prop}

The properties of the process $(e_t)_{t \in \mathbb{N}}$ from Proposition~\ref{prop:combine_seq_evalue} depend on the lag $h$. For the one-step-ahead case where $h=1$, the process simplifies to the product at \eqref{tag:illustration}. In contrast, when $h > 1$, the process $(e_t)_{t \in \mathbb{N}}$ is generally not a test supermartingale. This is because $e_t$ is an average of $h$ distinct sub-processes, $M^{[k]} = (\prod_{l \in I_k(t)} E_l)_{t \in \mathbb{N}}$. Each sub-process $M^{[k]}$ is a supermartingale, but only with respect to its own specific, lagged filtration. It is not a supermartingale with respect to the underlying filtration $(\mathcal{F}_t)_{t \in \mathbb{N}}$. An average of processes that are supermartingales relative to different filtrations is not guaranteed to be a supermartingale itself. As a result, we cannot apply Ville's inequality directly.

There are two strategies to address this issue for $h>1$. The first strategy constructs a valid level-$\alpha$ test by modifying the stopping rule. As Ville's inequality still applies to each sub-process $M^{[k]}$ individually, we can track the supremum of each sub-process over time. We average these $h$ supremum processes into a single process, $\frac{1}{h} \sum_{k=1}^h \sup_{s \le t} \prod_{l \in I_k(s)} E_l$, and apply a penalty factor of $e \log(h)$, where $e$ is Euler's number. The modified rule is:
\begin{equation}
\tau_{\alpha,h} = \inf\left\{ t \in \mathbb{N} : \frac{1}{h e \log(h)} \sum_{k=1}^h \sup_{s \le t} \prod_{l \in I_k(s)} E_l \ge 1/\alpha \right\}.
\label{tag:mod_stoppingrule}
\end{equation}
This stopping time ensures that the probability of a false rejection is no more than $\alpha$, i.e., $\mathbb{P}_{\mathcal{H}_0}(\tau_{\alpha,h} < \infty) \le \alpha$ \citep{arnold_sequentially_2023}.

A second strategy, employed by \citet{henzi_valid_2022}, redefines the constituent e-values $E_t$. This approach creates a new process $(\tilde{e}_t)_{t \in \mathbb{N}}$ that is guaranteed to be a test supermartingale. However, the method can be very conservative. It works by multiplying each e-value by a correction factor that depends on the minimum possible values of future betting functions. If any future function can produce a value close to zero, the correction factor becomes very small. This severely dampens the process and reduces the power of the test.

The practical challenge is to construct e-values that are powerful enough to detect miscalibration. We construct the e-value for each new observation as a likelihood ratio. This ratio compares the simple uniform null of a uniform distribution that results under calibration with an alternative distribution from the composite family, with parameters estimated from past data.

\subsubsection{E-values for the continuous uniform distribution}\label{ssec:cont_evals}
To test for probabilistic calibration, we assess if a sequence of PITs, $(z_t)_{t \in \mathbb{N}} \subseteq [0,1]$, consists of i.i.d.~draws from the standard uniform distribution. The null hypothesis is therefore the simple hypothesis $\mathcal{H}_0 := \{\text{UNIF}(0,1)\}$. For the alternative hypothesis, $\mathcal{H}_1$, we choose the family of beta distributions. This family can flexibly capture common forms of miscalibration, including bias (skewness) and dispersion errors (U-shaped or inverse-U-shaped distributions). Let $P_{\alpha, \beta}$ denote a beta distribution parametrized by \mbox{$\Theta = \{ (\alpha, \beta) \in \mathbb{R}^2 \mid \alpha > 0, \beta > 0 \}$}. The beta family nests the null hypothesis, since $P_{1,1}$ is the standard uniform distribution.

A fixed alternative, like a single beta distribution, is impractical because the nature of any miscalibration is unknown and can change over time. We therefore adopt the sequential, adaptive strategy from \citet{arnold_sequentially_2023}. This strategy uses the history of observations to inform the choice of parameters $\alpha$ and $\beta$ for the next test. This procedure constitutes a betting strategy: at each step, we bet on the alternative that best explains all data observed so far.

The fundamental component of this strategy is the likelihood ratio for a single observation $z \in [0,1]$. This ratio tests the simple null hypothesis that $z$ is drawn from $P_{1,1}$ against a simple alternative $P_{(\alpha, \beta)}$ for some $(\alpha, \beta) \neq (1,1)$. The likelihood ratio is:
\begin{equation}
    E^{\alpha,\beta}(z) = \frac{1}{B(\alpha, \beta)} z^{\alpha-1}(1-z)^{\beta-1}, \label{tag:betaevalue}
\end{equation}
where $B(\cdot, \cdot)$ is the beta function. By construction, $E^{\alpha,\beta}(z)$ is a valid e-value. In our sequential procedure, we use past data to select the alternative for the next observation. At each time $t \geq 2$, we compute the MLE of the beta parameters, $(\hat{\alpha}_{t}, \hat{\beta}_{t})$, using all previous observations $z_1, \dots, z_{t}$:
\begin{equation}
(\hat{\alpha}_t, \hat{\beta}_t) = \operatorname*{arg\,max}_{(\alpha,\beta) \in \Theta} \sum_{i=1}^t \log(p_{(\alpha, \beta)}(z_i)), \label{tag:mle}
\end{equation}
where $p_{(\alpha, \beta)}$ is the density of $P_{(\alpha, \beta)}$. We then use these parameters, which represent the data-driven estimate of the alternative, to form the e-value for the next observation, $z_{t+1}$:
\[
E_{t+1} = E^{\hat{\alpha}_{t},\hat{\beta}_{t}}(z_{t+1}).
\]
Setting $E_1 = E_2 = 1$, the procedure generates a sequence of e-values $(E_t)_{t \in \mathbb{N}}$. For a lag of $h=1$, the cumulative product of this sequence forms a test supermartingale, as in (\ref{tag:illustration}).

For sequential e-values at lag $h>1$, we must adapt the procedure to satisfy Definition~4.2 from \citet{arnold_sequentially_2023}. We perform the parameter estimation separately on $h$ disjoint subsequences of the data. For each $k \in \{1, \dots, h\}$, the $k$-th subsequence is defined by the indices $\{k + hs \mid s = 0,1,\dots\}$. For any time $t \ge 2h$, we compute $h$ separate MLEs on these subsequences:
\begin{equation}
(\hat{\alpha}_t^k, \hat{\beta}_t^k) = \operatorname*{arg\,max}_{(\alpha,\beta) \in \Theta} \sum_{s: k+hs \le t} \log(p_{(\alpha,\beta)}(z_{k+hs})), \quad k=1,\dots,h. \label{tag:laggedmle}
\end{equation}
The e-value for a new observation then uses the most recent lagged MLE from its corresponding subsequence. We initialize $E_1 = \dots = E_{2h} = 1$, and for $t = h,h+1, \dots$, we compute the sequential e-value as:
\[
E_{k+th} = E^{\hat{\alpha}_{t}^k,\hat{\beta}_{t}^k}(z_{k+th}), \quad k=1,\dots,h.
\]
The resulting sequence $(E_t)_{t\in\mathbb{N}}$ constitutes sequential e-values at lag $h$ for the uniform null hypothesis that $z_t \sim \text{UNIF}(0,1)$, conditional on $z_1, \dots, z_{t-h}$. As described in Proposition~\ref{prop:combine_seq_evalue}, we then combine these e-values using the formula in (\ref{tag:combine_evals}) to form an aggregated e-value at each time step.

In practice, we must address two potential pitfalls. First, the continuous uniform null assigns zero probability to the boundary points $\{0,1\}$, but empirical PITs can take these values. An observation of 0 or 1 can cause the MLE in (\ref{tag:mle}) to diverge. For the beta e-value in (\ref{tag:betaevalue}), this causes the likelihood ratio to become infinite or exactly zero. An e-value of zero is fatal for the cumulative product, as it permanently reduces the process to zero. In the betting interpretation of \citet{shafer_testing_2021}, this corresponds to losing all capital. We solve this problem by ignoring any observations in $\{0,1\}$ during parameter estimation and e-value calculation. This is a valid strategy because it does not alter the e-value's expectation under the null. As \citet{arnold_sequentially_2023} argue, this omission has little effect if boundary observations are rare; if they are frequent, the null hypothesis of a continuous uniform distribution is clearly false. Second, to ensure parameter estimates are stable, we begin the sequential estimation only after a warm-up period. Following \citet{arnold_sequentially_2023}, we set the first $n_0=10$ e-values to 1. This provides a buffer beyond the absolute minimum of two observations required to compute the MLE for the two-parameter beta distribution.

\subsubsection{E-values for the discrete uniform distribution}\label{ssec:discrete_evals}
To test for rank calibration, we assess if a sequence of ensemble ranks, $(r_t)_{t \in \mathbb{N}} \subseteq \{1, \dots, m\}$, is i.i.d.~uniform. For $m \geq 1$, the null hypothesis is therefore the simple hypothesis that $\mathcal{H}_0 := \{\text{UNIF}(\{1, \dots, m\})\}$. For the alternative hypothesis, $\mathcal{H}_1$, we use the family of beta-binomial distributions, parametrized by \mbox{$\Theta = \{ (\alpha, \beta) \in \mathbb{R}^2 \mid \alpha > 0, \beta > 0 \}$}. This choice is the discrete analogue of the beta distribution and can flexibly model common forms of miscalibration, like bias and dispersion errors in the rank histogram.

We construct the e-value using the same adaptive strategy as in the continuous case. The e-value for a single rank observation $r \in \{1, \dots, m\}$ is the likelihood ration of a specific alternative $P_{(\alpha, \beta)}$ to the uniform null:
\begin{equation}
E^{\alpha,\beta}(r) = \frac{p_{(\alpha, \beta)}(r)}{p_0(r)} = m \cdot p_{(\alpha, \beta)}(r), \label{tag:betabinomialevalue}
\end{equation}
where $p_0(r) = 1/m$ is the probability mass function (PMF) of the discrete uniform distribution and $p_{(\alpha, \beta)}(r)$ is the PMF of the beta-binomial distribution:
\[
p_{(\alpha, \beta)}(r) = \binom{m-1}{r-1} \frac{B(\alpha + r - 1, \beta + m - r)}{B(\alpha, \beta)}.
\]
By construction, the likelihood ratio in Equation~(\ref{tag:betabinomialevalue}) is an e-value for testing the uniform null against a simple beta-binomial alternative $P_{(\alpha, \beta)}$. We implement this test using a sequential procedure. For each new rank $r_{t+1}$ with $t \ge 2$, we first compute the MLE of the beta-binomial parameters, $(\hat{\alpha}_{t}, \hat{\beta}_{t})$, using all past ranks $r_1, \dots, r_{t}$:
\[
(\hat{\alpha}_t, \hat{\beta}_t) = \operatorname*{arg\,max}_{(\alpha,\beta) \in \Theta} \sum_{i=1}^t \log(p_{(\alpha, \beta)}(r_i)).
\]
We then form the e-value for the next rank $r_{t+1}$ using these estimated parameters:
\[
E_{t+1} = E^{\hat{\alpha}_{t},\hat{\beta}_{t}}(r_{t+1}).
\]
We initialize the sequence by setting $E_1 = E_2 = 1$ to have a sufficient sample for the first MLE. For a forecast horizon of $h=1$, the cumulative product of the resulting sequence $(E_t)_{t \in \mathbb{N}}$ forms a test supermartingale.

For sequential e-values at lag $h>1$, we perform the parameter estimation separately on the $h$ subsequences of the data, as in the continuous case. However, the beta-binomial MLE can be unstable with few observations, a problem exacerbated by partitioning the data. To ensure robust estimates, we again follow the recommendation of \citet{arnold_sequentially_2023} and introduce a warm-up period for each subsequence. Specifically, we set the first $n_0=10$ e-values for each of the $h$ subsequences to 1 before beginning their MLE procedure. After the warm-up, we compute the MLE for each $k$-th subsequence as in Equation~(\ref{tag:laggedmle}). The e-value for a subsequent observation is then calculated with the lagged estimates from its corresponding subsequence:
\[
E_{k+th} = E^{\hat{\alpha}_{t}^k,\hat{\beta}_{t}^k}(r_{k+th}), \quad k=1,\dots,h.
\]
The resulting sequence $(E_t)_{t\in\mathbb{N}}$ constitutes sequential e-values at lag $h$ for the discrete uniform null hypothesis, which we then combine using the formula in~(\ref{tag:combine_evals}).

\subsection{Implementation of sequential calibration tests}
\label{sec:application}
This section details how we apply the theoretical framework from Section~\ref{sec:evalsforcalibration}. We use this framework to evaluate the probabilistic inflation forecasts generated by the models in Section~\ref{sec:forecast_design}. Our evaluation is implemented in \texttt{R} using the \texttt{epit} package for the e-value computations~(\url{https://github.com/AlexanderHenzi/epit}). Our analysis simulates a realistic, real-time monitoring scenario, providing a dynamic assessment of forecast calibration. We structure the analysis by iterating through each combination of geographic region (United States, Euro Area, Switzerland), forecast horizon (in months, $h \in \{1, 3, 6, 12\}$), and forecasting model. To ensure a fair comparison, the evaluation period for each region begins only when all models provide valid, non-missing forecasts\footnote{For example, the series $(\hat{F}_{t,h}, \pi_t)_t$ is one month shorter for models requiring differencing of the time series (BVAR, DFM).}.

We group the models into three types. For the majority of models that produce a Gaussian predictive distribution -- the Gaussian baselines, ARIMA, BVAR, and DFM -- we test calibration using the PIT. We compute the PIT for each forecast by evaluating the predicted Gaussian CDF at the realized inflation rate. For the probabilistic no-change model, which generates a non-parametric ensemble of $m$ recent inflation values, we assess calibration using ranks. We calculate the rank of the true outcome relative to the ensemble members. Finally, the distributional random forest produces a non-parametric distribution as a set of weighted samples, which results in a discontinuous CDF. Therefore, we compute a randomized PIT for each DRF forecast, as detailed in Section~\ref{sec:notionscalibration}.

With these sequences of PITs and ranks, we sequentially test the null hypothesis of calibration ($\mathcal{H}_0: z_t \sim \text{UNIF}(0,1)$ or $\mathcal{H}_0: r_t \sim \text{UNIF}\{1, \dots, m\}$). We employ the adaptive betting strategy from Section~\ref{sec:evalsforcalibration}. This requires choosing a family of alternative distributions against which to test the uniform null. For PITs, we use the beta distribution family; for ranks, we use the beta-binomial family. While \citet{arnold_sequentially_2023} also explore a kernel density method for the alternative, they find that tests based on the beta family are more powerful in small samples. We therefore adopt their beta-based approach. Finally, to handle PIT values that are exactly 0 or 1, we follow \citet{arnold_sequentially_2023} and exclude these boundary observations from the parameter estimation step, but we record their frequency for each model and horizon.

The method for aggregating the sequential e-values depends on the forecast horizon $h$. For one-step-ahead forecasts ($h=1$), the underlying information sets are non-overlapping. The cumulative product of the sequential e-values is therefore a test supermartingale. This allows us to use the standard stopping rule for calibration testing, as described in the paragraph following Proposition~\ref{prop:combine_seq_evalue}. In contrast, for multi-step-ahead forecasts ($h>1$), the information sets overlap. To handle this, we follow the U-statistics approach from \citet{arnold_sequentially_2023}, see Proposition \ref{prop:combine_seq_evalue}. We split the sequence of test statistics into $h$ non-overlapping subsequences, calculate a separate e-value sequence for each, and then aggregate them into a single value at each time step. The resulting sequence of aggregated e-values is not a test supermartingale. Therefore, to conduct a formal test, we must apply the less powerful, modified stopping rule from~(\ref{tag:mod_stoppingrule}).

The cumulative e-values for each model and forecast horizon are the primary output of our analysis. The e-values track how evidence against the null hypothesis builds up over time. For $h=1$, we treat this process as a statistically anytime-valid test. We reject $\mathcal{H}_0$ once it crosses a pre-specified threshold, like $1/\alpha = 100$ for a significance level of $\alpha = 0.01$.\footnote{In traditional scientific hypothesis testing, there are (discipline-specific) conventions on thresholds for p-values, typically 0.05 or 0.01. There is not yet a consensus on thresholds for e-values. For our purposes, we regard an e-value of around 100 as sufficient for a clear rejection. We regard a p-value of 0.01 as a clear rejection as well, but we still report p-values between 0.01 and 0.1. For a more thorough discussion of the comparison between e-values and p-values, see \citet[Sec. 2.7]{ramdas_hypothesis_2025}.} We also use the path of the process to identify when miscalibration appears within the evaluation period. For $h > 1$, we can formally reject the null hypothesis of calibration by applying the modified stopping rule and its corresponding threshold defined in \eqref{tag:mod_stoppingrule}. To provide a complete picture, we display these formal rejection thresholds alongside the supremum process in our graphical diagnostics. Beyond formal binary rejections, a major benefit of the e-value-based methodology is the ability to monitor the tests sequentially over time to see exactly when miscalibration develops. Sustained increases in the e-values provide valuable insights into accumulating evidence against forecast calibration, even during periods where the formal threshold has not yet been crossed.

The cumulative e-values show when our forecasts are miscalibrated, but they do not show how. We therefore complement them with graphical tools. A common static tool is the PIT histogram, which provides end-of-sample information about the type of miscalibration. Following \citet{arnold_sequentially_2023}, we go further and exploit the e-value procedure itself. The e-value construction specifies an alternative density, determined by the data at each time point. We plot these estimated densities over time to see in what way the forecasts are misspecified. In practice, we use the beta densities that enter the e-value construction. We also add non-parametric, boundary-corrected kernel density estimates. This additional method investigates if the beta specification captures the main features of the empirical PIT density. To compare the e-value-based test with a traditional static test, we compute the one-sample Kolmogorov-Smirnov (KS) test to evaluate the uniformity of the PITs. We apply the KS test to the full sample of PITs for each model except the probabilistic no-change forecast (\url{https://stat.ethz.ch/R-manual/R-devel/library/stats/html/ks.test.html}). Our analysis, therefore, contrasts dynamic evidence from the e-values with static, end-of-sample evidence from a classical test. This approach lets us detect whether a model is miscalibrated overall and identify the periods in which the forecasts are not calibrated. These features illustrate the practical value of sequentially valid inference for continuous monitoring.


\section{Probabilistic inflation forecasting}
\label{sec:forecast_design}
This section presents the results of the empirical application, which focuses on an out-of-sample forecasting experiment for probabilistic inflation forecasts. We begin by briefly describing the datasets employed and outlining the overall forecast design.

\subsection{Data and forecast design}
\label{sec:data}
Our empirical analysis focuses on three economies: the United States (US), the Euro Area (EA), and Switzerland (CH). We take the US data from the FRED-MD database maintained by the Federal Reserve Bank of St. Louis \citep{mccracken_fred-md_2015}. For the EA, we rely on data from the Statistical Data Warehouse of the European Central Bank. For Switzerland, we compile publicly available data from the Swiss National Bank\citep{swiss_national_bank_snb_2025}\footnote{\url{https://data.snb.ch/en}}.

For each economic region, we construct a harmonized set of predictors that balances cross-country comparability with data availability. We retain only variables that are available at a monthly frequency and free of missing observations over the relevant sample period. Each predictor set comprises approximately 20 time series capturing monetary and credit aggregates, interest rate conditions, as well as consumer and producer price developments. A complete overview of the included series is provided in Appendix~\ref{app:data}.

The sample periods start at different dates but all end in April 2025. Our target variable is the year-over-year consumer price inflation rate, denoted by $\pi_t$, defined as
\begin{equation}
\pi_t \coloneq \frac{P_t}{P_{t-12}} - 1, \label{tag:inflation}
\end{equation}
where $P_t$ is the price index in month~$t$. We aim to forecast $\pi_{t+h}$ for forecast horizons $h \in \{1,3,6,12\}$. Forecasts at time~$t$ use only the information set $\mathcal{F}_t$, which contains all data available up to that month. This set includes past values of inflation $\{\pi_t, \pi_{t-1}, \dots\}$ and relevant macroeconomic predictors like interest rates and inflation components. The three datasets allow us to compare models and their calibration across economies with different inflation dynamics and sample sizes (see Figure~\ref{fig:infl}).

\begin{figure}[t]
  \centering
  \includegraphics[width=\textwidth]{Pictures/plots/eda/inflation_plot.pdf}
  \vspace{-2.25em}
  \caption{Monthly year-over-year inflation rates for the three monetary regions. The vertical dotted line marks the end of the initial 40\% training period and the beginning of the out-of-sample evaluation window.}
  \label{fig:infl}
\end{figure}

We mimic a realistic forecasting workflow as follows. We fix an initial training window at 40\% of each region's sample. At each evaluation month~$t$, we expand the training window to include all information up to $t$ and re-estimate the model parameters. We then issue the $h$-step predictive distribution $\hat{F}_{t,h}$ for inflation $\pi_{t+h}$. We update hyperparameters (like the number of lags) on a coarser schedule. We re-tune them whenever we observe at least a further 10\% of the full sample, that is, at 50\%, 60\%, $\dots$, and 90\% of each region's sample. Between these points, we keep hyperparameters fixed but still re-estimate model parameters each month using all data available at that time. Repeating this procedure yields a series ${(\hat{F}_{t,h}, \pi_{t+h})}_{t}$ for each horizon $h$.

The sample periods and sizes vary across the three regions. For the US, the data spans from January~1960 to April~2025 and contains 784 monthly observations. The out-of-sample evaluation begins after the initial training period ends in February~1986, leaving 470 monthly forecasts to evaluate. The Swiss dataset covers January~1983 to April~2025 (508 observations). The evaluation period starts after December~1999 and contains 305 forecasts. The Euro Area sample is the shortest. It runs from January~2001 to April~2025 (292 observations) and yields an evaluation window of 175 forecasts beginning after September~2010. Naturally, the evaluation windows are shorter for longer forecast horizons. Figure~\ref{fig:infl} plots the inflation series for each region, with a vertical line indicating the split between the initial training data and the evaluation period.

The target variable in each dataset is the year-over-year percentage change in the headline consumer price index: the Consumer Price Index for All Urban Consumers (CPIAUCSL) for the US, the Harmonized Index of Consumer Prices (HICP) for the EA, and the Swiss Consumer Price Index (CPI) for Switzerland. Some forecasting methods require stationary inputs, so we transform all predictor variables accordingly. We follow the transformations in \citet{mccracken_fred-md_2015} and the recommendations of each data provider. We report the exact transformations in Appendix~\ref{app:data}. To reduce outlier influence, we winsorize all stationary predictors at the 0.5th and 99.5th percentiles. We do not winsorize the inflation target series to preserve their original dynamics.

\subsection{Forecasting Models}

We consider a set of naive benchmark, standard time series, and machine learning models for computing the probabilistic inflation forecasts. Specifically, the set includes the naive last forecast, the rolling mean forecast, the probabilistic no-change forecast (PNC), the autoregressive integrated moving average model (ARIMA), Bayesian vector autoregressions with diffuse priors (BVAR (diffuse)), Minnesota priors (BVAR (Minnesota)), and Normal-Wishart priors (BVAR (NW)), the dynamic factor model (DFM), and the distributional random forest (DRF). All models are estimated recursively in an expanding-window framework and produce predictive distributions at horizons up to 12 months. Detailed descriptions of the models and their implementation are provided in Appendix~\ref{app:models}.

\begin{table}[!t]
  \centering
  \caption{Forecasting performance for the United States}
  \begin{threeparttable}
  \begin{tabular}{l cccccc}
\hline\hline
 & \multicolumn{3}{c}{$h=1$} & \multicolumn{3}{c}{$h=3$} \bigstrut[t]\\
\cmidrule(lr){2-4} \cmidrule(lr){5-7}
\textbf{Model} & \textbf{RMSE} & \textbf{MAE} & \textbf{CRPS} & \textbf{RMSE} & \textbf{MAE} & \textbf{CRPS} \bigstrut[b]\\
\hline
\multicolumn{7}{l}{\textit{Baseline Models}} \bigstrut[t]\\
Naive last & 0.398 & 0.276 & 0.206 & 0.865 & \textbf{0.575} & 0.441 \\
Rolling mean & 1.412 & 1.015 & 0.792 & 1.552 & 1.121 & 0.878 \\
PNC & 1.412 & 1.102 & 0.740 & 1.552 & 1.206 & 0.847 \\
ARIMA(1,1,0) & \textbf{0.362} & \textbf{0.259} & \textbf{0.191} & \textbf{0.849} & 0.577 & \textbf{0.436} \bigstrut[b]\\
\hline
\multicolumn{7}{l}{\textit{BVAR Models}} \bigstrut[t]\\
BVAR (diffuse) & 0.489 & 0.337 & 0.252 & 0.935 & 0.631 & 0.481 \\
BVAR (Minnesota) & 0.486 & 0.335 & 0.250 & 0.935 & 0.630 & 0.480 \\
BVAR (NW) & 0.483 & 0.332 & 0.249 & 0.920 & 0.617 & 0.471 \bigstrut[b]\\
\hline
\multicolumn{7}{l}{\textit{Multivariate \& ML Models}} \bigstrut[t]\\
DFM & 0.364 & 0.262 & 0.193 & 0.856 & 0.590 & 0.442 \\
DRF & 0.590 & 0.395 & 0.296 & 0.924 & 0.627 & 0.467 \bigstrut\\\hline\hline

 & \multicolumn{3}{c}{$h=6$} & \multicolumn{3}{c}{$h=12$} \bigstrut[t]\\
\cmidrule(lr){2-4} \cmidrule(lr){5-7}
\textbf{Model} & \textbf{RMSE} & \textbf{MAE} & \textbf{CRPS} & \textbf{RMSE} & \textbf{MAE} & \textbf{CRPS} \bigstrut[b]\\
\hline
\multicolumn{7}{l}{\textit{Baseline Models}} \bigstrut[t]\\
Naive last & 1.265 & 0.878 & 0.669 & 1.864 & 1.340 & 1.017 \\
Rolling mean & 1.716 & 1.246 & 0.983 & 1.907 & 1.419 & 1.125 \\
PNC & 1.716 & 1.316 & 0.975 & 1.907 & 1.489 & 1.135 \\
ARIMA(1,1,0) & 1.254 & 0.877 & 0.659 & 1.873 & 1.354 & 1.005 \bigstrut[b]\\
\hline
\multicolumn{7}{l}{\textit{BVAR Models}} \bigstrut[t]\\
BVAR (diffuse) & 1.317 & 0.921 & 0.715 & 1.933 & 1.376 & 1.103 \\
BVAR (Minnesota) & 1.318 & 0.926 & 0.715 & 1.931 & 1.376 & 1.100 \\
BVAR (NW) & 1.307 & 0.911 & 0.708 & 1.921 & 1.367 & 1.099 \bigstrut[b]\\
\hline
\multicolumn{7}{l}{\textit{Multivariate \& ML Models}} \bigstrut[t]\\
DFM & 1.266 & 0.888 & 0.665 & 1.911 & 1.374 & 1.033 \\
DRF & \textbf{1.173} & \textbf{0.785} & \textbf{0.575} & \textbf{1.341} & \textbf{0.864} & \textbf{0.620} \bigstrut\\\hline\hline
\end{tabular}
\vspace*{-0.5cm}
\begin{tablenotes}
    \singlespacing
    \item \leavevmode\kern-\scriptspace\kern-\labelsep
    Note: The minimum value achieved for a given evaluation metric (RMSE, MAE, and CRPS) is highlighted in bold.
\end{tablenotes}
  \end{threeparttable}
  \label{tab:us_performance_summary}
\end{table}

\subsection{Out-of-Sample Forecasting Results}
\label{sec:forecast_summary}
In the following, we summarize the out-of-sample forecasting accuracy measured in terms of the RMSE, MAE, and CRPS, for the US and the Euro Area. The results are illustrated in Table~\ref{tab:us_performance_summary} and Table~\ref{tab:eu_performance_summary} in the Appendix. For both regions, the ARIMA(1,1,0) model consistently achieves the lowest errors across short horizons ($h = 1, 3$), closely followed by the DFM and the BVARs. The DRF performs competitively at longer horizons ($h = 6, 12$), but remains less accurate for near-term forecasts. The improved performance at longer forecasting horizons is likely attributable to the use of longer inflation lags in the DRF, which allows for a more effective representation of persistent dynamic behavior. The baseline models, especially the Gaussian no-change forecast (``naive last''), perform surprisingly well, confirming the persistence of monthly inflation. Overall, the relative ranking of the models is stable across the countries, and the results for the Switzerland (see Table~\ref{tab:ch_performance_summary} in the Appendix) mirror those for United States at slightly lower error levels. These findings provide an empirical foundation for the calibration evaluation in Section~\ref{sec:results}.

\section{Evaluation of probabilistic inflation forecasts}
\label{sec:results}

This section reports the main empirical findings. The sequential testing framework introduced in Section~\ref{sec:evaluation_theory} is employed to assess the probabilistic inflation forecasts produced by the models described in Section~\ref{sec:forecast_design}. The analysis primarily relies on cumulative e-values, as elaborated in Subsections~\ref{ssec:cont_evals} and~\ref{ssec:discrete_evals}, which facilitate a dynamic assessment of forecast calibration. To underscore the additional insights offered by the sequential perspective, these results are compared with those obtained from the static one-sample Kolmogorov–Smirnov (KS) test. The e-value approach allows not only for the detection of forecast miscalibration but also for the identification of its timing and nature.

The results are presented separately for each model class and focus on the United States and the Euro Area. Outcomes for the United States are shown in Figure~\ref{fig:us_advanced_short} and in Figures~\ref{fig:us_baseline_short}--\ref{fig:us_bvar_long} in the Appendix, while those for the Euro Area are reported in Figure~\ref{fig:eu_advanced_short} and Figures~\ref{fig:eu_baseline_short}--\ref{fig:eu_bvar_long} in the Appendix. Results for Switzerland are deferred to Appendix~\ref{chap:results_ch}. Each figure displays probability integral transform (PIT) histograms for the full evaluation sample using 20 bins in panel~(a)\footnote{For the PNC method, the normalized rank histogram is reported instead.} and the evolution of cumulative e-values over time in panel~(b). For $h > 1$, the dashed grey line represents the supremum process used for the modified stopping rule from \eqref{tag:mod_stoppingrule}. Panels~(a) and~(b) are organized by forecast horizon (columns) and model specification (rows). To examine changes in forecast misspecification over time, panel~(c) is included. This panel presents estimates of the alternative PIT density based on observations from the highlighted subsample (solid line) and from the preceding sample excluding this period (dotted line). The column headers indicate the corresponding highlighted time intervals. Two approaches are used to estimate the alternative density. First, a beta distribution is fitted, which also serves as the basis for constructing the e-values (blue lines). Second, a boundary-corrected kernel density estimator is applied, shown in black, following \citet{arnold_sequentially_2023}.

\subsection{Results for the Baseline models}
\label{sec:baseline_results}
The baseline forecasts exhibit heterogeneous calibration properties. Some models maintain adequate calibration, mirroring their strong forecast accuracy. Others exhibit clear and significant miscalibration.

For the US at forecast horizon $h=1$, the ARIMA(1,1,0), rolling mean, and Gaussian no-change forecasts all exhibit notable miscalibration (Figures~\ref{fig:us_advanced_short} and~\ref{fig:us_baseline_short}). For the ARIMA(1,1,0) and Gaussian no-change approaches, the processes in Figures~\ref{fig:us_advanced_short} and~\ref{fig:us_baseline_short} show that most of the evidence against calibration accumulates during the 1990s. As shown in Table~\ref{tab:us_calibration_summary}, the static one-sample KS test misses these timing differences. It does not reject calibration for the ARIMA and Gaussian no-change forecasts for $h=1$, which is consistent with their approximately uniform full-sample PIT distributions.

\begin{figure}[!t]
  \centering
  \includegraphics[width=1\textwidth]{Pictures/plots/US/us_advanced_horizons_short_calibration.pdf}
    \vspace{-2.25em}
  \caption{Calibration diagnostics for ARIMA(1,1,0), DFM and DRF models at short horizons ($h=1,3$) for the United States. Panel (a) shows PIT histograms. Panel (b) shows cumulative e-values on a log scale; note the different y-scale across rows. The dotted horizontal lines indicate levels of 1 and the rejection thresholds. Panel (c) shows the evolution of the density estimates of the PIT; note that the y-axes are limited for better visibility, and the maximally attained value is given in the plot.}
  \label{fig:us_advanced_short}
\end{figure}

At longer horizons, these forecasts behave differently, as shown in Figure~\ref{fig:us_advanced_long} in the Appendix. For the ARIMA(1,1,0) model, the evidence against calibration weakens as the horizon increases. Both the one-sample KS test and the cumulative e-value process support this conclusion (Table~\ref{tab:us_calibration_summary} in the Appendix). For the Gaussian no-change forecast, however, the cumulative e-values show diminishing power, while the KS test produces p-values in the range of $0.001-0.005$ (Figure~\ref{fig:us_baseline_long}).

The rolling mean forecast is clearly miscalibrated across all horizons. The static KS test rejects calibration with a p-value that is effectively zero, and the sequential analysis reaches the same conclusion: for $h=1$, the process reaches values of order $10^{34}$ (Table~\ref{tab:us_calibration_summary}). The e-values also reveal a distinct dynamic pattern. They accumulate evidence against calibration during relatively calm economic periods but drop sharply during the major crises in 2008-2009 and 2021-2023 (Figure~\ref{fig:us_baseline_short}).


The Euro Area results illustrate the value of sequential testing in a relatively short sample with a pronounced regime shift (Figure~\ref{fig:infl}). For the one-month horizon, the processes for the baseline models rise sharply only at the end of the evaluation period, coinciding with the COVID-19 inflation shock (Figures~\ref{fig:eu_advanced_short} and~\ref{fig:eu_baseline_short}). This pattern is striking because the static one-sample KS test does not reject calibration for any of these methods at short horizons ($h \le 3$, see Table~\ref{tab:eu_calibration_summary}). The discrepancy suggests that the baseline forecasts suffered a transient but severe miscalibration during the pandemic shock. The sequential method detects this immediately, whereas the KS test, applied over the full sample, does not.

For longer horizons in the EA sample, the results are more complex. For the Gaussian no-change forecast, the cumulative e-values no longer provide evidence of miscalibration for $h \ge 3$. In contrast, for the ARIMA(1,1,0) model, the cumulative e-values continue to increase. They indicate miscalibration for $h=3$ and $h=6$ and reach values of order $10^{5}$ for $h=12$ (Figures~\ref{fig:eu_advanced_long} and~\ref{fig:eu_baseline_long}). The rolling mean method shows the familiar loss of power: its cumulative e-values no longer provide strong evidence of miscalibration (Figure~\ref{fig:eu_baseline_long}). The static KS test results at longer horizons are also difficult to interpret (Table~\ref{tab:eu_calibration_summary}). For example, for $h=6$ the one-sample KS test returns a small p-value only for the Gaussian no-change method ($p=0.04$), where the cumulative e-values do not grow. For $h=12$ it rejects calibration for the rolling mean forecast ($p=0.01$), where the e-values also remain flat.


\begin{figure}[!t]
  \centering
  \includegraphics[width=1\textwidth]{Pictures/plots/EU/eu_advanced_horizons_short_calibration.pdf}
    \vspace{-2.25em}
  \caption{Calibration diagnostics for ARIMA(1,1,0), DFM and DRF models at short horizons ($h=1,3$) for the Euro Area. Panel (a) shows PIT histograms. Panel (b) shows cumulative e-values on a log scale; note the different y-scale across rows. The dotted horizontal lines indicate levels of 1 and the rejection thresholds. Panel (c) shows the evolution of the density estimates of the PIT; note that the y-axes are limited for better visibility, and the maximally attained value is depicted in the plot.}
  \label{fig:eu_advanced_short}
\end{figure}

\subsection{Results for the BVAR models}
\label{sec:bvar_results}
Despite their sophistication, the BVAR models are not calibrated across most regions and horizons. Their PIT histograms consistently show that the forecasts are overconfident. The choice of prior specification does not alter the overall dynamics of the miscalibration. Among the three priors, the Normal-Wishart prior produces the poorest results. Its e-values grow fastest and reach the highest peaks, providing the strongest evidence of miscalibration.

For the United States, the global financial crisis (GFC) of 2008–2009 is the main episode of BVAR miscalibration. Across all forecast horizons, the cumulative e-values increase sharply during this period and reach levels of order $10^{15}$ (for the Minnesota prior at $h=3$) up to $10^{36}$ (for the Normal–Wishart prior at $h=12$). The PIT histograms support this conclusion: they show clear overconfidence for $h=1$, which becomes more pronounced at longer horizons. Despite the clear results from the sequential test, the static KS test does not reject calibration for $h=1$ (Table~\ref{tab:us_calibration_summary}). The test supermartingale, in contrast, makes the miscalibration at this horizon visible (Figure~\ref{fig:us_bvar_short}). It also provides a timeline for when the BVARs' probabilistic calibration deteriorates, which is valuable for forecasters and policymakers who monitor models in practice.


For the EA, the cumulative e-values for all BVAR specifications grow extremely rapidly during the 2021-2022 inflation surge. For the three-month horizon ($h=3$), the Normal-Wishart prior BVAR's cumulative e-value exceeds the rejection threshold by mid-2021 (Figure~\ref{fig:eu_bvar_short}). This constitutes evidence against calibration. While slightly less extreme, the Minnesota and diffuse prior BVARs also reach high values over the same period. The PIT histograms in Figure~\ref{fig:eu_bvar_short} show that many realized inflation rates lie in the tails of the predictive densities. The speed at which the e-values reach these levels confirms that the BVARs did not adjust their uncertainty estimates in time for the regime change. The sequential analysis highlights a dynamic that the static KS test misses, particularly at short horizons. While the one-sample KS test does not reject calibration for $h=1$ (Table~\ref{tab:eu_calibration_summary}), the process indicates miscalibration as it unfolds. The KS test does eventually detect the issue, but only for longer horizons.

Across all priors and forecast horizons, the e-values tell a consistent story. They show no evidence against calibration during the first 20 years of stable inflation. They then exhibit a sharp, sudden spike during the COVID-19 shock (Figure~\ref{fig:eu_bvar_long}). The miscalibration during this shock overcomes the power loss of the sequential test at longer horizons. This finding exemplifies the advantage of using e-values in a realistic monitoring context: the BVARs' misspecification was evident as it happened.


\subsection{Results for the DFM}
\label{sec:dfm_results}
The DFM forecasts present a mixed picture. They are generally better calibrated than the BVARs, and their PIT histograms and e-values show properties similar to the ARIMA(1,1,0) model. The DFM results for the US provide further insights into the model's dynamic calibration. For the one-month horizon ($h=1$), the process shows an evolution similar to the ARIMA(1,1,0) model. It reaches an order of magnitude of $10^4$ during the 1990s before decreasing until the end of the evaluation period. On the other hand, the PIT histogram in panel~(a) of Figure~\ref{fig:us_advanced_short} looks approximately uniform. To investigate this in detail, we look at the evolution of the estimated alternative densities over selected time windows in panel~(c1) of Figure~\ref{fig:us_advanced_short}. We see from the inverse-U-shape in the first subplot that the initial increase is due to an overdispersion of the forecasts. Later, in the early 2000s, the forecasts were too narrow because the estimated alternative during that period was U-shaped. However, the process decreases during that time period. This decrease illustrates that the e-values only have power if the type of miscalibration aligns with the currently chosen alternative distribution. The decrease does not mean that the forecasts are better calibrated during this time period, but the misspecification merely changes direction. There is a short rise in the process during the GFC due to underdispersion of the forecasts, which subsides after the GFC. Now, the test supermartingale has power for this miscalibration because the dotted line (the current alternative) is also U-shaped. From 2010 onwards, the forecasts seem to depict uncertainty accurately. The static KS test, however, gives a p-value of $0.97$ against uniformity (Table~\ref{tab:us_calibration_summary}). Hence, this is an exemplary case of the usefulness of the sequential method, as it allows insights that are not possible through traditional methods.


At longer horizons, the PIT histograms begin to show clear signs of overconfidence, with values bunching at 0 and 1. Still, the static KS test only finds evidence against calibration for $h=12$ with a p-value of $0.02$ (Table~\ref{tab:us_calibration_summary}). The e-values for $h > 3$ continue the pattern seen for $h=3$: a period of calibration in the 1990s followed by a steep increase during the 2008-2009 crisis. However, the e-values again show a loss of power as the horizon lengthens. While the GFC shock is visible for $h=6$, the increase is less pronounced than for $h=3$ and does not provide the same level of immediate evidence. Instead, strong evidence of miscalibration at this horizon only accumulates later, during the COVID-19 shock in 2021 (Figure~\ref{fig:us_advanced_long}).

For the EA, the DFM's one-month-ahead forecasts initially appear calibrated, and in fact, better calibrated than the ARIMA model for much of the evaluation period. The PIT histogram for the DFM for $h=1$ is fairly uniform, and the one-sample KS test returns a p-value of $0.815$, indicating no grounds for rejection. However, the sequential test tells a different story. The process for the DFM remains near 1 for most of the sample, but begins to rise sharply in late 2022. It ultimately reaches a level of order $10^7$, decisively rejecting calibration for the one-month horizon. To check the type of miscalibration, we look at the evolution of the estimated alternative distribution in panel~(c1) of Figure~\ref{fig:eu_advanced_short}. The first subplot confirms that the DFM was probabilistically calibrated before the COVID-19 shock, because the alternative densities align well with the uniform line. By contrast, in 2022 and 2023, the estimated alternative is U-shaped. Therefore, prediction uncertainty was underestimated. After 2023, the process levels off and based on the mild inverse-U-shaped density estimate from this period, calibration was approximately appropriate. This pattern is robust across the longer horizons, where the cumulative e-values provide evidence of miscalibration, growing to orders of magnitude between $10^{12}$ and $10^{25}$ (Figure~\ref{fig:eu_advanced_long}). This detailed analysis provides valuable insight, especially for the shorter horizons, as the static KS test only begins to clearly reject the null hypothesis of calibration for $h \geq 6$ (Table~\ref{tab:eu_calibration_summary}).

\subsection{Results for the DRF}
\label{sec:drf_results}
The DRF displays yet another distinct calibration profile. As noted in Section~\ref{sec:forecast_summary}, the DRF's design matrix, which includes longer inflation lags, leads to higher forecasting accuracy at longer horizons. This characteristic generally extends to its calibration, where we see fewer issues as the horizon increases.

For the US, where the DRF benefits from a longer training period, calibration issues are most pronounced at the short horizon. For $h=1$, the cumulative e-value reaches a high level of order $10^{14}$, providing evidence against calibration (Figure~\ref{fig:us_advanced_short}). This finding is supported by low p-values from the static KS tests (essentially zero for $h=1$) and by the clearly inverse U-shaped PIT histograms (Table~\ref{tab:us_calibration_summary}). For longer horizons, the picture becomes more nuanced. The e-values no longer indicate evidence against calibration, and the PIT histograms appear less inverse U-shaped (Figure~\ref{fig:us_advanced_long}). The one-sample KS test, however, continues to suggest miscalibration for $h=3$ (p-value of $0.016$) and only fails to reject the null hypothesis for $h=12$ clearly (p-value of $0.183$).

For the EA, the DRF's one-month-ahead forecasts show early signs of miscalibration that appear to be transient. The PIT histogram for $h=1$ is inversely U-shaped, suggesting overdispersion. The test supermartingale reflects this dynamic. It rises initially to a value of 100, but subsequently drops below one (Figure~\ref{fig:eu_advanced_short}). The process then rises and falls repeatedly, remaining below 1 for the majority of the time. In contrast, the static KS test rejects calibration for $h=1$ with a p-value of $0.004$ (Table~\ref{tab:eu_calibration_summary}).


For longer horizons ($h \geq 3$), the analysis is constrained by the very limited sample size, which makes definitive conclusions difficult. The cumulative e-values for these horizons do not provide any evidence against calibration. While the corresponding PIT histograms appear closer to uniform than for the one-month horizon, the picture from the one-sample KS test remains mixed. It produces p-values of $0.057$ and $0.098$ for horizons $h = 3$ and 6 but does not reject calibration for the longest horizon of $h=12$ with a p-value of $0.154$ (Figure~\ref{fig:eu_advanced_long}, Table~\ref{tab:eu_calibration_summary}).

\section{Conclusion}\label{sec:conclusion}
This paper investigates the calibration of probabilistic inflation forecasts for the United States, the Euro Area, and Switzerland, using a novel framework of sequentially valid inference to evaluate a range of forecasting models. By relying on e-values, we implement a dynamic assessment that reflects the realistic, sequential nature of macroeconomic forecasting and moves beyond traditional fixed-sample diagnostics. This allows us to identify not only whether a model is calibrated, but also when and under what conditions calibration deteriorates.

Our empirical results show that model complexity does not guarantee reliable uncertainty quantification. In fact, simpler models often provided the most trustworthy assessments of forecast uncertainty. The Gaussian no-change benchmark and the univariate AR models, were surprisingly well calibrated across many regions and horizons. By contrast, the Bayesian VAR models were systematically overconfident, with predictive distributions that were too narrow. Sequential testing was crucial in detecting this miscalibration, especially around the post-pandemic inflation surge in the Euro Area and the global financial crisis in the United States. The dynamic factor model performed better than the Bayesian VARs and was often close to the AR benchmark, though sequential diagnostics still revealed horizon-dependent miscalibration in the US forecasts.

Beyond these model-specific findings, the paper demonstrates the practical value of sequentially valid testing for forecast evaluation in macroeconomics. Traditional tools such as the Kolmogorov-Smirnov test assume a fixed evaluation period, an assumption that is often violated in practice when forecasters monitor models continuously and revisit evaluation decisions over time. Such optional stopping undermines classical p-values. E-values address this problem by providing a statistically rigorous way to test for miscalibration at any time, as often as needed, without inflating false positive rates. This has clear implications for model governance. Forecast users, central banks, and financial institutions can use this framework to monitor calibration continuously and detect when a forecasting model becomes unreliable. A sharp and sustained rise in the e-values signals that the model is misspecified and should be re-estimated, re-specified, or replaced. In this way, sequential calibration monitoring can support more robust forecasting systems and, ultimately, better-informed policy and decision-making.










{\singlespacing
\addtocontents{toc}{\vspace{.5\baselineskip}}
\phantomsection
\addcontentsline{toc}{section}{\protect\numberline{}{Bibliography}}
\bibliography{myReferences,references}
}

\newpage

\setcounter{table}{0}
\setcounter{figure}{0}


\begin{center}
{\bf \Large Appendix}
\end{center}

\addtocontents{toc}{\vspace{.5\baselineskip}}