EconBase
← Back to paper

Revision Risk in Real-Time Macroeconomic Forecasting

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

74,177 characters

Revision Risk in Real-Time Macroeconomic Forecasting



\maketitle

\begin{abstract}
Macroeconomic forecasts refer to outcomes that are first released and then
revised. A 90 percent interval for the first GDP release, a six-month value,
or a latest-value benchmark is not the same uncertainty statement. We ask how
revision risk evolves through the release cycle and what can be reported in
real time when later-outcome errors are scarce. We decompose later-outcome MSE
into preliminary forecast risk, revision risk, and their covariance. In SPF
data, first-release to roughly 180-day revisions account for 8.3 percent of
later-outcome MSE across real-activity targets, versus 3.6 percent across
inflation targets. We show that later-outcome uncertainty is partially
identified: released histories give early-error and revision marginals, but
not their dependence. This yields a sharp Fr\'echet--Makarov set and motivates
direct late calibration, dependence-robust transport, and signed or
revision-model transport. Out-of-sample results support method choice rather
than a universal transport rule: coverage and stability determine when
transport gains are usable.

\end{abstract}

\noindent\textbf{Keywords:} revision risk, real-time data, forecast uncertainty,
partial identification.

\section{Introduction}
\label{sec:introduction}

A macroeconomic forecast interval is not a complete uncertainty statement
until it says which vintage of the outcome it is meant to cover. A GDP nowcast
may be evaluated against the advance estimate that moves markets, a value
available several months later, or a latest available value used in a
historical track record. A central-bank or SPF interval may be intended to
cover the first published value of output growth, or a one-year value after
benchmark and source-data revisions have arrived. These are all legitimate
objects, but they are not interchangeable. The same numerical 90 percent
interval has different content depending on the outcome vintage it covers.

This is not only a reporting convention. Macroeconomic measurements mature
over time. Source data arrive with delay, seasonal factors are updated,
benchmark revisions change levels and growth rates, and national accounts are
reconciled after early releases. Advance estimates matter for decisions,
market reactions, and nowcast evaluation. Later vintages matter for policy
reviews, model comparisons, and historical forecast records. A latest-value
track record can be useful, but it answers an ex-post question. A real-time
interval can use only forecast errors whose relevant outcome vintages have
already been released. A forecast interval without a named outcome vintage is
not a complete real-time uncertainty statement.

This paper asks two connected questions. First, how does the risk attached to
an already-issued forecast change as the measured outcome moves through the
release cycle? Second, what uncertainty can be reported in real time when the
desired later-outcome errors have not yet accumulated? The first question
defines the macroeconomic object: release-cycle revision risk. The second is
the information problem created by that object. If an interval is meant to
cover a one-year outcome, then past one-year errors are the direct calibration
history, but those errors arrive later than first-release errors and may be
scarce at the forecast origin.

The empirical facts make the release horizon substantive rather than cosmetic.
In the SPF application, revisions between first release and the value available
roughly 180 days later account for 8.3 percent of later-outcome mean squared
forecast error for real-activity forecasts, compared with 3.6 percent for
inflation forecasts. The covariance term also differs across groups. At 180
days it is positive for SPF real activity and negative for SPF inflation, so
revisions tend to amplify real-activity risk while offsetting inflation risk.
Latest-value benchmarks can be much larger for some real-activity targets, but
they are ex-post historical assessments unless the relevant latest-value
errors had already been released at the original forecast dates. The release
cycle therefore tells us when first-release and later-outcome uncertainty are
close, and when they diverge.

The econometric issue is sharper than a revision adjustment. Suppose released
histories contain first-release forecast errors and revisions from the first
release to a later outcome vintage. Those histories identify the marginal
distribution of early errors and the marginal distribution of revisions. They
do not identify whether large early errors tend to arrive with large revisions,
offsetting revisions, or no stable dependence. That missing dependence is
exactly what is needed to determine the distribution of the later-outcome
forecast error. Later-outcome uncertainty is therefore partially identified
from released histories unless additional dependence restrictions are imposed.

The closest literatures motivate this problem from different angles.
Real-time macro data research shows why later revised series cannot be treated
as the information available to forecasters
\citep{croushoreStark2001,starkCroushore2002,croushore2011,orphanides2001}.
Studies of revisions analyze their information content and predictability
\citep{koenigDolmasPiger2003,faustRogersWright2005,aruoba2008,
jacobsVanNorden2011,corradiFernandezSwanson2009}. Closely related work asks
whether forecasters target first or later national-account releases
\citep{clements2019ForecastersTarget}, and Clements, Galv\~ao, and coauthors
develop revision-aware approaches to real-time macro uncertainty, forecasting,
and predictive distributions \citep{clements2017MacroUncertainty,
clementsGalvao2013JAE,clementsGalvao2019DataRevisionsForecasting,
clementsGalvao2023BVARDataUncertainty,clementsGalvao2017Uncertainty,
carrieroClementsGalvao2015}. Forecast
evaluation and density-forecast evaluation provide the scoring language
\citep{dieboldMariano1995,west1996,giacominiWhite2006,gneitingRaftery2007}.

The distinction here is the object and the information set. Existing work
often asks which release forecasters target or how a revision-aware forecasting
system or predictive density should be built. We instead start from a forecast
that has already been issued and ask three questions: how its risk changes
across outcome vintages, what later-outcome uncertainty is identified by
histories actually released at the forecast origin, and when direct
calibration should be replaced by a transport construction that combines early
errors with revision histories. This contrast is central relative to
Clements--Galv\~ao-style revision-aware forecasting, and it is especially
useful for survey forecasts, institutional nowcasts, and flexible forecast
producers whose predictive densities are unavailable or not comparable across
methods.

We make three primary contributions. First, we introduce the release-cycle
structure of revision risk as a macro empirical object. For each outcome
vintage, later-outcome risk is decomposed into preliminary forecast risk,
revision risk, and the covariance between preliminary errors and revisions.
This decomposition is useful because large revisions alone do not determine
whether later-outcome risk rises or falls. The covariance term shows whether
revisions amplify or offset preliminary forecast errors, and the release-cycle
structure shows when the distinction becomes economically relevant.

Second, we characterize what is identified about later-outcome uncertainty from
released histories. With an early outcome vintage \(v\) and a later vintage
\(w\), the later forecast error combines the early forecast error and the
revision from \(v\) to \(w\). Released histories identify the marginal laws of
these two pieces, but not their joint coupling. Using the classical
Fr\'echet--Makarov bounds for sums with fixed marginals
\citep{makarov1982,frankNelsenSchweizer1987}, we give the sharp
dependence-robust identified set for the later-error distribution. This is a
partial-identification result in the sense of
\citet{manski2003,imbensManski2004,chernozhukovHongTamer2007,tamer2010},
applied to release-indexed forecast errors and real-time admissible histories.
See also \citet{molinari2020} for a survey.

Third, we compare feasible interval constructions for later outcome vintages.
Direct late calibration is natural when a sufficient history of errors for the
intended outcome vintage is available. Dependence-robust transport uses
released early-error and revision histories together with the identified set
when late errors are scarce. Signed and revision-model transport can be
narrower when dependence is stable or revisions are predictable, but they
require stronger restrictions. The resulting error-budget comparison explains
why late-error scarcity alone is not enough to justify transport. Transport is
attractive only when the delay cost of direct late calibration exceeds the
sampling, drift, dependence, and identification costs of using earlier
released information.

Released-error conformal calibration supplies the model-agnostic
implementation. It converts past forecast errors into intervals, which is
useful when forecasts come from surveys, institutional nowcasts, econometric
models, or flexible forecast producers without a common predictive density
\citep{vovkGammermanShafer2005,shaferVovk2008,
chernozhukovWuthrichZhu2018,xuXie2021,barberEtAl2023,
oliveiraOrensteinRamosRomano2024NonExchangeable}. The release-indexed
restriction is essential: a past error can enter the calibration set only
after the outcome vintage relevant for that error has been released. The main
objects remain the release-cycle structure of revision risk, the identified
set for later-outcome uncertainty, and the direct-versus-transport comparison.

The Monte Carlo analysis clarifies the tradeoffs, while the empirical
applications test where the information restrictions bind. In SPF and
national-account applications, the target pairs and interval methods are fixed
before evaluation on an out-of-sample period. SPF shows that transport can
lower normalized interval score, but coverage depends on local stability.
National accounts show the complementary case: when later-outcome histories
are already informative, direct or revision-aware calibration can remain
preferable. The empirical evidence therefore supports method choice under
release-cycle information constraints, not a universal transport rule. The
paper also compares the intervals with standard residual-based alternatives,
including Gaussian bands and a simple revision-adjusted Gaussian analogue.
These comparisons provide historical context, while the out-of-sample
evaluation is the main evidence on prospective performance.

The practical implication is direct. A macro forecast report should state the
outcome vintage being evaluated, the outcome vintage used to form calibration
errors, the revision-risk and covariance diagnostics relevant for that release
horizon, and whether the interval is direct, dependence-robust, or
assumption-dependent. This reporting discipline separates cases where
preliminary and later outcomes are effectively close from cases where revision
risk is part of the forecast uncertainty statement.

The paper proceeds as follows. Section~\ref{sec:revision-risk-term-structure}
develops the release-cycle structure of revision risk.
Section~\ref{sec:identification-inference} gives the released-error
calibration and partial-identification results.
Section~\ref{sec:transport-usefulness} derives the transport error-budget
comparison. Section~\ref{sec:mc-transport} reports Monte Carlo evidence.
Section~\ref{sec:empirical-applications} gives the SPF and national-account
applications and the closest-method benchmark. Section~\ref{sec:conclusion}
concludes.

\section{Release-Indexed Design and Revision Risk}
\label{sec:revision-risk-term-structure}
\label{sec:forecast-risk-accounting}

This section measures when the outcome-release horizon matters economically.
The object is the risk of an already-issued forecast as the realized outcome
moves from first release to later vintages. The decomposition below separates
preliminary forecast risk, revision risk, and the covariance term that decides
whether revisions amplify or offset preliminary errors.

Let \(s\) denote the forecast origin and \(t\) the economic period being
forecast. The information available at the origin is \(\mathcal I_s\), and a
forecast feasible at that origin must be \(\mathcal I_s\)-measurable. An outcome
vintage \(v\) is the released estimate of the realized outcome at a stated
release stage: first release, fixed delay, one-year value, or latest value.
Let \(A_t(v)\) be the date on which \(Y_t(v)\) is released.

Release-indexed forecasting exercises involve four distinct choices. The
information available at the forecast origin determines what predictors and
historical releases the forecaster can use. The outcome vintage used for
training determines which realized values enter estimation or refitting. The
outcome vintage used for evaluation determines the loss and risk object. The
outcome vintage used for calibration determines which past forecast errors can
set uncertainty bands. These choices often coincide in simple textbook
settings, but they need not coincide in real-time macroeconomic data. A model
trained on latest values is an ex-post benchmark unless those values had been
released at the relevant forecast origin, and a final-error calibration rule is
real-time feasible only for forecast origins at which those final errors had
already been released.

For an early vintage \(v\) and a later vintage \(w\), define
\[
  e_t(v)=Y_t(v)-f_s(t),
  \qquad
  \Delta_t(v,w)=Y_t(w)-Y_t(v).
\]
Then \(e_t(w)=e_t(v)+\Delta_t(v,w)\). Unless conditioning is stated
explicitly, risk expectations in this section are unconditional over the
evaluation population of forecast origins and target periods. They are the
population counterparts of sample forecast-evaluation averages.

\begin{proposition}[Outcome-vintage risk decomposition]
\label{prop:outcome-vintage-risk}
Suppose a fixed forecast \(f_s(t)\) is evaluated against outcome vintages
\(v\) and \(w\). Let \(f_s(t)\), \(Y_t(v)\), and \(Y_t(w)\) be square
integrable, with \(f_s(t)\) \(\mathcal I_s\)-measurable. Under squared loss,
write \(R_v(f)=E[(Y_t(v)-f_s(t))^2]\). Then
\[
  R_w(f)-R_v(f)
  =
  -2E\!\left[
    \{f_s(t)-Y_t(v)\}\Delta_t(v,w)
  \right]
  +
  E\!\left[\Delta_t(v,w)^2\right].
  \label{eq:tvpa-risk-decomposition}
\]
For two forecast rules \(f_a\) and \(f_b\), define
\(D_v(a,b)=R_v(f_a)-R_v(f_b)\). Then
\[
  D_w(a,b)-D_v(a,b)
  =
  -2E\!\left[
    \{f_{a,s}(t)-f_{b,s}(t)\}\Delta_t(v,w)
  \right].
  \label{eq:tvpa-pairwise-decomposition}
\]
\end{proposition}

The proposition is a fixed-forecast result. It changes the outcome vintage used
for evaluation while holding the forecasts themselves fixed. It does not say
that every observed change in a forecasting exercise is caused by evaluation
vintage. If the forecasting rule is re-estimated on a different training
outcome vintage, the observed difference also contains a training or refitting
effect.

\begin{corollary}[Ranking reversal condition]
\label{cor:tvpa-ranking-reversal}
Let
\[
  G_{ab}(v,w)=
  E\!\left[
    \{f_{a,s}(t)-f_{b,s}(t)\}\Delta_t(v,w)
  \right].
\]
Then \(D_w(a,b)=D_v(a,b)-2G_{ab}(v,w)\). A pairwise ranking reversal occurs if
\[
  D_v(a,b)\{D_v(a,b)-2G_{ab}(v,w)\}<0.
\]
A sufficient condition for no reversal is
\[
  2|G_{ab}(v,w)|<|D_v(a,b)|.
\]
If \(E[\Delta_t(v,w)\mid\mathcal I_s]=0\) and
\(f_{a,s}(t)-f_{b,s}(t)\) is \(\mathcal I_s\)-measurable, then
\(G_{ab}(v,w)=0\). In that case revisions can raise loss levels but do not
change pairwise rankings through the covariance mechanism in
Proposition~\ref{prop:outcome-vintage-risk}.
\end{corollary}

The same identity gives the release-cycle structure of revision risk. For a
release horizon \(\tau\), write \(Y_t(\tau)\) for the outcome vintage available
at that horizon and \(e_t(\tau)=Y_t(\tau)-f_s(t)\). Relative to the first
release,
\[
R_f(\tau)=E[e_t(\tau)^2]
=E[e_t(0)^2]+E\{\Delta_t(0,\tau)^2\}
+2E[e_t(0)\Delta_t(0,\tau)] .
\]
The second term is revision risk at release horizon \(\tau\). The third term records
whether revisions amplify or offset preliminary forecast errors.

The SPF estimates use observed vintage histories for RGDP,
INDPROD, EMP, CPI, and PCE. Fixed-day horizons are as-of values: the latest released
value available 30, 60, 90, 180, or 365 days after first release. The
latest/final value is an ex-post historical benchmark. The national-account
application uses exact first, second, and third NIPA release stages where
available, the RTDSM QvQd one-year fixed-delay value, and the latest/final
value.

\begin{table}[!htbp]
\centering
\caption{Revision risk through the release cycle with block-bootstrap uncertainty}
\label{tab:revision-risk-term-structure-inference}
\footnotesize
\begin{tabular}{@{}llrrrr@{}}
\toprule
Group & Release horizon & Rev. share & Cov. share & Amp. ratio & Late-error \(n\) \\
\midrule
SPF real activity & 180 days & 8.3 (6.5, 16.7) & 3.4 (-2.6, 7.3) & 1.14 (1.08, 1.25) & 60.0 \\
SPF real activity & latest/final & 29.0 (22.7, 45.2) & -3.1 (-15.2, 3.2) & 1.37 (1.29, 1.55) & 0.0 \\
SPF inflation & 180 days & 3.6 (2.8, 4.6) & -11.0 (-16.7, -5.5) & 0.93 (0.89, 0.98) & 52.9 \\
SPF inflation & latest/final & 16.7 (13.0, 22.9) & -17.9 (-32.0, -9.2) & 1.00 (0.91, 1.07) & 0.0 \\
NA all components & one year & 0.6 (0.2, 3.6) & -1.2 (-7.4, -0.4) & 0.99 (0.96, 1.00) & 26.2 \\
NA all components & latest/final & 1.3 (0.6, 8.5) & -3.0 (-15.1, -1.6) & 0.98 (0.93, 0.99) & 0.0 \\
NA real\_output & one year & 0.7 (0.3, 4.4) & -2.0 (-8.2, 0.7) & 0.99 (0.94, 1.03) & 26.2 \\
NA real\_output & latest/final & 2.6 (1.1, 12.5) & -7.4 (-31.8, -0.9) & 0.95 (0.82, 1.05) & 0.0 \\
NA prices & one year & 0.5 (0.2, 4.3) & -0.8 (-8.4, 0.0) & 1.00 (0.95, 1.02) & 26.2 \\
NA prices & latest/final & 1.1 (0.4, 7.9) & -0.9 (-8.8, 0.9) & 1.00 (0.97, 1.06) & 0.0 \\
\bottomrule
\end{tabular}
\begin{minipage}{0.94\linewidth}
\footnotesize Notes: Parentheses report 95\% moving-block bootstrap intervals. Revision and covariance shares are percentages of later-outcome MSE. The amplification ratio is later-outcome MSE divided by first-release MSE. Late-error \(n\) is the average of \(n_{\ell}(s,w)=|\{u:A_u(w)\le s\}|\) over forecast origins \(s\). It measures the direct-calibration history available in real time, not the risk-decomposition sample size. Latest/final rows use the ex-post latest database value rather than a dated final-release vintage, so their late-error \(n\) is set to zero.
\end{minipage}
\end{table}


Table~\ref{tab:revision-risk-term-structure-inference} summarizes this
release-cycle risk structure. The entries report
revision-risk shares, covariance shares, risk-amplification ratios, and the
released late-error history available at each release horizon. The SPF rows average
target-level shares within real activity and inflation so that no target
drives the comparison by scale.

\begin{figure}[!t]
\centering
\includegraphics[width=0.74\linewidth]{figures/revision_risk_term_structure_spf_bands.pdf}
\caption{SPF revision risk through the release cycle with block-bootstrap bands}
\label{fig:revision-risk-term-structure-spf}
\end{figure}

Figure~\ref{fig:revision-risk-term-structure-spf} shows why selected
first-versus-late comparisons are too narrow. The relevant object is a curve
over the release cycle, with uncertainty bands around the target-averaged
revision-risk shares. The target-averaged real-activity
curve lies above the inflation curve at intermediate release horizons,
especially around 180 days. The covariance channel also differs across the two
groups: revisions tend to amplify real-activity forecast risk at these
horizons, while inflation revisions more often offset preliminary forecast
errors. These are group-level patterns in the SPF sample, not a universal
target-by-target or horizon-by-horizon ordering. The latest/final row in
Table~\ref{tab:revision-risk-term-structure-inference} is an ex-post
historical benchmark. It is useful for describing how much historical
assessments can move, but it should not be treated as a real-time calibration
target unless the relevant latest-value errors had already been released at the
forecast origin.

The value of this release-cycle view is not only to identify large revision-risk cases.
It also identifies cases in which first-release and later-outcome risk are
empirically close, or in which revision variance is offset by covariance with
preliminary forecast errors. For forecast reporting, both outcomes matter:
large release-horizon gaps call for separate later-outcome uncertainty
statements, while small or offsetting gaps justify simpler first-release
reporting for that target and release horizon. Online Appendix Table 2 reports
the complete 180-day target-level table behind this summary.

National accounts provide a complementary application with a richer revision calendar. The national-account
release-cycle pattern is flatter at the one-year horizon than the SPF real-activity
pattern, and covariance terms often offset revision variance. The online
appendix reports the corresponding release-horizon figure with
block-bootstrap bands. The main lesson is not that fixed-forecast ranking
changes are common. It is that release choices, training choices, and
calibration-error choices must be reported separately because national-account
vintages can move through several economically meaningful stages.

The release-cycle evidence also explains the role of later-outcome interval methods.
Direct late-outcome calibration becomes harder as the later-error history
thins out, while early-release errors and revision histories can be observed
sooner. Revision-risk transport is most relevant when the release-cycle evidence
indicates nontrivial later-outcome revision risk and the direct later-error
history is short.

\section{Identification and Real-Time Uncertainty}
\label{sec:identification-inference}

For a later outcome vintage, released early-error and revision histories do not
by themselves determine the forecast-error distribution. At forecast origin
\(s\), a forecaster can use only forecast errors and revision histories that
have already been released. For an
early outcome vintage, this restriction determines which past errors can enter
a real-time calibration sample. For a later outcome vintage, the forecaster
may observe many early-vintage errors and many revisions, but substantially
fewer late-vintage errors. The later-error distribution is not pinned down by
the two released marginal histories alone, because the dependence between
early errors and revisions also matters.

The results below formalize this information restriction. First,
released-error calibration gives uncertainty bands with a time-series coverage
bound that reflects effective calibration size, local drift, and discreteness
of the score distribution. Second, later-vintage uncertainty is partially
identified from released early-error and revision histories unless additional
dependence or revision-model restrictions are imposed. This identification
result explains the interval comparison in the applications: direct late
calibration uses the right errors when enough of them have been released,
dependence-robust transport is valid under marginal information, and signed or
revision-model transport can be narrower under stronger dependence
assumptions. The local-drift terms allow the relevant score law for a current
forecast to differ from the long historical distribution, as in work on local
stationarity and forecast instability
\citep{dahlhaus1997,giacominiRossi2009,rossi2013}.
The conformal step is not the source of the release-cycle object. It is the
device that turns released forecast-error histories into model-agnostic
intervals once the relevant outcome vintage has been named.

\subsection{Released errors and probability statements}
\label{subsec:theory-information}

Using the notation from Section~\ref{sec:revision-risk-term-structure}, all
random variables are defined on \((\Omega,\mathcal F,P)\). The same
information set \(\mathcal I_s\) that determines forecast feasibility is also
the conditioning information for the coverage statements below. For outcome
vintage \(v\) and a past target period \(u\), the released absolute error is
\[
  S_u(v)=|Y_u(v)-f_{s_u}(u)|,
\]
where \(s_u\) is the origin at which the forecast for \(u\) was made. At
origin \(s\), only scores satisfying \(A_u(v)\le s\) may be used. A
calibration rule \(m\) selects past target periods
\(\mathcal U_s^m(v)\subset\{u:A_u(v)\le s,\ u<t\}\) and nonnegative weights
\(\omega_{u,s}^m\). In the result below, this rule is fixed before the
calibration quantile is computed. It may depend on quantities known at origin
\(s\), including release dates, forecast origins, target types, and
pre-specified window lengths. The released error magnitudes \(S_u(v)\) are
also observed at \(s\), but they enter the result through the empirical CDF
below, not through the choice of \(m\).\footnote{If several rules are compared
on the same error history and the best-performing rule is then used to
construct the interval, the same calibration sample is doing two jobs. That
adaptive step requires an additional correction, such as sample splitting or a
multiple-rule adjustment.}

Let \(W_s^m=\sum_{u\in\mathcal U_s^m(v)}\omega_{u,s}^m\) and define the
weighted released-score CDF
\[
  \widehat H_s^m(x;v)=
  \frac{1}{W_s^m}\sum_{u\in\mathcal U_s^m(v)}
  \omega_{u,s}^m\mathbf 1\{S_u(v)\le x\}.
\]
The half-width is the weighted empirical quantile
\[
  q_s^m(v)=\inf\{x:\widehat H_s^m(x;v)\ge 1-\alpha\},
\]
and the released-error interval is
\[
  \Pi_s^m(v,t)=
  [f_s(t)-q_s^m(v),\,f_s(t)+q_s^m(v)].
\]
After conditioning on \(\mathcal I_s\), the selected scores, weights,
quantile, and interval are fixed. The remaining coverage probability is over
the new outcome \(Y_t(v)\). The proposition below is therefore a high-probability
statement over the random released histories that generate \(\mathcal I_s\);
on the good-history event, it gives a conditional coverage bound.

Let
\[
  F_{v,s}(x)=P\{S_t(v)\le x\mid\mathcal I_s\}
\]
be the current conditional score law for the new forecast and outcome vintage
\(v\). Let \(H_s^m(\cdot;v)\) be the population CDF represented by the
released scores selected by rule \(m\), with the same predictable weights. The
local drift term is
\[
  B_s^m(v)=\sup_x |H_s^m(x;v)-F_{v,s}(x)|.
\]

\begin{proposition}[Released-error calibration under dependent histories]
\label{prop:released-error-referee-hardened}
Fix \(s,t,v\) and a predictable released-error calibration rule \(m\).
Assume:
\begin{enumerate}[label=(\roman*),leftmargin=*]
\item \emph{Release admissibility.} Each selected score satisfies
\(A_u(v)\le s\) and \(u<t\), and the selected set is nonempty with
\(W_s^m>0\).
\item \emph{Predictable bounded weights.} The selected index set and weights
are predictable: they do not depend on the selected score magnitudes and are
fixed before those magnitudes are used to form the empirical CDF and quantile,
and
\[
  n_{\mathrm{eff},s}^m(v)=
  \frac{(W_s^m)^2}{\sum_{u\in\mathcal U_s^m(v)}(\omega_{u,s}^m)^2},
  \qquad
  \max_{u}
  \frac{\omega_{u,s}^m}{W_s^m}
  \le
  \frac{\kappa_s}{n_{\mathrm{eff},s}^m(v)} .
\]
\item \emph{Weak dependence.} Ordered by target period, the selected score
indicator array \(\{\mathbf 1(S_u(v)\le x):x\in\mathbb R\}\) is beta-mixing
with coefficients \(\beta_s(k)\), uniformly over thresholds.
\item \emph{Local drift and atoms.}
\[
  \sup_x|H_s^m(x;v)-F_{v,s}(x)|\le B_s^m(v),
  \qquad
  F_{v,s}(q_s^m(v))-F_{v,s}(q_s^m(v)^-)\le \rho_{v,s}.
\]
\end{enumerate}
Let \(n_s^m(v)=|\mathcal U_s^m(v)|\), choose a block length \(b_s\), and set
\[
  N_s(b_s)=
  \left\lfloor\frac{n_{\mathrm{eff},s}^m(v)}{2b_s}\right\rfloor,
  \qquad
  r_s^m(\delta;v)=
  C\kappa_s
  \left[
    \sqrt{
      \frac{\log(n_s^m(v)+1)+\log(4/\delta)}
           {N_s(b_s)}
    }
    +\frac{b_s}{n_{\mathrm{eff},s}^m(v)}
  \right],
\]
where \(C\) is a universal constant. Then, with outer probability at least
\[
  1-\delta-2N_s(b_s)\beta_s(b_s)
\]
over the released calibration history,
\[
  \left|
  P\{Y_t(v)\in\Pi_s^m(v,t)\mid\mathcal I_s\}-(1-\alpha)
  \right|
  \le
  r_s^m(\delta;v)+B_s^m(v)+\rho_{v,s}.
\]
\end{proposition}

The proposition should be read as a high-probability statement about the
released calibration history. Conditional on a realized \(\mathcal I_s\), the
interval and quantile are fixed, and coverage is evaluated under the current
target-score law \(F_{v,s}\). The three terms in the bound have different
interpretations. The term
\(r_s^m(\delta;v)\) is the sampling error from estimating the selected
released-error distribution. It decreases as the effective calibration size
\(n_{\mathrm{eff},s}^m(v)\) increases. Under short-range dependence, one can
choose blocks with \(b_s\to\infty\), \(b_s/n_{\mathrm{eff},s}^m(v)\to0\), and
\(N_s(b_s)\beta_s(b_s)\to0\), so this term converges to zero. Its leading
rate is the usual empirical-CDF rate with an effective sample size penalty,
roughly \(\left(b_s\log n_s^m(v)/n_{\mathrm{eff},s}^m(v)\right)^{1/2}\) up to
constants and logarithmic factors. The term \(B_s^m(v)\) is different. It is local
drift: the distance between the score law represented by the selected past
errors and the current forecast's score law. More historical observations do
not automatically remove this term. It is small only when the selected
released errors are locally representative for the current forecast origin,
which is why the empirical work reports stability and period-sensitivity
diagnostics. The term \(\rho_{v,s}\) is atom or quantile slack at the selected
threshold. It is zero for continuous score distributions with no mass at the
threshold and can be positive with rounded, discrete, or tied forecast errors.
Thus the proposition gives a convergence statement only when the sampling
term vanishes and the drift and atom terms are themselves small.

The proposition is not an exact finite-sample exchangeability result. It is a
time-series calibration bound: effective sample size, weak dependence, local
drift, and atom slack are the costs of using released macroeconomic forecast
errors. Exact conformal validity is recovered only in the same-outcome,
exchangeable special case with the usual conformal quantile convention.
The most common empirical use in this paper is the primitive case in which
the calibration window and weights are fixed before the evaluation outcome is
observed. More adaptive rules require either sample splitting, a finite-class
adjustment, or additional complexity control; they are not covered by the
primitive proposition as stated.

\begin{corollary}[Average coverage implication]
\label{cor:released-error-average}
Let
\[
  R_s^m(v)=r_s^m(\delta;v)+B_s^m(v)+\rho_{v,s}.
\]
Then the unconditional coverage error is bounded by
\[
  \mathbb E\{R_s^m(v)\}
  +\delta+2N_s(b_s)\beta_s(b_s).
\]
In particular, if \(R_s^m(v)\le \bar R_s\) uniformly over the evaluation
design, the unconditional coverage error is at most
\(\bar R_s+\delta+2N_s(b_s)\beta_s(b_s)\).
\end{corollary}

The proposition gives conditional coverage on good released histories. The
corollary then converts this into an average coverage statement by accounting
for the probability of bad history events.

\subsection{Mismatched calibration errors}
\label{subsec:score-mismatch-final}

A score may be released but correspond to the wrong outcome vintage. Let
\(F_{v,s}\) and \(F_{w,s}\) be the current conditional score laws for outcome
vintages \(v\) and \(w\), and define
\[
  M_s(v,w)=\sup_x|F_{v,s}(x)-F_{w,s}(x)|.
\]
If a released-error rule uses scores from outcome vintage \(w\) to form an
interval for \(Y_t(v)\), the bound in
Proposition~\ref{prop:released-error-referee-hardened} holds with \(B_s^m(w)\)
replaced by \(B_s^m(w)+M_s(v,w)\), with atom slack evaluated for the target
score law \(F_{v,s}\) at the selected cutoff. A mismatched score vintage
therefore adds a score-distribution mismatch term even when its own released
score law is estimated accurately.

For absolute scores, the mismatch is controlled by revisions. Since
\[
  |S_t(v)-S_t(w)|\le |Y_t(v)-Y_t(w)|=|\Delta_t(v,w)|,
\]
for any \(r\ge0\),
\[
  M_s(v,w)\le \omega_s(r;w)+\eta_s(r;v,w),
\]
where
\[
  \eta_s(r;v,w)=P(|\Delta_t(v,w)|>r\mid\mathcal I_s)
\]
and
\[
  \omega_s(r;w)=
  \sup_x\max\{F_{w,s}(x+r)-F_{w,s}(x),\,
              F_{w,s}(x)-F_{w,s}(x-r)\}.
\]
Thus calibrating on final or latest errors before those errors were released is
not feasible at the forecast origin, and released but mismatched errors
carry a revision-drift penalty.

\subsection{Later-outcome uncertainty}
\label{subsec:pi-final}

All objects in this subsection are conditional on \(\mathcal I_s\). For an
early outcome vintage \(v\) and later vintage \(w\),
\[
  e_t(w)=e_t(v)+\Delta_t(v,w),
  \qquad
  e_t(v)=Y_t(v)-f_s(t),\quad
  \Delta_t(v,w)=Y_t(w)-Y_t(v).
\]
Released histories may identify or estimate the marginal law \(F_{e,s}\) of
early forecast errors and the marginal law \(F_{\Delta,s}\) of revisions. They
do not in general identify the joint coupling between \(e_t(v)\) and
\(\Delta_t(v,w)\). For an admissible coupling class \(\Gamma_s\), the
identified set for the later-error CDF at \(z\) is
\[
  \mathcal J_s^{\Gamma}(z)=
  \left\{
  P_{\gamma}(e+\Delta\le z):
  \gamma\in\Gamma_s,\
  \gamma_e=F_{e,s},\
  \gamma_{\Delta}=F_{\Delta,s}
  \right\}.
\]
This is a partial-identification object in the sense of
\citet{manski2003}, \citet{imbensManski2004}, and
\citet{chernozhukovHongTamer2007}. See also \citet{molinari2020} for a survey.
The coupling bounds below are classical
Fr\'echet--Makarov bounds for sums with fixed marginals
\citep{makarov1982,frankNelsenSchweizer1987,nelsen2006}. Here they are applied
to later-vintage forecast errors using only histories released by origin \(s\).

Define
\[
  \underline F_s(z)=
  \sup_x \max\{F_{e,s}(x)+F_{\Delta,s}(z-x)-1,0\},
\]
and
\[
  \overline F_s(z)=
  \inf_x \min\{F_{e,s}(x)+F_{\Delta,s}(z-x),1\}.
\]

\begin{proposition}[Sharp identified set under released marginals]
\label{prop:pi-fm-final}
Fix \(s,v,w\). Suppose the released histories identify the marginal laws
\(F_{e,s}\) and \(F_{\Delta,s}\), but impose no restriction on their
dependence. Let \(\Gamma_s^{FM}\) be the set of all couplings with these
marginals. Then, for every \(z\),
\[
  \mathcal J_s^{\Gamma^{FM}}(z)
  =
  [\underline F_s(z),\overline F_s(z)].
\]
The bounds are pointwise sharp: for each \(z\) and each value in this interval
there exists an admissible coupling with the released marginals that attains
that value.
\end{proposition}

The proposition uses a classical coupling result. The object here is the
later-outcome forecast-error distribution that can be identified from
histories actually released at forecast origin \(s\).

\begin{corollary}[Dependence-robust later-outcome interval]
\label{cor:pi-robust-interval-final}
Let \(\ell_s\) and \(r_s\) satisfy
\[
  \overline F_s(\ell_s)\le \alpha/2,\qquad
  \underline F_s(r_s)\ge 1-\alpha/2.
\]
Then \([f_s(t)+\ell_s,\ f_s(t)+r_s]\) covers \(Y_t(w)\) with probability at
least \(1-\alpha\), conditional on \(\mathcal I_s\), uniformly over all
couplings in \(\Gamma_s^{FM}\). Within this equal-tail construction, moving
either endpoint inward is not uniformly justified by the released marginals
alone.
\end{corollary}

\begin{proposition}[Estimated identified sets]
\label{prop:pi-estimated-envelope-final}
Let \(\widehat F_{e,s}\) and \(\widehat F_{\Delta,s}\) be released empirical or
weighted CDFs. Suppose that, on a released-history event with probability at least
\(1-\delta\),
\[
  \sup_x|\widehat F_{e,s}(x)-F_{e,s}(x)|\le \varepsilon_{e,s},
  \qquad
  \sup_x|\widehat F_{\Delta,s}(x)-F_{\Delta,s}(x)|
  \le \varepsilon_{\Delta,s}.
\]
Let \(\widehat{\underline F}_s\) and \(\widehat{\overline F}_s\) be the
Fr\'echet--Makarov envelopes computed from the estimated marginals. Then
\[
  \sup_z|\widehat{\underline F}_s(z)-\underline F_s(z)|
  \le \varepsilon_{e,s}+\varepsilon_{\Delta,s},
  \qquad
  \sup_z|\widehat{\overline F}_s(z)-\overline F_s(z)|
  \le \varepsilon_{e,s}+\varepsilon_{\Delta,s}.
\]
\end{proposition}

The estimated envelope is continuous in the two marginal CDFs: an error
\(\varepsilon_{e,s}\) in the early-error law and an error
\(\varepsilon_{\Delta,s}\) in the revision law enlarge the envelope error by at
most their sum. For released time series, these two terms include sampling
error, local drift, and release-delay effective sample size. When the envelope
is converted into interval endpoints, atoms or flat spots in the relevant CDFs
add the usual endpoint slack. If both released histories grow and the drift
terms are small, the estimated envelope converges to the sharp population
envelope. This convergence does not remove the identification width between
\(\underline F_s\) and \(\overline F_s\); that width comes from unknown
dependence between early errors and revisions.

\begin{proposition}[Point identification under additional structure]
\label{prop:signed-transport-final}
Let \(Z_s(t)\) be \(\mathcal I_s\)-measurable released state information.
Suppose the conditional marginal laws \(F_{e,s}(\cdot\mid Z_s)\) and
\(F_{\Delta,s}(\cdot\mid Z_s)\) are identified from released histories. If
\(e_t(v)\) and \(\Delta_t(v,w)\) are conditionally independent given
\(Z_s(t)\), then the later-error CDF \(G_{w,s}\) conditional on \(Z_s(t)\) is
point identified by the convolution
\[
  G_{w,s}(z\mid Z_s)
  =
  \int F_{e,s}(z-d\mid Z_s)\,dF_{\Delta,s}(d\mid Z_s).
\]
If a revision forecast \(r_s(t)\) is \(\mathcal I_s\)-measurable and
\[
  u_t(v,w)=\Delta_t(v,w)-r_s(t)
\]
has a conditional marginal law identified from released histories, and if
\(u_t(v,w)\) is conditionally independent of \(e_t(v)\) given \(Z_s(t)\), then
the distribution of \(Y_t(w)-f_s(t)-r_s(t)\) conditional on \(Z_s(t)\) is the
convolution of the early-error law and the residual-revision law.
\end{proposition}

Proposition~\ref{prop:signed-transport-final} gives the point-identification
case. Conditional independence, or a revision model that makes the residual
revision independent of the early error after conditioning on \(Z_s(t)\), pins
down a single transported distribution instead of a Fr\'echet--Makarov
identified set. The cost is that validity now rests on the maintained
dependence restriction and, for revision-model transport, on the revision
model. In finite samples the remaining errors are marginal
estimation error, local drift, atom slack, dependence approximation error, and
revision-model error. The empirical diagnostics therefore focus on dependence
between early errors and revisions and on whether the revision model reduces
that dependence.

\begin{corollary}[No sharp nonconservative transport from marginals alone]
\label{cor:pi-no-free-lunch-final}
If the Fr\'echet--Makarov envelope is nondegenerate at a tail probability, an
equal-tail interval obtained by moving at least one dependence-robust endpoint
inward fails to achieve uniform \(1-\alpha\) coverage over \(\Gamma_s^{FM}\)
for some admissible coupling.
\end{corollary}

Corollary~\ref{cor:pi-no-free-lunch-final} is the boundary of what can be
learned from released marginals alone. It does not rule out narrower transport.
Rather, it says that tightening the dependence-robust equal-tail endpoints is
not uniformly valid from the released marginals alone. A forecaster who wants
narrower later-outcome intervals must either impose a dependence restriction,
use and validate a revision model, or report a sensitivity class. This is the
reason the empirical section treats dependence-robust transport as the
validity benchmark and signed or
revision-model transport as efficiency-oriented methods that require
diagnostics.

\subsection{Calibration or Transport}
\label{sec:transport-usefulness}

The preceding results characterize what released histories can identify. This
part turns that result into a practical comparison. Direct late calibration
uses errors for the intended later outcome vintage, but those errors arrive
late and may be few. Transport uses more released early-error and revision
information, but it pays either an identification cost or an assumption cost.

The empirical and simulation sections use five interval procedures. Direct
late-vintage calibration uses released forecast errors for the intended later
outcome vintage. Revision-aware released-error calibration remains a direct
released-error procedure, but updates the calibration history as same-outcome
errors for the relevant vintage become available. Dependence-robust transport
uses the early-error and revision marginal histories together with the
Fr\'echet--Makarov envelope from
Corollary~\ref{cor:pi-robust-interval-final}. Signed convolution transport
uses the signed identity \(e_t(w)=e_t(v)+\Delta_t(v,w)\) and forms the
later-error distribution by convolving early errors with revisions under the
dependence restriction in Proposition~\ref{prop:signed-transport-final}.
Revision-model transport first recenters the interval at a released revision
forecast and then applies the same logic to early errors and revision
residuals. Thus the relevant comparison is an error budget: transport can use
longer released histories, but it pays either an identification cost or an
assumption cost.

Let \(n_{\ell,s}\) be the effective number of released late-error observations,
\(n_{e,s}\) the effective number of released early-error observations, and
\(n_{\Delta,s}\) the effective number of released revisions or revision
residuals. With a common forecast panel and window, late errors are the most
delayed object, so their raw count is no larger than the corresponding
early-error and revision counts. Effective counts can differ after weighting,
blocking, or adding revision histories outside the forecast panel, but the
usual case is \(n_{\ell,s}<n_{e,s}\) and often \(n_{\ell,s}<n_{\Delta,s}\).
The drift notation in this section follows the same template as the object in
Proposition~\ref{prop:released-error-referee-hardened}. There,
\(B_s^m(v)\) compares the distribution represented by selected past scores
with the current score distribution. Here the same comparison is applied to
the three scalar inputs used by the competing procedures: the direct
late-error score, the signed early error, and the signed revision. Let
\(H_{\ell,s}\), \(H_{e,s}\), and \(H_{\Delta,s}\) be the population laws
represented by the released late-error, early-error, and revision histories
used in the comparison. Let \(F_{\ell,s}\), \(F_{e,s}\), and
\(F_{\Delta,s}\) be the corresponding current laws. Then
\[
  B_{\ell,s}=\sup_x|H_{\ell,s}(x)-F_{\ell,s}(x)|,\quad
  B_{e,s}=\sup_x|H_{e,s}(x)-F_{e,s}(x)|,\quad
  B_{\Delta,s}=\sup_x|H_{\Delta,s}(x)-F_{\Delta,s}(x)|.
\]
If revision-model transport uses residual revisions, \(H_{\Delta,s}\) and
\(F_{\Delta,s}\) are read as residual-revision laws. The same mapping applies
to atom and quantile slack. The term \(\rho_{v,s}\) in
Proposition~\ref{prop:released-error-referee-hardened} is the mass or
flat-spot slack at the selected cutoff. Here \(\rho_{\ell,s}\) is the
corresponding slack for the direct late-error cutoff,
\(\rho_{\mathrm{rob},s}\) for the dependence-robust interval endpoints, and
\(\rho_{\mathrm{sgn},s}\) for the signed or revision-model transport
endpoints. With
\(\mathfrak s(n,\delta)=C_s\sqrt{\log(1/\delta)/n}\), a
schematic direct late-calibration error budget is
\[
  \mathcal E_{\mathrm{direct},s}
  =
  \mathfrak s(n_{\ell,s},\delta)+B_{\ell,s}+\rho_{\ell,s}.
\]
Dependence-robust transport has
\[
  \mathcal E_{\mathrm{robust},s}
  =
  \mathfrak s(n_{e,s},\delta)
  +\mathfrak s(n_{\Delta,s},\delta)
  +B_{e,s}+B_{\Delta,s}
  +W_s^{FM}(\alpha)+\rho_{\mathrm{rob},s},
\]
where \(W_s^{FM}(\alpha)\) is the tail or quantile cost induced by the
Fr\'echet--Makarov identified set. Signed or revision-model transport replaces
that identified-set width with dependence and revision-model approximation
costs:
\[
  \mathcal E_{\mathrm{signed},s}
  =
  \mathfrak s(n_{e,s},\delta)
  +\mathfrak s(n_{\Delta,s},\delta)
  +B_{e,s}+B_{\Delta,s}
  +\Pi_s+M_s+\rho_{\mathrm{sgn},s}.
\]
Here \(\Pi_s\) measures conditional-independence or stable-dependence
approximation error, and \(M_s\) is revision-model error when the interval is
centered at \(f_s(t)+r_s(t)\).

The sufficient comparison is simply
\[
  \mathcal E_{m,s}<\mathcal E_{\mathrm{direct},s},
  \qquad m\in\{\mathrm{robust},\mathrm{signed}\}.
\]
When the relevant error CDF has density bounded away from zero near the target
quantiles, the same comparison translates into a smaller worst-case quantile
error up to the inverse-density constant. This condition is deliberately only
an error-budget comparison. It is useful because it exposes the economic and
statistical components of method choice, but it is not a new automatic
selector.

The comparison has a simple interpretation. Direct late calibration is most
appealing when the released late-error history is already informative and
stable. Dependence-robust transport is useful only when late errors are scarce
and the Fr\'echet--Makarov identified set remains reasonably narrow. Signed
or revision-model transport can be more efficient, but only when revision
predictability or stable error--revision dependence makes the extra
identifying restriction credible. When late histories are scarce, robust
transport is wide, and the signed restrictions are weak, the appropriate
conclusion is weak identification rather than a sharper interval. Figure
\ref{fig:transport-method-choice} summarizes this interpretation.

\begin{figure}[!htbp]
\centering
\begin{tikzpicture}[
  node distance=0.45cm and 0.65cm,
  decision/.style={draw, align=center, font=\small, inner sep=3pt,
    text width=0.43\linewidth},
  outcome/.style={draw, align=center, font=\small, inner sep=3pt,
    text width=0.28\linewidth},
  arrow/.style={->, >=stealth}
]
\node[decision] (q1) {Adequate and stable released late-error history?};
\node[outcome, right=of q1] (direct) {Direct late calibration};
\node[decision, below=of q1] (q2) {Late errors scarce and Fr\'echet--Makarov set reasonably narrow?};
\node[outcome, right=of q2] (robust) {Dependence-robust transport};
\node[decision, below=of q2] (q3) {Credible revision predictability or stable error--revision dependence?};
\node[outcome, right=of q3] (signed) {Signed or revision-model transport};
\node[outcome, below=of q3] (weak) {Weak identification};
\draw[arrow] (q1) -- node[above,font=\scriptsize] {yes} (direct);
\draw[arrow] (q1) -- node[left,font=\scriptsize] {no} (q2);
\draw[arrow] (q2) -- node[above,font=\scriptsize] {yes} (robust);
\draw[arrow] (q2) -- node[left,font=\scriptsize] {no} (q3);
\draw[arrow] (q3) -- node[above,font=\scriptsize] {yes} (signed);
\draw[arrow] (q3) -- node[left,font=\scriptsize] {no} (weak);
\end{tikzpicture}
\caption{Error-budget interpretation for choosing between direct calibration and revision-risk transport}
\label{fig:transport-method-choice}
\end{figure}

A feasible empirical index based on these components is reported in the
online appendix as a diagnostic. In the current applications it is
conservative and is not used as an automatic selector. The main text
therefore uses the error-budget comparison qualitatively: it identifies the
forces that favor direct late calibration, dependence-robust transport, or
assumption-dependent transport.

\section{Monte Carlo Evidence}
\label{sec:mc-transport}

The revision-transport theory separates conservative validity from sharper
methods that use additional structure. The Monte Carlo exercises ask when that
structure is useful. They are not designed to show that transport uniformly
improves on direct released-error calibration. They are designed to identify
boundary conditions: late-error scarcity, revision predictability, benchmark
shocks, score-scale drift, and unstable dependence between early errors and
revisions.

\subsection{Design}
\label{subsec:mc-transport-design}

All designs share the same release-indexed structure. A latent macroeconomic
target \(Z_t\) generates an early released outcome \(Y_t(v)\) and a later
outcome \(Y_t(w)\):
\[
  Y_t(v)=Z_t+\eta_t^{v},
  \qquad
  Y_t(w)=Y_t(v)+\Delta_t(v,w),
\]
where the revision process is written as
\[
  \Delta_t(v,w)=\mu_\Delta(X_t)+\sigma_\Delta(R_t)\zeta_t+B_t .
\]
Here \(X_t\) is a real-time covariate, \(R_t\) is a volatility or crisis
state, and \(B_t\) is an occasional benchmark-revision shock. Forecasts are
generated recursively from information available at the forecast origin. The
forecast errors are
\[
  e_t(v)=Y_t(v)-f_t,
  \qquad
  e_t(w)=Y_t(w)-f_t=e_t(v)+\Delta_t(v,w).
\]
Release delays are imposed by allowing \(S_u(w)=|Y_u(w)-f_u|\) to enter direct
late-vintage calibration only after the later outcome has been released. Early
errors and revision histories enter transport rules only after their required
outcome vintages have been released.

The specifications vary only the revision process, the score process, and the
delay length. Stable small revisions use a low-variance revision shock with a
stationary score distribution. Large unpredictable revisions increase revision
variance but keep revisions independent of real-time forecast information.
Predictable revisions set \(\mu_\Delta(X_t)=\beta_r r_t\), where \(r_t\) is a
persistent real-time signal observed at the forecast origin. The coefficient
\(\beta_r\) is large enough that a simple revision model can reduce residual
revision dispersion, while an idiosyncratic revision shock remains.
State-dependent revisions raise revision dispersion and shift the revision
mean during crisis states. Benchmark-shock designs add a persistent shift
\(B_t\) after a fixed date. Delayed-final designs lengthen the release delay
for \(Y_t(w)\), reducing the effective sample for direct late-vintage
calibration. Scale-drift designs change the forecast-error scale over time.
Local-drift designs gradually move the score distribution. The realistic
mixture combines a milder predictable component, crisis-state revision shifts,
extra crisis volatility, a benchmark shift, gradual scale drift, local score
drift, and a nonlinear latent target. It is meant to mimic a setting where
revisions are partly predictable and sometimes state dependent, but no single
feature makes transport automatically dominate released-error calibration.

The main table focuses on four regimes that are most informative for the
transport methods: benchmark shock, delayed final values, predictable
revisions, and realistic mixture. Online Appendix Table 11 reports the full
Monte Carlo grid.

\subsection{Simulation Procedures}
\label{subsec:mc-transport-methods}

The compact table compares five feasible interval procedures in the same
order across regimes: direct late-vintage calibration, revision-aware
released-error calibration, dependence-robust transport, signed convolution
transport, and revision-model transport. Section~\ref{sec:transport-usefulness}
defines these procedures and explains the direct-versus-transport comparison.
Online Appendix Table 11 also reports first-release intervals, absolute bridge
intervals, optimized bridge intervals, and an oracle diagnostic. Those
additional rows are useful diagnostics, but they are not needed for the main
comparison.

\begin{table}[t]
\centering
\scriptsize
\caption{Monte Carlo regimes for revision-risk transport}
\label{tab:revision-transport-mc-compact}
\begin{tabular}{@{}llrrrr@{}}
\toprule
Regime & Method & \(N\) & Cov. & Width & Score \\
\midrule
Benchmark shock & Direct late & 4110 & 0.787 & 2.72 & 4.82 \\
Benchmark shock & Revision aware & 4500 & 0.797 & 2.67 & 4.61 \\
Benchmark shock & Robust transport & 4500 & 0.946 & 3.82 & 4.23 \\
Benchmark shock & Signed transport & 4500 & 0.823 & 2.68 & 4.32 \\
Benchmark shock & Revision-model & 4500 & 0.823 & 2.52 & 4.10 \\
\addlinespace
Delayed final values & Direct late & 3030 & 0.941 & 3.21 & 3.59 \\
Delayed final values & Revision aware & 3600 & 0.902 & 2.84 & 3.49 \\
Delayed final values & Robust transport & 3600 & 0.996 & 4.93 & 4.95 \\
Delayed final values & Signed transport & 3600 & 0.934 & 3.24 & 3.65 \\
Delayed final values & Revision-model & 3600 & 0.923 & 2.89 & 3.43 \\
\addlinespace
Predictable revisions & Direct late & 4290 & 0.913 & 3.20 & 3.84 \\
Predictable revisions & Revision aware & 4500 & 0.900 & 3.06 & 3.81 \\
Predictable revisions & Robust transport & 4500 & 0.998 & 5.71 & 5.72 \\
Predictable revisions & Signed transport & 4500 & 0.954 & 3.74 & 4.04 \\
Predictable revisions & Revision-model & 4500 & 0.930 & 2.87 & 3.32 \\
\addlinespace
Realistic mixture & Direct late & 4110 & 0.913 & 4.45 & 5.54 \\
Realistic mixture & Revision aware & 4500 & 0.907 & 4.31 & 5.45 \\
Realistic mixture & Robust transport & 4500 & 0.984 & 6.39 & 6.58 \\
Realistic mixture & Signed transport & 4500 & 0.920 & 4.44 & 5.41 \\
Realistic mixture & Revision-model & 4500 & 0.912 & 4.25 & 5.26 \\
\bottomrule
\end{tabular}
\begin{minipage}{0.96\textwidth}
\footnotesize
Notes: Nominal 90 percent intervals for first-release to final-outcome
transport. \(N\) is the number of evaluated forecast origins. The compact table
uses the same five methods in each regime. The full Monte Carlo grid,
including first-release, absolute bridge, optimized bridge, and oracle
diagnostic rows, appears in Online Appendix Table 15.
\end{minipage}
\end{table}


The benchmark-shock rows show the tail-risk failure mode: direct,
revision-aware, signed, and revision-model intervals are too narrow, while
robust transport restores coverage at the cost of width. The delayed-final
and predictable-revision rows show where transport can pay off, especially
when a revision model reduces the dispersion of later-outcome errors. In the
realistic mixture, the simple revision-aware benchmark is already strong, so
revision-model transport yields only a modest score improvement. These
patterns motivate the diagnostic interpretation used in the applications.

\subsection{Results}
\label{subsec:mc-transport-results}

Table~\ref{tab:revision-transport-mc-compact} reports nominal 90 percent
intervals. The benchmark-shock design shows the main failure mode for signed
transport. Direct late calibration and revision-aware calibration do not cover
because released histories do not represent the new benchmark tail. Signed and
revision-model transport are also too narrow. Robust transport is wider, but
it restores coverage by protecting against dependence and tail behavior that
the signed methods miss.

The delayed-final and predictable-revision designs show when transport can
help. In delayed-final designs, direct late calibration has fewer usable
origins, and revision-model transport improves the interval score relative to
direct late calibration while maintaining above-nominal coverage. In
predictable-revision designs, the revision-model method has the strongest
score--coverage tradeoff among the feasible near-nominal methods shown: it
uses revision predictability to reduce width and improve interval score.
Robust transport covers in both regimes, but its envelope is too conservative
when the signed structure is benign.

The realistic mixture combines mild predictability, state variation, and
occasional revision shocks. In this regime, the simple revision-aware benchmark
is already strong. Revision-model transport still has the lowest score among
the feasible methods shown, while signed transport gives only a modest gain.
This is the pattern to expect in empirical macro applications: transport is
one possible response when late errors are scarce and revisions are stable or
predictable, but simple released-error benchmarks can be hard to beat when
calibration histories are already informative.

The Monte Carlo evidence therefore supports a diagnostic use of revision
transport. Robust transport is the validity-oriented construction when
dependence or tail behavior is uncertain. Signed and revision-model transport
are efficiency-oriented constructions for regimes with stable dependence or
predictable revisions. The empirical sections below use the same logic:
dependence diagnostics and common-origin comparisons determine whether
transport is interpreted as a safe interval, an efficiency improvement, or a
method that should be treated cautiously.

\section{Empirical Applications: SPF and National Accounts}
\label{sec:empirical-applications}

The empirical applications ask when the information restrictions characterized
in Section~\ref{sec:identification-inference}, especially the error-budget
comparison in Section~\ref{sec:transport-usefulness}, favor transport, and
when direct or revision-aware calibration remains preferable. SPF and national accounts
provide complementary release-cycle settings. The out-of-sample evaluation
period is the primary prospective evidence, while target-level
results, period decompositions, and common-origin benchmarks explain why
performance differs across variables and time.

\subsection{Design and scoring}
\label{subsec:empirical-design}

The SPF application uses Federal Reserve Bank of Philadelphia mean and median
forecasts for CPI, EMP, INDPROD, PCE, and RGDP. We use the current-period
forecast and the one-period-ahead forecast, denoted SPF horizons 0 and 1. The
outcome pair is first release to the value available about 180 days after
first release. Forecast origins run from 1996-03-02 to 2025-11-11, with
1996-03-02 to 2018-05-08 used as the pre-evaluation sample and 2018-08-07
to 2025-11-11 reserved for out-of-sample evaluation. For CPI, EMP, INDPROD,
and PCE, we match forecasts to real-time vintages of the corresponding source
series: CPIAUCSL for CPI, PAYEMS for EMP, INDPRO for INDPROD, and PCEPI for
PCE. RGDP is matched to Philadelphia Fed RTDSM real-output values. The online
appendix documents the source-series mapping, release-date construction, and
the RGDP release-date caveat.

The national-account application uses real-time RTDSM values with day-level
BEA release dates where the release calendar can be matched. The outcome pair
is first release to a one-year QvQd fixed-delay value. Eligible components
are P, RCON, REX, RG, RIMP, RINVBF, RINVRESID, ROUTPUT, and YRGDI. Forecast
rules are fixed within each comparison and consist of AR-lag, pooled-ridge
component, and simple forecast-combination rules. Forecast origins run from
2016-04-27 to 2025-04-29, with 2016-04-27 to 2023-01-25 used as the
pre-evaluation sample and 2023-04-26 to 2025-04-29 reserved for out-of-sample evaluation.
Online Appendix Table 3 reports the compact design table with target cells,
forecast-origin counts, and scoring definitions.

The methods are direct later-outcome calibration, revision-aware
released-error calibration, dependence-robust transport, signed transport,
and revision-model transport. A forecast-origin observation enters a comparison
only when every method is feasible for the same forecast origin and target
cell. The nominal interval level is 90 percent. The primary score is the
interval score divided by a pre-evaluation RMSFE scale for the
relevant target cell. That scale is fixed using the pre-evaluation sample only.
The normalized score is a comparative coverage--width measure, not
a welfare magnitude: within a target cell, a lower value means a better
coverage--width tradeoff relative to the forecast difficulty observed in the
pre-evaluation sample. SPF summaries give equal weight to each
target--horizon--forecast-type cell, and national-account summaries give
equal weight to each target--model cell.

\subsection{Out-of-Sample Evaluation}
\label{subsec:transport-holdout}

The out-of-sample evaluation is the primary evidence about prospective
performance. It uses the last calendar-time block of each sample after the
target pairs and interval methods have been fixed.
Table~\ref{tab:normalized-holdout-primary} reports these results. In SPF,
dependence-robust transport lowers normalized
interval score relative to direct late calibration, from 22.92 to 22.04. It
also raises coverage from 80.0 to 84.5 percent. Yet it still undercovers
relative to the nominal 90 percent level. Revision-aware
calibration gives a smaller score improvement and similar coverage to direct
late calibration. Signed and revision-model transport are narrower, but do
not resolve the coverage shortfall. The SPF evaluation period therefore shows
the central score--coverage tradeoff: early-error and revision histories can
improve score, but the improvement is usable only when coverage remains close
to the nominal target.

\begin{table}[t]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\caption{Out-of-sample interval performance}
\label{tab:normalized-holdout-primary}
\begin{tabular}{@{}llrrrrr@{}}
\toprule
Application & Method & Cells & Cov. & Norm. score & $\Delta$Score & Norm. width \\
\midrule
SPF & Direct late & 20 & 80.0\% & 22.92 & 0.00 [0.00,0.00] & 4.19 \\
 & Revision aware & 20 & 80.9\% & 22.72 & -0.20 [-0.34,-0.01] & 4.23 \\
 & Signed transport & 20 & 79.3\% & 23.12 & 0.20 [-0.04,0.42] & 4.15 \\
 & Robust transport & 20 & 84.5\% & 22.04 & -0.88 [-2.27,0.59] & 5.41 \\
 & Revision-model & 20 & 79.3\% & 23.11 & 0.19 [-0.05,0.40] & 4.16 \\
\addlinespace
National accounts & Direct late & 27 & 97.9\% & 2.16 & 0.00 [0.00,0.00] & 1.85 \\
 & Revision aware & 27 & 98.4\% & 2.09 & -0.08 [-0.15,0.01] & 1.78 \\
 & Signed transport & 27 & 98.4\% & 2.26 & 0.09 [-0.04,0.24] & 2.05 \\
 & Robust transport & 27 & 99.2\% & 2.99 & 0.83 [0.56,1.15] & 2.94 \\
 & Revision-model & 27 & 98.4\% & 2.25 & 0.09 [-0.05,0.24] & 2.05 \\
\addlinespace
\bottomrule
\end{tabular}
\begin{minipage}{0.96\textwidth}
\footnotesize Notes: Scores and widths are normalized by pre-evaluation cell RMSFE. The primary aggregation gives equal weight to each target--horizon--forecast-type cell for SPF and each target--model cell for national accounts. $\Delta$Score is method minus direct late calibration, so negative values favor the method. Parentheses report moving-block bootstrap 95\% intervals.
\end{minipage}
\end{table}


The national-account evaluation period gives a different lesson.
Direct late calibration and revision-aware released-error calibration already
cover well, and revision-aware calibration has the lowest normalized score.
Dependence-robust transport is safer but much wider, while signed and
revision-model transport cover well without improving on direct or
revision-aware calibration. Here the later-error histories are informative
enough that transport's identification and dependence costs outweigh its use
of revision histories.

\subsection{Target-Level Heterogeneity}
\label{subsec:holdout-target-heterogeneity}

The target-level figure explains why pooled results are not sufficient for
method choice. This heterogeneity is the empirical reason for treating revision
risk as target- and release-horizon-specific rather than reporting a single
revision correction.
Figure~\ref{fig:holdout-coverage-score-by-target} plots every pre-specified
target by later-period coverage and normalized score gain relative to direct late
calibration. In SPF, robust transport improves score for CPI, PCE, EMP, and
RGDP, but those points have coverage below the nominal level. INDPROD is the most favorable SPF
transport case because revision-model transport improves score modestly while
keeping coverage close to 90 percent. Thus the aggregate SPF score gain is
partly driven by cells where coverage remains binding.

\begin{figure}[!htbp]
\centering
\includegraphics[width=0.98\linewidth]{figures/holdout_coverage_score_by_target.pdf}
\caption{Target-level coverage and score tradeoffs in the out-of-sample evaluation}
\label{fig:holdout-coverage-score-by-target}
\begin{minipage}{0.94\linewidth}
\footnotesize Notes: Each non-red point is a pre-specified target and
non-direct method. Red circles mark direct late calibration coverage at zero
score gain. Text labels mark the target-method points discussed in the main
text. The horizontal dashed line is nominal 90 percent coverage, the dotted line
is a visual reference below nominal coverage, and the vertical line marks no
score gain relative to direct late calibration.
\end{minipage}
\end{figure}

A compact reading of the figure is that transport has useful cases, and their
value is target-specific. For SPF, INDPROD is the cleanest transport example,
while CPI and PCE illustrate score gains that remain constrained by coverage.
For national accounts, RIMP and YRGDI are the most favorable transport cases,
but many components are better described as cases where direct or
revision-aware calibration is already informative.

National-account heterogeneity is also substantial. RIMP is the clearest
dependence-robust transport case, and YRGDI is the clearest signed or
revision-model transport case. For many other components, revision-aware
released-error calibration is the strongest non-direct method, while robust
transport mainly adds width. Online Appendix Table 4 gives the compact
target-level summary, and Online Appendix Table 7 reports the target-level
out-of-sample detail for every pre-specified target, including cases where transport
does not improve on direct calibration.
Online Appendix Section F.2 reports the corresponding revision-transport
dependence diagnostics, including early-error--revision correlations, residual
correlations, and tail-dependence measures. These diagnostics are used to
interpret signed and revision-model transport, not to turn them into
assumption-free procedures.

The target-level patterns should also be read together with the period
decomposition in Online Appendix Table 6. In the SPF later evaluation sample,
pre-COVID coverage is near or above nominal for all methods, and post-COVID
coverage is again close to nominal. The COVID block is the exception: direct
late calibration covers 46.3 percent, revision-aware calibration covers
47.5 percent, and robust transport covers 57.5 percent. Robust transport
still reduces raw interval score in that block, but no method provides
adequate coverage. The SPF deterioration in the later evaluation period is
therefore primarily a local scale and tail episode, not a simple shortage of
late errors. In the same evaluation period, average SPF late-error histories
are larger than in the full common-origin benchmark \((99\) versus \(71)\),
and average revision histories are also larger. National-account late-error
histories similarly rise from about \(38\) to \(52\). Later-error scarcity is
therefore not mechanically worse in the later evaluation period. The issue is
whether the released histories are stable and informative for the current
release horizon. Modeling this period dependence is an important direction for
future work.

\subsection{Common-Origin Benchmark}
\label{subsec:closest-method-benchmark}

The closest-method benchmark has a different role. It asks whether the
paper's methods remain competitive with simple parametric residual bands on the
full common-origin historical sample. Table
\ref{tab:closest-method-benchmark-compact} compares direct late
released-error calibration, revision-aware released-error calibration,
dependence-robust transport, and revision-model transport with a Gaussian
same-outcome residual band and a revision-adjusted Gaussian band. The
revision-adjusted Gaussian band is a simplified analogue motivated by
\citet{clements2017MacroUncertainty} and
\citet{clementsGalvao2013JAE,clementsGalvao2023BVARDataUncertainty}.\footnote{A full
replication would change the forecast producer and density model. The benchmark
instead holds the issued point forecast fixed and compares interval
construction.} Secondary parametric benchmarks, including Student-\(t\)
residual bands, are reported in the online appendix.

\begin{table}[!htbp]
\centering
\small
\caption{Closest-method benchmark on common origins}
\label{tab:closest-method-benchmark-compact}
\begin{tabularx}{\textwidth}{@{}llXrrr@{}}
\toprule
Application & Method & Status & Cov. & Score & Width \\
\midrule
SPF & Direct late released-error & Paper method & 0.866 & 12.30 & 4.18 \\
SPF & Revision-aware released-error & Paper method & 0.868 & 12.19 & 4.17 \\
SPF & Dependence-robust transport & Paper method & 0.901 & 11.96 & 4.88 \\
SPF & Revision-model transport & Paper method & 0.845 & 12.32 & 3.93 \\
SPF & Gaussian same-outcome band & Generic benchmark & 0.886 & 13.03 & 5.81 \\
SPF & Revision-adjusted Gaussian & Literature & 0.886 & 13.04 & 5.82 \\
\addlinespace
National accounts & Direct late released-error & Paper method & 0.865 & 4.81 & 1.67 \\
National accounts & Revision-aware released-error & Paper method & 0.887 & 4.63 & 1.67 \\
National accounts & Dependence-robust transport & Paper method & 0.933 & 4.94 & 2.67 \\
National accounts & Revision-model transport & Paper method & 0.902 & 4.62 & 1.87 \\
National accounts & Gaussian same-outcome band & Generic benchmark & 0.888 & 4.75 & 1.96 \\
National accounts & Revision-adjusted Gaussian & Literature & 0.888 & 4.75 & 1.96 \\
\bottomrule
\end{tabularx}
\begin{minipage}{0.94\linewidth}
\footnotesize Notes: The table uses common forecast origins, pre-evaluation
RMSFE normalization, and equal target-cell weighting. Score is the normalized
interval score. Width is normalized average width. Student-\(t\) and other
secondary parametric bands are reported in the online appendix.
\end{minipage}
\end{table}


In the full common-origin benchmark, robust transport has the lowest
normalized SPF score among the compact methods and near-nominal coverage. In
national accounts, revision-model transport and revision-aware released-error
calibration have the lowest scores, while Gaussian bands are competitive but
not dominant.

Online Appendix Table 8 reconciles the historical benchmark with the
out-of-sample evidence. The common-origin benchmark describes rolling
performance under the full historical mix. The out-of-sample period is
the stronger prospective evaluation. In SPF, robust transport looks stronger
in the common-origin benchmark than in the later evaluation period because that
period contains the COVID scale shift. In national accounts, the later-period
late-error histories are already informative enough that direct and
revision-aware calibration are strong, so robust transport mainly adds width.
The difference between the two panels is therefore evidence about stability
and method choice, not a contradiction.
This reconciliation is central to the empirical interpretation: common-origin
performance describes historical fit under a broad mix of periods, while the
later evaluation period tests prospective stability.

\subsection{Empirical implications}
\label{subsec:empirical-lessons}

The results imply four practical lessons for method choice.

\begin{enumerate}[label=\arabic*.,leftmargin=*]
\item Naming the later outcome is necessary because revision risk differs by
target and release horizon. The release-cycle structure shows when first-release, 180-day,
one-year, and latest-value uncertainty statements diverge.

\item Later-error scarcity is only one input into method choice. SPF shows
that transport can lower normalized score, but coverage depends on stability
of recent errors and revisions.

\item Dependence-robust transport and signed transport answer different
questions. Robust transport trades width for protection against unknown
dependence. Signed and revision-model transport trade robustness for
efficiency when dependence and revision-model diagnostics are favorable.

\item Direct or revision-aware calibration remains the appropriate benchmark
when later-error histories are informative. The national-account later evaluation
period illustrates this case: once the intended later-outcome errors are
available in sufficient quantity and stability, simpler released-error methods
can beat transport on normalized score while maintaining coverage.
\end{enumerate}

\section{Conclusion}
\label{sec:conclusion}

Macroeconomic forecast risk changes as measured outcomes move through the
release cycle. A first release, an intermediate vintage, and a latest-value
benchmark can answer different uncertainty questions for the same forecast.
The empirical evidence shows that revision risk is not uniformly large, but it
is measurable, target-specific, and sometimes amplified by its covariance with
preliminary forecast errors. This makes the outcome vintage part of the
forecast-risk statement rather than a detail to be supplied after evaluation.

For uncertainty quantification, later-outcome risk is only partly identified
from histories available at the forecast origin unless additional dependence
restrictions are imposed. Released early-error histories and revision histories
identify marginal laws, but not the dependence between early errors and
revisions. Fr\'echet--Makarov bounds give the dependence-robust identified
set. Signed and revision-model transport can be narrower, but their
interpretation depends on stable dependence or revision predictability.
Released-error calibration provides a feasible model-agnostic way to turn
these objects into intervals.

The practical implication is a reporting discipline for real-time forecast
uncertainty. A forecast report should name the outcome vintage being covered,
state which outcome vintage is used to form calibration errors, report the
revision-risk and covariance diagnostics relevant for that release horizon,
and identify whether the interval is direct, dependence-robust, or
assumption-dependent. When later-error histories are sufficiently informative,
direct or revision-aware calibration can be preferable. When later errors are
scarce and the dependence diagnostics are favorable, revision-risk transport
can use released early-error and revision histories to make later-outcome
uncertainty statements feasible.


\clearpage