EconBase
← Back to paper

Revision Risk in Real-Time Macroeconomic Forecasting

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

74,177 characters · 18 sections · 17 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Revision Risk in Real-Time Macroeconomic Forecasting

abstractMacroeconomic forecasts refer to outcomes that are first released and then revised. A 90 percent interval for the first GDP release, a six-month value, or a latest-value benchmark is not the same uncertainty statement. We ask how revision risk evolves through the release cycle and what can be reported in real time when later-outcome errors are scarce. We decompose later-outcome MSE into preliminary forecast risk, revision risk, and their covariance. In SPF data, first-release to roughly 180-day revisions account for 8.3 percent of later-outcome MSE across real-activity targets, versus 3.6 percent across inflation targets. We show that later-outcome uncertainty is partially identified: released histories give early-error and revision marginals, but not their dependence. This yields a sharp Fr\'echet--Makarov set and motivates direct late calibration, dependence-robust transport, and signed or revision-model transport. Out-of-sample results support method choice rather than a universal transport rule: coverage and stability determine when transport gains are usable.

\noindentKeywords: revision risk, real-time data, forecast uncertainty, partial identification.

Introduction

A macroeconomic forecast interval is not a complete uncertainty statement until it says which vintage of the outcome it is meant to cover. A GDP nowcast may be evaluated against the advance estimate that moves markets, a value available several months later, or a latest available value used in a historical track record. A central-bank or SPF interval may be intended to cover the first published value of output growth, or a one-year value after benchmark and source-data revisions have arrived. These are all legitimate objects, but they are not interchangeable. The same numerical 90 percent interval has different content depending on the outcome vintage it covers.

This is not only a reporting convention. Macroeconomic measurements mature over time. Source data arrive with delay, seasonal factors are updated, benchmark revisions change levels and growth rates, and national accounts are reconciled after early releases. Advance estimates matter for decisions, market reactions, and nowcast evaluation. Later vintages matter for policy reviews, model comparisons, and historical forecast records. A latest-value track record can be useful, but it answers an ex-post question. A real-time interval can use only forecast errors whose relevant outcome vintages have already been released. A forecast interval without a named outcome vintage is not a complete real-time uncertainty statement.

This paper asks two connected questions. First, how does the risk attached to an already-issued forecast change as the measured outcome moves through the release cycle? Second, what uncertainty can be reported in real time when the desired later-outcome errors have not yet accumulated? The first question defines the macroeconomic object: release-cycle revision risk. The second is the information problem created by that object. If an interval is meant to cover a one-year outcome, then past one-year errors are the direct calibration history, but those errors arrive later than first-release errors and may be scarce at the forecast origin.

The empirical facts make the release horizon substantive rather than cosmetic. In the SPF application, revisions between first release and the value available roughly 180 days later account for 8.3 percent of later-outcome mean squared forecast error for real-activity forecasts, compared with 3.6 percent for inflation forecasts. The covariance term also differs across groups. At 180 days it is positive for SPF real activity and negative for SPF inflation, so revisions tend to amplify real-activity risk while offsetting inflation risk. Latest-value benchmarks can be much larger for some real-activity targets, but they are ex-post historical assessments unless the relevant latest-value errors had already been released at the original forecast dates. The release cycle therefore tells us when first-release and later-outcome uncertainty are close, and when they diverge.

The econometric issue is sharper than a revision adjustment. Suppose released histories contain first-release forecast errors and revisions from the first release to a later outcome vintage. Those histories identify the marginal distribution of early errors and the marginal distribution of revisions. They do not identify whether large early errors tend to arrive with large revisions, offsetting revisions, or no stable dependence. That missing dependence is exactly what is needed to determine the distribution of the later-outcome forecast error. Later-outcome uncertainty is therefore partially identified from released histories unless additional dependence restrictions are imposed.

The closest literatures motivate this problem from different angles. Real-time macro data research shows why later revised series cannot be treated as the information available to forecasters croushoreStark2001,starkCroushore2002,croushore2011,orphanides2001. Studies of revisions analyze their information content and predictability koenigDolmasPiger2003,faustRogersWright2005,aruoba2008, jacobsVanNorden2011,corradiFernandezSwanson2009. Closely related work asks whether forecasters target first or later national-account releases clements2019ForecastersTarget, and Clements, Galv\ ao, and coauthors develop revision-aware approaches to real-time macro uncertainty, forecasting, and predictive distributions clements2017MacroUncertainty, clementsGalvao2013JAE,clementsGalvao2019DataRevisionsForecasting, clementsGalvao2023BVARDataUncertainty,clementsGalvao2017Uncertainty, carrieroClementsGalvao2015. Forecast evaluation and density-forecast evaluation provide the scoring language dieboldMariano1995,west1996,giacominiWhite2006,gneitingRaftery2007.

The distinction here is the object and the information set. Existing work often asks which release forecasters target or how a revision-aware forecasting system or predictive density should be built. We instead start from a forecast that has already been issued and ask three questions: how its risk changes across outcome vintages, what later-outcome uncertainty is identified by histories actually released at the forecast origin, and when direct calibration should be replaced by a transport construction that combines early errors with revision histories. This contrast is central relative to Clements--Galv\ ao-style revision-aware forecasting, and it is especially useful for survey forecasts, institutional nowcasts, and flexible forecast producers whose predictive densities are unavailable or not comparable across methods.

We make three primary contributions. First, we introduce the release-cycle structure of revision risk as a macro empirical object. For each outcome vintage, later-outcome risk is decomposed into preliminary forecast risk, revision risk, and the covariance between preliminary errors and revisions. This decomposition is useful because large revisions alone do not determine whether later-outcome risk rises or falls. The covariance term shows whether revisions amplify or offset preliminary forecast errors, and the release-cycle structure shows when the distinction becomes economically relevant.

Second, we characterize what is identified about later-outcome uncertainty from released histories. With an early outcome vintage \(v\) and a later vintage \(w\), the later forecast error combines the early forecast error and the revision from \(v\) to \(w\). Released histories identify the marginal laws of these two pieces, but not their joint coupling. Using the classical Fr\'echet--Makarov bounds for sums with fixed marginals makarov1982,frankNelsenSchweizer1987, we give the sharp dependence-robust identified set for the later-error distribution. This is a partial-identification result in the sense of manski2003,imbensManski2004,chernozhukovHongTamer2007,tamer2010, applied to release-indexed forecast errors and real-time admissible histories. See also molinari2020 for a survey.

Third, we compare feasible interval constructions for later outcome vintages. Direct late calibration is natural when a sufficient history of errors for the intended outcome vintage is available. Dependence-robust transport uses released early-error and revision histories together with the identified set when late errors are scarce. Signed and revision-model transport can be narrower when dependence is stable or revisions are predictable, but they require stronger restrictions. The resulting error-budget comparison explains why late-error scarcity alone is not enough to justify transport. Transport is attractive only when the delay cost of direct late calibration exceeds the sampling, drift, dependence, and identification costs of using earlier released information.

Released-error conformal calibration supplies the model-agnostic implementation. It converts past forecast errors into intervals, which is useful when forecasts come from surveys, institutional nowcasts, econometric models, or flexible forecast producers without a common predictive density vovkGammermanShafer2005,shaferVovk2008, chernozhukovWuthrichZhu2018,xuXie2021,barberEtAl2023, oliveiraOrensteinRamosRomano2024NonExchangeable. The release-indexed restriction is essential: a past error can enter the calibration set only after the outcome vintage relevant for that error has been released. The main objects remain the release-cycle structure of revision risk, the identified set for later-outcome uncertainty, and the direct-versus-transport comparison.

The Monte Carlo analysis clarifies the tradeoffs, while the empirical applications test where the information restrictions bind. In SPF and national-account applications, the target pairs and interval methods are fixed before evaluation on an out-of-sample period. SPF shows that transport can lower normalized interval score, but coverage depends on local stability. National accounts show the complementary case: when later-outcome histories are already informative, direct or revision-aware calibration can remain preferable. The empirical evidence therefore supports method choice under release-cycle information constraints, not a universal transport rule. The paper also compares the intervals with standard residual-based alternatives, including Gaussian bands and a simple revision-adjusted Gaussian analogue. These comparisons provide historical context, while the out-of-sample evaluation is the main evidence on prospective performance.

The practical implication is direct. A macro forecast report should state the outcome vintage being evaluated, the outcome vintage used to form calibration errors, the revision-risk and covariance diagnostics relevant for that release horizon, and whether the interval is direct, dependence-robust, or assumption-dependent. This reporting discipline separates cases where preliminary and later outcomes are effectively close from cases where revision risk is part of the forecast uncertainty statement.

The paper proceeds as follows. Section (ref) develops the release-cycle structure of revision risk. Section (ref) gives the released-error calibration and partial-identification results. Section (ref) derives the transport error-budget comparison. Section (ref) reports Monte Carlo evidence. Section (ref) gives the SPF and national-account applications and the closest-method benchmark. Section (ref) concludes.

Release-Indexed Design and Revision Risk

This section measures when the outcome-release horizon matters economically. The object is the risk of an already-issued forecast as the realized outcome moves from first release to later vintages. The decomposition below separates preliminary forecast risk, revision risk, and the covariance term that decides whether revisions amplify or offset preliminary errors.

Let \(s\) denote the forecast origin and \(t\) the economic period being forecast. The information available at the origin is \(\mathcal I_s\), and a forecast feasible at that origin must be \(\mathcal I_s\)-measurable. An outcome vintage \(v\) is the released estimate of the realized outcome at a stated release stage: first release, fixed delay, one-year value, or latest value. Let \(A_t(v)\) be the date on which \(Y_t(v)\) is released.

Release-indexed forecasting exercises involve four distinct choices. The information available at the forecast origin determines what predictors and historical releases the forecaster can use. The outcome vintage used for training determines which realized values enter estimation or refitting. The outcome vintage used for evaluation determines the loss and risk object. The outcome vintage used for calibration determines which past forecast errors can set uncertainty bands. These choices often coincide in simple textbook settings, but they need not coincide in real-time macroeconomic data. A model trained on latest values is an ex-post benchmark unless those values had been released at the relevant forecast origin, and a final-error calibration rule is real-time feasible only for forecast origins at which those final errors had already been released.

For an early vintage \(v\) and a later vintage \(w\), define \[ e_t(v)=Y_t(v)-f_s(t), \qquad \Delta_t(v,w)=Y_t(w)-Y_t(v). \] Then \(e_t(w)=e_t(v)+\Delta_t(v,w)\). Unless conditioning is stated explicitly, risk expectations in this section are unconditional over the evaluation population of forecast origins and target periods. They are the population counterparts of sample forecast-evaluation averages.

proposition[Outcome-vintage risk decomposition] Suppose a fixed forecast \(f_s(t)\) is evaluated against outcome vintages \(v\) and \(w\). Let \(f_s(t)\), \(Y_t(v)\), and \(Y_t(w)\) be square integrable, with \(f_s(t)\) \(\mathcal I_s\)-measurable. Under squared loss, write \(R_v(f)=E[(Y_t(v)-f_s(t))^2]\). Then \[ R_w(f)-R_v(f) = -2E\!\left[ \{f_s(t)-Y_t(v)\}\Delta_t(v,w) \right] + E\!\left[\Delta_t(v,w)^2\right]. \label{eq:tvpa-risk-decomposition} \] For two forecast rules \(f_a\) and \(f_b\), define \(D_v(a,b)=R_v(f_a)-R_v(f_b)\). Then \[ D_w(a,b)-D_v(a,b) = -2E\!\left[ \{f_{a,s}(t)-f_{b,s}(t)\}\Delta_t(v,w) \right]. \label{eq:tvpa-pairwise-decomposition} \]

The proposition is a fixed-forecast result. It changes the outcome vintage used for evaluation while holding the forecasts themselves fixed. It does not say that every observed change in a forecasting exercise is caused by evaluation vintage. If the forecasting rule is re-estimated on a different training outcome vintage, the observed difference also contains a training or refitting effect.

corollary[Ranking reversal condition] Let \[ G_{ab}(v,w)= E\!\left[ \{f_{a,s}(t)-f_{b,s}(t)\}\Delta_t(v,w) \right]. \] Then \(D_w(a,b)=D_v(a,b)-2G_{ab}(v,w)\). A pairwise ranking reversal occurs if \[ D_v(a,b)\{D_v(a,b)-2G_{ab}(v,w)\}<0. \] A sufficient condition for no reversal is \[ 2|G_{ab}(v,w)|<|D_v(a,b)|. \] If \(E[\Delta_t(v,w)\mid\mathcal I_s]=0\) and \(f_{a,s}(t)-f_{b,s}(t)\) is \(\mathcal I_s\)-measurable, then \(G_{ab}(v,w)=0\). In that case revisions can raise loss levels but do not change pairwise rankings through the covariance mechanism in Proposition (ref).

The same identity gives the release-cycle structure of revision risk. For a release horizon \(\tau\), write \(Y_t(\tau)\) for the outcome vintage available at that horizon and \(e_t(\tau)=Y_t(\tau)-f_s(t)\). Relative to the first release, \[ R_f(\tau)=E[e_t(\tau)^2] =E[e_t(0)^2]+E\{\Delta_t(0,\tau)^2\} +2E[e_t(0)\Delta_t(0,\tau)] . \] The second term is revision risk at release horizon \(\tau\). The third term records whether revisions amplify or offset preliminary forecast errors.

The SPF estimates use observed vintage histories for RGDP, INDPROD, EMP, CPI, and PCE. Fixed-day horizons are as-of values: the latest released value available 30, 60, 90, 180, or 365 days after first release. The latest/final value is an ex-post historical benchmark. The national-account application uses exact first, second, and third NIPA release stages where available, the RTDSM QvQd one-year fixed-delay value, and the latest/final value.

table[table omitted — 1,914 chars of source]

Table (ref) summarizes this release-cycle risk structure. The entries report revision-risk shares, covariance shares, risk-amplification ratios, and the released late-error history available at each release horizon. The SPF rows average target-level shares within real activity and inflation so that no target drives the comparison by scale.

figure[figure omitted — 245 chars of source]

Figure (ref) shows why selected first-versus-late comparisons are too narrow. The relevant object is a curve over the release cycle, with uncertainty bands around the target-averaged revision-risk shares. The target-averaged real-activity curve lies above the inflation curve at intermediate release horizons, especially around 180 days. The covariance channel also differs across the two groups: revisions tend to amplify real-activity forecast risk at these horizons, while inflation revisions more often offset preliminary forecast errors. These are group-level patterns in the SPF sample, not a universal target-by-target or horizon-by-horizon ordering. The latest/final row in Table (ref) is an ex-post historical benchmark. It is useful for describing how much historical assessments can move, but it should not be treated as a real-time calibration target unless the relevant latest-value errors had already been released at the forecast origin.

The value of this release-cycle view is not only to identify large revision-risk cases. It also identifies cases in which first-release and later-outcome risk are empirically close, or in which revision variance is offset by covariance with preliminary forecast errors. For forecast reporting, both outcomes matter: large release-horizon gaps call for separate later-outcome uncertainty statements, while small or offsetting gaps justify simpler first-release reporting for that target and release horizon. Online Appendix Table 2 reports the complete 180-day target-level table behind this summary.

National accounts provide a complementary application with a richer revision calendar. The national-account release-cycle pattern is flatter at the one-year horizon than the SPF real-activity pattern, and covariance terms often offset revision variance. The online appendix reports the corresponding release-horizon figure with block-bootstrap bands. The main lesson is not that fixed-forecast ranking changes are common. It is that release choices, training choices, and calibration-error choices must be reported separately because national-account vintages can move through several economically meaningful stages.

The release-cycle evidence also explains the role of later-outcome interval methods. Direct late-outcome calibration becomes harder as the later-error history thins out, while early-release errors and revision histories can be observed sooner. Revision-risk transport is most relevant when the release-cycle evidence indicates nontrivial later-outcome revision risk and the direct later-error history is short.

Identification and Real-Time Uncertainty

For a later outcome vintage, released early-error and revision histories do not by themselves determine the forecast-error distribution. At forecast origin \(s\), a forecaster can use only forecast errors and revision histories that have already been released. For an early outcome vintage, this restriction determines which past errors can enter a real-time calibration sample. For a later outcome vintage, the forecaster may observe many early-vintage errors and many revisions, but substantially fewer late-vintage errors. The later-error distribution is not pinned down by the two released marginal histories alone, because the dependence between early errors and revisions also matters.

The results below formalize this information restriction. First, released-error calibration gives uncertainty bands with a time-series coverage bound that reflects effective calibration size, local drift, and discreteness of the score distribution. Second, later-vintage uncertainty is partially identified from released early-error and revision histories unless additional dependence or revision-model restrictions are imposed. This identification result explains the interval comparison in the applications: direct late calibration uses the right errors when enough of them have been released, dependence-robust transport is valid under marginal information, and signed or revision-model transport can be narrower under stronger dependence assumptions. The local-drift terms allow the relevant score law for a current forecast to differ from the long historical distribution, as in work on local stationarity and forecast instability dahlhaus1997,giacominiRossi2009,rossi2013. The conformal step is not the source of the release-cycle object. It is the device that turns released forecast-error histories into model-agnostic intervals once the relevant outcome vintage has been named.

Released errors and probability statements

Using the notation from Section (ref), all random variables are defined on \((\Omega,\mathcal F,P)\). The same information set \(\mathcal I_s\) that determines forecast feasibility is also the conditioning information for the coverage statements below. For outcome vintage \(v\) and a past target period \(u\), the released absolute error is \[ S_u(v)=|Y_u(v)-f_{s_u}(u)|, \] where \(s_u\) is the origin at which the forecast for \(u\) was made. At origin \(s\), only scores satisfying \(A_u(v)\le s\) may be used. A calibration rule \(m\) selects past target periods \(\mathcal U_s^m(v)\subset\{u:A_u(v)\le s,\ u<t\}\) and nonnegative weights \(\omega_{u,s}^m\). In the result below, this rule is fixed before the calibration quantile is computed. It may depend on quantities known at origin \(s\), including release dates, forecast origins, target types, and pre-specified window lengths. The released error magnitudes \(S_u(v)\) are also observed at \(s\), but they enter the result through the empirical CDF below, not through the choice of \(m\).\footnote{If several rules are compared on the same error history and the best-performing rule is then used to construct the interval, the same calibration sample is doing two jobs. That adaptive step requires an additional correction, such as sample splitting or a multiple-rule adjustment.}

Let \(W_s^m=\sum_{u\in\mathcal U_s^m(v)}\omega_{u,s}^m\) and define the weighted released-score CDF \[ \widehat H_s^m(x;v)= \frac{1}{W_s^m}\sum_{u\in\mathcal U_s^m(v)} \omega_{u,s}^m\mathbf 1\{S_u(v)\le x\}. \] The half-width is the weighted empirical quantile \[ q_s^m(v)=\inf\{x:\widehat H_s^m(x;v)\ge 1-\alpha\}, \] and the released-error interval is \[ \Pi_s^m(v,t)= [f_s(t)-q_s^m(v),\,f_s(t)+q_s^m(v)]. \] After conditioning on \(\mathcal I_s\), the selected scores, weights, quantile, and interval are fixed. The remaining coverage probability is over the new outcome \(Y_t(v)\). The proposition below is therefore a high-probability statement over the random released histories that generate \(\mathcal I_s\); on the good-history event, it gives a conditional coverage bound.

Let \[ F_{v,s}(x)=P\{S_t(v)\le x\mid\mathcal I_s\} \] be the current conditional score law for the new forecast and outcome vintage \(v\). Let \(H_s^m(\cdot;v)\) be the population CDF represented by the released scores selected by rule \(m\), with the same predictable weights. The local drift term is \[ B_s^m(v)=\sup_x |H_s^m(x;v)-F_{v,s}(x)|. \]

proposition[Released-error calibration under dependent histories] Fix \(s,t,v\) and a predictable released-error calibration rule \(m\). Assume: \begin{enumerate}[label=(\roman*),leftmargin=*] • Release admissibility. Each selected score satisfies \(A_u(v)\le s\) and \(u<t\), and the selected set is nonempty with \(W_s^m>0\). • Predictable bounded weights. The selected index set and weights are predictable: they do not depend on the selected score magnitudes and are fixed before those magnitudes are used to form the empirical CDF and quantile, and \[ n_{\mathrm{eff},s}^m(v)= \frac{(W_s^m)^2}{\sum_{u\in\mathcal U_s^m(v)}(\omega_{u,s}^m)^2}, \qquad \max_{u} \frac{\omega_{u,s}^m}{W_s^m} \le \frac{\kappa_s}{n_{\mathrm{eff},s}^m(v)} . \] • Weak dependence. Ordered by target period, the selected score indicator array \(\{\mathbf 1(S_u(v)\le x):x\in\mathbb R\}\) is beta-mixing with coefficients \(\beta_s(k)\), uniformly over thresholds. • Local drift and atoms. \[ \sup_x|H_s^m(x;v)-F_{v,s}(x)|\le B_s^m(v), \qquad F_{v,s}(q_s^m(v))-F_{v,s}(q_s^m(v)^-)\le \rho_{v,s}. \] \end{enumerate} Let \(n_s^m(v)=|\mathcal U_s^m(v)|\), choose a block length \(b_s\), and set \[ N_s(b_s)= \left\lfloor\frac{n_{\mathrm{eff},s}^m(v)}{2b_s}\right\rfloor, \qquad r_s^m(\delta;v)= C\kappa_s \left[ \sqrt{ \frac{\log(n_s^m(v)+1)+\log(4/\delta)} {N_s(b_s)} } +\frac{b_s}{n_{\mathrm{eff},s}^m(v)} \right], \] where \(C\) is a universal constant. Then, with outer probability at least \[ 1-\delta-2N_s(b_s)\beta_s(b_s) \] over the released calibration history, \[ \left| P\{Y_t(v)\in\Pi_s^m(v,t)\mid\mathcal I_s\}-(1-\alpha) \right| \le r_s^m(\delta;v)+B_s^m(v)+\rho_{v,s}. \]

The proposition should be read as a high-probability statement about the released calibration history. Conditional on a realized \(\mathcal I_s\), the interval and quantile are fixed, and coverage is evaluated under the current target-score law \(F_{v,s}\). The three terms in the bound have different interpretations. The term \(r_s^m(\delta;v)\) is the sampling error from estimating the selected released-error distribution. It decreases as the effective calibration size \(n_{\mathrm{eff},s}^m(v)\) increases. Under short-range dependence, one can choose blocks with \(b_s\to\infty\), \(b_s/n_{\mathrm{eff},s}^m(v)\to0\), and \(N_s(b_s)\beta_s(b_s)\to0\), so this term converges to zero. Its leading rate is the usual empirical-CDF rate with an effective sample size penalty, roughly \(\left(b_s\log n_s^m(v)/n_{\mathrm{eff},s}^m(v)\right)^{1/2}\) up to constants and logarithmic factors. The term \(B_s^m(v)\) is different. It is local drift: the distance between the score law represented by the selected past errors and the current forecast's score law. More historical observations do not automatically remove this term. It is small only when the selected released errors are locally representative for the current forecast origin, which is why the empirical work reports stability and period-sensitivity diagnostics. The term \(\rho_{v,s}\) is atom or quantile slack at the selected threshold. It is zero for continuous score distributions with no mass at the threshold and can be positive with rounded, discrete, or tied forecast errors. Thus the proposition gives a convergence statement only when the sampling term vanishes and the drift and atom terms are themselves small.

The proposition is not an exact finite-sample exchangeability result. It is a time-series calibration bound: effective sample size, weak dependence, local drift, and atom slack are the costs of using released macroeconomic forecast errors. Exact conformal validity is recovered only in the same-outcome, exchangeable special case with the usual conformal quantile convention. The most common empirical use in this paper is the primitive case in which the calibration window and weights are fixed before the evaluation outcome is observed. More adaptive rules require either sample splitting, a finite-class adjustment, or additional complexity control; they are not covered by the primitive proposition as stated.

corollary[Average coverage implication] Let \[ R_s^m(v)=r_s^m(\delta;v)+B_s^m(v)+\rho_{v,s}. \] Then the unconditional coverage error is bounded by \[ \mathbb E\{R_s^m(v)\} +\delta+2N_s(b_s)\beta_s(b_s). \] In particular, if \(R_s^m(v)\le \bar R_s\) uniformly over the evaluation design, the unconditional coverage error is at most \(\bar R_s+\delta+2N_s(b_s)\beta_s(b_s)\).

The proposition gives conditional coverage on good released histories. The corollary then converts this into an average coverage statement by accounting for the probability of bad history events.

Mismatched calibration errors

A score may be released but correspond to the wrong outcome vintage. Let \(F_{v,s}\) and \(F_{w,s}\) be the current conditional score laws for outcome vintages \(v\) and \(w\), and define \[ M_s(v,w)=\sup_x|F_{v,s}(x)-F_{w,s}(x)|. \] If a released-error rule uses scores from outcome vintage \(w\) to form an interval for \(Y_t(v)\), the bound in Proposition (ref) holds with \(B_s^m(w)\) replaced by \(B_s^m(w)+M_s(v,w)\), with atom slack evaluated for the target score law \(F_{v,s}\) at the selected cutoff. A mismatched score vintage therefore adds a score-distribution mismatch term even when its own released score law is estimated accurately.

For absolute scores, the mismatch is controlled by revisions. Since \[ |S_t(v)-S_t(w)|\le |Y_t(v)-Y_t(w)|=|\Delta_t(v,w)|, \] for any \(r\ge0\), \[ M_s(v,w)\le \omega_s(r;w)+\eta_s(r;v,w), \] where \[ \eta_s(r;v,w)=P(|\Delta_t(v,w)|>r\mid\mathcal I_s) \] and \[ \omega_s(r;w)= \sup_x\max\{F_{w,s}(x+r)-F_{w,s}(x),\, F_{w,s}(x)-F_{w,s}(x-r)\}. \] Thus calibrating on final or latest errors before those errors were released is not feasible at the forecast origin, and released but mismatched errors carry a revision-drift penalty.

Later-outcome uncertainty

All objects in this subsection are conditional on \(\mathcal I_s\). For an early outcome vintage \(v\) and later vintage \(w\), \[ e_t(w)=e_t(v)+\Delta_t(v,w), \qquad e_t(v)=Y_t(v)-f_s(t),\quad \Delta_t(v,w)=Y_t(w)-Y_t(v). \] Released histories may identify or estimate the marginal law \(F_{e,s}\) of early forecast errors and the marginal law \(F_{\Delta,s}\) of revisions. They do not in general identify the joint coupling between \(e_t(v)\) and \(\Delta_t(v,w)\). For an admissible coupling class \(\Gamma_s\), the identified set for the later-error CDF at \(z\) is \[ \mathcal J_s^{\Gamma}(z)= \left\{ P_{\gamma}(e+\Delta\le z): \gamma\in\Gamma_s,\ \gamma_e=F_{e,s},\ \gamma_{\Delta}=F_{\Delta,s} \right\}. \] This is a partial-identification object in the sense of manski2003, imbensManski2004, and chernozhukovHongTamer2007. See also molinari2020 for a survey. The coupling bounds below are classical Fr\'echet--Makarov bounds for sums with fixed marginals makarov1982,frankNelsenSchweizer1987,nelsen2006. Here they are applied to later-vintage forecast errors using only histories released by origin \(s\).

Define \[ \underline F_s(z)= \sup_x \max\{F_{e,s}(x)+F_{\Delta,s}(z-x)-1,0\}, \] and \[ \overline F_s(z)= \inf_x \min\{F_{e,s}(x)+F_{\Delta,s}(z-x),1\}. \]

proposition[Sharp identified set under released marginals] Fix \(s,v,w\). Suppose the released histories identify the marginal laws \(F_{e,s}\) and \(F_{\Delta,s}\), but impose no restriction on their dependence. Let \(\Gamma_s^{FM}\) be the set of all couplings with these marginals. Then, for every \(z\), \[ \mathcal J_s^{\Gamma^{FM}}(z) = [\underline F_s(z),\overline F_s(z)]. \] The bounds are pointwise sharp: for each \(z\) and each value in this interval there exists an admissible coupling with the released marginals that attains that value.

The proposition uses a classical coupling result. The object here is the later-outcome forecast-error distribution that can be identified from histories actually released at forecast origin \(s\).

corollary[Dependence-robust later-outcome interval] Let \(\ell_s\) and \(r_s\) satisfy \[ \overline F_s(\ell_s)\le \alpha/2,\qquad \underline F_s(r_s)\ge 1-\alpha/2. \] Then \([f_s(t)+\ell_s,\ f_s(t)+r_s]\) covers \(Y_t(w)\) with probability at least \(1-\alpha\), conditional on \(\mathcal I_s\), uniformly over all couplings in \(\Gamma_s^{FM}\). Within this equal-tail construction, moving either endpoint inward is not uniformly justified by the released marginals alone.
proposition[Estimated identified sets] Let \(\widehat F_{e,s}\) and \(\widehat F_{\Delta,s}\) be released empirical or weighted CDFs. Suppose that, on a released-history event with probability at least \(1-\delta\), \[ \sup_x|\widehat F_{e,s}(x)-F_{e,s}(x)|\le \varepsilon_{e,s}, \qquad \sup_x|\widehat F_{\Delta,s}(x)-F_{\Delta,s}(x)| \le \varepsilon_{\Delta,s}. \] Let \(\widehat{\underline F}_s\) and \(\widehat{\overline F}_s\) be the Fr\'echet--Makarov envelopes computed from the estimated marginals. Then \[ \sup_z|\widehat{\underline F}_s(z)-\underline F_s(z)| \le \varepsilon_{e,s}+\varepsilon_{\Delta,s}, \qquad \sup_z|\widehat{\overline F}_s(z)-\overline F_s(z)| \le \varepsilon_{e,s}+\varepsilon_{\Delta,s}. \]

The estimated envelope is continuous in the two marginal CDFs: an error \(\varepsilon_{e,s}\) in the early-error law and an error \(\varepsilon_{\Delta,s}\) in the revision law enlarge the envelope error by at most their sum. For released time series, these two terms include sampling error, local drift, and release-delay effective sample size. When the envelope is converted into interval endpoints, atoms or flat spots in the relevant CDFs add the usual endpoint slack. If both released histories grow and the drift terms are small, the estimated envelope converges to the sharp population envelope. This convergence does not remove the identification width between \(\underline F_s\) and \(\overline F_s\); that width comes from unknown dependence between early errors and revisions.

proposition[Point identification under additional structure] Let \(Z_s(t)\) be \(\mathcal I_s\)-measurable released state information. Suppose the conditional marginal laws \(F_{e,s}(\cdot\mid Z_s)\) and \(F_{\Delta,s}(\cdot\mid Z_s)\) are identified from released histories. If \(e_t(v)\) and \(\Delta_t(v,w)\) are conditionally independent given \(Z_s(t)\), then the later-error CDF \(G_{w,s}\) conditional on \(Z_s(t)\) is point identified by the convolution \[ G_{w,s}(z\mid Z_s) = \int F_{e,s}(z-d\mid Z_s)\,dF_{\Delta,s}(d\mid Z_s). \] If a revision forecast \(r_s(t)\) is \(\mathcal I_s\)-measurable and \[ u_t(v,w)=\Delta_t(v,w)-r_s(t) \] has a conditional marginal law identified from released histories, and if \(u_t(v,w)\) is conditionally independent of \(e_t(v)\) given \(Z_s(t)\), then the distribution of \(Y_t(w)-f_s(t)-r_s(t)\) conditional on \(Z_s(t)\) is the convolution of the early-error law and the residual-revision law.

Proposition (ref) gives the point-identification case. Conditional independence, or a revision model that makes the residual revision independent of the early error after conditioning on \(Z_s(t)\), pins down a single transported distribution instead of a Fr\'echet--Makarov identified set. The cost is that validity now rests on the maintained dependence restriction and, for revision-model transport, on the revision model. In finite samples the remaining errors are marginal estimation error, local drift, atom slack, dependence approximation error, and revision-model error. The empirical diagnostics therefore focus on dependence between early errors and revisions and on whether the revision model reduces that dependence.

corollary[No sharp nonconservative transport from marginals alone] If the Fr\'echet--Makarov envelope is nondegenerate at a tail probability, an equal-tail interval obtained by moving at least one dependence-robust endpoint inward fails to achieve uniform \(1-\alpha\) coverage over \(\Gamma_s^{FM}\) for some admissible coupling.

Corollary (ref) is the boundary of what can be learned from released marginals alone. It does not rule out narrower transport. Rather, it says that tightening the dependence-robust equal-tail endpoints is not uniformly valid from the released marginals alone. A forecaster who wants narrower later-outcome intervals must either impose a dependence restriction, use and validate a revision model, or report a sensitivity class. This is the reason the empirical section treats dependence-robust transport as the validity benchmark and signed or revision-model transport as efficiency-oriented methods that require diagnostics.

Calibration or Transport

The preceding results characterize what released histories can identify. This part turns that result into a practical comparison. Direct late calibration uses errors for the intended later outcome vintage, but those errors arrive late and may be few. Transport uses more released early-error and revision information, but it pays either an identification cost or an assumption cost.

The empirical and simulation sections use five interval procedures. Direct late-vintage calibration uses released forecast errors for the intended later outcome vintage. Revision-aware released-error calibration remains a direct released-error procedure, but updates the calibration history as same-outcome errors for the relevant vintage become available. Dependence-robust transport uses the early-error and revision marginal histories together with the Fr\'echet--Makarov envelope from Corollary (ref). Signed convolution transport uses the signed identity \(e_t(w)=e_t(v)+\Delta_t(v,w)\) and forms the later-error distribution by convolving early errors with revisions under the dependence restriction in Proposition (ref). Revision-model transport first recenters the interval at a released revision forecast and then applies the same logic to early errors and revision residuals. Thus the relevant comparison is an error budget: transport can use longer released histories, but it pays either an identification cost or an assumption cost.

Let \(n_{\ell,s}\) be the effective number of released late-error observations, \(n_{e,s}\) the effective number of released early-error observations, and \(n_{\Delta,s}\) the effective number of released revisions or revision residuals. With a common forecast panel and window, late errors are the most delayed object, so their raw count is no larger than the corresponding early-error and revision counts. Effective counts can differ after weighting, blocking, or adding revision histories outside the forecast panel, but the usual case is \(n_{\ell,s}<n_{e,s}\) and often \(n_{\ell,s}<n_{\Delta,s}\). The drift notation in this section follows the same template as the object in Proposition (ref). There, \(B_s^m(v)\) compares the distribution represented by selected past scores with the current score distribution. Here the same comparison is applied to the three scalar inputs used by the competing procedures: the direct late-error score, the signed early error, and the signed revision. Let \(H_{\ell,s}\), \(H_{e,s}\), and \(H_{\Delta,s}\) be the population laws represented by the released late-error, early-error, and revision histories used in the comparison. Let \(F_{\ell,s}\), \(F_{e,s}\), and \(F_{\Delta,s}\) be the corresponding current laws. Then \[ B_{\ell,s}=\sup_x|H_{\ell,s}(x)-F_{\ell,s}(x)|,\quad B_{e,s}=\sup_x|H_{e,s}(x)-F_{e,s}(x)|,\quad B_{\Delta,s}=\sup_x|H_{\Delta,s}(x)-F_{\Delta,s}(x)|. \] If revision-model transport uses residual revisions, \(H_{\Delta,s}\) and \(F_{\Delta,s}\) are read as residual-revision laws. The same mapping applies to atom and quantile slack. The term \(\rho_{v,s}\) in Proposition (ref) is the mass or flat-spot slack at the selected cutoff. Here \(\rho_{\ell,s}\) is the corresponding slack for the direct late-error cutoff, \(\rho_{\mathrm{rob},s}\) for the dependence-robust interval endpoints, and \(\rho_{\mathrm{sgn},s}\) for the signed or revision-model transport endpoints. With \(\mathfrak s(n,\delta)=C_s\sqrt{\log(1/\delta)/n}\), a schematic direct late-calibration error budget is \[ \mathcal E_{\mathrm{direct},s} = \mathfrak s(n_{\ell,s},\delta)+B_{\ell,s}+\rho_{\ell,s}. \] Dependence-robust transport has \[ \mathcal E_{\mathrm{robust},s} = \mathfrak s(n_{e,s},\delta) +\mathfrak s(n_{\Delta,s},\delta) +B_{e,s}+B_{\Delta,s} +W_s^{FM}(\alpha)+\rho_{\mathrm{rob},s}, \] where \(W_s^{FM}(\alpha)\) is the tail or quantile cost induced by the Fr\'echet--Makarov identified set. Signed or revision-model transport replaces that identified-set width with dependence and revision-model approximation costs: \[ \mathcal E_{\mathrm{signed},s} = \mathfrak s(n_{e,s},\delta) +\mathfrak s(n_{\Delta,s},\delta) +B_{e,s}+B_{\Delta,s} +\Pi_s+M_s+\rho_{\mathrm{sgn},s}. \] Here \(\Pi_s\) measures conditional-independence or stable-dependence approximation error, and \(M_s\) is revision-model error when the interval is centered at \(f_s(t)+r_s(t)\).

The sufficient comparison is simply \[ \mathcal E_{m,s}<\mathcal E_{\mathrm{direct},s}, \qquad m\in\{\mathrm{robust},\mathrm{signed}\}. \] When the relevant error CDF has density bounded away from zero near the target quantiles, the same comparison translates into a smaller worst-case quantile error up to the inverse-density constant. This condition is deliberately only an error-budget comparison. It is useful because it exposes the economic and statistical components of method choice, but it is not a new automatic selector.

The comparison has a simple interpretation. Direct late calibration is most appealing when the released late-error history is already informative and stable. Dependence-robust transport is useful only when late errors are scarce and the Fr\'echet--Makarov identified set remains reasonably narrow. Signed or revision-model transport can be more efficient, but only when revision predictability or stable error--revision dependence makes the extra identifying restriction credible. When late histories are scarce, robust transport is wide, and the signed restrictions are weak, the appropriate conclusion is weak identification rather than a sharper interval. Figure (ref) summarizes this interpretation.

figure[figure omitted — 1,405 chars of source]

A feasible empirical index based on these components is reported in the online appendix as a diagnostic. In the current applications it is conservative and is not used as an automatic selector. The main text therefore uses the error-budget comparison qualitatively: it identifies the forces that favor direct late calibration, dependence-robust transport, or assumption-dependent transport.

Monte Carlo Evidence

The revision-transport theory separates conservative validity from sharper methods that use additional structure. The Monte Carlo exercises ask when that structure is useful. They are not designed to show that transport uniformly improves on direct released-error calibration. They are designed to identify boundary conditions: late-error scarcity, revision predictability, benchmark shocks, score-scale drift, and unstable dependence between early errors and revisions.

Design

All designs share the same release-indexed structure. A latent macroeconomic target \(Z_t\) generates an early released outcome \(Y_t(v)\) and a later outcome \(Y_t(w)\): \[ Y_t(v)=Z_t+\eta_t^{v}, \qquad Y_t(w)=Y_t(v)+\Delta_t(v,w), \] where the revision process is written as \[ \Delta_t(v,w)=\mu_\Delta(X_t)+\sigma_\Delta(R_t)\zeta_t+B_t . \] Here \(X_t\) is a real-time covariate, \(R_t\) is a volatility or crisis state, and \(B_t\) is an occasional benchmark-revision shock. Forecasts are generated recursively from information available at the forecast origin. The forecast errors are \[ e_t(v)=Y_t(v)-f_t, \qquad e_t(w)=Y_t(w)-f_t=e_t(v)+\Delta_t(v,w). \] Release delays are imposed by allowing \(S_u(w)=|Y_u(w)-f_u|\) to enter direct late-vintage calibration only after the later outcome has been released. Early errors and revision histories enter transport rules only after their required outcome vintages have been released.

The specifications vary only the revision process, the score process, and the delay length. Stable small revisions use a low-variance revision shock with a stationary score distribution. Large unpredictable revisions increase revision variance but keep revisions independent of real-time forecast information. Predictable revisions set \(\mu_\Delta(X_t)=\beta_r r_t\), where \(r_t\) is a persistent real-time signal observed at the forecast origin. The coefficient \(\beta_r\) is large enough that a simple revision model can reduce residual revision dispersion, while an idiosyncratic revision shock remains. State-dependent revisions raise revision dispersion and shift the revision mean during crisis states. Benchmark-shock designs add a persistent shift \(B_t\) after a fixed date. Delayed-final designs lengthen the release delay for \(Y_t(w)\), reducing the effective sample for direct late-vintage calibration. Scale-drift designs change the forecast-error scale over time. Local-drift designs gradually move the score distribution. The realistic mixture combines a milder predictable component, crisis-state revision shifts, extra crisis volatility, a benchmark shift, gradual scale drift, local score drift, and a nonlinear latent target. It is meant to mimic a setting where revisions are partly predictable and sometimes state dependent, but no single feature makes transport automatically dominate released-error calibration.

The main table focuses on four regimes that are most informative for the transport methods: benchmark shock, delayed final values, predictable revisions, and realistic mixture. Online Appendix Table 11 reports the full Monte Carlo grid.

Simulation Procedures

The compact table compares five feasible interval procedures in the same order across regimes: direct late-vintage calibration, revision-aware released-error calibration, dependence-robust transport, signed convolution transport, and revision-model transport. Section (ref) defines these procedures and explains the direct-versus-transport comparison. Online Appendix Table 11 also reports first-release intervals, absolute bridge intervals, optimized bridge intervals, and an oracle diagnostic. Those additional rows are useful diagnostics, but they are not needed for the main comparison.

table[table omitted — 2,081 chars of source]

The benchmark-shock rows show the tail-risk failure mode: direct, revision-aware, signed, and revision-model intervals are too narrow, while robust transport restores coverage at the cost of width. The delayed-final and predictable-revision rows show where transport can pay off, especially when a revision model reduces the dispersion of later-outcome errors. In the realistic mixture, the simple revision-aware benchmark is already strong, so revision-model transport yields only a modest score improvement. These patterns motivate the diagnostic interpretation used in the applications.

Results

Table (ref) reports nominal 90 percent intervals. The benchmark-shock design shows the main failure mode for signed transport. Direct late calibration and revision-aware calibration do not cover because released histories do not represent the new benchmark tail. Signed and revision-model transport are also too narrow. Robust transport is wider, but it restores coverage by protecting against dependence and tail behavior that the signed methods miss.

The delayed-final and predictable-revision designs show when transport can help. In delayed-final designs, direct late calibration has fewer usable origins, and revision-model transport improves the interval score relative to direct late calibration while maintaining above-nominal coverage. In predictable-revision designs, the revision-model method has the strongest score--coverage tradeoff among the feasible near-nominal methods shown: it uses revision predictability to reduce width and improve interval score. Robust transport covers in both regimes, but its envelope is too conservative when the signed structure is benign.

The realistic mixture combines mild predictability, state variation, and occasional revision shocks. In this regime, the simple revision-aware benchmark is already strong. Revision-model transport still has the lowest score among the feasible methods shown, while signed transport gives only a modest gain. This is the pattern to expect in empirical macro applications: transport is one possible response when late errors are scarce and revisions are stable or predictable, but simple released-error benchmarks can be hard to beat when calibration histories are already informative.

The Monte Carlo evidence therefore supports a diagnostic use of revision transport. Robust transport is the validity-oriented construction when dependence or tail behavior is uncertain. Signed and revision-model transport are efficiency-oriented constructions for regimes with stable dependence or predictable revisions. The empirical sections below use the same logic: dependence diagnostics and common-origin comparisons determine whether transport is interpreted as a safe interval, an efficiency improvement, or a method that should be treated cautiously.

Empirical Applications: SPF and National Accounts

The empirical applications ask when the information restrictions characterized in Section (ref), especially the error-budget comparison in Section (ref), favor transport, and when direct or revision-aware calibration remains preferable. SPF and national accounts provide complementary release-cycle settings. The out-of-sample evaluation period is the primary prospective evidence, while target-level results, period decompositions, and common-origin benchmarks explain why performance differs across variables and time.

Design and scoring

The SPF application uses Federal Reserve Bank of Philadelphia mean and median forecasts for CPI, EMP, INDPROD, PCE, and RGDP. We use the current-period forecast and the one-period-ahead forecast, denoted SPF horizons 0 and 1. The outcome pair is first release to the value available about 180 days after first release. Forecast origins run from 1996-03-02 to 2025-11-11, with 1996-03-02 to 2018-05-08 used as the pre-evaluation sample and 2018-08-07 to 2025-11-11 reserved for out-of-sample evaluation. For CPI, EMP, INDPROD, and PCE, we match forecasts to real-time vintages of the corresponding source series: CPIAUCSL for CPI, PAYEMS for EMP, INDPRO for INDPROD, and PCEPI for PCE. RGDP is matched to Philadelphia Fed RTDSM real-output values. The online appendix documents the source-series mapping, release-date construction, and the RGDP release-date caveat.

The national-account application uses real-time RTDSM values with day-level BEA release dates where the release calendar can be matched. The outcome pair is first release to a one-year QvQd fixed-delay value. Eligible components are P, RCON, REX, RG, RIMP, RINVBF, RINVRESID, ROUTPUT, and YRGDI. Forecast rules are fixed within each comparison and consist of AR-lag, pooled-ridge component, and simple forecast-combination rules. Forecast origins run from 2016-04-27 to 2025-04-29, with 2016-04-27 to 2023-01-25 used as the pre-evaluation sample and 2023-04-26 to 2025-04-29 reserved for out-of-sample evaluation. Online Appendix Table 3 reports the compact design table with target cells, forecast-origin counts, and scoring definitions.

The methods are direct later-outcome calibration, revision-aware released-error calibration, dependence-robust transport, signed transport, and revision-model transport. A forecast-origin observation enters a comparison only when every method is feasible for the same forecast origin and target cell. The nominal interval level is 90 percent. The primary score is the interval score divided by a pre-evaluation RMSFE scale for the relevant target cell. That scale is fixed using the pre-evaluation sample only. The normalized score is a comparative coverage--width measure, not a welfare magnitude: within a target cell, a lower value means a better coverage--width tradeoff relative to the forecast difficulty observed in the pre-evaluation sample. SPF summaries give equal weight to each target--horizon--forecast-type cell, and national-account summaries give equal weight to each target--model cell.

Out-of-Sample Evaluation

The out-of-sample evaluation is the primary evidence about prospective performance. It uses the last calendar-time block of each sample after the target pairs and interval methods have been fixed. Table (ref) reports these results. In SPF, dependence-robust transport lowers normalized interval score relative to direct late calibration, from 22.92 to 22.04. It also raises coverage from 80.0 to 84.5 percent. Yet it still undercovers relative to the nominal 90 percent level. Revision-aware calibration gives a smaller score improvement and similar coverage to direct late calibration. Signed and revision-model transport are narrower, but do not resolve the coverage shortfall. The SPF evaluation period therefore shows the central score--coverage tradeoff: early-error and revision histories can improve score, but the improvement is usable only when coverage remains close to the nominal target.

table[table omitted — 1,482 chars of source]

The national-account evaluation period gives a different lesson. Direct late calibration and revision-aware released-error calibration already cover well, and revision-aware calibration has the lowest normalized score. Dependence-robust transport is safer but much wider, while signed and revision-model transport cover well without improving on direct or revision-aware calibration. Here the later-error histories are informative enough that transport's identification and dependence costs outweigh its use of revision histories.

Target-Level Heterogeneity

The target-level figure explains why pooled results are not sufficient for method choice. This heterogeneity is the empirical reason for treating revision risk as target- and release-horizon-specific rather than reporting a single revision correction. Figure (ref) plots every pre-specified target by later-period coverage and normalized score gain relative to direct late calibration. In SPF, robust transport improves score for CPI, PCE, EMP, and RGDP, but those points have coverage below the nominal level. INDPROD is the most favorable SPF transport case because revision-model transport improves score modestly while keeping coverage close to 90 percent. Thus the aggregate SPF score gain is partly driven by cells where coverage remains binding.

figure[figure omitted — 722 chars of source]

A compact reading of the figure is that transport has useful cases, and their value is target-specific. For SPF, INDPROD is the cleanest transport example, while CPI and PCE illustrate score gains that remain constrained by coverage. For national accounts, RIMP and YRGDI are the most favorable transport cases, but many components are better described as cases where direct or revision-aware calibration is already informative.

National-account heterogeneity is also substantial. RIMP is the clearest dependence-robust transport case, and YRGDI is the clearest signed or revision-model transport case. For many other components, revision-aware released-error calibration is the strongest non-direct method, while robust transport mainly adds width. Online Appendix Table 4 gives the compact target-level summary, and Online Appendix Table 7 reports the target-level out-of-sample detail for every pre-specified target, including cases where transport does not improve on direct calibration. Online Appendix Section F.2 reports the corresponding revision-transport dependence diagnostics, including early-error--revision correlations, residual correlations, and tail-dependence measures. These diagnostics are used to interpret signed and revision-model transport, not to turn them into assumption-free procedures.

The target-level patterns should also be read together with the period decomposition in Online Appendix Table 6. In the SPF later evaluation sample, pre-COVID coverage is near or above nominal for all methods, and post-COVID coverage is again close to nominal. The COVID block is the exception: direct late calibration covers 46.3 percent, revision-aware calibration covers 47.5 percent, and robust transport covers 57.5 percent. Robust transport still reduces raw interval score in that block, but no method provides adequate coverage. The SPF deterioration in the later evaluation period is therefore primarily a local scale and tail episode, not a simple shortage of late errors. In the same evaluation period, average SPF late-error histories are larger than in the full common-origin benchmark \((99\) versus \(71)\), and average revision histories are also larger. National-account late-error histories similarly rise from about \(38\) to \(52\). Later-error scarcity is therefore not mechanically worse in the later evaluation period. The issue is whether the released histories are stable and informative for the current release horizon. Modeling this period dependence is an important direction for future work.

Common-Origin Benchmark

The closest-method benchmark has a different role. It asks whether the paper's methods remain competitive with simple parametric residual bands on the full common-origin historical sample. Table (ref) compares direct late released-error calibration, revision-aware released-error calibration, dependence-robust transport, and revision-model transport with a Gaussian same-outcome residual band and a revision-adjusted Gaussian band. The revision-adjusted Gaussian band is a simplified analogue motivated by clements2017MacroUncertainty and clementsGalvao2013JAE,clementsGalvao2023BVARDataUncertainty.\footnote{A full replication would change the forecast producer and density model. The benchmark instead holds the issued point forecast fixed and compares interval construction.} Secondary parametric benchmarks, including Student-\(t\) residual bands, are reported in the online appendix.

table[table omitted — 1,610 chars of source]

In the full common-origin benchmark, robust transport has the lowest normalized SPF score among the compact methods and near-nominal coverage. In national accounts, revision-model transport and revision-aware released-error calibration have the lowest scores, while Gaussian bands are competitive but not dominant.

Online Appendix Table 8 reconciles the historical benchmark with the out-of-sample evidence. The common-origin benchmark describes rolling performance under the full historical mix. The out-of-sample period is the stronger prospective evaluation. In SPF, robust transport looks stronger in the common-origin benchmark than in the later evaluation period because that period contains the COVID scale shift. In national accounts, the later-period late-error histories are already informative enough that direct and revision-aware calibration are strong, so robust transport mainly adds width. The difference between the two panels is therefore evidence about stability and method choice, not a contradiction. This reconciliation is central to the empirical interpretation: common-origin performance describes historical fit under a broad mix of periods, while the later evaluation period tests prospective stability.

Empirical implications

The results imply four practical lessons for method choice.

enumerate[label=\arabic*.,leftmargin=*] • Naming the later outcome is necessary because revision risk differs by target and release horizon. The release-cycle structure shows when first-release, 180-day, one-year, and latest-value uncertainty statements diverge. • Later-error scarcity is only one input into method choice. SPF shows that transport can lower normalized score, but coverage depends on stability of recent errors and revisions. • Dependence-robust transport and signed transport answer different questions. Robust transport trades width for protection against unknown dependence. Signed and revision-model transport trade robustness for efficiency when dependence and revision-model diagnostics are favorable. • Direct or revision-aware calibration remains the appropriate benchmark when later-error histories are informative. The national-account later evaluation period illustrates this case: once the intended later-outcome errors are available in sufficient quantity and stability, simpler released-error methods can beat transport on normalized score while maintaining coverage.

Conclusion

Macroeconomic forecast risk changes as measured outcomes move through the release cycle. A first release, an intermediate vintage, and a latest-value benchmark can answer different uncertainty questions for the same forecast. The empirical evidence shows that revision risk is not uniformly large, but it is measurable, target-specific, and sometimes amplified by its covariance with preliminary forecast errors. This makes the outcome vintage part of the forecast-risk statement rather than a detail to be supplied after evaluation.

For uncertainty quantification, later-outcome risk is only partly identified from histories available at the forecast origin unless additional dependence restrictions are imposed. Released early-error histories and revision histories identify marginal laws, but not the dependence between early errors and revisions. Fr\'echet--Makarov bounds give the dependence-robust identified set. Signed and revision-model transport can be narrower, but their interpretation depends on stable dependence or revision predictability. Released-error calibration provides a feasible model-agnostic way to turn these objects into intervals.

The practical implication is a reporting discipline for real-time forecast uncertainty. A forecast report should name the outcome vintage being covered, state which outcome vintage is used to form calibration errors, report the revision-risk and covariance diagnostics relevant for that release horizon, and identify whether the interval is direct, dependence-robust, or assumption-dependent. When later-error histories are sufficiently informative, direct or revision-aware calibration can be preferable. When later errors are scarce and the dependence diagnostics are favorable, revision-risk transport can use released early-error and revision histories to make later-outcome uncertainty statements feasible.