Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
54,850 characters · 12 sections · 0 citation commands
Forecasts with Bayesian vector autoregressions under real time conditions
\thispagestyle{empty}
\onehalfspacing
Forecasting economic and financial series is of crucial interest to the private sector and for policy makers in central banks and governments. Many contributions in specialized academic journals thus develop new econometric methods, aiming to improve forecast accuracy of existing approaches. The proposed models are benchmarked against established frameworks, and evaluated in terms of their performance for point and density forecasts mostly by means of so called pseudo out-of-sample exercises.
In these simulations, one splits available data at a specific point in time into an estimation and holdout period, produces forecasts for the relevant horizon based on the estimation period and assesses these forecasts with realized values in the holdout sample. Subsequently, an additional observation is added to the information set, and the procedure is iterated until the end of data in the holdout is reached. Results for vast sets of models can subsequently be compared, and a “winner” in terms of point and density forecasts can be chosen on the grounds of various applicable statistics \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, for instance,][]{geweke2010comparing}.
A key difference for forecasting under real time conditions is, however, that data are frequently subject to revisions and that observations for many series are not yet available to the forecaster until the present date upon first release of the data. This implies that practitioners must rely on incomplete information sets, the “vintages” published at a specific point in time. By contrast, pseudo out-of-sample forecasts are produced and evaluated using truncated series from a single vintage. Handling real time data and accounting for revisions and missing values is a burdensome task. Is it necessary to address such concerns, or do pseudo out-of-sample exercises suffice to establish the relative performance ordering of different modeling approaches?
In this paper, we assess the sensitivity of forecasts to taking a real time perspective, and seek to identify models that are most robust to data revisions, and situations where data for the most recent periods are not yet available \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[for related studies, see][]{diron2008short,giannone2008nowcasting,molodtsova2008taylor,SCHUMACHER2008386,BANBURA2011333,banbura2013now,ghysels2014forecasting,krippner2018real}.\footnote{For pioneering work in the context of real time analysis of data, see \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{diebold1991forecasting} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{croushore2001real}. A comprehensive survey of the related literature is given by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{croushore2006forecasting,croushore2011frontiers}.} Rather than focusing on developing methods for predictable measurement errors in initial releases of statistical agencies \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{aruoba2008data,STROHSAL2020} as in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{kishor2012var} or \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{cogley2015price}, our interest centers on assessing the robustness of popular modeling approaches that do not explicitly account for data revisions. We ask whether forecast metrics derived from pseudo out-of-sample and real time studies establish the same ordering in terms of forecast performance, and thus agree on model selection.
We investigate datasets for the United States (US) and the Euro Area (EA). The Federal Reserve Bank of St. Louis maintains monthly vintages for the US, where the dataset for a given month only contains observations until the previous month and data are revised frequently \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[Federal Reserve Economic Data, FRED-MD, see][]{doi:10.1080/07350015.2015.1086655}. A comparable database is available for the EA, with even longer lags in the publication of available series \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[Euro Area Real Time Database, EA-RTD, see][]{giannone2012area}.
Using a set of popular models reflecting recent trends in macroeconomic forecasting, that is, Bayesian time-varying parameter (TVP) vector autoregressive (VAR) models with stochastic volatility (SV) of different sizes equipped with a flexible shrinkage prior, we investigate relative differences in forecast performance relying on real time information and pseudo out-of-sample exercises. To gauge the role of large information sets while keeping computation feasible, we also use variants of these models augmented by principal components extracted from the high-dimensional datasets that are available \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see also][]{bernanke2003monetary}. Moreover, we employ methods for in-sample imputation of missing values.
Our results suggest three main conclusions. First, we observe differences in the relative ordering of model performance for point and density forecasts depending on whether real time data or truncated final vintages in pseudo out-of-sample simulations are used for evaluating forecasts. This finding suggests that when developing econometric frameworks for forecasting, special attention has to be paid to the robustness of the proposed models to missing observations and data revisions. Providing methods for this case is crucial for practitioners, since initial data releases are typically incomplete and imperfect. Second, depending on the release schedule of the data, we detect differences in the severity of this problem between the US and the EA. Missing values and data revisions play a much more prominent role for the case of the EA. Finally, and perhaps unsurprisingly, producing accurate forecasts is substantially harder when relying on real time data. Depending on the variables and forecast horizon of interest, no clearly superior model specification arises. Relatedly, we cannot identify a unique specification that performs best for the US and EA economy. Differences in forecast performance between real time and pseudo out-of-sample designs are typically smaller for larger and more complex models featuring TVPs.
The rest of the paper is structured as follows. Section (ref) briefly presents the econometric framework. Section (ref) introduces the datasets and discusses imputation methods for missing data alongside details on model specification. The main findings are discussed in Section (ref). Section (ref) concludes.
Recent approaches in macroeconomic forecasting usually involve high-dimensional multivariate models to extract information from many available series to forecast key variables of interest. Examples are variants of factor or large Bayesian VAR models \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{forni2000generalized,stock2002macroeconomic,banbura2010large,giannone2015prior,doi:10.1080/07350015.2016.1256217,carriero2019large}. In addition, forecasters rely on methods to account for nonlinear dynamics and structural breaks. Key features identified to achieve gains in forecast performance, especially for predictive densities, are SVs, and to a lesser degree, TVPs \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, for instance,][]{d2013macroeconomic,KOOP2013185,carriero2016common,doi:10.1002/jae.2555,feldkircher2017sophisticated,huber2020Bayesian}. For the purposes of this paper, we consider TVP-VARs with SV, equipped with a flexible shrinkage prior as in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{feldkircher2017sophisticated} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{HKO}.\footnote{To incorporate high-dimensional information while keeping the computational burden such models entail feasible, we also consider variants augmented by principal components. This amounts to extracting principal components from the full information set other than the series treated as observed (depending on the model size) and including them in the vector $\bm{y}_t$.}
The baseline TVP-VAR-SV specification for an $M$-dimensional vector of endogenous variables $\bm{y}_t$ for $t=1,\hdots,T$ is
with $M\times M$-matrices of time-varying regression coefficients $\bm{A}_{pt}$ (for lags $p=1,\hdots,P$), and a Gaussian error term $\bm{\epsilon}_t$ with zero mean and time-varying $M\times M$-covariance matrix $\bm{\Omega}_t$. In the empirical application, we also include an intercept term that we omit here for brevity.
The covariance matrix can be decomposed as $\bm{\Omega}_t = \bm{H}_t \bm{\Sigma}_t \bm{H}_t'$, with $\bm{H}_t$ denoting a lower unitriangular matrix, and $\bm{\Sigma}_t = \text{diag}(\sigma_{1t}, \hdots, \sigma_{Mt})$. The model can be recast in triangular regression form using $\bm{A}_t = (\bm{A}_{1t},\hdots,\bm{A}_{Pt})$ and $\bm{x}_t = (\bm{y}_{t-1}',\hdots,\bm{y}_{t-P}')'$,
Reparameterizing the model and augmenting the individual VAR equations with the preceeding residuals allows to treat the covariances as regression coefficients \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see also][]{carriero2019large}, and for equation-by-equation estimation. Here, we exploit the fact that $\bm{H}_{t}^{-1}\bm{\epsilon}_t = \bm{\eta}_t$.
Denoting the $i$th row of $\bm{A}_t$ by $\bm{A}_{i\bullet,t}$ and the $j$th element in the $i$th row of $\bm{H}_t$ by $h_{ij,t}$, the $i$th equation for $i>1$ is
Equivalently, using $\bm{\beta}_{it} = \left(\bm{A}_{i\bullet,t},\{h_{ij,t}^{-1}\}_{j=1}^{i-1}\right)'$ of dimension $K_i\times1$ with $K_i=pM+i-1$ and corresponding $\bm{z}_{it}=\left(\bm{x}_t',\{-\epsilon_{jt}\}_{j=1}^{i-1}\right)'$,
The states are assumed to follow a random walk process with zero mean Gaussian innovations and a diagonal $K_i\times1$-covariance matrix $\bm{V}_i=\text{diag}\left(\{v_{ij}\}_{j=1}^{K_i}\right)$,
Using a non-centered parameterization of the model \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{FRUHWIRTHSCHNATTER201085} allows to treat the square root of the innovation variances in $\bm{V}_i$ as additional regressors,
Here, $\sqrt{\bm{V}_i}=\text{diag}\left(\{\sqrt{v_{ij}}\}_{j=1}^{K_i}\right)$, and $\bm{\tilde{\beta}}_{it}$ with $\tilde{\beta}_{ij,t} = \frac{\beta_{ij,t}-\beta_{ij,0}}{\sqrt{v_{ij}}}$. This feature allows for inducing shrinkage not only on the constant part of the parameters, but also on the amount of time variation in the states.\footnote{Setting the $v_{ij}=0$ and accounting for the reduced number of free parameters in the posterior of the shrinkage prior yields estimates for constant parameter specifications.}
We rely on equation-specific global-local shrinkage via the popular horseshoe prior \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{carvalho2010horseshoe}. Note that many shrinkage priors can be used on these coefficients \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{cadonna2019triple,HKO}. Popular choices also include the Bayesian Lasso \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{doi:10.1198/016214508000000337}, the Normal-Gamma \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{griffin2010} or the Dirichlet-Laplace prior \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{doi:10.1080/01621459.2014.960967}. For a comparison of the empirical shrinkage properties of various priors in VAR models, see \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{cross2020macroeconomic}.
On the log of the time-varying variances, we impose the following independent AR(1) processes and use the prior setup and algorithm proposed in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KASTNER2014408}:
All competing models feature stochastic volatilities, based on inferior forecast performance identified in many studies assuming constant variances \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, for instance,][]{clark2011real}. Appendix (ref) provides details on prior specification and the estimation algorithm.
In this paper we consider two available real time datasets for the US and the EA. Both are publicly available for download at \href{https://research.stlouisfed.org/}{research.stlouisfed.org} (Federal Reserve Economic Data, FRED-MD) and \href{http://sdw.ecb.europa.eu/}{sdw.ecb.europa.eu} (Euro Area Real Time Database, EA-RTD), respectively. Detailed descriptions of the data and a priori transformations to achieve approximate stationarity are provided in Appendix (ref).
For the US, the Federal Reserve Bank of St. Louis maintains many macroeconomic and financial series on a monthly frequency starting in 1959:01, with monthly data vintages available from 1999:08. We preselect a set of $99$ series with consistent availablity of all relevant variables starting from 1980:01. This establishes the initial sampling period for the forecast exercise with approximately $240$ vintages. The dataset referred to as FRED-MD is described in detail in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{doi:10.1080/07350015.2015.1086655}. The publication schedule of this data implies that the vintage in 2000:01, for instance, contains data up to 1999:12, with the current months' values not yet available. Taking a real time perspective, this necessitates these missing values to be “nowcasted,” or imputed, to enable forecasts further into the future.
A similar database for the EA is constructed using information from the Monthly Bulletin published by the European Central Bank (ECB). This dataset, EA-RTD, is described in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{giannone2012area}. The Monthly Bulletin provides the ECB Governing Council with the most recent macroeconomic and financial data available, and thus establishes a historical record of vintages. Since many of the variables of interest were only established after the inception of the Euro in 1999, consistent coverage can be achieved starting from 1999:01. We preselect a set of $94$ relevant variables. The first available vintage was published on January 3, 2001, which, accounting for some months where no new data became available or revisions took place twice, results in approximately $180$ vintages per series. To obtain a reasonably long initial estimation period for the forecast exercise, we only rely on vintages published after 2003:12.
Due to the mode of publication of the underlying Monthly Bulletin, the release schedule of the vintages in the EA is less strictly organized than for the case of the US data. Consequently, individual series may exhibit a different number of missing values, depending on the date of the respective release, implying a “ragged edge” of data availability at the end of the sample \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see also][]{jarocinski2018inflation}. In selected periods, this is also the case for FRED-MD. For instance, while oil prices may be already available for the full length of the vintage sample, inflation data may not yet have been released for the current month. This calls for a conditional nowcast or data imputation scheme, described in the next subsection.
Both datasets range until 2019:01. We preprocess the vintages (when applicable) to account for seasonality using the methods established by the US Census Bureau \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[X-13-ARIMA-SEATS, see][]{monsell2007x13,sax2018seasonal}, and obtain data on interest rates from the ECB's financial market database. For pseudo out-of-sample simulations, we rely on truncated samples from the final vintage, while for real time simulations we use data available at the specific months. In both cases, we evaluate the forecasts using the final available vintage.
The Bayesian approach to estimation allows for a fully probabilistic treatment of missing endogenous variables. There are two relevant cases to be considered. First, a subset of values may be missing, while other variables are already available. Second, at a specific point in time, values for all series may be missing, which allows for sampling these values similarly to unconditional forecasts.
Consider the case where all series are available for $\bm{y}_t$ at time $t$, but a subset of $q$ series at $t+1$ is missing. Let $\bm{y}_{t+1}^{\ast} = (y_{1,t+1}^{\ast},\hdots,y_{q,t+1}^{\ast},y_{q+1,t+1},\hdots,y_{M,t+1})'$ with asterisks indicating missing values and note that the series can always be reordered to yield such a structure. We partition the vector as $\bm{\tilde{y}}_{t+1}^{\ast} = (y_{1,t+1}^{\ast},\hdots,y_{q,t+1}^{\ast})'$ of dimension $q\times1$ and $\bm{\tilde{y}}_{t+1} = (y_{q+1,t+1},\hdots,y_{M,t+1})'$ of size $(M-q)\times1$.
The fitted values at time $t+1$ are $\bm{A}_{t+1}\bm{x}_{t+1} = (\bm{\mu}_1',\bm{\mu}_2')'$, with $\bm{\mu}_1$ being a $q\times1$ vector corresponding to the missing values, and $\bm{\mu}_2$ the $(M-q)\times1$ vector related to the available series. Similarly, we partition the covariance matrix $\bm{\Sigma}_{t+1}$, denoting the upper left $q\times q$ block by $\bm{\Sigma}_{11}$, the upper right $q\times(M-q)$ block by $\bm{\Sigma}_{12}$, the bottom left $(M-q)\times q$ block by $\bm{\Sigma}_{21}$, and the bottom right $(M-q)\times (M-q)$ block by $\bm{\Sigma}_{22}$.
The distribution of the missing values conditional on the realizations, the endogenous variables up to $t$ and all other model parameters, follows from the properties of the multivariate Gaussian distribution. In particular,
Conditioning on the realized values alters both the mean $\bm{\bar{\mu}}$ and variance $\bm{\bar{\Sigma}}$ of the distribution of the missing values, although the variance does not depend upon the particular value of the realizations. For the case where all values at a specific point in time are missing, the distribution of $\bm{\tilde{y}}_{t+1}^{\ast}$ is
Nowcasting the missing values in this way is related to data augmentation techniques, and moreover allows for drawing the model coefficients conditional on the synthetic data in each iteration of the algorithm. Consequently, posterior uncertainty surrounding both the missing data and model parameters is accounted for during estimation.
This subsection describes the differently sized information sets used for estimation. We focus on forecasting inflation, a short-term interest rate and the unemployment rate as key variables of interest (for abbreviations, see Appendix (ref)). We rely on three different model sizes, where the information sets are specified to approximately correspond to similar studies employing VAR models for forecasting, conditional on the availability of the respective series in all vintages:
To circumvent scaling issues arising from the data and to improve stability of the models, we standardize each vintage prior to estimation to have zero mean and unit variance by subtracting the mean and dividing all series by their standard deviation. We denote the sample mean and standard deviation by $\bm{m}_{y}$ and $\bm{s}_{y}$, respectively.
Augmented variants of the baseline specification to use the full available information sets are constructed as follows. Depending on the respective model size, we extract five principal components from the remaining (demeaned and standardized) series and include them in the vector of endogenous variables $\bm{y}_t$. Varying the number of principal components does not significantly alter the forecasting results. Rather than taking a fully Bayesian approach, we rely on this approximation to reduce the computational burden. This implies that the largest model features $M=16$ endogenous variables. We use $P=2$ lags.
One of the main questions this paper raises is whether pseudo out-of-sample forecasts suffice to establish a clear ranking in terms of predictive performance that also applies in the real time context. We assess both moments of the predictive distributions and evaluate point and density forecasts. Details on the employed metrics are provided in Appendix (ref).
To gauge overall forecast performance of the competing models in real time and pseudo out-of-sample designs, we rely on cumulative joint log predictive scores (LPS) over the full holdout. Higher scores indicate superior performance, and allow for constructing a ranking of the models on a monthly frequency. We obtain these ranks for the real time and pseudo out-of-sample exercises, and study their correlation over time to see whether they agree on the ordering of the models in terms of relative forecast performance.
As a summary statistic, we calculate Kendall's rank correlation coefficient $\tau$ for all holdout periods.\footnote{Popular nonparametric correlation estimators are Spearman's $\rho$ and Kendall's $\tau$. We choose Kendall's $\tau$ due to its superior robustness properties \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{croux2010influence} and note that estimates for the correlation are almost identical in terms of both Spearman's and Kendall's coefficient.} Values of $\tau$ close to one signal that real time and pseudo out-of-sample exercises produce similar rankings, values close to zero indicate no association, while values close to negative one imply reversed orderings. The resulting correlation coefficient over time at the one-month, one-quarter and one-year ahead horizon is depicted in Fig. (ref) for the US (a) and the EA (b).\footnote{For the US, recessions are dated by the National Bureau of Economic Research (NBER) Business Cycle Dating Committee and published at \href{https://www.nber.org/cycles/cyclesmain.html}{nber.org/cycles/cyclesmain.html}. For the EA, recessions are dated by the Centre for Economic Policy Research (CEPR) Euro Area Business Cycle Dating Committee and published at \href{https://eabcn.org/dc/chronology-euro-area-business-cycles}{eabcn.org/dc/chronology-euro-area-business-cycles}. Since recessions are only dated on a quarterly basis in the EA, we pick the first month of the respective quarter as the reference period.}
Figure (ref) shows that real time and pseudo out-of-sample information sets produce different rankings in terms of model performance depending on the forecast horizon, and more importantly, specific features of the data releases. Comparing the US and the EA, differences in the release schedule of real time data for the EA lead to substantially more disagreement between real time and pseudo out-of-sample simulations regarding the relative performance of the competing models.
Zooming in on the forecasts for the US, we detect a high concordance of real time and pseudo out-of-sample performance rankings. The rank correlation coefficient is close to one, especially after $2010$. Depending on the forecast horizon, the measure detects subtle differences over time. For one-month ahead forecasts, we observe that the relative ordering seems to change between the two recessions in the holdout, and during the Great Recession. These changes are even more pronounced for the one-quarter ahead forecasts, with Kendall's $\tau$ indicating values close to zero during and just after the Great Recession. Interestingly, for one-year ahead forecasts, differences in the ranking of the models occur mainly between the two recessions.
For the EA, using real time information appreciably changes the relative performance ordering of the competing specifications, especially at short forecast horizons. At the one-month ahead horizon, we observe periods over the full holdout where the model ordering between real time and pseudo out-of-sample information sets is reversed, indicated by negative values of Kendall's $\tau$. During and between the two recessions, the correlation fluctuates around zero. For the one-quarter ahead horizon, the rankings tend to agree with each other early in the holdout. After the European debt crisis, the rank correlation coefficient turns negative, indicating that real time information and pseudo out-of-sample rankings diverge. By contrast, for one-year ahead forecasts, real time and pseudo out-of-sample designs suggest similar rankings of relative performance, with correlations close to one apart from a brief period early in the holdout.
We proceed with investigating the relative forecast performance of the competing specifications in real time and based on the pseudo out-of-sample simulations. Figure (ref) shows the ranking over time, featuring the two best and worst specifications at the end of the holdout for visualization purposes. Starting with the US, panel (a) indicates that small and medium models featuring TVPs without principal components perform well for most of the holdout sample at one-month and one-quarter ahead horizons. Specifications of the baseline model augmented by principal components occupy the lowest ranks. As suggested in the context of our discussion of rank correlations, the best and worst performing models tend to be similar for both information sets in the case of the US. Some differences occur for one-year ahead forecasts. First, we find that the large model featuring TVPs but no principal components outperforms all other specifications for most of the holdout period. Second, models that perform relatively poorly for shorter horizons exhibit significant gains in the Great Recession.
Considering the results for the EA, Fig. (ref)(b) draws a rather different picture. While in a real time environment the large model augmented with principal components featuring TVPs performs best for most of the holdout at the one-month and one-quarter ahead horizon, smaller models with constant parameters and without principle components dominate all other specifications. This finding translates to values of the rank correlation coefficient close to zero or even negative, as described previously. A striking example is provided by the small constant parameter model without principal components, that performs best for one-month ahead forecasts in pseudo out-of-sample simulations, but is second to last when using real time information. For one-year ahead forecasts, a different picture emerges. The larger constant parameter models without principal components perform best, while the models featuring TVPs are ordered last for most of the holdout for both real time and pseudo out-of-sample contexts.
The final part of this subsection assesses in detail how forecast performance differs across models when adopting a real time and pseudo out-of-sample simulation. For this purpose, each real time model is benchmarked against its complementary specification estimated using the pseudo out-of-sample information set. Relative LPS larger (smaller) than zero indicate that predictive accuracy based on the pseudo out-of-sample information set is superior (inferior). These relative differences can be interpreted as measures capturing the distance between real time and pseudo out-of-sample forecasts, and do not necessarily indicate superior forecast performance, since each real time model is benchmarked against its pseudo out-of-sample counterpart.
Figure (ref) shows the resulting relative cumulative LPSs over time. We find that producing forecasts at the one-month and one-quarter ahead horizons using real time information is substantially harder for both the US and EA, indicated by negative relative joint LPSs for all models considered. At the one-month ahead horizon, the distance between the performance metric in real time and pseudo out-of-sample exercise is minimal for medium and large models with TVPs.
Reconsidering Fig. (ref), this implies that such models can be considered comparatively robust to data revisions and imputations, and they tend to perform well in both forecast evaluation contexts. Interestingly, we do not observe clear breaks in relative real time and pseudo out-of-sample performance, but rather consistent differences over time. For one-quarter ahead forecasts, differences are smallest for smaller constant parameter models, and substantial especially for principal component augmented variants with TVPs. At the one-year ahead horizon, we find that some models estimated using real time data outperform their pseudo out-of-sample counterparts, with small to medium constant parameter models showing the most pronounced gains. In contrast to the shorter forecast horizons, clear differences in relative performance emerge during the Great Recession.
A clearer pattern in terms of relative performance using real time and pseudo out-of-sample data is visible for the EA in Fig. (ref)(b). In particular, for all considered forecast horizons, the constant parameter models show the largest differences between the two forecast evaluation designs. While most specifications estimated using real time data perform worse consistently, we observe deteriorating relative performance measures for one-month and one-quarter ahead forecasts especially during the Great Recession. Similar to the US, the picture is different for one-year ahead forecasts. Again, there are several TVP real time models that outperform their pseudo out-of-sample counterparts.
This discussion of overall forecast performance can be summarized by noting three key observations. First, and perhaps unsurprisingly, producing accurate forecasts in real time is substantially harder than when relying on a pseudo out-of-sample forecast exercise. This is mainly the case at short forecast horizons, and differs for longer horizons. Second, real time and pseudo out-of-sample information sets do not necessarily result in the same relative performance ordering among models. This appears to be a significant problem in the case of the EA. Finally, differences occur both in terms of the specifics of the dataset and release schedule of the vintages, and which forecast horizon is considered.
Before turning to a more detailed analysis of differences between forecasts in real time and pseudo out-of-sample exercises for different variables and addressing point and density forecasts, we pick the small model with time-varying parameters but without principal components as an example to demonstrate both data features over the different vintages, and discuss differences in the obtained predictive distribution for the real time and pseudo out-of-sample forecast exercise.
Figure (ref) shows the full dataset of the focus variables over time, with the blue line marking the final vintage. To visualize the vast amount of data captured in all new releases and the imputed values, we proceed as follows. We use the posterior median of the nowcasted observations and calculate the minimum and maximum value of the respective series per period (black lines) across all vintages. The figure thus indicates both data revisions and uncertainty surrounding imputed values.
We start with the US in panel (a). Since historical vintage data starts only in late 1998, data imputation only plays a role past this date. Interestingly, differences across vintages are also visible prior to this date, implying that earlier data are also subject to revisions. This notion is especially important for inflation. After 1998, the main differences originate from the nowcasting scheme, with largest deviations observable during recessionary episodes for inflation and unemployment. Interest rates, on the other hand, are comparatively unaffected by revisions and data imputation.
The same is true for the EA. Data revisions and imputations in the EA, shown in Fig. (ref)(b), play only a minor role for interest rates, evidenced by the narrow bounds surrounding the final vintage. For inflation and unemployment, on the other hand, we find substantial differences of the vintage data over time. The largest changes are visible during the Great Recession starting in 2008, and the European debt crisis between 2011 and 2013. This observation is particularly striking for unemployment, with negative deviations up to 0.2 percentage points between nowcasts and the final series.
To assess out-of-sample features both in a real time and pseudo out-of-sample exercise, we depict the respective one-step ahead predictive distributions alongside the $68$ percent posterior coverage interval in Fig. (ref). The blue shaded area captures pseudo out-of-sample forecasts and the black lines are based on real time information sets.
In Fig. (ref)(a), it is clearly visible that taking a real time perspective yields wider credible sets for the predictive distribution for the US, originating from the need of nowcasting some of the involved quantities. Interestingly, even though interest rates are less affected by data revisions, we also detect differences in the predictive distribution stemming from missing values in the other variables. In general, the benchmark model performs well for forecasting, with the predictive intervals covering the realized series in most cases. Exceptions are various months during recessions, where some values lie outside the credible set.
For the EA (see Fig. (ref)(b)), a key notion again is that the posterior distribution for predictions based on real time data are wider than for pseudo out-of-sample forecasts for all focus variables, even more so than for the US. This increased uncertainty surrounding forecasts stems from both the imputed values and data revisions. While forecasts track the actual evolution of interest rates satisfactorily, we observe that some peaks and troughs of inflation lie outside the bounds of the credible set. An interesting finding for unemployment is that the model consistently predicts lower unemployment rates in real time during the Great Recession, while for the pseudo out-of-sample exercise, the posterior more accurately tracks the evolution of the realized series. The same is visible during the European debt crisis, albeit to a lesser extent.
In Sections (ref) and (ref), we provide a more detailed discussion of differences in forecasts when adopting a real time perspective individually for the US and the EA. In particular, we assess specific findings for point and density forecasts across focus variables and horizons. Additional results to identify the best performing models and differences between real time and pseudo out-of-sample contexts are provided in Appendix (ref).
Figure (ref) shows Kendall's rank correlation coefficient $\tau$ for the forecast performance at different horizons and across variable types based on both point and density forecasts. For point forecasts, we assign ranks based on minimal cumulative absolute forecast errors (FEs), while ranks for predictive densities are constructed from cumulative marginal LPSs in descending order. A striking observation is that rank correlations for density forecasts and point forecasts exhibit different patterns over time across most variable types and forecast horizons.
At the one-month ahead horizon, we find that the relative performance in terms of density forecasts is similar in real time and pseudo out-of-sample contexts for interest rates and inflation, with the rank correlation coefficient close to one for most of the holdout. Here, especially models with TVPs appear to perform well. For unemployment, there is essentially no correlation in rank orders until the Great Recession. During the recovery period, we observe a slightly higher agreement in terms of relative forecast performance for density forecasts. However, this robustness to data revisions and imputations fades slowly in the later part of the holdout, fluctuating around $0.3$ at the end of the considered period. In a real time context, introducing additional information via principal components tends to pay off.
Conversely, for point forecasts, we find almost no association in the rank ordering until the Great Recession for all variables. Afterwards, coinciding with the period when the zero lower bound was reached, the relative orderings tend to agree for interest rates, and to a slightly lesser extent for inflation. The rank correlation coefficients over time visibly show comovement for point and density forecasts of unemployment, albeit the correlation is higher for point forecasts. It is noteworthy that models that perform well for density forecasts do not necessarily perform well for point forecasts, especially for inflation. Interestingly, the large constant parameter model featuring principal components produces accurate point and density forecasts in real time and pseudo out-of-sample contexts for unemployment.
For one-quarter ahead forecasts of interest rates, we find that rank orderings consistently agree for point and density forecasts, although we observe a brief period of lower correlations just after the Great Recession in terms of point forecasts. The relative performance for point forecasts of inflation is rather stable and strongly positively correlated for real time and pseudo out-of-sample exercises. Conversely, we find that although the performance ordering agrees early in the sample for one-quarter ahead density forecasts of inflation, the rank correlation coefficient is zero during the Great Recession. After that, we again observe a clearly positive correlation between real time and pseudo out-of-sample relative performance ranking. Here, TVPs seem to be crucial to provide accurate density forecasts, while best point forecasts are achieved by small scale constant parameter models.
Turning to one-quarter ahead forecasts of unemployment, we observe fluctuating correlations between almost zero and $0.8$, with point and density rank correlations exhibiting a substantial degree of comovement. This relationship decouples during the Great Recession, with high correlations in terms of densities, and zero association between real time and pseudo out-of-sample rank ordering for point forecasts afterwards. Larger TVP specifications without principal components appear to produce the best point and density forecasts for both real time and pseudo out-of-sample simulations, on average.
One-year ahead forecasts appear to be remarkably robust in terms of relative performance orderings. For point forecasts, the rank correlation coefficient is close to one for all three focus variables for most of the holdout, indicating that real time and pseudo out-of-sample exercises produce almost identical rankings of model performance. For density forecasts, we observe a similar situation, although we identify lengthy periods before $2010$ with correlations around $0.5$, especially for the interest rate and inflation. In general, we find that constant parameter models produce superior forecasts at longer horizons, except for density forecasts of inflation.
Variable specific differences for relative orderings of point and density forecast performance in the EA are shown in Fig. (ref). As suggested in Section (ref), taking a real time perspective is even more consequential for relative forecast performance for the EA than the US. This observation is true for both point and density forecasts.
Starting with the one-month ahead horizon, point and density forecasts for interest rates exhibit a similar ordering for real time and pseudo out-of-sample simulations until the Great Recession. Overall, TVPs appear to improve forecast accuracy. In the period between the two recessions, we observe negative rank correlations, indicating that using real time information reverses the ordering of the considered models in terms of their relative forecast accuracy. In the period after the European debt crisis when the zero lower bound was reached, density forecasts are again similar for real time and pseudo out-of-sample information, while this is not the case for point forecasts. This is mainly due to the relatively poor performance of smaller constant parameter models augmented with principal components in the pseudo out-of-sample simulation, specifications that appear to perform satisfactorily in real time.
A different picture emerges for short-horizon inflation forecasts, with no clear patterns in terms of rank correlation, albeit changes in relative performance orderings appear to occur especially during and between the two recessions. While larger models with TVPs and without principal components perform well for density forecasts, small to medium size models with constant parameters indicate superior performance for point forecasts.
For unemployment, similar as in the US, we detect comovements in correlations over time for point and density forecasts. Point forecast accuracy orderings between real time and pseudo out-of-sample contexts tend to be concordant after the Great Recession. Density forecasts, however, exhibit substantially lower and decreasing correlation measures over time, with values close to zero at the end of the holdout. Large fluctuations for inflation and unemployment early in the sample may be explained by the relatively short initial estimation period. No clear best performing specification in real time and pseudo out-of-sample contexts can be identified, although constant parameter models appear to be the most robust.
Interest rate density forecasts for the one-quarter horizon show similar concordance as in the one-month ahead case, again with gains for models featuring TVPs. Point forecasts mostly agree on model ordering, different to the shorter horizon. Principal components seem to improve forecast accuracy in both the TVP and constant parameter cases. For inflation, we observe higher correlations especially after the Great Recession, indicating that pseudo out-of-sample exercises establish a similar model ordering when compared to relying on real time information. Here, TVP models are superior for density forecasts, while constant parameter models are superior for point predictions.
Interestingly, for unemployment, correlation measures show a similar path as for the US. Early in the holdout, the measure fluctuates substantially, increasing during the Great Recession. This implies that for density forecasts, both real time and pseudo out-of-sample simulations produce similar orderings of relative performance. The small constant parameter model featuring principal components performs relatively similar in both evaluation contexts, while TVP models are ranked last in both cases. Point forecast correlations deteriorate afterwards, and are close to zero or even negative after the first recession.
Long-horizon forecasts exhibit different patterns for rank correlations. For interest rates, consequences for the relative performance orderings derived from real time and pseudo out-of-sample simulations are present but muted, indicated by a coefficient that fluctuates just above $0.5$ for most of the holdout. TVPs play an important role for point and density forecasts both using real time and truncated vintage data. Inflation forecast performance orderings are discordant for point forecasts early in the sample, however, tend to agree on the rank order after the Great Recession. Small scale constant parameter models featuring principal components perform well. Density forecasts, on the other hand, tend to be ordered similarly early in the sample. This changes after the Great Recession, with negative $\tau$ suggesting inverse performance orderings of the competing models. At the end of the holdout sample, the association between performance orderings are close to zero. The best performing model in real time here is the large TVP-VAR without principal components, while the medium scale constant parameter model works best in the pseudo out-of-sample exercise.
For unemployment, we find consistently high values of correlations throughout the holdout period in terms of density forecasts. This implies that pseudo out-of-sample evaluation is sufficient to establish a useful ordering of model performance. Regarding point forecasts, we observe some periods until $2010$ where substantial rank changes occur. From $2010$ onwards, Kendall's rank correlation is close to one, again indicating that relying on real time information does not affect relative performance orderings. Here, constant parameter specifications appear to be superior for point and density forecasts in both information sets.
In this paper, we systematically assess differences in forecast performance between a set of differently sized models using pseudo out-of-sample versus real time simulations. We rely on constant and TVP-VAR models with SV, equipped with a global-local shrinkage prior. We also consider variants augmented by principal components to capture high-dimensional information from available datasets, and discuss imputing missing values. Our results suggest differences in the relative ordering of model performance for point and density forecasts. No clearly superior specification for the US or the EA across variable types and forecast horizons can be identified, although larger models featuring TVPs appear to be affected the least by missing values and data revisions. This finding suggests that pseudo out-of-sample simulations are insufficient to establish a clear ordering of relative model performance, and attention in the development of new methods should always be paid to real time features of data in order to be of value for forecasters in central banks and governments.
{\setstretch{0.85} \addcontentsline{toc}{section}{References} }