Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
64,519 characters · 12 sections · 71 citation commands
Regime-Switching Density Forecasts Using Economists' Scenarios
In recent years, it has become increasingly important for economic agents and policymakers to take account of the uncertainty around macroeconomic outlooks. In particular, two distinct approaches to this issue have emerged. The first one consists in generating economic forecasts in the form of (continuous) probability distributions, or density forecasts. This is now common practice among forecasters (elliottandtimmermann2016), and many economic agents, such as financial institutions, routinely evaluate their potential losses as random draws from predictive density functions (e.g., for the computation of value at risk and expected shortfall, see jorion2006). The second approach is to define a small number of (discrete) macroeconomic scenarios (e.g., see moodys2018). This approach facilitates communication to the public regarding economic uncertainty and finds important applications in financial supervision and risk management, most notably in the design of bank stress tests, which are now integral part of the financial regulatory framework in advanced economies (e.g., fed2018).
This paper develops a novel forecasting approach that combines these two perspectives. Using a regime-switching model, we construct macroeconomic density forecasts that explicitly incorporate information on discrete scenarios defined by economists. We consider alternative sets of scenarios, which we label as views, and use them to define Bayesian priors on economic regimes in the model. Next, as different views result in different density forecasts, we combine these forecasts in a way that optimizes standard evaluation criteria of density forecasting. The approach is illustrated through an empirical application to U.S. GDP growth forecasts, in which we exploit the information contained in the macroeconomic scenarios defined by the Federal Reserve for its bank stress tests. We show that the approach achieves good forecast accuracy and good calibration of forecast distributions. Moreover, we find that its forecast performance is further improved when we combine it with a beta transformation of density forecasts (GneitingRanjan2010,GneitingRanjan2013).
On the one hand, the proposed approach is able to produce highly flexible predictive distributions. Such flexibility results both from the regime-switching framework and from the forecast combination procedure. Density forecasts from regime-switching models are mixture distributions, by construction: they are weighted averages of regime-specific densities, where the weights consist in the probabilities of the economy ending up in the various regimes. Likewise, density forecast combinations also produce mixture distributions. Thus, the approach can easily accommodate a wide range of departures from normality, as mixture distributions are in general non-normal even if their individual components are. This is an important property when it comes to forecasting macroeconomic variables, whose empirical distributions often deviate from Gaussianity. In particular, a number of studies, including fagioloetal2008, curdiaetal2014, and ascari_fagiolo_roventini_2015, have documented a non-normal distribution of the GDP growth rate in the U.S. and other advanced economies, mainly as a result of large downturns. acemogluetal2017 find similar results and propose a theoretical model in which systematic departures of output from the normal distribution are explained by the interaction of idiosyncratic microeconomic shocks and sectoral heterogeneity. Compared to econometric approaches that impose non-normal distributions of errors (e.g., hansen1994), regime-switching models allow for a more transparent economic explanation of non-normality, based on the transition between different regimes.
On the other hand, regime-switching models appear as a natural framework for dealing with discrete scenarios such as those used in bank stress tests. First of all, these scenarios are typically constructed in a way that directly relates to economic regimes. For instance, the Fed defines its scenarios as “sets of economic and financial conditions", with baseline scenarios reflecting the most likely conditions and adverse/severely adverse scenarios reflecting conditions that prevail in recessions (fed2017, Appendix A of Part 252).\footnote{ As explained in fed2017 (Appendix A of Part 252), “the baseline scenario is designed to represent a consensus expectation" and is constructed using the forecasts of private sector forecasters (e.g., Blue Chip Consensus Forecasts and the Survey of Professional Forecasters), government agencies, and other public-sector organizations (e.g., the International Monetary Fund and the Organization for Economic Co-operation and Development). To develop the severely adverse scenario, the Fed relies on a recession approach, i.e., it specifies the “future paths of variables to reflect conditions that characterize post-war U.S. recessions". The severity of recessions is mainly determined by the level and increase of the unemployment rate, as well as the occurrence of a housing recession. In post-war U.S. recessions that are classified as severe, GDP has dropped about 3.5 percent and unemployment has increased by a total of about 4 percentage points, on average (fed2017, Appendix A of Part 252, Table 1). The adverse scenario reflects “a set of economic and financial conditions that are more adverse than those associated with the baseline scenario but less severe than those associated with the severely adverse scenario". } Also, many contributions in the literature on stress testing call for an important role of regime-switching models in the design of macro scenarios, arguing that realistic scenarios should reflect the non-linearities that derive from the state dependency of macro-financial conditions (adrianIMF2020, hanleika2019, borio2014, gross_henry_rancoita_2022, BidderMcKenna2015). As highlighted by hamilton2016, Markov-switching models provide a sufficiently parsimonious and yet robust characterization of the transition into and out of recession regimes, at least for the U.S. economy.
This paper contributes to several strands of the forecasting literature. First, it relates to other papers that incorporate external --possibly judgmental-- forecasts into econometric model-based forecasts. Kruegeretal2017 combine forecasts from Bayesian VARs with forecasts from other sources using entropic tilting, a method for modifying a distribution in a way that satisfies specific moment conditions (Robertsonetal2005). FaustWright2009 use the Federal Reserve Board's Greenbook forecasts as data in autoregressive and factor-augmented autoregressive models, in order to produce forecasts of GDP growth and inflation. SchorfheideSong1025 and Wolters2015 also use nowcasts from the Fed's Greenbook as data to generate forecasts from a Bayesian VAR (BVAR) and DSGE models, respectively. The main difference between our approach and these approaches is that we exploit information on multiple economic scenarios from external sources, mapping them into different regimes, while these approaches only target point forecasts or forecast variance.
Second, the paper relates to other contributions on forecasting using regime-switching models. The most closely related approaches are Bayesian ones, such as those by pesaranetal2006, who use a break point model (a generalization of regime-switching models) with hyperparameter uncertainty, and by bauwensetal2017, who estimate a Markov-switching model with an unknown and potentially infinite number of regimes. Unlike these studies, our paper exploits experts' views to define priors on economic regimes. A broad literature has investigated the forecast performance of regime-switching models. While the available evidence on the accuracy of point forecasts is mixed (elliottandtimmermann2016), these models have proved very useful for density forecasting and forecasts of tail events. chauvetpotter2013 show that Markov switching models can achieve high accuracy in forecasting GDP, especially with respect to the timing and depth of recessions. gewekeandamisano2011 show the usefulness of Markov mixtures for density forecasts of the stock market. bauwensetal2017 use an infinite Markov-switching autoregressive moving average model to produce density forecasts of GDP. alessandrimumtaz2017 find that a Threshold VAR in which regime shifts depend on financial conditions produces good density forecasts of U.S. GDP during the Great Recession.
Third, the paper relates to the large literature on forecast combinations. In particular, it has been shown that gains in density forecast performance are often achieved by combining different predictive distributions. Contributions in the field of macroeconomic forecasting include hallandmitchell2007, gewekeandamisano2011, and ganics2017, among others. A vast body of literature in other fields, such as management science, risk analysis, and meteorology, also deals with the combination of probabilistic forecasts from experts into “consensus" distributions (e.g., GenestZidek1986; clemenwinkler1999). For comprehensive reviews, see aastveitetal2018 and wangetal2022. Several ways to further enhance density forecasts of non-normal variables have been proposed in the literature on forecast combinations, such as beta opinion pools (GneitingRanjan2010,GneitingRanjan2013) and empirically-transformed opinion pools (garrattetal2023), which will be explored in the empirical part of the paper.
We combine forecasts from different experts' views using two main criteria of density forecast evaluation. The first one is the sum of log scores, which measures the ability to assign high probabilities to outcomes that are truly likely to be observed. The second is a uniformity test on the probability integral transforms (PITs) of the forecasts, which measures the degree of calibration of the forecast distribution (dieboldetal1998).\footnote{A well-calibrated forecast is one that does not make systematic errors: if p is the predicted probability assigned to a given random event, then that event should empirically occur with frequency p.} Both measures have been used in the literature to compute “optimal" forecast combinations (hallandmitchell2007; gewekeandamisano2011; ganics2017).
The remainder of the paper is organized as follows: Section (ref) explains the methodology, Section (ref) presents the empirical application, and Section (ref) concludes.
Our reference model is a Markov-switching autoregressive (MSAR) model in which the intercept and the error variance depend on the unobserved state of the economy. Let $y_t$ denote a macroeconomic variable of interest at time $t$. The MSAR can be written as:
where $S_t$ denotes the unobserved state variable at time $t$, $\beta_{S_t}$ is the intercept in regime $S_t$, $\alpha_j$ for $j=1,\dots,p$ is a state-independent autoregressive coefficient, $p$ is the maximum lag, $\varepsilon_t$ is the error term and $\sigma^2_{S_t}$ is the regime-dependent variance of the error.\footnote{We assume state-independent autoregressive coefficients (see, e.g., hamilton1989) to avoid the overfitting problems that easily arise in macro applications of regime-switching models. In particular, as explained in hamilton2016, “inference about parameters that only show up in regime $i$ can only come from observations within that regime. With postwar quarterly data that would mean about 50 observations from which to estimate all the parameters operating during recessions. One or two parameters could be estimated fairly well, but overfitting is again a potential concern in models with many parameters. For this reason researchers may want to limit the focus to a few of the most important parameters that are likely to change, such as the intercept and the variance".
Also, as pointed out by hamilton2016, while there is no theoretical reason to assume that all regime-specific densities are normal, the same concerns of overfitting suggest using a normal distribution, which is defined by just two parameters. The assumption of normality is generally made in the literature on regime-switching models, with very few exceptions, such as dueker1997. Some empirical evidence may be used to corroborate this assumption. For instance, acemogluetal2017 find that, once large downturns are excluded from the data, the U.S. GDP growth distribution appears to be well approximated by a normal. } In particular, $S_t$ is a Markov chain characterized by a transition matrix $\boldsymbol{\xi}$, where the element $\xi_{kj}$ in row $k$ and column $j$ represents the probability of transition from state $k$ to state $j$:
with $k,j=1,\dots,K$, where $K$ is the number of regimes in the economy. Accordingly, the MSAR captures the typical autocorrelation of macro variables in two ways: by means of the autoregressive coefficients in (ref) and through the persistence in the state variable $S_t$, as expressed by the transition matrix. Finally, let $\boldsymbol{\vartheta}$ denote the vector of parameters of the MSAR model, i.e. $\boldsymbol{\vartheta} = (\beta_1,\dots,\beta_K,\sigma_1,\dots,\sigma_K,\alpha_1,\dots,\alpha_p,\boldsymbol{\xi})$.
The MSAR model (ref) is estimated using Bayesian methods (see the Appendix for details). The priors on the parameters follow conventional distributions (fruehwirthschnatter2006), which are:
for $j=1,\dots,p$, where $\mathcal{N}$ and $\mathcal{G}^{-1}$ denote Normal and inverse Gamma distributions, respectively, and $a_{j,0},A_{j,0}, b_{0,k},B_{0,k},$ $c_0,C_0$ are hyperparameters to be selected by the researcher. In addition, for the transition matrix $\boldsymbol{\xi}$ it is assumed that the rows are independent and each row follows a Dirichlet distribution, denoted by $\mathcal{D}$, i.e.:
where $e_{k1},\ldots,e_{kK}$ are hyperparameters, for $k=1,\dots,K$.
We define a view as a set of assumptions concerning the number of regimes $K$ and the model parameters, i.e., a specific set of values for $K$ and for the hyperparameters of the model's prior.
Views can incorporate information on discrete economic scenarios defined by experts. In particular, we consider priors in which each regime is “centered" on a corresponding scenario using the following rule. Consider an AR($p$) model where the coefficients are given by the $k$-state-specific prior hyperparameters, i.e.:
In this model, the unconditional expectation of $y_t$, which we denote by $E_k \left( y_t \right)$, is:
Then, given an assumption on the state-independent autoregressive coefficients $a_j$, each regime-specific hyperparameter $b_{0,k}$ is chosen in such a way that expectation (ref) matches a specific value derived from a scenario provided by external sources. The hyperparameter $B_{0,k}$ determines the tightness of the prior around these expectations or, in other words, the strength of economists' views. In principle, all the other hyperparameters can also be set to values suggested by economists, if available. Otherwise, uninformative or diffuse priors can be used for the remaining parameters of the MSAR.
When several alternative views are considered, Bayesian averaging across views can be performed. Let $\boldsymbol{\vartheta}_{K,i}^0$ denote the generic $i$-th view assuming $K$ states and let $\boldsymbol{\pi}^{0}$ be a vector containing discrete prior probabilities assigned to all views considered. The posterior probability of view $\boldsymbol{\vartheta}^0_{K,i}$, which we denote as $\pi_{K,i}$, will depend on the prior probability vector $\boldsymbol{\pi}^0$ and on the values of the marginal likelihood of the MSAR model associated with the different views. Note that the letter $\pi$ is used throughout the text to denote discrete probability distributions.
Further details are provided in the Appendix.
Let us define the vector containing all observations up to time $t$ as $\mathbf{y}_t$, i.e., $\mathbf{y}_t=(y_1,\dots,y_t)$. Also, let us assume that the current time period is $T$ and the forecast horizon is one period. Finally, let us denote the density forecast produced by any given view as $ p\left(y_{T+1}|\mathbf{y}_{T},\boldsymbol{\vartheta}_{K,i}^0 \right)$ (see the Appendix A.2 for details on how density forecasts are calculated), and let us assume that, for any given number of states $K =1, \dots, \overline{K}$, a number $P_K$ of alternative views are available. We consider two methods for pooling forecasts across different views:
We find the “optimal" priors $\boldsymbol{\pi}^0$ and weights $\mathbf{w}$ by maximizing two alternative objective functions, based on statistics that are commonly used to evaluate density forecast performance:
Both the optimization based on $f_1$ and the one based on $f_2$ are solved numerically. For each $f_i$, with $i=1,2$, the optimization algorithm delivers two vectors at time $\tau$: the vector of optimal forecast weights $ \mathbf{w}^\ast_{i,\tau}$ for the set of alternative views, i.e.:
and the vector of optimal prior probabilities $\boldsymbol{\pi}^{0\ast}_{i,\tau}$:
The former represents the typical problem explored in the literature on density forecast combination, whereas the latter can be seen as an empirical method for eliciting priors in the context of Bayesian model averaging. The optimal prior $\boldsymbol{\pi}_{i,\tau}^{0\ast}$ represents the discrete prior probability distribution of views such that the resulting posterior $\boldsymbol{\pi}_{i,\tau}^\ast$, when used as a vector of forecast weights, maximizes the density forecast performance, based on the selected objective function. In practice, the main difference between (ref) and (ref) is that the first problem directly delivers weights for forecast combination, while in the second case the actual forecast weights are the posterior probabilities, and so will also depend on the marginal likelihoods of all views, i.e. $p(\mathbf{y}|\boldsymbol{\vartheta}^0_{K,i}) $ $\forall K,i$ (see the Appendix).
This section illustrates the approach through an empirical application to density forecasts of U.S. GDP growth. We use real-time quarterly data from 1948Q1 to 2019Q4\footnote{Source: U.S. Bureau of Economic Analysis, Real Gross Domestic Product (series code: GDPC1). We use the complete set of real-time data vintages provided in the Archival FRED (ALFRED) database compiled by the Federal Reserve Bank of St. Louis. The first vintage was released in December 1991. For all sample windows ending before 1991Q4, we use the data contained in the first available vintage. As vintages are generally released at a monthly frequency, for each quarter we take the last vintage released within that quarter. } and consider the year-on-year growth rate of real GDP (expressed in percentage points in what follows).\footnote{Our choice of using the year-on-year (YoY) rather than the quarter-on-quarter (QoQ) growth rate of GDP in the MSAR model follows the same reasoning as in bindergross103: since the YoY rate is more persistent than the QoQ rate, changes in the mean and variance of the YoY rate are more likely to reflect the transition between different economic regimes compared to changes in the QoQ rate, which are more affected by noisy events specific to a given quarter. } Over this period, the (unconditional) distribution of GDP growth is non-normal (the Jarque-Bera test rejects normality at the 1% level of significance) and, in particular, fat-tailed (kurtosis is close to 4), in line with the findings of prior literature (e.g., fagioloetal2008; ascari_fagiolo_roventini_2015).
We set the lag length $p$ of the model to 5, in consideration of the quarterly frequency of the variable. As explained below in more detail, we first produce pseudo-out-of-sample forecasts using a recursive window scheme. Then, we calculate optimal weights\ $\mathbf{w}^\ast$ and priors $\boldsymbol{\pi}^{0\ast}$ over time, following the procedure from section (ref). Finally, we assess the performance of pooled forecasts on an evaluation sample, i.e., using observations of the target variable that have not been used in the optimization procedure.
We consider a total of 13 alternative views on U.S. GDP growth regimes. Eight views impose strongly informative priors derived from the scenarios of the Fed stress tests 2015-2018.\footnote{The scenarios are available https://www.federalreserve.gov/supervisionreg/dfast-archive.htm.} The remaining five views are “vague" views, defined by imposing uninformative priors on MSAR parameters. We consider different assumptions on the number of regimes $K=1,2,3,4,5$.\footnote{Our choice of the maximum number of regimes is supported by the results of bauwensetal2017. These authors develop a model that allows for an infinite number of regimes, using a nonparametric Dirichlet process. However, when estimating the model for U.S. GDP growth (with different breaks for the mean and variance parameters), they find that the posterior probability of the number of regimes being at most 5 lies between 98% and 100% for the mean parameters and between 74% and 100% for the variance, depending on the prior used for estimation.}
Let us first consider the Fed-based views. For each of the four stress tests under consideration, two views are constructed, one with $K=3$ and the other with $K=5$. In the view with $K=3$, one of the regimes (which may be called the “normal times" regime), is derived from the Fed baseline scenario, another (“adverse regime") from the adverse scenario and the last one (“severely adverse regime") from the severely adverse scenario.\footnote{Although the Fed stress scenarios represent hypothetical paths and not forecasts, they are intended to be plausible even when severe. Therefore, they can legitimately be assigned predictive probabilities (see e.g., Yuen 2013) and used to form density forecasts.} In particular, we “center" each regime on the corresponding Fed scenario using the rule described in section (ref). Specifically, we consider an AR(5) model where the coefficients are given by the $k$-state-specific hyperparameters of the prior $\boldsymbol{\vartheta}^0_{K,i}$, so equation (ref) implies $ E_k \left( y_t \right) = b^{(K,i)}_{0,k} / ( 1- \sum_{j=1}^{5} a_j^{(K,i)} ) $. Then, after making an assumption on the state-independent $a_j^{(K,i)}$, with $j=1,\dots,5$, each regime-specific $b^{(K,i)}_{0,k}$ is chosen in such a way that this expectation matches a specific value derived from the relevant scenario of the Fed stress test. For the normal times regime, this value is the average growth rate in the last four quarters of the baseline scenario, which is assumed to be close to the convergence value of the year-on-year growth rate in the absence of shocks.\footnote{The stress test scenarios are defined in terms of annualized quarter-on-quarter growth rates, so that averaging over the last four quarters approximates the year-on-year growth rate in the last quarter.} For both the adverse and the severely adverse regimes, the value to be matched is the average growth rate in the first four quarters of the corresponding scenario, as the first quarters are those when the negative shocks are assumed to occur and the growth rates are lowest.
An example may help. Let us consider the view with $K=3$ derived from the 2018 Fed stress test. The average growth rate of GDP in the last four quarters of the baseline scenario is 2.1%, while the average growth rates in the first four quarters of the adverse and severely adverse scenarios are -2.125% and -6.275% respectively. As customary in the related Bayesian literature (e.g., alessandrimumtaz2017, fruehwirthschnatter2006), we set the prior on the autoregressive component of the model using a parsimonious specification. In particular, we set the prior mean of the autoregressive coefficients to the (approximate) OLS estimate of an AR(1), i.e., to 0.9 for the first lag and 0 for the higher-order lags (recall, however, that we allow up to 5 lags in the posterior estimates of the model). Then, $\sum_{j=1}^{5} a_j^{(K,i)} = 0.9$. Accordingly, the prior means for the regime-specific intercepts are set to $b_{0,1} = 2.1/(1-0.9)=0.21$ for the normal times regime, $b_{0,2} = -2.125/(1-0.9)=-0.2125$ for the adverse regime and $b_{0,3} = -6.275/(1-0.9)=-0.6275$ for the severely adverse regime.
The four stress test-based views with $K=5$ expand the views with $K=3$ by adding two regimes: a regime which we may call “recovery from adverse shock", designed to match the last four quarters of the adverse scenario, and a regime of “recovery from severely adverse shock", which matches the last four quarters of the severely adverse scenario. This is done in consideration of the fact that growth rates in the last four quarters of the Fed's adverse and severely scenarios are assumed to be higher than the baseline rates, implying a rebound of the economy after a negative shock. Of course, such regimes may be more generally interpreted as “favorable regimes" characterized by positive shocks and not necessarily as recoveries from recessions.
In the five vague views, all priors on the intercepts are centered on 0 and have a variance of 1 percentage point, while the priors on the autoregressive coefficients are centered on 0.5 for the first lag, on 0 for the higher-order lags, and have a variance of 1. Taken jointly, these assumptions imply a large prior variance on the regime-specific means of the GDP growth rate. Conversely, in the Fed-based views, the priors for both $\beta$ and $\alpha$ are strongly informative, so as to ensure that the regime-specific means are tightly centered on the stress test values. In particular, both priors are assumed to have minimal variance, equal to $ 10^{-5}$. For the autoregressive coefficients $\alpha$, the prior mean is assumed to be 0.9 for the first lag and 0 for higher-order lags, as mentioned in the previous example.
No strong assumption is made regarding the regime-switching error variance $\sigma_k^{2}$. Instead, a diffuse hierarchical prior is assumed in all views.\footnote{ Specifically, a Gamma hyper-prior is defined for $C_0$:
To make the prior on $\sigma_k^{2}$ diffuse, the following values are selected for the hyperparameters: $c_0=3$, $g_0=0.5$ and $G_0=0.5$. These imply that $\sigma_{k}^2$ has a prior expected value of 0.5 percentage points of GDP and a high prior variance of 1.25 percentage points (see Appendix A.3). }
Finally, the hyperparameters for the $k$-th row of the transition matrix $ \boldsymbol{\xi}$ are set in such a way that the (prior) expected probability of remaining in the same state $k$ in the next period is $ \text{E}(\xi_{kk}) = 2/3$ regardless of the number of regimes $K$, while the probability of moving to a different, specific state $j$ decreases with the number of regimes, $ \text{E}(\xi_{kj}) = 1/[3(K-1)]$ (see fruehwirthschnatter2006).\footnote{ Specifically, the hyperparameters for the $k$-th row of the transition matrix $ \boldsymbol{\xi}$ are $e_{kk} = 2$ and $e_{kj}=1/(K-1)$ if $k \neq j$, $\forall k,j$. Given the properties of the Dirichlet distribution, $ \text{E}(\xi_{kj}) = e_{kj}/(\sum_{l=1}^{K}e_{kl})$. }
The summary of the alternative views is provided in Table (ref), where views 1-5 are the vague ones while views 6-13 are those derived from the Fed stress tests 2015-2018. Table (ref) (in the Appendix) reports the GDP scenarios of the Fed stress tests.
We generate a sequence of pseudo-out-of-sample density forecasts using a recursive-window estimation scheme.\footnote{In this context, the choice of using expanding windows, as opposed to rolling windows, increases the probability that the variable “visits" most or all the regimes within the sample, thus greatly facilitating the estimation of the regime-switching model.} Next, the forecasts are used to carry out the optimization of weights/priors, which is iterated over time. The procedure can be described as follows. Let us assume that we are at time $T_w$ and the forecast horizon is $h$. For each view under consideration, the MSAR model is recursively estimated using observations between an initial time $t_0$ and time $t$, with $t=T_0,T_0+1,\dots,T_w-h$. $T_0$ is therefore the end period of the shortest estimation sample. Estimates at $T_0$ are used to make forecasts for period $T_0+h$, estimates at $T_0+1$ are used to make forecasts for $T_0+1+h$, and so on. Thus, at time $T_w$, a sequence of past forecasts is available for each view. At this point, the algorithm computes the optimal weights/priors based on the last $R$ forecasts, i.e., maximizes the relevant objective function between $T_w-R+1$ and $T_w$. Once the optimal weights/priors are retrieved, they are used to combine the different view-specific forecasts for the future period $T_w+h$, which is out of the optimization sample. When the actual value of the variable of interest is observed, at time $T_w+h$, the performance of the combined forecast is measured. The index $T_w$ runs from $T_0+h+R-1$ to $\overline{T}+h$, where $\overline{T}$ is the end of the largest estimation sample. $\overline{T}+2h$ is the last available observation for the target variable. Therefore, the period from $T_0+2h+R-1$ to $\overline{T}+2h$ defines the evaluation sample, or test set. Figure (ref) summarizes the procedure, which closely follows ganics2017.
More specifically, the application to U.S. GDP growth sets $t_0$=1948Q1, $T_0$=1967Q4, $R$=40 quarters, $h$=1 quarter and $\overline{T}$=2019Q2. Accordingly, the evaluation sample runs from 1978Q1 to 2019Q4.\footnote{We estimate the MSAR model using the MATLAB package bayesf Version 2.0 by fruehwirthschnatter2008. For each MSAR estimate, the Markov Chain Monte Carlo (MCMC) algorithm uses 1000 iterations as burn-in and 1000 iterations to store the results. Starting from the sample of forecasts produced by the MCMC algorithm, a complete probability density function is fitted using a normal kernel smoothing function.} We focus on the forecast horizon $h=1$ because, as previously mentioned, the result of PIT uniformity of well-behaved forecasts, used the optimization, only holds for $h=1$ (see RossiSekhposyan2014).
Table (ref) shows the performance of our composite regime-switching forecasts using optimal forecast weights and optimal priors, and compares it with benchmark approaches. We label our forecasts as scenario-augmented MSAR (SA-MSAR) forecasts. As mentioned in section (ref), weights $\mathbf{w}_{1}^{*}$ and priors $\boldsymbol{\pi}_{1}^{*}$ result from the optimization based on the sum of log scores, while $\mathbf{w}_{2}^{*}$ and $\boldsymbol{\pi}_{2}^{*}$ are obtained by maximizing the p-value of the Kolmogorov-Smirnov (KS) test of uniformity for the PITs. The first benchmark considered is a simple AR model (corresponding to view 1 in Table (ref)). Next, we consider three models that allow for non-normal and heteroskedastic errors: an AR with Student-$t$ errors, an AR with ARCH errors and an AR with GARCH errors. As with the MSAR models, the lag length for the AR component is set to 5 for all models, while the ARCH and GARCH components have a lag length of 1. The remaining two benchmarks consist in uniform combination schemes over the alternative views, assigning equal forecast weights/prior probabilities to different values of $K$ and, for any given $K$, equal weights/probabilities to the alternative views defined using $K$ regimes.\footnote{For instance, in the case of equal prior probabilities, it is assumed that $\pi_K^0=1/\overline{K}$ for each $K$ and that $\pi(\boldsymbol{\vartheta}_{K,i}^0|K)=1/P_K$ for each view $\boldsymbol{\vartheta}_{K,i}^0$.}
The table shows the average predictive density (APD), i.e., the average of the exponential of log scores, and the p-value of the KS test.
The results indicate that our optimized SA-MSAR forecasts achieve good forecast accuracy and good calibration of forecast distributions. When optimized by log scores, the SA-MSAR forecasts outperform all benchmarks in terms of APDs. When optimized by PITs, they lead to non-rejection of the null hypothesis of uniformity at the 5% level in the KS test, indicating a reliable specification of the conditional predictive distribution of GDP growth. Thus, optimized SA-MSAR forecasts forecasts achieve well-behaved PITs.\footnote{ Since correct calibration implies that the PITs are realizations of i.i.d U(0,1) variables, we also test for the serial independence of the PITs. Specifically, following RossiSekhposyan2014, we perform a Ljung–Box test of independence for both the first and the second moment of the PITs. Both tests do not reject the null hypothesis of serial independence, implying that forecasts are well-calibrated also in this respect.} By contrast, all benchmarks lead to rejection of the PIT uniformity hypothesis.
The approach can be used to evaluate the time-varying contribution of different views to the composite forecasts. Figures (ref)-(ref) display the evolution over time of the optimal forecast weights and of the posterior probabilities resulting from the optimal priors. In each figure, the area chart in the left panel shows the time-varying weights for all views from 1978Q1 to 2019Q4. The right panel plots the cumulative weight assigned to the views derived from the Fed supervisory scenarios. Figures (ref) and (ref) show the results of the optimization based on log scores, while Figures (ref) and (ref) show the results of the optimization based on the PITs. As can be seen from Figures (ref) and (ref), the vague views tend to dominate in the case of log-score optimization. In terms of optimal weights $\mathbf{w}_1^{*}$, the cumulative weight of the Fed-based views lies in the range 7%-42% until 1990 and is zero afterwards. As for optimized posteriors, the Fed-based views provide non-zero contributions only between 1982 and 1987. Overall, the results indicate a minor role of Fed-based views in boosting density forecast accuracy.
When the PIT-based optimization is considered, the contribution of the Fed-based views is much higher. On average, they account for about 35% of the combined forecasts in the case of optimal weights and 27% in the case of posterior probabilities. In particular, they dominate in the period following the Global Financial Crisis, in terms of both $\mathbf{w}_2^{*}$ and $\boldsymbol{\pi}_{2}^*$. The also provide a substantial contribution in the early 1980s and during the 1990s. It is important to remark that using a single view is not sufficient to achieve well-calibrated forecasts. None of these views, when considered individually, leads to non-rejection of the PIT uniformity hypothesis in the KS test. Instead, the combination of different views is what drives the good results in terms of calibration.
To illustrate the outcome of our procedure in more detail, Figure (ref) shows density forecasts for a specific quarter, 2016Q4, taken as an example. The figure displays three probability density functions (PDFs): the density forecast from an AR model (red line), the forecast generated by a 3-regime MSAR model using a prior derived from the Fed stress test scenarios (green line; view 9 in Table (ref)), and the optimized MSAR forecast (blue line), i.e., the final outcome of our procedure, using combination weights $\mathbf{w}_{2}^{*}$. Three dashed vertical lines indicate the three discrete scenarios for 2016Q4 contained in the Fed's 2016 stress test. The figure provides an example of what can be obtained from our approach, compared to what is obtained from the single use of the two continuous/discrete approaches to forecast uncertainty. The (normal) density forecast of the AR model is a basic example of the classical density forecast outcome. The predictive density from the 3-regime model is an example of scenario-augmented forecasts: it has a highly non-standard left-skewed profile, clearly reflecting a mixture of three different regime-specific normals, strongly influenced by the tight prior centered on external scenarios. The final (blue) forecast is an example of the “optimal" combination of different densities (including the red and green ones reported in the figure). It has an irregular and leptokurtic shape with two bumps in the tails, merging different forecasts as well as the prior information provided by external views. Finally, for comparison, the figure also reports a histogram representing the probability distribution of the annual growth rate provided in 2016Q3 (the forecast origin) by the Survey of Professional Forecasters.\footnote{ Note that the SPF does not provide a continuous density function of the quarterly GDP growth rate, but only probabilities associated with discrete intervals of annual GDP growth (e.g., the probability of that annual GDP growth in 2016 will be between 1% and 2%). } Our “optimal" density forecasts are broadly in line with the SPF histogram, but of course allow for a more detailed characterization of the forecast distribution.
Next, we examine the behavior of density forecasts across different parts of the GDP distribution. To this aim, Figure (ref) plots the entire empirical distribution of the PITs, along with the 95% confidence intervals of the uniform distribution, as calculated by RossiSekhposyan2017, which account for sample uncertainty.\footnote{For this analysis, we use Matlab code shared by RossiSekhposyan2017. As mentioned by AdrianBoyarchenkoGiannone2019, confidence bands should be taken as approximations when forecasts are computed using an expanding estimation windows.} The figure shows both the cumulative distribution function (CDF) and a histogram of the PITs (normalized), with deciles along the horizontal axis. For perfectly calibrated forecasts, the empirical CDF of the PITs would lie on the 45-degree line, and the histogram bars would all have a height of 1. We report the empirical distribution of the PITs associated with the optimized SA-MSAR forecasts using combination weights $\mathbf{w}_{2}^{*}$. As the figure shows, the empirical CDF always lies within the 95% confidence interval of the uniform distribution, in line with the result of the KS test in Table (ref). In the histogram, all bars are within the confidence bands, except for the highest decile. This indicates that the density forecasts are well-behaved in most parts of the GDP growth distribution. Focusing on the tails, we note that the bar at the lowest decile lies on the lower bound of the 95% confidence interval, suggesting that the forecasts capture left tail risk quite adequately. The bar at the highest decile lies below the lower bound, indicating some over-dispersion of the forecasts in the upper tail.
Lastly, we extend our approach by combining it with two types of density transformations that have been proposed in the literature to further improve density forecast performance: the beta-transformed linear pool (BLP; see GneitingRanjan2010,GneitingRanjan2013) and the empirically-transformed linear opinion pools (ETLOP, garrattetal2023). The BLP applies a transformation of the conditional forecast densities using the beta distribution. This can improve the calibration and, in particular, the dispersion of density forecasts produced by simple linear combinations. The ETLOP transforms the density forecasts so that it matches the empirical cumulative distribution function of the target variable (GDP growth, in our case), in order to better accommodate non-Gaussianity.
In the case of BLP, we select the values of the two parameters of the beta distribution using the same optimization scheme described before. In the case of ETLOP, at any time t considered as the forecast origin, we use the empirical distribution of GDP growth up to time t.
Table (ref) shows the results. We report results using optimal weights, but very similar conclusions are obtained using optimal priors. As a benchmark, we also report the results of BLP and ETLOP transformations applied to AR forecasts. As the table shows, the beta transformation further improves the density forecast performance of our scenario-augmented MSAR (SA-MSAR) forecasts, in terms of both log scores and PITs. The ADP is now 0.4-0.42 and the p-value of PIT uniformity test jumps to 0.68-0.81, depending on the specific weights used. Also, our optimized SA-MSAR forecasts, once transformed by BLP, remain clearly superior to the benchmark BLP-transformed AR forecasts, in terms of both accuracy and calibration, although benchmark forecasts have also improved compared to their untransformed counterparts from Table (ref). Conversely, the ETLOP methodology does not appear to provide any improvement in this specific empirical application.
Finally, Figure (ref) shows the entire empirical distribution of the PITs in the case of beta-transformed optimal SA-MSAR forecasts (using combination weights $\mathbf{w}_2^{*}$). In this case, also the highest decile lies within the 95% confidence bands of the uniform distribution.
We have developed an approach for generating density forecasts of macroeconomic variables using a variety of discrete economic scenarios provided by external sources. The approach is based on a Bayesian regime-switching model in which experts' views are translated into priors on economic regimes, and different views are pooled together to enhance density forecast performance.
We have presented an empirical application in which density forecasts of U.S. GDP are formed using the supervisory scenarios defined by the Fed. We have shown that the approach is able to achieve both good forecast accuracy and correct calibration of predictive distributions, merging the flexibility of mixture predictive densities provided by regime-switching models with the benefits of forecast combination, which are well-established in the literature.
Importantly, this methodology allows to evaluate the usefulness of economists' views for density forecasting. To illustrate this possibility, the empirical application tracks the contribution of Fed's scenarios to the optimized forecasts over time.
This approach appears particularly valuable in all contexts in which tail risk has a clear economic interpretation and when economic projections have to comply with external, possibly judgmental views. Researchers and practitioners interested in this kind of analysis may fine-tune the approach by tailoring the range of views to be considered and by selecting different objective functions in the optimization procedure.