Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
52,739 characters · 13 sections · 28 citation commands
A Factor-Augmented Markov Switching (FAMS) Model
\onehalfspacing
Since the late 1980s, Markov switching (MS) models have been widely employed by economic researchers to analyze various quantities of interest. Early studies analyze for instance stock market returns (pagan1990alternative) or asymetries over the business cycle (hamilton1989). It has been widely acknowledged that MS models are rather useful in capturing nonlinear time series characteristics as compared to more classical approaches such as standard ARMA models (hamilton1994time). As a result, MS models have gained in popularity and are commonly found in empirical macroeconomic research (hamilton2010). \\
Despite the popularity of the model family, the standard MS framework does not come without criticism. As pointed out by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KAUFMANN201582}, a main critique is the assumption of exogenous transition probabilities. This assumption implies that there is no explicit interpretation of the process driving the switching behavior of the model. Moreover, the most commonly applied prior framework a priori favors the model switching to the states that are visited most often. This approach leads to a model that neglects further information as for instance macroeconomic conditions that are readily available to the researcher.\\
This criticism has lead to a rather popular extension of the standard MS model that uses time-varying transition probabilities arising from a prior structure taking the form of a (multinomial) logit or probit specification (filardo1994business; Meligkotsidou2011; KAUFMANN201582). This modified prior setup effectively enables relevant exogenous variables to inform the switching behavior of the model. \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KAUFMANN201582} discusses Bayesian estimation and models a Philipps curve in a MS framework.\\
While in general this method is a viable and useful extension of the standard MS framework, it comes with some unpleasant downsides, especially when dealing in data rich environments like macroeconomics or finance. The (multinomial) probit or logit regression that forms the prior distribution of the transition probabilities is known to perform rather poorly when the set of predictors becomes too large as discussed in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{zahid2013multinomial}, \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{ranganathan2017common} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{dejong2019}. This often destabilizes the modeling framework and makes the effort to model regime switching using exogenous regressors cumbersome. In addition, it is often not clear a priori which variables are relevant to include when modeling switching behavior in a MS setting. In principal, this problem can be overcome by shrinkage priors on the coefficients of the logit/probit regression as demonstrated in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{Zens2019} for the related family of mixture-of-experts models. However, although variable selection might work well for a medium sized set of predictors, it is likely to become problematic when the number of regressors becomes too large compared to the available information in the data. Even more importantly, a variable selection approach does not resolve multicollinearity issues that commonly arise in macroeconomics and finance, where various time series describe very similar underlying processes of an economy. For instance, quarterly gross domestic product and quarterly industrial production are likely to show a large amount of comovement. Variable selection is prone to be not well behaved in such cases (george2010dilution). Thus, the researcher is required to come up with a pre-selection of relevant variables. This process is usually rather arbitrary and likely to severely reduce the information set available to inform the switching behavior of the model.\\
As a result, most articles dealing with MS models in economics and finance restrict the number of variables that inform the switching process to a relatively small number.\footnote{For instance, \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KAUFMANN201582} uses one exogenous variable and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{Meligkotsidou2011} uses six variables to inform the switching process.} As pointed out by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{bernanke2005favar}, a small number of variables will most probably not be able to span the information set consisting of hundreds of time series available to the researcher, to central banks and to financial market participants. This leads to two potential problems when employing these ”restricted” data sets. The first problem relates to policy analysis and forecasting exercises. It is well known that despite good in-sample predictions, out of sample forecasts of MS models regularly fail to consistently beat simple benchmark processes (engel1994can; boot2017macroeconomic). In these cases, the models might be missing important information due to restrictions in the estimation process. However, as shown by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{stock2002forecasting}, the forecasting ability of models using factor augmentation increases significantly as compared to models using ”restricted” information sets. Second, the effect of certain variables of interest to researchers and policy makers cannot be evaluated in the ”restricted” framework as they are not explicitly included in the model.\\
To overcome above mentioned issues, we offer a modeling framework that combines standard MS models with recent developments in factor modeling. In spirit, this article is closely related to the FAVAR model developed in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{bernanke2005favar} and builds on the sparse factor model outlined in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KASTNER2018}. The factor augmented Markov switching model (FAMS) that we derive in this article effectively summarizes large amounts of information about the economy and financial markets in a small number of estimated factors. Augmenting the model to inform the switching process in a MS framework provides a rather appealing solution to computational and statistical problems arising when using a large number of variables in this modeling framework.\\
This article summarizes the key points of the statistical framework necessary to estimate FAMS models in a fully Bayesian setup through Gibbs sampling. We demonstrate the ability of this model using autoregressive processes with possibly switching means and switching variances. Applying the model to an artificial data set, we find that the large information set included in FAMS models results in improved in-sample fit and more precise state allocations as compared to standard MS models making use of the full data set or random subsets of the data. The real world abilities of FAMS models are demonstrated through an application, where we aim to identify US business cycles through informing the switching process with a vast amount of macroeconomic time series.\\
Our contribution is thus twofold. The FAMS model enables it to conveniently estimate MS models where the switching process is influenced by a lower dimensional representation of a large number of possibly collinear candidate variables. Moreover, a shrinkage prior is imposed to conduct variable selection in cases where some factors might be irrelevant for the regime switching behavior of the model.\\
The remainder of this article is organized as follows. In Section 2 we formulate the general modeling framework for the FAMS model. Section 3 discusses Bayesian estimation of the model using MCMC methods. Section 4 presents the benefits of the FAMS model via a simulation study. In Section 5, a high-dimensional macroeconomic data set is used in an application to estimate business cycles in the United States. A brief conclusion is provided in Section 6.\\
We consider a general Markov switching autoregressive model (MS-AR) of order $p$ with time-varying transition probabilities. This model has been thoroughly discussed in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet[Ch. 4; Ch. 9]{kim1999state} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet[Ch. 12]{fruhwirth2006finite}. Let $\{y_{t}\}_{t=1}^T$ be an observed time series arising from the data-generating process (DGP)
where $\{S_{t}\}_{t=1}^T$ is assumed to be a hidden discrete-time state process with finite state space $\mathcal{H} = \{1,\hdots,H\}$. In the most general setup, the intercept parameter $\mu_{S_{t}}$, the persistence parameters $\phi_{j,S_{t}}$ and the error variance parameter $\sigma^2_{S_{t}}$ are state-dependent. First, we collect all our parameters as $\theta = \{\theta_h | \theta_h = (\mu_h,\phi_{1,h},\hdots,\phi_{p,h},\sigma_h^2)^\top, h = 1,\hdots,H\}$. Thus, it is straightforward to see that if the process is in state $k \in \mathcal{H}$ in time period $t$, parameter values change accordingly:
The peculiarity of Markov switching models is that their switching is determined completely by a stochastic process. The state indicator vector $S$ can thus be seen as an irreducible, aperiodic Markov chain describing a sequence of events where the probability of each state only depends on the state attained in previous rounds. We assume a first order Markov process, where the process has only a memory of one period:
Therefore, the transitions between two states are described by probabilities that are summarized in a transition matrix $\Xi_t$ of dimensions $H \times H$. In this transition matrix, each element $\xi_{jk,t}$ sufficiently describes the stochastic behaviour, i.e. the transition from state $k$ to state $j$. Note that the matrix $\Xi_t$ is row-standardized. Furthermore, the transition matrix $\Xi_t$ features a time index $t$ since we assume that the switching process is not constant over time. It is modeled as dependent on a set of $r$ exogenous factors captured in $\{Z_t\}_{t=1}^T$. These variables are not entering the model directly in the autoregressive process, but are rather expected to drive the transition dynamics indirectly. Therefore, they are responsible for altering the state allocation and thus the estimation of our state-dependent parameters in $\theta$. Moreover, a delay parameter $d$ is introduced for convenience in cases where the exogenous factors $Z_t$ are used as leading indicators. The estimation of $Z_t$ is covered in section (ref).\\
The multinomial logit link is used to model the influence of the factors $Z_t$ on the transition probabilities $\xi_{jk,t}$ in a way such that
where $j,k = 1,\hdots,H$ and $\gamma_{jk}$ and $\beta_{jk}$ are the coefficients of the intercepts and the factors, respectively. Moreover, we follow \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KAUFMANN201582} and specify $Z_{t} = \tilde{Z}_{t}-\bar{Z}_i$ in a centered way for two reasons. Besides the fact that it defines the mean $\bar{Z}_i$ as an arbitrary threshold level, more importantly, the time-invariant part of the transition probabilities $\gamma_{jk}$ does then not depend on the scale of $\tilde{Z}_{t}$. Otherwise, this would have to be taken into account when choosing a prior distribution for the parameter.\\
For identification purposes, it is necessary to define a baseline state $h_0 \in \mathcal{H}$ in which the parameters are assumed to be zero, i.e. $(\gamma_{jh_0},\beta_{jh_0}) = 0$. This yields
where $\mathcal{H}_{-h_0}$ denotes all states but the reference state $h_0$. \\
In order to improve the efficiency of the estimates, we follow AMISANO2013118 in imposing the restriction of common slope coefficients across states, $\beta_{jk} = \beta_{k}$. This translates into the assumption that the effect of $Z_{t}$ differs only by the state attained in the previous time period. For notational convenience, we collect the multinomial logit coefficients in the vectors $\gamma = \{\gamma_h \:\vert\:\gamma_h = (\gamma_{1h},\hdots,\gamma_{Hh})^\top, h = 1,\hdots,H\}$ and $\beta = \{\beta_h, h = 1,\hdots,H\}$ and gather all parameters in $\vartheta = \{\theta, \gamma, \beta\}$.
In the above mentioned modeling framework, it is possible to proceed by estimating the model using a given subset of the full set of exogenous covariates $X_t$ to inform the switching process through the prior specification. However, there is a variety of scenarios where additional relevant information may be available that is not incorporated in this subset. Suppose we can find a set of latent factors $Z_t$ that condenses a large part of the available information into a small number of time series. These factors might represent more diffuse concepts such as ”political instability” or ”credit market sentiments”. These concepts are generally difficult to capture in one or two time series. The task of finding $Z_t$ given $X_t$ can be dealt with using factor analysis frameworks (see lopes2004bayesian or lopes2014modern for a comprehensive overview). The setup used in this article is closely following \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KASTNER2018} who proposes a time varying sparse factor model. It can be summarized as
where $Z_t$ is the set of latent factors we want to recover for use in the MS model, $\Lambda$ is a $m \times r$ factor loadings matrix and $U_t$ and $V_t$ are diagonal matrices discussed in more detail below. The $m$ scaled and centered time series $X_t$ = $(X_{1t},\ldots,X_{mt})$ describing the economy are assumed to follow a conditional Gaussian distribution, i.e.
where $N_m(\cdot)$ denotes a $m$-dimensional normal distribution. To achieve dimensionality reduction, the $m \times m$ covariance matrix $\Sigma_t$ is assumed to decompose into a $m \times r$ factor loadings matrix $\Lambda$ as well as two diagonal matrices $V_t$ $(r \times r)$ and $U_t$ $(m \times m)$ as follows:
Following \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KASTNER2018}, $\Lambda$ is assumed to be constant over time. The factor variances in $V_t$ and the idiosyncratic variances in $U_t$ are evolving over time through stochastic volatility models (kastner2014ancillarity). Specifically, let $U_t = \text{diag}(\text{exp}(g_{1t}), \ldots, \text{exp}(g_{mt}))$ and $V_t = \text{diag}(\text{exp}(h_{1t}), \ldots, \text{exp}(h_{rt}))$. Then the log variances are modeled through autoregressive processes of the form
and
That is, the idiosyncratic variances follow a centered AR(1) process and the factor variances follow an autoregressive process with zero mean to identify the scaling of the factors. For further details on the factor model described here, refer to \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KASTNER2018}. For further topics in factor modeling including estimation and identification issues, see for instance \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{lopes2014modern} or \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{kaufmann2019bayesian}.
Bayesian estimation requires the specification of adequate prior distributions. This section gives an overview of prior distributions both in the Markov switching model ((ref)) and the factor model ((ref)) that form the proposed model setup.\\
For the Markov switching AR(p) model, independent prior distributions for all parameters $\pi(\vartheta) = \pi(\alpha)\pi(\phi)\pi(\sigma^2)\pi(\gamma)\pi(\beta)$ are assumed. Since the specification in (ref) is piece-wise linear, we specify a normal prior distribution for the mean and persistence parameters and an inverse Gamma distribution for the variance parameters
Regarding the hyperparameters, we choose $m_0 = r_0 = 0$, $M_0 = 10$ and $R_0 = 4$ to stay rather uninformative. In setting $c_0=d_0=1$, we use an informative prior distribution on the variance parameters in order to regularize the variances sufficiently away from zero to circumvent singularities.
In the multinomial logit, seperate prior distributions for the intercepts $\gamma$ and the coefficients $\beta$ are specified. For the intercepts, we assume
where $g_{0,h}=0$ and $G_0=4$.
Generally speaking, some factors might be more prone to capture information relevant to the switching process, whereas other factors might mainly introduce noise to the model. To overcome these issues, a normal gamma shrinkage prior (polson2010) is chosen for the coefficients $\beta$ of the exogenous factors. Therefore, the prior on each element for each state of the coefficient vector can be written as
Now we turn to the specification of prior distributions for the factor model. For the elements of the factor loadings matrix, again a normal gamma prior is employed to achieve sparsity:
where the global shrinkage parameter $\lambda^2_{\tau,i}$ is defined per row of the factor loadings matrix. Thus, this setup implies that each time series has a high a priori probability of not loading on any factor. Similarly, the global shrinkage parameter $\lambda^2_{\psi,h}$ is defined per equation of the multinomial logistic prior to allow varying levels of shrinkage across states in the MS model. In choosing prior distributions on the stochastic volatility parameters, we strongly follow \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{KASTNER2018}. Since both the stochastic volatility process of the factor variances and the idiosyncratic variances have own persistence and variance parameters, a subscript $j \in \{g,h\}$ indicates to which group of processes they belong. The priors can then be summarized as
where the function of persistence parameters $\phi_j$ follows a Beta distribution to ensure stationarity.\\
Finally, it is useful to impose hyperpriors on the hyperparameters in (ref) and (ref). In general, normal gamma type shrinkage priors \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{polson2010,griffin2010} enable implicit variable selection in a sophisticated way by specifying a continous prior distribution with a rather large probability mass on zero and fat tails. The degree of shrinkage is influenced by two parameters: a global shrinkage parameter $\lambda_j$ ($j \in \{\psi,\tau\})$ and a local shrinkage parameter $\psi$ or $\tau$. As seen in (ref) and (ref), an additional layer of prior distributions on the local shrinkage parameters is introduced. Both $\psi$ and $\tau$ are assumed to follow a Gamma distribution with shape and scale hyperparameters $\omega_j$ ($j \in \{\psi, \tau\})$. Note that by deviating from specifying the shape and scale parameters to be equal, it is not possible to identify either the global shrinkage parameter or the parameters of the prior for the local shrinkage component. A choice of $\omega_j = 1$ leads to the Bayesian LASSO (park2008bayesian). Smaller values are associated with a more severe penalty function and thus impose stronger shrinkage. To complete the prior setup, a Gamma prior is specified on the global shrinkage component $\lambda_j^2$, $j \in \{\psi,\tau\}$:
where the hyperparameters for the prior on $\beta$ are set to $c_{\psi0}=c_{\psi1}=0.01$. This implies heavy shrinkage on $\beta$. For reference, see for instance \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{bitto2019achieving}. The hyperparameter values in the factor model are chosen in accordance with the standard values proposed by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{kastnerfacstochvol}. They imply less heavy shrinkage on the elements of the factor loadings matrix $\Lambda$.
Estimation of the model is carried out via MCMC sampling using a variety of data augmentation techniques. Bayesian estimation and inference with a Gibbs sampling algorithm requires the combination of the likelihood and the proposed priors to produce conditional posterior distributions for all parameters. The employed algorithms closely follow the algorithms described in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet[Ch. 11]{fruhwirth2006finite} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet[Ch. 9]{kim1999state}. Since the FAMS model is a variant of a state space model, the data augmentation techniques introduced by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{carter1994gibbs} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{fruhwirth1994data} are viable candidates for performing estimation. Sequentially sampling from the conditional posterior distributions after convergence of the algorithm is achieved then allows posterior parameter inference to take place. Generally speaking, the sampling scheme employed in this paper is designed to iterate over the following three steps:
{
}
For estimating the factor stochastic volatility related quantities, we use the R-package factorstochvol provided by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{kastnerfacstochvol}.\footnote{In this preliminary version of the paper we choose to implement the FAMS model in a two step procedure. In a first step, we simulate from the posterior distributions of the factors. The resulting posterior means are then used as explanatory variables that inform the switching process in the MS framework. A fully Bayesian approach that includes the uncertainty around the factor posterior means will thus result in slightly higher uncertainty around other model estimates and will be included in future versions of the paper.} This allows us to make use of the efficient implementation using interweaving techniques. For ease of reference, estimation of the multinomial logistic prior specification using the partial dRUM sampler is discussed in detail in (ref). Posterior simulation resulting from a normal gamma prior setup is briefly discussed in (ref). The remaining simulation steps are standard and thus not discussed in detail.
Parameter estimation in the family of mixture and Markvo switching models can suffer from various difficulties, especially in a Bayesian framework. Label switching is a common issue in mixture modeling (hurn2003estimating; jasra2005markov). It is the result of the likelihood function being invariant to relabeling the components due to multimodality as discussed in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{redner1984mixture}. This can lead to problems as label switching during MCMC sampling might result in possibly distorted, multimodal posterior distributions that are in general difficult to summarize. Deriving posterior means or other point estimates based on these posteriors then becomes inappropriate (stephens2000dealing). The same rationale holds true for Markov switching models.\\
Early references for relabeling algorithms include \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{celeux1996stochastic}. However, this algorithm requires known true parameter values, which renders it not very useful in real data applications. \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{stephens2000bayesian} suggests an algorithm that relabels the draws in a way such that the (marginal) parameter posterior distributions are as unimodal as possible. \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{stephens2000dealing} provides a literature review as well as a decision theoretic framework to deal with label switching.\\
In the simulation studies and application presented below, we make use of identifying restrictions to identify the sampler (see for instance lenk2000bayesian).\footnote{Knowing that this is not ideal (fruhwirth2004estimating), a random permutation sampler in combination with a post-processing procedure to relabel the draws will be implemented in future versions of this paper.} This completes the simulation setup and the description of the employed estimation techniques of the FAMS model. The following sections discuss the model in the context of artificial data as well as an application to US business cycle estimation.
To test the performance of the proposed modeling framework, synthetic data sets are generated using a recursive procedure. The factor model outlined in (ref) is used to simulate data under the assumption that the true DGP is driven by the factors.\footnote{We proceed like this for two reasons. The first argument is a theoretical one, where we argue that it is the factors which convey the information on whether transition probabilities should change or not. Second, from an econometric point of view generating probabilities from multinomial logit models using about 200 time series is neither advisable nor applicable. The stability of the exercise depends strongly on the amplitude of the effects and it is extremely tedious to generate a setup where a simulation study can be conducted in a controlled setup.} Hence, we proceed in the following manner: In a first step, we generate $r=3$ AR(1) processes
centered around zero. From these, 200 time series are generated using the DGP outlined in section (ref). The AR coefficients $\phi_j$ $(j \in \{g,h\})$ of the log variances are simulated from U(-0.8,0.8) and the means $\mu_g$ are assumed to follow a normal distribution with N(0.2,0.2). The variances $\sigma^2_j$ $(j \in \{g,h\})$ are simulated from |N(0.2,0.2)|.\\
After that, three factors are estimated from the generated time series. Estimation is based on 80.000 draws, where the first 40.000 draws are discarded as burn-in. The true factors $F_t = (f_{1,t},f_{2,t},f_{3,t})^T$ are used to generate transition probabilities
which are then employed to draw the states $S_t$ of the process. Conditional on knowing the states, the final time series $y_t$ can be generated using the process
In this setup, $H=2$ and thus our finite state space is $\mathcal{H} = \{1,2\}$. We choose $\gamma_{j1} = (1.5,-1.5)^T$ and set $\gamma_{j2} = (0,0)^T$ for identification. Furthermore, $\beta_1 = (-1.2,1.1,0.9)^T$. Again, the second column of $\beta$ is set to zero due to identification reasons. The parameters in the AR model are set to $\mu = (-0.25,0.25)^T$, $\phi=0.55$ and $\sigma^2 = (0.1,0.05)^T$. Note that this setup allows to identify the sampler using identifying restrictions. Furthermore, the parameters are carefully chosen in a way such that a lot of time variation in the transition probabilities results and that the data is comparable to real world examples in economics and finance. This setting is used to generate $N=100$ different datasets, each with $T=250$.\\
Five different modeling approaches are implemented and compared. The first model imposes constant transition probabilities by using no external covariates in the MNL part of the MS model. This ”intercept only” model serves as baseline model in what follows below. In the second model, the full information set of 200 time series is used to model the transition probabilities.\footnote{A Gaussian prior with zero mean and a variance of 4 is assumed on $\beta$ for the models with no shrinkage prior.} The third model uses all 200 time series as well, however, has a shrinkage prior placed on $\beta$ to induce sparsity. The fourth model corresponds to the FAMS model without normal gamma prior. Finally, the fifth model is the FAMS model including a normal gamma prior. To gain even more insight, all models are estimated with varying degrees of shrinkage induced by the normal gamma prior. For the sake of brevity, we fix the number of factors in the estimation to the true number of factors.\\
For each run of the MCMC sampler, $M=50,000$ draws are kept after a burn-in phase of $50,000$ draws. Two measures are employed to compare the performance of the models. The first one is the average median root mean squared error (RMSE) of the in-sample fit through
The second criterion is a measure of the misclassification rate of the states through
where $\mathcal{I}(\cdot)$ denotes the indicator function. In both cases, $q_{50}$ denotes the posterior median.\\
(ref) reports the average in-sample median RMSE and MCR across 100 runs relative to the results of the baseline model. Results are presented for both modeling approaches and varying degrees of shrinkage as indicated by various values of $\omega_{\psi}$. It is noticeable that the FAMS model performs considerably better than the baseline model. Moreover, the FAMS model performs better than the full data set in most cases. A few interesting patterns are worth mentioning. Unsurprisingly, the full information set without a shrinkage prior performs even worse than the basline model (bear in mind that the full information set consists of 200 time series). Applying shrinkage therefore definitely makes sense and improves in-sample fit and classification. Furthermore, the degree of shrinkage matters. While the performance of $\theta_\psi=0.6$ and $\theta_\psi=1$ is relatively similar, inducing too much shrinkage (say, $\theta_\psi < 0.5$) seems not like an advisable strategy. In these cases, the prior tends to shrink too much information to zero. Hence, the performance of the full information set model approaches the baseline model as $\beta$ gets close to $0$. In our experience, the coefficients in a multinomial logistic prior in a MS context are rather sensitive and in general difficult to estimate precisely. Often, this leads to the normal gamma setup applying too much shrinkage. This furthermore underlines the benefits of using estimated factors in a MS context. It drastically reduces multicollinearity issues and maximizes the amount of information contained in a given time series used to estimate the multinomial logistic regression. Both channels may counteract problematic and unintended shrinkage effects.\\
In an additional exercise, we tried estimating the model using random subsamples from the full information set. This is supposed to roughly emulate theory-guided a priori variable selection. Each subsample consists of ten time series. Comparing the results of the best performing subsample with the FAMS model, two results are noticable: In-sample fit as measured by RMSEs become comparable to the FAMS model. Interestingly, model performance in terms of MCRs is rather bad. Here, further investigation is indicated.\\
Focusing on the FAMS model, the best performance is achieved in the setup without shrinkage prior. Again, we argue that the prior might shrink necessary information toward zero. However, more experimentation and simulation studies will be necessary to systematically investigate this issue. Apart from this observation, it has to be noted that the differences between various FAMS setups are very modest in general. We suspect that performance will strongly depend on the application at hand. Moreover, since the focus here lies not on variable selection, but on the benefits of projecting large information sets into lower dimensions, we conclude that the FAMS model does a pretty decent job.
Since the seminal paper of \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{hamilton1989}, Markov switching models have been used to detect and explore the nature of business cycles. An excellent overview is provided by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{hamilton2016macroeconomic}. It is nowadays widely acknowledged that the nature of macroeconomic behaviour is usually not well captured by linear models.
From a theoretical point of view, there is a lively debate on why and how regime changes take place. One school of thought posits that the alternation of expansions and recessions results from different paces of technological innovations, which are assumed to be exogenous to the economy. Another paradigm starts by assuming that the market economy is inherently instable and produces business cycles endogenously. A recent theoretical contribution by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{matsuyama2013good} thus endogenizes credit markets and shows how an economy itself can drive recurrent booms and busts, with each bust sowing the seeds of the next boom.
For the basic model of the macroeconomy of United States, we obtain publicly available data from the Federal Reserve Bank of St. Louis \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{doi:10.1080/07350015.2015.1086655}. This very well known data set has been used in a variety of macroeconomic studies, see for instance \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{ludvigson2015uncertainty} or \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{stock2016dynamic}. It covers a variety of macroeconomic series capturing eight core areas of the US economy such as real activity and financial market. Over 200 time series in this data set span an information set that is too large to be employed in standard MS models. Hence, the FRED database is a proper candidate to demonstrate the benefits of the FAMS model.\\
As the goal is to model business cycles, the quarterly log differenced time series of industrial production is modeled as uncentered AR(4) process with switching intercept:
This approach closely follows the seminal paper by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{hamilton1989}, who however uses a centered MS-AR model. We deviate from the centered parameterization, because in a Markov switching context the uncentered parametrization does not induce immediate mean level shifts. Instead, the mean level approaches the new value smoothly over various time periods. In our eyes, this behavior is a better description of large, complex and dynamic systems such as an economy.\\
After transforming all remaining variables to ensure stationarity and discarding time series with too many NA values, the full data set employed contains 209 time series and covers 234 time periods from 1959Q3 to 2017Q4. This data set is assumed to contain all necessary information to inform the switching mean process of industrial production growth.\\
The factor model outlined in (ref) is then applied to this data set. Comparing BICs for models with 1-25 factors points into the direction of a model with seven factors. Thus, the 7-factor model is described in more detail below. Guided by economic theory, a MS model with two states -- interpreted as recessionary and expansionary periods -- is implemented as in previous literature.\footnote{A more sophisticated approach would implement a marginal likelihood based grid search over various combinations of number of factors, number of states and different prior values.}. The estimated factors are then used as exogenous predictors in the MS framework to model the regime switching behavior interpreted as US business cycles.\\
Identification of the MS model is achieved by imposing the identifying restriction $\mu_1 > \mu_2$. Identification of the factor model is achieved by restricting the factor loadings matrix $\Lambda$ to have zeros above the diagonal. In theory, this might lead to problems when the first $r$ variables gain too much influence and weight when employed in estimating the factors. However, a large enough number of time series should prohibit this scenario and estimating 7 factors on more than 200 time series minimizes the risk to a certain degree. In addition, the estimated factor loadings matrix is not extremely sparse, again counteracting these possible threats to identification (see kaufmann2019bayesian).
The sampling algorithm is iterated $50,000$ times after discarding the first $50,000$ draws as burn-in. Convergence of the MCMC sampler is usually very good. (ref) and (ref) show the full information set used to estimate the factors as well as the factors posterior means used to inform the switching process.\footnote{To be clear, in this exercise the factor means in period $t$ inform the switching process in period $t$. This corresponds to a delay parameter $d=0$. It is possible to look at leading indicators by setting $d>0$. However, this leaves the results qualitatively unchanged in this setup. Thus, we do not show these additional results for brevity reasons.}\\
The estimated factors explain a variance share of 34% when averaging over time and series. This is comparable to related literature. For instance, \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{kaufmann2019bayesian} explain around 38% of the variance of 282 time series using a model with 20 factors. However, the explained variance share varies significantly over time and between series. The interquartile range of the explained variance share over time lies between 23% and 45%. The factors explain over 90% of the variance of 6% of the time series, over 80% of 10% of the series and more than 50% of the variance of 36% of the series.\\
Interpreting the factors is more of an art than a science. However, to give a brief intuition on how the factors could be interpreted, a table with the time series with the highest absolute factor loadings is provided in (ref). This table is used to name the factors in the results presented below. However, when interpreting factors, there is usually no clear cut and distinct interpretation possible as we do not restrict time series to only load on one factor. All in all, naming the factors is mainly done to achieve increased ease of presentation, but should be taken with a grain of salt.\\
(ref) provides the industrial production growth rate time series used in the analysis. (ref) plots the estimated posterior median of the transition probabilities. Both plots show recession dates as published by NBER as shaded areas. The FAMS model is able to capture the recessions rather well. Interestingly, the model detects a recession in 2014 that is not classified as a recession by NBER. The corresponding estimates of the parameters of the AR process can be found in (ref). The intercept estimates point into the direction of one state with rather volatile growth rates of industrial production. A large part of the intercept posterior density of this state also lies in the negative spectrum. Thus, we proceed to label this state Recession. On the other hand, the second state can be characterized by consistently positive growth rates and is thus labeled Expansion.\\
The estimated uncentered AR(4) process seems to be sufficiently stationary. This can be easily inspected visually by rewriting the AR(4) process as AR(1) process. This results in the so called state space or companion form of the process. If the eigenvalues of the resulting companion matrix lie well inside the unit circle, the process is considered stationary. This exercise is depicted in (ref), where the abscissa denotes the real part and the ordinate the imaginary part of the eigenvalues. Bold black dots denote the median eigenvalues.\\
Finally, the results of the multinomial logit part of the FAMS model are provided graphically in (ref).\footnote{Numerical results are available from the authors upon request.} The states are labeled ”Expansion” and ”Recession” corresponding to the estimated posterior mean growth rates. ”Expansion” is the baseline state in this setup, thus all coefficients are set to 0 for identification reasons as discussed earlier. The black dots indicate posterior means and the grey lines show the 90% HPD region of the posterior density of the MCMC draws. For this estimation run, we set $\omega_{\psi}=0.6$ to apply a medium amount of shrinkage on the factors.\\
A few observations are worth to be mentioned: First, not many of the coefficients are heavily shrunk to zero, indicating a rather large amount of information in the factors that is relevant to inform the switching process in the MS model. Secondly and not surprisingly, the highest correlation of recessions as measured per a real activity indicator is shown by the factors representing real economic activity (Business & Employment and Real Activity). Third, Interest Rate Spreads seem to have quite some predictive power when it comes to explaining recessions in this data set. This is in line with a large strain of literature connecting real economic recessions and risk premiums on the financial markets (gilchrist; salido).\\
For future applications, it will be interesting to estimate the FAMS with switching variances in stock market or interest rate applications as well as to experiment with various amounts of shrinkage. It is also possible to estimate $\omega_{\psi}$ from the data as demonstrated for instance in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{malsiner2016model}.\\
In this article, we develop a factor-augmented Markov switching model with time varying transition probabilities. We assume that a set of latent factors informs the switching process through a multinomial logistic prior setup. This alleviates problems arising from a high-dimensional data set where a large amount of variables might potentially be relevant to model the Markov switching process. Estimation of the factors and the MS model is carried out in a Bayesian framework through Gibbs sampling. In a first step, the FAMS model is applied to artificial data sets. These simulation studies underline the benefits of the FAMS when compared to similar regime switching frameworks. In addition, a real data application highlights the potential of the FAMS model in business cycle research.\\ Future avenues for research include for instance a thorough cross-model comparison of forecasting abilities in different environments and additional applications that test the usefulness of the FAMS framework in the context of switching variances. It will be interesting to see how the model performs in other real world applications where the goal is for instance to model interest rates (ang2002regime) or stock market returns (schaller1997regime).\\ To summarize, we hope the FAMS model will provide a proper modeling framework for fields where Markov switching models are commonly employed in data rich environments. Rather commonly, applications will thus be found in economics and finance. However, the modeling framework might also be useful in fields like medicine, where Markov switching models are used as an early detection system for influenza (beneito2008) or in accident analysis where MS models are used to model vehicle accicdent frequencies (malyshkina2009markov).
{\fontfamily{ptm}\selectfont\setstretch{0.85} \addcontentsline{toc}{section}{References} }