EconBase
← Back to paper

Identification of structural shocks in Bayesian VEC models with two-state Markov-switching heteroskedasticity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

52,874 characters · 12 sections · 117 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification of structural shocks in Bayesian VEC models with two-state Markov-switching heteroskedasticity

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 140 chars of source]

} \fi

abstractWe develop a Bayesian framework for cointegrated structural VAR models identified by two-state Markovian breaks in conditional covariances. The resulting structural VEC specification with Markov-switching heteroskedasticity (SVEC-MSH) is formulated in the so-called B-parameterization, in which the prior distribution is specified directly for the matrix of the instantaneous reactions of the endogenous variables to structural innovations. We discuss some caveats pertaining to the identification conditions presented earlier in the literature on stationary structural VAR-MSH models, and revise the restrictions to actually ensure the unique global identification through the two-state heteroskedasticity. To enable the posterior inference in the proposed model, we design an MCMC procedure, combining the Gibbs sampler and the Metropolis-Hastings algorithm. The methodology is illustrated both with a simulated as well as real-world data examples.

{\it Keywords:} cointegration, stock price fundamentals, structural VAR models

Declarations of interest: none

\spacingset{1.8}

Introduction

Since the seminal work by sims1980macroeconomics, both structural and reduced-form Vector Autoregressive (VAR) models have been used extensively in empirical macroeconomic studies. A reduced-form VAR provides a potent statistical tool for capturing the dynamics of multivariate macroeconomic time series, with the model's parameters being uniquely identified solely by data. However, when structural analyses are of interest (such as forecast error variance decomposition, impulse response functions, historical decompositions, and forecast scenarios), then a VAR model in its structural form is required. In a structural VAR (SVAR), it is identification of orthogonal (or, structural) shocks that is of vital importance, with the task hopefully delivering results that are not only statistically valid, but also meaningful from an economic perspective.

The identification of structural shocks boils down to identification of the SVAR model's parameters, a problem that a reduced-form VAR does not suffer from. Traditional approaches to the task include linear (typically, zero) or sign restrictions, which can be imposed on the system variables' reactions at given horizons (see, e.g., arias2018inference, rubio2010structural, but also rubio2005markov in the context of SVAR models identified through heteroskedasticity). Most often, such a number of the restrictions is introduced that ensures the so-called just identification (as opposed to over-identification), which, however, precludes testing thereof. On the other hand, the over-identified models do admit testing the over-identifying restrictions, but the approach proves viable only under the validity of the remaining constraints.

The above-listed problems can be circumvented by extending a standard homoskedastic and Gaussian VAR model to account for additional features of the shocks, such as heteroskedasticity (modeled by various approaches) and non-Gaussianity (see, e.g., kilian2017structural, lutkepohl2017structural for reviews). Such generalizations may (although do not have to) successfully aid the identification, so that the traditional restrictions (imposed a priori) become over-identifying, and thus testable. Nonetheless, only statistical identification is directly facilitated then, with no guarantee of economically sound interpretations whatsoever. Therefore, figuring out the economic meaning of obtained results is an additional, follow-up issue, which even under statistically identified shocks can still pose quite a challenge.

The line of research on identification through regime-switching heteroskedasticity (although not in VAR, but simultaneous equations models) started with the works by sentana2001identification, rigobon2003identification, and rigobon2003measuring. However, in the latter, the Authors notice that the very idea of linking the errors' variances with estimation and identification of the model parameters traces back to wright1928tariff.

In this paper, we adopt the idea of structural shocks being identified through heteroskedasticity for identification of Bayesian cointegrated (as opposed to stationary) VAR systems with volatility breaks governed by a two-state Markov chain. As such, our study contributes to the area of research devoted to VAR models with Markov-switching heteroskedasticity (VAR-MSH), started by lanne2010structural, and further developed by netsunajev2013reaction, velinov2013can, herwartz2014structural, lutkepohl2014disentangling, lutkepohl2017structural, velinov2015stock, lanne2016data, lutkepohl2016structural, lutkepohl2018relation, velinov2018importance, lutkepohl2021testing (see also lutkepohl2020bayesian for a more comprehensive literature review). The cited works utilized the frequentist statistical framework for estimation and inference, whereas kulikov2013identifying, wozniak2015assessing, lanne2016data lutkepohl2020bayesian resort to the Bayesian approach. To the best of our knowledge, out of all of the papers cited above, only the four: velinov2013can, herwartz2014structural, lutkepohl2016structural, and velinov2018importance consider cointegrated (as opposed to stationary) VARs with the structural shocks identified through regime-switching volatility, with the methodology developed within the frequentist estimation setting. Naturally, stationary VAR systems are methodologically more straightforward to handle (which obviously, does not preclude them from still proving empirically valid in many cases). However, taking long-run relationships into account can not only enhance the model's capability of capturing data dynamics, but may also provide additional restrictions aiding the identification by enabling one to isolate permanent shocks from ones that exert only transient impact on modeled variables (see, e.g., king1991stochastic, GonzaloNg2001, lutkepohl2005new).

As can be inferred from the above, the current state of pertinent research misses out Bayesian (as opposed to frequentist) structural VAR models featuring both cointegration and Markov-switching volatility. Hence, filling this gap is the main aim of this paper\footnote{Note that Bayesian Markov-switching VAR models with (possibly regime-changing) cointegration have already been considered in the literature, although only in their reduced form, and thus without the structural context; see cui2012bayesian, jochmann2015regime, sugita2016bayesian, and hauzenberger2021stochastic.}.

Another aim of our present work is to revisit the conditions under which a VAR (whether stationary or cointegrated) model with Markovian volatility breaks is uniquely and globally identified. The discussion builds on Theorem 1 stated in lutkepohl2020bayesian, for which we deliver a counterexample to show that yet another constraint (beyond the ones required by the theorem) is actually required to ensure the global identification of a two-state VAR-MSH system, and thus the theorem warrants only a local identification (up to the ordering of structural innovations). Then, we formalize our finding through a relevant theorem (with the proof presented in Appendix (ref)).

The remainder of the paper is organized as follows. The next section presents a two-state SVEC-MSH model underlying our methodological framework, and revisits a central issue of structural shocks identification therein. In Section 3, we develop a Bayesian approach to estimation of the model in question, further illustrated by an empirical analysis of a dividend discount model for the U.S. economy (Section 4), previously considered by binswanger2004stock, louis2010stock, velinov2013can, and lutkepohl2016structural. Section 5 concludes. Finally, our work is accompanied by a supplementary material presenting i) details on the prior structure, ii) a Markov Chain Monte Carlo algorithm for posterior sampling, iii) a simulated data study showcasing the methodology developed in the paper, and iv) additional empirical results.

Model framework

The sampling model

The basic structure underlying our modeling framework is a linear $n$-variate vector autoregressive process $y_t=(y_{t1}, ..., y_{tn})'$ of order $p$ ($y_t \sim$VAR(p)), in the structural vector error correction (SVEC) form, with deterministic terms and a regime-switching conditional covariance matrix:

equation[equation omitted — 174 chars of source]
equation[equation omitted — 124 chars of source]

where $\tilde{y}_{t-1}' = (y_{t-1}',d_{t-1}')$, $d_{t-1}$ denotes deterministic components restricted to cointegrating relations, $\alpha$ in an $n\times r$ matrix of adjustment coefficients, $\beta' = (\tilde{\beta}',\phi')$ is an $r\times \tilde{n}$ cointegration matrix $(\tilde{n}\ge n)$ of a full rank $r$, $r$ is the number of cointegration relationships ($0<r<n$, if they exist), or $r=0$ (for a non-stationary but non-cointegrated system), or $r=n$ (for a stationary VAR), and $D_t$ comprises deterministic variables (e.g., an unrestricted constant, inducing linear trends in $y_t$, and possibly seasonal dummies; see juselius2006cointegrated). Contemporaneously and serially uncorrelated error terms (i.e. structural shocks) are arrayed in $\varepsilon_t=(\varepsilon_{t1}, ..., \varepsilon_{tn})'$, with their regime-dependent, diagonal covariance matrix $\Lambda_{S_t}=diag(\lambda_{S_t})$, where $\lambda_{S_t}=(\lambda_{S_t,1}, ..., \lambda_{S_t,n})'$ denotes an $n$-variate vector of the conditional variances associated with volatility state $S_t$, i.e. $\lambda_{S_t,i}=Var(\varepsilon_{ti}|S_t,\lambda_{S_t,i})$ (note that, from a Bayesian perspective, conditioning on $\lambda_{S_t,i}$ as well as on $\lambda_{S_t}$ in Eq. ((ref)) is needed). The state-dependent covariance matrix of the reduced-form errors, $u_t=B\varepsilon_t$, is then given as $\Sigma_{S_t}=V(u_t|S_t, \lambda_{S_t})=B\Lambda_{S_t}B'$. Finally, $B=[b_{ij}]_{i,j=1,...,n}$, termed as the structural matrix, comprises the instantaneous responses $b_{ij}$ of the $i^{th}$ variable, $y_{ti}$, on the $j^{th}$ structural shock, $\varepsilon_{tj}$. For the sake of normalization, we set $b_{ii}=1$ for $i=1, 2, ..., n$, so that the signs of the shocks are fixed, leaving their identification hinged on heteroskedasticity and possibly also additional (linear) restrictions (see lutkepohl2020bayesian)\footnote{Note that lutkepohl2020bayesian assume that the normalization restriction, $b_{ii}=1$ for $i=1, 2, ..., n$, leaves the identification only to the changes of the relative variances, thereby resolving also the issue of the ordering of the shocks. However, as demonstrated in this paper, it is not the case.}. Simultaneously, $B^{-1}$ reflects the contemporaneous relationships between the endogenous variables.

Similarly to lanne2010structural and lutkepohl2020bayesian (among others), we assume that the variables $S_1, S_2, ..., S_T$ form a two-state ergodic and homogeneous Markov chain, $S_t\in{\{1,2\}}$, that governs the regime changes of structural shocks' variances\footnote{Notice that we restrict our considerations to a two-state model only. Generalization of the theoretical results presented in this paper, for systems with more regimes is deferred for our future work.}. The properties of the chain are determined by the transition matrix $\mathbf{P}=[p_{ij}]_{i,j=1,2}$, where $p_{ij}=\Pr( S_t=j|S_{t-1}=i)$, for $i,j\in\{1,2\}$, is the transition probability from state $i$ at time $t-1$ to state $j$ at time $t$. Thus, the elements of each row $p_i=(p_{i1},\ p_{i2})$, $i=1,2$, sum up to 1.

Note that, according to Eq. ((ref)), we decide to work directly with a structural VEC (SVEC) model instead of embarking on with a VAR structure and then recovering its VEC form parameters. This allows us to apply the B-model parameterization to the SVEC form, which is a common choice for a structural VEC framework (see lutkepohl2005new, p. 369). For the model at hand, the Beveridge-Nelson MA representation of $y_t$ takes the form:

equation[equation omitted — 244 chars of source]

where $y_0^*$ depends on the initial values, matrices $\Xi_j^*$, $j=0, 1, ...$, are some functions of the model parameters and are absolutely summable, thus forming such a convergent sequence that $\sum_{j=0}^{\infty}{\Xi^*_j \varepsilon_{t-j}}$ is an I(0) process. Finally, $\Xi=\beta_{\bot}(\alpha_{\bot}'\Gamma\beta_{\bot})^{-1}\alpha_{\bot}'$, with $\Gamma = I_n-\sum_{i=1}^{k-1}\Gamma_i$, is the matrix of the long-run multipliers so that $\Xi B$ contains the long-run effects of the structural innovations\footnote{Matrices $\alpha_{\bot}$ and $\beta_{\bot}$ span the orthogonal complements of $sp(\alpha)$ and $sp(\tilde{\beta})$, respectively, so they are $n\times(n-r)$ full column rank matrices, such that $\alpha'\alpha_{\bot} = 0$, $rank(\alpha, \alpha_{\bot}) = n$, $\tilde{\beta}'\beta_{\bot} = 0$, and $rank(\tilde{\beta}, \beta_{\bot}) = n$.}.

Notice that the rank of $\Xi$ equals $n-r$, which is also the rank of $\Xi B$, since $B$ is non-singular (see, e.g., johansen1996likelihood, and lutkepohl2005new for details). The property can by utilized to identify the shocks. Since $rank(\Xi B)= n-r$, this matrix can contain $r$ zero columns at most, allowing for maximally $r$ shocks to be only of a transient nature. To identify $B$, as many as $n(n-1)/2$ restrictions are required. Assuming that exactly $r$ shocks are transient, we impose $n(n-r)$ restrictions. Remaining restrictions that are needed to identify the shocks in both of the groups, need to be imposed on: $B$ (for the transient shocks; note that they correspond to the zero columns in $\Xi B$), and the non-zero columns of $\Xi B$ (for the permanent shocks); see, e.g., king1991stochastic, GonzaloNg2001, lutkepohl2005new.

Identification of the SVEC-MSH model

In this research, we identify structural shocks via Markovian volatility breaks, which contributes to the line of research initiated in a VAR framework by lanne2010structural, and later developed by, among others, lutkepohl2016structural, lutkepohl2020bayesian. In the first of the cited works, a stationary SVAR-MSH is considered within a frequentist estimation setting. On the other hand, lutkepohl2016structural shift the focus on cointegrated systems in a SVEC-MSH form (again, within the frequentist framework). Finally, lutkepohl2020bayesian develop a Bayesian approach to estimation and inference of a SVAR-MSH model, limiting their attention, however, only to the stationary case.

The point of departure for our considerations presented below, is Theorem 1 formulated in lutkepohl2020bayesian, providing conditions for ensuring the identification of a stationary structural VAR system with Markov-switching heteroskedasticity. Of note, in the cited paper, the so-called A-parameterization of a VAR system is employed (so that the instantaneous dependencies between the endogenous variables are explicitly modeled), as opposed to the B-parameterization followed in our present work (see Eq. ((ref))), so that the immediate responses of the endogenous variables to the shocks are explicitly modeled. Such a parameterization is a typical choice for cointegrated VARs in the VEC form, for being conducive to disentangling the permanent from only transient shocks (see lutkepohl2005new).

Finally, before we venture into the key matters underlying this section, let us invoke a remark often raised in the literature on various approaches to structural VAR identification, namely that distilling structural shocks through regime-switching heteroskedasticity (among other approaches, which use for example non-Gaussianity of the errors or more elaborate volatility structures such as GARCH; see kilian2017structural) is an attempt of merely a statistical identification of the shocks, and thus do not necessarily provide results that are of valid economic meaning, which remains to be figure out following up the statistical procedure (see, e.g., lutkepohl2016structural, lutkepohl2020bayesian). Nevertheless, a statistically identified VAR/VEC system enables testing traditionally imposed identifying restrictions (see, e.g., lanne2010structural, lutkepohl2016structural, kilian2017structural, lutkepohl2020bayesian).

To identify a two-state SVEC-MSH model, one could think of applying Theorem 1 formulated by lutkepohl2020bayesian, although formulated therein for a stationary (rather than cointegrated) structural VAR model with finite-state Markov-switching heteroskedasticity (see also lanne2010structural). Upon adaptation to our present notation, and assuming only a two-state system, we quote the theorem below.

theoremLet $\Sigma_m$, $m=1,2$, be positive-definite symmetric $n\times n$ matrices and $\Lambda_m=diag(\lambda_{m,1}, ..., \lambda_{m,n})$, $m=1,2$, be $n\times n$ diagonal matrices with positive diagonals. Suppose there exists a non-singular $n\times n$ matrix $B$ with unit main diagonal such that $\Sigma_m = B\Lambda_m B'$, $m=1,2$. Let $\omega_{2,i}=\lambda_{2,i}/\lambda_{1,i}$, $i=1,...,n$, be the i\textsuperscript{th} structural shock's (state 2) variance relative to state 1. Then, the $k^{th}$ row of $B$ is unique if $\forall_{i\in\{1,...,n\}\setminus \{k\}}\ \omega_{2,k} \ne \omega_{2,i}$.

Note that if the conditions stipulated in the above theorem hold for each of the rows of $B$, then the entire matrix $B$ is identified.

Since under the model parameterization considered in our paper, matrix $B$ features ones on its main diagonal (thereby restricting the signs of the shocks), Theorem (ref) should ensure the global identification of the system. However, consider the following (counter)example with $n=2$ variables only: $\lambda_1=\left(

array[array omitted — 22 chars of source]

\right)$, $\lambda_2=\left(

array[array omitted — 24 chars of source]

\right)$, and $B=\left(

array[array omitted — 42 chars of source]

\right)$, so that the following reduced-form state-specific covariance matrices are obtained: $\Sigma_1=\left(

array[array omitted — 50 chars of source]

\right)$ and $\Sigma_2=\left(

array[array omitted — 50 chars of source]

\right)$. The resulting relative (to state 1) structural variances are as follows: $\omega_{2,1}= \lambda_{2,1}/\lambda_{1,1}=\frac{1}{5}$ for the first variable, and $\omega_{2,2}= \lambda_{2,2}/\lambda_{1,2}=\frac{1}{7}$, respectively, so that $\omega_2=\left(

array[array omitted — 42 chars of source]

\right)'$. The identification condition stated in Theorem~\ref{Theorem_Lutke_Wozniak} would simply require that the difference (contrast) between these two ratios be non-zero, which is the case here, indeed: $\omega_{2,2}-\omega_{2,1}=-\frac{2}{35}\ne 0$. However, and to one's surprise, simple calculations show that the very same reduced-form covariance matrices, $\Sigma_1$ and $\Sigma_2$, can be obtained also for a different set of $\lambda_1$, $\lambda_2$ and $B$, namely: $\tilde{\lambda}_1=\left(

array[array omitted — 27 chars of source]

\right)$, $\tilde{\lambda}_2=\left(

array[array omitted — 27 chars of source]

\right)$ and $\tilde{B}=\left(

array[array omitted — 38 chars of source]

\right)$. The corresponding vector of the relative variances is now $\tilde{\omega}_2=\left(

array[array omitted — 109 chars of source]

\right)' = \left(

array[array omitted — 45 chars of source]

\right)'$, which indicates that $\tilde{\omega}_{2,2}-\tilde{\omega}_{2,1}=\frac{2}{35}\ne 0$, so the condition underlying Theorem~\ref{Theorem_Lutke_Wozniak} is also met here. Notice that $\tilde{\omega}_{2,2}-\tilde{\omega}_{2,1}=-(\omega_{2,2}-\omega_{2,1})$, which actually arises from nothing else but $\tilde{\omega}_2$ being equal to the reordered $\omega_2$. Indeed, $\tilde{\omega}_{2,1}=\omega_{2,2}$ and $\tilde{\omega}_{2,2}=\omega_{2,1}$.

It follows from the example above that: i) the theorem provided by lutkepohl2020bayesian actually does not ensure the global identification, enabling one only to identify the shocks locally up to their order; ii) apart from the condition already stipulated in Theorem (ref), the global identification requires restricting also the order of the relative changes of the state-dependent structural variances, $\omega_{2,i}$, $i=1, ..., n$. The latter remark leads to formulation of the following theorem (which can be viewed as a 'revised' version of the one formulated by lutkepohl2020bayesian) of the global identification of the structural two-state VAR/VEC-MSH model (we provide the proof in Appendix (ref)).

theoremLet $\Sigma_m$, $m=1,2$, be positive-definite symmetric $n\times n$ matrices and $\Lambda_m=diag(\lambda_{m,1}, ..., \lambda_{m,n})$, $m=1,2$, be $n\times n$ diagonal matrices with positive diagonals. Suppose there exists a non-singular $n\times n$ matrix $B$ with unit main diagonal such that $\Sigma_m = B\Lambda_m B'$, $m=1,2$. Let $\omega_{2,i}=\lambda_{2,i}/\lambda_{1,i}$, $i=1,...,n$, be the i\textsuperscript{th} structural shock's (state 2) variance relative to state 1. Then, under a fixed order of the elements of the vector $\omega_2=(\omega_{2,1},\omega_{2,2},\dots,\omega_{2,n})'$, the $k^{th}$ row of $B$ is unique if $\forall_{i\in\{1,...,n\}\setminus \{k\}}\ \omega_{2,k} \ne \omega_{2,i}$.

The only difference between the two theorems is this additional fixing of the order of $\omega_{2,i}$, $i=1, ..., n$. As we point it out in the proof of Theorem (ref) (see Appendix (ref)), the reason for which Theorem (ref) actually fails to guarantee the global identification is that in its proof it has been assumed that the state-dependent variances, $\lambda_{m,i}$ ($m=1,2$ and $i=1,\ldots,n$) are unique. Indeed, under such an assumption, it can be shown that there exists also a unique matrix $B$. However, the restriction of unique structural regime-specific variances does not appear justifiable (contrary to the uniqueness of the reduced-form regime-specific covariance matrices, obviously). And so, as the proof in Appendix (ref) presents, under free $\lambda_{m,i}$s, the global identification of $B$ (and $\lambda_{m,i}$s) requires fixing the order of the relative (to state 1) structural variances.

In view of the above, one can ponder on possible severity with which the failure to restrict the order of $\omega_{2,i}$s may affect estimation results. Obviously, the issue remains of a rather empirical nature, with particular outcomes hinging conceivably on assumed prior distributions, data set at hand, and even the properties of designed posterior distribution sampler. We illustrate some aspects of this otherwise general problem in Section 3 of the Supplementary material, through simulated-data-based studies (also using the case delivered above as the counterexample for Theorem (ref)).

Bayesian two-state SVEC-MSH model

This section is devoted to specification of a Bayesian two-state SVEC-MSH model. In Section (ref), we write down the likelihood function, preceded by introduction of a suitable matrix notation. Section (ref) displays the prior structure of our choice. A detailed presentation of the MCMC posterior sampling routine (combining the Gibbs and Metropolis-Hastings algorithms), along with the underlying full conditional posteriors, is deferred to the Supplementary material (see Section 2 therein).

The likelihood function

Writing down the likelihood function for the model at hand requires introduction of some notation conducive thereto. In that regard, our framework draws on jochmann2015regime, who developed a Bayesian Markov-switching VEC model in the reduced (instead of structural) form. Also note that their approach enables regime changes of all the model parameters (i.e. including the ones of the conditional mean), whereas in the structural VEC-MSH model considered here, only switches of the conditional covariance matrix are allowed, with the remaining parameters (and thus, the shocks transmission mechanism, in particular) held constant throughout (see also sugita2016bayesian). Such a limitation, however, is typically entertained in the literature on structural shocks identification through heteroskedasticity (see, e.g., lutkepohl2020bayesian in the context of Bayesian, although stationary VAR-MSH models; and lutkepohl2016structural in the context of cointegrated, although frequentist VAR-MSH framework).

For each of the two states ($m=1, 2$), define the matrix of the observations allocated to state $m$:

equation[equation omitted — 112 chars of source]

along with its vectorization

equation[equation omitted — 215 chars of source]

where $T_m$ is the number of the observations assigned to state $m$. All observations are then collected in

equation[equation omitted — 165 chars of source]

where $\tilde z_0 = \left(

array[array omitted — 44 chars of source]

\right)$, $\tilde z_1 = \left(

array[array omitted — 56 chars of source]

\right)$, $\tilde z_2 = \left(

array[array omitted — 56 chars of source]

\right)$,\\ $\gamma=vec(\Gamma)$, $\tilde u = \left(

array[array omitted — 30 chars of source]

\right)$, $u^{(m)} = vec\left(U^{(m)}\right) = \left(B\otimes I_{T_m}\right)vec(E^{(m)})$, for $m=1,2$, and $S=(S_1, S_2, \dots, S_T)$, denotes the vector of the Markov chain states $S_t\in\{1, 2\}$ at $t=1, 2, \dots, T$. By $\tilde\Sigma$ we denote the (block-diagonal) conditional covariance matrix of $\tilde u$ (given the state vector, $S$), i.e. $\tilde\Sigma = diag\left(B\Lambda^{(1)} B'\otimes I_{T_1},\, B\Lambda^{(2)} B'\otimes I_{T_2}\right)$.

Using the above notation, we can write down the likelihood function (conditional on the vector of states):\footnote{Initial conditions: $y_{-p+1},\dots,y_{-1}, y_0$ (fixed at pre-sample values) and $S_0$ (a random variable), are suppressed from the notation.}

equation[equation omitted — 657 chars of source]

with $\theta$ collecting all model parameters, i.e. $\theta = (\alpha, \beta, \Gamma, b, \lambda_1, \lambda_2, p_{11}, p_{22})$, where $b$ represents only the free entries of matrix $B$ (so that $vec(B) = Qb +q$, where $Q$ is an $n^2\times d_b$ known matrix, and $q$ is an $n^2\times 1$ known vector), and $\lambda_m=diag(\Lambda_m)$, $m=1, 2$. Notice that $\theta$ involves directly the structural variances for both states ($\lambda_1$ and $\lambda_2$) rather than only the first-state variances ($\lambda_1$) along with their relative changes in the second state ($\omega_2$). Although the latter parameterization was of choice in lutkepohl2020bayesian, where it proved conducive to a testing setup developed therein, we decide to work here directly with structural variances in both states instead, as we prefer to introduce prior information directly on the regime-changing parameters. That way, we can ensure that the joint prior distribution for these parameters is symmetric (with respect to relabeling the states), and thereby ensure the shocks identification to arise only from an order restriction imposed explicitly on $\frac{\lambda_{2,i}}{\lambda_{1,i}}$, and not from asymmetric prior information for $\lambda_1$ and $\lambda_2$ (introduced inadvertently when parameterizing the model through $\lambda_1$ and $\omega_2$, instead).

Obviously, under no further constraints on the structural variances, the likelihood function given by Eq. ((ref)) suffers from a perfect symmetry with respect to relabeling of the states. To ensure a unique identification of the regimes, we impose relevant restrictions on the prior distribution, as discussed in the following section.

The prior distribution

Specification of our Bayesian two-state SVEC-MSH statistical model is completed with a prior distribution of the following structure (with specific densities of our choice provided later on):

equation[equation omitted — 206 chars of source]

where the priors for $\alpha_*$ (i.e. the non-normalized matrix of adjustment coefficients), $\beta_*$ (the non-normalized matrix of cointegrating vectors) and $\Gamma$ are independent, but only up to the restriction on the spectral radius $\rho(\mathbf{A})$ of the companion matrix $\mathbf{A}=\mathbf{A}(\alpha_*,\beta_*,\Gamma)$ to be less or equal one, thereby preventing the explosiveness of the underlying VAR process (though still allowing for its non-stationarity): $$p(\alpha_*,\beta_*,\Gamma)\propto p(\alpha_*)p(\beta_*)p(\Gamma)\mathbb{I}(\rho(\mathbf{A})\leq 1),$$ with $\mathbb{I}(\cdot)$ denoting the indicator function.

Notice that, according to Eq. ((ref)), for the three: $b$, $\lambda_1$ and $\lambda_2$, which are the key parameters to model identification, we specify hierarchical priors, with the approach commonly recognized for its general merits in terms of facilitation of Bayesian estimation. In particular, the joint prior distribution of the structural shocks' state-specific variances takes the form:

equation[equation omitted — 296 chars of source]

where $s^\lambda_m=(s^\lambda_{m,1},s^\lambda_{m,2},\dots,s^\lambda_{m,n})$ collects the hyperparameters for state $m\in\{1, 2\}$, and $\mathcal{U}^n=\left\{(x_1, x_2, ..., x_n)'\in \mathbb{R}^{n}: x_1<x_2<\dots<x_n \right\}$ so that the expression $\mathbb{I}( \omega_2\in \mathcal{U}^n)$ imposes a uniqueness restriction of the form: $\omega_{2,1}<\omega_{2,2}<\dots,<\omega_{2,n}$. Note, however, that from a formal perspective any other order defining the $\mathcal{U}^n$ set (and thus, the form of the inequality between $\omega_{2,1}, ..., \omega_{2,n}$) is equally acceptable, leaving any specific choice of the ordering to be resolved empirically in a 'data-driven' way. The other indicator function term in Eq. ((ref)), i.e. $\mathbb{I}\left(\lambda_{1,l}>\lambda_{2,l} \right)$ for a given $l\in\{1, 2, ..., n\}$ (a particular choice of which remains of empirical nature), intends to ensure the regime identification, thereby preventing the label switching of structural variances. Finally, notice that both restrictions introduce some dependence between the otherwise conditionally (given the hyperparameters) a priori independent variances.

With details deferred to Section 1 of the Supplementary material, we only notice here that specific prior densities of our choice are the ones typically employed in VAR/VEC-MSH modeling, with their conditional conjugacy (except for $b$) enabling a use of a Gibbs sampler for posterior simulation (see Section 2 of the Supplementary material for a detailed presentation of the MCMC algorithm combining the Gibbs sampler with a Metropolis-Hastings step for sampling $b$). Specifically, the priors for the cointegration parameters, $\alpha_*$ and $\beta_*$, are standard in Bayesian cointegration literature (see koop2009efficient), and have earlier been used also by jochmann2015regime in reduced-form VEC-MSH models. A matrix normal prior distribution for the remaining VAR parameters, collected in $\Gamma$, is also a typical choice, with $\Omega_\Gamma$ admitting an optional introduction of some form of shrinkage to facilitate parsimony of the VAR lags structure (we apply the approach in Section (ref) of this paper as well as in Section 3 of the Supplementary material; see also jochmann2015regime, koop2009efficient, lutkepohl2020bayesian).

Empirical illustration

Data and model setting

For an empirical illustration of our methodology, we use one of the models analyzed previously by velinov2013can (inspired by the works of binswanger2004stock and louis2010stock) and later by lutkepohl2016structural, who focus on SVAR-MSH and SVEC-MSH specifications derived for a dividend discount model (DDM) for US data, relating the expected future discounted payoffs (dividends) to the real economic activity quantities such as the real GDP, industrial production, company earnings etc., thus contributing to the line of research initiated by ball1968empirical. The model of our choice is the one denoted as Model VI in binswanger2004stock, louis2010stock, Model IV in velinov2013can, and Model II in lutkepohl2016structural, and consists of $n=3$ variables: real earnings ($E_t$), real interest rates ($r_t$), and real stock prices ($s_t$), so that $y_t=(E_t, r_t, s_t)'$. Comprising only three variables, the model may arguably appear rather simple (if not simplistic). Nevertheless, it is still quite popular in asset pricing, and as such it would serve well for our illustrative purposes.

Following velinov2013can and lutkepohl2016structural, we use data from Robert Schiller's website\footnote{http://www.econ.yale.edu/shiller/data.htm} for earnings, but also for stock prices proxied by S&P500 quotations (note that S&P500 is also used in the two cited works, but with the data drawn from the Federal Reserve Economic Database, FRED). Both series are available already in real terms on the page, deflated by the CPI inflation rate. For the interest rate, we take the effective federal funds rate (% p.a.), with the data obtained from FRED\footnote{Board of Governors of the Federal Reserve System (US), Federal Funds Effective Rate [FEDFUNDS], retrieved from FRED, Federal Reserve Bank of St. Louis; https://fred.stlouisfed.org/series/FEDFUNDS, July 25, 2022.}, deflated by using the CPI inflation rate. The data is quarterly (last business day of quarter), seasonally unadjusted, and in logs (except for the real interest rate), ranging from 1960:I to 2021:II, thus covering a series of the US economy recessions, including the global financial crisis of 2007-2009 and the outbreak of the Covid-19 pandemic. Note that the data range here differs markedly from the one in velinov2013can and lutkepohl2016structural, who use series from 1947:1 to 2012:I. Finally, to facilitate accounting for any potential seasonality not captured by seasonal dummies, we arbitrarily set $p=5$ lags in the underlying VAR model (thus, four lags in its VEC form), which necessitates sparing the first five data points available (i.e. from 1960:I to 1961:I) for the sake of initial conditions, which yields the actual sample's size of $T=241$. The specified number of lags is supported by a visual inspection of the ACF and PACF plots for the data (particularly the interest rate), and auxiliary examination of the residuals from non-Bayesian, homoskedastic VAR models of orders 1-6, with the results left unreported here for the sake of brevity.

General estimation results

All results presented below are based on $500\,000$ MCMC draws, preceded by as many burn-in iterations. The hyperparameters of the prior distributions are set the same as in the simulation study presented in the Supplementary material (see Section 3 therein). The states are identified through a restriction of the form: $\lambda_{11}>\lambda_{21}$, i.e. the variance of the first structural shock in the first state is higher than the shock's variance in the second state: $Var(\varepsilon_{t1}|S_t=1, \theta)>Var(\varepsilon_{t1}|S_t=2, \theta)$. For unique identification of the model, the relative (to state 1) structural variances are subject to the constraint: $\omega_{2,1}<\omega_{2,2}<\omega_{2,3}$, with the choice of the ordering based on an ancillary examination of the posterior results obtained for a model without the constraint (yet retaining the state-identification restriction), strongly suggesting that particular order. Therefore, it appeared the most natural choice, remaining in tune with the data.

For choosing the number of cointegration relationships, we resort to the Savage-Dickey density ratio (SDDR; see, e.g., verdinelli1995computing, mulder2022generalization). The results (deferred to the Supplementary material; see Section 4.1 therein) indicate a huge superiority of the cointegrated models, with the specification featuring only one cointegration relation outperforming the others (although $r=2$ is a very close second), and therefore being selected here for further analysis.

Figure (ref) presents the modeled data (with the initial conditions trimmed off) along with the posterior probabilities of the system residing in the first state (featured by a higher volatility of structural shocks), i.e. $Pr(S_t=1\vert y)$, $t=1,...,T$, and additionally, with indicated NBER-dated recessions of the US economy\footnote{https://www.nber.org/research/data/us-business-cycle-expansions-and-contractions}. The identification of the regimes is rather clear-cut, with apparent domination of the lower-volatility state, which also translates into its higher persistence ($Me(p_{22}\vert y)\approx 0.948$). Not surprisingly, the higher-volatility regime is of a more transient nature ($Me(p_{11}\vert y)\approx 0.784$, with a more diffuse posterior distribution, see panels (a,b) in Figure (ref)), largely corresponding to the NBER-dated contractions. Apparent additional switches to the first state, indicating a rise in the structural shocks' volatility, took place in 1987:I-IV and 2002:IV-2004:I. The former can be attributed to a relatively strong drop in stock prices accompanied by a relatively large increase in earnings (although both changes can only vaguely be discerned in Figure (ref)), while the latter -- to some abrupt changes in earnings. As can be inferred from the posterior distributions of $\lambda_{m,i}$, $m=1,\,2$, $i=1,\, 2,\, 3$ (see panels (c-h) in Figure (ref)), all the three structural shocks are conditionally heteroskedastic, exhibiting markedly higher volatilities in the first regime, with the posterior medians of $\lambda_{1,i}$, $i=1,\, 2,\, 3$, exceeding the ones of $\lambda_{2,i}$ approximately by one to two orders of magnitude. This, in turn, leads to pronounced relative differences between these sets of parameters, reflected in the posterior probability of $\omega_{2,i}$, $i=1,\,2,\,3$, amassed entirely far from (and less than) unit. As noticed in Figure (ref), the posterior distributions of the contrasts: $\omega_{2,2}-\omega_{2,1}$ and $\omega_{2,3}-\omega_{2,1}$, display a definitive separation from zero, thereby implying a very clear distinction between the first and the second shock. The results obtained for the last contrast, $\omega_{2,3}-\omega_{2,2}$, may appear somewhat less convincing due to a part of the posterior adhering to zero, the value of which, however, still falls beyond the 95% highest posterior density interval. Moreover, the p-value of a Lindley-type test for $\omega_{2,3}-\omega_{2,2}=0$ equals 0.046 (with the test statistics at 3.986). Therefore, with a high posterior probability one may still assume that the second and the third shock are disentangled.

With respect to the variables' instantaneous reactions to the structural shocks (see Figure 4 in the Supplementary material), it appears that only the stock prices' reaction to the first shock is significant (and positive), which will emerge highly intuitive in view of the economic identification of the shocks attempted below. On the other hand, the attention should be paid not only to the $B$ matrix itself, but also on the long-run matrix, $\Xi B$, which is key in this empirical example (see the following subsection).

Structural analyses and economic interpretation of the shocks

Typically, in the model at hand, the identifying restrictions imposed on the long-run matrix assume the form (see, e.g., lutkepohl2016structural):

equation[equation omitted — 136 chars of source]

with the first two shocks identified as fundamental and related to the real earnings and the real interest rate, respectively (therefore, it is assumed that the second shock leaves the earnings unaffected in the long term). On the other hand, the third one is identified as non-fundamental (speculative), thereby of only a transient impact on the system. Below, we attempt to name the shocks identified through our statistical model, by means of typical structural analysis tools, and to assess the consistence of the results with the above-mentioned traditional long-run restrictions.

Interestingly enough, pairing ((ref)) with the impulse response functions (IRFs; see Figure (ref) as well as Figures 5-6 in Section 4.2 of the Supplementary material) indicates that the traditionally postulated restrictions are rejected by our empirical results, and are therefore of little use for the purpose of figuring out the economic meaning of the shocks\footnote{Notice, however, that the results seen in these figures do not need to be ordered column-wise in strict adherence with $\Xi B$ given by ((ref)).}. Arguably, it could already put into question the very possibility of providing here sound economic interpretations for the otherwise statistically identified shocks (see, e.g., lutkepohl2020bayesian). However, other structural analysis tools, including forecast error variance decomposition (FEVD) and the point estimates of structural shocks, may still be of assistance.

According to the FEVD results in Figure (ref), the first shock's contribution to the earnings' forecast error variance is overwhelming both in the first and second state, which implies a fundamental role of the shock. Further, regarding the stock prices forecast error, it is interesting to observe a huge contribution of the third shock in the second (i.e. lower-volatility) regime, which indicates a speculative nature of the shock (see, e.g., binswanger2004stock (binswanger2004stock, binswanger2004important, binswanger2004G7), louis2010stock). Such a classification is further corroborated by a decreasing impact of the shock for forecasts of longer horizons (see, e.g., lee1998permanent).

Based on the above, we can identify the first two shocks as fundamental (to the real earnings and the real interest rate, respectively), while the third one -- as speculative. This line of reasoning seems supported also by visual inspection of the shocks' point estimates, presented in Figure (ref). The path for the second shock reveals the strongest swings at the turn of the 1970s and 1980s, which relates 'neatly' to the extremely sharp movements in the interest rate featuring that turbulent times in the US economy. This second type of fundamental impulses does not seem to contribute meaningfully to the forecast error variance of the stock prices (see Figure (ref)), which is consistent with earlier findings by lee1995response (lee1995fundamentals, lee1995response, lee1998permanent), and binswanger2004stock (notice, however, that louis2010stock point the opposite).

As seen in Figure (ref), the first shock manifests the most prominent values over the years 2008-2009, indicating their apparent association with the real earnings that suffered severe drops back then, followed soon by a swift rebound (see Figure (ref)). However, these strong movements in the earnings were accompanied by highly volatile, both downward and upward shifts in the interest rates (see Figure (ref)). In particular, a strong negative impulse of 2010, seen for the first shock in Figure (ref), may be related to a sudden drop in the interest rate (see Figure (ref)). These findings effectively prevent us from a clear-cut classification of the first shock as one that strictly relates to the earnings only. Instead, it appears more fair to conclude that these two economic types of shocks: to the real earnings and the interest rate, are mixed up to some degree in this otherwise statistically identified shock.

Finally, a rather irregular path of the third shock's estimates (see Figure (ref)) seems to resonate with its only speculative nature.

Recall here that the traditional long-run zero restrictions are not met in our model. As argued by lutkepohl2016structural, this may hinder valid IRF interpretations, thus adding further trouble in attempts to deliver sound and unique economic 'identities' of the shocks. The 68% posterior quantile bands around the posterior medians of impulse response functions are displayed in Figure (ref), with the blue and yellow bands corresponding to the first and second state, respectively (notice that the first- and second-state IRFs actually coincide up to a scaling, which is an artifact of only time-invariant autoregressive parameters, and thus the transmission mechanism in our model's specification). It appears that the first shock, identified earlier as fundamental and earnings-related, exerts a positive and lasting influence on both the earnings and stock prices, while remaining immaterial to the interest rate. Obviously, the result is intuitive (see also, e.g., lee1998permanent).

The other of the two fundamental shocks, related to the interest rate, positively affects the rate permanently and throughout the horizons, while remaining neutral to the earnings for about 2.5 years, yet causing them decline slowly afterwards. The result seems economically valid and to generally agree with one's expectations.

On the whole, despite the results breaching the traditional restrictions to some extent, the first two shocks can still be regarded as fundamental. The third one, on the other hand, may raise some doubts. As implied by Figure (ref), a positive impulse exerts a sustained, flat-lining throughout the horizons, and positive influence on the stock market, possibly indicating a permanent discrepancy (bubble) between the stock prices and fundamentals (although such a hypothesis should be formally investigated and tested within a model incorporating also dividends; see, e.g., chung1998fundamental). Concerns are raised also by the earnings' response to the shock, with no significant reaction for about 1.5 year, followed by a systematic increase, perpetuating also in the long run (see Figures 5-6 in Section 4.2 of the Supplementary material), which simply questions only a speculative nature of the shock. Therefore, it may be the case that the shock still includes some 'admixture' of the fundamental impulses. On that note, binswanger2004stock points out that the contribution of fundamental shocks to forecast variance is underestimated in a model featuring earnings, and thus, including in the model some real activity variable(s) instead, such as GDP or industrial production, could deliver more conclusive results.

Conclusions

In the paper, we developed a Bayesian framework for estimation and inference in a two-state structural VEC model with Markov-switching heteroskedasticity. In the process, we revisited the identification conditions stated previously in Theorem 1 by lutkepohl2020bayesian in the context of (Bayesian) stationary VARs, to find them warranting only a local identification. We pinpointed the lacking restriction and showed that for the global identification of structural impulses, the ordering of the relative variances across the system's variables needs to be fixed.

Results of the simulated-data-based study (see Section 3 of the Supplement material) indicate that disregarding the additional restraint may manifest in more or less pronounced irregularities in the marginal posteriors of key model parameters, with all of the former being removed once a relevant ordering restriction is imposed on $\omega_{2,1}, \omega_{2,2}, \ldots, \omega_{2,n}$. On the other hand, the empirical illustration (Section (ref)) shows that introducing all relevant restrictions on the prior distribution does not necessarily translates into fully convincing and economically sound identification of shocks that are statistically identified nonetheless. Thereby, combining our current setup with conventional approaches to identification, much in the spirit of herwartz2014structural, could prove remedial here.