EconBase
← Back to paper

Structural Gaussian mixture vector autoregressive model with application to the asymmetric effects of monetary policy shocks

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

74,713 characters · 12 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
titlepage\begin{center} \linespread{1.2} {Structural Gaussian mixture vector autoregressive model with application to the asymmetric effects of monetary policy shocks}\\ {Savi Virolainen}\\ {University of Helsinki}\\ \begin{abstract} A structural Gaussian mixture vector autoregressive model is introduced. The shocks are identified by combining simultaneous diagonalization of the reduced form error covariance matrices with constraints on the time-varying impact matrix. This leads to flexible identification conditions, and some of the constraints are also testable. The empirical application studies asymmetries in the effects of the U.S. monetary policy shock and finds strong asymmetries with respect to the sign and size of the shock and to the initial state of the economy. The accompanying CRAN distributed R package gmvarkit provides a comprehensive set of tools for numerical analysis.\\[1cm] \center{Acknowledgements}\\ This work was supported by the Academy of Finland under Grant 308628. The author thanks Markku Lanne, Mika Meitz, and Pentti Saikkonen who helped to improve this paper substantially. The author also thanks Henri Nyberg and Antti Ripatti for the useful comments.\\[1cm] \noindentKeywords: structural nonlinear autoregression, structural mixture VAR, regime-switching, Gaussian mixture, mixture autoregression, monetary policy shock\%[2.0cm] \end{abstract} \begingroup \footnote{Contact address: Savi Virolainen, Faculty of Social Sciences, University of Helsinki, P. O. Box 17, FI–00014 University of Helsinki, Finland; e-mail: [email removed].} \addtocounter{footnote}{-1} \endgroup \begingroup \footnote{The author has no conflict of interest to declare.} \addtocounter{footnote}{-1} \endgroup \end{center}

Introduction

Tracing out the effects of an economic shock is a major task in econometrics. A popular approach is to consider a set of key variables and utilize a structural vector autoregressive (SVAR) or structural error correction (SVEC) model for the purpose. They have well established theoretical grounds Kilian+Lutkepohl:2017 and are accommodated by many of the popular statistical software packages. Linear SVAR and SVEC models are not, however, suitable for modelling series in which the underlying data generating dynamics are nonlinear or the shocks have asymmetric effects in different states of the economy. Models capable of capturing such features include mixture models, such as the mixture vector autoregressive model Fong+Li+Yau+Wong:2007, the mixture periodic vector autoregressive model Bentarzi+Djeddou:2014, the Gaussian mixture vector autoregressive (GMVAR) model Kalliovirta+Meitz+Saikkonen:2016, and the logit mixture vector autoregressive model Burgard+Neuenkirch+Nockel:2019.

This paper introduces a structural version of the GMVAR model. In the structural GMVAR (SGMVAR) model of autoregressive order $p$, the regime-switching dynamics are endogenously determined by the full distribution of the previous $p$ observations. Specifically, the greater the relative weighted likelihood of a regime is, the more likely the process is to generate an observation from it. This facilitates associating statistical characteristics and economic interpretations to the regimes. The SGMVAR model thus allows the regime-switches to depend on a richer set of statistical characteristics of the data than many of the popular threshold VAR Tsay:1998 and smooth transition VAR Anderson+Vahid:1998 models in which regime-transitions often depend only on the level of the transition-variables. The specific formulation of the mixing weights also leads to attractive theoretical properties, such as ergodicity and fully known stationary distribution of $p+1$ consecutive observations.

The effects of the structural shocks depend on the initial values of the included variables, and they are also allowed to vary according to the sign and size of the shock due to possibly resulting regime-switches. Consequently, the (generalized) impulse response functions reflect the prevailing macroeconomic conditions that are transmitted to the regime-switching probabilities through the level, variability, and temporal as well as contemporaneous dependence of the past observations. Because the shocks may have asymmetric effects with respect to their size, the conditional heteroskedasticity of the reduced form error needs to be controlled for. Therefore, the impact matrix of the SGMVAR model is time-varying and constructed so that it captures the conditional heteroskedasticity of the reduced form error, thereby enabling to standardize the conditional variance of each structural shock to a constant. The initial effects of a constant-sized structural shock are, hence, amplified according to the conditional variance of the reduced form error, also reflecting the prevailing state of the economy.

Identification of the shocks requires that they are simultaneously orthogonalized in all regimes. We show that together with any constant standardization of the structural shock's conditional variance, this condition generally leads to a unique identification of the impact matrix up to ordering of its columns and changing all signs in a column. Thus, as long as one is willing to impose the assumption of a single (time-varying) impact matrix, the columns of the impact matrix unambiguously characterize the estimated impact effects of the shocks without further constraints. The identification does not, however, reveal which column of the impact matrix is related to which shock. Since the impact matrix is also subject to estimation error, further constraints may be needed for labelling the shocks. The constraints are testable, as they are overidentifying.

In order to formulate the impact matrix and the identification conditions, it is convenient to utilize the well known matrix decomposition Muirhead:1982 proposed by Lanne+Lutkepohl:2010 and Lanne+Lutkepohl+Maciejowska:2010 for a similar identification problem. Lanne+Lutkepohl:2010 assume that the reduced form error covariance matrices admit this decomposition, then show that the shocks are statistically identified, and finally test conventional zero constraints that lead to economically interpretable shocks. Lanne+Lutkepohl+Maciejowska:2010, in turn, note that the shocks are readily identified when the matrix decomposition is imposed to the reduced form error covariance matrices. Our approach differs from them in that we obtain locally identified structural shocks by directly investigating the properties of the impact matrix. We also provide a general set of conditions for identifying any subset of the shocks that allows for using sign constraints alone or together with zero constraints. Moreover, we (partially) relax a technical condition required for statistical identification of the model and allow identification of a subset of the shocks when the model is only partially identified.

Our empirical application studies asymmetries in the expected effects of monetary policy shocks in the U.S. using a quarterly series covering the period from 1954Q3 to 2021Q4. Our SGMVAR model identifies two regimes: a stable inflation regime and an unstable inflation regime. The unstable inflation regime is characterized by high or volatile inflation, and it mainly prevails in the 1970's, early 1980's, during the Financial crisis, and in the COVID-19 crisis from 2020Q3 onwards. The stable inflation regime, in turn, is characterized by moderate inflation, and it prevails when the unstable inflation regime does not. We find the effects of the monetary policy shock relatively symmetric in the unstable inflation regime, as it rarely causes a switch to the stable inflation regime. A contractionary (expansionary) monetary policy shock appears to first increase (decrease) inflation after which the inflation significantly decreases (increases) for several years. The strong contraction (expansion) in the cyclical component of the GDP lasts for roughly three years and is followed by a small short-term expansion (contraction) before the response decays to zero.

In the stable inflation regime, the (generalized) impulse responses are strongly asymmetric with the respect to the sign and size of the monetary policy shock as well as to the initial state of the economy. A contractionary shock causes, on average, roughly a three-year hump-shaped contraction of the GDP, but it also seems to increase inflation by driving the economy towards the unstable inflation regime. A small expansionary shock does not move prices much on average, but a large expansionary shock often drives the economy towards the unstable inflation regime and propagates high and persistent inflation. The high inflation is followed by a significant monetary policy tightening and a persistent contraction of the GDP after the initial expansion. On average, the real effects of the monetary policy shock are found somewhat stronger in the stable inflation regime than in the unstable inflation regime.

The GMVAR model has been previously applied in impulse response analysis by Kalliovirta+Malinen:2020, who identify the shocks by constraining the reduced form error covariance matrices and allow the impact responses of the variables to vary relative to each other across the regimes. Our assumption of a common (time-varying) impact matrix for all the regimes constraints the relative magnitudes of the impact responses of the variables to be time-invariant (for each shock), but it leads to flexible identification conditions and enables to test the validity of the identifying constraints. Kalliovirta+Malinen:2020 estimate the impulse response functions for each regime of the GMVAR model separately as if each of them was a linear VAR. In contrast, we allow the regime to switch as a result of a shock and estimate the true (generalized) impulse response functions of the non-linear VAR.

Structural mixture VARs, in general, have been previously applied for studying to the effects of monetary policy shocks at least by Burgard+Neuenkirch+Nockel:2019, who proposed a mixture VAR with logistic mixing weights and Cholesky identified shocks. As opposed to Burgard+Neuenkirch+Nockel:2019, our identification scheme is more flexible in the sense that it does not require many (or necessarily any) zero constraints on the impact effects of the shocks. Moreover, in our model the regime-switching probabilities depend on the full distribution of the preceding $p$ observations instead of just on the level of the switching-variables.

The rest of this paper is organized as follows. Section (ref) defines the reduced form GMVAR model. In Section (ref), the structural GMVAR model is first introduced. Then, identification of the shocks and estimation of the model parameters are discussed. Section (ref) discusses impulse response analysis and describes the generalized impulse response function (GIRF) Koop+Pesaran+Potter:1996. Section (ref) presents the empirical application and Section (ref) summarizes. Appendices provide proofs for the stated lemma and propositions, a Monte Carlo algorithm for estimating the GIRF, and details on the empirical application. Finally, we have accompanied this paper with the CRAN distributed R package gmvarkit gmvarkitnormal, which is comprehensively documented and provides a comprehensive set of tools for numerical analysis of the model.

Reduced form GMVAR model

To build theory and notation, consider first the reduced form GMVAR model introduced by Kalliovirta+Meitz+Saikkonen:2016. Let $y_t$ ($t=1,2,...$) be the $d$-dimensional time series of interest and $\mathcal{F}_{t-1}$ denote the $\sigma$-algebra generated by the random vectors $\lbrace y_{t-j}, j>0 \rbrace$. For a GMVAR model with $M$ mixture components and autoregressive order $p$, we have

align[align omitted — 207 chars of source]

where $\phi_{m,0}\in\mathbb{R}^{d}$ are intercept parameters, $\Omega_m$ are positive definite covariance matrices, and for each $m$, the coefficient matrices $A_{m,i}$, $i=1,...,p$, are assumed to satisfy the usual stability condition

equation[equation omitted — 113 chars of source]

which guarantees stationarity of the component processes. The unobservable regime variables $s_{1,t},...,s_{M,t}$ are such that at each $t$, exactly one of them takes the value one and the others take the value zero according to the conditional probabilities $\mathrm{P}(s_{m,1}=1|\mathcal{F}_{t-1})\equiv\alpha_{m,t}$ that satisfy $\sum_{m=1}^{M}\alpha_{m,t}=1$. The normally and independently distributed (NID) errors $u_{m,t}$ are assumed independent of $\mathcal{F}_{t-1}$, and conditional on $\mathcal{F}_{t-1}$, $(s_{1,t},...,s_{M,t})$ and $u_{m,t}$ are independent.

The definition ((ref))-((ref)) implies that at each $t$, the process generates an observation from one of its mixture components, a linear VAR process, that is randomly selected according to the probabilities given by the mixing weights $\alpha_{m,t}$. Denoting $\boldsymbol{y}_{t-1}=(y_{t-1},...,y_{t-p})$, the mixing weights are defined as Kalliovirta+Meitz+Saikkonen:2016

equation[equation omitted — 268 chars of source]

where $\alpha_1,...,\alpha_M$ are mixing weight parameters that satisfy $\sum_{m=1}^M \alpha_m=1$ and $n_{dp}(\cdot;\boldsymbol{1}_p\otimes \mu_m, \boldsymbol{\Sigma}_m)$ is the density function of the $dp$-dimensional normal distribution with mean $\boldsymbol{1}_p\otimes \mu_m$ and covariance matrix $\boldsymbol{\Sigma}_m$. The symbol $\boldsymbol{1}_p$ denotes a $p$-dimensional vector of ones, $\otimes$ is Kronecker product, $\mu_m=(I_d - \sum_{i=1}^pA_{m,i})^{-1}\phi_{m,0}$, and the covariance matrix $\boldsymbol{\Sigma}_m$ is given in Lutkepohl:2005, Equation (2.1.39), but using the parameters of the $m$th component process. That is, $n_{dp}(\cdot;\boldsymbol{1}_p\otimes \mu_m, \boldsymbol{\Sigma}_m)$ corresponds to the density function of the stationary distribution of the $m$th component process.

The mixing weights are thus weighted ratios of the component process stationary densities corresponding to the preceding $p$ observations. This implies that the greater the weighted relative likelihood of a regime is, the more likely the process is to generate an observation from it. This facilitates associating statistical characteristics and economic interpretations to the regimes. In addition to the (generalized) impulse response functions of the observable variables, the responses of the mixing weights may therefore be of interest. The definition of the mixing weights also leads to attractive theoretical properties such as ergodicity and full knowledge of the stationary distribution of $p+1$ consecutive observations Kalliovirta+Meitz+Saikkonen:2016. Specifically, the stationary distribution of the process $\boldsymbol{y}_{t}=(y_{t},...,y_{t-p+1})$ is a mixture of $dp$-dimensional normal distributions that is characterized by the density

equation[equation omitted — 132 chars of source]

The knowledge of the stationary distribution is taken advantage of in the impulse response analysis in Section (ref) and Appendix (ref).

Structural GMVAR model

The model setup

Consider the GMVAR model defined in ((ref))-((ref)). We focus on the "B-model" setup and write the structural GMVAR model as

equation[equation omitted — 123 chars of source]

and

equation[equation omitted — 400 chars of source]

where the probabilities are expressed conditionally on $\mathcal{F}_{t-1}$ and $e_t$ is an orthogonal structural error. Unlike in the conventional SVAR analysis, the invertible $(d\times d)$ "B-matrix" (or impact matrix) $B_t$, which governs the contemporaneous relations of the shocks, is time-varying and a function of $y_{t-1},...,y_{t-p}$. This enables to amplify a constant-sized structural shock according to the conditional variance of the reduced form error, which varies according to the mixing weights. Appropriate modelling of conditional heteroskedasticity in the B-matrix is of interest, because the (generalized) impulse response functions may be asymmetric with respect to the size of the shock.

We have $\Omega_{u,t}\equiv\text{Cov}(u_t|\mathcal{F}_{t-1})=\sum_{m=1}^M\alpha_{m,t}\Omega_m$, while the conditional covariance matrix of the structural errors $e_t=B_t^{-1}u_t$ (which have a mixture normal distribution and are not IID but martingale differences and therefore uncorrelated) is obtained as

equation[equation omitted — 121 chars of source]

The B-matrix $B_t$ should therefore be chosen so that the structural shocks are orthogonal regardless of which regime they come from. We will next discuss the properties of any such B-matrix that solves the diagonalization problem. Then, we present a locally unique solution under a constant normalization of the structural error's conditional variance. After that, in the following two subsections, we will discuss global identification of the shocks, allowing also only partial identification of the model.

Specifically, we show that our model readily identifies the B-matrix up to ordering of its columns and changing all signs in a column, but it is not revealed which column of the B-matrix is related to which shock. The identification follows from the assumption $e_t=B_t^{-1}u_t$, which (as we show in this section) implies that for each shock the relative magnitudes of the impact responses of the variables stay constant over time.\footnote{See Kilian+Lutkepohl:2017 for a discussion on identification by heteroskedasticity in a linear VAR model.} This is different to the conventional SVAR setup, where the identification of the B-matrix requires further constraints to be imposed on model. Conventionally, the shocks are often identified, for instance, by placing economically motivated zero constraints on the impact or the long-run effects of the shocks Kilian+Lutkepohl:2017. Sign constraints, in turn, are commonly used to obtain a set identification with less restrictive or economically more plausible constraints Kilian+Lutkepohl:2017.

In Section (ref), also we make use of zero and sign constraints, but we do it in order to formally label the already locally identified columns of the B-matrix by the shocks of interest. The required conditions are, nevertheless, flexible, and allow for using sign constraints alone or together with zero constraints. Some of the constraints are also testable, as they are overidentifying. Section (ref) additionally takes advantage of zero constraints to identify the shock of interest when the condition for identification through conditional heteroskedasticity fails.

It turns out that any invertible B-matrix that simultaneously diagonalizes the covariance matrices $\Omega_1,...,\Omega_M$, i.e., produces a diagonal conditional covariance matrix ((ref)) of the structural error, has linearly independent eigenvectors of the matrix $\Omega_m\Omega_1^{-1}$ as its columns. If $M > 2$, the matrices $\Omega_m\Omega_1^{-1}$, $m=2,...,M$, thus need to share the common eigenvectors in $B_t$, which restricts the parameter space for the covariance matrices. In that case, the existence of such B-matrix can be tested with a likelihood ratio test, for example. Denoting the eigenvalues of $\Omega_m\Omega_1^{-1}$ as $\lambda_{mi}$, the B-matrix is also unique up to scalar multiples and ordering of its columns if none of the pairs of $\lambda_{mi}$, $i=1,...,d$, is identical for all $m=2,...,M$. These results are formalized in the following assumption and lemma.

assumptionConsider $M$ positive definite $(d\times d)$ covariance matrices $\Omega_m$, $m = 1,...,M$, and denote the strictly positive eigenvalues of the matrices $\Omega_m\Omega_1^{-1}$ as $\lambda_{mi}$, $i=1,...,d$, $m=2,...,M$. Suppose that for all $i\neq j\in\lbrace 1,...,d\rbrace$, there exists an $m\in\lbrace 2,...,M\rbrace$ such that $\lambda_{mi}\neq\lambda_{mj}$.
lemmaConsider $M$ positive definite $(d\times d)$ covariance matrices $\Omega_m$, $m = 1,...,M$, and an invertible $(d\times d)$ matrix $B_t$ such that $B_t^{-1}\Omega_m B_t'^{-1}$ are diagonal matrices with strictly positive diagonal elements. Then, $B_t$ has eigenvectors of $\Omega_m\Omega_1^{-1}$ as its columns. Moreover, $B_t$ is unique up to scalar multiples and ordering of its columns if Assumption (ref) holds.

Under Assumption (ref), the columns of $B_t$ are unique up to scalar multiples and ordering, implying that the shocks are identified up to sign, size, and ordering. Normalizing the conditional covariance matrix of the structural error to a constant diagonal matrix then identifies the B-matrix up to sign and ordering of the shocks. This is formalized in the following proposition.

propositionConsider $M$ positive definite $(d\times d)$ covariance matrices, $\Omega_m$, $m=1,...,M$, and an invertible $(d\times d)$ matrix $B_t$ such that $B_t^{-1}\Omega_m B_t'^{-1}$ are diagonal matrices with strictly positive diagonal elements. Suppose that Assumption (ref) holds. Then, if the conditional covariance matrix of the structural error, $\text{Cov}(e_t|\mathcal{F}_{t-1})=\sum_{m=1}^M\alpha_{m,t}B_t^{-1}\Omega_mB_t'^{-1}$, is normalized to a constant diagonal matrix with strictly positive diagonal entries, the B-matrix $B_t$ is unique up to ordering of its columns and changing all signs in a column.

That is, by fixing an ordering and signs for the columns of the B-matrix, the solution to the diagonalization problem is unique for any given (constant) normalization of the structural error's conditional covariance matrix, say, an identity matrix. In order to find the related B-matrix, it is then convenient to utilize the following matrix decomposition for the reduced form error covariance matrices, which was also employed by Lanne+Lutkepohl:2010 and Lanne+Lutkepohl+Maciejowska:2010 to solve a similar identification problem. We decompose the reduced form error covariance matrices as

equation[equation omitted — 110 chars of source]

where the diagonal of $\Lambda_m=\text{diag}(\lambda_{m1},...,\lambda_{md})$, $\lambda_{mi}>0$ ($i=1,...,d$), contains the eigenvalues of the matrix $\Omega_m\Omega_1^{-1}$ and the columns of the nonsingular $W$ are the related eigenvectors (that are the same for all $m$ by construction). When $M=2$, the decomposition ((ref)) always exists Muirhead:1982, but for $M>2$ its existence requires that the matrices $\Omega_m\Omega_1^{-1}$ share the common eigenvectors in $W$. This is, however, testable and relates to our earlier discussion on the existence of a B-matrix that simultaneously diagonalizes the reduced form error covariance matrices.

Any scalar multiples of linearly independent eigenvectors of $\Omega_m\Omega_1^{-1}$ comprise an appropriate B-matrix, but only specific scalar multiples comprise the locally unique B-matrix associated with a given normalization of structural error's conditional covariance matrix. Direct calculation shows that the B-matrix associated with the normalization $\text{Cov}(e_t|\mathcal{F}_{t-1})=I_d$ is obtained as

equation[equation omitted — 94 chars of source]

where $B_tB_t'=\Omega_{u,t}$. Since $B_t^{-1}\Omega_mB_t'^{-1}= \Lambda_m \left( \sum_{n=1}^M\alpha_{n,t}\Lambda_n \right)^{-1}$ where $\Lambda_1\equiv I_d$, the B-matrix ((ref)) simultaneously diagonalizes $\Omega_1,...,\Omega_M$, and $\Omega_{u,t}$ for each $t$ so that the structural error's conditional covariance matrix is normalized to an identity matrix:

equation[equation omitted — 149 chars of source]

Our specification of the B-matrix differs from Lanne+Lutkepohl+Maciejowska:2010 who assume that the instantaneous effects of the shocks are time-invariant (and specify $B_t=W$), but it extends the one in Lanne+Lutkepohl:2010 to accommodate time-varying mixing weights.

The SGMVAR model assumes a single B-matrix that varies continuously in time according to the conditional covariance matrix of the reduced form error, which in turn varies according to the mixing weights. We established that under Assumption (ref) and a normalization of the structural error's conditional variance, the B-matrix is unique up to ordering of its columns and switching all signs in a column. Hence, as long as one is willing to impose the assumption of a single (time-varying) B-matrix, the columns of the B-matrix unambiguously characterize the estimated impact effects of the shocks, but they do not reveal which column is related to which shock. Since the impact matrix is also subject to estimation error, further constraints may be needed for labelling the shocks.\footnote{As opposed to our single B-matrix, an alternative specification of the structural model would incorporate a separate B-matrix for each of the regimes. Our B-matrix allows the magnitude of the impact effects of a constant sized shock to vary according to the mixing weights, but unlike our model, the alternative specification would allow variation also in the impact effects relative to the other variables. That is, our model imposes structure already in the model Equations ((ref)) and ((ref)).}

Identification of the shocks

We derived a locally unique solution for the B-matrix ((ref)) under Assumption (ref). However, global identification requires fixing the signs and the ordering of its columns. The signs can be fixed by placing a single strict sign constraint in each of the columns of $W$, whereas the ordering of the columns can be fixed by fixing an ordering for the eigenvalues $\lambda_{mi}$ in the diagonals of $\Lambda_m$. This leads to statistical identification of the model with any arbitrary ordering, but it does not reveal which column of the B-matrix is related to which shock.

A structural shock relates to an economic shock through the specific constraints in the corresponding column of $W$ (or equally of the B-matrix) that only the shock of interest satisfies. If such constraints are readily satisfied in the (unrestricted) estimate of $W$, the identification amounts to labelling the structural shocks by the appropriate economic shocks, as long as the constraints are strong enough to pin down a unique ordering for the columns of $W$ (this argument will be formalized in Proposition (ref) below). If the unrestricted estimate of $W$ is such that the shocks of interest cannot be uniquely associated to it, the appropriate constraints can be placed for their identification.\footnote{For a more thorough discussion on economic shocks and their identification, see Ramey:2016 and Uhlig:2017, for example.}

As in practice the interest is often in identifying only some specific shock or shocks, it is of interest to consider only partial identification of the B-matrix as well. Specifically, we say that the $j$th structural shock is uniquely identified if the $j$th column of the B-matrix ((ref)) is unique for given mixing weights $\alpha_{1,t},...,\alpha_{M,t}$. This requires that the $j$th columns of $W$ and $\Lambda_m$, $m=2,..,M$, are unique. The following proposition gives sufficient conditions for global identification of the last $d_1$ shocks when the related pairs of $\lambda_{mi}$ are distinct for some $m$ (which is always the case under Assumption (ref) but does not require Assumption (ref) if $d_1<d$).

propositionSuppose $\Omega_1=WW'$ and $\Omega_m=W\Lambda_mW', \ \ m=2,...,M,$ where $\Lambda_m=\text{diag}(\lambda_{m1},...,\lambda_{md})$, $\lambda_{mi}>0$ ($i=1,...,d$), contains the eigenvalues of $\Omega_m\Omega_1^{-1}$ in the diagonal and the columns of the nonsingular $W$ are the related eigenvectors. Then, the last $d_1$ structural shocks are uniquely identified if \begin{enumerate}[label={(\arabic*)}] • for all $j > d - d_1$ and $i\neq j$ there exists an $m\in \lbrace 2,...,M \rbrace$ such that $\lambda_{mi} \neq \lambda_{mj}$, • the columns of $W$ are constrained in a way that for all $i\neq j > d - d_1$, the $i$th column cannot satisfy the constraints of the $j$th column as is nor after changing all signs in the $i$th column, and • there is at least one (strict) sign constraint in each of the last $d_1$ columns of $W$. \end{enumerate}

Condition (ref) of Proposition (ref) fixes the signs in the last $d_1$ columns of $W$ and therefore the signs of the instantaneous effects of the corresponding shocks. Changing the signs of the columns is effectively the same as changing the signs of the corresponding shocks, so Condition (ref) is not restrictive, however (as the structural shock has a distribution that is symmetric about zero). The assumption that the identified shocks are the last $d_1$ shocks is neither restrictive as one may always reorder the structural shocks accordingly.

For example, if $d=3$, $\lambda_{m1}\neq\lambda_{m3}$ for some $m$, and $\lambda_{m2}\neq\lambda_{m3}$ for some $m$, the third structural shock can be identified with the following constraints:

equation[equation omitted — 290 chars of source]

and so on, where "$*$" signifies that the element is not constrained, "$+$" denotes a strict positive and "$-$" a strict negative sign constraint, and "$0$" means that the element is constrained to zero. In the first example, Condition (ref) is satisfied because the last shock is assumed to move the last two variables to the opposite directions when the first two shocks are assumed to move them to the same direction, implying that the first two shocks cannot satisfy the constraint imposed on the last shock (as is nor after changing all signs of the impact responses). Similarly in the second example, the last shock moves to opposite directions the variables that the first two shocks move to the same direction. The last example imposes a zero constraint for the impact response of the first variable to the last shock, while the first two shocks impose strict sign constraints. Since the non-zero impact responses of the first two shocks cannot satisfy the zero constraint of the last shock, Condition (ref) is satisfied. By using sign and zero constraints in this manner, it is easy to produce further examples that lead to the identification of the last shock.

Imposing sign or zero constraints on $W$ equals to placing them on $B_t$, so they can be justified economically. Under Assumption (ref), the model is statistically identified prior to imposing the constraints, making the parameter constraints required in Condition (ref) also testable. This different to the conventional SVAR setup in which the identifying constraints cannot be validated statistically Kilian+Lutkepohl:2017. Similarly to the conventional SVAR model, labelling the shocks formally with the economic shocks of interest, however, requires the identification constraints to be economically motivated. As Proposition (ref) shows and the examples in ((ref)) demonstrate, our method facilitates finding economically plausible identification constraints by flexibly using sign constraints alone or in combination with zero constraints. A point identification can be obtained even with only sign constraints, while in the conventional SVAR setup, sign constraints alone lead to a set identification only Kilian+Lutkepohl:2017. If Assumption (ref) fails, the structural GMVAR model is not fully identified and the problem of testing the parameter constraints is non-standard, which is briefly addressed in the next section.

Identification of the shocks under partial identification of the model

If Assumption (ref) is violated and the structural GMVAR model is thus not statistically identified, the shocks of interest can still be identified with Proposition (ref) if Condition (ref) is satisfied. When the shocks of interest do not satisfy Condition (ref), their identification requires stronger constraints than in Proposition (ref). Therefore, we present the following proposition that provides sufficient criteria for global identification of the last $d_1$ shocks when Condition (ref) fails; specifically, when exactly one of the eigenvalues $\lambda_{mi}$ with $i\neq j > d - d_1$ is identical to $\lambda_{mj}$ for all $m$. For simplicity, we assume that only one of the shocks with identical eigenvalues is to be identified, i.e., $i\leq d - d_1$ above.

propositionLet $d_1<d$. Consider the matrix decomposition of Proposition (ref) and further suppose that for $j=d-d_1+1$ and some $i\leq d-d_1$, we have $\lambda_{mi}=\lambda_{mj}$ for all $m$, but for all $l\centernot\in\lbrace i,j \rbrace$, $\lambda_{ml}\neq\lambda_{mj}$ for some $m$. Then, the last $d_1$ structural shocks are uniquely identified if Conditions (ref)-(ref) of Proposition (ref) are otherwise satisfied, and in addition \begin{enumerate}[label={(\arabic*)}] \setcounter{enumi}{3} • the column $i\leq d-d_1$ of $W$ such that $\lambda_{mi}=\lambda_{mj}$ for all $m$ has at least one (strict) sign constraint and the $j$th column has a zero constraint where the $i$th column has the (strict) sign constraint. \end{enumerate}

Note that the assumption $j=d-d_1+1$ is made without loss of generality, as the structural shocks can always be reordered accordingly by also reordering the columns of $W$ (including the constraints) and the eigenvalues $\lambda_{mi}$ correspondingly.

To exemplify, if $d=4$, $\lambda_{m1}\neq\lambda_{m4}$ for some $m$, $\lambda_{m2}\neq\lambda_{m4}$ for some $m$, and $\lambda_{m3}=\lambda_{m4}$ for all $m$, the following constraints lead to global identification of last shock:

equation[equation omitted — 359 chars of source]

and so on. Condition (ref) is satisfied in each of the above examples, because the third shock has a strict sign constraint for the variable that the last shock imposes a strict zero constraint. As is demonstrated above, the structural shocks can often be identified with flexible constraints even when some of the eigenvalues are identical for all regimes.

Under Conditions (ref) and (ref) of Proposition (ref), the additional constraints on $W$ were stated testable because they are overidentifying and statistical identification of the model can always be achieved by fixing the ordering of the eigenvalues $\lambda_{mi}$, as long as none of the pairs of $\lambda_{mi}$, $i=1,...,d$, is identical for all $m=2,...,M$ (Assumption (ref)). In the setup of Proposition (ref), however, when $\lambda_{mi}=\lambda_{mj}$ for all $m$ and some $i\neq j$, the model is not generally identified even when one fixes a unique ordering for the eigenvalues and the columns of $W$. Also, even if Condition (ref) of Proposition (ref) is satisfied, only partial identification of the B-matrix is obtained since nothing guarantees unique identification of the $i$th column of $W$, which would require stronger conditions. Consequently, the model is not identified under the null nor the alternative hypothesis when testing for the constraints in Conditions (ref) and (ref), making the testing problem non-standard and the conventional asymptotic distributions of the likelihood ratio and Wald test statistics unreliable. The same applies when one tests the equality of the eigenvalues in order to assess the validity of Condition (ref) of Proposition (ref), as the model is not identified under the null. Deriving formal tests under no identification is, however, a major task and beyond the scope of this paper.\footnote{Lutkepohl+Meitz+Netsunajev+Saikkonen:2021 discussed a related testing problem under no identification and developed an asymptotic Wald type test for testing equality of the $\lambda_{mi}$ parameters in a linear SVAR model incorporating two volatility regimes with a known change point and (reduced form) shocks arriving from a class of elliptical distributions. Meitz+Saikkonen:2021, on the other hand, studied the asymptotic properties of a likelihood ratio test statistic under no identification when testing for the number of regimes in mixture models with Gaussian conditional densities. One of the studied models is the GMAR model Kalliovirta+Meitz+Saikkonen:2015, which is the univariate counterpart of the GMVAR model Kalliovirta+Meitz+Saikkonen:2016.}

If more than two eigenvalues are identical for all $m=2,...,M$ but they are not all identical, it may still be possible to find flexible conditions for identification of the shocks. Specifically, the idea utilized in the proof of Proposition (ref) (presented in Appendix (ref)) can be applied to larger numbers of identical eigenvalues. If all the eigenvalues are identical for all covariance matrices, then $\Omega_m=\lambda_{m1}\Omega_1$ and the identification condition is the same as for the conventional SVAR model Lutkepohl:2005. As a general remark, observe that constraining an element of $B_t$ to be any constant other than zero is infeasible, because all elements on the right side of ((ref)) are either zero or time-varying due to the time-varying mixing weights.\footnote{We have focused on the B-model, where the structure is imposed on the contemporaneous relations of the shocks. Alternatively, one may consider the "A-model" setup in which the structure is placed on the contemporaneous relations of the observable variables governed by the "A-matrix" Lutkepohl:2005. The A-model is obtained implicitly from the B-model ((ref)) and ((ref)) by defining the A-matrix as $A_t\equiv B_t^{-1}$, where $B_t$ is given by ((ref)). In this case, the structural model Equation ((ref)) becomes $$ A_ty_t = \sum_{m=1}^M(A_t\phi_{m,0}+\sum_{i=1}^pA_tA_{m,i}y_{t-i})+e_t, $$ thereby incorporating continuously varying intercepts and coefficient matrices due to the constant variance normalization of the structural shocks. In practice, however, one needs to carefully derive how any specific constraint on $A_t$ can be imposed by restricting $W$.}

Maximum likelihood estimation

The parameters of the reduced form GMVAR model are collected to the vector $\boldsymbol{\theta}=(\boldsymbol{\vartheta}_1,...,\boldsymbol{\vartheta}_M,\alpha_1,...,\alpha_{M-1})$ $((M(d^2p + d + d(d+1)/2) + 1) \times 1)$, where $\boldsymbol{\vartheta}_m=(\phi_{m,0},\text{vec}(A_{m,1}),....,\text{vec}(A_{m,p}),\text{vech}(\Omega_m))$, vec is a vectorization operator that stacks the columns of a matrix on top of each other, and vech stacks the columns of a matrix from the main diagonal downwards (including the main diagonal). The last mixing weight parameter $\alpha_M$ is omitted because it is obtained from the constraint $\sum_{m=1}^M \alpha_m=1$.

Using the notation described in Section (ref), indexing the observed data as $y_{-p+1},...,y_0,y_1,...,y_T$, and assuming that the initial values $\boldsymbol{y}_0=(y_{-p+1},...,y_0)$ are stationary, the exact log-likelihood function of the reduced form GMVAR model takes the form Kalliovirta+Meitz+Saikkonen:2016

equation[equation omitted — 220 chars of source]

where

equation[equation omitted — 231 chars of source]

If it does not seem reasonable to assume that the initial values are stationary, one may condition on them and base the estimation on the conditional log-likelihood function, which is obtained by dropping the first term on the right side of ((ref)).

The reduced form GMVAR model can be estimated by maximizing the exact or conditional likelihood function in ((ref)) and ((ref)) with respect to the parameter $\boldsymbol{\theta}$. To ensure identification, the parameter space should be constrained so that the mixture components cannot be 'relabelled', for instance, by assuming that the mixing weight parameters are in a decreasing order, $\alpha_M>\cdots\alpha_1>0$, and $\boldsymbol{\vartheta}_i=\boldsymbol{\vartheta}_j$ only if $i=j$ Kalliovirta+Meitz+Saikkonen:2016. If $M=2$, the structural GMVAR model is then obtained by simultaenously diagonalizing the reduced form error covariance matrices as discussed Section (ref). However, should overidentifying restrictions be imposed on $B_t$ through $W$ or if $M\geq 3$, it is more convenient to reparametrize the model with $W$ and $\Lambda_m$, $m=2,...,M$, instead of $\Omega_1,...,\Omega_M$ and maximize the log-likelihood function subject to the new set of parameters and constraints. In this case, the decomposition ((ref)) is plugged in to the log-likelihood function and the $\text{vech}(\Omega_1),...,\text{vech}(\Omega_M)$ are replaced with $\text{vec}(W),\lambda_2,...,\lambda_M$, where $\lambda_m=(\lambda_{m1},...,\lambda_{md})$, in the parameter vector $\boldsymbol{\theta}$.

Maximizing the complex and highly multimodal log-likelihood function can be challenging in practice, particularly if there are more than two regimes. Following Dorsey+Mayer:1995, Meitz+Preve+Saikkonen:2018,Meitz+Preve+Saikkonen:2021, and Virolainen:2021,uGMAR, we employ a two-phase estimation procedure where, in the first phase, a genetic algorithm is used to find starting values for a gradient based method which then, in the second phase, often converges to a nearby local maximum or saddle point. The genetic algorithm in the accompanying R package gmvarkit gmvarkitnormal has been modified to improve its performance significantly, and it functions similarly to the one described in Virolainen:2021 for the univariate GMAR Kalliovirta+Meitz+Saikkonen:2015, StMAR Meitz+Preve+Saikkonen:2021, and G-StMAR Virolainen:2021 models. In order to obtain reliable results, a (sometimes very large) number of estimation rounds should be performed, for which gmvarkit makes use of parallel computing.

Impulse response analysis

The expected effects of the structural shocks in the SGMVAR model generally depend on the initial values as well as on the sign and size of the shock, which makes the conventional way of calculating impulse responses unsuitable Kilian+Lutkepohl:2017. Following Koop+Pesaran+Potter:1996 and Kilian+Lutkepohl:2017, we therefore consider the generalized impulse response function (GIRF) defined as

equation[equation omitted — 158 chars of source]

where $h$ is the chosen horizon and $\mathcal{F}_{t-1}=\sigma\lbrace y_{t-j},j>0\rbrace$ as before. The first term on the right side of ((ref)) is the expected realization of the process at time $t+h$ conditionally on a structural shock of size $\delta_j \in\mathbb{R}$ in the $j$th element at time $t$ and the previous observations. The latter term on the right side is the expected realization of the process conditionally on the previous observations only. The GIRF thus expresses the expected difference in the future outcomes when the structural shock of size $\delta_j$ in the $j$th element hits the system at time $t$ as opposed to all shocks being random.

It is easy to see that the SGMVAR model has a $p$-step Markov property, so conditioning on (the $\sigma$-algebra generated by) the $p$ previous observations $\boldsymbol{y}_{t-1}=(y_{t-1},...,y_{t-p})$ is effectively the same as conditioning on $\mathcal{F}_{t-1}$ at time $t$ and later. The history $\boldsymbol{y}_{t-1}$ can be either fixed or random, but with random history the GIRF becomes a random vector, however. Using fixed $\boldsymbol{y}_{t-1}$ makes sense when one is interested in the effects of the shock at a particular point of time, whereas more general results are obtained by assuming that $\boldsymbol{y}_{t-1}$ follows the stationary distribution of the process. If one is, on the other hand, interested in a specific regime, $\boldsymbol{y}_{t-1}$ can be assumed to follow the stationary distribution of the corresponding component model.

The GIRF and its distributional properties can be estimated with a Monte Carlo algorithm that generates (partial) realizations of the process and then takes the sample mean for point estimate. If $\boldsymbol{y}_{t-1}$ is random and follows the distribution $G$, the GIRF should be estimated for different values of $\boldsymbol{y}_{t-1}$ generated from $G$, and then the sample mean and sample quantiles can be taken to obtain the point estimate and confidence intervals that reflect the uncertainty about the initial value. Such an algorithm, adapted from Koop+Pesaran+Potter:1996 and Kilian+Lutkepohl:2017, is given in Appendix (ref).

Because the SGMVAR model facilitates associating statistical characteristics and economic interpretations to the regimes, and because asymmetries in the GIRFs are caused by regime-switches, it may be of interest to also examine the effects of a structural shock to the mixing weights $\alpha_{m,t}$, $m=1,...,M$. We then consider the related GIRFs

equation[equation omitted — 167 chars of source]

for which point estimates and confidence intervals can be constructed similarly to ((ref)).

Empirical application

Our empirical application studies asymmetries in the effects of U.S. monetary policy shocks. Asymmetric effects of U.S. monetary policy shocks have been studied, among others, by Weise:1999, Garcia+Schaller:2002, Lo+Piger:2005, and Hoppner+Melzer+Neumann:2008, who all found the effects of monetary policy shocks to production stronger during recessions (or low growth periods) than booms (or high growth periods). Weise:1999 also found evidence in favor of large and small shocks having different effects, and large positive and negative monetary shocks having different effects. Hoppner+Melzer+Neumann:2008 concluded that the real effects of monetary policy shocks have decreased over their sample period from 1962 to 2002. Tenreyro+Thwaites:2016, on the other hand, found the effects of U.S. monetary policy shocks less powerful in recessions.

We consider the quarterly U.S. data covering the period from 1954Q3 to 2021Q4 ($270$ observations) and consisting of four variables: real GDP, GDP implicit price deflator, producer price index (all commodities), and an interest rate variable. Our policy variable is the interest rate variable, which is the effective federal funds (FF) rate from 1954Q3 to 2008Q2. After that we replaced it with the Wu+Xia:2016 shadow rate, which is not constrained by the zero lower bound and also quantifies unconventional monetary policy measures. \footnote{The Wu+Xia:2016 shadow rate series was retrieved from the Federal Reserve Bank of Atlanta's website and the rest of the data were retrieved from the Federal Reserve Bank of St. Louis database.}

The GMVAR model requires stationary data, so the logarithms of the real GDP, GDP deflator, and producer price index need to be detrended. We detrend the logarithm of the real GDP by separating its cyclical component from the trend with the backward-looking Hodrick-Prescott (HP) filter and then considering the cyclical component.\footnote{The backward-looking HP filter was obtained from the two-sided HP filter by applying the filter up to horizon $t$, taking the last observation, and repeating this procedure for the full sample $t=1,...,T$. In order to allow the series to start from any phase of the cycle, we applied the backward-looking filter to the full available sample from 1947Q1 to 2021Q4 before extracting our sample period from it. We computed the two-sided HP filter with the R package lpirfs lpirfs by using the standard smoothing parameter value of $1600$. } It is thereby implicitly assumed that the monetary policy shock does not have permanent effects on real output. The logarithms of the price variables are detrended by taking the first difference and multiplying it by hundred, so the resulting series approximate the percentage growth rates. The interest rate variable is treated as stationary.

figure[figure omitted — 803 chars of source]

The series are presented in the first four top panels of Figure (ref) with the shaded areas indicating the periods of NBER based U.S. recessions. Throughout, we refer to the variables as GDP (output), GDPDEF (prices), PPI (commodity prices), and RATE (interest rate) without making it explicit that some of them are detrended. The (S)GMVAR model of autoregressive order $p$ and $M$ mixture components is referred to as (S)GMVAR($p,M$) model.

We select the order of our GMVAR model by first finding a suitable autoregressive order for a linear Gaussian VAR; that is, a GMVAR($p,1$) model. The AIC is minimized by the order $p=3$, suggesting that this might be the appropriate lag order for modelling autocorrelation. So we estimate a GMVAR($3,2$) model, which we find superior to the linear VAR. Graphical quantile residual diagnostics reveal that our GMVAR($3,2$) model adequately captures the autocorrelation structure of the series, but some of the conditional heteroskedasticity and excess kurtosis is not captured. In our view, the overall adequacy of the model is, nevertheless, reasonable enough for further analysis. Details on the model selection and quantile residual diagnostics are given in Appendix (ref).

The estimated mixing weights of the two regimes are presented in the bottom panel of Figure (ref). The second regime (red) mainly dominates during periods of high inflation and interest rate in the 1970's and 1980's, after the collapse of Lehman Brothers in the Financial crisis until the end of 2009, and finally during the COVID-19 crisis from the third quarter of 2020 onwards. We refer to this regime as the unstable inflation regime, as it generally exhibits high or volatile inflation. The first regime (blue) prevails when the second one does not: before 1970's, short periods during 1970's, and from the mid 1980's onwards but excluding the Financial crisis and the COVID-19 crisis (but including the first two quarters of 2020). We refer to this regime as the stable inflation regime, as it is characterized by moderate inflation. Details about the characteristics of the regimes are provided in Appendix (ref).

Identification of the monetary policy shock

Decomposing the covariance matrices of the reduced form GMVAR($3,2$) model as in ((ref)) gives the following estimates for the structural parameters:

equation[equation omitted — 785 chars of source]

where the ordering of the variables is $y_t=(\text{GDP}_t, \text{GDPDEF}_t, \text{PPI}_t, \text{RATE}_t)$, the estimates $\hat{\lambda}_{2i}$ are in an increasing order (which fixes an arbitrary ordering for the columns of $\hat{W}$), and approximate standard errors are given in parentheses next to the estimates. The estimates that deviate from zero by more than two times their approximate standard error are bolded. We proceed by assuming that all the $\lambda_{2i}$, $i=1,...,4$, are different to each other, i.e., that Assumption (ref) holds, which leads to statistical identification of the model. After identifying the monetary policy shock, the robustness of our identification with respect to this unjustified assumption is discussed.

Based on the estimates and their standard errors in ((ref)), the first shock moves GDP and inflation to the opposite directions, whereas the second shock moves GDP and commodity price inflation to the opposite directions. Since the instantaneous movements of the interest rate variable are insignificant, these two shocks do not seem plausible candidates for the monetary policy shock. The third shock moves the interest rate variable on impact more significantly than the first two shocks, but since production and both prices move significantly to the same direction, its characteristics appear similar to an aggregate demand shock and not a monetary policy shock. The last shock moves the interest rate variable significantly, while it also moves GDP, inflation, and commodity price inflation to the opposite direction, which is consistent with many of the standard the economic theories Gali:2015. The impact effects of GDP, inflation, and commodity price inflation are, however, statistically insignificant and the response of inflation is very weak. Nevertheless, among the four structural shocks obtained for the model, the characteristics of the last shock mostly resemble those of a monetary policy shock, so we deem it as the monetary policy shock.

Identifying the monetary policy shock formally by Proposition (ref) requires such constraints to be imposed on $W$ that it can be unambiguously distinguished from the other shocks. We assume that the monetary policy shock moves the GDP and commodity price inflation to the opposite direction from the interest rate variable. In addition, we impose a zero constraint on the instantaneous movement of inflation, as the unrestricted estimated is very close to zero compared to the approximate standard error, and it allows us to avoid making restrictive assumptions about the first and third shocks. The Wald test produces the $p$-value $0.92$ for the zero constraint, so it is not rejected.

To distinguish the monetary policy shock from the other shocks, we assume that the first and third shocks move inflation at impact. This is not economically restrictive (since the responses can be very small) but it is a statistically reasonable assumption, as the Wald test rejects the hypotheses that the impact responses are zero (jointly or individually) with $p$-values less than $10^{-7}$. The Wald test produces the $p$-value $0.056$ for the hypothesis that the second shock does not move inflation at impact, so we cannot reject it. Therefore, we assume that the second shock moves the GDP and commodity price inflation to the opposite directions, as the corresponding estimates are large compared to their approximate standard errors, making the constraints statistically sensible.

The estimates obtained for the structural parameters under the above-described identification changed only slightly from the unrestricted ones in ((ref)), and they are presented in Appendix (ref). The estimates for $\lambda_{23}$ and $\lambda_{24}$ are somewhat close to each other relative to their standard errors, but due to our zero constraint on the inflation, Proposition (ref) identifies the monetary policy shock even if $\lambda_{23}=\lambda_{24}$ (or $\lambda_{21}=\lambda_{24}$). The monetary policy shock is identified also if additionally $\lambda_{2i}=\lambda_{2j}$ for any $i,j=1,2,3$, so our identification is not particularly sensitive to the validity of the unjustified Assumption (ref) (while the approximate standard errors and the Wald test results are invalid if the assumption fails).

Generalized impulse response functions

Due to the endogenously determined regime-switching probabilities and the fact that we allow the regime to switch as a result of a shock, there are multiple types of possible asymmetries. The impulse responses can vary depending on the initial value as well as on the sign and size of the shock. We study the state-dependence of the (generalized) impulse response functions by drawing initial values from the stationary distribution of each regime separately. Then, we calculate the $90\%$ confidence intervals that reflect uncertainty about the initial value within the given regime as is described in Section (ref) and Appendix (ref). Asymmetries related to the sign and size of the shock are studied by estimating the GIRFs for positive (contractionary) and negative (expansionary) one-standard-error (small) and two-standard-error (large) shocks. After estimating the GIRFs, they are scaled so that the peak effect of the interest variable is $25$ basis points within the first four quarters, making the responses to shocks of different sign and size comparable.\footnote{The GIRFs are scaled based on the peak response instead of the initial response, because the peak response is much higher compared to the initial response in the first regime than in the second regime. Scaling the GIRFs based on the initial response would then shift the response of the interest rate variable significantly higher in the first regime than in the second regime.}

Figure (ref) presents the GIRFs $h=0,1,...,32$ quarters ahead estimated for the identified monetary policy shock.\footnote{We use $R_1=R_2=2500$ in the Monte Carlo algorithm, i.e., for each regime, size, and sign of the shock we draw $2500$ initial values, and for each of those initial values the GIRF is estimated based on $2500$ different sample paths.} The GIRFs of inflation rate and commodity price inflation rate are not accumulated to levels. From top to bottom, the responses of GDP, inflation rate, commodity price inflation rate, interest rate, and the first regime's mixing weights are depicted in each row, respectively. The first [third] column shows the responses to small contractionary (blue solid line) and expansionary (red dashed line) shocks with the initial values generated from the stationary distribution of the first [second] regime. The second [fourth] column shows the responses to large contractionary and expansionary shocks with the initial values generated from the first [second] regime. The shaded areas are the $90\%$ confidence intervals that reflect uncertainty about the initial value within the given regime. The responses of the second regime's mixing weights are not depicted because they are the negative of those of the first regime.

figure[figure omitted — 1,450 chars of source]

In the first regime (the stable inflation regime; the first and second columns of Figure (ref)), contractionary (expansionary) monetary policy shock causes a significant contraction (expansion) in the GDP, with the peak effect occurring after three to four quarters. On average, the response is hump shaped and decays to zero roughly after three years from impact.\footnote{By zero, we mean the expected observation if all the shocks were random. Accordingly, by positive we mean expected observations larger than that and by negative expected observations smaller than that.} As the confidence bounds show, the response switches sign from some of the starting values, while some of the starting values display persistent contraction. Particularly large shocks seem to cause a delayed expansion (contraction) from many of the starting values - shortly after the impact when the shock is contractionary and later when the shock is expansionary.

The prices mostly rise in response to a contractionary monetary policy shock, although one would often expect a contractionary monetary policy shock to decrease inflation due to the decreased aggregate demand. This is often referred to as the price puzzle, and it has been discussed recently, for instance, in Ramey:2016.\footnote{A popular explanation is that the Fed uses more information in predicting the future inflation than the autoregressive system of the variables included in the model Sims:1992. Consequently, the identified monetary policy shock also contains a component that incorporates the Fed's endogenous response to the prediction of the future inflation that is not captured by the autoregressive system of the included variables. If the endogenous response is not strong enough to offset the predicted inflation, the impulse responses may then display a rise in the inflation. Another explanation proposes that the prices increase due to the cost-push effect of the monetary policy shock. An increase in the nominal interest rate increases the marginal cost of production of the firms who operate on borrowed money, and thereby decreases the aggregate supply and increases the price level Barth+Ramey:2001, Ravenna+Walsh:2006. Several authors have, however, argued that the cost-channel is not likely strong enough to cause a price puzzle even in the short-run Rabanal:2007, Kaufmann+Scharler:2009, Castelnuovo:2012. Nonetheless, we find our empirical results interesting, as the long-run price puzzle arises only from some of the starting values, while its occurrence is also sensitive to the sign and size of the shock. Moreover, as is discussed in Appendix (ref), enforcing linearity to the autoregressive dynamics makes the price puzzle worse.} As is explained in Appendix (ref), the prices rise because the shock drives the economy towards the unstable inflation regime, which has higher long-run inflation.

On average, inflation does not move much in response to a small expansionary monetary policy shock, while the interest rate stays low relatively persistently and is accompanied with a roughly three years long expansion of the GDP. A large expansionary shock drives the economy towards the unstable inflation regime relatively more than a small expansionary shock, as the responses of the first regime's mixing weights show. Consequently, the inflation mostly increases and the interest rate variable increases relatively fast towards zero, and as the confidence bounds show, from many of the starting values the interest rate variable overshoots significantly. It is shown in Appendix (ref) that these GIRFs are the ones that also display particularly high peak inflation, and that the significant monetary policy tightening is accompanied with a persistent contraction of the GDP after the initial expansion.

In the second regime (the unstable inflation regime; the third and fourth columns of Figure (ref)), contractionary (expansionary) monetary policy shock causes a strong contraction (expansion) of the GDP, with the peak effect occurring after two quarters. After roughly three years, from most of the starting values the response overshoots and becomes expansionary (contractionary) before decaying to zero. Inflation rate and commodity price inflation rate rise after the impact and then decrease significantly for several years before the price levels stabilize.

The interest rate decays towards zero for roughly two years after which, on average, it decreases (increases) slightly below (above) zero before returning to zero. The response of the interest rate switches its sign less significantly when the shock is contractionary, when also the inflationary effects of the shock are slightly weaker. The scaled GIRFs are almost identical for small and large shocks, so there does not appear to be much asymmetries with respect to the size of the shock. The asymmetries are weak, because the monetary policy shock has relatively weak effect on the regime-switching probabilities, as the scaled responses of the first regime's mixing weights show (the third and fourth bottom panels of Figure (ref)).

Overall, expansionary and contractionary monetary policy shocks seem to both mostly increase the probability of the unstable inflation regime, significantly more so if the economy is in the stable inflation regime when the shock arrives (the bottom panels in Figure (ref)). In the stable inflation regime, a large shock increases the probability of entering the unstable inflation regime relatively more than a small shock, and often propagates high and persistent inflation, which is followed by a significant monetary policy tightening and a persistent contraction of the GDP. On average, the real effects of the monetary policy shock are somewhat stronger in the stable inflation regime than in the unstable inflation regime, but the (average) effects die out equally fast.

Summary

We introduced a structural version of the Gaussian mixture vector autoregressive model Kalliovirta+Meitz+Saikkonen:2016 that incorporates endogenously determined mixing weights and a time-varying B-matrix. We showed that our model generally identifies the structural shocks up to ordering and sign, but does not reveal which column of the B-matrix is related to which shock. Since the B-matrix is also subject to an estimation error, we made use of the matrix decomposition proposed by Lanne+Lutkepohl:2010 and Lanne+Lutkepohl+Maciejowska:2010 and derived general conditions for formally identifying any subset of the shocks. This led to flexible identification conditions, and some of the constraints are also testable. For impulse response analysis, we utilized the generalized impulse response function Koop+Pesaran+Potter:1996 and proposed a Monte Carlo algorithm for its estimation by making use of the known stationary distribution of the SGMVAR process. The paper is accompanied with the CRAN distributed R package gmvarkit gmvarkitnormal, which provides a comprehensive set of tools for numerical analysis of the model.

Our empirical application studied asymmetries in the expected effects of monetary policy shocks in the U.S. using a quarterly series covering the period from 1954Q3 to 2021Q4. Our SGMVAR model identified two regimes: a stable inflation regime and an unstable inflation regime. The unstable inflation regime is characterized by high or volatile inflation, and it mainly prevails in the 1970's, early 1980's, during the Financial crisis, and in the COVID-19 crisis from 2020Q3 onwards. The stable inflation regime, in turn, is characterized by moderate inflation, and it prevails when the stable inflation regime does not. We found the effects of the monetary policy shock relatively symmetric in the unstable inflation regime, as it rarely causes a switch to the stable inflation regime. A contractionary (expansionary) monetary policy shock appears to first increase (decrease) inflation after which the inflation significantly decreases (increases) for several years. The strong contraction (expansion) in the cyclical component of GDP lasts for roughly three years and is followed by a relatively mild expansion (contraction) along with the interest rate variable overshooting to the negative (positive) side.

The effects of the monetary policy shock were found strongly asymmetric in the stable inflation regime with respect to the initial state of the economy as well as to the sign and size of the shock. A large shock often causes relatively stronger inflationary effects than a small shock, while both contractionary and expansionary shocks seem to increase inflation by driving the economy towards the unstable inflation regime. A small expansionary shock does not move prices much on average, but a large expansionary shock often drives the economy towards the unstable inflation regime and propagates high and persistent inflation. The high inflation is followed by a significant monetary policy tightening and a persistent contraction of the GDP after the initial expansion. On average, the real effects of the monetary policy shock were found somewhat stronger in the stable inflation regime than in the unstable inflation regime.