EconBase
← Back to paper

Gaussian and Student's $t$ mixture vector autoregressive model with application to the effects of the Euro area monetary policy shock

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

73,961 characters · 11 sections · 95 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
titlepage\begin{center} \linespread{1.2} {Gaussian and Student's t mixture vector autoregressive model with application to the effects of the Euro area monetary policy shock}\\ {Savi Virolainen}\\ {University of Helsinki}\\ \begin{abstract} A new mixture vector autoregressive model based on Gaussian and Student's $t$ distributions is introduced. As its mixture components, our model incorporates conditionally homoskedastic linear Gaussian vector autoregressions and conditionally heteroskedastic linear Student's $t$ vector autoregressions. For a $p$th order model, the mixing weights depend on the full distribution of the preceding $p$ observations, which leads to attractive practical and theoretical properties such as ergodicity and full knowledge of the stationary distribution of $p+1$ consecutive observations. A structural version of the model with statistically identified shocks is also proposed. The empirical application studies the effects of the Euro area monetary policy shock. We fit a two-regime model to the data and find the effects, particularly on inflation, stronger in the regime that mainly prevails before the Financial crisis than in the regime that mainly dominates after it. The introduced methods are implemented in the accompanying R package gmvarkit. \\[0.5cm] \noindentKeywords: regime-switching, Student's t mixture, mixture vector autoregressive model, mixture VAR, Euro area monetary policy shock\%[2.0cm] \end{abstract} \begingroup \footnote{The author thanks Markku Lanne, Mika Meitz, and Pentti Saikkonen for discussions and comments, which helped to improve this paper substantially. The author also thanks Leena Kalliovirta for the useful comments and the Research Council of Finland for the financial support (Grants 308628 and 347986).} \addtocounter{footnote}{-1} \endgroup \begingroup \footnote{Contact address: Savi Virolainen, Faculty of Social Sciences, University of Helsinki, P. O. Box 17, FI–00014 University of Helsinki, Finland; e-mail: [email removed]. ORCiD ID: 0000-0002-5075-6821.} \addtocounter{footnote}{-1} \endgroup \end{center}

Introduction

Nonlinear vector autoregressive (VAR) models are useful for modelling series in which the data generating dynamics vary in time. Such variation may arise due to wars, crises, business cycle fluctuations, or policy shifts, for example. Markov-switching vector autoregressive (MS-VAR) models Krolzig:1997 have been particularly popular (e.g., Garcia+Schaller:2002; Peersman+Smets:2002; Lo+Piger:2005; and Dolado+MariaDolores:2006) due to the capability to flexibly model series exhibiting discrete regime-switches. Also logistic smooth transition vector autoregressive (LSTVAR) models Anderson+Vahid:1998 have been widely employed (e.g., Weise:1999; Auerbach+Gorodnichenko:2012; and Caggiano+Castelnuovo+Colombo+Nodari:2015), as they allow to capture gradual shifts in the dynamics of the data with clear interpretations of the regimes.

Although useful, MS-VAR and LSTVAR models have limited switching dynamics: in the former, they depend only on the preceding regime, while in the latter, they depend solely on the levels of the switching variables. As a result, these models cannot capture more complex switching dynamics that may depend on a variety of statistical properties of the data. The Gaussian mixture vector autoregressive (GMVAR) model of Kalliovirta+Meitz+Saikkonen:2016, in contrast, incorporates switching dynamics that, for a $p$th order model, depend on the full distribution of the previous $p$ observations. Specifically, the greater the relative weighted likelihood of a regime is, the more likely the process is to generate an observation from it. This more data-driven approach facilitates capturing complex switching dynamics that do not depend on the choice of switching variables, and it leads to attractive theoretical properties such as ergodicity and full knowledge of the stationary distribution of $p+1$ consecutive observations. On the other hand, it is not always clear how to interpret the regimes of the GMVAR model. Moreover, since the regimes of the GMVAR model are conditionally homoskedastic linear Gaussian VARs, its capability to capture strong conditional heteroskedasticity and high kurtosis are rather limited.

In order to address the limitations of the extant models, this paper introduces a new mixture vector autoregressive model that is closely related to the GMVAR model but allows the regimes to be conditionally heteroskedastic Student's $t$ VARs. This model, which we refer to as the Gaussian and Student's $t$ mixture vector autoregressive (G-StMVAR) model, thereby accommodates conditionally homoskedastic linear Gaussian VARs and conditionally heteroskedastic linear Student's $t$ VARs as its regimes (or mixture components). The G-StMVAR model can be described as the multivariate counterpart of the Gaussian and Student's $t$ mixture autoregressive (G-StMAR) model of Virolainen:2022, and if all of its regimes are Student's $t$ VARs, the multivariate counterpart of the Student's $t$ mixture autoregressive (StMAR) model of Meitz+Preve+Saikkonen:2023 is obtained as a special case. At each time point, the G-StMVAR model generates an observation from one of its mixture components that is randomly selected according to the probabilities given by the mixing weights.

Both types of mixture components in the G-StMVAR model have the same form for the conditional mean, a linear function of the preceding $p$ observations, but their conditional covariance matrices are different. The linear Gaussian VARs have constant conditional covariance matrices, while those of the linear Student's $t$ VARs are products of a constant covariance matrix and a time-varying scalar that depends on the quadratic form of the previous $p$ observations. In contrast to the GMVAR model, our model is, hence, able to capture excess kurtosis and conditional heteroskedasticity also within the regimes.

Our specification of the conditional covariance matrix is not as general as the conventional multivariate autoregressive conditional heteroskedasticity (ARCH) process that allows the entries of the conditional covariance matrix to vary relative to each other Lutkepohl:2005. It is, nonetheless, convenient for establishing stationarity properties similar to the linear Gaussian VARs. Our specification is also parsimonious, as it only depends on the covariance matrix, degrees of freedom, and autoregressive parameters. This parsimony is particularly advantegeous in mixture VARs, where the large number of parameters can often be a problem even without an ARCH component.

In addition to the reduced form model, we propose a structural version of the G-StMVAR model with statistically identified shocks. Specifically, the shocks are identified (up to ordering and sign) by simultaneously orthogonalizing them in all regimes similarly to Virolainen:2024, who applied the identification to a linear SVAR model with regime-switching volatility. Our identification method restricts the relative magnitudes of the impact responses of the variables to each shock time-invariant, but in contrast to Virolainen:2024, the impulse responses are allowed to vary across the regimes after the impact period. Consequently, our model accommodates state-dependent impulse responses, which may additionally vary depending on the sign and size of the shock. Our model also has the interesting property that it allows to assess the likelihood of future regime switches due to each shock.

Compared to conventional recursive identification, our method is particularly advantageous when the shock of interest would be ordered last or nearly last, which is typical in monetary policy shock applications involving nonlinear SVARs, for example (e.g., Weise:1999; Peersman+Smets:2002; Garcia+Schaller:2002; Lo+Piger:2005; Hoppner+Melzer+Neumann:2008; Pellegrino:2018; and Burgard+Neuenkirch+Nockel:2019). This is because when the shock of interest (and the corresponding variable) is ordered last, recursive identification restricts the impact responses of the other variables to zero and thus time-invariant, while our methods allows them to deviate from zero. In line with the statistical identification literature, labelling the identified shocks requires external information, for example, in the form of economic short-run restrictions. Due to the statistical identification, the additional economic restrictions are also testable.

Mixture VARs have been previously applied in structural analysis by Kalliovirta+Malinen:2020, who identify the shocks in the GMVAR model by restricting the reduced form error covariance matrices and allow the impact responses of the variables to vary relative to each other across the regimes. They estimate the impulse response functions for each regime separately as if each of them was a linear VAR. In contrast, we allow the regime to switch as a result of a shock and compute the true (generalized) impulse response functions Koop+Pesaran+Potter:1996 of the nonlinear SVAR. Burgard+Neuenkirch+Nockel:2019, in turn, propose a mixture VAR with logistic mixing weights and recursively identified shocks. As opposed to Burgard+Neuenkirch+Nockel:2019, our identification scheme is more flexible in the sense that it does not require many (or necessarily any) zero restrictions on the impact effects of the shocks. Moreover, in our model, the regime-switching probabilities depend on the full distribution of the preceding $p$ observations instead of just on the levels of the switching-variables.

To illustrate the use of our model, we study the effects of the Euro area monetary policy shock. We fit a G-StMVAR model with two Student's $t$ regimes to monthly data covering the period from January 1999 to December 2021 and consisting of a number of macroeconomic variables. We find that one regime (Regime 1) is characterized by a negative (but volatile) output gap, and it mainly prevails after the Financial crisis, whereas the other regime (Regime 2) is characterized by a positive output gap and mainly dominates before the Financial crisis. The effects of the monetary policy shock are stronger in Regime 2 than in Regime 1, but asymmetries with respect to the sign and size of the shock are weak. In both regimes, a contractionary monetary policy shock decreases output gap significantly and persistently. However, while the price level decreases significantly and permanently in Regime 2, the decrease is insignificant in Regime 1.

The rest of this paper is organized as follows. Section (ref) introduces the linear Student's $t$ VAR and establishes its stationarity properties. Section (ref) introduces the G-StMVAR model and discusses its properties. Section (ref) introduces the structural G-StMVAR model, and Section (ref) discusses impulse responses analysis. Section (ref) discusses estimation of the model parameters by the method of maximum likelihood (ML) and establishes the asymptotic properties of the ML estimator. Section (ref) discusses a strategy for building a G-StMVAR model, and Section (ref) presents the empirical application. Appendix (ref) provides technical details related to Section (ref), and in Appendix (ref), the density functions and some properties of the Gaussian and Student's $t$ distributions are discussed. Appendix (ref) gives proofs of the stated theorems, Appendix (ref) describes a Monte Carlo algorithm for estimating the generalized impulse response function, and Appendix (ref) provides details on the empirical application. Finally, the introduced methods are implemented to the accompanying CRAN distributed R package gmvarkit gmvarkit, which provides tools for estimation and other numerical analysis of the introduced model.

Throughout this paper, we use the following notation. We write $x=(x_1,...,x_n)$ for the column vector $x$ where the components $x_i$ may be either scalars or (column) vectors. The notation $x\sim n_d(\mu,\Sigma)$ signifies that the random vector $x$ has a $d$-dimensional Gaussian distribution with mean $\mu$ and (positive definite) covariance matrix $\Sigma$, and $n_d(\cdot;\mu,\Sigma)$ denotes the corresponding density function. Similarly, $x\sim t_d(\mu,\Sigma,\nu)$ signifies that $x$ has a $d$-dimensional $t$-distribution with mean $\mu$, (positive definite) covariance matrix $\Sigma$, and degrees of freedom $\nu$ (assumed to satisfy $\nu>2$), and $t_d(\cdot;\mu,\Sigma,\nu)$ denotes the corresponding density function. The $(d\times d)$ identity matrix is denoted by $I_d$, $\otimes$ denotes the Kronecker product, and $\boldsymbol{1}_d$ denotes a $d$-dimensional vectors of ones.

Linear Gaussian and Student's t Vector Autoregressions

The G-StMVAR model accommodates two types of mixture components: conditionally homoskedastic linear Gaussian vector autoregressions and conditionally heteroskedastic linear Student's $t$ vector autoregressions. This section defines these linear vector autoregressions and establishes their stationarity properties. Consider the $d$-dimensional linear VAR model defined as

equation[equation omitted — 106 chars of source]

where the error process $\varepsilon_t$ is independent and identically distributed (IID), $\Omega_t^{1/2}$ is a symmetric square root matrix of the positive definite $(d\times d)$ covariance matrix $\Omega_t$ for all $t$, and $\phi_0\in\mathbb{R}^d$ is an intercept parameter. The $(d \times d)$ autoregression matrices are assumed to satisfy $\boldsymbol{A}_p \equiv [A_1:...:A_p]\in\mathbb{S}^{d\times dp}$, where

equation[equation omitted — 177 chars of source]

defines the usual stability condition of a linear VAR. The linear Gaussian VAR is obtained from Equation ((ref)) by assuming that $\varepsilon_t$ follows the $d$-dimensional standard normal distribution and that the conditional covariance matrix is a constant, $\Omega_t=\Omega$. We will first establish the stationarity properties of the linear Gaussian VAR, and by making use of the introduced notation, we then introduce the linear Student's $t$ VAR.

Under the stability condition, the linear Gaussian VAR is stationary, and the following properties are obtained. Denoting $\boldsymbol{z}_t=(z_t,...,z_{t-p+1})$ and $\boldsymbol{z}_t^{+}=(z_t,\boldsymbol{z}_{t-1})$, it is well known that the stationary solution to ((ref)) satisfies

align[align omitted — 317 chars of source]

where the last line defines the conditional distribution of $z_t$ given $\boldsymbol{z}_{t-1}$. Formulas of the quantities $\mu,\Sigma_p,\Sigma_{p+1}$ are given in Appendix (ref).

We show in Appendix (ref) (Theorem (ref)) that there exists linear Student's $t$ VARs with stationarity properties analogous ((ref)). Specifically, such Student's $t$ VARs are obtained by assuming that $\varepsilon_t$ follows the $d$-dimensional Student's $t$ distribution with mean zero, identity covariance matrix and $\nu+dp$ degrees of freedom and by defining conditional covariance matrix of $z_t$ as

equation[equation omitted — 198 chars of source]

Using the same notation as with the linear Gaussian VAR, the stationary solution to ((ref)) for the above-defined Student's $t$ VAR satisfies (Theorem (ref) in Appendix (ref))

align[align omitted — 338 chars of source]

Our Student's $t$ VAR has a conditional mean identical to the Gaussian VAR, but unlike the Gaussian VAR, it is conditionally heteroskedastic. Specifically, the conditional covariance matrix ((ref)) consists of a constant covariance matrix that is multiplied by a time-varying scalar that depends on the quadratic form of the preceding $p$ observations through the autoregressive parameters. In this sense, the model has a ‘VAR($p$)–ARCH($p$)’ representation, but the ARCH type conditional variance is not as general as in the conventional multivariate ARCH process Lutkepohl:2005 that allows the entries of the conditional covariance matrix to vary relative to each other. Our model is, however, more parsimonious than the conventional VAR-ARCH model, as the conditional covariance depends only on the degrees of freedom and autoregressive parameters (in addition to the parameters in the constant covariance matrix). Student's $t$ VARs similar to ours have previously appeared at least in Heracleous:2003 and Poudyal:2012.

The Gaussian and Student's t Mixture Vector Autoregressive Model

The G-StMVAR model can be described as a collection of linear autoregressive models each of which is a linear Gaussian VAR or a linear Student's $t$ VAR defined in Section (ref). At each time point, the process generates an observation from one of its mixture components that is randomly selected according to the probabilities given by the mixing weights. This definition is formalized in the following.

Let $y_t$, $t=1,2,...$, be the real valued $d$-dimensional time series of interest, and let $\mathcal{F}_{t-1}$ denote $\sigma$-algebra generated by the random vectors $\lbrace y_s, s<t \rbrace$. In a G-StMVAR model with autoregressive order $p$ and $M$ mixture components (or regimes), the observations $y_t$ are assumed to be generated by

align[align omitted — 179 chars of source]

where the following conditions Kalliovirta+Meitz+Saikkonen:2016 hold.

condition\ \begin{enumerate}[label=(\alph*)] • For $m=1,...,M_1\leq M$, the random vectors $\varepsilon_{m,t}$ are IID $n_d(0, I_d)$ distributed, and for $m=M_1+1,..., M$, they are IID $t_d(0, I_d,\nu_m + dp)$ distributed. For all $m$, $\varepsilon_{m,t}$ are independent of $\mathcal{F}_{t-1}$. • For each $m=1,...,M$$\phi_{m,0}\in\mathbb{R}^d$, $\boldsymbol{A}_{m,p} \equiv [A_{m,1}:...:A_{m,p}]\in\mathbb{S}^{d\times dp}$ (the set $\mathbb{S}^{d\times dp}$ is defined in ((ref))), and $\Omega_m$ is positive definite. For $m=1,...,M_1$, the conditional covariance matrices are constants, $\Omega_{m,t}=\Omega_m$. For $m=M_1+1,...,M$, the conditional covariance matrices $\Omega_{m,t}$ are as in ((ref)), except that $\boldsymbol{z}_{t-1}$ is replaced with $\boldsymbol{y}_{t-1}=(y_{t-1},...,y_{t-p})$ and the regime specific parameters $\phi_{m,0}$, $\boldsymbol{A}_{m,p}$, $\Omega_m$, $\nu_m$ are used to define the quantities therein. For $m=M_1+1,...,M$, also $\nu_m>2$. • The unobservable regime variables $s_{1,t},...,s_{M,t}$ are such that at each $t$, exactly one of them takes the value one and the others take the value zero according to the conditional probabilities expressed in terms of the ($\mathcal{F}_{t-1}$-measurable) mixing weights $\alpha_{m,t}\equiv \mathrm{P}(s_{m,t}=1|\mathcal{F}_{t-1})$ that satisfy $\sum_{m=1}^M\alpha_{m,t}=1$. • Conditionally on $\mathcal{F}_{t-1}$, $(s_{1,t},...,s_{M,t})$ and $\varepsilon_{m,t}$ are assumed independent. \end{enumerate}

The conditions $\nu_m>2$ in (ref) are made to ensure the existence of second moments. This definition implies that the G-StMVAR model generates each observation from one of its mixture components, a linear Gaussian or Student's $t$ vector autoregression discussed in Section (ref), and that the mixture component is selected randomly according to the probabilities given by the mixing weights $\alpha_{m,t}$.

Without loss of generality, the first $M_1$ mixture components are assumed to be linear Gaussian VARs, and the last $M_2\equiv M - M_1$ mixture components are assumed to be linear Student's $t$ VARs. If all the component processes are Gaussian VARs ($M_1=M$), the G-StMVAR model reduces to the GMVAR model of Kalliovirta+Meitz+Saikkonen:2016. If all the component processes are Student's $t$ VARs ($M_1=0$), we refer to the model as the StMVAR model.

Equations ((ref)) and ((ref)) and Condition (ref) lead to a model in which the conditional density function of $y_t$ conditional on its past, $\mathcal{F}_{t-1}$, is given as

equation[equation omitted — 192 chars of source]

The conditional densities $n_d(y_t;\mu_{m,t},\Omega_{m,t})$ and $t_d(y_t;\mu_{m,t},\Omega_{m,t},\nu_m+dp)$ are obtained from ((ref)) and ((ref)), respectively. The explicit expressions of the density functions are given in Appendix (ref). To fully define the G-StMVAR model it is then left to specify the mixing weights $\alpha_{m,t}$.

Analogously to Kalliovirta+Meitz+Saikkonen:2015, Kalliovirta+Meitz+Saikkonen:2016, Meitz+Preve+Saikkonen:2023, and Virolainen:2022, Virolainen:2024, we define the mixing weights as weighted ratios of the stationary densities of the regimes corresponding to the previous $p$ observations. To formally specify the mixing weights, we first define the following function for notational convenience. Let

equation[equation omitted — 336 chars of source]

where the $dp$-dimensional densities $n_{dp}(\boldsymbol{y};\mathbf{1}_p\otimes\mu_m,\Sigma_{m,p})$ and $t_{dp}(\boldsymbol{y};\mathbf{1}_p\otimes\mu_m,\Sigma_{m,p},\nu_m)$ correspond to the stationary distribution of the $m$th regime (given in Equation ((ref)) for the Gaussian regimes and in Theorem (ref) for the Student's $t$ regimes). Denoting $\boldsymbol{y}_{t-1}=(y_{t-1},...,y_{t-p})$, the mixing weights of the G-StMVAR model are defined as

equation[equation omitted — 238 chars of source]

where $\alpha_m\in (0,1)$, $m=1,...,M$, are mixing weights parameters assumed to satisfy $\sum_{m=1}^M\alpha_m = 1$, $\mu_m = (I_d - \sum_{i=1}^pA_{m,i})^{-1}\phi_{m,0}$, and covariance matrix $\Sigma_{m,p}$ is given in Equations ((ref)), ((ref)), and ((ref)) (in Appendix (ref)) but using the regime specific parameters to define the quantities therein.

Because the mixing weights are weighted ratios of the stationary densities of the regimes corresponding to the previous $p$ observations, the greater the relative weighted likelihood of a given regime is, the more likely the process it to generate an observation from it. This is a convenient feature for forecasting, and it also facilitates associating statistical characteristics and economic interpretations to the regimes. Moreover, it turns out that this specific formulation of the mixing weights leads to attractive theoretical properties such as full knowledge of the stationary distribution of $p+1$ consecutive observations and ergodicity of the process. These properties are summarized in Theorem (ref) below.

Before stating the theorem, a few notational conventions are provided. We collect the parameters of a G-StMVAR model to the $((M(d + d^2p + d(d+1)/2 + 2) - M_1 - 1)\times 1)$ vector $\boldsymbol{\theta}=(\boldsymbol{\vartheta}_1,...,\boldsymbol{\vartheta}_M,\alpha_1,...,\alpha_{M-1},\boldsymbol{\nu})$, where $\boldsymbol{\vartheta}_m=(\phi_{m,0},vec(\boldsymbol{A}_{m,p}),vech(\Omega_m))$ and $\boldsymbol{\nu}=(\nu_{M_1+1},...,\nu_M)$. The last mixing weight parameter is obtained as $\alpha_M=1-\sum_{m=1}^{M-1} \alpha_m$. The G-StMVAR model with autoregressive order $p$, and $M_1$ Gaussian and $M_2$ Student's $t$ mixture components is referred to as the G-StMVAR($p,M_1,M_2$) model, whenever the order of the model needs to be emphasized.

theoremConsider the G-StMVAR process $y_t$ generated by ((ref)), ((ref)), and ((ref)) with Condition (ref) satisfied. Then, $\boldsymbol{y}_t=(y_t,...,y_{t-p+1})$ is a Markov chain on $\mathbb{R}^{dp}$ with stationary distribution characterized by the density \begin{equation} f(\boldsymbol{y};\boldsymbol{\theta}) = \sum_{m=1}^{M_1}\alpha_m n_{dp}(\boldsymbol{y};\boldsymbol{1}_p\otimes\mu_{m},\Sigma_{m,p}) + \sum_{m=M_1+1}^M\alpha_mt_{dp}(\boldsymbol{y};\boldsymbol{1}_p\otimes\mu_{m},\Sigma_{m,p},\nu_m). \end{equation} Moreover, $\boldsymbol{y}_t$ is ergodic.

The stationary distribution is a mixture of $M_1$ $dp$-dimensional Gaussian distributions and $M_2$ $dp$-dimensional $t$-distributions with constant mixing weights $\alpha_m$. The proof of Theorem (ref) in Appendix (ref) shows that the marginal stationary distributions of $1,...,p+1$ consecutive observations are likewise mixtures of Gaussian and $t$-distributions. This gives the mixing weight parameters $\alpha_m$$m=1,..,M$, interpretation as the unconditional probabilities of an observation being generated from the $m$th component process. The unconditional mean, covariance, and first $p$ autocovariances are hence easily obtained similarly to the GMVAR model of Kalliovirta+Meitz+Saikkonen:2016.

Our StMVAR model features ARCH type conditional heteroskedasticity in the sense that the conditional covariance matrix of each regime depends on the quadratic form of the preceding $p$ observations through the covariance matrix, autoregressive, and degrees of freedom parameters (see the discussion at the end of Section (ref)). This property arises from utilizing the multivariate Student's $t$-distribution as the stationary distribution of the regimes. The use of the $t$-distribution thereby allows for parsimonious modelling of series that display fat tails and conditional heteroskedasticity within the regimes. This is particularly advantageous in the context of mixture VARs, as the large number of parameters may often be problematic even without an ARCH component. Appropriate modelling of kurtosis and conditional heteroskedasticity is important, because they may affect the endogenously determined regime-switching probabilities. Ignoring these factors would leave out potentially important dynamics that may affect the results of an empirical investigation. If some of the regimes have a constant conditional covariance matrix and zero excess kurtosis, we allow them to be conditionally homoskedastic linear Gaussian VARs, which leads to the G-StMVAR model.

Structural G-StMVAR Model

To facilitate structural analysis, the structural shocks must be identified by orthogonalizing the reduced form errors so that they conditional covariance matrix is a diagonal matrix. Consider the G-StMVAR model defined by ((ref)), ((ref)), and ((ref)) with Condition (ref) satisfied. We write the structural G-StMVAR model as

align[align omitted — 220 chars of source]

where $e_t=(e_{1t},...,e_{dt})$ $(d \times 1)$ is an orthogonal structural error. The definition ((ref)) of the structural error is similar to the definition (3.1) and (3.2) in Virolainen:2024, who studied a linear SVAR model incorporating switching in the volatility regime, but with shocks arriving from Student's $t$ regimes in addition to (or instead of) Gaussian ones.

For the Gaussian regimes ($m=1,...,M_1$)‚ $\Omega_{m,t}=\Omega_m$, whereas for the Student's $t$ regimes ($m=M_1+1,...,M$), $\Omega_{m,t}=\omega_{m,t}\Omega_m$, where

equation[equation omitted — 191 chars of source]

The invertible $(d\times d)$ impact matrix $B_t$, which governs the contemporaneous relationships of the shocks, is time-varying and a function of $y_{t-1},..., y_{t-p}$. Following Virolainen:2024, we define the impact matrix so that it captures the conditional heteroskedasticity of the reduced form error, and thereby amplifies a constant-sized structural shock accordingly. Appropriate modelling of conditional heteroskedasticity in the impact matrix is of interest because the (generalized) impulse response functions may be asymmetric with respect to the size of the shock.

We have $\Omega_{u,t}\equiv\text{Cov}(u_t|\mathcal{F}_{t-1})=\sum_{m=1}^{M_1}\alpha_{m,t}\Omega_m + \sum_{m=M_1+1}^{M}\alpha_{m,t}\omega_{m,t}\Omega_m$, while the conditional covariance matrix of the structural error $e_t=B_t^{-1}u_t$ (which are not IID but martingale differences and thereby uncorrelated) is obtained as

equation[equation omitted — 190 chars of source]

The impact matrix $B_t$ should be chosen so that the right side of Equation ((ref)) is a diagonal matrix. Virolainen:2024 shows that any such impact matrix that simultaneously diagonalizes $\Omega_1,...,\Omega_M$ (and thus also $\Omega_{1,t},...,\Omega_{M,t}$) has linearly independent eigenvectors of the matrix $\Omega_m\Omega_1^{-1}$ as its columns. Moreover, he shows that under the following assumption and a constant normalization of the structural error's conditional covariance matrix, say, $\Omega_{u,t}=I_d$, the impact matrix is unique up to ordering of its columns and changing all signs in a column.\footnote{Virolainen:2024 shows the uniqueness of the impact matrix for a linear SVAR model with shocks arriving from a mixture of Gaussian distributions, but the results directly apply to our structural G-StMVAR model as well.}

assumptionConsider $M$ positive definite $(d\times d)$ covariance matrices $\Omega_m$, $m = 1,...,M$, and denote the strictly positive eigenvalues of the matrices $\Omega_m\Omega_1^{-1}$ as $\lambda_{mi}$, $i=1,...,d$, $m=2,...,M$. Suppose that for all $i\neq j\in\lbrace 1,...,d\rbrace$, there exists an $m\in\lbrace 2,...,M\rbrace$ such that $\lambda_{mi}\neq\lambda_{mj}$.

Following Virolainen:2024 (and Lanne+Lutkepohl:2010 and Lanne+Lutkepohl+Maciejowska:2010), we utilize the following matrix decomposition that is convenient for specifying the impact matrix. We decompose the error term covariance matrices as:

equation[equation omitted — 116 chars of source]

where the diagonal of $\Lambda_m=\text{diag}(\lambda_{m1},...,\lambda_{md})$, $\lambda_{mi}>0$ ($i=1,...,d$), contains the eigenvalues of the matrix $\Omega_m\Omega_1^{-1}$ and the columns of the nonsingular $W$ are the related eigenvectors (which are identical for all $m$ by construction). When $M=2$, decomposition ((ref)) always exists Muirhead:1982, but for $M\geq 3$ its existence requires that the matrices $\Omega_m\Omega_1^{-1}$ share the common eigenvectors in $W$. If this is not the case, the impact matrix does not exist Virolainen:2024; however, its existence is testable.

Similarly to Virolainen:2024, any scalar multiples of the columns of $W$ comprise an appropriate impact matrix, but only specific scalar multiples comprise the locally unique impact matrix associated with a given normalization of the structural error's conditional covariance matrix. Direct calculation shows that the impact matrix associated with the normalization $\text{Cov}(e_t|\mathcal{F}_{t-1})=I_d$ is obtained as

equation[equation omitted — 144 chars of source]

where $B_tB_t'=\Omega_{u,t}$. Since $B_t^{-1}\Omega_mB_t'^{-1}=\Lambda_m(\sum_{n=1}^{M_1}\alpha_{n,t}\Lambda_n + \sum_{n=M_1+1}^M\alpha_{n,t}\omega_{n,t}\Lambda_n)^{-1}$, the impact matrix ((ref)) simultaneously diagonalizes $\Omega_{1},...,\Omega_{M}$, and $\Omega_{u,t}$ (and thereby also $\Omega_{1,t},...,\Omega_{M,t}$) for each $t$ so that $\text{Cov}(e_t|\mathcal{F}_{t-1}) = I_d$.

Our model identifies the shocks up to ordering and sign under Assumption (ref), and therefore, global identification of the shocks is obtained by fixing the signs and ordering of the columns of $B_t$. The ordering of the columns can be fixed by choosing an arbitrary ordering for the eigenvalues in the diagonals of $\Lambda_m$, $m=2,..,M$. The signs, in turn, can be normalized by placing a single strict sign constraint in each column of $B_t$. The identification is, however, a statistical one, and it does not reveal which column of the impact matrix is related to which shock, nor does it necessarily lead to structural shocks with economic interpretations. Moreover, as the structure of the impact matrix in ((ref)) shows, it follows from the simultaneous diagonalization of the covariance matrices that, for each shock, the relative impact responses of the variables are restricted to be constant over time. More specifically, this identifying restriction implies that the impact responses to the $i$th shock $e_{1t}$ are $w_i(\sum_{m=1}^{M_1}\alpha_{m,t}\lambda_{mi} + \sum_{m=M_1 + 1}^{M}\alpha_{m,t}\omega_{m,t}\lambda_{mi})^{1/2}e_{1t}$, where $w_i$ $(d\times 1)$ is the $i$th column of $W$.

Labelling the identified structural shocks by economic shocks requires external information, for instance, in the form of economically motivated overidentifying restrictions as in Lanne+Lutkepohl:2010 and Virolainen:2024. Such restrictions may readily be satisfied by the unrestricted estimate of $B_t$ (as in Virolainen:2024), and if not, the validity of the overidentifying economic restrictions can be tested for (as in Lanne+Lutkepohl:2010). However, the validity of Assumption (ref) required for full identification of the shocks is difficult to justify formally. If Assumption (ref) fails, further restrictions such as (untestable) zero restrictions on $B_t$ are required for the identification Virolainen:2024.

Impulse response analysis

The expected effects of the structural shocks in the the structural G-StMVAR model generally depend on the initial values, as well as on the sign and size of the shock. Additionally, the regime may also switch as a result of a shock. This makes the conventional way of calculating impulse responses unsuitable Kilian+Lutkepohl:2017. However, the aforementioned features of our SVAR model can be captured with the generalized impulse response function (GIRF) Koop+Pesaran+Potter:1996, defined as:

equation[equation omitted — 163 chars of source]

where $h$ is the horizon, $e_{it}$ is $j$th element of $e_t$, and $\mathcal{F}_{t-1}=\sigma\lbrace y_{t-j},j>0\rbrace$ as before. The first term on the right side of ((ref)) is the expected realization of the process at time $t+h$ conditionally on a structural shock of sign and size $\delta_i \in\mathbb{R}$ in the $i$th element of $e_t$ at time $t$ and the previous observations. The latter term on the right side is the expected realization of the process conditionally on the previous observations only. The GIRF thus expresses the expected difference in the future outcomes when the structural shock of sign and size $\delta_i$ in the $i$th element of $e_t$ hits the system at time $t$ as opposed to all shocks being random.

An interesting property of our SVAR model is that it also facilitates estimating the effects of shocks on the mixing weights. In other words, different from most nonlinear SVAR models, in our setup, it is possible to assess the likelihood of future regime switches due to each shock. The related impulse response functions are obtained by replacing $y_{t+h}$ by $\alpha_{m,t+h}$ in Equation ((ref)).

The G-StMVAR model has a $p$-step Markov property, so conditioning on (the $\sigma$-algebra generated by) the $p$ previous observations $\boldsymbol{y}_{t-1}=(y_{t-1},...,y_{t-p})$ is effectively the same as conditioning on $\mathcal{F}_{t-1}$ at time $t$ and later. The history $\boldsymbol{y}_{t-1}$ can be either fixed or random, but with a random history, the GIRF becomes a random vector. Using a fixed $\boldsymbol{y}_{t-1}$ is useful when the interest is in the effects of a shock at a particular point of time, whereas more general results are obtained by assuming that $\boldsymbol{y}_{t-1}$ follows some distribution. For example, the overall effects of a shock in the G-StMVAR model can be estimated by assuming that $\boldsymbol{y}_{t-1}$ follows the stationary distribution of the model (characterized in Theorem (ref)). On the other hand, if the interest is in the effects of the shock in particular state of the economy represented by a regime, $\boldsymbol{y}_{t-1}$ can be assumed to follow the stationary distribution of the corresponding regime (presented in Equations ((ref)) and ((ref))).

The GIRF and its distributional properties can be estimated with a Monte Carlo algorithm that generates (partial) realizations of the process and then takes the sample mean for a point estimate. If $\boldsymbol{y}_{t-1}$ is random and follows distribution $G$, the GIRF should be estimated for different values of $\boldsymbol{y}_{t-1}$ generated from $G$, and then the sample mean and sample quantiles can be taken to obtain the point estimate and confidence intervals that reflect the uncertainty about the initial value. Such an algorithm, adapted from Koop+Pesaran+Potter:1996 and Kilian+Lutkepohl:2017, is given in Appendix (ref).

Our SVAR model defined in Section (ref) assumes a single impact matrix that captures the conditional covariance matrix of the reduced form error, which varies according to the mixing weights. This specification allows the impact responses of the variables to vary in magnitude, but due to the simultaneous diagonalization of the error term covariance matrices, the impact response of any variable to each shock is restricted to be constant relative to the impact responses of the other variables (see the discussion in Section (ref)). It would, of course, be possible to relax the restriction of constant relative impact effects, but doing so would require additional restrictions to be imposed to achieve identification. Bacchiocchi+Fanelli:2015, and Angelini+Bacchiocchi+Caggiano+Fanelli:2019 have proposed such a model with exogenous change points of the volatility regime, where identification is based on combining heteroskedasticity with (untestable) zero restrictions on the impact matrices in different regimes. In addition to the untestable restrictions being potentially challenging to justify economically, their setup is not as straightforward to apply as ours.

Through regime switches, our SVAR model still accommodates impulse responses that vary in the periods after the impact depending on the initial state of the economy as well as on the sign and size of the shock. Specifically, as the regime may switch as a result of a shock, different signs and sizes of the shock may induce different responses of the mixing weights from different initial values $\boldsymbol{y}_{t-1}$, leading to nonlinear impulse responses in the periods after the impact period. This follows from the fact that the intercepts and AR matrices in ((ref))-((ref)) are not restricted to be constant across the regimes. While allowing for regime-specific autoregressive parameters is different from much of the previous statistical identification literature, it is always possible to test for the constancy of these parameters, and in case of non-rejection, proceed with a linear SVAR model with $M$ volatility regimes. This test and estimation of the linear SVAR model are also implemented in the accompanying R package gmvarkit gmvarkit. For instance, in our empirical application in Section (ref), the null hypothesis of constancy is clearly rejected (see Appendix (ref)).

Estimation

The parameters of the G-StMVAR model can be estimated by the method of maximum likelihood (ML), and the exact log-likelihood function is available, as we have established the stationary distribution of the process in Theorem (ref). Suppose the observed time series is $y_{-p+1},...,y_0,y_1,...,y_T$ and that the initial values are stationary. Then, the log-likelihood function of the G-StMVAR model takes the form

equation[equation omitted — 211 chars of source]

where $d_{m,dp}(\cdot;\boldsymbol{1}_p\otimes\mu_m,\Sigma_{m,p},\nu_m)$ is defined in ((ref)) and

equation[equation omitted — 219 chars of source]

If stationarity of the initial values seems unreasonable, one can condition on the initial values and base the estimation on the conditional log-likelihood function, which is obtained by dropping the first term on the right side of ((ref)). The rest of this section assumes that estimation is based on the conditional log-likelihood function divided by the sample size, $L_T^{(c)}(\boldsymbol{\theta})=T^{-1}\sum_{m=1}^M l_t(\boldsymbol{\theta})$, i.e., the ML estimator $\hat{\boldsymbol{\theta}}_T$ maximizes $L_T^{(c)}(\boldsymbol{\theta})$.

If there are two regimes in the model ($M=2$), the structural G-StMVAR model discussed in Section (ref) is obtained from the estimated reduced form model by decomposing the covariance matrices $\Omega_1,...,\Omega_M$ as in ((ref)). If $M\geq 3$ or overidentifying constraints are imposed on $B_t$ through $W$, the model can be reparametrized with $W$ and $\Lambda_m$ ($m=2,...,M$) instead of $\Omega_1,...,\Omega_M$, and the log-likelihood function can be maximized subject to the new set of parameters and constraints. In this case, the decomposition ((ref)) is plugged into the log-likelihood function and $vech(\Omega_1),...,vech(\Omega_M)$ are replaced with $vec(W)$ and $\boldsymbol{\lambda}_2,...,\boldsymbol{\lambda}_M$ in the parameter vector $\boldsymbol{\theta}$, where $\boldsymbol{\lambda}_m=(\lambda_{m1},...,\lambda_{md})$. Instead of constraining $vech(\Omega_1),...,vech(\Omega_M)$ so that $\Omega_1,...,\Omega_M$ are positive definite, we impose the constraints $\lambda_{mi}>0$ for all $m=2,...,M$ and $i=1,...,d$.

Establishing the asymptotic properties of the ML estimator requires that it is uniquely identified. In order to achieve unique identification, the parameters need to be constrained so that the mixture components cannot be 'relabelled' to produce the same model with a different parameter vector. To that end, we assume that

align[align omitted — 376 chars of source]

In the case of the structural G-StMVAR model, identification also requires that Assumption (ref) is satisfied (see Section (ref)).\footnote{With the appropriate zero restrictions on $W$, this condition can be relaxed, however Virolainen:2024.} Then, identification of the structural model follows from the identification of the reduced form model.

We summarize the constraints imposed on the parameter space in the following assumption.

assumptionThe true parameter value $\boldsymbol{\theta}_0$ is an interior point of $\boldsymbol{\Theta}$, which is a compact subset of $\lbrace \boldsymbol{\theta}=(\boldsymbol{\vartheta}_1,...,\boldsymbol{\vartheta}_M,\alpha_1,...,\alpha_{M-1},\boldsymbol{\nu})\in\mathbb{R}^{M(d + d^2p + d(d+1)/2)}\times (0,1)^{M-1}\times (2,\infty)^{M_2}:\boldsymbol{A}_{m,p}\in\mathbb{S}^{d\times dp}, \Omega_m$ is positive definite, for all $m=1,...,M$, and ((ref)) holds$\rbrace$.

Asymptotic properties of the ML estimator under the conventional high-level conditions are stated in the following theorem (which is proven in Appendix (ref)). Denote $\mathcal{I}(\boldsymbol{\theta})=E\left[\frac{\partial l_t(\boldsymbol{\theta})}{\partial \boldsymbol{\theta}}\frac{\partial l_t(\boldsymbol{\theta})}{\partial \boldsymbol{\theta}'} \right]$ and $\mathcal{J}(\boldsymbol{\theta})=E\left[\frac{\partial^2 l_t(\boldsymbol{\theta})}{\partial \boldsymbol{\theta}\partial\boldsymbol{\theta}'} \right]$.

theoremSuppose that $y_t$ are generated by the stationary and ergodic G-StMVAR process of Theorem (ref) and that Assumption (ref) holds. Then, $\hat{\boldsymbol{\theta}}_T$ is strongly consistent, i.e., $\hat{\boldsymbol{\theta}}_T\rightarrow \boldsymbol{\theta}_0$ almost surely. Suppose further that (i) $T^{1/2}\frac{\partial}{\partial\boldsymbol{\theta}_0}L_T^{(c)}(\boldsymbol{\theta}_0) \overset{d}{\rightarrow}N(0,\mathcal{I}(\boldsymbol{\theta}_0))$ with $\mathcal{I}(\boldsymbol{\theta}_0)$ finite and positive definite, (ii) $\mathcal{J}(\boldsymbol{\theta}_0)=-\mathcal{I}(\boldsymbol{\theta}_0)$, and (iii) $E[\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}_0}|\frac{\partial^2 l_t(\boldsymbol{\theta})}{\partial\boldsymbol{\theta}\partial\boldsymbol{\theta}'}|]<\infty$ for some $\boldsymbol{\Theta}_0$, compact convex set contained in the interior of $\boldsymbol{\Theta}$ that has $\boldsymbol{\theta}_0$ as an interior point. Then $T^{1/2}(\hat{\boldsymbol{\theta}}_T - \boldsymbol{\theta}_0)\overset{d}{\rightarrow} N(0,-\mathcal{J}(\boldsymbol{\theta}_0)^{-1})$.

Given consistency, conditions (i)-(iii) of Theorem (ref) are standard for establishing asymptotic normality of the ML estimator, but their verification can be tedious. If one is willing to assume the validity of these conditions, the ML estimator has the conventional limiting distribution, implying that the approximate standard errors for the estimates are obtained as usual. Furthermore, the standard likelihood based tests are applicable as long as the number of mixture components is correctly specified.\footnote{This condition is important, because if the number of Gaussian or Student's $t$ type mixture components is chosen too large, some of the parameters are not identified causing the result of Theorem (ref) to break down. This particularly happens when one tests for the number of regimes, as under the null some of the regimes are removed from the model. Meitz+Saikkonen:2021 have, however, recently developed such tests for mixture autoregressive models with Gaussian conditional densities. Developing a test for the number of a regimes in the G-StMVAR model is a major task and beyond the scope of this paper. Likewise, when testing whether a regime is a Gaussian VAR against the alternative that it is a Student's $t$ VAR, under the null, $\nu_m=\infty$ for the Student's $t$ regime $m$ to be tested, which violates Assumption (ref).}

Finding the ML estimate amounts to maximizing the log-likelihood function defined in ((ref)) and ((ref)) over a high dimensional parameter space satisfying the constraints in Assumption (ref). Due to the complexity of the log-likelihood function, numerical optimization methods are required. The maximization problem can be challenging in practice due to complex dependence of the mixing weights on the preceding observations, which induces a large number of modes to the surface of the log-likelihood function, and large areas to the parameter space, where it is flat in multiple directions. Also, the popular EM algorithm Redner+Walker:1984 is virtually useless here, as at each maximization step one faces a new optimization problem which not much simpler than the original one. Following Meitz+Preve+Saikkonen:2023 and Virolainen:2022, Virolainen:2024, we therefore employ a two-phase estimation procedure in which a genetic algorithm is used to find starting values for a gradient based variable metric algorithm. The accompanying R package gmvarkit gmvarkit employs a modified genetic algorithm that works similarly to the one described in Virolainen:2022.

Building a G-StMVAR Model

Building a G-StMVAR model amounts to finding a suitable autoregressive order $p$, the number of Gaussian regimes $M_1$, and the number of Student's $t$ regimes $M_2$. Following the model selection strategy of Virolainen:2022 concerning the univariate counterpart of the G-StMVAR model, we propose taking advantage of the observation that the G-StMVAR model is a limiting case of the StMVAR model (in which all the mixture components are linear Student's $t$ VARs).

It is easy to check that the linear Gaussian vector autoregression defined in Section (ref) is a limiting case of the linear Student's $t$ vector autoregression when the degrees of freedom parameter tends to infinity. As the mixing weights ((ref)) are weighted ratios of the stationary densities of the regimes, it then follows that a G-StMVAR($p,M_1,M_2$) model is obtained as a limiting case of the StMVAR($p,M$) model (or equivalently the G-StMVAR($p,0,M$) model) with the degrees of freedom parameters of the first $M_1$ regimes tending to infinity. Since a StMVAR($p,M$) model that is fitted to data generated by a G-StMVAR($p,M_1,M_2$) process is, therefore, asymptotically expected to get large estimates for the degrees of freedom parameters of the $M_1$ Gaussian regimes, we propose selecting the model by finding a suitable StMVAR model. If the fitted StMVAR model contains overly large estimates of the degrees of freedom parameters, the corresponding regimes should be switched to Gaussian VARs by estimating the appropriate G-StMVAR model.

For a strategy to find a suitable StMVAR model, we follow Kalliovirta+Meitz+Saikkonen:2015, and suggest first considering the linear version of the model, that is, a StMVAR model with one mixture component. Partial autocorrelation functions, information criteria, and quantile residual diagnostics Kalliovirta+Saikkonen:2010 can be made use of for selecting the appropriate autoregressive order $p$. If the linear model is found inadequate, models with multiple mixture components can be examined. One should, however, be conservative with the choice of $M$, because if the number of regimes is chosen too large, some of the parameters are not identified. Adding new regimes to the model also vastly increases the number of parameters, and moreover, due to the increased complexity, it might be difficult to obtain the ML estimate if there are many regimes in the model.

Overly large degrees of freedom parameters are redundant in the model, but their weak identification also causes numerical problems. Specifically, they induce a numerically nearly singular Hessian matrix of the log-likelihood function when evaluated at the estimate, which makes the approximate standard errors and the quantile residual diagnostic tests of Kalliovirta+Saikkonen:2010 often unavailable. Switching to the appropriate G-StMVAR model has little effect of on the fit of the model in the case of overly large estimates of degrees of freedom parameters, so it is then advisable.

Empirical Application

We employ our SVAR model to study the effects of the Euro area monetary policy shock. Its asymmetric effects of have been studied, among others, by Peersman+Smets:2002 and Dolado+MariaDolores:2006, who found that it has larger effects on production during recessions than expansions. Pellegrino:2018 found the real effects of the monetary policy shock weaker during uncertain times than tranquil times, whereas Burgard+Neuenkirch+Nockel:2019 found the effects of contractionary monetary policy shocks stronger but less enduring during "crisis" than during "normal times".

We consider monthly Euro area data covering the period from January 1999 to December 2021 ($276$ observations) and consisting of four variables: industrial production index (IPI), harmonized consumer price index (HCPI), Brent crude oil price (Europe, OIL), and an interest rate variable (RATE). Our policy variable is the interest rate variable, which is the Euro overnight index average (EONIA) from January 1999 to October 2008 and the Wu+Xia:2016 shadow rate from November 2008 to December 2021. The Wu+Xia:2016 shadow rate is not bounded by the zero-lower-bound and also quantifies unconventional monetary policy measures.\footnote{The IPI, HCPI, and EONIA were obtained from the European Central Bank Statistical Data Warehouse; the Brent crude oil prices were retrieved from the Federal Reserve Bank of St. Louis database; and the Wu+Xia:2016 shadow rate was obtained from the first author's website. The dataset used is included the CRAN distributed R package gmvarkit gmvarkit.}

The IPI is detrended by first separating its cyclical component from the trend with the linear projection filter proposed by Hamilton:2018 and then considering the cyclical component.\footnote{Denoting the univariate, non-stationary time series as $y_t$, the filter defines its transient component at the time $t+h$ ($h>0$) as the ordinary least squares residuals from regressing $y_{t+h}$ on a constant and $y_t,...,y_{t-s+1}$. When $s$ is chosen larger than the order of integration, the residual process is stationary. We used the parameter values $h=24$ and $s=12$, as suggested by Hamilton:2018 for monthly data.} It is thereby implicitly assumed that the monetary policy shock does not have permanent effects on real industrial production. Hereafter, we refer to the IPI's deviation from the trend as the output gap. The logs of HCPI and the oil price are detrended by taking first differences, whereas the interest rate variable is assumed stationary. For numerical reasons, we multiply the cyclical component of the IPI and the log-difference of HCPI by $100$, and the log-difference of OIL by $10$. The series are presented in Figure (ref), where the shaded areas indicate the periods of Euro area recessions defined by the OECD OECD:2022.

figure[figure omitted — 1,030 chars of source]

For selecting the order of our G-StMVAR model, we start by estimating one-regime StMVAR models with autoregressive orders $p=1,...,12$ and find that AIC is minimized at $p=1$. Next, we estimate a two-regime StMVAR model with $p=1$ but find this model somewhat inadequate. Consequently, we increase the autoregressive order to $p=2$, which increases the AIC but improves the model's adequacy. The overall adequacy of the StMVAR($2, 2$) model, i.e., G-StMVAR($p=2$, $M_1=0$, $M_2=2$) model, is found reasonable, so we employ it for the further analysis. Because the model does not contain large degrees of freedom parameter estimates, we do not consider incorporating Gaussian mixture components by switching to a G-StMVAR model with $M_1>0$. Details on model selection and the adequacy of the selected model are provided in Appendix (ref).

The estimated mixing weights of the StMVAR($2,2$) model are presented in the bottom panel of Figure (ref). Regime 1 mainly prevails after the Financial crisis, but it obtains large mixing weights also before and during the early 2000's recession. Regime 2 dominates in the remaining time periods, that is, mainly before the Financial crisis. Since the prevailing regime starts switching sharply in October 2008, our model is consistent with the evidence that the ECB changed its reaction function after the bankruptcy of Lehman Brothers in September 2008 Gerlach+Lewis:2014.\footnote{Gerlach+Lewis:2014 found that the ECB was cutting the interest rates faster at the time of the crisis, and that the ECB started a policy shift back in the late 2010. According our StMVAR model, however, the dominating regime never switches back to the pre-Financial crisis regime (in our sample period), although the second regime obtains mixing weights clearly larger than zero in the late 2010 and several relatively large mixing weights in the early 2011.}

Based on the unconditional means and marginal standard deviations of the regimes (presented in Table (ref) in Appendix (ref)), the post-Financial crisis regime is characterized by negative but volatile output gap as well as lower inflation, oil price inflation, and interest rate interest rate than the pre-Financial crisis regime, which is characterized by positive output gap. The post-Financial crisis regime also exhibits higher kurtosis and overall volatility than the pre-Financial crisis regime (in all of the variables). Details on the characteristics of the regimes are discussed in Appendix (ref).

Identification of the Monetary Policy Shock

Decomposing the covariance matrices of the reduced form StMVAR($2,2$) model as in ((ref)) gives the following estimates for the structural parameters:

equation[equation omitted — 613 chars of source]

where the ordering of the variables is $y_t=(\text{IPI}_t, \text{HCPI}_t, \text{OIL}_t, \text{RATE}_t)$, the estimates $\hat{\lambda}_{2i}$ are in decreasing order (which fixes an arbitrary ordering for the columns of $\hat{W}$), and approximate standard errors are given in parentheses next to the estimates. The estimates that deviate from zero by more than two times their approximate standard error are bolded. Many of the estimates $\hat{\lambda}_{2i}$ are somewhat close to each other, suggesting that it might not be clear whether Assumption (ref) is satisfied. The validity of Assumption (ref) is difficult to justify formally Virolainen:2024; however, the robustness of our identification to its failure is discussed at the end of this section.

The estimates and approximate standard errors in ((ref)) show that the fourth shock is the only shock that moves the interest variable significantly on impact, and it is also the only shock that moves production in the opposite direction. Therefore, we deem it as the monetary policy shock. The fourth shock, however, appears to move inflation and oil price inflation in the same direction as the interest rate variable, which is contrary to the standard economic theory that states that an increase in the nominal interest rate should decrease inflation by decreasing aggregate demand Gali:2015.

Since the instantaneous increase of prices in response to a contractionary monetary policy shock is arguably implausible, we test for the zero restrictions that the instantaneous responses of inflation and oil price inflation are zero. The zero restrictions obtain the $p$-values $0.22$ and $0.15$ in a Wald test individually and the $p$-value $0.34$ jointly, so they are not rejected.\footnote{Interestingly, our results (suggesting that the zero restriction improves plausibility of the response of inflation) are contrary to Castelnuovo:2016, who argued that muted response of inflation in the Euro area could be caused by misspecified zero restrictions in the impact matrix.} In addition to dampening the implausible impact responses of prices, the zero restrictions are useful because by Propositions 2 and 3 of Virolainen:2024, they make the identification robust to the failure of Assumption (ref), whose validity is difficult to justify formally. Specifically, the monetary policy shock is still identified if $\lambda_{2i}=\lambda_{2j}$ for any $i,j=1,2,3$ and additionally $\lambda_{2i}=\lambda_{24}$ for any one of $i=1,2,3$ (but the approximate standard errors and the Wald test results are valid only if $\lambda_{2i}$ are all distinct). Therefore, we proceed with the model involving these zero restrictions.

Generalized impulse response function

We compute the GIRFs to the monetary policy shock to study its macroeconomic effects. As discussed in Section (ref), since we allow the regime to switch as a result of a shock, the structural StMVAR model accommodates asymmetries with respect the initial state of the economy as well as to the sign and size of the shock. Therefore, we compute the GIRFs conditionally on the initial values generated from the stationary distribution of each regime separately to positive (contractionary) and negative (expansionary) as well as to one-standard-error (small) and two-standard error (large) shocks. However, as the GIRFs to small and large shocks are similar, only the responses to small shocks are reported. We scale the GIRFs so that they correspond to a $25$ basis point instantaneous increase in the interest rate variable, making the responses to shocks of different signs and sizes comparable.

figure[figure omitted — 1,241 chars of source]

Figure (ref) presents the GIRFs $h=0,1,...,96$ months ahead computed for the identified monetary policy shock. The GIRFs to two-standard-error shocks are not reported for simplicity, since they seem similar to the GIRFs to one-standard-error shocks. In the post-Financial crisis Regime 1 (the left column in Figure (ref)), the GIRFs to contractionary and expansionary shocks are practically identical. Output gap decreases strongly at impact in response to a contractionary monetary policy shock. On average, the peak response is in the first period, and then the response starts to slowly decay towards zero. The confidence intervals show that with some of the starting values the peak effect occurs later, however. The price level barely moves, but there is an insignificant permanent decrease. The oil price increases slightly for roughly fifteen months before it decreases persistently, and the response of the interest rate variable, in turn, is very persistent. The weak response of mixing weights of Regime 1 shows that the monetary policy shock has little effect on the regime switching probabilities.

In the pre-Financial crisis Regime 2 (the right column in Figure (ref)), the responses of the observable variables are quite symmetric with respect to the sign of the shock, but there are some differences in the responses of the mixing weights. Output gap decreases strongly at impact in response to a contractionary monetary policy shock. The response then decreases and peaks after several years before slowly decaying towards zero. The price level starts to steadily decrease after the impact period in response to a contractionary shock and keeps decreasing over our horizon of eight years. The oil price moves similarly to the consumer prices, whereas the response of the interest rate variable is very persistent. Interestingly, an expansionary shock significantly increases the probability of (the low-growth) Regime 1 in the first period after impact, but in the following periods, it significantly increases the probability of Regime 2. A contractionary shock, in turn, increases the probability of Regime 1 from the first period onwards.

Overall, we find strong asymmetries with respect to the initial state of the economy, but the asymmetries seem weak with respect to the sign and size of the shock. The effects of the monetary policy shock on output gap and prices are stronger in the pre-Financial crisis Regime 2 than in the post-Financial crisis Regime 1. The real effects are significant in both of the regimes, but the inflationary effects are only minor in Regime 1, while they are substantial in Regime 2. Finally, the effects on the probabilities of future regime switches are clearly stronger in Regime 2.\footnote{For robustness, we compute the GIRFs also for the StMVAR($1,2$) suggested by AIC. We find the estimated mixing weights quite similar to the main specification: one of the regimes mainly prevails prior to the Financial crisis and one after it. and the effects of the monetary policy shock, particularly on inflation, are stronger in the pre-Financial crisis regime. For comparison, we also consider GMVAR of Kalliovirta+Meitz+Saikkonen:2016 identified by heteroskedasticity as in Virolainen:2024 and the standard linear recursively identified Gaussian SVAR model. We find that the structural GMVAR model produces both price puzzles and output puzzles (the latter one meaning that the output gap increases in response a contractionary shock or decreases in response to an expansionary shock). The linear recursively identified Gaussian SVAR model (with $p=3$ based on AIC) performs even worse: in response to a contractionary shock, the output gap increases strongly and persistently, and prices rise permanently (the results are similar with the autoregressive order $p=1,2$).}

Conclusion

We introduced a new mixture vector autoregressive model, referred to as the G-StMVAR model, which has attractive theoretical and practical properties. The G-StMVAR model accommodates conditionally homoskedastic Gaussian VARs and conditionally heteroskedastic Student's $t$ VARs as its mixture components. The mixing weights are defined as weighted ratios of the stationary densities of the regimes corresponding to the previous $p$ observations. This specification is appealing, as it states that the greater the relative weighted likelihood of a regime is, the more likely the process is to generate an observation from it, which facilitates associating economic interpretations to the regimes. The specific formulation of the mixing weights also leads to attractive theoretical properties such as ergodicity and full knowledge of the stationary distribution of $p+1$ consecutive observations. Moreover, the maximum likelihood estimator of a stationary G-StMVAR model is strongly consistent, and therefore, it has the conventional limiting distribution under conventional high level conditions.

The G-StMVAR model is a multivariate version of the G-StMAR model of Virolainen:2022. As a special case, by assuming that all the mixture components are linear Student's $t$ VARs, we have also introduced a multivariate version of the StMAR model of Meitz+Preve+Saikkonen:2023, which we call the StMVAR model. In addition to the reduced form model, we introduced a structural version of the G-StMVAR model with statistically identified shocks. The structural G-StMVAR model identifies the shocks by heteroskedasticity similarly to the SVAR model of Virolainen:2024, but our model is nonlinear and accommodates asymmetries in the (generalized) impulse response functions. The introduced methods are implemented to the CRAN distributed R package gmvarkit gmvarkit that provides a comprehensive set of tools for maximum likelihood estimation and other numerical analysis of the introduced models.

Our empirical application studied the effects of the Euro area monetary policy shock in a monthly data consisting of four macroeconomic variables. We fitted a two-regime StMVAR to the data and found that Regime 1 is characterized by a negative (but volatile) output gap, and it mainly prevails after the Financial crisis, whereas Regime 2 is characterized by a positive output gap and it mainly dominates before the Financial crisis. We found that the effects of the monetary policy shock are stronger in Regime 2 than in Regime 1. In both regimes, a contractionary monetary policy shock decreases the output gap significantly and persistently. However, while the price level decreases significantly and permanently in Regime 2, the decrease is insignificant in Regime 1. Asymmetries with respect to the sign and size of the shock were found weak in both regimes.