Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,479 characters · 17 sections · 85 citation commands
Theory coherent shrinkage of Time-Varying Parameters in VARs
Over the past four decades vector autoregressive models have become the leading tool for description, forecasting, structural inference and policy analysis of macroeconomic data sims1980,SW2012. A natural progression in the literature was to allow for time-varying parameters to capture changes in the complex dynamic interrelationship among the variables in the system CS2002,primiceri2005, COGLEY2005262. On the one side, this class of models known as Time Varying Parameters VARs (TVP-VARs) can be flexible enough to fit many different forms of structural instabilities and evolving nonlinear relationships among the macroeconomic variables. On the other side, due to the growing number of parameters, the model can become overly parameterized, which may negatively affect the precision of inference for typical objects of interest, such as impulse response functions, and reduce the reliability of the forecasts.
In this paper, I propose using economic theory to sharpen inference in TVP-VARs. The approach consists in leveraging on prior information coming from an underlying economic theory regarding the macroeconomic variables in the system. This prior information is incorporated into a shrinkage prior for the time-varying parameters, guiding the time-varying coefficients towards a predetermined path implied by the chosen economic theory. The resulting model, equipped with this prior, which I label Theory Coherent TVP-VAR (TC-TVP-VAR), remains a flexible statistical model for the data that exploits economic theory to enhance inference about the time varying parameters. The model features two crucial hyper-parameters governing the behavior of the time varying coefficients: the first one determines their intrinsic time variation, while the other one determines their degree of theory coherence stemming from the amount of shrinkage towards the path for those coefficients dictated by the economic theory. Both the optimal degree of time variation and of theory coherence of the time varying parameters can be optimally tuned by maximizing the marginal data density of the model which is available in closed form. The TC-TVP-VAR can leverage prior information from economic theories that imply both constant and time varying paths for the time varying parameters, thereby extending the DSGE-VAR framework developed in DNS2004 to a time-varying framework. The TC-TVP-VAR can also serve as a tool for estimating the deep parameters from the underlying economic theory. These parameters are as another set of hyper-parameters in the model and are indirectly estimated by mapping the TVP-VAR estimates onto the theoretical restrictions imposed by the structural model.
In the paper I show that incorporating information from conventional economic theory into a prior for the coefficients of TVP-VARs can be beneficial to improve forecast accuracy and to obtain more precise estimates of typical objects of interest such as the impulse response functions. In particular, I find that exploiting a basic three equations New-Keynesian model to form a prior for the parameters of a trivariate TVP-VAR for output growth, inflation rate and the interest rate improves both point and density forecast accuracy of both output growth and the inflation rate at all the horizons considered (one quarter ahead, two quarters ahead and one year ahead). Then, I exploit the TC-TVP-VAR to investigate changes in the propagation of risk premium shocks inside and outside the Zero lower Bound (ZLB) period in the US economy. According to a standard NK, the economy is expected to exhibit a different response to demand and supply shocks when the ZLB constraint is in effect. However, and more importantly, the short length of the ZLB period in the US makes the standard TVP-VAR unfit to detect the change in the responses predicted by the NK model lubikbenati. In other words, whether or not there was a change in the propagation of risk premium shocks during the ZLB as predicted by a standard NK model, cannot be directly inferred by using a standard TVP-VAR. Based on a simulation study, I show that the TC-TVP-VAR can be used to solve this inferential problem. In particular, I exploit the time varying restriction functions implied by a medium scale NK model that accounts for forward guidance and the ZLB period to parametrize my prior in the TC-TVP-VAR. I show that this approach allows to estimate more precisely the response of the economy to macroeconomic shocks inside and outside the ZLB period, solving the inferential problems of the standard TVP-VAR. Estimating the model on US data, I find that there are convincing evidences supporting a different propagation of risk premium shocks inside the ZLB period similar to the one predicted by a standard NK model. This finding has clearly important policy implications for the conduct of fiscal and macroprudential policies at the ZLB. \\
Related Literature This paper shows how to exploit prior information grounded on the basis of an economic theory to impose parsimony on the coefficients of TVP-VARs. In this aspect, the contribution conceptually borrows from ideas from the seminal work of WHITEMAN and operationally from the insights in DNS2004 which show how to exploit the non-linear cross equation restrictions implied by a DSGE to form a prior for the parameters of a constant parameters VAR model. Extending the framework of DNS2004 to TVP-VARs is important at least for two reasons. First, because in macroeconomic applications the assumption of constant coefficients is often restrictive. Indeed, instabilities in the autoregressive coefficients of VARs used to model the dynamics of key macroeconomic indicators such as output and inflation have been widely documented in the literature CS2002,primiceri2005,giannonegambetti. In this setting, as the model becomes more flexible, additional shrinkage can be particularly beneficial to reduce overfitting. Second, because economic theories themselves might imply time-varying paths for the coefficients. For example, macroeconomic theories assuming rational expectations extended so as to allow some parameters to vary according to Markov process with given transition probabilities, lead to state space representation with time varying coefficients FARMER20091849. At the same time, solutions for linear stochastic rational expectations models in the face of a finite sequence of anticipated structural changes lead to state space representation with time varying coefficients kulishcagliarini.\footnote{Another notable case within the rational expectations framework are models log-linearized around time varying trend for inflation cogleysbordone,sbordoneascari,ascari2023long.} Likewise, outside the rational expectation framework, macroeconomic theories that assume learning can lead to state space representation with time varying coefficients MILANI20072065. In all those cases, the proposed TC-TVP-VAR can be used both to incorporate the implied time varying restriction functions into a prior for the time varying coefficients of the VAR and to estimate the deep parameters of the underlying economic model. In this sense, the paper is also related to the strand of literature that exploits an auxiliary flexible statistical model for the data to make indirect inference on the deep parameters of a structural model from the economic theory (see for example gallant kasyfessler2019).\footnote{In the TC-TVP-VAR the structural parameters from the underlying economic theory are estimated by implicitly minimizing the weighted discrepancy between the unrestricted TVP-VAR estimates and the restriction functions. This approach can be thought as a Bayesian version of smith1993 as in DNS2004.}
The paper is also related to the increasing number of studies that has recently focused on the issue of mitigating complexity and over-parametrization in TVP-VARs. One strand of literature has focused on identifying fixed versus time varying coefficients, by concentrating on the variance selection problem in the generic state equations of each of the TVP-VARs’ coefficients. FRUHWIRTHSCHNATTER201085,belmontekoopkorobilis,kalli,BITTO201975,onorante. This literature has produced shrinkage priors for variances aimed at “automatically” reducing time-varying coefficients to static ones if the model is overfitting.\footnote{Similarly, from a frequentist perspective, coulombe2021timevarying showed that time varying parameters can be framed as ridge regressions problems and used cross validation to tune the optimal amount of time variation in each of the state equations of the coefficients of the TVP-VAR.} While treating the coefficients of the model as independent stochastic processes and just focusing on the problem of tuning the optimal amount of time variation of the single coefficients, this strand of literature typically entirely neglects co-movement and correlation among the coefficients. However, in macroeconomic applications, the high degree of co-movement in the parameters is an empirical regularity. This fact was already found and stressed by COGLEY2005262 in one of the papers that introduced TVP-VARs in the field. In the same paper, the authors envisaged that the reduced-form parameters should move in a highly structured way because of the cross-equation restrictions suggesting that “a formal treatment of cross-equation restrictions with parameter drift is a priority for future work” (p. 274). More in line with these considerations, a smaller number of studies gambettidewind,stevanovic,CHAN2020105 proposed to use a factor structure to model the time variation of the parameters. Despite being compatible with the idea of the coefficients varying in a highly structured way because of the cross-equation restrictions associated with macroeconomic equations, this approach is purely statistical and abstracts from any macroeconomic theory disciplining the behavior of the coefficients.\footnote{An exception to this literature is the TVP-VAR in leon where the identified structural innovations are allowed to influence the dynamics of the coefficients.} This paper fills this gap in the literature by showing how to exploit prior information grounded on the basis of an economic theory to state a priori a plausible correlation structure among the time varying parameters of the model. While focusing on economic theory as a source of potentially useful information on the time varying parameters, the paper methodologically contributes to literature on priors for TVP-VARs, by specifying a joint shrinkage prior for the entire path of the time varying parameters. This approach allows to model the shape of the path the time varying coefficients are expected to follow. Operationally this is done by writing the TVP-VAR in static compact form and exploiting band matrices to specify a joint prior for the time varying parameters.
Finally, the paper is more broadly related to the econometric literature showing that moment conditions from economic theory can successfully be exploited for forecasting macroeconomic and financial variables (GIACOMINI2014145,ccm most notably).\\
Outline The paper is organized as follows. In Section (ref) I introduce the TC-TVP-VAR, presenting an analytical derivation of the proposed theory coherent prior and a simulation from the prior to showcase its main properties. In the section, I discuss the fit-complexity trade-off linked to the calibration of the hyper-parameters determining the intrinsic time variation of the time varying-coefficients and the degree of shrinkage towards the restriction functions from the economic theory. I conclude section (ref) by presenting an MCMC sampler used for estimation of the TC-TVP-VAR. In section (ref) I also discuss how to exploit the proposed prior in TVP-VARs featuring stochastic volatility. Then, in section (ref) I exploit the well known 3-equations New-Keynesian (NK) block to form a prior for the parameters of a TVP-VAR for GDP growth, inflation and the interest rate for the US economy. I compare the forecasts from a TC-TVP-VAR that encodes the restrictions from the NK model as a prior for the time varying coefficients to the forecasts from a standard TVP-VAR, showing that the former outperforms the latter both in terms of point and density forecast accuracy. Then, in section (ref) I focus on impulse response analysis at the ZLB. I conduct a simulation study and show that a standard TVP-VAR struggles to detect the change in the response of the economy to macroeconomic shocks during the ZLB period as predicted by a standard NK model. I show that the TC-TVP-VAR can be used to solve this inferential problem. Finally, I use the TC-TVP-VAR to investigate the propagation of risk premium shocks inside and outside the ZLB period. Section (ref) concludes.
Notation Before moving on, I introduce some notations conventions used in the paper. Scalars are in lowercase and normal weight. Vectors are in lowercase and in bold. Matrices are in uppercase and bold.
A TVP-VAR is given by:
where $\boldsymbol{y_t}$ is an $N$ dimensional random vector observed for $t = 1, \ldots,T$ periods, $\boldsymbol{x_t} = [1,\boldsymbol{y}_{t-1}', \ldots, \boldsymbol{y}_{t-p}']'$, $p$ is the lag order and $k = 1+Np$ is the number of coefficients in each equation of the VAR. For convenience, we can rewrite the TVP-VAR in ((ref)) in “static” compact form as
where
For now it is assumed a constant error covariance matrix $\boldsymbol{\Sigma_u}$, while in section (ref) the model is extended to the case of a time-varying error covariance. The likelihood implied by ((ref)) is
Note that in this representation, the matrix $\boldsymbol{\Phi}$ stores the matrices with the time varying coefficients one on the top of the other, for all the time periods $t=1, \ldots, T$ (see (ref)). This paper introduces a prior for $(\boldsymbol{\Phi},\boldsymbol{\Sigma_u})$, with its final form presented in equations ((ref)). This prior centers the time-varying coefficients in $\boldsymbol{\Phi}$ on a path predefined by an underlying economic theory. The aim of this section is to guide the reader through the construction of this prior, which begins by considering a simple dynamic model for the time-varying coefficients, namely
This dynamics can be rewritten in compact form for all the coefficients as
which graphically is
For now, I think of just an hyperparameter $\lambda$ affecting the variance of the innovations in the random walk dynamics of all the elements in $\boldsymbol{\Phi_1}, \ldots, \boldsymbol{\Phi_T}$. In particular, I assume that the variance covariance matrix $\boldsymbol{\Omega}$ has the Kronecker structure $\boldsymbol{\Omega} = \boldsymbol{\Sigma_u} \otimes ( \lambda^2 \boldsymbol{I_k})^{-1}$. This assumption implies that the variance in the state equation of the coefficients is proportional to the variance of the innovations of the equation to which the coefficients appertain $\sigma^2_{ii}$. In practice this means that the generic state equation of the coefficient attached to the $j^{th}$ regressor in the $i^{th}$ equation reads as follows
with $i=1, \ldots, N$ and $j=1, \ldots, k$. As $\lambda \rightarrow 0$ the variance of the Normal prior centering a coefficient at time $t$ on the realization at time $t-1$ increases, meaning that this prior becomes more and more diffuse and the coefficients of two consecutive time periods are not forced a priori to be close to each other. Conversely as $\lambda \rightarrow \infty$ the prior centering the coefficient at time $t$ on the realization at time $t-1$ becomes more and more tight. Clearly the calibration of $\lambda$ involves a fit complexity trade off since by letting $\lambda \rightarrow \infty$ we allow the coefficient at time $t$ to be arbitrary distant from the coefficient at $t-1$ up to the point that we can almost fit the data perfectly in sample.\footnote{This point will be covered more in detail in Section (ref), through the lenses of the marginal likelihood of the model.} The assumption of a unique hyperparameter $\lambda$ determining the amount of time variation of the time varying coefficients can be easily replaced by assuming $\lambda_j$'s hyper-parameters with $j = 1, \ldots, k$ as it will be clarified in section (ref). Exploiting representation ((ref)) and the Kronecker structure of $\boldsymbol{\Omega}$, we can conclude that equation ((ref)) implies a joint normal distribution for $\boldsymbol{\Phi}$, that is
since $|\boldsymbol{H}_{Tk}| = 1$, $\boldsymbol{H}_{Tk}$ is always invertible and the determinant of Jacobian matrix of the linear transformation cancels out.\footnote{Notice that this is true also if we consider an AR(1) dynamics for the time varying coefficients. In that case the elements off the main diagonal of the band matrix $\boldsymbol{H_{Tk}}$ would store the AR coefficients of the state equations.} Combining this conditional distribution for the time varying parameters with an Inverse-Wishart distribution for $\boldsymbol{\Sigma}_u$
leads to
Equation ((ref)) states a prior for the time varying parameters and the variance covariance matrix of the innovations of the TVP-VAR which is conditional on $\lambda$, the crucial hyper-parameter determining the amount of time variation of the coefficients.\footnote{To be precise this prior is also conditional on $\boldsymbol{\Phi_0}$, $\boldsymbol{\underline{S}}$ and $\underline{\nu}$. However, for the purpose of the paper I will consider $\boldsymbol{\Phi_0}$, $\boldsymbol{\underline{S}}$ and $\underline{\nu}$ as fixed.} Importantly, this prior is conjugate to the Gaussian likelihood. We exploit the conjugacy to update equation ((ref)) with the likelihood of some artificial observations from a structural model coming from an underlying economic theory. As a resulting distribution, we obtain a Normal-Inverse-Wishart distribution which is now theory coherent, meaning that it is centered on the moment restrictions from the economic theory. \footnote{Note that also assuming $p(\boldsymbol{\Sigma_u}) \propto |\boldsymbol{\Sigma_u}|^{-\frac{N +1}{2}}$ would lead to a Normal-Inverse Wishart distribution.} In the specific, updating ((ref)) with observations from the theory we get
where $p(\boldsymbol{Y(\theta)}| \boldsymbol{\Phi}, \boldsymbol{\Sigma_u}, \gamma) $ is the likelihood of $\gamma$ simulated samples of $\boldsymbol{Y(\theta)}$ from the theory and $\boldsymbol{\theta}$ are the deep parameters from the economic theory while
is an integrating constant which ensures that the prior density is proper and integrates to one.\footnote{In the appendix (ref) I report the details on the integrating constant.} Equation ((ref)) makes clear that the theory-coherent prior distribution for $(\boldsymbol{\Phi}, \boldsymbol{\Sigma_u})$ is just the posterior distribution of $(\boldsymbol{\Phi}, \boldsymbol{\Sigma_u})$ obtained by updating the hierarchical prior $p(\boldsymbol{\Phi, \Sigma_u}|\lambda)$ with the likelihood of $\gamma$ artificial samples of observations generated by the model from the theory. As in DNS2004, in $ p(\boldsymbol{Y(\theta)}| \boldsymbol{\Phi}, \boldsymbol{\Sigma_u})$ I replace the sample moments $\boldsymbol{Y(\theta)'Y(\theta)}$, $\boldsymbol{Y(\theta)'X(\theta)}$, and $\boldsymbol{X(\theta)'X(\theta)}$ by their expected values $\boldsymbol{\Gamma_{yy}(\theta)} \equiv \operatorname*{\mathbb{E}}[\boldsymbol{Y'Y}|\boldsymbol{\theta}]$ , $\boldsymbol{\Gamma_{xy}(\theta)} \equiv \operatorname*{\mathbb{E}}[\boldsymbol{X'Y}|\boldsymbol{\theta}]$ , $\boldsymbol{\Gamma_{xx}(\theta)} \equiv \operatorname*{\mathbb{E}}[\boldsymbol{X'X}|\boldsymbol{\theta}]$. Note that since they relate to $\boldsymbol{Y}$ and $\boldsymbol{X}$ in the static representation of the TVP-VAR in ((ref)), these moments matrices are storing the moments $\operatorname*{\mathbb{E}}[\boldsymbol{y}_t\boldsymbol{y}_t'|\boldsymbol{\theta}]$, $\operatorname*{\mathbb{E}}[\boldsymbol{y}_t\boldsymbol{x}_t'|\boldsymbol{\theta}]$, $\operatorname*{\mathbb{E}}[\boldsymbol{x}_t\boldsymbol{x}_t'|\boldsymbol{\theta}]$, for $t=1,\ldots,T$. \footnote{The structure of these matrices is provided in the Appendix (ref).} These theory implied population moments are assumed to be available in closed form as a function of the deep parameters from the economic theory $\boldsymbol{\theta}$ and in most of the cases can be derived from the state space representation of the structural model from the economic theory. This assumption let us write the likelihood of the artificial observations directly as a function of the vector of deep parameters from the economic theory $\boldsymbol{\theta}$. To sum up, the Normal-Inverse-Wishart distribution in equation ((ref)) is de-facto obtained as the posterior distribution that combines a Normal-Inverse Wishart prior for $(\boldsymbol{\Phi},\boldsymbol{\Sigma_u})$ with a Gaussian likelihood for the artificial data from the theory. Therefore the theory coherent prior in equation ((ref)) takes the form
This prior for $\boldsymbol{\Phi}$ encompasses two different pieces of information about the time varying coefficients. The first piece of information is their intrinsic time variation determined by the shrinkage hyperparameter $\lambda$. The second, is instead the degree of theory coherence implied by the shrinkage hyperparameter $\gamma$ which defines the tightness of the Normal prior around the restriction function defined by the economic theory
This restriction function implies that at each time period $t$ the matrix of coefficients is centered on
where $\boldsymbol{\Gamma_{xx,t}(\theta)} \equiv \operatorname*{\mathbb{E}} [\boldsymbol{x_tx_t'}|\boldsymbol{\theta}]$ and $\boldsymbol{\Gamma_{xy,t}(\theta)} \equiv \operatorname*{\mathbb{E}} [\boldsymbol{x_ty_t'}|\boldsymbol{\theta}]$ are respectively of dimension $k \times k$ and $k \times N$. Equation ((ref)) shows that the prior is centering the time varying coefficients on $\boldsymbol{\Phi_t(\theta)} ^*$, which is a local OLS formula evaluated at the theory implied population moments. In general, these population moments are assumed to be available in closed form as function of the deep parameters from the economic theory $\boldsymbol{\theta}$. However, in the case they are not directly available in closed form from the chosen economic theory, they can easily be obtained by stochastic simulation as in LORIA2022105. Importantly, as the moments stored in the matrices $\boldsymbol{\Gamma_{xx}}$, $\boldsymbol{\Gamma_{xy}}$, $\boldsymbol{\Gamma_{yy}}$ can be either constant or time varying conditionally on the deep parameters from the theory, ((ref)) can define both constant and time varying paths for the time varying coefficients. Finally, it is worth to mention that this structure of the prior can also be used more in general to model any a-priori beliefs about the evolution of the time-varying coefficients over time, even in the case in which these beliefs are not coming from an explicit economic theory.
To showcase the properties of the proposed theory coherent prior, figure (ref) shows random draws for the time series of a generic $\phi_{jt}^{(i)}$ coefficient of the TVP-VAR for $t=1, \ldots, 250$ from the prior. The dashed black line represents the path i.e. the restriction function $(\ref{restriction_function})$ defined by the economic theory for that coefficient, that I label $\phi_{jt}^{(i)}(\boldsymbol{\theta})^*$. Just for now, for illustrative purposes I consider the case of a constant restriction function, however having time varying moments stored in the matrices $\boldsymbol{\Gamma_{xx}}$ and $\boldsymbol{\Gamma_{xy}}$, would directly imply a time varying restriction function through ((ref)). The first row shows five draws from the theory coherent prior for different values of $\lambda$ conditioning on a positive $\gamma$, meaning that for a given degree of theory coherence I change the value of hyper-parameter which determines the intrinsic time variation of the coefficient. When $\lambda = 0$, the draws for the coefficient $\phi_{jt}^{(i)}$ are independent for $t = 1, \ldots, T$ and centered on the restriction function coming from the theory. Hence, nothing prevents a coefficient at time $t$ to be arbitrary distant from a coefficient at time $t-1$ except the fact both coefficients are centered, with a precision determined by $\gamma$, on the restriction function defined by the theory. As $\lambda$ increases the time series of the coefficient becomes less and less time-varying up to the point that when $\lambda$ is very big, the coefficient becomes almost constant over time. Instead, in the second row, I let the degree of theory coherence change conditionally on a given intrinsic time variation of the time varying coefficients. In other words, I show random draws for $\phi_{jt}^{(i)}$ for different values of $\gamma$ conditionally on $\lambda > 0$. When $\gamma = 0$, the coefficients are just random walks with variance equal to $\frac{\sigma_{ii}}{\lambda^2}$, and they are completely unrelated to the restriction function implied by the economic theory $\phi_{jt}^{(i)}(\boldsymbol{\theta})^*$.\footnote{Note that the draws fore the time varying coefficients were initiated near the restriction functions just for visualization purposes.} As $\gamma$ increases the draws for the coefficient are centered on the restriction function defined by the theory with an increasing precision and on the limit, with $\gamma\rightarrow \infty$, they go to $\phi_{jt}^{(i)}(\boldsymbol{\theta})^*$.\footnote{The coefficients would have a degenerate distribution, with point mass on the restriction function.} Importantly, the hyperparameter $\gamma$ is contemporaneously shrinking all the time varying coefficients $\phi_{t,j}^{(i)}$ in the equations of the VAR towards the restrictions implied by the economic theory. This reflects the idea that it is economic theory that a priori postulates a plausible correlation structure among the coefficients of VAR model through the restriction function in ((ref)).
It is important to remark that this framework can accommodate both constant and time-varying restrictions for the coefficients through ((ref)). This allows to extend the DSGE-VAR framework developed in DNS2004 to a broader set of economic theories, which might imply time varying restriction functions for the parameters of the VAR. To show this property, I exploit the medium scale New-Keynesian model in delnegroschorfheide2005 to parametrize a prior for a medium scale TVP-VAR for output growth, consumption growth, investment growth, real wage growth, hours worked, inflation and the Fed Fund rate. The NK model accounts for the ZLB and forward guidance and it is solved using the method proposed by kulishcagliarini for linear rational equation expectation systems in the face of anticipated structural changes. The solution of the model implies a state space representation that exhibits time varying coefficients. Once the population moments implied by the state space representation are used to parameterize ((ref)) we obtain time varying restrictions functions for the time varying coefficients.\footnote{The details on the derivation of the moments are available in the section (ref).} Figure (ref) makes this point by showing the implied path for the coefficient of the first lag of the inflation rate in the equation of the Fed Fund Rate. Together with the time varying restriction function, the figure shows draws from the prior for different values of the hyper-parameter $\gamma$ conditioning on a given value of $\lambda$ and a set of structural parameters of the NK model. More in the specific, while outside the zero lower bound the Fed Fund rate is expected to increase when the lagged value of inflation increases, inside the zero lower bound the interest rate is not expected to respond to the lagged value of the inflation rate and therefore the prior mean for this coefficients becomes centered around zero. As expected, as $\gamma$ increases the draws for the time varying coefficient are centered with increasing precision on the path defined by the economic theory.
It is straightforward to derive the posterior distribution of $(\boldsymbol{\Phi},\boldsymbol{\Sigma}_u)$ conditional on the hyper-parameters $\lambda$, $\gamma$ and on the deep parameters from the theory $\boldsymbol{\theta}$ since this is just a Normal-Inverse-Wishart distribution obtained by updating a Gaussian likelihood with the Normal-Inverse-Wishart prior ((ref)), namely
Equation ((ref)) makes it clear that thanks to the conjugacy of the prior the formula for the mean of the conditional posterior of $\boldsymbol{\Phi}$ translates in a standard OLS regression formula based on a sample augmented with a set of dummy observations that determines the degree of time variation of the time varying coefficients and another set of dummy observations that centers the coefficients on the restriction function implied by the theory. While the first set of dummy observations shapes the correlation of the time varying coefficients over time by imposing a set of linear restrictions as a function of the hyper-parameter $\lambda$, the second set of dummy observations induces at each time period a correlation structure implied by the non-linear cross equation restrictions coming from the theory, as a function of the deep parameters from the theory $\boldsymbol{\theta}$ and of the shrinking parameter $\gamma$.\footnote{In the appendix (ref) I show that thanks to the Kronecker structure of the prior ((ref)) the time variation of the coefficients can be modelled by dummy observations.} It is worth to remark that from a computational perspective equation ((ref)) makes the estimation of the time varying coefficients very efficient since it allows to draw all the history of the time varying coefficients in all the equations of the VAR in a single step, avoiding forward filtering and backward smoothing algorithms à la kc1994. \footnote{Indeed, as in the precision sampler by chan2009efficient, we can draw all the latent states from $t=1, \ldots, T$ in a single step and thanks to the Kronecker structure of the posterior, we can do it for all the $N$ equations of the TVP-VAR jointly. More in general, the Kronecker structure ((ref)) coupled with the precision sampler by chan2009efficient can be exploited to estimate medium to large scale TVP-VARs. }
Tuning of the optimal degree of theory coherence and the intrinsic amount of time variation of the coefficients is a delicate matter, since it clearly involves a fit-complexity trade off. As a matter of fact, very low values of both $\gamma$ and $\lambda$ imply a priori that the coefficients in two consecutive time periods can potentially be very distant from each others and from the restriction functions defined by the theory.\footnote{When both $\gamma=0$ and $\lambda=0$ the model is left totally unrestricted, with the prior variance covariance of the coefficients being equal to infinity. Clearly, in this case there are more parameters than you can feasibly estimate with a flat prior meaning that the conditional posterior of $\boldsymbol{\Phi}$ cannot be computed due to the non-invertibility of $\boldsymbol{X'X}$ (this can be seen from equation ((ref))).} Intuitively, this model will fit the data very well in sample but will perform badly for forecasting out-of-sample. Indeed, decreasing $\gamma$ and $\lambda$ will in general increase in-sample fit of the model at the expense of out-of-sample accuracy. Based on this argument, I recommend to base the optimal choice of both the hyper-parameters on the maximization of the marginal likelihood of the model or equivalently on the maximization of the posterior of the hyper-parameters $\lambda$ and $\gamma$ under a flat prior for these hyper-parameters. This translates into maximizing the one-step-ahead out-of-sample forecasting ability of the model. Indeed, the log-marginal likelihood (or Bayesian evidence) can be interpreted as the sum of the one step-ahead predictive scores, since it is equivalent to the scoring rule of the form
An attractive feature of the TC-TVP-VAR is that the marginal likelihood $p(\boldsymbol{Y}|\gamma, \lambda, \boldsymbol{\theta})$ obtained by integrating out $\boldsymbol{\Phi},\boldsymbol{\Sigma}_u $ from the conditional posterior $p(\boldsymbol{\Phi, \Sigma_u} |\lambda,\boldsymbol{\theta}, \gamma,\boldsymbol{Y}) $ is available in closed form and it is equal to
As a consequence, calibrating $\gamma$ and $\lambda$ to maximize ((ref)) corresponds to finding $\gamma$ and $\lambda$ maximizing the one-step-ahead out-of-sample forecasting ability of the model. This strategy of estimating hyper-parameters by maximizing the marginal likelihood is an empirical Bayes method which has a clear frequentist interpretation. In what follows, and in particular in the estimation algorithm detailed in the next section (ref) I will regard $\gamma$ and $\lambda$ as random variables and perform full posterior inference on the hyper-parameters, but analogously maximizing the posterior of the hyper-parameters will correspond to maximizing the one-step-ahead out of sample forecasting ability of the model. Following the same steps as in giannonelenzaprimicieri, we can rewrite equation ((ref)) as
where $\boldsymbol{V_{\varepsilon}^{post}}$ and $\boldsymbol{V^{prior}_{\varepsilon}}$ are the posterior and prior means (or modes) of the residual variance, while $\boldsymbol{V_{t|t-1}}$ is equal to the variance (conditional on $\boldsymbol{\Sigma_u}$) of the one-step-ahead forecast of $\boldsymbol{y}$, averaged across all possible a-priori realizations of $\boldsymbol{\Sigma_u}$. Complete formulas of these objects are reported in the appendix (ref) and are the analog of the formulas reported in giannonelenzaprimicieri. Equation ((ref)) makes clear that the marginal likelihood involves two terms: a reward for model fit, $| \boldsymbol{(V_{\varepsilon}^{post})^{-1}V^{prior}_{\varepsilon}}|^{\frac{T + \underline{\underline{\nu}}}{2}}$ and a penalty term for model complexity $\prod_{t=1}^{T}|\boldsymbol{V_{t|t-1}}|^{-\frac{1}{2}}$. Figure (ref) plots the model fit term and the penalty term of the marginal likelihood as a function of the two hyper-parameters $\gamma$ and $\lambda$ conditionally on a set of deep parameters from the theory $\boldsymbol{\theta}$. As $\lambda$ decreases the tightness of the prior centering each coefficient on its realization in the previous period eases and in sample model fit improves, meaning that $\boldsymbol{V}^{post}$ decreases. However, at the same time, as $\lambda$ decreases $\boldsymbol{V_{t|t-1}}$ increases, since the variance (conditional on $\boldsymbol{\Sigma}$) of the one-step-ahead forecast of $\boldsymbol{y}$ increases. Analogously, as $\gamma$ decreases the restriction functions from the theory become less and less binding, meaning that the model will fit better in sample, namely $\boldsymbol{V}^{post}$ decreases, but the variance of the one step ahead forecast error $\boldsymbol{V_{t|t-1}}$ will increase.
In this section I describe a posterior simulation sampler to make inference on the TVP-VAR parameters $\boldsymbol{\Phi}$ and $ \boldsymbol{\Sigma}_u$ and on the shrinkage hyper-parameters $\lambda$ and $\gamma$ together with the deep parameters from the theory $\boldsymbol{\theta}$ .\footnote{For the purpose of the paper we fix $\boldsymbol{\Phi}_0 = \boldsymbol{\hat{\Phi}_0}$ and $\boldsymbol{\underline{S}} = \boldsymbol{\hat{S}} = diag(\hat{s}_1^2, \ldots, \hat{s}_N^2)$ and $\nu = N+2$. Alternatively inference on these hyper-parameters can be made by extending the Random Walk Metropolis step in the MCMC sampler described below.} We can simulate draws from the posterior $p(\boldsymbol{\Phi, \Sigma_u}, \lambda, \boldsymbol{\theta}, \gamma|\boldsymbol{Y})$ exploiting the factorization:
This factorization makes clear that we can sample from the posterior of $\boldsymbol{\Phi}, \boldsymbol{\Sigma_u}, \boldsymbol{\theta}, \gamma, \lambda$ building a MCMC algorithm that iterates the following two steps:
In the first step, estimation of both the shrinkage parameters $\gamma$ and $\lambda$ and of deep parameters from the theory $\boldsymbol{\theta}$ happens by posterior evaluation of
where $p(\gamma)$ and $p(\lambda)$ are the priors for the shrinkage hyper-parameters \footnote{For both $\gamma$ and $\lambda$ I consider the Uniform prior $\gamma \sim \mathcal{U} (0, b_{\gamma}) $ and $\lambda \sim \mathcal{U} (0, b_{\lambda})$ where $b_{\gamma} = 10^{10}$ and $b_{\lambda} = 10^{10}$} while $p(\boldsymbol{\theta})$ is the prior of the deep parameters from the theory. Hence, learning about the structural parameters happens implicitly by projecting the VAR estimates onto the restrictions implied by the model from the theory. More precisely, the estimates of the deep parameters minimizes the weighted discrepancy between the TVP-VAR unrestricted estimates and the restriction function ((ref)). This approach can be thought as a Bayesian version of smith1993 and was pioneered by DNS2004. In particular, the TVP-VAR is used to summarize the statistical properties of both the observed data and the theory-simulated data and an estimate of the deep parameters from the theory is obtained by matching as close as possible TVP-VAR parameters from observed data and from the simulated data.\footnote{See Proposition 1 and Proposition 2 in DNS2004 for further details on this. } To sample from $p(\lambda, \boldsymbol{\theta}, \gamma|\boldsymbol{Y})$ I consider a two blocks random walk metropolis algorithm. This is basically a Gibbs Sampler where in the first block, I draw the hyper-parameters $\gamma$ and $\lambda$ conditionally on $\boldsymbol{\theta}$ while in the second block I draw $\boldsymbol{\theta}$ conditionally on the shrinking hyper-parameters $\gamma$ and $\lambda$. Step 2, instead, consists just on Monte Carlo draws from the posterior of $(\boldsymbol{\Sigma_u}, \boldsymbol{\Phi})$ conditional on $\lambda,\gamma$ and $ \boldsymbol{\theta}$ which is the Normal-Inverse-Wishart distribution in ((ref)) that is:
In this approach, the time varying parameters and the variance covariance matrix are drawn in a single step from their Normal-Inverse-Wishart conditional posterior distribution. As a last note, to let the estimated path of time varying coefficients $\boldsymbol{\Phi_1}, \ldots,\boldsymbol{\Phi_T}$ and the estimate of the hyper-parameter $\lambda$ less affected by the choice of the initial condition $\boldsymbol{\Phi_0}$ I assume that in equation ((ref)) $\boldsymbol{\xi_1} \sim \mathcal{N}(\boldsymbol{0}, \boldsymbol{\Sigma_u} \otimes 5 (\sum_{t=1}^{T_{pre}}\boldsymbol{x}_t\boldsymbol{x}_t')^{-1} ) $ such that $\boldsymbol{\Phi_1} \sim \mathcal{N}(\boldsymbol{\hat{\Phi}_0}, \boldsymbol{\Sigma_u} \otimes 5 (\sum_{t=1}^{T_{pre}}\boldsymbol{x}_t\boldsymbol{x}_t')^{-1} ) $ where $T_{pre}$ is the size of a pre-sample of observations while $\boldsymbol{\hat{\Phi}_0}=0$.
The non-linear rational expectation equations are solved using the method based on matrix eigenvalue decomposition by sims2002 leading to a solution which has the form:
The state equation ((ref)) is complemented with the observation equations:
Equations ((ref)) ((ref)) ((ref)) ((ref)) are used to derive the population moments encoded in the TVP-VAR $\boldsymbol{\Gamma_{xx}}$, $\boldsymbol{\Gamma_{xy}}$, $\boldsymbol{\Gamma_{yy}}$ as a function of the deep DSGE parameters
Table (ref) in the appendix (ref) presents details on the priors for the DSGE hyper-parameters, based on DNS2004 while Table (ref) in the appendix (ref) reports the posterior median estimate for the deep DSGE parameters together with the $15^{th}-85^{th}$ credible sets. Figure (ref) shows the estimated time varying parameters $\boldsymbol{\Phi}$. Figure (ref), presents the impulse response functions for GDP growth, inflation and interest rate to a monetary policy shock. The monetary policy shocks are identified using the identification strategy proposed by DNS2004. The strategy consists in exploiting the QR factorization of the impact matrix $\boldsymbol{\Omega(\theta)}$ in the state space representation of the DSGE model to obtain the rotation matrix $\boldsymbol{Q(\theta)}$. Then the rotation matrix $\boldsymbol{Q(\theta)}$ is used to identify the monetary policy shocks in the Structural TVP-VAR. In the specific, considering the linear relationship between the structural shocks $\boldsymbol{\varepsilon_t}$ and the reduced form innovations $\boldsymbol{u_t}$ in equation ((ref)), namely:
$\boldsymbol{A}$ is computed as the product of the Cholesky decomposition of $\boldsymbol{\Sigma_{u}}$ and $\boldsymbol{Q} = \boldsymbol{Q(\theta)}$. \fi
For simplicity of exposition in the previous section I assumed a unique hyper-parameter i.e. $\lambda$ determining the intrinsic time variation of the coefficients. This assumption can easily be replaced by assuming $\lambda_j$'s hyper-parameters with $j = 1, \ldots, k$ such that $\boldsymbol{\Omega}$ can be factorized as
where $\boldsymbol{\Lambda_k} = diag(\lambda_1, \ldots, \lambda_k) $ and the Kronecker structure of the Normal-Inverse-Wishart prior ((ref)) is preserved. In terms of estimation, in the case we allow for regressor specific shrinkage hyper-parameters $\lambda_j$ for $j = 1, \ldots, k$ nothing changes conceptually, since the algorithm should just be adapted to draw all the $\lambda_j$'s hyper-parameters together with $\gamma$ in the second block of the random walk metropolis in Step 1. Formulas for the marginal likelihood and the conditional posterior distribution of $\boldsymbol{\Phi}$ and $\boldsymbol{\Sigma}_u$ can be found in the appendix (ref). In general, keeping the assumption of a Kronecker structure for $\boldsymbol{\Omega}$ allows to have a closed form expression for the likelihood of $\gamma, \lambda$ and $\boldsymbol{\theta}$ marginally of the time varying parameters $\boldsymbol{\Phi_1,}, \ldots, \boldsymbol{\Phi_T}$ and of $\boldsymbol{\Sigma_u}$. In practice this allows to make inference on $\boldsymbol{\Phi}, \boldsymbol{\Sigma}_u$, the shrinking hyper-parameters $\gamma$ and $\lambda$ and also on the deep parameters from the theory $\boldsymbol{\theta}$ jointly as in DNS2004. Considering a more general structure for $\boldsymbol{\Omega}$, would in general lead to a non-conjugate structure of the prior. This would complicate inference on the parameters of the model, slowing down the convergence of the algorithm and produce more correlated draws, especially when drawing the deep parameters from the economic theory $\boldsymbol{\theta}$ jointly with the parameters of the TVP-VAR.
As remarked by sims2002 commenting on CS2002, a model with only time variation in parameters could mistakenly result in a substantial amount of time variation even though the true data generating process only features stochastic volatility. In this section I show how to extend the theory coherent prior introduced in the previous section to a TVP-VAR featuring heteroskedasticity. More in the specific, to allow for time-varying error covariance, I replace the assumption $vec(\boldsymbol{U}) \sim \mathcal{N}( \boldsymbol{0}, \boldsymbol{\Sigma}_u \otimes \boldsymbol{I}_T)$ with
where $\boldsymbol{D}$ is a $T \times T$ diagonal matrix storing the time-varying volatilities, namely $\boldsymbol{D} = diag(d_{1}, \ldots, d_{T})$. This class of VARs with flexible Kronecker error covariance structure is considered in jcc_kron. The likelihood of the model becomes
This common specification for the volatility process was introduced in ccm2016. For the elements in $\boldsymbol{D}$, i.e. the time varying volatilities, I follow jcc_kron and consider the following log-normal prior
The choice of the priors on $\rho$ and $\sigma^2_d$ also follows jcc_kron with $\rho \sim \mathcal{N}(\rho_0,V_{\rho})\mathbbm{1}(|\rho|<1)$ and $\sigma^2_{\eta} \sim \mathcal{IG}(\nu_{\eta}, S_{\eta})$.\footnote{Following jcc_kron I set $\rho_0 = 0.9$, $V_{\rho}= 0.2^2$ $\nu_{\eta} = 5$ and $S_{\eta} = 0.04$.} The theory coherent prior for $(\boldsymbol{\Phi},\boldsymbol{\Sigma}_u)$ is still the Normal-Inverse-Wishart prior in equation ((ref)). In order to draw from the posterior distribution of $(\boldsymbol{\Phi,\Sigma_u},\lambda,\gamma,\boldsymbol{\theta}, \boldsymbol{D}, \sigma^2_d, \rho )$ I consider an MCMC sampler that iterates the following steps:
This sampler generalizes the algorithm in ccm to a VAR that features both time varying volatility and time varying parameters. As in the algorithm used for the estimation of the homoskedastic model, in step 1 I resort to the random walk Metropolis to draw $\lambda, \boldsymbol{\theta}$ and $\gamma$. \footnote{Note that the conditional posterior distribution of $\lambda, \boldsymbol{\theta}$ and $\gamma$ is now combining the prior distribution $p(\lambda,\boldsymbol{\theta},\gamma)$ with $p(\boldsymbol{Y}|\boldsymbol{D},\lambda,\boldsymbol{\theta},\gamma)$. The formula $p(\boldsymbol{Y}|\boldsymbol{D},\lambda,\boldsymbol{\theta},\gamma)$ is available in the appendix (ref) .} In step 2, the formulas of the conditional posteriors of $(\boldsymbol{\Phi},\boldsymbol{\Sigma_u})$ modify as follows
Finally, in step 3, to draw time varying volatilities stored in $\boldsymbol{D}$, I compute an independence Metropolis step, following ccm. Finally, the draws for $\alpha$ and $\sigma^2_{\eta}$ are standard draws respectively from the Normal and the Inverse Gamma distributions.
In this section I consider the problem of forecasting the rate of growth of GDP, the inflation rate and the Fed Fund rate using a trivariate TVP-VAR model. In the specific, I estimate a trivariate TVP-VAR for the US economy using data from 1970 up to 2019 and I compare the forecast accuracy of a standard TVP-VAR model to the forecasts from a TC-TVP-VAR.
In the TC-TVP-VAR I exploit the New-Keynesian model in DNS2004 to parametrize the theory coherent prior. The conceptual framework commonly denoted as the 3-equation New-Keynesian model constitutes the nucleus of Michael Woodford's book “Interest and Prices" woodford2003interest and underpins most of modern monetary macroeconomics models.\footnote{See DNS2004 for the more in depth details on the New-Keynesian Model.} More specifically, the structural model is composed by an IS curve ((ref)), a New-Keynesian Phillips curve ((ref)), a monetary policy rule ((ref)) and two equations that describe the dynamics for the log-deviation from the steady state of technological process ((ref)) and government spending ((ref)), namely
The population moments needed to parametrize the prior are derived from the state-space representation of the New-Keynesian model obtained by solving the system of non-linear rational expectation equations. In particular, the non-linear rational expectation equations are solved using the method based on matrix eigenvalue decomposition by sims2002 leading to a solution which has the form
that is complemented with the set of observation equations
which look like:
and relate the unobservable latent states in ((ref)) to the observed time series of output growth, inflation rate and the Fed Fund rate. Equations ((ref)) and ((ref)) are used to derive the population moments needed to parametrize the prior. The population moments $\boldsymbol{\Gamma_{xx}}(\boldsymbol{\theta})$, $\boldsymbol{\Gamma_{xy}}(\boldsymbol{\theta})$, $\boldsymbol{\Gamma_{yy}}(\boldsymbol{\theta})$ are derived conditioning on $\boldsymbol{\theta}$, the vector of the deep parameters of the NK model
Table (ref) shows the comparison of the forecasts from a standard TVP-VAR model and the TC-TVP-VAR model. The forecasting exercise is designed such that I compute the recursive one quarter, two quarters, and one year ahead forecasts starting from 1985-Q1 up to 2019-Q4.\footnote{Details on the data can be found in the appendix (ref)} To compare relative point forecast accuracy, in the table I report the Root Mean Squared Error (RMSE) while for evaluating density forecast accuracy I report the average Cumulative Ranked Probability Scores (CRPS). In the table I also include the results concerning the forecasts from a constant parameters VAR with flat prior and a constant parameters Bayesian VAR (BVAR) with Minnesota type of prior.\footnote{Appendix (ref) the reports details on the competing forecasting models.} Overall, the TC-TVP-VAR provides the most accurate point and density forecasts for both output growth and inflation rate, outperforming the standard TVP-VAR model at all the horizons considered. For forecasting output growth, the standard TVP-VAR performs poorly relative to the TC-TVP-VAR, but also to the constant parameters BVAR with Minnesota prior, suggesting that the model tends to fit some noise in the time series of output growth. In line with the previous forecasting literature, allowing for time variation of the parameters of the VAR is important for obtaining accurate forecasts of the inflation rate, as the standard TVP-VAR outperforms both the VAR with flat prior and the BVAR with the Minnesota prior. However, economic shrinkage is helpful to obtain more reliable point and density forecasts when modelling the time variation of the coefficients. Indeed the TC-TVP-VAR outperforms the standard TVP-VAR at all the horizons. As a caveat, the standard Minnesota prior centering the autoregressive coefficients on a random walk process outperforms the other competitors, including the TC-TVP-VAR, for forecasting the Fed Fund Rate. This result is consistent with the results of the forecasting exercise in DNS2004 which use the same small scale NK model to parametrize a prior for a constant parameters VAR. Table (ref) presents the same metrics, comparing forecasts from a standard TVP-VAR model with stochastic volatility, as described in Section (ref), to those from the theory-coherent TVP-VAR model with stochastic volatility. In terms of both point and density forecast accuracy, and across all horizons, the theory-coherent model consistently delivers more accurate predictions of GDP growth and inflation.
TVP-VARs are extensively used in applied research not only to make forecasts, but also to infer changes in the response of the economy to macroeconomic shocks. In this section, I show that the proposed shrinkage prior can be useful also to enhance inference on the impulse response functions estimated from a TVP-VAR. Recent studies in empirical macroeconomics have used TVP-VARs to assess whether the US economy's performance was affected by a binding ZLB constraint galigambetti,lubikbenati. As a matter of fact, according to a standard New-Keynesian model, the economy is expected to exhibit different responses when the ZLB constraint is in effect. For example, the model predicts a distinct response of output and inflation following both demand and supply shocks when the conventional stabilizing monetary policy response to aggregate shocks is constrained as a consequence of the Federal Funds Rate hitting the ZLB. Figure (ref) makes this point, by showing the responses to a pure demand shock - the risk premium shock \footnote{The risk premium shock in SWauters is an exogenous term which affects the intertemporal margins entering both the consumption and investment Euler equation. In contrast to a discount factor shock, the risk premium shock helps explaining the co-movement of consumption and investment. More details on the risk premium shocks are presented in the Appendix (ref).} - in the SWauters model version considered in delnegroschorfheide2005.\footnote{In the model, the ZLB period is treated as in delnegroschorfheide2005, and more details on the model will be explained in the next subsection.} The figure plots the cumulative response of output growth and the response of the inflation rate to an exogenous risk premium shock that reduces households' required return of assets, decreasing firms' cost of capital. As expected, when monetary policy is constrained by the policy rate hitting the zero lower bound, the response of both output and inflation to the risk premium shock is magnified on impact with the effects of the shock taking much more quarters to be reabsorbed by the economy.
Despite the sharp difference in the propagation of the shocks in the economy inside and outside the zero lower bound period predicted by a standard NK model, empirical evidences investigating this issue are mixed and many studies do not find substantial evidence supporting a different response of the US economy during the zero lower bound period. For example, galigambetti, support the irrelevance hypothesis i.e. the hypothesis that the economy's performance has not been affected by a binding ZLB constraint, embracing the view that unconventional monetary policy have been effective at getting around the zero lower bound (ZLB) constraint. lubikbenati recently showed that, given the short length of the zero-lower-bound period, inference based on a standard TVP-VAR where the time varying parameters are regarded as slow moving stochastic processes, doesn't allow to capture the changing relationship among the macroeconomic variables during the ZLB period predicted by the New-Keynesian model. In what follows, based on a simulation study, I also find that a standard TVP-VAR struggles to detect the change in the responses of the economy in the ZLB period generated by a NK model. I show that the TC-TVP-VAR can in principle be used to solve this inferential issue and therefore to recover the distinct response of the economy during the zero lower bound period.
The New-Keynesian model is the version of the SWauters model considered in delnegroschorfheide2005. The model features price and wage stickiness, investment adjustment costs, habit formation in consumption and 7 shocks being monetary policy shocks, technology shocks, price mark-up shocks, wage mark-up shocks, risk premium shocks, fiscal policy shock and shocks to the marginal efficiency of capital. In the model the monetary policy rule accounts for ZLB period and forward guidance.\footnote{The log-linearized equilibrium conditions of the model can be found in the appendix (ref) The model is labelled ”SW" in delnegroschorfheide2005 and assumes a constant inflation target.} More specifically, when accounting for the zero lower bound and forward guidance the solution of the model implies a state state space representation which exhibits time varying coefficients. The solution method follows the approach developed by kulishcagliarini for linear stochastic rational expectations models in the face of a finite sequence of anticipated structural changes. In particular, it is assumed that at a given period $t$, agents expect the nominal interest rate to be at the ZLB for $\bar{H}$ periods. That is, the monetary policy rule
becomes
for $\tau = t, \ldots, t + \bar{H}$ while it is determined by ((ref)) for $\tau > t + \Bar{H} $. This implies that the equilibrium conditions
differ over time depending on whether $\tau \leq t + \Bar{H}$, with the row corresponding to the policy rule differing across $\tau$. Following kulishcagliarini the solution for the linear rational expectations model under anticipated structural variations takes the form
where $\mathcal{C}_{\tau}^{(\tau, \Bar{H})}$ $ \mathcal{T}_{\tau}^{(\tau, \Bar{H})}$ and $\mathcal{R}_{\tau}^{(\tau, \Bar{H})}$ are computed by recursion
starting from $\mathcal{T}_{t+\Bar{H} + 1}^{(\tau, \Bar{H})} = \mathcal{T}$ , $\mathcal{C}_{t+1+\Bar{H}}^{(\tau, \Bar{H})}=0$ and where the superscript $(t,\Bar{H})$ is used to indicate that the solution is obtained under the assumption that the announcement of zero interest rates for a duration of $\Bar{H}$ periods was made in period $t$. Following delnegroschorfheide2005 to measure the number of quarters $\bar{H}$ that the Federal Funds Rate is expected to remain at the ZLB I exploit information based on the overnight index swap (OIS) rates. In particular I identify the ZLB period as the quarters in which the OIS rate is lower then 0.35. This classification leads to the same ZLB period considered in lubikbenati and galigambetti namely 2009Q1-2015Q3 (28 quarters). Following chencurdiaferrero I assume that the number of quarters such that the policy rate is expected to stay fixed, is at most equal to four. According to the model's solution, the matrices $\mathcal{C}_{t}^{(\tau, \Bar{H})},\mathcal{T}_{t}^{(\tau, \Bar{H})},\mathcal{R}_{t}^{(\tau, \Bar{H})} $ characterize the transition equation of a time varying coefficients state space model
where the constant matrices $\mathcal{D}$, $\mathcal{B}$ and time varying matrices $\mathcal{C}_{t}$, $\mathcal{T}_{t}$, $\mathcal{R}_{t}$ depend on the structural parameters $\boldsymbol{\theta}$ while $\boldsymbol{v}_t$ is a measurement error. \footnote{Details for the observation equations are in the appendix (ref) together with the list of the deep structural parameters $\boldsymbol{\theta}$ (appendix (ref)).} Therefore, while the space representation of the model's solution features time varying coefficients, the deep structural parameters of the NK model $\boldsymbol{\theta}$ are constant. Defining $\underline{T}^{zlb}$ the first period inside the ZLB and $\bar{T}^{zlb}$ the last period inside the ZLB, for $t < \underline{T}^{zlb}$ and $t > \bar{T}^{zlb}$ we have $\mathcal{R}_t = \bar{\mathcal{R}}$ and $\mathcal{T}_t = \bar{\mathcal{T}}$ and $\mathcal{C}_t = \boldsymbol{0}$. Therefore the state space model becomes
To parameterize the prior using the moments from this state space representation we proceed as follows. For $ t = 1, \ldots, \underline{T}^{zlb}-1$ and $ t = \bar{T}^{zlb}+1, \ldots, T$ the moments $\boldsymbol{\Gamma_{xx,t}} = \operatorname*{\mathbb{E}} [\boldsymbol{x_tx_t'}|\boldsymbol{\theta}]$, $\boldsymbol{\Gamma_{xy,t}} = \operatorname*{\mathbb{E}} [\boldsymbol{x_ty_t'}|\boldsymbol{\theta}]$ and $\boldsymbol{\Gamma_{yy,t}} \equiv \operatorname*{\mathbb{E}} [\boldsymbol{y_ty_t'}|\boldsymbol{\theta}]$ are computed assuming $\operatorname*{\mathbb{E}}[\boldsymbol{s}_t\boldsymbol{s}_t'] = \operatorname*{\mathbb{E}}[\boldsymbol{s}_{t-1}\boldsymbol{s}_{t-1}']$ and solving the Lyapunov equation
For the periods inside the ZLB, i.e. for $ \underline{T}^{zlb} \leq t \leq \bar{T}^{zlb}$ , $\operatorname*{\mathbb{E}}[\boldsymbol{s}_t\boldsymbol{s}_t']$ and the implied population moments are computed recursively according to the low of motion implied by ((ref)) and ((ref)). These moments are then used to compute $\boldsymbol{\Gamma_{xx}(\theta)}$, $\boldsymbol{\Gamma_{xy}(\theta)}$, and $\boldsymbol{\Gamma_{yy}(\theta)}$ as detailed in (ref) and parameterize the prior.
As anticipated above and shown in figure (ref), the NK model predicts a distinct response of the economy to a risk premium shock inside and outside the ZLB period. Conditioning on the vector of deep parameters of the NK model, I simulate data from the state space representation ((ref)) and ((ref)).\footnote{Parameters are calibrated according to the posterior mode in delnegroschorfheide2005} In the simulation I consider artificial samples with $T=139 $ mimicking quarterly observations for the period 1985Q1-2019Q3. The length of the ZLB period is 28 quarters, covering the period 2009Q1-2015Q3. I estimate a standard TVP-VAR model chan2009efficient on the simulated data to understand whether the model is able to recover the change in the response of the economy during the ZLB period generated by the NK model.\footnote{The standard TVP-VAR model is the model in chan2009efficient. To estimate the model I use the MATLAB codes kindly made available by Joshua Chan on his personal website.} Figure (ref) shows the estimated responses of output and inflation to a risk premium shock obtained from a standard TVP-VAR. The figure plots the responses of output growth (cumulative) and the inflation rate to a one standard deviation risk premium shock in two reference dates, one outside and the other one inside the ZLB period. In order to identify the shocks in the structural TVP-VAR I exploit the true impact matrix of the NK model given by
where $\boldsymbol{\Omega}(\boldsymbol{\theta})= diag(\sigma_g^2, \sigma_b^2, \sigma_{\mu}^2, \sigma_z^2, \sigma_{\lambda_f}^2, \sigma_{\lambda_w}^2, \sigma_{r}^2)$ is the diagonal matrix containing the volatility of the structural shocks. The idea is that, by conditioning on the correct identification of risk premium shocks, we aim to determine whether the standard TVP-VAR model can reveal a distinct propagation of these shocks both inside and outside the ZLB. As figure (ref) shows, inference based on a standard TVP-VAR struggles to provide convincing evidences supporting a distinct response of both output growth and the inflation rate to a risk premium shock inside and outside the ZLB period. As a matter of fact, the estimates of the impulse responses are so imprecise that the model doesn't allow to detect any change in the response of the economy inside and outside the ZLB. \footnote{As shown in the appendix in figure (ref), this result is not driven by this specific simulated sample of artificial observations.}
Figure (ref), instead, shows the responses obtained by estimating a TC-TVP-VAR that exploits the NK model as a prior for the time varying coefficients. The moment matrices $\boldsymbol{\Gamma_{xx}}$, $\boldsymbol{\Gamma_{xy}}$, $\boldsymbol{\Gamma_{yy}}$ used to parametrize the Normal-Inverse-Wishart prior are derived from the equations of the state representation ((ref)) and ((ref)). Clearly, due to the time variation of the coefficients in the state space representation, the moments of the structural model are time varying with $\operatorname*{\mathbb{E}}[\boldsymbol{x_tx_t'}|\boldsymbol{\theta}]$, $\operatorname*{\mathbb{E}}[\boldsymbol{x_ty_t'}|\boldsymbol{\theta}]$ and $\operatorname*{\mathbb{E}}[\boldsymbol{y_ty_t'}|\boldsymbol{\theta}]$ changing over time. This, in turns, implies time varying restriction functions for the coefficients of the TVP-VAR encoded in the prior, resulting from $\boldsymbol{\Phi(\theta)^*} = \boldsymbol{\Gamma_{xx}(\theta)^{-1}\Gamma_{xy}(\theta)}$. Since the time varying restrictions incorporate the change in the response of the economy foreseen by the NK model, the responses of output growth and the inflation rate to the risk premium shock are more precisely estimated as the time varying coefficients are shrinked towards these restrictions. This is shown in figure (ref), which plots the responses of output growth and of the inflation rate to a risk premium shock for different fixed values of $\gamma$, that is for a different amount of shrinkage towards the restrictions implied by the NK model.\footnote{Also in this case, in the TC-TVP-VAR we condition on the correct identification scheme for the structural shocks.} As the figure shows, a sufficiently high $\gamma$ is needed to detect a precise and distinct response of the economy inside and outside the ZLB period as predicted by the NK model. As $\gamma$ increases and gets big enough, we get more and more precise estimates of the time varying coefficients. This allows to detect the different propagation of the shocks inside and outside the ZLB period, as predicted by the New-Keynesian model (panels (ref), (ref), (ref)). Letting $\gamma \rightarrow \infty$ the estimated model is a restricted TVP-VAR in which the time varying coefficients exactly satisfy $\boldsymbol{\Phi(\theta)^*} = \boldsymbol{\Gamma_{xx}(\theta)^{-1}\Gamma_{xy}(\theta)}$. As expected, the estimated responses from this model almost exactly resemble the responses from the NK model (panel (ref)).
When analyzing real data, it is reasonable to think of the NK model just as an approximate (most likely misspecified) tightly parameterized representation of the true data generating process. In other words, we do not expect the restrictions implied by the NK model to hold exactly. However, when encoded into a prior, these restrictions might prove to be useful to get more precise estimates of the time varying coefficients and of nonlinear functions of these coefficients, such as the impulse response functions. In particular, in our context, shrinkage might turn out to be particularly useful to detect the change in the response of the economy during the ZLB period predicted by the NK model. In order to analyze to what extent the data support the changing behavior during the ZLB, my approach consists in exploiting the New-Keynesian model as a prior for the TC-TVP-VAR. The extent to which the predictions from the New-Keynesian model will be supported by the data will depend on the estimated posterior distribution of $\gamma$, the hyper-parameter governing the degree of shrinkage of the parameters towards the restriction implied by the New-Keynesian model. If the posterior distribution of $\gamma$ is concentrated relatively far from zero, the coefficients of the TVP-VAR will be shrinked towards the restrictions implied by the NK model and the estimated responses from the TC-TVP-VAR will resemble the predictions of the NK model. This would happen in practice if the restrictions from the NK model find enough support on the data. Conversely, if the restrictions from the NK are deemed implausible by the data, the posterior distribution of $\gamma$ will be concentrated around zero and the estimates of the time varying coefficients will not reflect the restrictions coming from the prior. The estimated model is a 7-variable TC-TVP-VAR for the US economy including output growth, consumption growth, investment growth, real wage growth, hours worked, inflation and the Fed Fund rate and it is estimated over the sample 1985Q1-2019Q3.\footnote{Details on the variables and their transformation are available in the appendix (ref).} In order to identify the shocks in the structural TVP-VAR, I exploit the impact matrix of the NK model. The deep parameters of the NK model are treated as unknown and estimated along with the other parameters of the model, which implies that on impact there is uncertainty on the effect of the shocks on the variables of the system. Figure (ref) shows the estimated responses from the TC-TVP-VAR where the hyper-parameter $\gamma$ is estimated along with the other parameters of the model. The posterior distribution of $\gamma$ is estimated to be concentrated far from zero, meaning that the time varying restriction functions from the NK model find sufficient support on the data. This, in turns, is reflected on the estimates of the $10^{th}-90^{th}$ credible sets which provide some evidences supporting a distinct responses of the economy inside and outside the ZLB period. As predicted by the NK model, the estimates suggest that when monetary policy is constrained by the nominal rate hitting the ZLB, the risk premium shock is reabsorbed at a much slower pace by the economy. Following the positive demand shock both output growth and the inflation rate increase, with the effect on output growth being more precisely estimated at shorter horizons. The effect on output peaks after almost two quarters outside the ZLB period, while about after five quarters inside the ZLB period. As for inflation, the effect peaks after one quarter outside the ZLB period and after two quarters inside the ZLB period. After the peak, the effect of the risk premium shock on both output and inflation dies at a much slower pace inside the ZLB period, as foreseen by the NK model. Clearly, given the symmetric nature of the impulse response functions in our econometric model, figure (ref) implies that in the face of an increase in the risk premium, the economy would experience more persistent decreases of both inflation and output inside the ZLB period. Hence, risk premium shocks become more important when the ZLB binds as found in gourio2020risk. More in general, this finding is compatible with the findings in scorfZLB, which building on mavroeidis2021identification develop a structural VAR in which an occasionally-binding constraint generates censoring of one of the dependent variables. They find that the presence of the ZLB is empirically relevant for the propagation of macroeconomic shocks. Differently from both mavroeidis2021identification and scorfZLB, but similarly to galigambetti and lubikbenati our approach assumes the ZLB period to be completely observable. As well, while both mavroeidis2021identification and scorfZLB leverage censoring to find identifying information on the propagation of macroeconomic shocks we directly resort to economic theory encoded in the NK model to identify the macroeconomic shocks. One key difference is that while in scorfZLB the structural coefficients switch across unobserved regimes, in our TVP-SVAR the autoregressive time varying parameters are allowed to slowly drift inside, outside and in the transition to and from the ZLB period. \footnote{As for the impact matrix $A_{0,t}(\boldsymbol{\theta})$, in our framework deterministically changes over time (conditioning on the NK parameters $\boldsymbol{\theta}$) as a function of $\bar{H}_t$, the number of periods agents think the ZLB remains binding.}
TVP-VARs are flexible statistical models used both for prediction and semi-structural analysis in macroeconomics. Despite their flexibility allows to capture changes in the dynamic relationship among the macroeconomic variables, inference on the time varying parameters is often very imprecise when using standard priors. This translates into poor forecasting performances and imprecise inference on typical objects of interests such as the impulse responses. On the other side, models from the economic theory typically provide a more tightly parameterized representation of the macroeconomy and therefore have the opposite tendency of fitting the data rather poorly. This paper exploits economic theory to formulate a prior for the parameters of TVP-VARs. More specifically, the paper introduces a novel shrinkage prior that centers the time varying coefficients on a path implied by an underlying economic theory about the variables in the system. The paper shows that “economic shrinkage” can be successfully used to obtain more accurate forecasts and more precise estimates of the impulse response functions. In an application, I exploit a medium scale New-Keynesian model that accounts for forward guidance and the Zero Lower bound period, to formulate a prior for a medium scale TVP-VAR. Using this theory coherent shrinkage prior allows to estimate more precisely the response of the economy to macroeconomic shocks inside and outside the ZLB period, helping to solve the inferential problems faced by the standard TVP-VAR. On US data, the paper finds indeed a distinct propagation of a pure demand shock - the risk premium shock - inside and outside the ZLB period, with the effects of the shock reabsorbing at a much slower pace inside the ZLB period. This finding has important implications for the conduct of both fiscal and macroprudential policy at the ZLB. \\
Future research For future research the prior proposed in this paper could be modified to accommodate for multiple competing theories à la LORIA2022105. As well the prior could be adapted to allow for different forms of heteroskedasticity but this extension is non-trivial since in TVP-VAR-SV a la primiceri2005 the specification of the model breaks the Kronecker structure of the likelihood. More in general, restrictions from economic theory could successfully be used to reduce overfitting and sharpen inference in non-parametric VARs and Gaussian Procesess VARs hauzenberger2022gaussian. \printbibliography