EconBase
← Back to paper

On Time-Varying VAR Models: Estimation, Testing and Impulse Response Analysis

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

61,152 characters · 12 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\newtheorem{corollary}{Corollary} \newtheorem{definition}{Definition} \newtheorem{lemma}{Lemma} \newtheorem{proposition}{Proposition} \newtheorem{remark}{Remark} \newtheorem{theorem}{Theorem} \newtheorem{assumption}{Assumption} \newtheorem{example}{Example}

\numberwithin{corollary}{section} \numberwithin{definition}{section} \numberwithin{equation}{section} \numberwithin{lemma}{section} \numberwithin{proposition}{section} \numberwithin{remark}{section} \numberwithin{theorem}{section}

\allowdisplaybreaks[4]

titlepage\begin{center} { On Time-Varying VAR Models: \\Estimation, Testing and Impulse Response Analysis} \begingroup \footnote{{\it Corresponding author}: Jiti Gao, Department of Econometrics and Business Statistics, Monash University, Caulfield East, Victoria 3145, Australia. Email: \url{[email removed]}. The authors of this paper would like to thank George Athanasopoulos, Rainer Dahlhaus, David Frazier, Oliver Linton, Gael Martin, Peter CB Phillips and Wei Biao Wu for their constructive comments on earlier versions of this paper. Gao acknowledges financial support from the Australian Research Council Discovery Grants Program under Grant Numbers: DP170104421 and DP200102769. Peng also acknowledges the Australian Research Council Discovery Grants Program for its financial support under Grant Number DP210100476.} \addtocounter{footnote}{-1} \endgroup {\sc Yayi Yan, Jiti Gao and Bin Peng } Monash University, Melbourne, Australia \today \end{center} \begin{abstract} Vector autoregressive (VAR) models are widely used in practical studies, e.g., forecasting, modelling policy transmission mechanism, and measuring connection of economic agents. To better capture the dynamics, this paper introduces a new class of time-varying VAR models in which the coefficients and covariance matrix of the error innovations are allowed to change smoothly over time. Accordingly, we establish a set of theories, including the impulse responses analyses subject to both of the short-run timing and the long-run restrictions, an information criterion to select the optimal lag, and a Wald-type test to determine the constant coefficients. Simulation studies are conducted to evaluate the theoretical findings. Finally, we demonstrate the empirical relevance and usefulness of the proposed methods through an application to the transmission mechanism of U.S. monetary policy. {\bf Keywords}: Multivariate Dynamic Time Series; Time-Varying Impulse Response; Testing for Parameter Stability {\bf JEL Classification}: C14, C32, E52 \end{abstract}

Introduction

Vector autoregressive (VAR) models as well as their extensions are among some of the most popular frameworks for modelling dynamic interactions among multiple variables. These models arise mainly as a response to the “incredible” identification conditions embedded in the large-scale macroeconomic models sims1980. VAR modelling begins with minimal restrictions on the multivariate dynamic models. Gradually armed with identification information, VAR models and their statistical tool-kits like impulse response functions become powerful tools for conducting policy analysis. We refer interested readers to stock2001vector for a comprehensive review. Despite the popularity, linear VAR models can always be rejected by data in empirical studies (tsay1998testing). For example, stock2016dynamic point out, “changes associated with the Great Moderation go beyond reduction in variances to include changes in dynamics and reduction in predictability”.

To go beyond linear VAR models, various parametric time-varying VAR models have been proposed (e.g., tsay1998testing, sims2006, and the references therein) in order to allow for certain changes in economic relationship. However, model misspecification and parameter instability may undermine the performance of the proposed parametric models. As pointed out by hansen2001new, it may seem unlikely that structural change could be immediate and might seem more reasonable to allow structural change to take a period of time to take effect. Practically, it may be more reasonable to allow smooth structural changes over a period of time rather than in an abrupt manner. To model the dynamic transit, an important strand of the VAR literature assumes that the coefficients of VAR models evolve in a random way (e.g., primiceri2005time, petrova2019quasi), and the estimation procedure relies on extensive Markov Chain Monte Carlo (MCMC) draws plus the use of a variety of filters, such as Kalman, or related filters. However, by doing so, the asymptotic properties of the estimated model coefficients as well as the corresponding impulse responses are still unclear (giraitis2014inference).

Meanwhile, there is a separate literature about using nonparametric methods to estimate deterministically unknown time--varying parameters for autoregressive models. Up to this point, it is worth bringing up the terminology “local stationarity", as a bulk of the literature is carried out using the local stationarity technique, which at least dates back to the seminal work by dahlhaus1996. Recent developments focus on univariate autoregressive models (dahlhaus2006statistical,zhang2012inference,richter2019cross,shlwz20). To the best of our knowledge, it has had very little success to extend the local stationarity technique to multivariate settings, such as VAR models with deterministic time--varying parameters. In some specific cases where different locally stationary univariate time series may be approximated by their stationary versions on the same segments, the local stationarity technique may be applicable to such specific multivariate settings. However, univariate time series of general multivariate time series may have quite different patterns and behaviours, such as the three univariate time series plotted in Figure (ref) of Section (ref).

To address the aforementioned issues, this paper therefore proposes a class of deterministic time-varying VAR models where both VAR coefficients and covariance matrix of the model's error innovations are allowed to change smoothly over time. We develop a time-varying vector moving average infinity (VMA$(\infty)$) representation for a class of VAR models before we are able to establish uniform consistency results and a joint central limit theory for kernel-based estimators of both VAR coefficients and covariance matrix, which facilitate the inference on time-varying structural impulse responses. This type of impulse responses is of major importance in typical VAR applications ik13,ik20,paul2019time. In addition, inferences on structural impulse responses subject to both short-run timing and long-run restrictions are considered, so that the proposed model can better capture the simultaneous relations among multiple variables of interest over time. Such modelling strategy is especially useful for analyzing multivariate time series over a long horizon, since it offers a comprehensive treatment on tracking interests which are affected by frequently updated policies, environment, system, etc. In an economy system consisting of inflation, unemployment and interest rates, we discuss time-varying impacts of the interest rate change, which helps stabilize fluctuations in inflation and unemployment in the long-run. Under the proposed framework, it is achieved by investigating the corresponding time-varying impulse response functions.

We now comment on the literature closely related to the parameter stability testing problem we are considering in this paper. Detecting and estimating parametric components in univariate time-varying autoregressive models has been studied by zhang2012inference. Recently, truquet2017parameter considers parameter stability testing for univariate time-varying autoregressive conditional heteroscedasticity. More recently, LS2021 discuss testing for time--varying impact effects through a change--point mechanism. We develop a simple test for checking whether some of the time--varying coefficients (if not all) reduce to constant coefficients involved in the VAR models. It can be used to test whether the policy transmission mechanism is changing over time, which is of great importance in the macroeconomic literature primiceri2005time,paul2019time. To give another example, consider our empirical study on the ongoing debate about the high inflation during 1970-1980, i.e., whether the high inflation is due to bad monetary policy (primiceri2005time, sims2006). Mathematically, the discrepancy comes down to specification testing on the coefficient matrices of the VAR models by the proposed test.

In summary, our contributions are in three-fold. First, we propose a class of time-varying VAR($p$) models, and develop a time-varying VMA$(\infty)$ representation for the VAR($p$) models before we establish an estimation theory for time-varying structural impulse responses, which are often of great interest to describe how the economy reacts over time to structural shocks. The proposed estimation is conducted subject to both short-run timing and long-run restrictions, which have attracted considerable attention in the literature (kilian2017structural). Second, we develop a Wald--type test statistic for detecting time-invariant parameters in time-varying VAR models before we show that the proposed test statistic is asymptotically normally distributed under both the null hypothesis and a sequence of local alternatives. Third, in an empirical study, we investigate the changing dynamics of three key U.S. macroeconomic variables (inflation, unemployment, and interest rate), and uncover a fall in the volatilities of exogenous shocks. We observe that there exists a substantial time-variation in the policy transmission mechanism, and the “price puzzle” is limited to periods of bad monetary policy.

The organization of this paper is as follows. Section (ref) considers a class of time-varying VAR models, and establish asymptotic properties for the estimated time-varying impulse response functions. Section (ref) describes the proposed test and establishes the corresponding asymptotic theory. Section (ref) discusses some implementation issues and presents comprehensive simulation studies. Section (ref) provides a case study to demonstrate the empirical relevance of the proposed models and estimation theory. Section (ref) concludes. The proofs of the main results are given in Appendix A. Preliminary lemmas and their proofs are given in Appendix B.

Before proceeding further, it is convenient to introduce some notations: $\| \cdot \|$ denotes the Euclidean norm of a vector or the Frobenius norm of a matrix; $\otimes$ denotes the Kronecker product; $\bm{I}_a$ stands for an $a\times a$ identity matrix; $\bm{0}_{a\times b}$ stands for an $a\times b$ matrix of zeros, and we write $\bm{0}_a$ for short when $a=b$; for a function $g(w)$, let $g^{(j)}(w)$ be the $j^{th}$ derivative of $g(w)$, where $j\ge 0$ and $g^{(0)}(w) \equiv g(w)$; $K_h(\cdot) =K(\cdot/h)/h$, where $K(\cdot)$ and $h$ stand for a nonparametric kernel function and a bandwidth respectively; let $\tilde{c}_k =\int_{-1}^{1} u^k K(u) du$ and $\tilde{v}_k= \int_{-1}^{1} u^k K^2(u) du$ for integer $k\ge 0$; $\mathrm{vec}(\cdot)$ stacks the elements of an $m\times n$ matrix as an $mn \times 1$ vector; for any $a\times a$ square matrix $\bm{A}$, $\mathrm{vech}\left(\bm{A}\right)$ denotes the $\frac{1}{2}a(a+1)\times 1$ vector obtained from $\mathrm{vec}\left(\bm{A}\right)$ by eliminating all supra--diagonal elements of $\bm{A}$; $\mathrm{tr}\left(\bm{A}\right)$ denotes the trace of $\bm{A}$. Finally, $\to_P$ and $\to_D$ denote convergence in probability and convergence in distribution, respectively.

The Time-Varying VAR($p$) Model

Suppose that we observe $\{\bm{x}_{-p+1},\ldots,\bm{x}_0,\bm{x}_1,\ldots,\bm{x}_T\}$ from the following data generating process:

equation[equation omitted — 181 chars of source]

where $\tau_t = t/T$, $\bm{a}(\cdot)$ and $ \bm{A}_{j}(\cdot)$ are respectively a $d\times 1$ vector and $d\times d$ matrices of unknown functional coefficients, and $\bm{\omega}(\tau)$ is a $d\times d$ matrix of unknown functions capturing the heteroskedasticity of the error components. Allowing $\bm{\omega}(\cdot)$ to vary over time is important theoretically and practically, because a constant covariance matrix implies that the shock to the $i^{th}$ variable of $\bm{x}_{t}$ has a time-invariant effect on the $j^{th}$ variable of $\bm{x}_{t}$, restricting simultaneous interactions among multiple variables to be time-invariant (primiceri2005time). For the time being, we assume $p$ is known, and shall come back to the choice of $p$ in Section (ref).

The following conditions are necessary for our development.

assumption\begin{enumerate} • The roots of $\bm{I}_d-\bm{A}_{1}(\tau)L-\cdots -\bm{A}_{p}(\tau)L^p=\bm{0}_d$ all lie outside the unit circle uniformly in $\tau \in [0,1]$. • Each element of $\bm{A}(\tau)=\left[\bm{a}(\tau),\bm{A}_1(\tau),\ldots,\bm{A}_{p}(\tau)\right]$ is second order continuously differentiable on $[0,1]$ and $\bm{A}(\tau)=\bm{A}(0)$ for $\tau<0$. • Each element of $\bm{\omega}(\tau)$ is second order continuously differentiable on $[0,1]$. Moreover, $\bm{\Omega}(\tau)=\bm{\omega}(\tau)\bm{\omega}(\tau)^\top$ is positive definite uniformly in $\tau \in [0,1]$ and $\bm{\omega}(\tau)=\bm{\omega}(0)$ for $\tau<0$. \end{enumerate}
assumption$\{\bm{\epsilon}_t\}_{t=-\infty}^{\infty}$ is a martingale difference sequence (m.d.s.) adapted to the filtration $\left\{\mathcal{F}_t\right\}$, where $\mathcal{F}_t=\sigma\left(\bm{\epsilon}_t,\bm{\epsilon}_{t-1},\ldots\right)$ is the $\sigma$-field generated by $\left(\bm{\epsilon}_t,\bm{\epsilon}_{t-1},\ldots\right)$, $E[\bm{\epsilon}_t \bm{\epsilon}_t^\top | \mathcal{F}_{t-1}]=\bm{I}_d$ almost surely (a.s.), and $ \max_{t}E \left\|\bm{\epsilon}_t\right\|^\delta < \infty$ for some $\delta > 4$.

Assumption (ref).1 ensures the eigenvalues of the companion matrix $\bm{\Phi}(\tau)$ of (ref) below all lie inside the unit circle uniformly over $\tau \in [0,1]$. As a consequence, (ref) cannot include any unit-root or explosive components. Similar treatments have also been adopted to investigate univariate locally stationary models in the literature (e.g., Assumption T3 of zhang2012inference). Assumptions (ref).2 and (ref).3 allow the underlying data generating process to evolve over time in a smooth manner. The conditions $\bm{A}(\tau)=\bm{A}(0)$ and $\bm{\omega}(\tau)=\bm{\omega}(0)$ for $\tau<0$ gives

eqnarray[eqnarray omitted — 121 chars of source]

which basically assumes that $\bm{x}_t$ behaves like a parametric VAR$(p)$ model for $t\le 0$. A similar condition can be found in vogt2012nonparametric for a nonparametric time-varying time series model.

Assumption (ref) imposes some conditions on the innovation error terms, and are standard in the VAR literature lutkepohl2005new.

With these conditions, the following proposition says (ref) admits a time-varying vector moving average infinity (VMA$(\infty)$) representation, which sheds a light on how to recover the time-varying structural impulse responses.

propositionUnder Assumptions (ref) and (ref), there exists a time-varying VMA$(\infty)$ process of the form: $$ \widetilde{\bm{x}}_t= \bm{\mu}(\tau_t)+\bm{B}_0(\tau_t)\bm{\epsilon}_t+\bm{B}_1(\tau_t)\bm{\epsilon}_{t-1}+\bm{B}_2(\tau_t)\bm{\epsilon}_{t-2}+\cdots $$ such that $\bm{x}_t$ of (ref) satisfies $\max_{t\geq 1} \{E\left\|\bm{x}_t-\widetilde{\bm{x}}_t\right\|^\delta\}^{1/\delta}=O(T^{-1})$, where \begin{eqnarray} &&\bm{\mu}(\tau)=\bm{a}(\tau)+\sum_ {j=1}^{\infty}\bm{\Psi}_j(\tau)\bm{a}(\tau),\quad \bm{\Psi}_j(\tau)=\bm{J}\bm{\Phi}^j(\tau) \bm{J}^\top for j\geq 1,\nonumber \\ &&\bm{J}=\left[\bm{I}_d,\bm{0}_{d\times d(p-1)}\right],\quad \bm{B}_0(\tau)=\bm{\omega}(\tau),\quad \bm{B}_j(\tau)=\bm{\Psi}_j(\tau)\bm{\omega}(\tau), \nonumber \\ &&\bm{\Phi}(\tau)=\left[\begin{matrix} \bm{A}_{1}(\tau) & \cdots & \bm{A}_{p-1}(\tau) & \bm{A}_{p}(\tau) \\ \bm{I}_d & \cdots& \bm{0}_d & \bm{0}_d\\ \vdots & \ddots&\vdots & \vdots\\ \bm{0}_d &\cdots &\bm{I}_d & \bm{0}_d\\ \end{matrix} \right]. \end{eqnarray}

From Proposition (ref), we can see that the $d \times 1$ vector of the orthogonalized impulse response function of a unit shock at time $t$ to the $j^{th}$ equation on $\bm{x}_{t+n}$ is given by $\bm{B}_n(\tau_{t+n})\bm{e}_j$, where $\bm{e}_j$ is a $d \times 1$ selection vector with unity as its $j^{th}$ element and zeros elsewhere. Hence, the impulse responses produced by our model are deterministic functions of rescaled time, so that the TV-VAR model captures potential drifts in the transmission mechanism and produces impulse responses which are not history- and shock-dependent.

Estimation

To estimate $\{\bm{B}_j(\tau): j\geq 0\}$, we need a joint central limit theorem for the estimators of the coefficients and the innovation covariance matrix. That said, we consider the estimation of $\bm{A}(\cdot)$ and $\bm{\Omega}(\cdot)$ using the local linear kernel method\footnote{The local linear kernel method allows us to address the so-called boundary effects of the kernel estimation when establishing the first result of Theorem (ref) below. Alternatively, one may consider some boundary adjustment approaches as mentioned in HongLi and CHL2012.}. Intuitively, when $\tau_t$ is in a small neighbourhood of $\tau$, we can write (ref) as

eqnarray[eqnarray omitted — 162 chars of source]

where $\bm{z}_{t-1}=\left[1,\bm{x}_{t-1}^\top,\ldots,\bm{x}_{t-p}^\top\right]^\top$ and $\bm{z}^*_{t-1} = \left[\bm{z}_{t-1}^\top,\frac{\tau_t-\tau}{h}\bm{z}_{t-1}^\top\right]^\top$. The local linear estimators\footnote{It is worth pointing out that the estimation of the covariance matrix using the local linear kernel method (such as the second estimator of (ref)) is a non-trivial problem, and even has its own literature. We refer interested readers to zhang2012inference for more details.} of $\bm{A}(\tau)$ and $\bm{\Omega}(\tau)$ are then respectively given by

eqnarray[eqnarray omitted — 383 chars of source]

where $\bm{Z}_{t}^*=\bm{z}_{t}^*\otimes \bm{I}_d$, $\bm{\widehat{\eta}}_t=\bm{x}_t-\bm{\widehat{A}}(\tau_t)\bm{z}_{t-1}$, $\omega_{t}(\tau)=K_h(\tau_t-\tau)\frac{ P_{h,2}(\tau)-\frac{\tau_t-\tau}{h}P_{h,1}(\tau)}{P_{h,0}(\tau)P_{h,2}(\tau)-P_{h,1}^2(\tau)}$ is the local linear weight, and $P_{h,k}(\tau)=\frac{1}{T}\sum_{t=1}^{T}\left(\frac{\tau_t-\tau}{h}\right)^k K_h(\tau_t-\tau)$ for $k=0,1,2$.

We require the following conditions to hold for the kernel function and the bandwidth.

assumptionLet $K(\cdot)$ be a symmetric and positive kernel function defined on $[-1,1]$ with $\int_{-1}^{1}K(u)\mathrm{d}u = 1$. Moreover, $K(\cdot)$ is Lipschitz continuous on $[-1,1]$. As $(T,h) \to (\infty, 0)$, $Th\to \infty$.

With Assumption (ref) in hand, we summarize the first theorem of this paper below.

theoremLet Assumptions (ref)-(ref) hold. Suppose that $\max_{t\geq1} E [\|\bm{\epsilon}_t \|^4 |\mathcal{F}_{t-1} ] < \infty $ a.s., and $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$. Then \begin{enumerate} • $\sup_{\tau \in [0,1]} \| \bm{\widehat{A}}(\tau)-\bm{A}(\tau) \|=O_P \left(h^2+ (\frac{\log T}{Th} )^{1/2} \right)$. \end{enumerate} In addition, suppose that conditional on $\mathcal{F}_{t-1}$, the third and fourth moments of $\bm{\epsilon}_t$ are identical to the corresponding unconditional moments a.s., and $Th^{5} \to \alpha \in [0,\infty)$. Then the following two results also hold for $\forall\tau \in (0,1)$: \begin{enumerate} • $\sqrt{Th} \widehat{\bm{V}}^{-1/2}(\tau)\left[\begin{matrix} \mathrm{vec}\left(\bm{\widehat{A}}(\tau)-\bm{A}(\tau)-\frac{1}{2}h^2\tilde{c}_2\bm{A}^{(2)}(\tau)\right) \\ \mathrm{vech}\left(\bm{\widehat{\Omega}}(\tau)-\bm{\Omega}(\tau)-\frac{1}{2}h^2\tilde{c}_2\bm{\Omega}^{(2)}(\tau)\right) \end{matrix} \right]\to_D N\left(\bm{0},\bm{I}\right),$ \end{enumerate} where $\bm{V}(\tau) $ and $\widehat{\bm{V}}(\tau)$ are defined in (ref) and (ref) for the sake of presentation.

The first result of Theorem (ref) establishes a uniform convergence rate for $\bm{\widehat{A}}(\tau)$, which further allows us to establish a joint asymptotic distribution in the second result. If $\delta>5$, the usual optimal bandwidth $h_{opt}=O\left(T^{-1/5}\right)$ satisfies the condition $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$.

On Impulse Responses

Having established the joint CLT in Theorem (ref), we are now ready to study the impulse responses. As $\bm{\Omega}(\cdot)=\bm{\omega}(\cdot)\bm{\omega}^\top(\cdot)$, we cannot infer the elements of $\bm{\omega}(\cdot)$ unless certain identification restrictions are imposed. In the following, we consider two types of identification conditions: (i) the short-run timing restrictions, and (ii) the long-run restrictions. The economic interpretations of the two types of identification conditions can be found in kilian2017structural, and we do not repeat them here for the sake of space.

Under the short-run timing restrictions, $\bm{\omega}(\cdot)$ is a lower-triangular matrix. Thus, $\widehat{\bm{\omega}}(\tau)$ is chosen as the lower triangular matrix from the Cholesky decomposition of $\bm{\widehat{\Omega}}(\tau)$, i.e., $\bm{\widehat{\Omega}}(\tau)=\bm{\widehat{\omega}}(\tau)\bm{\widehat{\omega}}^\top(\tau)$. Alternatively, one can impose the conditions on the long-run impacts of the shocks (i.e., $\bm{B}(\tau)$ defined below). Specifically, define

eqnarray[eqnarray omitted — 277 chars of source]

where the last equalities of both lines follow in an obvious matter. Thus, the elements of $\bm{B}(\tau)$ may be recovered from $\bm{B}(\tau)\bm{B}^\top(\tau) = \bm{\Psi}(\tau)\bm{\Omega}(\tau)\bm{\Psi}^\top(\tau)$. It is then convenient to assume that $\bm{B}(\tau)$ is a lower-triangular matrix, so $\widehat{\bm{B}}(\tau)$ is chosen as the lower triangular matrix from the Cholesky decomposition of $\widehat{\bm{\Psi}}(\tau)\widehat{\bm{\Omega}}(\tau)\widehat{\bm{\Psi}}^\top(\tau)$, where $\widehat{\bm{\Psi}}(\tau)$ is defined in the same way as $\bm{\Psi}(\tau)$ in (ref) but replacing $\bm{A}_i(\tau)$ with $\widehat{\bm{A}}_i(\tau)$. Under the long-run restrictions, $\widehat{\bm{\omega}}(\tau) = \widehat{\bm{\Psi}}^{-1}(\tau)\widehat{\bm{B}}(\tau)$.

Either way, the estimator of the impulse response function $\bm{B}_j(\tau)$ for each given $j\ge 0$ is given by

eqnarray[eqnarray omitted — 109 chars of source]

where $\bm{\widehat{\Psi}}_j(\tau)=\bm{J} \bm{\widehat{\Phi}}^j(\tau) \bm{J}^\top$, $\bm{J}$ is defined in Proposition (ref), and $ \bm{\widehat{\Phi}}^j(\tau)$ is defined under (ref).

We summarize the asymptotic properties of the estimation of the impulse responses in the following theorem.

theoremUnder the conditions of Theorem (ref). For any fixed integer $j\ge 0$ \begin{eqnarray*} \sqrt{Th}\left(\mathrm{vec}\left(\bm{\widehat{B}}_j(\tau)-\bm{B}_j(\tau)\right)-\frac{1}{2}h^2\tilde{c}_2\bm{B}_j^{(2)}(\tau)\right)\to_D N\left(0,\bm{\Sigma}_{\bm{B}_j}(\tau)\right), \end{eqnarray*} where \begin{eqnarray*} \bm{B}_j^{(2)}(\tau)&=&\bm{C}_{j,1}(\tau)\mathrm{vec}\left(\bm{A}^{(2)}(\tau)\right) +\bm{C}_{j,2}(\tau)\mathrm{vech}\left(\bm{\Omega}^{(2)}(\tau)\right),\\ \bm{\Sigma}_{\bm{B}_j}(\tau) &=& \left[\bm{C}_{j,1}(\tau),\bm{C}_{j,2}(\tau)\right]\bm{V}(\tau)\left[\bm{C}_{j,1}(\tau),\bm{C}_{j,2}(\tau)\right]^\top. \end{eqnarray*} Specifically, we have the following expressions. \begin{enumerate} • Under the short-run timing restrictions, we have \begin{eqnarray*} \bm{C}_{0,1}(\tau)&=&0,\\ \bm{C}_{j,1}(\tau)&=&\left(\bm{\omega}^\top(\tau)\otimes \bm{I}_d\right) \left(\sum_{m=0}^{j-1} \bm{J}(\bm{\Phi}^\top(\tau))^{j-1-m}\otimes \bm{\Psi}_m(\tau) \right) \left[\bm{0}_{d^2p\times d},\bm{I}_{d^2p}\right],\ j \geq 1,\\ \bm{C}_{j,2}(\tau)&=&\left(\bm{I}_d\otimes\bm{\Psi}_j(\tau)\right) \bm{L}_d^\top\left(\bm{L}_d\bm{N}_1(\tau)\bm{L}_d^\top \right)^{-1},\ j \geq 0, \end{eqnarray*} in which $\bm{N}_1(\tau)=(\bm{I}_{d^2}+\bm{K}_{d,d})(\bm{\omega}(\tau)\otimes \bm{I}_d)$, the elimination matrix $\bm{L}_d$ satisfies that $\mathrm{vech}(\bm{F})=\bm{L}_d\mathrm{vec}(\bm{F})$ for any $d\times d$ matrix $\bm{F}$, and the commutation matrix $\bm{K}_{m,n}$ satisfies $\bm{K}_{m,n}\mathrm{vec}(\bm{G})=\mathrm{vec}(\bm{G}^\top)$ for any $m\times n$ matrix $\bm{G}$. • Under the long-run restrictions, we have \begin{eqnarray*} \bm{C}_{0,1}(\tau)&=&(\bm{I}_d \otimes \bm{\Psi}_j(\tau))\left(\bm{N}_1^\top(\tau)\bm{N}_1(\tau)+\bm{N}_2^\top(\tau)\bm{N}_2(\tau)\right)^{-1}\bm{N}_2^\top(\tau)\bm{D}_2(\tau),\\ \bm{C}_{j,1}(\tau)&=&\left(\bm{\omega}^\top(\tau)\otimes \bm{I}_d\right) \left(\sum_{m=0}^{j-1} \bm{J}(\bm{\Phi}^\top(\tau))^{j-1-m}\otimes \bm{\Psi}_m(\tau) \right)\left[\bm{0}_{d^2p\times d},\bm{I}_{d^2p}\right]+(\bm{I}_d \otimes \bm{\Psi}_j(\tau))\\ &&\times\left(\bm{N}_1^\top(\tau)\bm{N}_1(\tau)+\bm{N}_2^\top(\tau)\bm{N}_2(\tau)\right)^{-1}\bm{N}_2^\top(\tau)\bm{D}_2(\tau)\left[\bm{0}_{d^2p\times d},\bm{I}_{d^2p}\right],\ j \geq 1,\\ \bm{C}_{j,2}(\tau)&=&(\bm{I}_d \otimes \bm{\Psi}_j(\tau))\left(\bm{N}_1^\top(\tau)\bm{N}_1(\tau)+\bm{N}_2^\top(\tau)\bm{N}_2(\tau)\right)^{-1}\bm{N}_1^\top(\tau)\bm{D}_1,\ j \geq 0, \end{eqnarray*} in which $\bm{N}_2(\tau)=\bm{Q}\left(\bm{I}_d\otimes \bm{A}_\tau^{-1}(1)\right)$, $\bm{D}_{2}(\tau)=\bm{Q}\left(\bm{B}^\top(\tau)\otimes\bm{A}_\tau^{-1}(1)\right)\triangledown_{\bm{\alpha}(\tau)}\bm{A}_{\tau}(1)$ with $\bm{A}_{\tau}(1)=\bm{I}_d-\sum_{i=1}^{p}\bm{A}_i(\tau)$ and $\triangledown_{\bm{\alpha}(\tau)}\bm{A}_{\tau}(1)=-\left[\bm{I}_{d^2},...,\bm{I}_{d^2}\right]$ $(d^2\times d^2p)$, the duplication matrix $\bm{D}_1$ satisfies $\mathrm{vec}\left(\bm{\Omega}(\tau)\right)=\bm{D}_1\mathrm{vech}\left(\bm{\Omega}(\tau)\right)$, and $\bm{Q}$ is a $d(d-1)/2 \times d^2$ selection matrix of $0$ and $1$ such that $\bm{Q}\mathrm{vec}\left(\bm{B}(\tau)\right)=0$. \end{enumerate}

For ease of presentation, we provide the definitions of the respective estimators of $ \bm{V}(\tau)$ and $\bm{\Phi}(\tau)$ (i.e., $\widehat{\bm{V}}(\tau)$ and $\widehat{\bm{\Phi}}(\tau)$) in Appendix (ref). It is easy to see that $\widehat{\bm{\Phi}}(\tau)\to_P \bm{\Phi}(\tau)$, $\widehat{\bm{\omega}}(\tau)\to_P\bm{\omega}(\tau)$, and $\widehat{\bm{V}}(\tau)\to_P \bm{V}(\tau)$ by Theorem (ref). As a result, $\widehat{\bm{\Sigma}}_{\bm{B}_j}(\tau)\to_P\bm{\Sigma}_{\bm{B}_j}(\tau)$, where $\widehat{\bm{\Sigma}}_{\bm{B}_j}(\tau)$ has a form identical to $\bm{\Sigma}_{\bm{B}_j}(\tau)$ but replacing $\bm{\Phi}(\tau)$, $\bm{\omega}(\tau)$ and $ \bm{V}(\tau)$ with their estimators, respectively.

Selection of the Optimal Lag

We next consider the choice of the optimal number of lags, i.e., the estimation of $p$. Specifically, we consider the minimization of the next information criterion.

eqnarray[eqnarray omitted — 266 chars of source]

where $\text{RSS}(\mathsf{p})=\frac{1}{T}\sum_{t=1}^{T}\widehat{\bm{\eta}}_{\mathsf{p},t}^\top \widehat{\bm{\eta}}_{\mathsf{p},t}$, $\chi_T$ is the penalty term, $\widehat{\bm{\eta}}_{\mathsf{p},t}$ is the value of $\widehat{\bm{\eta}}_{t}$ by letting the lag be $\mathsf{p}$, and $\mathsf{P}$ is a sufficiently large fixed positive integer. The next theorem shows the validity of (ref).

theoremLet Assumptions (ref)-(ref) hold. Suppose $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$, $\max_{t\geq1} E (\|\bm{\epsilon}_t \|^4 |\mathcal{F}_{t-1} ) < \infty $ a.s., $\chi_T\to 0$, and $c_T^{-2}\chi_T\to \infty$, where $c_T=h^2+\left(\frac{\log T}{Th}\right)^{1/2}$. Then $\Pr\left(\widehat{\mathsf{p}}=p\right)\to 1$.

In view of the conditions on $\chi_T$, a natural choice is

eqnarray*[eqnarray* omitted — 87 chars of source]

To this end, we have established the asymptotic properties of the proposed estimators. In the following, we discuss some model specification issues when building (ref).

Testing for Parameter Stability

To close our investigation, we consider testing whether some (if not all) components of the coefficient matrices are time-invariant, which as explained in the introduction may help settle the discrepancy on the ongoing debate about the high inflation in the U.S. during 1970-1980.

For model (ref), we are specifically interested in testing

equation[equation omitted — 121 chars of source]

where $\bm{\beta}(\tau) := \mathrm{vec}\left(\bm{A}(\tau)\right)$ and $\bm{C}$ is a selection matrix. Practically, the choice of $\bm{C}$ should be theory/application driven, and $\bm{c}$ needs to be estimated. For example, in the context of monetary policy analysis primiceri2005time, one can let $\bm{C} = \left[\bm{0}_{d^2p\times d},\bm{I}_{d^2p}\right]$ to test whether the policy transmission mechanism is varying over time.

The test statistic is constructed based on the weighted integrated squared errors:

equation[equation omitted — 232 chars of source]

where $\widehat{\bm{\beta}}(\cdot) :=\text{vec}[\widehat{\bm A}(\cdot)]$ should be obvious, and $\widehat{\bm{c}}$ is the estimator of $\bm{c}$ to be constructed below. In (ref), $\bm{H}(\cdot)$ is an $s\times s$ positive definite weighting matrix, and is typically set as the precision matrix associated with $\widehat{\bm{\beta}}(\cdot)$.

Let $\bm{C}$ be a pre-specified matrix and $\bm{c} = \int_{0}^{1}\bm{C}\bm{\beta}(\tau)\mathrm{d}\tau$. Then a natural estimator of $\bm{c}$ is given by

equation[equation omitted — 107 chars of source]

of which the asymptotic properties are summarized in the following lemma.

lemmaLet Assumptions (ref)-(ref) hold. Suppose further that $\max_{t\geq1} E [\|\bm{\epsilon}_t \|^4 |\mathcal{F}_{t-1} ] < \infty $ a.s., $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$, each element of $\bm{A}(\cdot)$ has finite third-order derivative, $Th^2/(\log T)^2 \to \infty $, and $Th^6 \to 0$. As $T \to \infty$, we have $$ \sqrt{T}\left(\widehat{\bm{c}}-\bm{c}-\frac{1}{2}h^2\tilde{c}_2\int_{0}^{1}\bm{C}\bm{\beta}^{(2)}(\tau)\mathrm{d}\tau\right)\to_D N\left(\bm{0},\int_{0}^{1}\bm{C}\bm{V}_{\bm{\beta}}(\tau)\bm{C}^\top\mathrm{d}\tau\right), $$ where $\bm{V}_{\bm{\beta}}(\tau) := \bm{\Sigma}^{-1}(\tau)\otimes\bm{\Omega}(\tau)$ and $\bm{\Sigma}(\tau)$ is defined in (ref).

Lemma (ref) indicates that under the null, $\widehat{\bm{c}}$ will achieve a parametric rate (i.e., $\sqrt{T}$). This is important for the finite sample study, since more reliable information can be retrieved from the estimators.

Having established Lemma (ref), we move on to investigate the test statistic (ref), and examine the local alternatives. After some tedious development, one actually can show that

eqnarray[eqnarray omitted — 235 chars of source]

where $\bm{Z}_{t} = \bm{z}_t\otimes \bm{I}_d$, $\bm{H}_{0}(\tau)=\bm{\Sigma}_{\bm{Z}}^{-1}(\tau)\bm{C}^\top\bm{H}(\tau)\bm{C}\bm{\Sigma}_{\bm{Z}}^{-1}(\tau)$, $\bm{\Sigma}_{\bm{Z}}(\tau) = \bm{\Sigma}(\tau) \otimes \bm{I}_d$. Clearly,

eqnarray[eqnarray omitted — 146 chars of source]

converges to a fixed value in probability, so the first term on the right hand side of (ref) is the leading one. Moreover, $ \frac{\widetilde{v}_0}{T}\sum_{t=1}^T \bm{\eta}_t^\top \bm{Z}_{t-1}^\top \bm{H}_{0}(\tau_t) \bm{Z}_{t-1} \bm{\eta}_t $ can be considered as a weighted average of the second moment of $\bm\eta_t$, it thus does not yield any distribution for us to conduct the hypothesis testing. In this regard, it can be considered a bias term which should be removed. The term which really yields an asymptotic distribution is in fact the second part on the right hand side of ((ref)). Thus, we present the asymptotic distribution of the test statistic as follows.

theoremLet the conditions of Lemma (ref) hold. Under the null hypothesis, if $Th^{5.5}\to 0$ and $E [\|\bm{\epsilon}_t \|^\delta |\mathcal{F}_{t-1} ] < \infty $ a.s., we have \begin{eqnarray*} T\sqrt{h}\left(\widehat{Q}_{\bm{C}, \widehat{\bm{H}}} -\frac{1}{Th}s\widetilde{v}_0\right)\to_D N(0,4s C_B), \end{eqnarray*} where $\widehat{\bm{H}}(\tau) = (\bm{C}\widehat{\bm{V}}_{\bm{\beta}}(\tau)\bm{C}^\top )^{-1}$, $\widehat{\bm{V}}_{\bm{\beta}}(\tau) = \widehat{\bm{\Sigma}}^{-1}(\tau)\otimes\widehat{\bm{\Omega}}(\tau)$, $C_B = \int_{0}^{2}\left(\int_{-1}^{1-v}K(u)K(u+v) \mathrm{d}u\right)^2\mathrm{d}v$, and $\widehat{\bm{\Sigma}}(\tau)$ is defined in (ref).

Theorem (ref) states that the test statistic converges to a normal distribution. The term $s\widetilde{v}_0$ is the limit associated with (ref). Due to the use of $\widehat{\bm{H}}(\cdot)$, the second moment of $\bm\eta_t$ disappears from the asymptotic distribution automatically.

Here, we would like to emphasize that the proposed test is in fact a one-sided test, which is reflected in the investigation for the local alternative below. An intuitive explanation is that due to the quadratic form of (ref), any departure from the true value will eventually yield a squared term when analysing the asymptotic power. Therefore, the null of (ref) will be rejected at the level $\alpha$ if

eqnarray[eqnarray omitted — 203 chars of source]

where $q_{1-\alpha}$ stands for the $(1-\alpha)^{th}$ quantile of the standard normal distribution.

In what follows, we consider a sequence of local alternatives of the form:

eqnarray[eqnarray omitted — 95 chars of source]

where $\bm{f}(\tau)$ is a twice continuously differentiable vector of functions, and $d_T \to 0$. The term $d_T \bm{f}(\tau)$ characterizes the departure of the time-varying coefficient $\bm{C}\bm{\beta}(\tau)$ from the constant $\bm{c}$. Following the development of Theorem (ref), it is straightforward to obtain the following corollary.

corollaryLet the conditions of Theorem (ref) hold. Under the $\mathbb{H}_1$ of (ref), if $d_T = T^{-1/2}h^{-1/4}$, then \begin{eqnarray*} T\sqrt{h}\left(\widehat{Q}_{\bm{C}, \widehat{\bm{H}}} -\frac{1}{Th}s\widetilde{v}_0\right)\to_D N\left(\delta_1,4s C_B\right), \end{eqnarray*} where $\delta_1 = \int_{0}^{1}\bm{f}(\tau)^\top (\bm{C} \bm{V}_{\bm{\beta}}(\tau)\bm{C}^\top )^{-1}\bm{f}(\tau) \mathrm{d}\tau$. Moreover, we have $\Pr\left(\widehat{Q}_{\bm{C}, \widehat{\bm{H}}}^*>q_{1-\alpha}\right) \to \Phi\left(q_{\alpha} + \frac{\delta_1}{2\sqrt{ s C_B}} \right)$.

Corollary (ref) shows that the test has a non-trivial power against $\mathbb{H}_1$ when $d_T = T^{-1/2}h^{-1/4}$. If $T^{-1/2}h^{-1/4} = o(d_T)$, the power of the test converges to 1, i.e.,

eqnarray*[eqnarray* omitted — 82 chars of source]

Before we conclude this section, we would like to add some comments on the assumptions imposed and the main results established in relation to the relevant literature. The construction of the proposed test is similar to those discussed in zhang2012inference and then truquet2017parameter. Because we have developed and then employed Proposition 2.1 for the time--varying VMA($\infty$) representation, the assumptions, such as requiring $Th^{5.5} \to 0$, are less restrictive than those assumed in the relevant literature, see, for example, $Th^{3.5} \to 0$ by truquet2017parameter. As a consequence, the main techniques employed in our proofs may be of general interest and applicability in dealing with similar problems.

In what follows, we will examine the finite sample performance of the asymptotic properties for the proposed estimators and test statistic by simulation studies.

Simulation

In this section, we first provide some details of the numerical implementation in Section (ref), and then respectively examine the estimation and hypothesis testing in Sections (ref) and (ref).

Numerical Implementation

Throughout the numerical studies, Epanechnikov kernel $K(u)=0.75(1-u^2)I(|u|\leq 1)$ is adopted.

When selecting the optimal lag by (ref), the bandwidth $\widehat{h}_{cv}$ is always chosen by minimizing the following cross-validation criterion function for each $\mathsf{p}$.

equation[equation omitted — 185 chars of source]

where $\widehat{\bm{a}}_{-t}(\cdot)$, and $\widehat{\bm{A}}_{j,-t}(\cdot)$ are obtained using (ref) but leaving the $t^{th}$ observation out. Once $\widehat{\mathsf{p}}$ and $\widehat{h}_{cv}$ are obtained, the estimation procedure is relatively straightforward. It is pointed out that it might be possible to extend the Bayesian bandwidth selection method proposed for the conventional time--varying regression model by cgz2019 to the time--varying VAR setting under study for an optimal choice of the bandwidth parameter.

We then comment on the testing procedure. To improve the finite sample performance of the test, we propose a simulation-assisted testing procedure. A similar procedure has also been adopted by zhang2012inference and truquet2017parameter to for the same purpose in the context of univariate time-varying models. Alternatively, one may consider the moving block bootstrap as in CK2021. For simplicity, we adopt the former as follows.

{\bf Algorithm --- a simulation-assisted testing procedure}

itemize• Step 1: Use the sample $\{\bm{x}_t\}$ to estimate the unrestricted model and the restricted model, and then compute $\widehat{Q}_{\bm{C}, \widehat{\bm{H}}}$ based on (ref). Step 2: Generate i.i.d. standard multivariate normal random vectors $\{\bm{x}_t^*\}$. Step 3: Compute the bootstrap statistic $\widetilde{Q}_{\bm{C}, \widehat{\bm{H}}}^b$ in the same way as $\widehat{Q}_{\bm{C}, \widehat{\bm{H}}}$, with $\{\bm{x}_t^*\}$ replacing the original sample $\{\bm{x}_t\}$. Step 4: Repeat Steps 2-3 $B$ times to obtain $B$ bootstrap test statistics $\{\widetilde{Q}_{\bm{C}, \widehat{\bm{H}}}^b\}_{b=1}^{B}$, as well as its empirical quantile $\widehat{q}_{1-\alpha}$. We reject the null hypothesis (ref) at level $\alpha$ if $\widehat{Q}_{\bm{C}, \widehat{\bm{H}}}>\widehat{q}_{1-\alpha}$.

Examining the Model Estimation

We now examine the finite sample performance of the theoretical findings. The data generating process (DGP) is as follows.

equation[equation omitted — 222 chars of source]

where $\bm{\epsilon}_t$'s are i.i.d. draws from $N(\bm{0}_{2\times 1}, \bm{I}_2)$,

eqnarray*[eqnarray* omitted — 751 chars of source]

Let the sample size be $T\in \{200, 400, 800\}$, and conduct 1000 replications for each choice of $T$.

First, we check whether the coefficient matrices $\bm{A}_1(\tau)$ and $\bm{A}_2(\tau)$ satisfy Assumption (ref).1. Thus, for each generated dataset, we compute the largest eigenvalue of the true companion matrix $\bm{\Phi}(\tau)$ in $\tau \in [0,1]$, which varies from 0.54 to 0.88 over replications indicating the validity of Assumption (ref).1.

Next, we report the percentages of $\widehat{\mathsf{p}} < 2$, $\widehat{\mathsf{p}} = 2$, and $\widehat{\mathsf{p}} > 2$ respectively based on 1000 replications. Table (ref) shows that the information criterion (ref) performs reasonably well, as the percentages associated with $\widehat{\mathsf{p}}=2$ are sufficiently close to 1 even for $T=200$.

table[table omitted — 456 chars of source]

Finally, we evaluate the estimates of $\bm{A}(\tau)$, $\bm{\Omega}(\tau)$, as well as the estimates of the impulse responses (say, $\bm{B}_1(\tau)$ and $\bm{B}_5(\tau)$ without loss of generality) based on the short-run timing restrictions. For each parameter of interest, we calculate the root mean square error (RMSE) as follows:

equation*[equation* omitted — 149 chars of source]

where $\bm{\theta}(\cdot)\in\left\{\bm{A}(\cdot),\bm{\Omega}(\cdot),\bm{B}_1(\tau),\bm{B}_5(\tau)\right\}$, and $\widehat{\bm{\theta}}^{(n)}(\tau)$ is the estimate of $\bm{\theta}(\tau)$ for the $n^{th}$ replication. Of interest, we also report the finite sample coverage probabilities of the confidence intervals. The nominal coverage is 95%. Given $\bm{\theta}(\cdot)$, for each generated dataset, the coverage probability is first calculated for each functional component of $\bm{\theta}(\cdot)$ over the grid points $\{\tau_t,t=1,\ldots,T\}$, and then we further take an average across the elements of $\bm{\theta}(\cdot)$. After 1000 replications, we present the averaged value of these coverage probabilities in Table (ref).

As shown in Table (ref), the RMSEs decrease as the sample size increases. The finite sample coverage probabilities are smaller than their nominal level when $T=200$, but are fairly close to 95% as $T=800$.

table[table omitted — 495 chars of source]

Examining the Parameter Stability Testing

To evaluate the size and local power of the proposed test statistic, we consider the following DGP:

equation[equation omitted — 109 chars of source]

where $\bm{A}_2(\cdot)$ and $\bm{\eta}_t$ are generated in the same way as in Section (ref), and $$ \bm{A}_1(\tau)=\left[

matrix[matrix omitted — 43 chars of source]

\right]+b\times d_T\times \left[

matrix[matrix omitted — 91 chars of source]

\right] $$ in which $d_T = T^{-1/2}h^{-1/4}$ and $b$ is set to be $0$, $2$ or $4$ in order to investigate the size and local power of the proposed test. We use the proposed testing procedure to test whether the coefficient $\bm{A}_1(\cdot)$ is time-varying.

Again, we let $T\in \{200,400,800\}$ and conduct $1000$ replications for each choice of $T$. We use the simulation-assisted testing procedure to get the empirical critical value $\widehat{q}_{1-\alpha}$ after $1000$ bootstrap replications. We consider a sequence of bandwidths to check the robustness of the proposed test with respect to the bandwidth choice:

equation[equation omitted — 89 chars of source]

Table (ref) reports the rejection rates at the $5\%$ and $10\%$ nominal levels. A few facts emerge from the table. First, our test has reasonable sizes using empirical critical values obtained by the bootstrap procedure. Second, the size behaviour of our test is not sensitive to the choices of bandwidths. As discussed in gao2008bandwidth, the estimation-based optimal bandwidths may also be optimal for testing purposes, so for simplicity one can use the rule-of-thumb in practice. Third, the local power of our test increases rapidly as $b$ increases. {

table[table omitted — 2,662 chars of source]

}

Empirical Study

In this section, we discuss the transmission mechanism of the monetary policy. As well documented, inflation is higher and more volatile in the U.S. during 1970-1980, but substantially decreases in the subsequent period, which is often referred to as the Great Moderation (primiceri2005time). Two explanations (i.e., bad policy and bad luck) have been debated repeatedly in the literature. The first explanation focuses on the changes in the transmission mechanism cogley2005drifts, while the second regards it as a consequence of changes in the size of exogenous shocks sims2006. In what follows, we revisit the arguments associated with the Great Moderation using our approach. The estimation is conducted in exactly the same way as in Section (ref), so we no longer repeat the details.

First, we estimate the time-varying VAR$(p)$ model using three commonly adopted macroeconomic variables of the literature (primiceri2005time,cogley2010inflation), which are the inflation rate (measured by 100 times the year-over-year log change in the GDP deflator), the unemployment rate representing the non-policy block, and the interest rate (measured by the average value for the Federal funds rates over the quarter) representing the monetary policy block. Following primiceri2005time, the short-run timing restrictions are employed to identify the monetary policy shocks. The interest rate is included as the third variable and being treated as the monetary policy instrument. The identification requirement is that the monetary policy actions affect the inflation and the unemployment with at least one period of lag ({i.e., short-run timing restrictions}). The data are quarterly observations measured at an annual rate from 1954:Q3 to 2015:Q4, which are collected from the Federal Reserve Bank of St. Louis economic database. Figure (ref) plots the three variables.

figure[figure omitted — 199 chars of source]

For the time-varying VAR$(p)$ model, the optimal lag is $\widehat{\mathsf{p}}=3$ by our approach, while it is often assumed to be known with the values varying from $2$ to $4$ in the literature. Thus, the following analyses focus on the time-varying VAR$(3)$ model (referred to as TV-VAR$(3)$ hereafter). We then conduct a robustness check to see whether the innovation process $\bm{\epsilon}_t$ exhibits serial correlation. We use the Breusch-Godfrey LM test breusch1978testing,godfrey1978testing for testing the first-order autocorrelation of $\bm{\epsilon}_t$'s. The null hypothesis is $H_0: E\left(\bm{\epsilon}_t\bm{\epsilon}_{t-1}^\top\right) = 0$. Based on the estimates $\widehat{\bm{\epsilon}}_t = \widehat{\bm{\omega}}^{-1}(\tau_t)\widehat{\bm{\eta}}_t$, the $p$-value is 0.81 suggesting that the TV-VAR$(3)$ model fits the data quite well.

Before investigating the time-varying effects of monetary policy on real economy, it is reasonable to check whether the VAR coefficients (i.e., the policy transmission mechanism) are time-varying. To this end, we employ the proposed test statistic to examine the constancy of model coefficients, and summarize the results in Table (ref). Specifically, we first set $\bm{C}$ to be an identity matrix for selecting between parametric and time-varying VAR models. From Table (ref) (the row labelled “Constancy”), we conclude that the VAR coefficients are not constant, implying that there exists significant time-variations in policy transmission mechanism. We then apply the testing procedures to examine whether the selected variables (i.e., intercept, $\bm{x}_{t-1}$, $\bm{x}_{t-2}$, $\bm{x}_{t-3}$) have time-varying contributions. By Table (ref), at the 5% significance level we conclude that all $\bm{a}(\cdot)$, $\bm{A}_1(\cdot)$, $\bm{A}_2(\cdot)$ and $\bm{A}_3(\cdot)$ are time-varying by examining them individually\footnote{Certainly, one may examine each element of these matrices. However, it will lead to a quite lengthy presentation. In order not to deviate from our main goal, we no longer conduct more testing along this line.}, suggesting that the TV-VAR(3) is more appropriate than a constant VAR model.

table[table omitted — 458 chars of source]

We now consider the changes in the size of exogenous shocks. Figure (ref) plots the estimated residuals $\widehat{\bm{\eta}}_t$. As can be seen, the magnitudes of estimated residuals change over time. For example, the magnitudes of the estimated residuals in the interest rate component increase from 1970 to 1980, but decrease quickly after 1980. The evidence favours the TV-VAR$(3)$ model which accounts for heteroscedasticity. More precisely, Figure (ref) plots the time-varying standard deviation of the estimated residuals $\widehat{\bm{\eta}}_t$ as well as the associated 95% confidence intervals. We can see that a decline in unconditional volatilities of exogenous shocks after 1980 in every subplot of Figure (ref). Thus, our results support the explanation of “bad luck” primiceri2005time,sims2006.

figure[figure omitted — 236 chars of source]

We then discuss how inflation and unemployment respond to monetary policy shocks. Figure (ref) plots the time-varying cumulative impulse responses\footnote{The cumulative impulse responses are constructed in exactly the same way as in primiceri2005time, and the standard errors are computed based on Theorem (ref).} of inflation and unemployment to a monetary shock subject to the short-run timing restrictions. Clearly, these responses vary over time, indicating a substantial time-variation in the policy transmission mechanism. The result differs from primiceri2005time that has found no evidence of time variation in the response of the economy to monetary policy shocks using a Bayesian approach. Interestingly, our results show that the responses of inflation exhibit a price puzzle during 1970-1980, which suggests that inflation increases in response to a monetary tightening although standard macroeconomic theory predicts the opposite. Therefore, our results also support the explanation of “bad policy” cogley2005drifts.

figure[figure omitted — 440 chars of source]
figure[figure omitted — 320 chars of source]

Additionally, Figure (ref) plots the impulse responses of inflation to a monetary policy shock in 1978:Q1 and 1994:Q1. The left panel in Figure (ref) exhibits the price puzzle, while the right panel shows a response that corresponds to the mainstream prior: the price level declines soon after a tightening. This result shows that the price puzzle is limited to periods of bad monetary policy, and is consistent with a stream of literature elbourne2006financial,rusnak2013solve, which suggests that the price puzzle emerges when using data for different monetary policy regimes and VARs estimated within a period of single monetary policy rarely encounter the price puzzle. Hence, the existence of time-varying policy transmission mechanism favours the TV-VAR(3) model.

figure[figure omitted — 237 chars of source]

Conclusions and Discussion

\setcounter{equation}{0} In this paper, we have proposed a new class of time-varying VAR models where the VAR coefficients and covariance matrix of the error innovations are allowed to evolve over time. Accordingly, we have established a set of asymptotic results, including the impulse responses analyses subject to both short-run timing and long-run restrictions, an information criterion to select the optimal lag, and a Wald--type test to determine the constant coefficients. Simulation studies are conducted to evaluate the theoretical findings. Finally, we demonstrate the empirical relevance and usefulness of the proposed methods through an application to the transmission mechanism of U.S. monetary policy. We have revealed that there exists a substantial time-variation in the policy transmission mechanism, and the “price puzzle” is limited to periods of bad monetary policy.

There are several directions for possible extensions. The first one is about how to consistently estimate the $d$-dimensional components of the VAR$(p)$ process for the case where the dimensionality, $d$, and the number of lags, $p$, may diverge along with the sample size, $T$. The second one is to allow for some time-varying structure in cointegrated dynamic models. We wish to leave such issues for future study.