Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
49,605 characters · 11 sections · 56 citation commands
\newtheorem{corollary}{Corollary} \newtheorem{definition}{Definition} \newtheorem{lemma}{Lemma} \newtheorem{proposition}{Proposition} \newtheorem{remark}{Remark} \newtheorem{theorem}{Theorem} \newtheorem{assumption}{Assumption} \newtheorem{example}{Example}
\numberwithin{corollary}{section} \numberwithin{definition}{section} \numberwithin{lemma}{section} \numberwithin{proposition}{section} \numberwithin{remark}{section} \numberwithin{theorem}{section}
\allowdisplaybreaks[4]
\setcounter{equation}{0}
Vector autoregressions (VARs), as well as their extensions like vector autoregressive moving average (VARMA) models and VARs with exogenous variables (VARX), are among some of the most popular frameworks for modelling dynamic interactions among multiple variables. These models arise mainly as a response to the “incredible” identification conditions embedded in the large--scale macroeconomic models sims1980. VAR modelling begins with minimal restrictions on the multivariate dynamic models. Gradually armed with identification information, VARs plus their statistical tool--kits like impulse response functions, are powerful approaches for conducting policy analysis. Also, VARs can be applied to other important tasks including data description and forecasting, see stock2001vector for a detailed review. Despite the popularity, linear VAR models can always be rejected by data in empirical studies tsay1998testing. For example, stock2016dynamic point out, “changes associated with the Great Moderation go beyond reduction in variances to include changes in dynamics and reduction in predictability.”
To go beyond linear VAR models, various parametric time--varying VAR models have been proposed (e.g., tsay1998testing, sims2006, and references therein) in order to allow for abrupt structural breaks in economic relationships and obtain efficient estimation. However, model misspecification and parameter instability may undermine the performance of parametric time--varying VAR models. Usually, researchers do not know the true functional forms of the time--varying parameters, so the choice on the functional forms might be somewhat arbitrary. In addition, as pointed out by hansen2001new, “it may seem unlikely that a structural break could be immediate and might seem more reasonable to allow a structural change to take a period of time to take effect”. Therefore, it is reasonable to allow smooth structural changes over a period of time rather than abrupt structural changes. Another strand of the VAR literature assumes that structural coefficients evolve in a random way, such as primiceri2005time and giannone2019priors. Recently, giraitis2014inference point out that “the theoretical asymptotic properties of estimating such processes via the Kalman, or related filters are unclear”. Along this line, giraitis2014inference and giraitis2018inference have achieved some useful asymptotic results. However, estimation theory for time--varying impulse response functions, which are of interest in interpreting multivariate dynamic models, are not yet established in the random walk case.
It is worth pointing out that although nonparametric estimation for deterministic time--varying models has received much attention initially on time series regression models (robinson1989nonparametric, cai07, lcg11, cgl12, zhang2012inference) over the past three decades and then on univariate autoregressive models (dahlhaus2006statistical, richter2019cross) in recent years, little study about multivariate autoregressive models with deterministic time--varying coefficients has been done. One exception is dahlhaus1999nonlinear, who use a wavelet method to transform the time--varying VAR model to a linear approximation with an orthonormal wavelet basis, and then show that the wavelet estimator attains the usual near--optimal minimax rate of $L_2$ convergence. Up to this point, it is worth bringing up the terminology “local stationarity", which dates back to the seminal work by dahlhaus1996kullback at least. While several studies have been conducted along this line (dahlhaus1996kullback and zhang2012inference on time--varying AR; dahlhaus2009empirical on time--varying ARMA; rohan2013nonparametric and truquet2017parameter on time--varying ARCH/GARCH), the literature has not ventured much outside the univariate setting. A commonly used method is to approximate a locally stationary process by a stationary approximation on each of the segments dahlhaus2019towards. However, it remains unclear to us how to extend this approximation method for the univariate setting to the multivariate case where the segments on which stationarity approximations for each univariate time series may be quite different.
This paper is to show the versatility of an alternative approach that is especially designed for a wide class of time--varying VMA$(\infty)$ processes. Our approach relies on an explicit decomposition of time--varying VMA$(\infty)$ processes into long--run and transitory elements, which is known as the Beveridge--Nelson (BN) decomposition (beveridge1981new,phillips1992asymptotics). The long--run component in the decomposition yields a martingale approximation to the sum of time--varying VMA$(\infty)$ processes. We are then able to deal with a wide class of multivariate dynamic models with smooth time--varying coefficients, which have a general time--varying VMA$(\infty)$ representation, nesting VAR, VARMA, VARX and so forth as special cases. Specifically, the structural coefficients are unknown functions of the re--scaled time, so that the proposed models can better capture the simultaneous relations among variables of interest over time. Such a modelling strategy is especially useful for analysing time series over a long horizon, since it offers a comprehensive treatment on tracking interests which are affected by frequently updated policies, environment, system, etc. In an economy system consisting of inflation, unemployment and interest rates, one priority in Section (ref) is inferring time--varying impacts of the interest rate change, which helps stabilize fluctuations in inflation and unemployment in long--run. Under the proposed framework, it is achieved by investigating the corresponding time--varying impulse response functions.
In summary, our contributions are in threefold. First, we propose a wide class of time--varying VMA$(\infty)$ models, which covers several classes of multivariate dynamic models. Second, we develop a time--varying counterpart of the conventional BN decomposition and propose a unified estimation method for a class of unknown time--varying functions. We then establish the corresponding asymptotic theory for the proposed models and estimators. Third, in the empirical study of Section (ref), we study the changing dynamics of three key U.S. macroeconomic variables (i.e., inflation, unemployment, and interest rate), and uncover a fall in the volatilities of exogenous shocks. In addition, we find that (i) monetary policy shocks have less influence on inflation before and during the so--called Great Moderation; (ii) inflation is more anchored recently; and (iii) the long--run level of inflation is below, but quite close to the Federal Reserve's target of two percent after the beginning of the Great Moderation period.
The organization of this paper is as follows. Section (ref) proposes a class of time--varying VMA($\infty$) models. In Section (ref), we discuss a class of time--varying VAR models and then establish an estimation theory for the unknown quantities. Section (ref) presents an empirical study to show the practical relevance and applicability of the proposed models and estimation theory. Section (ref) gives a short conclusion. The main proofs of the theorems are given in Appendix A. In the online supplementary material, simulation results are given in Appendix B.1 and some technical lemmas and proofs are given in the rest of Appendix B.
Before proceeding further, it is convenient to introduce some notation: $\| \cdot \|$ denotes the Euclidean norm of a vector or the Frobenius norm of a matrix; $\otimes$ denotes the Kronecker product; $\bm{I}_a$ and $\bm{0}_a$ are $a\times a$ identity and null matrices respectively, and $\bm{0}_{a\times b}$ stands for a $a\times b$ matrix of zeros; for a function $g(w)$, let $g^{(j)}(w)$ be the $j^{th}$ derivative of $g(w)$, where $j\ge 0$ and $g^{(0)}(w) \equiv g(w)$; $K_h(\cdot) =K(\cdot/h)/h$, where $K(\cdot)$ and $h$ stand for a nonparametric kernel function and a bandwidth respectively; let $\tilde{c}_k =\int_{-1}^{1} u^k K(u) du$ and $\tilde{v}_k= \int_{-1}^{1} u^k K^2(u) du$ for integer $k\ge 0$; $\mathrm{vec}(\cdot)$ stacks the elements of an $m\times n$ matrix as an $mn \times 1$ vector; for any $a\times a$ square matrix $\bm{A}$, $\mathrm{vech}\left(\bm{A}\right)$ denotes the $\frac{1}{2}a(a+1)\times 1$ vector obtained from $\mathrm{vec}\left(\bm{A}\right)$ by eliminating all supra--diagonal elements of $\bm{A}$; $\mathrm{tr}\left(\bm{A}\right)$ denotes the trace of $\bm{A}$; finally, let $\to_P$ and $\to_D$ denote convergence in probability and convergence in distribution, respectively.
\setcounter{equation}{0}
We start our investigation by considering a class of time--varying VMA$(\infty)$ model:
for $t=1,\ldots,T$, where $\bm{x}_t$ is a $d$--dimensional vector of observable variables, $\bm{\mu}_t$ is a $d$--dimensional unknown trending function, $\bm{\epsilon}_t$ is a vector of $d$--dimensional random innovations, and $d$ is fixed throughout this paper. Moreover, $\bm{\mathbb{B}}_t(L)=\sum_{j=0}^{\infty}\bm{B}_{j,t}L^j$, where $L$ is the lag operator, and $\bm{B}_{j,t}$ is a matrix of $d \times d$ unknown deterministic coefficients.
We first comment on the usefulness of the structure in (ref), and the corresponding BN decomposition. An application of the BN decomposition gives
where we have used the decomposition of $\bm{\mathbb{B}}_t(L)$ as $\bm{\mathbb{B}}_t(L)= \bm{\mathbb{B}}_t(1) -(1-L)\widetilde{\mathbb{B}}_t(L)$, in which $\widetilde{\mathbb{B}}_t(L) = \sum_{j=0}^{\infty}\bm{\widetilde{B}}_{j,t}L^j$ and $\bm{\widetilde{B}}_{j,t}=\sum_{k=j+1}^{\infty}\bm{B}_{k,t}$. Equation ((ref)) indicates that one may establish some general asymptotic properties for partial sums and quadratic forms of $\bm{x}_t$ with minor restrictions on $\{\bm{\epsilon}_t\}$. For example, one can show that the simple average of $\bm{x}_t$ becomes
Similar to (ref), asymptotic properties for partial sums and quadratic forms mainly depend on the probabilistic structure of $\{\bm{\epsilon}_t\}$ and regularity conditions on $\{\bm{B}_{j,t}\}$. In other words, there is no need to impose any further structure on $\bm{x}_t$, such as requiring $\bm{x}_t$ to be locally stationary time series in a similar way to what has been done in the relevant literature for the univariate case. As a consequence, it facilitates to develop general theory for the multivariate case. Moreover, it should be added that our settings in ((ref)) and ((ref)) considerably extend similar treatments by phillips1992asymptotics for the univariate linear process case where both $\bm{\mu}_t$ and $\bm{B}_{j,t}$ reduce to constant scalars: $\bm{\mu}_t = \mu$ and $\bm{B}_{j,t} = B_j$.
Let us now stress that (ref) covers a wide range of models, which are of general interest in both theory and practice. Below, we list a few examples, of which the parametric counterparts can be seen in lutkepohl2005new.
To facilitate the development of our general theory, we introduce the following assumptions.
Assumption (ref) regulates the matrices $\bm{B}_{j,t}$'s, and ensures the validity of the BN decomposition under the time--varying framework. It covers cases such as (i) the parametric setting of phillips1992asymptotics, and (ii) $\bm{B}_{j,t}:= \bm{B}_j(\tau_t)$, where $\bm{B}_j(\cdot)$ satisfies Lipschitz continuity on $[0,1]$ for all $j$. Assumption (ref) imposes conditions on the innovation error terms by replacing the commonly used independent and identically distributed (i.i.d.) innovations (e.g., dahlhaus2009empirical) with a martingale difference structure.
We are now ready to present a summary of useful results for Examples 1--3, which explains why model (ref) serves as a foundation of the examples given above.
We now move on to investigate asymptotic properties for model (ref). We first propose some estimates for several population moments of $\bm{x}_t$ in (ref), which help derive the asymptotic theory throughout this paper. To conserve space, we present the rates of the uniform convergence below, while extra results on point--wise convergence are given in the supplementary Appendix B.
Theorem (ref) is readily used for studying many useful cases, including weighted kernel estimators (see Lemma (ref) of Appendix B for example), and will be repeatedly used in many of the theoretical derivations of this paper. Theorem (ref) is also helpful to a broad range of studies, such as those mentioned in HLG20, FanYao03, Gao2007, LR2007, hansen2008uniform, wang2014uniform, and li2016uniform.
As modelling time--varying means is always an important task in time series analysis (e.g., wu2007inference,friedrich2020autoregressive), we infer $\bm{\mu}_t$ of the model (ref) below. Up to this point, we have not imposed any specific form on the components $\bm{\mu}_t $ and $\bm{B}_{j,t}$ of $\bm{x}_t$. To carry on with our investigation, we suppose further that $\bm{\mu}_t = \bm{\mu}(\tau_t)$ and $\bm{B}_{j,t}=\bm{B}_{j}(\tau_t)$ with $\tau_t=t/T$, so (ref) can be expressed by
The challenge then lies in the fact that “residuals” are time--varying linear processes. Some detailed explanations can be found in dahlhaus2012locally.
The following assumptions are necessary for the development of our trend estimation.
Assumption (ref) can be considered as a stronger version than Assumption (ref). Assumption (ref) is standard in the literature of kernel regression (LR2007).
For $\forall\tau\in (0,1)$, we recover $\bm{\mu}(\tau)$ by the next estimator.
We are now ready to establish an important and useful theorem.
Note that $\bm{\Sigma}_{\bm{\mu}}(\cdot)$ is the long--run covariance matrix, which in general cannot be estimated directly. To construct confidence intervals practically, we use a dependent wild bootstrap (DWB) method which is initially proposed by shao2010dependent for stationary time series. For the sake of space, the detailed procedure with the associated asymptotic properties is presented in Appendix (ref).
\setcounter{equation}{0}
In this section, we pay particular attention to one of the most popular models of the VMA$(\infty)$ family --- VAR. Many multivariate time series exhibit time--varying simultaneous interrelationships and changes in unconditional volatility justiniano2008time,coibion2011monetary. Along this line, time--varying VAR models have proven to be especially useful for describing the dynamics of the multivariate time series. Majority time--varying VAR models are investigated under the Bayesian framework, while little has been done using a frequentists' approach. Building on Section (ref), we consider a time--varying VAR model under the nonparametric framework, and establish the corresponding estimation theory.
Suppose that we observe $\left(\bm{x}_{-p+1},\ldots,\bm{x}_0,\bm{x}_1,\ldots,\bm{x}_T\right)$ from the following data generating process. Accounting for heteroscedasticity, we consider the next model.
where $\bm{\omega}(\tau)$ is a matrix of unknown functions of $\tau$. The model (ref) allows dynamic variations for both the coefficients and the covariance matrix. We infer $\bm{a}(\cdot)$ and $ \bm{A}_{j}(\cdot)$'s below, which are respectively $d\times 1$ vector and $d\times d$ matrices of unknown smooth functions. In addition, we are interested in the $d\times d$ dimension $\bm{\omega}(\cdot)$ which governs the dynamics of the covariance matrix of $\{\bm{\eta}_t\}_{t=1}^T$. As mentioned in primiceri2005time, allowing $\bm{\omega}(\cdot)$ to vary over time is important theoretically and practically, because a constant covariance matrix implies that the shock to the $i$th variable of $\bm{x}_{t}$ has a time--invariant effect on the $j$th variable of $\bm{x}_{t}$, restricting simultaneous interactions among multiple variables to be time--invariant.
To facilitate the development, we first impose the following conditions.
Assumption (ref).1 ensures that model (ref) is neither unit--root nor explosive, while Assumption (ref).2 allows the underlying data generate process to evolve over time in a smooth manner. In addition, the conditions $\bm{A}(\tau)=\bm{A}(0)$ and $\bm{\omega}(\tau)=\bm{\omega}(0)$ for $\tau<0$ yield
for all $t\leq0$, which ensures (ref) behaves like a stationary parametric VAR$(p)$ model for $t\le 0$. Similar treatments can be found in vogt2012nonparametric for the univariate setting. With the above assumptions in hand, the following proposition shows that model (ref) can be approximated by a time--varying VMA$(\infty)$ process satisfying Assumption (ref).
Under Assumption (ref), when $\tau_t$ is in a small neighbourhood of $\tau$, we have
where $\bm{Z}_{t-1}= \bm{z}_{t-1}\otimes \bm{I}_d$ and $\bm{z}_{t-1}=(1,\bm{x}_{t-1}^\top,\ldots,\bm{x}_{t-p}^\top)^\top$. The estimators of $\bm{A}(\tau)$ and $\bm{\Omega}(\tau)$ are then sequentially given by
where $\bm{\widehat{\eta}}_t=\bm{x}_t-\bm{\widehat{A}}(\tau_t)\bm{z}_{t-1}$. The asymptotic properties associated with (ref) are summarized in the next theorem.
The first result of Theorem (ref) provides the uniform convergence rate for $\bm{\widehat{A}}(\tau)$. As a consequence, it allows us to establish a joint asymptotic distribution for the estimates of the coefficients and innovation covariance in the second result. For $\delta>5$, the usual optimal bandwidth $h_{opt}=O\left(T^{-1/5}\right)$ satisfies the condition $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$. The third result ensures the confidence interval can be constructed practically.
Before moving on to impulse responses, we consider a practical issue --- the choice of the lag $p$, which is usually unknown in practice and needs to be decided by the data. We select the number of lags by minimizing the next information criterion:
where $\text{IC}(\mathsf{p})=\log \left\{\text{RSS}(\mathsf{p})\right\}+\mathsf{p}\cdot\chi_T$, $\text{RSS}(\mathsf{p})=\frac{1}{T}\sum_{t=1}^{T}\widehat{\eta}_{\mathsf{p},t}^\top \widehat{\eta}_{\mathsf{p},t}$, $\chi_T$ is the penalty term, and $\mathsf{P}$ is a sufficiently large fixed positive integer. The next theorem shows the validity of (ref).
In view of the conditions on $\chi_T$, a natural choice is
In the supplementary Appendix (ref), we conduct intensive simulations to examine the finite sample performance of the above information criterion.
We now focus on the impulse responses below, which capture the dynamic interactions among the variables of interest in a wide range of practical cases.
By Proposition (ref), the impulse response functions of $\bm{x}_t$ is asymptotically equivalent to those of $\widetilde{x}_t$. Hence, recovering the impulse response functions requires estimating $\bm{\Psi}_j(\cdot)$'s and $\bm{\omega}(\cdot)$, which is then down to the estimation of $\bm{\Phi}(\cdot)$ and $\bm{\omega}(\cdot)$ by construction. Note further that $\bm{\Phi}(\cdot)$ is a matrix consisting of the coefficients of (ref). The estimator of $\bm{\Phi}(\cdot)$ is intuitively defined as $\widehat{\bm{\Phi}}(\cdot)$, in which we replace $\bm{A}_j(\cdot)$ of (ref) with the corresponding estimator obtained from (ref). Furthermore, we require $\bm{\omega}(\tau)$ to be a lower--triangular matrix in order to fulfil the identification restriction. Thus, $\widehat{\bm{\omega}}(\tau)$ is chosen as the lower triangular matrix from the Cholesky decomposition of $\bm{\widehat{\Omega}}(\tau)$ such that $\bm{\widehat{\Omega}}(\tau)=\bm{\widehat{\omega}}(\tau)\bm{\widehat{\omega}}^\top(\tau)$.
With the above notation in hand, we are now ready to present the estimator of $\bm{B}_j(\tau) $ by
where $\bm{\widehat{\Psi}}_j(\tau)=\bm{J} \bm{\widehat{\Phi}}^j(\tau) \bm{J}^\top$. The corresponding asymptotic results are summarized in the following theorem.
To close this section, we comment on how to construct the confidence interval. Since $\widehat{\bm{\Phi}}(\tau)\to_P \bm{\Phi}(\tau)$, $\widehat{\bm{\omega}}(\tau)\to_P\bm{\omega}(\tau)$ and $\widehat{\bm{V}}(\tau)\to_P \bm{V}(\tau)$ by Theorem (ref), it is straightforward to have $\widehat{\bm{\Sigma}}_{\bm{B}_j}(\tau)\to_P\bm{\Sigma}_{\bm{B}_j}(\tau)$, where $\widehat{\bm{\Sigma}}_{\bm{B}_j}(\tau)$ has a form identical to $\bm{\Sigma}_{\bm{B}_j}(\tau)$ but replacing $\bm{\Phi}(\tau)$, $\bm{\omega}(\tau)$ and $ \bm{V}(\tau)$ with their estimators, respectively.
We next show in Section (ref) about how to apply the proposed model and estimation method to an empirical data. Our findings show that the estimated coefficient matrices and impulse response functions capture various time--varying features.
\setcounter{equation}{0}
In this section, we study the transmission mechanism of the monetary policy, and infer the long--run level of inflation (i.e., trend inflation) and the natural rate of unemployment (NAIRU). The trend inflation and NAIRU are of central position in setting monetary policy since the Federal Reserve Bank aims to mitigate deviations of inflation and unemployment from their long--run targets. See stock2016core for more relevant discussions.
As well documented, inflation is higher and more volatile during 1970--1980, but substantially decreases in the subsequent period, which is often referred to as the Great Moderation (primiceri2005time). The literature has considered two main classes of explanations: bad policy or bad luck. The first type of explanations focuses on the changes in the transmission mechanism cogley2005drifts, while the second regards it as a consequence of changes in the size of exogenous shocks sims2006. In what follows, we revisit the arguments associated with the Great Moderation using our approach. Also, we use the VMA$(\infty)$ representation of the VAR$(p)$ model to infer the path of trend inflation and NAIRU over time.
First, we estimate the time--varying VAR$(p)$ model using three commonly adopted macroeconomic variables of the literature (primiceri2005time,cogley2010inflation), which are the inflation rate (measured by the 100 times the year--over--year log change in the GDP deflator), the unemployment rate, representing the non--policy block, and the interest rate (measured by the average value for the Federal funds rates over the quarter), representing the monetary policy block. To isolate the monetary policy shocks, the interest rate is ordered last in the VAR model, and is treated as the monetary policy instrument. The identification requirement is that the monetary policy actions affect the inflation and the unemployment with at least one period of lag primiceri2005time. The data are quarterly observations measured at an annual rate from 1954:Q3 to 2020:Q1, which are taken from the Federal Reserve Bank of St. Louis economic database. Figure (ref) plots the three macro variables.
The order of the VAR$(p)$ model and the optimal bandwidth are determined by the information criterion (ref) and the cross validation criterion (ref), respectively. We obtain $\widehat{\mathsf{p}}=3$ and $\widehat{h}_{cv}=0.435$. In the literature, the lag length is often assumed to be known with the values varying from $2$ to $4$, while our data--driven method indicates that 3 is the optimal value.
We first consider measuring the changes in the size of exogenous shocks. Figure (ref) plots the estimated volatilities of the innovations as well as the associated 95% confidence intervals. The figure exhibits evidence for a general decline in unconditional volatilities. Our results thus support the “bad luck” explanations primiceri2005time,sims2006.
We then consider the responses of the inflation to the monetary policy shocks. Figure (ref) plots the time--varying impulse responses of the inflation to a structural monetary shock as well as the associated 95% confidence intervals. It is clear that the confidence intervals are much wider at the beginning of the sample period implying higher uncertainty before 1970. On the other hand, the structural responses of the inflation seem to be statistically insignificant from 1970 to 2010. Thus, we conclude that the monetary shocks have less influence on the inflation before and during the period of the Great Moderation.
Finally, we investigate the trend inflation and the NAIRU. petrova2019quasi considers a Bayesian time--varying VAR(2) model, and induces the long--run mean of $\bm{x}_t$ by $\bm{\mu}_t=\lim_{p\rightarrow\infty}E_t(\bm{x}_{t+p})=(\bm{I}-\bm{A}_t)^{-1}\bm{a}_t$, where $\bm{a}_t$ is the intercept term and $\bm{A}_t$ is the autoregressive coefficients. The main difference between our method and Petrova's method is that we invert time--varying VAR$(p)$ model to the time--varying MA$(\infty)$ model, and then explicitly estimate the underlying trends of the inflation and the NAIRU using model (ref).
Figure (ref) plots the estimates of the trend inflation and the NAIRU. It is obvious that the underlying trend of the inflation is high in the 1970s, but decreases in the subsequent period. After the Great Moderation, the long--run level of the inflation is below, but quite close to the Federal Reserve's target of 2%, which indicates that the inflation is more anchored now than in the 1970s. However, the NAIRU is less persistent and fluctuates over time.
In this subsection, we focus on the out--of--sample forecasting, and compare the forecasts of a Bayesian time--varying parameter VAR with stochastic volatility (TVP--SV) primiceri2005time, as well as a VAR model with constant parameters (CVAR).
Specifically, we consider the 1--8 quarters ahead forecasts. That is to forecast
for $h=1$, $2$, $4$ and $8$, where $\bm{x}_t$ includes the values of the three aforementioned macro variables at the date $t$. The forecasts of $\overline{\bm{x}}_{t+1|t+h}$ are constructed by $\widehat{\overline{\bm{x}}}_{t+1|t+h}=\widehat{\bm{A}}_t\bm{z}_t$, in which $\widehat{\bm{A}}_t$ is estimated from CVAR, TVP--SV and the time--varying VAR model using the available data at time $t$ and $\bm{z}_t=[1,\bm{x}_{t}^\top,\bm{x}_{t-1}^\top,\bm{x}_{t-2}^\top]^\top$. The expanding window scheme is adopted. For comparison, we compute the root of mean square errors (RMSE) for CVAR as a benchmark, and the RMSE ratios for others. The out--of--sample forecast period covers 1985:Q2--end, about 35 years.
The forecasting results are presented in Table (ref), in which the values represent the ratios of the RMSEs of the corresponding method over the RMSEs of the benchmark method (i.e., CVAR). The result shows that the time--varying VAR model and TVP--SV perform much better than CVAR, which implies the desirability of introducing variations in forecasting models. In addition, the time--varying VAR model has a better forecasting performance than TVP--SV with the increase of the forecast horizon.
{
}
\setcounter{equation}{0}
This paper has proposed a class of time--varying VMA($\infty$) models, which nest, for instance, time--varying VAR, time--varying VARMA, and so forth as special cases. Both the estimation methodology and asymptotic theory have been established accordingly. In the empirical study, we have investigated the transmission mechanism of monetary policy using U.S. data, and uncover a fall in the volatilities of exogenous shocks. Our findings include (i) monetary policy shocks have less influence on inflation before and during the so--called Great Moderation, (ii) inflation is more anchored recently, and (iii) the long--run level of inflation is below, but quite close to the Federal Reserve's target of two percent after the beginning of the Great Moderation period. In addition, in Appendix B of the online supplementary material, we have evaluated the finite--sample performance of the proposed model and estimation theory.
There are several directions for possible extensions. The first one is to test about whether the $d$--dimensional components of the {\rm VAR}$(p)$ process is cross--sectionally independent for the case where the dimensionality, $d$, and the number of lags, $p$, may diverge along with the sample size, $T$. The second one is about model specification testing to check whether some of the time--varying coefficient matrices, $\bm{A}_j(\tau)$, may just be constant matrices, $\bm{A}_{j0}$. Existing studies by Gao08, pgy14, and cw19 may be useful for both issues. The third one is to allow for some cointegrated structure in our settings. The recent work by zry19 provides us with a good reference. We wish to leave such issues for future study.