EconBase
← Back to paper

A Class of Time-Varying Vector Moving Average Models: Nonparametric Kernel Estimation and Application

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

49,605 characters · 11 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\newtheorem{corollary}{Corollary} \newtheorem{definition}{Definition} \newtheorem{lemma}{Lemma} \newtheorem{proposition}{Proposition} \newtheorem{remark}{Remark} \newtheorem{theorem}{Theorem} \newtheorem{assumption}{Assumption} \newtheorem{example}{Example}

\numberwithin{corollary}{section} \numberwithin{definition}{section} \numberwithin{lemma}{section} \numberwithin{proposition}{section} \numberwithin{remark}{section} \numberwithin{theorem}{section}

\allowdisplaybreaks[4]

titlepage\begin{center} { A Class of Time--Varying Vector Moving Average ($\infty$) Models:\\ Nonparametric Kernel Estimation and Application} \begingroup \footnote{{\it Corresponding author}: Jiti Gao, Department of Econometrics and Business Statistics, Monash University, Caulfield East, Victoria 3145, Australia. Email: \url{[email removed]}. The authors of this paper would like to thank George Athanasopoulos, Rainer Dahlhaus, David Frazier, Oliver Linton, Gael Martin, Peter CB Phillips and Wei Biao Wu for their constructive comments on earlier versions of this paper. The second author would also like to acknowledge financial support from the Australian Research Council Discovery Grants Program under Grant Numbers: DP170104421 and DP200102769.} \addtocounter{footnote}{-1} \endgroup {\sc Yayi Yan, Jiti Gao and Bin Peng } Monash University \today \end{center} \begin{abstract} Multivariate dynamic time series models are widely encountered in practical studies, e.g., modelling policy transmission mechanism and measuring connectedness between economic agents. To better capture the dynamics, this paper proposes a wide class of multivariate dynamic models with time--varying coefficients, which have a general time--varying vector moving average (VMA) representation, and nest, for instance, time--varying vector autoregression (VAR), time--varying vector autoregression moving--average (VARMA), and so forth as special cases. The paper then develops a unified estimation method for the unknown quantities before an asymptotic theory for the proposed estimators is established. In the empirical study, we investigate the transmission mechanism of monetary policy using U.S. data, and uncover a fall in the volatilities of exogenous shocks. In addition, we find that (i) monetary policy shocks have less influence on inflation before and during the so--called Great Moderation, (ii) inflation is more anchored recently, and (iii) the long--run level of inflation is below, but quite close to the Federal Reserve's target of two percent after the beginning of the Great Moderation period. {\bf Keywords}: Multivariate Time Series Model; Nonparametric Kernel Estimation; Trending Stationarity \end{abstract}

Introduction

\setcounter{equation}{0}

Vector autoregressions (VARs), as well as their extensions like vector autoregressive moving average (VARMA) models and VARs with exogenous variables (VARX), are among some of the most popular frameworks for modelling dynamic interactions among multiple variables. These models arise mainly as a response to the “incredible” identification conditions embedded in the large--scale macroeconomic models sims1980. VAR modelling begins with minimal restrictions on the multivariate dynamic models. Gradually armed with identification information, VARs plus their statistical tool--kits like impulse response functions, are powerful approaches for conducting policy analysis. Also, VARs can be applied to other important tasks including data description and forecasting, see stock2001vector for a detailed review. Despite the popularity, linear VAR models can always be rejected by data in empirical studies tsay1998testing. For example, stock2016dynamic point out, “changes associated with the Great Moderation go beyond reduction in variances to include changes in dynamics and reduction in predictability.”

To go beyond linear VAR models, various parametric time--varying VAR models have been proposed (e.g., tsay1998testing, sims2006, and references therein) in order to allow for abrupt structural breaks in economic relationships and obtain efficient estimation. However, model misspecification and parameter instability may undermine the performance of parametric time--varying VAR models. Usually, researchers do not know the true functional forms of the time--varying parameters, so the choice on the functional forms might be somewhat arbitrary. In addition, as pointed out by hansen2001new, “it may seem unlikely that a structural break could be immediate and might seem more reasonable to allow a structural change to take a period of time to take effect”. Therefore, it is reasonable to allow smooth structural changes over a period of time rather than abrupt structural changes. Another strand of the VAR literature assumes that structural coefficients evolve in a random way, such as primiceri2005time and giannone2019priors. Recently, giraitis2014inference point out that “the theoretical asymptotic properties of estimating such processes via the Kalman, or related filters are unclear”. Along this line, giraitis2014inference and giraitis2018inference have achieved some useful asymptotic results. However, estimation theory for time--varying impulse response functions, which are of interest in interpreting multivariate dynamic models, are not yet established in the random walk case.

It is worth pointing out that although nonparametric estimation for deterministic time--varying models has received much attention initially on time series regression models (robinson1989nonparametric, cai07, lcg11, cgl12, zhang2012inference) over the past three decades and then on univariate autoregressive models (dahlhaus2006statistical, richter2019cross) in recent years, little study about multivariate autoregressive models with deterministic time--varying coefficients has been done. One exception is dahlhaus1999nonlinear, who use a wavelet method to transform the time--varying VAR model to a linear approximation with an orthonormal wavelet basis, and then show that the wavelet estimator attains the usual near--optimal minimax rate of $L_2$ convergence. Up to this point, it is worth bringing up the terminology “local stationarity", which dates back to the seminal work by dahlhaus1996kullback at least. While several studies have been conducted along this line (dahlhaus1996kullback and zhang2012inference on time--varying AR; dahlhaus2009empirical on time--varying ARMA; rohan2013nonparametric and truquet2017parameter on time--varying ARCH/GARCH), the literature has not ventured much outside the univariate setting. A commonly used method is to approximate a locally stationary process by a stationary approximation on each of the segments dahlhaus2019towards. However, it remains unclear to us how to extend this approximation method for the univariate setting to the multivariate case where the segments on which stationarity approximations for each univariate time series may be quite different.

This paper is to show the versatility of an alternative approach that is especially designed for a wide class of time--varying VMA$(\infty)$ processes. Our approach relies on an explicit decomposition of time--varying VMA$(\infty)$ processes into long--run and transitory elements, which is known as the Beveridge--Nelson (BN) decomposition (beveridge1981new,phillips1992asymptotics). The long--run component in the decomposition yields a martingale approximation to the sum of time--varying VMA$(\infty)$ processes. We are then able to deal with a wide class of multivariate dynamic models with smooth time--varying coefficients, which have a general time--varying VMA$(\infty)$ representation, nesting VAR, VARMA, VARX and so forth as special cases. Specifically, the structural coefficients are unknown functions of the re--scaled time, so that the proposed models can better capture the simultaneous relations among variables of interest over time. Such a modelling strategy is especially useful for analysing time series over a long horizon, since it offers a comprehensive treatment on tracking interests which are affected by frequently updated policies, environment, system, etc. In an economy system consisting of inflation, unemployment and interest rates, one priority in Section (ref) is inferring time--varying impacts of the interest rate change, which helps stabilize fluctuations in inflation and unemployment in long--run. Under the proposed framework, it is achieved by investigating the corresponding time--varying impulse response functions.

In summary, our contributions are in threefold. First, we propose a wide class of time--varying VMA$(\infty)$ models, which covers several classes of multivariate dynamic models. Second, we develop a time--varying counterpart of the conventional BN decomposition and propose a unified estimation method for a class of unknown time--varying functions. We then establish the corresponding asymptotic theory for the proposed models and estimators. Third, in the empirical study of Section (ref), we study the changing dynamics of three key U.S. macroeconomic variables (i.e., inflation, unemployment, and interest rate), and uncover a fall in the volatilities of exogenous shocks. In addition, we find that (i) monetary policy shocks have less influence on inflation before and during the so--called Great Moderation; (ii) inflation is more anchored recently; and (iii) the long--run level of inflation is below, but quite close to the Federal Reserve's target of two percent after the beginning of the Great Moderation period.

The organization of this paper is as follows. Section (ref) proposes a class of time--varying VMA($\infty$) models. In Section (ref), we discuss a class of time--varying VAR models and then establish an estimation theory for the unknown quantities. Section (ref) presents an empirical study to show the practical relevance and applicability of the proposed models and estimation theory. Section (ref) gives a short conclusion. The main proofs of the theorems are given in Appendix A. In the online supplementary material, simulation results are given in Appendix B.1 and some technical lemmas and proofs are given in the rest of Appendix B.

Before proceeding further, it is convenient to introduce some notation: $\| \cdot \|$ denotes the Euclidean norm of a vector or the Frobenius norm of a matrix; $\otimes$ denotes the Kronecker product; $\bm{I}_a$ and $\bm{0}_a$ are $a\times a$ identity and null matrices respectively, and $\bm{0}_{a\times b}$ stands for a $a\times b$ matrix of zeros; for a function $g(w)$, let $g^{(j)}(w)$ be the $j^{th}$ derivative of $g(w)$, where $j\ge 0$ and $g^{(0)}(w) \equiv g(w)$; $K_h(\cdot) =K(\cdot/h)/h$, where $K(\cdot)$ and $h$ stand for a nonparametric kernel function and a bandwidth respectively; let $\tilde{c}_k =\int_{-1}^{1} u^k K(u) du$ and $\tilde{v}_k= \int_{-1}^{1} u^k K^2(u) du$ for integer $k\ge 0$; $\mathrm{vec}(\cdot)$ stacks the elements of an $m\times n$ matrix as an $mn \times 1$ vector; for any $a\times a$ square matrix $\bm{A}$, $\mathrm{vech}\left(\bm{A}\right)$ denotes the $\frac{1}{2}a(a+1)\times 1$ vector obtained from $\mathrm{vec}\left(\bm{A}\right)$ by eliminating all supra--diagonal elements of $\bm{A}$; $\mathrm{tr}\left(\bm{A}\right)$ denotes the trace of $\bm{A}$; finally, let $\to_P$ and $\to_D$ denote convergence in probability and convergence in distribution, respectively.

Structure of Time--Varying VMA\texorpdfstring{$(\infty)$}

\setcounter{equation}{0}

We start our investigation by considering a class of time--varying VMA$(\infty)$ model:

equation[equation omitted — 148 chars of source]

for $t=1,\ldots,T$, where $\bm{x}_t$ is a $d$--dimensional vector of observable variables, $\bm{\mu}_t$ is a $d$--dimensional unknown trending function, $\bm{\epsilon}_t$ is a vector of $d$--dimensional random innovations, and $d$ is fixed throughout this paper. Moreover, $\bm{\mathbb{B}}_t(L)=\sum_{j=0}^{\infty}\bm{B}_{j,t}L^j$, where $L$ is the lag operator, and $\bm{B}_{j,t}$ is a matrix of $d \times d$ unknown deterministic coefficients.

We first comment on the usefulness of the structure in (ref), and the corresponding BN decomposition. An application of the BN decomposition gives

eqnarray[eqnarray omitted — 174 chars of source]

where we have used the decomposition of $\bm{\mathbb{B}}_t(L)$ as $\bm{\mathbb{B}}_t(L)= \bm{\mathbb{B}}_t(1) -(1-L)\widetilde{\mathbb{B}}_t(L)$, in which $\widetilde{\mathbb{B}}_t(L) = \sum_{j=0}^{\infty}\bm{\widetilde{B}}_{j,t}L^j$ and $\bm{\widetilde{B}}_{j,t}=\sum_{k=j+1}^{\infty}\bm{B}_{k,t}$. Equation ((ref)) indicates that one may establish some general asymptotic properties for partial sums and quadratic forms of $\bm{x}_t$ with minor restrictions on $\{\bm{\epsilon}_t\}$. For example, one can show that the simple average of $\bm{x}_t$ becomes

eqnarray[eqnarray omitted — 177 chars of source]

Similar to (ref), asymptotic properties for partial sums and quadratic forms mainly depend on the probabilistic structure of $\{\bm{\epsilon}_t\}$ and regularity conditions on $\{\bm{B}_{j,t}\}$. In other words, there is no need to impose any further structure on $\bm{x}_t$, such as requiring $\bm{x}_t$ to be locally stationary time series in a similar way to what has been done in the relevant literature for the univariate case. As a consequence, it facilitates to develop general theory for the multivariate case. Moreover, it should be added that our settings in ((ref)) and ((ref)) considerably extend similar treatments by phillips1992asymptotics for the univariate linear process case where both $\bm{\mu}_t$ and $\bm{B}_{j,t}$ reduce to constant scalars: $\bm{\mu}_t = \mu$ and $\bm{B}_{j,t} = B_j$.

Examples and Useful Properties

Let us now stress that (ref) covers a wide range of models, which are of general interest in both theory and practice. Below, we list a few examples, of which the parametric counterparts can be seen in lutkepohl2005new.

example\normalfont Suppose that $\bm{x}_t$ is a $d$--dimensional time--varying VAR$(p)$ process: \begin{equation} \bm{x}_t=\bm{A}_{1,t}\bm{x}_{t-1}+\cdots +\bm{A}_{p,t}\bm{x}_{t-p}+\bm{\epsilon}_t, \end{equation} which has been widely studied in the literature with Bayesian framework being the dominant approach (e.g., benati2009var, paul2019time). Similar to hamilton1994time, (ref) can be expressed as a time--varying MA($\infty$) process $\bm{x}_t=\sum_{j=0}^{\infty}\bm{B}_{j,t}\bm{\epsilon}_{t-j}$, where $\bm{B}_{0,t}=\bm{I}_d$ and $\bm{B}_{j,t}=\bm{J}\prod_{i=0}^{j-1}\bm{\Phi}_{t-i}\bm{J}^\top$ for $j\geq 1$ with $\bm{J}=[\bm{I}_d,\bm{0}_{d\times d(p-1)}]$ and \begin{eqnarray*} \bm{\Phi}_t=\left(\begin{matrix} \bm{A}_{1,t} & \cdots & \bm{A}_{p-1,t} & \bm{A}_{p,t} \\ \bm{I}_d & \cdots& \bm{0}_{d} & \bm{0}_{d}\\ \vdots & \ddots&\vdots & \vdots\\ \bm{0}_{d} &\cdots &\bm{I}_d &\bm{0}_{d}\\ \end{matrix} \right). \end{eqnarray*}
example\normalfont Suppose that $\bm{x}_t$ is a $d$--dimensional time--varying VARMA$(p,q)$ process as follows: \begin{equation} \bm{x}_t=\bm{A}_{1,t}\bm{x}_{t-1}+...+\bm{A}_{p,t}\bm{x}_{t-p}+\bm{\epsilon}_t+\bm{\Theta}_{1,t}\bm{\epsilon}_{t-1}+...+\bm{\Theta}_{q,t}\bm{\epsilon}_{t-q}, \end{equation} which then can be expressed as $\bm{x}_t =\sum_{b=0}^{\infty}\bm{D}_{b,t}\bm{\epsilon}_{t-b}$ with $\bm{D}_{b,t}=\sum_{j=\mathrm{max}(0,b-q)}^{b} \bm{B}_{j,t}\bm{\Theta}_{b-j,t-j}$, $\bm{B}_{j,t}$ defined similarly as in Example (ref), and $\bm{\Theta}_{0,t}\equiv \bm{I}_d$ independent of $t$.
example\normalfont Let $\bm{x}_t$ be a $d$--dimensional time--varying double MA$(\infty)$ process: \begin{equation} \bm{x}_t=\sum_{j=0}^{\infty}\bm{\Psi}_{j,t}\bm{v}_{t-j}\quad and\quad \bm{v}_t=\sum_{l=0}^{\infty}\bm{\Theta}_{l,t}\bm{\epsilon}_{t-l}, \end{equation} in which the innovations $\bm{v}_{t}$'s also follow a time--varying MA$(\infty)$ process. Simple algebra shows that $\bm{x}_t =\sum_{j=0}^{\infty} \bm{B}_{j,t} \bm{\epsilon}_{t-j}$, where $\bm{B}_{j,t}=\sum_{l=0}^{j}\bm{\Psi}_{l,t}\bm{\Theta}_{j-l,t-l}$.

To facilitate the development of our general theory, we introduce the following assumptions.

assumption$\max_{t\ge 1} \sum_{j=1}^{\infty} j \|\bm{B}_{j,t}\| < \infty$, $\limsup_{T\to \infty}\sum_{t=1}^{T-1}\sum_{j=1}^{\infty}j \|\bm{B}_{j,t+1}-\bm{B}_{j,t}\|<\infty$, and $\limsup_{T\to \infty}\sum_{t=1}^{T-1}\|\bm{\mu}_{t+1}-\bm{\mu}_{t}\|<\infty$.
assumption$\{\bm{\epsilon}_t\}_{t=-\infty}^{\infty}$ is a martingale difference sequences (m.d.s.) adapted to the filtration $\left\{\mathcal{F}_t\right\}$, where $\mathcal{F}_t=\sigma\left(\bm{\epsilon}_t,\bm{\epsilon}_{t-1},\ldots\right)$ is the $\sigma$--field generated by $\left(\bm{\epsilon}_t,\bm{\epsilon}_{t-1},\ldots\right)$, $E\left(\bm{\epsilon}_t \bm{\epsilon}_t^\top | \mathcal{F}_{t-1}\right)=\bm{I}_d$ almost surely (a.s.), and $ \max_{t\ge 1}E\left[\left\|\bm{\epsilon}_t\right\|^\delta\right] < \infty$ for some $\delta > 2$.

Assumption (ref) regulates the matrices $\bm{B}_{j,t}$'s, and ensures the validity of the BN decomposition under the time--varying framework. It covers cases such as (i) the parametric setting of phillips1992asymptotics, and (ii) $\bm{B}_{j,t}:= \bm{B}_j(\tau_t)$, where $\bm{B}_j(\cdot)$ satisfies Lipschitz continuity on $[0,1]$ for all $j$. Assumption (ref) imposes conditions on the innovation error terms by replacing the commonly used independent and identically distributed (i.i.d.) innovations (e.g., dahlhaus2009empirical) with a martingale difference structure.

We are now ready to present a summary of useful results for Examples 1--3, which explains why model (ref) serves as a foundation of the examples given above.

proposition\begin{enumerate} • Consider Examples (ref) and (ref). Suppose that the roots of $\bm{I}_d-\bm{A}_{1,t}-\cdots -\bm{A}_{p,t}=\bm{0}_d$ all lie outside the unit circle uniformly over $t$, $\limsup_{T\to \infty}\sum_{t=1}^{T-1}\left\|\bm{A}_{m,t+1}-\bm{A}_{m,t}\right\| < \infty$ for $m=1,\ldots,p$ and $\bm{A}_{m,t}=\bm{A}_{m,1}$ for $t \leq 0$ and $m=1,\ldots,p$. In addition, suppose that in Example 2, $\limsup_{T\to \infty}\sum_{t=1}^{T-1}\left\|\bm{\Theta}_{m,t+1}-\bm{\Theta}_{m,t}\right\| < \infty$ for $m=1,\ldots,q$. Then both (ref) and (ref) are time--varying MA$(\infty)$ processes, in which the MA coefficients satisfy Assumption (ref). • For Example (ref), let $\limsup_{T\to \infty}\sum_{t=1}^{T-1}\sum_{j=1}^{\infty}j \|\bm{\Psi}_{j,t+1}-\bm{\Psi}_{j,t}\|<\infty$ and $\max_{t\ge 1} \sum_{j=1}^{\infty} j \|\bm{\Psi}_{j,t}\| < \infty$. Moreover, let $\{\bm{\Theta}_{j,t}\}$ satisfy the same conditions as those for $\{\bm{\Psi}_{j,t}\}$. Then (ref) is a time--varying MA$(\infty)$, in which the MA coefficients satisfy Assumption (ref). \end{enumerate}

We now move on to investigate asymptotic properties for model (ref). We first propose some estimates for several population moments of $\bm{x}_t$ in (ref), which help derive the asymptotic theory throughout this paper. To conserve space, we present the rates of the uniform convergence below, while extra results on point--wise convergence are given in the supplementary Appendix B.

theoremLet Assumptions (ref) and (ref) hold. In addition, let $\left\{\bm{W}_{T,t}(\cdot)\right\}_{t=1}^T$ be a sequence of $m \times d$ matrices of deterministic functions, in which $m$ is fixed, each functional component is Lipschitz continuous and defined on a compact set $[a,b]$. Moreover, suppose that \begin{enumerate} • $\sup_{\tau\in[a,b]}\sum_{t=1}^{T} \|\bm{W}_{T,t}(\tau) \|=O(1)$; • $\sup_{\tau\in[a,b]}\sum_{t=1}^{T-1} \|\bm{W}_{T,t+1}(\tau)-\bm{W}_{T,t}(\tau) \| = O(d_T)$, where $d_T=\sup_{\tau\in[a,b],t\ge 1} \|\bm{W}_{T,t}(\tau) \|$. \end{enumerate} Then as $T\to \infty$, \begin{enumerate} • $ \sup_{\tau\in[a,b]} \|\sum_{t=1}^{T}\bm{W}_{T,t}(\tau)\left(\bm{x}_t-E(\bm{x}_t)\right) \|=O_P(\sqrt{d_T\log T})$ provided $T^{\frac{2}{\delta}} d_T \log T \rightarrow0$; • $\sup_{\tau\in[a,b]} \|\sum_{t=1}^{T}\bm{W}_{T,t}(\tau)\left(\bm{x}_t\bm{x}_{t+p}^\top-E(\bm{x}_t\bm{x}_{t+p}^\top)\right)\|=O_P (\sqrt{d_T\log T} )$ for any fixed integer $p\geq0$ provided $T^{\frac{4}{\delta}} d_T \log T \rightarrow0$ and $\max_{t\ge 1} E(\|\bm{\epsilon}_t \|^4 |\mathcal{F}_{t-1} ) < \infty$ a.s., where $\delta>2$ is the same as in Assumption 2. \end{enumerate}

Theorem (ref) is readily used for studying many useful cases, including weighted kernel estimators (see Lemma (ref) of Appendix B for example), and will be repeatedly used in many of the theoretical derivations of this paper. Theorem (ref) is also helpful to a broad range of studies, such as those mentioned in HLG20, FanYao03, Gao2007, LR2007, hansen2008uniform, wang2014uniform, and li2016uniform.

On the Trending Term --- \texorpdfstring{$\bm{\mu}_t$}

As modelling time--varying means is always an important task in time series analysis (e.g., wu2007inference,friedrich2020autoregressive), we infer $\bm{\mu}_t$ of the model (ref) below. Up to this point, we have not imposed any specific form on the components $\bm{\mu}_t $ and $\bm{B}_{j,t}$ of $\bm{x}_t$. To carry on with our investigation, we suppose further that $\bm{\mu}_t = \bm{\mu}(\tau_t)$ and $\bm{B}_{j,t}=\bm{B}_{j}(\tau_t)$ with $\tau_t=t/T$, so (ref) can be expressed by

equation[equation omitted — 111 chars of source]

The challenge then lies in the fact that “residuals” are time--varying linear processes. Some detailed explanations can be found in dahlhaus2012locally.

The following assumptions are necessary for the development of our trend estimation.

assumption$\bm{\mu}(\cdot)$ and $\bm{B}_{j}(\cdot)$ are $d \times 1$ vector and $d \times d$ matrix respectively. Each functional component of $\bm{\mu}(\cdot)$ and $\bm{B}_{j}(\cdot)$ is second order continuously differentiable on $[0, 1]$. Moreover, $\sup_{\tau\in [0.1]} \sum_{j=1}^{\infty} j \|\bm{B}_j^{(\ell)}(\tau) \| < \infty$ for $\ell =0,1$.
assumptionLet $K(\cdot)$ be a symmetric probability kernel function and Lipschitz continuous on $[-1,1]$. Also, let $h \to 0$ and $Th\to \infty$ as $T\rightarrow \infty$.

Assumption (ref) can be considered as a stronger version than Assumption (ref). Assumption (ref) is standard in the literature of kernel regression (LR2007).

For $\forall\tau\in (0,1)$, we recover $\bm{\mu}(\tau)$ by the next estimator.

equation[equation omitted — 142 chars of source]

We are now ready to establish an important and useful theorem.

theoremLet Assumptions (ref)--(ref) hold. For $\forall\tau \in (0,1)$, as $T\to \infty$, \begin{equation*} \sqrt{Th}\left(\widehat{\bm{\mu}}(\tau)-\bm{\mu}(\tau)-\frac{1}{2}h^2\tilde{c}_2{\bm{\mu}}^{(2)}(\tau)\right) \to_D N\left(\bm{0}_{d\times 1},\bm{\Sigma}_{\bm{\mu}}(\tau)\right), \end{equation*} where $\bm{\Sigma}_{\bm{\mu}}(\tau) =\tilde{v}_0 \left(\sum_{j=0}^{\infty}\bm{B}_j(\tau) \, \sum_{j=0}^{\infty}\bm{B}_j^\top(\tau)\right)$, and $\tilde{c}_2$ and $\tilde{v}_0$ are defined in the end of Section (ref).

Note that $\bm{\Sigma}_{\bm{\mu}}(\cdot)$ is the long--run covariance matrix, which in general cannot be estimated directly. To construct confidence intervals practically, we use a dependent wild bootstrap (DWB) method which is initially proposed by shao2010dependent for stationary time series. For the sake of space, the detailed procedure with the associated asymptotic properties is presented in Appendix (ref).

Time--Varying VAR

\setcounter{equation}{0}

In this section, we pay particular attention to one of the most popular models of the VMA$(\infty)$ family --- VAR. Many multivariate time series exhibit time--varying simultaneous interrelationships and changes in unconditional volatility justiniano2008time,coibion2011monetary. Along this line, time--varying VAR models have proven to be especially useful for describing the dynamics of the multivariate time series. Majority time--varying VAR models are investigated under the Bayesian framework, while little has been done using a frequentists' approach. Building on Section (ref), we consider a time--varying VAR model under the nonparametric framework, and establish the corresponding estimation theory.

Suppose that we observe $\left(\bm{x}_{-p+1},\ldots,\bm{x}_0,\bm{x}_1,\ldots,\bm{x}_T\right)$ from the following data generating process. Accounting for heteroscedasticity, we consider the next model.

equation[equation omitted — 179 chars of source]

where $\bm{\omega}(\tau)$ is a matrix of unknown functions of $\tau$. The model (ref) allows dynamic variations for both the coefficients and the covariance matrix. We infer $\bm{a}(\cdot)$ and $ \bm{A}_{j}(\cdot)$'s below, which are respectively $d\times 1$ vector and $d\times d$ matrices of unknown smooth functions. In addition, we are interested in the $d\times d$ dimension $\bm{\omega}(\cdot)$ which governs the dynamics of the covariance matrix of $\{\bm{\eta}_t\}_{t=1}^T$. As mentioned in primiceri2005time, allowing $\bm{\omega}(\cdot)$ to vary over time is important theoretically and practically, because a constant covariance matrix implies that the shock to the $i$th variable of $\bm{x}_{t}$ has a time--invariant effect on the $j$th variable of $\bm{x}_{t}$, restricting simultaneous interactions among multiple variables to be time--invariant.

Estimation Method and Asymptotic Theory

To facilitate the development, we first impose the following conditions.

assumption\begin{enumerate} • The roots of $\bm{I}_d-\bm{A}_{1}(\tau)L-\cdots -\bm{A}_{p}(\tau)L^p=\bm{0}_d$ all lie outside the unit circle uniformly in $\tau \in [0,1]$. • Each element of $\bm{A}(\tau)=\left(\bm{a}(\tau),\bm{A}_1(\tau),\ldots,\bm{A}_{p}(\tau)\right)$ is second order continuously differentiable on $[0,1]$ and $\bm{A}(\tau)=\bm{A}(0)$ for $\tau<0$. • Each element of $\bm{\omega}(\tau)$ is second order continuously differentiable on $[0,1]$. Moreover, $\bm{\Omega}(\tau)=\bm{\omega}(\tau)\bm{\omega}(\tau)^\top$ is positive definite uniformly in $\tau \in [0,1]$ and $\bm{\omega}(\tau)=\bm{\omega}(0)$ for $\tau<0$. \end{enumerate}

Assumption (ref).1 ensures that model (ref) is neither unit--root nor explosive, while Assumption (ref).2 allows the underlying data generate process to evolve over time in a smooth manner. In addition, the conditions $\bm{A}(\tau)=\bm{A}(0)$ and $\bm{\omega}(\tau)=\bm{\omega}(0)$ for $\tau<0$ yield

equation[equation omitted — 123 chars of source]

for all $t\leq0$, which ensures (ref) behaves like a stationary parametric VAR$(p)$ model for $t\le 0$. Similar treatments can be found in vogt2012nonparametric for the univariate setting. With the above assumptions in hand, the following proposition shows that model (ref) can be approximated by a time--varying VMA$(\infty)$ process satisfying Assumption (ref).

propositionUnder Assumption (ref), there exists a VMA$(\infty)$ process \begin{eqnarray} \widetilde{\bm{x}}_t= \bm{\mu}(\tau_t)+\bm{B}_0(\tau_t)\bm{\epsilon}_t+\bm{B}_1(\tau_t)\bm{\epsilon}_{t-1}+\bm{B}_2(\tau_t)\bm{\epsilon}_{t-2}+\cdots \end{eqnarray} such that $\max_{t\geq 1} E\left\|\bm{x}_t-\widetilde{\bm{x}}_t\right\|=O(T^{-1})$, where $\bm{\mu}(\tau)=\bm{a}(\tau)+\sum_ {j=1}^{\infty}\bm{\Psi}_j(\tau)\bm{a}(\tau)$, $\bm{B}_0(\tau)=\bm{\omega}(\tau)$, $\bm{B}_j(\tau)=\bm{\Psi}_j(\tau)\bm{\omega}(\tau)$, $\bm{\Psi}_j(\tau)=\bm{J}\bm{\Phi}^j(\tau) \bm{J}^\top$ for $j\geq 1$, $\bm{\Phi}(\cdot)$ is defined as follows: \begin{eqnarray} \bm{\Phi}(\tau)=\left(\begin{matrix} \bm{A}_{1}(\tau) & \cdots & \bm{A}_{p-1}(\tau) & \bm{A}_{p}(\tau) \\ \bm{I}_d & \cdots& \bm{0}_d & \bm{0}_d\\ \vdots & \ddots&\vdots & \vdots\\ \bm{0}_d &\cdots &\bm{I}_d & \bm{0}_d\\ \end{matrix} \right), \end{eqnarray} and $\bm{J}=\left[\bm{I}_d,\bm{0}_{d\times d(p-1)}\right]$. Moreover, $\bm{\mu}(\cdot)$ and $\bm{B}_j(\cdot)$ fulfil Assumption (ref).

Under Assumption (ref), when $\tau_t$ is in a small neighbourhood of $\tau$, we have

eqnarray[eqnarray omitted — 102 chars of source]

where $\bm{Z}_{t-1}= \bm{z}_{t-1}\otimes \bm{I}_d$ and $\bm{z}_{t-1}=(1,\bm{x}_{t-1}^\top,\ldots,\bm{x}_{t-p}^\top)^\top$. The estimators of $\bm{A}(\tau)$ and $\bm{\Omega}(\tau)$ are then sequentially given by

eqnarray[eqnarray omitted — 379 chars of source]

where $\bm{\widehat{\eta}}_t=\bm{x}_t-\bm{\widehat{A}}(\tau_t)\bm{z}_{t-1}$. The asymptotic properties associated with (ref) are summarized in the next theorem.

theoremLet Assumptions (ref), (ref) and (ref) hold. Suppose further that $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$ as $T\rightarrow \infty$ and $\max_{t\geq1} E (\|\bm{\epsilon}_t \|^4 |\mathcal{F}_{t-1} ) < \infty $ a.s. Then the following results hold. \begin{enumerate} • $\sup_{\tau \in [h,1-h]} \| \bm{\widehat{A}}(\tau)-\bm{A}(\tau) \|=O_P \left(h^2+ (\frac{\log T}{Th} )^{1/2} \right)$; • In addition, conditional on $\mathcal{F}_{t-1}$, the third and fourth moments of $\bm{\epsilon}_t$ are identical to the corresponding unconditional moments a.s.. For $\forall\tau \in (0,1)$, \begin{equation*} \sqrt{Th}\left(\begin{matrix} \mathrm{vec}\left(\bm{\widehat{A}}(\tau)-\bm{A}(\tau)-\frac{1}{2}h^2\tilde{c}_2\bm{A}^{(2)}(\tau)\right) \\ \mathrm{vech}\left(\bm{\widehat{\Omega}}(\tau)-\bm{\Omega}(\tau)-\frac{1}{2}h^2\tilde{c}_2\bm{\Omega}^{(2)}(\tau)\right) \end{matrix} \right)\to_D N\left(\bm{0},\bm{V}(\tau)\right), \end{equation*} where $\bm{V}(\tau) $ is defined in (ref) for the sake of presentation. • $ \widehat{\bm{V}}(\tau)\to_P \bm{V}(\tau)$, where $\widehat{\bm{V}}(\tau)$ is defined in (ref) for the sake of presentation. \end{enumerate}

The first result of Theorem (ref) provides the uniform convergence rate for $\bm{\widehat{A}}(\tau)$. As a consequence, it allows us to establish a joint asymptotic distribution for the estimates of the coefficients and innovation covariance in the second result. For $\delta>5$, the usual optimal bandwidth $h_{opt}=O\left(T^{-1/5}\right)$ satisfies the condition $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$. The third result ensures the confidence interval can be constructed practically.

Before moving on to impulse responses, we consider a practical issue --- the choice of the lag $p$, which is usually unknown in practice and needs to be decided by the data. We select the number of lags by minimizing the next information criterion:

eqnarray[eqnarray omitted — 131 chars of source]

where $\text{IC}(\mathsf{p})=\log \left\{\text{RSS}(\mathsf{p})\right\}+\mathsf{p}\cdot\chi_T$, $\text{RSS}(\mathsf{p})=\frac{1}{T}\sum_{t=1}^{T}\widehat{\eta}_{\mathsf{p},t}^\top \widehat{\eta}_{\mathsf{p},t}$, $\chi_T$ is the penalty term, and $\mathsf{P}$ is a sufficiently large fixed positive integer. The next theorem shows the validity of (ref).

theoremLet Assumptions (ref), (ref) and (ref) hold. Suppose $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$, $\max_{t\geq1} E (\|\bm{\epsilon}_t \|^4 |\mathcal{F}_{t-1} ) < \infty $ a.s., $\chi_T\to 0$, and $(c_T\phi_T)^{-1}\chi_T\to \infty$, where $c_T=h^2+\left(\frac{\log T}{Th}\right)^{1/2}$ and $\phi_T=h+\left(\frac{\log T}{Th}\right)^{1/2}$. Then $\Pr\left(\widehat{\mathsf{p}}=p\right)\to 1$ as $T\to \infty$.

In view of the conditions on $\chi_T$, a natural choice is

eqnarray*[eqnarray* omitted — 120 chars of source]

In the supplementary Appendix (ref), we conduct intensive simulations to examine the finite sample performance of the above information criterion.

Impulse Response Functions

We now focus on the impulse responses below, which capture the dynamic interactions among the variables of interest in a wide range of practical cases.

By Proposition (ref), the impulse response functions of $\bm{x}_t$ is asymptotically equivalent to those of $\widetilde{x}_t$. Hence, recovering the impulse response functions requires estimating $\bm{\Psi}_j(\cdot)$'s and $\bm{\omega}(\cdot)$, which is then down to the estimation of $\bm{\Phi}(\cdot)$ and $\bm{\omega}(\cdot)$ by construction. Note further that $\bm{\Phi}(\cdot)$ is a matrix consisting of the coefficients of (ref). The estimator of $\bm{\Phi}(\cdot)$ is intuitively defined as $\widehat{\bm{\Phi}}(\cdot)$, in which we replace $\bm{A}_j(\cdot)$ of (ref) with the corresponding estimator obtained from (ref). Furthermore, we require $\bm{\omega}(\tau)$ to be a lower--triangular matrix in order to fulfil the identification restriction. Thus, $\widehat{\bm{\omega}}(\tau)$ is chosen as the lower triangular matrix from the Cholesky decomposition of $\bm{\widehat{\Omega}}(\tau)$ such that $\bm{\widehat{\Omega}}(\tau)=\bm{\widehat{\omega}}(\tau)\bm{\widehat{\omega}}^\top(\tau)$.

With the above notation in hand, we are now ready to present the estimator of $\bm{B}_j(\tau) $ by

eqnarray*[eqnarray* omitted — 97 chars of source]

where $\bm{\widehat{\Psi}}_j(\tau)=\bm{J} \bm{\widehat{\Phi}}^j(\tau) \bm{J}^\top$. The corresponding asymptotic results are summarized in the following theorem.

theoremLet Assumptions (ref), (ref) and (ref) hold, and let $T\to \infty$. Suppose further that $\frac{T^{1-\frac{4}{\delta}}h}{\log T} \to \infty$, $\max_{t\geq1} E\left(\left\|\bm{\epsilon}_t \right\|^4 |\mathcal{F}_{t-1}\right) < \infty $ a.s. and conditional on $\mathcal{F}_{t-1}$, the third and fourth moments of $\bm{\epsilon}_t$ are identical to the corresponding unconditional moments a.s.. Then for any fixed integer $j\ge 0$ \begin{eqnarray*} \sqrt{Th}\left(\mathrm{vec}\left(\bm{\widehat{B}}_j(\tau)-\bm{B}_j(\tau)\right)-\frac{1}{2}h^2\tilde{c}_2\bm{B}_j^{(2)}(\tau)\right)\to_D N\left(0,\bm{\Sigma}_{\bm{B}_j}(\tau)\right), \end{eqnarray*} where \begin{eqnarray*} \bm{\Sigma}_{\bm{B}_j}(\tau)&=&[\bm{C}_{j,1}(\tau),\bm{C}_{j,2}(\tau)]\bm{V}(\tau)[\bm{C}_{j,1}(\tau),\bm{C}_{j,2}(\tau)]^\top,\\ \bm{B}_j^{(2)}(\tau)&=&\bm{C}_{j,1}(\tau)\mathrm{vec}\left(\bm{A}^{(2)}(\tau)\right) +\bm{C}_{j,2}(\tau)\mathrm{vech}\left(\bm{\Omega}^{(2)}(\tau)\right),\\ \bm{C}_{0,1}(\tau)&=&0,\quad \bm{C}_{0,2}(\tau)= \bm{L}_d^\top\left(\bm{L}_d(\bm{I}_{d^2}+\bm{K}_{dd})(\bm{\omega}(\tau)\otimes \bm{I}_d)\bm{L}_d^\top \right)^{-1},\\ \bm{C}_{j,1}(\tau)&=&(\bm{\omega}^\top(\tau)\otimes \bm{I}_d) \left(\sum_{m=0}^{j-1} \bm{J}(\bm{\Phi}^\top(\tau))^{j-1-m}\otimes (\bm{J} \bm{\Phi}^m(\tau)\bm{J}^\top)\right)\cdot [\bm{0}_{d^2p\times d},\bm{I}_{d^2p}],\ j \geq 1,\\ \bm{C}_{j,2}(\tau)&=&(\bm{I}_d\otimes(\bm{J}\bm{\Phi}^j(\tau)\bm{J}^\top)) \bm{L}_d^\top\left(\bm{L}_d(\bm{I}_{d^2}+\bm{K}_{dd})(\bm{\omega}(\tau)\otimes \bm{I}_d)\bm{L}_d^\top \right)^{-1},\ j \geq 1, \end{eqnarray*} in which the elimination matrix $\bm{L}_d$ satisfies that $\mathrm{vech}(\bm{F})=\bm{L}_d\mathrm{vec}(\bm{F})$ for any $d\times d$ matrix $\bm{F}$, and the commutation matrix $\bm{K}_{mn}$ satisfies that $\bm{K}_{mn}\mathrm{vec}(\bm{G})=\mathrm{vec}(\bm{G}^\top)$ for any $m\times n$ matrix $\bm{G}$.

To close this section, we comment on how to construct the confidence interval. Since $\widehat{\bm{\Phi}}(\tau)\to_P \bm{\Phi}(\tau)$, $\widehat{\bm{\omega}}(\tau)\to_P\bm{\omega}(\tau)$ and $\widehat{\bm{V}}(\tau)\to_P \bm{V}(\tau)$ by Theorem (ref), it is straightforward to have $\widehat{\bm{\Sigma}}_{\bm{B}_j}(\tau)\to_P\bm{\Sigma}_{\bm{B}_j}(\tau)$, where $\widehat{\bm{\Sigma}}_{\bm{B}_j}(\tau)$ has a form identical to $\bm{\Sigma}_{\bm{B}_j}(\tau)$ but replacing $\bm{\Phi}(\tau)$, $\bm{\omega}(\tau)$ and $ \bm{V}(\tau)$ with their estimators, respectively.

We next show in Section (ref) about how to apply the proposed model and estimation method to an empirical data. Our findings show that the estimated coefficient matrices and impulse response functions capture various time--varying features.

Empirical Study

\setcounter{equation}{0}

In this section, we study the transmission mechanism of the monetary policy, and infer the long--run level of inflation (i.e., trend inflation) and the natural rate of unemployment (NAIRU). The trend inflation and NAIRU are of central position in setting monetary policy since the Federal Reserve Bank aims to mitigate deviations of inflation and unemployment from their long--run targets. See stock2016core for more relevant discussions.

As well documented, inflation is higher and more volatile during 1970--1980, but substantially decreases in the subsequent period, which is often referred to as the Great Moderation (primiceri2005time). The literature has considered two main classes of explanations: bad policy or bad luck. The first type of explanations focuses on the changes in the transmission mechanism cogley2005drifts, while the second regards it as a consequence of changes in the size of exogenous shocks sims2006. In what follows, we revisit the arguments associated with the Great Moderation using our approach. Also, we use the VMA$(\infty)$ representation of the VAR$(p)$ model to infer the path of trend inflation and NAIRU over time.

In--Sample Study

First, we estimate the time--varying VAR$(p)$ model using three commonly adopted macroeconomic variables of the literature (primiceri2005time,cogley2010inflation), which are the inflation rate (measured by the 100 times the year--over--year log change in the GDP deflator), the unemployment rate, representing the non--policy block, and the interest rate (measured by the average value for the Federal funds rates over the quarter), representing the monetary policy block. To isolate the monetary policy shocks, the interest rate is ordered last in the VAR model, and is treated as the monetary policy instrument. The identification requirement is that the monetary policy actions affect the inflation and the unemployment with at least one period of lag primiceri2005time. The data are quarterly observations measured at an annual rate from 1954:Q3 to 2020:Q1, which are taken from the Federal Reserve Bank of St. Louis economic database. Figure (ref) plots the three macro variables.

The order of the VAR$(p)$ model and the optimal bandwidth are determined by the information criterion (ref) and the cross validation criterion (ref), respectively. We obtain $\widehat{\mathsf{p}}=3$ and $\widehat{h}_{cv}=0.435$. In the literature, the lag length is often assumed to be known with the values varying from $2$ to $4$, while our data--driven method indicates that 3 is the optimal value.

We first consider measuring the changes in the size of exogenous shocks. Figure (ref) plots the estimated volatilities of the innovations as well as the associated 95% confidence intervals. The figure exhibits evidence for a general decline in unconditional volatilities. Our results thus support the “bad luck” explanations primiceri2005time,sims2006.

We then consider the responses of the inflation to the monetary policy shocks. Figure (ref) plots the time--varying impulse responses of the inflation to a structural monetary shock as well as the associated 95% confidence intervals. It is clear that the confidence intervals are much wider at the beginning of the sample period implying higher uncertainty before 1970. On the other hand, the structural responses of the inflation seem to be statistically insignificant from 1970 to 2010. Thus, we conclude that the monetary shocks have less influence on the inflation before and during the period of the Great Moderation.

Finally, we investigate the trend inflation and the NAIRU. petrova2019quasi considers a Bayesian time--varying VAR(2) model, and induces the long--run mean of $\bm{x}_t$ by $\bm{\mu}_t=\lim_{p\rightarrow\infty}E_t(\bm{x}_{t+p})=(\bm{I}-\bm{A}_t)^{-1}\bm{a}_t$, where $\bm{a}_t$ is the intercept term and $\bm{A}_t$ is the autoregressive coefficients. The main difference between our method and Petrova's method is that we invert time--varying VAR$(p)$ model to the time--varying MA$(\infty)$ model, and then explicitly estimate the underlying trends of the inflation and the NAIRU using model (ref).

Figure (ref) plots the estimates of the trend inflation and the NAIRU. It is obvious that the underlying trend of the inflation is high in the 1970s, but decreases in the subsequent period. After the Great Moderation, the long--run level of the inflation is below, but quite close to the Federal Reserve's target of 2%, which indicates that the inflation is more anchored now than in the 1970s. However, the NAIRU is less persistent and fluctuates over time.

Out--of--Sample Forecasting

In this subsection, we focus on the out--of--sample forecasting, and compare the forecasts of a Bayesian time--varying parameter VAR with stochastic volatility (TVP--SV) primiceri2005time, as well as a VAR model with constant parameters (CVAR).

Specifically, we consider the 1--8 quarters ahead forecasts. That is to forecast

eqnarray*[eqnarray* omitted — 77 chars of source]

for $h=1$, $2$, $4$ and $8$, where $\bm{x}_t$ includes the values of the three aforementioned macro variables at the date $t$. The forecasts of $\overline{\bm{x}}_{t+1|t+h}$ are constructed by $\widehat{\overline{\bm{x}}}_{t+1|t+h}=\widehat{\bm{A}}_t\bm{z}_t$, in which $\widehat{\bm{A}}_t$ is estimated from CVAR, TVP--SV and the time--varying VAR model using the available data at time $t$ and $\bm{z}_t=[1,\bm{x}_{t}^\top,\bm{x}_{t-1}^\top,\bm{x}_{t-2}^\top]^\top$. The expanding window scheme is adopted. For comparison, we compute the root of mean square errors (RMSE) for CVAR as a benchmark, and the RMSE ratios for others. The out--of--sample forecast period covers 1985:Q2--end, about 35 years.

The forecasting results are presented in Table (ref), in which the values represent the ratios of the RMSEs of the corresponding method over the RMSEs of the benchmark method (i.e., CVAR). The result shows that the time--varying VAR model and TVP--SV perform much better than CVAR, which implies the desirability of introducing variations in forecasting models. In addition, the time--varying VAR model has a better forecasting performance than TVP--SV with the increase of the forecast horizon.

{

table[table omitted — 1,140 chars of source]

}

figure[figure omitted — 196 chars of source]
figure[figure omitted — 387 chars of source]
figure[figure omitted — 234 chars of source]
figure[figure omitted — 208 chars of source]

Conclusion

\setcounter{equation}{0}

This paper has proposed a class of time--varying VMA($\infty$) models, which nest, for instance, time--varying VAR, time--varying VARMA, and so forth as special cases. Both the estimation methodology and asymptotic theory have been established accordingly. In the empirical study, we have investigated the transmission mechanism of monetary policy using U.S. data, and uncover a fall in the volatilities of exogenous shocks. Our findings include (i) monetary policy shocks have less influence on inflation before and during the so--called Great Moderation, (ii) inflation is more anchored recently, and (iii) the long--run level of inflation is below, but quite close to the Federal Reserve's target of two percent after the beginning of the Great Moderation period. In addition, in Appendix B of the online supplementary material, we have evaluated the finite--sample performance of the proposed model and estimation theory.

There are several directions for possible extensions. The first one is to test about whether the $d$--dimensional components of the {\rm VAR}$(p)$ process is cross--sectionally independent for the case where the dimensionality, $d$, and the number of lags, $p$, may diverge along with the sample size, $T$. The second one is about model specification testing to check whether some of the time--varying coefficient matrices, $\bm{A}_j(\tau)$, may just be constant matrices, $\bm{A}_{j0}$. Existing studies by Gao08, pgy14, and cw19 may be useful for both issues. The third one is to allow for some cointegrated structure in our settings. The recent work by zry19 provides us with a good reference. We wish to leave such issues for future study.