EconBase
← Back to paper

Identification and estimation of structural vector autoregressive models via LU decomposition

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

51,710 characters · 11 sections · 17 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and estimation of structural vector autoregressive models via LU decomposition

\numberwithin{equation}{section} \theoremstyle{plain} \newtheorem{thm}{Theorem}[section]

\newtheorem{lem}[thm]{Lemma} \newtheorem{prop}[thm]{Proposition} \theoremstyle{definition} \newtheorem{defi}[thm]{Definition} \newtheorem{assumption}[thm]{Assumption} \newtheorem{cor}[thm]{Corollary} \newtheorem{rem}[thm]{Remark} \newtheorem{eg}[thm]{Example}

\affil[1, 2]{Faculty of Economics and Law, Shinshu University.}

abstractStructural vector autoregressive (SVAR) models are widely used to analyze the simultaneous relationships between multiple time-dependent data. Various statistical inference methods have been studied to overcome the identification problems of SVAR models. However, most of these methods impose strong assumptions for innovation processes such as the uncorrelation of components. In this study, we relax the assumptions for innovation processes and propose an identification method for SVAR models under the zero-restrictions on the coefficient matrices, which correspond to sufficient conditions for LU decomposition of the coefficient matrices of the reduced form of the SVAR models. Moreover, we establish asymptotically normal estimators for the coefficient matrices and impulse responses, which enable us to construct test statistics for the simultaneous relationships of time-dependent data. The finite-sample performance of the proposed method is elucidated by numerical simulations. We also present an example of an empirical study that analyzes the impact of policy rates on unemployment and prices.

Introduction

Structural vector autoregressive (SVAR) models studied by Sims1980, Bernanke1986, and Blanchard1988 are widely used to analyze time-dependent macroeconomic data. A basic SVAR model is expressed as follows:

equation[equation omitted — 152 chars of source]

where $\bm{\mu} \in \mathbb{R}^k$ is an intercept, $\bm{A}_s, s=0,1,\ldots,p$ are $k \times k$ coefficient matrices, and $\{\bm{v}_t\}_{t \in \mathbb{Z}}$ is an innovation process. The term $\bm{A}_0 \bm{Y}_t$ in the right-hand side represents the simultaneous relationships between the components of $\bm{Y}_t$. When $\bm{I}_k - \bm{A}_0$ is non-singular, we have the following ordinal vector autoregressive (VAR) representation:

equation[equation omitted — 137 chars of source]

where \[ \bm{\eta} = \bm{Q} \bm{\mu},\quad \bm{B}_s = \bm{Q} \bm{A}_s,\quad s=1,\ldots,p, \] with $\bm{Q} = (\bm{I}_k - \bm{A}_0)^{-1}$, and $\bm{e}_t = \bm{Q} \bm{v}_t$. We call (ref) the reduced form of (ref). Under some assumptions such as the stationarity of the process $\{\bm{Y}_t\}_{t \in \mathbb{Z}}$, we can construct an asymptotically normal estimator for $\bm{B}_s, s=1,\ldots,p$ using methods such as ordinary least squares (OLS). Meanwhile, some restrictions are required to recover $\bm{A}_s, s=0,\ldots,p$ from observable structures. Such identification problems have been discussed by several researchers. Typically, we consider zero restrictions, e.g., some specific components of $\bm{A}_s, s=0,1,\ldots,p$ are fixed to zero. Under some additional assumptions, a general identification method based on zero restrictions was proposed by Rubio-Ramirez2010. Canova2002 and Uhlig2005 proposed identification methods for SVAR models under the sign restriction of the coefficient matrices, i.e., the signs of some specific impulse responses are known. Hyvarinen2010 and Lanne2017 relaxed the restrictions on the coefficient matrices and developed other methods for non-Gaussian processes to allow flexible identification of structures.

To correctly interpret the analysis of SVAR models, it is necessary to impose appropriate restrictions based on prior knowledge and data background. Regarding this point, previous studies have focused on the restrictions on the coefficient matrices. Meanwhile, it is often assumed that $\mathop{\rm Var}\nolimits[\bm{v}_t]$ is a diagonal matrix. However, the components of $\bm{v}_t$ might be correlated if $\bm{v}_t$ includes some unobservable exogenous variables and does not correspond to the unique shocks of $\bm{Y}_t$. Because an invalid assumption may lead to misunderstandings in causal interpretations, it is imperative to consider other identification methods under less restrictive conditions on the innovation process.

In this study, under mild conditions on the innovation process, where $\mathop{\rm Var}\nolimits[\bm{v}_t]$ is allowed to be a non-diagonal matrix, we propose an identification method based on zero restrictions on the coefficient matrices and LU decomposition for a sub-matrix of $\bm{B} = (\bm{\eta}, \bm{B}_1,\ldots,\bm{B}_p)$.

Moreover, we establish asymptotically normal estimators for the coefficient matrices $\bm{A}_s, s=0,\ldots,p$ and impulse responses, enabling us to construct test statistics for the hypothesis testing whose null hypothesis is that $\mathcal{H}_0: \bm{A}_0 = \bm{O}$, which can be used to verify whether the simultaneous relationships should be considered for the data $\{\bm{Y}_t\}_{t \in \mathbb{Z}}$.

The remainder of this article is organized as follows. In Section (ref), we describe the model setup of SVAR models and present a motivational example. In Section (ref), we propose the identification and estimation methods for $\bm{A}_s, s=0,1,\ldots,p$ via LU decomposition under appropriate restrictions. We also consider the impulse response estimation and hypothesis testing for $\bm{A}_0$ in this section. The numerical simulations used to verify the asymptotic behavior of the estimators and test statistics are presented in Section (ref). We further apply the proposed methods to analyze the impact of policy rates on employment and prices in this section. The proofs of the main theoretical results are presented in Section (ref). We discuss the causal interpretations of the statistical inference for SVAR models in the Appendix.

SVAR models

Model setup and a motivational example

We consider the following model:

equation[equation omitted — 146 chars of source]

where $\bm{A}_s \in \mathbb{R}^{k \times k}, s=0,1,\ldots, p$ are the coefficient matrices and $\{\bm{v}_t\}_{t \in \mathbb{Z}}$ is an i.i.d. innovation process such that $\,\mathrm{E}[\bm{v}_t] = \bm{0}$. We suppose that $\bm{A}_0$ is a $k \times k$ lower-triangular matrix with all diagonal components being zero, which means that $Y_{i, t}$ cannot be the direct cause of $Y_{j, t}$ for $i >j$. \if0 The innovation process $\{\bm{v}_t\}_{t \in \mathbb{Z}}$ can be regarded as direct causes of $\bm{Y}_t$ which cannot be written by components of $\bm{Y}_{t-s}, s=0,\ldots,p$. \fi Several researchers have assumed that $\mathop{\rm Var}\nolimits[\bm{v}_t]$ is a diagonal matrix. Meanwhile, we consider the existence of contemporaneous confounding, indicating that $\bm{v}_{i, t}$ and $\bm{v}_{j, t}, i \neq j$ are correlated. We now consider the following motivational example.

egLet $\{{W}_t\}_{t \in \mathbb{Z}}$ be an $\mathbb{R}$-valued unobservable i.i.d. sequence. We consider the following model: \begin{equation} \bm{Y}_t = \bm{A}_0 \bm{Y}_t + \sum_{s=1}^{p} \bm{A}_{s} \bm{Y}_{t-s} + \bm{A}_W{W}_t + \bm{u}_t,\quad t \in \mathbb{Z}, \end{equation} where $\bm{A}_s, s=1,\ldots,p \in \mathbb{R}^{3 \times 3}$, $\bm{A}_0 \in \mathbb{R}^{3 \times 3}$ is a lower-triangular matrix with diagonal elements being zero, and $\{\bm{u}_t\}$ is an $\mathbb{R}^3$-valued i.i.d. sequence independent of ${W}_t$ with $\mathop{\rm Var}\nolimits[\bm{u}_t] = \sigma_u^2 \bm{I}_3$, and $\bm{A}_W \in \mathbb{R}^{3 \times 1}$. Thus, we can rewrite the model (ref) as follows: \[ \bm{Y}_t = \bm{A}_0 \bm{Y}_t + \sum_{s=1}^{p} \bm{A}_{s} \bm{Y}_{t-s} + \bm{v}_t,\quad t \in \mathbb{Z}, \] where \[ \bm{v}_t =\bm{A}_W{W}_t + \bm{u}_t. \] Therefore, $\mathop{\rm Var}\nolimits[\bm{v}_t]$ is not diagonal unless $\bm{A}_W \bm{A}_W^\top$ is diagonal. The SVAR model can be regarded as a special case of linear structural equation models, as described in the Appendix. For a linear structural equation model, we often consider a graphical representation. Suppose that we observe $\bm{Y}_{1-p},\ldots,\bm{Y}_T$ for some $T \in \mathbb{N}$ and consider graph $\bm{G}$ with the vertex set $\tilde{\bm{V}} = \{Y_{i,s} : i =1,\ldots,k, s =1-p,\ldots,T\} \cup \{W_{s} : s =1-p,\ldots,T \}$. \begin{figure}[H] \caption{A graphical representation for the model (ref) with $p=1$} \end{figure} Figure (ref) shows graph $\bm{G}$, where the edges represent the nonzero components of $\bm{A}_0, \bm{A}_1$ and $\bm{A}_W$. Since the model structure does not depend on time, Figure (ref) only shows the components of $\bm{Y}_t$, their direct causes, $W_t$, and $W_{t-1}$ included in the vertex set. Observe that $W_t, t=1-p,\ldots,T$ are (unobservable) contemporaneous confounding factors, which implies correlations between the components of the innovation process $\{\bm{v}_t\}_{t \in \mathbb{Z}}$. If we suppose that the covariance matrix of the innovation process is non-diagonal, we cannot deal with such data using most of the conventional SVAR models.

If the matrix $\bm{I}_k - \bm{A}_0$ is non-singular, then we have the reduced form of (ref) as follows:

equation[equation omitted — 131 chars of source]

where \[ \bm{Q} = (\bm{I}_k - \bm{A}_0)^{-1}. \] This reduced form is a typical VAR model. Thus, to ensure the unique existence of a stationary and ergodic solution to (ref), it is sufficient to assume the following conditions.

assumption\begin{itemize} • $\bm{A}_0$ is a lower-triangle matrix with the diagonal elements being zero. • The following polynomial \[ \left| \bm{I}_{k} - \bm{Q}\bm{A}_1z - \bm{Q}\bm{A}_2z^2 - \cdots - \bm{Q}\bm{A}_pz^p \right|,\quad z \in \mathbb{C} \] is nonzero for all $z \in \{\zeta \in \mathbb{C} : |\zeta| \leq 1\}$. \end{itemize}

Condition (i) of Assumption (ref) guarantees that the model (ref) has a directed acyclic graph (DAG) structure. Next, we consider the impulse response functions. Consider the following VAR$(1)$ representation of (ref): \[ \bm{\xi}_t = \bm{\Lambda} \bm{\xi}_{t-1} + \bm{\epsilon}_t, \] where $\bm{\xi}_t = (\bm{Y}_t^\top,\ldots,\bm{Y}_{t-p+1}^\top)^\top$, $\bm{\epsilon}_t = (\bm{e}_t^\top,\bm{0}^\top,\ldots,\bm{0}^\top)^\top$, \[ \bm{\Lambda} = \left(

array[array omitted — 267 chars of source]

\right), \] and $\bm{B}_s = \bm{Q}\bm{A}_s$ for $s=1,\ldots,p$. We omit $\bm{\mu}$ here for simplicity because impulse responses do not depend on it. Under Assumption (ref), we have the following moving average (MA) representation:

equation[equation omitted — 100 chars of source]

For the top $k$ rows in (ref), there exist $\bm{\Psi}_0 = \bm{I}_k$ and $\bm{\Psi}_s \in \mathbb{R}^{k \times k}, s=1,2,\ldots$ such that \[ \bm{Y}_t = \sum_{s=0}^\infty \bm{\Psi}_s \bm{e}_{t-s} = \sum_{s=0}^\infty \bm{\Psi}_s \tilde{\bm{L}} \tilde{\bm{u}}_{t-s}, \] where $\tilde{\bm{L}} \in \mathbb{R}^{k \times k}$ is a lower-unitriangular matrix obtained by the LU decomposition of $\mathop{\rm Var}\nolimits[\bm{e}_t]$ and $\tilde{\bm{u}}_t = \tilde{\bm{L}}^{-1}\bm{e}_t$. Notably, $\bm{\Psi}_s$ coincides with the $k \times k$ sub-matrix of $\bm{\Lambda}^s$ corresponding to the first $k$ columns and rows. Therefore, the orthogonalized impulse response $\mathrm{OIRF}_{ij}(s)$ and non-orthogonalized impulse response $\mathrm{IRF}_{ij}(s)$ are, respectively, given by \[ \mathrm{OIRF}_{ij}(s) = \left( \bm{\Psi}_s \tilde{\bm{L}} \right)_{ij},\quad \mathrm{IRF}_{ij}(s) = \left( \bm{\Psi}_s \right)_{ij}. \] By considering the SVAR model as a special case of linear structural equation models, the coefficient matrices and impulse responses can be regarded as causal effects (see the Appendix for the detail of such interpretations).

Identification and statistical inference of SVAR models

Estimations of the coefficient matrices

In this section, we establish asymptotically normal estimators for the coefficient matrices of model (ref).

We begin by considering the estimation method for the reduced form (ref) of (ref). We rewrite the model as follows:

equation[equation omitted — 108 chars of source]

where $\bm{A} = (\bm{\mu}, \bm{A}_1,\ldots, \bm{A}_p)$ and $\bm{X}_{t-1} = (1, \bm{Y}_{t-1}^\top,\ldots,\bm{Y}_{t-p}^\top)^\top$. The reduced form can be represented as follows.

equation[equation omitted — 86 chars of source]

where $\bm{B} = \bm{Q} \bm{A}$ and $\bm{e}_t = \bm{Q}\bm{v}_t$. Suppose that we observe $(\bm{Y}_{1-p},\ldots,\bm{Y}_T)$. We define the data matrix $\bm{Y}$ and the design matrix $\bm{X}$ as follows: \[ \bm{Y} = (\bm{Y}_1,\ldots,\bm{Y}_T)^\top\quad \mbox{and}\quad \bm{X} = (\bm{X}_0,\ldots,\bm{X}_{T-1})^\top, \] respectively. Let $r=1+kp$ and $\bm{\Theta} \subset \mathbb{R}^{k \times r}$ be a compact parameter space of $\bm{B}$ and $\bm{B}_*$ be the true value of $\bm{B}$. Then, we consider the following least squares estimators $\hat{\bm{B}}_T$ and $\hat{\bm{b}}_T$ for $\bm{B}$ and $\,\mathrm{vec}(\bm{B})$: \[ \hat{\bm{B}}_T = \bm{Y}^\top \bm{X} (\bm{X}^\top \bm{X})^{-1} = \left( \sum_{t=1}^T \bm{Y}_t \bm{X}_{t-1}^\top \right) \left( \sum_{t=1}^T \bm{X}_{t-1} \bm{X}_{t-1}^\top \right)^{-1}, \] and \[ \hat{\bm{b}}_T = \,\mathrm{vec}(\hat{\bm{B}}_T) = \bm{b}_* + \left\{\left( \frac{1}{T} \sum_{t=1}^T \bm{X}_{t-1} \bm{X}_{t-1}^\top \right)^{-1} \otimes \bm{I}_k\right\} \left( \frac{1}{T} \sum_{t=1}^T \bm{X}_{t-1} \otimes \bm{e}_t \right), \] where $\bm{b}_* = \,\mathrm{vec}(\bm{B}_*)$. To establish the asymptotic behavior of the estimator $\hat{\bm{b}}_T$, we assume the following conditions.

assumption\begin{itemize} • It holds that \[ \,\mathrm{E}[\|\bm{v}_t\|_2^4] < \infty,\quad t \in \mathbb{Z}. \]$\mathop{\rm Var}\nolimits[\bm{v}_t]$ and $\,\mathrm{E}[\bm{X}_{t-1} \bm{X}_{t-1}^\top]$ are positive definite. • The true value $\bm{B}_*$ is an interior point in the parameter space $\bm{\Theta} \subset \mathbb{R}^{k \times r}$. \end{itemize}

Then, we have the following asymptotic normality of $\hat{\bm{b}}_T$.

propUnder Assumptions (ref) and (ref), it holds that \[ \sqrt{T}(\hat{\bm{b}}_T - \bm{b}_*) \to^d N(\bm{0}, \bm{\Gamma}^{-1} \otimes \bm{\Sigma}),\quad T \to \infty, \] where $\bm{\Sigma} = \,\mathrm{E}[\bm{e}_t \bm{e}_t^\top]$ and $\bm{\Gamma} = \,\mathrm{E}[\bm{X}_{t-1} \bm{X}_{t-1}^\top]$.

The proof can be found in, e.g., Lutkepohl2013 or Hamilton1994, therefore, we omit it.

Next, we introduce estimators for $\bm{Q}$, $\bm{A}_0$, and $\bm{A}$ using LU decomposition. Denote \[ \bm{A} = (\bm{a}_1,\bm{a}_2,\ldots,\bm{a}_r). \] where $\bm{a}_1= \bm{\mu}$. The following condition is sufficient to construct estimators for $\bm{Q}$, $\bm{A}_0$, and $\bm{A}$ via LU decomposition and derive its asymptotic behavior.

assumptionThere exists a (known) $k$-tuple of distinct positive numbers $(j_1,\ldots,j_k)$ such that $(\bm{a}_{j_1 *},\ldots,\bm{a}_{j_k *})$ is a non-singular upper-triangular matrix, where $\bm{A}_*$ is the true value of $\bm{A}$.

Let $g: \mathbb{R}^{k \times r} \to \mathbb{R}^{k \times k}$ be the map defined as follows \[ g(\bm{C}) = (\bm{c}_{j_1},\ldots,\bm{c}_{j_k}), \] where $\bm{c}_m$ is the $m$-th column of the matrix $\bm{C}$. Similarly to Rubio-Ramirez2010, we should impose zero restrictions for lower-triangular part of $g(\bm{A})$ based on the data background (see Section (ref) for a concrete example).

Note that $g(\bm{B}) = \bm{Q} g(\bm{A})$. For a nonsingular matrix $\bm{C} \in \mathbb{R}^{k \times k}$ which allows an LU-decomposition with a lower-unitriangular matrix, we introduce the following notation: \[ \bm{C} = \bm{L}(\bm{C}) \bm{U}(\bm{C}). \] Let $\bm{q}, \bm{a}_0$, and $\bm{a}$ be the vectorizations of $\bm{Q}, \bm{A}_0$, and $\bm{A}$, respectively, i.e., \[ \bm{q} = \,\mathrm{vec}(\bm{Q}),\quad \bm{a}_0 = \,\mathrm{vec}(\bm{A}_0),\quad \mbox{and} \quad \bm{a} = \,\mathrm{vec}(\bm{A}). \] We define the estimators for $\bm{q}, \bm{a}_0$, and $\bm{a}$ as follows. \[ \hat{\bm{q}}_T = f_1(\hat{\bm{b}}_T) := \,\mathrm{vec}(\bm{L}_g(\hat{\bm{B}}_T)), \] \[ \hat{\bm{a}}_{0T} = f_2(\hat{\bm{b}}_T) := \,\mathrm{vec}(\bm{I}_k - \bm{L}_g(\hat{\bm{B}}_T)^{-1}), \] and \[ \hat{\bm{a}}_{T} = f_3(\hat{\bm{b}}_T) := \left(\bm{I}_r \otimes \bm{L}_g(\hat{\bm{B}}_T)\right)^{-1}\hat{\bm{b}}_T, \] where $\bm{L}_g := \bm{L} \circ g$. The asymptotic normality of the estimators follows from the delta method.

thmSuppose that Assumptions (ref), (ref), and (ref) hold. Then, the following conditions hold: \begin{equation} \sqrt{T}(\hat{\bm{q}}_T - \bm{q}_*) \to^d N(\bm{0}, \bm{\Sigma}_1), \end{equation} \begin{equation} \sqrt{T}(\hat{\bm{a}}_{0T} - \bm{a}_{0*}) \to^d N(\bm{0}, \bm{\Sigma}_2), \end{equation} \begin{equation} \sqrt{T}(\hat{\bm{a}}_T - \bm{a}_*) \to^d N(\bm{0}, \bm{\Sigma}_3), \end{equation} as $T \to \infty$, where \[ \bm{\Sigma}_l = \bm{J}_l \bm{\Sigma}_b \bm{J}_l^\top \] with $\bm{\Sigma}_b = \bm{\Gamma}^{-1} \otimes \bm{\Sigma}$ and \[ \bm{J}_l = \left.\frac{\partial f_l(\bm{b})}{\partial \bm{b}}\right|_{\bm{b} = \bm{b}_*},\quad l=1,2,3. \]

The asymptotic covariance matrices are singular because some components are fixed at $0$ or $1$ by LU decomposition.

Estimation of impulse responses

Next, we construct estimators for the impulse response $\mathrm{IRF}_{ij}(h)$ for $i, j = 1,2,\ldots,k$ and $h>0$. The matrix $\bm{\Psi}_h = (\mathrm{IRF}_{ij}(h))_{i, j = 1,\ldots,k}$ satisfies \[ \bm{\Psi}_h = [\bm{\Lambda}^h]_k^k, \] where $[\bm{\Lambda}^h]_k^k$ is the $k \times k$ sub-matrix of $\bm{\Lambda}^h$ corresponding to the top $k$ columns and rows. Therefore, there exists a differentiable map $f_{4, h}$ such that $f_{4, h}(\bm{b}) = \bm{\psi}_h$, where $\bm{\psi}_h = \,\mathrm{vec}(\bm{\Psi}_h)$. We define an estimator for $\bm{\psi}_h$ by $\hat{\bm{\psi}}_{h T} = f_{4, h}(\hat{\bm{b}}_T)$. The following proposition was proved by lutkepohl1990.

propSuppose that assumptions (ref), (ref), and (ref) hold. For every $h >0$, it holds that \begin{equation} \sqrt{T}(\hat{\bm{\psi}}_{hT} - \bm{\psi}_{h*}) \to^d N(\bm{0}, \bm{\Sigma}_{4, h}),\quad T \to \infty, \end{equation} where $\bm{\psi}_{h*} = \,\mathrm{vec}([\bm{\Lambda}_*^h]_k^k)$, \[ \bm{\Lambda}_* =\left(\begin{array}{ccccc} \bm{B}_{1*} & \bm{B}_{2*} & \cdots & \bm{B}_{p-1*} & \bm{B}_{p*} \\ \bm{I}_k & \bm{O} & \cdots & \bm{O} & \bm{O} \\ \bm{O} & \bm{I}_k & \cdots & \bm{O} & \bm{O} \\ \vdots & \vdots & \ddots & \vdots & \vdots \\ \bm{O} & \bm{O} & \cdots & \bm{I}_k & \bm{O} \\ \end{array} \right), \] and $\bm{\Sigma}_{4, h} = \bm{J}_{4, h}\bm{\Sigma}_{\bm{b}} \bm{J}_{4, h}^\top$ with \[ \bm{J}_{4, h} = \left.\frac{\partial f_{4, h}(\bm{b})}{\partial \bm{b}}\right|_{\bm{b} = \bm{b}_*}. \]

The general expression of each component of $\bm{J}_{4, h}$ can be found in lutkepohl1990.

Let $\bm{\Psi}_{h*}^{\mathrm{o}}, h>0$ be a matrix defined by \[ \bm{\Psi}_{h*}^{\mathrm{o}} = \bm{\Psi}_{h*}\bm{Q}_{*}, \] which corresponds to a “total effect” for a linear structural equation model as described in the Appendix. We define the estimator for $\bm{\psi}_{h*, T}^{\mathrm{o}}= \,\mathrm{vec}(\bm{\Psi}_{h*}\bm{Q}_{*})$ as follows. \[ \hat{\bm{\psi}}_{hT}^{\mathrm{o}} = f_{5, h}(\hat{\bm{b}}_T) := \left( \bm{L}_g(\hat{\bm{B}}_T)^\top \otimes \bm{I}_k \right) f_{4, h}(\hat{\bm{b}}_T). \] Because the map $f_{5, h}$ for every $h$ is differentiable, we obtain the following proposition based on the delta method.

propSuppose that Assumptions (ref), (ref), and (ref) hold. Then, it holds that \begin{equation} \sqrt{T}(\hat{\bm{\psi}}_{hT}^{\mathrm{o}} - \bm{\psi}_{h *}^{\mathrm{o}}) \to^d N(\bm{0}, \bm{\Sigma}_{5, h}) ,\quad T \to \infty, \end{equation} where $\bm{\psi}_{h *}^{\mathrm{o}} = \,\mathrm{vec}(\bm{\Psi}_{h *} \bm{Q}_*)$, $\bm{Q}_*$ is the true value of $\bm{Q}$, and $\bm{\Sigma}_{5, h} = \bm{J}_{5, h}\bm{\Sigma}_{\bm{b}} \bm{J}_{5, h}^\top$ with \[ \bm{J}_{5, h} = \left.\frac{\partial f_5(\bm{b})}{\partial \bm{b}}\right|_{\bm{b} = \bm{b}_*}. \]

Tests for simultaneous relationships

In this section, we consider the following test:

equation[equation omitted — 101 chars of source]

Let $\hat{\bm{e}}_t = \bm{Y}_t - \hat{\bm{B}}_T \bm{X}_{t-1}$ and \[ \hat{\bm{\Sigma}}_{bT} = \left( \frac{1}{T} \sum_{t=1}^T \bm{X}_{t-1} \bm{X}_{t-1}^\top \right)^{-1} \otimes \left( \frac{1}{T} \sum_{t=1}^T \hat{\bm{e}}_t \hat{\bm{e}}_t^\top \right). \] The estimators for $\bm{\Sigma}_l, l=1,2,3$ and $\bm{\Sigma}_{l, h}, l=4, 5$ can be constructed as follows: \[ \hat{\bm{\Sigma}}_l = \hat{\bm{J}}_l \hat{\bm{\Sigma}}_{\bm{b}}\hat{\bm{J}}_l^\top,\quad l=1,2,3, \] and \[ \hat{\bm{\Sigma}}_{l, h} = \hat{\bm{J}}_{l, h} \hat{\bm{\Sigma}}_{\bm{b}}\hat{\bm{J}}_{l, h}^\top,\quad l=4, 5 \] with \[ \hat{\bm{J}}_{l T}=\left.\frac{\partial f_l(\bm{b})}{\partial \bm{b}}\right|_{\bm{b} = \hat{\bm{b}}_{lT}},\quad l=1,2,3 \] and \[ \hat{\bm{J}}_{l, h T}=\left.\frac{\partial f_{l, h}(\bm{b})}{\partial \bm{b}}\right|_{\bm{b} = \hat{\bm{b}}_{lT}},\quad l=4, 5. \] The following lemma is obtained from the asymptotic normality of the estimators.

lemSuppose that Assumptions (ref), (ref), and (ref) hold. For every $\bm{w}_1, \bm{w}_2 \in \mathbb{R}^{k^2}$, and $\bm{w}_3 \in \mathbb{R}^{k(1+kp)}$, let \[ z_{1T}(\bm{w}_{1}) := \sqrt{T} (\bm{w}_{1}^\top \hat{\bm{\Sigma}}_{1} \bm{w}_{1})^{-1/2} \bm{w}_{1}^\top (\hat{\bm{q}}_T - \bm{q}_*), \] \[ z_{2T}(\bm{w}_{2}) := \sqrt{T} (\bm{w}_{2}^\top \hat{\bm{\Sigma}}_{2} \bm{w}_{2})^{-1/2} \bm{w}_{2}^\top (\hat{\bm{a}}_{0T} - \bm{a}_{0*}). \] and \[ z_{3T}(\bm{w}_{3}) := \sqrt{T} (\bm{w}_{3}^\top \hat{\bm{\Sigma}}_{\bm{b}} \bm{w}_{3})^{-1/2} \bm{w}_{3}^\top (\hat{\bm{b}}_T - \bm{b}_*). \] Then, it holds that \begin{equation} z_{lT}(\bm{w}_{l}) \to^d N(0, 1), \quad T \to \infty,\quad l=1,2,3. \end{equation}

Under the null hypothesis $\mathcal{H}_0$, we have $ \bm{Q}_* = \bm{I}_k$, $\bm{A}_{0*} = \bm{O}$, and $g(\bm{B}_*) = g(\bm{A}_*)$ is a non-singular upper-triangular matrix. For a matrix $\bm{C} \in \mathbb{R}^{k\times (1+kp)}$, $g(\bm{C}) \in \mathbb{R}^{k \times k}$ is a sub-matrix of $\bm{C}$. We define the sub-vectors $\bm{q}_{\mathrm{sub}}$ and $\bm{a}_{0 \mathrm{sub}}$ of $\bm{q}$ and $\bm{a}_0$, which comprise the lower-triangular components of $\bm{Q}$ and $\bm{A}_0$, except for diagonal components, respectively. Similarly, let $\bm{\beta}_{\mathrm{sub}}$ be a sub-vector which comprises the lower-triangular components of $g(\bm{B})$, except for diagonal components. Then, we have \[ \bm{w}^\top \bm{q}_{*\mathrm{sub}} = \bm{w}^\top \bm{a}_{0*\mathrm{sub}} = \bm{w}^\top \bm{\beta}_{*\mathrm{sub}} =0 \] for every vector $\bm{w} \in \mathbb{R}^{k(k-1)/2}$. Thus, we obtain the following theorem.

thmSuppose that Assumptions (ref), (ref), and (ref) hold. For $l=1,2,3$ and $\bm{v} \in \mathbb{R}^{k(k-1)/2}$, consider the following test statistics $z_{lT}$ \[ z_{1T} = \sqrt{T} (\bm{v}^\top \hat{\bm{\Sigma}}_{1\mathrm{sub}} \bm{v})^{-1/2} \bm{v}^\top \hat{\bm{q}}_{T\mathrm{sub}}, \] \[ z_{2T} = \sqrt{T} (\bm{v}^\top \hat{\bm{\Sigma}}_{2\mathrm{sub}} \bm{v})^{-1/2} \bm{v}^\top \hat{\bm{a}}_{0 T\mathrm{sub}}, \] and \[ z_{3T} = \sqrt{T} (\bm{v}^\top \hat{\bm{\Sigma}}_{\bm{b}\mathrm{sub}} \bm{v})^{-1/2} \bm{v}^\top \hat{\bm{\beta}}_{T\mathrm{sub}}, \] where $\hat{\bm{\Sigma}}_{1\mathrm{sub}}, \hat{\bm{\Sigma}}_{2\mathrm{sub}},$ and $\hat{\bm{\Sigma}}_{\bm{b}\mathrm{sub}}$ are sub-matrices of $\hat{\bm{\Sigma}}_{1}, \hat{\bm{\Sigma}}_{2},$ and $\hat{\bm{\Sigma}}_{\bm{b}}$ corresponding to sub-vectors $\hat{\bm{q}}_{T\mathrm{sub}}$, $\hat{\bm{a}}_{0T\mathrm{sub}}$, and $\hat{\bm{\beta}}_{T\mathrm{sub}}$, respectively. \begin{itemize} • Under the null hypothesis $\mathcal{H}_0$, \[ z_{lT} \to^d N(0, 1),\quad T \to \infty,\quad l=1,2,3. \] • Assume in addition that $ \bm{v}_1^\top \bm{q}_{*}, \bm{v}_2^\top \bm{a}_{0*}, \bm{v}_3^\top \bm{b}_* \neq 0 $ under the alternative hypothesis $\mathcal{H}_1$, then, \[ \,\mathrm{P}(|z_{lT}| > c) \to 1,\quad T \to \infty,\quad l=1,2,3 \] for every $c>0$. \end{itemize}
remSimilarly to the example of the empirical study described in the next section, the number of observations of time series data that seems to satisfy stationarity tends to be small. However, SVAR models have at least $k^2p$ parameters, and we often analyze the data such that the $k^2p/T \not\asymp 0$. To handle such data, we should consider central limit theorems or Gaussian approximations under high-dimensional settings, e.g., the case where $k^2p/T \to \kappa \in (0, 1)$ as $T \to \infty$. See Koike2021, Chernozhukov2023 for the recent development of high-dimensional asymptotic theory, Fang2023 for the degenerate case, and Belloni2018 for martingale high-dimensional central limit theorems.

Numerical studies

Simulations

In this section, we verify the finite-sample performances of the proposed estimators and test statistics. Because each component of $f_l(\bm{b})$ and $f_{l, h}(\bm{b})$ for $l=1,\ldots,5$ can be represented as a rational function of the components of $\bm{b}$, we can calculate their derivatives using, e.g., “PyTorch” in Python. In the sequel, we consider a stationary model with $k=p=5$. The innovation process $\{\bm{v}_t\}_{t \in \mathbb{Z}}$ is generated by the following model: \[ \bm{v}_t = \bm{A}_W \bm{W}_t + \bm{u}_t, \] where $\{\bm{W}_t\}_{t \in \mathbb{Z}}$, $\{\bm{u}_t\}_{t \in \mathbb{Z}}$ are two and five-dimensional mutually uncorrelated i.i.d. Laplace distributed with mean $\bm{0}$ and variance $0.5$, respectively, and \[ \bm{A}_W =

pmatrix[pmatrix omitted — 104 chars of source]

. \] The true values of the coefficient matrices are shown with the heat maps (Figure (ref)), where the $x$-axis of the figure for $g(\bm{A})$ represents the column number of $\bm{A}$.

figure[figure omitted — 164 chars of source]

We calculate the proposed estimators and test statistics for $1000$ replications.

First, we evaluate the performances of the estimators $\hat{\bm{b}}_T$, $\hat{\bm{q}}_T$, $\hat{\bm{a}}_{0T}$, and $\hat{\bm{a}}_T$, using the mean of empirical bias (MB) and mean of empirical mean absolute error (MMAE). Tables (ref)--(ref) summarize the results.

table[table omitted — 371 chars of source]
table[table omitted — 373 chars of source]
table[table omitted — 373 chars of source]

We also calculate the estimators for impulse responses for $h=1, 2, 3$. The results are summarized in Table (ref)--(ref). In summary, we can observe the consistency of the proposed estimators.

table[table omitted — 520 chars of source]
table[table omitted — 520 chars of source]
table[table omitted — 520 chars of source]

Next, we show the asymptotic normality of the proposed estimators $\hat{\bm{q}}_T, \hat{\bm{a}}_{0T}, \hat{\bm{a}}_T$, and $\hat{\bm{\psi}}_{h, T}$ for $h=1,2,3$. To do this, we introduce the following random variables.

align*[align* omitted — 619 chars of source]

where $\bm{1}_{k^2}=(1,1,\ldots,1)^\top \in \mathbb{R}^{k^2}$ and $\bm{1}_{kr}=(1,1,\ldots,1)^\top \in \mathbb{R}^{kr}$. Figures (ref)--(ref) show the histograms of the random variables over 1000 replications; the $y$-axis indicates the relative frequency of the statistics over the replications and the dashed line is the density function of the standard normal distribution. Moreover, Tables (ref) and (ref) show the empirical tail probabilities of the random variables. Even for the relatively small observations $T=100$, the estimators are well approximated by the normal distribution.

figure[figure omitted — 160 chars of source]
figure[figure omitted — 160 chars of source]
figure[figure omitted — 160 chars of source]
table[table omitted — 358 chars of source]
figure[figure omitted — 166 chars of source]
figure[figure omitted — 166 chars of source]
figure[figure omitted — 166 chars of source]
table[table omitted — 362 chars of source]

Finally, we check the behaviors of the test statistics $z_{iT}, i=1,2,3$ with $\bm{v} = \bm{1}=(1,\ldots,1)^\top \in \mathbb{R}^{k(k-1)/2}$ under the alternative hypothesis $\mathcal{H}_1$. Figures (ref)--(ref) show the histograms of $z_{iT}, i=1,2,3$ under $\mathcal{H}_1$ and the dashed line indicates the density function of the standard normal distribution. Moreover, the powers of the test statistics are summarized in Table (ref). The consistency of the test also seems valid. In terms of the power, it may be better to use $z_{3T}$ than $z_{1T}$ and $z_{2T}$ for the test.

table[table omitted — 334 chars of source]
figure[figure omitted — 181 chars of source]
figure[figure omitted — 185 chars of source]
figure[figure omitted — 185 chars of source]

Empirical studies

We analyze the impact of policy rates on unemployment and prices in U.S. from 1992 to 2018, which are quarterly data. We use the seasonally adjusted unemployment level data (U.S. Bureau of Labor Statistics, Unemployment Level [UNEMPLOY], retrieved from FRED, Federal Reserve Bank of St. Louis; https://fred.stlouisfed.org/series/UNEMPLOY, February 18, 2025), the seasonally adjusted core consumer price index data (U.S. Bureau of Labor Statistics, Consumer Price Index for All Urban Consumers: All Items Less Food and Energy in U.S. City Average [CPILFESL], retrieved from FRED, Federal Reserve Bank of St. Louis; https://fred.stlouisfed.org/series/CPILFESL, February 18, 2025), and the federal funds effective rate (Board of Governors of the Federal Reserve System (US), Federal Funds Effective Rate [DFF], retrieved from FRED, Federal Reserve Bank of St. Louis; https://fred.stlouisfed.org/series/DFF, February 18, 2025). In the sequel, we modify the data and use the following notations.

itemize$Y_{1,t}$ : growth rate of the unemployment level (%). • $Y_{2,t}$ : inflation rate (growth rate of the core consumer price index) (%). • $Y_{3,t}$ : difference of federal funds effective rate (%).

Figure (ref) shows the plot of the time series $\{Y_{i, t}\}, i=1,2,3$.

figure[figure omitted — 157 chars of source]

We apply the proposed method to the modified data. We choose the lag $p=4$, which enables us to consider the correlations between the three time series up to one year ago. The estimated coefficient matrices of the reduced form are as follows (the number of observations is $T = 104$).

align*[align* omitted — 700 chars of source]

Figure (ref) shows the estimated curve of the impulse responses and their $0.95$ confidence intervals.

figure[figure omitted — 170 chars of source]

Note that lag $s=4$ means one year. Since we assume that $Y_{3,t}$ does not affect $Y_{1,t}$ and $Y_{2,t}$ (this assumption derives from the ordering of the components of $\bm{Y}_{t}$), non-orthogonalized impulse responses in Figure (ref) can be regarded as 'total effects' from Proposition (ref) in Appendix. We expect that a policy rate decrease will cause an improvement in unemployment. The upper curve of the significance interval in the left figure seems consistent with our intuition. Meanwhile, the lower curve in the right figure seems to well describe the impact of the policy rate on the prices because the prices may be depressed if the policy rate increases. However, given that the monetary policy from 1992 to 2018 had often been shifted about every three years, the long-term effects shown by estimated values are more reasonable than the edges of significance intervals. We checked the accuracy of the estimated model in terms of the standard deviation, root mean squared error, and adjusted R-squared coefficient, which are summarized in Table (ref). The results indicate that the model is applicable to fit the data.

table[table omitted — 428 chars of source]

To estimate the coefficient matrices $\bm{A}_s, s=0,1,\ldots,4$ of the SVAR model, we consider the following assumptions corresponding to Assumption (ref).

assumption\begin{itemize} • The difference of the policy rate $\{Y_{3, t}\}_{t}$ does not depend on $(Y_{1, t-4}, Y_{2, t-4})^\top$ directly. • $Y_{3, t}$ does not affect $Y_{2, t+1}$ directly. • $Y_{3, t}$ affects $Y_{3, t+1}$ directly. • $Y_{1, t}$ affects $Y_{2, t+4}$ directly. • $Y_{3, t}$ affects $Y_{1,t+4}$ directly. \end{itemize}
remEach assumption is considered based on the following background. \begin{itemize} • The policy rates are determined based on the behavior of the current economic indicators by the Federal Open Market Committee per about six weeks. • Prices of commodities are considered to be directly affected by production costs including wages of workers. In this study, we simply assume that the policy rates affect prices through unemployment level. • The policy rate tends to be gradually changing. Here, we attempt to approximate the causal structure by simplifying $Y_{3,t}$ as a direct cause of $Y_{3,t+1}$. • The condition is a standard assumption because the relationship between inflation and unemployment is widely observed. • Naturally, the policy rate is considered to affect unemployment through various economic activities. In this study, we simply assume that $Y_{3, t}$ is a direct cause of $Y_{1, t+4}$. \end{itemize}

Under Assumption (ref), we consider the following restrictions for $\bm{A}_s, 0 \leq s \leq 4$.

align*[align* omitted — 364 chars of source]

where $\dagger$ and $*$ mean nonzero parameters and arbitrary $\mathbb{R}$-valued parameters, respectively. Then, the estimated coefficient matrices of the SVAR model are as follows.

align*[align* omitted — 1,029 chars of source]

Moreover, we consider the following test \[ \mathcal{H}_0: \bm{A}_0 = \bm{O},\quad \mathcal{H}_1: \bm{A}_0 \neq \bm{O}. \] Based on the test statistics $z_{3T}$ defined in the previous section with $\bm{v} = \bm{1} = (1,\ldots,1)^\top \in \mathbb{R}^{k(k-1)/2}$, the results are summarized in Table (ref).

table[table omitted — 271 chars of source]

We cannot reject the null hypothesis at a significance level of $\alpha=0.05$. As described in Proposition (ref) in Appendix, when $\bm{A}_0 = \bm{O}$, all non-orthogonalized impulse responses can be regarded not only as partial effects, but also as “total effects". Therefore, if the null hypothesis is accepted, we can treat non-orthogonalized impulse responses from not only $e_{k,t}$ but also $e_{j,t}, j=1,...,k-1$ as indicators of the overall effects. However, given that we obtained a relatively small $p$-value for a small observations $T=104$, we continue our discussion on the estimated value of the SVAR model. Because one of the purposes of the policy rate is to stabilize employment and prices, the signs of the third row of $\hat{\bm{A}}_{0T}$ may be consistent with our intuition. The sign of the $(2,1)$ component of $\hat{\bm{A}}_{0T}$ also seems reasonable because unemployment may cause a decline in wages, and hence, have a negative impact on prices. However, the absolute value of the $(2, 1)$ component may be overestimated because prices are rigid. Although we simplified the causal structure, obtaining an $\hat{\bm{A}}_{0T}$ that aligns with intuition may serve as evidence that the model can approximate the structure to some extent.

In summary, the proposed model seems to describe the data structure. In particular, the non-orthogonalized impulse responses suggest that it may take about one year for the policy rate decrease to improve unemployment and policy rates increase may depress the prices in the short term. See the Appendix for the causal interpretation of the non-orthogonalized impulse responses.

Proofs

proof[Proof of Theorem (ref)] It follows from the basic properties of the $\,\mathrm{vec}$ operator and the matrix Kronecker product that \begin{eqnarray*} \bm{q}_* &=& \,\mathrm{vec}(\bm{Q}_*) = \,\mathrm{vec}(\bm{L}_g(\bm{B}_*)) = f_1(\bm{b}_*) \\ \bm{a}_{0*} &=& \,\mathrm{vec}(\bm{A}_{0*}) = \,\mathrm{vec}(\bm{I}_k - \bm{L}_g(\bm{B}_*)^{-1}) = f_2(\bm{b}_*) \\ \bm{a}_* &=& \,\mathrm{vec}(\bm{Q}_*^{-1}\bm{B}_*) = \left(\bm{I}_r \otimes \bm{L}_g(\bm{B}_*) \right)^{-1} \bm{b}_* = f_3(\bm{b}_*). \end{eqnarray*} We show that the map $\bm{L}$ is differentiable at $g(\bm{B}_*) = g(\,\mathrm{vec}^{-1}(\bm{b}_*))$. For every matrix $\bm{C}=(c_{ij})_{1 \leq i, j \leq k}$ that admits the LU decomposition, the corresponding lower-triangular matrix $\bm{L}(\bm{C})=(l_{ij})_{1 \leq i, j \leq k}$ and the upper-triangular matrix $\bm{U}(\bm{C}) = (u_{ij})_{1 \leq i, j \leq k}$ are solutions to the following equation: \[ \begin{pmatrix} c_{11} & c_{12} & c_{13} & \cdots & c_{1k} \\ c_{21} & c_{22} & c_{23} & \cdots & c_{2k} \\ \vdots & \vdots & \vdots & \ddots & \vdots \\ c_{k1} & c_{k2} & c_{k3} & \cdots & c_{kk} \end{pmatrix} = \begin{pmatrix} 1 & 0 & 0 & \cdots & 0 \\ l_{21} & 1 & 0 & \cdots & 0 \\ l_{31} & l_{32} & 1 & \cdots & 0 \\ \vdots & \vdots & \vdots & \ddots & \vdots \\ l_{k1} & l_{k2} & l_{k3} & \cdots & 1 \\ \end{pmatrix} \begin{pmatrix} u_{11} & u_{12} & u_{13} & \cdots & u_{1k} \\ 0 & u_{22} & u_{23} & \cdots & u_{2k} \\ 0 & 0 & u_{33} & \cdots & u_{3k} \\ \vdots & \vdots & \vdots & \ddots & \vdots \\ 0 & 0 & 0 & \cdots & u_{kk} \\ \end{pmatrix}. \] We can observe that for every $i, j =1,2,\ldots,k$, $l_{ij}$ and $u_{ij}$ can be written as a rational function of some components of $\bm{C}$. Under our assumptions, $\bm{Q}_*$ and $g(\bm{A}_*)$ are non-singular lower and upper-triangular matrices, respectively, which implies that all leading principal sub-matrices are non-singular, therefore, $g(\bm{B}_*) = \bm{Q}_* g(\bm{A}_*)$ admits the LU decomposition. Combining these facts, we obtain the differentiability of $\bm{L}$ at $g(\bm{B}_*)$. Thus, we can apply the delta method to derive the conclusion.

We omit the proof of Propositions (ref), and (ref) because they are direct consequences of the delta method.

proof[Proof of Theorem (ref)] We provide the proof of the assertion (ii), because (i) is a direct consequence of Lemma (ref). We consider the case where $l=1$. Note that \[ \{|z_{1T}| > c\} = \left\{ \left|\frac{z_{1T}}{\sqrt{T}}\right| > \frac{c}{\sqrt{T}} \right\} = \{ A_T^+ > 0 \} \cup \{ A_T^- <0 \}, \] where \[ A_T^+ = \frac{z_{1T}}{\sqrt{T}}-\frac{c}{\sqrt{T}}\quad \mbox{and}\quad A_T^- = \frac{z_{1T}}{\sqrt{T}}+\frac{c}{\sqrt{T}}. \] Noting that $\hat{\bm{\Sigma}}_1$ and $\hat{\bm{q}}_T$ are consistent estimators for $\bm{\Sigma}_1$ and $\bm{q}_*$, respectively, we have \begin{eqnarray*} \frac{z_{1T}}{\sqrt{T}} &=& (\bm{v}^\top \hat{\bm{\Sigma}}_{1\mathrm{sub}} \bm{v})^{-1/2} \bm{v}^\top (\hat{\bm{q}}_{\mathrm{sub}} -\bm{q}_{*\mathrm{sub}}) +(\bm{v}^\top \hat{\bm{\Sigma}}_{1\mathrm{sub}} \bm{v})^{-1/2} \bm{v}^\top \bm{q}_{*\mathrm{sub}} \\ &=& o_p(1) + (\bm{v}^\top {\bm{\Sigma}}_{1\mathrm{sub}} \bm{v})^{-1/2} \bm{v}^\top \bm{q}_{*\mathrm{sub}},\quad T \to \infty. \end{eqnarray*} Therefore, it holds that \[ A_T^+ = A_T^- = o_p(1) + A, \quad T \to \infty, \] with \[ A=(\bm{v}^\top {\bm{\Sigma}}_{1\mathrm{sub}} \bm{v})^{-1/2} \bm{v}^\top \bm{q}_{*\mathrm{sub}}. \] By the continuous mapping theorem, we have $1_{\{A_T^+ >0\}} \to^p 1_{\{A >0\}}$ and $1_{\{A_T^- <0\}} \to^p 1_{\{A <0\}}$ as $T \to \infty$. Therefore, it follows from the dominated convergence theorem that \begin{eqnarray*} \lim_{T \to \infty}\,\mathrm{P}(|z_{1T}| > c) &=& \lim_{T \to \infty} \left\{ \,\mathrm{P} (A_T^+>0) + \,\mathrm{P} (A_T^-<0) \right\} \\ &=& \lim_{T \to \infty} \left\{ \,\mathrm{E} \left[1_{\{A_T^+>0\}}\right] + \,\mathrm{E}\left[ 1_{\{A_T^-<0\}} \right] \right\} \\ &=& \,\mathrm{E}\left[1_{\{A^+ >0\}}\right] + \,\mathrm{E}\left[1_{\{A^- <0\}}\right]= 1, \end{eqnarray*} which ends the proof for $l=1$. The assertion for $l=2,3$ can be proved similarly.