Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,732 characters · 10 sections · 49 citation commands
Structural Vector Autoregressions and Higher Moments: Challenges and Solutions in Small Samples.
\thispagestyle{empty}
{\it JEL Codes:} C12, C32, C51 \\ {\it Keywords:} structural vector autoregression, non-Gaussian, independent, GMM
In non-Gaussian structural vector autoregressions (SVAR), higher-order moment conditions derived from mutually independent structural shocks can be used to identify the simultaneous relationship. These higher-order moment conditions can be used to estimate the SVAR with a generalized method of moments (GMM) estimator or similar moment-based estimators, see, e.g., lanne2021gmm, keweloh2020generalized, guay2021identification, mesters2022non, amengual2022moment, karamysheva2022we, lanne2022statistical, lanne2022identifying, keweloh2023monetary, or drautzburg2023refining.
This study focuses on the small sample behavior of asymptotically efficient SVAR-GMM estimators based on higher-order moment conditions derived from mutually independent structural shocks. The results show that in small samples, standard implementations of GMM estimators lead to volatile and biased estimates with distorted coverage and Wald test statistics. To improve the small sample performance of the estimators, two extensions are proposed. First, the assumption of mutually independent structural shocks is used not only to derive moment conditions but also to estimate the asymptotically efficient weighting matrix and the asymptotic variance of the estimator. Second, a continuously updated scaling term is added to the weighting matrix to prevent the small sample scaling bias.
The small sample behavior of GMM estimators in general has been studied extensively and GMM estimators are known to exhibit a small sample bias, see, e.g., han2006gmm and newey2009generalized. I prove that in the SVAR, the asymptotically efficient GMM estimator using higher-order moment conditions is biased towards solutions corresponding to innovations with a variance smaller than the normalizing unit variance assumption. This means that the normalizing unit variance moment conditions are systematically violated towards innovations with smaller variances. To address this scaling bias, I propose the continuous scale updating estimator (SVAR-CSUE), which eliminates the scaling bias by continuously updating the weighting matrix with a scaling term.
The estimation of the asymptotically efficient weighting matrix and the asymptotic variance of SVAR GMM estimators using higher-order moment conditions relies on estimates of the long-run covariance matrix of the sample average of the moment conditions, which is challenging to estimate due to the large number of higher-order moment conditions involved. For example, in an SVAR with two variables, independent shocks with unit variance imply eight moment conditions of order two to four and therefore, the covariance matrix of the moment conditions contains $35$ co-moments of order four to eight. However, in an SVAR with six variables, the number of moment conditions grows to $185$ and the covariance matrix explodes to $2919$ co-moments of order four to eight.
To solve the issue of the rapidly growing number of higher-order co-moments in the long-run covariance matrix, I propose leveraging the assumption of serially and mutually independent structural shocks, which is already used to derive the identifying moment conditions. This assumption enables the decomposition of higher-order co-moments of the covariance matrix into products of lower-order moments. For instance, in the SVAR with six variables and $185$ moment conditions of order two to four, the covariance matrix leveraging the assumption of serially and mutually independent shocks contains only $39$ moments of order one to six. Therefore, by leveraging this assumption to estimate the covariance matrix of the moment conditions, the researcher only needs to estimate moments up to order six instead of order eight, and the number of these moments increases linearly in the dimension of the SVAR, avoiding an explosion in the number of co-moments.
The study provides Monte Carlo simulations to demonstrate the effectiveness of the proposed measures. The simulations show that the proposed estimator of the efficient weighting matrix, leveraging the assumption of serially and mutually independent shocks, provides a substantially better approximation to the true asymptotically efficient weighting scheme than traditional approaches based on serially uncorrelated moment conditions. Moreover, the simulations show that the asymptotically efficient SVAR-GMM estimator exhibits a small sample scaling bias towards innovations with variances smaller than the normalizing unit variance assumption, which the proposed SVAR-CSUE, including a scale updating term, is shown to be effective in preventing. Additionally, the proposed SVAR-CSUE outperforms standard GMM implementations in terms of bias and volatility by correcting the scaling bias and additionally leveraging the assumption of serially and mutually independent shocks to estimate the efficient weighting matrix. Finally, leveraging the assumption of serially and mutually independent shocks to estimate the asymptotic variance of the estimator leads to substantially better coverage and rejection rates of Wald tests.
hoesch2022robust show that GMM-based Wald tests utilizing higher-order moment conditions tend to perform poorly and over-reject, particularly when dealing with weakly non-Gaussian shocks. To address this issue, hoesch2022robust and drautzburg2023refining propose Bonferroni-based approaches for hypothesis testing and constructing confidence sets, ensuring correct asymptotic coverage regardless of the level of non-Gaussianity. The simulations in this study show that GMM-based Wald tests relying on traditional approaches to estimate the asymptotic variance, i.e. based on serially uncorrelated moment conditions, may even severely over-reject in a setup with strongly non-Gaussian shocks featuring non-zero skewness and excess kurtosis. I demonstrate that a significant portion of the over-rejection issue stems from challenges in accurately estimating the asymptotic variance of the GMM estimator when employing higher-order moment conditions. Furthermore, the simulations reveal that by leveraging the assumption of serially and mutually independent shocks to estimate the asymptotic variance of the estimator, the issue can be alleviated, leading to considerably improved coverage and rejection rates for Wald tests. A notable advantage of the approach proposed in this study, compared to Bonferroni-based methods, is its computational efficiency, as it does not require inverting a test statistic. However, it is important to acknowledge that the proposed method may introduce potential distortions in setups with weak non-Gaussianity.
The remainder of this article is organized as follows. Section (ref) contains an overview of SVAR identification and estimation using higher-order moment conditions implied by the assumption of independent structural shocks. Section (ref) shows how the independence assumption can be leveraged to estimate the long-run covariance matrix of the moment conditions required for the asymptotically efficient weighting matrix and asymptotic variance of the SVAR-GMM estimator. Section (ref) shows that the asymptotically efficient SVAR-GMM estimator using higher-order moment conditions exhibits a small sample scaling bias and proposes the SVAR-CSUE which prevents the scaling bias. Section (ref) demonstrates the advantages of the two proposed measures in a Monte Carlo simulation. Section (ref) concludes.
This section briefly summarizes how higher-order moment conditions derived from mutually independent structural shocks can be used to identify and estimate the SVAR. A detailed description can be found in lanne2021gmm, keweloh2020generalized, guay2021identification, or lanne2022identifying.
Consider the SVAR $y_t = \sum_{p=1}^{P} A_p y_{t-p} + u_t$ with an $n$-dimensional vector of observable variables $y_t=[y_{1,t},...,y_{n,t}]'$, the reduced form shocks $u_t=[u_{1,t},...,u_{n,t}]'$, and
describing the impact of an $n$-dimensional vector of unknown structural shocks $\varepsilon_t=[\varepsilon_{1,t},...,\varepsilon_{n,t}]'$ with zero mean and unit variance. The matrix $ B_0 \in \mathbb{B} := \{B\in \mathbb{R}^{n \times n} | det(B) \neq 0 \}$ governs the simultaneous interaction and is assumed to be invertible. The reduced form shocks can be estimated consistently and this study focuses on the simultaneous interaction in Equation ((ref)). The reduced form shocks are equal to an unknown mixture of the unknown structural shocks, $u_t = B_0 \varepsilon_t$. Reversing this relationship yields the innovations $e(B)_t$, defined as the innovations obtained by unmixing the reduced form shocks with some invertible matrix $B$
If $B$ is equal to the true mixing matrix $B_0$, the innovations are equal to the structural shocks.
Imposing structure on the (in-)dependencies of the structural form shocks allows deriving moment conditions. For example, if the structural shocks are mutually uncorrelated with unit variance, the unmixing matrix $B$ should yield uncorrelated innovations with unit variance and therefore, satisfy the second-order moment conditions in Table (ref). However, there are not sufficiently many second-order moment conditions to identify the SVAR, which is the well-known identification problem. Imposing more structure on the (in-)dependencies of shocks allows for deriving additional higher-order moment conditions. In particular, if the structural form shocks are mutually independent, the innovations should additionally satisfy the third- and fourth-order moment conditions in Table (ref).
In general, the moment conditions implied by mutually independent shocks with unit variance can be written as
with variance and covariance conditions for $m \in \mathbf{2}$, coskewness conditions for $m \in \mathbf{3}$, and cokurtosis conditions for $m \in \mathbf{4}$ with
and $c(m) :=
$ .
In fact, all co-moment conditions except for the symmetric cokurtosis conditions $E[\varepsilon_{it}^2\varepsilon_{jt}^2]=1$ with $i \neq j$ can be derived under the weaker assumption of mutually mean independent shocks instead of the stronger assumption of mutually independent shocks. However, the weaker assumption may not be sufficient to ensure identification. For example, lanne2021gmm, keweloh2020generalized, and lanne2022statistical propose different sets of identifying fourth-order moment conditions, nevertheless, all three identification results require the assumption that all cokurtosis conditions implied by independent shocks hold. Therefore, all three approaches are equally restrictive in terms of the structure imposed on the (in)dependencies of the shocks, and none of the proposed identification results using fourth-order moment conditions necessarily hold under the weaker assumption of mutually mean independent shocks.\footnote{ Specifically, keweloh2020generalized uses all cokurtosis conditions implied by independent shocks, lanne2021gmm use only asymmetric cokurtosis conditions ($E[\varepsilon_{it}^3 \varepsilon_{jt}]=0$ for $i \neq j$), and lanne2022identifying use all symmetric cokurtosis conditions ($E[\varepsilon_{it}^2\varepsilon_{jt}^2]=1$ for $i \neq j$). Although the last two approaches use only subsets of the cokurtosis conditions implied by independent shocks, identification is only ensured if all cokurtosis conditions implied by independence hold, see Assumption 1 (ii) in lanne2022identifying and the proof of Proposition 1 in lanne2021gmm uses the same assumption. }
Recently, several studies have made advancements in the identification of non-Gaussian SVAR models under the weaker assumption of mean independent shocks. Notably, mesters2022non, keweloh2023uncertain, and anttonen2023bayesian have provided identification results based on moment-based approaches. keweloh2023uncertain and anttonen2023bayesian employ third-order moment conditions, which enable the identification of SVAR models when the shocks exhibit sufficient skewness and zero co-skewness, a property derived from the assumption of mean independent shocks. On the other hand, mesters2022non present more general identification results that are not restricted to third-order moment conditions. However, despite these advancements, the simulations in Section (ref) reveal that achieving efficient weighting of the higher-order moment conditions and accurate estimation of the asymptotic variances of the estimator can be challenging without relying on the assumption of independent structural shocks.
The remainder of the study uses the following assumptions:
The SVAR-GMM estimator is given by
where $g_T(B)= \frac{1}{T}\sum_{t=1}^{T} f(B,u_t)$, $W$ is a positive semi-definite weighting matrix, and $f(B,u_t)$ contains all or a subset of the moment conditions $ f_m(B,u_t)$ implied by independence to ensure identification for sufficiently non-Gaussian shocks. Consistency and asymptotic normality of the estimator follow from standard assumptions, such that
with $\hat{\beta}_T = vec(\hat{B}_T)$ and ${\beta}_0 = vec({B}_0)$, see hall2005generalized. In particular, asymptotic normality requires that the matrix $S$ exists and is finite. For the SVAR-GMM estimator based on second- to fourth-order moment conditions this requires that $\varepsilon_t$ has finite moments up to order eight. The weighting matrix $W^* := S^{-1}$ leads to the estimator $\hat{B}_T^*$ with the asymptotic variance $\sqrt{T}(vec(\hat{B}_T^*)-\beta_0) \overset{d}{\rightarrow} \mathcal{N}(0, (G' S^{-1} G)^{-1})$, which is the lowest possible asymptotic variance, see hall2005generalized.
In practice, the long-run covariance matrix of the moment conditions $S$ and the expected value of the derivative of the moment conditions $G$ are unknown and need to be estimated for inference and asymptotically efficient weighting. However, estimating these matrices for higher-order moment conditions is difficult in small samples. In a GMM setup not related to SVAR models burnside1996small propose to impose restrictions of the underlying economic model on the estimator of $S$ and $G$. In this section, I propose to exploit the assumption of serially and mutually independent shocks to improve the estimation of $S$ and $G$. This approach has also been used by amengual2022moment to derive a test for independent components.
In the SVAR, the long-run covariance of two arbitrary moment conditions $ f_m(B,u_t) = \prod_{i=1}^{n} e(B)_{i,t}^{m_i} - c $ and $ f_{\tilde{m}}(B,u_t) = \prod_{i=1}^{n} e(B)_{i,t}^{\tilde{m}_i} - \tilde{c} $ with $m,\tilde{m} \in \mathbf{2} \cup \mathbf{3} \cup \mathbf{4}$ , $c=c(m)$, and $ \tilde{c}=c(\tilde{m}) $ at $B=B_0$ is equal to
where the last equality follows from identically distributed shocks and $e(B_0)_t=\varepsilon_t$. Therefore, with fourth-order moments, i.e. $m,\tilde{m}\in \mathbf{4}$ such that $\sum_{i=1}^{n} m_i = \sum_{i=1}^{n} \tilde{m}_i = 4$, the long-run covariance matrix $S_{m,\tilde{m}}$ contains co-moments of the structural shocks up to order eight. In practice, $S_{m,\tilde{m}}$ in Equation ((ref)) can be estimated by replacing $\varepsilon_t$ with $e(B)_t$ and some initial estimate or guess $B$ of $B_0$ and a heteroscedasticity and autocorrelation consistent covariance (HAC) estimator, see newey1994automatic.
However, with serially independent structural shocks implied by Assumption (ref) the expression of $S_{m,\tilde{m}}$ simplifies to
where the superscript $ {SI}$ indicates that the equality $S_{m,\tilde{m}}= S_{m,\tilde{m}}^{SI} $ only holds for serially independent shocks. The long-run covariance matrix under the assumption of serially independent shocks denoted by $S ^{SI}$ can then be estimated based on Equation ((ref)). For the SVAR-GMM estimator, serially independent structural shocks imply serially uncorrelated moment conditions and therefore, the estimator of the long-run covariance based on Equation ((ref)) is equal to the frequently used estimator $\hat{S}(B) = \frac{1}{T}\sum_{t=1}^{T} f(B,u_t) f(B,u_t)'$ which estimates the long-run covariance matrix based on the covariance of the moment conditions.
With serially independent structural shocks the expression of the long-run covariance matrix $S$ simplifies to the covariance matrix $S^{SI}$. Nevertheless, the covariance matrix $S_{m,\tilde{m}}^{SI}$ of two fourth-order moments is still of order eight and remains difficult to estimate in small samples. For instance, consider the two moment conditions $E[\varepsilon_{1,t}^3 \varepsilon_{2,t}]=0$ and $E[\varepsilon_{3,t}^3 \varepsilon_{4,t}]=0$, such that the covariance of both moments conditions is equal to $ E[\varepsilon_{1,t}^3 \varepsilon_{2,t}\varepsilon_{3,t}^3 \varepsilon_{4,t}] $, a co-moment of order eight.
To further simplify the estimation of $S$, we can exploit the assumption of mutually independent shocks in addition to the assumption of serially independent shocks. This assumption is already used to derive the identifying moment conditions, and thus can also be used to simplify the estimation of $S$. Assuming serially and mutually independent shocks, as implied by Assumption (ref), allows simplifying the expression of the long-run covariance matrix to
where the superscript ${SMI}$ indicates that the equality $S_{m,\tilde{m}} = S_{m,\tilde{m}}^{SMI} $ only holds for serially and mutually independent shocks. Therefore, the long-run covariance matrix under serially and mutually independent shocks denoted by $S ^{SMI} $ can then be estimated based on Equation ((ref)).
By exploiting mutually independent shocks, higher-order co-moments can be transformed into a product of lower-order moments. When the shocks are mutually independent, the eighth-order covariance $E[\varepsilon_{1,t}^3 \varepsilon_{2,t} \varepsilon_{3,t}^3 \varepsilon_{4,t}]$ of two moment conditions, $E[\varepsilon_{1,t}^3 \varepsilon_{2,t}]=0$ and $E[\varepsilon_{3,t}^3 \varepsilon_{4,t}]=0$, simplifies to $ E[\varepsilon_{1,t}^3] E[\varepsilon_{2,t}]E[\varepsilon_{3,t}^3]E[\varepsilon_{4,t}] $, which is a product of moments of order one and three. In general, the covariance matrix under serially independent shocks $S^{SI}$ requires estimating co-moments of $\varepsilon_t$ of order four to eight, while the covariance matrix under serially and mutually independent shocks $S^{SMI}$ requires estimating moments of $\varepsilon_t$ of order one to six. Table (ref) shows the number of co-moments of $\varepsilon_t$ contained in $S^{SI}$ and $S^{SMI}$ for an SVAR-GMM estimator using all second- to fourth-order moment conditions. The number of higher-order co-moments increases quickly with the dimension of $\varepsilon_t$. For instance, in an SVAR with $n=2$ variables $S^{SI}$ requires estimating nine co-moments of order eight, but in an SVAR with $n=4$, this number grows to $156$, and with $n=6$ variables, $S^{SI}$ requires estimating $1287$ co-moments of order eight. In contrast to that, the number of higher-order moments in $S^{SMI}$ grows linearly in $n$. Thus, using mutually independent shocks to estimate $S$ appears particularly advantageous in larger SVARs.
An alternative approach to reduce the number of co-moments in the covariance matrix $S$ is to reduce the number of moments of the GMM estimator as proposed in lanne2021gmm and lanne2022statistical. However, omitting moment conditions may lead to a loss of efficiency. lanne2021gmm suggest a strategy involving the sequential application of an information criterion to select an appropriate set of moment conditions. However, this method becomes increasingly impractical in the context of large SVAR models, where the number of potential combinations of moment conditions escalates rapidly with the dimension of the SVAR. To circumvent the need to exhaustively explore all possible combinations of moment conditions, one could employ theoretical findings to identify redundant higher-order moment conditions. Unfortunately, such theoretical results are not available in the literature. An exception is Proposition 4.3 in keweloh2022structural, which establishes that in the special case of a recursive SVAR with independent shocks, all non-bivariate coskewness and cokurtosis conditions, meaning moment conditions involving more than two shocks like for instance $ E[\varepsilon_{1,t} \varepsilon_{2,t}\varepsilon_{3,t} ] =0$, are redundant.
The assumption of mutually independent structural shocks can also be used to estimate the $G$ matrix required to estimate the asymptotic variance of the GMM estimator. For an arbitrary moment condition $ f_m(B,u_t) = \prod_{i=1}^{n} e(B)_{i,t}^{m_i} - c(m) $ with $m \in \mathbf{2} \cup \mathbf{3} \cup \mathbf{4}$ the derivative with respect to $b_{pq}$ the element at row $p$ and column $q$ of $B$ evaluated at $B=B_0$ corresponds to an element of $G$ and is equal to
with $A=B_0^{-1}$ and $a_{jp}$ are the elements of $A$. The equality follows from $e(B_0)_t=\varepsilon_t$, the product rule, and $\frac{\partial e(B_0)_{i,t} }{\partial b_{pq}} = -a_{ip} \varepsilon_{q,t}$. Using the assumption of mutually independent shocks allows to decompose higher-order co-moments in Equation ((ref)) into a product of lower-order moments such that
where the superscript $SMI$ indicates that the equality $G_{m,b_{ql}} = G_{m,b_{ql}}^{SMI}$ holds for mutually independent shocks.
The assumption of serially and mutually independent shocks can also be used to derive a guess for the optimal weighting matrix $W^{*}$ without requiring an initial guess or estimate of the unknown simultaneous interaction $B_0$. Instead, the researcher can guess the distribution of each structural shock $\varepsilon_{i,t}$ for $i=1,...,n$ and if the guess is correct Equation ((ref)) directly yields the correct covariance matrix $S$, which can be used to calculate the optimal weighting matrix.
This section shows that the asymptotically efficient SVAR-GMM estimator has a small sample bias towards innovations with a variance smaller than the imposed unit variance assumption. To address this, I propose the continuous scale updating estimator (SVAR-CSUE), which overcomes the scaling bias by incorporating a continuously updated scaling term into the weighting matrix. Notably, the SVAR-CSUE remains asymptotically efficient, exhibits no scaling bias, and does not require to continuously re-estimate the long-run covariance matrix required for the asymptotically efficient weighting matrix.
han2006gmm and newey2009generalized show that for i.i.d. observations and a nonrandom weighting matrix $W$ the expected value of a GMM objective function is equal to
with $S(B):=E[ f(B,u_t)f(B,u_t)']$. Equation ((ref)) decomposes the expected value into a signal and a noise term, with the former being minimized at the true parameter value $B_0$ since $E\left[f(B_0,u_t)\right]=0$. The noise term, however, is not minimized at $B_0$ and in small samples, the noise term can dominate the signal term and lead to a bias, especially in models with many moment conditions.
The following proposition shows that deviating from the normalizing unit variance moment conditions towards innovations with a variance smaller than one reduces the noise term of the asymptotically efficient SVAR-GMM estimator at $B_0$, which leads to a small sample bias towards solutions that correspond to innovations with variance smaller than one.
A simple solution to reduce the scaling bias is to increase the weight of the variance conditions. Alternatively, one could impose the variance conditions as binding constraints, which corresponds to the situation where the weights of the variance conditions approach infinity and completely eliminate the scaling bias. However, these solutions put excessive weight on the variance conditions and no longer result in an asymptotically efficient estimator. An alternative solution to avoid the small sample bias without sacrificing asymptotic efficiency is the continuous updating estimator (CUE) proposed by han2006gmm
where $\hat{W}(B)=\hat{S}(B)^{-1}$ with an estimator $\hat{S}(B)$ for $S(B)$. With $W(B)=S(B)^{-1}$ the noise term in Equation ((ref)) collapses to $K/T$, where $K$ is equal to the number of moment conditions. Therefore, the noise term no longer depends on $B$ and hence leads to no bias. The CUE is a feasible version of this approach and replaces $W(B) = S(B)^{-1}$ with an estimator $\hat{W}(B) = \hat{S}(B)^{-1}$ and thus, the CUE depends on the ability to precisely estimate the long-run covariance matrix $S(B)$ which is difficult if the GMM estimator contains higher-order moment conditions.
The specific form of the noise term in the SVAR-GMM setup allows for eliminating the scaling bias without updating the long-run covariance matrix. To achieve this, define the scale updating weighting matrix
for a given weighting matrix $\tilde{W}$ and
for $i=1,...,n$. By construction, at $B=B_0 D$ for a scaling matrix $D=diag(d_1 ,...,d_n )$ it holds that $d(B_0 D)_i = d_i$. Therefore, a decrease of the variance of the $i$-th innovation, which is equivalent to an increase of the scaling coefficient $d_i$, leads to an increase of the weights of all moment conditions including the $i$-th innovation. The scale updating weighting matrix is chosen such that the noise term in Equation ((ref)) at $B=B_0 D$ for a scaling matrix $D=diag(d_1 ,...,d_n )$ and with $W=W(B, W^*)$ collapses to
which is minimized at $d_1 = \cdots = d_n = 1$. Thus, the scale updating weighting matrix eliminates the scaling bias without updating the long-run covariance matrix. In practice, the scaling matrix $D(B)$ can be estimated by $ \hat{D}(B) := diag \left( \prod_{i=1}^{n} \hat{d}(B)_i^{ m_{1,i}} ,..., \prod_{i=1}^{n} \hat{d}(B)_i^{ m_{K,i}} \right) $ and $ \hat{d}(B)_i := \frac{1}{ \sqrt{1/T \sum_{t=1}^{T}e(B)_{i,t}^2 }} $ which leads to the SVAR continuous scale updating estimator (SVAR-CSUE)
The two-step SVAR-CSUE uses the weighting matrix $ W(B, I)= \hat{D}(B) I \hat{D}(B)$ in the first step and $ \hat{D}(B) \hat{W}^* \hat{D}(B)$ in the second step where $\hat{W}^*$ is an estimator of the asymptotically efficient weighting matrix based on the first step estimator. Consistency of the first step estimator and the normalizing unit variance conditions ensure that the SVAR-CSUE using the weighting matrix $\hat{D}(B) W^* \hat{D}(B)$ remains asymptotically efficient. In particular, it holds that $\hat{D}(\hat{B}_T) \overset{p}{\rightarrow} I$ and $\hat{D}(\hat{B}_T) W^* \hat{D}(\hat{B}_T) \overset{p}{\rightarrow} W^* $ for a consistent first step estimator with $\hat{B}_T \overset{p}{\rightarrow} B_0$.
Alternatively, the SVAR-CSUE estimator can be seen as a transformation of the moment conditions into standardized moment conditions. Specifically, the SVAR-GMM estimator minimizes co-moments that measure the degree of dependency of the innovations. However, the covariance, coskewness, and asymmetric cokurtosis conditions can be reduced by decreasing the variance of the innovations. Therefore, to decrease the dependency measure, the SVAR-GMM may deviate from the normalizing unit variance assumption. The SVAR-CSUE transforms covariance conditions into correlation conditions, such that the dependency measure does not depend on the variance of the innovations. Specifically, for some weighting matrix $W$, the SVAR-CSUE can be written as
with standardized conditions $ \tilde{g}_T(B) := g_T(B)' \hat{D}(B)$. The standardization process transforms covariance conditions into correlation conditions. For example, let the $k$-th entry of $g_T(B)$ be the covariance condition for shock $i$ and $j$ such that
with $\hat{\sigma}(B)_i := \sqrt{1/T \sum_{t=1}^{T}e(B)_{i,t}^2 } $ and $\hat{\sigma}(B)_j := \sqrt{1/T \sum_{t=1}^{T}e(B)_{j,t}^2 } $ such that $\tilde{g}_{k,T}(B)$ measures the correlation of the $i$-th and $j$-th shock. The same transformation applies to coskewness and asymmetric cokurtosis conditions. For instance, let the $\tilde{k}$-th entry of $g_T(B)$ measure the coskewness of the $i$-th and squared $j$-th shock such that
with $\tilde{g}_{k,T}(B)$ being a standardized coskewness measure not affected by re-scaling the shocks.
This section studies the impact of leveraging the assumption of serially and mutually independent shocks to estimate the long-run covariance matrix as proposed in Section (ref) and adding the scale updating term to the weighting matrix as proposed in Section (ref) on the finite sample performance of SVAR estimators based on higher-order moment conditions.
The Monte Carlo simulation consists of two SVAR models $u_t=B_0\varepsilon_t$ with two and four variables and structural impact matrices
respectively. The structural shocks are drawn from a mixture of Gaussian distributions with mean zero, unit variance, skewness equal to $0.89$ and excess kurtosis of $2.35$. In particular, the shocks satisfy
where $\mathcal{B}(p)$ indicates a Bernoulli distribution and $\mathcal{N}(\mu, \sigma^2)$ indicates a normal distribution.
This section examines the influence of the estimated asymptotically efficient weighting matrix on the GMM loss. Figure (ref) compares the average and quantiles of the GMM objective function $g_T(B)' W g_T(B)$, where $g_T(B)$ contains all second-, to fourth-order moment conditions implied by mutually independent shocks with unit variance, and the objective function is evaluated at the true impact matrix $B_0$ for different weighting matrices $W$. The benchmark red loss employs the true but in practice unknown asymptotically efficient weighting matrix equal to the inverse of the long-run covariance matrix $S$. The blue loss uses the traditional estimator for the asymptotically efficient weighting matrix equal to the inverse of the sample covariance matrix of the moment conditions, which relies on the assumption of serially uncorrelated moment conditions or equivalently serially independent shocks. The green loss corresponds to the proposed estimator for the asymptotically efficient weighting matrix, which leverages the assumption of serially and mutually independent shocks, and estimates the long-run covariance matrix based on Equation ((ref)).
The simulation results demonstrate that the traditional estimator for the asymptotically efficient weighting matrix is inadequate for approximating the true asymptotically efficient weighting matrix, particularly for large SVAR models. Specifically, in the SVAR model with four variables and small sample sizes, the average GMM loss using the traditional estimator for the asymptotically efficient weighting matrix substantially differs from the average GMM loss based on the true asymptotically efficient weighting matrix. In contrast, the proposed weighting scheme, which leverages the assumption of serially and mutually independent shocks, provides a close approximation to the infeasible asymptotically efficient weighting scheme, thus demonstrating its effectiveness.
The following section compares the performance of three asymptotically efficient estimators in the SVAR with four variables. The three estimators use all second-, third-, and fourth-order moment conditions implied by mutually independent shocks, including six covariance, $16$ coskewness, and $31$ cokurtosis conditions, as well as four additional moment conditions normalizing the variance of the shocks to one. The estimators analyzed are:
The first estimator GMM$^*$ is the asymptotically efficient but infeasible estimator and serves as a benchmark. The second estimator GMM represents the standard implementation of a GMM estimator. The third estimator CSUE is the proposed GMM implementation tailored to address the SVAR specific GMM issues related to the difficulty of estimating the long-run covariance matrix and the small sample scaling bias.\footnote{ The simultaneous interaction is only identified up to sign and permutation. The sign and permutation of each estimator $\hat{B}$ is chosen based on a Wald test with $H_0: \hat{B}P=B_0$ for each for each suitable sign permutation matrix $P$ and the Wald test uses the true asymptotic variance of the estimator. }
Table (ref) reports the mean, $10$% quantile, and $90$% quantile of the variance of the first innovation, i.e. the mean and quantiles of $v_m := \frac{1}{T} \sum_{t=1}^{T} e(\hat{B})_{1t}^2$ for $m=1,...,2000$ simulations. The variance of the innovation is normalized to one by the unit variance moment conditions $\frac{1}{T} \sum_{t=1}^{T} e(B)_{1t}^2 - 1 = 0$ used by all three estimators. The infeasible, asymptotically efficient estimator, $\text{GMM}^*$, exhibits a downward variance scaling bias. Specifically, the mean of the variance of the first innovation is $0.88$, and even the upper $90$% quantile is $0.93$ and thus below the normalizing unit variance moment condition. This finding aligns with Proposition (ref), which shows that the noise term of the asymptotically efficient estimator leads to a small sample bias towards solutions with a variance below the normalizing unit variance condition. Moreover, the innovations corresponding to the traditional $\text{GMM} $ estimator exhibit a variance bias in the opposite direction, indicating that it poorly approximates the asymptotically efficient estimator. In contrast, the innovations corresponding to the proposed $\text{CSUE} $ estimator are centered around the unit variance condition, demonstrating the effectiveness of the variance scaling correction.
Table (ref) displays the mean, median, interquartile range (IQR), and standard deviation (SD) of the estimates for three representative elements of $B_0$ in Equation ((ref)): the lower left element $\hat{B}_{41}$, the first diagonal element $\hat{B}_{11}$, and the upper right element $\hat{B}_{14}$. The proposed CSUE exhibits the smallest mean bias, median bias, interquartile range, and standard deviation, and outperforms the alternative estimators. It is worth noting that the bias of the diagonal and lower triangular element is related to the variance scaling bias reported in Table (ref). This also explains why the upper triangular element is unbiased, as it is zero and hence not affected by the variance scaling bias.
The second part of the finite sample analysis focuses on the impact of the weighting scheme and the estimation of the asymptotic variance on the coverage of confidence bands and rejection frequencies for various Wald tests. The construction of these bands and tests requires estimates of $S$ and $G$, required for the asymptotic variance in Equation ((ref)). To this end, I present results in two sets of columns: SMI and SI. The SMI columns estimate $S$ and $G$ by assuming serially and mutually independent shocks, as proposed in Section (ref) and the SI columns follow the traditional approach and estimate $S$ and $G$ by only assuming serially independent shocks, which is equivalent to assuming serially uncorrelated moment conditions.
Table (ref) presents the coverage of $90 $% confidence bands for the three representative elements of the estimators. Notably, the traditional approach based on serially uncorrelated moment conditions exhibits a substantial shortfall in achieving the $90$% coverage level, with the coverage rate barely approaching $60$% even for the large sample case with $T=800$ observations. In contrast, the coverage rates of the bands of the CSUE, which are calculated based on the assumption of serially and mutually independent shocks, are substantially better and close to the $90$% level. The infeasible $\text{GMM}^* $ estimator does not rely on estimates of the asymptotically efficient weighting matrix and provides further insight into the effect of the estimated $S$ and $G$ matrices on the estimated asymptotic variance and its impact on the coverage of the confidence bands. The coverage of the $\text{GMM}^* $ bands using only serially independent shocks are worse than coverage rates using serially and mutually independent shocks. However, the $\text{GMM}^* $ bands using only serially independent shocks substantially improve compared to the GMM bands which are also constructed under the assumption of serially independent shocks, indicating that a large part of the poor coverage rate of the traditional GMM estimator can be attributed to unprecise estimates of the $B_0$ matrix.
Table (ref) reports the rejection rates at the $10$% level for three different Wald tests. The first test examines the joint null hypothesis $B=B_0$, the second tests the joint null hypothesis that the SVAR is recursive, meaning the upper triangular of $B_0$ contains only zeros, and the last test examines if the $B_{14}$ element is equal to zero. The table illustrates that the test performance decreases with the number of jointly tested coefficients. Moreover, the rejection rates of the GMM estimator are again far from the $10$% level, even in the large sample with $T=800$ observations. In contrast, the proposed CSUE estimator displays substantially improved performance, although it still rejects too often in small samples when multiple coefficients are jointly tested.
Lastly, Table (ref) analyzes the impact of the different approaches on the power of a single coefficient Wald test. The table shows the rejection rates at the $10$% level for the Wald test with $H_0: B_{41}=b$ and $b=2,...,8$ where the true value of the $B_{41}$ element is five. Again, the GMM estimator rejects the correct null hypothesis $B_{41}=5$ too often with rejection in more than $60$% of the simulations. Moreover, the rejection rates increase only by around $30$ percentage points for the tests with $H_0:B_{41}=2$ and $H_0:B_{41}=8$ in the small sample. In contrast, the proposed CSUE estimator exhibits much better performance, with the rejection rate for the correct null hypothesis $B_{41}=5$ being closer to the $10$% level. Additionally, the rejection rate increases more strongly, by around $65$ percentage points, for the tests with $H_0:B_{41}=2$ and $H_0:B_{41}=8$ in the small sample.
Appendix (ref) presents additional results for various asymptotically efficient moment-based estimators in the simulation discussed in the previous section. The simulations demonstrate that adding the scaling term proposed in Section (ref) to the infeasible $\text{GMM}^*$ estimator using the true but in practice unknown efficient weighting matrix yields the infeasible $\text{CSUE}^*$, which successfully eliminates the scaling bias and performs similar to the proposed CSUE analyzed in the previous section that leverages the assumption of mutually and serially independent shocks to estimate the weighting matrix. Moreover, the continuous updating estimator that uses the assumption of mutually and serially independent shocks to estimate the weighting matrix performs similarly to the proposed CSUE estimator, while the continuous updating estimator that only uses the assumption of serially independent shocks performs worse than the two-step GMM estimator analyzed in the main text. Furthermore, adding the scaling term to the traditional two-step GMM estimator that does not use the assumption of mutually independent shocks to estimate the weighting matrix is not enough to eliminate the variance scaling bias, but it leads to less volatile estimates.
Appendix (ref), (ref), and (ref) present results for the estimators examined in the previous sections in SVAR models with structural shocks generated from different distributions, including a skewed normal distribution, a t-distribution, and a truncated normal distribution. The findings are consistent with those in the previous section. Specifically, the infeasible asymptotically efficient GMM estimator exhibits a scaling bias towards innovations with variances smaller than one and the CSUE eliminates the bias by including a continuously updated scaling term. Moreover, the CSUE outperforms the GMM estimator in terms of bias and volatility, highlighting the importance of utilizing the assumption of serially and mutually independent shocks to estimate the efficient weighting matrix and of incorporating the scaling term to avoid the scaling bias. Lastly, using the assumption of serially and mutually independent shocks to estimate the asymptotic variance of the estimator leads to a significant improvement in the rejection rates of the considered Wald tests.
While the focus of this study is on estimating the simulations interaction $ u_t = B_0 \varepsilon_t$, applications of SVAR models typically also involve estimating a VAR with a lag structure, $y_t = \sum_{p=1}^{P} A_p y_{t-p} + u_t$. This can be done in a two-step approach: first, the VAR is estimated, and then, in the second step, the simultaneous interaction is estimated using the reduced form shocks obtained from the initial estimation. Employing such a two-step approach affects the asymptotic variance of the GMM estimator in Equation ((ref)), see Appendix Subsection 5.2 in gourieroux2020identification. In Appendix (ref), I present simulations for an SVAR with lags estimated using the two-step approach. The results reaffirm the core findings presented in this section. Specifically, even when employing the two-step approach, the $\text{GMM}^*$ estimator continues to display the same scaling bias and the CSUE eliminates the bias. Moreover, the CSUE outperforms the GMM estimator and using the assumption of serially and mutually independent shocks to estimate the asymptotic variance leads to a notable enhancement in the rejection rates of the Wald tests under consideration.
Finally, achieving more accurate estimates of the efficient weighting matrix and asymptotic variance hinges upon assuming mutually independent shocks. However, a commonly voiced concern regarding this assumption is that macroeconomic shocks may be influenced by a shared volatility process, as discussed in montiel2022svar and others. In this scenario, estimates of $S$ relying on the assumption of mutually independent shocks are no longer consistent and may lead to inefficient weighting and distorted inference. Appendix (ref) analyzes the impact of a common volatility process on the estimators considered in the previous section. First, the findings concerning the scaling bias and the superior performance of the CSUE to the GMM estimator in terms of bias and variance remain unaltered. Secondly, the rejection frequencies of the considered Wald test relying on estimates of asymptotic variance under the assumption of mutually independent shocks deteriorates, however, they still outperform the Wald tests not relying on estimates of asymptotic variance by a large margin.
This study investigates the small sample performance of asymptotically efficient SVAR-GMM estimators based on higher-order moment conditions derived from independent shocks. The simulations reveal that standard implementations of GMM estimators lead to a poor small sample performance. I propose two measures that significantly improve estimator performance. First, I propose the contentious scale updating estimator, which adds a continuously updated scaling term to the weighting matrix, thereby avoiding the small sample variance scaling bias of the asymptotically efficient GMM estimator. Second, I simplify the estimation of the long-run covariance matrix of the moment conditions by leveraging the assumption of serially and mutually independent shocks. This leads to more precise estimates of the asymptotically efficient weighting matrix and of the asymptotic variance of the estimator. The Monte Carlo simulation demonstrates that these measures lead to substantial improvements in the finite sample performance of the estimator.
A previous version of the paper is available under the title "A Feasible Approach to Incorporate Information in Higher Moments in Structural Vector Autoregressions" conducted with financial support from the German Science Foundation, DFG - SFB 823. Moreover, I gratefully acknowledge the computing time provided on the LiDO3 cluster at TU Dortmund, partially funded by the German Research Foundation (DFG) as project 271512359.