Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,621 characters · 13 sections · 86 citation commands
Uncertain Short-Run Restrictions and Statistically Identified Structural Vector Autoregressions
\thispagestyle{empty}
{\it Keywords:} non-Gaussianity, restrictions, penalty, oil market, stock market
Traditional approaches to identifying structural vector autoregressions (SVAR) typically involve imposing economically motivated restrictions, often restricting how structural shocks affect the variables in the SVAR simultaneously. More recently, alternative approaches imposing structure on the stochastic properties of the shocks, such as time-varying volatility or non-Gaussian and independent shocks, have been used for identification. Although statistical identification methods do not rely on economically motivated restrictions for identification, prior economic knowledge is still required to label the shocks. Put differently, some form of prior economic knowledge beyond the stochastic properties of the shocks remains necessary, even if it is not required for identification.
Consequently, the question is how do we utilize our prior economic knowledge in a statistically identified SVAR? The question relates to the critique of the "all-or-nothing approach" w.r.t. prior economic knowledge raised by baumeister2019structural. Traditional methods often treat prior knowledge as indisputable truth, enforcing restrictions without the ability to update them, while simultaneously ignoring other prior knowledge entirely. In this context, estimators relying on statistical identification approaches, which disregard any available restrictions, represent the extreme end of the "nothing approach."
This study proposes an approach to incorporate potentially invalid short-run restrictions on impulse responses into the estimation of a statistically identified SVAR. The estimator relies on non-Gaussian and (mean) independent shocks for identification and adds penalizes deviations from imposed zero restrictions on the simultaneous impulse responses. The study goes beyond merely proposing a test for overidentifying restrictions; instead, it advocates a non-dogmatic approach to incorporate restrictions. In contrast to traditional estimation approaches that treat restrictions as binding constraints, the proposed shrinkage estimator can stop shrinkage towards restrictions if the data present evidence against them, offering a non-dogmatic approach to incorporate restrictions. This approach seeks to enhance the efficiency of the statistically identified estimator through valid restrictions while mitigating the impact of invalid restrictions when the data provide evidence against them.
In this study, short-run zero restrictions are incorporated using a ridge penalty with adaptive weights (see, e.g., zou2006adaptive). The adaptive weights induce an important feature: It becomes cheap to deviate from invalid restrictions and costly to deviate from valid restrictions. As a result, the weights determine the importance of a given restriction and, consequently, the degree of shrinkage toward it in a data-driven manner. Therefore, in contrast to traditional estimators, which dogmatically rely on restrictions included as binding constraints, the ridge penalty offers a non-dogmatic alternative where the data determine the degree of shrinkage towards imposed restrictions. This approach is only possible when restrictions are not required for identification. Therefore, a separate identification approach is required to determine the weights by providing evidence in favor or against the imposed restrictions.
Identification in this study relies on non-Gaussian and (mean) independent shocks. The assumption of independent shocks often faces criticism, with the common objection that shocks driven by the same volatility process are not independent, as discussed in montiel2022svar. In response to this critique, recent developments in the non-Gaussian SVAR literature yield identification results under more relaxed assumptions regarding the (in)dependencies of shocks; see guay2021identification, mesters2022non, anttonen2023bayesian, or lewis2023identification for a comprehensive overview. This study adds to the literature by providing an identification result that allows for a two-stage identification approach based on the non-Gaussianity of the shocks. For skewed shocks, identification requires mean independent shocks and allows for a common volatility process. For shocks with zero-skewness but non-zero excess kurtosis, identification requires their independence.
The non-Gaussian estimator considered in this study aims at achieving robustness by relying as little as possible on structure imposed on the stochastic properties of the shocks. Specifically, the estimator only minimizes second- to fourth-order moment conditions implied by mean independent shocks, thus circumventing the need to impose a specific distribution on the shocks and enabling the identification of shocks driven by a common volatility process. Although this approach of imposing minimal structure on the stochastic properties of the shocks enhances robustness, it comes at the cost of efficiency.
The motivation of the proposed ridge estimator is to combine the statistical identification approach with short-run restrictions, leveraging prior economic knowledge on the simultaneous impulse responses, to enhance the efficiency of the estimator. Monte Carlo simulations show how economically motivated restrictions and a statistical identification approach complement each other; Valid restrictions improve the accuracy of the statistically identified estimator, and the impact of invalid restrictions decreases with evidence of the statistically identified estimator against them.
Bayesian approaches offer a natural way to incorporate economic knowledge using the prior distribution of the parameters. Moreover, Bayesian SVARs identified by independence and non-Gaussianity allow for updating economically motivated priors, see lanne2020identification, anttonen2021statistically, braun2021importance, and keweloh2023estimating. Nevertheless, the incorporation of prior economic knowledge, specifically the imposition of economically motivated zero restrictions, has deep roots in the frequentist SVAR literature as well. However, the frequentist approach to including restrictions often adopts a dogmatic stance, lacking the ability to gather and utilize evidence against a given restriction. This study introduces a non-dogmatic approach to include restrictions in the frequentist SVAR estimation framework, allowing for a more flexible and data-driven treatment of restrictions.
The application analyzes the interaction of the oil and stock market. kilian2009impact propose recursive restrictions to identify and estimate the effects of different oil and stock market shocks. The proposed restrictions are widely used to analyze the impact of oil market shocks on the stock market, see, e.g., apergis2009structural, abhyankar2013oil, kang2013oil, sim2015oil, ahmadi2016global, lambertides2017effects, mokni2020time, arampatzidis2021oil, kwon2022impacts, or arampatzidis2023identification. However, the impact and importance of stock market information shocks on the oil price is usually not analyzed. The application in this study fills this gap. I present evidence that oil and stock prices cannot be ordered recursively. By allowing both variables to interact simultaneously, the study reveals that information shocks originating from the stock market contain crucial information on oil prices, which explains approximately $25$ % of the fluctuations in oil prices.
The remainder of the paper is organized as follows: Section (ref) contains a brief overview on SVAR models. Section (ref) derives the non-Gaussian identification and estimation approach. Section (ref) introduces the ridge estimator to incorporate potentially invalid restrictions. Section (ref) uses simulations to illustrate the ability of the estimator to exploit correctly and discard falsely imposed restrictions. Section (ref) applies the estimator to an oil and stock market SVAR. Section (ref) concludes.
Consider an SVAR with $n$ variables
with $B_0 \in \mathbb{B} := \{B \in \mathbb{R}^{n \times n} | det(B)\neq 0 \}$ and $A_0:= B_0^{-1}$ and $n$-dimensional vectors of time series $y_t=[y_{1t} ,...,y_{nt} ]'$, reduced form shocks $u_t=[u_{1t} ,...,u_{nt} ]'$, and structural shocks $\varepsilon_t=[\varepsilon_{1t},...,\varepsilon_{nt}]'$ with mean zero and unit variance. The parameter matrices $A_1,...,A_p$ and the intercept term can be consistently estimated to obtain the reduced form shocks. To simplify, I treat the reduced form shocks as observable random variables and focus on identifying and estimating the simultaneous interaction $u_t = B_0 \varepsilon_{t}$.\footnote{ In practice, an SVAR can be estimated using a two-step approach where the VAR is estimated in the first step and the simultaneous interaction is estimated in the second step. Simulations analyzing the performance of the two-step approach can be found in Appendix (ref) and show little differences compared to the simulations in the main text. }
Define the innovations
equal to the innovations obtained by unmixing the reduced form shocks with a matrix $B\in \mathbb{B}$. For $B=B_0$, the innovations are equal to the structural shocks. Identification of the SVAR comes down to formulating a set of equations that guarantee the equivalence between innovations and structural shocks.
Typically, SVAR models are identified based on the assumption of uncorrelated structural shocks. Therefore, the matrix $B$ should generate uncorrelated innovations with unit variance, which yields $(n+1)n/2$ moment conditions. However, the matrix $B$ has $n^2$ coefficients. Consequently, infinitely many matrices $B \in \mathbb{B}$ generate uncorrelated innovations with unit variance, meaning that the assumption of uncorrelated structural shocks is not sufficient to identify the SVAR.
Traditional identification methods solve the identification problem by imposing structure on the interaction of the variables or impact of the shocks (e.g. short-run restrictions in sims1980macroeconomics, long-run restrictions in blanchard1989dynamic, sign restrictions in uhlig2005effects, or proxy variables in mertens2013dynamic). The structure probably most frequently imposed are short-run restrictions, meaning restrictions on coefficients of the $B$ matrix to reduce the number of free coefficients to $(n+1)n/2$ such that the remaining unrestricted coefficients are identified by the $(n+1)n/2$ moment conditions implied by uncorrelated shocks with unit variance. Note that identification requires at least $(n-1)n/2$ restrictions, and incorrect restrictions lead to inconsistent estimates. Additionally, with $(n-1)n/2$ restrictions, the SVAR is just identified. Therefore, even when the sample size goes to infinity, we are unable to detect incorrect restrictions.
More recently, identification approaches based on additional structure imposed on the stochastic properties of the structural shocks have been put forward in the literature. These approaches use properties such as time-varying volatility (see, e.g., rigobon2003identification, lanne2010structural, lutkepohl2017structural, lewis2021identifying, or bertsche2022identification) or the non-Gaussianity and independence of the shocks (see, e.g., matteson2017independent, herwartz2016macroeconomic, gourieroux2017statistical, lanne2017identification, maxand2020identification, lanne2021gmm, keweloh2020generalized, guay2021identification, mesters2022non, lanne2022identifying, herwartz2023point, drautzburg2023refining, or fiorentini2023discrete) to ensure identification.
This section first provides an intuition of how non-Gaussian and independent shocks allow to solve the identification problem and discusses different degrees of (in)dependence assumptions, emphasizing their economic significance. Subsequently, the section derives explicit conditions under which third- and fourth-order moment conditions derived from the assumption of mutually mean independent shocks identify the SVAR up to labeling of the shocks. Moreover, I propose an approach to label the shocks based on a first-step estimator. Finally, the last subsection introduces the non-Gaussian moment based estimator used in the remainder of the study.
Assumptions on the mutual (in)dependence of the structural shocks can be used to derive higher-order moment conditions and identify the SVAR. For example, the coskewness $E[\epsilon_{1t}^2 \epsilon_{2t}]$ of two independent shocks is zero. Figure (ref) illustrates how the coskewness can be used to identify skewed shocks. The left side shows plots of independent structural shocks $\varepsilon_{1t}$ and $\varepsilon_{2t}$, while the right side shows a rotation $e_{1t}$ and $e_{2t}$ of the shocks. In the upper row, the shocks $\varepsilon_{1t}$ and $\varepsilon_{2t}$ are independently drawn from a standard normal distribution. Any rotation of the shocks again leads to uncorrelated and independent innovations. Specifically, the covariance and coskewness are equal to zero for any rotation of the shocks. In the lower row, the first shock $\varepsilon_{1t}$ is drawn from a mixture of normal distributions, that is, the shock is drawn from a standard normal distribution with a probability of $99$% and with a probability of $1$% the shock is drawn from a normal distribution with mean four and variance one, which leads to a skewed distribution of the shock $\varepsilon_{1t}$. Rotating the skewed shocks leads to uncorrelated but dependent shocks. In particular, for the rotation depicted in the bottom right, the coskewness is positive, indicating that high absolute values of $e_{1t}$ are correlated with positive values of $e_{2t}$. Consequently, knowing the value of the first shock, $e_{1t}$, conveys information about the other shock, $e_{2t}$, although both shocks are uncorrelated. By utilizing the fact that the structural shocks are independent, the coskewness allows to immediately detect that the bottom right panel only shows a rotation of the structural shocks.
The key difference between non-Gaussian identification approaches and traditional restriction based approaches is that non-Gaussian approaches impose and utilize more structure on the dependency of the structural shocks. Traditional approaches typically only utilize the assumption of uncorrelated structural shocks, whereas non-Gaussian approaches rely on stronger assumptions, i.e., on the assumption of independent shocks. However, non-Gaussian identification does not necessarily require the assumption of fully independent shocks, but instead can work with weaker assumptions on the dependencies of the shocks, see lanne2021gmm, guay2021identification, mesters2022non, or anttonen2023bayesian. For instance, the illustration in Figure (ref) uses only a coskewness condition.
This raises the question of what constitutes an appropriate assumption regarding the dependencies among structural shocks. The assumption of independent shocks is often criticized for being overly restrictive, as it does not account for the possibility that shocks are influenced by a common volatility process, see montiel2022svar. In contrast, one could also argue that the assumption of uncorrelated shocks may be too weak. Shocks can be uncorrelated, but still be highly dependent in a manner that may not be suitable for structural shocks. For instance, consider the lower right panel in Figure (ref), which shows uncorrelated but evidently dependent shocks. In this scenario, if the first shock represents a demand shock and the second a supply shock, knowing the value of the demand shock would immediately provide information about the mean of the supply shock, despite the fact that both shocks are not correlated. Alternatively, consider an even more extreme example with $\varepsilon_1 \sim \mathcal{N}(0,1)$ and $\varepsilon_2 = \varepsilon_1^2-1$. Both random variables are uncorrelated but clearly dependent in a way that appears implausible for structural shocks.
In this study, I advocate for the assumption of mean independent shocks, i.e. $E[\varepsilon_{it} | \varepsilon_{-it} ]=0$ for $i=1,...,n$ and $ \varepsilon_{-it}:=[\varepsilon_{1t},...,\varepsilon_{(i-1)t},\varepsilon_{(i+1)t},...,\varepsilon_{nt}]$. This assumption is stronger than mere uncorrelated shocks, yet more lenient than assuming full independence. Imposing the condition of mean independent shocks excludes dependency structures where one shock provides information about the mean of another shock, while still allowing for dependency structures like a common volatility process.
Indeed, it is crucial to recognize that while traditional identification approaches may primarily utilize the assumption of uncorrelated shocks, their applications implicitly rest on the assumption of mean independent shocks for interpretation. This implicit assumption is vital when making causal statements about the expected responses of variables to shocks. To illustrate this, consider a bivariate SVAR without lags such that the first variable is equal to $y_{1t} = b_{11} \varepsilon_{1t} + b_{12} \varepsilon_{2t}$. The simultaneous impulse response of $y_{1t}$ to shocks $\varepsilon_{1t}$ is $b_{11}$ and is typically interpreted as the expected response of $y_{1t}$ to shocks $\varepsilon_{1t}$, that is, $b_{11} = E\left[ y_{1t } | \varepsilon_{1t} =1\right]$. Crucially, the assumption of uncorrelated shocks is not sufficient to guarantee that this equality holds. This is because $ E\left[ y_{1,t } | \varepsilon_{1t} =1\right] = b_{11} E\left[ \varepsilon_{1t} | \varepsilon_{1t}=1 \right] + b_{12} E\left[ \varepsilon_{2t} | \varepsilon_{1t}=1 \right] $ and $ E\left[ \varepsilon_{2t} | \varepsilon_{1t}=1 \right] =0$ does not follow solely from the assumption of uncorrelated shocks.\footnote{ One can also argue that impulse responses represents the thought experiment $b_{11} = E\left[ y_{1t } | \varepsilon_{1t}=1,\varepsilon_{2t}=0 \right]$. Although this is mathematically correct, it is not clear whether the thought experiment makes economically any sense for uncorrolated but dependent shocks. For example, the two shocks $\varepsilon_{1t} \sim N(0,1)$ and $\varepsilon_{2t}=\varepsilon_{1t}^3-3\varepsilon_{1t}$ are uncorrelated; however, the combination $\varepsilon_{1t}=1$ and $\varepsilon_{2t}=0 $ cannot even occur. } Instead, it requires mean independent shocks. Therefore, the assumption of mean independent shocks is used implicitly in any SVAR application that makes causal statements about the expected response of the variables to the shocks and thus, the assumption of mean independent shocks can also be used to identify the SVAR.
For $i,j,k,l \in \{1,...,n\}$, mutually mean independent shocks with mean zero and unit variance imply (co-)variance conditions
coskewness conditions
and cokurtosis conditions
In general, the moment conditions implied by mean independent shocks can be written as
with variance and covariance conditions for $M \in \mathbf{2}:= \{M=[ m_1,....m_n] \in \{0,1,2\}^n | \sum_{i=1}^{n} m_i = 2 \}$, coskewness conditions for $M \in \mathbf{3}:= \{M=[ m_1,....m_n] \in \{0,1,2\}^n | \sum_{i=1}^{n} m_i = 3 \}$, and cokurtosis conditions for $M \in \mathbf{4}:= \{ M= [ m_1,....m_n] \in \{0,1,2,3\}^n | \sum_{i=1}^{n} m_i = 4 , \exists 1 \in [ m_1,....m_n] \}$. For a given moment condition $E[f_{M}(B,u_t)]=0$, the indices $m_i$ simply denote the power of each innovation in the moment condition, i.e. for the moment condition $E[f_{M}(B,u_t)]= E [e(B)_{1t}^2 e(B)_{2t} ] = 0$ the indices are $m_1=2$, $m_2=1$, and $m_{3}=...=m_{n}=0$.
The following proposition establishes conditions under which the second- to fourth-order moment conditions implied by mutually mean independent shocks identify the SVAR.
The proposition shows that mutually mean independent shocks with non-zero skewness are identified by the moment conditions. This allows to identify skewed shocks even if they are driven by the same volatility process. Moreover, if shocks exhibit zero skewness, the moment conditions still identify the shocks with non-zero excess kurtosis if these shocks are mutually independent.\footnote{ Technically, the first statement only requires that the shocks satisfy all coskewness conditions implied by mean independent shocks and the second statement only requires that the shocks satisfy all cokurtosis conditions implied by independent shocks. Therefore, the second statement does not necessarily require independent shocks, however, it requires that all cokurtosis conditions resulting from independent shocks hold. However, the conditions do not follow from mean independent shocks, and finding an economically plausible process other than independent shocks that yields shocks satisfying all such conditions is not straightforward. } The first statement is a generalization of moment-based identification results in the literature (specifically for third moments in bonhomme2009consistent and for arbitrary moments in mesters2022non) to the case with multiple Gaussian shocks, similar to the partial identification results in maxand2020identification and guay2021identification. The second statement is related to the identification result in lanne2021gmm based on asymmetric fourth-order moment conditions. However, in contrast to the identification result in lanne2021gmm which only provides a local identification result, Proposition (ref) provides a global identification result up to sign and permutation. The second statement assumes that the shocks are independent and, therefore, satisfy the symmetric cokurtosis conditions $E[\varepsilon_{it}^2 \varepsilon_{jt}^2-1]=0$ for $i\neq j$. However, these symmetric conditions are not contained in the moment conditions $E[f(B,u_t)]$. The contribution of the second statement is to show that for independent shocks with sufficient excess kurtosis, the moment conditions $E[f(B,u_t)]$ guarantee global identification.\footnote{ Note that while the identification result in lanne2021gmm only uses asymmetric fourth-order moment conditions, it still requires the assumption that all cokurtosis conditions (not just the asymmetric ones) implied by mutually independent shocks hold. This assumption is used in the proof of the proposition in lanne2021gmm. Therefore, the identification result in lanne2021gmm cannot be used to identify an SVAR with shocks affected by the same volatility process. } Exuding the symmetric cokurtosis moment conditions is important to guarantee identification based on mean independent and skewed shock in the first statement. Specifically, if the moment conditions $E[f(B,u_t)]$ would contain the symmetric cokurtosis moment conditions, the first statement would not hold.
Proposition (ref) establishes identification up to sign and permutation, e.g. for any sign-permutation matrix $P$ the models $u_t = B_0 \epsilon_t$ and $u_t = \tilde{B} \tilde{\epsilon}_t$ with $\tilde{B}= B P^{-1}$ and $\tilde{\epsilon}_t= P \epsilon_t$ have the same dependency structure. Without additional guidance from the researcher, the shocks do not possess explicit structural labels. However, imposing restrictions on the impact of a structural shock of interest, as discussed in the next section, requires to label the shocks a priori.
One approach to address the indeterminacy of sign-permutations and label the shocks is to restrict the set of admissible $B$ matrices to a set containing a single representative of each sign-permutation class. This can be achieved, for instance, by constraining the set to:
where for almost all $B \in \mathbb{B}$ there exists a unique sign-permutation matrix $P$ such that $B P^{-1} \in \bar{\mathbb{B}}$ and $ B \tilde{P}^{-1} \notin \bar{\mathbb{B}}$ for all sign-permutation matrices $\tilde{P} \neq P$, compare lanne2017identification. Intuitively, the constrained set $\bar{\mathbb{B}}$ imposes the restriction that shocks $\varepsilon_{it}$ have a positive impact on $u_{it}$ and the simultaneous impact of shock $\varepsilon_{it}$ on $u_{it}$ is greater in absolute terms than the impact of all following shocks on $u_{it}$.
Constraining the set of admissible matrices to $\bar{\mathbb{B}}$ allows to a priori label the shocks. For example, suppose that the first variable measures government spending, and thus in $\bar{\mathbb{B}}$ the first shock always has the largest simultaneous impact on government spending and could be labeled as the government spending shock.\footnote{ Sign restrictions are an alternative approach to restrict and a priori label the shocks. Sign restrictions restrict the set of admissible $B$ matrices to a smaller set containing unique sign-permutation representatives, and each matrix in the constrained set corresponds to a shock labeled a priori based on the sign restrictions. However, labeling based on sign restrictions is not ideal for zero restrictions which are located at the boundary of the constrained set. } However, relying on $\bar{\mathbb{B}}$ to label the shocks can be problematic if $B_0$ is located at the boundary of the set, i.e., if in the example above government spending is equally driven by government spending and output shocks.
Figure (ref) illustrates potential labeling problems using $\bar{\mathbb{B}}$. To simplify, I consider bivariate models where labeling based on $\bar{\mathbb{B}}$ only relies on the two elements in the first row of a given $B_0$ matrix and thus can be easily visualized. Each panel shows a different $B_0$ matrix and plots the elements in the first row that represent the matrix as a green dot. In the bivariate SVAR, there is only one permutation of $B_0$ with positive diagonal elements, represented by the red dot that represents the elements in the first row of the permutation. The dotted epsilon balls around points illustrate a set of similar matrices. The shaded area displays the set $\bar{\mathbb{B}}$ which always contains one of both permutations.
The first and second panels display examples where $B_0$ is located in the inner area of $\bar{\mathbb{B}}$ and, therefore, the set $\bar{\mathbb{B}}$ also contains similar matrices in the epsilon balls. In contrast, the third and fourth panels display examples where $B_0$ is located at the boundary, resulting in some similar matrices in the epsilon ball not being contained within the set. For example, consider the third panel where
Now, consider an estimator $\hat{B}$ and its sign permutation $\tilde{B}$ with
Clearly, $\hat{B}$ corresponds to the same order of shocks as $B_0$ and is contained in the epsilon ball around $B_0$, while $\tilde{B}$ corresponds to the reverse order with a sign flip and is contained in the epsilon ball around the permutation of $B_0$. However, even though $B_0 \in \bar{\mathbb{B}}$ the estimator $\hat{B}$ that is close to $B_0$ is not contained in $\bar{\mathbb{B}}$, while the estimator $\tilde{B}$ corresponding to the reverse order is included. This illustrates how using $\bar{\mathbb{B}}$ to label shocks can yield misleading results when $B_0$ is located at the boundary of the set.
To avoid this, I propose to use an initial labeled estimator $\bar{B} $ of $B_0$ as a transformation to ensure that $B_0$ is located in the inner area of the set. Define the generalized set
For $\bar{B}=I$ the set $\bar{\mathbb{B}}_{\bar{B}=I} $ is equal to $\bar{\mathbb{B} }$. The key idea is to use an initial labeled estimator $\bar{B}$ as a transformation that re-centers each $B$ matrices such that $B_0$ is located at the inner of the set $\bar{\mathbb{B}}_{\bar{B}}$. Figure (ref) visualizes how the transformation moves the green points and circles representing the correct permutations and their neighborhood to the center of the set, while the red points and circles, representing the incorrect permutation and their neighborhood, are pushed away from the set. This ensures that the neighborhood of the correct permutation is contained in the set, and thus all solutions in the neighborhood receive the same labels. Importantly, the initial estimator $\bar{B} $ used for the transformation does not need to be equal to $B_0$. Figure (ref) in the appendix shows that using a $\bar{B} $ in proximity of $B_0$ is sufficient to move $ B_0 $ to the inner area of the set such that all matrices in the neighborhood still receive the same labels.
Using the second- to fourth-order moment conditions, $ E[f(B,u_t)]=0$, implied by mutually mean independent shocks, the SVAR-GMM estimator can be written as
with $g_T(B)= \frac{1}{T}\sum_{t=1}^{T} f(B,u_t)$ and a suitable weighting matrix $W$. Consistency and asymptotic normality of the estimator follow from standard assumptions and the identification result in Proposition (ref), see hall2005generalized.
The asymptotically efficient weighting matrix $W =S^{-1}$ with the long-run covariance matrix $S := \underset{T \rightarrow \infty }{lim} E \left[T g_T(B_0) g_T(B_0)' \right]$ leads to the lowest possible asymptotic variance of the estimator, see hall2005generalized. However, in small samples and with higher-order moment conditions, the efficient weighting matrix is difficult to estimate and the asymptotically efficient SVAR-GMM estimator exhibits a scaling bias towards innovations with a variance smaller than the normalizing unit variance assumption, keweloh2023structural.
To address both problems, keweloh2023structural proposes the two-step SVAR continuous scale updating estimator (SVAR-CSUE). The SVAR-CSUE assumes serially and mutually independent shocks to estimate the efficient weighting matrix and incorporates a continuously updated scaling term into the weighting matrix to eliminate the scaling bias. The two-step SVAR-CSUE is defined as follows:
with $W(B)= \left( \hat{D}(B) \hat{W} \hat{D}(B) \right)$ and the continuously updated scaling term $ \hat{D}(B) := diag \left( \prod_{i=1}^{n} \hat{d}(B)_i^{ m_{1,i}} ,..., \prod_{i=1}^{n} \hat{d}(B)_i^{ m_{K,i}} \right) $ where $ \hat{d}(B)_i := \frac{1}{ \sqrt{1/T \sum_{t=1}^{T}e(B)_{it}^2 }} $. The parameters $ m_{j,i}$ for $j=1,...,K$ and $i=1,...,n$ in the scaling term $\hat{D}(B) $ are equal to the power of the $i$-th innovation in the $j$-th moment condition and $\hat{d}(B)_i$ is equal to the inverse of the standard deviation of the $i$-th innovation. Consequently, the scaling term increases the weight of a given moment condition if $B$ leads to innovations with a variance smaller than one, which eliminates the scaling bias towards innovations with a variance smaller than the normalizing unit variance assumption, see keweloh2023structural.
In the first step, the weighting matrix $\hat{W}$ of the two-step SVAR-CSUE is equal to the identity matrix, and in the second step, it is equal to the inverse of the estimated long-run covariance matrix leveraging the assumption of serially and mutually independent shocks as proposed in keweloh2023structural.
Importantly, the assumption of serially and mutually independent shocks is only used to estimate the weighting matrix, which affects efficiency of the estimation. However, identification and consistency are ensured by the weaker assumption of mutually mean independent and sufficiently skewed shocks with Proposition (ref). Therefore, consistency of the SVAR-GMM and SVAR-CSUE require to impose only little structure on the stochastic properties of the shocks. Intuitively, the precision of the estimation tends to increase with the structure imposed on the SVAR. For example, a correctly specified maximum likelihood estimator would likely yield reduced bias, lower mean squared error, and narrower confidence intervals. However, the gain in precision of such an estimator must be weighed against the potential pitfall of misspecifying the distribution and dependence patterns of the shocks.\footnote{The simulation in summarized in Table (ref) in the appendix illustrates how a non-Gaussian maximum likelihood estimator which misspecifies the common volatility process of the shocks is biased whereas the proposed moment-based estimator remains unbiased.} Rather than imposing more structure on the stochastic properties of the shocks, the subsequent section shows how the researcher can complement the SVAR-CSUE by incorporating economically grounded short-run zero restrictions using a shrinkage approach to increase the estimator's precision.
This section proposes an approach to incorporate potentially invalid short-run zero restrictions using a ridge penalty with adaptive weights into the non-Gaussian SVAR estimation. Unlike conventional SVAR methods that rely on restrictions for identification, the proposed methodology leverages restrictions as a means of incorporating the researchers' a priori economic knowledge to increase the precision and efficiency of the estimation in comparison to an estimator relying only on the non-Gaussianity of the shocks. Moreover, combining a statistical identification approach with restrictions using a shrinkage mechanism enables the data to provide evidence against a given restriction and to stop shrinkage towards invalid restrictions.
The following notation is used to denote short-run zero restrictions on $B_0$. Let $\mathcal{R} $ be the set of all pairs $(i,j) \in \{1,...,n\}^2 $ corresponding to elements $B_{ij}$ of $B_0$ restricted to zero. For example, imposing a recursive order implies that all elements in the upper-triangular of $B_0$ are equal to zero, which corresponds to $ \mathcal{R}=\{(i,j) \in \{1,...,n\}^2 | j>i \}$.\footnote{ Proxy variables can also be implemented using short-run restrictions and the ridge estimator proposed in this section can be applied to proxy variables in augmented proxy SVAR models. Moreover, the ridge approach can be applied to restrictions in an $A$-type SVAR model with $A_0 u_t = \varepsilon_t$. Lastly, the approach can easily be extended to non-zero restrictions. }
Define the Ridge SVAR-CSUE as
with a tuning parameter $\lambda \geq 0$ and adaptive weights $v_{ij}$, compare zou2006adaptive for adaptive Lasso and dai2018broken for adaptive Ridge estimators.
Before discussing the construction of the weights and tuning parameter, it is worth highlighting the relation of the proposed estimator to existing approaches. Traditional restriction based estimators use restrictions as binding constraints, while estimators based on stochastic properties, like non-Gaussianity, typically ignore any potentially available restrictions. The proposed shrinkage estimator nests both approaches as special cases. Specifically, if the tuning parameter and the weights are manually set to converge to infinity, all restrictions become binding constraints, resembling a traditional restriction based estimator. Conversely, if the tuning parameter and weights are set to zero, deviating from restrictions induces no penalty, reducing the estimator to one based solely on non-Gaussianity. In contrast, this study's proposal involves using the data to determine the tuning parameter and weights. Consequently, the cost of deviating from restrictions is determined by empirical evidence rather than by relying on a dogmatic approach that assigns either zero or infinity to the tuning parameter and weights.
The adaptive weights play a crucial role in determining the cost associated with deviating from a given restriction. The proposal of the study is to allow these weights and consequently the cost of deviation to be guided by the data. Let $\hat{B}$ be an initial consistent estimator of $B_0$ obtained by the SVAR-CSUE without penalties. The proposed adaptive weights are defined as
where the first term, $ 1/\hat{B}_{ij}^2$, is driven by the distance of the first step estimator to the zero restriction, and the second term, $\left( 1 + \frac{1}{nG_j^2 + min_{(2)}(nG_1,...,nG_n)^2}\right)$, adjusts the weights depending on the Gaussianity of the shocks corresponding to the first step estimator. The term $nG_j$ quantifies the non-Gaussianity of the first step estimator's $j$-th shock, with $nG_j := S_j^2/6 + (K_j-3)^2/24$ where $S_j$ and $K_j$ denote the skewness and kurtosis of the $j$-th shock, and $min_{(2)}(nG_1,...,nG_n)$ selects the second smallest non-Gaussianity measure of all shocks.
With the adaptive weights, the cost of deviating from a given restriction is determined by the data through the first-step estimator. Consider the scenario where the shocks are sufficiently non-Gaussian, allowing the first-step estimator to consistently estimate all elements of $B_0$. If a restriction is correct, the consistency of the first-step estimator ensures that the first term of the adaptive weight for that restriction converges to infinity. Consequently, it becomes costly to deviate from correct restrictions, and the estimator shrinks towards those restrictions in line with the data. However, if a restriction is not correct, the first-step estimator will not converge to the restriction, and the first term of the adaptive weight for an incorrect restriction will not diverge to infinity. As a result, the estimator can deviate from incorrect restrictions.
The introduction of the second term, which adjusts the weights based on the Gaussianity of the shocks, is designed to ensure proper weights when multiple shocks are Gaussian. In such cases, the first-step estimator can only identify and consistently estimate the impact of non-Gaussian shocks. The remaining Gaussian shocks are only identified up to a rotation of the Gaussian shocks. Therefore, if there are more than two Gaussian shocks, the first-step estimator cannot provide evidence against restrictions on the impact of the Gaussian shocks. In this case, the second term of the adaptive weights leads to an increase in the weights of the restrictions corresponding to Gaussian shocks. This ensures that only those restrictions where the data actually provide evidence against them receive low weights.
Specifically, the Gaussianity correction is constructed such that if two or more (or even all) shocks are Gaussian, the Gaussianity measure of restrictions corresponding to elements in the columns of the Gaussian shocks converges to zero and the weights of these restrictions go to infinity. However, for the restrictions corresponding to non-Gaussian shocks, the non-Gaussianity measure remains finite and scales the weights based on the degree of non-Gaussianity.
The tuning parameter $\lambda$ scales the weights of all restrictions. The tuning parameter is determined by cross-validation with two folds and multiple repetitions. First, I define an increasing sequence of possible $\lambda$ values. Larger values lead to stronger weighting of the restrictions, which in turn induces more shrinkage. Essentially, higher values of $\lambda$ express greater confidence in the validity of the restrictions imposed. For each $\lambda$, I estimate $\hat{B}_T $ using training fold data and calculate the loss in the let-out fold at $\hat{B}_T$.\footnote{ Solving the optimization problems in the cross-validation can be challenging and may require a good starting value. Appendix (ref) describes the details of the cross-validation that involves a rule to derive reasonable starting values for each $\lambda$ and an additional regularization term. } The subsequent task is to select the tuning parameter based on the losses from the cross-validation. The selection process aims to select the largest feasible $\lambda$ value supported by the cross-validation results. To this end, I calculate the median, $60$% quantile, and $40$% quantile loss across repetitions for each $\lambda$ value. Subsequently, I calculate the smallest value $\lambda$ for which the median loss of all subsequent $\lambda$ values is smaller than $60$% quantile loss of all preceding $\lambda$ values and the smallest $\lambda$ value for which the $40$% quantile loss of all subsequent $\lambda$ values is smaller than the median loss of all preceding $\lambda$ values. The tuning parameter is set to the minimum of both values. This approach employs a systematic ascent of the tuning parameter, translating to increased shrinkage, until a pronounced surge in losses becomes apparent, indicating a transition point where additional shrinkage is unwarranted.
The following Monte Carlo study illustrates the benefits of correctly imposed restrictions via the penalty term and sheds light on the estimator's ability to distinguish between correct and incorrect restrictions. The simulations show that correctly imposed restrictions via the penalty term lead to an increase in precision, i.e. a reduced small sample bias and reduced MSE, and an increase in efficiency, i.e. narrower confidence bands. Moreover, the simulations show how the impact of false restrictions decreases with increasing sample size.
I simulate an SVAR with four variables
where the structural shocks are independently and identically drawn from the two-component mixture $ \epsilon_{it} \sim 0.79\; \mathcal{N}(-0.2,0.7^2) + 0.21 \; \mathcal{N}(0.75,1.5^2), $ where $\mathcal{N}(\mu, \sigma^2)$ indicates a normal distribution with mean $\mu$ and standard deviation $\sigma$. The shocks have skewness $0.9$ and excess kurtosis $2.4$.
The simulation compares three estimators. The first estimator, denoted by CSUE, is the two-step SVAR-CSUE estimator and does not use restrictions. The second estimator, denoted by RCSUE($\mathcal{R}_1$), is the Ridge SVAR-CSUE with a penalty on the correct zero restrictions $ \mathcal{R}_1=\{(i,j) \in \{1,...,n\}^2 | j>i \text{ and } i \leq 2 \}$. The third estimator, denoted by RCSUE($\mathcal{R}_2$), is the Ridge SVAR-CSUE with a penalty on the restrictions $ \mathcal{R}_2=\{(i,j) \in \{1,...,n\}^2 | j>i \}$ which impose a recursive structure and thus contain one incorrect restriction. The adaptive weights of the ridge estimators are calculated based on the unrestricted estimator. The tuning parameter $\lambda$ of the ridge estimators is chosen using a repeated cross-validation with two folds, $10$ repetitions, and a sequence of $40$ potential $\lambda$ values. The Appendix contains multiple additional simulations, including simulations with Gaussian shocks, a VAR with lags and shocks with a common volatility process, $A$-type restrictions, and augmented proxy VAR restrictions.
Table (ref) shows the average and MSE of each estimated element and Table (ref) displays the coverage and average length of bootstrap confidence intervals, see Appendix (ref) for details on the construction of bootstrap confidence intervals.
First, the results show how imposing correct restrictions using the ridge estimator leads to an increase of the estimator's precision, meaning a smaller bias and reduced MSE, and to an increase of efficiency, meaning a reduction in the length of the confidence bands, compared to the unpenalized estimator. The most notable improvements are observed in penalized elements, where the MSE and width of confidence bands of the RCSUE($\mathcal{R}_1$) estimator are substantially smaller compared to the unpenalized CSUE. Notably, the penalty also improves the performance of the unpenalized elements of the RCSUE($\mathcal{R}_1$), with MSE approximately three times smaller and the width of confidence bands reduced by half compared to the CSUE.
These findings underscore the ability of the ridge estimator to identify and shrink towards correct restrictions. Moreover, they highlight the utility of restrictions beyond their traditional role in ensuring identification. While CSUE relies solely on statistical properties and imposes minimal structural constraints on the SVAR, resulting in volatile estimates with large uncertainties, the ridge estimator leverages economically motivated restrictions to enhance precision and efficiency.
Secondly, the results shed light on the impact of incorrect restrictions. The $\mathcal{R}_2$ penalty contains a false restriction, which shrinks the $b_{34}$ element to zero, contrary to its true value of five. This invalid restriction induces an increase in bias, MSE, and distorted coverage of the confidence bands for the $b_{34}$ element. At the same time, the correct restrictions lead to a performance increase of the correctly penalized elements and result in a mostly positive impact on the unpenalized elements of the estimator. One crucial distinction from traditional approaches, where restrictions are treated as binding constraints, is that the ridge penalty mitigates the adverse effects of incorrect restrictions, especially as the sample size increases. In small samples, the statistical identification approach may not provide robust evidence against invalid restrictions, causing the estimator to shrink towards them. However, with more data, the approach can detect that shrinking towards the incorrect restriction leads to dependent shocks, resulting in smaller tuning parameters determined by cross-validation. This reduces the impact of incorrect restriction with increasing sample size. Further details on the selected tuning parameters are provided in the Appendix.
The ridge estimator, by design, cannot outright disregard an invalid restriction; instead, evidence against the restriction only decreases the degree of shrinkage towards it. The ability to completely dismiss restrictions can be implemented using an additional step to select restrictions. Table (ref) shows the performance of the RCSUE($\mathcal{R}_2$) using an additional restriction selection step. Initially, the estimator estimates the RCSUE($\mathcal{R}_2$), then identifies all restrictions with estimated elements below an absolute threshold of $0.5$, and subsequently repeats the RCSUE estimation process using only the selected restrictions. The results demonstrate a notable enhancement in the estimator's performance, attributable to its ability to entirely dismiss restrictions inconsistent with the data.
The simulation results provide several insights for practical applications. Firstly, if credible, economically grounded restrictions are available, incorporating them into the estimation process is recommended, as they can significantly enhance the performance of the statistically identified estimator. Second, the simulations highlight that the estimator's ability to detect and disregard invalid restrictions depends on the sample size. Therefore, in situations where specific restrictions are dubious, especially in applications with limited sample sizes, it may be more prudent to refrain from imposing such restrictions. Alternatively, an additional restriction selection step could be used to refine the set of restrictions based on their consistency with the data.
This section analyzes the effects of different approaches to incorporate short-run restrictions on the oil and stock market interaction. The traditional recursive SVAR approach suggests that the stock market does not provide additional information on the price of oil. In contrast, the ridge estimator reveals that the stock market and oil prices cannot be ordered recusively and that information shocks influencing the stock market play are important drivers of oil price fluctuations.
The SVAR uses monthly data from January $1974$ to August $2023$ with
where $q_t$ is $100$ times the log of world crude oil production, $y_t$ is $100$ times the log of global industrial production, $p_t$ is $100$ times the log of the real oil price, and $s_t$ is $100$ times the log of a monthly U.S. stock price index. The data sources can be found in Appendix (ref).
kilian2009impact estimate a similar oil and stock market SVAR and propose to identify four shocks using a recursive order.\footnote{ The model analyzed in kilian2009impact uses a slightly different specification. Specifically, the authors use an economic activity index based on shipping costs. However, as noted by baumeister2022energy, the shipping index may not always be a reliable indicator of changes in global economic activity. Therefore, I follow the approach taken by baumeister2019structural and use a conventional measure of economic activity based on industrial production.
} In the recursive SVAR, oil supply shocks $\varepsilon_{S,t} $ can simultaneously affect all variables, economic activity shocks $\varepsilon_{Y,t} $ cannot simultaneously affect oil supply, oil-specific demand shocks $\varepsilon_{D,t} $ cannot simultaneously affect oil supply and economic activity, and stock market information shocks $\varepsilon_{SM,t} $ cannot simultaneously affect oil supply, economic activity, and the oil price.
Recursive restrictions have two major limitations. First, they imply that oil supply cannot respond simultaneously to demand shocks. Secondly, the reduced form price shocks that cannot be explained by supply and economic activity shocks are, by construction, identified as oil-specific demand shocks. However, if the oil price responds immediately to information shocks affecting stock prices, these information shocks would end up in the oil-specific demand shock of the recursive model. The former issue regarding the response of oil supply to non-supply shocks received a lot of attention in the literature, see, e.g. kilian2012agnostic, kilian2014role, baumeister2019structural, caldara2019oil, and braun2021importance, while the latter issue on the impact of stock market information shocks on the oil price received little attention.
In contrast to the recursive estimator, the proposed ridge estimator does not use restrictions to ensure identification. As a result, it does not require to impose the two questionable assumptions. Instead, the ridge estimator employs the following short-run restrictions:
Labeling of the ridge estimator is determined by the solution of the recursive SVAR. Specifically, I estimate the recursive model using the Cholesky decomposition and use the resulting estimated simultaneous interaction to construct a set of unique-sign permutation representatives centered at the recursive solution, see Section (ref). This set restricts admissible $B$ matrices and determines the labeling: within the set and in line with the recursive labeling, the first shock represents an oil supply shock, the second shock is an economic activity shock, the third shock is an oil-specific demand shock, and the last shock is the stock market information shock. Furthermore, the tuning parameter $\lambda$ required for the ridge estimator is determined similarly to the previous section, using repeated cross-validation with two folds and $50$ repetitions. Lastly, the non-Gaussianity measured by the skewness, excess kurtosis, and Jarque-Bera test of the reduced form and estimated structural form shocks are shown in the appendix. The results indicate that three out of four shocks are left skewed with heavy tails, which is sufficient to ensure identification based on Proposition (ref).
Figure (ref) displays the impulse responses generated by the recursive estimator and the ridge estimator. Although both estimators arrive at similar conclusions on the effects of economic activity shocks, there are notable differences in their findings regarding responses to oil supply, oil demand, and stock market information shocks.
To begin, both estimators find an immediate increase in oil supply and a decrease in oil price in response to the oil supply shock. However, the recursive estimator suggests a smaller response of the oil price and no significant reactions in economic activity and stock prices. In contrast, the ridge estimator shows a larger initial oil price response, a more positive (though not statistically significant) long-run response of economic activity, and a positive and significant response of stock prices after the supply shock.
Secondly, both estimators find a positive response of the oil price to oil demand shocks. In the recursive model, the simultaneous response of oil production to demand shocks is zero by construction. The ridge estimator indicates a positive short-run response of oil production to demand shocks, however, the immediate response is not significant. Furthermore, both estimators find a negative long-term impact on economic activity and stock prices in response to oil demand shocks. However, the recursive model suggests a significant positive response in economic activity in the medium term and a positive reaction in stock prices in the short term. In contrast, the ridge estimator suggests an earlier negative response of economic activity and an immediate negative response of stock prices to the oil price increase caused by oil demand shocks.
Third, both estimators show that the stock market information shock is followed by a subsequent expansion of economic activity and oil production, along with an immediate positive response of stock prices. Nevertheless, the response of the oil price to the stock market information shock differs substantially between both estimators. In the recursive model, the initial response of the oil price to the information shocks is restricted to zero, and the model implies that the shock has no noteworthy impact on oil prices in the medium and long term. In contrast, the ridge estimator does not impose such constraints. In fact, the data provide evidence against the zero restriction and suggest a significant positive response of the oil price to the stock market information shock.
Imposing a restriction that confines the oil price response to the stock market information shock to zero has significant implications for shocks characterized as oil-specific demand and stock market information shocks within the recursive model. In the recursive model, a shock that simultaneously affects the oil price residual unexplained by supply and economic activity shocks is by construction identified as an oil-specific demand shock. Consequently, information shocks about future economic activity and, henceforth, future oil demand, which immediately affect the oil and stock market in the same direction, become subsumed within the category of oil-specific demand shocks in the recursive model. Similarly, oil-specific demand shocks, i.e. those not originating from economic activity shocks that immediately impact both the oil and stock markets in opposing directions, also end up within the oil-specific demand shock of the recursive model. Therefore, the recursive model identifies the oil-specific demand shock as a mixture of oil demand and information shocks. Given that both shocks have opposing impacts on the stock price, this mixture of shocks results in an immediate stock market response that nearly offsets, leading to the conclusion that the stock market exhibits minimal immediate responsiveness to the oil-specific demand shock. Consequently, the stock market information shock is also a mixture of oil-specific demand and information shocks in the recursive model, which leads to the conclusion that information shocks driving the stock market have almost no effect on the oil price.
Table (ref) displays the estimated simultaneous interaction transformed to an $A$-type SVAR comparable to baumeister2019structural, caldara2019oil, or braun2021importance with
where Equation ((ref)) models oil supply, Equation ((ref)) models economic activity, Equation ((ref)) models oil demand, and Equation ((ref)) models the stock market. The oil supply elasticity is equal to zero in the recursive model by construction, whereas the ridge estimator indicates an elasticity of $0.088$, close to the location of the prior used in baumeister2019structural. The oil demand elasticity in the recursive model is close to minus one, whereas the ridge estimator yields a demand elasticity of $-0.325$, which is close to the median posterior demand elasticity in baumeister2019structural. Turning to the income elasticity of oil demand, the ridge estimator suggests a value of approximately $0.7$, which is equal to the location of the corresponding prior in baumeister2019structural, while the recursive estimator indicates a significantly higher value. Moreover, the recursive estimator finds that the oil demand response to the stock market and the stock market response to economic activity do not differ significantly from zero, whereas the ridge estimator indicates a positive response of oil demand to the stock market, a positive stock market response to economic activity. Lastly, the recursive estimator suggests a positive effect of the oil price on the stock market, whereas the ridge estimator suggests the opposite.
Table (ref) provides insights into the effect of recursiveness restrictions on the forecast error variance decomposition. In the recursive model, oil-specific demand shocks are the primary driver of the oil price, explaining more than $80$% of the variation. Conversely, the ridge estimator unveils a less one-sided picture, indicating that oil-specific demand shocks, although still the primary influence, explain a reduced share of only $36$% of the variance. The remaining variance in oil prices is attributed to oil supply and stock market information shocks, both contributing approximately $25$% to the overall variation. Moreover, disentangling oil-specific demand and stock market information shocks based on their interdependence, rather than relying on restrictions, also has a substantial impact on the variation of stock prices explained by oil-specific demand shock. In the recursive model, oil-specific demand only explains $2$% of the variation in stock prices, while the ridge estimator finds a more pronounced influence of oil-specific demand shocks.
Figure (ref) shows the historical decomposition of the oil price and sheds light on the importance of supply, demand, and information shocks in different periods. In the recursive SVAR, oil-specific demand shocks are the primary driver of the oil price. This pattern is consistent during events such as the collapse of OPEC in $1985$, the Persian Gulf War in $1990$, the oil price surge in $2007-2008$, the subsequent decline and recovery in oil prices following the collapse of Lehman Brothers in $2008$, the oil price downturn in $2014-2016$, the oil price fluctuations at the start of the COVID-$19$ pandemic, and the recent oil price increase in $2022$ following the Ukraine invasion. In contrast, the ridge estimator provides a more nuanced picture. First, it suggests that supply shocks played a more prominent role during the collapse of OPEC in $1985$, the Persian Gulf War in $1990$, and the upswing of oil prices in $2007-2008$. Second, it suggests that information shocks extracted from stock prices contributed largely to the increase in oil prices before $2008$, the decrease in the oil price following the collapse of Lehman Brothers in $2008$, the decrease in the oil price in $2014-2016$, and also to the decrease and recovery of the oil price at the beginning of the COVID-$19$ pandemic.
Overall, the analysis suggests an immediate response of the oil price to stock market information shocks. In addition, these information shocks play a significant role in explaining oil price fluctuations. In general, the results highlight the importance of incorporating information from the stock market into the analysis of oil price movements.
Economically motivated short-run restrictions have been an integral part of identifying SVAR models in numerous applications since sims1980macroeconomics. Despite their popularity in applied work, the disadvantageous of restriction based identification methods are well known: incorrect restrictions lead to biased estimates. Novel identification approaches that rely on stochastic properties of the shocks no longer require short-run restrictions for identification and applications of these approaches oftentimes completely disregard available economically motivated short-run restrictions or only conduct hypothesis tests of restrictions.
This study proposes a new approach to combine the rich literature on short-run restrictions with the recent statistical identification literature. By incorporating restrictions via a shrinkage approach alongside non-Gaussian based identification, the estimator combines the strengths of both methodologies. Simulations show how valid restrictions improve the accuracy of the statistically identified SVAR estimator and that the estimator can detect and reduce the impact of incorrect restrictions as the sample size increases. Therefore, the study underscores the enduring value of over four decades of research into plausible short-run restrictions within the statistical identification framework, where such restrictions are no longer required for identification.