Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,088 characters · 22 sections · 128 citation commands
Wild inference for wild SVARs with application to heteroscedasticity-based IV
Unexpected actions by policymakers provide a valuable source of variation that helps to identify how the economy would react to future policy actions sims1980comparison. Unfortunately, it is often difficult to detect the unexpected component of the actions using publicly available data. The true plan of future actions and commitments of the policymakers can be kept secret from the public, and the pure unexpected part of the available policy announcements is always a matter of modelers' assumptions. We propose an identification approach that remains agnostic about the actual size and direction of change in the policy variable (for example, interest rate) due to policymakers' discretion. Instead, we only assume that the conditional variance of the unexpected change in the policy variable is higher during policymakers' meetings than at other times. We show that the information contained in the meeting dates is sufficient to uniquely identify both the size and direction of impulse responses to the structural policy shock in a vector autoregression model (VAR).
Our paper applies this novel approach to the identification of monetary policy shocks using the number of US Federal Reserve Open Market Committee (FOMC) meetings per month as an instrument for heteroscedasticity of a monetary policy shock in a standard six-variable VAR model. The information available on the meeting dates goes back to at least 1965, which gives us instruments with sample sizes that were unprecedented in the monetary SVAR literature. Such long samples come with an additional set of complications.\footnote{According to ramey2016macroeconomic, the longer sample period, on the one hand, includes more variation in interest rate and thus advantageous for monetary policy shock identification; on the other hand, it is accompanied by extra challenges such as structural breaks, regime shifts, state-dependence, and sign-dependence.} In particular, the inference procedure needs to be robust to the potential presence of time trends and unit roots, which makes the existing popular methods like the residual-based bootstrap gonccalves2004bootstrapping inapplicable. Moreover, the use of external instruments for conditional heteroscedasticity requires a novel procedure that allows joint inference on all the relevant parameters in the model.
We address this challenge by proposing a novel inference procedure for structural impulse responses (IRF) based on local projection (LP) estimators of jorda2005estimation. Specifically, we propose both a formula and a proof of joint asymptotic normality of the local projection estimator in the potential presence of trends, cointegration, any number of unit roots, and any weakly-dependent stationary conditionally heteroscedastic innovations. Next, we extend the proof of the consistency of the joint HAC covariance estimator of andrews1991heteroskedasticity to accommodate non-stationary regressors. Finally, we propose a version of the dependent wild bootstrap procedure of shao2010dependent that can mimic the joint distribution of the local projection estimators across multiple horizons and the estimates of structural parameters. The only regularity conditions that we impose on the non-explosive VAR model is that the one-step-ahead forecast errors, jointly with the instrumental variable, follow a stationary vector process with a mixing property and existence of certain higher-order moments, which are necessary for consistency of the HAC estimator. To our knowledge, this provides the first proof of the validity of inference on the IRF estimators that allows for the general case of unstable but nonexplosive VAR models with a general form of stationary conditional heteroscedasticity without pretests.\footnote{Our results are built upon the long literature on non-stationary time series including sims1990inference,tsay1990asymptotic,de2000functional,inoue2002bootstrapping,montiel2021local.}
The proposed procedure consists of three main steps. First, we estimate reduced-form impulse response functions (i.e. responses to one-step-ahead linear forecast errors $\hat\eta_{it}$ for each time series $i$). This step is computed using local projection regressions. Second, we estimate the contemporaneous impact of the structural shock of interest on time series $i$ as the sample covariance between pairwise products of one-step-ahead forecast errors $\hat\eta_{1t}\hat\eta_{it}$ and an instrumental variable that is correlated with the heteroscedasticity of the structural shock of interest. (The time series with index $1$ plays a special role since it is assumed to have a non-zero correlation of squared forecast errors with the instrumental variable.) Third, we compute structural IRF for other horizons as products of the reduced-form IRF and vector of contemporaneous structural responses after appropriate normalization. The dependent wild bootstrap is then applied to the relevant score functions.
The novel simultaneous inference on local projections has multiple uses. First, as emphasized in bruggemann2016inference, it is important to make joint draws of the vector of the IRF coefficients and the structural impact vector to obtain valid confidence limits. Second, the procedure allows simultaneous confidence bounds that cover the entire IRF across all horizons and, for example, allows for inference on the largest and smallest values of IRF. Third, one can compute smoothed local projection estimators and the corresponding confidence bounds using the minimum distance approach to improve the efficiency of the estimates.
We illustrate the proposed method in two applications. First, we show that the identification strategy can recover the true IRF of a monetary policy shock in simulated data based on a medium-sized DSGE model of the US economy of smets2007shocks. Then we apply the method to study the impact of monetary policy shocks in a six-variable VAR model using the data set of antolin2018narrative. The variables include Real GDP, GDP deflator, commodity price index, total reserves, non-borrowed reserves, and the Federal funds rate. The sample period is January 1965 to November 2007. Our primary instrumental variable for conditional heteroscedasticity of the monetary policy shocks is the number of FOMC meetings per month, including telephone conferences and unscheduled meetings. For reduced-form estimates, we use local projection with polynomial time trends with power up to $k\in\{0,4\}$ and the number of lags $p\in\{12,24\}$. We report various confidence bands: point-wise, simultaneous sup-t, smoothed point-wise, and smoothed simultaneous bands.
Our main empirical findings can be summarized as follows: (i) the identification through the heteroscedasticity IV results in IRFs that are nearly identical to the classical Cholesky identification a la christiano1999monetary; (ii) smoothed LP-IRFs are noticeably different from the conventional recursive VAR estimates at the long horizons; (iii) effects on real GDP are comparable with BVAR identified using the narrative sign restriction effects antolin2018narrative, while the effects of the monetary shocks on inflation are much weaker in the novel scheme.
The literature on identification through heteroscedasticity has a long history (for example, feenstra1994new,rigobon2003identification,lewis2021identifying,brunnermeier2021feedbacks). To our knowledge, we are the first to propose an IV-type procedure that assumes that the conditional heteroscedasticity is not persistent, or structural shocks are fully independent and non-Gaussian lutkepohl2017structural. Moreover, we do not assume any particular parametric model of the innovations and the instruments. In contrast to statistical identification procedures, our approach does not require full identification of all shocks and gives an intuitive interpretation of the identified shock of interest.
In this paper, we apply our identification method to monetary policy shocks, but the heteroscedasticity-based external IV can be proposed in other applications.\footnote{There is a long-standing literature devoted to the identification of monetary policy shocks, including but not limited to christiano1999monetary, romer2004new, gurkaynak2005sensitivity, gertler2015monetary, stock2018identification, nakamura2018high, bauer2021alternative, bauer2022reassessment, bu2021unified, swanson2023speeches.} For example, kanzig2021macroeconomic shows that OPEC meetings have high-frequency impacts on oil prices that can be used to identify oil supply shocks; janzen2018commodity show that cotton storage levels change the heteroscedasticity of precautionary demand shocks. Therefore, we believe that our approach can be used more broadly.
Our inference procedure is based on the insight that both the reduced-form impulse response and covariance estimators involved in the structural identification can be represented as a mean of vector-valued stationary time series of cross-products of the $1$-step-ahead and $h$-step-ahead forecast errors. This representation allows us to employ the dependent wild bootstrap of shao2010dependent to make draws from the joint asymptotic distribution of the estimators without explicitly computing the long-run variance-covariance matrix estimator. This novel inference procedure can be used for many other identification approaches that include Cholesky christiano1999monetary, sign restrictions gafarov2016projection, and external IV stock2008what,mertens2013dynamic.
The paper is structured as follows. Section (ref) discusses the setup and the novel identification idea. Section (ref) provides the estimation and inference algorithms with the corresponding theorems. Section (ref) discusses the DSGE experiment. Section (ref) has an empirical application to US monetary policy. Section (ref) concludes. All the proofs and additional figures are provided in the appendix.
Consider a simultaneous equations model
where observations $\eta_t\in \R^n$ and $\varepsilon_t$ is an i.i.d. zero-mean process with $\E \varepsilon_t \varepsilon_t'=I_n$ that is referred to as structural shocks.\footnote{The simultaneous equations model with lags was first proposed and studied in haavelmo1943statistical,mann1943statistical.} Suppose that we are interested in measuring the impact of the structural shock $\varepsilon_{1t}$ on the vector of observables $\eta_t$, which is given by the first column of $B=(b_{\cdot1},\dots, b_{\cdot n})$. We can estimate $\Sigma_\eta \bydef \E \eta_t \eta_t' $ using the sample analog and then solve equation $\Sigma_\eta= B B'$ for the first column of $B$. However, there is a continuum of solutions to this equation. Therefore, additional assumptions or data are required to uniquely select $B$. The paper aims to propose a novel way of identifying $b_{\cdot1}$ using external instruments for the conditional heteroscedasticity process of $\varepsilon_t$ and to develop the inference procedure. This identification problem occurs in the more general SVAR models, which we will study later in the paper.
Before we generalize the novel identification proposal, let us consider a simple motivating example.
The identification conditions in Example (ref) can be generalized to
This assumption implies that the heteroscedasticity of the innovations for this time series must be correlated with the instrument $Z_t$. In the monetary shock example, one can by definition assume that the impact $b_{11}$ is nonzero for the federal funds rate ($i=1$). The instrument can only affect the conditional variance of the shock of interest and cannot impact the conditional correlation of the other structural shocks.
A recent work lewis2021identifying showed that in the case when heteroscedasticity of a structural shock is persistent, one can use internal instruments $Z_t = \eta^2_{t-1}$. However, the approach can be applied more generally if appropriate external instruments are available, as in our empirical application.
In most practical applications, data $y_t$ have a time dependence and innovations $\eta_t$ have to be estimated, which is the topic of the next section.
Suppose that we observe a $n-$vector time series $y_t$ that follows a VAR(p) model,
where $y_t$ is $n$-dimensional time series, $ V \mu_t$ is a deterministic polynomial time trend (or time-varying drift for integrated processes) with $\mu_t=(1,t,t^2,\dots,t^k)'$ for some $k$; $ A_1,\dots,A_p,V$ are $(n\times n)$ parameter matrices, $\eta_t$ are serially uncorrelated forecast errors (innovations). We will assume that $p$ is finite and known.
Assumption (ref).1 implies that innovations are white noise, which is required for the consistency of LS estimators in stationary VAR models. The martingale difference requirement with respect to sigma field $\mathcal{F}_{t-1}=\sigma(y_{t-1},y_{t-2},\dots)$ is only necessary for interpretation of the impulse response coefficients that will be introduced later in this section. Assumption (ref).2 allows for many stationary models of conditional heteroscedasticity.\footnote{See, for example, Assumption 2.1.(i-iii) in bruggemann2016inference. This assumption, for example, allows for asymmetric conditional heteroscedasticity models like the ones considered in rabemananjara1993threshold. Such models are ruled out by the popular residual wild bootstrap methods gonccalves2004bootstrapping.} We do not consider unconditional heteroscedasticity in innovations, since it would eliminate the need for structural identification. The 8-th moment condition imposed in our paper is sufficient for the consistency of the HAC estimators, but a weaker 4-th moment condition would be sufficient for the asymptotic normality of the LP estimators. Assumption (ref).4 is standard in papers on structural identification of the shocks. Assumption (ref).5 is required for asymptotic normality of $\hat\gamma$ and would be satisfied, for example, if $\eta_t$ is MDS with respect to a sigma field $\mathcal{F}_{t-1}$ that also includes $Z_t$ and past values of $\eta_t$ (for example, it is possible that the number of meetings in month $t$ is fully predictable based on the past economic conditions and the past number of meetings). This assumption is not required for the consistency of $\hat\gamma$ and the structural IRF.
Using the VAR model, one can define impulse responses for individual innovations $\eta_t$ since
where $e_j\in R^n$ is a vector of zeros with $1$ in $j-$th position. These coefficients appear as regression coefficients in the $H$-step ahead autoregression representation of (ref),
where $C_0=I$, $ C_h = \sum_{\ell=1}^h A_\ell C_{h-\ell}$ (assuming $ A_\ell=0$ for $\ell>p$), matrices $(V{(H)},A_1{(H)},\dots,A_p{(H)})$ are functions of $(V,A_1,\dots,A_p)$ that can also be computed recursively.
In the absence of deterministic trend and under conditions $det(I_n-A_1z-\dots-A_pz^p)\neq0 $ for $|z|\leq1$ and $V=0$, the VAR model can be inverted to obtain the familiar causal $MA(\infty)$ representation wold1938study,
In this case, coefficients $C_h$ coincide with the coefficients in the Wold representation. We do not need to rely on this stability assumption throughout the paper. Instead, we allow for an arbitrary number of unit roots and cointegrating vectors.
Assumption (ref) allows for an arbitrary number of unit roots, including complex unit roots, and cointegrating relations between $y_t$. It only rules out explosive roots. In practical terms, it means that one can apply our inference procedure to local projections estimated with or without a trend in levels without seasonal adjustment. It can be done without any pretests for trend stationarity, absence of cointegration, or unit roots. We allow both a polynomial trend as in sims1990inference and complex unit roots as in tsay1990asymptotic. Under this assumption, the least squares estimator for the coefficients in equations (ref) and (ref) is consistent and has well-defined distribution limits, as shown in sims1990inference,tsay1990asymptotic under conditionally homoscedastic MDS innovations $\eta_t$. In contrast to the earlier studies, we allow for more general (weak) white noise innovations with mixing by using strong laws of large numbers and functional central limit theorems for strong mixing sequences de2000functional,rio2017asymptotic.
Using the coefficients $C_0,\dots,C_H$ under Assumption (ref).1, we can compute structural impulse response function of $y_{t+h}$ to $\varepsilon_{1t}$, which is given by
The coefficients $C_1,\dots,C_H$ can be estimated using local projection estimators of jorda2005estimation. However, our empirical strategy uses conditional heteroscedasticity in the structural shocks in the dataset that has unit root time series with potential cointegration and time trends. Furthermore, structural impulse responses depend on both estimators of $(\Sigma_\eta,C_1,\dots,C_H)$ and the covariances of external instruments $Z_t$. Existing inference methods cannot take into account all of these features together. Specifically, there are no proposals for joint inference of $\Sigma_\eta$ and $C_0,\dots,C_H$ without pre-tests for unit roots.\footnote{bruggemann2016inference provide an example where the commonly used residual wild-bootstrap is invalid for joint inference on $C_i$ and $\Sigma_\eta$ even in the stationary case. Their proposed procedure, residual block bootstrap, solves the issue for stationary VARs. jentsch2022asymptotically then extends this residual block bootstrap approach to allow external instrumental variables, also in the stationary case. inoue2002bootstrapping,mikusheva2012one allow inference on individual components of $C_h$ in the presence of a unit root of a particular type, but impose an i.i.d. assumption on the innovations which is violated in our application.}
Consider the following example that illustrates the distinction between the single coefficient and joint inference on impulse responses in the nonstationary case.
In the following sections, we propose an estimation algorithm based on LP, and establish consistency and joint asymptotic normality of LP estimators under Assumptions (ref)-(ref). Finally, we propose and prove the validity of a novel dependent wild bootstrap procedure under the maintained assumptions. This inference procedure provides a way to conduct simultaneous inference on structural IRF for all horizons without pre-tests.
Our goal is to estimate the vector of parameters
and then compute the functions of this parameter.
The first step for our analysis is to estimate $C_h$ for $h=1,\dots,H$, and $\Sigma_\eta$ using the following hybrid local projection estimator.
Algorithm 1
Algorithm 1 gives estimates of $C_h$ for the first $H_1$ horizons. These estimates are identical to the local projection estimator of jorda2005estimation as Theorem (ref) below shows. Since we are interested in simultaneous inference on all relevant parameters of an SVAR, this algorithm imposes the VAR(p) structure on impulse responses of longer horizons, $H_1\leq h\leq H_2$. In doing so, we reduce the problem of inference about IRF at long horizons to the problem of simultaneous inference on a finite set of parameters $C_1,\dots,C_{H_1} $. This is a much easier problem, and our bootstrap procedure, which will be outlined later, can address it.
Note that, in general, the backward recursion estimators of $A_i$ are different from the corresponding LS estimators. The LS estimates may be super-consistent for some linear combinations of $A$, but come at the expense of non-standard asymptotic distributions that are hard to use for confidence sets mikusheva2012one. In contrast, $ \hat A^{BR}_i $ are polynomial functions of asymptotically Gaussian estimators $\hat C_h$. Therefore, a bootstrap procedure can estimate the bias and variance of $ \hat A^{BR}_i $. This approach allows us to bypass any need to pre-test data for the presence of deterministic trends, unit roots, and cointegration relations and avoid unnecessary data transformations. Such pretests may introduce pretesting biases and reduce efficiency of the resulting estimators if the transformations turn out to be redundant.
Using estimates $ \hat C_h$ and $\hat\gamma$, we can apply the following algorithm to calculate the structural IRF.
Algorithm 2
Algorithm 2 is consistent for the estimation of structural IRF for the shock of interest as follows from
In the previous section, we obtained the asymptotic Gaussian distribution and the corresponding covariance matrix for $\hat \theta (H_1)$. This covariance matrix can be estimated using heteroscedasticity and autocorrelation consistent (HAC) covariance estimator as the following theorem shows.
The variance-covariance matrix $\Omega$ can be estimated using the HAC estimator along the lines in andrews1991heteroskedasticity. We introduce
and
for some of kernel function $K(\cdot)$ and bandwidth sequence $B_T$.
Even though we had to modify the original proofs of andrews1991heteroskedasticity to account for the nonstationarity of $y_t$, the asymptotic theory for optimal selection of kernels and bandwidths still holds. This is because the infeasible analog of $\hat\Omega(H)$ that uses the true $H$-step innovations $\eta_t(H)$ has the same properties regardless of the presence of unit roots and deterministic trends in the VAR model. In particular, one can use the conventional rule $ B_T= \{0.75 \sqrt[3]{T}\}$ that is appropriate for the Bartlett kernel.
Recent literature emphasized the fact that when data is very persistent, HAC estimators do not perform well lazarus2021size. In particular, in the near-to-unity asymptotics, one should use a very large bandwidth $B_T$ to control the size of $t$ tests. This is not the case under our Assumption (ref).2 (weak dependence) for the innovations $\eta_t$. If, in addition, the innovations are two-sided MDS, then montiel2021local showed that one can use $B_T=0$ for inference on individual coefficients $\hat C_h$. However, even under these stronger assumptions, the joint inference on the entire vector $\theta(H_1)$ still requires $B_T$ to grow with the sample size to account for the potential presence of conditional heteroscedasticity of unknown form.
Using the Delta method, an asymptotic normal distribution can be derived for other estimators, including $\hat A^{BR}_j$ and $\hat C^{BR}_h$ at long horizons $h$.\footnote{ Another function of interest is the FEVD. gorodnichenko2020forecast studied inference on FEVD using local projections and stationary VAR-bootstrap under assumption of homoscedasticity of shocks. They show that block bootstrap of kilian2011reliable can also be used for inference on FEVD under the assumption of stationary data. The dependent wild bootstrap approach can also be applied to and $\widehat{FEVD}(h)$ without requiring stationarity of the VAR model.} The dependent wild bootstrap approach of shao2010dependent can be used as a convincing numerical implementation of the delta method. A particular way to perform bootstrap is given in the following algorithm.
Algorithm 3
This algorithm is consistent for estimating the asymptotic distribution of $\hat\theta$ as the following theorem shows.
To our knowledge, this is the first joint bootstrap resampling procedure for LP estimators. The validity of the dependent wild bootstrap is based on the fact that it mimics the asymptotic Gaussian distribution of $\hat \theta(H_1)$ with a covariance matrix that coincides with the HAC estimator given in Theorem (ref). This argument does not impose any parametric model on $\eta_t$ and $Z_t$ and is generally applicable under the maintained Assumption (ref).2 (weak dependence of innovations). Note that the algorithm uses $\hat\Xi_t$ that depend on possibly nonstationary processes with strong dependence, $Y_t$. As a result, conventional block bootstrap procedures are not directly applicable to $\hat\Xi_t$. Unlike the popular recursive residual-based bootstrap procedures of gonccalves2004bootstrapping, we do not assume stability (or stationarity) of $Y_t$ and allow for asymmetric conditional heteroscedasticity in $\eta_t$.
The primary use of the joint inference procedures outlined in Theorems (ref)-(ref) is simultaneous confidence bounds that cover the entire IRF for a particular time series over all horizons with a given nominal probability. The problem has been extensively studied in stationary VAR models jorda2009simultaneous,inoue2016joint,montiel2019simultaneous. The dependent wild bootstrap can be used to construct $\sup t$ confidence bounds using the following algorithm. This procedure can be used for joint inference on any smooth functions $f_h(\theta)$ (for example, structural IRF($h$) over multiple $h=1,\dots,H$):
Algorithm 4
In large samples, by the delta method, the bootstrap distribution of $f_h(\hat{\theta}^s)$ is approximately Gaussian. As a result, the difference between the $q(f_h(\hat{\theta}^s),\Phi(1))$ and $q(f_h(\hat{\theta}^s),1-\Phi(1))$ quantiles is converging in probability to 2 standard deviations of $f_h(\hat{\theta})$ as $T,S\to\infty$. Step 1 can be replaced with a conventional bootstrap standard error estimator if the sample size is sufficiently large.\footnote{Substitution of interquartile range for standard deviation is frequently used in the bootstrap literature chernozhukov2020generic. See, for example, Remark 3.2 of chernozhukov2013inference which uses the insights from kato2011note. } Another advantage of using the s.e. estimator based on the quantiles is that the simultaneous bounds coincide with the $68\%$ point-wise bounds if we set the critical value $CV_\alpha=1$.
An alternative use of Algorithm 4 is a test of misspecification. Suppose that we have two competing identification assumptions, for example, heteroscedasticity IV and Cholesky. We can construct a joint confidence set for the vector of differences of impacts under these two schemes, $\hat b^{Cholesky}_{.1} - \hat b^{HetIV}_{.1}$. Alogrithm 4 can then be used to compute $\sup t$ critical values for a test of hypothesis $b^{Cholesky}_{.1} - b^{HetIV}_{.1}=0$.
We illustrate both uses of Algorithm 4 in Sections (ref) and (ref).
Throughout the paper, we assume that the underlying process that generates the data follows a finite order VAR($p$) model. This model imposes particular relations on the matrices $C_h$, $h=1,\dots,H$ that can be used to improve efficiency (i.e., $C_0=I$, $ C_h = \sum_{\ell=1}^h A_\ell C_{h-\ell}$ with $ A_\ell=0$ for $\ell>p$). The backward recursion estimator $ \hat A^{BR}_i$, introduced in Algorithm 1, imposes the VAR structure on $C_h$ with $h=1,\dots,p$. However, it does not use the overidentifying information about $C_h$ for $h>p$. One can improve efficiency of the LP estimates using the following minimum distance estimator. Take any positive definite matrix $\hat W $ that converges to a positive definite matrix $W$ of dimension $n^2\times H_1$ and minimize
where the vector of the local projection estimators is denoted as $\hat g^{LP} = vec (\hat C_1,\dots,\hat C_{H_1})$ and $ g(A_1,\dots,A_p) = vec (C_1(A_1,\dots,A_p),\dots,C_{H_1}(A_1,\dots,A_p))$ is the corresponding recursively computed IRF coefficients for any values of matrices $(A_1,\dots,A_p)$. From the theory of the minimum distance estimators (see, for example, hansen2022econometrics) the optimal choice of $\hat W$ corresponds to the inverse of the long-run covariance matrix, given in Theorem (ref) earlier in this paper. In small samples, however, the HAC estimator given in Theorem (ref) may result in a non-invertible matrix that needs to be regularized. In this paper, we instead use a diagonal matrix $\hat W$ with the vector of inverse variances of the components of $\hat g^{LP} $ on the diagonal. These variances are estimated using the dependent wild bootstrap.
Based on the estimates of minimum distance $ ( \hat A^{MD}_1,\dots, \hat A^{MD}_p)$, we can compute smoothed impulse responses for any horizon $H$, $$ (\hat C^{MD}_1,\dots,\hat C^{MD}_{H})\bydef C_1((\hat A^{MD}_1,\dots,(\hat A^{MD}_p),\dots,C_{H}((\hat A^{MD}_1,\dots,(\hat A^{MD}_p),$$ using the usual recursive formulas.
Under the correct VAR($p$) specification, the limit of (ref) coincides with the true parameters $A_1,\dots,A_p$. Moreover, the estimator $(\hat A^{MD}_1,\dots,\hat A^{MD}_p)$ is a smooth transformation of $\hat \theta$. So, by the delta method, the dependent wild bootstrap from Theorem (ref) remains valid for $\hat C^{MD}_i$. By implication, we can use it both for point-wise inference and simultaneous inference on $C_h$.
It will be evident in Section (ref), that the smoothed LP estimates of the IRF are substantially different from the conventional recursive VAR estimates at long horizons (see also Appendix Figure (ref)). It should also be noticed that conventional recursive VAR estimates of IRF based on the least squares approach are only asymptotically efficient under homoscedastic and Gaussian innovations $\eta_t$ hamilton2008macroeconomics. In the presence of conditional heteroscedasticity, the smoothed LP estimators may be more efficient than the conventional recursive VAR estimators of $C_h$.
The idea of smoothing the local projection is not new. barnichon2019impulse imposed a particular flexible form on the LP estimates within a ridge-regression framework without providing the corresponding inference procedure. In contrast to earlier studies, we use the VAR structure of the IRF for smoothing. This allows us to improve efficiency under correct specification and use the dependent wild bootstrap as a robust inference procedure.
In this section, we would like to illustrate how heteroscedasticity IV identification works with the dependent wild bootstrap using data created by a dynamic general equilibrium (DSGE) model.
For the data generation process, we will use the well-known model of smets2007shocks (SW later on), a medium-scale New Keynesian model with a wide range of nominal and real rigidities. In particular, the model introduces nominal price and wage rigidity, consumption habits, investment adjustment, and capital utilization costs. At the same time, the Taylor rule for the interest rate assumes that the monetary authority reacts to the deviation of output from the level that would prevail in an economy with flexible prices and wages, which further increases the dimension of the problem because this requires additionally describing the dynamics of a hypothetical economy with flexible prices.
The state vector in the model has a dimension of about 20 variables. As in the original paper, we introduce seven structural macroeconomic shocks into the model, including the monetary policy shock. Typically, the model is estimated using the Kalman filter with the observation vector consisting of seven variables (the number of variables cannot be less than the number of shocks): the log difference of real GDP, real consumption, real investment, and the real wage, log hours worked, the log difference of the GDP deflator, and the federal funds rate. The filter is used to estimate the linear state space model that corresponds to the solution of the log-linearized DSGE models. The unobserved state vector follows a SVAR model. Since the observation vector has a smaller dimension than the state vector (7 vs. 20), the observation vector is not generally a finite-order vector autoregression. However, in a fairly general case, the vector of observed variables can be represented as a VARMA process, and, assuming the invertibility of the MA part, a VAR of finite order can act as an approximation for the corresponding VAR($\infty$).\footnote{The quality of approximation by a finite-order VAR depends on the specific structure of the model, the macroeconomics variables used as observables, and on the persistence of shocks ravenna2007vector,Fernandez2007.}
In the experiment of this section, we generate samples of various lengths from the log-linearized SW DSGE model (equivalent to a VARMA model) calibrated using the posterior mean estimates obtained in smets2007shocks. The only exception was the standard deviation parameter of the monetary policy shock, for which we introduce two regimes described by a Markov chain. We consider a symmetric case where the probability of remaining in one of the regimes is 70%, and the probability of switching to the other regime is 30%. We introduce an auxiliary instrumental variable $Z_t$, which takes the value 0 in the first regime and 1 in the second. The value of the standard deviation of the monetary policy shock, $\sigma_{r,t}$, is set according to the following equation:
where $\sigma_{r,0}$ is set to a value of the posterior mean estimate for the standard deviation of the monetary policy shock in the smets2007shocks paper. This specification assumes that the standard deviation of the monetary shock increases by a factor of four in the second regime.
Next, using the simulated dataset of seven observable variables and the instrument $Z_t$, we apply our new identification method to extract the structural monetary policy shock and compute the corresponding IRFs. We include in the vector of observed variables in the local projection model the following variables: log deviations from the steady-state level for real GDP, household consumption, investment, labor, wage, GDP deflator index, and the gross nominal interest rate.\footnote{ We use the GDP deflator index rather than growth rates in the econometric exercise in order to emphasize that the method can handle both stationary and nonstationary variables.} Thus, the model will use six stationary variables and one nonstationary variable. The original calibration of the model is based on the quarterly frequency of observation. The model includes four lags, which corresponds to a lag of one year. $Z_t$ is used as an instrument for the heteroscedasticity of the monetary policy shock.
Finally, we compare the IRFs generated from the theoretical SW model (VARMA) to those obtained from the estimated VAR model to validate our method's accuracy and effectiveness in capturing the dynamics prescribed by the SW framework. Figure (ref) shows the estimation results for 160 periods (40 years), and Figure (ref) shows the results for 1600 periods. In Figure (ref), it can be seen that the method can reconstruct the IRF for monetary policy shocks within the approximate VAR(4) model using a typical sample size. The true IRFs are well within the simultaneous 68% confidence bounds. As we increase the sample size to 1600 periods (Figure (ref)), the LP estimates get even closer to the true values. This result shows that LP estimators based on finite VAR(4) can reasonably well approximate the underlying VARMA model. Moreover, the simultaneous confidence bounds successfully covered the true IRFs in all the figures, including the non-stationary GDP deflator index.
The identification of monetary policy shocks is a classic application of structural VAR models that goes back to sims1980comparison. This application has been extensively studied with many identification schemes being proposed, including the classical Cholesky decomposition sims1980comparison,christiano1999monetary, instrumental variables romer1989does,romer2004new,gertler2015monetary, sign restrictions uhlig2005effects,gafarov2018delta,antolin2018narrative, high-frequency identification faust2004identifying, and regime-switching models sims2006were,lutkepohl2017structural, to name a few.\footnote{For foundational context and a detailed examination of prior empirical efforts we refer the reader to seminal works by christiano1999monetary, boivin2010has, ramey2016macroeconomic, which provide comprehensive reviews on the subject.} Such an abundance of benchmarks provides an opportunity to test the novel identification approach and validate or reject selected existing alternative methods.
To compare our identification strategy with those commonly used in the literature, we focus on the reduced-form model studied in seminal papers including, but not limited to christiano1999monetary, uhlig2005effects, antolin2018narrative. Our specification includes six variables for the U.S.: real GDP, GDP deflator, commodity price index, total reserves, non-borrowed reserves, and federal funds rate, as described in Table (ref). The sample covers the period from January 1965 to November 2007, with monthly data similar to antolin2018narrative.\footnote{The monthly series for real GDP and the GDP deflator are interpolated from quarterly data using the chow1971best methodology. We use the industrial production index for the real GDP interpolation and the consumer price index for the GDP deflator.} Our specification includes a constant and a polynomial time trend with power up to $k\in\{0,4\}$ and the number of lags $p\in\{12,24\}$.\footnote{When comparing our results to those from the literature, we align the number of lags and the inclusion of trends to match their specifications. For example, christiano1999monetary use constant but without time trend, with lag lengths of 6 and 12, while uhlig2005effects, antolin2018narrative use 12 lags without a constant or deterministic trend.}
It is well documented in the high-frequency identification literature that announcements made after FOMC meetings immediately impact financial variables. These changes in financial indicators are associated with unforeseen changes in monetary policy, commonly referred to as monetary policy shocks.\footnote{ See, for example, cochrane2002fed, faust2004identifying, gurkaynak2005sensitivity, gertler2015monetary, ramey2016macroeconomic, stock2018identification, nakamura2018high, bauer2021alternative, bauer2022reassessment.}
We focus exclusively on the dates of FOMC meeting announcements, including meetings, telephone and video conferences, unscheduled meetings, and sequential day meetings. Our baseline instrumental variable indicates the number of meetings that occur in a given month, including telephone conferences and unscheduled meetings.\footnote{An alternative instrumental variable is a dummy with ones in the months when the meetings occur and zero otherwise. Results are available upon request.} Our instrument remains agnostic about the sign and size of the policy change. What matters is the event itself.\footnote{We construct series using the FRASER St. Louis Fed website https://fraser.stlouisfed.org/title/federal-open-market-committee-meeting-minutes-transcripts-documents-677?browse=1980s} Figure (ref) plots the monthly time series of our baseline instrument starting from 1965 till 2007. As can be seen, the number of meetings in a given month varies from zero to four.
Most of the papers that use FOMC meetings for identification purposes focus their analysis on the sample starting from 1994 onwards. This is because the FED released no public announcements about monetary policy decisions prior to 1994. Some scholars swanson2023speeches extended the sample back to 1988. As our identification assumption is agnostic about the direction and size of the change, we can extend our sample back to 1965.
Figure (ref) displays the impulse response functions of six endogenous variables to the structural monetary policy shock. We estimate those responses using the heteroscedasticity-IV LP estimator in our benchmark specification with $p=12$ and $k=0$ (without a trend). The dotted lines represent 68% point-wise confidence intervals, while the dashed lines are simultaneous sup-t bands. Although simultaneous confidence bands are wider relative to point-wise intervals, they are useful when, for example, one is interested in comparing short-run, medium-run, and long-run effects.\footnote{See section (ref) below for comparison with alternative simultaneous bounds based on the Bonferroni inequality. Recent contributions on simultaneous confidence intervals are jorda2009simultaneous, inoue2016joint, montiel2019simultaneous, inoue2022joint, arias2023uniform.}
Before we proceed to the estimation, it is worth noting that the hypothesis $\gamma_1=0$ is rejected in favor of the alternative $\gamma_1>0$ (Assumption (ref)) at 5% significance level. The corresponding test statistic is $t=1.77$, with a p-value equal to $0.039$ based on Algorithm 3. In other words, the variance of the Fed funds rate forecast errors is higher in the months with more FOMC meetings. This confirms earlier findings in the high-frequency identification literature that markets react to FOMC meeting announcements faust2004identifying.
As evident from Figure (ref), monetary policy tightening leads to a significant drop in output, which remains significant from 10 to 60 months using point-wise confidence bands. When considering joint significance, this decline is significant from approximately 15 to 35 months, a duration notably longer than previously reported results by inoue2016joint. The results on inflation are similar to those obtained using Wald-joint confidence intervals considered in that paper under a stationary VAR specification. In both cases, the price puzzle disappears once we consider simultaneous confidence intervals. It is important to keep in mind that our procedure for constructing simultaneous confidence intervals allows for unit roots without pre-testing, while inoue2016joint requires transforming the model into a stationary one, which may cause a pre-testing bias.
Our inference methodology allows for a flexible trend specification. Appendix Figures (ref)-(ref) compare the benchmark model results (blue lines) with more flexible models (golden lines): $(k=4,p=12)$, $(k=0,p=24)$, and $(k=4,p=24)$. The polynomial trend can account for slow and gradual changes in the mean values of stationary variables and the growth rates of unit root variables.
Including either a trend or additional lags (or both) results in shorter-lived responses of real GDP and weaker effects of the monetary shock on inflation and commodity prices. These results show that the long-run effects of monetary policy are sensitive to model specification. One possible explanation is that the trend (or additional lags) absorbs the long-run variation in real GDP that is otherwise attributed to monetary policy shocks. As a result, the more flexible specifications only detect the impact of monetary policy shocks that are neutral in the long run. This result is consistent with the panel regression findings at 10-year horizons in jorda2020long: monetary tightening shocks can be non-neutral in the long run, while monetary loosening shocks are typically short-lived.
In the current literature, it is conventional to exclude the time trend from monetary SVARs computed in levels uhlig2005effects,antolin2018narrative. To facilitate the comparison, we will report only the results for the benchmark specification $(k=0,p=12)$ in the subsequent sections.
Appendix Figure (ref) presents impulse response functions with smoothed point-wise and simultaneous bands for the benchmark specification. Compared to the raw LP estimates in Figure (ref), the smoothed IRF results in a stronger impact of the shock on GDP at longer horizons but a weaker impact on inflation at longer horizons. Notice that the confidence bounds become tighter as a result of the smoothing. This tightening occurs as a result of efficiency gains from imposing restrictions on the LP coefficients under the assumption of a correct specification of the VAR(12) model.
The smoothed LP estimators are based on a minimum distance estimator that imposes the VAR(12) model on the coefficients $C_{h}$. The distance between the VAR(12) model and the raw estimates was weighted with the inverse of the asymptotic variance of the individual LP estimators. Such weights are not fully optimal and could be further improved by using the inverse variance-covariance matrix of LP estimators, as in Theorem (ref). In practice, however, the optimal weights require some regularization in small samples. We leave this for future research.
There are alternative existing methods that allow simultaneous inference on LP estimates without knowledge of the joint covariance matrix using so-called union or Bonferroni bounds inoue2023significance. Such an approach assumes the worst-case dependence between the individual components of the IRF estimates. As discussed in montiel2019simultaneous, the simultaneous sup-t bounds will be tighter than the union bounds in most cases. Appendix Figure (ref) shows IRFs with sup-t versus union bands. Similar to montiel2019simultaneous, we find that Bonferroni confidence bands are noticeably wider relative to sup-t confidence bands.
Figure (ref) compares the IRFs to a monetary policy shock identified using heteroscedasticity-IV LP to IRFs from christiano1999monetary. Both identification schemes use the smoothed LP IRF. The only difference is the estimation of the impact vector. As can be seen, the results of the new identification approach are very close to those of christiano1999monetary, which is a standard Cholesky identification scheme with a priori zero impacts on real GDP, GDP deflator, and commodity price index. Specifically, the $\sup t$ test for the equality of the impact vectors for the two identification schemes given in Subsection (ref) has $\sup t= 1.2637$ with a p-value $0.53$ (critical value at 32% level is 1.8083). Without imposing the apriori zero restrictions, we obtain a similar conclusion.
Notice that the confidence bounds for the heteroscedasticity-IV estimators account for uncertainty in the structural rotation matrix (which depends on both $\hat\gamma$ and $\hat \Sigma_\eta$), while the Cholesky approach only accounts for uncertainty in $\hat \Sigma_\eta$. As a result, the Cholesky IRF has misleadingly tighter confidence bounds that are only correct under the a priori zero restrictions on real GDP, GDP deflator, and commodity prices implied by the identification scheme.
The classical paper christiano1999monetary uses recursive formulas for the computation of IRF. Appendix Figure (ref) compares the results of the smoothed LP Cholesky method with conventional VAR estimations under homoscedastic residuals. One can see that smoothed LP point estimates are similar to the conventional VAR estimates only at the short horizons. At the long horizons, the plots for real GDP, GDP deflator, and the commodity price index are substantially different. What is more surprising is that the smoothed LP point-wise confidence bounds are comparable to, and even tighter than the conventional VAR bootstrap bounds. This highlights the fact that the additional robustness of the proposed procedure does not automatically imply wider confidence bounds than its conventional alternative.
In the previous subsection, we found that the zero restrictions on impacts on real GDP, GDP deflator, and the commodity price index cannot be rejected simultaneously. However, these restrictions have been questioned in the literature. In particular, uhlig2005effects replaced these zero restrictions with sign restrictions on impulse responses at short horizons. Such an approach eliminates the price puzzle by construction. In a more recent study, antolin2018narrative further imposed additional restrictions on the signs of the monetary policy shock at particular dates to tighten the bounds on IRF. Both papers use Bayesian Gaussian homoscedastic SVAR models with uninformative priors over the structural rotation matrix for inference.\footnote{Alternatively, gafarov2018delta considered frequentist bounds on partially identified sets of IRF using moment inequality methods. This approach results in wider confidence bounds since it is designed to cover the worst-case values of the unidentified structural rotations as opposed to the prior distribution weighted average. }
Figure (ref) compares the impulse response functions to a monetary policy shock using a new identification approach \textemdash heteroscedasticity-IV smoothed LP \textemdash with the narrative sign restriction approach of antolin2018narrative. To closely match antolin2018narrative, we mimic the absence of a deterministic trend and use the same number of lags as the authors, that is, $p=12$.\footnote{The reason for not including a constant and a time trend goes back to uhlig2005effects and even further back to uhlig1994macroeconomists. As uhlig2005effects mentions, this leads to misspecification. However, making this choice is important for the Bayesian estimation.} The golden dashed lines represent a replication of the antolin2018narrative results: there is no price puzzle, as the sign restriction set to prevent it. There is a negative response of output to a monetary policy tightening shock. In response to 25 basis points (b.p) monetary policy tightening, real GDP drops by 0.1% after 2 years and to 0.12% after 5 years. According to our results, the interest rate stays higher for a longer period of time, which causes a more pronounced output drop: the real GDP drops by 0.25% after 2 years and by 0.12 % after 5 years. As one can see from Figure (ref), the difference in the medium run between the IRFs is statistically significant; confidence bands do not overlap. Moreover, we do not set the restriction on inflation, which results in a price puzzle, as found in earlier literature.
There is an interesting discussion in the literature where some researchers argue that the price puzzle may reflects the reality. For example, mishkin2007housing argues that an increase in the short-term interest rate raises the cost of building new houses and reduces housing activity. Through the supply cost channel this may cause an increase in prices in the short run. This example pertains to the construction sector. With monetary policy tightening, the cost of mortgages increases, which impacts production costs in the affected sectors.
This paper makes several contributions to the literature on SVARs with potential unit roots and cointegration. First, we propose a novel heteroscedasticity-baesd IV for monetary policy shocks that requires weaker assumptions than the conventional external IV approaches. As a result, we are able to construct much longer time series of instruments. Second, we show how one can perform point-wise and simultaneous inference on IRF with such heteroscedasticity-based external IV using the LP estimators. We specifically generalize the HAC estimator of LP account for the non-statioanry regressors and show equivalence of dependent wild bootstrap and the analytical HAC inference. Third, we use the joint draws of the LP estimates to construct minimum distance smoothed LP estimator and the corresponding confidence bounds.
We then illustrate the usefulness of these procedures in two applications: a simulated data based on a realistic DSGE model and a monetary SVAR for the U.S. We find that the heteroscedasticity IV based on the number of FOMC meetings in a given month results in IRF that are very close to the classical christiano1999monetary bounds and are different from those obtained using a comparable BVAR with narrative sign restrictions that rule out the price puzzle. In particular, our results suggest a much weaker long-run impact on inflation than the BVAR would imply.
\part*{Appendix}