Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,388 characters · 34 sections · 38 citation commands
Nonlinear Forecast Error Variance Decompositions with Hermite Polynomials
\setstretch{1}
The Forecast Error Variance Decomposition (FEVD) is a popular tool for modern empirical macroeconomic research. It allows a practitioner to assess the importance of shocks in a dynamic model by partitioning the contributions of its structural innovations on the conditional variance of its forecast. In the linear (Gaussian) Structural Vector Autoregressive (SVAR) framework, the construction of the FEVD is facilitated by the partial moving average representation. Indeed, the future trajectory of a stationary SVAR model can always be written as a sum of its independent structural innovations. Hence, the conditional variance of this sum is additively separable, and it partitions the effects of shocks by time horizon and by components of the structural innovations. \\
Although the assumption of linearity may be appealing in practice, it can oversimplify economic phenomena, leading to misguided conclusions for policymakers. Despite recent advancements in nonlinear modeling, extensions of the FEVD in these contexts remain largely unexplored. To date, only two methods have emerged in the current literature. The first method, proposed by LN2016, introduces the Generalized FEVD (GFEVD), based on the Generalized Impulse Response Function (GIRF) developed by KPP1996. The second method, introduced by IN2020, is a FEVD based on the Law of Total Variance. While these methods have steered the nonlinear FEVD literature into a promising direction, a clear consensus among practitioners is still lacking [KL2017, Section 18.2.2].\\
This paper fulfills the gap in the literature and proposes a new method for nonlinear FEVD analysis known as the Hermite FEVD (HFEVD) approach. It can be applied directly to parametric nonlinear SVAR models written on standard Gaussian innovations and to semiparametric nonlinear SVAR models with innovations that are standardized to Gaussian\footnote{While this paper works only with nonlinear SVAR models, the results can be extended to a more general class of nonlinear dynamic models.}. The strategy is to approximate the trajectory of these models using a Hermite polynomial expansion, resulting in a partial Volterra representation of the process at any future horizon. The orthogonality of Hermite polynomials under the standard normal density facilitates the additive separability of the conditional variance for this expansion, leading to the HFEVD. Moreover, its validity is supported by the fact that each term in the HFEVD is directly linked to the impact multipliers of the nonlinear Expected Impulse Response Function (EIRF)\footnote{The proposed methodology can also be generalized to models with non-Gaussian innovations by constructing orthogonal polynomials with respect to their densities, but the link to EIRFs is preserved only with Gaussian innovations and Hermite polynomials.}. This extends an analogous interpretation for the terms in the classical FEVD for the linear SVAR framework.\\
The proposed approach is unique in that it partitions the variance of the forecast along three dimensions: by time horizon, by the components of structural innovations, and by degree of nonlinearity. The latter dimension is particularly useful for practitioners, as it allows assessment of the nonlinearity in shock transmission channels within dynamic models. Moreover, because each term in the HFEVD is orthogonal to one another, they can be computed independently, allowing for a partial decomposition of contributions even when a trajectory is highly nonlinear. The HFEVD also accounts for the state of the economy, which potentially yields a different decomposition depending on the historical context it conditions upon. A practical exposition of these ideas are studied through the lens of the FRF2015 Threshold VAR framework. \\
The structure of this paper is as follows: Section 2 provides a review of the FEVD and it properties within a linear SVAR framework. Section 3 introduces the nonlinear SVAR model and its future trajectory. Section 4 proposes the Hermite Forecast Error Variance Decomposition and its properties. Section 5 provides details on numerical implementation of the HFEVD and Section 6 offers simulation results. Section 7 features the application of the HFEVD to the FRF2015 TVAR model and Section 8 concludes. Proofs, technical details, and the generalization of the method to non-Gaussian innovations are confined to the Appendix.
This section reviews the construction of the FEVD and its properties through the lens of the linear Gaussian Structural Vector Autoregressive model. The FEVD is readily available in closed form and is computed easily in this framework. The SVAR(1) model will be introduced alongside its moving average representation, detailing how the trajectory at horizon $h$ can be derived through recursive substitution. Subsequently, the discussion will address the FEVD and its identification challenges. Finally, the relationship between FEVD and Impulse Response Functions (IRFs) will be established.
Consider an $n$-dimensional linear Gaussian SVAR(1) process $(Y_t)$ given by:
where $A$ is an $n \times n$ matrix of autoregressive coefficients, $D$ is an invertible $n \times n$ structural weighting matrix, $(\varepsilon_t)$ are the structural innovations assumed to be independent multivariate standard normal and $\varepsilon_t$ is independent of $y_{t-1}$. When the eigenvalues of $A$ have a modulus strictly less than 1, the system (ref) has a unique stationary solution with an infinite moving average representation of the form:
where $\Theta_i = A^iD$ are the coefficient matrices of the moving average terms. Under the assumption of Gaussian innovations, only the matrix coefficients $A$ and $\Sigma= DD'$ are identifiable in this context, but not $D$ itself \footnote{More precisely, $D$ is identified up to an orthogonal transformation and the degree to which it is underidentified is $\frac{n(n-1)}{2}$. }. Hence, the structural innovations $\varepsilon_t$ and its components are not identifiable without imposing further assumptions on the matrix $D$\footnote{The most straightforward approach would be to order the components of $Y_t$ and assume the matrix $D$ to be lower triangular. However, this approach proposed by S1980 depends on the ordering and has no real economic interpretation in general. More complex identification methods are also available and are comprehensively reviewed in R2016.}. It is explicitly presumed that the practitioner has applied a sufficient and reasonable set of assumptions to resolve this identification issue such that the shocks on the structural innovations $\varepsilon_t$ have valid economic meaning.\\
The trajectory of the SVAR(1) process at horizon $h$ can be computed by means of recursion. Beginning at time $t+1$, we have:
Then, by recursion, at any horizon $h$, the formula becomes:
Thus, a partial moving average representation is obtained at horizon $h$, where the trajectory $Y_{t+h}$ is partitioned into two parts: [1] The observed history, $A^hY_t$, which contains the weighted path of all past innovations. [2] The weighted path of future innovations, $\sum_{i=0}^{h-1} \Theta_i\varepsilon_{t+h-i}$. By conditioning on $Y_t$, the trajectory of this process is stochastic only through the future innovations $\varepsilon_{t+i}$, for $i=1,...,h$. Hence, a variance decomposition in this context aims to separate the effects of these innovations on this stochastic component.
The FEVD is a partition of the (conditional) variance covariance matrix for the future trajectory. For instance, a decomposition by time (innovation) horizon is given by:
where $\Theta_i=A^{i-1}D$ is called the multiplier matrix. Each term in the decomposition is the contribution of the structural innovation at the $i$-th horizon, $\varepsilon_{t+i}$ \footnote{Another interpretation of equation (ref) is the term structure of total risk arising from the forecast over $h$ horizons. Each component on the right hand side corresponds to the updating of information sets between time $t+i-1$ and $t+i$, and measures the risk along the “maturity horizon". More generally, for any square integrable process $(Y_t)$ with information set $(I_t)$, the decomposition by maturity horizon is given by: $\mathbb{V}[Y_{t+h}|I_t]=\sum_{i=1}^{h} \mathbb{V}\{[\mathbb{E}(Y_{t+h}|I_{t+i})-\mathbb{E}(Y_{t+h}|I_{t+i-1})]|I_t\}$ [see GL2024, Lemma 2]. In the linear framework, the notions of decomposition by “innovation" horizon and “maturity" horizon are the same, but this will not be the case in the nonlinear framework. In this particular, this paper is focused on the innovation horizon which is of interest to macroeconomists.}. Notably, the horizonal decomposition is a function of $A$ and $DD'$, so it is always identifiable. The decomposition in (ref) is also independent of the conditioning value $Y_t$. This is known as the history invariance property of the FEVD in the linear dynamic model; regardless of what we have seen in the most recent time period, the decomposition of risk remains the same. Clearly, this is not a very realistic assumption in practice\footnote{For example, we would expect the decomposition of financial or macroeconomic risk to be different when the interest rates are low compared to when the interest rates are high.}. \\
Another decomposition of interest is the contribution to the variance-covariance matrix of $Y_{t+h}$ accounted for by the individual components of the structural innovations. Let us denote the $j$-th column of the multiplier matrix as $(\Theta_{i})_{*,j}=(A^{i}D)_{*,j}$ (an $n \times 1$ vector). Then we have:
The product term $(\Theta_{i})_{*,j}(\Theta_{i})_{*,j}'$ captures the effect of the $j$-th structural innovation at horizon $t+i$ on the trajectory of the process. For instance, the $p,q$-th entry in $(\Theta_{i})_{*,j}(\Theta_{i})_{*,j}'$ corresponds to the contribution of $\varepsilon_{j,t+i}$ on the covariance between $Y_{p,t+i}$ and $Y_{q,t+i}$. Hence, the FEVD in (ref) now provides a decomposition by time horizon and by structural innovation component. However $(\Theta_{i})_{*,j}$ depends on the parameter $D$ itself, so the decomposition by components is not identifiable without further restrictions on $D$. \\
In general, a practitioner is more interested in the scalar decomposition of some component of $Y_{t+h}$ or their linear combinations. For instance, the horizonal decomposition in (ref) for a linear combination $\alpha' Y_{t+h}$ can be written as:
or a decomposition by components in (ref) as:
In the special case where $\alpha$ is a vector equal to 1 in the $\ell$-th entry and $0$, otherwise, the FEVD is a decomposition of the conditional variance for component $Y_{\ell,t+h}$:
An important property of the FEVD in the linear framework is its relationship to Impulse Response Functions. Intuitively, it evaluates the consequence of a transitory shock to the structural innovation at date $t+1$ by comparing two trajectories of $Y_{t+h}$: [1] A perturbed path, where a shock magnitude of $\delta$ is applied to structural innovation $\varepsilon_{t+1}$. [2] A baseline path, where the process follows its natural trajectory defined in (ref). To formalize this idea, suppose the practitioner is interested in a shock of magnitude $\delta$ at time period $t+1$. The perturbed and baseline paths at horizon $h$ are given by:
To evaluate the consequence of this conceptual experiment, the difference between these two paths is considered:
Equation (ref) defines the Impulse Response Function (IRF), which evaluates the effect of the shock through $\Theta_{h-1}=A^{h-1}D=\frac{\partial Y_{t+h}}{\partial \varepsilon_{t+1}}$. This is called the impact multiplier, which captures “consequences of a one-unit increase in the [structural] innovation" on the trajectory [H1994, Definition 11.4.2]. Like the FEVD, the IRF is history invariant in the linear framework, since it is independent of the most recent value $Y_t$. There is a drift effect at each horizon, which is linear in the magnitude vector $\delta$. We also note that the IRF does not depend on the path of future innovations, and coincides with the (conditional) expected IRF, that is:
It follows from equation (ref) that the FEVD can be written as:
where $e_j$ is a vector of 0's except for the $j$-th entry which is equal to 1. Intuitively, the FEVD is a sum of squared IRFs following shocks of unit magnitude, where the term $IRF(i,e_j)$ corresponds to the effect of a unit shock on the element $\varepsilon_{j,t+1}$ at horizon $i$.
This section introduces the Nonlinear SVAR process and explains how the future trajectory can be imputed by means of recursive substitution. Some examples of nonlinear data generating processes that fall under this framework are also provided.
First, a general class of nonlinear models, broadly referred to as the Nonlinear SVAR process, can be introduced through the following proposition:
Proof: This is a direct consequence of the triangular representation of joint continuous distributions [R1952]. \qed \\
The nonlinear structural Gaussian innovations $(\varepsilon_t)$ are given by:
where $g^{-1}(\cdot,\underline{Y_{t-1}})$ denotes the inverse of the function $g(\underline{Y_{t-1}},\cdot)$. These innovations are nonlinear functions of $Y_t$ given its past, and are the basis for structural shocks in the construction of IRFs and FEVDs which will be defined below. There are a number of nonlinear models in the literature which fall under this paradigm\footnote{For a comprehensive review of such models, see KL2017 Chapter 19.}, including but not limited to Threshold and Smooth Transitions [e.g. HT2013, GM2014, FRF2015], Time-Varying Parameter [e.g. P2005, N2011] and Nonparametric SVARs [e.g.HTY1998, J2013]. \\
In general, the function $g$ and its structural innovations $(\varepsilon_t)$ are not identifiable unless the dimension of $(Y_t)$ is equal to 1. This problem is analogous to the identification issue of the matrix $D$ in the linear Gaussian SVAR framework (ref). Depending on the nature of nonlinearity selected and the information available to the practitioner, this identification issue can be resolved in different ways\footnote{For example, for nonlinear SVAR models with additive errors, one may consider Independent Component Analysis (ICA) as a solution to the identification issue, if at most one component of the structural innovations is Gaussian [see GMR2017 for a discussion and review on ICA in SVAR models]. In such a context, the non-Gaussian innovations will have to be re-normalized inT order to satisfy the representation (ref) as discussed below.}. However, to avoid detracting from the main ideas in this paper, it is explicitly assumed that the practitioner has selected an appropriate identification scheme with discretion, so that a shock on the structural Gaussian innovations $(\varepsilon_t)$ is indeed compatible with plausible economic interpretations. \\
Remark 1: The nonlinear representation in (ref) is not unique to Gaussian innovations. More specifically, a Markov process $(Y_t)$ may also be written as:
where $g^*$ is invertible in $u_t=(u_{1,t},...,u_{n,t})$, $u_t$ has independent components and each component has a conditional c.d.f. $u_{j,t}\sim F_j$ for $j=1,...,n$., which are not necessarily Gaussian. Proposition 1 can still be used by changing the sign of $u_t$ (to transform a decreasing function to an increasing one) and transforming the components of $u_t$ by $\Phi^{-1}F_j(\cdot)$ for $j=1,...,n$, where $\Phi$ denotes the standard normal c.d.f. (to transform $F_j$ into the Gaussian). The choice of a representation with Gaussian innovations is simply a normalization\footnote{While the main text will provide methodology and the key results for models with Gaussian innovations, Appendix D will extend the ideas for representations with non-Gaussian errors.}, in line with the convention set by the literature [see L2019, Section 6.3 for a discussion of normalization in identification], and to facilitate a link between IRFs and FEVDs in the nonlinear framework [see the discussion in Section 4.4].
Beginning at time $t$, we may recursively define the trajectory of the process at future horizons using recursive substitution. Let us denote the relevant history at time $t$ as $\underline{Y}_{t} = (Y_{t},...,Y_{t-p+1})$. Based on (ref), the process $(Y_{t})$ is such that:
Applying the recursive formula at $h=2$, we obtain:
In general, for a horizon $h$, this substitution procedure yields:
where $\varepsilon_{t+1:t+h}=(\varepsilon_{t+1},...,\varepsilon_{t+h})'$ denotes the future path of structural innovations between $t+1$ and $t+h$, and $g^{(h)}$ represents the iteration of function $g$ a total of $h$ times. Thus the value of the process at time $t+h$ is a function of the relevant history $\underline{Y}_{t}$ and the sequence of all future structural innovations between $t+1$ and $t+h$. For relatively simple functions $g$, it is possible to obtain $Y_{t+h}$ in closed form. However, this may be challenging when the process (and the function $g$) is highly nonlinear. In the latter case, this trajectory can be obtained by means of simulation (see Section 5 for a discussion).
The univariate Double Autoregressive model of order one [L2007], denoted DAR(1), is used to model conditional heteroscedasticity. It is given by the representation:
where $\varepsilon_t$ is iid N(0,1), $\alpha >0$, and $\beta \geq 0$\footnote{Moreover, the model admits a strictly stationary solution if $\mathbb{E}\left[\log|\phi + \sqrt{\beta}\varepsilon_t|\right]<0$.}. Therefore, function $g$ is nonlinear in $y_{t-1}$, linear in $\varepsilon_t$, with nonlinear cross effects between $y_{t-1}$ and $\varepsilon_{t}$. Its Gaussian innovation is given by the expression:
Through recursive substitution, we get the following representation for $h=2$:
Even at horizon $h=2$, the expression (ref) is quite complicated, which features interaction terms between innovations at different horizons. By continuing the substitution process, it would be difficult to obtain a general formula for $y_{t+h}$, since there would be a compounding of terms from the previous horizons. \\
Remark 2: It is also possible to consider a semi-parametric version of Example 1 of the form:
where the innovation process $u_t$ is a strong white noise process, not necessarily Gaussian. Let $F$ denote its corresponding cumulative distribution function of $u_t$ and $\Phi$ denote the standard normal cumulative distribution function. Then, the model (ref) can be written as:
where the nonlinear autoregression (ref) is now written as a function of the Gaussian innovations $\varepsilon_t \sim N(0,1)$. In this extension, the function $g$ is also nonlinear in $\varepsilon_t$. The model is semi-parametric with scalar parameters $\phi$, $\alpha$, $\beta$ and functional parameter $F$.
Consider a discrete time stochastic volatility model where the daily return $(y_t)$ and the within-day logged realized volatility $(z_t)$, i.e. the logged variance of 5 minute returns in day $t$, evolve according to the model:
where $\varepsilon_t=(\varepsilon_{1,t},\varepsilon_{2,t})'$ is a bivariate standard Gaussian innovation. There are two nonlinearities introduced in the autoregressive representation: [1] The endogenous threshold transition between the regimes $y_{t-1} \geq 0$ and $y_{t-1}<0$. [2] The interaction between $\exp(z_{t-1})$ and $\varepsilon_{1,t}$. For exposition, let us compute the trajectory of this process at horizon $h=1$:
The conditioning value (history) $y_t$ plays an important role on this trajectory - in particular, it influences the regime of $y_{t+1}$. Similarly, by recursive substitution, we obtain for horizon $h=2$:
$y_{t+2}$ is function $y_t$, $z_t$ and the future innovations $\varepsilon_{1,t+1}$, $\varepsilon_{1,t+2}$ and $\varepsilon_{2,t+1}$. There are four regimes, which depend on the values of $y_t$ and $y_{t+1}$. In general, $2^h$ such regimes exist at horizon $h$, and the trajectory of a threshold process can be viewed as a “tree" of nodes that depend on the evolution of the threshold variable throughout the time horizons. As such, when analyzing shocks in a threshold framework, it is important to take into account the fact that there can be a switching of regimes between horizons\footnote{In particular, by conditioning on $y_t$, you may govern in which regime the process starts, but you cannot control whether or not it stays in that regime, since it depends on the future innovations $\varepsilon_{t+i}=(\varepsilon_{1,t+i},\varepsilon_{2,t+i})$ for $i=1,...,h$.}. \\
Remark 3: A smoothed version of (ref) is given by:
where a transition, characterized by a first-order logistic function, governs the evolution of the risk premium coefficients.
In the linear framework, recursive substitution yields directly a partial moving average representation (ref). However, when nonlinearities are present, this is no longer the case, and there is no reason for the expression obtained in (ref) to have a “nice" closed form expression that can facilitate the construction of a valid FEVD. The approach in this paper, referred to hereafter as the Hermite Forecast Error Variance Decomposition (HFEVD), utilizes a Hermite polynomial expansion to approximate the future trajectory\footnote{Recall from Remark 1 that the nonlinear SVAR can be written on innovations that are non-Gaussian. Indeed, the ideas behind the HFEVD can be generalized to other sequences of orthogonal polynomials. However, not all distributions are associated with a natural set of orthogonal polynomials. Appendix D.1 will discuss how orthogonal polynomials for an arbitrary distribution may be constructed, and Appendix D.2 will extend the HFEVD approach for non-Gaussian errors.}. The resulting nonlinear representation is a sum of uncorrelated terms, which allows us to separate the variance along three dimensions: by the time horizon, by the components of the structural innovations and by the degree of nonlinearity.
A sequence of polynomials of degree $k$, denoted $f_k(\varepsilon)$ for $\varepsilon \in \mathbb{R}$, are orthogonal on the interval $(-\infty,\infty)$, with respect to the weight function $w(\varepsilon)$ if they satisfy:
where $k\neq m$ and $k,m \in \mathbb{N}$ [AS1964]. A special case of orthogonal polynomials are the Hermite polynomials, where the weight function $w(x)=\frac{1}{\sqrt{2\pi}}\exp[-x^2/2]$ coincides with the standard normal density. In the univariate case, the $k$-th Hermite polynomial is given by:
for $k \in \mathbb{N}$. Beginning with the initial value $H_0(\varepsilon)=1$, it follows that the sequence of Hermite polynomials can be defined by the recursive relation:
which is a consequence of applying the chain rule directly on the definition (ref). For example, the first six Hermite polynomials are:
This paper works exclusively with Hermite polynomials of multivariate normal random variables. Suppose $\varepsilon=(\varepsilon_1,...,\varepsilon_h)'$ such that $\varepsilon \sim MVN(0,Id)$. Then, the joint Hermite polynomial of $\varepsilon$ is equal to the product of the univariate Hermite polynomials for each component in $\varepsilon$:
where $k_i$ represents the degree of the marginal Hermite polynomials for $\varepsilon_i$, $1 \leq i \leq h$, and $K=(k_1,...,k_h)'$ is a vector in $\mathbb{N}^h$. These polynomials have the following first and second-order moments:
Proof: Property (1) follows by definition of orthogonality under the standard normal density. Property (2) follows directly from Property (1) since $\mathbb{E}\left[H_K(\varepsilon)\right]=\mathbb{E}\left[H_K(\varepsilon)H_0(\varepsilon)\right]=0$. Lastly, Property (3) is a special case of R2017, Proposition 8, for independent Gaussian random variables. \qed
The joint Hermite polynomials form an orthogonal basis of $L^2(\mathbb{R}^h,\mathcal{B},MVN(0,Id))$ ($\mathcal{B}$ denotes the Borel set of $\mathbb{R}^h$), that is, the Hilbert space of square integrable functions with respect to the standard multivariate normal distribution. Hence, any measurable function $g$ of standard Gaussian random variables $\varepsilon=(\varepsilon_1,...,\varepsilon_h)'$ in this space can be approximated by a Hermite polynomial expansion. This is formalized in the following lemma:
with $c_\textbf{0} = \mathbb{E}\left[g(\varepsilon)\right]$, $c_K = \mathbb{E}\left[g(\varepsilon)H_K(\varepsilon)\right]/\sqrt{\mathbb{V}(H_K(\varepsilon))}$, $\forall K\neq\textbf{0}$.\\
Proof: This is a direct application of R2017, Corollary 15. \qed \\
Each term in (ref) consists of the joint Hermite polynomial $H_K(\varepsilon)$, where $K=(k_1,...,k_h)$ captures the degrees of nonlinearity for the marginal Hermite polynomials of the product, and $c_K$ is the associated coefficient in the Hermite polynomial expansion. The sum is taken across all possibilities in $\mathbb{N}^h$. As an example, consider the first six terms of the expansion for a bivariate standard Gaussian vector $\varepsilon=(\varepsilon_1,\varepsilon_2)'$. We get:
The first term corresponds to the initial polynomial $H_{(0,0)}=1$ of degree 0 with coefficient equal to $c_{(0,0)}=\mathbb{E}[g(\varepsilon)]$. The second and third terms correspond to the univariate Hermite polynomials of degree 1 for $\varepsilon_1$ and $\varepsilon_2$, respectively. The fourth and fifth terms are the univariate Hermite polynomials of degree 2 for $\varepsilon_1$ and $\varepsilon_2$, respectively. The sixth term is a joint Hermite polynomial which is a product of the two univariate Hermite polynomials of degree 1 for both $\varepsilon_1$ and $\varepsilon_2$ (and so on).
As discussed in Sections 3.2 and 3.3, the trajectory of a nonlinear process can be highly complex and may not be available in closed form. As such, it is difficult to obtain a valid decomposition of the conditional variance on (ref) directly. Instead, we may use Lemma 2 to approximate the trajectory using Hermite polynomial expansions to obtain what is known as the partial Volterra representation for the process\footnote{Our approach follows closely to the representation theorem discussed in Section 2.2 of Gourieroux and Jasiak (2005). The objective of their paper is to define nonlinear Impulse Response Functions in the univariate nonlinear framework. In contrast, our goal is to facilitate the construction of a nonlinear FEVD.}; it is the nonlinear analogue to the partial moving average representation in (ref) for the linear framework [see P1978, Section 4]. The conditional variance on this expansion can be separated additively due to the orthogonality of Hermite polynomials which facilitates the construction of the FEVD. To formalize this idea, let us consider the univariate nonlinear AR(p) case. Its Hermite Forecast Error Variance Decomposition (HFEVD) is given by the following proposition:
Proof: From (ref), the trajectory of a nonlinear AR(p) process is given by:
where $g^{(h)}$ is square integrable function of the $h$ Gaussian innovations $\varepsilon_{t+1:t+h}=(\varepsilon_{t+1},...,\varepsilon_{t+h})$ at future horizons. Thus, it satisfies the assumptions of Lemma 2, and we get a Hermite polynomial expansion of $y_{t+h}$ as:
with $c^{(h)}_\textbf{0}(\underline{y_t}) = \mathbb{E}\left[g^{(h)}(\underline{y}_t,\varepsilon_{t+1:t+h})\right]$ and $c^{(h)}_K(\underline{y_t}) = \mathbb{E}\left[g^{(h)}(\underline{y}_t,\varepsilon_{t+1:t+h})H_K(\varepsilon_{t+1:t+h})\right]/\sqrt{\mathbb{V}(H_K(\varepsilon_{t+1:t+h}))}$, $\forall K\neq\textbf{0}$. The orthogonality property of Hermite polynomials (Lemma 1(1)) for Gaussian random variables means that the variance of their sum is equal to the sum of their variances. Hence, taking the variance of (ref) conditional on $\underline{y_t}$ we get:
where:
Since $H_0(\varepsilon_{t+1:t+h})=1$ is a constant term, it has no effect on the conditional variance and the summation in (ref) is taken over $\mathcal{K} = \left\{\mathbb{N}^4 \ \backslash \ \textbf{0}\right\}$ instead. \qed\\
In the univariate dynamic framework, the HFEVD is a decomposition by time horizon and by degree of nonlinearity. Note that the coefficients $c^{(h)}_{K}(\underline{y_t})$ contain the superscript ${(h)}$ to indicate that they depend on forecast horizon $h$. They are also functions of the observed history $\underline{y_t}$; hence, the HFEVD is in general, not history invariant. For a better understanding of formula (ref), let us consider an autoregressive process at horizon 2: $Y_{t+2} = g^{(2)}(\underline{y_t},\varepsilon_{t+1},\varepsilon_{t+2})$. Then:
Since the sum is taken over $\mathcal{K} = \left\{\mathbb{N}^h \ \backslash \ \textbf{0}\right\}$, the aggregation in (ref) is potentially infinite, with the total number of terms in the expression depending solely on the degree of nonlinearity in $g^{(h)}$. A highly nonlinear trajectory will likely have many terms, while a more linear one could have most of the terms equal to 0\footnote{While this may seem discouraging for practitioners to implement, it is important to note that every term is orthogonal to one another and can be computed independently. In particular, the approximation accuracy of the entire trajectory does not influence the accuracy of individual terms within the approximation. Thus, a partial decomposition can still be obtained up to a finite number of terms without compromising the accuracy of the results.}.
The extension to the multivariate case follows naturally from the discussion above, but there are two key differences. First, the process $(Y_{t+h})$ is now of dimension $n$, and we now have a decomposition of the conditional variance-covariance matrix instead of a scalar conditional variance. Second, there are now $n\times h$ structural innovations across $h$ horizons. The proposition below extends the HFEVD to the nonlinear SVAR($p$) framework:
Proof: The proof is identical to Proposition 2, but now written in matrix form since we have a vector of coefficients. \qed \\
The vector $K=(k_{1,1},...,k_{n,h})$, which captures the degrees of nonlinearity for each marginal Hermite polynomial in the product, is now of dimension $n \times h$. The first $h$ elements correspond to the degrees of nonlinearity for $\varepsilon_{1,t+1:t+h}$, the next $h$ elements correspond to the degrees of nonlinearity for $\varepsilon_{2,t+1:t+h}$, ..., and the last $h$ elements correspond to the degrees of nonlinearity for $\varepsilon_{n,t+1:t+h}$. Moreover, the HFEVD is now a decomposition of the conditional variance-covariance matrix along three dimensions: by time horizon, by structural innovation components and by degree of nonlinearity. To illustrate this point, let us consider a bivariate autoregressive process at horizon $h=2$. Then:
In practice, appropriate terms in the Hermite polynomial expansion can be selected in the decomposition to obtain an interpretation that is of interest to the practitioner. For exposition in this section, consider the bivariate nonlinear autoregressive process $Y_{t}=(Y_{1,t},Y_{2,t})'$. At horizon 2, its trajectory is given by:
By applying (ref), the conditional variance of (ref) is of the form:
where $\mathbb{VC}^{(2)}_K(y_t)=\left[C_K^{(2)}(\underline{y_t})C_K^{(2)}(\underline{y_t})'(k_{11}!k_{12}!k_{21}!k_{22}!)^2\right]$ denotes the variance contribution of the joint Hermite polynomial indexed by $K=(k_{11},k_{12},k_{21},k_{22})$, $K\neq \textbf{0}$, and $k_{11}$, $k_{12}$, $k_{21}$ and $k_{22}$ represent the degrees of nonlinearity for the marginal Hermite polynomials of the innovations $\varepsilon_{1,t+1}$, $\varepsilon_{1,t+2}$, $\varepsilon_{2,t+1}$ and $\varepsilon_{2,t+2}$, respectively.
Each term in (ref) can be classified as a contribution to the conditional variance from either specific innovations or interactions between multiple innovations. For instance, the $k_{11}$-th degree variance contribution specific to innovation $\varepsilon_{1,t+1}$ is given by:
When $k_{11}=1$, (ref) captures the linear contribution specific to innovation $\varepsilon_{1,t+1}$. Likewise, $k_{11}\geq 2$ captures specific higher-order contributions of $\varepsilon_{1,t+1}$, such as quadratic (for $k_{11}=2$) or cubic (for $k_{11}=3$). Contributions can also be defined jointly for multiple innovations and arise due to interactions between them. For instance, the joint contribution between $\varepsilon_{1,t+1}$ and $\varepsilon_{2,t+1}$ at degrees of nonlinearity $k_{11}$ and $k_{21}$, respectively, is given by:
for $k_{11},k_{21} \geq 1$. These contributions can also be aggregated in different ways.
The set $\mathcal{K}=\mathbb{N}^4 \ \backslash \ \left\{\textbf{0}\right\}$, which contains all possible indices in the expansion for (ref), can be partitioned in different ways to obtain interpretations of the HFEVD that are of interest to the practitioner. Moreover, since each term in the decomposition are orthogonal to one another, these partitions can be easily computed even if the trajectory has many high-order nonlinear terms. For instance, suppose $\mathcal{K} = \mathcal{K}_1 \bigoplus \mathcal{K}_2$ can be partitioned into two sets and consider the following partitions.
In Section 2.3, a relationship between IRFs and FEVDs was established. An analogous relationship exists between nonlinear EIRFs and the HFEVD. Akin to the linear framework, we can define a perturbed path $Y_{t+h}(\delta)$ and a baseline path $Y_{t+h}$ as:
The uncertainty of the nonlinear Impulse Response Function (IRF) can be summarized by the joint distribution of these two paths [GJ2005, Definition 4]. In the current literature, it is standard practice to consider the expected difference of these paths, conditional on a history $\underline{y_{t}}$:
For exposition, let us consider the EIRF of a univariate process at horizon 1. Then $Y_{t+1}=g(\underline{y_t},\varepsilon_{t+1})$ will be a function of a single Gaussian innovation $\varepsilon_{t+1}$. A Taylor expansion of the EIRF yields the following:
Proof: See Appendix A.1. \\
Lemma 3 provides an intepretation of the EIRF as implied by the perturbed and baseline paths in (ref). Each term in the summation is the product of the $k$-th order Gaussian marginal effect of $\varepsilon_{t+1}$ on $Y_{t+h}$ and a scaling of the shock magnitude. In the linear SVAR model discussed in (ref), $EIRF(1,\delta,\underline{y_t})=\delta AD$, and the marginal effect of $\varepsilon_{t+1}$ on $Y_{t+h}$ is the impact multiplier $\Theta_i=AD=\frac{dY_{t+h}}{d\varepsilon_{t+1}}$. Hence, we can interpret (ref) as a weighted sum of impact multipliers at all degrees of nonlinearity $k$. Morever, these marginal effects are linked directly to the terms in the HFEVD in the following way\footnote{Two alternative defintions of the IRF have also been considered in the past: [1] The “MIT" shock where a shock is characterized by $\varepsilon_{t+1}+\delta$ as in this paper, but all other innovations are set directly equal to 0. [2] The Generalized Impulse Response Function (GIRF), where a shock is characterized by setting the innovation directly equal to the magnitude of the shock, i.e. $\varepsilon_{t+1}=\delta$ [KPP1996]. Their approaches do not share the same interpretation as Lemma 3 and are not appropriate characterizations of shocks in nonlinear dynamic models [see Appendix B for details].}:
Proof: See Appendix A.2. \\
By analogy to the relationship in the linear framework, Proposition 4 shows that the terms in the HFEVD are equivalent to the squared impact multipliers of the IRF expansion in (ref). This result can be extended to multivariate processes as well. Without loss of generality, consider the case of a bivariate process at horizon 1. Then $Y_{t+1}= g(\underline{y_t},\varepsilon_{1,t+1},\varepsilon_{2,t+1})$ is a function of Gaussian innovations $\varepsilon_{1,t+1}$ and $\varepsilon_{2,t+2}$ and we get the following proposition:
Proof: See Appendix A.3. \\
The terms in the EIRF expansion now feature cross-derivatives between $\varepsilon_{1,t+1}$ and $\varepsilon_{2,t+1}$. They capture the interaction effects between shocks, and are linked to the joint Hermite polynomials between the structural innovations. Hence, we can interpret the contribution of marginal Hermite polynomials as the expected marginal effects of shocks, and the contribution of joint Hermite polynomials as their interaction effects. These results can be generalized to any dimension $n$ and horizon $h$, where the expansion of the EIRF will contain cross-derivatives of the $n\times h$ innovations. \\
Remark 4: As shown in Appendix D.4, the link between the nonlinear FEVD and EIRF only holds for Gaussian innovations. Hence, Gaussianity is a very special case of economic innovations, and non-Gaussianity is in fact a source of nonlinearity that differs from the nonlinearity introduced by the dynamics of a model (e.g. An AR(1) process with t(5) errors is linear in the modelling dynamic, but exhibits nonlinear behaviour due to the non-Gaussianity of the innovation). Moreover, this observation also tells us that any nonlinear FEVD constructed based on the definition of the nonlinear IRF, such as the GFEVD approach of LN2016, can have misleading interpretations when errors are non-Gaussian.
Consider an explicit calculation of the HFEVD for the nonlinear process given by:
where $\varepsilon_t$ is a bivariate standard Gaussian noise. At horizon $h$ the value of this process is:
Suppose we are interested in the decomposition at horizon $t+3$. Then:
There are 3 (horizons) $\times$ 2 (variables) innovations in the system. For $Y_{1,t+3}$, there will be at most quadratic effects of $\varepsilon_{1,t+l}$ and linear effects of $\varepsilon_{2,t+l}$. For $y_{2,t+3}$, it is a function of only linear terms of $\varepsilon_{2,t+l}$. Thus, it suffices to consider Hermite polynomials of degree 1 and 2 for this decomposition that only include the marginal and pairwise contributions in the total variance. The resulting expansion is given by:
which is equivalent to the partial nonlinear moving representation (ref). This expression separates the effects of shocks on $Y_{t+3}$ by time horizon, by structural innovation component and by degree of nonlinearity. For example, the term $(\varepsilon^2_{t+3}-1)$ captures the quadratic effect of innovation $\varepsilon_1$ at horizon $3$. Its influence on $y_{1,t+3}$ is $b^2y_{2,t}$, but it has no effect on $y_{2,t+3}$. We also see interaction effects between innovations, for instance, in the term $(\varepsilon^2_{1,t+3}-1)\varepsilon_{2,t+2}$. Since the expansion above is now a sum of orthogonal Hermite polynomials, the conditional variance-covariance matrix of $Y_{t+3}$ can be additively separated as the sum of the individual variance-covariance matrices corresponding to the difference terms. It can be shown that:
The diagonal elements of each matrix correspond to the contribution of each innovation to the variance of $y_{1,t+3}$ and $y_{2,t+3}$. For example, the linear effects of all $\varepsilon_2$ innovations on $y_{1,t+3}$ is given by $(a+b)^2+1$ (the fourth and fifth term of the FEVD). On the other hand, there are no linear effects from the $\varepsilon_1$ innovations. The non-diagonal elements represent the contribution of each innovation to the covariance between $y_{1,t+3}$ and $y_{2,t+3}$. For example, any term with $\varepsilon_1$ (whether joint or marginal) has no effect, since $\varepsilon_1$ only appears in the equation for $y_{1,t+3}$ but not $y_{2,t+2}$. We also note that the covariance contributions can be negative and depends on the parameters $(a,b)$. If both $a$ and $b$ are both negative, then the linear effects of $\varepsilon_2$ on the covariance is also negative, since $b^2(a+b)+b <0$.\\
Remark 5: IN2020 argue that $\mathbb{E}_t[\mathbb{V}(y_{1,t+h}|\varepsilon^{-1}_{t+1:t+3})]=\mathbb{E}_t[\mathbb{V}(y_{1,t+h}|\varepsilon_{2,t+1},\varepsilon_{2,t+2},\varepsilon_{2,t+3})]$ captures not only the marginal contribution of innovation 1, but the contribution of its interactions. Indeed:
These terms coincide with the marginal and joint Hermite polynomials of $\varepsilon_{1,t+k}$ for $k=1,2,3$. Hence, the Isakin and Ngo (2020) approach is a special case of the HFEVD.
This section focuses on how the HFEVD can be implemented in practice. There are two issues that may arise with the proposed approach. First, the future trajectory in (ref) may not be available in closed form. Second, the model may be semi-parametric in the sense that the i.i.d. innovations are not necessarily Gaussian. The parametric framework is discussed first followed by the semi-parametric framework. In both cases, it is assumed that the parameter $\theta$ is identifiable.
Consider the $n$ dimensional nonlinear SVAR(p) process described in Proposition 1:
where $\varepsilon_t$ is assumed to be standard Gaussian and $\theta$ denotes a parameter which governs the function $g$. To evaluate some of the contributions to the conditional total variance, the approach follows the steps below:
Consider the $n$ dimensional SVAR(p) process described in (ref):
We assume that this model is well-specified with a true value $\theta_0$ of the parameter, and that $u_t = u_t(\theta_0)$ is i.i.d. with a distribution $F_0$ which is unknown to the practitioner and not necessarily Gaussian.
In this section, we provide the HFEVD analysis of simulated data following the examples discussed in Section 3.3.
Consider the Double Autoregressive model in Section 3.3.1 with parameter values $\alpha =1$, $\phi = 0.5$ and $\beta=0.5$, which ensures a strictly stationary solution with finite second-order moments \footnote{Ling (2007) describes the stationarity region as the set $\{(\phi,a) \ s.t. \ \mathbb{E}|\ln\phi+\sqrt{a}\varepsilon_t|<0$\}.}. This yields the data generating process:
For the structural innovations we consider the cases $u_t \sim N(0,1)$ and $u_t \sim t(5)$. Then, the algorithms for the parameteric model (Section 5.1) and semi-parametric model (Section 5.2) can be applied to these two cases respectively. First, simulations of 200 observations for each process with initial value $y_0 = 0$ are presented in the two graphs of Figure 1. As expected, the variation for the DAR(1) model with Gaussian innovations exhibit much lower variability than the DAR(1) model with t(5) innovations.
Figure 2 examines the linear contributions of innovations on the conditional total variance of the DAR(1) at the conditioning value $Y_T=0$. The blue and red bars represent the results for processes with N(0,1) and (normalized) t(5) structural innovations, respectively. As the forecast horizon increases, the linear contributions decrease more rapidly for the processes with t(5) innovations. This difference arises from the presence of fatter tails in the t-distribution, which introduces a nonlinear effect even though the underlying dynamics of the processes remain the same. Consequently, if the data were incorrectly modeled using a linear AR(1) process, it could lead to significant misinterpretations of the FEVD and potentially erroneous conclusions.
Next, an HFEVD analysis of both processes at horizon 2 with conditioning value $Y_T=0$ is shown in Figure 3, which provides a granular separation of contributions by the time horizon and by the degree of nonlinearity. The processes are approximated by a sum of four Hermite polynomial functions of structural innovations:
that capture the linear effects of $u_{t+1}$, the linear effects of $u_{t+2}$, the interaction of linear effects between $u_{t+1}$ and $u_{t+2}$, and the interaction of the quadratic effect of $u_{t+1}$ and the linear effect of $u_{t+2}$. For the process with N(0,1) innovations (left pie chart of Figure 3), this approximation sufficiently captures all the variation in $y_{t+2}$, as the percentage contribution of the four terms sum up to 100%. The majority of the variation is explained by the marginal linear contributions of $u_{t+1}$ and $u_{t+2}$, accounting for 97% (15% + 82%) of the conditional variance in $y_{t+2}$. However, 3% of the remaining variation is due to the interaction term between the quadratic effect of $u_{t+1}$ and the linear effect of $u_{t+2}$. It is worth noting that there are no interacting linear effects between $u_{t+1}$ and $u_{t+2}$. For the model with t(5) innovations (right pie chart of Figure 3), the partial expansion only explains 97% (77%+12%+8%) of the variation in $y_{t+2}$ conditional on $y_T=0$. The remaining contributions are due to higher order effects, which is to be expected due to the non-Gaussianity of the structural innovations. The interaction term between the quadratic effect of $u_{t+1}$ and the linear effect of $u_{t+2}$ also plays a larger role in explaining the total variation.
The importance of the conditioning is shown in Figure 4, where the HFEVD is applied to both processes conditional on $Y_T=0$, $Y_T=5$ and $Y_T=-5$, respectively. Both graphs show that the decomposition can be sensitive to conditioning. In particular, at the mean value of the process $Y_T=0$, the linear contribution diminishes slowly as the horizon increases, while at more extreme values such as $Y_T=5$, or $Y_T=-5$, the decrease is much steeper. However, note that the HFEVD look similar for $Y_T=5$, or $Y_T=-5$. Hence, the decomposition analysis demonstrates that the DAR(1) exhibits symmetry with respect to the HFEVD under both Gaussian and t-distributed innovations.
The stochastic volatility models in (ref) and (ref) are supplied with the parameter values $a=0.5$, $a_1=0.7$, $a_2=0.4$, $\phi=0.6$ and $b=0.7$. This yields:
for the threshold transition and:
for the smooth transition. Several exercises with the HFEVD are performed, to study the linear, specific quadratic and interaction contributions of innovations to the total variance conditional on $\underline{y_t}=(0,0)$:
The results are summarized below in Table 1. As expected, the conditional variance contributions for both models are quite similar. This is not surprising, since they have a similar structure, with the only difference accounted for by the type of transition (threshold vs smooth). The degree of nonlinearity considered is at most up to order three for the threshold model. This accounts for $61.29\%$ ($\approx 21.53+3.38+21.96+10.82+3.60)$ of the total conditional variance, whereas for the smooth transition model, this accounts for $59.02\%$ ($\approx 20.82 + 2.15 + 21.39 + 10.88 + 3.78$) of the total conditional variance. Second, the HFEVD can also be used to inform the practitioner on whether the use of a linear model (or linearization of a nonlinear model) is appropriate in the context of the data. For instance, the linear contribution each innovation component only accounts for $24.91\% (\approx 21.53+3.38)$ and $22.97\% (\approx 20.82 + 2.15 )$ of the variation in $Y_{t+5}$. Hence, the use of a linear model or approximation would misrepresent the trajectory for the underlying data generating process. Third, the HFEVD demonstrates that it is insufficient to include only marginal (specific) higher order terms; in this example, these specific quadratic contributions of the structural innovations account for almost zero in both models. This emphasizes the point of Isakin and Ngo (2020), who argue that interactions are indeed important for IRF and nonlinear FEVD analysis\footnote{As a secondary example, it is suggested in J2005, equation 16, that a local projection (IRF) can be computed based on higher order polynomial terms, while explicitly omitting interactions “as a matter of choice and parismony". This example shows that such a choice would be misguided.} Lastly, although these exercises have included 35 ($7 \times 5$) different Hermite polynomial terms, only 60% of the total conditional variance has been accounted for. Indeed, since the Volterra expansion can potentially have an infinite amount of terms, a highly nonlinear process such as a model with threshold or smooth transitions can be comprised of many higher-order nonlinearities.
The use of the HFEVD in an applied setting is demonstrated through the lens of the FRF2015 Threshold VAR (TVAR) framework (FRF2015 hereafter). The goal of their paper is to study how credit conditions can influence the transmission of fiscal policy shocks in the US macroeconomy. In particular, FRF2015 postulate that fiscal policy channels on aggregate output should be more effective under tighter credit conditions. Indeed, government spending boosts aggregate demand, driving increased household consumption by providing financial support to those who face borrowing constraints. Moreover, while expansionary fiscal policies often crowd-out private investment, they can mitigate financial limitations for firms, making it easier for them to grow when access to credit is limited. To investigate these claims empirically, FRF2015 propose a TVAR model on US quarterly data from 1984 to 2010. This section will first present the FRF2015 methodology, followed by an exposition of how the HFEVD can be used in their framework.
Suppose the dynamics of the macroeconomy are captured by a six-dimensional vector of time series $Y_t=(Y_{1,t},...,Y_{6,t})'$, containing measures of fiscal policy, output, public debt, inflation, monetary policy and credit conditions, respectively. The FFR2015 model is given by:
where the process $(Y_t)$ evolves according to a self-exciting TVAR(1) process [T1983]. The two regimes depend on credit conditions in the previous period, $Y_{6,t-1}$, and the structural innovations $(\varepsilon_{t})$ are assumed to be i.i.d. standard Gaussian. Under “tight" credit conditions characterized by $Y_{6,t-1} \geq r$ (where $r$ denotes the threshold value), the process follows a linear SVAR(1) model governed by coefficients $(A_2,D_2)$. Under “ordinary" credit conditions such that $Y_{6,t-1} < r$, the economy follows a different linear SVAR(1) process, governed by coefficients $(A_1,D_1)$. \\
Table 1 provides a summary of the variables and the transformations applied on the data (see Section 4 of their paper for a detailed discussion). The model is parametric and a two step estimation procedure is used:
The effects of a fiscal policy shock (or its contribution) can then be studied by computing nonlinear IRFs (or nonlinear FEVDs) on its structural innovation, that is, $\varepsilon_{1,t}$.
The HFEVD can be used to assess the nonlinearity of the model and to study the contributions of fiscal policy shocks. This subsection will focus on on six distinct periods in history, representing the starts and ends of three recessions in the United States as defined by the NBER. Table 3 summarizes these dates and indicates whether the observed history is in an “ordinary" credit regime ($r < 0.17$) or “tight" credit regime ($r \geq 0.17$).
A first question of practical significance is how much variation in $Y_{t+h}$ is accounted for by linear contribution of all the innovations, conditional on vector $\underline{y_t}$. This is equivalent of performing HFEVD analysis on a linear combination of all first-order Hermite polynomials for the structural innovations (i.e. the partial Wold representation of the process at horizon $h$). Indeed, if the model is truly nonlinear, one would expect to see varying results for different values of $\underline{y_t}$ (i.e. dependent on history) and contributions of less than 100% (since 100% would imply linearity). In Figure 5, each plot contains six lines, which represent the linear contribution of all structural innovations on the variance of a particular variable, conditional on one of the historical periods. For example, the light blue line in the top left graph represents the linear contribution of structural innovations on the variance of credit conditions, conditional on the data observed in 1990 Q3. Note that in all the graphs, as the horizon increases, the contributions “converge" to a particular value. Indeed, since the data was treated by FRF2015 to ensure its stationarity and the variance will approach its long run value as the forecast horizon tends to infinity. \\
For the top left graph, the HFEVD is performed conditional on the observation at 1990Q3, which was in an ordinary credit regime. Both output and inflation exhibit fairly linear behaviour, with the linear contribution approaching 80% of the total variation in the long run. However, fiscal policy, credit, monetary policy and public debt exhibit more nonlinear behaviour, with the contribution of the linear structural innovations approaching 60% in the long-run. At the end of the Early 90's recession in 1991Q1, there is a change in credit conditions to a tighter regime. The linear behaviour of output, fiscal policy, public debt, and inflation seemingly increase after forecast horizon 5, and converge to a new long-run value. However, the linear contribution to credit and inflation remain the same. In the second row, the results for the Early 00's recession is shown. Output, inflation and fiscal policy tend to exhibit linear behaviour, while public debt, monetary policy and credit behave more nonlinearly. Although the long run linear contribution of the innovations converges to similar values in both 2001 Q2 and 2001 Q4, there are different short-run effects. In particular, for the latter time period, the linear contribution of all structural innovations is 100% for the first three horizons. This suggests that the economy stayed linear in one regime for three forecast horizons before nonlinear dynamics returned. The results for the '08 financial crisis are shown in the last two graphs of Figure 5. The start of the crisis was characterized by an ordinary credit regime\footnote{This is an ordinary credit regime in the framework of this model. However, at this time, there was a decrease of the quality of newly issued credits. This variable has not been introduced in the model and it is important to keep in mind that the contributions may also depend on the list of variables included in the system.}, with the linear behavior of all six variables similar to some of the other periods above. However, at the conclusion of the crisis, it was evident that there were significant short-run and long-run changes on how the structural innovations contributed linearly to the variation in the economy. For credit in particular, there is a significant decline in linear contributions at horizon 2 for 2009Q2, that persists until horizon 6 before returning to the long-run value. On the other hand, inflation and public debt seem to converge towards different values between 2007Q3 and 2009Q2, with a higher proportion of linear contributions for the latter history.
To address original research question of FRF2015, the HFEVD can be used in this context to understand how fiscal policy shocks influence economic output. However, it is important to note that FRF2015 uses the Generalized Impulse Response Function (GIRF) to characterize and compute nonlinear IRFs. As mentioned in Section 4.5 and Appendix B, there are issues with such a definition and the interpretation of the GIRF is not compatible with the HFEVD. Hence, it is important to keep in mind that the results of the HFEVD here are associated with shocks according to the perturbed and baseline paths in (ref), and the EIRF in (ref).
Each graph in Figure 6 contains the marginal contribution of fiscal policy shocks on variance of economic output conditional on the histories in Table 3. There are three insights from these results. First, there is little reason to suggest that the marginal contribution of fiscal policy shocks are nonlinear. Indeed, each graph shows that the linear contribution and the linear + specific higher order contribution (up to degree 5) are not very different\footnote{However, this says nothing about interactions between fiscal policy shocks with shocks on the other innovations. Under recursive ordering, only the fiscal policy shock has economic meaning, so the analysis of interactions are suppressed in this exposition.} Second, FRF2015 “find that the responses of output to fiscal policies significantly change according to the state of credit markets" (Section 1). Indeed, from the start to the end of the Early 90's recession, there is a shift from an ordinary credit regime to a tightening of credit markets. The first two graphs of Figure 6 show that the contribution of fiscal policy shocks on the variance of output is higher in most forecast horizons conditional on 1991 Q1, in agreement with the statement made by FRF2015. Similarly, for the Great Recession, the contribution of fiscal policy shocks seem to have higher contribution on output in 2007 Q3 under a tighter credit regime compared to the ordinary regime in 2009 Q2. On the other hand, the analysis of the Early 2000's crisis dampens their claim. In particular, both the start and end of this recession are characterized by tight credit regimes with similar BAA spreads (0.20 and 0.18 for 2001 Q2 and 2001 Q4, respectively). However, the influence of fiscal policy shocks on the variance of output is lower conditional on the observed history at 2001 Q4. This contradiction with FRF2015 can be explained by the type of conditioning done in their paper. For the GIRF analysis, FRF2015 conditions only on the histories of the credit variable, but average out the histories of all other variables. This means they do not condition on specific dates, but rather average out the impact of fiscal policy shocks on several sets of dates that fit under “tight" credit regimes (or “ordinary" credit regimes). For instance, the results for 1991Q1, 2001Q2, 2001Q4, and 2007Q3 are all tight credit regimes, so their histories would be averaged out to generate a single result. However, even if two histories are within the same regime, there can be very different HFEVDs. As such, there can be misleading insights resulting from averaging out histories.
This paper explores a new method for FEVD analysis in nonlinear SVAR models called the HFEVD. It exploits the properties of Hermite polynomials and represents the trajectory of a Markov processes written on Gaussian innovations using a Hermite polynomial expansion. This representation facilitates the construction of an HFEVD since each term is orthogonal under the standard Gaussian density, leading to a deocmposition of effects by time horizon, by components of the innovations and by the degree of nonlinearity. A link between the terms in the HFEVD and the impact multipliers of the nonlinear IRF is also established. This extends an important property of FEVDs that also exists in the linear framework. Simulation exercises are provided to demonstrate how the HFEVD can be implemented in practice for both parametric models written on Gaussian innovations, or semiparametric models with cross-sectionally independent innovations that can be normalized to Gaussian. An application to fiscal policy shocks and credit markets under the lens of the FRF2015 TVAR framework is also provided. \\
The HFEVD has several useful applications. First, it can be used by policymakers to gauge the importance of innovations in a wide range of nonlinear dynamic models, including the ones that were not discussed in this paper. The granular view of effects allow practioners to assess the properties of shocks in their models that are obscured by existing methods in the literature. Secondly, the HFEVD is an effective tool for model appraisal. For instance, it can be used to evaluate the consequences of linearization for nonlinear models or in determining the feasibility of more complex specifications.