EconBase
← Back to paper

Nonlinear Impulse Response Functions and Local Projections

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

75,878 characters · 29 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonlinear Impulse Response Functions and Local Projections

\setstretch{1}

abstractThe goal of this paper is to extend the nonparametric estimation of Impulse Response Functions (IRF) by means of local projections in the nonlinear dynamic framework. We discuss the existence of a nonlinear autoregressive representation for Markov processes and explain how their IRFs are directly linked to the Nonlinear Local Projection (NLP), as in the case for the linear setting. We present a fully nonparametric LP estimator in the one dimensional nonlinear framework, compare its asymptotic properties to that of IRFs implied by the nonlinear autoregressive model and show that the two approaches are asymptotically equivalent. This extends the well-known result in the linear autoregressive model by Plagborg-Moller and Wolf (2017). We also consider extensions to the multivariate framework through the lens of semiparametric models, and demonstrate that the indirect approach by the NLP is less accurate than the direct estimation approach of the IRF. \\ \noindentKeywords: Nonlinear Autoregressive Model, Impulse Response Function, Local Projection, Nonlinear Innovation, Recurrent Markov Process. \\ \\ \noindentJEL Codes: C01, C22, C32. \\

Introduction

The Impulse Response Function (IRF), a notion initially introduced to economics by Frisch (1933), is a popular tool for macroeconomists to study the effects of shocks in a dynamic framework. In the current literature, there are two popular methods in which the IRF can be estimated. The first is called the direct approach which requires the practitioner to specify an economic model in the form of a (linear) Structural Vector Autoregression (SVAR) [see Sims (1980), Pesaran and Shin (1998), Christiano (2012), Ramey (2016), Kilian and Lutkepohl (2017) for examples]. The IRF is then estimated by simulating a perturbed and a baseline path implied by the autoregressive model, and computing the average difference. The second, which has garnered significant interest in recent studies, is an indirect nonparametric approach by means of a local linear projection [see Jorda (2005) for the definition of local linear projection\footnote{Also see Dufour and Renault (1998) for an interpretation in terms of short and long run causality. See their discussion of Impulse Response Functions on page 1113.}, Chang and Sakata (2007) for its derivation from long run regressions]. Each method boasts its own benefits [e.g. efficiency gains in the direct approach, robustness to misspecification in the indirect approach] and are directly linked to one another in the linear dynamic framework [Plagborg-Moller and Wolf (2021)]. \\

However, linear methods are rather restrictive and constitute a major source of misspecification risk. This arises due to the omission of nonlinear dynamics - for instance, “linear models cannot adequately capture asymmetries that may exist in business cycle fluctuations" [Koop et al. (1996)]. Indeed, nonlinear behaviour has continued to feature in a number of applied settings, including but not limited to: [1] Cyclical behaviour due to a tent map or logistic effects [Frank and Stengos (1988)]. [2] Speculative bubbles and local trends that are captured by causal/noncausal models [Lanne and Saikkonen (2011), Jorda et al. (2015), Gouri\'eroux and Jasiak (2017), (2022), Bec et al. (2020)]. [3] Jumps and disasters [Christoffersen, Du and Elkawhi (2017), Wang (2019), Paul (2020), Gouri\'eroux, Monfort, Mouabbi and Renne (2021)]. [4] Regime switches, introduced to model asymmetric responses of output to oil prices for instance [Hamilton (2003), (2011), Kilian and Vigfusson (2011), (2017)]. [5] Modern macroeconomic models, which need to account for transition to low carbon economics [Metcalf and Stock (2020)], network effects due to supply chain failures [Kuiper and Lansink (2013)], fast technological changes arising from digitalization, potential epidemiological occurrences [Toda (2020)], asymmetric inflation expectations [Baqace (2020)], asymmetric effects of weather shocks [Cashin et al. (2017), DeTruchis et al. (2024)], or for nonlinear summaries of income inequality [Frost and Van Stralen (2018)]\footnote{Such phenomena cannot be captured by simply log-linearizing nonlinear dynamic models around an equilibrium value. In particular, ad hoc methods tend to linearize the first-order conditions of the intertemporal optimization problem, leading to linear rational expectation models. Among the multiplicity of solutions, only stationary fundamental solutions with linear dyanmics are selected, whereas other stationary fundamental solutions with speculative bubbles are rejected ex-ante [see Gourieroux, Jasiak and Monfort (2020) for the complete set of stationary solutions.]. }. \\

The econometric literature on nonlinear IRFs is rapidly developing, with a number of definitions that have been proposed and discussed [Gallant et al. (1993), Koop, Pesaran and Potter (1996), Pesaran and Shin (1998), Gouri\'eroux and Jasiak (2005), Goncalves et al. (2024a)]. Among these studies, it is widely accepted that a nonlinear IRF is characterized by a perturbation of the structural innovation at horizon 1, that is $\varepsilon_{t+1}+\delta$\footnote{See Lee (2025) for a discussion of the various definitions of the nonlinear IRF, and why only this current definition can be linked to nonlinear Forecast Error Variance Decompositions (FEVD).}. While estimation of nonlinear IRFs have been studied extensively in the nonlinear autoregressive framework, the development of estimation methods via nonlinear local projections in this context is rather limited. Moreover, it is an open question whether the link between the direct and indirect methods remains intact in the nonlinear framework. This apparent gap in the literature is what motivates the discussion in this paper. \\

We introduce the nonlinear autoregressive representation of a Markov process and assume that the practitioner has applied restrictions to ensure the identifiability of its structural innovations\footnote{ It is known that without additional parametric or semi-nonparametric restrictions, this representation is not unique as well as the associated nonlinear Gaussian innovation [Gourieroux and Lee (2025)]}. We then establish the link between the nonlinear IRFs implied by this autoregressive model, and the IRFs obtained by means of a nonlinear local projection. Furthermore, we propose nonparametric estimators for both approaches and demonstrate their asymptotic equivalence in the univariate nonlinear setting. For the multivariate setting, the curse of dimensionality severely inhibits the extension of our estimation results. Nonetheless, we also propose semi-parametric alternatives which allow us to mitigate some of the issues related to dimensionality. \\

This paper is organized as follows. Section 2 introduces the nonlinear autoregressive reprsentation. Section 3 provides definitions of shocks, nonlinear IRFs and the nonlinear local projection. We then propose nonparametric estimators of the IRF in Section 4 and discuss their asymptotic properties. The multivariate extension is considered in Section 5, and Section 6 concludes. All derivations of statistical properties of nonparametric and/or parametric estimators of nonlinear IRF are provided in the appendices and standard asymptotic properties of kernel estimators are reviewed in the online appendices.

Nonlinear Autoregressive Representation of a Markov Process

First we introduce the nonlinear autoregressive process and discuss how to obtain its future trajectories. Then, we provide some simple examples of univariate Markov processes that fall under this framework.

The Properties

Let us consider an $n$-dimensional Markov process of order 1, denoted $y = (y_t)$, with values in $\mathcal{Y}=\mathbb{R}^n$. If this process has a continuous distribution, the Markov condition can be written on its transition density as $f(y_t|\underline{y_{t-1}})=f(y_t|y_{t-1})$, where $\underline{y_{t-1}}=(y_{t-1},y_{t-2},...)$. There is an equivalent way of writing the Markov condition.

proposition$(y_t)$ is a Markov process of order 1 on $\mathcal{Y}=\mathbb{R}^n$ with a strictly positive transition density: $f(y_t|y_{t-1})>0$, $\forall y_t,y_{t-1}$, if and only if it admits a nonlinear autoregressive representation: \begin{equation} y_t = g(y_{t-1};\varepsilon_t), \ t \geq 1, \end{equation} where the $\varepsilon_t$'s are independent and identically distributed $N(0,Id)$ variables, with $\varepsilon_t$ being independent of $y_{t-1}$, and $g$ is a one-to-one transformation with respect to $\varepsilon_t$, that is continuously differentiable with a strictly positive Jacobian. The process $(\varepsilon_t)$ defines a Gaussian nonlinear innovation of the process $(y_t)$.

Proof: This is directly deduced from Rosenblatt (1952) [see also Gourieroux and Jasiak (2005) for the terminology of Gaussian nonlinear innovation]. Q.E.D. \\

The condition of Gaussianity on the nonlinear innovation is just a normalization condition. Then, the nonlinear dynamic features can be introduced since $y_t$ is a nonlinear function of $y_{t-1}$ for a given $\varepsilon_t$, and/or a nonlinear function of $\varepsilon_t$ for a given $y_{t-1}$, and/or by nonlinear cross-effects of $y_{t-1}$ and $\varepsilon_{t}$. In particular this nonlinear representation exists even if $y_t$ has marginal and/or conditional fat tails. \\

Remark 1: It is known that, without additional parametric or semi-nonparametric restrictions on the dynamic Markov process, the nonlinear autoregressive representation, i.e. function $g$ and Gaussian nonlinear innovations, are not defined in a unique way [see the literature on nonlinear independent component analysis (ICA) [Comon (1994), Hyvarinen and Pajunen (1999), Hyvarinen et al. (2019), and Gourieroux and Lee (2025) for identification of nonlinear autoregressive representation.]. We assume in the rest of the paper that such restrictions are introduced, that will imply the identification of the structural shocks and of the associated nonlinear IRF. \\

Remark 2: Proposition 1 is easily extended to a multidimensional Markov process of order $p$. The nonlinear autoregressive representation becomes:

equation[equation omitted — 72 chars of source]

We consider the case $p=1$ for ease of exposition and to avoid the curse of dimensionality in nonparametric analysis.\\

Remark 3: The literature has suggested an alternative definition of nonlinear innovation (in the one dimensional framework) as:

equation*[equation* omitted — 102 chars of source]

[see Blanchard and Quah (1989), Koop, Pesaran and Potter (1996), eq. (1), where $\varepsilon_t^*$ is denoted $v_t$.]. It is easily checked that $\varepsilon_t^*$ is not independent of $y_{t-1}$ in general and thus not appropriate for shocking $\varepsilon_t^*$ in the construction of nonlinear IRFs. \\

Future Trajectories

The nonlinear autoregressive representation (ref) is appropriate for describing and/or simulating the future values of the process. We proceed by recursive substitution as follows. First, we have:

equation*[equation* omitted — 209 chars of source]

By substituting equation $y_{t+1}$ into $y_{t+2}$, we obtain:

equation*[equation* omitted — 162 chars of source]

We can perform these substitutions repeatedly for $y_{t+j}$, $j=1,...,h$ to obtain:

equation[equation omitted — 60 chars of source]

where $\varepsilon_{t+1:t+h}=(\varepsilon_{t+1},...,\varepsilon_{t+h}),h\geq1$. In particular we have the recursive formula:

equation[equation omitted — 120 chars of source]

Examples

Example 1: Conditionally Gaussian Model\\

In the one dimensional case, it includes the Double Autoregressive (DAR) model of order one [Weiss (1984), Borkovec and Kluppelberg (2001), Ling (2007)], introduced to account for conditional heteroscedasticity and given by:

equation[equation omitted — 110 chars of source]

where $\varepsilon_t$ is $IIN(0,1)$. The DAR model has a strictly stationary solution if the Lyapunov coefficient is negative:

equation[equation omitted — 72 chars of source]

The conditional heteroscedasticity can create fat tails for the stationary distribution of process $(y_t)$. This process has second-order moments if moreover:

equation*[equation* omitted — 39 chars of source]

When $ \gamma^2 + \beta > 1$ and $\mathbb{E}\log |\gamma + \sqrt{\beta} \ \varepsilon|Z<0$, the process $(y_t)$ is strictly stationary with infinite marginal variance (but finite conditional variance). \\

Example 2: Time Discretized Diffusion Process\\

The discrete time process is $y_t=y(t),t=1,2,...$, where the underlying process $y(\tau)$ is defined in continuous time $\tau \in (0,\infty)$ by a (multivariate) diffusion equation:

equation[equation omitted — 80 chars of source]

where $W$ is a Brownian motion with $\mathbb{V}[dW(\tau)]=d\tau$. The diffusion equation (ref) is the analogue of the conditionally Gaussian model written in an infinitesimal time unit. The class of time discretized diffusion contains the Gaussian AR(1) process, the autoregressive gamma (ARG) process, that is the time discretized Cox, Ingersoll, Ross process [Cox, Ingersoll and Ross (1985)], or the time discretized Jacobi process [Karlin and Taylor (1981), Gour\'ieroux and Jasiak (2009)], with values in $(-\infty,+\infty),(0,+\infty),(0,1)$, respectively. This is easily extended to the multivariate framework, including stochastic volatility models [Hull and White (1987), Heston (1993)] and the Wishart Autoregressive models (WAR) [Gour\'ieroux, Jasiak and Sufana (2009)].

Shocks, Impulse Response Functions and Local Projections

The analysis of shocks and of their propagations in a nonlinear stochastic system, i.e. the impulse response functions, can serve different objectives: [i] It can be a technical tool to analyze the dynamic properties of the system and their robustness. [ii] It can be used as a counterfactual experiment: “What would have arisen if...?". [iii] In a more structural approach, they can be used to evaluate ex-ante the effects of a new policy. This latter objective is more structural and demands to define what can be controlled and at what magnitude. The literature agrees on the fact that a shock performed at date $t+1$ cannot have impacts on the past. It has to be written on an innovation that is on a (multidimensional) variable, independent of its past. It also agrees on the observation that these innovations are not necessarily uniquely defined and that this identification issue has to be taken into account. However, the definitions of the shocked innovations can vary.

Shocks

We consider a transitory shock of magnitude $\delta \in \mathbb{R}^n$ hitting the innovation $\varepsilon_{t+1}$. This is a system wide shock and not a variable specific shock. Thus, $\varepsilon_{t+1}$ is replaced by $\varepsilon_{t+1}+\delta$, while the past values and other future innovations are unchanged. Then we can compare the future trajectories before this shock:

equation*[equation* omitted — 94 chars of source]

and the future trajectories after the shock is applied:

equation*[equation* omitted — 112 chars of source]

These trajectories correspond to a conceptual experiment, since the innovations do not necessarily have economic interpretations. Thus, we trace out the effects of the shock on the outcome future values.

Impulse Response Function (IRF)

Both the initial trajectory and the shocked trajectory are stochastic since they depend on the unknown future values of the innovations $\varepsilon_{t+1},...,\varepsilon_{t+h}$. It is usual to replace the comparison of the stochastic trajectories by a comparison of their summaries. In a nonlinear dynamic framework, a drift of $\delta$ on $\varepsilon_{t+1}$ can have very complicated effects on the distribution of the process. To get enough information on these effects, several notions of Impulse Response Functions (IRFs) will have to be introduced. While our focus here is on the one dimensional case $n=1$, the extensions to the multidimensional framework are straightforward. In particular, it is known that the joint distribution of the components of $y_t$ is characterized by the knowledge of the one-dimensional distributions of the linear combinations $a'y_t$, with $a$ varying. Then, in the multidimensional framework, all the IRF's below can be applied to such linear combinations. \\

In the definitions below, $\mathbb{E}[\cdot|y_t]$ denotes a conditional expectation, that is the best approximation by a square integrable nonlinear function of $y_t$. In nonlinear dynamic models, it is not equal to the theoretical linear regression, that is the projection on the linear space generated by $y_t$. Thus, the distinction between conditional expectations and projection by linear regression is crucial in nonlinear dynamic frameworks \footnote{It also does not coincide with the linear regression on quadratic or cubic functions of $y_t$ [see Jorda (2005), Section II, for this suggestion, called “flexible local projection"]. To insist on this difference, Billingsley (1986) proposed a double bar convention for the conditional expectation $\mathbb{E}[y_{t+h}||y_t]$. We use a single bar later on.}. This conditioning on the past is especially important in nonlinear dynamic models where the effects of a shock (i.e. the IRF) depends on both the current environment (Are we close to a tipping point?) and the magnitude of the shock (Is the shock sufficiently large to cross this tipping point?)\footnote{In this respect, an unconditional definition of IRF is not appropriate and can lead to misleading interpretations of the IRF [see Goncalves et al. (2021), Definition 1, or Goncalves et al. (2024a,b)].}. \\

(i) IRF for pointwise prediction \\

It is defined by:

equation[equation omitted — 101 chars of source]

This is a functional parameter that depends on both the horizon h and magnitude $\delta$ of the shock. It also depends on the history by means of $y_t$. In particular, the IRF has to be updated at each new observation. It is called the Generalized IRF in Koop, Pesaran, and Potter (1996)\footnote{It is also called expected IRF in some other literatures, to highlight the pointwise predictor by conditional expectations.}. \\

(ii) IRF for pointwise prediction of the transformed $y$ \\

Let us consider a nonlinear transformation $a(y)$ of $y$. The IRF is defined by:

equation[equation omitted — 88 chars of source]

For instance, if $n=1$ and $a_y(y_{t+h})=\mathbbm{1}_{y_{t+h}<y}$, we get:

equation[equation omitted — 134 chars of source]

and the possibility to compare the predictive distributions at all horizons. By inverting these cumulative distribution functions, we can also consider the effect of shocks on the conditional quantiles, that are the conditional Value-at-Risk (VaR) [Gouri\'eroux and Jasiak (2005)]. These IRF based on VaR are used to evaluate the sensitivity of reserves for banks in prudential supervision. \\

(iii) IRF for dynamic features \\

It is also important to evaluate the effects on dynamic features, such as the conditional serial dependence at lag 1. This would lead to the IRF's of the type:

equation[equation omitted — 109 chars of source]

This can be applied to the autocorrelation function (ACF) as well as the squared ACF. Indeed, in nonlinear dynamic models, a small shock can significantly change the dynamics, including the linear ACF due to chaotic features. \\

(iv) Joint IRF \\

All IRF's above are computed without taking into account the dependence between the trajectories. Other tranformations can reveal this cross-dependence, such as:

equation[equation omitted — 85 chars of source]

that includes a cross term: $\mathbb{E}[y_{t+h}^{(\delta)}y_{t+h}|y_t]$.

Nonlinear Local Projections (NLP)

Let us consider the $IRF(h,\delta)$ defined in (ref), and introduce the pointwise prediction at horizon $h$:

equation[equation omitted — 69 chars of source]

By the Markov property, this pointwise prediction depends on the past by means of $y_t$ only. Such direct prediction has been called local projection in the linear dynamic framework [see Jorda (2005)]. By analogy with what is known in this linear framework [Plagborg-Moller and Wolf (2021), Montiel Olea and Plagborg-Moller (2021),(2022)], we will relate the IRF's and the NLP associated with prediction (ref).

proposition\begin{equation} \begin{split} IRF(h,\delta|y_t)&=\mathbb{E}\{m^{(h-1)}[g(y_t;\varepsilon_{t+1}+\delta)]-m^{(h-1)}[g(y_t;\varepsilon_{t+1})]|y_t\} \\ & = \int \{m^{(h-1)}[g(y_t;\varepsilon+\delta)]-m^{(h-1)}[g(y_t;\varepsilon)]\}\phi(\varepsilon)d\varepsilon,\\ \end{split} \end{equation} where $\phi$ is the density of $N(0,Id)$, for $h\geq 1$\footnote{For $h=1$, $m^{(0)}$ is the identity function.}.

Proof: We have:

equation*[equation* omitted — 557 chars of source]

by definition of $m^{(h-1)}$. The result follows. Q.E.D. \\

Proposition 3 is easily extended to the other types of IRF's. The right hand side of equation (ref) defines the NLP interpretation of the IRF. As seen below, it differs from the standard formula since the effect of the shock is nonlinear in $\delta$ and has to be integrated out with respect to $\varepsilon$. The integration in equation (ref) can be avoided in special cases.

corollaryLet us assume that: \begin{equation*} m^{(h-1)}[g(y_t,\varepsilon_{t+1})] = a^{(h-1)}(y_t)\varepsilon_{t+1}+b^{(h-1)}(y_t). \end{equation*} Then: \begin{equation*} IRF(h,\delta|y_t) = a^{(h-1)}(y_t)\delta. \end{equation*}

The condition in Corollary 1 is satisfied in the linear dynamic models usually considered in the literature. It provides the same IRF's as the IRF comparing two shocked trajectories, one with $\varepsilon_{t+1}$ replaced by $\delta$, and another with $\varepsilon_{t+1}$ replaced by $0$\footnote{This corresponds to the so-called MIT definition of IRF. However, this definition is not appropriate for more complicated nonlinear features as noted in Kolesar and Plagborg-Moller (2024).}. Indeed, only the difference matters. In this case it can also be normalized by focusing on $\delta=1$, since $IRF(h,\delta|y_t)$ becomes linear in $\delta$. However, without the strong linearity restriction in Corollary 1, the IRF will depend on the current environment $y_t$, and on the magnitude of the multivariate shock $\delta$, in a nonlinear way (in general). Other simplifcations of formula (ref) are expected in specific dynamic models as fully recursive structural models [see Goncalves et al. (2021), Section 5 with the i.i.d. assumption on the “regressors" and the discussion in Section 5.3]. As mentioned by the authors this assumption of full recursivity is not economically plausible in general.

Sequence of Shocks

We have considered in the subsections above the case of a transitory (i.e. isolated) shock performed at date $t+1$. In a linear dynamic framework, it is standard to also consider sequence of shocks and in particular permanent shocks applied after this date. However, the analysis of the IRF following a sequence of shocks is significantly different in a nonlinear dynamic framework. For illustration, let us consider below two consecutive shocks of magnitude $\delta_1$ and $\delta_2$ at dates $t+1$ and $t+2$ respectively. Then, the future trajectories are:

equation[equation omitted — 161 chars of source]

and the IRF becomes:

equation[equation omitted — 113 chars of source]

The IRF can be written in terms of nonlinear local projections as:

equation[equation omitted — 358 chars of source]

This expression extends formula (ref) in Proposition 3 to a sequence of two consecutive shocks [see Diercks et al. (2023) for another attempt to extend the local projection to this framework]. In the linear dynamic framework, this IRF can be decomposed as:

equation[equation omitted — 111 chars of source]

that is directly deduced from the IRF's of the isolated shocks. This equality (ref) is no longer valid in the nonlinear dynamic framework. In general, the $IRF(h;\delta_1,\delta_2)$ is not a deterministic function of the IRF's associated with the isolated shocks and, morever, the combined effect of the two shocks can have diminishing, neutral, or amplifying effects on the values $y_{t+h}$, due to cascading effects. Therefore “with nonlinear models considering the IRF that measures the effect of a shock of a given size hitting at a given period can be very misleading" [Koop (1996), p 136]. Instead of computing the IRF for all values of the pair $(\delta_1,\delta_2)$, it has also been proposed to draw them in some distribution in order to account for the uncertainty on $(\delta_1,\delta_2)$ [Koop (1996)]. \\

Note that the definition and expansion of IRF are written directly from the autoregressive representation without deriving the nonlinear moving average representation of the process [see the discussion in Adamek et al. (2024), p. 324, for the drawback of such an approach even in the linear dynamic framework]. Indeed this approach based on autoregressive representation is more appropriate for deriving by simulation the different IRFs.

Nonparametric Inference for n=1

An argument for local projection analysis of IRF's is that this is a nonparametric approach that is less sensitive to possible mispecification of the lag in the autoregressive parametric dynamic. This argument has been criticized [see Kilian and Kim (2009), p.1466]. Moreover, it is given assuming a linear dynamic model, whereas one of the main sources of misspecification is likely the omitted nonlinear dynamics\footnote{The omitted nonlinearities do not only concern omitted conditional heteroscedasticity [see e.g. Kilian and Kim (2011), Herbst and Johannsen (2021), Montiel-Olea and Plagborg Moller (2021) for bias correction of the IRF standard errors under conditional heteroscedasticity either by bootstrap, or by expansions.]}. Let us now compare the direct IRF estimation approach and the LP estimation approach in a nonlinear dynamic framework. Due to the curse of dimensionality of conditional nonparametric analysis, we consider the one dimensional case $n=1$ and the observations $y_1,...,y_T$ of the process. By the Markov property, we also assume the same lag equal to 1 for the direct and NLP estimation approaches. Then the two approaches will agree and nevertheless are both nonparametric.

Two Nonparametric Estimation Approaches

When $n=1$, the nonlinear AR(1) [NLAR(1)] model is identifiable (up to a change of sign on $\varepsilon_t$). Then we can consider a nonparametric functional estimator of function $g$ as:

equation[equation omitted — 85 chars of source]

based on a kernel estimation of the conditional quantile of $y_t$ given $y_{t-1}$:

equation[equation omitted — 171 chars of source]

where $K$ is the kernel, $b_T$ the bandwidth and $\alpha$ the critical level\footnote{The kernel and bandwidth could depend on the level $\alpha$ to treat differently the smoothing in the standard values and in the tails. This question is out of the scope in the present paper. }. Then, this estimated NLAR(1) model can be used to simulate future trajectories $y^s_{t+k}$, $k=1,..,h$, $s=1,2,...,S$, by applying recursively the nonlinear autoregressive equation:

equation[equation omitted — 92 chars of source]

where the $\varepsilon_t$'s are independently drawn from the standard normal distribution, with starting value $\hat{y}^s_t=y_t$ \footnote{These simulated values depend on $T$ by means of $\hat{g}_T$. This index $T$ is omitted below for exposition, and they also depend on $y_t$.}. For illustration we focus on the IRF for $\mathbb{E}(y_{t+h}|y_t)$. \\

(i) Direct estimation of $IRF(h,\delta|y_t)$\\

The direct approach follows the steps below: \\

Step 1: Simulate $\hat{y}_{t+h}^s$ and $\hat{y}_{t+h}^{(\delta),s}$, $s=1,...,S$, based on $\hat{g_T}$. \\

Step 2: Compute:

equation[equation omitted — 132 chars of source]

(ii) Estimation by means of Local Projection (Indirect Estimation)\\

The indirect approach is based on formula (ref) in Proposition 3. The steps become: \\

Step 1: Compute the Nadaraya-Watson estimate\footnote{This estimator is based on a smoothing with respect to the conditioning value $y$. This differs from the smooth local projections in which the smoothing is with respect to $h$ [Barnichon and Brownlees (2019), Plagborg-Moller and Wolf (2021)].} of the function $m^{(h)}(\cdot)$. This provides:

equation[equation omitted — 279 chars of source]

where $K$ is a kernel and $b_T$ a bandwidth. \\

Step 2: Simulate $\varepsilon^s_{t+1}$, $\hat{y}_{t+1}^s$ and $\hat{y}_{t+1}^{(\delta),s}$, $s=1,...,S$. \\

Step 3: Compute:

equation[equation omitted — 180 chars of source]

We will compare, in the next subsection, the asymptotic properties of these functional estimators given in (ref) and (ref). But clearly, to get $IRF(h,\delta|y_t)$, $h=1,...,H$ for a given $\delta$, the first approach requires the nonparametric estimation of function $g$ at $2HS$ values (that are the $\hat{y}^s_{t+k},\hat{y}_{t+k}^{(\delta),s},k=1,...,h, s=1,...,S)$. The second approach requires the nonparametric estimation of function $g$ at $2S$ values (that are the $\hat{y}^s_{t+1},\hat{y}_{t+1}^{(\delta),s},s=1,...,S)$, and of functions $\hat{m}^{(h)}$ at $2(h-1)S$ values. Therefore, they seem numerically equivalent in the nonlinear dynamic framework\footnote{Note that the direct and indirect estimators coincide for $h=1$.}. \\

Their asymptotic accuracies will depend on the nonparametric estimation method that is used and in particular on the choice of kernel and bandwidth. Nevertheless, the first approach estimates in an “efficient" way the transition at horizon 1 and we use this estimator to compute the IRF. The second approach is mixing an “efficient" estimator of transition $g$, with an estimator $\hat{m}_T^{(h-1)}$, that does not account for its known dependence with respect to $g$. Therefore, we expect the second approach to be less accurate asymptotically under coherent choices of kernel and bandwidth in the two approaches. We show in the next subsection that this is not the case.

Asymptotic Properties

Let us now derive the asymptotic behaviours of the functional estimators of interest that are: $\hat{g}_T(\cdot)$, $\widehat{IRF}_T(h,\cdot|y_t)$, $\widehat{\widehat{IRF}}_T(h,\cdot|y_t)$. All these estimators are consistent with respect to their theoretical counterparts and they converge at different speeds. We perform in Appendix A the asymptotic expansions of the estimated IRF's in terms of the nonparametric estimates of the primitive characteristic of the conditional transitions. This allows for deriving the asymptotic behaviour of the estimated IRFs. For expository purposes, the derivations are performed in the one dimensional case and for stationary processes\footnote{The asymptotic results could be extended to nonstationary Markov processes if they satisfy a null recurrence property [Karlsen and Tjostheim (2011), Gourieroux and Jasiak (2019)]. This is for instance the case of the Gaussian random walk. Then the speed of convergence is modified, will depend on the number of regenerations in the sampling period, but not on the initial value $y_0$ of the process. We do not develop this extension in this paper [see Wright (2000), Gospodinov (2004), Pesavento and Rossi (2007), Montiel-Olea and Plagborg-M\o ller (2021) for the effect of unit roots in the linear local projection, which is the analogue in the semi-parametric linear framework].}. Then, the asymptotic variances are functions of the conditional cumulative distribution functions (c.d.f.), $F(\mathfrak{z}|y)$, the transition density, $f(\mathfrak{z}|y)$, and the conditional quantile, $Q(\alpha|y)$. \\

propositionThe nonparametric direct and local projection estimators have the following asymptotic properties: \\ (i) Direct Estimation \\ Let us assume $T \rightarrow \infty$, $S \rightarrow \infty$, with $S/T \rightarrow 0$, $b_T \rightarrow 0$, with $Tb_T^{5/3} \rightarrow \infty$. Then, $\widehat{IRF}_T(h,\delta|y_t)$ converges to $IRF(h,\delta|y_t)$ for any $h,\delta$. Morever, we have the convergence in distribution: \begin{equation*} \sqrt{Tb_T}(\widehat{IRF}_T(h,\delta|y_t)-IRF(h,\delta|y_t)) \rightarrow^d N(0,\sigma^2(h,\delta|y_t)), \end{equation*} where the expression of $\sigma^2(h,\delta|y_t)$ is derived in Appendices A.1 for $h=1$ and A.2 for $h \geq 2$. \\ (ii) Local Projection \\ We have: \begin{equation*} \sqrt{Tb_T}[\widehat{\widehat{IRF}}_T(h,\delta|y_t)-\widehat{IRF}_T(h,\delta|y_t)] = o_p(1), \end{equation*} In particular, the NLP estimator $\widehat{\widehat{IRF}}_T(h,\delta|y_t)$ converges to $IRF(h,\delta|y_t)$, with the same speed of convergence as the direct estimator $\widehat{IRF}(h,\delta|y_t)$ and they have the same asymptotic distribution and asymptotic variance. \\

Proof: See Appendices A.1, A.3, A.4 for the expansions and online appendix D for regularity conditions. \\

Recall the standard arguments usually provided in semi-parametric linear dynamic models when comparing the direct and indirect approaches. If the semi-parametric model (i.e. the lag) is well specified, the direct approach is more efficient than the local projection approach. If the semi-parametric model (i.e. the lag) is misspecified (for instance, if the true model is a VAR$(\infty)$), the local projection approach (applied with a sufficiently large controlled lag order \footnote{See Xu (2023), Montiel-Olea et al. (2024), for statistical inference for local projections when the controlled lag order diverges.} \footnote{By increasing the number of lags in the nonlinear local projection approach, we will encounter the curse of dimensionality on nonparametric approaches (see discussion in Section 5).}) is better. The proposition above shows that, under the Markov assumption, both approaches are consistent and asymptotically equivalent in a pure nonparametric approach, for both linear and nonlinear dynamic models. Moreover, they have the same speed of convergence and the same asymptotic distribution. Therefore, neither nonlinear local projections, nor nonlinear autoregressions dominates\ the other in terms of asymptotic inference. They can differ however by their finite sample properties, and also by the computational time that they require. \\

The results of Proposition 3 can be compared with the standard results when the IRF is computed by linear local projection (LP). In this linear semi-parametric framework, the associated multi-step error forecasts are generally serially correlated [see Montiel-Olea and Plagborg-Moller (2021), p 1793] and the standard errors of the IRF are usually estimated by heteroscedasticity and autocorrelation (HAC) consistent approaches [Jorda (2005), Ramey (2016), Kilian and Lutkepohl (2017)]. Such an adjustment is not needed with the nonparametric approaches, since the functional estimators are not only local in $h$, but also local in the value of the conditioning variable.

An Illustration

This section illustrates the above concepts through the lens of the DAR(1) model in Example 1. In particular, we consider the data generating process:

equation*[equation* omitted — 78 chars of source]

where $y_0=0.2$ and $\varepsilon_t \sim N(0,1)$. These parameter values ensure that there exists a strictly stationary solution with second-order moments.

The Data Generating Process

We simulate 200 observations of the DAR(1) process described above and plot both its trajectory and empirical density in Figure 1. The empirical density features rather fat tails, that is a standard effect of conditional heteroscedasticity on the marginal (i.e. stationary) distribution.

figure[figure omitted — 167 chars of source]

Estimators

The DAR model is semi-parametric, since its parameters and error distribution are not known. For estimation of the IRF, we first need to estimate the function $g$. To do so, we first perform a Quasi Maximum Likelihood Estimation (QMLE) for the parameters $\rho$, $\alpha$ and $\beta$, that is a MLE as if the $\varepsilon_t$ were Gaussian. The solution to the quasi log-likelihood can be written as [Ling (2007)]:

equation*[equation* omitted — 241 chars of source]

We perform a three dimensional grid search on the cube $[0.01,1.20]^3$ at intervals of size 0.01. Our results yield estimates $(\hat{\rho}_{MLE},\hat{\alpha}_{MLE},\hat{\beta}_{MLE})'= (0.35,0.9,0.46)'$, to be compared with the true values of the DGP, that are $(0.5, 1, 0.5)$.

figure[figure omitted — 209 chars of source]

For the local projection approach, we will have to run the nonparametric regression of $y_{t+h}$ on $y_t$ to obtain the Nadaraya-Watson estimates $\hat{m}_T^{(h)}$ in the NLP equation. In Figure 2, we compare the estimated regression line $\hat{m}_T^{(1)}$ and the true function, as well as the QQ plot of the residuals.

IRF

We present shocks of magnitude $-1, -0.5, 0.5$ and $1$ below in Figure 3. Each graph features three IRFs: [1] The true IRF based on equation (ref); [2] The IRF obtained by the (semi-parametric) direct estimation based on equation (ref) with QMLE estimation; [3] The IRF obtained by means of nonlinear local projection based on equation (ref). We see that, in all cases, holding the simulation size equal, the direct estimation method provides a more accurate IRF. This is not surprising since the nonlinear local projection is based on the Nadaraya-Watson estimator, which takes longer to converge than the QMLE method.

figure[figure omitted — 153 chars of source]

Statistical Inference for Multivariate Case

The nonparametric approaches of Section 4 cannot be easily extended to the multivariate framework with a dimension $n$ larger or equal to 3. Indeed, we encounter the curse of dimensionality due to the conditioning variable in kernel estimation of conditional cdfs and conditional expectations. The curse of dimensionality will be reached even faster for Markov processes of order $p$, $p>1$. Indeed, the rate of convergence of these nonparametric estimators deteriorates quickly [i.e. at an exponential rate] as the dimension grows [see Geenens (2011), Conn and Li (2018)]. In this section, we consider semi-parametric models in which the nonparametric dimension is reduced\footnote{To circumvent the curse of dimenisonal for either $n$, or $p$ larger than 2, say, Goncalves et al. (2024 a,b) propose to modify the conditioning set. Instead of conditioning by $y_{t-1}$ (when $p=1$ for instance), they suggest to condition by a one dimensional function $a(y_{t-1})$ of $y_{t-1}$. This solves the curse of dimensionality issue, but at the cost of a loss in information, and a change of the definition of the IRF, which depends on the transformation $a(\cdot)$. }.

Semiparametric Models

Let us now provide additional examples of multivariate Markov processes. \\

Example 3: Nonlinear Conditionally Gaussian Model\\

It is written as:

equation[equation omitted — 74 chars of source]

where $\varepsilon_t\sim IIN(0,Id)$ and is assumed independent of $y_{t-1}$. This specification is used in Koop, Pesaran and Potter (1996, eq.(1)) for instance\footnote{See also Goncalves et al. (2021), Ballarin (2025) without conditional heteroskedasticity}. Except in the Gaussian linear VAR(1) framework: $m(y_{t-1};\theta)=A(\theta) y_{t-1}, D(y_{t-1};\theta)=D(\theta)$, independent of the past, the process is not conditionally Gaussian at horizons larger or equal to 2. The nonlinear effects will appear on the IRF's after $h \geq 2$. \\

For instance, the multivariate extension of the DAR model can be written as:

equation*[equation* omitted — 77 chars of source]

with $h_t=a + B(y^2_{1,t-1},...,y^2_{n,t-1})$ [see Zhu et al. (2017)]. \\

Example 4: Strong Linear Causal Structural SVAR(1) Model \\

The process is defined as the stationary solution to the linear difference equation:

equation[equation omitted — 55 chars of source]

where the eigenvalues of autoregressive matrix $A$ have a modulus strictly smaller than 1, $D$ is an invertible matrix, and the $u_t$'s are i.i.d, with independent components. If the cdf of $u_{it}$, $i=1,...,n$, is $F_i$, assumed to be invertible, then this model can be rewritten with Gaussian errors as:

equation[equation omitted — 93 chars of source]

where $\Phi$ is the cdf of the standard normal distribution. The strong linear SVAR(1) model (ref) has in general a nonlinear dynamic feature when it is written with respect to Gaussian errors, by the nonlinear transformation $ F^{-1}_i \circ \Phi=Q_i \circ \Phi, \ i=1,...,n$, where $Q_i=F^{-1}_i$ is the quantile function. However, it is linear in $y_{t-1}$ and contains no cross-effects of $y_{t-1}$ and $\varepsilon_t$. We get a semi-parametric model that includes vector parameter $\beta=((\text{vec A})',(\text{vec D})')'$ and functional parameters $Q_i$ for $i=1,...,n$. These functional parameters depend on one-dimensional arguments, which allows us to circumvent the nonparametric curse of dimensionality. \\

Let us now discuss the respective roles of specifications (ref) and (ref).

enumerate• If $u_t=\varepsilon_t$, it is well-known that the parameter $A$ is identifiable, whereas $D$ is not identifiable. Therefore, the IRF becomes identifiable only if $A(\theta)$, $D(\theta)$ are (jointly) parameterized with identifiable $\theta$. The IRF has the simple form: \begin{equation*} IRF(h,\delta)=D\delta + ... + A^hD\delta = (Id-A)^{-1}(Id-A^{h+1})D\delta. \end{equation*} • If $u_t$ differs from $\varepsilon_t$ and if at most one component of $u_t$ is Gaussian, it is known that the model (ref) is semi-parametrically identifiable (up to permutation of indexes), that is parameters $A,D$ and functional parameters $F_i(\cdot)$, $i=1,...,n$ are identifiable by applying the identification results in linear ICA [see Comon (1994) for the linear ICA, and Gourieroux, Monfort and Renne (2017) for the application to SVAR models].

The transformations of the error term $u_t$ into a Gaussian error $\varepsilon_t$, when the components of $u_t$ are independent, can also be used to extend the models of Example 3 to non-Gaussian errors and then render semi-parametric the models in Example 3. Therefore, the reduction of the curse of dimensionality is due to the assumption of independent components for $u_t$, that is the possibility to replace the nonparametric estimator of the joint density of $u$ by the nonparametric estimation of the marginal distributions $F_i$, or $Q_i$, for $i=1,...,n$.

Direct vs Indirect Estimation

Let us now introduce the different (semi-) parametric estimation approaches of the model parameters, then of the IRF.

Nonlinear Structural VAR (NSVAR)

Let us consider the nonlinear autoregressive model:

equation[equation omitted — 57 chars of source]

where the components of the errors are $u_{i,t}=F_i^{-1}\cdot \Phi (\varepsilon_{i,t})$, $i=1,...,n$, and $(\varepsilon_t)$ is $IIN(0,Id)$. Then, model (ref) can be written under the equivalent form:

equation[equation omitted — 59 chars of source]

where $u_t = F_i^{-1}\circ \Phi (\varepsilon_{i,t})=Q_i\circ \Phi(\varepsilon_{i,t})$, $G$ is a demixing transformation, that is, the inverse of function $g$ with respect to $u_t$, and $\beta$ is a parameter. Then we can consider a parametric and a semi-parametric version of this model: \\

1. Parametric Model: \\

When the distributions $F_i$, $i=1,...,n$, are parameterized, with parameter $\alpha$, the model (ref) and (ref) becomes parametric with global parameters $\theta'=(\alpha',\beta')$. It can be estimated by maximum likelihood for instance. We denote $\hat{\theta}_T$ the corresponding estimator. \\

2. Semi-Parametric Model:\\

Under the assumption that the parameter $\beta$ is identifiable, it is possible to construct a consistent and asymptotically normal estimator of $\beta$. Then, this estimator satisfies:

equation[equation omitted — 90 chars of source]

for instance by applying a minimization of the cross covariances and autocovariances of nonlinear functions $G(y_t,y_{t-1},\beta)$ [see Gouri\'eroux and Jasiak (2023), Velasco (2023), for Generalized Covariance (GCov) estimators]. Once the parameter $\beta$ is estimated, we can compute the estimated errors as: $\hat{u}_{t,T} = G(y_t,y_{t-1};\hat{\beta}_T)$, $t=1,...,T$, and deduce nonparametric estimators $\hat{Q}_{i,T}$ of $Q_i$, for $i=1,...,n$ by applying the kernel approach (ref) to the series of residuals.

Direct Estimation

1. Parametric Model: \\

The estimated IRF is obtained by plugging in the theoretical expression of the IRF the maximum likelihood estimator $\hat{\theta}_T$ of $\theta$:

equation[equation omitted — 82 chars of source]

2. Semi-Parametric Model:\\

For exposition, let us consider the case $h=1$. In the semi-parametric framework, the IRF becomes a function of parameter $\beta$ and functional parameters $Q_i$, $i=1,...,n$, $IRF(h,\delta;\beta, (Q_i)|y_t)$, that is:

equation[equation omitted — 560 chars of source]

where $\text{vec}(a_i)$ denotes the $n$-dimensional vector with components $a_i$. When the model is well-specified with true parameters $\beta_0$, $Q_{i,0}$, $i=1,...,n$, the true IRF is $IRF(1,\delta;\beta_0,(Q_{i,0}))$. It can be estimated by plugging in the estimator $\hat{\beta}_T$ of $\beta$ and the functional estimator $\hat{Q}_{i,T}$ of $Q_i$ for $i=1,...n$. The resulting estimator $IRF(1,\delta,\hat{\beta}_T,(\hat{Q}_{i,T}))$ is consistent, asymptotically normal, and its asymptotic distribution is obtained by the delta method applied to both parameter $\beta$ and functional parameters $Q_i$, $i=1,...,n$ [see Appendix A.5].

propositionUnder standard regularity conditions, the asymptotic distribution of the estimated IRF is such that: \begin{equation} \sqrt{Tb_t}\left(IRF\left[1,\delta;\hat{\beta}_T,(\hat{Q}_{i,T})\right]-IRF\left[1,\delta;\beta_0,(Q_{i,0})\right]\right)\xrightarrow[]{d}N(0,V(1,\delta,\beta_0,(Q_{i,0}))), \end{equation} where $V(1,\delta,\theta_0,(Q_{i,0}))$ is the asymptotic variance given in Appendix A.5.

Proof: See Appendix A.5.\\

Remark 4: The limiting distributions in (ref) have been written for any given pair $h$,$\delta$. They can be extended to several pairs by taking into account the asymptotic covariances between $\widehat{IRF}_T(h,\delta|y_t)$ and $\widehat{IRF}_T(h^*,\delta^*|y_t)$ for two pairs $(h,\delta)$ and $(h^*,\delta^*)$. This joint inference can be used to provide confidence bands for the term structure of the IRF for a given $\delta$, and/or for the confidence bands of the IRF function of the magnitude of the shock for a given horizon\footnote{See Inoue et al. (2024), Section 5.2, for “simultaneous inference" in a linear dynamic setting.}. \\

In the nonlinear autoregressive model, the transition at horizon larger than 2 has no closed form expression and has to be approximated by simulation based on the estimated model $y_t = g(y_{t-1},\varepsilon_t,\hat{\beta}_T,(\hat{Q}_{i,T}))\equiv \hat{g}_T(y_{t-1},\varepsilon_t)$. This is done along the same steps as in Section 4.1. The number $S$ is chosen by the econometrician. If $S$ is chosen much larger than the number of observations $T$, that is if $\frac{S_T}{T}$ tends to infinity with $T$, then the asymptotic behaviour of the estimated IRF based on simulations is the same as in (ref). \\

Remark 5: Instead of the recursion:

equation*[equation* omitted — 140 chars of source]

corresponding to the recursion (4.3), it would be possible to apply a “bootstrap" recursion of the form:

equation*[equation* omitted — 102 chars of source]

where the $\hat{u}^s_{t+k,T}$ are independently drawn among the residuals of $\hat{u}_{t,T}$, $t=1,...,T$. The asymptotic properties of this bootstrap version however, is beyond scope of this paper and left for future research.

Indirect Approach (Local Projection) for n =2

Due to the curse of dimensionality of nonparametric approaches, the nonlinear local projection approach will faces challenges for a dimension strictly greater than 2 given the number of observations usually available in macroeconomic applications. If $n=2$, we can use a local projection approach based on the formula of Proposition 2 in which the short term responses $g(y_t,\varepsilon_t+\delta;\theta)$, $g(y_t,\varepsilon_t;\theta)$ are estimated parametrically, whereas the conditional expectation $m^{(h-1)}(\cdot)$ is estimated nonparametrically by the Nadaraya-Watson approach, say. This approach avoids the simulation based parametric approach used in Section 5.1.1 to approximate $m^{(h-1)}(\cdot)$, but with the drawback of a nonparametric rate of convergence infinitely smaller than the rate of convergence in the direct parametric and semi-parametric approach in Section 5.2.2. Under the assumption of a well-specified model, the rates of convergence for the estimated IRF to their true values will be the parametric rate of $1/\sqrt{T}$ in a parametric direct approach, the nonparametric rate of $1/\sqrt{Tb_T}$ for one-dimensional kernel estimators in the semi-parametric direct approach and the nonparametric rate of two dimensional kernel estimators in the indirect local projections approach\footnote{This curse of dimensionality problem is also encountered in the multivariate VAR model in which the local projection implies a regression with a rather large number of control variables and is sometimes approximated by sparsified LASSO [see Adamek et al. (2024)].}. To summarize, the nonlinear local projection approach cannot profit from the reduced nonparametric dimensionality of the semi-parametric model (5.4).

Illustrations

Let us now illustrate the estimation approaches in a bivariate nonlinear dynamic model.

The Simulated Data

Let us consider a bivariate Double Autoregressive model defined by:

equation*[equation* omitted — 124 chars of source]

where $h_t = b+ A

bmatrix[bmatrix omitted — 46 chars of source]

$ and the same experiment as in Zhu et al. (2017), Section 3. The parameter $\beta=\left[(vec $\Phi$)',b',(vec A)'\right]'$ is fixed at $\beta = (0.4,0.1,-0.3,0.4,0.1,0.2,0.3,0.2,0.1,0.4)'$, $u_{1,t}$ and $u_{2,t}$ are independent with distributions $N(0,1)$ and student $t(4)$, respectively. The number of observations is fixed to $T=400$, and we plot the simulated series below in Figure 4.

figure[figure omitted — 101 chars of source]

Semi-Parametric Estimation

The parameter $\beta$ is estimated in two steps. First, we regress by OLS $Y_t$ on $Y_{t-1}$, that provides an estimator $\hat{\Phi}_T$ of $\Phi$ and deduce $\hat{\upsilon}_{t,T} = y_t -\hat{\Phi}_Ty_{t-1}$. In a second step, we regress by OLS the squared $\upsilon$'s, i.e. $

bmatrix[bmatrix omitted — 72 chars of source]

$ on $

bmatrix[bmatrix omitted — 44 chars of source]

$ with intercept. The results of these regressions are tabulated below in Tables 1 and 2, with the usual OLS standard errors. Note that these standard errors are not adjusted for conditional heteroscedasticity (in both regressions), for the effect of the first step estimation in the second regression, and the fact that the student distribution t(4) has no fourth-order moments. As a consequence, the standard errors in Table 2 are likely underestimated. \\

table[table omitted — 808 chars of source]

Using the results for $\hat{b}_T$, $\hat{A}_T$, and $\hat{v}_{t,T}$, we compute:

equation*[equation* omitted — 183 chars of source]

for $t=2,...,T.$ These are used to derive the estimated quantile functions $\hat{Q}_{i,T}$, $i=1,2$. We provide below in Figure 5 their Q-Q plots with respect to the standard normal distribution.

figure[figure omitted — 237 chars of source]

As expected, the values in the QQ plot for $\hat{u}_{1,t,T}$ lie close to the 45 degree line since the residuals are normally distributed. On the other hand, the QQ plot for $\hat{u}_{2,t,T}$ deviate from the 45 degree line further from the mean due to the fatter tails of the t-distribution.

The “True" and Estimated IRF

To visualize the nonlinear local projection in the context of the bivariate DAR model, we consider four different shock scenarios with the same magnitude but different signs: $\delta = (+0.5,+0.5)$, $\delta = (+0.5,-0.5)$, $\delta = (-0.5,+0.5)$ and $\delta = (-0.5,-0.5)$. The results are featured in Figure 6 below. The red lines and blue lines represent the responses of the first and second variables, respectively. The dotted and solid lines represent the “true" IRFs and the estimated IRFs, respectively. Note that a larger sample of 3000 observations were used to generate nonlinear local projections. This reveals that the nonparametric approach is computationally expensive and requires a rather large data set on hand. For instance, even in the case of a $\delta=(+0.5,-0.5)$ shock to the process, the estimated IRF is still far away from the truth.\\

The behaviours of the plots showcase the nonlinear dynamics at play for the bivariate DAR process. We again see the absence of shock symmetry, as a positive and negative shock of the same magnitude to not seem to mirror one another in all scenarios. We also see that the response of $Y_{1,t}$ and $Y_{2,t}$ are different for shocks of the same magnitude and direction. This reflects not only the different parameterization for each variable in the DGP, but also the presence of non-Gaussianity in the second component of the innovation.

figure[figure omitted — 536 chars of source]

Concluding Remarks

This paper has extended the notions of innovations, impulse response functions and local projections to the nonlinear dynamic framework. In particular, we have precisely defined the nonlinear Gaussian innovation and the IRFs corresponding to shocks on these innovations and derive an integral expression of the IRF that extends to a nonlinear dynamic framework the idea of local projection. Then this integral expression can be used to extend the idea of local projection at least for small dimension of the variable to be predicted. In the one-dimensional framework, we have precisely compared the asymptotic distributions of the IRF estimated nonparametrically by direct and an indirect local projection approach, respectively, and show that they are asymptotically equivalent. This generalizes a result already derived in the literature for linear dynamic models, based on the Frisch-Waugh-Lovell Theorem. For a larger dimension, we can encounter the curse of dimensionality of conditional nonparametric inference, especially to apply the nonlinear analogue of local projections. Then we extend the analysis to a special class of semi-parametric nonlinear dynamic models to circumvent the curse of dimensionality. This semi-parametric family includes, in particular, models with conditional means, conditional heteroscedasticity and non-Gaussian errors. Nevertheless, only the case of bivariate models seem feasible in practice for nonparametric nonlinear local projection and the extensions of local projections for dimension 2 and semi-parametric analyses are infinitely less precise compared to direct parametric and even semi-parametric approaches in nonlinear structural autoregressive models. \\

References

Adamek, R., Smeekes, S., and I., Wilms (2024). Local Projection Inference in High Dimension. Econometrics Journal. 27, 323-342. \\

Ait-Sahalia, Y. (1993). Nonparametric Functional Estimation with Applications to Financial Models. Doctoral Dissertation, Massachusetts Institute of Technology. \\

Ballarin, G. (2025): Impulse Response Analysis of Structural Nonlinear Time Series. DP University of St. Gallen. \\

Baqace, D. (2020). Asymmetric Inflation Expectations, Downward Rigidity of Wages, and Asymmetric Business Cycles. Journal of Monetary Economics, 114, 174-193.\\

Barnichon, R., and C., Brownlees. (2019). Impulse Response Estimation by Smooth Local Projections. Review of Economics and Statistics, 101, 522-530. \\

Bec, F., Nielsen, H., and S., Saidi. (2020). Mixed Causal–Noncausal Autoregressions: Bimodality Issues in Estimation and Unit Root Testing. Oxford Bulletin of Economics and Statistics, 82 , 1413-1428. \\

Billingsley, P. (1986). Probability and Measure, Wiley, New York. \\

Blanchard, O., and D., Quah. (1989). The Dynamic Effects of Aggregate Demand and Supply Disturbances. American Economic Review, 73, 655-673. \\

Borkovec, M., and C., Kluppelberg. (2001). The Tail of the Stationary Distribution of an Autoregressive Process with ARCH(1) Errors. Annals of Applied Probability, 11, 1220-1241. \\

Cashin, P., Mohaddes, K., and M., Raissi. (2017). Fair Weather or Foul? The Macroeconomic Effects of El Niño. Journal of International Economics, 106, 37-54. \\

Chang, P. and S., Sakata. (2007). Estimation of Impulse Response Functions Using Long Autoregressions. The Econometrics Journal, 10, 453-463. \\

Christiano, L. (2012). Christopher A. Sims and Vector Autoregressions. The Scandinavian Journal of Economics, 114, 1082-1104. \\

Christoffersen, P., Du, D., and R., Elkamhi. (2017). Rare Disasters, Credit, and Option Market Puzzles. Management Science, 63, 1341-1364.\\

Conn, D., and E., Li (2019). “An Oracle Property of the Nadaraya-Watson Kernel Estimator for High Dimensional Nonparametric Regression," Scandinavian Journal of Statistics, 46, 735-764. \\

Comon, P. (1994). Independent Component Analysis, a New Concept?. Signal Processing, 36, 287-314. \\

Cox, J., Ingersoll, J., and S., Ross. (2005). A Theory of the Term Structure of Interest Rates. Econometrica, 53, 385-407. \\

De Truchis, G., Fries, S., and A., Thomas (2024). Forecasting Extreme Trajectories Using Semi-Norm Representations. DP Paris Dauphine University. \\

Diercks, A., Hsu, A., and A., Tamoni. (2024). When it Rains it Pours: Cascading Uncertainty Shocks. Journal of Political Economy, 132(2), 694-720.\\

Dufour, J.M., and E., Renault. (1998). Short Run and Long Run Causality in Time Series: Theory. Econometrica, 66, 1099-1125. \\

Falk, M. (1985). Asymptotic Normality of the Kernel Quantile Estimator. The Annals of Statistics, 13, 428-433.\\

Frank, M., and T., Stengos. (1988). Chaotic Dynamics in Economic Time‐Series. Journal of Economic Surveys, 2, 103-133.\\

Frisch, R. (1933). Propagation Problems and Impulse Responses in Dynamic Economies. Economic Essays in Honour of Gustav Cassel. Allen and Unwin Ltd. \\

Frost, J., and R., van Stralen. (2018). Macroprudential Policy and Income Inequality. Journal of International Money and Finance, 85, 278-290.\\

Geenens, G. (2011). Curse of Dimensionality and Related Issues in Nonparametric Functional Regression. Statistics Survey, 5, 30-43. \\

Giacomini, R. (2013). The Relationship Between DSGE and VAR Models. VAR models in Macroeconomics–New Developments and Applications: Essays in Honor of Christopher A. Sims, 1-25. \\

Gallant, A., Rossi, P., and G., Tauchen. (1993). Nonlinear Dynamic Structures. Econometrica, 61, 871-907.\\

Goncalves, S., Herrera, A., Kilian, L., and E., Pesavento. (2021). Impulse Response Analysis for Structural Dynamic Models with Nonlinear Regressors. Journal of Econometrics, 225, 107-130. \\

Gonçalves, S., Herrera, A., Kilian, L., and E., Pesavento. (2024a). State-Dependent Local Projections. Journal of Econometrics, 244(2), 105702.\\

Goncalves, S., Herrera, A., Kilian, L., and E., Pesavento. (2024b). Nonparametric Local Projections. Working Paper, Federal Reserve Bank of Dallas. \\

Gospodinov, N. (2004). Asymptotic Confidence Intervals for Impulse Response Functions of Near-Integrated Processes. Econometrics Journal, 7, 505-527.\\

Gouri\'eroux, C., and J., Jasiak. (2005). Nonlinear Innovations and Impulse Responses with Application to VaR Sensitivity. Annales d'Economie et de Statistique, 78, 1-31. \\

Gouri\'eroux, C., and J., Jasiak. (2006). Multivariate Jacobi Process with Application to Smooth Transitions. Journal of Econometrics, 131(1-2), 475-505. \\

Gourieroux, C., and J., Jasiak. (2010). Local Likelihood Density Estimation and Value‐at‐Risk. Journal of Probability and Statistics, 2010(1), 754-851.\\

Gouri\'eroux, C., and J., Jasiak. (2017). Noncausal Vector Autoregressive Process: Representation, Identification and Semi-Parametric Estimation, Journal of Econometrics, 200, 118-134. \\

Gouri\'eroux, C., and J., Jasiak. (2019). Robust Analysis of the Martingale Hypothesis. Econometrics and Statistics, 9, 17-41. \\

Gouri\'eroux, C., and J., Jasiak. (2022). Nonlinear Forecasts and Impulse Responses for Causal-Noncausal (S)VAR Models. arXiv preprint arXiv:2205.09922. \\

Gouri\'eroux, C., and J., Jasiak. (2023). Generalized Covariance Estimator. Journal of Business and Economic Statistics, 41, 1315-1327. \\

Gouri\'eroux, C., Jasiak, J., and A., Monfort (2020). Stationary Bubble Equilibria in Rational Expectation Models. Journal of Econometrics, 218, 714-735. \\

Gouri\'eroux, C., Jasiak, J., and R., Sufana. (2009). The Wishart Autoregressive Process of Multivariate Stochastic Volatility. Journal of Econometrics, 150(2), 167-181.\\

Gourieroux, C., and Q., Lee. (2024). Forecast Relative Error Decomposition. arXiv:2406.17708.\\

Gourieroux, C., and Q., Lee. (2025). Identification of Impulse Response Functions for Nonlinear Dynamic Models. arXiv:2506.13531.\\

Gouri\'eroux, C., Monfort, A., Mouabbi, S., and J.P., Renne. (2021). Disastrous Defaults. Review of Finance, 25, 1727-1772.\\

Gouri\'eroux, C., Monfort, A., and J.P., Renne. (2017). Statistical Inference for Independent Component Analysis: Application to Structural VAR Models. Journal of Econometrics, 196, 111-126. \\

Hamilton, J. (2003). What is an Oil Shock?. Journal of Econometrics, 113, 363-398. \\

Hamilton, J. (2011). Nonlinearities and the Macroeconomic Effects of Oil Prices. Macroeconomic Dynamics, 15, 364-378. \\

Herbst, E. , and B., Johannsen. (2024). Bias in Local Projections. Journal of Econometrics, 240(1), 105655. \\

Heston, S. (1993). A Closed Form Solution for Options with Stochastic Volatility with Applications to Bond and Currency Options. Review of Financial Studies, 6, 327-343. \\

Hull, S., and A., White (1987). The Pricing of Options on Assets with Stochastic Volatilities. Journal of Finance, 42, 281-300.\\

Hyvarinen, A., and P., Pajonen (1999). Nonlinear Independent Component Analysis: Existence and Uniqueness Results. Neural Networks, 13, 411-430. \\

Hyvarinen, A., Sasaki, A., and R., Turner (2019): Nonlinear ICA Using Auxilary Variables and Generalized Contrastive Learning. Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, 859-868. \\

Inoue, A., Jorda, O., and G., Kuersteiner. (2025). Inference for Local Projections. The Econometrics Journal, February.\\

Jordà, Ò. (2005). Estimation and Inference of Impulse Responses by Local Projections. American Economic Review, 95, 161-182.\\

Jordà, Ò., Schularick, M., and A., Taylor. (2015). Leveraged Bubbles. Journal of Monetary Economics, 76, 1-20. \\

Karlin, S., and H., Taylor. (1981). A Second Course in Stochastic Processes. Elsevier. \\

Karlsen, M., and D., Tjostheim. (2001). Nonparametric Estimation on Null Recurrent Time Series. Annals of Statistics, 29, 372-416. \\

Kilian, L., and Y., Kim. (2011). How Reliable are Local Projection Estimators of Impulse Responses?. Review of Economics and Statistics, 93, 1460-1466. \\

Kilian, L., and H., Lutkepohl. (2017). Structural Vector Autoregressive Analysis. Cambridge University Press. \\

Kilian, L., and R., Vigfusson. (2011). Nonlinearities in the Oil Price–Output Relationship. Macroeconomic Dynamics, 15, 337-363.\\

Kilian, L., and R., Vigfusson. (2017). The Role of Oil Price Shocks in Causing US Recessions. Journal of Money, Credit and Banking, 49, 1747-1776.\\

Kolesar, M., and M., Plagb\o rg-Moller (2024). Dynamic Causal Effects in a Nonlinear World: The Good, the Bad and the Ugly. DP Princeton University. \\

Koop, G. (1996). Parameter Uncertainty and Impulse Response Analysis. Journal of Econometrics, 72, 135-149.\\

Koop, G., Pesaran, H., and S., Potter. (1996). Impulse Response Analysis in Nonlinear Multivariate Models. Journal of Econometrics, 74, 119-147. \\

Kuiper, W., and A., Lansink. (2013). Asymmetric Price Transmission in Food Supply Chains: Impulse Response Analysis by Local Projections Applied to US Broiler and Pork Prices. Agribusiness, 29, 325-343.\\

Lanne, M., and P., Saikkonen. (2011). Noncausal Autoregressions for Economic Time Series. Journal of Time Series Econometrics, 3, 1-39. \\

Lee, Q. (2025). Nonlinear Forecast Error Variance Decompositions with Hermite Polynomials. arXiv:2503.11416. \\

Ling, S. (2007). A Double AR(p) Model: Structure and Estimation. Statistica Sinica, 17, 161-175. \\

Metcalf, G., and J., Stock. (2020). Measuring the Macroeconomic Impact of Carbon Taxes. AEA papers and Proceedings, 100, 101-06. \\

Montiel Olea, J., and M., Plagborg‐Møller. (2021). Local Projection Inference is Simpler and More Robust Than You Think. Econometrica, 89, 1789-1823. \\

Montiel Olea, J., and M., Plagborg‐Møller. (2022). Corrigendum: Local Projection Inference is Simpler and More Robust Than You Think. Online Manuscript. \\

Montiel-Olea, J., Plagborg-Moller, M., Qian, E., and C., Wolf (2024). Double Robustness of Local Projections and Some Unpleasant VARithmetic. DP Princeton University. \\

Paul, P. (2020). A Macroeconomic Model with Occasional Financial Crises. Journal of Economic Dynamics and Control, 112, 103830.\\

Pesaran, H., and Y., Shin. (1998). Generalized Impulse Response Analysis in Linear Multivariate Models. Economics Letters, 58, 17-29. \\

Pesavento, E., and B., Rossi. (2007). Impulse Response Confidence Intervals for Persistent Data: What Have We Learned? Journal of Economic Dynamics and Control, 31, 2398-2412. \\

Plagborg‐Møller, M., and C., Wolf. (2021). Local Projections and VARs Estimate the Same Impulse Responses. Econometrica, 89, 955-980. \\

Ramey, V. (2016). Macroeconomic Shocks and their Propagation. Handbook of Macroeconomics, 2, 71-162. \\

Rosenblatt, M. (1952). Remarks on a Multivariate Transformation. The Annals of Mathematical Statistics, 23, 470-472.\\

Sims, C. (1980). Macroeconomics and Reality. Econometrica, 48, 1-48. \\

Toda, A. (2020). Susceptible-Infected-Recovered (SIR) Dynamics of Covid-19 and Economic Impact. arXiv Preprint. arXiv:2003.11221. \\

Tweedie, R. (1975). Sufficient Conditions for Ergodicity and Recurrence of Markov Chains on a General State Space. Stochastic Processes and their Applications, 3, 385-403. \\

Velasco, C. (2023). Identification and Estimation of Structural VARMA Models Using Higher Order Dynamics. Journal of Business & Economic Statistics, 41, 819-832.\\

Wang, X. (2019). A Long Run Risks Model with Rare Disaster: An Empirical Test in the American Consumption Data. Advances in Economics, Business and Management Research, 68, 237-243.\\

Weiss, A. (1984). ARMA Models with ARCH Errors. Journal of Time Series Analysis, 5, 129-143. \\

Wright, J. (2000). Confidence Intervals for Univariate Impulse Responses with Near Unit Root. Journal of Business and Economic Statistics, 18, 368-373. \\

Xu, K. (2023). Local Projection Based Inference Under General Conditions," DP Indiana University. \\

Zhu, H., Zhang, X., Liang, X., and Y., Li. (2017). On a Vector Double Autoregressive Model. Statistics & Probability Letters, 129, 86-95.\\