Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,844 characters · 16 sections · 80 citation commands
Identification in (Endogenously) Nonlinear SVARs Is Easier Than You Think
\global\long\def\uwrite#1#2{\underset{#2}{\underbrace{#1}} }
\global\long\def\blw#1{\ensuremath{#1}}
\global\long\def\abv#1{\ensuremath{\overline{#1}}}
\global\long\def\vect#1{\mathbf{#1}}
\global\long\def\smlseq#1{\{#1\} }
\global\long\def\seq#1{\left\{ #1\right\} }
\global\long\def\smlsetof#1#2{\{#1\mid#2\} }
\global\long\def\setof#1#2{\left\{ #1\mid#2\right\} }
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long\def\Ellp#1{\ensuremath{\mathcal{L}^{#1}}}
\global\long \global\long \global\long \global\long \global\long
\global\long\def\abs#1{\ensuremath{\left|#1\right|}}
\global\long\def\smlabs#1{\ensuremath{\lvert#1\rvert}}
\global\long\def\bigabs#1{\ensuremath{\bigl|#1\bigr|}}
\global\long\def\Bigabs#1{\ensuremath{\Bigl|#1\Bigr|}}
\global\long\def\biggabs#1{\ensuremath{\biggl|#1\biggr|}}
\global\long\def\norm#1{\ensuremath{\left\Vert #1\right\Vert }}
\global\long\def\smlnorm#1{\ensuremath{\lVert#1\rVert}}
\global\long\def\bignorm#1{\ensuremath{\bigl\|#1\bigr\|}}
\global\long\def\Bignorm#1{\ensuremath{\Bigl\|#1\Bigr\|}}
\global\long\def\biggnorm#1{\ensuremath{\biggl\|#1\biggr\|}}
\global\long\def\floor#1{\left\lfloor #1\right\rfloor } \global\long\def\smlfloor#1{\lfloor#1\rfloor}
\global\long\def\ceil#1{\left\lceil #1\right\rceil } \global\long\def\smlceil#1{\lceil#1\rceil}
\global\long \global\long \global\long \global\long \global\long \global\long\def\clsr#1{\ensuremath{\overline{#1}}}
\global\long \global\long \global\long \global\long
\global\long\def\smlinprd#1#2{\ensuremath{\langle#1,#2\rangle}}
\global\long\def\inprd#1#2{\ensuremath{\left\langle #1,#2\right\rangle }}
\global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long\def\sigf#1{\mathcal{#1}}
\global\long\global\long \global\long\def\flt#1{\mathcal{#1}}
\global\long\global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \def\independenT#1#2{\mathrel{\rlap{$#1#2$}\mkern2mu{#1#2}}}
\global\long \global\long \global\long \global\long \global\long \global\long\def\inprobu#1{\ensuremath{\overset{#1}{\ensuremath{\rightarrow}}}}
\global\long \global\long \global\long\def\inLp#1{\ensuremath{\overset{\Ellp{#1}}{\ensuremath{\rightarrow}}}}
\global\long \global\long \global\long \global\long\def\wkcu#1{\overset{#1}{\ensuremath{\rightsquigarrow}}}
\global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long\def\cv#1{\left\langle #1\right\rangle }
\global\long\def\smlcv#1{\langle#1\rangle}
\global\long\def\qv#1{\left[#1\right]}
\global\long\def\smlqv#1{[#1]}
\global\long \global\long \global\long \global\long \global\long\global\long \global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long\def\mset#1{\mathcal{#1}}
\global\long\def\largedec#1{\mathbf{#1}}
\global\long \newcommandx\Ican[1][usedefault, addprefix=\global, 1=]{I_{#1}^{\ast}}
\global\long \global\long \global\long \global\long \global\long\def\b#1{\boldsymbol{#1}}
\global\long \global\long \global\long \global\long \global\long \global\long\def\smldblangle#1{\ensuremath{\savebox{\@brx}{\(\m@th{\langle}\)} \mathopen{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}#1\savebox{\@brx}{\(\m@th{\rangle}\)} \mathclose{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}}}
\global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\def\set#1{\mathscr{#1}}
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\footnotetext[1]{Department\ of Economics and Corpus Christi College; [email removed]}
\footnotetext[2]{Department of Economics and University College; [email removed]}
\setcounter{footnote}{0}
(ref) supersedes a result that first appeared, with rather stronger assumptions, as Theorem A.1 in DM24.
\thispagestyle{plain}
\pagenumbering{roman}
\thispagestyle{plain}
\setcounter{tocdepth}{2}
{\singlespacing
}
\pagestyle{fancy}
\pagenumbering{arabic}
For more than four decades, following the seminal contribution of Sims80, the linear structural vector autoregression (SVAR)
has played a central role in empirical macroeconomics. This is a dynamic linear simultaneous equations model (SEM), in which the $p$ endogenous variables $z_{t}$ are jointly determined as a function of their past values and the $p$ (mutually uncorrelated) structural shocks $\varepsilon_{t}$. The latter are regarded as the exogenous inputs to the system, so that causality is understood to run from these shocks to current and future values of $z_{t}$, and a key object of interest is the mapping between $\varepsilon_{t}$ and $z_{t+h}$ for $h\geq0$: the (structural) impulse response function.
In this context, a fundamental result that characterises the extent to which the data is informative about the model parameters, and thus also about those impulse responses, may be phrased heuristically as follows:
In light of this, what might be termed the `SVAR identification problem' becomes one of finding sufficient additional restrictions on that matrix $Q$, so as to pin down, wholly or partially, the model parameters. The literature since has developed a variety of ways of using macroeconomic theory to generate such restrictions, based e.g.\ on the relative timing of shocks, the signs of their effects on impact, their medium- and long-run effects, and their correlation with external instruments (for a textbook discussion of which, see KL17book).
The linearity of (ref) is convenient, but inherently limiting as to the nature of the dynamics that can be modelled. In particular, it has the rather undesirable implication that the response of the economy to shocks must be the same irrespective of the phase of the business cycle: so that e.g.\ an aggregate demand shock has exactly the same effect on unemployment and inflation in the depths of a recession, when there is considerable slack in the labour market, as it would during periods of expansion. The substantial literature on nonlinear (S)VAR models has arisen partly to address these limitations (see e.g.\ Chan09; TTG10; HT13, for surveys). These allow the parameters of the SVAR at time $t$ to depend on an exogenous (or if not wholly exogenous, at least predetermined) regime-switching process $s_{t-1}$, as e.g.\ in\footnote{In many treatments of these models, the regime indicator in (ref) is denoted as $s_{t}$, rather than $s_{t-1}$. However, a feature of these models is that the regime is always determined prior to the realisation of $\varepsilon_{t}$, and may thus be regarded as measurable with respect to time-$(t-1)$ information; we have written $s_{t-1}$ to make this clearer.}
where often $s_{t}\in\{1,\ldots,L\}$ takes finitely many values, and each $\Phi_{i}(s)=\sum_{\ell=1}^{L}\pi_{\ell}(s)\Phi_{i}^{(\ell)}$ switches, or smoothly transitions, between the parameters of the $L$ `regimes'; here each $\pi_{\ell}(s)\in[0,1]$, with $\sum_{\ell=1}^{L}\pi_{\ell}(s)=1$. The evolution of $\{s_{t}\}$ may be modelled as an exogenous Markov chain (as in a Markov switching model), possibly with state-dependent transition probabilities, or as a function of certain predetermined variables (such as an element of $z_{t-i}$ for some $i\geq1$, as in a typical `threshold autoregressive' model); but in any case, $s_{t-1}$ must be determined prior to the realisation of $\varepsilon_{t}$. We therefore refer to these henceforth as exogenous regime-switching SVARs. (This characterisation applies to time-varying parameter VARs, in which $\{s_{t}\}$ is also some exogenous but possibly nonstationary process, such as a random walk.)
While models of the form (ref) enjoy greatly enriched dynamics relative to (ref), here the possibility of regime switching exacerbates the identification problem. Indeed the counterpart of (ref) for general Markov-switching models is that, conditional on the regime $s_{t-1}=s$, the parameters of (ref) are identified up to an orthogonal matrix $Q(s)$. Since this matrix may vary with $s\in\{1,\ldots,L\}$, the number of unidentified parameters, and thus the number of restrictions needed to deliver (exact) identification, scales proportionally with the number of states $L$. In practice, this may necessitate replicating a common set of $p(p-1)/2$ restrictions across all $L$ regimes (see e.g.\ RRWZ05, Sec.\ II; SZ06AER, Sec.\ III), yielding $Lp(p-1)/2$ restrictions in total. Similarly, in their two-regime STVAR model, AG12AEJ impose a Cholesky ordering on the elements of $z_{t}$ in each regime.
The exogeneity of the regime (i.e.\ of $s_{t-1}$) moreover restricts the kinds of nonlinearities that may be exhibited by the model's impulse responses. Notably, since each regime is itself a linear SVAR, the immediate effects of the shocks (i.e.\ the impact multipliers) must be linear in $\varepsilon_{t}$: which in particular rules out the possibility of sign-dependent asymmetries. It also renders (ref) unable to accommodate occasionally binding constraints, such as the zero lower bound (ZLB) constraint on short-term nominal interest rates, because the model requires the regime (whether `constrained' or `unconstrained') to be determined prior to realising the value of the potentially constrained variable -- whereas, as a matter of logic, it ought to be the value of that variable which determines whether the model is in fact in the constrained or unconstrained regime (see AMSV21).
Recently, SM21 and AMSV21 introduced the first SVAR models involving what we here refer to as endogenous regime switching, which are notably distinguished from the earlier literature on the nonlinear SVARs of the form (ref) in that they permit the autoregressive `regime' to be determined jointly with the values of the endogenous variables. For example, the `censored and kinked SVAR' (CKSVAR) of SM21 takes the form \[ \phi_{0}^{+}y_{t}^{+}+\phi_{0}^{-}y_{t}^{-}+\Phi_{0}^{x}x_{t}=c+\sum_{i=1}^{k}[\phi_{i}^{+}y_{t-i}^{+}+\phi_{i}^{-}y_{t-i}^{-}+\Phi_{i}^{x}x_{t-i}]+\varepsilon_{t}. \] where $y_{t}^{+}$ and $y_{t}^{-}$ denote the positive and negative parts of $y_{t}$ (a scalar), and $x_{t}$ is $(p-1)$-dimensional. In this model there are two contemporaneous regimes: one associated with $y_{t}>0$ (the `unconstrained' regime, in the ZLB setting), and the other with $y_{t}\leq0$ (the `constrained' regime), and in every period the model is solved simultaneously for the current values of $y_{t}$ and $x_{t}$, and for the applicable regime (as depends on the sign of $y_{t}$). Thus in situations where the $\varepsilon_{t}=0$ would entail a solution of $y_{t}=0$ (or approximately so), this allows the impact multipliers of $\varepsilon_{t}$ to be asymmetric, being dependent on which regime they push the model into.
Building on these developments, this paper proposes a new class of nonlinear SVAR models, which have the general form
where each $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$ is a continuous, possibly nonlinear function, with $f_{0}$ being invertible; we refer to these models as `endogenously nonlinear', in view of the nonlinearities on the l.h.s. Because it is not tied to any particular functional form, (ref) also offers a great deal of flexibility in its dynamics, comparable to that offered by (ref). This framework readily encompasses the CKSVAR, which corresponds to a special case in which each $f_{i}$ is piecewise linear. More general models with endogenous switching between several regimes may be straightforwardly encompassed within the framework (ref), by specifying
to be an invertible, (continuous) piecewise affine function, where $\{\set Z^{(\ell)}\}_{\ell=1}^{L}$ is a convex partition of $\mathbb{R}^{p}$, and the current regime $\ell_{t}$ corresponds to the element of that partition for which $z_{t}\in\set Z^{(\ell_{t})}$.
The principal contribution of this paper is to characterise observational equivalence in the setting of the following, slightly more general formulation of (ref),
where $\b z_{t-1}=(z_{t-1}^{\top},\ldots,z_{t-k}^{\top})^{\top}$, and $\b{f}_{1}:\mathbb{R}^{kp}\ensuremath{\rightarrow}\mathbb{R}^{p}$ need not be separable in the lags of $z_{t}$ ((ref)). Remarkably, despite the far greater flexibility afforded by the parametrisation of (ref) relative to the linear SVAR (ref), the fundamental identification result (ref) carries over to (ref) essentially unchanged. Under weak conditions on the functions $(f_{0},\b{f}_{1})$ and the distribution of the shocks, we have ((ref)):
This is a nonparametric identification result, in the sense that we do not suppose that $(f_{0},\b{f}_{1})$ have any particular (known) parametric form. While its proof draws upon the microeconometrics literature on nonlinear SEMs (see in particular Matz09Ecta,Matz15Ecta; BH18Ecta; CGHP21JPE), it constitutes a genuinely novel result within that setting. (ref) has the powerful implication that most of the existing identification results for linear SVARs apply directly to the endogenously nonlinear SVAR, since in both cases exact identification is a matter of imposing $p(p-1)/2$ restrictions sufficient to pin down $Q$.
There follows a discussion of the $L$-regime endogenous regime-switching SVAR, which arises when $f_{0}$ is specified to have the piecewise affine form (ref), and of how to verify the conditions of (ref) in this case ((ref)). (Here we also suppose, mostly to provide a practically convenient parametrisation, that the SVAR has the time-separable form (ref), with each $\{f_{i}\}_{i=1}^{k}$ also specified to have the same functional form as (ref).) In this context, our results imply that it is sufficient, for the purposes of exact identification, to impose identifying restrictions in only one of those $L$ regimes, or even to distribute these in some way across those regimes. To obtain smooth transitions between adjacent regimes, we propose to convolve $f_{0}$ with a smooth kernel. This has the considerable advantage of preserving the invertibility of $f_{0}$, whereas this may fail if one attempts to smooth $f_{0}$ by the usual device of replacing each indicator function in (ref) by a smooth, cdf-like function (as is very commonly done to produce `smooth transition' (S)VARs).
Our methodology is illustrated by estimating an endogenously regime-switching SVAR (in the log vacancy--unemployment ratio and inflation), to investigate the possibility of a nonlinear Phillips curve relationship ((ref)) that was recently proposed by BenignoEggertsson2023 to explain the recent post-pandemic inflation surge. In particular, our identification results allow us to examine the evidence for nonlinearity in a manner that is robust to alternative identification assumptions, thus shedding light on the recent debate between BenignoEggertsson2023 and BeaudryHouPortier2025.
Finally, we provide an extension of our results to the augmented model
where $(\b z_{t-1}^{(1)},\b z_{t-1}^{(2)})$ is some partitioning of $\b z_{t-1}$, and $\{v_{t}\}$ is a strictly exogenous process, in the sense of being independent of $(\b z_{0},\{\varepsilon_{t}\})$ ((ref)). Here $\sigma(\cdot)$ is a diagonal matrix (with strictly positive entries), which allows the conditional variances of the structural shocks to depend on certain predetermined variables. In this setting, we show that (ref) continues to provide a valid characterisation of the identification of $(f_{0},\b{f}_{1})$, and that $Q$ may moreover be subject to further restrictions, if there is sufficient variability in the (diagonal) entries of $\sigma(\b z_{t-1}^{(2)},v_{t-1})$; these correspond to exactly the restrictions familiar from the linear SVAR literature on `identification by heteroskedasticity'. Our main result here ((ref)) not only accommodates both: (i) ARCH-type heteroskedasticity; and (ii) the possible dependence of the r.h.s.\ of the SVAR on an exogenous process $\{v_{t}\}$; but also (iii) permits $\b{f}_{1}$ to be discontinuous in some of its arguments (specifically, $\b z_{t-1}^{(2)}$ and $v_{t-1}$).
Our point of departure is the linear SVAR, in which the observed series $\{z_{t}\}$ are regarded as being generated linearly from an underlying $p$-dimensional sequence of structural shocks $\{\varepsilon_{t}\}$, each of which have an economic interpretation (as e.g.\ an aggregate supply shock, a monetary policy shock, etc.) That is, for some $k\in\ensuremath{\mathbb{N}}$,
where $z_{t}$ and $\varepsilon_{t}$ are $\mathbb{R}^{p}$-valued, and to permit a more compact presentation we have defined $\b{\Phi}_{1}\coloneqq[\Phi_{1},\ldots,\Phi_{k}]$ and $\b z_{t-1}\coloneqq(z_{t-1}^{\top},\ldots,z_{t-k}^{\top})^{\top}$.
Observational equivalence in this setting being well understood (see e.g.\ Hamilton94, Ch. 11; Lut07, Ch. 9), our purpose here is to briefly review this in a manner that facilitates the comparison with our results for the endogenously nonlinear SVAR, which are developed in (ref) below. To simplify the problem, we suppose that $\{\varepsilon_{t}\}_{t\in\mathbb{Z}}$ is i.i.d., with a (Lebesgue) density $\varrho\in\mathscr{R}$, normalised to have $\ensuremath{\mathbb{E}}\varepsilon_{t}=0$ and $\ensuremath{\mathbb{E}}\varepsilon_{t}\varepsilon_{t}^{\top}=I_{p}$. By the Markov property the joint density of $\{z_{t}\}_{t=1}^{T}$, conditional on $\b z_{0}$, is simply the product of the conditional densities of $z_{t}\mid\b z_{t-1}$, for $t\in\{1,\ldots,T\}$. Under our assumptions, this density is time-invariant, and equals \[ \varphi_{z_{t}\mid\b z_{t-1}}(\xi\mid\b{\xi}_{-1})=\varrho(\Phi_{0}\xi-\b{\Phi}_{1}\b{\xi}_{-1})\cdot\smlabs{\det\Phi_{0}}, \] where $\xi\in\mathbb{R}^{p}$, $\b{\xi}_{-1}\in\mathbb{R}^{kp}$. Accordingly, we say that two alternative parameterisations of the linear SVAR, $(c,\Phi_{0},\b{\Phi}_{1},\varrho)$ and $(\tilde{c},\tilde{\Phi}_{0},\tilde{\b{\Phi}}_{1},\tilde{\varrho})$, are observationally equivalent if they imply identical conditional densities $\varphi_{z_{t}\mid\b z_{t-1}}$; in which case they also yield identical (conditional) likelihoods, for every possible realisation of $\{z_{t}\}$.
We then have the following well-known result, that data on $\{z_{t}\}$ identifies the SVAR coefficients $(c,\Phi_{0},\b{\Phi}_{1})$ up to, and only up to, an orthogonal transformation. Let $\mathbb{O}(p)$ denote the set of $p\times p$ orthogonal matrices.
We now seek to extend (ref) to the setting of the following (endogenously) nonlinear SVAR
where $f_{0}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$ is invertible, and $\b{f}_{1}:\mathbb{R}^{kp}\ensuremath{\rightarrow}\mathbb{R}^{p}$. (As a convenient location normalisation, we set $f_{0}(0)=0$.) This model evidently nests (ref), by taking $f_{0}(z_{t})=\Phi_{0}z_{t}$ and $\b{f}_{1}(\b z_{t-1})=c+\sum_{i=1}^{k}\Phi_{i}z_{t-i}$. Another important special case arises when $\b{f}_{1}(\b z_{t-1})=\sum_{i=1}^{k}f_{i}(z_{t-i})$ is additively time-separable, as considered in DM24. But while this separability facilitates an extension of the Granger--Johansen representation theorem to these nonlinear SVARs, it is not necessary for the results that follow. The only separability that we require here is between $z_{t}$, $\b z_{t-1}$ and $\varepsilon_{t}$.
We develop the following running example throughout the rest of the paper.
As noted in the introduction, we term (ref) an endogenously nonlinear SVAR, because of the possible nonlinearity on the l.h.s.\ of the model, i.e.\ in the endogenous variables $z_{t}$. Were the model instead required to be linear in the endogenous variables, so that $f_{0}(z)=\Phi_{0}z$, identification of the model parameters would be as straightforward as it is in the linear SVAR; and along the lines of (ref)(ref) above, the assumption that $\{\varepsilon_{t}\}$ is independent across time could be weakened to one of $\{\varepsilon_{t}\}$ being merely a martingale difference sequence (with respect to the filtration generated by $\{z_{t}\}$). However, in imposing linearity on the l.h.s., we would lose the possibilities for endogenous regime switching, asymmetric impact multipliers, and of handling occasionally binding constraints, which the general model (ref) affords. We accordingly want to permit $f_{0}$ to be nonlinear: a consequence of which is that independence across time of $\{\varepsilon_{t}\}$ becomes necessary for the parameters of (ref) to be identified. (But note that there is no requirement of contemporaneous independence between the elements of $\varepsilon_{t}$.)
We thus continue to maintain that the structural shocks $\{\varepsilon_{t}\}$ are i.i.d.\ with $\ensuremath{\mathbb{E}}\varepsilon_{t}=0$ and $\ensuremath{\mathbb{E}}\varepsilon_{t}\varepsilon_{t}^{\top}=I_{p}$, and a (Lebesgue) density $\varrho\in\mathscr{R}$. The parameter space for the model (ref) then consists of collections $\mathscr{F}_{0}\nif_{0}$ and $\pmb{\mathscr{F}}_{1}\ni\b{f}_{1}$ of functions $\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$ and $\mathbb{R}^{kp}\ensuremath{\rightarrow}\mathbb{R}^{p}$, and a collection of densities $\mathscr{R}\ni\varrho$ supported on $\mathbb{R}^{p}$, which under our regularity conditions, together determine the conditional density
We continue to regard two alternative parametrisations of the model as being observationally equivalent if they yield the same conditional density. For convenience, we shall suppose throughout that $\b z_{0}$ is continuously distributed, with a density that is a.e.\ strictly positive on $\mathbb{R}^{kp}$. Our assumptions below then ensure that this is also true for every successive $\b z_{t}$, and $\varphi_{z_{t}\mid\b z_{t-1}}(\xi\mid\b{\xi}_{-1})$ is thus well defined for almost every $\xi\in\mathbb{R}^{p}$ and $\b{\xi}_{-1}\in\mathbb{R}^{kp}$, for all $t\geq1$.
Our regularity conditions on the model parameter space, which are sufficient to ensure that the conditional density (ref) exists (and is unique up to the usual a.e.\ equivalence), are as follows.
\needspace{4\baselineskip}
{{{{\scalefont{0.76}PS}}}}
Regarding the parameters $(f_{0},\b f_{1},\varrho)$ that generated $\{z_{t}\}$ in (ref), as distinct from the entirety of the model parameter space, we also maintain the following.
{{{{\scalefont{0.76}DGP}}}}
Remarkably, despite the far greater flexibility afforded by the nonlinear parametrisation of (ref), under the foregoing conditions we obtain the following, effectively identical characterisation of observational equivalence to that of the linear SVAR (ref), the proof of which appears in (ref).
Analogously (though not identically) to the `orthogonal reduced-form parametrisation' (ARRW18Ecta, Sec.\ 2.3) of the linear SVAR, (ref) suggests the following convenient reparametrisation of the endogenously nonlinear SVAR. Let $z_{0}\in\mathbb{R}^{p}$ be fixed, and a point at which $f_{0}$ is (assumed to be) differentiable, with full rank Jacobian $Df_{0}(z_{0})$. By the QR decomposition, we have $Df_{0}(z_{0})=Q^{\top}L$, where $L$ is lower triangular, and $Q\in\mathbb{O}(p)$; define $(g_{0},\b{g}_{1})\coloneqq(Qf_{0},Q\b{f}_{1})$. Multiplying (ref) through by $Q$, we may reformulate the model as
where now $g_{0}$ is restricted such that $Dg_{0}(z_{0})$ is lower triangular (which need hold only at that chosen $z_{0}$), and $Q\in\mathbb{O}(p)$.
This yields an equivalent parametrisation of the model, in which the parameter spaces for $\b{g}_{1}\in\b{\mathscr{F}}_{1}$ and $\varrho\in\mathscr{R}$ remain as before, but now $\mathscr{F}_{0}$ is additionally restricted (beyond (ref){{{\scalefont{0.76}.}}}(ref){{{\scalefont{0.76}--}}}(ref)) to functions $g_{0}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$ for which $Dg_{0}(z_{0})$ is lower triangular (at the nominated $z_{0}\in\mathbb{R}^{p}$); let $\mathscr{F}_{0}^{(z_{0})}$ denote the resulting parameter space for $g_{0}$. To exactly offset this restriction, we now have the additional parameter $Q\in\mathbb{O}(p)$, so that we may equivalently regard the nonlinear SVAR as being parametrised by $(g_{0},\b{g}_{1},Q,\varrho)\in\mathscr{F}_{0}^{(z_{0})}\times\pmb{\mathscr{F}}_{1}\times\mathbb{O}(p)\times\mathscr{R}$. The import of (ref) here is that the parameters $(g_{0},\b{g}_{1})\in\mathscr{F}_{0}^{(z_{0})}\times\pmb{\mathscr{F}}_{1}$ are exactly identified by data on $\{z_{t}\}$, with the non-identified part of the structural parameters being transferred entirely to $Q$. The `nonlinear SVAR identification problem' can thus be framed precisely as one of finding sufficient restrictions to pin down $Q$, from which the structural parameters may then be recovered, via $(f_{0},\b{f}_{1})=(Q^{\top}g_{0},Q^{\top}\b{g}_{1})$.
The reparametrisation (ref) provides a convenient perspective from which to import various approaches to identifying $Q$ from the linear SVAR literature. For the most part, these apply directly to the present setting, with little modification required. The following example illustrates how it remains possible to identify impulse responses via external instruments, without requiring any additional assumptions relative to those needed to identify the linear VAR.
Here we introduce a class of endogenously regime-switching models, in which the conditions required for our results may be verified relatively straightforwardly. Models of this form have been used recently to study monetary policy under an occasionally binding constraint on nominal interest rates: see SM21, AMSV21, ILMZ20, and CCMM25.
Suppose now that the l.h.s.\ of the nonlinear SVAR
is specified as
for $\{\set Z^{(\ell)}\}_{\ell=1}^{L}$ a collection of convex sets that partition $\mathbb{R}^{p}$, $\{\bar{\phi}_{0}^{(\ell)}\}_{\ell=1}^{L}\subset\mathbb{R}^{p}$ and $\{\Phi_{0}^{(\ell)}\}_{\ell=1}^{L}\subset\mathbb{R}^{p\times p}$. When these parameters are such that $f_{0}$ is continuous, we shall say that $f_{0}$ is a piecewise affine function. (We do not consider cases in which $f_{0}$ may be discontinuous, so continuity should always be taken as implied.) The model may then be regarded as consisting of $L$ `regimes' demarcated by the sets $\{\set Z^{(\ell)}\}_{\ell=1}^{L}$. Which of those regimes is operative in period $t$, i.e.\ the value of $\ell_{t}\in\{1,\ldots,L\}$ such that \[ f_{0}(z_{t})=\bar{\phi}_{0}^{(\ell_{t})}+\Phi_{0}^{(\ell_{t})}z_{t} \] is determined jointly with the value of $z_{t}$. For this reason, we say that there is endogenous switching between the $L$ regimes, as distinct from the exogenous regime switching that would result if $\ell_{t}$ were determined prior to the realisation of $z_{t}$. The situation here is thus markedly different from the regime-switching SVARs considered in the previous literature, which as noted in the introduction, can generally be written in the form \[ \Phi_{0}(s_{t-1})z_{t}=c(s_{t-1})+\sum_{i=1}^{k}\Phi_{i}(s_{t-1})z_{t-i}+\varepsilon_{t}, \] where $s_{t-1}$ is determined prior to $\varepsilon_{t}$ and $z_{t}$ (see e.g.\ AG12AEJ; CCCN15EJ; BP24EctJ).
\setcounter{savedexnumber}{\value{examplex}}
{{\ref*{exa:phillips}}}
\setcounter{examplex}{\value{savedexnumber}} \numberwithin{examplex}{section}
We would like primitive conditions that ensure $f_{0}$ in (ref) satisfies the requirements (ref){{{\scalefont{0.76}.}}}(ref){{{\scalefont{0.76}-}}}(ref) and (ref){{{\scalefont{0.76}.}}}(ref) of (ref): namely, that both it and its inverse should be locally Lipschitz, and that it should be (globally) invertible, with $\det Df_{0}(z)\neq0$ a.e. Two important special cases of (ref), for which these conditions may be readily verified, are those of:
Because the boundaries between the regimes are then affine subspaces (of $\mathbb{R}^{p}$), ensuring the continuity of $f_{0}$ is a straightforward matter of linearly restricting the elements of $\{\bar{\phi}_{0}^{(\ell)}\}_{\ell=1}^{L}$ and $\{\Phi_{0}^{(\ell)}\}_{\ell=1}^{L}$ such that the values prescribed by adjacent regimes agree on those boundaries; see the next appearance of (ref) for an illustration. Regarding our remaining requirements on $f_{0}$, for these it is necessary and sufficient that
See (ref) below; we note that equivalence of the preceding with the invertibility of $f_{0}$ follows directly from Theorems 1 and 4 of GLM80Ecta, and that since $Df_{0}(z)=\sum_{\ell=1}^{L}\ensuremath{\mathbf{1}}\{z\in\set Z^{(\ell)}\}\Phi_{0}^{(\ell)}$ a.e., the Jacobian is then clearly invertible a.e.
\setcounter{savedexnumber}{\value{examplex}}
{{\ref*{exa:phillips}}}
\setcounter{examplex}{\value{savedexnumber}} \numberwithin{examplex}{section}
The SVAR specification (ref)--(ref) thus provides a flexible but tractable means of introducing nonlinearity into an SVAR model. This is especially the case if we also specify that the r.h.s.\ should be additively time-separable, and of the same functional form as the l.h.s., so that
where now, for every $i\in\{0,\ldots,k\}$,
is (continuous) piecewise affine. (Note that there is no need to additionally index the regimes $\set Z^{(\ell)}$ by $i$ here, since if the partitions $\{\set Z_{i}^{(\ell)}\}_{\ell=1}^{L_{i}}$ did vary across $i$, we could always find a mutual refinement such that (ref) held for all $i$.) We term this model a piecewise affine SVAR; with piecewise linear and threshold affine SVARs corresponding to those cases where the $f_{i}$'s are either all piecewise linear or all threshold affine functions, respectively.
The conditions (ref){{{\scalefont{0.76}.}}}(ref) and (ref){{{\scalefont{0.76}.}}}(ref) imposed by (ref) on $\b{f}_{1}(\b z_{t-1})=\sum_{i=1}^{k}f_{i}(z_{t-i})$ are rather less taxing than those imposed on $f_{0}$. Under the specification (ref), continuity is readily imposed, and then automatically implies Lipschitz continuity. Moreover, $D\b f_{1}(\b z_{t-1})$ a.e.\ exists and satisfies \[ D\b{f}_{1}(\b z)=
\] for some $\ell_{i}\in\{1,\ldots,L\}$ depending on $\b z$, and so it is easy to verify whether $\operatorname{rk} D\b{f}_{1}(\b z)=p$ a.e.\ (or, since this holds generically, to test the null hypothesis of a deficient rank). In practice, this may be analysed more straightforwardly on the basis of the coefficients on the first lag or two of $z_{t}$, which may themselves be sufficient to satisfy this condition.
Verifying the (global) surjectivity condition on $\b{f}_{1}$ is a little more challenging, because of the apparent absence of a counterpart to (ref) for this case. In the special case of a model with only one lag, surjectivity of $z_{t-1}\ensuremath{\mapsto}f_{1}(z_{t-1})$ is equivalent to (ref). Though easy to check, this is far more than is necessary for surjectivity when additional lags are present. Alternatively, if some elements of $z_{t}$ enter $f_{i}$ linearly, as will often be the case in practice (as in our next example), then surjectivity holds so long as the coefficient vectors associated with (at least) $p$ of these variables (drawn from across the $k$ lags of $z_{t}$ appearing on the r.h.s.)\ form a rank $p$ matrix.
The piecewise affine SVAR (ref)--(ref) may be extended to allow for `smooth transitions' between the $L$ regimes. In the literature on smooth transition (vector) autoregressive models, the conventional approach (e.g.\ HT13) is to replace the indicator functions $\ensuremath{\mathbf{1}}\{z\in\set Z^{(\ell)}\}$ in (ref) by smooth maps $\pi^{(\ell)}(z)$, so that now \[ f_{i}^{\mathrm{ST}}(z)=\sum_{\ell=1}^{L}\pi^{(\ell)}(z)(\bar{\phi}_{i}^{(\ell)}+\Phi_{i}^{(\ell)}z), \] where $\pi^{(\ell)}(z)\in[0,1]$ and $\sum_{\ell=1}^{L}\pi^{(\ell)}(z)=1$ for all $z\in\mathbb{R}^{p}$, so that $f_{i}^{\mathrm{ST}}(z)$ is always a smooth, convex combination of the affine functions $z\ensuremath{\mapsto}\bar{\phi}_{i}^{(\ell)}+\Phi_{i}^{(\ell)}z$, for $\ell\in\{1,\ldots,L\}$. However, the fact that the gradient of $f_{i}^{\mathrm{ST}}$ is not a convex combination of those underlying affine regimes makes it difficult to reduce the high-level conditions of (ref) to a set of verifiable conditions on the underlying regime-specific coefficient matrices, in the manner of (ref). Indeed, as the simple example in (ref) illustrates, it may well be the case that $f_{0}^{\mathrm{ST}}$ is not invertible, even though its unsmoothed counterpart $f_{0}$ is.
As an alternative specification that allows for smooth transitions between regimes, but which also retains the simplicity -- in terms of verifying the conditions for (ref) -- enjoyed by piecewise affine models, consider
where $f_{i}$ is a (continuous) piecewise affine function as in (ref) above, and $K$ is a smooth (kernel) density function with mean zero, with $m\geq1$ continuous derivatives that satisfy the integrability condition
where $\partial_{u_{i}}$ denotes the partial derivative with respect to the $i$th element of $u\in\mathbb{R}^{p}$, for $\alpha_{i}\in\{1,\ldots,p\}$ and $1\leq n\leq m$.
Our next result establishes that $f_{0,K}(z)$ is smooth (with as many continuous derivatives as $K$ has), and moreover invertible if the determinantal condition (ref) is satisfied (its proof appears in (ref)). Recall that a function is said to be bi-Lipschitz if both it and its inverse are Lipschitz continuous.
The inflation surge that followed the COVID-19 pandemic reignited academic interest in the possibility of nonlinearity in the transmission of supply shocks to inflation. However, views on the relevance of nonlinearity are divided. On the one hand, BallLeighMishra2022 and BenignoEggertsson2023 find evidence of significant nonlinearity in their formulations of the Phillips curve, and argue that nonlinearity is needed to account for the recent inflation surge. On the other hand, BeaudryHouPortier2025 caution that the evidence on nonlinearity is not robust to functional form assumptions, especially as pertains to the treatment of expectations. Reconsidering this debate, through the lens of an endogenous regime-switching SVAR, provides an illustrative application of the methodology developed in this paper.
Our identification result can be useful in this debate because it highlights the following: since all observationally equivalent structures are identified up to a (linear) orthogonal transformation, then if one finds no (statistically significant) evidence of nonlinearity under one specific identification scheme, then this will remain true irrespective of how the model is identified. Indeed, one can see from the orthogonal reduced-form parametrisation developed in (ref) above, that the structural parameters $f_{0}$ (and $\b{f}_{1}$) will be nonlinear if and only if their normalised (and exactly identified) counterparts $g_{0}$ (and $\b{g}_{1}$) are also nonlinear, as will be the case for $Qf_{0}$ (and $Q\b{f}_{1}$) for any $Q\in\mathbb{O}(p)$. Thus the presence of nonlinearity can be tested for in a way that is robust to the identifying scheme employed. To be clear, this is a consequence of modelling the joint determination of $z_{t}=(\log\theta_{t},\pi_{t})^{\top}$ in its entirety; the argument does not carry over to the methodology employed in the aforementioned papers, because these provide only a single-equation analysis of the Phillips curve, and so their findings are potentially contingent on the assumptions made in order to identify that equation.
Building on the development already given to this point in (ref), inspired by the recent work of BenignoEggertsson2023 we consider the following endogenous regime-switching SVAR for $z_{t}=(\log\theta_{t},\pi_{t})^{\top}$,
where $\theta_{t}=v_{t}/u_{t}$ is the vacancy--unemployment ratio, $\pi_{t}$ is consumer price inflation, $\varepsilon_{t}$ are the structural shocks, and
where $z_{1t}=\log\theta_{t}$. This model thus has two regimes, determined by the sign of $z_{1t}$. Following the arguments that led to (ref) above, to ensure continuity of the model in both $z_{t}$ and its lags, we parametrise the regime-dependent coefficient matrices non-redundantly as
so that only the coefficients of the regime-determining variable, $z_{1t}$, are permitted to vary across the two regimes. The model is then guaranteed to yield a solution for $z_{t}$, for every possible value of the r.h.s.\ of (ref), provided that $\det\Phi_{0}^{(1)}\cdot\det\Phi_{0}^{(2)}>0$.
To obtain a just-identified specification, by (ref) it suffices to impose $p(p-1)/2=1$ restrictions on the model parameters (see also the discussion in (ref) above). For some identifying schemes, this may involve imposing a restriction on only one of the two regimes. However, the identifying assumption in BenignoEggertsson2023 corresponds to the `recursive' or `Cholesky' restriction under which (a shock to) inflation $\pi_{t}=z_{2t}$ has no contemporaneous effect on tightness $\log\theta_{t}=z_{1t}$, and thus that the matrix $\Phi_{0}^{(\ell)}$ is lower triangular for $\ell\in\{1,2\}$. In view of (ref), this in fact constitutes only a single restriction on the model parameters, that $\Phi_{0,12}=0$, and so is exactly identifying rather than over-identifying. The second equation of the nonlinear SVAR (ref) can in this case be estimated by nonlinear regression (with $\pi_{t}$ as the dependent variable), as was done by BenignoEggertsson2023.
Let $\{\Gamma_{i}^{(\ell)}\}$ momentarily denote the SVAR parameters corresponding to the recursive identification scheme of BenignoEggertsson2023. In light of (ref), because of the lower-triangular structure imposed on the Jacobian $\Phi_{0}^{(1)}$ of $f_{0}$ (at some nominated point $z_{0}$ in the `normal' regime), these are the coefficients associated with the orthogonal reduced-form parametrisation (ref) of the SVAR, when the $g_{j}$ are modelled as piecewise linear. (ref), together with a sign-normalisation of the shocks, then implies that all observationally equivalent models can be obtained by a common rotation of the recursively identified model, i.e.\ $\Phi_{i}^{(\ell)}=Q\Gamma_{i}^{(\ell)}$ for $\ell\in\{1,2\}$ and $i\in\{0,\ldots k\}$, where $Q\in\mathbb{O}(p)$ with $\det Q>0$.
Because $Q$ is not regime dependent, every observationally equivalent parametrisation of the model obtained in this way will exhibit regime dependence if, and only if, this is also true of the parameters $\{\Gamma_{i}^{(\ell)}\}$ obtained under the BenignoEggertsson2023 identification scheme. The presence of some regime dependence in $\{\Gamma_{i}^{(\ell)}\}$ is thus a necessary condition for the existence of a nonlinear Phillips curve under any identification scheme. Since the null hypothesis of no regime dependence in $\{\Gamma_{i}^{(\ell)}\}$ is testable, a failure to reject it would provide evidence, in favour of a linear Phillips curve, that is robust to all possible identifying schemes. (In this respect, our imposition of the BenignoEggertsson2023, restrictions merely provides a convenient way to normalise the system, in the manner of (ref)).
Observe that the specification of (ref) allows for nonlinearities in all lags of the SVAR. This permits the dynamic response of inflation to labour market tightness shocks to be nonlinear, even if the impact responses are linear, i.e., even if $\Phi_{0}^{(\ell)}$ is regime-invariant. We therefore consider two separate tests of linearity. The first tests
The null hypothesis $H_{0}^{\mathrm{NS}}$ can be interpreted as saying that there is no endogenous regime switching, and implies that the impact effect of labour tightness shocks on inflation does not depend on the state of the labour market.
However, $H_{0}^{\mathrm{NS}}$ does not exclude the possibility that $\Phi_{i}^{(1)}\neq\Phi_{i}^{(2)}$ for some $i\in\{1,\ldots,k\}$, in which case the dynamic effects of tightness shocks may still be regime dependent, at longer horizons. This motivates our second, more restrictive hypothesis:
which under the null entails a linear SVAR. Failure to reject $H_{0}^{\mathrm{lin}}$ would suggest that a linear SVAR provides an adequate description of the dynamic causal effects (modulo the usual invertibility caveats), and thus that the Phillips curve is linear, in a very strong sense, under any identification scheme.
We use the data from the 2025 version of BenignoEggertsson2023, available on the authors' websites. Specifically, inflation $\pi_{t}$ is the quarterly, annualised core CPI inflation (excluding food and energy), constructed from monthly CPI data averaged to quarterly frequency and sourced from the BLS via FRED. The vacancy-to-unemployment ratio $\theta_{t}=v_{t}/u_{t}$ is the ratio of job vacancies to unemployed workers, using the Barnichon2010 vacancy series (as updated by the author), also averaged from a monthly to a quarterly frequency. We estimate the piecewise linear SVAR (ref) with two lags ($k=2$) over the sample periods 1960Q1-2024Q4 and 2008Q1-2024Q4, to mirror the analysis of BenignoEggertsson2023.
(ref) reports likelihood ratio (LR) tests of our two linearity hypotheses: $H_{0}^{\mathrm{NS}}$ (no endogenous regime switching) and $H_{0}^{\mathrm{lin}}$ (fully linear SVAR). The results clearly reject the linearity hypothesis, both in its weak and strong forms. The apparent deterioration in fit of the linear models is even stronger in the shorter, more recent, sample.
Even though failure to reject would have been conclusive evidence against nonlinearity, these results are not enough to conclude that the Phillips curve itself, being only one equation in our bivariate system, is nonlinear. They imply that impulse responses to identified structural inflation and tightness shocks will be significantly state-dependent under any identification scheme, but it remains to be seen what this state dependence looks like in the Phillips curve that emerges from any specific identification scheme. We turn to this question next.
Further evidence on the nonlinearity of the Phillips curve is obtained by computing estimates of its slope under both regimes. We do this in two different ways. First, we produce a kinked Phillips curve plot (equivalent to Figure 6(b) of BenignoEggertsson2023). This is shown in the left panel of (ref). The scatterplot shows inflation after removing the effects of all explanatory variables from the supply equation in model (ref) except $\log\theta_{t}$. The solid lines trace out the estimated Phillips curve in the $(\log\theta_{t},\pi_{t})$ space. In particular, the slope coefficient under each regime is computed as $-\Phi_{0,21}^{\left(\ell\right)}/\Phi_{0,22}$, which is given by the equation in the bottom row of (ref), solved for $z_{2,t}=\pi_{t}$, and using the fact that $\Phi_{0,22}$ is regime-independent, as per (ref). For the 2008Q1--2024Q4 sample, the estimated slopes are $\hat{\beta}^{(1)}=3.82$ ($\log\theta_{t}\leq0$ regime) and $\hat{\beta}^{(2)}=16.92$ ($\log\theta_{t}>0$ regime).
The right panel of (ref) shows a dynamic Phillips curve multiplier under each regime, computed from the state-dependent IRFs. We choose two starting points that are representative of the two regimes: 2009Q3 ($\log\theta_{t}=-1.84$, the Great Recession trough) for the $\log\theta_{t}\leq0$ regime, and 2022Q2 ($\log\theta_{t}=0.68$, the post-COVID peak) for the $\log\theta_{t}>0$ regime. The multiplier is the ratio of the cumulative inflation IRF (at horizon $h$) to the cumulative tightness IRF following a market tightness shock which raises $\log\theta_{t}$ by 1 unit over the next $h$ periods: \[ \text{Slope}_{h}^{PC}=\frac{\sum_{s=0}^{h}\frac{\partial\pi_{t+s}}{\partial\varepsilon_{\theta,t}}}{\sum_{s=0}^{h}\frac{\partial\log\theta_{t+s}}{\partial\varepsilon_{\theta,t}}}. \]
Both approaches show a substantially steeper Phillips curve in the tight labour market regime $\log\theta_{t}>0$ compared to the loose labour market regime $\log\theta_{t}\leq0$. The results are qualitatively and quantitatively consistent with BenignoEggertsson2023, which is not surprising given that we used the same identifying assumption as them.
The appearance of a nonlinear transformation on the l.h.s.\ of the (endogenously) nonlinear SVAR
entails that the model automatically accommodates certain forms of regime-dependent heteroskedasticity. This can be readily seen, for example, when $f_{0}$ has the piecewise linear form \[ f_{0}(z_{t})=\sum_{\ell=1}^{L}\ensuremath{\mathbf{1}}\{z_{t}\in\set Z^{(\ell)}\}\Phi_{0}^{(\ell)}z_{t}. \] In this case, whenever the r.h.s.\ of the model is such that $z_{t}\in\set Z^{(\ell_{t})}$, the model behaves locally like a linear SVAR, with reduced form \[ z_{t}=(\Phi_{0}^{(\ell_{t})})^{-1}\b{f}_{1}(\boldsymbol{z}_{t-1})+(\Phi_{0}^{(\ell_{t})})^{-1}\varepsilon_{t}, \] for all $\varepsilon_{t}$ such that $z_{t}$ continues to lie in $\set Z^{(\ell_{t})}$. (Note that unlike in a model with exogenous regimes, $\ell_{t}$ depends on $\varepsilon_{t}$, and so the preceding does not hold for all $\varepsilon_{t}$.)
Nonetheless, in some situations it may be desirable to augment the model to allow for ARCH-type conditional heteroskedasticity, in which the variances of the structural shocks depend on certain (observed) predetermined variables. To that end, consider the following extension of (ref), to
where $\b z_{t-1}^{(1)}$ and $\b z_{t-1}^{(2)}$ partition (into vectors of dimension $d_{(1)}+d_{(2)}=kp$) the elements of $\b z_{t-1}$, while $\{v_{t}\}$ is strictly exogenous in the sense of being independent of $(\b z_{0},\{\varepsilon_{t}\})$, and takes values in the (possibly discrete) set $\mathcal{V}\subset\mathbb{R}^{d_{v}}$. (Rather than requiring $\{v_{t}\}$ to be stationary, we suppose that there is a measure $\mu_{v}$ on $\mathcal{V}$ to which the distribution of $v_{t}$ is equivalent, for every $t\geq0$.)
The skedastic function, $\sigma(\cdot)$, allows the volatilities of the structural shocks \[ w_{t}\coloneqq\sigma(\boldsymbol{z}_{t-1}^{(2)},v_{t-1})\varepsilon_{t} \] to depend on $(\boldsymbol{z}_{t-1}^{(2)},v_{t-1})$; we require $\sigma(\cdot)$ to be a diagonal matrix (with strictly positive entries), so that the structural shocks $w_{t}$ remain mutually uncorrelated (cf.\ Section 14.2 of KL17book). By introducing $\{v_{t}\}$, we also extend the model so as to permit the r.h.s.\ to depend on processes that are exogenous to the SVAR (such as deterministic processes). We continue to maintain that $\{\varepsilon_{t}\}$ is i.i.d.\ with mean zero and variance $I_{G}$, and moreover that $\varepsilon_{t+1}$ is independent of $(\b z_{0},\{\varepsilon_{s},v_{s}\}_{s\leq t})$, for all $t\geq0$.
Under the assumptions given below, the augmented model (ref) yields the following (time-invariant) density for $z_{t}$ conditional on $(\b z_{t-1},v_{t-1})$, \[ \varphi_{z_{t}\mid\b z_{t-1},v_{t-1}}(\xi\mid\b{\xi}_{-1},\upsilon)=\varrho\{\sigma(\b{\xi}_{-1}^{(2)},\upsilon)^{-1}[f_{0}(\xi)-\b{f}_{1}(\b{\xi}_{-1}^{(1)},\b{\xi}_{-1}^{(2)},\upsilon)]\}\cdot\smlabs{\det Df_{0}(\xi)}, \] where $\b{\xi}_{-1}\in\mathbb{R}^{kp}$ is partitioned into $(\b{\xi}_{-1}^{(1)},\b{\xi}_{-1}^{(2)})$ conformably with that of $\b z_{t-1}$ into $(\b z_{t-1}^{(1)},\b z_{t-1}^{(2)})$. Since the likelihood for $\{z_{t}\}_{t=1}^{n}$ conditional on $(\b z_{0},\{v_{t}\}_{t=0}^{n-1})$ can be expressed entirely in terms of these conditional densities, we continue to regard two alternative parametrisations as being observationally equivalent if they yield the same $\varphi_{z_{t}\mid\b z_{t-1},v_{t-1}}$ (up to the usual a.e.\ equivalences), similarly to (ref) above.
The parameters $(f_{0},\b{f}_{1},\sigma)$ of the model (ref) are, in a quite trivial sense, indistinguishable from $(\Lambdaf_{0},\Lambda\b{f}_{1},\Lambda\sigma)$, if $\Lambda$ is a diagonal matrix with strictly positive entries. Such a rescaling has no effect on the (scale-normalised) impulse responses implied by the model parameters, and is merely a consequence of the lack of a scale normalisation in (ref) -- something that was previously delivered, in the context of (ref), by the requirement that $\ensuremath{\mathbb{E}}\varepsilon_{t}\varepsilon_{t}^{\top}=I_{p}$. Letting $\set S\ni\sigma$ denote the parameter space for the skedastic function, we may fix the overall scale of the model by requiring every $\tilde{\sigma}\in\set S$ to satisfy
at some (user specified) value of $(\b z^{(2)\ast},v^{\ast})\in\mathbb{R}^{d_{(2)}}\times\mathcal{V}$. (To prevent this from being satisfied simply by a modification of $\tilde{\sigma}$ on a null set, we further maintain that $\tilde{\sigma}$ is continuous at $(\b z^{(2)\ast},v^{\ast})$, and that $\mu_{v}$ has strictly positive measure in every neighbourhood of $v^{\ast}$.)
Here we shall also relax the requirement that $\b f_{1}$ be continuous in all of its arguments: in fact we only require continuity of $\b z^{(1)}\ensuremath{\mapsto}\b{f}_{1}(\boldsymbol{z}^{(1)},\boldsymbol{z}^{(2)},v)$, at the cost of a strengthening of the surjectivity condition given in (ref){{{\scalefont{0.76}.}}}(ref) above. This reflects the crucial role that the variables $\boldsymbol{z}_{t-1}^{(1)}$, which are excluded from the skedastic function, now play in delivering the identification of the model parameters.
{{{{\scalefont{0.76}EXT}}}}
We may thus state the main result of this section, which extends (ref) above by allowing for: (i) ARCH-type heteroskedasticity; (ii) dependence of the r.h.s.\ of the model on an exogenous process $\{v_{t}\}$, and (iii) $\b{f}_{1}$ to be discontinuous in some arguments.
Since the skedastic function must be a diagonal matrix, (ref) may provide further restrictions on $Q$; the extent of these will depend on the properties of the actual skedastic function $\sigma$. On the one hand, suppose that $\sigma(\boldsymbol{z}^{(2)},v)=\lambda(\boldsymbol{z}^{(2)},v)I_{p}$ is always a rescaling of the identity matrix. Then (ref) yields a diagonal matrix for every $Q\in\mathbb{O}(p)$, and no further restrictions on $Q$ are implied. On the other hand, suppose that $\sigma(\boldsymbol{z}^{(2)},v)$ varies in such a way that it is not always proportional to the identity matrix, so that the variances of some of the structural shocks may differ from each other, at least for certain values of $(\boldsymbol{z}^{(2)},v)$. In particular, if there exists a $(\boldsymbol{z}^{(2)\dagger},v^{\dagger})\in\mathbb{R}^{d_{(2)}}\times\mathcal{V}$ such that all the (diagonal) entries of $\sigma(\boldsymbol{z}^{(2)\dagger},v^{\dagger})$ are distinct, then $Q$ must be a signed permutation matrix (as follows from Theorem 2.5.4 in HJ13book; cf.\ Proposition 1 in LLM10JEDC), in which case the structural impulse response functions are identified, up to a signing and economic `labelling' of the shocks. In this way, we here obtain exactly the same kinds of restrictions that are familiar from the linear SVAR literature on `identification by heteroskedasticity' (see e.g.\ the discussion in Sections 14.2--14.3 of KL17book).
{\singlespacing
}