Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
526,155 characters · 60 sections · 192 citation commands
Inference on Common Trends in a Cointegrated Nonlinear SVAR
\global\long\def\uwrite#1#2{\underset{#2}{\underbrace{#1}} }
\global\long\def\blw#1{\ensuremath{#1}}
\global\long\def\abv#1{\ensuremath{\overline{#1}}}
\global\long\def\vect#1{\mathbf{#1}}
\global\long\def\smlseq#1{\{#1\} }
\global\long\def\seq#1{\left\{ #1\right\} }
\global\long\def\smlsetof#1#2{\{#1\mid#2\} }
\global\long\def\setof#1#2{\left\{ #1\mid#2\right\} }
\global\long\def\ensuremath{\rightarrow}{\ensuremath{\rightarrow}}
\global\long\def\ensuremath{\nrightarrow}{\ensuremath{\nrightarrow}}
\global\long\def\ensuremath{\uparrow}{\ensuremath{\uparrow}}
\global\long\def\ensuremath{\downarrow}{\ensuremath{\downarrow}}
\global\long\def\ensuremath{\upuparrows}{\ensuremath{\upuparrows}}
\global\long\def\ensuremath{\downdownarrows}{\ensuremath{\downdownarrows}}
\global\long\def\ensuremath{\nearrow}{\ensuremath{\nearrow}}
\global\long\def\ensuremath{\searrow}{\ensuremath{\searrow}}
\global\long\def\ensuremath{\rightarrow}{\ensuremath{\rightarrow}}
\global\long\def\ensuremath{\mapsto}{\ensuremath{\mapsto}}
\global\long\def\ensuremath{\circ}{\ensuremath{\circ}}
\global\long\defC{C}
\global\long\defD{D}
\global\long\def\Ellp#1{\ensuremath{\mathcal{L}^{#1}}}
\global\long\def\ensuremath{\mathbb{N}}{\ensuremath{\mathbb{N}}}
\global\long\def\mathbb{R}{\mathbb{R}}
\global\long\def\mathbb{C}{\mathbb{C}}
\global\long\def\mathbb{Q}{\mathbb{Q}}
\global\long\def\mathbb{Z}{\mathbb{Z}}
\global\long\def\abs#1{\ensuremath{\left|#1\right|}}
\global\long\def\smlabs#1{\ensuremath{\lvert#1\rvert}}
\global\long\def\bigabs#1{\ensuremath{\bigl|#1\bigr|}}
\global\long\def\Bigabs#1{\ensuremath{\Bigl|#1\Bigr|}}
\global\long\def\biggabs#1{\ensuremath{\biggl|#1\biggr|}}
\global\long\def\norm#1{\ensuremath{\left\Vert #1\right\Vert }}
\global\long\def\smlnorm#1{\ensuremath{\lVert#1\rVert}}
\global\long\def\bignorm#1{\ensuremath{\bigl\|#1\bigr\|}}
\global\long\def\Bignorm#1{\ensuremath{\Bigl\|#1\Bigr\|}}
\global\long\def\biggnorm#1{\ensuremath{\biggl\|#1\biggr\|}}
\global\long\def\floor#1{\left\lfloor #1\right\rfloor } \global\long\def\smlfloor#1{\lfloor#1\rfloor}
\global\long\def\ceil#1{\left\lceil #1\right\rceil } \global\long\def\smlceil#1{\lceil#1\rceil}
\global\long\def\ensuremath{\bigcup}{\ensuremath{\bigcup}}
\global\long\def\ensuremath{\bigcap}{\ensuremath{\bigcap}}
\global\long\def\ensuremath{\cup}{\ensuremath{\cup}}
\global\long\def\ensuremath{\cap}{\ensuremath{\cap}}
\global\long\def\ensuremath{\mathcal{P}}{\ensuremath{\mathcal{P}}}
\global\long\def\clsr#1{\ensuremath{\overline{#1}}}
\global\long\def\ensuremath{\Delta}{\ensuremath{\Delta}}
\global\long\def\operatorname{int}{\operatorname{int}}
\global\long\def\otimes{\otimes}
\global\long\def\bigotimes{\bigotimes}
\global\long\def\smlinprd#1#2{\ensuremath{\langle#1,#2\rangle}}
\global\long\def\inprd#1#2{\ensuremath{\left\langle #1,#2\right\rangle }}
\global\long\def\ensuremath{\perp}{\ensuremath{\perp}}
\global\long\def\ensuremath{\oplus}{\ensuremath{\oplus}}
\global\long\def\operatorname{sp}{\operatorname{sp}}
\global\long\def\operatorname{rk}{\operatorname{rk}}
\global\long\def\operatorname{proj}{\operatorname{proj}}
\global\long\def\operatorname{tr}{\operatorname{tr}}
\global\long\def\operatorname{vec}{\operatorname{vec}}
\global\long\def\operatorname{diag}{\operatorname{diag}}
\global\long\def\operatorname{col}{\operatorname{col}}
\global\long\def\ensuremath{\Omega}{\ensuremath{\Omega}}
\global\long\def\ensuremath{\omega}{\ensuremath{\omega}}
\global\long\def\sigf#1{\mathcal{#1}}
\global\long\def\ensuremath{\mathcal{F}}{\ensuremath{\mathcal{F}}} \global\long\def\ensuremath{\mathcal{G}}{\ensuremath{\mathcal{G}}}
\global\long\def\flt#1{\mathcal{#1}}
\global\long\def\mathcal{F}{\mathcal{F}} \global\long\def\mathcal{G}{\mathcal{G}}
\global\long\def\ensuremath{\mathcal{B}}{\ensuremath{\mathcal{B}}}
\global\long\def\ensuremath{\mathcal{C}}{\ensuremath{\mathcal{C}}}
\global\long\def\ensuremath{\mathcal{N}}{\ensuremath{\mathcal{N}}}
\global\long\def\mathfrak{g}{\mathfrak{g}}
\global\long\def\mathfrak{m}{\mathfrak{m}}
\global\long\def\ensuremath{\mathbb{P}}{\ensuremath{\mathbb{P}}}
\global\long\def\ensuremath{\mathbb{P}}{\ensuremath{\mathbb{P}}}
\global\long\def\mathcal{P}{\mathcal{P}}
\global\long\def\mathcal{M}{\mathcal{M}}
\global\long\def\ensuremath{\mathbb{E}}{\ensuremath{\mathbb{E}}}
\global\long\def\ensuremath{\mathbb{E}}{\ensuremath{\mathbb{E}}}
\global\long\def\ensuremath{(\ensuremath{\Omega},\mathcal{F},\ensuremath{\mathbb{P}})}{\ensuremath{(\ensuremath{\Omega},\mathcal{F},\ensuremath{\mathbb{P}})}}
\global\long\def\ensuremath{i.i.d.}{\ensuremath{i.i.d.}}
\global\long\def\ensuremath{a.s.}{\ensuremath{a.s.}}
\global\long\def\ensuremath{a.s.p.}{\ensuremath{a.s.p.}}
\global\long\def\ensuremath{\ensuremath{i.o.}}{\ensuremath{\ensuremath{i.o.}}}
\newcommand\mathpalette{\independenT}{\perp}{\mathpalette{\independenT}{\perp}} \def\independenT#1#2{\mathrel{\rlap{$#1#2$}\mkern2mu{#1#2}}}
\global\long\def\mathpalette{\independenT}{\perp}{\mathpalette{\independenT}{\perp}}
\global\long\def\ensuremath{\sim}{\ensuremath{\sim}}
\global\long\def\ensuremath{\sim_{\ensuremath{i.i.d.}}}{\ensuremath{\sim_{\ensuremath{i.i.d.}}}}
\global\long\def\ensuremath{\overset{a}{\ensuremath{\sim}}}{\ensuremath{\overset{a}{\ensuremath{\sim}}}}
\global\long\def\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}{\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}}
\global\long\def\inprobu#1{\ensuremath{\overset{#1}{\ensuremath{\rightarrow}}}}
\global\long\def\ensuremath{\overset{\ensuremath{a.s.}}{\ensuremath{\rightarrow}}}{\ensuremath{\overset{\ensuremath{a.s.}}{\ensuremath{\rightarrow}}}}
\global\long\def=_{\ensuremath{a.s.}}{=_{\ensuremath{a.s.}}}
\global\long\def\inLp#1{\ensuremath{\overset{\Ellp{#1}}{\ensuremath{\rightarrow}}}}
\global\long\def\ensuremath{\overset{d}{\ensuremath{\rightarrow}}}{\ensuremath{\overset{d}{\ensuremath{\rightarrow}}}}
\global\long\def=_{d}{=_{d}}
\global\long\def\ensuremath{\rightsquigarrow}{\ensuremath{\rightsquigarrow}}
\global\long\def\wkcu#1{\overset{#1}{\ensuremath{\rightsquigarrow}}}
\global\long\def\operatorname*{plim}{\operatorname*{plim}}
\global\long\def\operatorname{var}{\operatorname{var}}
\global\long\def\operatorname{lrvar}{\operatorname{lrvar}}
\global\long\def\operatorname{cov}{\operatorname{cov}}
\global\long\def\operatorname{corr}{\operatorname{corr}}
\global\long\def\operatorname{bias}{\operatorname{bias}}
\global\long\def\operatorname{MSE}{\operatorname{MSE}}
\global\long\def\operatorname{med}{\operatorname{med}}
\global\long\def\operatorname{avar}{\operatorname{avar}}
\global\long\def\operatorname{se}{\operatorname{se}}
\global\long\def\operatorname{sd}{\operatorname{sd}}
\global\long\defH_{0}{H_{0}}
\global\long\defH_{1}{H_{1}}
\global\long\def\mathcal{C}{\mathcal{C}}
\global\long\def\mathcal{R}{\mathcal{R}}
\global\long\def\mathcal{A}{\mathcal{A}}
\global\long\def\mathcal{H}{\mathcal{H}}
\global\long\def\ensuremath{\mathbb{W}}{\ensuremath{\mathbb{W}}}
\global\long\def\bullet{\bullet}
\global\long\def\cv#1{\left\langle #1\right\rangle }
\global\long\def\smlcv#1{\langle#1\rangle}
\global\long\def\qv#1{\left[#1\right]}
\global\long\def\smlqv#1{[#1]}
\global\long\def\top{\mathsf{T}}
\global\long\def\ensuremath{\mathbf{1}}{\ensuremath{\mathbf{1}}}
\global\long\def\mathcal{L}{\mathcal{L}}
\global\long\def\nabla{\nabla}
\global\long\def\ensuremath{\wedge}{\ensuremath{\wedge}} \global\long\def\ensuremath{\bigwedge}{\ensuremath{\bigwedge}}
\global\long\def\ensuremath{\vee}{\ensuremath{\vee}} \global\long\def\ensuremath{\bigvee}{\ensuremath{\bigvee}}
\global\long\def\operatorname{sgn}{\operatorname{sgn}}
\global\long\def\operatorname*{argmin}{\operatorname*{argmin}}
\global\long\def\operatorname*{argmax}{\operatorname*{argmax}}
\global\long\def\operatorname{Re}{\operatorname{Re}}
\global\long\def\operatorname{Im}{\operatorname{Im}}
\global\long\def\ensuremath{\mathrm{d}}{\ensuremath{\mathrm{d}}}
\global\long\def\ensuremath{\ensuremath{\mathrm{d}}}{\ensuremath{\ensuremath{\mathrm{d}}}}
\global\long\def\ensuremath{\,\ensuremath{\mathrm{d}}}{\ensuremath{\,\ensuremath{\mathrm{d}}}}
\global\long\def\ensuremath{\mathrm{i}}{\ensuremath{\mathrm{i}}}
\global\long\def\mathrm{e}{\mathrm{e}}
\global\long\def,\ {,\ }
\global\long\def\coloneqq{\coloneqq}
\global\long\def\eqqcolon{\eqqcolon}
\global\long\def\varepsilon{\varepsilon}
\global\long\def\mset#1{\mathcal{#1}}
\global\long\def\largedec#1{\mathbf{#1}}
\global\long\def\largedec z{\largedec z}
\newcommandx\Ican[1][usedefault, addprefix=\global, 1=]{I_{#1}^{\ast}}
\global\long{{ \mathrm{JSR}}}
\global\long{{ \mathrm{CJSR}}}
\global\long{{ \mathrm{RJSR}}}
\global\long\def\mathscr{M}{\mathscr{M}}
\global\long\def\b#1{\boldsymbol{#1}}
\global\long\def\ensuremath{\abv y}{\ensuremath{\abv y}}
\global\long\def\kappa{\kappa}
\global\long\def\operatorname{adj}{\operatorname{adj}}
\global\long\def\savebox{\@brx}{\(\m@th{\langle}\)} \mathopen{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}{\savebox{\@brx}{\(\m@th{\langle}\)} \mathopen{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}}
\global\long\def\savebox{\@brx}{\(\m@th{\rangle}\)} \mathclose{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}{\savebox{\@brx}{\(\m@th{\rangle}\)} \mathclose{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}}
\global\long\def\smldblangle#1{\ensuremath{\savebox{\@brx}{\(\m@th{\langle}\)} \mathopen{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}#1\savebox{\@brx}{\(\m@th{\rangle}\)} \mathclose{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}}}
{{{(i)}}}
{{{(ii)}}}
{{{(iii)}}}
\global\long\def\mathrm{(i)}{\mathrm{(i)}}
\global\long\def\mathrm{(ii)}{\mathrm{(ii)}}
\global\long\deff{f}
\global\long\def\b f{\b f}
\global\long\deff{f}
\global\long\defg{g}
\global\long\def\chi{\chi}
\global\long\def\mathrm{X}{\mathrm{X}}
\global\long\def\Theta{\Theta}
\global\long\def\theta{\theta}
\global\long\def\psi{\psi}
\global\long\def\mu{\mu}
\global\long{{\cal M}}
\global\long\def\chi{\chi}
\global\long\def\chi_{2}{\chi_{2}}
\global\long\def\chi_{1}{\chi_{1}}
\global\long\def\set#1{\mathscr{#1}}
\global\long\def\mathrm{H}{\mathrm{H}}
\global\long\defh{h}
\global\long\def\operatorname{co}{\operatorname{co}}
\global\long\def\operatorname{im}{\operatorname{im}}
\global\long\def\mathrm{lin}{\mathrm{lin}}
\global\long\def\mathfrak{z}{\mathfrak{z}}
\global\long\def\mathscr{F}{\mathscr{F}}
\global\long\def\mathscr{R}{\mathscr{R}}
\global\long\def\varrho{\varrho}
\global\long\def\rho{\rho}
{\fnsymbol{footnote}}
\footnotetext[1]{Department\ of Economics and Corpus Christi College; [email removed]}
\footnotetext[2]{Department\ of Economics; [email removed]}
{\arabic{footnote}}
\setcounter{footnote}{0}
We thank participants at the Oxford Bulletin of Economics and Statistics `40 years of Unit Roots and Cointegration' workshop, held in Oxford in April 2025, for their comments and advice.
\thispagestyle{plain}
\pagenumbering{roman}
\thispagestyle{plain}
\setcounter{tocdepth}{2}
\pagenumbering{arabic}
{2.1}
\global\long\def\sim1{\sim1}
\global\long\def\dmn#1{\bar{#1}}
\global\long\def\dtr#1{\ddot{#1}}
\global\long\def\tau{\tau}
\global\long\defT{T}
\global\long\def\mathbb{W}{\mathbb{W}}
\global\long\def\mathbb{V}{\mathbb{V}}
\global\long\def\mathfrak{z}{\mathfrak{z}}
\global\long\def\top{\top}
{\citetalias{DMW22}}
For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).
In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.
The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form
where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.
There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per
where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.
In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.
We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining
we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow
or, more compactly,
where
for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.
We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq
=
, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.
{{{{\scalefont{0.76}DGP}}}}
As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that
on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).
In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if
While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables
where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting
for $\mathfrak{z}\in\mathbb{C}$, where
we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).
To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.
{{{{\scalefont{0.76}DGP$^{\ast}$}}}}
DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq
, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define
{{{{\scalefont{0.76}CVAR}}}}
The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}
=\alpha
\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of
for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as
where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and
Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.
{{{{\scalefont{0.76}CO{{(ii)}}}}}}
Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that
or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.
Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per
{{{{\scalefont{0.76}DET}}}}
Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that
where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).
On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.
Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:
But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.
We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)
The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)
To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq
z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).
By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.
Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)
To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq
=\left[
\right]
=S_{p}(y_{t})z_{t} \] via
from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.
In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix
where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).
Thus the linear transformation
exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now
where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.
To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.
To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation
where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.
Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]
This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of
where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at
which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.
Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).
{{{{\scalefont{0.76}DET$^{\prime}$}}}}
The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[
=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.
Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as
where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to
ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for
This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.
To state the limiting distribution of the test statistic, define
where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process
For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).
In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.
The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form
where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.
There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per
where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.
In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.
We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining
we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow
or, more compactly,
where
for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.
We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq
=
, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.
{{{{\scalefont{0.76}DGP}}}}
As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that
on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).
In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if
While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables
where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting
for $\mathfrak{z}\in\mathbb{C}$, where
we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).
To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.
{{{{\scalefont{0.76}DGP$^{\ast}$}}}}
DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq
, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define
{{{{\scalefont{0.76}CVAR}}}}
The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}
=\alpha
\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of
for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as
where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and
Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.
{{{{\scalefont{0.76}CO{{(ii)}}}}}}
Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that
or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.
Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per
{{{{\scalefont{0.76}DET}}}}
Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that
where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).
On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.
Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:
But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.
We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)
The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)
To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq
z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).
By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.
Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)
To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq
=\left[
\right]
=S_{p}(y_{t})z_{t} \] via
from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.
In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix
where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).
Thus the linear transformation
exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now
where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.
To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.
To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation
where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.
Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]
This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of
where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at
which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.
Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).
{{{{\scalefont{0.76}DET$^{\prime}$}}}}
The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[
=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.
Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as
where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to
ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for
This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.
To state the limiting distribution of the test statistic, define
where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process
For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).
In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.
The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form
where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.
There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per
where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.
In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.
We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining
we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow
or, more compactly,
where
for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.
We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq
=
, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.
{{{{\scalefont{0.76}DGP}}}}
As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that
on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).
In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if
While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables
where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting
for $\mathfrak{z}\in\mathbb{C}$, where
we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).
To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.
{{{{\scalefont{0.76}DGP$^{\ast}$}}}}
DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq
, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define
{{{{\scalefont{0.76}CVAR}}}}
The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}
=\alpha
\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of
for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as
where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and
Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.
{{{{\scalefont{0.76}CO{{(ii)}}}}}}
Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that
or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.
Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per
{{{{\scalefont{0.76}DET}}}}
Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that
where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).
On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.
Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:
But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.
We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)
The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)
To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq
z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).
By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.
Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)
To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq
=\left[
\right]
=S_{p}(y_{t})z_{t} \] via
from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.
In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix
where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).
Thus the linear transformation
exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now
where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.
To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.
To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation
where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.
Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]
This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of
where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at
which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.
Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).
{{{{\scalefont{0.76}DET$^{\prime}$}}}}
The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[
=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.
Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as
where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to
ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for
This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.
To state the limiting distribution of the test statistic, define
where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process
For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).
In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.
The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form
where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.
There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per
where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.
In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.
We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining
we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow
or, more compactly,
where
for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.
We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq
=
, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.
{{{{\scalefont{0.76}DGP}}}}
As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that
on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).
In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if
While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables
where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting
for $\mathfrak{z}\in\mathbb{C}$, where
we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).
To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.
{{{{\scalefont{0.76}DGP$^{\ast}$}}}}
DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq
, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define
{{{{\scalefont{0.76}CVAR}}}}
The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}
=\alpha
\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of
for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as
where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and
Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.
{{{{\scalefont{0.76}CO{{(ii)}}}}}}
Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that
or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.
Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per
{{{{\scalefont{0.76}DET}}}}
Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that
where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).
On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.
Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:
But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.
We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)
The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)
To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq
z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).
By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.
Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)
To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq
=\left[
\right]
=S_{p}(y_{t})z_{t} \] via
from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.
In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix
where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).
Thus the linear transformation
exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now
where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.
To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.
To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation
where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.
Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]
This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of
where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at
which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.
Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).
{{{{\scalefont{0.76}DET$^{\prime}$}}}}
The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[
=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.
Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as
where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to
ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for
This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.
To state the limiting distribution of the test statistic, define
where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process
For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).
In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.
The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form
where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.
There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per
where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.
In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.
We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining
we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow
or, more compactly,
where
for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.
We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq
=
, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.
{{{{\scalefont{0.76}DGP}}}}
As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that
on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).
In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if
While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables
where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting
for $\mathfrak{z}\in\mathbb{C}$, where
we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).
To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.
{{{{\scalefont{0.76}DGP$^{\ast}$}}}}
DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq
, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define
{{{{\scalefont{0.76}CVAR}}}}
The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}
=\alpha
\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of
for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as
where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and
Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.
{{{{\scalefont{0.76}CO{{(ii)}}}}}}
Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that
or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.
Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per
{{{{\scalefont{0.76}DET}}}}
Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that
where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).
On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.
Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:
But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.
We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)
The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)
To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq
z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).
By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.
Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)
To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq
=\left[
\right]
=S_{p}(y_{t})z_{t} \] via
from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.
In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix
where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).
Thus the linear transformation
exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now
where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.
To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.
To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation
where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.
Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]
This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of
where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at
which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.
Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).
{{{{\scalefont{0.76}DET$^{\prime}$}}}}
The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[
=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.
Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as
where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to
ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for
This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.
To state the limiting distribution of the test statistic, define
where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process
For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).
In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.
The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form
where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.
There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per
where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.
In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.
We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining
we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow
or, more compactly,
where
for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.
We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq
=
, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.
{{{{\scalefont{0.76}DGP}}}}
As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that
on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).
In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if
While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables
where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting
for $\mathfrak{z}\in\mathbb{C}$, where
we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).
To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.
{{{{\scalefont{0.76}DGP$^{\ast}$}}}}
DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq
, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define
{{{{\scalefont{0.76}CVAR}}}}
The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}
=\alpha
\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of
for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as
where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and
Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.
{{{{\scalefont{0.76}CO{{(ii)}}}}}}
Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that
or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.
Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per
{{{{\scalefont{0.76}DET}}}}
Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that
where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).
On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.
Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:
But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.
We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)
The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)
To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq
z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).
By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.
Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)
To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq
=\left[
\right]
=S_{p}(y_{t})z_{t} \] via
from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.
In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix
where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).
Thus the linear transformation
exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now
where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.
To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.
To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation
where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.
Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]
This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of
where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at
which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.
Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).
{{{{\scalefont{0.76}DET$^{\prime}$}}}}
The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[
=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.
Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as
where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to
ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for
This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.
To state the limiting distribution of the test statistic, define
where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process