EconBase
← Back to paper

Inference on Common Trends in a Cointegrated Nonlinear SVAR

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

526,155 characters · 60 sections · 192 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
This text was truncated for display. The citation measures were computed over the complete text.

Inference on Common Trends in a Cointegrated Nonlinear SVAR

\global\long\def\uwrite#1#2{\underset{#2}{\underbrace{#1}} }

\global\long\def\blw#1{\ensuremath{#1}}

\global\long\def\abv#1{\ensuremath{\overline{#1}}}

\global\long\def\vect#1{\mathbf{#1}}

\global\long\def\smlseq#1{\{#1\} }

\global\long\def\seq#1{\left\{ #1\right\} }

\global\long\def\smlsetof#1#2{\{#1\mid#2\} }

\global\long\def\setof#1#2{\left\{ #1\mid#2\right\} }

\global\long\def\ensuremath{\rightarrow}{\ensuremath{\rightarrow}}

\global\long\def\ensuremath{\nrightarrow}{\ensuremath{\nrightarrow}}

\global\long\def\ensuremath{\uparrow}{\ensuremath{\uparrow}}

\global\long\def\ensuremath{\downarrow}{\ensuremath{\downarrow}}

\global\long\def\ensuremath{\upuparrows}{\ensuremath{\upuparrows}}

\global\long\def\ensuremath{\downdownarrows}{\ensuremath{\downdownarrows}}

\global\long\def\ensuremath{\nearrow}{\ensuremath{\nearrow}}

\global\long\def\ensuremath{\searrow}{\ensuremath{\searrow}}

\global\long\def\ensuremath{\rightarrow}{\ensuremath{\rightarrow}}

\global\long\def\ensuremath{\mapsto}{\ensuremath{\mapsto}}

\global\long\def\ensuremath{\circ}{\ensuremath{\circ}}

\global\long\defC{C}

\global\long\defD{D}

\global\long\def\Ellp#1{\ensuremath{\mathcal{L}^{#1}}}

\global\long\def\ensuremath{\mathbb{N}}{\ensuremath{\mathbb{N}}}

\global\long\def\mathbb{R}{\mathbb{R}}

\global\long\def\mathbb{C}{\mathbb{C}}

\global\long\def\mathbb{Q}{\mathbb{Q}}

\global\long\def\mathbb{Z}{\mathbb{Z}}

\global\long\def\abs#1{\ensuremath{\left|#1\right|}}

\global\long\def\smlabs#1{\ensuremath{\lvert#1\rvert}}

\global\long\def\bigabs#1{\ensuremath{\bigl|#1\bigr|}}

\global\long\def\Bigabs#1{\ensuremath{\Bigl|#1\Bigr|}}

\global\long\def\biggabs#1{\ensuremath{\biggl|#1\biggr|}}

\global\long\def\norm#1{\ensuremath{\left\Vert #1\right\Vert }}

\global\long\def\smlnorm#1{\ensuremath{\lVert#1\rVert}}

\global\long\def\bignorm#1{\ensuremath{\bigl\|#1\bigr\|}}

\global\long\def\Bignorm#1{\ensuremath{\Bigl\|#1\Bigr\|}}

\global\long\def\biggnorm#1{\ensuremath{\biggl\|#1\biggr\|}}

\global\long\def\floor#1{\left\lfloor #1\right\rfloor } \global\long\def\smlfloor#1{\lfloor#1\rfloor}

\global\long\def\ceil#1{\left\lceil #1\right\rceil } \global\long\def\smlceil#1{\lceil#1\rceil}

\global\long\def\ensuremath{\bigcup}{\ensuremath{\bigcup}}

\global\long\def\ensuremath{\bigcap}{\ensuremath{\bigcap}}

\global\long\def\ensuremath{\cup}{\ensuremath{\cup}}

\global\long\def\ensuremath{\cap}{\ensuremath{\cap}}

\global\long\def\ensuremath{\mathcal{P}}{\ensuremath{\mathcal{P}}}

\global\long\def\clsr#1{\ensuremath{\overline{#1}}}

\global\long\def\ensuremath{\Delta}{\ensuremath{\Delta}}

\global\long\def\operatorname{int}{\operatorname{int}}

\global\long\def\otimes{\otimes}

\global\long\def\bigotimes{\bigotimes}

\global\long\def\smlinprd#1#2{\ensuremath{\langle#1,#2\rangle}}

\global\long\def\inprd#1#2{\ensuremath{\left\langle #1,#2\right\rangle }}

\global\long\def\ensuremath{\perp}{\ensuremath{\perp}}

\global\long\def\ensuremath{\oplus}{\ensuremath{\oplus}}

\global\long\def\operatorname{sp}{\operatorname{sp}}

\global\long\def\operatorname{rk}{\operatorname{rk}}

\global\long\def\operatorname{proj}{\operatorname{proj}}

\global\long\def\operatorname{tr}{\operatorname{tr}}

\global\long\def\operatorname{vec}{\operatorname{vec}}

\global\long\def\operatorname{diag}{\operatorname{diag}}

\global\long\def\operatorname{col}{\operatorname{col}}

\global\long\def\ensuremath{\Omega}{\ensuremath{\Omega}}

\global\long\def\ensuremath{\omega}{\ensuremath{\omega}}

\global\long\def\sigf#1{\mathcal{#1}}

\global\long\def\ensuremath{\mathcal{F}}{\ensuremath{\mathcal{F}}} \global\long\def\ensuremath{\mathcal{G}}{\ensuremath{\mathcal{G}}}

\global\long\def\flt#1{\mathcal{#1}}

\global\long\def\mathcal{F}{\mathcal{F}} \global\long\def\mathcal{G}{\mathcal{G}}

\global\long\def\ensuremath{\mathcal{B}}{\ensuremath{\mathcal{B}}}

\global\long\def\ensuremath{\mathcal{C}}{\ensuremath{\mathcal{C}}}

\global\long\def\ensuremath{\mathcal{N}}{\ensuremath{\mathcal{N}}}

\global\long\def\mathfrak{g}{\mathfrak{g}}

\global\long\def\mathfrak{m}{\mathfrak{m}}

\global\long\def\ensuremath{\mathbb{P}}{\ensuremath{\mathbb{P}}}

\global\long\def\ensuremath{\mathbb{P}}{\ensuremath{\mathbb{P}}}

\global\long\def\mathcal{P}{\mathcal{P}}

\global\long\def\mathcal{M}{\mathcal{M}}

\global\long\def\ensuremath{\mathbb{E}}{\ensuremath{\mathbb{E}}}

\global\long\def\ensuremath{\mathbb{E}}{\ensuremath{\mathbb{E}}}

\global\long\def\ensuremath{(\ensuremath{\Omega},\mathcal{F},\ensuremath{\mathbb{P}})}{\ensuremath{(\ensuremath{\Omega},\mathcal{F},\ensuremath{\mathbb{P}})}}

\global\long\def\ensuremath{i.i.d.}{\ensuremath{i.i.d.}}

\global\long\def\ensuremath{a.s.}{\ensuremath{a.s.}}

\global\long\def\ensuremath{a.s.p.}{\ensuremath{a.s.p.}}

\global\long\def\ensuremath{\ensuremath{i.o.}}{\ensuremath{\ensuremath{i.o.}}}

\newcommand\mathpalette{\independenT}{\perp}{\mathpalette{\independenT}{\perp}} \def\independenT#1#2{\mathrel{\rlap{$#1#2$}\mkern2mu{#1#2}}}

\global\long\def\mathpalette{\independenT}{\perp}{\mathpalette{\independenT}{\perp}}

\global\long\def\ensuremath{\sim}{\ensuremath{\sim}}

\global\long\def\ensuremath{\sim_{\ensuremath{i.i.d.}}}{\ensuremath{\sim_{\ensuremath{i.i.d.}}}}

\global\long\def\ensuremath{\overset{a}{\ensuremath{\sim}}}{\ensuremath{\overset{a}{\ensuremath{\sim}}}}

\global\long\def\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}{\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}}

\global\long\def\inprobu#1{\ensuremath{\overset{#1}{\ensuremath{\rightarrow}}}}

\global\long\def\ensuremath{\overset{\ensuremath{a.s.}}{\ensuremath{\rightarrow}}}{\ensuremath{\overset{\ensuremath{a.s.}}{\ensuremath{\rightarrow}}}}

\global\long\def=_{\ensuremath{a.s.}}{=_{\ensuremath{a.s.}}}

\global\long\def\inLp#1{\ensuremath{\overset{\Ellp{#1}}{\ensuremath{\rightarrow}}}}

\global\long\def\ensuremath{\overset{d}{\ensuremath{\rightarrow}}}{\ensuremath{\overset{d}{\ensuremath{\rightarrow}}}}

\global\long\def=_{d}{=_{d}}

\global\long\def\ensuremath{\rightsquigarrow}{\ensuremath{\rightsquigarrow}}

\global\long\def\wkcu#1{\overset{#1}{\ensuremath{\rightsquigarrow}}}

\global\long\def\operatorname*{plim}{\operatorname*{plim}}

\global\long\def\operatorname{var}{\operatorname{var}}

\global\long\def\operatorname{lrvar}{\operatorname{lrvar}}

\global\long\def\operatorname{cov}{\operatorname{cov}}

\global\long\def\operatorname{corr}{\operatorname{corr}}

\global\long\def\operatorname{bias}{\operatorname{bias}}

\global\long\def\operatorname{MSE}{\operatorname{MSE}}

\global\long\def\operatorname{med}{\operatorname{med}}

\global\long\def\operatorname{avar}{\operatorname{avar}}

\global\long\def\operatorname{se}{\operatorname{se}}

\global\long\def\operatorname{sd}{\operatorname{sd}}

\global\long\defH_{0}{H_{0}}

\global\long\defH_{1}{H_{1}}

\global\long\def\mathcal{C}{\mathcal{C}}

\global\long\def\mathcal{R}{\mathcal{R}}

\global\long\def\mathcal{A}{\mathcal{A}}

\global\long\def\mathcal{H}{\mathcal{H}}

\global\long\def\ensuremath{\mathbb{W}}{\ensuremath{\mathbb{W}}}

\global\long\def\bullet{\bullet}

\global\long\def\cv#1{\left\langle #1\right\rangle }

\global\long\def\smlcv#1{\langle#1\rangle}

\global\long\def\qv#1{\left[#1\right]}

\global\long\def\smlqv#1{[#1]}

\global\long\def\top{\mathsf{T}}

\global\long\def\ensuremath{\mathbf{1}}{\ensuremath{\mathbf{1}}}

\global\long\def\mathcal{L}{\mathcal{L}}

\global\long\def\nabla{\nabla}

\global\long\def\ensuremath{\wedge}{\ensuremath{\wedge}} \global\long\def\ensuremath{\bigwedge}{\ensuremath{\bigwedge}}

\global\long\def\ensuremath{\vee}{\ensuremath{\vee}} \global\long\def\ensuremath{\bigvee}{\ensuremath{\bigvee}}

\global\long\def\operatorname{sgn}{\operatorname{sgn}}

\global\long\def\operatorname*{argmin}{\operatorname*{argmin}}

\global\long\def\operatorname*{argmax}{\operatorname*{argmax}}

\global\long\def\operatorname{Re}{\operatorname{Re}}

\global\long\def\operatorname{Im}{\operatorname{Im}}

\global\long\def\ensuremath{\mathrm{d}}{\ensuremath{\mathrm{d}}}

\global\long\def\ensuremath{\ensuremath{\mathrm{d}}}{\ensuremath{\ensuremath{\mathrm{d}}}}

\global\long\def\ensuremath{\,\ensuremath{\mathrm{d}}}{\ensuremath{\,\ensuremath{\mathrm{d}}}}

\global\long\def\ensuremath{\mathrm{i}}{\ensuremath{\mathrm{i}}}

\global\long\def\mathrm{e}{\mathrm{e}}

\global\long\def,\ {,\ }

\global\long\def\coloneqq{\coloneqq}

\global\long\def\eqqcolon{\eqqcolon}

comment\global\long\def\objlabel#1#2{\{\}} \global\long\def\objref#1{\{\}}

\global\long\def\varepsilon{\varepsilon}

\global\long\def\mset#1{\mathcal{#1}}

\global\long\def\largedec#1{\mathbf{#1}}

\global\long\def\largedec z{\largedec z}

\newcommandx\Ican[1][usedefault, addprefix=\global, 1=]{I_{#1}^{\ast}}

\global\long{{ \mathrm{JSR}}}

\global\long{{ \mathrm{CJSR}}}

\global\long{{ \mathrm{RJSR}}}

\global\long\def\mathscr{M}{\mathscr{M}}

\global\long\def\b#1{\boldsymbol{#1}}

\global\long\def\ensuremath{\abv y}{\ensuremath{\abv y}}

\global\long\def\kappa{\kappa}

\global\long\def\operatorname{adj}{\operatorname{adj}}

\global\long\def\savebox{\@brx}{\(\m@th{\langle}\)} \mathopen{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}{\savebox{\@brx}{\(\m@th{\langle}\)} \mathopen{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}}

\global\long\def\savebox{\@brx}{\(\m@th{\rangle}\)} \mathclose{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}{\savebox{\@brx}{\(\m@th{\rangle}\)} \mathclose{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}}

\global\long\def\smldblangle#1{\ensuremath{\savebox{\@brx}{\(\m@th{\langle}\)} \mathopen{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}#1\savebox{\@brx}{\(\m@th{\rangle}\)} \mathclose{\copy\@brx\kern-0.5\wd\@brx\usebox{\@brx}}}}

{{{(i)}}}

{{{(ii)}}}

{{{(iii)}}}

\global\long\def\mathrm{(i)}{\mathrm{(i)}}

\global\long\def\mathrm{(ii)}{\mathrm{(ii)}}

\global\long\deff{f}

\global\long\def\b f{\b f}

\global\long\deff{f}

\global\long\defg{g}

\global\long\def\chi{\chi}

\global\long\def\mathrm{X}{\mathrm{X}}

\global\long\def\Theta{\Theta}

\global\long\def\theta{\theta}

\global\long\def\psi{\psi}

\global\long\def\mu{\mu}

\global\long{{\cal M}}

\global\long\def\chi{\chi}

\global\long\def\chi_{2}{\chi_{2}}

\global\long\def\chi_{1}{\chi_{1}}

\global\long\def\set#1{\mathscr{#1}}

\global\long\def\mathrm{H}{\mathrm{H}}

\global\long\defh{h}

\global\long\def\operatorname{co}{\operatorname{co}}

\global\long\def\operatorname{im}{\operatorname{im}}

\global\long\def\mathrm{lin}{\mathrm{lin}}

\global\long\def\mathfrak{z}{\mathfrak{z}}

\global\long\def\mathscr{F}{\mathscr{F}}

\global\long\def\mathscr{R}{\mathscr{R}}

\global\long\def\varrho{\varrho}

\global\long\def\rho{\rho}

{\fnsymbol{footnote}}

\footnotetext[1]{Department\ of Economics and Corpus Christi College; [email removed]}

\footnotetext[2]{Department\ of Economics; [email removed]}

{\arabic{footnote}}

\setcounter{footnote}{0}

abstractWe consider the problem of performing inference on the number of common stochastic trends when data is generated by a cointegrated CKSVAR (a two-regime, piecewise affine SVAR; Mavroeidis, 2021), using a modified version of the Breitung (2002) multivariate variance ratio test that is robust to the presence of nonlinear cointegration (of a known form). To derive the asymptotics of our test statistic, we prove a fundamental LLN-type result for a class of stable but nonstationary autoregressive processes, using a novel dual linear process approximation. We show that our modified test yields correct inferences regarding the number of common trends in such a system, whereas the unmodified test tends to infer a higher number of common trends than are actually present, when cointegrating relations are nonlinear.

We thank participants at the Oxford Bulletin of Economics and Statistics `40 years of Unit Roots and Cointegration' workshop, held in Oxford in April 2025, for their comments and advice.

\thispagestyle{plain}

\pagenumbering{roman}

\thispagestyle{plain}

\setcounter{tocdepth}{2}

\pagenumbering{arabic}

{2.1}

\global\long\def\sim1{\sim1}

\global\long\def\dmn#1{\bar{#1}}

\global\long\def\dtr#1{\ddot{#1}}

\global\long\def\tau{\tau}

\global\long\defT{T}

\global\long\def\mathbb{W}{\mathbb{W}}

\global\long\def\mathbb{V}{\mathbb{V}}

\global\long\def\mathfrak{z}{\mathfrak{z}}

\global\long\def\top{\top}

{\citetalias{DMW22}}

Introduction

For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).

In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.

The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form

equation[equation omitted — 82 chars of source]

where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.

There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per

equation[equation omitted — 187 chars of source]

where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.

In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.

notation*$e_{m,i}$ denotes the $i$th column of the $m\times m$ identity matrix $I_{m}$; when $m$ is clear from the context, we write this simply as $e_{i}$. In a statement such as $f(a^{\pm},b^{\pm})=0$, the notation `$\pm$' signifies that both $f(a^{+},b^{+})=0$ and $f(a^{-},b^{-})=0$ hold; similarly, `$a^{\pm}\in A$' denotes that both $a^{+}$ and $a^{-}$ are elements of $A$. All limits are taken as $n\ensuremath{\rightarrow}\infty$ unless otherwise stated. $\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}$ and $\ensuremath{\rightsquigarrow}$ respectively denote convergence in probability and in distribution (weak convergence). We write `$X_{n}(\lambda)\ensuremath{\rightsquigarrow} X(\lambda)$ on $D_{\mathbb{R}^{m}}[0,1]$' to denote that $\{X_{n}\}$ converges weakly to $X$, where these are considered as random elements of $D_{\mathbb{R}^{m}}[0,1]$, the space of cadlag functions $[0,1]\ensuremath{\rightarrow}\mathbb{R}^{m}$, equipped with the uniform topology; we denote this as $D[0,1]$ whenever the value of $m$ is clear from the context. $\smlnorm{\cdot}$ denotes the Euclidean norm on $\mathbb{R}^{m}$, and the matrix norm that it induces. For $X$ a random vector and $p\geq1$, $\smlnorm X_{p}\coloneqq(\ensuremath{\mathbb{E}}\smlnorm X^{p})^{1/p}$. $C$, $C_{1}$, etc., denote generic constants that may take different values at different places of the same proof.

Model: the censored and kinked SVAR

Framework

We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining

align[align omitted — 111 chars of source]

we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow

equation[equation omitted — 193 chars of source]

or, more compactly,

equation[equation omitted — 100 chars of source]

where

align*[align* omitted — 158 chars of source]

for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.

We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq

bmatrix[bmatrix omitted — 55 chars of source]

=

bmatrix[bmatrix omitted — 118 chars of source]

, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.

{{{{\scalefont{0.76}DGP}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\{(y_{t},x_{t})\}$ are generated according to (ref)--(ref) with $b=0$, with (possibly random) initial values $(y_{i},x_{i})$, for $i\in\{-k+1,\ldots,0\}$; • $\operatorname{sgn}(\det\Phi_{0}^{+})=\operatorname{sgn}(\det\Phi_{0}^{-})\neq0$. • $\Phi_{0,xx}$ is invertible, and \[ \operatorname{sgn}\{\phi_{0,yy}^{+}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{+}\}=\operatorname{sgn}\{\phi_{0,yy}^{-}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{-}\}>0. \]$\{u_{t}\}_{t\in\mathbb{Z}}$ is an i.i.d.\ sequence in $\mathbb{R}^{p}$ with $\ensuremath{\mathbb{E}} u_{t}=0$, $\ensuremath{\mathbb{E}} u_{t}u_{t}^{\top}=\Sigma_{u}$ positive definite, and $\smlnorm{u_{t}}_{2+\delta_{u}}<\infty$ for some $\delta_{u}>0$. \end{enumerate}

As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that

equation[equation omitted — 141 chars of source]

on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).

Canonical form

In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if

equation[equation omitted — 119 chars of source]

While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables

equation[equation omitted — 401 chars of source]

where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting

equation[equation omitted — 275 chars of source]

for $\mathfrak{z}\in\mathbb{C}$, where

equation[equation omitted — 127 chars of source]

we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).

To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.

{{{{\scalefont{0.76}DGP$^{\ast}$}}}}

assumption$\{(y_{t},x_{t})\}$ are generated by a canonical CKSVAR, i.e.\ (ref) holds with $\Phi_{0}=[\phi_{0}^{+},\phi_{0}^{-},\Phi^{x}]=\Ican[p]$, so that (ref) may be equivalently written as \begin{equation} \begin{bmatrix}y_{t}\\ x_{t} \end{bmatrix}=c+\sum_{i=1}^{k}\begin{bmatrix}\phi_{i}^{+} & \phi_{i}^{-} & \Phi_{i}^{x}\end{bmatrix}\begin{bmatrix}y_{t-i}^{+}\\ y_{t-i}^{-}\\ x_{t-i} \end{bmatrix}+u_{t}. \end{equation}

The cointegrated CKSVAR

DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq

bmatrix[bmatrix omitted — 62 chars of source]

, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define

align*[align* omitted — 107 chars of source]

{{{{\scalefont{0.76}CVAR}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\det\Phi^{\pm}(\mathfrak{z})$ has $q^{\pm}\in\{1,\ldots,p\}$ roots at real unity, and all others outside the unit circle; and • $\operatorname{rk}\Pi^{\pm}=r^{\pm}=p-q^{\pm}$. \end{enumerate}

The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}

bmatrix[bmatrix omitted — 35 chars of source]

=\alpha

bmatrix[bmatrix omitted — 47 chars of source]

\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of

gather[gather omitted — 420 chars of source]

for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as

align[align omitted — 388 chars of source]

where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and

equation[equation omitted — 168 chars of source]

Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.

{{{{\scalefont{0.76}CO{{(ii)}}}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}},itemsep=1pt,topsep=2pt] • $r^{+}=r^{-}=\operatorname{rk}\Pi^{x}=r$, for some $r\in\{0,1,\ldots,p-1\}$. • $\rho_{{\scriptstyle \mathrm{JSR}}}(\{I+\tilde{\b{\beta}}(+1)^{\top}\tilde{\b{\alpha}},I+\tilde{\b{\beta}}(-1)^{\top}\tilde{\b{\alpha}}\})<1$. • $\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;+1)\beta_{\perp}(+1)=\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;-1)\beta_{\perp}(-1)\neq0$. • \begin{enumerate}[label={{{\scalefont{0.76}\alph*.}}}, ref={{{\scalefont{0.76}.\alph*}}}, leftmargin=0.40cm] • $\beta(y_{t})^{\top}z_{t}$, and $\Delta z_{t}$ have uniformly bounded $2+\delta_{u}$ moments, for $t\in\{-k+1,\ldots,0\}$. • $n^{-1/2}z_{0}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}{\cal Z}_{0}=[\begin{smallmatrix}{\cal Y}_{0}\\ {\cal X}_{0} \end{smallmatrix}]$, where ${\cal Z}_{0}$ is non-random, and satisfies $\beta(\mathcal{Y}_{0})^{\top}{\cal Z}_{0}=0$. \end{enumerate} \end{enumerate}

Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that

equation[equation omitted — 71 chars of source]

or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.

Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per

{{{{\scalefont{0.76}DET}}}}

assumption$c\in\operatorname{sp}\Pi^{+}\ensuremath{\cap}\operatorname{sp}\Pi^{-}$.

Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that

equation[equation omitted — 297 chars of source]

where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).

On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.

Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:

align*[align* omitted — 115 chars of source]

But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.

The modified Breitung (2002) test

Fundamental ideas

We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)

The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)

propSuppose that $\{w_{n,t}\}_{t=1}^{n}$ is a triangular array, taking values in $\mathbb{R}^{d_{w}}$, such that \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}\ensuremath{\rightsquigarrow}\int_{0}^{\lambda}\begin{bmatrix}\mathbb{W}(s)\\ 0_{d_{w}-\ell} \end{bmatrix}\ensuremath{\,\ensuremath{\mathrm{d}}} s\eqqcolon\begin{bmatrix}\mathbb{V}(\lambda)\\ 0_{d_{w}-\ell} \end{bmatrix} \end{equation} on $D_{\mathbb{R}^{d_{w}}}[0,1]$, where $\mathbb{W}$ is a random element of $D_{\mathbb{R}^{\ell}}[0,1]$ and \begin{align} \frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top} & \ensuremath{\rightsquigarrow}\begin{bmatrix}\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s & 0\\ 0 & \Omega \end{bmatrix} \end{align} where $\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$, $\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$ and $\Omega\in\mathbb{R}^{(d_{w}-\ell)\times(d_{w}-\ell)}$ are a.s.\ positive definite. Let $\{\lambda_{n,i}\}_{i=1}^{d_{w}}$ denote the solutions to \[ \det(\lambda\mathbb{B}_{n}-\mathbb{A}_{n})=0 \] ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,d_{w}}$, for \begin{align*} \mathbb{A}_{n} & \coloneqq\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}, & \mathbb{B}_{n} & \coloneqq\sum_{t=1}^{n}\sum_{i=1}^{t}w_{n,i}\sum_{j=1}^{t}w_{n,j}^{\top}. \end{align*} Then \begin{enumerate} • if $\ell_{0}=\ell$, \[ n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\rightsquigarrow}\operatorname{tr}\left[\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\left(\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\right)^{-1}\right]; \] • if $\ell_{0}>\ell$, $n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}\infty$. \end{enumerate}

To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq

bmatrix[bmatrix omitted — 57 chars of source]

z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).

By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.

Extension to nonlinearly cointegrated series

Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)

To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq

bmatrix[bmatrix omitted — 43 chars of source]

=\left[

array[array omitted — 110 chars of source]

\right]

bmatrix[bmatrix omitted — 27 chars of source]

=S_{p}(y_{t})z_{t} \] via

equation[equation omitted — 410 chars of source]

from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.

In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix

equation[equation omitted — 163 chars of source]

where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).

Thus the linear transformation

equation[equation omitted — 468 chars of source]

exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now

equation[equation omitted — 280 chars of source]

where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.

To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.

LLN for regime-switching processes

To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation

equation[equation omitted — 123 chars of source]

where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.

Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]

This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of

align*[align* omitted — 571 chars of source]

where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at

align*[align* omitted — 342 chars of source]

which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.

Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).

{{{{\scalefont{0.76}DET$^{\prime}$}}}}

assumption$e_{1}^{\top}P_{\beta_{\perp}}(+1)c=0$.

The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[

bmatrix[bmatrix omitted — 35 chars of source]

=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.

thmSuppose (ref), (ref), (ref) and (ref) hold. Let $\{A_{t}\}$, $\{B_{t}\}$ and $\{c_{t}\}$ be random sequences adapted to $\{\mathcal{F}_{t}\}$, respectively taking values in $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$, where $t\in\mathbb{Z}$. Suppose $\{v_{t}\}$ is i.i.d\ with $\ensuremath{\mathbb{E}} v_{t}=0$, and that $\{w_{t}\}$ satisfies \begin{equation} w_{t}=c_{t}+A_{t}w_{t-1}+B_{t}v_{t} \end{equation} for $t\geq-k$ and some given (random) $w_{-k}$ (with $w_{t}\coloneqq0$ for all $t\leq-k-1$); and: \begin{enumerate}[itemsep=2pt,topsep=3pt] • $A_{t}\in\mset A$, $B_{t}\in\mset B$ and $c_{t}\in{\cal C}$ for all $t\in\ensuremath{\mathbb{N}}$, where $\mset A$, $\mset B$ and ${\cal C}$ are bounded subsets of $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$ respectively, and $\rho_{{\scriptstyle \mathrm{JSR}}}(\mset A)<1$; • there exist $A^{\pm}\in\mset A$, $B^{\pm}\in\mset B$ and $c^{\pm}\in\mset C$ such that \begin{align*} y_{t-1}>0 and y_{t}>0 & \implies A_{t}=A^{+},\ B_{t}=B^{+},\ c_{t}=c^{+},\\ y_{t-1}<0 and y_{t}<0 & \implies A_{t}=A^{-},\ B_{t}=B^{-},\ c_{t}=c^{-}; \end{align*} • $m_{0}\geq1$ is such that $\smlnorm{w_{0}}_{m_{0}}+\smlnorm{v_{0}}_{m_{0}}<\infty$. • $g:\mathbb{R}^{d_{w}}\ensuremath{\rightarrow}\mathbb{R}^{d_{g}}$ is a continuous function satisfying \begin{equation} \smlnorm{g(w)-g(w^{\prime})}\leq C(1+\smlnorm w^{\ell_{0}}+\smlnorm{w^{\prime}}^{\ell_{0}})\smlnorm{w-w^{\prime}} \end{equation} for all $w,w^{\prime}\in\mathbb{R}^{d_{w}}$, for some $0\leq\ell_{0}<m_{0}-1$. \end{enumerate} Then $\ensuremath{\mathbb{E}}\smlnorm{g(w_{0}^{\pm})}<\infty$, and on $D[0,1]$, \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(w_{t})\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\int_{0}^{\lambda}\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} where \begin{equation} w_{0}^{\pm}=(I_{d_{w}}-A^{\pm})^{-1}c^{\pm}+\sum_{\ell=0}^{\infty}(A^{\pm})^{\ell}B^{\pm}v_{-\ell}. \end{equation} Moreover, \begin{equation} \frac{1}{n^{3/2}}\sum_{t=1}^{\smlfloor{n\lambda}}[g(w_{t})\otimes z_{t}^{d}]\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\otimes\int_{0}^{\lambda}Z(\mu)\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} jointly with $U_{n}\ensuremath{\rightsquigarrow} U$.

Limiting distribution and consistency

Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as

equation[equation omitted — 100 chars of source]

where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to

equation[equation omitted — 81 chars of source]

ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for

align[align omitted — 224 chars of source]

This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.

To state the limiting distribution of the test statistic, define

equation[equation omitted — 86 chars of source]

where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process

equation[equation omitted — 17,865 chars of source]

Introduction

For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).

In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.

The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form

equation[equation omitted — 82 chars of source]

where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.

There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per

equation[equation omitted — 187 chars of source]

where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.

In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.

notation*$e_{m,i}$ denotes the $i$th column of the $m\times m$ identity matrix $I_{m}$; when $m$ is clear from the context, we write this simply as $e_{i}$. In a statement such as $f(a^{\pm},b^{\pm})=0$, the notation `$\pm$' signifies that both $f(a^{+},b^{+})=0$ and $f(a^{-},b^{-})=0$ hold; similarly, `$a^{\pm}\in A$' denotes that both $a^{+}$ and $a^{-}$ are elements of $A$. All limits are taken as $n\ensuremath{\rightarrow}\infty$ unless otherwise stated. $\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}$ and $\ensuremath{\rightsquigarrow}$ respectively denote convergence in probability and in distribution (weak convergence). We write `$X_{n}(\lambda)\ensuremath{\rightsquigarrow} X(\lambda)$ on $D_{\mathbb{R}^{m}}[0,1]$' to denote that $\{X_{n}\}$ converges weakly to $X$, where these are considered as random elements of $D_{\mathbb{R}^{m}}[0,1]$, the space of cadlag functions $[0,1]\ensuremath{\rightarrow}\mathbb{R}^{m}$, equipped with the uniform topology; we denote this as $D[0,1]$ whenever the value of $m$ is clear from the context. $\smlnorm{\cdot}$ denotes the Euclidean norm on $\mathbb{R}^{m}$, and the matrix norm that it induces. For $X$ a random vector and $p\geq1$, $\smlnorm X_{p}\coloneqq(\ensuremath{\mathbb{E}}\smlnorm X^{p})^{1/p}$. $C$, $C_{1}$, etc., denote generic constants that may take different values at different places of the same proof.

Model: the censored and kinked SVAR

Framework

We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining

align[align omitted — 111 chars of source]

we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow

equation[equation omitted — 193 chars of source]

or, more compactly,

equation[equation omitted — 100 chars of source]

where

align*[align* omitted — 158 chars of source]

for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.

We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq

bmatrix[bmatrix omitted — 55 chars of source]

=

bmatrix[bmatrix omitted — 118 chars of source]

, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.

{{{{\scalefont{0.76}DGP}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\{(y_{t},x_{t})\}$ are generated according to (ref)--(ref) with $b=0$, with (possibly random) initial values $(y_{i},x_{i})$, for $i\in\{-k+1,\ldots,0\}$; • $\operatorname{sgn}(\det\Phi_{0}^{+})=\operatorname{sgn}(\det\Phi_{0}^{-})\neq0$. • $\Phi_{0,xx}$ is invertible, and \[ \operatorname{sgn}\{\phi_{0,yy}^{+}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{+}\}=\operatorname{sgn}\{\phi_{0,yy}^{-}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{-}\}>0. \]$\{u_{t}\}_{t\in\mathbb{Z}}$ is an i.i.d.\ sequence in $\mathbb{R}^{p}$ with $\ensuremath{\mathbb{E}} u_{t}=0$, $\ensuremath{\mathbb{E}} u_{t}u_{t}^{\top}=\Sigma_{u}$ positive definite, and $\smlnorm{u_{t}}_{2+\delta_{u}}<\infty$ for some $\delta_{u}>0$. \end{enumerate}

As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that

equation[equation omitted — 141 chars of source]

on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).

Canonical form

In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if

equation[equation omitted — 119 chars of source]

While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables

equation[equation omitted — 401 chars of source]

where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting

equation[equation omitted — 275 chars of source]

for $\mathfrak{z}\in\mathbb{C}$, where

equation[equation omitted — 127 chars of source]

we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).

To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.

{{{{\scalefont{0.76}DGP$^{\ast}$}}}}

assumption$\{(y_{t},x_{t})\}$ are generated by a canonical CKSVAR, i.e.\ (ref) holds with $\Phi_{0}=[\phi_{0}^{+},\phi_{0}^{-},\Phi^{x}]=\Ican[p]$, so that (ref) may be equivalently written as \begin{equation} \begin{bmatrix}y_{t}\\ x_{t} \end{bmatrix}=c+\sum_{i=1}^{k}\begin{bmatrix}\phi_{i}^{+} & \phi_{i}^{-} & \Phi_{i}^{x}\end{bmatrix}\begin{bmatrix}y_{t-i}^{+}\\ y_{t-i}^{-}\\ x_{t-i} \end{bmatrix}+u_{t}. \end{equation}

The cointegrated CKSVAR

DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq

bmatrix[bmatrix omitted — 62 chars of source]

, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define

align*[align* omitted — 107 chars of source]

{{{{\scalefont{0.76}CVAR}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\det\Phi^{\pm}(\mathfrak{z})$ has $q^{\pm}\in\{1,\ldots,p\}$ roots at real unity, and all others outside the unit circle; and • $\operatorname{rk}\Pi^{\pm}=r^{\pm}=p-q^{\pm}$. \end{enumerate}

The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}

bmatrix[bmatrix omitted — 35 chars of source]

=\alpha

bmatrix[bmatrix omitted — 47 chars of source]

\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of

gather[gather omitted — 420 chars of source]

for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as

align[align omitted — 388 chars of source]

where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and

equation[equation omitted — 168 chars of source]

Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.

{{{{\scalefont{0.76}CO{{(ii)}}}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}},itemsep=1pt,topsep=2pt] • $r^{+}=r^{-}=\operatorname{rk}\Pi^{x}=r$, for some $r\in\{0,1,\ldots,p-1\}$. • $\rho_{{\scriptstyle \mathrm{JSR}}}(\{I+\tilde{\b{\beta}}(+1)^{\top}\tilde{\b{\alpha}},I+\tilde{\b{\beta}}(-1)^{\top}\tilde{\b{\alpha}}\})<1$. • $\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;+1)\beta_{\perp}(+1)=\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;-1)\beta_{\perp}(-1)\neq0$. • \begin{enumerate}[label={{{\scalefont{0.76}\alph*.}}}, ref={{{\scalefont{0.76}.\alph*}}}, leftmargin=0.40cm] • $\beta(y_{t})^{\top}z_{t}$, and $\Delta z_{t}$ have uniformly bounded $2+\delta_{u}$ moments, for $t\in\{-k+1,\ldots,0\}$. • $n^{-1/2}z_{0}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}{\cal Z}_{0}=[\begin{smallmatrix}{\cal Y}_{0}\\ {\cal X}_{0} \end{smallmatrix}]$, where ${\cal Z}_{0}$ is non-random, and satisfies $\beta(\mathcal{Y}_{0})^{\top}{\cal Z}_{0}=0$. \end{enumerate} \end{enumerate}

Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that

equation[equation omitted — 71 chars of source]

or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.

Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per

{{{{\scalefont{0.76}DET}}}}

assumption$c\in\operatorname{sp}\Pi^{+}\ensuremath{\cap}\operatorname{sp}\Pi^{-}$.

Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that

equation[equation omitted — 297 chars of source]

where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).

On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.

Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:

align*[align* omitted — 115 chars of source]

But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.

The modified Breitung (2002) test

Fundamental ideas

We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)

The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)

propSuppose that $\{w_{n,t}\}_{t=1}^{n}$ is a triangular array, taking values in $\mathbb{R}^{d_{w}}$, such that \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}\ensuremath{\rightsquigarrow}\int_{0}^{\lambda}\begin{bmatrix}\mathbb{W}(s)\\ 0_{d_{w}-\ell} \end{bmatrix}\ensuremath{\,\ensuremath{\mathrm{d}}} s\eqqcolon\begin{bmatrix}\mathbb{V}(\lambda)\\ 0_{d_{w}-\ell} \end{bmatrix} \end{equation} on $D_{\mathbb{R}^{d_{w}}}[0,1]$, where $\mathbb{W}$ is a random element of $D_{\mathbb{R}^{\ell}}[0,1]$ and \begin{align} \frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top} & \ensuremath{\rightsquigarrow}\begin{bmatrix}\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s & 0\\ 0 & \Omega \end{bmatrix} \end{align} where $\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$, $\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$ and $\Omega\in\mathbb{R}^{(d_{w}-\ell)\times(d_{w}-\ell)}$ are a.s.\ positive definite. Let $\{\lambda_{n,i}\}_{i=1}^{d_{w}}$ denote the solutions to \[ \det(\lambda\mathbb{B}_{n}-\mathbb{A}_{n})=0 \] ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,d_{w}}$, for \begin{align*} \mathbb{A}_{n} & \coloneqq\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}, & \mathbb{B}_{n} & \coloneqq\sum_{t=1}^{n}\sum_{i=1}^{t}w_{n,i}\sum_{j=1}^{t}w_{n,j}^{\top}. \end{align*} Then \begin{enumerate} • if $\ell_{0}=\ell$, \[ n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\rightsquigarrow}\operatorname{tr}\left[\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\left(\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\right)^{-1}\right]; \] • if $\ell_{0}>\ell$, $n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}\infty$. \end{enumerate}

To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq

bmatrix[bmatrix omitted — 57 chars of source]

z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).

By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.

Extension to nonlinearly cointegrated series

Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)

To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq

bmatrix[bmatrix omitted — 43 chars of source]

=\left[

array[array omitted — 110 chars of source]

\right]

bmatrix[bmatrix omitted — 27 chars of source]

=S_{p}(y_{t})z_{t} \] via

equation[equation omitted — 410 chars of source]

from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.

In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix

equation[equation omitted — 163 chars of source]

where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).

Thus the linear transformation

equation[equation omitted — 468 chars of source]

exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now

equation[equation omitted — 280 chars of source]

where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.

To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.

LLN for regime-switching processes

To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation

equation[equation omitted — 123 chars of source]

where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.

Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]

This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of

align*[align* omitted — 571 chars of source]

where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at

align*[align* omitted — 342 chars of source]

which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.

Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).

{{{{\scalefont{0.76}DET$^{\prime}$}}}}

assumption$e_{1}^{\top}P_{\beta_{\perp}}(+1)c=0$.

The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[

bmatrix[bmatrix omitted — 35 chars of source]

=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.

thmSuppose (ref), (ref), (ref) and (ref) hold. Let $\{A_{t}\}$, $\{B_{t}\}$ and $\{c_{t}\}$ be random sequences adapted to $\{\mathcal{F}_{t}\}$, respectively taking values in $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$, where $t\in\mathbb{Z}$. Suppose $\{v_{t}\}$ is i.i.d\ with $\ensuremath{\mathbb{E}} v_{t}=0$, and that $\{w_{t}\}$ satisfies \begin{equation} w_{t}=c_{t}+A_{t}w_{t-1}+B_{t}v_{t} \end{equation} for $t\geq-k$ and some given (random) $w_{-k}$ (with $w_{t}\coloneqq0$ for all $t\leq-k-1$); and: \begin{enumerate}[itemsep=2pt,topsep=3pt] • $A_{t}\in\mset A$, $B_{t}\in\mset B$ and $c_{t}\in{\cal C}$ for all $t\in\ensuremath{\mathbb{N}}$, where $\mset A$, $\mset B$ and ${\cal C}$ are bounded subsets of $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$ respectively, and $\rho_{{\scriptstyle \mathrm{JSR}}}(\mset A)<1$; • there exist $A^{\pm}\in\mset A$, $B^{\pm}\in\mset B$ and $c^{\pm}\in\mset C$ such that \begin{align*} y_{t-1}>0 and y_{t}>0 & \implies A_{t}=A^{+},\ B_{t}=B^{+},\ c_{t}=c^{+},\\ y_{t-1}<0 and y_{t}<0 & \implies A_{t}=A^{-},\ B_{t}=B^{-},\ c_{t}=c^{-}; \end{align*} • $m_{0}\geq1$ is such that $\smlnorm{w_{0}}_{m_{0}}+\smlnorm{v_{0}}_{m_{0}}<\infty$. • $g:\mathbb{R}^{d_{w}}\ensuremath{\rightarrow}\mathbb{R}^{d_{g}}$ is a continuous function satisfying \begin{equation} \smlnorm{g(w)-g(w^{\prime})}\leq C(1+\smlnorm w^{\ell_{0}}+\smlnorm{w^{\prime}}^{\ell_{0}})\smlnorm{w-w^{\prime}} \end{equation} for all $w,w^{\prime}\in\mathbb{R}^{d_{w}}$, for some $0\leq\ell_{0}<m_{0}-1$. \end{enumerate} Then $\ensuremath{\mathbb{E}}\smlnorm{g(w_{0}^{\pm})}<\infty$, and on $D[0,1]$, \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(w_{t})\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\int_{0}^{\lambda}\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} where \begin{equation} w_{0}^{\pm}=(I_{d_{w}}-A^{\pm})^{-1}c^{\pm}+\sum_{\ell=0}^{\infty}(A^{\pm})^{\ell}B^{\pm}v_{-\ell}. \end{equation} Moreover, \begin{equation} \frac{1}{n^{3/2}}\sum_{t=1}^{\smlfloor{n\lambda}}[g(w_{t})\otimes z_{t}^{d}]\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\otimes\int_{0}^{\lambda}Z(\mu)\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} jointly with $U_{n}\ensuremath{\rightsquigarrow} U$.

Limiting distribution and consistency

Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as

equation[equation omitted — 100 chars of source]

where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to

equation[equation omitted — 81 chars of source]

ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for

align[align omitted — 224 chars of source]

This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.

To state the limiting distribution of the test statistic, define

equation[equation omitted — 86 chars of source]

where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process

equation[equation omitted — 17,910 chars of source]

Introduction

For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).

In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.

The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form

equation[equation omitted — 82 chars of source]

where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.

There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per

equation[equation omitted — 187 chars of source]

where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.

In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.

notation*$e_{m,i}$ denotes the $i$th column of the $m\times m$ identity matrix $I_{m}$; when $m$ is clear from the context, we write this simply as $e_{i}$. In a statement such as $f(a^{\pm},b^{\pm})=0$, the notation `$\pm$' signifies that both $f(a^{+},b^{+})=0$ and $f(a^{-},b^{-})=0$ hold; similarly, `$a^{\pm}\in A$' denotes that both $a^{+}$ and $a^{-}$ are elements of $A$. All limits are taken as $n\ensuremath{\rightarrow}\infty$ unless otherwise stated. $\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}$ and $\ensuremath{\rightsquigarrow}$ respectively denote convergence in probability and in distribution (weak convergence). We write `$X_{n}(\lambda)\ensuremath{\rightsquigarrow} X(\lambda)$ on $D_{\mathbb{R}^{m}}[0,1]$' to denote that $\{X_{n}\}$ converges weakly to $X$, where these are considered as random elements of $D_{\mathbb{R}^{m}}[0,1]$, the space of cadlag functions $[0,1]\ensuremath{\rightarrow}\mathbb{R}^{m}$, equipped with the uniform topology; we denote this as $D[0,1]$ whenever the value of $m$ is clear from the context. $\smlnorm{\cdot}$ denotes the Euclidean norm on $\mathbb{R}^{m}$, and the matrix norm that it induces. For $X$ a random vector and $p\geq1$, $\smlnorm X_{p}\coloneqq(\ensuremath{\mathbb{E}}\smlnorm X^{p})^{1/p}$. $C$, $C_{1}$, etc., denote generic constants that may take different values at different places of the same proof.

Model: the censored and kinked SVAR

Framework

We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining

align[align omitted — 111 chars of source]

we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow

equation[equation omitted — 193 chars of source]

or, more compactly,

equation[equation omitted — 100 chars of source]

where

align*[align* omitted — 158 chars of source]

for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.

We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq

bmatrix[bmatrix omitted — 55 chars of source]

=

bmatrix[bmatrix omitted — 118 chars of source]

, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.

{{{{\scalefont{0.76}DGP}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\{(y_{t},x_{t})\}$ are generated according to (ref)--(ref) with $b=0$, with (possibly random) initial values $(y_{i},x_{i})$, for $i\in\{-k+1,\ldots,0\}$; • $\operatorname{sgn}(\det\Phi_{0}^{+})=\operatorname{sgn}(\det\Phi_{0}^{-})\neq0$. • $\Phi_{0,xx}$ is invertible, and \[ \operatorname{sgn}\{\phi_{0,yy}^{+}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{+}\}=\operatorname{sgn}\{\phi_{0,yy}^{-}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{-}\}>0. \]$\{u_{t}\}_{t\in\mathbb{Z}}$ is an i.i.d.\ sequence in $\mathbb{R}^{p}$ with $\ensuremath{\mathbb{E}} u_{t}=0$, $\ensuremath{\mathbb{E}} u_{t}u_{t}^{\top}=\Sigma_{u}$ positive definite, and $\smlnorm{u_{t}}_{2+\delta_{u}}<\infty$ for some $\delta_{u}>0$. \end{enumerate}

As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that

equation[equation omitted — 141 chars of source]

on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).

Canonical form

In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if

equation[equation omitted — 119 chars of source]

While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables

equation[equation omitted — 401 chars of source]

where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting

equation[equation omitted — 275 chars of source]

for $\mathfrak{z}\in\mathbb{C}$, where

equation[equation omitted — 127 chars of source]

we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).

To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.

{{{{\scalefont{0.76}DGP$^{\ast}$}}}}

assumption$\{(y_{t},x_{t})\}$ are generated by a canonical CKSVAR, i.e.\ (ref) holds with $\Phi_{0}=[\phi_{0}^{+},\phi_{0}^{-},\Phi^{x}]=\Ican[p]$, so that (ref) may be equivalently written as \begin{equation} \begin{bmatrix}y_{t}\\ x_{t} \end{bmatrix}=c+\sum_{i=1}^{k}\begin{bmatrix}\phi_{i}^{+} & \phi_{i}^{-} & \Phi_{i}^{x}\end{bmatrix}\begin{bmatrix}y_{t-i}^{+}\\ y_{t-i}^{-}\\ x_{t-i} \end{bmatrix}+u_{t}. \end{equation}

The cointegrated CKSVAR

DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq

bmatrix[bmatrix omitted — 62 chars of source]

, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define

align*[align* omitted — 107 chars of source]

{{{{\scalefont{0.76}CVAR}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\det\Phi^{\pm}(\mathfrak{z})$ has $q^{\pm}\in\{1,\ldots,p\}$ roots at real unity, and all others outside the unit circle; and • $\operatorname{rk}\Pi^{\pm}=r^{\pm}=p-q^{\pm}$. \end{enumerate}

The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}

bmatrix[bmatrix omitted — 35 chars of source]

=\alpha

bmatrix[bmatrix omitted — 47 chars of source]

\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of

gather[gather omitted — 420 chars of source]

for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as

align[align omitted — 388 chars of source]

where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and

equation[equation omitted — 168 chars of source]

Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.

{{{{\scalefont{0.76}CO{{(ii)}}}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}},itemsep=1pt,topsep=2pt] • $r^{+}=r^{-}=\operatorname{rk}\Pi^{x}=r$, for some $r\in\{0,1,\ldots,p-1\}$. • $\rho_{{\scriptstyle \mathrm{JSR}}}(\{I+\tilde{\b{\beta}}(+1)^{\top}\tilde{\b{\alpha}},I+\tilde{\b{\beta}}(-1)^{\top}\tilde{\b{\alpha}}\})<1$. • $\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;+1)\beta_{\perp}(+1)=\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;-1)\beta_{\perp}(-1)\neq0$. • \begin{enumerate}[label={{{\scalefont{0.76}\alph*.}}}, ref={{{\scalefont{0.76}.\alph*}}}, leftmargin=0.40cm] • $\beta(y_{t})^{\top}z_{t}$, and $\Delta z_{t}$ have uniformly bounded $2+\delta_{u}$ moments, for $t\in\{-k+1,\ldots,0\}$. • $n^{-1/2}z_{0}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}{\cal Z}_{0}=[\begin{smallmatrix}{\cal Y}_{0}\\ {\cal X}_{0} \end{smallmatrix}]$, where ${\cal Z}_{0}$ is non-random, and satisfies $\beta(\mathcal{Y}_{0})^{\top}{\cal Z}_{0}=0$. \end{enumerate} \end{enumerate}

Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that

equation[equation omitted — 71 chars of source]

or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.

Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per

{{{{\scalefont{0.76}DET}}}}

assumption$c\in\operatorname{sp}\Pi^{+}\ensuremath{\cap}\operatorname{sp}\Pi^{-}$.

Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that

equation[equation omitted — 297 chars of source]

where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).

On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.

Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:

align*[align* omitted — 115 chars of source]

But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.

The modified Breitung (2002) test

Fundamental ideas

We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)

The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)

propSuppose that $\{w_{n,t}\}_{t=1}^{n}$ is a triangular array, taking values in $\mathbb{R}^{d_{w}}$, such that \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}\ensuremath{\rightsquigarrow}\int_{0}^{\lambda}\begin{bmatrix}\mathbb{W}(s)\\ 0_{d_{w}-\ell} \end{bmatrix}\ensuremath{\,\ensuremath{\mathrm{d}}} s\eqqcolon\begin{bmatrix}\mathbb{V}(\lambda)\\ 0_{d_{w}-\ell} \end{bmatrix} \end{equation} on $D_{\mathbb{R}^{d_{w}}}[0,1]$, where $\mathbb{W}$ is a random element of $D_{\mathbb{R}^{\ell}}[0,1]$ and \begin{align} \frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top} & \ensuremath{\rightsquigarrow}\begin{bmatrix}\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s & 0\\ 0 & \Omega \end{bmatrix} \end{align} where $\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$, $\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$ and $\Omega\in\mathbb{R}^{(d_{w}-\ell)\times(d_{w}-\ell)}$ are a.s.\ positive definite. Let $\{\lambda_{n,i}\}_{i=1}^{d_{w}}$ denote the solutions to \[ \det(\lambda\mathbb{B}_{n}-\mathbb{A}_{n})=0 \] ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,d_{w}}$, for \begin{align*} \mathbb{A}_{n} & \coloneqq\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}, & \mathbb{B}_{n} & \coloneqq\sum_{t=1}^{n}\sum_{i=1}^{t}w_{n,i}\sum_{j=1}^{t}w_{n,j}^{\top}. \end{align*} Then \begin{enumerate} • if $\ell_{0}=\ell$, \[ n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\rightsquigarrow}\operatorname{tr}\left[\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\left(\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\right)^{-1}\right]; \] • if $\ell_{0}>\ell$, $n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}\infty$. \end{enumerate}

To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq

bmatrix[bmatrix omitted — 57 chars of source]

z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).

By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.

Extension to nonlinearly cointegrated series

Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)

To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq

bmatrix[bmatrix omitted — 43 chars of source]

=\left[

array[array omitted — 110 chars of source]

\right]

bmatrix[bmatrix omitted — 27 chars of source]

=S_{p}(y_{t})z_{t} \] via

equation[equation omitted — 410 chars of source]

from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.

In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix

equation[equation omitted — 163 chars of source]

where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).

Thus the linear transformation

equation[equation omitted — 468 chars of source]

exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now

equation[equation omitted — 280 chars of source]

where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.

To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.

LLN for regime-switching processes

To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation

equation[equation omitted — 123 chars of source]

where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.

Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]

This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of

align*[align* omitted — 571 chars of source]

where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at

align*[align* omitted — 342 chars of source]

which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.

Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).

{{{{\scalefont{0.76}DET$^{\prime}$}}}}

assumption$e_{1}^{\top}P_{\beta_{\perp}}(+1)c=0$.

The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[

bmatrix[bmatrix omitted — 35 chars of source]

=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.

thmSuppose (ref), (ref), (ref) and (ref) hold. Let $\{A_{t}\}$, $\{B_{t}\}$ and $\{c_{t}\}$ be random sequences adapted to $\{\mathcal{F}_{t}\}$, respectively taking values in $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$, where $t\in\mathbb{Z}$. Suppose $\{v_{t}\}$ is i.i.d\ with $\ensuremath{\mathbb{E}} v_{t}=0$, and that $\{w_{t}\}$ satisfies \begin{equation} w_{t}=c_{t}+A_{t}w_{t-1}+B_{t}v_{t} \end{equation} for $t\geq-k$ and some given (random) $w_{-k}$ (with $w_{t}\coloneqq0$ for all $t\leq-k-1$); and: \begin{enumerate}[itemsep=2pt,topsep=3pt] • $A_{t}\in\mset A$, $B_{t}\in\mset B$ and $c_{t}\in{\cal C}$ for all $t\in\ensuremath{\mathbb{N}}$, where $\mset A$, $\mset B$ and ${\cal C}$ are bounded subsets of $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$ respectively, and $\rho_{{\scriptstyle \mathrm{JSR}}}(\mset A)<1$; • there exist $A^{\pm}\in\mset A$, $B^{\pm}\in\mset B$ and $c^{\pm}\in\mset C$ such that \begin{align*} y_{t-1}>0 and y_{t}>0 & \implies A_{t}=A^{+},\ B_{t}=B^{+},\ c_{t}=c^{+},\\ y_{t-1}<0 and y_{t}<0 & \implies A_{t}=A^{-},\ B_{t}=B^{-},\ c_{t}=c^{-}; \end{align*} • $m_{0}\geq1$ is such that $\smlnorm{w_{0}}_{m_{0}}+\smlnorm{v_{0}}_{m_{0}}<\infty$. • $g:\mathbb{R}^{d_{w}}\ensuremath{\rightarrow}\mathbb{R}^{d_{g}}$ is a continuous function satisfying \begin{equation} \smlnorm{g(w)-g(w^{\prime})}\leq C(1+\smlnorm w^{\ell_{0}}+\smlnorm{w^{\prime}}^{\ell_{0}})\smlnorm{w-w^{\prime}} \end{equation} for all $w,w^{\prime}\in\mathbb{R}^{d_{w}}$, for some $0\leq\ell_{0}<m_{0}-1$. \end{enumerate} Then $\ensuremath{\mathbb{E}}\smlnorm{g(w_{0}^{\pm})}<\infty$, and on $D[0,1]$, \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(w_{t})\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\int_{0}^{\lambda}\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} where \begin{equation} w_{0}^{\pm}=(I_{d_{w}}-A^{\pm})^{-1}c^{\pm}+\sum_{\ell=0}^{\infty}(A^{\pm})^{\ell}B^{\pm}v_{-\ell}. \end{equation} Moreover, \begin{equation} \frac{1}{n^{3/2}}\sum_{t=1}^{\smlfloor{n\lambda}}[g(w_{t})\otimes z_{t}^{d}]\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\otimes\int_{0}^{\lambda}Z(\mu)\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} jointly with $U_{n}\ensuremath{\rightsquigarrow} U$.

Limiting distribution and consistency

Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as

equation[equation omitted — 100 chars of source]

where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to

equation[equation omitted — 81 chars of source]

ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for

align[align omitted — 224 chars of source]

This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.

To state the limiting distribution of the test statistic, define

equation[equation omitted — 86 chars of source]

where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process

equation[equation omitted — 17,932 chars of source]

Introduction

For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).

In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.

The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form

equation[equation omitted — 82 chars of source]

where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.

There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per

equation[equation omitted — 187 chars of source]

where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.

In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.

notation*$e_{m,i}$ denotes the $i$th column of the $m\times m$ identity matrix $I_{m}$; when $m$ is clear from the context, we write this simply as $e_{i}$. In a statement such as $f(a^{\pm},b^{\pm})=0$, the notation `$\pm$' signifies that both $f(a^{+},b^{+})=0$ and $f(a^{-},b^{-})=0$ hold; similarly, `$a^{\pm}\in A$' denotes that both $a^{+}$ and $a^{-}$ are elements of $A$. All limits are taken as $n\ensuremath{\rightarrow}\infty$ unless otherwise stated. $\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}$ and $\ensuremath{\rightsquigarrow}$ respectively denote convergence in probability and in distribution (weak convergence). We write `$X_{n}(\lambda)\ensuremath{\rightsquigarrow} X(\lambda)$ on $D_{\mathbb{R}^{m}}[0,1]$' to denote that $\{X_{n}\}$ converges weakly to $X$, where these are considered as random elements of $D_{\mathbb{R}^{m}}[0,1]$, the space of cadlag functions $[0,1]\ensuremath{\rightarrow}\mathbb{R}^{m}$, equipped with the uniform topology; we denote this as $D[0,1]$ whenever the value of $m$ is clear from the context. $\smlnorm{\cdot}$ denotes the Euclidean norm on $\mathbb{R}^{m}$, and the matrix norm that it induces. For $X$ a random vector and $p\geq1$, $\smlnorm X_{p}\coloneqq(\ensuremath{\mathbb{E}}\smlnorm X^{p})^{1/p}$. $C$, $C_{1}$, etc., denote generic constants that may take different values at different places of the same proof.

Model: the censored and kinked SVAR

Framework

We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining

align[align omitted — 111 chars of source]

we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow

equation[equation omitted — 193 chars of source]

or, more compactly,

equation[equation omitted — 100 chars of source]

where

align*[align* omitted — 158 chars of source]

for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.

We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq

bmatrix[bmatrix omitted — 55 chars of source]

=

bmatrix[bmatrix omitted — 118 chars of source]

, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.

{{{{\scalefont{0.76}DGP}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\{(y_{t},x_{t})\}$ are generated according to (ref)--(ref) with $b=0$, with (possibly random) initial values $(y_{i},x_{i})$, for $i\in\{-k+1,\ldots,0\}$; • $\operatorname{sgn}(\det\Phi_{0}^{+})=\operatorname{sgn}(\det\Phi_{0}^{-})\neq0$. • $\Phi_{0,xx}$ is invertible, and \[ \operatorname{sgn}\{\phi_{0,yy}^{+}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{+}\}=\operatorname{sgn}\{\phi_{0,yy}^{-}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{-}\}>0. \]$\{u_{t}\}_{t\in\mathbb{Z}}$ is an i.i.d.\ sequence in $\mathbb{R}^{p}$ with $\ensuremath{\mathbb{E}} u_{t}=0$, $\ensuremath{\mathbb{E}} u_{t}u_{t}^{\top}=\Sigma_{u}$ positive definite, and $\smlnorm{u_{t}}_{2+\delta_{u}}<\infty$ for some $\delta_{u}>0$. \end{enumerate}

As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that

equation[equation omitted — 141 chars of source]

on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).

Canonical form

In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if

equation[equation omitted — 119 chars of source]

While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables

equation[equation omitted — 401 chars of source]

where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting

equation[equation omitted — 275 chars of source]

for $\mathfrak{z}\in\mathbb{C}$, where

equation[equation omitted — 127 chars of source]

we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).

To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.

{{{{\scalefont{0.76}DGP$^{\ast}$}}}}

assumption$\{(y_{t},x_{t})\}$ are generated by a canonical CKSVAR, i.e.\ (ref) holds with $\Phi_{0}=[\phi_{0}^{+},\phi_{0}^{-},\Phi^{x}]=\Ican[p]$, so that (ref) may be equivalently written as \begin{equation} \begin{bmatrix}y_{t}\\ x_{t} \end{bmatrix}=c+\sum_{i=1}^{k}\begin{bmatrix}\phi_{i}^{+} & \phi_{i}^{-} & \Phi_{i}^{x}\end{bmatrix}\begin{bmatrix}y_{t-i}^{+}\\ y_{t-i}^{-}\\ x_{t-i} \end{bmatrix}+u_{t}. \end{equation}

The cointegrated CKSVAR

DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq

bmatrix[bmatrix omitted — 62 chars of source]

, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define

align*[align* omitted — 107 chars of source]

{{{{\scalefont{0.76}CVAR}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\det\Phi^{\pm}(\mathfrak{z})$ has $q^{\pm}\in\{1,\ldots,p\}$ roots at real unity, and all others outside the unit circle; and • $\operatorname{rk}\Pi^{\pm}=r^{\pm}=p-q^{\pm}$. \end{enumerate}

The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}

bmatrix[bmatrix omitted — 35 chars of source]

=\alpha

bmatrix[bmatrix omitted — 47 chars of source]

\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of

gather[gather omitted — 420 chars of source]

for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as

align[align omitted — 388 chars of source]

where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and

equation[equation omitted — 168 chars of source]

Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.

{{{{\scalefont{0.76}CO{{(ii)}}}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}},itemsep=1pt,topsep=2pt] • $r^{+}=r^{-}=\operatorname{rk}\Pi^{x}=r$, for some $r\in\{0,1,\ldots,p-1\}$. • $\rho_{{\scriptstyle \mathrm{JSR}}}(\{I+\tilde{\b{\beta}}(+1)^{\top}\tilde{\b{\alpha}},I+\tilde{\b{\beta}}(-1)^{\top}\tilde{\b{\alpha}}\})<1$. • $\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;+1)\beta_{\perp}(+1)=\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;-1)\beta_{\perp}(-1)\neq0$. • \begin{enumerate}[label={{{\scalefont{0.76}\alph*.}}}, ref={{{\scalefont{0.76}.\alph*}}}, leftmargin=0.40cm] • $\beta(y_{t})^{\top}z_{t}$, and $\Delta z_{t}$ have uniformly bounded $2+\delta_{u}$ moments, for $t\in\{-k+1,\ldots,0\}$. • $n^{-1/2}z_{0}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}{\cal Z}_{0}=[\begin{smallmatrix}{\cal Y}_{0}\\ {\cal X}_{0} \end{smallmatrix}]$, where ${\cal Z}_{0}$ is non-random, and satisfies $\beta(\mathcal{Y}_{0})^{\top}{\cal Z}_{0}=0$. \end{enumerate} \end{enumerate}

Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that

equation[equation omitted — 71 chars of source]

or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.

Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per

{{{{\scalefont{0.76}DET}}}}

assumption$c\in\operatorname{sp}\Pi^{+}\ensuremath{\cap}\operatorname{sp}\Pi^{-}$.

Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that

equation[equation omitted — 297 chars of source]

where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).

On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.

Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:

align*[align* omitted — 115 chars of source]

But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.

The modified Breitung (2002) test

Fundamental ideas

We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)

The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)

propSuppose that $\{w_{n,t}\}_{t=1}^{n}$ is a triangular array, taking values in $\mathbb{R}^{d_{w}}$, such that \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}\ensuremath{\rightsquigarrow}\int_{0}^{\lambda}\begin{bmatrix}\mathbb{W}(s)\\ 0_{d_{w}-\ell} \end{bmatrix}\ensuremath{\,\ensuremath{\mathrm{d}}} s\eqqcolon\begin{bmatrix}\mathbb{V}(\lambda)\\ 0_{d_{w}-\ell} \end{bmatrix} \end{equation} on $D_{\mathbb{R}^{d_{w}}}[0,1]$, where $\mathbb{W}$ is a random element of $D_{\mathbb{R}^{\ell}}[0,1]$ and \begin{align} \frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top} & \ensuremath{\rightsquigarrow}\begin{bmatrix}\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s & 0\\ 0 & \Omega \end{bmatrix} \end{align} where $\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$, $\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$ and $\Omega\in\mathbb{R}^{(d_{w}-\ell)\times(d_{w}-\ell)}$ are a.s.\ positive definite. Let $\{\lambda_{n,i}\}_{i=1}^{d_{w}}$ denote the solutions to \[ \det(\lambda\mathbb{B}_{n}-\mathbb{A}_{n})=0 \] ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,d_{w}}$, for \begin{align*} \mathbb{A}_{n} & \coloneqq\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}, & \mathbb{B}_{n} & \coloneqq\sum_{t=1}^{n}\sum_{i=1}^{t}w_{n,i}\sum_{j=1}^{t}w_{n,j}^{\top}. \end{align*} Then \begin{enumerate} • if $\ell_{0}=\ell$, \[ n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\rightsquigarrow}\operatorname{tr}\left[\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\left(\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\right)^{-1}\right]; \] • if $\ell_{0}>\ell$, $n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}\infty$. \end{enumerate}

To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq

bmatrix[bmatrix omitted — 57 chars of source]

z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).

By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.

Extension to nonlinearly cointegrated series

Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)

To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq

bmatrix[bmatrix omitted — 43 chars of source]

=\left[

array[array omitted — 110 chars of source]

\right]

bmatrix[bmatrix omitted — 27 chars of source]

=S_{p}(y_{t})z_{t} \] via

equation[equation omitted — 410 chars of source]

from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.

In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix

equation[equation omitted — 163 chars of source]

where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).

Thus the linear transformation

equation[equation omitted — 468 chars of source]

exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now

equation[equation omitted — 280 chars of source]

where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.

To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.

LLN for regime-switching processes

To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation

equation[equation omitted — 123 chars of source]

where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.

Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]

This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of

align*[align* omitted — 571 chars of source]

where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at

align*[align* omitted — 342 chars of source]

which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.

Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).

{{{{\scalefont{0.76}DET$^{\prime}$}}}}

assumption$e_{1}^{\top}P_{\beta_{\perp}}(+1)c=0$.

The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[

bmatrix[bmatrix omitted — 35 chars of source]

=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.

thmSuppose (ref), (ref), (ref) and (ref) hold. Let $\{A_{t}\}$, $\{B_{t}\}$ and $\{c_{t}\}$ be random sequences adapted to $\{\mathcal{F}_{t}\}$, respectively taking values in $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$, where $t\in\mathbb{Z}$. Suppose $\{v_{t}\}$ is i.i.d\ with $\ensuremath{\mathbb{E}} v_{t}=0$, and that $\{w_{t}\}$ satisfies \begin{equation} w_{t}=c_{t}+A_{t}w_{t-1}+B_{t}v_{t} \end{equation} for $t\geq-k$ and some given (random) $w_{-k}$ (with $w_{t}\coloneqq0$ for all $t\leq-k-1$); and: \begin{enumerate}[itemsep=2pt,topsep=3pt] • $A_{t}\in\mset A$, $B_{t}\in\mset B$ and $c_{t}\in{\cal C}$ for all $t\in\ensuremath{\mathbb{N}}$, where $\mset A$, $\mset B$ and ${\cal C}$ are bounded subsets of $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$ respectively, and $\rho_{{\scriptstyle \mathrm{JSR}}}(\mset A)<1$; • there exist $A^{\pm}\in\mset A$, $B^{\pm}\in\mset B$ and $c^{\pm}\in\mset C$ such that \begin{align*} y_{t-1}>0 and y_{t}>0 & \implies A_{t}=A^{+},\ B_{t}=B^{+},\ c_{t}=c^{+},\\ y_{t-1}<0 and y_{t}<0 & \implies A_{t}=A^{-},\ B_{t}=B^{-},\ c_{t}=c^{-}; \end{align*} • $m_{0}\geq1$ is such that $\smlnorm{w_{0}}_{m_{0}}+\smlnorm{v_{0}}_{m_{0}}<\infty$. • $g:\mathbb{R}^{d_{w}}\ensuremath{\rightarrow}\mathbb{R}^{d_{g}}$ is a continuous function satisfying \begin{equation} \smlnorm{g(w)-g(w^{\prime})}\leq C(1+\smlnorm w^{\ell_{0}}+\smlnorm{w^{\prime}}^{\ell_{0}})\smlnorm{w-w^{\prime}} \end{equation} for all $w,w^{\prime}\in\mathbb{R}^{d_{w}}$, for some $0\leq\ell_{0}<m_{0}-1$. \end{enumerate} Then $\ensuremath{\mathbb{E}}\smlnorm{g(w_{0}^{\pm})}<\infty$, and on $D[0,1]$, \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(w_{t})\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\int_{0}^{\lambda}\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} where \begin{equation} w_{0}^{\pm}=(I_{d_{w}}-A^{\pm})^{-1}c^{\pm}+\sum_{\ell=0}^{\infty}(A^{\pm})^{\ell}B^{\pm}v_{-\ell}. \end{equation} Moreover, \begin{equation} \frac{1}{n^{3/2}}\sum_{t=1}^{\smlfloor{n\lambda}}[g(w_{t})\otimes z_{t}^{d}]\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\otimes\int_{0}^{\lambda}Z(\mu)\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} jointly with $U_{n}\ensuremath{\rightsquigarrow} U$.

Limiting distribution and consistency

Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as

equation[equation omitted — 100 chars of source]

where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to

equation[equation omitted — 81 chars of source]

ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for

align[align omitted — 224 chars of source]

This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.

To state the limiting distribution of the test statistic, define

equation[equation omitted — 86 chars of source]

where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process

equation[equation omitted — 17,835 chars of source]

Introduction

For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).

In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.

The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form

equation[equation omitted — 82 chars of source]

where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.

There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per

equation[equation omitted — 187 chars of source]

where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.

In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.

notation*$e_{m,i}$ denotes the $i$th column of the $m\times m$ identity matrix $I_{m}$; when $m$ is clear from the context, we write this simply as $e_{i}$. In a statement such as $f(a^{\pm},b^{\pm})=0$, the notation `$\pm$' signifies that both $f(a^{+},b^{+})=0$ and $f(a^{-},b^{-})=0$ hold; similarly, `$a^{\pm}\in A$' denotes that both $a^{+}$ and $a^{-}$ are elements of $A$. All limits are taken as $n\ensuremath{\rightarrow}\infty$ unless otherwise stated. $\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}$ and $\ensuremath{\rightsquigarrow}$ respectively denote convergence in probability and in distribution (weak convergence). We write `$X_{n}(\lambda)\ensuremath{\rightsquigarrow} X(\lambda)$ on $D_{\mathbb{R}^{m}}[0,1]$' to denote that $\{X_{n}\}$ converges weakly to $X$, where these are considered as random elements of $D_{\mathbb{R}^{m}}[0,1]$, the space of cadlag functions $[0,1]\ensuremath{\rightarrow}\mathbb{R}^{m}$, equipped with the uniform topology; we denote this as $D[0,1]$ whenever the value of $m$ is clear from the context. $\smlnorm{\cdot}$ denotes the Euclidean norm on $\mathbb{R}^{m}$, and the matrix norm that it induces. For $X$ a random vector and $p\geq1$, $\smlnorm X_{p}\coloneqq(\ensuremath{\mathbb{E}}\smlnorm X^{p})^{1/p}$. $C$, $C_{1}$, etc., denote generic constants that may take different values at different places of the same proof.

Model: the censored and kinked SVAR

Framework

We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining

align[align omitted — 111 chars of source]

we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow

equation[equation omitted — 193 chars of source]

or, more compactly,

equation[equation omitted — 100 chars of source]

where

align*[align* omitted — 158 chars of source]

for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.

We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq

bmatrix[bmatrix omitted — 55 chars of source]

=

bmatrix[bmatrix omitted — 118 chars of source]

, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.

{{{{\scalefont{0.76}DGP}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\{(y_{t},x_{t})\}$ are generated according to (ref)--(ref) with $b=0$, with (possibly random) initial values $(y_{i},x_{i})$, for $i\in\{-k+1,\ldots,0\}$; • $\operatorname{sgn}(\det\Phi_{0}^{+})=\operatorname{sgn}(\det\Phi_{0}^{-})\neq0$. • $\Phi_{0,xx}$ is invertible, and \[ \operatorname{sgn}\{\phi_{0,yy}^{+}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{+}\}=\operatorname{sgn}\{\phi_{0,yy}^{-}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{-}\}>0. \]$\{u_{t}\}_{t\in\mathbb{Z}}$ is an i.i.d.\ sequence in $\mathbb{R}^{p}$ with $\ensuremath{\mathbb{E}} u_{t}=0$, $\ensuremath{\mathbb{E}} u_{t}u_{t}^{\top}=\Sigma_{u}$ positive definite, and $\smlnorm{u_{t}}_{2+\delta_{u}}<\infty$ for some $\delta_{u}>0$. \end{enumerate}

As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that

equation[equation omitted — 141 chars of source]

on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).

Canonical form

In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if

equation[equation omitted — 119 chars of source]

While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables

equation[equation omitted — 401 chars of source]

where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting

equation[equation omitted — 275 chars of source]

for $\mathfrak{z}\in\mathbb{C}$, where

equation[equation omitted — 127 chars of source]

we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).

To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.

{{{{\scalefont{0.76}DGP$^{\ast}$}}}}

assumption$\{(y_{t},x_{t})\}$ are generated by a canonical CKSVAR, i.e.\ (ref) holds with $\Phi_{0}=[\phi_{0}^{+},\phi_{0}^{-},\Phi^{x}]=\Ican[p]$, so that (ref) may be equivalently written as \begin{equation} \begin{bmatrix}y_{t}\\ x_{t} \end{bmatrix}=c+\sum_{i=1}^{k}\begin{bmatrix}\phi_{i}^{+} & \phi_{i}^{-} & \Phi_{i}^{x}\end{bmatrix}\begin{bmatrix}y_{t-i}^{+}\\ y_{t-i}^{-}\\ x_{t-i} \end{bmatrix}+u_{t}. \end{equation}

The cointegrated CKSVAR

DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq

bmatrix[bmatrix omitted — 62 chars of source]

, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define

align*[align* omitted — 107 chars of source]

{{{{\scalefont{0.76}CVAR}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\det\Phi^{\pm}(\mathfrak{z})$ has $q^{\pm}\in\{1,\ldots,p\}$ roots at real unity, and all others outside the unit circle; and • $\operatorname{rk}\Pi^{\pm}=r^{\pm}=p-q^{\pm}$. \end{enumerate}

The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}

bmatrix[bmatrix omitted — 35 chars of source]

=\alpha

bmatrix[bmatrix omitted — 47 chars of source]

\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of

gather[gather omitted — 420 chars of source]

for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as

align[align omitted — 388 chars of source]

where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and

equation[equation omitted — 168 chars of source]

Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.

{{{{\scalefont{0.76}CO{{(ii)}}}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}},itemsep=1pt,topsep=2pt] • $r^{+}=r^{-}=\operatorname{rk}\Pi^{x}=r$, for some $r\in\{0,1,\ldots,p-1\}$. • $\rho_{{\scriptstyle \mathrm{JSR}}}(\{I+\tilde{\b{\beta}}(+1)^{\top}\tilde{\b{\alpha}},I+\tilde{\b{\beta}}(-1)^{\top}\tilde{\b{\alpha}}\})<1$. • $\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;+1)\beta_{\perp}(+1)=\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;-1)\beta_{\perp}(-1)\neq0$. • \begin{enumerate}[label={{{\scalefont{0.76}\alph*.}}}, ref={{{\scalefont{0.76}.\alph*}}}, leftmargin=0.40cm] • $\beta(y_{t})^{\top}z_{t}$, and $\Delta z_{t}$ have uniformly bounded $2+\delta_{u}$ moments, for $t\in\{-k+1,\ldots,0\}$. • $n^{-1/2}z_{0}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}{\cal Z}_{0}=[\begin{smallmatrix}{\cal Y}_{0}\\ {\cal X}_{0} \end{smallmatrix}]$, where ${\cal Z}_{0}$ is non-random, and satisfies $\beta(\mathcal{Y}_{0})^{\top}{\cal Z}_{0}=0$. \end{enumerate} \end{enumerate}

Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that

equation[equation omitted — 71 chars of source]

or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.

Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per

{{{{\scalefont{0.76}DET}}}}

assumption$c\in\operatorname{sp}\Pi^{+}\ensuremath{\cap}\operatorname{sp}\Pi^{-}$.

Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that

equation[equation omitted — 297 chars of source]

where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).

On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.

Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:

align*[align* omitted — 115 chars of source]

But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.

The modified Breitung (2002) test

Fundamental ideas

We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)

The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)

propSuppose that $\{w_{n,t}\}_{t=1}^{n}$ is a triangular array, taking values in $\mathbb{R}^{d_{w}}$, such that \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}\ensuremath{\rightsquigarrow}\int_{0}^{\lambda}\begin{bmatrix}\mathbb{W}(s)\\ 0_{d_{w}-\ell} \end{bmatrix}\ensuremath{\,\ensuremath{\mathrm{d}}} s\eqqcolon\begin{bmatrix}\mathbb{V}(\lambda)\\ 0_{d_{w}-\ell} \end{bmatrix} \end{equation} on $D_{\mathbb{R}^{d_{w}}}[0,1]$, where $\mathbb{W}$ is a random element of $D_{\mathbb{R}^{\ell}}[0,1]$ and \begin{align} \frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top} & \ensuremath{\rightsquigarrow}\begin{bmatrix}\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s & 0\\ 0 & \Omega \end{bmatrix} \end{align} where $\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$, $\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$ and $\Omega\in\mathbb{R}^{(d_{w}-\ell)\times(d_{w}-\ell)}$ are a.s.\ positive definite. Let $\{\lambda_{n,i}\}_{i=1}^{d_{w}}$ denote the solutions to \[ \det(\lambda\mathbb{B}_{n}-\mathbb{A}_{n})=0 \] ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,d_{w}}$, for \begin{align*} \mathbb{A}_{n} & \coloneqq\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}, & \mathbb{B}_{n} & \coloneqq\sum_{t=1}^{n}\sum_{i=1}^{t}w_{n,i}\sum_{j=1}^{t}w_{n,j}^{\top}. \end{align*} Then \begin{enumerate} • if $\ell_{0}=\ell$, \[ n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\rightsquigarrow}\operatorname{tr}\left[\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\left(\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\right)^{-1}\right]; \] • if $\ell_{0}>\ell$, $n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}\infty$. \end{enumerate}

To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq

bmatrix[bmatrix omitted — 57 chars of source]

z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).

By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.

Extension to nonlinearly cointegrated series

Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)

To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq

bmatrix[bmatrix omitted — 43 chars of source]

=\left[

array[array omitted — 110 chars of source]

\right]

bmatrix[bmatrix omitted — 27 chars of source]

=S_{p}(y_{t})z_{t} \] via

equation[equation omitted — 410 chars of source]

from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.

In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix

equation[equation omitted — 163 chars of source]

where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).

Thus the linear transformation

equation[equation omitted — 468 chars of source]

exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now

equation[equation omitted — 280 chars of source]

where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.

To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.

LLN for regime-switching processes

To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation

equation[equation omitted — 123 chars of source]

where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.

Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]

This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of

align*[align* omitted — 571 chars of source]

where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at

align*[align* omitted — 342 chars of source]

which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.

Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).

{{{{\scalefont{0.76}DET$^{\prime}$}}}}

assumption$e_{1}^{\top}P_{\beta_{\perp}}(+1)c=0$.

The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[

bmatrix[bmatrix omitted — 35 chars of source]

=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.

thmSuppose (ref), (ref), (ref) and (ref) hold. Let $\{A_{t}\}$, $\{B_{t}\}$ and $\{c_{t}\}$ be random sequences adapted to $\{\mathcal{F}_{t}\}$, respectively taking values in $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$, where $t\in\mathbb{Z}$. Suppose $\{v_{t}\}$ is i.i.d\ with $\ensuremath{\mathbb{E}} v_{t}=0$, and that $\{w_{t}\}$ satisfies \begin{equation} w_{t}=c_{t}+A_{t}w_{t-1}+B_{t}v_{t} \end{equation} for $t\geq-k$ and some given (random) $w_{-k}$ (with $w_{t}\coloneqq0$ for all $t\leq-k-1$); and: \begin{enumerate}[itemsep=2pt,topsep=3pt] • $A_{t}\in\mset A$, $B_{t}\in\mset B$ and $c_{t}\in{\cal C}$ for all $t\in\ensuremath{\mathbb{N}}$, where $\mset A$, $\mset B$ and ${\cal C}$ are bounded subsets of $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$ respectively, and $\rho_{{\scriptstyle \mathrm{JSR}}}(\mset A)<1$; • there exist $A^{\pm}\in\mset A$, $B^{\pm}\in\mset B$ and $c^{\pm}\in\mset C$ such that \begin{align*} y_{t-1}>0 and y_{t}>0 & \implies A_{t}=A^{+},\ B_{t}=B^{+},\ c_{t}=c^{+},\\ y_{t-1}<0 and y_{t}<0 & \implies A_{t}=A^{-},\ B_{t}=B^{-},\ c_{t}=c^{-}; \end{align*} • $m_{0}\geq1$ is such that $\smlnorm{w_{0}}_{m_{0}}+\smlnorm{v_{0}}_{m_{0}}<\infty$. • $g:\mathbb{R}^{d_{w}}\ensuremath{\rightarrow}\mathbb{R}^{d_{g}}$ is a continuous function satisfying \begin{equation} \smlnorm{g(w)-g(w^{\prime})}\leq C(1+\smlnorm w^{\ell_{0}}+\smlnorm{w^{\prime}}^{\ell_{0}})\smlnorm{w-w^{\prime}} \end{equation} for all $w,w^{\prime}\in\mathbb{R}^{d_{w}}$, for some $0\leq\ell_{0}<m_{0}-1$. \end{enumerate} Then $\ensuremath{\mathbb{E}}\smlnorm{g(w_{0}^{\pm})}<\infty$, and on $D[0,1]$, \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(w_{t})\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\int_{0}^{\lambda}\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} where \begin{equation} w_{0}^{\pm}=(I_{d_{w}}-A^{\pm})^{-1}c^{\pm}+\sum_{\ell=0}^{\infty}(A^{\pm})^{\ell}B^{\pm}v_{-\ell}. \end{equation} Moreover, \begin{equation} \frac{1}{n^{3/2}}\sum_{t=1}^{\smlfloor{n\lambda}}[g(w_{t})\otimes z_{t}^{d}]\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\otimes\int_{0}^{\lambda}Z(\mu)\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} jointly with $U_{n}\ensuremath{\rightsquigarrow} U$.

Limiting distribution and consistency

Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as

equation[equation omitted — 100 chars of source]

where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to

equation[equation omitted — 81 chars of source]

ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for

align[align omitted — 224 chars of source]

This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.

To state the limiting distribution of the test statistic, define

equation[equation omitted — 86 chars of source]

where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process

equation[equation omitted — 17,812 chars of source]

Introduction

For almost half a century, the structural vector autoregression (SVAR) has been the workhorse model of empirical macroeconomics. In addition to providing a tractable framework for the identification of causal relationships in the presence of simultaneity, the model succeeds in capturing many of the characteristic properties of macroeconomic time series: their temporal dependence, their trending and random wandering behaviour, and the tendency of related series to move together. In this regard, the emergence of the theory of cointegration (Granger86OBES; EG87Ecta) was of major significance: for by formalising that co-movement in terms of common stochastic trends, it made it possible to identify the precise conditions under which an SVAR could generate such common trends, as per the Granger--Johansen representation theorem (GJRT; Joh91Ecta,Joh95). This result has in turn provided the basis for a rich and fruitful theory of asymptotic inference in cointegrated SVARs, concerning the number of common stochastic trends in the system (or equivalently, the cointegrating rank), the coefficients on the cointegrating relations, and the model parameters (and implied impulse responses, etc.).

In its original conception, cointegration was inherently linear; there have since been multifarious efforts to extend it in a nonlinear direction, as reviewed by Tjo20EctRev. Paralleling those efforts has been the burgeoning of a literature on nonlinear SVARs, but which has been confined almost entirely to the modelling of stationary time series (see e.g.\ Tong90; TTG10; for the exceptional case of `nonlinear VECM' models, see KR10JoE). This unfortunately precludes the application of these nonlinear SVARs to settings where, for economic reasons, the nonlinearities relate to the level of a stochastically trending series, so that reformulating the model in terms of the (more approximately stationary) differenced series is not appropriate. A leading example arises in the context of the zero lower bound (ZLB) constraint on nominal interest rates, which refers to the level of a highly persistent -- and arguably integrated -- series, rather than to its first differences.

The development of a new class of `endogenous regime switching' piecewise affine SVARs -- and their successful application to highly persistent series that are subject to occasionally binding constraints (SM21; AMSV21; ILMZ20) -- has recently foregrounded the question of whether, and how, one can accommodate stochastic trends in nonlinear SVARs. By way of an answer, DMW22 and DM24 provide extensions of the GJRT to a broad class of nonlinear SVARs: in the former, to a two-regime piecewise affine SVAR (the `CKSVAR'), and in the latter, to more general, additively time-separable nonlinear SVARs of the form

equation[equation omitted — 82 chars of source]

where $z_{t}$ and $u_{t}$ are respectively the observed series and the innovations, both of which are $\mathbb{R}^{p}$-valued, and $f_{i}:\mathbb{R}^{p}\ensuremath{\rightarrow}\mathbb{R}^{p}$. Their results demonstrate that, alongside linear cointegration, nonlinear SVARs of the form (ref) are capable of accommodating much richer varieties of long-run behaviour than are linear SVARs, including nonlinear common stochastic trends and nonlinear cointegrating relations.

There remains the question of how to perform inference in the setting of (ref), in the presence of (linear or nonlinear) cointegration. In this paper, we consider this problem when (ref) is specialised to the two-regime piecewise affine model of DMW22, as per

equation[equation omitted — 187 chars of source]

where we have partitioned $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ such that $y_{t}$ is $\mathbb{R}$-valued and $x_{t}$ is $\mathbb{R}^{p-1}$-valued, and $y_{t}^{+}=\max\{y_{t},0\}$ and $y_{t}^{-}=\min\{y_{t},0\}$ respectively denote the positive and negative parts of $y_{t}$. We further suppose that this model is configured such that the cointegrating rank, $r$, is invariant to the sign of $y_{t}$, while permitting those $r$ cointegrating relations to be nonlinear: what is termed `case (ii)' in the typology of DMW22; see (ref) for a discussion. Even in this case, asymptotic inference is complicated by the fact that the processes generated by the model do not readily fall within any class previously considered in econometrics. Although $\{z_{t}\}$ behaves similarly, in large samples, to a (linear) integrated process, in the sense that $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a nondegenerate limiting process $Z(\lambda)$, neither its first differences nor the equilibrium errors will be stationary, but instead follow a (stable) time-varying autoregressive process, whose coefficients depend on the sign of the integrated process $\{y_{t}\}$. This renders any existing LLN-type results for `weakly dependent' processes inapplicable.

In this paper we take the first steps towards the development of valid asymptotic inference in the model (ref), in the presence of cointegration. We do so by considering the simpler problem of inference on the cointegrating rank of (ref), using a form of the Breitung02JoE multivariate variance ratio test statistic, modified so as to accommodate the possibility of nonlinear cointegration. This motivates the main technical contribution of the paper: a new LLN-type result for the class of time-varying, stable but nonstationary autoregressive processes that may be generated by (ref), which is provided in (ref) along with the asymptotics of our test statistic. This result is fundamental to the asymptotics of estimators of the parameters of (ref), the derivation of which is the subject of the authors' ongoing research. The finite-sample performance of our proposed test is investigated through simulation exercises reported in (ref), where it is shown that the conventional (i.e.\ unmodified) Breitung02JoE test tends to incorrectly interpret the presence of nonlinear cointegration as evidence in favour of additional stochastic trends being present in the data, a problem that is avoided by our proposed test. (ref) concludes.

notation*$e_{m,i}$ denotes the $i$th column of the $m\times m$ identity matrix $I_{m}$; when $m$ is clear from the context, we write this simply as $e_{i}$. In a statement such as $f(a^{\pm},b^{\pm})=0$, the notation `$\pm$' signifies that both $f(a^{+},b^{+})=0$ and $f(a^{-},b^{-})=0$ hold; similarly, `$a^{\pm}\in A$' denotes that both $a^{+}$ and $a^{-}$ are elements of $A$. All limits are taken as $n\ensuremath{\rightarrow}\infty$ unless otherwise stated. $\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}$ and $\ensuremath{\rightsquigarrow}$ respectively denote convergence in probability and in distribution (weak convergence). We write `$X_{n}(\lambda)\ensuremath{\rightsquigarrow} X(\lambda)$ on $D_{\mathbb{R}^{m}}[0,1]$' to denote that $\{X_{n}\}$ converges weakly to $X$, where these are considered as random elements of $D_{\mathbb{R}^{m}}[0,1]$, the space of cadlag functions $[0,1]\ensuremath{\rightarrow}\mathbb{R}^{m}$, equipped with the uniform topology; we denote this as $D[0,1]$ whenever the value of $m$ is clear from the context. $\smlnorm{\cdot}$ denotes the Euclidean norm on $\mathbb{R}^{m}$, and the matrix norm that it induces. For $X$ a random vector and $p\geq1$, $\smlnorm X_{p}\coloneqq(\ensuremath{\mathbb{E}}\smlnorm X^{p})^{1/p}$. $C$, $C_{1}$, etc., denote generic constants that may take different values at different places of the same proof.

Model: the censored and kinked SVAR

Framework

We consider a structural VAR($k$) model in $p$ variables, in which one series, $y_{t}$, enters with coefficients that differ according to whether it is above or below a time-invariant threshold $b$, while the other $p-1$ series, collected in $x_{t}$, enter linearly (SM21; DMW22). Defining

align[align omitted — 111 chars of source]

we specify that $z_{t}=(y_{t},x_{t}^{\top})^{\top}$ follow

equation[equation omitted — 193 chars of source]

or, more compactly,

equation[equation omitted — 100 chars of source]

where

align*[align* omitted — 158 chars of source]

for $\phi_{i}^{\pm}\in\mathbb{R}^{p\times1}$ and $\Phi_{i}^{x}\in\mathbb{R}^{p\times(p-1)}$, and $L$ denotes the lag operator. Through an appropriate redefinition of $y_{t}$ and $c$, we may take $b$ (which we treat here as being known) to be zero without loss of generality, and will do so throughout the sequel. In this case, $y_{t}^{+}$ and $y_{t}^{-}$ respectively equal the positive and negative parts of $y_{t}$, and $y_{t}=y_{t}^{+}+y_{t}^{-}$.\footnote{Throughout the following, the notation `$a^{\pm}$' connotes $a^{+}$ and $a^{-}$ as objects associated respectively with $y_{t}^{+}$ and $y_{t}^{-}$, or their lags. If we want to instead denote the positive and negative parts of some $a\in\mathbb{R}$, we shall do so by writing $[a]_{+}\coloneqq\max\{a,0\}$ or $[a]_{-}\coloneqq\min\{a,0\}$.} Following SM21, we term this model the `censored and kinked SVAR' (CKSVAR), even though we here suppose that $y_{t}$ is observed on both sides of zero, rather than being subject to censoring.

We follow SM21 and AMSV21 in maintaining the following conditions, which are necessary and sufficient to ensure that (ref) has a unique solution for $(y_{t},x_{t})$, for all possible values of $u_{t}$. Define \[ \Phi_{0}\coloneqq

bmatrix[bmatrix omitted — 55 chars of source]

=

bmatrix[bmatrix omitted — 118 chars of source]

, \] $\Phi_{0}^{+}\coloneqq[\phi_{0}^{+},\Phi_{0}^{x}]$ and $\Phi_{0}^{-}\coloneqq[\phi_{0}^{-},\Phi_{0}^{x}]$.

{{{{\scalefont{0.76}DGP}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\{(y_{t},x_{t})\}$ are generated according to (ref)--(ref) with $b=0$, with (possibly random) initial values $(y_{i},x_{i})$, for $i\in\{-k+1,\ldots,0\}$; • $\operatorname{sgn}(\det\Phi_{0}^{+})=\operatorname{sgn}(\det\Phi_{0}^{-})\neq0$. • $\Phi_{0,xx}$ is invertible, and \[ \operatorname{sgn}\{\phi_{0,yy}^{+}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{+}\}=\operatorname{sgn}\{\phi_{0,yy}^{-}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{-}\}>0. \]$\{u_{t}\}_{t\in\mathbb{Z}}$ is an i.i.d.\ sequence in $\mathbb{R}^{p}$ with $\ensuremath{\mathbb{E}} u_{t}=0$, $\ensuremath{\mathbb{E}} u_{t}u_{t}^{\top}=\Sigma_{u}$ positive definite, and $\smlnorm{u_{t}}_{2+\delta_{u}}<\infty$ for some $\delta_{u}>0$. \end{enumerate}

As discussed in DMW23stat, (ref)(ref) may be maintained without loss of generality, when the invertibility condition (ref)(ref) holds. Let $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ denote an underlying filtration to which the preceding processes are all adapted. When we say that a sequence is i.i.d.,\ as per $\{u_{t}\}_{t\in\mathbb{Z}}$ in (ref)(ref), we mean that this sequence is $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$-adapted, and additionally that $u_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$. An immediate implication of (ref)(ref) is that

equation[equation omitted — 141 chars of source]

on $D[0,1]$, where $U$ is a $p$-dimensional Brownian motion with variance $\Sigma_{u}$. All the weak convergences that are stated in this paper hold jointly with (ref).

Canonical form

In the terminology of DMW23stat and DMW22, we designate a CKSVAR as canonical if

equation[equation omitted — 119 chars of source]

While it is not always the case that the reduced form of (ref) corresponds directly to a canonical CKSVAR, by defining the canonical variables

equation[equation omitted — 401 chars of source]

where $\bar{\phi}_{0,yy}^{\pm}\coloneqq\phi_{0,yy}^{\pm}-\phi_{0,yx}^{\top}\Phi_{0,xx}^{-1}\phi_{0,xy}^{\pm}>0$ and $P^{-1}$ is invertible under (ref); and setting

equation[equation omitted — 275 chars of source]

for $\mathfrak{z}\in\mathbb{C}$, where

equation[equation omitted — 127 chars of source]

we obtain a canonical CKSVAR for $\tilde{z}_{t}\coloneqq(\tilde{y}_{t},\tilde{x}_{t}^{\top})^{\top}$ (see Proposition 2.1 in DMW23stat).

To distinguish between a general CKSVAR in which possibly $\Phi_{0}\neq\Ican[p]$, and its associated canonical form, we shall refer to the former as the `structural form' of the CKSVAR. Since the time series properties of a general CKSVAR are largely inherited from its derived canonical form, we shall occasionally work with this more convenient representation of the system, and indicate this as follows.

{{{{\scalefont{0.76}DGP$^{\ast}$}}}}

assumption$\{(y_{t},x_{t})\}$ are generated by a canonical CKSVAR, i.e.\ (ref) holds with $\Phi_{0}=[\phi_{0}^{+},\phi_{0}^{-},\Phi^{x}]=\Ican[p]$, so that (ref) may be equivalently written as \begin{equation} \begin{bmatrix}y_{t}\\ x_{t} \end{bmatrix}=c+\sum_{i=1}^{k}\begin{bmatrix}\phi_{i}^{+} & \phi_{i}^{-} & \Phi_{i}^{x}\end{bmatrix}\begin{bmatrix}y_{t-i}^{+}\\ y_{t-i}^{-}\\ x_{t-i} \end{bmatrix}+u_{t}. \end{equation}

The cointegrated CKSVAR

DMW22, henceforth \citetalias{DMW22}, develop conditions under which the CKSVAR is capable of generating cointegrated time series. Their work identifies three cases, which may be distinguished according to whether stochastic trends are imparted: (i) to $y_{t}^{+}$ only (or equivalently to $y_{t}^{-}$ only); (ii) to both $y_{t}^{+}$ and $y_{t}^{-}$; and (iii) to neither $y_{t}^{+}$ nor $y_{t}^{-}$. Here our focus is on case (ii), which entails that the system has a well-defined cointegrating rank $r$, but permits the $r$ cointegrating relationships that eliminate the ($p-r=q$) common trends to be nonlinear. The assumptions that characterise how the model needs to be configured for case (ii) are given below. To state these, define the autoregressive polynomials \[ \Phi^{\pm}(\mathfrak{z})\coloneqq

bmatrix[bmatrix omitted — 62 chars of source]

, \] and let $\Gamma_{i}^{\pm}\coloneqq-\sum_{j=i+1}^{k}\Phi_{j}^{\pm}\eqqcolon[\gamma_{i}^{\pm},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, so that $\Gamma^{\pm}(\mathfrak{z})\coloneqq\Phi_{0}^{\pm}-\sum_{i=1}^{k-1}\Gamma_{i}^{\pm}\mathfrak{z}^{i}$ is such that \[ \Phi^{\pm}(\mathfrak{z})=\Phi^{\pm}(1)\mathfrak{z}+\Gamma^{\pm}(\mathfrak{z})(1-\mathfrak{z}). \] We further define

align*[align* omitted — 107 chars of source]

{{{{\scalefont{0.76}CVAR}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}, itemsep=1pt,topsep=2pt] • $\det\Phi^{\pm}(\mathfrak{z})$ has $q^{\pm}\in\{1,\ldots,p\}$ roots at real unity, and all others outside the unit circle; and • $\operatorname{rk}\Pi^{\pm}=r^{\pm}=p-q^{\pm}$. \end{enumerate}

The preceding conditions are common to all three cases noted above. To specialise to case (ii), which has a constant cointegrating rank $r=r^{+}=r^{-}$, with a stochastic trend present being in $y_{t}$, we must additionally suppose that $\operatorname{rk}\Pi^{x}=r$, so that $\Pi^{\pm}$ may be written as \[ \Pi^{\pm}=\Pi^{x}

bmatrix[bmatrix omitted — 35 chars of source]

=\alpha

bmatrix[bmatrix omitted — 47 chars of source]

\eqqcolon\alpha\beta^{\pm\top}, \] where $\alpha\in\mathbb{R}^{p\times r}$, $\beta_{x}\in\mathbb{R}^{(p-1)\times r}$ and $\beta^{\pm}\in\mathbb{R}^{p\times r}$ have rank $r$, and $\theta^{\pm}\in\mathbb{R}^{p-1}$ is such that $\Pi^{x}\theta^{\pm}=\pi^{\pm}$ (see Section 4.2 of \citetalias{DMW22}). Letting $\ensuremath{\mathbf{1}}^{+}(y)\coloneqq\ensuremath{\mathbf{1}}\{y\geq0\}$ and $\ensuremath{\mathbf{1}}^{-}(y)\coloneqq\ensuremath{\mathbf{1}}\{y<0\}$, the (possibly nonlinear) $r$ cointegrating relationships among the elements of $z_{t}$ are given by \[ \beta(y)\coloneqq\beta^{+}\ensuremath{\mathbf{1}}^{+}(y)+\beta^{-}\ensuremath{\mathbf{1}}^{-}(y). \] Let $\alpha_{\perp}\in\mathbb{R}^{p\times q}$ be such that $\alpha_{\perp}^{\top}\alpha=0$, and $[\alpha,\alpha_{\perp}]$ is nonsingular. The limiting form of the stochastic trends will be a kind of (regime-dependent) projection of the $p$-dimensional Brownian motion $U$ onto a manifold of dimension $q=p-r$, where this projection is defined in terms of

gather[gather omitted — 420 chars of source]

for $\theta(y)\coloneqq\ensuremath{\mathbf{1}}^{+}(y)\theta^{+}+\ensuremath{\mathbf{1}}^{-}(y)\theta^{-}$. (Such objects as $P_{\beta_{\perp}}(y)$ take only two distinct values, depending on the sign of $y$, and we routinely use the notation $P_{\beta_{\perp}}(+1)$ and $P_{\beta_{\perp}}(-1)$ to indicate these.) Define $\b{\alpha},\b{\beta}(y)\in\mathbb{R}^{[k(p+1)-1]\times[r+(k-1)(p+1)]}$ as

align[align omitted — 388 chars of source]

where $\Gamma_{i}\coloneqq[\gamma_{i}^{+},\gamma_{i}^{-},\Gamma_{i}^{x}]$ for $i\in\{1,\ldots,k-1\}$, and

equation[equation omitted — 168 chars of source]

Finally, let $\rho(M)$ denote the spectral radius of $M\in\mathbb{R}^{m\times m}$, and for $\mathcal{A}\subset\mathbb{R}^{m\times m}$ a bounded collection of matrices, let \[ \rho_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{t\ensuremath{\rightarrow}\infty}\sup_{B\in\mathcal{A}^{t}}\rho(B)^{1/t} \] denote its joint spectral radius (JSR; e.g.\ Jungers09, Defn.\ 1.1), where $\mathcal{A}^{t}\coloneqq\{\prod_{s=1}^{t}M_{s}\mid M_{s}\in\mathcal{A}\}$ is the set of $t$-fold products of matrices in ${\cal A}$.

{{{{\scalefont{0.76}CO{{(ii)}}}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}},itemsep=1pt,topsep=2pt] • $r^{+}=r^{-}=\operatorname{rk}\Pi^{x}=r$, for some $r\in\{0,1,\ldots,p-1\}$. • $\rho_{{\scriptstyle \mathrm{JSR}}}(\{I+\tilde{\b{\beta}}(+1)^{\top}\tilde{\b{\alpha}},I+\tilde{\b{\beta}}(-1)^{\top}\tilde{\b{\alpha}}\})<1$. • $\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;+1)\beta_{\perp}(+1)=\operatorname{sgn}\det\alpha_{\perp}^{\top}\Gamma(1;-1)\beta_{\perp}(-1)\neq0$. • \begin{enumerate}[label={{{\scalefont{0.76}\alph*.}}}, ref={{{\scalefont{0.76}.\alph*}}}, leftmargin=0.40cm] • $\beta(y_{t})^{\top}z_{t}$, and $\Delta z_{t}$ have uniformly bounded $2+\delta_{u}$ moments, for $t\in\{-k+1,\ldots,0\}$. • $n^{-1/2}z_{0}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}{\cal Z}_{0}=[\begin{smallmatrix}{\cal Y}_{0}\\ {\cal X}_{0} \end{smallmatrix}]$, where ${\cal Z}_{0}$ is non-random, and satisfies $\beta(\mathcal{Y}_{0})^{\top}{\cal Z}_{0}=0$. \end{enumerate} \end{enumerate}

Condition (ref)(ref) is stated slightly differently from the form given in \citetalias{DMW22}, so as to more directly accommodate the case of a general (i.e.\ non-canonical) CKSVAR. In particular, $\tilde{\b{\beta}}(y)$ and $\tilde{\b{\alpha}}$ refer to the counterparts of (ref) constructed from the parameters of the canonical form of the CKSVAR, derived via the mapping (ref). (So if the CKSVAR is in fact canonical, the tildes are redundant.) See Remark 4.2(i) of \citetalias{DMW22} for further details. Regarding the history of the process prior to time $t=-k+1$, we henceforth adopt the (innocuous) convention that

equation[equation omitted — 71 chars of source]

or equivalently that $z_{t}=z_{-k}$ for all $t\leq-k$.

Finally, for the purposes of developing the asymptotics of our rank test ((ref) below), we shall maintain that the intercept $c$ is such that no deterministic trends are present in any of the model variables, as per

{{{{\scalefont{0.76}DET}}}}

assumption$c\in\operatorname{sp}\Pi^{+}\ensuremath{\cap}\operatorname{sp}\Pi^{-}$.

Under the preceding conditions ((ref), (ref), (ref) and (ref)), it follows by Theorem 4.2 in \citetalias{DMW22} that

equation[equation omitted — 297 chars of source]

where $U_{0}(\lambda)=\Gamma(1;\mathcal{Y}_{0})\mathcal{Z}_{0}+U(\lambda)$. (For a further heuristic discussion of the convergence in (ref) and the properties of the limiting process $Z(\lambda)$, see Section 3.3 of \citetalias{DMW22}.) Since $P_{\beta_{\perp}}(\pm1)$ are rank $q$ (oblique projection) matrices, we may regard $\{z_{t}\}$ as having $q$ common (stochastic) trends, and $r$ cointegrating relations given by the columns of $\beta(y)$, that eliminate those trends (since $\beta(y)^{\top}P_{\beta_{\perp}}(y)=0$).

On the basis of (ref), \citetalias{DMW22} (see their Defn.\ 3.1) classify $\{z_{t}\}$ as $I^{\ast}(1)$, because $n^{-1/2}z_{\smlfloor{n\lambda}}$ converges weakly to a non-degenerate process. By contrast, since the equilibrium errors $\xi_{t}\coloneqq\beta(y_{t})^{\top}z_{t}$ are purged of the common trends in $z_{t}$, these satisfy $\max_{1\leq t\leq n}\smlnorm{\xi_{t}}=o_{p}(n^{1/2})$, and so are of strictly smaller order than $\{z_{t}\}$; they accordingly classify $\{\xi_{t}\}$ as $I^{\ast}(0)$. These notions of $I^{\ast}(0)$ and $I^{\ast}(1)$ processes provide a means of distinguishing between processes whose magnitudes differ, because of the presence or absence of stochastic trends, in a setting where the usual definitions of $I(0)$ and $I(1)$ processes do not apply -- because in general neither $\xi_{t}$ nor $\Delta z_{t}$ will be stationary under the foregoing assumptions.

Although (ref) implies that $Z$ is not `globally' a linear projection of $U_{0}$ onto a $q$-dimensional linear subspace, the following relationships hold `locally', depending on the sign of the first component, $Y$, of $Z$:

align*[align* omitted — 115 chars of source]

But in general neither $\beta^{+\top}Z(\lambda)$ nor $\beta^{-\top}Z(\lambda)$ will be identically zero for all $\lambda\in[0,1]$, unless $\beta^{+}=\beta^{-}$. The fact that there may be no rank $r=p-q$ matrix whose columns force $Z$ to be identically zero significantly complicates the problem of inference on the cointegrating rank, and motivates our development of a modified form of the Breitung02JoE test below.

The modified Breitung (2002) test

Fundamental ideas

We seek to develop an (asymptotically valid) test on the cointegrating rank $r$ -- or equivalently, the number of common trends $q$ -- that is able to accommodate the possibility of data generated by a CKSVAR configured as per case (ii), by adapting the approach of Breitung02JoE. Henceforth, as per the discussion following (ref) above, the threshold $b$ that delineates the two regimes is assumed to be known, and normalised to zero: so that what we have denoted as $y_{t}^{+}$ and $y_{t}^{-}$ may be regarded as directly observed, rather than depending on some prior estimator of $b$. Estimation of $b$ may be undertaken in conjunction with the estimation of the other parameters of the SVAR (ref), e.g.\ by maximum likelihood, the asymptotics of which are deferred to future work. (We anticipate that use of a consistent estimator of $b$ would yield a test statistic with an identical null limiting distribution to that derived below: due to $y_{t}$ being integrated under the null, any misclassification that results from $\hat{b}_{n}\neq b$ would affect at most $o_{p}(n^{1/2})$ observations.)

The mathematical underpinnings of Breitung02JoE's Breitung02JoE test, itself a multivariate generalisation of the variance ratio test, may be conveniently summarised as follows. (The proof of which, together with those of all other results given in this section, appear in (ref).)

propSuppose that $\{w_{n,t}\}_{t=1}^{n}$ is a triangular array, taking values in $\mathbb{R}^{d_{w}}$, such that \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}\ensuremath{\rightsquigarrow}\int_{0}^{\lambda}\begin{bmatrix}\mathbb{W}(s)\\ 0_{d_{w}-\ell} \end{bmatrix}\ensuremath{\,\ensuremath{\mathrm{d}}} s\eqqcolon\begin{bmatrix}\mathbb{V}(\lambda)\\ 0_{d_{w}-\ell} \end{bmatrix} \end{equation} on $D_{\mathbb{R}^{d_{w}}}[0,1]$, where $\mathbb{W}$ is a random element of $D_{\mathbb{R}^{\ell}}[0,1]$ and \begin{align} \frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top} & \ensuremath{\rightsquigarrow}\begin{bmatrix}\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s & 0\\ 0 & \Omega \end{bmatrix} \end{align} where $\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$, $\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s$ and $\Omega\in\mathbb{R}^{(d_{w}-\ell)\times(d_{w}-\ell)}$ are a.s.\ positive definite. Let $\{\lambda_{n,i}\}_{i=1}^{d_{w}}$ denote the solutions to \[ \det(\lambda\mathbb{B}_{n}-\mathbb{A}_{n})=0 \] ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,d_{w}}$, for \begin{align*} \mathbb{A}_{n} & \coloneqq\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}, & \mathbb{B}_{n} & \coloneqq\sum_{t=1}^{n}\sum_{i=1}^{t}w_{n,i}\sum_{j=1}^{t}w_{n,j}^{\top}. \end{align*} Then \begin{enumerate} • if $\ell_{0}=\ell$, \[ n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\rightsquigarrow}\operatorname{tr}\left[\int_{0}^{1}\mathbb{W}(s)\mathbb{W}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\left(\int_{0}^{1}\mathbb{V}(s)\mathbb{V}(s)^{\top}\ensuremath{\,\ensuremath{\mathrm{d}}} s\right)^{-1}\right]; \] • if $\ell_{0}>\ell$, $n^{2}\sum_{i=1}^{\ell_{0}}\lambda_{n,i}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}\infty$. \end{enumerate}

To illustrate how (ref) provides the basis for a test of cointegrating rank, let us suppose initially that $\{z_{t}\}$ is generated by a linear cointegrated SVAR with $q$ common trends, or more generally by a CKSVAR satisfying the conditions above ((ref), (ref), (ref) and (ref)), but for which $\beta^{+}=\beta^{-}=\beta$ and $\Gamma^{+}(1)=\Gamma^{-}(1)=\Gamma(1)$. Then $P_{\beta_{\perp}}(y)$ no longer depends on (the sign of) $y$, and (ref) reduces to \[ n^{-1/2}z_{\smlfloor{n\lambda}}\ensuremath{\rightsquigarrow}\beta_{\perp}[\alpha_{\perp}^{\top}\Gamma(1)\beta_{\perp}]^{-1}\alpha_{\perp}^{\top}U_{0}(\lambda). \] It follows that by taking \[ w_{n,t}\coloneqq

bmatrix[bmatrix omitted — 57 chars of source]

z_{t} \] we may linearly separate $z_{t}$ into its $q$ `integrated' (i.e.\ $I^{\ast}(1)$) and $r=p-q$ `weakly dependent' (i.e.\ $I^{\ast}(0)$) components, with the result that the first $q$ components of $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}w_{n,t}$ will converge weakly to a (nondegenerate) limiting process, whereas the final $r$ components will converge to zero, exactly as in the manner of (ref). $\frac{1}{n}\sum_{t=1}^{n}w_{n,t}w_{n,t}^{\top}$ then also converges to an (invertible) block diagonal matrix, as in (ref).

By (ref), the sum of the first $q_{0}$ generalised eigenvalues of $\mathbb{A}_{n}$ with respect to $\mathbb{B}_{n}$ will then exhibit divergent asymptotic behaviour, depending on whether $q_{0}$ is equal to or strictly greater than $q$. This provides the basis for the use of this quantity as a statistic for testing hypotheses regarding the value of $q$, exactly as proposed in Breitung02JoE. Since these generalised eigenvalues are invariant to common linear transformations of $\mathbb{A}_{n}$ and $\mathbb{B}_{n}$, and $w_{n,t}$ is a linear transformation of $z_{t}$, they may be computed without knowledge of $[\beta_{\perp},\beta]$, simply by replacing each instance of $w_{n,t}$ by $z_{t}$ in the definitions of those matrices.

Extension to nonlinearly cointegrated series

Suppose that we now permit $\beta^{+}\neq\beta^{-}$ and/or $\Gamma^{+}(1)\neq\Gamma^{-}(1)$. In this case, $P_{\beta_{\perp}}(-1)$ and $P_{\beta_{\perp}}(+1)$ each have rank $q$, but may differ by a rank one matrix, and as a result there may only be $r-1$ distinct linear combinations of $z_{t}$ that will be $I^{\ast}(0)$. Accordingly, applying the usual Breitung test to $\{z_{t}\}$ directly would tend to yield the incorrect conclusion that there are $q+1$ common trends, rather than only $q$. (Thus for example, in a bivariate nonlinear SVAR with one common nonlinear trend, this test may tend to conclude that there are two common trends and no cointegrating relations.)

To address this problem, here we utilise the fact that the nonlinearity in the CKSVAR is entirely a function of the sign of the first component of $z_{t}=(y_{t},x_{t}^{\top})^{\top}$, such that the nonlinear cointegrating relationships $\beta(y)$ can be rewritten as linear cointegrating relationships between the elements of \[ z_{t}^{\ast}\coloneqq

bmatrix[bmatrix omitted — 43 chars of source]

=\left[

array[array omitted — 110 chars of source]

\right]

bmatrix[bmatrix omitted — 27 chars of source]

=S_{p}(y_{t})z_{t} \] via

equation[equation omitted — 410 chars of source]

from which it follows that \[ \beta(y_{t})^{\top}z_{t}=\beta^{\ast\top}S_{p}(y_{t})z_{t}=\beta^{\ast\top}z_{t}^{\ast} \] since $z_{t}^{\ast}=S_{p}(y_{t})z_{t}$; the r.h.s.\ thus gives the $r$ linear relationships that render $\beta^{\ast\top}z_{t}^{\ast}\ensuremath{\sim} I^{\ast}(0)$. As a corollary, there will be $q+1$ (linearly independent) vectors in $\mathbb{R}^{p+1}$ that extract distinct $I^{\ast}(1)$ components from $z_{t}^{\ast}$. We obtain an additional $I^{\ast}(1)$ component, because under case (ii) the common trends are present in both $y_{t}^{+}$ and $y_{t}^{-}$, which appear separately as the first two components of $z_{t}^{\ast}$.

In extracting those common trends, we are free to choose any $(q+1)$-dimensional basis in $\mathbb{R}^{p+1}$ whose span does not (non-trivially) intersect with $\operatorname{sp}\beta^{\ast}$. Here we take this basis to be the columns of the following $(p+1)\times(q+1)$ matrix

equation[equation omitted — 163 chars of source]

where the columns of $\beta_{x,\perp}\in\mathbb{R}^{(p-1)\times(q-1)}$ span the orthogonal complement of $\operatorname{sp}\beta_{x}$ in $\mathbb{R}^{p-1}$, and as shown in the proof of (ref) (see (ref), in particular), we are free to choose $\tau_{xy}^{\pm}\in\mathbb{R}^{q-1}$ so as to facilitate the convergence of our test statistic to a pivotal limiting distribution. The matrix $\tau^{\ast}$ plainly has rank $q+1$; moreover the $(p+1)\times(p+1)$ matrix $[\beta^{\ast},\tau^{\ast}]$ is nonsingular, irrespective of the values of $\tau_{xy}^{\pm}$ (see (ref)).

Thus the linear transformation

equation[equation omitted — 468 chars of source]

exhaustively separates $z_{t}^{\ast}$ into its $I^{\ast}(0)$ and (appropriately standardised) $I^{\ast}(1)$ components, and so renders the process $\{z_{t}^{\ast}\}$ into a form conformable with (ref) above. The decomposition (ref) provides the basis for applying what we term our modified Breitung (MB) test to the data generated by a cointegrated CKSVAR, under case (ii), `modified' in the sense that the test statistic will be constructed from $z_{t}^{\ast}$ rather than $z_{t}$. Indeed, if $c=0$, then it will follow from our results below that $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}\ensuremath{\rightsquigarrow}0$ on $D[0,1]$, and so the test could be applied directly to $z_{t}^{\ast}$ in this case. More generally, when $c\neq0$, we need to first extract any deterministic components whose presence would otherwise distort the distribution of the test statistic. If we suppose that (ref) holds, then no deterministic trends are present in $z_{t}$, and by analogy with the approach taken in the linear setting, we may project out any constant deterministic terms by applying the test not to $z_{t}^{\ast}$ but rather to \[ \dmn z_{t}^{\ast}\coloneqq z_{t}^{\ast}-\hat{\mu}_{n,z^{\ast}} \] where $\hat{\mu}_{n,z^{\ast}}\coloneqq\frac{1}{n}\sum_{t=1}^{n}z_{t}^{\ast}$, so that now

equation[equation omitted — 280 chars of source]

where $\hat{\mu}_{n,\xi}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\xi_{t}$ and $\hat{\mu}_{n,\varrho}\coloneqq\frac{1}{n}\sum_{t=1}^{n}\varrho_{n,t}$.

To obtain the limiting distribution of our proposed test, we shall verify that $w_{n,t}=T_{n}^{\top}\dmn z_{t}^{\ast}$ satisfies the requirements of (ref) above. In order for (ref) to conform with (ref), we must show that \[ \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\dmn{\xi}_{t}=\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}\xi_{t}-\lambda\hat{\mu}_{n,\xi}=o_{p}(1) \] uniformly in $\lambda\in[0,1]$. Similarly, for the purposes of (ref), require that $\frac{1}{n}\sum_{t=1}^{n}\dmn{\xi}_{t}\dmn{\xi}_{t}^{\top}$ converges weakly to an (a.s.)\ positive definite matrix. In other words, we require a fundamental law of large numbers (LLN) for sample averages of the form $\frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(\xi_{t})$. Since $\{\xi_{t}\}$ is not, in general, a stationary process, existing results do not apply here, and this motivates the development of the novel LLN given as (ref) below.

LLN for regime-switching processes

To illustrate the essential ideas, suppose for simplicity of exposition that $k=1$, and that the CKSVAR is canonical. Then by Lemma B.2 of \citetalias{DMW22}, $\xi_{t}=\beta(y_{t})^{\top}z_{t}$ admits the time-varying autoregressive representation

equation[equation omitted — 123 chars of source]

where $\{\beta_{t}\}$ is a random sequence that in general depends, nonlinearly, on the values of $y_{t}$ and $y_{t-1}$. Under (ref)(ref), which implies that $I_{r}+\beta_{t}^{\top}\alpha$ is drawn from a set of matrices whose joint spectral radius is strictly bounded by unity, $\{\xi_{t}\}$ will be a `stable' process in the sense that it is stochastically bounded; but the dependence of $\beta_{t}$ on $y_{t}$ prevents $\{\xi_{t}\}$ from being stationary.

Since $\beta_{t}=\beta^{+}$ whenever $y_{t-1}>0$ and $y_{t}>0$, it follows that if $y_{s}>0$ for all $s\in\{t-m,\ldots,t\}$, then \[ \xi_{t}=(I_{r}+\beta^{+\top}\alpha)^{m}\xi_{t-m}+\sum_{\ell=0}^{m-1}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}(c+u_{t-\ell}). \] Since $\{y_{t}\}$ has a stochastic trend, it will tend to make lengthy sojourns above the origin, during which periods $\xi_{t}$ will be well approximated by the stationary linear process, \[ \xi_{t}^{+}\coloneqq-(\beta^{+\top}\alpha)^{-1}\beta^{+\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{+\top}\alpha)^{\ell}\beta^{+\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{+}+w_{t}^{+} \] On the other hand, $\{y_{t}\}$ will also tend to spend lengthy epochs below the origin, permitting $\xi_{t}$ to then be approximated by \[ \xi_{t}^{-}\coloneqq-(\beta^{-\top}\alpha)^{-1}\beta^{-\top}c+\sum_{\ell=0}^{\infty}(I_{r}+\beta^{-\top}\alpha)^{\ell}\beta^{-\top}u_{t-\ell}\eqqcolon\mu_{\xi}^{-}+w_{t}^{-}. \]

This reasoning suggests a kind of `dual linear process' approximation to $\xi_{t}$, leading to an argument along the lines of

align*[align* omitted — 571 chars of source]

where $m_{Y}^{+}(\lambda)$ measures the fraction of the interval $[0,\lambda]$ for which $Y(s)\geq0$. We thus arrive at

align*[align* omitted — 342 chars of source]

which will in general be random (so that the convergence is merely in distribution), except in the special case where $\ensuremath{\mathbb{E}} g(\xi_{0}^{+})=\ensuremath{\mathbb{E}} g(\xi_{0}^{-})=\mu_{g}$ -- whereupon the r.h.s.\ collapses to $\lambda\mu_{g}$, since $m_{Y}^{+}(\lambda)+m_{Y}^{-}(\lambda)=\lambda$. (Importantly for the purposes of our test, such a case systematically arises under our assumptions, when $g(\xi)=\xi$.) The randomness of the limit provides another manifestation of the non-ergodicity of $\{\xi_{t}\}$, induced as by the dependence of its law of motion on the level of $y_{t}$.

Such arguments, in the more general setting of a (not necessarily canonical) CKSVAR($k$), lead to the main technical contribution of this paper, a LLN-type result for additive functionals of a class of time-varying autoregressive processes, of which (ref) is a special case. To facilitate its use in other contexts, we prove this result supposing that the following weaker condition holds in place of (ref).

{{{{\scalefont{0.76}DET$^{\prime}$}}}}

assumption$e_{1}^{\top}P_{\beta_{\perp}}(+1)c=0$.

The preceding permits the model to impart deterministic trends to $x_{t}$ (but not to $y_{t}$), and leads us to consider the linearly detrended process \[

bmatrix[bmatrix omitted — 35 chars of source]

=z_{t}^{d}\coloneqq z_{t}-[P_{\beta_{\perp}}(+1)c]t,\quad t\geq1 \] in place of $z_{t}$, with the convention that $z_{t}^{d}\coloneqq z_{t}$ for $t\leq0$; note that $y_{t}^{d}=y_{t}$ (see Section 4.4 in \citetalias{DMW22}). Recall that, as per the remarks following the statement of (ref) above, there is an underlying filtration $\{\mathcal{F}_{t}\}_{t\in\mathbb{Z}}$ to which $\{u_{t}\}$ and $\{z_{t}\}$ are adapted, and that an i.i.d.\ process $\{v_{t}\}$ is one that is both $\mathcal{F}_{t}$-adapted, and such that $v_{s}$ is independent of $\mathcal{F}_{t}$ for $s>t$.

thmSuppose (ref), (ref), (ref) and (ref) hold. Let $\{A_{t}\}$, $\{B_{t}\}$ and $\{c_{t}\}$ be random sequences adapted to $\{\mathcal{F}_{t}\}$, respectively taking values in $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$, where $t\in\mathbb{Z}$. Suppose $\{v_{t}\}$ is i.i.d\ with $\ensuremath{\mathbb{E}} v_{t}=0$, and that $\{w_{t}\}$ satisfies \begin{equation} w_{t}=c_{t}+A_{t}w_{t-1}+B_{t}v_{t} \end{equation} for $t\geq-k$ and some given (random) $w_{-k}$ (with $w_{t}\coloneqq0$ for all $t\leq-k-1$); and: \begin{enumerate}[itemsep=2pt,topsep=3pt] • $A_{t}\in\mset A$, $B_{t}\in\mset B$ and $c_{t}\in{\cal C}$ for all $t\in\ensuremath{\mathbb{N}}$, where $\mset A$, $\mset B$ and ${\cal C}$ are bounded subsets of $\mathbb{R}^{d_{w}\times d_{w}}$, $\mathbb{R}^{d_{w}\times d_{v}}$ and $\mathbb{R}^{d_{w}}$ respectively, and $\rho_{{\scriptstyle \mathrm{JSR}}}(\mset A)<1$; • there exist $A^{\pm}\in\mset A$, $B^{\pm}\in\mset B$ and $c^{\pm}\in\mset C$ such that \begin{align*} y_{t-1}>0 and y_{t}>0 & \implies A_{t}=A^{+},\ B_{t}=B^{+},\ c_{t}=c^{+},\\ y_{t-1}<0 and y_{t}<0 & \implies A_{t}=A^{-},\ B_{t}=B^{-},\ c_{t}=c^{-}; \end{align*} • $m_{0}\geq1$ is such that $\smlnorm{w_{0}}_{m_{0}}+\smlnorm{v_{0}}_{m_{0}}<\infty$. • $g:\mathbb{R}^{d_{w}}\ensuremath{\rightarrow}\mathbb{R}^{d_{g}}$ is a continuous function satisfying \begin{equation} \smlnorm{g(w)-g(w^{\prime})}\leq C(1+\smlnorm w^{\ell_{0}}+\smlnorm{w^{\prime}}^{\ell_{0}})\smlnorm{w-w^{\prime}} \end{equation} for all $w,w^{\prime}\in\mathbb{R}^{d_{w}}$, for some $0\leq\ell_{0}<m_{0}-1$. \end{enumerate} Then $\ensuremath{\mathbb{E}}\smlnorm{g(w_{0}^{\pm})}<\infty$, and on $D[0,1]$, \begin{equation} \frac{1}{n}\sum_{t=1}^{\smlfloor{n\lambda}}g(w_{t})\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\int_{0}^{\lambda}\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} where \begin{equation} w_{0}^{\pm}=(I_{d_{w}}-A^{\pm})^{-1}c^{\pm}+\sum_{\ell=0}^{\infty}(A^{\pm})^{\ell}B^{\pm}v_{-\ell}. \end{equation} Moreover, \begin{equation} \frac{1}{n^{3/2}}\sum_{t=1}^{\smlfloor{n\lambda}}[g(w_{t})\otimes z_{t}^{d}]\ensuremath{\mathbf{1}}^{\pm}(y_{t})\ensuremath{\rightsquigarrow}[\ensuremath{\mathbb{E}} g(w_{0}^{\pm})]\otimes\int_{0}^{\lambda}Z(\mu)\ensuremath{\mathbf{1}}^{\pm}[Y(\mu)]\ensuremath{\,\ensuremath{\mathrm{d}}}\mu, \end{equation} jointly with $U_{n}\ensuremath{\rightsquigarrow} U$.

Limiting distribution and consistency

Using (ref) and the representation theory of \citetalias{DMW22}, we are able to derive the limiting distribution of our modified Breitung (MB) statistic for testing the null of $q_{0}$ common trends (and $r_{0}=p-q_{0}$ cointegrating relations), which is defined as

equation[equation omitted — 100 chars of source]

where $\{\lambda_{n,i}\}_{i=1}^{p+1}$ are the solutions to

equation[equation omitted — 81 chars of source]

ordered as $\lambda_{n,1}\leq\lambda_{n,2}\leq\cdots\leq\lambda_{n,p+1}$, for

align[align omitted — 224 chars of source]

This statistic has the same form as that considered in (ref), though note that for testing the null of $q_{0}$ common trends we sum over the first $q_{0}+1$ generalised eigenvalues $\{\lambda_{n,i}\}_{i=1}^{q_{0}+1}$, reflecting the fact that $y_{t}^{+}$ and $y_{t}^{-}$ separately enter $z_{t}^{\ast}$.

To state the limiting distribution of the test statistic, define

equation[equation omitted — 86 chars of source]

where ${\cal W}_{0}\in\mathbb{R}$ is nonrandom, and $W$ is a $q$-dimensional standard Brownian motion. Define the $(q+1)$-dimensional process

equation[equation omitted — 8,324 chars of source]