EconBase
← Back to paper

Canonical correlation analysis of stochastic trends via functional approximation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,684 characters · 19 sections · 82 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Canonical correlation analysis of stochastic trends via functional approximation

{1.5\defbaselineskip}

abstractThis paper proposes a novel approach for semiparametric inference on the number $s$ of common trends and their loading matrix $\psi$ in $I(1)/I(0)$ systems. It combines functional approximation of limits of random walks and canonical correlations analysis, performed between the $p$ observed time series of length $T$ and the first $K$ discretized elements of an $L^2$ basis. Tests and selection criteria on $s$, and estimators and tests on $\psi$ are proposed; their properties are discussed as $T$ and $K$ diverge sequentially for fixed $p$ and $s$. It is found that tests on $s$ are asymptotically pivotal, selection criteria of $s$ are consistent, estimators of $\psi$ are $T$-consistent, mixed-Gaussian and efficient, so that Wald tests on $\psi$ are asymptotically Normal or $\chi^2$. The paper also discusses asymptotically pivotal misspecification tests for checking model assumptions. The approach can be coherently applied to subsets or aggregations of variables in a given panel. Monte Carlo simulations show that these tools have reasonable performance for $T\geq 10 p$ and $p\leq 300$. An empirical analysis of 20 exchange rates illustrates the methods.

\begingroup \def\uppercasenonmath#1 \endgroup

Introduction

The analysis of multiple time series with nonstationarity and comovements has been steadily developing since the introduction of the notion of cointegration in EG:87, and the seminal contributions of SW:88, Joh:88, Joh:91, AR:90, Phillips1990; see Wat:94 for an early survey. These contributions discussed asymptotic results for diverging sample size $T$ and fixed cross-sectional dimension $p$. Associated applications considered a cross-sectional dimension $p$ in the low single digits.

A central parameter in this context is the $p\times r$ cointegration matrix $\beta$ and its rank $r$, the cointegration rank; inference on these quantities can be approached either parametrically, as in the context of Vector Autoregressive (VAR) processes, see Joh:91, Joh:96, or semiparametrically, as in Phillips1990, and Bie:97.

Data with $p$ in the double digits have recently become pervasive. The label `moderate dimensional' has been used when such data are analyzed assuming a fixed $p$ in the asymptoptics for diverging $T$, see CS:24; the present paper falls into this framework. Contributions for the moderate dimensional case include modifications of parametric procedures, such as Bartlett corrections, see Joh:00,Joh:02b,Joh:02, bootstrap implementations, see Swe:06, CRT:12 and CNR:15, and Lasso-type procedures, see CS:24.

One alternative stream of literature, called `high dimensional', considers $p$ and $T$ diverging proportionally; leading contributions are OW:18,OW:19, BG:22,BG:24, LS:19. Finally, the cross sectional dimension $p$ can also be infinite from the start, as in functional time series, see PZK22, NSS:22. Combinations of the moderate and high dimensional cases have been studied in ZRY:19, BCT:22, FGP:23.

Contributions and relation with the literature

The present paper provides novel inferential tools for semiparametric inference on $I(1)/I(0)$ $p$-dimensional vectors $X_t$ that admit the Common Trends (CT) representation $$ X_t=\gamma+\psi\kappa'\sum_{i=1}^t\varepsilon_i +C_1(L)\varepsilon_t,\qquad t=1,2,\dots, $$ where $\gamma$ is a vector of initial values, $\kappa'\sum_{i=1}^t\varepsilon_i$ are $0\leq s\leq p$ stochastic trends, $\psi $ is a $p\times s$ full column rank loading matrix, $L$ is the lag operator and $C_1(L)\varepsilon_t$ is an $I(0)$ linear process.

The analysis can be coherently applied to subsets or aggregations of variables. This fact rests on the property, called dimensional coherence, that linear combinations of the observables $X_t$ also admit a common trends representation with a number of stochastic trends that is at most equal to the one of the original system.

The proposed approach is based on the empirical canonical correlations between the $p$ observed variables $\{X_t\}_{t=1}^T$ and $K$ deterministic variables $\{d_t\}_{t=1}^T$, constructed as the first $K$ elements of an orthonormal $L^2[0,1]$ basis discretized over the equispaced grid $1/T,2/T,\dots,1$. In the asymptotic analysis, the cross-sectional dimension $p$ is kept fixed while $T$ and $K$ diverge sequentially, $K$ after $T$, denoted by $(T,K)_{seq}\rightarrow\infty$.

The techniques in the present paper build on the results in Phi:98,Phi:05,Phi:14, where functional approximations with an increasing number $K$ of $L^2[0,1]$ basis elements are used, respectively, for understanding regressions between $I(1)$ and deterministic variables, for the heteroskedasticity and autocorrelation-consistent estimation of long-run variance (LRV) matrices and for efficient instrumental-variables (IV) estimation of cointegrating vectors.

Related approaches based on a fixed number $K$ of cosine functions are presented in Bie:97 and in MW:08,MW:17,MW:18, who respectively estimate the number of stochastic trends and the cointegrating relations using an IV approach, and analyze long-run variability and covariability of filtered time series.

The present study combines canonical correlation analysis (cca) and the functional approximation of limits of random walks. It is shown that the $s$ largest squared canonical correlations (scc) tend to one while the $p-s$ smallest scc\ tend to zero for any cross-sectional dimension $p$, for any number of stochastic trends $0\leq s\leq p$ and for any choice of $L^2[0,1]$ basis, all fixed, as $(T,K)_{seq}\rightarrow\infty$. This implies that a criterion based on the maximal gap between scc\ selects the correct number of common trends consistently. The same result applies to criteria designed following Bie:97 and AH:13.

The cause of these results is that the stochastic trends in $X_t$ converge to Brownian motions as $T$ diverges while the deterministic variables $d_t$ used in the cca\ converge to the first $K$ elements of an $L^2[0,1]$ basis; Brownian motions have an $L^2$-representation (unlike $I(0)$ components), and in the $K$-limit they are completely recovered. This implies that the largest scc\ are associated to the nonstationary directions in $X_t$ and the smallest ones to the stationary linear combinations.

This is the opposite of what happens in the cases analyzed in OW:18,OW:19 and in BG:22,BG:24, which involves the cca\ between differences $\Delta X_t$ and lagged levels $X_{t-1}$ (possibly correcting for lagged differences and deterministic terms) derived from the Gaussian Quasi Maximum Likelihood (QML) estimator in cointegrated VARs, see Joh:88,Joh:91 and AR:90.

The present approach also differs from the one in FGP:23, who discuss the cca\ between levels $X_t$ and their cumulations $\sum_{i=1}^{t}X_i$, and consider an asymptotic regime in which $T$ and $p$ diverge sequentially. The advantage of the proposal here is that it leads to efficient asymptotic inference on the loading matrix $\psi$, as well as to tests on $s$ and misspecification tests, as explained next.

The canonical variates in the cca\ are used to define estimators of the space spanned by the loading matrix $\psi$ and of the {cointegrating space} spanned by the cointegrating matrix $\beta$; these estimators are shown to be $T$-consistent for any cross-sectional dimension $p$, for any number of stochastic trends $0\leq s\leq p$ and for any basis of $L^2[0,1]$, all fixed. When the cca\ is iterated once in a specific way, for $(T,K)_{seq}\rightarrow\infty$ the associated estimators of $\psi$ and $\beta$ are asymptotically mixed-Gaussian and efficient, as in Joh:96 and in Phi:05,Phi:14. The mixed-Gaussian asymptotic distribution is then used to construct asymptotically Normal and $\chi^2$ Wald test statistics for hypotheses on $\psi$ and $\beta$. Identification of $\psi$ and $\beta$ is discussed and an inferential criterion for the validity of the chosen identification condition is proposed.

Further asymptotic results are derived for the specific basis in the Karhunen-Lo\`eve\ (KL) representation of a Brownian Motion, see LoII:78. Using the KL basis, canonical correlations are used (i) to estimate $s$ by a top-down sequences of asymptotically pivotal tests for $H_0:s=j$ versus $H_1:s<j$, for $j=p,p-1,\dots,1$ and (ii) to perform misspecification tests for model assumptions. The sequential tests are shown to select the correct number of stochastic trends with limit probability equal to 1 minus the size of each test in the sequence, similarly to what happens for sequences of likelihood ratio (LR) tests in the VAR case, see Joh:96.

Organization and notation

The remainder of the paper is organized as follows. Section (ref) discusses assumptions on the Data Generating Process (DGP) and Section (ref) defines the main estimators. Section (ref) discusses the asymptotic properties of estimators of the number of stochastic trends and Section (ref) those of the loading and the cointegrating matrices. Section (ref) contains the Monte Carlo study, Section (ref) presents an empirical application to exchange rates and Section (ref) concludes. The Supplementary material contains lemmas, proofs and additional material on simulations and on the empirical application.

Concerning notation, for any $a \in \mathbb{R}$, $\lfloor a \rfloor$ and $\lceil a \rceil$ denote the floor and ceiling functions; for any condition $c$, $1_{c}$ is the indicator function of $c$, taking value 1 if $c$ is true and 0 otherwise; $\iota_{n}$ indicates the $n \times 1$ vectors of ones. The space of right continuous functions $f:[0,1]\mapsto\mathbb{R}^n$ having finite left limits, endowed with the Skorokhod topology, is denoted by $D_n[0,1]$, see JS:03 for details, also abbreviated to $D[0,1]$ if $n=1$. A similar notation is employed for the space of square integrable functions $L_n^2[0,1]$, abbreviated to $L^2[0,1]$ if $n=1$. Riemann integrals $\int_0^1 a(u)b(u)'\ensuremath{\mathrm{d}} u$, Ito stochastic integrals $\int_0^1\ensuremath{\mathrm{d}} a(u)b(u)'$ and Stratonovich stochastic integrals $\int_0^1\partial a(u)b(u)'$ are abbreviated respectively as $\int ab'$, $\int\ensuremath{\mathrm{d}} a b'$ and $\int\partial a b'$.

Generating process, identification and $L^2$-representation

This section describes the class of processes studied in the paper and their limiting $L^2$-representation; it also discusses identification of the parameters of interest.

Data Generating Process

Let $\{X_t\}_{t\in \mathbb{Z}}$ be a $p\times 1$ linear process generated by

equation[equation omitted — 58 chars of source]

where $L$ is the lag operator, $\Delta:=1-L$, $C(z)=\sum_{n=0}^\infty C_n z^n$, $z\in \mathbb{C}$, satisfies Assumption $\ref{ass_DGP}$ below, and the innovations $\{\varepsilon_t\}_{t\in \mathbb{Z}}$ are independently and identically distributed (i.i.d.) with expectation $\mathbb{E}(\varepsilon_t)=0$, finite moments of order $2+\epsilon$, $\epsilon>0$, and positive definite variance-covariance matrix $\mathbb{V}(\varepsilon_t)=\Omega_\varepsilon$, indicated as $\varepsilon_t\:\text{i.i.d.}\: (0,\Omega_\varepsilon)$, $\Omega_\varepsilon>0$.

Assumption $\ref{ass_DGP}$ below concerns the first-order expansion of $C(z)$, $C(z)=C+C_1(z)(1-z)$, and the rank $s:=\operatorname{rank} C(1)=\operatorname{rank} C$. For $s>0$, let $\breve\psi$ denote a matrix whose columns form a basis of $\operatorname{col} C$; then $\breve\kappa$ is the unique matrix such that $C=\breve\psi\breve\kappa'$.\footnote{For any matrix $a$, $\operatorname{col} a$ indicates the linear space spanned by its columns; when $(\operatorname{col} a)^\bot\neq\{0\}$, $a_\bot$ is used to indicate any basis of $(\operatorname{col} a)^\bot$.} Moreover, for $s<p$, let $\breve\beta:=\breve\psi_\bot$ and $\breve\kappa_\bot$ be matrices whose columns form bases of $(\operatorname{col} C)^{\bot}$ and $(\operatorname{col} C')^{\bot}$ respectively; thus, for $0<s<p$, it holds that $(\operatorname{col} \breve \psi)^{\bot}=\operatorname{col}(\breve\psi_\bot)$ and likewise for $\breve\kappa$. The space $\operatorname{col} C$ is called the {attractor space} and its orthogonal complement $(\operatorname{col} C)^\bot$ is the {cointegrating space} of $X_t$, of dimension $0\leq r:=p-s\leq p$ (the cointegrating rank).

ass[Assumptions on $C(z)$] $C(z)$ in (ref) is of dimension $p\times n_\varepsilon$, $n_\varepsilon \geq p$, satisfying \begin{enumerate} • $C(z)=\sum_{n=0}^\infty C_n z^n$ converges for all $|z|<1+\delta$, $\delta >0$; • $\operatorname{rank} C(z)<p$ may only hold at isolated points $z$, where either $z=1$ or $|z|>1$; • if $s<p$, then $\breve\psi_\bot'C_1\breve\kappa_\bot$ has full row rank $r:=p-s$, with $C_1:=C_1(1)$. \end{enumerate}

Assumption $\ref{ass_DGP}$ includes possibly nonstationary VARMA processes (for $n_\varepsilon = p$) in line with e.g. SW:88 and Dynamic Factor Models (for $n_\varepsilon > p$) as in e.g. Bai:04.

Cumulating (ref), one finds the Common Trends (CT) representation

equation[equation omitted — 128 chars of source]

which gives a decomposition of $X_t$ into initial conditions $\gamma:=X_0-C_1(L)\varepsilon_0$, the $s$-dimensional random walk component $\breve\kappa'\sum_{i=1}^t\varepsilon_i$ with loading matrix $\breve\psi$ and the stationary component $C_1(L)\varepsilon_t$. Observe that $X_t$ is $I(0)$ if $s=0$, i.e. $r=p$; it is $I(1)$ and cointegrated if $0<r,s<p$ and it is $I(1)$ and non-cointegrated if $s=p$, i.e. $r=0$.

The aim of the paper is to make inferences about $s$ and $\breve\psi$ (and on their complements $r$ and $\breve\beta$) from a sample $\{X_t\}_{t = 0,1,\dots,T}$ observed from (ref).

remark[Initial conditions] Definition 3.4 in Joh:96 stipulates that a vector $v\neq 0$ is a cointegrating vector if $v' X_t$ can be made stationary by a suitable choice of its initial distribution; in the present case, this corresponds to setting $\breve\beta'\gamma=0$, i.e. $\breve\beta'X_0=\breve\beta'C_1(L)\varepsilon_0$, so that $\gamma:=X_0-C_1(L)\varepsilon_0=\breve\psi\gamma_0$ and hence \begin{equation} X_t=\breve\psi\left(\gamma_0+\breve\kappa'\sum_{i=1}^t\varepsilon_i\right)+C_1(L)\varepsilon_t,\qquad t=1,2,\dots\:. \end{equation} Alternatively, as in Ell:99, see also MW:08 and OW:18, one can consider \begin{equation} X_t-X_0=\breve\psi\breve\kappa'\sum_{i=1}^t\varepsilon_i +C_1(L)(\varepsilon_t-\varepsilon_0),\qquad t=1,2,\dots\:. \end{equation}

The statistical analysis is performed on $x_t:=X_t$ or $x_t:=X_t-X_0$ according to whether specification (ref) or (ref) is chosen.

Identification

The parameters $\breve\psi$, $\breve\kappa$ and $\breve\beta$ are not identified; in fact $C=\breve\psi\breve\kappa'= (\breve\psi a) (a^{-1}\breve\kappa')$ for any nonsingular $a\in \mathbb{R}^{s\times s}$, so that $\breve\psi$ and $\breve\psi a$ are bases of the {attractor space} $\operatorname{col} C$, and $\breve\kappa$ and $\breve\kappa (a^{-1})'$ of $\operatorname{col} (C')$; similarly, $\breve\beta$ and $\breve\beta a$ with $a\in \mathbb{R}^{r\times r}$ nonsingular are bases of the {cointegrating space} $(\operatorname{col} C)^\bot$.

As done in Joh:96 for $\breve\beta$, just-identification of the loading matrix $\breve\psi$ is achieved by means of linear restrictions, employing a known and full column rank matrix $b \in \mathbb{R}^{p\times s}$ such that $b'\breve\psi\in \mathbb{R}^{s\times s}$ is nonsingular. Similarly, just-identification of the cointegrating vectors is obtained using a known and full column rank matrix $c\in \mathbb{R}^{p\times r}$ such that $c'\breve\beta\in \mathbb{R}^{r\times r}$ is nonsingular, where $c$ is a basis of $(\operatorname{col} b)^\bot$. The following Lemma discusses the relationship between the identified parameters obtained in this way; here $\bar a:=a(a'a)^{-1}$ for a full column rank matrix $a$ and $\breve \psi$ and $\psi$ (respectively $\breve \beta$ and $\beta$) are bases of $\operatorname{col} C$ (respectively $\operatorname{col} (C)^\perp$).

lemma[Duality] Let $C=\breve\psi\breve\kappa'$ be a rank factorization of $C$, let $b \in \mathbb{R}^{p\times s}$ be such that $b'\breve\psi\in \mathbb{R}^{s\times s}$ is nonsingular and let $c \in \mathbb{R}^{p\times r}$ be a basis of $(\operatorname{col} b)^\bot$; then \begin{enumerate} • for $s>0$, $\psi:=\breve\psi( b'\breve\psi)^{-1}$ and $\kappa:=\breve\kappa(\breve\psi' b)=C' b$ are identified as the unique solution of the system of equations $C=\psi\kappa'$ and $b'\psi=I_s$; • for $s<p$, $c'\breve\beta\in \mathbb{R}^{r\times r}$ is nonsingular and $\beta:=\breve\beta( c'\breve\beta)^{-1}$ is identified as the unique solution of the system of equations $\beta'C=0$ and $c'\beta=I_r$; • for $0<s<p$, the unrestricted coefficients $\psi_*:=\bar c'\psi$ and $\beta_*:=\bar b'\beta$ satisfy $\psi_*=-\beta_*'$; • for any estimator $\wt \psi$ of $\psi$ normalized as $ b'\wt \psi = b'\psi = I_{s}$ and any estimator $\wt \beta$ of $\beta$ normalized as $c'\wt \beta = c'\beta = I_{r}$, one has $\bar c'(\wt \psi -\psi) = \beta'(\wt \psi -\psi)$ and similarly $\bar b'(\wt \beta -\beta) = \psi'(\wt \beta -\beta)$. \end{enumerate}

Point (iii) shows that the unrestricted coefficients in the identified parameters are linearly dependent, $\psi_*=-\beta_*'$; this implies a duality between hypotheses on $\psi$ and $\beta$, see (ref) below.

Importantly, one does not know a priori which restrictions are just-identifying, i.e., which $b$ delivers a nonsingular $b'\breve\psi$. Take for instance $b=(I_s,0)'$: because $\breve\psi$ has rank $s$, there is always at least one permutation of the order of variables for which these restrictions are just-identifying; however, one does not know which permutation should be used. A consistent decision rule on the nonsingularity of $b'\breve\psi$ is proposed in Section (ref) below.

The choice of a specific $b$ implies a unique space $\operatorname{col} (b)^\bot = \operatorname{col} c$, but leaves the choice of the specific basis $c$ open, so that $c$ can be chosen to best fit the application. Observe also that the choice of $b$ and then of $c$ can be reversed, by interchanging the role of $(\breve\psi, b)$ and $(\breve\beta, c)$ in Lemma $\ref{lem_ident}$. In the remainder of the paper it is assumed that $\psi$, $\kappa$ and $\beta$ in (ref) are identified by the choice of $b$ and $c=b_\bot$ such that $b' \psi = I_s$, $c' \beta = I_r$ and $\kappa= C' b$.

Dimensional coherence

This section shows that whenever $X_t$ has a CT representation (ref), also linear combinations of $X_t$ admit a CT representation, with a number of stochastic trends that is at most equal to that in $X_t$. This property, called dimensional coherence, is formally stated in the next theorem.

theorem[Dimensional coherence] Let $H$ be a $p\times m$ full column rank matrix; if $X_t$ satisfies (ref) with $C(z)$ fulfilling Assumption $\ref{ass_DGP}$, then the same holds for $H'X_t$, i.e. $\Delta H'X_t=G(L)\varepsilon_t$, where $G(z):=H'C(z)$ satisfies Assumption $\ref{ass_DGP}$ with $G(1)=H'C$ of rank $q\leq s$.

Special cases of $H$ are the selection matrix $H=(I_m,0)'$ of the first $m$ variables, and the group-aggregation matrix $H=(I_{q}\otimes\iota_{n})$, where $p=qn$, $m=q$ and $\iota_{n}$ is an $n \times 1$ vector of ones. One consequence of Theorem (ref) is that the semiparametric approach described in this paper applies to subsets or aggregations of the original variables $X_t$; this property is exploited in the empirical application in Section $\ref{sec_E}$.

Functional central limit theorem and long-run variance

By the functional central limit theorem, see PS:92, the partial sums $T^{-\frac12}\sum_{i=1}^t\varepsilon_i$ converge weakly to an $n_{\varepsilon}$-dimensional Brownian Motion $W_\varepsilon(u)$, $u\in[0,1]$, with variance $\Omega_\varepsilon>0$; by the CT representation (ref), one hence has that

equation[equation omitted — 441 chars of source]

where $t =\lfloor Tu\rfloor\in \mathbb{N}$ and the $o_p(1)$-term is infinitesimal uniformly in $u\in[0,1]$. Here $\stackrel{w}{\rightarrow}$ indicates weak convergence of probability measures on $D_p[0,1]$ and $W(u):=( W_1(u)', W_2(u)')'$ is a $p\times 1$ Brownian motion with variance $\Omega:=D\Omega_\varepsilon D'$, $D:=(\kappa,C_1'\beta)'$ and $\beta=\psi_{\bot}$.

Remark that long-run variance (LRV) $\Omega$ is positive definite because $D$ is nonsingular by Assumption $\ref{ass_DGP}$.(iii); in fact, $$ \operatorname{rank} D= \operatorname{rank}\left(

array[array omitted — 33 chars of source]

\right) (\bar\kappa,\kappa_\bot) =\operatorname{rank}\left(

array[array omitted — 68 chars of source]

\right)=s+\operatorname{rank}\beta'C_1\kappa_\bot. $$ Partition $\Omega$ conformably to $W(u)=( W_1(u)', W_2(u)')'$, namely

equation[equation omitted — 323 chars of source]

and define the independent Brownian motions

equation[equation omitted — 134 chars of source]

with nonsingular variance matrices $\Omega_{22.1}:=\Omega_{22}-\Omega_{21}\Omega_{11}^{-1}\Omega_{12}$ and $I_s$ respectively.

$L^2$-representation

Recall, see e.g. LoII:78 section 37.5B, that the $p$-dimensional Brownian motion $W(u)$ admits the representation

equation[equation omitted — 201 chars of source]

where $\{\phi_k(u)\}_{k=1}^\infty$, $\int_0^1\phi_j(u)\phi_k(u)\ensuremath{\mathrm{d}} u=1_{j=k}$, is an orthonormal basis of $L^2[0,1]$ and $\simeq$ indicates that the series in (ref) is a.s. convergent in the $L^2$ sense to the l.h.s..

In the special case where $(\nu_k^2,\phi_k(u))$, $k=1,2,\dots$, is an eigenvalue-eigenvector pair of the covariance kernel of the standard Brownian motion, i.e. $\nu_k^2\phi_k(u)=\int_0^1\min(u,v)\phi_k(v)\ensuremath{\mathrm{d}} v$, (ref) is the {Karhunen-Lo\`eve} (KL) representation of $W(u)$, for which one has

equation[equation omitted — 151 chars of source]

In this case the series in (ref) is a.s. uniformly convergent in $u$ and $\simeq$ is replaced by $=$. In the following, a basis $\{\phi_k(u)\}_{k=1}^\infty$ is indicated as the {KL basis} when $\phi_k(u)$ is chosen as in (ref). Expansion (ref) underlies the large-$K$ asymptotic theory in this paper; see Appendix (ref) for details and references.

Definition of the estimators

This section presents the proposed estimators of the number $s$ of stochastic trends and the cointegrating rank $r=p-s$, as well as of $\psi$ (the identified basis of the {attractor space}) and $\beta$ (the identified basis of the {cointegrating space}).

Let $\varphi_K(u):=(\phi_1(u),\dots,\phi_K(u))'$ be a $K \times 1$ vector function of $u \in[0,1]$, where $\phi_1(u),\dots,\phi_K(u)$ are the first $K$ elements of some fixed orthonormal c\`adl\`ag\ basis of $L^2[0,1]$, see (ref). Let $d_t$ be the $K\times 1$ vector constructed by evaluating $\varphi_K(\cdot)$ at the discrete sample points $1/T,\dots, (T-1)/T, 1$, i.e.

equation[equation omitted — 116 chars of source]

For a generic $p$-dimensional variable $f_t$ observed for $t=1,\dots,T$, the sample canonical correlation analysis of $f_t$ and $d_t$ in (ref), denoted as $\textsc{cca}(f_t,d_t)$, consists in solving the following generalized eigenvalue problem, see e.g. Joh:96 and references therein,

equation[equation omitted — 143 chars of source]

This delivers eigenvalues $1\geq \lambda_1\geq\lambda_2\geq\dots\geq\lambda_p \geq 0$, collected in $\Lambda=\operatorname{diag}(\lambda_1,\dots,\lambda_p)$, and corresponding eigenvectors $V=(v_1,\dots,v_p)$, organized as

equation[equation omitted — 215 chars of source]

with the convention that $(\Lambda_1,V_1)=(\Lambda,V)$ when $s=p$ and $(\Lambda_0,V_0)=(\Lambda,V)$ when $s=0$.

defn[Max-gap estimators of $s$ and $r$] Define the maximal gap $($max-gap$)$ estimator of $s$ as \begin{equation} \wh s :=\underset{i\in \{0,\dots, p\}}{\operatorname{argmax}}(\lambda_i-\lambda_{i+1}),\qquad\lambda_0:=1,\qquad\lambda_{p+1}:=0, \end{equation} where $1\geq \lambda_1\geq\lambda_2\geq\dots\geq\lambda_p \geq 0$ are the eigenvalues of $\textsc{cca}(x_t,d_t)$, $x_t$ is defined after (ref) in Remark $\ref{rem_init_cond}$; further define the estimator of $r$ as $\wh r:=p-\wh s$.

The choice $\lambda_0:=1$ and $\lambda_{p+1}:=0$ permits the calculation of $\lambda_i-\lambda_{i+1}$ also for $i=0$ and $i=p$, and it is motivated by the results in Theorem $\ref{theorem_lim_eig}$ below.

First stage estimators of $\psi$ and $\beta$ are defined in terms of the eigenvectors in $\textsc{cca}(x_t,d_t)$, and they are indicated as $\wh\psi^{(1)}$ and $\wh\beta^{(1)}$. The Iterated Canonical Correlation (ICC) estimators are denoted as $\wh\psi$ and $\wh\beta$, and they are defined in terms of the eigenvectors in $\textsc{cca}(e_t,d_t)$, where $e_t$ denotes the residual of $x_t$ regressed on the fitted values from the regression of $\wh\psi^{(1)'}\Delta x_t$ on $d_t$. The motivation for the ICC estimation is to obtain mixed-Gaussian limit distributions, see Theorem $\ref{eq_asy_iter}$ below.

defn[Estimators of $\psi$ and $\beta$] Let $0<s<p$.\footnote{For $s=0$ (resp. $s=p$), $\psi$ (resp. $\beta$) is undefined and $\beta=(c')^{-1}$ (resp. $\psi=(b')^{-1}$) is known.} Let $b \in \mathbb{R}^{p\times s}$ and $c= b_\bot \in \mathbb{R}^{p\times r}$ be a pair of matrices for the just-identification of $\breve\psi$ and $\breve\beta$, see Section $\ref{sec_ident}$, with $\psi:=\breve\psi( b'\breve\psi)^{-1}$, $\beta:=\breve\beta( c'\breve\beta)^{-1}$; for any $V$ as in (ref), the first stage estimators of $\psi$ and $\beta$ are defined as $$ \wh\psi^{(1)} := M_{xx} V_1( b'M_{xx} V_1)^{-1},\qquad\wh\beta^{(1)} := V_0( c' V_0)^{-1}, $$ where the eigenvectors $V$ are those of $\textsc{cca}(x_t,d_t)$. The ICC estimators are defined as $$ \wh\psi := M_{ee} V_1( b'M_{ee} V_1)^{-1},\qquad\wh\beta := V_0( c' V_0)^{-1}, $$ where the eigenvectors $V$ are those of $\textsc{cca}(e_t,d_t)$, with \begin{equation} e_t:=x_t-M_{xg}M_{gg}^{-1} g_t,\qquad g_t:=\wh\psi^{(1)'} M_{\Delta x d}M_{dd}^{-1}d_t. \end{equation}

The asymptotic analysis implies that the inverse matrices in Definition $\ref{def_phi_hat}$ are well-defined with probability approaching one. Note moreover that, (i), the estimators in Definition $\ref{def_phi_hat}$ are invariant to the normalization of the eigenvectors $V$ and, (ii), the ICC estimators $\wh\psi $ and $\wh\beta $ depend on $\wh\psi^{(1)}$ only through $\operatorname{col}(\wh\psi^{(1)})$.

Remark that the estimators in Definition $\ref{def_phi_hat}$ satisfy the identifying constraints of $\psi$ and $\beta$, i.e.

equation[equation omitted — 127 chars of source]

and that the duality in Lemma $\ref{lem_ident}$(iii) equally applies to them, i.e.

equation[equation omitted — 122 chars of source]

where $a_*:=\bar c'a$ and $h_*:=\bar b'h$, for $a=\psi,\wh\psi^{(1)},\wh\psi$ and $h=\beta,\wh\beta^{(1)},\wh\beta$.

Inference on the number of common trends

This section discusses the limits of the eigenvalues of $\textsc{cca}(x_t,d_t)$ and their implications for estimating the number $s$ of stochastic trends (and the number $r$ of cointegrating relations).

First, results are provided for the max-gap estimator and for alternative argmax estimators based on scc. Second, by employing the KL basis, several statistics are introduced (i) to perform misspecification tests for model assumptions and (ii) to define top-down sequences of asymptotically pivotal tests of $H_0:s=j$ versus $H_1:s<j$ for $j=p,p-1,\dots,1$, which provide alternative estimators of $s$.

Consistency

The asymptotic behavior of scc\ is given in the following theorem.

theorem[Limits of the eigenvalues of $\textsc{cca}(x_t,d_t)$] Let $W_2(u)$, $ B_1(u)$ be defined as in (ref), (ref), let $d_t$ in (ref) be constructed using any orthonormal c\`adl\`ag\ basis of $L^2[0,1]$ and let $1\geq \lambda_1\geq\lambda_2\geq\dots\geq\lambda_p \geq 0$ be the eigenvalues of $\textsc{cca}(x_t,d_t)$, see (ref) and (ref); then for fixed $K\geq p$ and $T\rightarrow\infty$, \begin{enumerate} • the $s$ largest eigenvalues $\lambda_1\geq\lambda_2\geq\dots\geq\lambda_s$ converge weakly to the ordered eigenvalues of \begin{equation} \Upsilon_K:=\left(\int B_1 B_1'\right)^{-1}\int B_1\varphi_K'\int\varphi_K B_1'\stackrel{a.s.}{>}0; \end{equation} • the $r$ smallest eigenvalues $\lambda_{s+1}\geq\lambda_{s+2}\geq\dots\geq\lambda_p$ multiplied by $T$ converge weakly to the ordered eigenvalues of \begin{equation} \Psi_K:=Q_{00}^{-1}\int\ensuremath{\mathrm{d}} W_2\varphi_K'Q_{\varphi\varphi}\int\varphi_K\ensuremath{\mathrm{d}} W_2'\stackrel{a.s.}{>}0, \end{equation} where $Q_{00}:=\mathbb{E}(\beta'x_t x_t'\beta)>0$ and $Q_{\varphi\varphi}:=I_K-\int\varphi_K B_1'\left(\int B_1\varphi_K'\int\varphi_K B_1'\right)^{-1}\int B_1\varphi_K'$. \end{enumerate} Moreover, \begin{enumerate} • for $K\rightarrow\infty$, $\Upsilon_K\stackrel{p}{\rightarrow} I_s$ and $K^{-1}\Psi_K\stackrel{p}{\rightarrow} Q_{00}^{-1}\Omega_{22}$; • for $(T,K)_{seq}\rightarrow\infty$, all the $s$ largest eigenvalues $\lambda_1\geq\lambda_2\geq\dots\geq\lambda_s$ converge in probability to $1$, while all the $r$ smallest eigenvalues $\lambda_{s+1}\geq\lambda_{s+2}\geq\dots\geq\lambda_p$ converge in probability to $0$. \end{enumerate}

Let $a \stackrel{p}\asymp g(T,K)$ for some function $g$ indicate that $a /g(T,K)$ is bounded and bounded away from 0 in probability as $(T,K)_{seq}\rightarrow\infty$; points (i), (ii) and (iii) in Theorem $\ref{theorem_lim_eig}$ imply that

equation[equation omitted — 147 chars of source]

for $(T,K)_{seq}\rightarrow\infty$. Because $K/T\rightarrow 0$ as $(T,K)_{seq}\rightarrow\infty$, this implies $\lambda_i \stackrel{p}{\rightarrow} 1 $ for $i=1,\dots,s$ while $\lambda_i \stackrel{p}{\rightarrow} 0 $ for $i=s+1,\dots,p$, see Theorem $\ref{theorem_lim_eig}$.(iv).

corollary[Consistency of the max-gap estimator $\wh s$] Let $\wh s$ be the max-gap estimator in (ref) computed on $\textsc{cca}(x_t,d_t)$ and let the assumptions of Theorem $\ref{theorem_lim_eig}$ hold; then $\mathbb{P}(\wh s = s)\rightarrow 1$ as $(T,K)_{seq}\rightarrow\infty$ for any $0\leq s \leq p$.

This result holds because Theorem $\ref{theorem_lim_eig}$.(iv) implies that $f_0(i):=\lambda_i-\lambda_{i+1}\stackrel{p}{\rightarrow} 1_{i=s}$ as $(T,K)_{seq}\rightarrow\infty$; note that the function $1_{i=s}$ provides the best discrepancy that the scc\ $\{\lambda_i\}_{i=1}^{p}$ can give between the $I(1)$ and $I(0)$ components.

In the context of eigenvalue problems different from the ones considered here, Bie:97 and AH:13 propose to estimate $s$ selecting the integer $i$ that maximizes some other function $f_j(i)$ of the eigenvalues, possibly using the rates of convergence of $\lambda_{i}$ in the definition of $f_j(i)$. These approaches, applied to the $\{\lambda_i\}_{i=1}^{p}$ eigenvalues from $\textsc{cca}(x_t,d_t)$, suggest the following additional estimators of $s$, called here alternative argmax estimators:

align[align omitted — 405 chars of source]

where $\mathcal{I}_1:=\{0,1,\dots, p\}$, $\mathcal{I}_2:=\{1,\dots, p-1\}$, $\mathcal{I}_3:=\{1,\dots, p-2\}$, empty sums (products) are equal to 0 (1), and $f_1(i)$ uses the rates in (ref) for $(T,K)_{seq}\rightarrow\infty$.

The estimators $\wt s^{(j)}$, $j=1,2,3$, optimize over the set of integers $\mathcal{I}_j$; note that $f_1$ is defined for $0\leq i \leq p$, $f_2$ for $1\leq i \leq p-1$ and $f_3$ for $1\leq i \leq p-2$. One may extend $f_2$ and $f_3$ to cover $i=0$ by defining $\lambda_0:=1$ as in (ref); this leads to replacing $\mathcal{I}_j$ in (ref) with the sets $\mathcal{I}_j^0:=\{0\}\cup\mathcal{I}_j$, $j=2,3$. Note however that adding $\lambda_{p+1}:=0$ is not an option, as $\lambda_{p+1}$ would appear in the denominators of $f_2$ and $f_3$. This implies that $f_2$ (respectively $f_3$) cannot be defined for $i=p$ (respectively $i=p-1,p$).

corollary[Consistency of the alternative argmax estimators of $s$] Let $\wt s^{(j)}$, $j=1,2,3$, be the alternative argmax estimators of $s$ in (ref) computed on $\textsc{cca}(x_t,d_t)$ and let the assumptions of Theorem $\ref{theorem_lim_eig}$ hold; then $\mathbb{P}(\wt s^{(j)}=s)\rightarrow 1$ for $(T,K)_{seq}\rightarrow\infty$ for any $s\in \mathcal{I}_j$, $j=1,2,3$. The same holds replacing $\mathcal{I}_j$ with $\mathcal{I}_j^0$ both in the optimization and in the class of DGPs.

These results are in line with the ones reported in Bie:97 and AH:13 in their respective setups.

Testing identification

Recall that the joint identification of $\psi$ and $\beta$ via $b$ and $c$ given in Lemma (ref) rests on the validity of the rank restriction $\operatorname{rank}( b'\breve\psi)=s$. Here a decision rule on the correctness of $\operatorname{rank}( b'\breve\psi)=s$ is proposed, using the property of dimensional coherence; let

equation[equation omitted — 148 chars of source]

The proposed decision rule consists of estimating the number of stochastic trends of $b'x_t$, and rejecting $H_0$ if the estimated number of stochastic trends is lower than the one estimated for $x_t$. The rationale for the criterion is the following: from (ref) one has $b'x_t= b'\breve\psi\breve\kappa'\sum_{i=1}^t\varepsilon_i + b'C_1(L)\varepsilon_t$, so that the number of stochastic trends in $b'x_t$ coincides with the one in $x_t$ if and only if $b'\breve\psi$ is nonsingular. Formally,

equation[equation omitted — 112 chars of source]

where $\wh s (x_t)$ (respectively $\wh s ( b' x_t)$) is the max-gap estimator calculated for $x_t$ (respectively $b'x_t$), see (ref).

theorem[Consistency of decision rule (ref)]If $H_0$ in (ref) holds, one has $\mathbb{P}(\wh s ( b'x_t) <\wh s (x_t))\rightarrow 0$ and $\mathbb{P}(\wh s ( b'x_t) =\wh s (x_t))\rightarrow 1$ as $(T,K)_{seq}\rightarrow\infty$. If $H_1$ in (ref) holds, one has $\mathbb{P}(\wh s ( b'x_t) <\wh s (x_t))\rightarrow 1$ and $\mathbb{P}(\wh s ( b'x_t) =\wh s (x_t))\rightarrow 0$ as $(T,K)_{seq}\rightarrow\infty$. The same statements are true substituting $\wh s$ with $\wt s^{(j)}$ provided $s\in \mathcal{I}_j$ (respectively $s\in \mathcal{I}_j^0$) both in the DGP and in the optimization.

Asymptotic distributions

This subsection discusses misspecification tests and test sequences for the estimation of $s$; these are derived using the specific {KL basis} in (ref) in the definition of $d_t$ in (ref).

theorem[Asymptotic distribution of the eigenvalues with the {KL basis}] Let $ B_1(u)$ be defined as in (ref), let $d_t$ in (ref) be constructed using the {KL basis} in (ref) and let $1\geq \lambda_1\geq\lambda_2\geq\dots\geq\lambda_p \geq 0$ be the eigenvalues of$\textsc{cca}(x_t,d_t)$; then for $(T,K)_{seq}\rightarrow\infty$, \begin{enumerate} • the $s$ largest eigenvalues $\lambda_1\geq\lambda_2\geq\dots\geq\lambda_s$ satisfy \begin{equation} K\pi^2\tau^{(s)}\stackrel{w}{\rightarrow}\zeta^{(s)}, \end{equation} where $\tau^{(s)}:=(1-\lambda_s,1-\lambda_{s-1},\dots,1-\lambda_1)'$, $\zeta^{(s)}:=(\zeta_1,\zeta_2,\dots,\zeta_s)'$, and $\zeta_1\geq\zeta_2\geq\dots\geq\zeta_s$ are the eigenvalues of $ \omega:=\left(\int B_1 B_1'\right)^{-1}; $ • the $r$ smallest eigenvalues $\lambda_{s+1}\geq\lambda_{s+2}\geq\dots\geq\lambda_p$ satisfy $K\pi^2 (1-\lambda_i)\stackrel{p}{\rightarrow}\infty$, $i=s+1,\dots,p$. \end{enumerate}

Observe that $\omega$ does not depend on nuisance parameters, and that tests based on $K\pi^2\tau^{(s)}$ are asymptotically pivotal. Further, the limit distribution in Theorem $\ref{theorem_asy_distr_s}$(i) depends on the tail $\sum_{k=K+1}^\infty \nu_k^2$, see (ref), and is therefore not invariant to the choice of an $L_2[0,1]$ basis.

The rest of this section discusses misspecification tests and sequences of tests on $s$ based on the results of Theorem $\ref{theorem_asy_distr_s}$.

remark[Confidence stripe for misspecification analysis] The results in Theorem (ref) can be used to construct confidence sets for misspecification analysis. For example, consider $f(a)=\log(a)-\mathbb{E}(\log\zeta^{(s)})$, $a \in \mathbb{R}^s$, where the $\log$ function is applied component-wise, and let $\|\cdot\|_n$ be the $n$-norm for vectors. Define the confidence misspecification stripe $\mathcal{B}:=\{a \in \mathbb{R}^s:\|f(a)\|_{\infty}<\delta \}$ with $\delta$, such that $\mathbb{P} (\zeta^{(s)}\in \mathcal{B})=1-\eta$ for some pre-assigned confidence level $1-\eta$. The continuous mapping theorem and Theorem (ref) then imply that $\mathbb{P}(K\pi^2\tau^{(s)}\in \mathcal{B})\rightarrow 1-\eta$ as $(T,K)_{seq}\rightarrow\infty$. The set $\mathcal{B}$ is a stripe around $\mathbb{E}(\log\zeta^{(s)})$ with equal deviations above or below $\mathbb{E}(\log(\zeta_i))$ across $i=1,\dots,s$, see Figure $\ref{fig_all_XR}$ below for an illustration. In practice, one can compute $\mathcal{B}$ and check if $K\pi^2\tau^{(\wh s)}\in \mathcal{B}$. If so, the model assumptions appear valid, including the selection of $s$; otherwise it is necessary to question which assumptions need to be modified. The same approach can be used for continuous transformations other than $\log\zeta_i$ and any well-defined location indicator, such as the median.
remark[Sequential tests on $s$] The results in Theorem (ref) can be used to test hypotheses of the following form \begin{equation} H_{0j}:s=j\qquadversus\qquad H_{1j} :s<j,\qquad j=p,p-1,\dots,1. \end{equation} Define the test statistics $F_{j,n}:=\|K\pi^2\tau^{(j)}\|_n$. Under $H_{0s}$, Theorem (ref) shows that $F_{s,n}\stackrel{w}{\rightarrow}\|\zeta^{(s)}\|_n$, while under $H_{1s}$, $F_{s,n}\rightarrow\infty$. For $n=1,\infty$, one finds \begin{equation} F_{j,1}=\|K\pi^2\tau^{(j)}\|_1=K\pi^2\sum_{i=1}^j(1-\lambda_i),\qquad F_{j,\infty}=\|K\pi^2\tau^{(j)}\|_\infty=K\pi^2(1-\lambda_j), \end{equation} which are similar to the trace and $\lambda$-max test statistics in Joh:91, and satisfy $ F_{s,1}\stackrel{w}{\rightarrow}\operatorname{tr} (\omega)=\sum_{i=1}^s\zeta_i^{(s)}$, and $F_{s,\infty}\stackrel{w}{\rightarrow}\max\operatorname{eig} (\omega)=\zeta_1^{(s)} $ as $(T,K)_{seq}\rightarrow\infty$. The $F_{j,n}$ test rejects for large values of the statistics, and critical values $c_{n,\eta}$ are defined as $ \Pr(\|\zeta^{(s) }\|_{n} \leq c_{n,\eta})= 1- \eta $, see (ref); they can be estimated by simulation of the limit distribution. The following test sequence is defined using these quantities: choose an asymptotic test size $\eta \in(0,1)$ and test $H_{0j}$ versus $H_{1j}$ in (ref) for $j=p, p-1,\dots, 1$ using test statistic $F_{j,n}$ with significance level $\eta$ until a non-rejection is found. The resulting estimated value for $s$ equals the first non-rejected value of $j$; this test sequence is indicated as $\{F_{j,n}\}_{j=p,\dots,1}$.
corollary[Asymptotic properties of test sequences] Under the assumptions of Theorem $\ref{theorem_asy_distr_s}$, as $(T,K)_{seq}\rightarrow\infty$, the test sequence $\{F_{j,n}\}_{j=p,\dots,1}$, has asymptotic probability to select the true $s$ equal to $1-\eta$ for $s>0$ and equal to $1$ if $s=0$, $n=1,\infty$.

Results for test sequences are similar to the ones for LR tests in a VAR context for cointegration rank determination, see Joh:96.

Inference on the attractor space

This section discusses the limit distributions of the first stage and ICC estimators of the identified parameters $\psi$ and $\beta$ as well as Wald tests of hypotheses on their unrestricted parameters $\psi_*$ and $\beta_*$, see (ref), valid for any choice of basis of $L^2[0,1]$. Because $b'(\wh\psi-\psi)=0$ and $c'(\wh\beta-\beta)=0$, the unrestricted parameters are in $\bar c'(\wh\psi-\psi)=\wh\psi_*-\psi_*$ and $\bar b'(\wh\beta-\beta)=\wh\beta_*-\beta_*$, which are linearly related by the duality in Lemma $\ref{lem_ident}$ and (ref); see also Theorem $\ref{theorem_asy_distr_psi}$ below.

Results can be summarized as follows. First stage and ICC estimators are $T$-consistent and the ICC estimators are asymptotically mixed-Gaussian, unlike the first stage estimators. Thus, asymptotically Normal or $\chi^2$ Wald test statistics for hypotheses on $\psi$ and $\beta$ can be constructed using the ICC estimators.

theorem[Asymptotic distributions of the first stage and ICC estimators] Let $d_t$ in (ref) be constructed using any orthonormal c\`adl\`ag\ basis of $L^2[0,1]$ and let the unrestricted coefficients be as in (ref) and (ref); then for $(T,K)_{seq}\rightarrow\infty$, the first stage estimators satisfy \begin{equation} T(\wh\psi^{(1)}_*-\psi_*)=- T(\wh\beta^{(1)}_*-\beta_*)' \stackrel{w}{\rightarrow} \int \partial W_2 W_1'\left(\int W_1 W_1'\right)^{-1}=:Z^{(1)}, \end{equation} where $\int \partial W_2 W_1'$ denotes a Stratonovich stochastic integral and $Z^{(1)}$ is not mixed-Gaussian in general. On the contrary, for the ICC estimators one has \begin{equation} T(\wh\psi_*-\psi_*)=- T(\wh\beta_*-\beta_*)' \stackrel{w}{\rightarrow} \int \ensuremath{\mathrm{d}} W_{2.1} W_1'\left(\int W_1 W_1'\right)^{-1}=:Z, \end{equation} where $\int \ensuremath{\mathrm{d}} W_{2.1} W_1'$ denotes an Ito stochastic integral and $Z$ is mixed-Gaussian because $W_{2.1}$ and $W_1$ are independent.
remark[Mixed Gaussianity] Note that $W_2$ and $W_1$ in (ref) are dependent because $\Omega_{21}$ in (ref) is in general different from 0; this implies that $Z^{(1)}$ is not mixed-Gaussian in general. On the contrary, $W_{2.1}$ and $W_1$ in (ref) are independent and thus $Z$ is mixed-Gaussian; specifically, conditionally on $W_1$, one has $\operatorname{vec}(Z)\sim N(0,\left(\int W_1 W_1'\right)^{-1}\otimes\Omega_{22.1})$, which is a variance mixture of Normals. This shows that the dependence between $W_2 $ and $W_1 $ is cleansed by correcting $x_t$ for the fitted values of $\psi'\Delta x_t$ on $d_t$, similarly to the heteroskedasticity and autocorrelation-consistent and LRV estimators in Phi:05 and the IV estimator in Phi:14.

The limiting mixed-Gaussianity of the ICC estimators in Theorem (ref) can be exploited for testing hypothesis on $\psi_*$ and $\beta_*$ with Wald-type statistics, using consistent estimators of the LRV $\Omega_{22.1}$, see (ref). Among others, in line with Phi:05,Phi:14, the estimator $\wh\Omega_{22.1}:=\wh\Omega_{22}-\wh\Omega_{21}\wh\Omega_{11}^{-1}\wh\Omega_{12}$ could be used, with

equation[equation omitted — 357 chars of source]
theorem[Consistency of $\wh\Omega$] Let $d_t$ in (ref) be constructed using any orthonormal c\`adl\`ag\ basis of $L^2[0,1]$ and let $\wh\Omega$ be as in (ref); then $\wh\Omega\stackrel{p}{\rightarrow}\Omega$ as $(T,K)_{seq}\rightarrow\infty$, and as a consequence also $\wh\Omega_{22.1}\stackrel{p}{\rightarrow}\Omega_{22.1}$.

Consider the following hypothesis on the unrestricted coefficients of $\psi$,

equation[equation omitted — 69 chars of source]

where $R$ is a given $sr\times m$ matrix of full column rank and $h$ is a given $m \times 1$ vector.

remark[Hypotheses on $\psi_*$ and $\psi$] Recall that $\psi_* := \bar c' \psi$ so that (ref) corresponds to a linear hypotheses on $\psi$ of the form $R_0'\operatorname{vec}(\psi)=h$ with $R_0':=R' (I_s\otimes\bar c')$; $R_0$ is also of full column rank $m$ because $R$ and $c$ are of full column rank. Symmetrically, one can write $-h=R'\operatorname{vec}(\beta_*') = R_1'\operatorname{vec}(\beta')$, where $R_1':=R' (\bar b'\otimes I_r)$ is of full row rank $m$.
remark[Duality between hypothesis on $\psi$ and $\beta$] Note that by Lemma (ref).(iii) and (iv), or (ref), one has $R'\operatorname{vec}(\psi_*) - h = R'\operatorname{vec}(\beta_*') + h$. This shows that \begin{equation} H_0:R'\operatorname{vec}(\psi_*) = h\qquad\Leftrightarrow\qquad H_0:R'\operatorname{vec}(\beta_*') = -h. \end{equation}

The next theorem discusses the Wald test of (ref).

theorem[Wald tests based on the ICC estimators] Let the assumptions of Theorem $\ref{theorem_asy_distr_psi}$ and the $m$ linear restrictions on $\psi_*$ and $\beta_*$ in (ref) hold. Let also $\wh\Omega_{22.1}$ be a consistent estimator of $\Omega_{22.1}$; then as $(T,K)_{seq}\rightarrow\infty$ the Wald test statistic $Q$ satisfies \begin{equation} \begin{array}l Q:=T^2(R'\operatorname{vec}(\wh\psi_*)-h)'(R'\wh U R)^{-1}(R'\operatorname{vec}(\wh\psi_*)-h)\\ \quad\:=T^2 (R'\operatorname{vec}(\wh\beta_*') +h)' (R'\wh U R)^{-1} (R'\operatorname{vec}(\wh\beta_*') +h)\stackrel{w}{\rightarrow}\chi^2_m, \end{array} \end{equation} where $\wh U:=(T^{-1}\bar a'M_{xx}\bar a)^{-1}\otimes\wh\Omega_{22.1}\stackrel{w}{\rightarrow}\left(\int W_1W_1'\right)^{-1}\otimes\Omega_{22.1}$ with $a=\wh\psi$. Similarly when $m=1$, the $t$-ratio statistic satisfies \begin{equation} \frac{T(R'\operatorname{vec}(\wh\psi_*)-h)}{\sqrt{R'\wh U R}}=\frac{T(R'\operatorname{vec}(\wh\beta_*')+h)}{\sqrt{R'\wh U R}}\stackrel{w}{\rightarrow} N(0,1). \end{equation}

Efficiency

This section compares $Z$ in (ref) with the distribution in Joh:96 of the ML estimator for VAR processes and with the one in Phi:14 of the optimal IV estimator in the semiparametric case; it is shown that $Z$ coincides with both of them, thus proving that the ICC estimators are asymptotically efficient.

theorem[ML in Joh:96] Assume that the DGP is a VAR satisfying the $I(1)$ condition in Joh:96; then the DGP has representation (ref) with $C(z)$ satisfying Assumption $\ref{ass_DGP}$ and the asymptotic distribution $Z$ in (ref) coincides with the asymptotic distribution of the corresponding ML estimator of $-\beta_\ast'$ in Joh:96 and of $\psi_\ast$ in Par:97.

Consider next the semiparametric triangular form in Phi:14

equation[equation omitted — 95 chars of source]

where $u_t:=(u_{0t}',u_{1t}')'=Q(L)\varepsilon_t\sim I(0)$ with $Q(1)$ nonsingular, $y_t$ is $r \times 1$ and $z_t$ is $s \times 1$. This corresponds to a partition of $x_t$ into $(y_t',z_t')':=x_t$.

theorem[IV in Phi:14] Assume (ref) holds with $C(z)$ satisfying Assumption $\ref{ass_DGP}$; then $Q(z)$ satisfies the summability condition $(\textbf{L})$ in Phi:14 and the asymptotic distribution $Z$ in (ref) coincides with the asymptotic distribution of the corresponding IV estimator of $A:=-\beta_\ast'$ in Phi:14.

Monte Carlo simulations

This section reports a set of Monte Carlo (MC) simulations. The DGP is taken from OW:18 and BG:22 and data is generated from

equation[equation omitted — 176 chars of source]

with $\beta=(I_{p-s},0)'$ and $\alpha=-a\beta$. The DGP depends on the values of $(p,T,s,a)$ and identification is achieved with $b=(0,I_s)'$ and $c=(I_r,0)'$ except for $s=p$, where $\alpha$ and $\beta$ are undefined and $a$ plays no role.

When $s<p$, the first $p-s$ coordinate processes of $X_t$ in (ref) are univariate AR(1)'s with coefficient $1-a$ and the remaining $s$ are pure random walks. In terms of (ref), this implies that $\psi = (0,I_s)'$, $ \kappa' = (0,I_s)$, $C_1(L)=\operatorname{diag}(0,(1-(1-a)L)^{-1}I_r)$, so that the relevant LRV is $\Omega_{22.1}= a^{-2}I_r$. One expects small-sample inference on $s$ and $\psi$ to be more difficult when $a$ is smaller, because one needs to distinguish unit roots from autoregressive roots equal to $1-a$.

The properties of the max-gap and sequential estimators of $s$, see (ref) and (ref), are simulated with $K=\lceil T^{3/4} \rceil$ for the following values of $(p,T,s,a)$: $p=10,20,50,100,300$, $T=(10,20,30)p$, $s=\lceil jp / 4 \rceil$, $j=0,1,2,3,4$, $a=0.25,0.5,0.75,1$. Tables $\ref{table_max_gap}$, $\ref{table_mx}$ and $\ref{table_tr}$ in the supplement report full MC results; Figure $\ref{fig_freq_meadow}$ reports a representative subset for $p=20$.

figure[figure omitted — 608 chars of source]

The top row in Figure $\ref{fig_freq_meadow}$ reports results for the max-gap estimator. In absence of stationary AR components, i.e. when $s=p$ or $a=1$, the max-gap estimator selects the correct number of common trends $s$ with frequency 1 from $T=10p$ onward. When $a=0.75$ and $a=0.5$ the stationary AR components are increasingly autocorrelated, and the correct selection of $s$ with frequency 1 is observed for higher values of $T$. Finally, when $a=0.25$ and there are strong stationary AR components, the max-gap estimator never select the correct $s$ up to $T=30p$.

Overall, for fixed $(p,T,s)$ the performance of the max-gap worsens for lower values of $a$ because convergence to 0 of the smaller scc\ $\lambda_{j}$, $s<j\leq p$, appears to be slower for lower values of $a$, see Fig. (ref).

The middle row in Fig. $\ref{fig_freq_meadow}$ shows that when the max-gap estimator does not perform well, the estimator based on the test sequence $\{F_{j,n}\}_{j=p,\dots,1}$, $n=1$, has more reliable behavior (similar results hold for $n=\infty$, see Table (ref)). Table (ref) and (ref) in the Supplement further document that the tests $F_{s,1}$ and $F_{s,\infty}$ are undersized in finite samples.

One can combine inference based on the max-gap estimator and on the test $F_{p,n}$, defining a hybrid estimator $\wh s _{n}$ that selects $p$ when $F_{p,n}$ does not reject, and uses the restricted max-gap estimator over $\{0,\dots,p-1\}$. The bottom row of Fig. $\ref{fig_freq_meadow}$ shows that the hybrid estimator with $n=1$ performs well (similar results hold for $n=\infty$, see Tables (ref) and (ref)).

Next consider the Wald $Q$ and $t$-ratio statistics for linear hypotheses on $\psi$, see Theorem $\ref{theorem_inf_psi}$. Two hypotheses of type (ref) are considered, namely (i) the element 1,1 of $\psi_*$ equals 0, tested with the $t$-statistic in (ref), and (ii) the first column of $\psi_*$ equals 0, tested using the Wald $Q$ statistic in (ref) with a $\chi^2_r$ weak limit; both hypotheses are true under the DGP.

The main challenge in the implementation of the tests lies in the precise estimation of the LRV $\Omega_{22.1}$, which is notoriously difficult by nonparametric means, see e.g. Haug:02. Three different estimators of $\Omega_{22.1}$ are considered in the computations of the test statistics: the first one is the true value in the DGP for $\Omega_{22.1}=a^{-2}I_r$, called LRV-U (unfeasible); the second option employs $\wh\Omega_{22.1}$ in (ref) following Phi:14, called LRV-P, and the third one (LRV-A) computes $\wh\Omega_{22.1}$ as in An:91 and AM:92 using Parzen's kernel.

The properties of the tests are simulated with $K=\lceil T^{3/4} \rceil$ for the following values of $(p,T,s,a)$: $p=10,20$, $T=(30,60,90)p$, $s=\lceil jp /4 \rceil$, $j=1,2,3$, $a=0.25,0.5,0.75,1$. Table $\ref{table_W1_Wr}$ in the supplement reports full Monte Carlo results, and Figure $\ref{fig_distr_p_20_T_600_s_5}$ reports the subset for $T=30p$ and varying $a$. It can be seen that the tests are largely oversized in finite samples, in line with the literature Haug:02. The unfeasible $t$-ratio LRV-U performs considerably better, while the behavior of LRV-P and LRV-A is comparable, thus suggesting that a major part of the size distortions are due to the imprecise LRV estimation. Finally, the $N(0,1)$ approximation improves with increasing $a$ (i.e. with decreasing level of autocorrelation of the stationary AR components.)

figure[figure omitted — 527 chars of source]

Empirical application

This section provides an illustration on a panel of daily exchange rates between Jan 4, 2022 and Aug 30, 2024 of the US dollar against 20 World Markets (WM) currencies, see OW:19 for a similar dataset and a review of the literature on the topic. The sample size is $T=667$ and $K=\lceil T^{3/4}\rceil = 132$. The illustration of the present methods is compared with the high dimensional VAR analysis of OW:18 and BG:24 for inference on $s$, and with the likelihood analysis of Joh:96 for inference on $s$ and $\psi$. Details are reported in the supplement.

Data (in logs and normalized to start at 0) are plotted in the first four panels of Figure $\ref{fig_data}$. Countries are grouped as follows: Emerging Markets (EM: Brazil, China, India, Malaysia, Mexico, South Africa, South Korea, Taiwan, Thailand) and Developed Markets (DM). DM are broken down into non European (Non-EU: Australia, Canada, Hong Kong, Japan, Singapore) and European (EU). EU is partitioned into UK and SZ (United Kindom and Switzerland), and Nordic and Eurozone (Nc and EZ: Denmark, Eurozone, Norway, Sweden).

figure[figure omitted — 347 chars of source]

The cca\ analysis applied to all the $p=20$ WM time series is summarized in the first two subgraphs in Figure (ref), which report the associated scc\ profile and the misspecification stripe of Remark $\ref{rem_stripe}$ at 95% confidence level. The largest drop in the scc\ profile is between eigenvalue 19 and 20 (red vertical line), so that the max-gap estimator selects $\wh s = 19$, indicating $\wh r = 1$ cointegrating relation (CI).

The result on the presence of cointegration among the 20 WM currencies is further investigated using Wachter plots as in OW:18 and BG:24, and the test of no cointegration ($s=p$) of BG:24 in a VAR$(k)$, $k=1,2,3,4$. For all values of $k$, Figure $\ref{fig_BG}$ in the supplement displays eigenvalues to the right of the support of the Wachter distribution, which is an indication of the presence of cointegration ($s<20$). This is further confirmed by the results of their test at 0.01 significance level, see Table $\ref{table_BG}$ in the supplement.

Thanks to the dimensional coherence of the present approach, the present analysis can be performed on subgroups separately, as indicated in Figure $\ref{fig_tree}$.

figure[figure omitted — 843 chars of source]

The max-gap estimator finds no CI relations in EM and 1 CI relation in DM, see subgraphs 3-4 and 5-6 in Figure (ref). In the next split into EU and Non-EU, see Figure $\ref{fig_tree}$, the max-gap estimator indicates no CI in Non-EU and 1 CI relation in Euro, see subgraphs subgraphs 7-8 and 9-10 in Figure (ref). Finally the max-gap estimator indicates no CI in the UK and SZ group and 1 CI relation in the Nc and EZ group, see subgraphs 11-12 and 13-14 in Figure (ref).

figure[figure omitted — 346 chars of source]

Next the Nc and EZ subgroup is used to estimate $\psi$ and $\beta$, with identification restrictions $c=e_2$ and $b=(e_1,e_3,e_4)$, where $e_i$ indicates the $i$-th unit vector; this identification restriction is tested using the criterion in (ref) and it is not rejected. The resulting $\wh\psi$ and $\wh \beta$ are\footnote{Danish Krone (DK), Eurozone (Euro), Norwegian Krone (NK), Swedish Krone (SK).}

equation[equation omitted — 516 chars of source]

with $p$-values reported in parenthesis for the unrestricted coefficients $\psi_*=e_2'\wh\psi=-\beta_*'$, which correspond to the second row of $\wh\psi$. The hypothesis $H_0:\psi_{2,1}=0$ and $H_0:\psi_{2,2}=0$ are strongly rejected while $H_0:\psi_{2,3}=0$ is not, i.e. the Euro loads the first and second stochastic trends but not the third trend, which only affects the Swedish Krone.

For the Nc and EZ subgroup, Table (ref) in the supplement reports the likelihood ratio test on $s$ and the QML estimate of $\beta$ in Joh:96. The estimate of $s$ is 3 as in (ref), and the estimates of $\beta$ are essentially indistinguishable.

The time series of the Nc and EZ common stochastic trends $\wh \psi'X_t$ and cointegrating relation $\wh \beta'X_t$ are reported in Figure (ref); the plot of $\wh \beta'X_t$ includes some large infrequent shocks, whose analysis, however, goes beyond the scope of the present illustration.

Conclusions

This paper proposes a unified semiparametric framework for inference on the number and the loadings of stochastic trends (and on their duals, the number and the coefficients of the cointegrating relations) for $I(1)/I(0)$ processes, which includes VARIMA and Dynamic Factor Model processes. The setup directly applies also to subsets of variables or to their aggregation, i.e. it is dimensionally coherent.

The approach employs a novel canonical correlation analysis between the observed multiple time series and a finite subset of an $L^2[0,1]$ basis. Canonical correlations deliver test sequences for the number of stochastic trends, as well as consistent estimators of it. Canonical variates provide estimators of the loading matrix of stochastic trends and of their duals, the matrix of cointegrating coefficients. The ICC estimators are $T$-consistent, mixed-Gaussian and efficient. Wald tests on the parameters are developed, as well as misspecification tests for checking model assumptions.

The finite sample properties of the estimators and tests investigated via a Monte Carlo simulation study display reasonable performance. An empirical analysis illustrates the range of available inferences in the proposed framework.

thebibliography\bibitem[\citeauthoryear{Ahn and Horenstein}{Ahn and Horenstein}{2013}]{AH:13} Ahn, S. and A. Horenstein (2013). \newblock Eigenvalue {R}atio {T}est for the {N}umber of {F}actors. \newblock {\em Econometrica\/} {\em 81}, 1203--1227. \bibitem[\citeauthoryear{Ahn and Reinsel}{Ahn and Reinsel}{1990}]{AR:90} Ahn, S. and C. Reinsel (1990). \newblock Estimation for partially nonstationary multivariate autoregressive models. \newblock {\em Journal of the American Statistical Association\/} {\em 85}, 813--23. \bibitem[\citeauthoryear{Andrews}{Andrews}{1991}]{An:91} Andrews, D. W. K. (1991). \newblock Heteroskedasticity and autocorrelation consistent covariance matrix estimation. \newblock {\em Econometrica\/} {\em 59\/}(3), 817--858. \bibitem[\citeauthoryear{Andrews and Monahan}{Andrews and Monahan}{1992}]{AM:92} Andrews, D. W. K. and J. C. Monahan (1992). \newblock An improved heteroskedasticity and autocorrelation consistent covariance matrix estimator. \newblock {\em Econometrica\/} {\em 60\/}(4), 953--966. \bibitem[\citeauthoryear{Bai}{Bai}{2004}]{Bai:04} Bai, J. (2004). \newblock Estimating cross-section common stochastic trends in nonstationary panel data. \newblock {\em Journal of Econometrics\/} {\em 122}, 137--183. \bibitem[\citeauthoryear{Barigozzi, Cavaliere, and Trapani}{Barigozzi et al.}{2022}]{BCT:22} Barigozzi, M., G. Cavaliere, and L. Trapani (2022). \newblock Inference in {H}eavy-{T}ailed {N}onstationary {M}ultivariate {T}ime {S}eries. \newblock {\em Journal of the American Statistical Association\/}, 565--581. \bibitem[\citeauthoryear{Bierens}{Bierens}{1997}]{Bie:97} Bierens, H. (1997). \newblock Nonparametric cointegration analysis. \newblock {\em Journal of Econometrics\/} {\em 77}, 379--404. \bibitem[\citeauthoryear{Bykhovskaya and Gorin}{Bykhovskaya and Gorin}{2022}]{BG:22} Bykhovskaya, A. and V. Gorin (2022). \newblock Cointegration in large {VAR}s. \newblock {\em The Annals of Statistics\/} {\em 50}, 1593 -- 1617. \bibitem[\citeauthoryear{Bykhovskaya and Gorin}{Bykhovskaya and Gorin}{2024}]{BG:24} Bykhovskaya, A. and V. Gorin (2024). \newblock Asymptotics of cointegration tests for high-dimensional {VAR}(k). \newblock {\em The Review of Economics and Statistics\/}, 1--38. \bibitem[\citeauthoryear{Cavaliere, Nielsen, and Rahbek}{Cavaliere et al.}{2015}]{CNR:15} Cavaliere, G., H. B. Nielsen, and A. Rahbek (2015). \newblock Bootstrap testing of hypotheses on co-integration relations in vector autoregressive models. \newblock {\em Econometrica\/} {\em 83}, 813--831. \bibitem[\citeauthoryear{Cavaliere, Rahbek, and Taylor}{Cavaliere et al.}{2012}]{CRT:12} Cavaliere, G., A. Rahbek, and R. Taylor (2012). \newblock Bootstrap determination of the co-integration rank in vector autoregressive models. \newblock {\em Econometrica\/} {\em 80}, 1721--1740. \bibitem[\citeauthoryear{Chen and Schienle}{Chen and Schienle}{2024}]{CS:24} Chen, S. and M. Schienle (2024). \newblock Large {S}pillover {N}etworks of {N}onstationary {S}ystems. \newblock {\em Journal of Business & Economic Statistics\/} {\em 42}, 422--436. \bibitem[\citeauthoryear{Elliott}{Elliott}{1999}]{Ell:99} Elliott, G. (1999). \newblock Efficient {T}ests for a {U}nit {R}oot {W}hen the {I}nitial {O}bservation is {D}rawn from {I}ts {U}nconditional {D}istribution. \newblock {\em International {E}conomic {R}eview\/} {\em 40}, 767--783. \bibitem[\citeauthoryear{Engle and Granger}{Engle and Granger}{1987}]{EG:87} Engle, R. and C. Granger (1987). \newblock Co-integration and {E}rror {C}orrection: {R}epresentation, {E}stimation, and {T}esting. \newblock {\em Econometrica\/} {\em 55}, 251--276. \bibitem[\citeauthoryear{Franchi, Georgiev, and Paruolo}{Franchi et al.}{2023}]{FGP:23} Franchi, M., I. Georgiev, and P. Paruolo (2023). \newblock Estimating the number of common trends in large {$T$} and {$N$} factor models via canonical correlations analysis. \newblock {\em Econometrics and Statistics\/}, forthcoming. \bibitem[\citeauthoryear{Haug}{Haug}{2002}]{Haug:02} Haug, A. A. (2002). \newblock Testing linear restrictions on cointegrating vectors: sizes and powers of {W}ald and likelihood ratio tests in finite samples. \newblock {\em Econometric Theory\/} {\em 18\/}(2), 505--524. \bibitem[\citeauthoryear{Jacod and Shiryaev}{Jacod and Shiryaev}{2003}]{JS:03} Jacod, J. and A. Shiryaev (2003). \newblock {\em Limit {T}heorems for {S}tochastic {P}rocesses}. \newblock Springer. \bibitem[\citeauthoryear{Johansen}{Johansen}{1988}]{Joh:88} Johansen, S. (1988). \newblock Statistical {A}nalysis of {C}ointegration {V}ectors. \newblock {\em Journal of Economic Dynamics and Control\/} {\em 12}, 231--254. \bibitem[\citeauthoryear{Johansen}{Johansen}{1991}]{Joh:91} Johansen, S. (1991). \newblock Estimation and {H}ypothesis {T}esting of {C}ointegration {V}ectors in {G}aussian {V}ector {A}utoregressive {M}odels. \newblock {\em Econometrica\/} {\em 59}, 1551--1580. \bibitem[\citeauthoryear{Johansen}{Johansen}{1996}]{Joh:96} Johansen, S. (1996). \newblock {\em Likelihood-based {I}nference in {C}ointegrated {V}ector {A}uto-{R}egressive {M}odels}. \newblock Oxford University Press. \bibitem[\citeauthoryear{Johansen}{Johansen}{2000}]{Joh:00} Johansen, S. (2000). \newblock A {B}artlett {C}orrection {F}actor for {T}ests on the {C}ointegrating {R}elations. \newblock {\em Econometric Theory\/} {\em 16}, 740--778. \bibitem[\citeauthoryear{Johansen}{Johansen}{2002a}]{Joh:02b} Johansen, S. (2002a). \newblock A small sample correction for tests of hypotheses on the cointegrating vectors. \newblock {\em Journal of Econometrics\/} {\em 111}, 195--221. \bibitem[\citeauthoryear{Johansen}{Johansen}{2002b}]{Joh:02} Johansen, S. (2002b). \newblock A {S}mall {S}ample {C}orrection for the {T}est of {C}ointegrating {R}ank in the {V}ector {A}utoregressive {M}odel. \newblock {\em Econometrica\/} {\em 70}, 1929--1961. \bibitem[\citeauthoryear{Liang and Schienle}{Liang and Schienle}{2019}]{LS:19} Liang, C. and M. Schienle (2019). \newblock Determination of vector error correction models in high dimensions. \newblock {\em Journal of Econometrics\/} {\em 208}, 418--441. \bibitem[\citeauthoryear{Loeve}{Loeve}{1978}]{LoII:78} Loeve, M. (1978). \newblock {\em Probability theory II\/} (4th ed.). \newblock Graduate Texts in Mathematics. Springer. \bibitem[\citeauthoryear{M\"{u}ller and Watson}{M\"{u}ller and Watson}{2008}]{MW:08} M\"{u}ller, U. and M. Watson (2008). \newblock Testing {M}odels of {L}ow-{F}requency {V}ariability. \newblock {\em Econometrica\/} {\em 76}, 979--1016. \bibitem[\citeauthoryear{M\"{u}ller and Watson}{M\"{u}ller and Watson}{2017}]{MW:17} M\"{u}ller, U. and M. Watson (2017). \newblock Low-{F}requency {E}conometrics. \newblock In B. Honor\'e, A. Pakes, M. Piazzesi, and L. E. Samuelson (Eds.), {\em Advances in Economics and Econometrics: Eleventh World Congress}, Econometric Society Monographs, pp.\ 53--94. Cambridge University Press. \bibitem[\citeauthoryear{M\"{u}ller and Watson}{M\"{u}ller and Watson}{2018}]{MW:18} M\"{u}ller, U. and M. Watson (2018). \newblock Long-{R}un {C}ovariability. \newblock {\em Econometrica\/} {\em 86}, 775--804. \bibitem[\citeauthoryear{Nielsen, Seo, and Seong}{Nielsen et al.}{2022}]{NSS:22} Nielsen, M., W. Seo, and D. Seong (2022). \newblock Inference on the dimension of the nonstationary subspace in functional time series. \newblock {\em Econometric Theory\/} {\em 39}, 443--480. \bibitem[\citeauthoryear{Onatski and Wang}{Onatski and Wang}{2018}]{OW:18} Onatski, A. and C. Wang (2018). \newblock Alternative {A}symptotics for {C}ointegration {T}ests in {L}arge {VAR}s. \newblock {\em Econometrica\/} {\em 86}, 1465--1478. \bibitem[\citeauthoryear{Onatski and Wang}{Onatski and Wang}{2019}]{OW:19} Onatski, A. and C. Wang (2019). \newblock Extreme canonical correlations and high-dimensional cointegration analysis. \newblock {\em Journal of Econometrics\/} {\em 212}, 307--322. \bibitem[\citeauthoryear{Paruolo}{Paruolo}{1997}]{Par:97} Paruolo, P. (1997). \newblock Asymptotic {I}nference on the {M}oving {A}verage {I}mpact {M}atrix in {C}ointegrated {I(1) V}ar {S}ystems. \newblock {\em Econometric Theory\/} {\em 13}, 79--118. \bibitem[\citeauthoryear{Petersen, Zhang, and Kokoszka}{Petersen et al.}{2022}]{PZK22} Petersen, A., C. Zhang, and P. Kokoszka (2022). \newblock Modeling probability density functions as data objects. \newblock {\em Econometrics and Statistics\/} {\em 21}, 159--178. \bibitem[\citeauthoryear{Phillips and Hansen}{Phillips and Hansen}{1990}]{Phillips1990} Phillips, P. and B. Hansen (1990). \newblock Statistical inference in {I}nstrumental {V}ariables regression with {I(1)} processes. \newblock {\em Review of Economic Studies\/} {\em 57\/}(1), 99--125. \bibitem[\citeauthoryear{Phillips}{Phillips}{1998}]{Phi:98} Phillips, P. C. B. (1998). \newblock New {T}ools for {U}nderstanding {S}purious {R}egressions. \newblock {\em Econometrica\/} {\em 66}, 1299--1326. \bibitem[\citeauthoryear{Phillips}{Phillips}{2005}]{Phi:05} Phillips, P. C. B. (2005). \newblock {HAC} {E}stimation by {A}utomated {R}egression. \newblock {\em Econometric Theory\/} {\em 21}, 116--142. \bibitem[\citeauthoryear{Phillips}{Phillips}{2014}]{Phi:14} Phillips, P. C. B. (2014). \newblock Optimal estimation of cointegrated systems with irrelevant instruments. \newblock {\em Journal of Econometrics\/} {\em 178}, 210--224. \bibitem[\citeauthoryear{Phillips and Solo}{Phillips and Solo}{1992}]{PS:92} Phillips, P. C. B. and V. Solo (1992). \newblock {Asymptotics for Linear Processes}. \newblock {\em The Annals of Statistics\/} {\em 20}, 971 -- 1001. \bibitem[\citeauthoryear{Stock and Watson}{Stock and Watson}{1988}]{SW:88} Stock, J. and M. Watson (1988). \newblock Testing for common trends. \newblock {\em Journal of the American Statistical Society\/} {\em 83}, 1097--1107. \bibitem[\citeauthoryear{Swensen}{Swensen}{2006}]{Swe:06} Swensen, A. (2006). \newblock Bootstrap {A}lgorithms for {T}esting and {D}etermining the {C}ointegration {R}ank in {VAR M}odels. \newblock {\em Econometrica\/} {\em 74}, 1699--1714. \bibitem[\citeauthoryear{Watson}{Watson}{1994}]{Wat:94} Watson, M. W. (1994). \newblock Chapter 47: {V}ector autoregressions and cointegration. \newblock In R. F. Engle and D. L. McFadden (Eds.), {\em Handbook of Econometrics}, Volume 4, pp.\ 2843--2915. Elsevier. \bibitem[\citeauthoryear{Zhang, Robinson, and Yao}{Zhang et al.}{2019}]{ZRY:19} Zhang, R., P. Robinson, and Q. Yao (2019). \newblock Identifying cointegration by eigenanalysis. \newblock {\em Journal of the American Statistical Association\/} {\em 114}, 916--927.