EconBase
← Back to paper

Functional Principal Component Analysis for Cointegrated Functional Time Series

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

47,502 characters · 10 sections · 59 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Functional principal component analysis for cointegrated functional time series

abstractFunctional principal component analysis (FPCA) has played an important role in the development of functional time series analysis. This note investigates how FPCA can be used to analyze cointegrated functional time series and proposes a modification of FPCA as a novel statistical tool. Our modified FPCA not only provides an asymptotically more efficient estimator of the cointegrating vectors, but also leads to novel FPCA-based tests for examining essential properties of cointegrated functional time series. \\[2pt] MSC 2020: 62M10. \\ Keywords: Functional principal component analysis; functional time series; unit roots; cointegration.

Introduction

Functional principal component analysis (FPCA) has been a central tool for functional time series (FTS) in various contexts encompassing FTS regression Park2012397,seong2022, prediction hyndman2007robust,aue2015prediction and long memory FTS li2020long to name a few. In the recent work by Chang2016152, FPCA is applied to nonstationary cointegrated FTS. Given that many economic time series are nonstationary but allow a long run stable relationship, their analysis paves the way for potential applications and extensions (see e.g., chang2020evaluating, LRS, seoshang2022).

This note further investigates how FPCA can be used in statistical analysis of cointegrated FTS $\{X_t\}_{t\geq 1}$ taking values in a Hilbert space $\mathcal H$. Cointegrated FTS considered in this paper can be decomposed by orthogonal projections $P^N$ and $P^S = I-P^N$ (where $I$ denotes the identity map) into a finite dimensional unit root process $\{P^NX_t\}_{t \geq 1}$ and a potentially infinite dimensional stationary process $\{P^SX_t\}_{t \geq 1}$. When such a time series is given, it is important to estimate $P^N$ or $P^S$ since either of them characterizes the cointegrating behavior. Chang2016152 earlier studied how the eigenelements of the sample covariance operator can be used to construct a consistent estimator of $P^N$. In this note, we first add some novel asymptotic results, including the convergence rate and the asymptotic limit of the existing estimator of $P^N$ (or $P^S$) obtained from FPCA. Based on the results, we propose a modified FPCA methodology that guarantees (i) a more asymptotically efficient estimator and (ii) a more convenient asymptotic limit which leads to novel statistical tests for examining hypotheses about cointegration. \textcolor{black}{Our modified FPCA is motivated from semiparametric modifications to obtain asymptotically efficient estimators of the cointegrating vectors in a multivariate cointegrated system as in phillips1990statistical and harris1997principal, where the latter article is more related to the present paper (this will be discussed in detail in Section (ref) of the Supplementary Material to this article}).

The rest of this paper is organized as follows. Section (ref) provides main assumptions. In Section (ref), we discuss on the ordinary FPCA and our proposed modification of it as statistical tools to estimate $P^N$ or $P^S$, and provide the relevant asymptotic theory. Based on our modified FPCA, Section (ref) gives statistical tests to examine various hypotheses on cointegration. In Section (ref), we conclude with some cautions about using our methodology in practice and present some idea how the methodology can be further extended. The Supplementary Material to this article contains six sections (Sections (ref)-(ref)) on some extensions of the theoretical results to be given, Monte Carlo simulations, empirical applications and mathematical proofs.

We review notation for the subsequent sections. Let $\mathcal H$ denote a real separable Hilbert space with inner product $\langle \cdot, \cdot \rangle$ and norm $\|\cdot\| = \langle \cdot,\cdot \rangle^{1/2}$. The space of bounded linear operator is denoted by $\mathcal{L}_{\mathcal H}$, which is assumed to be equipped with the operator norm $\|\cdot\|_{\mathcal L_{\mathcal H}}$. For any $A\in \mathcal L_{\mathcal H}$, its adjoint ($A^\ast$), range ($\operatorname{ran} A$), kernel ($\ker A$), rank ($\operatorname{rank} A$), and Moore-Penrose inverse ($A^\dag$) are introduced in Section (ref). Moreover, various properties of $A \in \mathcal L_{\mathcal H}$ such as compactness, positive (semi)definiteness and self-adjointness are also defined in detail. As discussed in Section (ref), if $A$ is self-adjoint, positive semidefinite, and compact, we may define its $m$-regularized inverse $A{\mid_{m}^\dagger}= \sum_{j=1}^m a_j^{-1} u_j \otimes u_j$ where $\{a_j\}_{j=1}^m$ are positive eigenvalues of $A$ and $\{u_j\}_{j=1}^m$ are the corresponding eigenvectors; $A{\mid_{m}^\dagger}$ is understood as the partial inverse of $A$ on the restricted domain $\operatorname{span}(\{u_j\}_{j=1}^m)$ {(for convenience, we let $A{\mid_{m}^\dagger} = 0$ if $A=0$)}. Random elements taking values in $\mathcal H$ and $\mathcal L_{\mathcal H}$ are introduced in Section (ref). We let $L^2_{\mathcal H}$ denote the space of $\mathcal H$-valued random variables $X$ satisfying $\mathbb{E}X=0$ and $\mathbb{E}\|X\|^2 < \infty$. For any $X \in L^2_{\mathcal H}$, its covariance operator is defined by $C_X = \mathbb{E} \left[ X \otimes X \right]$. If $A$ is a random linear operator taking values in $\mathcal L_{\mathcal H}$ and $x_1,\ldots,x_n,y_1,\ldots,y_n \in \mathcal H$ for $n \in \mathbb{N}$, we let the distribution of $(\langle Ax_1,y_1 \rangle, \ldots, \langle Ax_n,y_n\rangle)'$ be called the finite dimensional distribution (fdd) of $A$ with respect to the choice of vectors; if $n=m^2$ for $m \in \mathbb{N}$, $x_j=e_j$ and $y_k=e_k$ for all $j,k\leq m$, and $\{e_j\}_{j=1}^m$ is an orthonormal basis of a subspace $\mathcal H_{m}$, then the associated fdd can be viewed as the distribution of $A$ as a map from $\mathcal H_{m}$ to $\mathcal H_{m}$ If two random bounded linear operators $A$ and $B$ have the same fdd regardless of $n$ and $x_1,\ldots,x_n,y_1,\ldots,y_n \in \mathcal H$, we write $A=_{\operatorname{fdd}}B$. Moreover, if $\{A_j\}_{j\geq 1}$ is a sequence of random elements in $\mathcal L_{\mathcal H}$ and $\|A_j-A\|_{\mathcal L_{\mathcal H}}\to_p 0$ for some $A$ taking its value in $\mathcal L_{\mathcal H}$, we write $A_j \to_{\mathcal L_{\mathcal H}} A$.

Cointegrated FTS in Hilbert space

We let $\{\varepsilon_t\}_{t \in \mathbb{Z}}$ be an iid sequence in $L^2_{\mathcal H}$ with a positive definite covariance and let $\{\Phi_j\}_{j \geq 0}$ be a sequence in $\mathcal L_{\mathcal H}$ satisfying $\sum_{j=0}^\infty j \|\Phi_j\|_{\mathcal L_{\mathcal H}} < \infty$. We define $\{\zeta_t\}_{t \in \mathbb{Z}}$ and $\{\eta_t\}_{t \in \mathbb{Z}}$ as follows,

equation*[equation* omitted — 173 chars of source]

where $\widetilde{\Phi}_j = - \sum_{k\geq j+1} \Phi_k$. We consider a sequence $\{X_t\}_{t\geq 0}$ of which first differences $\Delta X_t =X_t - X_{t-1}$ satisfy the equations $\Delta X_t = \zeta_t$ for $t \geq 1$. {Then by the Phillips-Solo decomposition (and ignoring the initial values $X_0$ and $\eta_0$ that are unimportant in the development of our asymptotic theory), we find that }

equation[equation omitted — 95 chars of source]

\textcolor{black}{where $\varepsilon_t^c = \sum_{s=1}^t \varepsilon_s$ and $\Phi(1) = \sum_{j\geq 0} \Phi_j$.} Note that $\{X_t\}_{t \geq 1}$ is a unit-root-type nonstationary process unless $\Phi(1) = 0$; however it may have an element $x$ such that $\{\langle X_t,x \rangle\}_{t \geq 1}$ is stationary. Such an element is called a cointegrating vector, and the collection of the cointegrating vectors, denoted by $\mathcal H^S$, is called the cointegrating space. The attractor space, denoted by $\mathcal H^N$, is defined by the orthogonal complement to $\mathcal H^S$. When $X_t$ satisfies \hyperref[{eqcointeg2}]{\tagform@{\ref*{eqcointeg2}}}, $\mathcal H^S$ (resp.\ $\mathcal H^N$) is given by $ [\operatorname{ran} \Phi(1)]^\perp$ (resp.\ the closure of $\operatorname{ran} \Phi(1)$), in this case, $\mathcal H = \mathcal H^N \oplus \mathcal H^S$ and the orthogonal projection $P^N$ onto $\mathcal H^N$ is uniquely defined (see Proposition 3.3 of BSS2017). Then $P^S=I-P^N$ is the orthogonal projection onto $\mathcal H^S$. Note that $\{P^NX_t\}_{t\geq 1}$ is a unit root process which does not allow any cointegrating vector while $\{P^SX_t\}_{t\geq 1}$ is stationary.

$\{X_t\}_{t\geq 1}$ in \hyperref[{eqcointeg2}]{\tagform@{\ref*{eqcointeg2}}} has zero mean by construction, but which is not essentially required in this paper. A deterministic component such as a nonzero intercept or a linear trend may be allowed in our analysis with a proper modification, which will be discussed in Section (ref) of the Supplementary Material.

For our asymptotic analysis, we apply the following additional conditions throughout.

assumpMM$\mathrm{(i)}$ $\sum_{j=0}^\infty j\|\widetilde{\Phi}_j\|_{\mathcal L_{\mathcal H}} < \infty$ and $\mathrm{(ii)}$ $\varphi = \operatorname{rank} \Phi(1) \in [0, \infty)$.

Assumption (ref)-(i) is employed for mathematical proofs, which may not be restrictive in practice. Assumption (ref)-(ii) implies that $\dim(\mathcal H^N) <\infty$, which is commonly assumed in the recent literature on cointegrated FTS for statistical analysis, see e.g., Chang2016152 and SS2019.

remark\normalfont Finiteness of $\varphi$ seems to be reasonable in many empirical examples; see Chang2016152 and SS2019. Moreover, it is known that functional autoregressive (AR) processes with unit roots (including functional ARMA processes that are considered in e.g.\ klepsch2017prediction, klepsch2017prediction) with compact AR operators always satisfy Assumption (ref)-(ii); see BS2018, Franchi2017b and seo_2022. It is common to assume compactness of AR operators in statistical analysis of such time series, and thus Assumption (ref)-(ii) does not seem to be restrictive in practice.

The case with $\mathcal H^S =\{0\}$ for each $\varphi$ is uninteresting for our study of cointegrated FTS, and thus such a case is not considered throughout this paper. Moreover, if $\varphi = 0$, $\{X_t\}_{t\geq 1}$ is stationary and thus $\mathcal H^N = \{0\}$; this is also an uninteresting case except when we examine the null hypothesis of $\varphi =0$ in Section (ref). Thus, our asymptotic theory, to be developed in this section, concerns the case when $\varphi \geq 1$.

Related to a sequence $\{X_t\}_{t \geq 1}$ satisfying Assumption (ref), it will be convenient to introduce additional notation. Note that the identity $I=P^N + P^S$ implies the following operator decomposition: for $A \in \mathcal L_{\mathcal H}$,

equation[equation omitted — 142 chars of source]

We also define

equation[equation omitted — 147 chars of source]

note that each sequence of $\mathcal E_t^N$, $\mathcal E_t^S$ and $\mathcal E_t$ is stationary in $\mathcal H$. We then let $\Omega$ (resp.\ $\Gamma$) denote the long run covariance (resp.\ one-sided long run covariance) operator of $\{\mathcal E_t\}_{t \in \mathbb{Z}}$, which are defined by

align[align omitted — 219 chars of source]

$\Omega$ and $\Gamma$ are well defined under Assumption (ref). As in \hyperref[{eqdecom2}]{\tagform@{\ref*{eqdecom2}}}, we may also decompose those operators as $\Omega= \Omega^{NN}+\Omega^{NS}+ \Omega^{SN}+ \Omega^{SS}$ and $\Gamma = \Gamma^{NN}+\Gamma^{NS}+\Gamma^{SN}+\Gamma^{SS}$. We let $W = \{W(r)\}_{r \in [0,1]}$ denote Brownian motion in $\mathcal H$ whose covariance operator is given by $\Omega$ (see e.g., chen1998). We then let $W^N = P^N W$ and $W^S = P^S W$. Note that $W^N$ and $W^S$ are correlated unless $\Omega^{NS} = \Omega^{SN} = 0$.

For convenience in our asymptotic analysis, we employ the following assumption: in the assumption below, we let the long-run covariance $\Omega$ of $\mathcal E_t$ be positive definite on $\mathcal H$ and, for $x_1,\ldots,x_k \in \mathcal H$, $\mathcal E_{k,t} = (\langle \mathcal E_{t}, x_1 \rangle, \ldots, \langle \mathcal E_{t}, x_k \rangle)'$ and $W_k(s) = (\langle W(s), x_1 \rangle, \ldots, \langle W(s), x_k \rangle)'$.

assumpWWFor any $k \geq 1$ and $x_1,\ldots,x_k \in \mathcal H$, $T^{-1} \sum_{t=1}^T \left(\sum_{s=1}^t \mathcal E_{k,s}\right) \mathcal E_{k,t}'$ converges in distribution to $\int_{0}^1 W_k(s) d W_k(s)' + \sum_{j\geq 0} \mathbb{E}[\mathcal E_{k,t-j}\mathcal E_{k,t}']$.

Appropriate lower-level conditions for the above can be found in hansen1992convergence; specifically, Assumption (ref) is satisfied if (i) $\{T^{-1/2} \sum_{t=1}^{\lfloor Ts \rfloor} \mathcal E_{k,t}\}_{s\in [0,1]}$ converges weakly in the Skorohod topology to $W_k$, (ii) $\sup_{t\geq 1} \|\mathcal E_t\|^a < \infty$, and (iii) $\{\mathcal E_{t}\}_{t\geq 1}$ is strongly mixing of size $-ab(a-b)$ for some $a>b>2$. Among these conditions, the first is implied by Assumption (ref) (see Lemma (ref) in the Supplementary Material).

FPCA of cointegrated FTS and asymptotic results

Ordinary FPCA of cointegrated FTS

We now consider estimation of $P^N$ and $P^S$ from observations $\{X_t\}_{t=1}^T$. Throughout this section, we assume that $\varphi=\dim(\mathcal H^N)$ is known. Of course, this is not a realistic assumption in most applications. However, we already have a few available methods to determine $\varphi$ such as those in Chang2016152, SS2019, and LRS. In addition to those, it will be shown in Section (ref) that the asymptotic results to be given in this section lead to a novel FPCA-based testing procedure to determine $\varphi$.

The sample covariance operator $\widehat{C}$ of $\{X_t\}_{t=1}^T$ is the random operator given by $\widehat{C} = T^{-1}\sum_{t=1}^T X_t \otimes X_t$. Broadly speaking, FPCA reduces to solving the following eigenvalue problem associated with $\widehat{C}$,

equation[equation omitted — 110 chars of source]

where $\hat{\lambda}_{1} \geq \ldots \geq \hat{\lambda}_{T} \geq 0 = \hat{\lambda}_{T+1}= \hat{\lambda}_{T+2}=\ldots$ and $\{\hat{v}_{j}\}_{j \geq 1}$ constitutes an orthonormal basis of $\mathcal H$. From the estimated eigenvectors $\{\hat{v}_j\}_{j \geq 1}$, we construct our preliminary estimators of $P^N$ and $P^S$ as follows,

equation[equation omitted — 176 chars of source]

Note that $\varphi$ as a subscript in \hyperref[{eqprelim}]{\tagform@{\ref*{eqprelim}}} denotes the the number of eigenvectors used to construct $\widehat{P}^N_{\varphi}$; even if keeping such a subscript may increase notational complexity, it will be eventually helpful to avoid potential confusion when we need to consider $\widehat{P}^N_{\varphi_0}$, the projection onto $\operatorname{span}\{\hat{v}_j\}_{j=1}^{\varphi_0}$, for a hypothesized value $\varphi_0$ in Section (ref). To investigate the asymptotic properties of the preliminary estimators, we may establish the limiting behavior of either of $\widehat{P}^N_\varphi-P^N$ or $\widehat{P}^S_\varphi-P^S$ (one is simply given by the negative of the other). The following theorem deals with the former case: we hereafter write $\int A(r)$ to denote $\int_0^1 A(r)dr$ for any operator- or vector-valued function $A$ defined on $[0,1]$.

theoremSuppose that Assumptions (ref) and (ref) hold with $\varphi \geq 1$. Then \begin{align*} T(\widehat{P}^N_\varphi - P^N) \quad &\to_{\mathcal L_{\mathcal H}} \quad \mathrm{F} + \mathrm{F}^\ast, \end{align*} where $\mathrm{F}=_{\operatorname{fdd}} \left(\int W^N(r) \otimes W^N(r)\right)^\dag\left(\int d W^S(r)\otimes W^N(r) + \Gamma^{NS}\right)$.

Chang2016152 earlier established that $\|\widehat{P}^N_{\varphi}- P^N\|_{\mathcal L_{\mathcal H}} = \mathrm{O}_p(T^{-1})$. Theorem (ref) extends their result by providing a more detailed limiting behavior of $\widehat{P}^N_\varphi$. Even if Theorem (ref) implies consistency of $\widehat{P}^N_\varphi$, it does not facilitate a direct asymptotic inference since $\mathrm{F}$ depends on nuisance parameters. We particularly focus on $\Gamma^{NS}$ and $\Omega^{NS}$; as shown in Theorem (ref), $\Gamma^{NS} \neq 0$ makes the limiting random operator $\mathrm{F}+\mathrm{F}^\ast$ be not centered at zero, and $\Omega^{NS} \neq 0$ makes $W^N$ and $W^S$ be correlated. Therefore, the asymptotic limit of any statistic based on $\widehat{P}^N_\varphi$ or $\widehat{P}^S_\varphi$ is, if it exists, also dependent on $\Gamma^{NS}$ and $\Omega^{NS}$ in general. As will be shown throughout this paper, elimination of such dependence is the most important step to our modified FPCA and statistical tests based on it.

Modified FPCA for cointegrated FTS

We now propose a modification of the ordinary FPCA, which not only provides an asymptotically more efficient estimator of $P^N$ (or $P^S$) but establishes the foundation for our statistical tests that will be developed in Section (ref). Our modified FPCA may be viewed as an adjustment of the ordinary FPCA for cointegrated FTS rather than an alternative methodology; we actually take full advantage of the asymptotic properties of the preliminary estimators given in Section (ref). To introduce our methodology, we first define

equation[equation omitted — 143 chars of source]

where $\widehat{P}^N_\varphi$ and $\widehat{P}^S_\varphi$ are the preliminary estimators obtained from the ordinary FPCA in Section (ref). Then the sample counterparts of $\Omega$ and $\Gamma$ given in \hyperref[{omegat}]{\tagform@{\ref*{omegat}}} are defined by

align[align omitted — 678 chars of source]

where $\mathrm{k}(\cdot)$ (resp.\ $h$) is the kernel (resp.\ the bandwidth) satisfying the following conditions:

assumpK$\mathrm{(i)}$ $\mathrm{k}(0) = 1$, $\mathrm{k}(u) = 0$ if $u > \kappa$ with some $\kappa>0$, $\mathrm{k}$ is continuous on $[0,\kappa]$, and $\mathrm{(ii)}$ $h \to \infty$ and $h/T \to 0$ as $T \to \infty$.

\textcolor{black}{Note that the one-sided kernel $\mathrm{k}(\cdot)$ in Assumption (ref) satisfies the identical conditions given by horvath2013estimation for estimation of the long run covariance of a stationary FTS.} From the identity $I = \widehat{P}^N_\varphi + \widehat{P}^S_\varphi$, we may decompose $\widehat{\Omega}_\varphi$ and $\widehat{\Gamma}_\varphi$ as in \hyperref[{eqdecom2}]{\tagform@{\ref*{eqdecom2}}}, i.e., $\widehat{\Omega}_\varphi = \widehat{\Omega}^{NN}_\varphi + \widehat{\Omega}^{NS}_\varphi + \widehat{\Omega}^{SN}_\varphi + \widehat{\Omega}^{SS}_\varphi$ and $\widehat{\Gamma}_\varphi = \widehat{\Gamma}^{NN}_\varphi + \widehat{\Gamma}^{NS}_\varphi + \widehat{\Gamma}^{SN}_\varphi + \widehat{\Gamma}^{SS}_\varphi$. We then define a modified variable ${X}_{\varphi,t}$ as follows,

equation[equation omitted — 227 chars of source]

The $\varphi$-regularized inverse $\widehat{\Omega}^{NN}_{\varphi}{\mid_{\varphi}^\dagger}$ in \hyperref[{modvar}]{\tagform@{\ref*{modvar}}} can easily be computed from eigendecomposition of $\widehat{\Omega}^{NN}$ once we know $\varphi$. Let $\widehat{C}_\varphi$ be the sample covariance operator of $\{{X}_{\varphi,t}\}_{t=1}^T$, i.e., $\widehat{C}_\varphi = T^{-1}\sum_{t=1}^T {X}_{\varphi,t} \otimes {X}_{\varphi,t}$. Roughly speaking, replacing $\widehat{C}$ with $\widehat{C}_{\varphi}$ in \hyperref[{eigenp}]{\tagform@{\ref*{eigenp}}} has the effect of transforming $W^S$, appearing in the expression of $\mathrm{F}$ in Theorem (ref), into another Brownian motion that is independent of $W^N$; however, this does not resolve the issue that the center of the limiting operator depends on nuisance parameters. We therefore need a further adjustment, which is achieved by considering the following modified eigenvalue problem:

equation[equation omitted — 200 chars of source]

where

equation[equation omitted — 233 chars of source]

Note that the operator $\widehat{C}_{\varphi} - \widehat{\Upsilon}_{\varphi} - \widehat{\Upsilon}_{\varphi}^\ast$ is self-adjoint, so \hyperref[{eigenp2a}]{\tagform@{\ref*{eigenp2a}}} does not produce any complex eigenvalues. From the estimated eigenvectors $\{\hat{w}_j\}_{j \geq 1}$, we construct our proposed estimators of $P^N$ and $P^S$ as follows:

equation[equation omitted — 185 chars of source]

Note that $\widehat{\Pi}^N_\varphi$ and $\widehat{\Pi}^S_\varphi$ are simply obtained by replacing $\hat{v}_j$ with $\hat{w}_j$ in \hyperref[{eqprelim}]{\tagform@{\ref*{eqprelim}}}. The asymptotic properties of $\widehat{\Pi}^N_\varphi$ or $\widehat{\Pi}^S_\varphi$ are described by the following theorem.

theoremSuppose that Assumptions (ref), (ref) and (ref) hold with $\varphi \geq 1$. Then \begin{align*} T(\widehat{\Pi}^N_\varphi-P^N) \quad \to_{\mathcal L_{\mathcal H}} \quad \mathrm{G} + \mathrm{G}^\ast, \end{align*} \sloppy{where $\mathrm{G} \hspace{-0.1em}=_{\operatorname{fdd}} \hspace{-0.1em}\left(\int W^N(r) \hspace{-0.1em}\otimes\hspace{-0.1em} W^N(r)\right)^\dag \left(\int d {W}^{S.N}(r)\hspace{-0.1em}\otimes\hspace{-0.1em} W^N(r)\right)$, ${W}^{S.N}(r)\hspace{-0.1em} = \hspace{-0.1em}W^S(r)\hspace{-0.1em} -\hspace{-0.1em} \Omega^{SN}(\Omega^{NN})^\dag W^N(r)$, and ${W}^{S.N}$ is independent of $W^N$}.

The asymptotic limit of the proposed estimator $\widehat{\Pi}^N_\varphi$ is of a more convenient form than that of the preliminary estimator based on the ordinary FPCA. First note that the limiting operator $\mathrm{G}+\mathrm{G}^\ast$ is now centered at zero, whereas $\mathrm{F}+\mathrm{F}^\ast$ in Theorem (ref) is not so in general due to $\Gamma^{NS}$. In addition, $\mathrm{G}$ is characterized by two independent Brownian motions while $\mathrm{F}$ is not so in general due to $\Omega^{NS}$. These properties of $\widehat{\Pi}^N_\varphi$ not only help us develop statistical tests for examining various hypotheses about cointegration in Section (ref), but also make $\widehat{\Pi}^N_\varphi$ asymptotically more efficient than $\widehat{P}^N_\varphi$ in a certain sense, see Remark (ref).

remark\normalfont \sloppy{For any $k$, $x = (x_1,\ldots,x_k) \in \mathcal H^k$ and $y=(y_1,\ldots,y_k)\in \mathcal H^k$, let $\widehat{P}(k,x,y) = (\langle TP^N\widehat{P}^S_\varphi x_1,y_1 \rangle, \ldots, \langle TP^N\widehat{P}^S_\varphi x_k,y_k \rangle)'$ and $\widehat{\Pi}(k,x,y) = (\langle TP^N\widehat{\Pi}^S_\varphi x_1,y_1 \rangle, \ldots, \langle TP^N\widehat{\Pi}^S_\varphi x_k,y_k \rangle)'$. From similar arguments used in saikkonen1991asymptotically and harris1997principal, the following can be shown: for any choice of $k$, $x$, $y$, and $\Theta \subset \mathbb{R}^{k}$ that is convex and symmetric around the origin,} $ \lim_{T\to \infty} \text{Prob.}\{ \widehat{\Pi}(k,x,y) \in \Theta\} \geq \lim_{T\to \infty} \text{Prob.}\{ \widehat{P}(k,x,y) \in \Theta \}$ and the equality does not hold in general unless $\Omega^{NS} = \Gamma^{NS} = 0$. A detailed discussion is given at the end of Section (ref) in the Supplementary Material. This implies that the asymptotic distribution of $\widehat{\Pi}(k,x,y) $ is more concentrated at zero than that of $ \widehat{P}(k,x,y)$. A similar result can be shown for $P^S\widehat{\Pi}_{\varphi}^N$ and $P^S\widehat{P}_{\varphi}^N$. In this sense, we say that $\widehat{\Pi}_{\varphi}^N$ is more asymptotically efficient than $\widehat{P}_{\varphi}^N$.
remark\normalfont Our methodology can be applied in the case where $\dim(\mathcal H) < \infty$ without further adjustment, which is differentiated from that the existing methods (SS2019, Chang2016152, and LRS) are only discussed in a specific infinite dimensional Hilbert space setting. In this finite dimensional case, the modified FPCA is related to the PCA methodology of harris1997principal, but the latter cannot directly be applied to an infinite dimensional setting. For more details, see Section (ref) of the Supplementary Material.
remark\normalfont \textcolor{black}{We assumed that $\varphi=\dim(\mathcal H^N)$ is known in this section, but this is not generally possible for practitioners. Of course, $\varphi$ can be replaced by its consistent estimator and this does not affect the asymptotic result given by Theorem (ref). Nevertheless, as long as this replacement is necessary, our estimator is not guaranteed to be better (in the sense of Remark (ref)) than the ordinary FPCA-based estimator in finite samples; more detailed discussion on this issue is beyond the scope of the present article and hence we leave it as a future work. Despite this limitation, Theorem (ref) by itself can be a basis for statistical inference on cointegrated FTS, and one such example is our test to be presented in the next section.}

Applications: statistical tests based on the modified FPCA

We develop statistical tests to examine various hypotheses about $\mathcal H^N$ or $\mathcal H^S$ based on the asymptotic results given in Section (ref). As in the previous sections, we focus on the case without deterministic terms; the discussion is extended to allow a deterministic component in Section (ref) of the Supplementary Material.

FPCA-based test for the dimension of $\mathcal H^N$

In the previous sections, we need prior knowledge of $\varphi = \dim(\mathcal H^N)$. We here provide a novel FPCA-based test that can be applied sequentially to determine $\varphi$. Consider the following null and alternative hypotheses,

align[align omitted — 134 chars of source]

for $\varphi_0 \geq 0$. It is worth mentioning that stationarity of $\{X_t\}_{t\geq 1}$, i.e., $\varphi_0 = 0$, can be examined by the test to be developed. This means that our test can be used as an alternative to the existing tests of stationarity of FTS proposed by e.g.\ horvath2014test, kokoszka2016kpss, and aue2017testing. It should also be noted that our selection of hypotheses in \hyperref[{hypo0}]{\tagform@{\ref*{hypo0}}} is the opposite of that used for the existing tests proposed by Chang2016152 and SS2019 in the sense that the alternative hypothesis in those tests is set to $H_1 : \dim(\mathcal H^N) < \varphi_0$. Due to this difference, our test has its own advantage especially when it is sequentially applied to estimate $\varphi$; this will be more detailed in Remark (ref).

For each $\varphi_0$, we define $\widehat{P}^N_{\varphi_0}$ and $\widehat{P}^S_{\varphi_0}$ as in \hyperref[{eqprelim}]{\tagform@{\ref*{eqprelim}}} from the eigenproblem \hyperref[{eigenp}]{\tagform@{\ref*{eigenp}}}; if $\varphi_0 = 0$, $\widehat{P}^N_{\varphi_0}$ is set to zero. Given $\widehat{P}^N_{\varphi_0}$ and $\widehat{P}^S_{\varphi_0}$, define $Z_{\varphi_0,t}$, $\widehat{\Omega}_{\varphi_0}$, $\widehat{\Gamma}_{\varphi_0}$, $\widehat{\Upsilon}_{\varphi_0}$, ${X}_{\varphi_0,t}$, and $\widehat{C}_{\varphi_0}$ as in Section (ref); see \hyperref[{defitem1}]{\tagform@{\ref*{defitem1}}}-\hyperref[{eigenp21}]{\tagform@{\ref*{eigenp21}}}.

We expect from Theorem (ref) that $\{\widehat{P}^S_{\varphi_0}X_t\}_{t\geq 1}$ will behave as a stationary process if $\varphi_0 = \varphi$ (since $\|\widehat{P}^S_{\varphi_0} - P^S\|_{\mathcal L_{\mathcal H}} = \mathrm{O}_p(T^{-1})$), and the eigenvectors $\{\hat{v}_{j}\}_{j=\varphi_0+1}^{\varphi}$ are asymptotically included in $\mathcal H^N$ if $\varphi_0 < \varphi$. In the latter case, $(\langle X_t, \hat{v}_{\varphi_0+1} \rangle,\ldots, \langle X_t, \hat{v}_{\varphi} \rangle)'$ is expected to behave as a unit root process, hence $\{\widehat{P}^S_{\varphi_0}X_t\}_{t\geq 1}$ will not be stationary. Based on this idea, it may be reasonable to ask if we can distinguish the correct hypothesis from incorrect ones specifying $\varphi_0 < \varphi$ by examining stationarity of $\{\widehat{P}^S_{\varphi_0}X_t\}_{t\geq 1}$.

To simplify the discussion, we focus on the time series $\{\langle X_t, \hat{v}_{\varphi_0+1} \rangle \}_{t\geq 1}$, which is expected to be stationary under $H_0$. To examine stationarity of $\{\langle X_t, \hat{v}_{\varphi_0+1} \rangle \}_{t\geq 1}$, we consider the following test statistic,

align[align omitted — 187 chars of source]

where $\operatorname{LRV}(\langle X_t, \hat{v}_{\varphi_0+1}\rangle)=\frac{1}{T}\sum_{s=-T+1}^{T-1}\mathrm{k}({|s|}/{h})\sum_{t=|s|+1}^T \langle X_t, \hat{v}_{\varphi_0+1} \rangle \langle X_{t-|s|}, \hat{v}_{\varphi_0+1} \rangle$, and $\mathrm{k}(\cdot)$ and $h$ satisfy Assumption (ref). Note that the test statistic is given by the ratio of the sum of squared partial sums of the univariate time series ($\{\langle X_t, \hat{v}_{\varphi_0+1} \rangle\}_{t=1}^T$) to its sample long-run variance ($\operatorname{LRV}(\langle X_t, \hat{v}_{\varphi_0+1}\rangle)$). This is similar to the well-known KPSS test for examining the null hypothesis of stationarity; however, it should be noted that \hyperref[{eqkpssorigin}]{\tagform@{\ref*{eqkpssorigin}}} is not identical to the original LM statistic proposed by kwiatkowski1992testing unless $\{ \langle X_t, \hat{v}_{\varphi_0+1} \rangle\}_{t=1}^T$ is demeaned (see Appendix of their paper). If $\Omega^{NS} = \Gamma^{NS} = 0$, the asymptotic null distribution of \hyperref[{eqkpssorigin}]{\tagform@{\ref*{eqkpssorigin}}} does not depend on any nuisance parameters and is given by a functional of two independent standard Brownian motions. We thus may assess the plausibility of the null hypothesis with no significant difficulty. In general cases where $\Omega^{NS}=\Gamma^{NS}=0$ is not satisfied, the limiting distribution of \hyperref[{eqkpssorigin}]{\tagform@{\ref*{eqkpssorigin}}} depends on both of $\Omega^{NS}$ and $\Gamma^{NS}$, which are unknown in practice (see Theorem (ref) and its proof given in Section (ref) of the Supplementary Material). As expected from Theorem (ref), this issue is linked closely to the fact that the asymptotic limit of $\widehat{P}^S_{\varphi_0}$ depends on $\Omega^{NS}$ and $\Gamma^{NS}$.

The assumption that $\Omega^{NS} = \Gamma^{NS} = 0$ is too restrictive in modeling cointegrated FTS, hence a naive use of \hyperref[{eqkpssorigin}]{\tagform@{\ref*{eqkpssorigin}}} in practice to examine \hyperref[{hypo0}]{\tagform@{\ref*{hypo0}}} should be limited to very special circumstances. However, using the asymptotic results developed for our modified FPCA, the test statistic \hyperref[{eqkpssorigin}]{\tagform@{\ref*{eqkpssorigin}}} can be modified to have an asymptotic null distribution that is free of nuisance parameters in general cases. Consider the following eigenvalue problem:

equation[equation omitted — 209 chars of source]

We then replace $X_t$ and $\hat{v}_{\varphi_0+1}$ in \hyperref[{eqkpssorigin}]{\tagform@{\ref*{eqkpssorigin}}} with $X_{\varphi_0,t}$ and $\hat{w}_{\varphi_0+1}$, respectively, and obtain the following statistic:

equation[equation omitted — 235 chars of source]

There is a big gain from this simple replacement: different from \hyperref[{eqkpssorigin}]{\tagform@{\ref*{eqkpssorigin}}}, the asymptotic null distribution of \hyperref[{eqkpssorigin2}]{\tagform@{\ref*{eqkpssorigin2}}} does not depend on any nuisance parameters and is given by a functional of two independent standard Brownian motions without requiring $\Omega^{NS}=\Gamma^{NS}=0$; this is, of course, closely related to the asymptotic result given by Theorem (ref). We thus may asymptotically assess the plausibility of the null hypothesis relative to the alternative as follows: \textcolor{black}{below, we let $\mathsf{B}$ and $\mathsf{W}$ denote independent $\varphi_0$-dimensional and one-dimensional standard Brownian motions, respectively.}

propositionSuppose that Assumptions (ref), (ref) and (ref) hold. \textcolor{black}{Then $\widehat{Q}\to_d \int \mathsf{V}^2(s)ds$ under $H_0$ of \hyperref[{hypo0}]{\tagform@{\ref*{hypo0}}} and $\widehat{Q}\to_p \infty$ under $H_1$ of \hyperref[{hypo0}]{\tagform@{\ref*{hypo0}}}, where $V(r)=\mathsf{W}(r) - \int d\mathsf{W}(s) \mathsf{B}(s)' \left(\int \mathsf{B}(s)\mathsf{B}(s)' \right)^{-1}\int_{0}^{r} \mathsf{B}(s)$. The second term in the expression of $V(r)$ is regarded as zero if $\varphi_0 = 0$.}

\textcolor{black}{If $\varphi_0 = 0$ (and thus $X_{\varphi,t} = X_t$, $\widehat{C}_{\varphi_0}=\widehat{C}$, and $\widehat{\Gamma}_{\varphi_0} =\widehat{\Gamma}_{\varphi_0}^\ast = 0$), the above test is similar but not identical to the test of horvath2014test; their test is based on computation of the eigenelements of the sample long-run covariance operator of $X_t$ unlike ours is based on those of $\widehat{C}$.} The asymptotic test given in Proposition (ref) in fact corresponds to a special case of the general version of our FPCA-based test, which is discussed in the Supplementary Material (Section (ref)); see also Section (ref) for our extension of the test for the case where there is a nonzero intercept and/or a linear trend.

remark\normalfont To determine $\varphi$, we may apply our test for $\varphi_0 = 0,1,\ldots$ sequentially. Let $\hat{\varphi}$ denote the value under $H_0$ that is not rejected for the first time at a fixed significance level $\alpha$. We deduce from Theorem (ref) that $\text{Prob.\ }\{ \hat{\varphi} = \varphi\} \to 1- \alpha$ and $\text{Prob.\ }\{ \hat{\varphi} < \varphi\} \to\ 0$. Note that, even if the sequential procedure requires multiple applications of the proposed test, the correct asymptotic size, $\alpha$, is guaranteed without any adjustments; this property is shared by the existing sequential procedures developed in a Euclidean/Hilbert space setting (see e.g., Johansen1996, Johansen1996; nyblom2000tests, nyblom2000tests; Chang2016152, Chang2016152; SS2019, SS2019). If the significance level is chosen such that $\alpha \to 0$ as $T\to \infty$ then, $\text{Prob.\ }\{ \hat{\varphi} = \varphi\} \to 1$.
remark\normalfont In Chang2016152 and SS2019, $\varphi$ is determined by testing the following hypotheses: for a pre-specified positive integer $\varphi_{\max}$ and $\varphi_0 = \varphi_{\max},\varphi_{\max}-1,\ldots,1$, \begin{align*} H_0 : \dim(\mathcal H^N) = \varphi_0 \quad against \quad H_1 : \dim(\mathcal H^N) < \varphi_0, \end{align*} sequentially until $H_0$ is not rejected for the first time. Note that we need a prior information on an upper bound of $\varphi$, i.e., $\varphi_{\max}$, for this type of procedure. However, in an infinite dimensional setting, there is no natural upper bound of $\dim(\mathcal H^N)$. On the other hand, our sequential procedure described in Remark (ref) does not require such a prior information; the procedure first examines the null hypothesis $H_0 : \dim(\mathcal H^N) = 0$, and $0$ is the minimal possible value of $\dim(\mathcal H^N)$. \textcolor{black}{Even with this difference resulting from a different formulation of the alternative hypothesis, our method can be compared with the existing ones as a way to estimate $\varphi$. From our simulation study we find that our procedure seems to perform comparably with the recent method proposed by SS2019 and tends to work better if the stationary component is persistent; see Section (ref) and Table (ref).} If $\varphi \geq 1$ is known in advance, $\varphi$ can also be estimated by the eigenvalue ratio criterion proposed by LRS under their assumptions. Our procedure given in Remark (ref) is significantly differentiated from any of these existing procedures, and thus ours can complement those in practice.

Tests of hypotheses about cointegration

Practitioners may be interested in testing various hypotheses about $\mathcal H^N$ or $\mathcal H^S$. For example, we may want to test if a specific element $x_0$ is included in $\mathcal H^N$ or the span of a specified set of vectors contains $\mathcal H^N$. More generally, we here consider testing the following hypotheses: for a specified subspace $\mathcal M$,

align[align omitted — 401 chars of source]

Let $P^{{\mathcal M}}$ denote the projection onto ${{\mathcal M}}$ and assume that $\varphi$ is known; in practice, we may apply any statistical method discussed in Remarks (ref) and (ref) to determine $\varphi$. In this case, \hyperref[{hptest1}]{\tagform@{\ref*{hptest1}}} and \hyperref[{hptest2}]{\tagform@{\ref*{hptest2}}} can be tested by examining the dimension of the attractor space associated with the residuals $\{(I-P^{{\mathcal M}})X_t\}_{t\geq 1}$.

Specifically, let $\widehat{Q}_{\mathcal M}$ be the test statistic computed as in \hyperref[{eqkpssorigin2}]{\tagform@{\ref*{eqkpssorigin2}}} from $\{(I-P^{{\mathcal M}})X_t\}_{t=1}^T$ for $\varphi_0 = \varphi-\dim(\mathcal M)$ (resp.\ $\varphi_0 = 0$) if we examine \hyperref[{hptest1}]{\tagform@{\ref*{hptest1}}} (resp.\ \hyperref[{hptest2}]{\tagform@{\ref*{hptest2}}}). Then we may deduce from Proposition (ref) that, under $H_0$ of \hyperref[{hptest1}]{\tagform@{\ref*{hptest1}}}, $\widehat{Q}_{\mathcal M} \to_d \int \mathsf{W}^2(s)ds$ while it diverges to infinity under $H_1$ of \hyperref[{hptest1}]{\tagform@{\ref*{hptest1}}}.

remark\normalfont We assess the plausibility of $H_0$ in \hyperref[{hptest1}]{\tagform@{\ref*{hptest1}}} or \hyperref[{hptest2}]{\tagform@{\ref*{hptest2}}} by checking if the dimension of the attractor space associated with $\{(I-P^{{\mathcal M}})X_t\}_{t\geq1}$ is higher than that is implied by $H_0$. This is possible since our test is consistent against the alternative hypotheses with higher dimensional attractor spaces. Unlike our proposed test, the top-down tests proposed by Chang2016152 and SS2019 cannot be used in this way.

Numerical studies on the proposed tests: a brief summary

\textcolor{black}{In Sections (ref)-(ref) of the Supplementary Material, we examine the finite sample properties of our proposed tests for \hyperref[{hypo0}]{\tagform@{\ref*{hypo0}}}, \hyperref[{hptest1}]{\tagform@{\ref*{hptest1}}} and \hyperref[{hptest2}]{\tagform@{\ref*{hptest2}}}. The tests seem to perform reasonably but slightly over-reject the null hypothesis when the stationary component is persistent. We also compare the finite sample performance of our test as a way to estimate $\varphi=\dim(\mathcal H^N)$ (Remark (ref)) with the recent testing procedure of SS2019. We found that, overall, ours performs comparably with theirs in general; moreover, in our simulation setting with a persistent stationary component, ours tends to outperform theirs. The performance of our testing procedure seems to be dependent on the choice of $h$ while the method proposed by SS2019 does not require such a bandwidth choice (since theirs does not require any long-run covariance estimation). Despite this disadvantage, we may conclude from the simulation evidence that our testing procedure can be an attractive alternative or complement to theirs. Of course, as mentioned in Remark (ref), the proposed test in the present paper can also be used to examine various hypotheses on cointegration from a relevant residual FTS, which may be another advantage of ours over the test of SS2019. In addition to the simulation studies, we in Sections (ref) and (ref) illustrate our methodology with two empirical examples: U.S.\ age-specific employment rates and monthly earning densities.}

Concluding remarks : some cautions and future direction

This note adds some novel asymptotic results for the FPCA-based estimator of $P_N$ or $P_S$ of cointegrated FTS. We then propose a modification of FPCA for statistical inference on the cointegrating behavior. Some cautions need to be made for practitioners. First, in order for the modified FPCA to work well, we need to accurately estimate $P_N$ first; however, as remarked by SS2019, it is not easy particularly if $\dim(\mathcal H^N)$ is large but the sample size is small. Moreover, this note does not address in detail on the choice of the bandwidth parameter $h$ used to estimate long-run covariance operators of some residual time series considered in the note. Given that all such time series are expected to be stationary, the method proposed by rice2017plug may be used, but this issue obviously requires a further investigation. Despite these limitations, our proposed methodology gives us some useful results, such as statistical tests on $\dim(\mathcal H^N)$ and cointegrating properties. \textcolor{black}{To put it in more detail, our test as a way to estimate $\dim(\mathcal H^N)$ performs comparably with the existing methods (it is found to perform better for some simulation setting considered in Section (ref)), and the test by itself can be used to examine various hypotheses on cointegration (see Remark (ref)); see also Remarks (ref) and (ref).} The theoretical results in this paper may be used for more sophisticated statistical models requiring dimension-reduction of a FTS exhibiting unit-root-like behavior; a potential example may be the cointegrating regression model involving functional regressors. Provided that practitioners want to employ FPCA-based dimension-reduction, which has been widely used in various contexts, our results will be helpful to derive detailed asymptotic properties of estimators for such models.