EconBase
← Back to paper

Determining the dimension of factor structures in non-stationary large datasets

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

85,749 characters · 13 sections · 82 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Determining the dimension of factor structures in non-stationary large datasets

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 113 chars of source]

} \fi

abstractWe propose a procedure to determine the dimension of the common factor space in a large, possibly non-stationary, dataset. Our procedure is designed to determine whether there are (and how many) common factors (i) with linear trends, (ii) with stochastic trends, (iii) with no trends, i.e. stationary. Our analysis is based on the fact that the largest eigenvalues of a suitably scaled covariance matrix of the data (corresponding to the common factor part) diverge, as the dimension $N$ of the dataset diverges, whilst the others stay bounded. Therefore, we propose a class of randomised test statistics for the null that the $p$-th eigenvalue diverges, based directly on the estimated eigenvalue. The tests only requires minimal assumptions on the data, and no restrictions on the relative rates of divergence of $N$ and $T$ are imposed. Monte Carlo evidence shows that our procedure has very good finite sample properties, clearly dominating competing approaches when no common factors are present. We illustrate our methodology through an application to US bond yields with different maturities observed over the last 30 years. A common linear trend and two common stochastic trends are found and identified as the classical level, slope and curvature factors.

{\it Keywords:} Common factors; Unit roots; Common trends; Randomised tests.

Introduction and main ideas

In this paper, we propose a methodology to estimate the dimension of the space spanned by the common (non-stationary) factors in the large approximate factor model

equation[equation omitted — 67 chars of source]

where $\mathcal{F}_{t}$ is the $r\times 1$ vector of common factors and $\Lambda $ is an $N\times r$ matrix of factor loadings. We will also make use of the scalar version of ((ref))

equation[equation omitted — 91 chars of source]

with $1\leq i\leq N$ and $1\leq t\leq T$. Although the relevant assumptions are detailed in the remainder of the paper, in ((ref)) we are assuming that there are three possible categories of common factors in the vector $\mathcal{F}_{t}$: factors with a linear trend and an additional, either an $I\left( 1\right) $ or $I\left( 0\right) $, zero mean component; pure, zero mean $I\left( 1\right) $ factors with no trends; and, finally, stationary common factors. Each group may well have dimension zero, e.g. factors with linear trends may not be present, etc. We also assume, throughout the paper, that the idiosyncratic terms $u_{i,t}$ are $I\left( 0\right) $ for each $i$. Based on the classification above, we develop a technique to estimate the number of common factors which have linear trends, and, separately, the ones which are zero mean $I\left( 1\right) $. In particular, we use the eigenvalues of the second moment matrix of the data, checking whether they diverge to infinity as $\min \left( N,T\right) \rightarrow \infty $, due to the presence of common factors, or whether they are bounded. In order to construct tests for the asymptotic behaviour of the eigenvalues, (i) we derive bounds on their divergence rates, and (ii) based on those bounds we propose a randomisation procedure which produces a statistic for which we are able to derive the asymptotic behaviour under the null and the alternative hypotheses.

Determining the presence (or not), and the number of non-stationary factors in (ref) can be useful in a variety of applications. First, we can assess the presence of unit roots in a large panel $X_{t}$ - see e.g. mp04, baing04, bai2006estimation, bai2009panel, kapetanios2011panels, and pesaran2013panel. Second, it is easy to see that $I\left( 1\right) $\ common factors in ( (ref)) entail the presence of common trends and therefore cross-unit cointegration among the components of $X_{t}$ - see e.g. EP1994, SW1988, gengenbach2009, and zyr18. Indeed, if in (ref) there are, say, $r$ common $I(1)$ factors (common trends), then there are $(N-r)$ cointegration relations - see also onatski2016 for an alternative approach to cointegration in large VARs. Empirical applications considering common $I(1)$ factors include: bai04 on employment fluctuations across 60 industries in the US, moon2007 on a panel of interest rates at different maturities in the US\ and Canada, and engel2015, who use ((ref)) as part of a strategy to develop a forecasting technique applied to a panel of bilateral US dollar rates against 17 OECD countries. Panel models with linear trends have also been employed in the context of modelling macro-econocmic data - see maciejowska2010 - and have also proven helpful in modelling US temperature data - see chenwu. In all these applications, the first step of the analysis would be the determination of the number of non-stationary and stationary common factors.

Starting at least from chamberlainrothschild83, the literature has developed a plethora of contributions to determine the number of common factors in a large panel. Most methodologies focus on the case of stationary datasets, and existing approaches can be broadly grouped into two categories. Several studies rely on setting a threshold for the eigenvalues of the covariance or of other second moment matrices of the $ X_{i,t}$s - see baing02, hallinliska07, and ABC10. In addition to this, the literature has explored the possibility of using ratios of adjacent eigenvalues - see onatski09,onatski10, lam2012, and ahnhorenstein13. Although the two approaches have different merits, the rationale underpinning them is the same: in the presence of, say, $r$ common factors, the largest $r$ eigenvalues of the covariance matrix of the $X_{i,t}$s\ diverge to infinity as $\min \left( N,T\right) \rightarrow \infty $, whilst all the remaining eigenvalues stay bounded.

Fewer contributions are available to deal directly with factor models for non-stationary data. In particular, developing an inferential theory for $\Lambda $ and $\mathcal{F}_{t}$ in a model similar to ((ref)) has been paid significant attention by the statistical literature: examples include baing04, bai04, penaponcela, zhangpangao, and zyr18. More specifically, bai04 extends the results by penaponcela to the large dimensional setting, i.e. letting $N\to\infty$, and develops the inferential theory and a criterion, from which it is possible to estimate the number of common stationary and non-stationary factors. zyr18 propose a method based on the ratio of eigenvalues of a transformation of the long-run covariance matrix to find the number of $I(d)$ factors for $d\ge 0$, but imposing the constraint $\frac N{T^\kappa}\to c\in(0,\infty)$, for $\kappa\in\left(0,\frac 12\right)$, as $\min\left( N,T\right) \to \infty$. Note that none of these two approaches considers the case of common linear trends. To this end, maciejowska2010 develops an inferential theory for estimated common factors and loadings in a set-up like ((ref)), where also linear trends are allowed, but no criteria for the determination of the number of common factors are proposed. Moreover, none of these contributions deals with the case in which there are no common (non-stationary or stationary) factors: thus, these approaches cannot detect whether $X_t$ is stationary, i.e. it has no common trends, or not. Finally, baing04 propose a method to assess the presence or not of stochastic trends in large panels. However, in their setup it is assumed that at least one stationary factor is always present. Moreover, in order to assess the presence of non-stationary factors their procedure requires to estimate first the number of factors and the factors themselves using differenced data and then to test for unit roots or cointegration in the cumulated estimated factors. Although asymptotically valid, such an approach might suffer of efficiency loss due to its two-steps.

Our paper fills the gaps mentioned above. Although the main arguments are laid out in the remainder of the paper, here we present a heuristic preview of how the procedure works.

To begin with, in the presence of linear trends, it can be expected that the sample second moment matrix of $X_{t}$ will diverge as fast as $T^{3}$. Also, due to the well known eigenvalue separation property of large factor models, it can be expected that the eigenvalues corresponding to common factors should diverge as fast as $N$ (see maciejowska2010). This suggests considering the eigenvalues of $T^{-3}\sum_{t}X_{t}X_{t}^{\prime }$ (denoted as, say, $\nu _{1}^{\left( p\right) }$) to decide between

equation*[equation* omitted — 185 chars of source]

as $\min \left( N,T\right) \rightarrow \infty $; the test can be carried out for $p=1,2,...$, stopping as soon as the null is rejected. Similarly, considering the zero mean, $I\left( 1\right) $ common factors, the Functional Central Limit Theorem (FCLT) suggests that the second moment matrix of $X_{t}$ will diverge as fast as $ T^{2}$, again with the eigenvalues corresponding to the common factors diverging as fast as $N$ (see bai04). Thus, one could study the eigenvalues of $ T^{-2}\sum_{t}X_{t}X_{t}^{\prime }$ (denoted as, say, $\nu _{2}^{\left( p\right) }$), and decide between

equation*[equation* omitted — 185 chars of source]

as $\min \left( N,T\right) \rightarrow \infty $, carrying out the test as above. The output of these two steps is an estimate of the number of common factors which have a linear trend and of those which are genuinely zero mean $I\left( 1\right) $ processes, respectively. Note that, in both steps, if we reject the null-hypothesis when $p=1$, we are in fact saying that there are no common factors. This approach could be complemented by using $ T^{-1}\sum_{t}\Delta X_{t}\Delta X_{t}^{\prime }$ and determining the number of total common factors as suggested in trapani17, which would provide an indirect estimate of the number of common stationary factors.

From a technical point of view, the implementation of the algorithm described above presents one difficulty: we are unable to construct test statistics which converge to a distributional limit under the null hypotheses, and the best result we can obtain are rates. Thus, we base our tests on randomising the test statistic. This approach builds on an idea of pearson50, and it has been exploited in numerous contexts - see e.g. corradi2006, bandi2014 and trapani17. A major advantage of this procedure is that only (strong) rates are needed, and these can be derived under quite general assumptions. In particular, we derive our rates (and, thus, we are able to apply our test) under no restrictions on the relative rates of divergence of $N$ and $ T $ as they pass to infinity, which can be compared with the standard restriction that as $\min \left (N,T\right) \rightarrow \infty $, $\frac{N }{T}\rightarrow c\in \left( 0,\infty \right) $, often assumed in random matrix theory (see also onatski2016, where a similar restriction is needed); this entails that our procedure can be applied to virtually any dataset, being particularly useful when either dimension is much bigger than the other. Also, our theory requires milder restrictions on the finiteness of moments than other contributions in the literature - see e.g. bai04 - and allows for arbitrary levels of (weak) cross-correlation among the idiosyncratic errors $ u_{i,t}$.

The remainder of the paper is organised as follows. In Section (ref), we spell out the main assumptions and (in Section (ref)) we study the strong rates of convergence of the eigenvalues of various rescalings of the second moment matrix of $X_{t}$. The testing algorithm is presented in Section (ref). Numerical evidence from simulations is in Section (ref), where we also report an empirical illustration. Finally, Section (ref) concludes. Proofs and technical results are in appendix.

NOTATION. We define the Euclidean norm of a vector $a=\left[ a_{1},...,a_{n} \right] $ as $\left\Vert a\right\Vert =\left( \sum_{i=1}^{n}a_{i}^{2}\right) ^{1/2}$; \textquotedblleft a.s.\textquotedblright\ stands for \textquotedblleft almost surely\textquotedblright , with orders of magnitude for an a.s. convergent sequence (say $s_{T}$) being denoted as $O_{a.s.}\left( T^{\varsigma }\right) $ and $o_{a.s.}\left( T^{\varsigma }\right) $ when, for some $ \epsilon >0$ and $\tilde{T}<\infty $, $P\left[ \left\vert T^{-\varsigma }s_{T}\right\vert <\epsilon \text{ for all }T\geq \tilde{T}\right] =1$ and $ T^{-\varsigma }s_{T}\rightarrow 0$ a.s., respectively; $I_{A}\left( x\right) $ is the indicator function of a set $A$; finally, $C_{0}$, $C_{1}$, etc... denote positive, finite constants whose value may differ from line to line. Other relevant notation is introduced later on in the paper.

Theory

In this section, we (a) lay out our main model - equation ((ref)) - in more precise terms and spell out the relevant assumptions (Section (ref)), and (b) present the main results on the eigenvalues of various rescaled versions of the sample second moment matrix of $X_{t}$ (Section (ref)).

Model and assumptions

Recall the scalar version of our model ((ref)):

equation[equation omitted — 111 chars of source]

where $\lambda _{i}$ and $\mathcal{F}_{t}$ are $r\times 1$ vectors. We begin with a representation result which, essentially, states that the number of common factors with a linear trend (and, possibly, further components which may be $I\left( 0\right) $ or $I\left( 1\right) $) can be either zero - no common factors with linear trends - or $1$. This result is originally due to maciejowska2010, and we report it hereafter, as a lemma, for convenience. We assume that

equation[equation omitted — 83 chars of source]

where $A$ is a non-zero $r\times 1$ vector, $B$ an $r\times r$ matrix, and, more importantly, in ((ref)) $d_{1}$ is a dummy variable, which has the purpose to entertain the possibility that there are linear trends or not, according as $d_1=1$ or $0$, respectively. As far as the $r$-dimensional vector $\psi _{t}$ is concerned, its components are allowed to be a mixture of $I\left( 0\right) $ and $I\left( 1\right) $ processes, with no linear trends.

We consider the following assumption, which ensures that the $\mathcal{F}_{t}$s are fully identified.

assumptionIt holds that: (i) $A$ is non-zero; (ii) $rank\left( B\right) =r$; (iii) the vector $\psi _{t}$ can be rearranged and partitioned as $\left[ \psi _{at}^{\prime },\psi _{bt}^{\prime }\right] ^{\prime }$, where $\psi _{at}\sim I\left( 1\right) $ has dimension $r_{2}+d_{2}$ and $ \psi _{bt}\sim I\left( 0\right) $ has dimension $r_{3}+\left( 1-d_{2}\right) $, where $d_{2}$ is a dummy variable.

By part (ii) of Assumption (ref), $B$ has full rank, which ensures the identification of the vector $\mathcal{F}_{t}$ irrespective of whether there is a trend or not. When there are trends, that is when $d_{1}=1$, part (i) of the assumption ensures that they do have an impact on $\mathcal{F}_{t}$. Finally, by part (iii) there could be both $I\left( 1\right) $ and $I\left( 0\right) $ factors in the vector $\psi _{t}$, sorted in no particular order.

lemmaUnder Assumption (ref), model ((ref)) can be equivalently represented as \begin{equation} X_{i,t}=\lambda _{i}^{\left( 1\right) }f_{t}^{\left( 1\right) }+\lambda _{i}^{\left( 2\right) \prime }f_{t}^{\left( 2\right) }+\lambda _{i}^{\left( 3\right) \prime }f_{t}^{\left( 3\right) }+u_{i,t}, \quad 1\le i\le N, \end{equation} where $\lambda _{i}^{\left( 1\right) }$ and $f_{t}^{\left( 1\right) }$ are $r_1\times 1$ with $0\le r_1\le 1$, $\lambda _{i}^{\left( 2\right)}$ and $f_{t}^{\left( 2\right) }$ are $r_2\times 1$ vectors with $0\le r_2\le \min(N,T)$, $\lambda _{i}^{\left( 3\right)}$ and $f_{t}^{\left( 3\right) }$ are $r_3\times 1$ vectors with $0\le r_3\le \min(N,T)$, and such that $r=r_1+r_2+r_3$ and $\lambda_{i}'=(\lambda_{i}^{(1)'}\lambda_{i}^{(2)'}\lambda_{i}^{(3)'})$ for all $i$.\\ Moreover, the common non-stationary factors are defined by the following equations \begin{eqnarray} f_{t}^{\left( 1\right) } &=&d_{1}t+d_{2}f_{t}^{\left( 1\right) \dag }+\left( 1-d_{2}\right) g_{t}, \\ f_{t}^{\left( 1\right) \dag } &=&f_{0}^{\left( 1\right) \dag }+\sum_{j=1}^{t}e_{t}^{\left( 1\right) }, \\ f_{t}^{\left( 2\right) } &=&f_{0}^{\left( 2\right) }+\sum_{j=1}^{t}e_{t}^{\left( 2\right) }, \end{eqnarray} where in ((ref))-((ref)): $f_{t}^{\left( 1\right)\dag }$, $g_t$ and $e_t^{(1)}$ are $r_{1}\times 1$ vectors, $e_t^{(2)}$ is an $r_{2}\times 1$ vector, $e_{t}^{\left( 1\right) }$, $e_{t}^{\left( 2\right) }$, $g_{t}$ and $f_{t}^{\left( 3\right) }$ are $I\left( 0\right) $, and $d_1$ and $d_2$ are dummy variables.

Lemma (ref) states that the number of linear trends is either zero or one: if an identified $k$-dimensional vector of common factors has linear trends, this is tantamount to an identified $k$-dimensional vector of common factors where only the first factor has a linear trend. When $r_{1}=1$ and $d_1=1$, we show in Theorem (ref) below, that it does not matter whether the remainder $d_{2}f_{t}^{\left( 1\right) \dag }+\left( 1-d_{2}\right) g_{t}$ is $I\left( 1\right) $ or $I\left( 0\right) $: the trend component is the one that dominates. When $r_{1}=0$, there are no linear trends in the factor structure; in this case, $f_{t}^{\left( 1\right) }$ can be $I\left( 1\right) $ or $I\left( 0\right) $, according as $d_{2}=1$ or $0$.

Let us denote as $r^*$ the number of non-stationary factors, and as $r$ the total number of factors. Then, based on ((ref))-((ref)), the numbers of common factors in $X_{i,t}$ are summarised in the table below.

center[center omitted — 539 chars of source]

Recall that we allow for the possibility of having any of the numbers $r_1$, $r_2$, $r_3$, $r^*$, or even $r$, to be equal to zero. On the other hand, if there is no linear trend ($d_1=0$), we have at most $r_1+r_2$ zero-mean $I(1)$ factors and $r_1+r_3$ zero-mean $I(0)$ factors, while if there is a linear trend ($d_1=0$), we have at most $r_2$ zero-mean $I(1)$ factors and $r_3$ zero-mean $I(0)$ factors.

We now spell out the main assumptions. Consider the vector of zero-mean $I(1)$ factors: $f_{t}^{\ast }$, where $f_{t}^{\ast }=\left[ f_{t}^{\left( 1\right) \dag },f_{t}^{\left( 2\right) \prime }\right] ^{\prime }$, and consider the $I(0)$ vector $e_{t}$, where $e_{t}=\left[ e_{t}^{\left( 1\right) },e_{t}^{\left( 2\right) \prime }\right] ^{\prime }$. Both $f_{t}^{\ast }$ and $e_{t}$ are $[r_2+r_1(1-d_1)d_2]\times 1$ vectors.\footnote{Note that if $d_1=1$ or $d_2=0$ then these vectors have dimension $r_2$ and are given by $f_{t}^{\ast }=f_{t}^{\left( 2\right) }$ and $e_{t}=e_{t}^{\left(2\right) }$; on the other hand if $d_1=0$ and $d_2=1$ then the vectors have dimension $r_1+r_2$, thus become scalars if $r_2=0$ and $r_1=1$.}

We define the long-run covariance matrix associated with $f_{t}^{\ast }$ as

equation[equation omitted — 128 chars of source]
assumptionLet $\kappa >0$. It holds that (i) $E\left\Vert e_{t}\right\Vert ^{4+\kappa }<\infty $ for all $t$; (ii) $E\left\vert f_{0}^{* }\right\vert ^{4+\kappa }<\infty $; (iii) $\Sigma _{\Delta f^*}$ is positive definite; (iv) there exists, on a suitably enlarged probability space, an $ \left( r_{2}+d_{2}\right) $-dimensional standard Wiener process $W\left( t\right) $ such that, for some $\epsilon >0$, \begin{equation*} \sup_{1\leq j\leq t}\left\Vert f_{j}^{\ast }-\Sigma _{\Delta f^*}^{1/2}W\left( j\right) \right\Vert =O_{a.s.}\left( t^{1/2-\epsilon }\right); \end{equation*} (v) $E\left\Vert \sum_{t=1}^{T}e_{t}\right\Vert ^{2+\kappa }\leq C_{0}\left( \sum_{t=1}^{T}E\left\Vert e_{t}\right\Vert ^{2}\right) ^{\frac{2+\kappa }{2}}$; (vi) $E\left\Vert \sum_{t=1}^{T}f_{t}^{\ast }f_{t}^{\ast \prime }\right\Vert ^{2}$ $\leq $ $ C_{0}T^{4}$.

Assumption (ref) poses some restrictions on the common $I(1)$ factors. Parts (i) and (ii) require the existence of at least the second moment of the innovation $e_{t}$ and of the initial condition $ f_{0}^{\ast }$ respectively. Part (iii) is a standard requirement, which rules out that the common, zero mean $I(1)$ factors are cointegrated: in essence, this ensures that the number of $I(1)$ common factors is genuinely $r_{2}+d_{2}$. Part (iv) states that a strong approximation exists for the partial sums process $f_{t}^*$. Although this is a high-level assumption, we prefer to write it in this form as opposed to spelling out more primitive assumptions, since this makes the set-up more general. Part (v) is a Burkholder-type inequality (see e.g. linbai, p. 108). Finally, part (vi) can be verified e.g. under independence and finite fourth moments.

An important implication is that $e_{t}$ is allowed to be (weakly) dependent over time. Considering part (iv) in particular, starting from the seminal paper by berkes1979, the literature has developed several refinements of the Strong Invariance Principle (SIP) for random vectors. In particular, liu2009 derive the SIP for stationary causal processes, a wide class which includes e.g. conditional heteroskedasticity models, Volterra series, and data generated by dynamical systems - see wu07. Thus, part (iv) of the assumption accommodates for a wide variety of commonly considered DGPs.

assumptionIt holds that: (i) (a) $\max_{1\leq i\leq N,1\leq t\leq T}E\left\vert u_{i,t}\right\vert ^{4}<\infty $; (b) $\max_{1\leq t\leq T}E\left\Vert f_{t}^{\left( 3\right) }\right\Vert ^{4}$ $<$ $\infty $; and (c) $\max_{1\leq t\leq T}E\left\vert g_{t}\right\vert ^{4}$ $<$ $\infty $; (ii) (a) $\max_{1\leq i\leq N}E\left\Vert \sum_{t=1}^{T}f_{t}^{\ast }u_{i,t}\right\Vert ^{2}$ $\leq $ $C_{0}T^{2}$; (b) $E\left\Vert \sum_{t=1}^{T}f_{t}^{\ast }f_{t}^{\left( 3\right) \prime }\right\Vert ^{2}$ $ \leq $ $C_{0}T^{2}$; and (c) $E\left\Vert \sum_{t=1}^{T}f_{t}^{\ast }g_{t}\right\Vert ^{2}$ $\leq $ $C_{0}T^{2}$; (iii) $E\left\Vert \sum_{t=1}^{T}tf_{t}^{\ast }\right\Vert ^{2}$ $\leq $ $C_{0}T^{5}$; (iv) (a) $\max_{1\leq i\leq N}E\left\vert \sum_{t=1}^{T}tu_{i,t}\right\vert ^{2}$ $ \leq $ $C_{0}T^{3}$; (b) $E\left\Vert \sum_{t=1}^{T}tf_{t}^{\left( 3\right) }\right\Vert ^{2}$ $\leq $ $C_{0}T^{3}$; and (c) $E\left\vert \sum_{t=1}^{T}tg_{t}\right\vert ^{2}$ $\leq $ $C_{0}T^{3}$; (v) $E\left\Vert \sum_{t=1}^{T}f_{t}^{\ast }f_{t}^{\ast \prime }\right\Vert ^{2}$ $\leq $ $ C_{0}T^{4}$.

Assumption (ref) deals with the idiosyncratic terms $u_{i,t}$ and the stationary factors. Part (i) requires the existence of the $4$-th moments, which is a milder assumption than the customary $8$-th moment existence requirement - see bai04. Part (ii) could be shown from more primitive assumptions; indeed, a prototypical assumption would require $ e_{t} $ and $u_{i,t}$ to be independent of each other and i.i.d. over time - in such a case, explicit calculations would yield part (ii)(a). Parts (iii) and (iv) could again be shown from more primitive assumptions; for example, part \textit{(iv)(a)} would automatically follow if $Eu_{i,t}^{2}<\infty $ and $u_{i,t}$ is \textit{i.i.d.} across time. Similarly, it could be verified that part \textit{(iii)} holds whenever $E\left\Vert e_{t}\right\Vert ^{2}<\infty $ and $e_{t}$ is \textit{ i.i.d.} across $t$.

We now spell out the assumptions for the $N\times r$ loadings matrix $\Lambda =[\lambda _{1}|...|\lambda _{N}]^{\prime }$.

assumptionThe loadings $\Lambda $ are non-stochastic with (i) $\max_{1\leq i\leq N}\left\Vert \lambda _{i}\right\Vert <\infty $; (ii) $ \lim_{N\rightarrow \infty }\frac{\Lambda ^{\prime }\Lambda }{N}\rightarrow \Sigma _{\Lambda }$, where the matrix $\Sigma _{\Lambda }$ is positive definite.

Assumption (ref) is standard in this literature - see e.g. bai04. One consequence of part (ii) and Lemma (ref) is that every diagonal block of $\Sigma _{\Lambda }$, defined by the loadings of $f_t^{(1)}$, $f_t^{(2)}$ or $f_t^{(3)}$, is also positive definite. Note that the assumption requires the loadings to be non-stochastic; however, this could be relaxed to the case of random loadings, with no changes to the main arguments in the paper.

Another, important consequence of Assumption (ref) is that the common factors belonging in each category are \textquotedblleft strong\textquotedblright\ or \textquotedblleft pervasive\textquotedblright . We postpone a discussion of this aspect, and of the possibility of extending this set-up, until Section (ref).

Asymptotic behavior of eigenvalues

We base inference on the two matrices

eqnarray[eqnarray omitted — 179 chars of source]

We denote the $p$-th largest eigenvalues of $\Sigma _{1}$\ and $\Sigma _{2}$ \ as $\nu _{1}^{\left( p\right) }$and $\nu _{2}^{\left( p\right) }$ respectively. The the asymptotic behaviour of those eigenvalues is studied in the following Theorem.

theoremUnder Assumptions (ref)-(ref), it holds that, for every positive, bounded constants $C_{p}$, there is a slow varying sequence \begin{equation*} l_{N,T}=\left( \ln N\right) ^{1+\epsilon }\left( \ln T\right) ^{\frac{3}{2}+\epsilon }, \end{equation*} with $\epsilon >0$, and some random $N_{0}$ and $T_{0}$ such that, for all $N\geq N_{0}$ and $T\geq T_{0}$, \begin{eqnarray} \nu _{1}^{\left( p\right) } &\geq &C_{p}N,\quad\quad\quad\quad\quad\quad\; for p\leq r_{1}, \\ \nu _{1}^{\left( p\right) } &=&O_{a.s.}\left( \frac{N}{\sqrt{T}}l_{N,T}\right),\quad for p>r_{1}, \end{eqnarray} and \begin{eqnarray} \nu _{2}^{\left( p\right) } &\geq &C_{p}\frac{N}{\ln \ln T}, \quad\quad\quad\quad\; for 1\leq p\leq r_{2}+\max \left\{ r_{1},d_{2}\right\} , \\ \nu _{2}^{\left( p\right) } &=&O_{a.s.}\left( \frac{N}{\sqrt{T}} l_{N,T}\right),\quad for p>r_{2}+\max \left\{ r_{1},d_{2}\right\} . \end{eqnarray}

Theorem (ref) is a separation result for the eigenvalues corresponding to common factors in $\Sigma _{1}$ and $\Sigma _{2}$ and is our first contribution.

Equations ((ref)) and ((ref)) refer to the eigenvalues of $\Sigma _{1}$. The results state that the first $r_{1}$ eigenvalues diverge to infinity at a rate $N$; conversely, the remaining eigenvalues have a smaller magnitude. We pose no restrictions on the relative rate of divergence between $N$ and $T$ as they pass to infinity. Thus, the magnitude of $\nu _{1}^{\left( p\right) }$, when $p>r_{1}$, may be very large; it is however smaller - by a factor $T^{-1/2}$ - compared to that of $\nu _{1}^{\left( p\right) }$ when $p\leq r_{1}$. In the definition of $\Sigma _{1}$, there is a denominator given by $T^{3}$: intuitively, this is due to the fact that the presence of a drift in the common factor $f_{t}^{\left( 1\right) }$ creates a linear trend. Norming by $T^{3}$ is needed in order to make the trend component converge.

Equations ((ref)) and ((ref)) refer to the eigenvalues of $\Sigma_{2}$. This matrix is normalised by $T^{2}$: the main idea is that we wish to separate the eigenvalues corresponding to non-stationary factors from the other ones. The partial sums of $f_{t}^{\ast }f_{t}^{\ast \prime }$ should grow at least as fast as $T^{2}$ by the CLT in functional spaces; the result in ((ref)) follows from this intuition, although, since we need an a.s. rate, it is based on the Law of the Iterated Logarithm (see donsker1977). Similarly to $\Sigma _{1}$, the remaining eigenvalues may also diverge, but this will happen at a slower rate. Equation ((ref)) illustrates the separation result, through the $T^{-1/2}$ term. Following the proof of the theorem, it could be readily shown that, if the idiosyncratic components $u_{i,t}$ were $I(1)$, the upper bound for $\nu _{2}^{\left( p\right) }$ when $p>r_{2}+\max \left\{r_{1},d_{2} \right\}$ would be $O_{a.s.}\left(Nl_{N,T} \right)$ - in essence, in this case a separation result could not be shown, whence the need to assume that the $u_{i,t}$s are $I(0)$. On the other hand, one could envisage a situation where only a fraction of the $u_{i,t}$s are $I(1)$ - say $O\left( N^{\alpha_{0}}\right)$, with $\alpha_{0}<1$. In such a case, by adapting the proof of Lemma (ref) it can be shown that the upper bound in ((ref)) would become $O_{a.s.}\left(N^{\alpha_{0}}l_{N,T} \right)+O_{a.s.}\left(\frac{N}{\sqrt{T}}l_{N,T} \right)$, and thus a separation result would obtain.

Note that Theorem (ref) provides only rates: no distributional results are available. When data are stationary, wang16 derive an asymptotic distribution for the estimates of the diverging eigenvalues of the sample covariance matrix. We do not know, however, if this can also be done for the $\nu _{1}^{\left( p\right) }$s and the $\nu _{2}^{\left( p\right) }$s. Hence, in what follows we will rely only on rates.

Finally, in order to construct the relevant test statistics, we will also make use of the first differenced version of ((ref)):

equation[equation omitted — 118 chars of source]
assumptionIt holds that: (i) $E\left( \Delta \mathcal F_{j,t}\Delta u_{i,t}\right) =0$ for $1\leq j\leq r$ and $1\leq i\leq N$; (ii) $\max_{1\leq i\leq N,1\leq t\leq T}E\left\vert \Delta X_{i,t}\right\vert ^{4}\leq C_{0}$; (iii) $ E\max_{1\leq \widetilde{t}\leq T}\left\vert \sum_{t=1}^{\widetilde{t}}\Delta X_{h,t}\Delta X_{j,t}-E\left( \Delta X_{h,t}\Delta X_{j,t}\right) \right\vert ^{2}\leq C_{0}$; (iv) (a) $T^{-1}\sum_{t=1}^{T}E\left( \Delta \mathcal F_{t}\Delta \mathcal F_{t}^{\prime }\right) $ is a positive definite matrix; (b) the largest eigenvalue of $T^{-1}\sum_{t=1}^{T}E\left( \Delta u_{t}\Delta u_{t}^{\prime }\right) $ is finite; (c) $T^{-1}\sum_{t=1}^{T}E\left( \Delta u_{t}\Delta u_{t}^{\prime }\right) $ is a positive definite matrix.

Assumption (ref) is the same as Assumptions 1-3 in trapani17, and we refer to that paper for examples in which the assumption is satisfied. Essentially, these are the same examples that hold for $e_t$ in Assumption (ref).

Estimating the number of common factors

We now present the algorithm to estimate the dimension of the factor space. We begin by determining the presence or not of a common linear trend by estimating $r_{1}$ based on $\nu _{1}^{\left( p\right) }$, and then we determine the presence or not of zero-mean $I(1)$ common factors by estimating $r^*$, based on $\nu _{2}^{\left( p\right) }$.

Preliminary definitions

Consider the notation $\beta =\frac{\ln N}{\ln T}$, and define

equation[equation omitted — 195 chars of source]

The role played by $\delta$ is the following. In view of Theorem (ref), the largest eigenvalues are (modulo some slowly varying functions) proportional to $N$; the others, to $NT^{-1/2}$. When premultiplying eigenvalues by $N^{-\delta }$, the former will be proportional to $N^{1-\delta }$, thereby still diverging; the latter will be proportional to $N^{1-\delta }T^{-1/2}$, which, by construction, will drift to zero. Note that (ref) provides a general rule to set $\delta$, and we discuss specific choices in Section (ref).

In order to construct our test statistics, we make use the eigenvalues of the matrix

equation[equation omitted — 83 chars of source]

which with the same notation as before, are denoted as $\nu _{3}^{\left( p\right) }$ in decreasing order. In particular, when running our procedure for the $p$-th largest eigenvalues of $\Sigma_1$ or $\Sigma_2$, we will extensively use the quantities

equation[equation omitted — 138 chars of source]

for different values of $k$. Essentially, ${\overline{\nu }}_{3,p}(k)$ is the average of all (or some) eigenvalues of $\Sigma_3$ and will be employed in order to rescale the estimated eigenvalues, so as to render all our test statistics scale invariant. In the numerical analysis of Section (ref), we consider rescaling schemes with $k=1$, $k=p$, or $k=(p+1)$ and we discuss the impact of these choices on our results. For simplicity in the rest of this section we do not make explicit the dependence of (ref) on $k$. Finally, note the division by $4$ in (ref), which is done, heuristically, since it is possible that $\Delta X_{i,t}$ could inflate the variance by over-differencing, and the factor $4$ represents the largest inflation factor possible.

Determining the presence of factors with linear trends

Consider first $\Sigma_1$ defined in (ref), and its eigenvalues $\nu _{1}^{\left( p\right) }$. Based on ((ref))-( (ref)), the first $r_{1}$ eigenvalues of $\Sigma_1$ should diverge to positive infinity, as $\min\left( N,T\right) \rightarrow \infty $, at a faster rate than the $(N-r_1)$ remaining ones. Thus, the cornerstone of the algorithm to determine $ r_{1}$ is based on checking whether $\nu _{1}^{\left( p\right) }$ diverges sufficiently fast. In particular, as suggested by Theorem (ref), we want to construct a test for

equation[equation omitted — 245 chars of source]

for some positive bounded constant $C_p$. Thus, given $r_1$ we have that $H_{0,1}^{\left( p\right) }$ holds true for $p\le r_1$, while $H_{A,1}^{\left( p\right) }$ holds true for $p> r_1$.

Consider the following transformation of $\nu _{1}^{\left( p\right) }$

equation[equation omitted — 154 chars of source]

Then, based on (ref), equations ((ref)) and ((ref)), and given the definition (ref) of $\delta$, we have that

equation*[equation* omitted — 340 chars of source]

In principle, we could then use $\phi _{1}^{\left( p\right) }$ to test $H_{0,1}^{\left( p\right) }$. However, since $\phi _{1}^{\left( p\right) }$ either diverges to infinity or not, it does not have any randomness. Therefore, we propose to use the following randomisation algorithm - note that other randomisations schemes would also be possible, in principle; the one we propose, however, has been often considered in this type of literature (see e.g. corradi2006, and trapani17).

description• Step A1.1. Generate an i.i.d. sample $\left\{ \xi _{1,j}^{\left( p\right) }\right\} _{j=1}^{R_{1}}$ from a common distribution $G_{1}$, independently across $p$. • Step A1.2. For any $u$ drawn from a distribution $F_{1}\left( u\right) $, define, for $1\le j\le R_1$, $$ \zeta _{1,j}^{\left( p\right) }\left( u\right) =I\left[ \phi _{1}^{\left(p\right) }\times \xi _{1,j}^{\left( p\right) }\leq u\right] $$\underline{Step A1.3.} Compute \begin{equation*} \vartheta _{1}^{\left( p\right) }\left( u\right) =\frac{1}{\sqrt{R_{1}}} \sum_{j=1}^{R_{1}}\frac{\zeta _{1,j}^{\left( p\right) }\left( u\right) -G_{1}\left( 0\right) }{\sqrt{G_{1}\left( 0\right) \left[ 1-G_{1}\left( 0\right) \right] }}. \end{equation*} • \textit{\underline{Step A1.4.}} Compute \begin{equation*} \Theta _{1}^{\left( p\right) }=\int_{-\infty }^{+\infty }\left\vert \vartheta _{1}^{\left( p\right) }\left( u\right) \right\vert ^{2}dF_{1}\left( u\right) . \end{equation*}

The intuition for considering this approach is the following. Under the null, we know that $\phi _{1}^{\left( p\right) }$ diverges; thus, we can expect $\zeta _{1,j}^{\left( p\right) }\left( u\right) $ to be an i.i.d. Bernoulli sequence with expected value exactly equal to $G_{1}\left( 0\right) $, and variance $G_{1}\left( 0\right) \left[ 1-G_{1}\left( 0\right) \right] $. In such case, a CLT should ensure that $\vartheta _{1}^{\left( p\right) }\left( u\right) $ follows a Normal distribution, and consequently $\Theta _{1}^{\left( p\right) }$ should be expected to follow a Chi-squared distribution. By the same token, under the alternative $\phi _{1}^{\left( p\right) }$ is finite, and therefore $\zeta _{1,j}^{\left( p\right) }\left( u\right) $ should be an i.i.d. Bernoulli sequence with expected value different from $ G_{1}\left( 0\right) $; thus, $\vartheta _{1}^{\left( p\right) }\left( u\right) $ should diverge as fast as $\sqrt{R_{1}}$ by the LLN, and consequently $\Theta _{1}^{\left( p\right) }$ should also diverge at a rate $ R_{1}$. The random variable $\Theta _{1}^{\left( p\right) }$ is then the statistic that we are going to use.

In order to derive the asymptotic behavior of $\Theta _{1}^{\left( p\right) }$, we need some regularity conditions on the distributions $G_{1}(\cdot)$ and $F_{1}(\cdot)$ - see Section (ref) for a choice of these functions and of $R_1$.

assumptionIt holds that: (i) (a) $G_{1}(\cdot)$ has a bounded density function; (b) $G_{1}\left( 0\right) \neq 0$ and $G_{1}\left( 0\right) \neq 1$; (ii) $\int_{-\infty}^{\infty} u^{2}dF_{1}\left( u\right) <\infty $.

Let $P^{\ast }$ denote the conditional probability with respect to \{$X_{i,t}, 1 \leq t \leq T, 1 \leq i \leq N$\} ; we use the notation \textquotedblleft $\overset{D^{\ast }}{\rightarrow }$ \textquotedblright\ and \textquotedblleft $\overset{P^{\ast }}{\rightarrow }$ \textquotedblright\ to define, respectively, conditional convergence in distribution and in probability according to $P^{\ast }$. It holds that

theoremConsider $H_{0,1}^{\left( p\right) }$ and $H_{A,1}^{\left( p\right) }$ defined in ((ref)). Under Assumptions (ref)-(ref), if \begin{equation} \lim_{\min \left(N,R_{1}\right) \rightarrow \infty }\sqrt{R_{1}}\exp \left\{ -N^{1-\delta }\right\} =0, \end{equation} then, for almost all realisations of $\left\{ e_{t},u_{i,t},1\leq i\leq N,1\leq t\leq T\right\} $ and for all $p$, as $\min \left( N,T,R_{1}\right) \rightarrow \infty$, under $H_{0,1}^{\left( p\right) }$ it holds that \begin{eqnarray} &&\Theta _{1}^{\left( p\right) }\overset{D^{\ast }}{\rightarrow }\chi _{1}^{2}, \end{eqnarray} and under $H_{A,1}^{\left( p\right) }$ it holds that \begin{eqnarray} &&\frac{1}{R_{1}}\, \frac{\int_{-\infty }^{\infty }\left[ G_{1}\left( u\right) -G_{1}\left( 0\right) \right] ^{2}dF_{1}\left( u\right) }{G_{1}\left( 0\right) \left[ 1-G_{1}\left( 0\right) \right] }\,\Theta _{1}^{\left( p\right) }\overset{P^{\ast }}{\rightarrow }1. \end{eqnarray}

The determination of $r_{1}$ follows from an algorithm which is based on a single step.

description• Step T1.1. Set $p=1$ and run the test for $H_{0,1}^{(1)}:\nu_{1}^{\left( 1\right) }=\infty $ based on $\Theta _{1}^{\left( 1\right) }$. If the null is rejected, set $ \widehat{r}_{1}=0$\ and stop, otherwise set $\widehat{r}_{1}=1$.

The output of this step is $\widehat{r}_{1}$, which is an estimate of $r_{1}$. As discussed above, $r_{1}$ can be either $0$ or $1$, whence the test being stopped at $p=2$. The procedure based on the single Step T1.1 can therefore be viewed as a test for the presence of a common factor with a linear trend.

As can be expected, in order to ensure that $\widehat{r}_{1}$ is consistent, a pivotal role is played by the level of the test, $\alpha _{1}:=P^*( \Theta _{1}^{\left( p\right) }>c_{\alpha ,1})$, through the relevant critical value denoted as $c_{\alpha ,1}$.

lemmaUnder the assumptions of Theorem (ref), as $\min \left( N,T,R_{1}\right) \rightarrow \infty $, if $c_{\alpha ,1}\rightarrow \infty $ with $c_{\alpha ,1}=o\left( R_{1}\right) $, then it holds that $P^*\left( \widehat{r}_{1}=r_{1}\right) =1$, for almost all realisations of $\left\{ e_{t},u_{i,t},1\leq i\leq N,1\leq t\leq T\right\} $.

Requiring that $c_{\alpha ,1}\rightarrow \infty $ is necessary in order to have asymptotically zero Type I error probability, which ensures the consistency result in the lemma; an immediate implication of $c_{\alpha ,1}\rightarrow \infty $ is that the level of the test is such that

equation[equation omitted — 150 chars of source]

The fact that $c_{\alpha ,1}$ diverges has also an interesting consequence on the interpretation of the outcome of our testing procedure. It is well-known that randomised tests will yield different results for different researchers when applied to the same data, since the added randomness does not vanish asymptotically. However, this is not the case with our procedure, since, when $c_{\alpha ,1}\rightarrow \infty $, (ref) holds under $H_{0,1}^{(p)}$. Further, we show in the proof that having $c_{\alpha ,1}=o\left( R_{1}\right) $ affords that the probability of a Type II error is asymptotically zero, thus ensuring consistency.

Looking at this from a different angle, the results in Lemma (ref) are guaranteed by letting the level of the test $\alpha_1\to 0$ as $\min \left( N,T,R_{1}\right) \rightarrow \infty $ and we refer to Section (ref) for the choice of $\alpha_1$.

Determining the number of non-stationary common factors

Consider the matrix $\Sigma_2$ defined in (ref) and its eigenvalues $\nu _{2}^{\left( p\right) }$. Based on Theorem (ref), the $r^{\ast }$ largest eigenvalues of $\Sigma_2$ should diverge to positive infinity, as $\min\left( N,T\right) \rightarrow \infty $, at a faster rate than the $(N-r^*)$ remaining ones. Therefore, we can construct a the test for

equation[equation omitted — 262 chars of source]

for some positive bounded constant $C_p$. Thus, given $r^*$ we have that $H_{0,2}^{\left( p\right) }$ holds true for $p\le r^*$, while $H_{A,2}^{\left( p\right) }$ holds true for $p> r^*$.

We exploit this fact, as in the above, by considering the following transformation of $\nu _{2}^{\left( p\right) }$

equation[equation omitted — 179 chars of source]

which is very similar to ((ref)) except for the presence of the logarithmic term, which is a consequence of ((ref)). Then, based on (ref), equations ((ref)) and ((ref)), and given the definition (ref) of $\delta$, we have that

equation*[equation* omitted — 337 chars of source]

We consider the following randomisation procedure.

description• Step A2.1 Generate an i.i.d. sample $\left\{ \xi _{2,j}^{\left( p\right) }\right\} _{j=1}^{R_{2}}$ from a common distribution $G_{2}$, independently across $p$ and of $\left\{ \xi_{2,j}^{\left( p^{\prime }\right) }\right\} _{j=1}^{R_{2}}$ for all $p^{\prime } \neq p$. • Step A2.2 For any $u$ drawn from a distribution $F_{2}\left( u\right) $, define, for $1\le j\le R_2$, \begin{equation*} \zeta _{2,j}^{\left( p\right) }\left( u\right) =I\left[ \phi _{2}^{\left( p\right) }\times \xi _{2,j}^{\left( p\right) }\leq u\right] . \end{equation*} • \underline{Step A2.3.} Compute \begin{equation*} \vartheta _{2}^{\left( p\right) }\left( u\right) =\frac{1}{\sqrt{R_{2}}} \sum_{j=1}^{R_{2}}\frac{\zeta _{2,j}^{\left( p\right) }\left( u\right) -G_{2}\left( 0\right) }{\sqrt{G_{2}\left( 0\right) \left[ 1-G_{2}\left( 0\right) \right] }}. \end{equation*} • \textit{\underline{Step A2.4.}} Compute \begin{equation*} \Theta _{2}^{\left( p\right) }=\int_{-\infty }^{+\infty }\left\vert \vartheta _{2}^{\left( p\right) }\left( u\right) \right\vert ^{2}dF_{2}\left( u\right) . \end{equation*}

The same comments as in the previous algorithm apply: in essence, the procedure exploits the fact that under the null and the alternative, $\phi _{2}^{\left( p\right) }$ diverges or drifts to zero respectively: the former feature ensures (asymptotic) normality of $ \vartheta _{2}^{\left( p\right) }\left( u\right) $, whereas the latter entails that $\vartheta _{2}^{\left( p\right) }\left( u\right) $ diverges under the alternative.

assumptionIt holds that: (i) (a) $G_{2}$ has a bounded density function; (b) $G_{2}\left( 0\right) \neq 0$ and $G_{2}\left( 0\right) \neq 1$; (ii) (a) $\int_{-\infty}^{\infty} u^{2}dF_{2}\left( u\right) <\infty $.

It holds that

theoremConsider $H_{0,2}^{\left( p\right) }$ and $H_{A,2}^{\left( p\right) }$ defined in ((ref)). Under Assumptions (ref)-(ref) and (ref), if \begin{equation} \lim_{\min \left( N,R_{2}\right ) \rightarrow \infty }\sqrt{R_{2}}\exp \left\{ -N^{1-\delta }\right\} =0, \end{equation} then, for almost all realisations of $\left\{ e_{t},u_{i,t},1\leq i\leq N,1\leq t\leq T\right\} $ and for all $p$, as $\min \left( N,T,R_{2}\right) \rightarrow \infty$, under $H_{0,2}^{\left( p\right) }$ it holds that \begin{eqnarray} &&\Theta _{2}^{\left( p\right) }\overset{D^{\ast }}{\rightarrow }\chi _{1}^{2}, \end{eqnarray} and under $H_{A,2}^{\left( p\right) }$ it holds that \begin{eqnarray} &&\frac{1}{R_{2}}\,\frac{\int_{-\infty }^{\infty }\left( G_{2}\left( u\right) -G_{2}\left( 0\right) \right) ^{2}dF_{1}\left( u\right) }{G_{2}\left( 0\right) \left( 1-G_{2}\left( 0\right) \right) }\,\Theta _{2}^{\left( p\right) }\overset{P^{\ast }}{\rightarrow }1 under H_{1}^{\left( 2\right) }. \end{eqnarray}

Note that, conditionally on the sample, the sequence $\left\{\Theta _{2}^{\left( p\right) }\right\}_{p=1}^N$ is independent across $p$. We recommend the following algorithm for the determination of $r^{\ast }$.

description• Step T2.1. Run the test for $H_{0,2}^{(1)}:\nu _{2}^{\left( 1\right) }=\infty $ based on $\Theta _{2}^{\left( 1\right) }$. If the null is rejected, set $ \widehat{r}^{\ast }=0$\ and stop, otherwise go to the next step. • Step T2.2. Starting from $p=1$, run the test for $H_{0,2}^{(p+1)}:\nu _{2}^{\left( p+1\right) }=\infty $ based on $\Theta _{2}^{\left( p+1\right) }$, constructed using an artificial sample $\left\{ \xi _{2,j}^{\left( p+1\right) }\right\} _{j=1}^{R_{2}}$ generated independently of $\left\{ \xi _{2,j}^{\left( 1\right) }\right\} _{j=1}^{R_{2}}$, ..., $\left\{ \xi _{2,j}^{\left( p\right) }\right\} _{j=1}^{R_{2}}$. If the null is rejected, set $\widehat{r}^{\ast }=p$\ and stop; otherwise repeat the step until the null is rejected (or until a pre-specified maximum number of factors, say $r_{\max }^{\ast }$, is reached).

As can be expected, in this context a pivotal role is played by the level of the individual tests, which should be chosen so that $\widehat{r}^{\ast }$ is a good approximation of $r^{\ast }$, at least asymptotically. Similarly to the previous case, let $c_{\alpha ,2}$ denote the critical value of the test at each step.

lemmaUnder the assumptions of Theorem (ref), as $\min \left(N,T,R_{2}\right) \rightarrow \infty $, if $r^*_{\max }\geq r^*$ and $c_{\alpha ,2}\rightarrow \infty $ with $c_{\alpha ,2}=o\left( R_{2}\right) $, then it holds that $P\left( \widehat{r}^{\ast }=r^{\ast }\right) =1$ for almost all realisations of $\left\{ e_{t},u_{i,t},1\leq i\leq N,1\leq t\leq T\right\} $.

This lemma has the same interpretation - especially when it comes to the condition that $c_{\alpha ,2}\rightarrow \infty $ - as Lemma (ref).

Determining the number of zero-mean $I(1)$ and $I(0)$ factors

After estimating $r^{\ast }$, it is possible to estimate the number of common, zero-mean $I\left( 1\right) $ factors by subtracting the number of those with a linear trend from the total number of non-stationary factors, i.e. as $(\widehat{r}^{\ast }-\widehat{r}_{1})$. Under the conditions of Lemmas (ref) and (ref), it is immediate to verify that

equation*[equation* omitted — 111 chars of source]

As a final remark, on the grounds of Assumption (ref) it is possible to use the algorithm proposed in trapani17 to estimate the total number of common factors. The algorithm - based on first-differenced data - uses the eigenvalues $\nu_3^{(p)}$ of $\Sigma_3$ defined in (ref) in a similar way to the algorithms above. Denoting the estimate of the total number of factors as $\widehat{r}$, the number of common $I(0)$ factors can be estimated as $\widehat{r}-\widehat{r}^{\ast }$. Under the conditions in trapani17 and of Lemma (ref) above, it follows that

equation*[equation* omitted — 125 chars of source]

Determining the presence of weak factors

By Assumption (ref), all the common factors are assumed to be strongly pervasive. This is a direct consequence of having $\left\Vert \Lambda \right\Vert ^{2}=O\left( N\right) $. It is however possible to imagine a situation in which some of the common factors are \textquotedblleft weak\textquotedblright , or \textquotedblleft less pervasive\textquotedblright : this can arise from e.g. having genuinely weak factors, or from having strong factors which impact only on a small number of units - see, for example, onatski12 and the references therein.

In this section, we report some heuristic arguments (similar to trapani17), on the ability of our procedure to determine weak factors. For the sake of a concise discussion, but with no loss of generality, we consider the case where all $r$ factors are zero-mean $I(1)$, and $\Lambda ^{\prime }\Lambda $ is diagonal, with diagonal elements $ c_{p}(N)$ given by

equation*[equation* omitted — 166 chars of source]

Allowing for $\kappa _{p}\in \left( 0,1\right) $ corresponds to the case of having weak factors, and the larger $\kappa _{p}$ the weaker the corresponding factor. Suppose that the researcher is using $\Sigma _{2}$ and its eigenvalues $\nu_2^{(p)}$ in order to determine $r$. Repeating exactly the same arguments in the proof of Theorem (ref), it can be shown that

equation[equation omitted — 96 chars of source]

Equation ((ref)) entails that, whenever $p^{\prime }<p\leq r$,

equation[equation omitted — 105 chars of source]

Recall that, our procedure, essentially, is based on testing whether, as $\min \left (N,T\right) \rightarrow \infty$

equation*[equation* omitted — 268 chars of source]

with $\delta $ selected as per ((ref)). Thus, based on ((ref) ), weak factors can be determined if

equation*[equation* omitted — 113 chars of source]

which requires

equation[equation omitted — 55 chars of source]

On the grounds of ((ref)), the constraint in ((ref)) explains up to which extent weak factors can be detected. When $\beta \leq \frac{1}{2}$, that is $\frac{N}{\sqrt T}=O(1)$, then $\delta =0$, and we need $\kappa _{p}<1$. This entails that, when $N$ is much smaller than $T$, our procedure is able to detect even very weak factors. Conversely, when $\beta >\frac{1}{2}$, that is $\frac{\sqrt T}{N}=o(1)$, it is required that $\kappa _{p}<1-\frac{1}{2\beta }$: as $\beta $ increases, i.e. $N$ increases, the test is less and less able to detect weak factors. Note that when $N$ and $T$ have the same order of magnitude, and thus $\beta =1$, weak factors can be detected as long as $\kappa _{p}<\frac{1}{2}$ - that is, when the eigenvalues associated with that factor diverge to infinity a bit faster than $\sqrt N$.

Monte Carlo and empirical evidence

In our experiments, we use data generated as

align[align omitted — 702 chars of source]

The loadings in (ref) are simulated such that each entry is distributed as $\mathcal N(0,1)$ and such that the matrix $\Lambda$ satisfies the normalization constraint $\Lambda'\Lambda=N I_r$. In (ref) and (ref), we use $\rho_j\sim U[0, \bar \rho]$ with $\bar \rho\in\{0,0.4,0.8\}$, and $\alpha_j\sim U[-0.5,0.5]$ respectively. The vector $\epsilon_t=(\epsilon_{t}^{(1)}\epsilon_{1,t}^{(2)} \ldots \epsilon_{r_2,t}^{(2)}\,\epsilon_{1,t}^{(3)}\ldots\epsilon_{r_3,t}^{(3)})$ is simulated from $\mathcal N(0,\Gamma)$ independently at each $t$, with $\Gamma$ diagonal and such that \[ \frac {1}{NT}\sum_{i=1}^N\sum_{t=1}^T(\lambda_i^{(1)}\Delta f_t^{(1)})^2=\frac {1}{NT}\sum_{i=1}^N\sum_{t=1}^T(\lambda_i^{(2)'}\Delta f_t^{(2)})^2=\frac {1}{NT}\sum_{i=1}^N\sum_{t=1}^T(\lambda_i^{(3)'} f_t^{(3)})^2, \] so that in first differences each factor component has, on average, the same weight. In (ref) we allow both for serial and cross-sectional dependence in the idiosyncratic errors and for all $1\le i\le N$. We fix $a_i=0.5$, $b_i=0.5$ and $C_i=\min\left(\left\lfloor \frac N{20}\right\rfloor,10\right)$, and the errors $v_{i,t}$ are simulated from $\mathcal N(0,1)$. Note that this model for the idiosyncratic component is the same as in ahnhorenstein13. Last, we set the noise-to-signal as \[ \theta = 0.5\,\frac{\sum_{i=1}^N\sum_{t=1}^T(\lambda_i^{(1)}\Delta f_t^{(1)}+\lambda_i^{(2)'}\Delta f_t^{(2)}+\lambda_i^{(3)'}\Delta f_t^{(3)})^2}{\sum_{i=1}^N\sum_{t=1}^T(\Delta u_{i,t})^2}. \]

We consider the following cases:

enumerate• we fix $r_3=0$ and we let $r_1\in \{0,1\}$ and $r_2\in\{0,1,2\}$ and we use the test based on $\phi _{1}^{\left( p\right) }$ to compute $\widehat{r}_1$ (see Table (ref)); • we fix $r_1=0$ and we let $r_2\in\{0,1,2\}$ and $r_3\in\{0,1,2\}$ and we use the test based on $\phi _{2}^{\left( p\right) }$ to compute $\widehat{r}^*=\widehat{r}_2$ (see Table (ref)); • we fix $r_1=1$ and we let $r_2\in\{0,1,2\}$ and $r_3\in\{0,1,2\}$ and we use the test based on $\phi _{2}^{\left( p\right) }$ to compute $\widehat{r}^*=\widehat{r}_2+1$ (see Table (ref)).

For each case, we set $N\in\{50,100,200\}$ and $T\in\{100,200,500\}$, and we simulate model (ref)-(ref) 500 times, reporting the average value of $\widehat{r}_1$ or $\widehat{r}_2=(\widehat{r}^*-r_1)$, across simulations. Moreover, when computing $\widehat{r}_2$ we compare our results with the Information Criteria by bai04, denoted as $IC$ - this corresponds to $IC3$ in the original paper; we note that the other criteria, known as $IC_1$ and $IC_2$, deliver a similar (or worse) performance and are therefore not reported.

Our tests are run as follows. When computing $\phi_1^{(p)}$ and $\phi_2^{(p)}$, we rescale the $p$-th eigenvalue as (see (ref)) \[ \frac{\nu_i^{(p)}}{\bar{\nu}_{3,p}(k)}=\frac{\nu_i^{(p)}}{\frac 1{4(N-k+1)}\sum^{N}_{h= k} \nu_3^{(h)}}, \qquad i=1,2. \] For a given $p$, we consider three different rescaling schemes corresponding to three different choices for $k$:

itemize• when $k=1$, i.e. $\bar{\nu}_{3,p}(k) = \frac 1{4N}\sum^{N}_{h= 1} \nu_3^{(h)}$; • when $k=p$, i.e. $\bar{\nu}_{3,p}(k) = \frac 1{4(N-p+1)}\sum^{N}_{h= p} \nu_3^{(h)}$; • when $k=(p+1)$, i.e. $\bar{\nu}_{3,p}(k) = \frac 1{4(N-p)}\sum^{N}_{h= {p+1}} \nu_3^{(h)}$.

We then divide the eigenvalues by $N^{\delta}$, where (see (ref))

equation[equation omitted — 211 chars of source]

with $\delta^* = 10^{-5}$. Thence, for each $p$, in the first step of the randomisation algorithm, $\{\xi_{1,j}^{(p)}\}_{j=1}^{R_1}$ and $\{\xi_{2,j}^{(p)}\}_{j=1}^{R_2}$ are generated from a standard normal distribution, with $R_1 = N$ and $R_2 = N$, if $p=1$ or $R_2=\left\lfloor \frac N 3\right\rfloor$, for $p>1$. In the second step of the randomisation algorithm, we set $u=\pm \sqrt 2$. In the Appendix, we provide an analysis of our results when varying $R_1$, $R_2$, $\delta^*$ and $u$, showing that results are robust to these specifications. All tests are carried out at a significance level $\alpha_1=\alpha_2=\frac{0.05}{\min(N,T)}$, which corresponds to critical values growing logarithmically with $N$ or $T$, hence satisfying the conditions in Lemmas (ref) and (ref).

To save space, here we report only the results when, in (ref), we set $\bar \rho=0.4$ - results for $\bar \rho=0$ and $\bar \rho=0.8$ are in the Appendix. As an overall comment, results are in general unaffected but for two cases. The first case is when $r_1=0$ and $r_2=1$ and we compute $\widehat{r}_1$; in this case, we find that lower values of $\bar \rho$ improve the results. Conversely, higher values of $\bar \rho$ make the innovations of the zero-mean $I(1)$ factors more persistent, thus making the associated eigenvalues larger: in this case, we are therefore more likely to falsely detect trends. The second case arises when $r_2=0$, and we compute $\widehat{r}_2$. In this case, we find the exact opposite. This can be explained upon noting that, for lower values of $\bar\rho$, the two $I(1)$ factors become closer to two pure random walks which are highly collinear, thus making the second eigenvalue $\nu_2^{(2)}$ much smaller than the first one $\nu_2^{(1)}$: thus, in this case, we are less likely to detect the second factor. For the same reason, higher values of $\bar\rho$ make the two factors less collinear, so that then the second factor is detected more easily.

table[table omitted — 1,593 chars of source]
table[table omitted — 2,400 chars of source]
table[table omitted — 2,400 chars of source]

The tables lend themselves to drawing some general conclusions about the main features of our methodology. First, $BT1$ and $BT2$ are usually very good at finding no common factors - whether with a linear trend or genuinely $I\left( 1\right) $ with zero mean - when there are no common factors (see Tables (ref) and (ref)), which is also consistent with the results in trapani17. In the case of detecting the presence of genuine zero-mean $I(1)$ common factors, these results can be compared with the ones obtained using $IC$ which invariably finds one common $I(1)$ factor even when such factors are not present (see Table (ref)). Few exceptions are found in Table (ref), in the case where there is no common factor with a linear trend but there is one zero-mean $I(1)$ common factor. Even in this case, both $BT1$ and $BT2$ work extremely well as $T$ increases. Note that $BT3$ also works very well in this case, at least when no common factors with linear trends are present. Conversely, when there is one common factor with a linear trend, $BT3$ tends to overestimated more when one $I(1)$ factor is present, even when $T$ is large. Second, in general our criteria tend to understate, albeit slightly, as opposed to overstate the true number of common factors, this is particularly true when estimating the number of zero-mean $I(1)$ factors in presence of linear trends (see Table (ref)). In any case, the bias is of the same order as the bias of $IC$ and tends to vanish as $T$ increases. Overall, the performance of all our criteria improves dramatically as $T$ increases: although results are usually good whenever $T=100$, they markedly improve when $T\geq 200$ for all cases considered. The impact of $N$ is, in general, less clear.

An empirical investigation of the dimensions of the yield curve

In this section, we illustrate our methodology through an application to the High Quality Market (HQM) Corporate Bond Yield Curve, available from the Federal Reserve Economic Data (FRED)\footnote{\tt{https://fred.stlouisfed.org}.} - details on the construction of the yield curves are available from the US Department of Treasury.\footnote{\tt{https://www.treasury.gov/resource-center/economic-policy/corp-bond-yie}.} We use monthly data on HQM Corporate Bonds with maturities from 6 months up to 100 years ($N=196$), and spanning the period from January 1985 to September 2017 ($T=393$). The data are shown in Figure (ref), which shows evidence of non-stationarity and co-movements both cross-sectionally and across time.

figure[figure omitted — 171 chars of source]

We use the same settings as in Section (ref). In particular, when computing $\widehat{r}_1$, we set $R_1 = N$, while for $\widehat{r}^*$ we set $R_2 = N$ if $p=1$ and $R_2=\lfloor N/3\rfloor$ for $p>1$. The significance level is $\frac{0.05}{\min(N,T)}=0.0002551$. Finally, we note that when computing $\widehat r$, $BT1$ is equivalent to the test by trapani17.

Results are in Table (ref), where we have reported our three criteria, and, as a term of comparison, the information criterion $IC$ - when computing $\widehat r$, this is equivalent to $IC3$ in baing02. Based on our findings, there is borderline evidence of a common factor with a linear trend - indeed, this is picked up by $BT3$. Note that $BT3$, in view of our simulations, may have a tendency to overstate the presence of a common factor with a linear trend in small samples when there is one $I$(1), zero mean common factor. However, in our case the sample sizes are sufficiently large, and there is clear evidence of having several common factors, which suggests that $BT3$ may be correct in indicating the presence of a common factor with a linear trend. As far as the other factors are concerned, both $BT2$ and $BT3$ indicate that there are zero mean $I$(1) common factors; based on the discrepancy between these two criteria, it may be argued that two of such factors may be only borderline non-stationary. This evidence is in line with the findings from $IC$; conversely, $BT1$ seems to suggest only one common factor, which is at odds with the stylised factors in this literature where, usually, at least three factors are identified. To sum up, the results in Table (ref) indicate the presence of five common factors, which we estimate as the principal components of $X_t$, using the covariance $T^{-2}\sum_{t=1}^T X_t X_t'$ and imposing the identifying constraint $\Lambda' \Lambda= N I_r$, bai04,maciejowska2010.

table[table omitted — 547 chars of source]

The estimated factors are shown in Figure (ref) (solid red lines). The first three factors appear to be non-stationary; in particular, as indicated by $BT3$, the first one does seem to be driven by a linear trend. This evidence is consistent with our findings in Table (ref) (save for $BT1$), and it implies that the first three factors are highly persistent. In Figure (ref) we report the autocorrelation of each estimated factor and the median, 5th and 95th percentiles of the autocorrelations of the idiosyncratic errors together with 95% confidence bands (dashed lines) computed as $\pm \frac{1.96}{\sqrt T}$. These results suggest that the fourth and fifth factor are nearly stationary, whilst the idiosyncratic component is clearly stationary since it shows no residual autocorrelation. The presence of common unit roots, and the stationarity of the idiosyncratic error imply cointegration, which in turn implies the factor structure in bond yields - see dungey.

figure[figure omitted — 668 chars of source]
figure[figure omitted — 736 chars of source]

Our findings can be contrasted with the stylised facts which are typically found in this literature. In particular, following nelson1987, it is common to model yield curves by means of three common factors, which are usually interpreted as the level, slope, and curvature of the yield curve in a given time period $t$ -- see for example dai2000 and diebold2006. Moreover, when considering corporate bonds it common to find additional factors beyond the classical first three -- see for example duffie1999modeling, duffie2007multi, and clopez2008.

First we analyse the first three estimated common factors. Throughout we assume that at each point in time $t$, the $N$ elements of $X_t$ are ordered according to their maturity, thus $X_{1,t}$ is the shortest maturity (6 months), while $X_{N,t}$ is the longest maturity (100 years). We compare each estimated factor with a standard proxy as specified by diebold2006b. Results are in the first three panels of Figure (ref), where we show both the estimated factors (solid red lines) and the proxies (dashed black lines). In particular, in order to identify $\widehat{\mathcal{F}}_{1,t}$, we consider the proxy $\bar X_t=N^{-1}\sum_{i=1}^NX_{i,t}$; we found that $\mbox{Corr}(\bar X_t,\widehat{\mathcal {F}}_{1,t})\simeq 1$, which strongly suggests that $\widehat{\mathcal {F}}_{1,t}$ can be viewed as the level of the curve. Turning to $\widehat{\mathcal {F}}_{2,t}$, we use, as a proxy for the slope, $dX_t = N^{-1}\sum_{i=2}^N(\ln X_{i,t}-\ln X_{i-1,t}) = N^{-1}(\ln X_{N,t}-\ln X_{1,t})$. We find that $\mbox{Corr}(dX_t,\widehat{\mathcal {F}}_{2,t})=.82$, which suggests that $\widehat{\mathcal {F}}_{2,t}$ can be interpreted as the slope of the term structure. Finally, we compare $\widehat{\mathcal F}_{3,t}$ to $d^2X_t = (N-2)^{-1}\sum_{i=2}^{N-1}( X_{i+1,t}-2 X_{i,t}+ X_{i-1,t})$ as a proxy for the curvature; we find $\mbox{Corr}(d^2X_t,\widehat{\mathcal {F}}_{3,t})=.53$, which shows some evidence that $\widehat{\mathcal F}_{3,t}$ can be interpreted as the curvature. Furthermore, according to diebold2006, the first three elements of the $i$-th row of the loadings matrix $\Lambda$ should be given by

equation[equation omitted — 189 chars of source]

for some $c>0$. To confirm this finding, in Figure (ref), we plot the estimated loadings $(\widehat{\lambda}_{i,1},\widehat{\lambda}_{i,2},\widehat{\lambda}_{i,3})$ (left panel) together with the theoretical curves in (ref) computed for $c=0.2$.

figure[figure omitted — 343 chars of source]

As far as the remaining two estimated common factors are concerned, we note that, in addition to level, slope and curvature, macroeconomic and financial factors have also been incorporated in the study of yield curves -- see for example estrella1998, ang2003, diebold2006b, duffie2007multi, and coroneo2016unspanned. We evaluate the correlation between $S_t$ - the spread between the 10 years HQM bond rate and the Federal Funds rate - and the fourth factor finding that $\mbox{Corr}(S_t,\widehat{\mathcal {F}}_{4,t})=.51$, whence we propose to interpret $\widehat{\mathcal F}_{4,t}$ as the spread factor. Also, letting $R_t$ be the yearly returns of the Standard & Poor's index, we have $\mbox{Corr}(R_t,\widehat{\mathcal {F}}_{5,t})=.30$; this seems to suggest that $\widehat{\mathcal F}_{5,t}$ may be viewed as a financial factor, or that, at a minimum, $\widehat{\mathcal F}_{5,t}$ is intimately related to the financial market.\footnote{Data for $S_t$ are available at {\tt https://fred.stlouisfed.org}.\\ Data for $R_t$ are available at {\tt http://www.econ.yale.edu/\textasciitilde shiller/data.htm}.} These results are in line with the results by duffie2007multi. In the last two panels of Figure (ref) we report the fourth an fifth estimated factors (solid red lines) and the corresponding proxies (dashed black lines).

Conclusions

In this paper, we propose a methodology to estimate the dimension of the common factor space for a given dataset $X_{i,t}$. We do not assume that the data are stationary or that they have (or not) linear trends: our procedure estimates separately the number of common factors with a linear trend (which can be only 0 or 1), the number of zero mean, $I\left( 1\right) $ common factors, and the number of zero mean, $I\left( 0\right) $ common factors.

Since estimation of these dimensions is carried out via testing (as opposed to using an information criterion or some other diagnostic), the results provide several interesting interpretations. For example, having $ r_{1}=0$ means that the data have been tested for the presence of common linear trends, and none has been found; finding $r^*=0$ indicates that the data have been tested for (the null of) non-stationarity, and have been found to be stationary; etc. Our methodology thus complements the results recently derived by zhangpangao.

Technically, our approach exploits the well-known eigenvalue separation property that characterises the covariance matrix of data with a common factor structure: essentially, the eigenvalues associated to common factors diverge to positive infinity, whereas the other ones are bounded. On top of this, we exploit the also well-known fact that linear trends, unit roots and stationary processes all imply different rates of divergence of the eigenvalues: these two facts allow us not merely to check whether there are common factors (and how many these are) but also to discriminate between those that have a trend, those that have a unit root, and the stationary ones. In this respect, our procedure is akin to the one proposed by bai04 and zyr18, although it is based on tests rather than an information criterion, and it entertains the possibility that linear trends could be present.

Several interesting issues, in the analysis of high dimensional, possibly non-stationary, time series, remain outstanding. In model ((ref)), we made no attempt to allow for the idiosyncratic components, $u_{i,t}$, to be non-stationary, thus relegating all the possible non-stationarity to the common component $\widetilde{\Lambda }\widetilde{F}_{t}$. Still, it would be worth considering the case where $u_{i,t}\sim I\left( 1\right) $ for at least some $i$, so as to be able to disentangle common and idiosyncratic sources of non-stationarity. In such a case, our procedure would not be immediately applicable, since, chiefly, ((ref)) would no longer hold. Also, by proper rescaling of the covariance matrix, our approach can be readily generalized to $I(d)$ factors with $d\ge 1$. These extensions are under current investigation by the authors.

{{\ }}