Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
149,804 characters · 21 sections · 137 citation commands
The Canonical Decomposition of Factor Models: Weak Factors are Everywhere
\onehalfspacing
Keywords: {Approximate Factor Model, Generalized Dynamic Factor Model, Weak Factors.}
Consider an $n$-dimensional, zero-mean, covariance stationary time series $(y_t^{n}: t \in \mathbb Z)$. Often in econometric applications $n$ can be very large. Notable examples are panels of macroeconomic indicators mccracken2016fred, of stock returns ait2017using, or of volatility measures barigozzi2024fnets. However, as $n$ gets large parametric modelling of $(y_t^{n})$ can become quickly unfeasible. One of the most common and successful strategies in this case, and the focus of this paper, is to model $(y_{t}^n)$ via an approximate factor structure.
There are two approaches to factor models in the high-dimensional time series literature: the approximate static factor model (chamberlain1983funds, chamberlain1983arbitrage, stock2002forecasting, bai2003inferential, among many others) and the Generalized Dynamic Factor model (GDFM) (forni2000generalized, forni2001generalized, forni2015dynamic,forni2017dynamic, barigozzi2024inferential). Both assume that $(y_{t}^n)$ is decomposed into a common component driven by few pervasive latent factors plus an idiosyncratic component having weakly dependent elements.
The relation between dynamic and static common components has often been discussed in the literature, but only under very restrictive conditions including the assumption that the dynamic and static idiosyncratic components coincide forni2009opening,bai2007determining,stock2016dynamic, doz2011two. No general result encompassing both models exists.
The main contribution of this paper is to derive a canonical representation of $(y_t^n)$ encompassing both a static and dynamic factor representation. In doing so we show that the connection between the two approaches is much more rich and complex than often assumed and it naturally introduces a third component which is dynamically common but statically idiosyncratic. We call this component the “weakly common component”. We then propose an estimator of such component and prove its consistency with rates. To highlight the importance of our results, in an empirical analysis based on the standard dataset of US macroeconomic indicators mccracken2016fred we show that: (i) the weakly common component is non-negligible for most time series, meaning that the assumption of coinciding dynamic and static idiosyncratic components is violated; and (ii) accounting for the weakly common component can lead to improved forecasts.
With no doubt, the static approach has been extremely successful in high-dimensional time series analysis. It is most commonly used because the estimation (via static principal components) produces good results in forecasting and is also straightforward to implement stock2002forecasting,stock2002macroeconomic. It has also proved a useful tool in several other applications like, e.g., factor augmented regressions bernanke2005measuring. The dynamic approach, in its most general GDFM setting, has instead still seen a more limited number of applications barigozzi2017generalized,forni2018dynamic, often due to the more complex estimation techniques required.
Despite the wider success of the static approach over the dynamic one, in this paper, we argue how factor analysis may further benefit from being aware of the difference between the dynamic and the static approach. Although in our empirical analysis weak common components seem to be important, not every panel might have a large weak common component. Still, the existence of weak common components shall always be acknowledged whenever we model a high-dimensional autocorrelated time series panel with a factor structure. In some sense sense ignoring those components factors is like forecasting an AR$(2)$ process with an AR$(1)$ model. Finally, our results pave the way to new approaches to estimation of the dynamic factor model which are also briefly discussed.
The remainder of the paper is organised as follows. First, we review our main results and compare them with existing ones (Sections (ref) and (ref)). In Section (ref) we motivate our work by means of a detailed example where we show the main implications of our results. The first main contributions of this paper are in Section (ref), where we present the new canonical decomposition encompassing static and dynamic factor models (Theorems (ref) and (ref)). In Section (ref) we provide a complete proof of consistency, with rates, for two estimators of the dynamic common component and for the other components of the canonical decomposition (Propositions (ref), (ref), and (ref)). This is our second contribution as these results are new and fill a gap in the GDFM literature. In Sections (ref) and (ref) we provide numerical studies on simulated and US macroeconomic data. In Section (ref) we conclude. Proofs and additional examples, as well as numerical results are in the Appendix.
To give more details, we briefly review the static and dynamic factor model approaches. Consider a double indexed, zero-mean, covariance stationary infinite dimensional stochastic process $(y_{it}: i \in \mathbb N, t \in \mathbb Z)=(y_{it})$, where the index $i$ denotes the cross-sectional unit and $t$ denotes time. Let $(y_t^n)$ be an $n$-dimensional sub-process of $(y_{it})$.
The approximate static factor model was firstly introduced by chamberlain1983funds, chamberlain1983arbitrage, and then extended and studied in detail by stock2002forecasting, stock2002macroeconomic, bai2002determining, bai2003inferential, and fan2013large, among many others. It is characterized by a decomposition of the form
where $(F_t)$ is a vector process of latent pervasive factors of fixed small dimension $r$, which are loaded statically by loadings $\Lambda_i$ (an $r$-dimensional row vector) into the “common component”, denoted as $(C_{it})$. The factors $(F_t)$ are also often called “static factors”. The “idiosyncratic component”, $(e_{it})$, is assumed to be weakly correlated within the cross-section. In particular, the common component can be retrieved by orthogonal projection on the space contemporaneously spanned by the latent factors (often estimated via static principal components as in bai2003inferential). It follows that, being the residual of such projection, the idiosyncratic component is naturally assumed to be contemporaneously orthogonal to the factors. Note that the claim that some authors, as e.g. bai2003inferential, allow for weak dependence is correct but only at the level of fourth-order moments while contemporaneous orthogonality is always assumed (see Remark (ref) for details).
The dynamic approach, also referred to as the {Generalised Dynamic Factor model} (GDFM), was firstly introduced by forni2000generalized and forni2001generalized. It is characterized by a decomposition of the form:
where the “common component” $(\chi_{it})$ is driven by a vector of latent pervasive factors $(\varepsilon_t)$ of fixed small dimension $q$, which, without loss of generality can always be assumed to be an orthonormal vector white noise process, while the $K_i(j)$'s are $1\times q$ vectors of square-summable coefficients. The factors $(\varepsilon_t)$ are often called “dynamic factors” or even “common shocks”. The “idiosyncratic component”, $(\xi_{it})$, is assumed to be weakly correlated over time and cross-sectionally. In this case, the common component can be retrieved by orthogonal projection on the space spanned by the latent factors and their leads and lags (estimated for example as in forni2000generalized or in forni2017dynamic). It follows that, being the residual of such projection, the common and idiosyncratic components are naturally assumed to be orthogonal at all leads and lags.
The two approaches entail conceptually different types of what is “common” and “idiosyncratic”. On the one hand, we shall distinguish between the dynamic common and the static common component, on the other hand between the dynamic idiosyncratic and the static idiosyncratic component.
According to the definition of forni2001generalized, $(\xi_{it})$ is a dynamic idiosyncratic component if for weights $c_{i}^{(n, k)}(j)$ such that $\lim_{n,k\to\infty} \sum_{i = 1}^n \sum_{j = -k}^k \{c_{i}^{(n, k)}(j)\}^2 =0$, we have
that is $(\xi_{it})$ vanishes under dynamic aggregation. As is shown in forni2001generalized, this is equivalent to the largest eigenvalue of the spectral density matrix of $(\xi_{t}^n)$ being essentially bounded as $n\to \infty$ on the frequency band $[-\pi,\pi]$. A dynamic common component, being orthogonal to $(\xi_{it})$ at all leads and lags, must then have the $q$ eigenvalues of the spectral density matrix diverging as $n\to\infty$ on the frequency band $[-\pi,\pi]$. This definition implies the existence and uniqueness of ((ref)), as well as identification of the number of factors $q$ based on the asymptotic behavior, as $n\to\infty$, of the eigenvalues of the spectral density matrix of $(y_t^n)$ forni2001generalized.
In this paper, we derive an analogous result for the static approach. We say that the double sequence $(e_{it})$ is a static idiosyncratic component (at time $t$) if for weights $c_{i}^{(n)}$ such that $\lim_{n\to\infty} \sum_{i = 1}^n \{c_{i}^{(n)}\}^2 =0$, we have
that is $(e_{it})$ vanishes under static/contemporaneous aggregation. We show that this is equivalent to the first eigenvalue of the covariance matrix of $(e_t^n)$, being bounded as $n\to \infty$ (see Theorem (ref)). A static common component (at time $t$), being contemporaneaously orthogonal to $(e_{it})$, must then have the $r$ eigenvalues of the covariance matrix diverging as $n\to\infty$. This definition implies the existence and uniqueness of ((ref)), as well as identification of the number of factors $r$ based on the asymptotic behavior, as $n\to\infty$, of the eigenvalues of the covariance matrix of $(y_t^n)$ (see Theorem (ref)).
It is then clear that there is no reason to believe that both definitions of idiosyncratic component result in the same separation between common- and idiosyncratic component. Nevertheless, this assumption (if only stated implicitly) pervades the literature on time series factor models. In fact, the two definitions do not coincide and this leads us to derive the following unique canonical decomposition encompassing both (ref) and (ref) (see Theorem (ref)):
The term $e_{it}^\chi = \chi_{it} - C_{it} = e_{it} - \xi_{it}$ is what we call the weak common component, and it is the difference between the dynamic and the static common component or, equivalently, the static and the dynamic idiosyncratic component.
A first implication of (ref) is that, while the factors $F_t$ driving the static common component are pervasive, the weak common component is driven by (potentially infinitely many) non-pervasive factors, denoted as $F_t^w$, which live in the dynamically common space, i.e., the space spanned by the dynamic factors $(\varepsilon_{t})$. This means that we also have the decomposition (see Theorem (ref)):
Following onatski2012asymptotics, we call $F_t^w$ “weak” factors. They should not be confused with pervasive “rate-weak” factors which correspond to eigenvalues of the covariance matrix of $(y_t^n)$ diverging as $n\to\infty$ at a rate $n^\alpha$ for $\alpha \in (0, 1)$, and which are considered by, e.g., demol2008forecasting, lam2012factor, uematsu2022estimation, freyaldenhoven2022factor, bai2023approximate, and fan2024can, among many others. In our framework, such rate-weak factors are contained in $F_t$ together with “strong” factors, which correspond to eigenvalues diverging at a linear rate $n$, i.e., when $\alpha=1$. The weak factors are instead associated with {non-divergent eigenvalues} of the covariance matrix of $(y_t^n)$, hence, they correspond to the case $\alpha = 0$, and, as such, they are non-pervasive and should be regarded as part of the static idiosyncratic component. This distinction between weak and rate-weak factors matches the fact that, while principal component analysis allows to consistently retrieve the space spanned by both strong and rate-weak factors, as well as their total number, $r$ bai2023approximate, in general the space spanned by weak factors cannot be consistently estimated onatski2012asymptotics.
The literature has often considered a restricted version of the GDFM and discussed is equivalence with the static approach. Namely, the GDFM is assumed to be
for some $p<\infty$ and with $\lambda_{ij}$ and $f_t$ being $1\times q$ and $q\times 1$, $G(k)$ being $q\times q$ and square-summable, and with $(\xi_{it})$ and $(\varepsilon_t)$ as in (ref). By letting $F_t=(f_t'\cdots f_{t-p}')'$ and $\Lambda_i=(\lambda_{i0}\cdots \lambda_{ip})$, the GDFM in (ref) can be written as the static factor model
where $N(k)$ is $q(p+1)\times q$ and square-summable. This is the argument adopted by forni2005generalized,forni2009opening, bai2007determining, stock2005implications,stock2011theoxford,stock2016dynamic, doz2011two,doz2012quasi, dagostoni2012comparing, and many others, to justify the use of a static approach based on (ref) instead of a fully unrestricted GDFM as in (ref).
Considering (ref) as an equivalent representation to (ref) is, however, potentially restrictive for various reasons. First of all the static idiosyncratic component is now $e_{it}=\xi_{it}$, hence, it is also dynamically idiosyncratic, and, thus, we also have that the static common component coincides with the dynamic common component $C_{it}=\chi_{it}$, i.e., there is no weak common component, $e_{it}^\chi=0$ for all $i$. As a consequence, second, the above mentioned literature assumes that $(F_t)$ and $(\xi_{it})$ are uncorrelated at all leads and lags, if not even independent. Third, the impulse response functions of the observed variables $(y_{it})$ to the common shocks $(\varepsilon_t)$ are given by $\Lambda_i \sum_{k=0}^\infty N(k)L^k$ (here $L$ denotes the lag operator), thus the only source of dynamic heterogeneity is $\Lambda_i$. Fourth, the number of pervasive static factors has to be always greater or equal than the number of common shocks $q$. Fifth, the GDFM is restricted since the dynamic common component must have a covariance of reduced rank equal to $q(p+1)$ fixed and independent of $n$, and, hence, it is typically assumed that $r=q(p+1)$. Last, the mapping holds only if we introduce $f_t$ in the GDFM (ref) and we assume it to be loaded with a finite number of lags $p$, thus, implying, in turn, a restricted impulse response function given by the convolution $(\sum_{j=0}^{p} D_i(j) L^{j}) (\sum_{k=0}^\infty G(k) L^k)$.
There is no reason for any of these restrictions to hold in practice, so none of them is imposed in this paper, which is, therefore, much more general. Moreover, even under such equivalence, things can still be much more complex than usually assumed. For example, we can even have $q=1$ but $r=0$, or even $q=1$ and $r=\infty$ (see Section (ref) and Appendix (ref)). All this is made clear in the example considered in the next section, as well as in Example (ref) in Section (ref).
The following example illustrates the role of the weak common component in finite samples. Consider the GDFM
where $|a|<1$ and $\xi_{it} \sim iidN(0, 1)$ and assumed to be uncorrelated at all leads and lags of $(\varepsilon_t)$, thus $(\xi_{it})$ is both static and dynamic idiosyncratic. Clearly, as explained in Section (ref), we can cast model ((ref)) in static form simply by defining $F_t = (f_t\; f_{t-1})'$. However, it is convenient to have a static factor model with orthonormal factors, an identifying constraint often assumed in the literature (see, e.g., bai2013principal). To this end, let $\widetilde \lambda_{i1} = \lambda_{i1} + \lambda_{i2}a$, $\widetilde \lambda_{i2} = \lambda_{i2} \sigma_\varepsilon$, $\widetilde f_{1t} = f_t$, $\widetilde f_{2t} = \sigma_\varepsilon^{-1}(f_{t-1} - a f_t)$, with $\sigma_\varepsilon^2:= (1-a^2)$. Then, we can reformulate the dynamic factor model (ref) as the static factor model
where it is easy to verify that now the factors $\widetilde F_t=(\widetilde f_{1t} \; \widetilde f_{2t} )'$ are orthonormal, i.e., $\operatorname{\mathbb E} [\widetilde F_t \widetilde F_t' ]= I_2$.
Now suppose that the common dynamic factor $f_t$ is only contemporaneously pervasive so that most, if not all, elements of $(y_t^n)$ load it, while only a small/finite subset of the $n$ series loads $f_t$ also with a lag. Such situation is verified for example by the US macroeconomic time series analyzed in Section (ref). Consequently, we can think of the dynamic loadings in (ref) being such that
Letting $\chi_t^n = (\chi_{1t}, ..., \chi_{nt})'$, it follows that only the largest eigenvalue of the covariance matrix of the common component $\Gamma_\chi^n := \operatorname{\mathbb E} [\chi_t^n (\chi_t^n)']$ diverges. Indeed, from (ref) we have $\chi_{it}= \widetilde \lambda_{i1} \widetilde f_{1t} + \widetilde \lambda_{i2}\widetilde f_{2t}$ and, since the factors $\widetilde F_t$ are orthonormal, we can apply Theorem (ref) in this paper.
Therefore, only one factor in the equivalent static representation (ref) is pervasive (either strong or rate-weak), while the second factor is weak in the sense of this paper, meaning it is not pervasive. Thus, we started from a GDFM with one factor, i.e, $q=1$, and we wrote it as a static factor model also with just one factor, i.e., $r=1$. With reference to the canonical decomposition in ((ref)) we have in this case that the static common component is $C_{it} = \widetilde\lambda_{i1} \widetilde f_{1t}$ and the weak common component is $e_{it}^\chi = \widetilde \lambda_{i2} \widetilde f_{2t}$.
We now illustrate some key implications of this elementary model ((ref)). In each of the three following exercises, for different values of $(n,T)$ we simulate $B=500$ times the considered DGP.
Throughout, we consider stochastic double sequences, i.e., a family of random variables indexed in time and cross-section: $(y_{it}: i \in \mathbb N, t \in \mathbb Z) = (y_{it})$. Such a process can also be thought of as a nested sequence of multivariate stochastic processes: $\left(y_t^n: t \in \mathbb Z \right) = (y_t^n)$, where $y_t^n = (y_{1t},..., y_{nt})'$ and $y_t^{n+1} = (y_t^{n'}, y_{n+1, t})'$ for $n \in \mathbb N \cup \{\infty\}$. In general we write $(y_t: t \in \mathbb Z)=(y_t)$ for $n = \infty$.
We make the following assumptions.
We then suppose that $(y_{it})$ has a static factor structure in ((ref)) as formulated by chamberlain1983funds and chamberlain1983arbitrage. For this let the covariance matrices of $(C_t^n)$ and of $(e_t^n)$ be $\Gamma_C^n$ and $\Gamma_e^n$, respectively, with eigenvalues $\mu_j(\Gamma_C^n)$ and $\mu_j(\Gamma_e^n)$, $j=1,\ldots, n$, sorted in decreasing order.
We shall say that a double sequence for which A(ref) holds is an $r$-Static Factor Sequence ($r$-SFS). Alternatively, and equivalently to A(ref), we may specify the pervasiveness of the factors via the loadings and we may define idiosyncraticness via bounding the cross-correlations of the $e_{it}$'s bai2003inferential. Orthonormality of $(F_t)$ can be assumed without loss of generality.
In addition, we assume that $(y_{it})$ has also a dynamic factor structure ((ref)) as introduced by forni2000generalized and forni2001generalized. For this let $f_\chi^n(\theta)$ and $f_\xi^n(\theta)$ be the spectral density matrices of $(\chi_t^n)$ and $(\xi_t^n)$, respectively, at frequency $\theta \in [-\pi, \pi]$, with eigenvalues $\mu_j(f_\chi^n(\theta))$ and $\mu_j(f_\xi^n(\theta))$, $j=1,\ldots, n$, sorted in decreasing order.
We call a double sequence for which A(ref) holds a $q$-Dynamic Factor Sequence ($q$-DFS) forni2001generalized. Note that the original formulation of the $q$-DFS is stated in terms of two-sided filters. The existence of the innovation form in Assumption A(ref) with one-sided filters is discussed in forni2015dynamic and in more general terms in gersing2024on.
To gain more insight about dynamic factor models let us revisit them from the perspective of aggregation and let us show that the dynamic common space is the space spanned by all variables that can be represented as so called “dynamic aggregates” (we refer to Appendix (ref) for a technical treatment). Let us consider weighting schemes that employ both shifts in time and cross-sectional averages, i.e., defined by means of an infinite dimensional linear two-sided row filter $\underline c^{(n,k)}(L)$ (here $L$ denotes the lag-operator), such that
Furthermore, we normalize the weights in such a way that (here $^*$ denotes the transposed complex conjugate)
Note that the coefficients of the filter still depend on $n$ and $k$ since the sequence of weights is not nested. We call the sequence of filter coefficients $(c^{(n,k)}_{i}(j) : i\in\mathbb N, j\in\mathbb Z, n,k\in\mathbb N)$ a dynamic averaging sequence (DAS) forni2001generalized. A {\it dynamic aggregate} is then a random variable $\phi_t$ and $\lim_{n,k\to\infty} \operatorname{\mathbb E}[(\underline c^{(n,k)}(L) y_t -\phi_t )^2]=0$ (see also hallin2013factor).
The Hilbert space spanned by the set of all dynamic aggregates that can be produced from $(y_{it})$ is called the dynamic aggregation space or the dynamically common space and we denote it by $\mathbb D(y)$ forni2001generalized. It is a closed subspace of the Hilbert space $\mathbb H(y) = \operatorname{\overline{\operatorname{sp}}}(y_{it}: i \in \mathbb N, t \in \mathbb Z)$, which is the time domain of $(y_{it})$.
According to forni2001generalized a double sequence $(z_{it})$, satisfying Assumption A(ref), is idiosyncratic if it vanishes under every dynamic averaging sequence (see also (ref)), which is equivalent to the first eigenvalue of the spectral density matrix of $(z_t^n)$ being essentially bounded for all $n\in \mathbb N$ as expressed in Assumption A(ref)(ii). Here we call such process dynamically idiosyncratic in order to contrast it with a statically idiosyncratic process below.
Now, if $(y_{it})$ satisfies A(ref) and A(ref) then the common component $(\chi_{it})$ from ((ref)) is the orthogonal projection on the dynamic aggregation space (see Theorem (ref) which summarizes the main results in forni2001generalized):
Here we call it the {\it dynamic common} component in order to contrast it with the static common component below. Consequently the dynamic idiosyncratic component belongs to the orthogonal complement of $\mathbb D(y)$ in $\mathbb H(y)$.
Now, the canonical decomposition in (ref) originates from the idea that we can parallel these notions of aggregation and idiosyncraticness for the static approach. Analogously to a DAS, we define a static averaging sequence (SAS) by means of a sequence of fixed weights, $(c_i^{(n)}: i\in\mathbb N, n\in\mathbb N)$, collected into an infinite dimensional row vector $c^{(n)}$ such that \[ c^{(n)} y_t = \sum_{i=1}^n c^{(n)}_{i} y_{it}, \] with $\lim_{n\to\infty} \sum_{i=1}^n \{ c_i^{(n)}\}^2 =0$. Then, a {\it static aggregate} is a random variable $w_t$ such that $\lim_{n\to\infty} \operatorname{\mathbb E}[ ( c^{(n)} y_t^n - w_t)^2]=0.$
Accordingly, we call the {\it static aggregation space} or the statically common space, the closed Hilbert space spanned by all {static aggregates} that can be produced from $(y_{it})$ at time $t$, and we denote it as $\mathbb S_t(y)$.
A double sequence $(z_{it})$, satisfying Assumption A(ref), is {\it statically idiosyncratic} if it vanishes under every static averaging sequence (see also (ref)), which is equivalent to the first eigenvalue of the covariance matrix of $(z_t^n)$ being bounded for all $n\in \mathbb N$ as expressed in Assumption A(ref)(ii) (see Theorem (ref)). It follows that every dynamically idiosyncratic double sequence is also statically idiosyncratic but not the other way around (this is a direct implication of Theorem (ref) below).
Finally, paralleling ((ref)), we characterize the {\it static common} component $(C_{it})$ in (ref) as being such that (see Theorem (ref)):
Consequently the static idiosyncratic component belongs to the orthogonal complement of $\mathbb S_t(y)$ in $\operatorname{\overline{\operatorname{sp}}}\left(y_{it}: i \in \mathbb N \right)$.
Our main result is the following encompassing canonical decomposition.
Hereafter, we call $(e_{it}^\chi)$ in ((ref)) the {\it weak common component}. We now make a list of remarks aimed at explaining the meaning of Theorem (ref).
Since the weak common component is statically idiosyncratic, it is generated by potentially infinite contemporaneously non pervasive random variables, which we shall call {\it weak factors.} Therefore, we could construct a “canonical vector representation” in terms of pervasive and weak factors. While the factors are determined up to a non-singular transformation, the separation between pervasive (contained in $C_{it}$) and weak ones (contained in $e_{it}^\chi$) is unique due to the uniqueness of the orthogonal projection in the proof of Theorem (ref).
Under Assumption A(ref) and Theorem (ref) we know that the factors $(F_t)$ have elements $F_{jt} \in \operatorname{\overline{\operatorname{sp}}} (\chi_{it}: i \in \mathbb N)$ for $j = 1, ..., r$ and $\operatorname{sp}(F_t) = \mathbb S_t(y)$. As shown below, these are pervasive factors. We use the Gram-Schmidt-orthogonalisation procedure to find the weak factors $F_t^{w,n}$ that complete $F_t$ to a basis of $\operatorname{\overline{\operatorname{sp}}} (\chi_{it}: i \in \mathbb N)$.
Start by recalling that by Assumption A(ref) $\operatorname{\mathbb E} [F_t F_t'] = I_r$ which can always be assumed without loss of generality. Choose the first $i$ in order for which $\chi_{it} - \operatorname{proj}(\chi_{it} \mid \operatorname{sp}(F_t)) \neq 0$, set this to $i_1$. Set $v_{1t} = \chi_{{i_1}, t} - \operatorname{proj}(\chi_{i_1, t} \mid \operatorname{sp}(F_t))$ and set $F_{1t}^w = \@ifstar{\oldnorm}{\oldnorm*}{v_{1t}}^{-1} v_{1t}$. Let $i_2 > i_1$ be the next $i$ in order such that $\chi_{it} - \operatorname{proj}(\chi_{it} \mid \operatorname{sp}(F_t, F_{1t}^w)) \neq 0$ and set $v_{2t} = \chi_{i_2,t} - \operatorname{proj}(\chi_{i_2, t} \mid \operatorname{sp}(F_t, F_{1t}^w))$ and $F_{2t}^{w} = \@ifstar{\oldnorm}{\oldnorm*}{v_{2t}}^{-1} v_{2t}$. In this way we obtain indices $i_1, i_2, ..., i_{r_\chi^+(n)}$ with $ r_{\chi}(n)^+ \leq n$ along with $F_t^{w,n} = (F_{1t}^w \cdots F_{r_\chi^+(n), t}^w)'$ which is such that $\operatorname{\mathbb E}[F_t^{w,n}{F_t^{w,n}}']=I_{r_\chi^+(n)}$ and also $\operatorname{\mathbb E}[F_{it}^{w,n}F_{jt}]=0$ for all $i,j,t$.
Now, set $r_\chi(n) := r + r_\chi^+(n)$ and consider the stacked $r_\chi(n)$-dimensional vector $\widetilde F_t = (F_t'\; {F_t^{w,n}}')'$ of pervasive and weak factors. Then, for every given fixed $n$, we can write the decomposition of Theorem (ref) in vector form as
with $\operatorname{\mathbb E} [\widetilde F_t \widetilde F_t']= I_{r_\chi(n)}$ by construction. Furthermore since $\Lambda^{w,n} F_t^{w, n}$ are the residual from the projection of $\chi_t^n$ on $\mathbb S_t(y) = \mathbb S_t(\chi)$, Theorem (ref) implies that \[ \sup_{n\in\mathbb N}\mu_r({\Lambda^{n}}'\Lambda^n)= \infty, \;\mbox{ but }\; \sup_{n \in \mathbb N}\mu_1 ({\Lambda^{w, n}}' \Lambda^{w, n}) < \infty. \]
We shall use the term static factor for any basis coordinate of $\operatorname{\overline{\operatorname{sp}}} (\chi_{it}: i \in \mathbb N)$. According to ((ref)), these, should be further distinguished between:
Our definition of weak factors is consistent with the definition by onatski2012asymptotics. In particular onatski2012asymptotics considers the model $y_t^n = \Lambda_w^n F_t^w + e_t^n$ with loadings such that $\sup_{n \in \mathbb N}\mu_1 ({\Lambda_w^n}'\Lambda_w^n )<\infty$ and shows that in this case the standard principal components estimator is not consistent. Therefore, it is evident that whereas the static common component can be estimated via principal components even when it is driven by rate-weak factors bai2023approximate, the weak common component cannot be consistently estimated by standard methods.
In general, the dimension $r_\chi^+(n)$ of $F_t^{w,n}$ may increase with $n$, i.e., when we add new variables in ((ref)), meaning that $r_\chi(n)\to \infty$ for $n\to \infty$. This implies that the covariance matrix of the dynamic common component does not necessarily have a fixed reduced rank, as instead often assumed in the literature forni2009opening,doz2011two. To see why this is the case consider the following examples:
Note that none of the two cases just considered can be written as a static factor model by stacking as in Section (ref), as a consequence the dynamic common component cannot be estimated via standard static principal components, but requires other estimation approaches, as those in forni2000generalized, forni2017dynamic and barigozzi2024inferential.
The next theorem clarifies the relation between the GDFM and any stacking approach.
The proof of this theorem is in Appendix (ref). The first statement says that every finite dimensional $q$-DFS , i.e., with $\sup_{n \in \mathbb N} r_\chi(n) = r_\chi < \infty$, is an $r$-SFS. This does not hold if $\sup_{n \in \mathbb N} r_\chi(n) = \infty$ as shown in Example (ref) above. The second statement clarifies that the dynamic and the static common component coincide only if all loading columns of $\chi_t^n$ correspond to diverging eigenvalues of the covariance matrix $\Gamma_\chi^n$. This is an assumption often made in the literature forni2005generalized,forni2009opening but never properly justified.
The third statement provides sufficient conditions for a representation in terms of canonical decomposition (ref). Those conditions impose structural requirements on the loadings columns and factors. For a separation between static common and weakly common, it is sufficient to find a representation of two mutually orthogonal groups of factors $x_t^1$ (pervasive) and $x_t^{2,n}$ (weak) where the $r$ loadings of the first group correspond to diverging eigenvalues and the eigenvalues of the second group are bounded. Note that, by virtue of Theorem (ref), it should be clear that the separation between the two groups $x_t^1$ and $x_t^{2,n}$ is unique, although the factors themselves are not uniquely identified within their group.
In this section, we study estimation of the canonical decomposition given in Theorem (ref). Throughout, we assume to observe $n$ times series of length $T$, i.e., the realizations $(y_{it} : i=1,\ldots, n,\, t=1,\ldots, T)$ of the $n$-dimensional process $(y_t^n : t\in\mathbb N)$ satisfying Assumption A(ref).
In the literature there have been proposed estimators for the static and for the dynamic common component, based on static and dynamic Principal Component Analysis (PCA). We could simply estimate $e_{it}^\chi$ as the difference of the estimates of $\chi_{it}$ and $C_{it}$. However, according to Theorem (ref), any such estimator must also ensure that the three estimated components $\widehat{\chi}_{it}$, $\widehat C_{it}$, and $\widehat{e}_{it}^{\,\chi}$ are contemporaneously uncorrelated so that the sample variance of $\widehat\chi_{it}$ is greater or equal than the sample variance of $\widehat C_{it}$, their difference being the variance explained by the weak common component.
A possible approach to obtain a decomposition that satisfies the desired properties is as follows.
We now review in details all estimation steps of part I and II, part III being trivial. Consistency of the proposed estimators is in Section (ref).
We have these steps when using the approach by forni2000generalized.
To apply the approach by forni2017dynamic we have the following steps.
We can proceed in two equivalent ways.
Regarding part I.a we remark that forni2000generalized provide no consistency rates, while those given in forni2004generalized for the same estimator are incomplete. For part I.b, consistency results are available in forni2017dynamic, and barigozzi2024inferential,barigozzi2024fnets. However, these work make use of similar but not identical assumptions and are based on a series of results on consistency of the estimated spectral density which are not comparable with each other. In this paper, we unify these results and give a first complete treatment of consistency with rates for estimating the dynamic common component either as in forni2000generalized or as in forni2017dynamic.
Regarding part II, we remark that the na\"ive approach which would consist in computing an estimator of $C_{it}$ via standard PCA on $y_{it}$, i.e., given by $\widetilde C_t^n = \widehat \Pi_n'\widehat \Pi_n y^n_t$, does not work in general, since it does not ensure orthogonality. Therefore, our theoretical analysis of part II must use the estimated dynamic common component $\widehat{\chi}_t^n$ obtained in part I, and thus we must also account for the related estimation error. We refer also to Corollary (ref) in Appendix (ref) for further motivations underlying our choice of estimator of $C_{it}$.
In order to prove consistency of the proposed estimators, we need to make some additional assumption. First, we strengthen Assumption A(ref) to the following assumption taken from forni2017dynamic and barigozzi2024fnets,barigozzi2024inferential.
Notice that (i) and (ii) imply the behavior of eigenvalues assumed in Assumption A(ref). For the common component this is obvious. For the idiosyncratic component we refer to forni2017dynamic. In principle independence of innovations could be relaxed to allow for martingale difference processes (see, e.g., barigozzi2024fnets, Assumption 2.4). Part (iii) is redundant since we already assumed common and idiosyncratic to be uncorrelated at all leads and lags in Assumption A(ref), but for convenience it is repeated here so that, hereafter, A(ref) replaces entirely A(ref). Part (iv) is natural as it allows for estimation of second moments.
We then assume linear divergence of eigenvalues.
Linear divergence is reasonable and assumed for convenience, we refer to the comments in barigozzi2024dynamic, for a possible justification. The results in this section can also be derived for rate-weak dynamic factors, so when assuming a sub-linear divergence rate. The proofs are almost identical so simplicity of notation we omit this case (see also the results in barigozzi2024fnets, on spectral density estimation in presence of rate-weak factors).
Then, we state standard assumptions for the kernel and its related bandwidth in (ref), as well as for the truncation levels in parts I.a and I.b.
For estimation in part I we need two more assumptions. First, for part I.a only we assume.
This assumption seems quite technical but in fact it is just assuming summability of the coefficient of the filter $\underline g_n(L)=\sum_{\ell=-\infty}^{\infty} G_n(\ell) L^\ell$, which is the population counterpart of the dynamic PCA filter $\widehat {\underline d}_n(L) = \sum_{\ell=-\infty}^\infty \widehat D_n(\ell) L^\ell$ defined in step I.a.ii in part I.a. Notice that by definition $\Vert G_{n}(0) \Vert\le 2\pi$ for all $n\in\mathbb N$.
Second, for part I.b only we assume.
Essentially, here we assume the existence of a block-diagonal autoregressive representation for the dynamic common component. This standard in GDFM literature and it is also assumed in forni2017dynamic who in turn make this assumption based on the representation results for processes with singular rational spectral density by anderson2008generalized and forni2015dynamic. We refer to those works for further details.
For estimation in part II we make the following assumption.
As shown in our proofs, this assumption contains the minimal set of conditions needed to prove consistency of PCA. Although parts of this assumption are already in Assumption A(ref), they are repeated here for convenience, so that, hereafter, A(ref) replaces entirely A(ref).
Consistency of part I.a follows.
The proofs of this proposition and the next ones are in Appendix (ref). By construction the method by forni2000generalized is two-sided thus as shown in Proposition (ref) it can give consistent estimates of $\chi_{it}$ only for $t$ in the central part of the sample, provided we choose a truncation level $\mathcal M_T=[\log T]$ in agreement with Assumption (ref). Furthermore, the first term in the rates is always dominated by the second one, since $\nu\ge 4$. Typically we use the Bartlett kernel to estimate the spectral density, for which $\vartheta=1$, and this implies that the optimal bandwidth is $\mathcal B_T=[T^{1/3}]$. Clearly, we always need $n\to\infty$ to achieve consistency since the GDFM is identified only asymptotically (see Theorem (ref)).
Similarly, we can prove consistency of part I.b.
The method by forni2017dynamic is multi-step but one-sided. Still, because the dynamic common component is estimated by inverting its estimated autoregressive representation, we need to specify a truncation lag $\mathcal K_T$ in the past, which should be large enough to provide a good approximation of the true infinite MA representation but not too large to affect consistency since we need to estimate as many coefficients as $\mathcal K_T$. A good trade-off is obtained if we choose $\mathcal K_T=[\log T]$ in agreement with Assumption (ref). To discuss the first three terms in the rates, assume that $n=T^\gamma$ for some $\gamma\in(0,\infty)$. Then, the second and third term in the rates are always dominated by the fourth one, since $\nu\ge 4$ and obviously $T> T/\mathcal B_T$. Concerning the first term, we clearly need $\nu>4$ which is stronger than what needed for Proposition (ref), and for $\nu$ large enough the first term gets also dominated by the fourth one. Under these conditions, and when using the Bartlett kernel to estimate the spectral density, for which $\vartheta=1$, the optimal bandwidth is still $\mathcal B_T=[T^{1/3}]$. Clearly, we always need $n\to\infty$ to achieve consistency since the GDFM is identified only asymptotically (see Theorem (ref)).
Finally, we prove consistency of parts II.a and II.b.
As clearly seen in the proof, the rates derived in this proposition are dominated by the estimation error of parts I.a or I.b. The rates of the first step indeed dominate the PCA rate $\min(n,\sqrt T)$ that we would have if we applied PCA on the true dynamic common component $\chi_t^n$ instead of the estimated one $\widehat{\chi}_t^n$ (see Lemma (ref)).
By combining Proposition (ref) or (ref) with Proposition (ref), we have consistency of the estimated weak common component, which we state under some additional requirements that allow us to simplify the rates as discussed above.
We simulate data with factors and loadings of DGP1 from Section (ref) and consider the estimation of the canonical decomposition for unit $i = 1$ which has a non-trivial static- and weak common component, i.e.,
For the dynamic idiosyncratic component, we consider different DGPs, based on the specification
with $\varepsilon_{it}^\xi \sim iidN(0, \sigma_{ei}^2)$ and independent of $(\varepsilon_t)$ the shocks of the factor process, with $\operatorname{\mathbb E}(\varepsilon_{it}^\xi \varepsilon_{jt}^\xi) = \tau^{\@ifstar{\oldabs}{\oldabs*}{i-j}}, i, j = 1, ..., n$, with $\tau \in \{0, 0.5\}$ if $\@ifstar{\oldabs}{\oldabs*}{i-j}\leq 10$ and $\operatorname{\mathbb E}(\varepsilon_{it}^\xi \varepsilon_{jt}^\xi) = 0$ otherwise; last $\alpha_i = \{0, \delta_i\}$ with $\delta_i \sim iidU(0, \delta)$ and $\delta \in \{0, 0.5\}$. The parameters $\tau$ and $\delta$ are crucial to control the cross-sectional and serial correlation in the dynamic idiosyncratic component, respectively.
We consider $(n, T) \in \{(30, 60), (60, 120), (120, 240), (240, 480), (480, 900)\}$ and $B = 500$ replications. At each replication we estimate the dynamic common component with $q = 1$ and $\mathcal B_T = \lfloor 0.75 \sqrt{T} \rfloor$ using the lag window estimator for the spectral density with a Bartlett kernel.
We compare estimators $\widehat\chi_{it}^{type}$ with $type$ = Ia, Ib and $\widehat C_{it}^{type}$, $\widehat e_{it}^{\chi, type} = \widehat \chi_{it}^{type} - \widehat C_{it}^{type}$ with $type$ = IaIIa, IaIIb, IbIIa and IbIIb as described in Section (ref). For $ t \in \mathcal T = \{\mathcal B_{T} + 1, ..., T - \mathcal B_T\}$ we evaluate
with analogous definitions for $\widehat \chi_{it}^{type}$ and $\widehat e_{it}^{\chi, type}$. The results are shown in Table (ref) for the case $\tau = 0.5$ and $\delta = 0.5$ and in Tables (ref), (ref), and (ref) in Appendix (ref) for other combinations of $\tau$ and $\delta$. All of the proposed estimators converge as the MSE decreases with increasing sample size. The convergence of estimates involving part I.b is slower because the procedure is multi-step and thus more complicated. The performance of estimates involving parts II.a and II.b are quite similar in performance. Therefore, we conclude that II.b shall be preferred over II.a as it yields orthogonal estimates.
We consider the FRED-MD dataset comprising monthly observations of $n=123$ time series of US macroeconomic data mccracken2016fred. The data is transformed to stationarity following the transformations recommended by the authors. Moreover, a few series are removed due to missing values, and outliers are removed and interpolated using standard procedures. The final dataset contains $T = 778$ monthly observations from 1959:1 to 2023:10. A list of the variables contained in the dataset is implicitly given in Figure (ref).
We estimate the canonical decomposition ((ref)) using the procedures from Section (ref) with the stationarity transformed and pre-processed data. Before estimation the data is standardised to have zero mean and unit sample variance. Firstly, we determine the number of dynamic and static factors. Using the method by hallin2007determining yields $q = 4$ dynamic factors and using $IC1$ and $IC3$ from the method by bai2002determining with the penalty tuning suggested by alessi2010improved yields $r = 8$ static pervasive factors, while $IC2$ yields $r = 6$ factors. Hereafter, we use $r = 8$ factors.
It is important to notice that, given the results in bai2023approximate, among the 8 considered static factors are included all possible rate-weak factors. Indeed, when looking for the number of stronger factors, by means of the method by freyaldenhoven2022factor we find 6 factors corresponding to eigenvalues diverging at rate $\sqrt n$ or faster, meaning that two factors found with $IC1$ and $IC3$ of bai2002determining method correspond to eigenvalues diverging at rate $n^\alpha$, with $\alpha\in (0,1/2)$, i.e., they are still pervasive but very weakly pervasive.
All estimates are computed as in Section (ref) using a bandwidth $\mathcal B_T = \lfloor 0.75 \sqrt{T} \rfloor = 20$ and a Bartlett kernel. As stated above, the advantage of our approach to is that it yields three components which are orthogonal also in sample meaning they have zero sample correlation. In order to assess the importance of each component for each variable, we compute the variance share as follows.
First, recall that we are working with standardized data, thus each $y_{it}$ has unit sample variance. Then, the share of variance explained by the dynamic common component is computed as:
where $\widehat{f}_\chi^n(\theta_h)$ is computed as in (ref) in step I.b.ii above. Note that the quantity $EV^\chi_i$ is the same regardless whether we use the estimator in part I.a or I.b. The share of variance explained by the static common component is computed as:
with $\widehat P_n$ the $r\times n$ matrix having as rows the normalized eigenvectors of $\widehat{\Gamma}_\chi^n$ which in turn is computed as in (ref). This, corresponds to the sample variance of $\widehat C_{it}$ when computed as in part II.b. The share of variance explained by the weak common component is then:
Notice that this by construction is always a non-negative quantity.
The results are shown in Figure (ref). Overall we see, that the weak common component accounts for a non-negligible part of the variation for most of the series. It explains more than 5% of the total variation for 90 of the 123 series. The weak common component seems to play an important role in particular for many series of the labour market sector and the money & credit sector.
On the one hand, the labour market variable CES0600000008 (Avg Hourly Earnings in the Goods-Producing sector) amounts to $33\%$ share of explained variance and has the largest weak common component according to this estimate. The variance explained by the weak common component of a few noteworthy indicators is: UEMPLT5 (Civilians Unemployed - Less Than 5 Weeks) with $16\%$, UNRATE (Civilian Unemployment Rate) with $18.5\%$, RETAILx (Retail and Food Services Sales) with $22.7\%$ or BOGMBASE (Total Monetary Base) with $24.8\%$.
On the other hand, Inflation (CPIAUCSL) with $3.4\%$ and Industrial Production (INDPRO) with $1.24\%$ have a small to negligible weak common component shares. This makes sense from a theoretical point of view. Recall that the static common component accounts for the contemporaneously common part whereas the dynamic common component is the projection on the infinite past of the dynamically common shocks. Therefore, Inflation and Industrial Production which are contemporaneous aggregates (static idiosyncratic part vanishes under averaging) themselves are mainly driven by the contemporaneously pervasive factors and also have a small dynamic idiosyncratic component. Still, more sector specific series may have potentially very heterogenous responses to the dynamically common shocks (dynamic factors) and therefore may be driven to a large extent by weak factors.
Alternatively, we also compute $EV_i^C$ as in (ref) but by using the eigenvectors of the estimator in part II.a yielding similar results (see Figure (ref) in Appendix (ref)). However, if in part II we were to use the PCA estimator of $C_{it}$ computed from the observed data via standard PCA on $y_{it}$, i.e., given by $\widetilde C_t^n = \widehat \Pi_n'\widehat \Pi_n y^n_t$, then we would often obtain negative values for $EV_i^{e^\chi}$. This is due to the fact that in this case there is no guarantee that the estimator $\widetilde C_t^n$ is orthogonal to $\widehat{\chi}_{t}^n$.
We also compute the shares of explained variance for $r = 6$ factors as recommended by $IC2$ which are given in Figure (ref) in Appendix (ref). The results are similar, but the variance share explained by the weak common component is, of course, even larger.
Finally, although we already pointed out that according to bai2023approximate when setting $r=8$ we are already likely to have included all rate-weak factors, we still would like to examine how the empirical results change if we increase $r$. For instance, one might argue that the large part of variation of the weak common component obtained above, in fact stems from an underspecification of the number of statically pervasive factors. Results for the case $r=12$ are presented in Figure (ref) in Appendix (ref) when using the estimator in part II.b. Of course, the share of variance explained by weak common component is reduced but the overall picture remains similar. Even with $r= 12$, for instance 68 of the 123 series have a weak common component share larger than $5\%$, e.g., UEMPLT5 (Civilians Unemployed - Less Than 5 Weeks) with $19.2\%$. This is consistent with the example in Section (ref), with our theory and with the results by onatski2012asymptotics: we cannot consistently recover the weak common component from contemporaneous aggregation as, e.g., PCA.
To investigate the benefits of the dynamic approach in terms of forecasting, we conduct a recursive window pseudo real time forecast evaluation for the target variables industrial production (INDPRO = IP) and Inflation (CPIAUCSL = CPI). We consider $h$-step ahead predictions for $h = 1, 6, 12$ months.
Clearly, the estimator discussed in part I.a Section (ref) and proposed by forni2000generalized requires the use of two-sided filters of the data and, therefore, it is not applicable for forecasting. However, as noted above we know that $\mathbb H_t(F) \subset \mathbb H_t(\varepsilon)$. If we assume that $F_t$ is driven by the same shocks as $\chi_{it}$, we may approximate the information of the infinite past of the dynamic common component, by considering a finite number lags of the contemporaneously pervasive factors, which in turn can be estimated via standard PCA. This is the distributed lag approach (dlreg) already discussed in Section (ref) and whose theoretical properties are studied in gersing2024distributed. Since the aim of this section is to show that by considering a fully dynamic approach we can improve over the standard diffusion index approach, we just consider the dlreg as its application is straightforward. As shown by forni2018dynamic the dlreg approach has a forecasting performance comparable to the estimator discussed in part I.b Section (ref) and proposed by forni2017dynamic.
As in mccracken2016fred we suppose that IP is I(1) (stationary after first log differences) and CPI is I(2) (stationary after 2 log differences). Following the setup in stock2002macroeconomic our target variables are then defined as
For each variable we compare forecasts of the sort
where here $\widehat F_t$ are factors estimated via static principal components from the standardised and stationary transformed data. Accordingly $p_F = 0$ and $p_y = 0$ corresponds to a regression without including factors or endogenous lags respectively. We start the recursive evaluation with 30% of the time observations in $T_0$ = 1978:5 and construct the first $h$-step ahead forecast and extend the training data until the feasible limit depending on the forecasting horizon. Throughout, we consider mean square forecasting error (MSFE) as evaluation metric. Time observations which have been previously classified as outliers are removed from the evaluation.
Autoregressive Forecasts. As benchmark models we use univariate regressions which correspond to model ((ref)) without the inclusion of factors, i.e., $p_F = 0$. The results for IP and CPI are shown in Tables (ref) and (ref). In all cases there are considerable performance gains from autoregressions compared to the na\"ive method of simply regressing on a constant. In the following, we compute the MSFE always relative to the MSFE of the best performing autoregressive model with the respective horizon.
Diffusion Index Forecasts. We consider model ((ref)) for different $r = 1,\ldots, 15$ with $p_F = 1$ and for different lag orders of the endogenous variable $p_y = 1, \ldots, 15$. This model, introduced by stock2002forecasting,stock2002macroeconomic has been widely used in the literature. Albeit stock2002macroeconomic already included higher lag orders of the factors $p_F >1$, this is rarely found in the literature. The common rationale is that stacking the factors $f_t$ as in ((ref)) results in a dynamic complete specification of the common component forni2005generalized,bai2006confidence,boivin2006more,schumacher2007forecasting,demol2008forecasting,dagostoni2012comparing,gonccalves2014bootstrapping, kotchoni2019macroeconomic,fan2023bridging. The implicit assumptions here are that $f_t$ is autoregressive and that $q p_F \leq r$, such that the dynamics of $f_t$ can be recovered from contemporaneous aggregation. We have argued above why this is in general not the case. The augmentation with lags of the endogenous variable shall account for the dynamics in the idiosyncratic component. On the other hand, IP and CPI are aggregates themselves with small to negligible idiosyncratic component (see Figure (ref)). Instead lags of the endogenous variable are also noisy proxies for the weak non-pervasive factors. Results are given in Tables (ref) and (ref).
Dynamic, Distributed Lag, Approach Forecasts. We compute forecasts of model ((ref)) without including lag of the endogenous variable, so $p_y = 0$, for different $r = 1,\ldots, 15$ and $p_F = 1, \ldots, 15$. Note that $p_F = 1$ corresponds to most commonly used model in the literature, where contemporaneously pervasive factors but no lags thereof are employed in the forecasting model. With regard to the motivating example (Section (ref)) gains in forecasting can stem potentially from two sources. In Section (ref).(ii) we argue that higher lag orders are potentially needed to obtain a dynamic complete specification of the dynamic common component. In Section (ref).(iii) we state that relatively weaker factors (even though contemporaneously pervasive) may be estimated more precisely be using lags of strong factors instead of principal components. The results are presented in Tables (ref) and (ref). In all cases the best models outperform the diffusion index approach.
We show that weak non-pervasive factors, associated with non-divergent signal eigenvalues in the sense of onatski2012asymptotics, are the general case for high-dimensional time series panels with a dynamic and a static factor structure. The encompassing decomposition stated in this paper clarifies how the static and the dynamic approach are related. We derive important implications on model interpretation, estimation and forecasting. Two new estimators for the canonical decomposition are introduced. The empirical application reveals that most series have a non-trivial weak common component in a high-dimensional time series panel of US macroeconomic data. Finally, our pseudo real-time forecasting study shows that considering the dynamic approach can be beneficial for forecasting compared the diffusion index approach mostly present in the literature where the forecasting models employ only regression on the contemporaneously pervasive factors without including lags.
On behalf of all authors, the corresponding author states that there is no conflict of interest.
{{The authors would also like to thank Paul Eisenberg, Sylvia Frühwirth-Schnatter, Tobias Hartl and Dominik Liebl for helpful comments that lead to the improvement of the paper. The authors gratefully acknowledge financial support from the Austrian Central Bank under Anniversary Grant No. 18287 and the DOC-Fellowship of the Austrian Academy of Sciences (ÖAW). }}