EconBase
← Back to paper

On the Existence of One-Sided Representations for the Generalised Dynamic Factor Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

51,234 characters · 10 sections · 34 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the Existence of One-Sided Representations for the Generalised Dynamic Factor Model

\onehalfspacing

abstractWe show that the common component of the Generalised Dynamic Factor Model (GDFM) can be represented using only current and past observations basically whenever it is purely non-deterministic.

Index terms--- Generalized Dynamic Factor Model, Representation Theory, MSC: 91B84

Introduction

There are two main approaches to approximate factor models in time series: a) the dynamic approach, i.e., the Generalised Dynamic Factor Model (GDFM) based on dynamic principal components forni2000generalized, forni2001generalized, and b) the static approach based on static principal components chamberlain1983arbitrage, chamberlain1983funds, stock2002forecasting, stock2002macroeconomic, bai2002determining. Contrary to the common view in the literature, these are fundamentally different decompositions, each imposing distinct interpretations of what is “common” and “idiosyncratic.” For an in-depth discussion, see gersing2023reconciling, gersing2024weak which was extended to the time domain by barigozzi2024dynamic.

We say that a process is causally subordinated to the data if it can be expressed purely in terms of current and past values of observed variables. In the static approach, the static common component is a linear combination of only contemporaneous observed variables, making it trivially causally subordinated. By contrast, the dynamic common component of the GDFM is, in its original form forni2001generalized, the mean-square limit of a dynamic low-rank approximation using lags and leads.

This paper shows that the use of leads is only a matter of representation: under fairly general conditions, the dynamic common component can instead be written using only current and past variables. Specifically, we prove that its innovations (one-step ahead prediction errors) remain causally subordinated to the data, provided the component is purely non-deterministic and the transfer function meets a mild condition related to causal invertibility. Unlike low rank approximations via dynamic principal components in general, this one-sidedness is a distinctive feature of the GDFM which is, as shown in this paper, implied by the special behaviour of its spectral eigenvalues.

This result highlights why, from an economic perspective, the dynamic decomposition is of primary interest compared to the static decomposition. Interpreting innovations of the dynamic common component as “common structural shocks of the economy”, the dynamic common component is the projection of observed variables onto the infinite past of those shocks lippi2021validating, forni2025common. Consequently, impulse response analysis in time series factor models should focus on how observed variables respond to these structural shocks and therefore be concerned with the dynamic common component. In contrast, the static decomposition captures only the part that is contemporaneously common.

In summary, the apparent two-sidedness of the classical GDFM is not an inherent flaw but a choice of representation. Our result reinforces the theoretical foundation of the GDFM by proving that, under mild conditions, the dynamic common component can always be represented in a causally subordinated, forecasting-relevant form. This paper is intended as a theoretical contribution: (1) to establish the interpretation of the dynamic common component as the response to common structural shocks, and (2) to justify starting future research from a one-sided representation.

The remainder of the paper is structured as follows. We begin in Section (ref) by formally stating the main result and outlining the core idea of the proof. Next, Section (ref) introduces the assumptions and notation underlying the GDFM. In Section (ref), we define purely non-deterministic processes in the infinite-dimensional, rank-deficient case and discuss aspects of causal invertibility. The main proof is presented in Section (ref): we first show in Theorem (ref) how to construct infinitely many full-rank $q\times q$ transfer-function blocks, then establish causal subordination under a strict minimum phase condition in Theorem (ref), and finally relax this condition to include cases with the same zeros on the unit circle in infinitely many rows. The paper concludes with Section (ref).

The Main Result and Idea of the Proof

To fix ideas, consider an infinite-dimensional time series as a double indexed (zero-mean, stationary) stochastic process $(y_{it}: i \in \mathbb N, t \in \mathbb Z)=(y_{it})$, indexed by cross-section $i \in \mathbb N$ and time $t \in \mathbb Z$. The GDFM decomposes

align[align omitted — 191 chars of source]

where $(u_t)$ is a $q$-dimensional orthonormal white noise process driving the dynamic common component $(\chi_{it})$ via square-summable filters $\ubar b_i(L)$, and $(\xi_{it})$ is the dynamic idiosyncratic component, weakly correlated over time and cross-section.

While the decomposition into common and idiosyncratic parts is unique, there are infinitely many equivalent representations of the filters and factor process. For any orthonormal $q\times q$ filter $\ubar c(L)$, we have

align[align omitted — 149 chars of source]

with $\tilde u_t = \ubar c(L) u_t$ being orthonormal white noise. The most straightforward way to estimate the GDFM is via dynamic principal components, leading to two-sided filters and factors forni2001generalized, forni2004generalized, hallin2007determining, which cannot be used for forecasting.

We show that whenever the common component is purely non-deterministic (plus a mild regularity condition) - a standard assumption in time series analysis — it admits a representation

align[align omitted — 182 chars of source]

where $\varepsilon_t \in \operatorname{\overline{\operatorname{sp}}}(y_{is} : i \in \mathbb N, s \leq t) := \mathbb H_t(y)$ and $\sum_{j = 0}^\infty \@ifstar{\oldnorm}{\oldnorm*}{K_i (j)}^2 <\infty$, where $\operatorname{\overline{\operatorname{sp}}}(\cdot)$ denotes the closed linear span. Here $(\varepsilon_t)$ is the innovation process of $(\chi_{it})$. This innovation form of the GDFM is unique (up to a real orthogonal matrix) and naturally one-sided.

Earlier work addressed the one-sidedness problem by imposing linear models on the dynamic common component: forni2005generalized used a static factor structure with VAR dynamics, while forni2011general, forni2015dynamic, forni2017dynamic, barigozzi2024inferential modeled the common component as VARMA, achieving one-sided representations in the shocks and output. This paper generalises these results extending the idea used in forni2015dynamic: Write the GDFM in blocked form, with transfer-function blocks $\ubar k^{(j)}(L)$, $j = 1, 2,...$ of dimension $q\times q$:

align[align omitted — 672 chars of source]

If $(\chi_{it})$ is purely non-deterministic and $\ubar k(L)$ is the transfer-function of its Wold representation, the inverse of $\ubar k(L)$ is also causal. Note that $(\phi_t)$ resembles a static factor structure with $(\varepsilon_t)$ as its factors. Suppose that all inverse transfer-functions $\left(\ubar k^{(j)}(L)\right)^{-1}$ are causal and such that the second term on the RHS of equation (ref) is statically idiosyncratic. We can retrieve $\varepsilon_t$ causally from $(y_{it})$ by applying static principal components to $\phi_t$.

General Setup

Notation

Let $\mathcal P = (\Omega, \mathcal A, \operatorname{\mathbb P})$ be a probability space and $L_2(\mathcal P, \mathbb C)$ be the Hilbert space of square integrable complex-valued, zero-mean, random-variables defined on $\Omega$ equipped with the inner product $\langle u, v\rangle = \operatorname{\mathbb E}[u \bar v]$ for $u, v \in L_2(\mathcal P, \mathbb C)$. We suppose that $(y_{it})$ lives in $L_2(\mathcal P, \mathbb C)$ using the following abbreviations: $\mathbb H(y) := \operatorname{\overline{\operatorname{sp}}}(y_{it}: i \in \mathbb N, t \in \mathbb Z)$, the “time domain” of $(y_{it})$, $\mathbb H_t(y):= \operatorname{\overline{\operatorname{sp}}}(y_{is}: i \in \mathbb N, s \leq t)$, the “infinite past” of $(y_{it})$. We write $y_t^n = (y_{1t}, ..., y_{nt})'$ and by $f_y^n(\theta)$ we denote the “usual spectrum” of $(y_t^n)$ times $2\pi$, i.e., $\Gamma_y^n := \operatorname{\mathbb E}\left[y_t^n (y_t^n)^*\right] = (2\pi)^{-1} \int_{-\pi}^\pi f_y^n(\theta) d\theta$. For a stochastic vector $u$ with coordinates in $L_2(\mathcal P, \mathbb C)$, we write $\operatorname{\mathbb V}\left[u\right] := \operatorname{\mathbb E}\left[u u^*\right]$ to denote the variance matrix. Let $u$ be a stochastic vector with coordinates in $L_2(\mathcal P, \mathbb C)$, let $\mathbb M\subset L_2(\mathcal P, \mathbb C)$ be a closed subspace. We denote by $\operatorname{proj}(u \mid \mathbb M)$ the orthogonal projections of $u$ onto $\mathbb M$ deistler2022modelle (coordinate-wise). Furthermore we denote by $\mu_i(A)$ the $i$-th largest eigenvalue of a square matrix $A$. If $A$ is a spectral density $\mu_i(A)$ is a measurable function in the frequency $\theta \in [-\pi, \pi]$. More generally denote by $\sigma_i(A)$ the $i$-th largest singular value of a matrix (not necessarily square).

The Generalised Dynamic Factor Model

Throughout we assume stationarity of $(y_{it})$ in the following sense:

assumption[Stationary Double Sequence] The process $(y^n_t : t \in \mathbb Z)$ is real valued, weakly stationary with zero-mean and such that \begin{itemize} • $y_{it} \in L_2(\mathcal P, \mathbb C)$ for all $(i, t) \in \mathbb N \times \mathbb Z;$ • it has existing (nested) spectral density $f_{y}^n(\theta)$ for $\theta \in [-\pi, \pi]$ defined as the $n\times n$ matrix: \[ f_{y}^n(\theta)=\frac 1{2\pi}\sum_{\ell=-\infty}^{\infty} e^{-\iota \ell \theta} \operatorname{\mathbb E} [y_t^n y_{t-\ell}^{n'}],\quad \theta \in[-\pi,\pi]. \] \end{itemize}

In addition we assume that $(y_{it})$ has a $q$-dynamic factor structure as in forni2001generalized, hallin2011dynamic: Denote by “$\operatorname{ess\,sup}$” the essential supremum of a measurable function, we assume

assumption[$q$-Dynamic Factor Structure] The process $(y_{it})$ is such that there exists $q < \infty$, with \begin{itemize} • $\sup_{n\in \mathbb N} \mu_q\left(f_y^n\right) = \infty$ almost everywhere on $[-\pi, \pi]$; • $\operatorname{ess\,sup}_{\theta \in [-\pi, \pi]} \sup_{n\in \mathbb N} \mu_{q+1}(f_y^n)< \infty$. \end{itemize}

By forni2001generalized Assumption A(ref) is equivalent to the existence of the representation (ref) with $(\xi_{it})$ and $(\chi_{it})$ being orthogonal at all leads and lags $\sup_{n\in \mathbb N}\mu_q\left(f_\chi^n(\theta)\right) = \infty$ almost everywhere on $[-\pi, \pi]$ and $\operatorname{ess\,sup}_{\theta \in [-\pi, \pi]} \sup_{n\in \mathbb N} f_\xi^n(\theta) < \infty$. Note that this also implies that the shocks of $(\chi_{it})$ are orthogonal to $(\xi_{it})$ at all leads and lags and that $f_y^n(\theta) = f_\chi^n(\theta) + f_\xi^n(\theta)$ for $\theta$ almost everywhere in $[-\pi, \pi]$.

We may also describe (ref) as a representation rather than a “model,” since the existence of this dynamic decomposition follows from the characteristic eigenvalue behaviour of the double-indexed process $(y_{it})$. Specifically, the divergence of the first $q$ eigenvalues of $f_\chi^n$ captures the sense in which the filter loadings in (ref) are pervasive. Meanwhile, the essential boundedness of $\sup_{n \in \mathbb N}\mu_1 \left( f_\xi^n \right)$ defines dynamic idiosyncraticness, ensuring that the dynamic idiosyncratic component is only weakly correlated across the cross-section and over time.

The approach of hallin2013factor is even more general: the dynamic common component is defined as the projection onto the Hilbert space spanned by so-called “dynamic aggregates,” which arise as limits of weighted averages in the time domain forni2001generalized. The dynamic idiosyncratic component is then simply the residual from this projection. In principle, the GDFM framework requires only stationarity - without even assuming the existence of a spectral density - to state the decomposition directly. However, this also admits wired cases, such as $q = \infty$.

Infinite Dimensional PND-Processes and Causal invertibility

We begin with recalling some basic facts related to purely non-deterministic processes. Suppose for now that $(x_t) = (x_t^n)$ is a finite dimensional a zero-mean weakly stationary process. We call $\mathbb H^-(x) := \bigcap_{t\in \mathbb Z} \mathbb H_t(x)$ the remote past of $(x_t)$.

definition[Purely Non-Deterministic and Purely Deterministic Stationary Process] If $\mathbb H^-(x) = \{0\}$, then $(x_t)$ is called purely non-deterministic (PND) or regular. If $\mathbb H^-(x) = \mathbb H(x)$, then $(x_t)$ is called (purely) deterministic (PD) or singular.

Note that “singular” in the sense of being purely deterministic must not to be confused with processes that have rank deficient spectrum, like the common component of the GDFM, and are also called “singular” in the literature anderson2008generalized, deistler2010generalized, forni2024approximating. Therefore we shall use PND and PD henceforth to avoid confusion.

Of course there are processes between the two extremes of PD and PND. The future values of a PD process can be predicted perfectly (in terms of mean squared error). On the other hand, a PND process is entirely governed by random innovations. It can only be predicted with positive mean squared error and the further we want to predict ahead, the less variation we can explain: Set

align*[align* omitted — 87 chars of source]

which is the $h$-step ahead prediction error. If $(x_t)$ is PND, then $\lim_{h\to \infty} \operatorname{\mathbb V}\left[\nu_{h|t}\right] \to \operatorname{\mathbb V}\left[x_t\right]$.

Next, we recall some basic facts about the finite dimensional case. By Wold's representation Theorem hannan2012statistical, deistler2022modelle, any weakly stationary process can be written as sum of a PD and PND process, being mutually orthogonal at all leads and lags.

There are several characterisations for PND processes. Firstly, the Wold decomposition implies rozanov1967stationary, masani1957prediction that $(x_t)$ is PND if and only if it can be written as a causal infinite moving average

align[align omitted — 94 chars of source]

where $\nu_t:= \nu_{1|t-1}$ is the innovation of $(x_t)$, possibly with reduced rank $\operatorname{rk} \operatorname{\mathbb V}\left[\nu_t\right] := q \leq n$. If $q<n$ we may uniquely factorise $\operatorname{\mathbb V}\left[\nu_t\right] = bb'$. Assume without loss of generality that the first $q$ rows of $b$ have full rank (otherwise reorder), a unique factor is obtained choosing $b$ to be upper triangular with positive entries on the main diagonal. This results in the representation

align[align omitted — 96 chars of source]

with $K(j) = C(j)b$ and $\operatorname{\mathbb V}\left[\varepsilon_t\right] = I_q$.

Secondly, a stationary process is PND rozanov1967stationary, masani1957prediction if and only if the spectral density has constant rank $q\leq n$ almost everywhere on $[-\pi, \pi]$ and can be factored as

align[align omitted — 299 chars of source]

$\@ifstar{\oldnorm}{\oldnorm*}{\cdot}_F$ denotes the Frobenius norm and

align[align omitted — 189 chars of source]

here $z$ denotes a complex number. The entries of the spectral factor $\ubar k(z)$ are analytic functions in the open unit disc $D$ and belong to the class $L^2(T)$, i.e., are square integrable on the unit circle $T$. If we employ the normalisation $K(0) = b$ from above, the transfer-function $\ubar k(z)$ from (ref) which corresponds to the Wold representation (ref) is causally invertible, i.e. $\mathbb H_t(\varepsilon) \subset \mathbb H_t(x)$. Causal invertibility is equivalent to $\operatorname{rk} k(z) = q$ for all $\@ifstar{\oldabs}{\oldabs*}{z} < 1$; there are no zeros inside the unit circle. We say also that the shocks $(\varepsilon_t)$ are fundamental for $(x_t)$.

fact[szabados2022regular] If $(x_t)$ is PND with $\operatorname{rk} f_x = q < n$ almost everywhere on $[-\pi, \pi]$, then there exists a $q$-dimensional sub-vector of $x_t$, say $\tilde x_t = (x_{i_1,t}, ..., x_{i_q, t})'$ of full dynamic rank, i.e., $\operatorname{rk} f_{\tilde x} = q$ almost everywhere on $[-\pi, \pi]$.

To see why, we follow the proof of Theorem 2.1 in szabados2022regular. A principal minor $M(\theta) = \det \left[(f_{x})_{i_j, i_l}\right]_{j, l = 1}^q$ of $f_x$ can be expressed by means of equation ((ref)) in terms of

align*[align* omitted — 322 chars of source]

with the same row indices in the minor $M_{\underline{k}}(z)$ of $\ubar k(z)$ as in the principal minor $M(\theta)$ of $f_x$. We know that $M_{\underline{k}}(z) = 0$ almost everywhere or $M_{\underline{k}}(z) \neq 0$ almost everywhere because $M_{\underline{k}}(z)$ is analytic in $D$. Since $\operatorname{rk} f_x = q$ almost everywhere, the sum of all principal minors of $f_x$ of order $q$ is different from zero almost everywhere, so there exists at least one order $q$ principal minor of $f_x$ different almost everywhere from zero.

Let us now consider the case of an infinite dimensional rank deficient PND process $(x_{it}: i \in \mathbb N, t \in \mathbb Z)$, so $\operatorname{rk} f_x^n = q$ almost everywhere on $[-\pi, \pi]$ for all $n \geq n_0$.

definition[Purely non-deterministic rank-reduced stochastic double sequence] Let $(x_{it})$ be a stationary stochastic double sequence such that $\operatorname{rk} f_x^n = q <\infty$ almost everywhere on $[-\pi, \pi]$ for all $n\geq n_0$. We say that $(x_{it})$ is PND if there exists an $n_1\geq n_0$ together with a $q$-dimensional orthonormal white noise process $(\varepsilon_t)\sim WN(I_q)$ such that \begin{itemize} • $ \varepsilon_t \in \operatorname{sp}\left(x_t^n - \operatorname{proj}\left[x_t^n \mid \mathbb H_{t-1}(x^n)\right] \right)$ for all $n \geq n_1$ and $t\in \mathbb Z$; • $x_{it} \in \mathbb H_{t}(\varepsilon)$ for all $i\in \mathbb N$ and for all $i\in \mathbb N$ there is a causal transfer-function $\ubar k_i(L)$ such that \begin{align} x_{it} = \ubar k_i (L) \varepsilon_t = \sum_{j = 0}^\infty K_i(j) \varepsilon_{t-j}, \end{align} where $\sum_{j = 0}^\infty \@ifstar{\oldnorm}{\oldnorm*}{K_i(j)}^2 < \infty$ and $\ubar k_i(z)$ are analytic in the open unit disc $D$ for all $i \in \mathbb N$. \end{itemize}

We conjecture that Definitions (ref) and (ref) are equivalent also for the infinite dimensional rank deficient case, which is however not the objective of the present paper. Uniqueness of $(\varepsilon_t)$ can be achieved e.g. by selecting the first index set in order such that the process has full rank almost everywhere on $[-\pi, \pi]$ and imposing constraints as described below equation (ref). An example for a purely deterministic double sequence would be $x_{it} = \varepsilon_{t+i-1}$; here we can perfectly predict the infinite future at time $t$ from $(x_{it}: i\in \mathbb N)$.

Next, we discuss fundamentalness in the infinite dimensional, rank deficient case. Thinking of $(x_t) = (x_{1t}, x_{2t}, ....)'$ as an infinite dimensional vector process, a full rank $q$-dimensional sub-block as in Fact (ref) has a transfer-function that is invertible, but not necessarily causally invertible. For illustration, consider the following examples with $\varepsilon_t$ being scalar ($q = 1$) white noise with unit variance:\\

tabular[tabular omitted — 838 chars of source]

\\ Note that in all three cases $(x_{it})$ is a stationary PND double sequence.

itemize• Starting with example (ref), let $\ubar k^n(z)$ be the transfer-function of $(x_t^n) = (x_{1t}, ..., x_{nt})'$ for $n\in \mathbb N$ as in (ref). We note that $\operatorname{rk} \ubar k^n(z_0) = 0$ for $z_0 = 1/3$ for all $n \in \mathbb N$. Therefore $\ubar k^n(z)$ has a zero inside the unit circle and is not causally invertible for any $n\in \mathbb N$ and (ref) is not the Wold representation (a non-causal inverse representation is given by $\varepsilon_t = -1/3 \sum_{j = 1}^\infty (1/3)^{j-1} x_{1, t+j}$). By the spectral factorisation we can obtain a causally invertible factor of the spectrum of the univariate processes $x_{it}$ for $i \in \mathbb N$ by mirroring the zero on the unit circle: Rewrite $f(z) = (1 - 3z)(1-3z^{-1}) = (1-3z^{-1})z \ (1 - 3z)z^{-1} = 3(1 - 1/3z) 3(1-1/3 z^{-1})$. Consequently, there is a white noise unit variance scalar innovation process, say $(\eta_t)$ which is different from $(\varepsilon_t)$, such that $x_{it} = 3 \eta_t - \eta_{t-1}$ associated with the causally invertible transfer-function $3(1 - 1/3L)$. Here $(\eta_t)$ is the innovation for each individual univariate process and also for the entire multivariate infinite dimensional process $(x_t)$. • In example (ref), setting $\tilde x_t := x_{1t}$, the associated transfer-function is causally invertible. So is the transfer-function of any other sub-process of dimension $n > 1$ which includes the first coordinate $x_{1t}$. It follows that $(\varepsilon_t)$ is the innovation process of the multivariate process $(x_t)$. On the other hand setting $\tilde x_t := x_{it}$ for $i \geq 2$, the transfer-function of $(\tilde x_t)$ is not causally invertible. Hence, even though $(\varepsilon_t)$ is the innovation for the multivariate rank-deficient process $(x_t)$ it is in general not the innovation for its full rank sub-blocks (compare example (ref)). The first coordinate settles the innovation for the whole infinite dimensional process $(x_t)$ and (ref) is the Wold representation. • Finally in example (ref), the associated transfer-functions of all one-dimensional sub-processes are not causally invertible. However the transfer-functions of $2$-dimensional sub-blocks such as $\tilde x_t = (x_{1t}, x_{2t})'$ are causally invertible. They have full rank for all $z\in \mathbb C$ and therefore also for all $z$ inside the unit circle. For instance a causal inverse is given by $\varepsilon_t = -2 x_{1t} + 3 x_{2t}$. Consequently, potential non-fundamentalness of the shocks with respect to the output can be tackled by adding new cross-sectional dimensions which are driven by the same shocks, and therefore “remove” zeros inside the unit circle. For example forni2025common exploit this fact to make structural VAR analysis more robust. As has been shown by anderson2016structure, a rank deficient VARMA system (i.e. $n>q$) has an autoregressive representation, i.e. $\operatorname{rk} k^n(z) = q$ for all $z\in \mathbb C$ generically in the parameter space. For a different approach to non-fundamentalness see funovits2024identifiability.

To summarise, an infinite-dimensional, rank-deficient PND process typically contains many full-rank sub-blocks. While these sub-blocks are not necessarily causally invertible, they are “more likely” to be so as the dimension of the sub-block grows. This is because potential zeros inside the unit circle can be compensated for by the contribution of additional rows in the transfer function.

Econometric time series analysis (in the realm of stationarity) is almost exclusively concerned with the modelling and prediction of “regular” time series. As noted above VAR, VARMA and state space models are all PND deistler2022modelle. If we think of $(y_{it})$ as a process of (stationarity transformed) economic data, we would not expect that any part of the variation of the process could be explained in the far distant future given information up to now. The same should hold true for the common component which explains a large part of the variation of the observed process. Even more so, if we interpret the idiosyncratic component of the GDFM as measurement errors lippi2021validating, forni2025common.

Therefore, we impose the following assumption:

assumption[Purely Non-Deterministic Dynamic Common Component] The dynamic common component $(\chi_{it})$ of the GDFM is PND with orthonormal white noise innovation $(\varepsilon_t)$ (of dimension $q$) and innovation-form \begin{align*} \chi_{it} =\ubar k_i (L)\varepsilon_t = \sum_{j = 0}^\infty K_i (j) \varepsilon_{t-j}. \end{align*}

This assumption resolves only half of the one-sidedness issue as it is not clear whether $\varepsilon_t$ has a representation in terms of current and past $y_{it}$'s.

One-Sidedness of the Common Shocks in the Observed Process

Consider the sequence of $1\times q$ row transfer-functions $(k_i: i \in \mathbb N)$. For the proof of Theorem (ref), the main result of the paper, we rely on the following key property: By reordering and stacking, we can construct a sequence of blocks $(k^{(j)}: j\in \mathbb N)$ of dimension $q_j\geq q$ such that the left-inverses $(k^{(j)})^{\dag}$ (here “$\dag$” denotes the generalised inverse) are causal and absolutely summable filters. First, we show that we can build infinitely many full-rank blocks by reordering the sequence. Second, we argue why it is reasonable to assume that we can stack the blocks in such a way that each block is also causally invertible, i.e. there are no zeros inside the unit circle. Third, absolute summability requires that the blocks have no zeros on the unit circle - a condition we will relax in the discussion following the proof of Theorem (ref).

theoremUnder Assumptions A(ref)-A(ref), there exists a reordering $\left(k_{i_l}: l \in \mathbb N\right)$ of the sequence $(k_{i}: i \in \mathbb N)$ such that all consecutive $q \times q$ blocks $(k^{(j)})$ of $\left(k_{i_l}: l \in \mathbb N\right)$ have full rank $q$ almost everywhere on $[-\pi, \pi]$.
remarkBy Assumption A(ref), we know that \begin{align} \mu_q\left(f_\chi^n \right) = \mu_q \bigg( \left(k^n\right)^* k^n \bigg) \to \infty \quad almost everywhere on \ [-\pi, \pi], \end{align} with $k^n = (k_1', ...., k_n')'$. Since by Assumption A(ref), $\ubar k_i$ is analytic in the open unit disc, it follows that either $k_i(\theta) = 0$ or $k_i(\theta) \neq 0$ almost everywhere on $[-\pi, \pi]$. If $k_i(\theta) = 0$ almost everywhere, then $\chi_{it} = 0$ and therefore $\chi_{it} \in \mathbb H_t(y)$. By ((ref)) the number of non-zero rows $k_i$ must be infinite. Therefore Theorem (ref) holds if and only if it holds after removing all rows with $k_i = 0$ almost everywhere on $[-\pi, \pi]$.
proof[Proof of Theorem (ref)] Concurring with Remark (ref), we assume that $(k_i: i \in \mathbb N)$ has no zero rows and prove the statement by constructing the reordering using induction. By Assumption A(ref) and Fact (ref) and equation ((ref)), we can build the first $q\times q$ block, having full rank almost everywhere on $[-\pi, \pi]$ by selecting the first linearly independent rows $i_1, ..., i_q$ of the sequence of row transfer-functions $(k_i: n \in \mathbb N)$, i.e., set $k^{(1)} = (k_{i_1}', ... k_{i_q}')'$. Now look at the block $j+1$: We use the next $k_i$ available in order, as the first row of $k^{(j+1)}$, i.e., $k_{i_{jq + 1}}$. Suppose we cannot find $k_i$ with $i \in \mathbb N \setminus \{ i_l: l \leq jq +1 \}$ linearly independent of $k_{i_{jq + 1}}$. Consequently, having built already $j$ blocks of rank $q$, all subsequent blocks that we can build from any reordering are of rank $1$ almost everywhere on $[-\pi, \pi]$. In general, for $\bar q < q$, suppose we cannot find rows $k_{i_{jq + \bar q + 1}}, ..., k_{i_{jq + q}}$ linearly independent of $k_{i_{jq+1}}, ...,k_{i_{jq + \bar q}}$, then all consecutive blocks that we can obtain from any reordering have at most rank $\bar q$. For all $m = j+1, j+2, ...$ by the RQ-decomposition we can factorise $k^{(m)} = R^{(m)} (\theta) Q^{(m)}$, where $Q^{(m)} \in \mathbb C^{q \times q}$ is orthonormal and $R^{(m)}(\theta)$ is lower triangular $q\times q$ filter which is analytic in the open unit disc. For $n \geq i_{j}$ and without loss of generality that $n$ is a multiple of $q$, the reordered sequence looks like { \begin{align*} (k^n)^* k^n &= \sum_{l=1}^n k_{i_l}^* k_{i_l} \\[-1em] &=\begin{bmatrix} \left(k^{(1)}\right)^* & \cdots & \left(k^{(j)}\right)^* \end{bmatrix} \begin{bmatrix} k^{(1)} \\ \vdots\\ k^{(j)} \end{bmatrix} + \begin{bmatrix} \left(R^{(j+1)}\right)^* & \cdots & \left(R^{(J)}\right)^* \end{bmatrix} \begin{bmatrix} R^{(j+1)} \\ \vdots \\ R^{(J)} \end{bmatrix} \\ &= \begin{bmatrix} \left(k^{(1)}\right)^* & \cdots & \left(k^{(j)}\right)^* \end{bmatrix} \begin{bmatrix} k^{(1)} \\ \vdots\\ k^{(j)} \end{bmatrix} + \begin{pmatrix} \times & 0 \\ 0 & 0 \end{pmatrix} = A + B^n, \ say, \end{align*} } where $\times$ is a placeholder. By the structure of the reordering, there are $q - \bar q$ zero end columns/rows in $B^n$ for all $n \geq jq$ where $A$ remains unchanged. Now by lancaster1985theory, we have \begin{align*} \mu_q\bigg(\left(k^n\right)^* k^n\bigg) &= \mu_q(A + B^n) \\[-0.8em] &\leq \mu_1(A) + \mu_q (B^n) \\ &= \mu_1(A) < \infty \ for all n \in \mathbb N \ almost everywhere on \ [-\pi, \pi] . \end{align*} This also implies that for any reordering the $q$-th eigenvalue of the resulting inner product of the transfer-function as in equation ((ref)) is bounded by $\mu_1(A)$. This is a contradiction and completes the induction step and the proof.

Theorem (ref) shows that we can extract infinitely many sub-blocks of of dimension $q$ from $(\chi_{it})$, each with full-rank spectrum almost everywhere on $[-\pi, \pi]$. This provides us with an unlimited supply of such “variable stacks,” each capable of capturing the signal from all $q$ common shocks.

Next, we impose the following assumption:

assumption[Uniformly Strictly Minimum Phase after Blocking] There exists a sequence of blocks $(k^{(j)}: j \in \mathbb N)$ of dimension $q_j\times q$ with $q_j\geq q$ constructed from $(k_i: i \in \mathbb N)$ (by reordering/elimination and appropriate blocking), such that $\sigma_q\left(k^{(j)}\right) > \delta > 0$ almost everywhere on $[-\pi, \pi]$ for all $j \in \mathbb N$.

This is similar to the commonly employed assumption in linear systems theory that the transfer-function is strictly minimum-phase, i.e., has no zeros on the unit circle as assumed in deistler2010generalized, section 2.3 or in forni2017dynamic, Assumption 7.

Some further comments in order: Consider a sequence of full rank $q\times q$ blocks as in Theorem (ref). First let us remark that the lack of causal invertibility of a full rank transfer-function block is the non-standard case. So we could simply assume that all consecutive $q\times q$ blocks (or a subsequence thereof) are causally invertible. On the other hand this excludes examples (ref) and (ref). However, as demonstrated in the discussion of example (ref), zeros can be removed by adding additional linearly independent rows to a block. For instance, we might be able to paste blocks together, increasing to dimension to $q_j\times q$ with $q_j \geq q$, while $q_j$ can be arbitrarily large, such that all zeros inside the unit circle vanish, so Assumption A(ref) holds for (ref). Furthermore, note that example (ref) is also covered by A(ref) with $(\eta_t)$ as innovation instead of $(\varepsilon_t)$ (see the related discussion). Only cases like (ref) are ruled out by A(ref), where the innovation is determined by a finite number of transfer-function rows while all other rows have the same zeros inside the unit circle which cannot be removed by stacking. Even then, causal invertibility holds if we would ignore those rows, i.e. row $i = 1$ in (ref) which brings us back to the case of (ref).

Next, Assumption A(ref) requires the stacks to be not only minimum phase (no zeros inside the unit circle) but strictly minimum phase (no zeros inside and on the unit circle) with a uniform bound $\delta$. This excludes e.g. $\ubar k_i(L) = 1-L, i\in \mathbb N$ or $\ubar k_i(L) = (1 - (1-\frac{1}{i})L), i\in \mathbb N$. We will show how to incorporate those cases after the proof of Theorem (ref). On the other hand, we may argue that the uniformly strict minimum phase property with a global bound $\delta$ can be achieved by eliminating blocks or extending their size as described above.

In summary, we conclude that Assumption A(ref) is in fact a very mild restriction, excluding only rather contrived edge cases.

theoremSuppose A(ref)-A(ref) hold for $(y_{it})$, then the innovations $(\varepsilon_t)$ of $(\chi_{it})$ are causally subordinated to the observed variables $(y_{it})$, i.e., $\varepsilon_t \in \mathbb H_t(y)$.

We apply remark (ref) also for Theorem (ref). Trivially, since $(k_i: i \in \mathbb N)$ are also causal, Theorem (ref) directly implies that the dynamic common component is causally subordinated to the observed output, i.e. $\chi_{it} \in \mathbb H_t(y)$ for all $i\in \mathbb N, t \in \mathbb Z$. Furthermore, if we had to eliminate rows to satisfy Assumption A(ref), then after recovering the common shocks $(\varepsilon_t)$ causally from $(y_{it})$, we can also reconstruct the common component of the eliminated variables by projection $\chi_{it} = \operatorname{proj}(y_{it} \mid \mathbb H_t(\varepsilon_t))$ forni2001generalized, gersing2023reconciling.

proof[Proof of Theorem (ref)] Suppose $(k^{(j)} : i \in \mathbb N)$ is such that Assumption A(ref) is satisfied. Suppose $\sum_{j = 1}^J q_j = n$ without loss of generality. { \begin{align*} \chi_t^n = \begin{pmatrix} \chi_t^{(1)} \\ \chi_t^{(2)} \\ \vdots \\ \chi_t^{(J)} \end{pmatrix} = \begin{pmatrix} \ubar k^{(1)}(L) \\ \vdots \\ \ubar k^{(J)}(L) \end{pmatrix} \varepsilon_t & = \begin{pmatrix} \ubar k^{(1)}(L) & &\\ & \ddots & \\ & & \ubar k^{(J)}(L) \end{pmatrix} \begin{pmatrix} I_q \\ \vdots \\ I_q \end{pmatrix} \varepsilon_t . \end{align*} } By Assumption A(ref), we know that all left-inverse transfer-functions $(k^{(j)})^{\dag}, j = 1, ..., J$ are causal as well. Next we show that $(\varphi_{it})$ in { \begin{align} \varphi_t^{qJ} &:= \begin{pmatrix} \left(\ubar k^{(1)}\right)^{\dag}(L) & & \\ & \ddots & \\ & & \left(\ubar k^{(J)}\right)^{\dag}(L) \end{pmatrix} \begin{pmatrix} y_t^{(1)} \\ \vdots \\ y_t^{(J)} \end{pmatrix} \nonumber \\ &= \begin{pmatrix} I_q \\ \vdots \\ I_q \end{pmatrix} \varepsilon_t + \begin{pmatrix} \left(\ubar k^{(1)}\right)^{\dag}(L) & &\\ & \ddots & \\ & & \left(\ubar k^{(J)}\right)^{\dag}(L) \end{pmatrix} \begin{pmatrix} \xi_t^{(1)} \\ \vdots \\ \xi_t^{(J)} \end{pmatrix} = C_t^{\varphi, qJ} + e_t^{\varphi, qJ} , say, \ \end{align} } has a static factor structure (see definition (ref)). Then $\varepsilon_t$ can be recovered from static aggregation/ via static principal components applied to $(\varphi_{it})$ by Theorem (ref).2. Firstly, $q$ eigenvalues of $\Gamma_{C^\varphi}^{qJ} = \operatorname{\mathbb E} \left[C_t^{\varphi, qJ} C_t^{\varphi, qJ'}\right]$ diverge for $J(n) \to \infty$ as $n\to \infty$, so A(ref)(i) holds. We are left to show that the first eigenvalue of $\Gamma_{e^\varphi}^n = \operatorname{\mathbb E}\left[e_t^{\varphi, qJ}e_t^{\varphi, qJ'}\right]$ is bounded in $qJ$, i.e., A(ref)(ii) holds. Let $U_j \Sigma_j V_j^* = k^{(j)}(\theta)$ be the singular value decomposition while $U_j$ is $n\times q$ with orthonormal columns, $\Sigma_j = \operatorname{diag}(\sigma_1(k^{(j)}), ..., \sigma_q(k^{(j)}))$ is the diagonal matrix of singular values of $k^{(j)}$ and $V_j$ is a $q\times q$ unitary matrix, where we suppressed the dependence on $\theta$ in the notation on the LHS. Let $f_\xi^n(\theta) = P^* M P$ be the eigen-decomposition of $f_\xi^n$ with orthonormal eigenvectors being the rows of $P$ and eigenvalues in the diagonal matrix $M$ (omitting dependence on $n$). Then \begin{align*} f_{e^\varphi}^{qJ}(\theta) = \bigoplus_{j = 1}^J V_j \underbrace{\bigoplus_{j = 1}^J \Sigma_j^{-1} \bigoplus_{j = 1}^J U_j^* P^* M P \bigoplus_{j = 1}^J U_j \Sigma_j^{-1}}_{B^J(\theta)} \bigoplus_{j = 1}^J V_j^*, \end{align*} where we used $\bigoplus_{j = 1}^J A_j$ to denote the block diagonal matrix with the square matrices $A_j$ for $1\leq j \leq J$ on the main diagonal block. The largest eigenvalue of $f_{e^\varphi}^{qJ}(\theta)$ is equal to the largest eigenvalue of $B^J(\theta)$. Therefore by Jensen's inequality and Assumption A(ref) we have \begin{align*} \mu_1\left(\Gamma_{e^\varphi}^{qJ} \right) &= \mu_1\left(\int_{-\pi}^\pi f_{e^{\varphi}}^{qJ}\right) \leq \int_{-\pi}^\pi \mu_1\left(f_{e^{\varphi}}^{qJ}\right) \\[0.8em] &\leq 2\pi \sup_{J\in \mathbb N} \operatorname{ess\,sup}_{\theta \in [-\pi, \pi]} \mu_1 \left(f_{e^{\varphi}}^{qJ}\right) \\ & \leq 2 \pi \sup_{J\in \mathbb N} \left\{\operatorname{ess\,sup}_{\theta \in [-\pi, \pi]} \mu_1 \left(f_{\xi}^{n}\right) \times \sup_{1\leq j \leq J} \operatorname{ess\,sup}_{\theta \in [-\pi, \pi]} \sigma_q\left(k^{(j)}\right)^{-2} \right\} \\ & \leq 2 \pi \sup_{J\in \mathbb N} \operatorname{ess\,sup}_{\theta \in [-\pi, \pi]} \mu_1 \left(f_{\xi}^{n}\right) \delta^{-2} < \infty. \end{align*} This completes the proof.

Note that we employed the uniform strict minimum phase property of Assumption A(ref) in the proof of Theorem (ref) to ensure that the filtered idiosyncratic blocks $(k^{(j)})^{\dag}(L) \xi_t^{(j)} = e_t^{\varphi, (j)}$ have finite variance. This condition prevents the highly-nongeneric situation in which zeros appear on the unit circle in almost all blocks, irrespective of how they are stacked. Nevertheless, in what follows, we show that absolute summability of the left-inverse transfer function blocks is not in fact required causal subordination:

For instance, consider the model

align*[align* omitted — 98 chars of source]

where $(\xi_{it})$ is dynamically idiosyncratic. Clearly, the coefficients of $(1-L)^{1}$ are not absolutely summable. Still, we can recover $(\varepsilon_t)$ one-sided in the $(y_{it})$: The cross-sectional average $\bar y_t^n = n^{-1}\sum_{i = 1}^n y_{it}$ is a static aggregation of $(y_{it})$ and therefore $\bar y_t^n \to \zeta_t$ converges in mean square gersing2023reconciling, gersing2024distributed. Furthermore $(\zeta_t)$ is PND and by the Wold representation Theorem, the innovations are recovered from the infinite past $\mathbb H_t(\zeta) \subset \mathbb H_t(y)$: It suffices that the inverse of $(1-L)$ exist for the input $(\zeta_t)$, since

align*[align* omitted — 375 chars of source]

we know that $(1-e^{-\iota \theta})^{-1}$ is an element of the frequency domain of $(\zeta_t)$ and therefore has an inverse also in the time domain anderson2016multivariate.

More generally, we may employ this procedure by factoring out the zeros from the analytic functions $(\ubar k^{(j)}: j\in \mathbb N)$. Let $\ubar k^{(j)} = \ubar g_j \ubar h_j$, where $\ubar g_j$ is a polynomial defined by the zeros of $\ubar k^{(j)}$ which are on the unit circle. Recall that the zeros of an analytic function are isolated, so if $z_0$ is a zero, we have $\ubar h_j(z)\neq 0$ in a neighbourhood around $z_0$. Furthermore the degree of a zero can be only finite or the function is zero everywhere. Thus there can be only finitely many different zeros on the unit circle, since the unit circle is a compact set. Write $\ubar g_j(z) = \Pi_{k_j = 1}^{M_j} (z - z_{j k_j})^{m_{k_j}}$ with $|z_{j k_j}| = 1$ for $k_j = 1, ..., M_j$ and $j \in \mathbb N$. It follows that $g_j^{-1}(\theta) k^{(j)}(\theta) \neq 0$ almost everywhere on $[-\pi, \pi]$, where $g_j(\theta) := \ubar g_j (e^{-i\theta})$.

Consequently setting $\ubar k^{(j)}(L) = \ubar g_j(L) \ubar h^{(j)}(L)$, we have

align[align omitted — 409 chars of source]

with $\operatorname{rk} \ubar h_j(L) = q$ almost everywhere on $[-\pi, \pi]$. Now, if Assumption A(ref) holds for $(h^{(j)})$ instead of $(k^{(j)})$, with the same arguments as above, we obtain a one-sided representation of $(\varepsilon_t)$, if there exists a static averaging sequence $(\hat c_i^{(n)}:(i, n) \in \mathbb N \times \mathbb N)$ (see definition (ref)) such that

align[align omitted — 182 chars of source]

converges in mean square to a PND process, say $\zeta_t$, with innovations $(\varepsilon_t)$, with $\varphi_{it}, C_{it}^{\varphi}, e_{it}^\varphi$ from equation ((ref)).

Summing up, even in the highly non-generic case where zeros lie on the unit circle in almost all transfer-function blocks, regardless of how they are stacked, it remains possible to retrieve the common innovations $(\varepsilon_t)$ causally from $(y_{it})$ by factoring out the zeros, inverting the invertible part, and then aggregating to recover the common shocks. This requires the existence of an aggregate $(\zeta_t)$ as in (ref) with innovations $(\varepsilon_t)$. As formally proving this in complete generality may be challenging, we assume such pathological cases are of limited practical relevance and do not pursue them further here.

The inclusion of edge cases relating zeros inside and on the unit circle in the $q\times q$ transfer-function blocks, while mostly of theoretical interest, highlight the generality of the GDFM’s one-sidedness rather than suggesting a practical estimation method. In practice, achieving the required structure by reordering and stacking variables is non-trivial, but we may assume that robust approaches - like using blocks of size $q+1$ or $q+2$ as in forni2015dynamic, forni2017dynamic, barigozzi2024inferential — are sufficient. Alternatively, recent methods gersing2024distributed, gersing2024weak estimate the dynamic common component by projecting onto current and past factors (imposing additional assumptions) extracted via static principal components, avoiding concerns about non-fundamentalness or unit-circle zeros — provided a one-sided representation exists.

Conclusion

We conclude by highlighting that causal subordination is a distinctive feature of the GDFM decomposition, rooted in the specific behaviour of its diverging spectral eigenvalues. Unlike general dynamic low-rank approximations via dynamic principal components, e.g. of dimension $q+h$, for $h>1$, the GDFM’s $q$ diverging spectral eigenvalues allow us to construct infinitely many full-rank transfer-function blocks. By causally inverting these blocks (potentially after factoring out spectral zeros) and then aggregating, we can recover the common innovations causally from the observed process. Consequently, provided that the dynamic common component is purely non-deterministic, it is itself causally subordinated to the observed process.

Acknowledgements

{{I am deeply grateful to my PhD supervisor, Manfred Deistler, for his guidance and support. I would also like to thank the two anonymous referees for their valuable comments, particularly for highlighting the issue of causal invertibility, which significantly improved the paper. My thanks further go to Paul Eisenberg and Sylvia Frühwirth-Schnatter for their helpful suggestions. Financial support from the Austrian Central Bank under Anniversary Grant No.18287, the DOC Fellowship of the Austrian Academy of Sciences (ÖAW), and the University of Vienna is gratefully acknowledged. }}

Data Availability Statement

Data sharing is not applicable to this article as no datasets were generated or analysed during the current study.