The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
72,562 characters
A Distributed Lag Approach to the Generalised Dynamic Factor Model
\onehalfspacing
\title{
A Distributed Lag Approach \\
to the Generalised Dynamic Factor Model
}
\author{
\textsc{Philipp Gersing}\footnote{Department of Statistics and Operations Research, University of Vienna},
}
\maketitle
\begin{abstract}
We propose a new estimator for the Generalised Dynamic Factor Model (GDFM) that simplifies estimation by avoiding frequency-domain methods. Our key theoretical insight shows that under reasonable conditions the dynamic common component can be represented in terms of a finite number of lags of contemporaneously pervasive factors. In this case the dynamic factor decomposition of the GDFM reduces to the OLS regression of observed variables on estimated factors and their lags, with factors obtained via static principal components. The approach naturally accommodates weak (non-pervasive) factors within the dynamic common space addressing an important limitation of existing methods. We establish consistency and asymptotic normality for both the dynamic and weak common components. An application to a large European macroeconomic dataset demonstrates strong empirical performance and uncovers a sizeable weak common component - particularly in sentiment indicators and several other variables - revealing dynamics that standard methods overlook.
\end{abstract}
\textbf{\textit{Index terms---}} Approximate Factor Model, Generalized Dynamic Factor Model, Weak Factors
\section{Introduction}
We consider a high-dimensional time series panel as double indexed zero mean stationary process $(y_{it}: i \in \mathbb N, t \in \mathbb Z) \equiv (y_{it})$. The Generalised Dynamic Factor Model (GDFM), introduced by \cite{forni2000generalized, forni2001generalized, hallin2013factor}, relies on a decomposition of the form
\begin{align}
y_{it} = \chi_{it} + \xi_{it} = \sum_{j =0}^\infty \bm{B}_i(j) \bm{\varepsilon}_{t-j} + \xi_{it}, \quad \bm{\varepsilon}_t \sim WN(\bm{I}_q), \label{eq: GDFM rep}
\end{align}
where $(\bm{\varepsilon}_t)$ is an orthonormal white noise process of the common innovations, $\ubar \bm{B}_i(L)$ are pervasive causal square summable filters \citep[see][for details on causal subordination of $(\varepsilon_t)$ to $(y_{it})$]{forni2015dynamic, gersing2026existence}, $(\chi_{it})$ is the dynamic common component - driven by the same $q$ shocks for all units $i$ and $(\xi_{it})$ is the dynamic idiosyncratic component, which is weakly correlated over time and cross-section and orthogonal to $(\chi_{it})$ at all leads and lags. Note, that in view of \cite{hallin2013factor}, stationarity is sufficient for obtaining a dynamic factor model decomposition—with possibly an infinite number $q$ of common shocks.
The GDFM has proven to be a powerful tool in many applications \citep[see][]{forni2005generalized, forni2004generalized, forni2017dynamic, forni2018dynamic, barigozzi2020consistent, barigozzi2021time, barigozzi2024fnets, barigozzi2025general, anderson2008generalized} and allows for the presence of weak non-pervasive factors in the dynamically common space \citep{gersing2023weak}. The available tools for estimation \cite{forni2001generalized, hallin2007determining, forni2005generalized, forni2017dynamic, barigozzi2024inferential, barigozzi2024fnets, barigozzi2024algebraic} involve spectral density estimation and several steps to achieve a representation that is one-sided in the observed variables and therefore fit for forecasting.
In this paper, we argue that under restricted but reasonable conditions the GDFM decomposition \eqref{eq: GDFM rep} can be estimated by regressing the observed variables on factors estimated by static principal components \textit{and their lags}. We propose the restricted model
\begin{align}
y_{it} &= \chi_{it} + \xi_{it} \nonumber \\
&= \bm{\beta}_{i0} \bm{F}_t + \bm{\beta}_{i1} \bm{F}_{t-1} + ... + \bm{\beta}_{ip} \bm{F}_{t-p} + \xi_{it} = \bm{\beta}_i \bm{x}_t + \xi_{it} \label{eq: GDFM stat rep}
\end{align}
where $\bm{x}_t := (\bm{F}_t', ..., \bm{F}_{t-p}')'$ is of dimension $r_\chi = r (p+1)$ and $(\xi_{it})$ is dynamically idiosyncratic, orthogonal to $(\bm{F}_t)$ at all leads and lags. Here $\bm{F}_t$ is an $r\times 1$ vector of statically pervasive factors, while the dimension $r$ is uniquely determined by the number of diverging signal eigenvalues in the variance matrix of $(\bm{y}_t^n) = (y_{1t}, ..., y_{nt})'$ for $n\to \infty$ \citep[see][]{chamberlain1983arbitrage}. In addition to the dynamic factor structure \eqref{eq: GDFM rep}, we assume that $(y_{it})$ also has a static factor structure as standard in the literature commencing with \cite{chamberlain1983arbitrage, chamberlain1983funds, stock2002forecasting, bai2002determining}:
\begin{align}
y_{it} = C_{it} + e_{it} = \bm{\Lambda}_i \bm{F}_t + e_{it}, \label{eq: static factor rep}
\end{align}
where $C_{it}$ is the static common component and $(e_{it})$ is the static idiosyncratic component, weakly correlated in the cross-section and \textit{contemporaneously} orthogonal to $(e_{it})$.
The key point is that the dynamic common component $(\chi_{it})$ in model \eqref{eq: GDFM rep} captures richer dynamics than the static common component $(C_{it})$ in \eqref{eq: static factor rep}. In particular, a central insight of this paper is that $\chi_{it}$ may incorporate feedback from lagged values of $\bm{F}_t$. Notably, this differs from specifying dynamics for $\bm{F}_t$ while imposing the static decomposition \eqref{eq: static factor rep} as e.g. in \cite{forni2009opening, doz2011two, barigozzi2019quasi}. More broadly, the dynamic common component represents the part that is dynamically common - that is, the projection of observed variables onto the infinite past of the common structural shocks \citep[see][]{gersing2026existence}. By contrast, the static common component captures only contemporaneous co-movement.
The general theory linking \eqref{eq: GDFM rep} and \eqref{eq: static factor rep} is introduced in \cite{gersing2023reconciling, gersing2023weak} and subsequently considered in the time-domain-framework in \cite{barigozzi2026dynamic}. Both models can be nested within the following canonical decomposition:
\begin{align}
y_{it} &=\lefteqn{\underbrace{\phantom{C_{it} + e_{it}^\chi}}_{\chi_{it}}} C_{it} + \overbrace{e_{it}^\chi + \xi_{it}}^{e_{it}} = \bm{\Lambda}_i \bm{F}_t + \bm{\Lambda}_i^w \bm{F}_t^w + \xi_{it} \label{eq: can decomp univar} \\
\bm{y}_t^n &= \bm{C}_t^n + \bm{e}_t^{\chi, n} + \bm{\xi}_t^n \nonumber \\
&= \underbrace{\bm{\Lambda}^n \bm{F}_t + \bm{\Lambda}^{w, n} \bm{F}_t^w}_{\bm{\chi}_t^n} + \bm{\xi}_t^n
=\begin{bmatrix}
\bm{\Lambda}^n & \bm{\Lambda}^{w, n}
\end{bmatrix}\begin{bmatrix}
\bm{F}_t \\[0.5em] \bm{F}_t^w
\end{bmatrix} + \bm{\xi}_t^n , \label{eq: canonical rep weak and strong factors}
\end{align}
The static common component $(C_{it})$ is driven by statically pervasive, or briefly, strong factors, i.e. associated with thick loading columns, usually expressed in terms of $\bm{\Lambda}^{n'}\bm{\Lambda}^n / n \to \bm{\Gamma}_{\bm{\Lambda}} > \bm{0}$ for $n\to \infty$. On the other hand the weak common component $(e_{it}^\chi)$ is driven by weak non-pervasive factors, i.e. associated with non-divergent loadings $\sup_{n\in \mathbb N} \left(\bm{\Lambda}^{w, n'} \bm{\Lambda}^{w, n} \right) < \infty$ as pioneered in \cite{onatski2012asymptotics}. All three terms in \eqref{eq: can decomp univar} are contemporaneously orthogonal for all $i \in \mathbb N$. However, $(\xi_{it})$ is orthogonal to $(\bm{F}_t)$ and $(\bm{F}_t^w)$ at all leads and lags, but $(\bm{F}_t)$ and $(e_{it})$ may correlate at leads and lags via $(\bm{F}_t^w)$.
The weak factors $\bm{F}_t^w$ are to be distinguished from the case of \textit{rate-weak factors}, which are excluded here. Such rate-weak factors have loadings diverging at a slower rate, i.e. $\bm{\Lambda}^{w, n'} \bm{\Lambda}^{w, n} / n^\alpha \to \bm{\Gamma}_{\Lambda^w} > \bm{0}$ with $\alpha \in (0, 1)$ \citep[see][]{demol2008forecasting, freyaldenhoven2022factor, bai2023approximate} and would be part of $C_{it}$ in our setup. \cite{barigozzi2026dynamic} show that cross-sectional exchangeability - a natural assumption for factor models - rules out rate weak factors.
To bring the distributed lags representation \eqref{eq: GDFM stat rep} into the form of the canonical decomposition \eqref{eq: canonical rep weak and strong factors} we project out $\bm{F}_t$ from its lags: For this let $\bm{\delta}$ be the population regression parameter from a regression of $\bm{F}_t^- :=(\bm{F}_{t-1}', ..., \bm{F}_{t-p}')'$ onto $\bm{F}_t$ and set $\bm{\beta}_i^-:= (\bm{\beta}_{i1}, ...., \bm{\beta}_{ip})$, then from \eqref{eq: GDFM stat rep}
\begin{align}
y_{it} &= \bm{\beta}_i \bm{x}_t + \xi_{it}
= \underbrace{\left(\bm{\beta}_{i0} + \bm{\beta}_i^-\bm{\delta}\right)}_{\bm{\Lambda}_i} \bm{F}_t + \underbrace{\bm{\beta}_i^-}_{\bm{\Lambda}_i^w}\underbrace{\left(\bm{F}_t^- - \bm{\delta} \bm{F}_t\right)}_{\bm{F}_t^w} + \xi_{it} = \underbrace{\bm{\Lambda}_i \bm{F}_t}_{C_{it}} + \underbrace{\bm{\Lambda}_i^w \bm{F}_t^w}_{e_{it}^\chi} + \xi_{it}. \label{eq: FWL GDFM static rep}
\end{align}
Consequently, we solve \cite{onatski2012asymptotics}'s problem, who shows that principal component based estimates are not consistent in the case of weak factors associated to non-divergent eigenvalues. We can estimate the influence of such weak factors consistently, if they live in the dynamic common space, i.e. are driven by the common shocks $(\varepsilon_t)$ \citep[for a in-depth discussion on the theoretical aspects see][]{gersing2023weak}. In our case the weak factors are simply the influence of lags of strong factors which are estimated with a ``strong signal'' by static principal components. In this paper, we do not focus on the estimation and inference of the weak factors $\bm{F}_t^w$ themselves, but rather on the common components $\chi_{it}$ and $e_{it}^\chi = \chi_{it} - C_{it}$, for which explicit identification of the weak factors is not required.
An interesting interpretation of \eqref{eq: FWL GDFM static rep} is that weak factors are associated to factor loadings that ``taper off'' for lags of strong factors. While the first $r$ factors are pervasive as $\bm{\beta}_0^{n'}\bm{\beta}_0^n \to \infty$, the influence of lags becomes weaker going further in the past as $\bm{\beta}^{-, n'}\bm{\beta}^{-, n} = \bm{\Lambda}^{w, n'}\bm{\Lambda}^{w, n}< \infty$. It is plausible that fewer series, are influenced by information from the past, e.g. if they are more (or less) persistent relative to the majority of series in the panel. Implications to impulse response functions have been studied as well in \cite{gersing2023weak}.
The estimation theory of the GDFM commences with \cite{forni2001generalized, forni2004generalized}, while \cite{forni2005generalized} provide the first one-sided estimator. Infinite dimensional factor spaces, i.e. $\dim \operatorname{\overline{\operatorname{sp}}}(\chi_{it}: i \in \mathbb N) = \infty$, are covered in \cite{forni2015dynamic, forni2017dynamic, barigozzi2024inferential}. In contrast, here we look only at the restricted finite dimensional case.
The contributions of this paper are as follows: First, we establish consistency and derive asymptotic confidence intervals for the dynamic common component, $\chi_{it}$, in the restricted GDFM framework \eqref{eq: GDFM rep}, as well as for its weak common component, $e_{it}^\chi$, in \eqref{eq: can decomp univar}. While regression-based approaches using factors and their lags have been extensively studied commencing with the seminal work of \cite{bai2006confidence, bernanke2005measuring}, the asymptotic framework in this paper is different: In particular, \cite{bai2006confidence} relies on the assumption that $(e_{it})$ and $(\bm{F}_t)$ are mutually independent. The GDFM requires accommodating dependence across leads and lags, because the common shocks $(\varepsilon_t)$ drive both, the weak factors $\bm{F}_t^w$ which are absorbed in $e_{it}$ and the statically pervasive factors $\bm{F}_t$. Hence, we need to allow explicitly for the presence of weak factors in view of equation \eqref{eq: FWL GDFM static rep} in order to provide asymptotic confidence intervals for the dynamic- and weak common component.
Second, we propose a novel mathematical approach to consistency and asymptotic normality of the estimated factors and common components. This approach starts from the assumption of consistent sample covariance estimators for factors and idiosyncratic component, and exploits eigen-value and -vector convergence under an eigen-gap condition, drawing on \cite{yu2015useful, davis1970rotation}, as also employed in \cite{barigozzi2022principal}. It interprets normalised principal components as a static aggregation - in the spirit of \cite{chamberlain1983arbitrage, forni2001generalized, hallin2013factor} - that averages out the weakly cross-correlated idiosyncratic component (see the role of $\bm{\mathcal K}$ in the proofs).
Third, we develop the first asymptotic theory for the GDFM that entirely avoids frequency-domain methods and avoids estimation of the number of dynamic factors $q$ in \eqref{eq: GDFM rep}. Our framework also accommodates time-varying heteroskedasticity in the idiosyncratic component, extending the scope of existing GDFM estimation theory.
Finally, in the empirical application we present statistical evidence for a significant share of weak common components in key variables for three different macroeconomic time series panels of the Euro Area.
\subsection{Notation}
For a real valued matrix $\bm{A} \in \mathbb R^{n \times m}$ we denote by $\bm{A}'$ the transpose. By $\@ifstar{\oldnorm}{\oldnorm*}{\bm{A}}$ we denote the spectral norm, for a vector $\bm{v} \in \mathbb R^n$ we write $\@ifstar{\oldnorm}{\oldnorm*}{\bm{v}}$ to denote the Euclidean norm. We denote by $\mu_j(\bm{A})$ the $j$-th largest eigenvalue of a square matrix $\bm{A}$ and set $\bm{M}(\bm{A}) \equiv \operatorname{diag}(\mu_1(\bm{A}), ..., \mu_r(\bm{A}))$, while $r$ is the number of static factors (see equation \eqref{eq: static factor rep}). By $\bm{P}(\bm{A})$ the $r\times n$ matrix consisting of the first orthonormal row eigenvectors (corresponding to the $r$ largest eigenvalues of $\bm{A}$) and by $\bm{P}^i(\bm{A})$ the $1\times r$ vector consisting of the entries of the $i$-th row of $\bm{P}(\bm{A})'$.
Let $\mathcal P = (\Omega, \mathcal A, \mathbb P)$ be a probability space and $L_2(\mathcal P, \mathbb R)$ be the Hilbert space of square integrable real-valued, zero-mean, random-variables defined on $\Omega$ equipped with the inner product $\langle u, v\rangle = \operatorname{\mathbb E} [u v]$ for $u, v \in L_2(\mathcal P, \mathbb R)$. If $\mathbb M\subset L_2(\mathcal P, \mathbb R)$ is a linear subspace, we denote by $\operatorname{proj}(\bm{u} \mid \mathbb M)$ the orthogonal projection of $\bm{u}$ onto $\mathbb M$ \citep[see e.g.][Theorem 1.2]{deistler2022modelle}.
If $\bm{u}_t, \bm{v}_t$ are stochastic vector processes of zero-mean random variables with fixed dimensions $n_1, n_2$, we denote the $n_1\times n_2$-dimensional covariance matrix by $\operatorname{\mathbb E}\left[\bm{u}_t \bm{v}_{t-h}'\right]=: \bm{\Gamma}_{\bm{u}\bm{v}}(h)$ and the sample covariance by $\hat \bm{\Gamma}_{\bm{u} \bm{v}}(h) := (T-h)^{-1}\sum_{t = h+1}^T \bm{u}_t \bm{v}_{t-h}$ and $\hat \bm{\Gamma}_{\bm{u}\bm{v}}(0)=: \hat \bm{\Gamma}_{\bm{u}\bm{v}}$, $\hat \bm{\Gamma}_{\bm{u}}(-h)' = \hat \bm{\Gamma}_{\bm{u}}(h)$. If either $\bm{u}_t$ or $\bm{v}_t$ is of dimension $n$ with $n\to\infty$ we add a superscript $n$, and write e.g. $\bm{\Gamma}_{\bm{u}\bm{v}}^n$. Also, we abbreviate $\operatorname{\mathbb E}\left[\bm{v}_t \bm{v}_t'\right]:= \bm{\Gamma}_{\bm{v}\bm{v}} = \bm{\Gamma}_{\bm{v}}$. If $\operatorname{\mathbb E}\left[\bm{v}_t \bm{v}_t'\right]$ depends on time we write $\bm{\Gamma}_{\bm{v}_t}(h):=\operatorname{\mathbb E}\left[\bm{v}_t \bm{v}_{t-h}'\right]$ and $\bm{\Gamma}_{\bm{v}_t}:=\operatorname{\mathbb E}\left[\bm{v}_t \bm{v}_t'\right]$. For a stochastic vector $\bm{u}$ with coordinates in $L_2(\mathcal P, \mathbb R)$, we write $\operatorname{\mathbb V} \bm{u} := \operatorname{\mathbb E}\left[ \bm{u} \bm{u}'\right]$ to denote the variance matrix, which is symmetric and always diagonalisable.
We consider real valued stochastic double sequences, i.e. a family of random variables indexed in time and cross-section:
\begin{align*}
(y_{it}: i \in \mathbb N, t \in \mathbb Z) = (y_{it}) \qquad \mbox{where} \ y_{it} \in L_2(\mathcal P, \mathbb R) \quad \forall (i, t) \in \mathbb N \times \mathbb Z .
\end{align*}
Such a process can also be thought of as a nested sequence of multivariate stochastic processes: $\left(\bm{y}_t^n: t \in \mathbb Z \right) = (\bm{y}_t^n)$, where $\bm{y}_t^n = (y_{1t},..., y_{nt})'$ and $\bm{y}_t^{n+1} = (\bm{y}_t^{n'}, y_{n+1, t})'$ for $n \in \mathbb N \cup \{\infty\}$. In general we will write $(\bm{y}_t: t \in \mathbb Z)=(\bm{y}_t)$ for $n = \infty$. We write $\mathbb H(\bm{y})\equiv \operatorname{\overline{\operatorname{sp}}}\left(y_{it}: i \in \mathbb N, t \in \mathbb Z \right)$ for the ``time domain'' of $(y_{it})$ and $\mathbb H_t(\bm{y}) \equiv \operatorname{\overline{\operatorname{sp}}}\left(y_{is}: i \in \mathbb N, s \leq t \right)$ for the infinite past,
where $\operatorname{\overline{\operatorname{sp}}}(\cdot)$ denotes the closure of the linear span.
\section{Structure Theory: Identification from Static Factors}
In this section, we justify why it is reasonable to assume that the GDFM in \eqref{eq: GDFM rep} can be represented as in \eqref{eq: GDFM stat rep}. To elucidate the connection with the GDFM and its difference to the static decomposition, we introduce the classical assumptions. The assumptions on the spectral density are invoked solely for theoretical motivation and are not required for estimation, which proceeds entirely in the time domain and accommodates for time dependent heteroskedasticity in the dynamic idiosyncratic component.
\begin{theoryassumption}[Stationarity]\label{T: stat}
For all $n\in \mathbb N$, the observed process $\left(y_t^n\right)$ is zero mean, weakly stationary, has existing spectral density $f_{y}^n(\theta)$ for $\theta \in [-\pi, \pi]$ defined as the $n\times n$ matrix:
\[
\bm{f}_{y}^n(\theta):=\frac 1{2\pi}\sum_{\ell=-\infty}^{\infty} e^{-\iota \ell \theta}
\operatorname{\mathbb E} [\bm{y}_t^n \bm{y}_{t-\ell}^{n'}],\quad \theta \in[-\pi,\pi].
\]
\end{theoryassumption}
We impose the standard GDFM assumption from \cite{forni2001generalized}:
\begin{theoryassumption}[$q$-Dynamic Factor Structure]\label{T: q-DFS struct}\ \\[-3em]
\begin{itemize}
\item[(i)] $\sup_n \mu_q\left(\bm{f}_{\bm{y}}^n\right) = \infty$ a.e. on $[-\pi, \pi]$;
\item[(ii)] $\operatorname{ess\,sup}_\theta \sup_n \mu_{q+1}(\bm{f}_{\bm{y}}^n)< \infty$,
\end{itemize}
\end{theoryassumption}
\noindent where ``$\operatorname{ess\,sup}$'' denotes the essential supremum of a measurable function. Assumption T\ref{T: q-DFS struct} implies that $(y_{it})$ can be represented in terms of \eqref{eq: GDFM rep} (in general with a two-sided transfer-function) with $(\chi_{it})$ and $(\xi_{it})$ being orthogonal at all leads and lags, with $\sup_n \mu_q\left( \bm{f}_{\bm{\chi}}^n \right) = \infty$ a.e. on $[-\pi, \pi]$ and $\operatorname{ess\,sup}_\theta \sup_n \mu_{1}(\bm{f}_{\bm{\xi}}^n) < \infty$. For the latter property, we say that $(\xi_{it})$ is \textit{dynamically idiosyncratic}. The essential boundedness of the first spectral eigenvalue restricts the cross-correlation and serial correlation in $(\xi_{it})$. Note that under very general conditions there exists a GDFM-representation such that $\varepsilon_t, \chi_{it} \in \mathbb H_t(y)$ for all $i \in \mathbb N, t \in \mathbb Z$ \citep{gersing2026existence}.
In addition, we also assume that $(y_{it})$ has a static factor structure as in \cite{chamberlain1983arbitrage, gersing2023reconciling}:
\begin{theoryassumption}[$r$-Static Factor Structure]\ \\[-3em]\label{T: r-SFS struct}
\begin{itemize}
\item[(i)] $\sup_{n\in \mathbb N} \mu_r\left(\bm{\Gamma}_{\bm{y}}^n\right) = \infty$;
\item[(ii)] $\sup_{n\in \mathbb N} \mu_{r+1}\left(\bm{\Gamma}_{\bm{y}}^n\right) < \infty$.
\end{itemize}
\end{theoryassumption}
which implies the existence of representation \eqref{eq: static factor rep} with $(\bm{F}_t)$ and $(e_{it})$ being \textit{contemporaneously orthogonal}, while correlation at leads and lags is allowed, and $\sup_{n\in \mathbb N} \mu_r \left(\bm{\Gamma}_C^n\right) = \infty$ and $\sup_{n\in \mathbb N} \mu_1 \left(\bm{\Gamma}_e^n\right) = \infty$. For the latter property, we say that $(e_{it})$ is \textit{statically idiosyncratic}.
As shown in \cite{gersing2023weak}, imposing the typical eigenvalue behavior on the spectral density (T\ref{T: q-DFS struct}) vs the variance-covariance matrix (T\ref{T: r-SFS struct}) imply two different types of decompositions which are nested in \eqref{eq: can decomp univar}. In particular it can be shown that the weak common component in \eqref{eq: can decomp univar} is the static idiosyncratic component of the dynamic common component. Consequently, the dynamic common component $\chi_{it}$ in Assumption T\ref{T: q-DFS struct} retains in general more variation of the output compared to the static common component $C_{it}$ which is due, as argued below, to the feedback in $(y_{it})$ to lags of the static factors $\bm{F}_t$.
Recall that as shown in \cite{forni2001generalized}, the dynamic common component can be represented as the orthogonal Hilbert space projection on the space spanned by the common shocks. In view of the one-sided representation \eqref{eq: GDFM rep} \citep[see][for details]{gersing2026existence}, it is enough to project on the infinite past of what we may call the structural common shocks $(\bm{\varepsilon}_t)$ of the economy:
\begin{align*}
\chi_{it} = \operatorname{proj}\left(y_{it} \mid \mathbb H_t(\bm{\varepsilon})\right).
\end{align*}
Furthermore, since the static factors $(\bm{F}_t)$ of $(y_{it})$ are also the static factors of $(\chi_{it})$, we have $\operatorname{sp}\left(\bm{F}_t\right)\subseteq \operatorname{\overline{\operatorname{sp}}}\left(\chi_{it}: i \in \mathbb N\right) \subset \mathbb H_t(\bm{\varepsilon})$, so both $(\chi_{it})$ and $(\bm{F}_t)$ are completely driven by $(\bm{\varepsilon}_t)$. Now, it is natural to assume that the same common shocks which drive the whole economy, also drive the static factors $(\bm{F}_t)$ and are fundamental for them. This assumption has been implicitly imposed literally everywhere in the literature where the common shocks $(\bm{\varepsilon}_t)$ in combination with $(\bm{F}_t)$ play a role \cite{bernanke2005measuring, bai2007determining, stock2011theoxford, forni2009opening, forni2025common} among many others.
Formally, we require that there is exists a representation of the form:
\begin{align*}
\bm{F}_t = \sum_{j = 0}^\infty \bm{B}_F(j) \bm{\varepsilon}_{t-j} = \ubar \bm{b}_F(L) \bm{\varepsilon}_t,
\end{align*}
where $\ubar \bm{b}_F(L)$ is a causal transferfunction that has a causal left inverse, say $\ubar \bm{c}_F (L) = \ubar \bm{b}_F^{\dag}(L)$, while ``\dag'' denotes the generalised inverse, which is of dimension $q \times r$ such that
\begin{align*}
\bm{\varepsilon}_t = \sum_{j = 0}^\infty \bm{C}_{\bm{F}} (j) \bm{F}_{t-j} = \ubar \bm{c}_{\bm{F}}(L) \bm{F}_t.
\end{align*}
In this case we can retrieve the infinite past of the common shocks from the infinite past of the static factors. Therefore, in order for our procedure to work, we must impose the following assumption:
\begin{theoryassumption}[Identification from Static Factors]\ \\[-3em] \label{A: ident from static fac}
\begin{itemize}
\item[(i)] (Fundamentalness of the Static Factors) The common shocks are fundamental for the statically pervasive factors,
i.e.
\begin{align*}
\mathbb H_t(\bm{F}) = \mathbb H_t(\bm{\varepsilon})
\end{align*}
\item[(ii)] (Truncation at Finite Lag) For every $i$ the projection of $y_{it}$ on $\mathbb H_t(\bm{F})$ can be expressed by at most $p$ lags while $\bm{\Gamma}_{\bm{x}} = \operatorname{\mathbb E}\left[\bm{x}_t \bm{x}_t'\right] > 0$:
\begin{align*}
\chi_{it} = \operatorname{proj}\left(y_{it} \mid \mathbb H_t(\bm{\varepsilon})\right) = \operatorname{proj}\left(y_{it} \mid \mathbb H_t(\bm{F})\right) = \bm{\beta}_{i0} \bm{F}_t + \cdots + \bm{\beta}_{ip} \bm{F}_{t-p}
\end{align*}
\end{itemize}
\end{theoryassumption}
To further explore the role of T\ref{A: ident from static fac}, the prototypical ```dynamic factor model'' as e.g. in \citet{bai2007determining, stock2011theoxford}:
\begin{align}
y_{it} &= \bm{\lambda}_{i0}\bm{f}_t + \cdots \bm{\lambda}_{ip_\lambda} \bm{f}_{t-p_\lambda} + \xi_{it} = \ubar \bm{\lambda}_i(L) \bm{f}_t + \xi_{it} = \bar \bm{\Lambda}_i \bar \bm{F}_t + \xi_{it} \label{eq: SW dfm} \\
\bm{f}_t &= \bm{A}_1 \bm{f}_{t-1} + ... + \bm{A}_{p_f} \bm{f}_{t-p_f} + \bm{b} \bm{\varepsilon}_t \label{eq: SW dfm dynfac} \quad \bm{b} \bm{\varepsilon}_t \sim WN(\bm{I}_q)
\end{align}
where $\bm{f}_t$ are $q\times 1$ dimensional ``dynamic factors'' stacked in $\bar \bm{F}_t = (\bm{f}_t', ..., \bm{f}_{t-p_\lambda}')$ with loadings $\bar \bm{\Lambda}_i = (\bm{\lambda}_{i1}, ..., \bm{\lambda}_{ip_\lambda})$ and modelled as an autoregressive process. Setting $\ubar \bm{a}_f(L) = \bm{I}_q - \bm{A}_1 L - ... - \bm{A}_{p_f} L^{p_f}$, where $L$ is the lag operator, we can write \eqref{eq: SW dfm}, \eqref{eq: SW dfm dynfac} in terms of the GDFM representation as
\begin{align*}
\chi_{it} = \ubar \bm{\lambda}_{i}(L) \ubar \bm{a}_f^{-1}(L) \bm{\varepsilon}_t = \ubar \bm{B}_i(L) \bm{\varepsilon}_t,
\end{align*}
while the characteristic eigenvalue behaviour of the GDFM is satisfied if at least one column block of the loadings has $q$ divergent eigenvalues \cite[for a proof see][]{gersing2023weak}.
As opposed to what is commonly imposed in the literature, this model - though finite dimensional - can in general not be estimated by means of static principal components (through stacking the $\bm{f}_t$'s) but has weak non-pervasive factors. Considering the vector representation of the dynamic common component $\bm{\chi}_t^n = \bar \bm{\Lambda}^{n} \bar \bm{F}_t$, in general $r < r_\chi = q(p+1)$ eigenvalues of $\bar\bm{\Lambda}^{n'} \bar\bm{\Lambda}^n$ diverge as $n\to \infty$, which implies that the weak common component in \ref{eq: canonical rep weak and strong factors} is non-trivial and there exists $\bm{T} \in \mathbb R^{r\times r_\chi}$ such that $\bm{F}_t = \bm{T} \bar \bm{F}_t$ in \eqref{eq: static factor rep} \citep[for details see][]{gersing2023reconciling}.
Rewrite \eqref{eq: SW dfm}, \eqref{eq: SW dfm dynfac} as a state space system
\begin{align}
\bm{F}_t &= \bm{T} \bar \bm{F}_t \label{eq: obs eq Ft} \\
\underbrace{\begin{pmatrix}
\bm{f}_{t+1} \\
\bm{f}_t \\
\vdots \\
\bm{f}_{t-p+2}
\end{pmatrix}}_{\bar \bm{F}_{t+1}} &=
\underbrace{\begin{bmatrix}
\bm{A}_1 & \bm{A}_2 & \cdots & \bm{A}_p \\
\bm{I}_q & & & \bm{0} \\
& \ddots & & \\
& & \bm{I}_q & \bm{0}
\end{bmatrix}}_{\bm{A}}
\underbrace{\begin{pmatrix}
\bm{f}_t \\
\bm{f}_{t-1} \\
\vdots \\
\bm{f}_{t-p+1}
\end{pmatrix}}_{\bar \bm{F}_t} +
\underbrace{\begin{bmatrix}
\bm{b} \\
\bm{0} \\
\vdots \\
\bm{0}
\end{bmatrix}}_{\bm{G}} \bm{\varepsilon}_{t+1} \nonumber \\[0.8em]
\bar \bm{F}_{t+1} &= \bm{A} \bar \bm{F}_t + \bm{G} \bm{\varepsilon}_{t+1}. \label{eq: trans eq Ft}
\end{align}
The transferfunction of $(\bm{F}_t)$ is
\begin{align*}
\ubar \bm{B}_{\bm{F}}(L) &= \bm{T} \left(\bm{I}_{r_\chi} - \bm{A} L\right)^{-1} \bm{G} \bm{\varepsilon}_t \\
\mbox{and} \ \bm{F}_t &= \sum_{j = 0}^\infty \bm{B}_{\bm{F}}(j) \bm{\varepsilon}_{t-j} = \sum_{j = 0}^\infty \bm{T}\bm{A}^j \bm{G} \bm{\varepsilon}_{t-j}.
\end{align*}
Now consider the parameter space consisting of parameters $(\bm{A}_1, ..., \bm{A}_p, \bm{T}, \bm{b})$ where the autoregressive parameters are stable, i.e. $\det \ubar \bm{a}_f(z) \neq \bm{0}$, $\bm{T}$ is free and $\bm{b}$ is a lower triangular matrix of full rank describing the $q(q-1) /2$ free parameters of the innovation variance matrix of for all $\@ifstar{\oldabs}{\oldabs*}{z} \leq 1$.
\begin{theorem}\label{thm: identification conditions ft AR}
Consider the set of minimal stable systems $(\bm{A}_1, ..., \bm{A}_p, \bm{b}, \bm{T}) \equiv (\bm{A}, \bm{G}, \bm{T})$, then $(\bm{\varepsilon}_t)$ is fundamental for $(\bm{F}_t)$ in the following cases:
\begin{enumerate}
\item For $r = q$ if $\det\left[\ubar\bm{t}(z)\right] \neq 0$ for all $\@ifstar{\oldabs}{\oldabs*}{z} < 1$, where $\ubar\bm{t}(z) = \bm{T} \ubar\bm{S}(z)$ with $\ubar\bm{S}(z) := \left(\bm{I}_q, \bm{I}_qz, ..., \bm{I}_q z^{p-1}\right)'$.
\item For $r > q$ generically.
\end{enumerate}
\end{theorem}
For the special case that $r = q$, Theorem \ref{thm: identification conditions ft AR}.1 establishes A \ref{A: ident from static fac}(i) under a minimum phase condition, which standard in time series analysis \citep[][]{deistler2022modelle}. For $r > q$, Theorem \ref{thm: identification conditions ft AR}.2, guarantees A\ref{A: ident from static fac}(i) for a generic (open and dense) set in parameter space, which implies that the set of parameters where this condition does not hold has measure zero.
\section{Estimation}\label{sec: estimation}
We assume to work with pre-centered data. If data is not pre-centered then all the following applies to the centered realizations $(y_{it}-T^{-1}\sum_{t=1}^T y_{it} : i=1,\ldots, n,\, t=1,\ldots, T)$. Let $\operatorname{\mathbb E}\left[\bm{y}_t^n \bm{y}_t^{n'}\right] =:\bm{\Gamma}_{\bm{y}}^n$ with $\bm{y}_t^n = (y_{1t}, ..., y_{nt})'$.
The normalised principal components of $\bm{y}_t^n$ are
\begin{align*}
\bm{W}_t^{y, n} &:= \bm{M}^{-1/2}(\bm{\Gamma}_{\bm{y}}^n) \bm{P}(\bm{\Gamma}_{\bm{y}}^n) \bm{y}_t^n = \bm{\mathcal K}(\bm{\Gamma}_{\bm{y}}^n) \bm{y}_t^n \\
\bm{W}_t^{C, n} &:= \bm{M}^{-1/2}(\bm{\Gamma}_{\bm{C}}^n) \bm{P}(\bm{\Gamma}_{\bm{C}}^n) \bm{C}_t^n = \bm{\mathcal K}(\bm{\Gamma}_{\bm{C}}^n) \bm{y}_t^n \\
\widehat{\bm{W}}_t^{y, n} &:= \bm{M}^{-1/2}(\hat \bm{\Gamma}_{\bm{y}}^n) \bm{P}(\hat \bm{\Gamma}_{\bm{y}}^n) \bm{y}_t^n = \bm{\mathcal K}(\hat \bm{\Gamma}_{\bm{y}}^n) \bm{y}_t^n = \hat{\bm{\mathcal K}} \bm{y}_t^n.
\end{align*}
Clearly the normalised principal components of $\bm{y}_t^n$ are determined only up to sign. To resolve this sign indeterminacy, we assume (without loss of generality by changing the cross-sectional ordering) that the first $r$ rows of $\bm{P}'(\bm{\Gamma}_{\bm{y}}^n) \bm{M}^{1/2}(\bm{\Gamma}_{\bm{y}}^n)$ have full rank (from a certain $n$ onwards) and fix the diagonal elements to be positive. This fixes the sign of the eigenvectors and the normalised principal components. We use $\hat{\bm{\mathcal K}}_j$ is the $j$-th row of $\hat{\bm{\mathcal K}} := \bm{M}^{-1/2}(\hat \bm{\Gamma}_{\bm{y}}^n) \bm{P}(\hat \bm{\Gamma}_{\bm{y}}^n)$ and $\bm{\mathcal K}_j$ is the $j$-th row of $\bm{\mathcal K} := \bm{M}^{-1/2}(\bm{\Gamma}_{\bm{y}}^n) \bm{P}(\bm{\Gamma}_{\bm{y}}^n)$.
\subsection{Consistency of Factor- and Loadings-Space}
We employ the following assumptions for estimation which are in line with Assumptions T\ref{T: stat}-T\ref{T: r-SFS struct}, but in addition impose rates on the divergence, allow for heteroscedasticity over time in the dynamic idiosyncratic component and don't require the existence of the spectral density.
\begin{assumption}[Asymptotic Properties of the Loadings]\label{A: divergence rates eval}
We assume that model \eqref{eq: GDFM stat rep} holds, with $(\bm{F}_t)$ being a weakly stationary, zero mean process and orthogonal to $(\xi_{it})$ at all leads and lags. Furthermore $(\bm{\beta}_i: i \in \mathbb N)$ are deterministic loadings, while $\left(\bm{\Lambda}_i: i \in \mathbb N\right), \left(\bm{\Lambda}_i^w: i \in \mathbb N\right)$ are obtained in \eqref{eq: FWL GDFM static rep} with the following features:
\begin{itemize}
\item[(i)] Normalised Representation: We assume that $\operatorname{\mathbb E}\left[\bm{x}_t \bm{x}_t'\right] = \bm{\Gamma}_{\bm{x}} > \bm{0}$ and without loss of generality, we employ the normalisation $\operatorname{\mathbb E}\left[\bm{F}_t \bm{F}_t'\right] = \bm{\Gamma}_{\bm{F}} = \bm{I}_r > \bm{0}$ and $\operatorname{\mathbb E}\left[\bm{F}_t^w \bm{F}_t^{w'}\right] = \bm{\Gamma}_{\bm{F}^w} = \bm{I}_{r^w} > \bm{0}$;
\item[(ii)] Convergence of the ``loadings variance'': Set $n^{-1}\bm{\Lambda}^{n'}\bm{\Lambda}^n = n^{-1} \sum_{i = 1}^n \bm{\Lambda}_i' \bm{\Lambda}_i := \bm{\Gamma}_{\bm{\Lambda}}^n$ and suppose that $\@ifstar{\oldnorm}{\oldnorm*}{\bm{\Gamma}_{\bm{\Lambda}}^n- \bm{\Gamma}_{\bm{\Lambda}}} = \mathcal O(n^{-1/2})$ for some $\bm{\Gamma}_{\bm{\Lambda}} > \bm{0}$;
\item[(iii)] The eigenvalues of $\bm{\Gamma}_{\bm{\Lambda}}$ are distinct and contained in the diagonal matrix $\bm{D}_{\bm{\Lambda}}$ sorted from the largest to the smallest;
\item[(iv)] There exists $\mathcal B_\Lambda < \infty$, such that $\@ifstar{\oldnorm}{\oldnorm*}{\bm{\Lambda}_i} < \mathcal B_\Lambda$ and $\@ifstar{\oldnorm}{\oldnorm*}{\bm{\Lambda}_i^w} < \mathcal B_{\Lambda^w}$;
\item[(v)] There exists $\mathcal B_\xi$, such that $\operatorname{\mathbb E} \xi_{it}^2 < \mathcal B_\xi$ for all $i\in \mathbb N$;
\item[(vi)] Global bounds: $\sup_{t \in \mathbb Z} \sup_{n\in \mathbb N}\mu_1\left(\bm{\Gamma}_{\bm{\xi}_t}^n\right) < \mathcal B_\xi$ and $\sup_{n\in \mathbb N}\mu_1\left(\bm{\Lambda}^{w, n'} \bm{\Lambda}^{w, n}\right) < \mathcal B_{\Lambda^w}$.
\end{itemize}
\end{assumption}
Assuming that $\bm{\Gamma}_{\bm{x}} > \bm{0}$ is a simplification and excludes the case where the stack $\bm{x}_t$ has a singular variance-covariance matrix \citep[see][]{anderson2016structure, deistler2019singular}. Without loss of generality for the proofs one may use a sub-vector of $\bm{x}_t$ with full rank variance-covariance matrix instead.
First note that the assumption $\bm{\Gamma}_{\bm{F}} = \bm{I}_r$ is without loss of generality for the results in this paper (see remark \ref{rem: Gamma F = I} in the appendix for details) and is made for simplicity. A\ref{A: divergence rates eval}(ii) implies T\ref{T: r-SFS struct}(i) together with a rate of divergence. Assumption A\ref{A: divergence rates eval}(vi) allows that dynamic idiosyncratic variance to be time dependent but globally bounded in the first eigenvalue and imposes furthermore the non-pervasiveness of the weak factors.
In the following we use the eigen-decomposition $\bm{\Gamma}_{\bm{\Lambda}}^n := \bm{P}_{\bm{\Lambda}}^{n'} \bm{D}_{\bm{\Lambda}}^n \bm{P}_{\bm{\Lambda}}^n$ with eigenvalues sorted from the largest to the smallest in the diagonal matrix $\bm{D}_{\bm{\Lambda}}^n$ and $\bm{P}_{\bm{\Lambda}}^n$ being a matrix of orthonormal row eigen-vectors. Analogously we use $\bm{\Gamma}_{\bm{\Lambda}} = \bm{P}_{\bm{\Lambda}}' \bm{D}_{\bm{\Lambda}} \bm{P}_{\bm{\Lambda}}$. Assumption A\ref{A: divergence rates eval} implies also that $\sup_{t\in \mathbb Z} \sup_{n\in \mathbb N} \mu_1 \left( \bm{\Gamma}_{\bm{e}_t}^n \right) \leq \mathcal B_\xi +\mathcal B_{\Lambda^w}=: \mathcal B_e < \infty$.
Let $\bm{\eta}_t^n := n^{-1/2}\bm{\Lambda}^{n'} \bm{\xi}_t^n$ be the $r\times 1$ dimensional aggregate of the idiosyncratic component obtained from averaging with the static factor loadings. Note that by Assumption A\ref{A: divergence rates eval}, it holds that $\@ifstar{\oldnorm}{\oldnorm*}{\operatorname{\mathbb E}\left[\bm{\eta}_t^n \bm{\eta}_t^{n'}\right]} < \mathcal B_\xi \@ifstar{\oldnorm}{\oldnorm*}{n^{-1/2}\bm{\Lambda}^n}< \mathcal B_\xi \mathcal B_\Lambda$.
\begin{assumption}[Sample Covariances]\label{A: simply the sample covariances}
We suppose that the sample covariances can be estimated consistently (up to heteroskedasticity) as
\begin{itemize}
\item[(i)] $\operatorname{\mathbb E}\left[\@ifstar{\oldnorm}{\oldnorm*}{\sqrt{T}\left(\hat \bm{\Gamma}_{\bm{F}}(h) - \bm{\Gamma}_{\bm{F}}(h)\right)}^2\right] \leq \mathcal B_F$ for all $\@ifstar{\oldabs}{\oldabs*}{h} <\infty$;
\item[(ii)] $\operatorname{\mathbb E}\left[\left(\frac{1}{\sqrt{T}}\sum_{t = 1}^T \left\{\xi_{it}\xi_{is} - \operatorname{\mathbb E}\left[\xi_{it} \xi_{is}\right]\right\} \right)^2\right] \leq \mathcal B_\xi $ independent of $i, j \in \mathbb N$, $t, s\in \mathbb Z$;
\item[(iii)] $\operatorname{\mathbb E}\left[\@ifstar{\oldnorm}{\oldnorm*}{\frac{1}{\sqrt{T}} \sum_{t = 1}^T \left\{\bm{\eta}_t^n \bm{\eta}_{t-h}^{n'} - \operatorname{\mathbb E}\left[\bm{\eta}_t^n \bm{\eta}_{t-h}^{n'}\right]\right\}}^2\right] \leq \mathcal B_\xi$ for all $\@ifstar{\oldabs}{\oldabs*}{h} < \infty$;
\item[(iv)] $\operatorname{\mathbb E}\left[\@ifstar{\oldnorm}{\oldnorm*}{\sqrt{T}\hat \bm{\Gamma}_{\bm{F} \bm{\xi}}(h)\bm{s}_i}^2\right] \leq \mathcal B_{F\xi} / T$ independent of $i \in \mathbb N$ and $\@ifstar{\oldabs}{\oldabs*}{h} <\infty$;
\item[(v)]
$\operatorname{\mathbb E}\left[\@ifstar{\oldnorm}{\oldnorm*}{\frac{1}{\sqrt{T}}\sum_{t = 1}^T \bm{\eta}_t^n \bm{F}_{t-j}}^2\right] \leq \mathcal B_{F\xi}$ for $\@ifstar{\oldabs}{\oldabs*}{h} <\infty$.
\end{itemize}
\end{assumption}
These conditions establish mean square convergence of the sample covariances. Note that in the standard setup of \cite{chamberlain1983arbitrage}, implied by Assumption T\ref{T: r-SFS struct}, the variance of any cross-sectional average (scaled by $\sqrt{n}$), like $\bm{\eta}_t^n$ is bounded in $n$. Consider an aggregate of dynamic idiosyncratic terms with $\bm{\eta}_t^n := p^{n'}\bm{\xi}_t^n$ with any $p\in \mathbb R^n$, $\@ifstar{\oldnorm}{\oldnorm*}{p}<\infty$ for all $n \in \mathbb N$. Then $\operatorname{\mathbb E}\left[(\bm{\eta}_t^n)^2\right]\leq \sup_{n\in \mathbb N}\@ifstar{\oldnorm}{\oldnorm*}{p^n}\mathcal B_\xi$ independent of $n \in \mathbb N$.
\begin{theorem}[Consistency of Factors' and Loadings' Spaces and Common Components]\label{thm: consistency of spaces and CCs}
Under Assumptions T\ref{T: r-SFS struct} - A\ref{A: simply the sample covariances}, with $\widehat{\bm{H}}= \frac{1}{T}\sum_{t = 1}^T \hat \bm{F}_t \bm{F}_t' \left(\frac{1}{T} \sum_{t = 1}^T \bm{F}_t \bm{F}_t'\right)^{-1}$ it follows that
\begin{itemize}
\item[(i)] $\@ifstar{\oldnorm}{\oldnorm*}{\widehat{\bm{W}}_t^{y, n} - \widehat{\bm{H}}\bm{F}_t} = \mathcal O_{P}(\max(n^{-1/2}, T^{-1/2}))$ and $\@ifstar{\oldnorm}{\oldnorm*}{\hat \bm{x}_t - \widehat{\bm{\mathcal H}} \bm{x}_t} = \mathcal O_{P}(\max(n^{-1/2}, T^{-1/2}))$;
\item[(ii)] $\@ifstar{\oldnorm}{\oldnorm*}{\hat \bm{\Lambda}_i - \bm{\Lambda}_i \hat{\bm{H}}^{-1}} = \mathcal O_{P}(\max(n^{-1/2}, T^{-1/2}))$ and $\@ifstar{\oldnorm}{\oldnorm*}{\hat \bm{\beta}_i - \bm{\beta}_i \widehat{\bm{\mathcal H}}^{-1}} = \mathcal O_{P}(\max(n^{-1/2}, T^{-1/2}))$;
\item[(iii)] $\@ifstar{\oldnorm}{\oldnorm*}{\hat C_{it} - C_{it}} = \mathcal O_{P}(\max(n^{-1/2}, T^{-1/2}))$, $\@ifstar{\oldnorm}{\oldnorm*}{\hat \chi_{it} - \chi_{it}} = \mathcal O_{P}(\max(n^{-1/2}, T^{-1/2}))$ and $\@ifstar{\oldnorm}{\oldnorm*}{\hat e_{it}^\chi - e_{it}^\chi} = \mathcal O_{P}(\max(n^{-1/2}, T^{-1/2}))$ with $\hat e_{it}^\chi = \hat \chi_{it} - \hat C_{it}$.
\end{itemize}
\end{theorem}
\subsection{Asymptotic Normality}
To establish asymptotic normality we need slightly more restrictive assumptions which refer mostly to further limiting the serial- and cross-correlation in the idiosyncratic component.
\begin{assumption}[Central Limit Theorem]\label{A: CLTs}
Define $\bm{\Xi}_t^n:= (\bm{\xi}_t^{n'}, ..., \bm{\xi}_{t-p}^{n'})'$
\begin{itemize}
\item[(i)] $T^{-1/2} \bm{x}'\bm{\xi}^i \Rightarrow \mathcal N\left(0, \bm{\Omega}_{\bm{x}\bm{\xi}}(i)\right)$ and $\bm{\Gamma}_{\bm{x}} > \bm{0}$ and $T^{-1/2} \bm{F}' \bm{e}^i \Rightarrow \mathcal N\left(\bm{0}, \bm{\Omega}_{\bm{F}\bm{e}}(i)\right)$;
\item[(ii)] $\left(\bm{I}_{p+1} \otimes n^{-1/2}\bm{\Lambda}^{n'} \right) \bm{\Xi}_t^n \Rightarrow \mathcal N\left(\bm{0}, \bm{\Theta}_{\bm{\Lambda}\bm{\Xi}}(t)\right)$ and $n^{-1/2}\bm{\Lambda}^{n'} \bm{e}_t^n \Rightarrow \mathcal N\left(\bm{0}, \bm{\Theta}_{\bm{\Lambda} \bm{e}}(t)\right)$;
\item[(iii)] The asymptotic covariances are given by
\begin{align*}
&\lim_{T\to \infty} \operatorname{\mathbb E}\left[\left(T^{-1/2}\bm{x}'\bm{\xi}_t^n\right) \left(T^{-1/2}\bm{F}'\bm{e}_t^n\right)'\right] = \bm{\Omega}_{\bm{x}\bm{\xi}, \bm{F}\bm{e}}(i)\\
\mbox{and}\quad &\lim_{n \to \infty}\operatorname{\mathbb E}\left[\left(\left(\bm{I}_{p+1} \otimes n^{-1/2}\bm{\Lambda}^{n'} \right) \bm{\Xi}_t^n\right)\left(n^{-1/2}\bm{\Lambda}^{n'} \bm{e}_t^n\right)'\right] = \bm{\Theta}_{\bm{\Lambda}\bm{\Xi}, \bm{\Lambda} \bm{e}}(t),
\end{align*}
while $\bm{\Theta}_{\bm{\Lambda}\bm{\Xi}, \bm{\Lambda} \bm{e}}(t) = \bm{\Theta}_{\bm{\Lambda}\bm{\Xi}, \bm{\Lambda}\bm{\xi}}(t)$ due to orthogonality relations.
\end{itemize}
\end{assumption}
\begin{assumption}[Idiosyncratic Auto-Covariance]\label{A: more restrictive sample covarainces}
We reinforce the assumptions on the idiosyncratic component by
\begin{itemize}
\item[(i)]
Summability in the cross-section: Let
\begin{align*}
&\sup_{n\in \mathbb N} \max_{1 \leq i \leq n}\sum_{j = 1}^n \@ifstar{\oldabs}{\oldabs*}{\operatorname{\mathbb E}[\xi_{it}\xi_{j,t-k}]} < \mathcal B_\xi, \quad \mbox{and} \ \sup_{n\in \mathbb N} \max_{1 \leq j \leq n}\sum_{i = 1}^n \@ifstar{\oldabs}{\oldabs*}{\operatorname{\mathbb E}[\xi_{it}\xi_{j,t-k}]} < \mathcal B_\xi \ \mbox{for all} \ t, k \in \mathbb Z,
\end{align*}
which implies that $\bm{\mathcal K} \bm{\Gamma}_{\bm{\xi}}^n(h) = \mathcal O(n^{-1})$;
\item[(ii)] $\bm{\mathcal K} \bm{\Lambda}^{w, n} = \mathcal O(n^{-1})$ or $\bm{\Lambda}^{n'} \bm{\Lambda}^{w, n} = \mathcal O(1)$ which is implied by $\@ifstar{\oldnorm}{\oldnorm*}{\sum_{i = 1}^n \bm{\Lambda}_i^{w, n}} < \mathcal B_{\Lambda^w}$;
\item[(iii)] For all $1\leq j \leq n$, $1 \leq s \leq T$ and $i, t, n, T \in \mathbb N$
\begin{align*}
\operatorname{\mathbb E}\left[\@ifstar{\oldabs}{\oldabs*}{\frac{1}{\sqrt{nT}} \sum_{i = 1}^n \sum_{t = 1}^T \left\{\xi_{is} \xi_{jt} - \operatorname{\mathbb E}[\xi_{is}\xi_{jt}]\right\} }^2\right] < \mathcal B_\xi;
\end{align*}
\item[(iv)] EITHER $(\xi_{it})$ and $(\bm{F}_t)$ are independent OR
for all $t, n, T \in \mathbb N$ and $\@ifstar{\oldabs}{\oldabs*}{h} < \infty$,
\begin{align*}
&\operatorname{\mathbb E}[\xi_{it} \xi_{j, t-k}] \leq \@ifstar{\oldabs}{\oldabs*}{\rho}^k \mathcal B_\xi \ \mbox{for some} \ 0 < \rho < 1, \\
&\operatorname{\mathbb E}\@ifstar{\oldnorm}{\oldnorm*}{
\frac{1}{\sqrt{nT}} \sum_{i = 1}^n \sum_{s = 1}^T \bm{F}_s \left[\xi_{is} \xi_{it} - \operatorname{\mathbb E}\left[\xi_{is} \xi_{it}\right]\right]}^2 < \mathcal B_\xi, \quad \mbox{and} \ \operatorname{\mathbb E}\@ifstar{\oldnorm}{\oldnorm*}{\sum_{t = 1}^T \bm{F}_s \rho^{\@ifstar{\oldabs}{\oldabs*}{s-t}}} < \mathcal B_F.
\end{align*}
\end{itemize}
\end{assumption}
Assumption A\ref{A: more restrictive sample covarainces} is similar to the conditions used in \cite{barigozzi2022principal} who provides an excellent overview on how the usual Assumptions of the factor model literature are nested in Assumption A\ref{A: more restrictive sample covarainces}. Whereas for consistency, we only need $L^2$ boundedness of dynamic and static idiosyncratic variance, for asymptotic normality we need $L^1$ boundedness reflected in A\ref{A: more restrictive sample covarainces}(i), (ii). Furthermore, additional bounds are needed that restrict serial dependence of the dynamic idiosyncratic part A\ref{A: more restrictive sample covarainces}(iii), (iv).
\begin{theorem}[Asymptotic Normality of the $\hat \bm{\beta}_i$]\label{thm: asymptotic normality of hat beta}
Under Assumptions T\ref{T: q-DFS struct}-A\ref{A: more restrictive sample covarainces}, as $\sqrt{T} / n \to 0$, with $\widehat{\bm{\mathcal H}} = T^{-1}\sum_{t = 1}^T \hat \bm{x}_t \bm{x}_t' \left(T^{-1}\sum_{t = 1}^T \bm{x}_t \bm{x}_t'\right)^{-1}$, we have
\begin{align*}
\sqrt{T} \left(\hat \bm{\beta}_i - \bm{\beta}_i \widehat{\bm{\mathcal H}}^{-1}\right) \Rightarrow \mathcal N\left(\bm{0}, asy\bm{\Gamma}_{\hat \bm{\beta}_i}\right)
\end{align*}
The asymptotic variance $asy\bm{\Gamma}_{\hat \bm{\beta}_i}$ is given by
\begin{align}
&asy\bm{\Gamma}_{\hat \bm{\beta}_i} := \bm{\Gamma}_{\bm{x}}^{-1} \left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}}(i) \left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}'\right) \bm{\Gamma}_{\bm{x}}^{-1} .
\end{align}
\end{theorem}
Following \cite{bai2003inferential, bai2006confidence, barigozzi2022principal} for estimating the middle term, robust to heteroskedasticity, we may use either one of the following:
\begin{align}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}}(i)\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right)'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}}(i)\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right)'}{\tmpbox}
&= \frac{1}{T}\sum_{t = 1}^T \hat \xi_{it}^2\hat \bm{x}_t \hat \bm{x}_t' \label{eq: HAC1 asy beta}\\
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}}(i)\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right)'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}}(i)\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right)'}{\tmpbox}
&=\frac{1}{T}\sum_{t = 1}^T \sum_{s = 1}^T \hat \bm{x}_t \hat \xi_{it} \hat \xi_{is}\widehat{\bm{x}}_s' \kappa(t, s) \label{eq: HAC2 asy beta}
\end{align}
where $\kappa(t, s)$ is a suitable kernel with bandwidth $M_T$.
The final estimator for the asymptotic variance is then given by
\begin{align*}
asy\hat \bm{\Gamma}_{\hat \bm{\beta}_i} = \left(\frac{1}{T}\sum_{t = 1}^T \hat \bm{x}_t \hat \bm{x}_t'\right)^{-1}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}}(i)\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right)'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}}(i)\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right)'}{\tmpbox}
\left(\frac{1}{T}\sum_{t = 1}^T \hat \bm{x}_t \hat \bm{x}_t'\right)^{-1}.
\end{align*}
\begin{theorem}[Asymptotic Normality of the Factors]\label{thm: asymptotic normality of hat bsxt}
Under Assumptions T\ref{T: q-DFS struct}-A\ref{A: more restrictive sample covarainces}, as $\sqrt{n} / T \to 0$, we have
\begin{align*}
\sqrt{n}\left(\hat \bm{x}_t - \widehat{\bm{\mathcal H}} \bm{x}_t\right) \Rightarrow \mathcal N(\bm{0}, asy\bm{\Gamma}_{\hat \bm{x}_t})
\end{align*}
The asymptotic variance $asy\bm{\Gamma}_{\hat \bm{x}_t}$ is given by
\begin{align*}
asy\bm{\Gamma}_{\hat \bm{x}_t} := \left(\bm{I}_{p+1}\otimes \bm{D}_{\bm{\Lambda}}^{-1} \bm{P}_{\bm{\Lambda}}\right) \bm{\Theta}_{\bm{\Lambda}\bm{\Xi}}(t)\left(\bm{I}_{p+1}\otimes \bm{P}_{\bm{\Lambda}}' \bm{D}_{\bm{\Lambda}}^{-1}\right).
\end{align*}
\end{theorem}
Following \cite{bai2003inferential, bai2006confidence}, under cross-sectional independence, we may use:
\begin{align}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, h}(t)\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, h}(t)\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
&= \frac{1}{n}\sum_{i = 1}^n \hat \bm{\Lambda}_i' \hat \bm{\Lambda}_i \hat \xi_{it}\hat \xi_{i,t-h},
\end{align}
which allows heteroskedasticity over time.
On the other hand if $(\bm{\xi}_t^n)$ is stationary, we may account for cross-sectional correlation by \cite{bai2006confidence}
\begin{align*}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi},h}\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi},h}\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
&= \frac{1}{n}\sum_{i = 1}^n \sum_{j = 1}^n \hat \bm{\Lambda}_i' \hat \bm{\Lambda}_j \frac{1}{T - h} \sum_{t = h+1}^{T} \hat \xi_{it}\hat \xi_{j,t-h},
\end{align*}
or \cite{fresoli2024dealing}:
\begin{align}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda} \bm{\xi}, h}\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda} \bm{\xi}, h}\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
&= \frac{1}{n}\sum_{i = 1}^n \sum_{j = 1}^n \hat \bm{\Lambda}_i' \hat \bm{\Lambda}_j \frac{1}{T - h} \sum_{t = h+1}^{T} \hat \xi_{it}\hat \xi_{j,t-h} I\left(\@ifstar{\oldabs}{\oldabs*}{\hat \sigma_{ij, h}^\xi} \geq c_{ij, h}\right) \label{eq: est cov Lambda xi, h}
\end{align}
with $\hat \sigma_{ij, h}^\xi := (T-h)^{-1}\sum_{t = h+1}^T \hat \xi_{it} \hat \xi_{j,t-h}$, $\hat{\operatorname{\mathbb V}}\left[\hat \xi_{it} \hat \xi_{j,t-h}\right] := (T-h)^{-1}\sum_{t = h+1}^T \left[\hat \xi_{it} \hat \xi_{jt} - \hat \sigma_{ij, h}^\xi\right]^2$ and $c_{ij, h} := \delta \left[\hat{\operatorname{\mathbb V}}\left[\hat \xi_{it} \hat\xi_{j,t-h}\right] \log(n) / (T-h)\right]^{1/2}$, while $\delta = [2(2-\gamma)]^{1/2}$ with $\gamma \in (0, 2)$ is a parameter for controlling sparsity.
\begin{align*}
asy \hat \bm{\Gamma}_{\hat \bm{x}_t} = \left(\bm{I}_{p+1} \otimes \left(\hat \bm{D}_{\bm{\Lambda}}^n\right)^{-1}\right)
\begin{pmatrix}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, 0}\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, 0}\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
& \cdots &
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, p}\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, p}\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
\\
& \ddots & \\
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, -p}\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, -p}\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
& \cdots &
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, 0}\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda}\bm{\xi}, 0}\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
\end{pmatrix} \left(\bm{I}_{p+1} \otimes \left(\hat \bm{D}_{\bm{\Lambda}}^n\right)^{-1}\right).
\end{align*}
Using the previous Theorems, we can prove asymptotic normality for the common component estimator for $i = 1, ..., n$ and $t = p+1, ..., T$:
\begin{theorem}[Asymptotic Normality of the Dynamic Common Component]\label{thm: asy norm chi}
Under Assumptions T\ref{T: q-DFS struct}-A\ref{A: more restrictive sample covarainces}, as $\sqrt{n} / T \to 0$ and $\sqrt{T} / n \to 0$, we have
\begin{align*}
\frac{\hat \chi_{it} - \chi_{it}}{\sqrt{\frac{1}{T} U_{it} + \frac{1}{n}V_{it}}} \Rightarrow \mathcal N(0, 1),
\end{align*}
where $U_{it} := \bm{x}_t'\bm{\Gamma}_{\bm{x}}^{-1} \bm{\Omega}_{\bm{x}\bm{\xi}}(i) \bm{\Gamma}_{\bm{x}}^{-1}\bm{x}_t$ and $V_{it} =\bm{\beta}_i\left(\bm{I}_{p+1} \otimes \bm{\Gamma}_{\bm{\Lambda}}^{-1}\right) \bm{\Theta}_{\bm{\Lambda}\bm{\Xi}}(t) \left(\bm{I}_{p+1} \otimes \bm{\Gamma}_{\bm{\Lambda}}^{-1}\right) \bm{\beta}_i'$.
\end{theorem}
For estimating the standard deviations $U_{it}$ and $V_{it}$ in Theorem \ref{thm: asy norm chi}, we use:
\begin{align}
\hat U_{it} = \hat \bm{x}_t' asy \hat \bm{\Gamma}_{\hat \bm{\beta}_i} \hat \bm{x}_t, \qquad \hat V_{it} = \hat \bm{\beta}_i asy \hat \bm{\Gamma}_{\hat \bm{x}_t} \hat \bm{\beta}_i'. \label{eq: est Uit Vit}
\end{align}
\begin{theorem}[Asymptotic Normality of the Weak Common Component]\label{thm: asy norm $e_{1t}^chi$}
Under Assumptions T\ref{T: q-DFS struct}-A\ref{A: more restrictive sample covarainces}, as $\sqrt{n} / T \to 0$ and $\sqrt{T} / n \to 0$, we have
\begin{align*}
\frac{\hat e_{it}^\chi - e_{it}^\chi}{\sqrt{\frac{1}{T}U_{it} + \frac{1}{n} V_{it}}} \Rightarrow \mathcal N(0, 1),
\end{align*}
where
\begin{align*}
U_{it} &:= \bm{x}_t'\bm{\Gamma}_{\bm{x}}^{-1} \bm{\Omega}_{\bm{x}\bm{\xi}}(i) \bm{\Gamma}_{\bm{x}}^{-1}\bm{x}_t + \bm{F}_t' \bm{\Omega}_{\bm{F}\bm{e}}(i) \bm{F}_t -2 \bm{x}_t'\bm{\Gamma}_{\bm{x}}^{-1} \bm{\Omega}_{\bm{x}\bm{\xi}, \bm{F}\bm{e}}(i) \bm{F}_t \\
V_{it} &:= \bm{\beta}_i \left(\bm{I}_{p+1} \otimes \bm{\Gamma}_{\bm{\Lambda}}^{-1}\right) \bm{\Theta}_{\bm{\Lambda}\bm{\Xi}}(t) \left(\bm{I}_{p+1} \otimes \bm{\Gamma}_{\bm{\Lambda}}^{-1}\right) \bm{\beta}_i' + \bm{\Lambda}_i \bm{\Gamma}_{\bm{\Lambda}}^{-1}\bm{\Theta}_{\bm{\Lambda} \bm{e}}(t) \bm{\Gamma}_{\bm{\Lambda}}^{-1} \bm{\Lambda}_i' \\
& \qquad \qquad \qquad - 2\bm{\beta}_i \left(\bm{I}_{p+1} \otimes \bm{\Gamma}_{\bm{\Lambda}}^{-1}\right) \bm{\Theta}_{\bm{\Lambda} \bm{\Xi}, \bm{\Lambda} \bm{e}}(t) \bm{\Gamma}_{\bm{\Lambda}}^{-1} \bm{\Lambda}_i'.
\end{align*}
\end{theorem}
The covariance terms can be estimated using:
\begin{align*}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}, \bm{F}\bm{e}}(i)\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\left(\bm{I}_{p+1} \otimes \bm{P}_{\bm{\Lambda}}\right) \bm{\Omega}_{\bm{x}\bm{\xi}, \bm{F}\bm{e}}(i)\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
&=\frac{1}{T}\sum_{t = 1}^T \sum_{s = 1}^T \hat \bm{x}_t \hat \xi_{it} \hat e_{is} \hat \bm{F}_s' \kappa(t, s),
\end{align*}
while for $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda} \bm{\xi}, \bm{\Lambda} \bm{e}, h}\bm{P}_{\bm{\Lambda}}'}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{\bm{P}_{\bm{\Lambda}} \bm{\Theta}_{\bm{\Lambda} \bm{\xi}, \bm{\Lambda} \bm{e}, h}\bm{P}_{\bm{\Lambda}}'}{\tmpbox}
$ we use the same formula as in \eqref{eq: est cov Lambda xi, h},
so $(T-h)^{-1}\sum_{t = h+1}^{T}\hat \xi_{it}\hat \xi_{i, t-h}$ instead $(T-h)^{-1}\sum_{t = h+1}^{T}\hat \xi_{it}\hat e_{i, t-h}$ which is justified by the orthogonality between $\bm{F}_t^w$ and $\xi_{is}$ for all leads and lags and $i \in \mathbb N$. The standard deviations are estimated analogously to \eqref{eq: est Uit Vit} using the formulas of Theorem \ref{thm: asy norm $e_{1t}^chi$}.
\section{Simulation Experiments}\label{sec: simulation}
To assess the finite sample properties of the proposed procedures, we consider data generated as
\begin{align}
y_{it} &= \chi_{it} + \xi_{it} = \lambda_{i1} f_t + \lambda_{i2} f_{t-1} + \xi_{it}, \label{eq: innocent factor model} \\[0.8em]
\mbox{with} \quad f_t &= a f_{t-1} + \varepsilon_t , \quad \varepsilon_t \sim \mathcal N\left(0, (1-a^2)\right), \nonumber
\end{align}
with $a = 0.8$ and one statically pervasive factor and one weak factor in the dynamic common space: For this we set $\lambda_{i1} = 0$ for $1\leq i \leq 10$, $\lambda_{i1} = 1$ for $11 \leq i \leq 20$, and $\lambda_{i1} \sim iidN(1, 1)$ for $i \geq 21$, and we set $\lambda_{i2} = 1$ for $1\leq i \leq 10$ and $\lambda_{i2} = 0$ for $i \geq 11$. For the dynamic idiosyncratic component $(\xi_{it})$ we consider different versions, based on the specification
\begin{align*}
\xi_{it} = \alpha_i \xi_{i, t-1} + \varepsilon^\xi_{it}
\end{align*}
with $\varepsilon_{it}^\xi \sim iidN(0, 1)$, independent of $(\varepsilon_t)$, with $\operatorname{\mathbb E}(\varepsilon_{it}^\xi \varepsilon_{jt}^\xi) = \tau^{\@ifstar{\oldabs}{\oldabs*}{i-j}}, i, j = 1, ..., n$, $\tau \in \{0, 0.5\}$ if $\@ifstar{\oldabs}{\oldabs*}{i-j}\leq 10$ and $\operatorname{\mathbb E}(\varepsilon_{it}^\xi \varepsilon_{jt}^\xi) = 0$ otherwise; last $\alpha_i = \{0, \delta_i\}$ with $\delta_i \sim iidU(0, \delta)$ and $\delta \in \{0, 0.5\}$. The parameters $\tau$ and $\delta$ are crucial to control the cross-sectional and serial correlation in the dynamic idiosyncratic component, respectively.
We obtain the canonical decomposition \eqref{eq: canonical rep weak and strong factors} by projecting out the statically pervasive factor $f_t$ from the lag $f_{t-1}$ \citep[see details in][]{gersing2023weak}. Let $\widetilde \lambda_{i1} = \lambda_{i1} + \lambda_{i2}a$, $\widetilde \lambda_{i2} = \lambda_{i2} \sigma_\varepsilon$, $\widetilde f_{1t} = f_t$, $\widetilde f_{2t} = \sigma_\varepsilon^{-1}(f_{t-1} - a f_t)$, with $\sigma_\varepsilon^2:= (1-a^2)$. Then, we can rewrite model \eqref{eq: innocent factor model} as
\begin{align}
y_{it}&= \underbrace{\widetilde \lambda_{i1} \widetilde f_{1t}}_{C_{it}} + \underbrace{\widetilde \lambda_{i2}\widetilde f_{2t}}_{e_{it}^\chi} + \xi_{it}, \label{eq: first ex noncan_alt}
\end{align}
where the factors $\widetilde \bm{F}_t=(\widetilde f_{1t} \; \widetilde f_{2t} )'$ are orthonormal, i.e., $\operatorname{\mathbb E} [\widetilde \bm{F}_t \widetilde \bm{F}_t' ]= \bm{I}_2$. The first factor $\widetilde f_{1t} = F_t$ is pervasive while the second one $\widetilde f_{2t} = F_t^w$ is not. Indeed, letting $\bm{\chi}_t^n = (\chi_{1t}, ..., \chi_{nt})'$, it is easily verified that only one signal eigenvalue in $\bm{\Gamma}_{\bm{\chi}}^n := \operatorname{\mathbb E} [\bm{\chi}_t^n \bm{\chi}_t^{n'}]$ diverges. Therefore $r = 1$ in view of the static decomposition.
The first unit, $i = 1$, decomposes as
\begin{align*}
\chi_{1t} &= \tilde \lambda_{11} \tilde f_{1t} + \tilde \lambda_{12} \tilde f_{2t} = a f_t + (f_{t-1} - a f_t) = f_{t-1}\\
C_{1t} &= a f_t , \quad e_{1t}^\chi = f_{t-1} - a f_t.
\end{align*}
\subsection{Consistency of FDL versus Non-Consistency of SPCA}
We generate data according to \eqref{eq: innocent factor model}. At each replication $j = 1, ..., B$ and for each considered estimator of $\chi_{1t}^{[j]}$, generically denoted as $\widehat \chi_{1t}^{[j]}$, we define $MSE_1 = B^{-1}\sum_{j = 1}^B T^{-1}\sum_{t = 1}^T (\widehat \chi_{1t}^{[j]} - \chi_{1t}^{[j]})^2$ and look at results for different $(n, T)$ combinations with $B = 500$ replications.
The left plot in Figure \ref{fig: inconsistency of weak factors} shows, that the principal component estimator with $r = 2$ (\texttt{spca2}) is not consistent for $\chi_{1t}$, indeed, $MSE_1$ is not approaching zero as $n$ and $T$ becomes larger \citep[for a theoretical proof see][]{onatski2012asymptotics}. The other two dynamic estimators (\texttt{dpca} and \texttt{fdl}) are consistent (Theorem \ref{thm: consistency of spaces and CCs}) as $MSE_1$ is monotonically decreasing with increasing $(n,T)$, while the finite distributed lags approach \texttt{fdl} yields better performance. Furthermore, the right plot of figure \ref{fig: inconsistency of weak factors} shows that we cannot recover $\chi_{1t}$ by static principal components, even if we increase the number of factors $r$, i.e., with \texttt{spcar} for $r = 1, 2, 3, 5, 9$. In other words, the weak factor $\widetilde f_{2t}$ is \textit{too weak} to be recovered by means of static principal components - though it is important individually, as it explains a large part of the variation for the first ten units.
\begin{figure}
\centering
\begin{tabular}{cc}
\includegraphics[width=.4\textwidth]{plots/fdl_res_nonconsist_tau05del05_dpca_fdl.pdf}&
\includegraphics[width=.4\textwidth]{plots/fdl_res_nonconsist_tau05del05_spca.pdf}
\end{tabular}
\caption{
\footnotesize DGP1.
Mean Squared Error of $\widehat \chi_{1t}$ over 500 replications under idiosyncratic serial and cross-correlation $(\tau, \delta) = (0.5, 0.5)$. \texttt{spcar}: estimation with static principal components with \texttt{r} $=1,2,3,5,9$, \texttt{dpca}: estimation by dynamic principal components with $q = 1$, \texttt{fdl}: finite distributed lag regression computed by regression on the first principal component and its first lag.}
\label{fig: inconsistency of weak factors}
\end{figure}
Analogous results are obtained for other configurations of the sample size and the idiosyncratic component, see tables \ref{tab: res nonconsist T > n}, \ref{tab: res nonconsist n < T} and \ref{tab: cov rates n = T}.
\subsection{Coverage Rates for the Confidence Intervals}
Next we explore the coverage rates of the asymptotic confidence intervals provided by the CLTs of section \ref{sec: estimation}. For estimating the loadings variance, we use the HAC version from equation \eqref{eq: HAC1 asy beta}, for estimating the factors' variance we use the estimator from \eqref{eq: est cov Lambda xi, h} by \cite{fresoli2024dealing}. Table \ref{tab: cov rates T > n} shows the results for the case where $T>n$. In all cases the sample coverage approaches the nominal rate of 95\% for increasing sample size $(n, T)$, while convergence kicks in about $(n, T) = (120, 240)$ and slightly slower for the case with higher idiosyncratic cross-correlation $(\tau = 0.5)$.
\begin{table}[!htbp] \centering
\begin{tabular}{@{\extracolsep{5pt}} c|c|c|c|c}
\\[-1.8ex]\hline
\hline \\[-1.8ex]
\multicolumn{5}{c}{\textbf{Coverage Rates, $T > n$}} \\[0.5em]
\hline
$(n,T)$ & (30,60) & (60,120) & (120,240) & (240,480) \\
\hline \\[-1.8ex]
\hline \\[-1.8ex]
$e_{1t}^\chi$, $\tau = 0$, $\delta = 0$ & 0.719 & 0.862 & 0.946 & 0.940 \\
$\chi_{1t}$, $\tau = 0$, $\delta = 0$ & 0.864 & 0.890 & 0.950 & 0.942 \\
$C_{1t}$, $\tau = 0$, $\delta = 0$ & 0.826 & 0.910 & 0.930 & 0.936 \\
\hline \\[-1.8ex]
$e_{1t}^\chi$, $\tau = 0.5$, $\delta = 0$ & 0.659 & 0.848 & 0.908 & 0.926 \\
$\chi_{1t}$, $\tau = 0.5$, $\delta = 0$ & 0.798 & 0.884 & 0.900 & 0.930 \\
$C_{1t}$, $\tau = 0.5$, $\delta = 0$ & 0.818 & 0.908 & 0.920 & 0.912 \\
\hline \\[-1.8ex]
$e_{1t}^\chi$, $\tau = 0$, $\delta = 0.5$ & 0.724 & 0.868 & 0.932 & 0.940 \\
$\chi_{1t}$, $\tau = 0$, $\delta = 0.5$ & 0.868 & 0.894 & 0.926 & 0.920 \\
$C_{1t}$, $\tau = 0$, $\delta = 0.5$ & 0.800 & 0.890 & 0.908 & 0.924 \\
\hline \\[-1.8ex]
$e_{1t}^\chi$, $\tau = 0.5$, $\delta = 0.5$ & 0.696 & 0.844 & 0.902 & 0.948 \\
$\chi_{1t}$, $\tau = 0.5$, $\delta = 0.5$ & 0.822 & 0.904 & 0.878 & 0.930 \\
$C_{1t}$, $\tau = 0.5$, $\delta = 0.5$ & 0.784 & 0.908 & 0.892 & 0.916 \\
\hline \\
\end{tabular}
\caption{Coverage rates for asymptotic $1-\alpha = 95\%$-confidence intervals of $\chi_{1,10}$, $e_{1,10}^{\chi}$ and $C_{1,10}$ over $B = 500$ replications.}
\label{tab: cov rates T > n}
\end{table}
\section{Empirical Application}
To demonstrate our methods empirically, we use different high-dimensional macroeconomic time series panels for the Euro Area (EA) and major EA countries, based on the open source dataset of \cite{barigozzi2024large}. We consider three datasets: a) monthly EA and country-level series ($n = 381$, $T = 309$, $2000:01$ to $2025:09$) b) quarterly EA and country-level series ($n = 595$, $T = 103$, $2000:Q1$ to $2025:Q3$) c) quarterly series for Germany ($n = 63$, $T = 103$, $2000:Q1$ to $2025:Q3$). All series are standardised to zero mean and unit sample variance. Following \cite{barigozzi2024large}, the data are transformed to stationarity. As principal components are sensitive to outliers but the sample includes the COVID-19 and several other crisis periods, we preprocess the data by removing univariate and multivariate outliers using standard methods. Missing values are then interpolated by splines (multivariate outliers) and the EM algorithm for factor models \citep[see][appendix]{stock2002forecasting} (univariate outliers).
The number of factors is estimated using \cite{alessi2010improved}, which refines \cite{bai2002determining} in the spirit of \cite{hallin2007determining}. Following factor selection, model \eqref{eq: GDFM stat rep} is estimated by regressing on the estimated factors, with the lag order determined by the modified BICM of \cite{groen2013model}, specifically designed for regressions involving estimated factors. To retain information from crisis periods, parameter estimation is conducted on the cleaned data, while factor and common component estimation use the standardised raw data combined with eigenvectors and eigenvalues estimated from the cleaned sample covariance matrix. Finally, we obtain the estimates $\hat \chi_{it}$ from model \eqref{eq: GDFM stat rep} and $\hat e_{it}^\chi = \hat \chi_{it} - \hat C_{it}$.
Let us comment the results in order: a) For the monthly data, we estimate $\hat r = 4$ factors, with $130$ series exhibiting lag order $p>0$. The weak common component (WCC) explains $3.6\%$ of the total variance, while the static common component (SCC) accounts for $19.6\%$. For $55$ variables, the WCC explains more than $10\%$ of the total variance, reaching up to $25.5\%$ for ESENTIX (Economic Sentiment Indicator) of the Euro Area.
Recall that the SCC captures the \textit{contemporaneously common} variation, arising from linear combinations across all series at a fixed time $t$. In contrast, the dynamic common component (DCC) reflects responses to common structural shocks, i.e., the projection of observed variables on the infinite past of common shocks $\varepsilon_t$ (see \eqref{eq: GDFM rep}). The DCC thus accommodates heterogeneous persistence in responses to these shocks, whereas persistence in the SCC derives solely from statically pervasive factors.
For ESENTIX (Figure \ref{fig: ESENTIX EA}), the WCC allows the DCC to track the observed series more closely during the COVID period, while the SCC responds only weakly to the large shocks. Moreover, outside crisis periods, the WCC captures additional dynamics not accounted for by the SCC. Similarly, for the growth rate of French industrial production (Figure \ref{fig: IPMN France}), the DCC exhibits dynamics that differ significantly from those of the SCC.
\begin{figure}
\centering
\begin{tabular}{cc}
\includegraphics[width=.5\textwidth]{plots/chi_conf_ESENTIX_EA.pdf}&
\includegraphics[width=.5\textwidth]{plots/echi_conf_ESENTIX_EA.pdf}
\end{tabular}
\caption{
\footnotesize Monthly Data a): Estimates of the Canonical Decomposition of standardised $\Delta \text{ESENTIX}_t$ = Economic Sentiment Indicator of the Euro Area. \texttt{Black solid line}: \textit{dynamic common component} (left), or \textit{weak common component} (right), \texttt{red line}: static common component, \texttt{dotted line}: true observed standardised index. Blue area represents 90\% confidence intervals.}
\label{fig: ESENTIX EA}
\end{figure}
\begin{figure}
\centering
\begin{tabular}{cc}
\includegraphics[width=.5\textwidth]{plots/chi_conf_IPMN_FR.pdf}&
\includegraphics[width=.5\textwidth]{plots/echi_conf_IPMN_FR.pdf}
\end{tabular}
\caption{
\footnotesize Monthly Data a): Estimates of the Canonical Decomposition of standardised $\Delta \log \text{IPMN}_t$ = Industrial Production Index: Manufacturing of France. \texttt{Black solid line}: \textit{dynamic common component} (left), or \textit{weak common component} (right), \texttt{red line}: static common component, \texttt{dotted line}: true observed standardised index. Blue area represents 90\% confidence intervals.}
\label{fig: IPMN France}
\end{figure}
Next, consider dataset b) of quarterly series. Only 10 out of 595 variables exhibit $p>0$. The WCC explains 0.3\% of total variation, whereas the SCC accounts for 38.9\%. Estimates for total hours worked (THOURS) in the Euro Area - alongside unemployment, a key labour market indicator - are shown in Figure \ref{fig: TOTHOURS EA}. Here as well, the WCC appears offset the stronger mean reversion of the SCC relative to the observed series.
\begin{figure}
\centering
\begin{tabular}{cc}
\includegraphics[width=.5\textwidth]{plots/chi_conf_THOURS_EAQ.pdf}&
\includegraphics[width=.5\textwidth]{plots/echi_conf_THOURS_EAQ.pdf}
\end{tabular}
\caption{
\footnotesize Estimates of the Canonical Decomposition of standardised THOURS = Total Hours Worked for Spain. \texttt{Black solid line}: \textit{dynamic common component} (left), or \textit{weak common component} (right), \texttt{red line}: static common component, \texttt{dotted line}: true observed standardised index. Blue area represents 84\% confidence intervals.}
\label{fig: TOTHOURS EA}
\end{figure}
\begin{figure}
\centering
\begin{tabular}{cc}
\includegraphics[width=.5\textwidth]{plots/chi_conf_GNFCIR_DE.pdf}&
\includegraphics[width=.5\textwidth]{plots/echi_conf_GNFCIR_DE.pdf}
\end{tabular}
\caption{
\footnotesize Estimates of the Canonical Decomposition using a time series panel comprising only macroeconomic indicators from Germany of GNFCIR = Gross-Investment Share of Non-Financial Corporations. \texttt{Black solid line}: \textit{dynamic common component} (left), or \textit{weak common component} (right), \texttt{red line}: static common component, \texttt{dotted line}: true observed standardised index. Blue area represents 84\% confidence intervals.}
\label{fig: GNFCIR Germany}
\end{figure}
Finally, consider dataset c), comprising only German series. Only one variable has $p>0$, namely GNFCIR (Gross Investment Share of Non-Financial Corporations) with $p=2$. Despite this, GNFCIR is a key leading indicator of long-term economic growth, reflecting business confidence about future conditions. In this case, most of the dynamics are captured by the WCC (40\%), while the SCC accounts for a smaller share of variation (30\%) relative to the observed series (see Figure \ref{fig: GNFCIR Germany}).
\section{Conclusion}
We consider the generalised dynamic factor model under the assumption that the dynamic common component can be represented in terms of a projection of the output on the statically pervasive factors and a finite number of their lags. First, we justify why this assumption is sensible from a theoretical point of view and show that it holds under general conditions for the prototypical dynamic factor model. We set up an asymptotic framework that allows the factors to correlate at leads and lags with the static idiosyncratic component (in which the weak common component is absorbed) to accommodate for the existence of weak non-pervasive factors in the dynamic common space. We provide asymptotic confidence intervals for the dynamic, the weak and the static common component. In an empirical demonstration on Euro-Area Macroeconomic Data, we find that there are several important variables significantly driven by the weak common component.
\bibliographystyle{apalike}
\bibliography{references.bib}
\clearpage
\begin{center}
{\LARGE \bf Appendix}
\end{center}