Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,309 characters · 21 sections · 85 citation commands
Determining the Structure of Dynamic Factor Models
JEL classification: C10, C38, C51, C55 \\ Key words and phrases: Dynamic Factor Model, Number of Factors, Filter Length, Alternating Least Squares
Factor models are a central tool in empirical macroeconomics and finance for summarizing co-movements in large-dimensional time series. Their appeal lies in the idea that a small number of latent factors can explain the dynamics of many observed variables.
A fundamental distinction is between static and dynamic factor models. In a static factor model, the observables $x_t$ depend only contemporaneously on a low-dimensional vector of factors $f_t$, e.g., $x_t = \lambda f_t + \varepsilon_t$, whereas a dynamic factor model allows lags of the factors to affect the observables, e.g., $x_t = \sum_{k=0}^{m-1} \lambda_k f_{t-k} + \varepsilon_t$. The dynamic specification is more general and economically informative, as it captures delayed responses and propagation mechanisms that are central to macroeconomic and financial dynamics.\footnote{The dynamic factor model studied in this paper allows the factor process $(f_t)$ is allowed to exhibit arbitrary serial dynamics, provided that it satisfies the Assumption (ref) to be introduced in Section (ref). This is in contrast to the generalized dynamic factor model of Forni2000.}
While a large body of work has developed methods for estimating the number of static factors—see, for example, Bai02, Onatski-2010, and Ahn-Horenstein-2013—the literature on estimating the number of dynamic factors remains relatively limited. Notable exceptions are Bai2007 and Amengual-Watson-2007. These papers approach the problem by constructing a static representation of the dynamic factor model. Under the assumption that the static representations of the dynamic factors follow a VAR process with a lower-dimensional innovation, Bai2007 and Amengual-Watson-2007 develop consistent procedures for estimating the number of dynamic factors. Although VAR models provide a flexible and widely used framework for capturing time series dynamics, the literature on the study of the static factor models shows that a VAR specification for the static factors is not necessary for either estimating the factors or determining their number. This raises the question of whether it is possible to consistently estimate the number of dynamic factors without imposing a VAR model on the factors.
This paper proposes a new approach to estimating the number of dynamic factors that does not rely on the assumption that the factors follow a VAR process. Instead, our method requires that the second moment matrix of the static representations of the dynamic factors be positive definite, a substantially weaker condition than imposing a parametric VAR structure. Our approach is enabled by directly estimating the dynamic factors using an alternating least squares algorithm. The procedure iteratively updates factors and loadings via simple least squares steps, making it computationally efficient and scalable to large datasets. When applied to static factor models, the alternating least squares estimator reduces to the principal components estimator commonly used in static factor models. We establish consistency under double asymptotics when both the number of samples and the dimension of the observed variable diverge, obtaining the standard convergence rates typically associated with the estimation of static factors. By leveraging consistent estimators of the dynamic factors, we propose two procedures that jointly estimate the number of dynamic factors and the filter length, defined as the length of the horizon over which current and lagged factors exert a direct effect on the observed variables.
Our approach extends the criteria of Ahn-Horenstein-2013 and Bai02 to dynamic factor models. Under general Assumptions to be introduced in Sections (ref) and (ref), the number of dynamic factors and the filter length of the dynamic factor model are jointly identified asymptotically independently of any parametric specification for the stochastic process governing the dynamic factors. In particular, if the dynamic factors follow a VAR\((p)\) process, the identification of the filter length is invariant to the choice of the lag order \( p \). We show in Section (ref) that when the dynamic factors follow a VAR process, the lag order selection can be done using the standard BIC criterion with the unobserved factors replaced with their estimates.
The eigenvalue ratio test of Ahn-Horenstein-2013 identifies the number of factors by detecting the point at which the ratio of consecutive singular values of the data matrix diverges. We extend this idea to dynamic factor models by introducing a notion of 'dynamic' singular values, defined as the spectral norm of the residual matrix after removing a given number of estimated dynamic factors. Building on this concept, we propose the dynamic singular value ratio test, which generalizes the eigenvalue ratio test of Ahn-Horenstein-2013 and coincides with it when applied to a static factor model. Moreover, the proposed framework can be used to determine the filter length of the dynamic factor model by examining the behavior of ratios of consecutive dynamic singular values as the filter length varies for a fixed number of dynamic factors.
The information criterion of Bai02 selects the number of static factors by augmenting the residual variance with a penalty term that increases with the number of static factors. Our second procedure extends this approach to dynamic factor models by employing a penalty term that depends on both the number of dynamic factors and the filter length. When we restrict the filter length to one, the proposed criterion reduces to that of Bai02, up to a constant factor of two in the penalty term.
There is a large literature on dynamic factor models, which can be broadly divided into two classes. The first class comprises generalized dynamic factor models with infinite filter length, which do not admit a finite-dimensional static representation. The second class consists of models with a finite filter length, for which a static representation exists. Our work falls into the second category: the dynamic factor models with a finite filter length. What distinguishes our work from many existing ones is that we work directly with the dynamic factors themselves, rather than the static representations.
The generalized dynamic factor model with infinite filter length was introduced by Forni2000, who proposed estimation based on frequency-domain principal components. Subsequent work has developed one-sided representations, estimation methods, and asymptotic theory for these models: Forni2015, Forni-Hallin-Lippi-Zaffaroni-2017, and Barigozzi-Hallin-Luciani-Zaffaroni-2023. Tests for the number of dynamic factors in the generalized dynamic factor framework have also been proposed by Hallin2007 and Onatski-2010. However, the greater generality afforded by infinite-length filters typically requires the dynamic factors to be serially uncorrelated processes for identification. We focus our attention on dynamic factor models with finite length filter in this paper.
A separate strand of the literature studies the static representations of dynamic factor models. Bai-2003 and Bai-Li-2016 provide inferential theory for static factors estimated by principal components and quasi–maximum likelihood, respectively. Procedures for determining the number of static factors include information criteria Bai02 and eigenvalue-ratio tests Ahn-Horenstein-2013, among others. Despite this extensive literature, comparatively little attention has been paid to the direct estimation of dynamic factors. An exception is Bai2015, who studies Bayesian estimation of dynamic factor models. This paper is most closely related to Bai2007 and Amengual-Watson-2007, who study the estimation of the number of dynamic factors in models with finite filter length.
To summarize, our contributions are as follows.
First, we generalize the procedures of Bai02 and Ahn-Horenstein-2013 to estimate the number of dynamic factors in dynamic factor models. Our approach is more robust than those of Bai2007 and Amengual-Watson-2007, as it does not require imposing a VAR structure on the static factors.
Second, our procedure provides a consistent estimator of the filter length, jointly with the number of dynamic factors, which, to the best of our knowledge, is the first in the literature. Identifying the extent to which primitive shocks directly affect the economy is of independent interest. For instance, using the FRED--MD dataset of McCracken-Ng-2015 over the sample period March 1973 to July 2025, the filter length is estimated to be two.\footnote{The associated number of dynamic factors were found to be four, which increases to seven after the onset of COVID.}
Third, we examine the finite-sample performance of our testing procedure and show that they outperform those of Bai2007 and Amengual-Watson-2007 across a range of designs in determining $q$.
Finally, we establish the consistency of the proposed alternating least squares estimator in estimating the dynamic factors. Furthermore, when the dynamic factors follow a VAR$(p)$ process, we show that the lag order selection can be done using the standard BIC criterion with the unobserved factors replaced with their estimates.
We denote by $\|\cdot\|$ the Euclidean norm when applied to vectors and the associated spectral (operator) norm when applied to matrices. The Frobenius norm is denoted by ${\left\vert\kern-0.3ex\left\vert\kern-0.3ex\left\vert \cdot \right\vert\kern-0.3ex\right\vert\kern-0.3ex\right\vert}$. The row-major vectorization operator is denoted by $\operatorname*{vec}(\cdot)$. For a symmetric matrix $A$, the notation $A>0$ and $A\geq0$ indicates that $A$ is positive definite and positive semi-definite, respectively. We denote $P_A=A(A'A)^{-1}A'$ and $M_A=I-P_A$ for any matrices $A$. We write $\lceil x \rceil$ for the smallest integer greater than or equal to $x$ and $\lfloor x\rfloor$ for the largest integer greater than or equal to $x$. Throughout, $q$ denotes the number of dynamic factors, $m$ the filter length, and $r = qm$ the corresponding number of static factors. Finally, we denote $\mathbb{Z}$ and $\mathbb{R}$ the set of nonnegative integers and real numbers, respectively.
The remainder of the paper is organized as follows. Section (ref) presents the model and discusses representations of the dynamic factors. Section (ref) describes the alternating least squares estimation procedure and the consistency of the proposed method. Section (ref) develops two methods for determining the number of dynamic factors and the filter length. Section (ref) reports simulation evidence on the finite-sample performance of the proposed procedures. Section (ref) provides an empirical application. Section (ref) concludes. All technical proofs are collected in the Appendix.
We consider the dynamic factor model
where $f_t \in \mathbb{R}^q$ are the dynamic factors and $\lambda_{ik} \in \mathbb{R}^q$ are factor loadings. When $m=1$, the model reduces to a static factor model. When $m=\infty$, it corresponds to the generalized dynamic factor model of Forni2000. However, we restrict attention to the finite-filter case $m<\infty$, which can be viewed either as a finite-dimensional specialization of Forni2000 or as a dynamic generalization of the conventional static factor model of Bai-2003 and Stock02. The model admits a static representation,
where $g_t \in \mathbb{R}^{qm}$ is the static factors constructed using the current and lags of the dynamic factors. Define \[ x_t = (x_{1t},\ldots,x_{Nt})', \qquad \varepsilon_t = (\varepsilon_{1t},\ldots,\varepsilon_{Nt})', \] and stack observations over time as \[ X = (x_1,\ldots,x_T)', \qquad E = (\varepsilon_1,\ldots,\varepsilon_T)', \qquad G = (g_1,\ldots,g_T)'. \] The model can then be written compactly as
where $\Gamma = (\lambda_{0},\ldots,\lambda_{m-1})$ is the $N \times qm$ matrix of factor loadings. Let $F_{-k} = (f_{1-k},\ldots,f_{T-k})'$ collect the $k$-th lag of the factors so that $G = \big( F_{0}, \ldots, F_{-m+1} \big)$. Then, we may write the model in its dynamic form
We let $F = F_0$ for the $T \times q$ matrix collecting $(f_t)_{t=1}^T$.
The dynamic factor model can be regarded as a static factor model with the restrictions that $m$ blocks of the static factors are given as leads and lags of themselves. Thus, direct estimation of dynamic factors is always more parsimonious than estimation based on static representations.
We impose the following two assumptions, which are required to establish the consistency of the estimated factors and factor loadings. We'll introduce two additional assumptions in Section (ref), which are needed to establish consistency for tests of the number of dynamic factors and filter length.
Assumption (ref) requires that the static representations $g_t$ of the dynamic factors $f_t$ are strong factors. Assumption (ref) holds whenever the dynamic factors $(f_t)$ are stationary with a spectral density matrix that is positive definite on a set of frequencies with positive measure, which includes any stationary and invertible VARMA processes with non-degenerate innovation as a special case. In fact, non-degenerate VAR processes have strictly positive definite spectral density matrix for all frequencies, so that Assumption (ref) is a strictly weaker condition.\footnote{Assumption (ref) is discuss in more detail in the Supplemental Materials (ref).} Note that both Bai2007 and Amengual-Watson-2007 also impose Assumption (ref) so that the number of static factors $r$ can be consistently estimated. In contrast, imposing a simple VMA($1$) model for $(f_t)$ will invalidate both Bai2007 and Amengual-Watson-2007 because a VMA($1$) model cannot be written as finite lag VAR model.\footnote{Bai2007 mentions but does not pursue the possibility of extending their theory to infinite-order VAR processes with rapidly decaying coefficients.}
Assumption (ref) is standard in the factor models literature and used to limit the amount of cross-sectional and serial correlations in the errors, e.g., Onatski-2010, Ahn-Horenstein-2013, Moon_Weidner_2017, and Bai-Ng-2023.
Let $(q_0, m_0)$ denote the unknown true structure, i.e., the number of dynamic factors and filter length, of the dynamic factor model, and let $(q, m)$ be a general notation used to denote the number of dynamic factors and the filter length. It is well known that the dynamic factor model with structure $(q_0,m_0)$ has an observationally equivalent static representation with $(q,m)=(q_0m_0,1)$. Dynamic factor model may also admit multiple dynamic representations other than the fully static representation. In other words, there may exist multiple dynamic factor models with different structures $(q, m)$ whose common components are observationally equivalent.
For example, consider a model with $(q,m)=(2,3)$ given by \[ x_t = \sum_{k=0}^{2} \lambda_k f_{t-k} + \varepsilon_t, \quad f_t \in \mathbb{R}^2. \] It is well known that this model has a static representation given as
and $g_{t} =
\in\mathbb{R}^6$. This model can also equivalently be written as
where $g_{t} =
\in\mathbb{R}^3$. Here, $f_{t-k,j}$ and $\lambda_{k,j}$ denote the $j$-th element of $f_{t-k}$ and the $j$-th column of $\lambda_k$, respectively. The model \eqref{example2} with the structure $(q,m) = (6,1)$ and the model \eqref{example1} with the structure $(q,m) = (3,2)$ are observationally equivalent to the original model with structure $(q,m)=(2,3)$.
Since there may be multiple pairs $(q,m)$ with the correct specification, we denote the pair $(q,m)$ with the smallest $q$ as the true structure $(q_0, m_0)$. If $m < m_0$, more factors are needed to achieve correct specification, and if $q>q_0$, a smaller filter length may suffice to preserve correct specification. The following proposition shows that such representations are always possible and gives the minimal number of additional factors required, given $m<m_0$, and the smallest $m$ attainable when $q>q_0$.
Let $q_0\geq0,$ $m_0\geq0$ be given, and consider disjoint partition of $\mathbb{Z}^2$ given by
where under $q_0=0$ or $m_0=0$ we let $S_1=\mathbb{Z}^2$ and $S_2=S_3=\emptyset$. Proposition (ref) implies that $S_1$ is the set of structures $(q,m)$ that yield models observationally equivalent to the true model with $(q_0,m_0)$. Furthermore, structures in $S_2$ lead to mis-specification. We introduce Assumption (ref) in Section (ref) so that structures in $S_3$ also lead to mis-specification as long as $m$ is finite.
In the dynamic factor model, \[ x_t = \sum_{k=0}^{m-1} \lambda_k f_{t-k} + \varepsilon_t, \quad t=1,\ldots,T, \] the dynamic factors $(f_t)$ are indexed from $t = 2-m$ to $t = T$ to account for the lagged factors entering the measurement equation. We jointly estimate $(f_t)_{t=2-m}^{T}$ and $(\lambda_k)_{k=0}^{m-1}$ by minimizing the least squares objective
over $(f_t)$ and $(\lambda_k)$. The first-order conditions for the factor loadings $(\lambda_j)$ are
The first-order conditions for $(f_t)$ are
for $t=2-m,\ldots,T$.
We minimize (ref) using an alternating least squares procedure. We start with an initial guess for $(\hat \lambda_k)$. Given $(\hat \lambda_k)$, we may update $(f_t)$ by solving (ref). Given $(\hat f_t)$, we may update $(\lambda_k)$ by solving (ref). We repeat the updating steps until convergence. The first-order conditions (ref) defines a multivariate regression with rank $qm$, which has a closed-form solution. The first-order condition (ref) can be efficiently solved using the discrete Fourier transform, requiring inversion of only $q \times q$ matrices.\footnote{Derivation of the discrete Fourier transform needed to solve the first order condition is given in the Supplemental Materials (ref).} Since each step minimizes a least squares objective, the algorithm guarantees a monotone decrease of the objective function. The minimization problem is not convex, and the alternating least squares algorithm generally does not converge to the global minimum. We try multiple initial values and choose the solution that obtains the smallest objective function value in practice.
Before introducing the test for the structure $(q,m)$ of the dynamic factor models, we first show that the estimated dynamic factors are consistent.
Proposition (ref) shows that the factors $g_t$ estimated via the alternating least squares applied to the dynamic factors $f_t$ are consistent as long as $(q,m)\in S_1$. Theorem 1 of Bai02 implies $g_t$ estimated using the structure $(q,m)$ in the set $\{(q,m):qm\geq q_0m_0,\:m=1\}$, i.e., the set of structures associated with the correct or over-specified static representations with $m=1$, is consistent. Proposition (ref) extends this result to all the structures $(q,m)\in S_1$ that can lead to observationally equivalent common components by directly estimating models with $m>1$. The alternating least squares estimator for $f_t$ imposes a specific structure on $g_t$, whereby its $m$ blocks correspond to leads and lags of the same underlying dynamic factors. That is, \[ \hat g_t =
. \] In contrast, when $g_t$ is estimated directly via principal components, this temporal structure is not imposed, and the above relation need not hold. This distinction is crucial when determining the structural dimensions $(q_0, m_0)$ of the dynamic factor model.
To construct a test for the true model $(q_0,m_0)$, we employ a notion of dynamic singular values, which generalizes the classical singular value decomposition to incorporate temporal structure. Formally, let $\delta_{q+1,m}(X)$ denote the $(q+1)$-th dynamic singular value of a matrix $X$ with filter length $m$, defined as \[ \delta_{q+1,m}(X) = \min_{(f_t) \in \mathbb{R}^{q}, \, (\lambda_k) \in \mathbb{R}^{q}} \Big\| X - \sum_{k=0}^{m-1} F_{-k} \lambda_k' \Big\|, \] where $q \ge 0$ and $m \ge 0$ with the convention that $\delta_{1,m}(X)$ or $\delta_{q+1,0}(X)$ denotes spectral norm of $X$. Note that we let $(f_t)$ and $(\lambda_k)$ be optimization arguments, not the true factors and factor loadings, whenever they appear as arguments within a minimization problem. Dynamic singular value generalizes the notion of the standard singular values; When $m=1$, $\delta_{q+1,1}(X)$ coincides with the usual $(q+1)$-th singular value of $X$: \[ \delta_{q+1,1}(X) = \min_{F \in \mathbb{R}^{T \times q}, \, \lambda \in \mathbb{R}^{N \times q}} \| X - F \lambda' \|. \] Note that $\delta_{q+1,m}(X)$ can be interpreted as the minimization problem
with the additional constraints that $m$ blocks of $G\in\mathbb{R}^{T\times qm}$ with $q$ columns are given as leads and lags of each other. This implies that we have $\delta_{q_0m_0+1,1}(X)\leq \delta_{q_0+1,m_0}(X)$ for any $(q_0,m_0)$. More generally, we have
for any integer $(q,m)\in S_1$ because $\delta_{q+1,m}(X)$ minimizes the objective function over a larger space due to Proposition (ref).
Assumption (ref) requires that the leading singular values of $\frac{1}{\sqrt{\max(N,T)}}E$ be bounded away from zero in probability. Under general conditions, the number of singular values that remain bounded away from zero can be shown to grow at the rate $\min(N,T)$; see, for example, Ahn-Horenstein-2013. As a direct implication of Assumption (ref), for any $\epsilon>0$ there exists $c_0>0$ such that \[ \mathbb{P}\!\left( \frac{1}{\sqrt{N+T}} \, \delta_{q+1,m}(E) < c_0 \right) \le \epsilon \] for all pairs $(q,m)$ satisfying $qm \le k_{\max}$, and for all sufficiently large $N,T$.
Assumption (ref) imposes a lower bound on the dynamic singular values of the common component. Note that
where the inequality follows from the additional restrictions imposed on the minimization problem on the left hand side. The term $\frac{1}{\sqrt{NT}} \delta_{q_0 m_0 - m_0 + 1,1}(G\Gamma')$ is strictly positive in the limit provided that $m_0\ge1$. Indeed, the $k$-th singular value of $(NT)^{-1/2}G\Gamma'$ equals the square root of the $k$-th eigenvalue of \[ \left(\frac{1}{N}G'G\right)\left(\frac{1}{T}\Gamma'\Gamma\right), \] which is bounded away from zero for all $k=1,\ldots,q_0m_0$ by Assumption (ref). Consequently, \[ \frac{1}{\sqrt{NT}} \delta_{q,m}\!\left(\sum_{k=0}^{m_0-1}F_{-k}\lambda_k'\right) \] is bounded away from zero in probability for all $q=1,\ldots,q_0$ and $m=1,\ldots,m_0$.
Assumption (ref) further requires this lower bound to persist for $m=m_0+1,\ldots,m_{\max}$. This restriction rules out the possibility that omitted dynamic factors can be fully recovered by increasing the filter length of the remaining factors. In other words, when fewer than $q_0$ dynamic factors are extracted, enlarging the lag polynomial alone is insufficient to span the factor space associated with the omitted components.
This requirement is satisfied whenever the $q_0$ largest eigenvalues of the spectral density matrix of the common component diverge at the rate $N$ on a set of frequencies with positive measure, which is more lenient than the standard assumption of divergence almost everywhere (see, e.g., Forni2000, Forni2015, Forni-Hallin-Lippi-Zaffaroni-2017). If $(f_t)$ is given by a stationary and invertible VARMA process with non-degenerate innovation, $q_0$ largest eigenvalues diverge at the rate $N$ for all frequencies so that Assumption (ref) is satisfied.
Assumption (ref) is imposed so that the structures in the set $S_3$ defined in (ref) are mis-specified as long as $m\leq m_{max}$. This implies that given $q_{max}$ and $m_{max}$ satisfying $q_0\leq q_{max}, m_0\leq m_{max}$, the set of correctly specified structures $(q,m)$ are given by \[ \mathcal{C}=\left\{(q,m)\in \mathbb{Z}^2 : q_0\le q\leq q_{max},\ \left\lceil \frac{q_0 m_0}{q} \right\rceil \le m \leq m_{max}\right\}. \] We construct a test based on estimated dynamic singular values. Let $\hat\delta_{q+1,m}(X)$ denote the estimated $(q+1)$-th dynamic singular value, defined as the spectral norm of the residual matrix obtained after fitting a dynamic factor model with $q$ factors and filter length $m$. Specifically,
where $(\hat F_{-k}^{(q)},\hat\lambda_k^{(q)})_{k=1}^m$ are the least-squares estimators obtained by minimizing the squared error under $q$ dynamic factors with $m$ filter length, i.e., the objective function (ref).
It is important to note that the alternating least squares algorithm yields the minimizer of the Frobenius norm, whereas the dynamic singular value is defined as the minimum of the spectral norm. The two minimization problems coincide in the static case ($m=1$), in which case $\hat\delta_{q+1,1}(X)$ reduces to the usual $(q+1)$-th singular value of $X$. For $m>1$, the two minimizers generally differ. While dynamic singular values can in principle be computed by directly minimizing the spectral norm, we rely on the Frobenius-norm solution for computational tractability. This approximation does not affect the asymptotic validity of the proposed test, as seen in Theorems (ref) and (ref).
As can be seen from the proofs of Theorems (ref) and (ref), we may characterize the asymptotic behavior of $\hat\delta_{q+1,m}(X)$ under Assumptions (ref), (ref), and (ref). In other words, we may show that there exists $q_{max}$ and $m_{max}$ such that for all $0\leq q\leq q_{max}$ and $0\leq m\leq m_{max}$, we have
This behavior of the estimated dynamic singular values $\hat\delta_{q+1,m}(X)$ motivates the dynamic singular value ratio tests given in Theorem (ref) and (ref) below.
In Theorem (ref), we fix $m$ and estimate the number of dynamic factors using a ratio-type statistic based on adjacent estimated dynamic singular values.
Theorem (ref) shows that, when the filter length is at least as large as $m_0$, the ratio statistic consistently identifies the true number of dynamic factors. When the filter length smaller than $m_0$, the procedure selects the minimal number of factors required to span the true common component, reflecting the observational equivalence between $(q_0,m_0)$ and alternative $(q,m)$ representations discussed in Proposition (ref). When $m=1$, the test statistic coincides with the eigenvalue ratio (ER) test of Ahn-Horenstein-2013 and consistently estimates the number of static factors, $r=q_0m_0$. Theorem (ref) thus generalizes the ER test to dynamic factor models by allowing the filter length $m$ to exceed one.
In Theorem (ref), we fix $q$ and estimate the filter length using a similar ratio-type statistic.
Theorem (ref) provides a consistent estimator of the true filter length $m$ in dynamic factor models. If $q=q_0$, the test uncovers the true filter length. If $q>q_0$, the test select the smallest filter length required to span the true common component. The case $q<q_0$ is excluded. Under Assumption (ref), $\delta_{q+1,m}(X)=O_p(\sqrt{NT})$ for all $m\le m_{\max}$ when $q<q_0$. As a result, $DR_{q,m}(X)$ may not diverge at $m=m_0$, and ratio-based tests for the filter length fail when not enough dynamic factors are included.
To recover $(q_0,m_0)$ using Theorem (ref) and (ref), we may start with any $m$ sufficiently large so that $m\geq m_0$. Using $m\geq m_0$, we may recover $q_0$ using Theorem (ref). Then, using Theorem (ref), we may recover $m_0$ using $q=q_0$.
Let $V(q,m)$ denote the mean squared errors when the structure $(q,m)$ is used to estimate a dynamic factor model, i.e.,
where $(\hat F_{-k}^{(q)},\hat\lambda_k^{(q)})_{k=0}^{m-1}$ is the least squares estimator obtained under the structure $(q,m)$.
Following Bai02, we augment the residual variance with a penalty term. The key is to construct a penalty sequence that vanishes asymptotically, yet dominates the difference in the residual sum of squares between the true model and any overparameterized alternative. In this spirit, we define the following two information criteria:
where $r=qm$ and $g(N,T)$ is a vanishing penalty function. Compared to Bai02, the main differences are: (i) model fit is evaluated after estimating the dynamic factor model, not its static representation, and (ii) the penalty accounts for both the number of static factors $r=qm$, and the number of dynamic factors $q$.
The $DC$ criterion is motivated naturally from the static case. For static models, the estimated factors and loadings minimize both the Frobenius norm and the spectral norm of the residuals. Hence, the squared spectral norm can serve as an alternative objective used to estimate the factors and factor loadings. While this argument does not trivially extend to dynamic factors, the $DC$ criterion also succeeds in finding the correct number of dynamic factors.
The key to determining the true structure $(q,m)$ in $\mathcal{C}$, i.e., the structure $(q,m)$ with the smallest $q_0$, is penalizing both the number of static factors $r=qm$ and the number of dynamic factors $q$. Penalizing only the number of static factors would select the static representation of the dynamic factor model, since among the structures $(q,m)$ with the same static factors, the structure $(qm,1)$ has the smallest squared error. Including $q$ as a separate term in the penalty ensures recovery of the most parsimonious dynamic structure. In fact, we may use any penalty function of the form $\varphi(q,m)g(N,T)$ as long as $g(N,T)\to0$, $\left(\frac{NT}{N+T}\right)g(N,T)\to\infty$ as $N,T\to\infty$, and
for all $(q,m)\in\mathcal{C}$ with strict inequality when $(q,m)\neq (q_0,m_0)$. For example, $\varphi(q,m) = q+m$ can also be used to consistently recover $(q_0,m_0)$.
Theorem (ref) shows that Assumptions (ref), (ref), and (ref) suffices to asymptotically identify $(q,m)$ regardless of parametric specification of $(f_t)$. Assumptions (ref) and (ref) are commonly used to identify the number of static factors, e.g., Chamberlain_Rothschild_1983, Bai02, Bai-Ng-2023. To recover the number of dynamic factors, along with the corresponding filter length, the only additional assumption required is the Assumption (ref).
Note that when $q_0=0$ or $m_0=0$, so that there does not exist a common component with factor structure, the information criterion test consistently recovers $(q_0,m_0)=(0,0)$.
The proof is omitted, as it follows directly from Corollary 1 of Bai02. The $IC$ criterion is convenient because it does not require scaling the penalty term, so that the test becomes independent of $q_{\max}$ or $m_{\max}$. As in Bai02, we use three candidate penalty terms:
We also scale the penalty term with $\hat\sigma^2 = V(q_{max},m_{max})$ for PC criterion and $\hat\sigma^2 = \frac{1}{NT}\hat\delta_{q_{max}+1,m_{max}}^2(X)$ for DC criterion.
Both $PC$ and $IC$ can be viewed as dynamic extensions of the original criteria in Bai02: if $m=1$, they reduce numerically to the criteria in Bai02, up to a multiplicative factor of two in the penalty term.
Although our method does not require the dynamic factors \((f_t)\) to follow a VAR process, it may still be useful in practice to model their dynamics using a VAR model. For instance, one may be interested in estimating the dynamic responses of macroeconomic variables to structural shocks identified using residuals of a VAR model for \((f_t)\), as in Forni2009.
In this context, modeling \((f_t)\) directly using a VAR is substantially more parsimonious than working with the static representation \((g_t)\). As shown by Bai2007 and Forni2009, if the dynamic factors \((f_t)\) follow a VAR\((p)\) process, then the associated static factors \((g_t)\) follow a VAR\((\tilde p)\) process, where \(\tilde p = 1\) if \(m \ge p\) and \(\tilde p = p - m + 1\) if \(m < p\).
When \(m \ge p\), estimation based on the static representation requires fitting a VAR\((1)\) in dimension \(qm\), involving \(q^2 m^2\) parameters. In contrast, direct estimation of \((f_t)\) requires fitting a VAR\((p)\) in dimension \(q\), involving only \(q^2p\) parameters. Since \(m \ge p\), and \(m > 1\) in genuine dynamic factor models, the reduction in dimensionality can be substantial. When \(m < p\), the estimation of the static representation requires \((p - m + 1) q^2 m^2\) parameters, which again strictly exceed \(p q^2\) whenever \(m > 1\). This highlights the efficiency gains from directly modeling the dynamic factors.
In this section, we assume $(f_t)$ follows a VAR$(p)$ process, and show that the lag order $p$ can be selected consistently by replacing the unobserved factors with their estimates in the standard BIC criterion.
Proposition (ref) only gives consistency of static representations $g_t$ of dynamic factors $f_t$, i.e., consistency up to a $qm\times qm$ dimensional invertible matrix. Assumption (ref) is used to identify and consistently estimate the dynamic factors $f_t$ up to a $q\times q$ dimensional matrix. bai_wang_2013 uses a similar assumption for identification. After establishing consistency of the dynamic factors $f_t$, we use the result to establish consistency of the BIC lag order selection in the following Proposition (ref).
Proposition (ref) shows that lag order selection in a VAR model for $(f_t)$ can be done in the standard way, with the penalty term only a function of $T$, not both $N$ and $T$ as in Theorem (ref).
We conduct Monte Carlo simulations of dynamic factor models to evaluate the finite-sample performance of tests for the number of dynamic factors and the filter length. Specifically, we simulate the model \[ x_{t} = \sum_{k=0}^{m-1} \lambda_k f_{t-k} + \varepsilon_t, \] where the idiosyncratic errors satisfy \[ \varepsilon_{it} = \sqrt{\frac{\theta(1-\rho^2)}{1+2J\beta^2}}\, e_{it}, \qquad e_{it} = \rho e_{i,t-1} + v_{it} + \sum_{1 \le |j| \le J} \beta v_{i-j,t}. \] The innovations \(v_{it}\) and the factor loadings \(\lambda_k\) are drawn independently from a standard normal distribution. This data-generating process is similar to those considered in Bai02 and Ahn-Horenstein-2013, except that we allow the filter length \(m\) to exceed one. The common factors \(f_t\) follow a VARMA\((1,1)\) process, \[ f_t = A f_{t-1} + u_t + \Theta u_{t-1}, \] under the following four specifications:
By construction, the variance of \(\varepsilon_{it}\) equals \(\theta\). The variance of the common component \(\sum_{k=0}^{m-1} \lambda_k f_{t-k}\) is given by \(m \cdot \mathrm{tr}(\Sigma_f)\), where \(\Sigma_f = \mathbb{E}[f_t f_t']\). We therefore set \(\theta = m \cdot \mathrm{tr}(\Sigma_f)\), ensuring that the signal-to-noise ratio equals one in all specifications.
The first specification corresponds to an exact dynamic factor model, in which both the dynamic factors and idiosyncratic errors are serially uncorrelated. The second specification represents an approximate dynamic factor model, allowing for heteroskedasticity as well as serial and cross-sectional dependence in the idiosyncratic errors. The third and fourth specifications examine settings in which the dynamic factors follow VAR\((1)\) and VMA\((1)\) processes, respectively.
We use the four specifications described in the previous section with $q_0=3$ and $m_0=3$.
In Table (ref), we use DR test to estimate the number $q$ given $m$. For $m=2$, the correct number of factors are $\lceil\frac{q_0m_0}{m}\rceil = 5$. For $m=3,4$, the correct number of factors are $q_0=3$. We find that all the tests perform well in finding the true number of factors given $m$.
In Table (ref), we use DR test to estimate the filter length $m$ given $q$. For $q=3,4$, the implied correct filter length is $\lceil\frac{q_0m_0}{q}\rceil = 3$, whereas for $q=5$, it is $\lceil\frac{q_0m_0}{q}\rceil = 2$.
For DGP3, in which the factors follow a VAR$(1)$ process, the DR test for $m$ performs poorly in finite samples when $m=4$. To understand the source of this behavior, consider the dynamic factor model with true filter length $m_0=3$: \[ x_t = \lambda_0 f_t + \lambda_1 f_{t-1} + \lambda_2 f_{t-2} + \varepsilon_t. \] If the factors satisfy the VAR$(1)$ representation \[ f_t = A f_{t-1} + u_t, \] then we can rewrite $x_t$ as \[ x_t = (\lambda_0 A^2 + \lambda_1 A + \lambda_2) f_{t-2} + (\lambda_0 A + \lambda_1) u_{t-1} + \lambda_0 u_t + \varepsilon_t. \]
This representation shows that a model with true filter length $m_0=3$ can effectively mimic a specification with $m=2$, particularly when $q>q_0$ and the model is overparameterized. In finite samples, the additional flexibility in the factor dimension allows the shorter filter to mimic the dynamics generated by the VAR structure. This phenomenon does not contradict the asymptotic validity of the test. Although not reported here, the simulation result shows that $\lceil\frac{q_0m_0}{q}\rceil = 3$ is correctly recovered when larger $N$ and $T$ are used.
We see finite sample performance of the $PC$, $DC$, $IC$ criteria in estimating $(q,m)$. We compare the results with the two existing methods given by Bai2007 and Amengual-Watson-2007\footnote{In both Bai2007 and Amengual-Watson-2007, we used IC test whenever it was needed to test for the number of static factors. When the lag order of the VAR($p$) model needed to be specified, we used $p=1$ for the first three specifications and $p=2$ for the fourth specification. For Bai2007, we used the default parameters $(\delta,m)=(0.1,1)$ with covariance matrix method. For Amengual-Watson-2007, we used the residuals from direct regression, i.e., $\hat Y_t^B$ in Amengual-Watson-2007.}. All the specifications use the second specifications for the penalty term, i.e., $g(N,T) = \left(\frac{N+T}{NT}\right)\log\left(\min(N,T)\right)$. In Table (ref) and (ref), PC, DC, and IC denote the information criterion tests proposed in this paper. BN and AW denotes tests proposed in Bai2007 and Amengual-Watson-2007 respectively. Results for testing filter length are missing in BN and AW, as they only test for the number of dynamic factors.
{
} In Table (ref), we see that all the tests perform well in DGP1. The PC and DC test outperforms both BN and AW significantly, especially when sample size is small. Performance deteriorates in the DGP2, where error terms are allowed serial correlation and heteroskedasticity. However, the PC test still performs well in finding the correct structure $(q_0,m_0)$.
{
} In Table (ref), we see that PC and DC criteria outperform both BN and AW in DGP3 and DGP4. In summary, PC test outperforms both BN and AW in all specifications considered. DC test outperforms both BN and AW in DGP1, DGP3, and DGP4.
In this section, we estimate the factor structure $(q,m)$ of dynamic factor models estimated using a large panel of U.S. macroeconomic variables. The Supplemental Materials (ref) additionally reports the estimated impulse responses of these variables to monetary policy shocks identified within the dynamic factor model framework.
The empirical analysis is based on the FRED--MD dataset of McCracken-Ng-2015, covering the period from March 1959 to March 2025. The sample comprises $T = 797$ observations on $N = 126$ macroeconomic time series. All variables are transformed in accordance with the procedures recommended by McCracken-Ng-2015.
To determine the factor structure, we implement the PC2, DC2, and IC2 criteria using a rolling window of fixed length equal to 10 years. The evaluation period begins in March 1969 and extends through March 2025. For the March of each year in this period, the criteria are computed using the information set consisting of the preceding 10 years of observations.
The DC2 criterion tends to select a larger number of factors, whereas the IC2 criterion typically yields a more parsimonious specification relative to the PC2 criterion. Based on the PC2 results, we find evidence in favor of four dynamic factors prior to the COVID-19 period. In contrast, the post-COVID-19 period is characterized by an increase in the estimated number of factors, with the criterion indicating the presence of seven dynamic factors.
The PC2 criterion consistently selects $m = 2$ across the sample. The DC2 criterion likewise predominantly indicates $m = 2$ in the pre-COVID-19 period, but points to an increase to $m = 3$ in the post-COVID-19 period. In contrast, the IC2 criterion generally favors a more parsimonious specification, selecting shorter filter lengths.
Based on the PC2 criterion, we find evidence supporting the specification $(q,m) = (4,2)$ for the majority of the sample period. The corresponding common component explains approximately 52% of the total variation in average, with this share increasing to about 65% in the post-COVID-19 period. A similar, but less drastic, increase in explained variation is also observed during the 2008 global financial crisis.
When adopting the static specification, $(q,m) = (8,1)$, the explained variation increases by approximately 3.6 percentage points. This gain can be interpreted as reflecting improved in-sample fit due to the greater flexibility of the model.
In contrast, holding the number of factors fixed while reducing the filter length to $(q,m) = (4,1)$ leads to a substantial decline in explained variation, on the order of 11.1 percentage points. This deterioration can be interpreted as evidence of misspecification.
This paper develops methods for determining the structure of dynamic factor models with finite filter length. As an intermediate step, we propose a computationally efficient alternating least squares (ALS) estimator that directly targets the dynamic factors and establish its consistency.
We introduce two procedures for jointly selecting the number of dynamic factors and filter length: a dynamic singular value ratio (DR) test extending Ahn-Horenstein-2013, and an information criterion method similar to Bai02 that penalizes both the number of dynamic factors and its implied static dimensions. Both procedures are shown to be consistent.
Empirical results using the FRED-MD dataset indicate evidence of filter length $m=2$ with four dynamic factors prior to COVID-19 period, which increase to seven post-COVID-19.
Proof of Proposition (ref) and Proposition (ref) are in Supplemental Materials.
Proof of Proposition (ref)
In the following Lemmas, we let $\hat X_{q,m}$ be defined as
where $(\hat F_{-k}^{(q)}, \hat \lambda_k^{(q)})_{k=0}^{m-1}$ is a minimizer of the least squares objective function given a structure $(q,m)$.
Let $q_0\geq0,$ $m_0\geq0$ be given, and consider disjoint partition of $\{(q,m):0\leq q\leq q_{max}, 0\leq m\leq m_{max}\}$ given by
where under $q_0=0$ or $m_0=0$ we let $S_1=\{(q,m):0\leq q\leq q_{max}, 0\leq m\leq m_{max}\}$ and $S_2=S_3=\emptyset$. Also, let $\mathcal{C}=S_1$ and $\mathcal{I}=S_2\cup S_3$.
Proof of Theorem (ref)
Proof of Theorem (ref)
Proof of Theorem (ref)