EconBase
← Back to paper

Quasi Maximum Likelihood Estimation of Non-Stationary Large Approximate Dynamic Factor Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

113,231 characters · 13 sections · 87 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Quasi Maximum Likelihood Estimation of Non-Stationary Large Approximate Dynamic Factor Models

center[center omitted — 294 chars of source]
abstractThis paper considers estimation of large dynamic factor models with common and idiosyncratic trends by means of the Expectation Maximization algorithm, implemented jointly with the Kalman smoother. We show that, as the cross-sectional dimension $n$ and the sample size $T$ diverge to infinity, the common component for a given unit estimated at a given point in time is $\min(\sqrt n,\sqrt T)$-consistent. The case of local levels and/or local linear trends trends is also considered. By means of a MonteCarlo simulation exercise, we compare our approach with estimators based on principal component analysis. Keywords: Dynamic Factor Model; EM Algorithm; Kalman Smoother; Stochastic trends.

\thispagestyle{empty}

\footnotetext{ A preliminary version of the results in this paper was made available with the title “Common factors, trends, and cycles in large datasets” (2017), by M. Barigozzi and M. Luciani, arXiv:1709.01445, and 2017-111, Board of Governors of the Federal Reserve System.\\

Disclaimer: the views expressed in this paper are those of the authors and do not necessarily reflect the views and policies of the Board of Governors or the Federal Reserve System. }

Introduction

In the last fifteen years, large dimensional stationary factor models have achieved great success in the economic profession, especially in forecasting macroeconomic variables Nowcasting, and are now a common tool in several policy institutions. However, macroeconomic time series are typically non-stationary due to the presence of common and idiosyncratic stochastic trends, and the practice of differencing the data to achieve stationarity is a problem that not always has a clear-cut solution. Take for example the case of the unemployment rate, which is a highly-persistent time series, but at the same time economic theory forbids it to have a unit root; or, take as another example the case of inflation, which shows periods of high-persistence in the late 70s early 80s, while more recently displays clear mean reversion. To avoid the risk of over- or under-differencing data, a Non-Stationary Dynamic Factor Model (NS-DFM) is then desirable, and it is studied in this paper.

The NS-DFM proposed in this paper captures several features of macroeconomic data as it takes into account the presence of common trends generating permanent fluctuations in the economy, as well as common transitory forces generating cyclical fluctuations. More technically, in our model, the common factors are a cointegrated vector process, thus containing both $I(1)$ trends and stationary components. Moreover, the NS-DFM addresses the possible presence of idiosyncratic trends, as well as the presence of secular (linear) trends, which can have either a constant slope (deterministic linear trends) or a time-varying slope (local linear trends).

In this paper, we study estimation of the NS-DFM by Quasi Maximum Likelihood (QML) implemented through the Expectation Maximization (EM) algorithm and the Kalman smoother (KS). Specifically, we extend the results in BLqml for the stationary case, to prove that when the common factors are the only source of non-stationarity, the common component estimated at a given point in time and for a given unit is $\min(\sqrt n,\sqrt T)$-consistent. We also discuss extensions to the cases of (i) unit roots in the idiosyncratic components, and (ii) local levels and local linear trends.

Estimation is implemented in two steps. First, given the observed data, by means of the KS we estimate the conditional mean of the latent factors, which, together with its associated conditional covariance matrix, we use to compute the expected log-likelihood of the model (E-step).\footnote{In a non-stationary setting the existence of the conditional mean of the factor as a minimizer of the mean-squared prediction error has been proved by hannan67 and sobel67.} Second, we maximize the expected log-likelihood with respect to the loadings and the other parameters of the model (M-step). The use of an iterative procedure to extract unobserved components in the case of non-stationary data was proposed since the original work by kalman60. Although this is not the first paper using these techniques for non stationary data, this is the first paper to address consistency of factors. Moreover, QML estimation of autoregressive processes with unit roots is a classical problem studied at length by the literature simsstockwatson,johansen91.\footnote{Solutions based on spectral analysis are also in bell84 and CT76.}

Estimation of NS-DFMs has also been studied by baing04 and BLL2 by PC analysis on differenced data. Both approaches are designed to account for non-stationary idiosyncratic components; however, only the latter is designed to deal with linear deterministic trends. bai04 has used a factor model to estimate common trends via PC on data in levels. However, because of its nature, that approach is valid only if all idiosyncratic components are stationary, i.e., only if data are cointegrated.

Compared to those PC based estimators, our approach has a number of practical advantages. First, it allows estimating the model even in the presence of missing values, which is crucial when using the model in real-time because macroeconomic data are published with delays and at non-synchronized dates. Second, it allows estimating jointly stochastic trends as well as (deterministic or local) linear trends, whereas baing04 and BLL2 are forced to remove the deterministic trends before running PC analysis. Third, it allows having time-varying parameters, such as, for example, the slope of linear trends. Fourth, it allows putting restrictions on the parameters, such as national accounts identities, or restrictions coming from economic theory.

From a theoretical point of view, our estimator converges at a faster rate than those of baing04 and BLL2. However, this faster convergence does not come for free. Indeed, our estimator is based on stronger assumptions than those of PC analysis: namely, it is derived under the assumption that we know which idiosyncratic components are $I(1)$ and which ones are stationary, and which series have a linear trend component. Under this assumption, we can model the $I(1)$ idiosyncratic components, and the time-varying slopes or means, as additional latent states in the KS, thus allowing to simultaneously estimate the entire model. This strategy is shown to work well in practice, provided the number of additional latent states is not too large.

The use of the EM in time series dates back to SS77, shumwaystoffer82, watsonengle83, quahsargent93, and SAZ13, among others. However, with the exception of the last two, none of the above works has considered the case of non-stationary data. Moreover, to the best of our knowledge, no asymptotic theory exists for the setting considered in this paper.

The rest of the paper proceeds as follows: in Section (ref), we present the NS-DFM and its assumptions. Estimation is outlined in Section (ref) where we also prove consistency. The extension to non-stationary idiosyncratic states is discussed in Section (ref). Numerical results are in Section (ref). Section (ref) concludes.

Model and assumptions

We define a NS-DFM driven by $q$ factors as

align[align omitted — 335 chars of source]

for $i=1,\ldots, n$, and $t=1,\ldots, T$. We let $\chi_{it}=\bm b_i^\prime(L) \bm f_t$. Then, $\bm\chi_{nt}=(\chi_{1t}\cdots\chi_{nt})^\prime$ is the common component, $\bm\xi_{nt}=(\xi_{1t}\cdots\xi_{nt})^\prime$ the idiosyncratic component, $\bm {\mathcal B}_n(L)=(\bm b_1(L)\cdots\bm b_n(L))^\prime$ the $n\times q$ polynomial matrix of factor loadings, $\bm f_t=(f_{1t}\cdots f_{qt})^\prime$ the $q$ factors, $\bm u_t=(u_{1t}\cdots u_{qt})^\prime$ the $q$ common shocks, $\mathbf e_{nt}=(e_{1t}\cdots e_{nt})^\prime$ the idiosyncratic shocks, and we also define $\bm\omega_{nt}=(\omega_{1t}\cdots\omega_{nt})^\prime$ and $\bm\eta_{nt}=(\eta_{1t}\cdots\eta_{nt})^\prime$.

We make the following assumptions.

ass\begin{inparaenum}[(a)] • for all $i\in\mathbb N$ and $z\in\mathbb C$, $\bm b_i(z)=\sum_{k=0}^s \bm b_{ik}z^k$, such that $\bm b_{ik}$ are $q\times 1$ and $s$ is a finite integer with $s\ge 0$; • for all $n\in\mathbb N$, let $\bm{\mathcal B}_{kn}=(\bm b_{1k}\cdots \bm b_{nk})^\prime$, then $\lim_{n\to\infty}\Vert n^{-1}\bm{\mathcal B}_{kn}^\prime\bm{\mathcal B}_{kn}-\bm\Sigma_{k}\Vert=0$, with $\bm\Sigma_{k}$ being $q\times q$, and $\bm \Sigma_0$ positive definite, while $\mbox{rk}(\bm\Sigma_k)\le q$ for $k=1,\ldots,s$, moreover, for all $i\in\mathbb N$ and $k=0,\ldots, s$, $\Vert\bm b_{ik}\Vert \le M_B$ for some finite positive real $M_B$ independent of $i$ and $k$; • $\bm\Gamma^{\Delta f}=\mathrm{E}_{\varphi_n}[\Delta\bm f_t\Delta\bm f_t^\prime]$ is $q\times q$ positive definite and there exists a finite positive real $M_f$, such that $\Vert\bm\Gamma^{\Delta f}\Vert\le M_f$; • $q$ is a finite positive integer, such that $q<n$ and is independent of $n$; • $\bm{\mathcal A}(z)=\sum_{k=1}^{p} \bm{\mathcal A}_{k}z^{k-1}$, such that $\bm{\mathcal A}_{k}$ are $q\times q$ and $p$ is a finite positive integer, and $\det(\mathbf I_q-\bm{\mathcal A}(z))\ne 0$ for all $z\in\mathbb C$ such that $|z|< 1$; • $\mbox{\upshape rk}(\bm{\mathcal A}(1))=d$ with $0<d\le q$; • $|\rho_i|\le 1$ for all $i\in\mathbb N$; • $\alpha_{i0}$ and $\beta_{i0}$ are finite reals. \end{inparaenum}
ass\begin{inparaenum}[(a)] • for all $t\in\mathbb Z$, $\bm u_t\sim \mathcal N(\mathbf 0_q,\bm\Gamma^u)$, such that $\bm\Gamma^u$ is $q\times q$ and positive definite, and $\mathrm{E}_{\varphi_n}[\bm u_{t}\bm u_{t-k}^\prime]=\mathbf 0_{q\times q}$ for all $k\ne 0$; • for all $t\in\mathbb Z$ and all $n\in\mathbb N$, $\mathbf e_{nt}\sim\mathcal N(\mathbf 0_n,\bm\Gamma_n^e)$, such that $\bm\Gamma_n^e$ is $n\times n$ and positive definite, and $\mathrm{E}_{\varphi_n}[\mathbf e_{nt}\mathbf e_{nt-k}]=\mathbf 0_{n\times n}$ for all $k\ne 0$; • for all $n\in\mathbb N$, $\Vert\bm\Gamma_n^e \Vert \le M_e$, for some positive real $M_e$ independent of $n$; • $\mathrm{E}_{\varphi_n}[\mathbf e_{nt}\bm u_{s}^\prime]=\mathbf 0_{n\times q}$ for all $n\in\mathbb N$ and $t,s\in\mathbb Z$; • for all $t\in\mathbb Z$ and all $n\in\mathbb N$, $\bm \omega_{nt}\sim\mathcal N(\mathbf 0_n,\bm\Gamma_n^\omega)$ and $\bm \eta_{nt}\sim\mathcal N(\mathbf 0_n,\bm\Gamma_n^\eta)$, such that $\bm\Gamma_n^\omega$ and $\bm\Gamma_n^\eta$ are diagonal, respectively with entries $0\le \sigma_{i\omega}^{2} < M_\omega$ and $0\le \sigma_{i\eta}^{2} < M_\eta$, for some positive reals $M_\omega$ and $M_\eta$ independent of $i$, and $\mathrm{E}_{\varphi_n}[\bm \omega_{nt}\bm \omega_{nt-k}]=\mathbf 0_{n\times n}$ and $\mathrm{E}_{\varphi_n}[\bm \eta_{nt}\bm \eta_{nt-k}]=\mathbf 0_{n\times n}$ for all $k\ne 0$; • $\mathrm{E}_{\varphi_n}[\bm\omega_{nt}\bm u_{s}^\prime]=\mathbf 0_{n\times q}$, $\mathrm{E}_{\varphi_n}[\bm\eta_{nt}\bm u_{s}^\prime]=\mathbf 0_{n\times q}$, for all $n\in\mathbb N$ and $t,s\in\mathbb Z$; • $\mathrm{E}_{\varphi_n}[\bm\omega_{nt}\mathbf e_{ns}^\prime]=\mathbf 0_{n\times n}$, $\mathrm{E}_{\varphi_n}[\bm\eta_{nt}\mathbf e_{ns}^\prime]=\mathbf 0_{n\times n}$, and $\mathrm{E}_{\varphi_n}[\bm\omega_{nt}\bm \eta_{ns}^\prime]=\mathbf 0_{n\times n}$, for all $n\in\mathbb N$ and $t,s\in\mathbb Z$. \end{inparaenum}
assFor any given $n\in\mathbb N$, there exists sets $\mathcal I_1\in\{1,\ldots,n\}$, $\mathcal I_a\in\{1,\ldots,n\}$, and $\mathcal I_b\in\{1,\ldots,n\}$, such that: \begin{inparaenum}[(a)] • if $i\in\mathcal I_1$ then $\rho_i=1$, while $\rho_i=0$ otherwise, moreover $\# \mathcal I_1=n_1$ such that $n_1n^{-1}\to 0$, as $n\to\infty$; • if $i\in\mathcal I_a$ then $\sigma_i^{2\omega}\ge C_\omega$ for some positive real $C_\omega$, while $\sigma_i^{2\omega}=0$ otherwise, moreover $\# \mathcal I_a=n_a$ such that $n_an^{-1}\to 0$, as $n\to\infty$; • if $i\in\mathcal I_b$ then $\sigma_i^{2\eta}\ge C_\eta$ for some positive real $C_\eta$, while $\sigma_i^{2\eta}=0$ otherwise, moreover $\# \mathcal I_b=n_b$ such that $n_bn^{-1}\to 0$, as $n\to\infty$. \end{inparaenum}

By Assumption (ref)(a) we are considering the case in which factors are loaded dynamically with a finite number of lags. We do not consider here the case of autoregressive filters, which has been studied in FHLZ17 in the stationary case. By Assumption (ref)(b) we are assuming for simplicity that all $q$ factors are pervasive at lag-zero, while at higher lags they might or might not have a pervasive effect depending on the rank of $\bm\Sigma_k$. In other words, (ref) can be seen as a factor model with $q(s+1)$ factors of which $q$ are strong factors, i.e., having an effect on all series, and the remaining $qs$ are weak factors, i.e., having an effect only on a subset of series.

By Assumptions (ref)(d) and (ref)(e), when $d<q$ we allow the dynamics of the factors to be driven by $(q-d)<q$ unit roots implying the presence of $(q-d)$ common trends stockwatson88JASA. Clearly, in this setting, the factors are cointegrated with cointegration rank $d$, thus representing the permanent and transitory aspects of macroeconomic dynamics. When $d=0$---i.e., the dynamics of the factors are driven by $q$ unit roots---the VAR for the common factors in levels in (ref) does not exist; instead, it exists a VAR for $\Delta \bm f_t$, or the factors can be modeled as $q$ independent random walks as in bai04. That said, the case $d>0$ is the relevant one, as there is full agreement in the economic profession that while some fluctuations in the economy are permanent (common trends), some others are only temporary.

Assumption (ref) characterizes the innovations of the model. In particular, by part (c) the idiosyncratic innovations $e_{it}$ are allowed to be mildly cross-correlated, thus implying that $\Delta x_{it}$ follows an approximate factor model. Moreover, by part (e) we allow some series to be driven by a time-varying intercept and/or a trend with time-varying slope, modeled as in a local level and local linear trend model, respectively harvey90. Notice that, if we set $\sigma_{i\eta}^2=0$, then the trend becomes deterministic with slope $\beta_{i0}$, which is fixed to a constant by Assumption (ref)(h), and similarly if we set $\sigma_{i\omega}^2=0$, we have a deterministic, hence constant, intercept term equal to $\alpha_{i0}$. Finally, by parts (d) and (f) all innovations are independent. Notice that gaussianity is not strictly needed, but it is a reasonable assumption in macroeconomics.

Under these assumptions, it can be shown that the covariance matrix of the differenced common component $\Delta\bm\chi_{nt}$ has at least $q$ and at most $q(s+1)$ eigenvalues that diverge linearly as $n\to\infty$. In particular, letting the covariance matrix of $\Delta \bm \chi_n$ be $\bm\Gamma_n^{\Delta\chi}$, and denoting as $\mu_{jn}^{\Delta\chi}$ the $j$-th largest eigenvalue of $\bm\Gamma_n^{\Delta\chi}$, Assumptions (ref)(b) and (ref)(c) imply that, for $j=1,\ldots, q$,

equation[equation omitted — 177 chars of source]

for some positive reals $\underline K_j$ and $\overline K_j$. Moreover, letting the covariance matrix of $\Delta \bm \xi_n$ be $\bm\Gamma_n^{\Delta\xi}$, and denoting as $\mu_{jn}^{\Delta\xi}$ the $j$-th largest eigenvalue of $\bm\Gamma_n^{\Delta\xi}$, Assumption (ref)(c), implies that

equation[equation omitted — 91 chars of source]

for some positive real $M_\xi$. From (ref) and (ref), and Assumption (ref)(e), by Weyl's inequality, the $q$ largest eigenvalues of the covariance matrix of $\Delta \mathbf x_{nt}$ diverge linearly in $n$, while all other $(n-q)$ eigenvalues stay bounded for all $n\in\mathbb N$.

Moreover, it can be shown that the $q$ largest eigenvalues of the spectral density of $\Delta \mathbf x_{nt}$ diverge with $n$ at all frequencies, but at zero-frequency, where, due to the presence of common trends, only $(q - d)$ eigenvalues diverge, all the others eigenvalues being bounded for all $n$ and all frequencies. Hence, by looking at the eigenvalues of the spectral density matrix of $\Delta \mathbf x_{nt}$ we can determine $q$ and $d$ (see hallinliska07, and BLL2, respectively). Moreover, notice that when all factors are pervasive at all lags, i.e., in Assumption (ref)(b) we let $\mbox{rk}(\bm\Sigma_k)=q$ for all $k=0,\ldots,s$, then (ref) holds for all $j=1,\ldots,q(s+1)$. Therefore, by looking at the eigenvalues of the covariance matrix of $\Delta \mathbf x_{nt}$, we can also determine $s$ dagostinogiannone12.

The model defined in (ref)-(ref) has $q$ latent states, given by the common factors $\bm f_t$, and additional latent states given by those idiosyncratic components that are autocorrelated as in (ref), and by the time-varying intercepts and trend slopes as in (ref) and (ref). These additional latent states are such that they satisfy the following assumption.

In other words, we are assuming that some, but not all, idiosyncratic components are $I(1)$, and that some, but not all, series have a time-varying intercept and/or a linear trend with time-varying slope. For simplicity, we are also assuming that stationary idiosyncratic components are serially uncorrelated.

We then make the following identifying assumptions.

assLet $\mathbf M_n^{\Delta \chi}$ be the $q\times q$ diagonal matrix with elements $\mu_{1n}^{\Delta \chi},\ldots,\mu_{qn}^{\Delta \chi}$, and let $\mathbf V_n^{\Delta\chi}$ be the $n\times q$ matrix having as columns the corresponding normalized eigenvectors. Then: \begin{inparaenum}[(a)] • $\Delta \bm f_t=(\mathbf M_n^{\Delta\chi})^{-1/2}\mathbf V_n^{\Delta\chi\prime} \Delta\bm\chi_{nt}$; • the entries of $\mathbf M_n^{\Delta\chi}$ are such that they satisfy (ref) and $\overline K_{j+1}<\underline K_j$ for $j=1,\ldots, q-1$; • the entries of $\mathbf V_n^{\Delta\chi}$ are such that $[\mathbf V_n^{\Delta\chi}]_{1j}>0$ for all $j=1,\ldots, q$. \end{inparaenum}

Parts (a) and (b) are standard in factor model literature for stationary processes and allow to identify the differenced factors up to a multiplication by a sign FGLR09,FLM13. We identify the first difference of the factors with the $q$ normalized principal components of $\Delta\bm\chi_{nt}$ and this implies in Assumption (ref)(b) that $\bm\Gamma^{\Delta f}=\mathbf I_q$. It can then be seen that the following must hold for the loadings

align[align omitted — 99 chars of source]

therefore we can choose $\bm{\mathcal B}_{0n}=\mathbf V_n^{\Delta\chi}(\mathbf M_n^{\Delta\chi})^{1/2}$, and in Assumption (ref)(a) we have that $\bm\Sigma_0$ is diagonal with entries given by $\lim_{n\to\infty} (n^{-1}\mu_{jn}^{\Delta \chi})$, which as requested are finite and positive because of (ref). Part (c) is a way to fix the sign indeterminacy in the identification of the factors. Once $\Delta\bm f_t$ and $\bm{\mathcal B}_{0n}$ are identified, then the remaining loadings are obtained by projecting $\Delta\mathbf x_{nt}$ onto the lagged factors.

The identifying restrictions in Assumption (ref) are particularly useful for initializing the EM algorithm with the PC estimator (see the next section). However, it has to be stressed that this identification does not provide any economic meaning to the factors. In other words we are not interested here in giving any interpretation of the factors, but we are just interested in the common component, which is always identified.

Estimation and asymptotic properties

Throughout the rest of the section we assume to observe the $nT$-dimensional vector $\bm X_{nT}=(\mathbf x_{n1}^\prime\cdots\mathbf x_{nT}^\prime)^\prime$ satisfying (ref)-(ref). In order to derive an estimator of the common component, we need to estimate the factors vector $\bm f_T=(\bm f_{1}^\prime\cdots \bm f_T^\prime)^\prime$ and the vector containing the true values of all parameters is \[ \bm\varphi_n=\left(\text{vec}(\bm{\mathcal B}_{0n}\cdots \bm{\mathcal B}_{sn})^\prime, \text{vech}(\bm\Gamma_n^e)^\prime, \rho_1,\ldots, \rho_{n_1} ,\text{vec}(\bm{\mathcal A}_1\cdots \bm{\mathcal A}_p)^\prime, \text{vech}(\bm\Gamma^u)^\prime,\sigma_{1\omega}^{2}\cdots\sigma^{2}_{n_a\omega},\sigma_{1\eta}^{2}\cdots\sigma^{2}_{n_b\eta}\right)^\prime, \] where, without loss of generality, we assumed that $\mathcal I_1=\{1,\ldots, n_1\}$, $\mathcal I_a=\{1,\ldots, n_a\}$, and $\mathcal I_b=\{1,\ldots, n_b\}$.

In this Section, we provide asymptotic results when $n_1=0$, $n_a=0$, and $n_b=0$, thus assuming that all idiosyncratic component are stationary and that no time-varying term is present. At first sight this might seem as a strong requirement, but notice that in our framework introducing non-stationary idiosyncratic components and/or local levels and/or local linear trends implies just adding latent states. We discuss this extension in Section (ref). Moreover, in (ref), we give all details of the EM algorithm together with explicit expressions for all estimators in the general case.

Without loss of generality, we fix $s=1,$ and we fix the VAR order in (ref) to $p=2$, thus $\bm{\mathcal A}(L)\equiv(\bm{\mathcal A}_1 L+\bm{\mathcal A}_2L^2)$, and, in this way the stationary component of $\bm f_t$ follows a non-trivial dynamics. For simplicity, we also assume that $\alpha_{i0}=0$ and $\beta_{i0}=0$.

The EM algorithm is an iterative procedure, which starts with an initial value of the parameters $\widehat{\bm\varphi}_n^{(0)}$, and at each iteration $k\ge 0$ produces estimates of the factors $\bm f_{t|T}^{(k)}$ (KS and E-step) and of the parameters $\widehat{\bm\varphi}_n^{(k+1)}$ (M-step). When the EM algorithm converges, say at iteration $k^*$, it gives the estimated common component $\widehat{\chi}_{it}=\widehat{\bm b}_{i0}^{(k^*+1)\prime}\bm f_{t|T}^{(k^*+1)}+\widehat{\bm b}_{i1}^{(k^*+1)\prime}\bm f_{t-1|T}^{(k^*+1)}$.

More in detail, the NS-DFM in (ref)-(ref) can be written as

align[align omitted — 521 chars of source]

By defining $\mathbf F_t=(\bm f_t^\prime\; \bm f_{t-1}^\prime)^\prime$ and $\bm\lambda_i=(\bm{b}_{0i}^\prime\;\bm{b}_{1i}^\prime)^\prime$, we see that, for given values of the parameters $\widehat{\bm\varphi}_n^{(k)}$, we can easily estimate the factors via the KS applied to the state-space form in (ref)-(ref). The estimated states are then $\mathbf F_{t|T}^{(k)}=\mathrm{E}_{\widehat{\varphi}_n^{(k)}}[\mathbf F_t|\bm X_{nT}]$, the first $q$-components of which give $\bm f_{t|T}^{(k)}=\mathrm{E}_{\widehat{\varphi}_n^{(k)}}[\bm f_t|\bm X_{nT}]$. Then, using the output of the KS, we can compute the expected log-likelihood, which is maximized by the loadings estimator $\widehat{\bm \lambda}_{i}^{(k+1)}\equiv(\widehat{\bm b}_{i0}^{(k+1)\prime} \; \widehat{\bm b}_{i1}^{(k+1)\prime})^\prime$, such that

align[align omitted — 525 chars of source]

The initial value of the parameters $\widehat{\bm\varphi}_n^{(0)}$ is determined as follows. For the loadings and the factors we use the approach proposed in BLL2, which makes use of the $q$ leading PCs of the model in first differences. Two comments are worth making. First, it important to stress that initializing the model in first differences (including when determining $q$ and $s$) is crucial, since it allows us to use PCs without incurring in spurious effects due to the presence of idiosyncratic unit roots (OW19), or linear trends (ngCG). Second, in light of the previous comment, this approach provides consistent estimates of the loadings, even in the case in which Assumption (ref) is satisfied with $n_1>0$ and $n_b>0$, but for constant intercepts and trend slopes baing04. In particular, our initialization delivers estimates of $\alpha_{i0}$ and $\beta_{i0}$, which, together with a given small initial value of the variances $\widehat{\sigma}_{i\omega}^{2(0)}$ and $\widehat{\sigma}_{i\eta}^{2(0)}$, can be used to update the slope state in (ref). Notice that the pre-estimators of those initial conditions do not need to be consistent for our results to hold. The initialization is completed by estimating the parameters of (ref) from an unrestricted VAR fitted on the estimated factors. This is a valid procedure when estimating an autoregressive model for cointegrated data (see simsstockwatson). Consistency of the pre-estimators of the loadings and VAR coefficients is proved in BLL2 (see also (ref)).

Finally, notice also that we initialize the KF by setting the initial value of the covariance of the factors, ${\mathbf P}_{0|0}$, to a very large value, as suggested by harvey90.

Consistency of the estimated common component follows.

propUnder Assumptions (ref), (ref), (ref), and if $\mbox{rk}(\bm\Sigma_k)=q$ for all $k=0,\ldots,s$, and $n_1=0$, $n_a=0$, and $n_b=0$, as $n,T\to\infty$, for any given $i=1,\ldots, n$ and $k=0,1$, $\min(\sqrt n,\sqrt T)\Vert \widehat{\bm b}_{ki}-\bm b_{ki}\Vert = O_p(1)$, and, for any given $t=\bar t,\ldots, T$, $\min(\sqrt n,\sqrt T)\Vert \widehat{\bm f}_{t|T}-\bm f_{t}\Vert= O_p(1)$. Moreover, $\min(\sqrt n,\sqrt {T})\,\Vert \widehat{\chi}_{it}-{\chi}_{it} \Vert = O_p(1)$, for any given $i=1,\ldots,n$ and $t=\bar t,\ldots, T$, with $\bar t\ge 2$.

The convergence rate depends on different ingredients. First, we show that the KS reaches a steady state within $\bar t$ periods, where $\bar t$ depends on the initial value $\mathbf P_{0|0}$ and, as shown in Section (ref), $\bar t$ is typically very small. Then, for the KS we show that, given the true parameters, the factors are $\sqrt N$-consistent. Third, given the true factors, the loadings estimator are consistent, with convergence rate $T$ for the loadings of the $I(1)$ components of the factors and convergence rate $\sqrt T$ for the loadings of the stationary component of the factors. As a result, for any given $i$, the whole loadings vector is $\sqrt T$-consistent, unless $d=q$, in which case each all $q$ factors are random walks and then the loadings vector would be $T$-consistent.

Under the assumption $n_1=0$, $n_a=0$, and $n_b=0$, our model is equivalent to the model studied in bai04, who considers estimation by means of PCs in levels. In this respect, we notice that the rates in Proposition (ref) are very similar to those in bai04. In other words, the QML estimator converges at the same rate than the estimator based on PC on the levels, which resembles the result in BLqml for the stationary case. However, in the simulation study in Section (ref) show that our estimator behave much better in finite samples.

$I(1)$ idiosyncratic components, local levels, local linear trends

If some idiosyncratic components are non-stationary, we can no longer use the EM algorithm described in the previous section. Indeed, when the residuals of (ref) are non-stationary, the M-step estimator of the loadings cannot be obtained by regressing $x_{it}$ is $I(1)$ onto $\bm f_t$ and $\bm f_{t-1}$. However, the case in which some idiosyncratic components are $I(1)$ is the relevant one for large macroeconomic datasets, since otherwise all data would be cointegrated. This is, for example, shown by the empirical results in OGAP, where the methodology proposed by baing04 for testing for idiosyncratic unit roots is applied on a standard US macroeconomic dataset.

In this Section, we adapt the EM algorithm to model non-stationary idiosyncratic components as well as local levels and local linear trends. In particular, we borrow from the literature on nowcasting with stationary factor models which models autocorrelated idiosyncratic components by treating them as additional latent states (see, e.g., banburamodugno14; and BGMR13).

Let us define $m=(n_1+n_a+n_b)$, as the number of additional latent states and recall that by Assumption (ref), $n^{-1}m\to 0$, as $n\to\infty$. Define also the set $\mathcal I_{m}=\mathcal I_1\cup\mathcal I_a\cup\mathcal I_b$, and notice that $\#\mathcal I_m\le m$, since it is possible that a variable has both a non-stationary idiosyncratic component as well as, for example, a linear trend. Then for all $i\in\mathcal I_{m}$, we replace the measurement equation (ref) with

align[align omitted — 115 chars of source]

such that Assumption (ref) still hold, and, moreover, letting $\bm\nu_{mt}=(\nu_{1t}\cdots\nu_{mt})^\prime$, for all $t\in\mathbb Z$, we assume $\bm\nu_{mt}\sim\mathcal N(\mathbf 0_m,\phi\mathbf I_m)$, with $\phi>0$, and $\mathrm{E}_{\varphi_n}[\bm\nu_{mt}\bm\nu_{mt-k}] = 0_{m\times m}$ for all $k\ne 0$. If $i\notin\mathcal I_m$, then (ref) stays the same. Moreover, we leave the dynamics of the factors in (ref) unchanged, while we change (ref) to

align[align omitted — 168 chars of source]

where Assumptions (ref)(b) and (ref)(c) still hold, and $\mathrm{E}_{\varphi_n}[\nu_{it} e_{js}]=0$, for all $t,s\in\mathbb Z$, all $i\in\mathcal I_m$ and all $j=1,\ldots, n$. Finally, according to (ref) and (ref), we have the state equations

align[align omitted — 202 chars of source]

such that $\mathrm{E}_{\varphi_n}[\nu_{it} \omega_{js}]=0$, and $\mathrm{E}_{\varphi_n}[\nu_{it} \eta_{js}]=0$, for all $t,s\in\mathbb Z$, all $i\in\mathcal I_m$, and all $j\in\mathcal I_a$ or $j\in\mathcal I_b$.

The model, which has as measurement equation either (ref) or (ref) if $i\in\mathcal I_m$, and which has as state equations (ref), and, if needed, also equations (ref), (ref) and (ref), has a compact state space form which is given in (ref), together with the details on its estimation via the EM algorithm. In particular, letting $w_{it}=\alpha_{it}+\beta_{it}t+\xi_{it}$, for all $i\in\mathcal I_m$ we show that, at a given iteration $k\ge 0$ of the EM algorithm, the M-step gives the loadings estimators:

align[align omitted — 534 chars of source]

where $\bm\lambda_i=(\bm{b}_{0i}^\prime\;\bm{b}_{1i}^\prime)^\prime$, while for $i\notin\mathcal I_m$ the loadings estimator is the same as in (ref). Formulas for all other estimators are given in (ref). In order to be able to compute $\widehat{\bm \lambda}_{i}^{(k+1)}$, we have to estimate the $m$ additional latent states $w_{it}$ and therefore we also need modify the KS accordingly (see (ref) for details).

In (ref), we provide an overview of the challenges involved by this task and we provide an informal derivation of the conditions necessary for consistent estimation, together with the related convergence rates. Three main results emerge. First, the new latent states can be recovered only if they display also some degree of cross-sectional correlation, as if they were driven by some common factor which is weakly pervasive for the whole panel. The intuition is that, if the additional latent states are completely uncorrelated across the components of $\mathbf x_{nt}$, then pooling many series does not help in recovering them, since their effect is always dominated by the factors.

Second, when the previous condition is verified, then we can still achieve $\sqrt n$-consistency for the estimated factors (as in the proof of Proposition (ref)), regardless of $m$, but provided that the variance of the measurement error $\nu_{it}$ in (ref) is fixed in such a way that $\phi=o(n^{-1})$, that is, it is asymptotically negligible. Indeed, the presence of $\nu_{it}$ represents a mis-specification of the original model in (ref), which needs to be introduced only as a numerical device, since the KF is not be defined if $\phi=0$. The smaller is $\phi$, the smaller the effect of the mis-specification is, and, therefore, the estimation of the factors is unaffected by the additional states.

As a consequence of this result, our estimator converges at a faster rate than those proposed by baing04 and BLL2, which are based on PC analysis on the differenced data. This faster convergence rate comes from the fact that we distinguish a priori between $I(1)$ and stationary idiosyncratic components. By contrast, due to differencing the estimator of baing04 and BLL2 essentially treat all idiosyncratic components as if they were $I(1)$. Of course, for the implementation of our estimator, it is crucial to be able to determine consistently which idiosyncratic component is $I(1)$---for example, using the test for idiosyncratic unit roots proposed by baing04.

Third, to achieve consistency of the additional latent states a necessary condition is $mn^{-1}\to 0$. This reflects the obvious intuition that the more latent states we need to estimate, the worse the performance of our estimator is going to be. Moreover, $\sqrt n$-consistency for the new states can be obtained for any $m$, but only if we choose an even smaller value of $\phi$, namely $\phi=o((m\sqrt n)^{-1})$.

We conclude with three remarks. First, the requirement that the new latent states display some degree of cross-sectional correlation is perfectly in line with Assumption (ref)(c) according to which the idiosyncratic components can be cross-correlated. Moreover, we can relax Assumptions (ref)(e) and (ref)(g) to allow for some correlation across the innovations $e_{it}$, $\omega_{it}$, and $\eta_{it}$ in (ref), (ref) and (ref). Indeed, it is reasonable to assume that local linear trends are shared by real variables (e.g., GDP and GDI), or that local levels are more apt to capture time-varying mean of groups of variables belonging, for example, to the labor market. Nevertheless, as shown in the proof of Proposition (ref), the fact that we estimate the $I(1)$ idiosyncratic components without modeling the cross-correlation between their innovations, will add miss-specification to our model, but will not affect the consistency of our estimates.

Second, as a far as estimation of the parameters given estimates of the states is concerned, we conjecture that nothing changes with respect to the results used in the proof of Proposition (ref), provided the states estimators are $\sqrt n$-consistent. Third, since the above are just asymptotic arguments, the choice of $\phi$ is not straightforward. A common way to proceed consists in initializing $\phi$ to be very small for all $m$ additional states and then update its estimate at each iteration of the EM algorithm, thus adding $m$ additional parameters. This is the way we implement the EM algorithm in the next section (see also (ref)).

MonteCarlo results

Throughout, we let $n\in\{75,100,200,300\}$, $T\in\{75,100,200,300\}$, $q\in\{2,4\}$, and $s\in\{0,1\}$, and we simulate data according to (ref), (ref), (ref), and (ref) as follows.

First, the factor loadings are such that $[\bm{\mathcal B}_{kn}]_{ij} \sim N(1,1)$ for $k=0,\ldots, s$, and then if $s=1$, for all $j=1,\ldots q$, we take $n/2$ randomly selected elements of $[\bm{\mathcal B}_{1n}]_{\cdot j}$ and we set them to zero. Second, for the common factors we set the VAR order $p=2$, and to generate $\bm{\mathcal A}(L)$ we use the Smith-McMillan factorization according to which $\bm{\mathcal A}(L)=\mathbfcal{U}(L) \mathbfcal{M}(L) \mathbfcal{V}(L)$, where $\mathbfcal{M}(L)= \mbox{diag} \left( (1-L)\mathbf I_{q-d}, \mathbf I_d\right)$, $\mathbfcal{V}(L)=\mathbf I_q$, and $\mathbfcal{U}(L)=(\mathbf I_q-\mathbfcal{U}_1 L)$, where $\mathbfcal{U}_1=\mu\,\widetilde{\mathbfcal{U}}_1(\nu^{(1)}(\widetilde{\mathbfcal{U}}_1))^{-1}$, where the diagonal elements of $\widetilde{\mathbfcal{U}}_1$ are drawn from a uniform distribution on $[0.5,0.8]$, while the off-diagonal elements from a uniform distribution on $[0,0.3]$, and $\mu=0.5$. In this way, $\bm f_t$ follows a VAR(2) with $q-d$ unit roots, or, equivalently, a VECM(1), where the number cointegration relations is set to $d=1$. The common innovations are such that $\bm u_t\stackrel{iid}{\sim} \mathcal{N}(\mathbf 0_q,\mathbf I_q)$, or $\bm u_t\stackrel{iid}{\sim} t_4(\mathbf 0_q,\mathbf I_q)$.

Third, each idiosyncratic component follows an AR(2) with roots $\rho_{i1}$ and $\rho_{i2}$, such that $\rho_{i1}=1$ if $\xi_{it}\sim I(1)$, while $\rho_{i1}=0$ otherwise, and $\rho_{i2}$ is drawn from a uniform distribution on $[0.2,0.6]$. We randomly select $n_1$ idiosyncratic components to have a unit root, with $n_1\in\{0,25,50,75,100\}$, provided $n_1<n$. The innovations are such that $\mathbf e_t\stackrel{iid}{\sim} \mathcal{N}(\mathbf 0_n,\bm\Gamma^e_n)$, or $\mathbf e_t\stackrel{iid}{\sim} t_4(\mathbf 0_n,\bm\Gamma^e_n)$ with $[\bm\Gamma^e_n]_{ij}=\tau^{|i-j|}$ if $\tau>0$, while, if $\tau=0$, $\bm\Gamma^e_n$ is diagonal with entries drawn from a uniform distribution on $[0.5,1.5]$. We set $\tau\in\{0,0.5\}$.

Fourth, we randomly select $n_b$ variables to have a non-zero linear trend, with $n_b\in\{0,25,50,75,100\}$, provided $n_b<n$. For those variables we draw $\beta_{i0}$ from a uniform distribution on ${[0.3,0.5]}$, but we set $\sigma_i^2=0$, thus considering only linear trends with constant slopes.

Last, we rescale the first differences of each common and idiosyncratic component in such a way that the share of variance of the $i$-th variable explained by the common component is $\theta(1+\theta)^{-1}$. We set $\theta=0.5$.

We consider $B=1000$ replications and we run the EM algorithm to estimate the NS-DFM by running the KS with $(q+n_1)$ latent states and estimating in the M-step only the diagonal terms of the idiosyncratic covariance matrix even when $\tau>0$. Similarly we do not add idiosyncratic states even when $\delta>0$. In other words, we always estimate a mis-specified model and, in this way, we are able to assess how robust our-estimators are with respect to mis-specifications.

In Table (ref), we report for different values of $n$ and for $t=1,\ldots, 10$, the trace of the one-step-ahead, KF, and KS MSEs when $q=2$, $s=1$, $T=100$, $\tau=0.5$, and $\delta=0.2$ (serially and cross-correlated idiosyncratic components). The MSEs are computed using the true simulated value of the parameters in order to verify numerically convergence to the steady-state. First, as $n$ grows, the one-step-ahead MSE reaches a steady state within maximum five time periods and $\mbox{tr}(\mathbf P_{t|t-1})/q\simeq 1$. This is consistent with the fact that due to the presence of unit roots we inizialize the filter with a vary large value of $\mathbf P_{0|0}$. Second, the KF and KS MSEs are very similar and both decrease to zero as $n$ grows and $\mbox{tr}(\mathbf P_{t|t})n/q$ and $\mbox{tr}(\mathbf P_{t|T})n/q$, computed when $t=10$, stabilize as $n$ grows thus showing that the rate of decrease is $n$.

table[table omitted — 4,458 chars of source]

In Table (ref) and in Table (ref), we report the relative MSE of our estimator over the MSE of the common component estimators obtained by PC as in bai04, and by PC in first differences as in baing04 and BLL2. Overall our estimator outperforms the others with the exception of the latter, which is show to perform better when $n_1$ becomes very large and about the same order of magnitude as $n$. This reflects the additional computational burden of our estimator which requires increasing the number of latent states when the idiosyncratic components are non-stationary and therefore we must include their dynamics in the model.

table[table omitted — 2,164 chars of source]
table[table omitted — 2,212 chars of source]

Some clarifications on the competing methods considered are necessary in order to interpret the results in Table (ref) and in Table (ref) (we refer to the original papers for details). First, notice that all alternative approaches considered here do not allow for dynamic loadings, so here they are implemented by computing the first $q(s+1)$ PCs.

Second, despite the common practice in the literature, bai04 did not propose its approach for factor model estimation, but rather to estimate common trends, and it is based on the crucial assumption of all idiosyncratic components being stationary. Indeed, we see from Table (ref) that when $n_1>0$ this approach fails completely.

Third, the baing04 approach delivers estimates of the common component which are obtained

inparaenum[($i$)] • by detrending the data by estimating the slope of the trend with the mean of the data in first difference; then • by estimating the factors in first differences; and, finally, • by cumulating the differenced estimator to obtain an estimate of the levels.

As such, this estimator is always subject to a location shift---it can be shown to converge to a Brownian bridge. Notice that this approach was introduced to test for the presence of unit roots rather than for factor model estimation, and, while the test is unaffected by location shifts, the use of the cumulated estimator for other scopes is not justified in general. As we see from Table (ref), this approach fails to consistently reconstruct the common component in all cases considered.

Fourth, the approach in BLL2 is based on the same ideas of baing04, but it takes care of the above mentioned issues related to detrending and cumulation, and, therefore, it is a valid alternative.

Concluding remarks

This paper considers estimation of large non-stationary approximate dynamic factor models by means of the Expectation Maximization algorithm, implemented jointly with the Kalman smoother. In our model the factors are a cointegrated vector process, thus containing both common $I(1)$ trends and stationary (cyclical) components. We show that, as the cross-sectional dimension $n$ and the sample size $T$ diverge to infinity, the common factors, the factor loadings, and the common component estimated are $\min(\sqrt n,\sqrt T)$-consistent at each $i$ and $t$.

Furthermore, we show that the model can be extended to account for the possible presence of idiosyncratic trends, as well as the presence of secular (linear) trends, which can have either a constant slope (deterministic linear trends) or a time-varying slope (local linear trends). Consistent estimation of this case is also considered.

Finally, the results in this paper provides the theoretical background for the application considered in OGAP, where the NS-DFM is used to estimate the output gap in the US.

\singlespacing {{ {.2cm} }} \setcounter{section}{0} \setcounter{subsection}{0} \setcounter{equation}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{footnote}{0} \gdef\thesection{Appendix \Alph{section}} \gdef\thesubsection{\Alph{section}.\arabic{subsection}} \gdef\thefigure{\Alph{section}\arabic{figure}} \gdef\theequation{\Alph{section}\arabic{equation}} \gdef\varphible{\Alph{section}\arabic{table}} \gdef\arabic{footnote}{\Alph{section}\arabic{footnote}}

Estimation in practice

Throughout, for simplicity, and without loss of generality, we let $p=2$ in the VAR for the factors (ref).

State space representation

Define $\mathbf F_t=(\bm f_t^\prime\cdots \bm f_{t-s}^\prime)^\prime$ be the $r$-dimensional vector of factors and $\bm\Lambda_n = (\bm {\mathcal B}_{0n}\cdots \bm {\mathcal B}_{sn})$ be the $n\times r$ matrix containing the factor loadings at all $s$ lags. Define also \[ \mathbf A = \left(

array[array omitted — 96 chars of source]

\right), \qquad \mathbf H =\left(

array[array omitted — 51 chars of source]

\right), \qquad \mathbf R_n=\left(

array[array omitted — 119 chars of source]

\right) \] Define also $\bm\alpha_{nt}=(\alpha_{1t}\cdots \alpha_{nt})^\prime$ and $\bm\beta_{nt}=(\beta_{1t}\cdots \beta_{nt})^\prime$. Then, using the process $\bm \nu_{nt}=(\nu_{1t}\cdots \nu_{nt})^\prime$ (see Section (ref)), we have the state space form for any $t=1,\ldots, T$:

align[align omitted — 1,321 chars of source]

where the innovations are such that

align[align omitted — 855 chars of source]

where $\bm{\mathcal S}_{1n}$, $\bm{\mathcal S}_{an}$, $\bm{\mathcal S}_{bn}$ are $n\times n$ diagonal matrices with entries $\{0,1\}$, $\bm\Gamma_n^\nu$, $\bm\Gamma_n^\omega$, and $\bm\Gamma_n^\eta$ are $n\times n$ diagonal matrices with entries $\sigma_{i\nu}^2$, $\sigma_{i\omega}^2$, and $\sigma_{i\eta}^2$, respectively, $\bm\Gamma^u$ satisfies Assumption (ref)(a), and $\bm\Gamma_n^{e}$ satisfies Assumptions (ref)(b) and (ref)(c). Specifically, letting $\mathcal I_m=\mathcal I_1\cup \mathcal I_a\cup \mathcal I_b$, the following constraints apply:

center[center omitted — 742 chars of source]

In a more compact form equation the state space model (ref) can be rewritten as

align[align omitted — 156 chars of source]

with obvious definitions of $\bm \Upsilon_n$, $\bm s_{nt}$, $ \bm\Theta_n$, and $\bm\zeta_{nt}$. This model is equivalent to (ref)-(ref) up to the error term $\bm\nu_{nt}$ which is needed to run the KS and notice that for those series such that $i\notin \mathcal I_1\cup \mathcal I_a\cup \mathcal I_b$, then we are setting $\nu_{it}=0$ and therefore $\xi_{it}=e_{it}$ is the measurement equation error. In other words $\bm\Gamma_n^\nu$ is always positive definite. As explained in Section (ref), this term is controlled by means of its variance $\phi$ and the smaller this is the better rate of convergence.

Initialization

Hereafter, for simplicity and without loss of generality, we let $s=1$, so that $r=q(s+1)=2q$, $\mathbf F_t=(\bm f_t^\prime\; \bm f_{t-1}^\prime)^\prime$ and $\bm\Lambda_n = (\bm{\mathcal B}_{0n}\; \bm {\mathcal B}_{1n})$.

The pre-estimators are defined as follows. Let $\widehat{\bm\Gamma}_n^{\Delta x}$ be the sample covariance matrix of the differenced data $\Delta \mathbf x_{nt}$ and denote as $\widehat{\mathbf M}_n^{\Delta x}$ the diagonal matrix with entries the $q$-largest eigenvalues of $\widehat{\bm\Gamma}_n^{\Delta x}$, and as $\widehat{\mathbf V}_n^{\Delta x}$ the $n\times q$ matrix of the corresponding normalized eigenvectors. We have the following pre-estimator of the loadings:

align[align omitted — 135 chars of source]

For all $i\in\mathcal I_a\cup\mathcal I_b$, let $\check{\alpha}_i$ and $\check{\beta}_i$ be the estimated parameter obtained by least squares of $x_{it}$ onto a constant and a time trend, and let $\check{x}_{it}=x_{it}-\check{\alpha}_i-\check{\beta}_it$. If $i\notin\mathcal I_a\cup\mathcal I_b$ define $\check {\alpha}_i=0$ and $\check{\beta}_i=0$. Then define: $\check{\mathbf x}_{nt}=(\check x_{1t}\cdots \check x_{nt})^\prime$. The pre-estimator of the factors is given by

equation[equation omitted — 151 chars of source]

Moreover, we define

equation[equation omitted — 313 chars of source]

Then, letting $\widetilde{\mathbf F}_{t}=(\widetilde{\bm f}_t^\prime\,\widetilde{\bm f}_{t-1}^\prime)^\prime$, we define

align[align omitted — 625 chars of source]

and $\widehat{\bm {\mathcal A}}_1^{(0)}$ is top left $q\times q$ block of $\widehat{\mathbf A}^{(0)}$, while $\widehat{\bm {\mathcal A}}_2^{(0)}$ is top right $q\times q$ block of $\widehat{\mathbf A}^{(0)}$. Also, we define \[ \widehat{\bm\Gamma}^{u(0)}=\frac 1T\sum_{t=3}^T(\widetilde{\bm f}_t-\widehat{\bm {\mathcal A}}_1^{(0)}\widetilde{\bm f}_{t-1}-\widehat{\bm {\mathcal A}}_2^{(0)}\widetilde{\bm f}_{t-2})(\widetilde{\bm f}_t-\widehat{\bm {\mathcal A}}_1^{(0)}\widetilde{\bm f}_{t-1}-\widehat{\bm {\mathcal A}}_2^{(0)}\widetilde{\bm f}_{t-2})^\prime. \] Moreover, letting $\widehat{\bm b}_{0i}^{(0)\prime}$ and $\widehat{\bm b}_{1i}^{(0)\prime}$ be the $i$-th row of $\widehat{\bm{\mathcal B}}_{0n}^{(0)}$ and of $\widehat{\bm{\mathcal B}}_{1n}^{(0)}$, respectively,

align[align omitted — 524 chars of source]

while $[\widehat{\bm\Gamma}_n^{e(0)}]_{ij}=0$ if $i\ne j$.\footnote{Alternatively, when $ i\notin\mathcal I_1$, we can set $[\widehat{\bm\Gamma}_n^{e(0)}]_{ii}= T^{-1}\sum_{t=2}^T ( x_{it} - \widehat{\bm b}_{0i}^{(0)\prime} \widetilde{\bm f}_{t}-\widehat{\bm b}_{1i}^{(0)\prime}\widetilde{\bm f}_{t-1} )^2$. }

Finally, if $i\in\mathcal I_a$ we define $\widehat{\alpha}_{i0}^{(0)}=\check{\alpha}_i$ and if $i\in\mathcal I_b$ we define $\widehat{\beta}_{i0}^{(0)}=\check{\beta}_i$, while $\widehat{\sigma}^{2(0)}_{i\omega}=10^{-2}$ and $\widehat{\sigma}^{2(0)}_{i\eta}=10^{-2}$, while if $i\in\mathcal I_m$, we fix $\widehat{\sigma}_{i\nu}^{2(0)}=10^{-5}$.

All the above quantities are collected into the vector of initial estimates of the parameters $\widehat{\bm\varphi}_n^{(0)}$.

E-step

To compute the expected log-likelihood of the model we run the KF-KS for the model in (ref) or (ref). The iterations of the KF-KS are standard and not reported. We just notice that, at iteration $k=0$ of the EM algorithm, the KF is inizialized as follows: we set $\bm f_{0|0}=\widetilde{\bm f}_0$ and, letting $\check{\mathbf A}^{(0)}=0.99 {\widehat{\mathbf A}^{(0)}}({\Vert\widehat{\mathbf A}^{(0)}\Vert})^{-1}$, we set

equation[equation omitted — 184 chars of source]

Then, at each iteration $k\ge 0$ the EM algorithm produces estimates of all states are computed using the parameters $\widehat{\bm\varphi}_n^{(k)}$ via KS. We obtain a vector $\bm s_{nt|T}^{(k)}=(\bm f_{t|T}^{(k)\prime}\, \bm\xi_{nt|T}^{(k)\prime}\,\bm\alpha_{nt|T}^{(k)\prime}\, \bm\beta_{nt|T}^{(k)\prime})^\prime$ with $(q+n_1+n_a+n_b)$ elements, such that

center[center omitted — 572 chars of source]

Finally, let $\bm w_{nt|T}^{(k)}=\bm\alpha_{nt|T}^{(k)}+\bm\beta_{nt|T}^{(k)}t+\bm \xi_{nt|T}^{(k)}$, which is $n$-dimensional with components $w_{it|T}^{(k)}$ for $i\in\mathcal I_{m}$ and zero otherwise.

We also define the $q\times q$ matrices

align[align omitted — 546 chars of source]

and the $r\times r$ matrices

equation[equation omitted — 389 chars of source]

We also define the $n_1\times n_1$ matrices $\mathbf P_{t|T}^{1(k)}$ and $\mathbf P_{t,t-1|T}^{1(k)}$ with entries

align[align omitted — 404 chars of source]

the $n_a\times n_a$ diagonal matrices $\mathbf P_{t|T}^{a(k)}$ and $\mathbf P_{t,t-1|T}^{a(k)}$ with entries

align[align omitted — 391 chars of source]

the $n_b\times n_b$ diagonal matrices $\mathbf P_{t|T}^{b(k)}$ and $\mathbf P_{t,t-1|T}^{b(k)}$ with entries

align[align omitted — 385 chars of source]

and the $\#\mathcal I_m\times \#\mathcal I_m$ diagonal matrix $\mathbf P_{t|T}^{w(k)}$ with entries

align[align omitted — 170 chars of source]

All those matrices are obtained from the KS. After the first iteration, for any $k\ge 1$ the KF is initialized with $\bm f_{0|0}^{(k)}=\bm f_{0|T}^{(k-1)}$ and $\mathbf P_{0|0}^{(k)}=\mathbf P_{0|T}^{(k-1)}$, which is defined as in (ref).

Denoting as $\underline{\bm\varphi}_n$ the generic values of the parameters, at each iteration $k\ge 0$, the expected log-likelihood is the given by (using the notation of (ref))

equation[equation omitted — 357 chars of source]

where $\bm X_{nT}$ is the $nT$-dimensional vector containing all data and $\bm S_{nT}$ is the $(r+n_1+n_a+n_b)T$-dimensional vector containing all latent states. In particular, denoting as $\bm F_T$ the $qT$-dimensional vector containing the $q$ factors, $\bm \Xi_{nT}$ the vector of all $I(1)$ idiosyncratic components, $\bm A_{nT}$ the vector of all time-varying intercepts, and $\bm B_{nT}$ the vector of all time-varying trend slopes, we have \[ \ell(\bm S_{nT};\underline{\bm\varphi}_n)= \ell(\bm F_{T};\underline{\bm\varphi}_n)+\ell(\bm \Xi_{nT};\underline{\bm\varphi}_n)+\ell(\bm A_{nT};\underline{\bm\varphi}_n)+\ell(\bm B_{nT};\underline{\bm\varphi}_n), \] since all groups of states are independent by assumption. Then,

align[align omitted — 524 chars of source]

M-step

As it is well known, the expected log-likelihood is maximized just by maximizing the first two terms in (ref). Therefore, at any iteration $k\ge 0$ of the EM algorithm, we have the following estimators. For the loadings (recall (ref)):

align[align omitted — 488 chars of source]

such that, $\widehat{\bm b}_{0i}^{(k+1)}$ is given by the first $q$-rows of $\widehat{\bm\lambda}^{(k+1)}_i$ and $\widehat{\bm b}_{1i}^{(k+1)}$ is given by the other $q$-rows.

For the VAR parameters:

align[align omitted — 646 chars of source]

and, letting, $\widehat{\bm {\mathcal A}}_1^{(k+1)}$ be the top-left $q\times q$ block of $\widehat{\mathbf A}^{(k+1)}$, and $\widehat{\bm {\mathcal A}}_2^{(k+1)}$ be the top-right $q\times q$ block of $\widehat{\mathbf A}^{(k+1)}$, we have

align[align omitted — 999 chars of source]

Moreover, the variances of the state residuals are given by:

align[align omitted — 902 chars of source]

while the variances of the residuals of the measurement equation are given by

align[align omitted — 899 chars of source]

Finally, we set $[\widehat{\bm\Gamma}_n^e]_{ij} =0$ for all $i,j=1,\ldots,n$ such that $i\ne j$.

Proof of Proposition (ref)

The proof follows the same steps as the proof of consistency in Theorem 1 in BLqml, and unless substantial differences emerge, we refer to results therein for detailed proofs of all the intermediate steps.

Throughout, for simplicity, and without loss of generality, we let $s=1$, so that $r=q(s+1)=2q$, and we let also $p=2$. Recall also that we are considering the case in which $n_1=0$, $n_a=0$, and $n_b=0$.\\

\it Stabilizability and detectability. Recall the state space form (ref)-(ref) of the NS-DFM

align[align omitted — 517 chars of source]

Then, (ref)-(ref) define a linear system with $r=2q$ latent states $(\bm f_t'\;\bm f_{t-1}')'$.

A linear system is stabilizable if its unstable states are controllable and all uncontrollable states are stable, and it is detectable if its unstable states are observable and all unobservable states are stable (see AM79).

Let us first show that (ref)-(ref) is stabilizable. Stability is dictated by the eigenvalues of the matrix of VAR coefficients,

align[align omitted — 162 chars of source]

Because of cointegration, ${\mathbf A}$ has $(q-d)$ unit eigenvalues corresponding to $(q-d)$ unstable states. Moreover, $(\mathbf I_q-\bm{\mathcal A}_1-\bm{\mathcal A}_2)=\mathbf a\mathbf b'$, where $\mathbf a$ and $\mathbf b$ have full column-rank $q\times d$ matrices, so that $\text{rk}(\mathbf a\mathbf b')=d$. Define the $q\times (q-d)$ matrices $\mathbf a_{\perp}$ and $\mathbf b_{\perp}$ such that $\mathbf a_{\perp}'\mathbf a=\mathbf b_{\perp}'\mathbf b=\mathbf 0_{(q-d)\times d}$. Then, since $\text{rk}(\mathbf a_{\perp}'\mathbf I_q)=(q-d)$, the unstable states are controllable because they satisfy the Popov-Belevitch-Hautus rank test (see franchi, Theorem 2.1, and AM07, Corollary 6.11, page 249). Clearly, ${\mathbf A}$ has also $(r-q+d)=(q+d)$ eigenvalues which are smaller than one in absolute value. Of these $q$ correspond to states which are uncontrollable because they are not driven by any shock, but are also stable since have no dynamics (see (ref)). The remaining $d$ states follow a stable VAR, hence are controllable.

Let us now show that (ref)-(ref) is detectable. First, notice that $\text{rk}(\bm{\mathcal B}_{0n})=q$ and $\text{rk}(\bm{\mathcal B}_{1n})=q$, because of Assumption (ref)(a) and we are assuming pervasive factors at all lags. Therefore, $\text{rk}(\bm{\mathcal B}_{0n} \mathbf b_{\perp})=(q-d)$ and $\text{rk}(\bm{\mathcal B}_{1n} \mathbf b_{\perp})=(q-d)$, which implies that the unstable states are observable because they satisfy the Popov-Belevitch-Hautus rank test (see franchi, Theorem 2.1, and AM07, Corollary 6.11, page 249). Since $\bm{\mathcal B}_{0n}$ and $\bm{\mathcal B}_{1n}$ have full column-rank there are no unstable unobservable states. \\

\it Estimation of factors given parameters. For the linear system in (ref)-(ref), define

equation[equation omitted — 159 chars of source]

Then, using the definitions in (ref) and (ref) and by setting $\mathbf K=\mathbf I_r$, the results in Lemmas 4, 5, and 6 of BLqml, still hold. In particular, since the system is stabilizable and detectable, the matrix $\mathbf P_{t|t-1}$ has a steady state denoted as ${\mathbf P}$, and there exists a positive integer $\bar n$, such that, for any $n\ge \bar n$, $$ \left\Vert{\mathbf P} - \left(

array[array omitted — 98 chars of source]

\right) \right\Vert \le Mn^{-1}, $$ for some positive real $M$. Moreover, notice that in the proof of Lemma 6 of \citet{BLqml} it is enough that $\Vert \mathbf A \Vert\le 1$, which is always satisfied because of Assumption \ref{ass:dynamic}(e). The definition of $\bar t$ is also unchanged.

Consistency can then be proved as in Proposition 1 of BLqml. By letting $\bm f_{t|T}$ be the KS estimate of $\bm f_t$ (given by the first $q$ components of $\mathbf F_{t|T}$), as $n\to \infty$, for any given $t\ge \bar t$, we have

equation[equation omitted — 76 chars of source]

This proves the analogous of Proposition 1 of BLqml for the NS-DFM.\\

\it QML estimation of parameters given factors. Recalling the definitions (ref) and (ref), the QML estimator of the loadings, for any $i=1,\ldots,n$, is given by

align[align omitted — 172 chars of source]

Because of Assumption (ref)(f), $\mathbf F_t$ is cointegrated and admits a common trends representation with $(q-d)$ common trends stockwatson88JASA. Therefore, we can find an orthonormal linear basis of dimension $(q-d)$ such that the projection of $\mathbf F_t$ onto this basis span the same space as the common trends. Collect the elements of this basis in the $r\times (q-d)$ matrix $\bm\gamma$, and denote as $\bm\gamma_\perp$ the $r\times(r-q+ d)$ matrix such that $\bm\gamma_\perp^\prime\bm\gamma=\mathbf 0_{(r-q+d)\times (q-d)}$. Then, consider the $r\times r$ linear transformation

equation[equation omitted — 242 chars of source]

where $\mathbf Z_{1t}$ has all $(q-d)$ components which are $I(1)$ while $\mathbf Z_{0t}\sim I(0)$ and is of dimension $(r-q+d)$. Moreover, for $\mathbf Z_{1t}$ we have the MA representation

equation[equation omitted — 92 chars of source]

with $\bm z_t$ being a $(r-q+d)$-dimensional vector with $\mathrm{E}_{\varphi_n}[\bm z_t]=\mathbf 0_{(q-d)}$, $\mathrm{E}_{\varphi_n}[\bm z_t\bm z_t^\prime]=\bm\Sigma_{z}$ positive definite and with finite norm, and $\mathrm{E}_{\varphi_n}[\bm z_s\bm z_t^\prime]=\mathbf 0_{(q-d)\times (q-d)}$, for any $s\ne t$. Moreover, $\text{rk}({\mathbf Q}(1))=(q-d)$, and $\sum_{k=0}^\infty\Vert \mathbf Q_k\Vert^2<\infty$.

Because of orthonormality $\bm{\mathcal D}^\prime\bm{\mathcal D}=\mathbf I_r$. Then, let $\bm\lambda_i^\prime\bm{\mathcal D}^\prime=(\bm\lambda_{i1}^\prime\; \bm\lambda_{i0}^\prime)$ such that (ref) reads (recall we are considering the case $\xi_{it}=e_{it}$),

equation[equation omitted — 116 chars of source]

and define also $\widehat{\bm\lambda}_i^{*\prime}\bm{\mathcal D}^\prime=(\widehat{\bm\lambda}_{i1}^{*\prime}\;\widehat{\bm\lambda}_{i0}^{*\prime})$. Since by construction $\mathbf Z_{1t}\mathbf Z_{0t}^\prime=\mathbf 0_{(q-d)\times(r-q+d)}$ and $\mathbf Z_{0t}\mathbf Z_{1t}^\prime=\mathbf 0_{(r-q+d)\times(q-d)}$, from (ref) and (ref), we have

align[align omitted — 566 chars of source]

First, consider the top left term on the rhs of (ref). From Hamilton and (ref), as $T\to\infty$,

equation[equation omitted — 250 chars of source]

where $\bm{\mathcal W}(\cdot)$ is a $(q-d)$-dimensional standard Wiener process. Thus this term is $O_p(1)$ and positive definite therefore invertible. Furthermore, for all $i=1,\ldots, (q-d)$ and all $t=1,\ldots, T$,

equation[equation omitted — 193 chars of source]

for some positive real $C_i$ and because of square summability of the MA coefficients in (ref) and since $\bm\Sigma_z$ has finite norm. Thus, since $\mathbf F_t$ and $e_{it}$ are gaussian and uncorrelated by Assumptions (ref)(a), (ref)(b), and (ref)(d), then they are also mutually independent, and we have

align[align omitted — 559 chars of source]

where we used (ref) and the fact that $\bm z_t$ is a white noise process with finite variance. From, (ref) and (ref), we have

equation[equation omitted — 106 chars of source]

Second, consider the bottom right term on the rhs of (ref). By the same arguments used to prove Lemma 8(i) in BLqml, we have

align[align omitted — 156 chars of source]

which imply

equation[equation omitted — 108 chars of source]

By substituting (ref) and (ref) into (ref) and since since $\bm{\mathcal D}$ does not depend on $T$, as $T\to\infty$, for any given $i=1,\ldots,n$, we have

equation[equation omitted — 102 chars of source]

Turning to estimation of the VAR coefficients, the QML estimator is given by

align[align omitted — 186 chars of source]

From (ref), we can also write

equation[equation omitted — 160 chars of source]

such that $\bm{\mathcal D}\bm u_t=(\mathbf v_{1t}^\prime\;\mathbf {v}_{0t}^\prime)^\prime$ where $\mathbf v_{1t}$ and $\mathbf v_{0t}$ are zero mean white noise processes of dimensions $(q-d)$ and $(r-q+d)$, respectively. Then, similarly to (ref), from (ref) and (ref), we have

align[align omitted — 782 chars of source]

Then, using the fact that $\mathbf v_{1t}$ and $\mathbf v_{0t}$ are gaussian white noise and therefore are martingale difference sequences, from Hamilton, it follows that

align[align omitted — 184 chars of source]

and, from Hamilton, it follows that

align[align omitted — 190 chars of source]

Substituting (ref), (ref), (ref) and (ref) into (ref), and since $\bm{\mathcal D}$ does not depend on $T$, we have

equation[equation omitted — 77 chars of source]

Finally, the process $T^{-2}\Vert\sum_{t=1}^T \mathbf Z_{1t} e_{it}\Vert$ is gaussian, with zero-mean and variance $O(T^{-2})$. Then, using Bonferroni inequality, and noticing that the rhs of (ref) does not depend on $i$, there exists a finite positive real $K_1$, independent of $i$, such that for all $\epsilon >0$

equation[equation omitted — 324 chars of source]

and, similarly, there exists a finite positive real $K_0$, independent of $i$, such that for all $\epsilon >0$

equation[equation omitted — 322 chars of source]

From (ref), (ref) and (ref), $\max_{i=1,\ldots,n}\sqrt T\Vert \widehat{\bm\lambda}_{i}^{*}-\bm\lambda_{i}\Vert = O_p(\sqrt{\log n})$. Then, for the estimator $\widehat{\bm\Gamma}_n^{e*}$ of $\bm\Gamma_n^e$ the same consistency proof given in Lemma 8(ii) in BLqml still holds. This proves the analogous of Lemma 8 in BLqml for the NS-DFM.\\

\it Estimation of factors given QML estimates of parameters. First, notice that using the notation of (ref) we have $\Vert\mathbf Z_{1t}\Vert=O_p(\sqrt T)$ and $\Vert\mathbf Z_{0t}\Vert=O_p(1)$. Then, the same steps leading to the proof of Lemma 9 in BLqml still hold, where, in particular, we can make use of the following relations (recall that $\bm{\mathcal D}^\prime\bm{\mathcal D}=\mathbf I_r$):

align[align omitted — 457 chars of source]

where (ref) holds because of (ref), (ref) and (ref), and, similarly, (ref) holds because of (ref), (ref) and (ref). Therefore, as $n,T\to\infty$, for any given $t\ge \bar t$, we have

equation[equation omitted — 93 chars of source]

This proves the analogous of Lemma 9 in BLqml for the NS-DFM.\\

\it Consistency of pre-estimator of parameters. Under Assumption (ref), the pre-estimators defined in Section (ref) are such that, for any given $i=1,\ldots, n$,

align[align omitted — 209 chars of source]

see BLL2 and also baing04. Moreover, it is easy to show that

equation[equation omitted — 243 chars of source]

Then,

align[align omitted — 870 chars of source]

which follows from (ref) and (ref), and noticing that, since $\Delta\bm f_{t-1}$ and $\Delta e_{it}$ are gaussian and uncorrelated by Assumptions (ref)(a), (ref)(b), and (ref)(d), then they are also mutually independent, and therefore

align[align omitted — 286 chars of source]

since $\mathrm{E}_{\varphi_n}[\Delta e_{is}\Delta e_{it}]>0$ only if $|t-s|\le 1$ because $\Delta e_{it}$ is an MA(1). For the same reason $\mathrm{Var}_{\varphi_n}[\Delta e_{it}]=2[\bm\Gamma_n^e]_{ii}$, thus from (ref) we have consistency of the diagonal elements of $\widehat{\bm\Gamma}_n^{e(0)}$, while for the off-diagonal terms \[ n^{-2}\sum_{\substack {i,j=1\\i\ne j} }^n\gamma_{ij}^2 \le n^{-2} \Vert \bm\Gamma_n^e\Vert_F^2 = n^{-2}\mbox{tr}(\bm\Gamma_n^e\bm\Gamma_n^e) \le n^{-1} \nu^{(1)}(\bm\Gamma_n^e) = n^{-1}\Vert \bm\Gamma_n^e\Vert^2\le n^{-1} \Vert \bm\Gamma_n^e\Vert_1^2\le n^{-1}M_e^2, \] because of Assumption (ref)(c).

Last, $\Vert\widehat{\mathbf A}^{(0)}-\mathbf A\Vert= O_p(\max(n^{-1/2},T^{-1/2}))$, because of BLL2. This proves the analogous of Lemma 10 in BLqml for the NS-DFM.\\

\it Convergence of EM estimator. The proof of Lemma 11 parts (i), (iii), and (iv) in BLqml holds also in the NS-DFM with no modifications. Thus, as $n,T\to\infty$, for any given $i=1,\ldots, n$,

equation[equation omitted — 184 chars of source]

From (ref), using the same reasoning leading to (ref), we also have, as $n,T\to\infty$, for any given $t=\bar t,\ldots, T$,

equation[equation omitted — 102 chars of source]

Therefore, from (ref) and (ref), as $n,T\to\infty$, for any given $i=1,\ldots,n$ and $t=(\bar t+1),\ldots, T$, \[ \min(\sqrt n,\sqrt T)\vert \widehat{\chi}_{it}-\chi_{it}\vert=O_p(1). \] This proves Proposition (ref). $\Box$

The case of additional latent states

For simplicity assume a static one factor model, thus $q=1$, $d=0$ and $s=0$, and also assume that $n_1=m$ while $n_a=0$ and $n_b=0$. Moreover, for any given $n\in\mathbb N$, assume that we always order the variables in such a way that $\mathcal I_m=\{1,\ldots,m\}$. Furthermore, define the $m\times 1$ vector $\mathbf e_{1t}=(e_{1t}\cdots e_{mt})^\prime$ with covariance matrix $\bm\Gamma_1^e$, and the $(n-m)\times 1$ vector $\mathbf e_{0t}=(e_{m+1t}\cdots e_{nt})^\prime$ with covariance matrix $\bm\Gamma_0^e$. Notice that $\bm\Gamma_1^e$ and $\bm\Gamma_0^e$ still satisfy Assumptions (ref)(b) and (ref)(c).

Recall also that $\mathrm{E}_{\varphi_n}[e_{it}u_{t}]=0$ for all $i=1,\ldots, n$, because of Assumption (ref)(d). Then, partition the $n\times 1$ loadings vector as $\bm{\mathcal B}_n=(\bm{\mathcal B}_{1}^\prime\ \bm{\mathcal B}_{0}^\prime)$, where $\bm{\mathcal B}_{1}=(\lambda_1\cdots \lambda_m)^\prime$ is $m\times 1$, and $\bm{\mathcal B}_{0}=(\lambda_{m+1}\cdots \lambda_n)^\prime$ is $(n-m)\times 1$. Consistently with Assumption (ref)(a) let us assume that $\lim_{m\to\infty}m^{-1}\bm{\mathcal B}_{1}^\prime\bm{\mathcal B}_{1}=1$ and $\lim_{n,m\to\infty}(n-m)^{-1}\bm{\mathcal B}_{0}^\prime\bm{\mathcal B}_{0}=1$. Let $\bm \xi_{t}=(\xi_{1t}\cdots \xi_{mt})^\prime$, while $\xi_{it}=e_{it}$ for $i=m+1,\ldots, n$, and $\bm\nu_t=(\nu_{1t}\cdots\nu_{mt})^\prime$. Last, let $\bm P=\mbox{diag}(\rho_1\cdots \rho_n)$. To avoid heavy notation we omit the dependence on $n$ and $m$ of the matrices and vectors considered.

We have the state space form

align[align omitted — 690 chars of source]

Moreover, the error terms in (ref) are such that

align[align omitted — 618 chars of source]

where $\sigma_u^2=\mathrm{E}_{\varphi_n}[u_{t}^2]$. Notice that $\bm\nu_t$, $\mathbf e_{0t}$, $\mathbf e_{1t}$, and $u_t$ are all white noise processes. Denote the $(m+1)$-dimensional state vector as $\bm s_t = (\bm\xi_t^\prime \ f_t)^\prime$, such that $\bm S_T=(\bm s_1^\prime\cdots \bm s_T^\prime)^\prime$. The parameter vector becomes $\bm\varphi_n=(\text{vec}(\bm{\mathcal B}_n)^\prime, \text{vech}(\bm\Gamma_n^e)^\prime, \rho_1,\ldots, \rho_n ,\text{vec}(\mathbf A)^\prime,\text{vec}(\mathbf H)^\prime)^\prime$. Notice that as shown later we need to control $\phi$ exogenously in order to achieve consistency, hence here we consider $\phi$ as given.

The log-likelihood of the data given the factors and the idiosyncratic states and for generic values of the parameters, $\underline{\bm\varphi}_n$, is

align[align omitted — 456 chars of source]

By maximizing (ref) we have the QML estimator of the loadings:

align[align omitted — 290 chars of source]

The formulas for the estimators obtained in the M-step at iteration $k\ge 0$ are obtained by repeating the same reasoning but when taking the expected log-likelihood.

Following the same reasoning leading in (ref), it is possible to show that at time $\bar t$ the one-step-ahead MSE of the linear system (ref) reaches a steady-state given by \[ \bm A = \left(

array[array omitted — 88 chars of source]

\right). \] Now, define the matrices

align[align omitted — 315 chars of source]

where $\bm C$ is $n\times (m+1)$ and $\bm B$ is $n\times n$, and notice also that $\bm B^{-1}$ is well defined because of Assumption (ref)(b). Then, given the true value of the parameters, $\bm\varphi_n$, for any given $t=\bar t,\ldots, T$, the KF estimator of the states is given by:

align[align omitted — 173 chars of source]

To prove consistency we need to apply Woodbury formula. However, this is not possible, indeed

align[align omitted — 432 chars of source]

and this is a singular matrix, because $\nu^{(m+1)}(\bm{\mathcal C}_1)= 0$ for all $m\in\mathbb N$ and $\nu^{(j)}(\bm{\mathcal C}_2)= 0$, for $j=2,\ldots, (m+1)$ and all $m\in\mathbb N$. Moreover, notice also that if $\phi=0$ then $\bm C^\prime\bm C$ will not be defined and the KF would have no sense.

However, since the idiosyncratic components are weakly cross-correlated by Assumption (ref)(c), it is reasonable to assume that they have a common factor. For simplicity, let us assume that

equation[equation omitted — 59 chars of source]

with $\mathrm{E}_{\varphi_n}[w_t^2]=1$ and $ \bm\beta=(\beta_1\cdots\beta_m)^\prime$ is such that $m^{-\alpha}\bm\beta^\prime\bm\beta=1$ for some real $\alpha$ and $\vert \beta_i\vert\le M_\beta$ for some positive real $M_\beta$ independent of $i$. In particular, notice that for $\bm\xi_t$ to be idiosyncratic, thus with $\Delta\bm\xi_t$ satisfying (ref), we must have $\alpha\in[0,1)$. This is equivalent to saying that the system is driven by a pervasive factors $f_t$ and a local factor $w_t$, which affects weakly only for the first $m$ units. Moreover, since $\bm\beta$ has full column rank, we can write $w_t=m^{-\alpha}\bm\beta^\prime \bm\xi_t$, and therefore

equation[equation omitted — 128 chars of source]

Using (ref) and (ref), the state space formulation in (ref) becomes:

align[align omitted — 714 chars of source]

where the errors have the same distribution as in (ref).

As a consequence, \[ \bm A=\left(

array[array omitted — 83 chars of source]

\right), \qquad \bm C= \left(

array[array omitted — 114 chars of source]

\right), \] while $\bm B$ is unchanged. The state vector is now $\bm s_t=(w_t\ f_t)^\prime$ and with these new definitions (ref) still holds, while (ref) becomes

align[align omitted — 470 chars of source]

which is not singular since $\nu^{(2)}(\bm C^\prime\bm C)=\phi^{-1}m^{\alpha}$. Moreover, from MK04 and Assumptions (ref)(b) and (ref)(c) we can show that $\nu^{(2)}(\bm A)= Mm^{-\alpha}$ for some positive real $M$. By following the same steps of the proof of Lemma 14 in BLqml, we have

equation[equation omitted — 136 chars of source]

from which we see that a necessary condition for consistency is $\phi\to0$ as $n\to\infty$. By the same arguments we also have

equation[equation omitted — 106 chars of source]

Substituting (ref) and (ref) into (ref), we have

align[align omitted — 459 chars of source]

for some positive real $K$. Then,

align[align omitted — 443 chars of source]

since $\mathrm{E}_{\varphi_n}[\nu_{it}\nu_{jt}]=0$ for all $i\ne j$, and where in the second relation we also used Assumption (ref)(a). Moreover,

align[align omitted — 257 chars of source]

by Assumptions (ref)(b) and (ref)(c). Therefore,

align[align omitted — 162 chars of source]

Consider the simplest case $\alpha=0$, then we need at least $\phi=o(n^{-1/2})$ to achieve convergence. In particular, if we set $\phi=n^{-1}$ we have that $f_{t|t}$ is $\sqrt n$-consistent, whereas $\vert w_{t|t}- w_t\vert=O_p(n^{-1}\sqrt m)$.

The previous result holds for any $m$ and $n$. However, what we are really interested in is the estimation of the vector $\bm\xi_t=\bm\beta w_t$. For given $\bm\beta$, letting $\bm\xi_{t|t}=\bm\beta w_{t|t}$, from (ref), we have \[ \Vert \bm\xi_{t|t}-\bm \xi_t\Vert = \sqrt{\sum_{i=1}^m \beta_i^2(w_{t|t}- w_t)^2} \le M_\beta \sqrt m \vert w_{t|t}- w_t\vert = O_p(\phi m^{-\alpha} m)+O(\phi\sqrt m). \] Hence, when $\alpha=0$, if we still set $\phi=n^{-1}$, we must have $mn^{-1}\to 0$ in order to have consistency. Furthermore, to achieve $\sqrt n$-consistency we would need either $mn^{-1/2}\to 0$ or an even smaller value of $\phi$.

To conclude, notice that in practice, although the model in (ref) is equivalent to the model in (ref), the latter has fewer states but more parameters to estimate and moreover estimation of $\bm\beta$ is not straightforward. In view of this comment the above derivations can just be seen as providing an intuition of the complexity involved by adding $m$ idiosyncratic latent states.