Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
668,380 characters · 45 sections · 239 citation commands
Quasi Maximum Likelihood Estimation and Inference .2cm of Large Approximate Dynamic Factor Models .2cm via the EM algorithm -.2cm
\thispagestyle{empty}
\footnotetext{A previous version of some results in this paper appeared in the paper “Common factors, trends, and cycles in large datasets”, available at \href{https://arxiv.org/abs/1709.01445}{arXiv:1709.01445} or at \href{ https://doi.org/10.17016/FEDS.2017.111}{ FEDS.2017.111}.\\
M. Barigozzi gratefully acknowledges financial support from MIUR (PRIN2020, Grant 2020N9YFFE).\\[.03cm]
Disclaimer: the views expressed in this paper are those of the authors and do not necessarily reflect the views and policies of the Board of Governors or the Federal Reserve System.\\[.03cm]
We thank for helpful comments: Majid Al-Sadoon, Leopoldo Catania, Giuseppe Cavaliere, Manfred Deistler, Massimo Franchi, Ivan Petrella, Esther Ruiz, Haihan Tang, Lorenzo Trapani, and Luca Trapin. }
Factor analysis can be considered a pioneering technique in unsupervised statistical learning GZ96. It originally gained popularity in the early decades of the twentieth century as a dimension-reduction technique used in psychometrics spearman04. Since then, it has become a classical method used for the statistical analysis of complex datasets in many human, natural, and social sciences (see, e.g., lawleymaxwell71). In the last thirty years, factor analysis has seen significant success in financial and macroeconometrics because it allows to analyze and predict economic activity by summarizing large panels of economic time series in a simple and effective way (see, e.g., the survey by stockwatson16 and references therein).
An $r$-factor model is defined by
where $x_{it}$ is the observation for the $i$th cross-section at time $t$, $\mu_i$ is a constant, and $\mathbf F_t$ and $\bm\lambda_i$ are $r$-dimensional latent column vectors of factors and factor loadings, with $r\ll n$. We call $\bm\lambda_i^\prime \mathbf F_t$ the common component and $\xi_{it}$ the idiosyncratic component. Throughout, we consider the standard case in which all $\{x_{it}\}$ are zero-mean weakly stationary processes or are the result of a transformation to stationarity.
Furthermore, in the case of time series, the factors are likely to be autocorrelated. For example, we can assume simple first order autoregressive dynamics:
with $\mathbf v_{t}$ being an $r$-dimensional vector of innovations. Likewise, the idiosyncratic components might be autocorrelated. The measurement equation (ref) and the state equation (ref) form a state-space model, or, equivalently, a Dynamic Factor Model (DFM) (this is a restricted version of the more general model by FHLR00, where factors can be loaded also with lags). Thanks to its simplicity and empirical success, the DFM is the most common approach to factor analysis of high-dimensional time series.
In large dimensional macroeconomic and financial datasets, the idiosyncratic components are likely to be also cross-correlated. Indeed, although macroeconomic or financial market dynamics are the main drivers of the comovement in these datasets, sectoral and local comovements are non-negligible sources of fluctuations. In the case of correlated idiosyncratic components the factor model is called approximate as opposed to an exact factor model having uncorrelated idiosyncratic components.
In an exact factor model a small number of variables is enough to estimate the loadings by Quasi Maximum Likelihood (QML), but we cannot consistently estimate the factors lawleymaxwell71. In an approximate factor model, we can disentangle the common and idiosyncratic components only in the extreme case when $n\to\infty$ chamberlainrothschild83---in other words, we do not suffer the typical “curse of dimensionality” but rather benefit from the “blessing of dimensionality”. In particular, when $n\to\infty$, we can consistently estimate the factors by some form of linear projection onto the estimated loadings. However, QML estimation of the loadings is now unfeasible since, in principle, it requires to also jointly estimate all the $T$ idiosyncratic (auto)covariance matrices, each of size $n\times n$, as well as the $T$ (auto)covariance matrices of the factors. Thus, we need to explore alternative approaches.
There are three main solutions to this problem. The first is Principal Component (PC) analysis, which delivers the optimal non-parametric estimator of a large approximate factor model, see, e.g., stockwatson02JASA and Bai03. The second is QML estimation based on a mis-specified exact model with no autocorrelations in the idiosyncratic components and the factors, and a diagonal or sparse idiosyncratic covariance matrix, see, e.g., baili16 and bailiao16. These two solutions only consider equation (ref), thus effectively estimating a static factor model, not a DFM. There is a third solution, which is the focus of this paper, that considers joint estimation of approximate DFM defined in (ref)-(ref): the Expectation Maximization (EM) algorithm (quahsargent93, DGRqml).
The EM algorithm is fully parametric and based on an iteration of two steps: (E-step) given the loadings and all the model's parameters, the factors and their second moments are estimated via the Kalman smoother and are used to compute the expected log-likelihood conditional on the observed data; (M-step) the expected log-likelihood is maximized to obtain a new estimate of the loadings and all the model's parameters. These steps are always implemented by considering a mis-specified likelihood with uncorrelated idiosyncratic components, thus making estimation feasible and producing estimators with closed-form expression As such, the EM algorithm should be regarded as an approximation of the QML estimation method because it maximizes a mis-specified likelihood of an exact DFM using an iterative procedure.
This paper focuses on the theoretical properties of the EM and Kalman smoother estimators. In a couple of breakthrough studies, DGRfilter,DGRqml provided the first fundamental theoretical treatment of these estimators, which quickly became popular in empirical macroeconomic research. Indeed, this approach allows the user to easily deal with data irregularities and missing values and impose restrictions that reflect any prior knowledge about the data on the model.\footnote{PC analysis in presence of missing data and parameter constraints has been studied by, e.g., baing19, fan2022we, XP20, among others.}
We make three main contributions. First, we prove that, as $n,T\to\infty$, the estimator of the loadings obtained via the EM algorithm converges asymptotically to a unique maximum of the likelihood. It is well known that in the Gaussian quasi-likelihood case, the EM algorithm produces approximate QML estimators wu83,BWB17. In this paper, we refine this result by showing that, as $n,T\to\infty$, the approximation error not only depends on the number of iterations but also becomes negligible. Moreover, the EM estimator of the loadings is asymptotically equivalent to the unfeasible Ordinary Least Squares estimator we would obtain if we had observed the factors. This result holds for any reasonable, but not necessarily consistent, pre-estimator of the loadings used to initialize the EM algorithm. We derive similar results for all the estimated parameters, namely, the idiosyncratic variances, the VAR coefficients, and the covariance matrix of the VAR residuals in equation (ref).
Second, we prove that, as $n,T\to\infty$, the estimator of the factors obtained via the Kalman smoother computed using the parameters estimated via the EM algorithm is equivalent to the unfeasible Weighted Least Squares estimator we would obtain if we had observed the loadings and idiosyncratic variances. As a by-product of this result, we also show that the Kalman smoother and filter estimators are asymptotically equivalent.
Third, we show that the EM estimator of the loadings is asymptotically equivalent to the PC estimator; hence, it has the same consistency rate ($\min(n,\sqrt T)$), is asymptotically normal and equally efficient. Likewise, the Kalman smoother has the same consistency rate ($\min(\sqrt n,T)$) as the PC estimator, it is asymptotically normal, and if the idiosyncratic covariance matrix is sufficiently sparse, it is more efficient than the PC estimator.
This paper is the first to fully characterize the asymptotic properties of the EM and Kalman smoother estimators. Other papers provide results that are close to ours, using more restrictive approaches or deriving only partial asymptotic results. baili16 considered QML estimation of the loadings for the static factor model (ref) only and did not study the convergence of the employed maximization algorithm. DGRfilter considered the Kalman smoother obtained using the PC estimator of the loadings but did not derive its asymptotic distribution and obtained a slower consistency rate. Last, DGRqml proved consistency of the Kalman smoother obtained from the EM algorithm but derived a slower rate and did not prove its asymptotic normality, nor did they prove consistency of the EM estimator of the loadings.
Our results lay the theoretical foundations for the wide empirical success of the EM algorithm for estimating large dimensional DFMs (see the next section for a list of applications) and answer two long-standing critiques. First, by showing the equivalence of the consistency rates of the EM and PC estimators, we reverse the belief that PC is a superior approach. Second, by providing the asymptotic distributions, we answer the call by TW87 and geweke93, who advocated a Bayesian approach based on the Gibbs sampler because the EM algorithm provides only point estimates.
The paper is organized as follows. In Section (ref), we briefly review the main applications of the EM algorithm and KS in factor analysis and alternative methods proposed to estimate DFM. In Section (ref), we describe the estimation and give a guide for implementing it. All assumptions are in Section (ref). The asymptotic results are in Section (ref). In Section (ref), we discuss efficiency of the EM estimator and the KS and compare them with the PC estimator. In Section (ref), we propose estimators of the asymptotic covariance matrices. In Section (ref), we present an extensive MonteCarlo study, and in Section (ref), we apply the EM algorithm to US macroeconomic data. Section (ref) concludes. The proofs of all theoretical results are in the Appendix.
The EM approach is arguably the most popular for conducting QML estimation of high-dimensional DFMs. This approach dates back to the 1970s when it was introduced in a low-dimensional setting by, e.g., SS77, shumwaystoffer82, watsonengle83, and harveypeters90, while its use in a high-dimensional setting was first suggested by quahsargent93 and then formalized by DGRqml.
The EM approach for high-dimensional DFMs has been extensively employed by empirical macroeconomic researchers, particularly those in central banks. Its most successful applications include (see also the survey by PRM20):
In addition to the EM approach, the literature has proposed several multi-step approaches to estimate the DFM in (ref)-(ref).
Approaches (a)-(d) consider estimation of the loadings, either via PC or QML, based only on (ref), while estimation of (ref) is in a second step. Approaches (e)-(h) consider joint estimation of (ref)-(ref). However, none of these fully develops an asymptotic theory for the proposed estimators. Moreover, (f) and (h) do not study the convergence of their numerical algorithms.
Last, an alternative to the EM algorithm is represented by Bayesian estimation of large DFMs using Gibbs sampling, see, e.g., by KOW03, lucianiricci, BW15, DGLM16, and koopman2017empirical, among many others.
To conclude, classical references for QML estimation of an exact factor model with no autocorrelations are, e.g., AR56 and AFP87, while tippingbishop99 suggest to simplify the maximization problem by considering a mis-specified likelihood with homoskedastic idiosyncratic components. baili12 extend classical QML estimation to the high-dimensional case. Alternatively, stockwatson89 combine QML estimation with the Kalman filter by using the prediction error likelihood. A review of these methods is in MBQML.
An $m\times m$ identity matrix is denoted as $\mathbf I_m$. Vectors are always considered as one-column matrices. An $m$ dimensional vector of ones is denoted as $\bm\iota_m$. An $m$ dimensional vector of zeros is denoted as $\mathbf 0_m$, an $m\times p$ matrix of zeros is denoted as $\mathbf 0_{m\times p}$.
The generic $(i,j)$ entry of a matrix $\mathbf A$ is denoted as $[\mathbf A]_{ij}$. Unless otherwise specified, we denote as $\nu^{(k)}(\mathbf A)$ the $k$-th largest eigenvalue of a generic squared matrix $\mathbf A$. The spectral norm for a real $p\times m$ matrix $\mathbf A$ is defined by $\Vert \mathbf A\Vert=(\nu^{(1)}(\mathbf A\mathbf A^\prime))^{1/2}$. The Frobenius norm is defined by $\Vert \mathbf A\Vert_F=(\mbox{tr}(\mathbf A\mathbf A^\prime))^{1/2}$. For a generic $p$-dimensional vector $\bm v=(v_1\cdots v_p)^\prime$, we consider the norms: $\Vert \bm v\Vert=(\sum_{j=1}^p v_j^2)^{1/2}$, and $\Vert \bm v\Vert_{\max}=\max_{j=1,\ldots, p} \vert v_j\vert$.
The indicator function on the event $A$ is denoted as $\mathbb I(A)$, i.e, $\mathbb I(A)=1$ if $A$ is true and 0 otherwise.
All random variables (scalars, vectors and matrices) are assumed to belong to $L_2((\Omega,\mathcal F, \mathrm P))$, where $(\Omega,\mathcal F, \mathrm P)$ is a common probability space. For a generic $p$-dimensional process $\{\mathbf y_t\}$ we adopt the following definitions.
Limits are always taken as $\min(n,T)\to\infty$ unless otherwise specified. We adopt the Landau $O(\cdot)$ and $o(\cdot)$ notation and the “in probability” $O_p(\cdot)$ and $o_p(\cdot)$ analogues. We denote convergence in probability and in distribution by $\stackrel{p}{\to}$ and $\stackrel{d}{\to}$, respectively.
For all quantities having a dimension growing with $n$ and/or $T$, we highlight such dependence. The true scalars, vectors or matrices are denoted as, e.g., $\sigma_i^2$, $\bm\Lambda_n$, $\bm\phi_n$, $\bm\theta$, $\bm F_T$,. The corresponding scalars, vectors or matrices containing generic values of parameters are underlined, so: $\underline{\sigma}_i^2$, $\underline{\bm\Lambda}_n$, $\underline{\bm\phi}_n$, $\underline{\bm\theta}$, $\underline{\bm F}_T$.
For a generic process $\{\mathbf y_t\}$ and any $T\in\mathbb N$, letting $\mathcal F^0_{-\infty}$ and $\mathcal F_T^\infty$ be the $\sigma$-algebras generated by $\{\mathbf y_t, \, t\le 0\}$ and $\{\mathbf y_t, \, t\ge T\}$, respectively, we define the strong mixing coefficients of $\{\mathbf y_t\}$ as $\alpha_y(T)=\sup_{A\in\mathcal F^0_{-\infty},\, B\in \mathcal F_T^\infty}\vert \mathrm P(A)\mathrm P(B)-\mathrm P(AB)\vert$.
In this section, we present the EM algorithm described by shumwaystoffer82 and implemented in the codes by DGRqml, which are available from theirs or ours webpages.\footnote{See the replication codes available at: \href{https://sites.uw.edu/dgiannon/domenico-giannone-s-homepage/}{https://sites.uw.edu/dgiannon/domenico-giannone-s-homepage/}, or\\ \href{http://www.barigozzi.eu/codes.html}{http://www.barigozzi.eu/codes.html}, or \href{https://sites.google.com/site/lucianimatteo/matlab-codes}{https://sites.google.com/site/lucianimatteo/matlab-codes}} This is the approach typically followed by applied researchers.
Let us assume to observe an $n$-dimensional stochastic process $\mathbf x_{nt}=(x_{1t}\cdots x_{nt})^\prime$ over $T$ periods. The DFM (ref)-(ref) reads:
where $\bm\Lambda_n=(\bm\lambda_1\cdots\bm\lambda_n)^\prime$ is the $n\times r$ matrix of factor loadings, $\mathbf v_t$ is an $r$-dimensional vector of factor innovations with covariance $\bm\Gamma^v=\mathbb{E}[\mathbf v_t\mathbf v_t^\prime]$, and $p_F$ is a finite integer such that $p_F\ge 1$. Throughout, we assume that the $r$-dimensional process of factors, $\{\mathbf F_t\}$, and the $n$-dimensional process of idiosyncratic components $\{\bm\xi_{nt}\}$ are zero-mean covariance stationary processes, so that $\{\mathbf x_{nt}\}$ is also covariance stationary and $\mathbb{E}[\mathbf x_{nt}]=\bm\mu_n$, with $\bm\mu_n=(\mu_1\cdots\mu_n)^\prime$.
Let us introduce the following notation. Define the $nT$-dimensional vectors $\bm X_{nT}=(\mathbf x_{n1}^\prime\cdots\mathbf x_{nT}^\prime)^\prime$ and $\bm \Xi_{nT}=(\bm\xi_{n1}^\prime\cdots\bm\xi_{nT}^\prime)^\prime$, and the $rT$-dimensional vector $\bm F_T=(\mathbf F_1^\prime\cdots\mathbf F_T^\prime)^\prime$. Moreover, let $\bm\Sigma_{n}^\xi$ be the diagonal matrix containing the $n$ diagonal terms of $\mathbb{E}[\bm\xi_{nt}\bm\xi_{nt}^\prime]$, which we denote as $\sigma_i^2$, $i=1,\ldots, n$, and ${\bm\Omega}^F_T=\mathbb{E}[\bm F_T\bm F_T^\prime]$.
In principle, to achieve QML, we should estimate (i) $nr$ loadings, (ii) $\simeq n^2T^2$ elements of the covariance matrix of $\bm \Xi_{nT}$, (iii) $\simeq r^2T^2$ elements of the covariance matrix of $\bm F_T$, and (iv) the $n$ constants in $\bm\mu_n$. Clearly the QML estimator of $\bm\mu_n$ is the sample mean $\bar{\mathbf x}_n=T^{-1}\sum_{t=1}^T\mathbf x_{nt}$ and we can then work with centered data: $\mathbf x_{nt}-\bar{\mathbf x}_n$ baili12. However, all other parameters depend on each other, so they must be estimated jointly---see, e.g., the first-order conditions in baili12,baili16 when the autocorrelation of the factors is not modeled. This is an unfeasible task since we have only $nT$ data points; therefore, we need some regularization to reduce the number of parameters to estimate.
To this end, first, we adopt the extreme form of regularization possible for the full-covariance matrix of $\bm\Xi_{nT}$ by replacing it with the diagonal matrix $\mathbf I_T\otimes \bm\Sigma_n^\xi$. Second, we assume that the parametric model (ref) describes the dynamics of the factors---thus, we write ${\bm\Omega}^F_T\equiv {\bm\Omega}^F_T(\bm{\mathcal A},\bm\Gamma^v)$, with $\bm{\mathcal A}=(\mathbf A_1\cdots \mathbf A_{p_F})$, in order to highlight its dependence on the VAR parameters. Thanks to these assumptions, we reduce the number of unknown parameters that need to be estimated to $Q_n=nr+n+r^2p_F+r(r+1)/2$, the elements of the vector $\bm{\varphi}_n=(\bm \phi_n^\prime\;\bm\theta^\prime)^\prime$, where $\bm\phi_n=(\text{vec}(\bm\Lambda_n)^\prime\; {\sigma}^2_{1}\cdots {\sigma}^2_{n})^\prime$ and $\bm \theta=(\text{vec}(\bm{\mathcal A})^\prime\; \text{vech}(\bm\Gamma^v)^\prime)^\prime$.
Our starting point is then the following log-likelihood, computed in the generic values of the parameters, $\underline{\bm\phi}_n$ and $\underline{\bm\theta}$:
where we removed the constant terms to simplify the notation. Since the assumptions in Section (ref) allow for correlations among idiosyncratic components, the expression in (ref) is a mis-specified or quasi log-likelihood. Such mis-specification is appealing because it coincides with the classical factor analysis under the exact factor structure, and is standard in DFM estimation DGRqml,baili16. Thus, the maximization of (ref) is QML estimation rather than ML estimation. As long as the idiosyncratic components are weakly correlated, the mis-specification introduced by maximizing (ref) has no effect on consistency but only on efficiency of the estimators (see Section (ref)). Hereafter, we denote the vector of QML estimators, which are the maximizers of (ref), as $\widehat{\bm\varphi}_n^*=(\widehat{\bm\phi}_n^{*\prime}\;\widehat{\bm\theta}^{*\prime})^\prime$.
Despite reducing the number of parameters to be estimated by introducing mis-specifications, direct maximization of (ref) is still unfeasible because $\bm\Omega_T^F$ is a full matrix, and to estimate its entries, we need to estimate the factors as well. Therefore, we still face a curse of dimensionality problem because we need to jointly estimate the $rT$ values of the factors and the $Q_n$ parameters using just the $nT$ observations of $\bm X_{nT}$.
There are three solutions to this problem. The first consists of rewriting the log-likelihood (ref) using its prediction error formulation, where the prediction errors and their covariance are obtained via the Kalman filter (see, e.g., harvey90 or DK01). This is the standard practice in low-dimensional DFMs stockwatson89. However, since there is no closed form solution for the QML estimator of the parameters obtained in this way, numerical maximization is required and this approach becomes quickly unfeasible even for moderate values of $n$. In high dimensions, JK15 propose to follow this approach subject to a preliminary step in which the data are projected onto a lower dimensional space but do not derive the theoretical properties of this estimator. In particular, it is not clear what are the effects of this first step on the asymptotic properties of the final estimator.
The second solution, proposed by baili16, further mis-specify the model by treating the factors as if they were serially uncorrelated. Thus, by imposing the standard identifying constraint $\mathbb{E}[\mathbf F_t\mathbf F_t^\prime]=\mathbf I_r$ and by replacing in (ref) the full-matrix ${\bm\Omega}_T^F(\underline{\bm{\mathcal A}}, \underline{\bm\Gamma}^v)$ with just $\mathbf I_T\otimes \mathbf I_r$, the log-likelihood is considerably simplified, and its maximization becomes feasible. However, because it does not exist a closed form solution, we still need a numerical approach to estimate the model, e.g., the iterative algorithm proposed by RT82 and reintroduced by baili12. To our knowledge, the convergence of the RT82 algorithm to the QML estimator has never been formally proved.
A third solution consists of computing an approximation of the QML estimator using the EM algorithm, as formalized by DGRqml. We consider this approach in this paper, and we refer to Table (ref) for its implementation.
The EM algorithm is an iterative procedure which allows for QML estimation in presence of missing data DLR77. In a nutshell, consider a given iteration $k\ge 0$ and assume to have an estimate of the parameters $\widehat{\bm\varphi}_n^{(k)}$. By taking expectations of the log-likelihood (ref) with respect to the conditional distribution of $\bm F_T$ given $\bm X_{nT}$ and computed using $\widehat{\bm\varphi}_n^{(k)}$, we get:
A maximum of the log-likelihood is then a maximum of the right hand side of (ref). Now, by definition of Kullback-Leibler divergence, for any $k\ge0$ it holds that
i.e., $\mathcal H(\underline{\bm\varphi}_n,\widehat{\bm\varphi}_{n}^{(k)})$ is maximum at $\underline{\bm\varphi}_n=\widehat{\bm\varphi}_{n}^{(k)}$. It is then enough to look for the maximum only of the expected full-information log-likelihood $\mathcal Q(\underline{\bm\varphi}_n,\widehat{\bm\varphi}_{n}^{(k)})$. This is accomplished in two steps: in the first step, for given estimated parameters $\widehat{\bm\varphi}_{n}^{(k)}$, we compute $\mathcal Q(\underline{\bm\varphi}_n,\widehat{\bm\varphi}_{n}^{(k)})$ using an estimate of the factors with their associated MSE; in the second step, for a given estimate of the factors, we maximize such log-likelihood to compute a new estimate of the parameters $\widehat{\bm\varphi}_{n}^{(k+1)}$. As shown below, this approach solves the curse of dimensionality problem, and its computational burden is minimal because all estimates have an explicit expression. Below we detail the main features of the two steps.
We obtain an initial estimate of the loadings and the idiosyncratic variances using the PC estimator, and of the VAR parameters by fitting a VAR on the PC estimator of the factors (see Appendix (ref)). Then, for any iteration $k\ge 0$, in the E-step, we compute the expected full-information log-likelihood, which we can decomposed as:
Consistently with the mis-specified log-likelihood (ref), the first term on the right-hand side of (ref) is:
which depends only on $\underline{\bm\phi}_n$ and is well defined as long as all the idiosyncratic components have finite positive variances.
As for the second term on the right-hand side of (ref), we assume that $\mathbf F_{t}=\mathbf 0_r$ for $t\le 0$ and consider $\ell(\bm F_T;\underline{\bm\theta})= \sum_{t=1}^T \ell(\mathbf F_t|\mathbf F_{t-1},\ldots, \mathbf F_{t-p_F};\underline{\bm\theta})$, which depends only on $\underline{\bm\theta}$. Then we have
which is well defined provided that the VAR innovations $\{\mathbf v_t\}$ have a finite full-rank covariance matrix.
Given (ref) and (ref), in order to compute the expected log-likelihood (ref) we need to compute the sufficient statistics:
Although exact expressions for the quantities in (ref) might be hard to compute, we can approximate them by using the output of the Kalman smoother, which gives the linear projection ${\mathbf F}_{t|T}^{(k)}=\mathrm {Proj}_{\widehat{\varphi}_{n}^{(k)}}[{\mathbf F}_t|\bm X_{nT}]$ and the associated conditional covariance and lag-$h$ autocovariance matrices ${\mathbf P}^{(k)}_{t|T}$ and ${\mathbf C}^{(k)}_{t,t-h|T}$, respectively (see Appendix (ref), for the explicit expressions). This approximation does not affect consistency of our estimators (see Proposition (ref)).
Hence, in the EM algorithm we let:
Summing up, in the E-step, we compute $rT$ values of the factors for a given value, $\widehat{\varphi}_{n}^{(k)}$, of the parameters. In general, this step is feasible because we have $T(n-r)\gg 0$ degrees of freedom. Moreover, since the Kalman filter and smoother are linear procedures entailing the inversion of a positive definite $r\times r$ matrix, we just have to compute $\simeq T$ recursions.
In the M-step, we have to maximize (ref) with respect to $\underline{\bm\varphi}_n$ to obtain a new estimate of the parameters $\widehat{\bm\varphi}_n^{(k+1)}$. This maximization has a closed form solution for all elements of $\widehat{\bm\varphi}_n^{(k+1)}$. Specifically, at a given iteration $k\ge 0$, by using (ref), we obtain the loadings estimators as
and $\widehat{\bm\Lambda}_n^{(k+1)}=(\widehat{\bm\lambda}_{1}^{(k+1)} \cdots \widehat{\bm\lambda}_{n}^{(k+1)} )^\prime$. Similarly, the estimator of the idiosyncratic variances is:
Because we consider a mis-specified log-likelihood, we do not estimate the out-of-diagonal terms of the idiosyncratic covariance matrix $\bm\Gamma_n^\xi$, and we act as if those terms are equal to zero.
For simplicity, let $p_F=1$ and denote $\mathbf A\equiv \mathbf A_1$. The estimator of $\mathbf A$ is then given by:
For $p_F>1$ we can simply write the VAR in companion form and derive the analogous of (ref).
Finally, the estimator of the covariance matrix of the VAR innovations $\{\mathbf v_t\}$ is:
Summing up, in the M-step, we need to compute $Q_n$ values of the parameters for a given estimator of the factors, $\mathbf F_{t|T}^{(k)}$, $t=1,\ldots, T$. This step is feasible because we used the mis-specified log-likelihood (ref) to estimate $\bm\phi_n$. Therefore, we decomposed our estimation problem into $n$ separate maximizations, each requiring estimating $r+1$ parameters using $T$ observations. In this way, estimating the high-dimensional parameter vector $\bm\phi_n$ becomes straightforward, and estimating $\bm\theta$ poses no problem because it is a low-dimensional problem with a closed-form solution.
Given the Gaussian quasi-likelihoods (ref) and (ref), it is easy show that, for any fixed $n$, there exists an $\omega>0$ such that for any $k\ge0$
where the first inequality follows from DLR77 and the second is due to strong concavity of $\mathcal Q(\cdot,\underline{\bm\varphi}_n)$ for any $\underline{\bm\varphi}_n$ wu83. Moreover, the left-hand-side of (ref) tends to zero as $k\to\infty$ wu83. Therefore, the EM algorithm defines a contractive map. Consequently, the sequence $\{\widehat{\bm\varphi}_n^{(k)}\}$ will converge to a maximum of the log-likelihood, as $k\to\infty$ (see Lemma (ref) for a formal proof when $n\to\infty$).
In practice, we stop the EM algorithm when the log-likelihood shows no further appreciable increase. This is ensured according to a standard convergence rule (see Appendix (ref)), depending on a pre-specified threshold $\varepsilon$.
We denote the last iteration of the EM algorithm as $k^*$. Upon convergence, the EM estimator of the parameters is $\widehat{\bm\varphi}_n\equiv \widehat{\bm\varphi}_n^{(k^*+1)}$. By running the Kalman smoother one last time using $\widehat{\bm\varphi}_n$, we have the estimator of the factors $\widehat{\mathbf F}_t=\mathbf F_{t|T}^{(k^*+1)}$, $t=1,\ldots, T$. Finally, we estimate the common components as $\widehat{\chi}_{it}=\widehat{\bm\lambda}_i^\prime\widehat{\mathbf F}_t$, where $\widehat{\bm\lambda}_i\equiv \widehat{\bm\lambda}_i^{(k^*+1)}$, $i=1,\ldots, n$.
This section presents the assumptions under which we can consistently estimate the DFM given in (ref)-(ref). These assumptions are stated for an infinite-dimensional stochastic process $\{ x_{it},\, i\in\mathbb N,\,t\in\mathbb Z\}$ of which $\{\mathbf x_{nt}=(x_{1t}\cdots x_{nt})^\prime,\,t\in\mathbb Z\}$ is an $n$-dimensional subprocess and the $T\times n$ matrix $(\mathbf x_{n1}\cdots\mathbf x_{nT})^\prime$ is an observed realization. Likewise, $\{\bm\xi_{nt}=(\xi_{1t}\cdots \xi_{nt})^\prime,\,t\in\mathbb Z\}$ is an $n$-dimensional sub-process of the infinite-dimensional stochastic process of idiosyncratic components $\{\xi_{it},\, i\in\mathbb N,\,t\in\mathbb Z\}$, and $\{\mathbf F_t=(F_{1t}\cdots F_{rt})^\prime,\, t\in\mathbb Z\}$ and $\{\mathbf v_t=(v_{1t}\cdots v_{rt})^\prime,\,t\in\mathbb Z\}$ are the $r$-dimensional processes of common factors and the corresponding VAR innovations, respectively. The $n\times r$ matrix of factor loadings $\bm\Lambda_n=(\bm\lambda_1\cdots\bm\lambda_n)^\prime$ forms a nested sequence as $n$ increases. Finally, as already noticed, the QML estimator of $\bm\mu_n$ is immediately obtained as the sample mean $\bar{\mathbf x}_n$. Thus, hereafter, for simplicity, we consider the DFM for pre-centered data, or, equivalently, we set $\bm\mu_n=\mathbf 0_n$.
Parts (a) and (b) imply that the loadings matrix has asymptotically maximum column rank $r$ (part (a)), and the factors have a finite full-rank covariance matrix (part (b))---these assumptions are similar to the requirements in baili16 and Bai03. Moreover, because of part (a), for any given $n\in\mathbb N$, all the factors have a finite contribution to each series (upper bound on $\max_{i=1,\ldots, n}\Vert\bm\lambda_i\Vert$), and there is at least one factor that contributes to at least one series (lower bound on $\max_{i=1,\ldots, n}\Vert\bm\lambda_i\Vert$). While the former condition is common, the latter is less standard but very mild as it simply guarantees that at least one loading is non-zero for any fixed $n\in\mathbb N$.
Part (c) implies the existence of a finite number of common factors $r$. In particular, $r$ is identified only for $n\to\infty$ (see also the next section). Hereafter, in parts (a) and (c), we can assume $N_0=N_1=N$, say, without loss of generality.
The remaining conditions of Assumption (ref) characterize the VAR for the factors in (ref). Part (d) implies that $\{\mathbf F_t\}$ is a weakly stationary process with a causal autoregressive representation. And in parts (e), (f), and (g), we assume that $\{\mathbf v_t\}$ is a zero-mean $r$-dimensional independent process with finite positive definite covariance matrix and finite summable 4th order cumulants. Parts (d) and (e) imply also that $\bm\Gamma^F$ is finite, as required in part (b).
Part (h) is an integral Lipschitz condition, which is satisfied by most continuous densities. This assumption guarantees that $\{\mathbf F_t\}$ is a strong mixing, or equivalently $\alpha$-mixing, process with mixing coefficients $\alpha_F(T)\le \exp\left\{-c_F T^{\gamma_F}\right\}$, for all $T\in\mathbb N$ and some finite positive reals $c_F$ and $\gamma_F$ independent of $T$ (PT85).\footnote{Independence in part (f) is not strictly necessary for having $\{\mathbf F_t\}$ strong mixing, as we could allow for GARCH effects by assuming geometric ergodicity of $\{\mathbf v_t\}$ instead (FZ06). Indeed, geometric ergodicity implies $\beta$-mixing, which implies strong mixing. } Strongly mixing factors with exponentially decaying mixing coefficients are directly assumed by FLM13.
Part (i) implies $\mathbf F_t=\mathbf 0_r$ for $t\le 0$. This assumption is standard and it fixes the initial conditions for the solution of the VAR in (ref).
Part (a) imposes that the idiosyncratic components have zero mean, and the idiosyncratic variances are positive and finite. We strengthen this assumption in part (f) by requiring that the whole idiosyncratic covariance matrix is positive definite.
Part (b) has a twofold purposes. First, it limits the degree of serial correlation of the idiosyncratic components by assuming standard geometric decay of the autocovariances. Second, it limits the degree of cross-sectional correlation between idiosyncratic components, a standard assumption for approximate DFMs. This assumption implies the usual conditions required by Bai03, FLM13, and baili16 (see Lemma (ref)).
In part (c), we assume that each idiosyncratic component is strongly mixing with exponentially decaying coefficients. This requirement is quite general since it allows the idiosyncratic to be non linear processes---FLM13 make the same assumption.
In part (d), we require finite summable fourth-order cumulants---a standard requirement found in the literature (see, e.g., Bai03 for the first condition, and baili16 for the second one---which, jointly with the mixing assumption in part (c), allows for consistent estimation of covariances by means of the sample covariances hannan.
In part (e), we assume a Central Limit Theorem, a standard assumption in the literature (e.g., Bai03, and baili16). This is a high-level requirement and in order to properly derive it, we should introduce some notion of dependence for random fields, which typically requires some ordering of the cross-sectional units. Now, in most applications, there is not a natural ordering of the variables. For this reason, we avoid to spell out primitive conditions that guarantee part (e) to hold.
We then add the natural requirement of independence between common shocks and idiosyncratic components.
Assumption (ref) implies that the factors and the common components are independent of the idiosyncratic components at all leads and lags and across all units. This assumption is compatible with the idea that the structural macroeconomic shocks driving the common component are independent of the idiosyncratic components representing measurement errors or local dynamics.\footnote{In principle, we could relax this requirement to allow for weak dependence as in, e.g., Bai03.}
Last, Assumption (ref) jointly with Assumptions (ref)(f), (ref)(g), (ref)(h), (ref)(c), and (ref)(d), implies that the process $\{\mathbf F_t\xi_{it}\}$ is strongly mixing with finite fourth moments bradley05. Then, from ibra62, we have the following Central Limit Theorem, for all $i\in\mathbb N$, as $T\to\infty$,
This last result is typically assumed in the literature---see, e.g., Bai03, and baili16.
We now state two more assumptions. These are needed only to derive some of our asymptotic results.
Assumption (ref) would hold if we directly assume joint Gaussianity of $\bm X_{nT}$ and $\bm F_T$. However, assuming Gaussianity might be too stringent, while Assumption (ref) is more general. For example, consider the case $r=1$, let $\mathrm f_t$ and $\bm{\mathrm X}_{nT}$ be realizations of the factor and the data, and let $f:\mathbb R^{nT+1}\to\mathbb R$ be the joint pdf of $F_t$ and $\bm X_{nT}$. Then, from steyn1960regression and kotz2004continuous, we see that if $f$ belongs to the family of multivariate Pearson-type distributions, and $\lim_{\mathrm f_t\to\pm \infty}\mathrm f_t^2 f(\mathrm f_t,\bm{\mathrm X}_{nT})=0$, then $\mathbb{E}[F_t|\bm X_{nT}=\bm{\mathrm X}_{nT}]$ is linear in $\bm{\mathrm X}_{nT}$. quahsargent93 and DGRqml made similar assumptions.
In part (a), we assume an exponential-type tail inequality for the common shocks, which implies that for all $s>0$, $t\in\mathbb Z$, and $j=1,\ldots,r$, $\mathrm P(|F_{jt}|>s)\le \exp\big\{-K_F s^{\delta_v}\big\}$, for some finite positive reals $K_F$ and $\delta_v\le 2$ independent of $t$ and $j$ (see Lemma (ref) and maleki). The same applies to part (b), which, by setting $n=1$, implies that for all $s>0$, $t\in\mathbb Z$, and $i\in\mathbb N$, $\mathrm P(\vert \xi_{it}\vert \ge s)\le \exp\big\{-K_\xi s^{\delta_\xi}\big\}$, for some finite positive real $K_\xi$ and $\delta_\xi \le 2$ independent of $t$ and $j$. This is because $\Vert \bm\lambda_i\Vert\le M_\lambda$ and $\sigma_i^2\ge C_\xi^{-1}$, for all $i\in\mathbb N$, by Assumptions (ref)(a) and (ref)(a), respectively. FLM13 also assume the factors and idiosyncratic components to belong to a distribution having exponentially decaying tails.
Depending on the values of $\delta_v$ and $\delta_\xi$, we are able to consider not only distributions with sub-Gaussian ($\delta_v,\delta_\xi=2$) or sub-exponential tails ($\delta_v,\delta_\xi=1$), which include the Laplace and the Generalized Error distribution vershynin18, but also distributions with sub-Weibull tails ($\delta_v,\delta_\xi<1$), which can mimic a heavy tail behavior even if all moments exist MN98,KC18,vladimirova2020sub. This is clarified in Remark (ref) below.
In general, part (b) is a Bernstein-type inequality implying that the weighted sums of the idiosyncratic components have Gaussian tails, as $n\to\infty$. This is a high-level requirement and, as for Assumption (ref)(e) in order to properly derive it, we should introduce some notion of dependence for random fields. Here, instead, we just notice that in the simplest case of cross-sectionally independent idiosyncratic components, this condition would be a direct consequence of maleki (see also vladimirova2020sub for a similar result, and vershynin18).
Finally, as a consequence of parts (a) and (b), jointly with Assumptions (ref)(f), (ref)(h), (ref)(c), and (ref), we can show that not only $\{\mathbf F_t\xi_{it}\}$ is a strongly mixing process but its components have also exponentially decaying tails. Therefore, for all $T\in\mathbb N$ and all $i\in\mathbb N$, the following Bernstein-type inequality holds:
for some finite positive reals $\kappa_{3}$, $\kappa_{4}$, and $\beta<1$ independent of $i$ and $T$ (MPR11, bosq12, and Lemma (ref)).
To identify the DFM, we need to address four issues: first, we need to identify the number of factors; second, we need to identify the true parameters of the model; third, we need to ensure that the linear system is identified; and fourth, we need to guarantee the existence of the maxima of the log-likelihood.
Starting with the number of factors $r$, let the covariance matrix of $\{\bm \chi_{nt}\}$ be $\bm\Gamma_n^{\chi}=\bm\Lambda_n\bm\Gamma^F\bm\Lambda_n^\prime$ and denote as $\mu_{jn}^\chi$ the $j$-th largest eigenvalue of $\bm\Gamma_n^{\chi}$, then Assumptions (ref)(a) and (ref)(b) imply that, for any $j=1,\ldots, r$,
for some finite positive reals $\underline C_j$ and $\overline C_j$. Furthermore, from Assumption (ref)(b) it follows that the largest eigenvalue of the idiosyncratic covariance matrix, $\bm\Gamma_n^\xi$, denoted as $\mu_{1n}^\xi$, is such that
From conditions (ref) and (ref) and Weyl's inequality, it follows that the $r$ largest eigenvalues of the covariance matrix of $\{\mathbf x_{nt}\}$ diverge linearly in $n$, whereas all remaining eigenvalues stay bounded for all $n\in\mathbb N$. Lemma (ref) proves this result, condition (ref), and condition (ref)---these conditions are directly imposed by DGRqml.
The asymptotic behavior of the eigenvalues of the covariance matrix allows for the identification of the common and idiosyncratic component when $n\to\infty$ and it is the basis for all existing methods for determining $r$ (see, e.g., baing02).
All the structures equivalent to (ref)-(ref) can be obtained through an $r\times r$ invertible matrix $\mathbf R$, as follows: \[ \mathbf F_t^o=\mathbf R^{-1}\mathbf F_t,\;\; \bm{\Lambda}_n^{o}=\bm{\Lambda}_n\mathbf R, \;\; \mathbf A^o(L)=\mathbf R^{-1}\mathbf A(L)\mathbf R, \;\; \bm\Gamma^{vo} = \mathbf R^{-1}\bm\Gamma^v,\;\; \bm\Gamma^{\xi o}_n=\bm\Gamma^{\xi}_n. \] Under such relationships, using only first- and second-moment information we cannot distinguish the model specified by $\bm{\Lambda}_n^{o}$, $\mathbf A^o(L)$, $\bm\Gamma^{vo}$, and $\bm\Gamma^{\xi o}$, from the one given by $\bm{\Lambda}_n$, $\mathbf A(L)$, $\bm\Gamma^v$, and $\bm\Gamma^{\xi}$. Once the loadings and the factors are identified, then $\mathbf A(L)$ and $\bm\Gamma^v$ are also identified; $\bm\Gamma_n^\xi$ is always identified.
To identify the model, we need enough a priori structure to preclude any but the trivial transformation $\mathbf R=\mathbf I_r$. This can be achieved by imposing additional $r^2$ identifying constraints. Let $\mathbf M_n^\chi$ be the $r\times r$ diagonal matrix with as elements the eigenvalues $\mu_{jn}^\chi$, $j=1,\ldots, r$, of the covariance matrix of the common component, $\bm\Gamma_n^\chi$, sorted in descending order. Let $\mathbf V_n^\chi$ be the $n\times r$ matrix with the corresponding normalized eigenvectors as columns. Then, we assume the following identifying constraints.
Part (a) is standard. Since the eigenvalues of $\bm\Sigma_\Lambda\bm\Gamma^F$ are equal to the $r$ non-zero eigenvalues of $\lim_{n\to\infty}n^{-1}\bm\Gamma_n^\chi$, given by $\lim_{n\to\infty}n^{-1}\mathbf M_n^\chi$, it implies that in (ref) we have $\overline C_{j}<\underline C_{j-1}$ for any $j=2,\ldots, r$, and, thus, it avoids the uninteresting difficulties related with asymptotically multiple eigenvalues, which would require more restrictions to identify the space spanned by the columns of $\bm\Lambda_n$.
Part (b) is similar to what is usually imposed in PC estimation (see Remark (ref) below). It implies that $\bm\Gamma_n^\chi=\bm\Lambda_n\bm\Lambda_n^\prime=\mathbf V_n^\chi\mathbf M_n^\chi\mathbf V_n^{\chi\prime}$. Hence, the $r$ non-zero eigenvalues of $\lim_{n\to\infty}n^{-1}\bm\Gamma_n^\chi$, given by $\lim_{n\to\infty} n^{-1}\mathbf M_n^\chi$, coincide with the diagonal entries of $\bm\Sigma_\Lambda$, which are then distinct because of part (a). Since part (b) concerns second moments and sums of squares, it allows us to identify $\bm\Lambda_n$ only up to a sign. To achieve global identification we must fix also the column sign of $\bm\Lambda_n$ which is done through part (c) (baing13).
Summing up, in part (b), we are imposing $r^2$ restrictions: $r(r + 1)/2$ by requiring orthonormality of the factors, and $r(r-1)/2$ by requiring that $\bm\Sigma_\Lambda$ is diagonal. Consistently with the fact that in the typical empirical applications the focus is on the common component only, these assumed identification conditions do not provide economic meaning to the factors; in this sense ours is an exploratory rather than confirmatory factor analysis. Moreover, and most importantly, the restrictions are imposed only in the limit $n,T\to\infty$. Thus, in our setting, the model is only asymptotically identified. This is enough to derive our asymptotic theory.
From our Assumptions (ref)(a), (ref)(d), and (ref)(e), we can show that the state space formulation (ref)-(ref) is minimal and stable, for all $n>N$, where $N= N_0$ in Assumption (ref)(a). It is convenient to consider here the factorization $\bm \Gamma^{v}=\mathbf H\mathbf H^\prime$ for some $r\times r$ matrix $\mathbf H$ having full rank. Consider again for simplicity the case $p_F=1$ with $\mathbf A \equiv \mathbf A_1$. Then, stability, which requires $\vert \nu^{(1)}(\mathbf A)\vert <1$, is a direct consequence of stationarity in Assumption (ref)(d). While minimality holds because the couple $(\mathbf A, \mathbf H)$ is controllable due to Assumption (ref)(e) and the couple $(\mathbf A, \bm\Lambda_n)$ is observable for all $n>N$ because Assumption (ref)(b) implies that $\text{rk}(\bm\Lambda_n)=\text{rk}(\mathbf V_n^\chi)=r$, for all $n>N$.\footnote{A linear system with $p_F=1$, is controllable if and only if $\text{rk}[\mathbf H\; (\mathbf A\mathbf H)\cdots (\mathbf A^{(r-1)}\mathbf H)]=r$ and it is observable if and only if $\text{rk}[\bm\Lambda_n^\prime\; (\bm\Lambda_n\mathbf A)^\prime\cdots (\bm\Lambda_n\mathbf A^{r-1})^\prime ]=r$ (AM79).} This implies that, for all $n>N$, the linear system satisfies the mini-phase condition: \[ rk\left(
\right)=2r, \; for \; |z|\ge1. \] In other words, the linear system (ref)-(ref) is the minimal state-space representation of a DFM having as McMillan degree the number of factors $r$ andersondeistler08.
This result, together with Assumption (ref), guarantees that the transfer function matrix $\bm W(z)=\bm\Lambda_n(\mathbf I_r-\mathbf Az)^{-1} \mathbf H$ is identified for all $z\in\mathbb C$, with the exception of a zero-measure set (LDA, and heatonsolo04, for similar results). This implies generic identifiability of the linear system (ref)-(ref).
Finally, we consider the issue of identification of the maxima of the log-likelihood. For any given $i\in\mathbb N$, define
where $M_\lambda$ and $C_\xi$ are defined in Assumptions (ref)(a) and (ref)(a), respectively. Likewise, define
where $M_A$ and $M_v$ are defined in Assumptions (ref)(d) and (ref)(e), respectively. Moreover, for any given $n\in\mathbb N$, let also
where $\underline C_r$ and $\overline C_1$ are defined in (ref), while $M_\xi$ and $L_\xi$ are defined in Assumptions (ref)(b) and (ref)(f), respectively. Then, the search for the maximum of the expected log-likelihood in the EM algorithm and the one of the log-likelihood takes place on the set $\mathcal O_n=\{\mathcal O_{\lambda_i}^n\cap \mathcal E_{\Lambda_n}\}\times\{\mathcal O_{\sigma_i^2}^n\cap\mathcal E_{\Gamma^\xi_n}\} \times \mathcal O_{\mathcal A}\times \mathcal O_{\Gamma^v}$, which has dimension $Q_n=n(r+1)+r^2p_F+r(r+1)/2$ growing with $n$.
Because of Assumptions (ref)(a) and (ref)(a), the rows of the loadings matrix $\bm\lambda_i$ and the idiosyncratic variances $\sigma_i^2$ belong to $\mathcal O_{\lambda_i}\times\mathcal O_{\sigma_i^2} \subset \mathbb R^{r+1}$, which is a compact set for any given $i\in\mathbb N$. Similarly, because of Assumptions (ref)(d) and (ref)(e), the entries of $\mathbf A_k$, $k=1,\ldots, p_F$, and $\mathbf H$ belong to $\mathcal O_{\mathcal A}\times \mathcal O_{\Gamma^v}\subset \mathbb R^{r^2p_F+r(r+1)/2}$, which is also a compact set. These properties are crucial as they ensure the existence of a maximum which is a solution of the EM algorithm. Indeed, for any iteration $k\ge 0$ the expected log-likelihood $\mathcal Q(\underline{\bm\varphi}_n,\widehat{\bm\varphi}_{n}^{(k)})$ in (ref), which we maximize in the M-step, is made of two terms: a term in (ref) associated to the VAR for the factors (ref), which is defined on the finite-dimensional set $\mathcal O_{\mathcal A}\times \mathcal O_{\Gamma^v}$, and a term (ref) associated to the factor equation (ref) and thus defined over a set $\mathcal O_{\lambda_i}^n\times\mathcal O_{\sigma_i^2}^n$ of dimension growing with $n$. Now, the former term poses no problem because we can use compactness of $\mathcal O_{\mathcal A}\times \mathcal O_{\Gamma^v}$ to show that the maxima $\widehat{\bm\theta}^{(k+1)}$ exist. Regarding the latter term, in order to prove the existence of $\widehat{\bm\phi}_n^{(k+1)}$ we can still use the same compactness argument by noticing that maximizing this term amounts to separately maximizing of $n$ terms, each depending only of $\underline{\bm\lambda}_i$ and $\underline{\sigma}_i^2$ for given $i\in\mathbb N$, and thus defined on the finite-dimensional compact set $\mathcal O_{\lambda_i}\times\mathcal O_{\sigma_i^2}$. It is also easy to show that the elements of $\widehat{\bm\varphi}_n^{(k+1)}$ are unique because the M-step gives a closed form solution for each of them. This reasoning guarantees the existence and uniqueness of the EM estimators (see Lemma (ref)).
Let us turn to the QML estimator that maximizes the full log-likelihood (ref). The existence of the QML estimator $\widehat{\bm\theta}^*$ poses no problem because being a finite-dimensional vector, we can use the compactness argument. Direct proof of the existence of the QML estimator $\widehat{\bm\phi}_n^*$ is instead more challenging, as in this case, we cannot rely on compactness because the full log-likelihood is defined on a set of increasing dimension $n$.
Nevertheless, we give an indirect proof of the existence of $\widehat{\bm\phi}_n^*$ by noticing that, under our identification Assumptions (ref)(a)-(ref)(c), as $n,T\to\infty$, the elements of $\widehat{\bm\phi}_n^*$ are asymptotically equivalent to the unfeasible Ordinary Least Squares estimators we would obtain if the factors were observed (see Lemma (ref)). Therefore, this argument ensures, at least asymptotically, both the existence and also the uniqueness of the QML estimators. We also refer to the next section for more details.
This section presents the asymptotic properties of the EM estimator of the parameters---i.e., of the factor loadings, idiosyncratic variances, VAR coefficients, and the covariance matrix of the VAR innovations---and of the Kalman smoother estimator of the factors. We assume that $r$, the number of common factors, is known; without loss of generality, we fix the VAR order in (ref) to $p_F=1$, and we let $\mathbf A\equiv \mathbf A_1$ so that $\mathbf A(L)\equiv\mathbf A L$. We briefly discuss the estimation of $r$ and $p_F$ at the end of the section.
We start by proving the consistency of the EM algorithm and the Kalman smoother under the most general case where we neither impose linearity of the conditional mean (Assumption (ref)) nor exponentially decaying tails (Assumption (ref)).
The consistency rate of the estimated parameter is the same as the PC estimator Bai03. This is because by initializing the algorithm with the consistent PC estimator, we can consider the EM estimator a “one-step” estimator (LC06). As for the the estimated factors, they converge at a slower rate than the PC estimator, due to the $\sqrt T$ term Bai03.
If we also assume that Assumptions (ref) and (ref) hold, we can refine the previous result because we can now guarantee that the EM algorithm converges to the QML estimator.
The rate of consistency of the estimated loadings, $\min(n/{\log^{2/\delta_v} T},\sqrt T)$, given in Proposition (ref), is new to the EM literature. This rate is the same, up to a logarithmic factor, as the one of the PC estimator (Bai03), which, in turn, is equivalent to the unfeasible Ordinary Least Squares (OLS) we would obtain if the factors were observed. The EM estimator is also asymptotically equivalent to the QML estimator considered by baili16 when imposing no autocorrelation for the factors. For an explanation of the logarithmic term we refer to Remark (ref) below. Efficiency is discussed in Section (ref).
It is important to stress that Proposition (ref) requires not only $T\to\infty$, as in classical QML estimation theory, but also $n\to\infty$ otherwise no consistency can be proved. In particular, as $n\to\infty$ the factors can be treated as observed, therefore, there is no more an issue of missing information and the QML estimator of the loadings must coincide with the unfeasible OLS. This is a manifestation of the blessing of dimensionality which is a fundamental feature of approximate factor models.
The proof of Proposition (ref) is based on the following decomposition of the estimation error into four terms:
which shows that the EM estimator $\widehat{\bm\lambda}_i$ is asymptotically equivalent to the OLS estimator we would obtain had we observed the factors.
To prove our result, we first show that the QML estimator of the loadings, $\widehat{\bm\lambda}_i^*$ is asymptotically equivalent to the PC estimator, which, in turn, is asymptotically equivalent to the unfeasible $\sqrt T$-consistent OLS estimator ${\bm\lambda}_{i}^{\text{\tiny OLS}}$ (second term on the rhs of (ref) proved in Lemma (ref), see also MBPCAQML). In particular, both approximation errors are $o_p(T^{-1/2})$, whenever $n^{-1}\sqrt {T}\log^{2/\delta_v} T\to 0$. This result extends to the DFM the result by baili16 obtained for QML estimation of a static factor model i.e., when we replace the full-matrix ${\bm\Omega}_T^F(\underline{\bm{\mathcal A}}, \underline{\mathbf H})$ with just $\mathbf I_{rT}$ in the log-likelihood (ref). Notice that, differently from the proofs in baili16 and, as mentioned in Remark (ref), this result does not depend on the QML estimator $\widehat{\sigma}_i^{*2}$.
Next, we show that the EM estimator converges to a global maximum of the likelihood (third term on the rhs of (ref)). As we discussed in Section (ref), we know that the sequence of estimators $\{\widehat{\bm\lambda}_i^{(k)},\, k\ge 0\}$ converges to a local maximum of the likelihood, say $\widehat{\bm\lambda}_i^{**}$, as $k\to\infty$ (see Lemma (ref)). However, in general, the likelihood might have many maxima due to the identification indeterminacy of the loadings. Nevertheless, once we make the identifying Assumptions (ref)(b) and (ref)(c), there is only a unique maximum, $\widehat{\bm\lambda}_i^*$, which is asymptotically equivalent to the unique OLS estimator, whenever $n^{-1}\sqrt {T}\log^{2/\delta_v} T\to 0$ (see Lemma (ref) and ruud91, for a similar result in the case of one-to-one mapping from the factors to the data, corresponding to the case of no idiosyncratic component). This is also clear from the asymptotic expansions of $\widehat{\bm\lambda}_i^*-\bm\lambda_i$ obtained by baili12,baili16 for the static model (i.e., the factors have no dynamics) and using identification schemes different than the one we use here.
Third, we show that asymptotically the error coming from running the EM a finite number of times vanishes (fourth term on the rhs of (ref)). Indeed, due to the finite number of iterations, $k^*$, the EM algorithm delivers an estimator $\widehat{\bm\lambda}_i\equiv\widehat{\bm\lambda}_i^{(k^*+1)}$, which is just an approximation of the local maximum $\widehat{\bm\lambda}_i^{**}$ that we would attained after an infinite number of iterations. We show that the error entailed by such approximation depends on the ratio of the Hessians of the complete and incomplete log-likelihoods, i.e., on how much information is missing because the factors are not observed (MR94, MLT07, and sundberg2019statistical). In this case the approximation error is $o_p(T^{-1/2})$, provided $n^{-1}\sqrt {T}\log^{2/\delta_v}T\to 0$ (see Lemma (ref)). This last result is a refinement of the results by BWB17 on the convergence of the EM algorithm. We refer to Section (ref) below for more details.
Last, ${\bm\lambda}_{i}^{\text{\tiny{OLS}}}$ is a $\sqrt T$-consistent estimator of the factor loadings $\bm\lambda_{i}$ (first term on the rhs of (ref)).
The rate of consistency of the estimated factors given in Proposition (ref) is faster than the rate originally derived by DGRqml for the same estimator---$\min(\sqrt n,T/\sqrt{\log n})$ vs. $\min(\sqrt n,T^{1/4}/\sqrt{\log n})$. Moreover, this consistency rate is the same (up to a logarithmic factor) as that of the PC estimator (Bai03). However, while the PC estimator is equivalent to the unfeasible Ordinary Least Squares (OLS) we would obtain if the loadings were observed, the Kalman smoother estimator we are considering is equivalent to the unfeasible Weighted Least Squares (WLS) we would obtain if the loadings were observed and we knew the idiosyncratic variances. As such, the Kalman smoother is also equivalent to the feasible WLS studied by baili16 and computed using the QML estimator of the loadings for a static factor model. For an explanation of the logarithmic term we refer to Remark (ref) below. Efficiency is discussed in Section (ref).
The proof of Proposition (ref) is based on the following decomposition of the estimation error into four terms
which shows that the Kalman smoother estimator $\widehat{\mathbf F}_t$ is asymptotically equivalent to the WLS estimator we would obtain had we observed the loadings and had we known the idiosyncratic variances.
To prove our result, we first show that our estimator of the factors, which is obtained via the Kalman smoother computed using the EM estimators of the parameters, is asymptotically equivalent to the Kalman filter, $\widehat{\mathbf F}_{t|t}$, which in turn is asymptotically equivalent to the WLS estimator $\widehat{\mathbf F}_{t}^{\text{\tiny WLS}}$ (fourth and third term on the rhs of (ref), respectively). Both approximation errors are $O_p(n^{-1})$ (see Lemma (ref) and PR22, for the one-factor case).
Then, we take into account the estimation error of the parameters given in Proposition (ref), which implies that the WLS estimator $\widehat{\mathbf F}_{t}^{\text{\tiny WLS}}$ converges to its unfeasible counterpart computed using the true value of the parameters, ${\mathbf F}_{t}^{\text{\tiny WLS}}$, with a rate which is $o_p(n^{-1/2})$, provided $T^{-1}\sqrt {n\log n}\to 0$ (second term on the rhs of (ref)). This result refines the result of Proposition (ref) because, thanks to Assumption (ref), we are now able to derive a tighter bound for $\Vert \widehat{\bm\Sigma}_n^\xi-\bm\Sigma_n^\xi \Vert=\max_{i=1,\ldots, n}\vert \widehat{\sigma}_i^2-\sigma_i^2\vert$, as shown in Proposition (ref) below.
Last, ${\mathbf F}_{t}^{\text{\tiny WLS}}$ is a $\sqrt n$-consistent estimator of the realizations of the factors $\mathbf F_t$ (first term on the rhs of (ref)).
Proposition (ref) does not require a limit for $T/n$ or $n/T$, so it holds without any constraint between the rates of divergence of $n$ and $T$. That said, Proposition (ref) has two special cases: (a) if $n/T\to 0$, then $\sqrt n(\widehat{\chi}_{it}-\chi_{it})\stackrel{d}{\to}\mathcal N(0,\mathcal C^F_{it})$; and, (b) if $T/n\to 0$, then $\sqrt T(\widehat{\chi}_{it}-\chi_{it})\stackrel{d}{\to}\mathcal N(0,\mathcal C^\lambda_{it})$. This is the same rate of consistency we obtain for the PC estimator of the common component (Bai03).
All other estimated parameters are also consistently estimated.
We conclude with a series of general remarks.
In this section, we discuss how our asymptotic results change if we initialize the EM algorithm with a generic initial estimator of the parameters, say $\check{\bm\varphi}_n^{(0)}=(\text{vec}(\check{\bm\Lambda}_n^{(0)})^\prime\; \check{\sigma}^{2(0)}_{1}\cdots \check{\sigma}^{2(0)}_{n}\; \text{vec}(\check{\mathbf A}^{(0)})^\prime\;\text{vech}(\bm\Gamma^{v(0)})^\prime)^\prime$, having still elements belonging to $\mathcal O_n$ as defined in Section (ref), and, thus, satisfying Assumptions (ref), (ref), and (ref).
For fixed $n$ and in a general setting, BWB17 prove that the EM algorithm defines a contraction path towards a local maximum of the likelihood, $\widehat{\bm\varphi}_n^{**}$. To give more details and understand the relation with our results, we need to introduce some general definitions. First, consider any initial estimator belonging to a {closed} neighborhood of the local maximum of given Euclidean radius $\varrho>0$, i.e., $\check{\bm\varphi}_n^{(0)}\in\mathcal B(\varrho; \widehat{\bm\varphi}_n^{**})\subset \mathbb R^Q$. In our setting, we can think of $\mathcal B(\varrho; \widehat{\bm\varphi}_n^{**})\equiv \mathcal O_n$. Then, define the EM operator $\bm M_T:\mathbb R^{Q}\to\mathbb R^{Q}$ such that $\bm M_T(\widehat{\bm\varphi}_n^{(k)})=\widehat{\bm\varphi}_n^{(k+1)}$. We have $\Vert \mathbb{E}[M_T(\underline{\bm\varphi}_n)]-\widehat{\bm\varphi}_n^{**}\Vert \le \beta\Vert \underline{\bm\varphi}_n-\widehat{\bm\varphi}_n^{**}\Vert $ for some $\beta\in(0,1)$ and all $\underline{\bm\varphi}_n\in \mathcal B(\varrho; \widehat{\bm\varphi}_n^{**})$ (BWB17). Second, for any given $T$ and $\delta\in(0,1)$, let $\varepsilon_{T,\delta}$ be the smallest scalar such that $\mathrm P(\sup_{\underline{\bm\varphi}_n\in \mathcal B(\varrho;\widehat{\bm\varphi}_n^{**})} \Vert M_T(\underline{\bm\varphi}_n)- \mathbb{E}[M_T(\underline{\bm\varphi}_n)]\Vert \le \varepsilon_{T,\delta})\ge 1-\delta$.
It follows that, if $T$ is large enough such that $\varepsilon_{T,\delta} \le (1-\beta)\varrho$, then, for any $k\ge 0$, the EM operator defines a contraction towards the maximum of the likelihood with high-probability (BWB17):
Now, under our mixing Assumptions (ref)(d)-(ref)(h) and (ref)(c), $\varepsilon_{T,\delta}\to 0$, as $T\to\infty$. This, jointly with (ref), has two implications, as $T\to\infty$:
These two results apply if $n$ is fixed. But when $n\to\infty$, we might conjecture that those results will hold for each component of $\bm\varphi_n$ separately. And this is, indeed, what we verify in this paper. Consider the loadings $\bm\lambda_i$. As mentioned in the previous section, as $n\to\infty$ the likelihood has a unique global maximum, $\widehat{\bm\lambda}_i^{*}$, which is consistent because it is the QML estimator. Therefore, $\Vert \widehat{\bm\lambda}_i^{**}-{\bm\lambda}_i\Vert\le \Vert \widehat{\bm\lambda}_i^{**}-\widehat{\bm\lambda}_i^*\Vert+\Vert \widehat{\bm\lambda}_i^{*}-{\bm\lambda}_i\Vert=o_p(1)$, as $n,T\to\infty$. If we initialize the EM algorithm with the consistent PC estimator $\widehat{\bm\lambda}_i^{(0)}$, we prove that, as $n,T\to\infty$, result (a) above still applies, i.e., there exists a $\beta_\lambda\in(0,1)$ such that, for all $k\ge 0$,
Actually, we are also able to show that the contraction factor is such that, as $n,T\to\infty$:
with $\mathbf F_{t|T}^{*}$ and $\mathbf P^{*}_{t|T}$ being the factors and their associated MSEs obtained from the Kalman smoother when using the QML estimator of the parameters---we refer to Lemma (ref) for a proof of (ref) and (ref). It follows that the convergence rate in (ref) is faster than the one of the initial PC estimator, and we can treat the EM estimator as if it were the QML estimator. Moreover, (ref) means that (ref) holds for any initial estimator (see Lemma (ref)). The intuition for this result is that, provided $n^{-1}\check{\bm\Lambda}^{(0)\prime}\check{\bm\Lambda}^{(0)}$ has full-rank, any cross-sectional averaging is likely to recover in high-dimensions a factors' space not too different from the true one, an argument often used when considering estimation methods based on aggregations schemes alternative to PCs WU15,fan2022learning.
For all other parameters, a relation like (ref) holds as well, with possibly different contraction factors. Hence, result (a) still applies (see Lemma (ref)). However, we could not derive a result like (ref) in this case because computing analytic expressions of all those contraction factors is difficult. Thus, when we initialize the algorithm with any generic initial estimator, we have to rely on result (b); that is, we can guarantee convergence of the EM algorithm to the QML estimator, as $n,T\to\infty$, only if we let the algorithm run for a number of iterations $k\ge k_T$, where the larger $k$ is, the faster is the convergence rate.
In this section, we compare the asymptotic covariances of the EM and Kalman smoother with those of the PC estimators, which are the optimal non-parametric estimators.
From Propositions (ref)(b) and (ref)(b), we see that consistency is not affected by estimating a mis-specified model with uncorrelated idiosyncratic components, but there is an efficiency loss due to this mis-specification, as shown by the sandwich forms of the asymptotic covariance matrices. In the case of uncorrelated idiosyncratic components, the log-likelihood we maximize is correctly specified. If the model is correctly specified, the EM estimator is the most efficient one because it is asymptotically equivalent to the Maximum Likelihood estimator. Thus, its asymptotic covariance attains the classical lower bound of the OLS estimator (see Propositions (ref)(c)). Likewise, the asymptotic covariance of the factors attains the WLS lower bound (see Propositions (ref)(c)).
Although, in general, the EM algorithm and the Kalman smoother do not provide the most efficient estimators, they can provide advantages with respect to the PC estimators.
Part (a) follows immediately once we impose the identifying Assumption (ref) to the results about PC estimation of the loadings (see MBPCAQML, and Bai03). Therefore, although we estimate a mis-specified model, the EM estimator of the loadings is as efficient as the PC estimator.
Turning to part (b), by imposing the identifying Assumption (ref), the asymptotic covariance of the PC estimator of the factors is $ \bm{\mathcal W}^{\text{\tiny PC}}_t=(\bm\Sigma_\Lambda)^{-1}\{\lim_{n\to\infty} n^{-1} \bm\Lambda_n^\prime \bm\Gamma_n^\xi \bm\Lambda_n\}(\bm\Sigma_\Lambda)^{-1} $ (see Lemma (ref), and Bai03), and if the true model were an exact factor model, i.e., $\mathbb{E}[\xi_{it}\xi_{jt}]=0$ if $i\ne j$ so that $\bm\Gamma_n^\xi=\bm\Sigma_n^\xi$, then we would have $ \bm{\mathcal W}^{\text{\tiny PC}}_t=(\bm\Sigma_\Lambda)^{-1}\{\lim_{n\to\infty} n^{-1} \bm\Lambda_n^\prime \bm\Sigma_n^\xi \bm\Lambda_n\}(\bm\Sigma_\Lambda)^{-1} $, which is the unfeasible OLS asymptotic covariance matrix in presence of heteroskedastic errors. For the Kalman smoother, from Proposition (ref) we have $ \bm{\mathcal W}_t =(\bm\Sigma_{\Lambda\Sigma\Lambda})^{-1} \{\lim_{n\to\infty} n^{-1} \bm\Lambda_n^\prime(\bm\Sigma_n^{\xi})^{-1} \bm\Gamma_n^\xi (\bm\Sigma_n^{\xi})^{-1}\bm\Lambda_n\} (\bm\Sigma_{\Lambda\Sigma\Lambda})^{-1} $, and if the true model were an exact factor model, this would reduce to $ \bm{\mathcal W}_t=(\bm\Sigma_{\Lambda\Sigma\Lambda})^{-1} $, which is the unfeasible WLS asymptotic covariance matrix. In this case, the Kalman smoother estimator is more efficient because it takes into account heteroskedasticity, whereas the PC estimator ignores the possibility of individual-specific variances.
Here, we go one step further and show that if the idiosyncratic covariance matrix ${\bm\Gamma}_n^\xi$ is sparse enough, we can still have efficiency gains compared to PC analysis. Indeed, if the total contribution of the off-diagonal elements ${\bm\Gamma}_n^\xi$ is negligible compared to $n$, we can expect the asymptotic covariance of the estimated factors to be close to the one we would have for an exact factor model, which we know is smaller than the PC asymptotic covariance. The sparsity condition we assume is the same as the one on bailiao16. Although this sparsity condition is hard to verify in practice and is seldom exactly satisfied by the data, in our MonteCarlo results in Section (ref), we show that the Kalman smoother estimator tends to perform better than the PC estimator even under more general idiosyncratic covariance structures like banded matrices.
To conduct inference, we need asymptotic covariances of the loadings and factors matrices and their estimators.
For any given $i,j=1,\ldots,n$, the most general estimator of the asymptotic covariance between $\widehat{\bm\lambda}_i$ and $\widehat{\bm\lambda}_j$ is given by:\footnote{As discussed in Remark (ref), the sample covariance of $\widehat{\mathbf F}_t$ is equal to the $r$-dimensional identity matrix only asymptotically; hence, it must be included in the estimator of the asymptotic covariance.}
where $\widehat{\xi}_{it}=x_{it}-\widehat{\chi}_{it}$ is the estimated idiosyncratic component of the $i$th variable at time $t$. $\mathrm K(t,s) = 1-\frac{|t-s|}{M_T+1}$, if $|t-s|\le M_T$ and zero otherwise, with $M_T$ is such that $M_T\to\infty$ and $M_T/T\to 0$, as $T\to\infty$. Consistency of (ref), as $n,T\to\infty$, follows from Bai03 combined with Propositions (ref) and (ref).
For any given $t=1,\ldots, T$ and $k=0,\ldots, M_T$, with $M_T$ defined above, the most general estimator of the asymptotic covariance between $\widehat{\mathbf F}_t$ and $\widehat{\mathbf F}_{t-k}$ is given by:
where $\widehat{\gamma}^\xi_{ij, k} = T^{-1}\sum_{t=k+1}^T \widehat{\xi}_{it}\widehat{\xi}_{j,t-k}$ and $\mathrm K(i,j)=1$ if $1\le i,j\le m_{n,T}$ and zero otherwise, with $m_{n,T}\to\infty$ and $m_{n,T}/\min(n,T)\to 0$, as $n,T\to\infty$. Consistency of (ref), as $n,T\to\infty$, follows from baing06 combined with Propositions (ref), (ref), and (ref). For larger values of $k$, $\mathbb{E}[\xi_{it}\xi_{j,t-k}]$ is likely to be small due to our Assumption (ref)(b), and, thus, we can consider $\widehat{\mathbf F}_t$ and $\widehat{\mathbf F}_{t-k}$ as asymptotically uncorrelated. An alternative kernel function of the correlation based distance between $i$ and $j$ is considered by kim2022robust who considers also averages of (ref) when computed choosing different random permutations of the selected $m_{n,T}$ cross-sectional units. Alternatively, rather that smoothing via the use of a kernel, BR24 consider an estimator based on thresholding of the idiosyncratic sample covariance matrix (see also FLM13).
Finally, we can obtain an estimator of $\bm{\mathcal W}_t$ that accounts also for the autocorrelation of the factors from the true MSE of the Kalman filter given in (ref) when using the estimated parameters. Once again this requires an estimator of the idiosyncratic covariance matrix like the estimators discussed above. The asymptotic covariance between $\widehat{\mathbf F}_t$ and $\widehat{\mathbf F}_{t-k}$ can be obtained by considering the Kalman filter with the augmented state vector $(\mathbf F_t^\prime\cdots \mathbf F_{t-k}^\prime)^\prime$ and by then considering the corresponding $r\times r$ off-diagonal block of the resulting $r(k+1)\times r(k+1)$ MSE.
Having the estimated loadings $\widehat{\bm\lambda}_i$, the estimated factors $\widehat{\mathbf F}_t$, and any of the above estimators of $\bm{\mathcal V}_{i,j}$ and $\bm{\mathcal W}_{t,t-k}$, we can estimate the variance of the estimated common component by plugging these quantities into the expression in Proposition (ref)(b).
To conclude, the covariance estimators defined in this section can be used, together with the asymptotic distributions derived in Propositions (ref) and (ref), for inferential purposes. Examples are in Section (ref).
Throughout, we consider a model with $r=4$ factors, and we simulate the data according to
where ${\bm\ell}_{i}$ has entries ${\ell}_{ij}\stackrel{iid}{\sim}\mathcal{N}(1,1)$, $i=1,\ldots, n$, $j=1,\ldots, r$; ${\bm A}=\mu\check{\bm A} \{\nu^{(1)}({\check{\bm A}})\}^{-1}$, where $[\check{\bm A}]_{jj}\sim U[0.5,0.8]$, $[\check{\bm A}]_{jk}\sim U[0,0.3]$, $j,k=1,\ldots,r$, and $\mu=0.7$; $u_{jt}\stackrel{iid}{\sim}(0,1)$, $j=1,\ldots, r$ following either a Gaussian, an Asymmetric Laplace, or a Skew-$t$ distribution, and with $\mathbb{C}\mathrm{ov}(u_{it},u_{jt})=0$, for $i\ne j$; $\alpha_i=\{0,\delta_i\}$, $i=1,\ldots, n$, where $\delta_i\stackrel{iid}{\sim}\mathcal{U}(0,\delta)$, and $\delta\in\{0,0.5\}$; $e_{it}\stackrel{iid}{\sim}(0,\sigma_{ei}^2)$, $i=1,\ldots, n$, following either a Gaussian distribution with $\sigma_{ei}^2\sim U[0.5, 1.5]$, or an Asymmetric Laplace distribution with $\sigma_{ei}^2=1$, or a Skew-$t$ distribution with $\sigma_{ei}^2=1$; $\mathbb{C}\mathrm{ov}(e_{it},e_{jt})=\tau^{\vert i-j\vert }$, $i,j=1,\ldots, n$, with $\tau\in\{0,0.5\}$ if $\vert i-j\vert \le 10$, and $\mathbb{C}\mathrm{ov}(e_{it},e_{jt})=0$ otherwise; and, last, $\phi_i=\sqrt{\theta_i{\widehat{\mathbb{V}\mathrm{ar}}(\chi_{it})}(\widehat{\mathbb{V}\mathrm{ar}}(\xi_{it}))^{-1}}$, $i=1,\ldots, n$, where $\widehat{\mathbb{V}\mathrm{ar}}(\cdot)$ denotes the sample variance, and $\theta_i\stackrel{iid}{\sim}\mathcal{U}(\bar{\theta}-0.25,\bar{\theta})$, and $\bar{\theta}=0.5$. The parameters $\mu$, $\tau$, $\delta$, and $\bar\theta$ are crucial and control: the persistence of the factors, the degrees of cross-sectional and serial idiosyncratic correlation in the idiosyncratic components, and the noise-to-signal ratio, respectively.\footnote{In the case of the Asymmetric Laplace distribution, all the innovations have location 0, scale index $\lambda$, and asymmetry index $\kappa$, with $\kappa \sim \mathcal{U}(.9,1.1)$ and $\lambda=\sqrt{(1+\kappa^4)\kappa^{-2}}$, so that all the shocks have variance 1. In the case of the Skew-$t$ distribution, all the shocks have location 0, dispersion 1, skewness index $\gamma$, and tail index $\nu$, with $\nu_u\sim {U}(4,12)$, $\gamma_u \sim {U}(-.1,.1)$, $\nu_e\sim {U}(3,13)$, and $\gamma_e \sim {U}(-.15,.15)$.}
Finally, in order to satisfy Assumptions (ref)(b) and (ref)(c), after we have generated the common component $\bm\chi_{nt}$ as in (ref), we construct the factors as $\mathbf F_t=(\mathbf M_n^{\chi})^{-1/2}\mathbf V_n^{\chi\prime} \bm\chi_{nt}$ and the loadings as $\bm \Lambda_n=\mathbf V_n^{\chi} (\mathbf M_n^{\chi})^{1/2}$, where $\mathbf M_n^\chi$ is the $r\times r$ diagonal matrix containing the eigenvalues of $T^{-1}\sum_{t=1}^T \bm\chi_{nt}\bm\chi_{nt}^\prime$, and $\mathbf V_n^\chi$ is the $n\times r$ matrix having as columns the corresponding normalized eigenvectors and with sign fixed such that it has non-negative entries in the first row.
We consider $n\in\{100,200,300, 500\}$, $T\in\{100,200,300, 500\}$ and $B=1000$ replications. At each replication $b$, we run the EM algorithm as described in Algorithm (ref), thus obtaining an estimate of the loadings $\widehat{\bm\lambda}_i^{(b)}$, the factors $\widehat{\mathbf F}_t^{(b)}$, and the common component $\widehat{\chi}_{it}^{(b)}=\widehat{\bm\lambda}_i^{(b)\prime}\widehat{\mathbf F}_t^{(b)}$. We initialize the EM algorithm either through the PC estimators as explained in Appendix (ref), or through a contaminated version of the PC estimator obtained using contaminated eigenvectors of the data. In particular, let $\widehat{\mathbf V}_n^x$ be the $n\times r$ matrix of the $r$ leading eigenvectors of $T^{-1}\sum_{t=1}^T \mathbf x_{nt}\mathbf x_{nt}^\prime$, then the contaminated eigenvector is $\widehat{\mathbf V}_n^{x,c}=\widehat{\mathbf V}_n^x+\bm {\mathcal Z}\bm \Upsilon^{1/2}$, where $[\bm {\mathcal Z}]_{ij}\stackrel{iid}{\sim} \mathcal N(0,1)$, $i=1,\ldots, n$, $j=1,\ldots,r$, and $[\bm \Upsilon]_{ij}=\iota^{\vert i-j\vert}$, $i,j=1,\ldots, r$, with $\iota \in (0,1)$---the bigger $\iota$ is, the stronger is the contamination.
The upper plots of Figure (ref) show the log-likelihood $\ell(\bm X_{nT};\widehat{\bm{\varphi}}_n^{(k)})$ (blue line, left scale) as a function of the iteration $k$ of the EM algorithm, and the convergence criterion $\Delta\ell_k$ defined in (ref) (red line, right scale) when we initialize the EM algorithm with the PC estimator. The results in Figure (ref) show that the log-likelihood is an increasing function in the number of iterations $k$ (as it should be). Moreover, the algorithm converges very fast: within two iterations $\Delta\ell_k\le 10^{-3}$. This is to be expected since we initialize the algorithm with the consistent PC estimator.
The lower plots of Figure (ref) show the percentage deviation of the log-likelihood of the algorithm initialized with the contaminated estimator ($\ell^c_k$) from the one initialized with the PC estimator ($\ell_k$), that is $\ell^c_k/\ell_k-1$. Contaminating the initialization implies starting from a much lower likelihood. However, with just a few iterations, the log-likelihood initialized with the contaminated estimator is the same as the one properly initialized. This result shows that initializing the EM with a non-consistent estimator is fine, as it just requires running the algorithm for a few more iterations. Therefore, hereafter, we focus on the case in which we inizialize the EM algorithm with the PC estimator.
Moving to the performance of our estimators, we begin with Proposition (ref) part (a), which gives consistency and rates for the common component's estimator. The left plot in Figure (ref) shows the root mean squared error:
Instead, the right plot shows the relative RMSE of our estimator over the RMSE of the PC estimator---values smaller than one indicate a better performance of our estimators. We show results for several DGPs, which allow us to disentangle the effects of each single mis-specification on the performance of our estimator.
Two main results emerge from Figure (ref). First, as $n$ and $T$ grow, the RMSE of all DGPs converge toward zero, thus indicating that the mis-specification introduced by estimating a model with uncorrelated and possibly non-Gaussian idiosyncratic components, even when that is not the case, does not affect our estimator. In particular, between serial correlation (the orange line) and cross-sectional correlation (the yellow line), the mis-specification that hurts the most is cross-sectional correlation. This is good news for practitioners because the cross-sectional correlation between the idiosyncratic components can somehow be limited by avoiding to include in the dataset variables that are too similar with one another (see, e.g., the discussions in BN06 and lucioADFM). Moreover, the model is consistently estimated when the shocks are asymmetric and have heavy tails even when they come from a distribution that does not meet Assumption (ref) of exponentially decaying tails of the distribution.\footnote{The dark gray line for the Skew-t distribution cannot be seen in the left plot because it is underneath, thus it coincides with, the red line.} This is also good news for practitioners because this result tells us that we can use this model in settings likely to be non-Gaussian, as is the case in datasets of disaggregated inflation rates (see, e.g., reiswatson10, and AL), which are notoriously skewed and fat-tailed.
Second, overall, our estimator behaves very similarly to but slightly better than the PC estimator, despite the latter being non-parametric and, thus, not affected by mis-specifications. This result confirms the conjectures based on extensive numerical studies made by DGRqml and baili16, showing that, in a high-dimensional setting, the EM estimator is “as good as” the PC estimator.
To evaluate the estimates of the factors and the loadings, at each replication $b$, we consider a multivariate version of the $R^2$ (see also DGRqml):
where $\widehat{\bm{ \mathcal F}}_T^{(b)}$ and $\widehat{\bm \Lambda}_n^{(b)}$ are the $T\times r$ and $n\times r$ matrices of the estimated factors and loadings, respectively. These trace statistics are smaller than one, and they tend to one when the empirical canonical correlations between the true quantities and their estimates tend to one.
Figure (ref) reports the values of $\mathrm{TR}_F^{(b)}$ and $\mathrm{TR}_\Lambda^{(b)}$, relative to the same measures computed for the PC estimator, averaged over all $B$ replications (values larger than one indicate a better performance of our estimators). The results in Figure (ref) mirror those of Figure (ref): our estimator behaves very similarly to but slightly better than the PC estimator.
To derive our asymptotic results, we assumed that the true factors and the loadings are asymptotically identified as in Assumption (ref)(b). As explained in Remark (ref) the estimated factors and loading satisfy this assumption asymptotically. To verify that this is the case, in Figure (ref), we show the two identification errors:
where $\widehat{\bm \Gamma}^{F{(b)}}=T^{-1}\widehat{\bm {\mathcal F}}_T^{(b)\prime} \widehat{\bm {\mathcal F}}_T^{(b)}$, and $\widehat{\bm \Sigma}_\Lambda^{(b)}=n^{-1}\widehat{\bm \Lambda}_n^{(b)\prime} \widehat{\bm \Lambda}_n^{(b)}$. According to our asymptotic results, both these quantities should tend to zero as $n$ and $T$ increase because, given our simulation design, the true factors are orthonormal and the true loadings are orthogonal in agreement with Assumption (ref)(b). Figure (ref) shows $I_F^{(b)}$ and $I_\Lambda^{(b)}$ averaged over all $B$ replications. The identifying constraints are more and more precisely satisfied as $n$ and $T$ grow.
Next, we move to the asymptotic distribution of the common component. To this end, for each replication $b$ and any $i,t$, we compute
where we use the robust estimators of the asymptotic covariance matrices defined in (ref) and (ref), respectively. For comparison we also consider $Z_{it}^{(0,b)}$ defined as in (ref), but when using estimators of the asymptotic covariance matrices in the case of non-correlated idiosyncratic components (henceforth, non-robust covariance matrices).
According to Proposition (ref), $Z_{it}^{(b)}\stackrel{d}{\to} \mathcal N(0,1)$ as $n,T\to\infty$. To evaluate the goodness of our theoretical results, we compute the average coverage \[ C(1-\alpha)=\frac 1{nTB}\sum_{i=1}^n\sum_{t=1}^T\sum_{b=1}^B\, \mathbb I\left(\mathcal Z_{\alpha/2} \le Z_{it}^{(b)}\le \mathcal Z_{1-\alpha/2}\right), \] where $\mathcal Z_\alpha$ is the $\alpha$-quantile of the standard normal distribution. In Table (ref), we report the coverage $C(1-\alpha)$, for selected values of $\alpha\in(0,1)$, while for illustration purposes, in Figure (ref), we show histograms of $\{Z_{it}^{(b)}: i=1,\ldots, n,\, t=1,\ldots, T, \, b=1,\ldots, B\}$, for some of the cases considered in Table (ref). We stress that throughout this exercise the chosen bandwidths for all robust covariance estimators are not data driven, but rather fixed a priori.\footnote{To estimate $\widehat{\bm{\mathcal V}}_i^{\text{\tiny(HAC)}}$ we set $M_T=\lfloor T^{1/4}\rfloor$ and to estimate $\widehat{\bm{\mathcal W}}^{\text{\tiny(HAC)}}_t$ we set $m=\lfloor n^{4/5}\rfloor$.}
Results confirm the derived asymptotic distribution. When the idiosyncratic components are uncorrelated, the non-robust estimators of the covariance matrices offers almost perfect coverage, while the robust estimators give a slight over-coverage. In the relevant cases of serially and cross-correlated idiosyncratic components, the considered robust estimators work very well. For comparison we show also results for $Z_{it}^{(b)}$ when $\chi_{it}$ is estimated with the PC estimator, and the asymptotic covariances are estimated using the estimators in baing06. In this case, deviations from Gaussianity seem to lead to serious under-coverage.
Last, we move to Proposition (ref), which states that under a sparsity condition on the covariance matrix of the idiosyncratic components, the Kalman smoother estimator of the factors is more efficient than the PC estimator. To verify this result, we compare the theoretical asymptotic covariance of the factors, $\bm{\mathcal W}_t$ and $\bm{\mathcal W}^{\text{\tiny\upshape PC}}_t$. Specifically, for each simulation, we look at the sign of the smallest eigenvalue of the matrix $(\bm{\mathcal W}^{\text{\tiny\upshape PC}}_t-\bm{\mathcal W}_t)$, computed using the true simulated parameters, which should always be positive.
Table (ref) reports the percentage of times out of 5000 simulations in which $(\bm{\mathcal W}^{\text{\tiny\upshape PC}}_t-\bm{\mathcal W}_t)$ is positive definite. The DGP used for this exercise is the one described at the beginning of this section, but for the case indicated as $\tau=0.5^\ast$ in which we set $\mathbb{C}\mathrm{ov}(e_{it},e_{jt})=0.5^{\vert i-j\vert}$ if $i,j\le \lfloor \sqrt{n}\rfloor$, and $\mathbb{C}\mathrm{ov}(e_{it},e_{jt})=0$, otherwise, in order to better proxy the assumed sparsity condition. The results in Table (ref) confirm those in Proposition (ref). When the sparsity condition is verified (this is the case of columns $\tau=0$ and $\tau=0.5^\ast$), the Kalman smoother is more efficient than the PC estimator. That said, even when the sparsity condition is not verified (column $\tau=0.5$), the Kalman smoother tends to be more efficient than the PC estimator. We also note that the two largest eigenvalues (not shown) are always positive, meaning that, in our simulations, the PC estimator is never more efficient than the EM estimator.
In this section, we consider the dataset used by OGAP, a typical panel of $n=103$ US macroeconomic quarterly indicators observed from 1960:Q1 to 2018:Q4, thus $T = 236$. All variables are transformed to stationarity. In particular, we follow the common approach of taking first differences of price indexes and keeping interest rates in levels (see, e.g., BBE05). The information criterion by baing02 indicates $\widehat r=6$ common factors. We estimate that the order of the VAR for the factors is $\widehat p_F=2$. As shown in Figure (ref), the EM algorithm converges in 22 iterations when setting the threshold in (ref) to $\varepsilon=10^{-4}$ and the log-likelihood increases monotonically.
As we said in the Introduction, there is extensive literature showing the effectiveness of the EM algorithm in estimating large DFMs. Therefore, the purpose of this section is not to show that this method works or that it is superior to the PC estimator. Rather, we concentrate on the innovations brought about in this paper that have an impact on empirical applications, namely: the confidence bands for the common components and the factors, and a test of hypothesis on the factor loadings.
Throughout, we compute the asymptotic covariances by using $\widehat{\bm{\mathcal V}}_{i}^{\text{\tiny HAC}}$ for the loadings, as given in (ref), with bandwidth $M_T=\lceil T^{1/4}\rceil$, and $\widehat{\bm{\mathcal W}}_t^{\text{\tiny KF}}=n\widehat{\bm \Pi}_{t|t}$ for the factors, computed as in (ref) by using the estimated parameters and applying local thresholding to the sample idiosyncratic covariance matrix FLM13,BR24.
The first row of Figure (ref) shows the common components of a few variables of interest estimated with the EM algorithm (the red line) with their 95% confidence bands, together with the observed data (the black line). Specifically, for any given $i=1,\ldots, n$ and all $t=1,\ldots, T$, a $(1-\alpha)\%$ confidence interval for the estimated common component is given by
Moreover, in order to have confidence bands valid for all $T$ observations we apply a Bonferroni correction to the critical values. For comparison, in the second row of Figure (ref), we report the estimated common components obtained with the PC estimator (the blue line) with their 95% confidence bands computed using the HAC estimators in baing06.
The common components of core CPI inflation and the Fed funds rate estimated with the EM track the observed series better than those estimated with the PC estimator. This result is possibly due to local departures from stationarity in those series, which create problems for PC analysis but not for the EM, as the Kalman smoother is able to track changes in the dynamics due to its recursive character. In particular, the confidence band of the common component of the Fed funds rate for the PC estimator is much wider than that of the EM estimator because the variance of the idiosyncratic component obtained with the PC estimator is nearly eight times larger than that of the EM estimator and it is also far more persistent.
In addition to computing confidence bands for the in-sample estimate of the common component, we can also do so for the unconditional and conditional forecasts obtained from the model. In other words, our results open the possibility of computing the uncertainty around GDP nowcast, which is nothing else than a short-term conditional forecast, and for scenario analysis performed with DFMs. Here we give a couple of simplified examples.
The left chart of Figure (ref) shows the 1-step ahead forecast of 2018:Q1. That is, we estimate the model up to 2017:Q4, and then we produce four different forecasts of GDP growth conditioning on the observations of an increasing number of variables for 2018:Q1. The conditional forecasts are obtained using the Kalman filter estimator of the factors, $\mathbf F_{t|t}^{(k^*+1)}$, while the unconditional forecast is obtained using their one-step-ahead prediction, $\mathbf F_{t|t-1}^{(k^*+1)}=\widehat{\mathbf A}\mathbf F_{t-1|t-1}^{(k^*+1)}$. This exercise mimics a simplified nowcasting setting---we are omitting the aspect of mixed frequency---as the variables we are conditioning on are published earlier than GDP, and the sequence of conditioning mimics the calendar of data releases. As shown in the left chart of Figure (ref), the model adjusts the forecast in the right direction as more hard data becomes available.
In the second exercise, we produce forecasts for GDP growth for each quarter of 2018 (blue line). Then, we adjust them based on different scenarios for payroll employment. We have two scenarios: one where employment grows at the same pace as the previous year (red line)---this is a slightly lower average pace than the model expected; another where employment grows at half the previous year's pace, resulting in a more pessimistic forecast (green line). As shown in the right chart of Figure (ref), one and two quarters ahead, the forecasts from these scenarios differ significantly.
Finally, we can test for linear restrictions on the loadings. Consider testing for $s$ linear restrictions $\text H_0:\; \bm R^\prime\,\text{\upshape vec}({\bm\Lambda}_{n})=\bm q$, against the alternative $\text H_1:\;\bm R^\prime\,\text{\upshape vec}({\bm\Lambda}_{n})\ne \bm q$, where $\bm R$ is $nr\times s$ and $\bm q$ is $s\times 1$. Then, the usual Wald-type test statistic is computed as
where $\bm R$ selects only $sr$ rows of $\widehat{\bm{\mathcal V}}_n^{\text{\tiny HAC}}$. Under $\text H_0$, from Proposition (ref) it follows that ${\mathsf W}_{\widehat{\bm\Lambda}_n}\stackrel{d}{\to}\chi^2_{(r)}$, as $n,T\to\infty$. This test is the analogous of the test derived in the case of PC estimation by baing13. Testing for equal loadings is equivalent to testing for equal common components for all $t=,1\ldots, T$.
Table (ref) shows the result of the test of five different null hypotheses. The first hypothesis we test (column (A)) is whether GDP and GDI, both measures of US aggregate output, have equal loadings and, as a consequence, have the same common component. Our test does not reject the hypothesis of equal loadings. This result supports the idea recently explored in the literature that combining GDP and GDI can better estimate aggregate output ADNSS,BLgdo.
The second hypothesis (column (B)) that we test is whether CPI inflation and PCE price inflation, which are two alternative measures of consumer price inflation, have equal loadings. These indexes usually differ because they are constructed differently.\footnote{The CPI, which captures the headlines in newspapers, determines the return on Treasury Inflation-Protected Securities, or TIPS, while the inflation objective of the Federal Reserve is specified in terms of PCE price inflation.} The test does not reject the null whenever $r<6$, which we read as signaling that, indeed, CPI inflation and PCE price inflation respond in the same way to the common factors, and thus, their difference is just idiosyncratic or, perhaps, weak/local factors. Columns (C), (D), and (E) in Table 5 test whether the difference between the PCE and CPI core sub-index, energy sub-index, and food sub-index are just idiosyncratic. The test suggests this to be the case.
Finally, in column (F) of Table (ref), we verify that if we test a non-sense hypothesis, the test rejects it. Specifically, we test whether the loadings of GDP growth and the growth in total non-farm employment are the same, which should not be the case. The test unequivocally reaches the same conclusion.
This paper provides the asymptotic properties of Quasi Maximum Likelihood (QML) estimation for large approximate dynamic factor models, implemented via the Kalman smoother and the Expectation Maximization (EM) algorithm. Our results provide the statistical foundations of one of the most popular and successful methods for estimating high-dimensional factor models for time series commonly used in many public and private institutions to track and predict economic activity.
From a technical point of view, we show and prove that the EM approach is feasible even in the high-dimensional case, i.e., when the cross-sectional size $n$ can be much larger than the sample size $T$, a point also made by DGRqml. Then, we show that the EM estimator converges at the same rate as the Principal Components (PC) estimator. Moreover, we show that the EM estimator of the loadings is always as efficient as the PC estimator, while the Kalman smoother estimator of the factors is more efficient than the PC estimator if the idiosyncratic covariance is sparse enough.
Compared to the standard PC estimator, the EM approach has the main advantage of allowing the user to easily impose restrictions that reflect any prior knowledge about the data on the model. The user can impose these restrictions because the state-space formulation and the Kalman Smoother allow explicit modeling and estimation of the dynamic evolution of the latent factors and deal with data irregularly spaced in time. Moreover, in the M-step, the user can impose restrictions on the parameters, thus allowing for constrained QML estimation. In contrast, the user cannot use the PC estimator to model the latent factors's dynamics. In an application on a dataset of US macroeconomic time series, we show that the EM algorithm can produce estimates of the common component that track the dynamics of the observed series better than the PC estimator, especially those series displaying periods of high persistence and regime changes, like inflation and interest rates. This result suggests that the Kalman smoother might be more robust to local deviations from stationarity, a feature already highlighted by kalman60 and kalmanbucy61.
\pagestyle{fancy} \fancyhf \chead{Supplementary material for the paper: QMLE of Large Approximate DFM via the EM algorithm} \cfoot{Page \thepage}
\setcounter{table}{0} \setcounter{figure}{0} \setcounter{page}{1}
\footnotetext{ M. Barigozzi gratefully acknowledges financial support from MIUR (PRIN2020, Grant 2020N9YFFE).\\[.03cm]
Disclaimer: the views expressed in this paper are those of the authors and do not necessarily reflect the views and policies of the Board of Governors or the Federal Reserve System.}
\gdef {(\roman{footnote})}
\contentsline {section}{\numberline {A}Further details on estimation}{2}{appendix.A} \contentsline {subsection}{\numberline {A.1}Principal Component estimators}{2}{subsection.A.1} \contentsline {subsection}{\numberline {A.2}Kalman filter and smoother}{3}{subsection.A.2} \contentsline {subsection}{\numberline {A.3}Stopping rule for the EM algorithm}{4}{subsection.A.3} \contentsline {section}{\numberline {B}Proof of main results}{4}{appendix.B} \contentsline {subsection}{\numberline {B.1}Proof of Proposition (ref)}{4}{subsection.B.1} \contentsline {subsection}{\numberline {B.2}Proof of Proposition (ref)}{8}{subsection.B.2} \contentsline {subsection}{\numberline {B.3}Proof of Proposition (ref)}{9}{subsection.B.3} \contentsline {subsection}{\numberline {B.4}Proof of Proposition (ref)}{14}{subsection.B.4} \contentsline {subsection}{\numberline {B.5}Proof of Proposition (ref)}{15}{subsection.B.5} \contentsline {subsection}{\numberline {B.6}Proof of Corollary (ref)}{18}{subsection.B.6} \contentsline {subsection}{\numberline {B.7}Proof of Proposition (ref)}{19}{subsection.B.7} \contentsline {section}{\numberline {C}General lemmas}{22}{appendix.C} \contentsline {section}{\numberline {D}Lemmas necessary for proving Proposition (ref)}{31}{appendix.D} \contentsline {section}{\numberline {E}Lemmas necessary for proving Proposition (ref)}{47}{appendix.E} \contentsline {section}{\numberline {F}Lemmas necessary for proving Proposition (ref)}{78}{appendix.F} \contentsline {section}{\numberline {G}Lemmas necessary for proving Proposition (ref)}{80}{appendix.G} \contentsline {section}{\numberline {H}Lemmas necessary for proving Proposition (ref)}{84}{appendix.H} \contentsline {section}{\numberline {I}Derivation of the Kalman filter MSE}{85}{appendix.I}
An analogous notation is used for the sub-vectors of parameters ${\bm\phi}_n$ and ${\bm\theta}$ and all their elements.
Hereafter, we assume without loss of generality that $\bm \mu_n=\mathbf 0_n$ and $p_F=1$ with $\mathbf A\equiv\mathbf A_1$. When $p_F>1$, it is enough to write the VAR in (ref) in companion form and to modify the estimation accordingly, using the augmented state vector $(\mathbf F_t^\prime\cdots \mathbf F_{t-p_F+1}^\prime)^\prime$.
Let $\widehat{\bm\Gamma}_n^x$ be the sample covariance matrix of the data and denote as $\widehat{\mathbf M}_n^x$ the diagonal matrix with entries the $r$-largest eigenvalues of $\widehat{\bm\Gamma}_n^x$, and as $\widehat{\mathbf V}_n^x$ the $n\times r$ matrix of the corresponding normalized eigenvectors. Moreover, let $\widehat{\bm{\mathcal S} }^{(0)}$ be a $r\times r$ diagonal matrix with entries $\mathbb I([\widehat{\mathbf V}_n^x]_{1j}\ge 0)- \mathbb I([\widehat{\mathbf V}_n^x]_{1j}< 0)$, $j=1,\ldots,r$. Then,
Finally, letting $\widehat{\bm\lambda}_i^{(0)\prime}$ be the $i$-th row of $\widehat{\bm\Lambda}_n^{(0)}$, and $\widetilde{\xi}_{it}$ the $i$th component of $\widetilde{\bm\xi}_{nt}$, then
The vector of initial estimates of parameters is then:
and it is used to run the first iteration of the EM algorithm.
The following iterations are stated for given initial conditions $\mathbf F_{0|0}$ and $\mathbf P_{0|0}$ and given the true parameters $\bm\varphi_n$.
The Kalman filter is based on the forward iterations for $t=1,\ldots, T$:
Moreover, by combining (ref) and (ref), we obtain the Riccati difference equation:
The Kalman smoother is then based on the backward iterations for $t=T,\ldots, 1$:
Finally, ${\mathbf C}_{t,t-1|T}$ can be obtained from a state space model with an augmented state vector containing both $\mathbf F_t$ and $\mathbf F_{t-1}$, by taking the $r\times r$ off-diagonal block of the $2r\times 2r$ matrix defined in (ref) but for the augmented model.
An equivalent way of implementing (ref), which does not require matrix inversion is in DK01, which is defined by the backward iterations for $t=T,\ldots, 1$
where $\mathbf r_T=\mathbf 0_{r}$, $\mathbf N_T=\mathbf 0_{r}$ and by construction $\mathbf A\mathbf P_{t|t}=\mathbf L_t\mathbf P_{t|t-1}$.
The Kalman filter is initialized as follows. At the first iteration of the EM algorithm, i.e., when $k=0$, we set ${\mathbf F}^{(0)}_{0|0}=\mathbf 0_r$ and ${\mathbf P}^{(0)}_{0|0}=\mathbf I_r$, consistently with Assumption (ref)(b). Other initializations as ${\mathbf P}^{(0)}_{0|0}=\kappa_0 \mathbf I_r$ for some finite real $\kappa_0>0$ are also possible. At any successive iteration of the EM algorithm, i.e, when $k>0$, we set ${\mathbf F}^{(k)}_{0|0}={\mathbf F}_{0|T}^{(k-1)}$ and ${\mathbf P}^{(0)}_{0|0}=\mathbf I_r$.
To run the Kalman smoother we start using the last predictions of the Kalman filter. Thus, for any $k\ge 0$, we set ${\mathbf F}^{(k)}_{T+1|T}= \widehat{\mathbf A}^{(k)} {\mathbf F}^{(k)}_{T|T}$ where ${\mathbf F}^{(k)}_{T|T}$ is obtained from (ref), and we set ${\mathbf P}^{(k)}_{T+1|T}=\widehat{\mathbf A}^{(k)}\mathbf P_{T|T}^{(k)} \widehat{\mathbf A}^{(k)\prime} +\widehat{\bm\Gamma}^{v(k)}$, where ${\mathbf P}^{(k)}_{T|T}$ is obtained from (ref).
To stop the EM algorithm we adopt the following convergence rule. We fix a maximum finite number of iterations $k_{\max}$, and we stop it at the first iteration $k^*\le k_{\max}$ such that:
where $\varepsilon$ is a pre-specified tolerance level. In this case, the log-likelihood is computed using its prediction error formulation obtained from the Kalman filter:
where $\underline{\mathbf F}_{t|t-1}$ and $\underline{\mathbf P}_{t|t-1}$ are computed using (ref) and (ref), respectively, when using generic values of the parameters. Similar convergence criteria can be found in boothhobert99 and MLT07.
\setcounter{equation}{0} \setcounter{lem}{0}
Hereafter, we assume without loss of generality that $\bm \mu_n=\mathbf 0_n$ and $p_F=1$ with $\mathbf A\equiv\mathbf A_1$.
Consider the EM algorithm initialized using the PC estimators of the parameters as defined in Section (ref). At $k^*=0$, from (ref), we have
Now,
and
Throughout, let $\bm y_t=\mathbf F_t$ or $\bm y_t = x_{it}$. Then, we have to consider
Let us consider each term in (ref). First,
by Lemmas (ref), (ref), and (ref), and since $\Vert\widehat{\mathbf A}^{(0)}\Vert\le \Vert{\mathbf A}\Vert+\Vert\widehat{\mathbf A}^{(0)}-\mathbf A\Vert =O_p(1)$, by Assumption (ref)(d) and Lemma (ref)(i). Second, from (ref) and (ref) in the proof of Lemma (ref)
because of Lemma (ref) and since
by Assumption (ref)(a), Lemmas (ref), (ref)(i), (ref)(ii), (ref)(iii), and (ref)(iv), and because $n^{-1}\Vert\mathbb{E}[\bm\xi_{nt}\xi_{it}]\Vert^2 \le n^{-1}\sum_{j=1}^n \vert\mathbb{E}[\xi_{jt}^2\xi_{it}^2] \vert\le \mathrm K_\xi$, by Assumption (ref)(d). Note that the first and third relations in (ref) cover also the case $\bm y_t=\mathbf F_t^\prime$.
Finally, let us consider the last term in (ref). From (ref) in the proof of Lemma (ref)
Then,
by (ref) and (ref)-(ref) in the proof of Lemma (ref). Moreover,
by (ref) and Lemmas (ref)(ii) and (ref)(v). Regarding $III_c$, if $\bm y_t=\mathbf F_t$, we have
by Lemmas (ref)(iii) and (ref)(iv). If $\bm y_t=x_{it}$, we have
by Assumption (ref)(a), and Lemmas (ref)(iii), (ref)(iv), and (ref)(v). Last,
by (ref) and Lemma (ref)(v). From (ref), (ref), (ref), (ref), (ref), and (ref)
Combining (ref), (ref), and (ref) we have
which once substituted into (ref) and (ref), jointly with Lemma (ref) give
and
Therefore, from (ref)
with $\bm\lambda_i^{\text{\tiny OLS}} = (T^{-1}\sum_{t=1}^T\mathbf F_t\mathbf F_t^\prime)^{-1}(T^{-1}\sum_{t=1}^T\mathbf F_tx_{it})$. And by Lemma (ref)(i) we also have
indeed, recalling that $\bm\Gamma^F=\mathbf I_r$ by Assumption (ref)(b), by Lemma (ref)(i) and Weyl's inequality MK04 we have $\vert \nu^{(r)}(T^{-1}\sum_{t=1}^T \mathbf F_t\mathbf F_t^\prime)^{-1}\vert = O_p(T^{-1/2})$ which implies $\Vert (T^{-1}\sum_{t=1}^T \mathbf F_t\mathbf F_t^\prime)^{-1} \Vert = O_p(1).$ From (ref) and (ref)
Moreover, by letting $\bm y_t=n^{-1/2} \mathbf x_{nt}$ the above proof leads to
and, therefore, using also Lemma (ref)(ii), we have
This proves parts (a.1) and (a.2) when $k^*=0$.
Following the same reasoning leading to (ref) and by Lemma (ref)(iv) we can easily prove that
where $\widehat{\mathbf A}^{(1)}$ is defined in (ref) and $\mathbf A^{\text{\tiny OLS}} = (T^{-1}\sum_{t=1}^T \mathbf F_t\mathbf F_{t-1}^\prime )(T^{-1}\sum_{t=1}^T \mathbf F_{t-1}\mathbf F_{t-1}^\prime )^{-1}$ (recall that $\mathbf F_0=\mathbf 0_r$ by Assumption (ref)(i)). To prove (ref) we also use the fact that $\max_{t=1,\ldots, T}\Vert \mathbf C_{t,t-1|T}^{(0)}\Vert=O(n^{-1})$ since it can be obtained by the upper right block of $\mathbf P_{t|T}^{(0)}$ when this is computed from the the Kalman smoother having the augmented state vector $(\mathbf F_t^\prime\,\mathbf F_{t-1}^\prime)^\prime$. This proves part (a.4) when $k^*=0$.
Likewise, using (ref) and the same reasoning leading to (ref), by Lemma (ref)(v), we can easily prove also that
where $\widehat{\bm\Gamma}^{v(1)}$ is defined in (ref) and $\bm\Gamma^{v\text{\tiny OLS}}=T^{-1}\sum_{t=1}^T (\mathbf F_t-\mathbf A^{\text{\tiny OLS}}\mathbf F_{t-1}) (\mathbf F_t-\mathbf A^{\text{\tiny OLS}}\mathbf F_{t-1})^\prime$. To prove (ref) we need to use also the intermediate quantity $T^{-1}\sum_{t=1}^T (\mathbf F_t-\widehat{\mathbf A}^{(1)}\mathbf F_{t-1}) (\mathbf F_t-\widehat{\mathbf A}^{(1)}\mathbf F_{t-1})^\prime$, which is $\min(n,\sqrt T)$-consistent because of (ref). This proves part (a.5) when $k^*=0$.
Finally, using again the same reasoning leading to (ref), by Lemma (ref)(iii), we can prove that
where $\widehat{\sigma}_i^{2(1)}$ is defined in (ref) and $\sigma_i^{2\text{\tiny OLS}}=T^{-1}\sum_{t=1}^T (x_{it}-{\bm\lambda}_i^{\text{\tiny OLS}\prime}\mathbf F_t)^2$. To prove (ref) we need to use also the intermediate quantity $T^{-1}\sum_{t=1}^T (x_{it}-\widehat{\bm\lambda}_i^{(1)\prime}\mathbf F_t)^2$, which is $\min(n,\sqrt T)$-consistent because of (ref). This proves part (a.3) when $k^*=0$.
Now, from (ref) and (ref) using the same reasoning of the proof of Lemma (ref)(ii), which in turn requires (ref) and (ref), again it follows that
And, using the same reasoning as in the proof of Lemma (ref) but using now (ref), (ref), and (ref),
From (ref), (ref), (ref), and the relations in (ref) and following the same steps as in the proofs of Lemmas (ref), (ref), and (ref), we get
and also $\Vert\mathbf F_{t|t}^{(1)}\Vert=O_p(1)$ and $\Vert\mathbf F_{t|T}^{(1)}\Vert=O_p(1)$ by the same arguments in Lemma (ref). It follows that we can apply the same steps as in the proofs of Lemmas (ref), (ref), and (ref) to get $$ \Vert {\mathbf F}_{t|T}^{(1)}-{\mathbf F}_{t}\Vert=O_p(\max(n^{-1/2},T^{-1/2})). $$ This proves part (b) when $k^*=0$.
Then we can show that (ref) and (ref) still hold when using ${\mathbf F}_{t|T}^{(1)}$ in place of ${\mathbf F}_{t|T}^{(0)}$ and using also the last of (ref) we prove $\min(n,\sqrt T)$-consistency of $\widehat{\bm\lambda}_i^{(2)}$. Similarly we can prove $\min(n,\sqrt T)$-consistency of $n^{-1/2}\widehat{\bm\Lambda}_n^{(2)}$, $\widehat{\mathbf A}^{(2)}$, $\widehat{\bm\Gamma}^{v(2)}$, and $\sigma_i^{2(2)}$. It is then clear that we can repeat the same reasoning leading to (ref), (ref), and (ref) but when $k^*=1$. So these arguments hold for all $k^*\ge 0$. This completes the proof. $\Box$
For part (a.1), for any $k^*\ge 0$, we have (recall that $\widehat{\bm\lambda}_{i}\equiv \widehat{\bm\lambda}_{i}^{(k^*+1)}$)
From Lemma (ref)
From Lemma (ref)(i)
From Lemma (ref)(i)
From Lemma (ref)(i)
By using (ref), (ref), (ref), and (ref) into (ref), we prove part (a.1).
For part (a.2), for any $k^*\ge 0$, we have (recall that $\widehat{\bm\Lambda}_{n}\equiv \widehat{\bm\Lambda}_{n}^{(k^*+1)}$)
and the proof follows from Lemmas (ref)(ii), (ref)(ii), and (ref)(i), and since
by Lemma (ref)(i).
For part (b), from part (a.1) and (ref), if $n^{-1}\sqrt T\log^{2/\delta_v} T\to 0$, as $n,T\to\infty$, we have
Now, since $\{\mathbf F_t \xi_{it}\}$ is strongly mixing with exponentially decaying coefficients by bradley05 (see also (ref) in the proof of Lemma (ref)), and given that by Assumption (ref) the following Cram\'er condition holds $$ \sup_{m\ge 1} r^{-1/\delta} (\mathbb{E}[\vert \xi_{it}F_{jt} \vert^m])^{1/m}\le K, $$ for some finite positive reals $\delta\in\left(0,\frac{\delta_v\delta_\xi}{\delta_v+\delta_\xi}\right)$ and $K$ independent of $t$, $i$, and $j$ KC18, then the Central Limit Theorem by ibra62 applies, i.e.,
Therefore, by Lemmas (ref)(i) and (ref), and Assumption (ref)(b)
Thus from (ref), (ref), and (ref), by Slutsky's Theorem, we have
where,
since $\bm\Gamma^F=\mathbf I_r$ because of Assumption (ref)(b) and $\{\mathbf F_t\}$ and $\{\xi_{it}\}$ are independent processes because of Lemma (ref). This proves part (b). Part (c) is straightforward. This completes the proof. $\Box$
Recall the definitions $\widehat{\mathbf F}_t\equiv \widehat{\mathbf F}_{t|T}\equiv \mathbf F_{t|T}^{(k^*+1)}$, for any $k^*\ge 0$. From Lemmas (ref)(ii) and (ref)(iii),
where $\widehat{\mathbf F}_{t}^{\text{\tiny \upshape {WLS}}}=(\widehat{\bm\Lambda}_n^{\prime}(\widehat{\bm\Sigma}_n^{\xi})^{-1}\widehat{\bm\Lambda}_n)^{-1}\widehat{\bm\Lambda}_n^{\prime}(\widehat{\bm\Sigma}_n^{\xi})^{-1}\mathbf x_{nt}.$
Now,
Let us consider each term in (ref). First, consider term $\rm A$ and notice that
by Lemmas (ref)(ii), (ref)(i), and (ref)(i) (see also (ref) in the proof of Proposition (ref)).
Therefore, from (ref)
Then,
by Lemmas (ref)(iii), (ref)(iv), and (ref). Moreover, ${\rm A}.2 = O_p(\max(n^{-1}\log^{2/\delta_v}T,n^{-1/2}T^{-1/2}\sqrt{\log n},T^{-1}))$, because of (ref) and Lemmas (ref), (ref)(iii), and Assumption (ref)(a) which implies $\Vert({\bm\Sigma}_n^{\xi})^{-1}\Vert\le C_\xi$. This, jointly with (ref) and (ref) implies that
since $\Vert \mathbf F_t\Vert = O_p(1)$ because $\mathbb{E}[F_{jt}^2]=1$, $j=1,\ldots, r$, by Assumption (ref)(b).
Second, by Proposition (ref)(a) and Lemma (ref)(v)
and again since $\Vert \mathbf F_t\Vert = O_p(1)$ because $\mathbb{E}[F_{jt}^2]=1$, $j=1,\ldots, r$, by Assumption (ref)(b).
Third,
by Lemmas (ref)(iii) and (ref)(i).
Fourth, and last,
Then, ${\rm D.1}= O_p(\max(n^{-3/2}\log^{2/\delta_v}T,n^{-1/2}T^{-1/2}\sqrt{\log n}))$, by Lemmas (ref)(i) and (ref)(iv). Moreover,
We then have the following results. First,
by Lemmas (ref)(i) and (ref)(i). Second,
by Lemmas (ref)(i) and (ref)(ii). Third, clearly by (ref) and (ref)
By (ref), (ref), and (ref), and Lemma (ref) we have ${\rm D.2}= O_p(\max(n^{-3/2}\log^{2/\delta_v}T,n^{-1/2}T^{-1/2}\sqrt{\log n})$. Last, ${\rm D.3}=O_p(\max(n^{-2}\log^{4/\delta_v}T, T^{-1}\log n)){\rm D.2}$, by Lemma (ref)(ii), thus it is dominated by ${\rm D.2}$. Therefore,
By substituting (ref), (ref), (ref), and (ref), into (ref) we have
which, once substituted into (ref), proves part (a.1).
For part (a.2), let $\widehat{\bm{\mathcal F}}_T^{\text{\tiny KF}}=(\widehat{\mathbf F}_{1|1} \cdots \widehat{\mathbf F}_{T|T})^\prime$ and $\widehat{\bm{\mathcal F}}_T^{\text{\tiny WLS}}=(\widehat{\mathbf F}_{1}^{\text{\tiny WLS}} \cdots \widehat{\mathbf F}_{T}^{\text{\tiny WLS}})^\prime$, and recall that, by definition, $\widehat{\bm{\mathcal F}}_T=(\widehat{\mathbf F}_{1|T} \cdots \widehat{\mathbf F}_{T|T})^\prime=(\widehat{\mathbf F}_{1} \cdots \widehat{\mathbf F}_{T})^\prime$. From (ref), we have
By Lemma (ref) the first two terms on the rhs of (ref) are such that
While, for the last term on the rhs of (ref), letting $\bm{\mathcal E}_{nT}=(\bm\xi_{n1}\cdots \bm\xi_{nT})^\prime$, we have
For the first term on the rhs of (ref) we have
where
by (ref), (ref), Lemma (ref), and Assumption (ref)(a), while
by Proposition (ref)(a) and (ref)(ii). By substituting (ref) and (ref) into (ref)
Moving to the second term on the rhs of (ref), we have
Then, considering each term on the rhs of (ref),
where
by Lemmas (ref)(vi) and (ref). Whereas $\mathcal B_{1.b}=O_p(\max(n^{-1}\log^{2/\delta_v}T, n^{-1/2}T^{-1/2}\sqrt{\log n},T^{-1}))$ by (ref) and Lemma (ref)(vi). Therefore,
Moreover, letting $\bm\zeta_i=(\xi_{i1}\cdots \xi_{iT})^\prime$,
by Assumption (ref)(a) and Lemmas (ref)(ii), (ref)(ii), and Last,
by Proposition (ref)(a) and Lemmas (ref)(vi) and (ref)(v). By substituting (ref), (ref), and (ref) into (ref)
Finally, for the third term on the rhs of (ref), we have
by Lemma (ref)(v).
By substituting (ref), (ref), and (ref) into (ref), and since $n\Vert(\widehat{\bm\Lambda}^\prime(\widehat{\bm\Sigma}_n^\xi)^{-1}\widehat{\bm\Lambda}_n)^{-1} \Vert=O_p(1)$ by Lemma (ref)(iii),
which, once substituted in (ref) together with (ref), proves part (a.2).
Turning to part (b), from (ref), (ref), (ref) and (ref), if $T^{-1}\sqrt {n\log n}\to 0$, as $n,T\to\infty$, we have
Then, by Assumption (ref)(e), as $n\to\infty$, it holds that
Moreover, from Lemma (ref)(iii) we have that $$ \Vert n^{-1} {\bm\Lambda}_n^\prime({\bm\Sigma}_n^{\xi})^{-1}{\bm\Lambda}_n-\bm\Sigma_{\Lambda\Sigma\Lambda}\Vert=o_p(1), $$ for some finite and positive definite $r\times r$ matrix $\bm\Sigma_{\Lambda\Sigma\Lambda}$, which, jointly with Lemma (ref)(v), implies
From (ref), (ref), and (ref), by Slutsky's Theorem, we have \[ \sqrt n (\widehat{\mathbf F}_t-\mathbf F_t)\stackrel{d}{\to}\mathcal N(\mathbf 0_r,\bm{\mathcal W}_t), \] where \[ \bm{\mathcal W}_t =(\bm\Sigma_{\Lambda\Sigma\Lambda})^{-1}\left\{\lim_{n\to\infty} n^{-1} \sum_{i,j=1}^n \bm\lambda_i\bm\lambda_j^\prime\mathbb{E}_{}[\xi_{it}\xi_{jt}] (\sigma_i^2\sigma_j^2)^{-1}\right\}(\bm\Sigma_{\Lambda\Sigma\Lambda})^{-1}. \] This proves part (b). Part (c) is straightforward. This completes the proof. $\Box$
First notice that
Then,
by Propositions (ref)(a) and (ref)(a), Assumption (ref)(a), and since $\Vert \mathbf F_t\Vert=O_p(1)$ because $\mathbb{E}[F_{jt}^2]=1$, $j=1,\ldots,r$, by Assumption (ref)(b). This proves part (a).
For part (b), let us denote $\delta_{nT}=\min(\sqrt n,\sqrt T)$, for simplicity of notation. Consider the first term on the rhs of (ref). Define $\bm K_\Lambda=n(\bm\Lambda_n^\prime(\bm\Sigma_n^\xi)^{-1}\bm\Lambda_n)^{-1}$. Then, from (ref) in the proof of Proposition (ref)
since $\delta_{nT}T^{-1}\le \delta_{nT}\max(n^{-1},T^{-1})=\delta_{nT}\delta_{nT}^{-2}\to 0$. Similarly, consider the second term on the rhs of (ref) and define $\bm K_F=(T^{-1}\sum_{t=1}^T\mathbf F_t\mathbf F_t^\prime)^{-1}$. Then, from (ref) in the proof of Proposition (ref)
since $\delta_{nT}n^{-1}\le \delta_{nT}\max(n^{-1},T^{-1})=\delta_{nT}\delta_{nT}^{-2}\to 0$. From Propositions (ref)(b) and (ref)(b), as $n,T\to\infty$,
where $\mathcal C^F_{it}=\bm\lambda_i^\prime\bm{\mathcal W}_t\bm\lambda_i$ and $\mathcal C^\lambda_{it}=\mathbf F_t^\prime \bm{\mathcal V}_i\mathbf F_t$. Moreover, $\mathcal A_{it}$ and $\mathcal B_{it}$ are asymptotically independent, since the former is a cross-sectional sum of random variables, while the latter is the sum of a given time series and under Lemmas (ref)(i)-(ref)(iii) are weakly serially and cross-sectionally correlated in the same sense as assumed by Bai03.
Define $a_{nT}=\delta_{nT}n^{-1/2}$ and $b_{nT}=\delta_{nT}T^{-1/2}$. Then, substituting (ref) and (ref) into (ref), we obtain
because of (ref). From (ref) and (ref), by Slutsky's theorem and following the same reasoning as in Bai03, as $n,T\to\infty$, we have \[ \delta_{nT}{\left(a_{nT}^2\mathcal C^F_{it} + b_{nT}^2\mathcal C^\lambda_{it}\right)^{-1/2}} { (\widehat{\chi}_{it}-\chi_{it}) } =(n^{-1}\mathcal C^F_{it} + T^{-1}\mathcal C^\lambda_{it})^{-1/2}(\widehat{\chi}_{it}-\chi_{it}) \stackrel{d}{\rightarrow}\mathcal N(0,1), \] which completes the proof. $\Box$
For part (a.1), for any $k^*\ge 0$, we have (recall that $\widehat{\sigma}_{i}^2\equiv \widehat{\sigma}_{i}^{2(k^*+1)}$)
by Lemmas (ref)(iii), (ref)(iii), (ref)(ii), and (ref)(i).
For part (a.2), for any $k^*\ge 0$, we have (recall that $\widehat{\bm\Sigma}_{n}^\xi\equiv \widehat{\bm\Sigma}_{n}^{\xi(k^*+1)}$)
and the proof follows from Lemma (ref)(iii) and since
by Lemma (ref)(i), and
by Lemma (ref)(ii).
For part (a.3), for any $k^*\ge 0$, we have (recall that $\widehat{\mathbf A}\equiv \widehat{\mathbf A}^{(k^*+1)}$)
by Lemmas (ref)(iv), (ref)(i), (ref)(iii), and (ref)(ii).
For part (a.4), for any $k^*\ge 0$, we have (recall that $\widehat{\bm\Gamma}^v\equiv \widehat{\bm\Gamma}^{v(k^*+1)}$)
by Lemmas (ref)(v), (ref)(ii), (ref)(iv), and (ref)(iii). This completes the proof of part (a).
To prove part (b), we need sharper rates and, to this end, we use the closed form expressions of the estimated parameters in (ref), (ref), and (ref).
Start with part (b.1). From (ref), for any $k^*\ge 0$, we have
Therefore, since by construction $T^{-1}\sum_{t=1}^T \mathbf F_t (x_{it}-{\bm\lambda}_i^{\text{\tiny OLS}\prime}\mathbf F_t) =0$, from (ref)
For term $\mathcal A$ on the rhs of (ref) we have
by Lemma (ref)(i). Term $\mathcal B$ is dominated by term $\mathcal A$. For term $\mathcal C$ on the rhs of (ref) we have
by Lemma (ref)(iv) when $k^*\ge 1$ and by Lemma (ref) when $k^*= 0$. For term $\mathcal D$ on the rhs of (ref) we have
by Lemmas (ref)(i) and (ref)(ii), and Assumption (ref)(a).
Now by substituting (ref), (ref), and (ref) into (ref), we have
by (ref)-(ref) in the proof of Proposition (ref), Lemma (ref)(i) combined with Assumption (ref)(b), and since $\Vert \widehat{\bm\lambda}_i\Vert\le \Vert \widehat{\bm\lambda}_i-\bm\lambda_i\Vert+\Vert {\bm\lambda}_i\Vert=O_p(1)$ by Proposition (ref) and Assumption (ref)(a).
Therefore, from (ref) and Lemma (ref)(i), if $n^{-1}\sqrt T\log^{2/\delta_v} T\to 0$, as $n,T\to\infty$, we have
where we used also (ref) in the proof of Lemma (ref). Now, by Assumption (ref)(c) and davidson, we have that $\{\xi_{it}^2-\sigma_i^2\}$ is strongly mixing with exponentially decaying coefficients and such that by Assumption (ref) $\sup_{m\ge 1} r^{-1/\delta_\xi} (\mathbb{E}[\vert \xi_{it}^2-\sigma_i^2\vert^m])^{1/m}\le K_1$ for some finite positive real $K_1$ independent of $t$ and $i$ KC18. Then, the Central Limit Theorem by ibra62 applies, i.e.,
From (ref), (ref), and Slutsky's Theorem, and by noticing that the excess kurtosis is given by $\kappa_i = \mathbb{E}[\xi_{it}^4]/\sigma_i^4-3$, so that $\mathbb{E}[\xi_{it}^4]=\sigma_i^4(\kappa_i+3)$, we prove part (b.1).
For part (b.2), from (ref), for any $k^*\ge 0$, we have
Now,
and
where we used: Lemma (ref)(i), Lemma (ref)(iv) when $k^*\ge 1$ or Lemma (ref) when $k^*= 0$ which can be both applied to $\mathbf C_{t,t-1|T}^{(k^*)}$ can be obtained by the upper right block of $\mathbf P_{t|T}^{(k^*)}$ when this is computed from the the Kalman smoother having the augmented state vector $(\mathbf F_t^\prime\,\mathbf F_{t-1}^\prime)^\prime$, and we also used the fact that $\Vert(T^{-1}\sum_{t=2}^T \mathbf F_{t-1}\mathbf F_{t-1}^\prime )^{-1}\Vert=O_p(1)$ by Lemma (ref). For $\mathfrak B$ we also used Lemma (ref)(i) combined with the fact that $\Vert \bm\Gamma_1^F\Vert\le 1$ by Cauchy-Schwartz inequality and Assumption (ref)(b). Clearly $\mathfrak C$ is dominated by $\mathfrak A$ and $\mathfrak B$.
By substituting (ref) and (ref) into (ref), we have
Therefore, if $n^{-1}\sqrt T\log^{2/\delta_v} T\to 0$, as $n,T\to\infty$, we have
and the proof of part (b.2) follows directly from Hamilton and Slutsky's Theorem.
For part (b.3), following a decomposition analogous to the one in (ref), from (ref), for any $k^*\ge 0$, we have
Therefore, since by construction $T^{-1}\sum_{t=2}^T \mathbf F_{t-1} (\mathbf F_{t}-{\mathbf A}^{\text{\tiny OLS}\prime}\mathbf F_{t-1})^\prime =T^{-1}\sum_{t=2}^T (\mathbf F_{t}-{\mathbf A}^{\text{\tiny OLS}\prime}\mathbf F_{t-1}) \mathbf F_{t-1}^\prime =\mathbf 0_{r\times r}$, by using again Lemmas (ref) and Lemma (ref)(iv) when $k^*\ge 1$ or Lemma (ref) when $k^*= 0$ which can be both applied to $\mathbf C_{t,t-1|T}^{(k^*)}$, as argued above, and using also (ref), we obtain
Therefore, if $n^{-1}\sqrt T\log^{2/\delta_v} T\to 0$, as $n,T\to\infty$, we have
and the proof of part (b.3) follows directly from Hamilton and Slutsky's Theorem. This completes the proof. $\Box$
For part (a), for ease of notation and without loss of generality let $s(j)=j$ for all $j=1,\ldots, \bar n$. Then, from (ref), since $\bar n$ is finite,
Moreover, since $\bar n$ is finite we can still apply the Central Limit Theorem by ibra62 so that, as $T\to\infty$,
with
where we used the fact that $\{\mathbf F_t\}$ and $\{\bm\xi_{nt}\}$ are independent processes because of Lemma (ref). The proof of part (a.1) follows from Lemmas (ref)(i) and (ref), (ref), (ref), and Slutsky's theorem. For part (a.2) just notice that, if $\mathbb{E}[(\bm\xi_{n1}^\prime\cdots\bm\xi_{nT}^\prime)^\prime(\bm\xi_{n1}^\prime\cdots\bm\xi_{nT}^\prime)]=\mathbf I_T\otimes \bm\Sigma_n^\xi$ for all $n,T\in\mathbb N$, then $\mathbb{E}\left[\bm\xi_{\bar n t}\bm\xi_{\bar n s}^\prime\right]= \mathbf 0_{\bar n\times\bar n}$ if $t\ne s$ while $\mathbb{E}\left[\bm\xi_{\bar n t}\bm\xi_{\bar n t}^\prime\right]= \bm\Sigma_{\bar n}^\xi$. By substituting into (ref) we get $\bm\Sigma_{\bar n}=\bm\Sigma_{\bar n}^\xi\otimes \bm\Gamma^F$, and $\bm{\mathcal V}_{\bar n}=(\mathbf I_{\bar n}\otimes\bm\Gamma^F)^{-1}(\bm\Sigma_{\bar n}^\xi\otimes \bm\Gamma^F)(\mathbf I_{\bar n}\otimes\bm\Gamma^F)^{-1}=\bm\Sigma_{\bar n}^\xi\otimes(\bm\Gamma^F)^{-1}$, thus proving part (a.2).
Turning to part (b), for ease of notation and without loss of generality let $s(j)=j$ for all $j=1,\ldots, \bar T$. Then, from (ref) in the proof of Proposition (ref), since $\bar T$ is finite
where $\bm \Xi_{n\bar T}=(\bm\xi_{n1}^\prime\cdots\bm\xi_{n\bar T}^\prime)^\prime$. Moreover, since $\bar T$ is finite we can still apply the Central Limit Theorem in Assumption (ref)(e), so that as $n\to\infty$
with
and recall that $\bm\zeta_{i\bar T}=(\xi_{i1}\cdots\xi_{i\bar T})^\prime$. The proof of part (b.1) follows from (ref) in the proof of Proposition (ref), (ref), (ref), and Slutsky's theorem. For part (b.2) just notice that, if $\mathbb{E}[(\bm\zeta_{1T}^\prime\cdots \bm\zeta_{nT}^\prime)^\prime(\bm\zeta_{1T}^\prime\cdots \bm\zeta_{nT}^\prime)]=\bm\Sigma_n^\xi\otimes\mathbf I_T$ for all $n,T\in\mathbb N$, then $\mathbb{E}_{}[\bm\zeta_{i\bar T}\bm\zeta_{j\bar T}^\prime]= \mathbf 0_{\bar T\times\bar T}$ if $i\ne j$ while $\mathbb{E}_{}[\bm\zeta_{i\bar T}\bm\zeta_{i\bar T}^\prime]= \sigma_i^2\mathbf I_{\bar T}$. By substituting into (ref) we get $\bm\Omega_{\bar T}=\mathbf I_{\bar T}\otimes \bm\Sigma_{\Lambda\Sigma\Lambda}$, and $\bm{\mathcal W}_{\bar T}=(\mathbf I_{\bar T}\otimes \bm\Sigma_{\Lambda\Sigma\Lambda})^{-1}(\mathbf I_{\bar T}\otimes \bm\Sigma_{\Lambda\Sigma\Lambda})(\mathbf I_{\bar T}\otimes \bm\Sigma_{\Lambda\Sigma\Lambda})^{-1}=\mathbf I_{\bar T}\otimes (\bm\Sigma_{\Lambda\Sigma\Lambda})^{-1}$, thus proving part (b.2). This completes the proof. $\Box$
Under Assumptions (ref), (ref), (ref), and (ref), from MBPCAQML, the asymptotic covariances of the PC estimator of the loadings is
where $\bm\Gamma^F=\mathbf I_r$, by Assumption (ref)(b). Therefore, the expression of $\bm{\mathcal V}_i^{\text{\tiny PC}}$ in (ref) coincides with the expression of $\bm{\mathcal V}_i$ for the asymptotic covariances of the EM estimator given in Proposition (ref)(b). This proves part (a).
Turning to part (b). Since, because of Proposition (ref), $\lim_{n,T\to\infty} \sqrt n\mathbb{E}[(\widehat{\mathbf F}_t-\mathbf F_t)]=\mathbf 0_r$, then
Similarly, because of Lemma (ref), $\lim_{n,T\to\infty} \sqrt n\mathbb{E}[(\widetilde{\mathbf F}_t-\mathbf F_t)]=\mathbf 0_r$, which implies
where we used the fact that $\lim_{n\to\infty} n^{-1}\mathbf M_n^\chi=\lim_{n\to \infty} n^{-1}\bm\Lambda_n^\prime\bm\Lambda_n$, because of Assumption (ref)(b).
Moreover, if $\sqrt{n\log n}/T\to 0$ (which implies also $\sqrt n/T\to0$), Proposition (ref) (see in particular (ref) in its proof) and Lemma (ref) (see in particular (ref) in its proof) jointly imply that
where
From (ref) and (ref) it follows that
Let us now define $\bm{\mathcal O}_n^\xi=\bm\Gamma_n^\xi-\bm\Sigma_n^\xi$, which is the $n\times n$ matrix of off-diagonal entries of the idiosyncratic covariance. Because of (ref), (ref), and Lemma (ref) we can write
Now, let $\bm{\mathfrak V}_n=\mathbf V_n^{\chi\prime}(\bm\Sigma_n^\xi)^{-1}\bm{\mathcal O}_n^\xi (\bm\Sigma_n^\xi)^{-1}\mathbf V_n^\chi$, then, for any $h,k=1,\ldots,r$,
Now, let $\bm\iota_{ni}$ be the $n$-dimensional vector with $i$th entry equal one and all other equal zero, and let $\bm s_j$ the $r$-dimensional vector with $j$th entry equal one and all other equal zero. Then, since $\mathbf V_n^{\chi} =\bm\Gamma_n^\chi\mathbf V_n^{\chi}(\mathbf M_n^\chi)^{-1}$, there exists a finite positive integer $\bar n$, such that, for all $n \ge \bar n$
where we used Assumptions (ref)(a) and (ref)(b), and Lemmas (ref)(iv) and (ref).
Therefore, from (ref) and (ref), and Assumption (ref)(a),
Moreover, since by assumption we have \[ \lim_{n\to\infty}n^{-1} \sum_{i,\ell=1}^n\vert [\bm{\mathcal O}_n^\xi]_{i\ell}\vert=\lim_{n\to\infty}n^{-1} \sum_{\substack{i,\ell=1\\ i\ne \ell}}^n\vert [\bm{\Gamma}_n^\xi]_{i\ell}\vert=0, \] then, from (ref), we have $\lim_{n\to\infty}\max_{h,k=1,\ldots r}\vert [\bm{\mathfrak V}_n]_{hk}\vert=0$. Therefore,
Following the same reasoning it also holds that
By substituting, (ref) and (ref) into (ref), we obtain $\bm{\mathcal C}_2=\mathbf 0_{r\times r}.$ Finally,
Therefore, $\bm{\mathcal C}_1$ is positive definite since $\bm{\mathcal W}_t$ and $\bm{\mathcal W}_t^{\text{\tiny PC}}$ are positive definite by Assumptions (ref)(a), (ref)(a), and (ref)(f), and, from (ref) we prove part (ii). This completes the proof. $\Box$
\setcounter{lem}{0} \numberwithin{lem}{section}
\noindentProof. Using Assumptions (ref)(a) and (ref)(b), we have:
Similarly,
and
Defining, $M_1=\frac{(C_\xi+M_\xi)(1+\rho)}{1-\rho}$, $M_2=C_\xi+M_\xi$, and $M_3= \frac{C_\xi (1+\rho)}{1-\rho}$, we prove parts (i), (ii), and (iii).
For part (iv), first notice that, for all $n\in\mathbb N$, the $r$ non-zero eigenvalues of $\bm\Gamma_n^\chi$ are also the $r$ eigenvalues of $n^{-1}\bm\Lambda_n^\prime\bm\Lambda_n\bm\Gamma^F$. Thus, by MK04, for all $j=1,\ldots, r$ and all $n\in\mathbb N$, we have
The proof then follows from Assumptions (ref)(a) and (ref)(b). Indeed, by continuity of eigenvalues, Assumption (ref)(a) implies that, for all $j=1,\ldots, r$, and all $n>N_0$
and there exist finite positive reals $m_\lambda$ and $M_\lambda$ such that $0<m_{\lambda}^2\le \nu^{(r)}(\bm\Sigma_\Lambda)\le \nu^{(1)}(\bm\Sigma_\Lambda)\le M_{\lambda}^2<\infty$. Similarly, Assumption (ref)(b), implies that there exist finite positive reals $m_F$ and $M_F$ such that $0<m_{F}\le \nu^{(r)}(\bm\Gamma^F)\le \nu^{(1)}(\bm\Gamma^F)\le M_F<\infty$.
For part (v), by Assumptions (ref)(a) and (ref)(b):
Part (vi) follows from parts (iv) and (v) and Weyl's inequality MK04. This completes the proof. $\Box$
\noindentProof. By definition and using Assumption (ref)(a), $$ \lim_{n\to\infty} n^{-1/2}\Vert\bm\Lambda_n\Vert = \lim_{n\to\infty} n^{-1/2} \sqrt{ \nu^{(1)}(\bm\Lambda_n^\prime\bm\Lambda_n)} =\sqrt{\nu^{(1)}(\bm\Sigma_\Lambda)}\le M_\lambda. $$ This completes the proof. $\Box$
Proof. By MK04
by Assumption (ref)(a) and since $\bm P$ finite by assumption. Then, as shown in (ref) in the proof of Lemma (ref)(iv),
Thus, we have $\lim_{n\to\infty} n^{-1}\nu^{(r)}(\bm\Lambda_n^\prime\bm\Lambda_n) \ge \underline C_r$, which, once substituted in (ref), proves part (i).
For part (ii) the proof is the same as part (i), but in (ref) we use Lemma (ref)(v) instead of Assumption (ref)(a), thus replacing $C_\xi$ with $M_2$. Parts (iii) and (iv) are obvious by just setting $\bm P^{-1}=\mathbf 0_{r\times r}$ in the first step of (ref).
For part (v) we have
by Assumption (ref)(a). Then, by (ref) in the proof of Lemma (ref), we have $\lim_{n\to\infty} n^{-1}\nu^{(1)}(\bm\Lambda_n^\prime\bm\Lambda_n) \le \overline C_1$, which, once substituted in (ref), proves part (v).
For part (vi) the proof is the same as part (v), but in (ref) we use Assumption (ref)(f) instead of Assumption (ref)(a), thus replacing $C_\xi^{-1}$ with $L_\xi$.
For part (vii)
by Lemma (ref) and Assumption (ref)(a). And for part (viii) the proof is the same as for part (vii) but using Assumption (ref)(f) instead of Assumption (ref)(a). This proves parts (vii) and (viii) and completes the proof. $\Box$
\noindentProof. We have
This completes the proof. $\Box$
\noindentProof. In Lemma (ref) set $\bm K=\bm C^\prime\bm B^{-1}\bm C$ and $\bm H=\bm A^{-1}$, then,
which implies
Then, by Weyl's inequality MK04:
From (ref), we have
For first term on the rhs of (ref), the $m$ eigenvalues of $\bm C'\bm B^{-1}\bm C$ are also the $m$ largest non-zero eigenvalues of $\bm C\bm C' \bm B^{-1}$, and the $m$ largest non-zero eigenvalues of $\bm C\bm C'$ are also the $m$ eigenvalues of $\bm C'\bm C$. Therefore, because of MK04: \[ n^{-1} \nu^{(m)}(\bm C\bm C^\prime \bm B^{-1} )\ge n^{-1} \nu^{(m)}(\bm C\bm C^\prime) \nu^{(m)}(\bm B^{-1})=n^{-1} \nu^{(m)}(\bm C^\prime\bm C)\left\{{\nu^{(1)}(\bm B)}\right\}^{-1}. \] Thus, by conditions (b) and (c) \[ n\left\{\nu^{(m)}(\bm C^\prime \bm B^{-1} \bm C)\right\}^{-1}\le n \left\{\nu^{(m)}(\bm C^\prime\bm C)\right\}^{-1} \nu^{(1)}(\bm B)\le \frac {M_B}{\underline M_{C_m}}. \] Moreover, by condition (a), $\nu^{(1)}(\bm A)>0$ and $\nu^{(1)}(\bm A)\le M_A$, i.e., $\{\nu^{(1)}(\bm A)\}^{-1}\ge M_A$, and \[ \nu^{(m)}(\bm C^\prime \bm B^{-1} \bm C)+\{\nu^{(1)}(\bm A)\}^{-1}\ge \nu^{(m)}(\bm C^\prime \bm B^{-1} \bm C)\ge \frac {M_B}{\underline M_{C_m}}. \]
For the second term on the rhs of (ref), by condition (a), \[ \left\{\nu^{(m)}(\bm A)\right\}^{-1} \leq \frac 1 L_A, \] for some finite positive real $L_A$. Hence, from (ref) \[ n \Vert (\bm A^{-1}+\bm C^\prime\bm B^{-1}\bm C)^{-1}\bm A^{-1}\Vert\le\frac {M_B}{\underline M_{C_m}L_A}. \]
and by using it in (ref) we complete the proof. $\Box$
Proof. Both results follow from Lemma (ref) since: $\bm P$ satisfies condition (a) by assumption, $\bm \Sigma_n^\xi$ satisfies condition (b) because of Assumption (ref)(a) and $\bm \Gamma_n^\xi$ satisfies condition (b) because of Lemma (ref)(v) and Assumption (ref)(f), and ${\bm\Lambda}_{n}^\prime{\bm\Lambda}_{n}$ satisfies condition (c) since it is positive definite because of Assumptions (ref)(a), and, moreover, its eigenvalues are such that $\underline C_{j}\le \lim\inf_{n\to\infty} n^{-1} \nu^{(j)}(\bm \Lambda_n^\prime\bm \Lambda_n) \le \lim\sup_{n\to\infty} n^{-1} \nu^{(j)}(\bm\Lambda_n^\prime\bm \Lambda_n) \le \overline C_j$, for $j=1,\ldots, r,$ as shown in (ref) in the proof of Lemma (ref).
Turning to part (iii),
by part (i) and Lemma (ref)(iii). This proves part (iii).
Part (iv) is proved in the same way but using Lemma (ref)(iv) instead of Lemma (ref)(iv). This completes the proof. $\Box$
Proof. Throughout, let $\lambda_{ij}$ be the $(i,j)$the entry of $\bm\Lambda_n$. For part (i), we have
where in the third step we used Assumption (ref)(a) (since $\max_{j=1,\ldots, r}\vert \lambda_{ij}\vert\le \Vert \bm\lambda_i\Vert\le M_\lambda$, for all $i=1,\ldots, n$), and Assumption (ref)(a), and in the last step we used Lemma (ref)(ii). By Chebychev's inequality and since the constants in (ref) do not depend on $t$, we prove part (i).
For part (ii), we have
by Lemma (ref)(vi). By Chebychev's inequality and since the constants in (ref) do not depend on $t$, we complete the proof of part (ii).
For part (iii), we have
by Assumptions (ref)(a), (ref)(a), and Lemma (ref)(i). By Chebychev's inequality we prove part (iii).
For part (iv), we have
by Assumption (ref)(a) and Lemma (ref)(ii). By Chebychev's inequality we prove part (iv).
Part (v) is proved as parts (iv), but using also Assumption (ref)(a).
For part (vi), we have
by Assumption (ref)(a). By Chebychev's inequality we prove part (vi). This completes the proof. $\Box$
Proof. Throughout, let $\lambda_{ij}$ be the $(i,j)$the entry of $\bm\Lambda_n$. For part (i),
by Lemma (ref)(i) and because $\vert \mathbb{E}[F_{kt}F_{ks}]\vert\le 1$ by Cauchy-Schwarz inequality and Assumption (ref)(b). By Chebychev's inequality we prove part (i).
For part (ii),
by Assumptions (ref)(a) and (ref)(d). By Chebychev's inequality we prove part (ii).
For part (iii), following the proof of part (ii),
by Assumptions (ref)(a) and (ref)(d). By Chebychev's inequality we prove part (iii).
Parts (iv) and (v) are proved as parts (i) and (ii), respectively, but using also Assumption (ref)(a).
For part (vi)
by Assumptions (ref)(a) and (ref)(d), Lemma (ref), and Cauchy-Schwarz inequality jointly with Assumption (ref)(b). By Chebychev's inequality we prove part (vi). This completes the proof. $\Box$
\noindentProof. By Assumption (ref)(b), for all $n\in\mathbb N$,
First notice that, the $r$ non-zero eigenvalues of $n^{-1}\bm\Gamma_n^\chi$ are the $r$ eigenvalues of $n^{-1}\bm\Lambda_n^\prime\bm\Lambda_n$ and for all $n>N_0$, $\Vert\bm\Lambda_n^\prime\bm\Lambda_n -\bm\Sigma_\Lambda\Vert =0$, by Assumption (ref)(a). Moreover, $\bm\Sigma_\Lambda$ is diagonal and positive definite by Assumptions (ref)(b) and (ref)(a), respectively. Hence, for all $n>N_0$,
Furthermore, it must be that the columns of $\bm\Lambda_n$ span the same space as the columns of $\mathbf V_n^\chi$. Since the eigenvectors are normalized and for all $n>N_0$, $\Vert n^{-1} \mathbf M_n^\chi -\bm\Sigma_\Lambda\Vert=0$ by (ref), there exist two $r\times r$ matrices $\mathbf K_{1n}$ and $\mathbf K_{2n}$ such that, for all $n>N_0$,
Let $$ \mathbf K_{1}= \lim_{n\to\infty} n^{-1/2} (\mathbf V_n^{\chi\prime}\mathbf V_n^{\chi})^{-1}\mathbf V_n^{\chi\prime}\bm\Lambda_n= \lim_{n\to\infty} n^{-1/2} \mathbf V_n^{\chi\prime}\bm\Lambda_n $$ and $$ \mathbf K_2= \lim_{n\to\infty} n(\bm\Lambda_n^\prime\bm\Lambda_n)^{-1} n^{-1/2} \bm\Lambda_n^\prime \mathbf V_n^\chi = \lim_{n\to\infty} \sqrt n(\bm\Lambda_n^\prime\bm\Lambda_n)^{-1} \bm\Lambda_n^\prime \mathbf V_n^\chi. $$ Then, by linear projection from (ref) we have that
which is positive definite since the columns of $\mathbf V_n^\chi$ are linear combinations of the columns of $\bm\Lambda_n$ so, for all $n>N_0$, $\text{rk} (n^{-1/2}\mathbf V_n^{\chi\prime}\bm\Lambda_n)=\text{rk} (n^{-1}\bm\Lambda_n^\prime \bm\Lambda_n)=\text{rk}(\bm\Sigma_\Lambda)=r$ by Assumption (ref)(a). Moreover, for all $n>N_0$, $\Vert \mathbf K_1\Vert\le n^{-1/2}\Vert \bm\Lambda_n\Vert$, which is finite by Lemma (ref).
Similarly, from (ref) we also have that
which exists and is positive definite, since $\Vert n(\bm\Lambda_n^\prime\bm\Lambda_n)^{-1}-n\bm\Sigma_\Lambda^{-1}\Vert = 0$, for all $n>N_0$, and $\bm\Sigma_\Lambda$ is finite and positive definite by Assumption (ref)(a). Moreover, for all $n>N_0$, $\Vert \mathbf K_2\Vert \le n\Vert(\bm\Lambda_n^\prime \bm\Lambda_n)^{-1}\Vert\,n^{-1/2}\Vert \bm\Lambda_n\Vert= \Vert \bm\Sigma_\Lambda^{-1}\Vert n^{-1/2}\Vert \bm\Lambda_n\Vert$, which is finite by Assumption (ref)(a) and Lemma (ref).
By using (ref) into the rhs of (ref), we get, for all $n>N_0$, $$ n^{-2} \bm\Lambda_n\mathbf K_{2n} \mathbf M_n^\chi \mathbf K_{2n}^\prime \bm\Lambda_n^\prime= \mathbf V_n^\chi\mathbf M_n^\chi \mathbf V_n^{\chi\prime}, $$ which, since eigenvectors are normalized, implies, that, for all $n>N_0$, we can write
From (ref) we must have $\mathbf I_r=\lim_{n\to\infty} n^{-1/2} \mathbf V_n^{\chi\prime}\bm\Lambda_n\mathbf K_{2n} = \mathbf K_1\mathbf K_2$, so, as expected $\mathbf K_1=\mathbf K_2^{-1}$ and $\mathbf K_2=\mathbf K_1^{-1}$. It follows that, for all $n>N_0$,
Now, from (ref), we can also write that for all $n>N_0$
for some $r\times r$ matrix $\bm R_n$. Let,
Hence, for all $n>N_0$, $\bm R = \mathbf K_{2} (\bm\Sigma_\Lambda)^{1/2}$ by (ref). So $\bm R$ is finite by Assumption (ref)(a) and since $\mathbf K_2$ is finite.
Moreover, from (ref)
From (ref) and (ref), and using Assumption (ref)(a), we also have
Hence, for all $n>N_0$, by (ref), $\bm R^{-1}=(\bm\Sigma_\Lambda)^{-1}\bm R^\prime \bm\Sigma_\Lambda$. Thus, $\bm R^{-1}$ is finite by Assumption (ref)(a) and since $\bm R_2$ is finite. It follows that $\bm R$ is positive definite.
Moreover, $\bm R$ is orthogonal, indeed, from (ref) and (ref)
By substituting (ref) into (ref), for all $n>N_0$, by (ref),
By right-multiplying (ref) by $\bm R$ and left-multiplying by $\bm\Sigma_\Lambda$ we have \[ \bm\Sigma_\Lambda=\bm R^{-1} \bm\Sigma_\Lambda \bm R, \] which implies $\bm R=\bm{\mathcal J}$ where $\bm{\mathcal J}$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$. Therefore, from (ref), for all $n>N_0$, \[ \bm\Lambda_n \bm{\mathcal J} = n^{-1/2} \mathbf V_n^\chi (\mathbf M_n^\chi)^{1/2} \] or, equivalently, for all $n>N_0$, \[ n^{-1/2}\bm\Lambda_n = n^{-1/2} \mathbf V_n^\chi \bm{\mathcal J} (\mathbf M_n^\chi)^{1/2}. \] Finally, by Assumption (ref)(c) it must be that $\bm{\mathcal J}=\bm{\mathcal S}$. This completes the proof. $\Box$
\noindentProof. We have,
by Assumption (ref), Assumptions (ref)(a) and (ref)(b), and Assumption (ref)(a). The proof of part (i) follows by Chebychev's inequality and by noticing that the bound in (ref) does not depend on $t$. This completes the proof. $\Box$
\noindentProof. It is enough to notice that $\mathbf F_t=\sum_{k=0}^\infty \mathbf A^k\mathbf H\mathbf u_{t-k}$, then, by Assumption (ref), we complete the proof. $\Box$
\noindentProof. Part (i) follows since $\{\mathbf F_t\}$ is ergodic, because of Assumption (ref)(d) which implies that $\{\mathbf F_t\}$ has summable autocovariances, and therefore $\{\mathbf F_t \mathbf F_{t-k}^\prime\}$ is also ergodic (white01, and stout1974almost). In particular, $\bm\Gamma_k^F$ is finite because of Assumptions (ref)(b), (ref)(d) and (ref)(e), and $\mathbb{E}[\Vert\mathbf F_t\Vert^4\Vert]$ is also finite because of Assumptions (ref)(d)-(ref)(g). See also Hamilton, which can be applied using the fact that $\{\mathbf v_t\}$ is an independent process by Assumption (ref)(f) and thus it is a martingale difference process. This proves part (i).
For part (ii), as $T\to\infty$,
because of Lemma (ref)(iii), and where we also used Lemma (ref), Cauchy-Schwarz inequality, and the fact that $F_{jt}$ is weakly stationary by Assumptions (ref)(b), (ref)(d), and (ref)(e). By noticing that the constants on the rhs of (ref) do not depend on $i$ we prove part (ii). Part (iii) follows directly from part (ii), indeed
For part (iv), notice that, since $\{\xi_{it}\}$ is a strongly mixing process with exponentially decaying coefficients, because of Assumption (ref)(c), then it is also ergodic (white01, and rosenblatt1972), so that also $\{\xi_{it}\xi_{jt}\}$ is ergodic (white01, and stout1974almost). In particular, $\mathbb{E}[\xi_{it}\xi_{jt}]$ is finite by Assumption (ref)(a) and $\mathbb{E}[\vert \xi_{it}\xi_{jt}\xi_{is}\xi_{js}\vert]$ is also finite because by Assumption (ref)(d), and both are bounded by constant independent of $i$ and $j$.
This proves part (iv). Part (v) follows directly from parts (i), (ii), and (iv). Part (vi) follows directly from part (iv), and part (vii) follows directly from part (v). This completes the proof. $\Box$
Proof. From Lemma (ref)(i), and MK04 which is Weyl's inequality
This implies (note that $x-y\ge -|x-y|$ for any $x,y\in\mathbb R$)
by Assumption (ref)(b). Thus, $\Vert (T^{-1}\sum_{t=1}^T\mathbf F_t\mathbf F_t^\prime)^{-1}\Vert = O_p(1)$. This completes the proof. $\Box$
Proof. Both results are direct consequences of MBPCAQML, see also Bai03 under similar assumptions. This completes the proof. $\Box$
Proof. For part (i), by definition
Using (ref) the first term on the rhs of (ref) is such that
by Lemma (ref)(ii), which does not depend on $t$, and Lemmas (ref) (jointly with Assumption (ref)(a)), (ref)(i), and (ref)(iii), and part (i). The case $k=1$ is proved in the same way and this proves part (iii).
For part (ii), as in part (i), we have
by Lemma (ref)(ii), which does not depend on $t$, and Lemmas (ref) (jointly with Assumption (ref)(a)), (ref)(iii), and (ref)(vi), and Lemma (ref)(ii). This proves part (ii).
Part (iii) follows directly from part (ii) but using Lemma (ref)(iii).
For part (iv), as in part (i), we have
by the same arguments used to prove part (ii) and noticing that now $\vert\xi_{it}\vert =O_p(1)$ by Assumption (ref)(a). Part (v) is proved as part (iv) setting $\xi_{it}=1$. This completes the proof. $\Box$
\noindentProof. For part (i), in agreement with Assumption (ref)(i), we can always set $\widetilde{\mathbf F}_0=\mathbf 0_r$, so it follows that, by construction $T^{-1}\sum_{t=2}^T \widetilde{\mathbf F}_{t-1}\widetilde{\mathbf F}_{t-1}^\prime=T^{-1}\sum_{t=1}^T \widetilde{\mathbf F}_{t-1}\widetilde{\mathbf F}_{t-1}^\prime=\mathbf I_r$. Then, by Assumption (ref)(b), we have
Now,
Now, by Lemma (ref)(i) the first and second term in the rhs of (ref) are $O_p(\max(n^{-1},T^{-1/2}))$, while the third term on the rhs is dominated by the first two. Hence, for the first term on the rhs of (ref) we have
The second term on the rhs of (ref) is $O_p(T^{-1/2})$, by Lemma (ref)(i). This proves part (i).
Part (ii) follows from the results in FGLR09 combined with part (i). This completes the proof. $\Box$
\noindentProof. Start from
since $\bm\Gamma^F=\mathbf I_r$ by Assumption (ref)(b), $\mathbb{E}[{\mathbf F}_t\xi_{it}]=\mathbf 0_{r}$ by Lemma (ref), and $T^{-1}\sum_{t=1}^T \widetilde{\mathbf F}_{t}\widetilde{\mathbf F}_{t}^\prime=\mathbf I_r$ by construction. For part (i), from (ref) we have
where, we used multiple times Assumption (ref)(a) and Lemma (ref)(i), and, we also used: Lemma (ref)(v) for the first term, Lemma (ref)(i) for the fourth term, Lemma (ref)(i) for the fifth term, Lemma (ref)(ii) for the sixth and seventh term, and Lemma (ref)(iv) for the last term. This proves part (i).
As for part (ii), from (ref) we have
The result in (ref) follows from repeated use of Lemmas (ref), (ref)(i), (ref)(ii), and (ref)(i) as well as the following results. First,
by Lemma (ref)(vii). Second,
by Lemma (ref)(i). Third,
by Lemma (ref)(iii). Fourth,
by Lemmas (ref)(iii) and (ref)(ii). And, last
by Lemmas (ref)(iii) and (ref)(ii). This completes the proof. $\Box$
\noindentProof. Start with
Consider each term on the rhs of (ref). The first term is
by Lemma (ref)(ii), and also Assumption (ref)(a) for which $\Vert({\bm\Sigma}_n^\xi)^{-1}\Vert\le C_\xi$, and Lemma (ref) for which $n^{-1/2}\Vert {\bm\Lambda}_n\Vert=O(1)$. The third term is
by Lemmas (ref) and (ref)(ii), Assumption (ref)(a), and since, by Lemma (ref)(i), for all $j=1,\ldots, n$,
By the same arguments, the fourth term is
Finally, the second term on the rhs of (ref) is
by Lemma (ref)(ii), (ref), and Assumptions (ref)(a) and (ref)(a). By substituting (ref), (ref), (ref), and (ref) into (ref), we prove part (i).
For part (ii), we have
by Lemmas (ref), (ref)(ii), (ref)(ii), and Assumptions (ref)(a) and (ref)(f). This proves part (ii).
For part (iii), by part (ii) and MK04 which is Weyl's inequality, we have
Moreover (note that $x-y\ge -|x-y|$ for any $x,y\in\mathbb R$),
thus, by Lemma (ref)(iv), which implies $\lim_{n\to\infty}n^{-1}\nu^{(r)} ({\bm\Lambda}_n^{\prime}({\bm \Sigma}_n^{\xi})^{-1}{\bm\Lambda}_n)>0$, and (ref), it follows that, with probability tending to one as $n,T\to\infty$, we have $\det (n^{-1} \widehat{\bm\Lambda}_n^{(0)\prime}(\widehat{\bm \Sigma}_n^{\xi(0)})^{-1}\widehat{\bm\Lambda}_n^{(0)})>0$, or, equivalently $n^{-1} \widehat{\bm\Lambda}_n^{(0)\prime}(\widehat{\bm \Sigma}_n^{\xi(0)})^{-1}\widehat{\bm\Lambda}_n^{(0)}$ is positive definite, i.e., $n\Vert(\widehat{\bm\Lambda}_n^{(0)\prime}(\widehat{\bm \Sigma}_n^{\xi(0)})^{-1}\widehat{\bm\Lambda}_n^{(0)})^{-1}\Vert=O_p(1)$. This proves part (iii).
For part (iv), we have
because of parts (i) and (iii) and Lemma (ref)(iii).
Part (v) follows directly from parts (ii) and (iv). This completes the proof. $\Box$
\noindentProof. Since $\mathbf P_{0|0}=\mathbf I_r$ is deterministic, then, for all $t=1,\ldots, T$, $\mathbf P_{t|t-1}$, $\mathbf P_{t|t}$, and $\mathbf P_{t|T}$ do not depend on the actual observations because of (ref) and (ref). This proves part (i).
As for part (ii), since $\mathbf F_{t|t-1}$ is based on less information than $\mathbf F_{t+1|t}$ and since ${\mathbf P}_{t|t-1}$ and $\mathbf P_{t+1|t}$ are deterministic, then $({\mathbf P}_{t|t-1} - {\mathbf P}_{t+1|t})$ is a positive definite matrix for all $t=1,\ldots, T-1$ (see, e.g., harvey90). As consequence, $\Vert \mathbf P_{t+1|t}\Vert\le \Vert \mathbf P_{t|t-1}\Vert$ for all $t=1,\ldots, T-1$ (see, e.g., MOA11).
The proof of parts (iii) and (iv) is identical to parts (i) and (ii), respectively, since also in this case $\mathbf P_{0,0|0}=\mathbf I_r$. This completes the proof. $\Box$
\noindentProof. Given that $\mathbf P_{0|0}=\mathbf I_r$ is obviously positive definite, by (ref) and Weyl's inequality MK04 it follows that
for some finite positive real $M_v^{-1}$, since $\bm\Gamma^v$ has full rank by Assumption (ref)(e) and $\nu^{(r)}(\mathbf A\mathbf A^\prime)$ is real and such that $\nu^{(r)}(\mathbf A\mathbf A^\prime)\ge (\nu^{(r)}(\mathbf A))^2\ge 0$. From $\mathbf P_{0|0}=\mathbf I_r$ and (ref) it follows also that
By Lemma (ref)(ii) and (ref), we have \[ \max_{t=2,\ldots, T}\Vert\mathbf P_{t|t-1}\Vert\le \Vert\mathbf P_{1|0}\Vert\le M_P, \] since $M_P$ is independent of $t$ and $\mathbf P_{t|t-1}$ is deterministic because of Lemma (ref)(i). This proves part (i).
For part (ii), by MK04 and (ref)
by the same arguments leading to (ref) and since $\nu_{\min}(\mathbf P_{t|t})\ge 0$ because $\mathbf P_{t|t}$ is at least positive semidefinite by construction. By letting $\underline M_P= M_v^{-1}$, we prove part (ii). Parts (iii) and (iv) are proved exactly as parts (i) and (ii), respectively, since also in this case $\mathbf P_{0,0|0}=\mathbf I_r$. This completes the proof. $\Box$
Proof. For part (i),
by Lemma (ref)(i) and since the second term on the rhs depends only on the estimation error of $\widehat{\mathbf A}^{(0)}$, $\widehat{\bm\Gamma}^{v(0)}$, $n^{-1/2}\widehat{\bm\Lambda}_n^{(0)}$, $n^{-1/2}\widehat{\bm\Lambda}_n^{(0)\prime}(\widehat{\bm\Sigma}_n^{\xi(0)})^{-1}$ and $n^{-1}(\widehat{\bm\Lambda}_n^{(0)\prime} (\widehat{\bm\Sigma}_n^{\xi(0)})^{-1} \widehat{\bm\Lambda}_n^{(0)})^{-1}$, which are all bounded by Lemmas (ref)(ii), (ref), (ref)(ii), and (ref)(iv). This proves part (i).
Part (ii) is proved in the same way as part (i) but using Lemma (ref)(ii). This completes the proof. $\Box$
Proof. Both results are proved by HS81. $\Box$
Proof. In Lemma (ref)(i) set $\bm A={\bm\Sigma}_n^{\xi}$, $\bm B=\mathbf I_r$, $\bm U={\bm\Lambda}_n\bm P$, and $\bm V={\bm\Lambda}_n^{\prime}$. Then, by noticing that the assumptions of Lemma (ref) are satisfied because of Assumptions (ref)(a) and (ref)(a), it follows that: \[ (\bm\Lambda_n\bm P \bm\Lambda_n^\prime+\bm\Sigma_n^\xi)^{-1} = ({\bm\Sigma}_n^{\xi})^{-1}-({\bm\Sigma}_n^{\xi})^{-1}\bm\Lambda_n\bm P(\mathbf I_r+ \bm\Lambda_n^\prime({\bm\Sigma}_n^{\xi})^{-1}\bm\Lambda_n\bm P)^{-1}\bm\Lambda_n^\prime({\bm\Sigma}_n^{\xi})^{-1}. \] Therefore,
Now, in Lemma (ref)(ii) set $\bm A=(\bm\Lambda_n^\prime({\bm\Sigma}_n^{\xi})^{-1}\bm\Lambda_n)^{-1}$, $\bm B=\bm P$, $\bm U=\mathbf I_r$, and $\bm V=\mathbf I_r $, and notice that the assumptions therein are satisfied because of Lemmas (ref)(iii) and (ref)(v). Then, for the last line of (ref) we have
By substituting (ref) into (ref) we complete the proof. $\Box$
Proof. From (ref) by using Lemma (ref), but with $\widehat{\bm\Lambda}_n^{(0)}$ and $\widehat{\bm\Sigma}_n^{\xi(0)}$ in place of $\bm\Lambda_n$ and $\bm\Sigma_n^\xi$, it holds that:
Then, by setting in Lemma (ref) $\bm K= {\mathbf P}_{t|t-1}^{(0)}$ and $\bm H=(\widehat{\bm\Lambda}_n^{(0)\prime}(\widehat{\bm\Sigma}_n^{\xi(0)})^{-1}\widehat{\bm\Lambda}_n^{(0)})^{-1}$, for the last line of (ref) we have
By substituting (ref) into (ref) we get
Finally, by using again (ref) into (ref)
Notice that we could use Lemmas (ref) and (ref) to derive (ref) since all inverses used are well defined because of Lemmas (ref)(iii), (ref)(i), and (ref)(ii) and (ref) in the proof of Lemma (ref).
Therefore, from (ref)
because of Lemmas (ref)(iii), (ref)(i), and (ref)(ii), and since, by MK04 which is Weyl's inequality,
again by Lemmas (ref)(iii), (ref)(i), and (ref)(ii). This completes the proof. $\Box$
Proof. From (ref), we get
Start with $t=T-1$, then from (ref),
by Lemmas (ref)(i), (ref)(ii), and (ref), and since $\Vert\widehat{\mathbf A}^{(0)}\Vert\le \Vert{\mathbf A}\Vert+\Vert\widehat{\mathbf A}^{(0)}-\mathbf A\Vert =O_p(1)$, by Assumption (ref)(d) and Lemma (ref)(i). From (ref) it follows that
Thus, at $t=T-2$, from (ref) and (ref),
From (ref) it follows that
Since all the bounds in (ref)-(ref) are the same for all $t$, from Lemma (ref) and (ref) we have
This completes the proof. $\Box$
\noindentProof. Recall the Woodbury forumla
Denote $\bm D=(\bm A^{-1}+\bm C^\prime\bm B^{-1}\bm C)^{-1}$ then from (ref) the lhs of (ref) is equivalent to
Then, (ref) becomes \[ \bm A\left[\bm I-\bm C^\prime\bm B^{-1}\bm C\bm D\right]\bm C^\prime\bm B^{-1}=\bm D\bm C^\prime\bm B^{-1}, \] or equivalently multiplying both sides on the right by $\bm B\bm C(\bm C^\prime\bm C)^{-1}$
Now multiplying (ref) on the left by $\bm A^{-1}$ and on the right by $\bm D^{-1}$ \[ \left[\bm D^{-1}-\bm C^\prime\bm B^{-1}\bm C\right]=\bm A^{-1}, \] which is equivalent to \[ \bm A^{-1}+\bm C^\prime\bm B^{-1}\bm C-\bm C^\prime\bm B^{-1}\bm C-\bm A^{-1}=\bm 0_{m\times m}, \] which is always true. $\Box$
Proof. Let $\widehat{\bm\Omega}^{F(0)}_s=\mathbb{E}_{\widehat{\varphi}^{(0)}}[\bm F_s\bm F_s^\prime]$, which is $rs\times rs$ having the $r\times r$ generic $(t_1,t_2)$ block denoted by $[\widehat{\bm\Omega}^{F(0)}_s]_{t_1,t_2}$ and such that $[\widehat{\bm\Omega}^{F(0)}_s]_{t_2,t_1}=[\widehat{\bm\Omega}^{F(0)}_s]_{t_1,t_2}^\prime$ and \[ \text{vec}\left([\widehat{\bm\Omega}^{F(0)}_s]_{t_1,t_2}\right) = (\mathbf I_r\otimes \widehat{\mathbf A}^{(0)})^{|t_1-t_2|} (\mathbf I_{r^2}-\{\widehat{\mathbf A}^{(0)}\otimes\widehat{\mathbf A}^{(0)}\})^{-1}\text{vec}(\widehat{\bm\Gamma}^{v(0)}), \quad t_1,t_2= 1,\ldots, s. \] Notice that although $\widehat{\bm\Omega}^{F(0)}_s$ depends on $\widehat{\mathbf A}^{(0)}$ and $\widehat{\bm\Gamma}^{v(0)}$, for simplicity of notation, hereafter, we omit such dependence. Let also ${\bm\Omega}^{F}_s=\mathbb{E}[\bm F_s\bm F_s^\prime]$, clearly ${\bm\Omega}^{F}_s$ is positive definite, since by Assumptions \ref{ass:com